Methods and systems for analyzing sequence reads
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2026-03-12
AI Technical Summary
Existing methods for determining nucleic acid sequences from sequence reads suffer from high error rates and difficulty in distinguishing nucleic acid molecules with identical start and stop positions, particularly in identifying distinct nucleic acid molecules.
A method involving obtaining and mapping sequence reads to a reference genomic sequence, grouping them into families based on identical start and stop positions, and further categorizing into subgroups to form consensus sequences for distinct nucleic acid molecules, reducing errors through nucleotide recurrence analysis.
This approach significantly reduces error rates to below 0.5 parts per million, enabling accurate differentiation between nucleic acid molecules and identifying somatic mutations with high precision.
Smart Images

Figure US2025040970_12032026_PF_FP_ABST
Abstract
Description
Attorney Docket No. 58626-723601METHODS AND SYSTEMS FOR ANALYZING SEQUENCE READS
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 752,383 filed January 31, 2025, and U.S. Provisional Patent Application No. 63 / 681,029 filed August 8, 2024, which are herein incorporated by reference in their entireties for all purposes.BACKGROUND OF THE INVENTION
[0002] Sequence reads may be used to determine the sequence of nucleic acid molecules from which the sequence reads are derived. A variety of methods for determining the sequence of nucleic acid molecules from sequence reads may have one or more drawbacks, such as (1) a relatively high error rate and / or (2) the difficultly or inability to identify distinct nucleic acid molecules with the same start and stop positions.SUMMARY OF THE INVENTION
[0003] Aspects disclosed herein provide methods for determining, from sequence reads, a sequence for a first nucleic acid molecule and a sequence for second nucleic acid molecule, wherein the first nucleic acid molecule and the second nucleic acid molecule have identical start and stop positions Such methods may comprise: (a) obtaining a plurality of sequence reads for each of a plurality of nucleic acid molecules from a biological sample of a subject; (b) mapping the plurality of sequence reads for each of the plurality of nucleic acid molecules to a reference genomic sequence; (c) grouping the plurality of sequence reads into groups of sequence reads that map to identical start and stop positions on the reference genomic sequence, thereby forming a plurality of families; (d) further grouping sequence reads of each family into a plurality of subgroups comprising a first subgroup and a second subgroup, each subgroup comprising a plurality of sequence reads; wherein: (i) the sequence reads of the first subgroup comprise a first nucleotide at a first position; and (ii) the sequence reads of the second subgroup comprise a second nucleotide at the first position, wherein the second nucleotide differs from the first nucleotide; (e) collapsing the sequence reads of the first subgroup to form a first consensus sequence corresponding with the sequence of a first nucleic acid molecule of the plurality of nucleic acid molecules; and (f) collapsing the sequence reads of the second subgroup to form a second consensus sequence corresponding with the sequence for a second nucleic acid molecule from the plurality of nucleic acid molecules, wherein the first nucleic acid molecule differs in sequence from the second nucleic acid molecule.
[0004] In some embodiments, the first subgroup comprises a plurality of non-identical sequence reads, wherein the non-identical sequence reads of the first subgroup have differing levels of recurrence of a nucleotide at a second position. In some embodiments, the first nucleotide at theAttorney Docket No. 58626-723601 first position is present in the plurality of sequence reads at a level of recurrence that is higher than the level of recurrence of the nucleotide at the second position. In some embodiments, the differing levels of recurrence of the nucleotide at the second position arise from one or more PCR errors. In some embodiments, the second subgroup comprises a plurality of non-identical sequence reads, wherein the plurality of non-identical sequence reads of the second subgroup have differing levels of recurrence of a nucleotide at a third position. In some embodiments, the first nucleotide at the first position is present in the plurality of sequence reads at a level of recurrence that is higher than the level of recurrence of the nucleotide at the third position. In some embodiments, the differing levels of recurrence of the nucleotide at the third position arise from one or more PCR errors. In some embodiments, the plurality of nucleic acid molecules from the biological sample of the subject are cell-free nucleic acid molecules. In some embodiments, the cell-free nucleic acid molecules comprise cell-free DNA molecules. In some embodiments, the reference genomic sequence is at least 100 kb in length. In some embodiments, the reference genomic sequence is at least 1 Mb in length. In some embodiments, the reference genomic sequence is a human exome. In some embodiments, the reference genomic sequence is from a human genome. In some embodiments, the reference genomic sequence is selected from the group consisting of the hgl9 human genome, the hgl8 genome, the hgl7 genome, the hgl6 genome, and the hg38 genome. In some embodiments, the sequence reads of the first subgroup comprising the first nucleotide at the first position comprise sequence reads originally derived from both the positive and negative strands of the reference genomic sequence. In some embodiments, sequence reads of the second subgroup comprising the second nucleotide at the first position comprise sequence reads originally derived from both the positive and negative strands of the reference genomic sequence. In some embodiments, the subject is a human subject. In some embodiments, the subject is suffering from a condition. In some embodiments, the condition is a cancer. In some embodiments, the biological sample comprises a blood sample, a serum sample, or a plasma sample. In some embodiments, the sequence reads were generated by next-generation sequencing. In some embodiments, the first nucleotide at the first position is a single-nucleotide variant. In some embodiments, the sequence reads of the first subgroup further comprise a second single-nucleotide variant. In some embodiments, the first nucleotide at the first position and the second single nucleotide variant are within 170 nucleotides of each other. In some embodiments, the first nucleotide at the first position and the second single nucleotide variant are separated by at least one nucleotide. In some embodiments, the method is implemented via a computer or computer system. In some embodiments, the reference genomic sequence is at least 10 kilobases in length. In some embodiments, theAttorney Docket No. 58626-723601 reference genomic sequence is at least 50 kilobases in length. In some embodiments, the reference genomic sequence is at least 100 kilobases in length. In some embodiments, the plurality of sequence reads comprises at least 10,000 sequence reads. In some embodiments, the plurality of sequence reads comprises at least 50,000 sequence reads. In some embodiments, the plurality of sequence reads comprises at least 100,000 sequence reads. In some embodiments, the plurality of nucleic acid molecules comprises at least 1,000 nucleic acid molecules. In some embodiments, the plurality of nucleic acid molecules comprises at least 5,000 nucleic acid molecules. In some embodiments, the plurality of nucleic acid molecules comprises at least 10,000 nucleic acid molecules. In some embodiments, (b) comprises performing a Burrows- Wheeler alignment (BWA) algorithm. In some embodiments, (b) comprises performing a Bowtie algorithm. In some embodiments, the subject previously had a condition. In some embodiments, the subject was previously treated for the condition. In some embodiments, the condition is a cancer. In some embodiments, the cancer is a lymphoma. In some embodiments, the subject was previously treated with a chemotherapeutic agent. In some embodiments, the subject was previously treated with a chimeric antigen receptor T cell therapy.
[0005] Another aspect disclosed herein provides a computer-implemented systems. Such computer-implemented systems may comprise a computing device comprising at least one processor, an operating system configured to perform executable instructions, a memory, and a computer program including instructions executable by the computing device for performing the method disclosed herein.
[0006] Another aspect disclosed herein provides a non-transitory computer-readable medium comprising machine-executable code that, upon execution by one or more computer processors, implements a method comprising: (a) obtaining a plurality of sequence reads for each of a plurality of nucleic acid molecules from a biological sample of a subject; (b) mapping the plurality of sequence reads for each of the plurality of nucleic acid molecules to a reference genomic sequence; (c) grouping the plurality of sequence reads into groups of sequence reads that map to identical start and stop positions on the reference genomic sequence, thereby forming a plurality of families; (d) further grouping sequence reads of each family into a plurality of subgroups comprising a first subgroup and a second subgroup, each subgroup comprising a plurality of sequence reads, wherein: (i) the sequence reads of the first subgroup comprise a first nucleotide at a first position; and (ii) the sequence reads of the second subgroup comprise a second nucleotide at the first position, wherein the second nucleotide differs from the first nucleotide; (e) collapsing the sequence reads of the first subgroup to form a first consensus sequence corresponding with the sequence of a first nucleic acid molecule of the plurality ofAttorney Docket No. 58626-723601 nucleic acid molecules; and (f) collapsing the sequence reads of the second subgroup to form a second consensus sequence corresponding with the sequence for a second nucleic acid molecule from the plurality of nucleic acid molecules, wherein the first nucleic acid molecule differs in sequence from the second nucleic acid molecule.
[0007] Another aspect disclosed herein provides a system comprising one or more computer processors and computer memory coupled thereto, wherein the computer memory comprises machine-executable code that, upon execution by the one or more computer processors, implements a method comprising: (a) obtaining a plurality of sequence reads for each of a plurality of nucleic acid molecules from a biological sample of a subject; (b) mapping the plurality of sequence reads for each of the plurality of nucleic acid molecules to a reference genomic sequence; (c) grouping the plurality of sequence reads into groups of sequence reads that map to identical start and stop positions on the reference genomic sequence, thereby forming a plurality of families; (d) further grouping sequence reads of each family into a plurality of subgroups comprising a first subgroup and a second subgroup, each subgroup comprising a plurality of sequence reads, wherein: (i) the sequence reads of the first subgroup comprise a first nucleotide at a first position; and (ii) the sequence reads of the second subgroup comprise a second nucleotide at the first position, wherein the second nucleotide differs from the first nucleotide; (e) collapsing the sequence reads of the first subgroup to form a first consensus sequence corresponding with the sequence of a first nucleic acid molecule of the plurality of nucleic acid molecules; and (f) collapsing the sequence reads of the second subgroup to form a second consensus sequence corresponding with the sequence for a second nucleic acid molecule from the plurality of nucleic acid molecules, wherein the first nucleic acid molecule differs in sequence from the second nucleic acid molecule.
[0008] Another aspect disclosed herein provides a method for identifying a somatic mutation, the method comprising: (a) obtaining a plurality of sequence reads for each of a plurality of nucleic acid molecules from a biological sample of a subject; (b) mapping the plurality of sequence reads for each of the plurality of nucleic acid molecules to a reference genomic sequence; (c) grouping the plurality of sequence reads into groups of sequence reads that map to identical start and stop positions on the reference genomic sequence, thereby forming a plurality of families; (d) further grouping sequence reads of each family into a plurality of subgroups comprising a first subgroup and a second subgroup, each subgroup comprising a plurality of sequence reads; wherein: (i) the sequence reads of the first subgroup comprise a first nucleotide at a first position; and (ii) the sequence reads of the second subgroup comprise a second nucleotide at the first position, wherein the second nucleotide differs from the first nucleotide;Attorney Docket No. 58626-723601(e) collapsing the sequence reads of the first subgroup to form a first consensus sequence corresponding with the sequence of a first nucleic acid molecule of the plurality of nucleic acid molecules; (f) collapsing the sequence reads of the second subgroup to form a second consensus sequence corresponding with the sequence for a second nucleic acid molecule from the plurality of nucleic acid molecules, wherein the first nucleic acid molecule differs in sequence from the second nucleic acid molecule, and (g) identifying the first nucleotide at the first position as a somatic mutation comprising a base-change selected from the group consisting of: an adenine-to- thymine (A>T) mutation, an adenine-to-cytosine (A>C) mutation, an adenine-to-guanine (A>G) mutation, a thymine-to-adenine (T>A) mutation, a thymine-to-cytosine (T>C) mutation, a thymine-to-guanine (T>G) mutation, a cytosine-to-adenine (C>A) mutation, a cytosine-to- thymine (C>T) mutation, a cytosine-to-guanine (OG) mutation, a guanine-to-adenine (G>A) mutation, a guanine-to-thymine (G>T) mutation, and a guanine-to-cytosine (G>C) mutation, wherein the identifying comprises an error rate of no more than 0.5 parts per million.
[0009] Another aspect disclosed herein provides a method for treating a condition in a subject, the method comprising: (a) obtaining a plurality of sequence reads for each of a plurality of nucleic acid molecules from a biological sample of the subject; (b) mapping the plurality of sequence reads for each of the plurality of nucleic acid molecules to a reference genomic sequence; (c) grouping the plurality of sequence reads into groups of sequence reads that map to identical start and stop positions on the reference genomic sequence, thereby forming a plurality of families; (d) further grouping sequence reads of each family into a plurality of subgroups comprising a first subgroup and a second subgroup, each subgroup comprising a plurality of sequence reads; wherein: (i) the sequence reads of the first subgroup comprise a first nucleotide at a first position; and (ii) the sequence reads of the second subgroup comprise a second nucleotide at the first position, wherein the second nucleotide differs from the first nucleotide;(e) collapsing the sequence reads of the first subgroup to form a first consensus sequence corresponding with the sequence of a first nucleic acid molecule of the plurality of nucleic acid molecules; (f) collapsing the sequence reads of the second subgroup to form a second consensus sequence corresponding with the sequence for a second nucleic acid molecule from the plurality of nucleic acid molecules, wherein the first nucleic acid molecule differs in sequence from the second nucleic acid molecule, and (g) identifying the first nucleic acid molecule as mutation containing based on the first consensus sequence; and (h) administering an effective amount of a therapeutic agent, based at least on part on (g).
[0010] In some embodiments, the condition is a cancer.Attorney Docket No. 58626-723601
[0011] Another aspect disclosed herein provides a non-transitory computer-readable storage media. Such a non-transitory computer- readable storage media may be encoded with a computer program including instructions executable by a processor for performing the methods disclosed herein.
[0012] Another aspect disclosed herein provides a method comprising: (a) sequencing or having sequenced a plurality of cell-free DNA molecules obtained or derived from a subject to produce a plurality of sequence reads; (b) aligning the plurality of sequence reads to a reference genomic sequence to produce a plurality of aligned sequence reads; (c) processing the plurality of aligned sequence reads to detect a presence of a somatic mutation in one or more cell-free DNA molecules from among the plurality of cell-free DNA molecules, wherein the somatic mutation comprises a base-change mutation selected from the group consisting of: an adenine-to-thymine (A>T) mutation, an adenine-to-cytosine (A>C) mutation, an adenine-to-guanine (A>G) mutation, a thymine-to-adenine (T>A) mutation, a thymine-to-cytosine (T>C) mutation, a thymine-to-guanine (T>G) mutation, a cytosine-to-adenine (C>A) mutation, a cytosine-to- thymine (C>T) mutation, a cytosine-to-guanine (OG) mutation, a guanine-to-adenine (G>A) mutation, a guanine-to-thymine (G>T) mutation, and a guanine-to-cytosine (G>C) mutation, wherein the processing in (c) comprises filtering the plurality of aligned sequence reads based at least in part on an expected error rate of detection of the base-change mutation in individual aligned sequence reads, thereby generating a set of filtered sequence reads; and (d) processing the set of filtered sequence reads to determine whether at least one of the set of filtered sequence reads is mutation-containing, wherein the determining in (d) has an error rate of no more than 0.5 parts per million.
[0013] In some embodiments, the somatic mutation comprises a single nucleotide variant (SNV). In some embodiments, the method further comprises obtaining a biological sample from the subject and extracting the plurality of cell-free DNA molecules from the biological sample. In some embodiments, the biological sample comprises blood, serum, plasma, tumor cells, saliva, urine, cerebrospinal fluid, lymphatic fluid, prostatic fluid, seminal fluid, milk, sputum, stool, tears, vaginal secretion, semen sample, bone marrow, or a combination thereof, or derivatives thereof. In some embodiments, the sequencing in (a) further comprising amplifying the plurality of cell-free DNA molecules to produce amplified DNA molecules and sequencing the amplified DNA molecules. In some embodiments, the amplifying comprises polymerase chain reaction. In some embodiments, the sequencing in (a) further comprises targeted sequencing. In some embodiments, the targeted sequencing further comprises performing probe enrichment of the plurality of cell-free DNA molecules. In some embodiments, the sequencing in (a) furtherAttorney Docket No. 58626-723601 comprises whole genome sequencing or whole exome sequencing. In some embodiments, the aligning in (b) further comprises aligning at least 1 thousand sequence reads, at least 10 thousand sequence reads, at least 100 thousand sequence reads, at least 1 million sequence reads, at least 10 million sequence reads, or at least 100 million sequence reads. In some embodiments, the aligning in (b) further comprises aligning sequence reads representative of at least 1 thousand cell-free DNA molecules, at least 10 thousand cell-free DNA molecules, at least 100 thousand cell-free DNA molecules, at least 1 million cell-free DNA molecules, at least 10 million cell-free DNA molecules, or at least 100 million cell-free DNA molecules. In some embodiments, the reference genomic sequence is a human genome assembly. In some embodiments, the human genome assembly is an HG19 human genome assembly or an HG38 human genome assembly. In some embodiments, the reference genomic sequence may have a length of at least 10 kilobases. In some embodiments, the reference genomic sequence may have a length of at least 50 kilobases. In some embodiments, the reference genomic sequence may have a length of at least 100 kilobases. In some embodiments, the aligning in (b) further comprises performing a Burrows-Wheeler alignment (BWA) algorithm or a Bowtie alignment algorithm. In some embodiments, the aligning in (b) may comprise use of a start position of the sequence reads. In some embodiments, the aligning in (b) may comprise use of a stop position of the sequence reads. In some embodiments, the sequence reads may be grouped based on a start position and a stop position. In some embodiments, the sequence reads may be grouped without use of molecular barcodes. In some embodiments, the method further comprises in (c) pre-processing the plurality of aligned sequence reads to filter out at least a portion of the plurality of aligned sequence reads. In some embodiments, the pre-processing comprises demultiplexing FASTQ files, extracting unique molecular identifiers (UMIDs), performing molecular barcode-mediated error suppression (e.g., integrated digital error suppression), performing background polishing, or deduplicating barcodes. In some embodiments, the detecting in (c) further comprises processing the plurality of aligned sequence reads using a somatic mutation calling algorithm. In some embodiments, the somatic mutation calling algorithm comprises MuTect2, VarScan2, or Strelka2. In some embodiments, the filtering in (c) further comprises comparing the expected error rate of detection of the base-change mutation to a pre-determined error rate threshold. In some embodiments, the filtering in (c) further comprises processing the base-change mutation using a trained algorithm. In some embodiments, the trained algorithm comprises a machine learning model or a statistical model. In some embodiments, the machine learning model comprises a natural language processing model, an artificial neural network, a decision tree, a random forest, a naive bayes classifier, a boosting algorithm, a k-nearest neighbor algorithm, or aAttorney Docket No. 58626-723601 clustering model. In some embodiments, the statistical model comprises a regression model, a classification model, a Monte Carlos simulation, or a polynomial model. In some embodiments, the determining in (d) is further based at least in part on a number of copies of the one or more cell-free DNA molecules having the somatic mutation. In some embodiments, the number of copies is compared to a threshold based at least in part on the expected error rate of detection of the base-change mutation. In some embodiments, the determining in (d) is further based at least in part on a number of independent copies of both Watson and Crick complementary strands present in the one or more cell-free DNA molecules having the somatic mutation. In some embodiments, the number of independent copies is compared to a threshold based at least in part on the expected error rate of detection of the base-change mutation. In some embodiments, the determining in (d) is further based at least in part on a strand location of the base-change mutation in the one or more cell-free DNA molecules. In some embodiments, the strand location is at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, or at least 10 nucleotides away from a 5’ end or a 3’ end of the one or more cell-free DNA molecules. In some embodiments, the somatic mutation comprises the A>T mutation. In some embodiments, the somatic mutation comprises the A>C mutation. In some embodiments, the somatic mutation comprises the A>G mutation. In some embodiments, the somatic mutation comprises the T>A mutation. In some embodiments, the somatic mutation comprises the T>C mutation. In some embodiments, the somatic mutation comprises the T>G mutation. In some embodiments, the somatic mutation comprises the OA mutation. In some embodiments, the somatic mutation comprises the OT mutation. In some embodiments, the somatic mutation comprises the OG mutation. In some embodiments, the somatic mutation comprises the G>A mutation. In some embodiments, the somatic mutation comprises the G>T mutation. In some embodiments, the somatic mutation comprises the G>C mutation. In some embodiments, the filtering in (c) further comprises selecting a subset of the plurality of aligned sequence reads, and analyzing the subset of the plurality of aligned sequence reads. In some embodiments, the subset is selected from among the plurality of aligned sequence reads based at least in part on having a lowest expected error rate of detection of the base-change mutation. In some embodiments, the subject has been diagnosed with cancer, and wherein the method further comprises determining the Minimal Residual Disease (MRD) status in the subject, based at least in part on the determining in (d). In some embodiments, the subject has been administered a treatment for cancer, and wherein the method further comprises assessing a therapeutic response of the subject in response to the treatment for cancer, based at least in part on the determining in (d). In some embodiments, the subject has been administered a treatment for cancer, and whereinAttorney Docket No. 58626-723601 the method further comprises assessing a risk of cancer recurrence or relapse in the subject, based at least in part on the determining in (d). In some embodiments, the subject has been administered a first treatment for cancer, and wherein the method further comprises treating the subject with a second treatment for cancer, based at least in part on the determining in (d). In some embodiments, the method further comprises administering, to the subject, an effective amount of a treatment for cancer, based at least in part on the determining in (d). In some embodiments, the method further comprises manufacturing a medicament for treating cancer in the subject, based at least in part on the determining in (d). In some embodiments, the cancer is a blood cancer. In some embodiments, the cancer is a solid tumor. In some embodiments, the cancer is a B cell malignancy. In some embodiments, the cancer is leukemia or lymphoma. In some embodiments, the cancer is large B-cell lymphoma (LBCL) or a Diffuse Large B-Cell Lymphoma (DLBCL). In some embodiments, the cancer is selected from acute myeloid (or myelogenous) leukemia (AML), chronic myeloid (or myelogenous) leukemia (CML), acute lymphocytic (or lymphoblastic) leukemia (ALL), chronic lymphocytic leukemia (CLL), hairy cell leukemia (HCL), small lymphocytic lymphoma (SLL), Mantle cell lymphoma (MCL), Marginal zone lymphoma, Burkitt lymphoma, Hodgkin lymphoma (HL), non-Hodgkin lymphoma (NHL), Anaplastic large cell lymphoma (ALCL), follicular lymphoma, refractory follicular lymphoma, diffuse large B-cell lymphoma (DLBCL) and multiple myeloma (MM), adult ALL. In some embodiments, the cancer is a bladder cancer, colorectal cancer, breast cancer, prostate cancer, renal cancer, hepatocellular cancer, lung cancer, ovarian cancer, cervical cancer, pancreatic cancer, rectal cancer, thyroid cancer, uterine cancer, gastric cancer, esophageal cancer, head and neck cancer, melanoma, neuroendocrine cancers, CNS cancers, brain tumors, bone cancer, or soft tissue sarcoma. In some embodiments, the treatment or the medicament is a first-line therapy. In some embodiments, the treatment or the medicament is a second-line therapy. In some embodiments, the treatment comprises radiographic imaging of the subject. In some embodiments, the treatment comprises computed tomography (CT) imaging, positron emission tomography (PET) imaging, and / or magnetic resonance imaging (MRI) of the subject. In some embodiments, the treatment or medicament comprises a cell therapy. In some embodiments, the treatment comprises a CAR-T cell therapy. In some embodiments, the CAR-T cell therapy is an anti-CD19 cell therapy. In some embodiments, the cell therapy comprises genetically engineered cells. In some embodiments, the genetically engineered cells are T cells. In some embodiments, the genetically engineered T cells comprise a Chimeric Antigen Receptor (CAR). In some embodiments, the CAR specifically binds to an antigen associated with a disease or condition and / or is expressed by cells associated with a disease or condition. In someAttorney Docket No. 58626-723601 embodiments, the CAR specifically binds to two antigens associated with a disease or condition and / or is expressed by cells associated with a disease or condition. In some embodiments, the antigen is selected from the group consisting of 5T4, 8H9, avb6 integrin, B7-H6, B cell maturation antigen (BCMA), CA9, a cancer-testes antigen, carbonic anhydrase 9 (CAIX), CCL- 1, CD19, CD20, CD22, CEA, hepatitis B surface antigen, CD23, CD24, CD30, CD33, CD38, CD44, CD44v6, CD44v7 / 8, CD 123, CD 138, CD171, carcinoembryonic antigen (CEA), CE7, a cyclin, cyclin A2, c-Met, dual antigen, EGFR, epithelial glycoprotein 2 (EPG-2), epithelial glycoprotein 40 (EPG-40), EPHa2, ephrinB2, erb-B2, erb-B3, erb-B4, erbB dimers, EGFR vIII, estrogen receptor, Fetal AchR, folate receptor alpha, folate binding protein (FBP), FCRL5, FCRH5, fetal acetylcholine receptor, G250 / CAIX, GD2, GD3, gplOO, Her2 / neu (receptor tyrosine kinase erbB2), HMW-MAA, IL-22R-alpha, IL-13 receptor alpha 2 (IL-13Ra2), kinase insert domain receptor (kdr), kappa light chain, Lewis Y, LI -cell adhesion molecule (LI -CAM), Melanoma-associated antigen (MAGE)-Al, MAGE-A3, MAGE-A6, MART-1, mesothelin, murine CMV, mucin 1 (MUC1), MUC16, NCAM, NKG2D, NKG2D ligands, NY-ESO-1, O- acetylated GD2 (OGD2), oncofetal antigen, Preferentially expressed antigen of melanoma (PRAME), PSCA, progesterone receptor, survivin, ROR1, TAG72, tEGFR, VEGF receptors, BAFF-R, VEGF-R2, Wilms Tumor 1 (WT-1), and a pathogen-specific antigen. In some embodiments, the antigen is CD 19. In some embodiments, the antigen is CD20. In some embodiments, the CAR comprises an extracellular antigen-recognition domain that specifically binds to the antigen and an intracellular signaling domain comprising an IT AM. In some embodiments, the intracellular signaling domain comprises an intracellular domain of a CD3- zeta (CD3 chain. In some embodiments, the CAR further comprises a costimulatory signaling region. In some embodiments, the costimulatory signaling region comprises a signaling domain of CD28 or 4-1BB. In some embodiments, the signaling domain is a domain of 4-1BB. In some embodiments, the T cells are CD4+. In some embodiments, the T cells are CD4+ or CD8+. In some embodiments, the T cells are primary T cells obtained from a subject. In some embodiments, the genetically engineered cells are autologous to the subject. In some embodiments, the genetically engineered cells are allogeneic to the subject. In some embodiments, the subject is a human. In some embodiments, the subject has Stage Eli disease. In some embodiments, the subject has Stage III / IV disease. In some embodiments, the subject has Minimal Residual Disease (MRD). In some embodiments, the subject is refractory to treatment with one or more prior therapies for the cancer. In some embodiments, the subject achieved an insufficient response to one or more prior therapies for the cancer. In some embodiments, the subject achieves a durable response to the treatment or the medicament. In some embodiments,Attorney Docket No. 58626-723601 the durable response is defined as an absence of relapse or remission of the cancer for up to 3 months 6 months, or 12 months. In some embodiments, the subject has a higher rate of survival. In some embodiments, the subject’s disease baseline characteristics are determined. In some embodiments, the baseline characteristics comprise international prognosis index (IPI) score, serum lactate dehydrogenase (LDH), or sum of the product of diameters (SPD), disease stage, or any combination of the above. In some embodiments, the baseline characteristics are defined by Lugano 2014 criteria. In some embodiments, the subject is classified under an international prognosis index (IPI) score. In some embodiments, the subject is classified as low, low- intermediate, or high-intermediate risk on the IPI score. In some embodiments, the subject has no risk factor or one risk factor and is considered to be in an IPI low risk group. In some embodiments, the subject has two risk factors and is considered to be in an IPI low-intermediate risk group. In some embodiments, the subject has three risk factors and is considered to be in an IPI high-intermediate risk group. In some embodiments, the subject has a higher or lower IPI score. In some embodiments, a volumetric measure of tumor burden of the subject is measured. In some embodiments, the volumetric measure of tumor burden of the subject is a sum of products of diameter (SPD). In some embodiments, the volumetric measure of tumor burden is measured using computed tomography (CT), positron emission tomography (PET), and / or magnetic resonance imaging (MRI) of the subject. In some embodiments, an SPD threshold value is or is about 30 per cm2, is or is about 40 per cm2, is or is about 50 per cm2, is or is about 60 per cm2, or is or is about 70 per cm2. In some embodiments, a level of an inflammatory marker is measured in the subject. In some embodiments, the level of the inflammatory marker is or is about 300 units per liter, is or is about 400 units per liter, is or is about 500 units per liter or is or is about 600 units per liter. In some embodiments, the inflammatory marker is lactate dehydrogenase (LDH). In some embodiments, the determining in (d) has an error rate of no more than 0.2 parts per million.
[0014] Another aspect disclosed herein provides a system comprising one or more computer processors and computer memory coupled thereto, wherein the computer memory comprises machine-executable code that, upon execution by the one or more computer processors, implements a method comprising: (a) sequencing or having sequenced a plurality of cell-free DNA molecules obtained or derived from a subject to produce a plurality of sequence reads; (b) aligning the plurality of sequence reads to a reference genomic sequence, to produce a plurality of aligned sequence reads; (c) processing the plurality of aligned sequence reads, to detect a presence of a somatic mutation in one or more cell-free DNA molecules from among the plurality of cell-free DNA molecules, wherein the somatic mutation comprises a base-changeAttorney Docket No. 58626-723601 mutation selected from the group consisting of: an adenine-to-thymine (A>T) mutation, an adenine-to-cytosine (A>C) mutation, an adenine-to-guanine (A>G) mutation, a thymine-to- adenine (T>A) mutation, a thymine-to-cytosine (T>C) mutation, a thymine-to-guanine (T>G) mutation, a cytosine-to-adenine (C>A) mutation, a cytosine-to-thymine (OT) mutation, a cytosine-to-guanine (OG) mutation, a guanine-to-adenine (G>A) mutation, a guanine-to- thymine (G>T) mutation, and a guanine-to-cytosine (G>C) mutation, wherein the processing in (c) comprises filtering the plurality of aligned sequence reads based at least in part on an expected error rate of detection of the base-change mutation in individual aligned sequence reads, thereby generating a set of filtered sequence reads; and (d) processing the set of filtered sequence reads to determine whether at least one of the set of filtered sequence reads is mutationcontaining, wherein the determining in (d) has an error rate of no more than 0.5 parts per million.
[0015] Another aspect disclosed herein provides a non-transitory computer-readable medium comprising machine-executable code that, upon execution by one or more computer processors, implements a method comprising: (a) sequencing a plurality of cell-free DNA molecules obtained or derived from a subject to produce a plurality of sequence reads; (b) aligning the plurality of sequence reads to a reference genomic sequence, to produce a plurality of aligned sequence reads; (c) processing the plurality of aligned sequence reads, to detect a presence of a somatic mutation in one or more cell-free DNA molecules from among the plurality of cell-free DNA molecules, wherein the somatic mutation comprises a base-change mutation selected from the group consisting of: an adenine-to-thymine (A>T) mutation, an adenine-to-cytosine (A>C) mutation, an adenine-to-guanine (A>G) mutation, a thymine-to-adenine (T>A) mutation, a thymine-to-cytosine (T>C) mutation, a thymine-to-guanine (T>G) mutation, a cytosine-to- adenine (OA) mutation, a cytosine-to-thymine (C>T) mutation, a cytosine-to-guanine (OG) mutation, a guanine-to-adenine (G>A) mutation, a guanine-to-thymine (G>T) mutation, and a guanine-to-cytosine (G>C) mutation, wherein the processing in (c) comprises filtering the plurality of aligned sequence reads based at least in part on an expected error rate of detection of the base-change mutation in individual aligned sequence reads, thereby generating a set of filtered sequence reads; and (d) processing the set of filtered sequence reads to determine whether at least one of the set of filtered sequence reads is mutation-containing, wherein the determining in (d) has an error rate of no more than 0.5 parts per million.
[0016] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the presentAttorney Docket No. 58626-723601 disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure.Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE
[0017] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. If a definition set forth herein is contrary to or otherwise inconsistent with a definition set forth in the patents, applications, published applications and other publications that are herein incorporated by reference, the definition set forth herein prevails over the definition that is incorporated herein by reference.BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Various features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which:
[0019] FIG. 1 shows a process for obtaining, mapping, grouping, and collapsing sequence reads to generate consensus sequences.
[0020] FIG. 2A-2D shows a schematic for grouping sequence reads based on features of the sequence reads.
[0021] FIG. 3 shows a computer system that is programmed or otherwise configured to implement methods provided herein.
[0022] FIG. 4 shows a plot of background error in control samples for three samples.
[0023] FIG. 5 shows a plot of the observed fraction relative to the expected fraction for three samples.
[0024] FIG. 6 shows a plot of background error in control samples for three samples.
[0025] FIG. 7 shows a plot of the observed fraction relative to the expected fraction for three samples.
[0026] FIG. 8 shows example data of the error rates of different base changes.
[0027] FIG. 9 shows example data of the error rates of different base changes with different number of copies.
[0028] FIG. 10 shows a schematic for identifying sequences using molecular barcodes.Attorney Docket No. 58626-723601
[0029] FIG. 11 shows a schematic for identifying sequences without using molecular barcodes.
[0030] FIG. 12A-12B shows schematics for identifying sequences without using molecular barcodes.
[0031] FIG. 13A shows a plot of expected value vs. SNV tumor fraction for different analysis methods of sequence reads.
[0032] FIG. 13B shows a plot of expected value vs. PV tumor fraction for different analysis methods of sequence reads.
[0033] FIG. 14A shows a plot of expected dilution vs. estimated actual fraction of tumor DNA based on SNV detection using two analysis conditions.
[0034] FIG. 14B shows a plot of expected dilution vs. estimated actual fraction of tumor DNA based on PV detection using two analysis conditions.
[0035] FIG. 15 shows example data of the error rates of different base changes.
[0036] FIG. 16 shows example data of the error rates of different base changes with different number of copies.DETAILED DESCRIPTION OF THE INVENTION
[0037] As used herein, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. For example, “a” or “an” means “at least one” or “one or more.”
[0038] Throughout this disclosure, various aspects of the claimed subject matter are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the claimed subject matter. Accordingly, the description of a range should be considered to have specifically disclosed all the possible sub-ranges as well as individual numerical values within that range. For example, where a range of values is provided, it is understood that each intervening value, between the upper and lower limit of that range and any other stated or intervening value in that stated range is encompassed within the claimed subject matter. The upper and lower limits of these smaller ranges can independently be included in the smaller ranges, and are also encompassed within the claimed subject matter, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the claimed subject matter. This applies regardless of the breadth of the range.
[0039] As used herein, a “subject” generally refers to a human or an animal, such as a mammal. In some embodiments, the subject, e.g., patient, to whom the agent or agents, cells, cell populations, or compositions are administered, is a mammal, typically a primate, such as aAttorney Docket No. 58626-723601 human. In some embodiments, the primate is a monkey or an ape. The subject can be male or female and can be any suitable age, including infant, juvenile, adolescent, adult, and geriatric subjects. In some embodiments, the subject is a non-primate mammal, such as a rodent.
[0040] As used herein, “treatment” (and grammatical variations thereof such as “treat” or “treating”) generally refers to complete or partial amelioration or reduction of a disease or condition or disorder, or a symptom, adverse effect or outcome, or phenotype associated therewith. Desirable effects of treatment include, but are not limited to, preventing occurrence or recurrence of disease, alleviation of symptoms, diminishment of any direct or indirect pathological consequences of the disease, preventing metastasis, decreasing the rate of disease progression, amelioration or palliation of the disease state, and remission or improved prognosis. The terms do not imply complete curing of a disease or complete elimination of any symptom or effect(s) on all symptoms or outcomes.
[0041] As used herein, an “effective amount” of an agent, e.g., a pharmaceutical formulation, cells, or composition, in the context of administration, generally refers to an amount effective, at dosages / amounts and for periods of time necessary, to achieve a desired result, such as a therapeutic or prophylactic result.
[0042] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.
[0043] The present disclosure provides methods and systems for determining sequences of two or more distinct nucleic acid molecules from a plurality of sequence reads. The sequences may represent the sequences of two or more nucleic acid molecules in a biological sample. The methods described herein may be used for grouping sequence reads based on certain features of the sequence reads and collapsing the sequence reads to form one or more consensus sequences. The systems described herein may be computer-implemented systems for performing the methods related to determining one or more sequences.
[0044] Methods and systems disclosed herein may reduce an error rate (e.g. a background error rate) often involved during detection and / or analysis of cell-free nucleic acid molecules (e.g., cell-free DNA molecules (cfDNA) or cell-free RNA molecules (cfRNA)). In some aspects, methods and systems for cell-free nucleic acid sequencing and detection of diseases (e.g., cancer) are provided. In some embodiments, cell-free nucleic acids (e.g., cfDNA or cfRNA) can be obtained from a subject and prepared for sequencing. Sequencing results of the cell-free nucleic acids can be analyzed to detect somatic mutations (e.g., single nucleotide variants or phased variants) as an indication of circulating-tumor nucleic acid (ctDNA or ctRNA) sequences (i.e., sequences that derived or are originated from nucleic acids of a cancer cell). Accordingly, inAttorney Docket No. 58626-723601 some cases, cancer can be detected in the subject by extracting a liquid biopsy from the subject and sequencing the cell-free nucleic acids derived from that liquid biopsy to detect circulatingtumor nucleic acid sequences, and the presence of circulating-tumor nucleic acid sequences can indicate that the individual has the disease (e.g., cancer). In some cases, a treatment or a medicament can be determined and / or performed on the subject based on the detection of the disease.
[0045] The methods and systems disclosed herein may be useful for differentiating between variants of nucleic acid molecules and errors introduced during amplification (e.g. PCR amplification). In some cases, amplifying nucleic acid molecules (e.g. cell-free DNA molecules) may introduce errors into sequences of the amplified nucleic acid molecules. The errors due to amplification of nucleic acid molecules may be difficult to differentiate from variants (e.g. single nucleotide variants (SNVs)) of the original nucleic acid molecules that are amplified. The methods and systems disclosed herein may enable identification of variants of the original nucleic acid molecules that are amplified without including errors resulting from amplification.MethodsA) Methods for determining molecular origin in molecules sharing start and stop positions
[0046] In some aspects provided herein are methods for determining, from sequence reads, a sequence for a first nucleic acid molecule and a sequence of a second nucleic acid molecule. The first nucleic acid molecule and the second nucleic acid molecule may have identical start positions, identical stop positions, or both identical start and stop positions. The methods described herein may comprise obtaining a plurality of sequence reads for each of a plurality of nucleic acid molecules from a biological sample of a subject. The plurality of sequence reads for each of the plurality of nucleic acid molecules may be mapped to a reference genomic sequence. The plurality of sequence reads may be grouped into groups of sequence reads that map to identical start and stop positions on the reference genomic sequence. Grouping the plurality of sequence reads into groups may thereby form a plurality of families. The sequence reads of each family may be grouped into a plurality of subgroups. The plurality of subgroups may comprise a first subgroup and a second subgroup. Additional subgroups may also be generated. The sequence reads of the first subgroup may comprise a nucleotide (e.g., a first nucleotide) at a position (e.g., a first position). The sequence reads of the second subgroup may comprise a nucleotide (e.g., a second nucleotide) at the same position (e.g., the first position). The second nucleotide may differ from the first nucleotide. The methods may comprise collapsing the sequence reads of the first subgroup to form a consensus sequence (e.g., a first consensusAttorney Docket No. 58626-723601 sequence) corresponding with the sequence of a nucleic acid molecule (e.g., a first nucleic acid molecule) of the plurality of nucleic acid molecules. The methods may comprise collapsing the sequence reads of the second subgroup to form a consensus sequence (e.g., a second consensus sequence) corresponding with the sequence for a nucleic acid molecule (e.g., a second nucleic acid molecule) from the plurality of nucleic acid molecules. The first nucleic acid molecule may differ in sequence from the second nucleic acid molecule.
[0047] FIG. 1 schematically illustrates an example of determining a first consensus sequence and a second consensus sequence from sequence reads. In this example, a plurality of sequence reads are obtained (101). The plurality of sequence reads may be mapped to a reference genomic sequence (102). In some cases, a start position and / or stop position may be identified for the plurality of sequence reads by mapping the plurality of sequence reads to the reference genomic sequence. The plurality of sequence reads may be grouped into groups of sequence reads that map to identical start and / or stop positions on the reference genomic sequence to generate a plurality of families (103). Each family may be grouped into subgroups comprising a first subgroup and a second subgroup (104). The sequence reads of the first subgroup may be collapsed to form a first consensus sequence (105). The sequence reads of the second subgroup may be collapsed to form a second consensus sequence (106).
[0048] In some aspects, the methods described herein may not comprise use of molecular barcodes. The methods of the present application may not comprise using nucleic acid molecules (e.g. cell-free DNA molecules) with molecular barcodes to provide information regarding a source (e.g. origin) of the nucleic acid molecules. The methods described herein, may comprise use of methods other than use of molecular barcodes to determine a source of a nucleic acid molecule. In some cases, information about the nucleic acid molecule (independent of any molecular barcode) may be used to determine a source (e.g., a molecular origin) of the nucleic acid molecule. For example, a start position of a sequence of a nucleic acid molecule, stop position of a sequence of a nucleic acid molecule, length of a sequence of a nucleic acid molecule, a presence and / or an absence of single nucleotide variants (SNVs) and / or phased variants (PVs) within a sequence of a nucleic acid molecule, or a combination thereof may be used to determine a source of the nucleic acid molecule. The terms “phased variants” or “PVs,” as used interchangeably herein, generally refers to two or more mutations (e.g., SNVs or indels) that occur in cis (i.e., on the same strand of a nucleic acid molecule) within a single cell-free nucleic acid molecule. In some cases, a cell-free nucleic acid molecule can be a cell-free deoxyribonucleic acid (cfDNA) molecule. In some cases, a source of a nucleic acid molecule after amplification may be determined based on a start position of a sequence read of the nucleicAttorney Docket No. 58626-723601 acid molecule after amplification and a stop position of the sequence read of the nucleic acid molecule after amplification. This information about the start position and stop position may be used to determine the source of the molecule from which the nucleic acid molecule was amplified from. In some cases, a presence of a SNV may be used in combination with start position and stop position information to determine a molecular origin of a nucleic acid molecule.
[0049] An example of methods described herein that do not comprise using molecular barcodes (e.g. unique molecule indexes (UMIs)) is shown in FIG. 12B. FIG. 12B shows a schematic of cell-free DNA molecules and copies thereof: cfDNA molecule 1, cfDNA molecule 2, cfDNA molecule 3, and cfDNA molecule 4. Each of these cfDNA molecules comprise the same start and stop positions: start position 990 and stop position 1140. Both cfDNA molecule 1 and cfDNA molecule 2 and copies thereof do not comprise any variants. These molecules are grouped in a single group (‘Family 1’). The molecules corresponding to cfDNA molecule 3 comprise a variant at position 1000 (a T to A variant). These molecules are grouped in a single group (‘Family 2’). The molecules corresponding to cfDNA molecule 4 comprise a variant at position 1000 and a variant at position 1100 (T to A, and G to A, respectively). These molecules are grouped in a single group (‘Family 3’). A consensus sequence for each of the groups of molecules (Family 1, Family 2, and Family 3) is generated (e.g. as part of a deduplication process) to generate three consensus sequences representing the sequences corresponding to cfDNA molecules 1 and 2, cfDNA molecule 3, and cfDNA molecule 4. An example workflow of using start and stop positions instead of molecular barcodes (e.g. UMIs) and without using analysis of variants within sequences that share a start and a stop position is shown in FIG. 12A. Here, the same cfDNA molecules as described in FIG. 12B are shown. Identification of a consensus sequence in this case is based only on start and stop positions. In this case, a single consensus sequence is generated.
[0050] There may be certain advantages to not using molecular barcodes for determining a molecular origin of a nucleic acid molecule. In some cases, not using molecular barcodes may result in a higher yield of nucleic acid material that is analyzed. For example, a workflow may comprise amplifying cell-free DNA molecules without tagging the cell-free DNA molecules with molecular barcodes. In some cases, tagging nucleic acid molecules with molecular barcodes may lead to sample loss (e.g., fewer cell-free DNA molecules). Not including a molecular tagging process may avoid sample loss prior to or after amplification. In some cases, not using molecular barcodes may be helpful in removing or limiting errors arising from amplification and / or sequencing. Such errors may be introduced during amplification and / or sequencing of nucleicAttorney Docket No. 58626-723601 acid molecules. Nucleic acid molecules (e.g., cell-free DNA molecules) may be amplified to generate amplified nucleic acid molecules. Amplifying the nucleic acid molecules may be performed using PCR. PCR may introduce one or more errors into the amplified nucleic acid molecules. The one or more errors introduced into the amplified nucleic acid molecules may comprise one or more nucleotides that are different than the nucleic acid molecule used as the basis of the amplification. For example, a nucleic acid molecule may comprise a sequence comprising 100 nucleotides in length. The sequence of the nucleic acid molecule may comprise a guanine (G) at the 50thposition of the nucleic acid molecule. The nucleic acid molecule may be amplified to generate a plurality of amplified nucleic acid molecules. A first amplified nucleic acid molecule of the plurality of amplified nucleic acid molecule may comprise an adenosine (A) at the 50thposition of the sequence of the amplified nucleic acid molecule due to an amplification error. The amplified nucleic acid molecule comprising the A at the 50thposition may be amplified to generate additional copies of amplified nucleic acid molecules comprising an A at the 50thposition. Other amplified nucleic acid molecules of the amplified nucleic acid molecules may (correctly) comprise a G at the 50thposition like the original molecule. The A at the 50thposition of the amplified nucleic acid molecule comprising the A at the 50thposition may be an example of an error introduced during PCR amplification. In some cases, use of information about the nucleic acid molecule may be used to remove errors introduced during amplification (e.g. during PCR) from analysis of sequencing data. Information regarding a start position of a nucleic acid molecule sequence, stop position of a nucleic acid molecule sequence, one or more SNVs and / or PVs, or a combination thereof may enable identifying errors introduced during amplification (e.g. PCR). Sequences associated with amplified nucleic acid molecules may be grouped based on their start position and / or stop positions to identify a group of sequences, thereby forming a family. Each group of sequences (or family) may be analyzed to identify nucleotide positions within the sequences where different nucleotides are present at the same position across sequences of the group of sequences. In some cases, positions comprising different nucleotides across sequences of the group of sequences may reflect different nucleic acid molecules in the biological sample. In other cases, positions comprising different nucleotides across sequence of the group of sequences may instead reflect an error introduced during amplification. By analyzing the frequency of sequences comprising different nucleotides at a position (after the initial grouping of molecules into families based on start and stop positions), variants reflective of different original molecules and errors introduced during PCR may be identified and distinguished. In some cases, errors introduced during PCR may be identified based on a lower frequency relative to variants present in original molecules. ForAttorney Docket No. 58626-723601 example, a group (or family) may comprise 20 sequences. About 50% of the sequences may comprise an A at position 10 of the sequences and about 50% of the sequences may comprise a G at position 10 of the sequence. About 10% of the sequences may comprise a T at position 15 of the sequences and about 90% of the sequence may comprise a C at position 15 of the sequences. In some cases, the A at position 10 and the G at position 10 may be identified as variants of two original nucleic acid molecules within the group. In some cases, the T at position 15 may be identified as an error introduced during PCR based on a of its decreased prevalence relative to the variants at position 10. Sequences comprising the T at position 15 may be corrected (e.g., by collapsing to a consensus sequence) to include a C at position 15.
[0051] Examples of workflows where errors are introduced during amplification are shown in FIGs. 10 and 11. FIG. 10 shows an example of using molecular barcodes for determining a molecular origin of amplified nucleic acid molecules. In this case, three cfDNA molecules (cfDNA molecule 1, cfDNA molecule 2, and cfDNA molecule 3 are PCR amplified. Each cfDNA molecule comprises a unique molecule barcode using a combination of two UMIs, as shown. The cfDNA molecules are amplified. Each cfDNA molecule is grouped into a group of sequences (e.g., a family of sequences) based on the UMI sequences. An amplification error is shown in one of the sequences associated with cfDNA molecule 1 after amplification. A consensus sequence may be generated for each group of sequences based on sequences that share UMIs. The consensus sequence for cfDNA molecule does not include the error introduced during amplification.
[0052] An example of a workflow that does not use molecular barcodes for determining a molecular origin of amplified nucleic acid molecules is shown in FIG. 11. In this case, three cfDNA molecules (cfDNA molecule 1, cfDNA molecule 2, and cfDNA molecule 3) are PCR amplified. The amplified nucleic acid molecules are grouped into families based on a start and stop position. In this case, further analysis is not performed to look for variants that may exist across sequences within the group. FIG. 11 shows that an amplification error is introduced, and variants are present. A consensus sequence is generated for each of the groups of sequences. In this case, a separate sequence comprising the variants of the group of sequences is not generated. Certain methods described herein have advantages over the workflow shown in FIG. 11 because the methods described herein enable identification of consensus sequences corresponding to variant-containing sequences. Additionally, certain methods described herein have the advantage that molecular barcodes may not be needed to identify consensus sequences comprising variants within groups of sequences with the same start and / or stop position.Attorney Docket No. 58626-723601
[0053] An identical sequence may refer to a sequence read of the plurality of sequence reads comprising 100% identity to a portion of the one or more reference genomic sequences. In some cases, the sequence read of the plurality of sequence reads may be identical to or substantially similar to the first portion of the one or more reference genomic sequences and be aligned to (e.g., mapped to) the first portion of the one or more reference genomic sequences. In some cases, the sequence read of the plurality of sequence reads may be determined to be identical to or substantially similar to the first portion of the one or more reference genomic sequences and the sequence read of the plurality of sequence reads may be compared to a second portion of the one or more reference genomic sequences to determine whether the sequence read of the plurality of sequence reads is better aligned to the second portion of the or more reference genomic sequences compared to the first genomic sequences (e.g., has a higher % identity to a second portion as compared to a first portion of the one or more reference genomic sequences). In some cases, a sequence read of the plurality of sequence reads may be compared to one or more portions of one or more reference genomic sequences to determine the portion of the one or more reference genomic sequences with the best alignment to the sequence reads of the plurality of sequence reads. For example, the sequence read of the plurality of sequence reads may be compared to more than 1000 portions (e.g., regions) of a human reference genomic sequence, and multiple portions may be considered substantially similar to the sequence read and one portion may be considered the most substantially similar (i.e., best aligned). The sequence read of the plurality of sequence reads may be mapped to the portion of the one or more reference genomic sequences with the best alignment (i.e., highest percentage sequence identity to the sequence read).
[0054] Mapping a sequence read of the plurality of sequence reads to one or more reference genomic sequences may comprise identifying one or more features of the sequence read and / or one or more reference genomic sequences. For example, a sequence read of the plurality of sequence reads may be mapped to a reference genomic sequence and a start position may be identified. The start position may refer to the nucleotide position of the reference genomic sequence corresponding to the first nucleotide of the sequence read of the plurality of sequence reads. For example, a sequence read may comprise a nucleotide sequence of 150 nucleotides which aligns with a portion of a human reference genomic sequence corresponding to nucleotide positions 756000-756149 on chromosome 1. The start position of the nucleotide sequence comprising 150 nucleotides may be nucleotide position 756000 of chromosome 1. The start position may correspond to the first nucleotide of the nucleotide sequence of 150 nucleotides. The start position may be given as a number corresponding to a nucleotide position within aAttorney Docket No. 58626-723601 reference genomic sequence. The number corresponding to a nucleotide position within a reference genomic sequence of a start position may be position information, chromosome information, or a combination thereof.
[0055] A stop position of a sequence read of the plurality of sequence reads may be identified. The stop position may refer to the nucleotide position of the reference genomic sequence corresponding to the last nucleotide of the sequence read of the plurality of sequence reads. For example, a sequence read may comprise a nucleotide sequence of 150 nucleotides which aligns with a portion of a human reference genomic sequence corresponding to nucleotide positions 756000-756149 on chromosome 1. The stop position of the nucleotide sequence comprising 150 nucleotides may be nucleotide position 756149 of chromosome 1. The stop position may be given as a number corresponding to a nucleotide position within a reference genomic sequence. The number corresponding to a nucleotide position within a reference genomic sequence of a stop position may be position information, chromosome information, or a combination thereof.
[0056] Mapping a sequence read of the plurality of sequence reads to one or more reference genomic sequences may comprise identifying a sequence identity of the sequence read of the plurality of sequence reads to the one or more reference genomic sequences (e.g., % identity of the sequence read relative to the at least a portion of a reference genomic sequence). The sequence identity may refer to the amount of the sequence read of the plurality of sequence reads that aligns to (e.g., maps to) the one or more reference genomic sequences. For example, a sequence read comprising 100 nucleotides may comprise 95 nucleotides of the 100 nucleotides that are identical to a portion of a human reference genomic sequence. In this case, the sequence read comprising 100 nucleotides would have 95% identity to the portion of the human reference genomic sequence.
[0057] Mapping a sequence read of the plurality of sequence reads to one or more reference genomic sequence may further comprise identifying one or more variants of sequence reads within the plurality of sequence reads. A sequence read variant may refer to a nucleotide at a position that is different than the nucleotide at the same position of a reference genomic sequence. For example, a sequence read may align to a portion of a human reference genomic sequence. The sequence read that aligns to the portion of the human reference genomic sequence may comprise a thymine (T) at nucleotide 30 of the sequence read and a cytosine (C) at the corresponding position of the human reference genomic sequence. A variant may comprise a single nucleotide variant, or a single nucleotide polymorphism. A variant may comprise an adenosine (A) to T substitution, an A to C substitution, an A to guanine (G) substitution, a T to A substitution, a T to C substitution, a T to G substitution, a C to A substitution, a C to TAttorney Docket No. 58626-723601 substitution, a C to G substitution, a G to A substitution, a G to C substitution, a G to T substitution, or a combination thereof. The one or more variants may comprise a singlenucleotide variant. One or more variants may be identified in a sequence read of the plurality of sequence reads by mapping the sequence read of the plurality of sequence reads to one or more reference genomic sequences. In some cases, at least about 1 variant, at least about 2 variants, at least about 3 variants, at least about 4 variants, at least about 5 variants, or more variants may be identified by mapping a sequence read of the plurality of sequence reads to one or more reference genomic sequences. In some cases, at most about 1 variant, at most about 2 variants, at most about 3 variants, at most about 4 variants, at most about 5 variants may be identified by mapping a sequence read of the plurality of sequence reads to one or more reference genomic sequences.
[0058] Sequence reads of the plurality of sequence reads may be grouped into groups of sequence reads (e.g., families of sequence reads). Grouping the sequence reads of the plurality of sequence reads into groups (e.g., families) may be based on the start positions of the sequence reads, the stop positions of the sequence reads, or a combination thereof. In some cases, sequence reads of the plurality of sequence reads may be grouped into one or more families. Sequence reads of the plurality of sequence reads that have identical (i.e., the same) start position may be grouped into the same family. For example, two sequence reads from a plurality of sequence reads may have the same start position corresponding to a nucleotide position on a specific chromosome when mapped to a reference genomic sequence. The two sequence reads with the same start position may be grouped into the same family. Sequence reads of the plurality of sequence reads that have identical (i.e., the same) stop position may be grouped into the same family. For example, two sequence reads from a plurality of sequence reads may have the same stop position corresponding to a nucleotide position on a specific chromosome when mapped to a reference genomic sequence. The two sequence reads with the same stop position may be grouped into the same family. Sequence reads of the plurality of sequence reads that have identical start and stop positions may be grouped into the same family. For example, two sequence reads from a plurality of sequence reads may have the same start position corresponding to a nucleotide position on a specific chromosome and the same stop position corresponding to a different nucleotide position on the specific chromosome. The two sequence reads with the same start and stop positions (e.g. sequence reads that map to the same positions in a reference genomic sequence) may be grouped into the same family.
[0059] Grouping the sequence reads of the plurality of sequence reads into groups (e.g., families) may generate one or more groups (e.g., families). In some cases, at least about 2Attorney Docket No. 58626-723601 groups, at least about 3 groups, at least about 4 groups, at least about 5 groups, at least about 6 groups, at least about 7 groups, at least about 8 groups, at least about 9 groups, at least about 10 groups, at least about 11 groups, at least about 12 groups, at least about 13 groups, at least about 14 groups, at least about 15 groups, at least about 16 groups, at least about 17 groups, at least about 18 groups, at least about 19 groups, at least about 20 groups, at least about 25 groups, at least about 30 groups, at least about 35 groups, at least about 40 groups, at least about 45 groups, at least about 50 groups, at least about 55 groups, at least about 60 groups, at least about 65 groups, at least about 70 groups, at least about 75 groups, at least about 80 groups, at least about 85 groups, at least about 90 groups, at least about 95 groups, at least about 100 groups, at least about 200 groups, at least about 300 groups, at least about 400 groups, at least about 500 groups, at least about 600 groups, at least about 700 groups, at least about 800 groups, at least about 900 groups, at least about 1,000 groups, at least about 2,000 groups, at least about 3,000 groups, at least about 4,000 groups, at least about 5,000 groups, at least about 6,000 groups, at least about 7,000 groups, at least about 8,000 groups, at least about 9,000 groups, at least about 10,000 groups, at least about 20,000 groups, at least about 30,000 groups, at least about 40,000 groups, at least about 50,000 groups, at least about 60,000 groups, at least about 70,000 groups, at least about 80,000 groups, at least about 90,000 groups, at least about 100,000 groups, at least about 200,000 groups, at least about 300,000 groups, at least about 400,000 groups, at least about 500,000 groups, at least about 600,000 groups, at least about 700,000 groups, at least about 800,000 groups, at least about 900,000 groups, at least about 1,000,000 groups, or more groups may be generated by grouping the sequence reads of the plurality of sequence reads into groups (e.g., families). In some cases, at most about 2 groups, at most about 3 groups, at most about 4 groups, at most about 5 groups, at most about 6 groups, at most about 7 groups, at most about 8 groups, at most about 9 groups, at most about 10 groups, at most about 11 groups, at most about 12 groups, at most about 13 groups, at most about 14 groups, at most about 15 groups, at most about 16 groups, at most about 17 groups, at most about 18 groups, at most about 19 groups, at most about 20 groups, at most about 25 groups, at most about 30 groups, at most about 35 groups, at most about 40 groups, at most about 45 groups, at most about 50 groups, at most about 55 groups, at most about 60 groups, at most about 65 groups, at most about 70 groups, at most about 75 groups, at most about 80 groups, at most about 85 groups, at most about 90 groups, at most about 95 groups, at most about 100 groups, at most about 200 groups, at most about 300 groups, at most about 400 groups, at most about 500 groups, at most about 600 groups, at most about 700 groups, at most about 800 groups, at most about 900 groups, at most about 1,000 groups, at most about 2,000 groups, at most about 3,000 groups, at most about 4,000 groups, atAttorney Docket No. 58626-723601 most about 5,000 groups, at most about 6,000 groups, at most about 7,000 groups, at most about 8,000 groups, at most about 9,000 groups, at most about 10,000 groups, at most about 20,000 groups, at most about 30,000 groups, at most about 40,000 groups, at most about 50,000 groups, at most about 60,000 groups, at most about 70,000 groups, at most about 80,000 groups, at most about 90,000 groups, at most about 100,000 groups, at most about 200,000 groups, at most about 300,000 groups, at most about 400,000 groups, at most about 500,000 groups, at most about 600,000 groups, at most about 700,000 groups, at most about 800,000 groups, at most about 900,000 groups, at most about 1,000,000 groups, or fewer groups may be generated by grouping the sequence reads of the plurality of sequence reads into groups (e.g., families). In some cases about 1-1,000,000 groups, about 2-900,000 groups, about 3-800,000 groups, about 4-700,000 groups, about 5-600,000 groups, about 6-500,000 groups, about 7-400,000 groups, about 8- 300,000 groups, about 9-200,000 groups, about 10-100,000 groups, about 11-90,000 groups, about 12-80,000 groups, about 13-70,000 groups, about 14-60,000 groups, about 15-50,000 groups, about 16-40,000 groups, about 17-30,000 groups, about 18-20,000 groups, about 19- 10,000 groups, about 20-9,000 groups, about 25-8,000 groups, about 30-7,000 groups, about 35- 6,000 groups, about 40-5,000 groups, about 45-4,000 groups, about 50-3,000 groups, about 55- 2,000 groups, about 60-1,000 groups, about 65-900 groups, about 70-800 groups, about 75-700 groups, about 80-600 groups, about 85-500 groups, about 90-400 groups, or about 100-300 groups may be generated by grouping the sequence reads of the plurality of sequence reads into groups (e.g., families).
[0060] A group of sequence reads (e.g., a family of sequence reads) of a plurality of groups of sequence reads (e.g., a plurality of families of sequence reads) may be grouped into subgroups of sequence reads (e.g., subfamilies of sequence reads). Grouping sequence reads of a group of sequence reads (e.g., a family of sequence reads) into one or more subgroups of sequence reads (e.g., one or more subfamilies of sequence reads) may be based on one or more features of the sequence reads of the group of sequence reads (e.g., sequence reads of the family of sequence reads). In some cases, grouping sequence reads of a group of sequence reads (e.g., a family of sequence reads) into one or more subgroups of sequence reads (e.g., one or more subfamilies of sequence reads) may be based on one or more nucleotides at one or more positions of the sequence reads. For example, the sequence reads of a group of sequence reads (e.g., a family of sequence reads) may be grouped into two subgroups, for example a first subgroup and a second subgroup. The sequence reads of the first subgroup may comprise a first nucleotide at a first position. The sequence reads of the second subgroup may comprise a second nucleotide at the first position. In some cases, the second nucleotide differs from the first nucleotide. For example,Attorney Docket No. 58626-723601 the first nucleotide may comprise a C and the second nucleotide may comprise a G. The first nucleotide at the first position may comprise a single-nucleotide variant. In some cases, grouping sequence reads of a group of sequence reads (e.g., a family of sequence reads) into one or more subgroups of sequence reads (e.g., one or more subfamilies of sequence reads) may be based on the frequency of one or more recurrent non-reference nucleotides.
[0061] The sequence reads of a subgroup (e.g., a first subgroup) of the plurality of subgroups as described herein may comprise sequence reads originally derived from a positive strand of a reference genomic sequence, a negative strand of a reference genomic sequence, or a combination thereof. The sequence reads of a subgroup (e.g., a first subgroup) of the plurality of subgroups as described herein may comprise sequence reads originally derived from a Watson strand of a reference genomic sequence, a Crick strand of a reference genomic sequence, or a combination thereof.
[0062] In some cases at least about 2 subgroups, at least about 3 subgroups, at least about 4 subgroups, at least about 5 subgroups, at least about 6 subgroups, at least about 7 subgroups, at least about 8 subgroups, at least about 9 subgroups, at least about 10 subgroups, or more subgroups may be generated by grouping the sequence reads of a group of sequence reads (e.g. family of sequence reads) into subgroups. In some cases at most about 2 subgroups, at most about 3 subgroups, at most about 4 subgroups, at most about 5 subgroups, at most about 6 subgroups, at most about 7 subgroups, at most about 8 subgroups, at most about 9 subgroups, at most about 10 subgroups, at most about 11 subgroups, at most about 12 subgroups, at most about 13 subgroups, at most about 14 subgroups, at most about 15 subgroups, at most about 16 subgroups, at most about 17 subgroups, at most about 18 subgroups, at most about 19 subgroups, at most about 20 subgroups, at most about 25 subgroups, at most about 30 subgroups, at most about 35 subgroups, at most about 40 subgroups, at most about 45 subgroups, at most about 50 subgroups, or fewer subgroups may be generated by grouping the sequence reads of a group of sequence reads (e.g., family of sequence reads) into subgroups.
[0063] The sequence reads of a subgroup of the one or more subgroups may be collapsed to form one or more consensus sequences. A consensus sequence of the one or more consensus sequences may correspond with a sequence of a nucleic acid molecule of a plurality of nucleic acid molecules. The nucleic acid molecule of the plurality of nucleic acid molecules may be part of a biological sample used to obtain sequence reads that are analyzed as part of the methods described herein. Collapsing the sequence reads of a subgroup of sequence reads may comprise determining the most prevalent nucleotide at each position of the sequence reads. A consensus sequence may be generated based on collapsing the sequence reads of the subgroup of sequenceAttorney Docket No. 58626-723601 reads. The consensus sequence may comprise the predominant nucleotide at each position of the sequence reads of the subgroup. In some cases, collapsing the sequence reads may comprise selecting the most frequent nucleotide at each position observed within the sequence reads. In some cases, the consensus sequence generated by collapsing the sequence reads may comprise the most frequently observed nucleotide at each position within the sequence reads. In some cases, collapsing the sequence reads to generate a consensus sequence may comprise determining a nucleotide sequence shared by multiple (e.g., greater than a threshold number, such as greater than 2, greater than 3, greater than 4, greater than 5, greater than 6, greater than 7, greater than 8, greater than 9, or greater than 10) sequence reads within the subset that have the same sequence and excluding sequence reads that do not. Determining the most prevalent nucleotide at each position of the sequence reads may comprise identifying each occurrence of a nucleotide at a position of the sequence reads and identifying a nucleotide that occurs most frequently at the position of the sequence reads. For example, some sequence reads of a subgroup of sequence reads may comprise a C at a given position of the sequence reads and some other sequence reads of the subgroup of sequence reads may comprise a T at the given position of the sequence reads. The number of sequence reads of the subgroup of sequence reads that comprise a C at the given position may be greater than the number of sequence reads of the subgroup of sequence reads that comprise a T at the given position. The consensus sequence may comprise a C at the given position. In some cases, sequence reads of a subgroup (e.g., a first subgroup) may be collapsed to form a consensus sequence (e.g., a first consensus sequence). The consensus sequence (e.g., a first consensus sequence) may correspond to the sequence of a nucleic acid molecule (e.g., a first nucleic acid molecule) of a plurality of nucleic acid molecules. In some cases, sequence reads of another subgroup (e.g., a second subgroup) may be collapsed to form another consensus sequence (e.g., a second consensus sequence). The other consensus sequence (e.g., a second consensus sequence) may correspond to the sequence of another nucleic acid molecule (e.g., a second nucleic acid molecule) pf the plurality of nucleic acid molecules. In some cases, the subgroup and other subgroup (e.g., first subgroup and second subgroup) may be from the same group or family (e.g., may share the same start and / or stop positions). In some cases, the nucleic acid molecule (e.g., the first nucleic acid molecule) of the subgroup (e.g., the first subgroup) and the other nucleic acid molecule (e.g., the second nucleic acid molecule) of the other subgroup (e.g., the second subgroup) may differ in sequence. For example, the first nucleic acid molecule may comprise a sequence that is different from the second nucleic acid molecule by one or more nucleotides. The consensus sequences generated from collapsing sequence reads may result in a measurable error rate, as compared to a reference genome. The error rate mayAttorney Docket No. 58626-723601 refer to the frequency of the consensus sequence comprising an incorrect nucleotide (e.g. the consensus sequence comprising a PCR generated error), as measured against a reference genomic sequence (e.g. hgl9 or a personalized germline genome sequence). The error rate of the consensus sequence may be determined from averaging a rate of non-reference nucleotides. The error rate of the consensus sequence may be less than or equal to 1 ppm, less than or equal to 0.5 ppm, less than or equal to 0.2 ppm, less than or equal to 0.1 ppm, less than or equal to 0.09 ppm, less than or equal to 0.08 ppm, less than or equal to 0.07 ppm, less than or equal to 0.06 ppm, less than or equal to 0.05 ppm, less than or equal to 0.04 ppm, less than or equal to 0.03 ppm, less than or equal to 0.02 ppm, less than or equal to 0.01 ppm, or less.
[0064] A nucleic acid molecule of the plurality of nucleic acid molecules of the consensus sequence as described herein may be derived from a variety of sources. In some cases, the nucleic acid molecule of the plurality of nucleic acid molecule may be derived from a biological sample of a subject. The biological sample of the nucleic acid molecule may comprise a body fluid sample, a tissue sample, an exogenous sample, or a combination thereof. The body fluid sample may comprise a urine sample, a saliva sample, a blood sample, a plasma sample, a whole blood sample, a vaginal secretion sample, a semen sample, a serum sample, a cerebrospinal fluid sample, a bone marrow sample, or a combination thereof. The tissue sample may comprise a surgical resection, a needle-core biopsy, a tissue slice, or a combination thereof. The exogenous sample may comprise a tissue culture sample, a cell culture sample, or a combination thereof. The blood sample may comprise plasma, buffy coat, erythrocytes, cell-free nucleic acid molecules or a combination thereof. The cell-free nucleic acid molecules of the biological sample may comprise one or more cfDNA molecules.
[0065] In some cases, a subgroup (e.g., the first subgroup) may comprise a plurality of nonidentical sequence reads. The plurality of non-identical sequence reads of the first subgroup may comprise sequence reads that differ by at least one nucleotide. For example, the first subgroup may comprise a plurality of sequence reads that may comprise the same start position, the same stop position, the same nucleotide at a first position, or a combination thereof. In addition, the first subgroup may comprise sequences within the plurality of sequences of the first subgroup that may differ at a second position. For example, some sequence reads of the plurality of sequence reads of the first subgroup may comprise a first nucleotide at the second position and some other sequence reads of the plurality of sequence reads of the first subgroup may comprise a second nucleotide at the second position. The first nucleotide of the second position and the second nucleotide of the second position may differ (i.e., be different nucleotides). The number of sequence reads of the plurality of sequence reads of the first subgroup that comprises the firstAttorney Docket No. 58626-723601 nucleotide at the second position may differ from (e.g., be greater than) the number of sequence reads of the plurality of sequence reads of the first subgroup that comprises the second nucleotide at the second position (e.g., the first subgroup may have differing levels of recurrence of the first nucleotide at the second position and the second nucleotide at the second position). The different levels of recurrence of the first nucleotide at the second position and the second nucleotide at the second position may arise from one or more polymerase chain reaction (PCR) errors. For example, the sample used to generate sequence reads of the plurality of sequence reads may have undergone one or more rounds of PCR. Each round of PCR may generate amplified nucleic acid molecules based on the original nucleic acid molecules. The one or more rounds of PCR may introduce one or more variants into the amplified nucleic acid molecules (e.g., PCR errors). The one or more variants of the amplified nucleic acid molecules may represent differences at a nucleotide position in the sequence of the amplified nucleic acid molecules relative to the original nucleic acid molecules. The one or more variants of the amplified nucleic acid molecules may be detectable as differences in the sequence reads of subgroups of sequence reads in the methods described herein.
[0066] A subgroup (e.g., a first subgroup) of the one or more subgroups described herein may comprise a two or more single-nucleotide variants (e.g., a first single-nucleotide variant and a second nucleotide variant). In some cases, the subgroup (e.g., the first subgroup) may comprise a first nucleotide at a first position. The first nucleotide at the first position of the subgroup (e.g., the first subgroup) may be a single-nucleotide variant (e.g., the first single nucleotide variant). In some cases, the subgroup (e.g., the first subgroup) may comprise a second nucleotide at a second position. The second nucleotide at the second position of the subgroup (e.g., the first subgroup) may be a single-nucleotide variant (e.g., the second single nucleotide variant). The first nucleotide at the first position (e.g., first single-nucleotide variant) of the subgroup (e.g., the first subgroup) and the second nucleotide at the second position (e.g., second single-nucleotide variant) of the subgroup (e.g., the first subgroup) may be part of the same nucleic acid molecule. The first nucleotide at the first position (e.g., first single-nucleotide variant) of the subgroup (e.g., the first subgroup) and the second nucleotide at the second position (e.g., second single- nucleotide variant) of the subgroup (e.g., the first subgroup) may be within about 170 base pairs of each other. In some cases, the first nucleotide at the first position (e.g., first single-nucleotide variant) of the subgroup (e.g., the first subgroup) and the second nucleotide at the second position (e.g., second single-nucleotide variant) of the subgroup (e.g. the first subgroup) may be within about 160 nucleotides, about 165 nucleotides, about 170 nucleotides, about 175 nucleotides, or about 180 nucleotides of each other.Attomey Docket No. 58626-723601
[0067] In some cases, the first nucleotide at the first position (e.g., the first single-nucleotide variant) of the subgroup (e.g., the first subgroup) and the second nucleotide at the second position (e.g., the second single-nucleotide variant) of the subgroup (e.g., the first subgroup) may be separated by one or more nucleotides. In some cases, the first nucleotide at the first position (e.g. first single-nucleotide variant) of the subgroup (e.g., the first subgroup) and the second nucleotide at the second position (e.g., second single-nucleotide variant) of the subgroup (e.g., the first subgroup) may be separated by at least about at least about 1 nucleotide, at least about 2 nucleotides, at least about 3 nucleotides, at least about 4 nucleotides, at least about 5 nucleotides, at least about 6 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, at least about 50 nucleotides, at least about 55 nucleotides, at least about 60 nucleotides, at least about 65 nucleotides, at least about 70 nucleotides, at least about 75 nucleotides, at least about 80 nucleotides, at least about 85 nucleotides, at least about 90 nucleotides, at least about 95 nucleotides, at least about 100 nucleotides, at least about 105 nucleotides, at least about 110 nucleotides, at least about 115 nucleotides, at least about 120 nucleotides, at least about 125 nucleotides, at least about 130 nucleotides, at least about 135 nucleotides, at least about 140 nucleotides, at least about 145 nucleotides, at least about 150 nucleotides, at least about 155 nucleotides, or at least about 160 nucleotides.B) Use of low error rate somatic mutations
[0068] In another aspect, the present disclosure provides a method. The method may comprise sequencing a plurality of cell-free DNA molecules. In some cases, the method may comprise having sequenced a plurality of cell-free DNA molecules. The cell-free DNA molecules may be obtained from a subject. In some cases, the cell-free DNA molecules may be derived from a subject. Sequencing or having sequenced a plurality of cell-free DNA molecules may produce a plurality of sequence reads. The method may comprise aligning the plurality of sequence reads to produce a plurality of aligned sequence reads. The method may comprise processing the plurality of aligned sequence reads to detect a presence of a somatic mutation in one or more cell-free DNA molecules from among the plurality of cell-free DNA molecules. The somatic mutation may comprise a base-change mutation selected from the group consisting of an adenine-to- thymine (A>T) mutation, an adenine-to-cytosine (A>C) mutation, an adenine-to-guanine (A>G) mutation, a thymine-to-adenine (T>A) mutation, a thymine-to-cytosine (T>C) mutation, a thymine-to-guanine (T>G) mutation, a cytosine-to-adenine (C>A) mutation, a cytosine-to-Attorney Docket No. 58626-723601 thymine (C>T) mutation, a cytosine-to-guanine (OG) mutation, a guanine-to-adenine (G>A) mutation, a guanine-to-thymine (G>T) mutation, and a guanine-to-cytosine (G>C) mutation. In some cases, processing the plurality of aligned sequence reads to detect a presence of a somatic mutation in one or more cell-free DNA molecules from among the plurality of cell-free DNA molecules may comprise filtering the plurality of aligned sequence reads based at least in part on an expected error rate of detection of the base-change mutation in individual aligned sequence reads, thereby generating a set of filtered sequence reads. The method may comprise processing the set of filtered sequence reads to determine whether at least one of the set of filtered sequence reads is mutation-containing. In some cases, determining whether at least one of the set of filtered sequence reads is mutation-containing may have an error rate of no more than 0.5 parts per million (ppm).
[0069] In some embodiments, the method comprises sequencing a plurality of cell-free nucleic acid molecules. In some embodiments, the method comprises having sequenced a plurality of cell-free nucleic acid molecules. In some embodiments, the present disclosure provides methods for sequencing a plurality of cell-free DNA molecules. In some embodiments, the method comprises having sequenced a plurality of cell-free DNA molecules. The plurality of cell-free DNA molecules can comprise a plurality of mutations. The plurality of mutations can comprise a plurality of single nucleotide variants (SNVs). The plurality of mutations can comprise a plurality of phased variants (PVs). In some embodiments, “phased variants” or “PVs,” as used interchangeably herein, generally refers to two or more mutations (e.g., SNVs or indels) that occur in cis (i.e., on the same strand of a nucleic acid molecule) within a single cell-free nucleic acid molecule. In some cases, a cell-free nucleic acid molecule can be a cell-free deoxyribonucleic acid (cfDNA) molecule. In some cases, a cfDNA molecule can be derived from a diseased tissue, such as a tumor (e.g., a circulating tumor DNA (ctDNA) molecule). The plurality of cell-free DNA molecules can be obtained or derived from a subject. In some embodiments, the subject has cancer.
[0070] In some embodiments, the methods described herein may comprise aligning the plurality of sequence reads to a reference sequence. The reference sequence may be any one of the reference sequences described herein. In some embodiments, the reference sequence may be a reference genomic sequence. In some embodiments, the reference genomic sequence is derived from an animal. In some embodiments, the reference genomic sequence may be derived from a mammal. In some embodiments, the reference genomic sequence may be derived from a human. In some embodiments, aligning the plurality of sequence reads to a reference sequence can produce a plurality of aligned sequence reads. In some embodiments, the method furtherAttorney Docket No. 58626-723601 comprises processing the plurality of aligned sequence reads. In some embodiments, the method further comprises processing the plurality of aligned sequence reads to detect a presence of a mutation. In some embodiments, the mutation is a somatic mutation. In some embodiments, the somatic mutation can include a single nucleotide variant (SNV).
[0071] In some embodiments, the somatic mutation may be in one or more cell-free DNA molecules from among the plurality of cell-free DNA molecules. In some embodiments, the somatic mutation is a base-change mutation. In some embodiments, the base-change mutation can be an adenine-to-thymine (A>T) mutation, an adenine-to-cytosine (A>C) mutation, an adenine-to-guanine (A>G) mutation, a thymine-to-adenine (T>A) mutation, a thymine-to- cytosine (T>C) mutation, a thymine-to-guanine (T>G) mutation, a cytosine-to-adenine (C>A) mutation, a cytosine-to-thymine (OT) mutation, a cytosine-to-guanine (OG) mutation, a guanine-to-adenine (G>A) mutation, a guanine-to-thymine (G>T) mutation, or a guanine-to- cytosine (G>C) mutation. In some embodiments, processing the plurality of aligned sequence reads further comprises filtering the plurality of aligned sequence reads. In some embodiments, filtering the plurality of aligned sequence reads is based at least in part on an error rate (e.g. an expected error rate) of detection of the base-change mutation in individual aligned sequence reads. In some embodiments, the error rate (e.g. expected error rate) can be determined by one or more factors such as the specific base-change mutation (e.g., A>T, A>C, A>G, T>A, T>C, T>G, OA, OT, OG, G>A, G>T, G>C), the number of copies of cell-free DNA molecules observed with the same somatic mutation, the number of independent copies of Watson and Crick strands observed carrying the same somatic mutation, and the location of the somatic mutation among the sequencing read. Watson and Crick strands can mean the two complementary strands of a DNA molecule. In some cases, the error rate will be determined (or lowered) by excluding nucleic acid molecules from analysis. In some cases, nucleic acid molecules may be excluded from analysis based on a distance of a somatic mutation from an end of the nucleic acid molecule. For example, in some cases, a nucleic acid molecule may be excluded from analysis when the nucleic acid molecule comprises a somatic mutation within 1 nucleotide, 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, 20 nucleotides, 21 nucleotides, 22 nucleotides, 23 nucleotides, 24 nucleotides, 25 nucleotides, 26 nucleotides, 27 nucleotides, 28 nucleotides, 29 nucleotides, and / or 30 nucleotides from an end of the nucleic acid molecule. In some cases, nucleic acid molecules may be excluded from analysis based on a distance of a somatic mutation from an end of the nucleic acid molecule. In some cases, a nucleicAttorney Docket No. 58626-723601 acid molecule may be excluded from analysis when the nucleic acid molecule comprises a somatic mutation within 10 nucleotides from an end of the nucleic acid molecule. In some cases, a nucleic acid molecule may be excluded from analysis when the nucleic acid molecule comprises a somatic mutation within 25 nucleotides from an end of the nucleic acid molecule.
[0072] The methods described herein may comprise identifying a sequence as mutation containing. Identifying the sequence as mutation containing may comprise an error rate. The error rate of identifying a sequence as mutation containing may refer to a frequency of incorrectly identifying a variant within the sequence. For example, an error rate of 1 part per million sequence reads may refer to incorrectly identifying a variant within the sequence 1 time out of a million sequence reads. Identifying a variant in the methods described herein may comprise identifying one or more SNVs and / or PVs associated with the variant. The error rate of identifying a variant within a sequence may be sequence dependent. For example, a variant comprising a OT mutation may have a higher error rate relative to an A>C mutation. The error rate of identifying a variant within a sequence may depend on the number of sequences comprising a variant that are detected. In some cases, the error rate of identifying a variant within a sequence may depend on the sequence context of the variant and the number of sequences comprising the variant. For example, an error rate for detecting a mutation, e.g. a T>G mutation, may be a first value with just a few sequences comprising that mutation identified (e.g. two sequences comprising a T>G mutation). An error rate for detecting the mutation (e.g. the T>G mutation) may be a second value with detecting more sequences comprising that mutation identified (e.g. 10 sequences comprising a T>G mutation). In some cases, the second value may be lower than the first value (e.g. the error rate for a given mutation may be higher when detecting fewer sequences comprising the mutation compared with the error rate when detecting more sequences comprising the mutation). In some cases, the second value may be approximately the same as the first value (e.g. the error rate for a given mutation may be about the same when detecting fewer sequences comprising the mutation compared with the error rate when detecting more sequences comprising the mutation). Mutations representing different somatic mutations may comprise different error rates when the same number of sequences are detected. For example, the error rate for detecting a T>C mutation when 5 sequences comprising the T?C mutation are analyzed may be higher compared to the error rate for detecting a T to G mutation when 5 sequences comprising the T>G mutation are analyzed.
[0073] In some embodiments, filtering the plurality of aligned sequence reads can generate a set of filtered sequence reads. In some embodiments, processing the plurality of aligned sequence reads can include trimming an end of the sequencing read to generate the set of filteredAttorney Docket No. 58626-723601 sequencing read. In some embodiments, trimming an end of the sequencing read can include trimming at most 1 base pair, at most 2 base pairs, at most 3 base pairs, at most 4 base pairs, at most 5 base pairs, at most 6 base pairs, at most 7 base pairs, at most 8 base pairs, at most 9 base pairs, at most 10 base pairs, at most 11 base pairs, at most 12 base pairs, at most 13 base pairs, at most 14 base pairs, at most 15 base pairs, at most 16 base pairs, at most 17 base pairs, at most 18 base pairs, at most 19 base pairs, or at most 20 base pairs. In some embodiments, processing the plurality of aligned sequence reads can include trimming both ends of the sequencing read to generate the set of filtered sequencing read. In some embodiments, processing the plurality of aligned sequence reads can include trimming both ends of the sequencing read to generate the set of filtered sequencing read as different number of base pairs. In some embodiments, determining the number of base pairs to trim from either ends of an aligned sequencing in the plurality of aligned sequence reads is based at least in part on a strand location the base-change mutation on the aligned sequencing read. In some embodiments, the set of filtered sequence reads comprises variant types of somatic mutations (e.g., SNVs). In some embodiments, the set of filtered sequence reads comprises a uniformly low expected error rate of detection across at least 20% of the variant types, at least 30% of the variant types, at least 40% of the variant types, at least 50% of the variant types, at least 60% of the variant types, at least 70% of the variant types, at least 80% of the variant types, at least 90% of the variant types, or 100% of the variant types. In some embodiments, the uniformly low expected error rate is no more than 0.005 parts per million (ppm), 0.01 parts per million (ppm), 0.02 parts per million (ppm), 0.03 parts per million (ppm), 0.04 parts per million (ppm), 0.05 parts per million (ppm), 0.06 parts per million (ppm), 0.07 parts per million (ppm), 0.08 parts per million (ppm), 0.09 parts per million (ppm), 0.1 parts per million (ppm), or 0.2 parts per million (ppm).
[0074] In some embodiments, the methods described herein may comprise processing the set of filtered sequence reads. In some embodiments, the method comprises processing the set of filtered sequence reads to determine whether at least one of the set of filtered sequence reads is mutation-containing. In some embodiments, determining whether at least one of the set of filtered sequence reads is mutation-containing has an error rate of no more than 0.005 parts per million (ppm), 0.01 parts per million (ppm), 0.02 parts per million (ppm), 0.03 parts per million (ppm), 0.04 parts per million (ppm), 0.05 parts per million (ppm), 0.06 parts per million (ppm), 0.07 parts per million (ppm), 0.08 parts per million (ppm), 0.09 parts per million (ppm), 0.1 parts per million (ppm), or 0.2 parts per million (ppm). In some embodiments, the determination of whether the set of filtered sequence reads is mutation-containing determines whether a cell-free DNA molecule of the plurality of cell-free DNA molecules is mutation-containing.Attorney Docket No. 58626-723601
[0075] In some embodiments, processing the plurality of aligned sequence reads further comprises pre-processing the plurality of aligned sequence reads to filter out at least a portion of the plurality of aligned sequence reads. In some embodiments, the preprocessing comprises demultiplexing FASTQ files, extracting unique molecular identifiers (UMIDs), performing molecular barcode-mediated error suppression (e.g., integrated digital error suppression), performing background polishing, and / or deduplicating barcodes.
[0076] In some embodiments, processing the plurality of aligned sequence reads further comprises processing the plurality of aligned sequence reads using a somatic mutation calling algorithm. In some embodiments, the somatic mutation calling algorithm can be MuTect2, VarScan2, or Strelka2.
[0077] In some embodiments, processing the plurality of aligned sequence reads may comprise comparing the expected error rate of detection of the base-change mutation to a pre-determined error rate threshold. For example, a pre-determined error rate threshold may be determined by selecting sequencing reads that meet a criterion (e.g. a threshold number of molecules observed). In some embodiments, this predetermined set of criteria may consider only molecules that have more than a specific number of duplicate molecules observed. In some embodiments, processing the plurality of aligned sequence reads further comprises processing the base-change mutation using a trained algorithm. In some embodiments, the trained algorithm can be a machine learning model or a statistical model. The machine learning model can include a supervised learning model, and unsupervised learning model, a semi-supervised learning model, a reinforcement learning model, or a combination thereof. The machine learning model can include a natural language processing model, an artificial neural network, a decision tree, a random forest, a naive bayes classifier, a boosting algorithm, a k-nearest neighbor algorithm, or a clustering model (e.g., k-means clustering model, a hierarchical clustering model). The statistical model can include a regression model (e.g., linear regression model, a logistic regression model), a classification model, a Monte Carlos simulation, or a polynomial model (e.g., binomial model, ternary model).
[0078] In some embodiments, processing the set of filtered sequence reads to determine whether at least one of the set of filtered sequence reads is mutation-containing is further based at least in part on a number of copies of the one or more cell-free DNA molecules having the somatic mutation. In some embodiments, the number of copies is compared to a threshold based at least in part on the expected error rate of detection of the base-change mutation. In some embodiments, processing the set of filtered sequence reads to determine whether at least one of the set of filtered sequence reads is mutation-containing is further based at least in part on a number of independent copies of both Watson and Crick complementary strands present in theAttorney Docket No. 58626-723601 one or more cell-free DNA molecules having the somatic mutation. In some embodiments, the number of independent copies is compared to a threshold based at least in part on the expected error rate of detection of the base-change mutation. In some embodiments, processing the set of filtered sequence reads to determine whether at least one of the set of filtered sequence reads is mutation-containing is further based at least in part on a strand location of the base-change mutation in the one or more cell-free DNA molecules. In some embodiments, the strand location is at least 1 nucleotide, at least 2 nucleotides, at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 20 nucleotides, at least 25 nucleotides, or at least 30 nucleotides away from a 5’ end or a 3’ end of the one or more cell-free DNA molecules.
[0079] In some embodiments, the somatic mutation can be the adenine-to-thymine (A>T) mutation, the adenine-to-cytosine (A>C) mutation, the adenine-to-guanine (A>G) mutation, the thymine-to-adenine (T>A) mutation, the thymine-to-cytosine (T>C) mutation, the thymine-to- guanine (T>G) mutation, the cytosine-to-adenine (C>A) mutation, the cytosine-to-thymine (C>T) mutation, the cytosine-to-guanine (OG) mutation, the guanine-to-adenine (G>A) mutation, the guanine-to-thymine (G>T) mutation, or the guanine-to-cytosine (G>C) mutation.
[0080] In some embodiments, the somatic mutation is the adenine-to-thymine (A>T) mutation, the cytosine-to-guanine (OG) mutation, or the adenine-to-cytosine (A>C) mutation. In some embodiments, the somatic mutation does not comprise any of the adenine-to-thymine (A>T) mutation, the cytosine-to-guanine (OG) mutation, or the adenine-to-cytosine (A>C) mutation. In some embodiments, the somatic mutation does not comprise any of the cytosine-to-thymine (OT) mutation, the cytosine-to-adenine (OA) mutation, or the adenine-to-guanine (A>G) mutation.
[0081] In some embodiments, filtering the plurality of aligned sequence reads may comprise selecting a subset of the plurality of aligned sequence reads. In some embodiments, filtering the plurality of aligned sequence reads further comprises analyzing the subset of the plurality of aligned sequence reads. In some embodiments, the subset is selected from among the plurality of aligned sequence reads, based at least in part on having a base-change mutation. In some embodiments, the subset is selected from among the plurality of aligned sequence reads, based at least in part on having a lowest expected error rate of detection of the base-change mutation. In some embodiments, the subset is selected from among an aligned sequencing read of the plurality of aligned sequence reads, based at least in part on having a base-change mutationAttorney Docket No. 58626-723601 located at least 1 nucleotide, at least 2 nucleotides, at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 20 nucleotides, at least 25 nucleotides, or at least 30 nucleotides away from a 5’ end or a 3’ end from the aligned sequence reads. In some embodiments, the subset is selected from among the plurality of aligned sequence reads, based at least in part on the number of copies of the aligned sequencing read in the plurality of aligned sequence reads. In some embodiments, the subset is selected from among the plurality of aligned sequence reads, based at least in part on the number of copies of the aligned sequencing read in the plurality of aligned sequence reads having the somatic mutation. In some embodiments, the subset is selected from among the plurality of aligned sequence reads base at least in part on the number of independent copies of both Watson and Crick complementary strands of the aligned sequencing read present in the plurality of aligned sequence reads. In some embodiments, the subset is selected from among the plurality of aligned sequence reads base at least in part on the number of independent copies of both Watson and Crick complementary strands present in the plurality of aligned sequence reads having the somatic mutation.
[0082] In some aspects, the methods described herein may comprise use of molecular barcodes. Molecular barcodes may be useful for identifying a source (e.g. a molecular origin) of a nucleic acid molecule after amplification. For example, a sample comprising cell-free DNA molecules may be tagged with molecular barcodes such that each cell-free DNA molecule is tagged with a unique molecular barcode of the molecular barcodes. The tagged cell-free DNA molecules may be amplified. The amplified tagged cell-free DNA molecules may be sequencing using any one of the sequencing methods described herein to generate read pairs. The read pairs may be analyzed to identify molecular barcodes associated with each read pair. The molecular barcode associated with each read pair may be used to identify amplified cell-free DNA molecules originating from the same original cell-free DNA molecule before amplification. Identifying amplified cell-free DNA molecules originating from the same original cell-free DNA molecule may be useful for quantifying, analyzing, and / or counting cell-free DNA molecules of a particular sequence (e.g. cell-free DNA molecules comprising one or more SNVs and / or PVs). In some cases, molecular barcodes may be used in combination with sample barcodes. The molecular barcode may provide information about the origin of a nucleic acid molecule associated with the molecular barcode. The sample barcode may provide information about the sample from which a nucleic acid molecule comprising the sample barcode originates. For example, multiple biological samples may be obtained. Nucleic acid molecules (e.g. cell-freeAttorney Docket No. 58626-723601DNA molecules) may be extracted from each biological sample. The nucleic acid molecules extracted from each biological sample may be tagged with a sample barcode, where sample barcodes corresponding to different samples may be different. In some cases, one or more sample barcodes may be added to nucleic acid molecules of a sample to provide information about the sample source of the nucleic acid molecules. For example, two sample barcodes may be added to nucleic acid molecules associated with a sample. The combination of the two sample barcodes may provide information regarding the sample source (e.g. sample origin) of the nucleic acid molecules comprising the two sample barcodes.
[0083] In some cases, molecular barcodes may be added to one or more ends of nucleic acid molecules of the present disclosure. A molecular barcode may be added to a 5’ end, a 3’ end or a combination thereof of a nucleic acid molecule (e.g. a cell-free DNA molecule). The molecular barcode may be single-stranded. The molecular barcode may be double stranded. In some cases, the molecular barcode may be part of a Y-shaped adaptor. The Y-shaped adaptor may be added to one or both ends (e.g. a 5’ end and / or a 3’ end) of a nucleic acid molecule (e.g. a cell-free DNA molecule). The molecular barcode may comprise nucleic acid. In some cases, the molecular barcode may comprise DNA, RNA, or a combination thereof. The molecular barcode may have a variety of lengths. The molecular barcode may have a length of at least about 2 nucleotides, at least about 3 nucleotides, at least about 4 nucleotides, at least about 5 nucleotides, at least about 6 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 11 nucleotides, at least about 12 nucleotides, at least about 13 nucleotides, at least about 14 nucleotides, at least about 15 nucleotides, at least about 16 nucleotides, at least about 17 nucleotides, at least about 18 nucleotides, at least about 19 nucleotides, at least about 20 nucleotides, at least about 21 nucleotides, at least about 22 nucleotides, at least about 23 nucleotides, at least about 24 nucleotides, at least about 25 nucleotides, at least about 26 nucleotides, at least about 27 nucleotides, at least about 28 nucleotides, at least about 29 nucleotides, at least about 30 nucleotides, at most about 2 nucleotides, at most about 3 nucleotides, at most about 4 nucleotides, at most about 5 nucleotides, at most about 6 nucleotides, at most about 7 nucleotides, at most about 8 nucleotides, at most about 9 nucleotides, at most about 10 nucleotides, at most about 11 nucleotides, at most about 12 nucleotides, at most about 13 nucleotides, at most about 14 nucleotides, at most about 15 nucleotides, at most about 16 nucleotides, at most about 17 nucleotides, at most about 18 nucleotides, at most about 19 nucleotides, at most about 20 nucleotides, at most about 21 nucleotides, at most about 22 nucleotides, at most about 23 nucleotides, at most about 24 nucleotides, at most about 25Attorney Docket No. 58626-723601 nucleotides, at most about 26 nucleotides, at most about 27 nucleotides, at most about 28 nucleotides, at most about 29 nucleotides, at most about 30 nucleotides, about 2 nucleotides to about 30 nucleotides, about 3 nucleotides to about 29 nucleotides, about 4 nucleotides to about 28 nucleotides, about 5 nucleotides to about 27 nucleotides, about 6 nucleotides to about 26 nucleotides, about 7 nucleotides to about 25 nucleotides, about 8 nucleotides to about 24 nucleotides, about 9 nucleotides to about 23 nucleotides, about 10 nucleotides to about 22 nucleotides, about 11 nucleotides to about 21 nucleotides, about 12 nucleotides to about 20 nucleotides, about 13 nucleotides to about 19 nucleotides, about 14 nucleotides to about 18 nucleotides, or about 15 nucleotides to about 17 nucleotides.
[0084] In some cases, sample barcodes may be added to one or more ends of nucleic acid molecules of the present disclosure. A sample barcode may be added to a 5’ end, a 3’ end or a combination thereof of a nucleic acid molecule (e.g. a cell-free DNA molecule). The sample barcode may be single-stranded. The sample barcode may be double stranded. In some cases, the sample barcode may be part of a Y-shaped adaptor. The Y-shaped adaptor may be added to one or both ends (e.g. a 5’ end and / or a 3’ end) of a nucleic acid molecule (e.g. a cell-free DNA molecule). The sample barcode may comprise nucleic acid. In some cases, the sample barcode may comprise DNA, RNA, or a combination thereof. The sample barcode may have a variety of lengths. The sample barcode may have a length of at least about 2 nucleotides, at least about 3 nucleotides, at least about 4 nucleotides, at least about 5 nucleotides, at least about 6 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 11 nucleotides, at least about 12 nucleotides, at least about 13 nucleotides, at least about 14 nucleotides, at least about 15 nucleotides, at least about 16 nucleotides, at least about 17 nucleotides, at least about 18 nucleotides, at least about 19 nucleotides, at least about 20 nucleotides, at least about 21 nucleotides, at least about 22 nucleotides, at least about 23 nucleotides, at least about 24 nucleotides, at least about 25 nucleotides, at least about 26 nucleotides, at least about 27 nucleotides, at least about 28 nucleotides, at least about 29 nucleotides, at least about 30 nucleotides, at most about 2 nucleotides, at most about 3 nucleotides, at most about 4 nucleotides, at most about 5 nucleotides, at most about 6 nucleotides, at most about 7 nucleotides, at most about 8 nucleotides, at most about 9 nucleotides, at most about 10 nucleotides, at most about 11 nucleotides, at most about 12 nucleotides, at most about 13 nucleotides, at most about 14 nucleotides, at most about 15 nucleotides, at most about 16 nucleotides, at most about 17 nucleotides, at most about 18 nucleotides, at most about 19 nucleotides, at most about 20 nucleotides, at most about 21 nucleotides, at most about 22 nucleotides, at most about 23Attorney Docket No. 58626-723601 nucleotides, at most about 24 nucleotides, at most about 25 nucleotides, at most about 26 nucleotides, at most about 27 nucleotides, at most about 28 nucleotides, at most about 29 nucleotides, at most about 30 nucleotides, about 2 nucleotides to about 30 nucleotides, about 3 nucleotides to about 29 nucleotides, about 4 nucleotides to about 28 nucleotides, about 5 nucleotides to about 27 nucleotides, about 6 nucleotides to about 26 nucleotides, about 7 nucleotides to about 25 nucleotides, about 8 nucleotides to about 24 nucleotides, about 9 nucleotides to about 23 nucleotides, about 10 nucleotides to about 22 nucleotides, about 11 nucleotides to about 21 nucleotides, about 12 nucleotides to about 20 nucleotides, about 13 nucleotides to about 19 nucleotides, about 14 nucleotides to about 18 nucleotides, or about 15 nucleotides to about 17 nucleotides.
[0085] In some aspects, the methods described herein may not comprise use of molecular barcodes. The methods of the present application may not comprise adding (e.g. tagging) nucleic acid molecules (e.g. cell-free DNA molecules) with molecular barcodes to provide information regarding a source (e.g. origin) of the nucleic acid molecules. The methods described herein, may comprise use of methods other than use of molecular barcodes to determine a source of a nucleic acid molecule. In some cases, information about the nucleic acid molecule may be used to determine a source (e.g. origin) of the nucleic acid molecule. For example, a start position of a sequence of a nucleic acid molecule, stop position of a sequence of a nucleic acid molecule, length of a sequence of a nucleic acid molecule, a presence and / or an absence of SNVs and / or PVs within a sequence of a nucleic acid molecule, or a combination thereof may be used to determine a source of the nucleic acid molecule. In some cases, a source of a nucleic acid molecule after amplification may be determined based on a start position of a sequence read of the nucleic acid molecule after amplification and a stop position of the sequence read of the nucleic acid molecule after amplification. This information about the start position and stop position may be used to determine the source of the molecule from which the nucleic acid molecule was amplified from. In some cases, presence of a SNV may be used in combination with start position and stop position information to determine a molecular origin of a nucleic acid molecule.C) Methods for determining molecular origin using a combination of start and stop positions and low error rate somatic mutations
[0086] In some embodiments, methods described herein may comprise use of (1) information about nucleic acid molecules (e.g. cell-free DNA molecules) to determine a molecular origin of amplified nucleic acid molecules (e.g. as described above in (A)), (2) preferential use of low error rate somatic mutations and selection of molecules with lower error profiles in determining aAttorney Docket No. 58626-723601 mutation status of nucleic acid molecules (e.g., as described above in (B)) or (3) a combination thereof. As described herein, information about nucleic acid molecules (e.g. a start position, a stop position, one or more SNVs, one or more PVS, or any combination thereof) may be used to determine a molecular origin of amplified nucleic acid molecules. For example, cell-free nucleic acid molecules may be amplified to generate amplified cell-free nucleic acid molecules. The start and stop positions of sequence reads of the amplified nucleic acid molecules may be used to group the amplified cell-free nucleic acid molecules to generate groups of sequence reads. The groups of sequence reads may be further grouped based on the presence and / or absence of SNVs and or PVs to generate subgroups. Consensus sequences of sequences of the subgroups may be generated. The consensus sequences may represent a sequence of a nucleic acid molecule prior to amplification. The methods described herein may comprise analyzing the consensus sequences to determine whether the consensus sequence comprises a mutation (e.g. to determine whether the nucleic acid molecule of the consensus sequence is mutation containing).Determining whether the consensus sequence comprises a mutation may comprise analyzing the consensus sequence for the presence of a somatic mutation. The somatic mutation may comprise a base-change. The somatic mutation comprising a base-change may comprise an adenine-to- thymine (A>T) mutation, an adenine-to-cytosine (A>C) mutation, an adenine-to-guanine (A>G) mutation, a thymine-to-adenine (T>A) mutation, a thymine-to-cytosine (T>C) mutation, a thymine-to-guanine (T>G) mutation, a cytosine-to-adenine (C>A) mutation, a cytosine-to- thymine (C>T) mutation, a cytosine-to-guanine (OG) mutation, a guanine-to-adenine (G>A) mutation, a guanine-to-thymine (G>T) mutation, or a guanine-to-cytosine (G>C) mutation. In some cases, the consensus sequence may comprise one or more somatic mutations. Determining whether the consensus sequence comprises a mutation may comprise an error rate, which may incorrectly identify an artefactual mutation. The error rate may refer to a frequency of incorrectly identifying a variant within the consensus sequence. In some cases, the error rate (e.g. expected or desired error rate) can be determined and / or a desired error rate can be pre-specified by selectively analyzing nucleic acid molecules that meet one or more factors such as the specific base-change mutation (e.g., A>T, A>C, A>G, T>A, T>C, T>G, OA, OT, OG, G>A, G>T, G>C), the number of copies of sequence reads used to generate a consensus sequence cell-free DNA, the number of molecules observed with the same somatic mutation, the number of independent copies of Watson and Crick strands observed carrying the same somatic mutation, and the location of the somatic mutation among the sequence reads.
[0087] In another aspect, the present disclosure provides a method. The method may comprise obtaining a plurality of sequence reads for each of a plurality of nucleic acid molecules. TheAttorney Docket No. 58626-723601 plurality of nucleic acid molecules may be from a biological sample of a subject. The method may comprise mapping the plurality of sequence reads for each of the plurality of nucleic acid molecules to a reference genomic sequence. The method may comprise grouping the plurality of sequence reads into groups of sequence reads that map to identical start and stop positions on the reference genomic sequence. Grouping the plurality of sequence reads into groups of sequence reads that map to identical start and stop positions on the reference genomic sequence may thereby form a plurality of families. The method may comprise further grouping sequence reads of each family into a plurality of subgroups. The subgroups may comprise a first subgroup and a second subgroup. Each subgroup may comprise a plurality of sequence reads. The sequence reads of the first subgroup may comprise a first nucleotide at a first position. The sequence reads of the second subgroup may comprise a second nucleotide at the first position, wherein the second nucleotide differs from the first nucleotide. The method may comprise collapsing the sequence reads of the first subgroup to form a first consensus sequence corresponding with the sequence of a first nucleic acid molecule of the plurality of nucleic acid molecules. The method may comprise collapsing the sequence reads of the second subgroup to form a second consensus sequence corresponding with the sequence for a second nucleic acid molecule from the plurality of nucleic acid molecules, wherein the first nucleic acid molecule differs in sequence from the second nucleic acid molecule. The method may comprise identifying the first nucleotide at the first position as a somatic mutation. The somatic mutation may comprise an adenine-to-thymine (A>T) mutation, an adenine-to-cytosine (A>C) mutation, an adenine-to-guanine (A>G) mutation, a thymine-to-adenine (T>A) mutation, a thymine-to-cytosine (T>C) mutation, a thymine-to-guanine (T>G) mutation, a cytosine-to-adenine (C>A) mutation, a cytosine-to- thymine (C>T) mutation, a cytosine-to-guanine (OG) mutation, a guanine-to-adenine (G>A) mutation, a guanine-to-thymine (G>T) mutation, or a guanine-to-cytosine (G>C) mutation. Identifying the first nucleotide at the first position as a somatic mutation may comprise an error rate of no more than 0.5 parts per million.Biological samples
[0088] The methods described herein (e.g., methods (A), (B), or (C) as outlined above) may comprise obtaining a biological sample from a subject and / or obtaining sequence reads originating from a biological sample. The biological sample may comprise a body fluid sample, a tissue sample, an exogenous sample, or a combination thereof. The body fluid sample may comprise a urine sample, a saliva sample, a blood sample, a plasma sample, a whole blood sample, a vaginal secretion sample, a semen sample, a serum sample, a cerebrospinal fluid sample, a bone marrow sample, or a combination thereof. The tissue sample may comprise aAttorney Docket No. 58626-723601 surgical resection, a needle-core biopsy, a tissue slice, or a combination thereof. The exogenous sample may comprise a tissue culture sample, a cell culture sample, or a combination thereof. The blood sample may comprise plasma, buffy coat, erythrocytes, cfDNA, cell-free RNA (cfRNA), or a combination thereof. In some embodiments, the biological sample comprises a liquid biopsy. In some embodiments, the biological sample comprises a whole blood sample. The whole blood sample may comprise plasma, buffy coat, erythrocytes, cell-free DNA, cell-free RNA (cfRNA), or a combination thereof. In some embodiments, the biological sample comprises plasma, serum, blood, cerebrospinal fluid, lymph fluid, saliva, urine, or stool.
[0089] The subject of the biological sample may be a human subject. The human subject may be suffering from a condition, including but not limited to a cancer, heart disease, a neurological disorder (e.g., Alzheimer’s Disease), diabetes, an infectious disease (e.g., a bacterial infection), kidney disease, or a combination thereof. In some embodiments, the subject may have been diagnosed with cancer. The cancer may comprise acute lymphoblastic leukemia (all), acute myeloid leukemia (ami), adrenocortical carcinoma, astrocytomas, childhood (brain cancer), atypical teratoid / rhabdoid tumor, childhood (brain cancer), atypical teratoid / rhabdoid tumor, childhood, central nervous system (brain cancer), basal cell carcinoma of the skin , bile duct cancer, bladder cancer, bone cancer (includes ewing sarcoma and osteosarcoma and malignant fibrous histiocytoma), brain tumors, breast cancer, bronchial tumors (lung cancer), burkitt lymphoma , carcinoma of unknown primary, cervical cancer, childhood cancers, childhood cardiac tumors treatment, cholangiocarcinoma, chordoma, childhood (bone cancer), chronic lymphocytic leukemia (ell), chronic myelogenous leukemia (cml), chronic myeloproliferative neoplasms, colorectal cancer, craniopharyngioma, childhood (brain cancer), cutaneous t-cell lymphoma , diffuse intrinsic pontine glioma (dipg) (brain cancer), ductal carcinoma in situ (dcis) , embryonal tumors, medulloblastoma and other central nervous system, childhood (brain cancer), endometrial cancer (uterine cancer), ependymoma, childhood (brain cancer), esophageal cancer, esthesioneuroblastoma (head and neck cancer), ewing sarcoma (bone cancer), extracranial germ cell tumor, childhood, extragonadal germ cell tumor, eye cancer, fallopian tube cancer, gallbladder cancer, gastric (stomach) cancer, gastrointestinal neuroendocrine tumors, gastrointestinal stromal tumors (gist) (soft tissue sarcoma, germ cell tumor, childhood (brain cancer), gestational trophoblastic disease, hairy cell leukemia, head and neck cancer, heart tumors, childhood, hepatocellular (liver) cancer, histiocytosis, langerhans cell, hodgkin lymphoma, hypopharyngeal cancer (head and neck cancer), intraocular melanoma, intraocular melanoma, islet cell tumors, pancreatic neuroendocrine tumors, kaposi sarcoma (soft tissue sarcoma), kidney (renal cell) cancer, langerhans cell histiocytosis, laryngeal cancer (head andAttorney Docket No. 58626-723601 neck cancer), leukemia, lip and oral cavity cancer (head and neck cancer), liver cancer, lung cancer (non-small cell, small cell, pleuropulmonary blastoma, pulmonary inflammatory myofibroblastic tumor, and tracheobronchial tumor), lymphoma, male breast cancer, medulloblastoma and other cns embryonal tumors, childhood (brain cancer), melanoma, melanoma, intraocular (eye), merkel cell carcinoma (skin cancer), mesothelioma, malignant, metastatic cancer, metastatic squamous neck cancer with occult primary (head and neck cancer), midline tract carcinoma with nut gene changes, mouth cancer (head and neck cancer), multiple endocrine neoplasia syndromes, multiple myeloma / plasma cell neoplasms, mycosis fungoides (lymphoma), myelodysplastic syndromes, myelodysplastic / myeloproliferative neoplasms, myelogenous leukemia, chronic (cml), myeloid leukemia, myeloproliferative neoplasms, chronic, nasal cavity and paranasal sinus cancer (head and neck cancer), nasopharyngeal cancer (head and neck cancer), neuroblastoma, neuroendocrine tumors (gastrointestinal), non-hodgkin lymphoma, non-small cell lung cancer, oral cancer, lip and oral cavity cancer and oropharyngeal cancer (head and neck cancer), ovarian cancer, pancreatic cancer, pancreatic neuroendocrine tumors (islet cell tumors), papillomatosis (childhood laryngeal), paraganglioma, paranasal sinus and nasal cavity cancer (head and neck cancer), parathyroid cancer, penile cancer, pharyngeal cancer (head and neck cancer), pheochromocytoma, pituitary tumor, plasma cell neoplasm / multiple myeloma, pleuropulmonary blastoma (lung cancer), pregnancy and breast cancer, primary central nervous system (cns) lymphoma, primary cns lymphoma, primary peritoneal cancer, prostate cancer, pulmonary inflammatory myofibroblastic tumor (lung cancer), rectal cancer, recurrent cancer, renal cell (kidney) cancer, retinoblastoma, retinoblastoma, rhabdomyosarcoma, childhood (soft tissue sarcoma), salivary gland cancer(head and neck cancer), sezary syndrome (lymphoma), skin cancer, small cell lung cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma of the skin - see skin cancer, squamous neck cancer with occult primary, metastatic (head and neck cancer), stomach (gastric) cancer, t- cell lymphoma, cutaneous - see lymphoma (mycosis fungoides and sezary syndrome), testicular cancer, thymoma and thymic carcinoma, thyroid cancer, tracheobronchial tumors (lung cancer), transitional cell cancer of the renal pelvis and ureter (kidney (renal cell) cancer), ureter and renal pelvis, transitional cell cancer (kidney (renal cell) cancer), urethral cancer, uterine cancer, endometrial, uterine sarcoma, vaginal cancer, vascular tumors (soft tissue sarcoma), vulvar cancer, or a combination thereof. In some embodiments, the cancer is a blood cancer, a solid tumor, a B cell malignancy, leukemia, or lymphoma. In some embodiments, the cancer is large B-cell lymphoma (LBCL), Diffuse Large B-Cell Lymphoma (DLBCL), acute myeloid (or myelogenous) leukemia (AML), chronic myeloid (or myelogenous) leukemia (CML), acuteAttomey Docket No. 58626-723601 lymphocytic (or lymphoblastic) leukemia (ALL), adult ALL, chronic lymphocytic leukemia (CLL), hairy cell leukemia (HCL), small lymphocytic lymphoma (SLL), Mantle cell lymphoma (MCL), Marginal zone lymphoma, Burkitt lymphoma, Hodgkin lymphoma (HL), non-Hodgkin lymphoma (NHL), Anaplastic large cell lymphoma (ALCL), follicular lymphoma, refractory follicular lymphoma, or multiple myeloma (MM). For example, the biological sample may be extracted from a human subject suffering from lymphoma. The human subject may not be suffering from a condition. The subject may not be a human subject (e.g., a non-human subject). The non-human subject may comprise a non-human primate, a mouse, a rat, a rabbit, a fruit fly, or a combination thereof.
[0090] In some embodiments, the subject may have been administered a treatment for cancer before analyzing nucleic acid molecules of the biological sample described herein. The treatment for cancer may comprise any one of the treatments described herein.
[0091] The biological sample may comprise a plurality of nucleic acid molecules (e.g. cell-free DNA molecules). The plurality of nucleic acid molecules (e.g. cell-free DNA molecules) may be sequenced to generate sequence reads using any one of the methods described herein.
[0092] The plurality of nucleic acid molecules (e.g. cell-free DNA molecules) may comprise at least about 1,000, at least about 2,000, at least about 3,000, at least about 4,000, at least about 5,000, at least about 6,000, at least about 7,000, at least about 8,000, at least about 9,000, at least about 10,000, at least about 20,000, at least about 30,000, at least about 40,000, at least about 50,000, at least about 60,000, at least about 70,000, at least about 80,000, at least about 90,000, at least about 100,000, at least about 200,000, at least about 300,000, at least about 400,000, at least about 500,000, at least about 600,000, at least about 700,000, at least about 800,000, at least about 900,000, at least about 1,000,000, at least about 10,000,000, at least about 50,000,000, at least about 100,000,000, at least about 500,000,000, at least about 1,000,000,000, or more nucleic acid molecules (e.g. cell-free DNA molecules). The plurality of nucleic acid molecules (e.g. cell-free DNA molecules) may comprise at most about 1,000, at most about 2,000, at most about 3,000, at most about 4,000, at most about 5,000, at most about 6,000, at most about 7,000, at most about 8,000, at most about 9,000, at most about 10,000, at most about 20,000, at most about 30,000, at most about 40,000, at most about 50,000, at most about 60,000, at most about 70,000, at most about 80,000, at most about 90,000, at most about 100,000, at most about 200,000, at most about 300,000, at most about 400,000, at most about 500,000, at most about 600,000, at most about 700,000, at most about 800,000, at most about 900,000, at most about 1,000,000, at most about 10,000,000, at most about 50,000,000, at most about 100,000,000, at most about 500,000,000, at most about 1,000,000,000, or less nucleic acid molecules (e.g. cell-Attorney Docket No. 58626-723601 free DNA molecules). The plurality of nucleic acid molecules (e.g. cell-free DNA molecules) may comprise about 100,000 to about 1,000,000,000, about 200,000 to about 900,000,000, about 300,000 to about 80,0000,000, about 400,000 to about 70,0000,000, about 500000 to about 600,000,000, about 600,000 to about 500,000,000, about 700,000 to about 400,000,000, about 800,000 to about 300,000,000, about 900,000 to about 200,000,000, about 1,000,000 to about 100,000,000, or about 5,000,000 to about 50,000,000 nucleic acid molecules (e.g. cell-free DNA molecules).Nucleic acid sequencing
[0093] The methods described herein (e.g., methods (A), (B), or (C) as outlined above) may comprise extracting nucleic acids from a biological sample. The nucleic acids may comprise DNA, RNA, or a combination thereof. In some cases, the nucleic acids may comprise cell-free DNA molecules. Extracting nucleic acids from the biological sample may comprise processing the biological sample to remove cells and / or cellular debris. Processing the sample to remove cells and / or cellular debris may comprise magnetic bead-based extraction, silica membranebased extraction, liquid-phase extraction, or a combination thereof. In some cases, extracting nucleic acid from a biological sample may comprise centrifuging the biological sample. The biological sample may separate into more than one layers (e.g. two layers) after the centrifuging. A layer of the more than one layers of the centrifuged biological sample may be isolated from the other layers of the centrifuged biological sample to extract nucleic acid from the biological sample. In some cases, plasma of the biological sample may be lysed to extract nucleic acid of the biological sample. The plasma of the biological sample may be lysed using chemical detergents, sonication, freeze-thaw treatment, enzymes (e.g. lysozymes), or a combination thereof. In some cases, a magnetic particle may be added to the biological sample. The magnetic particle may bind to nucleic acid (e.g. cell-free DNA molecules). Solution surrounding the magnetic particle may be removed and replaced with a buffer solution. In some cases, the buffer solution may be replaced one or more times. The nucleic acids (e.g. cell free DNA molecules) may be eluted from the magnetic particle using a solvent. For example, the magnetic particle may be soaked in ethanol (e.g. 75% ethanol). The nucleic acids may be eluted from the magnetic particle using a salt solution (e.g. a high salt solution). The extracted nucleic acids (e.g. cell-free DNA molecules) may be further treated.
[0094] Extracted nucleic acids (e.g. cell-free DNA molecules) may be treated to repair ends of the nucleic acids. For example, extracted cell-free DNA molecules may be mixed with a reaction mixture and incubated for a period of time. The reaction mixture may comprise one or more enzymes, nucleotides, one or more buffers, one or more salts, one or more detergents, or aAttorney Docket No. 58626-723601 combination thereof. In some cases, the one or more enzymes may comprise a polymerase, a polymerase fragment, a polynucleotide kinase, a polynucleotide fragment, or a combination thereof. The polymerase and / or polymerase fragment may comprise 5 ’to 3’ polymerase activity, 3’ to 5’ polymerase activity, or a combination thereof. The polynucleotide kinase and / or polynucleotide fragment may comprise 3’ phosphatase activity, 5’ kinase activity, or a combination thereof. In some cases the one or more enzymes may comprise a T4 polymerase, T7 polymerase, Klenow fragment, a T4 polynucleotide kinase, or a combination thereof. Treating the extracted nucleic acid to repair ends may produce nucleic acid with blunt ends. For example, extracted cell-free DNA molecules may be treated with a reaction mixture comprising a polymerase and a polynucleotide kinase to produce cell-free DNA molecules comprising blunt ends (e.g. ends without overhangs). The blunt ends of the end repair treated extracted nucleic acids may comprise phosphates (e.g. 5’ phosphates).
[0095] Extracted nucleic acids of the biological sample may be ligated to adaptors. In some cases, the adaptors may comprise a molecular barcode (e.g. UMIs). In some cases, the adaptors may not comprise a molecular barcode. The adaptors may comprise one or more primers. In some cases, the one or more primers may comprise sequencing primers, amplification primers, or a combination thereof. For example, the one or more adaptors may comprise primers compatible with a sequencing technology (e.g. Illumina sequencing primers). In some cases, the adaptors may comprise amplification primers for amplifying the extracted nucleic acid of the biological sample. In some cases, the adaptors may comprise sample barcodes. The sample barcodes may be added to extracted nucleic acid of a biological sample to differentiate nucleic acid from the biological sample from nucleic acid from a different biological sample. In some cases, the sample barcodes may be added to extracted nucleic acid of a biological sample to differentiate nucleic acid from the biological sample from other nucleic acid from the biological sample. For example, a first portion of nucleic acids from the biological sample may be treated with a first treatment. A second portion of the nucleic acids from the biological sample may be treated with a second treatment. The first portion of nucleic acids may be ligated to adaptors comprising a first sample barcode. The second portion of nucleic acids may be ligated to adaptors comprising a second sample barcode. The first sample barcode and the second sample barcode may be used to identify the source of nucleic acid sequences after sequencing. Ligating adaptors to extracted nucleic acids may comprise use of a ligase. In some cases, more than one ligation reaction may be performed to ligate adaptors and / or primers to the extracted nucleic acid. Primers may be added to the adaptors using amplification. For example, adaptors may be ligated to the extracted nucleic acid (e.g. cell-free DNA molecules). Primers that bind to theAttorney Docket No. 58626-723601 adaptors may be added to the extracted nucleic acid with ligated adaptors under conditions to promote binding of the primers to the adapters. An amplification reaction (e.g. polymerase chain reaction (PCR)) may be performed. The amplification reaction may produce nucleic acid molecules comprising the extracted nucleic acid or derivative thereof, the adapters or derivative thereof, the primers or derivative thereof, or a combination thereof. In some cases, the extracted nucleic acid may be end repaired, as described herein. The derivative of a nucleic acid may comprise a copy of the nucleic acid. The copy of the nucleic acid may comprise at least a portion of the same sequence of the nucleic acid. The derivative of a nucleic acid may comprise a revere complement of the nucleic acid. The derivative of a nucleic acid may comprise a complement of the nucleic acid.
[0096] The methods described herein may comprise amplifying nucleic acids before sequencing. In some cases, nucleic acids (e.g. cell-free DNA molecules) may be extracted from a biological sample, ligated with adaptors, and PCR amplified. Amplifying nucleic acids before sequencing may be useful for generating more nucleic acid material for sequencing. For example, cell-free DNA molecules may be extracted from a biological sample from a subject. The extracted cell- free DNA molecules may be amplified to produce additional copies of the cell-free DNA molecules for sequence analysis. Amplifying nucleic acids of the methods described herein may comprise increasing the input nucleic acid by at least about 2X, at least about 5X, at least about 10X, at least about 50X, at least about 100X, at least about 500X, at least about l,000X, at least about 5,000X, at least about 10,000X, at least about 100,000X, at least about 500,000X, at least about l,000,000X, or more. Amplifying nucleic acids of the methods described herein may comprise increasing the input nucleic acid by at most about 2X, at most about 5X, at most about 10X, at most about 50X, at most about 100X, at most about 500X, at most about l,000X, at most about 5,000X, at most about 10,000X, at most about 100,000X, at most about 500,000X, at most about l,000,000X, or less. Amplified nucleic acids produced from amplifying the nucleic acid may be purified. Purifying the amplified nucleic acids may comprise removing enzymes, primers, salts, detergents, or a combination thereof from the amplified nucleic acids. In some cases, purifying the amplified nucleic acids may comprise contacting the amplified nucleic acids with a substrate. In some cases, the substrate may comprise particles and / or beads. The particles and / or beads may be magnetic. The substrate may be used to perform solid-phase reversible immobilization of the amplified nucleic acids.
[0097] The methods described herein may comprise enriching nucleic acids (e.g. cell-free DNA molecules). In some cases, enriching nucleic acids may comprise contacting the nucleic acid with a plurality of probes. The plurality of probes may bind to sequences of the nucleic acids.Attorney Docket No. 58626-723601The plurality of probes may be associated with particles and / or beads (e.g. magnetic particles). The particles and / or beads comprising the plurality of probes bound to sequences of the nucleic acids (e.g. cell-free DNA molecules) may be isolated. The particles and / or beads comprising the plurality of probes bound to sequences of the nucleic acid (e.g. cell-free DNA molecules) may washed one or more times with a buffer. For example, a nucleic acid sample extracted form a biological sample may be contacted with a bead sample. The bead sample may comprise a plurality of probes comprising oligonucleotides attached thereto. The probes comprising oligonucleotides may comprise sequences associated with one or more SNVs and / or PVs. For example, the sequences associated with one or more SNVs and / or PVs may comprise sequences comprising and / or adjacent to sequences of a genome known to be associated with SNVs and / or PVs. In some cases, the sequences of a genome known to be associated with SNVs and / or PVs may be associated with a condition, as described herein (e.g. a cancer). In some cases, the sequences associated with one or more SNVs and / or PVs may comprise sequences comprising and / or adjacent to sequences of a subject known to be associated with SNVs and / or PVs of the subject. For example, a DNA sample from a subject may be analyzed to determine a patientspecific set of SNVs and / or PVs for the subject. The probes comprising oligonucleotides may comprise sequences associated with the patient-specific set of SNVs and / or PVs for the subject. The sequences associated with the patient-specific set of SNVs and / or PVs for the subject may comprise sequences comprising the SNVs and / or PVs for the subject and / or may comprise sequences adjacent to the SNVs and / or PVs for the subject. The bead sample comprising a plurality of probes comprising oligonucleotides attached thereto may bind to nucleic acid molecules (e.g. cell-free DNA molecules) of the nucleic acid sample extracted from the biological sample. The beads of the bead sample may be isolated from solution (e.g. using a magnet). The nucleic acid molecules (e.g. cell-free DNA molecules) of the nucleic acid sample extracted from the biological sample bound to the bead sample may be eluted (e.g. using heat, salt, solvent, or a combination thereof). The eluted enriched nucleic acid molecules (e.g. cell-free DNA molecules) may be sequenced using any one of the methods escribed herein.
[0098] The methods described herein may comprise sequencing nucleic acid (e.g. cell-free DNA molecules). In some cases, the sequence reads used for determining a sequence of a nucleic acid molecule (e.g., a first nucleic acid molecule) may be obtained by performing sequencing. In some cases, the methods may comprise sequencing the plurality of cell-free DNA molecules can produce a plurality of sequence reads. The sequencing of the methods described herein may be performed on one or more biological samples comprising nucleic acids. For example, sequencing may be performed on a cell-free DNA (cfDNA) sample extracted from a human subject. TheAttorney Docket No. 58626-723601 sequencing may comprise whole genome sequencing, targeted sequencing, whole exome sequencing, hybridization capture sequencing, amplicon sequencing, molecular inversion probe enrichment sequencing, whole transcriptome sequencing, targeted gene expression sequencing, sanger sequencing, next-generation sequencing, capillary electrophoresis sequencing, shotgun sequencing, or a combination thereof. In some cases, the sequencing may be performed on a sequencing platform, including but not limited to an Illumina HiSeq 2500 sequencer, an Illumina X10 sequencer, an Illumina MiSeq sequencer, an Illumina NovaSeq sequencers, an Infmium EPIC Array, an Infmium FSA Array, an Infmium Omni Express sequencer, an Oxford Nanopore PromethlON sequencer, a PacBio Sequel I sequencer, a PacBio Sequel II sequencer, an AVITI™ sequencer, an UGlOO™ sequencer, a DNBSEQ-E25 sequencer, a DNBSEQ-G99 sequencer, a DNBSEQ-G400 sequencer, a DNBSEQ-G800 sequencer, a DNBSEQ-T7 sequencer, a DNBSEQ-T20x2, a 454 sequencer, an lonTorrent sequencer, or a combination thereof. Sequencing the nucleic acid molecules (e.g. a plurality of cell-free DNA molecules) may comprise whole genome sequencing, whole exome sequencing, or a combination thereof.
[0099] Performing sequencing of nucleic acids from one or more biological samples may generate sequence reads. The sequence reads generated by performing sequencing may comprise information related to the nucleic acid sequences of the nucleic acid molecules within the one or more biological samples. For example, a biological sample comprising cfDNA molecules may be sequenced and the sequence reads generated may include the DNA sequence of each of the cfDNA molecules analyzed. In some cases, multiple sequence reads are generated from each analyzed cfDNA molecule. The sequence reads may be listed in a variety of formats. For example, the sequence reads may be in one or more tables, one or more arrays, one or more files, one or more lists, or a combination thereof. The one or more files comprising sequence reads may comprise a variety of formats including but not limited to a BAM file, a CRAM file, a SFF file, a HDF5 file, a FASTQ file, a CSFASTA file, a QU AL file, a TXT file, or a combination there of. For example, sequence reads from sequencing a cfDNA sample from a human subject with lung cancer may be listed in a FASTQ file. The one or more files comprising sequence reads may comprise information related to other features of the sequencing, including information related to the sequencing platform, sequencing run information, experimental details (e.g., sample type), analysis parameters, or a combination thereof. In some cases, the sequence reads reflect raw sequence information (e.g., sequence information generated directly from the sequencing platform), processed sequence information (e.g., sequence information that has been filtered and / or manipulated, mapped to a reference genome, preprocessed to remove optical duplicates,), or a combination thereof.Attorney Docket No. 58626-723601
[0100] The plurality of sequence reads obtained as described herein may comprise multiple sequence reads. The plurality of sequence reads may comprise at least about 1,000, at least about 2,000, at least about 3,000, at least about 4,000, at least about 5,000, at least about 6,000, at least about 7,000, at least about 8,000, at least about 9,000, at least about 10,000, at least about 20,000, at least about 30,000, at least about 40,000, at least about 50,000, at least about 60,000, at least about 70,000, at least about 80,000, at least about 90,000, at least about 100,000, at least about 200,000, at least about 300,000, at least about 400,000, at least about 500,000, at least about 600,000, at least about 700,000, at least about 800,000, at least about 900,000, at least about 1,000,000, at least about 10,000,000, at least about 50,000,000, at least about 100,000,000, at least about 500,000,000, at least about 1,000,000,000, or more sequence reads. The plurality of sequence reads may comprise at most about 1,000, at most about 2,000, at most about 3,000, at most about 4,000, at most about 5,000, at most about 6,000, at most about 7,000, at most about 8,000, at most about 9,000, at most about 10,000, at most about 20,000, at most about 30,000, at most about 40,000, at most about 50,000, at most about 60,000, at most about 70,000, at most about 80,000, at most about 90,000, at most about 100,000, at most about 200,000, at most about 300,000, at most about 400,000, at most about 500,000, at most about 600,000, at most about 700,000, at most about 800,000, at most about 900,000, at most about 1,000,000, at most about 10,000,000, at most about 50,000,000, at most about 100,000,000, at most about 500,000,000, at most about 1,000,000,000, or less sequence reads. The plurality of sequence reads may comprise about 100,000 to about 1,000,000,000, about 200,000 to about 900,000,000, about 300,000 to about 80,0000,000, about 400,000 to about 70,0000,000, about 500000 to about 600,000,000, about 600,000 to about 500,000,000, about 700,000 to about 400,000,000, about 800,000 to about 300,000,000, about 900,000 to about 200,000,000, about 1,000,000 to about 100,000,000, or about 5,000,000 to about 50,000,000 sequence readsSequence alignment
[0101] The methods described herein may comprise aligning (e.g. mapping) the plurality of sequence reads to one or more reference genomic sequences. In some cases, aligning (e.g. mapping) the plurality of sequence reads to one or more reference genomic sequences may produce a plurality of aligned sequence reads. Aligning (e.g. mapping) the plurality of sequence reads may comprise comparing a sequence read of the plurality of sequence reads to one or more reference genomic sequences. Comparing a sequence read of the plurality of sequence reads to one or more reference genomic sequences may comprise overlaying the sequence read of the plurality of sequence reads to at least a first portion of the one or more reference genomic sequences and determining whether the sequence of the sequence read of the plurality ofAttorney Docket No. 58626-723601 sequence reads is identical to or substantially similar to the first portion of the one or more reference genomic sequences. Substantially similar to may refer to a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or more sequence identity to a at least a portion of one or more reference genomics sequences.
[0102] In some embodiments, aligning or mapping the plurality of sequence reads to a reference genomic sequence may comprise aligning at least 1 thousand sequence reads, at least 10 thousand sequence reads, at least 100 thousand sequence reads, at least 1 million sequence reads, at least 10 million sequence reads, or at least 100 million sequence reads. In some embodiments, aligning the plurality of sequence reads to a reference genomic sequence further comprises aligning sequence reads representative of at least 1 thousand cell-free DNA molecules, at least 10 thousand cell-free DNA molecules, at least 100 thousand cell-free DNA molecules, at least 1 million cell-free DNA molecules, at least 10 million cell-free DNA molecules, or at least 100 million cell-free DNA molecules. In some embodiments, the reference genomic sequence is a human genome assembly, a non-human genome assembly, a mammalian genome assembly, or a portion thereof. The reference genomic sequence can be at least a portion of a nucleic acid sequence database (i.e., a reference genome), which database is assembled from genetic data and intended to represent the genome of a reference cohort. In some cases, a reference cohort can be a collection of individuals from a specific or varying genotype, haplotype, demographics, sex, nationality, age, ethnicity, relatives, physical condition (e.g., healthy or having been diagnosed to have the same or different condition, such as a specific type of cancer), or other groupings.
[0103] The one or more reference genomic sequences used for mapping the sequence reads of the plurality of sequence reads may comprise a human genome or a portion thereof. The human genome may comprise the NCBI Build 34 sequence, the NCBI Build 35 sequence, the NCBI Build 36.1 sequence, the GRCh37 sequence, the GRCh38 sequence, the T2T-CHM13 sequence, the GRCh39 sequence, the hgl6 sequence, the hgl7 sequence, the hgl8 sequence, the hgl9 sequence, the hsl sequence, or a combination thereof. The one or more reference genomic sequences may comprise a sequence derived from a sample of a human subject. The human subject of the sample may be a healthy subject. The human subject of the sample may be a subject suffering from a condition as described herein. The one or more reference genomic sequences may comprise a sequence derived from a sample of the human subject that is the source of the sequence reads of the plurality of sequence reads. The one or more reference genomic sequence may comprise a human exome. A reference genomic sequence of the one orAttorney Docket No. 58626-723601 more reference genomic sequences may have a length of at least about 10 kilobases (kb), at least about 20 kb, at least about 30 kb, at least about 40 kb, at least about 50 kb, at least about 60 kb, at least about 70 kb, at least about 80 kb, at least about 90 kb, at least about 100 kb, at least about 200 kb, at least about 300 kb, at least about 400 kb, at least about 500 kb, at least about 600 kb, at least about 700 kb, at least about 800 kb, at least about 900 kb, at least about 1,000 kb, at least about 2,000 kb, at least about 3,000 kb, at least about 4,000 kb, at least about 5,000 kb, at least about 6,000 kb, at least about 7,000 kb, at least about 8,000 kb, at least about 9,000 kb, at least about 10,000 kb, at least about 20,000 kb, at least about 30,000 kb, at least about 40,000 kb, at least about 50,000 kb, at least about 60,000 kb, at least about 70,000 kb, at least about 80,000 kb, at least about 90,000 kb, at least about 100,000 kb, at least about 200,000 kb, at least about 300,000 kb, at least about 400,000 kb, at least about 500,000 kb, at least about 600,000 kb, at least about 700,000 kb, at least about 800,000 kb, at least about 900,000 kb, at least about 1,000,000 kb, or more. A reference genomic sequence of the one or more reference genomic sequences may have a length of 10-4,000,000 kb, 10-1,000,000 kb, about 20-900,000 kb, about 30-800,000 kb, about 40-700,000 kb, about 50-600,000 kb, about 60-500,000 kb, about 70- 400,000 kb, about 80-300,000 kb, about 90-200,000 kb, about 100-100,000 kb, about 200-90,000 kb, about 300-80,000 kb, about 400-70,000 kb, about 500-60,000 kb, about 600-50,000 kb, about 700-40,000 kb, about 800-30,000 kb, about 900-20,000 kb, about 1000-10,000 kb, about 2,000- 9,000 kb, about 3,000-8,000 kb, about 4,000-7,000 kb, or about 5,000-6,000 kb.
[0104] In some embodiments, aligning the plurality of sequence reads to a reference genomic sequence further comprises performing Burrow-Wheeler alignment (BWA) algorithm, a Bowtie alignment algorithm, DRAGEN, minimap2, or a combination thereof. For example, the methods described herein for determining a molecular origin based on using a start and stop position may comprise aligning (e.g. mapping) a plurality of sequence reads for each of a plurality of nucleic acid molecules to a reference genomic sequence. Aligning (e.g. mapping) the plurality of sequence reads for each of the plurality of nucleic acid molecules to the reference genomic sequence may comprise performing Burrow- Wheeler alignment (BWA) algorithm, a Bowtie alignment algorithm, DRAGEN, minimap2, or a combination thereof. The methods described herein comprising use of low error rate somatic mutations may comprise performing Burrow- Wheeler alignment (BWA) algorithm, a Bowtie alignment algorithm, DRAGEN, minimap2, or a combination thereof. In some cases, the methods described herein for determining molecular origin using a combination of start and stop positions and low error rate somatic mutations may comprise performing Burrow- Wheeler alignment (BWA) algorithm, a Bowtie alignment algorithm, DRAGEN, minimap2, or a combination thereof.Attorney Docket No. 58626-723601Disease status
[0105] In some embodiments, the methods described herein may comprise determining a disease status (e.g., minimum residual disease status) of the subject. In some cases, determining a disease status (e.g., minimum residual disease status) of the subject is based at least in part on determining whether at least one of the set of filtered sequence reads is mutation-containing. In some embodiments, the set of filtered sequence reads contains one or more SNVs. In some embodiments, the one or more SNVs contained in the set of filtered sequence reads is associated with the disease. In some embodiments, determining the disease status is further based at least in part on determining whether one or more subsets of the plurality of aligned sequence reads contains one or more PVs. In some embodiments, the one or more PVs contained is associated with the disease. In some cases, determining a disease status (e.g., minimum residual disease status) of the subject is based at least in part on determining whether a consensus sequence formed using the methods described herein is mutation-containing. The consensus sequence may be formed using a start and stop position. The consensus sequence may comprise one or more SNVs, one or more PVs, or a combination thereof. The one or more SNVs, the one or more PVs, or a combination thereof may be associated with the disease. In some cases, determining a disease status (e.g., minimum residual disease status) of the subject is based at least in part on the sequence for the first nucleic acid molecule and the sequence for the second nucleic acid molecule using the methods described herein. For example, the first nucleic acid sequence and the second nucleic acid sequence of the methods described herein may be evaluated for the presence of one or more SNVs. The one or more SNVs may be associated with a disease state (e.g. a presence or absence of the one or more SNVs may indicate the subject has a disease).Based on the presence of the one or more SNVs, the subject may be determined to have a disease or minimum residual disease. In some cases, the first nucleic acid sequence and the second nucleic acid sequence of the methods described herein may be evaluated for the presence of one or more PVs. The one or more PVs may be associated with a disease state (e.g. a presence or absence of the one or more PVs may indicate the subject has a disease). Based on the presence of the one or more PVs, the subject may be determined to have a disease or minimum residual disease.
[0106] The methods described herein may comprise assessing a therapeutic response of the subject in response to a treatment for cancer administered to the subject. In some embodiments, assessing the therapeutic response of the subject in response to the treatment for cancer is based at least in part on determining whether at least one of the set of filtered sequence reads is mutati on-containi ng .Attorney Docket No. 58626-723601
[0107] In some embodiments, the method further comprises assessing a risk of cancer recurrence or relapse in the subject. In some embodiments, assessing a risk of cancer recurrence or relapse in the subject is based at least in part on determining whether at least one of the set of filtered sequence reads is mutation-containing.Treatment
[0108] In some embodiments, the subject of the biological sample of the methods described herein may have been administered a first treatment for cancer. In some embodiments, the methods may comprise treating the subject with a second treatment for cancer. In some embodiments, the methods may comprise treating the subject with a second treatment for cancer based at least in part on determining whether at least one of the set of filtered sequence reads is mutation-containing. In some embodiments, the methods may comprise treating the subject with a second treatment for cancer based at least in part on the sequence for the first nucleic acid molecule and the sequence for the second nucleic acid molecule using the methods described herein. For example, the first nucleic acid sequence and the second nucleic acid sequence of the methods described herein may be evaluated for the presence of one or more SNVs. The one or more SNVs may be associated with a disease state (e.g. a presence or absence of the one or more SNVs may indicate the subject has a disease). Based on the presence of the one or more SNVs, the subject may be treated with a second treatment for cancer. In some embodiments, the first treatment for cancer may include chemotherapy, radiotherapy, chemoradiotherapy, immunotherapy, adoptive cell therapy (e.g., chimeric antigen receptor (CAR) T cell therapy, CARNK cell therapy, modified T cell receptor (TCR) T cell therapy, etc.) hormone therapy, targeted drug therapy, surgery, transplant, transfusion, or medical surveillance. In some embodiments, the first treatment for cancer can include administering the subject with one or more therapeutic agents.
[0109] In some embodiments, the method described herein may comprise administering to the subject an effective amount of treatment (e.g. for cancer). The effective amount of the treatment (e.g. for cancer) may be a first treatment. In some embodiments, the methods described herein may comprise administering to the subject an effective amount of treatment for cancer based at least in part on determining whether at least one of the set of filtered sequence reads is mutationcontaining. In some embodiments, the methods described herein may comprise administering to the subject an effective amount of treatment for cancer based at least in part on the sequence for the first nucleic acid molecule and the sequence for the second nucleic acid molecule using the methods described herein. For example, the first nucleic acid sequence and the second nucleic acid sequence of the methods described herein may be evaluated for the presence of one or moreAttorney Docket No. 58626-723601SNVs. The one or more SNVs may be associated with a disease state (e.g. a presence or absence of the one or more SNVs may indicate the subject has a disease). Based on the presence of the one or more SNVs, the subject may be treated with an effective amount of treatment (e.g. for cancer).
[0110] In some embodiments, the method further comprises administering to the subject an effective amount of treatment for cancer based at least in part on determining whether at least the first nucleic acid sequence and / or the second nucleic acid sequence is mutation containing (e.g. PV or SNV containing). In some embodiments, the treatment for cancer can include chemotherapy, radiotherapy, chemoradiotherapy, immunotherapy, adoptive cell therapy (e.g., chimeric antigen receptor (CAR) T cell therapy, CARNK cell therapy, modified T cell receptor (TCR) T cell therapy, etc.) hormone therapy, targeted drug therapy, surgery, transplant, transfusion, or medical surveillance. In some embodiments, the treatment for cancer can include administering the subject with one or more therapeutic agents. The one or more therapeutic drugs can be administered to the subject by one or more of the following: orally, intraperitoneally, intravenously, intraarterially, transdermally, intramuscularly, liposomally, via local delivery by catheter or stent, subcutaneously, intraadiposally, and intrathecally. Non-limiting examples of the therapeutic drugs can include cytotoxic agents, chemotherapeutic agents, growth inhibitory agents, agents used in radiation therapy, anti-angiogenesis agents, apoptotic agents, anti-tubulin agents, and other agents to treat cancer.[oni] In some cases, the methods described herein may comprise administering one or more chemotherapeutic agents. The one or more chemotherapeutic agents may comprise one or more alkylating agents, one or more nitrosoureas, one or more antimetabolites, one or more anti-tumor antibiotics, one or more topoisomerase inhibitors, one or more mitotic inhibitors, one or more corticosteroids, or any combination thereof. In some cases, the methods described herein may comprise administering one or more of all-trans-retinoic acid, arsenic trioxide, asparaginase, eribulin, hydroxyurea, ixabepilone, mitotane, omacetaxine, pegaspargase, procarbazine, romidepsin, vorinostat, prednisone, methylprednisolone, dexamethasone, vinblastine, vincristine, vincristine liposomal, vinorelbine, cabazitaxel, docetaxel, nab-paclitaxel, paclitaxel, irinotecan, irinotecan liposomal, topotecan, etoposide, mitoxantrone, teniposide, irinotecan, irinotecan liposomal, topotecan, daunorubicin, doxorubicin, doxorubicin liposomal, epirubicin, idarubicin, valrubicin, bleomycin, dactinomycin, mitomycin-c, azacitidine, 5 -fluorouracil (5-fu), 6- mercaptopurine (6-mp), capecitabine, cladribine, clofarabine, cytarabine, decitabine, floxuridine, fludarabine, gemcitabine, hydroxyurea, methotrexate, nelarabine, pemetrexed, pentostatin, pralatrexate, thioguanine, carmustine, lomustine, streptozocin, bendamustine, busulfan,Attorney Docket No. 58626-723601 carboplatin, chlorambucil, cisplatin, cyclophosphamide, dacarbazine, ifosfamide, mechlorethamine, melphalan, oxaliplatin, temozolomide, thiotepa, trabectedin, or any combination thereof. In some cases, the methods describe herein may comprise administering to the subject rituximab, cyclophosphamide, doxorubicin, vincristine, prednisone, or any combination thereof. In some cases, the methods describe herein may comprise administering to the subject etoposide, prednisone, oncovin (vincristine), cyclophosphamide, hydroxydaunorubicin (doxorubicin), or any combination thereof.
[0112] In some embodiments, the method further comprises manufacturing a medicament for treating cancer in the subject. In some embodiments, the method may comprise manufacturing a medicament for treating cancer in the subject based at least in part on determining whether at least one of the set of filtered sequence reads is mutation-containing. In some embodiments, the method may comprise manufacturing a medicament for treating cancer in the subject based at least in part on whether at least the first nucleic acid sequence and / or the second nucleic acid sequence is mutation containing (e.g. PV or SNV containing).
[0113] In some embodiments, the cancer of the subject may be a blood cancer, a solid tumor, a B cell malignancy, leukemia, or lymphoma. In some embodiments, the cancer is large B-cell lymphoma (LBCL), Diffuse Large B-Cell Lymphoma (DLBCL), acute myeloid (or myelogenous) leukemia (AML), chronic myeloid (or myelogenous) leukemia (CML), acute lymphocytic (or lymphoblastic) leukemia (ALL), adult ALL, chronic lymphocytic leukemia (CLL), hairy cell leukemia (HCL), small lymphocytic lymphoma (SLL), Mantle cell lymphoma (MCL), Marginal zone lymphoma, Burkitt lymphoma, Hodgkin lymphoma (HL), non-Hodgkin lymphoma (NHL), Anaplastic large cell lymphoma (ALCL), follicular lymphoma, refractory follicular lymphoma, or multiple myeloma (MM). In some embodiments, the cancer is a pancreatic cancer, bladder cancer, colorectal cancer, breast cancer, prostate cancer, renal cancer, hepatocellular cancer, lung cancer, ovarian cancer, cervical cancer, rectal cancer, thyroid cancer, uterine cancer, gastric cancer, esophageal cancer, head and neck cancer, melanoma, neuroendocrine cancers, CNS cancers, brain tumors, bone cancer, or soft tissue sarcoma.
[0114] In some embodiments, the treatment or the medicament is a first-line therapy. In some embodiments, the treatment or the medicament is a second-line therapy. In some embodiments, the treatment comprises radiographic imaging, computed tomography (CT) imaging, positron emission tomography (PET) imaging, and / or magnetic resonance imaging (MRI) of the subject. In some embodiments, the treatment comprises further monitoring the subject.
[0115] In some embodiments, the treatment or the medicament comprises a cell therapy. In some embodiments, the treatment or the medicament comprises an immune cell therapy. In someAttorney Docket No. 58626-723601 embodiments, the treatment or the medicament comprises a T cell therapy. In some embodiments, the treatment or the medicament comprises a CAR-T cell therapy. In some embodiments, the CAR-T cell therapy is an anti-CD19 cell therapy or an anti-CD20 cell therapy. In some embodiments, the cell therapy comprises an adoptive therapy (e.g., an adoptive T cellbased therapy). Adoptive immunotherapy refers to a therapeutic approach for treating cancer or infectious diseases in which immune cells are administered to a host with the aim that the cells mediate either directly or indirectly specific immunity to (i.e., mount an immune response directed against) cancer cells. In some embodiments, the immune response results in inhibition of tumor and / or metastatic cell growth and / or proliferation, and in related embodiments, results in neoplastic cell death and / or resorption. The immune cells can be derived from a different organism / host (exogenous immune cells) or can be cells obtained from the subject organism (autologous immune cells). In some embodiments, the cell therapy comprises genetically engineered cells. In some embodiments, the cell therapy comprises genetically engineered immune cells. In some embodiments, the genetically engineered immune cells can be derived from autologous or allogeneic immune cells. In some embodiments, the genetically engineered cells are T cells (e.g., regulatory T cells, CD4+ T cells, CD8+ T cells, or gamma-delta T cells), NK cells, invariant NK cells, or NKT cells. In some embodiments, the genetically engineered cells can be genetically engineered to express a heterologous gene. In some embodiments, the genetically engineered cells can be genetically engineered to express antigen receptors such as engineered TCRs and / or chimeric antigen receptors (CARs). In some embodiments, the T cells are autologous or allogeneic to the subject. In some embodiments, the genetically engineered T cells are engineered to express a heterologous gene. In some embodiments, the genetically engineered T cells are engineered to express a receptor. In some embodiments, the receptor is a bispecific receptor. In some embodiments, the genetically engineered T cells are engineered to express a T cell receptor (TCR) having antigenic specificity for an antigen. In some embodiments, the genetically engineered T cells are engineered to express a T cell receptor (TCR) having antigenic specificity for more than one antigen. In some embodiments, the genetically engineered T cells comprise a Chimeric Antigen Receptor (CAR). In some embodiments, the genetically engineered T cells are engineered to express more than one TCRs and / or CARs. In some embodiments, the CAR specifically binds to an antigen. In some embodiments, the antigen is associated with a disease (e.g., cancer) or condition and / or is expressed by cells associated with a disease (e.g., cancer) or condition. In some embodiments, the CAR specifically binds to more than one antigen. In some embodiments, the CAR specifically binds to two antigens, 3 antigens, 4 antigens, 5 antigens, or 6 antigens. In someAttorney Docket No. 58626-723601 embodiments, two antigens, 3 antigens, 4 antigens, 5 antigens, or 6 antigens associated with a disease (e.g., cancer) or condition and / or is expressed by cells associated with a disease (e.g., cancer) or condition. In some embodiments, two antigens, 3 antigens, 4 antigens, 5 antigens, or 6 antigens are associated with the same disease (e.g., cancer) or condition and / or is expressed by cells associated with the same disease (e.g., cancer) or condition. In some embodiments, the antigen is selectively expressed or overexpressed by cells associated with the disease or condition.
[0116] In some embodiments, the antigen can be 5T4, 8H9, avb6 integrin, B7-H6, B cell maturation antigen (BCMA), CA9, a cancer-testes antigen, carbonic anhydrase 9 (CAIX), CCL- 1, CD19, CD20, CD22, CEA, hepatitis B surface antigen, CD23, CD24, CD30, CD33, CD38, CD44, CD44v6, CD44v7 / 8, CD 123, CD 138, CD171, carcinoembryonic antigen (CEA), CE7, a cyclin, cyclin A2, c-Met, dual antigen, EGFR, epithelial glycoprotein 2 (EPG-2), epithelial glycoprotein 40 (EPG-40), EPHa2, ephrinB2, erb-B2, erb-B3, erb-B4, erbB dimers, EGFR vIII, estrogen receptor, Fetal AchR, folate receptor alpha, folate binding protein (FBP), FCRL5, FCRH5, fetal acetylcholine receptor, G250 / CAIX, GD2, GD3, gplOO, Her2 / neu (receptor tyrosine kinase erbB2), HMW-MAA, IL-22R-alpha, IL-13 receptor alpha 2 (IL-13Ra2), kinase insert domain receptor (kdr), kappa light chain, Lewis Y, LI -cell adhesion molecule (LI -CAM), Melanoma-associated antigen (MAGE)-Al, MAGE-A3, MAGE-A6, MART-1, mesothelin, murine CMV, mucin 1 (MUC1), MUC16, NCAM, NKG2D, NKG2D ligands, NY-ESO-1, O- acetylated GD2 (OGD2), oncofetal antigen, Preferentially expressed antigen of melanoma (PRAME), PSCA, progesterone receptor, survivin, ROR1, TAG72, tEGFR, VEGF receptors, BAFF-R, VEGF-R2, Wilms Tumor 1 (WT-1), or a pathogen-specific antigen.
[0117] In some embodiments, the CAR comprises an extracellular antigen-recognition domain that specifically binds to the antigen. In some embodiments, the extracellular antigen-recognition domain can include a portion of an antibody (e.g., scFv, nanobody, variable heavy chain, variable light chain). In some embodiments, the CAR comprises an intracellular signaling domain. In some embodiments, the CAR comprises an extracellular antigen-recognition domain that specifically binds to the antigen and an intracellular signaling domain. In some embodiments, the CAR comprises an extracellular antigen-recognition domain that specifically binds to the antigen and one or more intracellular signaling domain. In some embodiments the extracellular antigen-recognition domain and an intracellular signaling domain are linked via one or more linkers and / or one or more transmembrane domain. In some embodiments, the intracellular signaling domain comprises an ITAM. In some embodiments, the CAR comprises an extracellular antigen-recognition domain that specifically binds to the antigen and anAttorney Docket No. 58626-723601 intracellular signaling domain comprising an IT AM. In some embodiments, the intracellular signaling domain comprises an intracellular domain of a CD3-zeta (CD3Q chain. In some embodiments, the CAR further comprises a costimulatory signaling region. In some embodiments, the costimulatory signaling region comprises a signaling domain. In some embodiments, the costimulatory signaling region comprises a signaling domain of CD28 or 4- 1BB. In some embodiments, the costimulatory domain is a domain of 4- IBB. In some embodiments, the CAR on the genetically engineered T cells can modulate the genetically engineered T cells’ activity (e.g., differentiation, homeostasis, longevity, survival, persistence in vivo).
[0118] In some embodiments, the genetically engineered cells are genetically engineered to bind to an antigen. In some embodiments, the genetically engineered cells are genetically engineered to bind to more than one antigen.
[0119] In some embodiments, the T cells are regulatory T cells, CD4+ T cells, CD8+ T cells, or gamma-delta T cells. In some embodiments, the T cells are primary T cells. In some embodiments, the T cells are primary T cells obtained from the subject.
[0120] In some embodiments, the genetically engineered cells are autologous to the subject. In some embodiments, the genetically engineered cells are allogeneic to the subject. In some embodiments, the subject is human.
[0121] In some embodiments, the subject has Stage I / II disease. In some embodiments, the subject has Stage III / IV disease. In some embodiments, the subject has Minimal Residual Disease (MRD). In some embodiments, the subject is refractory to treatment with one or more prior therapies for the cancer. In some embodiments, the subject achieved an insufficient response to one or more prior therapies for the cancer. In some embodiments, the subject achieves a durable response to the treatment or the medicament. In some embodiments, the durable response is defined as an absence of relapse or remission of the cancer for up to 1 month, 2 month, 3 months, 6 months, 12 months, 24 months or 36 months. In some embodiments, the subject was treatment-naive prior to administration of the treatment or the medicament. In some embodiments, the subject has a higher rate of survival.
[0122] In some embodiments, the subject’s disease baseline characteristics are determined. In some embodiments, the baseline characteristics comprise international prognosis index (IP I) score, serum lactate dehydrogenase (LDH), or sum of the product of diameters (SPD), disease stage, or any combination of the above. The International Prognostic Index (IPI) can be a prognostic index system. The IPI can identify four independent risk groups of patients with a combination of five clinical variables including age, serum lactate dehydrogenase (LDH) level,Attorney Docket No. 58626-723601 tumor stage, Eastern Cooperative Oncology Group (ECOG) Performance Status (PS), and extra- nodal sites of disease. In some embodiments, rate of survival of the subject is based on a score on the IPI. In some aspects, IPI score is based on the independent prognostic roles of age, histology, cancer stage, and serum LDH levels. In some embodiments, the baseline characteristics are defined by Lugano 2014 criteria. The Lugano 2014 criteria can include evaluation by imaging, tumor bulk measurements, and assessments of spleen, liver, and bone marrow involvement. In some embodiments, assessment by the Lugano 2014 criteria involves the use of positron emission tomography (PET)-computed tomography (CT) and / or CT as appropriate for imaging evaluation. In some embodiments, the subject is classified under an international prognosis index (IPI) score. In some embodiments, the subject is classified as low, low-intermediate, or high-intermediate risk on the IPI score. In some embodiments, the subject is classified as high risk on the IPI score. In some aspects, the IPI score can be used to characterize or predict overall survival rates of the subject. In some embodiments, the subject has no risk factor or one risk factor and is considered to be in an IPI low risk group. In some embodiments, the subject has two risk factors and is considered to be in an IPI low-intermediate risk group. In some embodiments, the subject has three risk factors and is considered to be in an IPI high- intermediate risk group. In some embodiments, the subject has a higher or lower IPI score (e.g., as compared to a reference, such as a threshold value or another IPI score at a different time point). In some embodiments, the subject has four or five risk factors and is considered to be in an IPI high risk group. In some embodiments, a higher IPI score is predictive of a worse outcome compared to a lower IPI score, and treating subjects with higher IPI scores typically is less successful than treating subjects with lower IPI scores.
[0123] In some embodiments, the volumetric measure of tumor burden of the subject is measured. In some embodiments, a volumetric measure of tumor burden in the subject is a sum of products of diameter (SPD). In some embodiments, the volumetric measure of tumor burden is measured using computed tomography (CT), positron emission tomography (PET), and / or magnetic resonance imaging (MRI) of the subject. In some embodiments, tumor burden is correlated with SPD. In some embodiments, higher or lower disease and / or tumor burden is correlated with higher or lower SPD. In some embodiments, an SPD threshold value is or is about 30 per cm2, is or is about 40 per cm2, is or is about 50 per cm2, is or is about 60 per cm2, or is or is about 70 per cm2.
[0124] In some embodiments, a level of an inflammatory marker is measured in the subject. In some embodiments, the level of the inflammatory marker is or is about 300 units per liter, is or is about 400 units per liter, is or is about 500 units per liter, or is or is about 600 units per liter. InAttomey Docket No. 58626-723601 some embodiments, the inflammatory marker is lactate dehydrogenase (LDH). In some embodiments, the level of LDH is measured using a colorimetric test or an in vitro enzyme- linked immunosorbent assay. In some aspects, the level of an inflammatory marker can be assessed alone and / or in combination with another measurement from the subject (e.g., volumetric measure of tumor burden).
[0125] In some embodiments, the level of an inflammatory marker and / or the volumetric measure of tumor burden of the subject is obtained prior to administration of the treatment or the medicament. In some embodiments, administration of the treatment or the medicament can result in a change in the level of an inflammatory marker and / or the volumetric measure of tumor burden of the subject.Systems
[0126] Systems, including computer-implemented systems may be used to perform one or more of the methods described herein. The computer-implemented systems may be configured to perform certain functions related to the methods described herein including but not limited to, receiving sequence reads, sorting sequence reads, storing sequence reads, modifying sequence reads, filtering sequence reads, aligning sequence reads, determining consensus sequences, or a combination thereof. The computer-implemented systems may comprise a variety of components used for performing the methods described herein. The computer-implemented systems may comprise one or more processors, one or more operating systems configured to perform executable instructions, one or more memories, one or more computer programs, or a combination thereof.
[0127] The methods described herein may provide certain improvements to computer systems described herein. In some cases, the computer systems described herein may comprise one or more processors, and a memory comprising a machine-executable code that may implement a method. The method implemented by the memory comprising a machine-executable code may be any one of the methods described herein. In some cases, the methods described herein may comprise identifying one more nucleic acid molecules (e.g. cell-free DNA molecules) that are mutation containing. Identifying one or more nucleic acid molecules that are mutation contain may involve differentiating between variants and errors introduced during an amplification process. Identifying one or more nucleic acid molecules that are mutation contain may comprise aligning sequences generated using sequencing. The methods described herein may enable identification of variants with a low error rate (e.g. less than 0.5 parts per million). The low error rate of identification of variants of the methods described herein may increase a processing speed of the computer systems described herein. For example, The low error rate of variantAttorney Docket No. 58626-723601 identification may comprise fewer processing steps by the computer system, including for example alignment, and result in an increased processing speed of the computer systems described herein.
[0128] In another aspect, the present disclosure provides a system. In some embodiments, the system may comprise one or more computer processors. In some embodiments, the systems further comprise computer memory coupled thereto. In some embodiments, the computer memory comprises machine-executable code that, upon execution by the one or more computer processors, implements a method for determining, from sequence reads, a sequence for a first nucleic acid molecule and a sequence of a second nucleic acid molecule. The first nucleic acid molecule and the second nucleic acid molecule may have identical start positions, identical stop positions, or both identical start and stop positions. The methods described herein may comprise obtaining a plurality of sequence reads for each of a plurality of nucleic acid molecules from a biological sample of a subject. The plurality of sequence reads for each of the plurality of nucleic acid molecules may be mapped to a reference genomic sequence. The plurality of sequence reads may be grouped into groups of sequence reads that map to identical start and stop positions on the reference genomic sequence. Grouping the plurality of sequence reads into groups may thereby form a plurality of families. The sequence reads of each family may be grouped into a plurality of subgroups. The plurality of subgroups may comprise a first subgroup and a second subgroup. Additional subgroups may also be generated. The sequence reads of the first subgroup may comprise a nucleotide (e.g., a first nucleotide) at a position (e.g., a first position). The sequence reads of the second subgroup may comprise a nucleotide (e.g., a second nucleotide) at the same position (e.g., the first position). The second nucleotide may differ from the first nucleotide. The methods may comprise collapsing the sequence reads of the first subgroup to form a consensus sequence (e.g., a first consensus sequence) corresponding with the sequence of a nucleic acid molecule (e.g., a first nucleic acid molecule) of the plurality of nucleic acid molecules. The methods may comprise collapsing the sequence reads of the second subgroup to form a consensus sequence (e.g., a second consensus sequence) corresponding with the sequence for a nucleic acid molecule (e.g., a second nucleic acid molecule) from the plurality of nucleic acid molecules. The first nucleic acid molecule may differ in sequence from the second nucleic acid molecule.
[0129] In another aspect, the present disclosure provides a system. In some embodiments, the system may comprise one or more computer processors. In some embodiments, the systems further comprise computer memory coupled thereto. In some embodiments, the computer memory comprises machine-executable code that, upon execution by the one or more computerAttorney Docket No. 58626-723601 processors, implements a method. In some embodiment, the method is described elsewhere herein. In some embodiments, the method comprises sequencing a plurality of cell-free nucleic acid molecules. In some embodiments, the method comprises having sequenced a plurality of cell-free nucleic acid molecules. In some embodiments, the method comprises sequencing a plurality of cell-free DNA molecules. In some embodiments, the method comprises having sequenced a plurality of cell-free DNA molecules. The plurality of cell-free DNA molecules can comprise a plurality of mutations. The plurality of mutations can comprise a plurality of single nucleotide variants (SNVs). The plurality of mutations can comprise a plurality of phased variants (PVs). The terms “phased variants” or “PVs,” as used interchangeably herein, generally refers to two or more mutations (e.g., SNVs or indels) that occur in cis (i.e., on the same strand of a nucleic acid molecule) within a single cell-free nucleic acid molecule. In some cases, a cell- free nucleic acid molecule can be a cell-free deoxyribonucleic acid (cfDNA) molecule. In some cases, a cfDNA molecule can be derived from a diseased tissue, such as a tumor (e.g., a circulating tumor DNA (ctDNA) molecule). The plurality of cell-free DNA molecules can be obtained or derived from a subject. In some embodiments, the subject has cancer. Sequencing the plurality of cell-free DNA molecules can produce a plurality of sequence reads. The plurality of sequence reads can comprise multiple sequence reads. In some embodiments, a sequencing read in the plurality of sequence reads is at least 100 nucleotides, at least 150 nucleotides, at least 160 nucleotides, at least 170 nucleotides, at least 180 nucleotides, at least 190 nucleotides, or at least 200 nucleotides in length. The plurality of sequence reads may comprise at least about 1,000, at least about 2,000, at least about 3,000, at least about 4,000, at least about 5,000, at least about 6,000, at least about 7,000, at least about 8,000, at least about 9,000, at least about 10,000, at least about 20,000, at least about 30,000, at least about 40,000, at least about 50,000, at least about 60,000, at least about 70,000, at least about 80,000, at least about 90,000, at least about 100,000, at least about 200,000, at least about 300,000, at least about 400,000, at least about 500,000, at least about 600,000, at least about 700,000, at least about 800,000, at least about 900,000, at least about 1,000,000, at least about 10,000,000, at least about 50,000,000, at least about 100,000,000, at least about 500,000,000, at least about 1,000,000,000, or more sequence reads. The plurality of sequence reads may comprise at most about 1,000, at most about 2,000, at most about 3,000, at most about 4,000, at most about 5,000, at most about 6,000, at most about 7,000, at most about 8,000, at most about 9,000, at most about 10,000, at most about 20,000, at most about 30,000, at most about 40,000, at most about 50,000, at most about 60,000, at most about 70,000, at most about 80,000, at most about 90,000, at most about 100,000, at most about 200,000, at most about 300,000, at most about 400,000, at most about 500,000, at most aboutAttorney Docket No. 58626-723601600,000, at most about 700,000, at most about 800,000, at most about 900,000, at most about 1,000,000, at most about 10,000,000, at most about 50,000,000, at most about 100,000,000, at most about 500,000,000, at most about 1,000,000,000, or less sequence reads. The plurality of sequence reads may comprise about 100,000 to about 1,000,000,000, about 200,000 to about 900,000,000, about 300,000 to about 80,0000,000, about 400,000 to about 70,0000,000, about 500000 to about 600,000,000, about 600,000 to about 500,000,000, about 700,000 to about 400,000,000, about 800,000 to about 300,000,000, about 900,000 to about 200,000,000, about 1,000,000 to about 100,000,000, or about 5,000,000 to about 50,000,000 sequence reads. In some embodiments, the method further comprises aligning the plurality of sequence reads to a reference sequence. In some embodiments, the reference sequence is a reference genomic sequence. In some embodiments, the reference genomic sequence is derived from an animal. In some embodiments, the reference genomic sequence is derived from a mammal. In some embodiments, the reference genomic sequence is derived from a human. In some embodiments, aligning the plurality of sequence reads to a reference sequence can produce a plurality of aligned sequence reads. In some embodiments, the method further comprises processing the plurality of aligned sequence reads. In some embodiments, the method further comprises processing the plurality of aligned sequence reads to detect a presence of a mutation. In some embodiments, the mutation is a somatic mutation. In some embodiments, the somatic mutation can include a single nucleotide variant (SNV), a phased variant (PV), an indel, a rearrangement, a fusion, a breakpoint, a structural variant, a variable number of tandem repeat, a hypervariable region, a mini satellite, a dinucleotide repeat, a trinucleotide repeat, a tetranucleotide repeat, a simple sequence repeat, a point mutation, a deletion mutation, a frameshift mutation, a silent mutation, a nonsense mutation, or a combination thereof. In some embodiments, the somatic mutation is selected from the Catalogue of Somatic Mutations in Cancer (COSMIC) database.
[0130] In some embodiments, the somatic mutation is in one or more cell-free DNA molecules from among the plurality of cell-free DNA molecules. In some embodiments, the somatic mutation is a base-change mutation. In some embodiments, the base-change mutation can be an adenine-to-thymine (A>T) mutation, an adenine-to-cytosine (A>C) mutation, an adenine-to- guanine (A>G) mutation, a thymine-to-adenine (T>A) mutation, a thymine-to-cytosine (T>C) mutation, a thymine-to-guanine (T>G) mutation, a cytosine-to-adenine (C>A) mutation, a cytosine-to-thymine (C>T) mutation, a cytosine-to-guanine (C>G) mutation, a guanine-to- adenine (G>A) mutation, a guanine-to-thymine (G>T) mutation, or a guanine-to-cytosine (G>C) mutation. In some embodiments, processing the plurality of aligned sequence reads further comprises filtering the plurality of aligned sequence reads. In some embodiments, filtering theAttorney Docket No. 58626-723601 plurality of aligned sequence reads is based at least in part on an expected error rate of detection of the base-change mutation in individual aligned sequence reads. In some embodiments, the expected error rate can be determined by factors such as the specific base-change mutation (e.g., A>T, A>C, A>G, T>A, T>C, T>G, OA, OT, OG, G>A, G>T, G>C), the number of copies of cell-free DNA molecules observed with the same somatic mutation, the number of independent copies of Watson and Crick strands observed carrying the same somatic mutation, and the location of the somatic mutation among the sequencing read. Watson and Crick strands can mean the two complementary strands of a DNA molecule. In some embodiments, filtering the plurality of aligned sequence reads can generate a set of filtered sequence reads.
[0131] In some embodiments, processing the plurality of aligned sequence reads can include trimming an end of the sequencing read to generate the set of filtered sequencing read. In some embodiments, trimming an end of the sequencing read can include trimming at most 1 base pair, at most 2 base pairs, at most 3 base pairs, at most 4 base pairs, at most 5 base pairs, at most 6 base pairs, at most 7 base pairs, at most 8 base pairs, at most 9 base pairs, at most 10 base pairs, at most 11 base pairs, at most 12 base pairs, at most 13 base pairs, at most 14 base pairs, at most 15 base pairs, at most 16 base pairs, at most 17 base pairs, at most 18 base pairs, at most 19 base pairs, or at most 20 base pairs. In some embodiments, processing the plurality of aligned sequence reads can include trimming both ends of the sequencing read to generate the set of filtered sequencing read. In some embodiments, processing the plurality of aligned sequence reads can include trimming both ends of the sequencing read to generate the set of filtered sequencing read as different number of base pairs. In some embodiments, determining the number of base pairs to trim from either ends of an aligned sequencing in the plurality of aligned sequence reads is based at least in part on a strand location the base-change mutation on the aligned sequencing read. In some embodiments, the set of filtered sequence reads comprises variant types of somatic mutations (e.g., SNVs). In some embodiments, the set of filtered sequence reads comprises a uniformly low expected error rate of detection across at least 20% of the variant types, at least 30% of the variant types, at least 40% of the variant types, at least 50% of the variant types, at least 60% of the variant types, at least 70% of the variant types, at least 80% of the variant types, at least 90% of the variant types, or 100% of the variant types. In some embodiments, the uniformly low expected error rate is no more than 0.005 parts per million (ppm), 0.01 parts per million (ppm), 0.02 parts per million (ppm), 0.03 parts per million (ppm), 0.04 parts per million (ppm), 0.05 parts per million (ppm), 0.06 parts per million (ppm), 0.07 parts per million (ppm), 0.08 parts per million (ppm), 0.09 parts per million (ppm), 0.1 parts per million (ppm), or 0.2 parts per million (ppm). In some embodiments, the method furtherAttorney Docket No. 58626-723601 comprises processing the set of filtered sequence reads. In some embodiments, the method comprises processing the set of filtered sequence reads to determine whether at least one of the set of filtered sequence reads is mutation-containing. In some embodiments, determining whether at least one of the set of filtered sequence reads is mutation-containing has an error rate of no more than 0.005 parts per million (ppm), 0.01 parts per million (ppm), 0.02 parts per million (ppm), 0.03 parts per million (ppm), 0.04 parts per million (ppm), 0.05 parts per million (ppm), 0.06 parts per million (ppm), 0.07 parts per million (ppm), 0.08 parts per million (ppm), 0.09 parts per million (ppm), 0.1 parts per million (ppm), or 0.2 parts per million (ppm). In some embodiments, the determination of whether the set of filtered sequence reads is mutationcontaining determines whether a cell-free DNA molecule of the plurality of cell-free DNA molecules is mutation-containing.Non-transitory computer-readable medium
[0132] Non-transitory computer-readable media may be used to perform one or more of the methods described herein. The non-transitory computer-readable media may be configured to perform certain functions related to the methods described herein including but not limited to, receiving sequence reads, sorting sequence reads, storing sequence reads, modifying sequence reads, filtering sequence reads, aligning sequence reads, determining consensus sequences, determining non-consensus sequences, or a combination thereof.
[0133] In another aspect, the present disclosure provides a non-transitory computer-readable medium. In some embodiments, the non-transitory computer-readable medium comprises machine-executable code that, upon execution by one or more computer processors, implements a method for determining, from sequence reads, a sequence for a first nucleic acid molecule and a sequence of a second nucleic acid molecule. The first nucleic acid molecule and the second nucleic acid molecule may have identical start positions, identical stop positions, or both identical start and stop positions. In some embodiments, the method may comprise obtaining a plurality of sequence reads for each of a plurality of nucleic acid molecules from a biological sample of a subject. The plurality of sequence reads for each of the plurality of nucleic acid molecules may be mapped to a reference genomic sequence. The plurality of sequence reads may be grouped into groups of sequence reads that map to identical start and stop positions on the reference genomic sequence. Grouping the plurality of sequence reads into groups may thereby form a plurality of families. The sequence reads of each family may be grouped into a plurality of subgroups. The plurality of subgroups may comprise a first subgroup and a second subgroup. Additional subgroups may also be generated. The sequence reads of the first subgroup may comprise a nucleotide (e.g., a first nucleotide) at a position (e.g., a first position). The sequenceAttorney Docket No. 58626-723601 reads of the second subgroup may comprise a nucleotide (e.g., a second nucleotide) at the same position (e.g., the first position). The second nucleotide may differ from the first nucleotide. The methods may comprise collapsing the sequence reads of the first subgroup to form a consensus sequence (e.g., a first consensus sequence) corresponding with the sequence of a nucleic acid molecule (e.g., a first nucleic acid molecule) of the plurality of nucleic acid molecules. The methods may comprise collapsing the sequence reads of the second subgroup to form a consensus sequence (e.g., a second consensus sequence) corresponding with the sequence for a nucleic acid molecule (e.g., a second nucleic acid molecule) from the plurality of nucleic acid molecules. The first nucleic acid molecule may differ in sequence from the second nucleic acid molecule.
[0134] In another aspect, the present disclosure provides a non-transitory computer-readable medium. In some embodiments, the non-transitory computer-readable medium comprises machine-executable code that, upon execution by one or more computer processors, implements a method. In some embodiments, the method is described elsewhere herein. In some embodiments, the method comprises sequencing a plurality of cell-free nucleic acid molecules. In some embodiments, the method comprises having sequenced a plurality of cell-free nucleic acid molecules. In some embodiments, the method comprises sequencing a plurality of cell-free DNA molecules. In some embodiments, the method comprises having sequenced a plurality of cell-free DNA molecules. The plurality of cell-free DNA molecules can comprise a plurality of mutations. The plurality of mutations can comprise a plurality of single nucleotide variants (SNVs). The plurality of mutations can comprise a plurality of phased variants (PVs). In some cases, a cell-free nucleic acid molecule can be a cell-free deoxyribonucleic acid (cfDNA) molecule. In some cases, a cfDNA molecule can be derived from a diseased tissue, such as a tumor (e.g., a circulating tumor DNA (ctDNA) molecule). The plurality of cell-free DNA molecules can be obtained or derived from a subject. In some embodiments, the subject has cancer. Sequencing the plurality of cell-free DNA molecules can produce a plurality of sequence reads. The plurality of sequence reads can comprise multiple sequence reads. In some embodiments, a sequencing read in the plurality of sequence reads is at least 100 nucleotides, at least 150 nucleotides, at least 160 nucleotides, at least 170 nucleotides, at least 180 nucleotides, at least 190 nucleotides, or at least 200 nucleotides in length. The plurality of sequence reads may comprise at least about 1,000, at least about 2,000, at least about 3,000, at least about 4,000, at least about 5,000, at least about 6,000, at least about 7,000, at least about 8,000, at least about 9,000, at least about 10,000, at least about 20,000, at least about 30,000, at least about 40,000, at least about 50,000, at least about 60,000, at least about 70,000, at least about 80,000, at leastAttorney Docket No. 58626-723601 about 90,000, at least about 100,000, at least about 200,000, at least about 300,000, at least about 400,000, at least about 500,000, at least about 600,000, at least about 700,000, at least about 800,000, at least about 900,000, at least about 1,000,000, at least about 10,000,000, at least about 50,000,000, at least about 100,000,000, at least about 500,000,000, at least about 1,000,000,000, or more sequence reads. In some embodiments, the plurality of sequence reads may comprise at most about 1,000, at most about 2,000, at most about 3,000, at most about 4,000, at most about 5,000, at most about 6,000, at most about 7,000, at most about 8,000, at most about 9,000, at most about 10,000, at most about 20,000, at most about 30,000, at most about 40,000, at most about 50,000, at most about 60,000, at most about 70,000, at most about 80,000, at most about 90,000, at most about 100,000, at most about 200,000, at most about 300,000, at most about 400,000, at most about 500,000, at most about 600,000, at most about 700,000, at most about 800,000, at most about 900,000, at most about 1,000,000, at most about 10,000,000, at most about 50,000,000, at most about 100,000,000, at most about 500,000,000, at most about 1,000,000,000, or less sequence reads. In some embodiments, the plurality of sequence reads may comprise about 100,000 to about 1,000,000,000, about 200,000 to about 900,000,000, about 300,000 to about 80,0000,000, about 400,000 to about 70,0000,000, about 500000 to about 600,000,000, about 600,000 to about 500,000,000, about 700,000 to about 400,000,000, about 800,000 to about 300,000,000, about 900,000 to about 200,000,000, about 1,000,000 to about 100,000,000, or about 5,000,000 to about 50,000,000 sequence reads. In some embodiments, the method further comprises aligning the plurality of sequence reads to a reference sequence. In some embodiments, the reference sequence is a reference genomic sequence. In some embodiments, the reference genomic sequence is derived from an animal. In some embodiments, the reference genomic sequence is derived from a mammal. In some embodiments, the reference genomic sequence is derived from a human. In some embodiments, aligning the plurality of sequence reads to a reference sequence can produce a plurality of aligned sequence reads. In some embodiments, the method further comprises processing the plurality of aligned sequence reads. In some embodiments, the method further comprises processing the plurality of aligned sequence reads to detect a presence of a mutation. In some embodiments, the mutation is a somatic mutation. In some embodiments, the somatic mutation can include a single nucleotide variant (SNV), a phased variant (PV), an indel, a rearrangement, a fusion, a breakpoint, a structural variant, a variable number of tandem repeat, a hypervariable region, a mini satellite, a dinucleotide repeat, a trinucleotide repeat, a tetranucleotide repeat, a simple sequence repeat, a point mutation, a deletion mutation, a frameshift mutation, a silent mutation, a nonsenseAttorney Docket No. 58626-723601 mutation, or a combination thereof. In some embodiments, the somatic mutation is selected from the Catalogue of Somatic Mutations in Cancer (COSMIC) database.
[0135] In some embodiments, the somatic mutation is in one or more cell-free DNA molecules from among the plurality of cell-free DNA molecules. In some embodiments, the somatic mutation is a base-change mutation. In some embodiments, the base-change mutation can be an adenine-to-thymine (A>T) mutation, an adenine-to-cytosine (A>C) mutation, an adenine-to- guanine (A>G) mutation, a thymine-to-adenine (T>A) mutation, a thymine-to-cytosine (T>C) mutation, a thymine-to-guanine (T>G) mutation, a cytosine-to-adenine (C>A) mutation, a cytosine-to-thymine (OT) mutation, a cytosine-to-guanine (C>G) mutation, a guanine-to- adenine (G>A) mutation, a guanine-to-thymine (G>T) mutation, or a guanine-to-cytosine (G>C) mutation. In some embodiments, processing the plurality of aligned sequence reads further comprises filtering the plurality of aligned sequence reads. In some embodiments, filtering the plurality of aligned sequence reads is based at least in part on an expected error rate of detection of the base-change mutation in individual aligned sequence reads. In some embodiments, the expected error rate can be determined by factors such as the specific base-change mutation (e.g., A>T, A>C, A>G, T>A, T>C, T>G, C>A, C>T, C>G, G>A, G>T, G>C), the number of copies of cell-free DNA molecules observed with the same somatic mutation, the number of independent copies of Watson and Crick strands observed carrying the same somatic mutation, and the location of the somatic mutation among the sequencing read. Watson and Crick strands can mean the two complementary strands of a DNA molecule. In some embodiments, filtering the plurality of aligned sequence reads can generate a set of filtered sequence reads.
[0136] In some embodiments, processing the plurality of aligned sequence reads can include trimming an end of the sequencing read to generate the set of filtered sequencing read. In some embodiments, trimming an end of the sequencing read can include trimming at most 1 base pair, at most 2 base pairs, at most 3 base pairs, at most 4 base pairs, at most 5 base pairs, at most 6 base pairs, at most 7 base pairs, at most 8 base pairs, at most 9 base pairs, at most 10 base pairs, at most 11 base pairs, at most 12 base pairs, at most 13 base pairs, at most 14 base pairs, at most 15 base pairs, at most 16 base pairs, at most 17 base pairs, at most 18 base pairs, at most 19 base pairs, or at most 20 base pairs. In some embodiments, processing the plurality of aligned sequence reads can include trimming both ends of the sequencing read to generate the set of filtered sequencing read. In some embodiments, processing the plurality of aligned sequence reads can include trimming both ends of the sequencing read to generate the set of filtered sequencing read as different number of base pairs. In some embodiments, determining the number of base pairs to trim from either ends of an aligned sequencing in the plurality of alignedAttorney Docket No. 58626-723601 sequence reads is based at least in part on a strand location the base-change mutation on the aligned sequencing read. In some embodiments, the set of filtered sequence reads comprises variant types of somatic mutations (e.g., SNVs). In some embodiments, the set of filtered sequence reads comprises a uniformly low expected error rate of detection across at least 20% of the variant types, at least 30% of the variant types, at least 40% of the variant types, at least 50% of the variant types, at least 60% of the variant types, at least 70% of the variant types, at least 80% of the variant types, at least 90% of the variant types, or 100% of the variant types. In some embodiments, the uniformly low expected error rate is no more than 0.005 parts per million (ppm), 0.01 parts per million (ppm), 0.02 parts per million (ppm), 0.03 parts per million (ppm), 0.04 parts per million (ppm), 0.05 parts per million (ppm), 0.06 parts per million (ppm), 0.07 parts per million (ppm), 0.08 parts per million (ppm), 0.09 parts per million (ppm), 0.1 parts per million (ppm), or 0.2 parts per million (ppm). In some embodiments, the method further comprises processing the set of filtered sequence reads. In some embodiments, the method comprises processing the set of filtered sequence reads to determine whether at least one of the set of filtered sequence reads is mutation-containing. In some embodiments, determining whether at least one of the set of filtered sequence reads is mutation-containing has an error rate of no more than 0.005 parts per million (ppm), 0.01 parts per million (ppm), 0.02 parts per million (ppm), 0.03 parts per million (ppm), 0.04 parts per million (ppm), 0.05 parts per million (ppm), 0.06 parts per million (ppm), 0.07 parts per million (ppm), 0.08 parts per million (ppm), 0.09 parts per million (ppm), 0.1 parts per million (ppm), or 0.2 parts per million (ppm). In some embodiments, the determination of whether the set of filtered sequence reads is mutationcontaining determines whether a cell-free DNA molecule of the plurality of cell-free DNA molecules is mutation-containing.Computer systems
[0137] The present disclosure provides computer systems that are programmed to implement methods of the disclosure. FIG. 3 shows a computer system 301 that is programmed or otherwise configured to analyze sequence read data. The computer system 301 can regulate various aspects of data analysis of the present disclosure. The computer system 301 can be an electronic device of a user or a computer system that is remotely located with respect to the electronic device. The electronic device can be a mobile electronic device.
[0138] The computer system 301 can include a central processing unit (CPU, also “processor” and “computer processor” herein) 305, which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 301 also includes memory or memory location 310 (e.g., random-access memory, read-only memory,Attorney Docket No. 58626-723601 flash memory), electronic storage unit 315 (e.g., hard disk), communication interface 320 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 325, such as cache, other memory, data storage and / or electronic display adapters. The memory 310, storage unit 315, interface 320 and peripheral devices 325 are in communication with the CPU 305 through a communication bus (solid lines), such as a motherboard. The storage unit 315 can be a data storage unit (or data repository) for storing data. The computer system 301 can be operatively coupled to a computer network (“network”) 330 with the aid of the communication interface 320. The network 330 can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network 330 in some cases is a telecommunication and / or data network. The network 330 can include one or more computer servers, which can enable distributed computing, such as cloud computing. The network 330, in some cases with the aid of the computer system 301, can implement a peer-to- peer network, which may enable devices coupled to the computer system 301 to behave as a client or a server.
[0139] The CPU 305 can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 310. The instructions can be directed to the CPU 305, which can subsequently program or otherwise configure the CPU 305 to implement methods of the present disclosure. Examples of operations performed by the CPU 305 can include fetch, decode, execute, and writeback.
[0140] The CPU 305 can be part of a circuit, such as an integrated circuit. One or more other components of the system 301 can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0141] The storage unit 315 can store files, such as drivers, libraries and saved programs. The storage unit 315 can store user data, e.g., user preferences and user programs. The computer system 301 in some cases can include one or more additional data storage units that are external to the computer system 301, such as located on a remote server that is in communication with the computer system 301 through an intranet or the Internet.
[0142] The computer system 301 can communicate with one or more remote computer systems through the network 330. For instance, the computer system 301 can communicate with a remote computer system of a user. Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC’s (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system 301 via the network 330.Attorney Docket No. 58626-723601
[0143] Methods as described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 301, such as, for example, on the memory 310 or electronic storage unit 315. The machine executable or machine readable code can be provided in the form of software. During use, the code can be executed by the processor 305. In some cases, the code can be retrieved from the storage unit 315 and stored on the memory 310 for ready access by the processor 305. In some situations, the electronic storage unit 315 can be precluded, and machine-executable instructions are stored on memory 310.
[0144] The code can be pre-compiled and configured for use with a machine having a processer adapted to execute the code, or can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a precompiled or as-compiled fashion.
[0145] Aspects of the systems and methods provided herein, such as the computer system 301, can be embodied in programming. Various aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. “Storage” type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.
[0146] Hence, a machine-readable medium, such as computer-executable code, may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium orAttorney Docket No. 58626-723601 physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as may be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0147] The computer system 301 can include or be in communication with an electronic display 335 that comprises a user interface (UI) 340 for providing, for example, nucleic acid sequence analysis parameters. Examples of UI’s include, without limitation, a graphical user interface (GUI) and web-based user interface.
[0148] Methods and systems of the present disclosure can be implemented by way of one or more algorithms. An algorithm can be implemented by way of software upon execution by the central processing unit 305. The algorithm can, for example, compare nucleic acid sequences to other nucleic acid sequences, including a reference genomic sequence.
[0149] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.NUMBERED EMBODIEMENTS
[0150] The following embodiments recite nonlimiting permutations of a combination of features disclosed herein. Other permutations of combinations of features are also contemplated. InAttorney Docket No. 58626-723601 particular, each of these numbered embodiments is contemplated as depending from or related to every previous or subsequent numbered embodiments, independent of their order listed. Embodiment 1: a method comprising: (a) sequencing or having sequenced a plurality of cell-free DNA molecules obtained or derived from a subject to produce a plurality of sequence reads; (b) aligning the plurality of sequence reads to a reference genomic sequence to produce a plurality of aligned sequence reads; (c) processing the plurality of aligned sequence reads to detect a presence of a somatic mutation in one or more cell-free DNA molecules from among the plurality of cell-free DNA molecules, wherein the somatic mutation comprises a base-change mutation selected from the group consisting of: an adenine-to-thymine (A>T) mutation, an adenine-to-cytosine (A>C) mutation, an adenine-to-guanine (A>G) mutation, a thymine-to- adenine (T>A) mutation, a thymine-to-cytosine (T>C) mutation, a thymine-to-guanine (T>G) mutation, a cytosine-to-adenine (C>A) mutation, a cytosine-to-thymine (OT) mutation, a cytosine-to-guanine (OG) mutation, a guanine-to-adenine (G>A) mutation, a guanine-to- thymine (G>T) mutation, and a guanine-to-cytosine (G>C) mutation, wherein the processing in (d) comprises filtering the plurality of aligned sequence reads based at least in part on an expected error rate of detection of the base-change mutation in individual aligned sequence reads, thereby generating a set of filtered sequence reads; and (d) processing the set of filtered sequence reads to determine whether at least one of the set of filtered sequence reads is mutationcontaining, wherein the determining in (e) has an error rate of no more than 0.5 parts per million. Embodiment 2: the method of embodiment 1, wherein the somatic mutation comprises a single nucleotide variant (SNV).Embodiment 3: the method of embodiment 1 or embodiment 2, further comprising obtaining a biological sample from the subject, and extracting the plurality of cell-free DNA molecules from the biological sample.Embodiment 4: the method of embodiment 3, wherein the biological sample comprises blood, serum, plasma, tumor cells, saliva, urine, cerebrospinal fluid, lymphatic fluid, prostatic fluid, seminal fluid, milk, sputum, stool, tears, vaginal secretion, semen sample, bone marrow, or a combination thereof, or derivatives thereof.Embodiment 5: the method of any one of embodiments 1-3, wherein the sequencing in (a) further comprising amplifying the plurality of cell-free DNA molecules to produce amplified DNA molecules, and sequencing the amplified DNA molecules.Embodiment 6: the method of embodiment 5, wherein the amplifying comprises polymerase chain reaction.Attorney Docket No. 58626-723601Embodiment 7: the method of any one of embodiments 1-6, wherein the sequencing in (a) further comprises targeted sequencing.Embodiment 8: the method of embodiment 7, wherein the targeted sequencing further comprises performing probe enrichment of the plurality of cell-free DNA molecules. Embodiment 9: the method of any one of embodiments 1-8, wherein the sequencing in (a) further comprises whole genome sequencing or whole exome sequencing.Embodiment 10: the method of any one of embodiments 1-9, wherein the aligning in (b) further comprises aligning at least 1 thousand sequence reads, at least 10 thousand sequence reads, at least 100 thousand sequence reads, at least 1 million sequence reads, at least 10 million sequence reads, or at least 100 million sequence reads.Embodiment 11: the method of any one of embodiments 1-10, wherein the aligning in (b) further comprises aligning sequence reads represented of at least 1 thousand cell-free DNA molecules, at least 10 thousand cell-free DNA molecules, at least 100 thousand cell-free DNA molecules, at least 1 million cell-free DNA molecules, at least 10 million cell-free DNA molecules, or at least 100 million cell-free DNA molecules.Embodiment 12: the method of any one of embodiments 1-11, wherein the reference genomic sequence is a human genome assembly.Embodiment 13: the method of embodiment 12, wherein the human genome assembly is an HG19 human genome assembly or an HG38 human genome assembly.Embodiment 14: the method of any one of embodiments 1-13, wherein the aligning in (b) further comprises performing a Burrows-Wheeler alignment (BWA) algorithm, DRAGEN, minimap2, or a Bowtie alignment algorithm.Embodiment 15: the method of any one of embodiments 1-14, wherein (c) further comprises pre-processing the plurality of aligned sequence reads to filter out at least a portion of the plurality of aligned sequence reads.Embodiment 16: the method of embodiment 15, wherein the pre-processing comprises demultiplexing FASTQ files, extracting unique molecular identifiers (UMIDs), performing molecular barcode-mediated error suppression (e.g., integrated digital error suppression), performing background polishing, deduplicating barcodes, using fastP, fastQC, DRAGEN, Trimmomatic, DRAGENm Pare, bowtie, BWA, or MEM.Embodiment 17: the method of any one of embodiments 1-16, wherein the detecting in (c) further comprises processing the plurality of aligned sequence reads using a somatic mutation calling algorithm.Attorney Docket No. 58626-723601Embodiment 18: the method of embodiment 17, wherein the somatic mutation calling algorithm comprises MuTect2, VarScan2, Strelka2, MUsE, DeepVariant, SomaticSniper or FreeBayes. Embodiment 19: the method of any one of embodiments 1-18, wherein the filtering in (c) further comprises comparing the expected error rate of detection of the base-change mutation to a pre-determined error rate threshold.Embodiment 20: the method of any one of embodiments 1-19, wherein the filtering in (c) further comprises processing the base-change mutation using a trained algorithm.Embodiment 21: the method of embodiment 20, wherein the trained algorithm comprises a machine learning model or a statistical model.Embodiment 22: the method of embodiment 21, wherein the machine learning model comprises a natural language processing model, an artificial neural network, a decision tree, a random forest, a naive bayes classifier, a boosting algorithm, a k-nearest neighbor algorithm, a clustering model, or a Principal Component Analysis.Embodiment 23: the method of embodiment 21, wherein the statistical model comprises a regression model, a classification model, a Monte Carlos simulation, or a polynomial model. Embodiment 24: the method of any one of embodiments 1-23, wherein the determining in (d) is further based at least in part on a number of copies of the one or more cell-free DNA molecules having the somatic mutation.Embodiment 25: the method of embodiment 24, wherein the number of copies is compared to a threshold based at least in part on the expected error rate of detection of the base-change mutation.Embodiment 26: the method of any one of embodiments 1-25, wherein the determining in (d) is further based at least in part on a number of independent copies of both Watson and Crick complementary strands present in the one or more cell-free DNA molecules having the somatic mutation.Embodiment 27: the method of embodiment 26, wherein the number of independent copies is compared to a threshold based at least in part on the expected error rate of detection of the basechange mutation.Embodiment 28: the method of any one of embodiments 1-27, wherein the determining in (d) is further based at least in part on a strand location of the base-change mutation in the one or more cell-free DNA molecules.Embodiment 29: the method of embodiment 28, wherein the strand location is at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9Attorney Docket No. 58626-723601 nucleotides, or at least 10 nucleotides away from a 5’ end or a 3’ end of the one or more cell-free DNA molecules.Embodiment 30: the method of any one of embodiments 1-29, wherein the somatic mutation comprises the A>T mutation.Embodiment 31: the method of any one of embodiments 1-29, wherein the somatic mutation comprises the A>C mutation.Embodiment 32: the method of any one of embodiments 1-29, wherein the somatic mutation comprises the A>G mutation.Embodiment 33: the method of any one of embodiments 1-29, wherein the somatic mutation comprises the T>A mutation.Embodiment 34: the method of any one of embodiments 1-29, wherein the somatic mutation comprises the T>C mutation.Embodiment 35: the method of any one of embodiments 1-29, wherein the somatic mutation comprises the T>G mutation.Embodiment 36: the method of any one of embodiments 1-29, wherein the somatic mutation comprises the OA mutation.Embodiment 37: the method of any one of embodiments 1-29, wherein the somatic mutation comprises the OT mutation.Embodiment 38: the method of any one of embodiments 1-29, wherein the somatic mutation comprises the OG mutation.Embodiment 39: the method of any one of embodiments 1-29, wherein the somatic mutation comprises the G>A mutation.Embodiment 40: the method of any one of embodiments 1-29, wherein the somatic mutation comprises the G>T mutation.Embodiment 41: the method of any one of embodiments 1-29, wherein the somatic mutation comprises the G>C mutation.Embodiment 42: the method of any one of embodiments 1-29, wherein the somatic mutation is selected from the group consisting of A>T, OG, A>C, T>A, G>C, and T>G.Embodiment 43: the method of any one of embodiments 1-29, wherein the somatic mutation does not comprise any of OT and G>AEmbodiment 44: the method of any one of embodiments 1-29, wherein the somatic mutation does not comprise any of OT, G>A, G>T, and OAAttorney Docket No. 58626-723601Embodiment 45: the method of any one of embodiments 1-44, wherein the filtering in (c) further comprises selecting a subset of the plurality of aligned sequence reads, and analyzing the subset of the plurality of aligned sequence reads.Embodiment 46: the method of embodiment 45, wherein the subset is selected from among the plurality of aligned sequence reads based at least in part on having a lowest expected error rate of detection of the base-change mutation.Embodiment 47: the method of any one of embodiments 1-46, wherein the subject has been diagnosed with cancer, and wherein the method further comprises determining the Minimal Residual Disease (MRD) status in the subject, based at least in part on the determining in (d). Embodiment 48: the method of any one of embodiments 1-47, wherein the subject has been administered a treatment for cancer, and wherein the method further comprises assessing a therapeutic response of the subject in response to the treatment for cancer, based at least in part on the determining in (d).Embodiment 49: the method of any one of embodiments 1-48, wherein the subject has been administered a treatment for cancer, and wherein the method further comprises assessing a risk of cancer recurrence or relapse in the subject, based at least in part on the determining in (d). Embodiment 50: the method of any one of embodiments 1-49, wherein the subject has been administered a first treatment for cancer, and wherein the method further comprises treating the subject with a second treatment for cancer, based at least in part on the determining in (d). Embodiment 51: the method of any one of embodiments 1-50, further comprising administering, to the subject, an effective amount of a treatment for cancer, based at least in part on the determining in (d).Embodiment 52: the method of any one of embodiments 1-51, further comprising manufacturing a medicament for treating cancer in the subject, based at least in part on the determining in (d).Embodiment 53: the method of any one of embodiments 47-52, wherein the cancer is a blood cancer.Embodiment 54: the method of any one of embodiments 47-52, wherein the cancer is a solid tumor.Embodiment 55: the method of any one of embodiments 47-52, wherein the cancer is a B cell malignancy.Embodiment 56: the method of any one of embodiments 47-52, wherein the cancer is leukemia or lymphoma.Attorney Docket No. 58626-723601Embodiment 57: the method of any one of embodiments 47-52, wherein the cancer is large B-cell lymphoma (LBCL) or a Diffuse Large B-Cell Lymphoma (DLBCL).Embodiment 58: the method of any one of embodiments 47-52, wherein the cancer is selected from acute myeloid (or myelogenous) leukemia (AML), chronic myeloid (or myelogenous) leukemia (CML), acute lymphocytic (or lymphoblastic) leukemia (ALL), chronic lymphocytic leukemia (CLL), hairy cell leukemia (HCL), small lymphocytic lymphoma (SLL), Mantle cell lymphoma (MCL), Marginal zone lymphoma, Burkitt lymphoma, Hodgkin lymphoma (HL), non-Hodgkin lymphoma (NHL), Anaplastic large cell lymphoma (ALCL), follicular lymphoma, refractory follicular lymphoma, diffuse large B-cell lymphoma (DLBCL) and multiple myeloma (MM), adult ALL.Embodiment 59: the method of any one of embodiments 47-52, wherein the cancer is a bladder cancer, colorectal cancer, breast cancer, prostate cancer, renal cancer, hepatocellular cancer, lung cancer, ovarian cancer, cervical cancer, pancreatic cancer, rectal cancer, thyroid cancer, uterine cancer, gastric cancer, esophageal cancer, head and neck cancer, melanoma, neuroendocrine cancers, CNS cancers, brain tumors, bone cancer, or soft tissue sarcoma.Embodiment 60: the method of any one of embodiments 47-52, wherein the treatment or the medicament is a first-line therapy.Embodiment 61: the method of any one of embodiments 47-52, wherein the treatment or the medicament is a second-line therapy.Embodiment 62: the method of any one of embodiments 47-52, wherein the treatment comprises radiographic imaging of the subject.Embodiment 63: the method of any one of embodiments 47-52, wherein the treatment comprises computed tomography (CT) imaging, positron emission tomography (PET) imaging, and / or magnetic resonance imaging (MRI) of the subject.Embodiment 64: the method of any one of embodiments 47-52, wherein the treatment or medicament comprises a cell therapy.Embodiment 65: the method of any one of embodiments 47-52, wherein the treatment comprises a CAR-T cell therapy.Embodiment 66: the method of embodiment 65, wherein the CAR-T cell therapy is an anti-CD19 cell therapy.Embodiment 67: the method of embodiment 64, wherein the cell therapy comprises genetically engineered cells.Embodiment 68: the method of embodiment 67, wherein the genetically engineered cells are T cells.Attorney Docket No. 58626-723601Embodiment 69: the method of embodiment 68, wherein the genetically engineered T cells comprise a Chimeric Antigen Receptor (CAR).Embodiment 70: the method of embodiment 69, wherein the CAR specifically binds to an antigen associated with a disease or condition and / or is expressed by cells associated with a disease or condition.Embodiment 71: the method of embodiment 69, wherein the CAR specifically binds to two antigens associated with a disease or condition and / or is expressed by cells associated with a disease or condition.Embodiment 72: the method of embodiment 70 or 71, wherein the antigen is selected from the group consisting of: 5T4, 8H9, avb6 integrin, B7-H6, B cell maturation antigen (BCMA), CA9, a cancer-testes antigen, carbonic anhydrase 9 (CAIX), CCL-1, CD 19, CD20, CD22, CEA, hepatitis B surface antigen, CD23, CD24, CD30, CD33, CD38, CD44, CD44v6, CD44v7 / 8, CD 123, CD 138, CD171, carcinoembryonic antigen (CEA), CE7, a cyclin, cyclin A2, c-Met, dual antigen, EGFR, epithelial glycoprotein 2 (EPG-2), epithelial glycoprotein 40 (EPG-40), EPHa2, ephrinB2, erb-B2, erb-B3, erb-B4, erbB dimers, EGFR vIII, estrogen receptor, Fetal AchR, folate receptor alpha, folate binding protein (FBP), FCRL5, FCRH5, fetal acetylcholine receptor, G250 / CAIX, GD2, GD3, gplOO, Her2 / neu (receptor tyrosine kinase erbB2), HMW-MAA, IL- 22R-alpha, IL- 13 receptor alpha 2 (IL-13Ra2), kinase insert domain receptor (kdr), kappa light chain, Lewis Y, Ll-cell adhesion molecule (Ll-CAM), Melanoma-associated antigen (MAGE)- Al, MAGE-A3, MAGE-A6, MART-1, mesothelin, murine CMV, mucin 1 (MUC1), MUC16, NCAM, NKG2D, NKG2D ligands, NY-ESO-1, O-acetylated GD2 (OGD2), oncofetal antigen, Preferentially expressed antigen of melanoma (PRAME), PSCA, progesterone receptor, survivin, ROR1, TAG72, tEGFR, VEGF receptors, BAFF-R, VEGF-R2, Wilms Tumor 1 (WT-1), and a pathogen-specific antigen.Embodiment 73: the method of embodiment 70 or 71, wherein the antigen is CD 19. Embodiment 74: the method of embodiment 70 or 71, wherein the antigen is CD20. Embodiment 75: the method of embodiment 69, wherein the CAR comprises an extracellular antigen-recognition domain that specifically binds to the antigen and an intracellular signaling domain comprising an IT AM.Embodiment 76: the method of embodiment 75, wherein the intracellular signaling domain comprises an intracellular domain of a CD3-zeta (CD3Q chain.Embodiment 77: the method of embodiment 69, wherein the CAR further comprises a costimulatory signaling region.Attorney Docket No. 58626-723601Embodiment 78: the method of embodiment 77, wherein the costimulatory signaling region comprises a signaling domain of CD28 or 4-1BB.Embodiment 79: the method of embodiment 78, wherein the signaling domain is a domain of 4- 1BB.Embodiment 80: the method of embodiment 78, wherein the T cells are CD4+.Embodiment 81: the method of embodiment 78, wherein the T cells are CD4+ or CD8+.Embodiment 82: the method of embodiment 78, wherein the T cells are primary T cells obtained from a subject.Embodiment 83: the method of embodiment 67, wherein the genetically engineered cells are autologous to the subject.Embodiment 84: the method of embodiment 67, wherein the genetically engineered cells are allogeneic to the subject.Embodiment 85: the method of any one of embodiments 1-84, wherein the subject is a human. Embodiment 86: the method of embodiment 85, wherein the subject has Stage I / II disease. Embodiment 87: the method of embodiment 85, wherein the subject has Stage III / IV disease. Embodiment 88: the method of any one of embodiments 1-85, wherein the subject has Minimal Residual Disease (MRD).Embodiment 89: the method of any one of embodiments 1-88, wherein the subject is refractory to treatment with one or more prior therapies for the cancer.Embodiment 90: the method of any one of embodiments 1-89, wherein the subject achieved an insufficient response to one or more prior therapies for the cancer.Embodiment 91: the method of any one of embodiments 1-88, wherein the subject achieves a durable response to the treatment or the medicament.Embodiment 92: the method of embodiment 91, wherein the durable response is defined as an absence of relapse or remission of the cancer for up to 3 months 6 months, or 12 months. Embodiment 93: the method of any one of embodiments 1-92, wherein the subject has a higher rate of survival.Embodiment 94: the method of any one of embodiments 1-93, wherein the subject’s disease baseline characteristics are determined.Embodiment 95: the method of embodiment 94, wherein the baseline characteristics comprise international prognosis index (IP I) score, serum lactate dehydrogenase (LDH), or sum of the product of diameters (SPD), disease stage, or any combination of the above.Embodiment 96: the method of embodiment 94 or embodiment 95, wherein the baseline characteristics are defined by Lugano 2014 criteria.Attorney Docket No. 58626-723601Embodiment 97: the method of any of embodiments 95-96, wherein the subject is classified under an international prognosis index (ZPI) score.Embodiment 98: the method of any of embodiments 95-97, wherein the subject is classified as low, low-intermediate, or high-intermediate risk on the IPI score.Embodiment 99: the method of any of embodiments 95-98, wherein the subject has no risk factor or one risk factor and is considered to be in an IPI low risk group.Embodiment 100: the method of any of embodiments 95-98, wherein the subject has two risk factors and is considered to be in an IPI low-intermediate risk group.Embodiment 101: the method of any of embodiments 95-98, wherein the subject has three risk factors and is considered to be in an IPI high-intermediate risk group.Embodiment 102: the method of any of embodiments 95-101, wherein the subject has a higher or lower IPI score.Embodiment 103: the method of any one of embodiments 1-102, wherein a volumetric measure of tumor burden of the subject is measured.Embodiment 104: the method of embodiment 103, wherein the volumetric measure of tumor burden of the subject is a sum of products of diameter (SPD).Embodiment 105: the method of embodiment 103, wherein the volumetric measure of tumor burden is measured using computed tomography (CT), positron emission tomography (PET), and / or magnetic resonance imaging (MRI) of the subject.Embodiment 106: the method of embodiment 104, wherein an SPD threshold value is or is about 30 per cm2, is or is about 40 per cm2, is or is about 50 per cm2, is or is about 60 per cm2, or is or is about 70 per cm2.Embodiment 107: the method of any one of embodiments 1-106, wherein a level of an inflammatory marker is measured in the subject.Embodiment 108: the method of embodiment 107, wherein the level of the inflammatory marker is or is about 300 units per liter, is or is about 400 units per liter, is or is about 500 units per liter or is or is about 600 units per liter.Embodiment 109: the method of embodiment 107, wherein the inflammatory marker is lactate dehydrogenase (LDH).Embodiment 110: the method of embodiment 1-109, wherein the determining in (d) has an error rate of no more than 0.2 parts per million.Embodiment 111: A system comprising one or more computer processors and computer memory coupled thereto, wherein the computer memory comprises machine-executable code that, upon execution by the one or more computer processors, implements a method comprising:Attorney Docket No. 58626-723601(a) sequencing or having sequenced a plurality of cell-free DNA molecules obtained or derived from a subject to produce a plurality of sequence reads; aligning the plurality of sequence reads to a reference genomic sequence, to produce a plurality of aligned sequence reads; (b) processing the plurality of aligned sequence reads, to detect a presence of a somatic mutation in one or more cell-free DNA molecules from among the plurality of cell-free DNA molecules, wherein the somatic mutation comprises a base-change mutation selected from the group consisting of: an adenine-to-thymine (A>T) mutation, an adenine-to-cytosine (A>C) mutation, an adenine-to-guanine (A>G) mutation, a thymine-to- adenine (T>A) mutation, a thymine-to-cytosine (T>C) mutation, a thymine-to-guanine (T>G) mutation, a cytosine-to-adenine (C>A) mutation, a cytosine-to-thymine (OT) mutation, a cytosine-to-guanine (OG) mutation, a guanine-to-adenine (G>A) mutation, a guanine-to- thymine (G>T) mutation, and a guanine-to-cytosine (G>C) mutation, wherein the processing in (c) comprises filtering the plurality of aligned sequence reads based at least in part on an expected error rate of detection of the base-change mutation in individual aligned sequence reads, thereby generating a set of filtered sequence reads; and (c) processing the set of filtered sequence reads to determine whether at least one of the set of filtered sequence reads is mutationcontaining, wherein the determining in (d) has an error rate of no more than 0.5 parts per million. Embodiment 112: A non-transitory computer-readable medium comprising machine-executable code that, upon execution by one or more computer processors, implements a method comprising: (a) sequencing a plurality of cell-free DNA molecules obtained or derived from a subject to produce a plurality of sequence reads; (b) aligning the plurality of sequence reads to a reference genomic sequence, to produce a plurality of aligned sequence reads; (c) processing the plurality of aligned sequence reads, to detect a presence of a somatic mutation in one or more cell-free DNA molecules from among the plurality of cell-free DNA molecules, wherein the somatic mutation comprises a base-change mutation selected from the group consisting of: an adenine-to-thymine (A>T) mutation, an adenine-to-cytosine (A>C) mutation, an adenine-to- guanine (A>G) mutation, a thymine-to-adenine (T>A) mutation, a thymine-to-cytosine (T>C) mutation, a thymine-to-guanine (T>G) mutation, a cytosine-to-adenine (C>A) mutation, a cytosine-to-thymine (C>T) mutation, a cytosine-to-guanine (OG) mutation, a guanine-to- adenine (G>A) mutation, a guanine-to-thymine (G>T) mutation, and a guanine-to-cytosine (G>C) mutation, wherein the processing in (c) comprises filtering the plurality of aligned sequence reads based at least in part on an expected error rate of detection of the base-change mutation in individual aligned sequence reads, thereby generating a set of filtered sequence reads; and (d) processing the set of filtered sequence reads to determine whether at least one ofAttorney Docket No. 58626-723601 the set of filtered sequence reads is mutation-containing wherein the determining in (d) has an error rate of no more than 0.5 parts per million.Embodiment 113: a method for identifying a somatic mutation, the method comprising: (a) obtaining a plurality of sequence reads for each of a plurality of nucleic acid molecules from a biological sample of a subject; (b) mapping the plurality of sequence reads for each of the plurality of nucleic acid molecules to a reference genomic sequence; (c) grouping the plurality of sequence reads into groups of sequence reads that map to identical start and stop positions on the reference genomic sequence, thereby forming a plurality of families; (d) further grouping sequence reads of each family into a plurality of subgroups comprising a first subgroup and a second subgroup, each subgroup comprising a plurality of sequence reads; wherein: (i) the sequence reads of the first subgroup comprise a first nucleotide at a first position; and (ii) the sequence reads of the second subgroup comprise a second nucleotide at the first position, wherein the second nucleotide differs from the first nucleotide; (e) collapsing the sequence reads of the first subgroup to form a first consensus sequence corresponding with the sequence of a first nucleic acid molecule of the plurality of nucleic acid molecules; (f) collapsing the sequence reads of the second subgroup to form a second consensus sequence corresponding with the sequence for a second nucleic acid molecule from the plurality of nucleic acid molecules, wherein the first nucleic acid molecule differs in sequence from the second nucleic acid molecule, and (g) identifying the first nucleotide at the first position as a somatic mutation comprising a base-change selected from the group consisting of: an adenine-to-thymine (A>T) mutation, an adenine-to-cytosine (A>C) mutation, an adenine-to-guanine (A>G) mutation, a thymine-to-adenine (T>A) mutation, a thymine-to-cytosine (T>C) mutation, a thymine-to- guanine (T>G) mutation, a cytosine-to-adenine (OA) mutation, a cytosine-to-thymine (C>T) mutation, a cytosine-to-guanine (OG) mutation, a guanine-to-adenine (G>A) mutation, a guanine-to-thymine (G>T) mutation, and a guanine-to-cytosine (G>C) mutation, wherein the identifying comprises an error rate of no more than 0.5 parts per million.EXAMPLESExample 1: Identification of multiple molecules after error-correction
[0151] In this example, two different methods are performed to demonstrate advantages of certain error correction strategies after grouping sequence reads by start and stop positions.
[0152] In both methods, a plurality of sequence reads from cell-free DNA from a subject are obtained. The sequence reads include reads derived from different (e.g., non-identical) DNA molecules that share the same start and stop positions when aligned with a reference genomicAttorney Docket No. 58626-723601 sequence, such as a reference human genome. The sequence reads were obtained by sequencing a cell-free DNA-containing biological sample of a human subject. As shown in FIG. 2 A, the nucleic acid molecules from the human subject include two cell-free nucleic molecules with the same start position and end (e.g., stop) position, but with different single nucleotide variants (SNV) at a particular nucleotide position (the “first position”). Specifically, FIG. 2A shows the two molecules having different SNVs, with one cell-free nucleic acid molecule comprising a T at a particular position and another cell-free nucleic acid molecule comprising a C at the same position. The biological sample comprising the cell-free nucleic acids is processed to generate sequence reads amenable for further analysis. Processing the nucleic acids to generate sequence reads may include a sample preparation step (e.g., to add adapters for sequence; adapters depicted in dashed lines), a PCR step, and / or a sequencing step (e.g., via next-generation sequencing technologies, such as sequencing by synthesis). As a result of such processing steps, more than one sequence read may be generated from one or more molecules from the biological sample, and the resulting sequence reads may contain one or more errors. For instance, as shown in FIG. 2B, the top three sequence reads may be derived from a first molecule having a “T” at the first position, and the bottom four sequence reads may be derived from a second molecule having a “C” at the first position. Further, as shown in FIG. 2B, the top sequence read includes an erroneous “A” at a second position, and the second-from-the-bottom sequence read includes an erroneous “G” at a third position. The generated sequence reads are mapped to a reference sequence, and a start and stop position for each sequence read (corresponding with the first nucleotide and last nucleotide of the corresponding nucleic acid molecule) is identified. For instance, FIG. 2B shows six sequence reads, each with the same start and stop positions that constitute a single family, even though this family includes non-identical sequence reads. For instance, as shown in step 2, the family includes some sequence reads that have a “T” at the first position (see top three nucleic acids) and some sequence reads that have a “C” at a first position. Note that these differences are a result of non-identical molecules in the biological sample of the subject. The family also includes some sequence reads that include errors that arise from sample processing (see the “A” in the top sequence read) and the “G” in the sequence read shown second from the bottom). Sequence reads are grouped into families based on sequences having the same start and stop positions.
[0153] In the first method, a single consensus sequence is identified from the entire family, as shown in FIG. 2C. As depicted in FIG. 2C, the single consensus sequence is not reflective of the biological reality of two non-identical sequences having the same start and stop positions, but differing due a SNV at position 1.Attorney Docket No. 58626-723601
[0154] In a second method, further error correction analysis is performed on the family of sequence reads having an identical start and stop position. In this second method, two different molecules are identified based on the presence of recurrent SNVs in two different subgroups of sequence reads of the group of sequence reads having identical start and stop positions (FIG. 2D). More specifically, sequence reads of a first subgroup have a first nucleotide (“T”) at a first position (e.g., the three top sequence reads having a “T” at the first position) are collapsed to form a first consensus sequence corresponding with the sequence of the first nucleic acid molecule in the biological sample that also had the “T” at the first position. Such collapse may involve suppression of processing errors (e.g., sample handling, PCR, or sequencing errors), such as the error in the top nucleic acid molecule including an erroneous “A” as shown in FIG. 2B, that are present in the sequence reads at a level of recurrence that is lower than the level of recurrence of the variant (“T”) at the first position. Similarly, the second subgroup of sequence reads having a second nucleotide at the first position (“C”)are collapsed to form a second consensus sequence corresponding with the sequence of the second nucleic acid molecule in the biological sample that also had the “C” at the first position. Such collapse may involve suppression of processing errors (e.g., sample handling, PCR or sequencing errors), such as the error in the second-to-the-bottom sequence read that include an erroneous “G” at the third position. Thus, the two different nucleic acid molecules from the biological sample that had identical start and stop positions are identified and distinguished from one another.Example 2: Analysis of background error rate with and without UMIs
[0155] In this example, three samples were prepared and analyzed to determine the background error rate of samples.
[0156] Tumor DNA was extracted from tumor FFPE blocks from three early-stage non-small cell lung cancer patients using a QIAGEN kit. Paired normal DNA was extracted from PBMCs from the same patients using a QIAGEN kit. The tumor and paired normal DNA was subjected to whole genome sequencing, which was analyzed to determine patient-specific DNA mutations to monitor in circulating tumor DNA from each patient. Three patient-specific oligonucleotide panels were designed and synthesized based on the patient-specific mutations identified.
[0157] From each of the three non-small cell lung cancer patients, pretreatment plasma DNA was obtained. The plasma was separated from the buffy coat and red blood cells and stored at - 80°C until cell-free DNA isolation. Cell-free DNA isolation was performed using either the QIAamp Circulating Nucleic Acid Kit or the QIAsymphony DSP Circulating Nucleic Acid Kit in accordance with the manufacturer’s instructions (QIAGEN).Attorney Docket No. 58626-723601
[0158] A dilution series from each of the three plasma cell-free DNA samples from each lung cancer patient was generated by adding patient cell-free DNA into healthy donor cell-free DNA. The dilution series ranged from 1 part of tumor-specific cell-free DNA in 10,000 to 1 part in 20,000,000. Cell-free DNA from three healthy donors was also analyzed without adding in cell- free DNA from each matched patient. For each healthy donor, each of the three lung cancer patient’s specific set of mutations were considered, for a total of nine data points.
[0159] Sequencing libraries were prepared using the KAPA HyperPrep Kit (Roche). Samples were prepared using adapter-sequences that contained UMIs (Chabon et al, Nature 2020). Approximately 80 ng of input DNA was sequenced per sample. Sequencing was performed using 2x150 paired end reads using the Illumina Nova X+ system.
[0160] Sample fastq files that were generated for each of the three samples during sequencing were pre-processed using fastp and aligned to the GRCh37 reference genome using bwa-mem2. These reads were processed in three separate ways: 1) without UMIs (i.e., removing the UMI sequences in silico and processing all reads based on the start and end positions without the use of UMIs); 2) with UMIs (i.e., considering the UMIs, and processing the reads through previously established methods for de-duplication and error suppression as described in Chabon et al, Nature 2020); and 3) via haplotype deduplication (in which the UMI sequences are removed in silico and then the resulting reads are processed as described herein). More specifically, for haplotype deduplication, UMIs were removed from each read, and reads were grouped based on their start and end positions to form a read-group. Each read-group was then analyzed for the number of supporting reads for a given variant within the read-group. A new read-group was generated from within the original read-group when 5 or more supporting reads within a single read-group contained the same variant.
[0161] Using these three separate data processing conditions for each sample, the amount of tumor DNA contained within each sample was determined by measuring the number of molecules with one or more tumor-specific mutations as determined by the initial tumor and paired normal sequencing, divided by the total number of sequence reads considered. Both the “no UMI” and “UMI” conditions considered all sequencing molecules as described in prior work (Newman et al, Nature Biotechnology 2016 and Chabon et al, Nature 2020). For the haplotype deduplication condition, only sequence reads that met a pre-specified set of criteria were considered, including having a pre-specified number of supporting read-group members, and having the locus of interest a sufficient distance from the edge of the molecule. For mutations A>C / T>G, C>G / G>C, and A>T / T>A, at least 3 supporting read-group members were required. For mutations A>G / T>C, at least 4 supporting read-group members were required. For mutationsAttomey Docket No. 58626-723601OA / G>T, and C>T / G>A, at least 5 supporting read-group members were required. For mutations A>C / T>G, C>G / G>C, A>T / T>A, A>G / T>C, and OA / G>1, only sequence reads with the mutation at least 10 nucleotides from the end of the nucleic acid molecule were considered. For mutation C>T / G>A, only sequence reads with the mutation at least 25 nucleotides from the end of the nucleic acid molecule were considered. The one or more tumorspecific mutations were assessed in a subset of all consensus sequence reads. The one or more tumor-specific mutations were assessed in a subset of all consensus sequence reads with a lower error rate compared to the error rate when all sequence reads were considered. Additional filters were applied to all sequence reads: 1) only duplex molecules were considered, 2) only sequences with a read mate present and duplex (same strand different ends) were considered, and 3) low quality reads (more than 5% of base pairs with phred score lower than 30) were removed from analysis.
[0162] Background error was determined for each sample, as shown in FIG. 4. The “Background Error in controls” describes the amount of sequence reads that show evidence of the patientspecific tumor DNA molecules in the nine samples from healthy donor cell-free DNA (three samples for each of the three lung cancer patient lists). Additionally, the observed fraction relative to the expected fraction was analyzed for each sample, as shown in FIG. 5. “Expected fraction” represents the pipetted, intended concentration of tumor DNA from the dilution experiment as described above, and the “observed fraction” represents the sequencing measurement. The small dots on the plots denote each experimental replicate, whereas the large dot represents the mean of the replicates.
[0163] An additional analysis was performed on the three samples to further demonstrate the performance of the haplotype deduplication analysis relative to the analysis with UMIs. The sequence reads from each of the conditions described above were again processed in three separate ways: 1) without UMIs; 2) with UMIs and 3) via haplotype deduplication, as described above. However, in this analysis, for both haplotype deduplication, and UMI mediated deduplication, only sequence reads with a pre-determined number of supporting fragments were considered. Specifically, only sequence reads with at least 3 read-group members were considered, where at least one read-group member from an original Watson strand, and one member from an original Crick strand, were included. This analysis was the same for both haplotype deduplication (without UMIs) and with UMI deduplication. Analysis of the sample without UMIs and without haplotype deduplication remained the same.
[0164] Background error was determined for each sample, as shown in FIG. 6. The “Background Error in controls” describes the amount of sequence reads that show evidence of the patient-Attorney Docket No. 58626-723601 specific tumor DNA molecules in the nine samples from healthy donor cell-free DNA (three samples for each of the three lung cancer patient lists). No difference was observed between the error rate with UMIs and haplotype deduplication, showing identical performance. Additionally, the observed fraction relative to the expected fraction was analyzed for each sample, as shown in FIG. 7. “Expected fraction” represents the pipetted, intended concentration of tumor DNA from the dilution experiment as described above, and the “observed fraction” represents the sequencing measurement. The small dots on the plots denote each experimental replicate, whereas the large dot represents the mean of the replicates.
[0165] The sequence reads from each of the conditions described above were again processed in three separate ways: 1) without UMIs, 2) with UMIs and with detection of low error rate nucleic acid molecules comprising SNVs, and 3) without UMIs and with detection of low error rate nucleic acid molecules comprising SNVs. Detection of low error rate nucleic acid molecules comprising SNVs involved detecting nucleic acid molecules comprising SNVs with lower predetermined error rates according to the data shown in FIG. 15. The nucleic acid molecules comprising SNVs with low error rate was based on a determination of a SNV, the number of detected molecules comprising the SNV (e.g. family size), and the location of the SNV on the nucleic acid molecule. SNVs within 25 (OT) nucleotides from the end of the nucleic acid molecule or 10 nucleotides from the end of the nucleic acid molecule for all remaining SNVs were not used for analysis. The first analysis method used start and stop positions to group sequence reads and determine a consensus sequence based on the most prevalent sequence of the sequence reads that shared a start and stop position. The second analysis method used UMIs to group sequence reads with the same UMI sequences. A SNV tumor fraction was determined based on a subset of tumor SNVs with low error rate, as described above. The third analysis method used start and stop positions to group sequence reads. The sequence reads were further separated into subgroups based on detecting sequence reads with different SNVs that had the same start and stop positions. The SNV tumor fraction considered a subset of the overall tumor SNVs with low error rate. The SNV tumor fraction for each of these conditions was plotted against the expected tumor fraction (FIG. 13 A). A similar analysis was performed based on the same three processing methods used in FIG. 13 A, but for the detection of phased variants. The data from this analysis is shown in FIG. 13B.Example 3: Differential error-rate reduction
[0166] Using systems and methods of the present disclosure, a process was developed for improved utilization of somatic mutations (e.g., SNVs, PVs) for ctDNA-MRD detection. Not all somatic variants have the same technical error profile. For example, phased variants can have anAttorney Docket No. 58626-723601 error profile below 1 part in 50 million, while SNVs without any error suppression may have an error profile of ~1 part per 10,000. Errors can come from a number of different sources, including biological variants, clonal hematopoiesis, amplification (e.g., PCR), oxidative damage, etc. This example describes systems and methods of choosing somatic mutations to track for ctDNA-MRD detection such that the chosen somatic mutations have low error profiles (e.g., error profiles of less than 1 part in 5 million or 1 part in 50 million). Such error profiles may be similar to or lower than the error profile of phase variants. The chosen SNV mutations can be used together with phased variants in a single assay for improved ctDNA-MRD detection.
[0167] Included in this example are factors to consider for selecting the somatic mutations (e.g., SNVs) to be considered for tracking for ctDNA-MRD detection. The motivation behind the selection of somatic mutations can be to reduce the error-rate of detection. To do so, the somatic mutations to consider can include base-change mutations or SNVs (e.g., A>T, A>C, A>G, T>A, T>C, T>G, OA, OT, OG, G>A, G>T, G>C).
[0168] Here, a biological sample comprising cell-free DNA molecules is obtained from a subject and the cell-free DNA molecules are extracted, sequenced, and analyzed to generate sequence reads. The cell-free DNA molecules include ctDNA content and are analyzed for somatic mutations (e.g., SNVs, PVs) that are indicative of ctDNA-MRD. However, not all sequence reads have the same error profile. A subset of sequence reads containing SNVs is identified to have a lower error profile - on the order of less than 1 part in 5,000,000. The subset of sequence reads containing SNVs is used, either in combination with sequence reads containing PVs or alone, for ultrasensitive detection of ctDNA and determination of ctDNA-MRD. The subset of sequence reads containing SNVs that is chosen for use in ctDNA-MRD detection is based on one or more factors.
[0169] A factor to consider when selecting sequence reads for ctDNA-MRD determination is the specific base change of the SNV (e.g., A>T, A>C, A>G, T>A, T>C, T>G, OA, OT, OG, G>A, G>T, G>C) in the sequencing read. Not all base-change mutations have the same error profile. For example, OT and G>A mutations have a higher background error rate than OG or G>C mutations.
[0170] FIG. 8 shows the background error rates in cfDNA from sequencing where both the Watson and Crick strands are identified for various specific base-change mutations. More specifically, FIG. 8 shows background error rates by mutation type for 18 cfDNA samples across 11 donors captured with two different panels. Samples were sequenced using Illumina SBS sequencing, fastq files were pre-processed using fastp, mapped to the hgl9 human genome using BWA MEM, and molecules were deduplicated using a previously described algorithm. For thisAttorney Docket No. 58626-723601 analysis, only fragments with family size of three or more and positions at least 20 basepairs from the edge of the molecule were considered. Genomic positions with depth lower than 300 were also skipped and background error was considered to be any allelic fraction below 5%. The sequence reads with base changes (A>T, OG, A>C), as shown in FIG. 8, exhibited a lower error rate when compared to sequence reads with other base changes (C>T, OA, A>G). FIG. 16 shows a similar trend and lists the somatic mutations corresponding to both stands of DNA in a duplex context. Therefore, instead of considering all sequence reads equally, the analysis to select a sequencing read for ctDNA-MRD detection is influenced by the specific base change of the SNV on the sequencing read. A base change mutation with a lower background error rate can be selected over a base-change mutation with a higher background error rate to provide a more accurate ctDNA-MRD determination.
[0171] Other factors to consider when selecting sequence reads for ctDNA-MRD determination are (1) the number of copies of each molecule seen, (2) whether both Watson and Crick strands are observed, (3) the number of independent copies of Watson and Crick strands that are observed, and / or (4) the location of the candidate mutation (e.g., SNV) in the sequencing read. These factors contribute to the error profile of a sequencing read.
[0172] One possible way to take these factors into account is to differentially require different number of duplicates for each base change to enable an “equally low error rate.” For example, as shown in FIG. 9, the error rates of each base change were observed to be different depending on the specific base change, the number of duplicates (family size), and if both Watson and Crick strands are observed. For example, as presented in FIG. 9, requiring only 4 copies (e.g., family size is 4) of a sequencing read resulted in an error rate close to 1 : le-6 when the base change is T>G or G>C for both cases of when both Watson and Crick strands are observed and when only the Watson or Crick strand is observed. However, to reach a similar error rate (e.g., 1 : le-6) when the base change is C>T or G>A, 10 copies and observation of both Watson and Crick strands were needed. These trends are also shown in FIG. 15, which was generated from the same dataset that was used to generate FIG. 9.
[0173] Similarly, the error rate can decrease from the edge of a sequencing read. For example, the first few bases and the last few bases of the sequencing read can have higher error rates. So, for some sequence reads with SNVs of certain base changes (e.g., those with lower error rates), the first 5 bp and / or the last 5 bp of the read can be trimmed while for other some sequence reads with SNVs of other base changes (e.g., those with higher error rates), the first 20 base pairs and / or last base pairs of the sequencing read can be trimmed, such that the resulting sequenceAttorney Docket No. 58626-723601 reads selected for use to determine ctDNA-MRD exhibit a uniformly low error rate across all variant types that are considered.Example 4: Dilution series analysis using start and stop positions and subgroup analysis
[0174] In this example, cfDNA samples were prepared with different concentrations of cfDNA from a lung cancer patient diluted into cfDNA from a normal subject to determine the detection limit of identifying SNVs and PVs in the tumor fraction using different analysis methods.
[0175] Blood samples were obtained from a subject with lung cancer. Cell-free DNA was extracted from the blood samples and added to cfDNA from a healthy subject. The cfDNA from the lung cancer patient was diluted at expected concentrations from 1 : 10,000 to 1 :20,000,000 (v / v). A subset of each sample with the different concentrations of cfDNA from the lung cancer patient was processed and labeled with UMIs. A sequencing library was generated and analyzed using Illumina sequencing. Another subset of each sample with the different concentrations of cfDNA from the lung cancer patient was processed to prepare a sequencing library. UMIs, if added to cfDNA of this other subset of each sample, were ignored during analysis. Sequence reads were obtained for both subsets at each concentration and analyzed. The sequence reads without use of UMIs were analyzed by grouping sequencing reads with the same start and stop positions. The sequencing reads were further sub-grouped based on identification of cfDNA molecules with the same SNVs within the group. The sequence reads with UMIs that were used were analyzed by grouping sequence reads based on the presence of the same UMIs. A plot of the expected dilution relative to the estimated actual fraction based on the sequence reads is shown in FIG. 14A. A similar analysis was performed for detecting PVs for both the sequence reads with and without use of UMIs. The results of that analysis are shown in FIG. 14B. Both sequence reads analyzed based on UMIs and without use of UMIs show similar concordance down to below IxlO'7dilution of cfDNA from the lung cancer subject into cfDNA from the healthy subject.
[0176] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
Attorney Docket No. 58626-723601WHAT IS CLAIMED IS:
1. A method for determining, from sequence reads, a sequence for a first nucleic acid molecule and a sequence for second nucleic acid molecule, wherein the first nucleic acid molecule and the second nucleic acid molecule have identical start and stop positions, the method comprising:(a) obtaining a plurality of sequence reads for each of a plurality of nucleic acid molecules from a biological sample of a subject;(b) mapping the plurality of sequence reads for each of the plurality of nucleic acid molecules to a reference genomic sequence;(c) grouping the plurality of sequence reads into groups of sequence reads that map to identical start and stop positions on the reference genomic sequence, thereby forming a plurality of families;(d) further grouping sequence reads of each family into a plurality of subgroups comprising a first subgroup and a second subgroup, each subgroup comprising a plurality of sequence reads; wherein:(i) the sequence reads of the first subgroup comprise a first nucleotide at a first position; and(ii) the sequence reads of the second subgroup comprise a second nucleotide at the first position, wherein the second nucleotide differs from the first nucleotide;(e) collapsing the sequence reads of the first subgroup to form a first consensus sequence corresponding with the sequence of a first nucleic acid molecule of the plurality of nucleic acid molecules; and(f) collapsing the sequence reads of the second subgroup to form a second consensus sequence corresponding with the sequence for a second nucleic acid molecule from the plurality of nucleic acid molecules, wherein the first nucleic acid molecule differs in sequence from the second nucleic acid molecule.
2. The method of claim 1, wherein the first subgroup comprises a plurality of non-identical sequence reads, wherein the non-identical sequence reads of the first subgroup have differing levels of recurrence of a nucleotide at a second position.
3. The method of claim 2, wherein the first nucleotide at the first position is present in the plurality of sequence reads at a level of recurrence that is higher than the level of recurrence of the nucleotide at the second position.Attorney Docket No. 58626-7236014. The method of claim 2 or claim 3, wherein the differing levels of recurrence of the nucleotide at the second position arise from one or more PCR errors.
5. The method of any one of claims 1-4, wherein the second subgroup comprises a plurality of non-identical sequence reads, wherein the plurality of non-identical sequence reads of the second subgroup have differing levels of recurrence of a nucleotide at a third position.
6. The method of claim 5, wherein the first nucleotide at the first position is present in the plurality of sequence reads at a level of recurrence that is higher than the level of recurrence of the nucleotide at the third position.
7. The method of claim 5 or claim 6, wherein the differing levels of recurrence of the nucleotide at the third position arise from one or more PCR errors.
8. The method of any one of claims 1-7, wherein the plurality of nucleic acid molecules from the biological sample of the subject are cell-free nucleic acid molecules.
9. The method of claim 8, wherein the cell-free nucleic acid molecules comprise cell-free DNA molecules.
10. The method of any one of claims 1-9, wherein the reference genomic sequence is at least 100 kb in length.
11. The method of any one of claims 1-9, wherein the reference genomic sequence is at least 1 Mb in length.
12. The method of any one of claims 1-11, wherein the reference genomic sequence is a human exome.
13. The method of any one of claims 1-11, wherein the reference genomic sequence is from a human genome (e.g., derived from a non-tumor control sample of the subject).
14. The method of claim 13, wherein the reference genomic sequence is selected from the group consisting of the hgl9 human genome, the hgl8 genome, the hgl7 genome, the hgl6 genome, and the hg38 genome.
15. The method of any one of claims 1-14, wherein the sequence reads of the first subgroup comprising the first nucleotide at the first position comprise sequence reads originally derived from both the positive and negative (e.g., Watson and Crick) strands of the reference genomic sequence.
16. The method of any one of claims 1-15, wherein the sequence reads of the second subgroup comprising the second nucleotide at the first position comprise sequence reads originally derived from both the positive and negative (e.g., Watson and Crick) strands of the reference genomic sequence.
17. The method of any one of claims 1-16, wherein the subject is a human subject.Attorney Docket No. 58626-72360118. The method of claim 17, wherein the subject is suffering from a condition.
19. The method of claim 18, wherein the condition is a cancer.
20. The method of any one of claims 1-19, wherein the biological sample comprises a blood sample, a serum sample, or a plasma sample.
21. The method of any one of claims 1-20, wherein the sequence reads were generated by nextgeneration sequencing.
22. The method of any one of claims 1-21, wherein the first nucleotide at the first position is a single-nucleotide variant.
23. The method of claim 22, wherein the sequence reads of the first subgroup further comprise a second single-nucleotide variant.
24. The method of claim 23, wherein the first nucleotide at the first position and the second single nucleotide variant are within 170 nucleotides of each other.
25. The method of claim 23 or claim 24, wherein the first nucleotide at the first position and the second single nucleotide variant are separated by at least one nucleotide.
26. The method of any one of claims 1-25, wherein the reference genomic sequence is at least 10 kilobases in length.
27. The method of any one of claims 1-25, wherein the reference genomic sequence is at least 50 kilobases in length.
28. The method of any one of claims 1-25, wherein the reference genomic sequence is at least 100 kilobases in length.
29. The method of any one of claims 1-28, wherein the plurality of sequence reads comprises at least 10,000 sequence reads.
30. The method of any one of claims 1-28, wherein the plurality of sequence reads comprises at least 50,000 sequence reads.
31. The method of any one of claims 1-28, wherein the plurality of sequence reads comprises at least 100,000 sequence reads.
32. The method of any one of claims 1-31, wherein the plurality of nucleic acid molecules comprises at least 1,000 nucleic acid molecules.
33. The method of any one of claims 1-31, wherein the plurality of nucleic acid molecules comprises at least 5,000 nucleic acid molecules.
34. The method of any one of claims 1-31, wherein the plurality of nucleic acid molecules comprises at least 10,000 nucleic acid molecules.
35. The method of any one of claims 1-34, wherein the first consensus sequence comprises an error rate of less than or equal to 0.5 parts per million (ppm).Attorney Docket No. 58626-72360136. The method of any one of claims 1-34, wherein the first consensus sequence comprises an error rate of less than or equal to 0.1 parts per million (ppm).
37. The method of any one of claims 1-34, wherein the second consensus sequence comprises an error rate of less than or equal to 0.5 parts per million (ppm).
38. The method of any one of claims 1-34, wherein the second consensus sequence comprises an error rate of less than or equal to 0.05 parts per million (ppm).
39. The method of any one of claims 1-38, wherein (b) comprises performing a Burrows- Wheeler alignment (BWA) algorithm.
40. The method of any one of claims 1-38, wherein (b) comprises performing a Bowtie alignment algorithm.
41. The method of any one of claims 1-40, wherein the method does not comprise analysis of a molecular barcode of a molecular barcode-containing adapter.
42. The method of any one of claims 1-40, wherein the plurality of sequence reads for each of the plurality of nucleic acid molecules are obtained from sequencing that does not utilize a molecular barcode-containing adapter.
43. The method of claim 42, wherein the sequencing that does not utilize a molecular barcodecontaining adapter does utilize a sample barcode-containing adapter.
44. The method of claim 42, wherein the sequencing that does not utilize a molecular barcodecontaining adapter does not utilize a sample barcode-containing adapter.
45. The method of claim 42, wherein the sequencing does not utilize an adapter.
46. The methods of any one of claims 1-45, wherein the subject previously had a condition.
47. The method of any claim 46, wherein the subject was previously treated for the condition.
48. The method of any one of claims 46-47, wherein the condition is a cancer.
49. The method of claim 48, wherein the cancer is a lymphoma.
50. The method of any one of claims 47-49, wherein the subject was previously treated with a chemotherapeutic agent.
51. The method of claims 47-49, wherein the subject was previously treated with a chimeric antigen receptor T cell therapy.
52. A method for treating a condition in a subject, the method comprising:(a) obtaining a plurality of sequence reads for each of a plurality of nucleic acid molecules from a biological sample of the subject;(b) mapping the plurality of sequence reads for each of the plurality of nucleic acid molecules to a reference genomic sequence;Attorney Docket No. 58626-723601(c) grouping the plurality of sequence reads into groups of sequence reads that map to identical start and stop positions on the reference genomic sequence, thereby forming a plurality of families;(d) further grouping sequence reads of each family into a plurality of subgroups comprising a first subgroup and a second subgroup, each subgroup comprising a plurality of sequence reads; wherein:(i) the sequence reads of the first subgroup comprise a first nucleotide at a first position; and(ii) the sequence reads of the second subgroup comprise a second nucleotide at the first position, wherein the second nucleotide differs from the first nucleotide;(e) collapsing the sequence reads of the first subgroup to form a first consensus sequence corresponding with the sequence of a first nucleic acid molecule of the plurality of nucleic acid molecules;(f) collapsing the sequence reads of the second subgroup to form a second consensus sequence corresponding with the sequence for a second nucleic acid molecule from the plurality of nucleic acid molecules, wherein the first nucleic acid molecule differs in sequence from the second nucleic acid molecule, and(g) identifying the first nucleic acid molecule as mutation containing based on the first consensus sequence; and(h) administering an effective amount of a therapeutic agent, based at least on part on (g).
53. The method of claim 52, wherein the condition is a cancer.
54. The method of any one of claims 1-49, wherein the method is implemented via a computer or computer system.
55. A computer-implemented system comprising: a computing device comprising at least one processor, an operating system configured to perform executable instructions, a memory, and a computer program including instructions executable by the computing device for performing the method of any one of claims 1-49.
56. A non-transitory computer-readable storage media encoded with a computer program including instructions executable by a processor for performing the method of any one of claims 1-49.
57. A system comprising one or more computer processors and computer memory coupled thereto, wherein the computer memory comprises machine-executable code that, uponAttorney Docket No. 58626-723601 execution by the one or more computer processors, implements the method of any one of claims 1-49.
58. A non-transitory computer-readable medium comprising machine-executable code that, upon execution by one or more computer processors, implements the method of any one of claims 1-49.