Non-invasive in vitro method for somatic mutation detection
By using a non-invasive in vitro method involving targeted enrichment and sequencing of TAC oligonucleotide pools in blood samples, the time delay and insufficient sample size issues in existing technologies for early cancer detection and MRD detection are resolved, achieving high sensitivity and high specificity for early cancer detection and MRD monitoring.
Patent Information
- Application Number
- CN202480042108.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-17
- Filing Date
- 2024-05-02
- Publication Date
- 2026-01-16
AI Technical Summary
Existing methods for early cancer detection and minimal residual disease (MRD) detection rely on imaging scans and tissue biopsies, which lead to problems such as detection time delays and insufficient samples. Furthermore, existing liquid biopsy methods lack sensitivity and specificity, making it difficult to accurately detect trace genetic abnormalities.
Using a non-invasive in vitro method, a nucleic acid sequencing library is prepared from blood samples. The TAC oligonucleotide pool is used for targeted enrichment and sequencing to isolate and analyze circulating tumor DNA in plasma. Combined with reference sequence alignment, somatic cell mutations are detected, achieving high-accuracy and low-cost MRD monitoring and early cancer detection.
It achieves highly sensitive and specific early cancer detection and MRD monitoring without tissue biopsy, reducing detection time and cost, and improving detection accuracy and reliability.
Smart Images

Figure CN121358876A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention is in the field of biology, medicine and chemistry, in particular in the field of molecular biology, more in particular in the field of molecular diagnostics. The present invention is in particular in the field of in vitro diagnostics using cell-free nucleic acids and DNA. The present invention is also in the field of diagnostic kits. BACKGROUND
[0002] Early detection of cancer and quantification of the risk of cancer recurrence and metastasis in patients who are under treatment or in remission remains one of the main concerns of patients and medical professionals in the field of oncology. Many of the current methods for detecting early stages of cancer or recurrence and metastasis rely heavily on imaging scans (at which point the tumor needs to grow to a sufficient size to be detected in the scan), thus resulting in a loss of critical treatment time.
[0003] Minimal Residual Disease (MRD) refers to the small amount of cancer cells that remain in a patient’s body during treatment or after curative treatment. Evidence suggests that MRD is an important factor leading to relapse or metastasis, and is therefore crucial for the prognosis and management of multiple cancer types. MRD detection allows the detection of molecular relapse of the disease before clinical symptoms are manifested, thus having the advantage of early intervention.
[0004] The discovery of free fetal DNA (ffDNA) in the maternal circulation was a milestone in the development of non-invasive prenatal testing (NIPT) for chromosomal abnormalities and opened new possibilities in the clinical setting (PMID 9529358). However, the direct analysis of limited amounts of ffDNA in the presence of a large amount of maternal DNA is a major challenge for NIPT of chromosomal abnormalities. The application of next-generation sequencing (NGS) technologies in the development of NIPT revolutionized the field. In 2008, two independent research groups demonstrated that NIPT of trisomy 21 could be achieved using next-generation massively parallel shotgun sequencing (MPSS) (PMID 18945714, PMID 18838674). The new era of NIPT of chromosomal abnormalities opened new possibilities for the application of these technologies in clinical practice. Biotechnology companies that are partially or entirely devoted to the development of NIPT tests have initiated large-scale clinical studies to drive their application.
[0005] Recently, targeted NGS methods for NIPT have been developed that sequence only specific target sequences. For example, a targeted NIPT method using Targeted Amplification of Sequences (TACS) has been described that utilizes a maternal blood sample to identify fetal chromosomal abnormalities (PCT publication WO2016 / 189388; U.S. Patent publication 2016 / 0340733; Koumbaris, G. et al. (2015) Clinical chemistry, 62(6), pp. 848-855).
[0006] The amount of sequencing required for this targeted approach is significantly less than the MPSS approach because sequencing is performed only on specific loci on the target sequences, rather than throughout the genome. There remains a need for additional methods for NGS-based approaches, particularly methods that are capable of targeting specific target sequences, thereby greatly reducing the amount of sequencing required compared to whole genome-based approaches, and increasing the sequencing depth of the target regions, thereby enabling detection of low signal-to-noise regions. In particular, there remains a need for additional methods that are capable of reliably detecting genetic abnormalities present in minute amounts in a sample.
[0007] Circulating tumor DNA (ctDNA) has also emerged as a dynamic biomarker for early detection of cancer or real-time assessment of MRD. To date, most MRD assessment methods using ctDNA are tumor-informed approaches that rely on initial genomic profiling of tumor tissue to identify tumor-derived variants for each individual patient. The rationale behind this approach is that knowing tumor-specific mutations for each patient can potentially increase the sensitivity of MRD detection, as the mutations can then be retrieved in ctDNA at predefined time points. However, despite the widespread use of tumor-informed approaches, there are several limitations. For example, tissue from surgical specimens can be insufficient for tissue sequencing due to limited tumor cell content, low DNA quality or yield. In addition to these limitations, surgical specimens can also fail to capture tumor heterogeneity adequately.
[0008] An alternative approach for early detection of cancer and MRD assessment using ctDNA is the tumor-agnostic approach that relies on non-invasive, plasma-only detection for MRD detection. Compared to tumor-informed approaches, tumor-agnostic approaches have several advantages. The advantages include faster turnaround time (as only a single (blood) sample is analyzed), lower processing cost and significantly reduced logistical complexity. Unlike traditional tissue biopsies, liquid biopsies are non-invasive, easily repeatable, and can provide valuable insights into tumor burden and treatment response. Furthermore, liquid biopsies can give a more complete molecular profile of the primary tumor, minimizing bias in biopsy results that is often caused by sampling bias and intratumoral heterogeneity.
[0009] Current liquid biopsy based tests fail to show sufficiently high accuracy due to their complexity and their limited sensitivity and specificity and can produce misleading results.
[0010] Currently available tests for detecting ctDNA from a patient's plasma detect a panel of fixed hotspots or actionable mutations. Given the heterogeneity of cancer, even large universal panels targeting more than a hundred genomic loci can detect only a small number of mutations from a particular individual's primary tumor. Mutations identified in these panels can not be of tumor origin, making this approach less specific. Likewise, in most available tests, primary tumor tissue and matched normal blood are collected from each patient. Genomic DNA from tumor tissue and buffy coat is extracted, sequenced, analyzed, and screened for patient-specific somatic mutations.
[0011] A very specific and preferred main object of the present application is the MRD monitoring and early detection of cancer using a tissue-agnostic approach customized to the patient's cancer biomarker signature, without the need for tissue biopsy analysis, complex methylation detection tests and / or the acquisition of large amounts of samples.
[0012] A more general problem to be solved and which has been solved by the present application is to provide a non-invasive in vitro molecular diagnostic method that avoids the need for solid tissue biopsy and can be performed using only blood.
[0013] A more general problem to be solved and which has been solved by the present application is to provide a non-invasive in vitro molecular diagnostic method that avoids the need for solid tissue biopsy and can be performed using only blood.
[0014] A more general problem to be solved and which has been solved by the present application is to provide a non-invasive in vitro molecular diagnostic method that avoids the need for solid tissue biopsy and can be performed using only blood.
[0015] In a more specific and preferred embodiment, another main object of the present application is the MRD monitoring and early detection of cancer using a tissue-agnostic approach customized to the patient's cancer biomarker signature, without the need for solid tissue biopsy analysis, complex methylation detection tests and / or the acquisition of large amounts of samples.
[0016] It will be fully apparent from the following description that the present application is generic in kind and specific in various forms of application. SUMMARY
[0017] The inventors have been able to solve the problems outlined above.
[0018] The present invention relates to an in vitro method of detecting somatic mutations in humans comprising the steps of:
[0019] (i) providing a human sample, preferably a blood sample, more preferably a plasma sample, a serum sample or a buffy coat sample, from a subject;
[0020] (ii) preparing a first nucleic acid sequencing library from nucleic acids present in said sample;
[0021] (iii) hybridizing one or more TAC oligonucleotides from a first pool of TAC oligonucleotides (TAC oligonucleotide-1) to said first nucleic acid library, thereby isolating a first subset of library nucleic acid molecules;
[0022] (iv) sequencing said first subset of library nucleic acid molecules and comparing the determined sequences to a human reference sequence, thereby creating a first set of putative informative sequence variants (PISV1);
[0023] (v) hybridizing one or more TAC oligonucleotides from a second, smaller pool of TAC oligonucleotides (TAC oligonucleotide-2) comprising TAC oligonucleotides specific for nucleic acid molecules of said first set of putative informative sequence variants to said first nucleic acid library, thereby isolating a second subset of library nucleic acid molecules;
[0024] (vi) sequencing said second subset of library nucleic acid molecules and comparing the determined sequences to a human reference sequence, thereby creating a second set of putative informative sequence variants (PISV2);
[0025] (vii) analyzing said putative informative sequence variants (PISV2), thereby detecting somatic mutations in said subject sample.
[0026] Steps (i) to (iv) are collectively referred to as step A. Steps (v) to (vii) are collectively referred to as step B.
[0027] The present invention relates to an in vitro diagnostic method without patient intervention. The human sample is provided previously.
[0028] Detecting herein whether a somatic mutation is present means identifying a mutation in a blood sample of a patient that originates from a somatic mutation event. Detecting herein also means identifying, which means unambiguously determining that it is a somatic mutation and characterizing the nucleotide type or mutation type.
[0029] In one aspect, the invention described herein relates to a method for determining the presence of cell-free circulating tumor DNA (ctDNA) nucleic acid molecules in a sample containing multiple cell-free DNA (cfDNA) nucleic acid molecules. Unlike previous studies (Jamshidi et al.
[2022] Cancer Cell 40, 1537-1549; Kotani, D., Oki, E., Nakamura, Y. et al.
[2023] NatMed, 29, 127-134; Reinert T, Henriksen TV, Christensen E et al.
[2019] JAMA Oncol., 5(8): 1124-1131; Coombes RC et al.
[2019] Clin Cancer Res, 25(14): 4255-4263), this invention achieves high accuracy without requiring significant sequencing costs for analyzing tissue or plasma samples. In one aspect, targeted deep resequencing of a small number of candidate somatic single nucleotide variant sites significantly reduces costs.
[0030] In one embodiment, a first subset of the library nucleic acid molecules from step A(iii) is amplified and sequenced to obtain at least 10,000, at least 20,000, at least 30,000, or at least 50,000 cfDNA nucleic acid molecules for each target region. In one embodiment, a second subset of the library nucleic acid molecules from step B(vi) is amplified and sequenced to obtain at least 100,000, at least 200,000, at least 300,000, or at least 500,000 cfDNA nucleic acid molecules for each target region.
[0031] As used in this article, “cell-free DNA” refers to DNA that is not contained within cells. Samples may contain cfDNA from normal or healthy cells and / or cancer cells. Cell-free DNA can be released into the blood or serum through secretion, apoptosis, or necrosis. If cfDNA is released from a tumor or cancer cell, it may be referred to as cell-free tumor DNA (cftDNA) or circulating tumor DNA (ctDNA).
[0032] The terms “nucleic acid” or “nucleic acid sequence” used in this article are used interchangeably and may refer to, but are not limited to, DNA, RNA, genomic DNA, cell-free DNA and / or RNA, tRNA, messenger RNA (mRNA), synthetic DNA, and synthetic RNA.
[0033] In the context of this invention, the terms "nucleic acid fragment" and "fragmented nucleic acid" are used interchangeably. In a preferred embodiment of the method, the nucleic acid fragment is circulating cell-free DNA or RNA, or circulating cell-free tumor DNA.
[0034] In the context of the present application, the term "subject" refers to an animal, preferably a mammal, more preferably a human or human patient. The term "subject" as used herein can refer to a subject having or suspected of having a tumor.
[0035] A "sample" as used herein refers to any biological material derived from a subject. The sample can be a liquid sample or a solid sample (e.g. a cell or tissue sample). The biological sample can be a body fluid, such as blood, plasma, serum, buffy coat, urine, pleural effusion, ascites, cerebrospinal fluid, saliva, sputum, bronchoalveolar lavage fluid, aspirate from different parts of the body (e.g. thyroid, breast), etc. The biological sample can be a cell-free sample. Thus, the sample can be a liquid biopsy sample obtained in a non-invasive manner from a blood sample of a subject comprising cell-free DNA (cfDNA), cell-free tumor DNA (cftDNA), circulating tumor DNA (ctDNA) or circulating cftDNA (ccftDNA), thereby enabling early detection of cancer before detectable or appreciable tumor formation is possible or enabling monitoring of disease progression, disease treatment or disease recurrence. The term "sample" can also refer to a stool sample. In one embodiment, the sample is selected from the group consisting of a plasma sample, a blood sample, a buffy coat sample, a urine sample, a sputum sample, a cerebrospinal fluid sample, an ascites sample and a pleural effusion sample of a subject having or suspected of having a tumor or having received a tumor resection. In one embodiment, the sample or DNA sample is from a tissue sample or a set of malignant cells of a subject having or suspected of having a tumor. In another embodiment, the sample is a stool sample.
[0036] In the context of the present application, the terms "tumor", "cancer" or "abnormality" can be used interchangeably. Herein, the terms "cancer" or "tumor" can also include early stage cancer or advanced cancer or metastasis. Herein, a "tumor" sample or "abnormality" sample can relate to a sample comprising (cell-free) DNA or RNA derived from a primary tumor or metastasis. Herein, a "normal" sample or "reference" sample can relate to a sample comprising (cell-free) DNA or RNA derived from non-cancerous, healthy or "normal" tissue or cells only. In the context of the present application, the terms "normal", "control" or "healthy" or "reference" can be used interchangeably.
[0037] A "tumor" herein refers to a general cancer, including but not limited to a solid tumor, an adenoma, a blood cancer, a liver cancer, a lung cancer, a pancreatic cancer, a prostate cancer, a breast cancer, a gastric cancer, a glioblastoma, a colorectal cancer, a head and neck cancer, an advanced cancer tumor, a benign or malignant tumor, a metastasis or a precancerous tissue.
[0038] Hybridization as used herein refers to the annealing of one or more probes to a target nucleotide sequence. Hybridization conditions typically include temperatures below the TAC oligonucleotide melting temperature but avoid temperatures that result in non-specific hybridization of the TAC oligonucleotide.
[0039] To achieve the desired separation of enriched sequences, the TAC oligonucleotide sequence is typically designed to enable separation of sequences that hybridize to the TAC oligonucleotide from sequences that do not bind to the TAC oligonucleotide. Typically, this is achieved by immobilizing the TAC oligonucleotide on a support. This allows those sequences that bind to the TAC oligonucleotide to be physically separated from those sequences that do not bind to the TAC oligonucleotide. For example, the individual sequences in the TAC oligonucleotide pool can be labeled with biotin, and the oligonucleotide pool can then be bound to beads coated with a biotin-binding substance, such as streptavidin or avidin. In a preferred embodiment, the TAC oligonucleotide is labeled with biotin and bound to streptavidin-coated magnetic beads, enabling separation using the magnetic properties of the beads. In one embodiment, the biotin can be chemically attached to the primers used to generate the TAC oligonucleotide. In a second embodiment, the latter can be generated by biotinylating a pool of sequences capable of hybridizing to the target region. However, it will be appreciated by those of ordinary skill in the art that other affinity binding systems are known in the art and can be used in place of biotin-streptavidin / avidin. This includes, but is not limited to, antibody-based methods, in which the TAC oligonucleotide is labeled with an antigen, and then bound to beads coated with an antibody. In addition, the TAC oligonucleotide can be integrated with a sequence tag on one end, and can be bound to a support through a complementary sequence on the support that hybridizes to the sequence tag. In addition to beads, other types of supports can be used, such as polymeric beads, and the like.
[0040] In one embodiment, the TAC oligonucleotide is provided in a form that enables it to bind to a support, such as a biotinylated TAC oligonucleotide. In another embodiment, the TAC oligonucleotide is provided with a support, such as a biotinylated TAC oligonucleotide provided with streptavidin-coated beads. In another embodiment, the TAC oligonucleotide is provided in a non-bound form and can exist freely in solution.
[0041] Sensitivity as used in the present invention refers to the number of true positives divided by the sum of the number of true positives and the number of false negatives. Sensitivity can characterize the ability of a test or method to correctly identify the proportion of people who truly have a condition. For example, sensitivity can characterize the ability of a method to correctly identify the number of subjects in a population who have cancer.
[0042] Specificity used in the present invention refers to the number of true negatives divided by the sum of true negatives and false positives. Specificity can characterize the ability of a detection method or method to correctly identify the proportion of people who are truly not suffering from a certain condition. For example, specificity can characterize the ability of a method to correctly identify the number of subjects in a population who are not suffering from cancer.
[0043] The "ratio of high quality to low quality aligned cfDNA nucleic acid molecules supporting a variant nucleotide" is defined as the number of aligned cfDNA nucleic acid molecules with alignment quality of 45 or more and supporting a variant nucleotide, divided by (the number of aligned cfDNA nucleic acid molecules with alignment quality below 45 and supporting a variant nucleotide plus 0.5) (to account for the case where the denominator is zero). In other embodiments of the present invention, alignment qualities of 20, 25, 30, 35 or 40 are used.
[0044] "Mapping quality" as used herein is defined as the probability that a read is mapped to the wrong position (Phred-scaled posterior probability that the mapping position of the read is incorrect).
[0045] Logarithmic likelihood value is defined as the logarithm of the ratio of the probability that a candidate somatic single nucleotide variant is a true event originating from the tumor tissue to the probability that the variant is an artifact / false signal.
[0046] In a preferred embodiment, preparing a DNA sequencing library comprises the step of adding unique molecular identifiers (UMIs) to uniquely label each molecule and create UMI families.
[0047] "Consensus sequence" refers to the order of the most frequently found residue (nucleotide or amino acid) at each position calculated in a sequence alignment. It serves as a unifying representation of each UMI family defined as aligned sequencing reads (cfDNA molecules) with identical start and end positions and identical UMI barcodes with respect to a reference genome. In the context of the present invention, the terms "mutation" or "variant nucleotide" are used interchangeably and generally refer to a mutation (e.g. single base pair mismatch) compared to a reference sequence.
[0048] As used herein, a single nucleotide variant is in concordance with a nearby single nucleotide polymorphism if it is only present on sequenced cfDNA molecules carrying the reference (or alternative) allele at the single nucleotide polymorphism site (i.e. they belong to the same haplotype).
[0049] As used herein, a somatic single nucleotide variant hotspot region is defined as a genomic region where the observed somatic variant frequency is statistically higher than the background frequency (in a preferred embodiment, the estimated value from the COSMIC dataset); the background frequency is estimated separately for each gene. In other embodiments, other databases are used, including but not limited to ICGC Data Portal, Genomic Data Commons (GDC), TP53 database, Cancer Hotspots, and cBioPortal.
[0050] As used herein, a "reference sequence" can be any nucleic acid sequence, genomic sequence, organism or subject's genomic sequence, preferably a sequence of the human genome (e.g. hgl9 or hgl8) or a sequence of a healthy individual or subject.
[0051] In the context of the present application, "frequency" is used interchangeably with abundance or occurrence number. In one embodiment of the present application, the variant allele frequency of a given locus describes the ratio of the number of aligned cfDNA molecules supporting a variant allele at the given locus to the total number of aligned cfDNA molecules covering the locus.
[0052] In the context of the present application, the term minimal residual disease (MRD) can refer to a very small amount of cancer cells remaining in the body during or after treatment.
[0053] Thus, in one embodiment, the method of the present application addresses and solves the problem of detecting MRD.
[0054] Herein, next generation sequencing (NGS) can be used for nucleic acid sequence analysis, although other sequencing technologies can also be employed which provide very accurate counts in addition to sequence information. Thus, other accurate counting methods can also be used instead of NGS, such as but not limited to digital PCR, single molecule sequencing, nanopore sequencing, DNA nanoball sequencing, ligation sequencing, pyrosequencing, ion semiconductor sequencing, semiconductor sequencing, sequencing by synthesis, and microarrays.
[0055] As used herein, the "functional impact" of a variant determines its impact on genes, transcripts and protein sequences, as well as regulatory regions. The functional impact can be deleterious, pathogenic, prognostically bad, pathogenic and predisposing.
[0056] The terms "Target Capture Oligonucleotides", "TAC oligonucleotides" or "TACs" used herein interchangeably refer to short DNA oligonucleotides complementary to a target region on a target genomic sequence, which are used as baits to capture and enrich the target region from a large sequence library, such as a whole genome sequencing library prepared from a biological sample. The TAC oligonucleotide pool is used for enrichment, wherein the oligonucleotides within the oligonucleotide pool are optimized for (i) the length of the oligonucleotides; (ii) the distribution of the TAC oligonucleotides over the target region; and (iii) the GC content of the TAC oligonucleotides. The number of oligonucleotides in the TAC oligonucleotide pool (pool size) is also optimized. BRIEF DESCRIPTION OF DRAWINGS
[0057] Figure 1 The basic principle of phasing candidate somatic single nucleotide variants with germline heterozygous single nucleotide polymorphisms is illustrated. This process provides a powerful approach to artifact elimination.
[0058] Figure 2 is a flow chart illustrating the workflow of the two-step non-invasive targeted re-sequencing method (nested TAC oligonucleotide method) of the present application for determining the presence of cell-free circulating tumor-derived DNA in plasma.
[0059] Figure 3 is a schematic diagram of TAC oligonucleotide-based target sequence enrichment using a single TAC oligonucleotide (left) and using a family of TAC oligonucleotides (right).
[0060] Figure 4 The number of somatic variants detected in each sample is illustrated. In Figure 4 the y-axis represents the number of mutations detected in each sample, with black bars corresponding to normal samples and grey bars corresponding to abnormal samples, and the x-axis indicates the status of each sample (normal or early stage cancer). The threshold for calling a sample positive is indicated by the black horizontal solid line. The clinical specificity of this detection was 100% (11 / 11; 95% confidence interval: 72-100%) and the clinical sensitivity was 71% (10 / 14; 95% confidence interval: 42-92%). DETAILED DESCRIPTION
[0061] The present application relates to an in vitro method for detecting human somatic mutations, comprising the following steps:
[0062] (i) providing a human sample, preferably a blood sample, more preferably a plasma sample, a serum sample or a buffy coat sample, from a subject;
[0063] (ii) preparing a first nucleic acid sequencing library from nucleic acids present in the sample;
[0064] (iii) hybridizing one or more TAC oligonucleotides from a first pool of TAC oligonucleotides (TAC oligonucleotide-1) to the first nucleic acid library, thereby isolating a first subset of library nucleic acid molecules;
[0065] (iv) sequencing the first subset of library nucleic acid molecules and comparing the determined sequences to human reference sequences, thereby creating a first set of putative informative sequence variants (PISV1);
[0066] (v) hybridizing one or more TAC oligonucleotides from a second, smaller pool of TAC oligonucleotides (TAC oligonucleotide-2) comprising TAC oligonucleotides specific for nucleic acid molecules of the first set of putative informative sequence variants to the first nucleic acid molecule library, thereby isolating a second subset of library nucleic acid molecules;
[0067] (vi) sequencing the second subset of library nucleic acid molecules and comparing the determined sequences to human reference sequences, thereby creating a second set of putative informative sequence variants (PISV2);
[0068] (vii) analyzing the putative informative sequence variants (PISV2), thereby detecting somatic mutations in the subject sample.
[0069] Steps (i) through (iv) are collectively referred to as Step A. Steps (v) through (vii) are collectively referred to as Step B.
[0070] The ability to detect somatic mutations is important and therefore primary in molecular research. It is of great interest to populate databases with somatic mutations that have been identified in different cell types, as such information can support subsequent diagnoses.
[0071] The present method enables the detection of somatic mutations in humans using any human sample. The present method preferably enables the detection of somatic mutations in humans using only a blood sample. Such detection of somatic mutations has applications in multiple diagnostic fields, such as, but not limited to, cancer detection. Somatic mutations refer to changes in DNA sequences in somatic cells of multicellular organisms that have specialized germ cells, i.e., mutations that occur in any cell other than a gamete, germ cell, or gametocyte. Unlike germline mutations, which can be passed on to the offspring of an organism, somatic mutations are typically not passed on to offspring. Although somatic mutations are not passed on to the offspring of an organism, somatic mutations are present in all offspring of cells within the same organism. Many cancers result from the accumulation of somatic mutations.
[0072] The term "somatic cell" generally refers to a cell of the body, as opposed to a germ cell (germline) that produces eggs or sperm. For example, in mammals, somatic cells make up all internal organs, skin, bone, blood, and connective tissue. There are approximately 220 types of somatic cells in the human body.
[0073] Thus, in one embodiment, cells are detected with cell type specificity.
[0074] In most animals, the separation of germ cells from somatic cells (germline development) occurs at an early stage of development. After this separation in the embryo, any mutation outside the germline cells cannot be passed on to the offspring of the organism. However, somatic cell mutations are passed on to the offspring of the mutated cell within the same organism. A large portion of the organism's tissues can carry the same mutation, especially when the mutation occurs at an early stage of development. Somatic mutations that occur later in the life of an organism can be difficult to detect because they can affect only a single cell, for example, post-mitotic neurons; thus, improvements in single-cell sequencing are important tools for studying somatic mutations. Both nuclear DNA and mitochondrial DNA of a cell can accumulate mutations, and somatic mitochondrial mutations have been implicated in the development of certain neurodegenerative diseases. Studies have shown that the frequency of mutations in somatic cells is generally higher than in germline cells, and, in addition, there are differences in the types of mutations observed in germline and somatic cells. Differences in mutation frequency exist between different somatic tissues within the same organism and between species, which is one application direction of the present application.
[0075] Somatic mutations accumulate within the cells of an organism as the organism ages and with each round of cell division; the role of somatic mutations in the development of cancer is well established and is related to the biology of aging.
[0076] Mutations in neural stem cells (especially during neurogenesis) and post-mitotic neurons can lead to genomic heterogeneity in neurons, which is referred to as the "somatic brain mosaic." The accumulation of age-related mutations in neurons can be related to neurodegenerative diseases, including Alzheimer's disease, but this association has not been confirmed. Most cells in the central nervous system of an adult human are post-mitotic cells, and somatic mutations in adults can affect only a single neuron. Unlike mutations in cancer that lead to clonal expansion, deleterious somatic mutations can lead to neurodegenerative diseases through cell death. Thus, accurately assessing the somatic mutation burden in neurons remains difficult, which is an important reason for the present application.
[0077] If a mutation occurs in a somatic cell of an organism, the mutation is present in all descendants of that cell within the same organism. The accumulation of certain mutations in somatic cell generations is part of the process of malignant transformation from normal cells to cancer cells.
[0078] Cells with heterozygous loss-of-function mutations (one copy of the gene is normal, one copy of the gene is mutated) can still function normally through the unmutated copy until the normal copy undergoes a spontaneous somatic mutation. Such mutations occur frequently in living organisms, but the frequency is difficult to measure. Measuring this frequency is of great importance in predicting the risk of cancer in humans.
[0079] Interestingly, somatic mutations are also being identified in other conditions besides cancer, including neurodevelopmental diseases. Somatic mutations can arise during prenatal brain development and lead to neurological diseases, for example, brain malformations associated with epilepsy and intellectual disability, even at low levels of mosaicism. The present invention will enable more accurate assessment of neurodevelopmental disorders and somatic mutations during normal brain development.
[0080] Rare conditions with a clear basis in somatic variation include hematopoietic disorders, in which stem cells can undergo mutation and expand, leading to a disease phenotype. These include paroxysmal nocturnal hemoglobinuria type 1 (PNH1) caused by mutations in the PIG-A gene, and X-linked alpha-thalassemia mental retardation syndrome caused by mutations in ATRX. PNH1 is an acquired hemolytic anemia that presents with hemoglobinuria, abdominal pain, smooth muscle tone disorders, fatigue, and thrombosis. The disease is caused by expansion of hematopoietic stem cells carrying a mutation in the PIG-A gene (a somatically acquired change). X-linked alpha-thalassemia mental retardation syndrome is sometimes associated with myelodysplastic syndrome, and such cases are often associated with somatic mutations. Interestingly, in the case of ATRX mutations, the myelodysplastic syndrome condition that seems to be triggered by somatic variation appears to be more severe than that triggered by germline mutations. Clearly, the clonal expansion capacity of hematopoietic stem cells provides a mechanism by which somatic mutations can trigger disease.
[0081] Neurofibromatosis type 1 (NF1) is a condition that maps to a segment of chromosome 17 long arm (17q) that presents with cafe au lait spots, ocular Lisch nodules, and cutaneous fibromas. Multiple studies have shown that the vast majority of NF1 cases are caused by somatic mutations, usually deletions or microdeletions of this chromosomal region (40% of cases). Other cases are caused by somatic mitochondrial DNA (mtDNA) mutations. In either case, somatic changes have been clearly identified as a common cause of NF1. Similarly, NF2 has been shown to be commonly caused by somatic mutations (25-30% of cases).
[0082] By careful characterization of resected tissue, it can be determined whether the condition of other tissues is somatic in origin, examples include heart and kidney disease. For example, mutations in connexin 40, a protein expressed by cardiomyocytes encoded by the GJA5 gene, have been shown to affect electrical signaling and are associated with the vast majority of cases of atrial fibrillation. Most GJA5 mutations found in patient cardiomyocytes are not present in blood, suggesting a somatic origin. A similar situation is found in some cases of Alport syndrome. Alport syndrome is an X-linked dominant disorder characterized by kidney disease, hearing loss, and eye abnormalities. The disease is caused by mutations in the collagen IV component, primarily COL4A5. While most cases of Alport syndrome are inherited through the germline, there are reports of male patients with milder phenotypes harboring somatic mutations in COL4A5. For many other X-linked diseases that are extremely severe or lethal in males, somatic mutations can manifest as milder disease forms.
[0083] Somatic mutations are also associated with certain neurological diseases, including epilepsy, autism spectrum disorders (e.g., Rett syndrome), and intellectual disability, although results from studies of identical twins with multiple sclerosis (MS) are largely negative. The latter example is based on whole genome data from identical twins with discordant phenotypes, but these data are derived from lymphocytes, which are clearly not ideal tissues for multiple sclerosis. Neurological diseases can be particularly susceptible to somatic mutations because even if cells harboring mutations are less than 10%, depending on the distribution of these cells in the brain, the phenotype can be affected. For example, hemimegalencephaly (HMG) is characterized by enlargement and malformation of an entire cerebral hemisphere and is associated with somatic mutations in AKT3 and other mutations in the PI3K-AKT3-mTOR pathway, even if cells harboring somatic mutations are as low as 8% (typically less than 35%). However, because cells harboring mutations are widely distributed, patients still exhibit HMG. Quite rare somatic mutations can have an impact due to the unique developmental patterns of the brain and its complex clonal migration patterns, such that clonality is not limited to adjacent or nearby cells.
[0084] Lissencephaly (also known as smooth brain) can be caused by mutations in two genes: Doublecortin X (DCX) or Lissencaphaly 1 (LIS1). Mutations in the LIS1 gene, located on the short arm of chromosome 17 (17pl), are usually lethal in males, but have been found in two patients with predominantly subcortical posterior band heterotopia, a milder form associated with somatic mosaicism
[40] . In these patients, 18-24% of blood cells and 21-34% of hair root cells were mutated. DCX1 somatic mutations have also been shown to be associated with a similar disease phenotype. For the above neurological diseases, not all neuronal cells carry these mutations, but these mutations are present in white blood cells, indicating the presence of early somatic mutations.
[0085] Mutations in the X-linked pyruvate dehydrogenase A1 gene (PDHA1) can manifest either metabolic or neurological features. The metabolic form usually leads to death in infancy due to lactic acidosis, while the neurological form presents symptoms including seizures, intellectual disability, and spasticity. There is a continuum between these two manifestations. A higher proportion of heterozygous females present with severe disease, but evidence of preferential X-inactivation and somatic mutation has been reported in a female with mild disease
[42] . Similarly, a male with the milder form of the disease had an exon skipping mutation in skin and muscle tissue, but this mutation was not detected in lymphocytes
[43] . Although limited to individual clinical cases, these instances suggest that somatic mutation of a single gene can affect disease risk. Notably, these cases caused by somatic variation manifest as the milder form of the disease.
[0086] Finally, autoimmune diseases can also be caused by somatic mutations. Autoimmune lymphoproliferative syndrome (ALPS) is a disease characterized by benign lymphocytosis, elevated immunoglobulins, plasma IL-10, and FAS-L, and accumulation of double negative T cells. A recent study on this disease showed that in some cases, the disease is caused by somatic mutations (Human Somatic Variation: It’s Not Just for Cancer Anymore, Chun Li & Scott M. Williams). In a preferred embodiment of the invention, the present invention relates to an in vitro method for diagnosing, prognosing, treatment response control, response prediction of a specific human disease, comprising the following steps:
[0087] (i) providing a human sample, preferably a blood sample, more preferably a plasma sample, a serum sample or a buffy coat sample, from a subject;
[0088] (ii) preparing a first nucleic acid sequencing library from nucleic acids present in the sample;
[0089] (iii) hybridizing one or more TAC oligonucleotides from a first pool of TAC oligonucleotides (TAC oligonucleotide-1) to the first nucleic acid library, thereby isolating a first subset of library nucleic acid molecules;
[0090] (iv) sequencing the first subset of library nucleic acid molecules and comparing the determined sequences to human reference sequences, thereby creating a first set of putative informative sequence variants (PISV1);
[0091] (v) hybridizing one or more TAC oligonucleotides from a second, smaller pool of TAC oligonucleotides (TAC oligonucleotide-2) comprising TAC oligonucleotides specific for nucleic acid molecules of the first set of putative informative sequence variants to the first nucleic acid molecule library, thereby isolating a second subset of library nucleic acid molecules;
[0092] (vi) sequencing the second subset of library nucleic acid molecules and comparing the determined sequences to human reference sequences, thereby creating a second set of putative informative sequence variants (PISV2);
[0093] (vii) analyzing the putative informative sequence variants (PISV2), thereby diagnosing, prognosticating, drug response controlling a particular human disease state.
[0094] Steps (i) through (iv) are collectively referred to as Step A. Steps (v) through (vii) are collectively referred to as Step B.
[0095] Providing a human sample herein means making a patient sample available. This step does not involve patient intervention.
[0096] Detection is not the same as diagnosis. Diagnosis involves more in-depth analysis of the quality of the mutation.
[0097] The inventors have coined the term "nested TAC oligonucleotide enrichment method" for this very general invention. This method differs from the one disclosed in the inventors' previous invention WO2016 / 189388.
[0098] When a human blood sample (preferably a plasma sample, a serum sample or a buffy coat sample) from a subject is provided, ideally the method is first performed on plasma and / or serum, and then on buffy coat.
[0099] After isolation, cell-free DNA from the sample is used to construct sequencing libraries, making the sample compatible with downstream sequencing technologies such as next-generation sequencing (NGS). Typically, this involves ligating adapters to the ends of cell-free DNA fragments, followed by amplification. Sequencing library preparation kits are commercially available. In one embodiment, one sequencing library is constructed. In other embodiments, two or more sequencing libraries are constructed in parallel.
[0100] In the past, in-solution hybridization enrichment prior to sequencing has been used to enrich specific target regions (see WO2016 / 189388). However, for the method of this invention, target oligonucleotides (referred to as targeted capture oligonucleotides, or TAC oligonucleotides) for enriching specific target regions have been optimized for maximum efficiency, specificity, and accuracy. Furthermore, these oligonucleotides are used in the form of a family of TAC oligonucleotides containing multiple members that bind to the same genomic sequence but have different start and / or termination positions, thus significantly improving the enrichment of the target genomic sequence compared to using a single TAC oligonucleotide that binds to the genomic sequence. The structure of this TAC oligonucleotide family is described in... Figure 3 The diagram illustrates that members of the TAC oligonucleotide family have different start and / or termination positions, which leads to an interleaved binding pattern when they bind to a target genomic sequence.
[0101] Using the TAC oligonucleotide family with pools of TAC oligonucleotides that bind to each target sequence significantly improves the enrichment of target sequences compared to using a single TAC oligonucleotide in a pool of TAC oligonucleotides that binds to each target sequence. This is demonstrated by the fact that the sequencing depth of the TAC oligonucleotide family is increased by an average of more than 50% compared to a single TAC oligonucleotide.
[0102] Each TAC oligonucleotide family contains multiple members that bind to the same target genomic sequence but have different start and / or termination positions relative to a reference coordinate system of that target genomic sequence. Typically, the reference coordinate system used for analyzing human genomic DNA is the human reference genome version hg19, which is publicly available in the art, but other versions (e.g., hg38) may also be used. Alternatively, the reference coordinate system can be a genome artificially created based on the hg19 version that contains only the target genomic sequence.
[0103] Each TAC oligonucleotide family contains at least two members that bind to the same target genomic sequence. In various embodiments, each TAC oligonucleotide family contains at least two member sequences, at least three member sequences, at least four member sequences, at least five member sequences, at least six member sequences, at least seven member sequences, at least eight member sequences, at least nine member sequences, or at least ten member sequences. In various embodiments, each TAC oligonucleotide family contains two member sequences, three member sequences, four member sequences, five member sequences, six member sequences, seven member sequences, eight member sequences, nine member sequences, or ten member sequences. In various embodiments, a plurality of TAC oligonucleotide families contains different families with different numbers of member sequences. For example, one TAC oligonucleotide pool can contain one TAC oligonucleotide family with three member sequences, another TAC oligonucleotide family with four member sequences, yet another TAC oligonucleotide family with five member sequences, and so on. In one embodiment, one TAC oligonucleotide family contains three to five member sequences. In another embodiment, one TAC oligonucleotide family contains four member sequences.
[0104] A TAC oligonucleotide pool contains a plurality of TAC oligonucleotide families. Thus, one TAC oligonucleotide pool contains at least two TAC oligonucleotide families. In various embodiments, one TAC oligonucleotide pool contains at least three different TAC oligonucleotide families, at least five different TAC oligonucleotide families, at least ten different TAC oligonucleotide families, at least fifty different TAC oligonucleotide families, at least one hundred different TAC oligonucleotide families, at least five hundred different TAC oligonucleotide families, at least one thousand different TAC oligonucleotide families, at least two thousand TAC oligonucleotide families, at least four thousand TAC oligonucleotide families, or at least five thousand TAC oligonucleotide families.
[0105] Each member of a TAC oligonucleotide family binds to the same target genomic region, but has a different starting and / or ending position relative to the reference coordinate system of the target genomic sequence, so that the binding patterns of the TAC oligonucleotide family members are staggered Figure 3). In various embodiments, the overlap of the starting and / or ending positions is at least 3 base pairs, or at least 4 base pairs, or at least 5 base pairs, or at least 6 base pairs, or at least 7 base pairs, or at least 8 base pairs, or at least 9 base pairs, or at least 10 base pairs, or at least 15 base pairs, or at least 20 base pairs, or at least 25 base pairs. Typically, the overlap of the starting and / or ending positions is 5 to 10 base pairs. In one embodiment, the overlap of the starting and / or ending positions is 5 base pairs. In another embodiment, the overlap of the starting and / or ending positions is 10 base pairs.
[0106] The nested TAC oligonucleotide enrichment-based methods of the present application can be used to detect a variety of genetic abnormalities. In one embodiment, the genetic abnormality is a chromosomal aneuploidy (e.g., trisomy, partial trisomy, or monosomy). In other embodiments, the genomic abnormality is a structural abnormality, including but not limited to copy number changes (including microdeletions and microduplications), insertions, translocations, inversions, and small fragment mutations (including point mutations) and mutational signatures. In another embodiment, the genetic abnormality is a chromosomal mosaicism.
[0107] In the methods of the present application, the second TAC oligonucleotide pool (TAC oligonucleotide-2) is ideally composed of a subset of the first TAC oligonucleotide pool. Alternatively, the second TAC oligonucleotide pool is newly synthesized and is preferably patient-specific or disease-specific. Fundamentally and generally, the first set of putative informative sequence variants (PISV1) will dictate which second TAC oligonucleotide pool is used. PISV1 lays the foundation for the diagnostic analysis and determines which loci need to be analyzed in more detail or depth.
[0108] In one embodiment, the second TAC oligonucleotide pool used is selected from the larger initial oligonucleotide pool used in the first TAC oligonucleotide pool of the present application, depending on the characteristics of each patient, i.e., the results of the first step of the plasma sample analysis as described in detail in Example 1 and Example 2 below. In various embodiments, the smaller second TAC oligonucleotide pool comprises a total percentage of the larger first TAC oligonucleotide initial pool of 0.1%, 0.25%, 0.5%, 0.75%, 1.0%, or 5%.
[0109] The first subset of the library nucleic acid molecules is amplified and sequenced such that at least 10,000, 20,000, 30,000, or 50,000 cfDNA nucleic acid molecules per target region are obtained. The second subset of the library nucleic acid molecules is amplified and sequenced such that at least 100,000 or 200,000 or 300,000 or 500,000 cfDNA nucleic acid molecules per target region are obtained.
[0110] The filtering statistics include a mapping quality score and a variant allele frequency threshold, which are calculated by estimating the distribution of false variant allele frequencies for each possible substitution in the regions covered by the TAC oligonucleotide pool using a set of normal reference samples that have not been previously diagnosed with cancer.
[0111] Among the many advantages of TAC oligonucleotides is that there are fewer sequence artifacts generated. However, there can be artificial sequence changes in the newly synthesized DNA strand. In the method of the present invention, creating a first set of putative informative sequence variants (PISV1) includes the step of distinguishing between sequence variants caused by experimental artifacts and sequence variants present in the sample.
[0112] In the method of the present invention, each sequence in the first set of putative informative sequence variants (PISV1) is ranked according to an information value based on likelihood statistics.
[0113] In the method, the likelihood statistics for a single nucleotide variant represent the probability that the single nucleotide variant is a true somatic single nucleotide variant, wherein the likelihood statistics are calculated using a regression model, wherein the regression model includes one or more of the following:
[0114] (i) the average mapping quality of the aligned cfDNA nucleic acid molecules;
[0115] (ii) the ratio of high quality to low quality aligned cfDNA nucleic acid molecules supporting the variant nucleotide;
[0116] (iii) the average distance of the single nucleotide variant site to one or more aligned cfDNA nucleic acid molecule endpoints;
[0117] (iv) the Levenshtein distance of the sequence of the one or more aligned cfDNA nucleic acid molecules to the reference genome sequence;
[0118] (v) the frequency of the one or more single nucleotide variant sites in a population of normal reference samples;
[0119] (vi) the proportion of the sequenced cfDNA nucleic acid molecules that support phasing of the single nucleotide variant with a nearby single nucleotide polymorphism, if present, wherein the single nucleotide variant and the single nucleotide polymorphism are separated by up to 110 bp;
[0120] (vii) the frequency of the single nucleotide variant in a disease database, such as the Catalogue of Somatic Mutations in Cancer;
[0121] (viii) the predicted functional impact of the single nucleotide variant.
[0122] Accordingly, the likelihood statistic for a single nucleotide variant is calculated by a regression model that includes the following covariates: (a) the average alignment quality of the aligned cfDNA nucleic acid molecules; (b) the ratio of high quality to low quality aligned cfDNA nucleic acid molecules supporting the variant nucleotide; (c) the average distance of the single nucleotide variant site to the end of one or more aligned cfDNA nucleic acid molecules; (d) the Levenshtein distance of the sequence of one or more aligned cfDNA nucleic acid molecules to the reference genome sequence; (e) the frequency of one or more single nucleotide variant sites in a normal reference population of samples; (f) the proportion of single nucleotide variants in the sequenced cfDNA nucleic acid molecules that support phasing with nearby single nucleotide polymorphisms, if present, where the interval between the single nucleotide variant and the single nucleotide polymorphism is at most 110 bp; (g) the frequency of the single nucleotide variant in a disease database, such as the Catalogue of Somatic Mutations in Cancer (COSMIC) or other cancer somatic mutation database; and (h) the functional impact of the candidate somatic single nucleotide variant.
[0123] The present invention relates to an in vitro method for predicting the therapeutic response of a specific human disease, comprising the steps of:
[0124] (i) providing a human sample, preferably a blood sample, more preferably a plasma sample, a serum sample or a buffy coat sample, from a subject;
[0125] (ii) preparing a first nucleic acid sequencing library from nucleic acids present in said sample;
[0126] (iii) hybridizing one or more TAC oligonucleotides from a first pool of TAC oligonucleotides (TAC oligonucleotide-1) to said first nucleic acid library, thereby isolating a first subset of library nucleic acid molecules;
[0127] (iv) sequencing said first subset of library nucleic acid molecules and comparing the determined sequences to a human reference sequence, thereby creating a first set of putative informative sequence variants (PISV1);
[0128] (v) hybridizing one or more TAC oligonucleotides from a second, smaller pool of TAC oligonucleotides (TAC oligonucleotide-2) comprising TAC oligonucleotides specific for nucleic acid molecules of said first set of putative informative sequence variants to said first nucleic acid molecule library, thereby isolating a second subset of library nucleic acid molecules;
[0129] (vi) sequencing the second subset of library nucleic acid molecules and comparing the determined sequences to a human reference sequence, thereby creating a second set of putative informative sequence variants (PISV2);
[0130] (vii) analyzing the putative informative sequence variants (PISV2), thereby making a drug response prediction for a particular human disease state.
[0131] Steps (i) to (iv) are collectively referred to as Step A. Steps (v) to (vii) are collectively referred to as Step B.
[0132] Cancer is a complex disease. Cancer tissue is often heterogeneous in terms of the composition of cancer mutations within the tissue. As described above, each cell can harbor different mutations. A selected drug will have responder specificity, but only for a selected set of mutations and a selected metabolic pathway. A cancer tissue can harbor only 20% of mutation "A" and 80% of a different cancer mutation "B". By using the method of the present application, this heterogeneity can be reflected and the mutations can be ranked in terms of their likelihood of response to various drugs acting on different pathways as well as the quantitative presence of the mutations and their oncogenic driving ability.
[0133] In a preferred embodiment, the method relates to a method of making a prognosis for a particular human disease, comprising the steps of:
[0134] (i) providing a human sample, preferably a blood sample, more preferably a plasma sample, a serum sample or a buffy coat sample, from a subject;
[0135] (ii) preparing a first nucleic acid sequencing library from nucleic acids present in the sample;
[0136] (iii) hybridizing one or more TAC oligonucleotides from a first pool of TAC oligonucleotides (TAC oligonucleotide-1) to the first nucleic acid library, thereby isolating a first subset of library nucleic acid molecules;
[0137] (iv) sequencing the first subset of library nucleic acid molecules and comparing the determined sequences to a human reference sequence, thereby creating a first set of putative informative sequence variants (PISV1);
[0138] (v) hybridizing one or more TAC oligonucleotides from a second, smaller pool of TAC oligonucleotides (TAC oligonucleotide-2) comprising TAC oligonucleotides specific for nucleic acid molecules of the first set of putative informative sequence variants to the first nucleic acid molecule library, thereby isolating a second subset of library nucleic acid molecules;
[0139] (vi) sequencing a second subset of the library nucleic acid molecules and comparing the determined sequences to a human reference sequence, thereby creating a second set of putative informative sequence variants (PISV2);
[0140] (v) analyzing the putative informative sequence variants (PISV2), thereby prognosticating a particular human disease state.
[0141] Prognosis is a medical term that means predicting the likely or expected course of a disease, including whether signs and symptoms will improve or worsen, the speed of improvement or worsening, or whether they will remain stable over time. When applied to large statistical populations, prognostic assessments can be very accurate: for example, the statement "45% of patients with severe septic shock will die within 28 days" is a statement that has some credibility because previous studies have found that that proportion of patients die.
[0142] This statistical information does not apply to the prediction of each individual patient, because patient-specific factors can significantly alter the expected course of the disease: additional information is needed to determine whether a certain patient belongs to the 45% of patients who will die, or to the 55% of patients who will survive. By using the method of the present invention and by developing a sample testing protocol, it is possible to make prognostic assessments. For example, a flexible and dynamic model of cancer mutations is established, which assigns differential weights to each mutation, thereby prognosticating disease progression and / or disease outcome.
[0143] The present invention relates to an in vitro method for therapeutic response control of a particular human disease, comprising the steps of:
[0144] (i) providing a human sample, preferably a blood sample, more preferably a plasma sample, a serum sample or a buffy coat sample, from a subject;
[0145] (ii) preparing a first nucleic acid sequencing library from nucleic acids present in the sample;
[0146] (iii) hybridizing one or more TAC oligonucleotides from a first pool of TAC oligonucleotides (TAC oligonucleotide-1) to the first nucleic acid library, thereby isolating a first subset of library nucleic acid molecules;
[0147] (iv) sequencing the first subset of the library nucleic acid molecules and comparing the determined sequences to a human reference sequence, thereby creating a first set of putative informative sequence variants (PISV1);
[0148] (v) hybridizing one or more TAC oligonucleotides from a second smaller pool of TAC oligonucleotides (TAC oligonucleotide-2) comprising TAC oligonucleotides specific for nucleic acid molecules comprising putative informative sequence variations of the first set to the first library of nucleic acid molecules, thereby isolating a second subset of library nucleic acid molecules;
[0149] (vi) sequencing the second subset of library nucleic acid molecules and comparing the determined sequences to human reference sequences, thereby creating a second set of putative informative sequence variations (PISV2);
[0150] (vii) analyzing the putative informative sequence variations (PISV2), thereby controlling drug response to a specific human disease state.
[0151] Steps (i) to (iv) are collectively referred to as Step A. Steps (v) to (vii) are collectively referred to as Step B.
[0152] Therapeutic drug monitoring (TDM) is a branch of clinical chemistry and clinical pharmacology that specializes in the determination of drug concentrations in blood. Its main focus is on drugs with a narrow therapeutic window, i.e. drugs that are prone to under- or overdosing. TDM aims to improve patient care by individualized adjustment of drug dosages, which has been shown by clinical experience or clinical trials to improve treatment outcome in the general or in a specific population.
[0153] In cancer, the situation is more complex. Cancer tissue is usually heterogeneous in terms of the composition of cancer mutations within the tissue. As described above, each cell can carry different mutations, thereby making the tissue a complex three-dimensional model of heterogeneous mutations. By using the method of the present application and by devising a sample testing protocol, monitoring / control can be achieved. For example, a flexible and dynamic model of cancer mutations is established, differentiating weights are assigned to individual mutations, thereby allowing prognosis of disease progression and / or disease outcome. The selected drug will have responder specificity, but only for a selected group of mutations and a selected metabolic pathway. The method will monitor the stress on this pathway and the quantitative outcome of the mutations at specific time points before, during and after treatment.
[0154] In the method of the present application, the interval between single nucleotide variations and single nucleotide polymorphisms is ideally at most 100 bp, 120 bp, 130 bp, 140 bp or 150 bp.
[0155] In a preferred embodiment, preparing a DNA sequencing library comprises the step of adding unique molecular identifiers (UMIs) to uniquely label each molecule and create UMI families.
[0156] Preferably, in said method, the sample is a plasma sample, a serum sample or a buffy coat sample and the nucleic acids in the sample are cell-free DNA (cfDNA).
[0157] Preferably, the putative informative nucleic acid sequence variant is a somatic DNA variant and is selected from the group consisting of a frameshift mutation, an indel mutation and a single nucleotide substitution (single nucleotide mutation). Most preferably, the analyzed mutation is a single nucleotide substitution (single nucleotide mutation).
[0158] In other embodiments, the method of the application can be used to detect a plurality of genetic abnormalities. In one embodiment, the genetic abnormality is a chromosomal aneuploidy (e.g., trisomy, partial trisomy, or monosomy). In yet other embodiments, the genomic abnormality is a structural abnormality including, but not limited to, copy number variations (including microdeletions and microduplications), insertions, translocations, inversions, and small fragment mutations (including point mutations) and mutational signatures. In another embodiment, the genetic abnormality is a chromosomal mosaicism.
[0159] Preferably, in the method of the application, the first set of putative informative sequence variants (PISV1) comprises 5 to 200 PISVs, 8 to 100 PISVs, 10 to 80 PISVs and 20 to 60 PISVs, preferably about 40 PISVs.
[0160] Ideally, in the method of the application, said diagnosis is a diagnosis of a disease selected from the group comprising a chronic disease, a congenital disease, a genetic disease, a familial genetic disease, an acute disease and an idiopathic disease.
[0161] In the method of the application, the first TAC oligonucleotide pool comprises oligonucleotides specific for known genetic disease loci of said disease to be diagnosed.
[0162] In the method of the application, said disease is preferably selected from the group comprising a cancer, a neurodegenerative disease, McCune-Albright syndrome, blood system and immune related disorders, paroxysmal nocturnal hemoglobinuria type 1, X-linked alpha-thalassemia mental retardation syndrome, Alport syndrome, a genetic disease, an autoimmune disease, a kidney disease, a cardiovascular disease, a psychiatric disease, a disease of aging, a neuromuscular disease, a disease of the reproductive system, a disease of the lung, organ transplant monitoring and sepsis.
[0163] In the method of the application, said cancer is selected from the group consisting of a carcinoma, a sarcoma, a melanoma, a lymphoma, an adenoma and a leukemia.
[0164] In various embodiments of the method, the top 20, 30, 40, 50, 60, 70, 80, 90, 96 single variant sites are selected from the plurality of ranked single nucleotide variant sites in the first set of putative informative sequence variants (PISV1).
[0165] In various embodiments of the method, the presence of cell-free tumor DNA nucleic acid molecules is determined if more than 1%, 10%, 15%, 20%, 25% of the plurality of single nucleotide variants selected in step A are detected as mutated in step B, or 1, 2, 3, or 4 mutations are detected.
[0166] In various embodiments, in addition to the plurality of single nucleotide variant sites, the method can also identify a plurality of small fragment insertions and deletions.
[0167] The method of the present application can be used to detect relapse and monitor MRD, wherein the detection of at least a portion of ctDNA remaining in the patient after treatment can indicate that the treatment was not sufficient or that the tumor has relapsed.
[0168] The method of detecting cell-free tumor DNA can also be used in a variety of different clinical scenarios in the field of oncology. For example, the method can be used to make a preliminary cancer diagnosis in a subject suspected of having cancer. Thus, in one embodiment, the method further comprises making a diagnosis in the subject based on the detection of at least a portion of cell-free tumor DNA in a sample taken from the subject.
[0169] The method can also be used to select an appropriate treatment regimen for a patient diagnosed with cancer, wherein the treatment regimen is designed to be effective against a tumor having a tumor biomarker detected in the patient's tumor (i.e., what is referred to in the art as personalized medicine). Thus, in another embodiment, the method further comprises selecting a treatment regimen for the subject based on the detection of the sequence of at least one tumor biomarker.
[0170] In one aspect, the method can be used to monitor the efficacy of a treatment regimen, wherein changes in the tumor biomarker detection results are used as an indicator of treatment efficacy. Thus, in another embodiment, the method further comprises monitoring the treatment efficacy of a treatment regimen in the subject based on the detection of the sequence of at least one tumor biomarker.
[0171] In a different embodiment, the method can be used to detect cancer-related germline (inherited) mutations in a cancer patient or an individual suspected of having a cancer-predisposing syndrome, wherein the detection of at least one germline mutation is used as an indicator of a predisposition to cancer. Thus, in another embodiment, the method further comprises diagnosing the patient or individual with an inherited cancer-predisposing syndrome, thereby enabling early medical intervention, treatment regimen selection, and close monitoring.
[0172] In a different embodiment, the method can be used to detect donor-derived cell-free DNA (dd-cfDNA) in transplantation as a potential biomarker of rejection. Donor-derived cell-free DNA is cfDNA derived from the transplanted organ that is foreign to the patient. The dd-cfDNA concentration increases even before creatinine levels start to rise, which can enable early diagnosis and adequate treatment of transplant injury, thus avoiding premature loss of the graft (Martuszewski A et al. (2021), J Clin Med, 10(2):192).
[0173] In one embodiment of the application, the present application relates to an in vitro method of detecting donor-derived cell-free DNA after transplantation, the method comprising the steps of:
[0174] (i) providing a human sample, preferably a blood sample, more preferably a plasma sample, a serum sample or a buffy coat sample, from a subject;
[0175] (ii) preparing a first nucleic acid sequencing library from nucleic acids present in the sample;
[0176] (iii) hybridizing one or more TAC oligonucleotides from a first pool of TAC oligonucleotides (TAC oligonucleotide-1) to the first nucleic acid library, thereby isolating a first subset of library nucleic acid molecules;
[0177] (iv) sequencing the first subset of library nucleic acid molecules and comparing the determined sequences to a human reference sequence, thereby creating a first set of putative informative sequence variants (PISV1);
[0178] (v) hybridizing one or more TAC oligonucleotides from a second, smaller pool of TAC oligonucleotides (TAC oligonucleotide-2) comprising TAC oligonucleotides specific for nucleic acid molecules of the first set of putative informative sequence variants to the first nucleic acid molecule library, thereby isolating a second subset of library nucleic acid molecules;
[0179] (vi) sequencing the second subset of library nucleic acid molecules and comparing the determined sequences to a human reference sequence, thereby creating a second set of putative informative sequence variants (PISV2);
[0180] (vii) analyzing the putative informative sequence variants (PISV2), thereby diagnosing, prognosing, drug response controlling a specific human disease.
[0181] The methods disclosed herein can also be used to detect sepsis in cell-free DNA from patients suspected of having a bloodstream infection. Plasma microbial cell-free DNA biomarkers can identify patients with sepsis. Cell-free DNA can originate from cells undergoing necrosis or apoptosis, as well as from pathogens. Assessing for sepsis through cell-free DNA detection can improve and expedite diagnosis in the ICU, facilitating medical intervention.
[0182] In step A of the method, in one embodiment, a DNA sequencing library is prepared from a sample of a subject having a tumor, suspected of having a tumor, or who has undergone tumor resection. Subsequently, the sequencing library is hybridized to a pool of targeted capture oligonucleotides (TAC oligonucleotides) targeting regions of interest, wherein the regions of interest are regions comprising mutations known to be associated with cancer. In one embodiment, one sequencing library is constructed. In another embodiment, two or more sequencing libraries are prepared in parallel and pooled together prior to hybridization.
[0183] In one embodiment, after preparing one sequencing library or more than two sequencing libraries, the region of interest is enriched by the following steps: (a) hybridizing the TAC oligonucleotide pool to the sequencing library; and (b) isolating the sequences of the sequencing library that bind to the TAC oligonucleotide. To facilitate the isolation of the desired enriched sequences, the TAC oligonucleotide is typically modified to enable the sequences that hybridize to the TAC oligonucleotide to be isolated from the sequences that do not hybridize to the TAC oligonucleotide. In one embodiment, this is achieved by immobilizing the TAC oligonucleotide on a solid support, thereby enabling the sequences that bind to the TAC oligonucleotide to be physically separated from the sequences that do not bind to the TAC oligonucleotide. For example, each sequence in the TAC oligonucleotide pool can be labeled with biotin, and then the oligonucleotide pool can be bound to beads coated with a biotin-binding material such as streptavidin or avidin. In a preferred embodiment, the TAC oligonucleotide is labeled with biotin and bound to streptavidin-coated magnetic beads. However, one skilled in the art will appreciate that other affinity binding systems are known in the art and can be used in place of biotin-streptavidin / avidin. For example, an antibody-based system can be used in which the TAC oligonucleotide is labeled with an antigen and then bound to beads coated with an antibody. In addition, the TAC oligonucleotide can incorporate a sequence tag at one end and can be bound to a solid support through a complementary sequence that hybridizes to the sequence tag on the solid support. In addition, other types of solid supports can be used in place of magnetic beads, such as polymeric beads. In one embodiment, the TAC oligonucleotide is not provided in a bound form and can exist freely in solution. In a preferred embodiment, the hybridization is a two-step hybridization in which the TAC oligonucleotide is hybridized to the enriched sequencing library after the first elution step. After the region of interest is enriched using the TAC oligonucleotide to form an enriched library, in one embodiment, the members of the enriched library are eluted from the solid support and amplified and sequenced using standard methods known in the art. Standard Illumina next generation sequencing (NGS) technology is typically used to provide very accurate counts in addition to sequence information, but other sequencing technologies can also be used. In one embodiment, the enriched library is amplified and sequenced to obtain at least 10,000 cfDNA nucleic acid molecules per region of interest. In other embodiments of the method, a sequencing DNA library is also prepared from the corresponding buffy coat, hybridized to the selected TAC oligonucleotide pool and sequenced to obtain at least 10,000 cfDNA nucleic acid molecules per region of interest.
[0184] In one embodiment, the sequencing data obtained from the sequencing step is processed by a computer system to align at least a portion of the sequenced cfDNA nucleic acid molecules to the hg19 human reference genome assembly. In another embodiment, alignment to any other reference genome, preferably a human genome, can be used.
[0185] After alignment, a consensus sequence is generated for each UMI family (http: / / fulcrumgenomics.github.io / fgbio / ). In another embodiment, the alignment in step B is performed using the patient-specific consensus sequence generated in step A.
[0186] Subsequently, variant calling is performed with ultra-high sensitivity by identifying a plurality of single nucleotide variant sites with respect to a reference genome sequence, wherein, in one embodiment, at least 1 and at most 25% of sequenced cfDNA nucleic acid molecules covering one or more single nucleotide variant sites support the variant nucleotide. As used herein, the term "variant nucleotide" generally refers to a single base pair mismatch compared to the reference sequence. In other embodiments of the method, a plurality of single nucleotide variant sites are identified with respect to a reference genome sequence, wherein at least 1 and at most 10%, 20%, 30%, or 40% of sequenced cfDNA nucleic acid molecules covering one or more single nucleotide variant sites support the variant nucleotide. Filtering is then performed using alignment and sequencing quality scores and a threshold calculated by estimating the distribution of false allele frequencies for each possible substitution covered by the TAC oligonucleotide pool using a set of normal reference samples previously diagnosed as not having cancer. In different embodiments, in addition to a plurality of single nucleotide variant sites, the method can also identify a plurality of small insertions and deletions.
[0187] After the filtering step, a likelihood statistic is calculated for each candidate somatic single nucleotide variant site using a regression model. In one embodiment, the regression model is a logistic regression model and the likelihood statistic is the log odds ratio of the probability that the candidate is a true somatic variant versus the probability that the candidate is a false positive. In one embodiment, the regression model is trained using at least 7,000,000 aligned cfDNA nucleic acid molecules from 71 normal reference samples previously diagnosed as not having cancer and samples from 45 cancer patients (non-small cell lung cancer, NSCLC) with at least three known somatic variants previously characterized using at least one tissue sample. The regression model includes at least one of the following covariates:
[0188] R1) the average alignment quality of the aligned cfDNA nucleic acid molecules;
[0189] R2) the ratio of high quality to low quality aligned cfDNA nucleic acid molecules supporting the variant nucleotide;
[0190] R3) the average distance of the single nucleotide variant site to one or more aligned cfDNA nucleic acid molecule endpoints;
[0191] R4) the Levenshtein distance of the sequence of one or more aligned cfDNA nucleic acid molecules to a reference genome sequence;
[0192] R5) the frequency of one or more single nucleotide variant sites in a normal reference population;
[0193] R6) the proportion of single nucleotide variants in sequenced cfDNA nucleic acid molecules that support phasing with nearby single nucleotide polymorphisms, if present, wherein the interval between the single nucleotide variants and the single nucleotide polymorphisms is at most 110 bp;
[0194] R7) the frequency of single nucleotide variants in the Catalogue of Somatic Mutations in Cancer (COSMIC);
[0195] R8) the predicted functional impact.
[0196] In one embodiment, the plurality of candidate somatic single nucleotide variant sites are first ranked in order of log-likelihood ratio from high to low, and then the top 40 candidate somatic single nucleotide variant sites are selected for subsequent analysis. In one aspect, a small number of single nucleotide variant sites are selected for targeted deep resequencing, thereby significantly reducing cost. In a preferred embodiment, the top 40 candidate somatic single nucleotide variant sites are selected for targeted deep resequencing. In various other embodiments, the top 20, 30, 40, 50, 60, 70, 80, 90, or 96 candidate somatic single nucleotide variant sites are selected for targeted deep resequencing. Selecting the top 40 single nucleotide variant sites marks the completion of the first step (Step A) of the method.
[0197] The second step (Step B) of the method uses the same sequencing library as Step A or the same two or more DNA sequencing libraries, wherein the DNA sequencing libraries are rehybridized with a smaller pool of TAC oligonucleotides derived from the larger initial pool of TAC oligonucleotides in Step A, which are designed to capture genomic regions covering the top 40 selected candidate somatic single nucleotide variant sites in Step A, thereby resulting in an enriched library. In a preferred embodiment, the rehybridization is a two-step rehybridization, wherein the TAC oligonucleotides are second hybridized with the sequencing library after a first elution step.
[0198] In a preferred embodiment of the application, the second TAC oligonucleotide pool (TAC oligonucleotide-2) comprises a defined fraction of the first TAC oligonucleotide pool (TAC oligonucleotide-1) and comprises 0.1%, 0.25%, 0.5%, 0.75%, 1.0% or 5%, most preferably about 0.5% of the first TAC oligonucleotide pool (TAC oligonucleotide-1). This means that a selected group of TAC oligonucleotides or a selected TAC oligonucleotide pool is chosen from the first oligonucleotide pool as the smaller second oligonucleotide pool. In this embodiment, the composition of the first and second TAC oligonucleotide pools is different from each other, which logically leads to different enrichment steps. Thus, in a preferred embodiment, the first TAC oligonucleotide pool (TAC oligonucleotide-1) is different from the second TAC oligonucleotide pool (TAC oligonucleotide-2).
[0199] Following the step of rehybridization with the smaller TAC oligonucleotide pool, the enriched library is amplified and sequenced, resulting in at least 200,000 cfDNA nucleic acid sequences per target region in one embodiment. A sequencing DNA library is also prepared from the corresponding buffy coat, hybridized with the selected TAC oligonucleotide pool and sequenced, resulting in at least 200,000 cfDNA nucleic acid sequences per target region. This is intended to remove somatic variations derived from clonal hematopoiesis. In other embodiments of the method, a sequencing DNA library is also prepared from the corresponding buffy coat, hybridized with the selected TAC oligonucleotide pool and sequenced, resulting in at least 100,000 or 200,000 or 300,000 or 400,000 or 500,000 cfDNA nucleic acid molecules per target region. The sequencing data is aligned to a reference genome sequence. In one embodiment, the reference genome is the hg19 version. In another embodiment, the reference genome is the hg38 version. In another embodiment, the alignment to any other reference genome, preferably a human genome, can be used.
[0200] With age, the number of somatic mutations accumulated in human tissues increases. While most of these mutations have little or no functional impact, mutations conferring a fitness advantage to the cell can arise. When this process occurs in the hematopoietic system, a significant fraction of circulating blood cells can be derived from a single mutated stem cell. This phenomenon of derivation is called "clonal hematopoiesis" and is extremely prevalent in the elderly population.
[0201] The sequencing data is aligned to a reference genome sequence. In one embodiment, the reference genome is the hg19 version. In another embodiment, the reference genome is the hg38 version.
[0202] After alignment to the reference genome sequence, a classification model is used to determine whether a candidate somatic single nucleotide variant in a list of selected candidate somatic single nucleotide variants exists in the sample after removing any variants derived from clonal hematopoiesis. The classification model is an ensemble learning method. In one embodiment, the method is a Decision Random Forest. In other embodiments, the method is a Supporting Vector Machine, a Naive Bayes Classifier, a logistic regression model, or a neural network. The classification method is trained with at least 7,000,000 aligned cfDNA molecules from normal reference samples (previously not diagnosed with cancer) and samples of cancer patients that have previously been characterized with somatic variants using at least one tissue sample. The classification model includes at least one of the following covariates:
[0203] C1) the number of cfDNA nucleic acid molecules that have the variant allele;
[0204] C2) the average alignment quality of sequenced cfDNA nucleic acid molecules relative to the reference genome sequence;
[0205] C3) the ratio of high quality to low quality sequenced cfDNA nucleic acid molecules that support the variant;
[0206] C4) the proportion of sequenced cfDNA nucleic acid molecules that support phase of the single nucleotide variant with nearby single nucleotide polymorphisms (if present), where the interval between the single nucleotide variant and the single nucleotide polymorphism is at most 110 bp;
[0207] C5) the frequency of the single nucleotide variant in the Catalogue of Somatic Mutations in Cancer (COSMIC).
[0208] In one embodiment, the output of the classifier is an estimated class / status of each candidate somatic single nucleotide variant, i.e., detected / not detected. The presence of cell-free tumor DNA nucleic acid molecules in a sample is determined by applying the following rule: if more than 5% of the selected plurality of candidate somatic single nucleotide variants are detected in the targeted resequencing step (step B) of the method, the sample is determined to be positive (i.e., cell-free tumor DNA nucleic acid molecules are detected in the sample), otherwise the sample is determined to be negative (i.e., cell-free tumor DNA nucleic acid molecules are not detected in the sample). In other embodiments of the method, the presence of cell-free tumor DNA nucleic acid molecules is determined if more than 1%, 10%, 15%, 20%, or 25% of the selected plurality of single nucleotide variants in step A are also detected in step B.
[0209] EMBODIMENTS
[0210] METHOD STEPS
[0211] SAMPLE COLLECTION AND PREPARATION
[0212] To detect cell-free circulating tumor DNA nucleic acid molecules, the sample is a biological sample obtained from a subject having a tumor or suspected of having a tumor. In one embodiment, the DNA sample is obtained from a human subject. In one embodiment, the sample comprises cell-free tumor DNA (ctDNA). In a preferred embodiment, the oncology sample is a patient plasma sample prepared from the patient's peripheral blood. Thus, the sample can be a liquid biopsy sample obtained from a patient blood sample by a non-invasive means, thereby enabling early detection of cancer before detectable or appreciable tumor formation. In another embodiment, the sample is the patient's serum, buffy coat, urine, sputum, ascites, cerebrospinal fluid, or pleural effusion. In one embodiment, the oncology sample is a tissue sample having cancer or suspected of having cancer (e.g., tissue from a tumor biopsy). In another embodiment, the sample is a stool sample. In yet another embodiment, the oncology sample is a healthy cell sample from the patient, such as a buffy coat, buccal swab, healthy tissue near a tumor, or other source of healthy cells prepared from the patient's peripheral blood. Thus, the healthy cells can provide a source of DNA that can be used to detect germline mutations and compared to tumor DNA.
[0213] In one embodiment of the application, the buffy coat of the subject biological sample is processed according to the steps described in step B, and candidate somatic single nucleotide variants detected in both steps A and the buffy coat are then removed from the classification model to determine whether cell-free tumor-derived DNA nucleic acid molecules are present.
[0214] In one embodiment, plasma samples are obtained from human subjects with tumors or suspected of having tumors. For biological sample preparation, plasma DNA is typically extracted using standard techniques known in the art, non-limiting examples of which are Qiagen DNeasy Blood & Tissue extraction protocol. In another embodiment, cell-free DNA is isolated from plasma using standard techniques, non-limiting examples of which are Qiasymphony (Qiagen) protocol or any other method known in the art.
[0215] Sequencing library preparation
[0216] Following isolation, in one embodiment, cell-free DNA in the sample is used to construct a sequencing library, enabling the sample to be compatible with downstream sequencing technologies such as NGS. Typically, this involves ligation of adapters to the ends of the cell-free DNA fragments, followed by amplification. Sequencing library preparation kits are commercially available. In another embodiment, nuclear DNA (non-limiting examples of which are DNA extracted from tissue or buffy coat) is fragmented using standard techniques, including but not limited to sonication. The fragmented nuclear DNA is then subjected to the downstream processing protocol for cell-free DNA described in this section. In one embodiment, one sequencing library is prepared. In another embodiment, two or more sequencing libraries are prepared in parallel.
[0217] DNA extracted from plasma samples is used to construct a sequencing library. Standard library preparation methods are used with the following modifications: a negative control extraction library is prepared separately to monitor for any contamination introduced during the process. In this step, 5’ and 3’ overhangs are filled in with the addition of 12 units of T4 polymerase (NEB) in a 10 μΐ reaction, 5 units of Taq polymerase is used to add an adenine to the 3’ end of each fragment; 40 units of T4 polynucleotide kinase (NEB) is used to ligate 5’ phosphate groups, followed by incubation at 20°C for 30 minutes and then 65°C for 30 minutes.
[0218] Subsequently, P5 and P7 adapters with double-stranded unique molecular identifiers (UMIs) (IDT CS Adapters) were ligated to both ends of the DNA using 5 units of T4 DNA ligase (NEB) in a 40 μΐ reaction at room temperature for 15 minutes at a dilution ratio of 1 :5, followed by purification using Ampure beads at a 1.0x ratio. Library amplification was performed in a 50 μΐ reaction using fusion polymerase (Herculase II Fusion DNA Polymerase (Agilent Technologies) or any other polymerase known in the art) under the following cycling conditions: 98°C for 3 minutes; followed by 11 cycles of 98°C for 30 seconds, 60°C for 30 seconds, 72°C for 30 seconds, and finally 72°C for 3 minutes (modified from Koumbaris, G. et al. (2016) Clinical Chemistry, 62(6), pp. 848-855). The final library product was purified using Ampure beads at a 1.5x ratio.
[0219] Design and preparation of target capture oligonucleotides (TAC oligonucleotides)
[0220] This example describes a method for the preparation of custom TAC oligonucleotides for the detection of cell-free tumor-derived DNA in plasma. The genomic target loci used for the design of TAC oligonucleotides were selected based on their GC content and distance to repetitive elements (at least 50 bp apart). The length of TAC oligonucleotides can vary.
[0221] In a preferred embodiment, each sequence in the TAC oligonucleotide pool is 150-260 base pairs in length. In various other embodiments, each sequence in the TAC oligonucleotide pool is 100-200 base pairs, 200-260 base pairs, 100-350 base pairs, 100-500 base pairs, or 100-1000 base pairs in length, or any combination thereof. However, one of ordinary skill in the art will appreciate that many more possible length ranges exist. The TAC oligonucleotides are prepared by singleplex polymerase chain reaction using standard Taq polymerase, primers designed for amplification of the target locus, and normal DNA as a template. All custom TAC oligonucleotides are generated using the following cycling conditions: 95 °C for 3 minutes; 40 cycles of 95 °C for 15 seconds, 60 °C for 15 seconds, 72 °C for 12 seconds, and finally 72 °C for 12 seconds. Subsequently, they are verified by agarose gel electrophoresis and purified using standard PCR purification kits, such as Qiaquick PCR purification kit (Qiagen) or NucleoSpin 96 PCR purification kit (Mackerey Nagel) or Agencourt AMPure XP PCR purification kit (Beckman Coulter). Concentrations are determined using a Nanodrop (Thermo Scientific).
[0222] One of ordinary skill in the art will recognize that the TAC oligonucleotides can be obtained by other means, such as but not limited to solid phase oligonucleotide synthesis, semiconductor-based DNA synthesis, silicon-based DNA synthesis, enzymatic DNA synthesis, or any commercially available method.
[0223] Biotinylation of TAC oligonucleotides
[0224] TAC oligos for hybridization were prepared following the previously described method (Koumbaris, G. et al. (2016) Clinical Chemistry, 62(6), pp. 848-855): blunt ends were first generated using the Quick Blunting kit (NEB) and incubated at room temperature for 30 minutes. The reaction product was then purified using the MinElute kit (Qiagen) and ligated to a biotin linker using the Quick Ligation kit (NEB) in a 40 μΐ reaction at room temperature for 15 minutes. The reaction product was purified using the MinElute kit (Qiagen) or Ampure beads (Beckman Coulter) and denatured to single stranded DNA before immobilization on streptavidin coated magnetic beads (Invitrogen). In one embodiment, the TAC discovery oligos are free in solution and not bound to any solid support. In another embodiment, the TAC oligos are readily presented as single stranded.
[0225] TAC oligos hybridization
[0226] The amplified libraries were mixed with blocking oligos (Koumbaris, G. et al. (2016) Clinical Chemistry, 62(6), pp. 848-855) (200 1-1M), 50 μg of Cot-1 DNA (Invitrogen), 50 μg of salmon sperm DNA (Invitrogen), 2x Agilent hybridization buffer, 10x Agilent blocking reagent or any other hybridization solution, and the DNA strands were denatured by heating at 98°C for 3 minutes. After denaturation, a 30 minutes incubation step at 37°C was performed to block repetitive elements and adapter sequences. The resulting mixture was then added to the biotinylated TAC oligos. All samples were incubated at 66°C for 4-48 hours in a rotating incubator. After incubation, the beads were washed following the previously described method and the DNA was eluted by heating (Koumbaris, G. et al. (2016) Clinical Chemistry, 62(6), pp. 848-855). The eluted product was amplified using the outer binding adapter primers. The enriched amplified product was recaptured using the same decoy pool and amplified again after elution following the same protocol described above. The double enriched amplified product was equimolar mixed and sequenced on a suitable platform.
[0227] In another embodiment, the amplified library is transferred to a well plate containing a pre-mixed solution and biotinylated TAC oligonucleotides in the absence of magnetic beads, the mixture is incubated in a thermal cycler at 98°C for 4 minutes to denature the DNA strands, followed by 65°C for 24 hours. After incubation, beads coated with streptavidin or avidin, or other biotin-binding material, are added to the mixture, and a second incubation step at 65°C for 35 minutes is performed. After incubation, the beads are washed according to the previously described method, and the DNA is eluted by heating (Koumbaris, G. et al. (2016) Clinical Chemistry, 62(6): 848-855). The elution product is amplified using the outer binding linker primers. In a second hybridization step, the enriched amplification product is recaptured using the same decoy pool, and after a second elution, amplification is performed using the same protocol described above. The double-enriched amplification product is mixed equimolarly, and sequenced on a suitable platform.
[0228] In one embodiment, the plurality of TAC oligonucleotide families used in the method bind to a plurality of regions known to be associated with cancer (referred to herein as target regions). As used herein, the target regions refer to regions carrying point mutations known to be associated with cancer. The regions are extracted from the COSMIC database. A large number of well-defined catalogs of cancer-associated mutations are known in the art, referred to as COSMIC (Catalogue of Somatic Mutations in Cancer), which are described, for example, in Forbes, S.A. et al. (2016) Curr. Protocol Hum. Genetic 91: 10.11.1-10.11.37, Forbes, S.A. et al. (2017) Nucl. Acids Res. 45:0777-0783, and Prior et al. (2012) Cancer Res. 72:2457-2467. The COSMIC database is publicly available at www.cancer.sanger.ac.uk. In addition to the COSMIC catalog, other compendia of tumor biomarker mutations are described in the art, non-limiting examples of which include the ENCODE project (which describes mutations in oncogene regulatory sites, see, e.g., Shar, N.A. et al. (2016) Mol. Cane. 15:76) and ClinVar (a database of genomic variations related to human health by the National Center for Biotechnology Information (NCBI)). The ClinVar database is publicly available at www.ncbi.nlm.nih.gov / clinvar.
[0229] The target region is enriched by hybridizing the pool of TAC oligonucleotides to the sequencing library, followed by isolating the sequences of the sequencing library that are bound to the TAC oligonucleotides. To facilitate isolation of the desired enriched sequences, the TAC oligonucleotide sequences are often modified so that those sequences that hybridize to the TAC oligonucleotides can be separated from those that do not. Typically, this is accomplished by immobilizing the TAC oligonucleotides to a solid support. This allows those sequences that are bound to the TAC oligonucleotides to be physically separated from those that are not. For example, the sequences in the pool of TAC oligonucleotides can be labeled with biotin, and then the pool of oligonucleotides can be bound to beads coated with a biotin-binding material such as streptavidin or avidin. In a preferred embodiment, the TAC oligonucleotides are labeled with biotin and bound to magnetic beads coated with streptavidin. In one embodiment, the biotin can be chemically attached to the primers used to generate the TAC oligonucleotides. In a second embodiment, the TAC oligonucleotides can be generated by biotinylating the pool of sequences that are capable of hybridizing to the target region. In another embodiment, the biotin can be added during the synthesis of the TAC oligonucleotides.
[0230] In certain embodiments, the members of the sequencing library that bind to the pool of TAC oligonucleotides are fully complementary to the TAC oligonucleotides. In other embodiments, the members of the sequencing library that bind to the pool of TAC oligonucleotides are partially complementary to the TAC oligonucleotides. For example, in certain cases, it can be desirable to utilize and analyze data from DNA fragments that result from the enrichment process but that do not necessarily belong to the target genomic region (i.e., these DNA fragments can bind to the TAC oligonucleotides due to the presence of partial homology (partial complementarity) to the TAC oligonucleotides, and when sequenced will produce very low coverage across the genome in non-TAC oligonucleotide coordinates).
[0231] After the target sequences are enriched using the TAC oligonucleotides to form an enriched library, the members of the enriched library are eluted from the solid support and amplified and sequenced using standard methods known in the art.
[0232] The TAC oligonucleotide pools and TAC oligonucleotide families used in the methods for detecting cell-free tumor-derived DNA in plasma can comprise any of the design features described herein. In various embodiments, the TAC oligonucleotide pool comprises at least 5, 10, 50, or 100 or more different TAC oligonucleotide families. In various embodiments, each TAC oligonucleotide family comprises at least 2, at least 3, at least 4, or at least 5 different member sequences. In one embodiment, each TAC oligonucleotide family comprises 4 different member sequences. In various embodiments, the stagger of the start and / or end positions of the member sequences within a TAC oligonucleotide family relative to a reference coordinate system of the target genomic sequence is at least 5 base pairs, at least 10 base pairs, or 5-10 base pairs.
[0233] Alignment to human genome and consensus sequence generation
[0234] In one embodiment, for each sample, the following bioinformatics analysis pipeline is applied to align the sequenced DNA fragments of the sample to the human reference genome. The targeted paired-end reads resulting from the NGS run are processed using the cutadapt software (Martin, M. et al. (2011) EMB.net Journal 17.1) to remove the adapter sequences and low quality reads (Q-value < 25). The quality of the raw and / or processed reads, as well as any descriptive statistics that aid in the assessment of the quality of the sample sequencing run, are obtained using the FastQC software (Babraham Institute (2015) FastQC) and / or other custom developed software. The processed reads of at least 25 bases in length are processed in the UMI-based bioinformatics suite of FGBIO (https: / / bio.tools / fgbio) to align and create UMI families (with the same start coordinate, end coordinate, and UMI sequence) and generate a consensus sequence for each UMI family in binary alignment format. Where applicable, sequencing output results belonging to the same sample but processed on different sequencing lanes are merged into a single sequencing output file.
[0235] Data analysis
[0236] Variant calling is performed in an ultra-sensitive manner by identifying a plurality of single nucleotide variant sites relative to a reference genome sequence, wherein at least 1 and at most 25% of sequenced cfDNA nucleic acid molecules covering one or more single nucleotide variant sites support the variant nucleotide (mismatch from the reference sequence). In other embodiments of the method, a plurality of single nucleotide variant sites are identified relative to a reference genome sequence, wherein at least 1 and at most 10%, 20%, 30%, or 40% of sequenced cfDNA nucleic acid molecules covering one or more single nucleotide variant sites support the variant nucleotide. Filtering is then performed using alignment quality scores, sequencing quality scores, and a threshold calculated by estimating the distribution of variant allele frequencies for false variants for each possible substitution covered by the TAC oligonucleotide pool using a set of normal reference samples previously diagnosed as not having cancer. Thereafter, a likelihood statistic is calculated for each candidate somatic single nucleotide variant using a regression model. In one embodiment, the regression model is a logistic regression model and the likelihood statistic is the log odds ratio of the probability that the candidate is a true somatic variant versus the probability that the candidate is a false positive. In one embodiment, the regression model is trained using at least 7 million aligned cfDNA nucleic acid molecules from 71 normal reference samples previously diagnosed as not having cancer and samples from 45 cancer patients having at least three known somatic variants previously characterized using at least one tissue sample. The regression model includes at least one of the following covariates:
[0237] R1) average alignment quality of aligned cfDNA nucleic acid molecules;
[0238] R2) ratio of high quality to low quality aligned cfDNA nucleic acid molecules supporting the variant nucleotide;
[0239] R3) average distance of the single nucleotide variant site to one or more aligned cfDNA nucleic acid molecule endpoints;
[0240] R4) Levenshtein distance of the sequence of one or more aligned cfDNA nucleic acid molecules to the reference genome sequence;
[0241] R5) frequency of the one or more single nucleotide variant sites in a population of normal reference samples;
[0242] R6) proportion of sequenced cfDNA nucleic acid molecules supporting a single nucleotide variant phased with a nearby single nucleotide polymorphism, if present, wherein the single nucleotide variant and single nucleotide polymorphism are separated by at most 110 bp;
[0243] R7) frequency of the single nucleotide variant in the Catalogue of Somatic Mutations in Cancer (COSMIC).
[0244] R8) predicted functional impact.
[0245] In one embodiment, the plurality of candidate somatic single nucleotide variant sites are first ranked in descending order of log-likelihood ratio value, and the top 40 sites are selected for subsequent analysis. In other embodiments of the method, the top 20, 30, 40, 50, 60, 70, 80, 90, or 96 candidate somatic single nucleotide variant sites are selected. This is the final part of the first step (Step A) of the method.
[0246] The second step (Step B) comprises the same DNA sequencing library, wherein, in one embodiment, the DNA sequencing library is re-hybridized with a smaller pool of TAC oligonucleotides designed to capture the genomic regions covering the top 40 candidate somatic single nucleotide variant sites selected in Step A, resulting in an enriched library, and further wherein the enriched library is amplified and sequenced to yield at least 200,000 cfDNA nucleic acid sequences per target region. In other embodiments of the method, a sequencing DNA library is also prepared for the corresponding buffy coat, hybridized with the selected pool of TAC oligonucleotides and sequenced to yield at least 100,000 or 200,000 or 300,000 or 400,000 or 500,000 cfDNA nucleic acid molecules per target region. A classification model is used to determine whether a candidate somatic single nucleotide variant in the list of selected candidate somatic single nucleotide variants is present in the sample after removal of variants resulting from clonal hematopoiesis, if any. The classification model is an ensemble learning method. In one embodiment, the method is a decision random forest. In other embodiments, the method is a support vector machine, a naive Bayes classifier, a logistic regression model, or a neural network. The classification method is trained with at least 7,000,000 aligned cfDNA nucleic acid molecules from normal reference samples that were not previously diagnosed with cancer, and samples from cancer patients that have previously been characterized with at least one tissue biopsy for somatic variants. The classification model comprises at least one of the following covariates:
[0247] C1) the number of cfDNA nucleic acid molecules that harbor the variant allele;
[0248] C2) the average alignment quality of sequenced cfDNA nucleic acid molecules relative to the reference genome sequence;
[0249] C3) the ratio of high-quality to low-quality sequenced cfDNA nucleic acid molecules supporting the variant;
[0250] C4) the proportion of single nucleotide variants in the sequenced cfDNA nucleic acid molecules that support phasing with nearby single nucleotide polymorphisms, if present, wherein the interval between the single nucleotide variant and the single nucleotide polymorphism is at most 110 bp;
[0251] C5) the frequency of the single nucleotide variant in the Catalogue of Somatic Mutations in Cancer (COSMIC).
[0252] The output of the classifier is an estimated class / status of each candidate somatic single nucleotide variant, i.e., detected / not detected. The presence of cell-free tumor DNA nucleic acid molecules in a sample is determined by applying the following rule: if more than 5% of the selected plurality of candidate somatic single nucleotide variants are detected in the targeted resequencing experiment, the sample is called positive (i.e., cell-free tumor DNA nucleic acid molecules are detected in the sample), otherwise the sample is called negative (i.e., cell-free tumor DNA nucleic acid molecules are not detected in the sample). In other embodiments of the method, the presence of cell-free tumor DNA nucleic acid molecules is determined if more than 1%, 10%, 15%, 20%, or 25% of the selected plurality of single nucleotide variants in step A are detected in step B.
[0253] In another embodiment, step B is performed using a second sequencing library prepared from the same human sample.
[0254] Phasing of candidate somatic single nucleotide variants with germline heterozygous single nucleotide polymorphisms
[0255] Phasing candidate somatic single nucleotide variants with germline heterozygous single nucleotide polymorphisms provides an effective method for artifact elimination Figure 1). In one embodiment of the method, a plurality of aligned sequencing reads (cfDNA molecules) are used to detect heterozygous germline single nucleotide polymorphisms (SNPs) in the sample if the variant allele frequency of the genomic site is between 0.3 and 0.7, and the p-value of the binomial test (variant count based on success probability = 0.5) is > 0.001. Afterwards, a list of candidate somatic single nucleotide variant sites is created for each SNP. The list contains candidate somatic single nucleotide variants that are less than or equal to 110 bp apart, such that a single aligned sequencing read can cover both genomic sites, and at least 3 aligned sequencing reads cover both genomic sites. Using a proprietary algorithm (Python v2.7), each SNP detected with a non-empty list is iteratively computed: (a) the ratio of aligned sequencing reads that support the variant allele at the candidate somatic single nucleotide variant site and the reference allele at the SNP site, out of the total number of aligned sequencing reads that cover the SNP and the candidate somatic single nucleotide variant; or (b) the proportion of aligned sequencing reads that support the variant allele at the candidate somatic single nucleotide variant site and the alternative / variant allele at the SNP site. The score H is defined as the maximum of (a) and (b). In one embodiment of the method, the candidate somatic single nucleotide variant is excluded if H is less than 0.95. In other embodiments of the method, the candidate somatic single nucleotide variant is excluded if H is less than 0.9 or 0.8.
[0256] Example 1
[0257] In one embodiment of the method, DNA sequencing libraries are prepared for 5 Seraseq reference samples, which contain 29 single nucleotide variants with very low variant allele frequencies (VAF) of 0.02-0.3% (sample A: 7 single nucleotide variants with VAF of 0.02-0.07%; sample B: 12 single nucleotide variants with VAF of 0.1-0.3%; sample C: 5 single nucleotide variants with VAF of 0.08-0.15%; sample D: 5 single nucleotide variants with VAF of 0.2-0.7%; sample E: wild type negative control). Each sample’s sequencing library contains unique molecular identifiers for uniquely labeling each molecule. The DNA sequencing libraries are hybridized in solution with a pool of TAC oligonucleotides with an average length of 250 bp covering the total length of 500,000 bp of human reference genome version 19 (hg19). The pool of TAC oligonucleotides is enriched for known hotspots of somatic single nucleotide variants in non-small cell lung cancer (adenocarcinoma and squamous cell carcinoma) tissue specimens. The estimated median mutation rate for stage I-III lung cancer tissue is 9 / Mb (range: 7-13 / Mb) (van de Heuvel et al.
[2021] Respir Res 22:302; The Cancer Genome Atlas (TCGA); https: / / www.cancer.gov / tcg / research / genome- sequencing / tcga). The enriched libraries are sequenced at an average unique depth of at least 5,000-fold with a sequencing data volume of 20 Gb on a Novaseq sequencer. The pipeline for alignment and consensus sequence generation to suppress errors is prepared using samtools, bwamem, and fgbio bioinformatics suite. A super-sensitive variant calling method is applied, which selects a substitution as a candidate true somatic single nucleotide variant site even if at least 1 read supports the variant allele (and at most 25% to avoid germline SNPs). A background artifact model is applied to remove false signals that are very likely false positive signals (not true somatic variants). The background model estimates the number of occurrences of each false substitution (artifact) in the normal reference sample population and the variant allele frequency distribution of each of the artifacts / false signals. Very common artifacts are removed / filter out. In one embodiment, a common artifact is defined as an artifact that is present in more than 5% of the normal reference samples. The list of candidate somatic variants is further reduced by removing candidate variants with alignment quality of 0 or removing candidate variants with a ratio of high-quality to low-quality aligned cfDNA nucleic acid molecules supporting the variant below a pre-set threshold. In one embodiment, the pre-set threshold is 2.The list of candidate somatic variants is further reduced by removing candidate somatic single nucleotide variants that are less than 110 bp away from at least one single nucleotide polymorphism detected in the sample and cannot be phased with the at least one single nucleotide polymorphism. The log odds ratio of each candidate somatic variant being a true variant is calculated using the trained logistic regression model including covariates R1-R7, and the top 40 candidates with the highest odds are selected for targeted resequencing. After alignment and consensus sequence generation using at least 200,000 raw read depth per targeted region, the number of variants detected in each sample is calculated using the trained decision random forest classification model including covariates C1-C5. A flow chart of the main steps of this embodiment of the present invention is shown in FIG. 1. Figure 2
[0258] The sensitivity achieved for sample A was 86% (6 out of 7 true somatic variants were detected; 95% confidence interval: 42-99.6%), for sample B was 100% (12 / 12; 95% confidence interval: 74-100%), for sample C was 100% (5 / 5; 95% confidence interval: 48-100%), and for sample D was 100% (5 / 5; 95% confidence interval: 48-100%). For the wild type negative control sample E, no known variants were detected.
[0259] Example 2
[0260] In another embodiment of the method, DNA sequencing libraries are prepared for 11 normal reference samples taken from healthy donors and 14 cancer samples taken from patients in stages I-IV (2 in stage I, 3 in stage II, 5 in stage III, 4 in stage IV) and unknown number of single nucleotide variations (NSCLC). The sequencing library of each sample contains unique molecular identifiers for uniquely labeling each molecule. The DNA sequencing libraries are hybridized in solution using a pool of TAC oligonucleotides with an average length of 250 bp covering the total length of 500000 bp of human reference genome version 19 (hg19). In one embodiment, the pool of TAC oligonucleotides is enriched in known hotspots of somatic single nucleotide variations in lung cancer (adenocarcinoma and squamous cell carcinoma) tissue specimens. The enriched libraries are sequenced at a unique average depth of at least 3000 fold with a sequencing data amount of 20 Gb on a Novaseq sequencer. The pipeline for alignment and consensus sequence generation to suppress errors is prepared using samtools, bwa mem and fgbio bioinformatics suite. A super-sensitive variant calling method is applied, i.e. a substitution is considered as a candidate true somatic single nucleotide variation even if at least 1 read supports the variant allele (and at most 25% to avoid germline SNPs). Very common artifacts are removed / filter out. In one embodiment, a common artifact is defined as an artifact present in more than 5% of normal reference samples. The list of candidate somatic variations is further reduced by removing candidate variations with alignment quality of 0, or removing candidate variations with a ratio of high quality to low quality aligned cfDNA nucleic acid molecules supporting the variation below a pre-set threshold. In one embodiment, the pre-set threshold is 2. The list of candidate somatic variations is further reduced by removing candidate somatic single nucleotide variations that are less than 110 bp away from at least one single nucleotide polymorphism detected in the sample and cannot be phased with the at least one single nucleotide polymorphism. The log odds ratio that each candidate somatic variation is true is calculated using a trained logistic regression model containing covariates R1-R7, and the top 40 candidate variations with the highest odds ratio are selected for targeted resequencing. Sequencing DNA libraries are also prepared for buffy coats of all samples, hybridized with the selected pool of TAC oligonucleotides, sequenced, and analyzed using the same pipeline to remove variations arising from clonal hematopoiesis. Subsequently, a trained decision random forest classifier (covariates C1-C5) calculates the number of variations detected in each sample. The results are as follows Figure 4The y-axis represents the number of somatic variants detected in each sample, with black bars corresponding to normal samples and grey bars corresponding to abnormal samples, and the x-axis represents the status of each sample (normal or cancer stage). The threshold for calling a sample positive is represented by the black horizontal solid line. The clinical specificity of this test was 100% (11 / 11; 95% confidence interval: 72-100%) and the clinical sensitivity was 71% (10 / 14; 95% confidence interval: 42-92%).
[0261] Example 3
[0262] A set of samples was processed using the method of the application, namely 40 normal reference samples taken from healthy donors and 20 cancer samples taken from patients in stages II-IV and for which the number of single nucleotide variants was not known (10 cases of non-small cell lung cancer, 10 cases of colorectal cancer). In one embodiment of the method, for each sample, two DNA sequencing libraries were prepared using DNA extracted from two independent aliquots (step A). The two sequencing libraries of the sample contained unique molecular identifiers for uniquely labeling each molecule. The two DNA sequencing libraries were hybridized in solution using a pool of TAC oligonucleotides with an average length of 250 bp covering a total length of 500000 bp of the human reference genome version 19. The enriched libraries were sequenced on a NovaSeq 6000 system (Illumina) with an average unique depth of at least 1500 times
[0263] The sequencing read processing pipeline was prepared using samtools, bwa mem and fgbio suite, respectively, including merging, alignment and generation of consensus sequences to suppress errors. The ultra-high sensitivity variant calling method was applied, i.e. a substitution was considered as a candidate true somatic single nucleotide variant even if at least 1 read supported the variant allele. In one embodiment of the method, variants satisfying a set of logical operators implemented in R were selected for targeted resequencing. The logical operators contained the following variables:
[0264] 1. Sequencing depth
[0265] 2. Variant count
[0266] 3. Variant allele frequency
[0267] 4. Strand bias test
[0268] 5. Average position of variant allele on reads
[0269] 6. Average base quality
[0270] 7. Ratio of high quality reads to low quality reads
[0271] 8. Average number of mismatches of reads containing variants
[0272] 9. Alignment quality
[0273] 10. Frequency of variants in COSMIC database
[0274] 11. Background noise level calculated from a set of normal samples previously not diagnosed with cancer
[0275] For each sample, two sets of candidate somatic variants are selected (each sequenced DNA sequencing library corresponds to a set). Thus, the presence of each unique candidate somatic variant in the two sets is assessed using a proportion statistical test. In various embodiments, a binomial test, a Fisher test or a Chi-square test is used. If the p-value of the statistical test is higher than a threshold, the set of candidate somatic variants is selected. In various embodiments, the threshold is set to 0.0001, 0.001, 0.005, 0.01, 0.05 or 0.1. The selected candidate somatic variants are ranked according to their frequency in the COSMIC database and up to the top 70 variants are selected for targeted resequencing (step B). The step B comprises the method of step A for identifying and filtering candidate somatic variants in two targeted resequenced DNA libraries of each sample. In one embodiment of the method, variants that satisfy a set of logical operators constructed using variables 1-11 listed above are selected. Meanwhile, a sequencing DNA library is also prepared for the buffy coat of all samples and hybridized with the selected TAC oligonucleotide pool (step B) together with the plasma library, sequenced and analyzed using the same pipeline to remove variants generated by clonal hematopoiesis. In various embodiments of the method, a sample is classified as positive or negative if at least 1, 2 or 3 candidate somatic variants are selected. In another embodiment, each sample is classified using a function of the number of selected candidate somatic variants, variant allele frequency or any other variable in the list 1-11 above. For a test population consisting of 40 normal samples and 20 cancer samples, the clinical specificity and sensitivity of the test are 100% (95% confidence interval: 91-100%) and 100% (95% confidence interval: 83-100%) respectively.
[0276] In another aspect, the present application provides kits for carrying out the methods of the present application. In one embodiment, the kit comprises a container holding a pool of TAC oligonucleotides and instructions for carrying out the method. In one embodiment, the TAC oligonucleotides are provided in a form that enables their binding to a solid support, for example biotinylated TAC oligonucleotides. In another embodiment, the TAC oligonucleotides are provided with a solid support, for example biotinylated TAC oligonucleotides are provided with streptavidin-coated magnetic beads.
[0277] In one embodiment, the kit comprises a container holding a pool of TAC oligonucleotides and instructions for carrying out the method, wherein the pool of TAC oligonucleotides comprises a plurality of families of TAC oligonucleotides, wherein each family of TAC oligonucleotides comprises a plurality of member sequences, wherein each member sequence binds the same target genomic sequence but has a different starting and / or ending position relative to a reference coordinate system of the target genomic sequence, and further wherein
[0278] (i) each member sequence in each family of TAC oligonucleotides is 100-500 base pairs in length, each member sequence having a 5' end and a 3' end;
[0279] (ii) each member sequence binds the same target genomic sequence and is at least 50 base pairs from a 5' end and a 3' end of a region containing a copy number variation (CNV), a segmental duplication, or a repetitive DNA element;
[0280] (iii) the GC content of the pool of TAC oligonucleotides is 19% to 80% as determined by calculating the GC content of each member of each family of TAC oligonucleotides.
[0281] In addition, any one or several of the various features described herein with respect to TAC oligonucleotide design and structure can be incorporated into the TAC oligonucleotides included in the kit.
[0282] In various other embodiments, the kit can comprise additional components for carrying out other aspects of the method. For example, in addition to the pool of TAC oligonucleotides, the kit can comprise one or more of the following components: (i) one or more components for isolating cell-free DNA from a biological sample; (ii) one or more components for preparing a sequencing library; (iii) one or more components for enriching a sequencing library; (iv) one or more components for amplifying and / or sequencing an enriched library; and / or (v) software for performing statistical analysis.
[0283] Preferably, the kit comprises TAC oligonucleotides for Step A and TAC oligonucleotides for Step B, wherein the TAC oligonucleotides for Step B are a subset of the TAC oligonucleotides for Step A.
Claims
1. An in vitro method of detecting somatic mutations in humans comprising the steps of: (i) providing a human sample, preferably a blood sample, more preferably a plasma sample, a serum sample or a buffy coat sample, from a subject; (ii) preparing a first nucleic acid sequencing library from nucleic acids present in said sample; (iii) hybridizing one or more TAC oligonucleotides from a first pool of TAC oligonucleotides (TAC oligonucleotide-1) to said first nucleic acid library, thereby isolating a first subset of library nucleic acid molecules; (iv) sequencing said first subset of library nucleic acid molecules and comparing the determined sequences to a human reference sequence, thereby creating a first set of putative informative sequence variants (PISV1); (v) hybridizing one or more TAC oligonucleotides from a second pool of TAC oligonucleotides (TAC oligonucleotide-2) comprising TAC oligonucleotides specific for nucleic acid molecules of said first set of putative informative sequence variants to said first nucleic acid molecule library, thereby isolating a second subset of library nucleic acid molecules; (vi) sequencing said second subset of library nucleic acid molecules and comparing the determined sequences to a human reference sequence, thereby creating a second set of putative informative sequence variants (PISV2); (vii) analyzing said putative informative sequence variants (PISV2), thereby detecting somatic mutations in said subject sample.
2. The method of claim 1, wherein, said second pool of TAC oligonucleotides (TAC oligonucleotide-2) is a subset of said first pool of TAC oligonucleotides.
3. The method of claim 1 or 2, wherein, each sequence in said first set of putative informative sequence variants (PISV1) is ranked according to an information value based on likelihood statistics.
4. The method of claims 1 to 3, wherein, each sequence in said first set of putative informative sequence variants (PISV1) is ranked according to an information value based on likelihood statistics, said likelihood statistics being an ensemble learning method, and wherein said method comprises: (i) the number of cfDNA nucleic acid molecules with variant alleles; (ii) the average alignment quality of sequenced cfDNA nucleic acid molecules to the reference genome sequence; (iii) the ratio of high quality to low quality sequenced cfDNA nucleic acid molecules supporting said variant; (iv) if a nearby single nucleotide polymorphism exists, the proportion of sequenced cfDNA nucleic acid molecules supporting a single nucleotide variant phased with a nearby single nucleotide polymorphism, wherein the interval between said single nucleotide variant and said single nucleotide polymorphism is at most 110 bp; (v) the frequency of a single nucleotide variant in a disease specific database.
5. The method of claims 1 to 4, wherein, TAC oligonucleotides in said first pool of TAC oligonucleotides (TAC oligonucleotide-1) and said second pool of TAC oligonucleotides (TAC oligonucleotide-2) have a length of 150 to 260 base pairs, wherein they are designed to bind to a region of interest, and wherein they have a 5’ end and a 3’ end.
6. The method of claims 1 to 5, wherein, The GC content of the TAC oligonucleotide pool is determined by calculating the GC content of each member of the TAC oligonucleotide pool and is between 19% and 80%.
7. The method of claims 1 to 6, wherein, The TAC oligonucleotide pool comprises a plurality of TAC oligonucleotide families each directed to a different target region, wherein each TAC oligonucleotide family comprises a plurality of member sequences, wherein each member sequence binds the same target region but has a different start and / or end position relative to a reference coordinate system of the target region, further wherein the start and / or end positions of the member sequences within a TAC oligonucleotide family are staggered by 5 to 10 base pairs relative to the reference coordinate system of the target region.
8. The method of claims 1 to 7, wherein, The second TAC oligonucleotide pool (TAC oligonucleotide-2) comprises a specified proportion of the first TAC oligonucleotide pool (TAC oligonucleotide-1) and is 0.1% or 0.25% or 0.5% or 0.75% or 1.0% of the first TAC oligonucleotide pool (TAC oligonucleotide-1), but most preferably is about 0.5%.
9. The method of any one of the preceding claims, wherein, The second TAC oligonucleotide pool is designed to capture regions covering somatic single nucleotide variant sites of the top 20 or 30 or 40 or 50 or 60 or 70 or 80 or 90 or 96 candidates selected in step A.
10. The method of any one of the preceding claims, wherein, Step vi) of claim 1 comprises amplifying and sequencing a second subset of the library nucleic acid molecules, such that at least 200,000 cfDNA nucleic acid sequences per target region are obtained.
11. The method of any one of the preceding claims, wherein, In addition to the plasma sample, the buffy coat fraction of the blood is analyzed to determine the proportion of sequences arising from clonal hematopoiesis.
12. The method of claim 1, wherein, The human sample is a blood sample or a serum sample or a buffy coat sample or a urine sample or a sputum sample or an ascites sample or a cerebrospinal fluid sample or a pleural effusion sample or a saliva sample or a bronchoalveolar lavage fluid, or an aspirate sample from different parts of the body.
13. The method of claim 1, wherein, The human sample is a tissue sample or a stool sample.
14. The method of claim 1, wherein, The method is used to detect donor-derived cell-free DNA.
15. The method of claim 1, wherein, The alignment in step B is performed using the patient-specific consensus sequence generated in step A.
16. The method of any one of the preceding claims, wherein, The single nucleotide variants are spaced apart from single nucleotide polymorphisms by at most 100 bp or 120 bp or 130 bp or 140 bp or 150 bp.
17. The method of any one of the preceding claims, wherein, The TAC oligonucleotide comprises a biotin modification.
18. A kit for performing the method of claim 1, wherein, The kit comprises a container comprising: i) TAC oligonucleotides; ii) one or more components for isolating cell-free DNA from a biological sample; iii) one or more components for preparing a sequencing library; iv) one or more components for amplifying and / or sequencing the enriched library; and optionally v) software for performing statistical analysis. The kit comprises a container comprising: i) TAC oligonucleotides; ii) one or more components for isolating cell-free DNA from a biological sample; iii) one or more components for preparing a sequencing library; iv) one or more components for amplifying and / or sequencing the enriched library; and optionally v) software for performing statistical analysis.
Citation Information
Patent Citations
Multiplexed parallel analysis of targeted genomic regions for non-invasive prenatal testing
US20160340733A1
Multiplexed parallel analysis of targeted genomic regions for non-invasive prenatal testing
WO2016189388A1