Nucleic acid molecule enrichment methods and compositions for sequencing
Through the adjustable target enrichment method of multiple sets of nucleic acid capture nucleic acids, the problems of low nucleic acid capture efficiency and insufficient detection sensitivity in early cancer in the prior art are solved, and efficient detection of circulating tumor DNA is achieved.
Patent Information
- Application Number
- CN202380061543.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-06-23
- Filing Date
- 2023-06-22
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art is difficult to effectively capture and enrich nucleic acid molecules, especially in early cancer detection, where the background noise of circulating tumor DNA is high, resulting in insufficient detection sensitivity.
By providing multiple sets of capture nucleic acids, nucleic acid molecules are enriched and sequenced separately, and the contact time and concentration of capture nucleic acids and nucleic acids are adjusted, so as to achieve adjustable sequencing depth for different nucleic acid sequences.
It improves the efficiency of nucleic acid capture and enrichment, enhances the detection sensitivity of circulating tumor DNA, and can identify cancer earlier.
Smart Images

Figure CN120167013A_ABST
Abstract
Description
[0001] Cross-reference
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 355,002, filed Jun. 23, 2022, which is hereby incorporated by reference in its entirety. BACKGROUND OF THE DISCLOSURE
[0003] The present disclosure generally relates to the capture or enrichment of nucleic acid molecules. Nucleic acid molecules can be captured or enriched and sequenced to determine the nucleic acid sequence. Based on the sequence, certain medical conditions can be analyzed. For example, sequencing can be used to screen for or monitor cancer. Such screening and monitoring can help improve outcomes because early detection leads to better outcomes as cancer can be eliminated before it has a chance to spread. SUMMARY OF THE DISCLOSURE
[0004] The present disclosure provides methods and systems for tunable target capture or enrichment of nucleic acid molecules.
[0005] In one aspect, the present disclosure provides a method comprising: (a) providing a sample from a subject, wherein the sample comprises a plurality of nucleic acids; (b) providing to the sample a first set of capture nucleic acids that enrich a first set of the plurality of nucleic acids to generate a sufficient amount of the first set of nucleic acids for sequencing the first set of nucleic acids to a first sequencing depth; (c) providing to the sample a second set of capture nucleic acids that enrich a second set of the plurality of nucleic acids to generate a sufficient amount of the second set of nucleic acids for sequencing the second set of nucleic acids to a second sequencing depth, wherein the first sequencing depth and the second sequencing depth are different; and (d) sequencing the first set of nucleic acids and the second set of nucleic acids to generate sequencing reads. In some embodiments, the plurality of nucleic acids are from a cell-free sample.
[0006] In some embodiments, the plurality of nucleic acids includes cell-free DNA (cfDNA) or cell-free RNA (cfRNA). In some embodiments, the plurality of nucleic acids includes circulating tumor DNA (ctDNA). In some embodiments, the first set of capture nucleic acids includes more nucleic acids than the second set of capture nucleic acids. In some embodiments, the concentration of the first set of capture nucleic acids in the sample is higher than the concentration of the second set of capture nucleic acids in the sample. In some embodiments, the method further includes contacting the first set of capture nucleic acids with the plurality of nucleic acids for a first contact duration and contacting the second set of capture nucleic acids with the plurality of nucleic acids for a second contact duration, wherein the first contact duration and the second contact duration are different. In some embodiments, the method further includes contacting the first set of capture nucleic acids with the plurality of nucleic acids for a first contact duration and contacting the second set of capture nucleic acids with the plurality of nucleic acids for a second contact duration, wherein the first contact duration and the second contact duration are the same or substantially the same.
[0007] In some embodiments, the first set of capture nucleic acids includes a first tiling density of 1x. In some embodiments, the first set of capture nucleic acids includes a first tiling density of 2x. In some embodiments, the first set of capture nucleic acids includes a first tiling density of 0.5x. In some embodiments, the first set of capture nucleic acids includes a first tiling density and the second set of capture nucleic acids includes a second tiling density, wherein the first tiling density and the second tiling density are different. In some embodiments, the first set of capture nucleic acids includes a first tiling density and the second set of capture nucleic acids includes a second tiling density, wherein the first tiling density and the second tiling density are the same or substantially the same. In some embodiments, the first tiling density is generated by overlapping sequences in the nucleic acids of the first set of capture nucleic acids.
[0008] In some embodiments, the first set of capture nucleic acids or the second set of capture nucleic acids comprises at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more nucleotides. In some embodiments, the first set of capture nucleic acids or the second set of capture nucleic acids comprises no more than 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or fewer nucleotides. In some embodiments, the nucleotide length of the first set of capture nucleic acids is shorter than the nucleotide length of the second set of capture nucleic acids. In some embodiments, the nucleotide length of the first set of capture nucleic acids is longer than the nucleotide length of the second set of capture nucleic acids. In some embodiments, the first set of capture nucleic acids has incomplete complementarity with the first set of nucleic acids. In some embodiments, the first set of capture nucleic acids has at least one mismatched base in the nucleic acid region with the first set of nucleic acids. In some embodiments, the first set of capture nucleic acids has at least two mismatched bases in the nucleic acid region with the first set of nucleic acids. In some embodiments, the first set of capture nucleic acids has at least three mismatched bases in the nucleic acid region with the first set of nucleic acids. In some embodiments, the first set of capture nucleic acids has complete complementarity with the first set of nucleic acids. In some embodiments, the first set of capture nucleic acids or the second set of capture nucleic acids comprises DNA. In some embodiments, the first set of capture nucleic acids or the second set of capture nucleic acids comprises RNA. In some embodiments, the first set of capture nucleic acids or the second set of capture nucleic acids comprises DNA and RNA. In some embodiments, the nucleic acid of the first set of capture nucleic acids comprises DNA and RNA. In some embodiments, the first set of capture nucleic acids includes a first nucleic acid comprising DNA and a second nucleic acid comprising RNA.
[0009] In some embodiments, the sequencing comprises performing a next-generation sequencing reaction. In some embodiments, the first sequencing depth is at least 10 reads. In some embodiments, the first sequencing depth is at least 100 reads. In some embodiments, the first sequencing depth is at least 1000 reads. In some embodiments, the first sequencing depth is no more than 10 reads. In some embodiments, the first sequencing depth is no more than 100 reads. In some embodiments, the first sequencing depth is no more than 1000 reads. In some embodiments, the second sequencing depth is at least 100 reads. In some embodiments, the second sequencing depth is at least 1000 reads. In some embodiments, the second sequencing depth is no more than 100 reads. In some embodiments, the second sequencing depth is no more than 1000 reads.
[0010] In some embodiments, the first set of nucleic acids comprises sequences associated with cancer or a cell proliferative disorder. In some embodiments, the cancer or cell proliferative disorder is colon cancer or a cell proliferative disorder. In some embodiments, the cancer or cell proliferative disorder is selected from colorectal cancer, prostate cancer, lung cancer, breast cancer, pancreatic cancer, ovarian cancer, uterine cancer, liver cancer, esophageal cancer, gastric cancer, and thyroid cancer or a cell proliferative disorder. In some embodiments, (b) and (c) are performed simultaneously or substantially simultaneously. In some embodiments, (b) and (c) are performed sequentially. In some embodiments, the method further comprises analyzing the sequencing reads to determine the presence of genetic parameters. In some embodiments, the genetic parameters are single nucleotide variants, copy number variants, deletions, insertions, or transversions. In some embodiments, the genetic parameters are associated with cancer or a cell proliferative disorder. In some embodiments, the method further comprises analyzing the sequencing reads to determine whether the subject has cancer or a cell proliferative disorder.
[0011] In another aspect, the present disclosure provides a method comprising: (a) providing a sample derived from a subject, wherein the sample comprises a plurality of nucleic acids; (b) differentially enriching at least a subset of the plurality of nucleic acids by contacting the plurality of nucleic acids with a plurality of oligonucleotides, wherein at least a subset of the plurality of oligonucleotides anneals to the subset of the plurality of nucleic acids, wherein the subset of the plurality of oligonucleotides has a different percentage of complementarity to the nucleic acids of the plurality of nucleic acids, and wherein a higher percentage of complementarity to the nucleic acids provides an increased enrichment ratio compared to a lower percentage of complementarity to the nucleic acids; and (c) sequencing the enriched subset of the plurality of nucleic acids to generate sequencing reads.
[0012] In some embodiments, the plurality of nucleic acids are derived from a cell-free sample. In some embodiments, the plurality of nucleic acids include cfDNA or cfRNA. In some embodiments, the plurality of nucleic acids include ctDNA. In some embodiments, the plurality of oligonucleotides include more oligonucleotides annealed to a first nucleic acid of the plurality of nucleic acids than oligonucleotides annealed to a second nucleic acid of the plurality of nucleic acids. In some embodiments, the plurality of oligonucleotides include a higher concentration of oligonucleotides annealed to a first nucleic acid of the plurality of nucleic acids than oligonucleotides annealed to a second nucleic acid of the plurality of nucleic acids. In some embodiments, the plurality of oligonucleotides include a tiling density of 1x. In some embodiments, the plurality of oligonucleotides include a tiling density of 2x. In some embodiments, the plurality of oligonucleotides include a tiling density of 0.5x. In some embodiments, a subset of the plurality of oligonucleotides configured to anneal to a first region of the nucleic acid of the plurality of nucleic acids has a different tiling density from a subset of the plurality of oligonucleotides configured to anneal to a second region of the nucleic acid of the plurality of nucleic acids. In some embodiments, a subset of the plurality of oligonucleotides configured to anneal to a first region of the nucleic acid of the plurality of nucleic acids has the same tiling density as a subset of the plurality of oligonucleotides configured to anneal to a second region of the nucleic acid of the plurality of nucleic acids. In some embodiments, the tiling density is generated by overlapping sequences in the oligonucleotides of the plurality of oligonucleotides. In some embodiments, the plurality of oligonucleotides include oligonucleotides of different lengths. In some embodiments, a subset of the plurality of oligonucleotides has at least one mismatched base with a nucleic acid region of the plurality of nucleic acids. In some embodiments, a subset of the plurality of oligonucleotides has at least two mismatched bases with a nucleic acid region of the plurality of nucleic acids. In some embodiments, a subset of the plurality of oligonucleotides has at least three mismatched bases with a nucleic acid region of the plurality of nucleic acids. In some embodiments, a subset of the plurality of oligonucleotides has perfect complementarity with the nucleic acid of the plurality of nucleic acids. In some embodiments, the plurality of oligonucleotides contain DNA. In some embodiments, the plurality of oligonucleotides contain RNA. In some embodiments, the plurality of oligonucleotides contain DNA and RNA. In some embodiments, the oligonucleotides of the plurality of oligonucleotides contain DNA and RNA. In some embodiments, a first oligonucleotide of the plurality of oligonucleotides contains DNA, and a second oligonucleotide of the plurality of oligonucleotides contains RNA. In some embodiments, the sequencing includes performing a next-generation sequencing reaction. In some embodiments, the sequencing generates at least 10 reads for a first region of the nucleic acid of the plurality of nucleic acids. In some embodiments, the sequencing generates at least 100 reads for a first region of the nucleic acid of the plurality of nucleic acids. In some embodiments, the sequencing generates at least 1000 reads for a first region of the nucleic acid of the plurality of nucleic acids.In some embodiments, the sequencing generates no more than 10 reads for a first region of the nucleic acids of the plurality of nucleic acids. In some embodiments, the sequencing generates no more than 100 reads for a first region of the nucleic acids of the plurality of nucleic acids. In some embodiments, the sequencing generates no more than 1000 reads for a first region of the nucleic acids of the plurality of nucleic acids. In some embodiments, the sequencing generates at least 100 reads for a second region of the nucleic acids of the plurality of nucleic acids. In some embodiments, the sequencing generates at least 1000 reads for a second region of the nucleic acids of the plurality of nucleic acids. In some embodiments, the sequencing generates no more than 100 reads for a second region of the nucleic acids of the plurality of nucleic acids. In some embodiments, the sequencing generates no more than 1000 reads for a second region of the nucleic acids of the plurality of nucleic acids.
[0013] In some embodiments, the subset of the plurality of nucleic acids comprises sequences associated with cancer or a cell proliferative disorder. In some embodiments, the cancer or cell proliferative disorder is colon cancer or a cell proliferative disorder. In some embodiments, the cancer or cell proliferative disorder is selected from colorectal cancer, prostate cancer, lung cancer, breast cancer, pancreatic cancer, ovarian cancer, uterine cancer, liver cancer, esophageal cancer, gastric cancer, and thyroid cancer or a cell proliferative disorder. In some embodiments, the method further comprises analyzing the sequencing reads to determine the presence of genetic parameters. In some embodiments, the genetic parameters are single nucleotide variants, copy number variants, deletions, insertions, or transversions. In some embodiments, the genetic parameters are associated with cancer or a cell proliferative disorder. In some embodiments, the method further comprises analyzing the sequencing reads to determine whether the subject has cancer or a cell proliferative disorder.
[0014] Another aspect of the present disclosure provides a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.
[0015] Another aspect of the present disclosure provides a system comprising one or more computer processors and a computer memory coupled thereto. The computer memory comprises machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.
[0016] In light of the following detailed description, additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art. Only illustrative embodiments of the present disclosure are shown and described in the following detailed description. As will be understood, the present disclosure is capable of other and different embodiments, and several details can be modified in various obvious aspects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not restrictive.
[0017] Incorporated by reference
[0018] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that the publications and patents or patent applications incorporated by reference conflict with the disclosure contained herein, the specification is intended to supersede and / or take precedence over any such conflicting material. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description and the accompanying drawings (also referred to herein as "figures" and "FIGs.") that illustrate exemplary embodiments that utilize the principles of the invention, in which:
[0020] Figure 1 A schematic diagram of a computer system programmed or otherwise configured to implement the methods provided herein is shown.
[0021] Figure 2 Shows the median prostate adenocarcinoma (PRAD) panel coverage for a cfDNA library.
[0022] Figure 3 Shows the percentage of bases covered in a cfDNA library.
[0023] Figure 4 Shows the variation in median PRAD panel coverage levels between different enrichments.
[0024] Figure 5 Shows the sequencing depth of regions with reduced coverage. DETAILED DESCRIPTION
[0025] Although various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Various variations, changes, and alternatives will be apparent to those skilled in the art without departing from the present invention. It should be understood that various alternative embodiments to those described herein for the present invention may be employed.
[0026] The present disclosure generally relates to the capture or enrichment of nucleic acid molecules. Nucleic acid molecules can be captured or enriched and sequenced to determine the nucleic acid sequence. Based on the sequence, certain conditions can be analyzed. For example, sequencing can be used to screen for or monitor cancer or other diseases. Such screening and monitoring can help improve outcomes because early detection leads to better outcomes as the disease can be identified before it progresses and worsens.
[0027] Next-generation sequencing (NGS) technologies can enable researchers or clinicians to survey the entire genomic landscape of an individual. This data can empower patients to understand their health status or disease risk. However, most of the DNA or RNA found in a sample from an object (e.g., a patient), such as tissue, blood, plasma, urine, etc., may be uninformative and thus does not need to be sequenced. Target capture (or target enrichment) can be used to select regions of interest from the entire nucleic acid pool to generate an NGS library enriched in informative sequences and thereby eliminate unwanted nucleic acid fragments. To capture or enrich the selected targets, nucleic acid molecules having sequences complementary to the regions of interest can be synthesized and then mixed with the sample. These nucleic acid molecules having sequences complementary to the regions of interest can hybridize with the nucleic acids from the original sample and can then be captured or amplified while non-target nucleic acids can be removed. In one embodiment, the capture method involves hybridizing biotinylated oligonucleotides with the nucleic acids from the regions of interest in the original sample and capturing these regions using streptavidin-coated beads.
[0028] Target capture can be designed to achieve uniform sequencing coverage among each region of interest in a sample. However, the amount of sequencing reads required for a locus depends on a number of factors specific to that region of interest. For example, when looking for signals from circulating tumor DNA (ctDNA) in plasma, deep sequencing (e.g., several hundred to a thousand reads per genomic region, or depth of coverage) may be necessary because the number of molecules originating from the tumor is low relative to DNA from other sources. However, in the exact same sample, low coverage (e.g., 10 reads) may be sufficient to genotype an individual's genes related to cancer risk. This represents one of many use cases that demonstrate the need for customizable sequencing depth specific to each individual region of interest. Having a method to achieve variable coverage in a purposeful manner in a single target capture reaction has the potential to increase data utility while reducing the overall cost of sequencing. For example, sequencing only certain regions at a specific coverage, rather than sequencing the entire library or genome at the same coverage, can allow for sequencing fewer bases, thereby reducing the overall cost of sequencing.
[0029] Particularly of interest may be capturing or enriching genes related to the detection of hyperplastic disorders and disease progression in lung, colon, liver, ovary, pancreas, prostate, rectum, and breast cells. For example, circulating tumor DNA can be a viable "liquid biopsy" for detecting and profiling tumors in a non-invasive manner. Identification of tumor-specific mutations in circulating tumor DNA can be used to diagnose colon, breast, and prostate cancers. However, the sensitivity of these techniques can be limited due to the high background of normal (e.g., non-tumor-derived) DNA present in the circulation.
[0030] I. Definitions
[0031] Unless the context clearly indicates otherwise, as used in the specification and claims, the singular forms "a / an" and "the" include plural referents. For example, the term "nucleic acid" includes a plurality of nucleic acids, including mixtures thereof.
[0032] As used herein, the term "subject" generally refers to an entity or agent having testable or detectable genetic information. A subject can be a person, individual, or object. A subject can be a vertebrate, such as a mammal. Non-limiting examples of mammals include humans, apes, farm animals, sport animals, rodents, and pets. A subject can be a human having cancer or suspected of having cancer. A subject can exhibit symptoms indicative of the subject's health or physiological state or condition, such as the subject's cancer or other disease, disorder, or condition. Alternatively, a subject can be asymptomatic with respect to such health or physiological state or condition.
[0033] As used herein, the term "sample" generally refers to a biological sample obtained from or derived from one or more subjects. A biological sample can be a cell-free biological sample or a substantially cell-free biological sample, or can be processed or fractionated to yield a cell-free biological sample. For example, cell-free biological samples can include cell-free ribonucleic acid (cfRNA), cell-free deoxyribonucleic acid (cfDNA), cell-free fetal DNA (cffDNA), plasma, serum, urine, saliva, amniotic fluid, and derivatives thereof. Ethylenediaminetetraacetic acid (EDTA) collection tubes, cell-free RNA collection tubes (e.g., ) or cell-free DNA collection tubes (e.g., ) can be used to obtain or derive cell-free biological samples from a subject. Cell-free biological samples can be derived from whole blood samples by fractionation (e.g., centrifugation into cellular and cell-free fractions). A biological sample or its derivative can contain cells. For example, a biological sample can be a blood sample or its derivative (e.g., blood collected by a collection tube or a blood droplet).
[0034] As used herein, the term "nucleic acid" generally refers to a polymeric form of nucleotides of any length, whether deoxyribonucleotides (dNTPs) or ribonucleotides (rNTPs), or analogs thereof. Nucleic acids can have any three-dimensional structure and can perform any known or unknown function. Non-limiting examples of nucleic acids include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), coding or non-coding regions of genes or gene fragments, loci defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant nucleic acids, branched nucleic acids, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Nucleic acids can contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure can be imparted before or after nucleic acid assembly. The nucleotide sequence of a nucleic acid can be interrupted by non-nucleotide components. Nucleic acids can be further modified after polymerization, such as by conjugation or binding to a reporter moiety.
[0035] As used herein, the term "target nucleic acid" generally refers to a nucleic acid molecule in an initial population of nucleic acid molecules, the presence, quantity, and / or sequence of the nucleotide sequence of which, or a change in one or more of them, needs to be determined. A target nucleic acid can be any type of nucleic acid, including DNA, RNA, and analogs thereof. As used herein, "target ribonucleic acid (RNA)" generally refers to a target nucleic acid that is RNA. As used herein, "target deoxyribonucleic acid (DNA)" generally refers to a target nucleic acid that is DNA.
[0036] As used herein, the terms “amplifying” and “amplification” generally refer to increasing the size or number of nucleic acid molecules. The nucleic acid molecules can be single-stranded or double-stranded. Amplification can include generating one or more copies or “amplification products” of the nucleic acid molecule. Amplification can be carried out, for example, by extension (such as primer extension) or ligation. Amplification can include performing a primer extension reaction to generate a strand complementary to a single-stranded nucleic acid molecule and, in some cases, generating one or more copies of that strand and / or the single-stranded nucleic acid molecule. The term “DNA amplification” generally refers to generating one or more copies of a DNA molecule or “amplified DNA product”. The term “reverse transcription amplification” generally refers to generating deoxyribonucleic acid (DNA) from a ribonucleic acid (RNA) template through the action of reverse transcriptase.
[0037] As used herein, the term “cell-free nucleic acid (cfNA)” generally refers to nucleic acids in a biological sample that are not contained within cells (such as cell-free RNA (“cfRNA”) or cell-free DNA (“cfDNA”)). cfDNA can circulate freely in body fluids such as in the bloodstream.
[0038] As used herein, the term “cell-free sample” generally refers to a biological sample that is substantially devoid of intact cells. This can be derived from a biological sample that is itself substantially cell-free or can be derived from a sample from which cells have been removed. Examples of cell-free samples include those derived from blood, such as serum or plasma; urine; or samples from other sources, such as semen, sputum, feces, catheter effluent, lymph, or recovered lavage fluid.
[0039] As used herein, the term “circulating tumor DNA (ctDNA)” generally refers to cfDNA that is derived from a tumor.
[0040] As used herein, the term “genomic region” generally refers to an identified region of nucleic acid that is identified based on its location in a chromosome. In some instances, a genomic region is referred to by a gene name and encompasses coding and non-coding regions associated with a physical region of nucleic acid. As used herein, a gene includes coding regions (exons), non-coding regions (introns), transcriptional control regions or other regulatory regions, and promoters. In another instance, a genomic region can incorporate an intron or exon or an intron / exon boundary within a named gene.
[0041] As used herein, the term "cell proliferative disorder" generally refers to a disorder or disease that includes a cellular disorder or abnormal proliferation, such as cancer. In some non-limiting examples, the disorder is selected from colorectal cell proliferation, prostate cell proliferation, lung cell proliferation, breast cell proliferation, pancreatic cell proliferation, ovarian cell proliferation, uterine cell proliferation, liver cell proliferation, esophageal cell proliferation, gastric cell proliferation, or thyroid cell proliferation. In some embodiments, the cell proliferative disorder is selected from colon adenocarcinoma, hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, ovarian cystadenocarcinoma, pancreatic adenocarcinoma, prostate adenocarcinoma, and rectal adenocarcinoma.
[0042] As used herein, the terms "normal" or "healthy" generally refer to cells, tissues, plasma, blood, biological samples, or subjects that do not have a cell proliferative disorder.
[0043] As used herein, the term "epigenetic parameter" generally refers to cytosine methylation. Further epigenetic parameters include, for example, acetylation of histones, although they may not be directly analyzable using the described methods, but conversely, are related to DNA methylation. Epigenetic parameters can also include other modifications of nucleotides, such as methylation, oxidation, deamination, fluorination, hydroxymethylation, formylation, glycosylation, amination of cytosine.
[0044] As used herein, the term "genetic parameter" generally refers to mutations and polymorphisms of genes and the sequences required for their further regulation. Examples of mutations include insertions, deletions, point mutations, inversions, and polymorphisms, such as SNPs (single nucleotide polymorphisms).
[0045] The terms cancer "type" and "subtype" are generally used relatively herein, whereby a "type" of cancer, such as breast cancer, can be a "subtype" based on, for example, stage, morphology, histology, gene expression, receptor profile, mutation profile, invasiveness, prognosis, malignant characteristics, etc. Similarly, "type" and "subtype" can be applied at a finer level, for example, differentiating a histological "type" into "subtypes", for example, defined according to the mutation profile or gene expression. Cancer "stage" is also used to refer to the classification of cancer types based on histological and pathological features related to disease progression.
[0046] II. Samples
[0047] The sample can be a biological sample. The sample can be derived from a biological sample. The biological sample can be, for example, blood, plasma, serum, urine, saliva, mucosal secretions, sputum, feces, or tears. The biological sample can be a fluid sample. The fluid sample can be a blood or plasma sample. The biological sample can be a tissue sample, such as a biopsy, core biopsy, needle aspiration, or fine needle aspiration. The biological sample can be a fluid sample, such as a blood sample, urine sample, or saliva sample. The biological sample can be a skin sample. The biological sample can be a buccal swab. The biological sample can be a plasma or serum sample. The biological sample can include one or more cells. The biological sample can be, for example, blood, plasma, serum, urine, saliva, mucosal secretions, sputum, feces, or tears. The biological sample can contain cell-free nucleic acids (e.g., cell-free RNA, cell-free DNA, etc.). The sample can contain circulating tumor DNA (ctDNA). The sample can be a cell-free biological sample. The nucleic acid target can be a nucleic acid suspected of containing one or more mutations.
[0048] The cell-free biological sample can be obtained or derived from a human subject. The cell-free biological sample can be stored under different storage conditions before processing, such as different temperatures (e.g., room temperature, refrigerated or frozen conditions, 25 °C, 4 °C, -18 °C, -20 °C, or -80 °C) or different suspensions (e.g., EDTA collection tube, cell-free RNA collection tube, or cell-free DNA collection tube).
[0049] The cell-free biological sample can be obtained from a subject with cancer, from a subject suspected of having cancer, or from a subject who has never had or is not suspected of having cancer. The cancer can be colon cancer.
[0050] The cell-free biological sample can be collected before and / or after treatment of a cancer subject. During treatment or a treatment regimen, a cell-free biological sample can be obtained from the subject. Multiple cell-free biological samples can be obtained from the subject to monitor the treatment effect over time. The cell-free biological sample can be taken from a subject known or suspected of having cancer who cannot be definitively diagnosed as positive or negative through a clinical trial. The sample can be taken from a subject suspected of having cancer. The cell-free biological sample can be taken from a subject presenting with unexplained symptoms such as fatigue, nausea, weight loss, pain, weakness, or bleeding. The cell-free biological sample can be taken from a subject with explained symptoms. The cell-free biological sample can be taken from a subject at risk of developing cancer due to factors such as family history, age, hypertension or prehypertension, diabetes or prediabetes, overweight or obesity, environmental exposure, lifestyle risk factors (e.g., smoking or alcohol consumption), or the presence of other risk factors.
[0051] Cell-free biological samples can contain one or more analytes that can be assayed, such as cell-free ribonucleic acid (cfRNA) molecules suitable for analysis to generate transcriptomic data, cell-free deoxyribonucleic acid (cfDNA) molecules suitable for analysis to generate genomic and / or epigenetic data, or mixtures or combinations thereof. One or more such analytes (e.g., cfRNA molecules and / or cfDNA molecules) can be isolated or extracted from one or more cell-free biological samples of a subject for downstream analysis using one or more suitable assays. Cell-free biological samples can contain methylated nucleic acids. Methylated nucleic acids can contain methylated cytosine. Methylated nucleic acids can be assayed to identify epigenetic parameters or correlations with disease states or conditions.
[0052] A nucleic acid sample or subset of nucleic acid molecules can contain one or more genomic regions. One or more genomic regions can contain genetic parameters, such as polymorphisms or portions thereof. Genetic parameters can be genetic aberrations. For example, genetic parameters can be mutations, single nucleotide polymorphisms, single nucleotide variants, insertions, deletions, fusions, copy number variations, copy number losses, or other changes in the sequence or copy number of a nucleic acid or nucleic acids. Genomic regions can contain methylated nucleotides or epigenetic parameters. Capturing nucleic acids containing genomic regions can permit determination of the nucleic acids in a sample or subject.
[0053] After obtaining a cell-free biological sample from a subject, the cell-free biological sample can be processed to generate a data set indicative of cancer in the subject. For example, the nucleic acid molecules of the cell-free biological sample can be evaluated for presence, absence, or quantification on a panel of loci of a cancer-related genome (e.g., a quantitative measure of an RNA transcript or DNA at a cancer-related genomic locus). Processing of the cell-free biological sample obtained from a subject can include: (i) placing the cell-free biological sample under conditions sufficient to isolate, enrich, or extract a plurality of nucleic acid molecules, and (ii) assaying the plurality of nucleic acid molecules to generate a data set.
[0054] In some embodiments, a plurality of nucleic acid molecules are extracted from a cell-free biological sample and sequenced to generate a plurality of sequencing reads. The nucleic acid molecules can include ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). The nucleic acid molecules (e.g., RNA or DNA) can be extracted from the cell-free biological sample by a variety of methods, such as from the kit protocol of MP , from the cell-free DNA Mini Kit of , or from Norgen Cell-free biological DNA isolation kit protocol. The extraction method can extract all RNA or DNA molecules from the sample. Alternatively, the extraction method can selectively extract a portion of the RNA or DNA molecules from the sample. The RNA molecules extracted from the sample can be converted to DNA molecules by reverse transcription (RT).
[0055] Sequencing can be performed by any suitable sequencing method, such as massively parallel sequencing (MPS), paired-end sequencing, high-throughput sequencing, next-generation sequencing (NGS), shotgun sequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, pyrosequencing, sequencing by synthesis (SBS), ligation-based sequencing, hybridization sequencing, and
[0056] Sequencing can include nucleic acid amplification (e.g., RNA or DNA molecules). In some embodiments, the nucleic acid amplification is polymerase chain reaction (PCR). Appropriate numbers of rounds of PCR (e.g., PCR, qPCR, reverse transcriptase PCR, digital PCR, etc.) can be performed to sufficiently amplify the initial amount of nucleic acid (e.g., RNA or DNA) to the desired input amount for subsequent sequencing. In some cases, PCR can be used for the overall amplification of target nucleic acids. This can include using an aptamer sequence that can first be ligated to a different molecule and then amplified by PCR using a universal primer. PCR can be performed using any of a number of commercial kits, such as those provided by Life etc. In other cases, only certain target nucleic acids within a nucleic acid population can be amplified. Specific primers (possibly in combination with adapter ligation) can be used to selectively amplify certain targets for downstream sequencing. PCR can include the targeted amplification of one or more genomic loci, such as cancer-related genomic loci. Sequencing can include the use of simultaneous reverse transcription (RT) and polymerase chain reaction (PCR), such as the OneStep RT-PCR kit protocol provided by Thermo Fisher or provided.
[0057] RNA or DNA molecules isolated or extracted from cell-free biological samples can be labeled, for example, with an identifiable tag to allow multiplexing of multiple samples. Any number of RNA or DNA samples can be multiplexed. For example, the multiplexing reaction can comprise RNA or DNA from at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more than 100 initial cell-free biological samples. For example, multiple cell-free biological samples can be tagged with sample barcodes so that each DNA molecule can be traced back to the sample (and subject) from which the DNA molecule originated. Such tags can be attached to the RNA or DNA molecule by ligation or primer PCR amplification.
[0058] After sequencing the nucleic acid molecules, the sequence reads can be subjected to appropriate bioinformatics processing to generate data indicative of the presence, absence, or relative assessment of cancer. For example, the sequence reads can be aligned to one or more reference genomes (e.g., the genomes of one or more species, such as the human genome). The aligned sequence reads can be quantified at one or more genomic loci to generate a dataset indicative of cancer. For example, quantifying sequences corresponding to multiple genomic loci with or without genetic or epigenetic parameters associated with cancer can generate a dataset indicative of cancer.
[0059] Cell-free biological samples can be processed without any nucleic acid extraction. For example, cancer can be identified or monitored by using probes or primers configured to selectively enrich nucleic acid (e.g., RNA or DNA) molecules corresponding to multiple cancer-associated genomic loci. The probes can have sequence complementarity to nucleic acid sequences from one or more of multiple cancer-associated genomic loci or genomic regions. The multiple cancer-associated genomic loci or genomic regions can comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, at least about 90, at least about 95, at least about 100 or more different cancer-associated genomic loci or genomic regions.
[0060] A probe can be a nucleic acid molecule (e.g., RNA or DNA) having sequence complementarity to the nucleic acid sequence (e.g., RNA or DNA) of one or more genomic or epigenetic loci (such as cancer-associated genomic loci). These nucleic acid molecules can be primers or enrichment sequences. Analysis of a cell-free biological sample using a probe that selectively targets one or more genomic loci (such as cancer-associated genomic loci) can include using array hybridization (e.g., microarray-based), polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., RNA sequencing or DNA sequencing). In some embodiments, DNA or RNA can be analyzed by one or more of the following: isothermal DNA / RNA amplification methods (such as loop-mediated isothermal amplification (LAMP), helicase-dependent amplification (HDA), rolling circle amplification (RCA), recombinase polymerase amplification (RPA)), immunoassays, electrochemical assays, surface-enhanced Raman spectroscopy (SERS), quantum dot (QD)-based assays, molecular inversion probes, droplet digital PCR (ddPCR), CRISPR / Cas-based detection (such as CRISPR typing PCR (ctPCR), specific high-sensitivity enzymatic reporter unlocking (SHERLOCK), DNA endonuclease-targeted CRISPR trans-reporter (DETECTR), and CRISPR-mediated analog multi-event recording apparatus (CAMERA)), and laser transmission spectroscopy (LTS).
[0061] The assay readout can be quantified at one or more genomic or epigenetic loci (such as, cancer-associated genomic loci) to generate data indicative of cancer. For example, quantification of array hybridization or polymerase chain reaction (PCR) corresponding to multiple genomic loci (such as, cancer-associated genomic loci) can generate data indicative of cancer. The assay readout can include quantitative PCR (qPCR) values, digital PCR (dPCR) values, digital droplet PCR (ddPCR) values, fluorescence values, etc., or their normalized values. The assay can be a home user test configured to be performed in a home environment.
[0062] III. Probe or Primer Panel
[0063] The present disclosure provides methods and systems for analyzing a biological sample to obtain sequencing data of nucleic acids of a subject. The sequencing data can include nucleic acids that have been captured or enriched by a set of one or more probes or primers.
[0064] The groups described herein generally refer to a collection of targeted regions of genomic DNA identified in a biological sample. In certain embodiments, the biological sample is a cell-free nucleic acid sample. The formation of a signature group allows for the rapid and specific analysis of regions associated with a disorder, condition, or specific genotype. The groups as described and used in the methods herein can be used to improve the diagnosis, prognosis, treatment selection, and monitoring (e.g., treatment monitoring) of a disorder or condition such as cancer.
[0065] The signature groups and methods provide a significant improvement over current methods as markers or signature groups are needed for the detection of early cell proliferative disorders from body fluid samples such as whole blood, plasma, or serum.
[0066] The present disclosure also provides sequencing methods to determine genetic or epigenetic parameters of one or more genes. The genetic parameters can be genetic aberrations. For example, the genetic parameters can be mutations, single nucleotide polymorphisms, single nucleotide variants, insertions, deletions, fusions, copy number variations, copy number losses, or other changes in the sequence or copy number of a nucleic acid or nucleic acids. The method can include obtaining a sample from a subject and sequencing the nucleic acid. The nucleic acid sequencing can include sequencing techniques and workflows as described elsewhere in the present disclosure.
[0067] As described herein, a tumor or cell proliferative disorder can be selected from colorectal, prostate, lung, breast, pancreas, ovary, uterus, liver, esophagus, stomach, or thyroid cell proliferation. In some embodiments, the cell proliferative disorder is selected from colon adenocarcinoma, hepatocellular carcinoma, lung adenocarcinoma, lung squamous cell carcinoma, ovarian serous cystadenocarcinoma, pancreatic adenocarcinoma, prostate adenocarcinoma, and rectal adenocarcinoma.
[0068] In some embodiments, the cell proliferative disorder is a colon cell proliferative disorder. In some embodiments, the colon cell hyperplastic disorders are selected from adenomas (adenomatous polyps), polyposis, Lynch syndrome, sessile serrated adenomas (SSA), advanced adenomas, colorectal dysplasia, colorectal adenomas, colorectal cancer, colon cancer, rectal cancer, colorectal carcinoma, colorectal adenocarcinoma, carcinoid tumors, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors (GIST), lymphomas, and sarcomas.
[0069] The hybridization methods provided herein can be used for various forms of nucleic acid hybridization, such as in-solution hybridization and hybridization on solid supports (e.g., Northern hybridization, Southern hybridization, and in situ hybridization on membranes, microarrays, and cell / tissue slides). Specifically, the method is applicable to in-solution hybridization capture for target enrichment of certain types of genomic DNA sequences (e.g., exons) for use in targeted next-generation sequencing. For hybridization capture methods, cell-free nucleic acid samples undergo library preparation. As used herein, "library preparation" includes end repair, A-tailing, adapter ligation, or any other preparation of cell-free DNA to allow subsequent DNA sequencing. In certain instances, the prepared cell-free nucleic acid library sequences contain adapters, sequence tags, index barcodes, UMIs, or combinations thereof, ligated to the cell-free nucleic acid sample molecules. Various commercially available kits can be utilized to facilitate library preparation for NGS methods. NGS library construction can include using a series of coordinated enzymatic reactions to prepare nucleic acid targets to generate a collection of random DNA fragments of a specific size for high-throughput sequencing. Advances and developments in various library preparation techniques have expanded the applications of NGS in fields such as transcriptomics and epigenetics.
[0070] Improvements in sequencing technology have led to changes and improvements in library preparation. By companies such as
[0071] Bioo Kapa New England
[0072] Life Pacific and developed NGS library preparation kits can be used to provide consistency and reproducibility for various molecular biology reactions, ensuring compatibility with the latest NGS instrument technologies.
[0073] In various instances of targeted capture gene panels, various library preparation kits can be selected from Nextera Flex DNA Prep Ion (Thermo Fisher ), (Thermo Fisher ), Agilent ClearSeq Capture Bioo xGen
[0074] and
[0075] In some embodiments, a hybridization capture method is performed on the prepared library sequences using specific probes. As used herein, the term "specific probe" generally refers to a probe specific for a region in some embodiments. In some embodiments, the specific probes are designed based on using the human genome as a reference sequence and using specified genomic regions of interest. Thus, when the specific probes of some embodiments are used for hybridization capture, sequences in the sample genome complementary to the target sequence can be effectively captured.
[0076] According to the principle of complementary base pairing, single-stranded capture probes can combine complementarily with single-stranded target sequences, thus successfully capturing the target regions. In some embodiments, the designed probes can be designed as solid capture chips (where the probes are immobilized on a solid support) or as liquid capture chips (where the probes are free in a liquid). However, due to various limitations such as probe length, probe density, and high cost, solid capture chips are rarely used, while liquid capture chips are used more frequently.
[0077] In some embodiments, compared with normal sequences (where the average content of A, T, C, and G bases is 25% each), sequences rich in GC in nucleic acids (where the GC base content is higher than 60%) can lead to an increase in capture efficiency due to the molecular structures of C and G bases.
[0078] The number of probes added for each region of interest can be a specific amount or concentration. The number of probes can be increased or decreased relative to the final sequencing depth for a given region. For example, changing the number of probes targeting a given region can result in a change in the resulting sequencing depth for each region. Compared with a second region, a first region can have more probes annealing to it. The more probes there are, the more nucleic acid sequences can be allowed to be captured, and the sequence depth of that region can be increased. Conversely, a region with a lower number of probes can allow less nucleic acid to be captured and result in a lower sequencing depth. In this way, the sequencing depth can be adjusted or regulated at least based on the number of probes for a given region.
[0079] The amount of time allowed for hybridization can be adjusted or otherwise varied. The hybridization step of the target capture reaction can range from a few minutes to several hours. Altering the amount of time that the complementary sequences are able to hybridize to each region of interest can result in a change in the coverage or depth of a given region. Probes with shorter hybridization times can result in lower recovery rates of their targeted regions and lower sequencing coverage or depth compared to probes with longer allowed hybridization times. The hybridization time can be adjusted by adding probes to the hybridization reaction at multiple time points to generate a specific sequencing depth. For example, in a 16-hour hybridization reaction, some probes can be allowed to hybridize for 16 hours, while other probes can be added to the reaction after 15 hours, resulting in an incubation time of 1 hour for the second set of molecules. Using this strategy, the target coverage between regions can be adjusted and customized in a single reaction, and different sequencing depths can be generated for different regions.
[0080] In certain embodiments, the temperature at which hybridization is allowed to occur can be adjusted or otherwise varied. The hybridization temperature of the target capture reaction can range from a few minutes to several hours. Altering the temperature at which the complementary sequences are able to hybridize to each region of interest can result in a change in the coverage or depth of a given region. The approximate probe hybridization temperature can be calculated by a computer. Using this method, the target coverage between regions can be adjusted and customized in a single reaction, and different sequencing depths can be generated for different regions.
[0081] The molecular density targeting regions of interest can be altered on a region-by-region basis. Assuming 1x coverage is achieved in an exemplary target capture reaction (each region of interest has exactly one synthetic molecule designed to be complementary to the region of interest), increasing the probe tiling to have more than one capture probe per region can result in higher coverage and higher sequencing depth. Alternatively, decreasing the tiling density such that only a fraction of the regions of interest are covered by probes (e.g., 0.5x) can result in lower sequencing coverage. In this way, each region of interest can have a customized tiling density to generate a specific sequencing coverage for each region, where the first region can have a different coverage compared to the second region.
[0082] The probe can have a specific length. For example, the length of the probe can be more than 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more nucleotides. The lengths of the probes in the reaction can be different from each other. For example, the first probe can be of a first length and the second probe can be of a length different from that of the first probe. The number of directly complementary bases between two molecules can affect the strength of their binding to each other, which in turn can affect the optimal temperature for molecular binding (annealing) or separation (melting). Changing the length of the probe (rather than having all probes of a set length) in a target capture reaction can result in different optimal hybridization conditions between regions. Due to the different annealing and melting temperatures between probes, targeting regions with probes of different lengths can result in differences in subsequent sequence coverage.
[0083] The probe can have a certain amount of complementarity to the target region. The efficiency of hybridization of two molecules can be affected by the degree of their sequence match. The probe can have perfect complementarity to the target region, where each base of the probe makes Watson-Crick pairing with the bases on the target region. The probe can have imperfect complementarity. For example, the probe can have mismatches with the bases of the target region such that not all bases pair with the target region. Such mismatches can result in lower hybridization efficiency. Compared with a perfectly complementary probe, a mismatched probe can capture fewer nucleic acid molecules. Introducing mismatched bases into the synthetic probe can reduce the hybridization efficiency in proportion to how many mismatches are present in each region. Adding mismatches in the selected regions of interest can result in lower target coverage or depth. By using probes with different complementarities, the coverage or depth can be partially adjusted such that regions requiring lower depth can use more mismatched probes.
[0084] The probe can also contain RNA or DNA or both. The target capture probe can be synthesized using both DNA and RNA. The target capture reaction can consist of a single class of molecules (DNA or RNA). Multiple probes can include probes containing RNA and probes containing DNA. The hybridization affinity of DNA and RNA probes and their optimal hybridization conditions (temperature, time, etc.) can be different. Using DNA probes in some regions of interest while using RNA probes in other regions can result in different coverages between the two sets, as there can be inherent differences in the behavior of the two molecules in a single reaction. A target capture panel consisting of both DNA and RNA probes can allow for differential coverage between regions in a single reaction. The probe can contain methylated or modified bases.
[0085] The probes can be used in a given reaction as a probe set or a probe group. The reactions can be carried out sequentially, simultaneously, or overlapping with a previous reaction. For example, a first set of probes can be added to a sample and annealed. After a period of time, a second set of probes can be added to the sample. The first set of probes can be removed before adding the second set of probes, or can be retained in the sample when adding the second set of probes.
[0086] The probes can allow enrichment such that a specific sequencing depth or range of sequencing depths is achieved for a given region or sub-region of the genome. The sequencing depth for a region can be at least 0.1x, 0.5x, 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 15x, 20x, 25x, 30x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 125x, 150x, 175x, 200x, 300x, 400x, 500x or more. The sequencing depth for a region can be no more than 0.1x, 0.5x, 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 15x, 20x, 25x, 30x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 125x, 150x, 175x, 200x, 300x, 400x, 500x or less.
[0087] Amplification of nucleic acids
[0088] A nucleic acid molecule or a fragment thereof can be amplified. Amplification can be used to enrich specific sequences of interest. For example, a set of primers can anneal to a target sequence and generate amplicons related to that sequence. Then, the targeted sequence can be present at a higher concentration and account for a larger fraction of the total molecules in the molecular pool. In this way, a set of nucleic acid sequences can be enriched. The amount of enrichment can be related to the sequencing coverage or depth during nucleic acid sequencing. Compared with unenriched molecules, the enriched molecules can have a higher depth or sequence coverage. An increase in molecular enrichment or amplification can be related to a higher sequencing depth or coverage.
[0089] In various instances, the source of the DNA is cell-free DNA from whole blood, plasma, serum, or genomic DNA extracted from cells or tissues. In some embodiments, the amplified fragment lengths are between about 100 and 200 base pairs (bp). In some embodiments, the DNA source is extracted from a cell source (e.g., tissue, biopsy, cell line), and the amplified fragment lengths are between about 100 and 350 bp. Amplification can be performed using a set of primer oligonucleotides according to the present disclosure and can use a thermostable polymerase. Amplification of several DNA segments can be performed simultaneously in the same reaction vessel. In some embodiments of the method, two or more fragments are amplified simultaneously. For example, amplification can be performed using polymerase chain reaction (PCR). In certain embodiments, the methods discussed herein can achieve differential recovery of nucleic acid fragments of different sizes. For example, by increasing the tiling density of regions that are more likely to have short (<100 nucleotides) fragments, these smaller fragments can be preferentially recovered over more difficult (e.g., 100 - 300 bp) fragments.
[0090] Primers are designed to target such sequences that are disease-related or corresponding. In some embodiments, the PCR primers are designed to be specific to genes associated with cancer. In some embodiments, the primers are designed to be specific to genes associated with colon cancer.
[0091] Primers can be designed to amplify DNA fragments based on the expected (e.g., typical) size range of the circular DNA. Optimizing primer design to account for the target size can increase the sensitivity of the method according to this instance. In some embodiments, the primers are designed to amplify DNA fragments that are 75 to 350 bp in length. Primers can be designed to amplify regions that are about 50 to 200 bp, about 75 to 150 bp, or about 100 or 125 bp in length.
[0092] Suitable tools (such as Primer3, Primer3Plus, Primer-BLAST, etc.) can be used to design primers for the target region. The design can include complementarity to a specific region or gene and can be designed to have specific properties, such as melting temperature, GC content, dimerization energy, or hairpin formation energy.
[0093] The number of primers added for each region of interest can be a specific amount or concentration. The number of primers can be increased or decreased relative to the final sequencing depth for a given region. For example, changing the number of primers targeting a given region can result in a change in the resulting sequencing depth for each region. The first region can have more primers annealing to the first region compared to the second region. The more primers there are, the more nucleic acid sequences can be captured, resulting in an increase in the sequence depth of that region. Conversely, a region with a lower number of primers can allow for the amplification of fewer nucleic acids and result in a lower sequencing depth. In this way, the sequencing depth can be adjusted or regulated at least based on the number of primers for a given region.
[0094] The amount of time allowed for hybridization, annealing, extension, or other reactions can be adjusted or otherwise changed. The hybridization of an amplification reaction can vary from a few seconds to several hours. A change in the amount of time that complementary sequences are able to hybridize to each region of interest can result in a change in the coverage or depth of a given region. Primers with a shorter hybridization time can result in a lower recovery rate and lower sequencing coverage or depth of their targeted regions compared to primers with a longer hybridization time. The hybridization time can be regulated by adding primers to the hybridization reaction at multiple time points. The extension time can be modified to change the amount of time that the enzyme has to generate an extension or amplification product. A change in the extension time of nucleic acids in a region of interest can result in a change in the coverage or depth of a given region. For example, an extension product generated in a shorter extension time can result in an incomplete product that cannot be amplified by a second primer. Primers can be designed such that a first extension product is generated and can be amplified within the extension time, while a second extension product is not amplified within the extension time.
[0095] The amount of amplification cycles can be regulated to differentially enrich for sequences of interest. Primers annealing to the first region can undergo a certain number of cycles to generate a certain amount of amplicons, while primers annealing to the second region can undergo a different number of cycles. For example, in an amplification reaction of 30 cycles, some primers can be added at the beginning and allowed to amplify for all 30 cycles, while other primers can be added to the reaction after 15 cycles, resulting in 15 cycles of amplification for the second set of molecules. Using this strategy, the target coverage between regions can be regulated and customized in a single reaction, and different sequencing depths can be generated for different regions
[0096] Primers can have a specific length. For example, the length of a primer can be more than 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more nucleotides. The lengths of the primers in a reaction can be different from each other. For example, the first primer can be of a first length and the second primer can be of a length different from the first primer. The number of directly complementary bases between two molecules can affect the strength of binding of the molecules to each other, which in turn can affect the optimal temperature for molecule binding (annealing) or separation (melting). Changing the length of the primers (rather than having all primers of a set length) in an amplification reaction can result in different optimal hybridization conditions between regions. Due to the different annealing and melting temperatures between primers, targeting regions with different length primers can result in differences in subsequent sequence coverage.
[0097] Primers can be designed to include a specific melting temperature or annealing temperature. For example, primers can include a GC content. Based on the annealing or melting temperature, the amplification or extension efficiency of some primers can be higher or lower at different temperatures. The conditions of an amplification reaction can include a temperature above the annealing or melting temperature of a set of primers. At this temperature, that set of primers may have lower efficiency or may not produce an extension, while a set of primers with a higher melting temperature may be able to produce an extension or amplification product more effectively at this temperature. The resulting amplification can result in more amplicons corresponding to a first region than to a second region.
[0098] Primers can be used in a given reaction as a primer set or primer group. The reactions can be carried out sequentially, simultaneously, or overlapping with a previous reaction. For example, a first set of primers can be added to a sample and allowed to anneal. After a period of time, a second set of primers can be added to the sample. The first set of primers can be removed before adding the second set of primers, or it can be retained in the sample when adding the second set of primers.
[0099] Primers can also contain RNA or DNA. Primers can be synthesized using both DNA and RNA. A target capture reaction can consist of a single class of molecules (DNA or RNA). Multiple primers can include primers containing RNA and primers containing DNA. The hybridization affinity of DNA and RNA primers and their optimal hybridization conditions (such as temperature, time, etc.) can be different. Using DNA primers in some regions of interest while using RNA primers in other regions can result in different coverage between the two sets, as there can be inherent differences in the behavior of the two molecules in a single reaction. Multiple primers consisting of both DNA and RNA primers can allow for differential coverage between regions in a single reaction. Primers can contain methylated or modified bases.
[0100] Primers can allow for enrichment such that a specific sequencing depth or range of sequencing depths is achieved for a given region or sub-region of the genome. The sequencing depth for a region can be at least 0.1x, 0.5x, 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 15x, 20x, 25x, 30x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 125x, 150x, 175x, 200x, 300x, 400x, 500x or more. The sequencing depth for a region can be no more than 0.1x, 0.5x, 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 15x, 20x, 25x, 30x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 125x, 150x, 175x, 200x, 300x, 400x, 500x or less.
[0101] In some embodiments, more than 100 primer pairs are used for amplification. Amplification can be performed using about 10, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, about 130, about 140, about 150 or more primer pairs. In some embodiments, the amplification is multiplex amplification. Multiplex amplification can allow for the collection of a large amount of sequence information in parallel from many target regions in the genome, even from cfDNA samples where DNA is typically not abundant. Multiplexing can scale to a platform such as ION where up to about 24,000 amplicons can be queried simultaneously. In some embodiments, the amplification is nested amplification. Nested amplification can increase sensitivity and specificity.
[0102] The amplification reaction can be performed on nucleic acids that have hybridized to a probe. Similarly, the amplicons and extension products generated by the primers can undergo hybridization reactions involving probes.
[0103] The methods and systems provided herein can be used to prepare cell-free polynucleotide sequences for downstream sequencing reactions. In some embodiments, the sequencing method is classical Sanger sequencing, nanopore sequencing, or long-read sequencing. Examples of sequencing methods can include, but are not limited to: high-throughput sequencing, pyrosequencing, sequencing by synthesis, single-molecule sequencing, long-read sequencing (PacBio), nanopore sequencing, semiconductor sequencing, ligation sequencing, hybridization sequencing, RNA-Seq Digital gene expression Next-generation sequencing, synthetic single-molecule sequencing (SMSS) Massively parallel sequencing, clonal single molecule arrays (Solexa), shotgun sequencing, Maxim-Gilbert sequencing, primer walking, and any other sequencing method.
[0104] The methods disclosed herein can include performing one or more enrichment reactions on one or more nucleic acid molecules in a sample. The methods disclosed herein can include performing differential enrichment reactions on two or more nucleic acid molecules in a sample, thereby generating different enrichment amounts for different nucleic acids. The enrichment reaction can include contacting the sample with one or more probes or probe sets. The enrichment reaction can include differential amplification of two or more nucleic acid molecules in the sample. The enrichment reaction can be based on genetic or epigenetic parameters of the nucleic acid for enrichment. For example, the enrichment can enrich nucleic acids belonging to a specific region of the genome. The enrichment can include enrichment for a specific mutation or suspected mutation region. The enrichment can include enrichment for a specific region that can be associated with copy number variation or copy number loss. The enrichment can include enrichment for a specific region that can be associated with cancer.
[0105] IV. Nucleic Acid Sequencing
[0106] In some embodiments, the generation of sequencing reads is performed by next-generation sequencing. This can allow for high-depth reads of a given region. These can be high-throughput methods, including, for example (Solexa) sequencing, DNB-Sequencer T7 or G400 (MGITech Co., Ltd), sequencing (GenapSys, Inc.), Roche 454 sequencing (Roche Sequencing Solutions, Inc.), Ion Torrent sequencing (ThermoFisher Scientific), and SOLiD sequencing (Thermo Fisher ). The number of sequencing reads can be adjusted according to the DNA input amount and the data depth required for the analysis.
[0107] In some embodiments, the generation of sequencing reads is performed simultaneously on samples obtained from multiple patients, where the cell-free nucleic acid fragments are barcoded for each patient. This allows for the parallel analysis of multiple patients in a single sequencing run.
[0108] In another aspect, the present disclosure provides a kit for detecting tumors, which includes reagents for implementing the above methods and instructions for detecting tumor signals. The reagents can include, for example, primer sets, PCR reaction components, and / or sequencing reagents.
[0109] A library can be prepared by adding adapters or adapter sequences. The adapter sequences can allow nucleic acids to attach to a flow cell or other solid support. The adapter sequences can include sequences that can allow amplification of the library. Sequencing primers or other primers can bind to the adapter sequences to generate additional copies of the nucleic acids and can allow sequencing to be performed. The adapters can be ligated to the nucleic acids. The adapters can be ligated to both ends of the nucleic acids. The adapters can have both single-stranded and double-stranded regions (e.g., Y-shaped adapters). The adapters can be double-stranded adapters. The adapters can include barcode sequences or unique molecular identifier sequences. The adapters can include methylated nucleotides. For example, the adapters can include methylated cytosine. The library can be generated by fragmentation, ligation, amplification, extension, polymerization, or other enzymatic conversions or other reactions. The reactions or enzymatic conversions can allow generation of nucleic acids suitable for sequencing by the sequencing methods and sequencers described elsewhere herein.
[0110] The sequencing depth can at least partially depend on or be related to the efficiency of nucleic acid enrichment. The greater the number of sequencing molecules corresponding to a certain region, the greater the associated sequencing depth can be. By adjusting the efficiency of the enrichment reaction for a specific region, the depth of a given region can be increased or decreased compared to another region. The ability to adjust or otherwise control the sequencing depth can allow for data customization.
[0111] The sequence depth of a specific region can be different from the sequencing depth for another region. As described elsewhere herein, the method can allow for the adjustment, tuning, or customization of the sequencing depth for a given region. The sequencing depth for a region can be at least 0.1x, 0.5x, 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 15x, 20x, 25x, 30x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 125x, 150x, 175x, 200x, 300x, 400x, 500x or more. The sequencing depth for a region can be no more than 0.1x, 0.5x, 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 15x, 20x, 25x, 30x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 125x, 150x, 175x, 200x, 300x, 400x, 500x or less.
[0112] Compared to the sensitivity of sequencing reactions that do not use the enrichment strategies described herein, the methods and systems disclosed herein can increase the sensitivity of one or more sequencing reactions. The sensitivity of one or more sequencing reactions can be increased by at least about 1%, 2%, 3%, 4%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, 10%, 10.5%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, 90%, 95%, 97% or more.
[0113] V. Computer Systems
[0114] The present disclosure provides computer systems programmed to implement the methods described herein. Figure 1 Shown is a computer system 101 that is programmed or otherwise configured to store, process, identify, or interpret object data, biological data, biological sequences, and reference sequences. The computer system 101 can process various aspects of the object data, biological data, biological sequences, or reference sequences of the present disclosure. The computer system 101 can be an electronic device of a user or a computer system located distal to the electronic device. The electronic device can be a mobile electronic device.
[0115] The computer system 101 includes a central processing unit (CPU, also referred to herein as “processor” and “computer processor”) 105, which can be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 101 also includes a memory or storage location 110 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 115 (e.g., hard disk), a communication interface 120 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 125 such as caches, other memories, data storage, and / or electronic display adapters. The memory 110, storage unit 115, interface 120, and peripheral devices 125 communicate with the CPU 105 via a communication bus (solid lines), such as a motherboard. The storage unit 115 can be a data storage unit (or data repository) for storing data. With the aid of the communication interface 120, the computer system 101 can be operably coupled to a computer network (“network”) 130. The network 130 can be the Internet, an intranet, and / or an extranet, or an intranet and / or extranet that communicates with the Internet. In some instances, the network 130 is a telecommunications and / or data network. The network 130 can include one or more computer servers, which can implement distributed computing, such as cloud computing. In some instances, with the aid of the computer system 101, the network 130 can implement a peer-to-peer network, which can enable devices coupled to the computer system 101 to act as clients or servers.
[0116] The CPU 105 can execute a series of machine-readable instructions, which can be embodied in a program or software. The instructions can be stored in a memory location (such as the memory 110). The instructions can be directed to the CPU 105, which can then be programmed or otherwise configured to implement the methods of the present disclosure. Examples of operations performed by the CPU 105 can include fetching, decoding, executing, and writing back.
[0117] The CPU 105 can be part of a circuit (such as an integrated circuit). One or more other components of the system 101 can be included in the circuit. In some instances, the circuit is an application specific integrated circuit (ASIC).
[0118] The storage unit 115 can store files, such as drivers, libraries, and saved programs. The storage unit 115 can store user data, e.g., user preferences and user programs. In some instances, the computer system 101 can include one or more additional data storage units external to the computer system 101, such as located on a remote server that communicates with the computer system 101 via an intranet or the Internet.
[0119] The computer system 101 can communicate with one or more remote computer systems via the network 130. For example, the computer system 101 can communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), slate / tablet PCs (e.g., iPad, Galaxy Tab), telephones, smartphones (e.g., iPhone, Android-supported devices, ) or personal digital assistants. The user can access the computer system 101 via the network 130.
[0120] The methods described herein can be implemented by machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 101 (e.g., stored on the memory 110 or the electronic storage unit 115). The machine executable or machine readable code can be provided in the form of software. During use, the code can be executed by the processor 105. In some instances, the code can be retrieved from the storage unit 115 and stored on the memory 110 for access by the processor 105. In some instances, the electronic storage unit 115 can be excluded and the machine executable instructions can be stored on the memory 110.
[0121] The code can be pre-compiled and configured to be used with a machine having a processor suitable for executing the code, or it can be interpreted or compiled at runtime. The code can be provided in a programming language, and the programming language can be selected to enable the code to be executed in a pre-compiled, interpreted, or compiled manner.
[0122] Aspects of the systems and methods provided herein, such as computer system 101, can be embodied in programming. Various aspects of the technology can be considered "products" or "articles of manufacture", typically in the form of machine (or processor) executable code and / or associated data, which are carried or contained in one type of machine-readable medium. The machine executable code can be stored in an electronic storage unit such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. The "storage" type of medium can include any or all of the tangible memories or their associated modules of a computer, processor, etc., such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time. All or part of the software can sometimes be communicated via the Internet or various other telecommunications networks. For example, such communication can enable the software to be loaded from one computer or processor to another, e.g., from an administrative server or a main computer to the computer platform of an application server. Thus, another type of medium that can carry software elements includes optical, electrical, and electromagnetic waves, such as those used on physical interfaces between local devices via wired and optical landline networks and various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be considered media that carry software. As used herein, unless restricted to non-transitory, tangible "storage" media, the term such as computer or machine "readable medium" refers to any medium that participates in providing instructions to a processor for execution.
[0123] Thus, machine-readable media, such as computer-executable code, can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media includes, for example, optical or magnetic disks, such as any storage device in any one or more computers, such as may be used to implement databases shown in the figures. Volatile storage media includes dynamic memory, such as the main memory of such a computer platform. Tangible transmission media includes coaxial cables; copper wire and fiber optics, including the wires that make up the buses within a computer system. Carrier transmission media can take the form of electrical or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example: floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched cards, paper tapes, any other physical storage media with hole patterns, RAM, ROM, PROM, and EPROM, FLASH-EPROM, any other memory chip or cartridge, carriers that carry data or instructions, cables or links that carry such carriers, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media can involve carrying one or more sequences of one or more instructions to a processor for execution.
[0124] Computer system 101 can include or communicate with an electronic display 135 that includes a user interface (UI) 140 for providing, for example, nucleic acid sequences, concentrated nucleic acid samples, methylation profiles, expression profiles, and analysis of methylation or expression profiles. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0125] The methods and systems of the present disclosure can be implemented by one or more algorithms. The algorithms can be implemented by software when executed by a central processing unit 105. For example, the algorithms can store, process, identify, or interpret object data, biological data, biological sequences, and reference sequences.
[0126] Although certain examples of methods and systems have been shown and described herein, those skilled in the art will recognize that these are provided by way of example only and are not intended to be limiting in the specification. Many variations, changes, and alternatives will now occur to those skilled in the art without departing from the scope described herein. In addition, it should be understood that all aspects of the methods and systems are not limited to the specific descriptions, configurations, or relative proportions set forth herein, which depend on a variety of conditions and variables, and the description is intended to include such alternatives, modifications, variations, or equivalents.
[0127] In some instances, the subject matter disclosed herein may include at least one computer program or its use. A computer program may be a sequence of instructions that is executed in a CPU, GPU, or TPU of a digital processing device and is written to perform a specified task. Computer-readable instructions may be implemented as program modules that perform specific tasks or implement specific abstract data types, such as functions, objects, application programming interfaces (APIs), data structures, and the like. Given the disclosure provided herein, computer programs can be written in various languages in various versions.
[0128] In various environments, the functions of computer-readable instructions can be combined or allocated as needed. In some instances, a computer program may include a sequence of instructions. In some instances, a computer program may include multiple sequences of instructions. In some instances, a computer program may be provided from one location. In some instances, a computer program may be provided from multiple locations. In some instances, a computer program may include one or more software modules. In some instances, a computer program may partially or wholly include one or more web applications, one or more mobile applications, one or more stand-alone applications, one or more web browser plug-ins, extensions, add-ons, or add-ins, or a combination thereof.
[0129] In some instances, computer processing can be a method of statistics, mathematics, biology, or a combination thereof. In some instances, computer processing methods include dimensionality reduction methods, for example, including logistic regression, dimensionality reduction, principal component analysis, autoencoders, singular value decomposition, Fourier basis, singular value decomposition, wavelets, discriminant analysis, support vector machines, tree-based methods, random forests, gradient boosting trees, logistic regression, matrix factorization, network clustering, and neural networks such as convolutional neural networks.
[0130] In some instances, computer processing methods are supervised machine learning methods, including, for example, regression, support vector machines, tree-based methods, and networks.
[0131] In some instances, computer processing methods are unsupervised machine learning methods, including, for example, clustering, networks, principal component analysis, and matrix factorization.
[0132] Digital processing device
[0133] In some instances, the subject matter described herein may include a digital processing device or its use. In some instances, the digital processing device may include one or more hardware central processing units (CPUs), graphics processing units (GPUs), or tensor processing units (TPUs) that implement device functionality. In some instances, the digital processing device may include an operating system configured to execute executable instructions. In some instances, the digital processing device may optionally be connected to a computer network. In some instances, the digital processing device may optionally be connected to the Internet. In some instances, the digital processing device may optionally be connected to a cloud computing infrastructure. In some instances, the digital processing device may optionally be connected to an intranet. In some instances, the digital processing device may optionally be connected to a data storage device.
[0134] Non-limiting examples of suitable digital processing devices include server computers, desktop computers, portable computers, notebook computers, sub-notebook computers, netbook computers, set-top box computers, handheld computers, Internet devices, mobile smartphones, and tablet computers. Suitable tablet computers may include, for example, tablet computers having book, tablet, and convertible configurations.
[0135] In some instances, the digital processing device may include an operating system configured to execute executable instructions. For example, the operating system may include software, including programs and data, that manages the device's hardware and provides services for executing applications. Non-limiting examples of operating systems include Ubuntu, FreeBSD, OpenBSD, Mac OS X Windows and Non-limiting examples of suitable personal computer operating systems include Mac OS and UNIX-like operating systems such as In some instances, the operating system may be provided by cloud computing, and cloud computing resources may be provided by one or more service providers.
[0136] In some instances, the device may include a storage and / or memory device. The storage and / or memory device may be one or more physical devices for storing data or programs on a temporary or permanent basis. In some instances, the device may be volatile memory and require power to maintain the stored information. In some instances, the device may be non-volatile memory and retain the stored information when the digital processing device is not powered on. In some instances, the non-volatile memory may include flash memory. In some instances, the non-volatile memory may include dynamic random access memory (DRAM). In some instances, the non-volatile memory may include ferroelectric random access memory (FRAM). In some instances, the non-volatile memory may include phase change random access memory (PRAM).
[0137] In some instances, the device may be a storage device, including for example CD-ROM, DVD, flash memory device, disk drive, tape drive, optical disk drive, and cloud computing-based storage. In some instances, the storage and / or memory device may be a combination of those devices as disclosed herein. In some instances, the digital processing device may include a display for sending visual information to a user. In some instances, the display may be a cathode ray tube (CRT). In some instances, the display may be a liquid crystal display (LCD). In some instances, the display may be a thin film transistor liquid crystal display (TFT-LCD). In some instances, the display may be an organic light emitting diode (OLED) display. In some instances, the OLED display may be a passive matrix OLED (PMOLED) or an active matrix OLED (AMOLED) display. In some instances, the display may be a plasma display. In some instances, the display may be a video projector. In some instances, the display may be a combination of those devices as disclosed herein.
[0138] In some instances, the digital processing device may include an input device for receiving information from a user. In some instances, the input device may be a keyboard. In some instances, the input device may be a pointing device, including for example a mouse, trackball, trackpad, joystick, game controller, or stylus. In some instances, the input device may be a touch screen or a multi-touch screen. In some instances, the input device may be a microphone for capturing voice or other sound input. In some instances, the input device may be a camera for capturing motion or visual input. In some instances, the input device may be a combination of those devices as disclosed herein.
[0139] Non-transitory computer-readable storage medium
[0140] In some instances, the subject matter disclosed herein can include one or more non-transitory computer-readable storage media encoded with a program including instructions executable by an operating system of an optionally networked digital processing device. In some instances, the computer-readable storage media can be a tangible component of the digital processing device. In some instances, the computer-readable storage media can optionally be removed from the digital processing device. In some instances, the computer-readable storage media can include, for example, CD-ROMs, DVDs, flash devices, solid state memories, disk drives, tape drives, optical disc drives, cloud computing systems and services, and the like. In some instances, the program and instructions can be encoded on the media permanently, substantially permanently, semi-permanently, or non-transitorily.
[0141] database
[0142] In some instances, the subject matter disclosed herein can include one or more databases, or the use of such databases to store object data, biological data, biological sequences, or reference sequences. The reference sequences can be derived from the databases. Given the disclosure provided herein, many databases can be suitable for storing and retrieving sequence information. In some instances, suitable databases can include, for example, relational databases, non-relational databases, object-oriented databases, object databases, entity-relationship model databases, associative databases, and XML databases. In some instances, the database can be Internet-based. In some instances, the database can be network-based. In some instances, the database can be cloud computing-based. In some instances, the database can be based on one or more local computer storage devices.
[0143] In one aspect, the present disclosure provides a non-transitory computer-readable medium including instructions that direct a processor to implement the methods described herein.
[0144] In one aspect, the present disclosure provides a computing device including a computer-readable medium VI. Kits
[0145] The present disclosure provides kits for identifying or monitoring one or more cancer types of an object. The kit can include probes for capturing sequences at multiple genomic loci in a cell-free biological sample of the object. The probes can be selectively directed to sequences at multiple cancer-related genomic loci in the cell-free biological sample. The kit can include primers for amplifying sequences at multiple genomic loci in a cell-free biological sample of the object. The primers can be selectively directed to sequences at multiple cancer-related genomic loci in the cell-free biological sample. The kit can include instructions for processing the cell-free biological sample using the probes or primers.
[0146] The probes in the kit can selectively target sequences at multiple cancer-related genomic loci in a cell-free biological sample. The probes in the kit can be configured to selectively enrich nucleic acid (e.g., RNA or DNA) molecules corresponding to multiple cancer-related genomic loci. The probes in the kit can be nucleic acid primers. The probes in the kit can have sequence complementarity with nucleic acid sequences from one or more of multiple cancer-related genomic loci or genomic regions. The multiple cancer-related genomic loci or genomic regions can include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55 or more different cancer-related genomic loci or genomic regions.
[0147] The primers in the kit can selectively target sequences at multiple cancer-related genomic loci in a cell-free biological sample. The primers in the kit can be configured to selectively enrich nucleic acid (e.g., RNA or DNA) molecules corresponding to multiple cancer-related genomic loci. The primers in the kit can have sequence complementarity with nucleic acid sequences from one or more of multiple cancer-related genomic loci or genomic regions. The multiple cancer-related genomic loci or genomic regions can include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55 or more different cancer-related genomic loci or genomic regions.
[0148] The instructions in the kit can include instructions for assaying a cell-free biological sample using probes that selectively target sequences at multiple cancer-associated genomic loci in the cell-free biological sample. These probes can be nucleic acid molecules (e.g., RNA or DNA) that have sequence complementarity to nucleic acid sequences (e.g., RNA or DNA) from one or more of the multiple cancer-associated genomic loci. These nucleic acid molecules can be primers or enrichment sequences. The instructions for assaying the cell-free biological sample can include instructions for performing hybridization in an array or solution, polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., DNA sequencing or RNA sequencing) to process the cell-free biological sample to generate a data set that indicates a sequence quantification metric (e.g., indicating presence, absence, or relative amount) at each of the multiple cancer-associated genomic loci in the cell-free biological sample. The sequence quantification metric (e.g., indicating presence, absence, or relative amount) at each of the multiple cancer-associated genomic loci in the cell-free biological sample can indicate one or more cancers.
[0149] Examples
[0150] Example 1: Capturing nucleic acid molecules using a set of tunable capture probes.
[0151] Experiments were performed using a methylation panel (example). The panel was 3.12 Mb in size and contained a 50:50 mixture of methylated and unmethylated probes. Approximately 4 μl of the panel was used for each target capture, and the concentration of each probe was 0.1 fM. In addition, for each target capture reaction, a second panel (prostate adenocarcinoma / PRAD panel) was added at different concentrations. The PRAD panel was 89 kB in size. The PRAD panel contained a 50:50 mixture of methylated and unmethylated probes. In the undiluted PRAD panel, the concentration of each probe was 0.1 fM. The PRAD probes were diluted and added at the following concentration ranges: Tunable 01: control DNA, 34x - 3,400x dilution; Tunable 03: control DNA, 500x - 1,500x dilution; Tunable 04: control DNA, 200x - 750x dilution; Tunable 07: cfDNA and control, 200x - 400x dilution. Figure 2 The median PRAD panel coverage for each of the tested cfDNA libraries is shown. The median PRAD panel coverage in the 1:1 treatment was 1,500. As the number of probes decreased, a decrease in the median coverage was observed. The range of the percentage of bait removal between samples in Example 7 was 12 - 24%.
[0152] Figure 3Shows the percentage of bases covered at sequencing depths of 30x (left), 50x (middle), or 100x (right) in the cfDNA library at 1:1 dilution, 1:200 dilution, 1:340 dilution, 1:400 dilution, and 1:0 dilution. Each point represents the percentage of bases at a given threshold in a library. At 1:200 and 1:340 dilutions, most bases are covered at 30 - 50x.
[0153] Figure 4 Shows the variation in coverage levels between individual experiments. Experiment 1 shows the highest amount of coverage variation, which can be due to the fact that Example 1 also has the highest percentage of bait removal (40 - 50%). All experiments except Example 7 were run on low - diversity sgDNA libraries, where the average methyl panel coverage was approximately 300 - 500x. Despite differences in sequencing depth, bait removal percentage, and input DNA type between experiments, each given treatment has a predictable coverage level.
[0154] Figure 5 Shows the sequencing depth of the coverage - reduced regions (calculated as total read mapping for each base in the PRAD region / total read mapping for each base to the methyl panel region * 100). The sequencing depth for the low - coverage regions is consistent between the two experiments, especially in the 1:200 treatment, where the average sequencing depth is 5.5% for Example 4 and 5.6% for Example 7. The reported numbers do not include any correction for bait - removed reads, which are on average 32% of the reads for Example 4 and 19% for Example 7.
[0155] Table 1 summarizes the data for all experiments and all DNA types. While reference control samples (sgDNA) are included, the same range of data still applies when looking only at the cfDNA libraries. Sequencing depth is calculated as the expected (e.g., typical or average) mapped coverage (total molecules, non - unique) for each region in the PRAD panel divided by the coverage in the methyl panel region for the same library. Both 1:200 and 1:340 consistently provide 30 - 50x coverage. Due to differences between replicate samples, experiments, and regions, a probe concentration slightly higher than expected can be used.
[0156] Table 1
[0157]
[0158] While the preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. The present invention is not intended to be limited to the specific examples provided within this specification. While the present invention has been described with reference to the foregoing specific description, the description and illustration of the embodiments herein are not intended to be construed in a limiting sense. Various variations, changes, and alternatives will now occur to those skilled in the art without departing from the present invention. In addition, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions described herein that depend on various conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be used to practice the present invention. Accordingly, it is contemplated that the present invention should also cover any such alternatives, modifications, variations, or equivalents. The appended claims are intended to define the scope of the present invention and are intended to thereby cover methods and structures within the scope of these claims and their equivalents.
Claims
1. A method, comprising: (a) Provide a sample derived from an object, wherein the sample comprises a plurality of nucleic acids; (b) Provide a first set of capture nucleic acids to the sample, the first set of capture nucleic acids enriching a first set of nucleic acids of the plurality of nucleic acids to generate a sufficient amount of the first set of nucleic acids for sequencing the first set of nucleic acids to a first sequencing depth; (c) Provide a second set of capture nucleic acids to the sample, the second set of capture nucleic acids enriching a second set of nucleic acids of the plurality of nucleic acids to generate a sufficient amount of the second set of nucleic acids for sequencing the second set of nucleic acids to a second sequencing depth, wherein the first sequencing depth and the second sequencing depth are different; and (d) Sequence the first set of nucleic acids and the second set of nucleic acids to generate sequencing reads.
2. The method according to claim 1, wherein the plurality of nucleic acids are derived from a cell-free sample.
3. The method according to claim 1, wherein the plurality of nucleic acids comprise cell-free DNA (cfDNA) or cell-free RNA (cfRNA).
4. The method according to claim 1, wherein the plurality of nucleic acids comprise circulating tumor DNA (ctDNA).
5. The method according to claim 1, wherein the first set of capture nucleic acids comprises more nucleic acids than the second set of capture nucleic acids.
6. The method according to claim 1, wherein the concentration of the first set of capture nucleic acids in the sample is higher than the concentration of the second set of capture nucleic acids in the sample.
7. The method according to claim 1, further comprising contacting the first set of capture nucleic acids with the plurality of nucleic acids for a first contact duration and contacting the second set of capture nucleic acids with the plurality of nucleic acids for a second contact duration, wherein the first contact duration and the second contact duration are different.
8. The method according to claim 1, further comprising contacting the first set of capture nucleic acids with the plurality of nucleic acids for a first contact duration and contacting the second set of capture nucleic acids with the plurality of nucleic acids for a second contact duration, wherein the first contact duration and the second contact duration are the same or substantially the same.
9. The method according to claim 1, wherein the first set of capture nucleic acids comprises a first tiling density of 1x.
10. The method according to claim 1, wherein the first set of capture nucleic acids comprises a first tiling density of 2x.
11. The method according to claim 1, wherein the first set of capture nucleic acids comprises a first tiling density of 0.5x.
12. The method according to claim 1, wherein the first set of capture nucleic acids comprises a first tiling density and the second set of capture nucleic acids comprises a second tiling density, wherein the first tiling density and the second tiling density are different.
13. The method according to claim 1, wherein the first set of capture nucleic acids comprises a first tiling density and the second set of capture nucleic acids comprises a second tiling density, wherein the first tiling density and the second tiling density are the same or substantially the same.
14. The method according to claim 10, wherein the first tiling density is generated by overlapping sequences in the nucleic acids of the first set of capture nucleic acids.
15. The method according to claim 1, wherein the first set of capture nucleic acids or the second set of capture nucleic acids comprises at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more nucleotides.
16. The method according to claim 1, wherein the first set of capture nucleic acids or the second set of capture nucleic acids comprises no more than 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or fewer nucleotides.
17. The method according to claim 1, wherein the nucleotide length of the first set of capture nucleic acids is shorter than the nucleotide length of the second set of capture nucleic acids.
18. The method according to claim 1, wherein the nucleotide length of the first set of capture nucleic acids is longer than the nucleotide length of the second set of capture nucleic acids.
19. The method according to claim 1, wherein the first set of capture nucleic acids has incomplete complementarity with the first set of nucleic acids.
20. The method according to claim 19, wherein the nucleic acid region of the first set of capture nucleic acids and the first set of nucleic acids has at least one mismatched base.
21. The method according to claim 19, wherein the nucleic acid region of the first set of capture nucleic acids and the first set of nucleic acids has at least two mismatched bases.
22. The method according to claim 19, wherein the nucleic acid region of the first set of capture nucleic acids and the first set of nucleic acids has at least three mismatched bases.
23. The method according to claim 1, wherein the first set of capture nucleic acids has complete complementarity with the first set of nucleic acids.
24. The method according to claim 1, wherein the first set of capture nucleic acids or the second set of capture nucleic acids comprises DNA.
25. The method according to claim 1, wherein the first set of capture nucleic acids or the second set of capture nucleic acids comprises RNA.
26. The method according to claim 1, wherein the first set of capture nucleic acids or the second set of capture nucleic acids comprises DNA and RNA.
27. The method according to claim 26, wherein the nucleic acids of the first set of capture nucleic acids comprise DNA and RNA.
28. The method according to claim 26, wherein the first set of capture nucleic acids comprises a first nucleic acid comprising DNA and a second nucleic acid comprising RNA.
29. The method according to claim 1, wherein the sequencing comprises performing a next-generation sequencing reaction.
30. The method according to claim 1, wherein the first sequencing depth is at least 10 reads.
31. The method according to claim 1, wherein the first sequencing depth is at least 100 reads.
32. The method according to claim 1, wherein the first sequencing depth is at least 1000 reads.
33. The method according to claim 1, wherein the first sequencing depth is no more than 10 reads.
34. The method according to claim 1, wherein the first sequencing depth is no more than 100 reads.
35. The method according to claim 1, wherein the first sequencing depth is no more than 1000 reads.
36. The method according to claim 30, wherein the second sequencing depth is at least 100 reads.
37. The method according to claim 31, wherein the second sequencing depth is at least 1000 reads.
38. The method according to claim 33, wherein the second sequencing depth is no more than 100 reads.
39. The method according to claim 34, wherein the second sequencing depth is no more than 1000 reads.
40. The method according to claim 1, wherein the first set of nucleic acids comprises sequences associated with cancer or a cell proliferative disorder.
41. The method according to claim 40, wherein the cancer or cell proliferative disorder is colon cancer or a cell proliferative disorder.
42. The method according to claim 40, wherein the cancer or cell proliferative disorder is selected from colorectal cancer, prostate cancer, lung cancer, breast cancer, pancreatic cancer, ovarian cancer, uterine cancer, liver cancer, esophageal cancer, gastric cancer, and thyroid cancer or a cell proliferative disorder.
43. The method according to claim 1, wherein (b) and (c) are performed simultaneously or substantially simultaneously.
44. The method according to claim 1, wherein (b) and (c) are performed sequentially.
45. The method according to claim 1, further comprising analyzing the sequencing reads to determine the presence of a genetic parameter.
46. The method according to claim 45, wherein the genetic parameter is a single nucleotide variant, a copy number variant, a deletion, an insertion, or a transversion.
47. The method according to claim 45, wherein the genetic parameter is associated with cancer or a cell proliferative disorder.
48. The method according to claim 1, further comprising analyzing the sequencing reads to determine whether the subject has cancer or a cell proliferative disorder.
49. A method comprising: (a) Provide a sample derived from an object, wherein the sample comprises a plurality of nucleic acids; (b) Differentially enrich at least a subset of the plurality of nucleic acids by contacting the plurality of nucleic acids with a plurality of oligonucleotides to generate an enriched subset of the plurality of nucleic acids, wherein at least a subset of the plurality of oligonucleotides anneals to the subset of the plurality of nucleic acids, wherein the subset of the plurality of oligonucleotides has different percentages of complementarity to the nucleic acids of the plurality of nucleic acids, and wherein a higher percentage of complementarity to the nucleic acids provides an increased enrichment ratio compared to a lower percentage of complementarity to the nucleic acids; and (c) Sequence the enriched subset of the plurality of nucleic acids to generate sequencing reads.
50. The method according to claim 49, wherein the plurality of nucleic acids are derived from a cell-free sample.
51. The method according to claim 49, wherein the plurality of nucleic acids comprise cfDNA or cfRNA.
52. The method according to claim 49, wherein the plurality of nucleic acids comprise ctDNA.
53. The method according to claim 49, wherein the plurality of oligonucleotides comprise more oligonucleotides annealing to a first nucleic acid of the plurality of nucleic acids than to a second nucleic acid of the plurality of nucleic acids.
54. The method according to claim 49, wherein the plurality of oligonucleotides comprise a higher concentration of oligonucleotides annealing to a first nucleic acid of the plurality of nucleic acids than to a second nucleic acid of the plurality of nucleic acids.
55. The method according to claim 49, wherein the plurality of oligonucleotides comprise a tiling density of 1x.
56. The method according to claim 49, wherein the plurality of oligonucleotides comprise a tiling density of 2x.
57. The method according to claim 49, wherein the plurality of oligonucleotides comprise a tiling density of 0.5x.
58. The method according to claim 49, wherein a subset of the plurality of oligonucleotides configured to anneal to a first region of the nucleic acids of the plurality of nucleic acids has a different tiling density from a subset of the plurality of oligonucleotides configured to anneal to a second region of the nucleic acids of the plurality of nucleic acids.
59. The method according to claim 49, wherein a subset of the plurality of oligonucleotides configured to anneal to a first region of the nucleic acids of the plurality of nucleic acids has the same tiling density as a subset of the plurality of oligonucleotides configured to anneal to a second region of the nucleic acids of the plurality of nucleic acids.
60. The method according to claim 55, wherein the tiling density is generated by overlapping sequences in the oligonucleotides of the plurality of oligonucleotides.
61. The method according to claim 49, wherein the plurality of oligonucleotides includes oligonucleotides of different lengths.
62. The method according to claim 49, wherein the subset of the plurality of oligonucleotides has at least one mismatched base with the nucleic acid region of the plurality of nucleic acids.
63. The method according to claim 49, wherein the subset of the plurality of oligonucleotides has at least two mismatched bases with the nucleic acid region of the plurality of nucleic acids.
64. The method according to claim 49, wherein the subset of the plurality of oligonucleotides has at least three mismatched bases with the nucleic acid region of the plurality of nucleic acids.
65. The method according to claim 49, wherein the subset of the plurality of oligonucleotides has perfect complementarity with the nucleic acids of the plurality of nucleic acids.
66. The method according to claim 49, wherein the plurality of oligonucleotides comprises DNA.
67. The method according to claim 49, wherein the plurality of oligonucleotides comprises RNA.
68. The method according to claim 49, wherein the plurality of oligonucleotides comprises DNA and RNA.
69. The method according to claim 68, wherein the oligonucleotides of the plurality of oligonucleotides comprise DNA and RNA.
70. The method according to claim 68, wherein the first oligonucleotide of the plurality of oligonucleotides comprises DNA, and the second oligonucleotide of the plurality of oligonucleotides comprises RNA.
71. The method according to claim 49, wherein the sequencing includes performing a next-generation sequencing reaction.
72. The method according to claim 49, wherein the sequencing generates at least 10 reads for a first region of the nucleic acids of the plurality of nucleic acids.
73. The method according to claim 49, wherein the sequencing generates at least 100 reads for a first region of the nucleic acids of the plurality of nucleic acids.
74. The method according to claim 49, wherein the sequencing generates at least 1000 reads for a first region of the nucleic acids of the plurality of nucleic acids.
75. The method according to claim 49, wherein the sequencing generates no more than 10 reads for a first region of the nucleic acids of the plurality of nucleic acids.
76. The method according to claim 49, wherein the sequencing generates no more than 100 reads for a first region of the nucleic acid of the plurality of nucleic acids.
77. The method according to claim 49, wherein the sequencing generates no more than 1000 reads for a first region of the nucleic acid of the plurality of nucleic acids.
78. The method according to claim 72, wherein the sequencing generates at least 100 reads for a second region of the nucleic acid of the plurality of nucleic acids.
79. The method according to claim 73, wherein the sequencing generates at least 1000 reads for a second region of the nucleic acid of the plurality of nucleic acids.
80. The method according to claim 75, wherein the sequencing generates no more than 100 reads for a second region of the nucleic acid of the plurality of nucleic acids.
81. The method according to claim 76, wherein the sequencing generates no more than 1000 reads for a second region of the nucleic acid of the plurality of nucleic acids.
82. The method according to claim 49, wherein the subset of the plurality of nucleic acids comprises sequences associated with cancer or a cell proliferative disorder.
83. The method according to claim 82, wherein the cancer or cell proliferative disorder is colon cancer or a cell proliferative disorder.
84. The method according to claim 82, wherein the cancer or cell proliferative disorder is selected from colorectal cancer, prostate cancer, lung cancer, breast cancer, pancreatic cancer, ovarian cancer, uterine cancer, liver cancer, esophageal cancer, gastric cancer, and thyroid cancer or a cell proliferative disorder.
85. The method according to claim 49, further comprising analyzing the sequencing reads to determine the presence of a genetic parameter.
86. The method according to claim 85, wherein the genetic parameter is a single nucleotide variant, a copy number variant, a deletion, an insertion, or a transversion.
87. The method according to claim 85, wherein the genetic parameter is associated with cancer or a cell proliferative disorder.
88. The method according to claim 49, further comprising analyzing the sequencing reads to determine whether the subject has cancer or a cell proliferative disorder.