Methods and compositions for enriching nucleic acid molecules for sequencing

JP2025525372A5Pending Publication Date: 2026-04-17FREENOM HLDG INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
FREENOM HLDG INC
Filing Date
2023-06-22
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing nucleic acid sequencing technologies lack the ability to achieve customizable and efficient sequencing depth across different genomic regions, leading to unnecessary sequencing of uninformative DNA or RNA and increased costs.

Method used

A method for tunable targeted capture or enrichment of nucleic acid molecules, allowing for different sequencing depths and durations for distinct sets of nucleic acids, enabling variable coverage specific to each region of interest.

Benefits of technology

This approach increases data availability while reducing sequencing costs by focusing sequencing efforts on informative regions, improving the detection of diseases like cancer through enhanced sensitivity and specificity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and system for capturing or enriching nucleic acid sequences. Probes or primers can be used to capture or enrich nucleic acids. The properties of the probes or primers can be adjusted or regulated to generate sequencing depth for a given region. The sequencing depth can be non-uniform across a genomic region.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] cross reference This application claims the benefit of U.S. Provisional Patent Application No. 63 / 355,002, filed June 23, 2022, which is incorporated herein by reference in its entirety. [Background technology]

[0002] The present disclosure generally relates to the capture or enrichment of nucleic acid molecules. Nucleic acid molecules can be captured or enriched and sequenced to determine the nucleic acid sequence. Based on the sequence, certain diseases can be analyzed. For example, sequencing can be used for cancer screening or monitoring. This screening and monitoring can help improve outcomes, as early detection can eliminate cancer before it has a chance to metastasize, resulting in a better outcome. Summary of the Invention

[0003] The present disclosure provides methods and systems for tunable targeted capture or enrichment of nucleic acid molecules.

[0004] In one aspect, the present disclosure provides a method comprising: (a) providing a sample from a subject, the sample comprising a plurality of nucleic acids; (b) providing a first set of capture nucleic acids enriched in a first set of nucleic acids from the plurality of nucleic acids to the sample, thereby generating a quantity of the first set of nucleic acids sufficient to sequence the first set of nucleic acids to a first sequencing depth; (c) providing a second set of capture nucleic acids enriched in a second set of nucleic acids from the plurality of nucleic acids to the sample, thereby generating a quantity of the second set of nucleic acids sufficient to sequence the second set of nucleic acids to a second sequencing depth, wherein the first sequencing depth and the second sequencing depth are different; and (d) sequencing the first set of nucleic acids and the second set of nucleic acids to generate sequencing reads. In some embodiments, the plurality of nucleic acids is derived from a cell-free sample.

[0005] In some embodiments, the plurality of nucleic acids comprises cell-free DNA (cfDNA) or cell-free RNA (cfRNA). In some embodiments, the plurality of nucleic acids comprises circulating tumor DNA (ctDNA). In some embodiments, the first set of capture nucleic acids comprises more nucleic acids than the second set of capture nucleic acids. In some embodiments, the concentration of the first set of capture nucleic acids in the sample is higher than the concentration of the second set of capture nucleic acids in the sample. In some embodiments, the method further comprises contacting the first set of capture nucleic acids with the plurality of nucleic acids for a first contact duration and contacting the second set of capture nucleic acids with the plurality of nucleic acids for a second contact duration, wherein the first contact duration and the second contact duration are different. In some embodiments, the method further includes contacting a first set of capture nucleic acids with the plurality of nucleic acids for a first contact duration and contacting a second set of capture nucleic acids with the plurality of nucleic acids for a second contact duration, wherein the first contact duration and the second contact duration are the same or substantially the same.

[0006] In some embodiments, the first set of capture nucleic acids comprises a first tiling density of 1x. In some embodiments, the first set of capture nucleic acids comprises a first tiling density of 2x. In some embodiments, the first set of capture nucleic acids comprises a first tiling density of 0.5x. In some embodiments, the first set of capture nucleic acids comprises a first tiling density and the second set of capture nucleic acids comprises a second tiling density, wherein the first tiling density and the second tiling density are different. In some embodiments, the first set of capture nucleic acids comprises a first tiling density and the second set of capture nucleic acids comprises a second tiling density, wherein the first tiling density and the second tiling density are the same or substantially the same. In some embodiments, the first tiling density is generated by overlapping sequences in the first set of capture nucleic acids.

[0007] In some embodiments, the first set of capture nucleic acids or the second set of capture nucleic acids comprises at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more nucleic acids. In some embodiments, the first set of capture nucleic acids or the second set of capture nucleic acids comprises no more than 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or fewer nucleic acids. In some embodiments, the first set of capture nucleic acids is shorter in nucleotide length than the second set of capture nucleic acids. In some embodiments, the first set of capture nucleic acids is longer in nucleotide length than the second set of capture nucleic acids. In some embodiments, the set of capture nucleic acids is imperfectly complementary to the first set of nucleic acids. In some embodiments, the first set of capture nucleic acids contains at least one mismatched base with a region of one nucleic acid in the first set of nucleic acids. In some embodiments, the first set of capture nucleic acids contains at least two mismatched bases with a region of one nucleic acid in the first set of nucleic acids. In some embodiments, the first set of capture nucleic acids contains at least three mismatched bases with a region of one nucleic acid in the first set of nucleic acids. In some embodiments, the first set of capture nucleic acids is perfectly complementary to the first set of nucleic acids. In some embodiments, the first set of capture nucleic acids or the second set of capture nucleic acids comprises DNA. In some embodiments, the first set of capture nucleic acids or the second set of capture nucleic acids comprises RNA. In some embodiments, the first set of capture nucleic acids or the second set of capture nucleic acids comprises DNA and RNA. In some embodiments, the nucleic acids of the first set of capture nucleic acids comprise DNA and RNA. In some embodiments, the first set of capture nucleic acids comprises a first nucleic acid comprising DNA and a second nucleic acid comprising RNA.

[0008] In some embodiments, the sequencing comprises performing a next-generation sequencing reaction. In some embodiments, the first sequencing depth is at least 10 reads. In some embodiments, the first sequencing depth is at least 100 reads. In some embodiments, the first sequencing depth is at least 1000 reads. In some embodiments, the first sequencing depth is 10 reads or less. In some embodiments, the first sequencing depth is 100 reads or less. In some embodiments, the first sequencing depth is 1000 reads or less. In some embodiments, the second sequencing depth is at least 100 reads. In some embodiments, the second sequencing depth is at least 1000 reads. In some embodiments, the second sequencing depth is 100 reads or less. In some embodiments, the second sequencing depth is 100 reads or less.

[0009] In some embodiments, the first set of nucleic acids comprises sequences associated with cancer or a cell proliferative disorder. In some embodiments, the cancer or cell proliferative disorder is colon cancer or a cell proliferative disorder of the colon. In some embodiments, the cancer or cell proliferative disorder is selected from the group consisting of colorectal, prostate, lung, breast, pancreatic, ovarian, uterine, liver, esophageal, gastric, and thyroid cancer or a cell proliferative disorder. In some embodiments, steps (b) and (c) are performed simultaneously or substantially simultaneously. In some embodiments, steps (b) and (c) are performed sequentially. In some embodiments, the method further comprises analyzing the sequencing reads to determine the presence of a genetic parameter. In some embodiments, the genetic parameter is a single nucleotide variant, copy number variant, deletion, insertion, or transversion. In some embodiments, the genetic parameter is associated with cancer or a cell proliferative disorder. In some embodiments, the method further comprises analyzing the sequencing reads to determine whether the subject has cancer or a cell proliferative disorder.

[0010] In another aspect, the disclosure provides a method for detecting nucleic acids comprising: (a) providing a sample from a subject, the sample comprising a plurality of nucleic acids; (b) differentially enriching at least a subset of the plurality of nucleic acids by contacting the plurality of nucleic acids with a plurality of oligonucleotides, wherein at least a subset of the plurality of oligonucleotides anneal to the subset of the plurality of nucleic acids, the subset of the plurality of oligonucleotides having different percentages of complementarity to nucleic acids of the plurality of nucleic acids, a higher percentage of complementarity to a nucleic acid providing an increased enrichment ratio compared to a lower percentage of complementarity to the nucleic acid; and (c) sequencing the enriched subset of the plurality of nucleic acids to generate sequencing reads.

[0011] In some embodiments, the plurality of nucleic acids is derived from an acellular sample. In some embodiments, the plurality of nucleic acids comprises cfDNA or cfRNA. In some embodiments, the plurality of nucleic acids comprises ctDNA. In some embodiments, the plurality of oligonucleotides comprises a greater number of oligonucleotides that anneal to a first nucleic acid of the plurality of nucleic acids than to a second nucleic acid of the plurality of nucleic acids. In some embodiments, the plurality of oligonucleotides comprises a higher concentration of oligonucleotides that anneal to a first nucleic acid of the plurality of nucleic acids than to a second nucleic acid of the plurality of nucleic acids. In some embodiments, the plurality of oligonucleotides comprises a tiling density of 1x. In some embodiments, the plurality of oligonucleotides comprises a tiling density of 2x. In some embodiments, the plurality of oligonucleotides comprises a tiling density of 0.5x. In some embodiments, the subset of the plurality of oligonucleotides configured to anneal to a first region of a nucleic acid of the plurality of nucleic acids comprises a different tiling density than the subset of the plurality of oligonucleotides configured to anneal to a second region of a nucleic acid of the plurality of nucleic acids. In some embodiments, the subset of the plurality of oligonucleotides configured to anneal to a first region of a nucleic acid of the plurality of nucleic acids has the same tiling density as the subset of the plurality of oligonucleotides configured to anneal to a second region of a nucleic acid of the plurality of nucleic acids. In some embodiments, the tiling density is generated by overlapping sequences among the oligonucleotides of the plurality of oligonucleotides. In some embodiments, the plurality of oligonucleotides comprises oligonucleotides of different lengths. In some embodiments, the subset of the plurality of oligonucleotides comprises at least one mismatched base with respect to a region of one nucleic acid of the plurality of nucleic acids. In some embodiments, the subset of the plurality of oligonucleotides comprises at least two mismatched bases with respect to a region of one nucleic acid of the plurality of nucleic acids.In some embodiments, a subset of the plurality of oligonucleotides comprises at least three mismatched bases with respect to a region of one nucleic acid of the plurality of nucleic acids. In some embodiments, the subset of the plurality of oligonucleotides is fully complementary to a nucleic acid of the plurality of nucleic acids. In some embodiments, the plurality of oligonucleotides comprises DNA, in some embodiments, the plurality of oligonucleotides comprises RNA, in some embodiments, the plurality of oligonucleotides comprises DNA and RNA. In some embodiments, one oligonucleotide of the plurality of oligonucleotides comprises DNA and RNA. In some embodiments, a first oligonucleotide of the plurality of oligonucleotides comprises DNA and a second oligonucleotide of the plurality of oligonucleotides comprises RNA. In some embodiments, sequencing comprises performing a next-generation sequencing reaction. In some embodiments, sequencing generates at least 10 reads for a first region of nucleic acids of the plurality of nucleic acids. In some embodiments, sequencing generates at least 100 reads for a first region of nucleic acids of the plurality of nucleic acids. In some embodiments, sequencing generates at least 1000 reads for a first region of nucleic acids of the plurality of nucleic acids. In some embodiments, sequencing generates 10 or fewer reads for a first region of nucleic acids of the plurality of nucleic acids. In some embodiments, sequencing generates 100 or fewer reads for a first region of nucleic acids of the plurality of nucleic acids. In some embodiments, sequencing generates 1000 or fewer reads for a first region of nucleic acids of the plurality of nucleic acids. In some embodiments, sequencing generates at least 100 reads for a second region of nucleic acids of the plurality of nucleic acids. In some embodiments, sequencing generates at least 1000 reads for a second region of nucleic acids of the plurality of nucleic acids. In some embodiments, sequencing generates 100 or fewer reads for a second region of nucleic acids of the plurality of nucleic acids.In some embodiments, the sequencing generates 1000 or fewer reads for a second region of nucleic acids of said plurality of nucleic acids.

[0012] In some embodiments, the subset of nucleic acids in the plurality comprises a sequence associated with a cancer or a cell proliferative disorder. In some embodiments, the cancer or cell proliferative disorder is colorectal cancer or a cell proliferative disorder of the colorectum. In some embodiments, the cancer or cell proliferative disorder is selected from the group consisting of colorectal, prostate, lung, breast, pancreatic, ovarian, uterine, liver, esophageal, gastric, and thyroid cancer or a cell proliferative disorder. In some embodiments, the method further comprises analyzing the sequencing reads to determine the presence of a genetic parameter. In some embodiments, the genetic parameter is a single nucleotide variant, copy number variant, deletion, insertion, or transversion. In some embodiments, the genetic parameter is associated with a cancer or a cell proliferative disorder. In some embodiments, the method further comprises analyzing the sequencing reads to determine whether the subject has cancer or a cell proliferative disorder.

[0013] Another aspect of the present disclosure provides a computer-readable medium containing machine-executable code that, when executed by one or more computer processors, performs any of the methods described above or elsewhere herein.

[0014] Another aspect of the present disclosure provides a system comprising one or more computer processors and a computer memory coupled thereto, the computer memory including machine-executable code that, when executed by the one or more computer processors, performs any of the methods described above or elsewhere herein.

[0015]

[0013] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure have been shown and described. As will be understood, the present disclosure is capable of other and different embodiments, and its various details can be modified in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.

[0016] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that the publications and patents or patent applications incorporated by reference conflict with the disclosure contained herein, the present specification is intended to supersede and / or take precedence over any such conflicting material. [Brief explanation of the drawings]

[0017] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "Figures" and "FIGs.")

[0018] [Figure 1] FIG. 1 illustrates a computer system that is programmed or configured to perform the methods provided herein. [Figure 2] Figure 2 shows the median coverage of the prostate adenocarcinoma (PRAD) panel for cfDNA libraries. [Figure 3] Figure 3 shows the percent of bases covered in the cfDNA library. [Figure 4]Figure 4 shows the variation in median PRAD panel coverage levels across different enrichments. [Figure 5] Figure 5 shows the sequencing depth of the reduced coverage region. DETAILED DESCRIPTION OF THE INVENTION

[0019] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It will be understood that various alternatives to the embodiments of the invention described herein may be utilized.

[0020] The present disclosure generally relates to the capture or enrichment of nucleic acid molecules. Nucleic acid molecules can be captured or enriched and sequenced to determine the nucleic acid sequence. Based on the sequence, certain diseases can be analyzed. For example, sequencing can be used for screening or monitoring cancer or other diseases. This screening and monitoring can help improve outcomes, as early detection can identify diseases before they progress, leading to better outcomes.

[0021] Next-generation sequencing (NGS) technology can enable researchers or clinicians to explore an individual's entire genomic landscape. Such data can inform patients about their health status or disease risk. However, the majority of DNA or RNA found within a subject's (e.g., patient) sample (e.g., tissue, blood, plasma, urine, etc.) may be uninformative and therefore unnecessary to sequence. Target capture (or target enrichment) can be used to select regions of interest from a total pool of nucleic acids to generate an NGS library enriched for informative sequences and thus depleted of undesired nucleic acid fragments. To capture or enrich selected targets, nucleic acid molecules with sequences complementary to the regions of interest can be synthesized and then mixed with the sample. These nucleic acid molecules with sequences complementary to the regions of interest can hybridize with nucleic acids from the original sample and be captured or amplified, while non-target nucleic acids can be removed. In one embodiment, the method for capture involves hybridizing biotinylated oligonucleotides to nucleic acids from regions of interest in the original sample and capturing these regions using streptavidin-coated beads.

[0022] Target capture can be designed to achieve even sequencing coverage across every region of interest in a sample. However, the amount of sequenced reads required per site depends on many factors specific to the region of interest. For example, when searching for signal from circulating tumor DNA (ctDNA) in plasma, deep sequencing (e.g., 100–1000 reads per genomic region, or depth of coverage) may be necessary due to the low number of tumor-derived molecules compared to DNA from other sources. However, to genotype individuals for genes associated with cancer risk in the exact same sample, lower coverage (e.g., 10 reads) may be sufficient. This represents one of many use cases demonstrating the need for customizable sequencing depth specific to each individual region of interest. Having a method for achieving variable coverage in a deliberate manner within a single target capture reaction has the potential to increase data availability while lowering overall sequencing costs. For example, sequencing only one region at a particular coverage, as opposed to a library or entire genome at the same coverage, can allow fewer bases to be sequenced, thereby reducing the overall cost of sequencing.

[0023] Of particular interest is the detection of cell proliferative disorders in the lung, colon, liver, ovary, pancreas, prostate, rectum, and breast, and the enrichment or capture of genes associated with disease progression. For example, circulating tumor DNA can be a viable "liquid biopsy" for noninvasive tumor detection and surveillance. Identification of tumor-specific mutations in circulating tumor DNA can be applied to the diagnosis of colon cancer, breast cancer, and prostate cancer. However, due to the high background of normal (e.g., non-tumor-derived) DNA present in the circulation, these techniques may have limited sensitivity.

[0024] I. Definition As used in this specification and claims, the singular forms "a," "an," and "the" include the plural forms unless the context clearly indicates otherwise. For example, the term "a nucleic acid" includes a plurality of nucleic acids, including mixtures thereof.

[0025] As used herein, the term "subject" generally refers to an entity or medium that has testable or detectable genetic information. A subject may be a human, an individual, or a patient. A subject may be, for example, a vertebrate, such as a mammal. Non-limiting examples of mammals include humans, monkeys, farm animals, sport animals, rodents, and pets. A subject may be a human who has cancer or is suspected of having cancer. A subject may exhibit symptoms indicative of a health or physiological condition or disease of the subject, such as cancer or other disease, disorder, or illness of the subject. Alternatively, a subject may be asymptomatic with respect to such health or physiological condition or disease.

[0026] As used herein, the term "sample" generally refers to a biological sample obtained or derived from one or more subjects. A biological sample may be a cell-free or substantially cell-free biological sample, or may be processed or fractionated to generate a cell-free biological sample. For example, a cell-free biological sample may include cell-free ribonucleic acid (cfRNA), cell-free deoxyribonucleic acid (cfDNA), cell-free fetal DNA (cffDNA), plasma, serum, urine, saliva, amniotic fluid, and their derivatives. A cell-free biological sample may be obtained or derived from a subject using an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube (e.g., Streck®), or a cell-free DNA collection tube (e.g., Streck®). A cell-free biological sample may be derived from a whole blood sample by fractionation. A biological sample or its derivative may contain cells. For example, a biological sample may be a blood sample or its derivative (e.g., blood or a blood drop collected by a blood collection tube).

[0027] As used herein, the term "nucleic acid" generally refers to a polymeric form of nucleotides of any length, either deoxyribonucleotides (dNTPs) or ribonucleotides (rNTPs), or their analogs. Nucleic acids may have any three-dimensional structure and may perform any function, known or unknown. Non-limiting examples of nucleic acids include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), coding or non-coding regions of a gene or gene fragment, loci defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, small interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant nucleic acids, branched nucleic acids, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Nucleic acids can contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure can be made before or after assembly of the nucleic acid. The sequence of nucleotides of a nucleic acid can be interrupted by non-nucleotide components. Nucleic acids can be further modified after polymerization, such as by conjugation or conjugation with a reporter agent.

[0028] As used herein, the term "target nucleic acid" generally refers to a nucleic acid molecule in a starting population of nucleic acid molecules having a nucleotide sequence, the presence, amount, and / or sequence of which, or one or more changes therein, are desired to be determined. The target nucleic acid can be any type of nucleic acid, including DNA, RNA, and analogs thereof. As used herein, "target ribonucleic acid (RNA)" generally refers to a target nucleic acid that is RNA. As used herein, "target deoxyribonucleic acid (DNA)" generally refers to a target nucleic acid that is DNA.

[0029] As used herein, the terms "amplify" and "amplification" generally refer to increasing the size or amount of a nucleic acid molecule. A nucleic acid molecule can be single- or double-stranded. Amplification can include generating one or more copies of a nucleic acid molecule or "amplification product." Amplification can be performed, for example, by extension (e.g., primer extension) or ligation. Amplification can include performing a primer extension reaction to generate a strand complementary to a single-stranded nucleic acid molecule, and optionally generating one or more copies of the strand and / or single-stranded nucleic acid molecule. The term "DNA amplification" generally refers to generating one or more copies of a DNA molecule or "amplified DNA product." The term "reverse transcription amplification" generally refers to generating deoxyribonucleic acid (DNA) from a ribonucleic acid (RNA) template via the action of reverse transcriptase.

[0030] The term "cell-free nucleic acid (cfNA)" as used herein generally refers to nucleic acid that is not enclosed in cells in a biological sample, such as cell-free RNA ("cfRNA") or cell-free DNA ("cfDNA"). cfDNA can circulate freely in bodily fluids, such as in the bloodstream.

[0031] The term "acellular sample," as used herein, generally refers to a biological sample that is substantially devoid of intact cells. It may be derived from a biological sample that is itself substantially devoid of cells, or from a sample from which cells have been removed. Examples of acellular samples include those derived from blood, such as serum or plasma, urine, or samples derived from other sources, such as semen, saliva, feces, fallopian tube exudates, lymph, or collection lavage fluid.

[0032] As used herein, the term "circulating tumor DNA (ctDNA)" generally refers to cfDNA derived from a tumor.

[0033] The term "genomic region" as used herein generally refers to an identified region of nucleic acid that is identified by its location on a chromosome. In some instances, a genomic region is referred to by gene name and encompasses coding and non-coding regions associated with that physical region of nucleic acid. As used herein, a gene includes coding regions (exons), non-coding regions (introns), transcriptional control regions or other regulatory regions, and promoters. In another example, a genomic region may incorporate introns or exons, or intron / exon boundaries, within a named gene.

[0034] As used herein, the term "cell proliferative disorder" generally refers to a disorder or disease involving unregulated or abnormal growth of cells, such as cancer. In some non-limiting examples, the disorder is selected from colorectal cell proliferation, prostate cell proliferation, lung cell proliferation, breast cell proliferation, pancreatic cell proliferation, ovarian cell proliferation, uterine cell proliferation, liver cell proliferation, esophageal cell proliferation, gastric cell proliferation, or thyroid cell proliferation. In some embodiments, the cell proliferative disorder is selected from colon adenocarcinoma, hepatocellular carcinoma of the liver, lung adenocarcinoma, lung squamous cell carcinoma, severe adenoid cystic carcinoma of the ovary, pancreatic adenocarcinoma, prostate adenocarcinoma, and rectal adenocarcinoma.

[0035] As used herein, the terms "normal" or "healthy" generally refer to a cell, tissue, plasma, blood, biological sample, or subject that does not have a cell proliferative disorder.

[0036] The term "epigenetic parameters" as used herein generally refers to cytosine methylation. Additional epigenetic parameters include, for example, histone acetylation, which are not directly analyzed using the above method, but may instead be correlated with DNA methylation. Examples of epigenetic parameters may include, for example, other modifications of nucleotides, such as cytosine methylation, oxidation, deamination, fluoridation, hydroxymethylation, formylation, glycosylation, and amination.

[0037] The term "genetic parameters" as used herein generally refers to mutations and polymorphisms of genes and sequences that are further involved in their regulation. Examples of mutations include polymorphisms such as insertions, deletions, point mutations, inversions, and SNPs (single nucleotide polymorphisms).

[0038] The terms "type" and "subtype" of cancer are generally used relatively herein, such that a single "type" of cancer, such as breast cancer, may have "subtypes" based on, for example, stage, morphology, histology, gene expression, receptor profile, mutation profile, aggressiveness, prognosis, malignant characteristics, etc. Similarly, "type" and "subtype" may be applied at a finer level, such as differentiating a single histological "type" into "subtypes" defined, for example, according to mutation profile or gene expression. Cancer "stage" is also used to refer to the classification of cancer types based on histological and pathological features related to disease progression.

[0039] II. Sample The sample can be a biological sample. The sample can be derived from a biological sample. The biological sample can be, for example, blood, plasma, serum, urine, saliva, mucosal discharge, saliva, feces, or tears. The biological sample can be a fluid sample. The fluid sample can be a blood or plasma sample. The biological sample can be a tissue sample such as a biopsy, core biopsy, needle aspirate, or fine needle aspirate. The biological sample can be a fluid biological sample such as a blood sample, urine sample, or saliva sample. The biological sample can be a skin sample. The biological sample can be a cheek swab. The biological sample can be a plasma or serum sample. The biological sample may contain one or more cells. The biological sample can be, for example, blood, plasma, serum, urine, saliva, mucosal discharge, sputum, feces, or tears. The biological sample can contain cell-free nucleic acids (e.g., cell-free RNA, cell-free DNA, etc.). The sample can contain circulating tumor DNA (ctDNA). The sample can be a cell-free biological sample. The nucleic acid target can be a nucleic acid suspected of containing one or more mutations.

[0040] Acellular biological samples can be obtained or derived from human subjects.Before processing, acellular biological samples can be stored under various storage conditions, such as different temperatures (for example, at room temperature, under refrigerated or frozen conditions, at 25°C, 4°C, -18°C, -20°C or -80°C), or can be stored in various suspensions (for example, EDTA collection tubes, acellular RNA collection tubes or acellular DNA collection tubes).

[0041] The acellular biological sample can be obtained from a subject having cancer, from a subject suspected of having cancer, or from a subject not having or not suspected of having cancer. The cancer can be colon cancer.

[0042] Acellular biological samples can be collected before and / or after treatment of a subject with cancer. Acellular biological samples can be obtained from a subject during treatment or a treatment regimen. Multiple acellular biological samples can be obtained from a subject to monitor the effectiveness of treatment over time. Acellular biological samples can be collected from subjects known to have or suspected of having cancer where clinical tests have not provided a definitive positive or negative diagnosis. Samples can be collected from subjects suspected of having cancer. Acellular biological samples can be collected from subjects experiencing unexplained symptoms, such as fatigue, nausea, weight loss, aches and pains, weakness, or bleeding. Acellular biological samples can be collected from subjects with the described symptoms. Acellular biological samples can be collected from subjects at risk of developing cancer due to factors such as family history, age, hypertension or prehypertension, diabetes or prediabetes, overweight or obesity, environmental exposure, lifestyle risk factors (e.g., smoking, alcohol consumption, or drug use), or the presence of other risk factors.

[0043] The cell-free biological sample may contain one or more analytes that can be assayed, such as cell-free ribonucleic acid (cfRNA) molecules suitable for assays to generate transcriptome data, cell-free deoxyribonucleic acid (cfDNA) molecules suitable for assays to generate genomic and / or epigenomic data, or a mixture or combination thereof. One or more such analytes (e.g., cfRNA molecules and / or cfDNA molecules) can be isolated or extracted from one or more subject cell-free biological samples for downstream assays using one or more suitable assays. The cell-free biological sample may contain methylated nucleic acids. The methylated nucleic acids may contain methylated cytosines. The methylated nucleic acids can be analyzed, such as to identify correlations with epigenetic parameters or disease states or disorders.

[0044] A subset of nucleic acid samples or nucleic acid molecules may comprise one or more genomic regions. One or more genomic regions may comprise genetic parameters (e.g., polymorphisms, or portions thereof). The genetic parameters may be genetic abnormalities. For example, the genetic parameters may be mutations, single nucleotide polymorphisms, single nucleotide variants, insertions, deletions, fusions, copy number variants, copy number defects, or other changes in sequence or copy number in one or more nucleic acids. The genomic regions may comprise methylated nucleotides or epigenetic parameters. Capture of nucleic acids comprising genomic regions may allow the determination of nucleic acids in a sample or subject.

[0045] After obtaining acellular biological sample from a subject, the acellular biological sample can be processed to generate a data set that indicates the cancer of the subject.For example, the presence, absence, or quantitative evaluation of the nucleic acid molecules of the acellular biological sample in a panel of cancer-related genomic loci (for example, the quantitative measurement of RNA transcripts or DNA in cancer-related genomic loci).Processing the acellular biological sample obtained from a subject can include: (i) subjecting the acellular biological sample to conditions sufficient to isolate, concentrate, or extract a plurality of nucleic acid molecules; and (ii) assaying a plurality of nucleic acid molecules to generate a data set.

[0046] In some embodiments, multiple nucleic acid molecules are extracted from a cell-free biological sample and subjected to sequencing to generate multiple sequencing reads. The nucleic acid molecules may include ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). Nucleic acid molecules (e.g., RNA or DNA) can be extracted from a cell-free biological sample by various methods, such as the FastDNA Kit® kit protocol from MP Biomedicals®, the QIAamp® DNA Cell-Free Biology Mini Kit from Qiagen®, or the Cell-Free Biology DNA Isolation Kit protocol from Norgen Biotek®. The extraction method can extract all RNA or DNA molecules from the sample. Alternatively, the extraction method can selectively extract a portion of the RNA or DNA molecules from the sample. The RNA molecules extracted from the sample can be converted into DNA molecules by reverse transcription (RT).

[0047] Sequencing can be performed by any suitable sequencing method, such as massively parallel sequencing (MPS), paired-end sequencing, high-throughput sequencing, next-generation sequencing (NGS), shotgun sequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, pyrosequencing, sequencing-by-synthesis (SBS), sequencing-by-ligation, sequencing-by-hybridization, and RNA-Seq® (Illumina®).

[0048] Sequencing can include amplification of nucleic acids (e.g., RNA or DNA molecules). In some embodiments, the amplification of nucleic acids is polymerase chain reaction (PCR). An appropriate number of PCR rounds (e.g., PCR, qPCR, reverse transcriptase PCR, digital PCR, etc.) can be performed to sufficiently amplify the initial amount of nucleic acid (e.g., RNA or DNA) to the desired input amount for subsequent sequencing. In some cases, PCR can be used for global amplification of target nucleic acids. This can involve first using adapter sequences that can be ligated to different molecules, followed by PCR amplification using universal primers. PCR can be performed using any of several commercially available kits, for example, provided by Life Technologies®, Affymetrix®, Promega®, Qiagen®, etc. In other cases, only specific target nucleic acids within a population of nucleic acids can be amplified. Specific primers, perhaps in conjunction with adapter ligation, can be used to selectively amplify specific targets for downstream sequencing. PCR can include targeted amplification of one or more genomic loci, such as genomic loci associated with cancer. Sequencing may involve the use of simultaneous reverse transcription (RT) and polymerase chain reaction (PCR), such as, for example, the OneStep RT-PCR kit protocol from Qiagen®, NEB®, Thermo Fisher Scientific®, or Bio-Rad®.

[0049] The RNA or DNA molecules isolated or extracted from acellular biological samples can be tagged with, for example, identifiable tags to enable multiplexing of multiple samples.Any number of RNA or DNA samples can be multiplexed.For example, a multiplexing reaction can contain at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more than 100 RNA or DNA molecules derived from the initial acellular biological sample.For example, multiple acellular biological samples can be tagged with sample barcodes so that each DNA molecule can be traced back to the sample (and subject) from which the DNA molecule originates.Such tags can be attached to RNA or DNA molecules by ligation or by PCR amplification using primers.

[0050] After nucleic acid molecule is subjected to sequencing, suitable bioinformatics process can be carried out on sequence read to generate data that indicates the existence, non-existence, or relative evaluation of cancer.For example, sequence read can be aligned with one or more reference genomes (for example, the genome of one or more species, such as human genome).The aligned sequence read can be quantified at one or more genomic loci to generate a data set that indicates cancer.For example, quantifying the sequences corresponding to multiple genomic loci with or without genetic or epigenetic parameters associated with cancer can generate a data set that indicates cancer.

[0051] Acellular biological sample can be processed without any nucleic acid extraction.For example, cancer can be identified or monitored in a subject by using probe or primer that is configured to selectively enrich the nucleic acid (for example, RNA or DNA) molecules corresponding to multiple cancer-related genomic loci.Probe can have sequence complementarity with the nucleic acid sequence from one or more of multiple cancer-related genomic loci or genomic regions. The plurality of cancer-associated genomic loci or genomic regions can include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, at least about 90, at least about 95, at least about 100, or more distinct cancer-associated genomic loci or genomic regions.

[0052] The probes can be nucleic acid molecules (e.g., RNA or DNA) that have sequence complementarity to the nucleic acid sequences (e.g., RNA or DNA) of one or more genomic or epigenomic loci (e.g., cancer-associated genomic loci). These nucleic acid molecules can be primers or enrichment sequences. Assays of cell-free biological samples using probes selective for one or more genomic loci (e.g., cancer-associated genomic loci) can include the use of array hybridization (e.g., microarray-based), polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., RNA sequencing or DNA sequencing). In some embodiments, DNA or RNA may be assayed by one or more of isothermal DNA / RNA amplification methods (e.g., loop-mediated isothermal amplification (LAMP), helicase-dependent amplification (HDA), rolling circle amplification (RCA), recombinase polymerase amplification (RPA)), immunoassays, electrochemical assays, surface-enhanced Raman spectroscopy (SERS), quantum dot (QD)-based assays, molecular inversion probes, droplet digital PCR (ddPCR), CRISPR / Cas-based detection (e.g., CRISPR typing PCR (ctPCR), specific highly sensitive enzyme reporter unlocking (SHERLOCK), DNA endonuclease-targeted CRISPR transreporter (DETECTR), and CRISPR-mediated analog multi-event recording device (CAMERA)), and laser transmission spectroscopy (LTS).

[0053] The assay readout can be quantified at one or more genomic or epigenomic loci (e.g., cancer-associated genomic loci) to generate data suggestive of cancer. For example, quantification of array hybridization or polymerase chain reaction (PCR) corresponding to multiple genomic loci (e.g., cancer-associated genomic loci) can generate data suggestive of cancer. The assay readout can include quantitative PCR (qPCR) values, digital PCR (dPCR) values, digital droplet PCR (ddPCR) values, fluorescence values, etc., or normalized values thereof. The assay can be a home use test configured to be performed in a home environment.

[0054] III. Panel of probes or primers The present disclosure provides methods and systems for analyzing biological samples to obtain sequencing data of nucleic acids of interest. The sequencing data may include nucleic acids captured or enriched by a panel of probes or primers, or by multiple probes or primers.

[0055] A panel, as described herein, generally refers to a collection of targeted regions of genomic DNA identified in a biological sample. In some embodiments, the biological sample is a cell-free nucleic acid sample. The formation of a signature panel allows for rapid and specific analysis of regions associated with a disorder, disease, or specific genotype. The panels described and employed in the methods herein can be used for improved diagnosis, prognosis, treatment selection, and monitoring (e.g., treatment monitoring) of disorders or diseases, such as cancer.

[0056] The signature panels and methods offer a significant improvement over current approaches in that there is a need for marker or signature panels that can be used to detect early stage cell proliferative disorders from bodily fluid samples such as whole blood, plasma, or serum.

[0057] The present disclosure also provides a method for sequencing to identify genetic or epigenetic parameters of one or more genes.The genetic parameter can be a genetic abnormality.For example, the genetic parameter can be a mutation, a single nucleotide polymorphism, a single nucleotide variant, an insertion, a deletion, a fusion, a copy number variant, a copy number loss, or other changes in sequence or copy number in one nucleic acid or multiple nucleic acids.The method can include obtaining a sample from a subject and subjecting the nucleic acid to sequencing.Nucleic acid sequencing can include the sequencing technology and workflow as described elsewhere in this disclosure.

[0058] The tumor or cell proliferative disorder described herein can be selected from cell proliferation of the colorectum, prostate, lung, breast, pancreas, ovary, uterus, liver, esophagus, stomach, or thyroid. In some embodiments, the cell proliferative disorder is selected from colon adenocarcinoma, hepatocellular carcinoma of the liver, lung adenocarcinoma, lung squamous cell carcinoma, ovarian adenoid cystic carcinoma, pancreatic adenocarcinoma, prostate adenocarcinoma, and rectal adenocarcinoma.

[0059] In some embodiments, the cell proliferative disorder is a cell proliferative disorder of the large intestine, hi some embodiments, the cell proliferative disorder of the large intestine is chosen from adenoma (adenomatous polyp), polyposis disorder, Lynch syndrome, sessile hyperplastic adenoma (SSA), advanced adenoma, colorectal dysplasia, colorectal adenoma, colorectal cancer, colon cancer, rectal cancer, colorectal carcinoma, colorectal adenocarcinoma, carcinoid tumor, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST), lymphoma, and sarcoma.

[0060] The hybridization method provided herein can be used in various formats of nucleic acid hybridization, such as in-solution hybridization and hybridization on solid supports (e.g., Northern, Southern, and in situ hybridization on membranes, microarrays, and cell / tissue slides). In particular, the method is suitable for in-solution hybrid capture for target enrichment of specific types of genomic DNA sequences (e.g., exons) used in targeted sequencing. In the hybrid capture approach, the cell-free nucleic acid sample may be subjected to library preparation. As used herein, "library preparation" includes end repair, A-tailing, adapter ligation, or any other preparation performed on cell-free DNA to enable subsequent DNA sequencing. In some examples, the prepared cell-free nucleic acid library sequences contain adapters, sequence tags, index barcodes, UMIs, or combinations thereof, ligated onto the cell-free nucleic acid sample molecules. Various commercially available kits are available to facilitate library preparation for NGS (next-generation sequencing) approaches. NGS library construction can involve preparing nucleic acid targets using a coordinated series of enzymatic reactions to generate a random collection of DNA fragments of specific sizes for high-throughput sequencing. Advances and developments in various library preparation techniques have expanded the application of NGS to fields such as transcriptomics and epigenetics.

[0061] The improvement of sequencing technology has led to changes and improvements in library preparation. To ensure compatibility with the latest NGS equipment technology, various molecular biology reactions are provided with consistency and reproducibility, and NGS library preparation kits developed by companies such as Agilent (R), Bioo Scientific (R), Kapa Biosystems (R), New England Biolabs (R), Illumina (R), Life Technologies (R), Pacific Biosciences (R), Takara (R) / Clontech (R), Qiagen (R) and Roche (R) can be used.

[0062] In various examples of target capture gene panels, various library preparation kits are available, including Nextera Flex (Illumina®), Illumina® DNA Prep (Illumina®), Ion AmpliSeq® (Thermo Fisher Scientific®), GeneXus® (Thermo Fisher Scientific®), Agilent ClearSeq (Illumina®), Agilent® SureSelect® Capture (Illumina®), Archer® FusionPlex® (Illumina®), Bioo Scientific® NEXTflex® (Illumina®), IDT® xGen (Illumina®), Illumina® TruSight® (Illumina®), NimbleGen® SeqCap® (Illumina®), and Qiagen®. GeneRead® (Illumina®).

[0063] In some embodiments, hybrid capture is performed on the library sequence prepared using a specific probe. In some embodiments, the term "specific probe" as used herein generally refers to a probe specific to a certain region. In some embodiments, the specific probe is designed based on the use of the human genome as a reference sequence and a specific genomic region of interest. Therefore, when hybrid capture is performed using the specific probe of some embodiments, it can efficiently capture the sequence in the sample genome that is complementary to the target sequence.

[0064] According to the principle of complementary base pairing, single-stranded capture probes can be combined with the complementarity of single-stranded target sequences to successfully capture target regions.In some embodiments, designed probes can be designed as solid capture chips (probes are fixed on solid support) or liquid capture chips (probes are free in liquid).However, due to the limiting factors such as probe length, probe density and high cost, solid capture chips are rarely used, while liquid capture chips are more frequently used.

[0065] In some embodiments, GC-rich sequences in nucleic acids (where the GC base content is 60% or more) compared to normal sequences (where the average content of A, T, C, and G bases is 25% each) may result in increased capture efficiency due to the molecular structure of the C and G bases.

[0066] The number of probes added for each region of interest can be a specific amount or concentration. The number of probes can increase or decrease the final sequencing depth for a given region. For example, changing the number of probes targeting a given region can result in a change in the resulting sequencing depth for each region. A first region can have a larger number of probes annealing to it than a second region. A larger number of probes can allow for the capture of more nucleic acid sequences, resulting in an increased sequencing depth for that region. Conversely, a region with a smaller number of probes can be expected to capture fewer nucleic acids, resulting in a lower sequencing depth. In this way, sequencing depth can be adjusted or regulated based at least on the number of probes for a given region.

[0067] The amount of time allowed for hybridization to occur can be adjusted or otherwise varied. The hybridization step of a target capture reaction can vary from a few minutes to several hours. Varying the amount of time that complementary sequences are allowed to hybridize to each region of interest can result in varying coverage or depth for a given region. Probes that hybridize for a shorter time can result in lower recovery rates and lower sequencing coverage or depth in the region they target compared to probes that are allowed to hybridize for a longer time. Hybridization time can be adjusted by adding probes to the hybridization reaction at multiple time points to produce a specific sequencing depth. For example, in a 16-hour hybridization reaction, some probes can be hybridized for the entire 16 hours, while other probes can be added to the reaction after 15 hours, resulting in a 1-hour incubation time for the second set of molecules. Using this strategy, adjustable and customizable target coverage across regions can be achieved in a single reaction, resulting in different sequencing depths for different regions.

[0068] In some embodiments, the temperature that allows hybridization to occur can be adjusted or otherwise varied.The hybridization temperature of target capture reaction can vary from several minutes to several hours.The temperature at which complementary sequences can hybridize to each region of interest can change the coverage or depth for a given region.Approximate probe hybridization temperature can be calculated computationally.Using this approach, it is possible to achieve adjustable and customizable target coverage across regions in a single reaction, and to achieve different sequencing depths for different regions.

[0069] The concentration of molecules targeting the region of interest can be varied for each region. Given that 1x coverage (each region of interest has exactly one synthetic molecule designed to be complementary to that region of interest) is achieved in a typical target capture reaction, increasing the probe tiling to have more than one capture probe can result in higher coverage and higher sequencing depth. Alternatively, reducing the tiling density (for example, 0.5x), in which only a portion of the region of interest is covered by the probe, can result in lower sequencing coverage. In this manner, every region of interest can have a tiling density customized to generate a specific sequencing coverage for each region, and the first region can have a different coverage compared to the second region.

[0070] Probes can be of a specific length. For example, probes can be 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 nucleotides or more in length. Probes in a reaction can be of different lengths. For example, a first probe can be of a first length, and a second probe can be of a different length than the first probe. The number of bases that two molecules are directly complementary to can affect how strongly the molecules bind to each other, which can in turn affect the optimal temperature at which the molecules can bind (anneal) or separate (melt). Varying the length of probes in a target capture reaction (rather than making all probes one set length) can result in different optimal hybridization conditions across regions. Due to differences in annealing and melting temperatures across probes, targeting a region with probes of different lengths can result in subsequent differences in sequence coverage.

[0071] A probe can have a certain amount of complementarity to a target region. The efficiency with which two molecules hybridize can be affected by how perfectly their sequences match. A probe can have perfect complementarity to a target region, where each base of the probe is Watson-Crick paired with the target region. A probe may also have imperfect complementarity. For example, a probe may have mismatches to bases in the target region, so that not all bases are paired with the target region. This mismatch can result in lower hybridization efficiency. A mismatched probe can capture fewer nucleic acid molecules than a perfectly complementary probe. Mismatched bases introduced into a synthetic probe can reduce hybridization efficiency in proportion to the number of mismatches present in each region. Adding mismatches to a selected region of interest can result in lower target coverage or depth. Coverage or depth can be adjusted, in part, by using probes with various complementarities, so that regions where lower depth is desired can use probes with more mismatches.

[0072] Probes can also contain RNA, DNA, or both. Target capture probes can be synthesized using both DNA and RNA. A target capture reaction can be composed of a single class of molecules (DNA or RNA). Multiple probes can include RNA-containing probes and DNA-containing probes. DNA and RNA probes can differ in their hybridization affinity and their optimal hybridization conditions (temperature, timing, etc.). Using DNA probes in some regions of interest simultaneously with RNA probes in other regions can result in different coverage between the two groups due to inherent differences in how the two molecules may behave in a single reaction. A target capture panel composed of both DNA and RNA probes can enable differential coverage across regions within a single reaction. Probes can contain methylated or modified bases.

[0073] Probes can be used as a group or set of probes for a given reaction. Reactions can be performed sequentially, simultaneously, or overlapping with previous reactions. For example, a first set of probes can be added to a sample and allowed to anneal. After a certain time, a second set of probes can be added to the sample. The first set of probes can be removed before the addition of the second set, or can remain in the sample while the second set is added.

[0074] The probes can enable enrichment such that a particular sequencing depth or range of sequencing depths can be achieved for a given region or subregion of the genome. The sequencing depth of the region can be at least 0.1x, 0.5x, 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 15x, 20x, 25x, 30x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 125x, 150x, 175x, 200x, 300x, 400x, 500x, or more. The sequencing depth of a region can be less than or equal to 0.1x, 0.5x, 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 15x, 20x, 25x, 30x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 125x, 150x, 175x, 200x, 300x, 400x, 500x.

[0075] Nucleic acid amplification Nucleic acid molecules or fragments thereof can be amplified. Amplification can be used to enrich for specific sequences of interest. For example, a set of primers can anneal to a target sequence and generate an amplicon related to the sequence. The target sequence can then be present at an increased concentration, representing a larger fraction of the total molecules in a pool of molecules. In this way, a set of nucleic acid sequences can be enriched. The amount of enrichment can be correlated with the sequencing coverage or depth when the nucleic acid is sequenced. Molecules that have been subjected to enrichment can have a higher depth or sequence compared to molecules that have not been enriched. Increased enrichment or amplification of molecules can be correlated with a higher sequencing depth or coverage.

[0076] In various examples, the source of DNA can be cell-free DNA derived from whole blood, plasma, serum, or genomic DNA extracted from cells or tissues. In some embodiments, the size of the amplified fragments is approximately 100-200 base pairs (bp) in length. In some embodiments, the DNA source is extracted from a cellular source (e.g., tissue, biopsy, cell line), and the amplified fragments are approximately 100-350 bp in length. Amplification can be performed using a set of primer oligonucleotides, and a thermostable polymerase can be used. Amplification of multiple DNA segments can be performed simultaneously in the same reaction vessel. In some embodiments of the method, two or more fragments are amplified simultaneously. For example, amplification can be performed using polymerase chain reaction (PCR). In certain embodiments, the methods discussed herein can enable differential recovery of nucleic acid fragments of different sizes. For example, by increasing the tiling density for regions likely to have short (<100 nucleotide) fragments, these smaller fragments can be preferentially recovered compared to more rigid (e.g., 100-300 bp) fragments.

[0077] Primers are designed to target such sequences associated with or corresponding to a disease. In some embodiments, PCR primers are designed to be specific to genes associated with cancer. In some embodiments, primers are designed to be specific to genes associated with colon cancer.

[0078] Primers can be designed to amplify DNA fragments based on the expected (e.g., typical) size range of circulating DNA. Optimizing primer design to take target size into account can increase the sensitivity of this example method. In some embodiments, primers are designed to amplify DNA fragments 75-350 bp in length. Primers can be designed to amplify regions that are approximately 50-200 bp, approximately 75-150 bp, or approximately 100 or 125 bp.

[0079] Primers can be designed for target regions using appropriate tools such as Primer3, Primer3Plus, Primer-BLAST, etc. The design can include complementarity to a specific region or gene and can be designed to have specific properties, such as melting temperature, GC content, dimerization energy, or hairpin formation energy.

[0080] The number of primers added to each region of interest can be a specific amount or concentration. The number of primers can increase or decrease in relation to the final sequencing depth for a given region. For example, changing the number of primers targeting a given region can result in a change in the resulting sequencing depth for each region. A first region can have a larger number of primers annealing to the first region than a second region. A larger number of primers can allow for the capture of more nucleic acid sequences, resulting in an increased sequencing depth for that region. Conversely, a region with a smaller number of primers can allow for less nucleic acid amplification, resulting in a lower sequencing depth. In this way, sequencing depth can be adjusted or regulated based on at least the number of primers for a given region.

[0081] The amount of time allowed for hybridization, annealing, extension, or other reactions to occur can be adjusted or otherwise varied. Hybridization in an amplification reaction can vary from a few seconds to several hours. Varying the amount of time complementary sequences are allowed to hybridize to each region of interest can result in varying coverage or depth for a given region. Primers that hybridize for a shorter time may result in lower recovery rates and lower sequencing coverage or depth in their target region compared to primers that hybridize for a longer time. Hybridization time can be adjusted by adding primers to the hybridization reaction at multiple time points. Extension time can be altered to change the amount of time enzymes may need to generate extension or amplification products. Varying the extension time of a nucleic acid in a region of interest can result in varying coverage or depth for a given region. For example, extension products generated under shorter extension times may produce incomplete products that cannot be amplified by a second primer. The primers can be designed so that a first extension product is generated and can be amplified during the extension period, but the second extension product cannot be amplified during the extension period.

[0082] The amount of amplification cycles can be adjusted to differentially enrich the sequence of interest.The primer that anneals to the first region can be subjected to a certain number of cycles to generate a certain amount of amplicon, while the primer that anneals to the second region can be subjected to a different number of cycles.For example, in a 30-cycle amplification reaction, some primers can be added first and amplify for all 30 cycles, and other primers can be added to the reaction after 15 cycles, resulting in 15 cycles of amplification for the second set of molecules.Using this strategy, adjustable and customizable target coverage can be achieved across regions in a single reaction, and different sequencing depths can be achieved for different regions.

[0083] Primers can be of a specific length. For example, primers can be 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 nucleotides or more in length. Primers in a reaction can be of different lengths. For example, a first primer can be of a first length, and a second primer can be of a different length than the first primer. The number of bases that two molecules are directly complementary to can affect how strongly the molecules bind to each other, which can in turn affect the optimal temperature at which the molecules can bind (anneal) or separate (melt). Varying the length of primers in an amplification reaction (rather than all primers being of the same length) can vary the optimal hybridization conditions across multiple regions. Due to differences in annealing and melting temperatures across multiple primers, targeting multiple regions with primers of different lengths can subsequently result in differences in sequence coverage.

[0084] Primers can be designed to have a specific melting or annealing temperature. For example, primers can contain GC content. Depending on the annealing or melting temperature, some primers may be more or less efficient at different temperatures for amplification or extension. The amplification reaction conditions can have a temperature higher than the annealing or melting temperature of a set of primers. A set of primers may be less efficient or unable to generate extension at this temperature, while a set of primers with a higher melting temperature may be able to generate extension or amplification products at this temperature and more efficiently. The resulting amplification may result in more amplicons corresponding to the first region than amplicons corresponding to the second region.

[0085] Primers can be used as a group or set of primers for a given reaction. Reactions can be performed sequentially, simultaneously, or overlapping with previous reactions. For example, a first set of primers can be added to a sample and allowed to anneal. After a certain amount of time, a second set of primers can be added to the sample. The first set of primers can be removed before the addition of the second set, or can remain in the sample while the second set is added.

[0086] Primers can also contain RNA or DNA. Primers can be synthesized using both DNA and RNA. A target capture reaction can be composed of a single class of molecules (DNA or RNA). Multiple primers can include RNA-containing primers and DNA-containing primers. DNA primers and RNA primers can differ in hybridization affinity and their optimal hybridization conditions (temperature, timing, etc.). Using DNA primers in some regions of interest simultaneously with RNA primers in other regions can result in different coverage between the two groups due to inherent differences in how the two molecules may behave in a single reaction. Multiple primers composed of both DNA probes and RNA primers can enable differential coverage across regions within a single reaction. Primers can contain methylated or modified bases.

[0087] The primers can enable enrichment so that a specific sequencing depth or range of sequencing depths can be achieved for a given region or subregion of the genome. The sequencing depth of the region can be at least 0.1x, 0.5x, 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 15x, 20x, 25x, 30x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 125x, 150x, 175x, 200x, 300x, 400x, 500x, or more. The sequencing depth of a region can be less than or equal to 0.1x, 0.5x, 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 15x, 20x, 25x, 30x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 125x, 150x, 175x, 200x, 300x, 400x, 500x.

[0088] In some embodiments, the amplification is performed using more than 100 primer pairs. The amplification can be performed using about 10, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, about 130, about 140, about 150, or more primer pairs. In some embodiments, the amplification is a multiplex amplification. Multiplex amplification allows a large amount of sequence information to be collected in parallel from many target regions in the genome, even from cfDNA samples that are generally not rich in DNA. Multiplexing can be scaled up to platforms such as ION AmpliSeq®, for example, and can simultaneously search for up to about 24,000 amplicons. In some embodiments, the amplification is a nested amplification. Nested amplification can improve sensitivity and specificity.

[0089] An amplification reaction can be performed on nucleic acids that have been subjected to hybridization with a probe. Similarly, amplicons and extension products generated via primers can be subjected to hybridization reactions that include a probe.

[0090] The methods and systems provided herein can be useful for preparing cell-free polynucleotide sequences for downstream application sequencing reactions. In some embodiments, the sequencing method is classical Sanger sequencing, nanopore sequencing, or long-read sequencing. Examples of sequencing methods include, but are not limited to, high-throughput sequencing, pyrosequencing, sequencing-by-synthesis, single-molecule sequencing, long-read sequencing (PacBio), nanopore sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, RNA-Seq (Illumina®), Digital Gene Expression (Helicos®), next-generation sequencing, Single Molecule Sequencing by Synthesis (SMSS) (Helicos®), massively parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Maxim-Gilbert sequencing, primer walking, and any other sequencing method.

[0091] The methods disclosed herein may include performing one or more enrichment reactions for one or more nucleic acid molecules in a sample. The methods disclosed herein may include performing differential enrichment reactions for two or more nucleic acid molecules in a sample, such as generating different amounts of enrichment for different nucleic acids. The enrichment reaction may include contacting the sample with one or more probes or sets of probes. The enrichment reaction may include differential amplification of two or more nucleic acid molecules in a sample. The enrichment reaction may enrich based on genetic or epigenetic parameters of the nucleic acids. For example, the enrichment may enrich for nucleic acids belonging to a specific region of the genome. The enrichment may include enrichment for a specific mutation or region of suspected mutation. The enrichment may include enrichment for a specific region that may be associated with copy number variation or copy number loss. The enrichment may include enrichment for a specific region that may be associated with cancer.

[0092] IV. Nucleic Acid Sequencing In some embodiments, the generation of sequencing reads is carried out by next-generation sequencing.This can enable high-depth reading of a given region to be achieved.These can be high-throughput methods, such as Illumina® (Solexa) sequencing, DNB-Sequencer T7 (DNBSEQ®) or G400 (MGI Tech Co., Ltd.), GenapSys sequencing (GenapSys, Inc.), Roche 454 sequencing (Roche sequencing Solutions, Inc.), Ion Torrent sequencing (Thermo Fisher Scientific), and SOLiD sequencing (Thermo Fisher Scientific®).The number of sequencing reads can be adjusted according to the amount of DNA input and the depth of data required for analysis.

[0093] In some embodiments, generation of sequencing reads is performed simultaneously on samples from multiple patients, and cell-free nucleic acid fragments are barcoded for each patient, allowing for parallel analysis of multiple patients in a single sequencing run.

[0094] In another aspect, the present disclosure provides a kit for detecting tumors, comprising reagents for carrying out the aforementioned method and instructions for detecting tumor signals. The reagents may include, for example, primer sets, PCR reaction components, and / or sequencing reagents.

[0095] Libraries can be prepared by adding adapters or adapter sequences. The adapter sequence may allow the nucleic acid to attach to a flow cell or other solid support. The adapter sequence may include a sequence that can enable library amplification. A sequencing primer or other primer can bind to the adapter sequence to generate additional copies of the nucleic acid, allowing sequencing to occur. The adapter can be ligated to the nucleic acid. The adapter can be ligated to both ends of the nucleic acid. The adapter can have both single-stranded and double-stranded regions (e.g., a Y-shaped adapter). The adapter can be a double-stranded adapter. The adapter can include a barcode sequence or a unique molecular identifier sequence. The adapter can include a methylated nucleotide. For example, the adapter can include a methylated cytosine. The library can be generated by fragmentation, ligation, amplification, extension, polymerization, or other enzymatic conversion or other reaction. The reaction or enzymatic conversion can allow for the generation of nucleic acids suitable for sequencing by the sequencing methods and sequencing devices described elsewhere herein.

[0096] The depth of sequencing can be at least partially dependent on or correlated with the efficiency of nucleic acid enrichment.The greater the number of sequenced molecules corresponding to a region, the greater the correlation of sequencing depth.By adjusting the efficiency of enrichment reaction of a specific region, the depth of a given region can be increased or decreased compared to another region.The ability to adjust or otherwise control the depth of sequencing can enable customizable data.

[0097] The sequencing depth of one region can be different from the sequencing depth of another region.As described elsewhere herein, this method can allow the sequencing depth of a given region to be adjusted, regulated or customized.The sequencing depth of region can be at least 0.1x, 0.5x, 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 15x, 20x, 25x, 30x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 125x, 150x, 175x, 200x, 300x, 400x, 500x or more. The sequencing depth of a region can be less than or equal to 0.1x, 0.5x, 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 15x, 20x, 25x, 30x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, 100x, 125x, 150x, 175x, 200x, 300x, 400x, 500x.

[0098] The methods and systems disclosed herein can increase the sensitivity of one or more sequencing reactions compared to the sensitivity of a sequencing reaction that does not use the enrichment strategies described herein. The sensitivity of one or more sequencing reactions can be increased by at least about 1%, 2%, 3%, 4%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, 10%, 10.5%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, 90%, 95%, 97%, or more.

[0099] V. Computer Systems The present disclosure provides a computer system programmed to perform the methods described herein. Figure 1 shows a computer system (101) programmed or otherwise configured to store, process, identify, or interpret subject data, biological data, biological sequences, and reference sequences. The computer system (101) can process various aspects of the patient data, biological data, biological sequences, or reference sequences of the present disclosure. The computer system (101) may be a user's electronic device or computer system located remotely from the electronic device. The electronic device may be a mobile electronic device.

[0100] The computer system (101) includes a central processing unit (CPU, herein referred to as "processor" and "computer processor") (105), which can be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system (101) also includes memory or memory locations (110) (e.g., random access memory, read-only memory, flash memory), an electronic storage unit (115) (e.g., a hard disk), a communication interface (120) (e.g., a network adapter) for communicating with one or more other systems, and peripherals (125), such as cache, other memory, data storage devices, and / or electronic display adapters. The memory (110), storage unit (115), interface (120), and peripherals (125) communicate with the CPU 105 through a communication bus (solid lines), such as a motherboard. The storage unit (115) may also be a data storage unit (or data repository) for storing data. The computer system (101) is operatively coupled to a computer network ("network") (130) with the aid of a communication interface (120). The network (130) may be the Internet, an Internet and / or extranet, or an intranet and / or extranet in communication with the Internet. In some examples, the network (130) is a telecommunications and / or data network. The network (130) may include one or more computer servers, which may enable distributed computing, such as cloud computing. In some examples, the network (130) may implement a peer-to-peer network with the aid of the computer system (101), which may enable devices coupled to the computer system (101) to act as clients or servers.

[0101] The CPU (105) can execute sequences of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory (110). The instructions may be directed to the CPU (105), which can then program or otherwise configure the CPU (105) to implement the methods of the present disclosure. Examples of operations performed by the CPU (105) may include fetch, decode, execute, and writeback.

[0102] The CPU 105 may be part of a circuit, such as an integrated circuit. One or more other components of the system 101 may be included in the circuit. In some examples, the circuit is an application-specific integrated circuit (ASIC).

[0103] The storage unit (115) can store files such as drivers, libraries, and saved programs. The storage unit (115) can store user data, such as user personalization settings and user programs. The computer system (101) can include one or more additional data storage units that are external to the computer system (101), such as located on a remote server that communicates with the computer system (101) over an intranet or the Internet, in some examples.

[0104] The computer system (101) can communicate with one or more remote computer systems via the network (130). For example, the computer system (101) can communicate with a user's remote computer system. Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access the computer system (101) via the network (130).

[0105] Methods as described herein may be implemented by machine (e.g., computer processor) executable code stored in an electronic storage location of the computer system (101), such as on the memory (110) or on the electronic storage unit (115). The machine-executable or machine-readable code may be provided in the form of software. During use, the code may be executed by the processor (105). In some examples, the code may be retrieved from the storage unit (115) and stored in the memory (110) for easy access by the processor (105). In some examples, the electronic storage unit (115) may be omitted, and the machine-executable instructions are stored in the memory (110).

[0106] The code may be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or may be compiled at run time. The code may be provided in a programming language that may be selected to enable the code to be executed in a pre-compiled or on-the-fly compiled manner.

[0107] Aspects of the systems and methods provided herein, such as the computer system (101), may be embodied in programming. Various aspects of this technology may be thought of as "products" or "articles of manufacture," typically in the form of machine- (or processor-) executable code and / or associated data executed or embodied on a type of machine-readable medium. The machine-executable code may be stored in electronic storage, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage" type media may include any or all of the tangible memory of a computer or processor, or its associated modules, such as various semiconductor memories, tape drives, disk drives, etc., which may provide non-transitory recording media at any time for programming the software. All or portions of the software may sometimes be communicated via the Internet or various other telecommunications networks. Such communication may enable loading of the software from one computer or processor to another, for example, from a management server or host computer to an application server computer platform. Thus, other types of media that may carry software elements include optical, electrical, and electromagnetic waves, such as those used over wired and terrestrial optical communication networks between local devices and various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered media that carry software. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.

[0108] Thus, a machine-readable medium such as a computer-executable code may take many forms, including, but not limited to, a tangible storage medium, a carrier wave medium, or a physical transmission medium. Non-volatile storage media include optical or magnetic disks, such as any of the storage devices in any computer(s), such as those that may be used to implement the databases shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wire, and fiber optics, including the conductors that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, other magnetic media, CD-ROMs, DVDs or DVD-ROMs, other optical media, punch cards, paper tapes, other physical storage media with patterns of holes, RAM, ROM, PROMs and EPROMs, FLASH-EPROMs, other memory chips or cartridges, carrier waves that transport data or instructions, cables or links that transmit such carrier waves, or other media from which a computer can read programming code and / or data. Many forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0109] The computer system (101) may include or be in communication with an electronic display (135) that includes a user interface (UI) (140) for providing, for example, nucleic acid sequences, enriched nucleic acid samples, expression profiles, or analysis of expression profiles. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.

[0110] The methods and systems of the present disclosure may be implemented by one or more algorithms. The algorithms may be implemented by software when executed by a central processing unit (105). The algorithms may, for example, store, process, identify, or interpret patient data, biological data, biological sequences, and reference sequences.

[0111] While certain examples of methods and systems have been shown and described herein, those skilled in the art will understand that these are provided by way of example only and are not intended to be limiting within the specification. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the scope described herein. Furthermore, it should be understood that all aspects of the described methods and systems are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables, and that the description is intended to include all such alternatives, modifications, variations, or equivalents.

[0112] In some examples, the subject matter disclosed herein may include at least one computer program or its use. A computer program may be a set of instructions that can be executed by a CPU, GPU, or TPU of a digital processing device and written to perform a particular task. Computer-readable instructions may be implemented as program modules, such as functions, objects, application programming interfaces (APIs), data structures, etc., that perform particular tasks or implement particular abstract data types. In the context of the disclosure provided herein, computer programs may be written in a variety of languages and versions.

[0113] The functionality of the computer-readable instructions may be combined or distributed as desired in various environments. In some examples, a computer program may include one sequence of instructions. In some examples, a computer program may include multiple sequences of instructions. In some examples, a computer program may be provided from one location. In some examples, a computer program may be provided from multiple locations. In some examples, a computer program may include one or more software modules. In some examples, a computer program may include, in part or in whole, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins or add-ons, or combinations thereof.

[0114] In some examples, the computational processing may be a method of statistics, mathematics, biology, or a combination thereof. In some examples, the computational processing method includes a dimension reduction method, including, for example, logistic regression, dimensionality reduction, principal component analysis, autoencoder, singular value decomposition, Fourier-based, singular value decomposition, wavelets, discriminant analysis, support vector machines, tree-based methods, random forests, gradient boosted trees, logistic regression, matrix factorization, network clustering, and neural networks such as convolutional neural networks.

[0115] In some examples, the computational methods are supervised machine learning methods including, for example, regression, support vector machines, tree-based methods, and networks.

[0116] In some examples, the computational methods are unsupervised machine learning methods including, for example, clustering, networks, principal component analysis, and matrix factorization.

[0117] Digital Processing Device In some embodiments, the subject matter described herein may include a digital processing device or use thereof. In some embodiments, the digital processing device may include one or more hardware central processing units (CPUs), graphics processing units (GPUs), or tensor processing units (TPUs) that perform the functions of the device. In some examples, the digital processing device may include an operating system configured to implement the executable instructions. In some examples, the digital processing device may optionally be connected to a computer network. In some examples, the digital processing device may optionally be connected to the Internet. In some examples, the digital processing device may optionally be connected to a cloud computing infrastructure. In some examples, the digital processing device may optionally be connected to an intranet. In some examples, the digital processing device may optionally be connected to a data storage unit.

[0118] Non-limiting examples of suitable digital processing devices include server computers, desktop computers, laptop computers, notebook computers, subnotebook computers, netbook computers, netpad computers, set-top computers, handheld computers, Internet appliances, mobile smartphones, and tablet computers. Suitable tablet computers may include, for example, those having booklet, slate, and convertible configurations.

[0119] In some examples, a digital processing device may include an operating system configured to execute executable instructions. For example, an operating system may include software, including programs and data, that controls the device's hardware and provides services for the execution of applications. For example, non-limiting examples of operating systems include Ubuntu, FreeBSD, OpenBSD, NetBSD®, Linux®, Apple® Mac OS X Server®, Oracle® Solaris®, Windows Server®, and Novell® NetWare®. Non-limiting examples of suitable personal computer operating systems include Microsoft® Windows®, Apple® Mac OS X®, UNIX®, and UNIX-like operating systems such as GNU / Linux®. In some examples, the operating system may be provided by cloud computing, where cloud computing resources may be provided by one or more service providers.

[0120] In some examples, the device may include a storage and / or memory device. The storage and / or memory device may be one or more physical devices used to temporarily or permanently store data or programs. In some examples, the device may be volatile memory and may require power to maintain the stored information. In some examples, the device may be non-volatile memory and may retain the stored information when power is not supplied to the digital processing device. In some examples, the non-volatile memory may include flash memory. In some examples, the non-volatile memory may include dynamic random access memory (DRAM). In some examples, the non-volatile memory may include ferroelectric random access memory (FRAM). In some examples, the non-volatile memory may include phase change random access memory (PRAM).

[0121] In some examples, the device may be a storage device, including, for example, a CD-ROM, a DVD, a flash memory device, a magnetic disk drive, an optical disk drive, and a cloud computing-based storage device. In some examples, the storage and / or memory device may be a combination of devices as disclosed herein. In some examples, the digital processing device includes a display that transmits visual information to a user. In some embodiments, the display may be a cathode ray tube (CRT). In some embodiments, the display may be a liquid crystal display (LCD). In some examples, the display may be a thin film transistor liquid crystal display (TFT-LCD). In some embodiments, the display may be an organic light emitting diode (OLED) display. In some embodiments, the OLED display may be a passive-OLED (PMOLED) or active-matrix OLED (AMOLED) display. In some embodiments, the display may be a plasma display. In some examples, the display may be a video projector. In some examples, the display may be a combination of devices as disclosed herein.

[0122] In some examples, the digital processing device may include an input device for receiving information from a user. In some examples, the input device may be a keyboard. In some examples, the input device may be a pointing device, including, for example, a mouse, trackball, trackpad joystick, game controller, or stylus. In some examples, the input device may be a touchscreen or multi-touchscreen. In some examples, the input device may be a microphone for capturing voice or other sound input. In some examples, the input device may be a video camera for capturing motion or visual input. In some examples, the input device may be a combination of devices, such as those disclosed herein.

[0123] Non-transitory computer-readable storage medium In some examples, the subject matter disclosed herein may optionally include one or more non-transitory computer-readable storage media encoded with a program including instructions executable by an operating system of a network-connected digital processing device. In some examples, the computer-readable storage medium may be a tangible component of the digital processing device. In some examples, the computer-readable storage medium may optionally be removable from the digital processing device. In some examples, the computer-readable storage medium may include, for example, CD-ROMs, DVDs, flash memory devices, solid-state memory, magnetic disk drives, magnetic tape drives, optical disk drives, cloud computing systems and services, etc. In some examples, the program and instructions may be encoded on the medium permanently, substantially permanently, semi-permanently, or non-transitoryly.

[0124] Database In some embodiments, the subject matter disclosed herein may include one or more databases or their use for storing subject data, biological data, biological sequences, or reference sequences. Reference sequences can be obtained from a database. In view of the disclosure provided herein, many databases may be suitable for storing and retrieving sequence information. In some examples, suitable databases may include, for example, relational databases, non-relational databases, object-oriented databases, object databases, entity-relationship model databases, associative databases, and XML databases. In some examples, the database may be internet-based. In some examples, the database may be web-based. In some examples, the database may be cloud computing-based. In some examples, the database may be based on one or more local computer storage devices.

[0125] In an aspect, the present disclosure provides a non-transitory computer-readable medium comprising instructions that direct a processor to perform the methods disclosed herein.

[0126] In an aspect, the present disclosure provides a computing device comprising a computer-readable medium.

[0127] VI. Kit The present disclosure provides a kit for identifying or monitoring one or more cancers in a subject. The kit may include probes for capturing sequences at multiple genomic loci in a cell-free biological sample from the subject. The probes may be selective for sequences at multiple cancer-associated genomic loci in the cell-free biological sample. The kit may include primers for amplifying sequences at multiple genomic loci in the cell-free biological sample from the subject. The primers may be selective for sequences at multiple cancer-associated genomic loci in the cell-free biological sample. The kit may include instructions for using the probes or primers to process the cell-free biological sample.

[0128] The probes in the kit may be selective for sequences at multiple cancer-associated genomic loci in a cell-free biological sample. The probes in the kit may be configured to selectively enrich nucleic acid (e.g., RNA or DNA) molecules corresponding to multiple cancer-associated genomic loci. The probes in the kit may be nucleic acid primers. The probes in the kit may have sequence complementarity with nucleic acid sequences from one or more of the multiple cancer-associated genomic loci or genomic regions. The multiple cancer-associated genomic loci or genomic regions may include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, or more distinct cancer-associated genomic loci or genomic regions.

[0129] The primers in the kit may be selective for sequences at multiple cancer-associated genomic loci in a cell-free biological sample. The primers in the kit may be configured to selectively enrich nucleic acid (e.g., RNA or DNA) molecules corresponding to multiple cancer-associated genomic loci. The primers in the kit may have sequence complementarity with one or more nucleic acid sequences from multiple cancer-associated genomic loci or genomic regions. The multiple cancer-associated genomic loci or genomic regions may include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, or more distinct cancer-associated genomic loci or genomic regions.

[0130] The instructions in the kit include instructions for assaying the cell-free biological sample using probes selective for sequences at multiple cancer-associated genomic loci in the cell-free biological sample. These probes can be nucleic acid molecules (e.g., RNA or DNA) having sequence complementarity to one or more nucleic acid sequences (e.g., RNA or DNA) of the multiple cancer-associated genomic loci. These nucleic acid molecules can be primers or enrichment sequences. The instructions for assaying the cell-free biological sample can include outlining performing array or in-solution hybridization, polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., DNA sequencing or RNA sequencing) to process the cell-free biological sample to generate a dataset indicating a quantitative measure (e.g., indicating presence, absence, or relative amount) of a sequence at each of the multiple cancer-associated genomic loci in the cell-free biological sample. A quantitative measure (e.g., indicating presence, absence, or relative amount) of a sequence at each of the multiple cancer-associated genomic loci in the cell-free biological sample can be indicative of one or more cancers. [Example]

[0131] Example 1: Capturing nucleic acid molecules using a set of tunable capture probes. An example experiment was performed using the Methyl Panel. This panel was 3.12 Mb in size and contained a 50:50 mix of methylated and unmethylated probes. Approximately 4 μl of the panel was used in each target capture, with each probe at a concentration of 0.1 fM. In addition, a second panel (prostate adenocarcinoma / PRAD panel) was added at various concentrations to each target capture reaction. The PRAD panel was 89 kB in size. The PRAD panel contained a 50:50 mix of methylated and unmethylated probes. Each probe was at a concentration of 0.1 fM in the undiluted PRAD panel. PRAD probes were added at a range of dilutions: Tunable 01: control DNA, 34x to 3,400x diluted; Tunable 03: control DNA, 500x to 1,500x diluted; Tunable 04: control DNA, 200x to 750x diluted; and Tunable 07: cfDNA and control, 200x to 400x diluted. Figure 2 shows the median PRAD panel coverage for each cfDNA library tested. The median PRAD panel coverage for the 1:1 treatment was 1500. Median coverage was observed to decrease with fewer probes. In example 7, the percent off bait ranged from 12 to 24% across samples.

[0132] Figure 3 shows the percentage of bases covered at sequencing depths of 30x (left), 50x (middle), and 100x (right) in cfDNA libraries diluted 1:1, 1:200, 1:340, 1:400, and 1:0, respectively. Each point represents the percent of bases at a given threshold within one library. At both the 1:200 and 1:340 dilutions, the majority of bases are covered between 30x and 50x.

[0133] Figure 4 shows the variation in coverage levels across each experiment. Experiment 1 showed the highest amount of variation in coverage, which can be attributed to the fact that experiment 1 also had the highest off-bait percentage (40-50%). With the exception of experiment 7, all experiments were performed with low-diversity sgDNA libraries, where the average Methyl Panel coverage was approximately 300-500x. Despite differences in sequencing depth, off-bait percentage, and input DNA type across experiments, there was a predictable coverage level for each given treatment.

[0134] Figure 5 shows the sequencing depth of regions with low coverage (calculated as the total number of reads mapping per base in the PRAD region / total number of reads mapping per base to the Methyl Panel region * 100). Sequencing depth in low-coverage regions was consistent between the two experiments, particularly in the 1:200 treatment, where the average sequencing depth was 5.5% in ex. 4 and 5.6% in ex. 7. The reported numbers do not include any correction for off-bait reads, which averaged 32% of reads in ex. 4 and 19% in ex. 7.

[0135] Data from all experiments and all DNA types are summarized in Table 1. A reference control sample (sgDNA) is included, but the same range of data applies when looking at cfDNA libraries alone. Sequencing depth was calculated as the expected (e.g., typical or average) mapped coverage (total molecules, not unique) for each region of the PRAD panel divided by the coverage in the Methyl Panel region of that same library. Both 1:200 and 1:340 consistently provided 30-50x coverage. Due to variation across replicate samples, experiments, and regions, slightly higher probe concentrations than expected may be used.

[0136] [Table 1]

[0137] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present invention be limited by the specific examples provided within the specification. While the present invention has been described with reference to the above specification, the description and illustration of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it will be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It will be understood that various alternatives to the embodiments of the present invention described herein are available for practicing the invention. It is therefore contemplated that the present invention also encompasses such alternatives, modifications, variations, or equivalents. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.

Claims

1. (a) A step of providing a sample derived from a target, wherein the sample contains a plurality of nucleic acids, (b) A step of supplying the sample with a first set of captured nucleic acids that concentrates the first set of nucleic acids from the plurality of nucleic acids, thereby generating a sufficient amount of the first set of nucleic acids to sequence the first set of nucleic acids to a first sequencing depth, (c) A step of supplying a second set of captured nucleic acids to the sample to concentrate the second set of nucleic acids from the plurality of nucleic acids, thereby generating a sufficient amount of the second set of nucleic acids to sequence the second set of nucleic acids to a second sequencing depth, wherein the first sequencing depth and the second sequencing depth are different. (d) A step of sequencing the first set of nucleic acids and the second set of nucleic acids to generate a sequencing read. Methods that include...

2. The method according to claim 1, wherein the plurality of nucleic acids are derived from a cell-free sample.

3. The method according to claim 1, wherein the plurality of nucleic acids include cell-free deoxyribonucleic acid (cfDNA), cell-free ribonucleic acid (cfRNA), circulating tumor deoxyribonucleic acid (ctDNA), or a combination thereof.

4. The method according to claim 1, wherein the first set of captured nucleic acids contains more nucleic acids than the second set of captured nucleic acids.

5. The method according to claim 1, wherein the concentration of the first set of captured nucleic acids in the sample is higher than the concentration of the second set of captured nucleic acids in the sample.

6. The method according to claim 1, further comprising the steps of contacting a first set of captured nucleic acids with the plurality of nucleic acids for a first contact duration, and contacting a second set of captured nucleic acids with the plurality of nucleic acids for a second contact duration, wherein the first contact duration and the second contact duration are different, the same, or substantially the same.

7. The method according to claim 1, wherein the first set of captured nucleic acids has a first tiling density of 1x, 2x, or 0.5x.

8. The method according to claim 1, wherein the first set of captured nucleic acids has a first tiling density, and the second set of captured nucleic acids has a second tiling density, wherein the first tiling density and the second tiling density are different, the same, or substantially the same.

9. The method according to claim 8, wherein the first tiling density is generated by overlapping the sequences in the nucleic acids of the first set of captured nucleic acids.

10. The method according to claim 1, wherein the first set of captured nucleic acids has incomplete complementarity with respect to the first set of nucleic acids.

11. The method according to claim 10, wherein the first set of captured nucleic acids includes at least one mismatch base for one region of one nucleic acid of the first set of nucleic acids.

12. The method according to claim 10, wherein the first set of captured nucleic acids includes at least two mismatch bases for one region of one nucleic acid of the first set of nucleic acids.

13. The method according to claim 10, wherein the first set of captured nucleic acids includes at least three mismatched bases for one region of one nucleic acid in the first set of nucleic acids.

14. The method according to claim 1, wherein the first set of captured nucleic acids is perfectly complementary to the first set of nucleic acids.

15. The method according to claim 1, wherein the first set of captured nucleic acids or the second set of captured nucleic acids comprises DNA, RNA, or DNA and RNA.

16. The method according to claim 1, wherein the first sequencing depth is at least 10 reads, at least 100 reads, at least 1000 reads, 10 reads or less, 100 reads or less, or 1000 reads or less, and the second sequencing depth is at least 100 reads, at least 1000 reads, 100 reads or less, or 1000 reads or less.

17. The method according to claim 1, wherein the first set of nucleic acids includes sequences associated with cancer or cell proliferation disorders.

18. The method according to claim 17, wherein the cancer or cell proliferation disorder is colorectal cancer or a cell proliferation disorder of the colon.

19. The method according to claim 1, wherein steps (b) and (c) are performed simultaneously, substantially simultaneously, or in succession.

20. The method according to claim 1, further comprising the step of analyzing the sequencing reads to determine the presence of genetic parameters.

21. The method according to claim 20, wherein the genetic parameter is a single nucleotide variant, a copy number variant, a deletion, an insertion, or a transversion.