Single-chain specific terminal pattern
The use of hairpin adapters with molecular barcodes for ligating cfDNA fragments addresses the issue of artificial modifications in sequencing library preparation, enabling precise analysis of both 5' and 3' end motifs and ragged ends, improving the accuracy of cell-free DNA analysis for disease detection and monitoring.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-23
- Publication Date
- 2026-03-10
AI Technical Summary
Existing methods for analyzing cell-free DNA fragments fail to accurately capture the natural 3' end motifs and 3' overhanging ragged ends due to artificial modifications during sequencing library preparation, limiting the precision of nucleic acid analysis in applications such as cancer detection and non-invasive prenatal testing.
A new method involving hairpin adapters with molecular barcodes is used to ligate double-stranded cfDNA fragments, allowing for the simultaneous detection of both 5' and 3' end motifs and ragged ends without an end-repair step, thereby preserving the natural termini of the DNA fragments.
This approach enables high-fidelity analysis of fragmentomic features, including fragment size, terminal motifs, and ragged ends, enhancing the precision of nucleic acid analysis for disease detection and monitoring.
Smart Images

Figure 2026508245000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a nonprovisional application of and claims the benefit of U.S. Provisional Patent Application No. 63 / 447,847, entitled "SINGLE-MOLECULE STRAND-SPECIFIC END MODALITIES," filed February 23, 2023, which is incorporated herein by reference in its entirety for all purposes. [Background technology]
[0002] Cell-free DNA has proven particularly useful in molecular diagnostics and monitoring. Cell-free based applications include non-invasive prenatal testing (Chiu RKW et al. Proc Natl Acad Sci USA. 2008;105:20458-63), cancer detection and monitoring (Chan KCA et al. Clin Chem. 2013;59:211-24; Chan KCA et al. Proc Natl Acad Sci USA. 2013;110:1876-8; Jiang P et al. Proc Natl Acad Sci USA. 2015;112:E1317-25), transplant monitoring (Zheng YW et al. Clin Chem. 2012;58:549-58), and tissue of origin tracking (Sun K et al. Proc Natl Acad Sci USA. 2015;112:E5503-12; Chan KCA; Snyder MW et al. al. Cell. 2016;164:57-68). Cell-free nucleic acid analysis methods developed to date include those based on analysis of single nucleotide variants (SNVs), copy number abnormalities (CNAs), cell-free DNA end positions in the human genome, or methylation markers. It would be beneficial to identify new nucleic acid analysis methods to detect novel traits and add precision to existing methods. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Chiu RKW et al.Proc Natl Acad Sci USA.2008;105:20458-63 [Non-patent document 2] Chan KCA et al.Clin Chem.2013;59:211-24 [Non-patent document 3] Chan KCA et al.Proc Natl Acad Sci USA.2013;110:1876-8 [Non-patent document 4] Jiang P et al.Proc Natl Acad Sci USA.2015;112:E1317-25 [Non-Patent Document 5] Zheng YW et al.Clin Chem.2012;58:549-58 [Non-patent document 6] Sun K et al.Proc Natl Acad Sci USA.2015;112:E5503-12 [Non-Patent Document 7] Snyder MW et al.Cell.2016;164:57-68 Summary of the Invention
[0004] Double-stranded cell-free DNA fragments contain two termini of each strand. One molecule can have four termini. Because the two strands are often not exactly complementary to each other, one strand may extend beyond the other, creating overhangs at the termini. These overhangs are often repaired to form blunt ends during analysis, which alters the termini of the cell-free DNA fragments. This document describes how the natural termini can be obtained from each cell-free DNA fragment and used in analysis. The method can simultaneously evaluate the original 5' and 3' end motifs of the Watson and Crick strands and the associated raggedness at single-base resolution. In some embodiments, the entire fragmentomic feature from a cfDNA molecule can be accurately analyzed, including, but not limited to, 5'-overhanging ragged ends, 3'-overhanging ragged ends, 5'-recessed ragged ends, 3'-recessed ragged ends, overhanging ragged end motifs, recessed ragged end motifs, genomic coordinates of the fragment ends, fragment sizes, methylation-related cfDNA fragmentomic features, and combinations thereof. In some embodiments, a terminal motif can be defined by one or more nucleotides spanning positions near the end of the molecule. A terminal motif can be defined by one or more nucleotides in the reference genome surrounding the genomic locus to which the end of the fragment is aligned. In other embodiments, a ragged end can be defined by overhanging single-stranded DNA at the end of a DNA fragment. The ragged ends can be separated into different groups according to the length and / or stranding of the overhanging single-stranded DNA.
[0005] In some embodiments, different fragmentomics features from a single DNA fragment can be combined. In some embodiments, the combined fragmentomics features can be used to detect or monitor cancer or other diseases. In other embodiments, the combined fragmentomics features can be used for non-invasive prenatal testing.
[0006] A better understanding of the nature and advantages of embodiments of the present invention may be obtained by reference to the following detailed description and accompanying drawings. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 shows a schematic diagram of simultaneous single molecule end-to-end analysis on a single molecule real-time sequencing platform, according to an embodiment of the present invention. [Figure 2A] 2A and 2B show a schematic diagram of simultaneous single molecule end-to-end analysis on a next generation sequencing (NGS) platform (e.g., an Illumina platform) according to an embodiment of the present invention. [Figure 2B] 2A and 2B show a schematic diagram of simultaneous single molecule end-to-end analysis on a next generation sequencing (NGS) platform (e.g., an Illumina platform) according to an embodiment of the present invention. [Figure 3] FIG. 3 is a flowchart of an exemplary process for analyzing a biological sample, according to an embodiment of the present invention. [Figure 4] FIG. 4 is a flowchart of an exemplary process for analyzing a biological sample, according to an embodiment of the present invention. [Figure 5A] 5A and 5B show the frequency of various ragged ends according to an embodiment of the present invention. [Figure 5B] 5A and 5B show the frequency of various ragged ends according to an embodiment of the present invention. [Figure 6A] FIG. 6A shows a graph of the overall size distribution of plasma DNA samples from healthy and HCC subjects, according to an embodiment of the present invention. [Figure 6B] FIG. 6B shows a graph of the frequency of fragments less than 150 bp in size across the combination ragged end categories, in accordance with an embodiment of the present invention. [Figure 6C] FIG. 6C shows a graph of the frequency of fragments greater than 280 bp in size across the combination ragged end categories, in accordance with an embodiment of the present invention. [Figure 7A] 7A and 7B are graphs including different ragged end size ratios, according to an embodiment of the present invention. [Figure 7B] 7A and 7B are graphs including different ragged end size ratios, according to an embodiment of the present invention. [Figure 8] FIG. 8 is a graph of CCCA end motifs across various types of ends, according to an embodiment of the present invention. [Figure 9A]Figures 9A and 9B show a technique that can combine ragged ends, 5' end motifs, and 3' end motifs to measure the end-style topology of cfDNA molecules, according to an embodiment of the present invention. [Figure 9B] Figures 9A and 9B show a technique that can combine ragged ends, 5' end motifs, and 3' end motifs to measure the end-style topology of cfDNA molecules, according to an embodiment of the present invention. [Figure 10A] FIG. 10 illustrates a technique for naming junction end motifs according to an embodiment of the present invention. [Figure 10B] FIG. 10 illustrates a technique for naming junction end motifs according to an embodiment of the present invention. [Figure 11A] FIG. 11A is a graph of the correlation of the overall 5′ end motif frequency between HCC and healthy subjects. [Figure 11B] FIG. 11B is a graph of the correlation of the frequency of topological end motifs between HCC and healthy subjects, according to an embodiment of the present invention. [Figure 11C] FIG. 11C is a graph of the correlation of the frequency of junction terminal motifs between HCC and healthy subjects, according to an embodiment of the present invention. [Figure 12A] 12A-12F are graphs of the frequency of various ragged end patterns for various nuclear activities, according to an embodiment of the present invention. [Figure 12B] 12A-12F are graphs of the frequency of various ragged end patterns for various nuclear activities, according to an embodiment of the present invention. [Figure 12C] 12A-12F are graphs of the frequency of various ragged end patterns for various nuclear activities, according to an embodiment of the present invention. [Figure 12D] 12A-12F are graphs of the frequency of various ragged end patterns for various nuclear activities, according to an embodiment of the present invention. [Figure 12E] 12A-12F are graphs of the frequency of various ragged end patterns for various nuclear activities, according to an embodiment of the present invention. [Figure 12F]12A-12F are graphs of the frequency of various ragged end patterns for various nuclear activities, according to an embodiment of the present invention. [Figure 13] FIG. 5 is a table of median frequencies and relative changes of 5′ A-, T-, C-, and G-termini in fragments with 5′-overhanging ragged ends, 3′-overhanging ragged ends, and blunt ends in WT, DNASE1L3− / −, DNASE1− / −, and DFFB− / − mice according to an embodiment of the present invention. [Figure 14A] 14A-14D are graphs of terminal motif rankings for DFFB- / - (DFFB knockout [KO]) mice and wild-type (WT) mice, according to embodiments of the present invention. [Figure 14B] 14A-14D are graphs of terminal motif rankings for DFFB- / - (DFFB knockout [KO]) mice and wild-type (WT) mice, according to embodiments of the present invention. [Figure 14C] 14A-14D are graphs of terminal motif rankings for DFFB- / - (DFFB knockout [KO]) mice and wild-type (WT) mice, according to embodiments of the present invention. [Figure 14D] 14A-14D are graphs of terminal motif rankings for DFFB- / - (DFFB knockout [KO]) mice and wild-type (WT) mice, according to embodiments of the present invention. [Figure 15] FIG. 15 is a flowchart of an exemplary process for analyzing a biological sample, according to an embodiment of the present invention. [Figure 16] FIG. 16 is a flowchart of an exemplary process for analyzing a biological sample, according to an embodiment of the present invention. [Figure 17] FIG. 17 is a flowchart of an exemplary process for analyzing a biological sample, according to an embodiment of the present invention. [Figure 18A] 18A-18C show the frequency of various protruding ragged ends in fetal-specific and shared cfDNA fragments, according to embodiments of the present invention. [Figure 18B]18A-18C show the frequency of various protruding ragged ends in fetal-specific and shared cfDNA fragments, according to embodiments of the present invention. [Figure 18C] 18A-18C show the frequency of various protruding ragged ends in fetal-specific and shared cfDNA fragments, according to embodiments of the present invention. [Figure 18D] FIG. 18D is a graph of fetal DNA fraction estimated from fragments with various types of protruding ragged ends, according to an embodiment of the present invention. [Figure 19] 1 is a graph of fetal DNA fraction versus various ragged end patterns, according to an embodiment of the present invention. [Figure 20A] 20A and 20B are graphs of fetal DNA fractions estimated from fragments with certain ragged end patterns and sequence end motifs, according to embodiments of the present invention. [Figure 20B] 20A and 20B are graphs of fetal DNA fractions estimated from fragments with certain ragged end patterns and sequence end motifs, according to embodiments of the present invention. [Figure 21] FIG. 21 is a flowchart of an exemplary process for enriching a biological sample for clinically relevant DNA, according to an embodiment of the present disclosure. [Figure 22A] FIG. 22A is a graph of DNASE1 mRNA expression levels in leukocytes and placenta, according to an embodiment of the present invention. [Figure 22B] FIG. 22B is a graph of DFFB mRNA expression levels in leukocytes and placenta, according to an embodiment of the present invention. [Figure 22C] Figure 22C is a graph of the correlation between fetal DNA fraction and the frequency of cfDNA fragments carrying 5'-overhanging ragged ends, according to an embodiment of the present invention. [Figure 22D] Figure 22D is a graph of the correlation between fetal DNA fraction and the frequency of cfDNA fragments carrying blunt ends, according to an embodiment of the present invention. [Figure 23]FIG. 23 is a flowchart of an exemplary process for determining the fraction of clinically relevant DNA in a biological sample, according to an embodiment of the present invention. [Figure 24] FIG. 24 shows a measurement system according to an embodiment of the present invention. [Figure 25] FIG. 25 illustrates a computer system according to an embodiment of the present invention.
[0008] term A "tissue" corresponds to a group of cells grouped together as a functional unit. More than one type of cell can be found within a single tissue. Different types of tissue can consist of different types of cells (e.g., liver cells, alveolar cells, or blood cells), but can also correspond to tissues from different organisms (mother vs. fetus) or healthy cells vs. tumor cells. A "reference tissue" can correspond to the tissue used to determine tissue-specific methylation levels. Multiple samples of the same tissue type from different individuals can be used to determine the tissue-specific methylation level of that tissue type.
[0009] An "organ" corresponds to a group of tissues with similar functions. One or more types of tissue may be found within a single organ. Organs may be part of different organ systems, including the cardiovascular system, digestive system, endocrine system, excretory system, lymphatic system, integumentary system, muscular system, nervous system, reproductive system, respiratory system, and skeletal system.
[0010] A "biological sample" refers to any sample obtained from a subject (e.g., a pregnant woman, a person with or suspected of having cancer, or an organ transplant recipient) or a subject suspected of having a disease process involving an organ (e.g., the heart in myocardial infarction, the brain in stroke, or the hematopoietic system in anemia) and containing one or more nucleic acid molecules of interest. A biological sample can be a bodily fluid such as blood, plasma, serum, urine, vaginal fluid, fluid from edema (e.g., of the testes), vaginal washings, pleural effusion, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple discharge, aspirates from different parts of the body (e.g., thyroid, mammary glands), etc. Stool samples can also be used. In various embodiments, the majority of the DNA in a biological sample enriched for cell-free DNA (e.g., a plasma sample obtained via a centrifugation protocol) can be cell-free, e.g., greater than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA can be cell-free. The centrifugation protocol can include, e.g., 3,000 g x 10 minutes, obtaining a fluid portion, and recentrifuging, e.g., at 30,000 g for an additional 10 minutes, to remove residual cells.
[0011] A "sequence read" refers to a chain of nucleotides sequenced from any portion or all of a nucleic acid molecule. For example, a sequence read can be a short chain (e.g., 20-150) of nucleotides sequenced from a nucleic acid fragment, a short chain of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of an entire nucleic acid fragment present in a biological sample. Sequence reads can be obtained in a variety of ways, for example, using sequencing techniques, or using probes, e.g., hybridization arrays or capture probes, or using single primer or isothermal amplification, amplification techniques such as polymerase chain reaction (PCR) or linear amplification.
[0012] An "ending position" or "end position" (or simply "end") can refer to the genomic coordinate or genomic identity or nucleotide identity of the outermost base, i.e., end, of a cell-free DNA molecule, e.g., a plasma DNA molecule. An end position can correspond to either end of a DNA molecule. Thus, when referring to the beginning and end of a DNA molecule, both correspond to ending positions. In practice, one end position is the genomic coordinate or nucleotide identity of the outermost base at one end of a cell-free DNA molecule, as detected or determined by analytical methods such as, but not limited to, massively parallel sequencing or next-generation sequencing, single-molecule sequencing, double- or single-stranded DNA sequencing library preparation protocols, polymerase chain reaction (PCR), or microarrays. Thus, each detectable end may represent a biologically true end, or the end may be one or more nucleotides inward from the original end of the molecule, or one or more nucleotides extended from the original end of the molecule, e.g., 5' blunting and 3' filling of the overhang of a non-blunt-ended double-stranded DNA molecule by Klenow fragment. The genomic identity or genomic coordinate of the end position can be derived from the results of aligning the sequence read to a human reference genome, e.g., hg19. It may also be derived from a catalog of indexes or codes representing the original coordinates of the human genome. It may refer to the position or nucleotide identity on a cell-free DNA molecule read by, but not limited to, target-specific probes, minisequencing, or DNA amplification.
[0013] A "sequence motif" can refer to a short repeating pattern of bases in a DNA fragment (e.g., a cell-free DNA fragment). A sequence motif can occur at the end of a fragment and therefore can be part of or include the end sequence. A "end motif" can refer to a sequence motif for end sequences that occur preferentially at the end of a DNA fragment, potentially for a particular type of tissue. A end motif can also occur just before or just after the end of a fragment, thereby still corresponding to an end sequence. A nuclease can have a particular cleavage preference for a particular end motif, as well as a second, most preferred cleavage preference for a second end motif.
[0014] The term "overhang length" between DNA strands can refer to a value that can be estimated by comparing the mismatch (e.g., mismatch index value) of plasma DNA overall or within a specific fragment size range between a reference sample (e.g., normal cells) and a differentially regulated nuclease sample (e.g., tumor cells). In some cases, the overhang length varies based on the specific DNA fragment size range (e.g., 130-160 bp, 200-300 bp) selected to characterize the biological sample.
[0015] In some embodiments, the length of an overhang in a DNA strand is a classification value that characterizes the length of an overhang between double-stranded DNA. For example, a "long" overhang can include a DNA strand overhang with a size of 5 nt, 6 nt, 7 nt, 8 nt, 10 nt, 15 nt, 20 nt, 30 nt, 40 nt, 50 nt, 100 nt, and a size greater than 100 nt. A "short" overhang can include a DNA strand overhang with a size of 0 nt, 1 nt, 2 nt, 3 nt, 4 nt, 5 nt. Additionally or alternatively, a specified length of an overhang in a DNA strand can be estimated based on the percentage of molecules with an overhang size exceeding a certain threshold. For example, the presence of a "long" overhang in plasma DNA can be expressed as the percentage of molecules greater than 5 nt, 6 nt, 7 nt, 8 nt, 10 nt, 15 nt, 20 nt, 30 nt, 40 nt, 50 nt, 100 nt, or a combination thereof.
[0016] A "calibration sample" may correspond to a biological sample in which the fractional concentration of clinically relevant DNA (e.g., tissue-specific DNA fraction) is known or determined via a calibration method using tissue-specific alleles, such as, for example, transplantation in a pregnant subject, where alleles present in the donor's genome but absent from the recipient's genome can be used as markers for the transplanted organ. As another example, a calibration sample may correspond to a sample in which terminal motifs can be determined. A calibration sample may be used for both purposes.
[0017] A "calibration data point" includes a "calibration value" and a measured or known characteristic of a sample or subject, such as an age- or tissue-specific fraction (e.g., fetal or tumor). A calibration value can be a relative abundance determined for a calibration sample with known characteristics. A calibration data point can include a calibration value (e.g., a random end value, also called an overhang index) and a known (measured) characteristic. Calibration data points can be defined in various ways, for example, as discrete points or as a calibration function (also called a calibration curve or calibration surface). A calibration function can be derived from additional mathematical transformations of the calibration data points. A calibration function can be linear or nonlinear.
[0018] A "site" (also called a "genomic site") corresponds to a single site, which may be a single base position, or a group of correlated base positions, e.g., a CpG site, or a larger group of correlated base positions. A "locus" may correspond to a region that includes multiple sites. A locus may contain only one site, which would make the locus equivalent to the site in that context.
[0019] A "separation value" corresponds to the difference or ratio involving two values, for example, two fractional contributions or two methylation levels. A separation value can be a simple difference or ratio. As an example, the direct ratio of x / y is a separation value, as is x / (x+y). A separation value can include other factors, for example, multiplicative factors. As another example, the difference or ratio of a function of the values can be used, for example, the difference or ratio of the natural logarithms (ln) of two values. A separation value can include differences and ratios.
[0020] As used herein, the term "classification" refers to any number or other feature associated with a particular property of a sample. For example, a "+" sign (or the word "positive") may indicate that the sample is classified as having a deletion or an amplification. Classifications can be binary (e.g., positive or negative) or have more levels of classification (e.g., a scale of 1 to 10 or 0 to 1). The terms "cutoff" and "threshold" refer to a predetermined number used in an operation. For example, a cutoff size may refer to a size above which a fragment is excluded. A threshold may be a value above or below which a particular classification is applied. Either of these terms can be used in either of these contexts.
[0021] As used herein, the term "parameter" refers to a numerical value that characterizes a quantitative data set and / or a numerical relationship between quantitative data sets. For example, a ratio (or a function of the ratio) between a first amount of a first nucleic acid sequence and a second amount of a second nucleic acid sequence is a parameter.
[0022] The terms "cutoff" and "threshold" refer to predetermined numbers used in an operation. For example, a cutoff size may refer to a size above which a fragment is excluded. A threshold may be a value above or below which a particular classification is applied. Either of these terms can be used in either of these contexts. A cutoff or threshold may be a "reference value" or may be derived from a reference value that represents a particular classification or distinguishes between two or more classifications. Such reference values can be determined in various ways, as will be understood by those skilled in the art. For example, a metric can be determined for two different cohorts of subjects known to have different classifications, and a reference value can be selected to represent one classification (e.g., a mean value) or a value between two clusters of the metric (e.g., selected to obtain a desired sensitivity and specificity). As another example, a reference value can be determined based on statistical analysis or sample simulation. Particular values for cutoffs, thresholds, references, etc. can be determined based on the desired accuracy (e.g., sensitivity and specificity).
[0023] "Pregnancy-related disorder" includes any disorder characterized by abnormal relative expression levels of genes in maternal and / or fetal tissues and / or by abnormal clinical features in the mother and / or fetus. These disorders include preeclampsia (Kaartokallio et al. Sci Rep. 2015;5:14107, Medina-Bastidas et al. Int J Mol Sci. 2020;21:3597), intrauterine growth restriction (Faxen et al. Am J Perinatol. 1998;15:9-13, Medina-Bastidas et al. Int J Mol Sci. 2020;21:3597), invasive placentation, preterm birth (Enquobahrie et al. BMC Pregnancy Childbirth. 2009;9:56), hemolytic disease of the newborn, placental insufficiency (Kelly et al. Endocrinology. 2017;158:743-755), and hydrops fetalis (Magor et al. al. Blood. 2015;125:2405-17), fetal malformations (Slonim et al. Proc Natl Acad Sci USA. 2009;106:9425-9), HELLP syndrome (Dijk et al. J Clin Invest. 2012;122:4003-4011), systemic lupus erythematosus (Hong et al. J Exp Med. 2019;216:1154-1169), and other maternal immune disorders.
[0024] "Level of pathology" (or level of injury or level of condition) may refer to the amount, degree, or severity of pathology associated with an organism. One example is cellular damage in the expression of nucleases. Another example of pathology is rejection of a transplanted organ. Other exemplary pathologies can include autoimmune attacks (e.g., lupus nephritis or multiple sclerosis, which damage the kidneys), inflammatory diseases (e.g., hepatitis), fibrotic processes (e.g., cirrhosis), fatty infiltration (e.g., fatty liver disease), degenerative processes (e.g., Alzheimer's disease), and ischemic tissue damage (e.g., myocardial infarction or stroke). A subject's healthy state can be considered a classification without pathology. The pathology can be cancer.
[0025] The term "level of cancer" may refer to whether cancer is present (i.e., present or absent), the stage of cancer, tumor size, whether there is metastasis, total tumor burden in the body, the response of cancer to treatment, and / or other measures of cancer severity (e.g., cancer recurrence). Cancer levels may be numbers or other indicators, such as symbols, alphabetic characters, and colors. A level may be zero. Cancer levels may also include premalignant or precancerous conditions (states). Cancer levels may be used in a variety of ways. For example, screening can confirm the presence of cancer in a person not previously known to have cancer. Evaluation can follow up on a person diagnosed with cancer to monitor the progression of cancer over time, study the effectiveness of a therapy, or determine a prognosis. In one embodiment, prognosis can be expressed as the likelihood that a patient will die from cancer, or the likelihood that the cancer will progress after a certain period or time, or the likelihood or extent that the cancer will metastasize. Detection can mean "screening" or determining whether a person with suggestive features of cancer (eg, symptoms or other positive test) has cancer.
[0026] The abbreviation "bp" refers to base pairs. In some cases, "bp" may be used to indicate the length of a DNA fragment even if the DNA fragment is single-stranded and does not contain base pairs. In the context of single-stranded DNA, "bp" may be interpreted as providing the length of the strand in nucleotides.
[0027] The abbreviation "nt" refers to a nucleotide. In some cases, "nt" may be used to indicate the length of a single-stranded DNA in base units. "nt" may also be used to indicate a relative position, such as upstream or downstream of the analyzed locus. In the case of double-stranded DNA, "nt" may still refer to the length of a single strand rather than the total number of nucleotides in the two strands, unless the context clearly dictates otherwise. In some contexts related to technical conceptualization, data presentation, processing, and analysis, "nt" and "bp" may be used interchangeably.
[0028] The term "ragged end" can refer to a sticky end of DNA, a DNA overhang, a strand overhang, or when double-stranded DNA contains a strand of DNA that is not hybridized to the other strand of DNA. A "ragged end value" is a measure of the extent of ragged ends. The ragged end value can be proportional to the length of one strand that overhangs the second strand of double-stranded DNA. The ragged end value of multiple DNA molecules can include consideration of blunt ends between the DNA molecules.
[0029] In some instances, the ragged end value can provide a collective measure of the strands that overhang one another in a plurality of cell-free DNA molecules. The collective measure of raggedness can be determined based on the estimated lengths of the overhangs in a plurality of cell-free DNA molecules, for example, the mean, median, or other collective measure of the individual measurements for each of the cell-free DNA molecules. In some instances, the collective measure of raggedness is determined for a particular fragment size range (e.g., 130-160 bp, 200-300 bp).
[0030] The term " size ratio " can refer to the amount of cell-free DNA molecules within a specific fragment size range.Size ratio can be proportional to the amount of cell-free DNA molecules within a specific fragment size range, normalized by the amount of cell-free DNA molecules within another specific fragment size range.When another specific fragment size range relates to all size ranges, the term " size frequency " can be used.
[0031] The term "alignment" and related terms can refer to matching a sequence to a reference sequence. The reference sequence can be a reference genome (e.g., the human genome) or the sequence of a specific molecule. Such a reference sequence can include at least 100 kb, 1 Mb, 10 Mb, 50 Mb, 100 Mb, and more. Such alignment methods cannot be performed manually, but rather by specialized computer software. Alignments can include long and numerous sequences (e.g., at least 1,000, 10,000, 100,000, 1,000,000, 10,000,000, or 100,000,000 sequences). Additionally, alignments can include variability within the sequences themselves or errors in sequence reads. Thus, alignments with such variability or errors may not require exact matching with the reference sequence.
[0032] The term "real-time" may refer to a computing operation or process that is completed within a certain time constraint, which may be one minute, one hour, one day, or seven days.
[0033] The term "subsequence" can refer to a series of bases that is less than the complete sequence corresponding to a nucleic acid molecule. For example, if the complete sequence of a nucleic acid molecule contains five or more bases, a subsequence can contain one, two, three, or four bases. In some embodiments, a subsequence can refer to a series of bases that form a unit, where the unit is repeated multiple times in tandem. Examples include a 3 nt unit or subsequence repeated at a locus associated with a trinucleotide repeat disorder, a 1 nt to 6 nt unit or subsequence repeated 5 to 50 times as a microsatellite, or a 10 nt to 60 nt unit or subsequence repeated 5 to 50 times as a microsatellite or in other genetic elements such as Alu repeats.
[0034] "Clinically relevant DNA" refers to DNA of a particular tissue source that is measured, for example, to determine the fractional concentration of such DNA or to phenotype a sample (e.g., plasma). Examples of clinically relevant DNA are fetal DNA in maternal plasma, or tumor DNA in a patient's plasma, or other sample containing cell-free DNA. Another example includes measuring the amount of graft-associated DNA in the plasma, serum, or urine of a transplant patient. Further examples include measuring the fractional concentrations of hematopoietic and non-hematopoietic DNA in a subject's plasma, or the fractional concentration of liver DNA fragments (or other tissues) in a sample, or the fractional concentration of brain DNA fragments in cerebrospinal fluid.
[0035] The term "simultaneous analysis" can refer to the use of more than one fragmentomics feature. Using only a 5'-end motif on one end of a nucleic acid molecule, or using only a ragged end pattern (e.g., a 5'-end overhang), would not be simultaneous analysis. However, using a combination of ragged end patterns from both ends of a molecule, a combination of one ragged end pattern and a sequence end motif on one end, or a combination of ragged end patterns and a sequence end motif on both ends, would be part of simultaneous analysis.
[0036] The term "about" or "approximately" can mean within an acceptable error range for a particular value, i.e., the limits of the measurement system, as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined. For example, "about" can mean within 1 or more than 1 standard deviation, according to convention within the art. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term "about" or "approximately" can mean within an order of magnitude, within 5-fold, or more preferably within 2-fold of a value. When a particular value is described in this application and claims, unless otherwise stated, the term "about" should be assumed to mean within an acceptable error range of the particular value. The term "about" can have the meaning commonly understood by one of ordinary skill in the art. The term "about" can refer to ±10%. The term "about" can refer to ±5%.
[0037] Where a range of values is provided, unless the context clearly dictates otherwise, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limit of that range is also specifically disclosed. Each smaller range between any stated or intervening value in a stated range and any other stated or intervening value in that stated range is encompassed within an embodiment of the disclosure. The upper and lower limits of these smaller ranges may independently be included or excluded within the range, and each range where either, neither, or both limits are included within the smaller range is also encompassed within the disclosure, subject to any specifically excluded limits in the stated range. When a stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included within the disclosure.
[0038] Standard abbreviations may be used, such as bp: base pairs, kb: kilobase, pi: picoliters, s or sec: seconds, min: minutes, h or hr: hours, aa: amino acids, nt: nucleotides, etc.
[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the present disclosure, some potential and exemplary methods and materials may be described here. DETAILED DESCRIPTION OF THE INVENTION
[0040] Cell-free DNA (cfDNA) molecules are fragmented non-randomly, and the fragmentation pattern of cfDNA molecules contains a wealth of molecular information. For example, the characteristic size profile of cfDNA shows a modal frequency at about 166 bp, and smaller molecules form a series of peaks with a periodicity of 10 bp (Lo et al. Sci Transl Med. 2010; 2: 61ra91). Such size patterns of plasma DNA fragments suggest the existence of both internucleosomal and intranucleosomal cleavage during the release of DNA molecules into the blood circulation during cell death and / or apoptosis. Furthermore, our group previously reported that a subset of genomic locations was found to be preferentially cleaved during the generation of plasma DNA molecules (Chan et al. Proc Natl Acad Sci USA. 2016;113:E8159-E8168; Jiang et al. Proc Natl Acad Sci USA. 2018;115:E10925-E10933), and such preferred cleavage could reflect the tissue of origin of the cfDNA (Jiang et al. Proc Natl Acad Sci USA. 2018;115:E10925-E10933; Sun et al. Proc Natl Acad Sci USA. 2018;115:E5106-E5114). Furthermore, the inventors have shown that various nucleases associate with cell-free DNA molecules with characteristic terminal signatures (i.e., 5'-end motifs and 5'-overhanging ragged ends) (Serpas et al. Proc Natl Acad Sci USA. 2019; 116:641-649, Han et al. Am J Hum Genet. 2020; 106:202-214, Ding et al. Clin Chem. 2022; 68:917-926). The 5'-end motif represents the sequence context of the 5'-end of the cfDNA fragment. The 5'-overhanging ragged ends represent the 5'-overhanging single-stranded DNA in the cfDNA molecule.Recently, cell-free DNA end signatures have shown promising results in functioning as liquid biopsy biomarkers (Jiang et al. Cancer Discov. 2020; 10: 664-673, Jiang et al. Genome Res. 2020; 30: 1144-1153). Ragged ends are also described in US2020 / 0056245A1 and US2022 / 0177971A1, the entire contents of both of which are incorporated herein by reference for all purposes.
[0041] In contrast to the widely studied 5'-end motif and 5'-overhanging ragged end, the actual 3'-end motif and 3'-overhanging ragged end have not been properly investigated, mainly due to artificial modifications that occur during sequencing library preparation. Typical library preparation methods include an end-repair step. During end-repair, the 3'-overhanging ragged end is removed, and the 3'-retreat end is extended using the opposite 5'-overhanging ragged end as a DNA template. Thus, the original 3' end is modified, resulting in the alteration of the nucleotide information proximal to the 3'-end motif and the loss of the 3'-overhanging ragged end. Furthermore, the 3'-overhanging ragged end is removed to form a blunt end. Due to such an end-repair step, the blunt-end information inferred from typical library preparation methods is unreliable.
[0042] Recently, one group developed an NGS library preparation method, named the XACTLY assay, to study ragged ends through a sequence adapter ligation-based approach (Harkins et al. Nucleic Acids Res. 2020;48:e47). For example, Harkins Kincaid et al. directly ligated Y-shaped adapters containing 7-nt barcodes (i.e., unique end identifiers (UEIs) indicating individual end types and lengths) to the original DNA template without an end-repair step (Harkins Kincaid et al. Nucleic Acids Res. 2020;48:e47). The ligated products were subjected to short-read sequencing (i.e., on an Illumina sequencing platform). It was believed that the natural ends of DNA molecules could be predicted according to the sequence information of the 7-nt barcodes ligated to the reads. However, as reported in the study by Harkins et al., the 5'UEI (i.e., Illumina P5 adapter) had much lower fidelity than the 3'UEI (i.e., Illumina P7 adapter) (Harkins Kincaid et al. Nucleic Acids Res. 2020;48:e47). Such inaccuracy associated with P5 adapters is due to the fact that the first ligation of the P5 adapter to the template DNA occurs regardless of whether the overhanging end of the adapter properly matches the overhanging end on the template DNA, whereas ligation of the P7 adapter can occur only if the first ligation event is correct. In other words, P5 adapter ligation occurs even when gaps (i.e., regions where one strand does not have complementary nucleotides on the other strand) and flaps (i.e., regions where nucleotides of one adapter do not hybridize to the strand) are present in the hybridization region between the template DNA and the adapter sequence. In the sequencing library prepared by the XACTLY assay, double-stranded DNA is denatured into two single-stranded DNA molecules, the 5' and 3' ends of which are tagged with P5 and P7, respectively.Therefore, the XACTLY assay has inherent limitations: 1. Using the XACTLY assay, at least one end cannot be accurately analyzed. 2. It is not possible to effectively analyze both strands of a DNA molecule simultaneously. 3. The actual length of the DNA molecule cannot be measured.
[0043] In this disclosure, we have developed a new approach for simultaneously detecting the natural fragmentomic features of cfDNA molecules with high fidelity, including fragment size, terminal motifs, and ragged ends. In one embodiment, DNA ligase can be used to appropriately ligate double-stranded cfDNA fragments with a pair of hairpin adapters to form circularized DNA molecules, depending on the terminal configuration of the double-stranded cfDNA molecules. Such hairpin adapters contain molecular barcodes and carry ragged or blunt ends of various lengths. Various cfDNA fragments are ligated with hairpin adapters containing various molecular barcodes corresponding to the ragged end length (e.g., 1-50 nt) and ragged type (blunt end, 5'-overhanging ragged end, 3'-overhanging ragged end, and combinations thereof). The ligation products can be treated with enzymes to remove incomplete circular DNA molecules, thereby enriching for the desired circular DNA molecules generated by hairpin adapter-mediated DNA ligation (i.e., a negative selection step).
[0044] The product enriched for circular DNA molecules can be further subjected to direct enrichment of circular DNA molecules (i.e., a positive selection step), such as single-molecule real-time sequencing (e.g., Pacific Biosciences) and rolling circle amplification, to minimize the impact of incorrect ligation. Because only complete circular DNA molecules can be sequenced multiple times and generate subreads during single-molecule real-time sequencing, selecting reads with three or more subreads allows for the exclusion of incomplete circular DNA molecules. Similarly, only complete circular DNA molecules can be amplified via rolling circle amplification. In some embodiments, the enzymes include, but are not limited to, exonuclease I, exonuclease II, exonuclease III, exonuclease IV, exonuclease V, exonuclease VI, exonuclease VII, or exonuclease VIII. In yet another embodiment, negative and positive selection steps can be performed alone or in combination. After sequencing, the naturally ragged ends of single cfDNA fragments can be estimated by analyzing the barcode sequences. Once the ragged ends of each end of the cfDNA molecule are determined, the entire fragmentomics features from the cfDNA molecule can be accurately analyzed at 1 nt resolution, including, but not limited to, 5'-overhanging ragged ends, 3'-overhanging ragged ends, 5'-recessed ragged ends, 3'-recessed ragged ends, overhanging ragged end motifs, recessed ragged end motifs, fragment end genomic coordinates, fragment sizes, methylation-related cfDNA fragmentomics features, and combinations thereof. In one embodiment, the length difference between Watson stands and Crick stands can be used as another type of fragmentomics feature.
[0045] In some embodiments, a terminal motif can be defined by one or more nucleotides spanning positions at or near the ends of a molecule. A molecule can have four ends. A terminal motif can be defined by one or more nucleotides in a reference genome surrounding the genomic locus to which the ends of the fragment are aligned. In other embodiments, ragged ends can be defined by overhanging single-stranded DNA at the ends of a DNA fragment. The ragged ends can be separated into different groups according to the length and stranding of the overhanging single-stranded DNA. In yet other embodiments, different fragmentomics features from a single DNA fragment can be combined. In one embodiment, the combined fragmentomics features can be used for detecting or monitoring cancer or other diseases. In another embodiment, the combined fragmentomics features can be used for non-invasive prenatal testing. Terminal motifs are described in US2021 / 0238668A1, the contents of which are incorporated herein by reference.
[0046] I. Principles of simultaneous analysis of single molecule end-to-end formats Figure 1 shows a schematic diagram of simultaneous single-molecule end-to-end analysis on a single-molecule real-time sequencing platform (e.g., the Pacific Biosciences (PacBio) platform). Stage 104 shows various cfDNA molecules containing various ragged or blunt ends. Exemplary fragment 108 has a blunt end on the left and a 3-nt 5'-overhanging ragged end on the right.
[0047] Stage 112 shows various hairpin adapters in a hairpin adapter pool. The hairpin adapter pool includes adapters with blunt ends and adapters with ragged ends (also called overhangs). Each ragged-end hairpin adapter has a protruding single-stranded end of varying length (indicated by the number of "N"s in the overhang 116). Barcode sequences synthesized with the hairpin adapters compatible with the PacBio sequencing platform can be used to indicate the ragged end type (e.g., 5' or 3' protruding end) and ragged end length (indicated by rectangles 120 and 124 filled with different patterns).
[0048] At stage 128, the cfDNA molecule is ligated with hairpin adapters. Fragment 108 has hairpin adapter 132 ligated to its blunt end (left side). Fragment 108 has hairpin adapter 136 ligated to its 3 nt 5' overhang ragged end (right side). Successful ligation results in molecule 140.
[0049] Other molecules may result from ligation. Molecule 144 represents a fragment with a hairpin adaptor ligated to only one end. Molecule 148 represents a fragment without a hairpin adaptor. Molecule 152 has a hairpin adaptor correctly ligated to the blunt end. However, molecule 152 has a hairpin adaptor incorrectly ligated to the 5'-overhanging end, resulting in a gap between the cfDNA fragment and the hairpin adaptor. Molecule 156 has a hairpin adaptor correctly ligated to the blunt end. However, molecule 156 has a hairpin adaptor incorrectly ligated to the 5'-overhanging end, resulting in the hairpin adaptor generating a flap where the nucleotides of the hairpin adaptor do not hybridize to the original cfDNA fragment.
[0050] At stage 160, the adaptor-ligated molecules may be treated with an enzyme(s) capable of digesting imperfect circular DNA molecules (e.g., molecule 152 and molecule 156). Enzymatic digestion of imperfect adaptor-ligated molecules may be referred to as negative selection, as incorrectly ligated molecules are selected for and removed.
[0051] At stage 164, the enzyme-treated ligation products can be sequenced on the PacBio platform. Only if cfDNA fragments with both ends are properly ligated with hairpin adapters corresponding to the naturally ragged ends (e.g., via rolling circle amplification) to form complete circular DNA can such circular DNA products be sequenced to generate multiple subreads for each strand. Because only correctly ligated molecules are selected and further analyzed, the amplification and / or sequencing of only molecules with complete adapters ligated can be referred to as positive selection.
[0052] In stage 168, the sequence is analyzed. After sequencing, the barcode sequence information of both ends can be read to estimate the presence of ragged ends and / or blunt ends, and if ragged ends exist, their type and length. Based on the estimated ends, the 5'-end motif, 3'-end motif, and / or the size of each strand of the cfDNA fragment can be further detected. The natural fragmentomics characteristics of the original cfDNA molecule can be evaluated.
[0053] 2A and 2B show schematic diagrams of simultaneous single-molecule end-to-end analysis on a next-generation sequencing (NGS) platform (e.g., an Illumina platform). Circular cfDNA molecules can be prepared according to embodiments of the present disclosure using modified hairpin adaptors. Stage 104 can be repeated in FIG. 2A. Stage 112 can include modified hairpin adaptors containing restriction enzyme cleavage sites. For example, cleavage sites 204 and 208 can be included in the hairpin adaptors.
[0054] In Figure 2B, similar to stage 128, DNA fragments are ligated with hairpin adapters. Similar to stage 160, the adapter-ligated molecules are treated with enzyme(s) capable of digesting incomplete circular DNA molecules (i.e., negative selection). Similar to stage 160, the enzyme-treated ligation products are amplified by rolling circle amplification (i.e., positive selection). Only cfDNA fragments with both ends properly ligated to hairpin adapters corresponding to the naturally ragged / blunt ends are amplified.
[0055] At stage 250, the rolling amplification products are treated with a specific restriction enzyme to cleave at the cleavage sites of the hairpin adapters, thus cleaving the large DNA molecules generated via rolling PCR into smaller DNA molecules suitable for Illumina sequencing or other similar sequencing.
[0056] Sequencing adaptors are ligated to the cleaved small DNA molecules at stage 254. The sequencing adaptors are configured for Illumina sequencing.
[0057] The analysis of FIG. 2B may be similar to stage 168 of FIG.
[0058] A. Exemplary Positive Selection Methods FIG. 3 is a flowchart of an exemplary process 300 for analyzing a biological sample. Process 300 can determine whether ragged ends are present at both ends of a cfDNA molecule, whether a 5' or 3' end is overhanging, the length of the overhang, and / or the sequence of the overhang. A strand that overhangs another strand can be understood to be an overhang. In some embodiments, one or more process blocks of FIG. 3 can be performed by a system, including system 2400. The biological sample can include multiple nucleic acid molecules. The nucleic acid molecules can be cell-free and double-stranded, having a first strand and a second strand.
[0059] In block 302, for each nucleic acid molecule of the plurality of nucleic acid molecules, a first hairpin adaptor is ligated to a first strand of the nucleic acid molecule and a second strand of the nucleic acid molecule at a first end of the nucleic acid molecule. The first hairpin adaptor may include a first sequence identifier. The first sequence identifier may specify a first length of zero or more nucleotides at the first end of the first hairpin adaptor that does not have a complementary portion at the second end of the first hairpin adaptor. The length of nucleotides that does not have a complementary portion at the hairpin corresponds to the length of the ragged end of the nucleic acid molecule. Zero nucleotides indicates a blunt end, and the ends of the hairpin adaptor are complementary. For example, the first sequence identifier may code for a length of 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. The first sequence identifier may code for whether the 3' strand or the 5' strand of the nucleic acid molecule to which the first hairpin ligates overhangs the other. In some embodiments, the first sequence identifier may encode a subsequence of zero or more nucleotides at the first end of the first hairpin adaptor that does not have a complementary portion at the second end of the first hairpin adaptor.
[0060] The first hairpin adaptor can include hairpin adaptors 132 and 136 in Figure 1. The first sequence identifier can be the nucleotides represented by rectangles 120 and 124. A length of zero or more nucleotides at the first end of the first hairpin adaptor that does not have a complementary portion at the second end of the first hairpin adaptor can include overhang 116.
[0061] In block 304, for each nucleic acid molecule of the plurality of nucleic acid molecules, a second hairpin adaptor is ligated to the first strand and the second strand at a second end of the nucleic acid molecule. The second hairpin adaptor may include a second sequence identifier. The second sequence identifier may specify a second length of zero or more nucleotides at the first end of the second hairpin adaptor that does not have a complementary portion at the second end of the second hairpin adaptor. The second sequence identifier may have similar characteristics to the first sequence identifier. The first sequence identifier and the second sequence identifier may use similar encoding. A particular predetermined subsequence in the sequence identifier may correspond to various numbers. With four nucleotides (A, T, G, C), the length may be expressed in numerical base four. A plurality of ligated nucleic acid molecules is generated after ligation.
[0062] In some examples, negative selection can be performed. After ligating a plurality of first hairpin adaptors and a plurality of second hairpin adaptors, an exonuclease can be added to the plurality of ligated nucleic acid molecules to remove an incorrectly ligated subset of the plurality of ligated nucleic acid molecules. For each nucleic acid molecule of the incorrectly ligated subset, the respective nucleic acid molecule is either not fully hybridized to the respective first hairpin adaptor or the respective second hairpin adaptor (e.g., a "gap" is present), or the respective first hairpin adaptor or the respective second hairpin adaptor is not fully hybridized to the respective nucleic acid molecule (e.g., a "flap" is present). The negative selection can be similar to stage 160 of FIG. 1, in which molecules 144, 148, 152, and 156 are removed.
[0063] In block 306, rolling circle amplification may be performed on a first subset of the plurality of ligated nucleic acid molecules to form a plurality of concatemers. The first subset may not contain any of the same nucleic acid molecules as the incorrectly ligated subset. Each nucleic acid molecule of the first subset may be ligated to a first hairpin adaptor of each of the plurality of first hairpin adaptors and a second hairpin adaptor of each of the plurality of second hairpin adaptors. Each nucleic acid molecule of the first subset may be correctly ligated to a hairpin adaptor without a gap or flap, similar to molecule 140 in FIG. 1. Each nucleotide of a strand of the first set of nucleic acid molecules may hybridize to a complementary nucleotide on the other strand.
[0064] Each nucleic acid molecule in the first portion of the first subset can have, at its first end, a respective first strand that overhangs a respective second strand. The first strand can be the 5' strand or the 3' strand of the first end. In some examples, each nucleic acid molecule in the second portion of the first subset can have, at its first end, a respective first strand that is equivalent to a respective second strand. In some examples, each nucleic acid molecule in the second portion of the first subset can have, at its first end, a respective second strand that overhangs a respective first strand. Each first strand can be the 5' strand. Each second strand can be the 3' strand.
[0065] The first subset may include portions corresponding to various combinations of ragged end properties: DNA molecules containing 5'-overhanging ragged ends and 3'-overhanging ragged ends (5-3), 5'-overhanging ragged ends and 5'-overhanging ragged ends (5-5), 3'-overhanging ragged ends and 3'-overhanging ragged ends (3-3), 5'-overhanging ragged ends and blunt ends (5-B), 3'-overhanging ragged ends and blunt ends (3-B), and blunt and blunt ends (BB).
[0066] In block 308, each concatemer of the plurality of concatemers is sequenced to identify a respective first sequence identifier and a respective second sequence identifier. The first sequence identifier and the second sequence identifier may each include a subsequence of nucleotides indicating that consecutive nucleotides are part of the identifier. Sequencing may be via single molecule, real-time sequencing, next-generation sequencing, or any suitable sequencing technique. Sequencing may be performed simultaneously with performing rolling circle amplification.
[0067] The length of an overhang present at a first end of a nucleic acid molecule of a first subset of the plurality of ligated nucleic acid molecules can be determined using a first sequence identifier. The first sequence identifier can include a subsequence corresponding to the length of the overhang. Additionally, the first sequence identifier can include a subsequence indicating whether the overhang is on the current strand or the complementary strand.
[0068] The length of the overhang present at the second end of the nucleic acid molecule of the first subset of the plurality of ligated nucleic acid molecules is determined using a second sequence identifier, which may be used in a similar manner as the first sequence identifier.
[0069] In some examples, a first sequence end motif of an overhang present at a first end of a nucleic acid molecule of a first subset of the plurality of ligated nucleic acid molecules can be determined using the sequence of a first sequence identifier. The first sequence identifier can indicate which strand of the end is overhanging, and an appropriate subsequence can be associated with the overhang. Additionally, because the first sequence identifier indicates the length of the overhang, the entire sequence of the overhang can be determined. In some embodiments, the entire sequence of the overhang need not be determined, and instead, a terminal motif (2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides) can be determined. In some examples, a second sequence end motif of an overhang present at a second end of a nucleic acid molecule of a first subset of the plurality of ligated nucleic acid molecules can be determined using the sequence of a second sequence identifier.
[0070] In some examples, whether the 5' strand or the 3' strand overhangs the other can be determined for each nucleic acid molecule having an overhang at a first end of the first subset using the sequence of each first sequence identifier. In some examples, whether the 5' strand or the 3' strand overhangs the other can be determined for each nucleic acid molecule having an overhang at a second end of the first subset using the sequence of each second sequence identifier.
[0071] In some examples, each first hairpin adaptor of the plurality of first hairpin adaptors may comprise a first cleavage site. Each second hairpin adaptor of the plurality of second hairpin adaptors may comprise a second cleavage site. The process may include cleaving each concatemer of the plurality of concatemers at the respective first cleavage site and the respective second cleavage site.
[0072] Process 300 can be used to determine length or terminal motifs in other processes disclosed herein. In some examples, each nucleic acid molecule of the plurality of molecules has a size greater than a first cutoff size. In some examples, each nucleic acid molecule of the plurality of molecules has a size less than a second cutoff size. The first cutoff size and the second cutoff size can independently be 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, or 350. The size of each nucleic acid molecule can be determined by aligning subsequences corresponding to the ends of each nucleic acid molecule with a reference genome.
[0073] The condition can be cancer (e.g., but not limited to, HCC and colorectal cancer [CRC]), an autoimmune disease (e.g., systemic lupus erythematosus), a pregnancy-related disorder, or any condition described herein. The reference value can be determined from one or more subjects with a particular level of the condition, or from one or more healthy subjects.
[0074] In some instances, the level of the condition is not determined. Instead, the fractional concentration of clinically relevant DNA can be determined using the comparison. The reference value can be determined from one or more subjects whose fractional concentrations of clinically relevant DNA are known. The reference value can be a calibration value determined using a calibration sample.
[0075] In some examples, reads corresponding to multiple nucleic acid molecules can be enriched for clinically relevant DNA. For example, a biological sample can be obtained from a female subject carrying a fetus. The method can further include selecting reads corresponding to a subset of nucleic acid molecules having a 5' or 3' strand overhanging the other end. The method can include analyzing the subset of nucleic acid molecules for a fetal characteristic. For example, the characteristic can be the presence of an abnormality (e.g., mutation, aneuploidy) in the fetal genome. As another example, reads can be enriched for a maternal sample by selecting reads with a blunt end at one end. Other clinically relevant DNA can be enriched by analyzing the concentration of such DNA among various ragged end patterns. A ragged end pattern with a higher concentration of clinically relevant DNA can be selected to result in an enriched dataset. The end pattern can include end patterns at the two ends of any given fragment.
[0076] Process 300 may include additional implementations, such as any single implementation or any combination of implementations described herein, and / or in conjunction with one or more other processes described elsewhere herein.
[0077] 3 illustrates example blocks of process 300, in some implementations, process 300 may include additional, fewer, different, or differently arranged blocks compared to the blocks illustrated in FIG 3. Additionally or alternatively, two or more of the blocks of process 300 may be performed in parallel.
[0078] B. Exemplary Negative Selection Methods 4 is a flowchart of an exemplary process 400 for analyzing a biological sample. Process 400 can determine whether ragged ends are present at both ends of a cfDNA molecule, whether a 5' or 3' end is overhanging, the length of the overhang, and / or the sequence of the overhang. A strand that overhangs another strand can be understood to be an overhang. In some embodiments, one or more process blocks of FIG. 4 can be performed by system 2400.
[0079] In block 402, for each nucleic acid molecule of the plurality of nucleic acid molecules, a first hairpin adaptor is ligated to a first strand of the nucleic acid molecule and to a second strand of the nucleic acid molecule at a first end of the nucleic acid molecule. Block 402 may be performed in a manner similar to block 302.
[0080] In block 404, for each nucleic acid molecule of the plurality of nucleic acid molecules, a second hairpin adaptor is ligated to the first strand and the second strand at a second end of the nucleic acid molecule. Block 404 may be performed in a manner similar to block 304.
[0081] At block 406, an exonuclease is added to the plurality of ligated nucleic acid molecules to remove a first subset of the plurality of ligated nucleic acid molecules, wherein for each nucleic acid molecule of the first subset, either the respective nucleic acid molecule is not fully hybridized to either a respective first hairpin adaptor or a respective second hairpin adaptor, or each first hairpin adaptor or each second hairpin adaptor is not fully hybridized to the respective nucleic acid molecule.
[0082] In block 408, each ligated nucleic acid molecule of a second subset of the plurality of ligated nucleic acid molecules may be sequenced to identify the respective first sequence identifier and the respective second sequence identifier. The second subset is the ligated nucleic acid molecules remaining in the biological sample after removing the first subset. Sequencing may be performed by next-generation sequencing, single-molecule real-time sequencing, or any sequencing technique described herein.
[0083] The length of the overhang present at the first end of the nucleic acid molecule of the second subset of the plurality of ligated nucleic acid molecules can be determined using the first sequence identifier. The first sequence identifier can include a subsequence corresponding to the length of the overhang. Additionally, the first sequence identifier can include a subsequence indicating whether the overhang is on the current strand or the complementary strand.
[0084] The length of the overhang present at the second end of the nucleic acid molecules of the second subset of the plurality of ligated nucleic acid molecules is determined using a second sequence identifier, which may be used in a similar manner as the first sequence identifier.
[0085] In some examples, the first sequence end motif of the overhang present at the first end of the nucleic acid molecules of the second subset of the plurality of ligated nucleic acid molecules can be determined using the sequence of the first sequence identifier. The first sequence identifier can indicate which strand of the end is overhanging, and an appropriate subsequence can be associated with the overhang. Additionally, because the first sequence identifier indicates the length of the overhang, the entire sequence of the overhang can be determined. In some embodiments, the entire sequence of the overhang need not be determined, and instead, a terminal motif (2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides) can be determined. In some examples, the second sequence end motif of the overhang present at the second end of the nucleic acid molecules of the second subset of the plurality of ligated nucleic acid molecules can be determined using the sequence of the second sequence identifier.
[0086] In some examples, whether the 5' strand or the 3' strand overhangs the other can be determined for each nucleic acid molecule having an overhang at its first end in the second subset using the sequence of each first sequence identifier. In some examples, whether the 5' strand or the 3' strand overhangs the other can be determined for each nucleic acid molecule having an overhang at its second end in the second subset using the sequence of each second sequence identifier.
[0087] In some examples, each first hairpin adaptor of the plurality of first hairpin adaptors may comprise a first cleavage site. Each second hairpin adaptor of the plurality of second hairpin adaptors may comprise a second cleavage site. The process may include cleaving each concatemer of the plurality of concatemers at the respective first cleavage site and the respective second cleavage site.
[0088] In some examples, the sample may be enriched for clinically relevant DNA, as described in process 300.
[0089] Process 400 may include additional implementations, such as any single implementation or any combination of implementations described herein, and / or in conjunction with one or more other processes (including process 300) described elsewhere herein.
[0090] 4 illustrates example blocks of process 400, in some implementations, process 400 may include additional, fewer, different, or differently arranged blocks compared to the blocks illustrated in FIG 4. Additionally or alternatively, two or more of the blocks of process 400 may be performed in parallel.
[0091] II. State Analysis and Detection A variety of conditions can be analyzed and detected. Cancer and nuclease activity deficiencies are examples of conditions that can be analyzed and detected using ragged end patterns and / or sequence end motifs. Other conditions, including conditions characterized by abnormal nuclease activity, can also be analyzed and detected.
[0092] A. Cancer For illustrative purposes, sequencing libraries were prepared from plasma DNA samples from one healthy subject and one hepatocellular carcinoma (HCC) patient. These sequencing libraries were sequenced using the PacBio sequencing platform, yielding 59,133 and 227,198 circular consensus sequencing (CCS) reads, respectively. Blunt-end, 5'-overhanging ragged ends (lengths ranging from 1 to 10 nt), and 3'-overhanging ragged ends (lengths ranging from 1 to 10 nt) hairpin adapters were used.
[0093] 1. Detection of ragged ends estimated from simultaneous analysis of single-molecule end patterns A higher occurrence of 5'-overhanging ragged ends has been reported in HCC plasma DNA samples compared to healthy controls (Jiang et al., Genome Res. 2020;30:1144-1153). In embodiments, 5'-overhanging ragged ends, 3'-overhanging ragged ends, and blunt ends can be estimated from simultaneous single-molecule end-to-end analysis in a more precise, accurate, and comprehensive manner, potentially improving diagnostic power.
[0094] Figure 5A shows the frequency of 5'-overhanging ragged ends, 3'-overhanging ragged ends, and blunt ends in HCC and healthy subjects. The y-axis shows frequency. The x-axis shows the type of ragged or blunt end. Two different bars represent healthy subjects versus subjects with HCC. Both ragged ends from one cfDNA fragment were analyzed separately. The frequency is based on the total number of ends (two ends per molecule), not the total number of molecules.
[0095] As shown in Figure 5A, the frequency of 5'-protruding ragged ends was higher in HCC cases compared with healthy controls (59.40% vs. 55.84%). A slight decrease in 3'-protruding ragged ends (23.51% vs. 25.57%) and blunt ends (17.09% vs. 18.59%) was observed in HCC cases.
[0096] Figure 5B shows the frequency of molecules across the combined ragged end categories. The y-axis shows the frequency as a percentage. The x-axis shows the different combinations of ragged end characteristics for both ends of each molecule: 5'-overhanging ragged end and 3'-overhanging ragged end (5-3), 5'-overhanging ragged end and 5'-overhanging ragged end (5-5), 3'-overhanging ragged end and 3'-overhanging ragged end (3-3), 5'-overhanging ragged end and blunt end (5-B), 3'-overhanging ragged end and blunt end (3-B), and blunt and blunt end (BB). Two different bars represent HCC subjects and healthy subjects. Both ragged ends from a single cfDNA fragment were analyzed simultaneously.
[0097] As shown in Figure 5B, HCC cases exhibited higher amounts of cfDNA fragments belonging to categories 5-5 (37.32% vs. 33.32%) and 5-B (18.75% vs. 17.65%) compared with healthy controls, but lower amounts of cfDNA fragments belonging to categories 5-3 (25.41% vs. 27.39%) and 5-BB (4.28% vs. 5.94%). The cumulative difference across these categories between HCC and healthy controls was higher when simultaneously analyzing both ends (10.21%) compared with a single end (7.12%). These results indicate that simultaneous analysis of ragged end patterns from both sides of cfDNA fragments can provide more detailed information not available with previously published techniques and improve diagnostic power.
[0098] 2. Detection with estimated fragment size from simultaneous analysis of single molecule end-type In one embodiment, ragged ends and fragment sizes estimated from simultaneous single molecule end-to-end analysis can be analyzed along with ragged end categories.
[0099] Figure 6A shows a graph of the overall size distribution of plasma DNA samples from healthy and HCC subjects. The y-axis shows frequency in percent. The x-axis shows size in bp. Fragment sizes are slightly shorter in HCC cases compared to healthy subjects.
[0100] Figure 6B shows a graph of the frequency of fragments less than 150 bp in size across the combined ragged end categories. The y-axis is the percentage frequency within that category of fragments less than 150 bp across all sizes of molecules within that category. The x-axis lists the combined ragged end categories: 5'-overhanging ragged end and 3'-overhanging ragged end (5-3), 5'-overhanging ragged end and 5'-overhanging ragged end (5-5), 3'-overhanging ragged end and 3'-overhanging ragged end (3-3), 5'-overhanging ragged end and blunt end (5-B), 3'-overhanging ragged end and blunt end (3-B), and blunt end and blunt end (BB). The x-axis also lists all cfDNA fragments. Both ends from a single cfDNA fragment were analyzed simultaneously.
[0101] Figure 6C shows a graph of the frequency of fragments greater than 280 bp in size across the combined ragged end categories. The y-axis is the frequency in percent. The x-axis lists the combined ragged end categories and all cfDNA fragments.
[0102] As shown in the "All" cfDNA fragment category in Figure 6B, the frequency of short cfDNA fragments (<150 bp) is higher in HCC cases (24.58% vs. 15.82%). In contrast, as shown in the "All" category in Figure 6C, the frequency of long cfDNA fragments (>280 bp) is lower in HCC cases compared to healthy subjects (14.81% vs. 24.94%).
[0103] Then, cfDNA fragments were separated into different groups according to the ragged end types at both ends.As shown in Figure 6B, compared with all cfDNA fragments (24.58% vs. 15.82%), the cfDNA fragments with blunt ends at both ends (BB, 24.44% vs. 10.27%), the cfDNA fragments with ragged ends and 3' overhangs at both ends (3-3, 33.45% vs. 19.78%), and the cfDNA fragments with ragged ends and blunt ends at 3' overhangs (3-B, 26.38% vs. 15.09%) showed a greater difference in the frequency of short cfDNA fragments between HCC cases and healthy cases.
[0104] As shown in Figure 6C, the populations of cfDNA fragments with blunt ends at both ends (BB, 20.29% vs. 57.60%), cfDNA fragments with 5'-overhanging ragged ends and blunt ends (5-B, 11.61% vs. 25.83%), and cfDNA fragments with 3'-overhanging ragged ends and blunt ends (3-B, 15.78% vs. 28.11%) showed larger differences in the frequency of long cfDNA fragments between HCC cases and healthy cases compared with all cfDNA fragments (14.81% vs. 24.94%).
[0105] Figure 7A is a graph of the ratio of short fragments to long fragments for different types of ragged ends. The y-axis is the ratio of short (less than 150 bp) fragments to long (more than 280 bp) fragments. The x-axis is the combined ragged end category and all cfDNA fragments. Different bars represent HCC cases and healthy cases. Both ends from one cfDNA fragment were analyzed simultaneously.
[0106] Figure 7B is a graph of the fold change in the short / long ratio of HCC cases versus healthy controls. The y-axis is the fold change calculated by dividing the short / long ratio of HCC cases by the short / long ratio of healthy controls. The x-axis is the combined jumbled end category and total cfDNA fragments.
[0107] The difference in short / long ratio (i.e., amount of fragments less than 150 bp / amount of fragments greater than 280 bp) between HCC cases and healthy cases was increased in cfDNA fragments with blunt ends at both ends (BB, 1.20 vs. 0.18, fold change: 6.75), cfDNA fragments with 5'-overhanging ragged ends and blunt ends (5-B, 1.90 vs. 0.61, fold change: 3.13), and cfDNA fragments with 3'-overhanging ragged ends and blunt ends (3-B, 1.67 vs. 0.53, fold change: 3.11) compared to all fragments (all, 1.65 vs. 0.63, fold change: 2.61). Figures 7A and 7B show that certain types of ragged ends or certain combinations of ragged ends can be as effective or more effective in distinguishing between healthy and HCC cases than using fragments without considering the ragged end type.
[0108] 3. Detection of single-molecule terminal patterns with deduced terminal motifs from simultaneous analysis The presence of the 5' CCCA end motif has been reported to be reduced in plasma DNA samples from patients with HCC compared to healthy subjects (Jiang et al. Cancer Discov. 2020;10:664-673). In embodiments, the 5' CCCA end motif can be calculated separately for 5'-overhanging ragged ends, 3'-overhanging ragged ends, and blunt ends.
[0109] Figure 8 is a graph of the CCCA end motif across different types of ends. The y-axis shows the CCCA frequency as a percentage. The x-axis shows where the CCCA end motif is found: 5'-overhanging ragged ends, 3'-overhanging ragged ends, blunt ends, and all fragments. Two bars represent healthy and HCC cases. Both ends from one cfDNA fragment were analyzed separately.
[0110] As shown in Figure 8, the frequency of the 5'CCCA end motif was decreased in HCC cases in 5'-overhanging ragged ends or 3'-overhanging ragged ends of cfDNA. For blunt ends, the frequency of the 5'CCCA end motif was increased in HCC cases compared with healthy subjects. For 5'-overhanging ragged ends, 3'-overhanging ragged ends, and blunt ends, the difference in the frequency of the 5'CCCA end motif between HCC cases and healthy subjects was greater than the 5'CCCA end motif estimated from all fragment ends. Figure 8 demonstrates that determining the type of ragged end can increase the accuracy of distinguishing HCC cases from healthy cases.
[0111] Figures 9A and 9B show a technique that can combine ragged ends, 5' end motifs, and 3' end motifs to measure the end-style topology of cfDNA molecules. In Figure 9A, there is a 5'-overhanging ragged end with a 5' "CCCA" end motif and a 3' "TTTT" end motif; the topology of the end motifs can be referred to as "CCCA_TTTT," where the 5' end motif is followed by the 3' end motif, both of which are capitalized and have an underscore (i.e., "_") as a separator.
[0112] In Figure 9B, there is a 3'-overhanging ragged end with a 5' "CCCA" end motif and a 3' "GAGG" end motif; the end motif topology can be referred to as "CCCA_gagg," with the 5' end motif in uppercase followed by the 3' end motif in lowercase. The lowercase letter indicates that the 3' end is an overhanging end. Different naming conventions can be used to indicate overhanging ends, non-overhanging ends, and whether the end is blunt. In some embodiments, the end motif can include only nucleotides from one strand, since information about the other strand can be inferred from only one strand. A separator can be used to indicate the location of the overhang. As an example, Figure 9B can be represented by "3-GAG-GGGT." The "3" indicates the 3' overhang, and the second "-" indicates where the 5' strand end begins.
[0113] Figures 10A and 10B show a technique that can combine ragged ends, 5'-end motifs, and 3'-end motifs from both sides of a fragment to measure the spliced end pattern of a cfDNA molecule. In Figure 10A, there is a DNA fragment with a 5'-overhanging ragged end with a 5' "C" end motif and a 3' "G" end motif on the left side, and a 3'-overhanging ragged end with a 5' "G" end motif and a 3' "T" end motif on the right side. The spliced end motif can be referred to as "5CG3GT," where the first three letters indicate the left end and the next three letters indicate the right end.
[0114] In Figure 10B, there is a DNA fragment with a 5'-overhanging ragged end with a 5' "C" end motif and a 3' "T" end motif on the left, and a blunt end with a 5' "A" end motif and a 3' "T" end motif on the right. The junction end motif can be referred to as "5CTBAT," where the first three letters indicate the left end and the next three letters indicate the right end. The first letter of the triplet indicates the type of ragged end: "5" for 5' ragged end, "3" for 3' ragged end, and "B" for blunt end. The second letter of the triplet indicates the 5' end motif, and the third letter indicates the 3' end motif.
[0115] Figure 11A is a graph of the correlation of overall 5'-end motif frequency between HCC and healthy subjects. The y-axis shows the frequency of 4-mer 5'-end motifs for HCC subjects. The x-axis shows the frequency of 4-mer 5'-end motifs for healthy subjects. Each point represents a different 4-mer end motif. The data show a high correlation with R = 0.98 and p < 2.2e-16. End motifs with dots that deviate further from the line yx may be useful in distinguishing HCC cases from healthy cases.
[0116] Figure 11B is a graph of the correlation of the frequency of topological end motifs between HCC and healthy subjects. The y-axis shows the frequency of 4-mer co-end motifs for HCC subjects. The x-axis shows the frequency of 4-mer co-end motifs for healthy subjects. Each dot represents a different co-end motif containing a 4-mer for both the 5' and 3' ends, distinguishing between 5'-overhanging and 3'-overhanging ends. The data show a correlation with R = 0.92 and p < 2.2e-16.
[0117] Figure 11C is a graph of the correlation of the frequency of junction end motifs between HCC and healthy subjects. The y-axis shows the frequency of junction end motifs for HCC subjects. The x-axis shows the frequency of junction end motifs for healthy subjects. Each dot represents a different junction end motif, including ragged end types and 1-mer motifs for both the 5' and 3' ends of cfDNA fragments. The data show a correlation with R = 0.91 and p < 2.2e-16.
[0118] Compared with a typical analysis based on the overall 5'-end motif in Figure 11A, the phase of the end motif showed a greater difference between HCC and healthy subjects in Figure 11B. The ranks of the top four motifs in the overall 5'-end motif remained the same between HCC and healthy subjects. In contrast, the ranks of the top four topological end motifs changed significantly. For example, the top-ranked topological end motif (CCCA_gagg) in healthy subjects dropped to fourth in HCC, while the second-ranked topological end motif (AAAA_TTTT) in healthy subjects rose to the top-ranked topological end motif in HCC patients. The difference in topological end motifs between HCC and healthy subjects indicates that different topological end motifs or combinations of different end motifs can be used to distinguish HCC cases from healthy cases.
[0119] Furthermore, compared with the overall 5'-end motifs and topological end motifs, the junction motifs further exaggerated the differences between HCC and healthy subjects (Figure 11C). The ranks of the top four motifs in the overall 5'-end motifs were the same between HCC and healthy subjects. Although the ranks of the top four topological end motifs were significantly altered, the top four topological end motifs were the same between HCC (top four topological motifs: AAAA_TTTT, CAAA_TTTT, CCCC_GGGT, and CCCA_gagg) and healthy subjects (CCCA_gagg, AAAA_TTTT, CAAA_TTTT, and CCCC_GGGT). In contrast, the top four junction end motifs were completely different between HCC (top four junction motifs: 5CT5CT, 5CG5CG, BCGBCG, and 5CA5CA) and healthy subjects (top four junction motifs: BATBAT, BGCBGC, BATBGC, and BGCBAT). The differences in the junction end motifs between HCC and healthy subjects indicate that information from ragged ends, a combination of different end motifs from both sides of the cfDNA fragments, can be used to distinguish HCC cases from healthy cases.
[0120] B. Nuclease activity DNASEs play different roles in cfDNA fragmentation: ragged end patterns and / or terminal motifs can be used to analyze nuclease activity.
[0121] 1. Uneven end style Our previous research has shown that various DNASEs play different roles in the generation of ragged ends in cfDNA. DNASE activity can be estimated by ragged ends (Ding et al. Clin Chem. 2022;68:917-926). However, previously, only 5'-overhanging ragged ends were analyzed, and 3'-overhanging ragged ends were not. Analysis of all types of ragged ends (e.g., 5'-overhanging ragged ends, 3'-overhanging ragged ends, and blunt ends), as well as simultaneous analysis of single-molecule end patterns, may provide more information about the activity of various DNASEs.
[0122] Figures 12A-12F show the analysis of wild-type, DNASE1 (DNASE1) using single-molecule end-mode analysis on the PacBio platform. - / - ), DNASE1L3(DNASE1L3 - / - ), and DFFB(DFFB - / - ) Analysis of plasma cfDNA samples from a knockout mouse model (median reads: 1,295,159, range: 176,285-2,624,708). The x-axis indicates categories of nuclease activity. The y-axis indicates the frequency of specific ragged end patterns.
[0123] DNASE1 - / - Mice showed a significant reduction in the frequency of fragments carrying 5'-overhanging ragged ends (8.76%) (Figure 12A), and a significant reduction in the frequency of fragments carrying 3'-overhanging ragged ends (52.80%) in DNASE1L3 - / - A significant reduction in the frequency of blunt-ended fragments (40.25%) can be observed in DFFB mice (Fig. 12B). - / - This can be observed in mice (Figure 12C). These results indicate that analysis of all types of ragged ends can provide more information about the activity of various DNASEs than analysis of 5'-overhanging ragged ends alone.
[0124] Furthermore, we classified the cfDNA fragments according to the ragged end pattern from both sides of the molecule (i.e., 5'-overhanging ragged end + 3'-overhanging ragged end (5-3), 5'-overhanging ragged end + 5'-overhanging ragged end (5-5), 3'-overhanging ragged end + 3'-overhanging ragged end (3-3), 5'-overhanging ragged end + blunt end (5-B), 3'-overhanging ragged end + blunt end (3-B), blunt end + blunt end (BB)). As shown in Figure 12D, DNASE1 - / -A greater decrease in the frequency of cfDNA fragments carrying 5-5 ragged ends was observed compared to the frequency of 5'-overhanging ragged ends in DNASE1- / - mice (5-5 ragged ends vs. 5'-overhanging ragged ends: 15.40% vs. 8.76%). Similar to DNASE1- / - mice, a greater decrease in the frequency of cfDNA fragments carrying 3-3 ragged ends and BB ragged ends was observed compared to the frequency of 3'-overhanging ragged ends and blunt ends in DNASE1L3- / - mice (reduced: 3-3 ragged ends vs. 3'-overhanging ragged ends: 71.45% vs. 52.80%) (Figure 12E) and DFFB- / - mice (reduced: BB ragged ends vs. blunt ends: 70.41% vs. 40.25%) (Figure 12F), respectively. These results demonstrate that simultaneous analysis of ragged ends on both sides of a single cfDNA fragment can improve the ability to distinguish between different DNASE activities. This technique can be used to improve the diagnostic power of diseases with abnormal DNASE activity, such as, but not limited to, systemic lupus erythematosus.
[0125] 2. Terminal motifs and irregular terminal patterns Our previous publications reported that cfDNA terminal motifs can be used to estimate the activity of various DNASEs (Han et al. Am J Hum Genet. 2020;106:202-214, Jiang et al. Cancer Discov. 2020;10:664-673). The results discussed in the previous section indicated that various DNASEs may be associated with various types of ragged ends. Analyzing the terminal motifs in various ragged end groups may improve the discriminatory power in detecting changes in DNASE activity.
[0126] Figure 13 is a table of various ragged end styles and terminal nucleotide types for various nuclease activities. Major columns 1304, 1308, and 1312 show data for various ragged ends. Major rows 1316, 1320, and 1324 show the various nuclease activities analyzed. Individual columns below each major column show the median frequencies for wild-type and specific nuclease knockout mice for that major row, as well as the relative change in median frequency between nuclease knockout and wild-type mice. Individual rows show the terminal nucleotide. Gray-shaded cells indicate the maximum change in each terminal nucleotide between different ragged end types. There is only one shaded cell per row. For example, in DNASE1L3- / - mice, the maximum change in A-terminus was seen in 5' ragged ends, so the well for the relative change in A-terminus of 5' ragged ends is shaded.
[0127] As shown in Figure 13, comparing fragments with 3'-overhanging ragged ends and fragments with blunt ends, the fragments with 5'-overhanging ragged ends showed a higher DNASE1L3 expression level compared to WT mice. - / - The largest increase in 5'-end motifs was observed in A (median increase: 5' vs. 3' vs. blunt: 39.28% vs. 13.86% vs. 25.68%) and G (median increase: 5' vs. 3' vs. blunt: 21.55% vs. 4.79% vs. 4.79%) in WT mice. Compared with fragments with 5'-overhanging ragged ends and fragments with 3'-overhanging ragged ends, fragments with blunt ends showed a significantly higher DNASE1L3 abundance compared to WT mice. - / - Mice showed the greatest reduction in 5'-end motifs: C (median reduction: 5' vs. 3' vs. blunt: 18.83% vs. 4.25% vs. 21.60%) and T (median reduction: 5' vs. 3' vs. blunt: 41.44% vs. 10.23% vs. 77.99%). - / -In DFFB mice, fragments with blunt ends showed the greatest reduction in 5' C-terminus (median reduction: 5' vs. 3' vs. blunt: 9.67% vs. 0.40% vs. 13.70%) and 5' T-terminus (median reduction: 5' vs. 3' vs. blunt: 7.68% vs. 1.64% vs. 45.90%), and the most significant increase in 5' A-terminus (median increase: 5' vs. 3' vs. blunt: 11.43% vs. 2.18% vs. 19.69%), whereas 5' protruding ragged ends showed the greatest increase in 5' G-terminus (median increase: 5' vs. 3' vs. blunt: 11.03% vs. -1.74% vs. -0.29%). Interestingly, in DFFB mice, fragments with blunt ends showed the greatest reduction in 5' C-terminus (median reduction: 5' vs. 3' vs. blunt: 9.67% vs. 0.40% vs. 13.70%) and 5' T-terminus (median reduction: 5' vs. 3' vs. blunt: 7.68% vs. 1.64% vs. 45.90%), and the most significant increase in 5' A-terminus (median increase: 5' vs. 3' vs. blunt: 11.43% vs. 2.18% vs. 19.69%) compared to WT mice. - / - In mice, the greatest changes in 5'-end motifs were observed in fragments with blunt ends (5'C-terminus (median reduction: 5' vs. 3' vs. blunt: 4.16% vs. -0.80% vs. 33.73%), 5'T-terminus (median reduction: 5' vs. 3' vs. blunt: 22.57% vs. 1.48% vs. 110.94%), 5'A-terminus (median reduction: 5' vs. 3' vs. blunt: 15.43% vs. 2.50% vs. 28.12%), and 5'G-terminus (median reduction: 5' vs. 3' vs. blunt: 4.70% vs. -2.35% vs. 8.68%)). Furthermore, DFFB - / - and WT mice, the 4-mer 5′-end motif was analyzed in all fragments, fragments with 5′-overhanging ragged ends, fragments with 3′-overhanging ragged ends, and fragments with blunt ends.
[0128] 14A to 14D show the DFFB - / - Graph of terminal motif rankings in (DFFB knockout [KO]) and wild-type (WT) mice. The figure has motif rankings in wild-type mice on the x-axis and DFFB on the y-axis. - / - Figure 14A shows the motif ranking in mouse. - / - Figure 14B shows the 5' end motif ranking of all pooled cfDNA fragments in DFFB and wild-type mice. - / - Figure 14C shows the 5' end motif ranking of pooled cfDNA fragments carrying 5' protruding ragged ends in DFFB mice and wild-type mice. - / -Figure 14D shows the 5' end motif ranking of pooled cfDNA fragments carrying 3' overhanging ragged ends in DFFB mice and wild-type mice. - / - 14A-14D show the 5'-end motif ranking of pooled cfDNA fragments bearing blunt ends in DFFB mice and wild-type mice. As shown in Figures 14A-14D, the 5'-end motif ranking of DFFB compared to all fragments, fragments with 5'-overhanging ragged ends, and fragments with 3'-overhanging ragged ends. - / - The greatest difference between DFFB and WT mice can be observed in fragments with blunt ends (R: all vs. 5' vs. 3' vs. blunt: 0.94 vs. 0.95 vs. 1 vs. 0.77). However, fragments with 5' protruding ragged ends also showed a significant difference in the DFFB - / - It has motifs that are either over- or under-represented in mice.
[0129] These data indicated that simultaneous single-molecule end-to-end analysis allows for more accurate deciphering of the characteristic cleavages caused by various DNA nucleases in plasma.
[0130] C. Exemplary Methods The level of the condition can be determined using any of the processes described herein. Examples can include determining the level of the disease or condition in the patient and then treating the disease or condition in the patient. Treatment can include any suitable therapy, drug, or surgery, including any treatment described in the references cited herein. Information regarding treatments in the references is incorporated herein by reference.
[0131] Treatment can be provided depending on the determined level of cancer, the identified mutation, and / or the tissue of origin. For example, the identified mutation (e.g., in a polymorphic embodiment) can be targeted by a specific drug or chemotherapy. The tissue of origin can be used to guide surgery or any other form of treatment. The level of cancer can also be used to determine how aggressive to be in any type of treatment, which can also be determined based on the level of cancer.
[0132] A statistically significant number of cell-free DNA molecules can be analyzed to accurately determine the proportional contribution from the first tissue type. In some embodiments, at least 1,000 cell-free DNA molecules are analyzed. In some examples, at least 1,000 cell-free DNA molecules are analyzed. In other embodiments, at least 10,000, 50,000, 100,000, 500,000, 1,000,000, or 5,000,000 cell-free DNA molecules or more can be analyzed.
[0133] 1. Uneven ends from simultaneous analysis FIG. 15 is a flowchart of an exemplary process 1500 for analyzing a biological sample obtained from an individual. The biological sample may include multiple nucleic acid molecules. The nucleic acid molecules may be cell-free and double-stranded, having a first strand and a second strand. At least one of the nucleic acid molecules may have an overhang, where the first strand or the second strand overlaps with the other. Process 1500 may use overhang information from the four ends of the two strands of the nucleic acid molecule to determine the level of the individual's condition. In some embodiments, one or more process blocks of FIG. 15 may be performed by system 10 or system 2400.
[0134] In block 1502, for each nucleic acid molecule of the plurality of nucleic acid molecules, a first strand-specific classification of a property of a first end of the nucleic acid molecule is measured. The strand-specific classification may indicate whether the first strand or the second strand overhangs the other, including cases where neither strand overhangs the other. The strand-specific classification may identify whether the first strand or the second strand is the 3' strand or the 5' strand. The strand-specific classification may also indicate the length of the overhang of either the first strand or the second strand. The strand-specific classification may include ragged end patterns as described herein. The property may be measured using process 300.
[0135] For each nucleic acid molecule of the plurality of nucleic acid molecules, a second strand-specific classification of a second end of the nucleic acid molecule can be determined.
[0136] In block 1504, a ragged end value is determined using a first strand-specific classification of the plurality of nucleic acid molecules. The ragged end value can be the amount of nucleic acid molecules having a particular type of ragged end, including a 5' overhang, a 3' overhang, and a blunt end (e.g., FIG. 5A). The amount can be number, total length, mass, or frequency. In some embodiments, the ragged end value can be the amount of nucleic acid molecules including an amount from one of the following classifications: a blunt end at a first end and a blunt end at a second end, a 5' overhang at a first end and a blunt end at a second end, a 3' overhang at a first end and a blunt end at a second end, a 5' overhang at a first end and a 3' overhang at a second end, a 5' overhang at a first end and a 5' overhang at a second end, and a 3' overhang at a first end and a 3' overhang at a second end (e.g., FIG. 5B). Determining the ragged end value can use a second strand-specific classification.
[0137] In some examples, the ragged end value can be an element of a vector. The vector can include a plurality of elements. The plurality of elements can include amounts of nucleic acid molecules in one or more of the following categories: a blunt end on a first end and a blunt end on a second end, a 5' overhang on a first end and a blunt end on a second end, a 3' overhang on a first end and a blunt end on a second end, a 5' overhang on a first end and a 3' overhang on a second end, a 5' overhang on a first end and a 5' overhang on a second end, and a 3' overhang on a first end and a 3' overhang on a second end.
[0138] The plurality of elements can include a class of nucleic acid molecules having sizes within one or more size ranges (e.g., Figures 5B and 5C). The one or more size ranges can be any size range described herein. The size range can include sizes smaller or larger than any of the following sizes: 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, or 350. Additionally, the size range can include sizes between and including any two sizes described herein. The size of each nucleic acid molecule can be determined by aligning subsequences corresponding to the ends of each nucleic acid molecule with a reference genome.
[0139] The ragged end value can include the ratio of the amount of nucleic acid molecules in one overhang class and having a size in a particular size range, and the amount of nucleic acid molecules with the same overhang class but with sizes in a different size range. A vector can include size ratios for each of multiple overhang classes (e.g., Figure 7A).
[0140] In block 1506, the unequal end value is compared to a reference value. The comparison may determine whether the unequal end value is statistically significantly different from the reference value. The reference value may be any reference value described herein. The vector may be compared to a reference vector (which may include multiple different reference values). The comparison may be between corresponding elements of the vector. In some embodiments, the comparison may be performed using a machine learning model. For example, the machine learning model may be trained using unequal end values determined from subjects with known condition levels.
[0141] In block 1508, the level of the individual's condition is determined using the comparison. The condition can be cancer, an autoimmune disease, a pregnancy-related disorder, a nuclease activity deficiency, or any condition described herein. The reference value can be determined from one or more subjects whose condition is at a particular level, or from one or more healthy subjects. If the random end value is statistically the same as the reference value, then the level of the condition can be determined to be the same as the subject(s) associated with the reference value.
[0142] In some instances, the level of the condition is not determined. Instead, the fractional concentration of clinically relevant DNA can be determined using the comparison. The reference value can be determined from one or more subjects whose fractional concentrations of clinically relevant DNA are known. The reference value can be a calibration value determined using a calibration sample.
[0143] Process 1500 may include additional implementations, such as any single implementation or any combination of implementations described herein, and / or in conjunction with one or more other processes described elsewhere herein.
[0144] Although Figure 15 illustrates example blocks of process 1500, in some implementations, process 1500 may include additional, fewer, different, or differently arranged blocks compared to the blocks illustrated in Figure 15. Additionally or alternatively, two or more of the blocks of process 1500 may be performed in parallel.
[0145] 2. Terminal motif at one end FIG. 16 is a flowchart of an exemplary process 1600 for analyzing a biological sample obtained from an individual. The biological sample may include a plurality of nucleic acid molecules. The nucleic acid molecules may be cell-free and double-stranded, having a first strand and a second strand. Process 1600 may use a terminal motif on at least one end of the molecule to determine the level of a condition. At least some of the nucleic acid molecules of the plurality of nucleic acid molecules have nucleotides on one strand that do not have a complementary portion on the other strand. In some embodiments, one or more process blocks of FIG. 16 may be performed by system 10 or system 2400.
[0146] At block 1602, for each nucleic acid molecule of the plurality of nucleic acid molecules, a first sequence end motif of a first strand at a first end of the nucleic acid molecule is determined.
[0147] In block 1604, a second sequence end motif of a second strand at the first end of the nucleic acid molecule is determined. In an example, the first strand may have a 5' end at the first end. In another example, the first strand may have a 3' end at the first end. The first strand may overhang the second strand, or the second strand may overhang the first strand. The first end may be blunt-ended. A partial sequence may be determined using process 300. In some embodiments, the second sequence end motif of the second strand may be determined by taking the complementary nucleotide of the corresponding nucleotide of the first strand.
[0148] In block 1606, a first quantity of nucleic acid molecules having a first combination of a first sequence end motif and a second sequence end motif at a first end is determined. The sequence end motif can have 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. The quantity can be number, total length, mass, or frequency. The first combination can be a topological end motif as illustrated in Figures 9A and 9B.
[0149] At block 1608, the first amount is used to generate a value for a terminal motif parameter. In some embodiments, the terminal motif parameter may be the first amount. In some embodiments, the terminal motif parameter may be a ratio of the first amount to another amount (e.g., the amount of all terminal motifs). As an example, the terminal motif parameter may be a frequency.
[0150] A second topological terminal motif can be used in addition to the first topological terminal motif. A second quantity of a nucleic acid molecule having a second combination of a third sequence terminal motif and a fourth sequence terminal motif at the first end can be determined. The third sequence terminal motif can be on the 5' strand or the 3' strand. As an example, the first quantity can be the quantity of AAAA_TTTT and the second quantity can be the quantity of CCCA_gagg. Generating a value for the terminal motif parameter can use the second quantity. The terminal motif parameter can be a vector of different quantities of a particular combination of terminal motifs.
[0151] In block 1610, the value of the terminal motif parameter is compared to a reference value. The reference value may be any reference value described herein. The comparison may be any comparison described herein, including block 1506.
[0152] In block 1612, the level of the individual's status is determined using the comparison. The determination may be performed similarly to block 1508.
[0153] The condition can be cancer, HCC, an autoimmune disease, a pregnancy-related disorder, a nuclease activity deficiency, or any condition described herein. The reference value can be determined from one or more subjects with a particular level of the condition, or from one or more healthy subjects.
[0154] In some instances, the level of the condition is not determined. Instead, the fractional concentration of clinically relevant DNA can be determined using the comparison. The reference value can be determined from one or more subjects whose fractional concentrations of clinically relevant DNA are known. The reference value can be a calibration value determined using a calibration sample.
[0155] Process 1600 may include using two sequence motifs from the ragged end of the other end of the molecule. Four sequence motifs from the same molecule may be used. In embodiments, the plurality of nucleic acid molecules is a first plurality of nucleic acid molecules. The biological sample may include a second plurality of nucleic acid molecules. The first plurality of nucleic acid molecules may include a subset of the second plurality of nucleic acid molecules. Process 1600 may further include, for each nucleic acid molecule of the second plurality of nucleic acid molecules, determining a third sequence terminal motif on a first strand of the second end of the nucleic acid molecule that does not have a complementary portion on the second strand, and determining a fourth sequence terminal motif on a second strand of the second end of the nucleic acid molecule. A second quantity of nucleic acid molecules having a second combination of the third sequence terminal motif and the fourth sequence terminal motif at the second end may be determined. A value for the terminal motif parameter may be generated using the second quantity. In some embodiments, the value for the terminal motif parameter may be the quantity of molecules having a particular combination of four sequence motifs present on the molecule.
[0156] In some embodiments, the ragged end pattern of one end can also be used to determine the level of the condition. Process 1600 can include, for each nucleic acid molecule of the plurality of nucleic acid molecules, determining a first strand-specific classification of a characteristic of a first end of the nucleic acid molecule. The strand-specific classification can indicate whether the first strand or the second strand overhangs the other. For example, at one end, the strand-specific classification can indicate a 3'-overhanging end, a 5'-overhanging end, a blunt end, or a ragged end (generally). Determining the first quantity can include determining a first quantity of nucleic acid molecules having the first combination and first strand-specific classification.
[0157] In some embodiments, a ragged end pattern of the second end may also be used. Process 1600 may include, for each nucleic acid molecule of the plurality of nucleic acid molecules, measuring a second strand-specific classification of the second end of the nucleic acid molecule. Determining the first amount may include determining a first amount of nucleic acid molecules having the first combination, the first strand-specific classification, and the second strand-specific classification.
[0158] Process 1600 may include additional implementations, such as any single implementation or any combination of implementations described herein, and / or in conjunction with one or more other processes described elsewhere herein.
[0159] Although Figure 16 illustrates example blocks of process 1600, in some implementations, process 1600 may include additional, fewer, different, or differently arranged blocks compared to the blocks illustrated in Figure 16. Additionally or alternatively, two or more of the blocks of process 1600 may be performed in parallel.
[0160] 3. 3'-terminal motif FIG. 17 is a flowchart of an exemplary process 1700 for analyzing a biological sample obtained from an individual. The biological sample may include a plurality of nucleic acid molecules. The nucleic acid molecules may be cell-free and double-stranded, having a first strand and a second strand. Process 1700 may use a terminal motif on at least one end of the molecule to determine the level of a condition. At least some of the nucleic acid molecules of the plurality of nucleic acid molecules have nucleotides on one strand that do not have a complementary portion on the other strand. In some embodiments, one or more process blocks of FIG. 17 may be performed by system 10.
[0161] In block 1702, for each nucleic acid molecule of the plurality of nucleic acid molecules, a first sequence end motif of a strand at a first end of the nucleic acid molecule is determined, the first end being the 3' end of the strand. The first sequence end motif is the actual end motif of the original molecule, not the end motif at the 3' end after the molecule is blunt-ended by either filling in nucleotides on the 3' strand or removing nucleotides on the 3' strand. The first sequence end motif can be determined using process 300.
[0162] At block 1704, a first amount of nucleic acid molecules having a first sequence end motif at a first end is determined. The first amount can be an absolute amount or a relative amount.
[0163] The sequence end motifs at both ends of the single strand can be determined. In some examples, process 1700 can include, for each nucleic acid molecule of the plurality of nucleic acid molecules, determining a second sequence end motif at a strand at a second end of the nucleic acid molecule. The first quantity is of nucleic acid molecules having a first sequence motif at a first end and a second sequence end motif at a second end.
[0164] In some embodiments, the size of a single strand can be determined. Both ends of the single strand can be aligned to a reference genome to determine the size. The size of the complementary strand can also be determined. The length difference between the two strands can be generated using the two sizes. A statistical value of the length difference between the two strands for multiple molecules can be determined and compared to a reference value. The comparison can be used to determine the level of the condition. The length difference can also be determined by adding or subtracting the length of the overhang at each end of the molecule, without having to determine the length of either strand.
[0165] At block 1706, the first amount is used to generate a value for a terminal motif parameter. In some embodiments, the terminal motif parameter may be a ratio of the first amount to another amount (e.g., the amount of all terminal motifs). As an example, the terminal motif parameter may be a frequency.
[0166] At block 1708, the value of the terminal motif parameter is compared to a reference value. The reference value may be any reference value described herein. For example, the reference value may be determined from a calibration sample in which the level of the individual's condition is known.
[0167] In block 1710, the level of status of the individual is determined using the comparison.
[0168] In some embodiments, terminal motifs at both 3' ends of a single molecule may be used. The strand may be a first strand. The terminal motif parameter may be a first terminal motif parameter. The reference value may be a first reference value. Process 1700 may further include, for each nucleic acid molecule of the plurality of nucleic acid molecules, determining a second sequence terminal motif of a second strand at a second end of the nucleic acid molecule. The second end is the 3' end of the second strand. Process 1700 may include determining a second quantity of nucleic acid molecules having the second sequence terminal motif at the second end. A value of the second terminal motif parameter may be generated using the second quantity. The value of the second terminal motif parameter may be compared to a second reference value. A level of the condition may be determined using the comparison.
[0169] In some embodiments, the ragged end pattern of one end can also be used to determine the level of the condition. Process 1700 can include, for each nucleic acid molecule of the plurality of nucleic acid molecules, determining a first strand-specific classification of a characteristic of a first end of the nucleic acid molecule. The strand-specific classification can indicate whether the first strand or the second strand overhangs the other. For example, at one end, the strand-specific classification can indicate a 3'-overhanging end, a 5'-overhanging end, a blunt end, or a ragged end (generally). The first (or second) strand-specific classification can be a particular ragged end pattern from a list of possible strand-specific classifications. Determining the first quantity can include determining a first quantity of nucleic acid molecules having the first combination and the first strand-specific classification.
[0170] In some embodiments, a ragged end pattern of the second end may also be used. Process 1700 may include, for each nucleic acid molecule of the plurality of nucleic acid molecules, measuring a second strand-specific classification of the second end of the nucleic acid molecule. Determining the first amount may include determining a first amount of nucleic acid molecules having the first combination, the first strand-specific classification, and the second strand-specific classification.
[0171] Process 1700 may include additional implementations, such as any single implementation or any combination of implementations described herein, and / or in conjunction with one or more other processes (e.g., process 1600) described elsewhere herein.
[0172] Although Figure 17 illustrates example blocks of process 1700, in some implementations, process 1700 may include additional, fewer, different, or differently arranged blocks compared to the blocks illustrated in Figure 17. Additionally or alternatively, two or more of the blocks of process 1700 may be performed in parallel.
[0173] III. Concentration Certain types of DNA, including clinically relevant DNA, may tend to be more highly represented among DNA with certain ragged end patterns or sequence end motifs. Thus, enrichment of certain ragged end patterns and / or sequence end motifs may result in a sample enriched for certain types of clinically relevant DNA. Enrichment may include physical enrichment of the sample or in silico enrichment of reads obtained from analysis of a biological sample.
[0174] Uneven ends To explore the potential application of simultaneous single-molecule end-pattern analysis in noninvasive prenatal testing (NIPT), end-pattern analysis has been performed on fetal-specific and shared cfDNA fragments in maternal plasma. Fetal-specific and shared cfDNA fragments were defined by genotypes in maternal buffy coat and placental tissue samples obtained using microarray-based genotyping technology (Illumina HumanOmni2.5 Genotyping Array). Informative SNPs were identified (i.e., the mother was homozygous (represented as the AA genotype) and the fetus was heterozygous (represented as the AB genotype)). Fetal-specific DNA fragments were identified according to the DNA fragments carrying the fetal-specific allele at the informative SNP site. In this scenario, the B allele was assumed to be fetal-specific, and DNA fragments carrying the B allele were assumed to originate from fetal tissue. Fetal-specific DNA fragments were identified according to the DNA fragments carrying the shared allele at the informative SNP site. In this scenario, the A allele was shared, and DNA fragments carrying the A allele were assumed to originate from fetal and maternal tissues (primarily maternal tissue). The number of fetal-specific molecules (p) carrying the fetal-specific allele (B) was determined. The number of molecules (q) carrying the shared allele (A) was determined. The fetal DNA fraction across all cell-free DNA samples was calculated by 2p / (p+q)*100%.
[0175] Figures 18A-18C show the simultaneous single molecule end-to-end analysis on the PacBio sequencing platform for a total of 10 plasma DNA samples from pregnant women (median number of reads: 1,305,115, range: 393,197-1,921,070).
[0176] Figure 18A is a graph of the frequency of 5'-overhanging ragged ends. The y-axis shows the frequency of 5'-overhanging ragged ends in percent. The x-axis shows fragments carrying shared alleles and fragments carrying fetal-specific alleles. Compared to shared cfDNA, fetal-specific cfDNA carries more 5'-overhanging ragged ends (shared vs. fetal-specific: median: 52.0% vs. 59.2%).
[0177] Figure 18B is a graph of the frequency of 3'-overhanging ragged ends. The y-axis shows the frequency of 3'-overhanging ragged ends in percent. The x-axis shows fragments carrying shared alleles and fragments carrying fetal-specific alleles. Compared to shared cfDNA, fetal-specific cfDNA carries more 3'-overhanging ragged ends (shared vs. fetal-specific: median: 23.0% vs. 33.5%).
[0178] Figure 18C is a graph of blunt end frequency. The y-axis shows the blunt end frequency in percent. The x-axis shows fragments carrying shared alleles and fragments carrying fetal-specific alleles. Compared to shared cfDNA, fetal-specific cfDNA carries fewer blunt ends (shared vs. fetal-specific: median: 23.5% vs. 8.7%).
[0179] Figure 18D shows the fetal DNA fraction percentage based on various end styles. The y-axis shows the fetal DNA fraction as a percentage. The x-axis shows the various end styles. Selective analysis of cfDNA carrying 5'-overhanging ragged ends (5'-overhanging ragged ends vs. all fragments: median: 16.61% vs. 15.41%) or 3'-overhanging ragged ends (3'-overhanging ragged ends vs. all fragments: median: 17.94% vs. 15.41%) showed a significant increase in fetal DNA fraction compared to all cfDNA fragments. In contrast, selective analysis of cfDNA carrying blunt ends (blunt ends vs. all fragments: median: 5.88% vs. 15.41%) showed a significant decrease in fetal DNA fraction compared to all cfDNA fragments, which indeed indicated a significant increase in the fraction of DNA of maternal origin. Figures 18A-18D show that ragged end types can be used to enrich for DNA of a particular origin.
[0180] In another embodiment, the cfDNA fragments were classified into six different groups according to the ragged end pattern on both sides of the molecule (i.e., 5'-overhanging ragged end + 3'-overhanging ragged end (5-3), 5'-overhanging ragged end + 5'-overhanging ragged end (5-5), 3'-overhanging ragged end + 3'-overhanging ragged end (3-3), 5'-overhanging ragged end + blunt end (5-B), 3'-overhanging ragged end + blunt end (3-B), blunt end + blunt end (BB)).
[0181] Figure 19 is a graph of fetal DNA fraction versus various ragged end styles. The y-axis shows the fetal DNA fraction estimated from cfDNA fragments. The x-axis shows various end styles (e.g., 5'-overhanging ragged end and 3'-overhanging ragged end (5-3), 5'-overhanging ragged end and 5'-overhanging ragged end (5-5), 3'-overhanging ragged end and 3'-overhanging ragged end (3-3), 5'-overhanging ragged end and blunt end (5-B), 3'-overhanging ragged end and blunt end (3-B), blunt end and blunt end (BB)) and total fragments. Selective analysis of cfDNA belonging to the 3-3 (3-3 vs. all fragments: median: 23.48% vs. 15.41%), 5-3 (5-3 vs. all fragments: median: 19.50% vs. 15.41%), and 5-5 (5-5 vs. all fragments: median: 17.93% vs. 15.41%) groups showed a significant increase in fetal DNA fraction compared to all cfDNA fragments. In contrast, selective analysis of cfDNA belonging to the 3-B (3-B vs. all fragments: median: 9.12% vs. 15.41%), 5-B (5-B vs. all fragments: median: 8.17% vs. 15.41%), and BB (BB vs. all fragments: median: 2.53% vs. 15.41%) groups showed a significant decrease in fetal DNA fraction compared to all cfDNA fragments. These results indicate that analysis of ragged end patterns from both sides of cfDNA fragments can enrich for clinically relevant DNA.
[0182] B. Simultaneous analysis of ragged ends and terminal motifs Simultaneous analysis of single-molecule end-specific sequences (e.g., combining end-specific motifs with ragged ends) can enrich for fetal DNA in maternal plasma. All sequenced reads from the 10 pregnant women listed above were pooled together (12,142,332 reads).
[0183] Figures 20A and 20B are graphs of the fetal DNA fraction estimated from fragments with certain ragged end patterns and sequence end motifs. The y-axis of the graphs is the fetal DNA fraction estimated from the cfDNA fragments. The x-axis shows the various categories of fragments: all fragments, ragged end patterns, end motifs, and combinations of ragged end patterns and end motifs.
[0184] Figure 20A shows fragments carrying 5'-overhanging ragged ends with 5' CCG end motifs on either side of the fragment (fetal DNA fraction: 29.3%), fragments with a substantial increase in fetal DNA fraction compared to all fragments (fetal DNA fraction: 16.3%), fragments with only 5'-overhanging ragged ends (fetal DNA fraction: 18.5%), or fragments with only the CCG 5' end motif (fetal DNA fraction: 25.2%).
[0185] Figure 20B shows fragments bearing 3'-overhanging ragged ends with 5' GCG end motifs on either side of the fragment (fetal DNA fraction: 35.6%), fragments with a substantial increase in fetal DNA fraction compared to all fragments (fetal DNA fraction: 16.3%), fragments with only 3'-overhanging ragged ends (fetal DNA fraction: 21.2%), or fragments with only a GCG 5' end motif (fetal DNA fraction: 16.2%). These results indicate that simultaneous analysis of end motifs with ragged ends at the same end can facilitate enrichment of fetal DNA in maternal plasma. Simultaneous analysis of single-molecule end formats can facilitate NIPT.
[0186] C. Exemplary Methods FIG. 21 is a flowchart of an exemplary process 2100 for enriching a biological sample for clinically relevant DNA. The biological sample may include clinically relevant DNA and other DNA. Each nucleic acid molecule of the plurality of nucleic acid molecules is double-stranded, having a first strand and a second strand. The clinically relevant DNA may be tumor DNA, transplant DNA, or fetal DNA. The biological sample may be obtained from a female subject carrying a fetus, and the clinically relevant DNA may be either fetal DNA or maternal DNA. In some embodiments, one or more process blocks of FIG. 21 may be performed by system 2400.
[0187] In block 2110, a first strand-specific classification of a first end of a nucleic acid molecule is determined for each nucleic acid of the plurality of nucleic acid molecules. The strand-specific classification indicates whether the first strand or the second strand overhangs the other. The strand-specific classification may include a first strand overhanging the second strand, a second strand overhanging the first strand, and / or a strand overhanging the other (blunt end). A subset of nucleic acid molecules may have a first strand-specific classification that is a first strand overhanging the second strand. The first strand may be the 3' strand or the 5' strand. In some embodiments, the first strand-specific classification of the subset of nucleic acid molecules indicates that the first strand of the nucleic acid molecule overhangs the second strand, and the second strand-specific classification of the subset of nucleic acid molecules indicates that the second strand of the nucleic acid molecule overhangs the first strand. For example, the 5' end may be overhanging on both ends.
[0188] In block 2120, reads corresponding to a subset of nucleic acid molecules having a first strand-specific classification are selected to form an enriched sample. The enriched sample may be an enriched in silico sample. In some embodiments, the enriched sample may be formed through a physical enrichment technique. For example, according to some embodiments, ragged end-specific hybridization-based targeted capture may be used to enrich for a certain number of ragged ends of interest. In one embodiment of physical enrichment analysis, ragged end-specific hybridization-based targeted capture may be used to enrich for ragged ends of interest. Biotinylated RNA probes are designed that can specifically hybridize to the ragged ends of interest. The ragged ends of interest that hybridize with the biotinylated probes can be pulled down by streptavidin-coated magnetic beads. The RNA probes will be degraded by ribonucleases such as RNase H. The ragged ends of interest will be enriched in the pulled-down material. In one embodiment, one or more different ragged ends are analyzed together, for example, the ratio or deviation between the read values of the different ragged ends is analyzed for practical applications.
[0189] In some embodiments, a second strand-specific classification of the second end of each nucleic acid molecule of the plurality of nucleic acid molecules. A subset of nucleic acid molecules may have the second strand-specific classification. For example, the enriched sample may include molecules with the same type of overhang on one end and the same type of overhang on the other end.
[0190] In embodiments, a first sequence end motif at a first end of the nucleic acid molecule can be determined for each nucleic acid molecule of the plurality of nucleic acid molecules. Selecting reads corresponding to a subset of the nucleic acid molecules can include selecting reads corresponding to nucleic acid molecules having the first sequence end motif.
[0191] In embodiments, a second sequence end motif at a second end of the nucleic acid molecule can be determined for each nucleic acid molecule of the plurality of nucleic acid molecules. Selecting reads corresponding to a subset of the nucleic acid molecules can include selecting reads corresponding to nucleic acid molecules having the same second sequence end motif.
[0192] In some embodiments, the method may further include analyzing a subset of nucleic acid molecules to determine a classification of the level of the disorder. For example, the method may include aligning the reads of the subset to a reference genome. Methylation-aware sequencing or other detection techniques may be performed to determine the methylation level or pattern (e.g., the methylation status at one or more genomic sites). The methylation level or pattern may be compared to a reference level or pattern of a control sample with a known level of the disorder. The level of the disorder may be determined using the comparison.
[0193] In some embodiments, the method can include determining a chromosomal abnormality or a fetal haplotype. The subset of reads can be aligned to a reference genome. The chromosomal abnormality (e.g., amplification or deletion) or fetal haplotype can be identified from the alignment.
[0194] In embodiments, process 2100 may further include determining a first quantity of reads. The first parameter may be determined using the first quantity of reads. The first parameter may be determined using the first quantity of sequence reads and another quantity (e.g., the total amount of reads with a particular strand-specific classification or sequence end motif). In some examples, both such quantities may be separate parameters. The other quantity may take various forms, for example, corresponding to the total number of sequence reads and / or analyzed DNA molecules. The first parameter may be a ratio of quantities.
[0195] A characteristic of the biological sample can be determined using a first parameter. The first characteristic can be the fractional concentration of clinically relevant DNA molecules in the biological sample. The characteristic of the biological sample can be the level of an abnormality in the biological sample. The first value of the characteristic of the biological sample is estimated by comparing the first parameter with one or more calibration values determined from one or more calibration samples whose values for the characteristic are known.
[0196] Therefore, the parameters generated based on each nuclease can be used to determine the characteristics of a biological sample. These respective parameters can be combined to form new combined parameters, for example, as a ratio, a ratio of functions of each parameter, and as two inputs to a more complex function, such as a machine learning model. Examples of parameter combinations include DNASE1L3 / DFFB, DNASE1 / DFFB, or other ratios of DNASE1L3:DNASE1:DFFB. Furthermore, parameters of more than two nucleases can be used, for example, relative parameters of three or more nucleases can be used.
[0197] In some embodiments, a first value for a feature of a biological sample is estimated based on an analysis of a set of parameters, each parameter corresponding to the amount of sequence reads, in combination with another amount (e.g., for normalization), each of which includes a terminal sequence corresponding to a particular sequence terminal signature. For example, the parameters may include a particular combination of frequency ratios between two sets of sequence reads having the respective terminal signatures. For example, a first parameter of the set of parameters may correspond to the ratio of strand-specific classifications between a first amount of sequence reads and another amount of sequence reads, each of which includes a strand-specific classification corresponding to a first nuclease, and a second parameter of the set of parameters may correspond to the ratio of strand-specific classifications between a second amount of sequence reads and a third amount of sequence reads, each of which includes a strand-specific classification corresponding to a second nuclease terminal signature. In some cases, the third amount of sequence reads is the other amount of sequence reads used to determine the first parameter.
[0198] The determined characteristic may include, for example, gestational age or range (e.g., 8 weeks, or 9-12 weeks) if nucleases are differentially regulated between fetal and maternal tissues. In another example, the determined characteristic may be a specific tissue type (e.g., hepatocytes) versus other tissue types (e.g., hematopoietic cells). The characteristic of a target tissue type may also indicate a specific pathology of the target tissue type (e.g., HCC, preeclampsia, preterm birth). In another example, the determined characteristic may be the size or nutritional status of an organ corresponding to a specific tissue type (e.g., hepatocytes). In yet another example, the determined characteristic may include the fraction of clinically relevant DNA in a biological sample.
[0199] The comparison can be to multiple calibration values. The comparison can be made by inputting the first parameter into a calibration function that is fitted to calibration data that provides the change in the first parameter relative to the change in the feature in the sample. As another example, one or more calibration values can correspond to other parameters in one or more calibration samples.
[0200] Generally, it is preferred that the one or more calibration values determined from one or more calibration samples are generated using an assay similar to that used for the biological (test) sample, e.g., a sequencing library can be generated in the same way.
[0201] Process 2100 may include additional embodiments, such as any single embodiment or any combination of embodiments described herein, and / or in conjunction with one or more other processes described elsewhere herein or in US 2022 / 0010353 A1, the entire contents of which are incorporated herein by reference for all purposes.
[0202] Although Figure 21 illustrates example blocks of process 2100, in some implementations, process 2100 may include additional, fewer, different, or differently arranged blocks compared to the blocks illustrated in Figure 21. Additionally or alternatively, two or more of the blocks of process 2100 may be performed in parallel.
[0203] IV. Fraction of clinically relevant DNA The results described in this document, including Section II.B, demonstrate a relationship between various ragged end types and the activity of various DNASEs. 5'-overhanging ragged ends have been shown to correlate with DNASE1 activity, and blunt ends have been shown to be associated with DFFB activity. Due to the altered expression of various DNASEs in different tissues, ragged end profiling can be used to infer the tissue of origin of cfDNA.
[0204] Figure 22A is a graph of DNASE1 mRNA expression levels in leukocytes and placenta. The y-axis shows normalized gene expression units (RPKM), i.e., reads per kilobase per million sequenced reads, estimated from RNA sequencing results (Trapnell et al. Nat Biotechnol. 2010;28:511-5). The x-axis shows leukocytes and placenta.
[0205] Figure 22B is a graph of DFFB mRNA expression levels in leukocytes and placenta. Axes are the same as in Figure 22A.
[0206] Figure 22C is a graph showing the correlation between fetal DNA fraction and the frequency of cfDNA fragments carrying 5'-protruding ragged ends.The x-axis shows SNP-based fetal DNA fraction.The y-axis shows the frequency of 5'-protruding ragged ends.
[0207] Figure 22D is a graph of the correlation between fetal DNA fraction and the frequency of cfDNA fragments carrying blunt ends. The x-axis shows SNP-based fetal DNA fraction. The y-axis shows the frequency of blunt ends.
[0208] As shown in the graphs, placental tissues showed higher DNASE1 expression (Figure 22A) and lower DFFB expression (Figure 22B) than leukocytes. Higher DNASE1 correlated with higher 5'-protruding ragged ends, and lower DFFB correlated with lower blunt ends in fragments from placental origin. The frequencies of 5'-protruding ragged ends and blunt ends can be used to reflect the fetal DNA fraction. Figure 22C shows the frequency of 5'-protruding ragged ends, which positively correlates with the fetal DNA fraction, while Figure 22D shows the frequency of blunt ends, which negatively correlates with fetal DNA. This further suggests that the ragged end pattern of plasma DNA may reflect the tissue of origin of these molecules.
[0209] FIG. 23 is a flowchart of an exemplary process 2300 for determining the fraction of clinically relevant DNA in a biological sample. The biological sample may include a plurality of nucleic acid molecules that are cell-free. Each nucleic acid molecule of the plurality of nucleic acid molecules may be double-stranded, having a first strand and a second strand. The biological sample may be obtained from an individual. At least some of the nucleic acid molecules of the plurality of nucleic acid molecules have nucleotides on one strand that do not have a complementary portion on the other strand. The biological sample may be any biological sample described herein. In some embodiments, one or more process blocks of FIG. 23 may be performed by system 2400.
[0210] The clinically relevant DNA can be fetal DNA, tumor DNA, or DNA from tissue types, including placenta, liver, white blood cells, colon, kidney, lung, or any other tissue type described herein.
[0211] At least two different steps are possible in block 2310. First, a first strand-specific classification of a first end of a nucleic acid molecule can be determined for each nucleic acid of the plurality of nucleic acid molecules. The strand-specific classification can indicate whether the first strand or the second strand overhangs the other, with the first strand being the 3' strand. Second, a first sequence end motif present at the first end of the nucleic acid molecule and a second sequence end motif present at the second end of the nucleic acid molecule can be determined for each nucleic acid molecule of the plurality of nucleic acid molecules. The sequence end motif can be an overhang and / or a blunt end.
[0212] In block 2320, a first amount of nucleic acid molecules having a first strand-specific classification of a first strand overhanging a second strand can be determined, or a second amount of a first sequence end motif and a third amount of a second sequence end motif can be determined.
[0213] In block 2330, a parameter may be determined using the first amount or both the second amount and the third amount. For example, the process may include determining a first amount, and determining the parameter may use the first amount. As another example, the process may include determining a second amount and a third amount, and determining the parameter uses both the second amount and the third amount.
[0214] In some embodiments, the amount of 5' overhanging ends can be used in addition to the amount of 3' overhanging ends. For example, process 2300 further includes determining a first amount and determining a fourth amount of nucleic acid molecules having the same first strand-specific classification of second strands overhanging the first strand, and determining the parameter uses the first amount and the fourth amount.
[0215] In some embodiments, the amount of blunt ends can be used in addition to the amount of 3' overhanging ends and / or the amount of blunt ends. For example, process 2300 further includes determining a fifth amount of nucleic acid molecules having the same first strand-specific classification of a first strand that is equivalent to a second strand, and determining the parameters uses the first amount, a fourth amount having a second strand that overhangs the first strand, and the fifth amount.
[0216] In some embodiments, overhangs at both ends of the nucleic acid molecule can be used. For example, process 2300 can further include, for each nucleic acid molecule of the plurality of nucleic acid molecules, measuring a second strand-specific classification of a second end of the nucleic acid molecule. The process can further include determining a fourth amount of nucleic acid molecules having the same second strand-specific classification, and determining the parameter uses the first amount and the fourth amount.
[0217] In some embodiments, a first strand-specific classification, a first sequence end motif, and a second sequence end motif may be used. For example, the process may include determining a first amount, a second amount, and a third amount. Determining the parameters may further include using the first amount, the second amount, and the third amount.
[0218] In some embodiments, the parameter can include a vector of amounts. The vector can include elements of any vector described herein, including elements specifying various combinations for overhangs. Determining the parameter can include using amounts other than the specific amounts mentioned. For example, the parameter can be a ratio or difference from the amount of all nucleic acid molecules.
[0219] At block 2340, the parameter may be compared to a reference value. The comparison may be performed similarly to any comparison described herein, including block 1508. The reference value may be a value determined from one or more control samples in which the fraction of clinically relevant DNA is known. The comparison of the parameter to the reference value may be performed using a machine learning model. The machine learning model may include linear regression, logistic regression, deep recurrent neural network, Bayesian classifier, hidden Markov model (HMM), linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering for applications with noise (DBSCAN), random forest algorithm, and support vector machine (SVM).
[0220] In block 2350, the comparison can be used to determine the fraction of clinically relevant DNA in the biological sample. If the parameter is statistically the same as the reference value, the level of the condition can be determined to be the same as the subject(s) associated with the reference value.
[0221] In some embodiments, the first nuclease can be identified as being differentially regulated in target tissue type compared with at least one other tissue type among multiple tissue types.Clinically relevant DNA molecules can be derived from target tissue type.For example, DNASE1 expression is relatively upregulated in placenta tissue compared with the DNASE1 expression level of leukocyte (Figure 22A).In another example, DNASE1L3 expression is relatively downregulated in HCC cells compared with liver tissue in healthy subjects.
[0222] The first nuclease can be determined to preferentially cleave DNA into DNA molecules with certain strand-specific classes and / or sequence terminal motifs. In some cases, the cleavage preference of the first nuclease is determined by analyzing biological samples from another organism (e.g., a mouse). These strand-specific classes and / or sequence terminal motifs can then be used to determine the fraction of clinically relevant DNA.
[0223] Process 2300 may include additional embodiments, such as any single embodiment or any combination of embodiments, in conjunction with one or more other processes described below and / or elsewhere herein. In a first embodiment, the reference values are determined from one or more calibration samples in which the fractional concentrations of clinically relevant DNA molecules are known.
[0224] Figure 23 shows example blocks of process 2300, although in some implementations, process 2300 may include additional, fewer, different, or differently arranged blocks compared to the blocks shown in Figure 23. Additionally or alternatively, two or more of the blocks of process 2300 may be performed in parallel.
[0225] V. Treatment A. Further Screening Modalities Based on any classification, for example, regarding the fractional concentration of pathology or clinically relevant DNA, the subject may be referred for additional screening modalities, for example, using chest x-ray, ultrasound, computed tomography, magnetic resonance imaging, or positron emission tomography. Such screening may be performed for cancer.
[0226] B. Treatment options Embodiments of the present disclosure can accurately predict disease recurrence (e.g., an increase in tumor DNA fraction after a decrease, a classification of cancer present after a classification of cancer absent), thereby facilitating early intervention and selection of appropriate treatment to improve a subject's disease outcome and overall survival. For example, intensified chemotherapy can be selected for a subject if the corresponding sample predicts disease recurrence. In another example, biological samples from a subject who has completed initial treatment can be sequenced to identify viral DNA that predicts disease recurrence. In such an example, the subject's cancer may be resistant to the initial treatment, and an alternative treatment plan (e.g., a higher dose) and / or a different treatment can be selected for the subject.
[0227] Embodiments may also include treating the subject in response to determining a classification of pathology recurrence. For example, if the prediction corresponds to locoregional failure, surgery may be selected as a possible treatment. In another example, if the prediction corresponds to distant metastasis, chemotherapy may additionally be selected as a possible treatment. In some embodiments, the treatment includes surgery, radiation therapy, chemotherapy, immunotherapy, targeted therapy, hormone therapy, stem cell therapy, or precision medicine. Based on the determined classification of recurrence, a treatment plan may be developed to reduce the risk of harm to the subject and increase overall survival. Embodiments may further include treating the subject according to the treatment plan.
[0228] C. Types of Treatment Embodiments may further include treating the pathology in the patient after determining the classification for the subject. Treatment can be provided according to the determined level of pathology, the fractional concentration of clinically relevant DNA, or the tissue of origin. For example, identified mutations can be targeted with specific drugs or chemotherapy. The tissue of origin can be used to guide surgery or any other form of treatment. The level of pathology can then be used to determine how aggressive any type of treatment should be, which can also be determined based on the level of pathology. Pathology (e.g., cancer) can be treated with chemotherapy, drugs, diet, therapy, and / or surgery. In some embodiments, the higher the value of a parameter (e.g., amount or size) exceeds a reference value, the more aggressive the treatment can be.
[0229] Treatment may include resection. In the case of bladder cancer, treatment may include transurethral resection of bladder tumor (TURBT). This procedure is used for diagnosis, staging, and treatment. During TURBT, a surgeon inserts a cystoscope through the urethra into the bladder. The tumor is then removed using tools with small wire loops, lasers, or high-energy electricity. For patients with non-muscle-invasive bladder cancer (NMIBC), TURBT may be used to treat or eliminate the cancer. Another treatment may include radical cystectomy and lymph node dissection. Radical cystectomy is the removal of the entire bladder and possibly surrounding tissues and organs. Treatment may also include urinary diversion. Urinary diversion is when a doctor creates a new pathway for urine to leave the body if the bladder is removed as part of treatment.
[0230] Treatment may include chemotherapy, which is the use of drugs to destroy cancer cells, usually by preventing their growth and division. Drugs may include, but are not limited to, for example, mitomycin-C (available as a generic drug), gemcitabine (Gemzar), and thiotepa (Tepadina) for intravesical chemotherapy. Systemic chemotherapy may include, but is not limited to, for example, cisplatin gemcitabine, methotrexate (Rheumatrex, Trexall), vinblastine (Velban), doxorubicin, and cisplatin.
[0231] In some embodiments, the treatment may include immunotherapy. The immunotherapy may include immune checkpoint inhibitors that block a protein called PD-1. Inhibitors may include, but are not limited to, atezolizumab (Tecentriq), nivolumab (Opdivo), avelumab (Bavencio), durvalumab (Imfinzi), and pembrolizumab (Keytruda).
[0232] Therapeutic embodiments may also include targeted therapy, which is a treatment that targets specific genes and / or proteins in cancer that contribute to cancer growth and survival. For example, erdafitinib is an orally administered drug approved to treat people with locally advanced or metastatic urothelial carcinoma with FGFR3 or FGFR2 gene mutations, which cause cancer cells to continue to grow or spread.
[0233] Some treatments may include radiation therapy. Radiation therapy is the use of high-energy X-rays or other particles to destroy cancer cells. In addition to each individual treatment, a combination of these treatments described herein may be used. In some embodiments, a combination of treatments may be used when the parameter value exceeds a threshold value, and the threshold value itself exceeds a reference value. Information regarding treatments in references is incorporated herein by reference.
[0234] VI. System 24 illustrates a measurement system 2400 according to one embodiment of the present disclosure. The system shown includes a biological object 2405, such as a biological sample from an organism (e.g., a human), within an analytical device 2410, and an emitter 2408 can transmit waves to the biological object 2405. For example, the biological object 2405 can receive magnetic fields and / or radio waves from the emitter 2408 to provide a signal of a physical feature 2415. The biological object 2405 may include an object treated with an enzyme, label, or primer, or other agent to facilitate detection. An example of an analytical device can be a sequencing device. The analytical device 2410 can include multiple modules.
[0235] A physical characteristic 2415 (e.g., light intensity, voltage, or current) from the biological object is detected by a detector 2420. The detector 2420 can take measurements at intervals (e.g., periodic intervals) to obtain data points that constitute a data signal. In one embodiment, an analog-to-digital converter converts the analog signal from the detector to digital form at multiple times. The analysis device 2410 and the detector 2420 can form an assay system, such as a sequencing system, that acquires data according to embodiments described herein. A data signal 2425 is transmitted from the detector 2420 to a logic system 2430. As an example, the data signal 2425 can be used to determine the identity of a nucleotide in the biological object. The data signal 2425 can include various measurements taken simultaneously, e.g., various signals for various regions of the biological object 2405, and thus the data signal 2425 can correspond to multiple signals. The data signal 2425 can be stored in a local memory 2435, an external memory 2440, or a storage device 2445.
[0236] Logic system 2430 may be or include a computer system, ASIC, microprocessor, graphics processing unit (GPU), etc. It may also include or be coupled to a display (e.g., a monitor, LED display, etc.) and user input devices (e.g., a mouse, keyboard, buttons, etc.). Logic system 2430 and other components may be part of a standalone or networked computer system, or they may be directly attached to or incorporated into a device (e.g., an imaging system) including detector 2420 and / or analysis device 2410. Logic system 2430 may also include software executing on processor 2450. Logic system 2430 may include a computer-readable medium that stores instructions for controlling measurement system 2400 to perform any of the methods described herein. For example, logic system 2430 may provide commands to a system including analysis device 2410 to perform a magnetic emission or other physical operation.
[0237] The measurement system 2400 may also include a treatment device 2460 that can provide treatment to a subject. The treatment device 2460 can be used to determine and / or administer a treatment. Examples of such treatments can include surgery, radiation therapy, chemotherapy, immunotherapy, targeted therapy, hormone therapy, stem cell transplantation, and radioactive seed implantation. The logic system 2430 can be connected to the treatment device 2460, for example, to provide results of the methods described herein. The treatment device can receive inputs from other devices, such as imaging devices (e.g., to control treatment, such as control for a robotic system), and user inputs.
[0238] Any of the computer systems referred to herein may utilize any suitable number of subsystems. An example of such a subsystem is shown in FIG. 14, computer system 10. In some embodiments, the computer system includes a single computer device, and the subsystems may be components of the computer device. In other embodiments, the computer system may include multiple devices, each of which is a subsystem, along with its internal components. Computer systems may include desktop and laptop computers, tablets, mobile phones, and other mobile devices.
[0239] The subsystems shown in FIG. 14 are interconnected via a system bus 75. Additional subsystems are shown, such as a printer 74, a keyboard 78, a storage device 79, a monitor 76 (e.g., a display screen such as an LED) coupled to a display adapter 82, and others. Peripherals and input / output (I / O) devices coupled to the I / O controller 71 can be connected to the computer system by any number of means known in the art, such as input / output (I / O) ports 77 (e.g., USB, Lightning). For example, the I / O ports 77 or external interface 81 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect the computer system 10 to a wide area network, such as the Internet, a mouse input device, or a scanner. The interconnection via the system bus 75 allows the central processor 73 to communicate with each subsystem and control the execution of instructions from the system memory 72 or storage device 79 (e.g., a fixed disk such as a hard drive or an optical disk) and the exchange of information between the subsystems. The system memory 72 and / or storage device 79 may embody computer-readable media. Another subsystem is a data collection device 85, such as a camera, microphone, accelerometer, etc. Any of the data mentioned herein can be output from one component to another and can be output to a user.
[0240] A computer system may include multiple identical components or subsystems connected together, for example, by an external interface 81, by an internal interface, or through a removable storage device that can be connected and disconnected from one component to another. In some embodiments, computer systems, subsystems, or devices may communicate over a network. In such cases, one computer may be considered a client and another computer may be considered a server, each of which may be part of the same computer system. The client and server may each include multiple systems, subsystems, or components.
[0241] Aspects of the embodiments can be implemented using hardware circuitry in the form of control logic (e.g., application-specific integrated circuits or field-programmable gate arrays) and / or using computer software with general-purpose programmable processors in a modular or integrated manner. As used herein, a processor can include a single-core processor, a multi-core processor on the same integrated chip, or multiple processing units on a single circuit board or networked, as well as dedicated hardware. Based on the disclosure and teachings provided herein, those skilled in the art will recognize and understand other means and / or methods of implementing embodiments of the present disclosure using hardware and combinations of hardware and software.
[0242] Any of the software components or functions described in this application may be implemented as software code to be executed by a processor using any suitable computer language, such as, for example, Java, C, C++, C#, Objective-C, Swift, or a scripting language such as Perl or Python, using, for example, conventional or object-oriented techniques. The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission. Suitable non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), magnetic media such as a hard drive, or optical media such as a compact disc (CD) or DVD (digital versatile disc) or Blu-ray disc, flash memory, etc. The computer-readable medium may be any combination of such storage or transmission devices.
[0243] Such programs may also be encoded and transmitted using carrier signals adapted for transmission over wired, optical, and / or wireless networks conforming to various protocols, including the Internet. Thus, computer-readable media may be created using data signals encoded with such programs. Computer-readable media encoded with program code may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer-readable medium may reside on or within a single computer product (e.g., a hard drive, CD, or entire computer system), or may reside on or within different computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing a user with any of the results mentioned herein.
[0244] Any of the methods described herein may be implemented, in whole or in part, using a computer system including one or more processors that may be configured to perform the steps. Accordingly, embodiments may be directed to a computer system configured to perform the steps of any of the methods described herein, potentially with different components performing each step or group of steps. While presented as numbered steps, steps of the methods herein may be performed simultaneously or at different times, or in different orders where logically possible. Additionally, portions of these steps may be used with portions of other steps from other methods. Also, all or portions of steps may be optional. Additionally, any of the steps of any of the methods may be performed using a module, unit, circuit, or other means of a system for performing these steps.
[0245] As will be apparent to those skilled in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has distinct components and features that may be readily separated from or combined with the features of any of the other embodiments without departing from the scope or spirit of the present disclosure.
[0246] The foregoing description of exemplary embodiments of the present disclosure has been presented for purposes of illustration and description and is written to provide those of ordinary skill in the art with a complete disclosure and description of how to make and use embodiments of the present disclosure. It is not intended to be exhaustive or to limit the disclosure to the precise form described, nor is it intended to represent all or the only experiments performed. Although the present disclosure has been described in some detail by way of illustration and example for purposes of clarity of understanding, it will be readily apparent to those skilled in the art in light of the teachings of the present disclosure that certain changes and modifications can be made to the present disclosure without departing from the spirit or scope of the appended claims.
[0247] Accordingly, the foregoing merely illustrates the principles of the present invention. It will be appreciated that those skilled in the art will be able to devise various arrangements, not explicitly described or shown herein, which embody the principles of the present invention and are within its spirit and scope. Moreover, all examples and conditional language recited herein are intended primarily to assist the reader in understanding that the principles of the disclosure are not limited to such specifically recited examples and conditions. Moreover, all statements herein reciting principles, aspects, and embodiments of the present invention, as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, such equivalents are intended to include both currently known equivalents and equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure. Thus, the scope of the present invention is not intended to be limited to the exemplary embodiments shown and described herein. Rather, the scope and spirit of the present invention are embodied by the appended claims.
[0248] The use of "a," "an," or "the" is intended to mean "one or more," unless specifically stated to the contrary. The use of "or" is intended to mean "inclusive or," and not "exclusive or," unless specifically stated to the contrary. A reference to a "first" element does not necessarily require that a second element be provided. Furthermore, a reference to a "first" or "second" element does not limit the referenced element to a particular location unless expressly stated. The term "based on" is intended to mean "based at least in part on."
[0249] The claims may be drafted to exclude any element that may be optional. Accordingly, this statement is intended to serve as a predicate for use of exclusive terminology such as "solely," "only," or the like in connection with the recitation of claim elements or for use of a "negative" limitation.
[0250] Where a range of values is provided, unless the context clearly dictates otherwise, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limit of that range is also specifically disclosed. Each smaller range between any stated or intervening value in a stated range and any other stated or intervening value in that stated range is encompassed within an embodiment of the disclosure. The upper and lower limits of these smaller ranges may independently be included or excluded within the range, and each range where either, neither, or both limits are included within the smaller range is also encompassed within the disclosure, subject to any specifically excluded limits in the stated range. When a stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included within the disclosure.
[0251] All patents, patent applications, publications, and descriptions mentioned in this specification are incorporated by reference herein in their entirety for all purposes as if each individual publication or patent was specifically and individually indicated to be incorporated by reference, and are incorporated by reference herein to disclose and describe methods and / or materials in connection with which the publications are cited. None are admitted to be prior art.
Claims
1. 1. A method for analyzing a biological sample containing a plurality of nucleic acid molecules, the biological sample being acellular, the method comprising: For each nucleic acid molecule of said plurality of nucleic acid molecules, ligating a first hairpin adaptor to a first strand of the nucleic acid molecule and a second strand of the nucleic acid molecule at a first end of the nucleic acid molecule, the first hairpin adaptor comprising a first sequence identifier specifying a first length of zero or more nucleotides at the first end of the first hairpin adaptor that does not have a complementary portion at the second end of the first hairpin adaptor; ligating a second hairpin adaptor to the first strand and the second strand at a second end of the nucleic acid molecule, the second hairpin adaptor comprising a second sequence identifier specifying a second length of zero or more nucleotides at the first end of the second hairpin adaptor that does not have a complementary portion at the second end of the first hairpin adaptor, thereby generating a plurality of ligated nucleic acid molecules; performing rolling circle amplification on a first subset of the plurality of ligated nucleic acid molecules to form a plurality of concatemers; and sequencing each concatemer of said plurality of concatemers to identify a respective said first sequence identifier and a respective said second sequence identifier.
2. The method of claim 1, wherein sequencing is performed simultaneously with performing rolling circle amplification.
3. using the first sequence identifier to determine a first length of an overhang present at the first end of nucleic acid molecules of the first subset of nucleic acid molecules of the plurality of nucleic acid molecules; 2. The method of claim 1, further comprising: using the second sequence identifier to determine second overhangs present at the second ends of nucleic acid molecules of the first subset of nucleic acid molecules.
4. adding an exonuclease to the plurality of ligated nucleic acid molecules to remove a second subset of the plurality of ligated nucleic acid molecules; the first subset of the plurality of nucleic acid molecules does not include any of the nucleic acid molecules in the second subset; For each nucleic acid molecule of said second subset, each of the nucleic acid molecules is not fully hybridized to either each of the first hairpin adaptors or each of the second hairpin adaptors; or Each of the first hairpin adaptors or each of the second hairpin adaptors does not completely hybridize to each of the nucleic acid molecules. The method according to claim 1, wherein
5. using the first sequence identifier to determine a first sequence end motif present at the first end of nucleic acid molecules of the first subset of nucleic acid molecules of the plurality of nucleic acid molecules; 2. The method of claim 1, further comprising using the second sequence identifier to determine a second sequence end motif present at the second end of the nucleic acid molecules of the first subset of the plurality of nucleic acid molecules.
6. using each of the first sequence identifiers to determine, for each nucleic acid molecule of the first subset having an overhang at the first end, whether the 5' strand or the 3' strand overhangs the other; 2. The method of claim 1, further comprising: using each of the second sequence identifiers to determine, for each nucleic acid molecule of the first subset having an overhang at the second end, whether the 5' strand or the 3' strand overhangs the other.
7. the biological sample is obtained from a female subject carrying a fetus; The method comprises: selecting reads corresponding to a subset of the plurality of nucleic acid molecules having the 5' strand or the 3' strand overhanging the other end; The method of any one of claims 1 to 6, further comprising analyzing said subset of nucleic acid molecules for characteristics of said fetus.
8. each first hairpin adaptor of the plurality of first hairpin adaptors comprises a first cleavage site; each second hairpin adaptor of the plurality of second hairpin adaptors comprises a second cleavage site; The method comprises:
2. The method of claim 1, further comprising cleaving each concatemer of the plurality of concatemers at a respective first cleavage site and a respective second cleavage site.
9. 2. The method of claim 1, wherein each nucleic acid molecule of the second portion of the first subset has a respective first strand at each first end that is equivalent to a respective second strand.
10. each nucleic acid molecule of the second portion of the first subset having, at each first end, a respective second strand that overhangs a respective first strand; each said first strand is a 5' strand; The method of claim 1 , wherein each of the second strands is a 3′ strand.
11. 1. A method for analyzing a biological sample containing a plurality of nucleic acid molecules, the biological sample being acellular, the method comprising: For each nucleic acid molecule of said plurality of nucleic acid molecules, ligating a first hairpin adaptor to a first strand of the nucleic acid molecule and a second strand of the nucleic acid molecule at a first end of the nucleic acid molecule, the first hairpin adaptor comprising a first sequence identifier specifying a first length of zero or more nucleotides at the first end of the first hairpin adaptor that does not have a complementary portion at the second end of the first hairpin adaptor; ligating a second hairpin adaptor to the first strand and the second strand at a second end of the nucleic acid molecule, the second hairpin adaptor comprising a second sequence identifier specifying a first length of zero or more nucleotides at the first end of the second hairpin adaptor that does not have a complementary portion at the second end of the first hairpin adaptor, thereby generating a plurality of ligated nucleic acid molecules; adding an exonuclease to the plurality of ligated nucleic acid molecules to remove a first subset of the plurality of ligated nucleic acid molecules; For each nucleic acid molecule of said first subset, each of the nucleic acid molecules is not fully hybridized to either each of the first hairpin adaptors or each of the second hairpin adaptors; or Each of the first hairpin adaptors or each of the second hairpin adaptors does not completely hybridize to each of the nucleic acid molecules. and and sequencing each ligated nucleic acid molecule of a second subset of the plurality of ligated nucleic acid molecules to identify each of the first sequence identifiers and each of the second sequence identifiers, wherein the second subset remains in the biological sample after removing the first subset.
12. using the first sequence identifier to determine a first length of an overhang present at the first end of nucleic acid molecules of the second subset of nucleic acid molecules of the plurality of nucleic acid molecules; 12. The method of claim 11, further comprising using the second sequence identifier to determine second overhangs present at the second ends of nucleic acid molecules of the second subset of nucleic acid molecules of the plurality of nucleic acid molecules.
13. using the first sequence identifier to determine a first sequence end motif of an overhang present at the first end of nucleic acid molecules of the second subset of nucleic acid molecules of the plurality of nucleic acid molecules; 12. The method of claim 11, further comprising using the second sequence identifier to determine a second sequence end motif of an overhang present at the second end of nucleic acid molecules of the second subset of nucleic acid molecules of the plurality of nucleic acid molecules.
14. using each of the first sequence identifiers to determine, for each nucleic acid molecule of the second subset having an overhang at the first end, whether the 5' strand or the 3' strand overhangs the other; 12. The method of claim 11, further comprising: using each of the second sequence identifiers to determine, for each nucleic acid molecule of the second subset having an overhang at the second end, whether the 5' strand or the 3' strand overhangs the other.
15. the biological sample is obtained from a female subject carrying a fetus; The method comprises: selecting reads corresponding to a subset of the plurality of nucleic acid molecules having the 5' strand or the 3' strand overhanging the other end; The method of any one of claims 11 to 14, further comprising analyzing said subset of nucleic acid molecules for characteristics of said fetus.
16. 1. A method for analyzing a biological sample comprising a cell-free plurality of nucleic acid molecules, wherein each nucleic acid molecule of the plurality of nucleic acid molecules is double-stranded having a first strand and a second strand, the biological sample being obtained from an individual, and at least some of the nucleic acid molecules of the plurality of nucleic acid molecules have nucleotides on one strand that do not have a complementary portion on the other strand, the method comprising: For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a first strand-specific classification of a property of a first end of the nucleic acid molecule, wherein the strand-specific classification indicates whether the first strand or the second strand overhangs the other; determining ragged end values using the first strand-specific classification of the plurality of nucleic acid molecules; comparing the staggered end values to a reference value; and using said comparison to determine a level of status of said individual.
17. For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a second strand-specific classification of the second end of the nucleic acid molecule; 17. The method of claim 16, wherein determining the ragged end value uses the second strand-specific classification.
18. 17. The method of claim 16, wherein the strand-specific classification further indicates the length of the overhang.
19. the ragged end values are elements of a vector, the vector comprises a plurality of elements, The plurality of elements are classified into the following categories: a blunt end at said first end and a blunt end at said second end; a 5' overhang at said first end and a blunt end at said second end; a 3' overhang at said first end and a blunt end at said second end; a 5' overhang at said first end and a 3' overhang at said second end; a 5' overhang at said first end and a 5' overhang at said second end; and a 3' overhang at said first end and a 3' overhang at said second end; and the amount of nucleic acid molecule in The method comprises: further comprising comparing the vector to a reference vector; 18. The method of claim 17, wherein determining the level of the condition of the individual uses the comparison of the vector with the reference vector.
20. 20. The method of claim 19, wherein the plurality of elements comprises a class of nucleic acid molecules having sizes in one or more size ranges.
21. The method of any one of claims 16 to 20, wherein the condition is a nuclease activity deficiency.
22. 1. A method of analyzing a biological sample comprising a plurality of nucleic acid molecules, wherein each nucleic acid molecule of the plurality of nucleic acid molecules is double-stranded having a first strand and a second strand, and at least some of the nucleic acid molecules of the plurality of nucleic acid molecules have nucleotides on one strand that do not have a complementary portion on the other strand, the biological sample being obtained from an individual, the method comprising: For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a first sequence end motif of the first strand at a first end of the nucleic acid molecule; and determining a second sequence end motif of the second strand of the first end of the nucleic acid molecule; determining a first amount of nucleic acid molecules having a first combination of the first sequence end motif and the second sequence end motif at the first end; generating a value for a terminal motif parameter using the first amount; comparing said value of said terminal motif parameter with a reference value; and using said comparison to determine a level of status of said individual.
23. 23. The method of claim 22, wherein the first strand has a 3' end at the first end, and the first strand overhangs the second strand.
24. 23. The method of claim 22, wherein the first end comprises a blunt end.
25. determining a second amount of nucleic acid molecules having a second combination of a third sequence end motif and a fourth sequence end motif at the first end; 23. The method of claim 22, wherein generating the value of the terminal motif parameter uses the second amount.
26. the plurality of nucleic acid molecules is a first plurality of nucleic acid molecules; the biological sample comprises a second plurality of nucleic acid molecules; the first plurality of nucleic acid molecules comprises a subset of the second plurality of nucleic acid molecules; The method comprises: For each nucleic acid molecule of the second plurality of nucleic acid molecules, determining a third sequence end motif on the first strand of the second end of the nucleic acid molecule; and determining a fourth sequence end motif on the second strand of the second end of the nucleic acid molecule; determining a second amount of nucleic acid molecules having a second combination of the third sequence end motif and the fourth sequence end motif at the second end; 23. The method of claim 22, wherein generating the value of the terminal motif parameter uses the second amount.
27. For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a first strand-specific classification of a first end of the nucleic acid molecule, wherein the strand-specific classification indicates whether the first strand or the second strand overhangs the other; 27. The method of any one of claims 22 to 26, wherein determining the first amount comprises determining the first amount of nucleic acid molecules having the first combination and the first strand-specific classification.
28. For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a second strand-specific classification of the second end of the nucleic acid molecule; 28. The method of Claim 27, wherein determining the first amount comprises determining the first amount of nucleic acid molecules having the first combination, the first strand-specific classification, and the second strand-specific classification.
29. For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a third sequence end motif of the first strand at a second end of the nucleic acid molecule; determining a fourth sequence end motif on the second strand at the second end of the nucleic acid molecule; 23. The method of claim 22, wherein the first combination is of the first sequence end motif at the first end, the second sequence end motif at the first end, the third sequence end motif at the second end, and the fourth sequence end motif at the second end.
30. For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a first strand-specific classification of a first end of the nucleic acid molecule, wherein the strand-specific classification indicates whether the first strand or the second strand overhangs the other; determining a second strand-specific classification of the second end of the nucleic acid molecule; 30. The method of claim 29, wherein determining the first amount comprises determining the first amount of nucleic acid molecules having the first combination, the first strand-specific classification, and the second strand-specific classification.
31. The method of any one of claims 22 to 28, wherein the condition is cancer.
32. The method of any one of claims 22 to 28, wherein the condition is a nuclease activity deficiency.
33. 1. A method for analyzing a biological sample comprising a plurality of nucleic acid molecules that are cell-free, the biological sample being obtained from an individual, each nucleic acid molecule of the plurality of nucleic acid molecules being double-stranded having a first strand and a second strand, at least some of the nucleic acid molecules of the plurality of nucleic acid molecules having nucleotides on one strand that do not have a complementary portion on the other strand, the method comprising: For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a first sequence end motif of a strand at a first end of the nucleic acid molecule, the first end being the 3' end of the strand; determining a first amount of nucleic acid molecules having the first sequence end motif at the first end; generating a value for a terminal motif parameter using the first amount; comparing said value of said terminal motif parameter with a reference value; and using said comparison to determine a level of status of said individual.
34. For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a second sequence end motif for said strand at a second end of said nucleic acid molecule; 34. The method of claim 33, wherein the first amount is of a nucleic acid molecule having the first sequence end motif at the first end and the second sequence end motif at the second end.
35. the strand is a first strand; the terminal motif parameter is a first terminal motif parameter, the reference value is a first reference value; The method comprises: For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a second sequence end motif of a second strand at a second end of the nucleic acid molecule; determining a second amount of nucleic acid molecules having the second sequence end motif at the second end; using said second amount to generate a value for a second terminal motif parameter; comparing the value of the second terminal motif parameter to a second reference value; 34. The method of claim 33, further comprising: using the comparison to determine the level of the condition.
36. For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a first strand-specific classification of the first end of the nucleic acid molecule, wherein the strand-specific classification indicates whether the first strand or the second strand overhangs the other; 36. The method of any one of claims 33 to 35, wherein determining the first amount comprises determining the first amount of nucleic acid molecules having the first combination and the first strand-specific classification.
37. 37. The method of claim 36, wherein the first strand-specific classification is that the 3' strand overhangs the 5' strand.
38. For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a second strand-specific classification of the second end of the nucleic acid molecule; 37. The method of Claim 36, wherein determining the first amount comprises determining the first amount of nucleic acid molecules having the first combination, the first strand-specific classification, and the second strand-specific classification.
39. 1. A method of enriching a biological sample for clinically relevant DNA, the biological sample comprising a plurality of nucleic acid molecules that are cell-free, the plurality of nucleic acid molecules comprising the clinically relevant DNA and other DNA, each nucleic acid molecule of the plurality of nucleic acid molecules being double-stranded having a first strand and a second strand, the method comprising: For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a first strand-specific classification of a first end of the nucleic acid molecule, wherein the strand-specific classification indicates whether the first strand or the second strand overhangs the other strand; selecting reads corresponding to a subset of nucleic acid molecules having said first strand-specific classification to form an enriched sample.
40. 40. The method of Claim 39, wherein the subset of nucleic acid molecules has the first strand-specific classification of the first strand overhanging the second strand.
41. 41. The method of claim 40, wherein the first strand is the 3' strand.
42. 40. The method of claim 39, wherein the first strand-specific classification comprises the first strand overhanging the second strand and the second strand overhanging the first strand.
43. For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a second strand-specific classification of the second end of the nucleic acid molecule; 40. The method of claim 39, wherein said subset of nucleic acid molecules has said second strand-specific classification.
44. the first strand-specific classification of the subset of nucleic acid molecules indicates that the first strand of the nucleic acid molecule overhangs the second strand; 44. The method of Claim 43, wherein the second strand-specific classification of the subset of nucleic acid molecules indicates that the second strand of the nucleic acid molecule overhangs the first strand.
45. 40. The method of claim 39, wherein the clinically relevant DNA is tumor DNA.
46. 40. The method of claim 39, wherein the biological sample is obtained from a female subject carrying a fetus, and the clinically relevant DNA is fetal DNA.
47. 47. The method of claim 45 or 46, further comprising analyzing said subset of nucleic acid molecules to determine a classification of the level of the disorder.
48. 40. The method of claim 39, further comprising analyzing said subset of nucleic acid molecules to determine a chromosomal abnormality or a haplotype of the fetus.
49. For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a first sequence end motif at the first end of the nucleic acid molecule; 49. The method of any one of claims 39 to 48, wherein selecting the reads corresponding to the subset of nucleic acid molecules comprises selecting reads corresponding to nucleic acid molecules having the first sequence end motif.
50. For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a second sequence end motif at the second end of the nucleic acid molecule; 50. The method of Claim 49, wherein selecting the reads corresponding to the subset of nucleic acid molecules comprises selecting reads corresponding to nucleic acid molecules having the second sequence end motif.
51. determining a first quantity of the lead; determining a first parameter using the first quantity of the lead; 51. The method of claim 49 or 50, further comprising determining a characteristic of the biological sample using the first parameter.
52. 52. The method of claim 51 , wherein the characteristic of the biological sample is the fractional concentration of clinically relevant DNA molecules in the biological sample.
53. 52. The method of Claim 51, wherein the first parameter is determined using the first amount and another amount of sequence reads.
54. 52. The method of claim 51 , wherein the characteristic of the biological sample is the level of an abnormality in the biological sample.
55. 1. A method for determining a fraction of clinically relevant DNA in a biological sample comprising a cell-free plurality of nucleic acid molecules, wherein each nucleic acid molecule of the plurality of nucleic acid molecules is double-stranded having a first strand and a second strand, the biological sample being obtained from an individual, and at least some of the nucleic acid molecules of the plurality of nucleic acid molecules have nucleotides on one strand that do not have a complementary portion on the other strand, the method comprising: For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a first strand-specific classification of a first end of the nucleic acid molecule, wherein the strand-specific classification indicates whether the first strand or the second strand overhangs the other, and the first strand is a 3′ strand; or determining a first sequence end motif present at the first end of the nucleic acid molecule and a second sequence end motif present at the second end of the nucleic acid molecule; determining a first amount of nucleic acid molecules having the first strand-specific classification of the first strand overhanging the second strand, or determining a second amount of the first sequence end motif and a third amount of the second sequence end motif; determining a parameter using the first amount or both the second amount and the third amount; comparing said parameter with a reference value; and using said comparison to determine said fraction of clinically relevant DNA in said biological sample.
56. 56. The method of claim 55, wherein the reference value is determined from one or more calibration samples in which the fractional concentrations of the clinically relevant DNA molecules are known.
57. 56. The method of claim 55, further comprising determining the first amount, wherein determining the parameter uses the first amount.
58. determining the first amount; determining a fourth amount of nucleic acid molecules having the first strand-specific classification of the second strand overhanging the first strand; 56. The method of claim 55, wherein determining the parameter uses the first amount and the fourth amount.
59. determining a fifth amount of nucleic acid molecules having the first strand-specific classification of the first strand that is equivalent to the second strand; 59. The method of claim 58, wherein determining the parameter uses the first amount, the fourth amount, and the fifth amount.
60. 56. The method of claim 55, further comprising determining the second amount and the third amount, wherein determining the parameter uses both the second amount and the third amount.
61. 61. The method of claim 60, further comprising determining the first amount, wherein determining the parameter uses the first amount, the second amount, and the third amount.
62. For each nucleic acid molecule of said plurality of nucleic acid molecules, determining a second strand-specific classification of the second end of the nucleic acid molecule; determining a fourth amount of nucleic acid molecules having the second strand-specific classification; 56. The method of claim 55, wherein determining the parameter uses the first amount and the fourth amount.
63. 56. The method of claim 55, wherein the comparison of the parameter with the reference value is performed using a machine learning model.
64. 64. The method of claim 63, wherein the machine learning model comprises linear regression, logistic regression, a deep recurrent neural network, a Bayesian classifier, a hidden Markov model (HMM), a linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering for noisy applications (DBSCAN), a random forest algorithm, or a support vector machine (SVM).
65. 56. The method of claim 55, wherein the clinically relevant DNA is fetal DNA.
66. 56. The method of claim 55, wherein the clinically relevant DNA is tumor DNA.
67. 56. The method of claim 55, wherein the clinically relevant DNA is DNA from a tissue type.
68. 66. The method of any one of claims 16 to 65, wherein the length of the overhang is determined by a method of any one of claims 1 to 10.
69. 10. The method of claim 1, wherein each nucleic acid molecule of the plurality of nucleic acid molecules has a size greater than a cutoff size.
70. 10. The method of claim 1, wherein each nucleic acid molecule of the plurality of nucleic acid molecules has a size smaller than a cutoff size.
71. 71. The method of Claim 69 or 70, further comprising determining the size of each nucleic acid molecule by aligning a subsequence corresponding to the end of each nucleic acid molecule with a reference genome.
72. 10. The method of any one of the above claims, wherein the condition is cancer, HCC, an autoimmune disease, or a pregnancy-related disorder.
73. 10. The method of any one of the above claims, wherein the reference value is determined from one or more subjects with a certain level of the condition or from one or more healthy subjects.
74. A computer product comprising a non-transitory computer readable medium storing a plurality of instructions that, when executed, control a computer system to perform a method according to any of the preceding claims.
75. 1. A system comprising:
75. A computer product according to claim 74; one or more processors for executing instructions stored on the computer-readable medium.
Citation Information
Patent Citations
US20131101876-8
US200810520458-63