Monomolecular chain-specific terminal morphology
By connecting hairpin adaptors to double-stranded cfDNA fragments and enzymatically treating them, the problem of inaccurate detection of end motifs of double-stranded free DNA fragments in the existing technology is solved, high-fidelity detection and analysis of double-stranded DNA is achieved, and the accuracy and completeness of the detection are improved.
Patent Information
- Application Number
- CN202480012432.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-23
- Filing Date
- 2024-02-23
- Publication Date
- 2025-09-12
AI Technical Summary
Existing cell-free DNA analysis methods have difficulty in accurately detecting and analyzing the original 5' and 3' end motifs of both chains of double-stranded cell-free DNA fragments, and often cause changes in the end information during the sequencing process, affecting the accuracy of the analysis.
By using hairpin adaptors to connect to double-stranded cfDNA fragments to form circular DNA molecules, and combining enzyme treatment and selective enrichment technology to ensure the accuracy of the terminal sequence, single-molecule real-time sequencing and rolling circle amplification technology are used to achieve high-fidelity detection of double-stranded DNA.
It can simultaneously and with high fidelity detect the natural fragment omics characteristics of cfDNA molecules, including fragment size, terminal motifs and jagged ends, improving the accuracy and completeness of the analysis.
Smart Images

Figure CN120641574A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is a non-provisional application of and claims the benefit of U.S. Provisional Patent Application No. 63 / 447,847, filed on February 23, 2023, entitled “SINGLE-MOLECULE STRAND-SPECIFIC ENDMODALITIES,” which is incorporated herein by reference in its entirety for all purposes. Background of the Invention
[0004] Cell-free DNA has proven to be particularly useful for molecular diagnosis and monitoring. Cell-based applications include noninvasive prenatal testing (Chiu RKW et al. Proc Natl Acad Sci USA. 2008;105:20458-63), cancer detection and monitoring (Chan KCA et al. Clin Chem. 2013;59:211-24; Chan KCA et al. Proc Natl Acad Sci USA. 2013;110:1876-8; Jiang P et al. Proc Natl Acad Sci USA. 2015;112:E1317-25), transplantation monitoring (Zheng YW et al. Clin Chem. 2012;58:549-58), and tissue-of-origin tracing (Sun K et al. Proc Natl Acad Sci USA. 2015;112:E5503-12; Chan KCA; Snyder MW et al. Cell. 2016;164:57-68). Cell-free nucleic acid analysis methods developed to date include those based on analysis of single nucleotide variants (SNVs), copy number aberrations (CNAs), cell-free DNA end positions, or methylation markers in the human genome. Identifying new nucleic acid analysis methods for detecting novel traits and increasing the accuracy of existing methods would be beneficial. SUMMARY OF THE INVENTION
[0006] Double-stranded, free-stranded DNA fragments contain two termini on each strand. A single molecule can have four termini. Because the two strands are usually not completely complementary, one strand may extend beyond the other, creating overhangs at the ends. These overhangs are often repaired during analysis, resulting in blunt ends that alter the information about the free-stranded DNA fragment termini. This document describes how to obtain native terminus information from each free-stranded DNA fragment and how it can be used in analysis.
[0007] This method can simultaneously evaluate the original 5' end and 3' end motif of Watson and Crick chains, as well as the associated jaggedness at single-base resolution. In some embodiments, the entire fragment omics features from cfDNA molecules can be accurately analyzed, including but not limited to 5' protruding jagged ends, 3' protruding jagged ends, 5' retracted jagged ends, 3' retracted jagged ends, terminal motifs of protruding jagged ends, terminal motifs of retracted jagged ends, genomic coordinates of fragment ends, fragment size, methylation-related cfDNA fragment omics features, and combinations thereof. In some embodiments, the terminal motif can be defined by one or more nucleotides spanning positions near the ends of the molecule. The terminal motif can be defined by one or more nucleotides of the genomic locus aligned around the end of the fragment in the reference genome. In other embodiments, the jagged ends can be defined by single-stranded DNA protruding from the ends of the DNA fragments. The jagged ends can be divided into different groups based on the length and / or chain of the protruding single-stranded DNA.
[0008] In some embodiments, different fragment-omics signatures from a single DNA fragment can be combined. In some embodiments, the combined fragment-omics signature can be used to detect or monitor cancer or other diseases. In other embodiments, the combined fragment-omics signature can be used for non-invasive prenatal testing.
[0009] A better understanding of the nature and advantages of embodiments of the present invention may be obtained by reference to the following detailed description and accompanying drawings.
[0010] BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 Shown is a schematic overview of the parallel analysis of single-molecule end morphology on a single-molecule real-time sequencing platform according to an embodiment of the present invention.
[0012] Figure 2A and 2B Shown is a schematic overview of the parallel analysis of single-molecule end morphology on a next-generation sequencing (NGS) platform (e.g., the Illumina platform) according to embodiments of the present invention.
[0013] Figure 3 is a flow chart of an exemplary process for analyzing a biological sample according to an embodiment of the present invention.
[0014] Figure 4 is a flow chart of an exemplary process for analyzing a biological sample according to an embodiment of the present invention.
[0015] Figure 5A and 5B The frequencies of different sawtooth tips according to embodiments of the present invention are shown.
[0016] Figure 6A Shown are graphs of the overall size distribution of plasma DNA samples from healthy and HCC subjects according to embodiments of the present invention.
[0017] Figure 6B Shown is a graph of the frequency of fragments less than 150 bp in size across combined jagged end classifications according to embodiments of the present invention.
[0018] Figure 6C Shown is a graph of the frequency of fragments larger than 280 bp in size across combined jagged end classifications according to embodiments of the present invention.
[0019] Figure 7A and 7B is a graph relating to size ratios of different saw tooth tips according to an embodiment of the present invention.
[0020] Figure 8 is a diagram of CCCA terminal motifs across different terminal types, according to an embodiment of the present invention.
[0021] Figure 9A and 9B A technique is shown in which jagged ends, 5' end motifs, and 3' end motifs can be combined to measure the terminal morphology stage of cfDNA molecules according to embodiments of the present invention.
[0022] FIG10 illustrates a technique for naming joint terminal motifs according to an embodiment of the present invention.
[0023] Figure 11A is a correlation plot of overall 5' end motif frequencies between HCC and healthy subjects.
[0024] Figure 11B is a correlation diagram of terminal motif frequencies by stage between HCC and healthy subjects according to an embodiment of the present invention.
[0025] Figure 11C is a correlation graph of joint terminal motif frequencies between HCC and healthy subjects according to an embodiment of the present invention.
[0026] Figures 12A-12F is a graph of the frequencies of different sawtooth end morphologies for different nuclear activities according to embodiments of the present invention.
[0027] Figure 13 TABLE 1 Table of median frequencies and relative changes of 5' A-, T-, C-, and G-termini in fragments with 5' overhanging jagged ends, 3' overhanging jagged ends, and blunt ends in WT, DNASE1L3- / -, DNASE1- / -, and DFFB- / - mice according to embodiments of the present invention.
[0028] Figures 14A-14Dis a DFFB according to an embodiment of the present invention - / - Terminal motif ranking diagram of (DFFB knockout [KO]) mice and wild-type (WT) mice.
[0029] Figure 15 is a flow chart of an exemplary process for analyzing a biological sample according to an embodiment of the present invention.
[0030] Figure 16 is a flow chart of an exemplary process for analyzing a biological sample according to an embodiment of the present invention.
[0031] Figure 17 is a flow chart of an exemplary process for analyzing a biological sample according to an embodiment of the present invention.
[0032] Figures 18A-18C The frequencies of different prominent jagged ends of fetal-specific and common cfDNA fragments according to embodiments of the present invention are shown.
[0033] Figure 18D is a graph of the fetal DNA fraction deduced from fragments having different types of protruding jagged ends, according to embodiments of the present invention.
[0034] Figure 19 is a graph of fetal DNA fraction relative to different jagged end morphologies, according to embodiments of the present invention.
[0035] Figure 20A and 20B is a graph of the fraction of fetal DNA deduced from fragments having certain jagged end morphologies and sequence end motifs, according to embodiments of the present invention.
[0036] Figure 21 is a flow chart of an exemplary process for enriching a biological sample for clinically relevant DNA, according to an embodiment of the present invention.
[0037] Figure 22A is a graph of DNASE1 mRNA expression levels in leukocytes and placenta according to an embodiment of the present invention.
[0038] Figure 22B is a graph of mRNA expression levels of DFFB in leukocytes and placenta according to embodiments of the present invention.
[0039] Figure 22C is a graph of the correlation between the fraction of fetal DNA carrying 5' overhanging jagged ends and the frequency of cfDNA fragments, according to embodiments of the present invention.
[0040] Figure 22D is a graph of the correlation between the fraction of fetal DNA carrying blunt ends and the frequency of cfDNA fragments according to an embodiment of the present invention.
[0041] Figure 23 is a flow chart of an exemplary process for determining the clinically relevant DNA fraction in a biological sample according to an embodiment of the present invention.
[0042] Figure 24 A measurement system according to an embodiment of the present invention is shown.
[0043] Figure 25 A computer system according to an embodiment of the present invention is shown.
[0044] the term
[0045] A "tissue" corresponds to a group of cells grouped together as a functional unit. More than one type of cell can be found in a single tissue. Different types of tissues can be composed of different types of cells (e.g., liver cells, alveolar cells, or blood cells), but can also correspond to tissues from different organisms (mother versus fetus) or to healthy cells versus tumor cells. A "reference tissue" can correspond to a tissue used to determine tissue-specific methylation levels. Multiple samples of the same tissue type from different individuals can be used to determine tissue-specific methylation levels for that tissue type.
[0046] An "organ" corresponds to a group of tissues with similar functions. One or more types of tissue can be found in a single organ. Organs can be part of different organ systems, including the cardiovascular, digestive, endocrine, excretory, lymphatic, integumentary, muscular, nervous, reproductive, respiratory, and skeletal systems.
[0047] A "biological sample" refers to any sample obtained from a subject (e.g., a human, such as a pregnant woman, a person with or suspected of having cancer, an organ transplant recipient, or a subject suspected of having a disease process involving an organ (e.g., the heart in myocardial infarction, the brain in stroke, or the hematopoietic system in anemia) and containing one or more target nucleic acid molecules. The biological sample can be a body fluid, such as blood, plasma, serum, urine, vaginal fluid, fluid from a hydrocele (e.g., testicular), vaginal washings, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple discharge, aspirates from various parts of the body (e.g., thyroid, breast), etc. Stool samples can also be used. In various embodiments, a majority of the DNA in a biological sample that has been enriched for cell-free DNA (e.g., a plasma sample obtained via a centrifugation protocol) can be cell-free, for example, greater than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA can be cell-free. The centrifugation protocol can include, for example, 3,000 g x 1000 rpm for 10 min. After 10 minutes, the fluid fraction is obtained and re-centrifuged again at, for example, 30,000 g for 10 minutes to remove residual cells.
[0048] A "sequence read" refers to a string of nucleotides sequenced from any portion or all of a nucleic acid molecule. For example, a sequence read can be a short string of nucleotides (e.g., 20-150) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of an entire nucleic acid fragment present in a biological sample. Sequence reads can be obtained in a variety of ways, for example, using sequencing techniques or using probes, such as in hybridization arrays or capture probes, or amplification techniques such as polymerase chain reaction (PCR) or linear amplification or isothermal amplification using a single primer.
[0049] "Ending position" or "end position" (or just "end") can refer to the genomic coordinates or genomic characteristics or nucleotide characteristics of the outermost bases, i.e., the extreme ends, of a free DNA molecule, such as a plasma DNA molecule. An end position can correspond to either end of a DNA molecule. In this way, if one refers to the start and end of a DNA molecule, both correspond to ending positions. In practice, an end position is the genomic coordinates or nucleotide characteristics of the outermost bases at one extreme end of a free DNA molecule as detected or determined by an analytical method such as, but not limited to, massively parallel sequencing or next generation sequencing, single molecule sequencing, double-stranded or single-stranded DNA sequencing library preparation protocols, polymerase chain reaction (PCR), or microarrays. Thus, each detectable end can represent a biologically true end, or an end that is one or more nucleotides inward, or one or more nucleotides extending from the original end of the molecule, for example, by 5' blunting and 3' filling of the overhang of a non-blunt-ended double-stranded DNA molecule by Klenow fragment. The genomic identity or genomic coordinate of the end position can be derived from alignment of sequence reads to a human reference genome, such as hg19. It can be derived from an index or code directory representing the original coordinates of the human genome. It can refer to the position or nucleotide identity on a free DNA molecule read by, but is not limited to, target-specific probes, mini-sequencing, or DNA amplification.
[0050] A "sequence motif" can refer to a short, repeating pattern of bases in a DNA fragment (e.g., a free DNA fragment). A sequence motif can occur at the end of a fragment and is therefore part of or includes a terminal sequence. A "terminal motif" can refer to a sequence motif of a terminal sequence that preferentially occurs at the end of a DNA fragment, perhaps for a particular type of tissue. A terminal motif can also occur just before or after the end of a fragment, thereby still corresponding to a terminal sequence. A nuclease can have a specific cleavage preference for a particular terminal motif, and a second most preferred cleavage preference for a second terminal motif.
[0051] The term "length of the overhang" between DNA strands can refer to a value that can be estimated by comparing the jaggedness (e.g., jaggedness index value) of total plasma DNA or plasma DNA within a certain fragment size range between a reference sample (e.g., normal cells) and a differentially regulated nuclease sample (e.g., tumor cells). In some cases, the length of the overhang varies based on a specific DNA fragment size range (e.g., 130-160 bp, 200-300 bp) selected for characterizing the biological sample.
[0052] In some embodiments, the length of the overhang in the DNA chain is a classification value that characterizes the overhang length between the two DNA chains. For example, a "long" overhang can include an overhang with 5 nt, 6 nt, 7 nt, 8 nt, 10 nt, 15 nt, 20 nt, 30 nt, 40 nt, 50 nt, 100 nt and a DNA chain greater than 100 nt size. A "short" overhang can include an overhang with 0 nt, 1 nt, 2 nt, 3 nt, 4 nt, 5 nt size. Additionally or alternatively, the clear length of the overhang in the DNA chain can be estimated based on the percentage of molecules with an overhang size exceeding a specific threshold. For example, the presence of a "long" overhang in plasma DNA can be expressed as a molecule greater than 5 nt, 6 nt, 7 nt, 8 nt, 10 nt, 15 nt, 20 nt, 30 nt, 40 nt, 50 nt, 100 nt, or a combination thereof.
[0053] A "calibrator sample" can correspond to a biological sample whose fractional concentration of clinically relevant DNA (e.g., tissue-specific DNA fraction) is known or determined via a calibration method, for example, using tissue-specific alleles, such as in transplantation in pregnant subjects, whereby alleles present in the donor genome but absent in the recipient genome can be used as markers for the transplanted organ. As another example, a calibrator sample can correspond to a sample from which end motifs can be determined. Calibrator samples can serve both purposes.
[0054] A "calibration data point" comprises a "calibration value" and a measured or known characteristic of a sample or object, e.g., age or tissue-specific fraction (e.g., fetal or tumor). A calibration value can be a relative abundance determined for a calibration sample whose characteristic is known. A calibration data point can comprise a calibration value (e.g., a sawtooth end value, also known as a protrusion index) and a known (measured) characteristic. Calibration data points can be defined in a variety of ways, e.g., as discrete points or as a calibration function (also known as a calibration curve or calibration surface). The calibration function can be derived from additional mathematical transformations of the calibration data points. The calibration function can be linear or nonlinear.
[0055] A "site" (also called a "genomic site") corresponds to a single position, which can be a single base position or a group of related base positions, e.g., a CpG site or a larger group of related base positions. A "locus" can correspond to a region that includes multiple sites. A locus can include only one site, which would make the locus equivalent to a site in this context.
[0056] "Isolation value" corresponds to a difference or ratio involving two values, for example, two fractional contributions or two methylation levels. Isolation value can be a simple difference or ratio. For example, the direct ratio of x / y is an isolation value, as well as x / (x+y). Isolation value can include other factors, for example, multiplication factors. As another example, the difference or ratio of a function of values can be used, for example, the difference or ratio of the natural logarithms (ln) of two values. Isolation value can include a difference or ratio.
[0057] As used herein, the term "classification" refers to any number or other character that relates to a particular characteristic of a sample. For example, a "+" symbol (or the word "positive") can indicate that a sample is classified as having a deletion or an amplification. Classifications can be binary (e.g., positive or negative) or have more levels of classification (e.g., a scale from 1 to 10 or from 0 to 1). The terms "cutoff" and "threshold" refer to predetermined numbers used in an operation. For example, a threshold size can refer to a size above which fragments are excluded. A threshold value can be a value above or below which a particular classification is applied. Either of these terms can be used in either context.
[0058] The term "parameter" as used herein means a numerical value that characterizes a quantitative data set and / or a numerical relationship between quantitative data sets. For example, the ratio (or a function of the ratio) between a first amount of a first nucleic acid sequence and a second amount of a second nucleic acid sequence is a parameter.
[0059] The terms "cutoff value" and "threshold value" refer to predetermined numbers used in an operation. For example, a threshold size can refer to a size above which fragments are excluded. A threshold value can be a value above or below which a particular classification is applied. Any of these two terms can be used in either context. A cutoff value or threshold value can be a "reference value," or derived from a reference value representing a particular classification or distinguishing between two or more classifications. As will be appreciated by those skilled in the art, such a reference value can be determined in various ways. For example, a metric can be determined for two different object cohorts with different known classifications, and a reference value can be selected as a representative of a classification (e.g., an average value) or as a value between two clusters of a metric (e.g., selected to obtain desired sensitivity and specificity). As another example, a reference value can be determined based on statistical analysis or simulation of a sample. Specific values of a cutoff value, threshold value, reference value, etc. can be determined based on desired accuracy (e.g., sensitivity and specificity).
[0060] "Pregnancy-related disorder" includes any disorder characterized by abnormal relative expression levels of genes in maternal and / or fetal tissues or abnormal clinical features in the mother and / or fetus. These conditions include, but are not limited to, preeclampsia (Kaartokallio et al. Sci Rep. 2015;5:14107; Medina-Bastidas et al. Int J Mol Sci. 2020;21:3597), intrauterine growth restriction (Faxén et al. Am J Perinatol. 1998;15:9-13; Medina-Bastidas et al. Int J Mol Sci. 2020;21:3597), placenta accreta, preterm birth (Enquobahrie et al. BMC Pregnancy Childbirth. 2009;9:56), hemolytic disease of the newborn, placental insufficiency (Kelly et al. Endocrinology. 2017;158:743-755), hydrops fetalis (Magor et al. Blood. 2015;125:2405-17), and fetal malformations (Slonim et al. et al. Proc Natl Acad Sci USA. 2009;106:9425-9), HELLP syndrome (Dijk et al. J Clin Invest. 2012;122:4003-4011), systemic lupus erythematosus (Hong et al. J Exp Med. 2019;216:1154-1169), and other maternal immune diseases.
[0061] "Pathology grade" (or disease grade or condition grade) can refer to the amount, extent or severity of the pathology associated with an organism. One example is a cellular disorder expressing a nuclease. Another example of pathology is the rejection of a transplanted organ. Other exemplary pathologies can include autoimmune attacks (e.g., lupus nephritis or multiple sclerosis that damages the kidneys), inflammatory diseases (e.g., hepatitis), fibrotic processes (e.g., cirrhosis), fatty infiltration (e.g., fatty liver disease), degenerative processes (e.g., Alzheimer's disease) and ischemic tissue damage (e.g., myocardial infarction or stroke). The health status of an object can be considered as a classification without pathology. The pathology can be cancer.
[0062] The term "cancer grade" can refer to whether cancer exists (i.e., exists or does not exist), the stage of cancer, the size of the tumor, whether there is metastasis, the total tumor burden of the body, the response of cancer to treatment, and / or other measurements of cancer severity (e.g., cancer recurrence). Cancer grade can be a number or other mark, such as a symbol, letter, and color. Grade can be zero. The level of cancer can also include pre-malignant or precancerous conditions (states). Cancer grade can be used in a variety of ways. For example, screening can be used to check whether cancer exists in people who were not previously aware of having cancer. Assessment can investigate people who have been diagnosed with cancer to monitor the progression of cancer over time, study the effectiveness of treatment, or determine prognosis. In one embodiment, prognosis can be expressed as the chance of a patient dying from cancer, or the chance of cancer progressing after a specific duration or time, or the chance or degree of cancer metastasis. Detection can mean "screening," or can mean checking whether a person with cancer-implying features (e.g., symptoms or other positive tests) suffers from cancer.
[0063] The abbreviation "bp" refers to base pairs. In some cases, "bp" can be used to indicate the length of a DNA fragment, even if the DNA fragment may be single-stranded and does not include base pairs. In the context of single-stranded DNA, "bp" can be interpreted as providing the length in units of nucleotides.
[0064] The abbreviation "nt" refers to nucleotides. In some cases, "nt" can be used to represent the length of single-stranded DNA in base units. In addition, "nt" can be used to represent relative positions, such as upstream or downstream of the analyzed locus. For double-stranded DNA, "nt" can still refer to the length of the single strand, rather than the total number of nucleotides in the two strands, unless the context clearly indicates otherwise. In some contexts involving technical conceptualization, data representation, processing and analysis, "nt" and "bp" can be used interchangeably.
[0065] The term "serrated end" can refer to the sticky end of DNA, an overhang of DNA, a protrusion of a strand, or that a double-stranded DNA includes a DNA strand that is not hybridized with another strand of DNA. A "serrated end value" is a measure of the degree of serrated end. The serrated end value can be proportional to the length of one strand of the double-stranded DNA that protrudes beyond the second strand. The serrated end value of multiple DNA molecules can include consideration of blunt ends in the DNA molecules.
[0066] In some cases, the jagged end value can provide a collective measure of the strands that protrude beyond other strands in a plurality of free DNA molecules. The collective measure of jaggedness can be determined based on the estimated lengths of the protruding strands in a plurality of free DNA molecules, for example, an average, median, or other collective measure of the individual measurements of each free DNA molecule. In some cases, the collective measure of jaggedness is determined for a specific fragment size range (e.g., 130-160 bp, 200-300 bp).
[0067] The term "size ratio" can refer to the amount of free DNA molecules within a specific fragment size range. The size ratio can be proportional to the amount of free DNA molecules within a specific fragment size range, which is normalized by another amount of free DNA molecules within another specific fragment size range. When another specific fragment size range refers to all size ranges, the term "size frequency" can be used.
[0068] The term "alignment" and related terms can refer to matching a sequence with a reference sequence. A reference sequence can be a sequence of a reference genome (e.g., the human genome) or a specific molecule. Such a reference sequence can include at least 100 kb, 1 Mb, 10 Mb, 50 Mb, 100 Mb and more. This alignment method cannot be performed manually and is performed by specialized computer software. The alignment may involve long and multiple sequences (e.g., at least 1,000, 10,000, 100,000, 1 million, 10 million or 100 million sequences). In addition, the alignment may involve variability in the sequence itself or errors in the sequence reading. Therefore, an alignment with this variability or error may not need to be precisely matched to the reference sequence.
[0069] The term "real time" can refer to a computing operation or process that is completed within a certain time limit. The time limit can be 1 minute, 1 hour, 1 day, or 7 days.
[0070] The term "subsequence" may refer to a string of bases that is smaller than the complete sequence corresponding to a nucleic acid molecule. For example, when the complete sequence of a nucleic acid molecule includes 5 or more bases, a subsequence may include 1, 2, 3, or 4 bases. In some embodiments, a subsequence may refer to a string of bases that forms a unit, wherein the unit is repeated multiple times in a tandem manner. Examples include 3-nt units or subsequences repeated at the locus associated with a trinucleotide repeat disorder, 1-nt to 6-nt units or subsequences repeated 5-50 times as microsatellites, 10-nt to 60-nt units or subsequences repeated 5-50 times as minisatellites, or in other genetic elements, such as Alu repeats.
[0071] "Clinically relevant DNA" can refer to DNA of a specific tissue origin to be measured, for example, to determine the fractional concentration of such DNA or to classify the phenotype of a sample (e.g., plasma). Examples of clinically relevant DNA are fetal DNA in maternal plasma or tumor DNA in patient plasma or other samples with free DNA. Another example includes measuring the amount of transplant-related DNA in the plasma, serum, or urine of a transplant patient. Another example includes measuring the fractional concentration of hematopoietic and non-hematopoietic DNA in the plasma of the subject, or the fractional concentration of liver DNA fragments (or other tissues) in the sample or the fractional concentration of brain DNA fragments in the cerebrospinal fluid.
[0072] The term "parallel analysis" can refer to the use of more than one fragmentomic feature. Using only the 5' end motif from one end of a nucleic acid molecule or only the serrated end morphology (e.g., 5' end overhang) would not be a parallel analysis. However, using a combination of serrated end morphologies from both ends of the molecule, a combination of a serrated end morphology and a sequence end motif from one end, or a combination of a serrated end morphology and a sequence end motif from both ends would be part of a parallel analysis.
[0073] The term "about" or "approximately" can mean within an acceptable error range for a particular value determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, according to the practice in the art, "about" can mean within 1 or more than 1 standard deviation. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term "about" or "approximately" can mean within an order of magnitude of a value, within 5 times, and more preferably within 2 times. Where specific values are described in the application and claims, unless otherwise indicated, it should be assumed that the term "about" means within an acceptable error range for a particular value. The term "about" can have the meaning commonly understood by one of ordinary skill in the art. The term "about" can refer to ± 10%. The term "about" can refer to ± 5%.
[0074] In the case of providing a range of values, it should be understood that unless the context clearly indicates otherwise, each intermediate value between the upper and lower limits of the scope is also specifically disclosed, to one-tenth of the lower limit unit. Any of the values described in the described range or the intermediate value, and each smaller range between any other described value in the described range or the intermediate value are included in the embodiments of the present disclosure. The upper and lower limits of these smaller ranges can be independently included in the scope or excluded from the scope, and each range in which any one, both or both limits are included in the smaller range is also included in the present disclosure, subject to the limitation of any clearly excluded limit in the described scope. In the case where the described scope includes one or two limits, the scope of excluding any one or both of the limits included is also included in the present disclosure.
[0075] Standard abbreviations may be used, eg, bp, base pairs; kb, kilobase; pi, picoliter; s or sec, seconds; min, minutes; h or hr, hours; aa, amino acid; nt, nucleotide; etc.
[0076] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the present disclosure, some potential and exemplary methods and materials are now described. Detailed Description of the Invention
[0078] Cell-free DNA (cfDNA) molecules are nonrandomly fragmented, and the fragmentation patterns of cfDNA molecules contain rich molecular information. For example, the characteristic size signature of cfDNA shows a pattern frequency of approximately 166 bp, with smaller molecules forming a series of peaks with a 10-bp periodicity (Lo et al. Sci Transl Med. 2010;2:61ra91). This size pattern of plasma DNA fragments suggests that both internucleosomal and intranucleosomal fragmentation occurs during the release of DNA molecules into the circulation during cell death and / or apoptosis. Furthermore, our group has previously reported that a subset of genomic locations are preferentially cleaved during the production of plasma DNA molecules (Chan et al. Proc Natl Acad Sci USA. 2016;113:E8159-E8168; Jiang et al. Proc Natl Acad Sci USA. 2018;115:E10925-E10933); this preferential cleavage may reflect the tissue of origin of the cfDNA (Jiang et al. Proc Natl Acad Sci USA. 2018;115:E10925-E10933; Sun et al. Proc Natl Acad Sci US A. 2018;115:E5106-E5114). Furthermore, we have shown that different nucleases associate with cell-free DNA molecules with characteristic terminal features, namely 5' terminal motifs and 5' overhanging jagged ends (Serpas et al. Proc Natl Acad Sci USA. 2019;116:641-649; Han et al. Am J Hum Genet. 2020;106:202-214, Ding et al. Clin Chem. 2022;68:917-926). The 5' terminal motif represents the sequence context at the 5' end of the cfDNA fragment. The 5' overhanging jagged ends represent the 5' overhanging single-stranded DNA in the cfDNA molecule. Recently, cell-free DNA end features have shown promising results as liquid biopsy biomarkers (Jiang et al. Cancer Discov. 2020;10:664-673; Jiang et al. Genome Res. 2020;30:1144-1153). Serrated tips are also described in US 2020 / 0056245 A1 and US 2022 / 0177971 A1, the entire contents of both patents are incorporated herein by reference for all purposes.
[0079] In contrast to the extensively studied 5' terminal motifs and 5' overhanging serrated ends, the actual 3' terminal motifs and 3' overhanging serrated ends have not been adequately studied, primarily due to artifactual modifications that occur during sequencing library preparation. Typical library preparation methods include an end-repair step. During this end-repair process, the 3' overhanging serrated ends are removed, and the opposing 5' overhanging serrated ends are used as DNA templates to extend the 3' retracted ends. Consequently, the original 3' ends are modified, resulting in alterations in the nucleotide information near the 3' terminal motif and loss of the 3' overhanging serrated ends. Furthermore, the 3' overhanging serrated ends are removed to form blunt ends. Due to this end-repair step, the blunt end information derived from typical library preparation methods is unreliable.
[0080] Recently, a group developed an NGS library preparation method, called the XACTLY assay, to investigate jagged ends via a sequence-based adaptor ligation approach (Harkins et al. Nucleic Acids Res. 2020; 48:e47). For example, Harkins Kincaid et al. ligated Y-shaped adaptors containing 7nt barcodes (i.e., unique end identifiers (UEIs) that indicate discrete end types and lengths) directly to the original DNA template without an end-repair step (Harkins Kincaid et al. Nucleic Acids Res. 2020; 48:e47). The ligated products were sequenced using short reads (i.e., an Illumina sequencing platform). Presumably, the native ends of the DNA molecules could be inferred from the sequence information of the 7nt barcodes attached to the reads. However, as reported in a study by Harkins et al., the 5' UEI (i.e., Illumina P5 adapter) has much lower fidelity than the 3' UEI (i.e., Illumina P7 adapter) (Harkins Kincaid et al. Nucleic Acids Res. 2020; 48: e47). This inaccuracy associated with the P5 adapter is because the first ligation of the P5 adapter to the template DNA will occur regardless of whether the overhang of the adapter correctly matches the overhang on the template DNA, while the P7 adapter ligation may only occur when the first ligation event is correct. In other words, P5 adapter ligation will occur even when gaps (i.e., regions where one strand does not have complementary nucleotides on the other strand) and flaps (i.e., regions where one nucleotide of the adapter does not hybridize to the strand) are present in the hybridization region between the template DNA and the adapter sequence. In sequencing libraries prepared by the XACTLY assay, the double-stranded DNA will be denatured into two single-stranded DNA molecules. The 5' and 3' ends of this single-stranded DNA molecule are labeled with P5 and P7, respectively. Therefore, the XACTLY assay has inherent limitations:
[0081] 1. At least one end cannot be analyzed in an accurate manner using the XACTLY assay.
[0082] 2. It is not possible to effectively analyze the two strands of a DNA molecule simultaneously.
[0083] 3. Unable to measure the true length of DNA molecules.
[0084] In the present disclosure, we have developed new methods to simultaneously and with high fidelity detect the natural fragment omics features of cfDNA molecules, including fragment size, terminal motifs, and jagged ends. In one embodiment, using DNA ligase, double-stranded cfDNA fragments can be appropriately ligated with a pair of hairpin adapters, depending on the terminal morphology of such double-stranded cfDNA molecules, to form circular DNA molecules. Such hairpin adapters contain molecular barcodes and carry jagged ends or blunt ends of different lengths. Different cfDNA fragments will be ligated with hairpin adapters containing different molecular barcodes, corresponding to jagged end lengths (e.g., 1-50 nt) and jagged types (blunt ends, 5' protruding jagged ends, 3' protruding jagged ends, and combinations thereof). The ligation products can be treated with enzymes to remove incomplete circular DNA molecules, thereby enriching the desired circular DNA molecules produced by hairpin adapter-mediated DNA ligation (i.e., negative selection step).
[0085] The product of the enrichment circular DNA molecule can further be subjected to direct enrichment of circular DNA molecules, such as single molecule real-time sequencing (for example, Pacific Biosciences) and rolling circle cycle amplification (i.e., positive selection step), to minimize the impact of inaccurate connection. Since only one complete circular DNA molecule can be sequenced multiple times, thereby generating sub-readings in the single molecule real-time sequencing process, the readout with three or more sub-readings is selected to allow the exclusion of incomplete circular DNA molecules. Similarly, only complete circular DNA molecules can be amplified via rolling circle amplification. In some embodiments, enzymes include but are not limited to exonuclease I, exonuclease II, exonuclease III, exonuclease IV, exonuclease V, exonuclease VI, exonuclease VII or exonuclease VIII. In another embodiment, the negative selection step and the positive selection step can be performed alone or in combination. After sequencing, the natural jagged ends of single cfDNA fragments can be derived by analyzing the barcode sequence. Once the jagged ends at each end of a cfDNA molecule are determined, the entire fragmentomic profile of the cfDNA molecule can be accurately analyzed at a 1-nt resolution, including but not limited to 5' protruding jagged ends, 3' protruding jagged ends, 5' retracted jagged ends, 3' retracted jagged ends, terminal motifs of protruding jagged ends, terminal motifs of retracted jagged ends, genomic coordinates of fragment ends, fragment size, methylation-related cfDNA fragmentomic features, and combinations thereof. In one embodiment, the length difference between the Watson chain and the Crick chain can be used as another type of fragmentomic feature.
[0086] In some embodiments, the terminal motif can be defined by one or more nucleotides that span the position at or near the end of the molecule. A molecule can have 4 ends. The terminal motif can be defined by one or more nucleotides of the genomic locus aligned around the end of the fragment in the reference genome. In other embodiments, the jagged end can be defined by the single-stranded DNA protruding from the end of the DNA fragment. The jagged end can be divided into different groups according to the length and chain of the protruding single-stranded DNA. In other embodiments, different fragment omics features from a DNA fragment can be combined. In one embodiment, the combined fragment omics features can be used to detect or monitor cancer or other diseases. In another embodiment, the combined fragment omics features can be used for non-invasive prenatal testing. The terminal motif is described in US 2021 / 0238668 A1, the entire content of which is incorporated herein by reference.
[0087] I. Principles of Parallel Analysis of Single-Molecular Terminal Morphology
[0088] Figure 1 A schematic overview of the parallel analysis of single-molecule end morphology on a single-molecule real-time sequencing platform (e.g., the Pacific Biosciences (PacBio) platform) is shown. Stage 104 shows different cfDNA molecules containing different jagged or blunt ends. Exemplary fragment 108 has a blunt end on the left and a 3-nt 5' overhanging jagged end on the right.
[0089] Stage 112 shows different hairpin aptamers in a library of hairpin aptamers. The library of hairpin aptamers contains aptamers with blunt ends and aptamers with jagged ends (also called overhangs). Each hairpin aptamer with a jagged end has a protruding single-stranded end of different length (indicated by the number of "N" in the overhang 116). A barcode sequence synthesized with the hairpin aptamer that is compatible with the PacBio sequencing platform can be used to indicate the jagged end type (e.g., 5' or 3' overhang) and the jagged end length (indicated by rectangles 120 and 124 filled with different patterns).
[0090] At stage 128, the cfDNA molecules are ligated to hairpin adapters. Fragment 108 is ligated at its blunt end (left) to hairpin adapter 132. Fragment 108 is ligated at its 3-nt 5' overhanging jagged end (right) to hairpin adapter 136. Appropriate ligation produces molecule 140.
[0091] Other molecules can result from ligation. Molecule 144 represents a fragment with a hairpin adapter attached to only one end. Molecule 148 represents a fragment without a hairpin adapter. Molecule 152 has a hairpin adapter correctly attached to the blunt end. However, molecule 152 has a hairpin adapter incorrectly attached to the 5' overhang, resulting in a gap between the cfDNA fragment and the hairpin adapter. Molecule 156 has a hairpin adapter correctly attached to the blunt end. However, molecule 156 has a hairpin adapter incorrectly attached to the 5' overhang, resulting in a curled-up hairpin, where the nucleotides of the hairpin adapter do not hybridize with the original cfDNA fragment.
[0092] At stage 160, the adaptor-ligated molecules can be treated with one or more enzymes capable of digesting incomplete circular DNA molecules (e.g., molecules 152 and 156). Enzymatic digestion of incomplete adaptor-ligated molecules can be referred to as negative selection because incorrectly ligated molecules are selected and removed.
[0093] At stage 164, the enzymatically treated ligation products can be sequenced on the PacBio platform. Only when the cfDNA fragments with two ends are properly ligated with hairpin adapters corresponding to naturally jagged ends to form intact circular DNA (e.g., by rolling circle amplification) can such circular DNA products be sequenced to generate multiple subreads per strand. Amplifying and / or sequencing only intact adapter-ligated molecules can be referred to as positive selection, as correctly ligated molecules are selected and further analyzed.
[0094] At stage 168, the sequence is analyzed. After sequencing, the barcode sequence information at both ends can be read to infer the presence of jagged and / or blunt ends, as well as the type and length of the jagged ends (if present). Based on the inferred ends, the 5' end motif, 3' end motif, and / or the size of each strand of the cfDNA fragment can be further detected. This allows the natural fragmentomic characteristics of the original cfDNA molecule to be assessed.
[0095] Figure 2A and 2B A schematic overview of the parallel analysis of single molecule end morphology on a next generation sequencing (NGS) platform (e.g., an Illumina platform) is shown. Circular cfDNA molecules can be prepared using modified hairpin adapters according to embodiments of the present disclosure. Stage 104 can be Figure 2A Stage 112 can include a modified hairpin aptamer containing a cleavage site for a restriction enzyme. For example, cleavage sites 204 and 208 can be included in the hairpin aptamer.
[0096] exist Figure 2BIn the process, similar to stage 128, DNA fragments are ligated with hairpin adapters, and similar to stage 160, the adapter-ligated molecules are treated with one or more enzymes capable of digesting incomplete circular DNA molecules (i.e., negative selection). Similar to stage 160, the enzyme-treated ligation products are amplified by rolling circle amplification (i.e., positive selection). Only cfDNA fragments whose ends are properly ligated to hairpin adapters corresponding to naturally jagged / blunt ends are amplified.
[0097] In stage 250, the rolling circle amplification product is treated with a designated restriction enzyme to cut at the cleavage site in the hairpin adaptor. Thus, the large DNA molecule generated via rolling circle PCR is cut into small DNA molecules, which are suitable for Illumina sequencing or other similar sequencing.
[0098] Sequencing adapters are ligated to the cleaved small DNA molecules at stage 254. The sequencing adapters are configured for Illumina sequencing.
[0099] Figure 2B The analysis in Figure 1 The stage 168 is similar.
[0100] A. Exemplary Positive Selection Methods
[0101] Figure 3 3 ' is a flow chart of an exemplary process 300 for analyzing a biological sample. The process 300 can determine whether jagged ends are present at both ends of a cfDNA molecule, whether the 5' or 3' end is an overhang, the length of the overhang, and / or the sequence of the overhang. A strand that overhangs beyond another strand can be considered an overhang. In some embodiments, Figure 3 One or more process blocks of can be performed by a system, including system 2400. The biological sample can include a plurality of nucleic acid molecules. The nucleic acid molecules can be free and double-stranded, having a first strand and a second strand.
[0102] At block 302, for each nucleic acid molecule in a plurality of nucleic acid molecules, a first hairpin adapter is connected to the first strand of the nucleic acid molecule and the second strand of the nucleic acid molecule at the first end of the nucleic acid molecule. The first hairpin adapter may include a first sequence identifier. The first sequence identifier may identify zero or more nucleotides of a first length at the first end of the first hairpin adapter that do not have a complementary portion at the second end of the first hairpin adapter. The length of the nucleotides in the hairpin that do not have a complementary portion corresponds to the length of the jagged end of the nucleic acid molecule. Zero nucleotides represent a blunt end and wherein the hairpin adapter ends are complementary. For example, the first sequence identifier may encode a length of 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. The first sequence identifier may encode whether the 3' strand or 5' strand of the nucleic acid molecule connected to the first hairpin protrudes beyond the other strand. In some embodiments, the first sequence identifier may encode a subsequence of zero or more nucleotides at the first end of the first hairpin adapter that do not have a complementary portion at the second end of the first hairpin adapter.
[0103] The first hairpin adaptor may include Figure 1 The first sequence identifier may be nucleotides represented by rectangles 120 and 124. The length of zero or more nucleotides at the first end of the first hairpin aptamer that does not have a complementary portion at the second end of the first hairpin aptamer may include an overhang 116.
[0104] At box 304, for each nucleic acid molecule in a plurality of nucleic acid molecules, a second hairpin adapter is connected to the first chain and the second chain at the second end of the nucleic acid molecule. The second hairpin adapter may include a second sequence identifier. The second sequence identifier may identify zero or more nucleotides of a second length that do not have a complementary portion at the second end of the second hairpin adapter at the first end of the second hairpin adapter. The second sequence identifier may have similar characteristics to the first sequence identifier. The first sequence identifier and the second sequence identifier may use similar encodings. Certain predetermined subsequences in the sequence identifier may correspond to different numbers. Utilizing four nucleotides (A, T, G, C), the length may be represented by a digital base 4. Multiple connected nucleic acid molecules are generated after connection.
[0105] In some examples, negative selection can be performed. An exonuclease can be added to the plurality of connected nucleic acid molecules after connecting the plurality of first hairpin aptamers and the plurality of second hairpin aptamers to remove an incorrectly connected subset of the plurality of connected nucleic acid molecules. For each nucleic acid molecule in the incorrectly connected subset, either the corresponding nucleic acid molecule does not completely hybridize with the corresponding first hairpin aptamer or the corresponding second hairpin aptamer (e.g., there is a "gap"), or the corresponding first hairpin aptamer or the corresponding second hairpin aptamer does not completely hybridize with the corresponding nucleic acid molecule (e.g., there is a "flip"). Negative selection can be similar to Figure 1 Stage 160, wherein molecules 144, 148, 152 and 156 are removed.
[0106] At block 306, rolling circle amplification can be performed on a first subset of the plurality of ligated nucleic acid molecules to form a plurality of concatemers. The first subset can include no nucleic acid molecules that are identical to nucleic acid molecules in the incorrectly ligated subset. Each nucleic acid molecule in the first subset can be ligated to a corresponding first hairpin adapter of the plurality of first hairpin adapters and a corresponding second hairpin adapter of the plurality of second hairpin adapters. Each nucleic acid molecule in the first subset can be correctly ligated to the hairpin adapter without gaps or flips, similar to Figure 1 Each nucleotide in the first set of nucleic acid molecule chains can hybridize with a complementary nucleotide in the other chain.
[0107] Each nucleic acid molecule in the first portion of the first subset can have a corresponding first strand at the corresponding first end that protrudes beyond the corresponding second strand. The first strand can be a 5' strand or a 3' strand at the first end. In some instances, each nucleic acid molecule in the second portion of the first subset can have a corresponding first strand even if it has a corresponding second strand at the corresponding first end. In some instances, each nucleic acid molecule in the second portion of the first subset has a corresponding second strand at the corresponding first end that protrudes beyond the corresponding first strand. The corresponding first strand can be a 5' strand. The corresponding second strand can be a 3' strand.
[0108] The first subset can include portions corresponding to different combinations of jagged end properties: DNA molecules containing 5' protruding jagged ends and 3' protruding jagged ends (5-3); 5' protruding jagged ends and 5' protruding jagged ends (5-5); 3' protruding jagged ends and 3' protruding jagged ends (3-3); 5' protruding jagged ends and blunt ends (5-B); 3' protruding jagged ends and blunt ends (3-B); and blunt ends and blunt ends (BB).
[0109] At box 308, each concatemer in the multiple concatemers is sequenced to identify a corresponding first sequence identifier and a corresponding second sequence identifier. The first sequence identifier and the second sequence identifier can each include a subsequence of nucleotides indicating that continuous nucleotides are a part of the identifier. Order-checking can be by unimolecular, real-time sequencing, next generation sequencing or any suitable sequencing technology. Order-checking can occur simultaneously with carrying out rolling circle amplification.
[0110] The first sequence identifier can be used to determine the length of the overhang present at the first end of the nucleic acid molecules in the first subset of a plurality of connected nucleic acid molecules. The first sequence identifier can include a subsequence corresponding to the overhang length. In addition, the first sequence identifier can include a subsequence indicating whether the overhang exists on the current strand or the complementary strand.
[0111] The second sequence identifier can be used to determine the length of the overhang present at the second end of the nucleic acid molecules in the first subset of the plurality of linked nucleic acid molecules.The second sequence identifier can be used in a similar manner to the first sequence identifier.
[0112] In some instances, the first sequence end motif of the overhang present at the first end of the nucleic acid molecule in the first subset of the nucleic acid molecules that are connected can be determined using the sequence of the first sequence identifier. The first sequence identifier can indicate which chain is overhanging at the end, and appropriate subsequences may be associated with the overhang. In addition, the first sequence identifier indicates the length of the overhang, so the entire sequence of the overhang can be determined. In some embodiments, the entire sequence of the overhang may not be determined, and on the contrary, an end motif (2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides) can be determined. In some instances, the second sequence end motif of the overhang present at the second end of the nucleic acid molecule in the first subset of the nucleic acid molecules that are connected can be determined using the sequence of the second sequence identifier.
[0113] In some examples, the sequence of the corresponding first identifier can be used to determine, for each nucleic acid molecule in the first subset having an overhang at the first end, whether the 5' strand or the 3' strand protrudes beyond the other strand. In some examples, the sequence of the corresponding second identifier can be used to determine, for each nucleic acid molecule in the first subset having an overhang at the second end, whether the 5' strand or the 3' strand protrudes beyond the other strand.
[0114] In some examples, each first hairpin aptamer in the plurality of first hairpin aptamers can include a first cleavage site. Each second hairpin aptamer in the plurality of second hairpin aptamers can include a second cleavage site. The process can include cleaving each concatemer in the plurality of concatemers at the corresponding first cleavage site and the corresponding second cleavage site.
[0115] Process 300 can be used for determining length or end motif in other processes disclosed herein.In some instances, each nucleic acid molecule in the multiple molecules has a size greater than a first threshold size.In some instances, each nucleic acid molecule in the multiple molecules has a size less than a second threshold size.The first threshold size and the second threshold size can independently be 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340 or 350.The size of each nucleic acid molecule can be determined by comparing the subsequence corresponding to the end of the corresponding nucleic acid molecule with the reference genome.
[0116] The condition can be cancer (such as, but not limited to, HCC and colorectal cancer [CRC]), an autoimmune disease (e.g., systemic lupus erythematosus), a pregnancy-related disorder, or any condition described herein. Reference values can be determined from one or more subjects with a certain grade of the condition or from one or more healthy subjects.
[0117] In some instances, the grade of the condition is not determined. Instead, a comparison can be used to determine the fractional concentration of clinically relevant DNA. A reference value can be determined from one or more subjects with known fractional concentrations of clinically relevant DNA. The reference value can be a calibration value determined using a calibration sample.
[0118] In some instances, the reading corresponding to a plurality of nucleic acid molecules can be enriched for clinically relevant DNA. For example, a biological sample can be obtained from a female subject who is pregnant with a fetus. The method can also include selecting the reading corresponding to a subset of nucleic acid molecules with a 5' chain or a 3' chain that protrudes beyond the other end. The method can include analyzing a subset of nucleic acid molecules for the characteristics of the fetus. For example, a feature can be that there is a distortion (e.g., a mutation, aneuploidy) in the fetal genome. As another example, readings can be enriched for maternal samples by selecting a reading with a blunt end at one end. Other clinically relevant DNAs can be enriched by analyzing the concentration of this DNA in different sawtooth end morphologies. The sawtooth end morphology of the clinically relevant DNA with higher concentration can be selected to produce an enriched data set. The end morphology can include the end morphology at both ends of any given fragment.
[0119] Process 300 may include additional implementations, such as any single implementation or any combination of implementations described herein and / or in conjunction with one or more other processes described elsewhere herein.
[0120] although Figure 3Example blocks of process 300 are shown, but in some implementations, process 300 may include Figure 3 The blocks may be more than, fewer than, different from, or arranged differently than those shown in process 300. Additionally or alternatively, two or more blocks of process 300 may be executed in parallel.
[0121] B. Exemplary Negative Selection Methods
[0122] Figure 4 4 is a flow chart of an exemplary process 400 for analyzing a biological sample. The process 400 can determine whether jagged ends are present at both ends of a cfDNA molecule, whether the 5' or 3' end is an overhang, the length of the overhang, and / or the sequence of the overhang. A strand that overhangs beyond another strand can be considered an overhang. In some embodiments, Figure 4 One or more process blocks of may be performed by system 2400 .
[0123] At block 402 , for each nucleic acid molecule in a plurality of nucleic acid molecules, a first hairpin adaptor is attached to a first strand of the nucleic acid molecule and a second strand of the nucleic acid molecule at a first end of the nucleic acid molecule. Block 402 may be performed in the same manner as block 302 .
[0124] At block 404 , for each nucleic acid molecule in the plurality of nucleic acid molecules, a second hairpin adaptor is attached to the first strand and the second strand at a second end of the nucleic acid molecule. Block 404 may be performed in the same manner as block 304 .
[0125] At block 406, an exonuclease is added to the plurality of ligated nucleic acid molecules to remove a first subset of the plurality of ligated nucleic acid molecules. For each nucleic acid molecule in the first subset, either the corresponding nucleic acid molecule does not fully hybridize to the corresponding first hairpin aptamer or the corresponding second hairpin aptamer, or the corresponding first hairpin aptamer or the corresponding second hairpin aptamer does not fully hybridize to the corresponding nucleic acid molecule.
[0126] At block 408, each of the connected nucleic acid molecules in the second subset of the plurality of connected nucleic acid molecules can be sequenced to identify a corresponding first sequence identifier and a corresponding second sequence identifier. The second subset is the connected nucleic acid molecules that remain in the biological sample after removing the first subset. Sequencing can be performed by next generation sequencing, single molecule real-time sequencing, or any sequencing technology described herein.
[0127] The first sequence identifier can be used to determine the length of the overhang present at the first end of the nucleic acid molecules in the second subset of a plurality of connected nucleic acid molecules. The first sequence identifier can include a subsequence corresponding to the overhang length. In addition, the first sequence identifier can include a subsequence indicating whether the overhang is on the current strand or the complementary strand.
[0128] The second sequence identifier can be used to determine the length of the overhang present at the second end of the nucleic acid molecules in the second subset of the plurality of linked nucleic acid molecules.The second sequence identifier can be used in a similar manner to the first sequence identifier.
[0129] In some instances, the first sequence end motif of the overhang that the first end of the nucleic acid molecule in the second subset of the nucleic acid molecules of multiple connections can be determined using the sequence of the first sequence identifier. The first sequence identifier can indicate which chain is overhanging at the end, and appropriate subsequence may be associated with the overhang. In addition, the first sequence identifier indicates the length of the overhang, so the entire sequence of the overhang can be determined. In some embodiments, the entire sequence of the overhang may not be determined, and instead the end motif (2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides) can be determined. In some instances, the second sequence end motif of the overhang that the second end of the nucleic acid molecule in the second subset of the nucleic acid molecules of multiple connections can be determined using the sequence of the second sequence identifier.
[0130] In some examples, the sequence of the corresponding first identifier can be used to determine, for each nucleic acid molecule of the second subset having an overhang at the first terminus, whether the 5' strand or the 3' strand protrudes beyond the other strand. In some examples, the sequence of the corresponding second identifier can be used to determine, for each nucleic acid molecule of the second subset having an overhang at the second terminus, whether the 5' strand or the 3' strand protrudes beyond the other strand.
[0131] In some examples, each first hairpin aptamer in the plurality of first hairpin aptamers can include a first cleavage site. Each second hairpin aptamer in the plurality of second hairpin aptamers can include a second cleavage site. The process can include cleaving each concatemer in the plurality of concatemers at the corresponding first cleavage site and the corresponding second cleavage site.
[0132] In some examples, the sample can be enriched for clinically relevant DNA, as explained with respect to process 300 .
[0133] Process 400 may include additional embodiments, such as any single embodiment or any combination of embodiments described herein and / or in conjunction with one or more other processes (including process 300) described elsewhere herein.
[0134] although Figure 4 Example blocks of process 400 are shown, but in some implementations, process 400 may include Figure 4 The blocks may be more than, fewer than, different from, or arranged differently than those shown in process 400. Additionally or alternatively, two or more blocks of process 400 may be executed in parallel.
[0135] II. Disease Analysis and Testing
[0136] A variety of conditions can be analyzed and detected. Cancer and nuclease activity defects are examples of conditions that can be analyzed and detected using jagged end morphology and / or sequence end motifs. Other conditions can also be analyzed and detected, including conditions characterized by abnormal nuclease activity.
[0137] A. Cancer
[0138] For illustrative purposes, sequencing libraries were prepared from plasma DNA samples from a healthy individual and a hepatocellular carcinoma (HCC) patient. These libraries were sequenced using the PacBio sequencing platform, yielding 59,133 and 227,198 cycle consensus sequencing (CCS) reads, respectively. Hairpin adapters with blunt ends, 5' overhanging jagged ends (length range 1-10 nt), and 3' overhanging jagged ends (length range 1-10 nt) were used.
[0139] 1. Detection of jagged ends using parallel analysis of single-molecule end morphology
[0140] It has been reported that the presence of 5' jagged overhangs is higher in HCC plasma DNA samples compared to healthy controls (Jiang et al., Genome Res. 2020;30:1144-1153). In embodiments, 5' jagged overhangs, 3' jagged overhangs, and blunt ends can be inferred from the parallel analysis of single-molecule end morphology in a more precise, accurate, and comprehensive manner, potentially improving diagnostic capabilities.
[0141] Figure 5A The frequency of 5' overhanging jagged ends, 3' overhanging jagged ends, and blunt ends in HCC and healthy subjects is shown. The y-axis shows the frequency. The x-axis shows the type of jagged or blunt end. Two different bar graphs show the difference between healthy subjects and subjects with HCC. Both sides of the jagged ends from a single cfDNA fragment were analyzed separately. The frequency is based on the total number of ends (two ends per molecule), not the total number of molecules.
[0142] like Figure 5AAs shown in Figure 3, the frequency of 5' protruding jagged ends was higher in HCC cases compared with healthy subjects (59.40% vs. 55.84%). A slight decrease was observed in 3' protruding jagged ends (23.51% vs. 25.57%) and blunt ends (17.09% vs. 18.59%) in HCC cases.
[0143] Figure 5B The frequency of molecules falling within the combined jagged end categories is shown. The y-axis displays the frequency as a percentage. The x-axis displays the different combined jagged end characteristics at each end of each molecule: populations of DNA molecules containing a 5' jagged overhang and a 3' jagged overhang (5-3); a 5' jagged overhang and a 5' jagged overhang (5-5); a 3' jagged overhang and a 3' jagged overhang (3-3); a 5' jagged overhang and a blunt end (5-B); a 3' jagged overhang and a blunt end (3-B); and a blunt end and a blunt end (BB). Two separate bar graphs are shown for HCC subjects and healthy subjects. Jagged ends from a single cfDNA fragment were analyzed simultaneously.
[0144] like Figure 5B As shown, compared with healthy samples, HCC cases showed higher amounts of cfDNA fragments belonging to the 5-5 (37.32% vs. 33.32%) and 5-B (18.75% vs. 17.65%) groups, but lower amounts of cfDNA fragments belonging to the 5-3 (25.41% vs. 27.39%) and BB groups (4.28% vs. 5.94%) groups. The cumulative difference between these categories between HCC and healthy subjects was higher when both ends were analyzed simultaneously (10.21%) compared to single ends (7.12%). These results suggest that parallel analysis of jagged end morphology from both sides of cfDNA fragments can provide more detailed information, not available with previously published techniques, and improve diagnostic performance.
[0145] 2. Detection of fragment sizes using parallel analysis of single-molecule terminal morphologies
[0146] In one embodiment, jagged end and fragment size inferred from parallel analysis of single molecule end morphology can be analyzed together with jagged end classification.
[0147] Figure 6A A graph showing the overall size distribution of plasma DNA samples from healthy and HCC subjects. The y-axis shows frequency in percentage. The x-axis shows size in base pairs. Fragment sizes are slightly shorter in HCC cases compared to healthy subjects.
[0148] Figure 6BA graph shows the frequency of fragments less than 150 bp across combined jagged end classifications. The y-axis represents the frequency of molecules within that classification with sizes less than 150 bp, as a percentage, compared to all sizes within that classification. The x-axis lists the combined jagged end classifications: 5' overhanging jagged end and 3' overhanging jagged end (5-3); 5' overhanging jagged end and 5' overhanging jagged end (5-5); 3' overhanging jagged end and 3' overhanging jagged end (3-3); 5' overhanging jagged end and blunt end (5-B); 3' overhanging jagged end and blunt end (3-B); and blunt end and blunt end (BB). The x-axis also lists all cfDNA fragments. Both ends from a single cfDNA fragment were analyzed simultaneously.
[0149] Figure 6C A plot showing the frequency of fragments larger than 280 bp across the combined jagged-end classification is shown. The y-axis is the frequency in percent. The x-axis lists the combined jagged-end classification and all cfDNA fragments.
[0150] like Figure 6B As shown in the “all” cfDNA fragments category of the , the frequency of short cfDNA fragments (<150 bp) was higher in HCC cases (24.58% vs. 15.82%). Figure 6C As shown in the “all” category of the data, the frequency of long cfDNA fragments (>280 bp) was lower in HCC cases compared with healthy subjects (14.81% vs. 24.94%).
[0151] We then classified cfDNA fragments into different groups based on the types of jagged ends at both ends. Figure 6B As shown, the populations with blunt-end cfDNA fragments (BB; 24.44% vs. 10.27%) and cfDNA fragments with 3' protruding jagged ends at both ends (3-3; 33.45% vs. 19.78%), as well as cfDNA fragments with both 3' protruding jagged ends and blunt ends (3-B; 26.38% vs. 15.09%) showed greater differences in the frequencies of short cfDNA fragments between HCC cases and healthy subjects compared with all cfDNA fragments (24.58% vs. 15.82%).
[0152] like Figure 6CAs shown, compared with all cfDNA fragments (14.81% vs. 24.94%), the populations with blunt ends at both ends (BB; 20.29% vs. 57.60%), cfDNA fragments with 5' protruding jagged ends and blunt ends (5-B; 11.61% vs. 25.83%), and cfDNA fragments with 3' protruding jagged ends and blunt ends (3-B; 15.78% vs. 28.11%) showed greater differences in the frequencies of long cfDNA fragments between HCC cases and healthy subjects.
[0153] Figure 7A This figure plots the ratio of short to long fragments for different types of jagged ends. The y-axis shows the ratio of short (<150 bp) to long (>280 bp) fragments. The x-axis shows the combined jagged end classification and all cfDNA fragments. Separate bar graphs show HCC cases and healthy individuals. Both ends of a single cfDNA fragment were analyzed simultaneously.
[0154] Figure 7B Figure 2 is a graph of the fold change in the short / long ratio of HCC cases compared to healthy controls. The y-axis is the fold change, calculated by dividing the short / long ratio of HCC cases by the short / long ratio of healthy controls. The x-axis is the combined jagged end classification and all cfDNA fragments.
[0155] When compared with all fragments (all; 1.65 vs. 0.63; fold change: 2.61), the difference in the short / long ratio (i.e., the amount of fragments <150 bp / the amount of fragments >280 bp) between HCC cases and healthy cases was increased in cfDNA fragments with blunt ends at both ends (BB; 1.20 vs. 0.18; fold change: 6.75), cfDNA fragments with 5' protruding jagged ends and blunt ends (5-B; 1.90 vs. 0.61; fold change: 3.13), and cfDNA fragments with 3' protruding jagged ends and blunt ends (3-B; 1.67 vs. 0.53; fold change: 3.11). Figure 7A and 7B It was shown that certain types of serrations or certain combinations of serrations can be as effective or more effective in distinguishing healthy cases from HCC cases than using segments regardless of their serration type.
[0156] 3. Detection of terminal motifs derived from parallel analysis of single-molecule terminal morphology
[0157] It has been reported that the presence of 5' CCCA terminal motifs is reduced in plasma DNA samples of patients with HCC compared to healthy subjects (Jiang et al. Cancer Discov. 2020;10:664-673). In the embodiment, the 5' CCCA terminal motifs in 5' protruding jagged ends, 3' protruding jagged ends, and blunt ends can be calculated separately.
[0158] Figure 8 This figure plots the CCCA terminal motif across different end types. The y-axis shows the CCCA frequency as a percentage. The x-axis shows the locations where the CCCA terminal motif was found: 5' jagged overhang, 3' jagged overhang, blunt ends, and all fragments. The two bar graphs show healthy individuals and HCC cases. Both ends from a single cfDNA fragment were analyzed separately.
[0159] like Figure 8 As shown, the frequency of the 5'CCCA motif decreased in HCC cases at either the 5' or 3' jagged ends of cfDNA. Regarding blunt ends, the frequency of the 5'CCCA motif increased in HCC cases compared to healthy subjects. The difference in the frequency of the 5'CCCA motif between HCC cases and healthy subjects for 5' jagged ends, 3' jagged ends, and blunt ends was greater than that for the 5'CCCA motif derived from all fragment ends. Figure 8 showed that determining the type of serrated end could improve the accuracy of distinguishing HCC cases from healthy ones.
[0160] Figure 9A and 9B A technique is shown that can combine jagged ends, 5' end motifs, and 3' end motifs to measure the stage of end morphology of cfDNA molecules. Figure 9A In the example, there is a 5' protruding serrated end with a 5' "CCCA" terminal motif and a 3' "TTTT" terminal motif. The stage of the terminal motifs can be referred to as "CCCA_TTTT", where the 5' terminal motif is followed by the 3' terminal motif, both of which are represented by uppercase letters, with an underscore (i.e., "_") as a separator.
[0161] exist Figure 9BIn the example, there is a 3' protruding jagged end with a 5' "CCCA" end motif and a 3' "GAGG" end motif, the stage of the end motif can be referred to as "CCCA_gagg", where the 5' end motif in uppercase letters is followed by the 3' end motif in lowercase letters. Lowercase letters indicate that the 3' end is a protruding end. Different naming conventions can be used to indicate protruding ends, non-protruding ends, and whether the end is blunt. In some embodiments, the end motif can only include nucleotides from one chain, because information about the other chain can be deduced from only one chain. Delimiters can be used to indicate the position of the protruding part. As an example, Figure 9B It can be represented by "3-GAG-GGGT." The "3" indicates that the 3' end is protruding, and the second "-" indicates the position where the 5' end of the chain begins.
[0162] Figure 10A and 10B A technique is shown that can combine jagged ends, 5' end motifs, and 3' end motifs from both sides of a fragment to measure the combined end morphology of a cfDNA molecule. Figure 10A In the present invention, there is a DNA fragment having a 5' protruding serrated end with a 5' "C" terminal motif and a 3' "G" terminal motif on the left side, and a 3' protruding serrated end with a 5' "G" terminal motif and a 3' "T" terminal motif on the right side. The combined terminal motif may be referred to as "5CG3GT", where the first 3 characters indicate the left end, and the subsequent 3 characters indicate the right end.
[0163] exist Figure 10B In the present invention, there is a DNA fragment having a 5' protruding serrated end with a 5' "C" terminal motif and a 3' "T" terminal motif on the left side, and a blunt end with a 5' "A" terminal motif and a 3' "T" terminal motif on the right side. The combined terminal motif can be referred to as "5CTBAT", where the first three characters represent the left end, and the following three characters represent the right end. The first of the three letters indicates the type of serrated end, i.e., "5" indicates a 5' serrated end, "3" indicates a 3' serrated end, and "B" indicates a blunt end. The second of the three letters indicates the 5' terminal motif, and the third letter indicates the 3' terminal motif.
[0164] Figure 11AThis is a correlation plot of the frequency of total 5'-terminal motifs between HCC and healthy subjects. The y-axis shows the frequency of 4-mer 5'-terminal motifs in HCC subjects. The x-axis shows the frequency of 4-mer 5'-terminal motifs in healthy subjects. Each dot represents a different 4-mer 5'-terminal motif. The data show a high correlation, R = 0.98, and p < 2.2e-16. Terminal motifs with points that deviate further from the line yx can be used to distinguish HCC cases from healthy subjects.
[0165] Figure 11B This figure shows the correlation between the frequencies of terminal motifs by stage between HCC and healthy subjects. The y-axis shows the frequency of co-occurring 4-mer terminal motifs in HCC subjects. The x-axis shows the frequency of co-occurring 4-mer terminal motifs in healthy subjects. Each dot represents a different co-occurring terminal motif, including both 5' and 3' 4-mers, distinguishing between 5' and 3' overhangs. The data show a correlation of R = 0.92 and p < 2.2e-16.
[0166] Figure 11C This figure shows the correlation between the frequencies of coterminal motifs in HCC and healthy subjects. The y-axis shows the frequency of coterminal motifs in HCC subjects, and the x-axis shows the frequency of coterminal motifs in healthy subjects. Each dot represents a different coterminal motif, including jagged end types and 1-mer motifs from both the 5' and 3' ends of cfDNA fragments. The data show a correlation of R = 0.91 and p < 2.2e-16.
[0167] Based on Figure 11A Compared to a typical analysis of the overall 5' end motifs in Figure 11B The terminal motif stages in the 5' terminal motifs showed greater differences between HCC and healthy subjects. The ranking of the first four motifs in the overall 5' terminal motifs remained the same between HCC and healthy subjects. In contrast, the ranking of the terminal motifs of the first four stages changed to a great extent. For example, the terminal motif (CCCA_gagg) of the previous stage of healthy subjects dropped to the 4th place in HCC, while the terminal motif (AAAA_TTTT) of the second-ranked stage in healthy subjects rose to the first-ranked stage in HCC patients. The difference in the terminal motifs of the stages of HCC subjects and healthy subjects shows that different terminal motifs of the stages or a combination of different terminal motifs can be used to distinguish HCC cases from healthy subjects.
[0168] In addition, compared with the overall 5' end motif and the phased end motif, the joint motif further expanded the differences between HCC and healthy subjects ( Figure 11C). The order of the first four motifs in the overall 5' terminal motifs was the same between HCC and healthy subjects. Although the order of the terminal motifs of the first four stages changed greatly, the terminal motifs of the first four stages were the same between HCC (the first four stage motifs: AAAA_TTTT, CAAA_TTTT, CCCC_GGGT and CCCA_gagg) and healthy subjects (CCCA_gagg, AAAA_TTTT, CAAA_TTTT and CCCC_GGGT). In contrast, the first four joint terminal motifs were completely different between HCC (the first four joint motifs: 5CT5CT, 5CG5CG, BCGBCG and 5CA5CA) and healthy subjects (the first four joint motifs: BATBAT, BGCBGC, BATBGC and BGCBAT). The difference in the joint terminal motifs between HCC subjects and healthy subjects suggests that the jagged end information, the combination of different terminal motifs from both sides of the cfDNA fragments, can be used to distinguish HCC cases from healthy cases.
[0169] B. Nuclease activity
[0170] DNASE plays different roles in cfDNA fragmentation. The presence of jagged end morphology and / or terminal motifs can be used to analyze nuclease activity.
[0171] 1. Sawtooth end shape
[0172] Our previous studies have shown that different DNASEs play different roles in the generation of cfDNA jagged ends. DNASE activity can be inferred from jagged ends (Ding et al. Clin Chem. 2022;68:917-926). However, only 5'-protruding jagged ends, not 3'-protruding jagged ends, have been analyzed previously. Analysis of all types of jagged ends (e.g., 5'-protruding jagged ends, 3'-protruding jagged ends, and blunt ends) and parallel analysis of single-molecule end morphology can provide more information about the activities of different DNASEs.
[0173] Figures 12A-12F Shown are the results of single-molecule end-point morphometric analysis of DNASE1 from wild-type, DNASE1 (DNASE1) on the PacBio platform (median reads: 1,295,159; range: 176,285-2,624,708). - / - )、DNASE1L3(DNASE1L3 - / - ) and DFFB (DFFB - / - Analysis of plasma cfDNA samples from a ) knockout mouse model. The x-axis shows the classification of nuclease activity. The y-axis shows the frequency of specific jagged end morphologies.
[0174] DNASE1 - / -Mice showed a significant decrease in the frequency of fragments carrying 5' protruding jagged ends (8.76%) ( Figure 12A ), and in DNASE1L3 - / - A significant decrease in the frequency of fragments with 3' protruding jagged ends (52.80%) was observed in mice ( Figure 12B ). In DFFB - / - A significant decrease in the frequency of fragments with blunt ends (40.25%) was observed in mice ( Figure 12C These results suggest that analyzing all types of jagged ends can provide more information about the activities of different DNASEs than analyzing 5' overhang jagged ends alone.
[0175] We further classified cfDNA fragments based on the serrated end morphology from both sides of the molecule (i.e., 5' overhanging serrated end + 3' overhanging serrated end (5-3), 5' overhanging serrated end + 5' overhanging serrated end (5-5), 3' overhanging serrated end + 3' overhanging serrated end (3-3), 5' overhanging serrated end + blunt end (5-B), 3' overhanging serrated end + blunt end (3-B), blunt end + blunt end (BB)). Figure 12D As shown in - / - Compared with the frequency of 5' overhanging jagged ends in WT mice, a greater reduction was observed in the frequency of cfDNA fragments carrying 5-5 jagged ends (5-5 jagged ends vs. 5' overhanging jagged ends: 15.40% vs. 8.76%). Similar to DNASE1- / - mice, and compared with DNASE1L3- / - mice (reduction: 3-3 jagged ends vs. 3' overhanging jagged ends: 71.45% vs. 52.80%), respectively ( Figure 12E ) and DFFB- / - mice (reduction: BB serrated end vs. blunt end: 70.41% vs. 40.25%) ( Figure 12F Compared to the frequencies of 3' overhanging jagged and blunt ends in cfDNA fragments (Figure 5—figure supplement 1), a greater reduction was observed in the frequencies of cfDNA fragments carrying 3-3 jagged and BB jagged ends. These results suggest that concurrent analysis of jagged ends on both sides of a cfDNA fragment can improve the ability to distinguish between distinct activities of different DNASEs. This technology could be used to enhance the diagnosis of diseases with abnormal DNASE activity, such as, but not limited to, systemic lupus erythematosus.
[0176] 2. Terminal motif and sawtooth terminal morphology
[0177] Our previous publications reported that cfDNA terminal motifs can be used to infer the activities of different DNASEs (Han et al. Am J Hum Genet. 2020;106:202-214; Jiang et al. Cancer Discov. 2020;10:664-673). The results discussed in the previous section suggest that different DNASEs may be associated with various types of jagged ends. Analyzing the terminal motifs of different jagged end groups may improve the ability to discriminate between changes in DNASE activity.
[0178] Figure 13 13 is a table of different serrated end morphologies and terminal nucleotide types for different nuclease activities. Main columns 1304, 1308, and 1312 show data for different serrated ends. Main rows 1316, 1320, and 1324 show the different nuclease activities analyzed. The individual columns under each main column show the median frequency for wild-type mice and specific nuclease knockout mice in the main row, as well as the relative change in median frequency between the nuclease knockout mice and wild-type mice. The individual rows indicate the terminating nucleotides. The gray-shaded cells indicate the maximum change for each terminating nucleotide between the different serrated end types. Each row has only one shaded cell. For example, for DNASE1L3- / - mice, the maximum change in the A-terminus was found in the 5' serrated end, so the small holes for the relative change in the A-terminus in the 5' serrated end are shaded.
[0179] like Figure 13 As shown, compared with the fragments with 3' protruding serrated ends and the fragments with blunt ends, the fragments with 5' protruding serrated ends were more abundant in DNASE1L3 compared with WT mice. - / - The largest increase in 5' end motifs was observed in mice, with A- (median increase: 5' vs. 3' vs. blunt: 39.28% vs. 13.86% vs. 25.68%) and G- (median increase: 5' vs. 3' vs. blunt: 21.55% vs. 4.79% vs. 4.79%). Fragments with blunt ends were significantly more abundant in DNASE1L3 compared to WT mice, compared to fragments with 5' protruding serrated ends and fragments with 3' protruding serrated ends. - / - Mice showed the greatest reduction in C- (median reduction: 5' vs. 3' vs. blunt: 18.83% vs. 4.25% vs. 21.60%) and T- (median reduction: 5' vs. 3' vs. blunt: 41.44% vs. 10.23% vs. 77.99%) 5' end motifs. - / -In mice, compared with WT mice, fragments with blunt ends showed the greatest decreases at the 5' C termini (median decrease: 5' vs. 3' vs. blunt: 9.67% vs. 0.40% vs. 13.70%) and 5' T termini (median decrease: 5' vs. 3' vs. blunt: 7.68% vs. 1.64% vs. 45.90%), and the most significant increase at the 5' A termini (median increase: 5' vs. 3' vs. blunt: 11.43% vs. 2.18% vs. 19.69%), while fragments with 5' protruding jagged ends showed the greatest increase at the 5' G termini (median increase: 5' vs. 3' vs. blunt: 11.03% vs. -1.74% vs. -0.29%). Interestingly, for DFFB - / - Mouse, the greatest changes in 5' terminal motifs were observed in fragments with blunt ends (5' C-terminus (median decrease: 5' vs. 3' vs. blunt: 4.16% vs. -0.80% vs. 33.73%); 5' T-terminus (median decrease: 5' vs. 3' vs. blunt: 22.57% vs. 1.48% vs. 110.94%); 5' A-terminus (median decrease: 5' vs. 3' vs. blunt: 15.43% vs. 2.50% vs. 28.12%); 5' G-terminus (median decrease: 5' vs. 3' vs. blunt: 4.70% vs. -2.35% vs. 8.68%)). We further analyzed DFFB - / - and 4-mer 5' terminal motifs in all fragments, fragments with 5' protruding serrated ends, fragments with 3' protruding serrated ends, and fragments with blunt ends in WT mice.
[0180] Figures 14A-14D It is DFFB - / - Figure 2. Schematic diagram of terminal motif ranking in DFFB knockout (KO) mice and wild-type (WT) mice. The graph has the motif ranking in wild-type mice on the x-axis, and DFFB - / - The order of motifs in mice is on the y-axis. Figure 14A DFFB is shown - / - 5' end motif sequencing of all cfDNA fragments pooled from mice and wild-type mice. Figure 14B Shows FFB - / - 5' end motif sequencing of pooled cfDNA fragments carrying 5' overhanging jagged ends in mice and wild-type mice. Figure 14C DFFB is shown - / - 5' end motif sequencing of pooled cfDNA fragments carrying 3' overhanging jagged ends in mice and wild-type mice. Figure 14D DFFB is shown - / -5' end motif sequencing of combined cfDNA fragments with blunt ends in mice and wild-type mice. Figures 14A-14D As shown in , DFFB can be observed in fragments with blunt ends compared to all fragments, fragments with 5' protruding serrated ends, and fragments with 3' protruding serrated ends. - / - The largest difference between the WT mice (R: all vs. 5' vs. 3' vs. blunt: 0.94 vs. 0.95 vs. 1 vs. 0.77). However, the fragments with 5' protruding jagged ends also had DFFB - / - Overrepresented or underrepresented motifs in mice.
[0181] These data demonstrate that parallel analysis of single-molecule end morphologies allows for more precise deciphering of the characteristic cleavages attributed to various DNA nucleases in plasma.
[0182] C. Exemplary Methods
[0183] Any of the processes described herein can be used to determine the grade of a condition. Examples can include treating a disease or condition in a patient after determining the grade of the disease or condition in the patient. Treatment can include any suitable therapy, medication, or surgery, including any treatment described in the references mentioned herein. Information about treatment in the references is incorporated herein by reference.
[0184] Treatment can be provided based on the determined cancer grade, determined mutation, and / or tissue of origin. For example, identified mutations (e.g., for polymorphism embodiments) can be targeted with specific drugs or chemotherapy. The tissue of origin can be used to guide surgery or any other form of treatment. Furthermore, the grade of cancer can be used to determine the aggressiveness of any type of treatment, which can also be determined based on the grade of cancer.
[0185] A statistically significant number of free DNA molecules can be analyzed to provide an accurate determination of the proportional contribution from the first tissue type. In some embodiments, at least 1,000 free DNA molecules are analyzed. In some instances, at least 1,000 free DNA molecules are analyzed. In other embodiments, at least 10,000 or 50,000 or 100,000 or 500,000 or 1,000,000 or 5,000,000 free DNA molecules or more can be analyzed.
[0186] 1. Sawtooth Ends from Parallel Analysis
[0187] Figure 1515 is a flow chart of an exemplary process 1500 for analyzing a biological sample obtained from an individual. The biological sample may include a plurality of nucleic acid molecules. The nucleic acid molecules may be free and double-stranded, having a first strand and a second strand. At least one of the nucleic acid molecules may have an overhang, where the first strand or the second strand overlaps the other strand. Process 1500 may use overhang information from the four ends of the two strands of the nucleic acid molecule to determine the grade of the individual's condition. In some embodiments, the process may be performed by system 10 or system 2400. Figure 15 One or more process boxes.
[0188] At block 1502, for each nucleic acid molecule in a plurality of nucleic acid molecules, a first strand-specific classification of a first end characteristic of the nucleic acid molecule is measured. The strand-specific classification can indicate whether the first strand or the second strand protrudes beyond the other strand, including if no strand protrudes beyond the other strand. The strand-specific classification can identify whether the first strand or the second strand is a 3' strand or a 5' strand. The strand-specific classification can also indicate the length of the protrusion of the first strand or the second strand. The strand-specific classification can include a sawtooth end morphology as described herein. The characteristics can be measured using process 300.
[0189] For each nucleic acid molecule in the plurality of nucleic acid molecules, a second strand-specific classification of the second end of the nucleic acid molecule can be measured.
[0190] At block 1504, a jagged end value is determined using the first strand-specific classification of the plurality of nucleic acid molecules. The jagged end value can be the amount of nucleic acid molecules having a certain type of jagged end, including 5' overhangs, 3' overhangs, and blunt ends (e.g., Figure 5A ). The quantity can be number, total length, mass, or frequency. In some embodiments, the sawtooth end value can be an amount of nucleic acid molecules, including an amount from one of the following categories: a blunt end at the first end and a blunt end at the second end, a 5' overhang at the first end and a blunt end at the second end, a 3' overhang at the first end and a blunt end at the second end, a 5' overhang at the first end and a 3' overhang at the second end, a 5' overhang at the first end and a 5' overhang at the second end, and a 3' overhang at the first end and a 3' overhang at the second end (e.g., Figure 5B ). The jagged end value can be determined using a second strand specific classification.
[0191] In some examples, the sawtooth end value can be an element in a vector. The vector can include multiple elements. The multiple elements can include the amount of nucleic acid molecules in one or more of the following categories: a blunt end at the first end and a blunt end at the second end, a 5' overhang at the first end and a blunt end at the second end, a 3' overhang at the first end and a blunt end at the second end, a 5' overhang at the first end and a 3' overhang at the second end, a 5' overhang at the first end and a 5' overhang at the second end, and a 3' overhang at the first end and a 3' overhang at the second end.
[0192] The plurality of elements may include an assortment of nucleic acid molecules having sizes within one or more size ranges (e.g., Figure 5B and 5C). The one or more size ranges can be any size range described herein. The size range can include sizes less than or greater than any of the following sizes: 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, or 350. In addition, the size range can include sizes between and including any two sizes described herein. The size of each nucleic acid molecule can be determined by aligning the subsequences corresponding to the ends of the corresponding nucleic acid molecule with a reference genome.
[0193] The sawtooth end value may include the ratio of the amount of nucleic acid molecules in one overhang classification having a size within a certain size range to the amount of nucleic acid molecules in the same overhang classification having a size within a different size range. A vector may include the size ratios for each of a plurality of overhang classifications (e.g., Figure 7A ).
[0194] At block 1506, the sawtooth end value is compared to a reference value. The comparison can determine whether the sawtooth end value is statistically significantly different from the reference value. The reference value can be any reference value described herein. The vector can be compared to a reference vector (which can include multiple different reference values). The comparison can be performed between corresponding elements in the vector. In some embodiments, the comparison can be performed by a machine learning model. For example, a machine learning model can be trained using sawtooth end values determined from subjects with a known grade of condition.
[0195] At block 1508, the comparison is used to determine the condition grade of the individual. The condition can be cancer, an autoimmune disease, a pregnancy-related disorder, a nuclease activity deficiency, or any condition described herein. A reference value can be determined from one or more subjects with a certain grade of the condition or from one or more healthy subjects. If the sawtooth end value is statistically identical to the reference value, the grade of the condition can be determined to be the same for the one or more subjects associated with the reference value.
[0196] In some instances, the grade of the condition is not determined. Instead, a comparison can be used to determine the fractional concentration of clinically relevant DNA. A reference value can be determined from one or more subjects with known fractional concentrations of clinically relevant DNA. The reference value can be a calibration value determined using a calibration sample.
[0197] Process 1500 may include additional embodiments, such as any single embodiment or any combination of embodiments described herein and / or in conjunction with one or more other processes described elsewhere herein.
[0198] although Figure 15 Example blocks of process 1500 are shown, but in some implementations, process 1500 may include Figure 15 The blocks shown in FIG1500 may be additional blocks, fewer blocks, different blocks, or blocks arranged differently than those shown in FIG1500. Additionally or alternatively, two or more blocks of process 1500 may be executed in parallel.
[0199] 2. Terminal motif at one end
[0200] Figure 16 16 is a flow chart of an exemplary process 1600 for analyzing a biological sample obtained from an individual. The biological sample may include a plurality of nucleic acid molecules. The nucleic acid molecules may be free and double-stranded, having a first strand and a second strand. Process 1600 may use terminal motifs at at least one end of the molecule to determine the grade of a condition. At least some of the plurality of nucleic acid molecules have nucleotides on one strand that do not have a complementary portion on the other strand. In some embodiments, Figure 16 One or more process blocks of may be performed by system 10 or system 2400 .
[0201] At block 1602, for each nucleic acid molecule in a plurality of nucleic acid molecules, a first sequence end motif of a first strand at a first end of the nucleic acid molecule is determined.
[0202] At block 1604, the second sequence end motif of the second strand of the first end of the nucleic acid molecule is determined. In an example, the first strand may have a 5' end at the first end. In other examples, the first strand may have a 3' end at the first end. The first strand may protrude beyond the second strand, or the second strand may protrude beyond the first strand. The first end may be a blunt end. The subsequence can be determined using process 300. In some embodiments, the second sequence end motif of the second strand can be determined by using the complementary nucleotides of the corresponding nucleotides in the first strand.
[0203] At block 1606, a first quantity of nucleic acid molecules having a first combination of a first sequence end motif and a second sequence end motif at a first end is determined. The sequence end motif can have 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. The quantity can be number, total length, mass, or frequency. The first combination can be about Figure 9A and Figure 9B The terminal modules of the phases described.
[0204] At block 1608, a value for a terminal motif parameter is generated using the first quantity. In some embodiments, the terminal motif parameter may be the first quantity. In some embodiments, the terminal motif parameter may be a ratio of the first quantity to another quantity (e.g., the quantities of all terminal motifs). As an example, the terminal motif parameter may be a frequency.
[0205] In addition to the terminal motifs of the first phase, the terminal motifs of the second phase can also be used. A second amount of a nucleic acid molecule having a second combination of a third sequence terminal motif and a fourth sequence terminal motif at the first end can be determined. The third sequence terminal motif can be at the 5' strand or the 3' strand. As an example, the first amount can be an amount of AAAA_TTTT, and the second amount can be an amount of CCCA_gagg. The second amount can be used to generate the value of the terminal motif parameter. The terminal motif parameter can be a vector of certain combinations of different amounts of terminal motifs.
[0206] At block 1610, the value of the terminal module parameter is compared to a reference value. The reference value can be any reference value described herein. The comparison can be any comparison described herein, including that described with respect to block 1506.
[0207] At block 1612 , the comparison is used to determine a level of the individual's condition. The determination may be performed similarly to block 1508 .
[0208] The condition can be cancer, HCC, an autoimmune disease, a pregnancy-related disorder, a nuclease activity defect, or any condition described herein. The reference value can be determined from one or more subjects suffering from a certain grade of the condition or one or more healthy subjects.
[0209] In some instances, the grade of the condition is not determined. Instead, a comparison can be used to determine the fractional concentration of clinically relevant DNA. A reference value can be determined from one or more subjects with known fractional concentrations of clinically relevant DNA. The reference value can be a calibration value determined using a calibration sample.
[0210] Process 1600 may include using two sequence motifs from the sawtooth end at the other end of the molecule. Four sequence motifs from the same molecule may be used. In an embodiment, the plurality of nucleic acid molecules is a first plurality of nucleic acid molecules. The biological sample may include a second plurality of nucleic acid molecules. The first plurality of nucleic acid molecules may include a subset of the second plurality of nucleic acid molecules. Process 1600 may further include determining, for each nucleic acid molecule in the second plurality of nucleic acid molecules, a third sequence end motif on the first strand that does not have a complementary portion on the second strand at the second end of the nucleic acid molecule, and determining a fourth sequence end motif on the second strand at the second end of the nucleic acid molecule. A second quantity of nucleic acid molecules having a second combination of the third sequence end motif and the fourth sequence end motif at the second end may be determined. The second quantity may be used to generate a value for an end motif parameter. In some embodiments, the value of the end motif parameter may be the quantity of molecules having a certain combination of the four sequence motifs present on the molecule.
[0211] In some embodiments, the jagged end morphology at one end can also be used to determine the grade of the condition. Process 1600 can include, for each nucleic acid molecule in a plurality of nucleic acid molecules, measuring a first strand-specific classification of a first end characteristic of the nucleic acid molecule. The strand-specific classification can indicate whether the first strand or the second strand protrudes beyond the other strand. For example, at one end, the strand-specific classification can indicate a 3' protruding end, a 5' protruding end, a blunt end, or a jagged end (usually). Determining a first amount can include determining a first amount of nucleic acid molecules having a first combination and a first strand-specific classification.
[0212] In some embodiments, a sawtooth end morphology of the second end can also be used. Process 1600 can include, for each nucleic acid molecule in the plurality of nucleic acid molecules, measuring a second strand-specific classification of the second end of the nucleic acid molecule. Determining the first amount can include determining a first amount of nucleic acid molecules having a first combination, a first strand-specific classification, and a second strand-specific classification.
[0213] Process 1600 may include additional embodiments, such as any single embodiment or any combination of embodiments described herein and / or in conjunction with one or more other processes described elsewhere herein.
[0214] although Figure 16 Example blocks of process 1600 are shown, but in some implementations, process 1600 may include Figure 16The blocks shown in FIG1600 may be additional blocks, fewer blocks, different blocks, or blocks arranged differently than those shown in FIG1600. Additionally or alternatively, two or more blocks of process 1600 may be executed in parallel.
[0215] 3.3' terminal motif
[0216] Figure 17 17 is a flow chart of an exemplary process 1700 for analyzing a biological sample obtained from an individual. The biological sample may include a plurality of nucleic acid molecules. The nucleic acid molecules may be free and double-stranded, having a first strand and a second strand. Process 1700 may use terminal motifs at at least one end of the molecule to determine the grade of a condition. At least some of the plurality of nucleic acid molecules have nucleotides on one strand that do not have a complementary portion on the other strand. In some embodiments, Figure 17 One or more process blocks of may be performed by system 10 .
[0217] At block 1702, for each nucleic acid molecule in a plurality of nucleic acid molecules, a first sequence terminal motif of a strand at a first end of the nucleic acid molecule is determined, where the first end is the 3' end of the strand. The first sequence terminal motif is the actual terminal motif of the original molecule and not the terminal motif at the 3' end after the molecule has been end-blunted, either by filling in nucleotides on the 3' strand or by removing nucleotides on the 3' strand. The first sequence terminal motif can be determined using process 300.
[0218] At block 1704, a first amount of a nucleic acid molecule having a first sequence end motif at a first end is determined. The first amount can be an absolute or relative amount.
[0219] Sequence end motifs can be determined at both ends of a single strand. In some examples, process 1700 can include, for each nucleic acid molecule in the plurality of nucleic acid molecules, determining a second sequence end motif of the strand at a second end of the nucleic acid molecule. The first quantity is nucleic acid molecules having a first sequence motif at a first end and a second sequence end motif at a second end.
[0220] In some embodiments, the size of a single strand can be determined. The two ends of a single strand can be compared to a reference genome to determine the size. The size of the complementary strand can also be determined. These two sizes can be used to generate a length difference between the two strands. A statistical value of the length difference between the two strands of multiple molecules can be determined and compared to a reference value. The comparison can be used to determine the grade of the condition. The length difference can also be determined by adding or subtracting the length of the overhang at each end of the molecule, without having to determine the length of any one strand.
[0221] At block 1706, a value of an end motif parameter is generated using the first quantity. In some embodiments, the end motif parameter can be a ratio of the first quantity to another quantity (e.g., the quantities of all end motifs). As an example, the end motif parameter can be a frequency.
[0222] At block 1708, the value of the terminal module parameter is compared to a reference value. The reference value can be any reference value described herein. For example, the reference value can be determined from a calibration sample of an individual with a known degree of a condition.
[0223] At block 1710, the comparison is used to determine a level of the individual's condition.
[0224] In some embodiments, terminal motifs at both 3' ends of a single molecule can be used. The strand can be a first strand. The terminal motif parameter can be a first terminal motif parameter. The reference value can be a first reference value. Process 1700 can also include, for each nucleic acid molecule in the plurality of nucleic acid molecules, determining a second sequence terminal motif of a second strand at a second end of the nucleic acid molecule. The second end is the 3' end of the second strand. Process 1700 can include determining a second quantity of nucleic acid molecules having the second sequence terminal motif at a second end. The second quantity can be used to generate a value for the second terminal motif parameter. The value for the second terminal motif parameter can be compared to a second reference value. The comparison can be used to determine the severity of the condition.
[0225] In some embodiments, the jagged end morphology at one end can also be used to determine the grade of the condition. Process 1700 can include, for each nucleic acid molecule in a plurality of nucleic acid molecules, measuring a first strand-specific classification of a first end characteristic of the nucleic acid molecule. The strand-specific classification can indicate whether the first strand or the second strand protrudes beyond the other strand. For example, at one end, the strand-specific classification can indicate a 3' protruding end, a 5' protruding end, a blunt end, or a jagged end (usually). The first (or second) strand-specific classification can be a specific jagged end morphology from a list of possible strand-specific classifications. Determining a first amount can include determining a first amount of nucleic acid molecules having a first combination and a first strand-specific classification.
[0226] In some embodiments, a jagged end morphology of the second end can also be used. Process 1700 can include, for each nucleic acid molecule in a plurality of nucleic acid molecules, measuring a second strand-specific classification of the second end of the nucleic acid molecule. Determining a first quantity can include determining a first quantity of nucleic acid molecules having a first combination, a first strand-specific classification, and a second strand-specific classification.
[0227] Process 1700 may include additional embodiments, such as any single embodiment or any combination of embodiments described herein and / or in conjunction with one or more other processes described elsewhere herein (eg, process 1600 ).
[0228] although Figure 17 Example blocks of process 1700 are shown, but in some implementations, process 1700 may include Figure 17 1700. In some embodiments, the process 1700 may include two or more blocks of process 1700, and two or more blocks of process 1700 may be executed in parallel.
[0229] III. Enrichment
[0230] Certain types of DNA, including clinically relevant DNA, may tend to be more highly represented in DNA having certain jagged end morphologies or sequence end motifs. Thus, enrichment for certain jagged end morphologies and / or sequence end motifs may result in samples enriched for certain types of clinically relevant DNA. Enrichment may include physical enrichment of samples or computational enrichment of reads obtained from analyzing biological samples.
[0231] A.Serrated end
[0232] In order to explore the potential application of single-molecule terminal morphology parallel analysis in non-invasive prenatal testing (NIPT), the terminal morphology in fetal-specific and shared cfDNA fragments in maternal plasma was analyzed. The genotypes of maternal buffy coat and placental tissue samples obtained by using microarray-based genotyping technology (HumanOmni2.5 genotyping array Illumina) were defined for fetal-specific and shared cfDNA fragments. Informative SNPs were identified (i.e., where the mother was homozygous (expressed as AA genotype) and the fetus was heterozygous (expressed as AB genotype)). Fetal-specific DNA fragments were identified based on DNA fragments carrying fetal-specific alleles at informative SNP sites. In this case, the B allele is fetal-specific, and the DNA fragments carrying the B allele are deduced to be derived from fetal tissue. Shared DNA fragments are identified based on DNA fragments carrying shared alleles at informative SNP sites. In this case, the A allele is shared, and DNA fragments carrying the A allele are inferred to originate from both fetal and maternal tissues (primarily from maternal tissue). The number of fetal-specific molecules carrying the fetal-specific allele (B) is determined (p). The number of molecules carrying the shared allele (A) is determined (q). The fetal DNA fraction across all cell-free DNA samples will be calculated using 2p / (p+q). Calculated at 100%.
[0233] Figures 18A-18CShown are parallel analyses of single-molecule end morphology of a total of 10 maternal plasma DNA samples on the PacBio sequencing platform (median number of reads: 1,305,115; range: 393,197-1,921,070).
[0234] Figure 18A Figure 2 is a plot of the frequency of 5' jagged overhangs. The y-axis shows the frequency of 5' jagged overhangs as a percentage. The x-axis shows the fractions carrying the shared allele and the fractions carrying the fetal-specific allele. Fetal-specific cfDNA carries more 5' jagged overhangs than shared cfDNA (shared vs. fetal-specific: median: 52.0% vs. 59.2%).
[0235] Figure 18B Figure 2 is a graph showing the frequency of 3' jagged overhangs. The y-axis shows the frequency of 3' jagged overhangs as a percentage. The x-axis shows the fractions carrying the shared allele and the fetal-specific allele. Fetal-specific cfDNA has more 3' jagged overhangs than shared cfDNA (shared vs. fetal-specific: median: 23.0% vs. 33.5%).
[0236] Figure 18C Figure 2 is a graph showing the frequency of blunt ends. The y-axis shows the frequency of blunt ends as a percentage. The x-axis shows the fractions carrying the shared allele and the fractions carrying the fetal-specific allele. Fetal-specific cfDNA has fewer blunt ends than shared cfDNA (median: 23.5% vs. 8.7%).
[0237] Figure 18D The percentage of fetal DNA fraction based on different end morphologies is shown. The y-axis shows the fetal DNA fraction as a percentage. The x-axis shows different end morphologies. Selective analysis of cfDNA with 5' overhanging jagged ends (5' overhanging jagged ends relative to all fragments: median: 16.61% vs. 15.41%) or 3' overhanging jagged ends (3' overhanging jagged ends relative to all fragments: median: 17.94% vs. 15.41%) showed a significant increase in the fetal DNA fraction compared to all cfDNA fragments. Conversely, selective analysis of cfDNA with blunt ends (blunt ends relative to all fragments: median: 5.88% vs. 15.41%) showed a significant decrease in the fetal DNA fraction compared to all cfDNA fragments, which indeed indicates a significant increase in the fraction of DNA of maternal origin. Figures 18A-18D Types displaying jagged ends can be used to enrich for DNA from specific sources.
[0238] In another embodiment, cfDNA fragments are classified into 6 different groups based on the jagged end morphology from both sides of the molecule (e.g., 5' overhanging jagged end + 3' overhanging jagged end (5-3), 5' overhanging jagged end + 5' overhanging jagged end (5-5), 3' overhanging jagged end + 3' overhanging jagged end (3-3), 5' overhanging jagged end + blunt end (5-B), 3' overhanging jagged end + blunt end (3-B), blunt end + blunt end (BB)).
[0239] Figure 19 The figure plots the fetal DNA fraction relative to different jagged end morphologies. The y-axis shows the fetal DNA fraction derived from cfDNA fragments. The x-axis shows different end morphologies (e.g., 5'-overhanging jagged end and 3'-overhanging jagged end (5-3), 5'-overhanging jagged end and 5'-overhanging jagged end (5-5), 3'-overhanging jagged end and 3'-overhanging jagged end (3-3), 5'-overhanging jagged end and blunt end (5-B), 3'-overhanging jagged end and blunt end (3-B), blunt end and blunt end (BB)) and all fragments. Selective analysis of cfDNA belonging to the 3-3 (3-3 relative to all fragments: median: 23.48% vs. 15.41%), 5-3 (5-3 relative to all fragments: median: 19.50% vs. 15.41%), and 5-5 (5-5 relative to all fragments: median: 17.93% vs. 15.41%) groups showed a significant increase in the fetal DNA fraction compared to all cfDNA fragments. In contrast, selective analysis of cfDNA belonging to the 3-B (3-B vs. all fragments: median: 9.12% vs. 15.41%), 5-B (5-B vs. all fragments: median: 8.17% vs. 15.41%), and BB (BB vs. all fragments: median: 2.53% vs. 15.41%) groups showed a significant reduction in fetal DNA fraction compared to all cfDNA fragments. These results suggest that analysis of jagged end morphology from both sides of cfDNA fragments can enrich for clinically relevant DNA.
[0240] B. Parallel Analysis of Sawtooth Terminals and Terminal Motifs
[0241] Parallel analysis of single-molecule end morphology (e.g., combined end motifs and jagged ends) can enrich fetal DNA in maternal plasma. We merged all sequencing reads from the 10 pregnant women mentioned above (12,142,332 reads).
[0242] Figure 20A and 20BFigure 2 is a graph of the fraction of fetal DNA derived from fragments with certain jagged end morphologies and sequence end motifs. The y-axis of the graph represents the fraction of fetal DNA derived from cfDNA fragments. The x-axis shows different classifications of fragments: all fragments, jagged end morphology, end motifs, and a combination of jagged end morphology and end motifs.
[0243] Figure 20A Fragments carrying 5' overhanging jagged ends as well as 5' CCG terminal motifs on either side of the fragment (fetal DNA fraction: 29.3%), fragments with only 5' overhanging jagged ends (fetal DNA fraction: 18.5%), or fragments with only CCG 5' terminal motifs (fetal DNA fraction: 25.2%) are shown with substantially increased fetal DNA fractions compared to all fragments (fetal DNA fraction: 16.3%).
[0244] Figure 20B The fragments carrying 3' protruding jagged ends and 5' GCG terminal motifs on either side of the fragment (fetal DNA fraction: 35.6%), fragments with only 3' protruding jagged ends (fetal DNA fraction: 21.2%), or fragments with only GCG 5' terminal motifs (fetal DNA fraction: 16.2%) showed a substantial increase in fetal DNA fraction compared to all fragments (fetal DNA fraction: 16.3%). These results indicate that parallel analysis of terminal motifs with jagged ends at the same end can promote the enrichment of fetal DNA in maternal plasma. Parallel analysis of single molecule terminal morphology can promote NIPT.
[0245] C. Exemplary Methods
[0246] Figure 21 21 is a flow chart of an exemplary process 2100 for enriching a biological sample for clinically relevant DNA. The biological sample may include clinically relevant DNA and other DNA. Each nucleic acid molecule in the plurality of nucleic acid molecules is double-stranded, having a first strand and a second strand. The clinically relevant DNA may be tumor DNA, transplant DNA, or fetal DNA. The biological sample may be obtained from a female subject who is pregnant with a fetus, and the clinically relevant DNA may be fetal DNA or maternal DNA. In some embodiments, the system 2400 may be performed Figure 21 One or more process boxes.
[0247] At box 2110, for each nucleic acid in a plurality of nucleic acid molecules, the first strand-specific classification of the nucleic acid molecule first end is measured. Whether the strand-specific classification indicates that the first strand or the second strand protrudes outside the other strand. The strand-specific classification can include the first strand protruding outside the second strand, the second strand protruding outside the first strand, and / or both strands do not protrude outside the other strand (blunt end). The first strand-specific classification that the subset of nucleic acid molecules has can be that the first strand protrudes outside the second strand. The first strand can be a 3' strand or a 5' strand. In some embodiments, the first strand of the strand-specific classification indication nucleic acid molecule of a subset of nucleic acid molecules protrudes outside the second strand, and the second strand of the strand-specific classification indication nucleic acid molecule of a subset of nucleic acid molecules protrudes outside the first strand. For example, the 5' end can protrude at both ends.
[0248] At box 2120, reads corresponding to a subset of nucleic acid molecules with a classification specific to the first strand are selected to form an enriched sample. The enriched sample can be an enriched computer sample. In some embodiments, the enriched sample can be formed by a physical enrichment technique. For example, according to some embodiments, a certain number of target sawtooth ends can be enriched using targeted capture based on sawtooth end-specific hybridization. In one embodiment of a physical enrichment analysis, a target sawtooth end can be enriched using targeted capture based on sawtooth end-specific hybridization. A biotinylated RNA probe that can specifically hybridize with the target sawtooth end is designed. The target sawtooth end hybridized with the biotinylated probe can be pulled down by streptavidin-coated magnetic beads. The RNA probe will be degraded by ribonucleases such as RNase H. The target jagged end will be enriched in the pull-down material. In one embodiment, one or more different sawtooth ends are analyzed together, for example, for practical applications, the ratio or deviation between the reads of different sawtooth ends.
[0249] In some embodiments, the second strand-specific classification of the second end of the nucleic acid molecule of each nucleic acid molecule in the plurality of nucleic acid molecules. A subset of nucleic acid molecules can have a second strand-specific classification. For example, an enriched sample can include molecules that have a certain type of overhang at one end and the same type of overhang at the other end.
[0250] In an embodiment, a first sequence end motif at a first end of the nucleic acid molecule can be determined for each nucleic acid molecule in the plurality of nucleic acid molecules. Selecting reads corresponding to the subset of nucleic acid molecules can include selecting reads corresponding to nucleic acid molecules having the first sequence end motif.
[0251] In an embodiment, a second sequence end motif at a second end of the nucleic acid molecule can be determined for each nucleic acid molecule in the plurality of nucleic acid molecules. Selecting reads corresponding to a subset of nucleic acid molecules can include selecting reads corresponding to nucleic acid molecules having the same second sequence end motif.
[0252] In some embodiments, the method can also include analyzing a subset of nucleic acid molecules to determine the classification of the disease grade. For example, the method can include comparing the reading of the subset with a reference genome. Methylation-aware sequencing or other detection techniques can be performed to determine the methylation level or methylation pattern (for example, the methylation state at one or more genomic sites). The methylation level or methylation pattern can be compared with the reference level or pattern of a control sample suffering from a known grade disease. The grade of the disease can be determined using comparison.
[0253] In some embodiments, the method can include determining chromosomal aberrations or fetal haplotypes. The reads of the subset can be compared with a reference genome. Chromosome aberrations (e.g., amplifications or deletions) or fetal haplotypes can be identified from the comparison.
[0254] In an embodiment, process 2100 may further include determining a first amount of readings. A first parameter may be determined using the first amount of readings. The first parameter may be determined using the first amount of sequence readings and another amount (e.g., the total amount of readings or readings with a certain chain-specific classification or sequence end motif). In some instances, both amounts may be separate parameters. Other amounts may take different forms, e.g., corresponding to the total number of sequence readings and / or DNA molecules analyzed. The first parameter may be a ratio of amounts.
[0255] A characteristic of the biological sample can be determined using a first parameter. The first characteristic can be a fractional concentration of clinically relevant DNA molecules in the biological sample. The characteristic of the biological sample can be a level of aberrations in the biological sample. A first value of the characteristic of the biological sample is estimated by comparing the first parameter to one or more calibration values determined from one or more calibration samples having known characteristic values.
[0256] Thus, parameters generated based on the corresponding nucleases can be used to determine the characteristics of a biological sample. These corresponding parameters can be combined to form new combined parameters, for example, as ratios, ratios of corresponding functions of the corresponding parameters, and as two inputs to more complex functions such as machine learning models. Exemplary combined parameters can include other ratios of DNASE1L3 / DFFB, DNASE1 / DFFB, or DNASE1L3:DNASE1:DFFB. In addition, parameters for more than two nucleases can be used, for example, relative parameters for three or more nucleases can be used.
[0257] In some embodiments, a first value of a biological sample characteristic is estimated based on analyzing a set of parameters, wherein each parameter corresponds to an amount of sequence reads combined with another amount (e.g., for normalization), each sequence read including a termination sequence corresponding to a particular sequence end feature. For example, a parameter can include a specific combination of frequency ratios between two sets of sequence reads and their corresponding end features. For example, a first parameter in the set of parameters can correspond to a ratio of chain-specific classifications between a first amount of sequence reads and another amount of sequence reads, each sequence read including a chain-specific classification corresponding to a chain-specific classification of a first nucleic acid enzyme, and a second parameter in the set of parameters can correspond to a ratio of chain-specific classifications between a second amount of sequence reads and a third amount of sequence reads, each sequence read including a chain-specific classification corresponding to an end feature of a second nucleic acid enzyme. In some cases, the third amount of sequence reads is another amount of sequence reads used to determine the first parameter.
[0258] The determined features can include gestational age or range (e.g., 8 weeks or 9-12 weeks), for example, when the nuclease is differentially regulated between fetal tissue and maternal tissue. In another example, the determined features can be specific tissue types (e.g., hematopoietic cells) relative to other tissue types (e.g., hematopoietic cells). The characteristics of the target tissue type can also indicate the specific condition (e.g., HCC, preeclampsia, premature birth) of the target tissue type. In another example, the determined features can be the size or nutritional status of an organ corresponding to a specific tissue type (e.g., hepatocytes). In another example, the determined features can include the score of clinically relevant DNA in the biological sample.
[0259] The comparison can be with multiple calibration values. The comparison can be performed by inputting the first parameter into a calibration function appropriate for the calibration data, the calibration function providing a change in the first parameter relative to a characteristic change in the sample. As another example, one or more calibration values can correspond to other parameters in one or more calibration samples.
[0260] Typically, it is preferred to generate the one or more calibration values determined from the one or more calibration samples using an assay similar to that used for the biological (test) samples.For example, sequencing libraries may be generated in the same manner.
[0261] Process 2100 may include additional embodiments, such as any single embodiment or any combination of embodiments described herein and / or in combination with one or more other processes described elsewhere herein or in US 2022 / 0010353 A1, the entire contents of which are incorporated herein by reference for all purposes.
[0262] although Figure 21Example blocks of process 2100 are shown, but in some implementations, process 2100 may include Figure 21 The blocks shown in FIG2100 may be additional blocks, fewer blocks, different blocks, or blocks arranged differently than those shown in FIG2100. Additionally or alternatively, two or more blocks of process 2100 may be executed in parallel.
[0263] IV. Fraction of Clinically Relevant DNA
[0264] This document, including the results described in Section II.B, demonstrates the relationship between different jagged end types and the activities of different DNASEs. 5' jagged overhangs have been shown to correlate with DNASE1 activity, and blunt ends have been shown to correlate with DFFB activity. Because different DNASEs are expressed differently in different tissues, jagged end signatures can be used to infer the tissue of origin of cfDNA.
[0265] Figure 22A This graph shows DNASE1 mRNA expression levels in leukocytes and placenta. The y-axis shows RPKM, a normalized gene expression unit derived from RNA sequencing results, representing reads per kilobase per million sequenced reads (Trapper et al. Nat Biotechnol. 2010;28:511-5). The x-axis shows leukocytes and placenta.
[0266] Figure 22B is a graph showing the mRNA expression levels of DFFB in leukocytes and placenta. Figure 22A same.
[0267] Figure 22C Figure 2 is a graph showing the correlation between the fetal DNA fraction and the frequency of cfDNA fragments with jagged 5' overhangs. The x-axis shows the SNP-based fetal DNA fraction. The y-axis shows the frequency of jagged 5' overhangs.
[0268] Figure 22D Figure 2 is a graph showing the correlation between the fetal DNA fraction and the frequency of cfDNA fragments with blunt ends. The x-axis shows the fetal DNA fraction based on the SNP, and the y-axis shows the frequency of blunt ends.
[0269] As shown in the figure, placental tissue showed higher DNASE1 expression compared to leukocytes ( Figure 22A ) and lower DFFB expression ( Figure 22B Higher DNASE1 is associated with higher 5' jagged ends, and lower DFFB is associated with fewer blunt ends in placental-derived fragments. The frequency of 5' jagged and blunt ends can be used to reflect the fetal DNA fraction. Figure 22CThe frequency of 5' overhanging jagged ends was positively correlated with the fetal DNA fraction, while Figure 22D The frequency of blunt ends was negatively correlated with that of fetal DNA, further suggesting that the jagged end pattern of plasma DNA may reflect the tissue of origin of those molecules.
[0270] Figure 23 23 is a flow chart of an exemplary process 2300 for determining a clinically relevant DNA fraction in a biological sample. The biological sample may include a plurality of free nucleic acid molecules. Each nucleic acid molecule in the plurality of nucleic acid molecules may be double-stranded having a first strand and a second strand. The biological sample may be obtained from an individual. At least some of the nucleic acid molecules in the plurality of nucleic acid molecules may have nucleotides on one strand that do not have a complementary portion on the other strand. The biological sample may be any biological sample described herein. In some embodiments, the process 2300 may be performed by system 2400. Figure 23 One or more process boxes.
[0271] Clinically relevant DNA can be fetal DNA, tumor DNA, or DNA from a tissue type. The tissue type can include placenta, liver, leukocytes, colon, kidney, lung, or any other tissue type described herein.
[0272] At block 2310, at least two different steps are possible. First, a first strand-specific classification of the first end of the nucleic acid molecule can be measured for each nucleic acid molecule in the plurality of nucleic acid molecules. The strand-specific classification can indicate whether the first strand or the second strand protrudes beyond the other strand, where the first strand is the 3' strand. Second, a first sequence end motif present at the first end of the nucleic acid molecule and a second sequence end motif present at the second end of the nucleic acid molecule can be determined for each nucleic acid molecule in the plurality of nucleic acid molecules. The sequence end motifs can have overhangs and / or blunt ends.
[0273] At block 2320, a first amount of first strand-specific classified nucleic acid molecules having a first strand protruding beyond a second strand can be determined, or a second amount of a first sequence terminal motif and a third amount of a second sequence terminal motif can be determined.
[0274] At block 2330, a parameter may be determined using the first quantity or the second and third quantities. For example, the process may include determining the first quantity, and determining the parameter may use the first quantity. As another example, the process may include determining the second and third quantities, wherein determining the parameter uses the second and third quantities.
[0275] In some embodiments, in addition to the amount of 3' overhangs, the amount of 5' overhangs can also be used. For example, process 2300 also includes determining a first amount, and determining a fourth amount of the same first-strand-specific classification of nucleic acid molecules having a second strand overhanging beyond the first strand, wherein the determining parameter uses the first amount and the fourth amount.
[0276] In some embodiments, the amount of blunt ends can be used in addition to the amount of 3' overhangs and / or the amount of blunt ends. For example, process 2300 also includes determining a fifth amount of nucleic acid molecules having the same first strand-specific classification as the first strand flush with the second strand, where the determination parameters use the first amount, the fourth amount of molecules having the second strand protruding beyond the first strand, and the fifth amount.
[0277] In some embodiments, overhanging ends at both ends of the nucleic acid molecule can be used. For example, process 2300 can also include, for each nucleic acid molecule in the plurality of nucleic acid molecules, measuring a second strand-specific classification at a second end of the nucleic acid molecule. The process can also include determining a fourth amount of nucleic acid molecules having the same second strand-specific classification, wherein the determining parameter uses the first amount and the fourth amount.
[0278] In some embodiments, a first strand-specific classification, a first sequence end motif, and a second sequence end motif can be used. For example, the process can include determining a first amount, a second amount, and a third amount. Determining a parameter can also include using the first amount, the second amount, and the third amount.
[0279] In some embodiments, the parameter can include a vector of quantities. The vector can include elements of any vector described herein, including elements of different combinations on the clear protrusion. Determining the parameter can include using quantities other than the specific quantities mentioned. For example, the parameter can be a ratio or difference to the quantities of all nucleic acid molecules.
[0280] At block 2340, the parameter can be compared to a reference value. The comparison can be performed similarly to herein, including any combination described in block 1508. The reference value can be a value determined from one or more control samples of clinically relevant DNA with a known fraction. A machine learning model can be used to perform the comparison of the parameter to the reference value. The machine learning model can include linear regression, logistic regression, deep recurrent neural network, Bayesian classifier, hidden Markov model (HMM), linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering of applications with noise (DBSCAN), random forest algorithm, or support vector machine (SVM).
[0281] The comparison can be used to determine the fraction of clinically relevant DNA in the biological sample at block 2350. If the parameter is statistically identical to the reference value, the grade of the condition can be determined to be the same as the one or more subjects associated with the reference value.
[0282] In some embodiments, the first nuclease can be identified as being differentially regulated in the target tissue type relative to at least one other tissue type in the plurality of tissue types. The clinically relevant DNA molecule can be from the target tissue type. For example, DNASE1 expression is relatively upregulated in placental tissue compared to the expression level of DNASE1 in leukocytes ( Figure 22A In another example, DNASE1L3 expression was relatively downregulated in HCC cells compared to liver tissues in healthy subjects.
[0283] The first nucleic acid enzyme can be determined to preferentially cut DNA into DNA molecules having a certain strand-specific classification and / or sequence end motif. In some cases, the cutting preference of the first nucleic acid enzyme is determined by analyzing a biological sample of another organism (e.g., a mouse). These strand-specific classifications and / or sequence end motifs can then be used to determine the fraction of clinically relevant DNA.
[0284] Process 2300 can include additional embodiments, such as any single embodiment or any combination of embodiments described below and / or in combination with one or more other processes described elsewhere herein. In a first embodiment, a reference value is determined from one or more calibration samples having known fractional concentrations of clinically relevant DNA molecules.
[0285] although Figure 23 Example blocks of process 2300 are shown, but in some implementations, process 2300 may include Figure 23 The blocks shown in FIG2300 may be additional blocks, fewer blocks, different blocks, or blocks arranged differently than those shown in FIG2300. Additionally or alternatively, two or more blocks of process 2300 may be executed in parallel.
[0286] V. Treatment
[0287] A. Other screening methods
[0288] Based on any classification, e.g., with respect to the concentration fraction of pathological or clinically relevant DNA, the subject can be referred to an additional screening modality, e.g., using chest X-ray, ultrasound, computed tomography, magnetic resonance imaging, or positron emission tomography. Such screening can be performed for cancer.
[0289] B. Treatment Options
[0290] Embodiments of the present disclosure can accurately predict disease recurrence (e.g., an increase in tumor DNA score after a decrease, classification of cancer as present after classification as absent), thereby facilitating early intervention and selection of appropriate treatment to improve the subject's disease outcome and overall survival. For example, if a subject's corresponding sample can predict disease recurrence, intensive chemotherapy can be selected for that subject. In another example, biological samples of a subject who has completed initial treatment can be sequenced to identify viral DNA that is predictive of disease recurrence. In such an example, an alternative treatment regimen (e.g., a higher dose) and / or a different treatment can be selected for the subject because the subject's cancer may be resistant to the initial treatment.
[0291] Embodiments may also include treating the subject in response to the classification of the determined pathological recurrence. For example, if the prediction corresponds to local regional failure, surgery may be selected as a possible treatment. In another example, if the prediction corresponds to distant metastasis, chemotherapy may be additionally selected as a possible treatment. In some embodiments, treatment includes surgery, radiotherapy, chemotherapy, immunotherapy, targeted therapy, hormone therapy, stem cell therapy, or precision medicine. Based on the determined recurrence classification, a treatment plan may be formulated to reduce the risk of injury to the subject and improve overall survival. Embodiments may also include treating the subject according to the treatment plan.
[0292] C. Types of Treatment
[0293] Embodiment can also be included in the pathology in the treatment patient after determining the classification of object.Can provide treatment according to the concentration score of the pathological grade, clinical relevant DNA or source tissue of determination.For example, can target identified mutation with specific medicine or chemotherapy.Source tissue can be used for guiding operation or any other form of treatment.And, pathological grade can be used for determining the aggressiveness of any type of treatment, and this can also be determined based on pathological grade.Can treat pathology (for example, cancer) by chemotherapy, medicine, diet, treatment and / or operation.In some embodiments, the value of parameter (for example, amount or size) exceeds reference value more, and treatment may be more aggressive.
[0294] Treatment may include resection. For bladder cancer, treatment may include transurethral resection of bladder tumors (TURBT). This procedure is used for diagnosis, staging, and treatment. During a TURBT, a surgeon inserts a cystoscope through the urethra and into the bladder. A tool with small coils, a laser, or high-energy electricity is then used to remove the tumor. For patients with non-muscle invasive bladder cancer (NMIBC), TURBT may be used to cure or eliminate the cancer. Another treatment may include radical resection and lymph node dissection. Radical resection is the removal of the entire bladder and possibly surrounding tissue and organs. Treatment may also include a urinary diversion. A urinary diversion is when the bladder is removed as part of treatment and the doctor creates a new path for urine to drain out of the body.
[0295] Treatment may include chemotherapy, which is the use of drugs to destroy cancer cells, usually by stopping them from growing and dividing. Drugs may include, for example, but not limited to, mitomycin-C (available as a generic drug), gemcitabine (Gemzar), and thiotepa (Tepadina), which are used for intravesical chemotherapy. Systemic chemotherapy may include, for example, but not limited to, cisplatin gemcitabine, methotrexate (Rheumatrex, Trexall), vinblastine (Velban), doxorubicin, and cisplatin.
[0296] In some embodiments, treatment may include immunotherapy. Immunotherapy may include immune checkpoint inhibitors that block a protein called PD-1. Inhibitors may include, but are not limited to, atezolizumab (Tecentriq), nivolumab (Opdivo), avelumab (Bavencio), durvalumab (Imfinzi), and pembrolizumab (Keytruda).
[0297] Treatment options may also include targeted therapy. Targeted therapy is a treatment that targets cancer-specific genes and / or proteins that contribute to cancer growth and survival. For example, erdafitinib is an orally administered drug approved to treat people with locally advanced or metastatic urothelial cancer in which mutations in the FGFR3 or FGFR2 genes contribute to continued growth or spread of cancer cells.
[0298] Some treatment methods can include radiation therapy. Radiation therapy uses high-energy x-rays or other particles to destroy cancer cells. In addition to each individual treatment, a combination of these treatments described herein can also be used. In some embodiments, when the value of a parameter exceeds a threshold value (which itself exceeds a reference value), a combination of treatments can be used. The information about treatment in the references is incorporated herein by reference.
[0299] VI. System
[0300] Figure 24A measurement system 2400 according to an embodiment of the present disclosure is shown. The system as shown includes a biological object 2405, such as a biological sample of an organism (e.g., a human), within an analytical device 2410, wherein a transmitter 2408 can transmit waves to the biological object 2405. For example, the biological object 2405 can receive a magnetic field and / or radio waves from the transmitter 2408 to provide a signal of a physical property 2415. The biological object 2405 may include an object treated with an enzyme, a marker, or a primer or other reagent to facilitate detection. An example of an analytical device may be a sequencing device. The analytical device 2410 may include multiple modules.
[0301] The physical characteristic feature 2415 (e.g., light intensity, voltage, or current) from the biological object is detected by detector 2420. Detector 2420 can measure at intervals (e.g., periodic intervals) to obtain data points that constitute the data signal. In one embodiment, an analog-to-digital converter converts the analog signal from the detector into a digital form multiple times. Analytical device 2410 and detector 2420 can form a determination system, for example, a sequencing system that obtains data according to the embodiments described herein. Data signal 2425 is sent from detector 2420 to logic system 2430. As an example, data signal 2425 can be used to determine the characteristics of nucleotides in a biological object. Data signal 2425 can include various measurements performed simultaneously, for example, different signals for different regions of biological object 2405, and therefore data signal 2425 can correspond to multiple signals. Data signal 2425 can be stored in local memory 2435, external memory 2440, or storage device 2445.
[0302] The logic system 2430 may be or may include a computer system, an ASIC, a microprocessor, a graphics processing unit (GPU), etc. It may also include or be coupled to a display (e.g., a monitor, an LED display, etc.) and a user input device (e.g., a mouse, a keyboard, buttons, etc.). The logic system 2430 and other components may be part of a standalone or network-connected computer system, or they may be directly attached to or integrated into a device (e.g., an imaging system) including the detector 2420 and / or the analysis device 2410. The logic system 2430 may also include software executed in the processor 2450. The logic system 2430 may include a computer-readable medium that stores instructions for controlling the measurement system 2400 to perform any of the methods described herein. For example, the logic system 2430 may provide commands to a system including the analysis device 2410 to perform a magnetic emission or other physical operation.
[0303] The measurement system 2400 may also include a treatment device 2460 that can provide treatment to the subject. The treatment device 2460 can determine the treatment and / or be used to perform the treatment. Examples of such treatments can include surgery, radiation therapy, chemotherapy, immunotherapy, targeted therapy, hormone therapy, stem cell transplantation, and implantation of radioactive seeds. The logic system 2430 can be connected to the treatment device 2460, for example, to provide the results of the methods described herein. The treatment device can receive input from other devices, such as imaging devices and user input (e.g., to control the treatment, such as controls on a robotic system).
[0304] Any computer system mentioned herein can utilize any suitable number of subsystems. An example of such a subsystem is shown in FIG14 as a computer system 10. In some embodiments, the computer system includes a single computer device, wherein the subsystem can be a component of the computer device. In other embodiments, the computer system can include multiple computer devices, each of which is a subsystem with internal components. The computer system can include desktop and laptop computers, tablet computers, mobile phones, and other mobile devices.
[0305] The subsystems shown in FIG14 are interconnected via a system bus 75. Subsystems such as a printer 74, a keyboard 78, a storage device 79, a monitor 76 coupled to a display adapter 82 (e.g., a display screen such as an LED), and other subsystems are shown. Peripheral devices and input / output (I / O) devices coupled to an I / O controller 71 can be connected to the computer system via any number of methods known in the art, such as an I / O port 77 (e.g., USB, Lightning). For example, an I / O port 77 or an external interface 81 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect the computer system 10 to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection via the system bus 75 allows the central processing unit 73 to communicate with each subsystem and control the execution of multiple instructions from the system memory 72 or the storage device 79 (e.g., a fixed disk such as a hard drive or optical disk), as well as the exchange of information between the subsystems. The system memory 72 and / or the storage device 79 can be embodied as computer-readable media. Another subsystem is a data collection device 85, such as a camera, microphone, accelerometer, etc. Any data mentioned in this article can be exported from one component to another and can be exported to the user.
[0306] The computer system may include multiple identical components or subsystems, for example, components or subsystems connected together via external interface 81, via internal interfaces, or via removable storage devices that can be connected from one component to another and removed from another component. In some embodiments, the computer systems, subsystems, or devices may communicate over a network. In this case, one computer may be considered a client and another computer may be considered a server, where each computer may be part of the same computer system. Each client and server may include multiple systems, subsystems, or components.
[0307] Aspects of the embodiments may be implemented in a modular or integrated manner using a general programmable processor, using hardware circuits (e.g., application specific integrated circuits or field programmable gate arrays) and / or using computer software in the form of control logic. As used herein, a processor may include a single-core processor, a multi-core processor on the same integrated chip, or multiple processing units on a single circuit board or networked, as well as dedicated hardware. Based on the disclosure and teachings provided herein, one of ordinary skill in the art will know and understand other ways and / or methods of implementing the embodiments of the present disclosure using hardware and combinations of hardware and software.
[0308] Any software component or function described in this application can be implemented as a software code executed by a processor using any suitable computer language, such as Java, C, C++, C#, Objective-C, Swift, or using a scripting language such as Perl or Python, for example, conventional or object-oriented technology. The software code can be stored on a computer-readable medium as a series of instructions or commands, for storage and / or transmission. Suitable non-transient computer-readable media can include random access memory (RAM), read-only memory (ROM), magnetic media such as hard drives, or optical media such as compact discs (CDs) or DVDs (digital versatile discs) or blue-ray discs, flash memories, etc. Computer-readable media can be any combination of such storage or transmission devices.
[0309] Alternatively, the computer readable medium may be used to encode and transmit such a program using a carrier signal suitable for transmitting via wired, optical and / or wireless network transmissions that meet the various protocols including the Internet. Thus, the data signal with such program encoding can be used to create a computer readable medium. The computer readable medium encoded with program code can be packaged with compatible devices, or separately provided (for example, via the Internet download) with other devices. Any such computer readable medium can be located on or in a single computer product (for example, hard drive, CD or entire computer system), and can be present on or in the different computer products in a system or network. The computer system can include a monitor, printer or other suitable displays, for providing any result mentioned herein to the user.
[0310] Any method described herein can be performed in whole or in part with a computer system comprising one or more processors, which can be configured to perform steps. Therefore, embodiments can be directed to a computer system configured to perform any method steps described herein, and may have different components that perform corresponding steps or corresponding step groups. Although presented as numbered steps, the method steps herein can be performed at the same time or at different times or in a logically possible different order. In addition, the parts of these steps can be used together with the parts of other steps from other methods. Moreover, all or part of the steps can be optional. In addition, any step of any method can be performed with the module, unit, circuit or other means of the system for performing these steps.
[0311] As will be apparent to those skilled in the art after reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features that can be readily separated or combined with the features of any of the other several embodiments without departing from the scope or spirit of the disclosure.
[0312] The above description of the exemplary embodiments of the present disclosure is presented for the purpose of illustration and description, and is set forth to provide a person of ordinary skill in the art with a complete disclosure and description of how to make and use the embodiments of the present disclosure. It is not intended to be exhaustive or to limit the present disclosure to the precise form described, nor is it intended to represent that the experiments are all or only the experiments performed. Although the present disclosure has been described in some detail by way of illustration and example for the purpose of clarity of understanding, it will be readily apparent to a person of ordinary skill in the art, based on the teachings of the present disclosure, that certain changes and modifications may be made thereto without departing from the spirit or scope of the appended claims.
[0313] Therefore, the above only illustrates the principle of the present invention. It should be understood that those skilled in the art will be able to design various arrangements, which, although not explicitly described or shown in this article, embody the principle of the present invention and are included in the spirit and scope of the present invention. In addition, all examples and conditional language listed herein are mainly intended to help readers understand the principle of the present disclosure, and are not limited to these specifically listed examples and conditions. In addition, all statements describing the principles, aspects and embodiments of the present invention and its specific examples herein are intended to cover their structure and function equivalents. In addition, it is intended that such equivalents include currently known equivalents and equivalents developed in the future, i.e., any element of the performance of the same function developed regardless of the structure. Therefore, the scope of the present invention is not intended to be limited to the exemplary embodiments shown and described herein. On the contrary, the scope and spirit of the present invention are embodied by the appended claims.
[0314] The recitation of "a," "an," or "the" is intended to mean "one or more," unless expressly stated otherwise. The use of "or" is intended to mean an "inclusive or," and not an "exclusive or," unless expressly stated otherwise. Reference to a "first" component does not necessarily require the presence of a second component. Furthermore, reference to a "first" or "second" component does not limit the referenced components to a particular location, unless expressly stated otherwise. The term "based on" is intended to mean "based at least in part on."
[0315] The claims may be drafted to exclude any possible optional element. Thus, this statement is intended to serve as antecedent basis for use of such exclusive terminology as "solely," "only," and the like in connection with the recitation of claim elements or use of a "negative" limitation.
[0316] Where a range of values is provided, it will be understood that unless the context clearly indicates otherwise, the intermediate value between the upper and lower limits of the range is also specifically disclosed, to one-tenth of the lower limit unit. Each smaller range between any of the values or intermediate values in the described range and any other described or intermediate values in the described range is included in the embodiments of the present disclosure. The upper and lower limits of these smaller ranges may be independently included in or excluded from the range, and each range in which any one, neither, or both limits are included in the smaller range is also included in the present disclosure, subject to the limitations of any clearly excluded limits in the described range. Where the range includes one or two limits, the scope excluding any one or both of those included limits is also included in the present disclosure.
[0317] All patents, patent applications, publications, and descriptions mentioned herein are incorporated by reference in their entirety for all purposes to the same extent as if each individual publication or patent was specifically and individually indicated to be incorporated by reference, and are incorporated herein by reference to disclose and describe the methods and / or materials in connection with the cited publication. No admission is made that any of the foregoing is prior art.
Claims
1. A method for analyzing a biological sample containing a plurality of free nucleic acid molecules, the method comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: ligating a first hairpin adaptor at a first end of the nucleic acid molecule to the first strand of the nucleic acid molecule and the second strand of the nucleic acid molecule, wherein the first hairpin adaptor comprises a first sequence identifier that identifies a first length of zero or more nucleotides at the first end of the first hairpin adaptor that has no complementary portion at the second end of the first hairpin adaptor, and ligating a second hairpin adaptor to the first strand and the second strand at a second end of the nucleic acid molecule, wherein the second hairpin adaptor comprises a second sequence identifier that identifies a second length of zero or more nucleotides at the first end of the second hairpin adaptor that has no complementary portion at the second end of the first hairpin adaptor, thereby generating a plurality of ligated nucleic acid molecules; performing rolling circle amplification on a first subset of the plurality of ligated nucleic acid molecules to form a plurality of concatemers; as well as Each concatemer in the plurality of concatemers is sequenced to identify a corresponding first sequence identifier and a corresponding second sequence identifier.
2. The method of claim 1, wherein sequencing occurs simultaneously with performing rolling circle amplification.
3. The method of claim 1 , further comprising: determining, using the first sequence identifier, that an overhang of a first length is present at a first end of a nucleic acid molecule in a first subset of the plurality of nucleic acid molecules, and The second sequence identifier is used to determine that a second overhang is present at a second end of a nucleic acid molecule in the first subset of the plurality of nucleic acid molecules.
4. The method of claim 1 , further comprising: adding an exonuclease to the plurality of ligated nucleic acid molecules to remove a second subset of the plurality of ligated nucleic acid molecules, wherein: the first subset of the plurality of nucleic acid molecules does not include any nucleic acid molecules in the second subset, For each nucleic acid molecule in the second subset, either: The corresponding nucleic acid molecule does not completely hybridize to the corresponding first hairpin aptamer or the corresponding second hairpin aptamer, or The corresponding first hairpin adaptor or the corresponding second hairpin adaptor does not completely hybridize to the corresponding nucleic acid molecule.
5. The method of claim 1 , further comprising: determining, using the first sequence identifier, that a first sequence end motif is present at a first end of a nucleic acid molecule in a first subset of the plurality of nucleic acid molecules; as well as The second sequence identifier is used to determine that a second sequence end motif is present at a second end of a nucleic acid molecule in the first subset of the plurality of nucleic acid molecules.
6. The method of claim 1 , further comprising: for each nucleic acid molecule in the first subset having an overhang at the first end, determining whether the 5' strand or the 3' strand overhangs the other strand using the corresponding first sequence identifier, and For each nucleic acid molecule in the first subset having an overhang at the second end, a determination is made using the corresponding second sequence identifier as to whether the 5' strand or the 3' strand overhangs the other strand.
7. The method of any one of claims 1 to 6, wherein: The biological sample is obtained from a female subject who is pregnant with a fetus, The method further comprises: selecting reads corresponding to a subset of the plurality of nucleic acid molecules in which either the 5' strand or the 3' strand protrudes beyond the other end, and A subset of the nucleic acid molecules is analyzed for characteristics of the fetus.
8. The method of claim 1, wherein: Each first hairpin aptamer in the plurality of first hairpin aptamers comprises a first cleavage site, and each second hairpin adaptor in the plurality of second hairpin adaptors comprises a second cleavage site, The method further comprises: Each concatemer in the plurality of concatemers is cleaved at a corresponding first cleavage site and a corresponding second cleavage site.
9. The method of claim 1, wherein each nucleic acid molecule in the second portion of the first subset has a respective first strand flush with a respective second strand at a respective first end.
10. The method of claim 1, wherein: Each nucleic acid molecule in the second portion of the first subset has a respective second strand at a respective first end that protrudes beyond the respective first strand, The corresponding first strand is the 5' strand, and The corresponding second strand is the 3' strand.
11. A method for analyzing a biological sample containing a plurality of free nucleic acid molecules, the method comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: ligating a first hairpin adaptor at a first end of the nucleic acid molecule to the first strand of the nucleic acid molecule and the second strand of the nucleic acid molecule, wherein the first hairpin adaptor comprises a first sequence identifier that identifies a first length of zero or more nucleotides at the first end of the first hairpin adaptor that has no complementary portion at the second end of the first hairpin adaptor, and ligating a second hairpin adaptor to the first strand and the second strand at a second end of the nucleic acid molecule, wherein the second hairpin adaptor comprises a second sequence identifier that identifies a first length of zero or more nucleotides at the first end of the second hairpin adaptor that has no complementary portion at the second end of the first hairpin adaptor, thereby generating a plurality of ligated nucleic acid molecules; adding an exonuclease to the plurality of ligated nucleic acid molecules to remove a first subset of the plurality of ligated nucleic acid molecules, wherein: For each nucleic acid molecule in the first subset, either: The corresponding nucleic acid molecule does not completely hybridize to the corresponding first hairpin adaptor or the corresponding second hairpin adaptor, or The corresponding first hairpin adaptor or the corresponding second hairpin adaptor is not completely hybridized to the corresponding nucleic acid molecule, Each ligated nucleic acid molecule in a second subset of the plurality of ligated nucleic acid molecules is sequenced to identify a corresponding first sequence identifier and a corresponding second sequence identifier, wherein the second subset remains in the biological sample after removing the first subset.
12. The method of claim 11, further comprising: determining, using the first sequence identifier, that an overhang of a first length is present at a first end of a nucleic acid molecule in a second subset of the plurality of nucleic acid molecules, and The second sequence identifier is used to determine that a second overhang is present at a second end of a nucleic acid molecule in a second subset of the plurality of nucleic acid molecules.
13. The method of claim 11, further comprising: determining, using the first sequence identifier, that a first sequence end motif of the overhang is present at a first end of a nucleic acid molecule in a second subset of the plurality of nucleic acid molecules; as well as The second sequence identifier is used to determine that a second sequence end motif of the overhang is present at a second end of a nucleic acid molecule in a second subset of the plurality of nucleic acid molecules.
14. The method of claim 11, further comprising: for each nucleic acid molecule in the second subset having an overhang at the first end, determining whether the 5' strand or the 3' strand overhangs the other strand using the corresponding first sequence identifier, and For each nucleic acid molecule in the second subset having an overhang at the second end, a determination is made using the corresponding second sequence identifier as to whether the 5' strand or the 3' strand overhangs the other strand.
15. The method of any one of claims 11 to 14, wherein: The biological sample is obtained from a female subject who is pregnant with a fetus, The method further comprises: selecting reads corresponding to a subset of the plurality of nucleic acid molecules in which either the 5' strand or the 3' strand protrudes beyond the other end, and A subset of the nucleic acid molecules is analyzed for characteristics of the fetus.
16. A method for analyzing a biological sample comprising a plurality of free nucleic acid molecules, each nucleic acid molecule in the plurality of nucleic acid molecules being double-stranded having a first strand and a second strand, the biological sample being obtained from an individual, at least some of the nucleic acid molecules in the plurality of nucleic acid molecules having a nucleotide on one strand that has no complementary portion on the other strand, the method comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: a first strand-specific classification measuring a property of a first end of the nucleic acid molecule, i.e., a strand-specific classification indicating whether the first strand or the second strand protrudes beyond the other strand; determining a jagged end value using the first strand-specific classification of the plurality of nucleic acid molecules; comparing the sawtooth end value with a reference value; and The comparison is used to determine the grade of the individual's condition.
17. The method of claim 16, further comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: measuring the second strand-specific classification of the second end of the nucleic acid molecule, in: The jagged end value is determined using the second strand-specific classification. The method of claim 16 , wherein the strand-specific classification further indicates the length of the overhang.
19. The method of claim 17, wherein: The sawtooth end values are the elements in the vector, The vector includes a plurality of elements, and The plurality of elements includes amounts of nucleic acid molecules in the following categories: the first end being blunt-ended and the second end being blunt-ended, a 5' overhang at the first end and a blunt end at the second end, a 3' overhang at the first end and a blunt end at the second end, a 5' overhang at the first end and a 3' overhang at the second end, a 5' overhang at the first end and a 5' overhang at the second end, and a 3' overhang at the first end and a 3' overhang at the second end; The method further comprises: comparing the vector to a reference vector; Wherein determining the level of the individual's condition uses a comparison of the vector with the reference vector.
20. The method of claim 19, wherein the plurality of elements comprises an assortment of nucleic acid molecules having sizes within one or more size ranges.
21. The method of any one of claims 16-20, wherein the condition is a defect in nuclease activity.
22. A method for analyzing a biological sample comprising a plurality of nucleic acid molecules, each nucleic acid molecule in the plurality of nucleic acid molecules being double-stranded having a first strand and a second strand, at least some of the nucleic acid molecules in the plurality of nucleic acid molecules having a nucleotide on one strand that has no complementary portion on the other strand, the biological sample being obtained from an individual, the method comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: determining a first sequence end motif of the first strand at a first end of the nucleic acid molecule, and determining a second sequence end motif of the second strand at the first end of the nucleic acid molecule; determining a first amount of nucleic acid molecules having a first combination of the first sequence end motif and the second sequence end motif at the first end; generating a value of a terminal module parameter using the first quantity; comparing the value of the terminal module parameter with a reference value; as well as The comparison is used to determine the grade of the individual's condition.
23. The method of claim 22, wherein the first strand has a 3' terminus at the first end, and the first strand protrudes beyond the second strand.
24. The method of claim 22, wherein the first end comprises a blunt end.
25. The method of claim 22, further comprising: determining a second amount of nucleic acid molecules having a second combination of a third sequence end motif and a fourth sequence end motif at said first end, in: The second quantity is used to generate the value of the terminal module parameter.
26. The method of claim 22, wherein: said plurality of nucleic acid molecules being a first plurality of nucleic acid molecules, The biological sample comprises a second plurality of nucleic acid molecules, and said first plurality of nucleic acid molecules comprises a subset of said second plurality of nucleic acid molecules, The method further comprises: For each nucleic acid molecule in the second plurality of nucleic acid molecules: determining a third sequence end motif on the first strand at the second end of the nucleic acid molecule, and determining a fourth sequence end motif on the second strand at the second end of the nucleic acid molecule, A second amount of nucleic acid molecules having a second combination of the third sequence end motif and the fourth sequence end motif is determined at the second end, wherein: The second quantity is used to generate the value of the terminal module parameter.
27. The method of any one of claims 22 to 26, further comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: measuring a first strand-specific classification of a first end of said nucleic acid molecule, i.e., a strand-specific classification indicating whether said first strand or said second strand protrudes beyond the other strand, Wherein determining the first amount comprises determining a first amount of classified nucleic acid molecules having the first combination and the first strand specificity.
28. The method of claim 27, further comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: measuring the second strand-specific classification of the second end of the nucleic acid molecule, Wherein determining the first amount comprises determining a first amount of nucleic acid molecules having the first combination, the first strand-specific classification, and the second strand-specific classification.
29. The method of claim 22, further comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: determining a third sequence end motif of the first strand at the second end of the nucleic acid molecule, and determining a fourth sequence end motif on the second strand at the second end of the nucleic acid molecule; in: The first combination is a combination of a first sequence terminal motif of the first end, a second sequence terminal motif of the first end, a third sequence terminal motif of the second end, and a fourth sequence terminal motif of the second end.
30. The method of claim 29, further comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: measuring a first strand-specific classification of a first end of the nucleic acid molecule, the strand-specific classification indicating whether the first strand or the second strand protrudes beyond the other strand, measuring the second strand-specific classification of the second end of the nucleic acid molecule, Wherein determining the first amount comprises determining a first amount of nucleic acid molecules having the first combination, the first strand-specific classification, and the second strand-specific classification.
31. The method of any one of claims 22-28, wherein the condition is cancer.
32. The method of any one of claims 22-28, wherein the condition is a deficiency in nuclease activity.
33. A method for analyzing a biological sample comprising a plurality of free nucleic acid molecules, the biological sample being obtained from an individual, each nucleic acid molecule in the plurality of nucleic acid molecules being double-stranded having a first strand and a second strand, at least some of the nucleic acid molecules in the plurality of nucleic acid molecules having a nucleotide on one strand that has no complementary portion on the other strand, the method comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: determining a first sequence end motif of the strand at a first end of the nucleic acid molecule, wherein the first end is the 3' end of the strand; determining a first amount of nucleic acid molecules having the first sequence end motif at the first end; generating a value of the terminal module parameter using the first quantity; comparing the value of the terminal module parameter with a reference value; as well as The comparison is used to determine the grade of the individual's condition.
34. The method of claim 33, further comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: determining a second sequence end motif of the strand at a second end of the nucleic acid molecule, in: The first amount is the amount of nucleic acid molecules having the first sequence terminal motif at the first end and the second sequence terminal motif at the second end.
35. The method of claim 33, wherein: said chain being a first chain, The terminal module parameter is a first terminal module parameter, and The reference value is a first reference value, The method further comprises: For each nucleic acid molecule in the plurality of nucleic acid molecules: determining a second sequence end motif of the second strand at the second end of the nucleic acid molecule, determining a second amount of nucleic acid molecules having the second sequence end motif at the second end, generating a value of a second terminal modulo parameter using the second quantity, comparing the value of the second terminal module parameter with a second reference value, and The comparison is used to determine the grade of the condition.
36. The method of any one of claims 33 to 35, further comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: measuring a first strand-specific classification of a first end of the nucleic acid molecule, the strand-specific classification indicating whether the first strand or the second strand protrudes beyond the other strand; Wherein determining the first amount comprises determining a first amount of classified nucleic acid molecules having the first combination and the first strand specificity.
37. The method of claim 36, wherein the first strand-specific classification is that the 3' strand protrudes beyond the 5' strand.
38. The method of claim 36, further comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: measuring a second strand-specific classification of a second end of the nucleic acid molecule; Wherein determining the first amount comprises determining a first amount of nucleic acid molecules having the first combination, the first strand-specific classification, and the second strand-specific classification.
39. A method for enriching a biological sample for clinically relevant DNA, wherein the biological sample comprises a plurality of free nucleic acid molecules, the plurality of nucleic acid molecules comprising the clinically relevant DNA and other DNA, each nucleic acid molecule in the plurality of nucleic acid molecules being double-stranded having a first strand and a second strand, the method comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: measuring a first strand-specific classification of a first end of the nucleic acid molecule, i.e., a strand-specific classification indicating whether the first strand or the second strand protrudes beyond the other strand; as well as Reads corresponding to a subset of classified nucleic acid molecules having the first strand specificity are selected to form an enriched sample.
40. The method of claim 39, wherein the subset of nucleic acid molecules has a first strand-specific assortment of the first strand protruding beyond the second strand.
41. The method of claim 40, wherein the first strand is the 3' strand.
42. The method of claim 39, wherein the first strand-specific classification comprises the first strand protruding beyond the second strand and the second strand protruding beyond the first strand.
43. The method of claim 39, further comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: measuring the second strand-specific classification of the second end of the nucleic acid molecule, wherein a subset of the nucleic acid molecules has a second strand-specific classification.
44. The method of claim 43, wherein: The first strand-specific classification of the subset of nucleic acid molecules indicates that the first strand of the nucleic acid molecules protrudes beyond the second strand, and The second strand-specific classification of the subset of nucleic acid molecules indicates that the second strand of the nucleic acid molecules protrudes beyond the first strand.
45. The method of claim 39, wherein the clinically relevant DNA is tumor DNA.
46. The method of claim 39, wherein the biological sample is obtained from a female subject pregnant with a fetus, and the clinically relevant DNA is fetal DNA.
47. The method of claim 45 or 46, further comprising analyzing a subset of the nucleic acid molecules to determine a classification of the disease stage.
48. The method of claim 39, further comprising analyzing a subset of the nucleic acid molecules to determine chromosomal aberrations or fetal haplotypes.
49. The method of any one of claims 39-48, further comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: determining a first sequence end motif at a first end of the nucleic acid molecule, Wherein selecting the reads corresponding to the subset of nucleic acid molecules comprises selecting reads corresponding to nucleic acid molecules having the first sequence end motif.
50. The method of claim 49, further comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: determining a second sequence end motif at a second end of the nucleic acid molecule, wherein selecting the reads corresponding to the subset of nucleic acid molecules comprises selecting reads corresponding to nucleic acid molecules having the second sequence end motif.
51. The method of claim 49 or 50, further comprising: determining a first quantity of said reading; determining a first parameter using a first quantity of said readings; as well as A characteristic of the biological sample is determined using the first parameter.
52. The method of claim 51, wherein the characteristic of the biological sample is the fractional concentration of clinically relevant DNA molecules in the biological sample.
53. The method of claim 51, wherein the first parameter is determined using the first quantity and another quantity of sequence reads.
54. The method of claim 51, wherein the characteristic of the biological sample is a level of aberration in the biological sample.
55. A method of determining the fraction of clinically relevant DNA in a biological sample comprising a plurality of free nucleic acid molecules, each nucleic acid molecule in the plurality being double-stranded having a first strand and a second strand, the biological sample being obtained from an individual, at least some of the nucleic acid molecules in the plurality having nucleotides on one strand that have no complementary portion on the other strand, the method comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: Measuring a first strand-specific classification of a first end of the nucleic acid molecule, i.e., a strand-specific classification indicating whether the first strand or the second strand protrudes beyond the other strand, wherein the first strand is the 3' strand, or determining that a first sequence terminal motif is present at a first end of the nucleic acid molecule and a second sequence terminal motif is present at a second end of the nucleic acid molecule; determining a first amount of first strand-specific assorted nucleic acid molecules having the first strand protruding beyond the second strand, or determining a second amount of the first sequence terminal motif and a third amount of the second sequence terminal motif; determining a parameter using either the first amount or both the second amount and the third amount; comparing the parameter with a reference value; as well as The comparison is used to determine the fraction of clinically relevant DNA in the biological sample.
56. The method of claim 55, wherein the reference value is determined from one or more calibration samples having known fractional concentrations of the clinically relevant DNA molecules.
57. The method of claim 55, further comprising determining the first quantity, wherein determining the parameter uses the first quantity.
58. The method of claim 55, further comprising: determining the first amount, and determining a fourth amount of classified nucleic acid molecules specific for the first strand wherein the second strand protrudes beyond the first strand, The first quantity and the fourth quantity are used in determining the parameter.
59. The method of claim 58, further comprising: determining a fifth amount of nucleic acid molecules of the first strand-specific classification having the first strand flush with the second strand, The first quantity, the fourth quantity and the fifth quantity are used in determining the parameter.
60. The method of claim 55, further comprising determining the second amount and the third amount, wherein determining the parameter uses the second amount and the third amount.
61. The method of claim 60, further comprising determining the first quantity, wherein determining the parameter uses the first quantity, the second quantity, and the third quantity.
62. The method of claim 55, further comprising: For each nucleic acid molecule in the plurality of nucleic acid molecules: measuring the second strand-specific classification of the second end of the nucleic acid molecule; as well as determining a fourth amount of classified nucleic acid molecules having said second strand specificity, The first quantity and the fourth quantity are used in determining the parameter.
63. The method of claim 55, wherein a machine learning model is used to perform the comparison of the parameter and the reference value.
64. The method of claim 63, wherein the machine learning model comprises linear regression, logistic regression, deep recurrent neural network, Bayesian classifier, hidden Markov model (HMM), linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering of applications with noise (DBSCAN), random forest algorithm, or support vector machine (SVM).
65. The method of claim 55, wherein the clinically relevant DNA is fetal DNA.
66. The method of claim 55, wherein the clinically relevant DNA is tumor DNA.
67. The method of claim 55, wherein the clinically relevant DNA is DNA from a tissue type.
68. The method of any one of claims 16-65, wherein the length of the protrusion is determined by the method of any one of claims 1-10.
69. The method of any of the above claims, wherein each nucleic acid molecule in the plurality of nucleic acid molecules has a size greater than a threshold size.
70. The method of any of the above claims, wherein each nucleic acid molecule in the plurality of nucleic acid molecules has a size that is less than a threshold size.
71. The method of claim 69 or 70, further comprising measuring the size of each nucleic acid molecule by aligning subsequences corresponding to the ends of the corresponding nucleic acid molecule to a reference genome.
72. The method of any of the above claims, wherein the condition is cancer, HCC, an autoimmune disease, or a pregnancy-related disorder.
73. The method of any of the above claims, wherein the reference value is determined from one or more subjects suffering from a certain grade of the condition or one or more healthy subjects.
74. A computer product comprising a non-transitory computer readable medium storing a plurality of instructions that, when executed, control a computer system to perform the method of any preceding claim.
75. A system comprising: The computer product of claim 74; as well as One or more processors for executing instructions stored on the computer-readable medium.
Citation Information
Patent Citations
Cell-free DNA damage analysis and its clinical applications
US20200056245A1
Biterminal DNA fragment types in cell-free samples and uses thereof
US20210238668A1
Nuclease-associated end signature analysis for cell-free nucleic acids
US20220010353A1
Methods using characteristics of urinary and other DNA
US20220177971A1