Fragmentomics in urine and plasma

Fragmentomic characterization methods enable comprehensive analysis of urinary and plasma samples to identify and evaluate transrenal and non-transrenal cell-free DNA characteristics and determine nuclease activity, addressing the lack of comprehensive tools in existing technologies.

JP2025539874APending Publication Date: 2025-12-09CENT FOR NOVOSTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025531090
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-29
Filing Date
2023-11-29
Publication Date
2025-12-09

Smart Images

  • Figure 2025539874000001
    Figure 2025539874000001
  • Figure 2025539874000002
    Figure 2025539874000002
  • Figure 2025539874000003
    Figure 2025539874000003
Patent Text Reader

Abstract

Fragmentomic signatures provide various characteristics of a sample (e.g., urine or plasma) and / or a subject. Fragmentomic signatures of urinary cell-free DNA can be used to determine the relative contribution or enrichment of clinically relevant DNA (e.g., transrenal and non-transrenal urinary cfDNA types). Such measurements reflect glomerular permeability and can be used to monitor various diseases, such as kidney abnormalities. Fragmentomic signatures can include corrected urinary DNA concentration, size, and terminal motifs of urinary DNA molecules, and cfDNA molecules derived from open chromatin regions (OCRs) of one or more tissues. In addition, nuclease activity or other cfDNA fragmentation processes can be determined based on the relative contribution of different cfDNA cleavage profiles, which can also be used to determine the contribution of tissue-derived cfDNA, the level of pathology, and gestational age.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a PCT application of and claims the benefit of U.S. Provisional Patent Application No. 63 / 428,694, entitled "FRAGMENTOMICS IN URINE AND PLASMA," filed November 29, 2022, which is incorporated herein by reference in its entirety for all purposes. [Background technology]

[0002] Urinary and plasma cell-free DNA molecules contain DNA molecules released from different normal tissues / organs and malignant cells, but they exhibit different fragmentation patterns. For example, compared to plasma cell-free DNA, urinary cell-free DNA molecules have a shorter size profile enriched with sharper 10-bp periodic peaks (Tsui et al., PLOS ONE 2012;7:e48319). Using a mouse model with a genetic deletion of DNase, it was demonstrated that deoxyribonuclease 1-like 3 (DNASE1L3) is a major contributor to plasma cell-free DNA fragmentation (Han et al., Am J Hum Genet. 2020;106:202-214). In contrast to plasma samples, deoxyribonuclease 1 (DNASE1) was involved in shaping the cell-free DNA fragmentation profile in urine (Chen et al. PLOS Genet. 2022;18:e1010262).

[0003] Urinary cell-free DNA may also contain different types of DNA molecules with their respective characteristics. For example, there are "transrenal" urinary cell-free DNA molecules released from non-urinary systems (e.g., blood cells, liver, lungs, colon, heart, brain, spleen, stomach, and placental tissues) that reach the urinary system via the glomeruli of the kidney. In addition to transrenal cell-free DNA, there are also "non-transrenal" urinary cell-free DNA molecules that originate from the urinary system, such as the renal tubules, bladder, and urethra, and are released directly from the urinary system. However, there is a lack of methods for identifying the characteristics or reflecting the degree of transrenal and non-transrenal cell-free DNA from a given urine sample. Additionally, numerous studies have demonstrated that plasma terminal motif usage can signal the presence of various diseases, ranging from autoimmune diseases to multiple cancer types (Chan et al. Am J Hum Genet. 2020;107:882-894, Jiang et al. Cancer Discov. 2020;10:664-73). Therefore, comprehensively determining the usage levels of nucleases such as DNASE1L3, DNASE1, and DNA fragmentation factor subunit beta (DFFB) may be clinically meaningful. We reasoned that the use of terminal motif profiles could estimate the extent of nucleases involved in the generation of cell-free DNA molecules (i.e., nuclease usage levels) and enable monitoring of nuclease activity across different pathophysiological conditions. However, tools that enable comprehensive evaluation of various DNA nucleases in a single analysis are lacking. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Tsui et al.,PLOS ONE 2012;7:e48319 [Non-patent document 2] Han et al.Am J Hum Genet.2020;106:202-214 [Non-patent document 3] Chen et al.PLOS Genet.2022;18:e1010262 Summary of the Invention [Means for solving the problem]

[0005] Methods, devices, and systems are provided for fragmentomic characterization of samples to determine various properties of the sample and / or subject. Various methods can be used for urine and / or plasma samples.

[0006] For example, for urine samples, the fractional concentration or enrichment of clinically relevant DNA types, such as transrenal and non-transrenal urinary cfDNA, can be provided using fragmentomic signatures of urinary cell-free DNA. Measuring the fragmentomic signature of urinary DNA can also reflect glomerular permeability and can be used to monitor various diseases, such as kidney abnormalities. Fragmentomic signatures can include corrected urinary DNA concentration, size, and terminal motifs of urinary DNA molecules, and cfDNA molecules from open chromatin regions (OCRs) of one or more tissues.

[0007] In other embodiments (e.g., for urine, plasma, or other acellular samples), nuclease activity or other fragmentation processes of cfDNA can be determined based on the relative contributions of different profiles of cfDNA cleavage, which can also be used to determine the fractional concentration of cfDNA from tissue, the level of pathology, and gestational age.

[0008] These and other embodiments of the present disclosure are described in detail below. For example, other embodiments are directed to systems, devices, and computer-readable media related to the methods described herein.

[0009] A better understanding of the nature and advantages of embodiments of the present disclosure may be obtained with reference to the following detailed description and accompanying drawings.

[0010]

[0013] Other features and advantages of the present disclosure will become apparent by reference to the remaining portions of the specification, including the drawings and claims. Further features and advantages of the present disclosure, as well as the structure and operation of various embodiments thereof, are described in detail below with reference to the accompanying drawings, in which like reference numbers may indicate identical or functionally similar elements. [Brief explanation of the drawings]

[0011] [Figure 1] 1 illustrates an exemplary scheme for characterizing transrenal and non-transrenal DNA in urine samples, according to some embodiments. [Figure 2] FIG. 1 shows a schematic illustrating the determination of transrenal urinary cell-free DNA contribution using fragmentomic signatures. [Figure 3] 1 illustrates a process for obtaining sequencing data from a urine sample, according to some embodiments. [Figure 4] 1 shows a set of graphs illustrating a comparison of urinary cell-free DNA concentrations before and after in vitro incubation between three urine collection groups. [Figure 5] 1 shows a set of graphs identifying a comparison of urinary cell-free DNA size profiles before and after in vitro incubation between control, EDTA, and stabilizer groups. [Figure 6] 1 shows a graph identifying the size difference between fetal DNA and maternal DNA in a urine sample, according to some embodiments. [Figure 7] FIG. 1 illustrates an exemplary schematic diagram showing the biological process of converting plasma DNA to transrenal DNA, according to some embodiments. [Figure 8] 1 shows a graph identifying the relationship between glomerular basement membrane permeability and transrenal DNA size, according to some embodiments. [Figure 9] Graphs identifying distinct terminal motifs identified from fetal-specific and shared cell-free DNA in maternal urine are shown. [Figure 10]1 shows a set of particular graphs illustrating the relationship between the fractional concentration of fetal DNA with CC-terminated fragments and the fractional concentration of fetal DNA with different sizes in urinary cell-free DNA, according to some embodiments. [Figure 11] FIG. 1 illustrates an exemplary diagram identifying various features of DNA molecules derived from open chromatin regions, according to some embodiments. [Figure 12] FIG. 1 illustrates an exemplary diagram for determining the amount of urinary cell-free DNA corresponding to open chromatin regions, according to some embodiments. [Figure 13] A set of box plots of the O / E ratios of fetal-specific and shared cell-free DNA molecules in urine and plasma samples are shown. [Figure 14] 1 illustrates a set of graphs identifying enrichment of DNA molecules within open chromatin regions in fetal-specific DNA from urine samples, according to some embodiments. [Figure 15] 1 illustrates a set of graphs identifying non-enrichment of DNA molecules within open chromatin regions in fetal-specific DNA from plasma samples, according to some embodiments. [Figure 16] 1 shows a graph identifying the correlation between fetal DNA fraction in maternal urine and the O / E ratio of all urinary cell-free DNA fragments. [Figure 17] 1 shows a graph illustrating the correlation between fetal DNA fraction in maternal urine and O / E ratio of urinary cfDNA fragments from placenta-specific DHS. [Figure 18] 1 shows a graph identifying the normalized end density of total urinary cell-free DNA in OCR of maternal urine and plasma samples. [Figure 19] 1 shows a set of graphs identifying a comparison between the end density of cell-free DNA in urine and the end density of cell-free DNA in plasma to determine the fractional concentration of fetal DNA, according to some embodiments. [Figure 20] 1 shows a set of graphs identifying the correlation between fetal DNA fraction and normalized end density of urinary cell-free DNA with different sizes, according to some embodiments. [Figure 21]1 shows a flowchart for estimating the fractional concentration of clinically relevant DNA molecules in a urine sample of a subject, according to some embodiments. [Figure 22] 1 shows a set of graphs identifying the correlation between fetal DNA fraction and the percentage of urinary cell-free DNA fragments carrying CC ends. [Figure 23] 1 illustrates a technique that uses probes to enrich for a set of one or more terminal motifs. [Figure 24] Another technique using probes and beads to enrich for a set of one or more terminal motifs is illustrated. [Figure 25] 1 shows a flowchart for enriching a urine sample for clinically relevant DNA based on terminal motif characteristics of urinary cell-free DNA, according to some embodiments. [Figure 26] 1 shows a set of graphs identifying enrichment of fetal DNA using urinary cell-free DNA with various fragmentomic properties, according to some embodiments. [Figure 27] 1 shows a bar graph identifying enrichment of transrenal urinary cell-free DNA using selective analysis of fragments with different fragmentomic features. [Figure 28] 1 shows a flowchart for enriching urine samples for clinically relevant DNA based on terminal motifs, open chromatin regions, and size of urinary cell-free DNA, according to some embodiments. [Figure 29] 1 shows a set of graphs identifying O / E ratio analysis in patients with RCC. [Figure 30] Fragmentomic analysis of transrenal DNA in patients with proteinuria using blood-specific DHS. [Figure 31] 1 shows fragmentomic analysis of transrenal DNA in pregnant women with preeclampsia using DHS. [Figure 32] 1 shows a flowchart for determining a classification of a kidney abnormality based on urinary cell-free DNA from open chromatin regions, according to some embodiments. [Figure 33]Fragmentomic analysis of transrenal DNA in patients with proteinuria and separately with preeclampsia. [Figure 34] 1 shows a flowchart for determining a classification of a kidney abnormality based on the size of urinary cell-free DNA, according to some embodiments. [Figure 35] Figure 1 shows urinary cfDNA concentration analysis of transrenal DNA in patients with separate proteinuria and preeclampsia. [Figure 36] 1 shows a flowchart for determining a classification of a kidney abnormality based on urinary cell-free DNA concentration, according to some embodiments. [Figure 37] Figure 1 shows ROC analysis in distinguishing patients with proteinuria and preeclampsia from healthy controls using fragmentomic features of transrenal DNA. [Figure 38] 1 shows a plot graph identifying the ranking of the frequency of certain terminal motifs present in urinary cell-free DNA molecules, according to some embodiments. [Figure 39] 1 shows a set of graphs identifying the observed terminal motif profiles of cell-free DNA molecules in mouse plasma and urine. [Figure 40] 1 shows a schematic workflow of an exemplary nuclease usage level analysis for cell-free DNA molecules. [Figure 41] Figure 1 shows the use of NMF analysis to identify the proportional contribution of each F profile (i.e., nuclease usage level) estimated from mouse cell-free DNA samples with different knockout genotypes. [Figure 42] Shown is a set of plots of six F profiles (A-F) estimated from mouse plasma and urinary cell-free DNA using NMF analysis. [Figure 43] 10 shows box plots identifying the proportional contribution of F profile I across different sample types, according to some embodiments. [Figure 44] FIG. 1 shows the relative frequency of cell-free DNA molecules across 256 terminal motifs for F profile I, according to some embodiments. [Figure 45]10 shows box plots identifying the proportional contribution of F profile II across different sample types, according to some embodiments. [Figure 46] FIG. 10 shows the relative frequencies of cell-free DNA molecules across 256 terminal motifs for F profile II, according to some embodiments. [Figure 47] 10 shows box plots identifying the proportional contribution of F profile III across different sample types, according to some embodiments. [Figure 48] FIG. 10 shows the relative frequencies of cell-free DNA molecules across 256 terminal motifs for F profile III, according to some embodiments. [Figure 49] 1 shows the relative frequencies of cell-free DNA molecules across 256 terminal motifs for F profiles IV-VI, according to some embodiments. [Figure 50] FIG. 1 shows a schematic diagram comparing terminal motif profiles of human subjects with reference F profiles determined based on mouse samples, according to some embodiments. [Figure 51] 1 shows the proportional contribution of F profiles across plasma and urine samples of human control subjects, according to some embodiments. [Figure 52] 1 shows proportional contributions of F profiles across plasma samples of normal and DNASE1L3-deficient human subjects, according to some embodiments. [Figure 53] 1 shows the proportional contribution of F profiles across urine samples of pregnant human subjects, according to some embodiments. [Figure 54] 1 shows a flowchart for determining a classification of nuclease activity based on the F profile of cell-free DNA molecules, according to some embodiments. [Figure 55] 1 shows a set of graphs identifying nuclease usage levels in urinary cell-free DNA of pregnant women. [Figure 56] 1 shows a flowchart for determining the fractional concentration of fetal DNA based on the F profile of cell-free DNA molecules, according to some embodiments. [Figure 57]1 shows a set of graphs identifying nuclease usage levels in cell-free DNA in plasma of pregnant women. [Figure 58] 1 shows a set of graphs identifying F profile analysis and oxidative stress levels in pregnant women. [Figure 59] 1 shows a flowchart for estimating gestational age based on an F profile of cell-free DNA molecules, according to some embodiments. [Figure 60] Box plots of F profile I (DNASE1L3) levels for healthy subjects, patients with DNASE1L3 deficiency, and parents of patients are shown. [Figure 61] 1 shows a set of graphs identifying nuclease usage levels in plasma cell-free DNA of subjects with systemic lupus erythematosus (SLE) and subjects without systemic lupus erythematosus. [Figure 62] 1 shows the proportional contribution of F profiles across normal, HBV, and HCC plasma samples, according to some embodiments. [Figure 63] 1 shows a set of graphs depicting nuclease utilization levels in plasma cell-free DNA of subjects with and without HCC. [Figure 64] 1 shows a bar graph identifying oxidative stress levels in blood samples from controls and HCC patients. [Figure 65] 1 shows a set of graphs providing F profile analysis and oxidative stress levels in CRC patients. [Figure 66] Box plots identifying the contribution of F-profile VI in NPC patients before and during chemoradiotherapy with cisplatin are shown. [Figure 67] 1 shows a flowchart for determining a classification of a level of pathology based on an F profile of cell-free DNA, according to some embodiments. [Figure 68] 1 illustrates a measurement system according to an embodiment of the present invention. [Figure 69] 1 shows a block diagram of an exemplary computer system usable with systems and methods according to embodiments of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0012] term A "tissue" corresponds to a group of cells grouped together as a functional unit. Two or more types of cells can be found within a single tissue. Different types of tissue can consist of different types of cells (e.g., liver cells, alveolar cells, or blood cells), but can also correspond to tissues from different organisms (mother vs. fetus) or healthy cells vs. tumor cells. A "reference tissue" can correspond to the tissue used to determine tissue-specific methylation levels. Multiple samples of the same tissue type from different individuals can be used to determine the tissue-specific methylation level of that tissue type.

[0013] A "biological sample" refers to any sample obtained from a subject (e.g., a human (or other animal), such as a pregnant woman, an individual with or suspected of having cancer or other disorder, an organ transplant recipient, or a subject suspected of having a disease process involving an organ (e.g., the heart in myocardial infarction, or the brain in stroke, or the hematopoietic system in anemia)) and containing one or more nucleic acid molecules of interest. A biological sample can be a bodily fluid, such as blood, plasma, serum, urine, vaginal fluid, fluid from edema (e.g., of the testes), vaginal washings, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple discharge, aspirates from various parts of the body (e.g., thyroid, mammary glands), intraocular fluid (e.g., aqueous humor), etc. Stool samples can also be used. In various embodiments, the majority of the DNA in a biological sample enriched for cell-free DNA (e.g., a plasma sample obtained via a centrifugation protocol) can be cell-free, e.g., greater than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA can be cell-free. The centrifugation protocol can include, for example, 1,600 g x 10 minutes, obtaining a fluid portion, and recentrifuging, e.g., at 16,000 g for an additional 10 minutes, to remove residual cells. As part of the analysis of the biological sample, a statistically significant number of cell-free DNA molecules can be analyzed for the biological sample (e.g., to provide an accurate measurement). In some embodiments, at least 1,000 cell-free DNA molecules are analyzed. In other embodiments, at least 10,000, or 50,000, or 100,000, or 500,000, or 1,000,000, or 5,000,000 or more cell-free DNA molecules can be analyzed. At least the same number of sequence reads can be analyzed.

[0014] A "sequence read" refers to a chain of nucleotides sequenced from any portion or all of a nucleic acid molecule. For example, a sequence read can be a short chain of nucleotides (e.g., 20-150 nucleotides) sequenced from a nucleic acid fragment present in a biological sample, a short chain of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of an entire nucleic acid fragment. Sequence reads can be obtained in a variety of ways, for example, using sequencing techniques, or using probes in hybridization arrays or capture probes, such as those used in microarrays, or using amplification techniques such as polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification. As part of the analysis of a biological sample, at least 1,000 sequence reads can be analyzed. As other examples, at least 10,000, or 50,000, or 100,000, or 500,000, or 1,000,000, or 5,000,000, or more sequence reads can be analyzed. The amount of sequence reads can be used as a proxy for the number of DNA fragments. To determine the number of DNA fragments from the amount of sequence reads, calculations can be performed to account for biases in paired-end sequencing and / or sequencing techniques.

[0015] A sequence read may include "end sequences" associated with the ends of a fragment. An end sequence may correspond to the outermost N bases of a fragment, e.g., 1 to 30 bases at the end of the fragment. If a sequence read corresponds to an entire fragment, the sequence read may include two end sequences. If paired-end sequencing provides two sequence reads corresponding to the ends of a fragment, each sequence read may include one end sequence.

[0016] A "sequence motif" can refer to a short repeating pattern of bases in a DNA fragment (e.g., a cell-free DNA fragment). A sequence motif can occur at the end of a fragment and therefore can be part of or include the end sequence. A "end motif" can refer to a sequence motif for end sequences that occur preferentially at the end of a DNA fragment, potentially for a particular type of tissue. A end motif can also occur just before or just after the end of a fragment, thereby still corresponding to an end sequence. A nuclease can have a particular cleavage preference for a particular end motif, as well as a second, most preferred cleavage preference for a second end motif.

[0017] The term "mapping" refers to the process of relating a sequence to a position or coordinate (e.g., a genomic coordinate) in a reference (e.g., a reference genome) with a known reference sequence, where the sequence is similar to the known reference sequence at the position in the reference. The degree of similarity can be measured or reported in terms of "mapping quality." As used herein, an example of mapping quality is that the mapping quality of a sequence X to a reported position or coordinate in the reference indicates that the probability of the sequence mapping to a different position is 10^(-X / 10) or less. For example, a mapping quality of 30 indicates that the probability of the sequence mapping to an alternative position is less than 0.1%.

[0018] A "reference genome" can be the entire genome sequence of a reference organism, a portion of a reference genome, a consensus sequence of many reference organisms, a compiled sequence based on different components of different organisms, or any other suitable reference sequence. The reference can also include information about variants of the reference known to be found within a population of organisms.

[0019] The "proportion" of DNA molecules terminating at a position relates to how frequently DNA molecules terminate at that position. Such a ratio can be referred to as "end density." The ratio can be based on the number of DNA molecules terminating at a position normalized to the number of DNA molecules analyzed. Normalization can also be based on the average, median, or total number of ends in a surrounding region. The surrounding region used for normalization can include, but is not limited to, 500, 1000, 3000, 5000 bp, etc., upstream and / or downstream of the position.

[0020] "Relative frequency" (also simply referred to as "frequency") can refer to a proportion (e.g., a percentage, fraction, or concentration). In particular, the relative frequency of a particular terminal motif (e.g., CCGA or only a single base) can provide the proportion of cell-free DNA fragments in a sample that are associated with the terminal motif CCGA, for example, by having a terminal sequence of CCGA.

[0021] The terms "control," "control sample," "background sample," "reference," "reference sample," "normal," and "normal sample" may be used interchangeably to generally describe a sample that does not have a particular pathological condition or that is otherwise healthy. In one example, a reference sample is a sample taken from a subject that does not have a pathological condition. A reference sample may be obtained from a subject or a database.

[0022] "Clinically relevant DNA" refers to DNA from a particular tissue source, which is measured, for example, to determine the fractional concentration of such DNA or to phenotype a sample (e.g., plasma). As used herein, clinically relevant DNA can refer to transrenal DNA present before passing through the kidney, as opposed to non-transrenal DNA (e.g., from the kidney or bladder). Examples of "clinically relevant" DNA include fetal DNA (e.g., from maternal plasma) and tumor DNA (e.g., from patient plasma). Another example includes measuring the amount of graft-associated DNA in the urine of a transplant patient. Further examples include measuring the fractional concentration of liver DNA fragments (or other non-hematopoietic or hematopoietic tissues, e.g., blood cells) in a sample or the fractional concentration of brain DNA fragments in cerebrospinal fluid.

[0023] A "calibration sample" can correspond to a biological sample in which the desired measurement (e.g., nuclease activity, fractional concentration of a clinically relevant nucleic acid, classification of a genetic disorder, or other desired characteristic) is known or determined via a calibration method, e.g., an ELISA to measure nuclease amount, or an assay that quantifies the rate of DNA digestion by a nuclease to measure nuclease activity (exemplary methods can include fluorimetric or spectrophotometric measurement of DNA amount before and after addition of a nuclease-containing sample, or in real time; another example is using a radial enzyme diffusion method). The fractional concentration of clinically relevant DNA (e.g., tissue-specific DNA fraction) can be known, as determined via a calibration method, e.g., using tissue-specific alleles. For example, in the case of tumors, fetuses, or transplants, alleles present in a tissue (e.g., the donor's genome) but absent in the healthy / maternal / recipient's genome can be used as a marker for the tissue corresponding to the clinically relevant DNA. As another example, tissue-specific methylation patterns can be used. Calibration samples can have discrete measurements (e.g., the amount of fragments with a particular terminal motif or having a particular size) to which it can be determined that the desired measurement can be correlated.

[0024] "Calibration data points" include "calibration values" (e.g., the amount of fragments having a particular terminal motif or having a particular size) and measurements or known values ​​desired to be determined for other test samples. Calibration values ​​can be determined from various types of data measured from the DNA molecules of a sample (e.g., the amount of fragments having a particular terminal motif or having a particular size). Calibration values ​​correspond to parameters that correlate with a desired characteristic, such as the classification of a genetic disorder, nuclease activity, or the effectiveness of anticoagulant administration. For example, calibration values ​​can be determined from measurements determined for a calibration sample whose desired characteristic is known. Calibration data points can be defined in various ways, for example, as discrete points or as a calibration function (also called a calibration curve or calibration surface). A calibration function can be derived from additional mathematical transformations of the calibration data points.

[0025] The term "fractional fetal DNA concentration" is used interchangeably with the terms "proportion of fetal DNA" and "fetal DNA fraction" and refers to the proportion of fetal DNA molecules present in a biological sample (e.g., a maternal plasma or serum sample) that originate from the fetus (Lo et al., Am J Hum Genet. 1998; 62: 768-775, Lun et al., Clin Chem. 2008; 54: 1664-1672). Similarly, tumor fraction or tumor DNA fraction can refer to the fractional concentration of tumor DNA in a biological sample.

[0026] A "site" (also referred to as a "genomic site") corresponds to a single site, which may be a single base position or a group of correlated base positions, e.g., a CpG site, a TSS site, a DNASE-hypersensitive site, or a larger group of correlated base positions. A "locus" may correspond to a region containing multiple sites. A locus may contain only one site, which would make the locus equivalent to the site in that context.

[0027] The term "open chromatin region (OCR)" refers to one or more sites corresponding to nucleosome-depleted regions (i.e., lack of histone-bound DNA). In some cases, an OCR contains one or more DNase 1 hypersensitive sites (DHSs) defined using DNase-seq (Meuleman et al. Nature. 2020;584:244-251). By way of example, an OCR can be defined based on sites identified using DNase-seq, sites identified using an assay for transposase-accessible chromatin using sequencing (ATAC-seq), transcription start sites (TSSs), CCCTC-binding factor (CTCF) sites, enhancer sites, histone modification-marked regions (e.g., H3K27ac, H3K4me3, etc.), and other nuclease-hypersensitive sites. In some cases, an OCR can be a region with a relative decrease in nucleosome occupancy. In some cases, an OCR can be tissue-specific. In various embodiments, at least 100, 500, 1,000, 5,000, or 10,000 OCRs may be used in the embodiments described herein.

[0028] The term "renal abnormality" refers to disorders that affect the kidneys and potentially other organs. By way of example, renal abnormalities can include renal cell carcinoma (RCC), renal syndrome, glomerulonephritis, Fabry disease, cystinosis, IgA nephropathy, IgM nephropathy, lupus nephritis, atypical hemolytic uremic syndrome (aHUS), polycystic kidney disease (PKD), Alport syndrome, interstitial nephritis, proteinuria, chronic kidney disease (CKD), acute kidney injury, preeclampsia, and the like.

[0029] A "terminal motif profile" can refer to the relationship between the terminal sequences (e.g., 1-30 bases) of cell-free DNA fragments (also simply referred to as DNA fragments) in a sample. Various relationships can be provided, such as the relationship between the amount of cell-free DNA and a specific terminal sequence (terminal motif), or the relationship between the relative frequency of cell-free DNA and a specific terminal sequence compared to one or more other terminal sequences. In some cases, terminal motif profiles are determined using other types of parameters, such as size. For example, terminal motif profiles can be provided in various ways that illustrate the amount of cell-free DNA fragments with one or more specific terminal sequences for a given size (single length or size range). A "reference terminal motif profile" or "F profile" refers to a terminal motif profile that can be generated by applying a factorization algorithm (e.g., nonnegative matrix factorization) to the relative frequencies of DNA molecules in a given biological sample across multiple terminal motifs (e.g., 256 terminal motifs).

[0030] The term "relative abundance" may generally refer to the ratio of a first amount of nucleic acid fragments having a particular characteristic (e.g., a specified length, terminating at one or more specified coordinates / end positions, or alignment (mapping) to a particular region of the genome) to a second amount of nucleic acid fragments having a particular characteristic (e.g., a specified length, terminating at one or more specified coordinates / end positions, or alignment (mapping) to a particular region of the genome). In one example, relative abundance may refer to the ratio of the number of DNA fragments terminating at a first set of genomic locations (e.g., open chromatin regions) to the number (e.g., mean or median) of DNA fragments terminating at a second set of genomic locations, which may be all genomic locations. Such relative abundance may be referred to as end density. In some embodiments, "relative abundance" is a type of separation value that relates the amount of cell-free DNA molecules terminating within one window of genomic locations (one value) to the amount of cell-free DNA molecules terminating within another window of genomic locations (the other value). The two windows may overlap but are of different sizes. In other implementations, the two windows do not overlap. Furthermore, the window may be one nucleotide long in width and therefore correspond to one genomic location. End density is a type of relative abundance. In some cases, the observed-to-expected (O / E) ratio is another type of relative abundance.

[0031] As used herein, the term "classification" refers to any number or other feature associated with a particular property of a sample. For example, a "+" sign (or the word "positive") may indicate that the sample is classified as having a deletion or amplification. Classification may be binary (e.g., positive or negative) or may have more levels of classification (e.g., a scale of 1 to 10 or 0 to 1).

[0032] As used herein, the term "parameter" refers to a numerical value that characterizes a quantitative data set and / or a numerical relationship between quantitative data sets. For example, a ratio (or a function of the ratio) between a first amount of a first nucleic acid sequence and a second amount of a second nucleic acid sequence is a parameter. The parameter can be used to determine any classification described herein, for example, for fetal, cancer, or transplantation analysis.

[0033] The terms "cutoff" and "threshold" refer to predetermined numbers used in an operation. For example, a cutoff size (or size threshold) can refer to a size above which a fragment is excluded. A threshold can be a value above or below which a particular classification is applied. Either of these terms can be used in either of these contexts. A cutoff or threshold can be a "reference value" or can be derived from a reference value that represents a particular classification or distinguishes between two or more classifications. A cutoff can be predetermined with or without reference to sample or subject characteristics. For example, a cutoff can be selected based on the age or sex of the subject being tested. A cutoff can be selected after and based on output of test data. For example, a particular cutoff can be used when sample sequencing reaches a certain depth. As another example, a reference subject with known classifications of one or more pathologies and measured characteristic values ​​(e.g., methylation levels, statistical size values, or counts) can be used to determine reference levels for distinguishing between different pathologies and / or pathology classifications (e.g., whether a subject has a pathology). The reference value can be selected to represent a value that lies between one classification (e.g., an average value) or two clusters of a metric (e.g., selected to obtain a desired sensitivity and specificity). As another example, the reference value can be determined based on statistical simulation of samples. Either of these terms can be used in any of these contexts. Such reference values ​​can be determined in various ways, as will be understood by those skilled in the art. For example, a metric can be determined for two different cohorts of subjects with different known classifications, and the reference value can be selected to represent a value that lies between one classification (e.g., an average value) or two clusters of a metric (e.g., selected to obtain a desired sensitivity and specificity). As another example, the reference value can be determined based on statistical simulation of samples. Particular values, such as cutoffs, thresholds, references, etc., can be determined based on the desired accuracy (e.g., sensitivity and specificity).

[0034] "Level of pathology" (or level of injury) can refer to the amount, degree, or severity of pathology associated with an organism. One example is cellular damage in the expression of nucleases. Another example of pathology is rejection of a transplanted organ. Other examples of pathology can include autoimmune attacks (e.g., lupus nephritis or multiple sclerosis, which damage the kidneys), inflammatory diseases (e.g., hepatitis), fibrotic processes (e.g., cirrhosis), fatty infiltration (e.g., fatty liver disease), degenerative processes (e.g., Alzheimer's disease), and ischemic tissue damage (e.g., myocardial infarction or stroke). A subject's healthy state can be considered a pathology-free category. The pathology can be cancer.

[0035] The term "level of cancer" may refer to whether cancer is present (i.e., present or absent), the stage of cancer, tumor size, whether there is metastasis, total tumor burden in the body, the cancer's response to treatment, and / or other measures of cancer severity (e.g., cancer recurrence). Cancer level may be a number or other indicator, such as a symbol, alphabetic letter, and color. A level may be zero. Cancer level may also include a premalignant or precancerous condition (state). Cancer level may be used in various ways. For example, screening may determine whether cancer is present in a person not previously known to have cancer. Evaluation may investigate a person diagnosed with cancer to monitor the progression of cancer over time, study the effectiveness of a therapy, or determine a prognosis. In one embodiment, prognosis may be expressed as the likelihood that a patient will die from cancer, or the likelihood that the cancer will progress after a certain period or time, or the likelihood or extent that the cancer will metastasize. Detection may mean "screening" or may mean checking whether a person with suggestive features of cancer (e.g., symptoms or other positive test) has cancer.

[0036] Gene names are typically written in italics. Human genes are typically written in all capital letters. Mouse genes may not be capitalized after the first letter. Proteins are conventionally written in all capital letters, without italics. As an example, a mouse may have a Dnase1l3 gene and a DNASE1L3 protein, while a human may have a DNASE1L3 gene and a DNASE1L3 protein.

[0037] A "machine learning model" (ML model) may refer to a software module configured to run on one or more processors and provide a classification or numerical value of one or more sample properties. ML models can be generated using sample data (e.g., training data) to predict test data. One example is an unsupervised learning model. Another example type of model is supervised learning, which can be used with embodiments of the present disclosure. Exemplary supervised learning models may include different approaches and algorithms, including analytical learning, statistical models, artificial neural networks, backpropagation, boosting (meta-algorithms), Bayesian statistics, case-based reasoning, decision tree learning, inductive logic programming, Gaussian process regression, genetic programming, group methods of data handling, kernel estimators, learning automata, learning classifier systems, minimum message length (decision trees, decision graphs, etc.), multiple linear subspace learning, naive Bayes classifiers, maximum entropy classifiers, conditional random fields, nearest neighbor algorithms, probabilistic approximately correct learning (PAC) learning, ripple down rules, knowledge acquisition methodologies, symbolic machine learning algorithms, less-than-symbolic machine learning algorithms, minimal complexity machines (MCM), random forests, ensembles of classifiers, ordinal classification, data preprocessing, handling imbalanced datasets, statistical relationship learning, or Proaftn, a multi-criteria classification algorithm. Models may include linear regression, logistic regression, deep recurrent neural networks (e.g., long short-term memory, LSTM), hidden Markov models (HMMs), linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering for applications with noise (DBSCAN), random forest algorithms, support vector machines (SVMs), or any of the models described herein. Supervised learning models can be trained in a variety of ways, using various cost / loss functions that define the error from known labels (e.g., least squares and absolute difference from known classifications), and various optimization techniques, e.g., backpropagation, steepest descent, conjugate gradient methods, and Newton and quasi-Newton techniques.

[0038] The term "about" or "approximately" can mean within an acceptable error range for a particular value, i.e., the limits of the measurement system, as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined. For example, "about" can mean within 1 or more than 1 standard deviation, according to convention within the art. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term "about" or "approximately" can mean within an order of magnitude, within 5-fold, or more preferably within 2-fold of a value. When a particular value is described in this application and claims, unless otherwise stated, the term "about" should be assumed to mean within an acceptable error range of the particular value. The term "about" can have the meaning commonly understood by one of ordinary skill in the art. The term "about" can refer to ±10%. The term "about" can refer to ±5%.

[0039] Various fragmentomic signatures of acellular samples (eg, urine and plasma) are used to determine various properties of the sample and / or subject.

[0040] For example, in the case of a urine sample, some embodiments can use fragmentomic signatures of urinary cell-free DNA to detect the contribution of transrenal and non-transrenal urinary cell-free DNA. Such measurements reflect glomerular permeability and can be used to monitor various disorders, such as, but not limited to, kidney cancer and kidney disease, as well as proteinuria and preeclampsia, which can be classified as types of kidney disorders.

[0041] In addition, such measurements can be used to determine the fractional concentration of clinically relevant cell-free DNA molecules and to enrich urine samples for clinically relevant DNA, including all types of transrenal DNA (e.g., liver DNA, lung DNA, colon DNA, heart DNA, and blood-derived DNA, e.g., leukocytes), fetal DNA, tumor DNA, or DNA from specific tissues other than the urinary tract, such as the kidney, ureter, and bladder. The relative contribution of transrenal cell-free DNA can be determined by determining the relative abundance of all or a representative sampling of open chromatin regions, e.g., OCRs (e.g., at least 100, 500, 1,000, 5,000, or 10,000 OCRs), from one or more tissues, or cell-free DNA molecules from one or more tissues that contribute any transrenal DNA. In other cases, the relative contribution or enrichment of transrenal cell-free DNA molecules can be determined based on cell-free DNA molecules from a urine sample having a particular size and / or terminal motif, and the corrected urine concentration.

[0042] For example, for terminal motifs, some embodiments can use the presence of C at the ends of cfDNA molecules to enrich samples for clinically relevant DNA. Thus, the contribution of transrenal cell-free DNA in urine can be determined using fragmentomic features such as terminal signatures from transrenal-specific open chromatin regions and cfDNA abundance in urinary cell-free DNA according to embodiments of the present disclosure.

[0043] Furthermore, different types of cell-free DNA cleavage can be analyzed simultaneously (i.e., together) using terminal motifs. Different types can be distinguished by representing different dimensions in the cleavage space, representing all possible nuclease activity in a subject. In the present disclosure, based on nuclease knockout mice and / or human subjects treated with various drugs, different types of cell-free DNA cleavage are linked to different fragmentation processes, including enzymatic and non-enzymatic breakage.

[0044] In contrast to techniques that focus on one specific nuclease activity each time using one terminal motif or several top-level terminal motifs (Serpas et al. Proc. Natl. Acad. Sci. USA 2019;116:641-649; Chan et al. Am. J. Hum. Genet. 2020;107:882-894; Chen et al. PLOS Genet. 2022;18:e1010262), embodiments of the present disclosure can simultaneously evaluate several nuclease activities or other fragmentation processes that may be involved (e.g., induced by chemoradiotherapy) based on the estimated relative contributions of different types of cell-free DNA breaks. The contribution of each type of cell-free DNA break can be determined by generating a set of F profiles representing the relative frequencies of terminal motifs for a given biological sample. In some cases, the set of F profiles can be generated by applying factorization (e.g., nonnegative matrix factorization) to the relative frequencies. Analysis of perturbation contributions may enable the detection and monitoring of a variety of diseases, including but not limited to cancer and immune disorders.

[0045] Thus, as described herein, terminal motifs (e.g., sample terminal motif profiles and reference profiles, referred to as reference F profiles) can be used in a variety of ways to determine characteristics of a sample and / or classification of a subject, such as determining the fractional concentration of clinically relevant DNA, the gestational age of a fetus, or the level of pathology in a subject.

[0046] I. Overview The glomerular basement membrane (GBM) allows plasma cell-free DNA to pass through the kidney and become transrenal cell-free DNA. Generally, smaller DNA molecules have greater GBM permeability than larger DNA molecules. For example, GBM permeability decreased as the size of molecules passing from plasma to urine increased (Lawrence et al. Proc Natl Acad Sci USA. 2017;114:2958-2963). Furthermore, nucleosome-depleted DNA molecules (e.g., DNA molecules from nucleosome-depleted regions) may have smaller molecular sizes than nucleosomal DNA of the same DNA length due to the attachment of histones to DNA. In some embodiments, the enrichment of nucleosome-depleted DNA in transrenal cell-free DNA is used to determine the level of glomerular permeability.

[0047] FIG. 1 illustrates an exemplary overview 100 for identifying characteristics of transrenal and non-renal DNA in a urine sample, according to some embodiments. As shown in FIG. 1 , urinary cell-free DNA includes transrenal DNA 102, which is from plasma and passes through the kidney into the urinary system. For example, transrenal DNA can be derived from liver, blood cells, tumor, or fetal DNA. Urinary cell-free DNA can also include non-renal DNA 104, which is from the urinary tract. For example, non-renal cell-free DNA can be derived from the urinary tract (e.g., the kidney and bladder). The two types of urinary cell-free DNA can be associated with different characteristics. By identifying differences between transrenal cell-free DNA (cfDNA) and non-renal cell-free DNA, the contribution of transrenal cell-free DNA or non-renal cell-free DNA can be determined from a given urine sample. Transrenal cfDNA generally, or certain types of transrenal DNA, can be considered clinically relevant DNA. Thus, the contribution of transrenal cfDNA, either total or of a specific type (e.g., fetal, tumor, or specific organ or blood cell type), can be determined from such differences.

[0048] FIG. 2 shows a schematic diagram 200 illustrating the determination of transrenal urinary cell-free DNA contribution using fragmentomic features. Both nucleosome-depleted cell-free DNA 202 (i.e., lacking associated proteins such as histones) and nucleosomal cell-free DNA molecules 204 are present in plasma 206. When plasma DNA molecules 206 pass through the GBM in the kidney, nucleosome-depleted cell-free DNA molecules 202 have higher permeability compared to nucleosomal DNA molecules 204, which have larger molecular sizes. On the other hand, when entering the urinary cell-free DNA pool, transrenal urinary cell-free DNA can still carry terminal signatures formed in plasma, for example, mediated by DNASE1L3. Thus, the contribution of transrenal cell-free DNA in urine can be determined using fragmentomic features, such as the abundance of cfDNA and terminal signatures from transrenal-specific open chromatin regions in urinary cell-free DNA, according to embodiments of the present disclosure.

[0049] 2 shows two illustrative examples for the urine side. A first urine sample 210 has a higher transrenal cfDNA contribution, as indicated by three of the seven DNA fragments being transrenal urinary cfDNA. A second urine sample 220 has a lower transrenal cfDNA contribution, as indicated by only one of the five DNA fragments being transrenal urinary cfDNA. Various embodiments can distinguish between such samples (and even between different types of transrenal DNA) as part of, for example, estimating the fractional concentration of clinically relevant DNA, determining the classification of pathologies such as kidney abnormalities, or detecting preeclampsia or proteinuria, which can be classified as types of kidney abnormalities.

[0050] However, challenges remain. The fragmentomechanisms of transrenal cell-free DNA are generally poorly understood. In addition, the fragmentation processes of urinary and plasma cell-free DNA may involve different nucleases (Han et al. Am J Hum Genet. 2020;106:202-214, Chen et al. PLOS Genet. 2022;18:e1010262). For example, DNASE1L3 is the predominant nuclease for generating C-terminal fragments in plasma (Serpas et al. Proc Natl Acad Sci USA. 2019;116:641-649), while DNASE1 is involved in generating T-terminal fragments in urine (Chen et al. PLOS Genet. 2022;18:e1010262).

[0051] Based on these differences, we hypothesized that transrenal urinary cell-free DNA molecules would carry the terminal motif signature of those present in plasma. In effect, analysis of terminal motifs in urinary cell-free DNA could be used to infer the contribution of transrenal cell-free DNA. For example, a high amount of urinary cell-free DNA bearing a C-terminus, a signature of plasma DNA, could suggest a high contribution of transrenal cell-free DNA.

[0052] Considering the above, it is necessary to understand the fragmentomic differences (e.g., size, terminal motifs) between transrenal and non-transrenal cell-free DNA. By identifying such differences, transrenal cell-free DNA contribution can be accurately estimated without any genetic or epigenetic information (e.g., SNPs from tumor tissue). The transrenal cell-free DNA contribution can then be applied to disease models. For example, subjects with renal function may have higher or lower transrenal contribution than normal subjects.

[0053] II. Example Urine Sample Preparation A variety of techniques can be used to prepare urine samples for cfDNA analysis. The techniques described below are merely examples, as will be understood by those skilled in the art.

[0054] A. Sample Collection Cell-free DNA molecules obtained from plasma and urine samples can be analyzed to determine differences between transrenal and non-transrenal DNA. For example, 192 human plasma and 18 urine cell-free DNA samples were sequenced using paired-end sequencing.In particular, plasma and urinary cell-free DNA samples were analyzed using (i) urinary cell-free DNA samples from pregnant women (n = 20), urinary cell-free DNA samples from preeclampsia patients (n = 5), and plasma cell-free DNA samples from pregnant women (n = 11) (median number of paired-end reads: 129.5 million, range: 30.1 million to 234.9 million); (ii) urinary cell-free DNA samples from renal cell carcinoma (RCC) (n = 16), proteinuria (n = 24), and controls (n = 34) (median number of paired-end reads: 25.03 million, range: 1,334 to 7, (iii) plasma cell-free DNA samples from eight healthy individuals, ten patients with DNASE1L3 disease-associated variants, and three parents of patients with mutant DNASE1L3 genes (median number of paired-end reads: 108 million, range: 40-162 million); (iv) plasma cell-free DNA samples from 24 SLE patients and 11 healthy individuals (median number of paired-end reads: 120 million, range: 18-208 million); (v) 38 healthy individuals, 17 chronic (vi) plasma cell-free DNA samples from patients with hepatitis B virus (HBV) but without hepatocellular carcinoma (HCC) (i.e., HBV carriers) and 34 patients with HCC (median number of paired-end reads: 38 million, range: 18-65 million); (vi) plasma cell-free DNA from 30 pregnant women across the first trimester (12-14 weeks, n=10), second trimester (20-24 weeks, n=10), and third trimester (38-40 weeks, n=10) (median number of paired-end reads: 103 million, range: 5, (vii) plasma cell-free DNA samples from 15 healthy control subjects, 25 colorectal cancer (CRC) patients without liver metastasis, and 24 CRC patients with liver metastasis (median number of paired-end reads: 40 million, range: 16-89 million); and (viii) plasma cell-free DNA samples from nasopharyngeal carcinoma (n=6) treated with cisplatin-based chemoradiotherapy and their pre-treatment paired patients (median number of paired-end reads: 5 million, range: 3-9 million).

[0055] B. Use of stabilizers in urine samples 3 illustrates a process 300 in which sequencing data is obtained from a urine sample, according to some embodiments. In block 302, a urine sample from a pregnant woman was used (e.g., a sample from the third trimester). Such a sample was used to determine whether transrenal contribution could be correlated with fetal DNA contribution.

[0056] Because DNASE1 activity is predominantly high in urine, sequencing urine samples can be challenging. If DNASE1 activity is not completely inhibited after urine collection, the in vitro sequential fragmentation caused by DNASE1 may disrupt the fragmentation pattern originally present in urine, potentially reducing the fragmentomic signal of urinary DNA fragments associated with certain diseases.

[0057] To address the above challenges, various collection and storage methods can be used to better preserve the original characteristics of urinary cell-free DNA. For example, as shown in block 304, a preservative can be added to obtain a sample that is preserved in block 306. Different urine collection methods can be used, including the addition of ethylenediaminetetraacetic acid (EDTA) and stabilizers. EDTA can inhibit the cleavage activity of DNASE1 family members by chelating magnesium and calcium, which are essential ions required for DNASE1 digestion. The stabilizer can potentially stabilize urinary DNA from degradation. Stabilizers include preservatives provided by Collipee, diazolidinyl urea (DU), dimethylol urea, 2-bromo-2-nitropropane-1,3-diol, 5-hydroxymethoxymethyl-1-aza-3,7-dioxabicyclo(3.3.0)octane and 5-hydroxymethyl-1-aza-3,7-dioxabicyclo(3.3.0)octane and 5-hydroxypoly[methyleneoxy]methyl-1-aza-3.7-dioxabicyclo(3.3.0)octane, bicyclic oxazolidines (e.g., Nuosept 95), DMDM ​​hydantoin, imidazolinone, methylparaben ... Other metal ion chelators may be zolidinyl urea (IDU), sodium hydroxymethylglycinate, hexamethylenetetramine chloroallyl chloride (quaternium-15), biocides (such as Bioban, Preventol, and Grotan), water-soluble zinc salts, EDTA, N,N'-bis(dithiocarboxy)piperazine (BDP), diethyldithiocarbamate (DDTC), iminodisuccinic acid (IDS), polyaspartic acid, S,S-ethylenediamine-N,N'-disuccinic acid (EDDS), methylglycine diacetic acid (MGDA), etc.

[0058] Compared with devices that do not contain stabilizers, urinary cell-free DNA can be better preserved in devices that contain stabilizers. As an illustrative example, urinary cell-free DNA samples from two control subjects were collected using different collection methods (without the addition of drugs, EDTA, and stabilizers) under different periods of room temperature in vitro incubation (e.g., 0 hours and 4 hours of incubation). The cell-free DNA concentration and size profile were compared between the three urine collection groups.

[0059] Figure 4 shows the results of three urine collection groups (a control group without any stabilizer added, 4 shows a set of graphs 400 illustrating a comparison of urinary cell-free DNA concentrations before and after in vitro incubation between the stabilizer groups (Sample 1, EDTA group, and stabilizer group). The term "GE" refers to genome equivalents. As shown in FIG. 4, the fold change in urinary cell-free DNA concentration at 4 hours of incubation versus 0 hours of incubation was lowest in the stabilizer groups (Sample 1: 1.61 GE / ml; Sample 2: 1.15 GE / ml) compared to the control group without any stabilizer added (Sample 1: 6.04 GE / ml; Sample 2: 1.98 GE / ml) and the EDTA group (Sample 1: 2.37 GE / ml; Sample 2: 2.09 GE / ml).

[0060] Autosomal DNA size profiles were further compared before and after in vitro incubation among the three collection groups.

[0061] FIG. 5 shows a set of graphs 500 identifying a comparison of urinary cell-free DNA size profiles before and after in vitro incubation between the control, EDTA, and stabilizer groups. The first set of graphs 502, 504, and 506 correspond to the cell-free DNA size profiles for the first sample, and the second set of graphs 508, 510, and 512 correspond to the cell-free DNA size profiles for the second sample. In the control group (graphs 502 and 508), the urinary cell-free DNA size profile changed significantly after 4 hours of incubation compared to 0 hours of incubation. In the EDTA group (graphs 504 and 510), the urinary cell-free DNA size profile changed only slightly after 4 hours of incubation compared to 0 hours of incubation. In the stabilizer group (graphs 506 and 512), the urinary cell-free DNA size profile showed the smallest (i.e., barely observable) change after 4 hours of incubation compared to 0 hours of incubation. The results shown in Figure 5 suggested that the use of a stabilizer-containing device may optimally preserve the fragmentation profile of urinary cell-free DNA molecules under room temperature conditions.

[0062] C. Cell-free DNA Extraction and Library Preparation As shown in Figure 3, cell-free DNA can be extracted from (stabilizer-treated) urine using the Wizard Plus Minipreps DNA Purification System (Promega) and guanidine thiocyanate (Sigma-Aldrich). Cell-free DNA was extracted from plasma samples using the QIAamp Circulating Nucleic Acid Kit (Qiagen) according to the manufacturer's protocol. Indexed DNA libraries were constructed using the TruSeq DNA Nano Library Prep Kit (Illumina) according to the manufacturer's instructions. Adapter-ligated DNA was enriched by PCR and then analyzed on an Agilent 4200 TapeStation (Agilent Technologies) for quality control and gel-based sizing. Libraries were quantified using the Qubit dsDNA High Sensitivity Assay Kit (Thermo Fisher Scientific) before sequencing.

[0063] D. DNA Sequencing and Alignment As further shown in Figure 3, the multiplexed DNA library was sequenced for paired-end reads on an Illumina platform. Other sequencing technologies, such as those described herein, may also be used. For example, single reads across entire DNA fragments may be determined. Sequences were assigned to corresponding samples based on 6-base index sequences. Paired-end reads from mouse plasma were aligned to the reference mouse genome (NCBI build 37 / UCSC mm9, non-repeat masked) or human reference genome (NCBI build 37 / hg19) using Short Oligonucleotide Alignment Program 2 (SOAP2) (Li et al. Bioinformatics 2009;25:1966-1967). As will be appreciated by those skilled in the art, any other alignment tool may also be used. In some implementations, mismatches of up to two nucleotides were allowed. Only paired-end reads aligned to the same chromosome in the correct orientation and spanning an insert size of less than 600 bp were retained for downstream analysis. Paired-end reads sharing the same start and end genomic coordinates were considered PCR duplicates and discarded from downstream analysis.

[0064] In some use cases, for example, in the case of a plasma sample, the genotype of buffy coat DNA from the mother can be paired with the corresponding placenta sample. This effectively determines the maternal and fetal genotypes. Genotyping is used to distinguish between fetal and maternal DNA molecules, allowing us to obtain an absolute standard for the fetal DNA fraction in a urine sample. This actual fetal DNA fraction also allows us to establish a recalibration curve to estimate the degree of transrenal DNA or renal permeability, assuming that higher renal permeability corresponds to more transrenal DNA.

[0065] III. Size characteristics of urinary cell-free DNA The size characteristics of urinary cfDNA were analyzed to illustrate the effect of fragment size on the ability of transrenal cfDNA fragments to pass through the kidney and enter the urine. Smaller sized molecules have been shown to have an increased ability to pass from the blood to the kidney.

[0066] FIG. 6 shows a graph 600 identifying the size difference between fetal DNA and maternal DNA in a urine sample, according to some embodiments. The size difference between fetal DNA and maternal DNA can be used as an example to determine transrenal DNA in a urine sample. For example, fetal DNA originates from the fetus and must pass through the kidney, so we can detect it in urine. In contrast, detected maternal DNA can include non-renal DNA, which may be contributed by the kidney, bladder, etc. Based on this distinction, we can extend the characteristics of fetal and maternal DNA to determine the size difference between transrenal DNA and non-renal DNA.

[0067] As shown in Figure 6, the majority of fetal-specific cell-free DNA molecules 620 (red) are less than 80 bp, significantly shorter than the shared cell-free DNA molecules 610 (blue). Shared cell-free DNA molecules 610 have the same allele shared between the mother's haplotype and one of the fetus's haplotypes. Fetal-specific cell-free DNA molecules 620 have a fetal-specific allele (inherited from the father) that is in one of the fetal haplotypes.

[0068] From the above, it is possible to consider whether the size characteristics of transrenal DNA can be correlated with those of fetal DNA.

[0069] FIG. 7 illustrates an example schematic 700 showing the biological process of converting plasma DNA to transrenal DNA, according to some embodiments. For example, to become transrenal DNA, plasma DNA from a blood vessel 702 passes through various tissue membranes to reach the kidney 704. For example, plasma DNA molecules pass through endothelial cells and glomerular basement membranes (GBM), as well as podocytes. As shown in FIG. 7, each of these biological structures has a different diameter. The kidney structure may be related to the size of the kidney's pores through which the plasma DNA molecules pass to become transrenal DNA. It is conceivable that plasma DNA molecules small enough to pass through the pores may ultimately become transrenal DNA.

[0070] Figure 8 shows a graph 800 identifying the relationship between glomerular basement membrane permeability and transrenal DNA size, according to some embodiments. The permeability of the glomerular basement membrane can be determined by molecular size. As shown in Figure 8, the x-axis is the radius of a given molecule, which corresponds to its size, and the y-axis identifies the renal permeability associated with the given molecule.

[0071] Specifically, the GBM permeability percentage, shown on the y-axis, identifies the percentage of molecules of a particular size that cross the GBM. For example, if the molecule is very small (e.g., 12 kDa), permeability is estimated to be approximately 50%. In contrast, as the molecule becomes larger (e.g., 150 kDa), renal permeability will drop significantly to approximately 10-15%. It is also known that nucleosomes typically have a size of 200 kDa / 5.5 nm (radius). Based on the size of nucleosomes, it can be assumed that DNA molecules wrapped in nucleosomes (and therefore attached to proteins) will be associated with lower GBM permeability compared to nucleosome-depleted DNA molecules. In effect, the size of transrenal DNA molecules that cross the GBM will likely be smaller compared to non-transrenal DNA molecules that originate directly from the urinary system.

[0072] IV. Terminal motif characteristics of urinary cell-free DNA In addition to size, the terminal sequences of urinary cell-free DNA were analyzed to determine that the terminal motifs of transrenal urinary cell-free DNA molecules differ from those of non-transrenal urinary cell-free DNA molecules. In some embodiments, a 4-mer terminal motif is defined as the terminal four nucleotides at the end of each 5' fragment of a cell-free DNA molecule, resulting in a total of 256 categories of 4-mer terminal motifs (i.e., 4 4 ) The median terminal motif frequency of the 256 terminal motifs was calculated separately for the fetal-specific and shared fragments in the maternal urine samples and ranked in descending order. Other terminal motifs may also be used, e.g., any K-mer terminal motif where K is 1, 2, 3, 4, 5, 6, 7, 8, 9, or more. As described herein, the terminal motifs (e.g., the sample terminal motif profile and the reference profile, referred to as the reference F profile) can be used in various ways to determine characteristics of a urine sample and / or classification of a subject, such as determining the fractional concentration of clinically relevant DNA, the gestational age of the fetus, or the level of pathology in the subject.

[0073] Figure 9 shows a graph 900 identifying different terminal motifs identified from fetal-specific and shared cell-free DNA in maternal urine. As shown in Figure 9, the median terminal motif frequency of 256 terminal motifs in fetal-specific shared cell-free DNA in maternal urine samples was calculated and ranked in descending order. The x-axis identifies the terminal motif ranking of the shared fragments. The y-axis identifies the motif ranking of the fetal-specific fragments. A higher ranking indicates a higher relative frequency for the corresponding terminal motif (e.g., CCTG). The colored areas indicate fetal-specific or shared DNA fragments, respectively, that have a preference for a particular 4-mer terminal motif.

[0074] Each of the top 10 motifs in both fetal-specific and shared cell-free DNA was labeled with the corresponding terminal motif sequence. The top 10 motifs in fetal-specific and shared urinary cell-free DNA were highlighted by red circles 902 and blue circles 904, respectively. The top 10 terminal motifs in fetal-specific cell-free DNA were dominated by C-terminal motifs (8 / 10), while the top 10 terminal motifs in shared cell-free DNA were enriched in T-terminal motifs (4 / 10). It has previously been identified that DNASE1L3 (which prefers to cleave C) is the predominant nuclease in plasma, and DNASE1 (which prefers to cleave T) is the predominant nuclease in urine. Based on the motif rankings above, it can be determined that fetal DNA corresponds to transrenal DNA. The data suggest that C-terminal-containing motifs can be used to represent transrenal urinary cell-free DNA, which can then be used to distinguish fetal DNA from maternal urine samples. As described below, some embodiments can use the presence of Cs at the ends of cfDNA molecules to enrich samples for clinically relevant DNA, such as total transrenal DNA, fetal DNA, tumor DNA, or DNA from specific tissues other than the kidney or bladder.

[0075] Figure 10 shows a set of graphs 1000 identifying the relationship between CC-terminated fragments in a urine sample and the fractional concentration of fetal DNA, according to some embodiments. As shown in Figure 10, the proportion of urinary cell-free DNA bearing CC termini increases proportionally to the fetal DNA fraction in the urine sample. This linear relationship is even more pronounced when the CC fragments are limited to fragments having a size of less than 80 base pairs. Therefore, the terminal motifs of a urine sample can be used to determine the fraction of fetal DNA.

[0076] V. Analysis of urinary cell-free DNA for OCR The open chromatin regions can be used to determine the characteristics of a urine sample and / or the classification of a subject. In some situations, the open chromatin regions can be associated with tissues that contribute to transrenal DNA, or with specific cell types (e.g., tissues from fetuses, tumors, transplanted organs, or the urinary tract, as well as other tissues such as blood, liver, and colon). For example, the abundance of cfDNA from a set of such regions can be used to estimate the fractional concentration of clinically relevant DNA, determine the classification of pathologies such as kidney abnormalities, or as part of detecting preeclampsia or proteinuria, which can be classified as types of kidney abnormalities.

[0077] The permeability of renal membranes (e.g., GBM) favors shorter DNA fragments. As a result, transrenal cell-free DNA that crosses the renal membrane is shorter than non-transrenal DNA fragments derived directly from the urinary system. In addition, cell-free DNA molecules bound to nucleosomes may have difficulty crossing the GBM because the nucleosome permeability is estimated to be approximately 10-15%. In contrast, nucleosome-depleted cell-free DNA molecules derived from open chromatin regions are not bound by any nucleosomes and can cross the GBM with higher permeability. Based on the above characteristics, identifying cell-free DNA molecules derived from open chromatin regions can be used to detect transrenal DNA in urine samples. In addition, the contribution of cell-free DNA molecules from open chromatin regions can be used to predict the classification of certain diseases as well as to determine the fractional concentration of fetal DNA.

[0078] A. Correlation between transrenal DNA and DNA from OCR Transrenal DNA can be correlated with DNA in open chromatin regions, according to some embodiments. Nucleosome-depleted cell-free DNA molecules in plasma have a smaller molecular size, allowing them to cross the GBM and convert to transrenal DNA. Based on this characteristic, it can be determined whether transrenal DNA is enriched in open chromatin regions corresponding to nucleosome-depleted regions. Such enrichment is described in later sections, for example, for all transrenal DNA or for certain tissue types.

[0079] Figure 11 illustrates an exemplary diagram 1100 for identifying various features of DNA molecules from open chromatin regions, according to some embodiments. Because DNase 1 has a cleavage preference within genomic regions that are relatively devoid of bound histones, open chromatin regions can be identified in various ways, for example, based on the location of DNase 1 hypersensitive sites (DHSs). For example, a given open chromatin region can be identified as a genomic region clustered with DNase 1-digested fragment ends. There are approximately 1 million DHS sites, and the median length of the region is approximately 200 base pairs. Such open chromatin regions contribute to 9.46% of the genome. Such DHSs (and corresponding OCRs) can be specific to a particular tissue, for example, placenta-specific DHSs.

[0080] As another example for identifying OCRs, DNase-Seq can be used to obtain DNA molecules from open chromatin regions. Specifically, DNA molecules from a urine sample can be digested using DNase 1, which preferentially cuts DNA molecules with hypersensitive sites, and then sequenced. The sequence reads can then be considered as DNA molecules from open chromatin regions. As a further example, OCRs can be identified from, but are not limited to, sites identified using DNase-Seq, sites identified using an assay for transposase-accessible chromatin using sequencing (ATAC-Seq), and transcription start sites (TSS).

[0081] B. Identification of tissue-specific OCRs OCRs can generally be determined and used for all tissues, tissues specific to renal tissues, or specific tissues. Different tissues generally have different regions of open chromatin within them. Therefore, a specific set of OCRs can be identified, for example, depending on the desired clinically relevant DNA. For example, whether a class of cfDNA is enriched or depleted in one or more OCRs can be determined, for example, to identify OCRs associated with one or more tissues.

[0082] FIG. 12 illustrates an exemplary diagram 1200 for determining the amount of urinary cell-free DNA corresponding to open chromatin regions, according to some embodiments. To analyze fragmentomic signatures based on urinary cell-free DNA from open chromatin regions, we collected urine samples from 14 pregnant women and examined the fetal cell-free DNA characteristics in maternal urine. The signatures of fetal-specific cell-free DNA molecules and shared cell-free DNA molecules can represent the signatures of transrenal cell-free DNA and non-renal cell-free DNA, respectively. Such representations are possible because transrenal cell-free DNA is shorter than non-renal cell-free DNA.

[0083] In the example shown in Figure 12, cfDNA containing any fetal-specific alleles (e.g., single nucleotide polymorphisms, SNPs) can be identified. The amount of such cfDNA within a genome-spanning window (e.g., 10, 20, 30, 40, 50, or 60 bp) can then be determined by aligning the sequence reads to the reference genome. The expected value can be determined as the number of fetal-specific SNPs in a region (where a region can include one or more windows) divided by the mean or median number of fetal-specific SNPs for all windows / regions. The observed value can be determined as the number of fetal-specific reads in a region divided by the mean or median number for all windows / regions. If the observed ratio is greater than the expected ratio, the region can be identified as an OCR because smaller fetal DNA fragments are more common. Such OCRs determined using fetal DNA will be specific to fetal tissue. However, DHS sites can be used for various tissues, or more generally, OCRs can be used for various tissues. Thus, DHS for a specific tissue or set of tissues can be used.

[0084] In an example implementation, to assess transrenal urinary cell-free DNA in pregnant women, we obtained a median fetal fraction of 0.31% (range: 0.20%–9.00%) in maternal urine samples. The fetal DNA fraction was estimated using a single-nucleotide polymorphism (SNP)-based approach in maternal urinary cell-free DNA (Yu et al. ClinChem. 2013;59:1228–1237). The contribution of nucleosome-depleted cell-free DNA can be indicated by the amount of DNA molecules derived from open chromatin regions (OCRs), where OCRs correspond to nucleosome-depleted regions (i.e., lack of histone-bound DNA). For illustration, we use DNase 1 hypersensitive sites (DHSs) defined from DNase-seq to represent OCRs (Meuleman et al. Nature. 2020;584:244–251).

[0085] From the OCR, the amount of nucleosome-depleted DNA molecules was determined by the number of sequenced cell-free DNA molecules aligned to the OCR. The amount of nucleosome-depleted DNA molecules can be a normalized value. For example, the amount of nucleosome-depleted DNA molecules can be translated into a percentage by dividing by the total number of sequenced molecules. Additionally or alternatively, the amount of nucleosome-depleted DNA molecules can be calculated by dividing the sequenced molecules (O) observed in the OCR (also referred to as the observed OCR-associated DNA contribution) by the expected OCR-associated value (E). This measurement is defined herein as the "O / E ratio."

[0086] The expected OCR-associated DNA contribution can correspond to the theoretical percentage of OCRs in a reference genome. For example, the observed value (O) can include the percentage of fragments aligned to the OCR among all fragments, and the expected value (E) as the theoretical percentage of OCRs in a reference genome (e.g., a human reference genome). In some cases, the expected OCR-associated DNA contribution for fetal-specific DNA can be calculated by the number of fetal-specific single nucleotide polymorphisms (SNPs) in the OCR normalized by the number of fetal-specific SNPs in all genomic regions. Such relative frequencies (e.g., percentages) can provide expected percentages that can be compared with the observed percentage of cell-free DNA molecules aligned to the OCR. SNPs can be obtained through genotyping. In some cases, the expected OCR-associated DNA contribution corresponds to the percentage of DNA molecules that fall within the OCR by random sampling.

[0087] The observed OCR-associated DNA contribution in fetal DNA can be calculated by the number of fetal-specific molecules aligned to the OCR normalized by the number of fetal-specific molecules aligned to all genomic regions. For O / E ratio analysis in non-transrenal urinary cell-free DNA, molecules carrying shared alleles between the fetal and maternal genomes were analyzed according to embodiments of the present disclosure. The O / E ratio was used to determine OCR enrichment. If the O / E ratio was close to 1, no OCR enrichment was observed. If the O / E ratio was greater than 1, OCR-associated DNA contribution increased. A higher O / E ratio may suggest a higher contribution of nucleosome-depleted DNA and is likely indicative of higher glomerular permeability. Other tissue-specific regions can be identified in a similar manner.

[0088] C. Quantification of transrenal cfDNA from OCR of urine and plasma The amount of cfDNA in the OCR region can be used to quantify transrenal DNA in urine samples. Fetal cfDNA is used as an example of transrenal DNA, but other examples would apply to tumors and other tissues that produce transrenal DNA.

[0089] Figure 13 shows a set of boxplots 1300 of the O / E ratios of fetal-specific and shared cell-free DNA molecules in urine (boxplot 1302) and plasma (boxplot 1304) samples in open chromatin regions. These OCRs were not tissue-specific; instead, they were a general sampling of OCRs across different tissues. Specifically, all known DHS sites were used.

[0090] As shown in boxplot 1302, the median O / E ratio of fetal-specific cell-free DNA (CFR) in transrenal urinary cell-free DNA was 1.84 (range: 1.68–2.13), which was 1.67-fold higher than the median O / E ratio of shared CFR DNA of primarily non-renal origin in urine samples with a fetal fraction greater than 0.44% (median: 1.10, range: 1.08–1.19). In contrast, no clear enrichment of O / E ratios was observed for both fetal-specific (median O / E ratio: 1.048, range: 1.023–1.126) and shared CFR DNA (median: 1.058, range: 1.033–1.124) in plasma samples (boxplot 1304).

[0091] In addition, OCR-associated DNA contribution was elevated in fetal DNA (an example of transrenal DNA molecules) compared with non-transrenal DNA molecules. These data demonstrated that the amount of OCR-associated DNA can be used to estimate the fractional concentration of transrenal cell-free DNA in urine. The higher the O / E ratio, the higher the fractional concentration of clinically relevant DNA, e.g., the tissue or tissues in which OCR was used. When OCRs from different transrenal tissues are used, the fractional concentration will correspond to the average concentration of those tissues (e.g., weighted by the number and size of the corresponding OCRs). The fractional concentration can approximate transrenal DNA concentration, and more OCRs from different tissues can provide greater accuracy for approximating transrenal DNA concentration.

[0092] To estimate fractional concentrations, calibration (training) samples with known fractional concentrations of clinically relevant DNA can be used. The calibration value can correspond to the relative abundance of the calibration sample, and the calibration value and the known fractional concentration comprise calibration data points. If the new sample has a higher relative abundance, the new sample has a higher fractional concentration than the calibration sample. If the new sample has a lower relative abundance, the new sample has a lower fractional concentration than the calibration sample. Multiple calibration samples can be used to determine a range of fractional concentrations. In other implementations, a calibration function (also called a calibration curve) can be determined via a function fit (e.g., linear or nonlinear regression) of the calibration data points.

[0093] 14 illustrates a set of graphs 1400 identifying the enrichment of DNA molecules within open chromatin regions in fetal-specific DNA from a urine sample, according to some embodiments. Graph 1402 identifies the expected and observed percentage of fetal-specific urinary DNA molecules that are from OCR regions. Graph 1404 identifies the expected and observed percentage of shared urinary DNA molecules that are from OCR regions. The expected percentages can be determined based on the size of the OCR regions as a proportion of the genome.

[0094] As shown in graph 1402, fetal-specific urinary DNA is enriched in open chromatin regions, as the observed values ​​are significantly greater than the expected values. Furthermore, graph 1404 shows that the expected and observed values ​​for shared urinary DNA molecules show a smaller decrease. Based on graphs 1402 and 1404, it is shown that urinary DNA is enriched in open chromatin regions. These results also suggest that the kidney's filtering mechanism contributes to the enrichment of transrenal DNA in open chromatin regions. Thus, embodiments can enrich urine samples for clinically relevant DNA, for example, by selecting cfDNA from open chromatin regions specific to one or more transrenal tissues (e.g., fetal, tumor, transplant, or transrenal tissues in general).

[0095] FIG. 15 shows a set of graphs 1500 identifying non-enrichment of DNA molecules within open chromatin regions in fetal-specific DNA from plasma samples, according to some embodiments. Graph 1502 identifies the expected and observed percentage of fetal-specific plasma DNA molecules that are from OCR regions. Graph 1504 identifies the expected and observed percentage of shared plasma DNA molecules that are from OCR regions. As shown in graphs 1502 and 1504, both fetal-specific plasma DNA and shared plasma DNA are not enriched in open chromatin regions. In addition, the O / E ratio shown in graph 1506 also suggests no enrichment of plasma DNA from open chromatin regions. Thus, while the observed enrichment of urinary cell-free DNA in open chromatin regions can be used to determine the fraction of fetal-specific DNA, such a determination would not be feasible for plasma cell-free DNA.

[0096] VI. Identification of clinically relevant DNA in urine Because differential fragmentation patterns between transrenal and non-transrenal urinary cell-free DNA can be identified, we hypothesized that transrenal cell-free DNA contributions could be enriched by selectively analyzing fragmentomic features of transrenal urinary cell-free DNA. Fragmentomic features may include, but are not limited to, terminal motifs (e.g., CC termini), genomic regions (e.g., OCRs), and size (e.g., 80 bp or less). Furthermore, the accuracy of determining transrenal cell-free DNA contributions can be further improved by identifying urinary cell-free DNA molecules that are from open chromatin regions.

[0097] A. Using abundance from OCR to estimate the amount of clinically relevant DNA As previously shown in Figure 13, greater OCR-associated DNA enrichment was observed in fetal-specific urinary cell-free DNA compared to shared cell-free DNA in urine samples, which differs from the non-enrichment of open chromatin regions for plasma DNA molecules.

[0098] Therefore, the fetal DNA fraction (or fraction of other clinically relevant DNA) can be determined in a urine sample using the O / E ratio of all urinary cell-free DNA fragments. To calculate the O / E ratio of all urinary cell-free DNA fragments, the observed OCR-associated DNA contribution can be determined as the percentage of fragments aligned to OCRs among all fragments. The expected OCR-associated DNA contribution can be defined as the theoretical percentage of OCRs in a reference genome (e.g., a human reference genome).

[0099] 1.O / E ratio Figure 16 shows a graph 1600 identifying the correlation between the fetal DNA fraction in maternal urine and the O / E ratio of all urinary cell-free DNA fragments from OCR regions. All OCRs corresponding to DHS sites were used. Here, we define the observed value (O) as the percentage of fragments aligned to OCRs among all fragments, and the expected value (E) as the theoretical percentage of OCRs in the reference genome. The higher the O / E ratio, the more fragments are enriched from the OCR region. As shown in Figure 16, the fractional concentration of fetal DNA in maternal urine increased proportionally to the O / E ratio of all urinary cell-free DNA fragments (Pearson's R = 0.866; P value < 0.001). Therefore, the fractional concentration of fetal DNA can be estimated in a urine sample by determining the O / E ratio of DNA molecules from OCRs.

[0100] Figure 17 shows a graph 1700 illustrating the correlation between the fetal DNA fraction in maternal urine and the O / E ratio of urinary cfDNA fragments from placenta-specific DHS. As with all DHS sites, the higher the O / E ratio, the more fragments are concentrated from the OCR region. As shown in Figure 17, the fractional concentration of fetal DNA in maternal urine increased proportionally to the O / E ratio of all urinary cell-free DNA fragments in placenta-specific DHS (Pearson's R = 0.820; P value < 0.001). Therefore, the fractional concentration of fetal DNA can be estimated in a urine sample by determining the O / E ratio of DNA molecules from tissue-specific OCRs, such as placenta-specific OCRs.

[0101] 2. Use of normalized terminal density and size As another example of using relative abundance to determine the fractional concentration of clinically relevant DNA, the end density of total urinary cell-free DNA located in the OCR can be used to determine the fetal fraction in a urine sample. End density can identify the proportion of DNA molecules that terminate at specific positions (e.g., DNase 1 hypersensitive sites). For example, for all DNase 1 hypersensitive sites, a normalized end density can be calculated at a distance of 0 bp to the central genomic position. A higher normalized end density at the OCR (a distance of 0 bp to the central genomic position of the OCR) can be associated with a higher fraction of transrenal cell-free DNA (e.g., fetal DNA) in urine.

[0102] To determine the end density in OCR regions, we analyzed 14 maternal urine samples and 11 maternal plasma samples. Both the 5' and 3' ends of DNA fragments within 1-kb upstream and 1-kb downstream of the central genomic location of the OCR were analyzed. Normalized end density was defined as the number of fragment ends located within a window (e.g., 1-kb upstream and 1-kb downstream) around the OCR divided by the median or average number across loci / regions adjacent (e.g., adjacent) to one or more of the OCRs used. Other windows, e.g., at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or more than 1,000 bp upstream or downstream, can also be used. For example, adjacent loci can be outside the window used to define the OCR and can be of various lengths, e.g., as described above.

[0103] 18 shows a graph 1800 identifying the normalized end density of total urinary cell-free DNA in the OCR of maternal urine and plasma samples. All OCRs corresponding to DHS sites were used. The maternal urine and maternal plasma samples are shown as red and yellow lines, respectively. As shown in graph 1800, the normalized end density of the OCR was substantially more concentrated in the maternal urine samples than in the maternal plasma samples.

[0104] Figure 19 shows a set of graphs 1900 identifying a comparison between the end density of urinary cell-free DNA and the end density of plasma cell-free DNA to determine the fractional concentration of fetal DNA, according to some embodiments. For urine samples, the end density of DNA molecules from OCR can be used to determine the fractional concentration of fetal DNA. Such a determination cannot be performed for plasma samples. Graph 1902 shows that the fetal DNA fraction increased proportionally to the normalized end density of urinary cell-free DNA, while graph 1904 shows no such increase for plasma cell-free DNA. Thus, the relative abundance in OCR in a urine sample can be used to estimate the fractional concentration of clinically relevant DNA in the urine sample.

[0105] Figure 20 shows a set of graphs 2000 identifying the correlation between fetal DNA fraction and normalized end densities of urinary cell-free DNA of different sizes, according to some embodiments. In Figure 20, graph 2002 shows the normalized end densities of all cell-free DNA fragments in maternal urine samples, and graph 2004 shows the normalized end densities of fragments with a size of 80 bp or less in maternal urine samples. As described herein, other size thresholds (size cutoffs) can be used in addition to 80 bp.

[0106] In Graph 2002, the fetal DNA fraction in maternal urine correlated significantly with the normalized end density in OCR (Pearson's R = 0.926; P value < 0.001). The correlation between the fetal DNA fraction in maternal urine and the normalized end density in OCR could be further improved by selecting fragments 80 bp or shorter (Pearson's R = 0.960; P value < 0.001) (Graph 2004). The results suggested that the use of molecules derived from OCR could inform the extent of transrenal urinary cell-free DNA.

[0107] 3. Method 21 is a flowchart of a method 2100 for estimating the fractional concentration of clinically relevant DNA molecules in a urine sample of a subject, according to some embodiments. The urine sample may contain clinically relevant DNA and other DNA that is acellular. In other examples, the biological sample may not contain clinically relevant DNA, and the estimated fractional concentration may show zero or a low percentage of clinically relevant DNA.

[0108] A urine sample may contain a mixture of cell-free DNA molecules from one or more tissue types, such as heart, lung, and liver. For example, a urine sample may be obtained from a pregnant woman and may contain maternal cell-free DNA molecules and fetal cell-free DNA molecules. A urine sample may contain tumor-specific cell-free DNA molecules as well as other tissue-specific cell-free DNA molecules. Clinically relevant DNA molecules may include fetal DNA. In some embodiments, clinically relevant DNA includes tumor DNA. Aspects of method 2100 and any other method described herein may be implemented by a computer system.

[0109] In some cases, urine samples are treated with a DNA stabilizing agent before obtaining cell-free DNA molecules. Different DNA stabilizing agents can be used, such as EDTA and Collipee stabilizers. EDTA can inhibit the cleavage activity of DNASE1 family members by chelating magnesium and calcium, which are essential ions required for DNASE1 digestion. The stabilizers can potentially stabilize urine DNA from degradation. Stabilizers include preservatives provided by Collipee, diazolidinyl urea (DU), dimethylol urea, 2-bromo-2-nitropropane-1,3-diol, 5-hydroxymethoxymethyl-1-aza-3,7-dioxabicyclo(3.3.0)octane and 5-hydroxymethyl-1-aza-3,7-dioxabicyclo(3.3.0)octane and 5-hydroxypoly[methyleneoxy]methyl-1-aza-3.7-dioxabicyclo(3.3.0)octane, bicyclic oxazolidines (e.g., Nuosept 95), DMDM ​​hydantoin, imidazolinone, methylparaben ... Other metal ion chelators may be zolidinyl urea (IDU), sodium hydroxymethylglycinate, hexamethylenetetramine chloroallyl chloride (quaternium-15), biocides (such as Bioban, Preventol, and Grotan), water-soluble zinc salts, EDTA, N,N'-bis(dithiocarboxy)piperazine (BDP), diethyldithiocarbamate (DDTC), iminodisuccinic acid (IDS), polyaspartic acid, S,S-ethylenediamine-N,N'-disuccinic acid (EDDS), methylglycine diacetic acid (MGDA), etc.

[0110] In block 2102, a plurality of cell-free DNA molecules from the urine sample is analyzed. In some cases, the plurality of cell-free DNA molecules from the urine sample is analyzed to obtain sequence reads. By way of example, the sequence reads may be obtained using sequencing or probe-based techniques, either of which may include, for example, enrichment via amplification or capture probes.

[0111] Sequencing can be performed in a variety of ways, for example, using massively parallel sequencing or next-generation sequencing, using single molecule sequencing, and / or using double-stranded or single-stranded DNA sequencing library preparation protocols. Those skilled in the art will understand the various sequencing techniques that can be used. As part of sequencing, it is possible that some of the sequence reads may correspond to cellular nucleic acids.

[0112] In some cases, analyzing the plurality of cell-free DNA molecules includes (i) determining the locations of the plurality of cell-free DNA molecules and (ii) identifying, based on the locations, a set of cell-free DNA molecules that are from open chromatin regions of one or more tissues associated with the clinically relevant DNA molecules. All OCRs or only a subset of OCRs may be used. For example, OCRs specific to tissues that produce (e.g., contribute to) transrenal DNA may be used. Any one or more of the transrenal-specific OCRs may be used in embodiments of the present disclosure. Such regions may be referred to as transrenal open chromatin regions.

[0113] The location of a cfDNA molecule can be determined by aligning (mapping) one or more corresponding sequence reads to a reference genome. Alternatively, the location can be defined based on the probe used, e.g., a probe identified by an emitted signal, such as the color of a fluorescent dye. In this way, it can be determined whether a cfDNA molecule is within a transrenal OCR.

[0114] OCRs can be identified in various ways, as will be understood by those skilled in the art in light of this disclosure. Open chromatin regions can include one or more DNase 1 hypersensitive sites (DHSs) defined using DNase-seq (Meuleman et al. Nature. 2020;584:244-251). Open chromatin regions can include sites identified using DNase-seq, sites identified using an assay for transposase-accessible chromatin using sequencing (ATAC-seq), transcription start sites (TSSs), CCCTC-binding factor (CTCF) sites, enhancer sites, and other nuclease hypersensitive sites.

[0115] In some embodiments, analyzing the plurality of cell-free DNA molecules further includes identifying a set of cell-free DNA molecules that are from open chromatin regions of one or more tissues and have a size below a specified size threshold. For example, as shown in graph 2004 of FIG. 20, the relative abundance of clinically relevant DNA molecules can be calculated based on shorter DNA fragments (e.g., fragments of less than 80 base pairs) that are from open chromatin regions. Because transrenal DNA (e.g., DNA molecules that pass through the GBM of the kidney) can be characterized by their shorter size, a size threshold can filter out shorter DNA fragments when determining the relative abundance of clinically relevant DNA molecules. As described herein, the amount of transrenal DNA can be used to identify clinically relevant DNA molecules in a urine sample. By way of example, the size threshold may be 40 base pairs, 50 base pairs, 60 base pairs, 70 base pairs, 80 base pairs, 90 base pairs, 100 base pairs, 110 base pairs, 120 base pairs, 130 base pairs, 140 base pairs, 150 base pairs, or 160 base pairs, which may be used in any embodiment using the size thresholds described herein.

[0116] A statistically significant number of cell-free DNA molecules can be analyzed to provide an accurate determination of fractional concentration. In some embodiments, at least 1,000 cell-free DNA molecules are analyzed. In other embodiments, at least 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, or 5,000,000 or more cell-free DNA molecules can be analyzed. In some cases, the set of cell-free DNA molecules includes at least 1,000, 2,000, 3,000, 4,000, 5,000, 10,000, 50,000, or 100,000 cell-free DNA molecules.

[0117] To identify a set of cell-free DNA molecules from open chromatin regions, a urine sample can be enriched for DNA fragments from OCRs (e.g., targeted sequencing), thereby creating an enriched sample. For example, a biological sample can be enriched for DNA fragments from open chromatin regions of one or more tissues, such as CTCF regions, TSS regions, DNase 1 hypersensitive sites, or Pol II regions. Enrichment can involve using capture probes that bind to a portion or the entire genome, as defined, for example, by a reference genome. As another example, enrichment can involve using primers to amplify specific regions of the genome (e.g., via PCR, rolling circle amplification, or multiple displacement amplification (MDA)). In some cases, enrichment involves using a set of cell-free DNA molecules from open chromatin regions of one or more tissues and having a size below a specified size threshold. In some embodiments, urine samples are enriched for cell-free DNA molecules with multiple fragmentological characteristics, including cell-free DNA molecules that: (i) are from open chromatin regions of one or more tissues, (ii) have a size that is less than a specified size threshold (e.g., 80 base pairs), and / or (iii) have one or more end sequences that correspond to a sequence end signature (e.g., CC ends).

[0118] In block 2104, the set of cell-free DNA molecules is used to determine the relative abundance of multiple cell-free DNA molecules from open chromatin regions of one or more tissues. In some cases, the relative abundance may include normalized end density. For example, normalized end density can be calculated based on the number of fragment ends of the set of DNA molecules located within various sized windows around an OCR (e.g., 1-kb upstream and 1-kb downstream, or others described herein) divided by the median or average number across all OCR-flanking loci. OCRs can be defined in various ways, for example, by CTCF sites, TSS sites, DNase 1 hypersensitive sites, or Pol II regions.

[0119] Thus, end density can include a first amount of a set of cell-free DNA molecules from an open chromatin region of one or more tissues divided by a second amount of cell-free DNA molecules from one or more other regions, e.g., regions adjacent to one or more of the OCRs, potentially all of the OCRs used. The second amount can be the amount of all of the cell-free DNA molecules, and thus the first amount can be a subset of the second amount.

[0120] In some embodiments, as previously shown in Figures 16-17, a urine sample from a healthy subject may be expected to exhibit a particular amount of DNA molecules from open chromatin regions, or a particular ratio (e.g., O / E ratio) between a first relative frequency (e.g., percentage) of urinary DNA molecules from OCRs (the observed value) and a second relative frequency (the expected value) of a reference sequence in a reference genome from open chromatin regions of one or more tissues. Thus, the expected OCR-associated DNA contribution can correspond to a theoretical percentage of OCRs in a reference genome (e.g., a human reference genome). For example, the observed value (O) can include the percentage of fragments aligned to OCRs among all fragments, and the expected value (E) as the theoretical percentage of OCRs in the reference genome. In some cases, the expected value can be determined based on the relative frequency of single-base variants in a reference genome that is from open chromatin regions of one or more tissues. In various examples, the relative abundance can be a ratio between the first relative frequency and the second relative frequency, or a ratio between one of the frequencies and the sum of both values.

[0121] In block 2106, the fractional concentrations of clinically relevant DNA molecules in the biological sample are estimated by comparing the relative abundances to one or more calibration values ​​determined from one or more calibration samples in which the fractional concentrations of clinically relevant DNA molecules are known. As shown in Figures 19 and 20, fetal DNA and maternal DNA have different relative abundances. A sample with a mixture of both will have a relative abundance that depends on the proportion of fetal / maternal DNA in the sample. The fractional concentrations of the calibration samples can be determined in other ways, for example, using a locus on the Y chromosome of a male fetus or a fetal-specific marker (e.g., an allele inherited from the father or a fetal-specific epigenetic marker).

[0122] The calibration data points include relative abundances and measured / known fractions of clinically relevant DNA. The comparison can involve comparing the relative abundances to a calibration curve (composed of calibration data points), and thus the comparison can identify points on the curve with the measured relative abundances for the test sample. For example, the relative abundances can be compared to a calibration curve by inputting the relative abundances into a calibration function that represents the calibration curve. The fractional concentrations corresponding to the identified points can then be used to estimate the fractional concentrations. For example, the relative abundances can be provided as inputs to a calibration function (e.g., a linear or nonlinear fit) to obtain an output of fractional concentrations.

[0123] Thus, comparing the relative abundances to one or more calibration values ​​can include comparing the relative abundances to a calibration curve including one or more calibration values. To obtain the calibration data points, some embodiments can measure, for each calibration sample of one or more calibration samples, the fractional concentration of clinically relevant DNA molecules in the calibration sample, and measure the relative abundance of cell-free DNA molecules from the calibration sample that are from open chromatin regions of one or more tissues. As described above, measuring the fractional concentration of clinically relevant DNA molecules can use tissue-specific alleles or tissue-specific methylation patterns.

[0124] A fractional concentration is a quantitative value and can be a range of values. For example, a fractional concentration can specify that the quantitative value is greater than or less than a specified value. In other implementations, a fractional concentration can have upper and lower limits that can correspond to the resolution at which the fractional concentration can be determined.

[0125] B. Enrichment of urine samples using terminal motif features In some embodiments, the transrenal urinary cell-free DNA fraction in a urine sample can be determined using certain terminal motifs, for example, as described in Section IV and elsewhere herein. Terminal motifs can include, but are not limited to, terminal sequences of certain lengths (e.g., 1-mer, 2-mer, 3-mer, 4-mer, 5-mer, 6-mer, 7-mer, 8-mer, 9-mer). Further data is provided below.

[0126] 1. Relationship between the C-terminal motif and the fetal fraction Figure 22 shows a set of graphs 2200 identifying the correlation between fetal DNA fraction and the proportion of urinary cell-free DNA fragments bearing CC ends. In Figure 22, graph 2202 shows the correlation between fetal DNA fraction and the proportion of all urinary cell-free DNA fragments bearing CC ends in a urine sample. Graph 2202 shows the correlation between fetal DNA fraction and the proportion of urinary cell-free DNA fragments with CC ends that are 80 bp or shorter.

[0127] In graph 2202, the fetal DNA fraction in maternal urine correlated significantly with the proportional contribution of urinary cell-free DNA fragments carrying a "CC end" among all fragments (Pearson's R = 0.637; P value = 0.006). In graph 2204, after selecting fragments 80 bp or shorter, we observed a further increase in the correlation between the fetal DNA fraction in maternal urine and the proportion of urinary cell-free DNA fragments carrying a "CC end" (Pearson's R = 0.807; P value < 0.001).

[0128] 2. Exemplary Enrichment Protocol Figure 23 illustrates a technique 2300 that uses probes to enrich for a set of one or more terminal motifs. As described in Figures 9, 10, and 22, technique 2300 can be used to enrich for C-terminal motifs to enrich urine samples for clinically relevant DNA.

[0129] As shown, cell-free DNA molecules 2302 have different terminal motifs, for example, 1-mer terminal motifs in this example.

[0130] In step 2304, cfDNA fragments with different terminal motifs were ligated with a common sequence 2305, e.g., an artificial sequence. Two or more sequences may be used, and using one common sequence may be more efficient. The length of the artificial sequence may be a specified length (4 bp), such as 16 bp (or 17, 18, 19, 20, 21, 22, 23, 24, 25, or 26 bp), to ensure specificity of probe binding. 16 >3×10 9 The length of the DNA fragment should be equal to or greater than the length of the human genome. Artificial sequences at the ends of the DNA fragment can facilitate probe recognition of specific DNA end motifs.

[0131] In step 2306, the DNA molecule having the consensus sequence 2305 is denatured to separate the two strands, resulting in single-stranded cfDNA 2308 having different terminal motifs and the consensus sequence 2305. As will be appreciated by one of skill in the art, a variety of denaturation protocols can be used, including, for example, the use of temperature.

[0132] In step 2310, a surface 2312 (e.g., a chip surface) is immobilized with many probe sequences 2316. The probe sequences 2316 have two components: a complementary sequence to the consensus sequence 2305 and a complementary motif sequence 2318 to a target terminal motif sequence (e.g., a "G" targeting a C-terminal motif). Only fragments with the targeted terminal motif (e.g., a terminal C motif) can bind to the probe; fragments terminating in other motifs remain unbound (i.e., unbound fragments 2320). The complementary motif sequence 2318 can be a set of different terminal motifs, for example, if more than two-mers are used. For example, in the case of two-mers, four different probes can be used for four different two-mers terminating in C.

[0133] In step 2314, unbound fragments 2320 are washed away. The remaining bound fragments 2322 can be detected or further analyzed in various ways. For example, only probes with the complementary motif sequence 2318 bound to the fragment can be extended (e.g., by a single nucleotide ligated with a fluorescent dye by a DNA polymerase). In this way, a fluorescent signal can be detected when a fragment bearing the targeted motif is present. As another example for detection, a reaction can extend the bound cfDNA fragment with a single nucleotide labeled with biotin. The biotin can be detected with streptavidin conjugated to a fluorophore. Alternatively, the reaction can extend a single nucleotide labeled with dinitrophenyl. The dinitrophenyl can be detected with an anti-DNP antibody labeled with a fluorophore. In other implementations, the bound fragments can be sequenced in a separate process.

[0134] Figure 24 illustrates another technique 2400 that uses probes and beads to enrich for a set of one or more terminal motifs. Similar to Figure 23, cfDNA fragments with different terminal motifs can be ligated with artificial sequences. The artificial sequences at the ends of the DNA fragments can facilitate probe recognition of specific DNA terminal motifs. Double-stranded DNA fragments can be denatured to single-stranded DNA.

[0135] In the example shown, a probe targeting a DNA fragment with a specific terminal motif has three components: biotin that can be bound to streptavidin beads, a sequence complementary to the common sequence, and a sequence complementary to the specific terminal motif sequence (e.g., "G" targeting the C-terminal motif). The probe is hybridized to the DNA fragment. Only fragments with the specific terminal motif can bind to the probe; fragments with other terminal motifs remain unbound.

[0136] Streptavidin beads can capture probes due to the high affinity between biotin and streptavidin. Only fragments with specific terminal motifs can be captured by streptavidin beads. Undiscovered fragments are washed away. As a result, cfDNA fragments with specific terminal motifs can be captured by such a design.

[0137] Fragments bound to complementary motif sequences using technique 2400 can be detected or further analyzed in the same manner as technique 2300.

[0138] Instead of washing away unbound fragments to enrich for the target end motif, the target end motif can be amplified. For example, a primer containing a consensus sequence and the target end motif can be added to the reaction along with nucleotides, and an amplification process (e.g., PCR or rolling circle) can be performed.

[0139] 3. Method 25 shows a flowchart of a method 2500 for enriching a urine sample for clinically relevant DNA based on terminal motif characteristics of urinary cell-free DNA, according to some embodiments. Method 2500 and aspects of other methods described herein can be performed in a similar manner to method 2100, e.g., sample preparation and analysis of DNA molecules. A urine sample can contain clinically relevant DNA and other DNA that is cell-free. In other examples, a biological sample may not contain clinically relevant DNA, and the estimated fractional concentration may show zero or a low percentage of clinically relevant DNA. Method 2500 and aspects of any other methods described herein can be performed by a computer system.

[0140] In block 2502, a plurality of cell-free DNA molecules from a urine sample are analyzed. This embodiment of block 2502 may be performed in a manner similar to block 2102 of method 2100. For example, a plurality of cell-free DNA molecules from a urine sample can be analyzed to obtain sequence reads. The sequence reads may include terminal sequences corresponding to the ends of a plurality of cell-free DNA fragments. By way of example, the sequence reads may be obtained using sequencing or probe-based techniques, either of which may include enrichment via, for example, amplification or capture probes.

[0141] Sequencing can be performed in a variety of ways, for example, using massively parallel sequencing or next-generation sequencing, using single molecule sequencing, and / or using double-stranded or single-stranded DNA sequencing library preparation protocols. Those skilled in the art will understand the various sequencing techniques that can be used. As part of sequencing, it is possible that some of the sequence reads may correspond to cellular nucleic acids.

[0142] In some embodiments, analyzing the plurality of cell-free DNA molecules further includes identifying a set of cell-free DNA molecules that fall within a set of one or more sequence motifs that include a C-terminal nucleotide. The sequence terminal signature can be part of a Kmer terminal motif, such as a 2-mer, 3-mer, 4-mer, etc. For example, a set of cell-free DNA molecules is further identified based on having a CC terminus. Furthermore, terminal sequences may be required at both ends of the DNA fragment, or specific pairs of different terminal motifs can be used to select specific sets of DNA fragments.

[0143] When sequencing is performed, identifying a set of multiple cell-free DNA molecules can include identifying sequence reads that have end sequences within one or more sequence motif sets.Therefore, enriched samples can correspond to sequence reads that have end sequences within one or more sequence motif sets.As an alternative to sequencing, to identify DNA molecules that have one or more end sequences, one or more probe molecules can be attached to a surface and detect the sequence motif in the end sequence by hybridization.

[0144] In some embodiments, the set of cell-free DNA molecules is further identified based on their respective sizes (e.g., fragments less than 80 base pairs). As shown in Figure 22, a size threshold can filter out shorter DNA fragments, since transrenal DNA (e.g., DNA molecules that pass through the GBM of the kidney) can be characterized by their shorter size. As described herein, the amount of transrenal DNA can be used to identify clinically relevant DNA molecules in a urine sample. Specified size thresholds can include 40 base pairs, 50 base pairs, 60 base pairs, 70 base pairs, 80 base pairs, 90 base pairs, 100 base pairs, 110 base pairs, 120 base pairs, 130 base pairs, 140 base pairs, 150 base pairs, or 160 base pairs.

[0145] A statistically significant number of cell-free DNA molecules can be analyzed as described herein.

[0146] In block 2504, an enriched sample can be created by using a set of cell-free DNA molecules that are in one or more sets of sequence motifs. Thus, the enriched sample contains a higher concentration of clinically relevant DNA compared to a urine sample. The enriched sample can be an in silico sample, in that only certain cfDNA molecules are measured. In another example, the enriched sample can be a physical sample.

[0147] Enrichment can involve using capture probes that bind to a set of one or more sequence motifs. For example, identifying a set of cell-free DNA molecules or creating an enriched sample can involve subjecting a plurality of cell-free DNA molecules to one or more probe molecules that detect a set of one or more sequence motifs in the terminal sequences of the plurality of cell-free DNA molecules. Through the use of such probe molecules, a set of cell-free DNA molecules can be obtained. As described with respect to Figures 23 and 24, some embodiments can attach a common sequence to a plurality of cell-free DNA molecules. One or more probe molecules can then contain a sequence complementary to the common sequence.

[0148] In some cases, creating an enriched sample includes capturing a set of cell-free DNA molecules using probe molecules and discarding other cell-free DNA molecules of the plurality of cell-free DNA molecules. In other examples, creating an enriched sample can include amplifying the set of cell-free DNA molecules using one or more probe molecules.

[0149] Capture probes can also bind to (target) part or all of a genome, e.g., as defined by a reference genome. As another example, enrichment can use primers to amplify specific regions of the genome (e.g., via PCR, rolling circle amplification, or multiple displacement amplification (MDA)).

[0150] In some embodiments, urine samples are enriched for cell-free DNA molecules with multiple fragmentological characteristics, including cell-free DNA molecules that: (i) are from open chromatin regions of one or more tissues; (ii) have a size below a specified size threshold (e.g., 80 base pairs); and / or (iii) have one or more end sequences corresponding to a sequence end signature (e.g., the C-terminus or CC-terminus of a Kmer end motif).

[0151] At block 2506, a characteristic associated with the clinically relevant DNA in the concentrated urine sample is determined. For example, the characteristic of the clinically relevant DNA in the urine sample can be (1) the fractional concentration of the clinically relevant DNA or (2) the level of pathology of the subject from which the biological sample was obtained, e.g., the level of pathology associated with the clinically relevant DNA molecule. Those skilled in the art will appreciate the variety of properties that can be determined using a set of cell-free DNA molecules having terminal sequences within a set of one or more sequence motifs that include a C-terminal nucleotide, as variously described in U.S. Publication Nos. 2009 / 0087847, 2009 / 0029377, 2011 / 0276277, 2011 / 0105353, 2013 / 0040824, 2014 / 0100121, 2014 / 0080715, and 2020 / 0199656, such as, for example, fetal inheritance of haplotypes, detection of mutations, copy number abnormalities (e.g., aneuploidy), methylation profiles, various base modifications, genomic interactions, protein binding status, fragmentomic features, etc.

[0152] C. Concentration of Urine Samples Using OCR and Other Features The above fragmentomic features (e.g., terminal motifs, size, and enrichment of open chromatin regions) can be combined to estimate the contribution of transrenal DNA. For example, fetal DNA molecules can be enriched in CC termini. Based on this correlation, the contribution of transrenal DNA can be estimated based on the proportion of urinary cell-free DNA with CC termini in a urine sample. When the proportion of urinary cell-free DNA with CC termini and size (e.g., fragments less than 80 bp) are used together, estimating the transrenal DNA contribution in a urine sample can be more accurate. In effect, the accuracy of estimating the fetal DNA fraction can also be improved.

[0153] Figure 26 shows a set of graphs identifying the enrichment of fetal DNA using urinary cell-free DNA with various fragmentomic features, according to some embodiments. As shown in Figure 26, DNA molecules filtered for CC ends (graph 2610), OCR (graph 2620), and sizes less than 80 base pairs (graph 2630) can result in a significant increase in fetal DNA in urine samples. The enrichment is even more pronounced using the combination (graph 2640) compared to using the above fragmentomic features individually.

[0154] Figure 27 shows a bar graph 2700 identifying enrichment of transrenal urinary cell-free DNA using selective analysis of fragments with different fragmentomic features. The percentage increase in fetal DNA fraction in urinary cell-free DNA with no selection 2702, selective analysis based on CC termini 2704, selective analysis based on OCR 2706, size-based selection (80 bp or less) 2708, and a combination of features corresponding to terminal motifs, OCR, and size 2710. To show the increase in fetal DNA fraction, we calculated the average increase in fractional concentration of fetal DNA after different criteria selection.

[0155] As shown in Figure 27, when fragments were filtered for having a CC end, being within an OCR region, or having a size of 80 bp or less, the fetal DNA fraction in a given urine sample increased by 78.6%, 60.1%, and 223.8%, respectively. In other words, filtering DNA molecules using CC end and OCR region criteria increased the fractional concentration of fetal DNA by approximately 1-fold. When size (fragments 80 bp or less) was used as a filter for DNA molecules, the fractional concentration of fetal DNA increased by approximately 2-fold. Combining these three fragmentomic features together can further increase the fetal DNA fraction in urine by 836.8% (more than 8-fold). Thus, the data in Figure 27 demonstrate that targeted transrenal urinary cell-free DNA can be enriched based on selective analysis of cell-free DNA molecules according to different combinations of fragmentomic features.

[0156] Additionally, a combination of two or more of these fragmentomic features can be used to estimate the contribution of fetal DNA in a urine sample. For example, method 2100 can further use size distribution statistics, as described in U.S. Patent No. 9,892,230. As another example, in addition to or instead of using OCR, a set of one or more sequence motifs containing a C-terminal nucleotide can be used. Each of these different features can be used together, for example, in a two- or three-dimensional calibration curve.

[0157] Figure 28 shows a flowchart of a method 2800 for enriching a urine sample for clinically relevant DNA based on terminal motifs and open chromatin regions, according to some embodiments. Method 2800 and aspects of the other methods described herein can be performed in a similar manner to method 2100 and / or method 2500, e.g., sample preparation and analysis of DNA molecules. A urine sample can contain clinically relevant DNA and other DNA that is cell-free. In other examples, a biological sample may not contain clinically relevant DNA, and the estimated fractional concentration may show zero or a low percentage of clinically relevant DNA. Method 2800 and aspects of any other methods described herein can be performed by a computer system.

[0158] In block 2802, a plurality of cell-free DNA molecules from a urine sample are analyzed. This embodiment of block 2802 may be performed in a manner similar to block 2102 or block 2502. For example, a plurality of cell-free DNA molecules from a urine sample are analyzed to obtain sequence reads. The sequence reads may include terminal sequences corresponding to the ends of the plurality of cell-free DNA fragments. By way of example, the sequence reads may be obtained using sequencing or probe-based techniques, either of which may include enrichment via, for example, amplification or capture probes, as described for method 2500.

[0159] In some embodiments, analyzing the plurality of cell-free DNA molecules further includes identifying a set of cell-free DNA molecules that: (i) are from open chromatin regions of one or more tissues, (ii) have a size below a specified size threshold (e.g., 80 base pairs), and / or (iii) have one or more end sequences corresponding to a sequence end signature (e.g., the C-terminus or CC-terminus of a Kmer end motif). Open chromatin regions can be identified in a similar manner as described herein.

[0160] As shown in Figures 26 and 27, a set of cell-free DNA molecules can be identified based on their respective sizes (e.g., fragments less than 80 base pairs). The size threshold can filter out shorter DNA fragments, since transrenal DNA (e.g., DNA molecules that pass through the GBM of the kidney) can be characterized by their shorter size. As described herein, the amount of transrenal DNA can be used to identify clinically relevant DNA molecules in a urine sample. Specified size thresholds can include 40 base pairs, 50 base pairs, 60 base pairs, 70 base pairs, 80 base pairs, 90 base pairs, 100 base pairs, 110 base pairs, 120 base pairs, 130 base pairs, 140 base pairs, 150 base pairs, or 160 base pairs.

[0161] In some embodiments, the size of multiple cell-free DNA molecules can be determined using various methods (e.g., gel electrophoresis). For example, the size of multiple cell-free DNA molecules can be measured using gel electrophoresis, filtration, size-selective precipitation, or hybridization. Additionally or alternatively, the size of multiple cell-free DNA fragments can be measured using sequence reads. For example, sequence reads can be obtained from sequencing multiple cell-free DNA molecules from a biological sample (e.g., massively parallel sequencing, single-molecule real-time sequencing, nanopore sequencing). To measure the size of a cell-free DNA molecule, several nucleotides can be counted for each sequence read. Those skilled in the art will understand various sequencing techniques that can be used. As part of sequencing, it is possible that some of the sequence reads may correspond to cellular nucleic acids. In some cases, a urine sample can be enriched for DNA fragments having a size below a predetermined size threshold (e.g., 80 bp).

[0162] A set of cell-free DNA molecules can be further identified based on having one or more terminal sequences corresponding to the sequence terminal signature. The sequence terminal signature can be part of a terminal motif, such as a 2-mer, a 3-mer, etc. For example, a set of cell-free DNA molecules can be further identified based on having a CC terminus. Furthermore, the terminal sequences can be required to be at both ends of the DNA fragment, or a specific pair of different terminal motifs can be used to select a specific set of DNA fragments. In addition to sequencing, to identify DNA molecules with one or more terminal sequences, one or more probe molecules can be attached to a surface or bead, and the sequence motif in the terminal sequence can be detected by hybridization.

[0163] A statistically significant number of cell-free DNA molecules can be analyzed as described herein.

[0164] In block 2804, an enriched sample can be created by using a set of cell-free DNA molecules that: (i) are from open chromatin regions of one or more tissues; (ii) have a size that is less than a specified size threshold (e.g., 80 base pairs); and / or (iii) have one or more end sequences that correspond to a sequence end signature (e.g., CC end). This embodiment of block 2804 can be performed in a manner similar to block 2504. Thus, the enriched sample contains a higher concentration of clinically relevant DNA compared to a urine sample. For example, a biological sample can be enriched for DNA fragments from open chromatin regions of one or more tissues, such as CTCF regions, TSS regions, DNase 1 hypersensitive sites, or Pol II regions. Enrichment can include using capture probes that bind to a portion or the entire genome, e.g., as defined by a reference genome. As another example, enrichment can use primers to amplify specific regions of the genome (e.g., via PCR, rolling circle amplification, or multiple displacement amplification (MDA)).

[0165] In some cases, creating an enriched sample includes capturing a set of cell-free DNA molecules using probe molecules and discarding other cell-free DNA molecules of the plurality of cell-free DNA molecules. The enriched sample can be an in silico sample.

[0166] In block 2806, a characteristic associated with the clinically relevant DNA in the concentrated urine sample is determined. This embodiment of block 2806 may be performed in a manner similar to block 2506. For example, the characteristic of the clinically relevant DNA in the urine sample may be (1) the fractional concentration of clinically relevant DNA, or (2) the level of pathology of the subject from which the biological sample was obtained, e.g., the level of pathology associated with the clinically relevant DNA molecules. Those skilled in the art will appreciate the variety of characteristics that may be determined, e.g., as described above for method 2500.

[0167] VII. Classification of Abnormalities Using Urinary cfDNA In some embodiments, cancer can be detected and monitored using transrenal urinary cell-free DNA molecules. For example, renal cell carcinoma (RCC) is a disease in which malignant cells are found in the lining of small tubules within the kidney. If renal function is affected, the fractional concentration of transrenal urinary cell-free DNA may change. In fact, patients with kidney cancer will exhibit abnormalities in the fractional concentration of transrenal urinary cell-free DNA when compared with subjects without kidney cancer. Other kidney abnormalities (besides RCC) may also affect transrenal urinary cell-free DNA in urine samples, for example, based on size or area, such as OCR. Other examples include proteinuria and preeclampsia.

[0168] A. OCR-based classification A urine sample from a healthy subject may be expected to exhibit a specific amount of DNA molecules from open chromatin regions of one or more tissues, or a specific ratio between the observed frequency of urinary DNA molecules from open chromatin regions and the expected frequency of reference sequences from the reference genome that are from open chromatin regions of one or more tissues. However, if GBM permeability is disrupted in some subjects (e.g., subjects with nephrotic syndrome or glomerulonephritis), the above ratio may increase or decrease. If such a change from the normal amount exceeds a predetermined threshold, the subject may be determined to have a disease or other abnormal condition that affects renal permeability.

[0169] The amount of DNA molecules from open chromatin regions of one or more tissues can be measured for a control subject. A significant deviation from the measured amount of DNA molecules can then be used to determine whether a given subject has a renal abnormality. For example, a blood sample can contain cell-free DNA molecules from different organs (e.g., heart, lung, liver). The amount of cell-free DNA molecules corresponding to open chromatin regions of the liver (for example) can be determined for a urine sample (e.g., using targeted sequencing of OCR). If there is a statistically significant difference between the determined amount of cell-free DNA molecules and a calibrated amount of cell-free DNA molecules corresponding to open chromatin regions of a healthy subject, a classification of renal abnormality can be determined. Two or more tissue-specific regions can be used. Collectively, measurements can be made for all OCRs, or for OCRs specific to one or more tissues that contribute to renal DNA, or for OCRs specific to one or more cell types, as described in the previous section.

[0170] 1. Renal cell carcinoma To illustrate, we analyzed the O / E ratio of urinary cell-free DNA from 15 control subjects and 16 patients with renal cell carcinoma (RCC). To calculate the O / E ratio of all urinary cell-free DNA fragments, the observed OCR-associated DNA contribution can be determined as the percentage of fragments aligned to OCRs among all fragments. The expected OCR-associated DNA contribution can be defined as the theoretical percentage of OCRs in the human genome.

[0171] Figure 29 shows a set of graphs 2900 identifying O / E ratio analysis in patients with RCC. Boxplot 2902 shows a boxplot of the O / E ratio between control subjects and patients with RCC. As shown in boxplot 2902, the O / E ratio of RCC patients (median: 1.378, range: 1.174-1.863) was significantly higher than that of control subjects (median: 1.288, range: 1.190-1.571) (Mann-Whitney, P value <0.001). In addition, we further performed receiver operating characteristic (ROC) analysis on these samples. ROC 2904 indicates the performance level of the ROC for distinguishing patients with RCC from control subjects. The area under the curve (AUC) of ROC 2904 was 0.964 in distinguishing RCC patients from control subjects.

[0172] 2. Proteinuria Proteinuria, also known as albuminuria, is the elevated level of protein in the urine and can be considered a kidney disorder. Because the kidneys are not functioning properly, it allows more protein to enter the urine, and therefore, it is a type of kidney disorder.

[0173] Because patients with proteinuria have excess protein in their urine, we hypothesized that fragmentomic features of urinary cfDNA could be used to identify proteinuric patients from healthy controls. We used abundance for OCR. Any transrenal-specific OCR could be used. As with other embodiments of the present disclosure, OCR alone would be used for any given use case (e.g., renal abnormality classification, fraction concentration estimation, or enrichment), but the two separate determinations could be performed and then combined.

[0174] Due to the absence of placental DNA in the urine of healthy controls and subjects with proteinuria, we use blood-associated regions to represent transrenal-associated genomic locations / regions.

[0175] Figure 30 shows fragmentomic analysis of transrenal DNA in patients with proteinuria. Box plot 3002 shows the O / E ratio of urinary cfDNA fragments in OCR (blood-specific DHS) in healthy controls and patients with proteinuria.

[0176] In the O / E ratio analysis, patients with proteinuria had significantly lower O / E ratios for fragments in the OCR (blood-specific DHS) (Mann-Whitney U test, P value = 0.0052). These results demonstrated a decreased proportion of fragments from the OCR in patients with proteinuria. ROC analysis is provided below.

[0177] 3. Preeclampsia We hypothesized that fragmentomic signatures of urinary cfDNA could be used to identify pregnant women with preeclampsia. Pregnant women with preeclampsia are usually diagnosed with elevated urinary protein levels, indicating impaired GBM function in the kidney. We speculated that if large plasma molecules such as proteins can cross the GBM and enter the urine, large DNA molecules from plasma (e.g., long DNA molecules or DNA molecules bound to histones) could also enter the urine.

[0178] We used DHS to represent OCR. Other methods of identifying OCR are described elsewhere in this disclosure.

[0179] Figure 31 shows fragmentomic analysis of transrenal DNA in pregnant women with preeclampsia. Boxplot 3110 shows the O / E ratio of urinary cfDNA fragments in OCR (all DHS) in healthy pregnant women and women with preeclampsia. Boxplot 3120 shows the O / E ratio of urinary cfDNA fragments in OCR (placenta-specific DHS) in healthy pregnant women and women with preeclampsia.

[0180] Using the O / E ratio for fragments, the performance of the placenta-specific DHS (boxplot 3120) in distinguishing between healthy pregnant women and women with preeclampsia was better than that of the total DHS (boxplot 3110) (Mann-Whitney U test, P value: 0.0011 vs. 0.0118). Therefore, we used tissue-specific regions for O / E ratio analysis in the urine of subjects with preeclampsia and proteinuria. Compared with healthy pregnant women, pregnant women with preeclampsia had significantly lower O / E ratios for fragments in the OCR (placenta-specific DHS) (Mann-Whitney U test, P value = 0.0011). These data indicated a decreased proportion of fragments from the OCR in patients with preeclampsia.

[0181] When renal abnormality is preeclampsia, additional factors can be used.For example, determining whether hypertension exists can also be used.For example, blood pressure can be compared with threshold to determine whether subject has hypertension.Another factor can be whether protein exists in urine, for example, whether proteinuria.

[0182] 4. Method FIG. 32 shows a flowchart of a method 3200 for determining a classification of a kidney abnormality based on urinary cell-free DNA from open chromatin regions, according to some embodiments. Method 3200 and aspects of the other methods described herein can be performed in a manner similar to the methods described above, e.g., sample preparation and analysis of DNA molecules. A urine sample can contain a mixture of cell-free DNA molecules from one or more tissue types, such as heart, lung, and liver. A urine sample can contain tumor-specific cell-free DNA molecules as well as other tissue-specific cell-free DNA molecules. For example, a urine sample can contain cell-free DNA molecules specific to renal cell carcinoma (RCC). Method 3200 and aspects of any other methods described herein can be performed by a computer system.

[0183] In block 3202, a plurality of cell-free DNA molecules from a urine sample are analyzed. Aspects of block 3202, as can be implemented in other methods herein, may be implemented in a manner similar to similar blocks of other methods, such as block 2102 of method 2100. For example, a plurality of cell-free DNA molecules from a urine sample are analyzed to obtain sequence reads. By way of example, the sequence reads may be obtained using sequencing or probe-based techniques, either of which may include, for example, enrichment via amplification or capture probes.

[0184] In some cases, analyzing the plurality of cell-free DNA molecules includes (i) determining the locations of the plurality of cell-free DNA molecules and (ii) identifying, based on the locations, a set of cell-free DNA molecules that are from open chromatin regions of one or more tissues associated with the clinically relevant DNA molecules. The one or more tissues can include at least one of the heart, lung, or liver. The open chromatin regions can include one or more DNase 1 hypersensitive sites (DHSs) defined using DNase-seq (Meuleman et al. Nature. 2020;584:244-251). The open chromatin regions can be identified as described herein.

[0185] A statistically significant number of cell-free DNA molecules can be analyzed as described herein.

[0186] To identify a set of cell-free DNA molecules from open chromatin regions, a urine sample can be enriched for DNA fragments from open chromatin regions (e.g., by targeted sequencing), thereby creating an enriched sample. For example, a biological sample can be enriched for DNA fragments from open chromatin regions of one or more tissues, such as CTCF regions, TSS regions, DNase 1 hypersensitive sites, or Pol II regions. Enrichment can involve using capture probes that bind to a portion or the entire genome, as defined, for example, by a reference genome. As another example, enrichment can involve using primers to amplify specific regions of the genome (e.g., via PCR, rolling circle amplification, or multiple displacement amplification (MDA)). In some cases, enrichment involves using a set of cell-free DNA molecules from open chromatin regions of one or more tissues and having a size below a specified size threshold.

[0187] In block 3204, the relative abundance of a plurality of cell-free DNA molecules from open chromatin regions of one or more tissues is determined. Embodiments of block 3204, as can be performed in other methods herein, may be performed in a similar manner to similar blocks of other methods, such as block 2104 of method 2100. In some cases, the relative abundance may include normalized end density. For example, normalized end density may be calculated based on the number of fragment ends of a set of DNA molecules located within 1 kb upstream and 1 kb downstream of an OCR (e.g., a CTCF site, a TSS site, a DNase 1 hypersensitive site, a Pol II region) divided by the median number across all loci flanking the OCR.

[0188] In some embodiments, as previously shown in Figures 29-31, a urine sample from a healthy subject may be expected to exhibit a particular amount of DNA molecules from open chromatin regions, or a particular ratio (e.g., O / E ratio) between a first relative frequency (e.g., percentage) of urinary DNA molecules from OCRs (observed value) and a second relative frequency (expected value) of reference sequences in a reference genome from open chromatin regions of one or more tissues. Thus, the expected OCR-associated DNA contribution can correspond to the theoretical percentage of OCRs in a reference genome (e.g., a human reference genome). For example, the observed value (O) can include the percentage of fragments aligned to OCRs among all fragments, and the expected value (E) as the theoretical percentage of OCRs in the reference genome. In some cases, the expected value can be determined based on the relative frequency of single-base variants in the reference genome that are from open chromatin regions of one or more tissues. However, if GBM permeability is disrupted in some subjects, the above ratio may increase or decrease. If such a change relative to normal exceeds a predetermined threshold, the subject can be determined to have a disease or other abnormal condition (eg, renal syndrome, glomerulonephritis) that affects renal permeability.

[0189] In block 3206, the relative abundance is compared to a reference value. The reference value can correspond to another relative abundance determined based on cell-free DNA molecules from open chromatin regions of one or more reference samples, where the one or more reference samples are associated with known classifications of kidney abnormalities. For example, the reference value can correspond to a relative abundance determined from a healthy subject. In some cases, the reference value is a calibration value or is determined from a calibration value of a calibration sample. As with other reference values, the particular value selected can depend on a trade-off between specificity and sensitivity. In some embodiments, a machine learning model can be used to perform the comparison.

[0190] In block 3208, a classification of the subject having a kidney abnormality is determined based on the comparison. In some embodiments, comparing the relative abundance to the reference value includes: (1) determining whether the relative abundance differs from the reference value by at least a threshold amount, or whether the difference is less than a threshold amount; (2) determining whether the relative abundance is less than the reference value by at least a threshold amount; or (3) determining whether the relative abundance is greater than the reference value by at least a threshold amount. By way of example, kidney abnormalities can include renal cell carcinoma (RCC), kidney syndrome, glomerulonephritis, Fabry disease, cystinosis, IgA nephropathy, IgM nephropathy, lupus nephritis, atypical hemolytic uremic syndrome (aHUS), polycystic kidney disease (PKD), Alport syndrome, interstitial nephritis, proteinuria, chronic kidney disease, acute kidney injury, proteinuria, preeclampsia, etc. In some cases, the classification of the subject having a kidney abnormality includes an increased level of permeability associated with the glomerular basement membrane of the kidney.

[0191] The classification of the kidney abnormality can be determined using machine learning trained using a training dataset. The training dataset can include training samples. The training samples can be related to known classifications of kidney abnormalities. As another example, the comparison can be performed via a machine learning model. The machine learning model can be applied to the relative abundances to generate a classification of the kidney abnormality. In some embodiments, the machine learning model may include, but is not limited to, a convolutional neural network (CNN), linear regression, logistic regression, a deep recurrent neural network (e.g., a fully connected recurrent neural network (RNN), a gated recurrent unit (GRU), a long short-term memory (LSTM)), a transformer-based method (e.g., XLNet, BERT, XLM, RoBERTa), a Bayesian classifier, a hidden Markov model (HMM), a linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering for applications with noise (DBSCAN), a random forest algorithm, adaptive boosting (AdaBoost), eXtreme Gradient Boosting (XGBoost), a support vector machine (SVM), or a hybrid model including one or more of the models proposed above.

[0192] B. Size-based classification In addition to or instead of classification using OCR, classification can be performed using the size of cfDNA in urine. Classification of kidney abnormalities can be performed in a similar manner, but instead, the size distribution statistics of cfDNA size in urine samples can be used.

[0193] 1. Proteinuria and preeclampsia FIG. 33 is a fragmentomic analysis 3300 of transrenal DNA in a patient with proteinuria and another patient with preeclampsia.

[0194] Boxplot 3310 shows the proportion of urinary cfDNA >80 bp in healthy controls and patients with proteinuria. We observed a higher proportion of long urinary cfDNA fragments (i.e., >80 bp) in patients with proteinuria than in healthy controls (Mann-Whitney U test, P value = 0.0256).

[0195] Boxplot 3320 shows the proportion of urinary cfDNA >80 bp in healthy pregnant women and women with preeclampsia. We observed a higher proportion of long urinary cfDNA fragments (i.e., >80 bp) in pregnant women with preeclampsia than in healthy pregnant women (Mann-Whitney U test, P value = 0.0021).

[0196] 2. Method FIG. 34 shows a flowchart of a method 3400 for determining a classification of a kidney abnormality based on the size of urinary cell-free DNA, according to some embodiments. Method 3400 and aspects of the other methods described herein can be performed in a manner similar to the methods described above, e.g., sample preparation and analysis of DNA molecules. A urine sample can contain a mixture of cell-free DNA molecules from one or more tissue types, such as heart, lung, and liver. A urine sample can contain tumor-specific cell-free DNA molecules as well as other tissue-specific cell-free DNA molecules. For example, a urine sample can contain cell-free DNA molecules specific to renal cell carcinoma (RCC). Method 3400 and aspects of any other methods described herein can be performed by a computer system.

[0197] In block 3402, a plurality of cell-free DNA molecules from a urine sample are analyzed. Aspects of block 3402, as can be implemented in other methods herein, may be implemented in a manner similar to similar blocks of other methods, such as block 2102 of method 2100. For example, a plurality of cell-free DNA molecules from a urine sample are analyzed to obtain sequence reads. By way of example, the sequence reads may be obtained using sequencing or probe-based techniques, either of which may include, for example, enrichment via amplification or capture probes.

[0198] In some embodiments, analyzing the plurality of cell-free DNA molecules includes determining the size of the plurality of cell-free DNA molecules. Various methods (e.g., gel electrophoresis) can be used to determine the size of the plurality of cell-free DNA molecules. For example, the size of the plurality of cell-free DNA molecules can be measured using gel electrophoresis, filtration, size-selective precipitation, or hybridization. Additionally or alternatively, the size of the plurality of cell-free DNA fragments can be measured using sequence reads. For example, sequence reads can be obtained from sequencing a plurality of cell-free DNA molecules from a biological sample (e.g., massively parallel sequencing, single-molecule real-time sequencing, nanopore sequencing). Then, to measure the size of the cell-free DNA molecule, several nucleotides can be counted for each sequence read. Those skilled in the art will understand various sequencing techniques that can be used. As part of sequencing, it is possible that some of the sequence reads may correspond to cellular nucleic acids. In some cases, a urine sample can be enriched for DNA fragments having a size below a predetermined size threshold (e.g., 80 bp).

[0199] A statistically significant number of cell-free DNA molecules can be analyzed as described herein.

[0200] In block 3404, a statistical value is determined for the set of cell-free DNA molecules. The statistical value can be determined based on the sizes of the plurality of cell-free DNA molecules. The sizes can form a size distribution. Various statistical values ​​can be used, for example, the average, mean, median, or mode of the size distribution. As another example, the ratio of cfDNA in a first size range to a second size range can be used, and the size ranges can be different but overlapping. The second size range can be all sizes, i.e., all cfDNA molecules.

[0201] In one example, the relative amount of transrenal DNA in a urine sample (an example of a statistical value) can be characterized by DNA fragments with a size of less than 80 base pairs. If GBM permeability is impaired in some subjects, the relative amount of shorter DNA fragments in the urine sample may increase or decrease. If such a change from the normal amount exceeds a threshold (reference value), the subject may be determined to have a disease or other abnormal condition (e.g., renal syndrome, glomerulonephritis) that affects renal permeability.

[0202] For example, the statistical value can be the size ratio of a first quantity of cell-free DNA molecules having a size less than a size threshold (e.g., 80 bp) to a second quantity corresponding to a plurality of cell-free DNA molecules. By way of example, the size threshold can be 40 base pairs, 50 base pairs, 60 base pairs, 70 base pairs, 80 base pairs, 90 base pairs, 100 base pairs, 110 base pairs, 120 base pairs, 130 base pairs, 140 base pairs, 150 base pairs, or 160 base pairs, which can be used in any embodiment using the size thresholds described herein.

[0203] In some cases, determining the statistical value includes a proportion of the set of cell-free DNA molecules having a size within a size range relative to the plurality of cell-free DNA molecules from the urine sample. The size range can have a lower limit and an upper limit selected from, for example, any of 0, 5, 10, 15, 20, 30, 35, 40, 45, 50, 55, or 60 bases for the lower limit, and 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, or 160 bases.

[0204] In block 3406, the statistical value is compared to a reference value. The reference value can correspond to another statistical value determined based on the measured sizes of cell-free DNA molecules of one or more reference samples, where the one or more reference samples are associated with known classifications of kidney abnormalities. For example, the reference value can be determined based on the sizes of cell-free DNA molecules in healthy urine samples. In some cases, the reference value is a calibration value or is determined from a calibration value of a calibration (training) sample.

[0205] In block 3408, a classification of the subject having a renal abnormality is determined based on the comparison. In some embodiments, comparing the statistical value to the reference value includes: (1) determining whether the statistical value differs from the reference value by at least a threshold amount, or whether the difference is less than a threshold amount; (2) determining whether the statistical value is less than the reference value by at least a threshold amount; or (3) determining whether the statistical value is greater than the reference value by at least a threshold amount. Renal abnormalities can include renal cell carcinoma (RCC), renal syndrome, glomerulonephritis, Fabry disease, cystinosis, IgA nephropathy, IgM nephropathy, lupus nephritis, atypical hemolytic uremic syndrome (aHUS), polycystic kidney disease (PKD), Alport syndrome, interstitial nephritis, proteinuria, chronic kidney disease, acute kidney injury, and the like. In some cases, the classification of the subject having a renal abnormality includes an increased level of permeability associated with the glomerular basement membrane of the kidney.

[0206] The classification of the kidney abnormality can be determined using machine learning trained using a training dataset. The training dataset can include training samples. The training samples can be related to known classifications of kidney abnormalities. As another example, the comparison to the reference value can be performed using a machine learning model. The machine learning model can be applied to the statistical values ​​to generate a classification of the kidney abnormality. In some embodiments, the machine learning model may include, but is not limited to, a convolutional neural network (CNN), linear regression, logistic regression, a deep recurrent neural network (e.g., a fully connected recurrent neural network (RNN), a gated recurrent unit (GRU), a long short-term memory (LSTM)), a transformer-based method (e.g., XLNet, BERT, XLM, RoBERTa), a Bayesian classifier, a hidden Markov model (HMM), a linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering for applications with noise (DBSCAN), a random forest algorithm, adaptive boosting (AdaBoost), eXtreme Gradient Boosting (XGBoost), a support vector machine (SVM), or a hybrid model including one or more of the models proposed above.

[0207] C. Classification based on urinary cfDNA concentration We also evaluated the difference in urinary cfDNA concentrations between healthy pregnant women and women with preeclampsia. Because cfDNA concentrations in urine samples depend on the subject's hydration state, urinary cfDNA concentrations were normalized.

[0208] In some embodiments, creatinine can be used to correct for urine concentration. For example, the amount of DNA in urine (e.g., measured by mass per volume, such as ng / mL) can be corrected by the amount of creatinine (e.g., mmol). In one implementation, the correction value was calculated by dividing the urinary cfDNA concentration per milliliter of urine sample (e.g., determined by Qubit assay) by the creatinine concentration expressed as nanograms of cfDNA per milliliter per millimole of creatinine (ng / ml / mmol Cr). Creatinine is produced at a constant rate by muscle cells, and all creatinine filtered through the glomerulus is excreted in urine. Therefore, expressing urinary cfDNA concentration per millimole of creatinine will minimize variations in urinary cfDNA concentration resulting from differences in the subject's hydration status.

[0209] Figure 35 shows urinary cfDNA concentration analysis of transrenal DNA in patients with proteinuria and preeclampsia, respectively.

[0210] Boxplot 5310 shows urinary cfDNA concentrations in healthy controls and patients with proteinuria. We observed higher urinary cfDNA concentrations (Mann-Whitney U test, P value = 0.0015) in patients with proteinuria than in healthy controls.

[0211] Boxplot 3520 shows urinary cfDNA concentrations in healthy pregnant women and women with preeclampsia. We observed higher urinary cfDNA concentrations (Mann-Whitney U test, P value = 0.0190) in pregnant women with preeclampsia than in healthy pregnant women.

[0212] 36 is a flowchart of a method 3600 for determining a classification of a kidney abnormality based on a urinary cell-free DNA concentration, according to some embodiments. The method 3600 can detect a kidney abnormality using a urine sample from a subject, the urine sample containing cell-free DNA molecules.

[0213] In block 3602, a first quantity of a plurality of cell-free DNA molecules in a urine sample is determined. For example, the first quantity may be determined using a fluorometer, spectrophotometer, PCR, or sequencing. The first quantity may be filtered to identify cfDNA molecules that meet one or more criteria. For example, the cfDNA may be of a specified size, e.g., greater than a size cutoff, such as 40-200 bp. Examples of size cutoffs are provided herein and include 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 bp. The specified size may have an upper and lower limit, a lower limit, or a size range with an upper limit.

[0214] In block 3604, an initial concentration is determined using a first amount and volume of the urine sample. If size is used as a criterion, the initial concentration can be the percentage of cell-free DNA molecules in the urine sample that fall within a specified range. For example, as described above, the specified range can be greater than a size cutoff.

[0215] In block 3606, a corrected concentration is determined using a second amount of the specific compound in the urine sample. The compound may be a waste product of digestion and therefore may be a compound naturally occurring in the subject. As an example, the specific compound may be creatinine. Creatinine is a waste product resulting from the digestion of protein in food and the normal breakdown of muscle tissue, such as creatine.

[0216] At block 3608, the corrected concentration is compared to a reference value, which can be determined from one or more reference subjects of known classification, e.g., the presence or absence of a renal abnormality, or a particular severity of the renal abnormality.

[0217] At block 3610, a classification of the subject as having a renal abnormality is determined based on the comparison. Examples of renal abnormalities are provided herein and include preeclampsia and proteinuria.

[0218] Additional details of an exemplary method for determining the first quantity are provided. The NanoDrop spectrophotometer is based on the principle that nucleic acids (i.e., DNA and RNA) absorb ultraviolet light with a peak at a wavelength of 260 nanometers (nm). A photodetector measures the light that passes through the sample. The more light absorbed by the nucleic acid, the less light that hits the photodetector, producing a higher optical density (OD), resulting in a higher nucleic acid concentration in the sample.

[0219] The Qubit fluorometer quantifies DNA concentration by detecting fluorescent dyes in samples. The fluorescent dyes, specific for DNA substrates, exhibit extremely low fluorescence before binding to their DNA targets. Upon binding to DNA, the dye molecules increase their fluorescence by several orders of magnitude by intercalating between DNA bases.

[0220] Quantitative PCR (qPCR) assays quantify DNA concentration by detecting the fluorescent signal of DNA products during real-time PCR. QPCR monitors the amplification of targeted DNA molecules during PCR by using DNA probes labeled with fluorescent dyes or fluorescent reporters. The amount of amplified product is then linked to the fluorescence intensity.

[0221] Digital polymerase chain reaction (dPCR) assays involve dividing a PCR solution into tens of thousands of nanoliter-sized droplets, each containing a separate PCR reaction of a single DNA molecule. DNA probes with fluorescent reporters facilitate detection of target DNA in the droplets, and the fraction of droplets containing target DNA can be translated into DNA quantity. Further details can be found in the following three publications, all of which use NanoDrop, Qubit, and qPCR for DNA concentration determination: Simbolo et al., PLOS ONE 2013;8:e62692; Heydt et al., PLOS ONE 2014;9:e104566; and Ponti et al., Clinica Chimica Acta 2018;479:14-19. Further details on dPCR for DNA concentration determination can be found in Gai et al., Clin. Chem. 2018;64:1239-1249.

[0222] D. Technology Comparison Figure 37 shows a ROC analysis of using fragmentomic features of transrenal DNA to distinguish patients with proteinuria and preeclampsia from healthy controls. Figure 37 shows a ROC analysis using the above technique.

[0223] ROC3710 shows AUCs of 0.76, 0.68, 0.73, and 0.75 in distinguishing patients with proteinuria from healthy controls using urinary cfDNA concentration, size (i.e., >80 bp), and O / E in OCR (blood-specific DHS). The AUC can be further improved to 0.85 in distinguishing proteinuria by combining these fragmentomic features with a support vector machine (SVM) method.

[0224] We further performed ROC analysis on these samples, where the AUCs were 0.84, 0.84, and 0.85 in distinguishing between pregnant women with preeclampsia and healthy pregnant subjects using the O / E of urinary DNA concentration, cfDNA size (i.e., >80 bp), and OCR (placenta-specific DHS), respectively. When these three fragmentomic features were combined using SVM, improved performance could be observed in distinguishing between preeclampsia and healthy subjects (AUC: 0.93).

[0225] SVM provides a higher dimensional separation of samples. The number of features input to the SVM will provide the number of dimensions within the SVM. In the example above, we used three features, so three dimensions were used. Additional features can be used, resulting in four or more dimensions.

[0226] VIII. F Profile of cfDNA and Nuclease Activity Different types of cell-free DNA cleavage have been linked to different fragmentation processes, including enzymatic and non-enzymatic breakage. Techniques exist that focus on one specific nuclease activity at a time, using one terminal motif or several top terminal motifs. While such approaches can be effective, they may not provide a comprehensive view of the nuclease activity occurring in a given sample (e.g., plasma sample, urine sample).

[0227] To address the above deficiencies, several nuclease activities or other fragmentation processes can be simultaneously assessed using the estimated relative contributions of different types of cell-free DNA cleavage. For example, the relative frequencies of DNA molecules corresponding to 256 terminal motifs can be determined for a subject with a known disease diagnosis (e.g., HCC). The relative frequencies of DNA molecules can be factorized into a set of "F profiles" that identify the relationships between the terminal sequences (e.g., 1-30 bases) of cell-free DNA fragments (also simply referred to as DNA fragments) in the sample. The set of F profiles can then be used to deconvolve the relative frequencies of DNA molecules obtained from another subject to predict the fraction of clinically relevant DNA molecules, disease classification, etc.

[0228] Figure 38 shows a plot graph 3800 identifying the ranking of the frequency of certain terminal motifs present in urinary cell-free DNA molecules. Figure 38 corresponds to Figure 9. As shown in Figure 38, fetal-specific transrenal DNA primarily contains fragments with C-termini, typically associated with DNASE1L3 cleavage preference. Covalent non-transrenal DNA primarily contains fragments with T-termini, typically associated with DNASE1 cleavage preference.

[0229] While focusing on certain terminal motifs can be beneficial in determining fetal DNA (for example), plot graph 3800 shows additional terminal motif information that can provide further insight; the relative frequency of DNA molecules across most of the 256 terminal motifs differs between fetal-specific and shared DNA. Thus, it can be advantageous to incorporate the relative frequency of DNA molecules across all 256 terminal motifs to determine the fetal DNA fraction or to determine a disease classification for a subject.

[0230] A. Terminal motif profile characteristics across different mouse samples Figure 39 shows a set of graphs 3900 identifying the observed terminal motif profiles of mouse plasma and urinary cell-free DNA molecules. As shown in Figure 39, the frequencies of 256 4-mer terminal motifs in both plasma and urinary cell-free DNA from mice with different nuclease knockout genotypes were alphabetized to form the terminal motif profiles. Motifs beginning with adenine (A), cytosine (C), guanine (G), and thymine (T) are highlighted in blue, red, green, and yellow, respectively.

[0231] We observed certain distinct patterns in the terminal motif profiles across different mice. WT mice, Dnase1l3 - / - Mouse, Dnase1 - / - Mouse, and Dffb - / - The observed terminal motif frequencies in plasma cell-free DNA from mice are shown in graphs 3902, 3904, 3906, and 3908, respectively. WT mice, Dnase1l3 - / - Mouse and Dnase1 - / - The observed terminal motif frequencies in urinary cell-free DNA from mice are shown in graphs 3910, 3912, and 3914, respectively.

[0232] Compared with WT mice, Dnase1l3 - / - Mouse plasma cell-free DNA typically showed periodic spikes in the terminal motif profile, with those terminal motifs having A-, C-, and G-termini. For urinary cell-free DNA from WT mice, the abundance of motifs with T-termini was significantly higher than that of Dnase1. - / - The plasma cell-free DNA and Dnase1 levels in WT mice were significantly increased compared to those in WT mice (P<0.0001, Mann-Whitney U test). - / - Mouse plasma cell-free DNA or WT mouse urine cell-free DNA and Dnase1l3 - / -Although it was visually difficult to discern differences when comparing mouse urinary cell-free DNA, we hypothesized that subtle differences in the 256-dimensional terminal motif profile could be delineated when a reference profile was used, e.g., via factorization into a reference profile. In some embodiments, nonnegative matrix factorization (NMF) was used to consider the 256 motifs as a whole, instead of focusing on one or a few specific motif types.

[0233] The terminal motif profile can be a Kmer, where K can have various values, for example, 1, 2, 3, 4, 5, 6, or more. As shown in Figure 39, a K of 4 is used.

[0234] B. NMF to determine the F profile of urinary cfDNA Figure 40 shows a schematic workflow of an exemplary nuclease usage level analysis for cell-free DNA molecules. 93 mouse cell-free DNA samples were sequenced, including 60 plasma cell-free DNA samples and 33 urinary cell-free DNA samples. Mouse plasma cell-free DNA samples were sequenced from 27 wild-type mice, 27 mice lacking the Dnase1 gene (Dnase1 - / - 10 mice carrying the Dnase1l3 gene deletion (Dnase1l3 - / - ), 18 mice carrying the Dffb gene deletion (Dffb - / - The median number of paired-end reads was 50 million (range: 16-243 million) from five mice with HIV-1. In addition, whole-genome sequencing data from mouse urinary cell-free DNA samples was collected from 14 WT mice and 10 Dnase1 mice. - / - mice, and 9 Dnase1l3 - / - obtained from mice (median number of paired-end reads: 43 million, range: 2–134 million).

[0235] In block 4002, WT mice and nuclease-deficient mice (e.g., Dnase1l3 - / - , Dnase1 - / - , Dffb - / -The terminal four nucleotides at each of the 5' fragment ends (i.e., 4-mer terminal motifs; n=256) were determined for 93 mouse cell-free DNA samples, including the 5' end of the 4-mer terminal motif;

[0236] For each mouse sample, the 256 4-mer terminal motifs of cell-free DNA molecules were then used to infer their respective nuclease usage levels.

[0237] In block 4004, six categories of reference terminal motif profiles, referred to as F profiles, were determined from the cell-free DNA molecules for each mouse sample. In some embodiments, the relative frequencies of DNA molecules terminating in 4-mer terminal motifs were subjected to non-negative matrix factorization (NMF) analysis to determine the underlying different types of cell-free DNA cleavage.

[0238] We applied NMF (Daniel et al. Nature 1999;401:788-791; Stein-O'Brien et al. Trends Genet. 2018;34:790-805) analysis to decompose the relative frequencies of cell-free DNA molecules into several F profiles. A total of 93 mouse cell-free DNA samples with different DNA nuclease knockout genotypes, including 60 plasma cell-free DNA samples and 33 urinary cell-free DNA samples, were used for such NMF analysis. After obtaining the terminal motif frequencies, a data matrix (M) was constructed in which each row represents a cell-free DNA sample (a total of 93 mouse cell-free DNA samples) and each column represents a terminal motif type (a total of 256 terminal motifs), thus having dimensions of 93 × 256. The data matrix was then subjected to NMF analysis to obtain two matrices, W and F. The mathematical relationship between M, W, and F is shown below. M=WF

[0239] M is the product of W and F, and W is the relative weight of each F profile in the 93×n matrix, where n corresponded to the number of F profiles. F represented the F profiles in the n×256 matrix. W and F were determined by minimizing the following objective function: ||M-WF||, subject to W≧0 and F≧0.

[0240] Singular value decomposition (SVD) was used to initialize the NMF procedure. Such factorization analysis was implemented in Python by using the sklearn.decomposition.NMF (v1.1.1) function (Pedregosa et al. J. Mach. Learn. Res. 2011;12:2825-2830).

[0241] To estimate the optimal number of F profiles, a 5-fold cross-validation pre-analysis was performed. Such factorization analysis can result in several different types of cell-free DNA cleavage. In this example, six F profiles (i.e., F profiles I, II, III, IV, V, and VI) were determined by considering the trade-off between the reproducibility of the factorized components and the value of the objective function (i.e., terminal motif profile reconstruction error). The number of different types of cell-free DNA cleavage can be, but is not limited to, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 100, etc., with a corresponding number of reference terminal motif profiles. In Figure 40, F profiles I, II, and III can be associated with the cleavage preferences of DNASE1, DNASE1L3, and DFFB, respectively.

[0242] In block 4006, the nuclease usage analysis framework learned from mouse cell-free DNA can be extrapolated to human cell-free DNA analysis to inform the proportional contribution of different nuclease activities in both mouse and human cell-free DNA samples. The observed terminal motif profiles can be reconstructed by iteratively adjusting the proportional contribution of each F profile. In other words, by using the F profiles generated from mouse cell-free DNA, the proportional contribution of the F profiles for any cell-free DNA sample can be estimated. In some embodiments, such estimated proportional contributions of the F profiles are used to reflect the nuclease activity or nuclease usage level in any cell-free DNA sample.

[0243] Additionally or alternatively, such estimated proportional contributions of the F profile can be used to reflect other types of fragmentation that may be involved in a patient, such as, but not limited to, oxidative stress-induced DNA damage, drug treatment-induced DNA damage, and radiation-induced DNA damage. The presence, absence, and changes in the F profile contributions can indicate disease or risk of developing disease. In some embodiments, other mathematical algorithms are for factorization, such as, but not limited to, component-based analysis (PCA), t-distributed stochastic neighborhood embedding, and uniform manifold approximation and projection.

[0244] C. F profiles in plasma and urine samples Figure 41 shows a diagram 4100 that uses NMF analysis to identify the contribution ratio of each F profile (i.e., nuclease usage level) estimated from mouse cell-free DNA samples with different knockout genotypes. Each F profile can indicate the pattern of relative frequency of cell-free DNA molecules across 256 terminal motifs. Six F profiles can be used as a signature of the terminal motif frequency of cell-free DNA molecules for the corresponding sample. The contribution ratio of each F profile in an individual cell-free DNA sample can be determined when the minimum error is achieved between the observed terminal motif profile and the sum of the F profiles weighted by their contribution ratios.

[0245] As shown in Figure 41, the proportional contributions of the F profiles in mouse samples tend to share similarities based on their respective nuclease activity levels. For example, WT samples showed a significant contribution of F profile I, but not Dnase113. - / - The sample showed little contribution from F profile I. In another example, Dnase1 - / - The sample shows a significantly smaller contribution of F profile II, and Dffb - / - The sample showed little contribution from F profiles III, IV, and V. Similar patterns of F profile contribution can be used to assess the nuclease activity level of another sample from a test subject.

[0246] Fragmentic features of DF profiles As described above, the six F profiles are linked to possible DNA nuclease activities. To illustrate such feasibility, we investigated typical terminal motifs in the F profiles and measured their change in proportional contribution when specific nuclease activities were depleted or enhanced.

[0247] Figure 42 shows a set of plots of six F profiles (A-F) estimated from mouse plasma and urinary cell-free DNA using NMF analysis. Each of F profile plots 4202-4212 includes the 4-mer terminal motif profile, the 1-mer terminal motif frequency, and the sequence preference at each position of the 4-mer terminal motif. F profile 14202 showed a predominance of C-terminal motifs (55%) and featured a "CC" initiation motif, consistent with the DNASE1L3 cleavage properties demonstrated in our previous studies (Serpas et al. Proc Natl Acad Sci USA. 2019;116:641-649, Jiang et al. Cancer Discov. 2020;10:664-73). We investigated the role of Dnase1L3 in the cleavage of Dnase1L3. - / -We observed that the contribution of F-profile I to plasma cell-free DNA in mice was significantly lower than that in WT mice (median: 2.7% vs. 35.4%, range: 0.0-4.6% vs. 19.5-47.9%) (P<0.0001, Mann-Whitney U test). Therefore, F-profile I was considered to be a DNASE1L3-associated F-profile that could be used to reflect the nuclease utilization level of DNASE1L3.

[0248] F-profile II4204 exhibited a significant preference for T-terminal motifs (51%) with a preference for "TG" initiation motifs. In WT mice, the contribution of F-profile II was significantly higher in urinary cell-free DNA compared with plasma cell-free DNA (median: 43.4% vs. 11.6%, range: 31.8-50.1% vs. 0.0-22.1%) (P<0.0001, Mann-Whitney U test). Notably, DNASE1 activity was much higher in urine than in plasma in WT mice (Chen et al. PLOS Genet. 2022;18:e1010262). Furthermore, Dnase1 - / - In both mouse plasma and urinary cell-free DNA, there was a median 8-fold decrease in F-profile II contribution compared with the WT counterpart, and therefore F-profile II was presumed to be related to DNASE1 activity.

[0249] F-profile III4206 contained a significant proportion of A-terminal motifs (40%) and was characterized by a preference for C and T nucleotides at the third and fourth positions, respectively, in 4-mer motifs in the 5' to 3' direction. The contribution of F-profile III was significantly higher in Dffb compared to its WT counterpart (median: 10.1%, range: 0.0-26.9%). - / - Cell-free DNA in mouse plasma was significantly reduced (median: 0.0%, range: 0.0-0.5%) (P = 0.0004, Mann-Whitney U test). Therefore, F-profile III was considered to be associated with DFFB activity.

[0250] F-profile IV4208 exhibited a high C-terminus preference (50%), somewhat reminiscent of F-profile I, but it also had some distinct features, such as the absence of a C-terminus preference. F-profile IV also exhibited a preference for "G" bases at the second, third, and fourth positions of 4-mer motifs. F-profile V4210 exhibited a strong G-terminus preference (50%). These results suggested that F-profiles IV and V were not directly attributable to known nucleases involved in cell-free DNA fragmentation and implied that several other enzymatic and / or nonenzymatic processes may play a role in the cell-free DNA fragmentation process. In addition, F-profile VI4212 showed a relatively uniform distribution across the 256 motifs without any apparent sequence preference, suggesting one possibility that other DNA nucleases or other factors may also cause nonspecific cleavage.

[0251] 43 shows a box plot 4300 identifying the proportional contribution of F profile I across different types of samples, according to some embodiments. The x-axis shows wild-type (WT) mouse samples (n=27), Dnase 1 - / - and Dnase1l3 - / - Mouse samples (n=2), Dnase1l3 - / - Mouse samples (n=18), and Dnase1l3 - / - and Cd40lg - / - The x-axis shows cell-free DNA molecules obtained from different types of samples, including WT mouse samples (n=5), fetal DNA (n=2), and Dnase113 (n=3). + / - Dnase1l3 - / - Maternal samples (n=4) and fetal DNA were analyzed for Dnase1l3. - / - Dnase1l3 - / - Various types of pregnant mouse samples are shown, including maternal samples (n=3). The contribution ratio of F profile I was determined relative to the contributions of the other F profiles II–VI.

[0252] As shown in Figure 43, the proportional contribution of F profile I was significantly reduced for DNASE1L3 knockout mouse samples. Furthermore, the significant reduction in the contribution of F profile I appears to be similar for pregnant samples. Interestingly, fetal Dnase1L3 + / - The contribution of F profile I to the sample is fetal Dnase1l3 - / - The sample shows a slight increase in the contribution of F profile I. Based on the proportional contribution of F profile I, it can be determined that F profile I corresponds to a signature associated with DNASE1L3.

[0253] Figure 44 shows the relative frequencies 4400 of cell-free DNA molecules across 256 terminal motifs for F profile I, according to some embodiments. Plot 4410 shows the terminal motif profile using 4-mers as terminal motifs. Plot 4420 shows the terminal motif profile using 1-mers (single nucleotides) as terminal motifs. As shown in the plots in Figure 44, F profile I has a cleavage preference for the C-terminus, followed by the T-terminus, and then the A-terminus and G-terminus. This cleavage preference is substantially similar to the cleavage preference of DNASE1L3, supporting the findings shown in Figure 43.

[0254] 45 shows a box plot 4500 identifying the proportional contribution of F-profile II across different sample types, according to some embodiments. The x-axis shows wild-type (WT) plasma samples (n=27), WT urine samples (n=14), Dnase 1 - / - Plasma samples (n = 10), Dnase1 / Dnase1l3 double-deficient plasma samples (n = 2), and Dnase1 - / - Cell-free DNA molecules obtained from different types of mouse samples, including urine samples (n=10), are shown. The contribution ratio of F-profile II was determined relative to the contribution ratios from the other F-profiles I and III-VI.

[0255] As shown in Figure 45, the proportional contribution of F profile II is significantly higher for urine samples compared to plasma samples. Furthermore, the proportional contribution of F profile II does not show significant changes across different types of plasma samples. However, Dnase 1 - / - The proportional contribution of F-profile II in the urine samples is significantly reduced from that in the WT urine samples. Based on this proportional contribution, it can be determined that F-profile II corresponds to a signature associated with DNASE1. Indeed, DNase1 nuclease activity appears to be more active in the urine samples, which correlates with the higher F-profile II contribution relative to the WT urine samples.

[0256] Figure 46 shows the relative frequency 4600 of cell-free DNA molecules across 256 terminal motifs for F profile II, according to some embodiments. As shown in Figure 46, F profile II has a cleavage preference for the T terminus, followed by the C terminus, and then the A terminus and the G terminus. This cleavage preference is substantially similar to that of DNASE1, supporting the findings shown in Figure 45.

[0257] 47 shows a box plot 4700 identifying the proportional contribution of F profile III across different types of samples, according to some embodiments. The x-axis shows the proportion of wild-type (WT) samples (n=27) and Dffb - / - Figure 47 shows cell-free DNA molecules obtained from samples containing WT samples (n=5). The proportional contribution of F profile III was determined relative to the contributions from the other F profiles I, II, and IV-VI. As shown in Figure 47, the proportional contribution of F profile III was significantly higher than that of the WT sample. - / - Based on the proportional contribution of F profile III, it can be determined that F profile III corresponds to the signature associated with DFFB.

[0258] Figure 48 shows the relative frequencies 4800 of cell-free DNA molecules across the 256 terminal motifs of F-profile III, according to some embodiments. As shown in Figure 48, F-profile III has a cleavage preference for the A-terminus, followed by the G-terminus, and then the C-terminus and the T-terminus. This cleavage preference is substantially similar to that of DFFB, supporting the findings shown in Figure 47.

[0259] In addition to F-profiles I–III, we also resolve F-profiles IV–VI.

[0260] Figure 49 shows the relative frequencies 4900 of cell-free DNA molecules across 256 terminal motifs for F-profiles VI-VI, according to some embodiments. Each of F-profiles IV-VI exhibits a unique cleavage preference. For example, F-profile IV 4902 exhibits a cleavage preference for the C-terminus, F-profile V 4904 exhibits a cleavage preference for the G-terminus, and F-profile VI 4906 exhibits no particular cleavage preference. Based on the above, it can be shown that F-profile VI may be associated with cleavage patterns caused not by the specific nucleases we investigated, but by other types of fragmentation agents.

[0261] E. Analysis of F profiles across different samples from human subjects Next, we investigated whether the murine F profile of DNASE-mediated cell-free DNA breaks could be applied to human subjects.

[0262] Figure 50 shows a schematic diagram 5000 comparing a terminal motif profile of a human subject with a reference F profile determined based on a mouse sample, according to some embodiments. To make motif patterns directly comparable between humans and mice, the frequencies of 4-mer terminal motifs associated with human and mouse cell-free DNA can be normalized by the genomic context of the human and mouse genomes, respectively. For example, expected 4-mer terminal motif frequencies can be used in the normalization process, where the expected terminal motif frequencies were determined by simulating 4-mer terminal motifs from the reference genome using a 4-bp sliding window across each chromosome. The normalized terminal motif frequency was calculated as the ratio of the observed frequency to the expected frequency, and then divided by the sum of all 256 normalized motif frequencies. All normalized terminal motif frequencies may be equal to 100%. The terminal motif frequencies mentioned in this NMF-based nuclease usage analysis were referred to as normalized terminal motif frequencies.

[0263] Once normalization is complete, the proportional contribution of the F profile can be determined for the normalized termination frequencies of the human sample. The proportional contribution can be determined by applying deconvolution to the normalized termination frequencies. For example, a data matrix M of dimension W by F can be used, where (i) M can represent the normalized termination frequency across 256 terminal motifs for a given biological sample, (ii) F can represent the termination frequency of a reference F profile obtained from a mouse sample, and (iii) W can represent the relative weight corresponding to the proportional contribution of each F profile.

[0264] The F termination frequency can be determined based on the proportion of cell-free DNA molecules in a set of reference F profiles. The proportional contribution can be determined by solving for the W relative weights based on the data matrix M and values ​​from the reference F profiles using non-negative least squares (NNLS). The proportional contribution determined using deconvolution can be used to identify the degree of nuclease activity in a particular human biological sample (e.g., the relative reduction in F profile I contribution).

[0265] Figure 51 shows the proportional contributions 5100 of F profiles across plasma and urine samples from a human subject, according to some embodiments. First, the deconvolution process of Figure 51 was applied to plasma and urine samples from a human subject. As shown in Figure 51, the plasma sample contained a relatively high contribution from F profile I, which is associated with DNASE1L3 activity. In contrast, the urine sample contained a relatively high contribution from F profile II, which is associated with DNASE1 activity.

[0266] The data shown in Figure 51 is consistent with the experiment that DNASE1L3 has a large contribution to the fragmentation pattern of cell-free DNA molecules in plasma samples, and DNASE1 has a large contribution to the fragmentation pattern of cell-free DNA molecules in urine samples.Therefore, it is shown that the deconvolution process using the reference F profile from mouse samples can be effectively used to identify the fragment cleavage pattern of cell-free DNA molecules in human samples.

[0267] Figure 52 shows the proportional contribution 5200 of F profiles across normal and DNASE1L3-deficient samples from a human subject, according to some embodiments. The deconvolution process of Figure 52 was applied to normal and DNASE1L3-deficient samples from a human subject. As shown in Figure 52, plasma samples from control subjects contained a relatively high contribution (greater than about 40%) of F profile I. In contrast, DNASE1L3-deficient samples contained a significantly lower contribution (less than about 15%) from F profile I. The data shown in Figure 52 is consistent with previous correlations between F profile I and the nuclease activity of DNASE1L3. It may be shown that the deconvolution process using reference F profiles from mouse samples can be effectively used to identify fragment cleavage patterns of cell-free DNA molecules in human samples.

[0268] Figure 53 shows the proportional contribution 5300 of F profiles across urine samples from pregnant human subjects, according to some embodiments. The deconvolution process of Figure 42 was applied to samples from pregnant human subjects. As shown in Figure 53, the urine samples contained a relatively high contribution from F profile II, which is associated with DNASE1 activity. The data shown in Figure 53 is consistent with experiments showing that DNASE1 has a large contribution to the fragmentation pattern of cell-free DNA molecules in urine samples. Furthermore, the higher proportional contribution of F profile IV may indicate a relatively large number of cell-free DNA fragments with C-termini. This observation is consistent with the plotted data of Figure 38, showing that fetal DNA in urine samples from pregnant subjects contains a greater proportion of cell-free DNA molecules with C-termini.

[0269] How to classify nuclease activity using FF profiles 54 shows a flowchart of a method 5400 for determining a classification of nuclease activity based on the F profile of cell-free DNA molecules, according to some embodiments. Exemplary biological samples can be cell-free samples containing cell-free DNA, such as blood, plasma, serum, urine, and saliva.

[0270] At block 5402, a set of reference F profiles is stored. Each reference F profile of the set identifies, for each nucleotide in the set of nucleotides, the proportion of cell-free DNA molecules that terminate at that nucleotide. In addition, each reference F profile is associated with a type of fragmentation factor. The type of fragmentation factor identifies a specific enzyme (e.g., DNASE1L3, DNASE1), protein (e.g., DFFB), or other biological component or process that causes fragmentation in cell-free DNA molecules. In some cases, the set of reference F profiles includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 45, 50, or more than 50 F profiles. For example, the set of reference F profiles may include six F profiles I-VI.

[0271] Each reference F profile and sample end motif profile can have a distinct proportion for each Kmer end motif in the set of Kmer end motifs. For example, Figure 39 shows a profile with K=4, resulting in 256 different proportion values. Plot 4420 in Figure 44 shows a profile with K=1, resulting in four different proportion values. Thus, each reference F profile in a set of reference F profiles can specify the proportion of cell-free DNA molecules that terminate at each Kmer end motif in the set of Kmer end motifs, where K is one or more.

[0272] In some cases, the set of reference F profiles is determined using one or more reference samples. The reference samples have a known genetic disorder classification (e.g., WT, DNASE1L3 - / - , DNASE1 - / -) can be obtained from a non-human subject (e.g., a mouse sample). To determine a reference F profile for a set of reference profiles, a factorization algorithm (e.g., NMF, PCA) is used to decompose the relative frequencies of cell-free DNA molecules in the reference sample into several F profiles. For example, reference cell-free DNA samples with different genotypes of DNA nuclease knockout were selected. After obtaining the terminal motif frequencies of the reference samples, a data matrix (M) is constructed in which each row represents a cell-free DNA sample (e.g., a total of 93 mouse cell-free DNA samples) and each column represents one type of terminal motif (e.g., a total of 256 4-mer terminal motifs), thus having dimensions of 93 x 256. The data matrix can then be subjected to NMF analysis to obtain two matrices, W and F. M=WF

[0273] M is the product of W and F, and W is the relative weight of each F profile in the 93×n matrix, where n corresponded to the number of F profiles. F represented the F profiles in the n×256 matrix. W and F were determined by minimizing the following objective function: ||M-WF||, subject to W≧0 and F≧0.

[0274] At block 5404, a plurality of cell-free DNA fragments from the biological sample are analyzed to obtain sequence reads. The sequence reads may include terminal sequences corresponding to the ends of a plurality of cell-free DNA molecules. The sequence reads may include terminal sequences corresponding to the ends of a plurality of cell-free DNA fragments. By way of example, the sequence reads may be obtained using sequencing or probe-based techniques, either of which may include enrichment via, for example, amplification or capture probes.

[0275] Sequencing can be performed in a variety of ways, for example, using massively parallel sequencing or next-generation sequencing, using single molecule sequencing, and / or using double-stranded or single-stranded DNA sequencing library preparation protocols. Those skilled in the art will understand the various sequencing techniques that can be used. As part of sequencing, it is possible that some of the sequence reads may correspond to cellular nucleic acids.

[0276] The sequencing can be, for example, targeted sequencing as described herein. For example, a biological sample can be enriched for DNA fragments from a specific region. Enrichment can include using a capture probe that binds to a portion or the entire genome, for example, as defined by a reference genome.

[0277] A statistically significant number of cell-free DNA molecules can be analyzed to provide an accurate determination of fractional concentration. In some embodiments, at least 1,000 cell-free DNA molecules are analyzed. In other embodiments, at least 10,000, or 50,000, or 100,000, or 500,000, or 1,000,000, or 5,000,000 or more cell-free DNA molecules can be analyzed.

[0278] In block 5406, a sample end motif profile of the subject is determined by determining the proportion of a plurality of cell-free DNA molecules that terminate at each nucleotide of the set of nucleotides based on the end sequences. The sample end motif profile identifies the relative frequencies of a plurality of end motifs corresponding to the end sequences of the plurality of cell-free DNA molecules. The plurality of end motifs can correspond to all possible combinations of N base positions. For example, if the plurality of end motifs correspond to 4-mers, the plurality of end motifs in the sample end motif profile can include 256 4-mer combinations (e.g., CCCA, TTCC).

[0279] To determine a sample end motif profile, for each of a plurality of cell-free DNA fragments, a sequence motif is determined for each of one or more end sequences of the cell-free DNA fragments. A sequence motif can include N base positions (e.g., 1, 2, 3, 4, 5, 6, etc.). By way of example, a sequence motif can be determined by analyzing sequence reads at ends corresponding to the ends of the DNA molecules and correlating the signals with particular motifs (e.g., if probes are used) and / or aligning the sequence reads to a reference genome.

[0280] For example, after sequencing by a sequencing device, sequence reads can be received by a computer system that can be communicatively coupled to the sequencing device, e.g., via wired or wireless communication or a removable storage device. In some implementations, one or more sequence reads including both ends of a nucleic acid fragment can be received. The location of a DNA molecule can be determined by mapping (aligning) one or more sequence reads of the DNA molecule to a respective portion, e.g., a specific region, of the human genome. Additionally or alternatively, a specific probe (e.g., after PCR or other amplification) can indicate a location or a specific terminal motif, such as via a specific fluorescent color. Identification can be that the cell-free DNA molecule corresponds to one of multiple terminal motifs.

[0281] Then, the relative frequencies of the plurality of terminal motifs corresponding to the terminal sequences of the plurality of cell-free DNA molecules are determined to determine the sample terminal motif profile of the subject. The relative frequencies of the sequence motifs can provide the proportion of the plurality of cell-free DNA fragments that have terminal sequences corresponding to the sequence motifs.

[0282] In block 5408, the proportional contribution for the set of reference F profiles, whose proportional aggregation provides the sample terminal motif profile, is determined. The proportional contributions of the set of reference F profiles sum to 1. The proportional contribution can be determined by applying deconvolution to the subject's sample terminal motif profile. For example, a data matrix M of dimension W by F can be used, where (i) M can represent the normalized termination frequency across 256 terminal motifs for the sample terminal motif profile, (ii) F can represent the termination frequency of the reference F profile obtained from the mouse sample, and (iii) W can represent the relative weight corresponding to the proportional contribution of each F profile. The proportional contribution can be determined by solving for W based on the data matrix M and the values ​​from the reference F profile F using non-negative least squares (NNLS). The proportional contribution determined using deconvolution can be used to identify the level of fragmentation factor activity in the subject (e.g., a relative decrease in the F profile I contribution).

[0283] In some cases, the frequency of 4-mer end motifs in the sample end motif profile of a subject (e.g., a human subject) and the frequency of 4-mer end motifs in the sample end motif profile of a reference sample (e.g., a mouse sample) are normalized by the genomic context of their respective genomes. For example, expected 4-mer end motif frequencies can be used in the normalization process, where the expected end motif frequencies are determined by simulating 4-mer end motifs from the reference genome using a 4-bp sliding window across each chromosome. The normalized end motif frequency is calculated as the ratio of the observed frequency to the expected frequency, and then divided by the sum of all 256 normalized motif frequencies. All normalized end motif frequencies may be equal to 100%.

[0284] In block 5410, a classification of the nuclease activity of a particular type of nuclease is determined based on the proportional contribution associated with a particular type of fragmentation factor. For example, the classification of the nuclease activity of a particular type of nuclease can include a classification of a decrease in nuclease activity associated with a particular nuclease. The classification of nuclease activity can be used to determine whether a subject has a nuclease activity deficiency or a genetic disorder for a gene associated with the nuclease. The genetic disorder can be a disorder of the DNASE1L3 gene. The genetic disorder may include a disorder of one or more of the following genes: DNASE1, DFFB, TREX1 (3 prime repair exonuclease 1), AEN (apoptosis-enhancing nuclease), EXO1 (exonuclease 1), DNASE2 (deoxyribonuclease 2), ENDOG (endonuclease G), APEX1 (apurinic / apyrimidinic site endodeoxyribonuclease 1), FEN1 (flap structure-specific endonuclease 1), DNASE1L1 (deoxyribonuclease 1-like 1), DNASE1L2 (deoxyribonuclease 1-like 2), and EXOG (exo / endonuclease G).

[0285] In some cases, the reduction in the level of nuclease activity associated with a particular type of nuclease is determined based on the proportional contribution of a set of reference F profiles.For example, the proportional contribution associated with one of a set of reference F profiles can be compared with a cutoff value.Based on the comparison (for example, if the proportional contribution exceeds the cutoff value), the reduction in the level of nuclease activity can be determined.In some cases, the cutoff value is determined using one or more reference samples with known classification of nuclease activity.

[0286] IX. Fraction Concentration Using cfDNA F Profile Because transrenal cell-free DNA molecules still preserve the DNASE1L3 cleavage signature of plasma cell-free DNA, we reasoned that urinary nuclease usage levels could potentially represent the amount of transrenal cell-free DNA. We hypothesized that NMF-based nuclease usage level analysis might be feasible to determine the fractional contribution of transrenal cell-free DNA in urine samples. To this end, we applied nuclease usage level analysis to 14 maternal urine samples.

[0287] Figure 55 shows a set of graphs 5500 identifying nuclease usage levels in urinary cell-free DNA of pregnant women. Graph 5502 shows the correlation between fetal DNA fraction and F Profile I (DNASE1L3) levels. Graph 5504 shows the correlation between fetal DNA fraction and F Profile IV levels. As shown in Figure 55, we found that the proportional contributions of F Profile I (DNASE1L3) and F Profile IV significantly correlated with the fetal DNA fraction in maternal urinary cell-free DNA estimated by a SNP-based approach (Pearson's r = 0.60, P = 0.025). Therefore, the nuclease usage level analysis presented in this disclosure may be useful for monitoring the proportion of transrenal cell-free DNA in urine samples.

[0288] Figure 56 shows a flowchart of a method 5600 for determining the fractional concentration of fetal DNA based on an F profile of cell-free DNA molecules, according to some embodiments. An exemplary biological sample can be an acellular sample containing cell-free DNA, such as blood, plasma, serum, urine, and saliva. The biological sample can contain clinically relevant DNA and other DNA that is acellular. In other examples, the biological sample may not contain clinically relevant DNA, and the estimated fractional concentration may show zero or a low percentage of clinically relevant DNA.

[0289] The biological sample may contain a mixture of cell-free DNA molecules from one or more tissue types, such as heart, lung, and liver. For example, the biological sample may be obtained from a pregnant woman and may contain maternal and fetal cell-free DNA molecules. The biological sample may contain tumor-specific cell-free DNA molecules as well as other tissue-specific cell-free DNA molecules. The clinically relevant DNA molecules may be from any of the tissue types described herein, e.g., fetal DNA, tumor DNA, or transplant DNA. Method 5600 and any other method embodiments described herein may be implemented by a computer system. Embodiments of method 5600 may be implemented in a manner similar to method 5400 of FIG. 54.

[0290] In block 5602, a set of reference F profiles is stored. Each reference F profile of the set can identify, for each nucleotide in the set of nucleotides, the proportion of cell-free DNA molecules that terminate at that nucleotide. In addition, each reference F profile can be associated with a type of fragmentation factor. Such fragmentation factors can identify specific enzymes (e.g., DNASE1L3, DNASE1), proteins (e.g., DFFB), or other biological components or processes that cause fragmentation in cell-free DNA molecules. In some cases, the set of reference F profiles includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 45, 50, or more than 50 F profiles. For example, the set of reference F profiles can include six F profiles I-VI. Block 5602 can be performed in a manner similar to block 5402 of FIG. 54.

[0291] In block 5604, a plurality of cell-free DNA fragments from the biological sample are analyzed to obtain sequence reads. The sequence reads may include terminal sequences corresponding to the ends of the plurality of cell-free DNA molecules. The sequence reads may include terminal sequences corresponding to the ends of the plurality of cell-free DNA fragments. Block 5604 may be performed in a manner similar to block 5404 of FIG. 54.

[0292] In block 5606, a sample end motif profile of the subject is determined by determining the proportion of a plurality of cell-free DNA molecules that terminate at each nucleotide of the set of nucleotides based on the end sequences. The sample end motif profile identifies the relative frequencies of a plurality of end motifs corresponding to the end sequences of the plurality of cell-free DNA molecules. The plurality of end motifs can correspond to all possible combinations of N base positions. For example, if the plurality of end motifs correspond to 4-mers, the plurality of end motifs of the sample end motif profile can include 256 4-mer combinations (e.g., CCCA, TTCC). Block 5606 can be performed in a manner similar to block 5406 of FIG. 54.

[0293] In block 5608, proportional contributions for a set of reference F profiles whose proportional aggregation provides a sample end motif profile are determined. The proportional contributions of the set of reference F profiles sum to 1. Block 5608 may be performed in a manner similar to block 5408 of FIG. 54.

[0294] The set of reference F profiles includes, for example, a first reference F profile that correlates with the fractional concentration of a clinically relevant DNA molecule determined using a calibration sample with a known fractional concentration. Figure 55 provides such an example, in which the first reference F profile is F profile I (corresponding to DNASE1L3) or F profile IV. The first proportional contribution for the set of reference F profiles can correspond to the first reference F profile.

[0295] In block 5610, the fractional concentration of clinically relevant DNA molecules in the biological sample is estimated by comparing a first proportional contribution corresponding to the first reference F profile to one or more calibration values ​​determined from one or more calibration samples in which the fractional concentrations of clinically relevant DNA molecules are known. An embodiment of block 5610 may be performed in a manner similar to block 2106 of method 2100. The first reference F profile may correspond to a particular type of nuclease, e.g., DNASE1L3.

[0296] As shown in Figure 55, the fetal DNA fraction increases as the proportional contribution of one or more reference F profiles (e.g., F profile I, F profile IV) increases. Any of the one or more reference F profiles can be the first reference F profile as long as it correlates with the fractional concentration of clinically relevant DNA. Therefore, the known proportional contribution of one or more reference F profiles in the calibration sample can be used as a calibration data point to determine the fractional concentration of clinically relevant DNA molecules in the biological sample.

[0297] Some embodiments can measure the fractional concentration of clinically relevant DNA molecules in one or more calibration samples for each calibration sample, and measure the proportional contribution of the first reference terminal motif profile for the calibration sample, thereby determining one or more calibration data points. The proportional contribution of the first reference terminal motif profile can be used as a calibration value. The proportional contribution of all of the sets of reference F profiles can be determined, thereby determining multiple calibration values, for example, when multiple reference F profiles are used to estimate fractional concentrations. Those skilled in the art will understand that fractional concentrations can be measured in a variety of ways, some of which, for example, using tissue-specific alleles or tissue-specific methylation patterns, are described herein.

[0298] In some cases, the proportional contributions determined for one or more of the reference F profiles (including the first proportional contribution for the first reference F profile) are compared to a calibration curve (composed of calibration data points), and thus the comparison can identify points on the curve with known proportional contributions for the biological sample. The fractional concentrations corresponding to the identified points can then be used to estimate the fractional concentrations. For example, the determined comparative contributions can be provided as inputs to a calibration function (e.g., a linear or nonlinear fit) to obtain an output of fractional concentrations.

[0299] In some embodiments, as described above, multiple proportional contributions can be used.In such an example, the calibration curve can be a calibration surface in two or more dimensions.Therefore, estimating the fractional concentration of clinically relevant DNA molecules in biological samples comprises comparing one or more additional proportional contributions with one or more additional calibration values ​​determined from one or more calibration samples in which the fractional concentration of clinically relevant DNA molecules is known.

[0300] X. Estimation of gestational age based on F profiles in cfDNA The nuclease usage level determined from the factorization analysis of cell-free DNA molecules can be used to estimate the gestational age of the fetus in the sample obtained from pregnant women.For example, the proportional contribution of the F profile obtained from the sample of known gestational age can be determined.Then, the proportional contribution determined can be used as a calibration data point to estimate the gestational age of the sample from another pregnancy.

[0301] As described further below, there was a correlation between the proportional contribution of F-profile I and fetal gestational age. The correlation may also indicate that because F-profile I represents the cleavage preference of DNASE1L3, gestational age may be influenced based on DNASE1L3 activity levels.

[0302] A. Estimation of gestational age Figure 57 shows a set of graphs 5700 identifying nuclease usage levels in plasma cell-free DNA of pregnant women. Boxplots 5702 and 5704 show DNASE1L3 expression levels in the placenta of pregnant women across different gestational stages. Boxplot 5706 shows F profile I (DNASE1L3) contribution in maternal plasma cell-free DNA of pregnant women across first, second, and third trimesters. Graph 5708 shows the correlation between fetal DNA fraction and F profile I (DNASE1L3) levels. As shown in boxplots 5702 and 5704, an upregulation of DNASE1L3 gene expression levels along gestational age was observed in placental tissues based on transcriptome data (Mikheev et al. Reprod. Sci. 2008; 15: 866-877, Sitras et al. PLOS One 2012; 7: e33294).

[0303] Nuclease usage level analysis based on NMF can be used to estimate gestational age based on specific F profiles in cell-free DNA. We analyzed nuclease usage levels based on terminal motifs in maternal plasma using a previously published cohort of 30 pregnant women (10 in each trimester) (Jiang et al. Clin. Chem. 2017;63:606-608). As shown in the boxplots, we observed a progressive increase in F profile I (DNASE1L3) levels in maternal plasma cell-free DNA across the first trimester (median: 40.2%, range: 38.5-42.7%), second trimester (median: 41.3%, range: 36.2-42.8%), and third trimester (median: 43.1%, range: 34.5-44.0%).

[0304] The nuclease usage level analysis disclosed herein can also be used to determine the fractional contribution of fetal DNA in plasma samples. As shown in graph 5708, the F profile I (DNASE1L3) level in maternal plasma cell-free DNA significantly correlated with the fetal DNA fraction estimated by the SNP-based approach (Pearson's r = 0.40, P = 0.027). Therefore, nuclease usage level analysis can be useful for monitoring physiological conditions such as pregnancy.

[0305] B. Relationship between oxidative stress and gestational age Apart from cancer patients, we also studied plasma from pregnant women from the first trimester (n=10), second trimester (n=10), and third trimester (n=10). Previous studies have revealed that oxidative stress in the placenta has been reported to decrease with increasing gestational age (Basu et al. Obstet Gynecol Int 2015;2015:276095).

[0306] Figure 58 shows a set of graphs 5800 identifying F-profile analysis and oxidative stress levels in pregnant women. Graph 5802 shows oxidative stress levels in placental tissue from pregnant women at different stages of pregnancy. Boxplot 5804 shows the F-profile VI contribution in fetal-specific DNA in the plasma of pregnant women across different stages of pregnancy. Boxplot 5806 shows the F-profile VI contribution in maternal-specific DNA in the plasma of pregnant women across different stages of pregnancy. As shown in the graphs in Figure 58, as the pregnancy stage progressed, the median F-profile VI levels in fetal-specific DNA significantly decreased (early: 26.7%, mid-term: 23.7%, late: 22.0%) (P = 0.014, Kruskal-Wallis test), while no significant change was observed in F-profile VI levels in maternal-specific DNA. These data indicated that F-profile VI levels can indicate the contribution of cell-free DNA due to oxidative stress-induced fragmentation.

[0307] C. Methods for estimating gestational age based on cell-free DNA F profiles Figure 59 shows a flowchart of a method 5900 for estimating gestational age based on an F profile of cell-free DNA molecules, according to some embodiments. A biological sample is obtained from a female subject carrying a fetus. The biological sample may be a sample having cell-free DNA molecules from the female subject and the fetus, such as a plasma, serum, urine, saliva, cerebrospinal fluid, pleural fluid, amniotic fluid, peritoneal fluid, or ascites sample. Aspects of method 5900 and any other method described herein may be implemented by a computer system. Aspects of method 5900 may be implemented in a manner similar to method 5400 of Figure 54.

[0308] In block 5902, a set of reference F profiles is stored. Each reference F profile of the set identifies, for each nucleotide in a set of nucleotides, the proportion of cell-free DNA molecules that terminate at that nucleotide. In addition, each reference F profile is associated with a type of fragmentation factor. The type of fragmentation factor identifies a specific enzyme (e.g., DNASE1L3, DNASE1), protein (e.g., DFFB), or other biological component or process that causes fragmentation in cell-free DNA molecules. In some cases, the set of reference F profiles includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 45, 50, or more than 50 F profiles. For example, the set of reference F profiles may include six F profiles I-VI. Block 5902 may be performed in a manner similar to block 5402 of FIG. 54.

[0309] In block 5904, a plurality of cell-free DNA fragments from the biological sample are analyzed to obtain sequence reads. The sequence reads may include terminal sequences corresponding to the ends of the plurality of cell-free DNA molecules. The sequence reads may include terminal sequences corresponding to the ends of the plurality of cell-free DNA fragments. Block 5904 may be performed in a manner similar to block 5404 of FIG. 54.

[0310] In block 5906, a sample end motif profile of the subject is determined by determining the proportion of a plurality of cell-free DNA molecules that terminate at each nucleotide of the set of nucleotides based on the end sequences. The sample end motif profile identifies the relative frequencies of a plurality of end motifs corresponding to the end sequences of the plurality of cell-free DNA molecules. The plurality of end motifs can correspond to all possible combinations of N base positions. For example, if the plurality of end motifs correspond to 4-mers, the plurality of end motifs of the sample end motif profile can include 256 4-mer combinations (e.g., CCCA, TTCC). Block 5906 can be performed in a manner similar to block 5406 of FIG. 54.

[0311] In block 5908, proportional contributions for a set of reference F profiles whose proportional aggregation provides a sample end motif profile are determined. The proportional contributions of the set of reference F profiles sum to 1. Block 5908 may be performed in a manner similar to block 5408 of FIG. 54.

[0312] The set of reference F profiles includes a first reference F profile correlated with gestational age, for example, determined using a calibration sample with a known gestational age. Figures 57 and 58 provide such examples in which the first reference F profile is F profile I (corresponding to DNASE1L3) or F profile IV. The first proportional contribution for the set of reference F profiles can correspond to the first reference F profile.

[0313] In block 5910, the gestational age of the fetus is estimated by comparing a first proportional contribution corresponding to the first reference F profile to one or more calibration values ​​determined from one or more calibration samples of known gestational age. Embodiments of block 5910 may be implemented in a manner similar to blocks 2106 and 5610. The first reference F profile may correspond to a particular type of nuclease, e.g., DNASE1L3, as shown in FIG. 57. A reference F profile that does not correspond to a nuclease, e.g., F profile VI, may also be used, as shown in FIG. 58.

[0314] As an example, the proportional contribution of the reference F profile (e.g., F profile I representing DNASE1L3) of the calibration data point can be plotted on a chart to form clusters for different gestational ages, and the determined proportional contribution of the biological sample can also be plotted on the chart to determine the cluster to which the biological sample belongs. Any of one or more reference F profiles can be the first reference F profile as long as it correlates with gestational age. Thus, the known proportional contribution of one or more reference F profiles in the calibration sample can be used as a calibration data point for determining gestational age.

[0315] Therefore, some embodiments can measure the gestational age in each calibration sample of one or more calibration samples and measure the proportional contribution of the first reference terminal motif profile determined for the calibration sample. For example, menstrual history and ultrasound are two methods for measuring gestational age. For example, gestational age can be estimated based on the date of the last menstrual period. Conception can be assumed to occur on day 14 of the cycle, which can be affected by variations in menstrual cycle and ovulation between individuals. Ultrasound measurement of the embryo or fetus in early pregnancy can be the most accurate method for establishing gestational age. Gestational age can be estimated from ultrasound using various parameters, such as mean sac diameter (MSD), crown rump length (CRL), biparietal diameter (BPD), and head circumference (HC).

[0316] The proportional contribution of the first reference terminal motif profile can be used as calibration value.The proportional contribution of all of the set of reference F profiles can be determined, thereby determining multiple calibration values, for example, when multiple reference F profiles are used to estimate gestational age.Those skilled in the art will understand that gestational age can be measured in various ways.

[0317] In some cases, the proportional contributions determined for one or more of the reference F profiles (including the first proportional contribution for the first reference F profile) can be compared to a calibration curve (comprised of calibration data points), and the comparison can then identify a point on the curve with a known proportional contribution for the biological sample. The gestational age corresponding to the identified point can then be used to estimate the gestational age. For example, the determined proportional contributions can be provided as input to a calibration function (e.g., a linear or nonlinear fit) to obtain an output of gestational age.

[0318] In some embodiments, as described above, multiple proportional contributions can be used. In such examples, the calibration curve can be a calibration surface in two or more dimensions. Thus, estimating gestational age can include comparing one or more additional proportional contributions to one or more additional calibration values ​​determined from one or more calibration samples with known gestational ages.

[0319] XI. Classification of pathologies based on F profiles in cfDNA The F profile can also be used to classify the level of pathology in a subject. Examples of pathologies are autoimmune diseases (e.g., SLE) and cancer.

[0320] A. Systemic lupus erythematosus (SLE) Nuclease usage level analysis can be used to distinguish between human subjects with and without DNASE1L3 deficiency based on a specific F profile of cell-free DNA. Human subjects with DNASE1L3 deficiency develop systemic lupus erythematosus (SLE)-like symptoms with childhood onset, also known as familial SLE (Chan et al. Am. J. Hum. Genet. 2020;107:882-894). We investigated nuclease usage levels by analyzing plasma cell-free DNA from patients (n=10) carrying both copies of the DNASE1L3 gene with genetic mutations (i.e., DNASE1L3 deficiency), parents of these patients (n=3) carrying one copy of the mutant DNASE1L3 gene (i.e., the other copy was functional), and healthy control subjects (n=8) (Chan et al. Am. J. Hum. Genet. 2020;107:882-894).

[0321] Figure 60 shows a boxplot 6000 of F profile I (DNASE1L3) levels in healthy subjects, patients with DNASE1L3 deficiency, and the patient's parents. As shown in Figure 60, F profile I (DNASE1L3) in plasma cell-free DNA of patients with DNASE1L3 deficiency appeared to be significantly reduced (median: 7.3%, range: 3.8-2.5%) compared to their parents (median: 51.4%, range: 47.4-51.9%) and healthy subjects (median: 52.9%, range: 47.3-58.2%) (P<0.0001, Kruskal-Wallis test).

[0322] Nuclease usage level analysis can distinguish between human subjects with SLE and those without SLE.

[0323] Figure 61 shows a set of graphs 6100 identifying nuclease usage levels in plasma cell-free DNA of subjects with and without systemic lupus erythematosus (SLE). Box plot 6102 shows F Profile I (DNASE1L3) levels in plasma cell-free DNA across healthy control subjects, patients with inactive SLE, and patients with active SLE. ROC curve 6104 shows the evaluation of differentiation between patients with and without SLE using F Profile I (DNASE1L3).

[0324] Graph 6106 shows the correlation between the Systemic Lupus Erythematosus Disease Activity Index (SLEDAI) and F-profile I levels (DNASE1L3) in patients with SLE. In a cohort including 10 healthy controls, 13 patients with active and inactive sporadic SLE, and 11 patients with inactive SLE (Chan et al. Proc Natl Acad Sci USA. 2014;111:E5302-E5311), boxplot 6102 shows that DNASE1L3 usage levels gradually decreased across healthy subjects (median: 39.8%, range: 38.0-42.3%), patients with inactive SLE (median: 33.3%, range: 31.4-41.0%), and patients with active SLE (median: 29.7%, range: 14.9-34.2%) (P<0.0001, Kruskal-Wallis test). As shown in ROC curve 6104, the DNASE1L3 usage level metric (F-profile I) allowed for differentiation between human individuals with and without SLE with an AUC of 0.97.

[0325] Additionally, as shown in graph 6106, DNASE1L3 usage levels showed a negative correlation with the Systemic Lupus Erythematosus Disease Activity Index (SLEDAI) (Pearson's r: -0.43; P=0.036). Therefore, a metric of DNASE1L3 usage levels (F-Profile I) may not only signal the presence of autoimmune disease, but also facilitate monitoring of disease progression.

[0326] B. Cancer In addition to SLE, nuclease usage level analysis can be used to distinguish between human subjects with and without hepatocellular carcinoma (HCC). Patients with HCC have been reported to be affected by DNASE1L3 activity (Jiang et al. Cancer Discov. 2020;10:664-73). Regarding the relationship between DNASE1L3 and HCC, nuclease usage level analysis was applied to a cohort of 38 healthy controls, 17 HBV carriers without HCC, and 34 patients with HCC from a previous study (Jiang et al. Cancer Discov. 2020;10:664-73).

[0327] FIG. 62 shows the proportional contribution 6400 of F profiles across normal, HBV, and HCC plasma samples, according to some embodiments. The deconvolution process of FIG. 50 was applied to normal, HBV, and HCC plasma samples from human subjects. As shown in FIG. 62, samples from control subjects contained a relatively high contribution of F profile I. Similarly, samples obtained from HBV patients also contained a relatively high contribution of F profile I. In contrast, HCC samples contained a relatively low contribution from F profile I. The data shown in FIG. 62 suggests that a diagnosis of HCC can be predicted by analyzing the proportional contribution from F profile I.

[0328] Figure 63 shows a set of graphs depicting nuclease usage levels in plasma cell-free DNA of subjects with and without HCC. Boxplots of F profiles I 6302 and VI 6304 show nuclease levels in plasma cell-free DNA of patients with and without HCC. ROC curve 6306 shows evaluation of differentiation between non-HCC and HCC groups using different metrics, including motif diversity scores and six F profiles. Compared to healthy controls, boxplot 6302 shows that usage levels of F profile I (DNASE1L3) were indeed found to be reduced by a median of 6.9% in HCC patients, but no significant changes were observed in HBV carriers.

[0329] As shown in the boxplot 6304, we also found a gradual increase in F-profile VI usage levels in HBV carriers and HCC patients. Additionally, the ROC curve 6306 indicates that among the six F-profiles, F-profile VI (AUC: 0.97) was the most discriminatory in detecting patients with HCC, which appeared to be randomly distributed across the 256 terminal motifs (i.e., no apparent preference for terminal motifs). This performance was superior to the previously reported motif diversity score (AUC: 0.86) (P = 0.019, DeLong test), which was used to quantify the uniformity of overall terminal motif frequency (Jiang et al. Cancer Discov. 2020;10:664-73). These data suggested that nuclease usage level analysis, by simultaneously considering the involvement of multiple nucleases, may have improved the signal-to-noise ratio in disease detection.

[0330] C. Relationship between disease and oxidative stress Because F-profile VI showed promising differentiation between patients with and without HCC, we considered whether any biological significance was associated with F-profile VI. Because of the nature of the F-profile, which shows a lack of clear preference in the frequency of the 256 4-mer motifs, one possible speculation is that cell-free DNA fragmentation occurring in cancer patients may preferentially involve DNA breaks that are different from DNA fragmentation induced by the classical apoptotic pathway.

[0331] Figure 64 shows a bar graph 6040 identifying oxidative stress levels in blood samples from controls and HCC patients. As shown in Figure 64, it has been reported that oxidative stress levels in blood samples with HCC are higher than those in normal controls (Arsian et al. J. Cancer Ther. 2014;5:192-197). Based on the above, it can be considered that F-profile VI is related to the degree of oxidative stress, and as a result, the F-profile VI contribution is significantly increased in patients with HCC (see boxplot 6304 in Figure 63).

[0332] To test the above hypothesis, we utilized a clinical model in which certain tissues have been reported to have higher / lower oxidative stress levels. We first analyzed plasma cell-free DNA from 15 control subjects, 25 colorectal cancer (CRC) patients without liver metastases, and 24 CRC patients with liver metastases.

[0333] Figure 65 shows a set of graphs 6500 providing F-profile analysis and oxidative stress levels in CRC patients. Boxplot 6502 shows F-profile VI contribution in controls, CRC patients with liver metastases, and CRC patients without liver metastases. Bar graph 6504 shows oxidative stress levels in colon tissues from controls and CRC patients at different stages (II-IV). As shown in boxplot 6502, we indeed found a significant increasing trend in F-profile VI levels from controls (median: 24.3%, range 16.8-33.3%) and CRC patients without liver metastases (median: 30.5%, range 23.4-34.5%) to CRC patients with liver metastases (median: 34.5%, range 18.1-43.5%) (Kruskal-Wallis test, P<0.0001). Such findings were consistent with reports that oxidative stress was increased in patients with colorectal cancer and further increased in patients with advanced stage cancer, as shown in bar graph 6504 (Skrzydlewska et al. World J Gastroenterol. 2005;11:403-406;).

[0334] D. Post-disease Treatment In addition to CRC patients, we also analyzed F profiles in plasma DNA from six patients with nasopharyngeal carcinoma (NPC) before and after cisplatin-based chemoradiotherapy.

[0335] Figure 66 shows a boxplot 6600 identifying F-profile VI contributions in NPC patients before and during chemoradiotherapy with cisplatin. Increased oxidative stress has been reported to be further enhanced during chemoradiotherapy (Conklin et al. Integr. Cancer Ther. 2004;3:294-300). As shown in Figure 66, F-profile VI levels in NPC patients increased during chemoradiotherapy with cisplatin (median: 22.1%, range: 18.5-22.8%) compared to paired patients before treatment (median: 23.6%, range: 21.8-25.5%) (P=0.04, Kruskal-Wallis test).

[0336] Classification of the level of pathology using EF profile cfDNA Figure 67 shows a flowchart of a method 6700 for determining a classification of a level of pathology based on an F profile of cell-free DNA, according to some embodiments. Exemplary biological samples can be acellular samples containing cell-free DNA, such as blood, plasma, serum, urine, and saliva. Pathologies can include cancer (e.g., hepatocellular carcinoma, lung cancer, breast cancer, gastric cancer, glioblastoma multiforme, pancreatic cancer, colorectal cancer, nasopharyngeal cancer, head and neck squamous cell carcinoma, etc.), and autoimmune disorders (e.g., systemic lupus erythematosus). Aspects of method 6700 can be performed in a manner similar to method 5400 of Figure 54.

[0337] In block 6702, a set of reference F profiles is stored. Each reference F profile of the set identifies, for each nucleotide in a set of nucleotides, the proportion of cell-free DNA molecules that terminate at that nucleotide. In addition, each reference F profile is associated with a type of fragmentation factor. The type of fragmentation factor identifies a specific enzyme (e.g., DNASE1L3, DNASE1), protein (e.g., DFFB), or other biological component or process that causes fragmentation in cell-free DNA molecules. In some cases, the set of reference F profiles includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 45, 50, or more than 50 F profiles. For example, the set of reference F profiles may include six F profiles I-VI. Block 6702 may be performed in a manner similar to block 5402 of FIG. 54.

[0338] In block 6704, a plurality of cell-free DNA fragments from the biological sample are analyzed to obtain sequence reads. The sequence reads may include terminal sequences corresponding to the ends of the plurality of cell-free DNA molecules. The sequence reads may include terminal sequences corresponding to the ends of the plurality of cell-free DNA fragments. Block 6704 may be performed in a manner similar to block 5404 of FIG. 54.

[0339] In block 6706, a sample end motif profile of the subject is determined by determining the proportion of a plurality of cell-free DNA molecules that terminate at each nucleotide of the set of nucleotides based on the end sequences. The sample end motif profile identifies the relative frequencies of a plurality of end motifs corresponding to the end sequences of the plurality of cell-free DNA molecules. The plurality of end motifs can correspond to all possible combinations of N base positions. For example, if the plurality of end motifs correspond to 4-mers, the plurality of end motifs of the sample end motif profile can include 256 4-mer combinations (e.g., CCCA, TTCC). Block 6706 can be performed in a manner similar to block 5406 of FIG. 54.

[0340] In block 6708, proportional contributions for a set of reference F profiles whose proportional aggregation provides a sample end motif profile are determined. The proportional contributions of the set of reference F profiles sum to 1. Block 6708 may be performed in a manner similar to block 5408 of FIG. 54.

[0341] In block 6710, a classification of a level of pathology may be determined for the subject based on a determination that at least one of the determined proportional contributions exceeds a predetermined threshold. The predetermined threshold may correspond to the proportional contribution of a particular reference F profile (e.g., F profile I, F profile IV). For example, a classification of a level of pathology may be determined for the subject based on a determination that one of the determined proportional contributions is less than a predetermined threshold, as shown in box plot 6302 of FIG. 63. In another example, a classification of a level of pathology may be determined for the subject based on a determination that one of the determined proportional contributions is greater than a predetermined threshold, as shown in box plot 6304 of FIG. 63.

[0342] The level of pathology may include non-cancerous, early stage, intermediate stage, or advanced stage. The classification may then select one of the levels. Thus, the classification may be determined from multiple levels of cancer, including multiple stages of cancer. For example, the cancer may be hepatocellular carcinoma, lung cancer, breast cancer, gastric cancer, glioblastoma multiforme, pancreatic cancer, colorectal cancer, nasopharyngeal carcinoma, and head and neck squamous cell carcinoma. For example, the autoimmune disorder may be systemic lupus erythematosus.

[0343] In a further example, the level of pathology corresponds to the fractional concentration of clinically relevant DNA associated with the pathology. For example, the level of pathology can be cancer, and the clinically relevant DNA can be tumor DNA. The reference value can be a calibration value determined from a calibration sample.

[0344] XII. Treatment A. Treatment options Embodiments of the present disclosure can accurately predict disease recurrence, thereby facilitating early intervention and selection of appropriate treatment, improving a subject's disease outcome and overall survival. For example, in cases where the corresponding sample predicts disease recurrence, intensified chemotherapy can be selected for the subject. In another example, biological samples from a subject who has completed initial treatment can be sequenced to identify viral DNA that predicts disease recurrence. In such an example, the subject's cancer may be resistant to the initial treatment, and an alternative treatment plan (e.g., a higher dose) and / or a different treatment can be selected for the subject.

[0345] Embodiments may also include treating the subject in response to determining a classification of pathology recurrence. For example, if the prediction corresponds to locoregional failure, surgery may be selected as a possible treatment. In another example, if the prediction corresponds to distant metastasis, chemotherapy may additionally be selected as a possible treatment. In some embodiments, the treatment includes surgery, radiation therapy, chemotherapy, immunotherapy, targeted therapy, hormone therapy, stem cell transplantation, or precision medicine. Based on the determined classification of recurrence, a treatment plan may be developed to reduce the risk of harm to the subject and increase overall survival. Embodiments may further include treating the subject according to the treatment plan.

[0346] B. Types of Treatment Embodiments may further include treating the pathology in the patient after determining the classification for the subject. Treatment can be provided according to the determined level of pathology, the fractional concentration of clinically relevant DNA, or the tissue of origin. For example, identified mutations can be targeted with specific drugs or chemotherapy. The tissue of origin can be used to guide surgery or any other form of treatment. The level of pathology can then be used to determine how aggressive any type of treatment should be, which can also be determined based on the level of pathology. Pathology (e.g., cancer) can be treated with chemotherapy, drugs, diet, therapy, and / or surgery. In some embodiments, the higher the value of a parameter (e.g., amount or size) exceeds a reference value, the more aggressive the treatment can be.

[0347] Treatment may include resection. In the case of bladder cancer, treatment may include transurethral resection of bladder tumor (TURBT). This procedure is used for diagnosis, staging, and treatment. During TURBT, a surgeon inserts a cystoscope through the urethra into the bladder. The tumor is then removed using tools with small wire loops, lasers, or high-energy electricity. For patients with non-muscle-invasive bladder cancer (NMIBC), TURBT may be used to treat or eliminate the cancer. Another treatment may include radical cystectomy and lymph node dissection. Radical cystectomy is the removal of the entire bladder and possibly surrounding tissues and organs. Treatment may also include urinary diversion. Urinary diversion is when a doctor creates a new pathway for urine to leave the body if the bladder is removed as part of treatment.

[0348] Treatment may include chemotherapy, which is the use of drugs to destroy cancer cells, usually by preventing their growth and division. Drugs may include, but are not limited to, for example, mitomycin-C (available as a generic drug), gemcitabine (Gemzar), and thiotepa (Tepadina) for intravesical chemotherapy. Systemic chemotherapy may include, but is not limited to, for example, cisplatin gemcitabine, methotrexate (Rheumatrex, Trexall), vinblastine (Velban), doxorubicin, and cisplatin.

[0349] In some embodiments, the treatment may include immunotherapy. The immunotherapy may include immune checkpoint inhibitors that block a protein called PD-1. Inhibitors may include, but are not limited to, atezolizumab (Tecentriq), nivolumab (Opdivo), avelumab (Bavencio), durvalumab (Imfinzi), and pembrolizumab (Keytruda).

[0350] Therapeutic embodiments may also include targeted therapy, which is a treatment that targets specific genes and / or proteins in cancer that contribute to cancer growth and survival. For example, erdafitinib is an orally administered drug approved to treat people with locally advanced or metastatic urothelial carcinoma with FGFR3 or FGFR2 gene mutations, which cause cancer cells to continue to grow or spread.

[0351] Some treatments may include radiation therapy. Radiation therapy is the use of high-energy X-rays or other particles to destroy cancer cells. In addition to each individual treatment, a combination of these treatments described herein may be used. In some embodiments, a combination of treatments may be used when the parameter value exceeds a threshold value, and the threshold value itself exceeds a reference value. Information regarding treatments in references is incorporated herein by reference.

[0352] XIII. Exemplary Systems FIG. 68 illustrates a measurement system 6800 according to one embodiment of the present disclosure. The system as shown includes a sample 6805, such as cell-free DNA molecules, in an assay device 6810, and an assay 6808 can be performed on the sample 6805. For example, the sample 6805 can be contacted with reagents for the assay 6808 to provide a signal of a physical characteristic 6815. An example of an assay device can be a flow cell containing assay probes and / or primers, or a tube through which droplets (along with droplets containing the assay) travel. The physical characteristic 6815 (e.g., fluorescence intensity, voltage, or current) from the sample is detected by a detector 6820. The detector 6820 can take measurements at intervals (e.g., periodic intervals) to obtain data points that constitute a data signal. In one embodiment, an analog-to-digital converter converts the analog signal from the detector to digital form at multiple times. The assay device 6810 and the detector 6820 can form an assay system, such as a sequencing system, that performs sequencing according to embodiments described herein. A data signal 6825 is transmitted from detector 6820 to logic system 6830. As an example, data signal 6825 can be used to determine the sequence and / or location in a reference genome of a DNA molecule. Data signal 6825 can include various measurements performed simultaneously, e.g., different colors of fluorescent dyes or different electrical signals for different molecules of sample 6805, and thus data signal 6825 can correspond to multiple signals. Data signal 6825 can be stored in local memory 6835, external memory 6840, or storage device 6845.

[0353] Logic system 6830 may be or include a computer system, ASIC, microprocessor, graphics processing unit (GPU), etc. It may also include or be coupled to a display (e.g., a monitor, LED display, etc.) and user input devices (e.g., a mouse, keyboard, buttons, etc.). Logic system 6830 and other components may be part of a standalone or networked computer system, or they may be directly attached to or incorporated into a device (e.g., a sequencing device) including detector 6820 and / or assay device 6810. Logic system 6830 may also include software executing on processor 6850. Logic system 6830 may include a computer-readable medium that stores instructions for controlling measurement system 6800 to perform any of the methods described herein. For example, logic system 6830 can provide commands to a system including assay device 6810 to perform sequencing or other physical operations. Such physical operations may be performed in a particular order, e.g., reagents are added and removed in a particular order. Such physical actions can be performed by a robotic system, including, for example, a robotic arm, such that the system can be used to obtain samples and perform assays.

[0354] System 6800 may also include a therapy device 6860 that can provide therapy to a subject. The therapy device 6860 can be used to determine and / or administer a therapy. Examples of such therapy may include surgery, radiation therapy, chemotherapy, immunotherapy, targeted therapy, hormone therapy, and stem cell transplant. Logic system 6830 may be connected to therapy device 6860, for example, to provide results of the methods described herein. The therapy device may receive input from other devices, such as imaging devices, and user input (e.g., to control therapy, such as control for a robotic system).

[0355] Any of the computer systems referred to herein may utilize any suitable number of subsystems. An example of such a subsystem is shown in FIG. 69, computer system 10. In some embodiments, the computer system includes a single computer device, and the subsystems may be components of the computer device. In other embodiments, the computer system may include multiple devices, each of which is a subsystem, along with its internal components. Computer systems may include desktop and laptop computers, tablets, mobile phones, and other mobile devices.

[0356] The subsystems shown in FIG. 69 are interconnected via a system bus 75. Additional subsystems are shown, such as a printer 74, a keyboard 78, a storage device 79, a monitor 76 (e.g., a display screen such as an LED) coupled to a display adapter 82, and others. Peripherals and I / O devices coupled to an input / output (I / O) controller 71 can be connected to the computer system by any number of means known in the art, such as input / output (I / O) ports 77 (e.g., USB, FireWire®). For example, the I / O ports 77 or an external interface 81 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect the computer system 10 to a wide area network, such as the Internet, a mouse input device, or a scanner. The interconnection via the system bus 75 allows the central processor 73 to communicate with each subsystem and control the execution of instructions from the system memory 72 or storage device 79 (e.g., a fixed disk such as a hard drive or an optical disk), as well as the exchange of information between the subsystems. The system memory 72 and / or storage device 79 may embody computer-readable media. Another subsystem is a data collection device 85, such as a camera, microphone, accelerometer, etc. Any of the data mentioned herein can be output from one component to another and can be output to a user.

[0357] A computer system may include multiple identical components or subsystems connected together, for example, by an external interface 81, by an internal interface, or through a removable storage device that can be connected and disconnected from one component to another. In some embodiments, computer systems, subsystems, or devices may communicate over a network. In such cases, one computer may be considered a client and another computer may be considered a server, each of which may be part of the same computer system. The client and server may each include multiple systems, subsystems, or components.

[0358] Aspects of the embodiments can be implemented in the form of control logic using hardware circuitry (e.g., application-specific integrated circuits or field-programmable gate arrays) and / or using computer software stored in memory with a processor that is generally programmable in a modular or integrated manner; thus, a processor can include a memory that stores software instructions that configure the hardware circuitry, and an FPGA or ASIC with the configuration instructions. As used herein, a processor can include a single-core processor, a multi-core processor on the same integrated chip, or multiple processing units on a single circuit board or networked, as well as dedicated hardware. Based on the disclosure and teachings provided herein, those skilled in the art will recognize and understand other means and / or methods of implementing embodiments of the present disclosure using hardware and combinations of hardware and software.

[0359] Any of the software components or functions described in this application may be implemented as software code to be executed by a processor using any suitable computer language, such as, for example, Java, C, C++, C#, Objective-C, Swift, or a scripting language such as Perl or Python, using, for example, conventional or object-oriented techniques. The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission. Suitable non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), magnetic media such as a hard drive or floppy disk, or optical media such as a compact disc (CD) or DVD (digital versatile disc) or Blu-ray disc, flash memory, etc. The computer-readable medium may be any combination of such devices. Additionally, the order of operations may be re-arranged. A process may terminate when its operation is completed, but may have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. If the process corresponds to a function, its termination may correspond to the invocation of the function or the return of the function to the main function.

[0360] Such programs may also be encoded and transmitted using carrier signals adapted for transmission over wired, optical, and / or wireless networks conforming to various protocols, including the Internet. Thus, computer-readable media may be created using data signals encoded with such programs. Computer-readable media encoded with program code may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer-readable medium may reside on or within a single computer product (e.g., a hard drive, CD, or an entire computer system) and may be present on or within different computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing a user with any of the results mentioned herein.

[0361] Any of the methods described herein may be implemented, in whole or in part, using a computer system including one or more processors, which may be configured to perform steps. Any operation (e.g., aligning, determining, comparing, calculating, computing) performed by a processor may be performed in real time. The term "real time" may refer to a computing operation or process that is completed within a certain time constraint. The time constraint may be one minute, one hour, one day, or seven days. Accordingly, embodiments may be directed to a computer system configured to perform steps of any of the methods described herein, potentially with different components performing each step or each group of steps. Although presented as numbered steps, steps of the methods herein may be performed simultaneously, at different times, or in different orders. Additionally, portions of these steps may be used with portions of other steps from other methods. Also, all or portions of the steps may be optional. Additionally, any of the steps of any of the methods may be performed using a system module, unit, circuit, or other means for performing these steps.

[0362] The specific details of particular embodiments may be combined in any suitable manner without departing from the spirit and scope of the embodiments of the present disclosure, however, other embodiments of the present disclosure may be directed to particular embodiments relating to each individual aspect, or particular combinations of these individual aspects.

[0363] The above description of example embodiments of the present disclosure has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the disclosure to the precise form described, and many modifications and variations are possible in light of the above teachings.

[0364] The use of "a," "an," or "the" is intended to mean "one or more" unless specifically stated to the contrary. The use of "or" is intended to mean "inclusive or," and not "exclusive or," unless specifically stated to the contrary. A reference to a "first" element does not necessarily require that a second element be provided. Furthermore, a reference to a "first" or "second" element does not limit the referenced element to a particular location unless expressly stated. The term "based on" is intended to mean "based at least in part on."

[0365] The claims may be drafted to exclude any element that may be optional. Accordingly, this statement is intended to serve as a predicate for use of exclusive terminology such as "solely," "only," or the like in connection with the recitation of claim elements or for use of a "negative" limitation.

[0366] It is understood that the present invention is not limited to the particular embodiments described, as such may, of course, vary. It is also understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention is limited only by the appended claims. Efforts have been made to ensure accuracy with respect to numbers used (e.g., amounts, temperatures, etc.), but some experimental error and deviation should be accounted for. Unless otherwise specified, parts are parts by weight, molecular weight is weight average molecular weight, temperature is in degrees Celsius, and pressure is at or near atmospheric.

[0367] All patents, patent applications, publications, and descriptions referred to herein are incorporated by reference in their entirety for all purposes. None is admitted to be prior art. In the event of a conflict between this application and the references provided herein, this application shall control.

Claims

1. 1. A method for estimating the fractional concentration of clinically relevant DNA molecules in a urine sample of a subject, the urine sample comprising the clinically relevant DNA molecules and other DNA molecules that are cell-free, the method comprising: analyzing a plurality of cell-free DNA molecules from the urine sample, wherein analyzing the plurality of cell-free DNA molecules comprises: determining the locations of the plurality of cell-free DNA molecules; and identifying a set of cell-free DNA molecules derived from open chromatin regions of one or more tissues associated with the clinically relevant DNA molecules based on the locations of the plurality of cell-free DNA molecules. using the set of cell-free DNA molecules to determine the relative abundance of the set of cell-free DNA molecules that are derived from open chromatin regions of the one or more tissues; and estimating the fractional concentration of the clinically relevant DNA molecule in the urine sample by comparing the relative abundance to one or more calibration values ​​determined from one or more calibration samples in which the fractional concentrations of the clinically relevant DNA molecule are known.

2. 2. The method of claim 1, wherein comparing the relative abundance to the one or more calibration values ​​comprises comparing the relative abundance to a calibration curve comprising the one or more calibration values.

3. For each calibration sample of the one or more calibration samples: determining the fractional concentration of the clinically relevant DNA molecules in the calibration sample; 3. The method of claim 1 or 2, further comprising measuring the relative abundance of cell-free DNA molecules from the calibration samples derived from the open chromatin regions of the one or more tissues.

4. 4. The method of claim 3, wherein measuring the fractional concentration of the clinically relevant DNA molecules uses tissue-specific alleles or tissue-specific methylation patterns.

5. 1. A method of enriching a urine sample for clinically relevant DNA molecules, said urine sample containing said clinically relevant DNA molecules and other DNA molecules that are cell-free, said method comprising: analyzing a plurality of cell-free DNA molecules from the urine sample, wherein analyzing the plurality of cell-free DNA molecules comprises: analyzing, including identifying from said plurality of cell-free DNA molecules a set of cell-free DNA molecules that are derived from open chromatin regions of one or more tissues associated with said clinically relevant DNA molecules; and creating an enriched sample using the set of cell-free DNA molecules derived from the open chromatin regions of the one or more tissues, wherein the enriched sample has a higher concentration of clinically relevant DNA compared to the urine sample.

6. 6. The method of claim 5, further comprising determining a characteristic associated with the clinically relevant DNA molecules in the enriched sample, wherein the characteristic associated with the clinically relevant DNA molecules in the urine sample is (1) the fractional concentration of the clinically relevant DNA molecules, or (2) the level of pathology of the subject from which the urine sample was obtained, which is associated with the clinically relevant DNA molecules.

7. 7. The method of any one of claims 5-6, wherein creating the enriched sample further comprises using a set of the cell-free DNA molecules derived from the open chromatin regions of the one or more tissues and having a size that is below a specified size threshold.

8. 8. The method of claim 7, wherein the specified size threshold is 40 base pairs, 50 base pairs, 60 base pairs, 70 base pairs, 80 base pairs, 90 base pairs, 100 base pairs, 110 base pairs, 120 base pairs, 130 base pairs, 140 base pairs, 150 base pairs, or 160 base pairs.

9. 7. The method of any one of claims 5-6, wherein creating the enriched sample further comprises using the set of cell-free DNA molecules derived from the open chromatin regions of the one or more tissues and having one or more end sequences corresponding to a sequence end signature.

10. identifying the set of cell-free DNA molecules or creating the enriched sample; 7. The method of claim 5, comprising subjecting the plurality of cell-free DNA molecules to probe molecules having sequences derived from the open chromatin region, thereby obtaining the set of cell-free DNA molecules.

11. creating the enriched sample, 11. The method of claim 10, comprising using the one or more probe molecules to amplify the set of cell-free DNA molecules.

12. creating the set of cell-free DNA molecules capturing the set of cell-free DNA molecules using the one or more probe molecules; and discarding other cell-free DNA molecules of the plurality of cell-free DNA molecules.

13. The method of claim 10, wherein one or more probe molecules are attached to a surface and detect a set of one or more sequence motifs in the terminal sequence by hybridization.

14. 1. A method of enriching a urine sample for clinically relevant DNA, wherein the urine sample contains the clinically relevant DNA and other DNA that is cell-free, the method comprising: analyzing a plurality of cell-free DNA molecules from the urine sample, wherein analyzing the plurality of cell-free DNA molecules comprises: analyzing, including identifying from the plurality of cell-free DNA molecules a set of cell-free DNA molecules having terminal sequences that fall within a set of one or more sequence motifs that include a C-terminal nucleotide; and creating an enriched sample using the set of cell-free DNA molecules having end sequences within the set of one or more sequence motifs, wherein the enriched sample has a higher concentration of clinically relevant DNA compared to the urine sample.

15. 15. The method of claim 14, wherein creating the enriched sample further comprises using a set of the cell-free DNA molecules that are within the set of one or more sequence motifs that include the C-terminal nucleotide and have a size that is below a specified size threshold.

16. 16. The method of claim 15, wherein the specified size threshold is 40 base pairs, 50 base pairs, 60 base pairs, 70 base pairs, 80 base pairs, 90 base pairs, 100 base pairs, 110 base pairs, 120 base pairs, 130 base pairs, 140 base pairs, 150 base pairs, or 160 base pairs.

17. identifying the set of cell-free DNA molecules or creating the enriched sample; 17. The method of any one of claims 14 to 16, comprising subjecting the plurality of cell-free DNA molecules to one or more probe molecules that detect the set of one or more sequence motifs in the terminal sequences of the plurality of cell-free DNA molecules, thereby obtaining the set of cell-free DNA molecules.

18. 18. The method of claim 17, further comprising attaching a common sequence to the plurality of cell-free DNA molecules, wherein the one or more probe molecules comprise a sequence complementary to the common sequence.

19. creating the enriched sample, 20. The method of claim 17, comprising using the one or more probe molecules to amplify the set of cell-free DNA molecules.

20. creating the set of cell-free DNA molecules capturing the set of cell-free DNA molecules using the one or more probe molecules; and discarding other cell-free DNA molecules of the plurality of cell-free DNA molecules.

21. 18. The method of claim 17, wherein one or more probe molecules are attached to a surface and detect the set of one or more sequence motifs in the terminal sequence by hybridization.

22. 22. The method of any one of claims 14-21, wherein analyzing the plurality of cell-free DNA molecules comprises receiving sequence reads obtained from sequencing the plurality of cell-free DNA molecules, and wherein identifying the set of the plurality of cell-free DNA molecules comprises identifying sequence reads having the end sequences that are within the set of one or more sequence motifs.

23. 23. The method of Claim 22, wherein the enriched samples correspond to the sequence reads having the end sequences that fall within the set of one or more sequence motifs.

24. 20. The method of any one of claims 1, 5, or 17, wherein the clinically relevant DNA molecule is a transrenal DNA molecule.

25. 20. The method of any one of claims 1, 5, or 17, wherein the clinically relevant DNA molecule comprises fetal DNA or tumor DNA.

26. 1. A method for detecting a kidney abnormality using a urine sample from a subject, the urine sample comprising cell-free DNA molecules, the method comprising: analyzing a plurality of cell-free DNA molecules from the urine sample, wherein analyzing the plurality of cell-free DNA molecules comprises: determining the location of the plurality of cell-free DNA molecules; identifying a set of cell-free DNA molecules that are derived from open chromatin regions of one or more tissues based on the locations of the plurality of cell-free DNA molecules; determining the relative abundance of the set of cell-free DNA molecules derived from the open chromatin region of the one or more tissues compared to other cell-free DNA molecules of the plurality of cell-free DNA molecules; comparing the relative abundance to a reference value; and determining a classification of the subject having the kidney abnormality based on the comparing.

27. 27. The method of claim 26, wherein the reference value corresponds to another relative abundance determined based on cell-free DNA molecules derived from open chromatin regions of one or more reference samples, the one or more reference samples being associated with a known classification of the kidney abnormality.

28. 27. The method of claim 26, wherein determining the classification further comprises applying a machine learning model to the relative abundances to generate an output indicative of the classification of the subject having the kidney abnormality, wherein the machine learning model has been trained using a training dataset determined from one or more training samples having known classifications of the kidney abnormality.

29. 27. The method of claim 26, wherein the set of cell-free DNA molecules terminate at one or more positions within a window around the open chromatin region of the one or more tissues.

30. determining the relative abundance determining a first relative frequency of the set of cell-free DNA molecules derived from the open chromatin regions of the one or more tissues; determining a second relative frequency of reference sequences in a reference genome that are derived from the open chromatin regions of the one or more tissues; and determining the relative abundance based on the first relative frequency and the second relative frequency.

31. 31. The method of Claim 30, wherein determining the second relative frequencies comprises identifying single base variants in the reference genome that are derived from the open chromatin regions of the one or more tissues.

32. 31. The method of claim 30, wherein the relative abundance is a ratio of the first relative frequency to the second relative frequency.

33. 27. The method of claim 1 or claim 26, wherein the relative abundance is the end density of the plurality of cell-free DNA molecules that terminate in the open chromatin regions of the one or more tissues.

34. 34. The method of Claim 33, wherein the end density comprises a first amount of the set of cell-free DNA molecules derived from the open chromatin region of the one or more tissues divided by a second amount of the plurality of cell-free DNA molecules derived from one or more other regions.

35. 27. The method of Claim 1 or Claim 26, wherein determining the locations of the plurality of cell-free DNA molecules comprises aligning sequence reads of the plurality of cell-free DNA molecules to a reference genome.

36. 27. The method of any one of claims 1, 5, and 26, wherein the one or more tissues include at least one of heart, lung, colon, liver, or white blood cells.

37. 27. The method of any one of claims 1, 5, and 26, wherein the set of cell-free DNA molecules terminate at one or more positions within a window around the open chromatin region of the one or more tissues.

38. 27. The method of any one of claims 1, 5, and 26, wherein the open chromatin region comprises a Dnase 1 hypersensitive site.

39. 39. The method of any one of claims 1 to 38, wherein the set of cell-free DNA molecules comprises at least 5,000 cell-free DNA molecules.

40. 1. A method for detecting a kidney abnormality using a urine sample from a subject, the urine sample comprising cell-free DNA molecules, the method comprising: analyzing a plurality of cell-free DNA molecules from the urine sample, wherein analyzing the plurality of cell-free DNA molecules comprises determining a size of the plurality of cell-free DNA molecules; determining a statistical value using the sizes of the plurality of cell-free DNA molecules; and comparing the statistical value to a reference value; and determining a classification of the subject having the kidney abnormality based on the comparing.

41. determining the statistical value identifying a set of cell-free DNA molecules from the plurality of cell-free DNA molecules having sizes within a size range; and determining the statistical value further based on the amount of the set of cell-free DNA molecules.

42. 42. The method of claim 41 , wherein the statistical value is a ratio of the set of cell-free DNA molecules to the plurality of cell-free DNA molecules from the urine sample.

43. 43. The method of claim 41 or claim 42, wherein the size range has an upper limit selected from one of at least 80 bases, at least 90 bases, at least 100 bases, at least 110 bases, or at least 120 bases.

44. The method of any one of claims 40 to 43, wherein the reference value is determined using one or more reference samples with a known classification of the kidney abnormality.

45. 45. The method of any one of claims 40-44, wherein determining the size of the plurality of cell-free DNA molecules comprises using gel electrophoresis, filtration, size-selective precipitation, or hybridization.

46. measuring the size of the plurality of cell-free DNA molecules; receiving sequence reads obtained from sequencing the plurality of cell-free DNA molecules from the urine sample; and and for each of the sequence reads, counting the number of nucleotides in the sequence read.

47. 47. The method of claim 46, wherein said sequencing of said plurality of cell-free DNA molecules comprises performing massively parallel sequencing, single molecule real-time sequencing, or nanopore sequencing of said plurality of cell-free DNA molecules.

48. 48. The method of any one of claims 26 to 47, wherein the renal abnormality is preeclampsia or proteinuria.

49. 48. The method of any one of claims 26 to 47, wherein the classification of the subject as having the kidney abnormality comprises an increased level of permeability associated with the glomerular basement membrane of the kidney.

50. 50. The method of any one of claims 1-49, wherein the urine sample has been treated with a DNA stabilizing agent prior to obtaining the plurality of cell-free DNA molecules from the urine sample.

51. 51. The method of any one of claims 1-50, wherein analyzing the plurality of cell-free DNA molecules comprises receiving sequence reads obtained from sequencing the plurality of cell-free DNA molecules.

52. 1. A method for detecting a kidney abnormality using a urine sample from a subject, the urine sample comprising cell-free DNA molecules, the method comprising: determining a first amount of cell-free DNA molecules in the urine sample; determining an initial concentration using the first amount and the volume of the urine sample; determining a corrected concentration using a second amount of the specific compound in the urine sample; and comparing the corrected concentration to a reference value; and determining a classification of the subject having the kidney abnormality based on the comparing.

53. 53. The method of claim 52, wherein determining the first amount comprises performing a measurement using a fluorometer, a spectrophotometer, or PCR.

54. 53. The method of claim 52, wherein the initial concentration is the percentage of cell-free DNA molecules in the urine sample that exceed a size cutoff.

55. 55. The method of claim 54, wherein the size cutoff is 40 to 200 bp.

56. 53. The method of claim 52, wherein the specific compound is creatinine.

57. 53. The method of claim 52, wherein the renal abnormality is preeclampsia or proteinuria.

58. 1. A method for determining the level of nuclease activity in a biological sample of a subject, said method comprising: storing a set of reference F profiles, each reference F profile of the set comprising: for each nucleotide of the set of nucleotides, determining the proportion of cell-free DNA molecules that terminate at that nucleotide; Memorization, which is related to a type of fragmentation factor, analyzing a plurality of cell-free DNA molecules from the biological sample to obtain sequence reads, wherein the sequence reads comprise terminal sequences corresponding to ends of the plurality of cell-free DNA molecules; determining a sample end motif profile by determining a proportion of the plurality of cell-free DNA molecules that terminate at each nucleotide of the set of nucleotides based on the end sequences; determining the percentage contributions of the set of reference F profiles, the aggregation of which percentages provides the sample end motif profile, and the percentage contributions sum to one; and determining a classification of nuclease activity of a particular type of nuclease based on a proportion contribution associated with a reference F profile among the set of reference F profiles.

59. determining the classification of nuclease activity 59. The method of claim 58, comprising determining a decrease in the level of nuclease activity associated with the particular type of nuclease.

60. Determining the decrease in the level of nuclease activity comparing a fractional contribution associated with one of the set of reference F profiles to a cutoff value; and determining a decrease in the level of the nuclease activity based on said comparing.

61. 61. The method of claim 60, wherein the cutoff value is determined using one or more reference samples with known classification of the nuclease activity.

62. 62. The method of any one of claims 58 to 61, further comprising determining a classification of the subject's genetic disorder based on the classification of nuclease activity of the specific type of nuclease.

63. 63. The method of any one of claims 58 to 62, wherein the specific type of nuclease comprises one of DNASE1, DNASE1L3, DFFB, TREX1 (3-prime repair exonuclease 1), AEN (apoptosis-enhancing nuclease), EXO1 (exonuclease 1), DNASE2 (deoxyribonuclease 2), ENDOG (endonuclease G), APEX1 (apurinic / apyrimidinic site endodenoxyribonuclease 1), FEN1 (flap structure-specific endonuclease 1), DNASE1L1 (deoxyribonuclease 1-like 1), DNASE1L2 (deoxyribonuclease 1-like 2), and EXOG (exo / endonuclease G).

64. 1. A method for estimating the fractional concentration of a clinically relevant DNA molecule in a biological sample of a subject, the biological sample comprising the clinically relevant DNA molecule and other DNA molecules that are acellular, the method comprising: storing a set of reference F profiles, each reference F profile of the set comprising: for each nucleotide of the set of nucleotides, determining the proportion of cell-free DNA molecules that terminate at that nucleotide; Memorization, which is related to a type of fragmentation factor, analyzing a plurality of cell-free DNA molecules from the biological sample to obtain sequence reads, wherein the sequence reads comprise terminal sequences corresponding to ends of the plurality of cell-free DNA molecules; determining a sample end motif profile by determining a proportion of the plurality of cell-free DNA molecules that terminate at each nucleotide of the set of nucleotides based on the end sequences; determining fractional contributions for the set of reference F profiles, the aggregation of which fractions provides the sample end motif profile, the set of reference F profiles including a first reference F profile having a first fractional contribution, the fractional contributions summing to one; and estimating the fractional concentration of the clinically relevant DNA molecule in the biological sample by comparing the first fraction contribution to one or more calibration values ​​determined from one or more calibration samples in which the fractional concentrations of the clinically relevant DNA molecule are known.

65. 65. The method of claim 64, wherein the clinically relevant DNA molecule comprises fetal DNA or tumor DNA.

66. 71. The method of claim 64 or 70, wherein comparing the first fractional contribution to the one or more calibration values ​​comprises comparing the first fractional contribution to a calibration curve comprising the one or more calibration values.

67. For each calibration sample of the one or more calibration samples: determining the fractional concentration of the clinically relevant DNA molecules in the calibration sample; 67. The method of any one of claims 64 to 66, further comprising measuring the proportion contribution of a first reference terminal motif profile determined for said calibration sample, thereby determining said one or more calibration values.

68. 68. The method of claim 67, wherein measuring the fractional concentration of the clinically relevant DNA molecule uses tissue-specific alleles or tissue-specific methylation patterns.

69. 78. The method of any one of claims 64 to 77, wherein estimating the fractional concentration of the clinically relevant DNA molecules in the biological sample comprises comparing one or more additional fractional contributions to one or more additional calibration values ​​determined from the one or more calibration samples in which the fractional concentrations of the clinically relevant DNA molecules are known.

70. 1. A method for analyzing a biological sample from a female subject carrying a fetus, wherein the biological sample comprises cell-free DNA molecules from the female subject and the fetus, the method comprising: storing a set of reference F profiles, each reference F profile of the set comprising: for each nucleotide of the set of nucleotides, determining the proportion of cell-free DNA molecules that terminate at that nucleotide; Memorization, which is related to a type of fragmentation factor, analyzing a plurality of cell-free DNA molecules from the biological sample to obtain sequence reads, wherein the sequence reads comprise terminal sequences corresponding to ends of the plurality of cell-free DNA molecules; determining a sample end motif profile by determining a proportion of the plurality of cell-free DNA molecules that terminate at each nucleotide of the set of nucleotides based on the end sequences; determining fractional contributions for the set of reference F profiles, the aggregation of which fractions provides the sample end motif profile, the set of reference F profiles including a first reference F profile having a first fractional contribution, the fractional contributions summing to one; and estimating the gestational age of the fetus by comparing the first fraction contribution to one or more calibration values ​​determined from one or more calibration samples of known gestational age.

71. 71. The method of claim 70, wherein comparing the first percentage contribution to the one or more calibration values ​​comprises comparing the first percentage contribution to a calibration curve comprising the one or more calibration values.

72. For each calibration sample of the one or more calibration samples: estimating the gestational age in the calibration sample; and 72. The method of claim 70 or claim 71, further comprising measuring the proportion contribution of a calibration terminal motif profile determined for the calibration sample.

73. 73. The method of any one of claims 70-72, further comprising identifying the plurality of cell-free DNA molecules as derived from the fetus.

74. 74. The method of claim 73, wherein the plurality of cell-free DNA molecules derived from the fetus are identified using fetal-specific alleles or fetal-specific epigenetic markers.

75. 75. The method of any one of claims 70 to 74, wherein estimating the gestational age comprises comparing one or more additional fractional contributions to one or more additional calibration values ​​determined from the one or more calibration samples of known gestational age.

76. 71. The method of any one of claims 64 and 70, wherein the first reference F profile corresponds to a particular type of nuclease.

77. 77. The method of claim 76, wherein the specific type of nuclease is DNASE1L3.

78. 1. A method of classifying a level of pathology in a biological sample of a subject, wherein the biological sample comprises cell-free DNA, the method comprising: storing a set of reference F profiles, each reference F profile of the set comprising: for each nucleotide of the set of nucleotides, determining the proportion of cell-free DNA molecules that terminate at that nucleotide; Memorization, which is related to a type of fragmentation factor, analyzing a plurality of cell-free DNA molecules from the biological sample to obtain sequence reads, wherein the sequence reads comprise terminal sequences corresponding to ends of the plurality of cell-free DNA molecules; determining a sample end motif profile by determining a proportion of the plurality of cell-free DNA molecules that terminate at each nucleotide of the set of nucleotides based on the end sequences, thereby determining a proportion; determining fractional contributions for the set of reference F profiles, the aggregation of which fractions provides the sample end motif profile, the fractional contributions summing to one; and determining a classification of the level of the pathology for the subject based on a determination that at least one of the percentage contributions exceeds a predetermined threshold.

79. 79. The method of claim 78, wherein the predetermined threshold is determined using one or more reference samples having a known classification of the level of the pathology.

80. 80. The method of claim 78 or claim 79, wherein the pathology is cancer or an autoimmune disorder.

81. 81. The method of claim 80, wherein the cancer is hepatocellular carcinoma, lung cancer, breast cancer, gastric cancer, glioblastoma multiforme, pancreatic cancer, colorectal cancer, nasopharyngeal cancer, and head and neck squamous cell carcinoma.

82. 81. The method of claim 80, wherein the classification is determined from multiple levels of cancer, including multiple stages of cancer.

83. 83. The method of any one of claims 58 to 82, wherein the biological sample is a plasma or urine sample.

84. 83. The method of any one of claims 58-82, wherein each reference F profile of the set of reference F profiles specifies a fraction of the cell-free DNA molecules that terminate at each Kmer terminal motif of a set of Kmer terminal motifs, and K is greater than or equal to 2.

85. 84. The method of any one of claims 58 to 83, wherein the set of reference F profiles is determined using cell-free DNA molecules obtained from one or more reference samples.

86. analyzing the plurality of cell-free DNA molecules from the biological sample, 86. The method of any one of claims 58-85, comprising receiving sequence reads obtained from sequencing the plurality of cell-free DNA molecules from the biological sample.

87. 87. The method of Claim 86, wherein said sequencing of said plurality of cell-free DNA molecules comprises performing massively parallel sequencing, single molecule real-time sequencing, or nanopore sequencing of said plurality of cell-free DNA molecules.

88. Determining the fractional contribution for the set of reference F profiles comprises:

88. The method of any one of claims 58 to 87, comprising applying deconvolution to the sample end motif profile of the subject to determine the fractional contribution for the set of reference F profiles.

89. applying the deconvolution for a data matrix M of dimension W×F, where (i) M represents the end frequencies across end motifs of the sample end motif profiles, where the end frequencies of the data matrix M are determined based on the proportions of the plurality of cell-free DNA molecules in the sample end motif profiles; (ii) F represents the end frequencies of the set of reference F profiles obtained from mouse samples, where the end frequencies of F are determined based on the proportions of the cell-free DNA molecules in the set of reference F profiles; and (iii) W represents a relative weight corresponding to the proportion contribution of each of the set of reference F profiles; 89. The method of claim 88, comprising determining the fractional contributions for the set of reference F profiles by solving for the relative weights of W from the data matrix M and the set of reference F profiles.

90. 90. The method of any one of claims 58-89, wherein analyzing the plurality of cell-free DNA molecules comprises receiving sequence reads obtained from sequencing the plurality of cell-free DNA molecules.

91. 91. The method of any one of claims 1 to 90, wherein the plurality of cell-free DNA molecules comprises at least 5,000 cell-free DNA molecules.

92. A computer product comprising a non-transitory computer readable medium storing a plurality of instructions that, when executed, cause a computer system to perform a method according to any one of the preceding claims.

93. 93. A computer product according to claim 92; one or more processors for executing instructions stored on the computer-readable medium.

94. A system comprising means for performing any of the above methods.

95. A system comprising one or more processors configured to perform any of the above methods.