Epigenetic analysis of cell-free DNA
By analyzing sequence motifs and sizes of cell-free DNA fragments, the method addresses the inefficiencies of conventional histone modification detection, enabling accurate tissue classification and disease detection with simpler and more efficient techniques.
Patent Information
- Application Number
- JP2025505427
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-29
- Filing Date
- 2023-07-31
- Publication Date
- 2025-08-26
AI Technical Summary
Conventional techniques for detecting histone modifications in cell-free DNA require large sample volumes and complex sample handling, making them inefficient and impractical for analyzing epigenomic states.
Measuring the relative frequency of sequence motifs and sizes of cell-free DNA fragments to determine the epigenomic state, using techniques that do not require large sample volumes or complex handling, such as cfChIP-seq, and enriching the sample for clinically relevant DNA based on specific histone modifications.
Enables accurate analysis of epigenomic states using smaller sample volumes with simpler handling, allowing for the classification of tissue types, determination of fetal gestational age, and detection of diseases like cancer and autoimmune disorders, while overcoming the limitations of traditional methods.
Smart Images

Figure 2025528058000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and is a non-provisional application of U.S. Provisional Patent Application No. 63 / 393,725, entitled "EPIGENETICS ANALYSIS OF CELL-FREE DNA," filed July 29, 2022, the disclosure of which is incorporated by reference in its entirety for all purposes. [Background technology]
[0002] Cell-free DNA (cfDNA) is a rich source of information that can be applied to the diagnosis and prognosis of many physiological and pathological conditions, such as pregnancy and cancer (Chan, KCA et al. (2017), New England Journal of Medicine 377, 513-522; Chiu, RWK et al. (2008), Proceedings of the National Academy of Sciences of the United States of America 105, 20458-20463; Lo, YMD et al. (1997), The Lancet 350, 485-487). Cell-free DNA molecules in various body fluids (e.g., plasma, serum, urine, saliva, semen, peritoneal fluid, cerebrospinal fluid) can contain a mixture of DNA molecules from various tissues. One mechanism by which such cfDNA molecules are released is through cell death (e.g., apoptosis or necrosis). Selected cell populations, such as lymphocytes and neutrophils, have also been shown to secrete DNA molecules into body fluids. cfDNA molecules consist of fragmented DNA molecules. Many studies have shown a correlation between cfDNA fragmentation patterns and nucleosome structure (Sun et al. Proc Natl Acad Sci USA. 2018;115:E5106, Snyder et al. Cell. 2016;164:57-68). Circulating cfDNA is now commonly used as a noninvasive biomarker and is known to circulate in the form of short fragments. However, the physiological factors governing cfDNA fragmentation and molecular profiles remain elusive. Cell-free DNA can be analyzed to understand the epigenomic state. The epigenomic state of DNA can indicate the regulation of genes, tissue origin, or disease. The amount of histone modifications is an epigenomic factor. Conventional techniques for detecting histone modifications involve the use of specific antibodies, relatively large sample volumes, and more complicated sample handling. Simpler and more efficient techniques for determining the epigenomic state of DNA are desirable. These and other needs are addressed. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Chan, KCA et al. (2017), New England Journal of Medicine 377, 513-522 [Non-patent document 2] Chiu, RWKet al. (2008), Proceedings of the National Academy of Sciences of the United States of America 105, 20458-20463 [Non-patent document 3] Lo,YMDet al.,(1997),The Lancet 350,485-487 [Non-patent document 4] Sun et al.Proc Natl Acad Sci USA.2018;115:E5106 [Non-patent document 5] Snyder et al.Cell.2016;164:57-68 Summary of the Invention
[0004] The present disclosure describes various techniques, such as measuring the amount (e.g., relative frequency) of sequence motifs and sizes of cell-free DNA fragments in a biological sample of an organism to determine sample characteristics (e.g., the fractional concentration of a tissue type or a tissue type characteristic), measuring the amount of histone modifications, determining the state of the organism based on such measurements, and enriching the biological sample for clinically relevant DNA. Different tissue types exhibit different patterns of chromatin structure. The present disclosure provides various uses for inferring chromatin structure, for example, based on measuring the relative frequency of sequence motifs and / or sizes of cell-free DNA in a mixture of cell-free DNA from various tissues. DNA from one of the specific tissues may be referred to as clinically relevant DNA.
[0005] Various examples may quantify the amount of sequence motifs representing terminal sequences (i.e., terminal motifs) of DNA fragments. For example, embodiments may determine the relative frequency of one or more sets of one or more sequence motifs for the terminal sequences of DNA fragments. In various implementations, a preferred set of terminal motifs may be determined by using another technique (e.g., cfChIP-seq [cell-free chromatin immunoprecipitation followed by sequencing]) to measure the epigenomic state of chromatin (e.g., histone modifications) in a particular region of interest. A preferred set of terminal motifs may be selected based on their more frequent occurrence in one or more regions with a particular epigenomic state compared to other terminal motifs. A particular epigenomic state may be associated with a particular tissue type or clinically relevant DNA.
[0006] In various implementations, the relative frequencies of the preferred sets can be used to classify characteristics of new samples (e.g., fractional concentrations of clinically relevant DNA), measure the state of an organism (e.g., fetal gestational age or level of pathology), or measure the epigenomic state (e.g., histone modification abundance). Thus, embodiments may provide measurements to inform physiological changes, including cancer, autoimmune disease, transplantation, and pregnancy.
[0007] As a further example, a preferred set of sequence end motifs can be used for physical and / or in silico enrichment of a biological sample for clinically relevant cell-free DNA fragments. Enrichment can use sequence end motifs that are preferred for one or more genomic regions with specific histone modifications. Particular histone modifications in one or more genomic regions may be preferred for certain clinically relevant tissues, such as fetuses, tumors, or transplants. Physical enrichment can use one or more probe molecules that detect a specific set of sequence end motifs, such that the biological sample is enriched for clinically relevant DNA fragments. For in silico enrichment, a group of sequence reads of cell-free DNA fragments that have one of a set of preferred end sequences for clinically relevant DNA can be identified. Certain sequence reads can be sorted based on the likelihood that they correspond to clinically relevant DNA, and the likelihood describes sequence reads that contain the preferred sequence end motif. The sorted sequence reads can be analyzed to determine characteristics of the clinically relevant DNA biological sample.
[0008] In some embodiments, the amount of DNA fragments in a particular size range can be used to determine the amount of histone modifications in cell-free DNA. The amount of histone modifications estimated through size information can be used to determine tissue fractions, classification of the level of injury, and the status of tissue or organ transplants.
[0009] In addition, histone modifications in certain genomic regions may indicate that DNA is from a specific type of tissue, but the histone modifications in many genomic regions may be the result of several different tissues.Using the histone modifications in genomic regions that are contributed by several different tissues may allow for more accurate analysis of biological samples than using only the histone modifications in genomic regions that originate from a single tissue.For example, using the histone modifications contributed by several different tissues may result in more accurate analysis of tissue origin and level of injury.
[0010] These and other embodiments of the present disclosure are described in detail below. For example, other embodiments are directed to systems, devices, and computer-readable media related to the methods described herein.
[0011]
[0013] Other features and advantages of the present disclosure will become apparent by reference to the remaining portions of the specification, including the drawings and claims. Further features and advantages of the present disclosure, as well as the structure and operation of various embodiments thereof, are described in detail below with reference to the accompanying drawings, in which like reference numbers may indicate identical or functionally similar elements. [Brief explanation of the drawings]
[0012] [Figure 1] An explanation of the structure of DNA is given. [Figure 2] 1 shows the use of immunoprecipitation to analyze plasma cfDNA molecules associated with histone modifications. [Figure 3] An illustration of the terminal motifs of the fragments is shown. [Figure 4] 1 is a graph defining categories of H3K4me3 regions with different levels of H3K4me3 ChIP signal, according to an embodiment of the present invention. [Figure 5] 1 is a table showing exemplary category definitions of H3K4me3 regions using H3K4me3 ChIP-seq analysis of pregnancy samples, according to embodiments of the present invention. [Figure 6] 1 shows a table illustrating exemplary category definitions of H3K27ac regions using H3K27ac ChIP-seq analysis of pregnancy samples, according to embodiments of the present invention. [Figure 7] 1 is a table showing exemplary category definitions of H3K4me3 regions using H3K4me3 ChIP-seq analysis of samples from healthy non-pregnant subjects, according to embodiments of the present invention. [Figure 8] 1 is a table showing exemplary category definitions of H3K27ac regions using H3K27ac ChIP-seq analysis of samples from healthy non-pregnant subjects, according to embodiments of the present invention. [Figure 9] 1 shows a heatmap of motif frequencies in regions with different levels of H3K4me3 ChIP signal for sequencing results of plasma DNA in which the immunoprecipitation step was omitted, according to an embodiment of the present invention. [Figure 10] 1 is a graph comparing terminal motif frequency rankings between plasma DNA sequencing results with and without H3K4me3-based immunoprecipitation, according to an embodiment of the present invention. [Figure 11] 1 shows a table of the 24 terminal motifs with the greatest ranking difference between traditional cfDNA sequencing and cfChIP-seq for the H3K4me3 histone modification, according to an embodiment of the invention. [Figure 12A] 1 shows the use of terminal motif patterns to estimate histone modification signals in plasma DNA for plasma DNA sequencing results without immunoprecipitation, according to an embodiment of the present invention. [Figure 12B] 1 shows the use of terminal motif patterns to estimate histone modification signals in plasma DNA for plasma DNA sequencing results without immunoprecipitation, according to an embodiment of the present invention. [Figure 13] 1 shows a graph of the correlation between the aggregated abundance of terminal motifs overrepresented in H3K4me3-based immunoprecipitated plasma DNA and H3K4me3 ChIP signal, according to an embodiment of the present invention. [Figure 14] 1 is a graph showing the correlation between cfChIP signal and terminal motif frequency for 11 peak groups, according to an embodiment of the present invention. [Figure 15A] 1 is a graph showing the correlation between cfChIP signal and terminal motif frequency for groups of 6 and 8 peaks, according to an embodiment of the present invention. [Figure 15B] 1 is a graph showing the correlation between cfChIP signal and terminal motif frequency for groups of 6 and 8 peaks, according to an embodiment of the present invention. [Figure 16]1 is a graph of the correlation between H3K4me3 ChIP signal in placenta-specific H3K4me3 regions predicted by terminal motifs and fetal DNA fraction determined by a SNP-based approach, according to an embodiment of the present invention. [Figure 17] 1 is a flowchart of an exemplary process associated with determining the fractional concentration of cell-free DNA fragments in a biological sample, according to an embodiment of the present invention. [Figure 18] 1 is a flowchart of an exemplary process associated with estimating a first value of a characteristic of a target tissue, in accordance with an embodiment of the present invention. [Figure 19] 1 is a flowchart of an exemplary process associated with using sequence motifs to determine the abundance of histone modifications in one or more genomic regions, according to an embodiment of the present invention. [Figure 20] 1 is a flowchart of an exemplary process associated with determining the abundance of histone modifications in one or more genomic regions using fragmentomic features, according to an embodiment of the present invention. [Figure 21] FIG. 1 shows the application of ChIP-seq (cell-free chromatin immunoprecipitation followed by sequencing) to determine contributions from different tissues, according to an embodiment of the present invention. [Figure 22] 1 is a ROC curve for distinguishing patients with and without HCC using predicted H3K4me3 signals in liver-specific H3K4me3 regions using terminal motifs according to an embodiment of the present invention. [Figure 23] 1 is a flowchart of an exemplary process associated with classifying a level of impairment, according to an embodiment of the present invention. [Figure 24A] 1 shows the percentage of cfDNA molecules with certain sizes in regional categories with different levels of H3K27ac signal, according to an embodiment of the present invention. [Figure 24B]1 shows the percentage of cfDNA molecules with certain sizes in regional categories with different levels of H3K27ac signal, according to an embodiment of the present invention. [Figure 24C] 1 shows the percentage of cfDNA molecules with certain sizes in regional categories with different levels of H3K27ac signal, according to an embodiment of the present invention. [Figure 25A] 1 shows that the correlation between size and ChIP signal of a histone modification according to an embodiment of the present invention can be generalized to other histone modifications. [Figure 25B] 1 shows that the correlation between size and ChIP signal of a histone modification according to an embodiment of the present invention can be generalized to other histone modifications. [Figure 25C] 1 shows that the correlation between size and ChIP signal of a histone modification according to an embodiment of the present invention can be generalized to other histone modifications. [Figure 26A] 1 shows the use of size information to estimate histone modifications of plasma DNA for plasma DNA sequencing results without immunoprecipitation, according to an embodiment of the present invention. [Figure 26B] 1 shows the use of size information to estimate histone modifications of plasma DNA for plasma DNA sequencing results without immunoprecipitation, according to an embodiment of the present invention. [Figure 27A] 1 shows the correlation between the percentage of cfDNA molecules within a size range and log-transformed H3K4me3 ChIP signal, according to embodiments of the present invention. [Figure 27B] 1 shows the correlation between the percentage of cfDNA molecules within a size range and log-transformed H3K4me3 ChIP signal, according to embodiments of the present invention. [Figure 27C] 1 shows the correlation between the percentage of cfDNA molecules within a size range and log-transformed H3K4me3 ChIP signal, according to embodiments of the present invention. [Figure 28A]1 shows an evaluation of the performance of estimated H3K4me3 ChIP signals in placenta-specific H3K4me3 regions for fetal DNA fraction estimation, according to an embodiment of the present invention. [Figure 28B] 1 shows an evaluation of the performance of molecules within a certain size range in placenta-specific H3K4me3 regions for fetal DNA fraction estimation, according to an embodiment of the present invention. [Figure 29] 1 is a graph evaluating the performance of predicted H3K27ac ChIP signals in placenta-specific H3K27ac regions for determining fetal DNA fraction, according to an embodiment of the present invention. [Figure 30] 1 is a graph of the Pearson correlation coefficient between size ranges with and without calibration for the amount of H3K27ac signal in placenta-specific regions and fetal DNA fraction determined by a SNP-based approach, according to an embodiment of the present invention. [Figure 31A] 1 is a graph showing the use of estimated H3K4me3 ChIP signals based on liver-specific H3K4me3 regions for HCC detection, according to an embodiment of the present invention. [Figure 31B] 1 is a graph showing the use of estimated H3K4me3 ChIP signals based on liver-specific H3K4me3 regions for HCC detection, according to an embodiment of the present invention. [Figure 32A] 1 shows the use of predicted H3K27ac ChIP signals based on H3K27ac regions for HCC detection, according to an embodiment of the present invention. [Figure 32B] 1 shows the use of predicted H3K27ac ChIP signals based on H3K27ac regions for HCC detection, according to an embodiment of the present invention. [Figure 33] 1 is a receiver operating characteristic (ROC) curve for distinguishing subjects with intermediate and advanced stage hepatocellular carcinoma from healthy subjects according to an embodiment of the present invention. [Figure 34] 1 is a graph showing the correlation between estimated H3K27ac ChIP signal and donor DNA fraction in liver-specific H3K27ac regions, according to an embodiment of the present invention. [Figure 35]10 is a graph of the Pearson correlation coefficient between size ranges with and without calibration for the amount of H3K27ac signal in liver-specific regions and donor DNA fraction determined by a SNP-based approach, according to an embodiment of the present invention. [Figure 36] 1 is a flowchart of an exemplary process associated with using fragment size to determine the abundance of histone modifications in one or more genomic regions, according to embodiments of the present invention. [Figure 37] 1 shows a table of tissue-specific histone modification regions according to an embodiment of the present invention. [Figure 38] 1 is a graph showing tissue mapping of plasma DNA based on H3K4me3 histone modification of cell-free DNA, according to an embodiment of the present invention. [Figure 39] 1 shows a graph illustrating the correlation between placental contribution estimated by H3K4me3 ChIP signal and fetal DNA fraction according to an embodiment of the present invention. [Figure 40] 1 is a graph of the percentage contribution of different tissues for both pregnant and non-pregnant samples based on the H3K27ac histone modification of cfDNA, according to an embodiment of the present invention. [Figure 41] 1 shows a heatmap of tissue contributions inferred from H3K27ac ChIP signals in pregnant and non-pregnant subjects, according to an embodiment of the present invention. [Figure 42] 1 is a graph of estimated H3K27ac ChIP signals across various tissue-specific regions, according to an embodiment of the present invention. [Figure 43A] 1 shows the correlation between placental contribution estimated by H3K27ac ChIP signal and fetal DNA fraction determined by a SNP-based approach, according to an embodiment of the present invention. [Figure 43B] 1 shows the correlation between normalized reads / kb in placenta-specific regions and fetal DNA fraction determined by a SNP-based approach, according to an embodiment of the present invention. [Figure 44]1 is a ROC curve for distinguishing between pregnant and non-pregnant subjects according to an embodiment of the present invention. [Figure 45] 1 shows a receiver operating characteristic (ROC) curve for distinguishing control subjects from subjects with colorectal cancer (CRC) using estimated colon contribution according to an embodiment of the present invention. [Figure 46A] 1 is a graph comparing erythroblast contribution estimated by H3K27ac ChIP signal between subjects with beta-thalassemia major and control subjects without beta-thalassemia major, according to an embodiment of the present invention. [Figure 46B] 1 is a ROC curve for using estimated erythroblast contribution to distinguish between subjects with and without beta-thalassemia major, according to an embodiment of the present invention. [Figure 47] 1 is a heat map of tissue contributions estimated using H3K27ac ChIP signals in subjects with beta thalassemia major and control subjects, according to an embodiment of the present invention. [Figure 48A] 1 shows the correlation between erythrocyte DNA percentage as determined by ddPCR assay and erythroblast contribution as determined by H3K27ac signal, according to an embodiment of the present invention. [Figure 48B] 1 shows the correlation between erythrocyte DNA percentage as determined by ddPCR assay and erythroblast contribution as determined by H3K27ac signal, according to an embodiment of the present invention. [Figure 48C] 1 shows the correlation between erythrocyte DNA percentage as determined by ddPCR assay and erythroblast contribution as determined by H3K27ac signal, according to an embodiment of the present invention. [Figure 49A] 1 is a graph of estimated H3K27ac signals across healthy subjects, subjects with colorectal cancer (CRC) but without liver metastases, and subjects with CRC and liver metastases, according to an embodiment of the present invention. [Figure 49B]1 is a graph of estimated H3K27ac signals across healthy subjects, subjects with colorectal cancer (CRC) but without liver metastases, and subjects with CRC and liver metastases, according to an embodiment of the present invention. [Figure 50] 1 is a graph of tissue contribution in urine and plasma DNA samples using the H3K27ac histone modification of cell-free DNA, according to an embodiment of the present invention. [Figure 51] 1 is a flowchart of an exemplary process associated with determining the fractional concentration of a tissue type, according to an embodiment of the present invention. [Figure 52] 1 is a flowchart of an exemplary process associated with determining a pregnancy or disease classification, according to an embodiment of the present invention. [Figure 53] 1 illustrates the nature of inputs in a machine learning model for determining a cancer classification, according to an embodiment of the present invention. [Figure 54A] 1 shows results from a machine learning model in determining a cancer classification, according to an embodiment of the present invention. [Figure 54B] 1 shows results from a machine learning model in determining a cancer classification, according to an embodiment of the present invention. [Figure 55] 1 shows area under the curve (AUC) results for distinguishing hepatocellular carcinoma (HCC) from non-HCC cases using machine learning models with different fragmentomics features according to embodiments of the present invention. [Figure 56] 1 is a flowchart of an exemplary process associated with analyzing a subject's biological sample to determine a classification of the subject's condition, according to an embodiment of the present invention. [Figure 57] 1 is a flowchart of an exemplary process associated with enriching a biological sample for clinically relevant DNA, according to an embodiment of the present disclosure. [Figure 58] 1 is a flowchart of an exemplary process associated with enriching a biological sample for clinically relevant DNA, according to an embodiment of the present disclosure. [Figure 59] 1 illustrates a measurement system according to one embodiment of the present invention. [Figure 60] 1 shows a block diagram of an exemplary computer system usable with systems and methods according to embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0013] term A "tissue" corresponds to a group of cells grouped together as a functional unit. More than one type of cell can be found within a single tissue. Different types of tissue can consist of different types of cells (e.g., liver cells, alveolar cells, or blood cells), but can also correspond to tissues from different organisms (maternal versus fetal) or healthy versus tumor cells.
[0014] A "biological sample" refers to any sample obtained from a subject (e.g., a human (or other animal), such as a pregnant woman, a human with or suspected of having cancer, an organ transplant recipient, or a subject suspected of having a disease process involving an organ (e.g., the heart in a myocardial infarction, the brain in a stroke, or the hematopoietic system in anemia)) and containing one or more nucleic acid molecules of interest. A biological sample can be a bodily fluid, such as blood, plasma, serum, urine, vaginal fluid, fluid from a hydrocele (e.g., of the testes), vaginal washings, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple discharge, aspirated fluid from different parts of the body (e.g., thyroid, breast), intraocular fluid (e.g., aqueous humor), etc. A stool sample can also be used. In various embodiments, the majority of DNA in a biological sample enriched for cell-free DNA (e.g., a plasma sample obtained via a centrifugation protocol) can be cell-free, e.g., greater than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA can be cell-free. The centrifugation protocol can include, for example, 3,000 g x 10 minutes, obtaining a fluid portion, and recentrifuging, for example, at 16,000 g for an additional 10 minutes to remove residual cells. As part of the analysis of a biological sample, at least 1,000 cell-free DNA molecules can be analyzed. As other examples, at least 10,000, 50,000, 100,000, 500,000, 1,000,000, or 5,000,000, or more cell-free DNA molecules can be analyzed.
[0015] "Clinically relevant DNA" can refer to DNA of a particular tissue source that is measured, for example, to determine the fractional concentration of such DNA or to phenotype a sample (e.g., plasma). Examples of clinically relevant DNA are fetal DNA in maternal plasma, or tumor DNA in a patient's plasma, or other sample containing cell-free DNA. Another example includes measuring the amount of graft-associated DNA in the plasma, serum, or urine of a transplant patient. Further examples include measuring the fractional concentrations of hematopoietic and non-hematopoietic DNA in a subject's plasma, or the fractional concentration of liver DNA fragments (or other tissues) in a sample, or the fractional concentration of brain DNA fragments in cerebrospinal fluid.
[0016] A "sequence read" refers to a chain of nucleotides sequenced from any portion or all of a nucleic acid molecule. For example, a sequence read can be a short chain of nucleotides (e.g., 20-150 nucleotides) sequenced from a nucleic acid fragment, a short chain of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of an entire nucleic acid fragment present in a biological sample. Sequence reads can be obtained in a variety of ways, for example, using sequencing techniques, or using probes, e.g., hybridization arrays or capture probes, or using single primer or isothermal amplification, amplification techniques such as polymerase chain reaction (PCR) or linear amplification. At least 1,000 sequence reads can be analyzed as part of the analysis of a biological sample. As other examples, at least 10,000, or 50,000, or 100,000, or 500,000, or 1,000,000, or 5,000,000, or more sequence reads can be analyzed.
[0017] A sequence read may include "end sequences" associated with the ends of a fragment. An end sequence may correspond to the outermost N bases of a fragment, e.g., 2 to 30 bases at the end of the fragment. If a sequence read corresponds to an entire fragment, the sequence read may include two end sequences. If paired-end sequencing provides two sequence reads corresponding to the ends of a fragment, each sequence read may include one end sequence.
[0018] The "sequence motif" of a "sequence end signature" can refer to a short repeating pattern of bases in a nucleic acid fragment (e.g., a cell-free DNA fragment). A sequence motif can occur at the end of a fragment and thus can be part of or include the end sequence. An "end motif" can refer to a sequence motif for end sequences that occur preferentially at the ends of nucleic acids, e.g., DNA fragments, potentially for a particular tissue type. An end motif can also occur just before or just after the end of a fragment, thereby still corresponding to an end sequence.
[0019] The term "fractional fetal DNA concentration" is used interchangeably with the terms "percent fetal DNA" and "fetal DNA fraction" and refers to the proportion of fetal DNA molecules present in a biological sample (e.g., a maternal plasma or serum sample) that originates from the fetus (Lo et al., Am J Hum Genet. 1998;62:768-775, Lun et al., ClinChem. 2008;54:1664-1672).
[0020] "Relative frequency" (also simply referred to as "frequency") can refer to a proportion (e.g., a percentage, fraction, or concentration). Specifically, the relative frequency of a particular pair of terminal motifs (e.g., A<>A) can provide the proportion of cell-free DNA fragments that have the terminal sequence of that particular pair.
[0021] An "aggregate value" can refer to, for example, a collective property of the relative frequencies of a set of terminal motifs. Examples include the mean, median, sum of the relative frequencies, the variation between the relative frequencies (e.g., entropy, standard deviation (SD), coefficient of variation (CV), interquartile range (IQR), or a particular percentile cutoff (e.g., the 95th or 99th percentile) among various relative frequencies), or the difference (e.g., distance) of the relative frequencies from a reference pattern, which may be implemented by clustering. As another example, the aggregate value can include an array / vector of relative frequencies, which can be compared to a reference vector (e.g., representing multidimensional data points).
[0022] A "calibration sample" may correspond to a biological sample in which the fractional concentration of a clinically relevant nucleic acid (e.g., tissue-specific DNA fraction) is known or determined via a calibration method using tissue-specific alleles, such as in transplants where alleles present in the donor's genome but absent from the recipient's genome may be used as markers for the transplanted organ. As another example, a calibration sample may correspond to a sample in which terminal motifs may be determined. A calibration sample may be used for both purposes. Multiple calibration samples may be used. As an example, a first calibration sample may correspond to a biological sample with measurable histone modification levels across various genomic regions of interest. A second calibration sample may correspond to a biological sample with measurable fragmentomics characteristics across various genomic regions of interest. The first and second calibration samples may be used together to determine a calibration value.
[0023] "Calibration data points" include "calibration values" and measured or known feature values of a target tissue type or the fractional concentration of a clinically relevant nucleic acid (e.g., DNA of a particular tissue type). Calibration values can be determined from various types of data measured from nucleic acid molecules of a sample, such as the amount of a terminal motif or fragment size. Calibration values correspond to parameters that correlate with a desired property, such as the feature value of a target tissue type or the fractional concentration of clinically relevant DNA. For example, calibration values can be determined from the relative frequencies (e.g., aggregate values) of terminal signatures determined for calibration samples in which the desired property is known. Calibration data points can be defined in various ways, for example, as discrete points or as a calibration function (also called a calibration curve or calibration surface). A calibration function can be derived from additional mathematical transformations of the calibration data points. In some embodiments, "calibration data points" can include "calibration values" and measured or known feature values (e.g., fragmentomics characteristics) of a group of genomic regions of interest (e.g., characterized by a certain level of histone modification).
[0024] A "separation value" corresponds to the difference or ratio involving two values, for example, two fractional contributions or two methylation levels. A separation value can be a simple difference or ratio. As an example, the direct ratio of x / y is a separation value, as is x / (x+y). A separation value can include other factors, for example, multiplicative factors. As another example, the difference or ratio of a function of the values can be used, for example, the difference or ratio of the natural logarithms (ln) of two values. A separation value can include differences and ratios.
[0025] "Separation values" and "aggregate values" (e.g., of relative frequency) are two examples of parameters (also called metrics) that provide a measure of the variation of a sample between different classifications (states) and can therefore be used to determine different classifications. Aggregate values can be separation values when the difference is taken between a set of relative frequencies of a sample and a reference set of relative frequencies, as can be done, for example, in clustering.
[0026] As used herein, the term "classification" refers to any number or other feature associated with a particular property of a sample. For example, a "+" sign (or the word "positive") may indicate that the sample is classified as having a deletion or amplification. The classification may be binary (e.g., positive or negative) or may have more levels of classification (e.g., a scale of 1 to 10 or 0 to 1). As a further example, the level of classification may correspond to, for example, a fractional concentration or feature value of a sample or target tissue type.
[0027] As used herein, the term "parameter" refers to a numerical value that characterizes a quantitative data set and / or a numerical relationship between quantitative data sets. For example, the ratio (or function of the ratio) between a first amount of a first nucleic acid sequence and a second amount of a second nucleic acid sequence is a parameter.
[0028] The terms "cutoff" and "threshold" refer to predetermined numbers used in operations. For example, a cutoff size may refer to a size above which a fragment is excluded. A threshold may be a value above or below which a particular classification is applied. Either of these terms can be used in either of these contexts. A cutoff or threshold may be a "reference value" or may be derived from a reference value that represents a particular classification or distinguishes between two or more classifications. Such reference values can be determined in various ways, as will be understood by those skilled in the art. For example, a metric (parameter) can be determined for two different cohorts of subjects with different known classifications, and a reference value can be selected as representative of one classification (e.g., the mean) or as a value between two clusters of the metric (e.g., selected to obtain a desired sensitivity and specificity). As another example, a reference value can be determined based on statistical simulation of samples. Specific values such as cutoffs, thresholds, and criteria can be determined based on a desired accuracy (e.g., sensitivity and specificity). A parameter can be compared to a cutoff value, threshold, reference value, or calibration value to determine a classification. The process for determining such values can be performed, for example, as part of training a machine learning model that receives a training vector of one or more sets of parameters. Comparison of a parameter to any such value can then be achieved by inputting the parameter into a machine learning model that was trained using parameter values determined from another subject, e.g., a subject with or without the condition, abnormality, or pathology, or a subject with known parameter values (e.g., calibration values).
[0029] The term "level of cancer" may refer to whether cancer is present (i.e., present or absent), the stage of cancer, tumor size, whether there is metastasis, total tumor burden in the body, the response of cancer to treatment, and / or other measures of cancer severity (e.g., cancer recurrence). Cancer levels may be numbers or other indicators, such as symbols, alphabetic characters, and colors. A level may be zero. Cancer levels may also include premalignant or precancerous conditions (states). Cancer levels may be used in a variety of ways. For example, screening can confirm the presence of cancer in a person not previously known to have cancer. Evaluation can follow up on a person diagnosed with cancer to monitor the progression of cancer over time, study the effectiveness of a therapy, or determine a prognosis. In one embodiment, prognosis can be expressed as the likelihood that a patient will die from cancer, or the likelihood that the cancer will progress after a certain period or time, or the likelihood or extent that the cancer will metastasize. Detection can mean "screening" or determining whether a person with suggestive features of cancer (e.g., symptoms or other positive test) has cancer. "Level of disease" is similar to "level of cancer," but can refer to disease rather than cancer.
[0030] "Abnormality level" may refer to the amount, degree, or severity of an abnormality associated with an organism, and this level may be as described above for cancer. An example of an abnormality is a pathology associated with an organism. Another example of an abnormality is the rejection of a transplanted organ. Other examples of abnormalities may include autoimmune attacks (e.g., lupus nephritis or multiple sclerosis, which damage the kidney), inflammatory diseases (e.g., hepatitis), fibrotic processes (e.g., cirrhosis), fatty infiltration (e.g., fatty liver disease), degenerative processes (e.g., Alzheimer's disease), and ischemic tissue damage (e.g., myocardial infarction or stroke). A healthy state of a subject may be considered as a normal classification.
[0031] The term "gestational age" may refer to a measure of the duration of pregnancy from the onset of a woman's last menstrual period (LMP), or the corresponding duration of pregnancy estimated, if possible, by more accurate methods, such as adding 14 days to the known duration from fertilization (possible with in vitro fertilization) or by obstetric ultrasound.
[0032] "Pregnancy-related disorder" includes any disorder characterized by abnormal relative expression levels of genes in maternal and / or fetal tissues and / or by abnormal clinical features in the mother and / or fetus. These disorders include preeclampsia (Kaartokallio et al. Sci Rep. 2015;5:14107, Medina-Bastidas et al. Int J Mol Sci. 2020;21:3597), intrauterine growth restriction (Faxen et al. Am J Perinatol. 1998;15:9-13, Medina-Bastidas et al. Int J Mol Sci. 2020;21:3597), invasive placentation, preterm birth (Enquobahrie et al. BMC Pregnancy Childbirth. 2009;9:56), hemolytic disease of the newborn, placental insufficiency (Kelly et al. Endocrinology. 2017;158:743-755), and hydrops fetalis (Magor et al. al. Blood. 2015;125:2405-17), fetal malformations (Slonim et al. Proc Natl Acad Sci USA. 2009;106:9425-9), HELLP syndrome (Dijk et al. J Clin Invest. 2012;122:4003-4011), systemic lupus erythematosus (Hong et al. J Exp Med. 2019;216:1154-1169), and other maternal immune disorders.
[0033] A "site" (also called a "genomic site") corresponds to a single site, which may be a single base position, or a group of correlated base positions, e.g., a CpG site, or a larger group of correlated base positions. A "locus" may correspond to a region that includes multiple sites. A locus may contain only one site, which would make the locus equivalent to the site in that context.
[0034] The term "about" or "approximately" can mean within an acceptable error range for a particular value, as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within 1 or more than 1 standard deviation, according to convention within the art. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term "about" or "approximately" can mean within an order of magnitude, within 5-fold, or in some forms within 2-fold of a value. When a particular value is described in this application and claims, unless otherwise stated, the term "about" meaning within an acceptable error range of the particular value should be assumed. The term "about" can have the meaning commonly understood by one of ordinary skill in the art. The term "about" can refer to ±10%. The term "about" can refer to ±5%.
[0035] Where a range of values is provided, unless the context clearly dictates otherwise, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limit of that range is also specifically disclosed. It is also understood that the endpoints of the provided ranges are included in the range. Each smaller range between any stated or intervening value in a stated range and any other stated or intervening value in that stated range is encompassed within embodiments of the present disclosure. The upper and lower limits of these smaller ranges may independently be included or excluded in the range, and each range where either, neither, or both limits are included in the smaller range is also encompassed within the present disclosure, subject to any specifically excluded limits in the stated range. When a stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure.
[0036] Standard abbreviations may be used, such as bp: base pairs, kb: kilobase, pi: picoliters, s or sec: seconds, min: minutes, h or hr: hours, aa: amino acids, nt: nucleotides, etc.
[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the present disclosure, some potential exemplary methods and materials can be described herein. (Mode for Carrying Out the Invention)
[0038] The epigenomic state (DNA and proteins) of different regions of chromatin can indicate gene expression activity, tissue origin, or disease. Histone modifications are an example of epigenomic factors, and measuring the amount of histones with specific epigenomic states can be used in various ways. Techniques for detecting histone modifications include cfChIP-seq (cell-free chromatin immunoprecipitation followed by sequencing), which has several drawbacks. cfChIP-seq requires a sample volume of 1-2 ml or more, which is large compared to the hundreds of microliters or less required for simple sequencing. In addition, cfChIP-seq uses more complex and time-consuming sample techniques than traditional plasma cfDNA-seq procedures. In the cfChIP-seq procedure, the target epigenome is linked to proteins (e.g., histone modifications). Proteins are less stable than DNA. Freezing, thawing, and storage conditions affect protein stability more than DNA stability.
[0039] This disclosure demonstrates that certain terminal motifs (i.e., sequences at the ends of naturally fragmented DNA), size, and / or other fragmentomic characteristics of cell-free DNA are highly correlated with histone modifications. The abundance of these terminal motifs can indicate the abundance of histone modifications in a sample, and therefore in a subject. As a result, terminal motifs can be used to indicate gene activity, tissue origin, or disease, circumventing the drawbacks of cfChIP-seq. Analysis of terminal motifs can use sequencing techniques that do not require the additional step of cfChIP-seq. As a result, embodiments of the present invention can use biological samples of less than 100 μl, which can contain approximately 500 pg of cell-free DNA. Sample handling for sequencing is much simpler than with cfChIP-seq techniques. Samples do not need to be frozen to temperatures below -80°C. Samples can be transported over longer distances from the clinic to the laboratory. In addition, analysis of terminal motifs can be applied to study multiple different epigenome types from a single measurement, rather than being limited by the specific histone modification bound by the specific antibody used in a particular cfChIP-seq assay.
[0040] Thus, measuring certain terminal motifs in cell-free DNA can provide an improved technique for determining the epigenomic state of specific regions of chromatin, for example, corresponding to specific regions of a reference genome. Additionally, measuring certain terminal motifs can also determine different characteristics of a sample, such as the fractional concentration of a tissue type, a classification of a disorder, gestational age, the nutritional status of an organ, organ size, or other characteristics, where such characteristics are associated with the epigenomic state of a specific region. These characteristics can be determined using the epigenomic state determined from the terminal motifs or directly from the terminal motifs.
[0041] A sample can be enriched, either physically or in silico, for certain terminal motifs that are more frequently associated with certain epigenomic states, including histone modifications. Enriching a sample can allow for more accurate measurement of a sample's characteristics, the amount of histone modification, or the state of an organism.
[0042] I. Epigenomic status Figure 1 shows an illustration of the structure of DNA. DNA within a cell is a large structure. In addition to the nucleotides in DNA, DNA is assembled with several different proteins, including chromatin remodelers, transcription factors, nucleosomes, and histones. Histones are proteins around which DNA wraps. DNA typically wraps around eight histone proteins (e.g., a histone octamer 104). The structural unit of DNA around a histone is a nucleosome. Histones can carry modifications that can affect gene transcription. Histone modifications include methylation and acetylation. Histone modifications are part of the epigenome. The epigenetic state differs for different types of cells. The structure of DNA and proteins within a cell is called chromatin. Within chromatin, DNA itself is also methylated. The protein structures that physically open and close chromatin and other DNA modifications that contribute to chromatin structure are also part of the epigenome. Chromatin remodelers are versatile tools that catalyze a wide range of chromatin-altering reactions, including octamer sliding across DNA (nucleosome sliding), changes in nucleosomal DNA conformation, and alterations in octamer composition (histone variant exchange). Additionally, chromatin remodelers can remove other chromatin proteins from chromatin.
[0043] Histone modifications have various functions in cells. One function is to regulate gene expression. Gene expression can be promoted or inhibited. For example, the amount of H3K4me3 correlates with transcriptional activity. In some cases, histone modifications can increase chromatin condensation and decrease transcription (e.g., H3K36me3).
[0044] II. Measurement of epigenetic status A. Histone modifications determined using cfChIP-seq Plasma DNA pools are a mixture of DNA molecules released from various tissues, some of which are bound to histone proteins with specific histone modifications. Histone proteins include H1 (linker histone), H2A / B, H3, and H4 (core histones). DNA molecules associated with histone proteins can form nucleosome structures (Zhou et al. Nat Struct Mol Biol. 2019;26:3-13). Coiling of DNA around histones is primarily due to electrostatic affinity between the positively charged histones and the negatively charged phosphate backbone of DNA. Histone modifications include, but are not limited to, histone methylation, acetylation, phosphorylation, and ubiquitination (Barth et al. Trends Biochem. Sci. 2010;35:618-626). Histone methylation can occur at different lysine residues of histones. Methylation of each lysine residue can involve one, two, or three methyl groups, such that the lysine residue is mono-, di-, or trimethylated, respectively. Examples of histone methylation include, but are not limited to, trimethylation of lysine (K) residue 4 at the N-terminus of histone H3 (H3K4me3) for transcriptional activation, monomethylation of lysine (K) residue 4 at the N-terminus of histone H3 (H3K4me1), H3K27me3 and H3K9me3 for transcriptional inactivation, and H3K36me3, which is associated with transcribed regions within gene bodies. H3K9me2 has been reported to signal heterochromatin formation in gene-poor chromosomal regions with tandem repeat structures, such as satellite repeats, telomeres, and pericentromeres. Histone acetylation includes, but is not limited to, H3K27ac, H3K9ac, and H3K14ac.
[0045] The plasma cfDNA molecules that histones with certain modifications are bound to can be isolated through chromatin immunoprecipitation.These immunoprecipitated plasma cfDNA molecules can be analyzed using different techniques.In one embodiment, they can be analyzed by DNA sequencing.
[0046] FIG. 2 illustrates the use of immunoprecipitation to analyze plasma cfDNA molecules associated with histone modifications. Step 204 shows the plasma portion of a blood sample. The plasma is isolated. Step 208 shows plasma components, including DNA, DNA surrounding histones, and DNA surrounding histones with histone modifications. Plasma cfDNA molecules associated with histone modifications, such as H3K27ac, are precipitated using magnetic beads conjugated with an H3K27ac antibody. Step 212 shows the precipitated plasma cfDNA molecules. Step 216 shows a DNA library prepared, in which DNA molecules are linked to barcoded adapters. The precipitated cfDNA molecules were analyzed by next-generation sequencing (e.g., Illumina NextSeq 500). Sequencing reads can be aligned to the human reference genome GRCh37 (hg19), for example, using Bowtie2 (Langmead et al. Nat Methods. 2012;9:357-359). In some embodiments, algorithms such as, but not limited to, SOAP2 (Li et al. Bioinformatics. 2009;25:1966-67), Burrows-Wheeler Aligner (BWA) (Li et al. Bioinformatics. 2009;25:1754-60), BLAT (Kent. Genome Res. 2002:12:656-664), BLAST (Zhang et al. J Comput Biol. 2000;7:203-14), BFAST (Homer N et al. PLoS One. 2009;4:e7767), MOSAIK (Lee et al. PLoS One. 2014;9:e90581), etc. may be used. Step 220 shows a plot of histone modification signal (y-axis) versus genomic position (x-axis). The sequencing depth (or sequencing read density) of a particular genomic region represents the extent of H3K27ac modification present in that region across different cell types. The greater the sequencing depth of a particular region, the more H3K27ac modifications can potentially be identified.If such H3K27ac modification is specific to a particular cell type in a particular region, the sequencing depth in such region can be used to determine the amount of cfDNA molecules carrying H3K27ac from that cell type.In one embodiment, the sequencing depth can be normalized and corrected for sequencing bias and / or noise resulting from non-specific binding.In some examples, the sequencing depth associated with chromatin immunoprecipitation assay followed by sequencing (i.e., ChIP-seq) can be used to define histone modification signals or ChIP signals.
[0047] B. Selected terminal motifs represent histone modifications. Using the characteristics of fragmentomics, including but not limited to the terminal motifs and size of plasma DNA, we developed a new approach to analyze histone modifications in plasma without the need for immunoprecipitation. Regions relatively enriched with histone modifications will generate differential fragment terminal motif patterns when compared to regions lacking histone modifications. Therefore, the pattern of fragment terminal motifs can be used to estimate histone modifications. A terminal motif can be defined as one or more nucleotides at one end of a cell-free DNA fragment. The number of nucleotides (nt) at each fragment end used for analysis can be, for example, but is not limited to, 1 nt, 2 nt, 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, and 10 nt or more. The fragment size of plasma DNA can be measured in various ways. In one embodiment, the fragment size of plasma DNA can be measured by the number of nucleotides present in the plasma DNA molecule. In another embodiment, the fragment size of plasma DNA can be measured using paired-end sequencing, aligning the sequence to the genome, and then estimating the size from the genomic coordinates of the aligned sequence. In embodiments, tissue- or disease-specific histone modification levels are estimated, such as from terminal motif or size frequencies in cfDNA, allowing for the detection of monitoring the physiology or pathology of one or more tissues, or monitoring a disease state.
[0048] Regions with histone modifications may include, but are not limited to, repeat regions, X-inactivated regions, chromatin structures (e.g., open and closed chromatin structures), pseudogenes, CTCF, DNase I hypersensitive sites [DHS], active and inactive transcribed regions, G-quadruplexes, etc. For example, terminal motifs selected in regions with DNase I hypersensitive sites can be used to inform the amount of histone modifications associated with the DNase I hypersensitive sites. As another example, the size of DNA fragments in X-inactivated regions can inform the amount of histone modifications of X-chromosome genes.
[0049] A particular region may be associated with a particular tissue type. In some cases, certain characteristics of a region may occur more frequently in a particular tissue type. As an example, a region with open chromatin (i.e., large gaps between histones) may occur more frequently in a particular tissue type than in other tissue types. Other characteristics may include regions that are repetitive, X-inactivated, closed chromatin structures, pseudogenes, CTCF, DHS, actively transcribed regions, inactively transcribed regions, or G-quadruplexes. A particular region may be specifically associated with a particular tissue type and not with other tissue types. In other embodiments, a particular region may be associated with several different tissue types. The spread of the region characteristic may relate to the contribution of a particular tissue type and the relative intensity of the particular tissue associated with the region characteristic. Deconvolution, similar to that described for histone modifications below, can be used to determine tissue contributions from these regions.
[0050] 1. Determining terminal motifs associated with histone modifications Different histone modifications can confer different accessibility to DNA nucleases, thus resulting in characteristic fragmentation. Selective cleavage of DNA by nucleases during cfDNA fragmentation occurs at TSSs and CpG islands, which have specific epigenetic states (Han et al., Genome Res. 2021:31:2008-2021). The fragmentation pattern of cell-free DNA can be useful for inferring the histone modifications present in plasma DNA molecules. In embodiments, we analyzed the nuclease cleavage preference for cfDNA within regions of interest, which can be indicated by the pattern of cfDNA terminal motifs. A fragment terminal motif can be defined as one or more nucleotides at the end of a cell-free DNA fragment. For example, we determined the percentage of cfDNA molecules carrying specific 4-mer terminal motifs (256 types in total).
[0051] Figure 3 shows an illustration of the terminal motifs of the fragments. Each nucleotide can be one of the four nucleotides A, C, G, T. For a terminal motif of four nucleotides, 4 There are 256 possible combinations (i.e., 4-mer terminal motifs were defined as four nucleotides at the 5' end of the cfDNA molecule.
[0052] Regions involved in histone modifications can be grouped into different categories depending on the magnitude of the ChIP signal. Figure 4 shows a graph defining the categories of H3K4me3 regions with different levels of H3K4me3 ChIP signal. The y-axis is log 10 The x-axis shows the ranking of genomic regions associated with H3K4me3. Higher rankings indicate stronger signals. Regions were first sorted by ChIP signal magnitude and then empirically classified into nine categories.
[0053] Figure 5 is a table showing exemplary category definitions of H3K4me3 regions using H3K4me3 ChIP-seq analysis of pregnancy samples. Column 1 indicates the category identification. Column 2 indicates the number of regions in the category. Column 3 indicates the percentile range of ChIP signal magnitude in the category's region. Column 4 indicates the average ChIP signal in the category's region. As shown in Figure 5, we empirically classified H3K4me3-associated regions into nine categories according to the percentile range in terms of ChIP signal intensity. The ChIP signal intensity of a region can be the average value of FPKM across 12 pregnancy samples subjected to H3K4me3 ChIP-seq analysis. For example, the percentile range of ChIP signals from 0 to 70 is defined as category 1 with a mean ChIP signal of 0.10, the percentile range of ChIP signals from 70 to 80 is defined as category 2 with a mean ChIP signal of 0.81, the percentile range of ChIP signals from 80 to 90 is defined as category 3 with a mean ChIP signal of 1.59, the percentile range of ChIP signals from 90 to 95 is defined as category 4 with a mean ChIP signal of 3.27, and the percentile range of ChIP signals from 95 to 97 is defined as category 5.84. A percentile range of 97-98 ChIP signals was defined as category 5 with an average ChIP signal of 9.93, a percentile range of 98-98.5 ChIP signals was defined as category 7 with an average ChIP signal of 14.63, a percentile range of 98.5-99 ChIP signals was defined as category 8 with an average ChIP signal of 18.81, and a percentile range of 99 or higher ChIP signals was defined as category 9 with an average ChIP signal of 31.68.
[0054] Figure 6 shows a table illustrating exemplary category definitions for H3K27ac regions using H3K27ac ChIP-seq analysis of pregnancy samples. The table in Figure 6 follows the same format as the table in Figure 5. As shown in Figure 6, we empirically classified regions associated with H3K27ac into nine categories according to percentile ranges in terms of ChIP signal intensity. The ChIP signal intensity of a region can be expressed as the average value of FPKM across 19 pregnancy samples subjected to H3K27ac ChIP-seq analysis. For example, a percentile range of ChIP signals from 0 to 70 is defined as category 1 with a mean ChIP signal of 0.45, a percentile range of ChIP signals from 70 to 80 is defined as category 2 with a mean ChIP signal of 0.99, a percentile range of ChIP signals from 80 to 90 is defined as category 3 with a mean ChIP signal of 1.31, a percentile range of ChIP signals from 90 to 95 is defined as category 4 with a mean ChIP signal of 1.84, and a percentile range of ChIP signals from 95 to 97 is defined as category 5 with a mean ChIP signal of 2.4 A percentile range of ChIP signals of 97-98 was defined as category 5 with an average ChIP signal of 2.93, a percentile range of ChIP signals of 98-98.5 was defined as category 7 with an average ChIP signal of 3.34, a percentile range of ChIP signals of 98.5-99 was defined as category 8 with an average ChIP signal of 3.74, and a percentile range of ChIP signals above 99 was defined as category 9 with an average ChIP signal of 5.33. We could also define region categories using other methods, including but not limited to K-means clustering analysis.
[0055] Figure 7 is a table showing exemplary category definitions of H3K4me3 regions using H3K4me3 ChIP-seq analysis of samples from non-pregnant healthy subjects. The table in Figure 7 follows the same format as the table in Figure 5. Figure 7 shows the construction of a reference using non-pregnant healthy samples subjected to ChIP-seq analysis. As shown in Figure 7, we empirically classified H3K4me3-associated regions into nine categories according to percentile ranges in terms of ChIP signal intensity. The ChIP signal intensity of a region can be the average value of FPKM across four healthy samples subjected to H3K4me3 ChIP-seq analysis. For example, a percentile range of ChIP signals from 0 to 70 is defined as category 1 with a mean ChIP signal of 0.00, a percentile range of ChIP signals from 70 to 80 is defined as category 2 with a mean ChIP signal of 0.15, a percentile range of ChIP signals from 80 to 90 is defined as category 3 with a mean ChIP signal of 0.69, a percentile range of ChIP signals from 90 to 95 is defined as category 4 with a mean ChIP signal of 2.71, and a percentile range of ChIP signals from 95 to 97 is defined as category 5 with a mean ChIP signal of 6.00. A percentile range of ChIP signals from 97 to 98 was defined as category 5 with an average ChIP signal of 11.39, a percentile range of ChIP signals from 98 to 98.5 was defined as category 7 with an average ChIP signal of 17.11, a percentile range of ChIP signals from 98.5 to 99 was defined as category 8 with an average ChIP signal of 21.95, and a percentile range of ChIP signals above 99 was defined as category 9 with an average ChIP signal of 35.44.
[0056] Figure 8 is a table showing exemplary category definitions of H3K27ac regions using H3K27ac ChIP-seq analysis of samples from healthy non-pregnant subjects. The table in Figure 8 follows the same format as the table in Figure 5. As shown in Figure 8, we empirically classified regions associated with H3K27ac into nine categories according to percentile ranges in terms of ChIP signal intensity. The ChIP signal intensity of a region can be the average value of FPKM across six healthy samples subjected to H3K27ac ChIP-seq analysis. For example, a percentile range of ChIP signals from 0 to 70 is defined as category 1 with a mean ChIP signal of 0.23, a percentile range of ChIP signals from 70 to 80 is defined as category 2 with a mean ChIP signal of 0.89, a percentile range of ChIP signals from 80 to 90 is defined as category 3 with a mean ChIP signal of 1.49, a percentile range of ChIP signals from 90 to 95 is defined as category 4 with a mean ChIP signal of 2.45, and a percentile range of ChIP signals from 95 to 97 is defined as category 5 with a mean ChIP signal of 3.3 A percentile range of ChIP signals of 9 was defined as category 5 with an average ChIP signal of 9, a percentile range of ChIP signals of 97-98 was defined as category 6 with an average ChIP signal of 4.07, a percentile range of ChIP signals of 98-98.5 was defined as category 7 with an average ChIP signal of 4.56, a percentile range of ChIP signals of 98.5-99 was defined as category 8 with an average ChIP signal of 5.01, and a percentile range of ChIP signals above 99 was defined as category 9 with an average ChIP signal of 6.54.
[0057] We analyzed 4mer end motif frequencies across nine categories defined according to different levels of H3K4me3 signal in samples without immunoprecipitation.
[0058] Figure 9 shows a heat map of motif frequency in regions with different levels of H3K4me3 ChIP signal for plasma DNA sequencing results. Graph 904 shows the average H3K4me3 ChIP signal. The y-axis shows the average H3K4me3 ChIP signal. The x-axis shows nine categories of H3K4me3 regions. The x-axis categories align with the regions in heat map 908. The y-axis of heat map 908 corresponds to different 4-mer terminal motifs. The more red the point, the higher the terminal motif frequency in a region of a sample compared to other combinations of region category and sample. The more blue the point, the lower the terminal motif frequency in a region of a sample compared to other combinations of region category and sample. As shown in Figure 9, the terminal motif frequency from the sequencing data of plasma DNA without immunoprecipitation varies according to the intensity of the ChIP signal obtained from the sequencing data of plasma DNA with immunoprecipitation, suggesting the possibility of inferring histone modifications in plasma DNA based on the terminal motifs of plasma DNA molecules without immunoprecipitation. Point 912 is the intersection of four unequal-sized quadrants. The upper right quadrant is redder. The upper left quadrant is bluer. The lower right quadrant is bluer. The lower left quadrant is redder.
[0059] Figure 10 shows a graph comparing terminal motif frequency rankings between plasma DNA sequencing results with and without H3K4me3-based immunoprecipitation. The y-axis shows the ranking of terminal motifs from cfChIP-seq for H3K4me3 from 256 to 1, with 1 representing the most frequent terminal motif. The x-axis shows the ranking of terminal motifs from conventional cfDNA sequencing of plasma samples without the addition of an antibody specific for the H3K4me3 modification from 256 to 1. The shape of the data points indicates the terminal nucleotide (circle for A, triangle for C, square for G, and plus for T).
[0060] Figure 10 shows that, compared to plasma DNA without immunoprecipitation, the sequencing results of plasma DNA immunoprecipitated with H3K4me3 showed that several 4-mer terminal motifs appeared to be overrepresented, including but not limited to GCGG, GCGC, CGCG, CCGC, CCGA, TCCG, CCGT, GGCG, CCGG, TGCG, GCCG, CTCG, GCGA, TCGG, CGGC, TCGC, CGGG, CGCC, ACCG, AGCG, CGGA, GGGC, GCGT, CACG, etc. (i.e., motifs whose y-axis ranking is lower than their x-axis ranking). A dominant terminal motif was considered to be one that crossed the diagonal line y = x, where x = 100. These dominant terminal motifs may suggest the presence of histone modifications (H3K4me3). In another embodiment, a minor motif (i.e., a motif whose y-axis ranking is higher than its x-axis ranking) can be used.
[0061] Figure 11 shows a table of the 24 terminal motifs with the largest ranking difference between conventional cfDNA sequencing and cfChIP-seq for the H3K4me3 histone modification. Column 1 shows the motif. Column 2 shows the nucleotide at the very end of the fragment (i.e., the first nucleotide listed in column 1). Column 3 shows the ranking of the motif in conventional cfDNA sequencing, with 1 being the highest frequency and highest ranking, and 256 being the lowest frequency and lowest ranking. Column 4 shows the ranking of the motif in cfChIP-seq for the H3K4me3 histone modification. Column 5 shows the ranking difference when taking the cfChIP-seq ranking and subtracting the conventional cfDNA sequencing ranking. Columns are ordered by the magnitude of the ranking difference. Data was obtained from multiple healthy subjects.
[0062] The results also show that many of the higher-ranking terminal motifs in cfChIP-seq have C and G nucleotides adjacent to each other. H3K4me3 sites appear to be enriched at CG sequences.
[0063] Thus, terminal motifs with the largest ranking differences occur more frequently in regions associated with H3K4me3 compared to the whole genome or a random group of DNA fragments without cfChIP.
[0064] 12A and 12B show the use of terminal motif patterns to estimate histone modification signals of plasma DNA from the results of sequencing plasma DNA without immunoprecipitation. FIG. 12A shows the construction of a recalibration formula using the frequency of dominant terminal motifs in nine categories and the level of H3K4me3 ChIP signal. In step 1204, regions involved in H3K4me3 are grouped into different categories according to the magnitude of the ChIP signal. In one embodiment, the regions can be divided into nine categories based on the magnitude of the ChIP signal in each region. After obtaining the region categories with different ChIP signals, the terminal motif patterns in each region category (e.g., the aggregated frequency of cfDNA molecules with dominant terminal motifs from the results of sequencing plasma DNA without immunoprecipitation) can be used to correlate with the H3K4me3 ChIP signal. In step 1208, a recalibration formula can be determined based on the correlation between fragment terminal motifs and ChIP signal. A linear formula is shown as an example of a recalibration formula, but nonlinear formulas can also be used.
[0065] Figure 12B shows how the recalibration formula can be used to predict ChIP signals in other regions (e.g., placenta-specific H3K4me3 regions) according to their corresponding terminal motif information (i.e., predicted ChIP signals). In step 1212, plasma DNA is sequenced without immunoprecipitation. In step 1216, molecules that overlap with tissue-specific (e.g., placenta) H3K4me3 regions are identified. In step 1220, the frequency of terminal motifs that account for a large proportion in H3K4me3-based immunoprecipitated plasma DNA is calculated. The terminal motif information is input into the recalibration formula, and in step 1224, H3K4me3 ChIP signals are predicted in tissue-specific (e.g., placenta) H3K4me3 regions.
[0066] 2. Correlation test with cfChIP-seq signals Figure 13 shows a graph of the correlation between the aggregate abundance of over-represented terminal motifs in H3K4me3-based immunoprecipitated plasma DNA and the H3K4me3 ChIP signal. The x-axis shows the frequency of over-represented terminal motifs as a percentage. The y-axis shows log 10 Figure 13 shows that the aggregate abundance of dominant terminal motifs was highly correlated with H3K4me3 ChIP signal (Pearson's r: 0.99, P-value: <0.0001). The data indicate that the use of terminal motifs in plasma DNA can be used to estimate the intensity of the signal associated with a particular histone modification. Thus, a linear regression model can be used to generate a calibration equation, facilitating the estimation of H3K4me3 ChIP signal based on the terminal motifs of plasma DNA molecules without the need for immunoprecipitation assays. Additionally, motif frequency can be used to predict H3K4me3 histone modification and any other characteristics in which H3K4me3 histone modification may be used, such as the percentage of DNA from a particular tissue type or the condition of a subject.
[0067] The higher frequency of the 24 terminal motifs from Figure 11 is expected to correlate with larger cfChIP-seq signals. To test this hypothesis, we divided the cfChIP-seq signals into different numbers of groups based on peak height.
[0068] Figure 14 is a graph showing the correlation between cfChIP signal and terminal motif frequency for 11 peak groups. Each dot (data point) corresponds to a different peak group among the 11 peak groups. Because peaks correspond to signal values, signal increases with the series of peak groups. The x-axis shows the aggregate frequency of terminal motifs in a peak group relative to all motifs in the particular genomic region analyzed. The terminal motif frequency of a peak group is relative to the particular genomic region associated with the peak group. The y-axis shows the average signal from cfChIP-seq for the H3K4me3 histone modification for each peak group. As an example, each peak group may include several peaks, such as those shown in Figures 5, 6, 7, and 8.
[0069] Figure 15A is a graph showing the correlation between cfChIP signal and terminal motif frequency for six peak groups. The y-axis shows the average signal from cfChIP-seq for the H3K4me3 histone modification for each peak group. The x-axis, similar to Figure 14, shows the frequency of terminal motifs. Each dot represents one of the six peak groups. The terminal motif frequency for a peak group is relative to the specific genomic region associated with the peak group. The terminal motifs in Figure 15A included all 24 terminal motifs identified in Figure 11. The graph shows a high correlation, with an R value of 0.98 and a p-value of 0.00059. This graph suggests that using the frequency of the top 24 terminal motifs correlates with the cfChIP-seq signal of the H3K4me3 histone modification. The graph also shows that grouping the terminal motifs into six peak groups maintains a correlation with the cfChIP-seq signal.
[0070] Figure 15B is a graph showing the correlation between cfChIP signal and terminal motif frequency for eight peak groups. The y-axis shows the average signal from cfChIP-seq for the H3K4me3 histone modification for each peak group. The x-axis, similar to Figure 14, shows the frequency of terminal motifs. Each dot represents one of the eight peak groups. The terminal motif frequency for a peak group is relative to the specific genomic region associated with the peak group. The terminal motifs in Figure 15A included all 24 terminal motifs identified in Figure 11. The graph shows a high correlation, with an R value of 0.97 and a p-value of 4.4e-05. This graph suggests that using the frequency of the top 24 terminal motifs correlates with the cfChIP-seq signal of the H3K4me3 histone modification. The graph also shows that grouping the terminal motifs into eight peak groups maintains a correlation with the cfChIP-seq signal.
[0071] Figures 14, 15A, and 15B show that terminal motif frequency correlates with the signal from cfChIP-seq signal peaks within a group. The correlation is high even when the number of peak groups varies.
[0072] III. Using Sequence Motifs to Analyze Epigenomic Status Terminal motif frequency can identify epigenomic states, and because different cells have different epigenomic states, terminal motif frequency can be used to identify tissue origin, determine the fractional concentration of tissue in a sample, estimate tissue characteristics, or determine the level of a disorder. Terminal motif frequency can also measure the amount of histone modifications.
[0073] A. Estimation of fractional concentrations in tissue of origin Genomic regions with high H3K4me3 signals are known for the placenta (Figure 4). Additionally, the terminal motif frequencies of these genomic regions are known for different peak groups (Figure 14). The overall terminal motif frequencies are determined for 24 terminal motifs in various genomic regions corresponding to 11 peak groups. Based on the terminal motif frequencies, the H3K4me3 signal is predicted. The equation describing the linear relationship in Figure 14 is log(average H3K4me3 signal) = a * (terminal motif frequency) + b.
[0074] 1. Results Figure 16 is a graph of the correlation between H3K4me3 ChIP signal at placenta-specific H3K4me3 regions estimated by terminal motifs and fetal DNA fraction determined by a SNP-based approach. The x-axis is the fetal DNA fraction as a percentage, as determined by a SNP-based approach. The y-axis is the H3K4me3 ChIP signal estimated using terminal motifs. The H3K4me3 ChIP signal estimated using terminal motifs correlated with the fetal DNA fraction in plasma DNA of pregnant women (Pearson's r: 0.67, P-value: <0.001).
[0075] 2. Exemplary Method for Determining Fractional Concentrations FIG. 17 is a flowchart of an exemplary process 1700 related to determining the fractional concentration of cell-free DNA fragments in a biological sample. The biological sample may include cell-free DNA fragments. The biological sample may be any biological sample described herein, including plasma or serum. In some implementations, one or more process blocks of FIG. 17 may be performed by a system (e.g., measurement system 5900). In some implementations, one or more process blocks of FIG. 17 may be performed by another device or devices separate from or including the system. Additionally or alternatively, one or more process blocks of FIG. 17 may be performed by one or more components of measurement system 5900, such as assay 5908, assay device 5910, detector 5920, logic system 5930, local memory 5935, external memory 5940, storage device 5945, and / or processor 5950.
[0076] A plurality of sequence reads for the cell-free DNA fragments are received at block 1710. The plurality of sequence reads includes end sequences corresponding to ends of the plurality of cell-free DNA fragments.
[0077] In some embodiments, process 1700 can include sequencing cell-free DNA fragments in a biological sample to obtain multiple sequence reads. In embodiments, the volume of the biological sample can be 100 μl or less, including 80-100 μl, 50-80 μl, or 30-50 μl. The biological sample can use a volume smaller than that used in cfChIP-seq.
[0078] In some embodiments, process 1700 may include probe-based techniques for measuring the amount of motifs. Techniques may include qPCR, digital PCR, digital droplet PCR, etc. As an example, cfDNA molecules may be subjected to the processes of DNA end pairing, A-tailing, and universal adapter ligation. The adapter-ligated molecules may be separated into different reactions, such as droplets. A pair of PCR primers may be designed so that one primer binds to the universal adapter region and the other primer binds to a specific region of interest. DNA molecules are amplified inside the reaction (e.g., droplets) by the pair of PCR primers. A fluorescent probe specific to a particular terminal motif may be hydrolyzed and emit a fluorescent signal, thereby enabling detection of the presence and quantification of the specific motif. For digital PCR, the number of reactions positive for a specific terminal motif can be counted and used to determine the amount of DNA fragments with that terminal motif in the analyzed region. For real-time PCR, the intensity of each signal can be used as a measure of the amount of DNA fragments terminating in the specific motif. The two intensities may be compared with each other.
[0079] At block 1720, a group of sequence reads located in one or more genomic regions is identified. Each of the one or more genomic regions has a histone modification associated with a target tissue type. The target tissue type may include placenta, liver, heart, neutrophils, monocytes, B cells, adipose tissue, NK cells, or any tissue type described herein. The histone modifications may be H3K4me3, H3K4me1, H3K4me2, H3K27me3, H3K27ac, H3K36me3, H3K9me2, H3K9me3, H3S10P, H3R2me, H3T2P, H3K14ac, H3K9ac, H3K79me2, H3K79me3, H4K5ac, H4K8ac, H4K12ac, H4K16ac, H4K20me, H2BK120ub, or H2AK119ub. The one or more genomic regions may include a transcription start site, a promoter region, an enhancer region, a super-enhancer region, a gene body, a repetitive sequence, a satellite repeat, a telomere, a pericentromere, a mitotic chromosome, a transcription end site, an exon, an intron, an insulator, etc. The one or more genomic regions may have an amount of histone modification that is statistically significantly different from the amount of histone modification in other genomic regions, or from the average amount of modification in other genomic regions or across all genomic regions. Sequence reads may be aligned to a reference genome (e.g., a human reference genome) to determine whether the sequence reads map to one or more genomic regions.
[0080] In block 1730, one or more sequence motifs corresponding to one or more end sequences of the corresponding cell-free DNA fragment are determined for each sequence read in the group of sequence reads. The one or more sequence motifs may correspond to a single nucleotide, a two-nucleotide sequence, a three-nucleotide sequence, a four-nucleotide sequence, a five-nucleotide sequence, a six-nucleotide sequence, a seven-nucleotide sequence, an eight-nucleotide sequence, or a sequence having more than eight nucleotides. The one or more sequence motifs may each have the same number of nucleotides. In some embodiments, the sequence motif comprises nucleotides at the ends of the cell-free DNA fragment. The sequence motif may be at the 5' end of the cell-free DNA fragment. In some embodiments, the sequence motif may be at the 3' end. In embodiments, the one or more sequence motifs may include sequence motifs at the 3' end and the 5' end. If the entire fragment is sequenced, two sequence motifs may be determined.
[0081] In block 1740, the relative frequencies of one or more of the sets of one or more sequence motifs are determined. The sets of one or more sequence motifs occur at a higher rate in chromatin immunoprecipitation followed by sequencing (ChIP-seq) for histone modifications associated with one or more genomic regions than in sequencing without chromatin immunoprecipitation. The chromatin immunoprecipitation may be cell-free chromatin immunoprecipitation followed by sequencing (cfChIP-seq) or cellular chromatin immunoprecipitation followed by sequencing. Sequencing without chromatin immunoprecipitation may include genome-wide sequencing. The sets of one or more sequence motifs correspond to sequence motifs with similar relative frequencies, such as the peaks in Figures 14, 15A, or 15B. The one or more sequence motifs may be, for example, any of the sequence motifs in Figure 11. The relative frequencies may be the motif frequencies of Figures 14, 15A, or 15B. The set of one or more sequence motifs can include 1 to 5, 5 to 10, 11 to 15, 15 to 20, or 20 to 25 sequence motifs. The relative frequency of each sequence motif can be determined. In other embodiments, a relative frequency can be determined for multiple sequence motifs, including one or more sets of sequence motifs. Determining the set of sequence motifs is described below.
[0082] At block 1750, one or more relative frequency summary values are determined. Exemplary summary values are described throughout this disclosure and include, for example, an entropy value (motif diversity score or variance), a sum of relative frequencies, and a multidimensional data point corresponding to a vector of counts for a set of motifs (e.g., a vector of 256 counts for 256 possible 4-mer motifs, or 64 counts for 64 possible 3-mer motifs). If one or more sets of sequence motifs include multiple sequence motifs, the summary value may include a sum of the relative frequencies of the set. In some embodiments, the summary value may be an estimate of a histone modification. The level of histone modification may be determined by various types of data, such as the amount of terminal motifs or fragment size.
[0083] At block 1760, the aggregate value is compared to one or more calibration values determined from one or more calibration samples with known fractional concentrations of cell-free DNA fragments from the target tissue type.
[0084] One or more calibration values can be determined by determining aggregate values of sequence motifs for one or more calibration samples. For example, the aggregate value determined from a biological sample can be a first aggregate value determined from one or more first relative frequencies. One or more second relative frequencies of a set of one or more sequence motifs in one or more genomic regions can be determined for each calibration sample of one or more calibration samples. A second aggregate value can be determined for each calibration sample of one or more calibration samples for one or more second relative frequencies. Each of the one or more second aggregate values can thereby be associated with a known concentration of the calibration sample. A calibration value can include one or more second aggregate values. For example, a calibration value can be a point along a line or curve relating a known concentration to the second aggregate value.
[0085] In some embodiments, one or more calibration values may be determined from a function relating known concentrations to a second aggregate value. The first aggregate value may be input into the function to return a fractional concentration. The first aggregate value is then used as a calibration value. Comparing the aggregate value is comparing the aggregate value to the calibration value used in the function and determining that the aggregate value is the same as the calibration value.
[0086] At block 1770, the fractional concentration of cell-free DNA fragments from the target tissue type is determined using the comparison. The fractional concentration may be a known fractional concentration associated with a calibration value, which may have a value close to or equal to the first summary value. In some embodiments, the fractional concentration may be determined from a function or line including one or more calibration values. The function or line may relate the known fractional concentration to one or more calibration values. The fractional concentration of the target tissue type can be used to determine the tissue type and / or a characteristic of the subject from which the biological sample was obtained.
[0087] The classification of the disorder or disease can be determined using the fractional concentration. For example, if the target tissue type is placenta, the method can further include using the fractional concentration to determine the classification or gestational age of the pregnancy-related disorder. The fractional concentration can be compared to a cutoff value determined from a sample from a reference subject having a certain classification of pregnancy-related disorder or a certain gestational age. Pregnancy-related disorders can include preeclampsia, intrauterine growth restriction, invasive placentation and preterm birth, hemolytic disease of the newborn, placental insufficiency, hydrops fetalis, fetal malformations, HELLP syndrome, systemic lupus erythematosus, and other maternal immune disorders. Pregnancy-related disorders can be associated with the fetus or the mother.
[0088] In some embodiments, the classification of the level of cancer can be determined using the fractional concentration, which can be compared to a cutoff value determined from samples from reference subjects with a particular classification of the level of cancer.
[0089] a) Fractional concentration of a second target tissue type In some embodiments, the fractional concentrations of multiple tissue types can be determined. Different tissues can exhibit different amounts of histone modifications in different genomic regions (e.g., as described in Section VA). A biological sample, such as a plasma sample, can contain DNA fragments from different tissues. Thus, the DNA fragments can include fragments associated with histone modifications in different genomic regions. Each genomic region can have a sequence motif associated with a histone modification. The sequence motifs in different genomic regions can be used to determine the fractional concentrations of different tissues in the biological sample. The amount of the sequence motif correlates with the fractional concentration of the tissue. The method can be repeated for a second target tissue to determine the fractional concentration of the second target tissue.
[0090] For example, the above steps may be for a first target tissue type. The one or more genomic regions associated with the first target tissue type may be one or more first genomic regions. The group of sequence reads located in the one or more first genomic regions may be a first group of sequence reads. The histone modifications in the one or more first genomic regions may be first histone modifications. The set of one or more sequence motifs may be one or more first set of sequence motifs. The relative frequencies may be first relative frequencies. The aggregate value may be a first aggregate value. The one or more calibration samples may be one or more first calibration samples. The fractional concentration may be a first fractional concentration.
[0091] The method may further include identifying a second group of sequence reads that map to one or more second genomic regions in a manner similar to block 1720. Each of the one or more second genomic regions may have a second histone modification associated with the second target tissue type. The one or more second genomic regions may be the same as or different from the one or more first genomic regions.
[0092] For each sequence read in the second group of sequence reads, one or more second sequence motifs corresponding to one or more end sequences of the corresponding cell-free DNA fragments can be determined, similar to block 1730.
[0093] A second relative frequency of one or more of the set of one or more second sequence motifs can be determined similarly to block 1740. The set of one or more second sequence motifs can occur at a higher rate in chromatin immunoprecipitation sequencing for a second histone modification associated with one or more second genomic regions than in sequencing without chromatin immunoprecipitation. Sequence motifs that appear more frequently in ChIP sequencing can be used because those sequence motifs can be associated with the second histone modification (similar to FIG. 10). Determining the set of sequence motifs is described below.
[0094] A second aggregate value of one or more second relative frequencies may be determined similar to block 1750 .
[0095] The one or more second aggregate values may be compared to one or more second calibration values in a manner similar to block 1760 .
[0096] One or more second calibration values may be determined from one or more second calibration samples in which the fractional concentration of DNA fragments from the second target tissue type is known. The second fractional concentration of cell-free DNA fragments from the second target tissue type may be determined using a comparison, similar to block 1770.
[0097] b) Determination of sequence motifs A set of one or more sequence motifs can be determined in a manner similar to the procedures described in Figures 3, 10, and 11. In cfChIP sequencing, a first proportion of each of one or more sequence motifs compared to other sequence motifs can be determined. The first proportion can be a ranking, as in Figure 10, or a frequency. The frequency can be determined by the ratio of the raw count of sequence motifs within the set to the count outside the set. In sequencing without chromatin immunoprecipitation, a second proportion of each of one or more sequence motifs compared to other sequence motifs can be determined. The second proportion can be of the same type as the first proportion (e.g., ranking, frequency). Each of the one or more sets of sequence motifs can be identified as having a first proportion higher than the second proportion. Identification can be through the use of a graphical representation (e.g., Figure 10) or through determining the difference between the rankings or frequencies (e.g., Figure 11). Each of the one or more sets of sequence motifs can have a difference greater than a threshold difference. Sequence motifs that are not in the set may have differences below a threshold difference.
[0098] Process 1700 may include additional implementations, such as any single implementation or any combination of implementations described herein, and / or in conjunction with one or more other processes described elsewhere herein.
[0099] 17 illustrates example blocks of process 1700, in some implementations, process 1700 may include additional, fewer, different, or differently arranged blocks compared to the blocks illustrated in FIG 17. Additionally or alternatively, two or more of the blocks of process 1700 may be performed in parallel.
[0100] B. Estimation of feature values of target tissue The values of various features of a target tissue can be estimated using sequence motifs associated with histone modifications. Features can describe the health of the tissue, the age of the tissue, or the level of disease within the tissue. For example, the features to be determined can include a particular gestational age or range (e.g., 8 weeks, 9-12 weeks). In another example, the features to be determined can be the size or nutritional status of an organ corresponding to a particular tissue type.
[0101] FIG. 18 is a flowchart of an example process 1800 associated with estimating a first value of a characteristic of a target tissue. In some implementations, one or more process blocks of FIG. 18 may be performed by a system (e.g., measurement system 5900). In some implementations, one or more process blocks of FIG. 18 may be performed by another device or devices separate from or including the system. Additionally or alternatively, one or more process blocks of FIG. 18 may be performed by one or more components of measurement system 5900, such as assay 5908, assay device 5910, detector 5920, logic system 5930, local memory 5935, external memory 5940, storage device 5945, and / or processor 5950. Process 1800 may include aspects described in process 1700.
[0102] In block 1810, a plurality of sequence reads for the cell-free DNA fragments are received. The plurality of sequence reads include terminal sequences corresponding to the ends of the plurality of cell-free DNA fragments. Block 1810 can be performed in a manner similar to block 1710.
[0103] In block 1820, a group of sequence reads that map to one or more genomic regions are identified, each of the one or more genomic regions having a histone modification associated with the target tissue type. Block 1820 can be performed in a manner similar to block 1720.
[0104] In block 1830, one or more sequence motifs corresponding to one or more end sequences of the corresponding cell-free DNA fragments are determined for each sequence read in the group of sequence reads. Block 1830 can be performed in a manner similar to block 1730.
[0105] In block 1840, the relative frequency of one or more sets of one or more sequence motifs is determined. The set of one or more sequence motifs occurs at a higher rate in chromatin immunoprecipitation followed by sequencing (ChIP-seq) of histone modifications associated with one or more genomic regions than in sequencing without chromatin immunoprecipitation. Block 1840 can be performed in a manner similar to block 1740.
[0106] One or more relative frequency aggregates are determined in block 1850. Block 1850 may be implemented in a manner similar to block 1750.
[0107] In block 1860, the aggregate value is compared to one or more calibration values. The one or more calibration values are determined from one or more calibration samples in which values for the characteristics of the target tissue type are known. The comparison may be performed using a machine learning model, which may be any of the machine learning models described herein. The calibration values may be determined using a machine learning model.
[0108] One or more calibration values may be determined in the same manner as in block 1760, but using a calibration sample in which values for the features of the target tissue type are known. For example, the aggregate value determined from the biological sample may be a first aggregate value determined from one or more first relative frequencies. One or more second relative frequencies of a set of one or more sequence motifs in one or more genomic regions may be determined for each calibration sample of the one or more calibration samples. A second aggregate value may be determined for each calibration sample of the one or more calibration samples for the one or more second relative frequencies. Each of the one or more second aggregate values may thereby be associated with a value of the feature of the calibration sample. The calibration value may include one or more second aggregate values. For example, the calibration value may be a point along a line or curve relating a known value of the feature to the second aggregate value.
[0109] At block 1870, a first value of a feature of the target tissue type is estimated using the comparison. The first value of the feature may be a known first value associated with a calibration value, which may have a value close to or equal to the aggregate value. In some embodiments, the first value of the feature may be determined from a function or line including one or more calibration values. The function or line may relate the known first value to one or more calibration values.
[0110] The target tissue type can be liver or hematopoietic cells. The target tissue type can be fetal tissue. In some embodiments, the biological sample can be obtained from a pregnant female subject, and the target tissue type can be placental tissue. In some embodiments, the target tissue type can be an organ with cancer. The target tissue type can be any organ described herein. The characteristic can be the level of cancer or the nutritional status of the organ. For example, the nutritional status of the organ can be whether the organ is healthy, including any intermediate levels that measure organ health. As another example, the characteristic can be gestational age. In another example, the characteristic to be determined can be the concentration of a particular tissue type (e.g., liver cells) relative to the concentration of another tissue type (e.g., hematopoietic cells).
[0111] In some embodiments, process 1800 may include using size frequencies along with the relative frequencies of sequence motifs. Process 1800 may include using sequence reads to measure the size of cell-free DNA fragments. Process 1800 may further include determining one or more size frequencies of sequence reads for one or more size ranges, which may be any size range described herein. One or more aggregated size frequencies may be determined. The aggregated value may be any value similar to the sum of size frequencies or the aggregated relative frequencies of sequence motifs. In some embodiments, the aggregated value may be an estimate of a histone modification. The level of histone modification may be determined by various types of data, such as the amount of terminal motifs or fragment size. One or more aggregated size frequencies may be compared to a calibration value determined using a calibration sample in which values for the feature of the target tissue type are known. Estimating a first value of the feature may include using a comparison of aggregated size frequencies, as well as a comparison of aggregated relative frequencies of sequence motifs.
[0112] Process 1800 may include additional implementations, such as any single implementation or any combination of implementations described herein, and / or in conjunction with one or more other processes described herein.
[0113] C. Measurement of histone modification levels Sequence motifs can be used to determine the amount of histone modification. As shown in Figures 14, 15A, and 15B, motif frequency can be correlated with the cfChIP-seq signal associated with H3K4me3 signal, which is proportional to the amount of H3K4me3. Therefore, motif frequency can be correlated with the amount of histone modification. In addition, the amount of histone modification in different regions can be used to determine the fractional concentration of multiple tissues in the same sample.
[0114] 1. Exemplary Method for Determining the Abundance of Histone Modifications Using Sequence Motifs 19 is a flowchart of an exemplary process 1900 related to determining the amount of histone modifications in one or more genomic regions. In some implementations, one or more process blocks of FIG. 19 may be performed by a system (e.g., measurement system 5900). In some implementations, one or more process blocks of FIG. 19 may be performed by another device or devices separate from or including the system. Additionally or alternatively, one or more process blocks of FIG. 19 may be performed by one or more components of measurement system 5900, such as assay 5908, assay device 5910, detector 5920, logic system 5930, local memory 5935, external memory 5940, storage device 5945, and / or processor 5950. Process 1900 may include aspects described in process 1700.
[0115] In block 1910, a plurality of sequence reads for the cell-free DNA fragments are received. The plurality of sequence reads include terminal sequences corresponding to the ends of the plurality of cell-free DNA fragments. Block 1910 can be performed in a manner similar to block 1710.
[0116] In block 1920, a group of sequence reads that map to one or more genomic regions are identified, each of the one or more genomic regions having a histone modification associated with the target tissue type. Block 1920 can be performed in a manner similar to block 1720.
[0117] In block 1930, one or more sequence motifs corresponding to one or more end sequences of the corresponding cell-free DNA fragments are determined for each sequence read of the group of sequence reads. Block 1930 can be performed in a manner similar to block 1730.
[0118] In block 1940, the relative frequency of one or more of the sets of one or more sequence motifs is determined, where the sets of one or more sequence motifs occur at a higher or lower rate in chromatin immunoprecipitation followed by sequencing (ChIP-seq) of histone modifications associated with one or more genomic regions than in sequencing without chromatin immunoprecipitation. Block 1940 can be performed in a manner similar to block 1740.
[0119] One or more relative frequency aggregates are determined in block 1950. Block 1950 may be implemented in a manner similar to block 1750.
[0120] In block 1960, the aggregated value is compared to one or more calibration values. The one or more calibration values are determined from one or more calibration samples in which the amount of histone modification is known. The amount of histone modification in the one or more calibration samples may be known from performing ChIP sequencing on each of the one or more calibration samples.
[0121] One or more calibration values may be determined in the same manner as block 1760 or block 1860, but using a calibration sample in which the amount of histone modification is known. For example, the aggregate value determined from the biological sample may be a first aggregate value determined from one or more first relative frequencies. One or more second relative frequencies of a set of one or more sequence motifs in one or more genomic regions may be determined for each calibration sample of one or more calibration samples. A second aggregate value may be determined for each calibration sample of one or more calibration samples for one or more second relative frequencies. Each of the one or more second aggregate values may thereby be related to the amount of the histone modification in the calibration sample. The calibration value may include one or more second aggregate values. For example, the calibration value may be a point along a line or curve relating a known value of a feature to the second aggregate value.
[0122] In block 1970, the amount of histone modification in one or more genomic regions is determined using the comparison. The amount of histone modification can be a known amount with a calibration value, which can have a value close to or equal to the aggregate value. In some embodiments, the amount of histone modification can be determined from a function or line including one or more calibration values. The function or line can relate the known amount of histone modification to one or more calibration values. The amount of histone modification can be of the target tissue type.
[0123] Process 1900 may include additional implementations, such as any single implementation or any combination of implementations described herein, and / or in conjunction with one or more other processes described herein.
[0124] 19 illustrates example blocks of process 1900, in some implementations, process 1900 may include additional, fewer, different, or differently arranged blocks compared to the blocks illustrated in FIG 19. Additionally or alternatively, two or more of the blocks of process 1900 may be performed in parallel.
[0125] 2. Exemplary Methods for Using Fragmentomics Features 20 is a flowchart of an exemplary process 2000 associated with determining the abundance of histone modifications in one or more genomic regions. In some implementations, one or more process blocks of FIG. 20 may be performed by a system (e.g., measurement system 5900). In some implementations, one or more process blocks of FIG. 20 may be performed by another device or devices separate from or including the system. Additionally or alternatively, one or more process blocks of FIG. 20 may be performed by one or more components of measurement system 5900, such as assay 5908, assay device 5910, detector 5920, logic system 5930, local memory 5935, external memory 5940, storage device 5945, and / or processor 5950.
[0126] A plurality of sequence reads of cell-free DNA fragments are received in block 2010. Block 2010 may be performed in a manner similar to block 1710.
[0127] In block 2020, a group of sequence reads that map to one or more genomic regions are identified, each of the one or more genomic regions having a histone modification associated with the target tissue type. Block 2020 can be performed in a manner similar to block 1720.
[0128] In block 2030, a value of a fragmentomics attribute of each cell-free DNA fragment corresponding to each sequence read in the group of sequence reads is determined. The fragmentomics attribute may include fragment size, terminal motif, jagged end (overhang of one strand over the other), terminal nucleotide, topological morphology, and / or nucleosome footprint. The fragmentomics attribute may be any fragmentomics attribute described herein.
[0129] For example, as depicted in Figure 19, a fragmentomics feature can be a sequence motif corresponding to the terminal sequence of the ends of the cell-free DNA fragments, and one or more value ranges are one or more sequence motifs.
[0130] As another example, the fragmentomics attribute may be size, and the one or more value ranges may be one or more size ranges, as described in Section IV.E.
[0131] As an example, a fragmentomics feature can be a topological form, and one or more value ranges can be one or more topological forms. The topological form can be cyclic or linear.
[0132] In one example, the fragmentomics feature is a nucleosome footprint, and the one or more value ranges are one or more nucleosome footprints. The nucleosome footprint represents the binding pattern of nucleosomes to genomic DNA. The spacing between nucleosomes can be a nucleosome footprint value.
[0133] In block 2040, the relative frequencies of one or more cell-free DNA fragments having values of fragmentomics traits within one or more sets of value ranges are determined. The one or more sets of value ranges occur at differential rates in chromatin immunoprecipitation followed by sequencing (ChIP-seq) for histone modifications associated with one or more genomic regions than in sequencing without chromatin immunoprecipitation. The differential rates may be higher or lower and may be statistically significant. Block 2040 may be performed in a manner similar to block 1740, but using one or more value ranges for the fragmentomics traits instead of one or more sequence motifs. In other embodiments, the one or more sets of value ranges determined by sequencing the sample without cell-free chromatin immunoprecipitation are determined by focusing on genomic regions containing differential rates with higher or lower histone modification signals, as previously determined from other reference samples or databases.
[0134] One or more aggregate values of the relative frequencies are determined at block 2050. The aggregate value may be a sum of one or more relative frequencies or a statistical measure of one or more relative frequencies (e.g., mean, median, mode, percentile).
[0135] In block 2060, the aggregate value is compared to one or more calibration values. The one or more calibration values are determined from one or more calibration samples in which the amount of histone modification is known. The amount of histone modification in the one or more calibration samples may be known from performing cfChIP sequencing on each of the one or more calibration samples. The one or more calibration values may be determined in the same manner as block 1960, but using the frequency of one or more value ranges of the fragmentomics feature instead of one or more sequence motifs.
[0136] In block 2070, the amount of the histone modification in the biological sample is determined using the comparison. The amount of the histone modification can be of the target tissue type. Block 2070 can be performed in a manner similar to block 1970.
[0137] The amount of histone modification can be used to determine the fractional concentration of the target tissue, the classification of the level of injury, or the classification of the transplant status of the target tissue type (e.g., as described in process 2000).
[0138] 20 illustrates example blocks of process 2000, in some implementations process 2000 may include additional, fewer, different, or differently arranged blocks compared to the blocks illustrated in FIG 20. Additionally or alternatively, two or more of the blocks of process 2000 may be performed in parallel.
[0139] 3. Determining Fraction Concentrations Using Deconvolution The fractional concentrations of multiple tissue types can be determined through the deconvolution process. Figure 21 illustrates the application of ChIP-seq to determine contributions from different tissues. Graph 2104 is a graph of histone modification signals from ChIP-seq on the y-axis and genomic location on the x-axis. Graphs 2108, 2112, and 2116 show tissue-specific regions for histone modification signals. Graph 2108 shows that region X carries a neutrophil-specific histone modification. Graph 2112 shows that region Y carries a liver-specific histone modification. Graph 2116 shows that region Z carries a monocyte-specific histone modification. The ChIP signals of plasma DNA across these informative genomic regions were compared with the patterns of ChIP signals across different tissues to estimate the proportional DNA contributions associated with H3K27ac from different tissues to plasma. Graph 2120 shows the estimated proportional DNA contributions of different tissues.
[0140] Based on Figure 21, a biological sample containing DNA from multiple tissues may have H3K4me3 cfChIP-seq signals in the same region from multiple tissues. For example, genomic region X, represented in graph 2108, has the highest H3K4me3 signal in neutrophils but lower signals in other tissues (e.g., liver and monocytes). Similarly, genomic region Y, represented in graph 2112, also has different signals across different tissues, including neutrophils, liver, and monocytes. Genomic region Z, represented in graph 2116, also has different signals across different tissues, including neutrophils, liver, and monocytes. Overlapping H3K4me3 signals within the same region may allow for the determination of tissue fractional concentrations.
[0141] A system of linear equations, one for each region, can be solved to determine the fractional concentration of each tissue in a cell-free mixture, such as a plasma sample.
number
[0142] The amount of histone modifications in a target tissue at a particular genomic region (e.g., h 1,A , h 1,B These amounts can be determined from calibration samples. For example, a calibration sample having half target tissue 1 and half target tissue 2 can exhibit a certain ratio of histone modification amounts, which can be expressed as h 1,A and h 1,B It can be used for.
[0143] The number of equations must be equal to or greater than the number of target tissues to solve for fractional concentrations. The number of equations can be equal to the number of genomic regions, and therefore the number of genomic regions can be equal to the number of target tissues. If the sum of the fractional concentrations is known (e.g., the sum is 1), the number of genomic regions can be equal to the number of regions minus 1. Using the histone modification amount in each genomic region measured using sequence motifs, the fractional concentration can be determined by solving the system of equations.
[0144] Thus, in some embodiments, multiple tissue types may have the same or similar sequence motifs associated with histone modifications in the same genomic region, and the fractional concentrations of each of these multiple tissue types may be determined by a deconvolution process, which may involve solving a set of linear or nonlinear equations, such as those described herein.
[0145] The amount of a histone modification may be determined as described in process 1900. In process 1900, the group of sequence reads is a first group of sequence reads. The one or more genomic regions are one or more first genomic regions. The set of one or more sequence motifs is one or more first set of sequence motifs. The one or more relative frequencies are one or more first relative frequencies. The aggregate value is a first aggregate value. The one or more calibration values are one or more first calibration values. The amount of a histone modification is a first amount of a histone modification. An example of a first amount is H in the above equation. A is.
[0146] A second amount of the histone modification in one or more second genomic regions can be determined with respect to a system of linear equations. An example of the second amount is H B The histone modifications can be associated with a first tissue type and a second tissue type at one or more first genomic regions.
[0147] The histone modifications can be associated with a first tissue type and a second tissue type in one or more second genomic regions. For example, the one or more first genomic regions can be regions associated with region X in Figure 21, and the one or more second genomic regions can be regions associated with region Y. As another example, the one or more first genomic regions and the one or more second genomic regions can be regions within the same box (e.g., region X or region Y).
[0148] A second group of sequence reads located in one or more second genomic regions are identified. The identification may be performed in a manner similar to that described in block 1920. Each of the one or more second genomic regions may have histone modifications associated with the first tissue type and the second tissue type. In some embodiments, the histone modifications in the one or more second genomic regions may be different from those in the one or more first genomic regions.
[0149] For each sequence read in the second group of sequence reads, one or more second sequence motifs corresponding to one or more end sequences of the corresponding cell-free DNA fragments are determined. The determining can be performed in a manner similar to that described in block 1930.
[0150] A second relative frequency of one or more of the set of one or more second sequence motifs is determined. The set of one or more second sequence motifs occurs at a higher rate in ChIP-seq for histone modifications associated with one or more second genomic regions than in sequencing without chromatin immunoprecipitation. The determination can be performed in a manner similar to that described in block 1940.
[0151] A second aggregate value of one or more second relative frequencies is determined, which may be performed in a manner similar to that described in block 1950.
[0152] The second aggregate value is compared to one or more second calibration values, which may be performed in a manner similar to block 1960.
[0153] A second amount of the histone modification in one or more second genomic regions is determined using the comparison. The determining can be performed in a manner similar to block 1970.
[0154] The first fractional concentration of the first tissue type and the second fractional concentration of the second tissue type are determined by solving a system of linear or non-linear equations. The system of linear equations can be a set of equations described herein ... A ) and a second amount of histone modification (e.g., H B ) and a parameter (e.g., h ) specifying the relative amount of each histone modification for each tissue type within one or more first genomic regions and one or more second genomic regions. 1,A , h1,B , h 2,A , h 2,B The first fractional concentration can be f1 and the second fractional concentration can be f2.
[0155] A biological sample may contain two or more target tissue types. The methods for determining the fractional concentrations of two target tissue types can be extended to three or more tissue types.
[0156] In embodiments, the histone modification may be associated with a third tissue type at one or more first genomic regions and one or more second genomic regions. The histone modification may be associated with the first tissue type, the second tissue type, and the third tissue type at one or more third genomic regions. The process may include performing similar steps as described for the second tissue type. The process may include determining a third amount of the histone modification (e.g., H) at one or more third genomic regions in the same manner as determining the second amount of the histone modification. m , where m is C). The third fractional concentration of the third tissue type can be determined by solving a system of linear or non-linear equations. The system of linear equations can include parameters for the third amount of the histone modification and the relative amount of each tissue type in the one or more third genomic regions.
[0157] D. Classification of levels of disability The sequence motif can be used to classify the level of a disorder. The disorder can be specific to a particular tissue type or can be applied to a subject. The sequence motif can indicate the amount or presence of a histone modification, and the amount or presence of a histone modification can be associated with a particular level of a disorder. However, the amount or presence of a histone modification may not need to be determined to classify the level of a disorder using the sequence motif.
[0158] Figure 22 shows the ROC curve for distinguishing between patients with and without hepatocellular carcinoma (HCC) using H3K4me3 signals estimated in liver-specific H3K4me3 regions using terminal motifs. Specificity is shown on the x-axis, and sensitivity is shown on the y-axis. Using plasma H3K4me3 ChIP signals estimated by terminal motifs and a cutoff, the AUC was 0.718 for distinguishing between patients with and without HCC. These results demonstrate that ChIP signals of histone modifications estimated by terminal motifs are clinically useful for non-invasive prenatal testing and cancer detection and monitoring.
[0159] FIG. 23 is a flowchart of an example process 2300 related to classifying a level of impairment. In some implementations, one or more process blocks of FIG. 23 may be performed by a system (e.g., measurement system 5900). In some implementations, one or more process blocks of FIG. 23 may be performed by another device or devices separate from or including the system. Additionally or alternatively, one or more process blocks of FIG. 23 may be performed by one or more components of measurement system 5900, such as assay 5908, assay device 5910, detector 5920, logic system 5930, local memory 5935, external memory 5940, storage device 5945, and / or processor 5950. Process 2300 may include aspects described in process 1700.
[0160] In block 2310, a plurality of sequence reads for the cell-free DNA fragments are received. The plurality of sequence reads include terminal sequences corresponding to the ends of the plurality of cell-free DNA fragments. Block 2310 can be performed in a manner similar to block 1710.
[0161] In block 2320, a group of sequence reads that map to one or more genomic regions are identified, each of the one or more genomic regions having a histone modification associated with one or more target tissue types. Block 2320 may be performed in a manner similar to block 1720.
[0162] In block 2330, one or more sequence motifs corresponding to one or more end sequences of the corresponding cell-free DNA fragments are determined for each sequence read of the group of sequence reads. Block 2330 can be performed in a manner similar to block 1730.
[0163] In block 2340, the relative frequency of one or more sets of one or more sequence motifs is determined. The set of one or more sequence motifs occurs at a higher rate in chromatin immunoprecipitation followed by sequencing (ChIP-seq) of histone modifications associated with one or more genomic regions than in sequencing without chromatin immunoprecipitation. Block 2340 can be performed in a manner similar to block 1740.
[0164] One or more relative frequency aggregates are determined in block 2350. Block 2350 may be implemented in a manner similar to block 1750.
[0165] At block 2360, the aggregate value is compared to one or more calibration values determined from one or more calibration samples with known classifications of the level of impairment.
[0166] One or more calibration values may be determined in the same manner as block 1760, block 1860, or block 1960, but using a calibration sample for which the classification of the level of the disorder is known. For example, the aggregate value determined from the biological sample may be a first aggregate value determined from one or more first relative frequencies. One or more second relative frequencies of a set of one or more sequence motifs in one or more genomic regions may be determined for each calibration sample of the one or more calibration samples. A second aggregate value may be determined for each calibration sample of the one or more calibration samples for the one or more second relative frequencies. Each of the one or more second aggregate values may thereby be associated with a classification of the level of the disorder. The calibration value may include one or more second aggregate values. For example, the calibration value may be a point along a line or curve relating a known classification of the level of the disorder to the second aggregate value.
[0167] At block 2370, a classification of the level of impairment is determined using the comparison. The classification of the level of impairment may be a known classification having a calibrated value, which may have a value close to or equal to the aggregate value. In some embodiments, the classification of the level of impairment may be determined from a function or line including one or more calibrated values. The function or line may relate the known classification to one or more calibrated values. In some embodiments, the classification may be a level of abnormality.
[0168] The disorder may be of a target tissue type. The disorder may be a cancer of a target tissue type. The cancer may include hepatocellular carcinoma (HCC), colorectal cancer (CRC), or any cancer described herein. In some embodiments, the disorder is a pregnancy-related disorder. The disorder may be a hematological disorder. The disorder may be any disorder described herein.
[0169] In an embodiment, the process 2300 may include using a size frequency as described in the process 1800.
[0170] Process 2300 may include additional implementations, such as any single implementation or any combination of implementations described herein, and / or in conjunction with one or more other processes described herein.
[0171] 23 illustrates example blocks of process 2300, in some implementations, process 2300 may include additional, fewer, different, or differently arranged blocks compared to the blocks illustrated in FIG 23. Additionally or alternatively, two or more of the blocks of process 2300 may be performed in parallel.
[0172] IV. Predicting histone modifications using size information A. Size information for estimating ChIP signals The size information of plasma DNA can be used to detect and quantify histone modifications present in plasma DNA molecules. Similar to the relationship between cfDNA terminal motif information and histone modification levels, the size information of cfDNA molecules can be influenced by histone modification levels, i.e., epigenetic status. We analyzed the size information of cfDNA molecules within regions of interest. These regions involved in histone modifications can be grouped into different categories based on the magnitude of their ChIP signal. For example, regions were first sorted by ChIP signal magnitude and then empirically classified into nine categories (e.g., Figure 4). After obtaining the region categories with different H3K27ac ChIP signals, we were able to compare the DNA size information of the sequencing results of plasma DNA without immunoprecipitation in the different region categories.
[0173] Figures 24A, 24B, and 24C show the percentage of cfDNA molecules with a particular size in regional categories with different levels of H3K27ac signal. The x-axis represents the size in base pairs from sequencing of plasma DNA without H3K27ac-based precipitation. The y-axis represents the percentage of cfDNA molecules with that size. Different colored lines in each graph represent different regional categories with H3K27ac ChIP signal. Figure 24A shows the size range of 50-140 bp. Figure 24B shows the size range of 150-200 bp. Figure 24C shows the size range of 250-350 bp. As shown in Figures 24A-24C, the size profile varies depending on the intensity of the ChIP signal. For example, the higher the ChIP signal, the more DNA molecules in the 270-300 bp range can be observed. Additionally, size differences showed different trends for different size ranges. For example, for sizes of approximately 165-200 bp, the higher the ChIP signal, the fewer DNA molecules there are. For sizes of approximately 60-100 bp, the higher the ChIP signal, the more DNA molecules there are. Therefore, it is possible to estimate the histone modifications in plasma DNA based on the size information of plasma DNA molecules without immunoprecipitation.
[0174] Figures 25A, 25B, and 25C show that the correlation between size and ChIP signal of histone modifications can be generalized to other histone modifications (e.g., H3K27ac). The x-axis is the cumulative size frequency as a percentage of a particular size fragment for plasma DNA without immunoprecipitation. The y-axis is log 10Figure 25 shows H3K27ac ChIP signals at different scales. Figure 25A is for the 50-140 bp fragment. Figure 25B is for the 150-200 bp fragment. Figure 25C is for the 250-350 bp fragment. For plasma DNA without immunoprecipitation, across nine categories, the percentage of cfDNA molecules within the 50-140 bp and 250-350 bp size ranges was positively correlated with the log-transformed ChIP signal obtained from the ChIP-seq data, with a Pearson's r of 0.99 (P value: <0.0001) (Figure 25A) and 0.99 (P value: <0.0001) (Figure 25C). The percentage of cfDNA molecules within the size range of 150-200 bp was negatively correlated with the log-transformed ChIP signal (Pearson's r: -0.99, P-value: <0.0001) (Figure 25B).
[0175] Figures 26A and 26B show the use of size information to estimate histone modifications in plasma DNA from the results of sequencing plasma DNA without immunoprecipitation. Figure 26A shows the construction of a recalibration equation using the percentage of cfDNA molecules in a certain size range in nine categories and the level of H3K4me3 ChIP signal. Step 2604 shows that regions involved in H3K4me3 were grouped into different categories based on the magnitude of the ChIP signal. As shown in Figure 26A, size information from each region category (e.g., the percentage of cfDNA molecules from the results of sequencing plasma DNA without immunoprecipitation within the 250-350 bp size range) can be used to determine the correlation with H3K4me3 ChIP signal. In step 2608, a recalibration equation can be determined based on the correlation between fragment size and ChIP signal (on a logarithmic scale). A linear equation is shown as an example of a recalibration equation, but nonlinear equations can also be used.
[0176] Figure 26B shows how the recalibration formula can be used to estimate (i.e., estimate) ChIP signals in other regions (e.g., placenta-specific H3K4me3 regions) according to their corresponding size information. In step 2612, plasma DNA is sequenced without immunoprecipitation. In step 2616, molecules overlapping with tissue-specific (e.g., placenta) H3K4me3 regions are identified. In step 2620, the percentage of molecules within a specific size range (e.g., 250-350 bp) in H3K4me3-based immunoprecipitated plasma DNA is calculated. The size information is input into the recalibration formula, and in step 2624, H3K4me3 ChIP signals are estimated at tissue-specific (e.g., placenta) H3K4me3 regions.
[0177] Figures 27A, 27B, and 27C show the correlation between the percentage of cfDNA molecules within a size range and the log-transformed H3K4me3 ChIP signal. The x-axis is the cumulative size frequency as a percentage of a particular size fragment for plasma DNA without immunoprecipitation. The y-axis is log 10 Figure 27 shows H3K4me3 ChIP signals at different scales. Figure 27A is for the 50-140 bp fragment. Figure 27B is for the 150-200 bp fragment. Figure 27C is for the 250-350 bp fragment. For plasma DNA without immunoprecipitation, across nine categories, the percentage of cfDNA molecules within the 50-140 bp and 250-350 bp size ranges positively correlated with the log-transformed ChIP signals obtained from ChIP-seq data, with a Pearson's r of 0.99 (P value: <0.0001) (Figure 27A) and 0.99 (P value: <0.0001) (Figure 27C). The percentage of cfDNA molecules within the 150-200 bp size range was negatively correlated with the log-transformed ChIP signal (Pearson's r: -0.99, P-value: <0.0001) (Figure 27B). The results indicate that the fragment size pattern can be used to estimate histone modifications in plasma DNA molecules (referred to as the estimated ChIP signal).
[0178] B. Putative ChIP signal and fetal fraction Furthermore, we used a linear regression model to construct a model (i.e., a recalibration equation) for estimating H3K4me3 ChIP signals in a region or set of regions of interest. As an example, we trained a model for each sample to estimate ChIP signals based on a size range of 250-350 bp. That is, Y = aX + b, where "Y" represents the log-transformed ChIP signal, and "X" represents the percentage of cfDNA molecules within the 250-350 bp size range from the specific genomic region or set of regions of interest for which histone modifications are being determined. "a" and "b" are the slope and intercept, respectively. In one embodiment, we determined the percentage of cfDNA molecules within the 250-350 bp size range from those placenta-specific regions in terms of H3K4me3. We analyzed 30 plasma DNA samples from pregnant women. The 250-350 bp size range was chosen for illustrative purposes. Other size ranges may also be used. The size ranges may be selected using machine learning models.
[0179] Figures 28A and 28B show an evaluation of the performance of estimated H3K4me3 ChIP signals in placenta-specific H3K4me3 regions for fetal DNA fraction estimation. The x-axis shows fetal DNA fraction as a percentage determined by the SNP-based approach. In Figure 28A, the y-axis is the estimated H3K4me3 ChIP signal using a size range of 250-350 bp. Using the size metric, the estimated H3K4me3 ChIP signal correlated with fetal DNA fraction (Pearson's r: 0.62, P-value: <0.0001).
[0180] In Figure 28B, the y-axis is cumulative size frequency as a percentage of 250-350 bp fragments. There was no significant correlation between the percentage of plasma DNA within the 250-350 bp size range (Pearson's r: -0.31, P-value: 0.096). These results in Figures 28A and 28B indicate that the use of estimated ChIP signals for plasma DNA samples without immunoprecipitation can enable analysis of tissue of origin for plasma DNA molecules.
[0181] Furthermore, we used a linear regression model to construct a model (i.e., a recalibration equation) for estimating H3K27ac ChIP signals in a region or set of regions of interest. As an example, we trained a model for each sample to estimate ChIP signals based on a size range of 250-350 bp. That is, Y = aX + b, where "Y" represents the log-transformed ChIP signal, and "X" represents the percentage of cfDNA molecules within the size range of 250-350 bp from a specific genomic region or set of regions of interest for which histone modifications are determined. "a" and "b" represent the slope and intercept, respectively. In one embodiment, we determined the percentage of cfDNA molecules within the size range of 250-350 bp from these placenta-specific regions in terms of H3K27ac. We analyzed 30 plasma DNA samples from pregnant women.
[0182] Figure 29 is a graph evaluating the performance of the estimated H3K27ac ChIP signal in placenta-specific H3K27ac regions for determining fetal DNA fraction. The x-axis is the fetal DNA fraction as a percentage, as determined by the SNP-based approach. The y-axis is the estimated H3K27ac ChIP signal using a size range of 250-350 bp. Based on such a size metric, the estimated H3K27ac ChIP signal showed a higher correlation with fetal DNA fraction (Pearson's r: 0.95, P-value: <0.0001) compared to H3K4me3-based analysis (Pearson's r: 0.62, P-value: <0.0001) (Figure 28A). These results highlight that different types of histone modifications can be used to determine the tissue of origin of plasma DNA molecules via the ChIP signal of histone modifications estimated by cfDNA size information.
[0183] We analyzed different size ranges to estimate H3K27ac ChIP signals and correlated the estimated H3K27ac ChIP signals with tissue DNA fractions determined by a SNP-based approach. We analyzed 30 plasma DNA samples from pregnant women. Size ranges of 50-150 bp, 160-225 bp, and 230-350 bp were used for illustrative purposes. Other size ranges may also be used in some other embodiments.
[0184] Figure 30 is a graph showing how well the size ranges, with and without considering histone modification levels, correlated to fetal DNA fraction. The y-axis shows the three size ranges tested. The x-axis shows the Pearson correlation coefficient. For each size range, two different bars are shown. The upper bar (gray) of each pair shows the Pearson correlation coefficient when using raw size frequencies. The lower bar (black) of each pair shows the Pearson correlation coefficient when using estimated H3K27ac signal levels in placenta-specific H3K27ac regions.
[0185] As shown in Figure 30, fetal DNA fractions determined by the SNP-based approach were strongly correlated with estimated H3K27ac signal levels in placenta-specific H3K27ac regions in the 230-350 bp size range (Pearson's r: 0.96, P-value: <0.0001). In contrast, no such correlation was observed with the raw cumulative size frequency (Pearson's r: -0.25, P-value = 0.18). Comparisons were also performed for other size ranges. For all tested size ranges, estimated H3K27ac levels in placenta-specific H3K27ac regions showed substantially higher correlations with fetal DNA fractions (Pearson's r: 0.76-0.96) compared with the respective raw cumulative size frequencies (Pearson's r: -0.25-0.53). In addition, the H3K27ac ChIP signal estimated based on molecules with a size range of 230–350 bp showed the best performance (Pearson's r = 0.96) compared to other tested size ranges (Pearson's r: 0.76).
[0186] C. Inferred ChIP signals and cancer In one embodiment, we investigated whether the predicted ChIP signals of histone modifications from plasma DNA without immunoprecipitation are useful for cancer detection. We analyzed samples from 34 patients with hepatocellular carcinoma (HCC), 17 patients with chronic hepatitis B virus (HBV), and 8 healthy controls.
[0187] Figures 31A and 31B are graphs showing the use of estimated H3K4me3 ChIP signals based on liver-specific H3K4me3 regions for HCC detection. H3K4me3 ChIP signals were estimated using the cumulative frequency of molecules within the size range of 250-350 bp. Figure 31A shows a boxplot of estimated H3K4me3 ChIP signals (y-axis) against subject type (x-axis). For liver-specific regions, estimated H3K4me3 ChIP signals were significantly higher in subjects with HCC (median: 0.21, range: 0-2.90) compared with subjects without HCC (median: 0.09, range: 0-5.36) (P-value: 0.015, Mann-Whitney U test).
[0188] Figure 31B is a receiver operating characteristic (ROC) curve. ROC analysis revealed that an AUC of 0.686 can be achieved in distinguishing subjects with HCC cancer from those without. These results indicate that the estimated ChIP signal can be used for cancer detection. This approach eliminates the need for immunoprecipitation assays before sequencing, thus reducing costs and experimental time, and can easily incorporate other technologies such as whole-genome random sequencing or targeted sequencing, or whole-genome random sequencing or targeted bisulfite sequencing.
[0189] Figures 32A and 32B show the use of H3K27ac ChIP signals estimated based on the H3K27ac region for HCC detection. The H3K27ac ChIP signals were estimated using the cumulative frequency of molecules within the size range of 250 to 350 bp. Figures 32A and 32B are the same as Figures 31A and 31B, respectively, except that the H3K27ac region was used instead of the H3K4me3 region. The use of the H3K27ac-related estimated ChIP signal improved the classification ability when distinguishing patients with and without HCC, increasing the AUC from 0.686 (Figure 31B) to 0.738 (Figure 32B).
[0190] Figure 33 is a graph showing how size selection affects performance for distinguishing patients with cancer from healthy controls. Figure 33 is an ROC curve with sensitivity on the y-axis and specificity on the x-axis. The ROC curve is for distinguishing subjects with intermediate and advanced hepatocellular carcinoma (HCC) from subjects without HCC by estimated H3K27ac ChIP signal for liver-specific regions. The black line is for molecules within the 230-350 bp size range. The gray line is for molecules within the 50-150 bp size range.
[0191] ROC analysis revealed that H3K27ac ChIP signals estimated using the cumulative frequency of molecules within the size range of 230–350 bp in the liver-specific H3K27ac region achieved a significantly larger area under the receiver operating characteristic curve (AUC) of 0.934 in distinguishing patients with intermediate- and advanced-stage HCC from those without HCC compared with those within the size range of 50–150 bp (AUC: 0.586) ( P = 0.001, DeLong test).
[0192] D. Inferred ChIP signals and transplantation Figure 34 is a graph showing the correlation between estimated H3K27ac ChIP signal in liver-specific H3K27ac regions and donor DNA fraction. H3K27ac ChIP signal was estimated using the cumulative frequency of molecules within the 250-350 bp size range. The y-axis shows the estimated H3K27ac ChIP signal. The x-axis shows the donor DNA fraction as a percentage. We used liver-specific regions to estimate H3K27ac ChIP signal in plasma DNA from liver transplant patients. The graph shows a high correlation between liver contribution determined by estimated ChIP signal of histone modifications in liver-specific regions according to embodiments of the present disclosure and donor DNA fraction using a SNP-based approach (Pearson's r: 0.9, P-value: <0.0001). The data demonstrate that estimated H3K27ac ChIP signal for liver-specific regions may enable monitoring of subjects undergoing organ transplantation.
[0193] Additionally, we analyzed the results of sequencing plasma DNA without immunoprecipitation for a cohort of 14 liver transplant patients. Size ranges of 50-150 bp, 160-225 bp, and 230-350 bp were used for illustrative purposes. Other size ranges may also be used in some other embodiments.
[0194] Figure 35 is a graph showing how well the size ranges, taking histone modification levels into account and not, correlate with the donor DNA fraction determined by the SNP-based approach. The y-axis shows the three size ranges tested. The x-axis shows the Pearson correlation coefficient. For each size range, two different bars are shown. The upper bar (gray) of each pair shows the Pearson correlation coefficient when using raw size frequencies. The lower bar (black) of each pair shows the Pearson correlation coefficient when using estimated H3K27ac signal levels in liver-specific H3K27ac regions.
[0195] As shown in Figure 35, the highest correlation was observed between donor DNA fraction and H3K27ac values estimated in liver-specific H3K27ac regions by molecules with a size range of 230-350 bp (Pearson's r: 0.91, P-value: <0.0001).
[0196] E. Exemplary Method for Determining Histone Modifications Using Size
[00130] Figure 36 is a flowchart of an exemplary process 3600 relating to determining the amount of histone modifications in one or more genomic regions using fragment size. In some implementations, one or more process blocks of Figure 36 may be performed by a system (e.g., measurement system 5900). In some implementations, one or more process blocks of Figure 36 may be performed by another device or devices separate from or including the system. Additionally or alternatively, one or more process blocks of Figure 36 may be performed by one or more components of measurement system 5900, such as assay 5908, assay device 5910, detector 5920, logic system 5930, local memory 5935, external memory 5940, storage device 5945, and / or processor 5950.
[0197] At block 3610, a plurality of sequence reads of the cell-free DNA fragments are received. The plurality of sequence reads may be obtained by random massively parallel sequencing. The plurality of sequence reads may be obtained using paired-end sequencing.
[0198] At block 3620, a group of sequence reads that map to one or more genomic regions are identified, each of the one or more genomic regions having a histone modification associated with one or more target tissue types. Block 3620 may be performed in a manner similar to block 1720.
[0199] In block 3630, the size of each cell-free DNA fragment corresponding to each sequence read in the group of sequence reads is measured. Fragment size can be measured using paired-end sequencing, aligning the sequence to the genome, and then estimating the size from the genomic coordinates of the aligned sequence. In some embodiments, fragment size can be measured by sequencing the entire fragment and then determining the size from the sequence.
[0200] At block 3640, one or more relative frequencies of cell-free DNA fragments having sizes within a set of one or more size ranges are determined. The set of one or more size ranges may occur at a differential rate in chromatin immunoprecipitation followed by sequencing (ChIP-seq) for histone modifications associated with one or more genomic regions than in sequencing without chromatin immunoprecipitation. The differential rate may be higher or lower, and may be a statistically significant amount. The one or more size ranges may include 50-100 bp, 100-150 bp, 150-200 bp, 200-250 bp, 250-300 bp, 300-350 bp, 350-400 bp, 400-450 bp, 450-500 bp, greater than 500 bp, or any combination thereof.
[0201] An aggregate value of one or more relative frequencies is determined at block 3650. The aggregate value may be a sum of one or more relative frequencies or a statistical measure of one or more relative frequencies (e.g., mean, median, mode, percentile).
[0202] In block 3660, the aggregate value is compared to one or more calibration values. The one or more calibration values are determined from one or more calibration samples in which the amount of histone modification is known. The amount of histone modification in the one or more calibration samples may be known from performing cfChIP sequencing on each of the one or more calibration samples. The one or more calibration values may be determined in the same manner as block 1960, but using the frequency of one or more size ranges instead of one or more sequence motifs.
[0203] In block 3670, the amount of the histone modification in the biological sample is determined using the comparison. The amount of the histone modification can be of the target tissue type. Block 3670 can be performed in a manner similar to block 1970.
[0204] The amount of histone modification can be used to determine the fractional concentration of the target tissue, the classification of the level of injury, or the classification of the transplant status of the target tissue type. In addition to size range, the amount of histone modification can be determined using sequence motifs, fragmentomics characteristics, or any other technique.
[0205] In some embodiments, the amount of the histone modification may be compared to one or more second calibration values. The one or more second calibration values may be determined from one or more second calibration samples in which the fractional concentration of the target tissue type and the amount of the histone modification are known. The fractional concentration of the target tissue type may be determined using a comparison of the amount of the histone modification to one or more second calibration values.
[0206] In some embodiments, the amount of histone modification can be compared with one or more third calibration values. The one or more third calibration values can be determined from one or more third calibration samples in which the level of the disorder and the amount of the histone modification are known. The classification of the level of the disorder is determined using one or more third calibration values. The disorder can be any disorder described herein.
[0207] In some embodiments, the amount of histone modification is compared to one or more fourth calibration values. The one or more fourth calibration values can be determined from one or more fourth calibration samples in which the transplant status and the amount of histone modification are known. A classification of the transplant status of the target tissue type is determined using the one or more fourth calibration values. The classification of the transplant status includes whether the transplanted organ is rejected by the subject.
[0208] Although Figure 36 illustrates example blocks of process 3600, in some implementations, process 3600 may include additional, fewer, different, or differently arranged blocks compared to the blocks illustrated in Figure 36. Additionally or alternatively, two or more of the blocks of process 3600 may be performed in parallel.
[0209] V. Tissue Contributions Inferred from Histone Modifications The characteristic size profile of cfDNA shows a peak frequency at approximately 166 bp, with smaller molecules forming a series of peaks with a periodicity of 10 bp (Lo et al. Sci Transl Med. 2010;2:61ra91). Such a size pattern of plasma DNA fragments suggests the presence of histone proteins bound to cfDNA molecules. One recent study used cell-free chromatin immunoprecipitation followed by sequencing (cfChIP-seq) to reveal the presence of histone modifications associated with cfDNA molecules in plasma (Sadeh et al. Nat Biotechnol. 2021;39:586-598). However, Sadeh et al.'s study did not provide any approach for estimating the percentage contribution of chromatin modifications from various tissues / organs.
[0210] Sadeh et al. analyzed the average number of reads per kilobase across genomic regions associated with tissue-specific histone modifications for a tissue as a signal indicating contribution from that tissue. Tissue-specific regions estimated from reference tissues were considered as an independent factor when analyzing these signals (Sadeh et al. 2021). One limitation of the method described by Sadeh et al. is that it cannot accurately estimate DNA contribution from a tissue if the tissue lacks tissue-specific histone modifications or if the number of regions showing tissue-specific histone modifications is insufficient in the tissue. Sadeh's method relies on the absolute signal of histone modifications in plasma with respect to tissue-specific regions. However, the relative intensity of the histone modification signal in each reference tissue is not taken into account in this approach by Sadeh et al., which is likely to lead to inaccurate or impossible analysis.
[0211] For example, the reads per kilobase within a genomic region associated with a tissue's histone modification may be governed by at least two factors: the percentage of DNA (including DNA unrelated to histone modifications) contributed by that tissue, and the level of histone modifications present in that tissue. Analysis adjusted for the level of histone modifications present in that tissue is important for tissue contribution analysis based on histone modifications. Sadeh et al. attempted to analyze the percentage contribution from the liver using linear regression. Plasma DNA from healthy subjects was considered to have 0% liver contribution, and DNA from liver tissue was considered to have 100% liver contribution. The difference in histone modifications between liver tissue and plasma DNA from healthy subjects was used to determine the liver contribution in other plasma DNA samples (Sadeh et al. Nat Biotechnol. 2021). Such analyses did not use histone modification signals from more than one tissue. Plasma DNA contains contributions from various tissues, and liver contribution to plasma may vary between healthy subjects. Therefore, the assumptions of linear regression analysis may not hold true under such circumstances.
[0212] Therefore, the contributions from two or more tissues analyzed could not be accurately estimated using Sadeh et al.'s approach. The intensity of the histone modification signal from each tissue is important when quantitatively analyzing the signal present in plasma cfDNA. The intensity of the histone modification signal can refer to the percentage of cells in a tissue that have the histone modification of interest, which can be measured by the depth of coverage of sequencing reads in ChIP-seq. By not using histone modification signals across different tissues, the approach would significantly worsen its performance when determining the contribution of cfDNA bearing histone modifications to plasma from different tissues.
[0213] In this disclosure, we developed an approach to estimate the percentage contribution of each cell type or tissue by comparing the relative signal of histone modifications in plasma DNA with signals from a reference tissue, herein referred to as tissue mapping of plasma DNA by histone modifications. In one embodiment, such comparisons consider the modified histone signals from various tissues as covariates and deconvolve the percentage contributions from various tissues to plasma using, for example, but not limited to, quadratic programming, non-negative least squares (NNLS), etc. Sun et al. demonstrated that by comparing the methylation signals of plasma DNA with those of various tissues, the use of quadratic programming enabled estimation of the percentage contribution of DNA molecules to plasma across tissues (Sun et al., Proc Natl Acad Sci USA. 2018;115:E5106). However, histone modifications can occur in the amino acid sequence of histone proteins, resulting in signal characteristics of modified signals that differ from those of DNA methylation. Signal processing procedures used in DNA methylation analysis could not be used for modified histones. Histone modifications involve post-translational modifications of histone proteins, which affect their interaction with DNA. In contrast, DNA methylation is a biochemical process in which DNA bases, usually cytosine, are enzymatically methylated at the 5-carbon position. Histone modifications and methylation involve different types of biochemical mechanisms. In some embodiments of the present disclosure, the contribution of histone modifications to plasma could be estimated by comparing the number of DNA immunoprecipitated via one or more antibodies of interest with measurements of counterparts across various reference tissues. In contrast to the approach used by Sadeh et al., which only provided information on tissue-specific histone modifications, the approach presented in this disclosure may be able to utilize both tissue-specific and tissue-variable histone modifications.
[0214] A. Tissue mapping of plasma DNA by histone modifications In embodiments, the percentage contribution of DNA from various cell types to plasma can be determined by comparing the histone modification profile of plasma DNA with the histone modification profiles from several organs, tissues, or cells. For example, H3K27ac ChIP-seq can be applied to several tissues, including but not limited to neutrophils, megakaryocytes, T cells, B cells, erythrocytes, monocytes, natural killer cells, or cells from the liver, colon, adipose tissue, brain, pancreas, placenta, heart, lung, kidney, spleen, bladder, stomach, etc. Informative genomic regions bearing tissue-specific histone modifications (e.g., H3K27ac) can be determined. Informative genomic regions refer to regions that preferentially enrich for a specific histone modification (e.g., H3K27ac) in a specific tissue (e.g., the liver) and are relatively depleted of such modifications in other tissues. Such regions can be referred to as tissue-specific histone modification regions (e.g., tissue-specific H3K27ac regions). In some embodiments, informative genomic regions refer to regions that show variable signals of a particular histone modification (e.g., H3K27ac) across tissues of interest. Variable signals can be defined as a coefficient of variation (CV) of histone signals exceeding, but not limited to, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, etc., where the difference between the maximum and minimum modified histone signals exceeds a certain cutoff, such as, but not limited to, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 50, 100, 500, 1000, 5000, 10,000 reads per kilobase. Such regions can be defined as tissue-variable histone modification regions (e.g., tissue-variable H3K27ac regions). As previously mentioned, Figure 3 illustrates the application of ChIP-seq to determine contributions from different tissues.
[0215] Because different pathological or physiological conditions can alter the chromatin state in certain cell types, we speculated that analysis of histone modifications on cfDNA molecules would enable the non-invasive detection and monitoring of diseases, such as fetal abnormalities in pregnant women, cancer, autoimmune diseases, the presence of transplant rejection, and blood disorders.
[0216] B. Example of tissue mapping of plasma DNA by histone modifications The signals of putative histone modifications can be used to determine the fetal DNA fraction, determine specific tissue contributions to a sample, classify subjects as pregnant or not pregnant, and classify subjects with a possible disorder (e.g., cancer).
[0217] 1. Biological samples from pregnant women We recruited 19 pregnant women with a median gestational age of 38 weeks. Plasma was isolated from whole blood within 6 hours of sample collection through sequential centrifugation steps (10 min at 1,600 g, followed by a further 10 min of re-centrifugation of the plasma fraction at 16,000 g). Plasma can be stored at -80°C. We used two types of histone modifications (H3K27ac and H3K4me3) as examples. Antibody-conjugated beads were incubated with plasma overnight by rotation at 4°C, washed with wash buffer, and immunoprecipitated DNA was ligated to barcoded adapters on the beads. DNA was eluted and subsequently amplified by PCR. DNA libraries were multiplexed with several other libraries using an Illumina platform (e.g., Nextseq 500 or NovaSeq 6000) to obtain a median of 4.3 million paired-end reads (range: 0.10–30.73). We performed H3K27ac ChIP-seq on 19 pregnant, 13 non-pregnant, and 12 hematological disease samples (10 beta-thalassemia major, 1 iron deficiency anemia, and 1 aplastic anemia samples). Additionally, we performed H3K4me3 ChIP-seq on 12 pregnant women, 4 non-pregnant healthy subjects, and 4 hematological disease patients (2 beta-thalassemia major, 1 iron deficiency anemia, and 1 aplastic anemia samples).
[0218] The fetal DNA fraction in maternal plasma from each pregnant woman was calculated based on a single nucleotide polymorphism (SNP)-based approach (Lo et al. Sci Transl Med. 2010;2:61ra91). Genotypes for maternal buffy coat and placental tissue samples were obtained using microarray-based genotyping technology (Illumina Infinium Omni 2.5-8 arrays) to identify informative SNPs (i.e., those in which the mother was homozygous (represented as the AA genotype) and the fetus was heterozygous (represented as the AB genotype)). Fetal-specific DNA fragments were identified according to the DNA fragments carrying the fetal-specific allele at the informative SNP site. In this scenario, the B allele was fetal-specific, and DNA fragments carrying the B allele were presumed to be derived from fetal tissue. The number of fetal-specific molecules (p) carrying the fetal-specific allele (B) was determined. The number of molecules (q) carrying the shared allele (A) was determined. The fetal DNA fraction across all cell-free DNA samples is calculated by 2p / (p+q)*100%.
[0219] For illustrative purposes, ChIP-seq data for various tissues were obtained from public databases. Public databases used herein include, but are not limited to, the Blueprint Project (blueprint-epigenome.eu / ), the ENCODE Project (encodeproject.org / ), and the Roadmap Project (roadmapepigenomics.org / ). In total, we obtained H3K27ac ChIP-seq results from 18 tissue types, including, but not limited to, neutrophils, monocytes, B cells, T cells, natural killer cells, erythroid cells, and megakaryocytes, liver, brain, pancreas, placenta, heart, colon, lung, adipose tissue, kidney, spleen, and bladder, with a median of 22.5 million paired-end / single-end reads (range: 12-45 million). Additionally, we obtained H3K4me3 ChIP-seq data from 19 tissues, including but not limited to neutrophils, monocytes, B cells, T cells, natural killer cells, erythroid cells, megakaryocytes, liver, brain, pancreas, placenta, heart, colon, lung, adipose, kidney, spleen, bladder, and stomach, with a median of 25 million paired-end reads (range: 7–32 million).
[0220] Based on ChIP-seq data from various tissues, we determined informative genomic regions carrying tissue-specific histone modifications. In one embodiment, several genomic regions known to be enriched with specific types of histone modifications can be analyzed. For example, H3K4me3 is known to occur preferentially in regions near transcription start sites (i.e., promoter regions). Therefore, ChIP signals can be determined across regions near transcription start sites (TSSs). In one embodiment, the ChIP signal for a region of interest can be determined by the percentage of sequencing reads that overlap with such a region among all mapped reads. In another embodiment, the ChIP signal for a region of interest can be determined by the percentage of sequencing reads that overlap with such a region among all mapped reads relative to all regions of interest. ChIP signals are adjusted for GC bias and mapping bias and expressed as fragments per kilobase per million fragments analyzed (i.e., FPKM).
[0221] In one embodiment, according to the ChIP signals identified from several tissues / organs, the human reference genome is classified into regions of interest (ROIs) containing a specific histone modification (e.g., H3K27ac) and regions lacking such histone modifications (background regions). Plasma DNA ChIP-seq reads present in the background regions were considered background noise because they may be due to nonspecific antibody (Ab) binding during the experimental process. The raw ChIP signal of an ROI was determined as the number of fragments whose 5' ends fell within the ROI. In some embodiments, the raw ChIP signal of an ROI was determined as the number of fragments whose molecules overlapped with the ROI by at least one or more nucleotides. The raw signal of an ROI can be subtracted by the background noise over the background region surrounding the ROI being examined.
[0222] Using H3K27ac as an example, we divided the genome into non-overlapping 5-Mb windows. For each 5-Mb window, we calculated the raw signal of the H3K27ac-bound ROIs (N regions) according to the ChIP results published in the ENCODE and Blueprint projects. The remaining regions (M regions) were considered background regions to determine noise. A Poisson distribution can be used to estimate the average sequencing depth per kilobase (kb) across the M background regions, referred to as the estimated background noise. The raw ChIP signal across the N ROIs, subtracted by the estimated background noise (i.e., noise-subtracted ChIP signal), is used for downstream analysis. To minimize the effect of sequencing depth on the comparison of ChIP signals between samples, we used sequencing reads from those regions shown to be bound by H3K27ac across various samples to determine a scaling factor for sequencing depth across samples. The noise-subtracted ChIP signal is adjusted by the corresponding scaling factor for sequencing depth. In one embodiment, the ChIP signal can be further expressed as fragments per kilobase per million (FPKM). In some embodiments, several overlapping windows can be used to estimate background noise. The window size can be, but is not limited to, 10 kb, 50 kb, 100 kb, 500 kb, 1 Mb, 2 Mb, 3 Mb, 4 Mb, 5 Mb, 10 Mb, etc.
[0223] Regions carrying tissue-specific histone modifications (i.e., tissue-specific regions) can be determined using the following criteria: 1. The first tissue of interest had the highest ChIP signal for each tissue-specific region across all tissues analyzed, with a normalized ChIP signal greater than 15. 2. The ratio of ChIP signals in the log2 scale in the tissue-specific region between the first tissue and the second tissue with the second highest ChIP signal is greater than 3. As a result, we identified a total of 4,245 tissue-specific regions for H3K27ac and 807 tissue-specific regions for H3K4me3.
[0224] Figure 37 shows a table of regions of tissue-specific histone modifications. Column 1 lists the tissue / cell type. Column 2 lists the number of regions in that tissue that show the H3K27ac modification. Column 3 lists the number of regions in that tissue that show the H3K4me3 modification.
[0225] In one embodiment, the selected regions were not necessarily limited to tissue-specific regions. Regions showing high variability in histone modification signals across a panel of tissues of interest for analysis (tissue variable regions) can be used. These regions can be determined using the following criteria: 1. The first tissue of interest had the highest ChIP signal for each region analyzed across all tissues, with a normalized ChIP signal greater than 15. Normalization can take into account background noise, sequencing depth, GC bias, and ROI length, as described above. Normalized ChIP signals can be expressed as fragments per kilobase per million (i.e., FPKM). 2. The relative percentage difference between the highest (denoted as H) and lowest (denoted as L) ChIP signals across all tissue types was required to be at least 20% (i.e., (HL) / L*100%≧20%). 3. The coefficient of variation (CV) of the ChIP signal across all tissue types was required to be at least 25%, where CV was defined as the ratio of the standard deviation to the mean time of 100%. As a result, we identified a total of 27,941 tissue-variable regions for H3K27ac and 17,321 tissue-variable regions for H3K4me3.
[0226] For plasma ChIP-seq data for H3K27ac, we determined the number of plasma DNA fragments whose 5' ends overlapped with each tissue-specific region of H3K27ac. The normalized ChIP signal in FPKM was calculated for each tissue-specific region. Similarly, for plasma DNA ChIP data for H3K4me3, we determined the number of plasma DNA fragments whose 5' ends overlapped with each tissue-specific region of H3K4me3. The normalized ChIP signal was calculated for each tissue-specific region. Comparing the plasma DNA ChIP signal with the ChIP signals from various tissues allowed us to estimate the DNA contribution to the plasma DNA pool associated with the histone modification of interest.
[0227] In one embodiment, the measured ChIP signal levels of DNA molecules were recorded in a vector (X), and the reference ChIP signal levels retrieved across different tissues were recorded in a matrix (M). The proportional contributions (P) from different tissues to the plasma DNA pool were estimated by quadratic programming.
number
number
[0228] The aggregated DNA contribution associated with a particular type of histone modification from all cell types is constrained to be 100%.
number
number
[0229] Figure 38 shows a graph of the percentage contribution of different tissues for both pregnant and non-pregnant samples based on the H3K4me3 histone modification of cfDNA. The x-axis indicates tissue type. The y-axis indicates the percentage contribution estimated by H3K4me3. Tissue types may include, but are not limited to, neutrophils, monocytes, B cells, T cells, natural killer cells, erythroid cells, megakaryocytes, liver, brain, pancreas, placenta, heart, colon, lung, fat, kidney, spleen, bladder, and stomach. It was observed that the main contributors associated with H3K4me3 in plasma DNA were blood-related cell types (i.e., megakaryocytes, neutrophils, and erythrocytes) in both pregnant and non-pregnant subjects, with median contributions of 61.74% and 82.13%, respectively. Notably, placental contribution of H3K4me3 was significantly higher in pregnant subjects (median: 27%, range: 0%-36.67%) compared with non-pregnant subjects, who contributed very little (P-value: 0.0081, Mann-Whitney U test). These results suggested that it was possible to use histone modifications to estimate the proportional contribution of various tissues to plasma DNA.
[0230] Figure 39 shows a graph of placental contribution determined by histone modifications relative to fetal DNA fraction. Placental contribution as a percentage estimated by H3K4me3 signal is on the y-axis. Fetal DNA fraction determined by SNP-based approach is on the x-axis. The placental contribution estimated by ChIP signal of histone modifications according to embodiments of the present disclosure correlates well with the fetal DNA fraction estimated by SNP-based approach (Pearson's r: 0.68, P-value: 0.031). These results suggested that the use of histone modifications enabled the determination of proportional contributions to plasma DNA from various tissues.
[0231] Figure 40 is a graph of the percentage contribution of different tissues for both pregnant and non-pregnant samples based on the H3K27ac histone modification in cfDNA. The x-axis indicates tissue. The y-axis indicates tissue contribution estimated by the H3K27ac histone modification. Figure 40 shows that the use of another type of histone modification, such as H3K27ac, can allow for the estimation of proportional DNA contributions of histone modifications from various tissues to plasma DNA. It could be observed that the main contributors associated with H3K27ac in plasma were blood-related cell types (i.e., megakaryocytes, neutrophils, and erythrocytes) in both pregnant and non-pregnant subjects, with median contributions of 89.51% and 58.67%, respectively. Notably, the placental contribution of H3K27ac was significantly higher in pregnant subjects (median: 14.45%, range: 0.42–28.19%) compared with non-pregnant subjects (median: 0%, range: 0–4.01%) ( P value: <0.0001, Mann-Whitney U test).
[0232] Figure 41 shows a heat map of tissue contributions estimated from H3K27ac ChIP signals in pregnant and non-pregnant subjects. Placental contribution of H3K27ac was significantly higher in pregnant subjects. However, tissues not typically associated with pregnant subjects have a higher contribution of H3K27ac ChIP signals in pregnant subjects than in non-pregnant subjects. Heat map and clustering analysis of tissue contributions revealed tissue clusters (e.g., placenta, lung, colon, spleen, pancreas, adipose, heart, kidney) that showed higher contributions in pregnant subjects compared to non-pregnant subjects. Figure 41 shows that tissues common to both pregnant and non-pregnant subjects may have different contributions of histone modifications. Other tissue clusters composed of blood-type cells (e.g., erythroblasts, megakaryocytes, neutrophils) showed relatively lower contributions in pregnant subjects.
[0233] 2. Concurrent organizational contribution analysis The contribution of a specific tissue of interest can be determined based on the estimated H3K27ac histone modification signal (ChIP signal in this disclosure) associated with tissue-specific histone modification regions. In one embodiment, the amount of histone modification can be estimated by fragmentomics analysis. In one embodiment, contributions from multiple tissue types can be analyzed simultaneously using various tissue-specific histone modification regions. As an example, we analyzed plasma DNA samples from eight healthy subjects. For each sample, we estimated the H3K27ac ChIP signal for regions carrying the H3K27ac histone modification specific to a different tissue. The H3K27ac ChIP signal was estimated using the cumulative frequency of molecules within the size range of 230-350 bp.
[0234] Figure 42 shows a graph of estimated H3K27ac signals for specific tissues. Estimated H3K27ac histone modification signals (also called ChIP signals) are shown on the y-axis. Tissue-specific regions are shown on the x-axis. Each dot represents one plasma DNA fragment.
[0235] When comparing the estimated H3K27ac ChIP signals across various tissue-specific regions, the neutrophil-specific region showed the highest median level compared to other tissues, suggesting neutrophils as the major contributor to plasma cfDNA. The contribution of each tissue can be related to the ChIP signal. For example, it can be determined that monocytes and megakaryocytes may be the next major contributors. The tissues with the lowest contributions may be the placenta and colon. These observations are consistent with previous studies on healthy individuals, which demonstrated that neutrophils are the major contributor to plasma DNA (K. Sun, et al., Proc Natl Acad Sci USA. 2015;112;E5503-E5512).
[0236] 3. Classification of pregnant subjects The ChIP signal can be used to determine the fetal DNA fraction or to distinguish between pregnant and non-pregnant subjects.
[0237] Figures 43A and 43B are graphs showing the correlation between H3K27ac ChIP signal and fetal DNA fraction determined by a SNP-based approach. The x-axis shows the fetal DNA fraction as a percentage determined by the SNP-based approach. As seen in Figure 43A, the use of H3K27ac signal enabled a higher correlation between placental contribution estimated by histone modifications and fetal DNA fraction estimated by the SNP-based approach (Pearson's r: 0.96, P-value: <0.0001). This result highlighted that, in some embodiments, selective use of different types of histone modifications can improve the performance of plasma DNA deconvolution for tissue DNA contributions related to histone modifications. As seen in Figure 43B, there was a weaker correlation between fetal DNA fraction and H3K27ac signal as reads / kb (at the million scale) in placenta-specific H3K27ac regions (Pearson's r: 0.64, P-value: <0.046).
[0238] Figure 44 shows ROC curves for distinguishing between pregnant and non-pregnant subjects. The x-axis shows specificity. The y-axis shows sensitivity. The solid line shows the use of placental contribution estimated from H3K27ac ChIP signals. The dashed line shows the use of reads (in million) / kb in placenta-specific H3K27ac regions. The estimated placental contribution technique has an AUC of 0.984 for distinguishing between pregnant and non-pregnant subjects. The reads / kb technique (i.e., the metric reported in the Sadeh et al. study) has an AUC of 0.785. The results suggested that the use of tissue contribution estimated using quadratic programming resulted in better classification performance compared to the use of reads / kb.
[0239] 4. Samples from subjects with cancer In one embodiment, although there were no colon-specific H3K4me3 regions (Figure 37), we were still able to estimate colon contributions using other tissue-specific and tissue-variable regions. We analyzed the raw sequencing data from the study of Sadeh et al. according to an embodiment of the present disclosure.
[0240] Figure 45 shows the receiver operating characteristic (ROC) curve when using estimated colon contribution to distinguish between control subjects and subjects with colorectal cancer (CRC). The ROC curve shows an area under the curve (AUC) of 0.7. Colon contribution can serve as an indicator for distinguishing subjects with CRC from control subjects. In some embodiments, only tissue variable regions can be used.
[0241] C. Disease Detection and Monitoring Histone modification levels measured by embodiments of the present disclosure can be used to determine classifications of the likelihood of a blood disorder and the level of cancer, including whether the cancer has metastasized. Biological samples from subjects with beta-thalassemia major were analyzed for histone modification levels. Beta-thalassemia major is an example of a blood disorder. Other blood disorders can at least be expected to have similar abnormal results, as blood disorders can have abnormal contributions from cells in the blood. Biological samples from subjects with colorectal cancer (CRC) were analyzed for histone modification levels. CRC is an example of a cancer. Other cancers can be expected to have similar histone modification levels if the cancer is localized to a tissue or if the cancer has metastasized to another tissue.
[0242] 1. Blood disorders To demonstrate the clinical utility of using histone modification-based tissue deconvolution of plasma DNA, we recruited patients with hematological disorders, including but not limited to beta-thalassemia major, iron deficiency anemia, aplastic anemia, and idiopathic thrombocytopenic purpura. We performed an H3K27ac-based immunoprecipitation assay followed by massively parallel sequencing on their plasma DNA samples.
[0243] Figure 46A is a graph comparing erythroblast contribution estimated by H3K27ac ChIP signal between subjects with beta-thalassemia major and control subjects without beta-thalassemia major. The x-axis indicates tissue category. The y-axis indicates erythroblast contribution in percent estimated by H3K27ac ChIP signal. In Figure 46A, subjects with beta-thalassemia major showed an abnormal contribution from erythroblasts (median: 34.97%, range: 6.89-68.44%) compared with healthy control subjects (median: 7.54%, range: 0-12.85%) (P-value: 0.00024, Mann-Whitney U test).
[0244] Figure 46B shows the ROC curve for using estimated erythroblast contribution to distinguish between subjects with and without beta-thalassemia major. The x-axis indicates specificity, and the y-axis indicates sensitivity. ROC analysis revealed that the estimated erythroblast contribution achieved an AUC of 0.923 in the erythroblast-specific region, suggesting that the use of histone modification-based tissue deconvolution of plasma DNA enables the detection and / or monitoring of hematological disorders (e.g., beta-thalassemia major). The estimated tissue contribution was superior to the regional signal measured by reads / kb, with an AUC of 0.892.
[0245] Figure 47 shows a heat map of tissue contributions estimated using H3K27ac ChIP signals in subjects with beta-thalassemia major and control subjects. Tissues were clustered and separated by tissue contribution under different pathological conditions. Compared with control subjects, erythroblasts, monocytes, brain, and others showed higher contributions in beta-thalassemia major subjects. T cells, neutrophils, and megakaryocytes showed lower contributions in beta-thalassemia major subjects. Additionally, we observed lower erythroblast contributions in subjects with aplastic anemia (1.62%) and higher erythroblast contributions in subjects with iron deficiency anemia (16.07%) compared with the median level of erythroblast contribution in control subjects (7.54%). These results are consistent with previous findings that observed similar trends using droplet digital PCR (ddPCR) assays using methylation markers (Lam, et al. ClinChem. 2017;63:1614-1623). These results suggest the potential clinical utility of histone modification-based tissue deconvolution of plasma DNA.
[0246] In addition, we measured erythrocyte DNA in those plasma DNA samples using a published ddPCR assay using differentially methylated regions that were hypomethylated in erythroblasts but hypermethylated in other cell types (Lam et al. ClinChem. 2017;63:1614-1623).
[0247] Figures 48A, 48B, and 48C show the correlation between erythroid DNA percentage determined by ddPCR assay and erythroid contribution determined by H3K27ac signal. The x-axis shows erythroid contribution determined by H3K27ac signal. The y-axis shows erythroid DNA percentage determined by ddPCR assay. Figure 48A shows the use of the FECH (chr18:55250563-55250585) marker with a Pearson's r of 0.87 and a P-value of <0.0001. Figure 48B shows the use of the Ery1 (chr12:48227688-48227701) marker with a Pearson's r of 0.90 and a P-value of <0.0001. Figure 48C shows the use of the Ery2 (chr12:48228144-48228167) marker with a Pearson's r of 0.90 and a P value of <0.0001. The data in these figures further suggested that the use of histone modifications allowed for accurate estimation of the proportional contributions to plasma DNA from various tissues.
[0248] 2. Cancer with metastasis The estimated ChIP signal of histone modifications from plasma DNA without immunoprecipitation can be used to distinguish between localized and metastatic cancers. We analyzed a cohort of samples from four localized colorectal cancer (CRC) patients, seven CRC patients with liver metastases, and eight healthy control subjects. For each sample, we estimated H3K27ac ChIP signals for colon- and liver-specific regions. H3K27ac ChIP signals were estimated using the cumulative frequency of molecules within the size range of 230–350 bp.
[0249] Figure 49A is a graph comparing plasma DNA results from healthy controls to subjects with CRC in the colon-specific H3K27ac region. The graph shows the estimated H3K27ac signal on the y-axis and the subject type (healthy, CRC without liver metastasis, and CRC with liver metastasis) on the x-axis. Each dot represents one plasma DNA fragment. The estimated H3K27ac signal in healthy subjects (median: 0.54, range: 0.27-1.08) was lower than the levels in patients with localized (i.e., without liver metastasis) CRC (median: 0.81, range: 0.47-1.09) and CRC patients with liver metastasis (median: 1.73, range: 0.93-22.28).
[0250] Figure 49B is a graph comparing plasma DNA results from healthy controls to subjects with CRC in the liver-specific H3K27A region. The graph shows the estimated H3K27ac signal on the y-axis and the subject type (healthy, CRC without liver metastasis, and CRC with liver metastasis) on the x-axis. Each dot represents one plasma DNA fragment. The estimated H3K27ac levels for the liver-specific H3K27ac region were shown to be exclusively increased in CRC patients with liver metastasis, indicating an increased liver contribution to cfDNA caused by liver metastasis. Data from both the colon-specific and liver-specific regions can be used to distinguish between patients with localized and metastatic cancer, which may be beneficial for clinical management.
[0251] D. Example of tissue mapping of urine DNA We have demonstrated that the relative tissue contributions to the plasma DNA pool can be estimated by comparing the histone modification profile of plasma DNA with the histone modification profiles from several organs, tissues, or cells. Furthermore, we have demonstrated that these methods presented in this disclosure can be extended to urine samples.
[0252] Figure 50 is a graph of tissue contribution in urine and plasma samples. The x-axis indicates tissue type. The y-axis indicates tissue percent contribution. Each tissue contains two boxplots. The first boxplot (gray) represents data from the plasma sample. The second boxplot (black) represents data from the urine sample. For urine samples, contribution was estimated by comparing the histone modification (e.g., H3K27ac) profile of urinary DNA with the histone modification profile derived from a reference organ, tissue, or cell. For plasma samples, contribution was similarly estimated by comparing the histone modification profile of plasma DNA with the histone modification profile derived from a reference organ, tissue, or cell.
[0253] Urine DNA samples showed significantly higher percentage contributions from kidney (median: 10.66%) and bladder (median: 4.98%) than their counterparts in plasma DNA samples (median: kidney: 0.00%, median: bladder: 0.00%), which is expected from urine samples. These results demonstrate that urine samples can be used to determine tissue contributions using estimated histone modification levels.
[0254] E. Exemplary Methods for Determining Fractional Concentrations Figure 51 is a flowchart of an example process 5100 associated with determining the fractional concentration of a tissue type. In some implementations, one or more process blocks of Figure 51 may be performed by a system (e.g., measurement system 5900). In some implementations, one or more process blocks of Figure 51 may be performed by another device or devices separate from or including the system. Additionally or alternatively, one or more process blocks of Figure 51 may be performed by one or more components of measurement system 5900, such as assay 5908, assay device 5910, detector 5920, logic system 5930, local memory 5935, external memory 5940, storage device 5945, and / or processor 5950.
[0255] In block 5110, N genomic regions are identified. N is an integer greater than 1. The N genomic regions may be regions known to carry tissue-specific histone modifications. The regions may be determined by the criteria described herein. For example, the regions may have tissue histone modification levels greater than a cutoff amount. The cutoff amount may be a normalized ChIP signal, may be based on a relative percentage difference, and / or may be based on a coefficient of variation across all tissue types. The regions may be any regions of interest described herein.
[0256] In block 5120, N tissue-specific histone modification levels are obtained for N genomic regions for each of the M tissue types. N is equal to or greater than M. The histone modification may be H3K27ac, H3K4me3, or any histone modification described herein. The tissue histone modification levels form an N×M dimensional matrix A. One of the M tissue types corresponds to a first tissue type. The first tissue type may be fetal, erythroblast, any tissue listed in FIG. 37 or FIG. 38, or any tissue described herein. At least one genomic region of the N genomic regions includes non-zero histone modification levels from at least two of the M tissue types. For example, at least one histone modification level may not be exclusive to a single tissue.
[0257] At block 5130, an input data vector b is received. The input data vector b may include N mixed histone modification levels at N genomic regions. The N mixed histone modification levels may be measured from a plurality of cell-free DNA molecules in a biological sample of a subject. The biological sample may be any biological sample described herein. The N mixed histone modification levels may be measured by cell-free chromatin immunoprecipitation followed by sequencing (cfChIP-seq) to determine the relative frequency of one or more sets of one or more sequence motifs in the plurality of cell-free DNA molecules, or by determining the relative frequency of one or more size ranges in the plurality of cell-free DNA molecules. The relative frequencies of fragmentomics features other than sequence motifs and size ranges may also be used. The mixed histone modification levels may be determined by any method described herein.
[0258] At block 5140, using a computer system, the fractional concentration of the first tissue type is determined using matrix A and input data vector b. The fractional contribution may be determined using quadratic programming.
[0259] Process 5100 may include determining a classification using the fractional concentration. For example, the first tissue type may be fetal tissue, and process 5100 may further include determining a classification of pregnancy in the subject using the fractional concentration of the first tissue type. The classification of pregnancy may be whether or not pregnant, the gestational age of the fetus (e.g., trimester length), or the level (e.g., presence) of a pregnancy-related disorder.
[0260] As another example, process 5100 may include determining a classification of a disease using the fractional concentration of the first tissue type. For example, the disease may be beta-thalassemia major, iron deficiency anemia, aplastic anemia, or idiopathic thrombocytopenic purpura. The first tissue type may be erythroblast, monocyte, brain, T cell, neutrophil, megakaryocyte, or any other tissue described herein. The level of disease may be whether the disease is present or the severity of the disease. The disease may be a disease of the first tissue type (e.g., cancer).
[0261] Although Figure 51 illustrates example blocks of process 5100, in some implementations, process 5100 may include additional, fewer, different, or differently arranged blocks compared to the blocks illustrated in Figure 51. Additionally or alternatively, two or more of the blocks of process 5100 may be performed in parallel.
[0262] F. Exemplary Methods for Determining Pregnancy or Disease Classification Figure 52 is a flowchart of an example process 5200 associated with determining the fractional concentration of a tissue type. In some implementations, one or more process blocks of Figure 52 may be performed by a system (e.g., measurement system 5900). In some implementations, one or more process blocks of Figure 52 may be performed by another device or devices separate from or including the system. Additionally or alternatively, one or more process blocks of Figure 52 may be performed by one or more components of measurement system 5900, such as assay 5908, assay device 5910, detector 5920, logic system 5930, local memory 5935, external memory 5940, storage device 5945, and / or processor 5950.
[0263] In block 5210, N genomic regions are identified. Block 5210 may be performed in a similar manner to block 5110.
[0264] In block 5220, N tissue-specific histone modification levels are obtained at N genomic regions for each of the M tissue types, where N is greater than or equal to M. Block 5220 may be performed in the same manner as block 5120.
[0265] An input data vector b is received in block 5230. Block 5230 may be implemented in a similar manner to block 5130.
[0266] At block 5240, either a classification of pregnancy in the subject or a classification of disease in the subject can be determined using the computer system, matrix A, and input data vector b. The classification of pregnancy or disease can be any classification described in process 5100. Process 5200 can determine the classification without determining the fractional concentration of the tissue type.
[0267] Determining the pregnancy classification or disease classification may include inputting the matrix A and the input data vector b into a model (e.g., a machine learning model). The model may be trained by receiving the matrix A and a plurality of training input data vectors b obtained from a plurality of biological samples of a plurality of training subjects. Each training subject may have a known classification of the training subject's condition. The condition may be a pregnancy status, a known classification of a disease, or any condition described herein. A plurality of training samples may be saved. Each training sample may include one of the plurality of training input data vectors b and a first label indicating the known classification of the condition. Parameters of the model may be optimized using the plurality of training samples when the matrix A and the plurality of training input data vectors b are input into the model, based on an output of the model that matches or does not match a corresponding label of the first label. The output of the model may specify a classification of the condition. The classification of the condition may be determined using the model.
[0268] The model may include a convolutional neural network (CNN). The CNN may include a set of convolutional filters configured to filter a plurality of input data vectors b. The filters may be any of the filters described herein. The number of filters in each layer may be 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, 150-200, or more. The filter kernel size may be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 15-20, 20-30, 30-40, or more. The CNN may include an input layer configured to receive the filtered plurality of input data vectors b. The CNN may also include multiple hidden layers including multiple nodes. The first of the multiple hidden layers is connected to the input layer. The CNN may further include an output layer coupled to the last of the multiple hidden layers and configured to output an output data structure, which may include the features.
[0269] The models may include supervised learning models, which may include different approaches and algorithms, such as analytical learning, artificial neural networks, backpropagation, boosting (meta-algorithms), Bayesian statistics, case-based reasoning, decision tree learning, inductive logic programming, Gaussian process regression, genetic programming, group methods of data processing, kernel estimators, learning automata, learning classifier systems, minimum message length (decision trees, decision graphs, etc.), multiple linear subspace learning, naive Bayes classifiers, maximum entropy classifiers, conditional random fields, nearest neighbor algorithms, probabilistic approximately correct learning (PAC) learning, ripple-down rules, knowledge acquisition methods, symbolic machine learning algorithms, sub-symbolic machine learning algorithms, support vector machines, minimal complexity machines (MCMs), random forests, ensembles of classifiers, regular classification, data preprocessing, processing of imbalanced datasets, statistical relationship learning or Proaftn, and multi-criteria classification algorithms. The model may be linear regression, logistic regression, a deep recurrent neural network (e.g., long short-term memory, LSTM), a Bayesian classifier, a hidden Markov model (HMM), a linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering for applications with noise (DBSCAN), a random forest algorithm, a support vector machine (SVM), or any model described herein.
[0270] As part of training the machine learning model, the parameters of the machine learning model (weights, thresholds, etc., which can be used, for example, in the activation function of a neural network) are optimized based on the training samples (training set) to provide optimized accuracy in classifying nucleotide modifications at target positions. Various forms of optimization can be performed, such as backpropagation, empirical risk minimization, and structural risk minimization. A validation set of samples (data structure and labels) can be used to verify the accuracy of the model. Cross-validation can be performed using different parts of the training set for training and validation. A model can include multiple submodels, thereby providing an ensemble model. The submodels can be weaker models, but when combined, provide a more accurate final model.
[0271] Although Figure 52 illustrates example blocks of process 5200, in some implementations, process 5200 may include additional, fewer, different, or differently arranged blocks compared to the blocks illustrated in Figure 52. Additionally or alternatively, two or more of the blocks of process 5200 may be performed in parallel.
[0272] VI. Lesion Detection Using Sequence Motifs and Fragment Sizes In embodiments, one or both of fragment size and sequence motifs can be used to classify a pregnancy or disorder. For example, terminal motifs can be used as described elsewhere in this application, including in Section III.D and process 2300 of Figure 23. Fragment sizes that may not be limited to certain terminal motifs can be used. A machine learning model can use terminal motifs and / or fragment size to classify a pregnancy or disorder.
[0273] A. Illustrative Results Figure 53 shows the characteristics of inputs included in a machine learning model to distinguish hepatocellular carcinoma (HCC) from non-HCC cases. Arrays 5304, 5308, 5312, and 5316 each contain data from tissue-specific regions. Tissue-specific regions include liver-specific, neutrophil-specific, megakaryocyte-specific, and erythroblast-specific regions. Each array contains fragment size and fragment end motif information. The frequency of all molecules within 230 to 350 nt (molecules are not limited to any particular fragment end motif when considering size) is included in each array. For example, in array 5304, fragments that align to the liver-specific region and have a size of 230 have a frequency of 0.1 compared to other sizes within the liver-specific region. Other size ranges are possible.
[0274] The array also includes the frequency of all molecules with nine H3K27ac-associated end motifs (molecules are not limited to any fragment size when considering end motifs). H3K27ac-associated end motifs include, but are not limited to, CCGG, CCGC, GCGG, TCGG, TCGC, CCGA, CCCG, GCGC, and / or CCGT. H3K27ac-associated end motifs can be defined as end motifs that are over-represented in regions with high H3K27ac signals compared to regions with low H3K27ac signals in the sequencing results of plasma DNA samples without immunoprecipitation. For example, a high proportion of occupancy can be a fold change in end motif frequency of 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 20x, 30x, 50x, etc. when comparing the results of plasma DNA samples in regions with high and low H3K27ac signals. In some embodiments, H3K27ac-associated terminal motifs can be defined as those motifs that occupy a greater proportion of the sequencing results of plasma DNA samples with immunoprecipitation compared to the results of plasma DNA samples without immunoprecipitation. For example, a greater proportion of occupancy can be a fold change in terminal motif frequency of 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 20x, 30x, 50x, etc. when comparing the results of plasma DNA samples with and without immunoprecipitation.
[0275] The data from all arrays (i.e., one large array or matrix) can be input into a machine learning model to distinguish between non-HCC subjects and HCC subjects. The machine learning model can include, but is not limited to, a support vector machine, a random forest, a convolutional neural network, or any model described herein. In this example, there are a total of 130 features for one type of tissue-specific H3K27ac-associated region. For four different tissue-specific regions, there are 520 features.
[0276] Figures 54A and 54B show results from a machine learning model using the features shown in Figure 53. Figure 54A shows the probability of HCC determined by the machine learning model for control subjects, subjects with chronic hepatitis B virus (HBV), and subjects with HCC. The y-axis shows the probability of HCC. The x-axis shows the type of subject. Figure 54A shows that the probability of HCC determined by the machine learning model was significantly higher in HCC patients compared to patients without HCC.
[0277] Figure 54B is a receiver operating characteristic (ROC) curve. Sensitivity is on the y-axis. Specificity is on the y-axis. ROC analysis revealed that an area under the curve (AUC) of 0.96 was achieved when distinguishing non-HCC from HCC cases by the probability of HCC.
[0278] Figure 55 shows the AUC values determined using different fragmentomics features when distinguishing non-HCC from HCC cases. The y-axis shows the AUC values. The x-axis shows the different fragmentomics features used in the machine learning model for distinguishing non-HCC from HCC cases. Column 1 shows an AUC of 0.93 using the frequency of molecule sizes within 230-350 bp. This model includes 484 features (121 different sizes and 4 tissue-specific regions). Column 2 shows an AUC of 0.95 using the frequency of molecules with H3K27ac-associated motifs. This model uses 36 features (9 motifs and 4 tissue-specific regions). Column 3 has an AUC of 0.96 using both the frequency of molecule sizes within 230-350 bp and the frequency of H3K27ac-associated motifs. This model uses 520 features and is described in Figures 53 and 54. Figure 55 shows that combining size frequency and motif frequency improves the accuracy of HCC case determination. Figure 55 also shows that size frequency and motif frequency can be used separately for different tissue-specific regions to distinguish HCC cases from non-HCC cases.
[0279] B. Exemplary Methods FIG. 56 is a flowchart of an exemplary process 5600 for analyzing a subject's biological sample to determine a classification of the subject's condition. The biological sample includes cell-free DNA. In some implementations, one or more process blocks of FIG. 56 may be performed by a system (e.g., measurement system 5900). In some implementations, one or more process blocks of FIG. 56 may be performed by another device or devices separate from or including the system. Additionally or alternatively, one or more process blocks of FIG. 56 may be performed by one or more components of measurement system 5900, such as assay 5908, assay device 5910, detector 5920, logic system 5930, local memory 5935, external memory 5940, storage device 5945, and / or processor 5950. Process 5600 may include aspects described in process 1700.
[0280] At block 5610, a plurality of sequence reads for the cell-free DNA fragments are received. The plurality of sequence reads can include terminal sequences corresponding to ends of the cell-free DNA fragments.
[0281] At block 5620, a group of sequence reads that map to one or more genomic regions is identified. Each of the one or more genomic regions may have a histone modification associated with one or more target tissue types. The one or more target tissue types may include an organ with cancer or fetal tissue. In some embodiments, the one or more target tissue types may include liver, neutrophils, megakaryocytes, or erythroblasts. The histone modification can be H3K4me1, H3K4me2, H3K27me3, H3K27ac, H3K36me3, H3K9me2, H3K9me3, H3S10P, H3R2me, H3T2P, H3K14ac, H3K9ac, H3K79me2, H3K79me3, H4K5ac, H4K8ac, H4K12ac, H4K16ac, H4K20me, H2BK120ub, or H2AK119ub.
[0282] At block 5630, one or more sequence motifs corresponding to one or more end sequences of the corresponding cell-free DNA fragments are determined for each sequence read in the group of sequence reads. The set of one or more sequence motifs can include 1 to 5, 5 to 10, 11 to 15, 15 to 20, or 20 to 25 sequence motifs. The cell-free DNA fragments can consist of fragments having sequence motifs from the set of one or more sequence motifs.
[0283] At block 5640, the size of the cell-free DNA fragments is measured using the sequence reads. The cell-free DNA fragments may have a size having a predetermined size range. The predetermined size range may be any size range described herein, including 230 to 350 nt.
[0284] At block 5650, one or more sequence motif frequencies of the one or more sets of sequence motifs are determined for each of the one or more target tissue types, wherein the one or more sets of sequence motifs occur at a higher rate in chromatin immunoprecipitation followed by sequencing (ChIP-seq) of histone modifications associated with one or more genomic regions than in sequencing without chromatin immunoprecipitation.
[0285] At block 5660, one or more size frequencies of sequence reads for one or more size ranges are determined for each of one or more target tissue types.
[0286] At block 5670, the one or more sequence motif frequencies and one or more size frequencies for each of the one or more target tissue types are input into a machine learning model. The machine learning model may include a support vector machine, a random forest, or a convolutional neural network. The machine learning model may be any machine learning model disclosed herein, including models similar to those described in process 5200.
[0287] The machine learning model can be trained by receiving a training dataset. The training dataset can include training sequence motif frequencies of one or more sets of sequence motifs for each of one or more target tissue types, and training size frequencies of cell-free DNA fragments from multiple biological samples of multiple training subjects. Each training subject can have a known classification of a condition.
[0288] The machine learning model can also be trained by storing multiple training samples. Each training sample can include one or more training sequence motif frequencies of one or more sets of sequence motifs occurring in cell-free DNA fragments in the training sample for each of one or more target tissue types. Each training sample can include training size frequencies of cell-free DNA fragments in the training sample for each of one or more target tissue types. Each training sample can also include a first label indicating a known classification of a condition.
[0289] The machine learning model can be trained by using a plurality of training samples to optimize parameters of the machine learning model based on the output of the machine learning model that matches or does not match the corresponding label of the first label, once the sequence motif frequency and size frequency for each of one or more target tissue types are input into the machine learning model. The output of the machine learning model can specify a classification of the condition.
[0290] In some embodiments, process 5600 may include, for each sequence motif of the set of one or more sequence motifs, determining a size parameter for fragments having the respective sequence motif. The size parameter may be a statistical value (e.g., mean, median, mode, percentile) of fragments having the respective sequence motif. Process 5600 may further include inputting the one or more size parameters into a machine learning model. The machine learning model in these embodiments may be trained with training samples containing the determined size parameters.
[0291] At block 5680, a classification of the subject's condition is determined using the machine learning model. The condition can be pregnancy. For example, the classification of pregnancy can provide gestational age or the presence or severity of a pregnancy-related disorder, including any pregnancy-related disorder described herein. The condition can be a disease. The classification of the disease can be the presence or severity of the disease. The disease can be hepatocellular carcinoma (HCC) or cancer, including any cancer described herein.
[0292] In some embodiments, process 5600 may be modified so that either sequence motif frequencies or size frequencies are used. For example, process 5600 may include using only size frequencies of molecules within a certain size range (e.g., the first column of FIG. 55). In this case, blocks 5630 and 5650 are optional. Block 5670 may be modified so that one or more size frequencies are input to the machine learning model and one or more sequence motif frequencies are not input thereto. As another example, process 5600 may include using only motif frequencies of molecules (e.g., the second column of FIG. 55). In this case, blocks 5640 and 5660 are optional. Block 5670 may be modified so that one or more sequence motif frequencies are input to the machine learning model and one or more size frequencies are not input thereto.
[0293] Process 5600 may include additional implementations, such as any single implementation or any combination of implementations described herein, and / or in conjunction with one or more other processes described elsewhere herein.
[0294] Although Figure 56 illustrates example blocks of process 5600, in some implementations, process 5600 may include additional, fewer, different, or differently arranged blocks compared to the blocks illustrated in Figure 56. Additionally or alternatively, two or more of the blocks of process 5600 may be performed in parallel.
[0295] VII. Region Enrichment The preference for DNA fragments associated with a particular epigenomic state, which exhibit a particular set of terminal motifs, can be used to enrich a sample for DNA with that particular epigenomic state. Thus, embodiments can enrich a sample for clinically relevant DNA, including DNA from a particular tissue. For example, only DNA fragments with a particular terminal sequence can be sequenced, amplified, and / or captured using the assay. As another example, filtering of sequence reads can be performed.
[0296] A. Physical concentration Physical enrichment can be performed in a variety of ways, for example, via targeted sequencing or PCR, such as by using specific primers or adapters. If a specific terminal motif in the terminal sequence is detected, an adapter can be added to the end of the fragment. Then, when sequencing is performed, only DNA fragments with the adapter are sequenced (or at least primarily sequenced), thereby providing targeted sequencing.
[0297] As another example, primers that hybridize to a set of specific terminal motifs can be used. These primers can then be used to perform sequencing or amplification. Capture probes corresponding to specific terminal motifs can also be used to capture DNA molecules with those terminal motifs for further analysis. Some embodiments can ligate short oligonucleotides to the ends of plasma DNA molecules. The probes can then be designed to recognize only sequences that are partly terminal motifs and partly ligated oligonucleotides.
[0298] Some embodiments may use CRISPR-based diagnostic techniques, for example, using guide RNAs to define sites corresponding to preferred terminal motifs in clinically relevant DNA, and then using nucleases to cleave the DNA fragments, as can be done using Cas-9 or Cas-12. For example, adapters can be used to recognize the terminal motif, and CRISPR / Cas9 or Cas-12 can be used to cleave the terminal motif / adapter hybrid to create universally recognizable ends for further enrichment of molecules at the desired ends.
[0299] FIG. 57 is a flowchart of an exemplary process 5700 related to enriching a biological sample for clinically relevant DNA. The biological sample may contain clinically relevant DNA and other DNA that is cell-free. In some implementations, one or more process blocks of FIG. 57 may be performed by a system (e.g., measurement system 5900). In some implementations, one or more process blocks of FIG. 57 may be performed by another device or devices that are separate from or include the system. Additionally or alternatively, one or more process blocks of FIG. 57 may be performed by one or more components of measurement system 5900, such as assay 5908, assay device 5910, detector 5920, logic system 5930, local memory 5935, external memory 5940, storage device 5945, and / or processor 5950. Process 5700 may include aspects described in process 1700.
[0300] At block 5710, a plurality of sequence reads for the cell-free DNA fragments are received. The plurality of sequence reads include terminal sequences corresponding to the ends of the plurality of cell-free DNA fragments. One or more sequence motifs may correspond to one or more terminal sequences for each cell-free DNA fragment. Block 5710 may be performed in a manner similar to block 1710.
[0301] At block 5720, a set of one or more sequence motifs is identified that occur at a higher rate in chromatin immunoprecipitation followed by sequencing (ChIP-seq) of histone modifications in clinically relevant DNA than in sequencing without chromatin immunoprecipitation. Identification of the sequence motifs can be similar to the procedures described in process 1700 and Figures 3, 10, 11, and 51.
[0302] At block 5730, the plurality of cell-free DNA fragments are subjected to one or more probe molecules that detect a set of one or more sequence motifs in the terminal sequences of the plurality of cell-free DNA fragments, thereby obtaining detected DNA fragments. Such use of probe molecules can result in obtaining detected DNA fragments. In one example, the one or more probe molecules can include one or more enzymes that interrogate the plurality of cell-free DNA fragments and add new sequences that are used to amplify the detected DNA fragments. In another example, the one or more probe molecules can be attached to a surface to detect sequence motifs in the terminal sequences by hybridization.
[0303] At block 5740, the detected DNA fragments are used to enrich the biological sample for clinically relevant DNA fragments. In some embodiments, enriching the biological sample using the detected DNA fragments may include amplifying the detected DNA fragments. In some embodiments, enriching the biological sample for clinically relevant DNA fragments using the detected DNA fragments may include capturing the detected DNA fragments and discarding the undetected DNA fragments.
[0304] Process 5700 may further include analyzing the enriched biological sample to determine a tissue of origin or disease level classification. Analysis of the enriched biological sample may include sequencing DNA fragments in the enriched biological sample.
[0305] Process 5700 may include additional implementations, such as any single implementation or any combination of implementations described herein, and / or in conjunction with one or more other processes described herein.
[0306] Although Figure 57 illustrates example blocks of process 5700, in some implementations, process 5700 may include additional, fewer, different, or differently arranged blocks compared to the blocks illustrated in Figure 57. Additionally or alternatively, two or more of the blocks of process 5700 may be performed in parallel.
[0307] B. In silico enrichment In silico enrichment can select or discard certain DNA fragments using various criteria. Such criteria include terminal motifs, open chromatin regions, size, sequence variation, methylation, and other epigenetic features. Epigenetic features include all modifications of the genome that do not involve changes in the DNA sequence. Criteria can specify cutoffs that require certain characteristics, such as a certain size range, a methylation metric above or below a certain amount, a combination of methylation states of two or more CpG sites (e.g., a methylation haplotype (Guo et al., Nat Genet. 2017;49:635-42)), or have a combined probability that exceeds a threshold. Such enrichment can also include weighting DNA fragments based on such probability.
[0308] For example, enriched samples can be used to classify pathologies (as described above), as well as to identify tumor or fetal mutations, or for tag counting to detect amplification / deletion of chromosomes or chromosomal regions. For example, if a particular terminal motif or set of terminal motifs is associated with liver cancer (i.e., a higher relative frequency than non-cancer or other cancers), an embodiment for performing cancer screening may weight such DNA fragments higher than DNA fragments that do not have this preferred one or this preferred set of terminal motifs.
[0309] FIG. 57 is a flowchart of an exemplary process 5800 related to enriching a biological sample for clinically relevant DNA. The biological sample may contain clinically relevant DNA and other DNA that is acellular. The clinically relevant DNA may be DNA from the tissue of origin or DNA from a diseased tissue. In some implementations, one or more process blocks of FIG. 58 may be performed by a system (e.g., measurement system 5900). In some implementations, one or more process blocks of FIG. 58 may be performed by another device or devices that are separate from or include the system. Additionally or alternatively, one or more process blocks of FIG. 58 may be performed by one or more components of measurement system 5900, such as assay 5908, assay device 5910, detector 5920, logic system 5930, local memory 5935, external memory 5940, storage device 5945, and / or processor 5950. Process 5800 may include aspects described in process 1700.
[0310] At block 5810, a plurality of sequence reads for the cell-free DNA fragments are received. The plurality of sequence reads include terminal sequences corresponding to the ends of the plurality of cell-free DNA fragments. One or more sequence motifs may correspond to one or more terminal sequences for each cell-free DNA fragment. Block 5810 may be performed in a manner similar to block 1710.
[0311] The plurality of sequence reads may be located in one or more predetermined genomic regions, each of which has a histone modification associated with one or more target tissue types. The sequence reads may be aligned to a reference genome to determine their locations. Identifying sequence reads at these locations may be performed in a manner similar to block 1720.
[0312] In block 5820, one or more sequence motifs corresponding to one or more end sequences of the cell-free DNA fragments are determined for each sequence read of the group of sequence reads. Block 5820 may be performed in a manner similar to block 1730.
[0313] At block 5830, a set of one or more sequence motifs is identified that occur at a higher rate in chromatin immunoprecipitation followed by sequencing (ChIP-seq) of histone modifications in clinically relevant DNA than in sequencing without chromatin immunoprecipitation. Identification of the sequence motifs can be similar to the procedures described in process 1700 and Figures 3, 10, 11, and 51.
[0314] At block 5840, a group of sequence reads having a set of one or more sequence motifs in the end sequences is identified, which may be considered the first stage of filtering.
[0315] At block 5850, a likelihood that the sequence read corresponds to clinically relevant DNA is determined for each sequence read of the group of sequence reads based on an end sequence of the sequence read that includes one or more sequence motifs from the set of sequence motifs. For example, for each sequence read of the group of sequence reads, a likelihood that the sequence read corresponds to clinically relevant DNA may be determined based on an end sequence of the sequence read that includes one or more sequence motifs from the set of sequence motifs.
[0316] In block 5860, the likelihood is compared to a threshold for each sequence read in the group of sequence reads. As an example, the threshold may be determined empirically. For example, various thresholds may be tested for samples in which the concentration of clinically relevant DNA may be measured for the group of sequence reads. An optimal threshold may maximize the concentration while maintaining a certain percentage of the total number of sequence reads. The threshold may be determined by one or more given percentiles (5, 10, 90, or 95) of the concentration of one or more terminal motifs present in healthy controls or control groups without the disease but exposed to similar etiological risk factors. The threshold may be a regression or probability score.
[0317] In block 5870, when the likelihood exceeds a threshold for each sequence read in the group of sequence reads, the sequence read is stored. The sequence read may be stored in memory (e.g., a file, table, or other data structure), thereby obtaining stored sequence reads. Sequence reads with likelihoods below the threshold may be discarded or not stored in the memory location of retained reads, or a database field may include a flag indicating a low threshold for reads so that subsequent analysis may exclude such reads. By way of example, likelihood may be determined using various techniques, such as odds ratios, z-scores, or probability distributions.
[0318] In block 5880, the stored sequence reads are analyzed to determine a characteristic of the clinically relevant DNA biological sample. For example, the characteristic may be any of those described herein, including other flowcharts. For example, the characteristic of the clinically relevant DNA biological sample may be the fractional concentration of clinically relevant DNA. As another example, the characteristic may be the level of pathology of the subject from which the biological sample was obtained, the level of pathology being associated with the clinically relevant DNA. As another example, the characteristic may be the gestational age of the fetus of the pregnant woman from whom the biological sample was obtained.
[0319] Other criteria can be used to determine likelihood.The size of multiple cell-free DNA fragments can be measured using sequence reads.The likelihood that a particular sequence read corresponds to clinically relevant DNA can also be based on the size of the cell-free DNA fragment that corresponds to a particular sequence read.
[0320] Methylation may also be used. Thus, embodiments may measure one or more methylation states at one or more sites in cell-free DNA fragments corresponding to a particular sequence read. The likelihood that a particular sequence read corresponds to clinically relevant DNA may be further based on one or more methylation states. As a further example, whether a read falls within a specified set of open chromatin regions may be used as a filter.
[0321] Process 5800 may include additional implementations, such as any single implementation or any combination of implementations described herein, and / or in conjunction with one or more other processes described herein.
[0322] Although Figure 58 illustrates example blocks of process 5800, in some implementations, process 5800 may include additional, fewer, different, or differently arranged blocks compared to the blocks illustrated in Figure 58. Additionally or alternatively, two or more of the blocks of process 5800 may be performed in parallel.
[0323] VIII. Exemplary Systems FIG. 59 illustrates a measurement system 5900 according to one embodiment of the present disclosure. The system shown includes a sample 5905, such as acellular nucleic acid molecules (e.g., DNA and / or RNA), in an assay device 5910, and an assay 5908 can be performed on the sample 5905. For example, the sample 5905 can be contacted with reagents for the assay 5908 to provide a signal of a physical characteristic 5915 (e.g., sequence information of the acellular nucleic acid molecule). An example of an assay device can be a flow cell containing assay probes and / or primers, or a tube through which droplets (along with droplets containing the assay) travel. The physical characteristic 5915 (e.g., fluorescence intensity, voltage, or current) from the sample is detected by a detector 5920. The detector 5920 can take measurements at intervals (e.g., periodic intervals) to obtain data points that constitute a data signal. In one embodiment, an analog-to-digital converter converts the analog signal from the detector to digital form at multiple times.
[0324] Assay device 5910 and detector 5920 may form an assay system, e.g., a sequencing system, that performs sequencing according to embodiments described herein. Data signal 5925 is transmitted from detector 5920 to logic system 5930. As an example, data signal 5925 can be used to determine the sequence and / or location in a reference genome of a nucleic acid molecule (e.g., DNA and / or RNA). Data signal 5925 can include various measurements performed simultaneously, e.g., different colors of fluorescent dyes or different electrical signals for different molecules in sample 5905; thus, data signal 5925 can correspond to multiple signals. Data signal 5925 can be stored in local memory 5935, external memory 5940, or storage device 5945. An assay system can be composed of multiple assay devices and detectors.
[0325] Logic system 5930 may be or include a computer system, ASIC, microprocessor, graphics processing unit (GPU), etc. It may also include or be coupled to a display (e.g., a monitor, LED display, etc.) and user input devices (e.g., a mouse, keyboard, buttons, etc.). Logic system 5930 and other components may be part of a standalone or networked computer system, or may be directly attached to or incorporated into a device (e.g., a sequencing device) including detector 5920 and / or assay device 5910. Logic system 5930 may also include software executing on processor 5950. Logic system 5930 may include a computer-readable medium that stores instructions for controlling measurement system 5900 to perform any of the methods described herein. For example, logic system 5930 may provide commands to a system including assay device 5910 such that sequencing or other physical operations are performed. Such physical operations may be performed in a particular order, e.g., reagents are added and removed in a particular order. Such physical actions can be performed by a robotic system, including, for example, a robotic arm, such that it can be used to obtain samples and perform assays.
[0326] The measurement system 5900 may also include a treatment device 5960 that can provide treatment to the subject. The treatment device 5960 can be used to determine and / or administer a treatment. Examples of such treatments may include surgery, radiation therapy, chemotherapy, immunotherapy, targeted therapy, hormone therapy, and stem cell transplant. The logic system 5930 may be connected to the treatment device 5960, for example, to provide results of the methods described herein. The treatment device may receive input from other devices, such as an imaging device, and user input (e.g., to control the treatment, such as control for a robotic system).
[0327] Any of the computer systems referred to herein may utilize any suitable number of subsystems. An example of such a subsystem is shown in FIG. 60, computer system 10. In some embodiments, the computer system includes a single computer device, and the subsystems may be components of the computer device. In other embodiments, the computer system may include multiple devices, each of which is a subsystem, along with its internal components. Computer systems may include desktop and laptop computers, tablets, mobile phones, and other mobile devices.
[0328] The subsystems shown in FIG. 60 are interconnected via a system bus 75. Additional subsystems are shown, such as a printer 74, a keyboard 78, a storage device 79, a monitor 76 (e.g., a display screen such as an LED) coupled to a display adapter 82, and others. Peripherals and input / output (I / O) devices coupled to the I / O controller 71 may be connected to the computer system by any number of means known in the art, such as input / output (I / O) ports 77 (e.g., USB, Lightning, Thunderbolt). For example, the I / O ports 77 or external interface 81 (e.g., Ethernet, Wi-Fi, etc.) may be used to connect the computer system 10 to a wide area network, such as the Internet, a mouse input device, or a scanner. The interconnection via the system bus 75 allows the central processor 73 to communicate with each subsystem and control the execution of instructions from the system memory 72 or storage device 79 (e.g., a fixed disk such as a hard drive or optical disk), as well as the exchange of information between the subsystems. The system memory 72 and / or storage device 79 may embody computer-readable media. Another subsystem is a data collection device 85, such as a camera, microphone, accelerometer, etc. Any of the data mentioned herein can be output from one component to another and can be output to a user.
[0329] A computer system may include multiple identical components or subsystems connected together, for example, by an external interface 81, by an internal interface, or through a removable storage device that can be connected and disconnected from one component to another. In some embodiments, computer systems, subsystems, or devices may communicate over a network. In such cases, one computer may be considered a client and another computer may be considered a server, each of which may be part of the same computer system. The client and server may each include multiple systems, subsystems, or components.
[0330] Aspects of the embodiments can be implemented in the form of control logic using hardware circuitry (e.g., application-specific integrated circuits or field-programmable gate arrays) and / or using computer software stored in memory in conjunction with a processor that is generally programmable in a modular or integrated fashion; thus, a processor may include a memory that stores software instructions that configure the hardware circuitry, as well as an FPGA or ASIC that includes the configuration instructions. As used herein, a processor may include a single-core processor, a multi-core processor on the same integrated chip, or multiple processing units on a single circuit board or networked, as well as dedicated hardware. Based on the disclosure and teachings provided herein, those skilled in the art will recognize and understand other means and / or methods of implementing embodiments of the present disclosure using hardware and combinations of hardware and software.
[0331] Any of the software components or functions described in this application may be implemented as software code to be executed by a processor using any suitable computer language, such as, for example, Java, C, C++, C#, Objective-C, Swift, or a scripting language such as Perl or Python, using, for example, conventional or object-oriented techniques. The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission. Suitable non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), magnetic media such as a hard drive or floppy disk, optical media such as a compact disc (CD) or DVD (digital versatile disc) or Blu-ray disc, flash memory, etc. The computer-readable medium may be any combination of such devices. Additionally, the order of operations may be re-arranged. A process may terminate when its operations are completed, but may have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. If the process corresponds to a function, its termination may correspond to the invocation of the function or the return of the function to the main function.
[0332] Such programs may also be encoded and transmitted using carrier signals adapted for transmission over wired, optical, and / or wireless networks conforming to various protocols, including the Internet. Thus, computer-readable media may be created using data signals encoded with such programs. Computer-readable media encoded with program code may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer-readable medium may reside on or within a single computer product (e.g., a hard drive, CD, or entire computer system), or may reside on or within different computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing a user with any of the results described herein.
[0333] Any of the methods described herein may be implemented, in whole or in part, using a computer system including one or more processors that may be configured to perform steps. Any operation (e.g., aligning, determining, comparing, calculating, computing) performed by a processor may be performed in real time. The term "real time" may refer to a computing operation or process that is completed within a certain time constraint. The time constraint may be one minute, one hour, one day, or seven days. Accordingly, embodiments may be directed to a computer system configured to perform steps of any of the methods described herein, potentially with different components performing each step or each group of steps. Although presented as numbered steps, steps of the methods herein may be performed simultaneously, at different times, or in different orders. Additionally, portions of these steps may be used with portions of other steps from other methods. Also, all or portions of the steps may be optional. Additionally, any of the steps of any of the methods may be performed using a system module, unit, circuit, or other means for performing these steps.
[0334] The specific details of particular embodiments may be combined in any suitable manner without departing from the spirit and scope of the embodiments of the present disclosure, however, other embodiments of the present disclosure may be directed to specific embodiments relating to each individual aspect, or specific combinations of these individual aspects.
[0335] The foregoing description of exemplary embodiments of the present disclosure has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the disclosure to the precise form described, and many modifications and variations are possible in light of the above teachings.
[0336] The use of "a," "an," or "the" is intended to mean "one or more," unless specifically stated to the contrary. The use of "or" is intended to mean "inclusive or," and not "exclusive or," unless specifically stated to the contrary. A reference to a "first" element does not necessarily require that a second element be provided. Furthermore, a reference to a "first" or "second" element does not limit the referenced element to a particular location unless expressly stated. The term "based on" is intended to mean "based at least in part on."
[0337] The claims may be drafted to exclude any element that may be optional. Accordingly, this statement is intended to serve as a predicate for use of exclusive terminology such as "solely," "only," or the like in connection with the recitation of claim elements or for use of a "negative" limitation.
[0338] All patents, patent applications, publications, and specifications referred to herein are incorporated by reference in their entirety for all purposes. None is admitted to be prior art. In the event of a conflict between this application and the references provided herein, this application shall control.
Claims
1. 1. A method for analyzing a biological sample, wherein the biological sample comprises cell-free DNA fragments, the method comprising: receiving a plurality of sequence reads for the cell-free DNA fragments; identifying a group of sequence reads that map to one or more genomic regions, each of the one or more genomic regions having a histone modification associated with a target tissue type; determining a value of a fragmentomic attribute of each cell-free DNA fragment corresponding to each sequence read in the group of sequence reads; determining a relative frequency of one or more cell-free DNA fragments having values of the fragmentomics trait within a set of one or more value ranges, wherein the set of one or more value ranges occurs at a differential rate in chromatin immunoprecipitation followed by sequencing for the histone modifications associated with the one or more genomic regions than in sequencing without chromatin immunoprecipitation; determining a summary value of the one or more relative frequencies; comparing the aggregated value to one or more calibration values; and using said comparison to determine the amount of said histone modification in said biological sample.
2. the fragmentomics feature is a sequence motif corresponding to the terminal sequence of the end of the cell-free DNA fragment; The method of claim 1 , wherein the one or more value ranges are one or more sequence motifs.
3. the fragmentomics characteristic is size; The method of claim 1 , wherein the one or more value ranges are one or more size ranges.
4. The fragmentomics feature is a topological form, The method of claim 1 , wherein the one or more value ranges are in one or more topological forms.
5. the fragmentomics feature is a nucleosome footprint; 2. The method of claim 1, wherein the one or more ranges of values are one or more nucleosome footprints.
6. comparing the amount of the histone modification to one or more second calibration values; using said comparison of said amount of said histone modification to said one or more second calibration values, determining the fractional concentration of said target tissue type; To determine the classification of the level of disability; or and determining a classification of the target tissue type implant status.
7. 1. A method for analyzing a biological sample, wherein the biological sample comprises cell-free DNA fragments, the method comprising: receiving a plurality of sequence reads for the cell-free DNA fragments, the plurality of sequence reads including terminal sequences corresponding to ends of the cell-free DNA fragments; identifying a group of sequence reads that map to one or more genomic regions, each of the one or more genomic regions having a histone modification associated with a target tissue type; determining, for each sequence read of said group of sequence reads, one or more sequence motifs that correspond to one or more end sequences of the corresponding cell-free DNA fragment; determining the relative frequency of one or more of the set of one or more sequence motifs, wherein the set of one or more sequence motifs occurs at a higher rate in chromatin immunoprecipitation followed by sequencing of the histone modifications associated with the one or more genomic regions than in sequencing without chromatin immunoprecipitation; determining a summary value of the one or more relative frequencies; comparing the aggregated value to one or more calibration values; and determining a fractional concentration of cell-free DNA fragments from the target tissue type using the comparison, wherein the one or more calibration values are determined from one or more calibration samples having known fractional concentrations of cell-free DNA fragments from the target tissue type.
8. the one or more relative frequencies are one or more first relative frequencies; the aggregated value is a first aggregated value, The one or more calibration values are: For each calibration sample of the one or more calibration samples: determining a second relative frequency of one or more of said sets of one or more sequence motifs in said one or more genomic regions; determining a second aggregate value of the one or more second relative frequencies; and correlating each of the one or more second aggregate values to a known fractional concentration, wherein the one or more calibration values comprise the one or more second aggregate values.
9. 8. The method of claim 7, wherein the aggregate value is a value selected from the group consisting of: (i) an entropy value, (ii) a sum of relative frequencies, (iii) a ratio of relative frequencies, and (iv) a multidimensional data point corresponding to a vector of counts for the set of one or more sequence motifs.
10. 8. The method of claim 7, wherein a sequence motif of the set of one or more sequence motifs corresponds to a single nucleotide, a two-nucleotide sequence, a three-nucleotide sequence, a four-nucleotide sequence, a five-nucleotide sequence, a six-nucleotide sequence, or a seven-nucleotide sequence.
11. 11. The method of claim 10, wherein the sequence motif comprises the nucleotides at the ends of the cell-free DNA fragments.
12. The method of claim 10, wherein the sequence motif is at the 5' end.
13. 8. The method of claim 7, wherein the target tissue type comprises placenta, liver, heart, neutrophils, monocytes, B cells, adipose, or NK cells.
14. the target tissue type is placenta; The method comprises:
8. The method of claim 7, further comprising using the fractional concentrations to determine a pregnancy-related disorder classification or gestational age.
15. 8. The method of claim 7, further comprising using the fractional concentration to determine a classification of the level of cancer.
16. the group of sequence reads is a first group of sequence reads; the one or more genomic regions are one or more first genomic regions; the histone modification is a first histone modification; the target tissue type is a first target tissue type; said set of one or more sequence motifs is a set of one or more first sequence motifs; the one or more relative frequencies are one or more first relative frequencies; the aggregated value is a first aggregated value, the one or more calibration samples are one or more first calibration samples; the fractional concentration is a first fractional concentration; The method comprises: identifying a second group of sequence reads that map to one or more second genomic regions, each of the one or more second genomic regions having a second histone modification associated with a second target tissue type; determining, for each sequence read in the second group of sequence reads, one or more second sequence motifs that correspond to the one or more end sequences of the corresponding cell-free DNA fragment; determining a second relative frequency of one or more of the set of one or more second sequence motifs, wherein the set of one or more second sequence motifs occurs at a higher rate in chromatin immunoprecipitation for the second histone modification associated with the one or more second genomic regions followed by sequencing than in sequencing without chromatin immunoprecipitation; determining a second aggregate value of the one or more second relative frequencies; comparing the second aggregated value to one or more second calibration values; 8. The method of claim 7, further comprising: using said comparison to determine a second fractional concentration of cell-free DNA fragments from said second target tissue type, wherein said one or more second calibration values are determined from one or more second calibration samples having known fractional concentrations of DNA fragments from said second target tissue type.
17. 1. A method for analyzing a biological sample, wherein the biological sample comprises cell-free DNA fragments, the method comprising: receiving a plurality of sequence reads for the cell-free DNA fragments, the plurality of sequence reads including terminal sequences corresponding to ends of the cell-free DNA fragments; identifying a group of sequence reads that map to one or more genomic regions, each of the one or more genomic regions having a histone modification associated with a target tissue type; determining, for each sequence read of said group of sequence reads, one or more sequence motifs that correspond to one or more end sequences of the corresponding cell-free DNA fragment; determining the relative frequency of one or more of the set of one or more sequence motifs, wherein the set of one or more sequence motifs occurs at a higher rate in chromatin immunoprecipitation followed by sequencing of the histone modifications associated with the one or more genomic regions than in sequencing without chromatin immunoprecipitation; determining a summary value of the one or more relative frequencies; comparing the aggregated value to one or more calibration values; and using the comparison to estimate a first value of a characteristic of the target tissue type, wherein the one or more calibration values are determined from one or more calibration samples having known values for the characteristic of the target tissue type.
18. the one or more relative frequencies are one or more first relative frequencies; the aggregated value is a first aggregated value, The one or more calibration values are: For each calibration sample of the one or more calibration samples: determining a second relative frequency of said set of one or more sequence motifs in said one or more genomic regions; determining a second aggregate value of the one or more second relative frequencies; and correlating each of one or more second aggregate values to a known value for the feature, whereby the one or more calibration values comprise the one or more second aggregate values.
19. 18. The method of claim 17, wherein the target tissue type is liver or hematopoietic cells.
20. 18. The method of claim 17, wherein the target tissue type is an organ with cancer.
21. The method of any one of claims 17 to 20, wherein the characteristic is the level of cancer.
22. 18. The method of claim 17, wherein the characteristic is the nutritional status of an organ.
23. 18. The method of claim 17, wherein the target tissue type is fetal tissue.
24. 18. The method of claim 17, wherein the biological sample is obtained from a pregnant woman and the target tissue type is placental tissue.
25. 18. The method of claim 17, wherein the target tissue type is placental tissue and the characteristic of the placental tissue comprises gestational age of a pregnant subject.
26. the aggregated value is a first aggregated value, the one or more calibration values are one or more first calibration values; The method comprises: measuring the size of the cell-free DNA fragments using the sequence reads; and determining one or more size frequencies of said sequence reads for one or more size ranges; determining a second aggregate value of the one or more size frequencies; comparing the second aggregated value to one or more second calibration values; 18. The method of claim 17, wherein estimating the first value of the characteristic comprises using the comparison of the second aggregated value with the one or more second calibration values, the one or more second calibration values being determined from the one or more calibration samples.
27. 1. A method for analyzing a biological sample, wherein the biological sample comprises cell-free DNA fragments, the method comprising: receiving a plurality of sequence reads for the cell-free DNA fragments, the plurality of sequence reads including terminal sequences corresponding to ends of the cell-free DNA fragments; identifying a group of sequence reads that map to one or more genomic regions, each of the one or more genomic regions having a histone modification; determining, for each sequence read of said group of sequence reads, one or more sequence motifs that correspond to one or more end sequences of the corresponding cell-free DNA fragment; determining the relative frequency of one or more of the set of one or more sequence motifs, wherein the set of one or more sequence motifs occurs at a higher rate in chromatin immunoprecipitation followed by sequencing of the histone modifications associated with the one or more genomic regions than in sequencing without chromatin immunoprecipitation; determining a summary value of the one or more relative frequencies; comparing the aggregated value to one or more calibration values; and using said comparison to determine the amount of histone modification in said one or more genomic regions, wherein said one or more calibration values are determined from one or more calibration samples in which the amount of histone modification is known.
28. the one or more relative frequencies are one or more first relative frequencies; the aggregated value is a first aggregated value, The one or more calibration values are: For each calibration sample of the one or more calibration samples: determining a second relative frequency of one or more of said sets of one or more sequence motifs in said one or more genomic regions; determining a second aggregate value of the one or more second relative frequencies; and correlating each of the one or more second aggregate values to a known amount of the histone modification, wherein the one or more calibration values comprise the one or more second aggregate values.
29. 29. The method of claim 27 or 28, wherein the amount of histone modification is known for the one or more calibration samples from performing chromatin immunoprecipitation followed by sequencing on each of the one or more calibration samples.
30. the group of sequence reads is a first group of sequence reads; the one or more genomic regions are one or more first genomic regions; said set of one or more sequence motifs is a set of one or more first sequence motifs; the one or more relative frequencies are one or more first relative frequencies; the aggregated value is a first aggregated value, the one or more calibration values are one or more first calibration values; the amount of histone modification is a first amount of histone modification; the histone modifications are associated with a first tissue type and a second tissue type at the one or more first genomic regions; The method comprises: identifying a second group of sequence reads that map to one or more second genomic regions, each of the one or more second genomic regions having the histone modifications associated with the first tissue type and the second tissue type; determining, for each sequence read in the second group of sequence reads, one or more second sequence motifs corresponding to one or more end sequences of the corresponding cell-free DNA fragment; determining a second relative frequency of one or more of the set of one or more second sequence motifs, wherein the set of one or more second sequence motifs occurs at a higher rate in chromatin immunoprecipitation followed by sequencing of the histone modifications associated with the one or more second genomic regions than in sequencing without chromatin immunoprecipitation; determining a second aggregate value of the one or more second relative frequencies; comparing the second aggregated value to one or more second calibration values; using the comparison to determine a second amount of the histone modification in the one or more second genomic regions, wherein the one or more second calibration values are determined from one or more second calibration samples having known amounts of the histone modification; 28. The method of claim 27, further comprising determining a first fractional concentration of the first tissue type and a second fractional concentration of the second tissue type by solving a system of linear or nonlinear equations comprising the first amount of a histone modification, the second amount of a histone modification, and parameters specifying the relative amount of the respective histone modification for each tissue type in the one or more first genomic regions and the one or more second genomic regions.
31. the histone modifications are associated with a third tissue type in the one or more first genomic regions and the one or more second genomic regions; the histone modifications are associated with the first tissue type, the second tissue type, and the third tissue type at one or more third genomic regions; The method comprises: determining a third amount of the histone modification in the one or more third genomic regions in the same manner as determining the second amount of the histone modification; 31. The method of claim 30, further comprising: determining a third fractional concentration of the third tissue type by solving the system of linear or nonlinear equations including parameters for the third amount of histone modification and the relative amount of each tissue type in the one or more third genomic regions.
32. 1. A method for analyzing a biological sample, wherein the biological sample comprises cell-free DNA fragments, the method comprising: receiving a plurality of sequence reads for the cell-free DNA fragments, the plurality of sequence reads including terminal sequences corresponding to ends of the cell-free DNA fragments; identifying a group of sequence reads that map to one or more genomic regions, each of the one or more genomic regions having a histone modification associated with one or more target tissue types; determining, for each sequence read of said group of sequence reads, one or more sequence motifs that correspond to one or more end sequences of the corresponding cell-free DNA fragment; determining the relative frequency of one or more of each sequence motif of the set of one or more sequence motifs, wherein the set of one or more sequence motifs occurs at a higher rate in chromatin immunoprecipitation of the histone modifications associated with the one or more genomic regions followed by sequencing than in sequencing without chromatin immunoprecipitation; determining a summary value of the one or more relative frequencies; comparing the aggregated value to one or more calibration values; and determining a classification of the level of the disorder using the comparison, wherein the one or more calibration values are determined from one or more calibration samples in which the classification of the level of the disorder is known.
33. the one or more relative frequencies are one or more first relative frequencies; the aggregated value is a first aggregated value, The one or more calibration values are: For each calibration sample of the one or more calibration samples: determining a second relative frequency of said set of one or more sequence motifs in said one or more genomic regions; determining a second aggregate value of the one or more second relative frequencies; and associating each of one or more second aggregate values with a known classification of the level of the impairment, whereby the one or more calibration values comprise the one or more second aggregate values.
34. 33. The method of claim 32, wherein each of the one or more genomic regions has a histone modification associated with one target tissue type.
35. 35. The method of claim 34, wherein the disorder is in the one target tissue type.
36. 35. The method of claim 34, wherein the disorder is a cancer of said one target tissue type.
37. 33. The method of claim 32, wherein the disorder is a pregnancy-related disorder.
38. the aggregated value is a first aggregated value, the one or more calibration values are one or more first calibration values; The method comprises: measuring the size of the cell-free DNA fragments using the sequence reads; and determining one or more size frequencies of said sequence reads for one or more size ranges; determining a second aggregate value of the one or more size frequencies; comparing the second aggregated value to one or more second calibration values; 33. The method of claim 32, wherein determining the classification of the level of the disorder comprises using the comparison of the second aggregate value to the one or more second calibration values, wherein the one or more second calibration values are determined from the one or more calibration samples.
39. 38. The method of any one of claims 7 to 37, wherein the set of one or more sequence motifs comprises 1 to 5, 5 to 10, 11 to 15, 15 to 20, or 20 to 25 sequence motifs.
40. said set of one or more sequence motifs determining a first ratio of each of the one or more sequence motifs relative to other sequence motifs in chromatin immunoprecipitation followed by sequencing; determining a second ratio of each of the set of one or more sequence motifs relative to other sequence motifs in sequencing without chromatin immunoprecipitation; and identifying each of said set of one or more sequence motifs as having a first proportion that is higher than said second proportion.
41. 41. The method of any one of claims 7 to 40, wherein the histone modifications are H3K4me1, H3K4me2, H3K27me3, H3K27ac, H3K36me3, H3K9me2, H3K9me3, H3S10P, H3R2me, H3T2P, H3K14ac, H3K9ac, H3K79me2, H3K79me3, H4K5ac, H4K8ac, H4K12ac, H4K16ac, H4K20me, H2BK120ub, and H2AK119ub.
42. 42. The method of any one of claims 7-41, further comprising sequencing the cell-free DNA fragments in the biological sample to obtain the plurality of sequence reads.
43. 43. The method of claim 42, wherein the volume of the biological sample is 100 μl or less.
44. The method of any one of claims 7 to 43, wherein the biological sample is plasma or serum.
45. 1. A method of enriching a biological sample for clinically relevant DNA, wherein the biological sample contains the clinically relevant DNA and other DNA that is acellular, the method comprising: receiving a plurality of sequence reads of cell-free DNA fragments from the biological sample, the sequence reads including terminal sequences corresponding to ends of the cell-free DNA fragments; determining, for each sequence read of the plurality of sequence reads, one or more sequence motifs corresponding to one or more end sequences of the cell-free DNA fragments; identifying a set of one or more sequence motifs that occur at a higher rate in chromatin immunoprecipitation followed by sequencing of histone modifications in the clinically relevant DNA than in sequencing without chromatin immunoprecipitation; identifying a group of said sequence reads having said set of said one or more sequence motifs in an end sequence; For each sequence read of the group of sequence reads, determining a likelihood that the sequence read corresponds to the clinically relevant DNA based on an end sequence of the sequence read that includes a sequence motif of the set of one or more sequence motifs; comparing the likelihood to a threshold; storing the sequence read when the likelihood exceeds the threshold, thereby obtaining a stored sequence read; and analyzing the stored sequence reads to determine a characteristic of the clinically relevant DNA of the biological sample.
46. 46. The method of Claim 45, wherein said plurality of sequence reads map to one or more predetermined genomic regions, each of said one or more predetermined genomic regions having said histone modification associated with a target tissue type.
47. 47. The method of claim 46, wherein the clinically relevant DNA is DNA from the target tissue type or DNA from diseased tissue.
48. 46. The method of Claim 45, wherein said plurality of sequence reads map to one or more predetermined genomic regions, each of said one or more predetermined genomic regions having said histone modification associated with any of a plurality of target tissue types.
49. 49. The method of any one of claims 45-48, further comprising using the sequence reads to measure sizes of the cell-free DNA fragments, and wherein determining the likelihood that a particular sequence read corresponds to the clinically relevant DNA is further based on sizes of the cell-free DNA fragments corresponding to the particular sequence read.
50. 50. The method of any one of claims 45-49, further comprising measuring one or more methylation states at one or more sites in the cell-free DNA fragments corresponding to a particular sequence read, wherein determining the likelihood that the particular sequence read corresponds to the clinically relevant DNA is further based on the one or more methylation states.
51. 1. A method of enriching a biological sample for clinically relevant DNA, wherein the biological sample contains the clinically relevant DNA and other DNA that is acellular, the method comprising: receiving a plurality of cell-free DNA fragments from the biological sample; the ends of the plurality of cell-free DNA fragments have terminal sequences; receiving, wherein the one or more sequence motifs correspond to one or more terminal sequences of each cell-free DNA fragment of the plurality of cell-free DNA fragments; identifying a set of one or more sequence motifs that occur at a higher rate in chromatin immunoprecipitation followed by sequencing of histone modifications in the clinically relevant DNA than in sequencing without chromatin immunoprecipitation; subjecting the plurality of cell-free DNA fragments to one or more probe molecules that detect the set of one or more sequence motifs in the terminal sequences of the plurality of cell-free DNA fragments, thereby obtaining detected DNA fragments; and using the detected DNA fragments to enrich the biological sample for clinically relevant DNA fragments.
52. 52. The method of claim 51, further comprising analyzing the enriched biological sample to determine a tissue of origin or disease level classification.
53. using the detected DNA fragments to enrich the biological sample for the clinically relevant DNA fragments; 52. The method of claim 51, comprising amplifying the detected DNA fragment.
54. 52. The method of Claim 51, wherein the one or more probe molecules comprise one or more enzymes that interrogate the plurality of cell-free DNA fragments and add new sequences that are used to amplify the detected DNA fragments.
55. using the detected DNA fragments to enrich the biological sample for the clinically relevant DNA fragments; capturing the detected DNA fragments; and discarding undetected DNA fragments.
56. 56. The method of claim 55, wherein one or more probe molecules are attached to a surface and detect the sequence motif in the terminal sequence by hybridization.
57. 1. A method for analyzing a biological sample, wherein the biological sample comprises cell-free DNA fragments, the method comprising: identifying N genomic regions, where N is an integer greater than 1; For each of the M tissue types, obtaining N tissue-specific histone modification levels at the N genomic regions, where N is greater than or equal to M; the tissue-specific histone modification levels form an N×M dimensional matrix A; one of the M tissue types corresponds to a first tissue type; obtaining at least one genomic region among the N genomic regions comprising a non-zero histone modification level from at least two of the M tissue types; receiving an input data vector b comprising N mixed histone modification levels at the N genomic regions, wherein the N mixed histone modification levels are measured from a plurality of cell-free DNA molecules in a biological sample of a subject; and using a computer system, determining the fractional concentration of the first tissue type using matrix A and input data vector b.
58. 58. The method of claim 57, wherein the levels of the N mixed histone modifications are measured by determining the relative frequency of one or more sets of one or more sequence motifs in the plurality of cell-free DNA molecules by chromatin immunoprecipitation followed by sequencing, or by determining the relative frequency of one or more size ranges in the plurality of cell-free DNA molecules.
59. 58. The method of claim 57, wherein the histone modification is H3K27ac or H3K4me3.
60. 58. The method of claim 57, wherein the first tissue type is fetal or erythroblastic tissue.
61. the first tissue type is fetal tissue; The method comprises:
58. The method of claim 57, further comprising using the fractional concentration of the first tissue type to determine a classification of pregnancy in the subject.
62. 58. The method of claim 57, further comprising using the fractional concentration of the first tissue type to determine a disease classification.
63. 1. A method for analyzing a biological sample, wherein the biological sample comprises cell-free DNA fragments, the method comprising: identifying N genomic regions, where N is an integer greater than 1; For each of the M tissue types, obtaining N tissue-specific histone modification levels at the N genomic regions, where N is greater than or equal to M; the tissue histone modification levels form an N×M dimensional matrix A; one of the M tissue types corresponds to a first tissue type; obtaining at least one genomic region among the N genomic regions comprising a non-zero histone modification level from at least two of the M tissue types; receiving an input data vector b comprising N mixed histone modification levels at the N genomic regions, wherein the N mixed histone modification levels are measured from a plurality of cell-free DNA molecules in a biological sample of a subject; Using a computer system, the matrix A, and the input data vector b, a classification of pregnancy in said subject; or and determining a classification of the disease in said subject.
64. determining the classification of the pregnancy or the classification of the disease, inputting the matrix A and the input data vector b into a machine learning model, wherein the machine learning model: storing a plurality of training samples, each of which comprises: one of a plurality of training input data vectors b, wherein the plurality of training input data vectors b are obtained from a plurality of biological samples of a plurality of training subjects; a first label indicating a known classification of the state of the training subject; optimizing parameters of the machine learning model using the plurality of training samples based on an output of the machine learning model that matches or does not match a corresponding label of the first label, when the matrix A and the plurality of training input data vectors b are input to the machine learning model, wherein the output of the machine learning model specifies the classification of the state; and determining the classification of the pregnancy or the classification of the disease using the machine learning model.
65. 1. A method for analyzing a biological sample, wherein the biological sample comprises cell-free DNA fragments, the method comprising: receiving a plurality of sequence reads for the cell-free DNA fragments; identifying a group of sequence reads that map to one or more genomic regions, each of the one or more genomic regions having a histone modification associated with one or more target tissue types; measuring the size of each cell-free DNA fragment corresponding to each sequence read in said group of sequence reads; determining the relative frequency of one or more cell-free DNA fragments having sizes within a set of one or more size ranges, wherein the set of one or more size ranges occurs at a differential rate in chromatin immunoprecipitation followed by sequencing for the histone modifications associated with the one or more genomic regions than in sequencing without chromatin immunoprecipitation; determining a summary value of the one or more relative frequencies; comparing the aggregated value to one or more calibration values; and using said comparison to determine the amount of said histone modification in said biological sample.
66. comparing the amount of the histone modification to one or more second calibration values; 66. The method of claim 65, further comprising determining a fractional concentration of the target tissue type using the comparison of the amount of the histone modification to the one or more second calibration values.
67. comparing the amount of the histone modification to one or more second calibration values; 66. The method of claim 65, further comprising: determining a classification of a level of impairment using the one or more second calibration values.
68. comparing the amount of the histone modification to one or more second calibration values; 66. The method of claim 65, further comprising determining a classification of the implant status of the target tissue type using the one or more second calibration values.
69. 1. A method of analyzing a biological sample from a subject, wherein the biological sample comprises cell-free DNA fragments, the method comprising: receiving a plurality of sequence reads for the cell-free DNA fragments, the plurality of sequence reads including terminal sequences corresponding to ends of the cell-free DNA fragments; identifying a group of sequence reads that map to one or more genomic regions, each of the one or more genomic regions having a histone modification associated with one or more target tissue types; determining, for each sequence read of said group of sequence reads, one or more sequence motifs that correspond to one or more end sequences of the corresponding cell-free DNA fragment; measuring the size of the cell-free DNA fragments using the sequence reads; and for each of said one or more target tissue types: determining a sequence motif frequency of one or more of the set of one or more sequence motifs, wherein the set of one or more sequence motifs occurs at a higher rate in chromatin immunoprecipitation followed by sequencing of the histone modifications associated with the one or more genomic regions than in sequencing without chromatin immunoprecipitation; determining one or more size frequencies of said sequence reads for one or more size ranges; inputting the one or more sequence motif frequencies and the one or more size frequencies for each of the one or more target tissue types into a machine learning model; and using the machine learning model to determine a classification of the subject's condition.
70. 70. The method of claim 69, wherein the condition is pregnancy.
71. 70. The method of claim 69, wherein the condition is a disease.
72. 70. The method of claim 69, wherein the condition is cancer.
73. The machine learning model: storing a plurality of training samples, each training sample comprising: for each of said one or more target tissue types: one or more training sequence motif frequencies of the set of one or more sequence motifs occurring in cell-free DNA fragments in the training sample; and training size frequencies of the cell-free DNA fragments in the training sample; and a first label indicating a known classification of the state; 70. The method of claim 69, wherein the method is trained by: optimizing parameters of the machine learning model based on an output of the machine learning model that matches or does not match a corresponding label of the first label when the sequence motif frequency and the size frequency are input into the machine learning model using the plurality of training samples, wherein the output of the machine learning model specifies the classification of the state.
74. 70. The method of claim 69, wherein the cell-free DNA fragments have a size having a predetermined size range.
75. 75. The method of claim 74, wherein the predetermined size range is 230-350 nt.
76. 70. The method of claim 69, wherein said cell-free DNA fragments consist of fragments having sequence motifs of said set of one or more sequence motifs.
77. for each sequence motif of said set of one or more sequence motifs, determining a size parameter of fragments having said each sequence motif; 70. The method of claim 69, further comprising inputting the one or more size parameters into the machine learning model.
78. 70. The method of claim 69, wherein the one or more target tissue types comprise a cancerous organ or fetal tissue.
79. 70. The method of claim 69, wherein the one or more target tissue types comprise liver, neutrophils, megakaryocytes, or erythroblasts.
80. 70. The method of claim 69, wherein the set of one or more sequence motifs comprises 1 to 5, 5 to 10, 11 to 15, 15 to 20, or 20 to 25 sequence motifs.
81. 70. The method of claim 69, wherein the histone modification is H3K4me1, H3K4me2, H3K27me3, H3K27ac, H3K36me3, H3K9me2, H3K9me3, H3S10P, H3R2me, H3T2P, H3K14ac, H3K9ac, H3K79me2, H3K79me3, H4K5ac, H4K8ac, H4K12ac, H4K16ac, H4K20me, H2BK120ub, or H2AK119ub.
82. 70. The method of claim 69, wherein the machine learning model comprises linear regression, logistic regression, a deep recurrent neural network, a Bayesian classifier, a convolutional neural network (CNN), a hidden Markov model (HMM), a linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering for applications with noise (DBSCAN), a random forest algorithm, or a support vector machine (SVM).
83. A computer product comprising a computer readable medium storing a plurality of instructions for controlling a computer system to perform a method according to any one of the preceding claims.
84. 84. A computer product according to claim 83; and one or more processors for executing instructions stored on the computer-readable medium.
85. A system comprising means for carrying out the method according to any one of the preceding claims.
86. 10. A system comprising one or more processors configured to perform a method according to any one of the preceding claims.
87. A system comprising modules for respectively implementing the steps of the method according to any one of the preceding claims.