Cell-free DNA end characteristics

TWI935564BActive Publication Date: 2026-08-11THE CHINESE UNIVERSITY OF HONG KONG
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
TW113146999
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-12-19
Filing Date
2019-12-19
Publication Date
2026-08-11
Estimated Expiration
2039-12-18

AI Technical Summary

Technical Problem

Existing methods struggle to accurately determine the tissue origin of cell-free DNA fragments in biological samples, particularly for applications such as cancer detection, pregnancy monitoring, and organ transplantation, due to the non-random nature of plasma DNA generation and the need for reference genomes.

Method used

Measuring the relative frequency of sequence terminal motifs in cell-free DNA fragments, which exhibit tissue-specific patterns, allows for the classification and enrichment of clinically relevant DNA without relying on reference genomes, using techniques like physical and electronic hybridization enrichment.

Benefits of technology

Enables accurate classification of tissue-specific DNA, improving cancer detection, pregnancy monitoring, and transplantation assessment by enhancing the specificity and reducing the reliance on reference genomes, thereby increasing accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TB001905499_001
    Figure TWG2TB001905499_001
  • Figure TWG2TB001905499_002
    Figure TWG2TB001905499_002
  • Figure TWG2TB001905499_003
    Figure TWG2TB001905499_003
Patent Text Reader

Abstract

This disclosure describes techniques for measuring the amount (e.g., relative frequency) of sequence terminal motifs of cell-free DNA fragments in a biological sample, for measuring the properties of the sample (e.g., fractional concentration of clinically relevant DNA) and / or for determining the condition of the organism based on such measurements. Different tissue types exhibit different patterns regarding these relative frequencies of sequence terminal motifs. This disclosure provides various uses for measuring, for example, the relative frequencies of sequence terminal motifs of cell-free DNA in a mixture of cell-free DNA from various tissues. DNA from one of these tissues may be referred to as clinically relevant DNA.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-reference to related applications This application is claimed in and is the subject of U.S. Provisional Patent Application No. 62 / 782,316, filed on December 19, 2018, entitled “Cell-FREE DNA End Characteristics,” which is incorporated herein by reference in its entirety for all purposes. Prior Technology

[0002] It is believed that plasma DNA is composed of cell-free DNA excreted from multiple tissues in the body, including but not limited to hematopoietic tissue, brain, liver, lungs, colon, pancreas, etc. (Sun et al., Proceedings of the National Academy of Sciences of the United States of America. 2015; 112: E5503-12; Lehmann-Werman et al., Proceedings of the National Academy of Sciences of the United States of America. 2016; 113: E1826-34; Moss et al., Nature Communications. 2018; 9: 5068). Plasma DNA molecules (a type of free DNA molecule) have been shown to be generated through a non-random process, for example, their size distribution shows a main peak of 166-bp and periodicity of 10-bp in smaller peaks (Lo et al., Sci Transl Med. 2010; 2: 61ra91; Jiang et al., Proceedings of the National Academy of Sciences. 2015; 112: E1317-25).

[0003] Recently, it has been reported that subsets of human genomic locations (e.g., at reference genomic locations) are preferentially cleaved to produce plasma DNA fragments with end positions associated with the tissue of origin (Chan et al., Proceedings of the National Academy of Sciences, 2016; 113:E8159-8168; Jiang et al., Proceedings of the National Academy of Sciences, 2018; doi:10.1073 / pnas.1814616115). Chandrananda et al. (BMC Med Genomics, 2015; 8:29) used the rediscovery software DREME (Bailey, Bioinformatics, 2011; 27:1653-9) to mine cell-free DNA data associated with nuclease breakage motifs regardless of tissue type. Summary of the Invention

[0004] This disclosure describes techniques for measuring the amount (e.g., relative frequency) of sequence terminal motifs in cell-free DNA fragments in a biological sample, to measure the properties of the sample (e.g., fractional concentration of clinically relevant DNA) and / or to determine the condition of the organism based on such measurements. Different tissue types exhibit different patterns regarding the relative frequency of sequence terminal motifs. This disclosure provides various uses for measuring, for example, the relative frequency of sequence terminal motifs in cell-free DNA in a mixture of cell-free DNA from various tissues. DNA from one of these tissues may be referred to as clinically relevant DNA.

[0005] Various examples can quantify the quantity of sequence motifs (terminal motifs) representing the terminal sequences of DNA fragments. For example, embodiments can determine the relative frequencies of a set of sequence motifs used for the terminal sequences of DNA fragments. In various embodiments, genotypic (e.g., tissue-specific paired genes) or phenotype-based methods (e.g., using samples with similar conditions) can be used to determine preferred sets of terminal motifs and / or patterns of terminal motifs. Preferred sets or relative frequencies of specific patterns can be used to quantify the classification of the nature of a new sample or organism's condition (e.g., gestational age or pathological grade of a fetus) (e.g., fractional concentration of clinically relevant DNA). Thus, embodiments can provide measurements to inform physiological changes, including cancer, autoimmune diseases, transplantation, and pregnancy.

[0006] As another example, sequence end motifs can be used for the physical enrichment and / or electronic hybridization enrichment of clinically relevant cell-free DNA fragments in biological samples. Enrichment can utilize sequence end motifs preferred for clinically relevant tissues (such as fetuses, tumors, or transplants). Physical enrichment can use one or more probe molecules that detect a specific set of sequence end motifs, thereby enriching the biological sample with clinically relevant DNA fragments. For electronic hybridization enrichment, a set of sequence reads can be identified of cell-free DNA fragments having one of a set of preferred end sequences for clinically relevant DNA. Certain sequence reads can be stored based on the probability corresponding to clinically relevant DNA, where the probability takes into account sequence reads containing preferred sequence end motifs. The stored sequence reads can be analyzed to determine the nature of the clinically relevant DNA in the biological sample.

[0007] These and other embodiments of the present disclosure are described in detail below. For example, other embodiments are directed to systems, apparatuses, and computer-readable media relating to the methods described herein.

[0008] A better understanding of the nature and advantages of the embodiments of this disclosure can be obtained by referring to the following detailed description and accompanying drawings. Simple Explanation of the Diagram

[0009] Figure 1 shows an example of an end unit according to an embodiment of this disclosure.

[0010] Figure 2 illustrates a schematic diagram of a method based on genotype differences according to an embodiment of this disclosure, which is used to analyze differential terminal motif patterns between fetal and maternal DNA molecules.

[0011] Figure 3 shows a bar graph of the terminal motif frequencies between fetal and maternal DNA molecules according to an embodiment of this disclosure.

[0012] Figure 4 shows 10 terminal motifs of the fetus and shared (i.e., fetus plus mother) sequence from Figure 3 according to an embodiment of this disclosure.

[0013] Figures 5A and 5B show box plots of the entropy between fetal and maternal DNA molecules in pregnant women according to embodiments of this disclosure.

[0014] Figures 6A and 6B illustrate hierarchical clustering analysis of fetal and maternal DNA molecules according to embodiments of this disclosure.

[0015] Figures 7A and 7B illustrate the entropy distribution of all primitives used for pregnant women at different stages of pregnancy according to an embodiment of this disclosure. Figures 7C and 7D illustrate the entropy distribution of 10 primitives used for pregnant women at different stages of pregnancy according to an embodiment of this disclosure.

[0016] Figure 8A shows the entropy of all fragments at different gestational ages. It is shown that the entropy of plasma DNA fragments in individuals with late pregnancy is lower than that in individuals with early and mid-pregnancy (p=0.06). Figure 8B shows the entropy of Y-chromosome-derived fragments at different gestational ages. It is shown that the entropy of Y-chromosome-derived fragments in individuals with late pregnancy is lower than that in individuals with early and mid-pregnancy (p=0.01).

[0017] Figures 9 and 10 show the distribution of the terminal motifs of the first 10 arrangements between fetal and maternal DNA molecules at different stages of pregnancy, according to an embodiment of this disclosure.

[0018] Figure 11 shows the combination frequency of the first 10 arrangements of the basic units between the fetus and the shared molecules at different stages of pregnancy, according to an embodiment of this disclosure.

[0019] Figure 12 illustrates a schematic diagram of a genotype-based method according to an embodiment of this disclosure, used to analyze differential terminal motif patterns between mutations and shared molecules in the plasma DNA of cancer patients.

[0020] Figure 13 shows a diagram of the plasma DNA terminal motifs of cancer-related mutations and common molecules in hepatocellular carcinoma according to an embodiment of this disclosure.

[0021] Figure 14 shows a radial diagram of plasma DNA terminal motifs of cancer-related mutations and common molecules in hepatocellular carcinoma according to an embodiment of this disclosure.

[0022] Figure 15A shows the top 10 end motifs in the arrangement differences of end motif frequencies between mutations and common sequences in the plasma DNA of HCC patients according to an embodiment of this disclosure.

[0023] Figure 15B shows the combined frequencies of eight terminal motifs in HCC patients and pregnant women according to embodiments of this disclosure.

[0024] Figures 16A and 16B illustrate the entropy values ​​of the shared and mutated fragments of different sets of terminal motifs for HCC cases according to embodiments of this disclosure.

[0025] Figure 17 is a graph of the primitive diversity score (entropy) relative to the measured circulating tumor DNA score according to an embodiment of this disclosure.

[0026] Figure 18A illustrates entropy analysis using donor-specific fragments according to an embodiment of this disclosure. Figure 18B illustrates hierarchical clustering analysis using donor-specific fragments.

[0027] Figure 19 is a flowchart of an embodiment according to this disclosure, illustrating a method for estimating the fractional concentration of clinically relevant DNA in an individual's biological sample.

[0028] Figure 20 is a flowchart according to an embodiment of the present disclosure, illustrating a method for determining the gestational age of a fetus by analyzing biological samples from a pregnant woman. [。]

[0029] Figure 21 shows a schematic diagram of a phenotypic method for plasma DNA terminal motif analysis according to an embodiment of this disclosure.

[0030] Figure 22 illustrates an example of the frequency distribution of tetramer terminal motifs among HCC and HBV individuals using all plasma DNA molecules, according to an embodiment of this disclosure.

[0031] Figure 23A shows a box plot of the combination frequencies of the top 10 plasma DNA tetramer terminal motifs in various individuals with different cancer grades according to embodiments of the present disclosure. These grades are: Control: healthy control individuals; HBV: chronic hepatitis B carriers; Cirr: individuals with cirrhosis; eHCC: early HCC; iHCC: intermediate HCC; aHCC: late HCC. Figure 23B shows the receiver operating characteristic (ROC) curves of the combination frequencies of the top 10 plasma DNA tetramer terminal motifs between HCC and non-cancer individuals according to embodiments of the present disclosure.

[0032] Figure 24A shows a box plot of the frequencies of CCA motifs in different groups according to an embodiment of the present disclosure. Figure 24B shows the ROC curves of the most frequent trimer motif (CCA) present in non-HCC individuals between the non-HCC and HCC groups according to an embodiment of the present disclosure.

[0033] Figure 25A shows a box plot of entropy values ​​for different groups using 256 tetramer terminal motifs according to an embodiment of the present disclosure. Figure 25B shows a box plot of entropy values ​​for different groups using 10 tetramer terminal motifs according to an embodiment of the present disclosure.

[0034] Figure 26A shows a box plot of the entropy values ​​of the trimer terminal motifs used in different groups according to an embodiment of this disclosure. It was found that the entropy of HCC individuals using trimer motifs (a total of 64 motifs) was significantly higher than that of non-HCC individuals (p < 0.0001). Figure 26B shows the ROC curves of the entropy of the 64 trimer motifs used in the non-HCC and HCC groups according to an embodiment of this disclosure. The AUC was found to be 0.872.

[0035] Figures 27A and 27B show box plots of primitive diversity (entropy) scores for different groups of 4-mers according to embodiments of this disclosure.

[0036] Figure 28 shows recipient operation curves for various techniques used to distinguish between healthy control groups and cancer according to embodiments of this disclosure.

[0037] Figure 29 shows receiver operating curves for MDS analysis using various k-mers according to an embodiment of this disclosure.

[0038] Figure 30 illustrates the performance of MDS-based cancer detection for various tumor DNA scores according to embodiments of this disclosure.

[0039] Figure 31 shows receiver operating curves for MDS, SVM, and logistic regression analysis according to embodiments of this disclosure. []

[0040] Figure 32 illustrates a hierarchical clustering analysis of the top 10 terminal motifs of individuals with different cancer grades in different groups, according to an embodiment of this disclosure. The different groups include: control group: healthy control individuals; HBV: chronic hepatitis B virus carriers; Cirr: individuals with cirrhosis; eHCC: early-stage HCC; iHCC: immediate-stage HCC; aHCC: late-stage HCC.

[0041] Figures 33A-33C illustrate, according to an embodiment of this disclosure, hierarchical clustering analysis of all plasma DNA molecules in different groups with different cancer grades.

[0042] Figure 34 illustrates an embodiment of this disclosure, based on hierarchical clustering analysis using trimer motifs of all plasma DNA molecules in different groups with different cancer grades.

[0043] Figure 35A illustrates an entropy analysis of all plasma DNA molecules between healthy control individuals and SLE patients according to an embodiment of this disclosure. Figure 35B illustrates a hierarchical clustering analysis of all plasma DNA molecules between healthy control individuals and SLE patients according to an embodiment of this disclosure.

[0044] Figure 36 illustrates an entropy analysis of plasma DNA molecules with 10 selected terminal motifs between healthy control individuals and SLE patients according to an embodiment of this disclosure.

[0045] Figure 37 shows the ROC curves for combined analysis of terminal motifs and copy number or methylation according to an embodiment of this disclosure.

[0046] Figure 38A illustrates a tetramer-based entropy analysis according to an embodiment of this disclosure, wherein the tetramer is constructed from the ends of ordered plasma DNA fragments and their adjacent genomic sequences in HCC and non-HCC individuals. Figure 38B illustrates a tetramer-based clustering analysis according to an embodiment of this disclosure, wherein the tetramer is constructed from the ends of ordered plasma DNA fragments and their adjacent genomic sequences in HCC and non-HCC individuals.

[0047] Figure 39 illustrates a ROC comparison of techniques 140 and 160 used in Figure 1 according to an embodiment of this disclosure to define the terminal motifs of plasma DNA.

[0048] Figure 40 shows a comparison of the accuracy of embodiments according to this disclosure, demonstrating that tissue-specific open chromatin regions improve the ability to distinguish plasma DNA terminal motifs.

[0049] Figure 41 illustrates plasma DNA end motif analysis based on size bands according to an embodiment of this disclosure.

[0050] Figure 42 is a flowchart according to an embodiment of this disclosure, illustrating a method for classifying pathological grades in an individual's biological sample. [。] []

[0051] Figure 43 is a flowchart of an embodiment of the present disclosure, illustrating a method for enriching clinically relevant DNA in biological samples. []

[0052] Figure 44 is a flowchart of an embodiment according to this disclosure, illustrating a method 3700 for enriching clinically relevant DNA in biological samples.

[0053] Figure 45 shows an example graph of an embodiment according to this disclosure, illustrating the increase in fetal DNA fraction using CCCA terminal motifs.

[0054] Figure 46 illustrates a measurement system according to an embodiment of the present invention.

[0055] Figure 47 shows a block diagram of an example computer system that can be used with the system and method according to an embodiment of the present invention. Implementation

[0056] the term A "tissue" corresponds to a group of cells that are collectively classified as a functional unit. More than one type of cell can be found in a single tissue. Different types of tissues can be composed of different types of cells (e.g., hepatocytes, alveolar cells, or blood cells), but can also correspond to tissues from different organisms (mother and fetus) or to healthy cells and tumor cells. A "reference tissue" corresponds to the tissue used to determine tissue-specific methylation levels. Multiple samples of the same tissue type from different individuals can be used to determine the tissue-specific methylation levels of that tissue type.

[0057] "Biological sample" refers to any sample obtained from an individual (such as a human (or other animal), such as a pregnant woman, a person with cancer or suspected of having cancer, an organ transplant recipient, or an individual suspected of having a disease process involving an organ (e.g., the heart in a myocardial infarction, the brain in a stroke, or the hematopoietic system in anemia), and containing one or more relevant nucleic acid molecules. Biological samples can be bodily fluids, such as blood, plasma, serum, urine, vaginal fluid, fluid from scrotal edema (e.g., testicles), vaginal douches, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple discharge, aspirated fluid from different parts of the body (e.g., thyroid gland, breast), intraocular fluid (e.g., aqueous humor), etc. Fecal samples may also be used. In various embodiments, the majority of DNA in a cell-free DNA-enriched biological sample (e.g., a plasma sample obtained via centrifugation) may be cell-free, for example, greater than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA may be cell-free. The centrifugation protocol may include, for example, obtaining a fluid fraction at 3,000 g for 10 minutes, and then centrifuging again at 30,000 g for an additional 10 minutes to remove residual cells. As part of the biological sample analysis, at least 1,000 cell-free DNA molecules may be analyzed. As other examples, at least 10,000, 50,000, 100,000, 500,000, 1,000,000, or 5,000,000 or more cell-free DNA molecules may be analyzed.

[0058] "Clinically relevant DNA" refers to DNA from a specific tissue source that is to be measured, for example, to determine the fractional concentration of such DNA or to classify the phenotype of a sample (e.g., plasma). Examples of clinically relevant DNA include fetal DNA in maternal plasma or tumor DNA in patient plasma or other cell-free DNA samples. Another example includes measuring the amount of graft-associated DNA in the plasma, serum, or urine of transplant patients. Another example includes measuring the fractional concentration of DNA from hematopoietic and non-hematopoietic tissues in an individual's plasma, or the fractional concentration of liver DNA fragments (or other tissues) in a sample, or the fractional concentration of brain DNA fragments in cerebrospinal fluid.

[0059] "Sequence read" refers to a nucleotide string sequenced from any part or all of a nucleic acid molecule. For example, a sequence read can be a short nucleotide string (e.g., 20-150 nucleotides) sequenced from a nucleic acid fragment, a short nucleotide string at one or both ends of a nucleic acid fragment, or the sequence of an entire nucleic acid fragment present in a biological sample. Sequence reads can be obtained in various ways, such as using sequencing techniques or probes, such as hybridization arrays or capture probes, or amplification techniques, such as polymerase chain reaction (PCR) or linear or isothermal amplification using a single primer. As part of a biological sample analysis, at least 1,000 sequence reads can be analyzed. As other examples, at least 10,000, 50,000, 100,000, 500,000, 1,000,000, or 5,000,000 or more sequence reads can be analyzed.

[0060] A read may contain a "terminal sequence" associated with the end of the fragment. The terminal sequence may correspond to the outermost N bases of the fragment, for example, the last 2-30 bases. If a read corresponds to the entire fragment, then the read may contain two terminal sequences. When bilateral sequencing provides two reads corresponding to the ends of a fragment, each read may contain one terminal sequence.

[0061] A "sequence motif" can refer to a short, repetitive pattern of bases in a DNA fragment (e.g., a free DNA fragment). Sequence motifs can appear at the ends of a fragment and are therefore part of or contain end sequences. A "terminal motif" can refer to a sequence motif of a terminal sequence that preferentially appears at the ends of a DNA fragment, possibly for a particular type of tissue. Terminal motifs can also appear exactly before or after the ends of a fragment and thus still correspond to end sequences.

[0062] The term "pair gene" refers to an alternative DNA sequence at a locus on the same genome, which may or may not result in different phenotypic traits. In any given diploid organism with two copies of each chromosome (except for the sex chromosome in male humans), the genotype of each gene includes the pair of pairs of genes present at that locus, which are identical in isozygotes and different in heterozygotes. A population or species of organism typically contains multiple pairs of genes at each locus in each individual. Genomic loci where more than one pair of genes is found in a population are called polymorphic loci. Pair gene variation at a locus can be measured as the number of pairs of genes present in the population (i.e., the degree of polymorphism) or the proportion of heterozygotes (i.e., the heterozygosity ratio). As used herein, the term "polymorphism" refers to any inter-individual variation in the human genome, regardless of its frequency. Examples of such variations include (but are not limited to) single nucleotide polymorphisms, simple tandem repeat polymorphisms, insertion-deletion polymorphisms, mutations (which can be pathogenic), and duplicate number variations. As used herein, the term "haplotype" refers to a combination of paired genes at multiple loci transmitted together on the same chromosome or chromosomal region. A haplotype can refer to as few as one pair of loci, a chromosomal region, or an entire chromosome or set of chromosomes.

[0063] The term "fetal DNA fraction concentration" is used interchangeably with the terms "fetal DNA proportion" and "fetal DNA fraction," and refers to the proportion of fetal DNA molecules present in a biological sample derived from the fetus (e.g., maternal plasma or serum sample) (Lo et al., *American Journal of Human Genetics*, 1998; 62: 768-775; Lun et al., *Clinical Chemistry*, 2008; 54: 1664-1672). Similarly, tumor fraction or tumor DNA fraction can refer to the fractional concentration of tumor DNA in a biological sample.

[0064] "Relative frequency" can refer to a proportion (e.g., percentage, fraction, or concentration). Specifically, the relative frequency of a particular terminal motif (e.g., CCGA) can be provided, for example, by the proportion of free DNA fragments associated with the terminal motif CCGA having a terminal sequence of CCGA.

[0065] "Total value" can refer to collective properties, such as the relative frequencies of a set of terminal primitives. Examples include the mean, median, sum of relative frequencies, variation between relative frequencies (e.g., entropy, standard deviation (SD), coefficient of variation (CV), interquartile range (IQR), or a cutoff value of a percentage point between different relative frequencies (e.g., the 95th or 99th percentage point)), or the difference from a reference pattern of relative frequencies (e.g., distance).

[0066] A "calibration sample" can correspond to a biological sample whose clinically relevant DNA fraction concentration (e.g., tissue-specific DNA fraction) is known or determined via calibration methods, such as using tissue-specific paired genes, as in transplantation, where the paired gene is present in the donor's genome but not in the recipient's genome and can be used as a marker for the transplanted organ. As another example, a calibration sample can correspond to a sample from which terminal motifs can be determined. Calibration samples can be used for two purposes.

[0067] A "calibration data point" comprises a "calibration value" and a measured or known fractional concentration of clinically relevant DNA (e.g., DNA of a specific tissue type). The calibration value can be determined from a relative frequency (e.g., total value) defined for the calibration sample, for which the fractional concentration of clinically relevant DNA is known. Calibration data points can be defined in various ways, such as as discrete points or as a calibration function (also known as a calibration curve or calibration surface). The calibration function may be derived from additional mathematical transformations of the calibration data points.

[0068] A "site" (also called a "genomic site") corresponds to a single location, which can be a single base position or a group of related base positions, such as a CpG site or a larger group of related base positions. A "locus" can correspond to a region containing multiple sites. A locus can contain only one site, which would make the locus equivalent to a single site in that case.

[0069] The "methylation index" of each genomic site (e.g., a CpG site) refers to the proportion of DNA fragments showing methylation at that site (e.g., as determined by a sequence read or probe) to the total number of reads covering that site. A "read" corresponds to information obtained from a DNA fragment (e.g., the methylation state at the site). Reads can be obtained using reagents (e.g., primers or probes) that preferentially hybridize to DNA fragments with specific methylation states. Typically, such reagents are applied after treatment with methods that modify or recognize DNA molecules differently depending on their methylation state, such as bisulfite conversion, methylation-sensitive restriction enzymes, methylation-binding proteins, or anti-methylcytosine antibodies, or single-molecule sequencing techniques that recognize, for example, methylcytosine and hydroxymethylcytosine.

[0070] The "methylation density" of a region can be defined as the number of reads at sites within the methylated region divided by the total number of reads at sites within the covered region. Sites can have specific characteristics, such as CpG sites. Therefore, the "CpG methylation density" of a region can be defined as the number of reads showing CpG methylation divided by the total number of reads at CpG sites within the covered region (e.g., specific CpG sites, CpG islands, or CpG sites within larger regions). For example, the methylation density per 100 kb facet in the human genome can be determined as the proportion of all CpG sites covered by a sequence read mapped to a 100 kb region, based on the total number of unconverted cytosines (corresponding to methylated cytosines) at CpG sites after bisulfite treatment. This analysis can also be performed for other facet sizes, such as 500 bp, 5 kb, 10 kb, 50 kb, or 1 Mb. A region can be the entire genome, a chromosome, or a portion of a chromosome (e.g., a set of chromosomes). When a region contains only CpG sites, the methylation index of the CpG sites is the same as the methylation density of the region. "The proportion of methylated cytosine" refers to the number of methylated (e.g., unconverted after bisulfite conversion) cytosine sites "C" compared to the total number of cytosine residues analyzed; that is, cytosine in the region excluding the CpG background. The methylation index, methylation density, and proportion of methylated cytosine are examples of "methylation level." Besides bisulfite conversion, other methods known to those skilled in this technique can be used to query the methylation state of DNA molecules, including (but not limited to) methylation-sensitive enzymes (e.g., methylation-sensitive restriction enzymes), methylation-binding proteins, single-molecule sequencing using methylation-sensitive platforms (e.g., nanopore sequencing (Schreiber et al., Proc Natl Acad Sci USA, 2013; 110: 18910-18915), and single-molecule real-time analysis by Pacific Biosciences (Flusberg et al., Nat Methods, 2010; 7: 461-465)). The methylation measure of a DNA molecule can correspond to the percentage of methylated sites (e.g., CpG sites). The methylation measure can be specified as an absolute number or a percentage and can be referred to as the methylation density of the molecule.

[0071] The term "sequencing depth" refers to the number of times a locus is covered by sequence reads aligned to that locus. A locus can be as small as a nucleotide, as large as a chromosome set, or as large as the entire genome. Sequencing depth can be expressed as 50×, 100×, etc., where "×" indicates the number of times the locus is covered by sequence reads. Sequencing depth can also be applied to multiple loci or the entire genome; in this case, "×" can refer to the average number of times the locus, haploid genome, or entire genome is sequenced separately. Ultra-deep sequencing can specify a sequencing depth of at least 100×.

[0072] A "separation value" corresponds to a difference or ratio involving two values, such as two fractional contributions or two methylation levels. A separation value can be a simple difference or ratio. As an example, the ratio of x / y to x / (x+y) is a separation value. Separation values ​​can include other factors, such as multiplicative factors. As another example, the difference or ratio can be a function of the equivalent value, such as the difference or ratio of the natural logarithms (ln) of two values. A separation value can include both differences and ratios.

[0073] "Separation value" and "total value" (e.g., relative frequency) are two instances of parameters (also called measures) that provide a measure of the variation of samples between different categories (states) and can therefore be used to determine different categories. The total value can be a separation value, for example, when taking the difference between a set of relative frequencies of a sample and a reference set of relative frequencies, as can be done in clustering.

[0074] As used herein, the term "classification" refers to any number or other character associated with a specific property of a sample. For example, a "+" symbol (or the word "positive") may indicate that a sample is classified as having a deletion or an amplification. Classifications can be binary (e.g., positive or negative) or have a higher level of classification (e.g., levels from 1 to 10 or 0 to 1).

[0075] The terms "cutoff value" and "threshold value" refer to predetermined numbers used in the operation. For example, a cutoff size may refer to a size to which a segment longer than this size is excluded. A threshold value may be a value above or below which a particular classification applies. Either of these terms may be used in either of these situations. A cutoff value or threshold value may be a "reference value" or derived from a reference value that represents a particular classification or distinguishes between two or more classifications. As those skilled in the art will understand, this reference value can be determined in various ways. For example, a measure may be determined for individuals in two different groups with different known classifications, and a reference value may be selected to represent a classification (e.g., the mean) or a value between two clusters of measures (e.g., selected to obtain the desired sensitivity and specificity). As another example, a reference value may be determined based on statistical simulation of a sample.

[0076] The term "cancer grade" refers to the presence or absence of cancer, cancer stage, tumor size, presence or absence of metastasis, total tumor burden in the body, response to treatment, and / or other measures of cancer severity (e.g., cancer recurrence). Cancer grades can be numbers or other markers, such as symbols, letters, and colors. Grades can be zero. Cancer grades can also include precancerous or precancerous conditions (statuses). Cancer grades can be used in various ways. For example, screening can check whether someone previously unaware of cancer has cancer. Assessment can investigate someone already diagnosed with cancer to monitor cancer development over time, study the effectiveness of treatments, or determine prognosis. In one embodiment, prognosis can be expressed as the probability of a patient dying from cancer or the probability or extent of cancer progression or metastasis after a specific period or time. Detection can mean "screening" or checking whether someone has cancer-indicating characteristics (e.g., symptoms or other positive tests).

[0077] "Pathological grade" can refer to the quantity, extent, or intensity of a biologically relevant pathology, where the grade can be as described above for cancer. Another example of pathology is transplant rejection. Other examples of pathology may include autoimmune attacks (e.g., lupus nephritis damaging the kidneys or multiple sclerosis), inflammatory diseases (e.g., hepatitis), fibrotic processes (e.g., cirrhosis), fatty infiltration (e.g., fatty liver disease), degenerative processes (e.g., Alzheimer's disease), and ischemic tissue damage (e.g., myocardial infarction or stroke). An individual's health status can be considered as a non-pathological classification.

[0078] The terms "about" or "approximately" may mean within an acceptable margin of error for a particular value as determined by someone generally skilled in the art, depending in part on how the value is measured or determined (i.e., the limitations of the measurement system). For example, according to practice in the art, "about" may mean within one or more standard deviations. Alternatively, "about" may mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Or, particularly with respect to biological systems or methods, the terms "about" or "approximately" may mean within a certain order of magnitude of the value, within 5 times, and more preferably within 2 times. If a particular value is described in the claims of this application and the claims, unless otherwise stated, the term "about" should be assumed to mean within an acceptable margin of error for the particular value. The term "about" may have the meaning as commonly understood by someone generally skilled in the art. The term "about" may refer to ±10%. The term "about" may refer to ±5%. Detailed description

[0079] This disclosure describes techniques for measuring the quantity (e.g., relative frequency) of terminal motifs of cell-free DNA fragments in a biological sample of an organism, to measure the properties of the sample and / or to determine the condition of the organism based on such measurements. Different tissue types exhibit different patterns of relative frequency of these sequence motifs. This disclosure provides various uses for measuring the relative frequency of terminal motifs of cell-free DNA, for example, in mixtures of cell-free DNA from various tissues. DNA from one of these tissues may be referred to as clinically relevant DNA.

[0080] Clinically relevant DNA from a specific tissue (e.g., fetus, tumor, or transplanted organ) exhibits a specific pattern of relative frequency, which can be measured as a total value. Other DNA in a sample may exhibit different patterns, allowing the amount of clinically relevant DNA in the sample to be measured. Thus, in one instance, the fractional concentration (e.g., percentage) of clinically relevant DNA can be determined based on the relative frequency of terminal motifs. The fractional concentration can be a number, a numerical range, or other classification (e.g., high, medium, or low), or whether the fractional concentration exceeds a threshold value. In various embodiments, the total value can be the sum of the relative frequencies of a set of terminal motifs, the variance (e.g., entropy, also known as a motif diversity score) of the relative frequencies of all or a set of terminal motifs, or the difference in a reference pattern (e.g., total distance), such as an array (vector) of relative frequencies for a calibration sample with known fractional concentrations. Such an array can be considered as a reference set of relative frequencies. Such differences can be used in classifiers, with hierarchical clustering, support vector machines, and logistic regression being examples. As an example, clinically relevant DNA can be DNA from fetus, tumor, transplanted organ, or other tissues (e.g., hematopoietic tissue or liver).

[0081] In another example, motif relative frequencies can be used to determine pathological grade. Organisms with different phenotypes can exhibit different patterns of motif relative frequencies of cell-free DNA fragments. The total value of the relative frequencies of terminal motifs can be compared to a reference value to classify the phenotype. In various implementations, the total value can be the sum of relative frequencies, the variance of relative frequencies, or the difference from a reference set of relative frequencies. Example pathologies include cancer and autoimmune diseases such as SLE.

[0082] In another instance, the relative frequencies of primitives can be used to determine the gestational age of the fetus. The total relative frequencies of terminal primitives in the maternal sample change as the gestational age of the fetus increases. Such totals can be determined as described above and elsewhere.

[0083] Given that cell-free DNA fragments from a particular tissue possess a specific set of preferred end motifs, these preferred end motifs can be used to enrich samples of DNA from certain specific tissues (clinically relevant DNA). Such enrichment can be performed through physical manipulation to enrich physical samples. Some embodiments, such as using primers or appendages, can capture and / or amplify cell-free DNA fragments with end sequences matching a set of preferred end motifs. Other examples are described herein.

[0084] In some embodiments, enrichment can be performed using electronic hybridization. For example, the system can receive sequence reads and then filter them based on end motifs to obtain a subset of sequence reads corresponding to DNA fragments with higher concentrations of clinically relevant DNA. If a DNA fragment has a terminal sequence containing a preferred end motif, it can be identified as having a higher probability of originating from the tissue of interest. As described herein, probability can be further determined based on the methylation and size of the DNA fragment.

[0085] Such uses of terminal motifs can avoid the need for a reference genome, which might be required when using terminal positions (Chan et al., Proceedings of the National Academy of Sciences of the United States of America, 2016; 113: E8159-8168; Jiang et al., Proceedings of the National Academy of Sciences of the United States of America, 2018; doi:10.1073 / pnas.1814616115). Furthermore, since the number of terminal motifs can be less than the number of preferred terminal positions in the reference genome, more statistical data can be collected for each terminal motif, potentially improving accuracy.

[0086] The ability to use terminal motifs in this way is surprising. For example, Chandrananda et al. found a high degree of similarity between maternal and fetal fragments in terms of position-specific nucleotide patterns, with regard to the single nucleotide frequency in a 51 bp region (upstream / downstream 20 bp) around the fragment start site (Chandrananda et al., BMC Med Genomics, 2015; 8:29), meaning that methods based on single nucleotide frequencies around the ends cannot indicate the origin of tissue-free DNA fragments. [I.] [free] [DNA] [Terminal primitive] []

[0087] A terminal motif relates to the end segment of a cell-free DNA, such as a terminal sequence representing a sequence of K bases at any end of that segment. The terminal sequence can be a k-mer with various numbers of bases, such as 1, 2, 3, 4, 5, 6, 7, etc. A terminal motif (or "sequence motif") refers to the sequence itself opposite to a specific location in a reference genome. Therefore, the same terminal motif may appear at many locations throughout the reference genome. A reference genome can be used to determine terminal motifs, for example, to identify bases exactly before or exactly after the start position. Such bases will still correspond to the ends of the cell-free DNA segment, for example, because this identification is based on the terminal sequence of the segment.

[0088] Figure 1 illustrates examples of terminal motifs according to embodiments of this disclosure. Figure 1 depicts two methods for defining the tetramer terminal motif to be analyzed. In technique 140, the tetramer terminal motif is constructed directly from the first 4-bp sequence at each end of a plasma DNA molecule. For example, the first 4 nucleotides or the last 4 nucleotides of a sequence fragment can be used. In technique 160, the tetramer terminal motif is constructed by utilizing the dimer sequence from the sequence ends of a fragment and other dimer sequences from adjacent genomic regions. In other embodiments, other types of motifs, such as 1-mer, 2-mer, 3-mer, 5-mer, 6-mer, and 7-mer terminal motifs, can be used.

[0089] As shown in Figure 1, the cell-free DNA fragment 110 is obtained, for example, by purification treatment of a blood sample, such as by centrifugation. Besides plasma DNA fragments, other types of cell-free DNA molecules can also be used, such as those from serum, urine, saliva, and other such cell-free samples mentioned herein. In one embodiment, the DNA fragment may be blunt-ended.

[0090] In step 120, bilateral sequencing is performed on the DNA fragment. In some embodiments, bilateral sequencing may generate two reads from both ends of the DNA fragment, for example, each read being 30-120 bases long. These two reads form a pair of reads for a DNA fragment (molecule), wherein each read contains the terminal sequences of the corresponding ends of the DNA fragment. In other embodiments, the entire DNA fragment may be sequenced to provide a single read containing the terminal sequences from both ends of the DNA fragment.

[0091] In step 130, the sequence read can be aligned with a reference genome. This alignment is used to illustrate different ways of defining sequence motifs and may not be used in some embodiments. Various software suites can be used to perform the alignment procedure, such as BLAST, FASTA, Bowtie, BWA, BFAST, SHRiMP, SSAHA2, NovoAlign, and SOAP.

[0092] Technique 140 illustrates a sequence read of sequence fragment 141 aligned with genome 145. With the 5' end considered as the starting point, a first end motif 142 (CCCA) is located at the start of sequence fragment 141. A second end motif 144 (TCGA) is located at the tail of sequence fragment 141. In one embodiment, such end motifs may occur when an enzyme recognizes CCCA and subsequently cleaves exactly before the first C. If this is the case, CCCA will preferentially be at the end of the plasma DNA fragment. For TCGA, the enzyme may recognize it and subsequently cleave after the A.

[0093] Technique 160 illustrates a sequence read of sequence fragment 161 aligned with genome 165. With the 5' end considered as the starting point, the first end motif 162 (CGCC) has a first portion (CG) appearing exactly before the starting point of sequence fragment 161 and a second portion (CC) representing the terminal sequence portion of the starting point of sequence fragment 161. The second end motif 164 (CCGA) has a first portion (GA) appearing exactly after the tail of sequence fragment 161 and a second portion (CC) representing one of the terminal sequences of the tail of sequence fragment 161. In one embodiment, such end motifs may appear when an enzyme recognizes CGCC and subsequently cleaves between G and C. If this is the case, CC will preferentially be at the end of the plasma DNA fragment, with CG appearing exactly before it, thus providing the end motif of CGCC. As for the second end motif 164 (CCGA), the enzyme can cleave between C and G. If this is the case, CC will preferentially be at the end of the plasma DNA fragment. For technique 160, the number of bases from adjacent genomic regions and ordered plasma DNA fragments can vary and is not limited to a fixed ratio. For example, instead of 2:2, the ratio can be 2:3, 3:2, 4:4, 2:4, etc.

[0094] The higher the number of nucleotides contained in the end marker of free DNA, the higher the specificity of the motif, because the probability of having 6 ordered bases in the precise configuration of the genome is lower than the probability of having 2 ordered bases in the precise configuration of the genome. Therefore, the choice of end motif length can be adjusted according to the required sensitivity and / or specificity for the intended application.

[0095] Because the terminal sequences are used to align the sequence reads with a reference genome, any sequence motif determined by the terminal sequences, or exactly before / after them, is still determined by the terminal sequences. Therefore, technique 160 establishes the relationship between the terminal sequences and other bases, with the reference used as the mechanism for establishing this relationship. The difference between techniques 140 and 160 lies in assigning specific DNA fragments to those two terminal motifs, which affects specific values ​​of relative frequency. However, overall results (e.g., fractional concentrations of clinically relevant DNA, pathological grade classifications, etc.) will not be affected by how DNA fragments are assigned as terminal motifs, as long as consistent techniques are used for data training and production.

[0096] The relative frequencies of DNA fragments having terminal sequences corresponding to specific terminal motifs are determined by counting (e.g., in an array stored in memory). As described in more detail below, the relative frequencies of terminal motifs in free DNA fragments can be analyzed. Differences in the relative frequencies of terminal motifs have been detected from different tissue types and different phenotypes, such as different pathological grades. Differences can be quantified within a set of terminal motifs (e.g., all possible combinations of k-mers corresponding to the lengths used) by the quantity or overall pattern of DNA fragments having specific terminal motifs, such as variance (e.g., entropy, also known as motif diversity score). [II.] [Methods based on genotype differences] []

[0097] We have identified different tissue types with different terminal motifs. In this paper, we describe how terminal motifs can be used to determine the fractional concentration of clinically relevant DNA, such as fetal DNA, tumor DNA, DNA from transplanted organs, or DNA from a specific organ.

[0098] To identify terminal motifs of clinically relevant DNA that favor a specific type, genotypic differences can be used to identify DNA fragments originating from clinically relevant tissues. Once a DNA fragment is detected as originating from a clinically relevant tissue, its terminal motif can be determined. Our analysis of the relative frequencies of terminal motifs shows that the relative frequencies vary across different tissues. As explained below, the quantification of relative frequency differences can be used in conjunction with calibration samples, such as those containing known fractional concentrations of clinically relevant DNA (e.g., measured by separate techniques, such as tissue-specific paired genes), to determine the classification of clinically relevant DNA fractional concentrations in a biological sample.

[0099] While it may be necessary to calibrate the measurement of fractional concentrations of clinically relevant DNA in samples, the resulting calibration values ​​(e.g., as part of a calibration function) can be used to determine fractional concentrations in new samples without identifying pairwise genes specific to clinically relevant DNA. This provides a more robust way to determine fractional concentrations. A. Pregnancy

[0100] Genotypic differences between the maternal and fetal genomes can be used to distinguish fetal and maternal DNA molecules. For example, we can use informative single nucleotide polymorphism (SNP) sites where the mother is homozygous (AA) and the fetus is heterozygous (AB).

[0101] Figure 2 illustrates a schematic diagram of genotypic differences based on a method for analyzing differential terminal motif patterns between fetal and maternal DNA molecules according to an embodiment of this disclosure. As illustrated in Figure 2, fetal-specific molecules 205 carrying fetal-specific paired genes (B) are identifiable. On the other hand, shared molecules 207 carrying common paired genes (A) are identifiable, which will represent DNA molecules primarily derived from the mother, since fetal DNA molecules are typically a minority in the maternal plasma DNA pool. Therefore, the molecular properties of any molecule derived from the shared molecule will reflect the characteristics of maternal background DNA molecules (i.e., DNA molecules derived from hematopoietic tissue). In addition to paired genes, other fetal-specific markers (e.g., epigenetic markers) may also be used.

[0102] We used technique 140 in Figure 1 to analyze the tetramer terminal motifs. 256 terminal motifs were analyzed. We calculated the proportion of each tetramer motif and compared the frequencies of the 256 motifs using a bar chart depicted as bar graph 220. Such bar charts provide the relative frequency (%) of each tetramer appearing as a terminal motif. For ease of illustration, only a few tetramers are shown. The relative frequency (sometimes called "frequency") can be determined by (number of DNA fragments with terminal motifs) / total number of DNA fragments analyzed (possibly with a factor of 2 in the denominator) to interpret the two ends. Such percentages can be considered relative frequencies because they are ratios of one quantity (e.g., number) of the first terminal motif to the quantity of one or more other motifs (which may contain the first terminal motif). As we can see, terminal motif 222 shows significant differences in relative frequency between DNA fragments in different tissue types. Such differences can be used for various purposes, such as enriching samples of fetal DNA or determining fetal DNA concentration.

[0103] The relative frequency values ​​displayed in bar chart 220 can be stored values ​​in an array of 256 values. Each end motif in a set of end motifs can have a counter, wherein the counter of a particular end motif increments each time a new DNA fragment has an end motif corresponding to that counter. The set of motifs can be chosen in various ways, such as as all end motifs or a smaller set, such as those occurring most frequently in the reference sample or those exhibiting the largest intervals in the reference sample.

[0104] Various quantitative techniques can be used to provide measurements of the relative frequencies of samples, and such techniques can be used to classify the amount of cell-free DNA from clinically relevant DNA. One example quantitative technique comprises the sum of the relative frequencies of a set of terminal motifs, also referred to herein as combined frequencies. For example, such a set could be the terminal motifs most frequently occurring in a particular tissue type or identified as having the largest interval between two tissue types. Weighted sums can also be used. The weights can be predetermined or variable; for example, the weights used for a given frequency can depend on the frequency itself. Entropy is such an example.

[0105] In another embodiment, to capture the lateral differences in terminal motifs between fetal and maternal DNA molecules, entropy-based analysis 230 can be used. Entropy is an instance of variance / diversity. To analyze the frequency distribution of motifs (e.g., a total of 256 motifs), one definition of entropy uses the following equation: in The frequency of a specific primitive; a higher entropy value indicates higher diversity (i.e., higher randomness).

[0106] In this example, the entropy reaches its maximum value (5.55) when 256 motifs are present in terms of their frequencies. In contrast, the entropy decreases when the 256 motifs have a skewed distribution in their frequencies. For example, if one particular motif constitutes 99% and other motifs make up the remaining 1%, the entropy would decrease to 0.11 in this configuration, but other configurations can be used, such as no logarithm or only logarithm. Therefore, the entropy of a decrease in motif frequency implies an increased skewness in the frequency distribution of terminal motifs. The entropy of an increase in motif frequency indicates that the frequencies in the motifs will shift toward equal probability towards those motifs. Therefore, the entropy of motif frequency measures the uniformity of terminal motif abundance in plasma DNA. The higher the uniformity of motif frequencies, the higher the entropy value will be expected. In other words, the entropy of a decrease in motif frequency implies an increased skewness in its frequency distribution among terminal motifs.

[0107] In various other instances, the standard deviation (SD), coefficient of variation (CV), interquartile range (IQR), or a percentile cutoff (e.g., the 95th or 99th percentile) across different motif frequencies can be used to assess the overall variation in terminal motif patterns between fetal and maternal DNA molecules. These various instances provide a measure of the variance / diversity of the relative frequencies of a set of terminal motifs. Given the definition of entropy in Figure 2, entropy will have a minimum if only one terminal motif has a non-zero count. If other terminal motifs are present in some DNA fragments, entropy will increase. If there is no selection (a random distribution of all terminal motifs, e.g., in a hypothetical case where all have the same frequency), entropy will become a maximum. In this way, entropy quantifies the overall selectivity of the terminal sequence of a free DNA fragment for terminal motifs.

[0108] Figure 235 illustrates the entropy values ​​of the common sequence (primarily maternal) and the fetal sequence. The common sequence comprises fetal DNA at a lower concentration than the fetal sequence (approximately 5% if the original sample has 10% fetal DNA), and these fetal sequences will have nearly 100% fetal DNA, within the error tolerance for genotype measurement. Given this separation, the greater the concentration of fetal DNA in the sample, the greater the difference in entropy values. This relationship between fetal DNA concentration and entropy can be used to determine fetal DNA concentration, for example, as measured using one or more calibration values. For instance, the clinically relevant DNA concentration of the calibration sample can be measured via another technique (generating calibration values), which may be generally unsuitable, such as using Y-chromosome DNA from male fetuses or previously identified mutations in tumor tissue. When an entropy measurement is given for the calibration sample, comparing two entropy values ​​(one for the test sample and one for the calibration sample) can provide a fractional concentration to the test sample using the measured concentration in the calibration sample. Further details of such uses of calibration values ​​and calibration functions will be described later.

[0109] In another embodiment, cluster-based analysis 240 can be used. The vertical axis corresponds to tetramer motifs, and the horizontal axis corresponds to different samples, such as different classifications based on fetal DNA concentration. Colors correspond to the relative frequency of a specific tetramer motif for a particular sample, such as red calibration sample 242 having a higher concentration than green calibration sample 244 (which has a lower value).

[0110] Cluster-based analysis utilizes the assumption that the similarity of the frequency profiles of the 256 tetramer terminal motifs will be relatively high within either fetal or maternal DNA molecules (i.e., intra-group molecular properties) compared to the similarity between fetal and maternal DNA molecules. Therefore, a calibration sample of an individual characterized by terminal motifs derived from a shared sequence (e.g., a higher concentration of the shared sequence) is expected to differ from a calibration sample of an individual characterized by terminal motifs derived from a fetal-specific sequence (e.g., a lower concentration of the shared sequence, and therefore a higher fetal concentration). Each individual corresponds to a vector comprising the 256 terminal motifs and their corresponding frequencies (i.e., a 256-dimensional vector). Instance clustering techniques include (but are not limited to) hierarchical clustering, center-based clustering, distribution-based clustering, and density-based clustering. Due to the frequency differences of terminal motifs between maternal and fetal DNA fragments, different clusters may correspond to different amounts of fetal DNA in samples, as they will have different relative frequencies of patterns.

[0111] To assess the differences in terminal motifs between fetal and maternal DNA molecules, we used a microarray platform (Human Omni2.5, Illumina) to genotype maternal leukocyte and fetal samples, and sequenced matched plasma DNA samples. Peripheral blood samples were obtained from 10 pregnant women from each of the early (12–14) week, mid (20–23) week, and late (38–40) week gestation periods, and plasma and maternal leukocyte samples were collected from each case. We obtained 195,331 informative SNPs (range: 146,428–202,800), of which the mothers were homozygous and the fetuses were heterozygous. Plasma DNA molecules carrying fetal-specific paired genes were identified as fetal-specific DNA molecules. Plasma DNA molecules carrying shared paired genes were identified and believed to be primarily derived from maternal DNA. The median fetal DNA fraction in these samples was 17.1% (range: 7.0%–46.8%). A median of 103 million (range: 52–186 million) mapped bilaterally sequenced reads were obtained across all cases. The terminal motifs of each plasma DNA molecule were determined by bioinformatics analysis of the tetramer sequence closest to the fragment ends. Results from the analysis of this sample set are provided below. [, 1. , ] [, Differences in relative frequencies arranged in order , ] [, , ]

[0112] We hypothesize that the apical motifs in the frequency grading of motifs between fetal and maternal DNA molecules will be useful for detecting or enriching both types of DNA. Therefore, we graded the apical motifs based on the frequency differences between fetal and maternal DNA molecules in a pregnant woman, with a sequencing depth of 270×. Using a method similar to that mentioned above, we identified fetal sequences and shared sequences based on informative SNPs.

[0113] Figure 3 shows a bar graph of the frequencies of terminal motifs between fetal and maternal DNA molecules according to an embodiment of this disclosure. Data was obtained from a pregnant woman with a sequencing depth of 270×. The vertical axis corresponds to the percentage frequency of a given tetramer motif, determined by dividing the number of DNA fragments having the given tetramer motif (as determined from the sequence reads) by the total number of terminal sequences of the analyzed DNA fragments (e.g., twice the number of DNA fragments). The horizontal axis corresponds to 256 different tetramers. For shared sequences, tetramers are categorized in decreasing frequency, with Figure 3 separated into two parts with different scale bars for the vertical axis. Differences in the frequencies of terminal motifs can be observed between fetal DNA molecules (fetal DNA molecules with fetal-specific paired genes) and maternal DNA molecules (maternal DNA molecules with shared paired genes).

[0114] Figure 4 shows the first 10 terminal motifs of the fetal and common (i.e., fetal plus maternal) sequences from the embodiment of this disclosure. The vertical axis is shifted and starts at a frequency of 1%. The first 10 terminal motifs are CCCA, CCAG, CCTG, CCAA, CCCT, CCTT, CCAT, CAAA, CCTC, and CCAC. As can be seen, some terminal motifs show greater differences between the common and fetal-specific sequences compared to other sequences. Therefore, to distinguish between maternal and fetal DNA, we might want to use the terminal motifs that show the greatest difference relative to only the terminal motifs with the highest frequency. [, 2. , ] [, Use of entropy , ] [, , ]

[0115] For each sample, the entropy of DNA molecules sharing a common paired gene and the entropy of DNA molecules sharing a fetal-specific paired gene were then analyzed. The former identified the mother, and the latter identified the fetus. For each sample, two data points were obtained: the entropy of fetal DNA molecules and the entropy of shared DNA molecules (labeled "mother").

[0116] Figure 5A shows that the entropy of terminal motifs in fetal DNA molecules is lower than that in maternal DNA molecules (p < 0.0001), indicating a higher skewness in the distribution of terminal motifs derived from maternal DNA molecules. For a given sample and a given library of fetal or maternal DNA molecules, the entropy in Figure 5A is determined using all 256 motifs, such as tetramers used in these examples.

[0117] Similar to curve 235 in Figure 2, the difference in entropy between the two tissue types demonstrates how entropy can be used to determine the fractional concentration of fetal DNA in a mixture of cell-free DNA fragments (e.g., plasma or serum). As explained above, the library identified as fetal DNA has a higher percentage (e.g., close to 100%) of fetal DNA than the maternal library. The entropy values ​​measured differ for different types of libraries. Therefore, a relationship exists between entropy and fetal DNA concentration. This relationship can be determined based on a calibration function using measurements of fetal DNA concentration in calibration samples (calibration values) and corresponding entropy values ​​(examples of relative frequencies), where calibration values ​​and relative frequencies form calibration data points. Calibration samples with different fetal DNA concentrations will have different entropy values. The calibration function can be fitted to the calibration data points such that recently measured relative frequencies (e.g., entropy) can be input to the calibration function to provide the output of fetal DNA concentration.

[0118] Figure 5B illustrates entropy when using the relative frequencies of the 10 motifs from Figure 4. As shown, for a given set of 10 terminal motifs, this relationship changes with fetal sequences having higher entropy. The fractional concentration of fetal DNA can still be determined, but a different calibration function will be used. Therefore, the set of motifs used for calibration should be the same as those subsequently used, i.e., when measuring fractional concentration based on entropy or other total values ​​of the relative frequencies of the set. 3. Clustering

[0119] We further conducted a hierarchical cluster analysis on pregnant women, each represented by a 256-dimensional vector comprising the frequencies of all tetramer terminal motifs. In effect, individuals characterized by terminal motifs derived from fetal-specific sequences and maternal DNA molecules can be divided into two groups.

[0120] Figures 6A and 6B show hierarchical clustering analysis of fetal and maternal DNA molecules in early pregnancy according to embodiments of this disclosure. Figure 6A illustrates hierarchical clustering analysis based on the frequencies of 256 tetramer terminal motifs. The vertical axis corresponds to tetramer motifs, and the horizontal axis corresponds to different portions of various samples (i.e., fetal-specific 620 (yellow) and common 610 (blue) sequences). The colors correspond to the relative frequencies of specific tetramer motifs in specific portions of a sample.

[0121] Different segments (fetal-specific and common) have different fetal DNA concentrations, and therefore will have different classifications based on fetal DNA concentration. When performing this type of clustering using calibration samples, fetal DNA concentrations can be measured, for example, as described in the entropy section above. Each calibration sample will have a corresponding vector of length equal to the number of primitives used (e.g., 256 for all tetramers or possibly only tetramers, since there is the greatest difference between fetal and common sequences, but other k-mers can be used).

[0122] Figure 6B shows a magnified visualization of hierarchical clustering analysis based on the frequencies of 256 tetramer terminal motifs. Each column represents a type of terminal motif (i.e., different terminal motifs). Each row represents pregnant individuals. The gradient colors indicate the frequency of the terminal motifs. Red represents the highest frequency and green represents the lowest frequency. As we can see, the two parts (fetal and common) representing samples with different fetal DNA concentrations cluster completely into two independent clusters, demonstrating good accuracy in distinguishing samples with different fetal DNA concentration levels. [, 4. , ] [, Samples at different stages of pregnancy , ] [, , ]

[0123] In addition to being able to distinguish between samples with different fractional concentrations, some embodiments can distinguish between samples from pregnant individuals at different gestational ages (e.g., during or just into their late pregnancy).

[0124] Figures 7A and 7B illustrate the entropy distribution of all primitives used by pregnant women at different stages of pregnancy according to an embodiment of this disclosure. Interestingly, the entropy value of the number of terminal primitives determined using fetal-specific data appears to be correlated with gestational age (p-value: 0.024, early pregnancy data relative to combined mid and late pregnancy data), but those from common segments (mainly maternal DNA) do not appear to be correlated with gestational age (p-value: 1, early pregnancy data relative to combined mid and late pregnancy data). Late pregnancy typically has a higher concentration of fetal DNA. Therefore, a correlation may exist between concentration and gestational age.

[0125] For fetal-specific segments, there is reduced entropy in mid- and late-pregnancy compared to early pregnancy. Therefore, fetal segments can convey gestational age. Furthermore, since shared segments essentially have constant entropy (e.g., because terminal motif changes primarily associated with maternal segments and / or maternal physiology cancel out such fetal signals), changes in entropy across all segments will reflect gestational age as fetal segments change. Due to the presence of maternal segments, this relationship of entropy across different gestational periods will show less variation, but the relationship will still exist. However, when fetal specificity is identified (e.g., a male fetus or by identifying a pair of genes present at a percentage similar to the expected fetal DNA concentration or using paternal genotype information), a more pronounced relationship will subsequently emerge (e.g., as shown in Figure 7B).

[0126] Figures 7C and 7D illustrate the entropy distribution of 10 primitives for pregnant women at different stages of pregnancy according to an embodiment of this disclosure. The 10 primitives were selected by ranking from common segments. These figures demonstrate that, for fetal-specific segments, the entropy still changes at different stages of pregnancy, even if the relationship can be decreasing (opposite to the increase in Figure 7B), due to the specific selection of the primitives.

[0127] Figure 8A shows the entropy of all fragments at different gestational ages according to the embodiments of this disclosure. Entropy was determined using all 256 tetramer terminal motifs. It was shown that the entropy of plasma DNA fragments in individuals with late pregnancy was lower than that in individuals with early and mid-pregnancy (p=0.06). Furthermore, the average entropy was lower in mid-pregnancy than in early pregnancy. Therefore, when all fetal fragments are included (as opposed to the shared fragments in Figure 7A), entropy does indeed provide gestational age.

[0128] Figure 8B shows the entropy of Y-chromosome-derived segments at different gestational ages. It is shown that the entropy of Y-chromosome-derived segments in individuals with late pregnancy is lower than that in individuals with early or mid-pregnancy (p=0.01). These samples, filtered for fetal molecules (using fetal-specific sequences from the Y chromosome), exhibit a larger interval between late and mid-pregnancy.

[0129] Figures 9 and 10 illustrate the distribution of the first 10 terminal motifs of the fetal and maternal DNA molecules at different stages of pregnancy, according to an embodiment of this disclosure. The first 10 terminal motifs in the differences in motif frequency arrangements between fetal and maternal DNA molecules were mined from a single deep-seated pregnancy case. These first 10 terminal motifs were subsequently used to analyze each of the samples.

[0130] The proportions of fetal and shared DNA molecules bearing these relevant terminal motifs were calculated in an independent group comprising 10 pregnant women from each of the first (12–14 weeks), second (20–23 weeks), and third (38–40 weeks) pregnancies. Fetal DNA molecules showed a higher proportion of many terminal motifs compared to shared molecules, indicating a relationship between these motifs and the tissue of origin. For example, the median percentage of CAAA in fetal DNA molecules was consistently higher than that in shared molecules (primarily maternal) in the first (1.26% vs. 1.11%), second (1.24% vs. 1.11%), and third (1.24% vs. 1.15%) pregnancies. Therefore, the terminal motif CAAA can be identified as a marker indicating an increased likelihood that a specific DNA fragment with CAAA-terminal sequences is from the fetus.

[0131] Some terminal motifs show a more pronounced relationship with gestational age. For example, fetal DNA molecules with the terminal motif CCCA show a continuous (monotonic) increase in gestational age, as do CCAG, CCTG, CCAA, CCCT, and CCAC. However, CCTT values ​​do not show a continuous increase with gestational age, as they decrease in mid-pregnancy and then increase in late pregnancy.

[0132] In another embodiment, we can combine the terminal motifs of the first 10 permutations to examine the differences between fetal and maternal DNA molecules at different stages of pregnancy.

[0133] Figure 11 illustrates the combination frequencies of the first 10 arrangements of terminal motifs between fetal and maternal DNA molecules at different stages of pregnancy, according to an embodiment of this disclosure. As shown in Figure 11, we found that the difference in the combination frequencies of the first 10 arrangements of terminal motifs between fetal and maternal DNA molecules was relatively greater in the second trimester (p: 0.013) and third trimester (p: 0.0019) compared to the first trimester (p: 0.92). The frequency of fetal molecules consistently increased from the first trimester to the second trimester to the third trimester, while no such continuity was observed for shared molecules. This demonstrates that different physiological conditions (e.g., gestational age) will affect terminal motifs derived from different tissue sources. B. Oncology

[0134] Genotypic components designed for pregnancy can also be used in oncology settings.

[0135] Figure 12 illustrates a schematic diagram of a genotype-based method according to an embodiment of this disclosure, used to analyze differentially expressed terminal motif patterns between mutations and shared molecules in the plasma DNA of cancer patients. As illustrated in Figure 12, a tumor-specific molecule 1205 carrying a tumor-specific paired gene (B) can be identified. On the other hand, a shared molecule 1207 carrying a shared paired gene (A) can be identified, which will represent DNA molecules primarily derived from healthy individuals, since tumor DNA molecules are typically few in number in the plasma DNA pool.

[0136] As an example, we can distinguish between mutated sequences (i.e., plasma DNA carrying cancer-related mutations) and common sequences (DNA primarily derived from hematopoietic tissue). Cancer-related mutations can be defined as mutations present in tumor tissue (hepatocellular carcinoma, HCC) but not in normal cells (e.g., leukocytes). For example, in an HCC patient, assuming the tumor tissue genotype is "AG" at a specific locus and the leukocyte cells are "AA", then the "G" specifically present in the tumor tissue would be considered a cancer-related mutation, and "A" would be considered a common wild-type pair. In various embodiments, mutated sequences can be obtained by sequencing tissue sections from the tumor or by analyzing free samples (such as plasma or serum), as described, for example, in U.S. Patent Publication 2014 / 0100121.

[0137] This study determined the frequency profile of terminal motifs between mutated and common sequences in HCC patients, using plasma DNA sequenced at a depth of 220×. Bar chart 1220 provides the relative frequency (%) of each tetramer appearing as a terminal motif for mutated and common sequences. Such relative frequencies can be determined with respect to bar chart 220 in Figure 2 as described above. As can be seen, terminal motif 1222 shows significant differences in relative frequency between DNA fragments in different tissue types. Such differences can be used for various purposes, such as enriching samples of tumor DNA or determining tumor DNA concentration.

[0138] In another embodiment, to capture the overall difference in terminal motifs between the tumor and the shared DNA molecule, an entropy-based analysis 1230, similar to that in Figure 2, can be used. Curve 1235 shows the entropy values ​​of the shared sequence and the tumor sequence. Differences in entropy or other variance measures can, for example, use a calibration function to provide tumor fraction concentration.

[0139] In another embodiment, cluster-based analysis 1240 can be performed, similar to the fetal analysis in Figure 2. The classification of the amount of tumor sequences in the sample can be determined based on a reference cluster to which the new sample belongs of a known tumor score classification. [, 1. , ] [, Differences in relative frequencies arranged in order , ] [, , ]

[0140] Figure 13 shows an overview of plasma DNA terminal motifs of cancer-associated mutations and shared molecules in hepatocellular carcinoma according to embodiments of this disclosure. Multiple terminal motifs with alterations observed between mutated and shared sequences are present, such as, but not limited to, CCCA, CCAG, CCAA, CCTG, CCTT, CCCT, CAAA, CCAT, TAAA, and AAAA motifs. Figure 13 shows information similar to Figure 3, but for clinically relevant DNA, it is tumor DNA, the opposite of fetal DNA.

[0141] Figure 14 shows a radial plot of plasma DNA terminal motifs of cancer-related mutations and shared molecules in hepatocellular carcinoma according to embodiments of this disclosure. Different terminal motifs are listed peripherally, and their frequencies are expressed as different radial lengths. Terminal motifs are categorized by the frequency of wild-type (wt) pairs in non-tumor (e.g., healthy) cells. Frequency value 1410 corresponds to the wt pair, and frequency value 1420 corresponds to the mutated (mut) pair. This radial plot demonstrates a significant difference in the relative frequency of terminal motifs in mutated sequences compared to wild-type (shared) sequences.

[0142] Figure 15A shows the top 10 end motifs showing the most significant differences in frequency between mutations and common sequences in the plasma DNA of HCC patients according to embodiments of this disclosure. The end motifs of the common sequences in the reference sample were identified. As shown, the end motifs are CCCA, CCAG, CCAA, CCTG, CCTT, CCCT, CAAA, CCAT, TAAA, and AAAA. The relative frequency differences varied among the end motifs. For example, the motifs showing the majority of the differences between mutations and common sequences (CCCA) were found to be 1.9% and 1.6%, respectively, indicating that the mutated sequences of these motifs were 15% lower than the common sequences (primarily hematopoietic wild-type sequences).

[0143] Figure 15B illustrates the combination frequencies of eight terminal motifs in HCC patients and pregnant women according to embodiments of this disclosure. The combination frequencies are total values ​​for the examples, for example, the sum of the relative frequencies of a set of terminal motifs. As can be seen, in each of these two cases, there is an interval in the combination frequencies between the two classes of sequences: between wild-type (WT) and mutant sequences, and between maternal and fetal sequences. The interval in the combination frequencies between wild-type (WT) and mutant sequences is greater than the interval between maternal and fetal sequences.

[0144] This combination of frequencies exhibits behavior similar to that of the entropy plot in fetal analysis. Therefore, Figure 15B shows another example of the total relative frequency that can be used to determine the fractional concentration of clinically relevant DNA. Furthermore, the WT relative to mutation in Figure 15B shows that the fractional concentration of other clinically relevant DNA (e.g., tumor DNA) can also be determined. [, 2. , ] [, The use of entropy , ] [, , ]

[0145] Figures 16A and 16B illustrate the entropy values ​​of shared and mutant fragments for different sets of terminal motifs in HCC cases according to embodiments of this disclosure. Similar to fetal sequences, the relationship between the entropies of the two types of sequences can vary depending on the set of terminal motifs used. Figure 16A uses all 256 terminal motifs of the tetramer. The entropy of the mutant fragments is higher due to their more uniform (e.g., flatter) frequency distribution. And the entropy of the shared fragments is lower due to the higher skewness of the frequency distribution.

[0146] Figure 16B uses the first 10 terminal motifs of the tetramers appearing in HCC individuals for shared fragments. The entropy relationship is the opposite of that of the first ten motifs. Figures 16A and 16B both show that the calibration analysis used to determine fetal DNA concentration can also be used to determine tumor DNA concentration.

[0147] As explained above, higher entropy values ​​indicate greater diversity in terminal motifs. Motif diversity scores (MDS) can be used to estimate the fractional concentration of clinically relevant DNA (e.g., fetus, transplant, or tumor) in biological samples containing circulating cell-free DNA.

[0148] Figure 17 is a graph of the primitive diversity score (entropy) relative to the measured circulating tumor DNA fraction according to an embodiment of this disclosure. For each of a plurality of calibration samples, calibration data point 1705 is measured. The calibration data point includes the primitive diversity score of the sample and the fractional concentration of clinically relevant DNA (in this case, tumor DNA fraction). The tumor DNA fraction is estimated based on ichorCNA, a suite of software that measures the tumor DNA fraction in plasma DNA by utilizing cancer-related duplicate number bias (Adalsteinsson et al. 2017).

[0149] The given samples can be healthy control samples without tumor DNA or samples from patients with tumors, where the tumor DNA score is non-zero, meaning that tumor DNA and other (e.g., healthy) DNA are present. A positive correlation was found between the MDS value of plasma DNA and the tumor DNA score in patients with HCC (Spearman's correlation coefficient (ρ): 0.597; p-value: 0.0002). This is presented using the calibration function 1710 (a linear function in this example).

[0150] The calibration function 1710 can be used to determine the tumor DNA score in a new test sample with a measured primitive diversity score. The calibration function 1710 can be determined, for example, by fitting a function to the calibration data points 1705 using regression.

[0151] In some instances, the calculated MDS value X for a new sample can be used as input to a function F(X), where F is a calibration function (curve). The output of F(X) is the fractional concentration. An error range can be provided, and the error range for each X value can be different, thereby providing a series of values ​​as the output of F(X). In other instances, the fractional concentration in a new sample corresponding to a measurement of 0.95 of the MDS can be determined as the average concentration calculated from the calibration data point at 0.95 of the MDS. As another example, calibration data point 1705 can be used to provide a series of fractional DNA concentrations for a specific calibration value, where this range can be used to determine whether the fractional concentration is above the threshold limit. C. Transplantation

[0152] Genotyping techniques can also be used to monitor transplantation, such as liver transplantation. SNP sites where the recipient is homozygous and the donor is xenozygous will allow for the identification of donor-specific DNA molecules in the transplant patient's plasma and DNA from major hematopoietic tissues.

[0153] Figure 18A illustrates entropy analysis using donor-specific fragments in an embodiment of this disclosure. Figure 18B illustrates hierarchical clustering analysis using donor-specific fragments. As shown in Figures 18A and 18B, in the case of liver transplantation, liver-specific DNA molecules were observed to have properties different from common sequences (primarily blood-derived DNA). Compared to common sequences, the entropy of plasma DNA terminal motifs was generally found to be lower in donor-specific DNA molecules (liver DNA) (Figure 18A). Individuals characterized by terminal motifs derived from liver-specific DNA molecules clustered together, while individuals characterized by terminal motifs of common DNA molecules clustered together. D. Classification score concentration

[0154] As described above, the relative frequencies of a set of single or multiple terminal motifs can be used to classify the fractional concentrations of clinically relevant DNA.

[0155] Figure 19 is a flowchart illustrating a method 1900 for estimating the fractional concentration of clinically relevant DNA in an individual's biological sample according to an embodiment of this disclosure. The biological sample may contain free clinically relevant DNA and other DNA. In other instances, the biological sample may not contain clinically relevant DNA, and the estimated fractional concentration may indicate zero or a lower percentage of clinically relevant DNA. Method 1900 and any other methods described herein can be performed by a computer system.

[0156] At step 1910, multiple cell-free DNA fragments from the biological sample are analyzed to obtain sequence reads. Sequence reads may contain terminal sequences corresponding to the ends of the multiple cell-free DNA fragments. As an example, sequence reads can be obtained using sequencing or probe-based techniques, either of which involves enrichment, for example, via amplification or capture probes.

[0157] Sequencing can be performed in various ways, such as using massively parallel sequencing or next-generation sequencing, single-molecule sequencing, and / or double-stranded or single-stranded DNA sequencing library preparation protocols. Those skilled in the art should understand that a variety of sequencing techniques are available. As part of the sequencing process, some sequence reads may correspond to cellular nucleic acids.

[0158] Sequencing can be targeted sequencing as described herein. For example, a biological sample can be enriched with DNA fragments from a specific region. Enrichment can involve using capture probes that bind to, for example, a portion or the entire genome as defined by a reference genome.

[0159] A significant number of cell-free DNA molecules can be analyzed to provide an accurate determination of fractional concentration. In some embodiments, at least 1,000 cell-free DNA molecules are analyzed. In other embodiments, at least 10,000, 50,000, 100,000, 500,000, 1,000,000, or 5,000,000 or more cell-free DNA molecules can be analyzed.

[0160] At step 1920, for each of the plurality of cell-free DNA fragments, a sequence motif is determined for each of one or more terminal sequences of the cell-free DNA fragment. The sequence motif may contain N base positions (e.g., 1, 2, 3, 4, 5, 6, etc.). As an example, the sequence motif may be determined by analyzing sequence reads corresponding to the ends of the DNA fragments, associating signals with specific motifs (e.g., when using probes), and / or aligning sequence reads with a reference genome, as illustrated in Figure 1.

[0161] For example, after sequencing by a sequencing device, the sequence reads can be received by a computer system communicatively coupled to the sequencing device, such as via wired or wireless communication or via a removable memory device. In some embodiments, one or more sequence reads comprising the two ends of a nucleic acid fragment can be received. The location of the DNA molecule can be determined by aligning one or more sequence reads of the DNA molecule with different portions of the human genome (e.g., specific regions). In other embodiments, specific probes (e.g., after PCR or other amplification) can indicate the location or specific end motifs, such as by a specific fluorescent color. Identification can be made by recognizing a free DNA molecule as one of a set of sequence motifs.

[0162] In step 1930, the relative frequencies of a set of one or more sequence motifs corresponding to the terminal sequences of a plurality of cell-free DNA fragments are determined. The relative frequencies of the sequence motifs provide the proportion of a plurality of cell-free DNA fragments having terminal sequences corresponding to the sequence motifs. A reference set of one or more reference samples can be used to identify the set of one or more sequence motifs. Although genotypic differences can be determined to identify differences between the terminal motifs of clinically relevant DNA and other DNA (e.g., healthy DNA of how an individual received a transplanted organ, maternal DNA, or other DNA), it is not necessary to know the fractional concentration of clinically relevant DNA for the reference samples. Specific terminal motifs can be selected based on differences (e.g., selecting the terminal motif with the highest absolute or percentage difference). Examples of relative frequencies are described throughout the disclosure.

[0163] In some implementations, a sequence motif comprises N base positions, wherein a set of one or more sequence motifs contains all combinations of N bases. In some instances, N may be an integer equal to or greater than two or three. The set of one or more sequence motifs may be the top M (e.g., 10) most common sequence motifs generated in one or more calibration samples or other reference samples not used for calibration fraction concentration.

[0164] In step 1940, the total relative frequency of a set of one or more sequence primitives is determined. The total value is described throughout the disclosure and may include, for example, an entropy value (primitive diversity score), the sum of relative frequencies, and multidimensional data points corresponding to a count vector of a set of primitives (e.g., vector 256 for counting 245 primitives of a possible 4-merger or 64 for counting 64 primitives of a possible 3-merger). When the set of one or more sequence primitives contains a plurality of sequence primitives, the total value may include the sum of the relative frequencies of that set.

[0165] As an example, when a set of one or more sequence primitives contains a complex number of sequence primitives, the total value may contain the sum of the relative frequencies of that set. As another example, the total value may correspond to the variance of the relative frequencies. For instance, the total value may contain an entropy term. The entropy term may contain the sum of terms, each containing the relative frequency multiplied by the logarithm of the relative frequencies. As yet another example, the total value may contain the final or intermediate output of a machine learning model (e.g., a clustering model).

[0166] In step 1950, the classification of the fractional concentration of clinically relevant DNA in the biological sample is determined by comparing the total value with one or more calibration values. One or more calibration values ​​can be determined from one or more calibration samples whose fractional concentration of clinically relevant DNA is known (e.g., measured). The comparison can be made with a plurality of calibration values. The comparison can be performed by fitting the total value as a calibration function to calibration data, which provides a change in the total value relative to a change in the fractional concentration of clinically relevant DNA in the sample. As another example, one or more calibration values ​​correspond to one or more total values ​​representing the relative frequencies of one or more sets of sequence motifs measured using cell-free DNA fragments in one or more calibration samples.

[0167] A calibration value can be calculated as the total value for each calibration sample. Calibration data points for each sample can be determined, where each calibration data point includes the calibration value of the sample and the measured fractional concentration. These calibration data points can be used in method 1900 or to determine the final calibration data points (e.g., as defined by functional fitting). For example, a linear function can be fitted to the calibration value that varies with fractional concentration. The linear function can define the calibration data points to be used in method 1900. As part of the comparison, the new total value of a new sample can be used as input to the function to provide the output fractional concentration. Therefore, one or more calibration values ​​can be multiple calibration values ​​of a calibration function determined using the fractional concentrations of clinically relevant DNA from multiple calibration samples.

[0168] As another example, the new total value can be compared with the average total value of samples in the same category (e.g., within the same range) that have fractional concentrations, and if the new total value is closer to a calibration value compared to the average of another category, then the new sample can be determined to have the same concentration. Such techniques can be used when performing clustering. For example, the calibration value could be a representative value of a cluster corresponding to a specific category of fractional concentrations.

[0169] Calibration data points may include, for example, the measurement of fractional concentrations. For each of one or more calibration samples, the fractional concentration of clinically relevant DNA can be measured in the calibration sample. One or more total values ​​can be determined by analyzing cell-free DNA fragments from the calibration sample as part of obtaining calibration data points, thereby determining the total relative frequencies of one or more sequence motifs. Each calibration data point can specify the measured fractional concentration of clinically relevant DNA in the calibration sample and the total value determined for the calibration sample. One or more calibration values ​​can be one or more total values, or can be determined using one or more total values ​​(e.g., when using a calibration function). The measurement of fractional concentrations can be performed in various ways as described herein, for example by using a pair of genes specific to clinically relevant DNA.

[0170] In various embodiments, tissue-specific paired genes or epigenetic markers may be used, or the size of the DNA fragment may be used to measure the fractional concentration of clinically relevant DNA, as described, for example, in U.S. Patent Publication 2013 / 0237431, which is incorporated herein by reference in its entirety. Tissue-specific epigenetic markers may be included in a sample as DNA sequences that exhibit tissue-specific DNA methylation patterns.

[0171] In various embodiments, clinically relevant DNA can be selected from the following groups: fetal DNA, tumor DNA, DNA from a transplanted organ, and a specific tissue type (e.g., from a specific organ). Clinically relevant DNA can be a specific tissue type, such as liver or hematopoietic tissue. When the individual is a pregnant woman, clinically relevant DNA can be placental tissue, which corresponds to fetal DNA. As another example, clinically relevant DNA can be tumor DNA derived from an organ with cancer.

[0172] Typically, the preferred method is to use an analysis similar to that used for the biological (test) sample to measure fractional concentrations to generate one or more calibration values ​​determined by one or more calibration samples. For example, sequenced libraries can be generated in the same manner. Two sample processing techniques are GeneRead (www.qiagen.com / us / shop / sequencing / generead-size-selection-kit / #orderinginformation) and SPRI (Solid-phase reversible immobilization, AMPure beads, www.beckman.hk / reagents_depr / genomic_depr / cleanup-and-size-selection / pcr—). GeneRead removes shorter DNA fragments, primarily tumor fragments, which can affect the relative frequencies of wild-type and mutant fragments, as well as terminal motifs in fetal and transplant cases. E. Determine gestational age

[0173] As described in Figures 7A, 7B and 8-10 above, fetal-specific fragment motifs can be used to estimate gestational age.

[0174] Figure 20 is a flowchart of an embodiment according to this disclosure, illustrating a method 2000 for determining gestational age of a fetus by analyzing biological samples from a pregnant woman. The biological samples contain cell-free DNA molecules from the woman and the fetus.

[0175] In step 2010, multiple cell-free DNA fragments from a biological sample are analyzed to obtain sequence reads. The sequence reads may contain terminal sequences corresponding to the ends of the multiple cell-free DNA fragments. Step 2010 can be performed in a similar manner to step 1910 in Figure 19.

[0176] Before, after, or as part of the analysis, multiple cell-free DNA fragments can be identified as originating from the fetus, for example, as described above with reference to Figures 2 and 5A. This can filter DNA fragments that are fetal or most likely to be fetal monosomy. As an example, fetal-specific paired genes or fetal-specific epigenetic markers can be used to identify multiple cell-free DNA fragments. As another example, for each of the sequence reads, the probability that the sequence read corresponds to the fetus can be determined based on the terminal sequence of the sequence read, which contains sequence motifs from a set of one or more sequence motifs. Other criteria can also be used, such as those described in Section II. E. The probability can be compared to a threshold value, and when the probability exceeds the threshold value, the sequence read can be identified as originating from the fetus. Further details regarding samples enriched with clinically relevant DNA can be found in Section IV.

[0177] In step 2020, for each of the plurality of cell-free DNA fragments, the sequence motif of each of one or more terminal sequences of the cell-free DNA fragment is determined. Step 2020 can be performed in a similar manner to step 2020 in Figure 19.

[0178] In step 2030, the relative frequencies of a set of single or multiple sequence motifs corresponding to the terminal sequences of a plurality of free DNA fragments are determined. The relative frequencies of the sequence motifs provide the proportion of a plurality of free DNA fragments having terminal sequences corresponding to the sequence motifs. Step 2030 can be performed in a similar manner to step 1930 of Figure 19.

[0179] In step 2040, the total relative frequency of the set of one or more sequence primitives is determined. Step 2040 can be performed in a similar manner to step 1940 in Figure 19.

[0180] In step 2050, one or more calibration data points are obtained. Each calibration data point may specify a gestational age (e.g., the gestational age as described in the above diagram) corresponding to the total value. As described above, one or more calibration data points may be determined from a plurality of calibration samples having a known gestational age and containing cell-free DNA molecules. In some embodiments, the one or more calibration data points may be a plurality of calibration data points that form a calibration function approximating the measured total value determined from cell-free DNA molecules in the plurality of calibration samples having a known gestational age.

[0181] In step 2060, the total value is compared with the calibration value of at least one calibration data point. For example, the new total value of the new sample can be compared with the mean value for late pregnancy as determined in Figure 8A. As another example, the calibration value of at least one calibration data point may correspond to the total value measured using the molecular weight of free DNA in at least one of a plurality of calibration samples. The comparison of the total values ​​may be a plurality of calibration values, for example, each corresponding to one of a plurality of calibration samples. The comparison may occur by fitting the total value to a function (calibration function) into calibration data that provides the change in the total value relative to gestational age. The comparison may be made, for example, with reference to step 1950 in a similar manner to that described in method 1900.

[0182] In step 2070, the gestational age of the fetus is estimated based on this comparison. For example, if the new total value is closest to the late-pregnancy mean (or another calibration value used), the new sample can be identified as being in late pregnancy. As another example, the new total value can be compared with a calibration function (e.g., a linear function) fitted to data in Figure 8A or other similar graphs. The function can output gestational age, for example, as the Y-value of a linear function. Other examples of calibration functions provided herein can also be used in the context of determining gestational age. [III.] [Representational Approach] []

[0183] Using genotype-based analysis, the presence of plasma DNA terminal motifs in pregnant individuals, cancer patients, and liver transplant recipients is related to the tissue of origin. We hypothesize that in cancer patients, the release of tumor DNA into the bloodstream alters the original, normal presentation of plasma DNA terminal motifs. However, we do not rule out the possibility that other aspects of cancer pathology, such as the tumor microenvironment (infiltrating T cells, B cells, neutrophils, etc.), may produce different terminal motifs, exerting a lateral influence on the terminal motifs. Therefore, analysis of plasma DNA terminal motifs between cancer individuals and non-cancer controls will reveal the ability of controls to classify HCC.

[0184] Figure 21 illustrates a schematic diagram of a phenotypic method for plasma DNA terminal motif analysis according to an embodiment of this disclosure. Figure 21 is similar to Figures 2 and 12, for example, it can plot relative frequencies, determine variance values ​​(e.g., entropy), and perform clustering.

[0185] In Figure 21, terminal motifs (e.g., tetramers) inferred from plasma DNA molecules are used for comparison between cancer and control individuals. This avoids the limitations of genotype markers and makes it widely applicable in many clinical scenarios, such as the detection of autoimmune diseases (e.g., systemic lupus erythematosus, SLE) and transplantation. Using a phenotype-based approach with all sequenced plasma DNA fragments, entropy and cluster analyses can be performed in very similar analytical steps as in genotype-based methods. In this case, entropy and cluster analyses will be compared between controls and diseased individuals.

[0186] Pathogenic molecule 2105 was derived from individuals identified as having one or more of the disease. Control molecule 2107 was derived from individuals not having one or more of the disease. The relative frequencies of one set of terminal motifs from the two molecular libraries were determined. Bar graph 1220 provides the relative frequency (%) of each tetramer appearing as a terminal motif for control and pathogenic sequences. Such relative frequencies can be determined with respect to bar graph 220 of Figure 2 as described above. As can be seen, terminal motif 2122 shows significant differences in relative frequency between DNA fragments in different tissue types. Such differences can be used for various purposes, such as classifying new samples as pathogenic or non-pathogenic, or some other degree of disease.

[0187] To capture the overall differences in terminal motifs between tumors and shared DNA molecules, an entropy-based analysis, similar to Figure 2, can be used. Curve 2135 shows the entropy values ​​for controls and diseased individuals. Differences in entropy or other measures of variance can provide a classification of the pathological grade associated with the disease.

[0188] In another embodiment, cluster-based analysis 2140 can be performed, similar to the fetal analysis in Figure 2 and the tumor analysis in Figure 12. The pathological grade classification can be determined based on new samples belonging to a reference cluster with a known classification.

[0189] Therefore, in one instance of the total relative frequency, the characteristics of each individual may include a vector (i.e., a 256-dimensional vector) of 256 frequencies relating to the terminal motifs of the tetramer. In other instances, the standard deviation (SD), coefficient of variation (CV), interquartile range (IQR), or a percentile cutoff (e.g., the 95th or 99th percentile) of the different motif frequencies can be used to assess the overall change in the terminal motif pattern between the disease and control groups. Other instances of total values ​​are provided in other sections and are applicable herein. A. Oncology

[0190] In some embodiments, the disease (pathology) may be cancer. Therefore, some embodiments may classify cancer grades. [, 1. , ] [, Differences in relative frequencies arranged in order , ] [, , ]

[0191] Figure 22 illustrates an example of the frequency profile of tetramer terminal motifs of all plasma DNA molecules between individuals with hepatocellular carcinoma (HCC) and hepatitis B virus (HBV) according to an embodiment of this disclosure. Figure 22 compares the frequencies of 256 terminal motifs in an HCC patient and an HBV individual. Like similar curves, the vertical axis represents motif frequencies and the horizontal axis corresponds to individual terminal motifs. In Figure 22, the motifs are arranged in ascending order based on the average motif frequencies in non-HCC individuals. The bottom curves follow the top curves, but are scaled differently for clarity.

[0192] Several terminal motifs exhibit abnormalities in HCC patients. For example, compared to HBV individuals, the mean fold change of the top 10 terminal motifs (TGGG, TAAA, AAAA, GAAA, GGAG, TAGA, GCAG, TGGT, GCTG, and GAGA) showing increased frequency in HCC patients was 1.22-fold, ranging from 1.12 to 1.35-fold; and the mean fold change of the top 10 terminal motifs showing decreased frequency in HCC patients (CCCA, CCAG, CCAA, CCCT, CCTG, CCAC, CCAT, CCCC, CCTC, and CCTT) was 1.23-fold, ranging from 1.16 to 1.29-fold. This set of leading motifs showing increased (or decreased) frequency in the HCC group relative to the non-cancer group can be used to classify new individuals with the relevant cancer. As another example, the ranking process can optionally display all primitives corresponding to HCC individuals, and then sort those primitives in descending order based on the AUC between HCC individuals and non-HCC individuals. The top 10 primitives are then selected based on the AUC values.

[0193] To assess the diagnostic potential of plasma DNA terminal motifs, we sequenced 20 healthy controls (controls), 22 chronic hepatitis B virus carriers (HBV), 12 patients with cirrhosis (Cirr), 24 early-stage HCC (eHCC), 11 intermediate-stage HCC (iHCC), and 7 late-stage HCC (aHCC), with a median paired read count of 215 million (range: 97-1681 million).

[0194] Figure 23A shows a box plot of the combined frequencies of the top 10 plasma DNA tetramer terminal motifs in various individuals with different cancer levels according to embodiments of this disclosure. The top 10 plasma DNA tetramer terminal motifs were selected based on the data in Figure 22, i.e., based on the frequencies of HBV individuals. The combined frequency is the sum of the frequencies of the 10 terminal motifs in a given individual. We found that the combined frequencies of the top 10 rating terminal motifs were significantly lower in HCC patients compared to non-cancer individuals (p < 0.0001). Importantly, using this terminal motif analysis, 58.3% of eHCC patients could be identified with 95% specificity. Furthermore, different stages of cancer could be detected. For example, advanced HCC had substantially lower values ​​than eHCC and iHCC.

[0195] Figure 23B shows the receiver operating characteristic (ROC) curves of the combination frequencies of the top 10 plasma DNA tetramer terminal motifs between HCC and non-cancer individuals according to an embodiment of this disclosure. The area under the ROC curve (AUC) was found to be 0.91, indicating the clinical potential of plasma DNA terminal motifs to effectively distinguish HCC from non-cancer individuals. In another embodiment, the combination frequencies of the seven terminal motifs with the largest interval between HCC and non-HCC individuals provided an AUC of 0.92.

[0196] Figure 24A shows a box plot of the frequencies of CCA motifs in different groups according to an embodiment of this disclosure. The most common trimer motif (CCA) in the non-HCC group shows a significant decrease in the HCC group (p < 0.0001). Figure 24B shows the ROC curves of the most frequent trimer motif (CCA) present in non-HCC individuals between the non-HCC and HCC groups according to an embodiment of this disclosure. The AUC was found to be 0.915. The most common tetramer (CCCA) also provided a similar AUC of 0.91. [, 2. , ] [, Use of entropy (primitive diversity scoring) , ] [, , ]

[0197] Figure 25A shows a box plot of entropy values ​​in different groups using 256 tetramer terminal motifs according to an embodiment of this disclosure. All 256 motifs of the tetramer are used. As shown in Figure 25A, the entropy values ​​of HCC patients (mean: 5.242; range: 5.164-5.29) were significantly increased compared to non-HCC individuals (mean: 5.203; range: 5.124–5.253) (p < 0.0001). Importantly, using this terminal motif analysis, 41.7% of eHCC patients could be identified with 95% specificity. Entropy was generally increased in the HCC, IHCC, and advanced HCC groups compared to the non-HCC group. Furthermore, different stages of cancer could be detected. For example, advanced HCC had substantially higher values ​​than eHCC and iHCC.

[0198] Figure 25B shows a box plot of entropy values ​​in different groups using 10 tetramer terminal motifs according to an embodiment of this disclosure. Here, HCC individuals have reduced entropy relative to non-HCC individuals. Therefore, the set of terminal motifs used can change the relationship from increasing to decreasing. For example, using the first 10 motifs, entropy decreases in the HCC group. In either case, there is diagnostic capability between the HCC and non-HCC groups, as well as late-stage HCC relative to early-stage HCC.

[0199] Figure 26A shows a box plot of the entropy values ​​of the trimer terminal motifs used in different groups according to an embodiment of this disclosure. It was found that the entropy of HCC individuals using trimer motifs (a total of 64 motifs) was significantly higher than that of non-HCC individuals (p < 0.0001). Figure 26B shows the ROC curves of the entropy of the 64 trimer motifs used in the non-HCC and HCC groups according to an embodiment of this disclosure. The AUC was found to be 0.872.

[0200] As explained above, higher entropy values ​​indicate greater diversity in terminal primitives. As further illustration of the ability to use diversity scores to differentiate between various cancer types and control (e.g., healthy) samples, data from published studies are used.

[0201] Figures 27A and 27B show box plots of primitive diversity scores using different groups of tetramers according to an embodiment of this disclosure. All 256 tetramers were used to determine the primitive diversity score. When we perform MDS analysis using plasma DNA sequencing results downloaded from publicly available studies, increased plasma DNA terminal diversity is typically observed in various cancer types (Song et al. 2017), which reflects the fact that different tumor cells from different structural sites release their DNA into the bloodstream (Bettegowda et al. 2014). The cancers analyzed were: hepatocellular carcinoma (HCC), lung cancer (LC), breast cancer (BC), gastric cancer (GC), glioblastoma multiforme (GBM), pancreatic cancer (PC), and colorectal cancer (CRC).

[0202] To further test the generalizability of MDS changes across different cancer types, we sequenced an independent group of 40 plasma DNA samples with other cancer types, including patients with colorectal cancer (n=10), lung cancer (n=10), nasopharyngeal carcinoma (n=10), and head and neck squamous cell carcinoma (n=10), with a median of 0.42 billion bilateral sequenced reads (range: 0.19–0.65 billion). As shown in Figure 27B, the MDS value in the cancer patient group (median: 0.943; range: 0.939–0.949) was significantly higher than that in the cancer-free control group (median: 0.941; range: 0.933–0.946; p < 0.0001, Wilcoxon sum-rank test)

[0203] Figure 28 shows receiver operating curves for various techniques used to distinguish between healthy controls and cancer according to embodiments of this disclosure. We had a total of 129 samples, including healthy controls (n=38), hepatitis B virus carriers (n=17), hepatocellular carcinoma patients (n=34), colorectal cancer patients (n=10), lung cancer patients (n=10), nasopharyngeal carcinoma patients (n=10), and head and neck squamous cell carcinoma patients (n=10). Interestingly, compared with fragment size 2803 (AUC=0.74, p=0.0040; DeLong test), fragment-biased end 2804 (AUC=0.52, p<0.0001) (Jiang et al. 2018) and... Compared to other fragmentation measures such as directionally conscious plasma fragmentation signal, OCF, and 2802 (AUC=0.68, p=0.0013) (Sun et al. 2019), the MDS-based method 2801 (AUC=0.85) appears to have the best power (Yu et al. 2017b). If any of the techniques classifies an individual as having cancer, then the combined analysis 2805 identifies the individual as having cancer.

[0204] For primitives of different lengths, the accuracy of MDS analysis in distinguishing between cancerous and non-cancer cells was maintained relatively well. MDS analysis was performed on 1- to 5-mers.

[0205] Figure 29 illustrates receiver operating curves for MDS analysis using various k-mers according to embodiments of this disclosure. The MDS values ​​inferred from the 1- to 5-mer motifs also enhance the ability to distinguish between patients with and without cancer. The 1-mer analysis 2901 provides 0.81 AUC. The 2-mer analysis 2902 provides 0.85 AUC. The 3-mer analysis 2903 provides 0.85 AUC. The 4-mer analysis 2904 provides 0.85 AUC. The 5-mer analysis 2905 provides 0.81 AUC.

[0206] We also explored the impact of tumor DNA score on the effectiveness of MDS-based cancer detection using computer simulations.

[0207] Figure 30 illustrates the performance of MDS-based cancer detection for various tumor DNA fractions according to embodiments of this disclosure. As shown in Figure 30, the effectiveness of cancer detection progressively improves with increasing tumor DNA fraction in plasma DNA. For example, for patients with a tumor DNA fraction of 0.1%, the area under the ROC curve (AUC) is only 0.52, while for patients with a tumor DNA fraction of 3%, the AUC increases to 0.9, and further increases at higher concentrations, but approaches its maximum at a tumor fraction of 5%. [, 3. , ] [, Machine learning ( , ] [, SVM , ] [, (Regression and Clustering) , ] [, , ]

[0208] To further explore whether a classifier could be constructed to detect cancer patients using plasma DNA terminal motifs, we used 256 plasma DNA terminal motifs to construct a classifier to distinguish between patients with (n=55) cancer and those without (n=74) cancer, employing Support Vector Machine (SVM) and logistic regression, which takes into account the magnitude and orientation of each terminal motif. SVM analysis identified the hyperplane that best distinguishes between cancer and non-cancer patients across 256 dimensional locations, where the training data points were the frequencies of each of the 256 motifs in the tetramer. Logistic regression determined the coefficients by which each of the 256 frequencies is multiplied, and also determined the cutoff value of the output of the logarithmic function, which could be the weighted sum of the frequencies multiplied or the input to receive the weighted sum. Those familiar with this technique will recognize that such a logarithmic function can be a sigmoid function or other initiation functions.

[0209] To minimize the fitting problem, we employ a leave-one-out procedure to evaluate its performance using receiver operating characteristic (ROC) curve analysis. The leave-one-out procedure is performed according to the following steps: In N sample sizes, we use one sample as the test sample and use the remaining samples (N-1) to train a classifier based on SVM and logistic regression using 256 plasma DNA terminal motifs. We then use the trained classifier to determine whether the remaining samples are classified as belonging to individuals with or without cancer. We systematically leave one sample as the test sample to test the classifier trained on the remaining samples. Therefore, we obtain the prediction results for each sample and calculate the accuracy based on the prediction results.

[0210] Figure 31 shows the receiver operating curves for MDS, SVM, and logistic regression analyses according to embodiments of this disclosure. We observed that the AUC increase using a classifier with 256 terminal primitives (AUC = 0.89 for both SVM and logistic regression) was less compared to the MDS-based analysis (AUC = 0.85).

[0211] As another machine learning technique, we use clustering based on the frequency of terminal primitives.

[0212] Figure 32 illustrates a hierarchical clustering analysis of the top ten terminal motifs for groups with different cancers and different cancer levels, according to an embodiment of this disclosure. As shown, HCC individuals (eHCC: early HCC 3205; iHCC: mid-stage HCC 3230; and aHCC: late-stage HCC 3225) typically cluster together, and non-HCC individuals (healthy controls; HBV: chronic hepatitis B carriers) typically cluster together. For example, the cluster on the right is early HCC 3205 (yellow). The left middle mainly consists of controls 3210, HBV 3215, and cirrhosis 3220. The different clustering patterns between the HCC and non-HCC groups indicate that terminal motifs will reflect disease-related preferences in plasma DNA terminal motifs and suggest the potential diagnostic capability of plasma DNA terminal motifs. In addition to connection-based hierarchical clustering, other clustering techniques can also be used as statistical methods, such as center-based clustering, distribution-based clustering, and density-based clustering.

[0213] Figures 33A-33C illustrate hierarchical clustering analysis of all plasma DNA molecules in groups with different cancers and different cancer levels, according to embodiments of this disclosure. Figure 33A shows hierarchical clustering analysis based on the frequencies of 256 tetramer terminal motifs. Figure 33B shows a magnified view of the hierarchical clustering analysis based on the frequencies of 256 tetramer terminal motifs. Each row represents a type of terminal motif. Each column represents an individual plasma DNA sample. Gradient colors indicate the frequency of terminal motifs. Red indicates the highest frequency and green indicates the lowest frequency. Figure 33C shows principal component analysis (PCA) of HCC and non-HCC individuals using terminal motifs. The principal components are linear combinations of the 256 motifs that provide the maximum variance, for example, in the weighted sum of the resulting frequencies.

[0214] Since HCC and non-HCC individuals exhibit two distinct clusters, the terminal motifs derived from all plasma DNA molecules are an important measure for distinguishing between HCC and non-HCC individuals. Figures 33A and 33B show that HCC individuals 3305 (red) tend to cluster into one group, while non-HCC individuals 3310 (blue) tend to cluster into another. In Figure 33C, PCA analysis also shows that HCC and non-HCC individuals tend to cluster into two distinct groups. PC1 and PC2 correspond to different linear combinations of relative frequencies (e.g., weighted averages), which can represent the pattern of a given histogram of relative frequencies. Figure 33C shows that linear combinations (or other transformations) can be performed before clustering or using cutoff values ​​or cutoff planes. Therefore, transformed relative frequencies can be used to determine the total value.

[0215] Figure 34 illustrates a hierarchical clustering analysis based on the trimer motifs of all plasma DNA molecules in different groups with different cancer levels, according to an embodiment of this disclosure. For ease of illustration, only the top portion of the heatmap is shown. As shown, HCC individuals (eHCC: early HCC 3405; iHCC: mid-stage HCC 3430; and aHCC: late-stage HCC 3425) typically cluster together, and non-HCC individuals (healthy control 3410; HBV 3415: chronic hepatitis B carrier; and cirrhosis 3420) typically cluster together.

[0216] Based on these findings, machine learning (e.g., deep learning) models can be used to train cancer classifiers using 256-dimensional vectors that include terminal motifs of plasma DNA. This includes, but is not limited to, Support Vector Machines (SVMs), decision trees, naive Bayes classification, logistic regression, clustering algorithms, PCA, single-valued factorization (SVD), t-distributed random neighborhood embeddings (tSNEs), artificial neural networks, and ensemble methods that construct an ensemble of classifiers and then classify new data points by weighted voting based on their predictions. Once the cancer classifier is trained on a matrix based on 256-dimensional vectors containing a range of cancer and non-cancer patients, it will be able to predict the probability of new patients developing cancer.

[0217] In such applications of machine learning algorithms, the total value may correspond to a probability or distance that can be compared to a reference value (e.g., when using SVMs). In other embodiments, the total value may correspond to an earlier output in the model (e.g., an earlier layer in a neural network) compared to a cutoff value between two classifications or to a representative value for a given classification. B. Monitoring of immune diseases

[0218] Figure 35A illustrates an entropy analysis of all plasma DNA molecules between healthy controls and SLE patients according to an embodiment of this disclosure. Figure 35B illustrates a hierarchical clustering analysis of all plasma DNA molecules between healthy controls and SLE patients according to an embodiment of this disclosure.

[0219] A comprehensive analysis of the overall abnormalities in plasma DNA terminal motifs, including entropy (Fig. 35A, p-value: 0.00014) and cluster analysis (Fig. 35B), demonstrated that SLE patients are distinguishable from healthy controls. For example, individuals with SLE showed increased entropy (Fig. 35A), and two clusters typically formed on the left (SLE 3510) and right (control / normal 3505). Therefore, autoimmune diseases alter plasma DNA fragmentation patterns, thereby demonstrating the ability to distinguish plasma DNA terminal motifs between SLE and control individuals.

[0220] Figure 36 illustrates an embodiment of this disclosure, using entropy analysis of plasma DNA molecules with 10 selected terminal motifs between healthy controls and SLE patients. The top 10 motifs with the highest relative frequencies for the controls are used. As with other phenotypes, the set of motifs can influence whether SLE entropy is higher or lower. Given that 10 motifs are selected as having the highest values ​​for the controls, the entropy is higher because the values ​​are similar to each other (i.e., due to permutation). And SLE entropy is lower due to the presence of more variation, for example, because it is not permuted for SLE individuals. The opposite relationship could exist if the top 10 motifs are selected from the SLE sample. Therefore, the grade of an autoimmune disease (e.g., SLE) can be determined using the total value of relative frequencies. C. Composite analysis of terminal primitives and traditional measures

[0221] We are testing whether combined analysis of plasma DNA terminal motifs and other metrics (duplicate number bias (CNA), hypomethylation, and hypermethylation) will improve the performance of non-invasive cancer detection. For example, decision tree-based classification can be used for combined analysis.

[0222] Figure 37 shows the ROC curves for combined analyses of terminal motifs and duplicate counts or methylation in individuals with and without HCC, according to embodiments of this disclosure. Terminal motif analysis used a motif diversity score determined using all 256 motifs of the tetramer. If either analysis resulted in a cancer classification, the combined analysis identified cancer. The combined analysis of terminal motifs and methylation (AUC: 0.94) or the combined analysis of terminal motifs and CNA (AUC: 0.93) was superior to the analysis using only terminal motifs (AUC: 0.86). Methylation analysis used a higher number of 1 Mb cells with low methylation (defined as a methylation density z-score < -3) than the normal control, where the cutoff number of aberrant cells distinguished between cancer and non-cancer. CNA analysis used the number of 1 Mb cells with a z-score greater than or less than 3, where the cutoff number of aberrant cells distinguished between cancer and non-cancer. Further details of the methylation analysis can be found in U.S. Patent Publication 2014 / 0080715 and further details of the CNA analysis can be found in U.S. Patent Publication US 2013 / 0040824.

[0223] This describes an instance-based decision tree classification. For example, we can use a random forest algorithm to infer cutoff values ​​for various metrics, including CNA, hypomethylation, hypermethylation, size (e.g., as described in U.S. Patent Publication 2013 / 0237431), terminal primitives, and fragmentation patterns (e.g., as described in U.S. Patent Publications 2017 / 0024513 and 2019 / 0341127 and U.S. Patent Application 16 / 519,912). Each metric will have a specific cutoff value. Taking a metric (hypomethylation) as an example, a case can be classified as cancer or non-cancer, depending on whether the metric value is below or above the cutoff value. A metric represents a node in the decision tree. After the sample has traversed all nodes in the entire tree, for example, a majority vote (e.g., the number of nodes indicating cancer is greater than the number of nodes indicating non-cancer) provides the final classification. D. Examples of alternative methods for defining the terminal motifs of plasma DNA

[0224] To demonstrate the feasibility of using an alternative method to define the terminal motif of plasma DNA, technique 160 in Figure 1 was adopted for analysis of individuals with HCC and those without HCC, including 20 sequenced healthy controls (controls), 22 chronic hepatitis B virus carriers (HBV), 12 individuals with cirrhosis (Cirr), 24 individuals with early-stage HCC (eHCC), 11 individuals with intermediate-stage HCC (iHCC), and 7 individuals with advanced-stage HCC (aHCC).

[0225] Figure 38A illustrates a tetramer-based entropy analysis according to an embodiment of this disclosure, where the tetramer is constructed from the ends of a sequenced plasma DNA fragment and its adjacent genomic sequences in HCC and non-HCC individuals. Entropy is determined using all 256 end motifs. As with the analysis using technique 140 of Figure 1 to define the motifs, the entropy of HCC individuals differs from that of non-cancer individuals. Furthermore, late-stage HCC shows substantial differences from eHCC and iHCC. Figure 38B illustrates a tetramer-based clustering analysis according to an embodiment of this disclosure, where the tetramer is constructed from the ends of a sequenced plasma DNA fragment and its adjacent genomic sequences in HCC individual 3810 and non-HCC individual 3805.

[0226] Figure 39 illustrates an ROC comparison of techniques 140 and 160 used in Figure 1 to define the terminal motifs of plasma DNA according to an embodiment of this disclosure. The same individuals as shown in Figure 38A were used, and entropy analysis using tetramers was performed for classification. Method (i) corresponds to technique 140, and method (ii) corresponds to technique 160. Slightly worse performance was observed when using technique 160 in Figure 1 compared to technique 140 (AUC: 0.815 vs. 0.856). E. Filtering to improve discrimination

[0227] Specific criteria can be used to filter specific DNA fragments (other than terminal motifs) to provide greater accuracy, such as sensitivity and specificity. As an example, terminal motif analysis can be limited to DNA fragments derived from open chromatin regions of a specific tissue, such as by determining reads that are entirely or partially aligned within one of several open chromatin regions. For instance, any read having at least one nucleotide overlapping with an open chromatin region can be defined as a read within an open chromatin region. A typical open chromosomal region is approximately 300 bp, depending on the DNase I hypersensitive site. The size of an open chromatin region can vary depending on the technique used to define it, for example, ATAC sequences (for chromatin sequencing analysis available from translocases) versus DNase I sequences.

[0228] As another example, DNA fragments of specific sizes can be selected for end motif analysis. As shown below, this increases the interval of the total relative frequencies of end motifs, thereby improving accuracy.

[0229] Another example is the use of the methylation properties of DNA fragments. Fetal and tumor DNA are typically hypomethylated. Examples can determine the methylation measure (e.g., density) of a DNA fragment (e.g., the proportion or absolute number of methylated sites on the DNA fragment). DNA fragments can be selected for end motif analysis based on the measured methylation density. For example, a DNA fragment can only be used if the methylation density is above a threshold value.

[0230] Regardless of whether the DNA fragment contains sequence variations relative to the reference genome (such as base substitutions, insertions, or deletions), it can also be used for filtering.

[0231] Various filtering criteria can be combined. For example, it may be necessary to meet each criterion individually, or at least a specific number of criteria may need to be met. In another embodiment, the probability that a fragment corresponds to clinically relevant DNA (e.g., fetus, tumor, or transplant) can be determined, and the DNA fragment can be determined to meet a threshold value imposed by probability before being used for end motif analysis. As another example, the proportion of a DNA fragment to a frequency counter for a specific end motif can be weighted based on probability (e.g., adding the probability of a value less than one, rather than adding one). Thus, DNA fragments with a specific end motif will be weighted more heavily and / or have a higher probability. Such enrichment is further described below. [1.] [Terminal motifs in tissue-specific chromatin regions] []

[0232] Because different tissues exhibit preferred fragment patterns during apoptosis (Chan et al., Proceedings of the National Academy of Sciences of the United States of America, 2016; 113: E8159-8168; Jiang et al., Proceedings of the National Academy of Sciences of the United States of America, 2018; doi:10.1073 / pnas.1814616115), we further infer that the selection of a specific genomic region for plasma DNA terminal motif analysis will further improve the ability to classify and differentiate between patients and controls. Using HCC patients as an example, we employed open chromatin regions from blood and liver.

[0233] Figure 40 illustrates a comparison of the accuracy of embodiments according to this disclosure, demonstrating that tissue-specific open chromatin regions improve the ability to distinguish terminal DNA motifs in plasma from HCC and non-cancer patients. Entropy analysis of all 256 motifs was performed using the combination frequencies of tetramers and the first 10 motifs. For liver open chromatin results, reads containing at least one nucleotide overlapping with one of the liver open chromatin regions were retained (i.e., not filtered out).

[0234] The ability to identify terminal motifs of plasma DNA molecules that overlap with open chromatin regions of the liver yielded the best performance with an AUC of 0.918, using the combination frequency of the first 10 permutation motifs. In contrast, without any selection, the ability to identify terminal motifs of plasma DNA molecules had a minimum AUC of 0.855.

[0235] Therefore, when screening for specific tissues for cancer, DNA fragments from open chromatin of that specific tissue (or at least where the terminal sequences are located in open chromatin regions) can be used for analysis, instead of DNA fragments not located in these identified regions. In this case, the liver is used because the cancer is HCC. The location of the DNA fragment can be determined by aligning the sequence reads to a reference genome, where open chromatin regions can be identified from literature or databases. [2.] [Analysis of End Elements Based on Size Zones] []

[0236] The frequency of certain terminal motifs varies depending on the analyzed size range (size band), as shown by the percentage of CCCA. This suggests that size band-based terminal motif analysis can affect the efficacy of using plasma DNA terminal motifs to distinguish between cancer patients and non-cancer individuals. To illustrate this possibility, we tested a range of size ranges, including but not limited to 50–80 bp, 81–110 bp, 111–140 bp, 141–170 bp, 171–200 bp, and 201–230 bp, to investigate how the analyzed size bands affect overall diagnostic efficacy.

[0237] Figure 41 illustrates plasma DNA end motif analysis based on size bands according to an embodiment of this disclosure. Classification using 256 motifs of the tetramer was determined using a motif diversity score (entropy). Various ranges are listed in Figure 41, but other ranges may be used. Analysis 4101 (50-80) provides 0.826 AUC. Analysis 4102 (81-110) provides 0.537 AUC. Analysis 4103 (111-140) provides 0.551 AUC. Analysis 4104 (141-170) provides 0.716 AUC. Analysis 4105 (171-200) provides 0.769 AUC. Analysis 4106 (201-230) provides 0.756 AUC.

[0238] This size range can be used for techniques to enrich clinically relevant DNA. For example, selecting DNA molecules of 50-80 bases will enrich a sample of tumor DNA. In contrast to a single size range, multiple disjoint size ranges can be used. This type of enrichment may occur because a better AUC occurs in the 50-80 base range compared to the 81-110 base range.

[0239] Terminal motifs from plasma DNA molecules in the 50 to 80 bp range exhibit optimal discriminative ability for detecting HCC in non-HCC individuals (AUC: 0.83). Therefore, examples can filter DNA fragments to select DNA fragments within a specific size range, and then use the selected DNA fragments (reads) to determine relative frequencies and subsequent operations. As an example, size filtering can be performed via physical spacing or by determining the size using sequence reads (e.g., if the entire fragment is sequenced or by aligning bilateral sequencing with a reference length). Examples of physical enrichment of shorter DNA include during gel electrophoresis, by collecting the lysate at a certain retention time during capillary electrophoresis, after liquid chromatography, or by band cleavage using microfluidics. F. Classification of Pathological Grades

[0240] Figure 42 is a flowchart illustrating a method 4200 for classifying pathological grades in an individual's biological sample according to an embodiment of the present disclosure. The biological sample contains cell-free DNA. Method 4200 can be performed in a similar manner to method 1900 of Figure 19 and method 2000 of Figure 20.

[0241] In step 4210, multiple cell-free DNA fragments from a biological sample are analyzed to obtain sequence reads. Each sequence read contains terminal sequences corresponding to the ends of the multiple cell-free DNA fragments. Step 4210 can be performed in a similar manner to step 1910 in Figure 19.

[0242] In step 4220, for each of the plurality of free DNA fragments, the sequence motif of each of one or more terminal sequences of the free DNA fragment is determined. Step 4220 can be performed in a similar manner to step 1920 of Figure 19.

[0243] In step 4230, the relative frequencies of a set of one or more sequence motifs corresponding to the terminal sequences of a plurality of free DNA fragments are determined. The relative frequencies of the sequence motifs provide the proportion of a plurality of free DNA fragments having terminal sequences corresponding to the sequence motifs. Step 4230 can be performed in a similar manner to step 1930 of Figure 19. For example, the set of one or more sequence motifs may contain N base positions. The set of one or more sequence motifs may contain all combinations of N bases. N may be an integer equal to or greater than three, and any other integer.

[0244] As another example, the set of one or more sequence motifs may be the top M sequence motifs showing the greatest difference between two types of DNA, as determined in one or more reference samples, such as all motifs showing the greatest positive difference (e.g., the top 10 or other numbers) or all motifs showing the greatest negative difference. M may be an integer equal to or greater than one. For methods 1900 and 2000, the two types of DNA may be clinically relevant DNA and another type of DNA. For method 4200, the two types of DNA may be from two reference samples with different classifications for pathological grades. As another example, the set of one or more sequence motifs may be the top M most common sequence motifs occurring in one or more reference samples, such as those shown in Figure 22, where the reference samples are non-cancer samples, such as HBV samples.

[0245] In step 4240, the total relative frequency of a set of one or more sequence primitives is determined. Step 4240 can be performed in a similar manner to step 1940 of Figure 19. Instances of total values ​​are described throughout the disclosure, and these instances include: entropy, combined frequency, differences (e.g., distance) with reference patterns of relative frequencies that can be implemented in clustering or using SVM or values ​​determined from the differences (e.g., probabilities), or outputs in machine learning models (e.g., intermediate or final layers in a neural network) compared to cutoff values ​​between two classifications or compared to representative values ​​for a given classification.

[0246] When a set of one or more sequence primitives contains a complex number of sequence primitives, the total value may contain the sum of the relative frequencies of that set. The sum may be a weighted sum. For example, the total value may contain an entropy term, which contains the sum of terms including the weighted sum. Each term may contain a relative frequency multiplied by the logarithm of the relative frequency. The total value may correspond to the variance of the relative frequencies.

[0247] In another instance, the total value includes the final or intermediate output of the machine learning model. In various implementations, the machine learning model uses clustering, support vector machines, or logistic regression.

[0248] In step 4250, the pathological grade classification of an individual can be determined based on a comparison of the total value and a reference value. For example, the pathology could be cancer or an autoimmune disease. For example, the grade could be non-cancerous, early, intermediate, or late. The classification can then be one of these grades. Therefore, the classification can be determined by a plurality of cancer grades encompassing a plurality of cancer stages. For example, cancer could be hepatocellular carcinoma, lung cancer, breast cancer, gastric cancer, glioblastoma multiforme, pancreatic cancer, colorectal cancer, nasopharyngeal carcinoma, and squamous cell carcinoma of the head and neck. For example, an autoimmune disease could be systemic lupus erythematosus.

[0249] In other instances, the pathological grade corresponds to the fractional concentration of clinically relevant DNA associated with the pathology. For example, the pathological grade could be cancer and the clinically relevant DNA could be tumor DNA. Reference values ​​could be calibration values ​​determined by self-calibrated samples, as described with respect to method 1900.

[0250] In some embodiments, cell-free DNA is filtered to identify multiple cell-free DNA fragments. Examples of filtering are provided in the preceding sections. For example, filtering may be based on methylation (density or whether a specific site is methylated), size, or the region of origin of the DNA fragment. Cell-free DNA may be filtered for DNA fragments originating from open chromatin regions of a particular tissue. [IV.] [Enrichment] []

[0251] Preferred selection of DNA fragments from a specific tissue that present a specific set of terminal motifs can be used to enrich samples of DNA from that specific tissue. Thus, examples can enrich samples of clinically relevant DNA. For instance, DNA fragments having only specific terminal sequences can be used for analytical sequencing, amplification, and / or capture. As another example, sequence read filtering can be performed, for example, in a manner similar to that described in Section III. E. A. Physical enrichment

[0252] Physical enrichment can be performed in various ways, such as via targeted sequencing or PCR, using specific primers or grafts. If a specific end motif of the terminal sequence is detected, a graft can then be added to the end of the fragment. Subsequently, during sequencing, only the DNA fragment with the graft is sequenced (or at least primarily sequenced), thereby providing targeted sequencing.

[0253] As another example, primers that hybridize to a specific set of terminal motifs can be used. These primers can then be used for sequencing or amplification. Capture probes corresponding to specific terminal motifs can also be used to capture DNA molecules having those terminal motifs for further analysis. Some embodiments involve ligating shorter oligonucleotides to the ends of plasma DNA molecules. Probes can then be designed to recognize sequences that are partly terminal motifs and partly conjugated oligonucleotides.

[0254] Some embodiments may utilize CRISPR-based diagnostic techniques, such as using guide RNA to locate sites corresponding to preferred end motifs for clinically relevant DNA, followed by DNA fragment cleavage using nucleases, such as Cas-9 or Cas-12. For example, appendix can be used to identify end motifs, and CRISPR / Cas9 or Cas-12 can then be used to cleave end motif / appendix hybrids and generate universally predictable ends for further enrichment of molecules with desired ends.

[0255] Figure 43 is a flowchart illustrating a method 4300 for enriching a biological sample containing clinically relevant DNA according to an embodiment of this disclosure. The biological sample contains clinically relevant DNA molecules and other free DNA molecules. Method 4300 may use specific assays to perform the enrichment.

[0256] In step 4310, a plurality of cell-free DNA fragments are received from the biological sample. Clinically relevant DNA fragments (e.g., fetal or tumor fragments) have terminal sequences containing sequence motifs that appear at a higher frequency than those of another DNA (e.g., maternal DNA, healthy DNA, or blood cells). As an example, data from Figures 3 and 13 can be used. Therefore, sequence motifs can be used to enrich clinically relevant DNA.

[0257] In step 4320, a plurality of cell-free DNA fragments are subjected to one or more probe molecules, which detect sequence motifs in the terminal sequences of the plurality of cell-free DNA fragments. Such use of probe molecules allows for the acquisition of the detected DNA fragments. In one example, the one or more probe molecules may contain one or more enzymes that query the plurality of cell-free DNA fragments and append new sequences for amplifying the detected DNA fragments. In another example, the one or more probe molecules may be attached to a surface to detect sequence motifs in the terminal sequences by hybridization.

[0258] In step 4330, a biological sample containing clinically relevant DNA fragments is enriched using the detected DNA fragments. As one example, enriching a biological sample containing clinically relevant DNA fragments using the detected DNA fragments may involve amplifying the detected DNA fragments. As another example, the detected DNA fragments may be captured, and undetected DNA fragments may be discarded. B. Electron hybridization enrichment

[0259] Electronic hybridization enrichment can use various criteria to select or discard certain DNA fragments. These criteria can include terminal motifs, open chromatin regions, size, sequence variations, methylation, and other epigenetic features. Epigenetic features encompass all modifications in the genome that do not involve changes in the DNA sequence. Criteria can specify cutoff values, such as requiring certain properties, like a specific size range, methylation measures above or below a certain amount, combinations of methylation states exceeding one CpG site (e.g., methylation haplotypes (Guo et al., *Nature Genetics*, 2017; 49: 635-42)), or the probability of combinations exceeding the threshold. Such enrichment can also involve weighting DNA fragments based on these probabilities.

[0260] As examples, enriched samples can be used for pathological classification (as described above), as well as for identifying tumors or fetal mutations, or for marking and counting amplifications / deletions of chromosomes or chromosomal regions. For instance, if a particular end motif or set of end motifs is associated with liver cancer (i.e., a higher relative frequency than for non-cancer or other cancers), then embodiments used for cancer screening may weight such DNA fragments above DNA fragments that do not have one of these preferred end motifs or this preferred set of end motifs.

[0261] Figure 44 is a flowchart illustrating a method 4400 for enriching clinically relevant DNA in a biological sample according to an embodiment of this disclosure. The biological sample contains clinically relevant DNA molecules and other free DNA molecules. Method 4400 may use specific sequence read criteria to perform enrichment.

[0262] In step 4410, multiple cell-free DNA fragments from a biological sample are analyzed to obtain sequence reads. Each sequence read contains terminal sequences corresponding to the ends of the multiple cell-free DNA fragments. Step 4410 can be performed in a similar manner to step 1910 in Figure 19.

[0263] In step 4420, for each of the plurality of free DNA fragments, the sequence motif of each of one or more terminal sequences of the free DNA fragment is determined. Step 4420 can be performed in a similar manner to step 1920 of Figure 19.

[0264] In step 4430, a set of one or more sequence motifs occurring in clinically relevant DNA at a relative frequency greater than that of another DNA is identified. The set of sequence motifs can be identified using the genotyping or phenotyping techniques described herein. Calibration or reference samples can be used to align and select sequence motifs that are selective for clinically relevant DNA.

[0265] In step 4440, a set of sequence reads with one or more sequence motifs in the terminal sequence is identified. This can be considered the first stage of filtering.

[0266] In step 4450, sequence reads corresponding to the probability of clinically relevant DNA exceeding a threshold value can be stored. The probability can be determined using a set of terminal motifs. For example, for each sequence read in a sequence read set, the probability of the sequence read corresponding to clinically relevant DNA can be determined based on the terminal sequence of the sequence read, which contains sequence motifs from a set of one or more sequence motifs. The probability can be compared to a threshold value. As an example, the threshold value can be determined empirically. For example, various threshold values ​​can be tested for a sample, and the concentration of clinically relevant DNA in the sample can be measured for a set of sequence reads. The optimal threshold value maximizes the concentration while maintaining a certain percentage of the total number of sequence reads. The threshold value can be determined by one or more given percentages (5th, 10th, 90th, or 95th) of the concentration of one or more terminal motifs present in healthy controls or control groups exposed to similar pathogen risk factors but without disease. The threshold value can be a regression or probability score.

[0267] When the probability exceeds a threshold, a read sequence can be stored in memory (e.g., in a file, table, or other data structure), thus obtaining the stored read sequence. Read sequences with a probability below the threshold can be discarded or not stored in the memory location where the read is stored, or a region of the database can contain markers indicating that a read has a lower threshold so that subsequent analysis can exclude such reads. As examples, various techniques (such as odds ratios, z-scores, or probability distributions) can be used to determine probability.

[0268] In step 4460, the stored sequence reads may be analyzed to determine the nature of the clinically relevant DNA in the biological sample, such as as described herein and in other flowcharts. Methods 1900, 2000, and 4200 are examples of this. For example, the nature of the clinically relevant DNA in the biological sample may be the fractional concentration of the clinically relevant DNA. As another example, the nature may be the pathological grade of the individual from whom the biological sample was obtained, wherein the pathological grade is associated with the clinically relevant DNA. As yet another example, the nature may be the gestational age of the fetus of the pregnant woman from whom the biological sample was obtained.

[0269] Other criteria can be used to determine probability. The size of multiple cell-free DNA fragments can be measured using sequence reads. The probability that a particular sequence read corresponds to clinically relevant DNA can be further based on the size of the cell-free DNA fragment corresponding to that particular sequence read.

[0270] Methylation can also be used. Therefore, examples can measure one or more methylation states at one or more sites of a cell-free DNA fragment corresponding to a specific sequence read. The likelihood that a specific sequence read corresponds to clinically relevant DNA can be further based on one or more methylation states. As another example, whether a read is within a set of identified open chromatin regions can be used as a filter.

[0271] Figure 45 shows an example graph of an embodiment according to this disclosure, illustrating the increase in fetal DNA fraction using the CCCA terminal motif. The vertical axis represents the fetal DNA fraction of the tested sample. Two sets of data are used for (1) all fragments overlapping with informative SNPs (i.e., fragments with fetal-specific paired genes) and (2) fragments with CCCA terminal motifs overlapping with informative SNPs. Thus, the data on the left provides the actual fetal DNA fragments in the entire sample, and the data on the right provides data for the electronically enriched sample. In this example, when the terminal motif is CCCA, the probability can be determined to be above the threshold value. More motifs can be used in a similar manner, for example, as a group indicating a probability above the threshold value.

[0272] The median increase in fetal DNA score was 3.2% (IQR: 1.3–6.4%). The relative increase in fetal DNA score was defined as (ba) / a × 100, where a is the original fetal DNA score calculated from all fragments overlapping with informative SNPs, where the mother is homozygous and the fetus is heterozygous, and b is the fetal DNA score calculated from fragments marked by CCCA motifs (i.e., enriched in the fetal DNA molecule).

[0273] For any of the methods described herein, the sequence motif for each of the one or more terminal sequences of the free DNA fragment can be determined using a reference genome (e.g., via technique 160 of Figure 1). Such techniques may include: aligning the one or more sequence reads corresponding to the free DNA fragment with a reference genome, identifying one or more bases in the reference genome adjacent to the terminal sequence, and using the terminal sequence and one or more bases to determine the sequence motif. [V.] [Instance System] []

[0274] Figure 46 illustrates a measurement system 4600 according to an embodiment of the present invention. As shown, the system contains a sample 4605, such as a free DNA molecule, within a sample holder 4610, wherein the sample 4605 can contact an analytical method 4608 to provide a signal of physical characteristic 4615. One example of a sample holder may be a channel or tube containing a probe and / or primer for testing, or a droplet for movement (in the case where the droplet contains a test). A detector 4620 detects the physical characteristic 4615 of the sample (e.g., fluorescence intensity, voltage, or current). The detector 4620 may perform measurements at time intervals (e.g., periodic time intervals) to obtain data points constituting a data signal. In one embodiment, an analog-to-digital converter converts the analog signal from the detector into digital form at multiple time intervals. The sample holder 4610 and the detector 4620 may form a testing device, such as a sequencing device for sequencing according to the embodiments described herein. Data signal 4625 is sent from self-detector 4620 to logic system 4630. Data signal 4625 can be stored in local memory 4635, external memory 4640, or storage device 4645.

[0275] The logic system 4630 may be or include a computer system, an ASIC, a microprocessor, etc. It may also include or be coupled to a display (e.g., a monitor, an LED display, etc.) and a user input device (e.g., a mouse, a keyboard, buttons, etc.). The logic system 4630 and other components may be part of a standalone or network-connected computer system, or may be directly connected to or incorporated into a device (e.g., a sequencing device) including the detector 4620 and / or the sample holder 4610. The logic system 4630 may also include software executing in the processor 4650. The logic system 4630 may include computer-readable media storing instructions for controlling the measurement system 4600 to perform any of the methods described herein. For example, the logic system 4630 may provide commands to a system including the sample holder 4610 to enable sequencing or other physical operations. Such physical operations may be performed in a specific order, such as in the case of reagents being added and removed in a specific order. Such physical operations can be performed by robotic systems (e.g., those containing robotic arms) that can be used to obtain samples and perform tests.

[0276] Any computer system mentioned herein may utilize any suitable number of subsystems. An example of such a subsystem is shown in computer system 10 in Figure 47. In some embodiments, the computer system comprises a single computer device, wherein a subsystem may be a component of the computer device. In other embodiments, the computer system may comprise multiple computer devices having internal components, each of which is a subsystem. The computer system may include desktop and laptop computers, tablet computers, mobile phones, and other mobile devices.

[0277] The subsystems shown in Figure 47 are interconnected via system bus 75. Other subsystems are shown, such as printer 74, keyboard 78, storage device 79, monitor 76 coupled to display adapter 82 (e.g., display screen, such as LED), and others. Peripheral devices and input / output (I / O) devices (coupled to I / O controller 71) can be connected to the computer system via any number of components known in this art (e.g., input / output (I / O) ports 77 (e.g., USB, FireWire®)). For example, I / O ports 77 or external interfaces 81 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect computer system 10 to a wide area network (e.g., Internet, mouse input device, or scanner). Interconnection via system bus 75 allows central processing unit 73 to communicate with the subsystems and control system memory 72 or storage device 79 (e.g., fixed disk, such as hard drive, or optical disk) to execute multiple instructions, as well as information exchange between subsystems. System memory 72 and / or storage device 79 may be implemented as computer-readable media. Another subsystem is a data collection device 85, such as a camera, microphone, accelerometer, and the like. Any of the data mentioned herein may be output from one component to another and may be output to a user.

[0278] A computer system may include a plurality of identical components or subsystems, for example, connected together via an external interface 81, an internal interface, or via a removable storage device that allows connection and removal from one component to another. In some embodiments, the computer system, subsystem, or device may communicate via a network. In such cases, one computer may be considered a client and another computer a server, each of which may be part of the same computer system. The client and server may each include multiple systems, subsystems, or components.

[0279] The embodiments can be implemented in a modular or integrated manner using hardware circuitry (e.g., application-specific integrated circuitry or field-programmable gate arrays) and / or computer software with a generally programmable processor. As used herein, the processor may include a single-core processor, a multi-core processor on the same integrated chip, a single circuit board or network hardware, or multiple processing units on dedicated hardware. Based on the disclosure and teachings provided herein, those skilled in the art will recognize and understand other ways and / or methods of implementing embodiments of the invention using hardware and combinations of hardware and software.

[0280] Any software component or function described in this application may be implemented in the form of software code using, for example, conventional or object-oriented techniques. This software code is executed by a processor using any suitable computer language (such as Java, C, C++, C#, Objective-C, Swift) or scripting language (such as Perl or Python). The software code may be stored on a computer-readable medium in the form of a series of instructions or commands for storage and / or transmission. Suitable non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), magnetic media (such as hard disk drives or floppy disk drives), or optical media such as optical discs (CDs) or DVDs (Digital Universal Discs) or Blu-ray discs, flash memory, and the like. Computer-readable media may be any combination of such storage or transmission devices.

[0281] Such programs can also be encoded and transmitted using carrier signals suitable for transmission over wired, optical, and / or wireless networks (including the Internet) conforming to various protocols. Therefore, computer-readable media can be constructed using data signals encoded with such programs. Computer-readable media encoded with program code can be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer-readable media can reside on or within a single computer product (e.g., a hard drive, CD, or an entire computer system) and can reside on or within different computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing the user with any of the results mentioned herein.

[0282] Any of the methods described herein can be performed wholly or partially using a computer system, which includes processors configured to perform one or more steps. Therefore, embodiments may potentially use different components for individual steps or groups of steps for a computer system configured to perform the steps of any of the methods described herein. Although presented as numbered steps, the steps of the methods herein may be performed simultaneously or at different times or in different orders. Furthermore, portions of such steps may be used by portions of other steps from other methods. Additionally, all or part of the steps may be selected as appropriate. Moreover, any step of any method may be performed using modules, units, circuits, or other components of a system used to perform such steps.

[0283] Specific details of a particular embodiment may be combined in any suitable manner without departing from the spirit and scope of the embodiments of the invention. However, other embodiments of the invention may be specific embodiments relating to each particular variant or a particular combination of such individual variants.

[0284] The foregoing description of exemplary embodiments of the present disclosure has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the present disclosure to the precise form described, and many modifications and variations are possible in light of the foregoing teachings.

[0285] Unless otherwise specified, the use of "a / an" or "the" means "one or more". Unless otherwise specified, the use of "or" means "inclusive or", not "mutually exclusive or". Referring to the "first" component does not necessarily require providing the second component. Furthermore, referring to the "first" or "second" component does not limit the mentioned component to a specific location unless explicitly stated otherwise. The term "based on" is intended to mean "at least partially based on".

[0286] For all purposes, all patents, patent applications, publications and descriptions mentioned herein are incorporated herein by reference in their entirety. None of them are recognized as prior art.

[0287] 10: Computer System 71: I / O Controller 72: System Memory 73: Central Processing Unit 74: Printer 75: System Bus 76: Monitor 77: I / O Ports 78: Keyboard 79: Storage device 81: External Interface 82: Monitor adapter 85: Data collection device 110: Free DNA fragment 120: Steps 130: Steps 140: Technology 141: Sequential Fragments 142: First terminal primitive 144: Second terminal primitive 145: Genome 160: Technology 161: Sequential Fragments 162: First terminal primitive 164: Second terminal primitive 165: Genome 205: Fetal-specific molecules 207: Shared molecule 220: Bar chart 222: Terminal primitive 230: Analysis Based on Entropy 235: Figure 240: Cluster-based Analysis 242: Red calibration sample 244: Green Calibration Sample 610: Shared 620: Fetal-specific 1205: Tumor-specific molecules 1207: Shared molecule 1220: Bar Chart 1222: Terminal primitive 1230: Analysis Based on Entropy 1235: Image 1240: Cluster-based Analysis 1410: Frequency value 1420: Frequency value 1700: 1705: Calibration Data Point 1710: Calibration Function 1900: Method 1910: Steps 1920: Steps 1930: Steps 1940: Steps 1950: Steps 2000: Method 2010: Steps 2020: Steps 2030: Steps 2040: Steps 2050: Steps 2070: Steps 2080: Steps 2105: Pathogen molecules 2107: Control molecule 2120: Bar Chart 2122: Terminal primitive 2130: Analysis Based on Entropy 2135: Figure 2140: Cluster-based Analysis 2801: MDS-based methods 2802:OCF 2803: Fragment Size 2804: Fragment Preference End 2805: Combinatorial Analysis 2901:1 Merger Analysis 2902: Dimer Analysis 2903: Trimer Analysis 2904: Tetramer Analysis 2905:5-mer analysis 3205: Early HCC 3210: Comparison 3215:HBV 3220: Cirrhosis 3225: Late-stage HCC 3230: Intermediate-term HCC 3305: HCC individual 3310: Non-HCC individuals 3405: ​​Early HCC 3410: Healthy control individuals 3415:HBV 3420: Cirrhosis 3425: Late-stage HCC 3430: Intermediate-term HCC 3505: Control / Normal 3510:SLE 3805: Non-HCC individuals 3810: HCC individual 4101:50-80 Analysis 4102:81-110 Analysis Analysis 4103:111-140 Analysis 4104:141-170 4105:171-200 Analysis Analysis of 4106:201-230 4200: Method 4210: Steps 4220: Steps 4230: Steps 4240: Steps 4250: Steps 4300: Method 4310: Steps 4320: Steps 4330: Steps 4400: Method 4410: Steps 4420: Steps 4430: Steps 4440: Steps 4450: Steps 4460: Steps 4600: Measurement System 4605: Sample 4608: Analysis 4610: Sample Holder 4615: Physical characteristics 4620: Detector 4625: Data Signal 4630: Logic System 4635: Memory 4640: External Memory 4645: Storage device 4650: Processor

Claims

1. A method for classifying the pathological grade of a biological sample of an individual, the biological sample containing cell-free DNA, the method comprising: The method involves receiving sequence reads from a plurality of cell-free DNA fragments of a biological sample, wherein the sequence reads contain terminal sequences corresponding to the ends of the plurality of cell-free DNA fragments; for each of the plurality of cell-free DNA fragments, determining a sequence motif for each of one or more terminal sequences of the cell-free DNA fragments; determining one or more relative frequencies of a set of one or more sequence motifs corresponding to the terminal sequences of the plurality of cell-free DNA fragments, wherein the relative frequencies of the sequence motifs provide the proportion of the plurality of cell-free DNA fragments having a terminal sequence corresponding to the sequence motif; and using the one or more relative frequencies and reference values ​​to determine the pathological grade classification of the individual.

2. The method of claim 1, wherein the classification is determined by using a machine learning model trained with reference samples of individuals having and not having the pathology.

3. The method of claim 1, wherein each of the plurality of free cellular DNA fragments is selected to have one of a plurality of specific size ranges.

4. The method of request item 1, further comprising: The cell-free DNA is filtered to identify the multiple cell-free DNA fragments.

5. The method of claim 4, wherein the filtering is based on the size of the free DNA fragment or the region of the derived DNA fragment.

6. The method of claim 5, wherein the free DNA is filtered for DNA fragments from open chromatin regions of a particular tissue.

7. The method of request item 1, wherein the pathology is cancer.

8. The method of claim 7, wherein the cancer is hepatocellular carcinoma, lung cancer, breast cancer, gastric cancer, glioblastoma multiforme, pancreatic cancer, colorectal cancer, nasopharyngeal carcinoma, or squamous cell carcinoma of the head and neck.

9. The method of claim 7, wherein the classification is determined from a plurality of cancer grades containing a plurality of cancer stages.

10. The method of claim 1, wherein the pathology is an autoimmune disease.

11. The method of claim 10, wherein the autoimmune disease is systemic lupus erythematosus.

12. The method of claim 1, wherein the pathological grade corresponds to the fractional concentration of clinically relevant DNA associated with the pathology.

13. A method for estimating the fractional concentration of clinically relevant DNA in a biological sample of an individual, the biological sample comprising the clinically relevant DNA and other cell-free DNA, the method comprising: The method involves receiving sequence reads from a plurality of cell-free DNA fragments of the biological sample, wherein the sequence reads contain terminal sequences corresponding to the ends of the plurality of cell-free DNA fragments; for each of the plurality of cell-free DNA fragments, determining a sequence motif for each of one or more terminal sequences of the cell-free DNA fragments; determining a relative frequency of a set of one or more sequence motifs corresponding to the terminal sequences of the plurality of cell-free DNA fragments, wherein the relative frequency of the sequence motifs provides a proportion of the plurality of cell-free DNA fragments having terminal sequences corresponding to the sequence motifs; and using the one or more relative frequencies and one or more calibration values ​​determined in a calibration sample of one or more known clinically relevant DNA fraction concentrations to determine the classification of the clinically relevant DNA fraction concentration in the biological sample.

14. The method of claim 13, wherein the clinically relevant DNA is selected from the group consisting of: fetal DNA, tumor DNA, DNA from transplanted organs, and specific tissue types.

15. The method of claim 13, wherein the clinically relevant DNA is a specific tissue type.

16. The method of claim 15, wherein the particular tissue type is liver or hematopoietic tissue.

17. The method of claim 13, wherein the individual is a pregnant woman and wherein the clinically relevant DNA is placental tissue.

18. The method of claim 13, wherein the clinically relevant DNA is tumor DNA derived from an organ with cancer.

19. The method of claim 13, wherein the one or more calibration values ​​are a plurality of calibration values ​​of a calibration function determined using the fractional concentrations of clinically relevant DNA from a plurality of calibration samples.

20. The method of claim 13, wherein the one or more calibration values ​​correspond to one or more total values ​​of relative frequencies of one or more sequence motifs, the one or more sequence motifs being measured using free DNA fragments in the one or more calibration samples.

21. The method of claim 13, further comprising: For each of the one or more calibration samples: the fractional concentration of clinically relevant DNA in the calibration sample is measured; and the total relative frequency of one or more sequence motifs is determined by analyzing cell-free DNA fragments from the calibration sample as part of obtaining calibration data points, thereby determining one or more total values, wherein each calibration data point specifies the measured fractional concentration of clinically relevant DNA in the calibration sample and the total value determined for the calibration sample, and wherein the one or more calibration values ​​are the one or more total values ​​or are determined using the one or more total values.

22. The method of claim 21, wherein the fractional concentration of clinically relevant DNA in the calibration sample is measured using a pair of genes specific to that clinically relevant DNA.

23. A method for determining the gestational age of a fetus by analyzing a biological sample from a pregnant woman, the biological sample comprising cell-free DNA molecules from the woman and the fetus, the method comprising: The method involves receiving sequence reads from a plurality of cell-free DNA fragments of the biological sample, wherein the sequence reads contain terminal sequences corresponding to the ends of the plurality of cell-free DNA fragments; for each of the plurality of cell-free DNA fragments, determining a sequence motif for each of one or more terminal sequences of the cell-free DNA fragments; determining one or more relative frequencies of a set of one or more sequence motifs corresponding to the terminal sequences of the plurality of cell-free DNA fragments, wherein the relative frequencies of the sequence motifs provide a proportion of the plurality of cell-free DNA fragments having a terminal sequence corresponding to the sequence motif; obtaining one or more calibration data points, wherein each calibration data point specifies the gestational age corresponding to the relative frequencies of the set of one or more sequence motifs, determined from one of a plurality of calibration samples having a known gestational age and containing cell-free DNA molecules; and estimating the gestational age of the fetus using the one or more relative frequencies and the one or more calibration data points.

24. The method of claim 23, wherein the one or more calibration data points are a plurality of calibration data points, which are formed by a calibration function approximating the relative frequencies determined from free DNA molecules in the plurality of calibration samples having a known gestational age.

25. The method of claim 23, wherein the method further comprises determining the total value of the one or more relative frequencies, and wherein the total value is compared with a plurality of calibration values, each of the plurality of calibration values ​​corresponding to one of the plurality of calibration samples.

26. The method of claim 23, wherein the calibration value of at least one calibration data point corresponds to the respective total value measured using the free DNA molecules of at least one of the plurality of calibration samples.

27. The method of claim 23, further comprising: The multiple cell-free DNA fragments were identified as originating from the fetus.

28. The method of claim 27, wherein the plurality of cell-free DNA fragments are identified using fetal-specific paired genes or fetal-specific epigenetic markers.

29. The method of claim 27, wherein the plurality of cell-free DNA fragments are identified by: for each of the sequence reads: determining the probability that the sequence read corresponds to the fetus based on the terminal sequence of the sequence read containing the set of one or more sequence motifs; comparing the probability with a threshold value; and identifying the sequence read as originating from the fetus when the probability exceeds the threshold value.

30. A method for enriching a biological sample with clinically relevant DNA, the biological sample comprising the clinically relevant DNA and other cell-free DNA, the method comprising: The method involves receiving analytical sequence reads from a plurality of cell-free DNA fragments of a biological sample, wherein the sequence reads contain terminal sequences corresponding to the ends of the plurality of cell-free DNA fragments; for each of the plurality of cell-free DNA fragments, identifying a sequence motif in one or more of the terminal sequences of the cell-free DNA fragments; identifying a set of sequence motifs that occur at a relatively higher frequency in clinically relevant DNA than in other DNAs; identifying a set of sequence reads that have the set of one or more sequence motifs in the terminal sequences, thereby obtaining stored sequence reads; and analyzing the stored sequence reads to determine the nature of the clinically relevant DNA in the biological sample.

31. The method of claim 30, wherein the nature of the clinically relevant DNA in the biological sample is: (1) the fractional concentration of the clinically relevant DNA; (2) the pathological grade of the individual from whom the biological sample was obtained, the pathological grade being related to the clinically relevant DNA; or (3) the gestational age of the fetus of the pregnant woman from whom the biological sample was obtained.

32. The method of claim 30, wherein the property is the genotype of the clinically relevant DNA.

33. The method of claim 30, further comprising: The size of the plurality of cell-free DNA fragments is measured using these sequence reads; and the stored sequence reads are filtered based on the size of the plurality of cell-free DNA fragments corresponding to the stored sequence reads.

34. The method of claim 30, further comprising: Measuring one or more methylation states at one or more sites of a cell-free DNA fragment corresponding to a specific sequence read; and filtering the specific sequence read from the stored sequence read based on the one or more methylation states.

35. The method of any one of claims 1 to 34, wherein determining the sequence motif of each of the one or more terminal sequences of the cell-free DNA fragment comprises: aligning the sequence reads corresponding to the one or more sequence reads of the cell-free DNA fragment with a reference genome; identifying one or more bases in the reference genome adjacent to the terminal sequence; and using the terminal sequence and the one or more bases to determine the sequence motif.

36. A computer-readable medium storing instructions that, when executed by a computer system, cause the computer system to perform any of the methods described in claims 1 to 35.

37. A system comprising: Computer-readable media; and one or more processors for executing instructions stored on the computer-readable medium as described in any of requests 1 to 35.

38. A system comprising components for performing the method as described in any one of claims 1 to 35.

39. A system comprising modules that respectively perform the steps of the method as described in any one of claims 1 to 35.

Citation Information

Patent Citations

  • Analysis of fragmentation patterns of cell-free DNA

    TW201718872A