Use of free DNA fragmentation pattern associated with epigenetic modification
By analyzing and comparing the nucleosome signal patterns around the target site with the reference pattern, the problem of insufficient accuracy in cancer detection in existing technologies is solved, enabling more precise determination of methylation levels and lesion grades, and improving the diversity and accuracy of detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CENT FOR NOVOSTICS
- Filing Date
- 2024-09-19
- Publication Date
- 2026-04-24
AI Technical Summary
Existing cell-free DNA analysis technologies lack sufficient accuracy in cancer detection, making it difficult to achieve efficient and diverse detection and measurement.
By analyzing the nucleosome signal patterns around the target site and comparing them with reference patterns, the methylation level, lesion grade, and DNA concentration ratio of the target CpG site in a specific tissue type can be determined. The detection and measurement are performed by comparing the nucleosome signal patterns with reference patterns.
It improves the accuracy and diversity of cancer detection, enabling more precise determination of methylation levels and lesion grades at target CpG sites, and measurement of DNA proportion concentrations in specific tissue types.
Smart Images

Figure CN121925480A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to, and is a PCT application for, the following: U.S. Patent Application No. 18 / 883,637, filed September 12, 2024, entitled "Uses of Cell-Free DNA Fragmentation Patterns Associated with Epigenetic Modifications"; U.S. Patent Application No. 63 / 539,980, filed September 22, 2023, entitled "Uses of Cell-Free DNA Fragmentation Patterns Associated with Epigenetic Modifications"; and U.S. Provisional Application No. 63 / 663,564, filed June 24, 2024, entitled "Uses of Cell-Free DNA Fragmentation Patterns Associated with Epigenetic Modifications". "EpigeneticModifications", the entire contents of which are incorporated herein by reference for all purposes. Background Technology
[0003] Cell-free DNA (cfDNA) analysis offers an attractive non-invasive method for detecting and monitoring diseases such as cancer. Nucleosome origins of plasma DNA are well-documented. For example, the dominant population of plasma DNA molecules shows a size of 166 bp, consistent with the size of nucleosome units (Lo et al. Sci Transl Med. 2010;2:61ra91). Straver et al. (Straver et al. Prenat Diagn. 2016:36(7):614-621) attempted to use plasma DNA by analyzing the frequency of cfDNA fragments whose ends were located 73 bp upstream and downstream of the nucleosome center inferred from aggregated maternal plasma DNA sequencing reads. Such values showed a low Pearson correlation of 0.654 with the proportion of fetal DNA.
[0004] Snyder et al. (Snyder et al. Cell. 2016;164:57-68) developed a metric called the window protection score, defined as the number of molecules that cross a specific genomic window minus those whose endpoints are within that window. The window protection score displays a wavy signal across the human genome. Snyder et al. correlated the spacing patterns (distance between peaks) inferred from cfDNA molecules ranging in size from 193 bp to 199 bp with RNA expression datasets. However, the correlation between the median nucleosome spacing in the transcriptome body and gene expression was only -0.17. Such low correlations cannot achieve practically meaningful accuracy for disease diagnosis. Furthermore, the discrepancy between the RNA profiles of a public dataset containing 76 cell lines and primary tissues and actual body organs in the test samples would further reduce the accuracy of cancer detection. Ulz et al. (Ulz et al. Nat Commun. 2019;10:4666) attempted to detect cancer and tumor subtypes using sequencing coverage around transcription factor binding sites.
[0005] Therefore, more accurate and diverse technologies are needed for this detection and measurement. Summary of the Invention
[0006] For various purposes, techniques have been provided for using fragmented nucleosome signaling patterns at locations surrounding a target site. For example, nucleosome signaling patterns can be used to determine the methylation level of a target site (e.g., a CpG site). The signal can be associated with nucleosome patterns of cfDNA molecules within a genomic region that exhibits differential methylation in a target tissue type (e.g., an organ) by having different methylation levels (or multiple levels, e.g., as patterns) relative to one or more other tissue types (e.g., blood cells). Nucleosome signaling patterns can be compared with one or more reference patterns having known methylation levels.
[0007] Another example method can determine the lesion grade in an object, for example, detecting a patient with a lesion (e.g., cancer); such examples can include determining the tissue type in which the lesion is present. Nucleosome signal patterns at one or more CpG sites can be compared to one or more reference patterns determined based on one or more training samples with known lesion grades. CpG sites may exhibit differential methylation relative to one or more other tissue types within the target tissue type.
[0008] Another example is determining the proportionate concentration of DNA from a specific tissue type, such as the proportionate concentration of DNA with differential methylation at CpG sites. In this example, nucleosome signaling patterns can be compared to one or more reference patterns determined based on one or more calibration samples having known proportionate concentrations of DNA from a specific tissue type.
[0009] One general aspect includes a method for measuring methylation of target CpG sites in the genome of an object using cell-free DNA molecules. The method may include analyzing multiple cell-free DNA molecules from a sample of the object organism, wherein analyzing each of the multiple cell-free DNA molecules includes identifying two genomic locations in a reference genome corresponding to the ends of the cell-free DNA molecules. The method may also include determining a nucleosome signaling pattern in a genomic region surrounding the target CpG site by: for each genomic location within the genomic region: determining a first amount of multiple cell-free DNA molecules spanning a window around the genomic location, said window being 2 bp or longer; determining a second amount of multiple cell-free DNA molecules ending within the window surrounding the genomic location; and determining a nucleosome signal at each location using the first and second amounts, wherein the genomic region is at least 140 bp in length. The method may also include determining the methylation level of the target CpG site in the genome of the object based on a comparison of the nucleosome signaling pattern with a reference pattern, wherein the reference pattern is determined based on one or more training samples with known methylation levels.
[0010] Another general aspect includes a method for analyzing biological samples to determine the cancer grade in a subject's biological sample. This method may include analyzing multiple cell-free DNA molecules from the subject's biological sample, wherein analyzing each of the multiple cell-free DNA molecules includes identifying two genomic locations in a reference genome corresponding to the ends of the cell-free DNA molecules. The method may also include determining a nucleosome signaling pattern in a genomic region surrounding a target CpG site by: for each genomic location within the genomic region: determining a first quantity of multiple cell-free DNA molecules spanning a window around the genomic location, said window being 2 bp or longer; determining a second quantity of multiple cell-free DNA molecules ending within the window surrounding the genomic location; and using the first and second quantities to determine a nucleosome signal at each location, wherein the target CpG site exhibits differential methylation relative to one or more other tissue types in the target tissue type. The method may also include determining a cancer grade classification of the subject based on a comparison of the nucleosome signaling pattern with a reference pattern, wherein the reference pattern is determined based on one or more training samples with known cancer grades, and wherein the cancer grade is determined relative to the target tissue type.
[0011] Another general aspect includes a method for measuring the proportionate concentration of DNA from a first tissue type in a target biological sample. This method may include analyzing a plurality of cell-free DNA molecules from the target biological sample, wherein analyzing each of the plurality of cell-free DNA molecules includes determining two genomic locations in a reference genome corresponding to the ends of the cell-free DNA molecules. The method may also include determining a nucleosome signaling pattern in a genomic region surrounding a target CpG site by the following steps: for each genomic location within the genomic region: determining a first amount of a plurality of cell-free DNA molecules spanning a window around the genomic location, said window being 2 bp or longer; determining a second amount of a plurality of cell-free DNA molecules ending within the window surrounding the genomic location; and determining a nucleosome signal at each location using the first and second amounts, wherein the target CpG site in the biological sample exhibits differential methylation relative to one or more other tissue types in the first tissue type. The method may also include determining the proportionate concentration of DNA from the first tissue type in the biological sample by comparing the nucleosome signaling pattern to a reference pattern, wherein the reference pattern is determined based on one or more calibration samples having known proportionate concentrations of DNA from the first tissue type.
[0012] These and other embodiments of this disclosure are described in detail below. For example, other embodiments relate to systems, apparatuses, and computer-readable media associated with the methods described herein.
[0013] A better understanding of the characteristics and advantages of the embodiments of this disclosure can be obtained by referring to the following detailed description and accompanying drawings.
[0014] Brief description of the attached figures
[0015] Figure 1 It is a chart illustrating the various molecular characteristics associated with the identification of cell-free DNA molecules.
[0016] Figure 2 An exemplary free DNA fragmentation pattern is shown regarding nucleosome signals around the genomic region of interest.
[0017] Figure 3 An example of nucleosome signaling at the CTCF-binding site is shown.
[0018] Figure 4A The differences in nucleosome signaling patterns at sites with different methylation levels were shown. Figure 4B Enhanced wavy nucleosome signals were observed around tissue-specific hypermethylated and hypomethylated CpG sites in HCC.
[0019] Figure 5A The percentages of the two types of DMS located in regions A and B of the genome are shown.
[0020] Figure 5B Receiver operating characteristic (ROC) analysis is shown to predict methylation status at sites in the training and test sets.
[0021] Figure 6 This is a flowchart illustrating a method for determining the methylation level at a site using nucleosome signaling patterns according to an embodiment of this disclosure.
[0022] Figure 7A and Figure 7B The mean fragmentation score (normalized nucleosome signal) of HCC-specific hypermethylated and hypomethylated sites is shown in healthy controls and patients with advanced HCC (aHCC).
[0023] Figure 8A The average peak-to-trough distance is shown for healthy controls and aHCC subjects. Figure 8B The average peak-to-trough distances for the following types of objects are shown: control, HBV, early HCC (eHCC), intermediate HCC (iHCC), and late HCC (aHCC). Figure 8C and 8D The study showed that the nucleosome amplitude at HCC-specific hypomethylated CpG sites was also reduced in HCC patients.
[0024] Figures 9A to 9B and Figure 10 An exemplary schematic diagram is shown for standardizing the nucleosome score around the CpG site.
[0025] Figures 11A to 11C A set of graphs is shown to evaluate the performance of using cfDNA nucleosome signaling patterns as biomarkers for cancer diagnosis.
[0026] Figures 12A to 12B The cancer detection method using optional standardized nucleosome signals is shown.
[0027] Figure 13A An SVM model trained on dataset A is provided to distinguish the HCC probability of different grades of HCC patients from all non-HCC subjects in dataset B. Figure 13B The ROC curves used to distinguish HCC patients from all non-HCC subjects are shown.
[0028] Figure 14A An SVM model trained on a dataset using one sequencing protocol is provided to distinguish the HCC probabilities of various grades of HCC patients from all non-HCC subjects on datasets using different sequencing protocols. Figure 14B The ROC curves used to distinguish HCC patients from all non-HCC subjects are shown.
[0029] Figures 15A to 15BThe results show the cancer detection (HCC probability) in dataset B based on an SVM model trained on dataset A using nucleosome signals.
[0030] Figure 16 Displayed based on Figure 15A The graph shown is a Kaplan-Meil analysis of the predicted HCC probability against the survival of HBV subjects classified into two risk groups.
[0031] Figures 17A to 17B A set of graphs is shown to evaluate the performance of nucleic acid body scores in dataset C, generated using nanopore sequencing, for cancer diagnosis.
[0032] Figures 18A to 18C A set of ROC curves (using cfDNA data from dataset B as an example) is shown to evaluate the effect of data normalization on nucleosome scoring patterns and cancer detection.
[0033] Figures 19A to 19C ROC curves for HCC detection using nucleosome signaling patterns based on differentially methylated sites (DMS), defined by different criteria for methylation levels between the leukocyte layer and tumor tissue, are shown.
[0034] Figures 20A to 20B A set of graphs shows the performance of nucleosome scoring analysis for evaluating random CpG sites.
[0035] Figure 21 This shows the nucleosome signaling pattern (FRAGMA) relative to other techniques in dataset A. XR ) performance.
[0036] Figure 22A and 22B The results of cancer detection for lung cancer, breast cancer, and ovarian cancer using nucleosome signaling patterns are shown.
[0037] Figures 23A to 23C It displays a probability score of having cancer predicted based on the proportion of tumor DNA in the plasma of cancer patients.
[0038] Figure 24 This is a flowchart illustrating a method for analyzing biological samples to determine the cancer grade in a target biological sample (2000).
[0039] Figure 25A and Figure 25B Nucleosome signaling patterns at placental-specific CpG sites with high and low methylation were shown.
[0040] Figure 26A The average peak-to-valley distance of hypermethylated sites with different proportions of fetal DNA was shown. Figure 26BThe average peak-to-valley distance of hypomethylation sites is shown for different proportions of fetal DNA. Figures 26C to 26D The study showed a good correlation between the proportion of fetal DNA predicted by nucleosome patterns and the proportion of actual fetal DNA inferred by single nucleotide polymorphism (SNP)-based methods.
[0041] Figure 27A Nucleosome signaling patterns at placental-specific CpG sites with low and high methylation were shown. Figures 27B to 27C This shows a comparison between the actual and predicted proportions of fetal DNA in the training and test sets. The actual proportions were inferred from single nucleotide polymorphism (SNP) analysis.
[0042] Figures 28A to 28B The results showed a good correlation between the proportion of donor-origin DNA predicted by nucleosome patterns and the proportion of actual donor-origin DNA inferred by single nucleotide polymorphism analysis, with Pearson correlations of 0.99 and 0.97 for the training and test sets, respectively.
[0043] Figures 29A to 29B The results showed a good correlation between the proportion of tumor-derived DNA predicted by nucleosome patterns and the proportion of tumor-derived DNA inferred by ichorCNA, with Pearson correlations of 0.96 and 0.94 for the training and test sets, respectively.
[0044] Figure 30 This is a flowchart illustrating a method for determining the proportionate concentration of DNA from a first tissue type in a biological sample.
[0045] Figure 31 A measurement system according to an embodiment of the present disclosure is shown.
[0046] Figure 32 A block diagram of an example computer system that can be used with systems and methods according to embodiments of this disclosure is shown.
[0047] the term
[0048] “ organize "This corresponds to a group of cells that function together as a single unit. More than one type of cell can be found in a single tissue. Different types of tissues can be composed of different types of cells (e.g., hepatocytes, alveolar cells, or blood cells), but can also correspond to tissues from different organisms (mother and fetus) or to non-tumor cells and tumor cells. A 'reference tissue' can correspond to the tissue used to determine tissue-specific methylation levels. Multiple samples of the same tissue type from different individuals can be used to determine the tissue-specific methylation levels for that tissue type."
[0049] the term" biological samples"A biological sample" refers to any sample obtained from an object (such as a human or other animal, such as a pregnant woman, an individual with cancer or other medical condition, or an individual suspected of having cancer or other medical condition, an organ transplant recipient, or an object suspected of having a disease process involving an organ (such as the heart in a myocardial infarction, or the brain in a stroke, or the hematopoietic system in anemia)) and containing one or more nucleic acid molecules of interest (e.g., DNA and / or RNA). Biological samples can be bodily fluids, such as blood, plasma, serum, urine, vaginal fluid, fluid from scrotal edema (e.g., testes), vaginal douches, pleural effusion, ascites, cerebrospinal fluid, saliva, sweat, etc. Fluids, tears, sputum, bronchoalveolar lavage fluid, breast milk, aspirated fluid from different parts of the body (e.g., thyroid, breast), intraocular fluid (e.g., aqueous humor), amniotic fluid, etc., can also be used. Fecal samples may also be used. In various embodiments, the majority of DNA in the biological sample (e.g., a biological sample enriched with cell-free DNA, such as a plasma sample obtained via centrifugation) can be cell-free, for example, more than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA can be cell-free. Centrifugation protocols for enriching cell-free DNA from biological samples may include, for example, centrifuging 1,600 g of the biological sample. The liquid fraction of the centrifuged sample is obtained after 10 minutes, followed by a further centrifugation at, for example, 16,000 g for 10 minutes to remove remaining cells. As part of the analysis of the biological sample, a statistically significant number of cell-free DNA molecules can be analyzed (e.g., to provide accurate measurement results). In some embodiments, at least 1,000 cell-free DNA molecules are analyzed. In other embodiments, at least 10,000, 50,000, 100,000, 500,000, 1,000,000, or 5,000,000 or more cell-free DNA molecules can be analyzed. At least the same number of sequence reads can be analyzed. Any quantity described herein can be any of the numbers listed above. Example sample sizes can include 30, 50, 100, 200, 300, 500, 1,000, 5,000, or 10,000 or more nanograms, or 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 ml.
[0050] the term" Comparison "", control sample "", Background Sample "", refer to "", Reference Sample "", normal "and" normal samples"This term can be used interchangeably to generally describe samples that do not have a specific disease condition or are otherwise healthy. In one example, a template-free control (NTC) sample with contaminating DNA can be considered a reference sample. In another example, a reference sample is a sample extracted from an uninfected object. Reference samples can be obtained from the object or from a database. A reference generally refers to a reference genome, which is used to map sequence readings obtained from sequencing samples from an object. A reference genome generally refers to a haploid or diploid genome that can be aligned and compared with sequence readings from a biological sample. For a haploid genome, there is only one nucleotide at each locus. For a diploid genome, heterozygous loci can be identified, such loci having two alleles, either of which allows for an alignment match with that locus. A reference genome (e.g., by including one or more microbial genomes) can be a reference microbial genome corresponding to a specific microbial species."
[0051] “ Reference genome The "reference sequence" or "reference sequence" can be the entire genome sequence of a reference organism, one or more segments of the reference genome that may be continuous or non-contiguous, a common sequence of multiple reference organisms, a compiled sequence based on different components of different organisms, or any other suitable reference sequence. As an example, the reference genome / sequence length can be at least 1,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 5,000,000, 10,000,000, 50,000,000, 100,000,000, 500,000,000, one billion, or three billion nucleotides, such as the complete human genome or a duplicated human genome. The reference may also include information about reference variations known to be found in a population of organisms.
[0052] “ Clinically relevant DNA "Clinically relevant DNA" refers to DNA from a specific tissue source being measured, for example, to determine the proportional concentration of such DNA or to classify the phenotype of a sample (e.g., plasma). Examples of clinically relevant DNA include fetal DNA in maternal plasma or tumor DNA in patient plasma or other samples containing cell-free DNA. Another example includes measuring the amount of graft-associated DNA in the plasma, serum, or urine of a transplant patient. Further examples include measuring the proportional concentration of hematopoietic and non-hematopoietic DNA in the recipient's plasma, or the proportional concentration of liver DNA fragments (or other tissues) in a sample, or the proportional concentration of brain DNA fragments in cerebrospinal fluid.
[0053] the term" Fetal DNA ratio concentration "Can be used with terminology" fetal DNA proportion )"and" fetal fetal DNA fractionThe term "fetal DNA" is used interchangeably and refers to the proportion of fetal DNA molecules present in biological samples (e.g., maternal plasma or serum samples) derived from the fetus (Lo et al.). Am J Hum Genet. 1998;62:768-775; Lun et al, Clin Chem. (2008; 54:1664-1672). Similarly, tumor proportion or tumor DNA proportion can refer to the proportional concentration of tumor DNA in a biological sample.
[0054] As used herein, the term "fragment" (e.g., DNA or RNA fragment) may refer to a portion of at least three consecutive nucleotides comprising a polynucleotide or polypeptide sequence. Nucleic acid fragments may retain the biological activity and / or some characteristics of the parent polypeptide. Nucleic acid fragments may be double-stranded or single-stranded, methylated or unmethylated, intact or nicked, and complexed or uncomplexed with other macromolecules (e.g., lipid particles, proteins). Nucleic acid fragments may be linear or circular. Tumor-derived nucleic acids may refer to any nucleic acid released from tumor cells, including pathogen nucleic acids from pathogens within tumor cells. As part of the analysis of a biological sample, a statistically significant number of fragments may be analyzed; for example, at least 1,000 fragments may be analyzed. As other examples, at least 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, or 5,000,000 or more fragments may be analyzed, and such fragments may be selected randomly or according to one or more criteria.
[0055] The term "assay" generally refers to a technique used to determine the characteristics of nucleic acids or a sample of nucleic acids (e.g., a statistically significant number of nucleic acids) and the characteristics of the object from which the sample was obtained. An assay (e.g., a first assay or a second assay) generally refers to a technique used to determine the amount of nucleic acids in a sample, the genomic identity of the nucleic acids in the sample, the copy number variation of the nucleic acids in the sample, the methylation state of the nucleic acids in the sample, the fragment size distribution of the nucleic acids in the sample, the mutation state of the nucleic acids in the sample, or the fragmentation pattern of the nucleic acids in the sample. Any assay known to those skilled in the art can be used to detect any characteristic of the nucleic acids mentioned herein. Characteristics of nucleic acids include sequence, amount, genomic identity, copy number, methylation state at one or more nucleotide positions, size of the nucleic acid, mutations in the nucleic acid at one or more nucleotide positions, and fragmentation pattern of the nucleic acid (e.g., nucleotide position at a nucleic acid fragment). The term "assay" may be used interchangeably with the term "method." Assays or methods may (e.g., based on the selection of one or more cutoff values) have specific sensitivity and / or specificity, and their relative usefulness as diagnostic tools can be measured using the area under the receiver operating characteristic (ROC) curve (AUC).
[0056] the term" Sequence readings "Sequence reads" refers to a sequence of nucleotides obtained from any part or all of a nucleic acid molecule. For example, a sequence read can be a short nucleotide sequence (e.g., 20 to 150 nucleotides) sequenced from a nucleic acid fragment, a short nucleotide sequence at one or both ends of a nucleic acid fragment, or the sequencing of an entire nucleic acid fragment present in a biological sample. Sequence reads can be obtained in various ways, such as using sequencing technologies or using probes (e.g., in hybridization arrays or capture probes that can be used in microarrays) or amplification technologies (e.g., polymerase chain reaction (PCR) or linear or isothermal amplification using a single primer). Example sequencing technologies include high-throughput parallel sequencing, targeted sequencing, Sanger sequencing, sequencing by ligation, ion semiconductor sequencing, and single-molecule sequencing (e.g., using nanopores), or single-molecule real-time sequencing (e.g., from Pacific...). Biosciences). This type of sequencing can be random sequencing or targeted sequencing (e.g., enriching such regions by using capture probes that hybridize with a specified region or by amplifying certain regions). Example probe-based techniques include real-time PCR and digital PCR (e.g., droplet digital PCR). As part of the analysis of a biological sample, a statistically significant number of sequence reads can be analyzed, for example, at least 1,000 sequence reads can be analyzed. As other examples, at least 5... 000, 10,000, 50,000, 100,000, 500,000, 1,000,000, or 5,000,000 or more sequence readings. Furthermore, the number of sequence readings determined for embodiments of this disclosure may be at least 1,000, 5,000, 10,000, 50,000, 100,000, 100,000, 500,000, 1,000,000, or 5,000,000.
[0057] the term" Mapping "or" Alignment "Associating a sequence with a location or coordinates (e.g., genomic coordinates) in a reference (e.g., a reference genome) having a known reference sequence, wherein the sequence is similar to the known reference sequence at that location in the reference. This can be done according to..." Mapping quality This is used to measure or report similarity. In one example of mapping quality used in this paper, the mapping quality X of a sequence relative to the reported position or coordinates in the reference indicates that the probability of the sequence mapping to a different position is no greater than 10^(-X / 10). For example, a mapping quality of 30 indicates that the probability of the sequence mapping to another position is less than 0.1%.
[0058] “ site "(also known as " Genomic sitesA "locus" corresponds to a single site, which can be a single base position or a group of related base positions, such as a CpG site, a TSS site, a DNase hypersensitive site, or a larger group of related base positions. A "locus" can correspond to a region that includes multiple sites. A locus can include only one site, which would make the locus equivalent to the site in the context. A region surrounding the site can be defined, such as a symmetrical or asymmetrical region around the site. As an example, a region can be included within the site. At least + / - 50 bases before and after (e.g., 101 bases), + / - 60 bases, + / - 70 bases, + / - 80 bases, + / - 90 bases, + / - 100 bases, + / - 150 bases, + / - 200 bases, + / - 300 bases, + / - 400 bases. + / - 500 bases, + / - 600 bases, + / - 700 bases, + / - 800 bases, + / - 900 bases. The number of bases and + / - 1,000 bases. As other examples, the region can be at least 100 bases, 140 bases, 147 bases, or 167 bases long. One or more regions can be analyzed, for example, to provide a lesion (e.g., cancer) grade or a proportion of a specific tissue. Various numbers of regions, sites, or loci can be analyzed, such as 50, 100, 200, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, one million, or more. Various techniques, such as aligning sequence reads to a reference genome or using position-specific probes, can determine one or more genomic locations of the DNA molecule within the reference genome. Position determination can be performed on some or all of the reference genome, for example, if only a portion of the genome is analyzed. As an example, the amount of genome analyzed can be greater than 0.01%, 0.1%, 1%, 5%, 10%, or 50%.
[0059] Nucleosome (nucleosome) signaling patterns may include values at each genomic location within a region. Values at genomic locations may be measures of the properties of cell-free DNA molecules within a window surrounding the genomic location. Exemplary window sizes are 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 90 bp, 100 bp, 110 bp, 120 bp, 130 bp, 140 bp, 150 bp, 160 bp, 170 bp, 180 bp, 190 bp, 200 bp, etc., or any size greater than or smaller than these. A window may be a specified value equal to or smaller than any of these aforementioned numbers. Values at genomic locations (e.g., nucleosome signals at individual locations) may depend on a first amount of cell-free DNA molecules ending within the window and / or a second amount of cell-free DNA molecules spanning the window. The sum of the two quantities is the total number of DNA molecules covering the genomic location. As an example, this value can be a relative quantity of any of these quantities (e.g., a separation value), such as the ratio of any two of the first quantity, the second quantity, and the sum of the first and second quantities. No such standardized nucleosome signal can be referred to as the primitive nucleosome signal or primitive nucleosome signal pattern for a region. Nucleosome signals can be determined using all cfDNA fragments in a plasma sample or cfDNA fragments of a specific size (e.g., 120 bp to 180 bp, where 120 is the lower boundary and 180 is the upper boundary). Other example lower boundaries for the size range are 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 110 bp, 120 bp, 130 bp, and 140 bp. Other examples of upper limits for the size range are 150 bp, 160 bp, 170 bp, 180 bp, 190 bp, 200 bp, 210 bp, 220 bp, 230 bp, and 240 bp.
[0060] "Nucleosome scoring" can refer to a standardized nucleosome signal obtained using one or more nucleosome signals from one or more regions (e.g., the same region, flanking regions, or one or more distant regions at the site of interest, including regions on different chromosomes). For example, statistical values (e.g., average, mean, median, etc.) of these nucleosome signals can be taken and applied to a nucleosome signal pattern (e.g., each value in the pattern is divided or subtracted from the statistical value, or a combination of these operations). The result can be a "nucleosome scoring pattern" that includes nucleosome scores at a set of locations within the region.
[0061] A "standardized nucleosome score" (also known as a background-adjusted nucleosome score or simply a background nucleosome score) can refer to a nucleosome score standardized using background (baseline) signals from one or more other regions of the same sample. A given nucleosome score (or nucleosome value, if using the original signal) at a given location can be standardized using a corresponding nucleosome score determined based on one or more other regions (i.e., as the same location relative to a target site, such as a CpG site). Therefore, this standardization can be a standardization at each location. A standardized nucleosome signal can refer to this type of standardization of the original nucleosome signal.
[0062] "Distribution-adjusted nucleosome score" can refer to a nucleosome score adjusted for the mean and / or variance (e.g., standard deviation) of the distribution of a reference nucleosome signal from the same region or other regions, which can serve as a baseline control (e.g., a healthy sample). An example is a z-score or t-score. This adjustment (standardization) can be performed on a per-location basis. Thus, the mean and variance of a particular location can be determined (using values from other regions), and these mean and variance are used to determine a distribution-adjusted value for a given location in the current sample and region. This adjustment can be applied to any nucleosome value described herein (e.g., the original signal, nucleosome score, and background-adjusted value). Exemplary distributions defining how to use the mean and variance include the normal distribution, t-distribution, Poisson distribution, gamma distribution, binomial distribution, beta distribution, and Cauchy distribution. For models that well generalize the same population (e.g., models trained on dataset A to predict dataset B), standardized nucleosome scores and distribution-adjusted nucleosome scores can provide improved accuracy. In one example, nucleosome scores from clustered regions associated with hypomethylated and hypermethylated CpG sites can undergo two data normalization steps. Randomly selected genomic regions (200,000) are each centered around a CpG site. Sequencing cfDNA molecules from these genomic regions can be clustered together to determine background nucleosome scores. The background nucleosome scores for hypomethylated and hypermethylated CpG sites are divided by the background nucleosome score relative to the central CpG, termed the normalized nucleosome score. Next, based on a set of healthy subjects, the mean (µ) and standard deviation (δ) of the normalized nucleosome score for each location relative to the central CpG can be determined. The normalized nucleosome score (S) can be converted to a nucleosome z-score as follows: .
[0063] For a dataset, half of the healthy objects can be used to determine the µ and δ values. Nucleosomes adjusted for other distributions can be used to determine statistics other than the mean (e.g., other aggregate statistics such as the median or mode) and standard deviation (e.g., other dispersion values).
[0064] Any one of the above nucleosome values can be collectively referred to as the nucleosome signal that contains a nucleosome signal pattern for the region.
[0065] "in the mammalian genome" DNA methylation "This typically refers to the addition of a methyl group to the 5' carbon of a cytosine residue (i.e., 5-methylcytosine) in a CpG dinucleotide. DNA methylation can occur in cytosine in other backgrounds, such as CHG and CHH, where H is adenine, cytosine, or thymine. Cytosine methylation can also be in the form of 5-hydroxymethylcytosine. Non-cytosine methylation, such as N6-methyladenine, has also been reported."
[0066] "Each genomic locus (e.g., CpG site) " Methylation index "This can refer to the proportion of the number of DNA fragments showing methylation at that site (e.g., as determined by sequence readings or probes) to the total number of readings covering that site." First based state "Reading" can refer to whether a specific site is methylated at a specific location on a specific DNA fragment. "Reading" can correspond to information obtained from a DNA fragment (e.g., the methylation state at the site). Readings can be obtained using reagents (e.g., primers or probes) that preferentially hybridize with DNA fragments of a specific methylation state. Typically, such reagents are applied after treatment with methods that differentially modify or recognize DNA molecules based on their methylation state, such as bisulfite conversion, methylation-sensitive restriction enzymes, methylation-binding proteins, anti-methylcytosine antibodies, or single-molecule sequencing techniques that recognize methylcytosine and hydroxymethylcytosine. Some embodiments can determine the methylation level that has not undergone such treatment.
[0067] "region or set of sites" Methylation density"CpG methylation density" can refer to the number of readings showing CpG methylation within a region (also called a bin) or group of sites divided by the total number of readings of sites covering that region or group of sites. A region can include one or more sites of interest, including at least 1, 2, 3, 4, 5, 10, 20, 50, 100, 200, 500, and 1,000 sites. Sites can have specific characteristics, such as CpG sites. Therefore, the "CpG methylation density" of a region can refer to the number of readings showing CpG methylation divided by the total number of readings of CpG sites (e.g., specific CpG sites, CpG islands, or CpG sites within a larger region) covering the region. For example, the methylation density per 100 kb bin in the human genome can be determined based on the proportion of the total number of unconverted cytosines (corresponding to methylated cytosines) at CpG sites after bisulfite treatment relative to all CpG sites covered by the sequence readings of a 100 kb region. This analysis can also be performed on other bin sizes, such as 500 kb. bp, 5 kb, 10 kb, 50 kb, or 1 Mb, etc. A region can be an entire genome or a chromosome or a portion of a chromosome (e.g., a chromosome arm). When a region includes only CpG sites, the methylation index of the CpG sites is the same as the methylation density of the region. "The proportion of methylated cytosine" can refer to the number of cytosine sites "C" that are methylated (e.g., unconverted after bisulfite conversion) relative to the total number of cytosine bases analyzed, i.e., including cytosine in the region excluding the CpG background. Methylation index, methylation density, and the proportion of methylated cytosine are examples of "methylation level." Other methods known to those skilled in the art, besides bisulfite conversion, can be used to detect the methylation state of DNA molecules, including (but not limited to) enzymes sensitive to methylation state (e.g., methylation-sensitive restriction enzymes), methylation-binding proteins, and single-molecule sequencing using platforms sensitive to methylation state (e.g., nanopore sequencing). (Schreiber et al. Proc Natl Acad Sci USA 2013; 110:) (18910-18915) or by single-molecule real-time sequencing analysis by Pacific Biosciences (Tse et al. Proc Natl Acad Sci USA 2021; 118: e2019768118).
[0068] “ methylation level"Methylation" is an example of relative abundance, such as the relative abundance of methylated DNA molecules (e.g., at one or more specific sites) compared to other DNA molecules (e.g., all other DNA molecules or DNA molecules that are unmethylated only at one or more specific sites). The amount of other DNA molecules can act as a normalization factor. As another example, the intensity (e.g., fluorescence or electrical intensity) of methylated DNA molecules can be determined relative to the intensity of all DNA molecules or unmethylated DNA molecules. Relative abundance may also include intensity per volume. Methylation levels can be determined using methylation-sensing analysis, such as methylation-sensing sequencing or PCR. Exemplary methylation-sensing sequencing may include bisulfite sequencing or, for example, single-molecule techniques using nanopores.
[0069] Differentially methylated regions (DMRs) are genomic regions (e.g., loci sets) spanning two or more biological samples and exhibiting different levels of DNA methylation. These different DNA methylation levels can be defined by specific differences in methylation index or density, such as (but not limited to) 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, etc. Differentially methylated sites (DMSs) can be defined in a similar manner.
[0070] the term" hypomethylation "Can refer to a site or set of sites (e.g., a region) having a methylation level below a specified threshold, such as 50%, 45%, 40%, 35%, 30%, 25%, or 20% of the methylation level. If the methylation level is below the threshold, the site in the genome can be considered unmethylated. The term " Hypermethylation "Can refer to a specified value above the methylation level, such as a site or set of sites (e.g., a region) that is equal to or higher than 95%, 90%, 80%, 75%, 70%, 65%, or 60% of the methylation level. If the methylation level is greater than the threshold, the site in the genome can be considered methylated."
[0071] “ Calibration Samples "A calibration sample may correspond to a biological sample whose proportion of clinically relevant DNA (e.g., the proportion of tissue-specific DNA) is known or determined by calibration methods, such as using tissue-specific alleles, as in transplantation, where alleles present in the donor genome but lacking in the recipient genome can be used as markers for the transplanted organ. As another example, a calibration sample may correspond to a sample from which terminal motifs can be determined. Calibration samples can be used for both purposes."
[0072] “ Calibration data points "include" Calibration valueThe calibration data points are the measured or known proportions of clinically relevant DNA (e.g., DNA from a specific tissue type). Calibration values can be determined based on a reference pattern, such as a peak-to-trough distance, as determined for a calibration sample where the proportions of clinically relevant DNA are known. Calibration data points can be defined in various ways, such as as discrete points or as a calibration function (also known as a calibration curve or calibration surface). The calibration function can be derived from additional mathematical transformations of the calibration data points. Calibration values can be part of a calibration reference pattern, such as a nucleosome reference pattern determined by one or more calibration samples known to have similar proportions. Proportional concentrations can be determined in various ways, such as using tissue-specific alleles, tissue-specific methylation values or patterns, and the size distribution of samples with known proportions.
[0073] As the term "as used in this article" Classification A plus sign ("+") refers to any number or other character associated with a specific characteristic of a sample. For example, the symbol "+" (or the word "positive") can indicate that a sample is classified as having a missing or amplified portion. Classification can be binary (e.g., positive or negative) or have more levels of classification (e.g., scales from 1 to 10 or 0 to 1), including probabilities. Different techniques used to determine the classification can be combined, for example by majority voting or requiring all initial / intermediate classifications to be the same (e.g., positive), to obtain the final classification from the initial or intermediate classifications used for each of the different techniques.
[0074] As used herein, the term "parameter" refers to a numerical value that characterizes a quantitative dataset and / or the numerical relationship between quantitative datasets. For example, the ratio (or a function of the ratio) between a first quantity of a first nucleic acid sequence and a second quantity of a second nucleic acid sequence is a parameter. Parameters can be used to determine any classification described herein, such as any classification concerning fetal, cancer, or transplant analysis.
[0075] “ Separation value "Separation value corresponds to a difference or ratio involving two values, such as the difference or ratio of two proportions or two methylation levels. A separation value is an example of a parameter. A separation value can be a simple difference or ratio. As an example, the proportionality of x / y and the proportionality of x / (x+y) are separation values. Other examples are y / x and y / (x + y). Separation values can include other factors, such as a multiplication constant. As another example, a difference or ratio can be a function of the following values, such as the difference or ratio of the natural logarithms (ln) of two values. Separation values can include both differences and ratios. Separation values can be compared to a threshold to determine whether the separation between two values is statistically significant. A separation value is an example of a relative quantity. Separation values can be compared to a threshold to determine whether the separation between two values is statistically significant. A nucleosome signal at a specific location can be a separation value."
[0076] the term" Cutoff value "and" threshold "Cutoff value" refers to a predetermined numerical value used in the operation. For example, a cutoff size can refer to a size that fragments exceeding which will be excluded. A threshold can be a value that applies to a specific classification, whether it is above or below that value. Any of these terms can be used in any of these situations. A cutoff value or threshold can be a "reference value" or derived from a reference value that represents a specific classification or distinguishes two or more classifications. A cutoff value can be predetermined with or without reference to the characteristics of the sample or object. For example, a cutoff value can be selected based on the age or sex of the test subject. A cutoff value can be selected after the output of the test data and based on the output of the test data. For example, a specific cutoff value can be used when the sequencing of a sample reaches a certain depth. As another example, a characteristic value with known classifications and measurements of one or more diseases (e.g., methylation level, etc.) A reference object (statistical size value or count) can be used to determine a reference level to distinguish different symptoms and / or symptom categories (e.g., whether an object has a symptom). A reference value can be chosen as a representative of a category (e.g., the mean) or a value between two clusters of the measure (e.g., chosen to obtain desired sensitivity and specificity). As another example, a reference value can be determined based on a statistical simulation of a sample. Any of these terms can be used in either of these contexts. Such reference values can be determined in various ways, as a person skilled in the art will understand. For example, a measure can be determined for two different groups of objects with different known classifications, and this can be chosen. As another example, a reference value can be determined based on a statistical simulation of a sample. Specific cutoff values, thresholds, reference values, etc., can be determined based on the desired accuracy (e.g., sensitivity and specificity).
[0077] the term" Cancer grading"Cancer grade can refer to the presence or absence of cancer (i.e., presence or absence), stage of cancer, size of the tumor, presence or absence of metastasis, total tumor burden in the body, response to treatment, and / or other measures of cancer severity (e.g., cancer recurrence). Cancer grade can be a number or other marker, such as symbols, letters, and colors. The grade can be zero. Cancer grade can also include the condition (state) before it worsened or became cancerous. Cancer grade can be used in a variety of ways. For example, screening can check for the presence of cancer in someone who previously had no known cancer. Assessment can investigate someone who has been diagnosed with cancer to monitor cancer progression over time, investigate the effectiveness of treatments, or determine prognosis. In one implementation, prognosis can be..." This indicates the probability that a patient will die from cancer, or the probability that cancer will progress after a specified period or time, or the probability or extent of cancer metastasis. Testing can mean 'screening' or it can mean checking whether someone has cancer-like characteristics (such as symptoms or other positive tests). Various types of cancer can be graded, such as carcinoma or sarcoma, melanoma, lymphoma, and leukemia, as well as cancers from various tissues, including, for example: breast, lung, liver, colon, pancreas, stomach, bone, blood, head and neck (e.g., squamous cell carcinoma of the head and neck), throat, bladder, kidney, prostate, uterus, rectum, bile duct, brain, eye, esophagus, ovary, oral cavity, nasopharynx, thyroid, urethra, testis, vagina, and pituitary gland.
[0078] “ Lesion grade "A disease can refer to the quantity, extent, or severity of a lesion associated with an organism, where the grade can be as described above for cancer. Another example of a lesion is rejection of a transplanted organ. Other examples of lesions can include autoimmune attacks (e.g., lupus nephritis that damages the kidneys or multiple sclerosis that damages the central nervous system), inflammatory diseases (e.g., hepatitis), fibrotic processes (e.g., cirrhosis), fatty infiltration (e.g., fatty liver disease), degenerative processes (e.g., Alzheimer's disease), and ischemic tissue damage (e.g., myocardial infarction or stroke). Pregnancy can be considered a lesion. The subject's health status can be considered a classification of being free of lesions."
[0079] “ Machine learning models"(ML model) can refer to a software module configured to run on one or more processors to provide categorical or numerical characteristics of one or more samples. An ML model can include various parameters (e.g., functions for coefficients, weights, thresholds, functional attributes, such as activation functions). As an example, an ML model can include at least 10, 100, 1,000, 5,000, 10,000, 50,000, 100,000, or one million parameters. Sample data (e.g., training samples) can be used to generate an ML model to make predictions on test data. Various numbers of..." Training samples, for example, at least 10, 100, 1,000, 5,000, 10,000, 50,000, 100,000, or at least 200,000 training samples. One example is an unsupervised learning model. Another example type of model is supervised learning that can be used with embodiments of this disclosure. Example supervised learning models can include different methods and algorithms, including analytical learning, statistical models, artificial neural networks (e.g., including convolutional layers and / or variable layers), boosting (a common algorithm), Bayesian statistics, etc. Statistics, case-based reasoning, decision tree learning, inductive logic programming, Gaussian process regression, genetic programming, data grouping, kernel estimator, learning automata, learning classification systems, minimum information length (decision trees, decision graphs, etc.), multilinear subspace learning, naive Bayes classifier, maximum entropy classifier, conditional random fields, nearest neighbor algorithm, probabilistic approximate correct learning (PAC), ripple down rule, knowledge acquisition methods, symbolic machine learning algorithms, sub-symbolic machine learning algorithms, minimum complexity machine (MCM), random forest, ordered classification, data preprocessing, handling imbalanced datasets, statistical relation learning, or Proaftn (a multi-criteria classification algorithm), or a combination of these types. Models may include linear regression, logistic regression, deep recurrent neural networks (e.g., Long Short-Term Memory, LSTM), hidden Markov models, etc. Supervised learning models can be trained using various methods, including Hidden Models (HMM), Linear Discriminant Analysis (LDA), k-means clustering, density-based spatial clustering with noise (DBSCAN), random forest algorithms, support vector machines (SVM), or any model described herein. Various cost / loss functions and optimization techniques can be used to define the error from known labels (e.g., least squares and absolute difference from the known classification), such as backpropagation, steepest descent, conjugate gradients, and Newton and quasi-Newton techniques.
[0080] The terms "about" or "approximately" may mean within an acceptable range of error for a particular value as determined by those skilled in the art, the range of error depending in part on the manner in which the value is measured or determined, i.e., limitations of the measurement system. For example, according to practice in the art, "about" may mean within one or more standard deviations. Alternatively, "about" may mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or methods, the terms "about" or "approximately" may mean within a certain order of magnitude of the value, within 5 times, or more preferably within 2 times. If a particular value is described in this application and claims, unless otherwise stated, it should be assumed that the term "about" means within an acceptable range of error for the particular value. The term "about" may have the meaning as commonly understood by those skilled in the art. The term "about" may mean ±10%. The term "about" may mean ±5%.
[0081] Where a range of values is provided, it should be understood that, unless the context explicitly specifies otherwise, each intermediate value between the upper and lower limits of that range, accurate to the tenths of the lower limit unit, is also specifically disclosed. Embodiments of this disclosure cover every smaller range between any stated values or intermediate values within the stated range and any other stated values or intermediate values within the stated range. The upper and lower limits of these smaller ranges may independently include or exclude them from the range, and each range including any boundary, no boundary, or two boundaries within a smaller range is also covered by this disclosure, subject to any specifically excluded boundaries in the stated range. Where a stated range includes one or both boundaries, ranges excluding any one or both of those included boundaries are also included in this disclosure.
[0082] Standard abbreviations can be used, such as bp, base pair; kb, kilobase; pi, picoliter; s or sec, second; min, minute; h or hr, hour; aa, amino acid; nt, nucleotide; etc.
[0083] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. While any methods and materials similar to or equivalent to those described herein may be used in practice and testing of embodiments of this disclosure, some potential and exemplary methods and materials may now be described.
[0084] Detailed description
[0085] In this disclosure, for various purposes, we have developed novel methods using fragmented nucleosome signaling patterns at locations surrounding a target site. For example, nucleosome signaling patterns can be used to determine the methylation level of a target site (e.g., a CpG site). Other examples may utilize signals associated with nucleosome patterns of cfDNA molecules within genomic regions that exhibit differential methylation in a target tissue type (e.g., a tissue type of an organ) by having different methylation levels (or multiple levels, e.g., as patterns) relative to one or more other tissue types (e.g., blood cells). Such nucleosome patterns may be referred to as CpG-associated cfDNA nucleosome patterns. Tissue types may include, but are not limited to, blood cells, liver, lung, bladder, kidney, spleen, pancreas, heart, stomach, intestine, etc. Nucleosome signaling patterns can be compared with one or more reference patterns having known methylation levels.
[0086] Another example method can determine the lesion grade in a subject, for example, detecting a patient with a lesion (e.g., cancer); such examples may include determining the tissue type in which the lesion is present. In this example, a nucleosome signal pattern can be compared with one or more reference patterns determined based on one or more training samples with known lesion grades. Example results evaluate the diagnostic performance of distinguishing patients with and without hepatocellular carcinoma (HCC) based on nucleosome signals from differentially methylated regions derived from hepatocellular carcinoma (HCC) tumor tissue and leukocyte layer cells.
[0087] Another example is determining the proportional concentration of DNA from a specific tissue type, such as the proportional concentration of DNA with differential methylation at a site. In this example, nucleosome signal patterns can be compared to one or more reference patterns determined from one or more calibrated samples having known proportional concentrations of DNA from a specific tissue type. Such techniques do not require tissue-specific alleles and therefore do not require additional initial measurements to determine tissue-specific alleles or other tissue-specific markers. For example, a tumor biopsy is not required to identify tumor-specific markers, such as tumor-specific alleles.
[0088] Such differences in tissue type methylation can cause fragmentation effects across long distances (e.g., across one or more nucleosomes). Fragmentomics-based methylation analysis can include extended regions with multiple nucleosomes. Using fragmented nucleosome signal patterns encompassing multiple nucleosomes can provide improved accuracy and / or alternative methods for such measurements that can be combined with other techniques. Such measurements can include methylation levels, lesion grade, and the proportional concentration of DNA from tissue type (e.g., tumor tissue type). For example, measurements of methylation-dependent nucleosome signaling can enable cancer detection because DNA methylation patterns in diseased cells (e.g., tumors) can be preferentially altered compared to other cell types (e.g., hematopoietic cells). Such cancer detection can achieve improved accuracy relative to other techniques, allowing for fewer false positives and false negatives and enabling early therapeutic intervention.
[0089] Signals associated with nucleosome patterns can be defined in various ways, for example, as the ratio of the number of cfDNA molecules spanning a genomic window (e.g., one of the various sizes described herein, such as 140 bp) to the number of cfDNA molecules whose ends lie within such a window. Nucleosome signals can be determined using all cfDNA fragments in a plasma sample or cfDNA fragments of specific sizes (e.g., 60 bp to 120 bp, 80 bp to 140 bp, 100 bp to 160 bp, 120 bp to 180 bp, 140 bp to 200 bp, etc.). Such signals can also be referred to as nucleosome signals, which can be methylation-dependent. The advantage of using nucleosome signals to report methylation changes is that it eliminates the need for processes that differentially modify or identify DNA molecules based on their methylation state (e.g., bisulfite treatment), processes that significantly degrade DNA (Grunau et al. Nucleic Acids Res. 2001;29, E65). Furthermore, the use of nucleosome signals can leverage other types of epigenetic modifications that affect nucleosome patterns, potentially enhancing detection capabilities. Such techniques can also be combined with other technologies, for example, to form integrated models for determining site methylation, lesion grade, or the proportion of cell-free DNA from a specific tissue type.
[0090] Various standardization methods can be applied to obtain a class of nucleosome signals that can be used for the various applications described in this paper.
[0091] As an example result below, we used a deep learning model to classify whether CpG sites were HCC-specific hypomethylation or hypermethylation, achieving an area under the receiver operating characteristic (AUC) of 0.87. Hepatocellular carcinoma (HCC) patients showed a reduced amplitude of nucleosome patterns compared to subjects without cancer, with the amplitude gradually decreasing with tumor stage. As another example result, using a machine learning model to detect patients with HCC using nucleosome patterns associated with differentially methylated CpG sites, the AUC was 0.93. We further validated the cancer detection method in independent populations. As another example result, using a machine learning model (e.g., a regression function) of nucleosome patterns to quantitatively measure the proportion of DNA from the first tissue type (e.g., the proportion of fetal DNA in pregnant women) showed good agreement with single nucleotide polymorphism analysis (Pearson r = 0.85).
[0092] This disclosure reveals the interaction between nucleosome signaling and methylation status. Example implementations can utilize methylation-associated nucleosome patterns to report tissue-derived cfDNA molecules, such as in cancer detection. This approach can incorporate methylation signals without requiring bisulfite sequencing, thus opening new possibilities for diagnostics using cfDNA sequencing.
[0093] In the following description, we first investigate whether those CpG sites also exhibit differentially fragmented signals across genomic regions. Then, we evaluate the diagnostic performance in distinguishing patients with and without hepatocellular carcinoma (HCC) based on nucleosome signals from differentially methylated regions derived from HCC tumor tissue and leukocyte layer cells. Furthermore, nucleosome signals associated with tissue of origin were validated in pregnant women.
[0094] I. The relationship between nucleosome signaling and methylation
[0095] This disclosure describes methods and techniques for correlating nucleosome signaling patterns with methylation (e.g., the methylation level at a specific site in the genome). The methylation state at a given site in a given cell can statistically influence how enzymes cleave DNA over a long range (e.g., for the length of a genomic region described herein) relative to a site (e.g., across one or more nucleosomes). Such changes in methylation affect nucleosome signaling at each location in the genomic region, thereby allowing nucleosome signaling patterns to be used to detect the root cause of fragmentation alterations. Such changes in methylation can occur in various tissue types (e.g., fetus / pregnancy and tumors). Aspects of fragmentation analysis, including nucleosome signaling patterns, are now discussed.
[0096] A. Cell-free DNA fragmentomics and methylation
[0097] Figure 1 It is a chart illustrating the various molecular characteristics associated with the identification of cell-free DNA molecules. Figure 1 Different biomarkers in cell-free DNA can be used alone or in combination, including methylation (e.g., methylation patterns, such as the location of hypermethylation and hypomethylation, global hypermethylation, etc.), fragmentomics (e.g., terminal motifs, serrated ends, fragment size, preferred end coordinates, window protection score, and nucleosome signal patterns), and topology (e.g., circular mitochondrial DNA, linear mitochondrial DNA, extrachromosomal circular DNA).
[0098] B. Nucleosome signaling pattern
[0099] Figure 2 An example of a free DNA fragmentation pattern regarding nucleosome signaling around a genomic region of interest 210 is shown. Nucleosome signaling at a genomic location can be measured using a genomic window with a specified width (e.g., 140 bp) centered at the genomic location (e.g., location 212). In various embodiments, windows of sizes of 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 110 bp, 120 bp, 130 bp, 140 bp, 150 bp, 160 bp, 170 bp, 180 bp, 190 bp, 200 bp, etc., or any value smaller than or larger than these example sizes, can be used.
[0100] The number of sequencing cfDNA fragments that completely (222) and partially (224) cover the window can be determined separately. The portion of the sequencing fragment that completely covers the window is called the nucleosome signal. For example... Figure 2 As shown, the number of cfDNA fragments that completely cover the window (denoted by "F") and the number of cfDNA fragments that partially cover the window (denoted by "P") are calculated respectively. The value of F / (F + P) can be used to reflect nucleosome signal.
[0101] DNA nucleases will preferentially cleave the linker DNA 232 between the two nucleosomes 230 to produce cfDNA fragments. The F / (F + P) value will be a peak of 240 when the center of the analyzed window corresponds to a bipartite nucleosome. The F / (F + P) value can be a trough of 242 when the center of the analyzed window corresponds to the linker DNA 232 between the two nucleosomes. The nucleosome signal can be determined for each genomic location 210, such as the location covered by the sequenced cfDNA molecule. This value (F / (F + P)) can be used as the nucleosome signal. Other types of separation values can be used, such as P / (F+P), P / (FP), F / (FP), PF, FP, F / P, or P / F. Any of these separation values using such quantities can be used to determine the nucleosome signal at each location.
[0102] If a location on the genome (e.g., in a dimorph) is highly protected, there are segments with more complete coverage, resulting in high nucleosome signal. If the location is exposed (less protected) and highly degraded, the nucleosome signal is low. Figure 2 This demonstrates the existence of this pattern. The signal peaks at the same location on the nucleosomes.
[0103] Fragmentation in serum and urine should be similar to that in plasma, as the DNA fragments are the same but collected in different ways. Other samples besides plasma and serum should show similar fragmentation, as such fragmentation is still primarily produced by the same enzymatic process of DNA cleavage and therefore has the same methylation effect at a given site.
[0104] C. Standardization using nucleosome signal values
[0105] Some implementations can correct for inter-sample variability by performing standardization. Signal standardization within a region can use signals from the region itself or from outside the region (e.g., flanking regions adjacent to the region). Signals from outside the region can originate from anywhere else in the genome and can be any number (e.g., 2, 3, 4, 5, and more) of external regions, each with a different length, such as 1 kb, 5 kb, 10 kb, 50 kb, 1 Mb, and longer. Such standardized signals can be referred to as background signals. Some sites within the region can be used, or all sites within the region can be used. Statistical values (e.g., averages) across all locations or based on each location can be used to standardize any given nucleosome signal, for example, by division or subtraction.
[0106] Figure 3 An example of nucleosome signaling at the CTCF binding site with a strong localization nucleosome is shown. Figure 3The standardization of nucleosome signals within the region using the average nucleosome signal of that region is shown. Figure 310 shows the nucleosome signal before standardization, and Figure 320 shows the region-level standardized nucleosome signal applicable to each signal value, relative to the standardization for each location described later. Therefore, region-level statistics can be used.
[0107] Nucleosome signal spectra were obtained from plasma DNA of healthy subjects. The horizontal axis represents the position relative to CTCF binding sites. The vertical axis represents the nucleosome signal or nucleosome score (normalized signal). Each line corresponds to a different sample. For a given position at a line, the value was determined as the sum across a set of CTCF binding sites. Each sample consistently displays the signal relative to the nucleosome signal. Figure 3 Similar patterns are shown, as demonstrated in Figures 310 and 320.
[0108] To minimize inter-sample variability regarding nucleosome signal, the nucleosome signal can be standardized using values (e.g., statistical values, such as the mean or median) of nucleosome signals derived from the region itself and / or other regions (e.g., flanking regions). The standardized nucleosome signal can be referred to as a nucleosome score. In one example, the nucleosome signal across the genomic region of interest minus the mean of nucleosome signals in that genomic region and / or one or more other regions can be considered the nucleosome score (standardized signal). In another example, division can be used in the signal standardization step, e.g., division by a statistical value or other baseline.
[0109] As can be seen, the changes in Figure 320 are much smaller than those in Figure 310. This reduction in such changes can provide higher accuracy for comparing with a reference nucleosome signal pattern, which can be used for the various purposes described herein.
[0110] In another example, the flanking regions are 2-kb windows upstream and downstream of the normalization factor. The mean or other statistics of the flanking regions can be used to standardize the curve. This type of normalization can provide a more comparable normalized signal across different samples. Using the normalization score (normalized signal value) at each location, we can calculate the nucleosome signal pattern along the genome.
[0111] II. Determine the methylation level at sites in the genome.
[0112] Nucleosome signaling patterns can detect biases in fragmentation patterns caused by varying methylation levels. Different reference patterns can be determined for different methylation levels. Comparing a measured nucleosome signaling pattern with one or more reference patterns having known methylation levels (e.g., measured from a sample where methylation is determined using another technique, such as bisulfite treatment) provides a measurement of the methylation levels at genomic loci in a new sample. Such measurements using nucleosome signaling patterns avoid the use of bisulfite treatment, which degrades DNA, resulting in fewer usable DNA fragments.
[0113] A. Methylation levels as generally measured
[0114] Figure 4A The differences in nucleosome signal patterns at sites with different methylation levels are shown. The horizontal axis shows the location relative to the CpG site. As shown, the region surrounding the CpG site is + / - 800 bp (approximately 1,600 bp total width). Other regions of different sizes are described, such as those with total widths (lengths) greater than or less than 400, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, etc. The vertical axis shows the nucleosome score (normalized nucleosome signal pattern) at each location within the region surrounding the CpG site. The region may include one or more CpG sites separated from the target CpG site at location 0.
[0115] Nucleosome signal pattern 410 corresponds to CpG sites with a methylation density greater than 70% (range: 70% to 100%, median: approximately 85%) in leukocyte layer cells. Lines of nucleosome signal pattern 410 correspond to a set of hypermethylated regions using pooled cfDNA data. Nucleosome signal pattern 420 corresponds to CpG sites with a methylation density less than 30% (range: 0% to 30%, median: approximately 15%) in leukocyte layer cells. Lines of nucleosome signal pattern 420 correspond to a set of hypomethylated regions.
[0116] As shown in the figure, the two nucleosome signal patterns (410 and 420) are different. In one example, a measured nucleosome signal pattern more similar to nucleosome signal pattern 410 can be classified as hypermethylated. In another example, a measured nucleosome signal pattern more similar to nucleosome signal pattern 420 can be classified as hypomethylated. Other classifications for different methylation levels, such as intermediate values or numerical values, such as ranges, can also be used.
[0117] The majority (e.g., 80% or more) of cell-free DNA in plasma originates from hematopoietic cells. Therefore, nucleosome methylation signaling patterns from plasma can be used to detect the overall methylation level at a specific site in a sample.
[0118] Hypomethylation can be defined, but is not limited to, below 10%, below 20%, below 30%, below 40%, below 50%, etc. Hypermethylation can be defined, but is not limited to, above 40%, above 50%, above 60%, above 70%, above 80%, above 90%, etc. Thresholds can be ranges, such as, but not limited to, 20% to 30%, 30% to 40%, 40% to 50%, 60% to 70%, 70% to 80%, etc. Figure 4A In this study, based on the bisulfite sequencing results of leukocyte layer samples, low-methylated CpG sites and high-methylated CpG sites were defined as CpG sites with methylation levels of <30% and >70%, respectively (Lun et al. Clin Chem. 2013;59(11):1583-1594; Tse et al. Proc Natl Acad Sci USA. 2021;118(5):e2019768118).
[0119] like Figure 4A As shown, cfDNA molecules were aligned relative to low-methylated CpG sites (denoted as position 0) to determine nucleosome scores at each genomic location surrounding the CpG site. Similarly, cfDNA molecules were aligned relative to high-methylated CpG sites (denoted as position 0) to determine nucleosome scores at each genomic location surrounding the CpG site. Nucleosome scores typically appear as wavy signals around low-methylated and high-methylated CpG sites. However, the nucleosome scoring patterns differ between high-methylated and low-methylated CpG sites in healthy subjects.
[0120] There is a shift (phase difference) in the positions of peaks and troughs related to nucleosome scores associated with high and low methylation. Therefore, the difference between two signal patterns is the phase, such as the location of the peak (maximum) and trough (minimum). This phase difference can be viewed as an offset in the positions of the maximum / minimum values of the two signals. A comparison can provide the relative phase difference between the measured signal pattern and one of the reference patterns, where the phase difference can be correlated with the methylation level. The reference pattern can also be relative to the target site (i.e., Figure 4A The sine or cosine of the specified phase (0) can be used to determine the phase difference for different signal patterns with known methylation levels at a site (or a group of sites with similar methylation levels, e.g., all hypomethylated or all hypermethylated sites). Therefore, the phase difference can be measured in a variety of ways.
[0121] B. Determine the methylation level at tissue-specific sites.
[0122] For a specific tissue type at a specific site, the methylation level can be determined even when cell-free DNA from that tissue type is typically in the minority of a sample (which is a mixture of cell-free DNA from multiple tissue types). This can be done by using sites that are typically hypomethylated in other tissues (e.g., hematopoietic cells) but hypermethylated in specific tissue types (e.g., liver tumors). In this way, any variation in the pattern can be attributed to the methylation level at a single site or a set of sites.
[0123] Figure 4B Enhanced wavy nucleosome signaling is shown around HCC tissue-specific hypermethylated CpG sites (440) and hypomethylated CpG sites (430). HCC is used as a specific tissue type, but healthy tissue types can also be used. In this example, HCC tumor-specific hypomethylated sites are defined by those CpG sites with a methylation density of less than 30% in HCC tissue but more than 70% in leukocyte layer samples.
[0124] Surprisingly, as Figure 4B As shown, when focusing on CpG sites exhibiting tissue-specific DNA methylation, the oscillations between peaks and troughs in the nucleosome score are significantly enhanced. In this analysis, we analyzed 118,544 HCC-specific hypermethylated CpG sites (e.g., methylation levels >70% in HCC tumors and <30% in leukocyte layers) and 842,892 HCC-specific hypomethylated sites (e.g., methylation levels <30% in HCC tumors and >70% in leukocyte layers). This oscillation pattern appears to be asynchronous between HCC-specific hypomethylation and hypermethylation. These results suggest that methylation patterns can influence nucleosome signaling in cfDNA molecules.
[0125] To further investigate how the two types of differentially methylated CpG sites (DMS) are associated with chromatin organization of the genome, we overlapped DMS with regions A or B rich in open and closed chromatin, respectively. Figure 5A As shown, type A DMS (HCC-specific hypermethylated CpG sites) was approximately 4-fold more enriched in region A than in region B (78% vs 20%). In contrast, type B DMS (HCC-specific hypomethylated sites) showed an inverse trend with a lower proportion in region A than in region B (35% vs 60%).
[0126] We split these hypomethylated and hypermethylated CpG sites into training and test sets and used a CNN model to predict the methylation index of individual CpG sites using cfDNA nucleosome scoring. In various implementations, the CNN model can use two two-dimensional (2D) convolutional layers, for example, each with 16 filters of kernel size 4; other values can also be used. Batch normalization layers can then be applied, followed by convolutional layers. The Rectified Linear Unit (ReLU) activation function can be used for these convolutional layers; other activation functions can also be used. Max pooling layers with a pooling size of 2 are used; other values can also be used. A spread layer can be further added, followed by a fully connected layer with 3200 neurons using the ReLU activation function. Finally, an output layer with two neurons can be applied using the softmax function to generate a hypermethylation probability score for the CpG site. Other values can be used for any of the parameters mentioned above.
[0127] Figure 5B Receiver operating characteristic (ROC) curve analysis is shown for predicting methylation levels (e.g., exponential) at sites in the training and test sets. We achieved an area under the ROC curve (AUC) of 0.87 in the test set. Therefore, nucleosome signaling patterns can be used to determine methylation levels at specific sites in the genome, including in specific tissue types.
[0128] We examined the overlap between liver-specific CpGs (normal liver vs. leukocyte layer) and HCC-specific CpGs (HCC tumor vs. leukocyte layer). There was 80,543 overlap sites between HCC hypermethylation sites (118,544) and liver hypermethylation sites (258,630). There were 78,755 overlap sites between HCC hypomethylation sites (842,892) and liver hypomethylation sites (226,417). Therefore, tissue-specific or disease-specific sites can be used to determine methylation levels in a specific tissue type.
[0129] Additionally or alternatively, methylation alterations in specific diseases can be reflected by analyzing these nucleosome signals associated with tissue-specific methylation patterns, as demonstrated by pattern differences in HCC tissues. The advantage of using nucleosome signals is that it allows for the elimination of bisulfite treatment, which significantly degrades DNA material. Such techniques are described in later sections.
[0130] C. Methods for determining methylation levels
[0131] Figure 6This is a flowchart illustrating a method 600 for determining methylation levels at a site using nucleosome signal patterns according to embodiments of this disclosure. Various examples of method 600 have been described above. Method 600 can be performed partially or entirely using a computer system. Method 600 can determine methylation levels without the need for bisulfite conversion of degradable DNA. Therefore, higher sequencing library yields can be obtained.
[0132] At box 610, multiple cell-free DNA molecules from a biological sample of the object are analyzed. Various techniques can be used to perform such analyses in any of the methods described in this disclosure. For example, sequencing can be used, such as high-throughput parallel sequencing, targeted sequencing, and single-molecule sequencing (e.g., using nanopores or using real-time single-molecule sequencing (e.g., from Pacific Biosciences)). The analysis may include physical steps of performing such assays and receiving measurement data obtained from such assays, or it may only include the step of receiving measurement data.
[0133] Analyzing cell-free DNA molecules can include determining genomic locations in a reference genome corresponding to at least one end of the cell-free DNA molecule. Therefore, analyzing each of a plurality of cell-free DNA molecules can include determining two genomic locations in a reference genome corresponding to both ends of the cell-free DNA molecule. For example, one or more sequence reads of a DNA molecule (e.g., paired reads at the ends or reads of the entire molecule) can be aligned to a reference genome using any of a variety of alignment techniques known to a person skilled in the art. Alignment can be performed with some or all of the reference genome. Location determination can be targeted at some or all of the reference genome, for example, if only a portion of the genome is analyzed. As an example, the amount of genome analyzed can be greater than 0.01%, 0.1%, 1%, 5%, 10%, or 50% of the reference genome. Such analyses can be performed with respect to other methods described herein.
[0134] At box 620, the nucleosome signal pattern of the genomic region surrounding the target CpG site is determined. The target CpG site may be located at the center of the genomic region. The nucleosome signal pattern can be determined in various ways, as described herein, such as using one or more normalization techniques. For example, the nucleosome signal pattern may include nucleosome scores, normalized nucleosome scores, or distribution-adjusted nucleosome scores (e.g., z-scores or t-scores). Such values are examples of position-by-position nucleosome signals. Position-by-position nucleosome signals can be determined using cell-free DNA molecules spanning a window around the genomic location and cell-free DNA molecules whose ends are located within the window around the genomic location.
[0135] Nucleosome signaling patterns can be determined in the following way: For each genomic location within a genomic region, a first quantity of multiple cell-free DNA molecules spanning a window around the genomic location can be determined. A second quantity of multiple cell-free DNA molecules with ends located within a window around the genomic location can be determined. Figure 2 Examples of such determinations are provided. The nucleosome signal at each location can be determined using the first and second quantities as described herein. For example, the first quantity F and the second quantity P can be determined using any of the following separation values: P / (F+P), P / (FP), F / (FP), PF, FP, F / P, or P / F.
[0136] This article provides example sizes for regions. For example, the length of a genomic region can be at least 140 bp. Other sizes of genomic regions include at least 100 bp, 110 bp, 120 bp, 130 bp, 150 bp, 160 bp, 170 bp, 180 bp, 190 bp, 200 bp, 250 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, and 1,000 bp, as well as other sizes described herein.
[0137] As an example of normalization, nucleosome signals at individual locations can be normalized using nucleosome signals from other genomic regions (e.g., flanking regions). This type of normalization can use the average value of nucleosome signals from other genomic regions, which may be adjacent to the original genomic region.
[0138] As another example of standardization, nucleosome signal patterns are standardized using region-level statistics determined based on one or more regions. This location-by-location nucleosome signal can be referred to as a nucleosome score. The one or more regions can include genomic regions or one or more other regions. The one or more regions can include one or more other regions adjacent to that genomic region. As an example, the region-level statistics can be the mean or the median.
[0139] As another example of normalization, the nucleosome signal at each location can be normalized using a background signal that depends on the distance from the genomic location to the target CpG site, as described in more detail later. The background signal can be determined based on one or more other genomic regions, each centered at the CpG site. One or more regions can be randomly selected. Normalization can be performed for each distance from the CpG site using the mean or median of the nucleosome signal from the other genomic regions.
[0140] As another example of standardization, the nucleosome signal at each location can be standardized using a distribution of a reference nucleosome signal determined from one or more regions of one or more reference samples. Standardization using a distribution may include, for each location, determining a total statistic and a dispersion of the reference nucleosome signal distribution, subtracting the total statistic from the nucleosome signal at each location, and dividing by the dispersion. For example, the distribution could be a normal distribution, where the total statistic is the mean of the reference nucleosome signal at the location, and the dispersion is the standard deviation.
[0141] At box 630, the methylation level of the target CpG site in the target genome is determined based on a comparison of the nucleosome signaling pattern with a reference pattern. The reference pattern can be determined from one or more training samples with known methylation levels. Examples of reference patterns are described in this paper. Comparisons can be made in various ways, such as using machine learning models, such as SVM or other models described herein. Known methylation levels can be in various forms, such as ranges (e.g., with 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10%), or can be considered as specified values.
[0142] The methylation level can be either highly methylated or lowly methylated at the target site. In another example, methylation is a range of methylation density at the target site. Examples of such ranges can include 0% to 10%, 10% to 20%, 20% to 30%, 30% to 40%, 40% to 50%, 50% to 60%, 60% to 70%, 70% to 80%, 80% to 90%, and 90% to 100%. Other examples include smaller or wider widths of said ranges, such as 5%, 10%, 15%, 20%, or 25%.
[0143] As an example, the biological sample could be plasma or serum. Differential methylation of the target site is known to exist between the first tissue type and blood cells, and the methylation level of the target CpG site in the first tissue type is determined therein. As an example, the first tissue type could be cancer tissue, fetal tissue, or tissue from a specific organ.
[0144] III. Cancer classification based on nucleosome signaling patterns
[0145] Nucleosome signals from surrounding sites (e.g., tissue-specific hypomethylated and / or hypermethylated sites) can be used to detect and monitor lesions (e.g., diseases such as cancer).
[0146] A. Nucleosome signaling patterns in cancer patients and healthy controls
[0147] We investigated the cancer detection capability using cfDNA-based nucleosome scoring across CpG sites, which exhibit differential methylation patterns across various tissues. Nucleosome scores were determined using cfDNA molecules from HCC-specific hypomethylated and hypermethylated CpG regions derived from healthy subjects and patients with cancer in Dataset A. Dataset A (Jiang et al. Cancer Discov. 2020;10:664-673) comprises paired-end sequencing data (75 bp x 2, Illumina) of plasma cfDNA obtained from healthy controls (n=38), subjects with chronic hepatitis B (HBV, n=17), hepatocellular carcinoma patients (HCC, n=34), and pregnant women (n=30), with a median of 38 million paired-end sequencing reads (range: 18 million to 65 million).
[0148] Figure 7A and Figure 7B The mean nucleosome scores (normalized nucleosome signal) for HCC-specific hypermethylated and hypomethylated sites are shown in healthy controls and patients with advanced HCC (aHCC). Nucleosome signal patterns 710 and 730 are for healthy controls. Nucleosome signal patterns 720 and 740 are for HCC patients. Differences between HCC patients and healthy controls can be seen. Nucleosome scores around HCC-specific hypomethylated and hypermethylated CpG sites are shown (relative to CpG sites -800 bp to 800 bp).
[0149] For HCC-specific hypermethylation sites, patients with advanced HCC (aHCC) showed weaker amplitude of nucleosome patterns compared to healthy controls. This can be seen in... Figure 7A This can be seen from the peak-valley distances shown for D1 to D8.
[0150] Figure 8A This displays the average peak-to-trough distance between healthy controls and aHCC subjects. The bar values are averages taken over the samples; for example, the blue bars are the averages taken over healthy samples, and the red bars are the averages taken over HCC samples. For example, the average peak-to-trough distance can be determined at the peaks / troughs (e.g., D1 to D8) of each sample, and then those averages can be averaged. Peak-to-trough pairs can be counted from the left or right. aHCC subjects show a significant reduction in distance. The use of such distances is one way to compare the nucleosome signal pattern of a sample (measured in a new sample) with a reference pattern for each known classification (e.g., healthy or cancer).
[0151] Figure 8BThe mean peak-to-trough distances are shown for various subjects: control, HBV, early HCC (eHCC), intermediate HCC (iHCC), and late HCC (aHCC). Each data point corresponds to a sample of subjects, where the mean peak-to-trough distance for the sample is determined by the number of peak-to-trough pairs in the graph. When comparing individual subjects, we found that nucleosome amplitude (mean) was significantly reduced in HCC patients compared to non-HCC subjects (P < 0.001, Wilcoxon rank-sum test). Furthermore, we observed that nucleosome amplitude gradually decreased with tumor stage, showing median amplitudes of 0.62, 0.51, and 0.40 for early HCC (eHCC), intermediate HCC (iHCC), and late HCC (aHCC), respectively, which further supports the contribution of tumor-derived DNA to attenuating nucleosome amplitude.
[0152] Similarly, Figure 8C and Figure 8D Nucleosome amplitude at HCC-specific hypomethylated CpG sites was also reduced in HCC patients (P = 0.0015, Wilcoxon rank-sum test), and HCC patients showed the smallest peak-valley distance.
[0153] B. Standardize using the background and mean of the healthy sample (e.g., z-score).
[0154] Standardization can be performed to make the differential signals more significant and intense. For illustrative purposes, differences between hypermethylation sites and hypomethylation sites can be connected or observed side-by-side, as shown below.
[0155] Figures 9A to 9B and Figure 10 An exemplary schematic diagram is shown for standardizing nucleosome scores around CpG sites. Nucleosome scores from clusters associated with hypomethylated and hypermethylated CpG sites are analyzed together (e.g., concatenated) to form a relatively large nucleosome score spectrum. Each curve corresponds to a different sample. Each curve can be standardized using a background signal. The background signal can be determined in various ways. For example, the signal from random CpG sites on the genome or a set of multiple random CpG sites can be used to determine the background signal. Mean bias standardization (e.g., z-score standardization) can then be performed. In addition to z-scores, other distribution-adjusted nucleosome scores can also be used.
[0156] Figure 9A This shows the signals of HCC-specific hypomethylation and hypermethylation sites. Curve 910 corresponds to the HCC sample. Curve 920 corresponds to the healthy sample. As mentioned above, curve 910 for HCC has smaller amplitudes (i.e., smaller distances) in the peaks and troughs compared to the healthy sample.
[0157] Using the same HCC-specific sites as described above, a nucleosome score can be determined for a given curve of a sample by summing data across the sites. One example is to align 842,892 HCC-specific hypomethylated CpG sites at position "0" and then pool all DNA fragments mapped relative to the "0" -800 bp to 800 bp region to construct a nucleosome curve for the sample. Another example is to determine the average of the 842,892 nucleosome curves at each position. These two techniques can be equivalent if there is sufficient DNA fragmentation at each CpG site.
[0158] Figure 9B The results show the normalization using background nucleosome scores. Background nucleosome scores for each sample were obtained by summarizing 1,000,000 genomic regions, each centered at a CpG site. HCC-specific nucleosome scores were then normalized using these background nucleosome scores. Curve 930 corresponds to HCC samples. Curve 940 corresponds to healthy samples. Compared to curve 940 for healthy samples, the smaller the amplitude of curve 930 for HCC, the more significant the normalization.
[0159] We randomly selected 1,000,000 genomic regions, each centered at a CpG site. Other numbers of randomized genomic regions can be used, for example, at least 500, 1,000, 5,000, 10,000, 50,000, 100,000, 200,000, 500,000, and 2,000,000 regions. Based on their relative position to the central CpG site (i.e., position 0), sequenced cfDNA molecules from these genomic regions were aggregated. Nucleosome scores spanning those relative positions were determined as described herein and referred to as background nucleosome scores. Since the number of background nucleosome scores corresponds to the size of the region, a background pattern for the background nucleosome scores can be referenced.
[0160] Based on relative position, nucleosome scores from clusters associated with hypomethylated and hypermethylated CpG sites are divided (or subtracted) from background nucleosome scores. That is, each score is divided or subtracted from the corresponding value. Nucleosome scores from clusters associated with hypomethylated and hypermethylated CpG sites are concatenated to form a spectrum of relatively large nucleosome scores.
[0161] Figure 10 z-score standardization of nucleosome scores was displayed. A set of healthy controls was used as a baseline to standardize the z-scores for each sample. Based on the set of healthy controls, we further determined the mean (µ) and standard deviation (δ) of the standardized nucleosome scores for each relative position. Figure 9BThe standardized nucleosome score(s) (i.e., standardized based on the background pattern on a sample-by-sample basis) is further transformed into a z-score value based on the following formula: (s-µ) / δ, which is called the nucleosome z-score. In other examples, the median or mode can be used instead of the mean or average. Therefore, various aggregate statistics can be used. Similarly, variance can be used instead of standard deviation. Therefore, various measures can be used for expanding the value (score).
[0162] z-score standardization can be achieved using the mean and standard deviation of the healthy samples. Therefore, the formula can be (healthy - mean signal) / (healthy - standard deviation). z-score standardization can be performed individually for each point. Therefore, there will be a different mean and standard deviation for each point. Thus, the entire curve can be standardized point-by-point.
[0163] like Figure 10 As shown, the HCC patient signal 1050 (normalized nucleosome z-score pattern) exhibits a unique nucleosome z-score pattern with significant fluctuations deviating from the relatively flat healthy control signal 1060. This unique pattern in HCC patients suggests that nucleosome scoring can be informative for cancer patients. For cell-free DNA from cancer patients, we see the theoretical peak of the nucleosome signal pattern from cancer. Cell-free DNA is a mixture of cell-free DNA molecules from blood cells and from tumors. However, after normalization, the blood signal is normalized, therefore the HCC patient signal primarily originates from the tumor.
[0164] The difference between the HCC patient signal 1050 and the healthy control signal 1060 is more significant between HCC-specific hypomethylated sites and hypermethylated sites. This difference may be due to the fact that the number of hypomethylated sites used is approximately seven times that of hypermethylated sites. Therefore, increasing the number of sites analyzed can improve accuracy. As an example, the number of sites could be 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 500,000, one million, or more.
[0165] C. Machine Learning Technology
[0166] Patterns of nucleosome signals from various regions (e.g., raw nucleosome signal patterns or normalized using the mean nucleosome signal (nucleosome score), background signals, and / or statistics from other samples, such as z-score normalization) can be used for downstream classification analysis. Various machine learning models can be used (e.g., as mentioned in the terminology section), such as, but not limited to, Support Vector Machines (SVM), analytical learning, artificial neural networks, backpropagation, boosting (a common algorithm), Bayesian statistics, case-based reasoning, decision tree learning, inductive logic programming, Gaussian process regression, genetic programming, data grouping, kernel estimation, learning automata, learning classification systems, minimum information length (decision trees, decision graphs, etc.), multilinear subspace learning, Bayesian classifiers, maximum entropy classifiers, conditional random fields, nearest neighbor algorithm, probabilistic approximate correct (PAC) learning, ripple descent rules, knowledge acquisition methods, symbolic machine learning algorithms, subsymmetric machine learning algorithms, minimum complexity machines (MCM), random forests, classifier ensembles, ordered classification, data preprocessing, handling imbalanced datasets, statistical relation learning, or Proaftn (a multi-criteria classification algorithm). Models can include linear regression, logistic regression, deep recurrent neural networks (e.g., Long Short-Term Memory, LSTM), hidden Markov models (HMM), linear discriminant analysis (LDA), k-means clustering, density-based spatial clustering with noise (DBSCAN), random forest algorithms, etc. Standardized nucleosome scores before z-score transformation, nucleosome scores before correction based on background nucleosome scores, or nucleosome signals before standardization based on other regions can be used as features for classification analysis.
[0167] The features of an ML model can be nucleosome signals at a set of sites within a region. In various implementations, different nucleosome signal values can be used. The entire nucleosome signal pattern can be used, but a subset of signal values at only some locations within the region can also be used. For example, each of the other locations or randomly selected; a reference pattern of samples with a known classification will have values at the same locations.
[0168] Different types of target sites can be used. For example, only hypomethylated sites or only hypermethylated sites can be used, or both sets can be used. Nucleosome values of sites of a given category (e.g., hypomethylated or hypermethylated) can be aggregated into a single pattern, or a single pattern can be used as a feature on its own.
[0169] 1. Example of SVM using HCC-specific differentially methylated sites
[0170] To illustrate the diagnostic capability of using nucleosome signaling patterns of cfDNA molecules (e.g., nucleosome scores that could be background-adjusted, distribution-adjusted, or raw nucleosome signals) to detect HCC patients, we used an example ML model of a support vector machine (SVM) model to determine the probability that a test sample had HCC. The input feature vector of a sample contained nucleosome score values at relative positions within 800-bp upstream and downstream of the CpG of interest. These CpG sites included 842,892 HCC-specific hypomethylated and 118,544 hypermethylated CpG sites. Therefore, a total of 3,202 features were used: the z-score for each position in this region.
[0171] like Figure 9A As shown, nucleosome scores at hypomethylation sites are summarized. For example, for 118,544 HCC-specific hypermethylation sites, all cfDNA fragments mapped from these 118,544 regions (each 1601 bp long) were used to determine the nucleosome signal value at the K position (any position within the region) of the HCC-specific hypermethylation pattern. The same applies to HCC-specific hypomethylation nucleosome signal patterns.
[0172] Figures 11A to 11C A set of graphs is shown to evaluate the performance of using cfDNA nucleosome signaling patterns as biomarkers for cancer diagnosis.
[0173] Figure 11A ROC curves for detecting HCC (i.e., distinguishing HCC from non-HCC) using various peak-valley pairs of HCC-hypermethylated sites and HCC-hypomethylated sites are shown. Figure 11A As shown, HCC and non-HCC objects can be classified directly using the peak-valley distance alone, with the receiver operating feature (ROC) AUC reaching as high as 0.82.
[0174] We further utilize the complete nucleosome model ( Figure 9A and Figure 9B The -800 bp to 800 bp region shown is used as the input feature for each sample, and a support vector machine (SVM) model is used to determine the probability of having HCC based on such patterns of nucleosome scores.
[0175] Figure 11B The probabilities of HCC in different subject groups (known to have different classifications) using an SVM model are shown. HCC probabilities were determined by nucleosome scores, without background or z-score standardization. Leave-one-out analysis showed that patients with HCC had a significantly higher probability of having HCC than those without (median: 0.81 vs 0.09). P< 0.001, Wilcoxon rank-sum test). Additionally, patients with advanced HCC tended to have a higher probability of having HCC compared to those with intermediate and early-stage HCC (median: 0.80 and 0.72, respectively) (median: 0.93). This is likely due to the higher proportion of cell-free DNA in advanced tumors. Determination of the proportion of cell-free DNA from specific tissue types is provided elsewhere in this disclosure.
[0176] Figure 11C The ROC curve of the SVM model is shown. ROC analysis indicates that the nucleosome pattern-based method based on cfDNA can distinguish between subjects with and without HCC, achieving an AUC of 0.93.
[0177] Figures 12A to 12B Cancer detection using alternatively standardized nucleosome signals is shown. Nucleosome scores are obtained by dividing the original nucleosome signal by the average nucleosome signal value of the flanking regions (-2000 bp to -1000 bp and 1000 bp to 2000 bp relative to CpG sites). The nucleosome scores are further standardized using background CpG nucleosome signals and z-score statistics to obtain the nucleosome z-score.
[0178] Figure 12A The probability of HCC (vertical axis) inferred from an SVM model using a nucleosome z-score pattern around HCC tissue-specific CpG sites (approximately 800 bp to 800 bp relative to the CpG site). The HCC probability was determined using the SVM model described above. The probability of having HCC was significantly assessed in HCC patients compared to those without HCC (median: 0.056 vs 0.934; P-value: 1.2e-11, Wilcoxon rank-sum test). Furthermore, patients with advanced HCC tended to have a higher probability of having HCC compared to those without advanced HCC (median: 0.98 vs 0.90). This is likely due to the higher proportion of advanced tumor DNA concentration. Determination of the proportion of free DNA from specific tissue types is provided elsewhere in this disclosure.
[0179] Figure 12B The ROC analysis shows that the predicted HCC probability was used to distinguish between HCC and non-HCC subjects. The ROC analysis showed that the nucleosome z-score method based on cfDNA can distinguish between subjects with and without HCC, achieving an area under the curve (AUC) of 0.97.
[0180] In this study, we analyzed fragmentation patterns in genomic regions with differentially methylated states. The majority of DMS were not located in the TSS (0.81%) and TFBS (25.5%) regions (i.e., 73.7% in other regions), indicating the presence of tissue-specific nucleosome fragmentation patterns in many other genomic regions. Therefore, CpG-associated nucleosome patterns in plasma DNA may contain molecular information independent of transcriptional activity and transcription factor binding events.
[0181] D. Universality (validation) of cancer detection models across different datasets
[0182] To test the universality of nucleosome-based cancer diagnosis, we even tested our method on separate cohort datasets using different sequencing protocols or platforms. We further hypothesize that a cancer detection model learned from one research cohort can be applied to other cohorts even with different sequencing platforms or protocols, without requiring retraining.
[0183] 1. Validation using the same sequencing platform but different protocols.
[0184] To test the generalizability of nucleosome pattern-based cancer diagnosis, we even tested our method on a separate cohort dataset B with different sequencing protocols. We validated nucleosome pattern-based cancer diagnosis in another cohort (dataset B). Dataset B was published in 2015 (Jiang et al. Proc Natl Acad Sci USA. 2015;112(11):E1317-1325). Dataset B (Jiang et al. Proc Natl Acad Sci USA. 2015;112(11):E1317-1325) includes paired-end sequencing data (75 bp × 2, Illumina) of plasma cfDNA obtained from healthy controls (n=32), cirrhotic subjects (n=36), HBV subjects (n=67), and HCC patients (n=90), with a median of 31 million paired-end sequencing reads (range: 18 million to 79 million).
[0185] The two datasets (A and B) used different experimental kits. Dataset A used a DNA library prepared using the TruSeq Nano DNA Library Prep Kit (Illumina). Dataset B used a DNA library prepared using the Kapa Library Preparation Kit (Kapa Biosystems).
[0186] a) Training and testing on dataset B
[0187] Figure 13A It provides a leave-one-out analysis using SVM to distinguish the HCC probability of HCC patients of various grades from all non-HCC subjects in dataset B. Figure 13B ROC curves are displayed to differentiate HCC patients from all non-HCC subjects. Nucleosome scoring is used for nucleosome signal patterns. Nucleosome scoring is used instead of background and z-score standardization.
[0188] As shown in the figure, we can use the probability of having HCC to achieve a good distinction between HCC patients and non-HCC subjects (including healthy controls, patients with cirrhosis, and HBV carriers) (median: 0.93 vs 0.10; P < 0.001, Wilcoxon rank-sum test). Figure 13A The ROC AUC for distinguishing HCC patients from non-HCC subjects was 0.93. Figure 13B Furthermore, we found that the predicted probability of HCC was correlated with tumor burden (). Figure 13A Patients with circulating tumor DNA (ctDNA) proportions below 3%, 3% to 10%, and above 10% had median probabilities of HCC of 0.96, 0.95, and 0.89, respectively. The proportion of ctDNA was also increased. The proportion was estimated based on copy number aberrations using ichorCNA (Adalsteinsson et al. Nat Commun. 2017;8:1324). Even when focusing on cirrhotic patients without HCC, the AUC could still be as high as 0.91 to distinguish between HCC and cirrhosis. Since cirrhosis is the most important risk factor for developing HCC, the ability of this method to accurately differentiate between HCC and cirrhosis is clinically important.
[0189] b) Train on dataset B and test on dataset A
[0190] We further challenged the robustness of this analytical framework by testing whether a cancer detection model trained on one research cohort could be generalized to another cohort (i.e., using different sequencing protocols). To this end, we attempted to distinguish between patients with and without HCC in dataset A using an SVM model for cancer detection trained on dataset B. Nucleosome scores were normalized using background CpG nucleosome signal and z-score statistics from each dataset to minimize inter-dataset bias.
[0191] In the analysis of dataset B, we used 32 healthy subjects, 36 subjects with cirrhosis, 67 HBV carriers, and 90 HCC patients. We used 16 healthy controls to infer the mean and standard deviation of nucleosome scores across locations relative to the CpG sites of interest. These means and standard deviations were used to determine the nucleosome z-scores in the independent test samples in dataset B. Similarly, in dataset A, a subset of healthy controls (n=19) was used to infer the mean and standard deviation to further determine the nucleosome z-scores for the remaining samples in dataset A. We standardized datasets A and B separately.
[0192] Figure 14A An SVM model trained on a dataset using one sequencing protocol is provided to distinguish the HCC probabilities of various grades of HCC patients from all non-HCC subjects on datasets using different sequencing protocols. An SVM model trained on dataset B is applied to predict the HCC probabilities of an independent dataset A. Figure 14B The ROC curves used to distinguish HCC patients from all non-HCC subjects in test dataset A are shown. Standardized z-scores were used for nucleosome signal patterns.
[0193] As shown in the figure, compared with non-HCC subjects in the independent test set, we can observe a significantly higher probability of having HCC in HCC patients. P < 0.001, Wilcoxon rank-sum test; Figure 14A We achieved an AUC of 0.93 to distinguish HCC patients from healthy controls, and an AUC of 0.86 to distinguish HCC patients from all non-HCC subjects, including HBV carriers. Figure 14B This indicates that nucleosome-based methods for cancer detection can be well generalized across different datasets.
[0194] c) Train on dataset A and test on dataset B.
[0195] Figures 15A to 15B The image shows cancer detection (HCC probability) in dataset B based on an SVM model trained on dataset A using nucleosome signals. In this example, the original nucleosomes were normalized to a nucleosome score by dividing by the average signal of the flanking regions (-2000 bp to -1000 bp and 1000 bp to 2000 bp relative to CpG sites) and then normalizing against the background CpG nucleosome signal and z-score statistics.
[0196] exist Figure 15AIn this study, the groups consisted of healthy controls, individuals with cirrhosis, HBV, and HCC. Using the probability of having HCC inferred from a model trained on dataset A (P=1.5e-20, Wilcoxon rank-sum test), we achieved good separation between non-HCC and HCC subjects. Furthermore, we observed an increasing trend in the probability of having HCC in patients with a higher proportion of tumor DNA.
[0197] Figure 15B The ROC analysis of the predicted HCC probability in distinguishing HCC and non-HCC subjects from dataset B is shown. The ROC curves show that the AUC for distinguishing HCC patients from healthy controls is 0.95, and the AUC for distinguishing HCC patients from all non-HCC subjects is 0.88.
[0198] In short, Figures 13A to 13B The cancer detection using leave-one-out analysis in dataset B is shown, demonstrating an implementation that works in another group. Figures 14A to 14B The example demonstrates cancer detection using an SVM model trained on dataset B in dataset A, and shows an implementation scheme that enables a model to generalize across different datasets. Figures 15A to 15B The image shows cancer detection in dataset B using an SVM model trained on dataset A.
[0199] 2. Survival rates of HBV subjects with high and low HCC risk
[0200] Figure 16 Displayed based on Figure 15A The chart shown is a Kaplan-Meil analysis of the survival of HBV subjects divided into two risk groups based on predicted HCC probabilities. HBV subjects with high HCC risk (HCC probability > 0.12) had significantly worse survival compared to those with low risk (HCC probability ≤ 0.12).
[0201] Importantly, follow-up clinical information analysis of dataset B showed that among the 7 HBV carriers predicted to be at high risk of HCC (HCC probability > 0.12), 4 HBV carriers (57.1%) developed HCC within 5 years after blood testing; however, among the 52 HBV carriers at low risk of HCC (HCC probability ≤ 0.12), only 6 HBV carriers (11.5%) developed HCC. The odds ratio for HCC development between the two groups was almost 5. Kaplan-Meil analysis of survival showed that HBV carriers at high risk of HCC had significantly worse survival rates compared to those at low risk of HCC (P < 0.0001).
[0202] The results indicate that patients can be grouped using the probability of developing HCC inferred from nucleosome signals, for example, by selecting a subgroup of HBV carriers with a high risk of HCC. Patients with a high risk of HCC will be recommended for more frequent clinical monitoring. These data also suggest that the method in this disclosure is more sensitive in screening HBV carriers who are likely to eventually develop HCC within a certain timeframe compared to conventional clinical monitoring and management, such as abdominal imaging (ultrasound, CT, or MRI).
[0203] Therefore, a viral infection can be determined, for example, by detecting viral DNA fragments, viral proteins, or other mechanisms. Cancer grading can be done even if the cancer is not present, but the subject is at higher risk than other subjects with a viral infection.
[0204] More generally, cancer grading can be categorized as the likelihood that an individual will develop cancer in the future. When the likelihood exceeds a threshold (determined by comparison), an individual can be monitored based on this likelihood exceeding the threshold. If an individual has a likelihood below the threshold, monitoring will be performed more frequently than tan. Examples of monitoring (also known as screening) are provided elsewhere in this article.
[0205] 3. Validation using different sequencing platforms
[0206] In another example, we validated our approach to cancer detection using cfDNA data with different sequencing platforms. Dataset C (Yu et al. Clin Chem. 2023;69(2):168-179) consists of single-molecule sequencing data (Oxford Nanopore) of plasma cfDNA from HBV subjects (n=6) and HCC patients (n=8), with a median of 9 million paired-end sequencing reads (range: 1.6 million to 31 million).
[0207] Figures 17A to 17B A set of graphs is shown to evaluate the performance of nucleosome scores in dataset C, generated using nanopore sequencing, for cancer diagnosis. Figure 17A The diagram shows the predicted HCC probabilities for samples in dataset C based on leave-one-out analysis using SVM. Figure 17A As shown, even though dataset C was prepared using nanopore sequencing technology (e.g., Oxford Nanopore Technologies), nucleosome scores remained effective in distinguishing HCC from non-HCC (P=0.0019, Wilcoxon rank-sum test). Figure 17BThe ROC analysis shows the predicted HCC probability in distinguishing HCC from non-HCC subjects from dataset C (where AUC is 1.0). This analysis is limited by the population size, as there are only 6 HBV and 8 HCC patients. However, we can see the promise of this method and the biomarkers.
[0208] 4. Utilizing the accuracy of a smaller number of segments
[0209] Downsampling analysis was performed on dataset A to determine the minimum number of fragments required for robust classification between individuals with and without HCC. As a result, using 5 million fragments achieved fairly accurate classification (AUC: 0.93 vs 0.92) compared to using all fragments (median: 38 million). The fragment numbers and AUC values were: all fragments (median: 38 million), AUC 0.931; 10 million fragments, AUC 0.9289; 5 million fragments, AUC 0.916; 2 million fragments, AUC 0.9091; 1 million fragments, AUC 0.8861; 500,000 fragments, AUC 0.7631; 200,000 fragments, AUC 0.7471; and 100,000 fragments, AUC 0.7283.
[0210] E. Performance of cancer classification under different settings
[0211] Different settings can be used, such as normalization or what constitutes a hypermethylated or hypomethylated site.
[0212] 1. Standardization
[0213] The universality of nucleosome signals across different sequencing protocols may be attributed to the data normalization steps used in this disclosure. To further demonstrate the impact of data normalization on classification analysis, we determined the AUC value corresponding to each normalization step.
[0214] Figures 18A to 18C A set of ROC curves is shown to evaluate the impact of data standardization on nucleosome scoring patterns and cancer detection. An SVM model trained on dataset B is tested on dataset A. Samples from dataset A are analyzed using the model trained on dataset B to evaluate the impact of data standardization on cancer diagnosis based on ROC analysis.
[0215] Figure 18A The analysis shows the results of cancer detection using nucleosome scoring. If a model trained on dataset B is used, the AUC is only 0.52 in distinguishing between subjects in dataset A who have and do not have HCC.
[0216] Figure 18BThe cancer detection analysis using nucleosome scores normalized with background CpG nucleosome scores is shown. The AUC for distinguishing HCC from non-HCC subjects rapidly improved to 0.75.
[0217] Figure 18C The results demonstrate a cancer detection analysis using a nucleosome score (referred to as the nucleosome z-score) standardized with background CpG nucleosome score and z-score statistics. The AUC distinguishing HCC from non-HCC subjects further increased to 0.86. These data indicate that standardization using a healthy sample distribution significantly improves the accuracy of the model's general applicability across platforms.
[0218] 2. Different methylation cutoff values define DMS
[0219] As mentioned above, hypermethylation sites and hypomethylation sites can be defined as the percentage of methylation density at that site relative to other tissue types or other sites. Hypermethylation sites and hypomethylation sites can be defined for a specific tissue type or more generally across all tissue types. HCC-specific hypermethylation sites and HCC-hypomethylation sites were used in dataset A. The effect of different cutoff values on the results was tested and shown.
[0220] We used two sets of cutoff values for analysis: cutoff values for <20% methylation in the leukocyte layer and >80% methylation in the tumor (tumor hypermethylation sites) and the opposite (tumor hypomethylation sites, i.e., <20% methylation in the tumor and >80% methylation in the leukocyte layer), and cutoff values for <10% methylation in the leukocyte layer and >90% methylation in the tumor and the opposite.
[0221] Figures 19A to 19C The ROC curves for HCC detection using nucleosome signal patterns based on differentially methylated sites (DMS), defined by different criteria for methylation levels between the leukocyte layer and tumor tissue, are shown. In dataset A, methylation cutoffs of 30% and 70% resulted in an AUC of 0.93 for cancer detection. When using methylation cutoffs with more stringent criteria, the AUCs for cancer detection were 0.90 (<20% in the leukocyte layer and >80% in the tumor, and <20% in the tumor and >80% in the leukocyte layer) and 0.80 (<10% in the leukocyte layer and >90% in the tumor, and <10% in the tumor and >90% in the leukocyte layer).
[0222] F. Results using random CpG sites
[0223] Conversely, this accurate prediction is not seen when using cfDNA nucleosome scoring based on CpG sites that do not show differential methylation patterns across tissues (e.g., a difference in methylation density of less than 5% between the leukocyte layer and HCC tumors).
[0224] When using a model trained on dataset B, we tested the use of random CpG sites from dataset A for cancer detection.
[0225] Figures 20A to 20B A set of graphs shows the performance of nucleosome scoring analysis for random CpG sites. Median subtraction was used for standardization. Figure 20A The probability of HCC is shown by inferring it using nucleosome z-scores around random CpG sites. Figure 20A As shown, there was no observable significant difference in nucleosome scores between subjects with and without HCC (P: 0.08, Wilcoxon rank-sum test), with an AUC of 0.62.
[0226] Figure 20B The performance of nucleosome scoring analysis based on random CpG sites for cancer detection is demonstrated. ROC analysis uses the predicted HCC probability to distinguish between HCC and non-HCC subjects. The model based on nucleosome scoring at random CpG sites failed to detect cancer regardless of whether it was unnormalized (AUC=0.53), normalized with background CpG (AUC=0.54), or normalized with both background CpG and z-score (AUC 0.62).
[0227] This data suggests that using fragmentation and methylation information from nucleosome signaling patterns (i.e., using signaling patterns at differentially methylated sites in cancer) provides improved accuracy.
[0228] G. Comparison with other technologies
[0229] We compared the use of nucleosome signaling patterns with other methods for cancer detection. Based on data from dataset A (Jiang et al., 2020), we evaluated three alternative methods: ichorCNA, short fragment ratios, and terminal motif patterns. IchorCNA is based on copy number aberrations. cfDNA fragments from cancer patients have been reported to tend to be shorter and exhibit greater terminal motif diversity. The ratios of fragments shorter than 150 bp and motif diversity scores (MDS) have been reported to provide information for cancer detection (Jiang et al., 2020).
[0230] Figure 21 The nucleosome signaling pattern (FRAGMA) for dataset A is shown. XRCompared to other techniques, the performance of FRAGMAXR 2210 (AUC=0.93) appears to outperform these methods (ichorCNA 2120 with an AUC of 0.70, short fragment ratio 2140 with an AUC of 0.82, and MDS 2130 with an AUC of 0.86).
[0231] H. Other cancer types
[0232] To demonstrate that this use of nucleosome signaling patterns can be applied to other cancer types, we conducted further analyses by applying nucleosome signaling patterns to detect different types of cancer, including lung cancer, ovarian cancer, and breast cancer. To this end, we identified differentially methylated sites (DMS) in these cancer types by comparing methylation levels between leukocytes and tumor tissues using bisulfite sequencing data generated by NovaSeq 6000 (Illumina).
[0233] We used the same criteria to define both types of DMS. In short, for type A DMS (tumor hypermethylated CpG sites), the methylation level of CpG in the leukocyte layer needed to be <30%, while the corresponding level in the tumor tissue of interest needed to be >70%. For type B DMS (tumor hypomethylated sites), the opposite requirements applied. As a result, we identified 45,848 type A DMS and 117,414 type B DMS in lung cancer; 133,123 type A DMS and 348,608 type B DMS in ovarian cancer; and 254,669 type A DMS and 588,133 type B DMS in breast cancer.
[0234] We further constructed and analyzed nucleosome patterns around DMS from publicly available cfDNA sequencing data from previous studies (Cristiano et al., Nature 570, 385-389, 2019), which included lung cancer (n=12), breast cancer (n=54), and ovarian cancer (n=28), as well as healthy individuals (n=245).
[0235] Figure 22A and Figure 22B The results of cancer detection using nucleosome signaling patterns are shown for lung cancer, breast cancer, and ovarian cancer. Figure 22A It displays the probability score of having cancer in datasets with different cancer types. Figure 22B ROC analysis using predicted cancer probability scores is shown. Using these nucleosome patterns, we demonstrated AUC values of 0.99, 0.81, and 0.93 for lung cancer 2210, breast cancer 2220, and ovarian cancer 2230, respectively.
[0236] We also analyzed the impact of an increased tumor proportion on the probability of different cancers. A higher probability score for cancer was typically observed in patients with a higher proportion of tumor DNA across all three cancer types.
[0237] Figures 23A to 23C This displays a probability score predicting the presence of cancer based on the proportion of tumor DNA in the plasma of cancer patients. Cancer patients are divided into three groups based on the tumor DNA proportion estimated using ichorCNA (copy number aberration-based).
[0238] These data truly demonstrate that the implementation scheme disclosed herein can be applied to multiple cancer detection.
[0239] To test whether the implementation scheme could distinguish between different cancer types, we trained a model for multi-cancer classification using plasma samples from cancer patients, including those with liver cancer (n=34), lung cancer (n=12), breast cancer (n=54), and ovarian cancer (n=28). Using nucleosome z-scores around DMS as input features, we employed a support vector machine (SVM)-based multi-class classification model to classify the different cancer types. Leave-one-out analysis showed that the SVM model correctly identified 79% of liver cancers (27 / 34), 83% of lung cancers (10 / 12), 70% of breast cancers (38 / 54), and 71% of ovarian cancers (20 / 28). Other model types, such as neural networks, decision trees, and other model types described in this paper, can also be used. This data demonstrates that the implementation scheme can distinguish between different cancers.
[0240] I. Methods of Cancer Detection
[0241] Figure 24 This is a flowchart illustrating method 2400 for analyzing biological samples to determine the cancer grade in the subject biological sample. Various examples of method 2400 are described above. Method 2400 can be performed partially or entirely using a computer system.
[0242] In box 2410, multiple free DNA molecules from the biological sample of the object are analyzed. Box 2410 can be performed in a similar manner to box 610 of method 600.
[0243] At box 2420, nucleosome signaling patterns in the genomic region surrounding the target CpG site are identified. Box 2420 can be performed in a similar manner to box 620 of method 600. The target CpG site may exhibit differential methylation relative to one or more other tissue types in the target tissue type (e.g., tissue of an organ), for example, as described above. As an example, the biological sample may be plasma or serum, and one or more other tissue types include blood cells.
[0244] In some implementations, nucleosome signaling patterns can be determined by clustering multiple genomic regions around multiple CpG sites. These multiple CpG sites may be hypermethylated relative to one or more other tissue types. These multiple CpG sites may also be hypomethylated relative to one or more other tissue types. In other examples, both may be used for their respective individual patterns.
[0245] At box 2430, the cancer grade of the object is determined based on a comparison of the nucleosome signal pattern with a reference pattern. Box 2430 can be performed in a similar manner to box 630 of method 600.
[0246] A reference pattern can be determined based on one or more training samples of cancers with known grades. Cancer grades can be determined for target tissue types. For example, cancer classification can detect liver cancer by selecting differentially methylated CpG sites in blood and liver tumors. Similarly, the model can detect colon cancer by selecting differentially methylated CpG sites in blood and colon tumors. Other example tissue types (e.g., lung, breast, etc.) are provided in this paper. Using all such patterns (including single-connection patterns), the implementation can provide a pan-cancer test. Such tissue types used to determine differentially methylated sites can be healthy or cancerous tissue.
[0247] This article describes an example of a reference pattern. Comparisons can be made in various ways, for example, using machine learning models such as SVM or other models described herein. Therefore, determining a cancer grade classification could involve feeding a nucleosome signal pattern into a machine learning model trained on a reference pattern using one or more training samples. As another example, the comparison of a nucleosome signal pattern with a reference pattern could use peak-valley distance, which could be a threshold representing the reference pattern.
[0248] For at least one of one or more training samples, the known cancer level can be that the object has no cancer.
[0249] In some implementations, the subject may have a viral infection, and the cancer grade classification may be that the cancer is absent, but the subject is at higher cancer risk than other subjects with a viral infection. For example, Figure 16 This indicates an HBV infection.
[0250] Cancer grading can be categorized as the likelihood (e.g., probability) that an individual will develop cancer in the future. In this case, the probability can be compared to a threshold. Based on the probability exceeding the threshold, the individual can be monitored. For example, when the probability exceeds the threshold, the individual can be monitored by performing screenings at a higher frequency than when the probability is below the threshold. Examples of monitoring are provided elsewhere, such as in the treatment and monitoring section.
[0251] According to some implementation schemes, one or more advantages of using nucleosome signaling patterns are that bisulfite sequencing is not required, thus enabling methylation-sensing fragmentomics signals from non-bisulfite sequencing data of free DNA molecules.
[0252] IV. Determining the proportion of DNA from a specific tissue type
[0253] The proportion of cell-free DNA from a specific tissue type can also be determined using nucleosome signal patterns. For example, nucleosome signal patterns at tissue-specific sites (e.g., differential methylation at a particular tissue type relative to other tissue types in a cell-free sample) can be used as input to an ML model that outputs the proportion of tissue. The ML model can be trained using training (calibration) samples with known proportions of tissue types. Such measurements can be used for a variety of purposes. For example, the proportion of fetal DNA in a pregnant woman can be used for non-invasive prenatal testing (NIPT).
[0254] Tissue types can be clinically relevant, such as tissue types not typically found in the recipient's body, like fetal tissue, tumor tissue, or transplanted tissue. For example, the proportion of liver DNA fragments in a recipient who has undergone a liver transplant can be determined. Other transplanted organs can be analyzed in a similar manner.
[0255] A. Fetal ratio
[0256] We further explored whether nucleosome patterns could be used to reflect DNA contributions from specific tissues in plasma. Pregnancy models were used to address this issue. Because placental tissue contains a large number of unique alleles present in the placental tissue but absent in the background maternal genome, placental contribution can be directly inferred using genotypic information between the fetal and maternal genomes, providing a gold standard for evaluating nucleosome-based methods for inferring placental contribution.
[0257] The nucleosome signal patterns we show at tissue-specific CpG sites can be used to accurately determine the proportion of tissue in a specific tissue type. This example uses fetal DNA to quantify the proportion of fetal DNA, but this analysis is equally applicable to other tissue types.
[0258] 1. Using nucleosome scoring
[0259] Figure 25A and Figure 25B Nucleosome signaling patterns at placental-specific CpG sites were observed: hypermethylation and hypomethylation. Specifically, Figure 25ANucleosome scores were displayed around placental-specific hypermethylated CpG sites (methylation levels <30% in the leukocyte layer and >70% in the placenta) using non-bisulfite cfDNA data from pregnant women. However, other nucleosome signaling patterns can also be used. Figure 25B Nucleosome scores were displayed around placental-specific hypomethylated CpG sites (methylation levels <30% in the placenta and >70% in the leukocyte layer). Similar to the reduced nucleosome amplitude in cancer patients with higher ctDNA, we observed attenuated amplitudes in individuals with 16% fetal DNA compared to individuals with 8% fetal DNA.
[0260] Figure 26A This shows the average peak-to-trough distance of hypermethylation sites in different fetal proportions. Figure 26B The average peak-to-trough distance of hypomethylation sites is shown for different fetal proportions. As shown in the figure, for hypermethylation (… P < 0.001, Wilcoxon rank-sum test; Figure 26A ) and hypomethylation ( P < 0.001, Wilcoxon rank-sum test; Figure 26B At the CpG site, individuals with a high proportion of fetal DNA (>10%) showed a reduction in nucleosome amplitude compared to individuals with a low proportion of fetal DNA (<10%).
[0261] Based on the ridge regression model, we used nucleosome score as the independent variable and fetal DNA score as the dependent variable to construct a linear regression function for predicting the proportion of fetal DNA.
[0262] Figures 26C to 26D The results showed a strong correlation between the proportion of fetal DNA predicted by nucleosome patterning and the proportion of actual fetal DNA inferred by single nucleotide polymorphism (SNP)-based methods, with Pearson correlations of 0.91 and 0.85 for the training and test sets, respectively. These results indicate that nucleosome patterning analysis can be used to infer the contribution of tissues to the plasma DNA pool. Other techniques can be used to determine tissue-specific proportions (e.g., fetal, tumor, or transplant), such as using the Y chromosome of male fetuses or tissue-specific methylation patterns.
[0263] Data point 2620 can correspond to the calibration data point of the calibration sample. Lines 2630 and 2640 correspond to the calibration curve (function) that can be determined from the calibration data point, for example, by performing linear or nonlinear regression. The tissue proportion of a new sample can be determined by comparing the measured nucleosome signal pattern with a reference pattern of any calibration sample. For example, the measured peak-to-trough distance can be compared with the calibration value (the peak-to-trough distance of the calibration sample), or the measured peak-to-trough distance can be input into the calibration curve.
[0264] 2. Using nucleosome z-score
[0265] We also show that nucleosome signal patterns at tissue-specific CpG sites can be used to accurately determine the proportion of tissue in a specific tissue type. This example uses fetal DNA to quantify the proportion of fetal DNA, but this analysis is equally applicable to other tissue types.
[0266] Figure 27A Nucleosome signaling patterns at placental-specific CpG sites are observed: hypomethylation and hypermethylation. Specifically, Figure 27A The z-scores of nucleosomes around placental-specific CpG sites are shown using cfDNA data from pregnant women 2710 and non-pregnant controls 2720 from dataset A, but other nucleosome signaling patterns can also be used.
[0267] Based on the Lasso regression model, we constructed a linear regression function using the nucleosome z-score as the independent variable and the fetal DNA proportion as the dependent variable to predict the fetal DNA proportion. To determine the nucleosome z-score, the mean was the average of 19 non-pregnant healthy controls.
[0268] Figures 27B to 27C This shows a comparison between the actual and predicted proportions of fetal DNA in the training and test sets. The actual proportions are inferred from single nucleotide polymorphism (SNP) analysis. Figure 27B The results showed that in a training set including 20 pregnant and 14 non-pregnant subjects, the proportion of fetal DNA predicted by nucleosome z-score was well correlated with the actual proportion of fetuses inferred by the SNP-based method (Pearson r: 0.98). Figure 27C The regression function was well validated in an independent test set comprising 10 pregnant and 5 non-pregnant subjects, with a Pearson r of 0.98 between the predicted and actual fetal DNA proportions. The results indicate that nucleosome signal analysis according to the embodiments of this disclosure can be used for non-invasive prenatal testing and to infer DNA contributions from various tissues.
[0269] For the distribution used to determine the z-score, the reference sample may not contain DNA fragments from the first tissue type. For example, the reference sample could be from a non-pregnant subject, and the first tissue type could be fetal / placental DNA. In other examples, the reference sample may contain DNA fragments from the first tissue type, but at normal proportions. For example, the reference sample could contain DNA fragments from the liver, and the subject being tested could have undergone a liver transplant, in which elevated levels of liver DNA fragments were detected.
[0270] A regression model using measured nucleosome signal patterns is one example of a machine learning model. Other example models include support vector machine regression, neural networks, or other examples described herein. This type of regression model is an example of a calibration function. Therefore, measured nucleosome signal patterns can be input as feature vectors into a machine learning model.
[0271] B. Graft ratio
[0272] It can also determine, for example, the proportion of donor DNA in liver transplant recipients or other organ transplant recipients.
[0273] We constructed nucleosome patterns of liver-specific hypermethylated sites (methylation level <30% in leukocyte layer and >70% in liver tissue) and hypomethylated sites (methylation level <30% in liver tissue and >70% in leukocyte layer) CpG sites and examined whether they could be used to estimate the proportion of donor-derived DNA in liver transplant recipients. Based on a ridge regression model, we used nucleosome scores as independent variables and the proportion of donor-derived DNA as the dependent variable to construct a linear regression function for predicting the proportion of donor-derived DNA.
[0274] Figures 28A to 28B The results showed a good correlation between the proportion of donor-origin DNA predicted by nucleosome patterns and the proportion of actual donor-origin DNA inferred by single nucleotide polymorphism analysis, with Pearson correlations of 0.99 and 0.97 for the training and test sets, respectively.
[0275] C. Tumor Ratio
[0276] It is also possible to determine, for example, the proportion of tumor-derived DNA from liver cancer (e.g., HCC). It is also possible to determine the proportion of DNA from other types of tumors, such as those described herein.
[0277] We used a ridge regression model to measure the proportion of circulating tumor DNA in plasma samples from dataset A (Jiang et al., 2020). We used nucleosome scores to build the ridge regression model to predict the proportion of tumor DNA in plasma samples from HCC patients (n=34).
[0278] Figures 29A to 29B The results showed a good correlation between the proportion of tumor-derived DNA predicted by nucleosome patterns and the proportion of actual tumor-derived DNA inferred by ichorCNA, with Pearson correlations of 0.96 and 0.94 for the training and test sets, respectively. Figures 29A to 29B As shown, the proportion of tumor DNA predicted by nucleosome patterns correlates well with those inferred by copy number aberration-based ichorCNA (Adalsteinsson et al., 2017), with Pearson correlations of 0.96 and 0.94 for the training and test sets, respectively. These results suggest that nucleosome pattern analysis can be used to infer the contribution of tissues to the plasma DNA pool.
[0279] D. Methods for determining organizational proportions
[0280] Figure 30 This is a flowchart illustrating a method for determining the proportional concentration of DNA from a first tissue type in a biological sample. Various examples of method 3000 are described above. Method 3000 can be performed partially or entirely using a computer system.
[0281] At box 3010, multiple free DNA molecules from a biological sample of the object are analyzed. Box 3010 can be performed in a manner similar to box 610 of method 600 and box 2410 of method 2400.
[0282] At box 3020, nucleosome signaling patterns in the genomic region surrounding the target CpG site are identified. Box 3020 can be performed in a manner similar to box 620 of method 600 and box 2420 of method 2400. The target CpG site may exhibit differential methylation relative to one or more other tissue types in the target tissue type, as described above. As an example, the biological sample may be plasma or serum, and one or more other tissue types include blood cells.
[0283] In some implementations, nucleosome signaling patterns can be determined by clustering multiple genomic regions around multiple CpG sites. These multiple CpG sites may be hypermethylated relative to one or more other tissue types. These multiple CpG sites may also be hypomethylated relative to one or more other tissue types. In other examples, both may be used for their respective individual patterns.
[0284] At box 3030, the proportional concentration of DNA from a first tissue type in the biological sample is determined by comparing the nucleosome signal pattern with a reference pattern. Box 3030 can be performed in a similar manner to box 630 of method 600 and box 2430 of method 2400. The reference pattern can be determined based on one or more calibration samples having known proportional concentrations of DNA from the first tissue type.
[0285] A reference pattern can be determined based on one or more calibration samples having a known proportionate concentration of DNA from a first tissue type (also known as tissue proportion). The proportionate concentration of DNA can have various resolutions, for example, above or below a specific percentage or within a range. Such ranges can have various resolutions, such as 5%, 10%, 15%, 20%, 25%, and 30% or lower. As an example, a 5% resolution could correspond to 35% to 40%.
[0286] The tissue proportions of a new sample can be determined by comparing it with a reference pattern of any calibration sample or as input to a calibration curve (e.g., using peak-to-valley distance). When one or more calibration samples are multiple calibration samples, comparing the nucleosome signal pattern with the reference pattern can include inputting the peak-to-valley distance of the nucleosome signal pattern into the calibration function. The calibration function can be determined by measuring the proportional concentrations of DNA from a first tissue type in multiple calibration samples and measuring the peak-to-valley distances of multiple calibration samples, thereby determining calibration data points that include proportional concentrations and peak-to-valley distances. The calibration function can then be fitted to the calibration data points, for example, using regression.
[0287] In other examples, determining the proportional concentration of DNA from a first tissue type may include using each nucleosome signal, for example, by comparing it with individual signals of a reference pattern, as may happen when using a machine learning model. Such nucleosome signal values can be used in the model's feature vector. Therefore, each signal value can be compared to a reference pattern by inputting the nucleosome signal pattern as part of the feature vector into a machine learning model (e.g., a regression model or other types described herein). Thus, determining the proportional concentration of DNA from a first tissue type may include inputting the nucleosome signal pattern into a machine learning model trained using a reference pattern with multiple calibration samples.
[0288] Objects and primary tissue types can take many forms. For example, an object can be pregnant with a fetus, and the primary tissue can be fetal tissue. As another example, the primary tissue type can originate from a specific organ.
[0289] As a further step, the cancer grade can be determined in the first tissue type of the subject based on the proportion of concentration.
[0290] V. Treatment methods
[0291] A. Further screening methods
[0292] Based on any classification, such as the proportion of lesion or clinically relevant DNA concentrations, individuals may be referred to additional screening methods, such as chest X-ray, ultrasound, computed tomography, magnetic resonance imaging, or positron emission tomography. Such screening can be performed for cancer.
[0293] B. Treatment options
[0294] The embodiments of this disclosure can accurately predict disease recurrence, thereby facilitating early intervention and selection of appropriate treatments to improve disease outcomes and overall survival in subjects. For example, enhanced chemotherapy can be selected for subjects whose corresponding samples predict disease recurrence. In another example, biological samples from subjects who have completed initial treatment can be sequenced to identify viral DNA that predicts disease recurrence. In such examples, alternative treatment regimens (e.g., higher doses) and / or different treatments can be selected for subjects because their cancer may have developed resistance to the initial treatment.
[0295] The implementation plan may also include treating the subject in response to a classification indicating lesion recurrence. For example, if the prediction corresponds to local failure, surgery may be selected as a possible treatment. In another example, if the prediction corresponds to distant metastasis, chemotherapy may be additionally selected as a possible treatment. In some implementation plans, treatments include surgery, radiation therapy, chemotherapy, immunotherapy, targeted therapy, hormone therapy, stem cell therapy, or precision medicine. To reduce the risk of harm to the subject and increase overall survival, a treatment plan may be developed based on the determined recurrence classification. The implementation plan may further include treating the subject according to the treatment plan.
[0296] C. Treatment type
[0297] The implementation plan may further include treating the patient's lesion after the classification of the subject has been determined. Treatment may be provided based on the determined lesion grade, the proportionate concentration of clinically relevant DNA, or the tissue of origin. For example, specific drugs or chemotherapy may be used to target the identified mutation. The tissue of origin may be used to guide surgery or any other form of treatment. Furthermore, the lesion grade can be used to determine the degree of invasiveness when using any type of treatment, and the degree of invasiveness may also be determined based on the lesion grade. Lesions (e.g., cancer) may be treated with chemotherapy, drugs, diet, therapy, and / or surgery. In some implementation plans, the more a parameter (e.g., amount or size) deviates from a reference value, the more invasive the treatment may be.
[0298] Treatment may include resection. For bladder cancer, treatment may include transurethral resection of bladder tumor (TURBT). This procedure is used for diagnosis, classification, and treatment. During TURBT, a surgeon inserts a cystoscope into the bladder through the urethra. The tumor is then removed using instruments with small wire loops, lasers, or high-energy electricity. For patients with non-muscle-invasive bladder cancer (NMIBC), TURBT can be used to treat or eliminate the cancer. Another treatment may include radical cystectomy and lymph node dissection. Radical cystectomy involves removing the entire bladder and any surrounding tissues and organs. Treatment may also include urinary diversion. Urinary diversion is when, during the removal of the bladder as part of treatment, a physician creates a new pathway for urine to drain out of the body.
[0299] Treatment may include chemotherapy, which uses drugs that typically destroy cancer cells by preventing them from growing and dividing. Drugs may involve, for example (but not limited to), mitomycin-C (available as a general medication), gemcitabine (Gemzar), and thiotepa (Tepadina) for intrabladder chemotherapy. Systemic chemotherapy may involve, for example (but not limited to), cisplatin, gemcitabine, methotrexate (Rheumatrex, Trexall), vinblastine (Velban), doxorubicin, and cisplatin.
[0300] In some implementations, treatment may include immunotherapy. Immunotherapy may include immune checkpoint inhibitors that block a protein called PD-1. Inhibitors may include (but are not limited to) atezolizumab (Tecentriq), nivolumab (Opdivo), avelumab (Bavencio), durvalumab (Imfinzi), and pembrolizumab (Keytruda).
[0301] Treatment implementation plans may also include targeted therapies. Targeted therapies are treatments that target specific genes in the cancer and / or proteins that contribute to cancer growth and survival. For example, erdafitinib is an orally administered drug approved for the treatment of people with locally advanced or metastatic urothelial carcinoma (with FGFR3 or FGFR2 genetic mutations) that has been continuously growing or spreading cancer cells.
[0302] Some treatments may include radiation therapy. Radiation therapy uses high-energy X-rays or other particles to destroy cancer cells. In addition to each individual treatment, combinations of the treatments described herein may also be used. In some implementations, combinations of treatments may be used when the value of a parameter exceeds a threshold (which itself exceeds a reference value). Information regarding treatments in the references is incorporated herein by reference.
[0303] VI. Example System
[0304] Figure 31 A measurement system 3100 according to one embodiment of the present disclosure is shown. The system includes a sample 3105, such as a free DNA molecule (e.g., DNA and / or RNA), within an analytical apparatus 3110, through which an analysis 3108 can be performed on the sample 3105. For example, the sample 3105 can be contacted with reagents of the analysis 3108 to provide a signal of physical characteristics 3115 (e.g., sequence information of a free nucleic acid molecule). An example of the analytical apparatus may be a flow channel comprising probes and / or primers or droplets through which the analysis moves (where the droplets comprise the analysis). Physical characteristics 3415 of the sample (e.g., fluorescence intensity, voltage, or current) are detected by a detector 3120. The detector 3120 can perform measurements at time intervals (e.g., periodic time intervals) to obtain data points constituting a data signal. In one embodiment, an analog-to-digital converter converts the analog signal from the detector into digital form at multiple times.
[0305] The analysis device 3110 and detector 3120 can form an analysis system, such as a sequencing system that performs sequencing according to the embodiments described herein. Data signal 3125 is transmitted from detector 3120 to logic system 3130. As an example, data signal 3125 can be used to determine the sequence and / or location of a nucleic acid molecule (e.g., DNA and / or RNA) in a reference genome. Data signal 3125 can include various measurements taken simultaneously, such as different colors of fluorescent dyes or different electrical signals for different molecules in sample 3105, and therefore data signal 3125 can correspond to multiple signals. Data signal 3125 can be stored in local memory 3135, external memory 3140, or storage device 3145. The analysis system can consist of multiple analysis devices and detectors.
[0306] As an example applicable to any of the above implementation schemes, cell-free DNA (e.g., plasma, DNA) samples can be processed and sequenced as follows: DNA is extracted from 1.6 mL of plasma samples using the QIAamp Cyclic Nucleic Acid Kit (QIAGEN) according to the manufacturer's protocol. Plasma DNA index libraries are constructed using the TruSeq NanoDNA Library Preparation Kit (Illumina) and xGen UDI-UMI aptamers (Integrated DNA Technologies) according to the manufacturer's instructions. The multiplexed DNA library is sequenced on a NovaSeq 6000 platform (Illumina) using paired-end mode (75 bp × 2). Sequence read alignment is performed using Short Oligonucleotide Alignment Procedure 2 (SOPAP2)(21), and paired-end reads are filtered to remove inconsistent and duplicate reads. Genomic DNA is extracted from 200 μL to 400 μL of maternal leukocyte layer samples using the QIAamp DNA Blood Mini Kit (QIAGEN) according to the manufacturer's protocol. Maternal genotype information was obtained from microarray analysis using an Infinium Omni 2.5 BeadChip (Illumina). Fetal QuantSD (22) was used to determine the proportion of fetal DNA.
[0307] The logic system 3130 may be a computer system, ASIC, microprocessor, graphics processing unit (GPU), etc., or may include a computer system, ASIC, microprocessor, graphics processing unit (GPU), etc. It may also include a display (e.g., a monitor, LED display, etc.) and user input devices (e.g., a mouse, keyboard, buttons, etc.), or be connected to a display and user input devices. The logic system 3130 and other components may be part of a stand-alone or network-connected computer system, or may be directly connected to or incorporated into a device including detector 3120 and / or analysis device 3110 (e.g., a sequencing device). The logic system 3130 may also include software executing in processor 3150. The logic system 3130 may include a computer-readable medium storing instructions for controlling the measurement system 3100 to perform any of the methods described herein. For example, the logic system 3130 may provide commands to a system including analysis device 3110 to perform sequencing or other physical operations. Such physical operations may be performed in a specific order, such as adding and removing reagents in a specific order. Such physical operations can be performed using robotic systems (e.g., robotic arms) that can be used to obtain samples and perform analysis.
[0308] The measurement system 3100 may also include a treatment device 3160 capable of providing treatment to an object. The treatment device 3160 may determine treatment and / or be used to perform treatment. Examples of such treatments may include surgery, radiation therapy, chemotherapy, immunotherapy, targeted therapy, hormone therapy, and stem cell transplantation. A logic system 3130 may be connected to the treatment device 3160, for example, to provide results of the methods described herein. The treatment device may receive input from, for example, imaging devices and other devices for user input (e.g., to control treatment, such as controlling a robotic system).
[0309] Any computer system mentioned in this article can utilize any suitable number of subsystems. Examples of such subsystems are shown in [the document / examples]. Figure 32 In computer system 10. In some embodiments, the computer system includes a single computer device, wherein a subsystem may be a component of the computer device. In other embodiments, the computer system may include multiple computer devices having internal components, each of which is a subsystem. The computer system may include desktop and laptop computers, tablet computers, mobile phones, and other mobile devices.
[0310] Shown Figure 32 The subsystems in the display are interconnected via system bus 75. Additional subsystems, such as printer 74, keyboard 78, one or more storage devices 79, monitor 76 (e.g., display screen, such as LED) connected to display adapter 82, and other devices, are displayed. Peripheral devices and I / O devices connected to input / output (I / O) controller 71 can be connected to the computer system via a variety of devices known in the art, such as input / output (I / O) port 77 (e.g., USB, FireWire®). For example, I / O port 77 or external interface 81 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect computer system 10 to a wide area network (e.g., the Internet), a mouse input device, or a scanner. The interconnection via system bus 75 allows central processing unit 73 to exchange information with each subsystem and control the execution of multiple instructions from system memory 72 or storage device 79 (e.g., fixed disk, such as hard disk drive, or optical disk), as well as information exchange between subsystems. System memory 72 and / or storage device 79 may include computer-readable media. Another subsystem is the data collection device 85, such as a camera, microphone, accelerometer, etc. Any data mentioned herein can be output from one component to another and can also be output to the user.
[0311] A computer system may include multiple identical components or subsystems, connected together, for example, via an external interface 81, an internal interface, or via a removable storage device that can be connected to and moved from one component to another. In some embodiments, the computer system, subsystem, or device may communicate via a network. In such cases, one computer may be considered a client and another a server, where each computer may be part of the same computer system. Clients and servers may each include multiple systems, subsystems, or components. In various embodiments, the method may involve a variety of numbers of clients and / or servers, including at least 10, 20, 50, 100, 200, 500, 1,000, or 10,000 devices. The method may include a variety of numbers of communication messages between devices, including at least 100, 200, 500, 1,000, 10,000, 50,000, 100,000, 500,000, or one million communication messages. Such communications can involve at least 1 MB, 10 MB, 100 MB, 1 GB, 10 GB, or 100 GB of data.
[0312] Aspects of the implementation may be implemented using hardware circuitry (e.g., application-specific integrated circuits or field-programmable gate arrays) and / or using computer software stored in a modular or integrated manner in memory with a general-purpose programmable processor in the form of logical control, and thus the processor may include memory storing software instructions for configuring the hardware circuitry and an FPGA or ASIC having configuration instructions. As used herein, the processor may include a single-core processor, a multi-core processor on the same integrated chip, or single board or network hardware, and multiple processing units on dedicated hardware. Based on this disclosure and the teachings provided herein, those skilled in the art will appreciate other ways and / or methods of implementing embodiments of this disclosure using hardware and combinations of hardware and software.
[0313] Any software component or function described in this application may be implemented in the form of software code using, for example, conventional or object-oriented techniques. This software code is executed by a processor using any suitable computer language (e.g., Java, C, C++, C#, Objective-C, Swift) or scripting language (e.g., Perl or Python). The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission. Suitable non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), magnetic media (e.g., hard disk or floppy disk) or optical media (e.g., closed-disk CD, digital optical disc (DVD), or Blu-ray disc), flash memory, and similar devices. The computer-readable medium may be any combination of such devices. Additionally, the order of operations may be reconfigured. A process may terminate upon completion of its operations, but may have additional steps not included in the figure. A process may correspond to a method, function, program, subroutine, subroutines, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.
[0314] Such programs can also be encoded and transmitted using carrier signals suitable for transmission over wired, optical, and / or wireless networks (including the Internet) conforming to various protocols. Therefore, computer-readable media can be constructed using data signals encoded by such programs. Computer-readable media encoded with program code (e.g., as software) can be packaged with compatible devices or provided separately from other devices (e.g., for download via the Internet). Any such computer-readable media can reside on or within a single computer product (e.g., a hard disk drive, CD, or an entire computer system) and can reside on or within different computer products within a system or network. Computer systems may include monitors, printers, or other suitable displays for providing users with any of the results mentioned herein.
[0315] Any method described herein can be performed wholly or partially using a computer system including one or more processors configured to perform the steps described. Any operation performed via the processor (e.g., comparison, determination, contrast, computer calculation, computation) can be performed in real time. The term "..." real time"This can refer to a computational operation or process completed within a certain time limit. The time limit can be 1 minute, 1 hour, 1 day, or 7 days. Therefore, an implementation can refer to a computer system configured to perform the steps of any method described herein, potentially having different components that perform the respective steps or groups of steps. Although presented in the form of numbered steps, the steps of the methods herein can be performed simultaneously or at different times or in different orders. Furthermore, a portion of these steps can be used in conjunction with a portion of other steps of other methods. Additionally, all or some steps can be optional. Moreover, any step of any method can be performed using modules, units, circuits, or other means of a system for performing these steps."
[0316] The specified details of a particular embodiment may be combined in any suitable manner without departing from the spirit and scope of the embodiments of this disclosure. However, other embodiments of this disclosure may refer to specific embodiments relating to each individual aspect or a particular combination of these individual aspects.
[0317] The foregoing description of exemplary embodiments of this disclosure has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit this disclosure to the precise forms described, and many modifications and variations are possible in light of the foregoing teachings.
[0318] Unless explicitly stated otherwise, the use of “a,” “an,” or “the” is intended to mean “one or more.” Unless explicitly stated otherwise, the use of “or” is intended to mean “inclusive of or” rather than “exclusive of or.” Referring to the “first” component does not necessarily require the provision of the second component. Furthermore, unless explicitly stated otherwise, referring to the “first” or “second” component does not limit the referenced component to a particular location. The term “based on” is intended to mean “at least partially based on.”
[0319] The claims are drafted to exclude any optional elements. Therefore, this statement is intended to serve as a basis for the use of exclusive terms such as “solely,” “only,” and similar terms, or for the use of negative limitations, in conjunction with the description of the claimed elements.
[0320] All patents, patent applications, publications, and specifications mentioned herein are incorporated herein by reference in their entirety for all purposes. No prior art is acknowledged. In the event of any conflict between this application and the references provided herein, this application shall prevail.
Claims
1. A method for analyzing a biological sample to determine the cancer grade in the biological sample of an object, the biological sample containing cell-free DNA, the method comprising: Analyzing a plurality of cell-free DNA molecules from a biological sample of the object, wherein analyzing each of the plurality of cell-free DNA molecules includes determining two genomic locations in a reference genome corresponding to the two ends of the cell-free DNA molecule; The following steps were used to determine the nucleosome signaling patterns in the genomic region surrounding the target CpG site: For each genomic location within the genomic region: A first amount of multiple free DNA molecules spanning a window around the said genomic location is determined, the window being 2 bp or longer; A second quantity of multiple free DNA molecules with ends located within the window surrounding the said genomic location is determined; and The first and second quantities are used to determine nucleosome signals at each location, wherein the target CpG sites are differentially methylated relative to one or more other tissue types in the target tissue type; and The cancer grade of the object is determined by comparing the nucleosome signal pattern with a reference pattern, wherein the reference pattern is determined based on one or more training samples with known cancer grades, and wherein the cancer grade is determined for the target tissue type.
2. The method of claim 1, wherein for at least one of the one or more training samples, the known cancer grade is that the subject does not have the cancer.
3. The method of claim 1, wherein determining the cancer grade classification of the object comprises inputting the nucleosome signal pattern into a machine learning model trained using the reference pattern of the one or more training samples.
4. The method of claim 1, wherein the subject has a viral infection, and wherein the cancer grade classification is that cancer is not present, but the subject is at a higher risk of cancer than other subjects with the viral infection.
5. The method of claim 1, wherein the cancer grade classification is the likelihood that the subject will develop cancer in the future.
6. The method of claim 5, further comprising: Compare the probability to a threshold; and The object is monitored based on the probability that it exceeds the threshold.
7. The method of claim 6, wherein when the probability exceeds the threshold, the object is monitored by screening at a higher frequency than when the probability is less than the threshold.
8. The method of claim 1, wherein the target tissue type is an organ tissue type.
9. A method for measuring the proportional concentration of DNA of a first tissue type in a biological sample from an object, the biological sample comprising cell-free DNA, the method comprising: Analyzing a plurality of cell-free DNA molecules from a biological sample of the object, wherein analyzing each of the plurality of cell-free DNA molecules includes determining two genomic locations in a reference genome corresponding to the two ends of the cell-free DNA molecule; The following steps were used to determine the nucleosome signaling patterns in the genomic region surrounding the target CpG site: For each genomic location within the genomic region: A first amount of multiple free DNA molecules spanning a window around the said genomic location is determined, the window being 2 bp or longer; A second quantity of multiple free DNA molecules with ends located within the window surrounding the said genomic location is determined; and The first and second quantities are used to determine nucleosome signals at each location, wherein the target CpG sites in the biological sample are differentially methylated relative to one or more other tissue types in the first tissue type; and The proportionate concentration of DNA from the first tissue type in the biological sample is determined by comparing the nucleosome signal pattern with a reference pattern, wherein the reference pattern is determined by one or more calibration samples having a known proportionate concentration of DNA from the first tissue type.
10. The method of claim 9, wherein the object is pregnant with a fetus, and wherein the first tissue type is fetal tissue.
11. The method of claim 9, wherein the first tissue type originates from a specific organ.
12. The method of claim 9, wherein the target CpG site in the biological sample is differentially methylated relative to one or more other tissue types in the first tissue type by having a difference of at least 30% methylation level.
13. The method of claim 9, further comprising: The cancer grade of the object in the first tissue type is determined based on the proportion and concentration of DNA from the first tissue type in the biological sample.
14. The method of any one of claims 9 to 13, wherein the one or more calibration samples are a plurality of calibration samples, wherein determining the proportional concentration of DNA from the first tissue type comprises inputting the nucleosome signal pattern into a machine learning model trained using a reference pattern of the plurality of calibration samples.
15. The method of any one of claims 9 to 13, wherein the one or more calibration samples are multiple calibration samples, and wherein comparing the nucleosome signal pattern with the reference pattern includes inputting the peak-valley distance of the nucleosome signal pattern into a calibration function.
16. The method of claim 15, wherein the calibration function is determined by the following steps: Measure the proportional concentration of DNA from the first tissue type in the plurality of calibration samples; Measure the peak-to-valley distance of the plurality of calibration samples to determine calibration data points that include the proportional concentration and the peak-to-valley distance; and The calibration function is fitted to the calibration data points.
17. The method according to any one of claims 1 to 13, wherein the comparison of the nucleosome signal pattern with the reference pattern uses peak-valley distance.
18. The method of any one of claims 1 to 17, wherein the biological sample is plasma or serum, and wherein one or more other tissue types include blood cells.
19. The method of any one of claims 1 to 18, wherein the nucleosome signaling pattern is determined by clustering multiple genomic regions around multiple CpG sites.
20. The method of claim 19, wherein the plurality of CpG sites are hypermethylated relative to the one or more other tissue types.
21. The method of claim 19, wherein the plurality of CpG sites are hypomethylated relative to the one or more other tissue types.
22. The method of any one of claims 1 to 21, wherein the length of the genomic region is at least 140 bp.
23. A method for measuring methylation of target CpG sites in the genome of a target organism using cell-free DNA molecules, the method comprising: Analyzing a plurality of cell-free DNA molecules from a biological sample of the object, wherein analyzing each of the plurality of cell-free DNA molecules includes determining two genomic locations in a reference genome corresponding to the two ends of the cell-free DNA molecule; The following steps were used to determine the nucleosome signaling patterns in the genomic region surrounding the target CpG site: For each genomic location within the genomic region: A first amount of multiple free DNA molecules spanning a window around the said genomic location is determined, the window being 2 bp or longer; A second quantity of multiple free DNA molecules with ends located within the window surrounding the said genomic location is determined; and The first and second quantities are used to determine nucleosome signals at each location, wherein the length of the genomic region is at least 140 bp; and The methylation level of the target CpG site in the genome of the object is determined by comparing the nucleosome signaling pattern with a reference pattern, wherein the reference pattern is determined based on one or more training samples with known methylation levels.
24. The method of claim 23, wherein the methylation level is either high or low methylation of the target CpG site.
25. The method of claim 23, wherein the methylation is a range of methylation density at the target CpG site.
26. The method of claim 23, wherein the biological sample is plasma or serum, wherein differential methylation of the target CpG site is known to exist between a first tissue type and blood cells, and wherein the methylation level of the target CpG site in the first tissue type is determined.
27. The method of claim 26, wherein the first tissue type is cancerous tissue.
28. The method of claim 26, wherein the first tissue type is fetal tissue or a specific organ.
29. The method of claim 26, wherein the target CpG site is known to be differentially methylated between the first tissue type and the blood cells by having at least 30% methylation difference.
30. The method of claim 23, wherein the target CpG site is located at the center of the genomic region.
31. The method of any of the preceding claims, wherein the nucleosome signal at each position comprises the difference or ratio between the first amount and the second amount.
32. The method of any one of claims 1 to 31, wherein the nucleosome signals at each location are normalized using nucleosome signals from other genomic regions.
33. The method of claim 32, wherein the normalization uses the mean or median of nucleosome signals at each location derived from the other genomic regions.
34. The method of claim 32 or claim 33, wherein the other genomic region is adjacent to the genomic region.
35. The method of any one of claims 1 to 31, wherein the nucleosome signal pattern is standardized using regional-level statistics determined based on one or more regions, resulting in a nucleosome score for each location of the nucleosome signal.
36. The method of claim 35, wherein the one or more regions include the genomic region or one or more other regions.
37. The method of claim 36, wherein the one or more regions include the one or more other regions adjacent to the genomic region.
38. The method of claim 35, wherein the regional statistic is the mean or median.
39. The method of any of the preceding claims, wherein the nucleosome signal at each location is normalized using a background signal that depends on the distance between the genomic location and the target CpG site.
40. The method of claim 39, wherein the background signal is determined based on one or more other genomic regions, each region centered on a CpG site.
41. The method of claim 40, wherein the one or more regions are randomly selected.
42. The method as claimed in any of the preceding claims, wherein the nucleosome signals at each location are normalized using a distribution of reference nucleosome signals determined according to one or more regions of one or more reference samples.
43. The method of claim 42, wherein standardization using the distribution comprises: For each location, determine the total statistics and dispersion of the reference nucleosome signal distribution; and Subtract the total statistic from each of the nucleosome signals at each location and divide by the dispersion value.
44. The method of claim 43, wherein the distribution is a normal distribution, wherein the total statistic is the average value of the reference nucleosome signal at the location, and wherein the dispersion value is the standard deviation.
45. The method as described in any of the preceding claims, wherein the length of the window is at least 5 bp.
46. The method of claim 45, wherein the length of the window is at least 50 bp.
47. The method of claim 46, wherein the length of the window is at least 100 bp.
48. The method of any of the preceding claims, wherein analyzing the plurality of cell-free DNA molecules comprises sequencing the plurality of cell-free DNA molecules to obtain sequence reads.
49. The method of claim 48, wherein determining the two genomic positions in the reference genome corresponding to the two ends of the free DNA molecule comprises aligning the sequence readings with the reference genome.
50. A computer product comprising a non-transitory computer-readable medium storing a plurality of instructions, which, when executed, cause a computer system to perform the method as described in any of the preceding claims.
51. A system comprising: The computer product as described in claim 50; and One or more processors configured to execute instructions stored on the computer-readable medium.
52. A system comprising means for performing any of the methods described above.
53. A system comprising one or more processors configured to perform any of the methods described above.
54. A system comprising modules that respectively perform the steps of any of the methods described above.