Method and system for identifying source of variation
By generating predictive models and using sequence and epigenetic data to distinguish between nucleic acid variations originating from tumors and CHIP, the problem of difficulty in distinguishing between them in existing technologies has been solved, achieving higher detection accuracy and sensitivity, and supporting personalized treatment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUARDANT HEALTH INC
- Filing Date
- 2024-07-26
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies are unable to effectively distinguish between nucleic acid variations originating from tumors and those originating from clonal hematopoiesis of undetermined potential (CHIP), making it difficult to determine the origin of variations observed in leukocytes.
By identifying sequence and epigenetic data associated with multiple genomic regions, a predictive model is generated. This model is then trained using machine learning algorithms to distinguish between nucleic acid variations of tumor and non-tumor origin. Accuracy is improved by combining feature space reduction and filtering criteria.
It enables effective differentiation between nucleic acid variants originating from tumors and those originating from CHIP, improving the accuracy and sensitivity of detection and supporting personalized treatment decisions.
Smart Images

Figure CN121844385A_ABST
Abstract
Description
Cross Reference to Related Applications
[0001] This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 516,207, filed July 28, 2023, which is incorporated by reference herein in its entirety for all purposes. BACKGROUND
[0002] Known liquid biopsy next generation sequencing (NGS) assays observe a confounding genomic signal from nucleic acid variants originating from white blood cells. Stem cells in the bone marrow “white blood cells” divide to produce new blood cells, and with each cell division, there is a probability that a DNA replication error can occur. The high rate of cell division in stem cells allows for the accumulation of mutations, producing daughter blood cells that share these mutations, even if the cells are non-cancerous. The accumulation of mutations in blood cells is known as clonal hematopoiesis of indeterminate potential (CHIP). While it is well understood that the variants observed in specific gene panels provide a large portion of the confounding CHIP signal, it is currently difficult to determine whether the variants observed in these genes are caused by white blood cells or a tumor.
[0003] Accordingly, there is a need for methods to distinguish tumor-derived nucleic acid variants and CHIP-derived nucleic acid variants from one another. SUMMARY
[0004] Described herein is a method comprising: determining sequence data of a plurality of sequence fragments associated with a plurality of genomic regions, wherein the sequence data comprises a plurality of sequence reads, wherein the plurality of sequence reads are sequenced from a plurality of sequence fragments from a plurality of samples, wherein each of the plurality of samples is labeled as tumor-derived or non-tumor-derived; determining epigenetic data associated with the plurality of sequence fragments; determining a plurality of features of a prediction model based on the sequence data and the epigenetic data; generating the prediction model from the plurality of features based on the sequence data and the epigenetic data. In other embodiments, determining the sequence data comprises obtaining a plurality of samples from a plurality of subjects, wherein the plurality of samples comprise a plurality of cell-free nucleic acids. In other embodiments, determining the plurality of features of the prediction model comprises selecting applied features from a candidate feature set. In other embodiments, selecting applied features from the candidate feature set comprises feature space reduction. In various embodiments, selecting applied features for feature space reduction comprises univariate model performance. In various embodiments, this comprises one or more of, for example: a false discovery rate of less than or equal to 50% (FDR <= 50%) or lower and a sensitivity of >= 1% or higher, and / or cancer-specific feature exclusion as a filtering criterion. In various embodiments, the candidate features are derived from one or more of: variant level information such as VAF, fragmentomics, methylation, and / or clinical test sample databases such as aggregated variant summary statistics from clinical patients, clonality, longitudinal variant variability, cancer type variability, etc. and / or public datasets such as COSMIC (Cancer of Somatic Cell Ontology), GnomAD (population level allele frequency), COSMIC mutational signatures (SBS signatures). In various embodiments, selecting applied features comprises excluding candidate features with greater than 50%, 60%, 70%, 80% or more correlation. In other embodiments, the plurality of features comprises at least one of: fragment length, variant VAF, variant CHIP to somatic ratio, APOBEC-related cancer marker, variant measurement variability, variant maximum clonality, age-related marker, variant clonality variance, population allele frequency, ratio of methylated to unmethylated fragments, genomic region associated with cancer type, genomic region associated with methylation status, genomic region associated with hypomethylation, or genomic region associated with therapy response. In other embodiments, the method comprises applying 1-5, 5-10, 10-15, 15-20, 20 or more features. In other embodiments, the applied features comprise a ratio of methylated fragments to unmethylated fragments. In other embodiments, the fragment length is mean length and / or length variance. In other embodiments, the fragment length is single-nucleosome and / or double-nucleosome associated.In other embodiments, the feature is COSMIC signature. In other embodiments, the age-related marker is SBS88. In other embodiments, the APOBEC-related cancer marker is SBS2. In other embodiments, the variant measure is one or more of: variability, variant maximum clonality, and variant clonality variance. In other embodiments, the epigenetic data comprises information about DNA methylation, histone states or modifications, inflammation-mediated cytosine damage products, or protein binding. In other embodiments, the epigenetic data associated with the plurality of sequence fragments comprises determining methylation states of the plurality of sequence fragments. In other embodiments, the methylation states of the plurality of sequence fragments comprise determining at least one of: a methylation state vector or a methylation CpG density. In other embodiments, the method comprises determining the methylation state vector, determining the methylation state vector comprising: aligning the plurality of sequence reads to a reference sequence; determining, based on the alignment, methylation status of one or more CpG sites in the sequence reads of the plurality of sequence reads and positions of the one or more CpG sites; and vectorizing the methylation status of the one or more CpG sites and the positions of the one or more CpG sites to generate a methylation state vector for the sequence reads of the plurality of sequence reads. In other embodiments, the method comprises determining the methylation CpG density, determining the methylation CpG density comprising: aligning the plurality of sequence reads to a reference sequence; determining, based on the alignment, methylation status of one or more CpG sites in the sequence reads of the plurality of sequence reads; determining, based on the methylation status of the one or more CpG sites in the sequence reads, whether the sequence reads are methylated or unmethylated; for the plurality of sequence reads, determining a count of methylated sequence reads and a count of unmethylated sequence reads; and determining, based on the count of methylated sequence reads and the count of unmethylated sequence reads, the methylation CpG density. In other embodiments, training the predictive model comprises applying a machine learning algorithm. In other embodiments, the machine learning method comprises at least one of: discriminant analysis, decision tree, nearest neighbor (NN) algorithm, Bayesian network, clustering algorithm, neural network, support vector machine (SVM), logistic regression algorithm, linear regression algorithm, Markov model, or principal component analysis (PCA). In other embodiments, the method comprises retraining the predictive model.In other embodiments, the method includes, for a subject, determining test sequence data, said test sequence data comprising more than one sequence read sequenced from a sample from the subject; generating test epigenetic data and / or test fragment omics data associated with more than one sequence fragment; providing the subject's test sequence data, test epigenetic data, and test fragment omics data to a predictive model; and, based on the subject's test sequence data, test epigenetic data, and test fragment omics data, determining the origin of at least one sequence fragment in the sequence data. In other embodiments, the method includes determining the origin of at least one sequence fragment in the sequence data. In other embodiments, the origin is one of tumor-derived or non-tumor-derived. In other embodiments, the method includes administering one or more therapies to the subject based on the origin being tumor-derived. In other embodiments, the therapy includes administering chemotherapy, administering radiation therapy, or performing surgery to remove all or part of a tumor.
[0005] This document describes a method comprising: obtaining sequence data, said sequence data including more than one sequence read associated with more than one genomic region, wherein the sequence data is generated from a sample from a subject; determining epigenetic data associated with the more than one sequence read; providing at least a portion of the sequence data and at least a portion of the epigenetic data to a trained predictive model; and determining, based on the predictive model, whether the sample is of tumor origin or non-tumor origin. In other embodiments, the method includes generating a predictive model. In other embodiments, generating a predictive model includes: determining sequence data for more than one sample, wherein each of the more than one sample is labeled as of tumor origin or non-tumor origin; determining epigenetic data associated with more than one sequence read; determining more than one feature of the predictive model based on a portion of the sequence data and a portion of the epigenetic data; training the predictive model according to the more than one feature based on a portion of the sequence data and a portion of the epigenetic data; and predicting the model based on the training output. In other embodiments, the method includes testing based on a portion of the sequence data and a portion of the epigenetic data. In other embodiments, the method includes determining that the sequence data comprises obtaining more than one sample from more than one subject, wherein the more than one sample contains more than one cell-free nucleic acid. In other embodiments, determining more than one feature of the predictive model comprises selecting an applicable feature from a set of candidate features. In other embodiments, selecting an applicable feature from a set of candidate features comprises feature space reduction. In various embodiments, selecting an applicable feature for feature space reduction comprises univariate model performance. In various embodiments, this comprises, for example, one or more of the following: a false discovery rate of less than or equal to 50% (FDR <= 50%) or lower and a sensitivity of >= 1% or higher, and / or exclusion of cancer-specific features as filtering criteria. In various embodiments, candidate features are derived from one or more of the following: variant level information such as VAF, fragmentomics, methylation and / or clinical test sample databases, such as aggregate variant summaries from clinical patients, clonalness, longitudinal variant variability, cancer type variability, etc., and / or public datasets such as COSMIC (variation frequency in cancer tissue), GnomAD (population-level allele frequency), COSMIC mutation imprinting (SBS imprinting). In various implementation schemes, the selection of application features includes excluding candidate features with a correlation greater than 50%, 60%, 70%, 80%, or more.In other embodiments, more than one feature includes at least one of the following: fragment length, variant VAF, variant CHIP to somatic cell ratio, APOBEC-related cancer marker, variant measurement variability, variant maximum clonal size, age-related marker, variant clonal variance, population allele frequency, methylated to unmethylated fragment ratio, genomic region associated with cancer type, genomic region associated with methylation status, genomic region associated with hypomethylation, or genomic region associated with therapy response. In other embodiments, the method includes applying 1-5, 5-10, 10-15, 15-20, 20, or more features. In other embodiments, the applied feature includes the ratio of methylated to unmethylated fragments. In other embodiments, fragment length is mean length and / or length variance. In other embodiments, fragment length is associated with mononucleosomes and / or binucleosomes. In other embodiments, the feature is a COSMIC-related imprint. In other embodiments, the age-related marker is SBS88. In other embodiments, the APOBEC-related cancer marker is SBS2. In other embodiments, the variation measure is one or more of the following: variability, maximum clonalness of the variant, and variance of the clonalness of the variant. In other embodiments, more than one genomic region includes at least one of the following: a genomic region known to be associated with cancer type, a genomic region associated with a known methylation status, a genomic region known to be associated with hypomethylation, or a genomic region known to be associated with therapy response. In other embodiments, the epigenetic data includes at least one of the following: information about DNA methylation, histone status or modification, inflammation-mediated cytosine damage products, or protein binding. In other embodiments, the method includes determining epigenetic data, which includes determining methylation status associated with more than one genomic region. In other embodiments, the method includes determining methylation status, which includes determining at least one of the following: a methylation status vector or a methylation CpG density. In other embodiments, the method includes determining a methylation state vector, which includes: aligning more than one sequence read with a reference sequence; determining, based on the alignment, the methylation status of one or more CpG sites and the location of one or more CpG sites in the sequence reads of the more than one sequence read; and vectorizing the methylation status of one or more CpG sites and the location of one or more CpG sites to generate a methylation state vector for the sequence reads of the more than one sequence read.In other embodiments, the method includes determining the methylated CpG density, which includes: aligning more than one sequence read with a reference sequence; determining the methylation status of one or more CpG sites in the more than one sequence read based on the alignment; determining whether the sequence read is methylated or unmethylated based on the methylation status of one or more CpG sites in the sequence read; for more than one sequence read, determining the count of methylated sequence reads and the count of unmethylated sequence reads; and determining the methylated CpG density based on the count of methylated sequence reads and the count of unmethylated sequence reads. In other embodiments, the method includes training a predictive model based on more than one feature using sequence data and epigenetic data, including training the predictive model using a machine learning algorithm. In other embodiments, the machine learning method includes at least one of the following: discriminant analysis, decision trees, nearest neighbor (NN) algorithms, Bayesian networks, clustering algorithms, neural networks, support vector machines (SVM), logistic regression algorithms, linear regression algorithms, Markov models, or principal component analysis (PCA). In other embodiments, the method includes retraining the predictive model. In other embodiments, one or more therapies are administered to the subject based on whether the sample is tumor-derived. In other embodiments, the therapy includes administering chemotherapy, administering radiation therapy, or performing surgery to remove all or part of the tumor.
[0006] This document describes a method for distinguishing between tumor-derived nucleic acid variants and clonal hematopoietic variants of undetermined potential (CHIP) origin from test samples obtained from test subjects, using at least part of a computer. The method includes: obtaining sequence data comprising more than one sequence read associated with more than one genomic region, wherein the sequence data is generated from samples from a test subject; determining epigenetic data associated with the more than one sequence read; providing at least a portion of the sequence data and at least a portion of the epigenetic data to a trained prediction model; and determining the presence or absence of tumor-derived nucleic acid variants and clonal hematopoietic variants of undetermined potential (CHIP) origin in the test sample based on the prediction model. In other embodiments, the method includes generating a prediction model, wherein generating the prediction model includes: determining sequence data for more than one sample, wherein each of the more than one sample is labeled as tumor-derived or non-tumor-derived; determining epigenetic data associated with the more than one sequence read; determining more than one feature of the prediction model based on a portion of the sequence data and a portion of the epigenetic data; training the prediction model based on the more than one feature based on the portion of the sequence data and a portion of the epigenetic data; and outputting the prediction model based on the training.
[0007] This article describes a method for treating cancer in test subjects, the method comprising: Obtaining sequence data, said sequence data including more than one sequence read associated with more than one genomic region, wherein the sequence data is generated from a sample from a subject; identifying epigenetic data associated with the more than one sequence read; providing at least a portion of the sequence data and at least a portion of the epigenetic data to a trained prediction model; and determining, based on the prediction model, the presence or absence of tumor-derived nucleic acid variants and CHIP-derived nucleic acid variants in the test sample. Treating the test subject's cancer by administering at least one therapy to the test subject based on one or more differentiated tumor-derived nucleic acid variants from a set of differentiated tumor-derived nucleic acid variants and CHIP-derived nucleic acid variants present in the test sample.
[0008] This document describes a method for treating cancer in a test subject, the method comprising administering at least one therapy to the test subject based on one or more distinct sets of tumor-derived nucleic acid variants and CHIP-derived nucleic acid variants present in the test sample, wherein the distinct sets of tumor-derived and CHIP-derived nucleic acid variants are generated by: identifying nucleic acid variants in a set of target genomic regions by a computer based on sequence information obtained from nucleic acids in the test sample obtained from the test subject, to generate an identified set of test nucleic acid variants; identifying at least one epigenetic imprint corresponding to a given test nucleic acid variant in the identified set of test nucleic acid variants based on epigenetic information obtained from the nucleic acids in the test sample, to generate a test nucleic acid variant-epigenetic imprint cluster; and using at least one trained classifier by a computer to distinguish tumor-derived and CHIP-derived nucleic acid variants in the test nucleic acid variant-epigenetic imprint cluster from each other. In other embodiments, the method includes generating a predictive model, wherein generating the predictive model includes: Determine the sequence data for more than one sample, where each of the more than one samples... The data is labeled as tumor-derived or non-tumor-derived; epigenetic data associated with more than one sequence read is identified; more than one feature of a predictive model is identified based on a portion of the sequence data and a portion of the epigenetic data; a predictive model is trained based on a portion of the sequence data and a portion of the epigenetic data, according to more than one feature; and a predictive model is based on the training output.
[0009] In other embodiments, epigenetic data includes distinct epigenetic states or conditions expressed through one or more epigenetic loci in a given target genomic region. In other embodiments, distinct corresponding epigenetic imprints include different cell-free nucleic acid (cfNA) fragment lengths, locations, and / or endpoint density distributions. In other embodiments, a given target genomic region includes two or more nucleic acid variant loci. In other embodiments, the nucleic acids in the sample include cell-free nucleic acid (cfNA) fragments and / or nucleic acid molecules obtained from one or more tissues or cells in the sample. In other embodiments, epigenetic data includes the presence or absence of methylation, hydroxymethylation, acetylation, ubiquitination, phosphorylation, ubiquitin-like methylation, ribosylation, citrullination, and / or histone post-translational modifications or other histone variants in one or more of more than one genomic region. In other embodiments, the method includes ignoring distinguishable CHIP-derived nucleic acid variants from further analysis. In other embodiments, the method includes generating at least one report listing tumor-derived nucleic acid variants and CHIP-derived nucleic acid variants distinguishable from each other in the test sample. In other embodiments, the method includes identifying at least one cancer type associated with the distinguishable tumor-derived nucleic acid variants. In other embodiments, the method includes administering at least one therapy to a test subject to treat the identified cancer type. In other embodiments, the method includes administering at least one therapy to a test subject based on one or more distinguishing tumor-derived nucleic acid variants.
[0010] This document describes one or more non-transitory computer-readable media having processor-executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the method according to any of the foregoing embodiments.
[0011] This document describes a system comprising: a computing device configured to perform the method according to any of the preceding claims; and an output device configured to output a prediction model.
[0012] This document describes an apparatus comprising: one or more processors; and a memory storing processor-executable instructions, which, when executed by one or more processors, cause the apparatus to perform the method according to any of the preceding claims.
[0013] This document describes one or more non-transitory computer-readable media having processor-executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the method according to any of the preceding claims.
[0014] This document describes a system including a computing device configured to perform the method according to any of the preceding claims; and an output device configured to output an indication of whether a sample is of tumor origin or non-tumor origin.
[0015] This document describes an apparatus comprising: one or more processors; and a memory storing processor-executable instructions, which, when executed by one or more processors, cause the apparatus to perform the method according to any of the preceding claims.
[0016] This document describes one or more non-transitory computer-readable media having processor-executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the method according to any of the preceding claims.
[0017] This document describes a system including a computing device configured to perform the method according to any of the preceding claims; and an output device configured to output an indication of whether a sample is of tumor origin or non-tumor origin.
[0018] This document describes an apparatus comprising: one or more processors; and a memory storing processor-executable instructions, which, when executed by one or more processors, cause the apparatus to perform the method according to any of the preceding claims. Brief description of the attached diagram
[0019] Figure 1. Figure 1A Model Design. Features were constructed from Guardant's internal and external public datasets, and multiple models were trained using 10-fold cross-validation. Only results from the logistic regression model are shown. Independent cohorts of paired late-stage plasma and white blood cell (WBC) samples sequenced on the epigenomic panel and healthy donors sequenced on existing genomic panels were used for model validation. Figure 1B The image shows measurements from WBC sequencing. Figure 1C The annotation workflow is illustrated. Here, the presence of CHIP (hereinafter also referred to as "CH") can be determined through a classification process including: determining WBC VAF > plasma VAF as a CH variant; if plasma VAF > WBC VAF*10, as a somatic variant; or if plasma VAF > WBC VAF, excluding the presence of WBC variants due to plasma contamination using a Fisher exact test. In the absence of detected WBC variants, additional filtering criteria can be applied based on coverage to determine the probability of detection in WBCs.
[0020] Figure 2 Model performance. Predictions of tumor and non-tumor status were compared with WBC confirmation in 713 somatic SNVs / insertions / deletions (indels) from 72 paired plasma and WBC epigenomic assay samples, and validation was performed in 243 somatic SNVs / insertions / deletions from 76 paired plasma and healthy donors on the genomic assay. The low confirmation rate of low VAF variants (<0.6%) observed in WBC sequencing may be attributed to the detection limit of WBC variant determination and / or possible non-WBC lineage origins.
[0021] Figure 3. Feature importance and examples. Figure 3A The top 10 features in the validation dataset, ranked by relative importance. Gene names are encapsulated in one-hot encoding. Figure 3B High-ranking construction features include clonalness, defined as VAF / tumor score as measured by methylation or maximum somatic VAF (left panel), VAF variation across time points (middle panel), and mean percentage (right panel); and Figure 3C Consistency of variation incidence across solid tumor cancer types in the Guardant plasma database. Figure 3D Further exemplary features and their applications are shown.
[0022] Figure 4 The number of variants predicted to be of non-tumor or tumor origin within each gene, as confirmed in the late validation cohort by WBC or model prediction. The variant counts for genes with cfDNA variations most frequently confirmed in WBC samples are shown, along with counts for clinically operable genes (BRCA1, BRAF, KRAS, ESR1, ATM, CHEK2). The most commonly confirmed WBC genes are consistent with previous reports, including the high incidence of clonal hematopoiesis in ATM and CHEK2 (*).
[0023] Figure 5 The proportion of variants detected or predicted as non-tumor in WBCs in late-stage validation cohorts categorized by age range. Variants predicted or confirmed as non-tumor are reported to be highly correlated with age.
[0024] Figure 6 The diagram shows the flowchart of the training method for the example.
[0025] Figure 7 This is a diagram illustrating an exemplary process flow using a machine learning-based classifier.
[0026] Figure 8 The instance method is shown.
[0027] Figure 9The instance method is shown.
[0028] Figure 10 The instance method is shown.
[0029] Figure 11 The instance method is shown.
[0030] Figure 12 The instance method is shown.
[0031] Figure 13 Feature selection and model building process. Correlation checks prevent correlations from exceeding a threshold (e.g., 70%). Feature space reduction includes univariate model performance. Here, a filtering criterion of a false detection rate (FDR <= 50%) and sensitivity >= 1% is applied, along with exclusion of cancer-specific features to remove bias from training.
[0032] Figure 14 ESR (erythrocyte sedimentation rate) buffy plasma data: clinically relevant genes from NHC. Instances of ground-truth states determined by matching the erythrocyte sedimentation rate buffy plasma data.
[0033] Figure 15. Figure 15A CHIP Classifier Features: The exemplary model uses 14 features and demonstrates best performance on both validation and test data. The exemplary sample-level features are depicted. Figure 15B Performance Receiver Operating Characteristics (ROC) include multi-feature models, exemplary conventional models without multi-feature generation using rule tables for inclusion / exclusion, and simplified feature models obtained by CH variant markers based on the presence of the same SNV / insertion / deletion in paired cfDNA and erythrocyte sedimentation rate (ESR) brown layer.
[0034] Figure 16 CHIP classifier features: All features capture the differences between CHIP variants and somatic variants.
[0035] Figure 17 Regarding the performance of clinically relevant genes: ESR brown-yellow layer-plasma matched samples, performance of multi-feature models, use of rule tables for inclusion / exclusion of exemplary conventional models and simplified feature models without multi-feature generation.
[0036] Figure 18 Regarding the performance of clinically relevant genes: ESR-plasma validation data depicting gene-level performance. High FP genes were not detected in the current classifier.
[0037] Figure 19Regarding the performance of clinically relevant genes: ESR-matched plasma samples were plotted to validate VAF bins. No high FP bins were detected in the current classifier.
[0038] Figure 20 Regarding the performance of clinically relevant genes: ESR-plasma-matched sample validation data. All cancer types were recorded. No high-FP cancer types were detected in the current classifier.
[0039] Figure 21 Regarding the performance of clinically relevant genes: tissue-plasma matched samples.
[0040] Figure 22 Performance of clinically relevant genes: tissue-plasma matched samples. Specificity assessment in low VAF bins. No high FP bins below 0.5% were detected in the current classifier.
[0041] Figure 23 ESR (erythrocyte sedimentation rate) brown-yellow layer - plasma data: clinically relevant genes from an exemplary group of over 700 genes. Instances of basic factual states determined by matching the ESR brown-yellow layer.
[0042] Figure 24 Regarding the performance of panel-wide variability: ESR brown-yellow layer-plasma matched samples, performance of multi-feature models, exemplary regular models and simplified feature models using rule tables for inclusion / exclusion without multi-feature generation.
[0043] Figure 25 Regarding the performance of the whole group of variant tissue-plasma matched samples.
[0044] Figure 26 Operable variation performance: ESR (erythrocyte sedimentation rate) brown layer - plasma-matched samples. The current classifier achieves 100% accuracy and 0% FRP (free rate of response).
[0045] Figure 27 Operable variation performance: tissue-plasma matched samples. The current classifier achieves 0% FRP.
[0046] Figure 28 The operable variant classification indicates a very low FPR; an exemplary group includes approximately 80 genes.
[0047] Figure 29 Operable variant performance, plasma ctDNA-. Most CHIP determinations are in the known KRAS CHIP variant K117N.
[0048] Figure 30 Current model performance summary.
[0049] Figure 31Overview of the CHIP classifier algorithm. Detailed description
[0050] While various embodiments of this disclosure have been shown and described herein, those skilled in the art will understand that such embodiments are provided by way of example only. Many variations, modifications, and substitutions will occur to those skilled in the art without departing from this disclosure. It should be understood that various alternatives may be adopted to the embodiments of this disclosure described herein.
[0051] The term “about” and its grammatical equivalents associated with a reference value can include a range of values that are plus or minus 10% of that value. For example, the quantity “about 10” can include quantities from 9 to 11. The term “about” associated with a reference value can include a range of values that are plus or minus 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%.
[0052] The term “at least” and its grammatical equivalents associated with a reference value can include the reference value and values greater than that value. For example, a quantity “at least 10” can include the value 10 and any value higher than 10, such as 11, 100, and 1,000.
[0053] The term “at most” and its syntactic equivalents associated with a reference value can include the reference value and be less than that value. For example, a quantity “at most 10” can include any value of 10 and below 10, such as 9, 8, 5, 1, 0.5, and 0.1.
[0054] Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” as used herein can include plural referents. Thus, for example, reference to “cell” can include more than one such cell, and reference to “culture” can include reference to one or more cultures and their equivalents known to those skilled in the art, and so on. Unless otherwise clearly indicated, all technical and scientific terms used herein may have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0055] Cancer can be indicated by epigenetic variations such as methylation. Examples of methylation changes in cancer include localized increases in DNA methylation at CpG islands at transcription start sites (TSS) of genes involved in normal growth control, DNA repair, cell cycle regulation, and / or cell differentiation. This hypermethylation may be associated with an abnormal loss of transcriptional capacity of the genes involved and occurs at least as frequently as point mutations and deletions that cause altered gene expression. DNA methylation profiling can be used to detect regions in the genome with different levels of methylation (“differentially methylated regions” or “DMRs”) that have changed during development or been perturbed by disease (e.g., cancer or any cancer-related disease). The genome of cancer cells has an imbalance in the aforementioned DNA methylation patterns, and therefore an imbalance in the functional packaging of DNA. Thus, the combination of chromatin organization abnormalities with methylation changes, when analyzed together, may help enhance cancer profiling. Combining MBD partitioning with fragmentomics data (such as the start and end positions of fragment mappings (related to nucleosome position), fragment length, and associated nucleosome occupancy) can be used for chromatin structure analysis in hypermethylation studies to improve biomarker detection rates.
[0056] Methylation profiling can include identifying methylation patterns across different regions of the genome. For example, after partitioning and sequencing molecules based on their methylation levels (e.g., the relative number of methylation sites per molecule), the sequences of molecules in different partitions can be mapped to a reference genome. This can reveal regions of the genome that are more or less methylated compared to other regions. In this way, genomic regions can differ in their methylation levels in contrast to individual molecules.
[0057] The properties of nucleic acid molecules can be modified, which can include various chemical or protein modifications (i.e., epigenetic modifications). Non-limiting examples of chemical modifications may include, but are not limited to, covalent DNA modifications, including DNA methylation. In some embodiments, DNA methylation includes adding a methyl group to cytosine (the cytosine followed by guanine in a nucleic acid sequence) at a CpG site. In some embodiments, DNA methylation includes adding a methyl group to adenine, such as N6-methyladenine. In some embodiments, DNA methylation is 5-methylation (modification of the fifth carbon of the 6-carbon ring of cytosine). In some embodiments, 5-methylation includes adding a methyl group to the 5C position of cytosine to produce 5-methylcytosine (m5c). In some embodiments, methylation includes derivatives of m5c. Derivatives of m5c include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-carboxycytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the third carbon of the 6-carbon ring of cytosine). In some embodiments, 3C methylation involves adding a methyl group to the 3C position of cytosine to generate 3-methylcytosine (3mC). Other examples include N6-methyladenine or glycosylation. DNA methylation involves adding a methyl group to DNA (e.g., CpG) and can alter the expression of methylated DNA regions. Methylation can also occur at non-CpG sites; for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, when DNA in a promoter region is methylated, gene transcription can be repressed. DNA methylation is essential for normal development, and abnormalities in methylation can disrupt epigenetic regulation. Disruptions in epigenetic regulation, such as repression, can lead to diseases such as cancer. DNA methylation of the promoter may indicate cancer.
[0058] CpG dinucleotides are dinucleotides CpG (cytosine-phosphate-guanine) on the sense strand of a double-stranded DNA molecule, i.e., at the 5' end of the nucleic acid sequence. In the 3' direction, cytosine is followed by guanine and its complementary CpG on the antisense chain. CpG dyads can be fully methylated or hemimethylated (methylated on only one chain).
[0059] CpG dinucleotides are underrepresented in the normal human genome, where most CpG dinucleotide sequences are transcriptionally inert (e.g., near-centromere regions of chromosomes and heterochromatin regions of DNA in repetitive elements) and are methylated. However, many CpG islands are protected from such methylation, especially around transcription start sites (TSS).
[0060] Specifically, based on the methods and techniques described herein, epigenomic measurements of tumor scores can improve negative predictions. Overestimation of tumor score predictions (for CNV, CHIP, etc., regarding the epi-MAF gene) based on data assumptions and the probability of the absence of variants with >30% clonalness can lead to overconfidence in negative predictions, while underestimation of tumor scores (TND, etc.) results in an inability to make a confident negative determination. Static parameters include the subclonal variant purity boundary (30%), the probability of previous variants, and mutual exclusivity or co-occurrence with other variants. Key derived parameters include tumor score and cancer tissue origin. Taking these parameters into account, the probability distribution of tumor scores based on methylation data can be measured.
[0061] Protein modifications include components that bind chromatin, particularly histones (including their modified forms), as well as components that bind other proteins, such as those involved in replication or transcription. This disclosure provides methods for processing and analyzing nucleic acids with varying degrees of modification, such that the nature of their original modifications is correlated with nucleic acid tags, which can be decoded during nucleic acid analysis by sequencing. Genetic variation in the nucleic acid modifications of a sample can then be correlated with the degree of modification (epigenetic variation) of that nucleic acid in the original sample, said nucleic acids including single-stranded (e.g., ssDNA or RNA) or double-stranded molecules (e.g., dsDNA).
[0062] DNA loss can reduce the presence of one or more types of DNA, making them difficult to detect (e.g., cfDNA). In one or more other cases, existing methods for measuring DNA methylation, such as enrichment or depletion methods, can have relatively high levels of resolution, such as from about 100 base pairs (bp) to about 200 bp, which can make it difficult to accurately determine the amount of DNA methylation. The accuracy of DNA methylation determination can affect the accuracy of tumor score estimation in a sample. Since tumor scores are used to determine whether a sample is derived from a subject with or without a tumor, the accuracy of tumor score estimation can influence individual diagnostic and / or treatment decisions.
[0063] Sample The sample can be any biological sample isolated from the subject. The sample can be a bodily sample. Samples can include body tissues such as known or suspected solid tumors, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells or leukocytes, endothelial cells, tissue biopsies, cerebrospinal fluid, synovial fluid, lymph, ascites, interstitial fluid or extracellular fluid, fluids in the intercellular spaces (including gingival crevicular fluid), bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. Samples are preferably bodily fluids, particularly blood and its fractions, and urine. Samples can be in the form initially isolated from the subject, or can be further processed to remove or add components, such as cells, or to enrich one component relative to other components. Therefore, preferred bodily fluids for analysis are plasma or serum containing cell-free nucleic acids. Samples can be isolated or obtained from the subject and transported to the sample analysis site. Samples can be stored and transported at desirable temperatures, such as room temperature, 4°C, -20°C, and / or -80°C. Samples may be isolated from or obtained from the subject at the sample analysis site. Subjects may be humans, mammals, animals, companion animals, service animals, or pets. Subjects may have cancer. Subjects may not have cancer or may have detectable symptoms of cancer. Subjects may have been treated with one or more cancer therapies, such as chemotherapy, antibodies, vaccines, or biologics. Subjects may be in remission. Subjects may be diagnosed or may not be diagnosed with a susceptibility to cancer or any cancer-related genetic mutations / disorders.
[0064] The volume of plasma can depend on the desired read depth for the sequenced region. Exemplary volumes are 0.4–40 ml, 5–20 ml, and 10–20 ml. For example, the volume can be 0.5 mL, 1 mL, 5 mL, 10 mL, 20 mL, 30 mL, or 40 mL. The volume of plasma sampled can be from 5 ml to 20 ml.
[0065] Samples can contain varying amounts of nucleic acids containing genomic equivalents. For example, a sample of approximately 30 ng DNA can contain approximately 10,000 (10^3) nucleotides. 4 The genome equivalent of 200 billion haploid human genomes, and in the case of cfDNA, it can contain approximately 200 billion (2 × 10⁻⁶) haploid human genomes. 11 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng DNA can contain about 30,000 haploid human genome equivalents, and in the case of cfDNA, about 600 billion individual molecules.
[0066] The sample may contain nucleic acids from different sources, such as nucleic acids and cell-free nucleic acids from cells of the same subject, or nucleic acids and cell-free nucleic acids from cells of different subjects. The sample may contain nucleic acids carrying mutations. For example, the sample may contain DNA carrying germline mutations and / or somatic mutations. Germline mutations refer to mutations present in the germline DNA of the subject. Somatic mutations refer to mutations originating from the subject's somatic cells, such as cancer cells. The sample may contain DNA carrying cancer-related mutations (e.g., cancer-related somatic mutations). The sample may contain epigenetic variations (i.e., chemical or protein modifications) that are associated with the presence of genetic variations (such as cancer-related mutations). In some embodiments, the sample contains epigenetic variations associated with the presence of genetic variations, wherein the sample does not contain said genetic variations.
[0067] Exemplary amounts of cell-free nucleic acid in the pre-amplification sample range from about 1 fg to about 1 µg, such as 1 pg to 200 ng, 1 ng to 100 ng, 10 ng to 1000 ng. For example, amounts can be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. Amounts can be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The quantity can be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acid molecules. The method may include obtaining 1 femtogram (fg) to 200 ng.
[0068] Cell-free nucleic acids are nucleic acids that are not contained within cells or otherwise bound to cells, or in other words, nucleic acids that remain in a sample after the removal of intact cells. Cell-free nucleic acids include DNA, RNA, and their hybrids, including genomic DNA, mitochondrial DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids can be released into body fluids through secretion or cell death procedures such as cell necrosis and apoptosis. Some cell-free nucleic acids are released into body fluids from cancer cells, such as circulating tumor DNA (ctDNA). Others are released from healthy cells. In some embodiments, cfDNA is cell-free fetal DNA (cffDNA). In some embodiments, cell-free nucleic acids are produced by tumor cells. In some embodiments, cell-free nucleic acids are produced by a mixture of tumor cells and non-tumor cells.
[0069] Cell-free nucleic acids have an example size distribution of approximately 100-500 nucleotides, with molecules of 110 to 230 nucleotides representing approximately 90% of the molecules, a mode of approximately 168 nucleotides, and a second small peak in the range of 240 to 440 nucleotides. Cell-free nucleic acids can be separated from body fluids via a fractionation or partitioning step, in which the cell-free nucleic acids present in solution are separated from intact cells and other insoluble components in the body fluid. Partitioning can include techniques such as centrifugation or filtration. Optionally, cells in the body fluid can be lysed, and cell-free and cellular nucleic acids are processed together. Typically, nucleic acids can be precipitated with alcohol after adding buffer and washing steps. Further cleaning steps, such as silica-based columns, can be used to remove contaminants or salts. Nonspecific bulk carrier nucleic acids (such as Cot-1 DNA) or DNA or proteins for bisulfite sequencing, hybridization, and / or ligation can be added throughout the reaction to optimize certain aspects of the procedure, such as yield.
[0070] Following such processing, the sample can include various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and single-stranded RNA. In some embodiments, single-stranded DNA and RNA can be converted into double-stranded forms, and therefore included in subsequent processing and analytical steps.
[0071] Analyte Analytes may include nucleic acid analytes and non-nucleic acid analytes. This disclosure provides a method for detecting genetic variations in biological samples from a subject. Biological samples may include polynucleotides from cancer cells. Polynucleotides may be DNA (e.g., genomic DNA, cDNA), RNA (e.g., mRNA, small RNA), or any combination thereof. Biological samples may include, for example, tumor tissue from a biopsy. In some cases, biological samples may include blood or saliva. In specific cases, biological samples may contain cell-free DNA (“cfDNA”) or circulating tumor DNA (“ctDNA”). Cell-free DNA may be present, for example, in blood.
[0072] Examples of non-nucleic acid analytes include, but are not limited to, lipids, carbohydrates, peptides, proteins, glycoproteins (N-linked or O-linked), lipoproteins, phosphoproteins, specific phosphorylation or acetylation variations of proteins, amidation variations of proteins, hydroxylation variations of proteins, methylation variations of proteins, ubiquitination variations of proteins, sulfation variations of proteins, viral proteins (e.g., viral capsid, viral envelope, viral outer shell, viral appendages, viral glycoproteins, viral spikes, etc.), extracellular and intracellular proteins, antibodies, and antigen-binding fragments. This also includes receptors, antigens, surface proteins, transmembrane proteins, differentiation protein clusters, protein channels, protein pumps, carrier proteins, phospholipids, glycoproteins, glycolipids, cell-cell interaction protein complexes, antigen-presenting complexes, major histocompatibility complexes, engineered T-cell receptors, T-cell receptors, B-cell receptors, chimeric antigen receptors, extracellular matrix proteins, post-translational modifications (e.g., phosphorylation, glycosylation, ubiquitination, nitrosation, methylation, acetylation, or lipidation) of cell surface proteins, gap junctions, and adhesion junctions.
[0073] Typically, systems, apparatus, methods, and compositions can be used to analyze any number of analytes, further including both nucleic acid analytes and non-nucleic acid analytes. For example, the number of analytes being analyzed can be at least about 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 100, at least 1,000, at least 10,000, at least 100,000, or more different analytes present in a region of the sample or a single feature of the substrate.
[0074] One or more nucleic acid analytes and / or non-nucleic acid analytes constitute a set of molecular interactions in the biological system under study (e.g., a cell), which can be considered as an “interaction set”—molecular interactions occurring between molecules belonging to different biochemical families (proteins, nucleic acids, lipids, carbohydrates, etc.) and also within a given family. In various embodiments, the interaction set is a protein-DNA interaction set (a network formed by transcription factors (and DNA or chromatin regulatory proteins) and their target genes). In other embodiments, the interaction set refers to a protein-protein interaction network (PPI) or a protein-protein interaction network (PIN). The methods described herein allow for the study and analysis of interaction sets. Techniques such as proteogenomics (whole genome sequencing, whole exome sequencing, and RNA-seq, as well as mass spectrometry, as examples) can support the study of interaction sets.
[0075] Analysis The methods of this invention can be used to diagnose the presence of a condition, particularly cancer, in a subject, to characterize the condition (e.g., to stage the cancer or determine its heterogeneity), monitor the condition's response to treatment, and achieve prognostic assessment of the risk of condition progression or subsequent disease development. This disclosure can also be used to determine the efficacy of a particular treatment option. If treatment is successful, a successful treatment option may increase the amount of copy number variations or rare mutations detected in the subject's blood as more cancer cells may die and shed DNA. In other instances, this may not occur. In yet another instance, some treatment options may be correlated with the genetic profile of the cancer over time. This correlation can be used to select a therapy. Additionally, if cancer is observed to be in remission after treatment, the methods of this invention can be used to monitor residual disease or disease recurrence.
[0076] The types and numbers of cancers that can be detected include leukemia, brain cancer, lung cancer, skin cancer, nasal cancer, laryngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, skin cancer, colorectal cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, solid tumors, heterogeneous tumors, and homogeneous tumors. Cancer type and / or stage can be detected based on genetic variations including: mutations, rare mutations, insertions / deletions, copy number variations, transversions, translocations, inversions, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structural alterations, gene fusions, gene truncation, gene amplification, gene duplication, chromosomal damage, DNA damage, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, and abnormal changes in 5-methylcytosine.
[0077] Genetic and other analyte data can also be used to characterize specific forms of cancer. Cancers are often heterogeneous in both composition and stage. Genetic profiling data can allow for the characterization of specific subtypes of cancer, which may be important in the diagnosis or treatment of that specific subtype. This information can also provide subjects or practitioners with clues about the prognosis of a specific type of cancer and allow them to adjust treatment options based on disease progression. Some cancers can progress and become more aggressive and genetically unstable. Other cancers may remain benign, inactive, or dormant. The systems and methods disclosed herein can be used to determine disease progression.
[0078] The analyses of this invention can also be used to determine the efficacy of a particular treatment option. If the treatment is successful, a successful treatment option may increase the amount of copy number variations or rare mutations detected in the subject's blood as more cancer cells may die and shed DNA. In other instances, this may not occur. In yet another instance, some treatment options may be correlated with the genetic profile of the cancer over time. This correlation can be used to select a therapy. Additionally, if cancer is observed to be in remission after treatment, the method of this invention can be used to monitor residual disease or recurrence of disease.
[0079] The methods of this invention can also be used to detect genetic variations in conditions other than cancer. Following the onset of certain diseases, immune cells, such as B cells, can undergo rapid clonal expansion. Copy number variation detection can be used to monitor clonal expansion and certain immune states. In this instance, copy number variation analysis can be performed over time to generate a spectrum of how a particular disease may progress. Copy number variation or even rare mutation detection can be used to determine how a pathogen population changes during the course of infection. This can be particularly important during chronic infections (such as HIV / AID or hepatitis infections), where the virus can alter its life cycle state and / or mutate into a more virulent form during the course of infection. When immune cells attempt to destroy transplanted tissue, the methods of this invention can be used to determine or analyze the host body's rejection activity to monitor the state of the transplanted tissue and to modify the process of rejection treatment or prevention.
[0080] Furthermore, the methods of this disclosure can be used to characterize the heterogeneity of anomalies in a subject. Such methods may include, for example, generating a genetic profile of extracellular polynucleotides derived from the subject, wherein the genetic profile includes more than one set of data obtained from copy number variation and rare mutation analysis. In some embodiments, the anomaly is cancer. In some embodiments, the anomaly may be a condition leading to a heterogeneous genomic population. In the example of cancer, some tumors are known to contain tumor cells at different stages of cancer. In other examples, heterogeneity may include multiple lesions of the disease. Again, in the example of cancer, multiple tumor lesions may be present, perhaps one or more of which are the result of metastases that have spread from the primary site.
[0081] The method of this invention can be used to generate or analyze a fingerprint or dataset that represents the sum of genetic information derived from different cells in heterogeneous diseases. This dataset can include individual or combined copy number variation and mutation analyses.
[0082] The methods of this invention can be used for the diagnosis, prognosis, monitoring, or observation of cancer or other diseases. In some embodiments, the methods herein do not involve the diagnosis, prognosis, or monitoring of the fetus, and therefore do not involve noninvasive prenatal testing. In other embodiments, these methods can be used in pregnant subjects to diagnose, prognose, monitor, or observe cancer or other diseases in unborn subjects whose DNA and other polynucleotides may co-circulate with maternal molecules.
[0083] Determination of 5-methylcytosine patterns of nucleic acids Bisulfite-based sequencing and its variations provide a means of determining the methylation patterns of nucleic acids. In some embodiments, determining the methylation pattern includes distinguishing between 5-methylcytosine (5mC) and unmethylated cytosine. In some embodiments, determining the methylation pattern includes distinguishing between N6-methyladenine and unmethylated adenine. In some embodiments, determining the methylation pattern includes distinguishing between 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxycytosine (5caC) and unmethylated cytosine. Examples of bisulfite sequencing include, but are not limited to, oxidized bisulfite sequencing (OX-BS-seq), Tet-assisted bisulfite sequencing (TAB-seq), and reduced bisulfite sequencing (redBS-seq).
[0084] Oxidized bisulfite sequencing (OX-BS-seq) is used to distinguish between 5mC and 5hmC by first converting 5hmC to 5fC, followed by bisulfite sequencing as described above. Tet-assisted bisulfite sequencing (TAB-seq) can also be used to distinguish between 5mC and 5hmC. In TAB-seq, 5hmC is protected by glycosylation. As mentioned earlier, 5mC is converted to 5caC using the Tet enzyme before bisulfite sequencing. Reduced bisulfite sequencing is used to distinguish 5fC from modified cytosine.
[0085] Typically, in bisulfite sequencing, nucleic acid samples are split into two aliquots, and one aliquot is treated with bisulfite. Bisulfite converts native cytosine and certain modified cytosine nucleotides (e.g., 5-formylcytosine or 5-carboxycytosine) to uracil, while other modified cytosines (e.g., 5-methylcytosine, 5-hydroxymethylcytosine) are not converted. Comparison of the nucleic acid sequences of molecules from the two aliquots indicates which cytosines were converted to uracil and which were not. Therefore, modified and unmodified cytosines can be identified. Initially splitting the sample into two aliquots is disadvantageous for samples containing only small amounts of nucleic acids and / or including heterogeneous cellular / tissue-derived samples such as body fluids containing cell-free DNA.
[0086] This disclosure provides methods for enabling bisulfite sequencing and its variations. These methods function by linking nucleic acids in a population to a capture motif (i.e., a tag that can be captured or immobilized). Capture motifs include, but are not limited to, biotin, avidin, streptavidin, nucleic acids containing a specific nucleotide sequence, haptens recognized by antibodies, and magnetically attractive particles. Extraction motifs can be members of binding pairs such as biotin / streptavidin or haptens / antibodies. In some embodiments, the capture motif attached to the analyte is captured by its binding pair, which is attached to a separable motif, such as magnetically attractive particles or large particles that can be precipitated by centrifugation. The capture motif can be any type of molecule that allows affinity separation of nucleic acids with the capture motif from nucleic acids lacking the capture motif. Example capture motifs are biotin or oligonucleotides, where biotin allows affinity separation by binding to streptavidin linked to or capable of being linked to a solid phase, and oligonucleotides allow affinity separation by binding to complementary oligonucleotides linked to or capable of being linked to a solid phase. After the capture motif is linked to the sample nucleic acid, the sample nucleic acid is used as an amplification template. After amplification, the original template remains connected to the captured portion, but the amplicon does not connect to the captured portion.
[0087] The capture portion can be attached to the sample nucleic acid as a component of an adaptor, which can also provide binding sites for amplification and / or sequencing primers. In some methods, the sample nucleic acid is attached to an adaptor at both ends, with both adaptors containing the capture portion. Preferably, any cytosine residues in the adaptor are modified, such as by 5-methylcytosine, to protect against bisulfite. In some cases, the capture portion is attached via a cleavable adapter (e.g., photocleavable dethiobiotin-TEG or uracil residues cleavable by the USER™ enzyme, Chem. Commun. (Camb. 2015 Feb21; 51(15): 3266-3269), in which case the capture portion can be removed if desired.
[0088] The amplicon is denatured and then contacted with an affinity reagent used to capture the tag. The original template binds to the affinity reagent, while the amplified nucleic acid molecules do not. Therefore, the original template can be separated from the amplified nucleic acid molecules.
[0089] After isolation or partitioning, the corresponding populations of nucleic acids (i.e., the original template and amplification products) can be subjected to bisulfite treatment, with the original template population receiving bisulfite treatment while the amplification products do not. Optionally, the amplification products can undergo bisulfite treatment while the original template population does not. After such treatment, the corresponding populations can be amplified (in the case of the original template population, this converts uracil to thymine). The populations can also undergo biotinylated probe hybridization for enrichment. The corresponding populations are then analyzed and sequences are compared to determine which cytosines are 5-methylated (or 5-hydroxymethylated) in the original sample. Detection of T nucleotides (corresponding to unmethylated cytosine converted to uracil) in the template population and C nucleotides at corresponding positions in the amplification population indicates unmodified C. The presence of C at corresponding positions in the original template and amplification population indicates the presence of modified C in the original sample.
[0090] In some embodiments, the method uses sequential DNA-seq and bisulfite-seq (BIS-seq) NGS library preparation with molecularly tagged DNA libraries. The process involves tagging of an adaptor (e.g., biotin), DNA-seq amplification of the entire library, parental molecule recovery (e.g., streptavidin bead pull-down), bisulfite conversion, and BIS-seq. In some embodiments, the method identifies 5-methylcytosine at single-base resolution through preparative sequential NGS amplification of parental library molecules with and without bisulfite treatment. This can be achieved by modifying the 5-methylated NGS adaptor used in BIS-seq (oriented adaptor; Y-shaped / forked, replaced with 5-methylcytosine) with a marker (e.g., biotin) on one of the two adaptor strands. Sample DNA molecules are linked adaptors and are amplified (e.g., by PCR). Since only parental molecules will have tagged adaptor ends, they can be selectively recovered from their amplified progeny using tag-specific capture methods (e.g., streptavidin magnetic beads). Because the parental molecules retain the 5-methylation marker, bisulfite conversion on the captured library will produce a 5-methylation state at single-base resolution during BIS-seq, thus preserving molecular information in the corresponding DNA-seq. In some embodiments, the bisulfite-treated library can be combined with the untreated library prior to enrichment / NGS by adding a sample-tagged DNA sequence in a standard multiplex NGS workflow. As with the BIS-seq workflow, bioinformatics analysis can be performed for genome alignment and 5-methylation base recognition. In summary, this method provides the ability to selectively recover parental, ligated molecules carrying the 5-methylcytosine marker after library amplification, allowing for parallel processing of bisulfite-converted DNA. This overcomes the detrimental nature of bisulfite treatment to the quality / sensitivity of DNA-seq information extracted from the workflow. With this method, the recovered ligated, parental DNA molecules (via tagged adaptors) allow for the amplification of a complete DNA library and the parallel application of treatments that induce epigenetic DNA modifications. This disclosure discusses the use of BIS-seq methods to identify 5-methylated cytosine (5-methylcytosine), but this should not be limiting. Variations of BIS-seq have been developed to identify hydroxymethylated cytosine (5hmC; OX-BS-seq, TAB-seq), formylcytosine (5fC; redBS-seq), and carboxycytosine. These methods can be implemented using sequential / parallel library preparation as described herein.
[0091] Alternative methods of analyzing modified nucleic acids This disclosure provides alternative methods for analyzing modified nucleic acids (e.g., methylated, histone-linked, and other modifications discussed above). In some such methods, a population of nucleic acids with varying degrees of modification (e.g., each nucleic acid molecule has 0, 1, 2, 3, 4, 5, or more methyl groups) is contacted with an adaptor, and the population is then stratified according to the degree of modification. The adaptor is attached to one or both ends of the nucleic acid molecules in the population. Preferably, the adaptor contains a sufficient number of different tags such that the number of tag combinations results in a high probability, e.g., 95%, 99%, or 99.9%, that two nucleic acids with the same start and end points receive different tag combinations. After attaching the adaptor, the nucleic acid is amplified from a primer that binds to a primer binding site within the adaptor. Adaptors with the same or different tags may contain the same or different primer binding sites, but preferably the adaptor contains the same primer binding site. After amplification, the nucleic acid is contacted with an agent that preferably binds the modified nucleic acid (such as those previously described). The nucleic acid is divided into at least two partitions, the difference between the at least two partitions being the degree to which the modified nucleic acid binds to the agent. For example, if the reagent has an affinity for the modified nucleic acid, the overrepresented modified nucleic acid (compared to the median representation in the population) preferentially binds to the reagent, while the underrepresented modified nucleic acid does not bind to the reagent or is more easily eluted from it. After separation, the different partitions can then undergo additional processing steps, which typically include parallel but separate additional amplification and sequence analysis. The sequence data from the different partitions can then be compared.
[0092] Nucleic acid molecules can be coupled to Y-adaptors containing primer binding sites and tags. The molecule is then amplified. The amplified molecule is then partitioned by contacting an antibody that preferentially binds to 5-methylcytosine to produce two partitions. One partition contains the unmethylated original molecule and the amplified copy that has lost methylation. The other partition contains the original DNA molecule with methylation. The two partitions are then processed and sequenced separately, and the methylated partition is further amplified. The sequence data of the two partitions can then be compared. In this example, the tag is not used to distinguish between methylated and unmethylated DNA, but rather to distinguish the different molecules within these partitions, allowing one to determine whether reads with the same start and end points are based on the same or different molecules.
[0093] This disclosure also provides methods for analyzing nucleic acid populations, wherein at least some nucleic acids contain one or more modified cytosine residues, such as 5-methylcytosine and any other modifications previously described. In these methods, the nucleic acid population is contacted with an adaptor comprising one or more cytosine residues, such as 5-methylcytosine, modified at the 5C position. Preferably, all cytosine residues in such an adaptor are also modified, or all such cytosine residues in the primer-binding region of the adaptor are modified. The adaptor is attached to both ends of the nucleic acid molecules in the population. Preferably, the adaptor contains a sufficient number of different tags such that the number of tag combinations results in a high probability, for example, 95%, 99%, or 99.9%, that two nucleic acids with the same start and end points receive different tag combinations. The primer-binding sites in such an adaptor may be the same or different, but are preferably the same. After the adaptor is attached, the nucleic acid is amplified by primers that bind to the primer-binding sites of the adaptor. The amplified nucleic acid is divided into a first aliquot and a second aliquot. The first aliquot is sequenced, with or without further processing. This determines the sequence data of molecules in the first aliquot regardless of the initial methylation state of the nucleic acid molecules. Nucleic acid molecules in the second aliquot are treated with bisulfite. This treatment converts unmodified cytosine to uracil. The bisulfite-treated nucleic acids then undergo amplification, initiated by primers targeting the original primer binding sites of the adaptor attached to the nucleic acid. Now only the nucleic acid molecules initially attached to the adaptor (unlike their amplification products) are amplified because these nucleic acids retain cytosine at the primer binding sites of the adaptor, while the amplification products have lost the methylation of these cytosine residues, which were converted to uracil during the bisulfite treatment. Therefore, only the original molecules in the population (at least some of which are methylated) undergo amplification. After amplification, these nucleic acids are sequenced. Comparison of the sequences determined from the first and second aliquots can particularly indicate which cytosine residues in the nucleic acid population have undergone methylation.
[0094] Partitioning a sample into more than one subsample; aspects of the sample; analysis of epigenetic characteristics Subjecting a first subsample to a procedure that differentially affects a first nucleobase in DNA and a second nucleobase in DNA of the first subsample In some embodiments described herein, different forms of nucleic acid populations (e.g., hypermethylated and hypomethylated DNA in a sample, such as a capture set of cfDNA as described herein) can be physically partitioned based on one or more properties of the nucleic acids, followed by further analysis, such as differential modification or isolation of nucleotides, tagging, and / or sequencing. This approach can be used to determine, for example, whether certain sequences are hypermethylated or hypomethylated. In some embodiments, hypermethylated variable epigenetic target regions are analyzed to determine whether they exhibit the hypermethylation properties of tumor cells, and / or hypomethylated variable epigenetic target regions are analyzed to determine whether they exhibit the hypomethylation properties of tumor cells. Additionally, by partitioning heterogeneous nucleic acid populations, one can increase rare signals, for example, by enriching rare nucleic acid molecules that are more prevalent in one fraction (or partition) of the population. For example, by partitioning a sample into hypermethylated and hypomethylated nucleic acid molecules, it is easier to detect genetic variations that are present in hypermethylated DNA but less so (or absent) in hypomethylated DNA. By analyzing more than one fraction of a sample, multidimensional analysis of individual loci or nucleic acid species in the genome can be performed, thus enabling greater sensitivity.
[0095] In some cases, heterogeneous nucleic acid samples are partitioned into two or more partitions (e.g., at least three, four, five, six, or seven partitions). In some implementations, each partition is differentially tagged. The tagged partitions can then be pooled together for collective sample preparation and / or sequencing. The partitioning-tag-pooling step can occur more than once, with each round of partitioning occurring based on a different attribute (as exemplified in this paper) and tagged using differential tags that distinguish them from other partitioning and partitioning methods.
[0096] Examples of attributes that can be used for partitioning include sequence length, methylation level, nucleosome binding, sequence mismatch, immunoprecipitation, and / or proteins that bind to DNA. The resulting partitions can include one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), shorter DNA fragments, and longer DNA fragments. In some embodiments, partitioning is typically performed based on cytosine modification (e.g., cytosine methylation) or methylation, and optionally combined with at least one additional partitioning step, which can be based on any of the aforementioned attributes or forms of DNA. In some embodiments, a heterogeneous population of nucleic acids is partitioned into nucleic acids having one or more epigenetic modifications and nucleic acids not having said one or more epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation, methylation level, methylation type (e.g., 5-methylcytosine with other types of methylation, such as adenine methylation and / or cytosine hydroxymethylation), and association with one or more proteins (such as histones) and the level of association. Optionally or additionally, the heterogeneous nucleic acid population can be partitioned into nucleosome-associated nucleic acid molecules and nucleosome-free nucleic acid molecules. Optionally or additionally, the heterogeneous nucleic acid population can be partitioned into single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Optionally or additionally, the heterogeneous nucleic acid population can be partitioned based on nucleic acid length (e.g., molecules with a maximum length of 160 bp and molecules with a length greater than 160 bp).
[0097] In some cases, each partition (representing a different nucleic acid form) is differentially labeled, and the partitions are pooled together and then sequenced. In other cases, the different forms are sequenced separately. In some implementations, different nucleic acid populations are partitioned into two or more distinct partitions. Each partition represents a different nucleic acid form, and the first partition (also called a subsample) contains DNA with a larger proportion of cytosine modifications than the second subsample. Each partition is tagged differently. The first subsample undergoes a procedure that differently affects the first and second nucleotides in the DNA of the first subsample, where the first nucleotide is modified or unmodified, and the second nucleotide is modified or unmodified, different from the first nucleotide, and the first and second nucleotides have the same base-pairing specificity. The tagged nucleic acids are pooled together and then sequenced. Sequence reads are obtained and analyzed, including computer simulation (in silico) to distinguish the first and second nucleotides in the DNA of the first subsample. Tags are used to sort reads from different partitions. Analysis can be performed at the level of individual partitions and at the level of the entire nucleic acid population to detect genetic variation. For example, the analysis may include computer simulation analysis to determine genetic variations, such as CNVs, SNVs, insertions / deletions, and fusions in the nucleic acids of each partition. In some cases, computer simulation analysis may include determining chromatin structure. For example, the coverage of sequence reads can be used to determine the location of nucleosomes in chromatin. Higher coverage may be associated with higher nucleosome occupancy in a genomic region, while lower coverage may be associated with lower nucleosome occupancy or nucleosome depleted regions (NDRs).
[0098] Samples may include nucleic acids with various modifications, including post-replication modifications of nucleotides and binding to one or more proteins (typically non-covalent).
[0099] In the implementation scheme, the nucleic acid population is a population of nucleic acids obtained from serum, plasma, or blood samples of subjects suspected of having vegetations, tumors, or cancer, or previously diagnosed with vegetations, tumors, or cancer. The nucleic acid population includes nucleic acids with different levels of methylation. Methylation can occur by any one or more post-replication or post-transcriptional modifications. Post-replication modifications include modifications to nucleotide cytosine, particularly at the 5-position of the nucleobase, such as 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine, and 5-carboxycytosine. The affinity agent can be an antibody with desired specificity, a natural binding partner or a variant thereof (Bock et al., Nat Biotech 28: 1106-1114 (2010); Song et al., Nat Biotech 29: 68-72 (2011)), or, for example, an artificial peptide selected for specificity to a given target via phage display.
[0100] Examples of capture fractions envisioned herein include methyl-binding domains (MBDs) and methyl-binding proteins (MBPs) as described herein, including proteins such as MeCP2 and antibodies that preferentially bind to 5-methylcytosine. Similarly, partitioning of different forms of nucleic acids can be performed using histone-binding proteins that can separate histone-bound nucleic acids from free or unbound nucleic acids. Examples of histone-binding proteins that can be used in the methods disclosed herein include RBBP4, RbAp48, and SANT domain peptides. For some affinity agents and modifications, although binding to the agent may occur substantially all-or-nothing depending on whether the nucleic acid is modified, separation may be to a certain extent. In such cases, nucleic acids overrepresented in a modification bind to the agent to a greater extent than nucleic acids underrepresented in the modification. Optionally, modified nucleic acids may bind in an all-or-nothing manner. However, various levels of modification can then be eluted sequentially from the binding agent.
[0101] For example, in some implementations, partitioning can be binary or based on the degree / level of modification. For instance, a methyl-binding domain protein (e.g., the MethylMiner methylated DNA enrichment kit (ThermoFisherScientific)) can be used to partition all methylated fragments with unmethylated fragments. Subsequently, additional partitioning can include eluting fragments with different methylation levels by adjusting the salt concentration of the solution containing the methyl-binding domain and the binding fragment. As the salt concentration increases, fragments with higher methylation levels are eluted. In some cases, the final partitioning represents nucleic acids with different degrees of modification (overrepresentation or underrepresentation). Overrepresentation and underrepresentation can be defined by the number of modifications a nucleic acid carries relative to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine residues in the nucleic acids in a sample is 2, then nucleic acids containing more than two 5-methylcytosine residues are overrepresented, while nucleic acids with one or zero 5-methylcytosine residues are underrepresented. The purpose of affinity separation is to enrich over-represented nucleic acids in the binding phase and under-represented nucleic acids in the non-binding phase (i.e., in solution). The nucleic acids in the binding phase can be eluted before subsequent processing.
[0102] When using the MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific), sequential elution can be used to separate methylated regions with different levels of methylation. For example, low-methylated regions (e.g., unmethylated) can be separated from methylated regions by contacting a group of nucleic acids with MBDs attached to magnetic beads from the kit. The beads are used to isolate methylated nucleic acids from unmethylated nucleic acids. Subsequently, one or more elution steps are performed sequentially to elute nucleic acids with different methylation levels. For example, the first group of methylated nucleic acids can be eluted at a salt concentration of 160 mM or higher, such as at least 150 mM, at least 200 mM, at least 300 mM, at least 400 mM, at least 500 mM, at least 600 mM, at least 700 mM, at least 800 mM, at least 900 mM, at least 1000 mM, or at least 2000 mM. After such methylated nucleic acids have been eluted, magnetic separation is again used to separate nucleic acids with higher levels of methylation from those with lower levels of methylation. The elution and magnetic separation steps can be repeated to produce various partitions, such as hypomethylated partitions (representing no methylation), methylated partitions (representing low methylation levels), and hypermethylated partitions (representing high methylation levels).
[0103] In some methods, nucleic acids bound to an affinity separator undergo a washing step. This washing step removes nucleic acids that are weakly bound to the affinity agent. Such nucleic acids can be enriched with a degree of modification close to the mean or median (i.e., an intermediate value between nucleic acids that remain bound to the solid and those that do not upon initial contact with the agent). Affinity separation results in at least two, and sometimes three or more, partitions of nucleic acids with different degrees of modification. While the partitions remain separate, nucleic acids from at least one partition, and usually two or three (or more) partitions, are attached to nucleic acid tags, which are typically provided as part of an adaptor, and the nucleic acids in different partitions receive different tags that distinguish members of one partition from members of another. Tags attached to nucleic acid molecules in the same partition can be the same or different from each other. However, if they are different, the tags can have a shared encoding to identify the molecules to which they are attached as belonging to a particular partition. For more details on partitioning nucleic acid samples based on properties such as methylation, see WO2018 / 119452, which is incorporated herein by reference. In some implementations, nucleic acid molecules can be hierarchically separated into different partitions based on whether they bind to a specific protein or a fragment thereof and whether they do not bind to that specific protein or a fragment thereof.
[0104] Nucleic acid molecules can be fractionated based on DNA-protein binding. Protein-DNA complexes can be fractionated based on specific properties of the protein. Examples of such properties include various epitopes, modifications (e.g., histone methylation or acetylation), or enzymatic activities. Examples of proteins that can bind DNA and be used as the basis for fractionation can include, but are not limited to, protein A and protein G. Any suitable method can be used for fractionation of nucleic acid molecules based on protein-binding regions. Examples of methods for fractionation of nucleic acid molecules based on protein-binding regions include, but are not limited to, SDS-PAGE, chromatin immunoprecipitation (ChIP), heparin chromatography, and asymmetric field flow fractionation (AF4).
[0105] In some implementations, nucleic acid fractionation is performed by contacting the nucleic acid with the methylation-binding domain (“MBD”) of a methylation-binding protein (“MBP”). The MBD binds to 5-methylcytosine (5mC). The MBD is coupled to paramagnetic beads (such as Dynabeads® M-280 streptavidin) via a biotin linker. Fractionation into fractions with different degrees of methylation can be performed by eluting the fractions with increasing NaCl concentrations.
[0106] An exemplary method for molecular tag identification of MBD bead partition libraries using NGS is as follows: Extracted DNA samples (e.g., plasma DNA extracted from human samples) are physically partitioned using a methyl-binding domain protein-bead purification kit, retaining all eluents from the process for downstream processing.
[0107] Differential molecular tags and NGS-feasible linker sequences were applied in parallel to each partition. For example, hypermethylated, residual methylated ('washed'), and hypomethylated partitions were linked to NGS linkers with molecular tags.
[0108] All molecularly tagged partitions were reassembled and subsequently amplified using adaptor-specific DNA primer sequences.
[0109] Enrich / hybridize the recombined and amplified total library to target genomic regions of interest (e.g., cancer-specific genetic variations and differentially methylated regions).
[0110] The enriched total DNA library was re-amplified and tagged with samples. Different samples were pooled and subjected to multiplex assays on an NGS instrument.
[0111] Bioinformatics analysis of NGS data was performed, using molecular tags to identify unique molecules and deconvolving samples into molecules with differentially expressed molecular dividers (MBDs). This analysis can simultaneously produce information on the relative 5-methylcytosine content of genomic regions with standard gene sequencing / variation detection.
[0112] Examples of MBPs envisioned in this article include, but are not limited to: (a) MeCP2, a protein that preferentially binds to 5-methyl-cytosine compared to binding to unmodified cytosine.
[0113] (b) RPL26, PRP8 and DNA mismatch repair protein MHS6 preferentially bind to 5-hydroxymethyl-cytosine compared to binding to unmodified cytosine.
[0114] (c) FOXK1, FOXK2, FOXP1, FOXP4 and FOXI3 preferentially bind 5-formyl-cytosine compared to binding unmodified cytosine (Iurlaro et al., Genome Biol. 14: R119 (2013)).
[0115] (d) Antibodies specific to one or more methylated nucleotide bases.
[0116] Typically, elution varies with the number of methylation sites per molecule, with molecules having more methylation eluting at increasing salt concentrations. To elute DNA into different populations based on the degree of methylation, a series of elution buffers with increasing NaCl concentrations can be used. Salt concentrations can range from about 100 nM to about 2500 mM NaCl. In one embodiment, the process produces three (3) partitions. Molecules are contacted with a solution of a first salt concentration, and this solution contains molecules containing methyl-binding domains that can attach to capture moieties such as streptavidin. At the first salt concentration, one population of molecules will bind to MBD, and one population will remain unbound. The unbound population can be separated into a “hypomethylated” population. For example, the first partition representing hypomethylated DNA is the partition that remains unbound at low salt concentrations (e.g., 100 mM or 160 mM). The second partition representing moderately methylated DNA is eluted using a moderate salt concentration (e.g., between 100 mM and 2000 mM). This is also separated from the sample. The third partition, representing the highly methylated form of DNA, was eluted with a high salt concentration (e.g., at least about 2000 mM).
[0117] This disclosure also provides methods for analyzing nucleic acid populations, wherein at least some nucleic acids contain one or more modified cytosine residues, such as 5-methylcytosine and any other modifications previously described. In these methods, after partitioning, a sample of nucleic acid subsamples is contacted with an adaptor containing one or more cytosine residues modified at the 5C position (such as 5-methylcytosine). Preferably, all cytosine residues in such an adaptor are also modified, or all such cytosine residues in the primer-binding region of the adaptor are modified. The adaptor is attached to both ends of nucleic acid molecules in the population. Preferably, the adaptor contains a sufficient number of different tags such that the number of tag combinations results in a high probability, for example, 95%, 99%, or 99.9%, that two nucleic acids with the same start and end points receive different tag combinations. The primer-binding sites in such an adaptor can be the same or different, but are preferably the same. After the adaptor is attached, the nucleic acid is amplified by primers that bind to the primer-binding sites of the adaptor. The amplified nucleic acid is divided into a first aliquot and a second aliquot. With or without further processing, the first aliquot is sequenced. This determines the sequence data of molecules in the first aliquot regardless of the initial methylation state of the nucleic acid molecules. Nucleic acid molecules in the second aliquot undergo a procedure that differently affects the first and second nucleobases in the DNA, where the first nucleobase includes cytosine modified at position 5, and the second nucleobase includes unmodified cytosine. This procedure can be bisulfite treatment or another procedure that converts unmodified cytosine to uracil. The nucleic acids that have undergone this procedure are then amplified using primers targeting the original primer binding sites of the adaptor attached to the nucleic acid. Now only the nucleic acid molecules initially attached to the adaptor (unlike their amplification products) are amplified because these nucleic acids retain cytosine at the primer binding sites of the adaptor, while the amplification products have lost the methylation of these cytosine residues, which have been converted to uracil during bisulfite treatment. Therefore, only the original molecules in the population (at least some of which are methylated) undergo amplification. After amplification, these nucleic acids are sequenced. Comparison of sequences determined from the first and second aliquots can particularly indicate which cytosines in the nucleic acid population have undergone methylation.
[0118] Such analysis can be performed using the following exemplary procedure. After partitioning, the two ends of the methylated DNA are ligated to a Y-shaped adaptor containing primer binding sites and a tag. The cytosine in the adaptor is modified at position 5 (e.g., 5-methylated). This modification of the adaptor serves to protect the primer binding sites in subsequent transformation steps (e.g., bisulfite treatment, TAP transformation, or any other transformation that does not affect the modified cytosine but affects the unmodified cytosine). After adaptor attachment, the DNA molecule is amplified. The amplified product is divided into two aliquots for sequencing with and without transformation. The untransformed aliquot may undergo sequence analysis with or without further treatment. The other aliquot undergoes a procedure that differently affects the first and second nucleotides in the DNA, where the first nucleotide includes the cytosine modified at position 5, and the second nucleotide includes the unmodified cytosine. This procedure may be bisulfite treatment or another procedure to convert the unmodified cytosine to uracil. When contacted with primers specific to the original primer binding site, only primer binding sites protected by cytosine modification can support amplification. Therefore, only the original molecule, not copies from the first amplification, undergoes further amplification. The further amplified molecules then undergo sequence analysis. The sequences from the two aliquots can then be compared. In the isolation scheme discussed above, the nucleic acid tag in the adaptor is not used to distinguish between methylated and unmethylated DNA, but rather to distinguish nucleic acid molecules within the same partition.
[0119] Enrichment / capture steps; amplification; adaptors; barcodes Capture set The method disclosed herein includes steps of subjecting a first subsample to procedures that differently affect a first nucleotide and a second nucleotide in the DNA of the first subsample, wherein the first nucleotide is a modified or unmodified nucleotide, the second nucleotide is a modified or unmodified nucleotide different from the first nucleotide, and the first and second nucleotides have the same base-pairing specificity. In some embodiments, if the first nucleotide is modified or unmodified adenine, then the second nucleotide is modified or unmodified adenine; if the first nucleotide is modified or unmodified cytosine, then the second nucleotide is modified or unmodified cytosine; if the first nucleotide is modified or unmodified guanine, then the second nucleotide is modified or unmodified guanine; if the first nucleotide is modified or unmodified thymine, then the second nucleotide is modified or unmodified thymine (wherein, for the purposes of this step, modified and unmodified uracil are included in modified thymine).
[0120] In some embodiments, the first nucleobase is a modified or unmodified cytosine, and then the second nucleobase is a modified or unmodified cytosine. For example, the first nucleobase may comprise unmodified cytosine (C), and the second nucleobase may comprise one or more of 5-methylcytosine (mC) and 5-hydroxymethylcytosine (hmC). Alternatively, the second nucleobase may comprise C, and the first nucleobase may comprise one or more of mC and hmC. Other combinations are also possible, such as those indicated, for example, in the overview above and the discussion below, such as where one of the first and second nucleobases comprises mC and the other comprises hmC.
[0121] In some implementations, the procedure that differently affects the first and second nucleobases in the DNA of the first subsample includes bisulfite conversion. Bisulfite treatment converts unmodified cytosine and certain modified cytosine nucleotides (e.g., 5-formylcytosine (fC) or 5-carboxycytosine (caC)) into uracil, while other modified cytosines (e.g., 5-methylcytosine and 5-hydroxymethylcytosine) are not converted. Therefore, in the case of bisulfite conversion, the first nucleobase comprises one or more of unmodified cytosine, 5-formylcytosine, 5-carboxycytosine, or other bisulfite-affected cytosine forms, and the second nucleobase may comprise one or more of mC and hmC, such as mC and optionally hmC. Sequencing of the bisulfite-treated DNA identifies the position read as a cytosine as the mC position or the hmC position. Simultaneously, positions read as T are identified as T or bisulfite-susceptible forms of C, such as unmodified cytosine, 5-formylcytosine, or 5-carboxycytosine. Therefore, bisulfite conversion of the first subsample as described herein facilitates the identification of positions containing mC or hmC using sequence reads obtained from the first subsample. For an exemplary description of bisulfite conversion, see, for example, Moss et al., Nat Commun. 2018; 9: 5068.
[0122] In some embodiments, the procedures that differently affect the first and second nucleotides of the DNA in the first subsample include oxidized bisulfite (Ox-BS) conversion. In some embodiments, the procedures that differently affect the first and second nucleotides of the DNA in the first subsample include Tet-assisted substituted borane reducing agent conversion, optionally wherein the substituted borane reducing agent is 2-methylpyridineborane, pyridineborane, tert-butylamineborane, or aminoborane. In some embodiments, the procedures that differently affect the first and second nucleotides of the DNA in the first subsample include chemically assisted substituted borane reducing agent conversion, optionally wherein the substituted borane reducing agent is 2-methylpyridineborane, pyridineborane, tert-butylamineborane, or aminoborane. In some implementations, procedures that differently affect the first and second nucleobases in the DNA of the first subsample include APOBEC-coupled epigenetic (ACE) transformation.
[0123] In some implementations, procedures that differently affect the first and second nucleotides in the DNA of the first subsample include enzymatic conversion of the first nucleotide, for example, as in EM-Seq. See, for example, Vaisvila R et al. (2019) Em-seq: Detection of DNA methylation at single base resolution from picograms of DNA. bioRxiv; DOI: 10.1101 / 2019.12.20.884692, available at www.biorxiv.org / content / 10.1101 / 2019.12.20.884692v1. For example, TET2 and T4-βGT can be used to convert 5mC and 5hmC into substrates that cannot be deaminated by deaminases (e.g., APOBEC3A), and then the deaminase (e.g., APOBEC3A) can be used to deaminate unmodified cytosine, converting it to uracil.
[0124] In some implementations, the procedure that differently affects the first nucleobase in the DNA of the first subsample and the second nucleobase in the DNA includes separating the DNA that initially contains the first nucleobase from the DNA that initially does not contain the first nucleobase.
[0125] In some embodiments, the first nucleobase is a modified or unmodified adenine, and the second nucleobase is a modified or unmodified adenine. In some embodiments, the modified adenine is N6-methyladenine (mA). In some embodiments, the modified adenine is one or more of N6-methyladenine (mA), N6-hydroxymethyladenine (hmA), or N6-formyladenine (fA).
[0126] Techniques including methylated DNA immunoprecipitation (MeDIP) can be used to separate DNA containing modified bases (such as mA) from other DNA. See, for example, Kumar et al., Frontiers Genet. 2018; 9: 640; Greer et al., Cell 2015; 161: 868-878. Antibodies specific to mA are described in Sun et al., Bioessays 2015; 37:1155-62. Antibodies against various modified nucleobases (such as thymine / uracil forms, including halogenated forms such as 5-bromouracil) are commercially available. Various modified bases can also be detected based on changes in their base pairing specificity. For example, hypoxanthine is a modified form of adenine that can be produced by deamination and is read as G in sequencing. See, for example, U.S. Patent 8,486,630; Brown, Genomes, 2nd Edition, John Wiley & Sons, Inc., New York, NY, 2002, Chapter 14, “Mutation, Repair, and Recombination”.
[0127] Epigenetic target region set In some embodiments, the methods disclosed herein include the step of capturing one or more target regions of DNA, such as cfDNA. Capture can be performed using any suitable method known in the art. In some embodiments, capture includes contacting the DNA to be captured with a set of target-specific probes. The target-specific probe set may have any of the characteristics of the target-specific probe set described herein, including but not limited to the characteristics set forth above and in the probe-related sections below. One or more subsamples prepared during the methods disclosed herein may be captured. In some embodiments, DNA is captured from at least a first subsample or a second subsample, e.g., at least a first subsample and a second subsample. If the first subsample undergoes a separation step (e.g., separating DNA initially containing a first nucleobase (e.g., hmC) from DNA initially not containing a first nucleobase, such as hmC-seal), capture can be performed on any one, any two, or all of the DNA initially containing a first nucleobase (e.g., hmC), the DNA initially not containing a first nucleobase, and the second subsample. In some embodiments, the subsamples are differentially tagged (e.g., as described herein) and then pooled prior to capture.
[0128] The capture step can be performed using conditions suitable for a specific nucleic acid hybridization, which are typically dependent to some extent on the characteristics of the probe, such as length, base composition, etc. Given the general knowledge of nucleic acid hybridization in the art, those skilled in the art will be familiar with appropriate conditions. In some embodiments, a complex of the target-specific probe and DNA is formed.
[0129] In some embodiments, the method described herein includes capturing more than one set of target regions of cfDNA obtained from a test subject. The target regions include epigenetic target regions that may exhibit differences in methylation levels and / or fragmentation patterns, depending on whether they originate from tumor cells or healthy cells. The target regions also include sequence-variable target regions that may exhibit sequence differences, depending on whether they originate from tumor cells or healthy cells. The capture step produces a capture set of cfDNA molecules, and within the capture set of cfDNA molecules, cfDNA molecules corresponding to the sequence-variable target region set are captured with a greater capture yield than cfDNA molecules corresponding to the epigenetic target region set. For further discussion of the capture steps, capture yield, and related aspects, see WO2020 / 160414, which is incorporated herein by reference for all purposes.
[0130] In some implementations, the method described herein includes contacting cfDNA obtained from a test subject with a target-specific probe set, wherein the target-specific probe set is configured to capture cfDNA corresponding to a sequence-variable target region set with a greater capture yield than cfDNA corresponding to a set of epigenetic target regions.
[0131] Capturing cfDNA corresponding to a set of sequence-variable target regions at a higher capture yield than that corresponding to the epigenetic target region set is advantageous, because analyzing sequence-variable target regions with sufficient confidence or accuracy may require a greater sequencing depth than analyzing epigenetic target regions. The amount of data required to determine fragmentation patterns (e.g., testing for perturbations of transcription start sites or CTCF binding sites) or fragment abundance (e.g., in hypermethylated and hypomethylated regions) is generally less than the amount of data required to determine the presence or absence of cancer-related sequence mutations. Capturing the target region set at different yields can facilitate sequencing the target regions to different sequencing depths within the same sequencing run (e.g., using pooled mixtures and / or within the same sequencing pool).
[0132] In various embodiments, the method further includes sequencing the captured cfDNA to, for example, different sequencing depths for epigenetic target sets and sequence-variable target sets, consistent with those discussed herein. In some embodiments, the complex of the target-specific probe and DNA is separated from DNA not bound to the target-specific probe. For example, in cases where the target-specific probe is covalently or nonvalently bound to a solid support, washing or aspiration steps can be used to separate the unbound material. Alternatively, chromatography can be used where the complex has different chromatographic properties than the unbound material (e.g., where the probe contains ligands that bind to chromatographic resins).
[0133] As discussed in detail elsewhere herein, target-specific probe sets may include more than one set, such as probes for a sequence-variable target region set and probes for an epigenetic target region set. In some such embodiments, the capture step is performed simultaneously using probes for sequence-variable targets and probes for epigenetic targets in the same container; for example, probes for both the sequence-variable and epigenetic target regions are in the same composition. This method provides a relatively more efficient workflow. In some embodiments, the concentration of probes for the sequence-variable target region set is greater than the concentration of probes for the epigenetic target region set.
[0134] Optionally, a capture step is performed in a first container with a sequence-variable target region probe set and in a second container with an epigenetic target region probe set, or a contact step is performed in the first time and the first container with a sequence-variable target region probe set and in a second time before or after the first time with an epigenetic target region probe set. This method allows for the preparation of separate first and second compositions comprising captured DNA corresponding to a sequence-variable target region set and captured DNA corresponding to an epigenetic target region set. The compositions can be processed individually as desired (e.g., graded based on methylation, as described elsewhere herein) and recombine in appropriate proportions to provide material for further processing and analysis, such as sequencing.
[0135] In some implementations, DNA is amplified. In some implementations, amplification occurs before the capture step. In some implementations, amplification occurs after the capture step.
[0136] In some implementations, the DNA contains an adaptor. This can be performed simultaneously with the amplification process, for example, by providing the adaptor in the 5' portion of the primer, as described above. Alternatively, the adaptor can be added by other methods such as ligation.
[0137] In some implementations, the DNA contains a tag, which may be a barcode or contain a barcode. The tag can aid in identifying the origin of the nucleic acid. For example, a barcode can be used to allow identification of the DNA's origin, such as a subject, after pooling more than one sample for parallel sequencing. This can be performed concurrently with the amplification procedure, for example, by providing a barcode in the 5' portion of the primer, as described above. In some implementations, the adaptor and the tag / barcode are provided by the same primer or primer set. For example, the barcode may be located at the 3' of the adaptor and the 5' of the target hybridization portion of the primer. Alternatively, the barcode can be added by other methods, such as ligation, optionally along with the adaptor in the same ligation substrate.
[0138] Further details regarding amplification, labeling, and barcodes are discussed in the following “General Characteristics of the Method” section, and these details can be combined, to a degree of feasibility, with any of the aforementioned implementation schemes and the implementation schemes described in the “Introduction and Overview” section.
[0139] Hypermethylated variable target region In some embodiments, a capture set of DNA (e.g., cfDNA) is provided. For the disclosed methods, the capture set of DNA may be provided, for example, by performing a capture step after a partitioning step as described herein. The capture set may include DNA corresponding to a set of sequence-variable target regions, DNA corresponding to a set of epigenetic target regions, or a combination thereof. In some embodiments, the amount of captured sequence-variable target region DNA is greater than the amount of captured epigenetic target region DNA when normalized for differences in target region size (footprint size).
[0140] Optionally, a first capture set and a second capture set may be provided, comprising DNA corresponding to a sequence-variable target region set and DNA corresponding to an epigenetic target region set, respectively. The first capture set and the second capture set may be combined to provide a combined capture set.
[0141] In some embodiments, where the capture set, which includes DNA corresponding to both sequence-variable target regions and epigenetic target regions, comprises a combination of capture sets as discussed above, the DNA corresponding to the sequence-variable target regions may be present at a higher concentration than the DNA corresponding to the epigenetic target regions, for example, 1.1 to 1.2 times higher, 1.2 to 1.4 times higher, 1.4 to 1.6 times higher, 1.6 to 1.8 times higher, 1.8 to 2.0 times higher, or 2.0 to 2.2 times higher. Concentration, 2.2 to 2.4 times larger concentration, 2.4 to 2.6 times larger concentration, 2.6 to 2.8 times larger concentration, 2.8 to 3.0 times larger concentration, 3.0 to 3.5 times larger concentration, 3.5 to 4.0 times larger concentration, 4.0 to 4.5 times larger concentration, 4.5 to 5.0 times larger concentration, 5.0 to 5.5 times larger concentration, 5.5 to 6.0 times larger concentration, 6.0 to 6.5 times larger concentration, 6.5 to 7.0 times larger concentration, 7.0 to 7. 5 times larger concentration, 7.5 to 8.0 times larger concentration, 8.0 to 8.5 times larger concentration, 8.5 to 9.0 times larger concentration, 9.0 to 9.5 times larger concentration, 9.5 to 10.0 times larger concentration, 10 to 11 times larger concentration, 11 to 12 times larger concentration, 12 to 13 times larger concentration, 13 to 14 times larger concentration, 14 to 15 times larger concentration, 15 to 16 times larger concentration, 16 to 17 times larger concentration, 17 to 18 times larger concentration, 18 times larger to Concentrations ranging from 19 to 20 times, 20 to 30 times, 30 to 40 times, 40 to 50 times, 50 to 60 times, 60 to 70 times, 70 to 80 times, 80 to 90 times, 90 to 100 times, 10 to 20 times, 10 to 40 times, 10 to 50 times, 10 to 70 times, or 10 to 100 times. The degree of concentration variation is calculated based on normalization for the target footprint size, as discussed in the definitions section.
[0142] Hypomethylated variable target region An epigenetic target set may include one or more types of target regions that can distinguish DNA from DNA from vegetative (e.g., tumor or cancer) cells from DNA from healthy cells (e.g., non-vegetative circulating cells). Example types of such regions are discussed in detail herein. An epigenetic target set may also include one or more control regions, such as those described herein. In some embodiments, the epigenetic target set has a footprint of at least 100 kb, for example, at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, the epigenetic target set has a footprint in the range of 100-1000 kb, for example, 100-200 kb, 200-300 kb, 300-400 kb, 400-500 kb, 500-600 kb, 600-700 kb, 700-800 kb, 800-900 kb, and 900-1000 kb.
[0143] Composition comprising captured DNA In some implementations, the epigenetic target set includes one or more hypermethylated variable target regions. Typically, a hypermethylated variable target region refers to a region in which, for example in a cfDNA sample, an observed increase in methylation levels indicates an increased likelihood that the sample (e.g., cfDNA) contains DNA produced by invasive cells (such as tumor cells or cancer cells). For example, hypermethylation of tumor suppressor gene promoters has been repeatedly observed. See, for example, Kang et al., Genome Biol. 18:53 (2017) and the references cited therein. In examples, a hypermethylated variable target region may include a region in which the methylation is not necessarily different from that of DNA from the same type of healthy tissue, but is indeed different from that of typical cfDNA in healthy subjects (e.g., having more methylation). For example, such a hypermethylated variable target region can be used at least in part to detect cancer when the presence of cancer leads to increased cell death (such as apoptosis corresponding to the tissue type of cancer). In some implementations, the hypermethylated variable target region includes one or more genomic regions in which the methylation status of cfDNA molecules in these regions is not different from that of cfDNA from healthy subjects, but the presence / increased amount of hypermethylated cfDNA in these regions indicates a specific tissue type (e.g., cancer origin) and is presented as cfDNA entering circulation due to increased apoptosis (e.g., tumor shedding).
[0144] Hypermethylated target regions can be obtained, for example, from the Cancer Genome Atlas. Kang et al., Genome Biology 18:53 (2017) describe a probabilistic approach to constructing a cancer locator using hypermethylated target regions from the breast, colon, kidney, liver, and lung. In some embodiments, the hypermethylated target regions can be specific to one or more types of cancer. Thus, in some embodiments, the hypermethylated target regions comprise one, two, three, four, or five subsets of hypermethylated target regions that collectively exhibit hypermethylation in one, two, three, four, or five of the following cancers: breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.
[0145] In some embodiments, probes targeting a set of epigenetic target regions include probes specific to one or more hypermethylated variable target regions. Hypermethylated variable target regions can be any of the hypermethylated variable target regions listed above. For example, in some embodiments, probes specific to hypermethylated variable target regions include probes specific to more than one locus listed in Table 1 (e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1). In some embodiments, probes specific to hypermethylated variable target regions include probes specific to more than one locus listed in Table 2 (e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 2). In some embodiments, probes specific to hypermethylated variable target regions include probes specific to more than one locus listed in Table 1 or Table 2 (e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the loci listed in Table 1 or Table 2). In some embodiments, for each locus included as a target region, there may be one or more probes having hybridization sites that bind between the transcription start site and the stop codon (or the final stop codon for alternatively spliced genes) of that gene. In some embodiments, one or more probes bind within 300 bp (e.g., within 200 bp or 100 bp) of the listed positions. In some embodiments, the probes have hybridization sites that overlap with the positions listed above. In some embodiments, probes specific to hypermethylated target regions include probes specific to one, two, three, four, or five subsets of hypermethylated target regions that collectively exhibit hypermethylation in one, two, three, four, or five of breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.
[0146] Computer system, processing of real world evidence (RWE) Universal hypomethylation is a phenomenon commonly observed in a variety of cancers. See, for example, Hon et al., GenomeRes. 22:246-258 (2012) (breast cancer); Ehrlich, Epigenomics 1:239-259 (2009) (a review article noting observations of hypomethylation in colon cancer, ovarian cancer, prostate cancer, leukemia, hepatocellular carcinoma, and cervical cancer). For example, regions that are normally methylated in healthy cells (such as repetitive elements (e.g., LINE1 elements, Alu elements, centromere tandem repeats, pericentromere tandem repeats, and satellite DNA) and intergenic regions) may show reduced methylation in tumor cells. Therefore, in some embodiments, the epigenetic target set includes hypomethylated variable target regions, where the observed reduction in methylation levels indicates an increased likelihood that the sample (e.g., cfDNA) contains DNA produced by cytoplasmic cells (such as tumor cells or cancer cells). In examples, a hypomethylated variable target region may include a region whose methylation state is not necessarily different in cancerous tissue relative to DNA from the same type of healthy tissue, but is indeed different in methylation (e.g., less methylated) relative to cfDNA typical of healthy subjects. For example, such a hypomethylated variable target region can be used at least in part to detect cancer when the presence of cancer leads to increased cell death (such as apoptosis corresponding to the tissue type of cancer). In some embodiments, the hypomethylated variable target region comprises one or more genomic regions in which the methylation state of cfDNA molecules in these regions is not different from cfDNA from healthy subjects, but the amount of hypomethylated cfDNA present / increased in these regions indicates a specific tissue type (e.g., cancer origin) and is presented as cfDNA entering circulation accompanied by increased apoptosis (e.g., tumor shedding).
[0147] In some embodiments, the hypomethylated variable target region includes repeating elements and / or intergenic regions. In some embodiments, repeating elements include one, two, three, four, or five of the following: LINE1 elements, Alu elements, centromere tandem repeat sequences, pericentromere tandem repeat sequences, and / or satellite DNA.
[0148] Example specific genomic regions exhibiting cancer-related hypomethylation include nucleotides 8403565-8953708 and 151104701-151106035 on human chromosome 1. In some embodiments, the hypomethylated variable target region overlaps with or includes one or both of these regions.
[0149] In some implementations, probes targeting a set of epigenetic target regions include probes specific to one or more hypomethylated variable target regions. Hypomethylated variable target regions can be any of the hypomethylated target regions listed above. For example, probes specific to one or more hypomethylated variable target regions may include probes targeting regions such as repetitive elements (e.g., LINE1 elements, Alu elements, centromere tandem repeats, pericentromere tandem repeats, and satellite DNA) and intergenic regions that are normally methylated in healthy cells but may exhibit reduced methylation in tumor cells.
[0150] In some embodiments, probes specific to hypomethylated variable target regions include probes specific to repetitive elements and / or intergenic regions. In some embodiments, probes specific to repetitive elements include probes specific to one, two, three, four, or five of the following: LINE1 elements, Alu elements, centromere tandem repeat sequences, pericentromere tandem repeat sequences, and / or satellite DNA.
[0151] Example probes specific to genomic regions exhibiting cancer-related hypomethylation include probes specific to nucleotides 8403565-8953708 and / or 151104701-151106035 of human chromosome 1. In some embodiments, probes specific to hypomethylated variable target regions include probes specific to regions that overlap with or contain nucleotides 8403565-8953708 and / or 151104701-151106035 of human chromosome 1.
[0152] Probes used to detect this set of regions can include probes for detecting genomic regions of interest (hotspot regions) and nucleosome-sensing probes (e.g., KRAS codons 12 and 13), and can be designed to optimize capture based on analysis of cfDNA coverage and fragment size variation and GC sequence composition influenced by nucleosome binding patterns. The regions used in this paper can also include non-hotspot regions optimized based on nucleosome location and GC models. Subjects In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject with cancer. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject suspected of having cancer. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject with a tumor. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject suspected of having a tumor. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject with a growth. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject suspected of having a growth. In some embodiments, the DNA (e.g., cfDNA) is obtained from a subject in remission from a tumor, cancer, or growth (e.g., after chemotherapy, surgical resection, radiation, or a combination thereof). In any of the foregoing embodiments, the cancer, tumor, or growth, or suspected cancer, tumor, or growth, may be of the lung, colon, rectum, kidney, breast, prostate, or liver. In some embodiments, the cancer, tumor, or growth, or suspected cancer, tumor, or growth, is of the lung. In some embodiments, the cancer, tumor, or growth, or suspected cancer, tumor, or growth, is of the colon or rectum. In some embodiments, the cancer, tumor, or growth, or suspected cancer, tumor, or growth, is of the breast. In some embodiments, the cancer, tumor, or growth, or suspected cancer, tumor, or growth, is of the prostate. In any of the foregoing embodiments, the subject may be a human subject.
[0153] In some embodiments, the sequence-variable target probe set has a footprint of at least 0.5 kb (e.g., at least 1 kb, at least 2 kb, at least 5 kb, at least 10 kb, at least 20 kb, at least 30 kb, or at least 40 kb). In some embodiments, the epigenetic target probe set has a footprint in the range of 0.5–100 kb (e.g., 0.5–2 kb, 2–10 kb, 10–20 kb, 20–30 kb, 30–40 kb, 40–50 kb, 50–60 kb, 60–70 kb, 70–80 kb, 80–90 kb, and 90–100 kb).
[0154] In some embodiments, probes specific to a sequence-variable target set include probes specific to targets from at least 10, 20, 30, or 35 cancer-related genes, such as AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53, and U2AF1.
[0155] Figure 6 This document provides a combination of a first group and a second group comprising captured DNA. The first group may comprise or be derived from DNA with a greater proportion of cytosine modifications than the second group. The first group may comprise a form of a first nucleobase initially present in the DNA with altered base pairing specificity and a second nucleobase without altered base pairing specificity, wherein the form of the first nucleobase initially present in the DNA before the alteration of base pairing specificity is a modified or unmodified nucleobase, and the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the form of the first nucleobase initially present in the DNA before the alteration of base pairing specificity and the second nucleobase have the same base pairing specificity. The second group does not comprise a form of the first nucleobase initially present in the DNA with altered base pairing specificity. In some embodiments, the cytosine modification is cytosine methylation. In some embodiments, the first nucleobase is modified or unmodified cytosine, and the second nucleobase is modified or unmodified cytosine. The first and second nucleobases can be any nucleobases discussed in the overview or regarding procedures that subject the first subsample to different effects on the first and second nucleobases in the DNA of the first subsample.
[0156] In some implementations, the first group includes sequence tags selected from one or more sequence tags in the first group, and the second group includes sequence tags selected from one or more sequence tags in the second group, wherein the sequence tags in the second group are different from the sequence tags in the first group. The sequence tags may include barcodes.
[0157] In some embodiments, the first population comprises protected hmC, such as glucosylated hmC. In some embodiments, the first population undergoes any of the transformation procedures discussed herein, such as bisulfite transformation, Ox-BS transformation, TAB transformation, ACE transformation, TAP transformation, TAPSβ transformation, or CAP transformation. In some embodiments, the first population is protected by hmC followed by deamination of mC and / or C. In some combined embodiments, the first population comprises or is derived from DNA having a larger proportion of cytosine modification than the second population, and the first population comprises both first and second subpopulations, and the first nucleotide is a modified or unmodified nucleotide, the second nucleotide is a modified or unmodified nucleotide different from the first nucleotide, and the first and second nucleotides have the same base-pairing specificity. In some embodiments, the second population does not contain the first nucleotide. In some embodiments, the first nucleotide is a modified or unmodified cytosine, and the second nucleotide is a modified or unmodified cytosine, optionally wherein the modified cytosine is mC or hmC. In some embodiments, the first nucleobase is a modified or unmodified adenine, and the second nucleobase is a modified or unmodified adenine, optionally wherein the modified adenine is mA.
[0158] In some embodiments, the first nucleobase (e.g., modified cytosine) is biotinylated. In some embodiments, the first nucleobase (e.g., modified cytosine) is a product of Huisgen cycloaddition of β-6-azido-glucosyl-5-hydroxymethylcytosine, the product containing an affinity tag (e.g., biotin).
[0159] In any of the combinations described herein, the captured DNA may include cfDNA. The captured DNA may have any of the characteristics described herein with respect to the capture set, including, for example, a higher concentration of DNA corresponding to a sequence-variable target region set than the concentration of DNA corresponding to an epigenetic target region set (as normalized for footprint size as discussed above). In some embodiments, the captured DNA contains a sequence tag, which may be added to the DNA as described herein. Typically, the inclusion of a sequence tag results in DNA molecules that differ from their naturally occurring, untagged form.
[0160] This combination may also include the probe set or sequencing primers described herein, each of which may differ from naturally occurring nucleic acid molecules. For example, the probe set described herein may contain a capture portion, and the sequencing primers may contain a non-naturally occurring marker.
[0161] Figure 9 The methods of this disclosure can be implemented using or by means of a computer system. For example, such a method may include: partitioning a sample into more than one subsample, said more than one subsample including a first subsample and a second subsample, wherein the first subsample contains DNA with a larger proportion of cytosine modification than the second subsample; subjecting the first subsample to procedures that differently affect a first nucleobase and a second nucleobase in the DNA of the first subsample, wherein the first nucleobase is a modified or unmodified nucleobase, the second nucleobase is a modified or unmodified nucleobase different from the first nucleobase, and the first and second nucleobases have the same base-pairing specificity; and sequencing the DNA in the first subsample and the DNA in the second subsample in a manner that distinguishes the first and second nucleobases in the DNA of the first subsample.
[0162] In one aspect, this disclosure provides a non-transitory computer-readable medium including computer-executable instructions that, when executed by at least one electronic processor, perform at least a portion of a method comprising: collecting cfDNA from a test subject; capturing more than one set of target regions from the cfDNA, wherein the more than one set of target regions includes a sequence-variable target region set and an epigenetic target region set, thereby generating a capture set of cfDNA molecules; sequencing the captured cfDNA molecules, wherein the captured cfDNA molecules of the sequence-variable target region set are sequenced to a deeper sequencing depth than the captured cfDNA molecules of the epigenetic target region set; obtaining more than one sequence read generated by a nucleic acid sequencer through sequencing the captured cfDNA molecules; mapping the more than one sequence read to one or more reference sequences to generate mapped sequence reads; and processing the mapped sequence reads corresponding to the sequence-variable target region set and the epigenetic target region set to determine the likelihood that the subject has cancer.
[0163] The code can be pre-compiled and configured for use with machines having processors suitable for executing the code, or it can be compiled during runtime. The code can be provided in a programming language, which can be selected to enable the code to be executed in a pre-compiled or just-in-time (JIT) compiled manner.
[0164] Additional details relating to computer systems and networks, databases, and computer program products are provided, for example, in the following: Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th edition (2011); Kurose, Computer Networking: A Top-Down Approach, Pearson, 7th edition (2016); Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th edition (2010); Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11th edition (2014); Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd edition (2006); and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), all of which are incorporated herein by reference in their entirety.
[0165] Figure 6 This is a flowchart illustrating an exemplary training method 900 for generating an ML module 830 using a training module 820. The training module 820 can implement supervised, unsupervised, and / or semi-supervised (e.g., reinforcement-based) machine learning-based classification models 840. Figure 7 The method 900 shown in the middle figure is an example of a supervised learning method; variations of this example of training methods are discussed below, however, other training methods can be similarly implemented to train unsupervised and / or semi-supervised machine learning models.
[0166] Training method 900 may determine (e.g., access, receive, retrieve, etc.) data in step 910. The data may include tumor-derived body fluid sample data / non-tumor-derived body fluid sample data. The data may include sequence data of one or more sequence fragments / reads and / or variants, epigenetic data, and / or fragmentomics data, each sequence fragment / read and / or variant having a specified tumor-derived or non-tumor-derived origin.
[0167] Training method 900 can generate training and testing datasets in step 920. Training and testing datasets can be generated by randomly assigning data to either the training or testing dataset. In some implementations, the computational parameters and associated experimental parameters can be assigned to the training or testing data in a less random manner. As an example, a majority of the computational parameters and associated experimental parameters can be used to generate the training dataset. For instance, 75% of the computational parameters and associated experimental parameters can be used to generate the training dataset, and 25% can be used to generate the testing dataset. In another example, 80% of the computational parameters and associated experimental parameters can be used to generate the training dataset, and 20% can be used to generate the testing dataset.
[0168] Training method 900 may determine (e.g., extract, select, etc.) one or more features in step 930, which may be used by, for example, a classifier to distinguish between tumor-derived and non-tumor-derived statuses. As an example, training method 900 may determine a feature set from tumor-derived body fluid sample data / non-tumor-derived body fluid sample data. In another example, a feature set may be determined from data different from tumor-derived body fluid sample data / non-tumor-derived body fluid sample data in the training dataset or test dataset. Such additional data may be used to determine an initial feature set, which may be further reduced using the training dataset.
[0169] Training method 900 can train one or more machine learning models using one or more features in step 940. In one instance, supervised learning can be used to train the machine learning models. In another instance, other machine learning techniques, including unsupervised learning and semi-supervised learning, can be employed. The machine learning model trained in 940 can be selected based on different criteria, depending on the problem to be solved and / or the data available in the training dataset. For example, machine learning classifiers may suffer from varying degrees of bias. Therefore, more than one machine learning model can be trained in 940, and optimized, improved, and cross-validated in step 950.
[0170] Training method 900 may select one or more machine learning models at 960 to build a predictive model. A test dataset may be used to evaluate the predictive model. At step 970, the predictive model may analyze the test dataset and generate predicted tumor / non-tumor origin states. The predicted tumor / non-tumor origins may be evaluated at step 980 to determine if such values have reached the desired level of accuracy. The performance of the predictive model may be evaluated in various ways based on multiple true positive, false positive, true negative, and / or false negative classifications indicated by the predictive model for more than one data point.
[0171] For example, a false positive in a predictive model can refer to the number of times a predictive model incorrectly classifies a sequence fragment / read and / or variant that is not actually of tumor origin as of tumor origin. Conversely, a false negative in a predictive model can refer to the number of times a machine learning model classifies a sequence fragment / read and / or variant as of non-tumor origin when it is actually of tumor origin. True negatives and true positives can refer to the number of times a predictive model correctly classifies one or more sequence fragments / reads and / or variants. Related to these measures are the concepts of recall and precision. Typically, recall is the ratio of true positives to the sum of true positives and false negatives, quantifying the sensitivity of the predictive model. Similarly, precision is the ratio of true positives to the sum of true positives and false positives. When the desired level of accuracy is reached, the training phase ends and the prediction model can be output in step 990 (e.g., ML module 830); however, if the desired level of accuracy is not reached, subsequent iterations of the training method 900 can be performed from step 910 and changes can be introduced, such as, for example, considering a larger dataset.
[0172] Figure 8 This is a diagram illustrating an exemplary process flow for classifying sequence fragments / reads and / or variants as of or not of tumor origin using a machine learning-based classifier. Figure 9As illustrated in the diagram, sequence data, epigenetic data, and / or fragmentomics data of unclassified sequence fragments / reads and / or variants 1010 can be provided as input to ML module 830. ML module 830 can use one or more machine learning-based classifiers to process the sequence data, epigenetic data, and / or fragmentomics data of unclassified sequence fragments / reads and / or variants 1010 to arrive at a prediction result 1020. Prediction result 1020 can identify one or more attributes of the sequence data, epigenetic data, and / or fragmentomics data of unclassified sequence fragments / reads and / or variants 1010. For example, classification result 1020 can identify the origin of the sequence fragments / reads and / or variants 1010 (e.g., whether the sequence fragments / reads and / or variants are of tumor origin or non-tumor origin). Therefore, in the implementation, a method implemented using a network-based computer system is disclosed, the computer system including one or more processors, a network interface and one or more memories, the method comprising retrieving sequence data, epigenetic data and / or fragment omics data having an indicated tumor origin or a non-tumor origin through the computer system; and training machine learning models by fitting one or more models to the sequence data, epigenetic data and / or fragment omics data through one or more processors, each of the one or more models being configured to receive sequence data, epigenetic data and / or fragment omics data of an individual as input and provide a prediction of whether the individual has or has developed a tumor as output.
[0173] In some aspects, this disclosure provides methods for coupling somatic genomic information with epigenetic imprints (e.g., methylation profiles, fragmentomics, etc.), which provide additional genomic signals to definitively identify a tumor or a known CHIP gene by excluding the background of clonal hematopoietic variants of indeterminate potential (CHIP) using bioinformatics. In some embodiments, the methylation and fragmentation profiles of normal leukocytes exhibiting CHIP differ from their pathogenic tumor counterparts. In some embodiments, incorporating targeted hybridization sets of known methylation sites or other epigenetic sites in genes that may interfere with CHIP (e.g., DNMT3A, TP53, LRP1B, KRAS, etc.) into the NGS workflow provides orthogonal information for CHIP determination. Similarly, in some embodiments, the incorporation of a bioinformatics module analyzing the distribution of ctDNA fragments of genes known to exhibit high CHIP incidence is used as orthogonal information to generate a CHIP determination decision-maker. Combinations of known CHIP incidence genes or other genomic regions and epigenetic profiles (e.g., methylation profiles, ctDNA fragment distribution (e.g., fragmentomics), bisulfite sequencing, etc.) provide technical solutions to improve diagnostic efficacy.
[0174] To illustrate, Figure 10 This is a flowchart, schematically depicting exemplary method steps according to some embodiments of the present invention, using a computer to distinguish between tumor-derived nucleic acid variants and undetermined clonal hematopoietic (CHIP)-derived nucleic acid variants in a test sample obtained from a test subject. As shown, method 1100 includes identifying nucleic acid variants in a target genomic region set based on sequence information obtained from nucleic acids in the test sample to generate an identified set of test nucleic acid variants (step 1101). The method further includes identifying at least one epigenetic imprint of a given test nucleic acid variant corresponding to more than one identified test nucleic acid variant in the identified set of test nucleic acid variants based on epigenetic information obtained from nucleic acids in the test sample to generate a test nucleic acid variant-epigenetic imprint cluster (step 1102). In some embodiments, the epigenetic imprint, such as methylation features, may be determined based on the methods and systems disclosed in PCT application PCT / US2021 / 025201. The method further includes matching a given test nucleic acid variant-epigenetic imprinting cluster in the test nucleic acid variant-epigenetic imprinting cluster with a reference nucleic acid variant-epigenetic imprinting cluster corresponding to a tumor-derived nucleic acid variant or with a reference nucleic acid variant-epigenetic imprinting cluster corresponding to a CHIP-derived nucleic acid variant, thereby distinguishing tumor-derived nucleic acid variants and CHIP-derived nucleic acid variants in the test sample obtained from the test subject from each other (step 1103). In some embodiments, method 1100 further includes using at least one trained classifier to distinguish tumor-derived nucleic acid variants and CHIP-derived nucleic acid variants in the test nucleic acid variant-epigenetic imprinting cluster from each other to generate a distinguished set of tumor-derived nucleic acid variants and CHIP-derived nucleic acid variants present in the test sample. In some embodiments, the method further includes administering at least one therapy to the test subject based on one or more distinguished tumor-derived nucleic acid variants in the distinguished set of tumor-derived nucleic acid variants and CHIP-derived nucleic acid variants present in the test sample, thereby treating the test subject's cancer.
[0175] To illustrate, Figure 11This is a flowchart schematically depicting exemplary method steps for using a computer-generated trained classifier according to some embodiments of the present invention. As shown, method 1200 includes identifying nucleic acid variants in at least one target genomic region set based on sequence information obtained from nucleic acids in more than one reference sample to generate an identified reference nucleic acid variant set (step 1201). Method 1200 further includes identifying at least one epigenetic imprint of a given nucleic acid variant corresponding to more than one identified reference nucleic acid variant in the identified reference nucleic acid variant set based on epigenetic information obtained from nucleic acids in the reference sample to generate a reference nucleic acid variant-epigenetic imprint cluster (step 1202). Method 1200 further includes training a machine learning algorithm using at least a portion of the reference nucleic acid variant-epigenetic imprint cluster to create at least one trained classifier configured to classify one or more test nucleic acid variant-epigenetic imprint clusters as containing tumor-derived nucleic acid variants and / or undetermined clonal hematopoietic (CHIP)-derived nucleic acid variants (step 1203).
[0176] To further illustrate, Figure 12 This is a flowchart schematically depicting exemplary method steps for using a computer-generated trained classifier according to some embodiments of the present invention. As shown, method 1300 includes identifying nucleic acid variants in at least one target genomic region set based on sequence information obtained from nucleic acids in more than one reference sample to generate an identified reference nucleic acid variant set (step 1301). Method 1300 further includes training a machine learning algorithm using at least a portion of the identified reference nucleic acid variant set to create at least a first model configured to classify nucleic acid variants in the target genomic region set based on sequence information obtained from nucleic acids in a test sample to generate an identified test nucleic acid variant set (step 1302). Method 1300 further includes identifying at least one epigenetic imprint corresponding to a given nucleic acid variant in more than one reference identified nucleic acid variant in the identified reference nucleic acid variant set based on epigenetic information obtained from nucleic acids in the reference sample to generate a reference epigenetic imprint set (step 1303). Method 1300 further includes training a machine learning algorithm using at least a portion of a reference epigenetic imprint set to create at least a second model, said at least a second model being configured to distinguish between tumor-derived nucleic acid variants and undetermined clonal hematopoietic (CHIP)-derived nucleic acid variants in the test nucleic acid variant-epigenetic imprint set to produce an identified test nucleic acid variant set (step 1304).
[0177] In some embodiments, the tested nucleic acid variant-epigenetic imprint cluster comprises at least a first member and a second member, the first and second members comprising the same nucleic acid variant and different corresponding epigenetic imprints. In some of these embodiments, the different corresponding epigenetic imprints comprise different epigenetic states or conditions expressed by one or more epigenetic loci within a given target genomic region. In some of these embodiments, the different corresponding epigenetic imprints comprise different cell-free nucleic acid (cfNA) fragment lengths, locations, and / or endpoint density distributions. In some embodiments, the tested nucleic acid variant-epigenetic imprint cluster comprises at least a first member and a second member, the first and second members comprising different nucleic acid variants and the same corresponding epigenetic imprint.
[0178] In some implementations, the matching step includes using at least one trained classifier to distinguish tumor-derived nucleic acid variants and CHIP-derived nucleic acid variants from test samples obtained from test subjects. In some implementations, the identified set of nucleic acid variants includes somatic nucleic acid variants. In some implementations, a given target genomic region includes two or more nucleic acid variant loci. In some implementations, the test nucleic acid variant-epigenetic imprint cluster includes at least one member, said member comprising one or more nucleic acid variants and one or more corresponding epigenetic imprints from different genomic regions within the target genomic region set. In some implementations, the test nucleic acid variant-epigenetic imprint cluster includes at least one member, said member comprising one or more nucleic acid variants and one or more corresponding epigenetic imprints located in the same genomic region within the target genomic region set. In some implementations, more than one targeted genomic region contains one or more genes selected from the group consisting of: DNMT3A, TP53, LRP1B, KRAS, MARCH11, TAC1, TCF21, SHOX2, p16, Casp8, CDH13, MGMT, MLH1, MSH2, TSLC1, APC, DKK1, DKK3, LKB1, WIF1, RUNX3, GATA4, GATA5, PAX5, E-cadherin, H-cadherin, VIM, SEPT9, CYCD2, TFPI2, GATA4, RARB2, p16INK4a, APC, NDRG4, HLTF, HPP1, hMLH1, RASSF1A, IGFBP3, ITGA4, PIK3CA, ERBB2. (HER2), BRCA1 / 2, NTRK1 / 2 / 3, MSI-High, ESR1, ATM, HRR, FGFR2 / 3, IDH1, KRAS, NRAS, BRAF, KIT, PDGFRA, EGFR, ALK, ROS1, MET, TMB, or RET. In some embodiments, the nucleic acids in the sample include cell-free nucleic acid (cfNA) fragments and / or nucleic acid molecules obtained from one or more tissues or cells in the sample. In some embodiments, the epigenetic imprinting includes cfNA fragment length, location, and / or endpoint density distribution.
[0179] In some embodiments, epigenetic imprinting includes epigenetic states expressed by one or more epigenetic loci in a given target genomic region. In some embodiments, epigenetic states include the presence or absence of methylation, hydroxymethylation, acetylation, ubiquitination, phosphorylation, ubiquitin-like methylation, ribosylation, citrullination, and / or histone post-translational modifications or other histone variations. In some embodiments, the method further includes ignoring CHIP-derived nucleic acid variations that are distinguishable from each other in further analysis. In some embodiments, the method further includes generating at least one report listing tumor-derived nucleic acid variations and CHIP-derived nucleic acid variations that are distinguishable from each other in the test sample.
[0180] In some embodiments, the method further includes identifying at least one cancer type associated with a distinguished tumor-derived nucleic acid variant. In some embodiments, the method further includes administering at least one therapy to a test subject to treat the identified cancer type. In some embodiments, the method further includes administering at least one therapy to a test subject based on one or more distinguished tumor-derived nucleic acid variants. In some embodiments, one or more cells contain nucleic acids from the test sample.
[0181] In some implementations, the method further includes identifying nucleic acid variants in a target genomic region set using a computer based on sequence information obtained from nucleic acids in test samples obtained from test subjects to generate an identified set of test nucleic acid variants; identifying at least one epigenetic imprint corresponding to a given test nucleic acid variant in the identified set of test nucleic acid variants using a computer based on epigenetic information obtained from nucleic acids in the test samples to generate a test nucleic acid variant-epigenetic imprint cluster; and using a trained classifier to distinguish tumor-derived nucleic acid variants and CHIP-derived nucleic acid variants in the test nucleic acid variant-epigenetic imprint cluster from the test samples obtained from test subjects from each other. In some implementations, the second model is a further trained version of the first model. In some implementations, the reference nucleic acid variant-epigenetic imprint cluster includes incidence data of epigenetic imprints corresponding to a given nucleic acid variant in the identified reference nucleic acid variant set.
[0182] In some embodiments, identifying at least one epigenetic imprint corresponding to a given nucleic acid variant includes: determining an epigenetic rate corresponding to the given nucleic acid variant, wherein at least a first epigenetic rate is generated based on a first sample obtained from a given subject at a first time point, and at least a second epigenetic rate is generated based on a second sample obtained from the given subject at a second time point different from the first time point; adjusting at least one epigenetic rate threshold based on at least the first epigenetic rate to produce an adjusted epigenetic rate threshold; and using the adjusted epigenetic rate threshold to identify the epigenetic imprint. In some embodiments, the first and second samples include a test sample. In some embodiments, the first and second samples include a reference sample. In some embodiments, the first sample includes a tumor tissue sample. In some embodiments, the second sample includes a body fluid sample. Some embodiments include using epigenetic rates to identify a tumor fraction in the sample. Some implementations optionally include determining more than one epigenetic rate for more than one genomic region in a first sample; determining the probability of one or more tumor fractions in more than one genomic region in the second sample based on a predetermined set of epigenetic rates for more than one genomic region in a second sample, a set of epigenetic attributes of a set of cell-free polynucleotides mapped to more than one genomic region in the second sample, and the epigenetic rate of more than one genomic region in the first sample; combining more than one probability for one or more of more than one genomic region to determine the overall posterior probability of cancer presence in the subject; and comparing the overall posterior probability of cancer presence in the subject with a predetermined threshold. Some of these implementations also include: (a) classifying the subject as circulating tumor DNA (ctDNA) positive if the overall posterior probability of cancer presence in the subject is greater than or equal to the predetermined threshold, or (b) classifying the subject as ctDNA negative if the overall posterior probability of cancer presence in the subject is less than the predetermined threshold. In some embodiments, methods and systems for analyzing epigenetic status can be found in International Patent Application No. PCT / US2020 / 035605 entitled “METHODS AND SYSTEMSFOR IMPROVING PATIENT MONITORING AFTER SURGERY”, filed June 1, 2020, which is incorporated herein by reference.
[0183] exist Example 1: Annotation of tumor and non-tumor SNV / insertions / deletions to allow filtering of non-tumor variantsThe embodiment shown discloses a method 1400 for generating a predictive model. Method 1400 may be performed wholly or partially by a single computing device, more than one computing device, etc. Method 1400 may include determining sequence data at 1401. Method 1400 may include determining at 1402 at at least one of: epigenetic data or fragment omics data. Method 1400 may include determining more than one feature of the predictive model at 1403. Method 1400 may include training and / or testing the predictive model based on more than one feature at 1404. Method 1400 may include outputting the predictive model at 1405.
[0184] More than one genomic region may contain at least one of the following: DNMT3A, TP53, LRP1B, KRAS, MARCH11, TAC1, TCF21, SHOX2, p16, Casp8, CDH13, MGMT, MLH1, MSH2, TSLC1, APC, DKK1, DKK3, LKB1, WIF1, RUNX3, GATA4, GATA5, PAX5, E-cadherin, H-cadherin, VIM, SEPT9, CYCD2, TFPI2, GATA4, RARB2, p16INK4a, APC, NDRG4, HLTF, HPP1, hMLH1, RASSF1A, IGFBP3, ITGA4, PIK3CA, ERBB2 (HER2), BRCA1 / 2, NTRK1 / 2 / 3, MSI-High, ESR1, ATM, HRR, FGFR2 / 3, IDH1, KRAS, NRAS, BRAF, KIT, PDGFRA, EGFR, ALK, ROS1, MET, TMB, or RET. Identifying sequence data may include obtaining more than one sample from more than one subject, wherein each sample contains more than one cell-free nucleic acid. More than one genomic region may include at least one of the following: a genomic region known to be associated with cancer type, a genomic region associated with a known methylation status, a genomic region known to be associated with hypomethylation, or a genomic region known to be associated with therapy response.
[0185] Epigenetic data may include at least one of the following: information about DNA methylation, histone status or modification, inflammation-mediated cytosine damage products, protein binding, or other molecular states reflected in the analyzed nucleic acid fragment that are not determined solely from the nucleotide base sequence (e.g., the methylation status of a given base or set of bases). Identifying epigenetic data associated with more than one sequence fragment includes determining the methylation status of more than one sequence fragment.
[0186] Determining the methylation state of more than one sequence fragment may include determining at least one of the following: a methylation state vector or a methylation CpG density. Determining the methylation state vector may include aligning more than one sequence read with a reference sequence; determining the methylation status of one or more CpG sites and the location of one or more CpG sites in the sequence reads within the more than one sequence read based on the alignment; and vectorizing the methylation status of one or more CpG sites and the location of one or more CpG sites to generate a methylation state vector for the sequence reads within the more than one sequence read. Determining the methylated CpG density may include aligning more than one sequence read with a reference sequence; determining the methylation status of one or more CpG sites in the sequence reads based on the alignment; determining whether the sequence read is methylated or unmethylated based on the methylation status of one or more CpG sites in the sequence read; for more than one sequence read, determining the count of methylated sequence reads and the count of unmethylated sequence reads; and determining the methylated CpG density based on the count of methylated sequence reads and the count of unmethylated sequence reads.
[0187] Fragmentomics data may include at least one of the following: information about fragment size, nucleotide motifs at fragment ends, single-stranded serrated ends, genomic location of the center point of fragment length, genomic location of fragment endpoints, and / or any values indicating fragment endpoints. Fragmentomics data for determining association with more than one sequence fragment may include at least one of the following: determining the size of the sequence fragments in more than one fragment or determining the amount of more than one sequence fragment having a specific size. The specific size may be a range. The range may be at least one of the following: 50-80, 50-100, 50-150, 100-200, 150-200, 150-230, 200-300, or 300-400 bases.
[0188] Determining fragmentomics data associated with more than one sequence fragment may include determining terminal motifs of more than one sequence fragment, wherein the terminal motifs relate to the terminal sequences of the sequence fragments. Determining terminal motifs of more than one sequence fragment may include aligning more than one sequence read sequenced from more than one sequence fragment with a reference sequence, and determining the terminal motif at each end of the more than one sequence fragment based on the alignment. The terminal sequence may contain a number of bases, wherein the number of bases is between 1 and 6 bases. The terminal sequence may contain a number of bases extending beyond the sequence fragment, wherein the number of bases is between 1 and 6 bases. Method 1400 may also include determining the frequency of occurrence of the terminal motif within more than one sequence fragment. Method 1400 may also include determining the terminal bases of the terminal motif and determining the frequency of occurrence of the terminal bases of the terminal motif. Determining fragmentomics data associated with more than one sequence fragment may include determining serrated ends of the sequence fragments in more than one sequence fragment. Determining serrated ends of the sequence fragments in more than one sequence fragment may include determining an overhang index. The sequence fragment can be double-stranded, i.e., a first strand and a second strand having a first part, and determining the protrusion index can include determining the methylation state of the first strand or the second strand, the protrusion index being proportional to the length of the first strand protruding from the second strand, and determining the protrusion index based on the methylation state, wherein the protrusion index provides a measure of one strand protruding from the other strand.
[0189] Identifying fragmentomics data associated with more than one sequence fragment can include determining the genomic location of the fragment endpoints. Determining the genomic location of the fragment endpoints can include determining the window protection score (WPS). Determining the WPS can include determining the number of sequence fragments that span the window, and adjusting the number of sequence fragments that span the window based on any sequence fragments that begin within the window.
[0190] Method 1400 may also include determining the source of the sequence fragment and assigning the source of the sequence fragment to the sequence data, epigenetic data, and fragmentomics data associated with the sequence fragment. The source may be tumor-derived or non-tumor-derived, the source may be a tissue type, or the source may be a cancer type.
[0191] Determining more than one feature of a predictive model based on at least a portion of sequence data and at least a portion of at least one of epigenetic data or fragmentomics data may include determining at least one of the following: methylation state vector, methylation density, fragment size, fragment size distribution, terminal motif, terminal motif frequency, presence of serrated ends, protrusion index, genomic location of the midpoint of fragment length, genomic location of fragment endpoints, genomic location of fragment endpoints, genomic location indicating the endpoint of a fragment, or window protection score; and determining at least one of the following, individually or in combination, that has a predictive value associated with the origin of the sequence fragment: methylation state vector, methylation density, fragment size, fragment size distribution, terminal motif, terminal motif frequency, presence of serrated ends, protrusion index, genomic location of the midpoint of fragment length, genomic location of fragment endpoints, any value indicating the fragment endpoint, or window protection score.
[0192] Training a predictive model based on at least one of the first portion of sequence data and epigenetic or fragment omics data, according to more than one feature, may include training the predictive model using a machine learning method. The machine learning method may include at least one of the following: discriminant analysis, decision tree, nearest neighbor (NN) algorithm, Bayesian network, clustering algorithm, neural network, support vector machine (SVM), logistic regression algorithm, linear regression algorithm, Markov model, or principal component analysis (PCA). Testing the predictive model based on at least one of the second portion of sequence data and epigenetic or fragment omics data may include retraining the predictive model.
[0193] Method 1400 may further include determining test sequence data for a subject, said test sequence data comprising more than one sequence fragment associated with more than one genomic region, wherein the more than one sequence fragment is sequenced from a sample from the subject; determining at least one of test epigenetic data or test fragment omics data associated with the more than one sequence fragment; providing the subject's test sequence data, test epigenetic data, and test fragment omics data to a predictive model; and determining the origin of at least one sequence fragment in the sequence data based on the subject's test sequence data, test epigenetic data, and test fragment omics data. The origin may be of either tumor origin or non-tumor origin.
[0194] Method 1400 may also include administering one or more therapies to the subject based on the source being a tumor. Therapies may include administering chemotherapy, administering radiation therapy, or performing surgery to remove all or part of the tumor. Therapies may include the administration of at least one of the following: ALECENSA®, ALUNBRIG®, BRAFTOVI®, ERBITUX®, GAVRETO™, GILOTRIF®, HERCEPTIN®, IRESSA®, KADCYLA®, KEYTRUDA®, LORBRENA®, LUMAKRAS™, LYNPARZA®, MEKINIST®, OPDIVO®, PERJETA®, PIQRAY®, RETEVMO™, ROZLYTREK™, RUBRACA®, TABRECTA™, TAFINLAR®, TAGRISSO®, TALZENNA®, TARCEVA®, TEPMETKO™, TYKERB®, VITRAKVI®, VIZIMPRO®, XALKORI®, YBREVANT™, YERVOY®, or ZYKADIA®.
[0195] In the implementation plan, such as Example 2: Study design The diagram illustrates a method 1500 for determining the origin of a sample. Method 1500 can be performed wholly or partially by a single computing device, more than one computing device, etc. Method 1500 may include determining the sequence data of the subject's sample at 1501. Method 1500 may include determining at 1502 at at least one of: epigenetic data or fragment omics data. Method 1500 may include providing the sequence data and at least one of the epigenetic data or fragment omics data to a predictive model. Method 1500 may include determining, based on the predictive model, whether the sample is of tumor origin or non-tumor origin. Method 1500 may also include generating the predictive model. Generating a predictive model may include identifying sequence data associated with more than one sequence fragment, wherein the sequence data includes more than one sequence read, wherein the more than one sequence read is sequenced from more than one sequence fragment from more than one sample, wherein each of the more than one sample is labeled as tumor-derived or non-tumor-derived; identifying at least one of epigenetic data or fragment omics data associated with more than one sequence fragment; determining more than one feature of the predictive model based on at least a portion of the sequence data and at least a portion of the epigenetic data or fragment omics data; training the predictive model based on the more than one feature; testing the predictive model based on a second portion of the sequence data and at least one of the epigenetic data or fragment omics data; and predicting the model based on the test output.
[0196] More than one genomic region may contain at least one of the following: DNMT3A, TP53, LRP1B, KRAS, MARCH11, TAC1, TCF21, SHOX2, p16, Casp8, CDH13, MGMT, MLH1, MSH2, TSLC1, APC, DKK1, DKK3, LKB1, WIF1, RUNX3, GATA4, GATA5, PAX5, E-cadherin, H-cadherin, VIM, SEPT9, CYCD2, TFPI2, GATA4, RARB2, p16INK4a, APC, NDRG4, HLTF, HPP1, hMLH1, RASSF1A, IGFBP3, ITGA4, PIK3CA, ERBB2 (HER2), BRCA1 / 2, NTRK1 / 2 / 3, MSI-High, ESR1, ATM, HRR, FGFR2 / 3, IDH1, KRAS, NRAS, BRAF, KIT, PDGFRA, EGFR, ALK, ROS1, MET, TMB or RET.
[0197] Determining sequence data may include obtaining more than one sample from more than one subject, wherein the more than one sample contains more than one cell-free nucleic acid. The more than one genomic region may include at least one of the following: a genomic region known to be associated with a cancer type, a genomic region associated with a known methylation status, a genomic region known to be associated with hypomethylation, or a genomic region known to be associated with a therapy response.
[0198] Epigenetic data may include at least one of the following: information about DNA methylation, histone status or modification, inflammation-mediated cytosine damage products, protein binding, or other molecular states reflected in the analyzed nucleic acid fragment that are not determined solely from the nucleotide base sequence (e.g., methylation status of a given base or set of bases). Determining epigenetic data associated with more than one sequence fragment may include determining the methylation status of more than one sequence fragment. Determining the methylation status of more than one sequence fragment may include determining at least one of the following: a methylation status vector or a methylation CpG density. Determining a methylation status vector may include aligning more than one sequence read to a reference sequence; determining, based on the alignment, the methylation status of one or more CpG sites and the location of one or more CpG sites in the sequence reads of more than one sequence read; and vectorizing the methylation status of one or more CpG sites and the location of one or more CpG sites to generate a methylation status vector for the sequence reads of more than one sequence read. Determining the methylated CpG density may include aligning more than one sequence read with a reference sequence; determining the methylation status of one or more CpG sites in the sequence reads based on the alignment; determining whether the sequence read is methylated or unmethylated based on the methylation status of one or more CpG sites in the sequence read; for more than one sequence read, determining the count of methylated sequence reads and the count of unmethylated sequence reads; and determining the methylated CpG density based on the count of methylated sequence reads and the count of unmethylated sequence reads.
[0199] Fragmentomics data may include at least one of the following: information about fragment size, nucleotide motifs at the fragment ends, single-stranded serrated ends, genomic location of the center point of the fragment length, genomic location of the fragment endpoints, and / or any values indicating the fragment endpoints. Determining fragmentomics data associated with more than one sequence fragment may include at least one of the following: determining the size of the sequence fragments in more than one fragment or determining the amount of more than one sequence fragment having a specific size. The specific size may be a range. The range may be at least one of the following: 50-80, 50-100, 50-150, 100-200, 150-200, 150-230, 200-300, or 300-400 bases. Determining fragmentomics data associated with more than one sequence fragment may include determining the terminal motifs of more than one sequence fragment, wherein the terminal motifs relate to the terminal sequences of the sequence fragments. Determining the terminal motif of more than one sequence fragment may include aligning more than one sequence read from the sequence fragment with a reference sequence, and determining the terminal motif at each end of the sequence fragment based on the alignment. The terminal sequence may contain a certain number of bases. The number of bases may be between 1 and 6 bases. The terminal sequence contains a certain number of bases extending beyond the sequence fragment, wherein the number of bases is between 1 and 6 bases. Method 1500 may further include determining the frequency of occurrence of the terminal motif within the more than one sequence fragment. Method 1500 may further include determining the terminal bases of the terminal motif and determining the frequency of occurrence of the terminal bases of the terminal motif.
[0200] Determining fragmentomics data associated with more than one sequence fragment includes identifying the serrated ends of the sequence fragments within that sequence fragment. Identifying the serrated ends of the sequence fragments within the multiple sequence fragments includes determining a protrusion index. The sequence fragment may be double-stranded, i.e., a first strand and a second strand having a first portion, and determining the protrusion index may include determining the methylation state of the first or second strand, the protrusion index being proportional to the length of the first strand protruding from the second strand, and the protrusion index being determined based on the methylation state, wherein the protrusion index provides a measure of how much one strand protrudes from the other.
[0201] Identifying fragmentomics data associated with more than one sequence fragment can include determining the genomic location of the fragment endpoints. Determining the genomic location of the fragment endpoints can include determining the window protection score (WPS). Determining the WPS can include determining the number of sequence fragments that span the window, and adjusting the number of sequence fragments that span the window based on any sequence fragments that begin within the window.
[0202] Method 1500 may also include determining the source of the sequence fragment and assigning the source of the sequence fragment to the sequence data, epigenetic data, and fragmentomics data associated with the sequence fragment. The source may be tumor-derived or non-tumor-derived, the source may be a tissue type, or the source may be a cancer type.
[0203] The prediction model is determined based on at least a portion of sequence data and at least a portion of at least one of epigenetic data or fragmentomics data, including determining at least one of the following: methylation state vector, methylation density, fragment size, fragment size distribution, terminal motif, terminal motif frequency, presence of serrated ends, protrusion index, genomic location of the center point of fragment length, genomic location of fragment endpoints, genomic location of fragment endpoints, genomic location of fragment endpoints, or window protection score; and determining at least one of the following, individually or in combination, that has a predictive value associated with the source of the sequence fragment: methylation state vector, methylation density, fragment size, fragment size distribution, terminal motif, terminal motif frequency, presence of serrated ends, protrusion index, genomic location of the center point of fragment length, genomic location of fragment endpoints, genomic location of fragment endpoints, any value indicating fragment endpoints, or window protection score.
[0204] Training a predictive model based on at least one of the first portion of sequence data and epigenetic or fragment omics data, according to more than one feature, may include training the predictive model using a machine learning method. The machine learning method may include at least one of the following: discriminant analysis, decision tree, nearest neighbor (NN) algorithm, Bayesian network, clustering algorithm, neural network, support vector machine (SVM), logistic regression algorithm, linear regression algorithm, Markov model, or principal component analysis (PCA). Testing the predictive model based on at least one of the second portion of sequence data and epigenetic or fragment omics data may include retraining the predictive model.
[0205] The method may also include administering one or more therapies to the subject based on whether the sample is tumor-derived. Therapies may include administering chemotherapy, administering radiation therapy, or performing surgery to remove all or part of the tumor. Therapies may include the administration of at least one of the following: ALECENSA®, ALUNBRIG®, BRAFTOVI®, ERBITUX®, GAVRETO™, GILOTRIF®, HERCEPTIN®, IRESSA®, KADCYLA®, KEYTRUDA®, LORBRENA®, LUMAKRAS™, LYNPARZA®, MEKINIST®, OPDIVO®, PERJETA®, PIQRAY®, RETEVMO™, ROZLYTREK™, RUBRACA®, TABRECTA™, TAFINLAR®, TAGRISSO®, TALZENNA®, TARCEVA®, TEPMETKO™, TYKERB®, VITRAKVI®, VIZIMPRO®, XALKORI®, YBREVANT™, YERVOY®, or ZYKADIA®.
[0206] More information can be found in PCT applications PCT / US2021 / 015837, PCT / US2021 / 047619 and PCT / US2021 / 021994, Fairchild et al., Science Trans. Med. (2023), each of which is incorporated herein by reference in its entirety. Example
[0207] Example 3: CHIP genes As described, the presence of clonal hematopoietic (CH) variants and biological noise, due to aging and therapy, can potentially obscure the interpretation of biomarkers.
[0208] Currently, comprehensive methods for filtering out non-tumor variants require genotyping of white blood cell (WBC) fractions from paired plasma samples, which is an expensive and complex workflow. Of interest are plasma-only bioinformatics solutions to identify precise biomarkers in cell-free DNA (cfDNA) for assessing the desired non-tumor variants.
[0209] The sensitivity of annotations using ESR and tannin sequencing can influence the interpretation of the clinical significance of variants, particularly in genes with mixed tumor / CH origins, such as TP53 and ATM, and influencing the specificity of products requiring CH filtering, such as genomic MR and TMB scores.
[0210] Example 4: CHIP variant list To address these challenges, the inventors identified variant determinations from >250,000 plasma samples, including healthy donors, early- and late-stage cancer patients, sequenced on genomic and / or epigenomic liquid biopsy sets, and further incorporated public tissue datasets. The model was trained on paired plasma and WBC datasets and optimized with 10-fold cross-validation to produce non-tumor and tumor variant classifiers. Validation was provided using a paired plasma WBC late-stage cancer sample cohort genotyped in epigenomic assays. Another cohort of healthy donor samples genotyped in genomic assays was also evaluated.
[0211] Example 5: Additional features In some cases, somatic variants in these genes associated with hematologic malignancies will be annotated as CHIP. These include ASXL1, DNMT3A, GNAS, JAK2, PPM1D, SF3B1, TET2, BCOR, CBL, CEBPA, CREBBP, ETV6, EZH2, FLT3, IDH1, IDH2, MPL, MYD88, NPM1, RUNX1, SH2B3, STAG2, STAT3, U2AF1, ZRSR2, PLCG2, CALR, CSF3R, DDX41, BRCC3, EZH1, KIR2DL1, KIR2DL3, and KMT2C.
[0212] Example 6: Findings In other cases, various techniques can be employed. One example includes a logistic regression (LR) classifier, trained on a curated, empirically determined dataset and an external public dataset, and applied to a larger, empirically determined dataset for the known biological characteristics of CHIP-identified CHIP. Another example involves incorporating variants already identified in heme malignancies in the COSMIC imprint of >2 samples, or variants similarly identified in heme malignancies in TCGA of >2 samples. Yet another example includes an LR classifier that uses the proportion of heme samples as a single feature for CHIP identification.
[0213] Example 7: Multi-feature model In some cases, to reduce tumor variants labeled CH, annotation labels can be adjusted based on the erythrocyte sedimentation rate (ESR) brown layer by filtering out variants that are not detected due to insufficient coverage. Another example would involve expanding the training dataset (additional internal / drug samples, publicly available data) to include assessments of the effects of tumor mutational burden (TMB) and genomic molecular response (MR). Finally, methylation data, including sample and / or variant level information, can be incorporated into model training. As shown, when applied to analysis, these techniques involved in aggregation, normalization, and encoding for feature determination result in the construction of features representing the most important ones. Ultimately, such factors can be weighted when establishing a scoring metric for “CH burden.”
[0214] Figure 13 Based on the implementation described herein, an exemplary bioinformatics model demonstrating high sensitivity and specificity to WBC in distinguishing between tumor and non-tumor variants using only cfDNA is shown. As illustrated herein, at low VAF (<0.6%), the bioinformatics model exhibits improved sensitivity for identifying non-tumor variants compared to WBC sequencing.
[0215] In paired plasma and WBC advanced cancer cohorts, most non-tumor variants were found among known clonal hematopoietic genes and variants of unknown significance. Apart from ATM and CHEK2, no clinically actionable variants were identified or annotated as non-tumor.
[0216] Figure 14 According to the implementation scheme described herein, multi-feature models can be generated using data feature sources including: variant level information (such as VAF, fragment omics, methylation), and / or clinical test sample databases (such as aggregate variant summaries from clinical patients, clonalness, longitudinal variant variability, cancer type variability, etc.), and / or public datasets (such as COSMIC (variation frequency in cancer tissues), GnomAD (population-level allele frequency), COSMIC mutation imprinting (SBS imprinting)).
[0217] like Example 8: Multi-feature model performance characteristics, buffy-coat-plasma matched samples The diagram illustrates the feature selection and model building process, in which a correlation check is applied to prevent the correlation from exceeding a threshold (e.g., 70%, 80%, etc.). Figure 15B An example of a baseline fact state determined by a matched ESR (erythrocyte sedimentation rate) brown-yellow layer for clinically relevant genes from NHC is shown. An example of a multi-feature model is shown in Figure 15. Here, the exemplary multi-feature model of the CHIP classifier uses 14 features and shows best performance on both validation and test data. Exemplary sample-level features are depicted.
[0218] Figure 16 Compared to exemplary conventional models without multi-feature generation and simplified feature models obtained by CH variant markers based on the presence of the same SNV / insertion / deletion in paired cfDNA and erythrocyte sedimentation rate (ESR) brown-yellow layer, the described exemplary model exhibits superior receiver-operator characteristics (ROC), such as... Figure 17 As shown. Figure 18 As shown, the CHIP classifier features can effectively capture almost all differences between CHIP and somatic variation. Figure 19 The results shown demonstrate superior performance of the multi-feature model, the exemplary conventional model without multi-feature generation, and the simplified feature model in ESR-plasma-matched samples for clinically relevant genes. Figure 20 As further shown, high FP genes were not detected in the current classifier, and as... Example 9: Multi-feature model performance characteristics, tissue-plasma matched samples The validation data using various VAF bins shown in the figure indicate that, for all recorded cancer types, no high FP bins were detected in the exemplary multi-feature classifier, and as... Figure 21 The results showed that no high-FP cancer types were detected.
[0219] Figure 22 The described multi-feature model was then applied to tissue-plasma-matched samples, and its superior performance was demonstrated. Figure 23 and Figure 24 Regarding the performance of clinically relevant genes: specificity assessment of tissue-plasma matched samples, including low VAF bins, showed no high FP bins detected below 0.5% in the current classifier. Figure 25 In this study, an exemplary multi-feature model was applied to clinically relevant genes from an exemplary group of over 700 genes, depicting instances of baseline factual states determined by matched erythrocyte sedimentation rate (ESR) brown-yellow layers. It was observed that the multi-feature model exhibited superior performance regarding group-wide variation compared to exemplary conventional models without multi-feature generation and simplified feature models. Example 10: Multi-feature model performance characteristics, actionable variants and Figure 27 As shown in the image.
[0220] Figure 28 The described multi-feature model was then applied to actionable variation performance: ESR-plasma-matched samples. The current classifier achieved 100% accuracy and 0% FRP. Figure 29 As shown, it is worth noting that the exemplary multi-feature classifier achieves 0% FRP. Further performance is shown when applied to operational variation classification indicating a very low FRP. Figure 30The exemplary group comprises approximately 80 genes. The operable variant performance in plasma ctDNA was assessed, such as... Figure 31 As shown, most CHIP decisions fall within the known KRAS CHIP variant K117N. A summary of the current model performance is shown in... In the tests, compared to the finite feature model, data sensitivity increased from 47% to 69%, and the FPR decreased from 2.9% to 1.4%. The deployment of the CHIP classifier algorithm is as follows: As shown in the image.
[0221] All patent applications, websites, other publications, registration numbers, etc., cited above or below are incorporated by reference in their entirety for all purposes, to the extent that each individual item is specifically and individually indicated by reference. If different versions of a sequence are associated with a registration number at different times, it refers to the version associated with that registration number on the effective filing date of this application. The effective filing date refers to the earlier of the actual filing date using that registration number or the filing date of the priority application (if applicable). Similarly, if different versions of publications, websites, etc., are published at different times, it refers to the most recently published version on the effective filing date of the application, unless otherwise indicated. Any feature, step, element, embodiment, or aspect of this disclosure may be used in combination with any other feature, step, element, embodiment, or aspect, unless otherwise specifically indicated. Although this disclosure has been described in considerable detail by way of illustration and example for purposes of clarity and understanding, it will be apparent that certain changes and modifications may be made within the scope of the appended claims.
Claims
1. A method comprising: Sequence data that identifies more than one sequence fragment associated with more than one genomic region, wherein the sequence data includes more than one sequence read, wherein the more than one sequence read is obtained from sequencing the more than one sequence fragment from more than one sample, wherein each of the more than one sample is labeled as tumor-derived or non-tumor-derived. Identify epigenetic data associated with the more than one sequence fragment; Based on the sequence data and epigenetic data, more than one feature is determined for the prediction model; Based on the sequence data and epigenetic data, a prediction model is generated according to more than one feature.
2. The method of claim 1, wherein determining the sequence data comprises obtaining more than one sample from more than one subject, wherein the more than one sample comprises more than one cell-free nucleic acid.
3. The method according to claims 1-2, wherein the more than one feature comprises at least one of the following: fragment length, variant VAF, variant CHIP to somatic cell ratio, APOBEC-related cancer markers, variant measurement variability, maximum variant clonalness, age-related markers, variant clonalness variance, population allele frequency, ratio of methylated to unmethylated fragments, genomic regions associated with cancer type, genomic regions associated with methylation status, genomic regions associated with hypomethylation, or genomic regions associated with therapy response.
4. The method according to claim 3, wherein the segment length is the average length and / or length variance.
5. The method of claim 4, wherein the fragment length is associated with single-nucleus bodies and / or dual-nucleus bodies.
6. The method of claim 4, wherein the age-related biomarker is SBS88.
7. The method of claim 4, wherein the APOBEC-associated cancer biomarker is SBS2.
8. The method of claim 4, wherein the variation measurement is one or more of variability, maximum clonalness of variation, and variance of clonalness of variation.
9. The method according to claims 1-9, wherein the epigenetic data includes at least one of the following: information on DNA methylation, histone status or modification, inflammation-mediated cytosine damage products, or protein binding.
10. The method of claims 1-10, wherein determining the epigenetic data associated with the more than one sequence fragment includes determining the methylation status of the more than one sequence fragment.
11. The method of claim 10, wherein determining the methylation state of the more than one sequence fragment comprises determining at least one of the following: a methylation state vector or a methylation CpG density.
12. The method of claim 11, wherein determining the methylation state vector comprises: Align the more than one sequence read with a reference sequence; Based on the alignment, the methylation status of one or more CpG sites in the sequence reads of the more than one sequence read and the location of the one or more CpG sites are determined; as well as The methylation status of the one or more CpG sites and the position of the one or more CpG sites are vectorized to generate a methylation status vector for the sequence reads in the more than one sequence read.
13. The method of claim 11, wherein determining the methylated CpG density comprises: Align the more than one sequence read with a reference sequence; Based on the alignment, the methylation status of one or more CpG sites in the sequence reads of the more than one sequence reads is determined; The methylation status of one or more CpG sites in the sequence read is used to determine whether the sequence read is methylated or unmethylated. For more than one sequence read, determine the count of methylated sequence reads and the count of unmethylated sequence reads; and The methylated CpG density is determined based on the counts of the methylated sequence reads and the counts of the unmethylated sequence reads.
14. The method according to claims 1-13, wherein training the prediction model includes applying a machine learning algorithm.
15. The method according to claim 14, wherein, The machine learning method includes at least one of the following: discriminant analysis, decision tree, nearest neighbor (NN) algorithm, Bayesian network, clustering algorithm, neural network, support vector machine (SVM), logistic regression algorithm, linear regression algorithm, Markov model or principal component analysis (PCA).
16. The method according to claims 1-15, comprising retraining the prediction model.
17. The method according to claims 1-16, further comprising: For the subject, test sequence data is determined, which includes more than one sequence read obtained from sequencing a sample from the subject; Generate test epigenetic data and / or test fragment omics data associated with the more than one sequence fragment; The predictive model is provided with the subject's test sequence data, test epigenetic data, and test fragment omics data; as well as Based on the test sequence data, the test epigenetic data, and the test fragment omics data of the subject, the origin of at least one sequence fragment in the sequence data is determined.
18. The method according to claims 1-17, further comprising determining the source of at least one sequence segment in the sequence data.
19. The method according to claims 1-18, wherein the source is either tumor-derived or non-tumor-derived.
20. The method of claims 1-19, further comprising administering one or more therapies to the subject based on the fact that the source is tumor-derived.
21. The method according to claims 1-20, wherein the therapy includes administering chemotherapy, administering radiotherapy, or performing surgery to remove all or part of the tumor.
22. A method comprising: Obtain sequence data, which includes more than one sequence read associated with more than one genomic region, wherein the sequence data is generated from a sample from the subject; Identify epigenetic data associated with the more than one sequence read; Provide at least a portion of the sequence data and at least a portion of the epigenetic data to a trained prediction model; and The prediction model determines whether the sample is of tumor origin or not.
23. The method of claim 22, further comprising generating the prediction model.
24. The method of claim 23, wherein generating the prediction model comprises: Sequence data for more than one sample are determined, wherein each of the more than one sample is labeled as tumor-derived or non-tumor-derived; Identify epigenetic data associated with the more than one sequence read; More than one feature of the prediction model is determined based on a portion of the sequence data and a portion of the epigenetic data; The prediction model is trained based on a portion of the sequence data and a portion of the epigenetic data, according to more than one feature. as well as The prediction model is output based on the training.
25. The method of claims 22-24, comprising testing based on a portion of the sequence data and a portion of the epigenetic data.
26. The method of claims 22-25, wherein determining the sequence data comprises obtaining more than one sample from more than one subject, wherein the more than one sample comprises more than one cell-free nucleic acid.
27. The method of claims 22-26, wherein the more than one feature comprises at least one of the following: fragment length, variant VAF, variant CHIP to somatic cell ratio, APOBEC-associated cancer markers, variant measurement variability, maximum variant clonalness, age-related markers, variant clonalness variance, population allele frequency, ratio of methylated to unmethylated fragments, genomic regions associated with cancer type, genomic regions associated with methylation status, genomic regions associated with hypomethylation, or genomic regions associated with therapy response.
28. The method of claim 27, wherein the segment length is the average length and / or length variance.
29. The method of claim 27, wherein the fragment length is associated with single-nucleus bodies and / or dual-nucleus bodies.
30. The method of claim 27, wherein the age-related biomarker is SBS88.
31. The method of claim 27, wherein the APOBEC-associated cancer biomarker is SBS2.
32. The method of claim 27, wherein the variation measurement is one or more of variability, maximum clonalness of variation, and variance of clonalness of variation.
33. The method according to any one of claims 22-32, wherein the more than one genomic region comprises at least one of the following: a genomic region known to be associated with a cancer type, a genomic region associated with a known methylation status, a genomic region known to be associated with hypomethylation, or a genomic region known to be associated with a therapy response.
34. The method according to any one of claims 22-33, wherein the epigenetic data includes at least one of the following: information on DNA methylation, histone status or modification, inflammation-mediated cytosine damage products, or protein binding.
35. The method according to any one of claims 22-34, wherein determining the epigenetic data includes determining the methylation status associated with the more than one genomic region.
36. The method according to claims 22-35, wherein determining the methylation state comprises determining at least one of the following: a methylation state vector or a methylation CpG density.
37. The method of claim 36, wherein determining the methylation state vector comprises: Align the more than one sequence read with a reference sequence; Based on the alignment, the methylation status of one or more CpG sites in the sequence reads of the more than one sequence read and the location of the one or more CpG sites are determined; as well as The methylation status of the one or more CpG sites and the position of the one or more CpG sites are vectorized to generate a methylation status vector for the sequence reads in the more than one sequence read.
38. The method of claim 36, wherein determining the methylated CpG density comprises: Align the more than one sequence read with a reference sequence; Based on the alignment, the methylation status of one or more CpG sites in the sequence reads of the more than one sequence reads is determined; The methylation status of one or more CpG sites in the sequence read is used to determine whether the sequence read is methylated or unmethylated. For more than one sequence read, determine the count of methylated sequence reads and the count of unmethylated sequence reads; and The methylated CpG density is determined based on the counts of the methylated sequence reads and the counts of the unmethylated sequence reads.
39. The method according to claims 22-28, wherein training the prediction model based on the sequence data and the epigenetic data according to the more than one feature comprises training the prediction model according to a machine learning algorithm.
40. The method according to claim 39, wherein, The machine learning method includes at least one of the following: discriminant analysis, decision tree, nearest neighbor (NN) algorithm, Bayesian network, clustering algorithm, neural network, support vector machine (SVM), logistic regression algorithm, linear regression algorithm, Markov model or principal component analysis (PCA).
41. The method according to claims 1-40, comprising retraining the prediction model.
42. The method of claims 1-40, further comprising administering one or more therapies to the subject based on the fact that the source is tumor-derived.
43. The method of claim 42, wherein the therapy comprises administering chemotherapy, administering radiotherapy, or performing surgery to remove all or part of the tumor.
44. A method, at least in part, using a computer to distinguish between tumor-derived nucleic acid variants and undetermined clonal hematopoietic (CHIP)-derived nucleic acid variants from a test sample obtained from a test subject, said method comprising: Obtain sequence data, which includes more than one sequence read associated with more than one genomic region, wherein the sequence data is generated from a sample from the subject; Identify epigenetic data associated with the more than one sequence read; Provide at least a portion of the sequence data and at least a portion of the epigenetic data to a trained prediction model; and The presence or absence of tumor-derived nucleic acid variants and undetermined clonal hematopoietic (CHIP)-derived nucleic acid variants in the test sample is determined based on the prediction model.
45. The method of claim 44, further comprising generating the prediction model, wherein, Generating the prediction model includes: Sequence data for more than one sample are determined, wherein each of the more than one sample is labeled as tumor-derived or non-tumor-derived; Identify epigenetic data associated with the more than one sequence read; More than one feature of the prediction model is determined based on a portion of the sequence data and a portion of the epigenetic data; Based on a portion of the sequence data and a portion of the epigenetic data, the prediction model is trained according to more than one feature; and The prediction model is output based on the training.
46. A method for treating cancer in a test subject, the method comprising: Obtain sequence data, which includes more than one sequence read associated with more than one genomic region, wherein the sequence data is generated from a sample from the subject; Identify epigenetic data associated with the more than one sequence read; Provide at least a portion of the sequence data and at least a portion of the epigenetic data to a trained prediction model; and Based on the prediction model, the presence or absence of tumor-derived nucleic acid variants and undetermined clonal hematopoietic (CHIP)-derived nucleic acid variants in the test sample is determined. At least one therapy is administered to the test subject based on one or more distinct tumor-derived nucleic acid variants and CHIP-derived nucleic acid variants present in the test sample, thereby treating the cancer of the test subject.
47. A method of treating a test subject for cancer, the method comprising administering at least one therapy to the test subject based on one or more distinct tumor-derived nucleic acid variants and a set of discriminative clonal hematopoietic (CHIP)-derived nucleic acid variants present in the test sample, wherein the distinct tumor-derived nucleic acid variants and the set of CHIP-derived nucleic acid variants are generated by: A set of identified test nucleic acid variants is generated by using a computer to identify nucleic acid variants in a target genomic region based on sequence information obtained from nucleic acids in test samples obtained from the test subjects. The computer identifies at least one epigenetic imprint corresponding to a given test nucleic acid variant from a set of identified test nucleic acid variants, based on epigenetic information obtained from the nucleic acids in the test sample, to generate a test nucleic acid variant-epigenetic imprint cluster; and... The computer uses at least one trained classifier to distinguish between tumor-derived nucleic acid variants and CHIP-derived nucleic acid variants in the tested nucleic acid variant-epigenetic imprint cluster.
48. The method of claim 47, further comprising generating the prediction model, wherein, Generating the prediction model includes: Sequence data for more than one sample are determined, wherein each of the more than one sample is labeled as tumor-derived or non-tumor-derived; Identify epigenetic data associated with the more than one sequence read; More than one feature of the prediction model is determined based on a portion of the sequence data and a portion of the epigenetic data; Based on a portion of the sequence data and a portion of the epigenetic data, the prediction model is trained according to more than one feature; and The prediction model is output based on the training.
49. The method of claim 48, wherein the epigenetic data includes different epigenetic states or conditions exhibited by one or more epigenetic loci in a given target genomic region.
50. The method according to claims 48-49, wherein the different corresponding epigenetic imprints include different cell-free nucleic acid (cfNA) fragment lengths, locations, and / or endpoint density distributions.
51. The method according to claims 48-50, wherein a given targeted genomic region comprises two or more nucleic acid variant loci.
52. The method according to any preceding claim, wherein the nucleic acid in the sample comprises cell-free nucleic acid (cfNA) fragments and / or nucleic acid molecules obtained from one or more tissues or cells in the sample.
53. The method according to any preceding claim, wherein the epigenetic data includes the presence or absence of methylation, hydroxymethylation, acetylation, ubiquitination, phosphorylation, ubiquitin-like methylation, ribosylation, citrullination and / or histone post-translational modifications or other histone variations in one or more of the more than one genomic region.
54. The method according to any of the preceding claims further includes CHIP-derived nucleic acid variants that are ignored from further analysis.
55. The method according to any of the preceding claims further comprises generating at least one report listing tumor-derived nucleic acid variants and CHIP-derived nucleic acid variants distinguished from each other in the test sample.
56. The method according to any of the preceding claims further comprises identifying at least one cancer type associated with the distinguished tumor-derived nucleic acid variants.
57. The method according to any of the preceding claims further comprises administering at least one therapy to the test subject to treat the identified type of cancer.
58. The method according to any of the preceding claims further comprises administering at least one therapy to the test subject based on one or more of the tumor-derived nucleic acid variants of the aforementioned distinction.
59. One or more non-transitory computer-readable media having processor-executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the method according to any of the preceding claims.
60. A system comprising: A computing device configured to perform the method according to any of the preceding claims; and An output device configured to output the prediction model.
61. An apparatus comprising: One or more processors; and A memory storing processor-executable instructions, which, when executed by the one or more processors, cause the device to perform the method of any of the preceding claims.
62. One or more non-transitory computer-readable media having processor-executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the method according to any of the preceding claims.
63. A system comprising: A computing device configured to perform the method according to any of the preceding claims; and An output device configured to output an indication of whether the sample is of tumor origin or non-tumor origin.
64. An apparatus comprising: One or more processors; and A memory storing processor-executable instructions, which, when executed by the one or more processors, cause the device to perform the method of any of the preceding claims.
65. One or more non-transitory computer-readable media having processor-executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the method according to any of the preceding claims.
66. A system comprising: A computing device configured to perform the method according to any of the preceding claims; and An output device configured to output an indication of whether the sample is of tumor origin or non-tumor origin.
67. An apparatus comprising: One or more processors; and A memory storing processor-executable instructions, which, when executed by the one or more processors, cause the device to perform the method of any of the preceding claims.
Citation Information
Patent Citations
Methods for accurate sequence data and modified base position determination
US8486630B2
Methods and systems for analyzing nucleic acid molecules
WO2018119452A2
Compositions and methods for isolating cell-free DNA
WO2020160414A1