Technologies for integrated epigenomic exposure signature discovery
The ESDA algorithm integrates epigenetic assays to identify universal diagnostic signatures based on host responses, addressing the limitations of traditional pathogen-specific tests and enhancing diagnostic accuracy and speed.
Patent Information
- Application Number
- PCT/US2025/024104
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-10
- Filing Date
- 2025-04-10
- Publication Date
- 2025-10-16
AI Technical Summary
Traditional medical diagnosis often relies on specific tests for individual pathogens, which can be time-consuming and unreliable, especially in cases where pathogens are fleeting or have low detection limits, leading to delayed or incorrect treatment regimens.
A computer-implemented method using a machine learning algorithm, Epigenomic Exposure Signature Discovery Algorithm (ESDA), that integrates multiple epigenetic assays to identify universal diagnostic signatures by analyzing host responses to various exposures, focusing on epigenetic changes in regulatory mechanisms.
Enables rapid and accurate differentiation between exposure states, facilitating timely and precise diagnosis of infectious and non-infectious diseases, reducing the need for specific pathogen testing and improving treatment outcomes.
Smart Images

Figure US2025024104_16102025_PF_FP_ABST
Abstract
Description
TECHNOLOGIES FOR INTEGRATED EPIGENOMIC EXPOSURE SIGNATURE DISCOVERYCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 632,390, filed April 10, 2024, the entire disclosure of which is incorporated herein by reference.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0002] This invention was made with government support under Contract Nos. W911NF19C0041 and W911NF24C0023 awarded by the Defense Advanced Research Projects Agency (DARPA). The government has certain rights in the invention.BACKGROUND
[0003] Traditional medical diagnosis depends on direct testing for the pathogen that is causing the patient’s symptoms. However, testing for individual pathogens requires methods that are specific to each causative agent, which are often fleeting targets. In time-critical cases, it is essential to identify the correct cause of disease to provide lifesaving treatment. However, tests for particular pathogens may not be able to determine the cause due to low detection limits or transient presence that initiated the disorder.
[0004] By way of example, the proper treatment of systemic inflammatory response syndrome (SIRS) depends on whether it is caused by an infection (sepsis). Antibiotics are effective for microbial caused sepsis but have no effect on non-infectious SIRS. For this reason, rapid differentiation of the cause of SIRS is necessary to enable an accurate and timely treatment regimen. Blood cultures used to identify the bacterial etiology of sepsis take over 24 hours and have a high incidence of false negatives, and molecular tests for specific pathogens are not sensitive enough to diagnose bloodstream infections.
[0005] Another example relates to lower respiratory tract infections (LRTI), which have a high morbidity. Since non-infectious respiratory syndromes often look symptomatically like LRTI, this can confound diagnosis when methods are used that search for specific pathogens. To improve recovery rates, it is therefore necessary to develop more sensitive and universal diagnostic tests for rapid identification and treatment of both infectious and non-infectious diseases.SUMMARY
[0006] According to one aspect, a computer-implemented method may comprise obtaining feature data from a plurality of epigenetic assays performed on samples from a population, wherein the samples comprise first samples corresponding to a first exposure state and second samples corresponding to a second exposure state different from the first exposure state, training a plurality of assay-level models to distinguish between the first and second exposure states, wherein each one of the plurality of assay-level models is trained using feature data from one of the plurality of epigenetic assays, and combining the plurality of assay-level models to build an exposure-level model configured to distinguish between the first and second exposure states based on a subset of the feature data.
[0007] In some embodiments, the samples further comprise third samples corresponding to a third exposure state different from both the first exposure state and the second exposure state, each one of the plurality of assay-level models is trained to distinguish between the first, second, and third exposure states, and wherein the exposure-level model is configured to distinguish between the first, second, and third exposure states.
[0008] In some embodiments, each one of the plurality of epigenetic assays provides different feature data for the same sample.
[0009] In some embodiments, the plurality of epigenetic assays comprises at least one assay for chromatin accessibility. In some embodiments, the plurality of epigenetic assays comprises an ATAC-seq assay.
[0010] In some embodiments, the plurality of epigenetic assays comprises at least one assay for transcription. In some embodiments, the plurality of epigenetic assays comprises an RNAseq assay. In some embodiments, the plurality of epigenetic assays comprises an miRNAseq assay.
[0011] In some embodiments, the plurality of epigenetic assays comprises at least one assay for histone modification. In some embodiments, the plurality of epigenetic assays comprises a MINT- ChlP-seq assay. In some embodiments, the plurality of epigenetic assays comprises a ChlPmentation assay.
[0012] In some embodiments, the plurality of epigenetic assays comprises at least one assay for methylation. In some embodiments, the plurality of epigenetic assays comprises an EPIC assay. In some embodiments, the plurality of epigenetic assays comprises whole genome bisulfite sequencing.
[0013] In some embodiments, training the plurality of assay-level models comprises performing recursive feature elimination to iteratively reduce a number of features used by each one of the plurality of assay-level models to distinguish between the exposure states. In some embodiments, each one the plurality of assay-level models is a ridge classifier.
[0014] In some embodiments, the method further comprises normalizing the feature data from the plurality of epigenetic assays by genome position prior to training the plurality of assay-level models. In some embodiments, normalizing the feature data comprises assigning counts to nonoverlapping bins defined across positions in the whole human genome. In some embodiments, normalizing the feature data comprises assigning counts to one or more loci associated with one of the plurality of epigenetic assays.
[0015] In some embodiments, combining the plurality of assay-level models to build the exposure-level model comprises iteratively removing selected features from the exposure-level model that are uninformative. In some embodiments, combining the plurality of assay-level models to build the exposure-level model comprises iteratively removing selected features from the exposure-level model that are redundant with remaining features.
[0016] In some embodiments, the exposure-level model is a Sparse-Group LASSO (SGL) model. In some embodiments, the SGL model treats each one of the plurality of epigenetic assays as a group in the model.
[0017] In some embodiments, combining the plurality of assay-level models to build the exposure-level model comprises calculating an overall posterior probability for each exposure state as a weighted average of assay -level posteriors associated with the plurality of assay-level models. In some embodiments, the weight average of assay-level posteriors is weighted based on crossvalidated error rates for the plurality of assay-level models.
[0018] In some embodiments, the method further comprises selecting the plurality of assay-level models used to build the exposure-level model based on an area-under-the-curve value for each of the plurality of assay-level models exceeding a threshold value.
[0019] In some embodiments, combining the plurality of assay-level models to build the exposure-level model comprises combining the plurality of assay-level models with one or more genomic models to build the exposure-level model. In some embodiments, the one or more genomic models comprise a model trained to distinguish between the exposure states based on single nucleotide polymorphisms.
[0020] In some embodiments, the samples comprise peripheral blood mononuclear cell (PMBC) samples. In some embodiments, the first and second samples are PMBC samples.
[0021] In some embodiments, the first exposure state is exposure to methicillin-resistant Staphylococcus aureus (MRSA), and wherein the second exposure state is exposure to methicillinsensitive Staphylococcus aureus (MSSA). In some embodiments, the first exposure state is exposure to methicillin-resistant Staphylococcus aureus (MRSA), and wherein the second exposure state is non-exposure to MRSA. In some embodiments, the first exposure state is exposure to methicillin-sensitive Staphylococcus aureus (MSSA), and wherein the second exposure state is non-exposure to MSSA. In some embodiments, the first exposure state is exposure to human immunodeficiency virus (HIV), and wherein the second exposure state is non- exposure to HIV. In some embodiments, the first exposure state is exposure to influenza A virus, and wherein the second exposure state is non-exposure to influenza A virus. In some embodiments, the first exposure state is exposure to SARS-CoV-2, and wherein the second exposure state is non- exposure to SARS-CoV-2. In some embodiments, the first exposure state is a vaccinated status, and wherein the second exposure state is a non-vaccinated status. In some embodiments, the first exposure state is exposure to an explosive agent, and wherein the second exposure state is non- exposure to the explosive agent. In some embodiments, the first exposure state is exposure to opioid synthesis, and wherein the second exposure state is non-exposure to opioid synthesis. In some embodiments, the first exposure state is exposure to an organophosphate, and wherein the second exposure state is non-exposure to the organophosphate. In some embodiments, the first exposure state is a first geographic origin, and wherein the second exposure state is second geographic origin. In some embodiments, the first exposure state is an occurrence of a disease, and wherein the second exposure state is a non-occurrence of the disease. In some embodiments, the first exposure state is an occurrence of a cancer, and wherein the second exposure state is a nonoccurrence of the cancer. In some embodiments, the first exposure state is an occurrence of opioid use disorder, and wherein the second exposure state is a non-occurrence of opioid use disorder. In some embodiments, the first exposure state is a disposition toward opioid use disorder, and wherein the second exposure state is a non-disposition toward opioid use disorder. In some embodiments, the first exposure state is an occurrence of a mood disorder, and wherein the second exposure state is a non-occurrence of the mood disorder.
[0022] In some embodiments, the method further comprises obtaining the subset of feature data from the plurality of epigenetic assays performed on a new sample from the population, and operating the exposure-level model on the subset of the feature data to predict one of the exposure states for the new sample.
[0023] In some embodiments, the method further comprises analyzing a biological function associated with one or more features of the subset of feature data used by the exposure-level model.
[0024] According to other aspects: a computer system may be configured to perform any of the foregoing methods; a computer system may comprise one or more processors and memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform any of the foregoing methods; one or more computer-readable media may store instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any of the foregoing method; and a computer-readable medium storing an exposure-level model obtained according to the method of any one of claims 1 -47.BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The aspects and embodiments of the present disclosure can be more fully understood from the following detailed description taken in conjunction with the accompanying drawings, in which:
[0026] FIG. 1 is a simplified diagram illustrating aspects of an Epigenetic Exposure Signature Discovery Algorithm (ESDA) in which an exposure-level model is built from a plurality of assaylevel models;
[0027] FIG. 2 is a simplified diagram illustrating a portion the ESDA in which an assay-level model is built;
[0028] FIG. 3 is a simplified diagram illustrating another portion of the ESDA in which assaylevel models are combined into an exposure-level model;
[0029] FIG. 4 is a simplified diagram illustrating an optional, alternative portion of the ESDA using an ensemble-based exposure model;
[0030] FIG. 5 is a graph illustrating features identified using the ESDA for a negative control;
[0031] FIG. 6A is a graph illustrating features identified using the ESDA for HIV samples from a first cohort;
[0032] FIG. 6B is a graph illustrating the same feature identified in FIG. 6A applied to HIV samples from a second cohort;
[0033] FIG. 7A is a graph illustrating the predictive posterior probabilities of a number of assays (“H3K27ac”...“miRNA-seq”) and of the ESDA-generated exposure-level model (“ESDA”) for acute HIV versus negative;
[0034] FIG. 7B is a graph illustrating the predictive posterior probabilities of a number of assays (“H3K27ac”...“miRNA-seq”) and of the ESDA-generated exposure-level model (“ESDA”) for chronic HIV versus negative;
[0035] FIG. 8A is a graph illustrating the predictive posterior probabilities of a number of assays (“EPIC”, “RNA-seq”) and of the ESDA-generated exposure-level model (“ESDA”) for confirmed positive SARS-CoV-2 versus prediagnosis;
[0036] FIG. 8B is a graph illustrating the predictive posterior probabilities of a number of assays (“EPIC”, “RNA-seq”) and of the ESDA-generated exposure-level model (“ESDA”) for confirmed positive SARS-CoV-2 versus convalescent infection;
[0037] FIG. 8C is a graph illustrating the predictive posterior probabilities of a number of assays (“EPIC”, “RNA-seq”) and of the ESDA-generated exposure-level model (“ESDA”)for convalescent infection versus prediagnosis;
[0038] FIG. 9A is a graph illustrating the predictive posterior probabilities of a number of assays (“Mint-ChIP H3K27ac”. . . “miRNA-seq”) and of the ESDA-generated exposure-level model (“ESDA”) for methicillin-resistant Staphylococcus aureus (MRSA) versus control;
[0039] FIG. 9B is a graph illustrating the predictive posterior probabilities of a number of assays (“Mint-ChIP H3K27ac”. . .“miRNA-seq”) and of the ESDA-generated exposure-level model (“ESDA”) for methicillin-sensitive Staphylococcus aureus (MSSA) versus control;
[0040] FIG. 9C is a graph illustrating the predictive posterior probabilities of a number of assays (“Mint-ChIP H3K27ac”. . .“miRNA-seq”) and of the ESDA-generated exposure-level model (“ESDA”) for MRSA versus MSSA; and
[0041] FIG. 10 is a simplified block diagram of one illustrative embodiment of a computer system that may be used in conjunction with the present disclosure.DETAILED DESCRIPTION
[0042] While the concepts of the present disclosure are susceptible to various modifications and alternative forms, specific embodiments thereof have been shown by way of example in the drawings and will be described herein in detail. It should be understood, however, that there is no intent to limit the concepts of the present disclosure to the particular forms disclosed, but on the contrary, the intention is to cover all modifications, equivalents, and alternatives consistent with the present disclosure and the appended claims.
[0043] References in the specification to “one embodiment,” “an embodiment,” “an illustrative embodiment,” etc., indicate that the embodiment described may include a particular feature, structure, or characteristic, but every embodiment may or may not necessarily include that particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is submitted that it is within the knowledge of one skilled in the art to effect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described. Additionally, it should be appreciated that items included in a list in the form of “at least one A, B, and C” can mean (A); (B); (C); (A and B); (A and C); (B and C); or (A, B, and C). Similarly, items listed in the form of “at least one of A, B, or C” can mean (A); (B); (C); (A and B); (A and C); (B and C); or (A, B, and C).
[0044] The disclosed embodiments may be implemented, in some cases, in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried by or stored on a transitory or non-transitory machine-readable (e.g., computer- readable) storage medium, which may be read and executed by one or more processors. A machine-readable storage medium may be embodied as any storage device, mechanism, or other physical structure for storing or transmitting information in a form readable by a machine (e.g., a volatile or non-volatile memory, a media disc, or other media device).
[0045] In the drawings, some structural or method features may be shown in specific arrangements and / or orderings. However, it should be appreciated that such specific arrangements and / or orderings may not be required. Rather, in some embodiments, such features may be arranged in a different manner and / or order than shown in the illustrative figures. Additionally, the inclusion of a structural or method feature in a particular figure is not meant to imply that suchfeature is required in all embodiments and, in some embodiments, may not be included or may be combined with other features.
[0046] The present disclosure relates to a machine learning algorithm, referred to herein as the Epigenomic Exposure Signature Discovery Algorithm (ESDA), that identifies the most important features present in multiple epigenomic and transcriptomic datasets to produce an integrated exposure signature. The epigenome influences gene regulation and phenotypes in response to exposures. Thus, epigenome assessment can determine exposure history aiding in diagnosis. The presently disclosed ESDA was used to identify integrated exposure signatures for several infectious diseases. Specifically, in the illustrative embodiments described below, integrated exposure signatures were developed for six exposures including Staphylococcus aureus, human immunodeficiency virus (HIV), SARS-CoV-2, and influenza A (H3N2) virus. These integrated exposure signatures differed in the assays and features selected, as well as predictive value, integrated exposure signatures may be utilized for diagnosis or forensic attribution in illustrative embodiments. For instance, the ESDA can identify the most distinguishing features enabling diagnostic panel development for precision health deployment. Using the ESDA, we can build a database of signatures and simplify diagnostic testing, reducing the time it takes to determine the cause of symptoms. Understanding the epigenomic changes from different exposures also enables development of a universal diagnostic test for infectious diseases.
[0047] An alternative to direct testing for the pathogen or chemical causing a patient’ s symptoms (as described above in the Background) is to instead focus on the host’s response to the exposure, providing the opportunity for a universal diagnostic by leveraging commonalities in each exposure response. In all cases, whether symptoms are caused by an autoimmune deficiency or an infection, there is a human physiological response due to changes in molecular regulatory mechanisms. The environmental stimulus elicits a response through changes in methylation, regulatory RNAs, chromatin accessibility, and / or histone modifications that in turn cause changes in gene expression or post translational modification of proteins. By focusing on the effects of environmental changes on these regulatory mechanisms, a universal diagnostic platform can be developed for exposures from one multi-omic assay. Conventional research studies focus on the study of individual regulatory mechanisms in response to exposure, providing extensive data on differential exposure profiles, seldom attempting to integrate results from multi-omic data or distilling this information down to the most highly relevant features for distinguishing different exposures by a diagnostictest that can be conducted in a healthcare setting. With the recent shift in the scientific field toward multi-omic studies, an informatic tool has been needed to combine and integrate differential omic datasets into a combined, holistic exposure signature focusing on the most important features for differentiation to enable development of an exposure signature database and universal diagnostic assay, which could be used to aid in diagnosis of disorders or forensic investigations. The presently disclosed ESDA addresses this need.
[0048] The ESDA allows for the automated identification of multi-modal epigenetic biomarkers associated with exposure to biological and chemical sources of concern. The methodology used within the algorithm is intentionally designed to avoid reliance on annotated regions of the epigenome in public databases, as these annotations are typically skewed by findings from well- funded and often-studied domains, creating a potential bias toward identification of such markers (instead of other, potentially more informative markers). Instead, the ESDA is data-driven in the identification of the markers, after which a secondary analysis may optionally be conducted to link those markers back to annotated (or new) biological functions for interpretation.
[0049] As illustrated in the simplified diagram of FIG. 1, the ESDA begins with samples 10 (e.g., peripheral blood mononuclear cell (PMBC) samples) drawn from one or more cohorts or populations 12 that represent multiple exposure states. One typical use case is a two-state model for “exposed” and “healthy” (i.e., “not exposed”), but the states could also represent “exposure Type A” and “exposure Type B,” “acute exposure” and “chronic exposure,” or other multiple-state combinations. The ESDA also accommodates embodiments with more than two (e.g., three, four, five, six, etc.) exposure states as well, allowing identification of biomarkers that collectively discriminate between several types of exposures at once.
[0050] The next steps in the ESDA are to analyze the samples using a plurality of epigenetic assays 14 and to use the set of features generated by each assay 14 to train a corresponding assaylevel model 16 to distinguish between the exposure states based on that set of features. Each of these models 16 captures the relationship between the exposure states (e.g., “exposed” and “healthy”) and the features generated by a specific assay 14, allowing for the identification of assay-specific biomarkers that could be used for diagnosis or serve as therapeutic targets. The creation of the assay-level models 16 is further described below with reference to FIG. 2.
[0051] Next, the ESDA combines the assay-level models 16 into a single exposure-level model 18. This exposure-level model 18 relates all the features across the assays 14 to the exposure states(e g., “exposed” and “healthy”). In an iterative process, the ESDA removes features from the exposure-level model 18 that are uninformative or redundant with other features. By examining the remaining features, one can then ascertain which of the features are providing the discriminatory power (and would be good candidates for secondary investigation for biological relevance). The creation of the exposure-level model 18 is further described below with reference to FIGS. 3 and 4. Finally, the exposure-level model 18 may then be used with a new sample 20 to make predictions regarding unknown exposure(s).
[0052] Referring now to FIG. 2, a portion of the ESDA in which the features generated by an assay 14 are used to build a corresponding assay-level model 16 is shown in a simplified diagram. In this illustrative embodiment, the data (output by the assay 14) is first normalized by genome position, such as through binning (for sequence-based assays) or by locus (for methylation panels), as shown in Step (1) of FIG. 2. The ESDA was designed to work with multi-omic assays for chromatin accessibility (e.g., ATAC-seq), transcription (e g., RNAseq, miRNAseq), histone modification (e.g., MINT-ChlP-seq, ChlPmentation), and methylation (e.g., EPIC, whole genome bisulfite sequencing). In most cases, these assays rely on sequencing and alignment of reads to the human genome to better understand regulation of coding regions of the DNA or gene expression. EPIC arrays, while not sequencing and alignment based, rely on specific loci in the human genome. As such, this first step of the ESDA represents these different data sources in a common format to allow them to be modeled in similar ways. For assays like RNAseq or MINT-ChlP-seq that produce sequence alignments at any position in the genome, this common format is a series of binned sequence counts. Illustratively, non-overlapping bins of a 500 bp were defined across the positions in the whole human genome, then each aligned read was mapped to one and only one bin by finding the bin that most overlaps the positions occupied by the read. The features then become the count of reads mapped to each bin, ranging from zero to the maximum reads in the sample, with one feature per bin. For assays like the EPIC Infinium array, which focuses on specific loci within the genome, the typical representation is as a series of proportions (e.g., beta values) at those loci. In methylation data, for example, these proportions represent the fraction of reads that aligned at the locus and were methylated. Here, the features range from 0 to 1 and there is one feature per locus.
[0053] After normalization, the features can optionally be ranked and reduced through a statistical test (e.g., ANOVA p-value or fold-change criterion), as shown in Step (2) of FIG. 2.These steps may involve standardizing the features to account for variation in the number of reads generated and / or batch effects in the analysis (e.g., different reagent stocks, technicians, protocol variations across labs, etc.). For FASTQ data, for example, a quality filter was applied, the data was aligned to the hg38 reference, the sorted .bam file was converted to a sorted bed format, and the bed file was binned using BEDOPS (Neph et al., “BEDOPS: high-performance genomic feature operations,” Bioinformatics 28(14), 1919-1920 (2012)). As another example, DESeq2 can be used to normalize count-based binned data. For proportional locus-level data, such as the EPIC methylation array data, the beta values can be processed directly by the ESDA.
[0054] As shown in Step (3) of FIG. 2, the assay-level model 16 is then built using recursive feature elimination (RFE). RFE is a procedure that iteratively builds a predictive model, assesses the importance of the predictors in the model, then removes the least important predictors from consideration in subsequent iterations. RFE is model-agnostic, so any kind of classification model can be used in this step. For example, a ridge classifier works well in this role, as the shrinkage of the coefficients helps to accentuate the gradient between feature importance and the features can be more easily ranked at each step than in a LASSO model (since the coefficients are non-zero and ties do not occur). One illustrative embodiment employs the RFE approach described in Guyon et al., “Gene Selection for Cancer Classification using Support Vector Machines,” Machine Learning 46(1), 389-422 (2002), in this step. This approach can triage the large number of features simultaneously (typically, the number of features p is much larger than the sample size / ?) and identify the features that carry the most information to help identify the exposure states for the samples.
[0055] The RFE procedure of the ESDA produces a succession of models, each of which contains fewer features than the last, allowing an exploration of the tradeoff between model prediction accuracy and the number of features in order to settle on a model that is parsimonious for the application and still has reasonable predictive power. The features for this model are candidate assay-level biomarkers associated with the exposure states of interest and can be prioritized for review based on the magnitudes of the coefficients, which provide a natural ranking of their importance. The illustrative embodiment used a step size of 0.01 for RFE, meaning that the least important 1% of features were removed from consideration at each iteration.
[0056] Referring now to FIG. 3, a portion of the ESDA in which assay-level models 16 are combined into an exposure-level model 18 is shown in a simplified diagram. In Step (1) of FIG.3, a plurality of assay-level models 16 (each built according to the process described in FIG. 2) are gathered. These assay-level models 16 produce many candidate features (i.e., bins or loci in the genome) that could have potential value in diagnostics or therapeutics. However, since all of the assays are measuring different aspects of the same biological mechanisms, it is likely that many of these features will carry redundant information. To facilitate streamlining the feature set to arrive at a more parsimonious model, the assay-level features are combined into an overall profile associated with each sample, as shown in Step (2) of FIG. 3.
[0057] The ESDA then builds the exposure-level model 18 using a Sparse-Group LASSO (SGL) model that combines those features, with the assay 14 as the group label, and selects the optimal feature set that is both parsimonious and predictive of the exposure state, as shown in Step (3) of FIG. 3. As described in Simon et al., “A Sparse-Group Lasso,” Journal of Computational and Graphical Statistics 22(2), 231-245 (2013), the SGL is similar to a LASSO model (see Tibshirani, “Regression Shrinkage and Selection via the Lasso,” Journal of the Royal Statistical Society, Series B (Methodological) 58(1), 267-288 (1996)) in that it is a linear model that can assign coefficients with a value of zero to the features, effectively selecting for features through the nonzero coefficients. However, the SGL contains an additional penalty term in the loss function that enforces a group-level selection as well. In the ESDA, the groups are the assays 14, so the SGL will simultaneously select the most important assays as well as the most important features within those assays.
[0058] The primary requirement for using the SGL is that there must be a full set of features for each observation, meaning all of the assays 14 need to be run on the same population 12. In many cases, due to budget or technology access limitations, this may not be true. In cases where the cohorts vary across one or more assays, the ESDA can alternatively employ an ensemble modeling approach (rather than SGL) to build the exposure-level model 18, as illustrated in FIG. 4. Each observation in the dataset will have a subset of features that allow one or more assay -level models 16 to be built, and each of these assay-level models 16 will produce a posterior probability (Pi, P2, P3, etc.) for each exposure state. In this alternative approach, the overall posterior probability for each exposure state is calculated as a weighted average of those assay-level posteriors (Pi, P2, P3, etc.), where the weights are larger for the assay-level models 16 that had the best cross-validated error rates. The overall posteriors for the different exposure states can then be used as an ensemble-based exposure model 18 in downstream algorithms for biomarker discovery or predictive modeling.
[0059] In some embodiments, the bins selected by the ESDA may be mapped to biological function based on genome annotation. In one illustrative embodiment, the Homer bioinformatic tool (Heinz et al., “Simple combinations of lineage-determining transcription factors prime cis- regulatory elements required for macrophage and B cell identities,” Mol Cell 38(4), 576-589 (2010)) was implemented using the annotatePeaks.pl script to generate a tab delimited output file containing all of the selected bins from the ESDA, plus additional gene information for each line in the ESDA document including Gene Ontology enrichment category data. The Homer output was processed through an R script to gather and extract relevant information to annotate the original features with biological process ontologies, export the results in readable format, and generate a visualization of the results.
[0060] When there are far more predictors than samples, one concern with biomarker discovery algorithms is that they may find noisy markers that happen to separate the samples in the dataset on which they are trained, but which do not generalize to new cohorts within the population. To assess the potential for the ESDA to “hallucinate” signatures for exposures that do not exist, a negative case study was conducted in which the ESDA was told that one group of samples was exposed when in fact it was not. The desired behavior was for the ESDA would fail to find a signature that allowed one to predict the exposure labels with better than coin-flip accuracy. Specifically, thirty-five control samples were utilized from a study of individuals who were vaccinated against anthrax and worked with Bacillus anthracis in a controlled facility while wearing personal protective equipment. RNAseq data from seventeen of the control samples were randomly assigned labels of “exposed” and the other eighteen controls were assigned labels of “normal.” RNAseq assay-level models 16 were built and evaluated using five-fold cross validation, where model building was repeated five times with different samples left out, and cross- validation accuracies were calculated. The assay-level models 16 for this negative case study produced cross-validation accuracies of 0.40, 0.43, 0.57, 0.51, and 0.54 across the five folds, demonstrating that without a true exposure present in the data, the models 16 generated by the ESDA cannot accurately differentiate unexposed from exposed and can only achieve coin-flip accuracy. FIG. 5 displays the top ten features (bins) from the assay -level model 16 generated bythe ESDA in this study. As expected, there is no differentiation between the two groups of samples, since they both contain control samples from the same study.
[0061] While the negative case study of FIG. 5 shows that the ESDA will not identify a signature when no exposure is present, the positive case study reflected in FIGS. 6A and 6B shows that the ESDA will find an integrated exposure signature when there are truly exposed samples in the dataset (FIG. 6A) and that the integrated exposure signature will generalize to other cohorts within the same population (FIG. 6B). In the positive case study, an RNAseq assay-level model 16 was built and evaluated using leave one out cross validation (LOOCV) on ten acute infected and ten control samples for HIV data from the iPrEx study described in Grant et al., “Preexposure Chemoprophylaxis for HIV Prevention in Men Who Have Sex with Men,” New England Journal of Medicine 363(27), 2587-2599 (2010) (patients in North and South America). The top ten features (bins) from this assay-level model 16 are shown in FIG. 6A. The RNAseq assay-level model 16 was then applied to an HIV cohort (11 infected and 6 control samples) from the SCOPE study (https: / / hividgm.ucsf.edu / scope-study, with patients from California). As shown in FIG. 6B, these same features showed similar differentiation between HIV exposed and controls in the SCOPE samples. This example demonstrates the ESDA finding epigenetic biomarkers that not only differentiate HIV samples from controls in the training dataset (the iPrEx samples, FIG. 6A) but also for an independent cohort of samples from a different HIV study (the SCOPE samples, FIG. 6B).
[0062] The ESDA was utilized to develop exposure-level models 18 for HIV (acute and chronic), SARS-CoV-2 (naive, acute, and convalescent), Staphylococcus aureus (MRSA and MSSA), and influenza A. Table 1 below summarizes these exposure-level models 18, including the number of samples per group, the assays utilized, the positive predictive value (PPV), the negative predictive value (NPV), and the area under the ROC curve (AUC) for each. Six of the models 18 identified epigenetic signatures that can differentiate exposure to a chemical or pathogen with an AUC over 0.85, demonstrating that some signatures are stronger than others, and further data may identify more subtle differentiating signatures.
[0063] As reflected in FIGS. 7A-B, the iPrEx study cohort (ten paired samples from donors who were negative, those with acute diagnosis of HIV, and those with chronic HIV (on treatment)) was extensively studied through multiple analyses including ChlPmentation on six histone modifications, RNAseq, and miRNAseq. Assay -level models 16 and two exposure-level models 18 (ESDA in FIGS. 7A-B) were developed comparing acute and chronic HIV to negative, matched samples. While both exposure-level models 18 utilized the same three data types, RNAseq, miRNAseq, and H3K27ac ChlPmentation, the exposure-level model 18 for acute HIV, which utilized 2518 features, was not strong enough for an accurate diagnosis of HIV with an AUC of 0.76 (FIG. 7A, Table 1). However, the exposure-level model 18 for chronic HIV, utilizing 139 features, produced an AUC of 0.91 (Figure 7B, Table 1), demonstrating a clear differentiation that could be utilized as part of a diagnostic algorithm when differentiating HIV (on treatment) from other conditions is desired.
[0064] As reflected in FIGS. 8A-C, samples of 101 prediagnosis, 72 acute infection, and 104 convalescent participants from the CHARM study cohort (see Chen et al., “Mapping disease regulatory circuits at cell-type resolution from single-cell multiomics data,” Nature ComputationalScience 3(7), 644-657 (2023) and Letizia et al., “SARS-CoV-2 seropositivity and subsequent infection risk in healthy young adults: a prospective cohort study,” Lancet Respir Med 9(7), 712- 720 (2021)) were analyzed by RNAseq and EPIC arrays. Prediagnosis samples were defined as being collected at least 14 days prior to the first positive SARS-CoV-2 test. Similarly, convalescent infections were defined as being collected from individuals who have tested negative for SARS- CoV-2 for at least 14 days after having had an infection. Assay-level models 16 and three exposure-level models 18 were developed comparing acute and naive (prediagnosis), acute and convalescent infection, and naive and convalescent infection samples. Each exposure-level model 18 utilized both EPIC and RNAseq data. The models that compared patients with acute (FIG. 8A, Table 1) or convalescent (FIG. 8C, Table 1) SARS-CoV-2 infections readily distinguished these samples from SARS-CoV-2 prediagnosis samples with AUC of 0.89 and 0.85, respectively. However, the exposure-level model 18 comparing acute and convalescent SARS-CoV-2 infection samples was not strong enough for an accurate differentiation with an AUC of 0.73 using 2366 features (FIG. 8B, Table 1), demonstrating the effect of SARS-CoV-2 on the epigenome and transcriptome of people who have been infected are similar to people who are CO VID positive confounding differentiation of these two groups. The exposure-level model 18 comparing naive, prediagnosis participants and acute, confirmed positive SARS-CoV-2 infected participants selected 34 features: four RNAseq features and thirty EPIC features, listed in Table 2 below. Several of the features were involved in apoptosis, the immune response to viruses, or regulation of gene expression. Interestingly, a methylation feature was found associated with OR5B3, an olfactory receptor. As one symptom of SARS-CoV-2 is a loss of smell, the epigenetic regulation of OR5B3 may be involved.TABLE 2
[0065] As reflected in FIGS. 9A-C, seven methicillin-sensitive S. aureus (MSSA), thirteen methicillin-resistant S. aureus (MRSA), and twelve control samples were analyzed by ChlPmentation on six histone modifications, RNAseq, and miRNAseq. Assay -level models 16 and three exposure-level models 18 were developed comparing control to MRSA or MSSA andcomparing MSSA to MRSA samples. Each exposure-level model 18 utilized the same three data types, RNAseq, miRNAseq, and H3K4me3 ChlPmentation data. While this cohort had low power, the signals were very strong, enabling the ESDA to accurately distinguish control from MRSA or MSSA, and MSSA from MRSA with an AUC of 1.0 (see Table 1). The MRSA model 18 utilized twenty-seven features: five H3K4me3, fifteen miRNA, and seven RNAseq (see Table 3 below). The MSSA model 18 utilized seventeen features: nine H3K4me3, six miRNA, and two RNAseq (see Table 4 below). While some of the features used in the MRSA and MSSA models 18 are on the same chromosome, none were identical loci. When comparing MRSA to MSSA directly, the model 18 identified six features: two H3K4me3 features, two miRNAseq features, and two RNAseq features (see Table 5 below), which could differentiate methicillin resistant from sensitive S. aureus infections based on host response. Interestingly, in the MRSA model 18, one of the RNAseq features selected functions in H3-K4 histone methylation, and this histone modification was the modification selected for incorporation in the model 18 as one of the top three significant assays. Furthermore, the functions of several of the features were identified to be involved in immune system regulation, cytoskeleton dynamics, protein turnover, and regulation, which aligns with a response to the S. aureus infection.
[0066] In another illustrative embodiment, thirteen matched individuals pre-exposure and postexposure to influenza A (H3N2) from the the BARD A-Vacci tech FLU010 Study cohort (Evans et al., “Efficacy and safety of a universal influenza A vaccine (MVA-NP+M1) in adults when given after seasonal quadrivalent influenza vaccine immunisation (FLU009): a phase 2b, randomised, double-blind trial,” The Lancet Infectious Diseases 22(6), 857-866 (2022)) were analyzed by MINT-ChlPseq on six histone modifications, RNAseq, and miRNAseq. Assay-level models 16 and an exposure-level model 18 were developed comparing pre- and post-exposure to the influenza challenge in the control group for the vaccine study. The MINT-ChlP data on histone H3K4mel was best at differentiating exposure from unexposed with an AUC of 0.83, PPV of 0.72, and NPV of 0.81. The exposure-level model 18 utilized 123 features from the histone H3K4mel, RNAseq, and miRNAseq assay-level models 16, and although the NPV increased to 0.85, the PPV and AUC dropped to 0.53 and 0.81, respectively. The slight drop in predictive power for the exposure-level model 18 is due to the lower predictive power of the second and third most accurate assay-level models 16 combined with the H3K4mel (see Table 1). This demonstrates that combining multiple low predictive power assay -level models 16 does not always improve exposure prediction, and an AUC threshold for inclusion in the exposure-level model 18 can improve its performance.
[0067] The presently disclosed ESDA identifies key diagnostic biomarkers in the host epigenome that are indicative of exposure to a variety of biological and non-biological materials that alter molecular regulatory mechanisms. By focusing on the commonalities of the host response to these materials, it is possible to determine the cause of a disease or disorder, including whether it is infectious or non-infectious in its origin. Using such a technique has powerful ramifications in terms of developing assays that can rapidly characterize and diagnose the cause of disease in a symptomatic patient, thereby ensuring that the proper treatment can be selected in time. The present disclosure expands the biomarkers that can be used in a host-based diagnostic to include miRNA, DNA methylation, chromatin remodeling, histone modifications, and others, enabling more precise diagnostics.
[0068] Computer-implementation of the present disclosure can include a central processing unit (CPU), memory, and / or support circuits (or I / O), among other features. In embodiments having a memory, that memory can be connected to the CPU, and may be one or more of a readily available memory, such as a read-only memory (ROM), a random access memory (RAM), floppy disk, hard disk, cloud-based storage, or any other form of digital storage, local or remote. Software instructions, algorithms, and data can be coded and stored within the memory for instructing the CPU. Support circuits can also be connected to the CPU for supporting the processor in a conventional manner. The support circuits may include conventional cache, power supplies, clock circuits, input / output circuitry, and / or subsystems, and the like. A non-limiting one embodiment of a computer system 100 with which the present disclosure can be used and / or implemented is illustrated in FIG. 10.
[0069] More particularly, FIG. 10 is a simplified block diagram of one exemplary embodiment of a computer system 100 upon which the presently disclosed ESDA can be built, performed, operated, trained, etc. The system 100 can include a processor 110, a memory 120, a storage device 130, and an input / output device 140. Each of the components 110, 120, 130, and 140 can be interconnected, for example, using a system bus 150. The processor 110 can be capable of processing instructions for execution within the system 100. The processor 110 can be a singlethreaded processor, a multi -threaded processor, or similar device. The processor 110 can be capable of processing instructions stored in the memory 120 or on the storage device 130. The processor 110 may execute one or more of the operations described herein.
[0070] The memory 120 can store information within the system 100. In some implementations, the memory 120 can be a computer-readable medium. The memory 120 can, for example, be a volatile memory unit or a non-volatile memory unit. In some implementations, the memory 120 can store information related to reaction variables and parameters, among other information. The storage device 130 can be capable of providing mass storage for the system 100. In some implementations, the storage device 130 can be a non-transitory computer-readable medium. The storage device 130 can include, for example, a hard disk device, an optical disk device, a soliddate drive, a flash drive, magnetic tape, and / or some other large capacity storage device. The storage device 130 may alternatively be a cloud storage device, e.g., a logical storage device including multiple physical storage devices distributed on a network and accessed using a network.In some implementations, the information stored on the memory 120 can also or instead be stored on the storage device 130.
[0071] The input / output device 140 can provide input / output operations for the system 100. In some implementations, the input / output device 140 can include one or more of network interface devices (e.g., an Ethernet card), a serial communication device (e.g., an RS-212 3 port), and / or a wireless interface device (e.g., a short-range wireless communication device, an 802.7 card, a 4G wireless modem, a 5G wireless modem). In some implementations, the input / output device 140 can include driver devices configured to receive input data and send output data to other input / output devices, e.g., a keyboard, a printer, and / or display devices. In some implementations, mobile computing devices, mobile communication devices, and other devices can be used. In some implementations, the system 100 can be a microcontroller. A microcontroller is a device that contains multiple elements of a computer system in a single electronics package. For example, the single electronics package could contain the processor 110, the memory 120, the storage device 130, and / or input / output devices 140.
[0072] Although an example computer system has been described above, implementations of the subject matter and the functional operations described above can be implemented in other types of digital electronic circuitry, or in computer software, firmware, and / or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Implementations of the subject matter described in this specification can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible program carrier, for example a computer-readable medium, for execution by, or to control the operation of, a processing system. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them.
[0073] Various embodiments of the present disclosure may be implemented at least in part in any conventional computer programming language. For example, some embodiments may be implemented in a procedural programming language (e.g., “C”), or in an object-oriented programming language (e.g., “C++”). Other embodiments of the invention may be implemented as a pre-configured, stand-along hardware element and / or as preprogrammed hardware elements(e.g., application specific integrated circuits, FPGAs, and digital signal processors), or other related components.
[0074] Further features and advantages based on the above-described embodiments are possible and within the scope of the present disclosure. Accordingly, the disclosure is not to be limited by what has been particularly shown and described. All publications and references cited herein are expressly incorporated herein by reference in their entirety, except for any definitions, subject matter disclaimers, or disavowals, and except to the extent that the incorporated material is inconsistent with the express disclosure herein, in which case the language in this disclosure controls.
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented method comprising: obtaining feature data from a plurality of epigenetic assays performed on samples from a population, wherein the samples comprise first samples corresponding to a first exposure state and second samples corresponding to a second exposure state different from the first exposure state; training a plurality of assay-level models to distinguish between the first and second exposure states, wherein each one of the plurality of assay-level models is trained using feature data from one of the plurality of epigenetic assays; and combining the plurality of assay-level models to build an exposure-level model configured to distinguish between the first and second exposure states based on a subset of the feature data.
2. The method of any preceding claim, wherein the samples further comprise third samples corresponding to a third exposure state different from both the first exposure state and the second exposure state, wherein each one of the plurality of assay-level models is trained to distinguish between the first, second, and third exposure states, and wherein the exposure-level model is configured to distinguish between the first, second, and third exposure states.
3. The method of any preceding claim, wherein each one of the plurality of epigenetic assays provides different feature data for the same sample.
4. The method of any preceding claim, wherein the plurality of epigenetic assays comprises at least one assay for chromatin accessibility.
5. The method of any preceding claim, wherein the plurality of epigenetic assays comprises an ATAC-seq assay.
6. The method of any preceding claim, wherein the plurality of epigenetic assays comprises at least one assay for transcription.
7. The method of any preceding claim, wherein the plurality of epigenetic assays comprises an RNAseq assay.
8. The method of any preceding claim, wherein the plurality of epigenetic assays comprises an miRNAseq assay.
9. The method of any preceding claim, wherein the plurality of epigenetic assays comprises at least one assay for histone modification.
10. The method of any preceding claim, wherein the plurality of epigenetic assays comprises a MINT-ChlP-seq assay.
11. The method of any preceding claim, wherein the plurality of epigenetic assays comprises a ChlPmentation assay.
12. The method of any preceding claim, wherein the plurality of epigenetic assays comprises at least one assay for methylation.
13. The method of any preceding claim, wherein the plurality of epigenetic assays comprises an EPIC assay.
14. The method of any preceding claim, wherein the plurality of epigenetic assays comprises whole genome bisulfite sequencing.
15. The method of any preceding claim, wherein training the plurality of assay-level models comprises performing recursive feature elimination to iteratively reduce a number of features used by each one of the plurality of assay-level models to distinguish between the exposure states.
16. The method of any preceding claim, wherein each one the plurality of assay-level models is a ridge classifier.
17. The method of any preceding claim, further comprising normalizing the feature data from the plurality of epigenetic assays by genome position prior to training the plurality of assay -level models.
18. The method of claim 17, wherein normalizing the feature data comprises assigning counts to non-overlapping bins defined across positions in the whole human genome.
19. The method of claim 17, wherein normalizing the feature data comprises assigning counts to one or more loci associated with one of the plurality of epigenetic assays.
20. The method of any preceding claim, wherein combining the plurality of assay-level models to build the exposure-level model comprises iteratively removing selected features from the exposure-level model that are uninformative.
21. The method of any preceding claim, wherein combining the plurality of assay-level models to build the exposure-level model comprises iteratively removing selected features from the exposure-level model that are redundant with remaining features.
22. The method of any preceding claim, wherein the exposure-level model is a Sparse-Group LASSO (SGL) model.
23. The method of claim 22, wherein the SGL model treats each one of the plurality of epigenetic assays as a group in the model.
24. The method of any one of claims 1-17, wherein combining the plurality of assay-level models to build the exposure-level model comprises calculating an overall posterior probability for each exposure state as a weighted average of assay-level posteriors associated with the plurality of assay-level models.
25. The method of claim 24, wherein the weight average of assay-level posteriors is weighted based on cross-validated error rates for the plurality of assay-level models.
26. The method of any preceding claim, further comprising selecting the plurality of assaylevel models used to build the exposure-level model based on an area-under-the-curve value for each of the plurality of assay-level models exceeding a threshold value.
27. The method of any preceding claim, wherein combining the plurality of assay-level models to build the exposure-level model comprises combining the plurality of assay-level models with one or more genomic models to build the exposure-level model.
28. The method of claim 27, wherein the one or more genomic models comprise a model trained to distinguish between the exposure states based on single nucleotide polymorphisms.
29. The method of any preceding claim, wherein the samples comprise peripheral blood mononuclear cell (PMBC) samples.
30. The method of any preceding claim, wherein the first and second samples are PMBC samples.
31. The method of any preceding claim, wherein the first exposure state is exposure to methicillin-resistant Staphylococcus aureus (MRSA), and wherein the second exposure state is exposure to methicillin-sensitive Staphylococcus aureus (MSSA).
32. The method of any one of claims 1-30, wherein the first exposure state is exposure to methicillin-resistant Staphylococcus aureus (MRSA), and wherein the second exposure state is non-exposure to MRSA.
33. The method of any one of claims 1-30, wherein the first exposure state is exposure to methicillin-sensitive Staphylococcus aureus (MSSA), and wherein the second exposure state is non-exposure to MSSA.
34. The method of any one of claims 1-30, wherein the first exposure state is exposure to human immunodeficiency virus (HIV), and wherein the second exposure state is non-exposure to HIV.
35. The method of any one of claims 1-30, wherein the first exposure state is exposure to influenza A virus, and wherein the second exposure state is non-exposure to influenza A virus.
36. The method of any one of claims 1-30, wherein the first exposure state is exposure to SARS-CoV-2, and wherein the second exposure state is non-exposure to SARS-CoV-2.
37. The method of any one of claims 1-30, wherein the first exposure state is a vaccinated status, and wherein the second exposure state is a non-vaccinated status.
38. The method of any one of claims 1-30, wherein the first exposure state is exposure to an explosive agent, and wherein the second exposure state is non-exposure to the explosive agent.
39. The method of any one of claims 1-30, wherein the first exposure state is exposure to opioid synthesis, and wherein the second exposure state is non-exposure to opioid synthesis.
40. The method of any one of claims 1-30, wherein the first exposure state is exposure to an organophosphate, and wherein the second exposure state is non-exposure to the organophosphate.
41. The method of any one of claims 1-30, wherein the first exposure state is a first geographic origin, and wherein the second exposure state is second geographic origin.
42. The method of any one of claims 1-30, wherein the first exposure state is an occurrence of a disease, and wherein the second exposure state is a non-occurrence of the disease.
43. The method of any one of claims 1-30, wherein the first exposure state is an occurrence of a cancer, and wherein the second exposure state is a non-occurrence of the cancer.
44. The method of any one of claims 1-30, wherein the first exposure state is an occurrence of opioid use disorder, and wherein the second exposure state is a non-occurrence of opioid use disorder.
45. The method of any one of claims 1-30, wherein the first exposure state is a disposition toward opioid use disorder, and wherein the second exposure state is a non-disposition toward opioid use disorder.
46. The method of any one of claims 1-30, wherein the first exposure state is an occurrence of a mood disorder, and wherein the second exposure state is a non-occurrence of the mood disorder.
47. The method of any preceding claim, further comprising: obtaining the subset of feature data from the plurality of epigenetic assays performed on a new sample from the population, andoperating the exposure-level model on the subset of the feature data to predict one of the exposure states for the new sample.
48. The method of any preceding claim, further comprising analyzing a biological function associated with one or more features of the subset of feature data used by the exposure-level model.
49. A computer system configured to perform the method of any preceding claim.
50. A computer system comprising: one or more processors; and memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the method of any one of claims 1-48.
51. One or more computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 1- 48.
52. The one or more computer-readable media of claim 51, wherein the one or more computer-readable media are non-transitory.
53. A computer-readable medium storing an exposure-level model obtained according to the method of any one of claims 1-48.
54. The computer-readable medium of claim 53, wherein the computer-readable medium is non-transitory.
Citation Information
Patent Citations
Machine learning model trained to determine a biochemical state and / or medical condition using DNA epigenetic data
US11817214B1
Model-based featurization and classification
US20200365229A1
Neural-Network-Based Classifier
WO2023026996A1