Prediction of chromatin state

By integrating 5-methylcytosine and 5-hydroxymethylcytosine data into chromatin state prediction models, the method addresses the limitations of existing DNA sequence-based predictions, achieving improved accuracy and efficiency in chromatin state forecasting.

WO2025158025A1PCT designated stage Publication Date: 2025-07-31BIOMODAL LTD

Patent Information

Application Number
PCT/EP2025/051846
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-04
Filing Date
2025-01-24
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing methods for predicting chromatin state from DNA sequences are data-hungry and lack accuracy, requiring multiple costly and time-consuming assays, and integrating data across these assays is technically challenging.

Method used

Incorporating epigenetic information, specifically the presence of 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC), into chromatin state prediction models using 6-letter sequencing data, allowing for improved model tailoring to specific samples without extensive retraining, and enabling the integration of genetic and epigenetic information from a single assay.

Benefits of technology

The approach achieves single-base resolution predictions with enhanced accuracy compared to 4-letter sequencing alone, reducing the need for large training datasets and improving computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025051846_31072025_PF_FP_ABST
    Figure EP2025051846_31072025_PF_FP_ABST
Patent Text Reader

Abstract

Methods of predicting a chromatin state metric associated with a genomic region in a sample are provided. The methods comprise: receiving sequence data comprising genetic data and epigenetic data indicative of the presence of one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine at one or more genomic positions, the one or more epigenetic bases, the genetic data and epigenetic data obtained from a single assay; and predicting for each base or set of bases of the genomic region and using the sequence data associated with the genomic region, a value of the chromatin state metric, wherein the predicting is performed using a machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Prediction of chromatin state

[0002] Field of the Disclosure

[0003] The present invention relates to methods of predicting chromatin state information from 6-letters sequence data and particularly, although not exclusively, to methods of predicting enhancer status and / or chromatin accessibility using sequencing data comprising information about the presence of methylcytosine and hydroxy methylcytosine, and methods of performing genetic tests and screening perturbations using said methods.

[0004] Background

[0005] A key challenge in genomics is the generation and integration of data across different modalities. This typically requires several assays to be performed, which is costly and time-consuming. Moreover, the subsequent integration of these data across assays is technically challenging.

[0006] A few attempts have been made to predict gene expression and chromatin states from DNA sequences alone. In Avsec et al. (2021), DNA sequence data (i.e. genetic information alone) was used to predict gene expression using a deep learning architecture termed “Enformer”. Enformer is trained to predict human and mouse genomic tracks at 128-bp resolution from 200 kb of input DNA sequence, where the genomic tracks include transcriptional activity measured by CAGE (Cap analysis of gene expression), histone modification (measured by histone modification ChIP sequencing), transcription factor binding (measured by transcription factor ChIP sequencing) and DNA accessibility (measured by DNase-seq or ATA-seq). Similarly, the model chromBPNet (Kundaje et al. 2024) is a deep learning model trained to predict chromatin accessibility profiles (DNASE-seq and ATAC-seq) from genetic sequence data.

[0007] These have shown some promise but are extremely data hungry to train and the predictions are still relatively far from the corresponding experimental data.

[0008] The present invention has been devised in light of the above considerations.

[0009] Summary of the Disclosure

[0010] The inventors postulated that improved models of chromatin state could be obtained by including epigenetic information (specifically the presence of 5-methylocytosine (5mC) and 5-hydroxy- methylocytosine (5hmC)) when making predictions of chromatin state from sequence data. They further hypothesised that because this information is dynamic (whereas the genetic sequence is not), its inclusion in chromatin state prediction models from sequence data would enable the prediction to be better tailored to a particular sample (e.g. a particular cell line, individual, condition, etc.) without necessitating retrained of models for each such condition, or complex conditional mechanisms requiring very large amounts of training data. The inventors further postulated that additional benefits would be achievable if the single assay can also provide genetic information about the same DNA molecule (i.e. combining genetic and epigenetic information on a single read) for which epigenetic information is obtained. Such methods can be referred to as “6-letters sequencing” (or “6L-sequencing”). To test this hypothesis, they leveraged an assay (described in Fullgrabe et al. 2023) that enables sequencing the complete genetic sequence and the DNA modifications, 5- methylocytosine (5mC) and 5-hydroxy- methylocytosine (5hmC), from low nanogram amounts of DNA, to provide 6-base genomic data. They trained and evaluated a series of machine-learning models to predict chromatin accessibility and enhancer state from 6-base sequence data. They surprisingly demonstrated that they were able to obtain single base resolution predictions that have improved accuracy compared to comparative predictions made from 4-base data alone.

[0011] Thus, according to a first aspect, the disclosure provides a computer-implemented method of predicting a chromatin state metric associated with a genomic region in a sample, the method comprising: receiving (110) sequence data associated with the genomic region obtained from the sample, the sequence data comprising genetic data and epigenetic data indicative of the presence of one or more epigenetic bases at one or more genomic positions, the one or more epigenetic bases including methylated cytosine and hydroxy methylated cytosine, the genetic data and epigenetic data obtained from a single assay; and predicting (130), for each base or set of bases of the genomic region and using the sequence data associated with the genomic region, a value of the chromatin state metric, wherein the predicting is performed using a machine learning model trained to: take as input encoded sequence data for a genomic sequence, the encoded sequence data including: (i) for each of the one or more positions of the genomic sequence, encoded genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxy methylated cytosine, or (ii) values of one or more features derived from data including, for each of the one or more positions of the genomic sequence, genetic data and epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine, and the one or more features including at least one feature indicative of the presence of hydroxymethylated cytosine; and produce as output one or more chromatin state metrics for each base or set of bases of the genomic sequence.

[0012] The methods according to the present aspect may have any one or more of the following optional features.

[0013] The one or more chromatin state metrics may be selected from: metrics indicative of chromatin accessibility, and metrics indicative of histone modification. The one or more chromatin state metrics may be metrics that are obtained by DNA sequencing of a sample that has been subject to a protocol to enrich accessible regions of chromatin or regions of chromatin that are associated with a histone that carries a specific modification. In embodiments, a chromatin state metric indicative of chromatin accessibility is a metric obtained by ATAC-seq or DNAse-seq. In embodiments, a chromatin state metric indicative of chromatin accessibility is a base specific number of reads per genomic content. In embodiments, a chromatin state metric indicative of histone modification is a metric obtained by modified histone ChlP-seq. In embodiments, the histone modification is selected from: H3K4me1 , H3K4me3, and H3K27ac. In embodiments, a chromatin state metric indicative of histone modification is a fold change of a number of reads relative to a control condition. The machine learning model may be a deep learning model trained to: take as input encoded sequence data for a genomic sequence, the encoded sequence data including, for each position of the genomic sequence, encoded genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxy methylated cytosine; and produce as output a chromatin state metric for each base of the genomic sequence. In embodiments, the encoded sequence data further comprises for each base in the sequence, information indicating whether the base is on the sense or antisense strand.

[0014] In embodiments, the encoded sequence data comprises encoded genetic data and encoded epigenetic data for each position of each strand of the sense and antisense strand of the genomic sequence. In embodiments, the encoded sequence data includes, for each position of the sequence, genetic data encoded using a one-hot encoding scheme.

[0015] In embodiments, the encoded sequence data comprises encoded epigenetic data that includes, for each position of the sequence that is a C: one or more of, optionally at least three of: a number of reads that have mC at the position, a number of reads that have hmC at the position, a number of reads that have C at the position, and a total number of reads at the position, each number of read obtained from a sequencing technology that distinguishes between C, mC and hmC. Instead or in addition to this, the encoded sequence data may comprise encoded epigenetic data that includes, for each position of the sequence that is a C: one or more of, optionally at least two of: a fraction of reads that have mC at the position, a fraction of reads that have hmC at the position, and a fraction of reads that have C at the position, each fraction of read obtained from counts of reads obtained using a sequencing technology that distinguishes between C, mC and hmC. Instead or in addition to this, the encoded sequence data may comprise encoded epigenetic data that includes, for each position of the sequence that is a C: one or more of, optionally all of: left and right boundaries of a confidence interval around an estimate of a fraction of reads that have mC at the position, left and right boundaries of a confidence interval around an estimate of a fraction of reads that have hmC at the position, and left and right boundaries of a confidence interval around an estimate of a fraction of reads that have C at the position, each confidence interval obtained from counts of reads obtained using a sequencing technology that distinguishes between C, mC and hmC. The confidence intervals may be estimated using a sampling method assuming that the fractions follow a Dirichlet distribution. Instead or in addition to this, the encoded sequence data may comprise encoded epigenetic data that includes, for each position of the sequence that is a C: one or more of, optionally all of: a statistical estimate of a fraction of reads that have mC at the position, a statistical estimate of a fraction of reads that have hmC at the position, and a statistical estimate of a fraction of reads that have C at the position, each statistical estimate obtained from counts of reads obtained using a sequencing technology that distinguishes between C, mC and hmC. The encoded epigenetic data may further comprise a standard deviation or variance estimate associated with each of said statistical estimates. The statistical estimates may be means of respective categories of a Dirichlet distribution.

[0016] The machine learning model may be a deep learning sequence model. The machine learning model may be a deep neural network. The machine learning model may be a deep neural network comprising one or more transformer layers and / or one or more convolutional layers. In embodiments, the deep neural network comprises a plurality of convolutional layers with one or more residual connections. In embodiments, the deep neural network comprises one or more convolutional layers each comprising a 1 D convolution followed by an activation function. In embodiments, the deep neural network comprises one or more convolutional layers each comprising a dilated convolution. In embodiments, the deep neural network comprises one or more convolutional layers and does not comprise any pooling operations in or between any convolutional layer(s).

[0017] In embodiments, the machine learning model has been trained using training data comprising, for each of a plurality of training genomic regions: (a) encoded sequence data including: for each position of the sequence of the training genomic region, encoded genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine; and (b) measured values of the one or more chromatin state metrics. In embodiments, the machine learning model has been trained using training data comprising, for each of a plurality of training genomic regions: (a) encoded sequence data including: values of one or more features derived from data including, for each position of the sequence, genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position , the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine, and the one or more features including at least one feature indicative of the presence of hydroxymethylated cytosine; and b) measured values of the one or more chromatin state metrics.

[0018] In embodiments, the plurality of training genomic regions comprise regions of a predetermined size comprising: a transcription start site, an enhancer sequence, or an ATAC-seq, DNAse-seq or histone modification ChlP-seq peak. In embodiments, the plurality of training genomic regions comprise background regions of the predetermined size that do not comprise: an ATAC-seq, DNAse-seq or histone modification ChlP-seq peak. In embodiments, the background regions are selected to include, for each region of a predetermined size comprising an ATAC-seq, DNAse-seq or histone modification ChlP-seq peak, a background region of matched GC content that does not comprise an ATAC-seq, DNAse-seq or histone modification ChlP-seq peak.

[0019] The sequence data may comprise or consist of reads in 6-letters code. In embodiments, the method further comprises: obtaining the encoded sequence data from the received sequence data. In embodiments, the method further comprises obtaining the sequence data from a sample by processing the sample using a sequencing technology from which epigenetic and genetic bases can be called on the same read. Thus, in embodiments of any aspect described herein, the sequence data may have been obtained using a sequencing technology from which epigenetic and genetic bases can be called on the same read. In embodiments of any aspect described herein, the sequence data may comprise or consist of reads in 6- letters code.

[0020] In embodiments, the genomic region comprises an enhancer region, the chromatin state metric is a metric indicative of histone modification, and the metric indicative of histone modification is an enhancer status. In some such embodiments, the machine learning model is a machine learning classifier that takes as input values of one or more features derived from data including, for each position of the sequence, genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine, and the one or more features including at least one feature indicative of the presence of hydroxymethylated cytosine. Thus, also described according to the present aspect is a computer-implemented method of predicting the status of one or more enhancers in a sample, the method comprising: receiving sequence data associated with the one or more enhancers obtained from the sample, the sequence data comprising epigenetic data indicative of the presence of one or more epigenetic bases at one or more positions in the genome, the one or more genetic bases including methylated cytosine and hydroxymethylated cytosine; and predicting, for each enhancer of the one or more enhancers and using the sequence data associated with the enhancer, the status of the enhancer, wherein the predicting is performed using a machine learning classifier trained to classify an enhancer between a plurality of classes each associated with a different enhancer status selected from active, repressed and poised, using as input sequence data for an enhancer or one or more features derived therefrom including at least one feature indicative of the presence of hydroxy methylated cytosine. In embodiments, the machine learning classifier is a binary classifier or a multiclass classifier. In embodiments, the machine learning classifier is configured to classify a genomic region comprising an enhancer between at least three classes comprising a class associated with active enhancers, a class associated with poised enhancers, and a class associated with repressed enhancers. In embodiments, the machine learning classifier is trained to classify an enhancer between a plurality of classes using as input one or more features derived from sequence data for a genomic region comprising an enhancer including at least one feature indicative of the presence of hydroxymethylated cytosine, and the method comprises determining the value of said one or more features for the one or more enhancers. In embodiments, the machine learning classifier is a random forest classifier, a gradient boosted tree classifier, a neural network classifier, K-nearest neighbour classifier, or a support vector machine classifier. The one or more features may include one or more of: a feature indicative of the presence of hydroxy methylated cytosines in the enhancer, a feature indicative of the presence of methylated cytosines in the enhancer, a feature indicative of the presence of cytosines that are either methylated or hydroxy methylated in the enhancer, a feature indicative of the length of the enhancer, and a feature indicative of the number of CpGs in the enhancer. The one or more features may include a feature indicative of the presence of hydroxymethylated cytosines in the sequence data associated with the enhancer and a feature indicative of the presence of methylated cytosines in the sequence data associated with the enhancer. In embodiments, the one or more features include: a feature indicative of the presence of hydroxy methylated cytosines in the sequence data associated with the genomic region and a feature indicative of the presence of methylated cytosines in the sequence data associated with the genomic region. The epigenetic data may comprise sequence reads. The feature indicative of the presence of hydroxy methylated cytosines in the sequence data associated with the genomic region may be a summarised value, over one or more CpGs in the region, of a fraction of reads indicative of the presence of a hydroxymethylated cytosine at the CpG. The feature indicative of the presence of methylated cytosines in the sequence data associated with the genomic region may be a summarised value, over one or more CpGs in the region, of a fraction of reads indicative of the presence of a methylated cytosine at the CpG. A summarised value may be a mean value.

[0021] In embodiments, the machine learning classifier has been trained using training data comprising, for each of a plurality of genomic regions comprising an enhancer: (i) sequence data associated with the genomic region obtained from one or more samples, the sequence data comprising epigenetic data indicative of the presence of one or more epigenetic bases at one or more positions in the enhancer, the one or more genetic bases including methylated cytosine and hydroxymethylated cytosine; and (ii) enhancer status data derived from histone modification data, optionally derived from ChiP-seq data.

[0022] In embodiments, the genomic region comprises a transcription start site, the chromatin state metric is a summarised metric indicative of chromatin accessibility over the genomic region, and the machine learning model is a machine learning model that takes as input values of one or more features derived from data including, for each position of the sequence, genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine, and the one or more features including at least one feature indicative of the presence of hydroxymethylated cytosine. In some such embodiments, the machine learning is a regressor model. The machine learning model may be a random forest regressor, a gradient boosted tree regressor, or a neural network. The one or more features may include: a feature indicative of the presence of hydroxymethylated cytosines in the sequence data associated with the genomic region and a feature indicative of the presence of methylated cytosines in the sequence data associated with the genomic region. In embodiments, the epigenetic data comprises sequence reads, the feature indicative of the presence of hydroxymethylated cytosines in the sequence data associated with the genomic region is a summarised value, over one or more CpGs in the region, of a fraction of reads indicative of the presence of a hydroxymethylated cytosine at the CpG; and the feature indicative of the presence of methylated cytosines in the sequence data associated with the genomic region is a summarised value, over one or more CpGs in the region, of a fraction of reads indicative of the presence of a methylated cytosine at the CpG; optionally wherein a summarised value is a mean value. In embodiments, the machine learning classifier has been trained using training data comprising, for each of a plurality of genomic regions comprising a transcription start site: (i) sequence data associated with the genomic region obtained from one or more samples, the sequence data comprising epigenetic data indicative of the presence of one or more epigenetic bases at one or more positions in the genomic region, the one or more genetic bases including methylated cytosine and hydroxymethylated cytosine; and (ii) measured chromatin status metrics derived from chromatin accessibility data, optionally derived from ATAC-seq or DNAse-seq.

[0023] Also provided according to a further aspect is a computer-implemented method of training a machine learning model for predicting a chromatin state metric associated with a genomic region in a sample, the method comprising: receiving training data comprising: (i) sequence data associated with a plurality of genomic regions, the sequence data comprising genetic data and epigenetic data indicative of the presence of one or more epigenetic bases at one or more positions in the genome, the one or more genetic bases including methylated cytosine and hydroxymethylated cytosine, the genetic data and epigenetic data obtained from a single assay; and (ii) measured values of corresponding one or more chromatin state metrics; and training a machine learning model, using said training data, to: take as input encoded sequence data for a genomic sequence, the encoded sequence data including: (i) for each position of the sequence, encoded genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxy methylated cytosine, or (ii) values of one or more features derived from data including, for each position of the sequence, genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxy methylated cytosine, and the one or more features including at least one feature indicative of the presence of hydroxymethylated cytosine; and produce as output one or more chromatin state metrics for each base or set of bases of the genomic sequence. Also described according to the present aspect is a computer-implemented method of training a machine learning model for predicting the status of one or more enhancers in a sample, the method comprising: receiving a training data set comprising (i) sequence data associated with a plurality of enhancers, the sequence data comprising epigenetic data indicative of the presence of one or more epigenetic bases at one or more positions in the enhancers, the one or more genetic bases including methylated cytosine and hydroxymethylated cytosine, and (ii) status data comprising for each of the plurality of enhancer, a status label derived from histone modification data associated with the enhancer; and training a machine learning classifier, using said training data, to take as input sequence data for an enhancer or one or more features derived therefrom including at least one feature indicative of the presence of hydroxymethylated cytosine, and produce as output a classification between a plurality of classes each associated with a different enhancer status selected from active, repressed and poised. The methods according to the present aspect may have any of the features described in relation to any preceding aspect. Also described according to a further aspect is a method of performing a genetic test for a subject, the method comprising: obtaining sequence data associated with a sample previously obtained from a subject, the sequence data comprising genetic data indicative of the presence of one or more genetic bases at one or more positions in the genome and epigenetic data indicative of the presence of one or more epigenetic bases at one or more positions in the genome, the one or more genetic bases including methylated cytosine and hydroxymethylated cytosine; and performing the method of any embodiment of the first or second aspect to obtain one or more predicted chromatin state metrics. In embodiments, the method further comprises analysing the one or more predicted chromatin state metrics and / or genetic data by obtaining the value of one or more metrics characterising the sample. In embodiments, the one or more metrics are selected from a diagnostic or prognostic metric.

[0024] According to a further aspect, there is provided a system comprising at least one processor and a non- transitory computer readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform the method of any embodiment of any of the preceding aspects, or any method described herein.

[0025] According to a further aspect, there is provided a non-transitory computer readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform the method of any embodiment of any of the first to sixth aspects, or any method described herein.

[0026] According to a further aspect, there is provided a computer program comprising code which, when the code is executed on a computer, causes the computer to perform the method of any embodiment of any of the first to sixth aspects, or any method described herein. The invention includes the combination of the aspects and preferred features described except where such a combination is clearly impermissible or expressly avoided.

[0027] Summary of the Figures

[0028] Embodiments and experiments illustrating the principles of the invention will now be discussed with reference to the accompanying figures in which:

[0029] Figure 1 is a flowchart illustrating a method of predicting one or more chromatin state metrics according to a general embodiment of the disclosure (A), and methods of performing a genetic test and / or screening one or more perturbations according to general embodiments of the disclosure (B).

[0030] Figure 2 is a flowchart illustrating a method of providing a trained model for predicting one or more chromatin state metrics, according to embodiments of the disclosure.

[0031] Figure 3 shows schematically a system for implementing methods of the disclosure.

[0032] Figure 4 shows results of prediction of chromatin accessibility using a method as described herein, in terms of relationship and R2metric between predicted log(accessibility) and observed log(accessibility). Results are shown for chromosome 8 in ES-E14 embryonic stem cells (test data) from a model trained on all chromosomes except chromosome 8.

[0033] Figure 5 shows results of prediction of enhancer status using a method as described herein. A support vector machine classifier was trained to predict the state of an enhancer (active, primed or repressed) based on the mean 5mC and the mean 5hmC fraction of the enhancer. The different shades shown in the plot represent the decision boundaries of the classifier. Each point represents an enhancer, coloured by known enhancers status from ChlP-seq data. Enhancers are cis-acting regulatory regions that have a profound effect on cell-type specific gene expression programs. These regions are typically classified by histone modifications in flanking nucleosomes. Active enhancers (H3K4me1 & H3K27ac) have low 5mC and 5hmC levels, primed enhancers (H3K4me1 but not H3K27ac) moderate 5mC and high 5hmC levels, and repressed enhancers (H3K9me3) have high 5mC and low 5hmC levels. An SVM model is able to classify the three groups with 85.5% accuracy based on their 5mC and 5hmC levels only.

[0034] Figure 6 shows schematically alternative approaches for predicting chromatin state metrics from 6 letters sequencing data. TSS=transcription start site.

[0035] Figure 7 shows examples of encoding schemes for encoding 4-base and 6-base sequence data.

[0036] Figure 7A shows an example of one-hot encoding of canonical 4-base data. An additional number can be used to indicate the strand (which can be e.g. 0 for the positive / sense strand and 1 for the negative I antisense strand) that the base is associated with.

[0037] Figure 7B shows examples of encoding of 6-base data which uses counts or fractional counts at C positions. As above, an additional number can be included to indicate the strand that the base is associated with. Figure 7C shows examples of encoding of 6-base data which uses statistics of model-based estimates of fractions of C, mC and 5hmC at C positions. sd=standard deviation. Cl=confidence interval.

[0038] Figure 8 shows and example of encoding of a sequence using one of the schemes on Fig. 7B.

[0039] Figure 9 shows an exemplary model architecture for prediction of chromatin state according to the disclosure, using convolutional neural networks.

[0040] Figure 10 shows results of prediction of chromatin accessibility using methods described herein.

[0041] Where the figures laid out herein illustrate embodiments of the present invention, these should not be construed as limiting to the scope of the invention. Where appropriate, like reference numerals will be used in different figures to relate to the same structural features of the illustrated embodiments.

[0042] Detailed Description

[0043] Aspects and embodiments of the present invention will now be discussed with reference to the accompanying figures. Further aspects and embodiments will be apparent to those skilled in the art. All documents mentioned in this text are incorporated herein by reference.

[0044] The methods described herein are computer-implemented unless context specifies otherwise (such as e.g. where measurement steps and / or wet steps are involved). Thus, the methods described herein are typically performed using a computer system or computer device. Any reference to an action such as “obtaining”, “processing”, “determining” may therefore refer to a processor performing the action, or a processor executing instructions that cause the processor to perform the action. Indeed, the methods of the present invention comprising sequence data analysis (requiring the analysis of thousands of reads per sample) and machine learning model training and / or deployment, is such that it cannot be performed in the human mind.

[0045] As used herein, the terms “computer system” of “computer device” includes the hardware, software and data storage devices for embodying a system or carrying out a computer implemented method. For example, a computer system may comprise one or more processing units such as a central processing unit (CPU) and / or a graphical processing unit (GPU), input means, output means and data storage, which may be embodied as one or more connected computing devices. Preferably the computer system has a display or comprises a computing device that has a display to provide a visual output display (for example in the design of the business process). The data storage may comprise RAM, disk drives or other computer readable media. The computer system may include a plurality of computing devices connected by a network and able to communicate with each other over that network. For example, a computer system may be implemented as a cloud computer. The term “computer readable media” includes, without limitation, any non-transitory medium or media which can be read and accessed directly by a computer or computer system. The media can include, but are not limited to, magnetic storage media such as floppy discs, hard disc storage media and magnetic tape; optical storage media such as optical discs or CD-ROMs; electrical storage media such as memory, including RAM, ROM and flash memory; and hybrids and combinations of the above such as magnetic / optical storage media. A computer system may comprise one or more networked computers, including e.g. a cloud computer. The methods described herein may be provided as computer programs or as computer program products or computer readable media carrying a computer program which is arranged, when run on a computer, to perform the method(s) described herein.

[0046] The present disclosure relates broadly to the use of epigenetic data (alone or in combination with genetic data - collectively, “sequence data”) and machine learning to predict one or more chromatin state metrics, such as e.g. chromatin accessibility, histone modification and / or enhancer status.

[0047] DNA encodes information in both genetic and epigenetic bases. Epigenetic information is closely associated with transcription, ultimately influencing cell fate, disease etc. DNA methylation, the covalent addition of a methyl group to cytosine residues (primarily at CpG sites) is the most widely studied epigenetic mark. DNA methylation is typically measured using whole genome bisulfite sequencing (WGBS, see Frommer et al. 1992), enzymatic methyl-sequencing (EM-seq, see Vaisvila et al. 2021) or DNA methylation microarrays. All of these technologies investigate methylation only and genetic information is acquired separately, typically through high-throughput sequencing (also referred to as next generation sequencing) (Villicana and Bell, 2021). Recently, approaches for the simultaneous sequencing of genetic and epigenetic bases have been proposed. Fullgrabe et al. (2023) described a single base-resolution sequencing methodology that identifies G, C, T and A as well as 5-methylcytosine and 5-hydroxymethylcytosine (either combined, also known as “5-letters sequencing”, or separately, also known as “6-letters sequencing”). Sequence data refers to information indicative of the presence of any of the 4 genetic bases (A, T, C, G) and / or any epigenetic base (e.g. 5mC and / or 5hmC) at one or more positions in the genome. Sequence data encompasses epigenetic data and genetic data. Genetic data refers to information indicative of the presence of any of the 4 genetic bases (A, T, C, G) at one or more positions in the genome. Epigenetic data refers to information indicative of the presence of one or more epigenetic bases at one or more positions in the genome. Epigenetic bases include methylated cytosine (5-mC or 5mC), hydroxymethylated cytosine (5-hmC or 5hmC), 5-carboxycytosine (5-caC), formylcytosine (5-fC), and methyladenosine (6-mA). The present disclosure relates in particular to the use of epigenetic data including information about the presence of 5-methylcytosine and 5-hydroxymethylcytosine (also known as “6-letters sequencing”). Thus, in embodiments, epigenetic bases include or are selected from methylated cytosine (5-mC or 5mC) and hydroxy methylated cytosine (5-hmC or 5hmC). In embodiments, sequence data is acquired using a sequencing technology from which epigenetic and genetic bases can be called on the same read. The technology used may be a technology that identifies A, T, C, G and methylated cytosine, also referred to as “5-letters sequencing”. Alternatively, the technology used may be a technology that identifies A, T, C, G, mC and hmC, also referred to as “6-letters” sequencing. Sequencing technologies from which epigenetic and genetic bases can be called on the same read include the technology provided by biomodal (see biomodal.com / product / ), described in e.g. WO 2022 / 023753 A1 and Fullgrabe et al. (2023), Oxford Nanopore technologies (e.g. using the PromethlON instrument) as described in Simpson et al. (2017), or Pacific Biosciences (e.g. using 5-base HiFi sequencing as described in PacBio 2022, “Measuring DNA methylation with 5-base HFi sequencing”, available at www.pacb.com / wp-content / uploads / application- brief-measuring-dna-methylation-with-5-base-hifi-sequencing.pdf). Sequencing technologies from which epigenetic and genetic bases can be called on the same read in 6 letters code include the technology provided by Biomodal (see biomodal.com / product / ), described in e.g. WO 2022 / 023753 A1 and Fullgrabe et al. (2023).

[0048] The present inventors have surprisingly discovered that 6-base sequence data can be used to obtain extremely accurate machine learning models predicting chromatin state information (also referred to as “chromatin state metrics”), such as chromatin accessibility data (e.g. from ATAC-seq), histone modification data (e.g. from histone modification chromatin immunoprecipitation and sequencing, histone ChlP-seq) and derived metrics such as enhancer status.

[0049] The term “chromatin state metric” refers to a metric indicative of the structure of chromatin at a particular base or set of bases, including chromatin accessibility and histone modification. A chromatin state metric may be base-specific. A measured chromatin state metric may be a metric that is obtained by DNA sequencing of a sample that has been subject to a protocol to enrich accessible regions of chromatin that may be present in the sample or regions of chromatin that are associated with a histone that carries a specific modification. A predicted chromatin state metric may be a metric that is predicted by a model as described herein, corresponding to a measured chromatin state metric (i.e. using a model that has been trained to predict said measured chromatin state metric for an input sequence). A chromatin state metric may be a metric associated with chromatin accessibility, such as e.g. a metric obtained by ATAC-seq (Assay for Transposase-Accessible Chromatin with Sequencing) or DNAse-seq (also known as “DNase I- hypersensitive site sequencing”). Any base-specific metric known in the art derived from these technologies may be used in the context of the present disclosure. Metrics obtained by DNAse-seq or ATAC-seq are typically derived from the number of reads that map to a position in a reference genome (or any suitable reference, the term “reference genome” being used here as a shorthand to refer to any DNA sequence reference to which genetic and epigenetic sequencing reads can be aligned, the choice of which may depend on the sample being analysed). Thus, reads obtained from sequencing are mapped to a reference genome and the number of reads mapped to a particular position (also referred to as count, read depth or simply depth) or normalised versions thereof are used. For example, a metric obtained by ATAC-seq or DNAse-seq may be a number of RPGC (reads per genomic content). The RPGC for a base can be calculated as the number of reads that map to the base, normalised to 1x sequencing depth, where sequencing depth is defined as: (number of mapped reads x fragment length) / genome size. Protocols for obtaining and analysing ATAC-seq data are known in the art (see e.g. Grandy et al. 2022). Protocols for obtaining and analysing DNAse-seq data are known in the art (see e.g. Liu et al. 2019). A chromatin state metric may be a metric that is associated with histone modification. A metric that is associated with histone modification may be a metric that is associated with the presence of a particular histone modification at the genomic position of a base. A histone modification may be selected from: H3K4me1 , H3K4me3, and H3K27ac. A metric that is associated with histone modification may be a metric that is obtained from histone modification chromatin immunoprecipitation and sequencing (ChlP-seq). Histone modification ChlP-seq comprises using an antibody that specifically binds to a particular modified form of a histone protein to pull down fragmented DNA that is physically associated with the modified histone, then sequencing this DNA. The sequenced material is typically compared to sequenced material from a control sample which has been processed in the same way but without the specific antibody. Metrics obtained by ChlP-seq are typically derived from the number of reads that map to a position in a reference genome. Thus, reads obtained from sequencing are mapped to a reference genome and the number of reads mapped to a particular position (also referred to as count, read depth or simply depth) or normalised versions thereof are used. For example, fold change of the number of reads mapping to a position relative to the number of reads mapped to the same position in a control sample may be used. Protocols for obtaining and analysing histone ChlP- seq data are known in the art (see e.g. O’Geen et al. 2014). Metrics derived from any of the above may also be used, such as e.g. classification labels associated with respective range of values of a chromatin state metric. For example, instead of using a metric as provided by an assay as described herein, ranges of observed values of the metric may be associated with respective label (such as e.g. low, medium, high accessibility). These labels may also be referred to as a chromatin state metric, and may be used to train models as described herein to provide corresponding predicted labels for input sequences. Peak calling algorithms may be used as known in the art to identify regions of interest from any of the types of data above that provide sequencing reads in regions of interest. Such algorithms may be used to identify regions of interest for training machine learning models as described herein. For example, machine learning models as described herein may be trained using training data comprising encoded genetic and epigenetic data for regions that encompass peaks in accessibility / histone modification data, and corresponding ground truth chromatin state metrics. The training data may further comprise regions of matched length and GC content in which peaks were not identified. These may be referred to as “background regions”. A region of matched GC content may refer to a region that has a GC content within a predetermined range from the GC content of another region. The predetermined range may be 0 (i.e. matched regions may have identical GC content), or e.g. 1 , 2, 3, 4, 5, 6, 7, 8, 9 or 10%. GC content can be expressed as count(G+C) / count(A+T+G+C)*100%. Genomic regions as used herein may have fixed (predetermined) lengths or variable lengths. Fixed lengths regions are particularly useful in embodiments in which a machine learning model takes as input encoded sequence data including for each position of the sequence, encoded genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine. Variable lengths regions are particularly useful in embodiments wherein the machine learning mode takes as input values of one or more features derived from data including, for each position of the sequence, genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxy methylated cytosine, and the one or more features including at least one feature indicative of the presence of hydroxymethylated cytosine. When using a predetermined length, the length may be e.g. 1 kbp, 1.5kbp, 2kbp or more.

[0050] Enhancers (also referred to as “enhancer sequences”) are genetic loci that have a regulatory role in gene expression. Enhancers are bound by specific transcription factors, resulting in enhanced transcription of one or more genes. The genes regulated by an enhancer can be located upstream or downstream of the enhancer, on the same chromosome. Enhancers are classified in multiple categories depending on the pattern of histone modifications at the genomic location of the enhancer. Each of these categories is referred to herein as a “status”. Thus, enhancer status refers to a category of enhancer state defined by the presence of one or more histone modifications. Enhancers are typically classified in 3 categories: active enhancers, primed enhancers (also referred to as “poised” enhancers) and repressed enhancers. Active enhancers are believed to actively enhance gene expression. They are typically marked by the presence of H3K4me1 and H3K27ac modifications in flanking nucleosomes. Primed enhancers are believed to be in a state that is between repressed and activate enhancers, and generally indicate that they are about to activate gene expression. They are marked by the presence of H3K4me1 , but not H3K27ac. Repressed enhancers do not enhance gene expression. They are typically marked by the presence of H3K27me3 or H3K9me2 / 3. Thus, prediction of enhancer status refers to predicting whether or not an enhancer is in a state selected from active, primed and repressed. Prediction of enhancer status also refers to classifying an enhancer between a plurality of categories selected from active, primed, repressed and “other”. For example, prediction of enhancer status can refer to determining whether or not an enhancer is active (which is equivalent to classifying an enhancer between “active” and “other”). The same principle can be applied to poised and repressed enhancers. Alternatively, prediction of enhancer status can refer to determining whether an enhancer is active, repressed or “other” (which encompasses any enhancer that does not possess the histone modification marks to be classified as either active or repressed). The same principle can be applied to any combination of active, repressed and primed. As another example, prediction of enhancer status can refer to classifying an enhancer between active, repressed and primed. These are all equivalent to predicting whether the enhancer shows the histone modification marks described above as associated with the enhancer. For example, determining whether an enhancer is active can also be described as determining whether the enhancer has H3K4me1 and H3K27ac modifications. Similarly, determining whether an enhancer is primed can also be described as determining whether the enhancer has H3K4me1 , but not H3K27ac modifications. Similarly, determining whether an enhancer is repressed can also be described as determining whether the enhancer has H3K27me3 or H3K9me2 / 3 modifications.

[0051] Enhancer status can be experimentally determined (for use as ground truth or test data when training a model as described herein) or predicted. Enhancer status can be experimentally determined using any method known in the art for detecting the presence of histone modifications. Histone modifications are commonly detected using chromatin immunoprecipitation (ChIP) combined with either microarrays (ChlP- on-chip) or high-throughput sequencing (ChlP-seq). Thus, reference to “measured enhancer status” or “experimentally determined enhancer status” or “ground truth enhancer status” may refer to an enhancer status category that has been associated with an enhancer by analysis of ChlP-seq or ChlP-on-chip data.

[0052] Enhancer status can be predicted using methods as described herein, which can either predict enhancer status directly from 6-base data (e.g. using a machine learning model trained to make such prediction using features derived from such data over the enhancer region, or using a machine learning model that is a deep learning sequence model trained to make such predictions for an enhancer region using encoded 6-base sequence data), or can predict histone modification data at enhancer regions using a deep learning sequence model trained to predict histone modification data using encoded 6-base sequence data.

[0053] The term “machine learning model” refers to a mathematical model that has been trained to predict one or more output values (predicted variables) based on input data (predictive variables, also referred to as values of predictive features). Training refers to the process of learning, using training data, the parameters of the mathematical model that result in a model that can predict outputs values that satisfy an optimality criterion or criteria. In the case of supervised learning, training typically refers to the process of learning, using training data, the parameters of the mathematical model that result in a model that can predict output values with minimal error compared to comparative (known) values associated with the training data (where these comparative values are commonly referred to as “labels” or “ground truth”). The are two major types of supervised learning models: classification models and regression models. Classification models aim to classify observations between a plurality of categories. They are typically suitable when the output to be predicted is a categorical label in the training data (e.g. enhancer status or class of chromatin accessibility). Regression models aim to predict the value of a continuous variable associated with observations (e.g. a continuous chromatin accessibility level). They are typically suitable when the output to be predicted is a continuous value in the training data. A classification model may provide as output a classification label and / or one or more probabilities of an observation belonging to respective one or more classes. A binary classification model may provide as output a single probability indicating the probability that the observation belongs to a positive class (e.g. accessible regions). A predetermined threshold may be applied to the one or more probabilities to assign a class label. The threshold may be determined based on desired characteristics of the classification, such as e.g. a desired level of specificity, sensitivity or any accuracy performance that combines aspects of specificity (precision) and sensitivity (recall), such as accuracy and F1 score or balanced versions thereof that take into account the proportions of observations in the training data in each of the classes.

[0054] The term “machine learning algorithm” or “machine learning method” refers to an algorithm or method that trains and / or deploys a machine learning model. The machine learning models of the present disclosure are trained by supervised learning. An optimality criterion may be the minimisation of a loss function that quantifies the model prediction error based on the observed (ground truth) and predicted values of the predicted variables. Suitable loss functions for use in training machine learning models are known in the art and include the mean squared error, and the mean absolute error. Any of these can be used according to the present disclosure. The mean squared error (MSE) can be expressed as £(•) = MS£’(xi,xi) = (x j - j)2where xtand xtare the predicted and observed (ground truth) values for a data instance (in the present case, a candidate drug combination in the training data), respectively. The mean absolute error (MAE) can be expressed as: £(•) = MAE x_i,x'_i ) = [ |xj J - x'_i |where xtand xtare the predicted and observed (ground truth) values for a data instance (e.g. a known chromatin state metric), respectively. The MAE is believed to be more robust to outlier observations than the MSE. The MAE may also be referred to as “L1 loss function”. However, MSE remains a very commonly used loss functions especially when a strong effect from outliers is not expected, as it can make optimization problems simpler to solve. Regularised loss functions are functions that include a loss function as described above, and one or more terms penalizing model complexity in order to reduce the risk of overfitting. Overfitting is a phenomenon that occurs where a machine learning model is trained to very closely reproduce the features of a training data set, resulting in poorer performance on other datasets that do not have the same characteristics (i.e. poor generalizability). L1 regularisation (also known as “Lasso” in the context of regression) add a regularization term to the loss function that penalizes models based on the sum of absolute value of the coefficients of the model. L2 regularisation (also known as “Ridge” in the context of regression) add a regularization term to the loss function that penalizes models based on the sum of squared value of the coefficients of the model. L1 regularisation can be used as a feature selection method as it minimizes the coefficients associated with less informative predictive features. In embodiments, the machine learning model is a regularized model. In embodiments, the machine learning model is a regularized tree-based model. In embodiments, the machine learning model is a regularized gradient boosted decision tree model. Examples of such models are available in the XGBoost software library (xgboost.ai / ). Such models may be referred to as “XGBoost” models, although any other implementation of regularized gradient boosted models may equally be used. A machine learning model as described herein may be selected from: decision trees and variants thereof including regularised and / or gradient boosted decision trees and random forest models, regularised discriminant analysis, logistic regression models, artificial neural networks (ANNs) including multilayer perceptrons (with linear or non-linear activation functions) and deep learning models (e.g. long short-term memory networks (LSTMs), Recurrent neural networks (RNNs), Generative Adversarial Networks (GANs), etc), naive Bayes classifiers, support vector machines (SVM, using linear or non-linear kernels such as radial basis function) and multivariate adaptive regression splines (MARS). For the classification of enhancer status, the present inventors have found it to be particularly beneficial to use non-linear models, such as decision trees and variants thereof (including in particular random forests and gradient boosted trees), SVM with a non-linear kernel, and ANNs (e.g. multilayer perceptrons) with nonlinear activation functions.

[0055] In embodiments, a machine learning model comprises an ensemble of models whose predictions are combined. Alternatively, a machine learning model may comprise a single model. Random forest models and gradient boosted tree models (such as XGBoost) are ensemble models. Ensemble versions of any models can be constructed. Ensemble models are expected to result in better prediction performance than single models. For example, the machine learning model may be a random forest classifier or regressor, or a gradient boosted decision tree or regressor model. A random forest classifier is a model that comprises an ensemble of decision trees and outputs a class that is the average prediction of the individual trees. Decision trees perform recursive partitioning of a feature space until each leaf (final partition sets) is associated with a single value of the target. Gradient boosting is a machine learning method that forms an ensemble of weak prediction models (e.g. decision trees) from which a combined strong prediction is obtained. The algorithm iteratively adds new weak predictors to improve the prediction obtained by combining the outputs of the weak predictors. By contrast, random forest iteratively trains a set number of trees using random subsets of the training data. Ensemble models can also be constructed for deep learning sequence models, such as e.g. by training multiple models for the same task using different random seeds, and combining the outputs of such models, e.g. by averaging.

[0056] The term “deep learning sequence model” refers to a deep learning model that is configured to take as input encoded sequence data, and learn an internal representation (also called embedding) from which a property of interest (e.g. chromatin accessibility metric) can be predicted. A deep learning sequence model is typically a deep neural network. Multiple architectures can be used. For example, a transformer based deep learning model architecture can be used due to its computational efficiency in handling large input strings. Alternatively, a convolutional neural network can be used. Further, mixtures of convolutional layers and transformer layers can be used. These may be referred to as a “convformer” model. A transformer layer is a layer that processes input data using only attention mechanisms, as described in Vaswani et al. 2017. A convolutional layer is the main building block of a convolutional neural network, that applies a set of filters (also referred to as kernels) to input data. A model comprising convolutional layers may comprise a plurality of convolutional layers with one or more residual connections. A residual connection adds the input of a layer (or a downsampled version thereof) to the output of a subsequent layer. This may also be referred to as a shortcut connection. A model comprising convolution layers may comprise one or more dilated convolutions. Dilated convolutions are described in Yu & Koltun 2016. Dilated convolutions can aggregate multi-scale contextual information without losing resolution. This is particularly useful to enable base level predictions that take the context of the rest of the sequence into account. In embodiments, the sequence model does not comprise any pooling operations in or between any convolutional layer(s). This preserves the single-base-resolution of the models. A model as described herein may comprise one or more convolutional layers, each comprising a 1 D convolution followed by an activation function (e.g. a ReLU activation function). A model comprising one or more transformer layers may comprise one or more transformer decoder layers. This may be referred to as a transformer decoder stack. Each transformer decoder layer may comprise a multi-head self-attention mechanism followed by a feedforward network with e.g. ReLU activation. A model comprising one or more transformer layer may use learnable positional encodings added to the input of the first transformer layer. The one or more transformer layers may use a causal attention mask to ensure that each position only attends to previous positions. The one or more transformer layers may maintain the sequence-resolution throughout the transformer processing.

[0057] A deep learning sequence model may further comprise a regression head. The regression head may take as input embeddings produced by transformer and / or convolutional layers and produce a numerical output for each base of the input sequence. Alternatively, when base level resolution is not desired, the regression head may take as input embeddings produced by transformer and / or convolutional layers and produce a numerical output for a plurality of bases of the input sequence (or the entire input sequence). A deep learning model may comprise a classification head. The classification head may take as input embeddings produced by transformer and / or convolutional layers and produce a classification output (e.g. classification label or probability for each of one or more classes) for each base of the input sequence. Alternatively, when base level resolution is not desired, the classification head may take as input embeddings produced by transformer and / or convolutional layers and produce a classification output for a plurality of bases of the input sequence (or the entire input sequence). A deep learning sequence model may comprise a plurality of regression and / or classification heads, each trained to predict a different chromatin state metric. In embodiments, an entire deep learning sequence model is trained for prediction of a particular chromatin state metric. In other embodiments, one or more layers of the model are trained using a first training data set (e.g. comprising data for a plurality of sample types, such as e.g. different cell lines, cell types, tissues, organisms, etc.) and then further trained or included as part of a model comprising additional trainable layers (also referred to as fine tuning) using a second data set (e.g. comprising data for a single sample type, such as e.g. a single organism, cell type, tissue, disease status, etc.) The second training data set may be a subset of the first training data set. The pretrained layers may be frozen (i.e. parameters may be set after the first round of training) or fine-tuned in the second round of training.

[0058] An encoded sequence (also referred to as “encoded sequence data”) refers to a numerical representation of a sequence, for providing as input to a machine learning model as described herein. For the purpose of machine learning models that are not deep learning sequence models, encoded sequence data may comprise values for a plurality of features derived from the sequence data. For the purpose of a deep learning sequence model, the encoded sequence data preserves base specific and positional information, i.e. it includes information about the base present at each position in a sequence (and evidence for the presence of mC and hmC when using 6-base sequencing data) as well as the order of these bases in the sequence. An encoded sequence can be obtained from sequence information comprising genetic sequence information (i.e. A, C, T, G, N, at each position in a sequence) and epigenetic information (i.e. information indicating the presence / absence or proportion of C, mC or hmC at each position in a sequence). The latter may be set to a default value for all positions where the genetic information does not indicate the presence of a C. Thus, an encoded sequence as described herein may comprise encoded genetic sequence information and encoded epigenetic information. Encoded genetic sequence information can be obtained from genetic sequence information using any method known in the art, such as e.g. one- hot encoding. An example one-hot encoding scheme for genetic sequence information is shown on Fig. 7A. Encoded epigenetic information can be obtained from mC / hmC sequencing information comprising, for each position comprising a C in a sequence, a number of reads that have mC at the position, a number of reads that have hmC at the position, and a number of reads that have C at the position (or two of these and a total number of reads, from which the third value can be obtained). Encoded epigenetic information can comprise one or more of these numbers of reads, or normalised versions thereof (i.e. fractions of reads that have mC, hmC or C at the position, respectively). Encoded epigenetic information can comprise one or more statistics associated with an estimate of the fraction of copies of the genomic region in the sample that have a C, mC and / or hmC, obtained from the mC / hmC sequencing information. For example, the statistics may be selected from: an estimate of the statistically most likely fraction of copies of the genomic region in the sample that have a C, mC and / or hmC, the boundaries of a confidence interval around an estimate of the statistically most likely fraction of copies of the genomic region in the sample that have a C, mC and / or hmC, and a mean and standard deviation (or variance) of an estimated distribution of the fraction of copies of the genomic region in the sample that have a C, mC and / or hmC. For example, this can be a posterior distribution for the respective fraction parameter. The statistics may be estimated using a model based on a Dirichlet process, which can capture uncertainty in multinomial count data. A confidence interval may be a confidence interval at a predetermined level of statistical confidence, such as e.g. 90%, 95%, 96%, or 97%, 98%.

[0059] An encoded sequence may also comprise, for each base in the sequence, information indicating whether the base is on the sense or antisense strand. This may be encoded as a single bit of information (i.e. a single binary value) for each base in the sequence. An encoded sequence for a genomic region may comprise encoded sequence information for the sense strand and encoded sequence information for the antisense strand. Thus, the total length of the encoded sequence for a genomic region may be equal to 2*L*E, where L is the length of the genomic region, and E is the length of the encoding for a single base. For example, in the scheme illustrated on Fig. 7A, E is 4, for the schemes illustrated on Fig. 7B, E is 7, and for the schemes illustrated on Fig. 7C, E is 10. Each of these schemes can be supplemented with strand information for each base, leading respectively to schemes that have E=5 on Fig. 7A, E=8 on Fig. 7B, and E=11 on Fig. 7C.

[0060] Figure 1A is a flowchart illustrating a method of predicting one or more chromatin state metrics according to a general embodiment of the disclosure.

[0061] At step 110, sequence data (also referred to herein as “sequence information”) is obtained for one or more samples. The term “sample” as used herein refers to a sample comprising genomic DNA or material from which genomic DNA can be obtained, such as e.g. cells, tissues etc, as well as samples comprising naturally fragmented genomic DNA such as cell free DNA (including circulating tumour DNA). For example, the sample may be a blood sample orsample derived therefrom (e.g. purified peripheral blood mononuclear cells, plasma), a cerebrospinal fluid sample or sample derived therefrom, a biopsy sample (e.g. a tissue biopsy sample or tumour sample), or a sample of cells (e.g. primary cells, which may be purified and / or cultured, cell lines, etc). The sample may be a sample from a subject, typically a eukaryote organism, such as a vertebrate, mammalian and / or a human subject. Similarly, the sample may be a sample of eukaryote cells, such as vertebrate or mammalian cells (e.g. mouse, rat, or human cells). The sample may be a sample from a human subject or model animal. The sample may be a sample of human or model animal cells.

[0062] The sequence data comprise epigenetic information (also referred to herein as “epigenetic data”) about the sample, i.e. information indicative of the presence of methylated cytosine and hydroxymethylated cytosine at a plurality of genomic positions. The sequence data further comprise genetic information about the sample, i.e. information indicative of the presence of each of the 4 genetic bases at a plurality of genomic positions. The genetic and epigenetic information may have advantageously been obtained from the same assay, such as e.g. a sequencing assay in which genetic and epigenetic bases can be called on the same read. The sequence data may comprise a plurality of sequence reads (e.g. in the form of a SAM or BAM file per sample) or information derived therefrom (such as e.g. information identifying, for each read overlapping a genetic or epigenetic locus, whetherthe read contains a variant or reference allele, a modified or unmodified base). According to any embodiment, the sequence data used for prediction and / or training may have been obtained using a sequencing technology from which epigenetic and genetic bases can be called on the same read. Thus, in any embodiment, the sequence data can comprise or consist of reads in 6-letters code. Reads in 6 letter code refers to reads in which each base is specified as A, T, C, G, mC or hMC (or N, indicating that the base could not be confidently called).

[0063] At step 120, encoded sequence data is obtained from the sequence data obtained at step 110. The encoded sequence data includes one or both of: (i) for each position of the sequence, encoded genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine, and (ii) values of one or more features derived from data including, for each position of the sequence, genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine, and the one or more features including at least one feature indicative of the presence of hydroxymethylated cytosine.

[0064] At step 130, the encoded sequence data is used to predict, for each base or set of bases of the genomic region, the value of a chromatin state metric, using a machine learning model trained to: take as input the encoded sequence data for a genomic sequence, and produce as output one or more chromatin state metrics for each base or set of bases of the genomic sequence. Multiple chromatin state metrics may be predicted using one model, or a plurality of models may be used to predict respective chromatin state metrics. Predictions may be made in a base specific manner (i.e. one prediction per base), for example using a deep learning sequence model as described herein. These models may take as input encoded sequence data including for each position of the sequence, encoded genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine. Alternatively, predictions may be made for a plurality of bases in the genomic region (this may also be done using a deep learning sequence model as described herein), or for the region as a whole. The latter may be done using a machine learning model that takes as input encoded sequence data comprising values of one or more features derived from data including, for each position of the sequence, genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine, and the one or more features including at least one feature indicative of the presence of hydroxymethylated cytosine. The values of said features may be summarised values over the whole genomic region for which a prediction is being made.

[0065] At step 140, the results of any one or more of the preceding steps can optionally be provided to a user (e.g. through a user interface) or data store.

[0066] In embodiments comprising predicting enhancer status, step 130 comprises predicting, for each genomic region comprising an enhancer and using the sequence data associated with the enhancer, the status of the enhancer, wherein the predicting is performed using a machine learning classifier trained to classify an enhancer between a plurality of classes each associated with a different enhancer status selected from active, repressed and poised, using as input sequence data for an enhancer or one or more features derived therefrom including at least one feature indicative of the presence of hydroxy methylated cytosine. The machine learning model may have been trained using training sequence data and enhancer status from the same organism, tissue and / or cell type or cell line as that of the sample for which enhancer status is predicted. The machine learning classifier may have been trained as explained further below by reference to Figure 2. The classifier can be a binary classifier or a multiclass classifier. The classifier can be configured to classify an enhancer between at least three classes comprising a class associated with active enhancers, a class associated with poised enhancers, and a class associated with repressed enhancers. In embodiments, the classifier is trained to classify a genomic region comprising an enhancer between a plurality of classes using as input one or more features derived from sequence data for an enhancer including at least one feature indicative of the presence of hydroxymethylated cytosine, and the method comprised determining at step 120 the value of said one or more features for the one or more enhancers. The one or more features may include one or more of: a feature indicative of the presence of hydroxy methylated cytosines in the genomic region, a feature indicative of the presence of methylated cytosines in the genomic region, a feature indicative of the presence of cytosines that are either methylated or hydroxymethylated in the genomic region, a feature indicative of the length of the genomic region, and a feature indicative of the number of CpGs in the genomic region. The epigenetic data may comprise sequence reads. In such embodiments, a feature indicative of the presence of hydroxy methylated cytosines in the genomic region can be selected from: (i) a summarised value, over one or more CpGs in the enhancer, of the fraction of reads indicative of the presence of a hydroxymethylated cytosine at the CpG; (ii) a density of hydroxymethylated cytosines in the genomic region (this can be calculated as the mean hmC*count / length, where mean hmC is the mean fraction of reads indicative of the presence of a hydroxy methylated cytosine over the CpGs in the genomic region, count is the number of CpGs in the enhancer, and length is the length of the genomic region (e.g. in base pairs); (iii) a standard deviation of the fraction of reads indicative of the presence of a hydroxymethylated cytosine at each CpG in the genomic region; and (iv) an entropy of the hmC read counts over the enhancer (this can be calculated as Sfc -Pfc(log(pfc)) where pk is a summarised hydroxymethylation fraction calculated for a bin k of the enhancer, where the genomic region of divided into bins of equal length). Similarly, a feature indicative of the presence of methylated cytosines in the enhancer can be selected from: (i) a summarised value, over one or more CpGs in the genomic region, of the fraction of reads indicative of the presence of a methylated cytosine at the CpG; (ii) a density of methylated cytosines in the enhancer (this can be calculated as the mean mC*count / length, where mean mC is the mean fraction of reads indicative of the presence of a methylated cytosine over the CpGs in the genomic region, count is the number of CpGs in the genomic region, and length is the length of the genomic region (e.g. in base pairs); (iii) a standard deviation of the fraction of reads indicative of the presence of a methylated cytosine at each CpG in the genomic region; and (iv) an entropy of the mC read counts over the genomic region (this can be calculated asSfc -Pfc(log(pfc)) where pk is a summarised methylation fraction calculated for a bin k of the genomic region, where the enhancer of divided into bins of equal length). Similarly, a feature indicative of the presence of modified cytosines in the genomic region can be selected from: (i) a summarised value, over one or more CpGs in the genomic region, of the fraction of reads indicative of the presence of a modified cytosine at the CpG; (ii) a density of modified cytosines in the genomic region (this can be calculated as the mean modC*count / length, where mean modC is the mean fraction of reads indicative of the presence of a modified cytosine over the CpGs in the genomic region, count is the number of CpGs in the genomic region, and length is the length of the genomic region (e.g. in base pairs); (iii) a standard deviation of the fraction of reads indicative of the presence of a modified cytosine at each CpG in the genomic region; and (iv) an entropy of the modC read counts over the genomic region (this can be calculated as £fc-pfc(log(pfc)) where pk is a summarised modified C fraction calculated for a bin k of the region, where the region of divided into bins of equal length).

[0067] In embodiments, the one or more features include a feature indicative of the presence of hydroxy methylated cytosines in the sequence data associated with the genomic region and a feature indicative of the presence of methylated cytosines in the sequence data associated with the genomic region. In embodiments, the epigenetic data comprises sequence reads, the feature indicative of the presence of hydroxymethylated cytosines in the sequence data associated with the genomic region comprising an enhancer is a summarised value, over one or more CpGs in the region, of the fraction of reads indicative of the presence of a hydroxymethylated cytosine at the CpG; and the feature indicative of the presence of methylated cytosines in the sequence data associated with the genomic region is a summarised value, over one or more CpGs in the region, of the fraction of reads indicative of the presence of a methylated cytosine at the CpG. A summarised value can be a mean value. The machine learning classifier may have been trained using training data comprising, for each of a plurality of genomic region comprising an enhancer: sequence data associated with the enhancer obtained from one or more samples, the sequence data comprising epigenetic data indicative of the presence of one or more epigenetic bases at one or more positions in the genomic region, the one or more genetic bases including methylated cytosine and hydroxymethylated cytosine; and status data derived from histone modification data, such as derived from ChiP-seq data. The machine learning classifier may be a random forest classifier, a gradient boosted tree classifier, a neural network classifier, K-nearest neighbour classifier, or a support vector machine classifier.

[0068] Also described herein is a computer-implemented method of predicting the status of one or more enhancers in a sample, the method comprising: receiving sequence data associated with the one or more enhancers obtained from the sample, the sequence data comprising epigenetic data indicative of the presence of one or more epigenetic bases at one or more positions in the genome, the one or more genetic bases including methylated cytosine and hydroxymethylated cytosine; and predicting, for each enhancer of the one or more enhancers and using the sequence data associated with the enhancer, the status of the enhancer, wherein the predicting is performed using a machine learning classifier trained to classify an enhancer between a plurality of classes each associated with a different enhancer status selected from active, repressed and poised, using as input sequence data for an enhancer or one or more features derived therefrom including at least one feature indicative of the presence of hydroxymethylated cytosine. In embodiments, the classifier is a binary classifier or a multiclass classifier. In embodiments, the classifier is configured to classify an enhancer between at least three classes comprising a class associated with active enhancers, a class associated with poised enhancers, and a class associated with repressed enhancers. In embodiments, the classifier is trained to classify an enhancer between a plurality of classes using as input one or more features derived from sequence data for an enhancer including at least one feature indicative of the presence of hydroxymethylated cytosine, and the method comprises determining the value of said one or more features for the one or more enhancers. The one or more features may include one or more of: a feature indicative of the presence of hydroxymethylated cytosines in the enhancer, a feature indicative of the presence of methylated cytosines in the enhancer, a feature indicative of the presence of cytosines that are either methylated or hydroxymethylated in the enhancer, a feature indicative of the length of the enhancer, and a feature indicative of the number of CpGs in the enhancer. The one or more features may include a feature indicative of the presence of hydroxymethylated cytosines in the sequence data associated with the enhancer and a feature indicative of the presence of methylated cytosines in the sequence data associated with the enhancer. The epigenetic data may comprise sequence reads, and the feature indicative of the presence of hydroxymethylated cytosines in the sequence data associated with the enhancer may be a summarised value, over one or more CpGs in the region, of the fraction of reads indicative of the presence of a hydroxymethylated cytosine at the CpG. The epigenetic data may comprise sequence reads, and the feature indicative of the presence of methylated cytosines in the sequence data associated with the enhancer may be a summarised value, over one or more CpGs in the region, of the fraction of reads indicative of the presence of a methylated cytosine at the CpG. A summarised value may be a mean value. In embodiments, the machine learning classifier has been trained using training data comprising, for each of a plurality of enhancers: sequence data associated with the enhancer obtained from one or more samples, the sequence data comprising epigenetic data indicative of the presence of one or more epigenetic bases at one or more positions in the enhancer, the one or more genetic bases including methylated cytosine and hydroxymethylated cytosine; and status data derived from histone modification data, optionally derived from ChiP-seq data. In embodiments, the machine learning classifier is a random forest classifier, a gradient boosted tree classifier, a neural network classifier, K-nearest neighbour classifier, or a support vector machine classifier. Also described herein is a computer-implemented method of training a machine learning model for predicting the status of one or more enhancers in a sample, the method comprising: receiving a training data set comprising (i) sequence data associated with a plurality of enhancers, the sequence data comprising epigenetic data indicative of the presence of one or more epigenetic bases at one or more positions in the enhancers, the one or more genetic bases including methylated cytosine and hydroxymethylated cytosine, and (ii) status data comprising for each of the plurality of enhancer, a status label derived from histone modification data associated with the enhancer; and training a machine learning classifier, using said training data, to take as input sequence data for an enhancer or one or more features derived therefrom including at least one feature indicative of the presence of hydroxy methylated cytosine, and produce as output a classification between a plurality of classes each associated with a different enhancer status selected from active, repressed and poised. Such methods may have any one or more of the features described above in relation to a method of predicting the status of one or more enhancers in a sample.

[0069] In embodiments of any aspect described herein, the sequence data may have been obtained using a sequencing technology from which epigenetic and genetic bases can be called on the same read. In embodiments of any aspect described herein, the sequence data may comprise or consist of reads in 6- letters code. The methods described herein find application in a variety of contexts. For example, the methods described herein can be used to predict chromatin status (e.g. chromatin accessibility and / or enhancer status) in a genetic assay. A genetic assay is an assay that determines genetic information about a subject. Genetic assays are typically used to inform the likelihood of a subject developing a disease, to characterise a subject’s disease, to determine a prognostic or treatment response for a subject, etc. In embodiments, a genetic assay is an assay that selectively determines the presence of genetic variants at a one or more predetermined loci in a sample from a subject (e.g. mutations in one or more predetermined genes). This may also be referred to as a targeted assay, gene panel assay or simply panel assay. The genetic variants may be somatic variants or germline variants. The predetermined loci may be selected as loci that are relevant to the likelihood of developing a disease, to the prognostic associated with a disease or disorder, to the likelihood of response to a particular therapy, or to the characterisation of a disease (e.g. to characterise the subject as having a disease of a particular subtype or severity). Thus, the genetic assay may be used to obtain a polygenic risk score for a particular disease or disorder, based on the genotype of a subject at each of the predetermined loci in the genetic assay. A polygenic risk score is a score that indicates the risk of developing a particular disease or disorder (which can refer to a disease or a particular subtype or severity of a disease) based on the genotype of a subject at a plurality of loci. In other words, a polygenic risk score is a score that is calculated based on the determined genotype of a subject at a set of predetermined genetic loci, and which is indicative of the probability or odds ratio of a subject having a particular phenotype (such as e.g. developing a disease, acquiring resistance to a therapy, having a poor prognosis, etc.). The disease or disorder may be referred to as “complex disease”, by contrast with singlegene diseases which can be associated with a single gene (i.e. having a genetic variant at the particular gene causes the disease). Complex diseases are influenced by multiple genes and environmental factors such that the genotype at multiple genes, optionally in combination with one or more additional factors such as environmental factors, enables the prediction of a likelihood of developing the disease. According to the present disclosure, it is possible to accurately predict chromatin state (which is typically closely related to gene expression but alterations of which have also been associated with multiple diseases including but not limited to e.g. aging, cancer, etc. - see e.g. Corces et al. 2018) from combined genetic and epigenetic (i.e. 6-base sequence data) information. As a result, it is possible to obtain insights obtainable by both a genetic assay (e.g. a polygenic risk score) and insights obtainable by chromatin state assays (or even new scores combining both genetic and chromatin structure alteration related effects) using a single source of sequence data that contains genetic and epigenetic information. In other words, according to the present disclosure, a genetic assay can be used to provide information about chromatin state, such as e.g. chromatin accessibility and / or enhancer status.

[0070] Further, the methods of the present disclosure find use in the screening of one or more perturbations for an effect on one or more of: chromatin accessibility, histone modification, enhancer status, cytosine methylation / hydroxymethylation. Indeed, the methods described herein provide an opportunity to characterise perturbations in relation to their effect on the genome, the epigenome and chromatin state using data from a single assay.

[0071] Indeed, the improved prediction accuracy provided by the methods described herein are particularly significant in the context of clinical assays and perturbation (e.g. drug) screening. In the context of clinical assay, any improvement in the accuracy of a prediction has meaningful implications in terms of the likelihood of patients being assigned to the correct clinical pathways and / or the likelihood of patients needing to undergo additional tests to confirm a clinical insight of the assay. In the context of drug screening, improved prediction accuracy (and especially in the context of the ability to obtain multiple types of information characterising the state of the perturbed system from a single assay) enables screening of larger sets of perturbations with lower risk of missing relevant candidates or spending additional resources on irrelevant candidates. Even small difference in accuracy can make large practical differences at when screening is performed at scale. Additionally, the methods described herein are particularly useful in the context of analysing samples from which chromatin state data cannot easily be obtained or where sample availability is limited, such as e.g. cell free DNA samples. Indeed, cfDNA samples such as those obtained from liquid biopsies are an increasingly common and convenient tool to diagnose, characterise and / or monitor diseases such as cancer. The present disclosure provides methods that augment the possibilities associated with analysis of such samples by enabling the quantification of chromatin state such as chromatin accessibility and enhancer status without the need to perform additional assays that may either not be possible to perform on this type of material or for which sufficient amounts of material is not available.

[0072] Further, the methods described herein have improved accuracy for prediction of chromatin state (e.g. accessibility and / or histone modifications or derived information such as enhancer status), compared to methods that make these predictions solely from genetic sequence data (i.e. 4-base data). As a result, it is possible to train machine learning models to make these predictions with similar or higher accuracy using smaller amounts of training data and / or smaller models, leading to improved computational efficiency in the training (and also deployment, in the case of smaller models) of the models.

[0073] Figure 1 B is a flowchart illustrating methods of performing a genetic test and / or screening one or more perturbations according to general embodiments of the disclosure.

[0074] At optional step 1100, sequence data is obtained from a sample, the sequence data comprising genetic and epigenetic information, including 5mC and optionally 5hmC information. Step 1100 may comprise analysing the sample using a sequencing protocol that identifies genetic and epigenetic bases (including mC and hmC) on the same reads. At step 1200, a method as described in relation to Fig. 1A is performed using the sequence data obtained at step 1 100. This may use sequence data that has been acquired using a genome-wide assay or a targeted assay that obtains sequence data about a selected set of genes. The selected set of genes may be disease associated genes. For example, in the context of a genetic test for cancer diagnosis or prognosis, genes (or specific genetic loci within genes) that are known to be associated with the risk of a subject developing cancer, the severity and / or drug response of a subject with cancer (e.g. mutations in tumour suppressor or tumour driver genes, mutations known to be associated with drug resistance or sensitivity, etc.) may be selected. At step 1300, the predictions obtained at step 1200, and optionally the sequence data obtained at step 1100 are analysed to obtain one or more metrics characterizing the sample. Obtaining sequence data may comprise sequencing material in a sample or receiving previously acquired sequence data from a memory, database or user interface.

[0075] In embodiments where the method is used to perform a genetic test, the sample may be a sample from a subject who has been diagnosed as having or being likely to have a disease or disorder. In such embodiments, the one or more metrics may include e.g. metrics indicative of a prognostic, treatment response, or diagnosis (including identification of a disease subtype such as a molecular subtype). For example, the genetic information (e.g. presence of one or more mutations at respective genetic loci) obtained at step 1100 may be used step 1300 to obtain a polygenic risk score. Defining a polygenic risk score for a particular phenotype of interest using a set of genetic loci has been identified as associated with this phenotype is within the capability of the skilled person. The epigenetic information (optionally in combination with the genetic information) is used at step 1200, to obtain one or more predicted chromatin state metrics (such as e.g. chromatin accessibility and / or predicted enhancer status). The predicted chromatin state metrics may then be used at step z to obtain a metric indicative of a prognostic, treatment response, or diagnosis (including identification of a disease subtype such as a molecular subtype). Alternatively, a multiomic metric may be obtained which combines information about the genotype of the subject as specific loci and the predicted chromatin state associated with e.g. one or more genes or regions.

[0076] In embodiments where the method is used to determine the effect of a perturbation (such as e.g. for drug screening, genetic KO screening, exposure to one or more physico-chemical or metabolic stresses etc.), the sample may be a sample of cells that have been exposed to the perturbation or a sample previously obtained from a subject that has been exposed to the perturbation. The one or more metrics may include metrics that compare the predictions obtained at step 1200 to predictions obtained for a suitable control, in order to characterize the effect of the perturbation on chromatin states. Methods according to these embodiments may comprise performing the method for a plurality of samples exposed to respective perturbations, and selecting one or more of the plurality of perturbations for further screening based on the values of the one or more metrics obtained at step 1300. Further, methods according to this embodiment may further comprise exposing the one or more samples to the one or more perturbations.

[0077] At optional step 1400, the results of any one or more of the preceding steps (including e.g. a predicted chromatin state metric obtained at step 1200, metric characterising the sample obtained at step 1300, genetic and / or epigenetic information obtained at step 1100) are provided to a user (e.g. through a user interface) or data store.

[0078] Figure 2 is a flowchart illustrating a method of providing a trained model for predicting one or more chromatin state metrics, according to embodiments of the disclosure.

[0079] At step 210, training data is received comprising: (i) sequence data associated with a plurality of genomic regions, the sequence data comprising genetic data and epigenetic data indicative of the presence of one or more epigenetic bases at one or more positions in the genome, the one or more genetic bases including methylated cytosine and hydroxymethylated cytosine, the genetic data and epigenetic data obtained from a single assay; and (ii) measured values of corresponding one or more chromatin state metrics (ground truth chromatin state metrics). At step 220, the training sequence data is encoded as described herein. The encoded training sequence data includes one or both of: (i) for each position of the sequence, encoded genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine, and (ii) values of one or more features derived from data including, for each position of the sequence, genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxy methylated cytosine, and the one or more features including at least one feature indicative of the presence of hydroxymethylated cytosine. At step 230, the training data is used to training a machine learning model as described herein, to take as input encoded sequence data for a genomic sequence, and produce as output one or more chromatin state metrics for each base or set of bases of the genomic sequence.

[0080] At step 240, the results of any preceding step may be provided to a user, computing device or data store. These may include one or more of: one or more trained models, one or more parameters of a trained model, one or more training and / or performance statistics that may be obtained as part of the train step, etc.

[0081] The training dataset may comprise data obtained from a plurality of samples. The training data set may comprise sequence data obtained from a first plurality of samples, and chromatin state data obtained from a second plurality of samples. In other words, the samples from which sequence data and chromatin state data was obtained may be different, although they are advantageously from the same organism, tissue and / or cell type or cell line (i.e. same type of sample). They may also have been subject to the same treatment (e.g. they may all be control samples that have not been perturbed, or they may all be samples that have been perturbed in the same way). When different types of samples or perturbations are used, it may be advantageous for samples of the same type and / or exposed to the same perturbation to be treated as forming pairs of ground truth data for the purpose of training. For example, values of chromatin state metrics obtained in one cell type may be used as ground truth for predictions from sequence data from the same cell type, whereas values obtained in another cell type may be used as ground truth for predictions from sequence data from the same other cell type. The machine learning model may be trained using training data comprising sequence data and chromatin state data from the same organism, tissue and / or cell type or cell line as that of a sample for which chromatin state metrics are to be predicted (i.e, intended use of the machine learning model). For example, the machine learning model may be trained using human data and deployed (used for predictions) on human data. The plurality of genomic regions may comprise at last 1000 regions, at least 2000 regions, at least 3000 regions, at least 5000 regions, at least 6000 regions, at least 7000 regions, at least 8000 regions, at least 9000 regions or at least 10000 regions. As explained above, the machine learning model may be a random forest model, a gradient boosted tree model, or a neural network model. Each of these can be implemented as a regressor or a classifier, and is particularly well suited to making predictions from features rather than sequences (strings). In other embodiments, the machine learning model is a deep learning model trained to take as input sequence data and produce as output position specific (also referred to herein as base specific) values.

[0082] In embodiments comprising training a machine learning model for predicting the status of one or more enhancers in a sample, the method comprises at step 210 receiving a training data set comprising (i) sequence data associated with a plurality of enhancers, the sequence data comprising epigenetic data indicative of the presence of one or more epigenetic bases at one or more positions in the enhancers, the one or more genetic bases including methylated cytosine and hydroxymethylated cytosine, and (ii) status data comprising for each of the plurality of enhancer, a status label derived from histone modification data associated with the enhancer.

[0083] The training dataset may comprise data obtained from a plurality of samples. The training data set may comprise sequence data obtained from a first plurality of samples, and enhancer status data obtained from a second plurality of samples. In other words, the samples from which sequence data and status data was obtained may be different, although they are advantageously from the same organism, tissue and / or cell type or cell line (i.e. same type of sample). They may also have been subject to the same treatment (e.g. they may all be control samples that have not been perturbed, or they may all be samples that have been perturbed in the same way). When different types of samples or perturbations are used, it may be advantageous for samples of the same type and / or exposed to the same perturbation to be treated as forming pairs of ground truth data for the purpose of training. For example, status data obtained in one cell type are used as ground truth for predictions from sequence data from the same cell type, whereas status data obtained in another cell type are used as ground truth for predictions from sequence data from the same other cell type. The machine learning model may be trained using training data comprising sequence data and status data from the same organism, tissue and / or cell type or cell line as that of a sample for which enhancer status is to be predicted (i.e, intended use of the machine learning model). For example, the machine learning model may be trained using human data and deployed (used for predictions) on human data. The plurality of enhancers may comprise at last 10000 enhancers, at least 15000 enhancers, at least 20000 enhancers, at least 25000 enhancers, or at least 30000 enhancers.

[0084] In embodiments comprising training a machine learning model for predicting the status of one or more enhancers in a sample, the method comprises training a machine learning classifier, using said training data, to take as input sequence data for an enhancer or one or more features derived therefrom including at least one feature indicative of the presence of hydroxymethylated cytosine, and produce as output a classification between a plurality of classes each associated with a different enhancer status selected from active, repressed and poised. The machine learning model, its inputs and outputs and the training data may have any of the features described in relation to Figure 1A. The machine learning model may be a random forest classifier, a gradient boosted tree classifier, a neural network classifier, K-nearest neighbour classifier, or a support vector machine classifier. The classifier may be trained to classify an enhancer between a plurality of classes using as input one or more features derived from sequence data for an enhancer including at least one feature indicative of the presence of hydroxymethylated cytosine. In such embodiments, the method comprises determining at step 220 the value of said one or more features for the one or more enhancers in the training dataset. The machine learning classifier may have been trained using training data comprising, for each of a plurality of enhancers: sequence data associated with the enhancer obtained from one or more samples, the sequence data comprising epigenetic data indicative of the presence of one or more epigenetic bases at one or more positions in the enhancer, the one or more genetic bases including methylated cytosine and hydroxymethylated cytosine; and status data derived from histone modification data. The histone modification data may be ChiP-seq data. Thus, receiving training data may comprise receiving ChiP-seq data and processing the data to identify enhancer status for each of a plurality of enhancers. Alternatively, the received training data may already comprise enhancer genomic coordinates and previously determined status.

[0085] Figure 3 shows an embodiment of a system for implementing methods of the disclosure, such as e.g. providing a trained model, using a trained model to predict gene expression and / or enhancer status, performing a genetic test and / or screening one or more perturbations. The system comprises a computing device 1 , which comprises a processor 101 and computer readable memory 102. In the embodiment shown, the computing device 1 also comprises a user interface 103, which is illustrated as a screen but may include any other means of conveying information to a user such as e.g. through audible or visual signals. The computing device 1 is communicably connected, such as e.g. through a network, to one or more databases 2 storing sequence data from a sample or a cohort of samples, and / or to one or more sequence data acquisition means 4. In embodiments, the sequence data acquisition means 4 is configured to obtain sequence data from samples, in the form of DNA sequencing reads or RNA sequencing reads (where RNA sequencing reads may be obtained by sequencing of cDNA derived from RNA). Thus, the sequence data acquisition means may comprise a sequencing machine, such as e.g. an Illumina sequencer when using the duet multiomics solution from Biomodal (such as e.g. the duet multiomics solution +modC which performs 5-letter sequencing or the duet multiomics solution evoC which performs 6-letter sequencing), the PromethlON sequencer from Oxford Nanopore Technologies, or the Revio or Sequel long-read sequencers from Pacific Biosciences. The samples may be samples that have been processed to enable the sequencing of reads in 5 or 6 letter code (i.e. A, C, T, G, modified C and / or methylated C and hydroxy methylated C). The sequence data acquisition means 4 may be connected to the database 2. Connection between the sequence data acquisition means 4 and the database 2 may be through a wired or wireless connection. The sequence data may be raw data (e.g. reads) and / or processed versions thereof, such as e.g. variant calling files, counts files, etc. The one or more databases 2 may further store one or more of: one or more reference sequences, one or more gene expression data sets and / or enhancer status datasets (e.g. ChlP-seq datasets), parameters (such as e.g. parameters of a trained, parameters of a sequence data preprocessing methods, etc.), etc. The computing device may be a smartphone, tablet, personal computer or other computing device. The computing device is configured to implement a method as described herein. In alternative embodiments, the computing device 1 is configured to communicate with a remote computing device (not shown), which is itself configured to implement a method as described herein. In such cases, the remote computing device may also be configured to send the result of the method to the computing device. Further, the various steps of the methods described herein may be split between the computing device 1 and the remote computing device. The remote computing device may be a cloud computing device, a server node, etc. Any processing device known in the art may be used for this purpose. Communication between the computing device 1 and the remote computing device may be through a wired or wireless connection, and may occur over a local or public network 3 such as e.g. over the public internet.

[0086] The features disclosed in the foregoing description, or in the following claims, or in the accompanying drawings, expressed in their specific forms or in terms of a means for performing the disclosed function, or a method or process for obtaining the disclosed results, as appropriate, may, separately, or in any combination of such features, be utilised for realising the invention in diverse forms thereof.

[0087] While the invention has been described in conjunction with the exemplary embodiments described above, many equivalent modifications and variations will be apparent to those skilled in the art when given this disclosure. Accordingly, the exemplary embodiments of the invention set forth above are considered to be illustrative and not limiting. Various changes to the described embodiments may be made without departing from the spirit and scope of the invention. For the avoidance of any doubt, any theoretical explanations provided herein are provided for the purposes of improving the understanding of a reader. The inventors do not wish to be bound by any of these theoretical explanations.

[0088] Any section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0089] Throughout this specification, including the claims which follow, unless the context requires otherwise, the word “comprise” and “include”, and variations such as “comprises”, “comprising”, and “including” will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps.

[0090] It must be noted that, as used in the specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from “about” one particular value, and / or to “about” another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by the use of the antecedent “about,” it will be understood that the particular value forms another embodiment. The term “about” in relation to a numerical value is optional and means for example + / - 10%.

[0091] Examples

[0092] EXAMPLE 1 - Prediction of chromatin accessibility

[0093] ATAC-seq is a technique used to assess chromatin accessibility in a biological sample, which reflects the regions of the genome that are open and accessible to transcription factors and other regulatory proteins. Chromatin accessibility is important because it influences gene expression. Open chromatin regions are more likely to be transcribed, while closed regions are less accessible to the transcriptional machinery. When a chromatin region is more accessible, it is more likely to be transcribed into RNA, leading to higher gene expression. In this example, the inventors set out to design a set of features based on methylated CpGs assessed using a 6-letter sequencing method, and assess whether these could be used to predict ATAC-seq data.

[0094] Methods

[0095] Data - methylation. The 6-letters sequencing data used is available from the Gene Expression Omnibus (GEO) database under accession GSE251668

[0096] (www.ncbi.nlm.nih.gov / geo / query / acc.cgi?acc=GSE251668). The data was obtained using duet multiomics solution evoC - a new sequencing technology that simultaneously derives all four genetic bases without ambiguity in C or T calls alongside the modification status of Cytosines (5mC and 5hmC, i.e, 6-Letter data) in a single read from a single DNA molecule. The technology consists of pre-sequencing library prep and post-sequencing analysis pipeline, providing single-base resolution of genetics and epigenetics at high accuracy. The technology is described in Fullgrabe et al. 2023 and was implemented as described. We sequenced DNA from a mouse embryonic stem cell-line, ES-E14TG2A at high depth to obtain simultaneous reading of the genome and epigenome (both 5mC and 5hmC). The data was processed using a software developed in house as a python package to efficiently analyse the quant files and extract relevant information regarding the methylation status of each CpG site that we cover.

[0097] Data - ATAC-Seq. The ATAC-seq data was obtained from the Gene Expression Omnibus database under accession number GSE209527 (www.ncbi.nlm.nih.gov / geo / query / acc.cgi?acc=GSE209527). This contains ATAC-seq data for 6 samples of untreated mouse embryonic stem cells. The data was summarised by taking the mean accessibility over TSS regions.

[0098] Feature engineering. Features were created for each gene analysed based on genomic and methylation information (e.g. mean methylation, number of CpG) at different genomic annotations. Gene annotations from Gencode for GRCm38.p6 were used (GENCODE - Mouse Release M25, see www.gencodegenes.org / mouse / release_M25.html). A set of regions of interest were identified for each gene around the TSS (region around the transcription start site). The region around the TSS of each gene was split into sub-regions of equal predetermined lengths. For each region, the following features were calculated: (i) mean mC fraction; (ii) mean hmC fraction; (iii) CpG count in the region. A complete list of features and description is provided in Table 2.

[0099] For each region, the following features were calculated: (i) mean mC fraction; (ii) mean hmC fraction; (iii) CpG count in the region. Additional features were tested but not used in the results shown in the present examples, including mean modC fraction, density of mC, hmC or modC (e.g. mean mC*count / length; this captures information already captured by the combination of the corresponding mean, CpG count and region length (which is fixed) and therefore could be used instead of one or more of these), entropy of mC, hmC or modC (which could be used to measure the degree of variability of the methylation across a region), and the standard deviation of the mC fraction, hmC fraction and / or modC fraction. Entropy can be calculated by calculating a probability density function 'p' for mC, hmC or modC over the region, which is essentially a histogram representing the methylation fraction (mC fraction), hydroxymethylation fraction (hmC fraction) or modified C fration (modC fraction) in each of a plurality of bins over the region. The entropy is then obtained as the sum over k of -p_k*log(p_k), where p_k would represent the probability of bin k (i.e. entropy=fc-pfc(log(pfc)) where pkis the probability for bin k).

[0100] The mean 5mC fraction is the mean over a defined region of the fraction of reads at each CpG in the region that have a 5mC relative to all reads mapping to the C. The mean 5hmC fraction is the mean over a defined region of the fraction of reads at each CpG in the region that have a 5hmC relative to all reads mapping to the C. The number of CpGs is the number of CpG dinucleotides in the region. The length of the region is the length of the region in number of bases (bp). This is fixed for regions of fixed length (e.g. bins of fixed length in promoter regions, TSS) but is variable for regions like exons, introns and UTRs.

[0101] For each region, a mean modC fraction feature can also be calculated as the mean over a defined region of the fraction of reads at each CpG in the region that have a 5mC or a 5hmC (i.e. conflating 5mC and 5hmC) relative to all reads mapping to the C. This feature is termed mean modC and can be used e.g. in comparative models that do not differentiate between mC and hmC (or as an additional feature in models that do include hmC information). It is not used in the results shown.

[0102] Table 2. Features used and their description. Individual features of the model are the “features calculated in the region”. Regions are defined individually for each gene based on annotations from Gencode as explained above. The total model comprises 3 features calculated in each of 5 regions, i.e. 15 features per gene.

[0103] Annotations and missing features. The GENCODE annotation is made by merging the manual annotation produced by the Ensembl-Havana team (levels 1 and 2 annotations, level 1 being validated annotations and level 2 being manual annotations) and the Ensembl-genebuild automated gene annotation (level 3). Regions that were annotated as a gene by manual annotation but where the corresponding annotation was from automated annotation were discarded, resulting in the loss of a small number of genes. Soft-masked genes were not included (i.e. genes in soft-masked regions of the genome which are regions that are difficult to sequence accurately due to repetitive sequences). These represent very small numbers of genes. Only manual annotations from Gencode (Havana) were kept (level 1 and level 2 annotations in Gencode), with priority for level 1 annotations when available. If a region has 0 CpGs, the mean of the region is set to 0 for all of 5mC and 5hmC.

[0104] Model training. Once a list of features is obtained, the data comprising values for each of the epigenetic features for each of the genes observed in the ATAC-seq data and corresponding ground truth log(Accessibility) values (log(mean accessibility over TSS region), where TSS regions was defined as a region encompassing -1500 bp to +1500bp around the TSS) was passed to a machine learning algorithm to predict chromatin accessibility. Two types of architectures were used: a random forest (RF) regressor from scikit learn (scikit- learn.org / stable / modules / generated / sklearn. ensemble. RandomForestRegressor.html) and a gradient boosted tree model (XGBoost) model from PyPI (xgboost 2.0.3, pypi.org / project / xgboost / ). All results shown here are using the gradient boosted regressor from XGBoost. Best hyperparameters were identified through a grid search using the following parameters: (i) max_depth (max depth of tree): searched between 3 and 7, selected value of 5; (ii) n_estimators (number of trees): searched between 100 and 600, selected 500; (iii) Eta (learning rate, the step size at each iteration while moving toward a minimum of the loss function): searched between 0.01 and 0.05, selected 0.03; (iv) Subsample (fraction of the training data to be randomly sampled for building each tree) searched between 0.2 and 0.6, selected 0.5; (v) colsample_bytree (fraction of features to be randomly sampled for building each tree), searched between 0.8 and 1 , selected 0.9. Results are shown for models with the “selected” values of these parameters. All genes on chromosome 8 were used as a test set (468 genes), and the rest were used as a training set (10981 genes). Note that classifiers could have been used instead of regressors. Both of the above types of models have classifier equivalents (see e.g. scikit- learn.org / stable / modules / generated / sklearn. ensemble. RandomForestClassifier.html). Further, any machine learning model that is suitable for regression (linear or non-linear, although models that are able to capture non-linear relationships, such as RF models, XGboost and neural networks, are preferred) or classification could have been used. When classification is used, TSS regions can be classified between a plurality of categories (e.g. 4), corresponding to non-overlapping ranges of the log(accessibility)values observed in the training data, such as e.g. ranges corresponding to the first, second, third and fourth quartiles of the distribution of log(accessibility) in the training data set. Alternative approaches for selecting ranges can be used, such as e.g. by clustering log(accessibility) values to identify groups of genes and boundaries between them.

[0105] Results

[0106] Figure 4 shows the results obtained by training the XGBoost regression model for ATAC-seq prediction from the complete set of features in Table 2. The model obtained a R2=0.83 and Spearman rank correlation^.93. This excellent prediction performance confirms the relevance of the mC and hmC data to the physiological mechanisms thought to influence gene expression.

[0107] EXAMPLE 2 - Prediction of enhancer status

[0108] Enhancers are DNA sequences that can increase the transcription of genes, and are therefore of interest to predict gene expression. Enhancers tend to be classified in 3 categories:

[0109] Active: actively enhance gene expression - marked by the presence of H3K4me1 and H3K27ac modifications in flanking nucleosomes.

[0110] Primed (also referred to as poised): are in an in-between state that generally indicate that they are about to activate gene expression - marked by the presence of H3K4me1 , but not H3K27ac Repressed: do not enhance gene expression - marked by the presence of H3K27me3 or H3K9me2 / 3.

[0111] In analysing their mC and hmC data over a set of regions identified as enhancers, the present inventors identified signs that enhancers may have different patterns of methylation depending on their state. They therefore set out to test whether it would be possible to predict enhancer state from 6-letters sequencing data.

[0112] Methods

[0113] Data - enhancer locations and status. A list of enhancers in ES-E14 cells were obtained from the study in Cruz-Molina et al. (2017) (available as Table S1 of the paper). The list includes, for each enhancer (1 line per enhancer), chromosome position start and end, and status.

[0114] Data - methylation. See Example 1 . Features engineering. Two predictive features were designed: mean 5mC fraction and mean 5hmC fraction, calculated over the annotated enhancer region. The mean 5mC fraction and mean 5hmC fraction were calculated as explained in Example 1. Further, in order to assess the impact of including the 5hmC information, a comparative model was trained using exactly the same procedure but using only mC (i.e. using a mean mC fraction feature alone instead of the mean mC fraction and mean hmC fraction features).

[0115] A length feature was not included here because the primary purpose of the exploratory work was to representing the enhancers in the mc / hmc plane and see how it differs from modC alone, not necessarily building the strongest classifier. However, the enhancers did differ significantly in length, from a few hundred bp to a few thousand bp, with a mean at 1269 bp. It would therefore be possible to include a length feature when building a classifier, or one or more other features that capture the combination of the above fractions and length, such as e.g. density. Enhancers with 0 CpGs or 0 modified CpGs were not included.

[0116] Model training. The data comprising values for each of the two features above for each of the enhancer in the list of annotated enhancers, and corresponding ground truth status (active, poised, repressed) was passed to a machine learning algorithm to predict gene expression. A support vector machine classifier with a linear kernel (SVM from the Python Scikit-learn library, see scikit- learn.org / stable / modules / svm. html#svm-classification) was trained to perform the multiclass classification. Chromosome 8 was used as a training set, consisting of 1793 enhancers, over 33108 enhancers in total, and the remaining ones were used in the training set (i.e. 31 ,315 enhancers in the training set). A model using only the mean modC fraction was also trained in the same way. Performance was assessed in terms of the percentage of enhancers in the test set that were correctly classified. Note that other multiclass classification methods could also be used, such as e.g. RF classifiers, neural network classifiers, K-nearest neighbour classifiers and decision trees. Classifiers that are able to handle non-linear relationships, such as RF classifiers, SVM and neural network classifiers are likely to perform better.

[0117] Results

[0118] The inventors noticed that enhancers seem to have different patterns of methylation depending on their state. Specifically, the trend is as follows: (a) active enhancers have low mC and hmC levels, (b) primed enhancers moderate mC and high hmC levels, (c) repressed enhancers have high mC and low hmC levels. The inventors therefore trained a classifier model to predict the state of an enhancer based on two variables: mean 5mC fraction and mean 5hmC fraction of the enhancer.

[0119] Figure 5 shows the results obtained by training the SVM model for classification of enhancer state, illustrated in the mean hmC fraction vs mean mC fraction space (space of predictive features). The model was able to correctly predict 86% of the enhancer states. The model using modC alone correctly classified only 82% of the enhancers. The data show a complex picture of relationship between mC, hmC and enhancer status, where active enhancers have low me and hmC levels, primed enhancers moderate mC and high mC levels, and repressed enhancers have high mC and low hmC levels. To the best of the inventors’ knowledge, this is the first demonstration that enhancer status (which is normally evaluated at the level of histone modification patterns) can be predicted from DNA methylation data, and even more accurately from 6-letters sequencing data. EXAMPLE 3 - Prediction of chromatin state using a sequence model

[0120] In the examples above, the inventors used feature-based classification and regression algorithms to predict orthogonal types of omics data related to chromatin state, using epigenetic features derived from 6-letters sequencing data (illustrated on Figure 6, right hand side). However, the data used (6-letters sequencing data) inherently contains more information than just epigenetic data, since it contains genetic and epigenetic information on a single read. Building on the findings in Examples 1 -2 above, the present inventors set out to investigate whether 6-letters sequencing data could be used to train improved models for predicting chromatin state.

[0121] The designed an approach illustrated on Figure 6, left hand side, where instead of using the 6 letters sequencing data from the indicated regions to calculate features, 6 letters sequence data in the entire region from 1 kb upstream of the TSS to 1 kb downstream of the TSS, or in enhancer regions, is used to obtain an input sequence provided as an input to a sequence based deep learning model trained to predict a chromatin state metric from ATAC-seq or histone ChlP-seq. Note that other / longer regions of the genome could be used instead, and there is also no requirement to be limited to regions around TSS or enhancer regions. These models can be trained using 6 letters sequence data and any chromatin state data for a plurality of samples from the same species, which can be from the same or a plurality of different cell lines and / or tissues.

[0122] Methods

[0123] E14 mESC cell-line evoC methylation data. The inventors generated four technical replicate samples from the mouse embryonic stem cell-line (ES-E14TG2A), extracted DNA (80ng input), prepared libraries using a duet multiomics solution evoC kit, and sequenced with an S4 2x 150bp kit on an Illumina Novaseq 6000. They performed initial post-sequencing analysis (read resolution, trimming, alignment to GRCm38, quantification and QC) using the duet multiomics analysis pipeline to generate methylation quantification files for each of the 4 samples. Mean coverage of samples ranged from 35 to 50X. For quantification they pooled these technical replicates to realise a single biological sample with approximately 100x CpG coverage. Pooling was performed only because the samples are technical replicates of the same cell pool, and therefore treating them as independent samples would have amounted to pseudo-replication. Such high levels of coverage are not necessary to train the models. Indeed, the inventors have done some downsampling experiments showing that the performance of the models is robust to lower coverage of methylation data at least as low as 15-20x coverage.

[0124] Generating feature and target data, the present examples focus on the task of predicting functional genomic readouts in 2000bp windows centred at TSSs (Transcription Start Sites). The inventors generated focal regions using GENCODE gene annotations for the M25 release (www.gencodegenes.org / mouse / release_M25). Specifically, they downloaded the “basic annotation” and parsed strands separately to identify the transcription start sites as the starts of protein coding genes, while matching the strand orientation. After identifying the start sites they then got the region coordinates for 2000bp regions centred around the start sites (+ / - 1000bp). In total, they generated focal region coordinates for 21 ,674 TSSs. Note that the approach is not limited to TSS, and could be applied to any genomic region. TSS regions are particularly interesting because of their functional relevance (to gene expression) and because they are regions that are known to be likely to contain peaks in the ATAC-seq training data. However, other functionally interesting regions can be used, such as e.g. enhancer regions, or even large spans of genome (i.e. not discriminating for functional annotations) using windows of fixed sizes. For regions of variable sizes such as enhancer regions, fixed sized windows can be used that span the region e.g. when the region is larger than the input size of the model, or a window of fixed size comprising the region (e.g. at the start, end or in the middle of the fixed size window, i.e. centring the region of interest in the input window) can be used e.g. when the region of interest is smaller than the size of the input window expected by the model. Further, regions defined from the chromatin state data itself may be used, such as e.g. regions which were identified in orthogonal peak calling from the chromatin state training data (e.g. peaks in ATAC-seq, DNAse-seq or ChlP-seq) can be used for training. In such cases, fixed sized windows spanning or encompassing peak regions (e.g. regions of fixed size centred around called peaks) can be used. Such training data can be complemented with ‘background regions’ which are a GC-content-matched set of non-peak regions.

[0125] The inventors next extracted the GRCm38 reference genome sequence at each of these focal genomic regions and then performed a one-hot encoding of these DNA sequences to generate inputs compatible with our model architecture (A=[1 , 0,0,0]; C=[0, 1 ,0,0]; G=[0, 0,1 ,0]; T=[0, 0,0,1]; N=[0, 0,0,0]). Alongside the genomic one-hot encoding, we added some additional information. As a standard data augmentation approach, the inventors appended the reverse complement for each one-hot sequence to the input in order to capture the alternative strand state (such that we represent sense and anti-sense sequences) and added an additional strand element to the one-hot arrays, to indicate from which strand transcription initiates. The above encoding captures the full extent of feature preparation for genomic-only models (4-base encoding). This is similar to the encoding illustrated on Fig. 7A, except with an additional dimension to encode the strand, i.e. 5 bits per base, 2000 bases per input sequence on each strand, i.e. 4000 bases, leading to an encoded dimension of (5, 4000) for each focal genomic region.

[0126] For the genomic + evoC methylation models (6-base encoding) the inventors additionally overlaid: (i) the number of C calls, (ii) the number of 5mC calls and (iii) the number of 5hmC calls, which were quantified in the E14 mESC cell-line data, at all CpG locations. They used strand-specific methylation quantification in order to specifically overlay these counts to each of the sense and anti-sense components of the input sequences. This is similar to the encoding illustrated on Fig. 7B (Methylated C example 1), except with an additional dimension to encode the strand, i.e. 8 bits per base, 2000 bases per input sequence on each strand, i.e. 4000 bases, leading to an encoded dimension of (8, 4000) for each focal genomic region. Alternative encodings are possible, such as e.g. using fractional counts instead of counts, as illustrated on Fig. 7B, Methylated C example 2.

[0127] Further, alternative encodings may use statistics from model based estimates of fractions of modified and unmodified bases, to account for non-biological noise owing to sampling variation and methylation miscalls that may influence counts / fractional counts data. For example, a model based on a Dirichlet process can be used, which can capture uncertainty in multinomial count data. For example, one possible encoding would be to use the 95% confidence interval for the fraction of each possible state (C, 5mC, 5hmC), or the mean and variance of fractional estimates, as illustrated on Fig. 7C. The mean (also known as expectation) and variance of the Dirichlet can be calculated exactly using statistical formulae, while confidence intervals can be be estimated empirically using a sampling approach. For a Dirichlet distribution with X=(Xi,..., Xk)-Dir(a), where Xi,..., Xk are K fractions that sum to 1 and a =(ai,..., ak) are concentration parameters for each of the k categories, where ai>0, the mean for each category can be and the variance can be calculated as Va ■ This is similar to the encoding illustrated on Fig. 7C (Methylated C example 1), except with an additional dimension to encode the strand, i.e. 11 bits per base (4 for A, C, G, T, and 2 for each of C, mC, hC fractions, and 1 for the strand), 2000 bases per input sequence on each strand, i.e. 4000 bases, leading to an encoded dimension of (11 , 4000) for each focal genomic region.

[0128] In methylated C example 3 on Fig. 7C, the order of the encoded information after the 4 bits encoding the genetic information (N, A, C, G, T) is set out as: [C fraction left boundary of Cl, C fraction right boundary of Cl, mC fraction left boundary of Cl, mC fraction right boundary of Cl, hmC fraction left boundary of Cl, hmC fraction right boundary of Cl], where “Cl” stands or confidence interval. In methylated C example 4 on Fig. 7C, the order of the encoded information after the 4 bits encoding the genetic information (N, A, C, G, T) is set out as: [C fraction mean, C fraction sd, mC fraction mean, mC fraction sd, hmC fraction mean, hmC fraction sd], where “sd” stands for standard deviation. However, the exact order of each of these bits of information is not essential and any order can be used provided that the order is consistently used. The same applies to methylated C examples 1 and 2 on Fig. 7B. It is also irrelevant whether the epigenetic information is provided in the first few bits or the last few bits of the encoding, as long as the location of each bit of information is consistent. Similarly, the specific choice of order for the one-hot encoding of N, A, C, G, T is also arbitrary, and any other order could be used (e.g. where A is [0,1 ,0,0] and C is [1 , 0, 0, 0]). The orders presented are intuitive and therefore advantageously easier to interpret and to encode with minimum risk of confusion.

[0129] Following the data preparation outlined above, the final encoded dimensions for each focal genomic region were (5, 4000) and (8, 4000) for 4-base and 6-base encodings, respectively. An example of an encoded sequence according to the scheme on Fig. 7B (Methylated C example 1), as used in the present examples (except not showing the strand bit) is shown on Fig. 8.

[0130] In order to generate training target data the inventors identified and downloaded publicly available BigWig files for different primary assay types, namely ATACseq data and CHIPseq data from the E14 mESC cell line. They queried these BigWig files at each of the focal genomic regions to generate target tensors consisting of functional genomic signals across each region. For chromatin accessibility models they used the following dataset: www.ncbi.nlm.nih.gov / geo / query / acc.cgi?acc=GSE209527. For histone modification models they focussed on three primary histone marks, with data from EE14 hosted on ENCODE: H3K4me1 (www.encodeproject.org / files / ENCFF531JKE), H3K4me3 (www.encodeproject.org / files / ENCFF806NDV), and H3K27ac (www.encodeproject.org / files / ENCFF163SBS). The accessibility data was expressed in RPGC (reads per genomic content), so data that has been scaled to 1x genome size (i.e. scaling coverage to the sum of length of the standard chromosomes). This normalises reads to 1x sequencing depth, where sequencing depth is defined as: (mapped reads x fragment length) / effective genome size. Data on Fig. 10 is shown as log(RPGC). The ChlP-seq data was expressed in fold change over control (of read counts in the pull down vs control). Convolutional neural network model architecture. A schematic representation of a model used in the present example is shown on Fig. 9. The Convolutional neural network (CNN) models consist of three primary components: (i) an initial 1 -layer convolutional block, (ii) a 9-layer residual convolutional stack with exponential dilation, and (iii) a final convolution and cropping output layer. Each convolutional layer consists of a 1 D convolution followed by a ReLU activation function. The initial ‘stem’ block is used to upscale the initial channel dimensions (5 initial channels for 4-base encoding, 8 initial channels for 6-base encoding) to 64 channels (kernel_size=17, stride=1 , padding =“same”). Following this initial upscaling, the data flows through a 9-layer 1 D convolutional stack; all layers within this stack have 64 channels (kernel_size=17, stride=1 , padding =“same”). This stack has residual connections throughout and also uses exponential dilation (exponent of 2), such that the span of the dilation increases from 0 to 512 throughout the stack. The final layer first uses a 1 D convolution to downscale the channels from 64 to 1 (kernel_size=3, stride= 1 ) , followed by a central cropping layer to restrict the output dimensions prior to flattening for a final output target tensor with dimensions (1 , 2000). Note, we do not use any pooling operations throughout, in order to maintain the single-base-resolution of the models.

[0131] Model training. The inventors divided the full datasets (21 ,674 training examples) by chromosome into training (chromosomes: [1 ,2,3,4,5,10,1 1 ,12,13,14,15,16,17,18, X, Y]), validation (chromosomes = [6, 7]) and test datasets (chromosomes = [8, 9]). Models were trained with the following configuration: initial_learning_rate=3e-5, batch_size=64; max_epochs=100; patience=17; optimiser- ’Adam”. The ivnentors used a ‘reduce on plateau’ learning rate scheduler with parameters: factor=0.75; patience=3, this scheduler used the validation loss to indicate when performance had plateaued. We used the mean squared error loss function for training and tracked validation loss and validation R2throughout training. Optimal hyperparameters were chosen based on previous hyperparameter sweep experiments. For the chromatin accessibility work, they trained a total of 20 different models, corresponding to 10 random initialisations (with different random seeds) for each base encoding (4-base; 6-base).

[0132] The inventors evaluated the trained models on the test dataset using three primary metrics, R2score, Spearman’s correlation coefficient, and Pearson correlation coefficient. These were implemented across the target sequence to return single values for each prediction -target array pair.

[0133] All models were implemented in PyTorch (v2.3.0) and Lightning (v2.2.1) and were trained via Vertex Al using g2-standard-8 (NVIDIA_L4) machines.

[0134] Transformer and ConvFormer model architectures. Alternative architectures were also tested successfully (results not shown), including a transformer and ConvFormer architecture. The Transformer models consist of three primary components: (i) a layer normalization block, (ii) a configurable-depth transformer decoder stack, and (iii) a final prediction layer with GELU activation. Each transformer decoder layer consists of multi-head self-attention followed by a feedforward network with ReLU activation. The initial layer normalization block processes the input tensor of shape (batch_size, sequencejength, num_channels) to standardize the features. Following this normalization, learnable positional encodings are added to the input to provide position information to the transformer (shape: 1 , sequencejength, num_channels). The data then flows through the transformer decoder stack; all layers within this stack maintain the input dimension (sequencejength * num_channels) and use a causal attention mask to ensure each position only attends to previous positions. The number of attention heads is dynamically determined based on the number of channels (we default to matching the number of attention heads to the number of channels). The feedforward networks within each transformer layer have 32 hidden units, use ReLU activation, and apply dropout (probability=0.5). The final prediction layer first applies mean pooling across the tracks dimension, followed by a linear transformation to predict the target at each position, producing an output tensor with dimensions matching the input sequence length. A GELU activation is applied as the final non-linearity. Note that the model maintains the sequence-resolution throughout the transformer processing, with dimensionality reduction only occurring by condensing across the sequence “channels” to yield a final 1 D prediction output across the length of the sequence, using mean pooling in the final layer.

[0135] ConvFormer models simply take a combination of convolutional and transformer blocks as described in the above sections.

[0136] Results

[0137] Figure 10 shows the results of the work described above. Specifically, the inventors trained dilated residual convolutional neural networks to predict base-resolved chromatin accessibility across 2000bp sequences, centered around TSS regions. These models consume a combined encoding of genomic bases (mm10 reference), 5mC and 5hmC (duet evoC) as input sequences. The inventors used public ATACseq data from E14 mESC as the target output. During training the inventors held out chromosomes 6 and 7 for validation and chromosomes 8 and 9 as the test set. Globally, the data show that models only trained using the genomic sequence (4-base encoding, R2=0.54, Spearman’s correlation coefficient= 0.52, Pearson correlation coefficients.63) perform significantly worse than models which also had access to 5mC and 5hmC as part of the inputs (6-base encoding, R2=0.62, Spearman’s correlation coefficient= 0.65, Pearson correlation coefficients.74). The figure shows four example regions from the test dataset with ensembled predictions (10 random initialisations) from 4-base (blue) and 6-base (red) models, together with the ground truth data for these regions (black). Note that these performance metrics are not directly comparable to other prior art models that only use 4-base data as models should be evaluated on the same data (as is done here, i.e. the data show that the 6-base models outperform 4 base models significantly). This is a particularly significant difference in performance when considering that the models were trained and evaluated on data for the exact same cell line in the exact same conditions (i.e. data from the same samples). The difference in performance is expected to be significantly larger when using the models on different samples, as the models including the 6-base data would be much better able to provide predictions that are tailored to the sample (which is essentially impossible for the 4-base data which is strictly limited to capturing the effect of polymorphisms). For example, the 4-base models would by definition provide the same predictions for any two samples of the same cell line even if the two samples are exposed to significant perturbations. By contrast, the 6-base models would be able to provide predictions that are tailored to the sample (to the extent that the perturbations impact C modification). References

[0138] A number of publications are cited above in order to more fully describe and disclose the invention and the state of the art to which the invention pertains. Full citations for these references are provided below. The entirety of each of these references is incorporated herein.

[0139] The ENCODE Project Consortium., Moore, J.E., Purcaro, M.J. et al. Expanded encyclopaedias of DNA elements in the human and mouse genomes. Nature 583, 699-710 (2020).

[0140] Alda-Catalinas C, Bredikhin D, Hernando-Herraez I, Santos F et al. A Single-Cell Transcriptomics CRISPR-Activation Screen Identifies Epigenetic Regulators of the Zygotic Genome Activation Program. Cell Syst 2020 Jul 22;11 (1 ) :25-41 ,e9. PMID: 32634384

[0141] ENCODE Project Consortium. An integrated encyclopedia of DNA elements in the human genome. Nature. 2012 Sep 6;489(7414):57-74. doi: 10.1038 / nature11247. PMID: 22955616; PMCID: PMC3439153.

[0142] Lin S, Lin Y, Nery JR, Urich MA et al. Comparison of the transcriptional landscapes between human and mouse tissues. Proc Natl Acad Sci U S A 2014 Dec 2;1 11 (48):17224-9. PMID: 25413365

[0143] Kraushar ML, Krupp F, Harnett D, Turko P et al. Protein Synthesis in the Developing Neocortex at Near- Atomic Resolution Reveals Ebp1 -Mediated Neuronal Proteostasis at the 60S Tunnel Exit. Mol Cell 2021 Jan 21 ;81 (2):304-322.e16. PMID: 33357414

[0144] Schwalb B, Michel M, Zacher B, Fruhauf K, Demel C, Tresch A, Gagneur J, Cramer P. TT-seq maps the human transient transcriptome. Science. 2016 Jun 3;352(6290):1225-8. doi: 10.1126 / science.aad9841 . PMID: 27257258.

[0145] Cruz-Molina S, Respuela P, Tebartz C, Kolovos P, Nikolic M, Fueyo R, van Ijcken WFJ, Grosveld F, Frommolt P, Bazzi H, Rada-Iglesias A. PRC2 Facilitates the Regulatory Topology Required for Poised Enhancer Function during Pluripotent Stem Cell Differentiation. Cell Stem Cell. 2017 May 4;20(5):689- 705. e9. doi: 10.1016 / j.stem.2017.02.004. Epub 2017 Mar 9. PMID: 28285903.

[0146] Avsec Z, Agarwal V, Visentin D, Ledsam JR, Grabska-Barwinska A, Taylor KR, Assael Y, Jumper J, Kohli P, Kelley DR. Effective gene expression prediction from sequence by integrating long-range interactions. Nat Methods. 2021 Oct;18(10):1196-1203. doi: 10.1038 / s41592-021 -01252-x. Epub 2021 Oct 4. PMID: 34608324; PMCID: PMC8490152.

[0147] Lou, S., Lee, HM., Qin, H. et al. Whole-genome bisulfite sequencing of multiple individuals reveals complementary roles of promoter and gene body methylation in transcriptional regulation. Genome Biol 15, 408 (2014). doi.org / 10.1186 / s13059-014-0408-0

[0148] Pongor LS, Tlemsani C, Elloumi F, Arakawa Y, Jo U, Gross JM, Mosavarpour S, Varma S, Kollipara RK, Roper N, Teicher BA, Aladjem Ml, Reinhold W, Thomas A, Minna JD, Johnson JE, Pommier Y. Integrative epigenomic analyses of small cell lung cancer cells demonstrates the clinical translational relevance of gene body methylation. iScience. 2022 Oct 12;25(11):105338. doi: 10.1016 / j.isci.2022.105338. PMID: 36325065; PMCID: PMC9619308.

[0149] Villicana and Bell. Genetic impacts on DNA methylation: research findings and future perspectives. Genome Biology (2021) 22:127

[0150] Frommer, M. et al. A genomic sequencing protocol that yields a positive display of 5-methylcytosine residues in individual DNA strands. PNAS 89, 1827-1831 (1992). Vaisvila, R. et al. Enzymatic methyl sequencing detects DNA methylation at single-base resolution from picograms of DNA. Genome Res. 31 : 1280-1289 (2021).

[0151] WO 2022 / 023753 A1 - COMPOSITIONS AND METHODS FOR NUCLEIC ACID ANALYSIS, assignee: CAMBRIDGE EPIGENETIX LIMITED.

[0152] Simpson JT, Workman RE, Zuzarte PC, David M, Dursi LJ, Timp W. Detecting DNA cytosine methylation using nanopore sequencing. Nat Methods. 2017 Apr;14(4):407-410. doi: 10.1038 / nmeth.4184. Epub 2017 Feb 20. PMID: 28218898.

[0153] Pacific Biosciences 2022, “Measuring DNA methylation with 5-base HFi sequencing”, available at www.pacb.com / wp-content / uploads / application-brief-measuring-dna-methylation-with-5-base-hifi- sequencing.pdf

[0154] Fullgrabe, J., Gosal, W.S., Creed, P. et al. Simultaneous sequencing of genetic and epigenetic bases in DNA. Nat Biotechnol 41 , 1457-1464 (2023). doi.org / 10.1038 / s41587-022-01652-0

[0155] Takahashi H, Kato S, Murata M, Carninci P. CAGE (cap analysis of gene expression): a protocol for the detection of promoter and transcriptional networks. Methods Mol Biol. 2012;786:181-200. doi: 10.1007 / 978-1 -61779-292-2_11. PMID: 21938627; PMCID: PMC4094367.

[0156] Kundaje et al.2024. ChromBPNet: Deep learning models of base-resolution chromatin profiles reveal cis- regulatory syntax and regulatory variation. Slides available at docs, google. com / presentation / d / 1Ow6K8TYN40u7T3ODdo-JRCLuv5fUUacA / edit#slide=id.p1 and github.com / kundajelab / chrombpnet

[0157] Grandi, F.C., Modi, H., Kampman, L. etal. Chromatin accessibility profiling by ATAC-seq. Nat Protoc 17, 1518-1552 (2022). doi.org / 10.1038 / s41596-022-00692-9

[0158] Yongjing Liu, Liangyu Fu, Kerstin Kaufmann, Dijun Chen, Ming Chen, A practical guide for DNase-seq data analysis: from data management to common applications, Briefings in Bioinformatics, Volume 20, Issue 5, September 2019, Pages 1865-1877, doi.org / 10.1093 / bib / bby057

[0159] O'Geen H, Echipare L, Farnham PJ. Using ChlP-seq technology to generate high-resolution profiles of histone modifications. Methods Mol Biol. 2011 ;791 :265-86. doi: 10.1007 / 978-1-61779-316-5_20

[0160] Vaswani et al. 2017. Attention Is All You Need. arXiv:1706.03762

[0161] Yu & Koltun 2016. Multi-scale context aggregation by dilated convolutions. arXiv:1511 .07122v3

[0162] For standard molecular biology techniques, see Sambrook, J., Russel, D.W. Molecular Cloning, A Laboratory Manual. 3 ed. 2001 , Cold Spring Harbor, New York: Cold Spring Harbor Laboratory Press

Claims

Claims:

1. A computer-implemented method of predicting a chromatin state metric associated with a genomic region in a sample, the method comprising: receiving (110) sequence data associated with the genomic region obtained from the sample, the sequence data comprising genetic data and epigenetic data indicative of the presence of one or more epigenetic bases at one or more genomic positions, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine, the genetic data and epigenetic data obtained from a single assay; and predicting (130), for each base or set of bases of the genomic region and using the sequence data associated with the genomic region, a value of the chromatin state metric, wherein the predicting is performed using a machine learning model trained to: take as input encoded sequence data for a genomic sequence, the encoded sequence data including: (i) for each of the one or more positions of the genomic sequence, encoded genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine, or (ii) values of one or more features derived from data including, for each of the one or more positions of the genomic sequence, genetic data and epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine, and the one or more features including at least one feature indicative of the presence of hydroxymethylated cytosine; and produce as output one or more chromatin state metrics for each base or set of bases of the genomic sequence.

2. The method of claim 1 , wherein the one or more chromatin state metrics are selected from: metrics indicative of chromatin accessibility, and metrics indicative of histone modification.

3. The method of claim 1 or claim 2, wherein the one or more chromatin state metrics are metrics that are obtained by DNA sequencing of a sample that has been subject to a protocol to enrich accessible regions of chromatin or regions of chromatin that are associated with a histone that carries a specific modification, optionally wherein a chromatin state metric indicative of chromatin accessibility is a metric obtained by ATAC-seq or DNAse-seq, and / or a chromatin state metric indicative of histone modification is a metric obtained by modified histone ChlP-seq and / or wherein the histone modification is selected from: H3K4me1 , H3K4me3, and H3K27ac and / or wherein a chromatin state metric indicative of chromatin accessibility is a base specific number of reads per genomic content, and / or wherein a chromatin state metric indicative of histone modification is a fold change of a number of reads relative to a control condition.

4. The method of any preceding claim, wherein the machine learning model is a deep learning model trained to:take as input encoded sequence data for a genomic sequence, the encoded sequence data including, for each position of the genomic sequence, encoded genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine; and produce as output a chromatin state metric for each base of the genomic sequence.

5. The method of claim 4, wherein the encoded sequence data further comprises for each base in the sequence, information indicating whether the base is on the sense or antisense strand.

6. The method of claim 4 or claim 5, wherein the encoded sequence data comprises encoded genetic data and encoded epigenetic data for each position of each strand of the sense and antisense strand of the genomic sequence.

7. The method of any preceding claim, wherein the encoded sequence data includes, for each position of the sequence, genetic data encoded using a one-hot encoding scheme.

8. The method of any preceding claim, wherein the encoded sequence data comprises encoded epigenetic data that includes, for each position of the sequence that is a C: a. one or more of, optionally at least three of: a number of reads that have mC at the position, a number of reads that have hmC at the position, a number of reads that have C at the position, and a total number of reads at the position, each number of read obtained from a sequencing technology that distinguishes between C, mC and hmC; and / or b. one or more of, optionally at least two of: a fraction of reads that have mC at the position, a fraction of reads that have hmC at the position, and a fraction of reads that have C at the position, each fraction of read obtained from counts of reads obtained using a sequencing technology that distinguishes between C, mC and hmC; and / or c. one or more of, optionally all of: left and right boundaries of a confidence interval around an estimate of a fraction of reads that have mC at the position, left and right boundaries of a confidence interval around an estimate of a fraction of reads that have hmC at the position, and left and right boundaries of a confidence interval around an estimate of a fraction of reads that have C at the position, each confidence interval obtained from counts of reads obtained using a sequencing technology that distinguishes between C, mC and hmC, optionally wherein the confidence intervals are estimated using a sampling method assuming that the fractions follow a Dirichlet distribution; and / or d. one or more of, optionally all of: a statistical estimate of a fraction of reads that have mC at the position, a statistical estimate of a fraction of reads that have hmC at the position, and a statistical estimate of a fraction of reads that have C at the position, each statistical estimate obtained from counts of reads obtained using a sequencing technology that distinguishes between C, mC and hmC, optionally wherein the encoded epigenetic data further comprises a standard deviation or variance estimate associated with each of saidstatistical estimates and / or wherein the statistical estimates are means of respective categories of a Dirichlet distribution.

9. The method of any preceding claim, wherein the machine learning model is a deep learning sequence model, optionally a deep neural network.

10. The method of claim 9, wherein the machine learning model is a deep neural network comprising one or more transformer layers and / or one or more convolutional layers.11 . The method of claim 10, wherein the deep neural network comprises a plurality of convolutional layers with one or more residual connections.

12. The method of claim 10 or claim 11 , wherein the deep neural network comprises one or more convolutional layers each comprising a 1 D convolution followed by an activation function, and / or wherein the deep neural network comprises one or more convolutional layers each comprising a dilated convolution.

13. The method of any of claims 9 to 12, wherein the deep neural network comprises one or more convolutional layers and does not comprise any pooling operations in or between any convolutional layer(s).

14. The method of any preceding claim, wherein the machine learning model has been trained using training data comprising, for each of a plurality of training genomic regions: a. encoded sequence data including: (i) for each position of the sequence of the training genomic region, encoded genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine, or (ii) values of one or more features derived from data including, for each position of the sequence, genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine, and the one or more features including at least one feature indicative of the presence of hydroxymethylated cytosine; and b. measured values of the one or more chromatin state metrics.

15. The method of claim 14, wherein the plurality of training genomic regions comprise regions of a predetermined size comprising: a transcription start site, an enhancer sequence, or an ATAC- seq, DNAse-seq or histone modification ChlP-seq peak.

16. The method of claim 15, wherein the plurality of training genomic regions comprise background regions of the predetermined size that do not comprise: an ATAC-seq, DNAse-seq or histone modification ChlP-seq peak, optionally wherein said background regions are selected to include, for each region of a predetermined size comprising an ATAC-seq, DNAse-seq or histonemodification ChlP-seq peak, a background region of matched GC content that does not comprise an ATAC-seq, DNAse-seq or histone modification ChlP-seq peak.

17. The method of any preceding claim, wherein the sequence data comprises or consists of reads in 6-letters code and / or wherein the method further comprises: a. obtaining the encoded sequence data from the received sequence data; and / or b. obtaining the sequence data from a sample by processing the sample using a sequencing technology from which epigenetic and genetic bases can be called on the same read.

18. The method of any preceding claim, wherein the genomic region comprises an enhancer region, the chromatin state metric is a metric indicative of histone modification, and the metric indicative of histone modification is an enhancer status, optionally wherein the machine learning model is a machine learning classifier that takes as input values of one or more features derived from data including, for each position of the sequence, genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine, and the one or more features including at least one feature indicative of the presence of hydroxymethylated cytosine.

19. The method of claim 18, wherein the machine learning classifier is a binary classifier or a multiclass classifier, or wherein the machine learning classifier is configured to classify a genomic region comprising an enhancer between at least three classes comprising a class associated with active enhancers, a class associated with poised enhancers, and a class associated with repressed enhancers; and / or wherein the machine learning classifier is trained to classify an enhancer between a plurality of classes using as input one or more features derived from sequence data for a genomic region comprising an enhancer including at least one feature indicative of the presence of hydroxy methylated cytosine, and the method comprises determining the value of said one or more features for the one or more enhancers; and / or wherein the machine learning classifier is a random forest classifier, a gradient boosted tree classifier, a neural network classifier, K-nearest neighbour classifier, or a support vector machine classifier.

20. The method of claim 19, wherein the one or more features include: a feature indicative of the presence of hydroxymethylated cytosines in the sequence data associated with the genomic region and a feature indicative of the presence of methylated cytosines in the sequence data associated with the genomic region, optionally wherein the epigenetic data comprises sequence reads, the feature indicative of the presence of hydroxymethylated cytosines in the sequence data associated with the genomic region is a summarised value, over one or more CpGs in the region, of a fraction of reads indicative of the presence of a hydroxymethylated cytosine at the CpG; and the feature indicative of the presence of methylated cytosines in the sequence data associated with the genomic regionis a summarised value, over one or more CpGs in the region, of a fraction of reads indicative of the presence of a methylated cytosine at the CpG; optionally wherein a summarised value is a mean value.

21. The method of any of claims 18 to 20, wherein the machine learning classifier has been trained using training data comprising, for each of a plurality of genomic regions comprising an enhancer: sequence data associated with the genomic region obtained from one or more samples, the sequence data comprising epigenetic data indicative of the presence of one or more epigenetic bases at one or more positions in the enhancer, the one or more genetic bases including methylated cytosine and hydroxy methylated cytosine; and enhancer status data derived from histone modification data, optionally derived from ChiP-seq data.

22. The method of any of claims 1 to 3, wherein the genomic region comprises a transcription start site, the chromatin state metric is a summarised metric indicative of chromatin accessibility over the genomic region, and the machine learning model is a machine learning model that takes as input values of one or more features derived from data including, for each position of the sequence, genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine, and the one or more features including at least one feature indicative of the presence of hydroxymethylated cytosine.

23. The method of claim 22, wherein the machine learning is a regressor model; and / or wherein the machine learning model is a random forest regressor, a gradient boosted tree regressor, or a neural network, and / or wherein the one or more features include: a feature indicative of the presence of hydroxymethylated cytosines in the sequence data associated with the genomic region and a feature indicative of the presence of methylated cytosines in the sequence data associated with the genomic region, optionally wherein the epigenetic data comprises sequence reads, the feature indicative of the presence of hydroxymethylated cytosines in the sequence data associated with the genomic region is a summarised value, over one or more CpGs in the region, of a fraction of reads indicative of the presence of a hydroxymethylated cytosine at the CpG; and the feature indicative of the presence of methylated cytosines in the sequence data associated with the genomic region is a summarised value, over one or more CpGs in the region, of a fraction of reads indicative of the presence of a methylated cytosine at the CpG; optionally wherein a summarised value is a mean value; and / or wherein the machine learning classifier has been trained using training data comprising, for each of a plurality of genomic regions comprising a transcription start site: sequence data associated with the genomic region obtained from one or more samples, the sequence data comprising epigenetic data indicative of the presence of one or more epigenetic bases at one or more positions in the genomic region, the one or more genetic bases including methylated cytosine and hydroxy methylated cytosine; and measured chromatin status metrics derived from chromatin accessibility data, optionally derived from ATAC-seq or DNAse-seq.

24. A computer-implemented method of training a machine learning model for predicting a chromatin state metric associated with a genomic region in a sample, the method comprising: receiving (210) training data comprising: (i) sequence data associated with a plurality of genomic regions, the sequence data comprising genetic data and epigenetic data indicative of the presence of one or more epigenetic bases at one or more positions in the genome, the one or more genetic bases including methylated cytosine and hydroxymethylated cytosine, the genetic data and epigenetic data obtained from a single assay; and (ii) measured values of corresponding one or more chromatin state metrics; and training (230) a machine learning model, using said training data, to: take as input encoded sequence data for a genomic sequence, the encoded sequence data including: (i) for each position of the sequence, encoded genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxymethylated cytosine, or (ii) values of one or more features derived from data including, for each position of the sequence, genetic data and encoded epigenetic data indicative of the presence of one or more epigenetic bases at the position, the one or more epigenetic bases including methylated cytosine and hydroxy methylated cytosine, and the one or more features including at least one feature indicative of the presence of hydroxymethylated cytosine; and produce as output one or more chromatin state metrics for each base or set of bases of the genomic sequence.

25. The method of claim 24, wherein the method has the features of any of claims 1 to 23.

26. A method of performing a genetic test for a subject, the method comprising: obtaining (1100) sequence data associated with a sample previously obtained from a subject, the sequence data comprising genetic data indicative of the presence of one or more genetic bases at one or more positions in the genome and epigenetic data indicative of the presence of one or more epigenetic bases at one or more positions in the genome, the one or more genetic bases including methylated cytosine and hydroxymethylated cytosine; and performing (1200) the method of any of claims 1 to 23 to obtain one or more predicted chromatin state metrics, optionally wherein the method further comprises analysing (1300) the one or more predicted chromatin state metrics and / or genetic data by obtaining the value of one or more metrics characterising the sample, optionally wherein the one or more metrics are selected from a diagnostic or prognostic metric.

27. A system comprising at least one processor and a non-transitory computer readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform the method of any of claims 1 to 26.

28. A non-transitory computer readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform the method of any of claims 1 to 27.

Citation Information

Patent Citations

  • Compositions and methods for nucleic acid analysis

    WO2022023753A1

  • Cell type-specific prediction of 3D chromatin architecture

    WO2023168079A2

Cited By

  • Application of recombinant promoter in improvement of bovine muscle performance

    CN121687205A