Chromatin structure biomarkers

EP4724601A2Pending Publication Date: 2026-04-15ALTOS LABS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
ALTOS LABS INC
Filing Date
2024-06-11
Publication Date
2026-04-15

AI Technical Summary

Technical Problem

Current methods for analyzing gene expression data do not effectively capture the influence of chromatin structure on transcription, particularly in understanding age-related disorders and rejuvenation processes, as they fail to incorporate structural and functional causes of expression differences.

Method used

A data-driven approach using correlation length as a biomarker metric, calculated from transcript abundance data, to characterize chromatin states by measuring long-range correlations between gene expression, distinguishing between healthy, diseased, and rejuvenated states.

Benefits of technology

This approach allows for accurate classification of cellular states, identification of treatment effects on chromatin structure, and measurement of cellular age, providing a robust metric for distinguishing between natural and pathological aging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024033444_19122024_PF_FP_ABST
    Figure US2024033444_19122024_PF_FP_ABST
Patent Text Reader

Abstract

Described herein are methods of determining the chromatin state of a sample from transcript abundance data from the sample. The methods comprise receiving transcript abundance data for a plurality of transcripts from the sample; determining, for each of a plurality of genomic distance bins, a correlation metric between transcript abundance for pairs of transcripts that are separated by a genomic distance within the respective bin; and generating a value of a biomarker metric derived from the correlation metric for the plurality of bins, the biomarker metric characterising the range of genomic distance at which pairs of transcripts have correlated expression, wherein samples with different chromatin states have different values of the biomarker metric. Related methods, systems and products are also described.
Need to check novelty before this filing date? Find Prior Art

Description

CHROMATIN STRUCTURE BIOMARKERSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the priority of co-pending U.S. patent application titled “CHROMATIN STRUCTURE BIOMARKERS”, filed on June 12, 2023, and having application number 63 / 507,720. The subject matter of this related application is hereby incorporated herein by reference.FIELD OF THE INVENTION

[0002] present disclosure relates to methods for characterising samples using biomarkers indicative of chromatin structure derived from expression data. Related methods, systems and products are also described.BACKGROUND OF THE INVENTION

[0003] Regulation of multiple genes and other transcribe-able elements are known to be the most fundamental of all genome functions. Traditional methods like bulk RNA- seq analysis provide data that can be useful to group or classify genes (or samples) based on transcript abundance. This is typically accomplished by using feature extraction and dimension reduction algorithms like Partial Least Squares (see e.g. Boulesteix and Strimmer, 2007), Sliced Inverse Regression (see e.g. Li and Lon, 2008), Principal Component Analysis (PCA) (see e.g. Yeung and Ruzzo, 2001 ) and more recently using Word Embeddings (see e.g. Choy, Wong and Chan, 2018). These approaches can be useful for identifying diseased phenotypes or for identifying genes that are co-regulated and may have similar functions. Another important technique constructs a network of genes based on their expression patterns across multiple samples (see e.g. Van Dam et al., 2018). Once the network is constructed, it identifies groups of genes that are co-expressed and can help identify key regulatory genes, or pathways, that may drive observed expression patterns. These methods greatly improved our understanding of co-expressing sets of genes involved in a shared biological pathway or process.

[0004] Regulation of expression depends on a gene’s local chromatin environment (Li, Carey and Workman, 2007). Genes located within dense heterochromatic regions generally exhibit lower levels of expression compared to genes located in more open euchromatin regions. In this regard, the chromatin environment constrains transcription. Heterochromatin is a more condensed, transcriptionally attenuatedstructure that is less accessible to RNA polymerase and other transcription factors. Euchromatin is more open, more accessible, and more easily transcribed. Epigenetic modifications including DNA methylation and histone acetylation, play an essential role in gene regulation and the formation of functionally and structurally distinct regions of heterochromatin and euchromatin. The changes in local chromatin structure due to such modification can influence gene expression by regulating the accessibility of DNA to transcription factors (Venkatesh and Workman, 2015). On a relatively smaller scale, repositioning of nucleosomes near promoter regions can up-regulate some genes and down regulate others in response to stress or changes in a cell’s environment (Lai and Pugh, 2017). A much recent study analysed human lung cells and umbilical vein cells and found that ageing cells contained fewer nucleosomes, facilitating the movement of RNA polymerase II (Pol II), allowing it to travel faster which shows that structural changes to DNA can dramatically influence the transcription machinery to be less precise and more error-prone as it makes a copy of RNA (Debes et al. , 2019). Globally, breakdown of heterochromatin domains can cause widespread alterations in genome organization, activate genes that were previously silenced, and result in abnormal gene expression, leading to genome instability. These structural changes in the heterochromatin have been linked to aging and related genetic disorders such as Hutchinson-Gilford progeria syndrome (HGPS) and Werner syndrome (WRN) (Tsurumi and Li, 2012, Lee et al., 2020).

[0005] Nevertheless, the dependence of gene expression on chromatin structure or local environment is far from understood.SUMMARY

[0006] Given that chromatin structure, by definition, creates distinct regions of more open or more densely packed DNA along a chromosome, and that this local genome packing affect transcription, the present inventor hypothesised that gene expression will be spatially correlated given the relative location of genes within these distinct regions. The present inventors further hypothesised that the degree of correlation between neighbouring genes would depend on the similarity of their chromatin structure and accessibility. The present inventors further hypothesised that if this is indeed the case, then transcript abundance, as measured by RNA-seq, should be correlated along the chromosome, with genes in euchromatin domains being up-regulated relative to genes in heterochromatin domains. They further recognised that these euchromatinand heterochromatin domains need not be continuous along a chromosome. Rather the average lengths of these domains should be determined by the rate at which these correlations decay for genes with higher or lower expression respectively.

[0007] They further postulated that loss of heterochromatin associated with age related disorders cause changes to the structure of chromatin, which in turn should affect the scale of correlations and hence that this scale of correlation could serve as an effective biomarker for classifying diseased states from healthy states. The present inventors therefore set out to develop a consolidated, data-driven approach to validate the hypothesis that epigenetic modification resulting from aging or pre-aging conditions may alter chromatin structure which ultimately affects long-range correlations between gene expression. Using a top-down approach, they systematically explored the pairwise correlation of transcript abundance between distant genes, as a function of their proximity. This method was applied to various whole genome (bulk RNA) transcriptom ic datasets from human fibroblast to establish transcript-transcript correlation based on their abundance. They then utilized correlation length as a feature to group cells into different states (healthy vs. diseased), demonstrating that this metric can indeed be used as a biomarker of diseases, disorders and conditions that affect the structure of chromatin, such as age-related disorders. In particular, in the context of aging related disorder, they utilized correlation length as a feature to group cells from diseased I aged samples treated with a test agent (rejuvenated), samples that are not treated with the test agent (diseased / aged), and healthy samples, into different states, showing clear patterns with biological replicates of the same state clustering together and closer to biologically more similar samples (e.g. rejuvenated samples being closer to healthy samples). This demonstrated that this metric can indeed be used as a biomarker of treatment effects for diseases, disorders and conditions that affect the structure of chromatin, such as age-related disorders.

[0008] Thus, according to a first aspect, there is provided a method of determining the chromatin state of a sample comprising: receiving, by a processor, transcript abundance data for a plurality of transcripts from the sample; determining, by said processor, for each of a plurality of genomic distance bins, a correlation metric between transcript abundance for pairs of transcripts that are separated by a genomic distance within the respective bin; and generating, by said processor, a value of a biomarker metric derived from the correlation metric for the plurality of bins, the biomarker metric characterising the range of genomic distance at which pairs of transcripts havecorrelated expression, wherein samples with different chromatin states have different values of the biomarker metric.

[0009] As the skilled person understands, the complexity of the operations described herein (due at least to the complexity of analysing expression data on genome or chromosome scales, and the amount of data that is typically generated by sequencing transcriptom ic material) are such that they are beyond the reach of a mental activity. Thus, unless context indicates otherwise (e.g. where sample preparation or acquisition steps are described), all steps of the methods described herein are computer implemented.

[0010] Thus, also described herein according to the present aspect is a computer implemented method of determining the chromatin state of a sample, comprising: receiving transcript abundance data for a plurality of transcripts from the sample; determining, for each of a plurality of genomic distance bins, a correlation metric between transcript abundance for pairs of transcripts that are separated by a genomic distance within the respective bin; and generating a value of a biomarker metric derived from the correlation metric for the plurality of bins, the biomarker metric characterising the range of genomic distance at which pairs of transcripts have correlated expression, wherein samples with different chromatin states have different values of the biomarker metric.

[0011] The method of the present aspect may have one or more of the following features.

[0012] The transcript abundance data may comprise a normalised abundance for each of the plurality of transcripts. The normalised abundance for a transcript may be obtained by dividing the abundance of a transcript by the mean abundance for all transcripts from the sample thereby obtaining a scaled abundance. Instead or in addition to this, the normalised abundance for a transcript may be obtained by dividing the abundance or scaled abundance for the transcript by the sum of abundances or scaled abundances for all transcripts from the sample thereby obtaining a fractional abundance.

[0013] Each of the plurality of transcripts may be associated with a genomic locus in a reference genome, such as a human reference genome. The sample may be a human sample. The sample may be a sample comprising a plurality of cells, preferably human cells. The sample may be a blood or tissue sample, or a cell sample (e.g. cell line).

[0014] The plurality of transcripts may be high abundance transcripts, wherein high abundance transcripts are obtained by identifying, in transcript abundance data for a plurality of transcripts, a first group of transcripts and a second group of transcripts, the first group of transcripts having higher abundance in the sample than the second group of transcripts. The method may comprise: identifying, by said processor, using the transcript abundance data, a first group of transcripts and a second group of transcripts, the first group of transcripts having higher abundance in the sample than the second group of transcripts, and selecting, by said processor, the first group of transcripts before determining the correlation metric for the plurality of bins. Identifying first and second group of transcripts may be performed using a data driven method. The data driven method may comprise: a clustering method, and a curve characteristics-based algorithms applied to a data series comprising the ranked abundance of transcripts as a function of rank. The curve-characteristics based algorithm may be the kneedie algorithm. The clustering method may be selected from spectral clustering, k-means, k-medoid, and density-based spatial clustering of applications with noise (DBSCAN). The clustering method may be spectral clustering. The clustering method may be DBSCAN. The spectral clustering may be performed using a kernel selected from: Laplacian affinity, Radial Basis Function or nearest neighbour. The data driven method may comprise a first data driven method used for outlier detection (e.g. DBSCAN) and a second data driven method used for identifying transcripts in the first and second groups from the set of transcripts after outlier detection (e.g. spectral clustering). The first and second data driven methods may be selected from: a clustering method, and a curve characteristics-based algorithms applied to a data series comprising the ranked abundance of transcripts as a function of rank.

[0015] The genomic distance bins may be based on the chromosomal positions of the transcripts, where transcripts in a bin are associated with genes that are the same number of genes apart from each other. The genomic distance bins may each comprise pairs of transcripts associated with genes that are n genes apart on the same chromosome, where n is a natural number or range of natural numbers, where n is different for each bin. The genomic distance bins may comprise a first bin comprising pairs of transcripts associated with genes that n1 genes apart on the same chromosome, and a second bin comprising pairs of transcripts that are associated with genes that are n2 genes apart on the same chromosome. In embodiments, n1 is 0 and n2 is not 0. The genomic distance bins may comprise a first bin comprising pairs oftranscripts associated with genes that n1 genes apart on the same chromosome, and a second bin comprising pairs of transcripts that are associated with genes that are n2 genes apart on the same chromosome. n1 and n2 are non-overlapping natural numbers or ranges of natural numbers. For example, n1 may be 0 (i.e. the genes are direct neighbours on the chromosome) and n2 may be 1 (i.e. the genes are separated by 1 gene on the chromosome).

[0016] The correlation metric may be a non-linear correlation coefficient. The correlation metric may be a distance correlation. The correlation metric for a bin may be calculated aswhere:(x, y) are vectors of transcript abundance for all possible pairs of transcripts in the bin; andV is the empirical distance covariance for the transcripts in the bin. The empirical distance may be calculated using the following equation:where D are matrices of all centred Euclidian pairwise distances between transcript abundance for transcripts in the bin, centred by the row mean and column mean.

[0017] The biomarker metric derived from the correlation metric may be a value derived from the auto-correlation of the correlation metric. The autocorrelation of the correlation metric may be calculated as a function of a lag in the same units as the units of genomic distance associated with the bins. The autocorrelation of the correlation metric may be calculated as a function of a lag in number of genes apart on the chromosome. The biomarker metric derived from the correlation metric may be a value proportional to a lag at which the auto-correlation of the correlation metric falls below a predetermined confidence limit. The biomarker metric derived from the correlation metric may be a value proportional to a lag at which the auto-correlation of the correlation metric is no longer significantly different from zero at a chosen confidence level. A confidence limit or level may be selected between 90% and 99%. A confidence level or limit may be 90%, 91 %, 92%, 93%, 94%, 05%, 96%, 97%, 98%, or 99%. The correlation metric may be the lag at which the auto-correlation of the correlation metric falls below apredetermined confidence limit, multiplied by 1000. The correlation metric may be the lag at which the auto-correlation of the correlation metric is no longer significantly different from zero at a chosen confidence level, multiplied by 1000.

[0018] The plurality of transcripts may be located on the same chromosome. The plurality of transcripts may be located on a plurality of chromosomes, and the correlation metric and biomarker value may be determined separately for each chromosome of the plurality of chromosomes, or for one or more chromosomes of the plurality of chromosomes, using the transcripts located on the respective chromosome. The plurality of transcripts may be located on a plurality of chromosomes, and the method may comprise selecting, by said processor, a plurality of transcripts located on the same chromosome prior to the processor determining the correlation metric. Selecting, by said processor, a plurality of transcripts located on the same chromosome may comprise selecting a plurality of transcripts located anywhere on the chromosome, or selecting a plurality of transcripts located on a portion of the chromosome. The sample may be a human sample. The chromosome may be chromosome 6 or chromosome 7. In particular, the chromosome may be chromosome 6. The method may further comprise repeating the steps of selecting a plurality of transcripts located on the same chromosome for a further chromosome or a plurality of further chromosomes and determining the correlation metric and biomarker value for the further chromosome or plurality of further chromosomes. The further chromosomes or plurality of chromosomes may comprise chromosomes 6 and / or 7.

[0019] The transcript abundance data may have been obtained from whole transcriptome RNA sequence data. The transcript abundance data may comprise abundance data from each transcript that is detectable in the sample using a transcriptome wide transcriptom ic analysis technique, such as RNA sequencing. The biomarker metric may be generated from transcript abundance data from a plurality of transcripts located on the same chromosome, and the value of the biomarker metric may be scaled by the number of transcripts on the chromosome. The transcript abundance data may comprise data from a plurality of transcripts located on the same chromosome and generating, by said processor, the value of the biomarker metric may comprise the processor scaling the generated value by the number of transcripts on the chromosome, and optionally determining the number of transcripts located on the chromosome.

[0020] The method may further comprise generating, by said processor, a value of the biomarker metric for each of a plurality of samples and clustering the values obtained. Said clustering may be performed using hierarchical clustering and a Manhattan distance. The method may comprise determining, by said processor, an intercluster distance based on the values of the biomarker metrics for samples in at least one pair of clusters. The pair of clusters may comprise a cluster comprising normal samples and a cluster comprising test samples. The method may comprise comparing samples in a first cluster and samples in a second cluster using the intercluster distance between the first cluster and the second cluster or between the first cluster and a third cluster and between the second cluster and the third cluster. The intercluster distance may be determined by calculating for each pair of samples, the distance between the biomarker values for each pair of samples comprising a sample from each of the pair of clusters, thereby obtaining a plurality of distances, and obtaining a summarised distances from the plurality of distances. The distance may be a Manhattan distance. The summarised distance may be an average distance.

[0021] The method may comprise obtaining transcript abundance data from the sample. Obtaining transcript abundance data from the sample may comprise processing RNA sequence data comprising RNA sequencing reads to determine an abundance for each of the plurality of transcripts based on the reads of the plurality of reads that map to the respective transcripts. Obtaining transcript abundance data from the sample comprises obtaining RNA sequencing reads from the sample by RNA sequencing. The transcript abundance data may have been obtained from RNA sequence data. The RNA sequence data may comprise a plurality of sequencing reads. The sequence data may comprise a plurality of aligned sequencing reads, for example in the form of a SAM or BAM file. The sequence data may comprise a plurality of sequencing reads each associated with coordinates in a reference genome or transcriptome to which the sequencing reads align.

[0022] The method may comprise obtaining RNA sequence data from the sample, and processing the RNA sequence data to obtain the transcript abundance data for the plurality of genes. The step of processing RNA sequence data may comprise one or more of: selecting one or more RNA sequence reads based on one or more quality metrics, removing one or more transcripts based on one or more criteria on the number of reads mapping to the transcripts in one or more samples, removing batch effects, and normalising the RNA sequence data.

[0023] The step of obtaining RNA sequence data from one or more samples from the subject may comprise or consist of receiving sequence data from a user (for example through a user interface), from one or more computing device(s), or from one or more data stores or databases. The step of obtaining RNA sequence data may further comprise sequencing (or otherwise determining the sequence composition of transcriptom ic material present in a sample) the sample.

[0024] The method may comprise identifying genomic coordinates for the plurality of genes.

[0025] The method may further comprise providing to a user, for example through a user interface, one or more of: value of the biomarker metric, information identifying a chromatin status derived from said value, the value of the correlation metric for each of one or more of the plurality of genomic distance bins, information derived from clustering of a plurality of samples comprising the sample based on the values of the biomarker metric for each sample and optionally one or more further metrics (such as e.g. a cluster associated with a sample of the plurality of samples, a distance between one or more clusters, etc.)

[0026] According to a second aspect, there is provided a method of observing the effects of one or more test agents on aging or disease in cells , the method comprising: optionally combining the test agent(s) with cells; obtaining transcript abundance data from a sample of the cells exposed to the test agent(s); determining the chromatin state of the sample using the method of any embodiment of the first aspect; and comparing the value of the biomarker value for the sample to one or more corresponding control values.

[0027] The methods according to the present aspect may have any one or more of the following optional features.

[0028] Comparing the biomarker value for the sample to one or more corresponding control values may comprise clustering the biomarker value for the sample with a biomarker value determined using the method of any embodiment of the first aspect for one or more control samples. The control samples may comprise one or more samples selected from: aged or diseased samples, healthy samples and samples treated with a rejuvenating agent.

[0029] According to a third aspect, there is provided a method of determining the age, rejuvenation state or reprogramming state of cells in a sample, the method comprising: obtaining transcript abundance data from the sample; determining the chromatin stateof the sample using the method of any embodiment of the first aspect; and comparing the value of the biomarker value for the sample to one or more corresponding control values. The control values may have been obtained from one or more samples with known ages, rejuvenation state or reprogramming state, using the methods according to any embodiment of the first aspect.

[0030] According to a fourth aspect, there is provided a method of treating a subject who has been diagnosed as having an aging-related disorder, the method comprising: administering to the subject a therapeutically effective amount of an agent that has been identified as effective in treating the aging related disorder, wherein the agent has been identified as effective in treating the aging-related disorder using a method according to any embodiment of the second aspect.

[0031] The methods according to the present aspect may have any one or more of the following optional features.

[0032] The method may further comprise identifying the agent that is effective in treating the aging-related disorder using a method according to any embodiment of the second aspect, wherein the one or more test agents include the agent that is identified as effective in treating the aging-related disorder.

[0033] The control values may include the value of a biomarker metric obtained using the method of any embodiment of the first aspect for at least one healthy sample and at least one sample comprising cells that have an aging-related disorder (e.g. cells no treated with the test agent(s)).

[0034] The aging related disorder may be HGPS or WRN.

[0035] According to a fifth aspect, there is provided a method of performing a reprogramming experiment, the method comprising: subjecting a sample of cells to be reprogrammed to an in vitro reprogramming protocol (e.g. incubation with a plurality of rejuvenation factors according to a predetermined protocol, such as e.g. OSKM), and performing the method of any embodiment of the third aspect. The comparison to control values may be used to identify a state of reprogramming of the cells at any particular point in the in vitro reprogramming protocol, thereby enabling to select protocol parameters (such as e.g. harvest day, quantity and / or time of addition of one or more rejuvenation factors) that result in cells having a desired reprogramming state. For example, different cells are known to respond differently to the same reprogramming protocol, making it labour intensive and difficult to adapt a protocol to obtain cells in the desired reprogrammed state. It is additionally difficult to ensure thatsuch adaptations have been successful. The methods described herein provides a solution to this by providing a biomarker that is indicative of chromatin state underlying reprogramming state, and which can be measured using a simple, commonly available data source (e.g. RNA sequencing). Thus, the method may comprise using the comparison to control value to select a time point for harvesting of cells and / or to select a quantity and / or timing of exposure of the cells to one or more reprogramming factors.

[0036] According to a further aspect, there is provided one or more non-transitory computer readable media comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of any method described herein, such as a method according to any embodiment of the first, second, or third aspects above.

[0037] According to a further aspect, there is provided a computer program comprising code which, when the code is executed on a computer, causes the computer to perform the steps of any method described herein, such as a method according to any embodiment of the first, second, or third aspects above.BRIEF DESCRIPTION OF THE FIGURES

[0038] Figure 1 is a flowchart illustrating schematically a method of characterising a sample as described herein.

[0039] Figure 2 is a flowchart illustrating schematically a method of performing an intervention screening according to embodiments of the disclosure.

[0040] Figure 3 shows an embodiment of a system for characterising a sample and / or for performing an intervention screen according to embodiments of the disclosure.

[0041] Figure 4 illustrates schematically the concept of the existence of long-range correlations between genes belonging to high (HAT, high abundance transcripts) and low abundance class (LAT, low abundance transcripts) in red and green colour respectively. Both the high and low abundance transcripts can exhibit strong correlations.

[0042] Figure 5 illustrates an unbiased clustering procedure used to segregate all transcripts into one of two groups in chromosome 1 : HAT and LAT. The x-axis is the order of genes after ranking their abundance in decreasing order.

[0043] Figure 6 shows a distance correlation coefficient measured based on abundances for binned pairs of transcripts as a function of Am, where pairs in each binare separated by Am between them. Red is HAT transcripts, green is LAT transcripts. The inset shows the same data but zoomed into the range of Am between 0-500.

[0044] Figure 7 shows an autocorrelation function as a function of lag for HAT (red, left) and LAT (green, right) class. The correlation length is defined as the point where the transcript-transcript autocorrelation falls below the 95% confidence limit.

[0045] Figure 8 shows a hierarchical clustering of the feature matrix LHAT (correlation length for HAT genes, for each sample (row) and chromosome (column)). The data shows a clear clustering of the samples based on treatment and disease, a consistent inner level of merging at the level of replicates and eventual merging of the samples. WT = Wild Type; HGPS SCR = Hutchinson-Guilford Progeria Syndrome (HGPS) treated with scrambled ASO; HGPS NT = HGPS with no treatment; HGPS L1 ASO = HGPS treated with Line-1 ASO; WRN SCR = Werner Syndrome (WRN) treated with scrambled ASO; WRN NT = WRN with no treatment; and WRN L1 ASO = WRN treated with Line-1 ASO. See Della Valle et al (2022).

[0046] Figure 9A shows pairwise cluster distances of the samples in Fig. 8 with respect to the wildtype. The data shows that ASO (antisense oligonucleotide) treated samples are much closer to the wildtype than their diseased counterparts. WT = Wild Type; HGPS SCR = Hutchinson-Guilford Progeria Syndrome (HGPS) treated with scrambled ASO; HGPS NT = HGPS with no treatment; HGPS L1 ASO = HGPS treated with Line- 1 ASO; WRN SCR =Werner Syndrome (WRN) treated with scrambled ASO; WRN NT = WRN with no treatment; and WRN L1 ASO = WRN treated with Line-1 ASO. See Della Valle et al (2022).

[0047] Figure 9B shows normalized correlation length EHAT measured for diseased samples taken from chromosome 1 and 6. The figure illustrates that diseased samples exhibit a higher correlation length than the ASO treated / wildtype samples, and this effect is particularly noticeable on chromosome 6. WT = Wild Type; HGPS SCR = Hutchinson-Guilford Progeria Syndrome (HGPS) treated with scrambled ASO; HGPS NT = HGPS with no treatment; HGPS L1 ASO = HGPS treated with Line-1 ASO; WRN SCR =Werner Syndrome (WRN) treated with scrambled ASO; WRN NT = WRN with no treatment; and WRN L1 ASO = WRN treated with Line-1 ASO. See Della Valle et al (2022).

[0048] Figure 10 shows a comparison of clustering using correlation length in different scenarios: considering only LAT genes in the dataset of Figures 8-9. (A), considering only HAT genes (B), considering all genes (C), considering only HAT genes but withoutchromosome 6 (D), considering only HAT genes but without chromosome 1 (E), and considering HAT genes multiplied by PCA1 of normalised transcript abundance (F).

[0049] Figures 11 A, 11 B and 11 C show results of hierarchical clustering using PCA (principal component analysis) loadings or values from SVD (singular value decomposition) of expression levels in the dataset of Figures 8-9. Fig. 11A shows results of hierarchical clustering using only the first principal component (PC). The data shows that diseased samples (NT and SCR) merge with each other rather than merging with their own types. Fig. 11 B shows results of hierarchical clustering using first 10 PCs. Again the diseased samples are not separated. Fig. 11 C shows results of hierarchical clustering using SVD.

[0050] Figure 11 D shows mean I* values calculated for chromosome 6 for samples in the different groups in the data from Della Valle et al.

[0051] Figure 12A shows results of hierarchical clustering using correlation length on a dataset from Fleisher et al. (2018).

[0052] Figure 12B shows results of hierarchical clustering using the first 10 PCs of normalised transcript abundance on the same dataset as Figure 12A.

[0053] Figure 12C shows boxplots of the mean I* across chromosomes for each age bracket in the Fleischer dataset.

[0054] Figure 12D shows mean I across chromosomes for discrete ages in the Fleischer dataset rather than age bins, with a fitted linear regression model (linear regression coefficients.0173, p=2.01769*1 O’10).

[0055] Figure 13A illustrates schematically the spatial distribution (along genomic coordinates) of gene expression (bottom) associated with RNA polymerase (black circles) transcribing DNA in a healthy young sample, with intact closed chromatin and intact histone density.

[0056] Figure 13B illustrates schematically the spatial distribution (along genomic coordinates) of gene expression (bottom) associated with RNA polymerase (black circles) transcribing DNA in an aged sample, with loss of heterochromatin.

[0057] Figure 14A shows gene expression data as a function of genomic coordinates on chromosome 6 in WT cells in the data from Dalia Valle et al.

[0058] Figure 14B shows gene expression data as a function of genomic coordinates on chromosome 6 in HGPS-non treated cells in the data from Dalia Valle et al.

[0059] Figure 14C shows gene expression data as a function of genomic coordinates on chromosome 6 in ASO treated cells in the data from Dalia Valle et al.

[0060] Figure 15A shows gene expression data as a function of genomic coordinates on chromosome 12 in a 51 years old donor (data from Fleischer et al. 2019).

[0061] Figure 15B shows gene expression data as a function of genomic coordinates on chromosome 12 in a 94 years old donor (data from Fleischer et al. 2019).

[0062] Figure 16 shows plots illustrating the relationship between autocorrelation (bottom plots) and variance for multiple single spatial distributions of gene expression with different variances (shown on the top plots).

[0063] Figure 17 illustrates schematically a method according to embodiments of the disclosure.

[0064] Figure 18A shows results of hierarchical clustering using I* on a reprogramming dataset from Gil et al. 2022.

[0065] Figure 18B shows absolute I* (mean across chromosomes) for different cellular states identified in Gil et al. 2022.

[0066] Figure 18C shows the mean I* on chr 7 (error bars indicate standard deviation across samples), by state on day 13, for the same samples as Figs. 18A and B.

[0067] Figure 19A shows mean I* values for chromosome 6 in a dataset where HUVEC cells (human umbilical vein endothelial cells) were cultured in the presence of chemicals being tested for their effect on cellular rejuvenation. Figure 19A shows results for two compounds shown to have a rejuvenation effect based on gene expression and methylation data.

[0068] Figure 19B shows mean I* values for chromosome 6 from the same dataset as Fig. 19A but showing data for a compound that was shown to have an aging effect based on gene expression and methylation data.

[0069] Figure 19C and Figure 19D show the distribution of mean I* for HUVEC cells cultured for multiple passages as an in vitro model of aging. Fig. 19C shows data for a first cell line, HUVEC B, and Fig. 19D shows data for a second cell line, HUVEC D.

[0070] Figures 20A, B, C show hierarchical clustering results using an intrachromosomal correlation length metric (I*=IHAT) calculated for high expression transcript selected using spectral clustering with different values of the gamma parameter. Data is from liver samples separated by age, from GTEx. Fig. 20A shows results for gamma=3.5. Fig. 20B shows results with gamma=2.5, and Fig. 20C shows results for gamma=1 .DETAILED DESCRIPTION

[0071] The present disclosure provides methods to determine the chromatin state of a sample using transcript abundance data, such as e.g. data from bulk RNA-seq. Different chromatin states may refer to different amounts or proportions of open vs closed chromatin in one or more regions of the genome, or different lengths of regions of open chromatin in one or more regions of the genome. Different chromatin states may be associated with diseases, with natural aging, pathological aging, rejuvenation, and / or reprogramming. Thus, the methods described herein can be used to determine the state of cells in terms of natural age, pathological age, rejuvenation status, reprogramming status, disease status, as well as the effects of any perturbation (e.g. drug treatment) on any of these. Traditionally, biologists have used RNA-seq data to analyze gene expressions and had to rely on other techniques like atac-seq to understand changes in the chromatin structure. ATAC-seq is an expensive procedure compared to relatively cheaper and more prevalent bulk RNA-seq. The present inventors recognised that changes in the chromatin structure not only affect the expression of a gene but also the number of genes that can express themselves - certain genes can be silenced if they are part of heterochromatin, while certain genes in the euchromatin regions are transcriptionally active. Thus, they postulated that if one was able to capture the average length of correlation between genes in the transcribeable region, this would provide a measure on the changes in the chromatin structure for a particular sample (e.g. cell line). Because there is a non-linear relationship between genes, the inventors recognised that simply and only apply metrics like Pearson correlation would not adequately capture this phenomenon. “To explain away”, the non-linear relationship between genes, the genes are bucketized based on their relative distance between each other. Thus, a bucket_k, contains a list of pairs of genes that are ‘k genes apart’ on the chromosome. A non-linear correlation metric, such as distance correlation (aka Brownian correlation, in the present examples) is used to capture the correlation between the pairs of genes in a bucket. Having explained away the non-linear relationship between genes, the inventors propose to develop a metric that captures the dependency between buckets of genes. In particular, they used autocorrelation to measure the dependency between the buckets of genes - i.e., is correlation value of the bucket containing genes separated by ‘2 genes apart’ dependent on bucket containing genes separated by '1 gene apart’. Similarly is the correlation of the bucket containing genes separated by ‘3 genes apart’ dependent onthe bucket containing genes separated by '2 genes apart’. So on and so forth. Because each bucket advances over the previous bucket by 1 gene distance, the significant lag measure by the intersection autocorrelation and 95% confidence interval envelope is indicative of the maximum average gene influence length. Finally, the inventors further recognised that not all genes are equally important in these analyses. They therefore proposed a data driven approach to segregate signals from the noise, by identifying a group of genes (HAT) that are more informative for the above process.

[0072] The methods described can be used to quantify similarity between different cellular states (e.g. diseased vs treated vs healthy), to measure the changes in a cellular state due to an intervention (e.g. test agent, candidate therapeutic, etc), to determine more accurately the true age (cellular age) of a subject, and for early detection of a disorder that has an effect on chromatin state (e.g. some Mendelian disorders) - for example by comparing a sample from a subject suspected of having the disorder to a control value (e.g. a metric obtained using the methods described herein for a control sample, e.g. healthy sample).

[0073] In the present disclosure, the following terms will be employed, and are intended to be defined as indicated below.

[0074] A “sample” as used herein may be a cell or tissue sample, a biological fluid, an extract (e.g. a RNA extract obtained from a subject), from which transcriptom ic material can be obtained for transcriptom ic analysis, such as RNA sequencing (also referred to as “RNAseq” or “RNA-seq”). The sample may be a cell, tissue or biological fluid sample obtained from a subject (e.g. a biopsy). Such samples may be referred to as “subject samples” or “individual samples”. In particular, the sample may be a sample comprising skin and / or blood cells. The skin and / or blood cells may be selected from: fibroblasts, keratinocytes, buccal cells, endothelial cells, lymphoblastoid cells, and / or cells obtained from blood skin, dermis, epidermis or saliva. For the purposes of obtaining RNA sequence data (also referred to herein as “transcript abundance data”), a “sample” as used herein may be a cell or tissue sample, or an extract (e.g. a RNA extract obtained from a subject) from which transcriptom ic material can be obtained. The sample may be one which has been freshly obtained from a subject or may be one which has been processed and / or stored prior to transcriptom ic analysis (e.g. frozen, fixed or subjected to one or more purification, enrichment or extraction steps). The sample may be a cell or tissue culture sample. As such, a sample as described herein may refer to any type of sample comprising cells or transcriptom ic material derived therefrom, whether froma biological sample obtained from a subject, or from a sample obtained from e.g. a cell line. In embodiments, the sample is a sample obtained from a subject, such as a human subject. The sample is preferably from a mammalian (such as e.g. a mammalian cell sample or a sample from a mammalian subject, such as a cat, dog, horse, donkey, sheep, pig, goat, cow, mouse, rat, rabbit or guinea pig), preferably from a human (such as e.g. a human cell sample or a sample from a human subject). Further, the sample may be transported and / or stored, and collection may take place at a location remote from the sequence data acquisition (e.g. sequencing) location, and / or any computer- implemented method steps described herein may take place at a location remote from the sample collection location and / or remote from the sequence data acquisition (e.g. sequencing) location (e.g. the computer-implemented method steps may be performed by means of a networked computer, such as by means of a “cloud” provider).

[0075] A “normal sample”, “healthy sample” or “wild type sample” refers to a sample that is assumed not to be in a disease state, such as e.g. a sample derived from a healthy subject. A control sample may be a normal sample. A control sample may be a sample comprising cells at a specific, known developmental stage, such as e.g. an embryonic cell or an adult or mature cell. A “test” or “diseased” sample, or simply a sample that is being characterised using a method as described herein (for example by determining biomarker values as described herein and optionally comparing those to control values or values determined from one or more control samples), may be a sample comprising diseased cells, such as e.g. a sample derived from an individual with a disease or disorder. A disease or disorder may be an aging-related disorder. A disease or disorder may be a premature aging syndrome. A disease or disorder may be selected from: Hutchinson-Gilford progeria syndrome (HGPS) and Werner Syndrome (WS). A test sample may be a sample comprising cells at a specific developmental stage, for example cells obtained from an embryo or adult or mature cells, or cells from an elderly person. A “test” sample may be a sample comprising cells that have been treated with one or more test agents, or cells from a subject that has been treated with one or more test agents. A “control” sample may be a sample comprising cells that have not been treated with the one or more test agents, or cells from a subject that has not been treated with the one or more test agents. A test or control sample may be derived from a single cell. A test agent may be a reprogramming factor. A reprogramming factor may be selected from Oct3 / 4, Sox2, Klf4 and c-Myc (the “Yamanaka factors” or “OSKM”). A test agent may be radiation, chemical compoundsand drugs, antibodies, proteins or peptides, or antisense compounds (eg, ASO or siRNA). A test agent may be a genetic perturbation, such as a CRISPR-mediated decrease or increase in gene expression. The test sample may be from a perturb-seq library.

[0076] The term “sequence data” or “transcript abundance data” refers to information that is indicative of the presence of transcripts in a sample that have a particular sequence. Such information may be obtained using sequencing technologies, such as e.g. RNA sequencing or sequencing of captured genomic loci (targeted or panel sequencing), or using array technologies, such as e.g., expression arrays or other molecular counting assays. In the context of the present disclosure, the sequence data is typically obtained by RNA sequencing, and particularly next generation sequencing. Thus, the sequence data may comprise sequencing reads. Information such as the count of sequencing reads that have a particular sequence or that cover a particular genomic location can be derived from sequencing reads using methods known in the art. When non-digital technologies are used such as array technology, the sequence data may comprise a signal (e.g. an intensity value) that is indicative of the number of sequences in the sample that have a particular sequence, for example by comparison to an appropriate control. Sequence data may be mapped to a reference sequence, for example a reference genome or transcriptome, using methods known in the art. This may result in aligned sequencing reads, for example in the form of a SAM or BAM file. Thus, sequencing reads or equivalent non-digital signals may be associated with a particular genomic location (where the “genomic location” refers to a location in the reference genome to which the sequence data was mapped).

[0077] A composition as described herein may be a pharmaceutical composition which additionally comprises a pharmaceutically acceptable carrier, diluent or excipient. The pharmaceutical composition may optionally comprise one or more further pharmaceutically active polypeptides and / or compounds. Such a formulation may, for example, be in a form suitable for intravenous infusion.

[0078] As used herein "treatment" refers to reducing, alleviating or eliminating one or more symptoms of the disease which is being treated, relative to the symptoms prior to treatment. "Prevention" (or prophylaxis) refers to delaying or preventing the onset of the symptoms of the disease. Prevention may be absolute (such that no disease occurs) or may be effective only in some individuals or for a limited amount of time.

[0079] As used herein, the terms “computer system” includes the hardware, software and data storage devices for embodying a system or carrying out a method according to the above-described embodiments. For example, a computer system may comprise a central processing unit (CPU) and / or a graphic processing unit (GPU), input means, output means and data storage, which may be embodied as one or more connected computing devices. A computer system may comprise a display or comprises a computing device that has a display to provide a visual output display. The data storage may comprise RAM, disk drives or other non-transitory computer readable media. The computer system may include a plurality of computing devices connected by a network and able to communicate with each other over that network. It is explicitly envisaged that computer system may consist of or comprise a cloud computer.

[0080] As used herein, the term “computer readable media” includes, without limitation, any non-transitory medium or media which can be read and accessed directly by a computer or computer system. The media can include, but are not limited to, magnetic storage media such as floppy discs, hard disc storage media and magnetic tape; optical storage media such as optical discs or CD-ROMs; electrical storage media such as memory, including RAM, ROM and flash memory; and hybrids and combinations of the above such as magnetic / optical storage media.Characterisation of chromatin state

[0081] The present disclosure provides methods for determining the chromatin state of a sample. An illustrative method will be described by reference to Figure 1 .

[0082] The method may comprise optional step 100 of obtaining one or more samples comprising transcriptom ic material (such as e.g. one or more cell or tissue samples). Optionally, the samples may be sequenced at step 120, to obtain at least RNA sequence data. The RNA sequence data may be obtained by RNA sequencing. Alternatively, the sequence data may have been previously obtained and may be received from a user interface, computing device or database.

[0083] At optional step 140, transcript abundance data for a plurality of transcripts is obtained from the RNA sequence data. This may comprise obtaining a normalised and / or fractional transcript abundance. A normalised transcript abundance may be obtained for a sample by dividing the transcript abundance for each transcript by the mean transcript abundance across the plurality of transcripts. A fractional transcript abundance may be obtained for a sample by dividing the (optionally normalised)transcript abundance for each transcript by the sum of (optionally normalised) abundances for all transcripts in the plurality of transcripts.

[0084] At step 160, a first group of transcripts and a second group of transcripts are identified using the transcript abundance data, the first group of transcripts (HAT, high abundance transcripts) having higher abundance in the sample than the second group of transcripts (LAT, low abundance transcripts). This may be performed using any data driven approach, such as e.g. using clustering (e.g. spectral clustering). This may further comprise selecting the transcripts in the second group of transcripts.

[0085] At step 180, the transcripts that map to a first selected chromosome are selected. This may comprise obtaining genomic coordinates for each transcript by mapping the sequences of the transcripts to a reference genome. Alternatively, the transcripts may already be associated with genomic coordinates or genomic coordinates for transcripts may be obtained from a database (e.g. Ensembl).

[0086] At step 200, all possible pairs of transcripts are created from the plurality of transcripts selected at step 180 are created, and separated between a plurality of genomic distance bins. The genomic distance bins each comprise pairs of transcripts associated with genes that are n genes apart on the same chromosome, where n is a natural number or range of natural numbers, where n is different for each bin.

[0087] At step 220, a correlation metric between transcript abundance for all pairs of transcripts is calculated, for each bin obtained at step 200. The correlation metric may be a distance correlation coefficient, calculated aswhere:(x, y) are vectors of transcript abundance for all possible pairs of transcripts in the bin; andV is the empirical distance covariance for the transcripts in the bin, calculated aswhere D are matrices of all centred Euclidian pairwise distances between transcript abundance for transcripts in the bin, centred by the row mean and column mean.

[0088] At step 240, a biomarker value is derived from the correlation metric calculated for the plurality of bins at step 220. This may comprise calculating the autocorrelationof the correlation metric as a function of lag in number of genes part on the chromosome (i.e. the units used for the bins, where each bin is associated with a value of the lag, e.g. Iag= 1 , 2, 3, 4, 5, etc. up to the largest distance in terms of number of genes between any pair of transcripts on the chromosome). This may further comprise determining a confidence interval around the autocorrelation metric at each lag value. The biomarker value may be the largest lag value at which the autocorrelation metric is not significantly different from zero, at the level of confidence associated with the confidence interval. This value may be further scaled by the total number of genes considered on the chromosome. This may help to make the biomarker values comparable between chromosomes. This is not necessary when the same chromosome is compared across samples.

[0089] Steps 180 to 240 may be repeated for one or more further chromosomes. Thus, a biomarker value may be obtained for each of a plurality of chromosomes.

[0090] Steps 140 to 240 may be repeated for one or more further samples. At step 260, the biomarker value (which may comprise a plurality of biomarker values, one for each of a plurality of chromosomes) for a plurality of samples are compared, for example by clustering in the illustrated embodiment.

[0091] At step 280, an intercluster distance is obtained between a pair of cluster (e.g. a cluster comprising test samples and a cluster comprising normal samples). This may be performed by computing for each pair of samples comprising a sample in each of the pair of clusters, a Manhattan distance between the biomarker value for the samples. The intercluster distance may be used to compare the samples in the two clusters, e.g. to establish whether treated samples are closer to normal samples than untreated samples (by computing the cluster distance between (i) a cluster of treated samples and a cluster of normal samples, and (ii) a cluster of untreated samples and a cluster of normal samples).

[0092] At optional step 300, the results of one or more of the preceding steps are provided to a user, for example through a user interface.

[0093] The above methods find applications in the context of characterising perturbations and / or samples from diseased subject that have an effect on chromatin structure, particularly but not limited to aging and aging related diseases and disorders. Indeed, as shown in the examples below, the metrics described herein can be used toidentify samples from genetic aging disorders from those samples after treatment with a therapy that has been shown to have an effect in these diseases, and to show that the treated samples are closer to the healthy state when characterised using the metrics described herein. In other words, the methods described herein can identify samples that are rejuvenated from aged I diseased samples prior to treatment, and can quantify their closeness to a healthy state. The metrics described herein can also be used to identify the cellular ages of samples, as demonstrated using chronological age as a proxy. Further, the metrics described herein can also be used to diagnose a disease or disorder that has an effect on chromatin structure, for example by quantifying the metrics described herein for a sample from a subject suspected of having such a disorder, and comparing these to metrics obtained using the methods described herein for one or more control samples.

[0094] Thus, the present disclosure also provides methods of observing the effects of one or more test agents on aging or disease in human cells. These methods may comprise combining the test agent(s) with human cells (e.g. for specified period of time such as at least one day, one week or one month), and then observing the chromatin state of the human cells using the methods described herein. The methods then compare the observations from cells exposed to the test agent with observations of the chromatin state determined using the methods described herein for control cells not exposed to the test agent or a control or reference value such that effects of the test agent on aging or disease of the human cells is observed. Optionally, the test agent is a compound having a molecular weight less than 3,000, 2,000, 1 ,000 or 500 g / mol, an antibody, a polypeptide, a polynucleotide, an antisense compound (e.g., ASO, siRNA) or the like. The cells may be fibroblasts. The cells may be human cells.

[0095] Also described herein are methods of observing the effects of one or more test agents on aging or disease in human cells and comparing the effects of the one or more test agents to the effects of a different test agent, using the methods described here. For example, the methods of the disclosure can be used to compare effects of a rejuvenating agent, e.g., OSKM, with the effects of a test agent, to determine whether the test agent has similar properties to OSKM.

[0096] In some embodiments, the method comprises administering the one or more test agent to a human subject. In further embodiments, the method comprises observing the effect of the one or more test agent by removing a cell or cells or tissues from the human subject, then observing the effects of the test agent. Thus, also described hereinare methods of observing the effects of one or more test agents on aging or disease in a human subject. The methods may comprise administering the test agent(s) to the subject, and then observing the chromatin state of a sample comprising cells from the subject using the methods described herein. The methods then compare the observations from cells from the subject treated with the test agent with observations of the chromatin state determined using the methods described herein for a a sample from a subject not treated with the test agent and / or a sample from a healthy subject, or a control or reference value such that effects of the test agent on aging or disease of the human subject is observed.

[0097] Also described herein are methods of determining the cellular age of an individual by comparing the value of a biomarker metric as described herein obtained from a sample of cells of the individual to corresponding values obtained from healthy subjects with known chronological age. The age of the individual may be determined as the age associated with one or more samples from healthy subjects that are associated with a biomarker value that is / are closest to the biomarker value obtained for the individual. In other words, the biomarkers described herein can be used to measure the age of an individual based on transcript abundance data (or in other words, based on RNA extracted from a sample) from a sample of cells or tissues, e.g., skin (eg fibroblasts).

[0098] The biomarkers described herein are also useful for ex vivo studies of anti-aging interventions, thus allowing interventions to be quickly evaluated based on real-time measures of aging. Embodiments of the invention are also useful for applications in personalized medicine, as it allows for evaluation of accelerated or decelerated aging effects based on transcript abundance data (i.e. RNA expression data).

[0099] Also described herein are methods of observing biomarkers in a sample, eg, skin sample, that correlate with an age of an individual, the method comprising observing the individual’s chromatin status by determining the value of a biomarker metric as described herein. Because the biomarker metrics described herein are associated with the age of the individual, determining the chromatin state of a sample as described herein comprises observing biomarkers associated with the age of the individual. Typically, the skin and blood cells are human fibroblasts, keratinocytes, buccal cells, endothelial cells, lymphoblastoid cells, and / or cells obtained from blood skin, dermis, epidermis or saliva. Embodiments of this method further comprise using the observations to estimate the age of the individual (e.g. using a regression analysisor the like, or by comparing the biomarkers to biomarker values obtained for control samples with known age).

[0100] Thus, the present disclosure provides methods for identifying agents that can treat a chromatin-related disorder such as aging-related disorders, and methods of treating a subject who has been diagnosed as having an aging-related disorder. An illustrative method will be described by reference to Figure 2.

[0101] At optional step 20, a sample comprising cells that have an aging-related disorder is obtained (diseased sample). Optionally, one or more samples comprising healthy cells may also be obtained (normal sample). At optional step 22, the cells in the diseased sample are exposed to one or more test agents. At optional step 24, transcriptom ic data is obtained from a sample of the cells exposed to each test agent, and optionally also from a diseased sample not exposed to any test agent and / or the normal sample. At step 26, transcript abundance data is obtained for each sample, from the transcriptom ic data. At step 28, the chromatin state of each sample is determined using a method as described herein, such as e.g. by reference to Figure 1. At step 30, the one or more test agents are compared based on the determined chromatin state. For example, test agents that bring the biomarker value for a diseased sample treated with the test agent closer to the biomarker value for a normal sample (compared to the initial distance between the biomarker value obtained for a disease sample not treated with the agent) may be considered to be effective in treating the aging related disorder. At optional step 32, the subject is treated with a test agent identified at step 30 as effective in treating the aging related disorder.

[0102] Figure 3 shows an embodiment of a system for characterising a sample, and / or for screening one or more perturbations, according to the present disclosure. The system comprises a computing device 1 , which comprises a processor 101 and computer readable memory 102. In the embodiment shown, the computing device 1 also comprises a user interface 103, which is illustrated as a screen but may include any other means of conveying information to a user such as e.g. through audible or visual signals. The computing device 1 is communicably connected, such as e.g. through a network 6, to sequence data acquisition means 3, such as a sequencing machine, and / or to one or more databases 2 storing sequence data. The one or more databases may additionally store other types of information that may be used by thecomputing device 1 , such as e.g. reference sequences, parameters, etc. The computing device may be a smartphone, tablet, personal computer or other computing device. The computing device is configured to implement a method as described herein. In alternative embodiments, the computing device 1 is configured to communicate with a remote computing device (not shown), which is itself configured to implement a method as described herein. In such cases, the remote computing device may also be configured to send the result of the method to the computing device. Communication between the computing device 1 and the remote computing device may be through a wired or wireless connection, and may occur over a local or public network such as e.g. over the public internet or over WiFi. The sequence data acquisition 3 means may be in wired connection with the computing device 1 and / or the database 2, or may be able to communicate through a wireless connection, such as e.g. through a network 6, as illustrated. The connection between the computing device 1 and the sequence data acquisition means 3 may be direct or indirect (such as e.g. through a remote computer). The sequence data acquisition means 3 are configured to acquire sequence data from nucleic acid samples, for example RNA samples extracted from cells and / or tissue samples. In some embodiments, the sample may have been subject to one or more preprocessing steps such as RNA purification, fragmentation, library preparation, target sequence capture (such as e.g. polyA capture, exon capture and / or panel sequence capture). The sequence data acquisition means is preferably a next generation sequencer. The sequence data acquisition means 3 may be in direct or indirect connection with one or more databases 2, on which sequence data (raw or partially processed) may be stored.

[0103] The following is presented by way of example and is not to be construed as a limitation to the scope of the claims.EXAMPLESIntroduction

[0104] In this work, the inventors present and exemplify a new method to characterise samples and perturbations based on spatial transcript-transcript abundance correlations (i.e. correlations and changes in correlation along genomic coordinates).

[0105] In particular, the objective of the study described in the examples below was to develop a consolidated, data-driven approach to validate the hypothesis that epigeneticmodification resulting from aging or pre-aging conditions may alter chromatin structure which ultimately affects long-range correlations between gene expression. The work below systematically explores the pairwise correlation of transcript abundance between distant genes, as a function of their proximity. This method is then applied to various whole genome (bulk RNA) transcriptom ic datasets from human fibroblast to establish transcript-transcript correlation based on their abundance. A new metric termed “correlation length” is then developed and proposed as a feature to group cells into different states (healthy vs. diseased).

[0106] The methods described herein provide a solution to the problem of characterising the chromatin state of samples, using gene expression data.

[0107] In particular, the inventors set out to develop a new RNA-seq biomarker that probes chromatin structure. The biomarker aimed to capture changes in chromatin (e.g. to provide information about restoration of chromatin integrity in rejuvenation treatments as well as effects of aging on chromatin - where aging is associated with heterochromatin loss), changes in gene expression (e.g. to provide information about restoration of cell functionalities in rejuvenation treatments), and to provide a metric that can differentiate patterns across different cell states (e.g. able to discriminate between young vs. old, diseased vs. healthy vs. treated). This work is based on the assumption that if chromatin structure constraints transcription, there is loss of heterochromatin as one ages and one of the effects of rejuvenation is its restoration, it would be extremely useful to come up with a metric that is a proxy to the length of the open chromatin in order to monitor aging and rejuvenation.Example 1 - Correlation length as a measure of co-expression

[0108] In this example, the inventors propose and describe the concept of correlation length as a way of characterising samples.Methods

[0109] Data. Two publicly available whole genome bulk RNA-seq data sets from human fibroblast cells were used for the autocorrelation analysis.

[0110] The first dataset (Line-1 or L1 ), originally published in Della Valle et. Al., contains 7 human fibroblast samples each containing 3 biological replicates - Wild Type (WT), Hutchinson-Guildford Progeria Syndrome (HGPS; sometimes termed progeria in these Examples) cells treated with scrambled ASO (HGPS SCR), HGPS cells with notreatment (HGPS NT), HGPS cells treated with Line-1 ASO (HGPS L1 ASO) and Werner Syndrome (WRN; sometimes termed “Werner” in these Examples) cells treated with scrambled ASO (WRN SCR), WRN cells with no treatment (WRN NT), and WRN cells treated with Line-1 ASO (WRN L1 ASO). This is available on the NCBI SRA (www.ncbi.nlm.nih.gov / sra) under the accession numbers PRJNA704498 and GSE1 98675.

[0111] HGPS and WS are premature aging genetic disorders that are characterized by disorganized heterochromatin (Scaffid and Misteli, 2005; Osorio et al., 2011 ; Ocampo et al., 2016; Zhang et al., 2015). HGPS causes children to age rapidly, starting in their first two years of life. Heart problems or strokes are the eventual cause of death in most children with HGPS. The average life expectancy for a child with HGPS is about 15 years. Werner’s syndrome patients are marked by rapid aging that begins in early adolescence or young adulthood and an increased risk of cancer. Signs and symptoms include shorter-than-average height, thinning and graying hair, skin changes, thin arms and legs, voice changes, and unusual facial features. Della Valle et al. (2022) demonstrated that treatment with an anti-LINE 1 ASO can ameliorate aging hallmarks, e.g. reducing DNA methylation age and restored heterochromatin histone marks, in cells from HGPS and WS patients.

[0112] The second dataset was obtained from Fleischer et al. and is comprised of human fibroblast cells with 143 samples from human subjects of different ages ranging from 1 year to 94 years. This is available on Gene Expression Omnibus under accession number GSE113957 (www.ncbi.nlm.nih.gov / geo / ). Amongst the 143 donors, there are 10 donors suffering from Hutchinson-Gilford Progeria Syndrome (HGPS), a premature aging disease. These were removed from the dataset to avoid misleading analysis and conclusion.

[0113] These two datasets are particularly useful in testing the hypothesis of identifying correlations between large-scale changes in gene expression and modification in chromatin structure induced by epigenetic alterations, such as heterochromatin loss in aging and related conditions.

[0114] Transcriptom ics data analysis. The transcripts were annotated using Ensembl identifiers and their genomic loci were mapped from the reference human genome GRCh38.p13 using the BioMart - Ensembl tool. Transcripts that do not express in any of the samples (i.e. , their abundance values are 0 in the entire dataset) were dropped. In the case of Fleischer dataset, we additionally sub-divide the dataset based on theage ranges: 0-20, 21 -40,41 - 60, 61 - 80 and 81 - 100 to demonstrate trends in the correlation length as a function of ages and infer some characteristic properties.

[0115] For all subsequent analysis including calculation of the correlation length, each sample in the dataset is normalized by its mean and divided by the sum to obtain the fraction of each transcript contribution to the total abundance.

[0116] Identification of high and low abundance transcripts. For each sample individually, transcripts are sorted and assigned a rank based on their abundance value such that the transcript with the highest abundance will be assigned rank 1 and the one with the lowest abundance will be assigned the highest rank (i.e. total number of transcripts considered in that sample). Rearranging the samples will give us a smooth monotonic decay of transcript abundance (see Figure 5) which then can be separated into two classes: high abundance transcript (HAT) and low abundance transcript (LAT) using a spectral clustering algorithm with Laplacian affinity. This step is performed to remove the noise occurring due to sequencing depth errors (low read counts) across different samples. Note that clustering is performed on the normalised abundance data, and the ranks are shown for illustrative purposes, to show the noise in the data highlighted by departure of the ranked data from a linear relationship in log-log scale.

[0117] Alternative kernels for spectral clustering can be used, such as RBF (Radial Basis Function) kernel or Nearest Neighbour kernel. Further alternative clustering methods may also be used, such as e.g. k-means, k-medoid, etc. In general, any data driven method to separate LAT and HAT genes (or more generally, a first group of data points I region in a graph from a second group of data points I region in a graph, such as e.g. the “Kneedie” algorithm which identifies knee points I elbow points in a curve, see Satopaa et al. 2011 ) may be used. The present inventors found spectral clustering, particularly with Laplacian affinity, to perform particularly well (by comparing the clustering of samples from HAT derived correlation lengths - see Example 2 - obtained with data that has been clustered using different algorithms to identify HAT / LAT transcripts).

[0118] Distance correlation. Given the division of the genes into the HAT and LAT classes, we quantify the positional dependence of gene expression by calculating the average autocorrelation function of the pairwise transcript abundance from the normalized RNA-seq data. For each sample and each chromosome, the pairwise normalised transcript abundances are binned based on the chromosomal positions of the genes in a pair: all gene pairs that are 1 gene apart (direct neighbours) are in bin 1 ,all gene pairs that are 2 genes apart are in bin 2, etc. The distance between genes is referred to as Am. Thus, bin 1 is Am=1 , bin 2 is Am=2, etc. For example, considering a hypothetical arrangement of genes on the chromosome as follow: A B C D E F G, then pairs of genes that are in Am=1 would be <A,B>, <BA>, <B,C>, <C,B>, <C,D>, <D,C>, <D,E>, <E,D>, <E,F>, <F,E>, <F,G>, <G,F>. Pairs of genes that are in Am=2 would be <A,C>, <C,A>, <B,D>, <D,B>, <C,E>, <E,C>, <D,F>, <F,D>, <E,G>, <G,E> etc. Bin i contains transcript abundance of all possible pairs of genes m where m is the total number of pairs of genes that are separated by Am=i in the group of genes considered (e.g. all HAT genes in the chromosome considered). The total number of pairs ni in each bin vary and is typically decreasing as i (i.e. Am) increases. We calculate a distance correlation coefficient C(Am) for each bin, using the following equation (similar to what is described in Szekely, Rizzo and Bakirov, 2007):where:(x, y) are vectors of transcript abundance for all possible pairs of genes in the bin (i.e. all pairs of genes in the data for the sample where the genes are separated by Am number of genes on the genome); andV denotes the empirical distance covariance, which can be written in terms of centred Euclidian distances D using the following equation:where D are matrices of all pairwise distances between normalised transcript levels for genes in the bin, i.e. distance matrices containing all pairwise distances (||Xj-Xk|| and l|Yj-Yj||, respectively, where ||.|| denotes Euclidian norm), centred by the row mean and column mean. For each bin, there is a pair of vectors (X, Y). One matrix D is calculated for vector X and one for vector Y. The distance correlation is a measure of dependence between two paired random vectors, where the population distance correlation coefficient is zero if and only if the random vectors are independent such that both linearand non-linear associations can be captured. By contrast, Pearson correlation only detects linear association between two random variables.

[0119] Distance correlation is a useful metric in identifying correlation between two random variables when there exists a non-linear relationship between them. Since transcript abundance or gene expression is highly non-linear (i.e. transcript abundance is a non-linear function of gene position, i.e. genes are not necessarily expected to correlate linearly as a function of distance in gene position), other measures like Pearson correlation may not be sufficient to capture the complex relationship.

[0120] Correlation length. The inventors further perform the auto-correlation of C(Am) which measures the relationship of a variable with lagged values of itself. This provides a measure to obtain the length scale of the correlation phenomenon. Performing autocorrelation is useful to capture the scale of distance correlation decay till the point of sharp increase for high (Am). We then measure the correlation length for each chromosome as the point where the transcript-transcript autocorrelation falls below the 95% confidence limit (the shaded regions in Figure 7). Note that a 95% confidence limit was used as a commonly used threshold for looking at confidence of metrics, but the results described herein work equally for other choice of confidence limits such as e.g. any confidence between 90% and 99%. When calculating the autocorrelation on white noise data, the correlation for all the lags will fall inside the envelope (95% confidence interval, i.e. a white noise process has zero autocorrelations for all values of the lag, in other words white noise is serially uncorrelated). In other words, with white noise data the value of the data at one point in a series is not predictive of the value at a later point in the series. By contrast, in data where there is high (significantly different from zero) autocorrelation, the autocorrelation for at least some of the values of the lag will be outside of the confidence Interval envelope (i.e. significantly different from 0). When plotting the envelope of the 95% confidence interval as a function of lag, the intersection between the autocorrelation data series and the envelope (i.e. where the value of the autocorrelation comes outside of the envelope) determines the maximum lag that still leads to significant autocorrelation. The confidence interval (Cl) of autocorrelation AC at each lag k (ACk) can be calculated as CIACk=where ACSE.K is the standard error of the autocorrelation at lag k, calculated as ACSE k= where Aci is the autocorrelation estimate at lag i<=k and N isthe number of time steps (here gene lags) in the sample.

[0121] Finally, we construct 2-dimensional matrices LHAT and LLAT of size Nsampies ><Mchr (number of samples by number of chromosomes), the coefficients of which are the correlation lengths for each sample and chromosome, scaled by total number of genes on the respective chromosome. For hierarchical clustering of the samples shown in Figure 8, we utilised the correlation length obtained from the transcripts in the high abundance class which led to segregation of samples accurately.Results

[0122] First, the inventors evaluated whether the use of RNA-seq expression values alone (i.e. not taking into account chromosomal positions in the analysis) could be used as a biomarker of aging. Figure 11 C shows the clustering results of SVD on LINE-1 dataset. Figures 11 A and 11 B show the clustering results using 1 principal components and 10 principal components, respectively. These results show that SVD shows lot of interleaving cell lines. While clustering using principal components shows clearer segregations at a high level, (e.g., WT and ASO merging while on all progeria merging together and all WRN merging together), there is still interleaving of NT and SCR lines. Figure 12B shows clustering results using the Fleischer dataset on 10 principal components (which cover similar variance). This shows that when using PCA loadings on gene expression data alone, there is heavy misclustering. Note that the data was bucketized in 5 age groups as shown to ensure that there is enough samples in each age bracket. This shows that conventional RNA-seq analyses techniques are limited. While they are helpful in analyzing the downstream resultant effects of an expression value, the inventors postulated that their limitations were at least in part due to the fact that they do not incorporate any modelling and understanding the underlying causes of expression differences between samples. Consequently, it has limited interpretability in explaining the observations. The present inventors postulated that a new measure that does incorporate some elements pf functional and structural causes could be used in conjunction with these traditional approaches to get better and more complete picture of underlying dynamics.

[0123] Gene transcription is one of many factors that is age and cell state dependant. Figure 13A illustrates schematically RNA polymerase (black circles) transcribing DNA in a healthy young sample, with intact closed chromatin and intact histone density. For a certain histone density and epigenetic state, this will result in an inherent gene expression which is plotted against the position of the gene on the plot in the lowerpanel (where the x axis here represents the genome coordinates and the plot shows a spatial distribution of gene expression, rather than a histogram). As the RNA polymerase traverses through the DNA various more genes are transcribed and will result in a series of different types of distribution of normalised gene expression along the genome. This can be of any form but is illustrated here as normal distributions with different parameters at each location. The horizontal arrow illustrates the variance (spread) of the respective distributions. Figure 13B illustrates the same locations for aged cells, where it is postulated that there is a loss of heterochromatin. If that is the case, this should result in an epigenetic state where there is more open chromatin, relatively lower histone density and the polymerase can access more functional DNA to transcribe. This will then open up adjacent sites and activate repressed genes. Now there are more genes being expressed this would result in a flatter and wider distribution compared to more healthy young cells. Again the shape of the distributions shown is irrelevant and for illustrative purposes, the point being that for aging cells, the present inventors postulated that there may be a higher variance of gene expression not only globally but also locally. In particular, the inventors observed that in aged cells, there is an increased number of activated genes and this leads to a higher variance in gene expression with respect to genomic coordinates (i.e. negative kurtosis - flatter and wider probability distributions of gene expression along genomic coordinates in aged cells compared to healthy / young cells).

[0124] This is demonstrated in real data sets (LINE-1 dataset from Della Valle et al.) on Figure 14. Figure 14A shows data for WT cells, Figure 14B shows data for HGPS NT cells, and Figure 14C shows data for the ASO treated cells. All data shown focusses on a single chromosome (chromosome 6). The scatter plot at the bottom of each subfigure is the original dataset where the x axis is the genomic coordinates and the y axis is normalized RNAseq data. A mixture of Gaussian distributions was fitted (using a Dirichlet process algorithm)) to the RNAseq data to construct the contour plots shown on the scatter plots. The corresponding Gaussian distributions are shown on the top panel in each subfigure. The dashed lines is the resultant distribution coming from the mixture of individual Gaussian distributions shown as solid lines. By inspecting these plots, it becomes apparent (especially in the contour plots) that there is a slight increase in the variance of the distributions (flatter shapes) for a diseased Progeria state when compared with WT and the ASO treated samples. The ASO treatment brings the “fitted” distributions closer to WT but the means and the variance of eachdistribution is still quite different from WT, because of the uncertainties and quality of fitting. This motivated the inventors to look for a metric other than variance of the spatial distribution that can capture this mechanism. Figure 15 shows similar data for the aging data set in Fleischer 2019, for chromosome 12 for donors with different ages (Fig. 15A shows data for a 51 years old donor, Fig. 15B shows data for a 94 years old donor). This shows that there is a higher variance for an aged sample compared to a younger sample.

[0125] In summary, the inventors postulated that unravelling of heterochromatin leads to an increase in the number of genes that are expressed. If we believe that heterochromatin unravelling happens due to aging, the aging process should open up a lot of adjacent regions and change the shape of the distribution of gene expression along chromosome coordinates - resulting in an increase in variance (lower kurtosis). A challenge in capturing this phenomenon is that we do not want to assume some kind of distribution to start with, and even if we do it would be challenging to measure the increase in the variance “reliably” (as shown above) because of the sheer complexity of the data, the presence of high noise and the fact that the data is expected to be a mixture of different distributions. Therefore, the inventors set out to design a methodology that can provide an alternative to capture this phenomenon, and ideally that is usable as a macroscopic global measure, that takes into account global changes in the characteristics nature of gene expression distribution.

[0126] The inventors first tested the evidence of positional dependence of gene expression by measuring pairwise correlation of transcript abundance. Figure 4 shows a graphical representation of genes which are epigenetically similar (or belongs to same chromatin environment) exhibiting higher correlations. Since genes which are lowly expressed can also be likely correlated, there must exist a correlation between transcripts with low abundance. The existence of such correlations between neighbouring genes which are part of same regions should be evident from the correlation length between transcripts.

[0127] As explained in the methods above, the inventors begin with mean normalizing the RNA-seq data for all samples. They then segregate transcripts in each chromosome into two classes based on their abundance values - high abundance transcripts (HAT) and low abundance transcripts (LAT). Figure 5 shows the mean normalized transcript abundance for chromosome 1 of wildtype (3 replicates) sample on a log-log scale in decreasing order of their abundance from the Line-1 data. Theranks (x-axis) are based on the magnitude of transcript abundance. The two classes, HAT and LAT, are shown in red and green respectively. The data shows points for all 3 replicates (i.e. there are in fact 3 points at any rank, one for each replicate).

[0128] The inventors then binned the pairwise transcript abundances based on their chromosomal positions where pairs in each bin were separated by Am number of transcripts between them (see methods for more details). Examining the spatial dependence of the correlation coefficient C(Am), measured using distance correlation, for HAT transcripts (in red, data series that increases first) shows a sharp decay for Am = 10 - 20 (see Figure 6 inset - note the peak at Am=0 where C(Am)=1 is very close to the y axis) and fluctuating around a mean value until Am = 200 beyond which there aren’t enough pair of genes to quantify the correlation with confidence. A similar pattern can be seen in Figure 6 (main) for transcripts in the LAT class (in green, data series that increases at higher values of Am).

[0129] By definition, the distance correlation coefficient is a non-negative number and a value of 0 signifies independence between pairs of transcripts. The correlation would be expected to decrease to zero as Am increases (i.e. the gene expression between pairs of genes is less and less correlated as pairs that are further apart are considered). However, due to fewer number of transcripts for high Am , the correlation shows a sharp increase towards the end. In order to obtain the scale over which the distance correlation is monotonically decaying, the inventors calculated the autocorrelation of C(Am). The autocorrelation of a data series is the correlation between the time series and its lagged version over time. Autocorrelation is typically talked about in the context of time series data but here time is replaced with space. In particular, in this case the lag is in terms of a number of genes distance. Autocorrelation depicts the influence / correlation of a point separated by various distances. A gentler sloped distribution is associated with a longer influence. Performing autocorrelation allowed to identify the range of Am over which the distance correlation decays, up until the point where it sharply increases. Figure 7 shows the spatial decay (i.e. decay as a function of lag in terms of number of genes) of the transcript-transcript autocorrelation function I for the two classes (HAT on the left and LAT on the right), for chromosome 1 , for the WT samples (3 replicates). Note that a similar pattern is observed for all chromosomes. The inventors combined the concept of autocorrelation with the concept of a confidence interval envelope to find out the maximum significant lag, which is indicative of the length of correlation of the gene expression signal along genomic coordinates. Asmentioned above, the underlying hypothesis is that reduced heterochromatin / histone density (due to e.g. aging or disease) would result in increased co-transcription, which would be identifiable as a longer length of correlation, leading to the hypothesis that it would be possible to correlate chromatin structural changes to the length of correlation of gene expression. In other words, the inventors hypothesised that the change in chromatin structure due to aging or disease should result in a longer length of correlation of gene expression. The inventors defined the correlation length (EHAT and LAT) as the point (i.e. gene lag) where the transcript-transcript autocorrelation falls below the 95% confidence limit (the shaded regions on Figure 7 which shows the 95% confidence limit of the autocorrelation for each value of the gene lag) and units of it is defined in terms of number of genes. For example, 27 genes belonging to the HAT class are correlated on average on chromosome 1 while 154 genes are correlated in the LAT class. Thus, the data shows that transcript abundance is spatially correlated along the chromosome for both the HAT and LAT transcripts, and decays faster for the former. This likely primarily stems from the fact that majority of the transcripts in the low abundance class have similar values (often close to 0). This can be confirmed from Figure 6 (main), where one can observe the relative difference in not only on the y-axis (the absolute value of the correlation) but also along the x-axis. The length of the x-axis depends on the number of transcripts identified in the HAT and LAT class. As the range of x-axis is markedly different across the two regions, the inset in Figure 6 illustrates the difference in the correlation for the same range. Since correlation length varies across all chromosomes, the inventors defined a metric that is the correlation length scaled by the total number of genes on each chromosome. For each sample, they then construct the 2-dimensional matrices LHAT and LLAT, each for HAT and LAT class. The matrices comprise, for a class of genes (either HAT or LAT), the scaled correlation length for each chromosome and sample (i.e. a vector for each sample, comprising scaled correlation lengths for each of a plurality of chromosomes).

[0130] A theoretical proof of the concept of autocorrelation being linked to spatial gene expression distribution variance is illustrated on Figure 16, which illustrates the relationship between autocorrelation and variance in a univariate scenario. This considers a simple example with distributions with a mean centered at 500. All these curves have different variances to illustrate the activation of suppressed genes resulting from opening up of heterochromatin. By performing autocorrelation on these curves we can see an increasing trend on the maximum significant lag (or significant correlationlength), indicated here by the red arrow that shows the point where the autocorrelation function intersects the 95% confidence interval. In a multivariate scenario (multiple distributions for each gene), the points in each distribution are correlated differently such that it is not possible to measure the autocorrelation length on the “cumulative” distribution. The inventors therefore came up with a way to aggregate different distributions preserving intra and inter distribution characteristics. This was done by sweeping the multivariate distribution using a pair of genes with varying genomic distance between them as illustrated on Figure 4. For example, delta = 0 accounts for all possible pairs that are immediate neighbours, delta = 1 accounts for all pairs of genes with one gene in between them, etc. for all deltas up to n-1 . Having collected all pairs of genes with different separation distance in their corresponding buckets, it is then possible to calculate the distance correlation for each bucket. As explained above, the distance correlation function is similar to other correlation functions, except that it additionally captures non-linear relationships. This distribution is then passed to the autocorrelation function to find out the maximum significant lag (i.e. calculating the value of the autocorrelation as a function of lag and identifying the value of the lag at which the autocorrelation is no longer significant at a chosen confidence level). The end to end pipeline is illustrated on Figure 17, which shows that, for each chromosome individually, RNA-seq data (gene expression as a function of genomic coordinates) is denoised, then binned (each bin comprising genes that are separated by the same number of genes along the chromosome), then a distance correlation is calculated on each bin, and finally an autocorrelation of the distance correlation is obtained, from which the correlation length I* is determined and used as a proxy for open chromatin length. In short typical RNAseq data is passed through this pipeline to obtain a I* value per chromosome per sample. This is then normalized by the number of genes per chromosome. In some of the results shown below, this value is then multiplied by 1000 to obtain a metric that has values in integer ranges (therefore easier to handle) and that can be interpreted as the intrachromosomal correlation length per 1000 genes. Example 2 shows a systematic validation of this method.

[0131] Thus, the data in this example shows that there exists a spatial correlation between genes and their transcript abundance.2 - Systematic Validation in natural aqinq,aqinq, and

[0132] In this example, the methods described in Example 1 are used to obtain metrics that are used to characterise samples to verify that these have the potential to identify differences between samples and treatment. In particular, this example demonstrates the use of the algorithm described in Example 1 on two datasets - LINE-1 from Della Valle et al. and the data from Fleischer et al. The objective here is to first validate whether I* has discriminating ability and then try to dig deeper into discriminating drivers to understand what I* is capturing. Also during this process, the inventors hoped to be able to highlight similarities and differences between natural and pathological aging.Methods

[0133] See Example 1. As explained above, the correlation length (I* also referred to as intrachromosomal correlation length) can be calculated using all transcripts, using high expression transcripts (£HAT) or using low expression transcripts (£LAT). Unless indicated otherwise, the results shown are obtained using the correlation length for high expression transcripts (£HAT). Thus, unless indicated otherwise, I* corresponds to £HAT. Figures 9, 10, and 12A shows results for the correlation length calculated without multiplication by 1000. All other figures showing I* values use scales where the correlation length is multiplied by 1000.

[0134] Identification of high and low abundance transcripts. For the data on Figures 8, 9, 10 and 12A, spectral clustering was used as explained above. For the data on Figures 11 D, 12C, 12D and 18 onwards, DBSCAN (Density-based spatial clustering of applications with noise) was used in addition to spectral clustering, to remove outliers (denoising) prior to spectral clustering. DBSCAN can also be used instead of spectral clustering for identifying high vs low abundance transcripts. DBSCAN is advantageously robust to changes in the way that the RNA-seq data was processed, at least in part because it is deterministic rather than stochastic.

[0135] Clustering of samples. To achieve optimal results, we used correlation length obtained from the transcripts in the high abundance class which led to classification of samples accurately. Classification was performed using hierarchical clustering with average linkage and Manhattan distance metric. Our approach utilized clustering to divide abundance into high and low classes (see Figure 5), leading to an improved accuracy of clustering. However, when performing clustering based on the correlation length derived from all transcripts or transcripts with low abundance, the inventors found that the classification of diseased and treated samples from the wildtype controlwas suboptimal, as shown in Figure 10b,c. While it is true that excluding low abundance transcripts from analysis can result in a loss of important biological information, such as those encoding transcription factors or signalling molecules, restricting our analysis to high abundance transcripts increased the accuracy and reliability of our clustering algorithm at the global or chromosomal level. Further, the inventors performed extensive validation of the effects of the parameters of the clustering on the results of the method. In particular, the applied spectral clustering as described above to data from GTEx (gtexportal.org / home / ) comprising age stratified samples in different tissues, using different values for the parameter gamma, influencing the cutoff for selection of high expression transcripts. The high expression genes were then used to calculate I* as explained above, and these were then used to cluster the samples using hierarchical clustering. The results of this are show on Figure 20 for an example tissue (here liver, Fig. 20A shows results for gamma=3.5, Fig. 20B shows results for gamma=2.5 and Fig. 20C shows results for gamma=1 ), showing that the clustering is robust since values of gamma between 3.5 and 1 still produced good clusters (and better separation than when applying PCA to gene expression data, data not shown), although the more low expression transcripts are included (lower gamma) the more interleaving of clusters can be observed. This shows that the method is robust and also provides a method to assess clustering parameters used to select high expression transcripts for calculation of I* enabling a choice of parameters leading to most informative I*

[0136] Note that the above steps were performed in addition to a standard processing pipeline involving appropriate quality control, normalization, and statistical analysis that the authors in their respective papers have incorporated as a pre-processing step to improve the accuracy and reliability of the data.

[0137] To calculate the cluster distance used in Figure 9A, the inventors use the feature matrix LHAT and construct the cluster distance matrix (Cp) based on the Manhattan distance between each pair of samples using wildtype as a reference to compare the similarity between two experimental phenotype. This is then averaged across chromosomes. For 3 replicates (L1 dataset) in each phenotype, this will result in a 3 x 3 matrix (average distance between LHAT for each chromosome for each pair of one treated sample (columns or rows) and one wild type sample (rows or columns)) which can then be used to plot the average cluster distance along with an error bar (standard deviation obtained from the replicate information). Note that other distancemetrics are possible but Manhattan distances were found to work best. Without wishing to be bound by theory, this is believed to be because the Manhattan distance uses the range of values more effectively than e,g. Eucledian distances.Results

[0138] To test whether the metrics developed in Example 1 have potential to identify difference between samples or treatments, the inventors use the measured correlation length (by experiment and by chromosome) for the HAT class. Note that the clustering performed based on correlation length derived using all transcripts or transcripts with low abundance resulted in sub-optimal clustering of samples as explored further below. The data indicates that the transcripts in the LAT class may act as a uniform background noise and focusing only on the high abundance transcripts does not seem to drop key signals. Figure 8 illustrates a hierarchical clustering of the feature matrix LHAT in which replicates of samples are grouped together at the firs level of merging. At the next level of merging, it is worth noting that the “scrambled treated” (i.e. treated with an ASO with scrambled sequence - negative control for the ASO treatment) and “no treatment” samples are merged together, and confirm the existence of similarities between the diseased samples. This is because these samples did not receive Line-1 antisense oligonucleotides (ASO) treatment and can be associated with each other as controls to the treatment. The clustering of LINE-1 ASO treated samples with wild-type samples is consistent with experimental observations that a LINE1 ASO can help to bring diseased cells closer to the state of healthy cells (Della Valle et al., 2022). Depletion of L1 RNA from diseased samples patients using ASOs restored heterochromatin histone marks and reversed DNA methylation age further pointing towards a potential modification in the underlying chromatin structure with intervention (Della Valle et al., 2022).

[0139] In the previous sections, we discuss the use of correlation length as a metric for segregating cells in different states (such as diseased (e.g. with premature aging disorders), treated, and samples of different ages). To ensure accurate analysis and avoid errors related to sequencing depth, we employ a clustering algorithm for signal extraction. The inventors further evaluated the effectiveness of hierarchical clustering by analysing the correlation length of genes in the HAT class, LAT class, and all transcripts depicted in Figures 10A-C, respectively. These results outline the significance of focusing on the key targets that contribute to the classification andeliminating the noise. Indeed, clustering using correlation length on HAT genes performed significantly better than using either LAT genes only or all genes (although both of these still worked to some extent - the data shows that removing the LAT genes is beneficial in focusing on key informative signal). To further demonstrate the robustness of the methodology and highlight the significance of genome wide analysis, they conducted sample clustering analysis after removing chromosome 6 (Figure 10D) and 1 (Figure 10E) from the analysis. These two chromosomes can potentially have a significant impact on metrics characterising changes in chromatin structure as a result of disorders or interventions, because of their high density of histone genes.

[0140] The inventors further propose that EHAT can be integrated with P1 (the first principal component obtained by PCA of the complete transcript level dataset) as illustrated in Figure 10F to complement an existing workflow. This integration leads to a relatively improved clustering performance, as compared to using only PCA (see Figure 11 A).

[0141] The inventors also compared their method to existing techniques by presenting a clustering based on the principal components of the transcript abundance data. Figure 11A depicts the clustering of samples using the feature matrix (P1 ), which is constructed using the firs principal component of transcript abundance, for each chromosome. They observe that the HGPS NT and HGPS SCR (non-treated and scramble treated) samples merge with each other rather than merging with their own type which also hold true for diseased Werner Syndrome replicates. Since one principal component accounts for 44% - 92% of the variance across all chromosomes, it is advantageous to have additional principal components. Figure 11 B presents a more concrete comparison with a feature matrix comprising 10 principal components, averaged across all chromosomes, covering ~ 99.9% of the variance and it is evident that most of the replicates in the diseased state do not merge well. Furthermore, we observed that chromosome X influence the first principal component and removal of this resulted in further deterioration of the clustering. This shows that while through our approach, one can identify which chromosomes are most affected or are important, as stated earlier, removal of these chromosomes does not adversely affect the output. Therefore, the data shows that the proposed approach extracts the patterns of changes in the structure across all chromosomes and is highly robust. A Mann-Whitney ll-test did not identify a significant difference between the distribution of mean I* across chromosomes in the WT and L1 -ASO treated cells in the LINE-1 dataset. However,when looking at individual chromosomes, the inventors found a strong trend on chromosome 6 (see Figure 11 D; Mann-Whitney U test indicates significant difference between WT and L1 -ASO vs non-treated (NT) and SCR), indicating that if I* is capturing chromatin structural changes, then pathological aging does not affect uniformly all chromosomes. The inventors hypothesised that the importance of chr 6 may be related to the fact that the largest cluster of histone genes (~80%) HIST1 , is located on Chr 6, and Chromosome 6 is also the house of FOXO3 gene that defines an aging hub. The importance of chr6 was also confirmed by gene expression analysis in the LINE-1 dataset. Among the top 200 differentially expressed genes identified by comparing diseased vs WT (i.e. HGPS-NT vs WT and WRN-NT vs WT), chromosome 6 has the highest fraction of differentially expressed genes. This is also true overall, see Table 1 . Thus, pathological aging seems to induce varying changes among chromosomes, with Chr 6 standing out prominently. The data on Figure 11 D shows that ASO treatment demonstrates enhanced effectiveness in aligning cell states of HGPS and WRN more closely to WT.

[0142] Figure 12A highlights the hierarchical clustering for the Fleischer dataset. The Fleischer dataset is an aging dataset and contains samples from newborns to adult senior people (>90 years). The objective of using this dataset is to show that theproposed measure works even in natural aging situation. Further, as the common underlying effect of pathological as well as natural aging is changes in chromatin (mainly loss of heterochromatin), this shows that the proposed measure acts as “lens" into the chromatin dynamics. However, the age in Fleischer dataset is continuous number and so the ages were bucketed so that there are sufficient samples per “category”. The 0-20, 21 -40, etc is an arbitrary choice and one could pre-process it as 0-25, 26-50 and so on as well so long as there are sufficient samples in each bucket. Because of this arbitrary bucketing, and the fact that categories reflect chronological, there can be mismatch between the clustering and the age brackets (the clustering is based on RNA-seq data which shows cellular information). The data shows that the younger groups (0-20 and 21 -40 age bracket) merge together first. Concurrently, on the other side we see the elder population merging 61 - 80 age group with that of 81 and beyond. It is worthwhile to note here that the presence of a small group of senior population that directly merges with the 61 - 80 age bracket indicating that the algorithm learns nuances (all of these samples are early 80s) and is not biased exclusively by the age buckets. In other words, the clustering shows a portion of the senior population (81 -100 age bracket) merging with 61 -80 bucket. This group are the early 80s senior people who are likely grouped with the 61 -80 bucket primarily because they are cellularly younger meaning healthier. This validates that the method is not deriving signals from the way the data is age-bucketed. If it was then there would not have been this "cross-contamination".

[0143] The middle age group (41 - 60 bracket), representing a transition state between the two sub-groups stands distinctively and merges in the end due to its heterogeneity. On the other hand, clustering by PCA (Figure 12B) shows extremely poor performance, further reinforcing that although the principal components capture the variances, they neither have the potential to capture the underlying cause nor the subtle nuances resulting from them. Figure 12C shows boxplots of the mean I* across chromosomes for each age bracket in the Fleischer dataset. This shows that there is an overall increasing trend of I* across age buckets, which was shown to be statistically significant (Mann Kendall test, trend=increasing, p=0.027). Note that the inventors tested in this and other datasets the use of the mean I* across chromosomes and sum of I* across chromosomes, both of which were found too correlate with natural aging. Figure 12D shows the same data for discrete ages rather than age bins, with a fitted linear regression model (linear regression coefficients.0173, p=2.01769*10’10). The showsthat the metric described herein increases significantly with age, and can therefore be used as an indicator of natural aging.

[0144] The data above show that f* demonstrates sensitivity in detecting and capturing chromosomal structural changes. The data further show that noticeable variations exist in the trajectory of DNA damage between natural and pathological aging. In particular, natural aging seems to impact all chromosomes, unlike pathological aging. The data further show that L1 -treated ASO appears to bring diseased cell lines closer to the Wild Type.

[0145] Based on this work on the LINE1 and Fleischer dataset validating the I* metric, the inventors set out to test the robustness of I* and see if we can get more insight on partial reprogramming experiments using OSKM, using a dataset from Gill et al., 2022. These reprogramming experiments use dermal fibroblast from three middle-aged donors (38,53 and 53) with epigenetic age of 49, 45 and 55 (based on methylation patterns). Cells received the OSKM treatment along with doxycycline for different lengths of time and cells were sorted at day 10,13,15 and 17. The ones that have SSEA4 positive, CD13 negative markers were labeled as ‘transient reprogramming intermediate’ and the other were labeled as ‘failing to transiently reprogram intermediate’: CD13 positive, SSEA4 negative. The authors also performed complete reprogramming using the Sendai reprogramming protocol (this data was not used here). RNAseq analysis was performed on these cells, see Fig. 1 e of Gill et al. 2022. This showed the restoration of original identity of the old fibroblast cells. In particular, Fig. 1 e of Gill et al. 2022 shows a PCA plot performed on full reprogramming dataset + the transient reprogramming all genes showing a reprogramming trajectory along the first principal component PC1 from old fibroblasts to iPSC and other samples are clustered along this trajectory. Interestingly, transiently reprogrammed samples clustered at the beginning of this trajectory showing that these samples once again transcriptionally resemble fibroblasts rather than reprogramming intermediates or iPSCs. A similar trajectory was also seen for partial reprogramming dataset but now clearly showing 5 distinct states. The inventors set out to analyse this data using the I* metric. The started with clustering of samples based on the days of treatment and wanted to see if I* can predict the transcriptional / epigenetic state of the cell. The results of this are shown on Figure 18A, which show clustering using I* This shows that in all the days of treatment, transient reprogramming intermediate and iPSC cluster together for all days which makes sense since they are the closest. Additionally, the cells thatare intermediate failing to transiently reprogram and negative control intermediate cluster together which also makes sense since they represent their own state in Gill et al. 2022. Next, the inventors plotted the absolute I* (mean value across chromosomes) for different states (Figure 18B, where states are sorted based on day 13 since that’s the one that is found to be the most optimum day of treatment by Gill et. al.) This shows that the transiently reprogrammed state is the most closed state, iPSC is the most open state, transient reprogramming intermediate is the next more open state, and fibroblast is somewhere in the middle of the two. The behaviour shown on Fig. 18B can be understood with a simple picture (or model) for how euchromatin state varies over time. It starts with iPSC which has a lot of euchromatin and slowly with differentiation process, the heterochromatin starts to appear and euchromatin goes down. Then it goes back up again because of heterochromatin loss that happens during aging. When looking at individual chromosomes (data not shown), some chromosomes have higher I* for iPSC, some have lower, but on an average the iPSC have a higher I* compared to any other state. By performing differential gene expression analysis it is possible to identify the most important chromosome which has the most differentially expressed genes and then look at the variation of I* on that chromosome. The results of this are shown in Table 2, which shows that chromosome 7 stood out at the top for transiently reprogrammed vs iPSC and close second for fibroblast vs iPSC. Chromosome 7 comprises the elastin gene which codes for elastin, a major component of elastic fibers, which are a major component of the extra cellular matrix. Figure 18C shows the I* on chr 7, by state on day 13. This shows that the I* metric strongly captured the expected trends: iPSC and transient intermediate is probably the most open and hence a larger I* whereas the transiently reprogrammed which would have recovered its heterochromatin following OSKM has the lowest I*, and Fib is somewhere in the middle. This shows that f* is able to capture different states in the OSKM data based on the underlying epigenetic changes to chromatin.

[0146] Next, the inventors set out to investigate whether the new metric could be useful in the context of a chemical screen performed on HUVEC cells. The screen tested 18 chemicals and 1 control (DMSO vehicle) in 2 cell lines (HUVEC B and HUVEC D) (156 samples in total including multiple biological replicates per condition). The 18 chemicals were carefully selected spanned among different categories including Oxidative Stress / inflammation, Protein homeostasis, Epigenetic modifiers, and kinase / protease inhibitors. The cells were exposed to the chemicals from Day 0 to day 3, then allowed to recover until Day 7, at which point they were harvested for RNAseq. The similarity distance between each treatment and DMSO was calculated using gene expression, and using I*, separately for each cell lines. The same two compounds were found to deviate the most from DMSO in both cell lines using both gene expression and I*. Analysis of methylation data for these two compounds using 3 different methylation clocks (AgeSkinBloodClock, DNAm Age clock, DNAm AgeHannum clock) showed that coumpound A rejuvenated the clock in HUVEC B significantly in all 3 clocks, and in 2 out of 3 clocks in HUVEC D. For compound B, the rejuvenation was not as prominent as for compound A and was only seen in DNAmAge HUVEC B panel.

[0147] The inventors then looked at the I* in both treatments, compared to DMSO. The data on Figure 19A shows significant change in I* values in chromosome 6 brought by the chemicals. On Chromosome 6, compound A and compound B prominently reducethe I* values in HUVEC B, indicating the restoration of heterochromatin. No effect was observed on chr 6 in HUVEC D. Compound A was also independently tested using methylation and gene expression changes, to confirm its rejuvenation effect. This showed that compound A decreases H3K27me3 and rescues age related transcriptome changes significantly (reducing log fold changes in expression observed with aging). By contrast, another compound, compound C, aggravated this age-related transcriptome changes in a significant manner. Looking at I* values for compound C treated sample they inventors found that on chromosome 6, compound C increases I* value (statistically significant, p<0.05 Mann Whitney U test, in HUVEC B), compared to DMSO sample, in both HUVEC B and HUVEC D cell lines (Figure 19B). Thus, at the individual chromosome level, chromosome 6 showed significant differences for all three chemicals introduced here, in the expected direction which is decreased I* values in compound A and compound B and increased I* value in Compound C compared to DMSO, at least in HUVEC B cell line. The inventors then set out to further investigate the difference in behaviour between the two cell lines.

[0148] The chemical screening dataset used also includes a range of HUVEC cells that are experiencing passaging, to mimic the effects of aging. This can be used to investigate the in vitro aging effect in the HUVEC cells. Figure 19C and Figure 19D show the distribution of I* (mean I* across chromosomes and variation across chromosomes) for each passage from 5 to 21 , young to old, for the HUVEC B and HUVEC D cells, respectively. As in the Fleischer dataset, Figure 19C shows a noticeable trend in the increasing of the I* in HUVEC B passing the Mann Kendall test, reaching a plateau when the cells are getting relatively old (after passage 17). However, this is not observed in HUVEC D, where there is no systematic trend of I*. The inventors believe that this can provide an explanation for why the chemical intervention are more pronounced in HUVEC B than HUVEC D. This data shows that HUVEC B and HUVEC D show different patterns, indicating that there is a donor effect in chromatin structure associated with aging and the effects of chemicals on this. Nevertheless, the data also shows that the approach described herein can be used to identify drugs that have an effect on rejuvenation, as validated by orthogonal (methylation and gene expression based) assays. The data also show that chr 6 seems to be an important chromosome in this process.Conclusions

[0149] The study described above reveals the existence of long-range spatial correlation between genes based on the underlying alterations in chromatin structure due to epigenetic modifications. From an enzymatic point of view, the process of gene expression entails the binding of RNA polymerase to a chromatin segment which must be accessible before this recruitment occurs. Changes to chromatin accessibility mediated by histone modification and DNA methylation caused by diseases or treatments alters the regulation of gene expression. Despite the richness and complexity of underlying mechanism of gene-regulation, the work described here shows that a signature of correlation between gene expression based on their chromosomal location can be formally quantified using distance correlation of pairwise transcript abundance. The correlation length measured across different chromosomes and samples proved to be a useful measure to separate healthy and diseased state, and different aged states, as well as a useful marker of rejuvenation and reprogramming, which were shown to be more robust that traditional methods like PCA. This analysis suggests that there exists a long-scale correlation between geneexpression driven by an intricate interplay between the local variations in chromatin structure and gene activity.References

[0150] Boulesteix, A.-L., Strimmer, K.: Partial least squares: a versatile tool for the analysis of high-dimensional genomic data. Briefings in bioinformatics 8(1 ), 32-44 (2007).

[0151] Li, L., Yin, X.: Sliced inverse regression with regularizations. Biometrics 64(1 ), 124-131 (2008).

[0152] Yeung, K.Y., Ruzzo, W.L.: Principal component analysis for clustering gene expression data. Bioinformatics 17(9), 763-774 (2001 ).

[0153] Choy, C.T., Wong, C.H., Chan, S.L.: Infer related genes from large scale gene expression dataset with embedding. bioRxiv, 362848 (2018).

[0154] Van Dam, S., Vosa, U., van der Graaf, A., Franke, L., de Magalhaes, J.P.: Gene co-expression analysis for functional classification and gene-disease predictions. Briefing in bioinformatics 19(4), 575-592 (2018).

[0155] Li, B., Carey, M., Workman, J.L.: The role of chromatin during transcription. Cell 128(4), 707-719 (2007).

[0156] Venkatesh, S., Workman, J.L.: Histone exchange, chromatin structure and the regulation of transcription. Nature reviews Molecular cell biology 16(3), 178-189 (2015).

[0157] Lai, W.K., Pugh, B.F.: Understanding nucleosome dynamics and their links to gene expression and dna replication. Nature reviews Molecular cell biology 18(9), 548- 562 (2017).

[0158] Debes, C., Papadakis, A., Gronke, S., Karalay, O., Tain, L., Mizi,A. , Nakamura, S., Hahn, O., Weigelt, C., Josipovic, N., et al.: Aging-associated changes in transcriptional elongation influence metazoan longevity. bioRxiv, 719864 (2019).

[0159] Tsurumi, A., Li, W.: Global heterochromatin loss: a unifying theory of aging? Epigenetics 7(7), 680-688 (2012).

[0160] Lee, J.-H., Kim, E.W., Croteau, D.L., Bohr, V.A.: Heterochromatin: an epigenetic point of view in aging. Experimental & Molecular Medicine 52(9), 1466-1474 (2020)

[0161] Fleischer, J.G., Schulte, R., Tsai, H.H., Tyagi, S., Ibarra, A., Shokhirev, M.N., Huang, L., Hetzer, M.W., Navlakha, S.: Predicting age from the transcriptome of human dermal fibroblasts. Genome biology 19(1 ), 1-8 (2018).

[0162] Szekely, G.J., Rizzo, M.L., Bakirov, N.K.: Measuring and testing dependence by correlation of distances. Annals of statistics 35(6), 2769-2794 (2007).

[0163] V. Satopaa, J. Albrecht, D. Irwin and B. Raghava", "Finding “a "Kneedie" in a Haystack: Detecting Knee Points in System Behavior," 201131 st International Conference on Distributed Computing Systems Workshops, Minneapolis, MN, USA, 2011 , pp. 166-171.

[0164] Della Valle F, Reddy P, Yamamoto M, Liu P, Saera-Vila A, Bensaddek D, Zhang H, Prieto Martinez J, Abassi L, Celii M, Ocampo A, Nunez Delicado E, Mangiavacchi A, Aiese Cigliano R, Rodriguez Esteban C, Horvath S, Izpisua Belmonte JC, Orlando V. LINE-1 RNA causes heterochromatin erosion and is a target for amelioration of senescent phenotypes in progeroid syndromes. Sci Transl Med. 2022 Aug 10;14(657):6057.

[0165] P. Scaffidi, T. Misteli, Reversal of the cellular phenotype in the premature aging disease Hutchinson-Gilford progeria syndrome. Nat. Med. 11 , 440-445 (2005).

[0166] F. G. Osorio, C. L. Navarro, J. Cadi anos, I. C. Lopez-Mejia, P. M. Quiros, C. Bartoli, J. Rivera, J. Tazi, G. Guzman, I. Varela, D. Depetris, F. de Carlos, J. Cobo, V. Andres, A. De Sandre-Giovannoli, J. M. P. Freije, N. Levy, C. Lopez-Otin, Splicing-directed therapy in a new mouse model of human accelerated aging. Sci. Transl. Med. 3, 106 (2011 ).

[0167] A. Ocampo, P. Reddy, P. Martinez-Redondo, A. Platero-Luengo, F. Hatanaka, T. Hishida, M. Li, D. Lam, M. Kurita, E. Beyret, T. Araoka, E. Vazquez-Ferrer, D. Donoso, J. L. Roman, J. Xu, C. Rodriguez Esteban, G. Nunez, E. Nunez Delicado, J. M. Campistol, I. Guillen, P. Guillen, J. C. Izpisua Belmonte, In vivo amelioration of age- associated hallmarks by partial reprogramming. Cell 167, 1719-1733. (2016).

[0168] W. Zhang, J. Li, K. Suzuki, J. Qu, P. Wang, J. Zhou, X. Liu, R. Ren, X. Xu, A. Ocampo, T. Yuan, J. Yang, Y. Li, L. Shi, D. Guan, H. Pan, S. Duan, Z. Ding, M. Li, F. Yi, R. Bai, Y. Wang, C. Chen, F. Yang, X. Li, Z. Wang, E. Aizawa, A. Goebl, R. D. Soligalla, P. Reddy, C. R. Esteban, F. Tang, G.-H. Liu, J. C. I. Belmonte, Aging stem cells. A Werner syndrome stem cell model unveils heterochromatin alterations as a driver of human aging. Science 348, 1160-1163 (2015).

[0169] Gill D, Parry A, Santos F, Okkenhaug H, Todd CD, Hernando-Herraez I, Stubbs TM, Milagre I, Reik W. Multi-omic rejuvenation of human cells by maturation phase transient reprogramming. Elife. 2022 Apr 8;11 :e71624.

[0170] All references cited herein are incorporated herein by reference in their entirety and for all purposes to the same extent as if each individual publication or patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety.

[0171] The specific embodiments described herein are offered by way of example, not by way of limitation. Various modifications and variations of the described compositions, methods, and uses of the technology will be apparent to those skilled in the art without departing from the scope and spirit of the technology as described. Any sub-titles herein are included for convenience only and are not to be construed as limiting the disclosure in any way. Unless context dictates otherwise, the descriptions and definitions of the features set out above are not limited to any particular aspect or embodiment of the invention and apply equally to all aspects and embodiments which are described.

[0172] The methods of any embodiments described herein may be provided as computer programs or as computer program products or computer readable media carrying a computer program which is arranged, when run on a computer, to perform the method(s) described above.

[0173] Throughout the specification and claims, the following terms take the meanings explicitly associated herein, unless the context clearly dictates otherwise. The phrase “in one embodiment” as used herein does not necessarily refer to the same embodiment, though it may. Furthermore, the phrase “in another embodiment” as used herein does not necessarily refer to a different embodiment, although it may. Thus, as described below, various embodiments of the invention may be readily combined, without departing from the scope or spirit of the invention.

[0174] It must be noted that, as used in the specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from “about” one particular value, and / or to “about” another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by the use of the antecedent “about,” it will be understood that the particular value forms another embodiment. The term “about” in relation to a numerical value is optional and means for example + / - 10%.

[0175] Throughout this specification, including the claims which follow, unless the context requires otherwise, the word “comprise” and “include”, and variations such as “comprises”, “comprising”, and “including” will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps.

[0176] Other aspects and embodiments of the invention provide the aspects and embodiments described above with the term “comprising” replaced by the term “consisting of” or “consisting essentially of”, unless the context dictates otherwise.

[0177] “and / or” where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. For example “A and / or B” is to be taken as specific disclosure of each of (i) A, (ii) B and (iii) A and B, just as if each is set out individually herein.

[0178] The features disclosed in the foregoing description, or in the following claims, or in the accompanying drawings, expressed in their specific forms or in terms of a means for performing the disclosed function, or a method or process for obtaining the disclosed results, as appropriate, may, separately, or in any combination of such features, be utilised for realising the invention in diverse forms thereof.

Claims

CLAIMS1 . A method of determining the chromatin state of a sample comprising: receiving, by a processor, transcript abundance data for a plurality of transcripts from the sample; determining, by said processor, for each of a plurality of genomic distance bins, a correlation metric between transcript abundance for pairs of transcripts that are separated by a genomic distance within the respective bin; and generating, by said processor, a value of a biomarker metric derived from the correlation metric for the plurality of bins, the biomarker metric characterising the range of genomic distance at which pairs of transcripts have correlated expression, wherein samples with different chromatin states have different values of the biomarker metric.

2. The method of claim 1 , wherein the transcript abundance data comprise a normalised abundance for each of the plurality of transcripts, optionally wherein the normalised abundance for a transcript is obtained by dividing the abundance of a transcript by the mean abundance for all transcripts from the sample thereby obtaining a scaled abundance and / or wherein the normalised abundance for a transcript is obtained by dividing the abundance or scaled abundance for the transcript by the sum of abundances or scaled abundances for all transcripts from the sample thereby obtaining a fractional abundance.

3. The method of any preceding claim, wherein each of the plurality of transcripts is associated with a genomic locus in a reference genome, optionally a human reference genome.

4. The method of any preceding claim, wherein the plurality of transcripts are high abundance transcripts, wherein high abundance transcripts are obtained by identifying, in transcript abundance data for a plurality of transcripts, a first group of transcripts and a second group of transcripts, the first group of transcripts having higher abundance in the sample than the second group of transcripts; and / or wherein the method comprises: identifying, by said processor, using the transcript abundance data, a first group of transcripts and a second group of transcripts, the first group of transcripts having higher abundance in the sample than the second group oftranscripts, and selecting, by said processor, the first group of transcripts before determining the correlation metric for the plurality of bins.

5. The method of claim 4, wherein identifying first and second group of transcripts is performed using a data driven method.

6. The method of claim 5, wherein the data driven method comprises: a clustering method, and a curve characteristics-based algorithms applied to a data series comprising the ranked abundance of transcripts as a function of rank, optionally wherein the curve-characteristics based algorithm is the kneedie algorithm; and / or wherein the data driven method comprises a first data driven method used for outlier detection and a second data driven method used to identify the first and second groups of transcripts from a set of transcripts obtained as a result of outlier detection, optionally wherein the first and / or second data driven methods are clustering methods.

7. The method of claim 6, wherein the clustering method is selected from spectral clustering, k-means, k-medoid, and density-based spatial clustering of applications with noise (DBSCAN), optionally wherein the spectral clustering is performed using a kernel selected from: Laplacian affinity, Radial Basis Function or nearest neighbour.

8. The method of any preceding claim, wherein the genomic distance bins are based on the chromosomal positions of the transcripts, where transcripts in a bin are associated with genes that are the same number of genes apart from each other.

9. The method of any preceding claim, wherein the genomic distance bins each comprise pairs of transcripts associated with genes that are n genes apart on the same chromosome, where n is a natural number or range of natural numbers, where n is different for each bin.

10. The method of claim 9, wherein the genomic distance bins comprise a first bin comprising pairs of transcripts associated with genes that n1 genes apart on the same chromosome, and a second bin comprising pairs of transcripts that are associated with genes that are n2 genes apart on the same chromosome, optionally wherein n1 is 0 and n2 is not 0.

11. The method of any preceding claim, wherein the correlation metric is a non-linear correlation coefficient, optionally wherein the correlation metric is the distance correlation.

12. The method of claim 11 , wherein the correlation metric for a bin is calculated aswhere:(x, y) are vectors of transcript abundance for all possible pairs of transcripts in the bin; andV is the empirical distance covariance for the transcripts in the bin.

13. The method of claim 12, wherein the empirical distance is calculated using the following equation:where D are matrices of all centred Euclidian pairwise distances between transcript abundance for transcripts in the bin, centred by the row mean and column mean.

14. The method of any preceding claim, wherein the biomarker metric derived from the correlation metric is a value derived from the auto-correlation of the correlation metric.

15. The method of claim 14, wherein the autocorrelation of the correlation metric is calculated as a function of a lag in the same units as the units of genomic distance associated with the bins.

16. The method of claim 14 or claim 15, wherein the autocorrelation of the correlation metric is calculated as a function of a lag in number of genes apart on the chromosome.

17. The method of any of claims 14 to 16, wherein the biomarker metric derived from the correlation metric is a value proportional to a lag at which the auto-correlation of the correlation metric falls below a predetermined confidence limit, and / or wherein the biomarker metric derived from the correlation metric is a value proportional to a lag at which the auto-correlation of the correlation metric is no longer significantly different from zero at a chosen confidence level, optionally wherein the confidence limit or level is selected between 90% and 99%, and / or wherein the confidence level or limit is 90%, 91 %, 92%, 93%, 94%, 05%, 96%, 97%, 98%, or 99% and / or wherein the correlation metric is the lag at which the autocorrelation of the correlation metric falls below a predetermined confidence limit and / or the lag at which the auto-correlation of the correlation metric is no longer significantly different from zero at a chosen confidence level, in each case multiplied by 1000.

18. The method of any preceding claim, wherein the plurality of transcripts are located on the same chromosome, or wherein the plurality of transcripts are located on a plurality of chromosomes, and the correlation metric and biomarker value are determined separately for each chromosome of the plurality of chromosomes using the transcripts located on the respective chromosome.

19. The method of any preceding claim, wherein the plurality of transcripts are located on a plurality of chromosomes, and the method comprises selecting, by said processor, a plurality of transcripts located on the same chromosome prior to the processor determining the correlation metric, optionally wherein the method comprises repeating the steps of selecting a plurality of transcripts located on the same chromosome for a further chromosome or a plurality of further chromosomes and determining the correlation metric and biomarker value for the further chromosome or plurality of further chromosomes, optionally wherein the sample is a human sample and the chromosome or further chromosome is chromosome 6 or chromosome 7.

20. The method of any preceding claim, wherein the transcript abundance data have been obtained from whole transcriptome RNA sequence data, and / or wherein the transcript abundance data comprises abundance data from each transcript that is detectable in the sample using a transcriptome wide transcriptom ic analysis technique, such as RNA sequencing.21 . The method of any preceding claim, wherein the biomarker metric is generated from transcript abundance data from a plurality of transcripts located on the same chromosome, and the value of the biomarker metric is scaled by the number of transcripts on the chromosome, and / or wherein the transcript abundance data comprise data from a plurality of transcripts located on the same chromosome and generating, by said processor, the value of the biomarker metric comprises the processor scaling the generated value by the number of transcripts on the chromosome, and optionally determining the number of transcripts located on the chromosome.

22. The method of any preceding claim, further comprising generating, by said processor, a value of the biomarker metric for each of a plurality of samples and clustering the values obtained.

23. The method of claim 22, wherein said clustering is performed using hierarchical clustering and a Manhattan distance.

24. The method of claim 22 or claim 23, further comprising determining, by said processor, an intercluster distance based on the values of the biomarker metrics for samples in at least one pair of clusters, optionally wherein the pair of cluster comprises a cluster comprising normal samples and a cluster comprising test samples and / or wherein the method comprises comparing samples in a first cluster and samples in a second cluster using the intercluster distance between the first cluster and the second cluster or between the first cluster and a third cluster and between the second cluster and the third cluster.

25. The method of claim 24, wherein the intercluster distance is determined by calculating for each pair of samples, the distance between the biomarker values for each pair of samples comprising a sample from each of the pair of clusters, thereby obtaining a plurality of distances, and obtaining a summarised distances from the plurality of distances, optionally wherein the distance is a Manhattan distance and / or the summarised distance is an average distance.

26. The method of any preceding claim, wherein the method comprises obtaining transcript abundance data from the sample, wherein obtaining transcript abundance data from the sample comprises processing RNA sequence data comprising RNA sequencing reads to determine an abundance for each of the plurality of transcripts based on the reads of the plurality of reads that map to the respective transcripts, and / or wherein obtaining transcript abundance data from the sample comprises obtaining RNA sequencing reads from the sample by RNA sequencing.

27. A method of observing the effects of one or more test agents on aging or disease in cells, the method comprising: combining the test agent(s) with cells;Obtaining transcript abundance data from a sample of the cells; determining the chromatin state of the sample using the method of any of claims 1 to 26; andComparing the value of the biomarker value for the sample to one or more corresponding control values.

28. The method of claim 27, wherein comparing the biomarker value for the sample to one or more corresponding control values comprises clustering the biomarker value for the sample with a biomarker value determined using the method of any of claims 1 to 25 for one or more control samples.

29. The method of claim 27 or claim 28, wherein the control samples comprises one or more samples selected from: aged or diseased samples, healthy samples and samples treated with a rejuvenating agent.

30. A method of determining the age, rejuvenation state or reprogramming state of cells in a sample, the method comprising: obtaining transcript abundance data from the sample; determining the chromatin state of the sample using the method of any of claims 1 to 26; and comparing the value of the biomarker value for the sample to one or more corresponding control values, optionally wherein the control values are obtained from one or more samples with known ages, rejuvenation state or reprogramming state.31 . A method of treating a subject who has been diagnosed as having an aging-related disorder, the method comprising: administering to the subject a therapeutically effective amount of an agent that has been identified as effective in treating the aging related disorder, wherein the agent has been identified as effective in treating the aging-related disorder using a method according to any of claims 27 to 29.

32. The method of claim 31 , further comprising identifying the agent that is effective in treating the aging-related disorder using a method according to any of claims 27 to 29, wherein the one or more test agents include the agent that is identified as effective in treating the aging-related disorder.

33. The method of claim 31 or claim 32, wherein the control values include the value of a biomarker metric obtained using the method of any of claims 1 to 26 for at least one healthy sample and at least one sample comprising cells that have an aging-related disorder.

34. The method of any of claims 31 to 33, wherein the aging related disorder is HGPS or WRN.

35. A system comprising: a processor; and a computer readable medium comprising instructions that, when executed by the processor, cause the processor to perform the steps of any method described herein, such as a method according to any of claims 1 to 26.

36. One or more non-transitory computer readable media comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of any method described herein, such as a method according to any of claims 1 to 26.