Haplotype approach

By using pre-mRNA for phasing genome analysis, combined with RNA sequencing and other genomics technologies, the problem of identifying allelic skew and disease genetic signals in phasing genomes has been solved, enabling accurate identification of gene expression disorders and disease-related analysis.

CN114245828BActive Publication Date: 2026-01-02OXFORD UNIVERSITY INNOVATION LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080058031.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-22
Filing Date
2020-08-21
Publication Date
2026-01-02
Estimated Expiration
2040-08-21

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify allelic biases and disease-related genetic signals in the context of phased genomes, particularly the link between regulatory elements of non-coding genomes and gene expression, leading to uncertainty in disease genetics research.

Method used

Phased genomic analysis using pre-mRNA (polyA-mRNA) can identify allelic skewness. Combined with RNA sequencing and sequence-based genomics analysis, such as ATAC-seq and ChIP-seq, it can determine gene expression dysregulation and changes in regulatory elements, and associate gene expression with disease-related genetic signals.

Benefits of technology

This technology enables accurate identification of allele expression dysregulation in phased genomes, can be correlated with in vitro sequence variations of genes, provides candidate drug targets and genetic causes for disease-related genes, and improves the statistical robustness and accuracy of disease research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114245828B_ABST
    Figure CN114245828B_ABST
Patent Text Reader

Abstract

The present invention relates to the use of haplotype phasing to detect abnormal expression of genes that can be associated with a disease or condition. In particular, the present invention relates to a method of obtaining an indication of a dysregulation between the expression levels of at least two alleles of a gene in a target eukaryotic cell. The method comprises the steps of, for a plurality of genes from one or more target eukaryotic cells, (a) obtaining pre-mRNAs of at least two alleles of the same gene; and (b) determining a ratio (Ri,j) between the amounts of pre-mRNAs of one or more pairs of alleles (i,j) of the same gene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates to the use of haplotype phasing to detect abnormal expression of genes that can be associated with a disease or condition. In particular, the present invention relates to a method of obtaining an indication of a dysregulation between the expression levels of at least two alleles of a gene in a target eukaryotic cell. BACKGROUND

[0002] Common variant genome-wide association studies (GWAS) have identified tens of thousands of genetic variants associated with complex traits (e.g. body mass index, height) and susceptibility to common diseases (e.g. diabetes, heart disease, immune disorders and cancer). Despite the statistical significance of genetics, identifying the causal variants and genes underlying GWAS signals has been hindered by two major factors. Firstly, causal sequences are difficult to pinpoint precisely due to the extensive linkage of sequence variation within the genome; and secondly, the vast majority of variants, regardless of the specific GWAS in question, lie outside the coding genome.

[0003] The ability to interpret the non-coding genome is a priority for biomedical science. Over the past decade, our understanding of non-coding DNA has evolved from “junk” to complex regulatory circuits controlling the expression and function of the coding and non-coding transcriptomes. Central to this function are regulatory non-coding elements: promoters, enhancers and intron / exon boundaries. Although the existence of certain regulatory elements has been known for decades, their ubiquitous presence, with key tissue-specific roles in development and routine cellular processes, has only become clear in the last two decades. The large overlap between these elements and the genetics of common diseases has begun to enable understanding of the confounding non-coding distribution of GWAS signals.

[0004] In particular, there is a need for further research into the transcriptome (i.e. the collection of all RNA molecules in a cell or group of cells) to help identify genes that are abnormally expressed in a particular disease (or its associated biochemical pathways), and therefore determine candidate drug targets.

[0005] Abnormal expression of genes can be due to changes in elements that regulate gene expression. Changes in these elements can result in over-expression or under-expression of genes, changes in their temporal expression or their tissue specificity.

[0006] It is also known that such changes can result in a dysregulation in the expression levels of different alleles of the same gene within a single cell or group of cells; and that this dysregulation can also be associated with a particular disease or condition.

[0007] Determining the linkage between sequence variations of alleles (i.e. haplotypes) within a cell is known as “phasing”; and the finding of different expression levels of alleles is known as “allele skewing”.

[0008] Traditionally, biochemical and genetic analysis of most mRNAs has been performed on polyA + mRNAs. Total RNA is first extracted from cells; then passed down an oligo-dT column to bind polyA + mRNAs; and then the polyA + mRNAs are selectively eluted from the column. PolyA + RNA contains the coding sequence of the mRNA, i.e. the (non-coding) introns have been removed, and this is the form of mRNA that researchers are generally most interested in. There is also an advantage to using poly-dT columns in that rRNA, which is an unwanted contaminant and is expressed in large amounts in all cells, can be removed by this method.

[0009] The inventors have now recognised that additional information from the transcriptome can be obtained by using mRNAs that are not limited to polyA + RNAs. More specifically, they have recognised that if pre-mRNAs (e.g. polyA - mRNAs, which still contain introns) are used, natural variations in the genomic sequence that are found in introns and downstream sequences and removed in polyA + RNAs can be used to determine allelic skewing, and thus identify dysregulated genes. Importantly, when performed on phased genomic sequences, all sequence variants can be attributed to a particular allele to determine the allelic skewing of the entire gene and link these to sequence variations outside the gene.

[0010] While the isolation of polyA - mRNAs is known previously (e.g. Kowalczyk, 2012), it has not previously been used in the context of gene haplotypes or allelic skewing.

[0011] EP 1829979 Al relates to a method of searching for genetic polymorphisms (e.g. SNPs) in cDNA as a means of identifying genes whose expression levels differ between alleles. The inventors recognised that mature polyA+ mRNA has only one exon sequence after splicing and therefore such sequences are "too short" to contain enough SNPs to be evaluated. Therefore, the inventors chose nuclear RNA to provide "long chains", which were expected to contain many genetic polymorphisms that would enable genes whose expression levels differ between alleles to be distinguished (see paragraphs

[0006] and

[0018] ).

[0012] James et al. (2013) used pre-RNA to detect skewing of expression of one allele over the other in immune cells under stimulated conditions, but did not do so in the context of phased genomes to link to genetics or test for functional skewing in the genome, within or distal to the gene.

[0013] Sigurdsson et al. (2008) used allele skewing to detect allele skewing with risk haplotypes located within the gene body in the context of unphased genomes, but did not do so in the context of phased genomes to link to genetics or test for functional skewing of the epigenome within or distal to the gene.

[0014] Thomas et al. (2011) combined the use of epigenetic skewing and gene expression skewing to study epigenetic effects associated with monoallelic genes, but only within the gene body, as they did not use it in the context of phased genomes.

[0015] Rainbow et al. (2008) used pre-mRNA forms to observe allele-specific differences in IL-2 gene expression, but this was not in the context of phased genomes relating to the distal regulatory landscape or genetics. SUMMARY

[0016] It is an object of the present invention to provide a method of obtaining an indication of dysregulation between the expression levels of different alleles of the same gene in a target eukaryotic cell, which method comprises the use of pre-mRNA. In a phased genome, this dysregulation can then be linked to sequence variations outside the gene body, leading to dysregulation of the gene on the affected allele.

[0017] Without phasing, each SNP detected in the RNA analysis cannot be assigned to the same allele (unless they are very close, i.e. within a short distance of < 300bp) under the assumption that they are skewed in the same direction. Therefore, they cannot be used to add statistical robustness to the allele skewing for a given allele. By phasing, all SNPs can be assigned to a given allele and their behavior can be seen to be similar statistically in terms of representation changes and skewing direction in the RNA.

[0018] More importantly, without phasing, this skewed RNA expression cannot be linked to changes outside the gene. Regulatory elements controlling gene expression in cis on the same chromosome can be located up to 2 million base pairs away from the genes they control. With phasing, the genotype genetically linking a sequence change to a given trait or disease can be considered to be on the same allele as the gene that shows reproducible skewing of expression in its alleles. This also allows one to link epigenetic changes at distal SNPs to changes in expression of alleles on the same chromosome. These epigenetic changes can be detected using allele skewing, sequence-based genomic assays to measure enhancer activity, such as open chromatin assays (e.g. DNAse-seq or ATAC-seq), chromatin marks associated with regulatory activity (e.g. ChIP-seq for chromatin marks like H3k27ac, H3k4mel or for protein binding, such as RNA polymerase, transcription factor binding) or binding of important regulatory structural proteins, such as CTCF.

[0019] The present invention relates to the use of haplotype phasing to detect abnormal expression of genes that can be associated with a disease or condition.

[0020] Using the methods of the present invention, skewing in gene expression can be linked to genetic signals associated with trait and disease traits outside the gene, where they are typically enriched. It also allows skewing in gene expression to be linked to skewing in sequence-based genomics analyses associated with regulatory activity, and can be done on a genome scale, in large scale, to validate potential mechanisms of gene expression changes.

[0021] The information obtained from the methods of the present invention can be used to identify genes that are abnormally expressed in a particular disease or condition (or its associated biochemical pathways), and thus to identify candidate drug targets.

[0022] The information can also be useful in determining potential genetic causes of a disease or condition.

[0023] In one embodiment, the present invention provides a method of obtaining an indication of dysregulation between the expression levels of at least two alleles of the same gene in a target eukaryotic cell, the method comprising the steps of:

[0024] for a plurality of genes from one or more target eukaryotic cells,

[0025] (a) obtaining pre-mRNAs of at least two alleles of the same gene; and

[0026] (b) determining a ratio (R i,j ) between the pre-mRNA amounts of one or more pairs of alleles (i,j) of the same gene;

[0027] where if for one or more pairs of alleles (i,j) of the same gene, R i,j ≠ 1, then this indicates that there is a misregulation between the expression levels of these two alleles of the gene in the target eukaryotic cell.

[0028] Preferably, if for one or more pairs of alleles (i,j) of the same gene, if R i,j < 0.9 or if R i,j > 1.1, then this indicates a misregulation between the expression levels of these two alleles of the gene in the target eukaryotic cell.

[0029] In another embodiment, the present application provides a method of identifying a gene whose alleles are misregulated in a target eukaryotic cell, the method comprising the steps of:

[0030] for a plurality of genes from one or more target eukaryotic cells,

[0031] (a) obtaining pre-mRNAs of at least two alleles of the gene; and

[0032] (b) determining the ratio (R i,j ) between the pre-mRNA amounts of a pair of alleles (i,j) of the gene;

[0033] where if for a pair of alleles (i,j) of the gene, R i,j ≠ 1, then this indicates a gene whose alleles are misregulated in the target eukaryotic cell.

[0034] Preferably, if for a pair of alleles (i,j) of the gene, R i,j < 0.9 or if R i,j > 1.1, then this indicates a gene whose alleles are misregulated in the target eukaryotic cell.

[0035] In yet another embodiment, the present application provides a method of obtaining an indication of a disease or condition, the method comprising the steps of:

[0036] for a plurality of genes from one or more target eukaryotic cells,

[0037] (a) obtaining pre-mRNAs of at least two alleles of the gene; and

[0038] (b) determining the ratio (R i,j ) between the pre-mRNA amounts of a pair of alleles (i,j) of the gene;

[0039] where the target eukaryotic cell is a cell associated with or characteristic of the disease or condition,

[0040] and wherein if for a pair of alleles (i,j) of a gene, R i,j ≠ 1, then this indicates that the disease or disorder is caused by a misregulation between the expression levels of the alleles of the gene.

[0041] Preferably, if for a pair of alleles (i,j) of a gene, R i,j < 0.9 or if R i,j > 1.1, then this indicates that the disease or disorder is caused by a misregulation between the expression levels of the alleles of the gene.

[0042] In yet another embodiment, the present application provides a method of identifying an allelic mutation of a gene that can cause a misregulation of the expression levels of the alleles of the gene in a target eukaryotic cell, the method comprising the steps of:

[0043] for a plurality of genes from one or more target eukaryotic cells,

[0044] (a) obtaining pre-mRNAs of at least two alleles of a gene; and

[0045] (b) determining a ratio (R i,j ) between the amounts of pre-mRNAs of one or more pairs of alleles (i,j) of the gene;

[0046] wherein when R i,j ≠ 1 for a pair of alleles (i,j) of the gene, or in response to determining that R i,j ≠ 1 for a pair of alleles (i,j) of the gene, the method further comprises the steps of:

[0047] (c) determining the nucleotide sequences of the pair of alleles; and

[0048] (d) comparing the nucleotide sequences of the pair of alleles to identify a difference between the nucleotide sequences of the pair of alleles;

[0049] wherein the one or more differences between the nucleotide sequences of the pair of alleles of the gene can be a mutation that causes a misregulation of the expression levels of the two alleles of the gene in the target eukaryotic cell.

[0050] Preferably, when R i,j < 0.9 for a pair of alleles (i,j) of the gene, or if R i,j > 1.1, or in response to determining that R i,j < 0.9 for a pair of alleles (i,j) of the gene, or if R i,j > 1.1, the method comprises the steps of:

[0051] (c) determining the nucleotide sequences of the pair of alleles; and

[0052] (d) comparing the nucleotide sequences of the pair of alleles to identify differences between the nucleotide sequences of the pair of alleles;

[0053] wherein the one or more differences between the nucleotide sequences of the pair of alleles of the gene are likely to be mutations that cause the expression levels of the two alleles of the gene in the target eukaryotic cell to be dysregulated.

[0054] As used herein, the term "dysregulated" refers to a difference in regulation between one or more alleles of a gene. Typically, alleles of a gene are expressed at similar levels. However, mutations in regulatory regions that control expression of individual alleles can cause the expression of these alleles to be enhanced or reduced. Thus, the term "dysregulated" also refers to a difference in the expression levels of alleles of a gene.

[0055] The target cell is a eukaryotic cell (i.e. a cell whose nuclear DNA has introns). Preferably, the eukaryotic cell is a mammalian cell, e.g. a human, monkey, mouse, rat, pig, goat, horse, sheep or bovine cell. Most preferably, the cell is a human cell.

[0056] The cell can be a primary cell or a cell from a cell line. Preferably, the cell is a primary cell.

[0057] In some embodiments of the application, the target cell is a hematopoietic cell, e.g. a red blood cell, a lymphocyte (e.g. a T cell, a B cell and a natural killer cell), a granulocyte, a megakaryocyte and a macrophage. Preferably, the target cell is a primary lymphoid cell.

[0058] In other embodiments, the target cell is a brain cell; preferably, the target cell is a primary neuronal cell.

[0059] Although the cell is preferably a diploid cell, the methods of the application are applicable to cells of other ploidy, e.g. tetraploid cells.

[0060] In some embodiments of the application, the target eukaryotic cell is one that is associated with or characterised by a disease or condition. Examples of eukaryotic cells that are associated with or characterised by a disease or condition are given below:

[0061]

[0062] Examples of diseases and conditions include those in the above table.

[0063] The pre-mRNA is obtained from one or more target eukaryotic cells. In some preferred embodiments, the pre-mRNA is obtained from a single target cell. In other embodiments, the RNA is obtained from a population of target cells. In such embodiments, the population of target cells comprises substantially or completely the same type of cell, i.e. the population is substantially or completely homogenous.

[0064] The gene is a gene present in the target eukaryotic cell. As used herein, the term "gene" is not limited to RNA- or protein-encoding sequences. It includes associated regulatory elements, such as enhancers, promoters, and terminator sequences. The term "gene" can be defined to include its transcribed region, its introns, the promoter region, and all regulatory or structural elements that determine its activity or expression level within 2 million base pairs of the promoter of the transcribed part of the gene.

[0065] The method of the application is particularly suitable for genes affected by genetic variations known to be associated with a disease or trait. The genetic variations are often already determined by genome-wide association studies (GWAS), but the genetic variations can also be mutations in the individual subject or in cancerous subclones.

[0066] Although the application enables detection of dysregulation of any gene, certain types of dysregulated genes are preferred. These include genes that are specifically expressed in the cell type of interest compared to genes expressed in all cell types. Similarly, genes that are important for the function of the cell type of interest, such as genes of the immune synapse in immune cells, are also preferred. Also preferred are genes that encode transcription factors that determine the cell type properties and / or behavior of a given cell type. Furthermore, genes that affect the overall fitness of a cell, such as cell cycle regulatory genes or genes associated with apoptosis, are preferred. In cancer cells, genes that affect processes such as cell proliferation, motility, and / or adhesion are preferred, including but not limited to membrane receptors, secondary messenger molecules, transcription factors, and epigenetic regulators.

[0067] As used herein, the term "a plurality of genes" refers to 2 or more genes, such as 2-10, 10-100, 100-500, 500-1000, 1000-5000, 5000-10000, or 10000 or more.

[0068] There are at least two alleles of the same gene present in each target eukaryotic cell. For example, there can be 2, 3, or 4 alleles of the same gene. Preferably, there are 2 alleles of the same gene present in each target eukaryotic cell.

[0069] The sequence of one or more or all exons in the pre-mRNA alleles can be the same or different. The sequence of one or more of all introns in the pre-mRNA alleles can be the same or different.

[0070] The target cells are preferably heterozygous for the one or more alleles of interest.

[0071] Step (a) of the method of the application comprises a step of obtaining pre-mRNAs. Preferably, the pre-mRNAs are from at least two alleles of the same gene on a given haplotype.

[0072] Pre-mRNA is the first form of RNA produced by transcription during protein synthesis. Pre-mRNA is transcribed from a DNA template in the target nucleus. Pre-RNA is usually 5 '-capped (i.e., with a 7-methylguanosine) because this capping occurs within a few nucleotides after the start of RNA synthesis. In the context of the present invention, pre-mRNA can or can not contain the 5 '-cap.

[0073] Pre-mRNA and mRNA in eukaryotic cells differ in two major ways: (i) a poly-A tail is added to the 3' end of pre-mRNA; and (ii) introns are spliced out of pre-mRNA. Once pre-mRNA is processed to contain these features, it is referred to as "mature messenger RNA" or simply "messenger RNA" (mRNA). Polyadenylated mRNA is also referred to as "polyA + RNA".

[0074] In the context of the present invention, it is important that all or substantially all introns are retained in the pre-mRNA, i.e., they are not all spliced out. As used herein, in one embodiment, the term "substantially all introns are retained" is used to mean that at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the intron nucleotide sequences (that would normally be spliced out) are retained in the pre-mRNA. In another embodiment, the term "substantially all introns are retained" means that at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the total number of introns are retained in the pre-mRNA.

[0075] In the context of the present invention, it is also important (but not essential) that polyadenylation does not occur. This allows information to be obtained from the 3' end of the target gene. Thus, the pre-mRNA can contain small or trace amounts of polyadenylated mRNA.

[0076] In some embodiments of the present invention, it is also important that sequence variations around each allele are known. This allows sequence variations that affect gene regulation to be linked to allele skewing and gene deregulation.

[0077] In the methods of the present invention, the optimal amount of pre-mRNA is obtained from the target eukaryotic cell (i.e., as much as possible).

[0078] Pre-mRNA can be obtained from cells by cell lysis and then solvent extraction and removal of rRNA and polyA + RNA species. Pre-mRNA can also be obtained from cells by cell lysis and then binding the RNA to an affinity column and removal of rRNA and polyA + RNA species.

[0079] Preferably, the pre-mRNA is obtained by cell lysis, followed by solvent extraction and total RNA precipitation. For example, total RNA can be extracted using TRIzol reagent (Sigma), PhaseLock gel tubes (5Prime) and centrifugation. RNA can be precipitated by mixing with an equal volume of propan-2-ol, followed by centrifugation. The precipitated RNA can be washed with 75% ethanol, dried and dissolved in water (e.g. Kowalczyk, 2012).

[0080] The pre-mRNA used in the method of the application can also be mixed with small amounts of other RNA, such as rRNA and / or tRNA, or DNA.

[0081] Preferably, the pre-mRNA does not comprise ribosomal RNA (rRNA) or the pre-mRNA has been depleted of rRNA, e.g. prior to use. For example, rRNA can be removed using the RiboMinus Eukaryote Kit for RNA sequencing (Invitrogen).

[0082] Preferably, the pre-mRNA is polyA - mRNA. Preferably, the pre-mRNA does not comprise DNA; this can be removed with a DNAse.

[0083] Step (b) of the method of the application comprises determining the ratio (R i,j ) between the amounts of pre-mRNA of the pair (or one or more pairs) of alleles (i,j) of the (same) gene.

[0084] In this respect, the method of the application comprises (1) a method involving determining the absolute amounts of pre-mRNA of the pair of alleles (and their ratio) and (2) a method determining the relative amounts of pre-mRNA of the pair of alleles (not necessarily determining the absolute amounts of the alleles pre-mRNA).

[0085] The method of obtaining the absolute amounts of pre-mRNA of the alleles comprises RNA-Seq, followed by genomic alignment and counting of the number of aligned sequences. Sequences originating from a particular allele are identified and counted by identifying sequence variations in the introns, exons and downstream transcriptional regions in the RNA-Seq data that are known to be specific for that allele.

[0086] Preferably, the absolute quantity of pre-mRNA of the first and second alleles is determined from strand-specific RNA-Seq, followed by genomic alignment and counting of the number of aligned sequences. Sequences from non-transcribed strands are discarded to remove genomic contamination and antisense transcription. Sequences originating from a particular allele are identified and counted by recognizing sequence variations in introns, exons and downstream transcriptional regions in the RNA-Seq data that are known to be specific for that allele (e.g. Quinn et al., 2013).

[0087] Preferably, the relative or absolute quantity of pre-mRNA of the first and second alleles is determined using RNA-Seq.

[0088] The relative difference in the quantity of the first and second alleles can be determined using hybridization of cDNA to SNP microarrays and determining the relative signal from each allele. The relative difference in the quantity of the first and second alleles can also be determined using SNP-specific PCR and determining the relative signal from each allele.

[0089] R i,j The ratio of the quantity of pre-mRNA of a pair of alleles (i,j) of a gene, where i is the first allele and j is the second allele, is defined herein as R

[0090] At normal expression, the quantity of pre-mRNA of the first allele (i) of a gene should be approximately equal to the quantity of pre-mRNA of the second allele (j) of the gene, i.e. the expression of both alleles should be equal.

[0091] When R i,j ≠ 1 for one or more pairs of alleles (i,j) of the same gene, this indicates a dysregulation between the expression levels of these two alleles of the gene in the target eukaryotic cell.

[0092] In one embodiment, the term "R i,j ≠ 1" means that R ij is substantially different from 1, e.g. R ij is not equal to 1 when random variations and experimental errors are taken into account.

[0093] In another embodiment, the term "R i,j ≠ 1" means that the quantity of pre-mRNA of the first allele (i) of a gene is statistically different from the quantity of pre-mRNA of the second allele (j) of the gene (e.g. p < 0.05 using a Student's t-test).

[0094] Preferably, the term "R i,j ≠ 1" means that R i,jless than 0.95, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1, 0.05 or 0.01 ; or R i,j greater than 1.05, 1.1, 1.2, 1.4, 1.6, 2, 2.5, 3, 5, 10, 20 or 100. More preferably, R i,j less than 0.9 or R i,j greater than 1.1.

[0095] For example, if R i,j < 0.9, this indicates a dysregulation between the expression of the first and second alleles of the gene. In this case, the expression level of the first allele of the gene is lower than the expression level of the second allele of the gene.

[0096] For example, if R i,j > 1.1, this indicates a dysregulation between the expression of the first and second alleles of the gene. In this case, the expression level of the first allele of the gene is higher than the expression level of the second allele of the gene.

[0097] The quantification of R ij between the affected allele (i) and the normal allele (j) allows a reproducible and statistically reliable identification of genes associated with a human disease or trait. The variation in R ij gives the direction of the genetics, i.e. whether an increase or decrease in gene activity is associated with the disease or trait.

[0098] Furthermore, the magnitude of R ij indicates to what extent the gene activity needs to be changed to produce a measurable physiological effect.

[0099] If a dysregulation is found between a pair of alleles (i,j) of a gene, the sequence of the pair of alleles can be determined in order to try to determine the cause of the dysregulation, for example using RNA-Seq.

[0100] RNA-Seq (RNA sequencing), also known as whole transcriptome shotgun sequencing (WTSS), uses next-generation sequencing (NGS) to reveal the presence and quantity of RNAs in a biological sample at a given moment.

[0101] In the context of the present invention, RNA-Seq comprises the following steps:

[0102] (i) the precursor mRNAs are fragmented in vitro and copied into ds-cDNA (for example using a reverse transcriptase); and

[0103] (ii) the ds-cDNA is then sequenced, preferably using a high-throughput, short read sequencing method (for example NGS).

[0104] These sequences can then be aligned to the reference genome sequence to reconstruct the regions of the genome being transcribed. This data can be used to annotate the location of expressed genes, their relative expression levels, and any alternative splicing variants.

[0105] If the expression levels of the two alleles are different (e.g., R i,j < 0.9 or R i,j > 1.1), this provides an indication that there is a change in the regulatory elements controlling gene expression on one allele. The regulatory elements can be present within introns of the gene or extragenically.

[0106] Methods such as ATAC-seq and NG Capture-C can also be used to further identify the exact genetic cause of the allelic imbalance, particularly in cases where the underlying cause is due to a change in regulatory elements.

[0107] The methods of the present invention can also include the step of performing a sequence-based assay (preferably ATAC-seq, DNase-seq, or ChIP-seq) that measures the activity of regulatory elements to detect skew on the same allele of the discovered gene in RNA-seq skew. The methods of the present invention include the use of haplotype phasing. The disclosure of each of the references set forth herein is specifically incorporated herein by reference. BRIEF DESCRIPTION OF DRAWINGS

[0108] Figure 1 Showing the use of pre-mRNA in phasing haplotypes and pre-mRNA identification of gene imbalance.

[0109] Figure 2 Showing the use of pre-mRNA in phasing genomes to detect imbalance in the IKZF1 gene associated with specific sequence variants in regulatory elements.

[0110] Figure 3 Showing the frequency of informative heterozygotes in the general population that find risk alleles associated with common diseases. DETAILED DESCRIPTION

[0111] EMBODIMENT

[0112] The application is further illustrated by the following examples, in which all parts and percentages are by weight and degrees are Celsius, unless otherwise stated. It should be understood that, although the examples illustrate preferred embodiments of the application, they are merely for illustration. It can be apparent to one of ordinary skill in the art, based on the discussion and examples provided herein, that modifications and other implementations of the application can be made without departing from the spirit and scope thereof and that the scope of the application should be determined only by reference to the claims that follow. Accordingly, various modifications can be made to the application in light of the above description and the examples, without departing from the spirit and scope of the application. Such modifications are intended to be within the scope of the claims.

[0113] Example 1: Using pre-mRNA to detect gene misregulation associated with a specific haplotype in phased genomes

[0114] Figure 1 A shows a schematic of two genomic alleles, each containing two genes and a regulatory element. The sequence variation that distinguishes the two alleles is shown as X (e.g., a single nucleotide polymorphism (SNP), small insertion, or deletion). Exons of the two genes are shown as boxes, and the promoter element of the two genes is shown as a vertical line and horizontal arrow associated with the first exon of the genes. The location of the regulatory element is shown as a triangle between the two genes.

[0115] In haplotype A, the regulatory element contains a sequence variation (shown as lighter shading) that alters its activity. Regulatory interactions between the regulatory element and the genes are shown as arrowed arcs drawn by 3C methods, such as Capture-C. The sequence variation that distinguishes the allele of origin of the pre-mRNA from the two genes is located within the transcribed portion of the genes (e.g., introns, exons, and downstream regions). In this example, the gene misregulation is caused by a damaged regulatory element, but the same use of phased sequence variation in combination with sequencing of pre-mRNA can be used for any other mechanism (e.g., gain of function caused by sequence variation or larger-scale structural changes).

[0116] Figure 1 B shows that, in the example genes, the coverage of the pre-mRNA form of RNA-Seq is increased, which retains the transcribed introns and downstream regions of the genes and increases the amount of sequence variation that can detect gene misregulation. Exons of the genes are shown as vertical lines of different thickness, while introns are shown as horizontal hatching.

[0117] Figure 1 C shows loss of binding of the GATA1 transcription factor (ChIP qPCR) on the erythroid regulatory element containing a single base pair variation (rs10758656) that is homozygously edited into Hudep2 cells.

[0118] Figure 1 D shows loss of open chromatin signal at the same regulatory element using ATAC-seq in homozygous presence of sequence variant rs10758656.

[0119] Figure 1 E shows erythroid-specific interaction between this regulatory element and the promoter of RCL1 and JAK2 in wild-type cells.

[0120] Figure 1 F shows loss of expression of the JAK2 gene only in cells edited homozygous for rs10758656.

[0121] Figure 1 G shows allelic skewing towards the wild-type allele in primary erythrocytes for ATAC-seq at the regulatory element (squares) and pre-mRNA expression of the JAK2 gene (circles) only in individuals heterozygous for rs10758656.

[0122] Example 2: Using pre-mRNAs in phasing genomes to detect dysregulation of the IKZF1 gene associated with specific sequence variants in regulatory elements

[0123] Figure 2 A shows sequence variants in regulatory elements that impair the binding potential of transcription factors using the Sasquatch algorithm (Schwessinger R. et al., 2017).

[0124] Figure 2 B shows that this regulatory element specifically interacts with the promoter of the IKZF1 gene using NG Capture-C.

[0125] Figure 2 C shows allelic skewing of open chromatin signal in primary erythrocytes of 3 individuals heterozygous for the disruptive sequence variant (Hap-A) towards the wild-type (Hap-B) as determined by ATAC-seq. The corresponding signal in the same sample in haplotype B is associated with the signal in haplotype A by a dashed line.

[0126] Figure 2 D shows cumulative reduction of pre-mRNAs of the IKZF1 gene haplotype A summarizing all transcriptional sequence variants that distinguish between the two haplotypes. The corresponding signal in the same sample in haplotype B is associated with the signal in haplotype A by a dashed line.

[0127] Example 3: Using allelic skewing in pre-mRNAs in phasing genomes allows analysis of regulatory variants in primary cells at an unprecedented scale

[0128] Figure 3The number of informative individuals expected in a random sample of the general population for a given minor allele frequency (MAF) of a sequence variant. The grey bars indicate the number of expected individuals that are heterozygous at a given MAF. The black line shows the average distribution of minor allele frequencies in typical genome-wide association (GWA) studies for human diseases (type 1 diabetes, ankylosing spondylitis, red blood cell traits, and multiple sclerosis). This shows that at a MAF of 0.3 this would provide more than 20 independent observations of genetic perturbation and would cover more than half of a typical GWA study. Likewise, at a MAF of 0.1 this would provide 5 or more independent observations of genetic perturbation and would cover more than 90% of a typical GWA study.

[0129] References:

[0130] James C et al., Cell, vol. 155, 2013, "Human SNP Links Differential Outcomes in Inflammatory and Infectious Disease to a FOX03-Regulated Pathway", pages 57-69.

[0131] Kowalczyk, M.S. et al. Intragenic enhancers act as alternative promoters. Mol Cell 45, 447-58 (2012).

[0132] Quinn EM, et al. (2013) Development of Strategies for SNP Detection in RNA-Seq Data: Application to Lymphoblastoid Cell Lines and Evaluation Using 1000 Genomes Data. PLoS ONE 8(3): e58815. https: / / doi.org / 10.1371 / journal.pone.0058815.

[0133] Rainbow et al., BIOCHEMICAL SOCIETY TRANSACTIONS, vol. 36, 2008, "Commonality in the genetic control of Type 1 diabetes in humans and NOD mice: variants of genes in the IL-2 pathway are associated with autoimmune diabetes in both species", page 312.

[0134] Schwessinger R, et al. (2017) Sasquatch: predicting the impact of regulatory SNPs on transcription factor binding from cell- and tissue-specific DNase footprints. Genome Res. 2017 Oct; 27(10): 1730-1742. PMCID:PMC5630036.

[0135] Sigurdsson et al., HUMAN MOLECULAR GENETICS, vol. 17, 2008, "A risk haplotype of STAT4 for systemic lupus erythematosus is over-expressed, correlates with anti-dsDNA and shows additive effects with two risk alleles of IRF5", pages 2868-2876.

[0136] Thomas et al., EPIGENETICS & CHROMATIC, vol. 4, 2011, "Allele-specific transcriptional elongation regulates monoallelic expression of the IGF2BP1 gene", page 14.

Claims

1. A method of identifying a mutation in an allele of a gene that can result in a misregulation of the expression level of the allele of the gene in a target eukaryotic cell, the method comprising the steps of: for a plurality of genes from one or more target eukaryotic cells, (a) obtaining pre-mRNAs of at least two alleles of the genes; and (b) determining a ratio R between the pre-mRNA amounts of one or more pairs of alleles (i,j) of said gene i,j ; wherein when R i,j ≠ 1 for allele pair (i,j) of the gene, or in response to determining that R i,j ≠ 1 for allele pair (i,j) of the gene, the method further comprises the steps of: (c) determining the nucleotide sequences of the pair of alleles; and (d) comparing the nucleotide sequences of the pair of alleles to identify differences between the nucleotide sequences of the pair of alleles; wherein one or more differences between the nucleotide sequences of the pair of alleles of the gene can be a mutation that results in a misregulation of the expression level of the two alleles of the gene in a target eukaryotic cell; wherein the method is performed on phased genomic sequences, wherein all sequence differences are attributed to a particular allele to determine allele skewing of the entire gene, and these sequence differences are linked to sequence variation outside the gene; wherein the method comprises linking skewing in gene expression to skewing in sequence-based genomics analysis related to regulatory activities; and wherein the method is not a diagnostic method for a disease or condition.

2. The method of claim 1, wherein when for an allele pair (i,j) of the gene, R i,j < 0.9 or R i,j > 1.1, or in response to determining for an allele pair (i,j) of the gene, R i,j < 0.9 or R i,j > 1.1, the method comprises the step of: (c) determining the nucleotide sequences of the pair of alleles; and (d) comparing the nucleotide sequences of the pair of alleles to identify differences between the nucleotide sequences of the pair of alleles; wherein one or more differences between the nucleotide sequences of the pair of alleles of the gene can be a mutation that results in a misregulation of the expression level of the two alleles of the gene in a target eukaryotic cell.

3. The method of claim 1 or 2, wherein step (c) is performed using RNA-Seq.

4. The method of claim 3, wherein sequences derived from a particular allele are identified and counted by identifying sequence changes within the RNA-Seq data in introns, exons, and downstream transcriptional regions known to be specific to that allele.

5. The method of any one of the preceding claims, wherein if Ri,j < 0.9 or Ri,j > 1.1, it is indicative of a change in a regulatory element on one allele that controls expression of the gene.

6. The method of claim 5, wherein the regulatory element is a regulatory element present within an intron of the gene or outside the gene.

7. The method of claim 6, wherein the regulatory element is a regulatory element present outside the coding region of the gene.

8. The method of claim 3 or 4, wherein the method additionally comprises the further step of performing a sequence-based assay that measures activity of a regulatory element to detect skewing on the same allele of a gene found to be skewed using RNA-seq.

9. The method of claim 8, wherein the sequence-based assay is ATAC-seq, DNase-see, or ChIP-seq.

10. The method of any one of the preceding claims, wherein the eukaryotic cell is a human primary lymphoid cell or a primary neuronal cell.

11. The method of any one of the preceding claims, wherein the plurality of genes is 2-10 genes.

12. The method of any one of claims 1-10, wherein the plurality of genes is 10-100 genes.

13. The method of any one of claims 1-10, wherein the plurality of genes is 100-500 genes.

14. The method of any one of claims 1-10, wherein the plurality of genes is 500-1000 genes.

15. The method of any one of claims 1-10, wherein the plurality of genes is 1000-5000 genes.

16. The method of any one of claims 1-10, wherein the plurality of genes is 5000-10000 genes.

17. The method of any one of claims 1-10, wherein the plurality of genes is 10000 or more genes.

18. The method of any one of the preceding claims, wherein there are 2 alleles of the same gene present in each target eukaryotic cell.

19. The method of any one of the preceding claims, wherein the pre-mRNA is a polyA - mRNA.

20. The method of claim 19, wherein the pre-mRNA is a polyA+ mRNA obtained from total cellular RNA. - mRNA.

21. The method of any one of the preceding claims, wherein R i,j ≠ 1 means that R i,j is less than 0.

95.

22. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is less than 0.

9.

23. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is less than 0.

8.

24. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is less than 0.

7.

25. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is less than 0.

6.

26. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is less than 0.

5.

27. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is less than 0.

4.

28. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is less than 0.

3.

29. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is less than 0.

2.

30. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is less than 0.

1.

31. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is less than 0.

05.

32. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is less than 0.

01.

33. The method of any one of claims 1-20, wherein R i,j ≠ 1 means R i,j is greater than 1.

05.

34. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is greater than 1.

1.

35. The method of any one of claims 1-20, wherein R i,j ≠ 1 means R i,j is greater than 1.

2.

36. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is greater than 1.

4.

37. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is greater than 1.

6.

38. The method of any one of claims 1-20, wherein R i,j ≠ 1 means R i,j is greater than 2.

39. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is greater than 2.

5.

40. The method of any one of claims 1-20, wherein R i,j ≠ 1 means R i,j is greater than 3.

41. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is greater than 5.

42. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is greater than 10.

43. The method of any one of claims 1-20, wherein R i,j ≠ 1 means R i,j is greater than 20.

44. The method of any one of claims 1-20, wherein R i,j ≠ 1 means that R i,j is greater than 100.

Citation Information

Patent Citations

  • Method of identifying gene with variable expression

    EP1829979A1