Method for the diagnosis of multiple esclerosis
Detecting TAF1 isoforms with C-terminal deletions in biological samples addresses the challenge of MS progression by offering a non-invasive diagnostic and therapeutic screening method.
Patent Information
- Application Number
- PCT/EP2025/073132
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-14
- Filing Date
- 2025-08-12
- Publication Date
- 2026-02-19
AI Technical Summary
Current MS therapies fail to slow long-term disease progression due to unknown molecular events in the CNS, necessitating a simple, non-invasive method for diagnosis and screening of therapeutic agents.
Detecting specific isoforms of TAF1 with C-terminal deletions in biological samples to diagnose and prognose MS, and screen for therapeutic agents using in vitro and computer-implemented methods.
Provides a reliable method for diagnosing MS and predicting its progression, enabling effective screening of therapeutic agents.
Smart Images

Figure IMGF000068_0001 
Figure IMGF000069_0001 
Figure IMGF000077_0001
Abstract
Description
[0001] METHOD FOR THE DIAGNOSIS OF MULTIPLE ESCLEROSIS
[0002] DESCRIPTION
[0003] The present invention is comprised within the field of biomedicine. The present invention relates to a method for diagnosis of Multiple Sclerosis (MS) and methods of screening for therapeutic agents.
[0004] BACKGROUND ART
[0005] Multiple Sclerosis (MS) is a disabling neurological disease characterized by neuroinflammation and demyelination affecting the brain, the spinal cord and the optic nerves, profoundly reducing quality of life for the majority of affected individuals, which is about 2.8 million worldwide. Most individuals debut with partially or fully reversible episodes of neurological deficits (relapses) in the clinical form known as relapsing-remitting MS (RRMS). A great proportion of these evolve to the secondary progressive MS (SPMS) clinical form, with the development of permanent neurological deficits and the continuous progression of clinical disability. Approximately 15% of individuals with MS have a progressive course from disease onset, which is referred to as primary progressive MS (PPMS).
[0006] Genetic variants that confer MS risk implicate genes involved in immune function, while variants related to severity of the disease are associated with genes preferentially expressed within the CNS. Current MS therapies decrease relapse rates by preventing immune-mediated damage of myelin, but they ultimately fail to slow long-term disease progression, which apparently depends on CNS intrinsic processes. The molecular events that trigger progressive MS are still unknown.
[0007] Although the precise etiology of MS remains unknown, epidemiology indicates that it is a complex disease influenced by both environmental and genetic factors. Environmental factors include Epstein-Barr virus (EBV) infection in adolescence and early adulthood, smoking, lack of sun exposure and low vitamin D levels. The genetic architecture of MS consists of hundreds of small-effect common variants that increase disease risk or severity. While risk variants implicate genes strongly enriched for immune relevance, MS severity is associated with variation in genes preferentially expressed within the CNS. This suggests that processes related to severity leading to permanent disability originate from the CNS, explaining the lack of efficacy of the current available immunological therapies to prevent or halt the final progression of the disease independent of relapse activity.
[0008] Therefore, there is a need for biomarkers which allow the diagnosis of MS, the prognosis of disease progression by means of a simple, effective and non-invasive method and the screening for therapeutic agents.
[0009] BRIEF DESCRIPTION OF THE INVENTION
[0010] In this invention, inventors show that the C-terminal region of TAF1 (the scaffolding subunit of the general transcription factor TFIID) is underrepresented in postmortem brain tissue from individuals with MS. Furthermore, it is demonstrated in vivo, in genetically modified mice, that C-terminal alteration of TAF1 suffices to induce an MS-like brain transcriptomic signature, including increased expression of proinflammatory genes, and accompanying MS-like clinical symptoms. This transcriptional profile is accompanied by CNS-resident inflammation, robust demyelination and MS-like motor phenotypes. It is identified numerous interactors of C-terminal TAF1 that show evidence for genetic association to MS. This invention shows that TAF1 dysfunction converges with genetic susceptibility to cause transcriptional dysregulation in CNS cell types, such as oligodendrocytes, to ultimately trigger MS.
[0011] In a first aspect, the invention relates to an in vitro method for diagnosing multiple sclerosis in a subject that comprises: a) detecting an isoform of the TATA-box binding protein associated factor 1 (TAF1 ) as established in SEQ ID NO: 1 , or an isoform of the corresponding mRNA, in a biological sample from a subject, wherein i) the isoform of SEQ ID NO: 1 comprises at least one deletion from residues: 1709, 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893 of SEQ ID NO: 1 , or ii) the isoform of the mRNA encodes the isoform i) of SEQ ID NO: 1 , wherein the presence of the isoform i) or the presence of the isoform ii) indicates that the subject suffers from, or is at risk of suffering from, multiple sclerosis.
[0012] In a second aspect, the invention relates to an in vitro method for the prognosis of multiple sclerosis in a subject that suffers from primary progressive multiple sclerosis (PPMS) or secondary progressive multiple sclerosis (SPMS) that comprises: a) detecting an isoform of the TAF1 as established in SEQ ID NO: 1 , or an isoform of the corresponding mRNA, in a biological sample from a subject, wherein i) the isoform of SEQ ID NO: 1 comprises at least one deletion from residues: 1709, 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893 of SEQ ID NO: 1 , or ii) the isoform of the mRNA encodes the isoform i) of SEQ ID NO: 1 ; b) calculating the abundance or concentration of isoform at step a) with a predetermined reference value for the same biological marker; wherein the abundance or concentration of the isoform i) or the isoform ii) indicates that the subject suffers poor prognosis for primary progressive multiple sclerosis (PPMS) or secondary progressive multiple sclerosis (SPMS).
[0013] In a third aspect, the invention relates to a method for screening of an active compound for treating and / or preventing multiple sclerosis that comprises: a) administering a potentially active compound for the treatment and / or prevention of multiple sclerosis to a cell, a tissue or a non-human animal model that expresses an isoform of the TAF1 as established in SEQ ID NO: 1 that comprises at least one deletion from residues: 1709, 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893 of SEQ ID NO: 1 , b) evaluating the change in the gene expression, phenotype change and / or histological change of the cell, tissue or non-human animal model in response to the potentially active compound.
[0014] In a fourth aspect of the invention, it relates to a computer implemented method for diagnosing multiple sclerosis in a subject that comprises the following steps: a) receiving or storing or having access to data comprising the levels of an isoform of the TAF1 as established in SEQ ID NO: 1 or an isoform of the corresponding mRNA in a biological sample from the subject, wherein i) the protein isoform comprises at least one deletion from residues: 1709,
[0015] 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893 of SEQ ID NO: 1 , or ii) the isoform of the mRNA encodes the isoform i) of the SEQ ID NO: 1 , b) calculating a score value based on the comparison of the concentration of said isoform i) or ii) compared to a reference value, and c) returning as a result a risk or an indication of whether a subject has got multiple sclerosis or is at risk of suffering said disease based on the calculated score.
[0016] In a fifth aspect of the invention, it relates to a computer implemented method for the prognosis of multiple sclerosis in a subject that suffers from primary progressive multiple sclerosis (PPMS) or secondary progressive multiple sclerosis (SPMS) that comprises the following steps: a) receiving or storing or having access to data comprising the levels of an isoform of the TAF1 as established in SEQ ID NO: 1 or an isoform of the corresponding mRNA in a biological sample from the subject, wherein i) the protein isoform comprises at least one deletion from residues: 1709, 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893 of SEQ ID NO: 1 , or ii) the isoform of the mRNA encodes the isoform i) of the SEQ ID NO: 1 , b) calculating a score value based on the comparison of the concentration of said isoform i) or ii) compared to a reference value, and c) returning the result as a prognosis classification of said disease based on the calculated score. DETAILED DESCRIPTION OF THE INVENTION
[0017] DEFINITIONS
[0018] The meaning of some terms and expressions as they are used in the present description are indicated below to aid in understanding. The inventions illustratively described herein may suitably be practiced in the absence of any element or elements, limitation or limitations, not specifically disclosed herein. Thus, for example, the terms "comprising", "including", "containing", etc. shall be read expansively and without limitation. Additionally, the terms and expressions employed herein have been used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention has been specifically disclosed by preferred embodiments and optional features, modification and variation of the inventions embodied therein herein disclosed may be resorted to by those skilled in the art. and that such modifications and variations are considered to be within the scope of this invention.
[0019] By "consisting of” is meant including, and limited to, whatever follows the phrase "consisting of”. Thus, the phrase "consisting of” indicates that the listed elements are required or mandatory, and that no other elements may be present.
[0020] The articles "a" and "an" are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, "an element" means one element or more than one element.
[0021] As used herein, "and / or" refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (or).
[0022] As used herein, the terms "diagnosis," "diagnosing" and the like are used interchangeably herein to encompass determining the likelihood that a subject will develop or has a condition or clinical state (e.g., responsiveness or non-responsiveness to therapy). These terms also encompass, for example, determining the level of clinical state (e.g. the level of responsiveness to therapy), as well as in the context of rational therapy, in which the diagnosis guides therapy, including initial selection of therapy, modification of therapy (e.g., adjustment of dose or dosage regimen), and the like. The term "diagnosis”, as used herein, refers both to the process of attempting to determine and / or identify a possible disease in a subject, i.e. the diagnostic procedure, and to the opinion reached by this process, i.e. the diagnostic opinion.
[0023] By "likelihood" is meant a measure of whether a subject with particularly measured or derived biomarker values actually has a condition or clinical state (or not) based on a given mathematical model. An increased likelihood for example may be relative or absolute and may be expressed qualitatively or quantitatively. For instance, an increased likelihood may be determined simply by determining the subject's measured or derived biomarker values for one or more cancer therapy biomarkers and placing the subject in an "increased likelihood" category, based upon previous population studies. The term "likelihood" is also used interchangeably herein with the term "probability".
[0024] As used herein, the terms "prognosis," "prognosing" and the like are used interchangeably herein to encompass determining the likely or expected development of a disease or of the chances of getting better or worse, and also the velocity or progression of a particular disease. The prognosis of a disease encompasses, for example, determining the level of clinical state (e.g. the level of responsiveness to therapy), as well as in the context of rational therapy, in which the diagnosis guides therapy, including initial selection of therapy, modification of therapy (e.g., adjustment of dose or dosage regimen), and the like.
[0025] The terms "subject", or "individual" are used herein interchangeably to refer to all the animals classified as mammals and includes but is not limited to domestic and farm animals, primates and humans, for example, human beings, non-human primates, cows, horses, pigs, sheep, goats, dogs, cats, or rodents. Preferably, the subject is a male or female human being of any age or race.
[0026] The phrase "non-human animal" as used herein refers to any vertebrate organism that is not a human. In some embodiments, the non-human animal is a mammal. In specific embodiments, the non-human animal is a rodent such as a rat or a mouse. As used herein, “screening” refers to the process used to evaluate and identify candidate agents that affect such disease.
[0027] As used herein, a “potentially active compound", “candidate agent” or “test compound” are used interchangeable and refer to a compound selected for screening to determine if it can function as a therapeutic agent for treating and / or preventing a particular disease. “Administering” includes any mean or time that is considered sufficient for an agent to interact with a cell or tissue non-human animal model and provoke a response in such cell, tissue or non-human animal model; it encompasses other terms such as “Incubating”, “Contacting” and / or “Treating”.
[0028] The term "biomarker" broadly refers to any detectable compound, such as a protein, a peptide, a proteoglycan, a glycoprotein, a lipoprotein, a carbohydrate, a lipid, a nucleic acid (e.g., DNA, such as cDNA or amplified DNA, or RNA, such as mRNA), an organic or inorganic chemical, a natural or synthetic polymer, a small molecule (e.g., a metabolite), or a discriminating molecule or discriminating fragment of any of the foregoing, that is present in or derived from a sample. "Derived from" as used in this context refers to a compound that, when detected, is indicative of a particular molecule being present in the sample. For example, detection of a particular cDNA can be indicative of the presence of a particular RNA transcript in the sample. As another example, detection of or binding to a particular antibody can be indicative of the presence of a particular antigen (e.g., protein) in the sample. For example, a discriminating molecule or fragment is a molecule or fragment that, when detected, indicates presence or abundance of an above-identified compound. A biomarker can, for example, be isolated from a sample, directly measured in a sample, or detected in or determined to be in a sample. A biomarker can, for example, be functional, partially functional, or non-functional. In specific embodiments, the "biomarkers" include "cancer therapy biomarkers", which are described in more detail below.
[0029] Multiple Sclerosis (MS) refers to an autoimmune disease in which the insulating covers of nerve cells in the brain and spinal cord are damaged. This damage disrupts the ability of parts of the nervous system to transmit signals, resulting in a range of signs and symptoms, including physical, mental, and sometimes psychiatric problems. Symptoms include double vision, vision loss, eye pain, muscle weakness, and loss of sensation or coordination. MS debut either as partially or as fully reversible episodes of neurological deficits (relapses) in the clinical form known as relapsing-remitting MS (RRMS). A great proportion of these evolve to the secondary progressive MS (SPMS) clinical form, with the development of permanent neurological deficits and the continuous progression of clinical disability. Approximately 15% of individuals with MS have a progressive course from disease onset, which is referred to as primary progressive MS (PPMS).
[0030] By “isoform” is meant any of two or more functionally similar proteins that have a similar but not identical amino acid sequence and are either encoded by different genes or by RNA transcripts from the same gene which have had different exons removed or that are the result of different post-transcriptional or post-translational processes (e.g. enzymatic cleavage by proteases, phosphorylations, acetylations, glycosylations and / or methylations). It includes sequences showing at least at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with respect to reference amino acid sequences of peptides or proteins. The degree of identity between two amino acid sequences can be determined by conventional methods, for example, by means of standard sequence alignment algorithms known in the state of the art, such as BLAST for example.
[0031] TAF1 as referred herein is a protein that in humans is encoded by the TAF1 gene, it stands for “TATA-box binding protein associated factor 1 (TAF1 )”. This protein is the largest and scaffolding subunit of transcription factor IID (TFIID), the general transcription factor that makes contacts with core promoter DNA elements to aid to define the transcription start site (TSS) and coordinates the formation of the RNAPII transcriptional preinitiation complex (PIC) on all protein-coding genes. It is also known as “Transcription initiation factor TFIID subunit 1 ”, also known as “Transcription initiation factor TFIID 250 kDa subunit (TAFII-250)” or “TBP-associated factor 250 kDa (p250)”. It can also be found as Gene ID: 6872, UniProt P21675. SEQ ID NO: 1 is the canonical sequence of TAF1 that can be found at UniProt P21675.
[0032] TAF1 L as referred herein is a protein that in humans is encoded by the TAF1 L gene, it stands for “Transcription initiation factor TFIID subunit 1 -like”. It arose in the primate lineage from retrotransposition of the transcript from the multi-exon TAF1 locus on the X chromosome. The gene is expressed in male germ cells, and the product has been shown to function interchangeably with the TAF1 product. The product of the gene TAF1 L is an isoform of TAF1 , that lacks the C-terminal from residues 1800 to 1893 and has got a 95% identity with the canonical sequence of TAF1 . It can also be found as Gene ID: 138474, UniProt Q8IZX4. SEQ ID NO: 2 is the canonical sequence of TAF1 L that can be found at UniProt Q8IZX4.
[0033] The person skilled in the art will understand that variants of the sequences named and shown in the examples of this invention are included, particularly the sequences of genes and proteins TAF1 and TAF1 L (SEQ ID No1 , SEQ ID NO: 2) are encompassed in this invention. The term "variant" refers to a protein or peptide substantially homologous to another protein or peptide, for example, to the peptides the amino acid sequences of which are shown in SEQ ID NOU to 3, to TAF1 and TAF1 L proteins, etc. A variant generally includes additions, deletions or substitutions of one or more amino acids. According to the present disclosure, said variants are recognized by autoantibodies against the protein or peptide in question. Variants of said peptides or proteins include peptides or proteins showing at least 25%, at least 40%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity with respect to certain amino acid sequences of peptides or proteins. The degree of identity between two amino acid sequences can be determined by conventional methods, for example, by means of standard sequence alignment algorithms known in the state of the art, such as BLAST for example.
[0034] The person skilled in the art will understand that the amino acid sequences (variants and isoforms) referred to in this description can be chemically modified, for example, by means of physiologically relevant chemical modifications, such as phosphorylations, acetylations, glycosylations or methylations.
[0035] As referred herein the “C-terminus” (also known as the carboxyl-terminus, carboxy-terminus, C-terminal tail, carboxy tail, C-terminal end, or COOH-terminus) is the end of an amino acid chain (protein or polypeptide), terminated by a free carboxyl group (-COOH). When the protein is translated from messenger RNA, it is created from N-terminus to C-terminus. The convention for writing peptide sequences is to put the C-terminal end on the right and write the sequence from N- to C-terminus. The C-terminus of TAF1 (SEQ ID NO: 1 ) begins at the region downstream of mid-exon 36, which is highly conserved across species suggesting the presence of functional domains at the C-terminal end of TAF1 (see Example 1 ).
[0036] The term "screening" is understood as the examination or testing of a group of individuals pertaining to the general population, at risk of suffering from Multiple Sclerosis (MS), with the objective of discriminating healthy individuals from those who are suffering from MS or other neurological or neurodegenerative disease or those who are at high, risk of suffering from said indications.
[0037] The expression "biological sample" refers to any sample which is taken from the body of the subject. Specifically, biological sample refers in the present invention to: brain tissue, nerve tissue, cerebrospinal fluid, cerebrospinal biopsy, urine, tears, sweat, faeces / stool, blood, serum or plasma, preferably is a serum sample.
[0038] The isoform TAF1 L (SEQ ID NO:2) has been found naturally in the testes, as it is expressed in male germ cells, so in a preferred embodiment the biological sample is not semen, testes biopsy, or epididymal fluid, or a sample from the reproductive male tract.
[0039] The term "up-regulated" or "over-expressed" of any of the bio markers or combinations thereof described in the present invention, refers to an increase in their expression level with, respect to a given “reference value”, "threshold value" or "cutoff value" by at least 5%, by at least 10%, by at least 15, by at least 20%, by at least
[0040] 25%, by at least 30%, by at least 35%, by at least 40%, by at least 45%, by at least
[0041] 50%, by at least 55%, by at least 60%, by at least 65%, by at least 70%, by at least
[0042] 75%, by at least 80%, by at least 85%, by at least 90%, by at least 95%, by at least
[0043] 100%, by at least 110%, by at least 120%, by at least 130%, by at least 140%, by at least 150%, or more. In addition, the term “up-regulated" or "over-expressed" of any of the biomarkers or combinations thereof described in the present invention, also refers to an increased in their expression level with respect to a given "threshold value" or "cutoff value" by at least about 1 .5-fold, about 2-fold, about 5-fold, about 10-fold, about 15-fold, about 20-fold, about 50-fold, or of about 100-fold.
[0044] The term “down-regulated” or "reduced expression" of any of the biomarkers or combinations thereof described in the present invention, refers to a reduction in their expression level with respect to a given “reference value”, "threshold value" or "cutoff value" by at least 5%, by at least 10%, by at least 15%, by at least 20%, by at least
[0045] 25%, by at least 30%, by at least 35%, by at least 40%, by at least 45%, by at least
[0046] 50%, by at least 55%, by at least 60%, by at least 65%, by at least 70%, by at least
[0047] 75%, by at least 80%, by at least 85%, by at least 90%, by at least 95%, by at least
[0048] 100%, by at least 110%, by at least 120%, by at least 130%, by at least 140%, by at least 150%, or more. In addition, the term "reduced expression" of any of the biomarkers or combinations thereof described in the present invention, also refers to a decreased in their expression level with respect to a given “reference value”, "threshold value" or "cutoff value" by at least about 1.5-fold, about 2-fold, about 5- fold, about 10-fold, about 15 -fold, about 20-fold, about 50-fold, or of about 100-fold.
[0049] The term “reference value”, "threshold value" or "cutoff value", when referring to the expression levels of the isoform of the invention described in the present invention, refers to a reference expression level indicative that a subject is likely to suffer from a Multiple Sclerosis with a given sensitivity and specificity if the expression levels of the patient are above or below said threshold or cut-off or reference levels, in the context of the present invention, said "threshold value" or "cutoff value" is a reference expression level taken from a healthy subject.
[0050] A variety of statistical and mathematical methods for establishing the reference value, threshold or cutoff level of expression are known in the prior art. A reference, threshold or cutoff expression level for a particular biomarker may be selected. The person skilled in the art will appreciate that these reference values or expression levels can be varied, for example, by moving along the ROC plot for a particular biomarker or combinations thereof, to obtain different values for sensitivity or specificity thereby affecting overall assay performance. For example, if the objective is to have a robust diagnostic method from a clinical point of view, we should try to have a high sensitivity. However, if the goal is to have a cost-effective method we should try to get a high specificity. The best cutoff refers to the value obtained from the ROC plot for a particular biomarker that produces the best sensitivity and specificity. Sensitivity and specificity values are calculated over the range of thresholds (cutoffs). Thus, the threshold or cutoff values can be selected such that the sensitivity and / or specificity are at least about 70 %, and can be, for example, at least 75 %, at least 80 %, at least 85 %, at least 90 %, at least 95 %, at least 96 %, at least 97 %, at least 98 %, at least 99 % or at least 100% in at least 60 % of the patient population assayed, or in at least 65 %, 70 %, 75 % or 80 % of the patient population assayed.
[0051] In Vitro Method for diagnosing Multiple Sclerosis (MS)
[0052] Inventors show in the example 1 that Multiple Sclerosis patients have got an increased presence of the isoform of TAF1 (SEQ ID NO:1 ) with a deletion at the C-terminus. The presence of this isoform has been shown to induce an MS-like disease in transgenic mice overexpressing such isoforms (examples 2 and 3) and it is shown to interact with proteins encoded by MS related genes (examples 4, 5, 6, 7 and 8).
[0053] It is disclosed a method for diagnosing multiple sclerosis in a subject, hereinafter first method of the disclosure, comprising: a) detecting an isoform of the TATA-box binding protein associated factor 1 (TAF1 ) as established in SEQ ID NO: 1 , or an isoform of the corresponding mRNA, in a biological sample from a subject, wherein i) the isoform of SEQ ID NO: 1 comprises at least one deletion from residues: 1709, 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893 of SEQ ID NO:1 , or ii) the isoform of the mRNA encodes the isoform i) of SEQ ID NO: 1 , wherein the presence of the isoform i) or the presence of the isoform ii) indicates that the subject suffers from, or is at risk of suffering from, multiple sclerosis.
[0054] The isoform of the step a) is a variant of the SEQ ID NO: 1 that corresponds to the canonical sequence of TAF1 (UniProt P21675) that comprises at least one deletion on the C terminus of the protein, from residues: 1709, 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893.
[0055] In a preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises a deletion on the C-terminus from residues from residues 1771 to 1893. In a preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises a deletion on the C-terminus from residues from residues 1788 to 1893.
[0056] In a preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises a deletion on the C-terminus from residues from residues 1800 to 1893.
[0057] In a preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises a deletion on the C-terminus from residues from residues 1821 to 1893.
[0058] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1709 to residue 1893.
[0059] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1724 to residue 1893.
[0060] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1760 to residue 1893.
[0061] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0062] 1764 to residue 1893.
[0063] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0064] 1765 to residue 1893.
[0065] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1774 to residue 1893. In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0066] 1775 to residue 1893.
[0067] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0068] 1776 to residue 1893.
[0069] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0070] 1777 to residue 1893.
[0071] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0072] 1787 to residue 1893.
[0073] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0074] 1788 to residue 1893. In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1793 to residue 1893.
[0075] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1800 to residue 1893.
[0076] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1802 to residue 1893.
[0077] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1803 to residue 1893. In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1804 to residue 1893.
[0078] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0079] 1820 to residue 1893.
[0080] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0081] 1821 to residue 1893.
[0082] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1825 to residue 1893.
[0083] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0084] 1838 to residue 1893.
[0085] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0086] 1839 to residue 1893.
[0087] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0088] 1840 to residue 1893.
[0089] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0090] 1841 to residue 1893.
[0091] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1845 to residue 1893. In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1871 to residue 1893.
[0092] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1880 to residue 1893.
[0093] In a more preferred embodiment, isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0094] 1882 to residue 1893.
[0095] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0096] 1883 to residue 1893.
[0097] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0098] 1884 to residue 1893.
[0099] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1889 to residue 1893.
[0100] As the skilled person knows, a decreased detection of the C-terminal part of TAF1 can be observed with an antibody, said antibody having an epitope comprehended in that region. Examples of those antibodies are: an antibody that recognises the amino acid sequence from 1771 to 1821 , or an antibody that recognises the amino acid sequence from 1821 to 1871.
[0101] In a more preferred embodiment, the isoform of the step a) of the first method of the disclosure comprises, preferably consist of, the SEQ ID NO: 2. SEQ ID NO: 2 corresponds to the canonical sequence of TAF1 L (UniProt Q8IZX4) that comprises a deletion on the C terminus of the protein, from residues 1788 to 1893 and has got a 95% identity with SEQ ID NO: 1 .
[0102] As the skilled person knows, TAF1 L can be determined with an antibody that recognises the specific TAF1 L amino acid sequences from 1347 to 1356 or from 1815 to 1826.
[0103] In a preferred embodiment of the first method of the invention, step a) further comprises quantifying the expression levels of at least one biomarker selected from the list consisting of Ctsb, Atp13a2, C1 qc, Ina, C1 qa, C1 qb, H2-K1 , B2m, C4b, H2-D1 , Oasl2, 1117rb, Ifit3, Oas2, Spp1 , Gfap, Oasl a, Cd52, Serpina3n, Ccl4, CxcH O, Cst7, Klk6, Naaladl2, Phex, Ptgds, Hapln2, Anin, Ninj2, Enpp6, II33, Padi2, Trim59, Tppp3, Galnt6 and Abca8b, wherein:
[0104] - the increased expression levels of at least one biomarker selected from Ctsb, Atp13a2, C1 qc, Ina, C1qa, C1 qb, H2-K1 , B2m, C4b, H2-D1 , Oasl2, 1117rb, Ifit3, Oas2, Spp1 , Gfap, Oasl a, Cd52, Serpina3n, Ccl4, CxcH O and Cst7 with respect to a reference value, or
[0105] - the decreased expression levels of at least one biomarker selected from Klk6, Naaladl2, Phex, Ptgds, Hapln2, Anin, Ninj2, Enpp6, 1133, Padi2, Trim59, Tppp3, Galnt6 and Abca8b with respect to a reference value, indicates that the subject suffers from, or is at risk of suffering from, multiple sclerosis.
[0106] In a more preferred embodiment of the method of the invention, step a) further comprises quantifying the expression levels of at least one biomarker selected from the list consisting of: cathepsin B (CTSB), secreted phosphoprotein 1 (SPP1 ), ninjurin 2 (Ninj2), interleukin 33 ( II33) and / or beta-2 microglobulin (B2M).
[0107] In the methods of the present invention, the detection and quantification methods may be used nucleic acid amplification techniques, sequencing platforms, array and hybridization platforms, microscopy, flow cytometry, immunoassays, mass spectrometry, or a combination thereof.
[0108] Methods for quantifying protein expression are well known in the art. Suitable methods for determining the levels of a given protein include, without limitation, those described herein below. Preferred methods for determining the protein expression levels in the methods of the present invention are immunoassays. Various types of immunoassays are known to one skilled in the art for the quantitation of proteins of interest. These methods are based on the use of affinity reagents, which may be any antibody or other ligand specifically binding to the target protein or to a fragment thereof, wherein said affinity reagent is preferably labelled. For instance, the affinity reagent may be enzymatically labelled, or labelled with a radioactive isotope or with a fluorescent agent.
[0109] Affinity reagents may be any antibody or ligand specifically binding to the target protein or to a fragment thereof. Affinity ligands may include proteins, peptides, nucleic acid or peptide aptamers, and other target specific protein scaffolds, like antibody-mimetics. Specific antibodies against the protein markers used in the methods of the invention may be produced for example by immunizing a host with a protein of the present invention or a fragment thereof. Likewise, peptides specific against the protein markers used in the methods of the invention may be produced by screening synthetic peptide libraries.
[0110] Western blot or immunoblotting techniques allow comparison of relative abundance of proteins separated by an electrophoretic gel (e.g., native proteins by 3-D structure or denatured proteins by the length of the polypeptide). Immunoblotting techniques use antibodies (or other specific ligands in related techniques) to identify target proteins among a number of unrelated protein species. They involve identification of protein target via antigen-antibody (or protein-ligand) specific reactions. Proteins are typically separated by electrophoresis and transferred onto a sheet of polymeric material (generally nitrocellulose, nylon, or polyvinylidene difluoride). Dot and slot blots are simplified procedures in which protein samples are not separated by electrophoresis but immobilized directly onto a membrane.
[0111] Traditionally, quantification of proteins in solution has been carried out by immunoassays on a solid support. Said immunoassay may be for example an enzyme-linked immunosorbent assay (ELISA), a fluorescent immunosorbent assay (FIA), a chemiluminescence immunoassay (CIA), or a radioimmunoassay (RIA), an enzyme multiplied immunoassay, a solid phase radioimmunoassay (SPROA), a fluorescence polarization (FP) assay, a fluorescence resonance energy transfer (FRET) assay, a time -resolved fluorescence resonance energy transfer (TR-FRET) assay, a surface plasmon resonance (SPR) assay. Multiplex and any next generation versions of any of the above, are specifically encompassed. In a preferred embodiment, said immunoassay is an ELISA assay or any multiplex version thereof.
[0112] The term "solid support" as used herein refers to a solid inert surface or body to which a molecular species, such as a nucleic acid and polypeptides can be immobilized. Non-limiting examples of solid supports include glass surfaces, plastic surfaces, latex, dextran, polystyrene surfaces, polypropylene surfaces, polyacrylamide gels, gold surfaces, and silicon wafers. In some embodiments, the solid supports are in the form of membranes, chips or particles. For example, the solid support may be a glass surface (e.g., a planar surface of a flow cell channel). In some embodiments, the solid support may comprise an inert substrate or matrix which has been "functionalized", such as by applying a layer or coating of an intermediate material comprising reactive groups which permit covalent attachment to molecules such as polynucleotides. By way of non-limiting example, such supports can include polyacrylamide hydrogels supported on an inert substrate such as glass. The molecules (e.g., polynucleotides) can be directly covalently attached to the intermediate material (e.g., a hydrogel) but the intermediate material can itself be non-covalently attached to the substrate or matrix (e.g., a glass substrate). The support can include a plurality of particles or beads each having a different attached molecular species.
[0113] Flow cytometry is particularly valuable as it can determine not only whether a cell is expressing the protein of interest but also indicates the amount of protein expressed by a single cell based on intensity of fluorescence. Flow cytometry can also be used for quantification of proteins in liquid samples when capture antibodies are bound to beads or similar. Thus, the person skilled in the art knows how to adapt the technique to the detection of the sequences of the invention.
[0114] Other methods that can be used for quantification of proteins in solution are techniques based on mass spectrometry (MS), also called mass spectroscopy. The term “mass spectrometry (MS)-based methods” as used herein refers to mass spectrometry alone or coupled to other detection or separation methods, including gas chromatography combined with mass spectroscopy, liquid chromatography combined with mass spectroscopy, supercritical fluid chromatography combined with mass spectroscopy, ultra-performance liquid chromatography combined with mass spectrometry, MALDI combined with mass spectroscopy, ion spray spectroscopy combined with mass spectroscopy, capillary electrophoresis combined with mass spectrometry, NMR combined with mass spectrometry and IR combined with mass spectrometry. These MS-based methods may include single MS or tandem MS. In another preferred embodiment, protein expression levels are determined by liquid chromatography coupled to tandem mass spectrometry (LC-MS / MS) analysis.
[0115] Mass spectrometers operate by converting the analyte molecules to a charged (ionized) state, with subsequent analysis of the ions and any fragment ions that are produced during the ionization process, based on their mass to charge ratio (m / z). Several different technologies are available for both ionization and ion analysis, resulting in many different types of mass spectrometers with different combinations of these two processes. On the one hand, examples of ion sources include electrospray ionization source, atmospheric pressure chemical ionization source, atmospheric pressure photoionization or matrix-assisted laser desorption ionization (MALDI). On the other hand, mass spectrometers analyzers may be, but are not limited to, quadrupole analyzers, time-of- flight (TOF) analyzers, ion trap analyzers, orbitrap analyzers or hybrid analyzers, such as hybrid quadrupole time-of-flight (QTOF) analyzers, hybrid quadrupole-orbitrap analyzers, hybrid ion trap-orbitrap analyzers, hybrid triple quadrupole linear ion trap analyzers or trihybrid quadrupole-ion trap-orbitrap analyzers. In a preferred embodiment, the levels of the protein markers are determined by using a hybrid quadrupole-orbitrap analyzer.
[0116] Internal standards may be used in the MS analysis as it enables to correct for any losses or inefficiencies in the sample preparation process or for alterations in ionization efficiency, for instance those due to ion suppression. Stable or isobaric isotope versions of the analyte are ideal internal standards as they have almost identical chemical properties but are easily distinguished during MS. In a preferred embodiment, optionally in combination with one or more of the embodiments described above or below, the determination of the levels of the protein marker is conducted by an MS-based method using isotope / isobaric-labelled versions of the protein marker as internal standards. In addition, or alternatively, the extraction solvent may be spiked with compounds not detected in unspiked biological samples.
[0117] In one preferred example, the biomarker values relate to a level of abundance or activity of an expression product or other measurable molecule, quantified using a technique such as quantitative RT-PCR, sequencing, or the like. In this case, the biomarker values can be in the form of amplification amounts, or cycle times, which are a logarithmic representation of the concentration of the biomarker within a sample, as will be appreciated by persons skilled in the art and as will be described in more detail below. In other preferred examples, the biomarker values are quantified using immunofluorescence of cells containing the expression product.
[0118] Biomarkers can be assessed by determining biomarker nucleic acid transcript levels. In illustrative nucleic acid-based assays, nucleic acid is isolated from cells contained in the biological sample according to standard methodologies. The nucleic acid is typically fractionated (e.g., poly A RNA) or whole cell RNA. Where RNA is used as the subject of detection, it may be desired to convert the RNA to a complementary DNA. In some embodiments, the nucleic acid is amplified by a template-dependent nucleic acid amplification technique. A number of template dependent processes are available to amplify the cancer therapy biomarker sequences present in a given template sample. An exemplary nucleic acid amplification technique is the polymerase chain reaction (referred to as PCR). Briefly, in PCR, two primer sequences are prepared that are complementary to regions on opposite complementary strands of the biomarker sequence. An excess of deoxynucleotide triphosphates are added to a reaction mixture along with a DNA polymerase, e.g., Taq polymerase. If a cognate cancer therapy biomarker sequence is present in a sample, the primers will bind to the biomarker and the polymerase will cause the primers to be extended along the biomarker sequence by adding on nucleotides. By raising and lowering the temperature of the reaction mixture, the extended primers will dissociate from the biomarker to form reaction products, excess primers will bind to the biomarker and to the reaction products and the process is repeated. A reverse transcriptase PCR amplification procedure may be performed in order to quantify the amount of mRNA amplified. Alternative methods for reverse transcription utilize thermostable, RNA-dependent DNA polymerases. Polymerase chain reaction methodologies are well known in the art. In specific embodiments in which whole cell RNA is used, cDNA synthesis using whole cell RNA as a sample produces whole cell cDNA.
[0119] As known to the skilled person, the template-dependent amplification involves quantification of transcripts in real-time. For example, RNA or DNA may be quantified using the real-time PCR technique. By determining the concentration of the amplified products of the target DNA in PCR reactions that have completed the same number of cycles and are in their linear ranges, it is possible to determine the relative concentrations of the specific target sequence in the original DNA mixture. If the DNA mixtures are cDNAs synthesized from RNAs isolated from different tissues or cells, the relative abundance of the specific mRNA from which the target sequence was derived can be determined for the respective tissues or cells. This direct proportionality between the concentration of the PCR products and the relative mRNA abundance is only true in the linear range of the PCR reaction. The final concentration of the target DNA in the plateau portion of the curve is determined by the availability of reagents in the reaction mix and is independent of the original concentration of target DNA. In specific embodiments, multiplexed, tandem PCR (MT-PCR) is employed, which uses a two-step process for gene expression profiling from small quantities of RNA or DNA. In the first step, RNA is converted into cDNA and amplified using multiplexed gene specific primers. In the second step each individual gene is quantitated by real time PCR. Real-time PCR is typically performed using any PCR instrumentation available in the art. Typically, instrumentation used in real-time PCR data collection and analysis comprises a thermal cycler, optics for fluorescence excitation and emission collection, and optionally a computer and data acquisition and analysis software.
[0120] In some embodiments of RT-PCR assays, a TAQMAN® probe is used for quantitating nucleic acid. Such assays may use energy transfer ("ET"), such as fluorescence resonance energy transfer ("FRET"), to detect and quantitate the synthesized PCR product. Typically, the TAQMAN® probe comprises a fluorescent label (e.g., a fluorescent dye) coupled to one end (e.g., the 5'-end) and a quencher molecule is coupled to the other end (e.g., the 3'-end), such that the fluorescent label and the quencher are in close proximity, allowing the quencher to suppress the fluorescence signal of the dye via FRET. When a polymerase replicates the chimeric amplicon template to which the fluorescent labelled probe is bound, the 5'-nuclease of the polymerase cleaves the probe, decoupling the fluorescent label and the quencher so that label signal (such as fluorescence) is detected. Signal (such as fluorescence) increases with each PCR cycle proportionally to the amount of probe that is cleaved.
[0121] In addition to the TAQMAN® assays, other real-time PCR chemistries useful for detecting PCR products in the methods presented herein include, but are not limited to, Molecular Beacons, Scorpion probes and intercalating dyes, such as SYBR Green, EvaGreen, thiazole orange, YO-PRO, TO-PRO, etc. For example, Molecular Beacons, like TAQMAN® probes, use FRET to detect and quantitate a PCR product via a probe having a fluorescent label (e.g., a fluorescent dye) and a quencher attached at the ends of the probe. Unlike TAQMAN® probes, however, Molecular Beacons remain intact during the PCR cycles.
[0122] In some embodiments, labels that can be used on the FRET probes include colorimetric and fluorescent dyes such as Alexa Fluor dyes, BODIPY dyes, such as BODIPY FL; Cascade Blue; Cascade Yellow; coumarin and its derivatives, such as 7-amino-4- methylcoumarin, aminocoumarin and hydroxycoumarin; cyanine dyes, such as Cy3 and Cy5; eosins and erythrosins; fluorescein and its derivatives, such as fluorescein isothiocyanate; macrocyclic chelates of lanthanide ions, such as Quantum Dye™; Marina Blue; Oregon Green; rhodamine dyes, such as rhodamine red, tetramethylrhodamine and rhodamine 6G; Texas Red; fluorescent energy transfer dyes, such as thiazole orange-ethidium heterodimer; and, TOTAB.
[0123] Specific examples of dyes include, but are not limited to, those identified above and the following: Alexa Fluor 350, Alexa Fluor 405, Alexa Fluor 430, Alexa Fluor 488, Alexa Fluor 500. Alexa Fluor 514, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor 555, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 610, Alexa Fluor 633, Alexa Fluor 647, Alexa Fluor 660, Alexa Fluor 680, Alexa Fluor 700, and, Alexa Fluor 750; amine-reactive BODIPY dyes, such as BODIPY 493 / 503, BODIPY 530 / 550, BODIPY 558 / 568, BODIPY 564 / 570, BODIPY 576 / 589, BODIPY 581 / 591 , BODIPY 630 / 650, BODIPY 650 / 655, BODIPY FL, BODIPY R6G, BODIPY TMR, and, BODIPY-TR; Cy3, Cy5, 6-FAM, Fluorescein Isothiocyanate, HEX, 6-JOE, Oregon Green 488, Oregon Green 500, Oregon Green 514, Pacific Blue, REG, Rhodamine Green, Rhodamine Red, Renographin, ROX, SYPRO, TAM RA, 2', 4', 5', 7'- Tetrabromosulfonefluorescein, and TET.
[0124] Examples of dye / quencher pairs (i.e., donor / acceptor pairs) include, but are not limited to, fluorescein / tetramethylrhodamine; lAEDANS / fluorescein; EDANS / dabcyl; fluorescein / fluorescein; BODIPY FL / BODIPY FL; fluorescein / QSY 7 or QSY 9 dyes. When the donor and acceptor are the same, FRET may be detected, in some embodiments, by fluorescence depolarization. Certain specific examples of dye / quencher pairs (i.e., donor / acceptor pairs) include, but are not limited to, Alexa Fluor 350 / Alexa Fluor488; Alexa Fluor 488 / Alexa Fluor 546; Alexa Fluor 488 / Alexa Fluor 555; Alexa Fluor 488 / Alexa Fluor 568; Alexa Fluor 488 / Alexa Fluor 594; Alexa Fluor 488 / Alexa Fluor 647; Alexa Fluor 546 / Alexa Fluor 568; Alexa Fluor 546 / Alexa Fluor 594; Alexa Fluor 546 / Alexa Fluor 647; Alexa Fluor 555 / Alexa Fluor 594; Alexa Fluor 555 / Alexa Fluor 647; Alexa Fluor 568 / Alexa Fluor 647; Alexa Fluor 594 / Alexa Fluor 647; Alexa Fluor 350 / QSY35; Alexa Fluor 350 / dabcyl; Alexa Fluor 488 / QSY 35; Alexa Fluor 488 / dabcyl; Alexa Fluor 488 / QSY 7 or QSY 9; Alexa Fluor 555 / QSY 7 or QSY9; Alexa Fluor 568 / QSY 7 or QSY 9; Alexa Fluor 568 / QSY 21 ; Alexa Fluor 594 / QSY 21 ; and Alexa Fluor 647 / QSY 21 .
[0125] In some realizations, for example, in a multiplex reaction in which two or more moieties (such as amplicons) are detected simultaneously, each probe comprises a detectably different dye such that the dyes may be distinguished when detected simultaneously in the same reaction. One skilled in the art can select a set of detectably different dyes for use in a multiplex reaction. In some embodiments, multiple target cancer therapy biomarker polynucleotides are detected and / or quantitated in a single multiplex reaction. In some embodiments, each probe that is targeted to a different cancer therapy biomarker polynucleotide is spectrally distinguishable when released from the probe. Thus, each target cancer therapy biomarker polynucleotide is detected by a unique fluorescence signal.
[0126] Specific examples of fluorescently labeled ribonucleotides useful in the preparation of real-time PCR probes for use in some embodiments of the methods described herein are available from Molecular Probes (Invitrogen), and these include, Alexa Fluor 488-5- UTP, Fluorescein- 12-UTP, BODIPY FL- 14-UTP, BODIPY TMR- 14-UTP, Tetramethylrhodamine-6- UTP, Alexa Fluor 546- 14-UTP, Texas Red-5-UTP, and BODIPY TR-14-UTP.
[0127] Examples of fluorescently labeled deoxyribonucleotides useful in the preparation of real-time PCR probes for use in the methods described herein include Dinitrophenyl (DNP)- I'-dUTP, Cascade Blue-7-dUTP, Alexa Fluor 488-5-dUTP, Fluorescein-12- dUTP, Oregon Green 488-5-dUTP, BODIPY FL-14-dUTP, Rhodamine Green-5-dUTP, Alexa Fluor 532-5-dUTP, BODIPY TMR- 14-dUTP, Tetramethylrhodamine-6-dUTP, Alexa Fluor 546- 14- dUTP, Alexa Fluor 568-5-dUTP, Texas Red- 12-dUTP, Texas Red-5-dUTP, BODIPY TR- 14-dUTP, Alexa Fluor 594-5-dUTP, BODIPY 630 / 650- 14-dUTP, BODIPY 650 / 665- 14-dUTP; Alexa Fluor 488-7-OBEA-dCTP, Alexa Fluor 546-16-OBEA-dCTP, Alexa Fluor 594-7-OBEA-dCTP, Alexa Fluor 647-12-OBEA-dCTP.
[0128] In certain embodiments, target nucleic acids are quantified using blotting techniques, which are well known to those of skill in the art. Southern blotting involves the use of DNA as a target, whereas Northern blotting involves the use of RNA as a target. Each provides different types of information, although cDNA blotting is analogous, in many embodiments, to blotting or RNA species. Briefly, a probe is used to target a DNA or RNA species that has been immobilized on a suitable matrix, often a filter of nitrocellulose. The different species should be spatially separated to facilitate analysis. This often is accomplished by gel electrophoresis of nucleic acid species followed by "blotting" on to the filter. Subsequently, the blotted target is incubated with a probe (usually labeled) under conditions that promote denaturation and rehybridization. Because the probe is designed to base pair with the target, the probe will bind a portion of the target sequence under renaturing conditions. Unbound probe is then removed, and detection is accomplished as described above. Following detection / quantification, one may compare the results seen in a given subject with a control reaction or a statistically significant reference group or population of control subjects as defined herein. In this way, it is possible to correlate the amount of cancer therapy biomarker nucleic acid detected with the progression or severity of the disease.
[0129] Chip hybridization utilizes biomarker specific oligonucleotides attached to a solid substrate, which may consist of a particulate solid phase such as nylon filters, glass slides or silicon chips designed as a microarray. Microarrays are known in the art and consist of a surface to which probes that correspond in sequence to gene products (such as cDNAs) can be specifically hybridized or bound at a known position for the detection of biomarker gene expression.
[0130] Quantification of the hybridization complexes is well known in the art and may be achieved by any one of several approaches. These approaches are generally based on the detection of a label or marker, such as any radioactive, fluorescent, biological or enzymatic tags or labels of standard use in the art. A label can be applied to either the oligonucleotide probes or the RNA derived from the biological sample. In certain embodiments, the cancer therapy biomarker is a target RNA (e.g., mRNA) or a DNA copy of the target RNA whose level or abundance is measured using at least one nucleic acid probe that hybridizes under at least low, medium, or high stringency conditions to the target RNA or to the DNA copy, wherein the nucleic acid probe comprises at least 15 (e.g., 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, or more) contiguous nucleotides of MS biomarker polynucleotide. In some embodiments, the measured level or abundance of the target RNA or its DNA copy is normalized to the level or abundance of a reference RNA or a DNA copy of the reference RNA. Suitably, the nucleic acid probe is immobilized on a solid or semi-solid support. In illustrative examples of this type, the nucleic acid probe forms part of a spatial array of nucleic acid probes. In some embodiments, the level of nucleic acid probe that is bound to the target RNA or to the DNA copy is measured by hybridization (e.g., using a nucleic acid array). In other embodiments, the level of nucleic acid probe that is bound to the target RNA or to the DNA copy is measured by nucleic acid amplification (e.g., using a polymerase chain reaction (PCR)). In still other embodiments, the level of nucleic acid probe that is bound to the target RNA or to the DNA copy is measured by nuclease protection assay.
[0131] In general, mRNA quantification is suitably effected alongside a calibration curve so as to enable accurate mRNA determination. Furthermore, quantifying transcript(s) originating from a biological sample is preferably effected by comparison to a control sample, which sample is characterized by a known expression pattern of the examined transcript(s).
[0132] The method of the invention, as it is understood by a person skilled in the art, does not claim to be correct in 100% of the analyzed samples. However, it requires that a statistically significant amount of the analyzed samples are classified correctly. The amount that is statistically significant can be established by a person skilled in the art by means of using different statistical significance measures obtained by statistical tests; illustrative, non-limiting examples of said statistical significance measures include determining confidence intervals, determining the p-value, etc. Preferred confidence intervals are at least 90%, at least 95%, at least 97%, at least 98%, at least 99%. The p-values are, preferably less than 0.1 , less than 0.05, less than 0.01 , less than 0.005 or less than 0.0001. The teachings of the present invention preferably allow correctly classifying at least 60%, at least 70%, at least 80%, or at least 90% of the subjects of a determining group or population analyzed.
[0133] It is further noted that the accuracy of the method of the invention can be further increased by additionally considering biochemical and environmental parameters and / or clinical characteristics of the patients like age, sex, a previous infection by Epstein-Barr virus (EBV), smoking, sun exposure / vitamin D, adolescent obesity and HLA profile, are considered traditional MS risk factors and included in risk scores.
[0134] In Vitro Method for prognosis of Multiple Sclerosis (MS)
[0135] Inventors have shown that there is a diminished detection of the C-terminal end of TAF1 in normal-appearing grey matter (NAGM) and normal-appearing white matter (NAWM) of individuals with MS and suggest an even more pronounced TAF1 alteration associated with shorter disease duration as well as to demyelinated lesions thus, indicating a correlation between the levels of the isoform and damage of the tissue (see Example 1 ).
[0136] It is disclosed a method for the prognosis of multiple sclerosis in a subject that suffers from primary progressive multiple sclerosis (PPMS) or secondary progressive multiple sclerosis (SPMS), hereinafter second method of the disclosure, comprising: a) detecting an isoform of the TAF1 as established in SEQ ID NO: 1 , or an isoform of the corresponding mRNA, in a biological sample from a subject, wherein i) the isoform of SEQ ID NO: 1 comprises at least one deletion from residues: 1709, 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893 of SEQ ID NO: 1 , or ii) the isoform of the mRNA encodes the isoform i) of SEQ ID NO: 1 ; b) calculating the abundance or concentration of isoform at step a) with a predetermined reference value for the same biological marker; wherein the abundance or concentration of the isoform i) or the isoform ii) indicates that the subject suffers poor prognosis for primary progressive multiple sclerosis (PPMS) or secondary progressive multiple sclerosis (SPMS).
[0137] The isoform of the step a) is a variant of the SEQ ID NO: 1 that corresponds to the canonical sequence of TAF1 (UniProt P21675) that comprises at least one deletion on the C terminus of the protein, from residues: 1709, 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893.
[0138] In a preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises a deletion on the C-terminus from residues from residues 1771 to 1893.
[0139] In a preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises a deletion on the C-terminus from residues from residues 1788 to 1893.
[0140] In a preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises a deletion on the C-terminus from residues from residues 1800 to 1893.
[0141] In a preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises a deletion on the C-terminus from residues from residues 1821 to 1893.
[0142] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1709 to residue 1893.
[0143] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1724 to residue 1893.
[0144] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1760 to residue 1893.
[0145] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1764 to residue 1893.
[0146] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1765 to residue 1893.
[0147] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1774 to residue 1893.
[0148] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1775 to residue 1893.
[0149] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1776 to residue 1893.
[0150] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1777 to residue 1893.
[0151] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1787 to residue 1893.
[0152] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1788 to residue 1893.
[0153] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1793 to residue 1893.
[0154] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1800 to residue 1893.
[0155] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1802 to residue 1893.
[0156] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1803 to residue 1893.
[0157] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1804 to residue 1893.
[0158] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1820 to residue 1893.
[0159] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1821 to residue 1893.
[0160] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1825 to residue 1893.
[0161] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1838 to residue 1893.
[0162] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1839 to residue 1893. In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1840 to residue 1893.
[0163] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1841 to residue 1893.
[0164] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1845 to residue 1893.
[0165] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1871 to residue 1893.
[0166] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1880 to residue 1893.
[0167] In a more preferred embodiment, isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1882 to residue 1893.
[0168] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1883 to residue 1893.
[0169] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1884 to residue 1893.
[0170] In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1889 to residue 1893. In a more preferred embodiment, the isoform of the step a) of the second method of the disclosure comprises, preferably consist of, the SEQ ID NO: 2. SEQ ID NO: 2 corresponds to the canonical sequence of TAF1 L (UniProt Q8IZX4) that comprises a deletion on the C terminus of the protein, from residues 1788 to 1893 and has got a 95% identity with SEQ ID NO: 1 .
[0171] In a preferred embodiment of the second method of the invention, step a) further comprises quantifying the expression levels of at least one biomarker selected from the list consisting of Ctsb, Atp13a2, C1qc, Ina, C1qa, C1qb, H2-K1 , B2m, C4b, H2-D1 , Oasl2, 1117rb, Ifit3, Oas2, Spp1 , Gfap, Oasl a, Cd52, Serpina3n, Ccl4, CxcH O, Cst7, Klk6, Naaladl2, Phex, Ptgds, Hapln2, Anin, Ninj2, Enpp6, II33, Padi2, Trim59, Tppp3, Galnt6 and Abca8b, wherein:
[0172] - the increased expression levels of at least one biomarker selected from Ctsb, Atp13a2, C1qc, Ina, C1qa, C1qb, H2-K1 , B2m, C4b, H2-D1 , Oasl2, 1117rb, Ifit3, Oas2, Spp1 , Gfap, Oasla, Cd52, Serpina3n, Ccl4, CxcH O and Cst7 with respect to a reference value, or
[0173] - the decreased expression levels of at least one biomarker selected from Klk6, Naaladl2, Phex, Ptgds, Hapln2, Anin, Ninj2, Enpp6, 1133, Padi2, Trim59, Tppp3, Galnt6 and Abca8b with respect to a reference value, indicates that the subject suffers poor prognosis for primary progressive multiple sclerosis (PPMS) or secondary progressive multiple sclerosis (SPMS).
[0174] In a more preferred embodiment of the method of the invention, step a) further comprises quantifying the expression levels of at least one biomarker selected from the list consisting of: cathepsin B (CTSB), secreted phosphoprotein 1 (SPP1 ), ninjurin 2 (Ninj2), interleukin 33 ( II33) and / or beta-2 microglobulin (B2M).
[0175] In a more preferred embodiment of the first and second method of the invention the subject is a human being.
[0176] In a more preferred embodiment of the first and second method of the invention the biological sample isolated from the subject is brain tissue, nerve tissue, cerebrospinal fluid, cerebrospinal biopsy, urine, tears, sweat, faeces / stool, blood, serum or plasma, and even more preferably is a serum sample. Method for screening of an active compound for treating and / or preventing Multiple Sclerosis (MS)
[0177] Inventors have shown in example 2, 3, 5, 6 and 7 that the expression of the isoform in a non-human model of the invention (a mouse) and in cell lines (example 7) reproduces a MS-like model of the disease and thus it can be used to identify compounds that alleviate, reduce or even eliminate the progression of the disease.
[0178] It is disclosed a method for screening of an active compound for treating and / or preventing Multiple Sclerosis (MS), hereinafter third method of the disclosure, comprising: a) administering a potentially active compound for the treatment and / or prevention of multiple sclerosis to a cell, a tissue or a non-human animal model that expresses an isoform of the TAF1 as established in SEQ ID NO: 1 that comprises at least one deletion from residues: 1709, 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893 of SEQ ID NO: 1 , b) evaluating the change in the gene expression, phenotype change and / or histological change of the cell, tissue or non-human animal model in response to the potentially active compound.
[0179] The isoform of the step a) is a variant of the SEQ ID NO: 1 that corresponds to the canonical sequence of TAF1 (UniProt P21675) that comprises at least one deletion on the C terminus of the protein, from residues: 1709, 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893.
[0180] In a preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises a deletion on the C-terminus from residues from residues 1771 to 1893.
[0181] In a preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises a deletion on the C-terminus from residues from residues 1788 to 1893. In a preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises a deletion on the C-terminus from residues from residues 1800 to 1893.
[0182] In a preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises a deletion on the C-terminus from residues from residues 1821 to 1893.
[0183] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1709 to residue 1893.
[0184] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1724 to residue 1893.
[0185] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1760 to residue 1893.
[0186] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0187] 1764 to residue 1893.
[0188] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0189] 1765 to residue 1893.
[0190] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1774 to residue 1893.
[0191] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1775 to residue 1893.
[0192] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0193] 1776 to residue 1893.
[0194] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0195] 1777 to residue 1893.
[0196] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0197] 1787 to residue 1893.
[0198] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0199] 1788 to residue 1893.
[0200] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1793 to residue 1893.
[0201] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1800 to residue 1893.
[0202] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1802 to residue 1893.
[0203] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1803 to residue 1893.
[0204] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1804 to residue 1893.
[0205] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0206] 1820 to residue 1893.
[0207] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0208] 1821 to residue 1893.
[0209] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1825 to residue 1893.
[0210] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0211] 1838 to residue 1893.
[0212] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0213] 1839 to residue 1893.
[0214] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0215] 1840 to residue 1893.
[0216] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0217] 1841 to residue 1893.
[0218] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1845 to residue 1893. In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1871 to residue 1893.
[0219] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1880 to residue 1893.
[0220] In a more preferred embodiment, isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0221] 1882 to residue 1893.
[0222] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0223] 1883 to residue 1893.
[0224] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0225] 1884 to residue 1893.
[0226] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1889 to residue 1893.
[0227] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, the SEQ ID NO: 2. SEQ ID NO: 2 corresponds to the canonical sequence of TAF1 L (UniProt Q8IZX4) that comprises a deletion on the C terminus of the protein, from residues 1788 to 1893 and has got a 95% identity with SEQ ID NO: 1 .
[0228] In a more preferred embodiment, the isoform of the step a) of the third method of the disclosure comprises, preferably consist of, the SEQ ID NO: 3.
[0229] A variety of cells and tissues of different origins can be used for the screening method of the invention, they may be isolated from the “non-human animal model” or they might be isolated from a primate such as monkeys, chimpanzees, orangutans, gorillas and humans. In a preferred embodiment, the cell and / or tissue of the third method is of primate origin an even more preferred embodiment is from human origin.
[0230] A variety of animals can be used for generating the “non-human animal model” of the method. It includes, for example, non-human mammals commonly used in conventional experiments like mammals such as rodents (including mice, rats, hamsters and guinea pigs), cats, dogs, rabbits, farm animals including cows, horses, goats, sheep, pigs, etc., and primates (including monkeys, chimpanzees, orangutans and gorillas) are included within the definition. Tissues, cells and their progeny of a biological entity obtained in vivo or cultured in vitro are also encompassed. Preferred animals include, for example, rodents such as mice, nude mice, rats, nude rats, guinea pigs, hamsters, and rabbits. For the convenience of breeding and manipulation, more preferred animals include, for example, mice, nude mice, rats, and nude rats. In a more preferred embodiment, the non-human animal model is a rodent, preferably a mouse.
[0231] The “non-human animal” can be obtained by methods already known in the state of the art like gene editing (e.g. BE3, CRISPR / Cas9, TALEN or ZFN).
[0232] Some of the phenotype changes include, but are not limited to, measuring tail tremor, lowered pelvis during walking, kyphosis, body shivering, abnormal hind limb clasping, CNS neuroinflammation, demyelination.
[0233] In a preferred embodiment of the third method of the invention, step b) further comprises quantifying the expression levels of at least one biomarker selected from the list consisting of Ctsb, Atp13a2, C1 qc, Ina, C1 qa, C1 qb, H2-K1 , B2m, C4b, H2-D1 , Oasl2, 1117rb, Ifit3, Oas2, Spp1 , Gfap, Oasl a, Cd52, Serpina3n, Ccl4, CxcH O, Cst7, Klk6, Naaladl2, Phex, Ptgds, Hapln2, Anin, Ninj2, Enpp6, II33, Padi2, Trim59, Tppp3, Galnt6 and Abca8b, wherein:
[0234] - the increased expression levels of at least one biomarker selected from Ctsb, Atp13a2, C1 qc, Ina, C1qa, C1 qb, H2-K1 , B2m, C4b, H2-D1 , Oasl2, 1117rb, Ifit3, Oas2, Spp1 , Gfap, Oasl a, Cd52, Serpina3n, Ccl4, CxcH O and Cst7 with respect to a reference value, or
[0235] - the decreased expression levels of at least one biomarker selected from Klk6, Naaladl2, Phex, Ptgds, Hapln2, Anin, Ninj2, Enpp6, II33, Padi2, Trim59, Tppp3, Galnt6 and Abca8b with respect to a reference value, indicates that there is a response to the potentially active compound.
[0236] In a more preferred embodiment of the third method of the invention, step b) further comprises quantifying the expression levels of at least one biomarker selected from the list consisting of: cathepsin B (CTSB), secreted phosphoprotein 1 (SPP1 ), ninjurin 2 (Ninj2), interleukin 33 ( II33) and / or beta-2 microglobulin (B2M).
[0237] In a more preferred embodiment of the third method of the invention, the method is in vitro and the cell or tissue is from primate origin, preferably human.
[0238] Computer implemented method for diagnosing Multiple Sclerosis (MS)
[0239] The methods of the present invention or any of the steps thereof might be implemented by a computer. Therefore, a further aspect of the invention refers to a computer implemented method, wherein the method is any of the methods disclosed herein or any combination thereof.
[0240] It is disclosed herein a computer implemented method for diagnosing multiple sclerosis (MS), hereinafter the fourth method of the disclosure, comprising: a) receiving or storing or having access to data comprising the levels of an isoform of the TAF1 as established in SEQ ID NO: 1 or an isoform of the corresponding mRNA in a biological sample from the subject, wherein i) the protein isoform comprises at least one deletion from residues: 1709, 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893 of SEQ ID NO:1 , or ii) the isoform of the mRNA encodes the isoform i) of the SEQ ID NO: 1 , b) calculating a score value based on the comparison of the concentration of said isoform i) or ii) compared to a reference value, and c) returning as a result a risk or an indication of whether a subject has got multiple sclerosis or is at risk of suffering said disease based on the calculated score.
[0241] It is noted that any computer program capable of implementing any of the methods of the present invention or used to implement any of these methods or any combination thereof, also forms part of the present invention. This computer program is typically directly loadable into the internal memory of a digital computer, comprising software code portions for performing the steps of comparing the levels of the protein markers as described in the invention, from the one or more biological samples of a subject, with a reference value and determining the presence or likelihood of having subclinical atherosclerosis, when said product is run on a computer.
[0242] It is also noted that any device or apparatus comprising means for carrying out the steps of any of the methods of the present invention or any combination thereof, or carrying a computer program capable of, or for implementing any of the methods of the present invention or any combination thereof, is included as forming part of the present specification.
[0243] The methods of the invention may also comprise the storing of the method results in a data carrier, preferably wherein said data carrier is a computer readable medium. The present invention further relates to a computer-readable storage medium having stored thereon a computer program of the invention or the results of any of the methods of the invention. As used herein, “a computer readable medium” can be any apparatus that may include, store, communicate, propagate, or transport the results of the determination of the method of the invention. The medium can be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system (or apparatus or device) or a propagation medium.
[0244] In a preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises a deletion on the C-terminus from residues from residues 1771 to 1893.
[0245] In a preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises a deletion on the C-terminus from residues from residues 1788 to 1893.
[0246] In a preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises a deletion on the C-terminus from residues from residues 1800 to 1893. In a preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises a deletion on the C-terminus from residues from residues 1821 to 1893.
[0247] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1709 to residue 1893.
[0248] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1724 to residue 1893.
[0249] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1760 to residue 1893.
[0250] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0251] 1764 to residue 1893.
[0252] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0253] 1765 to residue 1893.
[0254] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0255] 1774 to residue 1893.
[0256] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0257] 1775 to residue 1893.
[0258] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0259] 1776 to residue 1893. In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1777 to residue 1893.
[0260] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0261] 1787 to residue 1893.
[0262] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0263] 1788 to residue 1893.
[0264] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1793 to residue 1893.
[0265] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1800 to residue 1893.
[0266] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0267] 1802 to residue 1893.
[0268] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0269] 1803 to residue 1893.
[0270] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0271] 1804 to residue 1893.
[0272] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1820 to residue 1893.
[0273] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1821 to residue 1893.
[0274] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1825 to residue 1893.
[0275] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0276] 1838 to residue 1893.
[0277] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0278] 1839 to residue 1893.
[0279] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0280] 1840 to residue 1893.
[0281] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0282] 1841 to residue 1893.
[0283] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1845 to residue 1893.
[0284] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1871 to residue 1893.
[0285] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1880 to residue 1893.
[0286] In a more preferred embodiment, isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0287] 1882 to residue 1893.
[0288] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0289] 1883 to residue 1893.
[0290] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0291] 1884 to residue 1893.
[0292] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1889 to residue 1893.
[0293] In a more preferred embodiment, the isoform of the step a) of the fourth method of the disclosure comprises, preferably consist of, the SEQ ID NO: 2. SEQ ID NO: 2 corresponds to the canonical sequence of TAF1 L (UniProt Q8IZX4) that comprises a deletion on the C terminus of the protein, from residues 1788 to 1893 and has got a 95% identity with SEQ ID NO: 1 .
[0294] In a preferred embodiment of the fourth method of the invention, step a) further comprises quantifying the expression levels of at least one biomarker selected from the list consisting of Ctsb, Atp13a2, C1 qc, Ina, C1 qa, C1 qb, H2-K1 , B2m, C4b, H2-D1 , Oasl2, 1117rb, Ifit3, Oas2, Spp1 , Gfap, Oasl a, Cd52, Serpina3n, Ccl4, CxcH O, Cst7, Klk6, Naaladl2, Phex, Ptgds, Hapln2, Anin, Ninj2, Enpp6, II33, Padi2, Trim59, Tppp3, Galnt6 and Abca8b, wherein:
[0295] - the increased expression levels of at least one biomarker selected from Ctsb, Atp13a2, C1 qc, Ina, C1qa, C1 qb, H2-K1 , B2m, C4b, H2-D1 , Oasl2, 1117rb, Ifit3, Oas2, Spp1 , Gfap, Oasl a, Cd52, Serpina3n, Ccl4, CxcH O and Cst7 with respect to a reference value, or - the decreased expression levels of at least one biomarker selected from Klk6, Naaladl2, Phex, Ptgds, Hapln2, Anin, Ninj2, Enpp6, II33, Padi2, Trim59, Tppp3, Galnt6 and Abca8b with respect to a reference value, indicates that the subject suffers from, or is at risk of suffering from, multiple sclerosis.
[0296] In a more preferred embodiment of the method of the invention, step a) further comprises quantifying the expression levels of at least one biomarker selected from the list consisting of: cathepsin B (CTSB), secreted phosphoprotein 1 (SPP1 ), ninjurin 2 (Ninj2), interleukin 33 ( II33) and / or beta-2 microglobulin (B2M).
[0297] The processing device of the computer implemented method can then generate a representation of the indicator, for example by generating a sign or alphanumeric indication of the indicator, a graphical indication of a comparison of the indicator to one or more indicator references or an alphanumeric indication of the likely responsiveness of the subject to the MS therapy. Thus, in another preferred embodiment, the computer implemented method for diagnosing multiple sclerosis, further comprises the following step: d) returning a treatment decision based on the calculated score.
[0298] Computer implemented method for prognosis of Multiple Sclerosis (MS)
[0299] The methods of the present invention or any of the steps thereof might be implemented by a computer. Therefore, a further aspect of the invention refers to a computer implemented method, wherein the method is any of the methods disclosed herein or any combination thereof.
[0300] It is disclosed herein a computer implemented method for diagnosing multiple sclerosis (MS), hereinafter the fifth method of the disclosure, comprising: a) receiving or storing or having access to data comprising the levels of an isoform of the TAF 1 as established in SEQ ID NO: 1 or an isoform of the corresponding mRNA in a biological sample from the subject, wherein i) the protein isoform comprises at least one deletion from residues: 1709, 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893 of SEQ ID NO: 1 , or ii)the isoform of the mRNA encodes the isoform i) of the SEQ ID NO: 1 , b) calculating a score value based on the comparison of the concentration of said isoform i) or ii) compared to a reference value, and c) returning as a result a risk or an indication of whether a subject suffers poor prognosis for primary progressive multiple sclerosis (PPMS), or secondary progressive multiple sclerosis (SPMS) based on the calculated score.
[0301] In a preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises a deletion on the C-terminus from residues from residues 1771 to 1893.
[0302] In a preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises a deletion on the C-terminus from residues from residues 1788 to 1893.
[0303] In a preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises a deletion on the C-terminus from residues from residues 1800 to 1893.
[0304] In a preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises a deletion on the C-terminus from residues from residues 1821 to 1893.
[0305] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1709 to residue 1893.
[0306] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1724 to residue 1893.
[0307] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1760 to residue 1893.
[0308] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1764 to residue 1893.
[0309] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1765 to residue 1893.
[0310] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0311] 1774 to residue 1893.
[0312] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0313] 1775 to residue 1893.
[0314] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0315] 1776 to residue 1893.
[0316] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0317] 1777 to residue 1893.
[0318] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0319] 1787 to residue 1893.
[0320] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0321] 1788 to residue 1893.
[0322] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1793 to residue 1893. In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1800 to residue 1893.
[0323] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0324] 1802 to residue 1893.
[0325] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0326] 1803 to residue 1893.
[0327] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0328] 1804 to residue 1893.
[0329] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0330] 1820 to residue 1893.
[0331] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0332] 1821 to residue 1893.
[0333] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1825 to residue 1893.
[0334] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0335] 1838 to residue 1893.
[0336] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0337] 1839 to residue 1893. In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0338] 1840 to residue 1893.
[0339] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0340] 1841 to residue 1893.
[0341] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1845 to residue 1893.
[0342] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1871 to residue 1893.
[0343] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1880 to residue 1893.
[0344] In a more preferred embodiment, isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0345] 1882 to residue 1893.
[0346] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0347] 1883 to residue 1893.
[0348] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue
[0349] 1884 to residue 1893.
[0350] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, a deletion on the C-terminus from residue 1889 to residue 1893.
[0351] In a more preferred embodiment, the isoform of the step a) of the fifth method of the disclosure comprises, preferably consist of, the SEQ ID NO: 2. SEQ ID NO: 2 corresponds to the canonical sequence of TAF1 L (UniProt Q8IZX4) that comprises a deletion on the C terminus of the protein, from residues 1788 to 1893 and has got a 95% identity with SEQ ID NO: 1 .
[0352] In a preferred embodiment of the fifth method of the invention, step a) further comprises quantifying the expression levels of at least one biomarker selected from the list consisting of Ctsb, Atp13a2, C1 qc, Ina, C1 qa, C1 qb, H2-K1 , B2m, C4b, H2-D1 , Oasl2, 1117rb, Ifit3, Oas2, Spp1 , Gfap, Oasl a, Cd52, Serpina3n, Ccl4, CxcH O, Cst7, Klk6, Naaladl2, Phex, Ptgds, Hapln2, Anin, Ninj2, Enpp6, II33, Padi2, Trim59, Tppp3, Galnt6 and Abca8b, wherein:
[0353] - the increased expression levels of at least one biomarker selected from Ctsb, Atp13a2, C1 qc, Ina, C1qa, C1 qb, H2-K1 , B2m, C4b, H2-D1 , Oasl2, 1117rb, Ifit3, Oas2, Spp1 , Gfap, Oasl a, Cd52, Serpina3n, Ccl4, CxcH O and Cst7 with respect to a reference value, or
[0354] - the decreased expression levels of at least one biomarker selected from Klk6, Naaladl2, Phex, Ptgds, Hapln2, Anin, Ninj2, Enpp6, 1133, Padi2, Trim59, Tppp3, Galnt6 and Abca8b with respect to a reference value, indicates that the subject suffers poor prognosis for primary progressive multiple sclerosis (PPMS) or secondary progressive multiple sclerosis (SPMS).
[0355] In a more preferred embodiment of the method of the invention, step a) further comprises quantifying the expression levels of at least one biomarker selected from the list consisting of: cathepsin B (CTSB), secreted phosphoprotein 1 (SPP1 ), ninjurin 2 (Ninj2), interleukin 33 ( II33) and / or beta-2 microglobulin (B2M).
[0356] The processing device of the computer implemented method can then generate a representation of the indicator, for example by generating a sign or alphanumeric indication of the indicator, a graphical indication of a comparison of the indicator to one or more indicator references or an alphanumeric indication of the likely responsiveness of the subject to the MS therapy. Thus, in another preferred embodiment, the computer implemented method for diagnosing multiple sclerosis, further comprises the following step: d) Returning a treatment decision based on the calculated score.
[0357] The methods of the present invention, and particularly the second and fifth method of the invention, also extend to the treatment of a subject with a MS. Thus, also provided is a method for treating MS in a subject, the method comprising, consisting or consisting essentially of performing the method described above and herein for determining an indicator used in assessing a likelihood of a subject with MS responding to MS therapy; and exposing the subject to a MS therapy on the basis that the indicator is at least partially indicative of a positive response to MS therapy.
[0358] MS therapy includes treatment with steroids, immunosuppressants, baclofen, gabapentin, duloxetine, gabapentin or carbamazepine, and / or amitriptyline.
[0359] “Treatment” and “treating” refer to administration or application of a therapeutic agent to a subject or performance of a procedure or modality on a subject for the purpose of obtaining a therapeutic benefit of a disease or health-related condition. For example, a treatment may include administrating chemotherapy, immunotherapy, radiotherapy, performance of surgery, or any combination thereof.
[0360] The term “therapeutic benefit” or “therapeutically effective” as used throughout this application refers to anything that promotes or enhances the well-being of the subject with respect to the medical treatment of this condition. This includes, but is not limited to, a reduction in the frequency or severity of the signs or symptoms of a disease. An effective response of a patient or a patient’s “responsiveness” to treatment refers to the clinical or therapeutic benefit imparted to a patient at risk for, or suffering from, a disease or disorder. Such benefit may include cellular or biological responses, a complete response, a partial response, a stable disease (without progression or relapse), or a response with a later relapse. For example, treatment of MS may involve, for example, a reduction in the frequency of relapses, the intensity and / or severity of symptoms: fatigue, clumsiness, dizziness, difficulty with bladder regulation, loss of balance and coordination, difficulty with cognitive function (thinking, memory, concentration, learning and judgment), mood changes, muscle stiffness and muscle spasms (tremors). Treatment of MS may also refer to prolonging survival of a subject with MS.
[0361] Kit of the invention and applications
[0362] In another aspect, the invention relates to a kit, hereinafter kit of the invention, useful for the first and second methods of the invention.
[0363] In a particular embodiment, said kit of the invention is useful for the first method of the invention. In another particular embodiment, said kit of the invention is the second method of the invention.
[0364] In another aspect, the invention relates to the in vitro use of kit of the invention for:
[0365] • detecting and / or quantifying the isoform of the TAF1 as established in SEQ ID NO:
[0366] 1 that comprises a deletion from residue 1709 to residue 1893 of SEQ ID NO: 1 ,
[0367] • diagnosing whether a subject has MS,
[0368] • determining the risk of a subject developing MS,
[0369] • monitoring MS progression in a subject,
[0370] • evaluating the efficacy of a treatment against MS, or
[0371] • predicting survival of a subject who has MS.
[0372] For said applications, the kit of the invention will include the reagents necessary for detecting the isoform of SEQ ID NO: 1 that comprises at least one deletion from residues: 1709, 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893 of SEQ ID NO:1. Or in a more preferred embodiment, the kit of the invention will include detecting the isoform of SEQ ID NO:2 for use in the diagnosis and / or prognosis of MS.
[0373] The kit of the invention can further contain all those reagents necessary for detecting the amount of isoform of SEQ ID NO: 1 that comprises at least one deletion from residues: 1709, 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893 of SEQ ID NO:1 as defined previously, their mRNA, proteins or their variants, such as but not being limited to the following for example
[0374] • secondary antibodies labeled with a marker specifically recognizing the isoforms; substrates for the markers present in said labeled secondary antibodies; and positive and / or negative controls.
[0375] Likewise, said kit of the invention can further include, without any type of limitation, buffers, agents for preventing contamination, protein degradation inhibitors, etc. In addition, the kit of the invention can include all the supports and containers necessary for being put into practice and for optimization. Preferably, the kit further comprises instructions for use.
[0376] It is also noted that the term "kit" as used herein is not limited to any specific device and includes any device suitable for working the invent ion such as but not limited to microarrays, bioarrays, biochips or biochip arrays.
[0377] SEQ ID NO: 1
[0378] MGPGCDLLLRTAATITAAAIMSDTDSDEDSAGGGPFSLAGFLFGNINGAGQLEGESVL DDECKKHLAGLGALGLGSLITELTANEELTGTDGALVNDEGWVRSTEDAVDYSDINEV AEDESRRYQQTMGSLQPLCHSDYDEDDYDADCEDIDCKLMPPPPPPPGPMKKDKD QDSITGVSENGEGIILPSIIAPSSLASEKVDFSSSSDSESEMGPQEATQAESEDGKLTL PLAGIMQHDATKLLPSVTELFPEFRPGKVLRFLRLFGPGKNVPSVWRSARRKRKKKH RELIQEEQIQEVECSVESEVSQKSLWNYDYAPPPPPEQCLSDDEITMMAPVESKFSQ STGDIDKVTDTKPRVAEWRYGPARLWYDMLGVPEDGSGFDYGFKLRKTEHEPVIKSR MIEEFRKLEENNGTDLLADENFLMVTQLHWEDDIIWDGEDVKHKGTKPQRASLAGWL PSSMTRNAMAYNVQQGFAATLDDDKPWYSIFPIDNEDLVYGRWEDNIIWDAQAMPRL LEPPVLTLDPNDENLILEIPDEKEEATSNSPSKESKKESSLKKSRILLGKTGVIKEEPQQ NMSQPEVKDPWNLSNDEYYYPKQQGLRGTFGGNIIQHSIPAVELRQPFFPTHMGPIK LRQFHRPPLKKYSFGALSQPGPHSVQPLLKHIKKKAKMREQERQASGGGEMFFMRT PQDLTGKDGDLILAEYSEENGPLMMQVGMATKIKNYYKRKPGKDPGAPDCKYGETVY CHTSPFLGSLHPGQLLQAFENNLFRAPIYLHKMPETDFLIIRTRQGYYIRELVDIFVVGQ QCPLFEVPGPNSKRANTHIRDFLQVFIYRLFWKSKDRPRRIRMEDIKKAFPSHSESSIR KRLKLCADFKRTGMDSNWWVLKSDFRLPTEEEIRAMVSPEQCCAYYSMIAAEQRLKD AGYGEKSFFAPEEENEEDFQMKIDDEVRTAPWNTTRAFIAAMKGKCLLEVTGVADPT GCGEGFSYVKIPNKPTQQKDDKEPQPVKKTVTGTDADLRRLSLKNAKQLLRKFGVPE EEIKKLSRWEVIDVVRTMSTEQARSGEGPMSKFARGSRFSVAEHQERYKEECQRIFD LQNKVLSSTEVLSTDTDSSSAEDSDFEEMGKNIENMLQNKKTSSQLSREREEQERKE LQRMLLAAGSAASGNNHRDDDTASVTSLNSSATGRCLKIYRTFRDEEGKEYVRCETV
[0379] RKPAVIDAYVRIRTTKDEEFIRKFALFDEQHREEMRKERRRIQEQLRRLKRNQEKEKLK
[0380] GPPEKKPKKMKERPDLKLKCGACGAIGHMRTNKFCPLYYQTNAPPSNPVAMTEEQE
[0381] EELEKTVIHNDNEELIKVEGTKIVLGKQLIESADEVRRKSLVLKFPKQQLPPKKKRRVGT
[0382] TVHCDYLNRPHKSIHRRRTDPMVTLSSILESIINDMRDLPNTYPFHTPVNAKVVKDYYKI
[0383] ITRPMDLQTLRENVRKRLYPSREEFREHLELIVKNSATYNGPKHSLTQISQSMLDLCDE
[0384] KLKEKEDKLARLEKAINPLLDDDDQVAFSFILDNIVTQKMMAVPDSWPFHHPVNKKFV
[0385] PDYYKVIVNPMDLETIRKNISKHKYQSRESFLDDVNLILANSVKYNGPESQYTKTAQEIV
[0386] NVCYQTLTEYDEHLTQLEKDICTAKEAALEEAELESLDPMTPGPYTPQPPDLYDTNTS
[0387] LSMSRDASVFQDESNMSVLDIPSATPEKQVTQEGEDGDGDLADEEEGTVQQPQASV
[0388] LYEDLLMSEGEDDEEDAGSDEEGDNPFSAIQLSESGSDSDVGSGGIRPKQPRMLQE
[0389] NTRMDMENEESMMSYEGDGGEASHGLEDSNISYGSYEEPDPKSNTQDTSFSSIGGY
[0390] EVSEEEEDEEEEEQRSGPSVLSQVHLSEDEEDSEDFHSIAGDSDLDSDE
[0391] SEQ ID NO: 2
[0392] MRPGCDLLLRAAATVTAAIMSDSDSEEDSSGGGPFTLAGILFGNISGAGQLEGESVLD
[0393] DECKKHLAGLGALGLGSLITELTANEELTGTGGALVNDEGWIRSTEDAVDYSDINEVA
[0394] EDESQRHQQTMGSLQPLYHSDYDEDDYDADCEDIDCKLMPPPPPPPGPMKKDKDQ
[0395] DAITCVSESGEDIILPSIIAPSFLASEKVDFSSYSDSESEMGPQEATQAESEDGKLTLPL
[0396] AGIMQHDATKLLPSVTELFPEFRPGKVLRFLHLFGPGKNVPSVWRSARRKRKKHRELI
[0397] QEEQIQEVECSVESEVSQKSLWNYDYAPPPPPEQCLADDEITMMVPVESKFSQSTGD
[0398] VDKVTDTKPRVAEWRYGPARLWYDMLGVSEDGSGFDYGFKLRKTQHEPVIKSRMM
[0399] EEFRKLEESNGTDLLADENFLMVTQLHWEDSIIWDGEDIKHKGTKPQGASLAGWLPSI
[0400] KTRNVMAYNVQQGFAPTLDDDKPWYSIFPIDNEDLVYGRWEDNIIWDAQAMPRLLEP
[0401] PVLALDPNDENLILEIPDEKEEATSNSPSKESKKESSLKKSRILLGKTGVIREEPQQNMS
[0402] QPEVKDPWNLSNDEYYFPKQQGLRGTFGGNIIQHSIPAMELWQPFFPTHMGPIKIRQF
[0403] HRPPLKKYSFGALSQPGPHSVQPLLKHIKKKAKMREQERQASGGGELFFMRTPQDLT
[0404] GKDGDLILAEYSEENGPLMMQVGMATKIKNYYKRKPGKDPGAPDCKYGETVYCHTS
[0405] PFLGSLHPGQLLQALENNLFRAPVYLHKMPETDFLIIRTRQGYYIRELVDIFVVGQQCP
[0406] LFEVPGPNSRRANMHIRDFLQVFIYRLFWKSKDRPRRIRMEDIKKAFPSHSESSIRKRL
[0407] KLCADFKRTGMDSNWWVLKSDFRLPTEEEIRAKVSPEQCCAYYSMIAAKQRLKDAGY
[0408] GEKSFFAPEEENEEDFQMKIDDEVHAAPWNTTRAFIAAMKGKCLLEVTGVADPTGCG
[0409] EGFSYVKIPNKPTQQKDDKEPQAVKKTVTGTDADLRRLSLKNAKQLLRKFGVPEEEIK
[0410] KLSRWEVIDVVRTMSTEQAHSGEGPMSKFARGSRFSVAEHQERYKEECQRIFDLQN KVLSSTEVLSTDTDSISAEDSDFEEMGKNIENMLQNKKTSSQLSREWEEQERKELRR
[0411] MLLVAGSAASGNNHRDDVTASMTSLKSSATGHCLKIYRTFRDEEGKEYVRCETVRKP
[0412] AVIDAYVRIRTTKDEKFIQKFALFDEKHREEMRKERRRIQEQLRRLKRNQEKEKLKGPP
[0413] EKKPKKMKERPDLKLKCGACGAIGHMRTNKFCPLYYQTNVPPSKPVAMTEEQEEELE
[0414] KTVIHNDNEELIKVEGTKIVFGKQLIENVHEVRRKSLVLKFPKQQLPPKKKRRVGTTVH
[0415] CDYLNIPHKSIHRRRTDPMVTLSSILESIINDMRDLPNTHPFHTPVNAKVVKDYYKIITRP
[0416] MDLQTLRENVRKCLYPSREEFREHLELIVKNSATYNGPKHSLTQISQSMLDLCDEKLK
[0417] EKEDKLARLEKAINPLLDDDDQVAFSFILDNIVTQKMMAVPDSWPFHHPVNKKFVPDY
[0418] YKMIVNPVDLETIRKNISKHKYQSRESFLDDVNLILANSVKYNGPESQYTKTAQEIVNIC
[0419] YQTITEYDEHLTQLEKDICTAKEAALEEAELESLDPMTPGPYTSQPPDMYDTNTSLSTS
[0420] RDASVFQDESNLSVLDISTATPEKQMCQGQGRLGEEDSDVDVEGYDDEEEDGKPKP
[0421] PAPEGGDGDLADEEEGTVQQPEASVLYEDLLISEGEDDEEDAGSDEEGDNPFSAIQL
[0422] SESGSDSDVGYGGIRPKQPFMLQHASGEHKDGHGK
[0423] SEQ ID NO: 3
[0424] MGPGWAGLLQDKGGGSPSVVMSDTDSDEESAGGGPFSLTGFLFGNINGAGQLEGE
[0425] SVLDDECKKHLAGLGALGLGSLITELTANEELSGSDGALVNDEGWIRSREDAVDYSDI
[0426] NEVAEDESRRYQQTMGSLQPLCHTDYDEDDYDADCEDIDCKLMPPPPPPPGPLKKE
[0427] KDQDDITGVSEDGEGIILPSIIAPSSLASEKVDFSSSSDSESEMGPQDAAQSESKDGQL
[0428] TLPLAGIMQHDATKLLPSVTELFPEFRPGKVLRFLRLFGPGKNVPSVWRSARRKRKKK
[0429] HRELIQEGQVQEEECSVELEVNQKSLWNYDYAPPPLPDQCLSDDEITMMAPVESKFS
[0430] QSTGDTDKVMDTKPRVAEWRYGPARLWYDMLGVPEDGSGFDYGFKMKKTEHESTI
[0431] KCNIMKKLRKLEENSGVDLLADENFLMVTQLHWEDDIIWDGEDVKHKGTKPQRASLA
[0432] GWLPSSMTRNAMAYNVQQGFTATLDDDKPWYSIFPIDNEDLVYGRWEDNIIWDAQN
[0433] MPRILEPPVLTLDPNDENLILEIPDEKEEATSNSPSKENKKESSLKKSRILLGKTGVIKEE
[0434] PQQNMSQPEVKDPWNLSNDEYYYPKQQGLRGTFGGNIIQHSIPAVELRQPFFPTHM
[0435] GPIKLRQFHRPPLKKYSFGALSQPGPHSVQPLLKHIKKKAKMREQERQASGGGEMFF
[0436] MRTPQDLTGKDGDLILAEYSEENGPLMMQVGMATKIKNYYKRKPGKDPGAPDCKYG
[0437] ETVYCHTSPFLGSLHPGQLLQAFENNLFRAPIYLHKMPESDFLIIRTRQGYFIRELVDIF
[0438] VVGQQCPLFEVPGPNSKRANTHIRDFLQVFIYRLFWKSKDRPRRIRMEDIKKAFPSHS
[0439] ESSIRKRLKLCADFKRTGMDSNWWVLKSDFRLPTEEEIRAMVSPEQCCAYYSMIAAE
[0440] QRLKDAGYGEKSFFAPEEENEEDFQMKIDDEVRTAPWNTTRAFIAAMKGKCLLEVTG
[0441] VADPTGCGEGFSYVKIPNKPTQQKDDKEPQPVKKTVTGTDADLRRLSLKNAKQLLRK
[0442] FGVPEEEIKKLSRWEVIDVVRTMSTEQARSGEGPMSKFARGSRFSVAEHQERYKEEC QRIFDLQNKVLSSTEVLSTDTDSSSAEDSDFEEMGKNIENMLQNKKTSSQLSREREE QERKELQRMLLAAGSAAAGNNHRDDDTASVTSLNSSATGRCLKIYRTFRDEEGKEYV RCETVRKATVIDAYVRIRTTKDEEFIRKFALFDEQHREEMRKERRRIQEQLRRLKRNQ EKEKLKGPPEKKPKKMKERPDLKLKCGACGAIGHMRTNKFCPLYYQTNAPPSNPVAM TEEQEEELEKTVIHNDNEELIKVEGTKIVLGKQLIESADEVRRKSLVLKFPKQQLPPKKK RRVGTTVHCDYLNRPHKSIHRRRTDPMVTLSSILESIINDMRDLPNTYPFHTPVNAKVV KDYYKIITRPMDLQTLRENVRKRLYPSREEFREHLELIVKNSATYNGPKHSLTQISQSM LDLCDEKLKEKEDKLARLEKAINPLLDDDDQVAFSFILDNIVTQKMMAVPDSWPFHHP VNKKFVPDYYKVIVSPMDLETIRKNISKHKYQSRESFLDDVNLILANSVKYNGPESQYT KTAQEIVNVCHQTLTEYDEHLTQLEKDICTAKEAALEEAELESLDPMTPGPYTPQAKPP DLYDNNTSLSVSRDASVYQDESNLSVLDIPSATSEKQLTQEGGDGDGDLADEEEGTV QQPQASVLYEDLLMSEGEDDEEDAGSDEEGDNPFFAIQLSESGSDSDVESGSLRPK QPRVLQENTRMGMENEESMMSYEGDGGDASRGLEDSNIR
[0443] DESCRIPTION OF THE DRAWINGS
[0444] FIGURE 1 Schematic view of human TAF1 gene and protein (Uniprot P21675), with the following domains depicted in the protein: TAF1 N-terminal domain (TAND), DNA binding domain (DUF3591 ), zinc knuckle (ZnK) and the bromodomains (BrD1 and BrD2). Also indicated the positions of the protein sequences used as immunogens to generate the antibodies employed in the Western blot analyses, each mapping to a different exon (3, 37 and 38).
[0445] FIGURE 2 TAF1 detection with antibodies against sequences encoded by exons 3, 37 or 38 in nuclear cortical extracts (NAG+WM) from controls (n = 10) and individuals with PPMS (n = 10) or SPMS (n = 8), quantified by normalizing respect to SMC1 levels (one-way ANOVA followed by Tukey’s post hoc test). Graphs show mean ± SEM. n.s., non-significant.
[0446] FIGURE 3 Disorder prediction of TAF1 determined with PrDOS webservice. Maximum false positive rate was stablished at 5% for the prediction, and a probability of 0.5 was taken as the threshold of disorder (red line). All amino acids in the C-terminal region (1740-1893 residues) have a disorder score higher than 0.6 (blue dashed line).
[0447] FIGURE 4 Conservation according to JalView of the last 476 amino acids of TAF1 protein (1418-1893 residues) across 15 species (the ones listed in the sequence alignment plus the reptile Notechis scutatus and the bony fishes Takifugu rubripes, Xiphophorus maculatus and Amphiprion ocellaris). The multiple sequence alignment corresponds to the 1794-1893 sequence encoded by part of exon 37 and the entire 38, residues are coloured according to Clustal X colour scheme. The protein consensus sequence is shown below.
[0448] FIGURE 5 Multiple sequence alignment also of the 1794-1893 sequence, but across 30 species (including Saccharomyces cerevisiae, Drosophila melanogaster and Caenorhabditis elegans) evidences that conservation extends beyond vertebrates. Note that interruptions in the alignment are caused by insertions in some sequences, which shift the rest in the alignment. The protein consensus sequence is shown below.
[0449] FIGURE 6 Immunohistochemistry in control and PPMS cortical sections with the TAF1 ex37+38 antibody (raised against the sequence encoded by exons 37 and 38); top: control white matter (WM) and PPMS NAWM, bottom: control gray matter (GM) and PPMS NAGM (controls, n = 3; PPMS, n = 6; SPMS, n = 6). Histograms show mean optical density quantification (mean ± SEM, one-way ANOVA followed by Tukey’s post hoc test) and graph shows simple linear regression (Spearman’s rank correlation coefficient) with significant correlation in GM between decreased C-terminal TAF1 staining and lower disease duration of both PPMS and SPMS cases. Scale bars: 50 pm (top); 100 pm (bottom).
[0450] FIGURE 7 Cortical sections (from two different individuals with SPMS) containing a WM inactive lesion (top) or a WM active-inactive lesion (bottom) as well as NAGM and NAWM. Left panels: lesions are delimited based on loss of eriochrome cyanine staining of myelin (blue) and increased HLA-DR immunoreactivity. Right panels: the same lesions immunostained, in sections proximal to those in left panels, with a TAF1 monoclonal antibody raised against the whole protein. Scale bar: 500 pm (top); 2 mm (bottom).
[0451] FIGURE 8 Transcript structure (3’ from exon 36) of canonical TAF1 RNA and of reported isoforms lacking exon 38 sequences. The Canonical form includes the whole exon 38 and the canonical 3’ UTR. The Short exon 38 isoform lacks the sequence encoding the last 52 amino acids encoded by exon 38 and is predicted to add 21 C-terminal amino acids encoded by part of the canonical 3’ UTR. The Intron 37 retention isoform includes intron 37 and is predicted to result in a truncated protein lacking the whole exon 38, due to the highly conserved stop codon at the beginning of the intron (see FIGURE 9). The Distal exon isoform splices a large region comprising exon 38 and the 3’ UTR and incorporate distal exons that are predicted to add only three additional amino acids (RYQ). The Distal exon isoform has only been reported in humans, but the rest are common to both human and mouse (h / m).
[0452] FIGURE 9 Multiple sequence alignment of the genomic sequence comprising the last 20 nucleotides of exon 37 and the beginning of intron 37 across 25 species. A dashed yellow line marks the border between exonic and intronic sequence. The amino acid sequence for Homo sapiens is shown below. Note the high conservation of the stop codon (red) if intron 37 is retained.
[0453] FIGURE 10 Schematic view of the 3’ region of TAF1 gene comprising exon 37 to 3’UTR. The scissors show the targets of the guide RNA in a region inside intron 37 with low conservation and inside the alternative 3’UTR, leading to a deletion of 4,246 bp.
[0454] FIGURE 11 Integrative Genomics Viewer (IGV) profiles of RNA-seq data from a wild-type (wt) and a hemizygous Taf1d38 mouse at the 3’region of TAF1. In the wt mouse, exons 37 and 38 and the canonical 3’UTR (blue) are present, minimal intron 37 retention (gray) is seen and the alternative 3’UTR (light blue) has few reads. In the Taf1d38 mouse, exon 37 (blue) is followed by the initial part of intron 37 (gray) fused to the distal part of the alternative 3’UTR (light blue) which provides a polyA sequence.
[0455] FIGURE 12 Detection of TAF1 protein by Western blot in brain of 1 ,5-month-old wt or Taf1d38 mice with the antibodies against sequences encoded by exons 38 or 3. Two-tailed unpaired t-test. Graphs show mean ± SEM.
[0456] FIGURE 13 Time course of the evident motor phenotype of Taf1 d38 mice from early phases (3 months) to end-stage disease (20 months).
[0457] FIGURE 14 Ambulatory distance traveled and number of vertical activity counts in open field test by 2-month-old wt (n = 10) or Taf1d38 mice (n = 16). Two-tailed unpaired t-test. Graphs show mean ± SEM. FIGURE 15 Limb clasping score evolution from early to symptomatic stages of wt (n = 10 for 2-8months and n= 9 after 8 months) or Taf1d38 mice (n = 16).
[0458] Figure 16 Mean latency to fall in the rotarod test for 10-month-old wt (n = 9) or Taf1d38 mice (n = 16). Two-tailed unpaired t-test. Graphs show mean ± SEM.
[0459] FIGURE 17 Photograph of the entire spinal cord of two 12-month-old mice, wt on the left and Taf1d38 on the right, showing severe atrophy and white matter loss in the latter.
[0460] FIGURE 18 FluroMyelin staining of lumbar spinal cord of 12-month-old wt or Taf1 d38 mice showing atrophy and profound white matter loss in both lateral and anterior columns. Insets (boxed areas) in lateral columns are shown enlarged on the right. Histogram shows quantification of the mean signal in the lateral columns of the lumbar spinal cord (n=2). Scale bars: 200 pm (left panels), 10 pm (right panels). Two-tailed unpaired t-test. Graphs show mean ± SD.
[0461] FIGURE 19 Transmission electron microscopy of spinal cord white matter of 7-month-old mice showing marked myelin loss in Taf1 d38 mice with respect to wt. G-ratio quantification is shown for the three main white matter tracts of the spinal cord at thoracolumbar level (left histogram). Axonal diameter vs G-ratio plot (right down) shows that increased G-ratio in Taf1d38 mice occurs independently of axonal size. Twenty microphotographs per tract and per animal (n = 3) were analyzed. Scale bar: 1 pm. Two-tailed unpaired t-test. Graphs show mean ± SD.
[0462] FIGURE 20 Black Gold myelin staining (top) allowing discern myelin loss, and immunohistochemistry staining for CD68 (bottom left) and GFAP (bottom right) showing glial activation, in sagittal sections from 6-month-old Taf1d38 mice with as compared to wt ones. Scale bars: 200 pm.
[0463] FIGURE 21 Examples of FluroMyelin staining of coronal brain slices of 12-month-old wt and Taf1d38 mice (left). The boxes in corpus callosum and subcortical white matter (cingulum bundle and medial external capsule) show the areas whose signal was quantified, as shown in the histograms (right). Quantifications of the mean intensity values were performed on 3 and 12-month-old wt and Taf1 d38 mice (n = 2). Scale bar: 500 pm.
[0464] FIGURE 22 Electrophysiological recording of propagated compound action potentials (CAPs) in the fornix of 14-month-old wt (n = 4) or Taf1d38 mice (n = 4), showing increased latency in the conduction through myelinated fibers. Latencies are defined as the time between the onset of the stimulus artifact and each of the two main negativities, N1 and N2, corresponding to conduction through myelinated and unmyelinated fibers respectively (see inset). Two-tailed unpaired t-test. Graphs show mean ± SD.
[0465] FIGURE 23 Immunohistochemistry for GFAP (left) and Iba1 (right) in brainstem sagittal sections from 6-month-old wt or Taf1 d38 mice, showing marked glial activation. Scale bars: 500 pm.
[0466] FIGURE 24 Spp1 (osteopontin) immunofluorescence in corpus callosum (sagittal sections) from 6-month-old wt and Taf1d38 mice, exemplifying the upregulation in Taf1d38 white matter. Scale bar: 50 pm.
[0467] FIGURE 25 CD3 or BC220 immunofluorescence (red) together with DAPI nuclear counterstaining (blue) in corpus callosum (coronal sections) from 3-month-old wt and Taf1d38 mice, evidencing the absence of lymphocyte infiltration. Scale bar: 20 pm.
[0468] FIGURE 26 Percentage of differentially expressed genes (DEG) identified through bulk tissue polyA-i- RNA-seq analysis of striata from 2-month-old wt (n = 4) or Taf1d38 mice (n = 4).
[0469] FIGURE 27 Gene ontology analysis, using Ingenuity Pathway Analysis (IPA) tool, of data in panel a. The most significant enriched canonical pathways for the top (log2FC > |0.5|) DEG in Taf1d38 mice are shown.
[0470] FIGURE 28 Volcano plot of Taf1 d38 DEG in figure 25.
[0471] FIGURE 29 Bulk RNA-seq analyses of human MS NAG+WM (polyA+) or WM lesions (ribo-depleted) used in panel e. For our polyA-i- RNA-seq of cortical NAG+WM from controls or individuals with PPMS, the principal component analysis of expression data of all detected genes is shown.
[0472] FIGURE 30 Principal component analysis of all detected genes in the polyA+ RNA-seq experiment on human cortical samples, that was performed on 3 individuals with PPMS and 4 controls. The CT4 control sample was atypical as it mapped closer to PPMS samples than to the rest of control samples. Accordingly, we decided to exclude this sample from the final analysis.
[0473] FIGURE 31 Representation factors (RF) showing the DEG overlap between Taf1d38 mice and MS samples. Left: overlap between the polyA+ RNA-seq data of Taf1d38 mice and PPMS cortical NAG+WM. Right: overlap between ribo-depleted RNA-seq data of Taf1 d38 mice and active lesions of individuals with progressive MS. The top DEGs (up or down) were defined for Taf1 d38 mice as those with log2FC > |0.25|, and for MS samples as those with log2FC > |1 .25|.
[0474] FIGURE 32 Bubble plot of top enriched categories from the Human Phenotype Ontology (HPO) after gene-set enrichment analysis (GSEA) of bulk polyA-i- RNA-seq data. The enriched categories that were validated for MS with a gene-set level analysis via MAGMA (p<0.05) are coloured and labeled in purple. CNS, central nervous system; ESR, erythrocyte sedimentation rate; IFN, interferon; MCV, mean corpuscular volume.
[0475] FIGURE 33 In the mouse snRNA-seq experiment, the UMAP plot of cell clusters in resolution 1 yield a total of 27 clusters.
[0476] FIGURE 34 UMAP plot of cell clusters sorted by cell population after single nuclei RNA-seq of the striatum (n = 10,351 nuclei from 3 wt and 3 Taf1d38 mice). AS, astrocytes; END, endothelium; IN, interneurons; MG, microglia; MOL, mature oligodendrocytes; MSN, medium spiny neurons; NFOL, newly formed oligodendrocytes; OPC, oligodendrocyte precursor cells.
[0477] FIGURE 35 Proportion of nuclei corresponding to each genotype (wt or Taf1d38) for each of the eight major cell populations: AS, astrocytes; END, endothelium; IN, interneurons; MG, microglia; MOL, mature oligodendrocytes; MSN, medium spiny neurons; NFOL, newly formed oligodendrocytes; and OPC, oligodendrocyte precursor cells. FIGURE 36 Top: violin plots of significantly dowregulated genes in the MOL population (n = 1372 nuclei) identified by Wilcoxon rank sum test, most of which were also identified in bulk RNA-seq. Bottom: IGV track for Ninj2 gene in MOL (merged tracks of 3 replicates are shown).
[0478] FIGURE 37 UMAP plot highlighting in red the most perturbed clusters (6 and 16 in FIG 32), which correspond to MOL and MG, respectively.
[0479] FIGURE 38 Proteomics and differential analysis experimental design. Oli-neu or N2a cells were transfected (in triplicates) with either EGFP, EGFP-TAF1 full length (FL) or EGFP-TAF1 d38 plasmids and processed for GFP-Trap. The resulting co-immunoprecipitates were analysed by liquid chromatography-mass spectrometry (LC-MS / MS). EGFP and TAF1 structures were taken from AlphaFold and drawings were imported from Biorender.
[0480] FIGURE 39 Top significant canonical pathways (from IPA gene ontology tool) for TAF1 FL>TAF1 d38 interactors found in Oli-neu cells (n =74) and N2a cells (n =18).
[0481] FIGURE 40 TAF1 FL>TAF1d38 interactors with transcriptional functions found in Oli-neu and N2a cells. Shown in purple the proteins encoded by genes associated to MS, according to gene-based analysis by MAGMA of the latest IMSGC GW AS data.
[0482] FIGURE 41 Total RNAPII signal (RPGC, spike-in normalized) after bulk ChlP-seq performed on striata of 2-month-old wt or Taf1d38 mice. Focus on the promoter region is shown below. The box plot shows the distribution of Iog2(pausing index) based on two biological replicates (right). Bottom: normalized BigWig track of total RNAPII signal along Npas2 gene showing increased promoter occupancy in Taf1d38. Mann-Whitney U test. Box plots show interquartile range; center line represents the median value.
[0483] FIGURE 42 Top significant IPA canonical pathways for genes whose promoter occupancy by total RNAPII is higher in Taf1d38 mice than in controls.
[0484] FIGURE 43 Ser2P RNAPII signal (RPGC, spike-in normalized) at genes larger than
[0485] 30 kb after bulk ChlP-seq. Focus on the gene body is shown below. The box plot shows the distribution of normalized RPGCs on gene bodies (right). Bottom: normalized BigWig track of Ser2P RNAPII signal along the Anin gene showing decreased gene body occupancy in Taf1d38. Mann-Whitney U test. Box plots show interquartile range; center line represents the median value.
[0486] FIGURE 44 Top significant IPA canonical pathways for genes whose gene body occupancy by Ser2P RNAPII is lower in Taf1d38 mice than in controls.
[0487] FIGURE 45 Western blot analysis of SREK1 protein levels in cortex (control and NAG+WM) from controls (n = 10) and individuals with PPMS (n = 10) or SPMS (n = 8), and their quantifications normalized to p-actin (one-way ANOVA followed by Tukey’s post hoc test). Graphs show mean ± SEM. n.s., non-significant.
[0488] FIGURE 46 Experimentally determined (PhosphoSite) and predicted (NetPhos and ELM) phosphorylation sites at human exon 38-encoded protein sequence. CK1 , casein kinase 1 ; CK2, casein kinase 2; CK1 / 2, both casein kinases; Uns, unspecific kinase.
[0489] FIGURE 47 Cathepsin B (CTSB) cleavage sites predicted with ProsperousPlus tool (http: / / prosperousplus.unimelb-biotools.cloud.edu.au / ) at human TAF1 protein sequence encoded by exon 38 are shown with their respective probability scores in parentheses (range 0-1 , threshold 0.8).
[0490] FIGURE 48 JalView alignment of C-terminal TAF1 and TAF1 L showing the loose of the amino acids 1788 to 1893 of TAF1 , with the addition of 12 amino acids (SEQ ID NO: 78 HASGEHKDGHGK) due to the frameshift in TAF1 L.
[0491] Examples
[0492] Example 1. Decreased detection of the C-terminal end of TAF1 in brains from individuals with MS
[0493] To explore whether TAF1 is altered in CNS tissue where the neurodegenerative process is still not evident, we investigated its protein levels on postmortem brain cortical samples containing normal-appearing grey plus white matter (NAG+WM) from individuals with PPMS or SPMS, with antibodies against different parts of the TAF1 protein. The human TAF1 gene is located on chromosome X and encompasses 38 exons that encode the 1893-amino acid canonical form of the protein (Uniprot P21675-2) (Fig.1 ). We used three commercial antibodies which were raised against the following amino acid sequences: 103 to 123 (encoded by exon 3), 1771 to 1821 (encoded by exon 37), or 1821 to 1871 (encoded by exon 38) (Fig. 1 ). Strikingly, we found a specifically reduced immunodetection at the C-terminal end of TAF1 with the antibody raised against amino acids encoded by the last exon (exon 38), in both PPMS and SPMS samples (Fig. 2).
[0494] We then performed a bioinformatics analysis of the C-terminal sequence of TAF1. In agreement with a previous prediction using AlphaFold (Bernardini A, et al. Nat Struct Mol Biol 2023; 30(8): 1141 -1152), the disorder prediction tool PrDOS depicts the secondary structure of the entire sequence after the tandem bromo domains (BrDs) as disordered (Fig. 3). We observed a high degree of conservation at the BrDs (encoded by exons 28-32), followed by a valley of low homology. Notably, conservation reappears at the particularly unstructured (disorder probability > 0.6) region downstream of mid-exon 36, suggesting the existence of functional domains at the C-terminal end of TAF1. Finally, the sequence encoded by exon 38 is particularly rich in conserved glutamic and aspartic residues (Fig. 4 and Fig. 5). Overall, our results show decreased detection of the C-terminal end of TAF1 (which is highly conserved, intrinsically disordered and acidic), with no changes in TAF1 total levels, in postmortem brain tissue from individuals with progressive MS.
[0495] We then performed immunohistochemistry on sections from snap-frozen brain tissue containing normal-appearing grey matter (NAGM), normal-appearing white matter (NAWM) and demyelinating lesions (Fig. 6 and Fig. 7). We used the two Human Protein Atlas-validated TAF1 antibodies, raised against the C-terminal region (ex37+ex38 Ab) or the whole protein (whole TAF1 Ab). In control individuals, the ex37+ex38 Ab recognized the nuclei and (to a lesser extent) the cytoplasm of both glial cells and neurons. The signal intensity of ex37+ex38 Ab showed a tendency to decrease in both NAGM and NAWM of PPMS samples. Notably, lower labelling intensity in NAGM significantly correlated with shorter disease duration (Fig. 6). In contrast, the whole TAF1 Ab selectively stained the entire surface of the core of lesions (Fig. 7). These results further support diminished detection of the C-terminal end of TAF1 in NAGM and NAWM of individuals with MS and suggest an even more pronounced TAF1 alteration associated to demyelinated lesions.
[0496] Example 2. Progressive disability in mice lacking the C-terminal end of TAF1
[0497] To explore the function of this conserved sequence at the C-terminal end of TAF1 and its possible pathological relevance in vivo, we generated a genetically modified mouse model that mimics the decreased detection of this region in MS brains. For this, we emulated one of the human alternative splicing events that generates TAF1 isoforms lacking all or most of exon 38-encoded sequences (Fig. 8). More precisely, we mimicked the retention of intron 37, which results in truncation of the TAF1 protein after the exon 37-encoded section (Fig. 9), with the only change of the last amino acid (S1820R). We designed a CRISPR-Cas9 strategy to eliminate the reported splicing acceptor sites in exon38 and in the canonical 3’UTR. For this, we generated a guide RNA mapping to intron 37 and another mapping downstream of the canonical 3’UTR (the latter within an adjacent and longer cryptic alternative 3’UTR) (Fig. 10). Simultaneous injection of both guide RNAs into single-cell mouse embryos led to a mouse line with a Taf1d38 allele in which the 4,246 base-pair sequence flanked by both RNA guides sites was deleted, with no point mutations (Fig. 10). RNA-seq profiles of hemizygous males or homozygous females for the Taf1d38 allele showed that the remaining intron 37 sequence and the distal UTR formed a new 3’UTR immediately downstream of the canonical exon 37 (Fig. 11 ). By Western blot, we verified that there was no signal with the exon38 antibody but that the total TAF1 levels remain unaltered (using the N-terminal [exon3] antibody) (Fig. 12).
[0498] Mice carrying at least one Taf1d38 allele were born at the expected frequencies, and they looked normal until the age of 3 months, when hemizygous males and homozygous females began to show a visible progressive disability phenotype mainly affecting the hindlimbs, mimicking some of the symptomatology associated with mouse MS models such as the experimental autoimmune encephalomyelitis (EAE) model. At 3 months, Taf1d38 hemizygous and homozygous mice start showing episodes of tail tremor and of lowered pelvis during walking. At 6 months, the action tremor of the tail became permanent and the episodes of lowered pelvis more frequent. Kyphosis was evident in some mice at 6 months and increased in the percentage of mice in the following months. At 10 months, the tail tremor was accompanied by lower body shivering that extended to the whole trunk at 12 months. The phenotype then progressed to paraplegia, which was frequent in 20-month-old mice (Fig. 13). For this reason, we established 18 months as the humane endpoint for the entire colony of mice with the Taf1del38 allele (including heterozygote females whose visible phenotype is subtle).
[0499] We decided to continue only with mice with full suppression of exon 38 (i.e., hemizygous males and homozygous females), hereby termed Taf1d38 mice. In line with their evident phenotype of progressive disability, Taf1d38 mice also showed decreased activity (both ambulatory and vertical) in the open field test (Fig. 14), progressive hindlimb clasping (Fig. 15) and motor coordination deficit in the rotarod test (Fig. 16). Together, these results indicate that deletion of the C-terminal portion of TAF1 , which is underdetected in human MS brains, results in a progressive disabling phenotype.
[0500] Example 3. Massive demyelination in Taf1d38 mice
[0501] \Ne observed a striking loss of the whitish appearance of spinal cords of 12-month-old Taf1d38 mice (Fig. 17), indicating massive demyelination, which was further corroborated by FluoroMyelin staining (Fig. 18). Indeed, demyelination was already observed in the spinal cord of 7-month-old Taf1d38 mice, as evidenced by transmission electron microscopy (with significant increases of the G-ratios in anterolateral, lateral and dorsal columns; Fig. 19). In the brain, we observed an apparent decreased staining with BlackGold reagent in corpus callosum and superior colliculus at 6 months (Fig. 20), and a significant decrease of FluoroMyelin staining at 12 months (Fig. 21 ).
[0502] We next explored whether demyelination in Taf1 d38 mice was mirrored by a decrease in conduction velocity in white matter tracts, as reported for the EAE mouse model. Electrophysiological recordings of the fornix of Taf1d38 mice showed a significant increase in the latency of the first peak, which corresponds to conduction through myelinated fibers, and an unaltered latency of the second peak (unmyelinated fibers) (Fig. 22). This demonstrated a reduced conduction velocity in the fornix that specifically affected the myelinated axons.
[0503] Example 4. White matter neuroinflammation in Taf1d38 mice is driven by resident glia
[0504] Immunohistochemistry of sagittal brain sections of Taf1d38 mice with the microglia / macrophage lineage markers CD68 and Iba1 revealed a profuse staining in most white matter tracts, such as those in the brainstem, the superior colliculus and the corpus callosum (Fig. 20 and Fig. 23). Immunostainings for two additional markers of MS lesions — namely, GFAP and Spp1 (OPTN, osteopontin) — yielded patterns similar to those of CD68 and Iba1 (Fig. 20, Fig. 23 and Fig. 24). The microglial activation in Taf1d38 mice did not appear to be accompanied by infiltration of T or B lymphocytes, as these were not detected by immunofluorescence using the CD3 and B220 markers, not even in the brain regions of Taf1d38 mice with profuse CD68 and Iba1 staining (Fig. 25). Together, these results demonstrate CNS resident neuroinflammation, demyelination and decreased conduction velocity affecting the brain and the spinal cord of Taf1d38 mice.
[0505] Example 5. Taf1d38 causes transcriptomic alteration of neuroinflammation and demyelination pathways
[0506] To investigate the molecular mechanism underlying the MS-like phenotype of Taf1d38 mice, we first examined whether eliminating the conserved C-terminal domain in the scaffolding subunit of the general transcription factor TFIID would affect transcription. We performed bulk tissue RNA-seq analysis of brain samples from 8-week-old (before onset of apparent disability) Taf1d38 mice. We analysed the striatum, a grey and white matter structure that shows atrophy and demyelination in MS (Vercellino M, et al. J Neuropathol Exp Neurol 2009; 68(5): 489-502). Strikingly, only 1.8% of the 15,627 detected genes showed altered transcript levels in Taf1d38 mice with respect to wild-type littermates (Fig. 26 and Table 1 ), with 163 genes downregulated and 131 upregulated (adjusted p-value <0.05).
[0507] Table 1. Relevant differential expressed genes in Taf1d38 mice. The gene name, the logarithmic fold change and the adjusted p-value of the top and relevant differentially expressed genes in Taf1d38 vs wild-type mice.
[0508] Gene ontology analysis (with Ingenuity Pathway Analysis, IPA) on the differentially expressed genes (DEG) with log2FC > |0.5| revealed that the canonical pathways with the most significant enrichments [-loglO(p-value) > 3] were related to inflammation, pathogenesis of MS and vitamin D receptor / retinoid X receptor activation (Fig. 27), thus coinciding with key pathways that have been implicated in MS neuropathology and as environmental risk factors. These results suggested that the observed progressive disability phenotype of Taf1d38 mice is due to a very selective alteration in the expression of a restrictive set of genes within MS related pathways.
[0509] Upregulated genes included some of the markers of the neuroinflammation associated to demyelination in MS plaques, such as Spp1 (osteopontin, OPTN) and GFAP, as well as many other immunoinflammation-related genes, including H2-D1, H2-K1, Ccl4, CxcllO, Cd52, B2m and C4b (Fig. 4c). In contrast, oligodendroglial-enriched genes2, such as Ptgds, Anin, 1133 and Klk6, were among the downregulated transcripts (Fig. 28).
[0510] We next specifically analysed the similarities between the DEG transcriptomic signature of Taf1d38 mice and the signatures of NAG+WM or active lesions from individuals with MS. For NAG+WM, we performed bulk tissue RNA-seq analysis on samples from controls and individuals with PPMS (Fig. 29, Fig. 30); for active lesions, we analysed results from a previous report (Elkjaer ML, et al. Acta Neuropathol Common 2019; 7(1 ): 205.). Notably, the Taf1d38-DEG signature showed overrepresentation of both the NAW+GM and the active lesion-DEG signatures, and this was particularly remarkable for the top downregulated genes, with a 7-fold enrichment (P<1.8 x 10'5) for the NAW+GM signature, and a 9.6-fold enrichment (P< 1.7 x 10-4) for active lesions one (Fig. 31 ).
[0511] Using the GWAS summary statistics of the largest meta-analysis thus far (of 47,429 individuals with MS and 68,374 controls) from the International Multiple Sclerosis Genetics Consortium (IMSGC) to perform a gene-based association analysis via MAGMA (de Leeuw CA, et al. PLoS Comput Biol 2015; 11 (4): e1004219.), we found that 21.7% of the Taf1d38 DEGs were nominally associated with MS. We then performed a gene set enrichment analysis (GSEA) of the Taf1d38 DEG dataset against different databases. In the context of the Human Phenotype Ontology (HPO) database we detected interesting enriched categories and, after validation in silico based on MS genetic associations via gene-set MAGMA analysis, we found abnormal CNS myelination, abnormality of complement and abnormal serum IFN among the top significant ones (Fig. 32). Together, these results indicate that C-terminal TAF1 modification selectively affects the expression of a relatively small set of genes, including many potential MS susceptibility genes, resulting in an MS-like transcriptomic signature.
[0512] Example 6. Specific transcriptional disturbances of mature oligodendrocytes in Taf1d38 mice
[0513] To investigate the cell types underlying the transcriptional changes unveiled by bulk tissue RNA-seq, we performed single-nuclei RNA-seq (snRNA-seq) on striatal samples from 8-week-old wild-type or Taf1 d38 mice (n=3 in two batches). After quality control filtering, we identified 27 clusters (Fig. 33) that were annotated on the basis of the expression of lineage marker genes and particularly those from a mouse striatal snRNA-seq study (Munoz-Manchado AB, et al. Cell Rep 2018; 24(8): 2179-2190 e2177.). As expected, we could identify and analyze the proportion of eight major cell populations: medium spiny neurons, interneurons, microglia, astrocytes, endothelial cells, oligodendrocyte precursor cells (OPCs), newly formed oligodendrocytes (NFOLs) and mature oligodendrocytes (MOLs) (Fig. 34 and Fig. 35).
[0514] We then performed differential gene expression analyses and we could not detect at the single-cell level the increase of inflammation-related genes observed in the bulk RNA-seq. Thus, these inflammatory responses might be restricted to specific areas within the striatum, as the anterior commissure white matter, which might not be captured in the single cell analysis. On the other hand, we found that many downregulated genes identified in the bulk RNA-seq analysis, such as Ninj2, 1133, Anin, Padi2 and Ptgds, were indeed found downregulated in snRNA-seq specifically in mature oligodendrocytes, together with other oligodendrocyte markers, such as Sox10 (Fig. 36). Finally, to interrogate which cell clusters showed the highest transcriptional disturbance (and as we found no significant perturbations in any of the eight major cell populations shown in Fig. 34), we performed perturbation-based analyses on the 27 initially identified subclusters. We found a small (240 cells) microglial subcluster (16, AUC 0.8055480) and a large (502 cells) mature oligodendrocytes subcluster (6, AUC 0.6917725) to be the most perturbed in Taf1d38 mice (Fig 37). Together, these analyses detected oligodendrocyte populations among those most affected transcriptionally by the lack of exon 38 of TAF1 , while also pointing to microglia affectation.
[0515] Example 6. TAF1 C-terminal end interacts with MS-associated transcriptional release factors
[0516] \Ne next explored the mechanism by which the lack of the C-terminal end of TAF1 results in a restricted MS-like transcriptional alteration with associated neuroinflammation and demyelination and leads to an MS-like progressive disability. For this, we used proteomics to identify the physiological interactors of TAF1 affected by the presence or absence of its exon 38-encoded C-terminal end. We first performed GFP-Trap immunoprecipitation from cells transfected with full-length TAF1 (TAF1 FL) or TAF1d38, each with EGFP fused to the N-terminal, or only EGFP (three replicates each), followed by liquid chromatography-mass spectrometry and differential analysis (Fig. 38). We used two different cell lines, one oligodendroglial (Oli-neu) and the other neuronal (N2a neuroblastoma), and determined the interactors of TAF1 FL or TAF1d38, and their relative efficiencies, for each cell type. Consistent with the restricted transcriptomic alteration in Taf1d38 mice, we detected that TAF7, which dimerizes with TAF1 in TFIID to drive PIC assembly at core promoter sequences (Bernardini A, et al. Nat Struct Mol Biol 2023; 30(8): 1141 -1152), interacted with TAF1 FL as well as with TAF1d38, with the same efficiency, in both Oli-neu and N2a cells. This validated the efficacy of our approach at capturing physiological interactors of TAF1 .
[0517] For interactors that bound more efficiently to TAF1 FL than to TAF1d38 (TAF1 FL>TAF1d38 interactors), gene ontology analysis detected assembly of RNAPII complex among top significant terms, as well as nuclear receptor (androgen- and glucocorticoid-) signaling (Fig. 39). Among these differential interactors, we noticed multiple components of the general transcription factors TFIIH and TFIIB, which are also part of the PIC and are involved in RNAPII promoter escape. Specifically, the detected subunits are Gtf2h1 (p62), Gtf2h2 (p44), Gtf2h4 (p52) and Ercc2 (XPD, p80) as components of TFIIH, and Gtf2b, which is the only component of TFIIB (Fig. 40). Other transcription initiation and elongation factors detected as TAF1 FL>TAF1d38 interactors included: i) Instl 2 (in Oli-neu cells), which is a subunit of the Integrator complex that, among other functions, regulates transcriptional initiation and pause release following activation; and ii) EloB (in N2a cells), which is a subunit of the trimeric complex Elongin (Sill), an elongation factor that is thought to stimulate transcription by suppressing transient pausing of RNAPII. Finally, other oligodendroglial differential interactors also involved in transcriptional regulation included Flii, Hmgn2, H2ac21 and Csnk2a2 (Fig. 40).
[0518] The identification of these TAF1 FL>TAF1d38 interactors indicates that the C-terminal end of TAF1 interplays with multiple transcriptional regulators involved in release of RNAPII from the promoter itself, from promoter proximal pausing and even from pausing events distant from the promoter. Strikingly, a similar mechanism regarding promoter proximal pausing had previously been associated to MS through a cell type-specific intralocus interaction analysis of MS GWAS data, which specifically involves loci pathogenic in the oligodendrocyte lineage. In addition, upon gene-based MAGMA analysis of the GWAS data from the IMSGC, we found that the oligodendroglial TAF1 FL>TAF1d38 interactors Gtf2h4 and Gtf2b showed genetic nominal association with MS risk and severity, respectively; the same applies to Flii (Flightless), an actin remodeling protein and transcriptional coactivator of hormone-activated nuclear receptors, for MS risk. In summary, these data provide a possible mechanism by which C-terminal alteration of TAF1 observed in brain tissue from individuals with MS synergizes with the MS-associated common variants to cause transcriptional, histological and functional alterations, leading to progressive disability.
[0519] Example 7. RNAPII-promoter pausing and decreased elongation at myelination and MS related genes
[0520] \Ne next tested whether the C-terminal of TAF1 regulates RNAPII promoter escape and proximal pausing release, selectively affecting myelination related genes. For this, we performed ChlP-seq analyses (two replicates) in brain tissue of 8-week-old Taf1d38 or wild-type mice and analysed the distribution of total and elongating (Ser2-P) RNAPII. Globally, we observed that Taf1d38, as compared to wild-type mice, had an increase of total RNAPII specifically at promoter regions (TSS + / - 1 .5 kb) (Fig. 41 ), but an unaltered occupancy profile along gene bodies and transcription termination sites (TTS). Together, this resulted in increased RNAPII pausing index in Taf1d38 mice (Fig. 41 ). Gene ontology analysis of the genes displaying higher increased RNAPII levels at the promoter regions in Taf1d38 mice as compared to wild-type showed MS signaling to be one of the top affected pathways, together with lipid metabolism (adipogenesis and triacylglycerol biosynthesis), IFN signaling and nucleotide excision pathway (Fig. 42). In line with the increase in RNAPII promoter pausing, we observed decreased Ser2P RNAPII levels along gene bodies, indicating a general deficit in RNAPII elongation. Of note, this defect in productive RNAPII elongation was more pronounced in longer genes (P<10-4for genes >30 kb) (Fig. 43). Remarkably, gene ontology analysis of top genes with decreased Ser2P RNAPII gene body occupancy in Taf1d38 as compared to wild-type mice detected the myelination signaling pathway as the most significant term, followed by various protein kinase signaling pathways (Fig. 44). These results further support decreased RNAPII release as the most plausible molecular mechanism leading to MS upon C-terminal TAF1 modification. Example 8. Possible mechanisms by which immunodetection of C-terminal TAF1 decreases in MS brains
[0521] The C-terminal end (encoded by exon 38) of TAF1 is rich in acidic amino acids and is predicted to be intrinsically disordered, with none of the functional or structural domains previously reported for TAF1. However, it has a degree of evolutionary conservation of sequence similar to that observed for known functional domains of TAF1 , and its alternative splicing events are also conserved in multiple species to generate forms of TAF1 with or without exon 38-encoded sequences, suggesting a relevant and regulation-prone physiological role for the C-terminal end of TAF1. We found that the C-terminal end of TAF1 is under-detected in the brain of individuals with MS. Moreover, the Taf1d38 mouse model, which only expresses the splice variant of TAF1 lacking this C-terminal end, exhibits progressive motor disability, starting at the age of 3 months with tremor and walking abnormalities and leads to paraplegia at around 20 months. This strongly suggested that C-terminal TAF1 modification does contribute to MS etiology. Indeed, histological analyses of Taf1 d38 mice confirmed CNS-resident neuroinflammation in white matter tracts together with progressive demyelination. We gained mechanistic insight through transcriptomic and ChlP-seq analyses of Taf1d38 mice, together with proteomics analyses of full length-TAF1 and TAF1 d38 differential interactors, to reveal a regulatory role of the C-terminal TAF1 in RNAII-promoter escape, which especially affects elongation of oligodendroglial myelination-related genes. To the best of our knowledge, this study provides the first reported alteration of the transcription machinery in the brain of individuals with MS that, when placed into an animal model, suffices to induce a full, progressive MS-like phenotype. Further, by identifying TAF1 to be a key player in MS pathogenesis, this study has important etiological and translational implications for MS and provides a novel therapeutic target. Finally, our genetic mouse model of MS -with construct and face validity- can be used to preclinically test new therapies that have the potential to prevent or halt the progressive phase of the disease in individuals with PPMS or SPMS.
[0522] Inefficient transcriptional elongation particularly affecting oligodendrocytes, has already been proposed to contribute to the etiology of MS, based on genetic variants that affect expression of regulators of RNAPII promoter proximal pausing release in oligodendrocytes. More precisely, BRD3, which facilitates transcriptional elongation, and HEXIM1 / 2, which inhibits RNAPII release upon promoter proximal pausing. Besides, MS lesions were found to have decreased BRD3 RNA and increased HEXIM1 RNA and protein. Therefore, the seminal study by Corradin and coworkers (Factor DC, et al. Cell 2020; 181 (2): 382-395 e321 ) identified oligodendroglial RNAPII promoter-proximal pausing as a potential mechanism in the etiology of MS, involving two modifiers -BRD3 and HEXIM1 / 2-, affected by common genetic variants. Our study reveals that TAF1 C-terminal alteration, which is observed in MS brains and suffices to induce progressive demyelination and disability in mice, converges with genetic risk into the disruption of transcription elongation in CNS cell types as oligodendrocytes. Interestingly, the transcriptional regulators GTF2B and GTF2H4 that were detected in our study as C-terminal TAF1 interactors with genetic association to MS are also involved in RNAPII release, but at an earlier point in transcription, namely promoter escape. It therefore seems that oligodendroglial cells are particularly sensitive to transcriptional pausing, with this having a likely role in progression of the disease.
[0523] The determination of C-terminal TAF1 interactors in this study might also provide a connection with an environmental risk factor, namely EBV-infection. EBNA2 is the EBV protein that regulates early viral transcription and is one of the six EBV-nuclear antigens expressed during latent infection. Interestingly, the intrinsically disordered C-terminal acidic region of EBNA2 recruits the basal transcription factors TFIIH and TFIIB, specifically by binding the subunits Gtf2h1 , Ercc2 and Gtf2b, which we detected as TAF1 FL>TAF1d38 interactors. The fact that both TAF1 and EBNA2 have intrinsically disordered C-terminal acidic domains that share these functional interactors opens the striking possibility of an etiological connection between TAF1 and EBV viral infection, the latter having recently been reported to be causal for MS.
[0524] TAF1 alterations previously linked to the neurodegenerative diseases XDP and Huntington’s disease essentially consist of decreased protein levels, and the prominent affectation of striatal medium spiny neurons in both diseases suggests that these cells might be particularly vulnerable to decreased TAF1 levels. Diminished TAF1 in Huntington’s disease may be caused by a degron sequence in an alternative exon 5 due to decreased SREK1 levels. In XDP, the causing haplotype contains multiple variants in noncoding regions within and around TAF1 , and the retrotransposon in intron 32 of TAF1 is believed to be the etiologically relevant one, as its variable number of hexameric repeats correlates with decreased TAF1 transcript levels and earlier age at disease onset. However, the retrotransposon has also been reported to cause TAF1 splicing alterations, and specifically by increasing partial retention of intron 32, and diminished usage of exons downstream of intron 32, such as the 34’ microexon, although the latter has recently been ruled out through Nanopore long-read sequencing of TAF1 mRNAs from XDP and control brains tissues. Reduced expression levels of TAF1 exons downstream of the retrotransposon have been reported in XDP-derived cell systems, possibly due to diminished passage of RNAP II through the SVA hexameric repeats. It would be therefore intriguing to know whether C-terminal TAF1 antibodies can detect any imbalance in XDP brain tissue similar to the one reported here regarding MS, which particularly affects oligodendroglial cells. It is worth noting that XDP neuropathology, apart from the mentioned loss of striatal medium spiny neurons, also includes overall changes of the structural integrity within the white matter.
[0525] Our study also opens numerous questions regarding the mechanism by which immunodetection of C-terminal TAF1 decreases in MS brains. We can think of at least four possibilities: i) alternative splicing favoring transcript isoforms lacking exon 38, ii) post-translational modifications that preclude recognition by antibodies raised against the exon 38-encoded native sequence iii) endoproteolysis at C-terminal motifs and iv) abnormal expression of TAF1 L gene. Regarding splicing, we analysed our bulk tissue RNA-seq data of NAG+WM (n=3) from controls and individuals with PPMS with three independent splicing analysis programs, and none of them detected significant differences in TAF1 C-terminal exons. Unfortunately, the previous bulk RNA-seq study with higher numbers of analysed individuals and brain structures (Elkjaer ML, et al. Acta Neuropathol Commun 2019; 7(1 ): 205) does not have enough sequencing depth for efficient interrogation of alternative splicing differences. We however observed by Western blot that the levels of the splicing factor SREK1 are increased in NAG+WM brain tissue of SPMS cases (Fig. 45). Indeed, SREK1 levels have been linked to the degree of usage of TAF1 exon 38, with increased SREK1 levels correlating to decreased usage of exon 38 (Solis AS, et al. RNA Biol 2010; 7(4): 486-494). It is therefore possible that attenuation of SREK1 levels in MS could be a useful strategy to increase C-terminal TAF1 detection.
[0526] Post-translational modifications, such as phosphorylation, at the C-terminal end of TAF1 could also account for its decreased detection and altered function in MS. In this regard, we observed that the exon 38 encoded C-terminal end of TAF1 contains numerous conserved casein kinase I and II (CK1 and CK2) canonical phosphorylation sites (Fig. 46). This fits with the fact that the CK1 epsilon isoform (CSNK1 E) and the alpha' (or alpha 2) catalytic subunit of CK2 (CSNK2A2) were detected in our differential proteomics analysis as interactors of TAF1 FL but not of TAF1d38. Notably, another TAF1 FL selective interactor is the phosphatase Pgam5, which has been reported to establish a signaling loop with CK2 (Chen G, et al. Mol Cell 2014; 54(3): 362-377), and autoantibodies against Pgam5 have been detected in MS (Ayoglu B, et al. Mol Cell Proteomics 2013; 12(9): 2657-2672).
[0527] Regarding proteolysis, cathepsin B is a good candidate protease to be responsible for the decreased detection of C-terminal TAF1 in MS brain tissue, as i) cathepsin B activity is increased in MS brain tissue (Bever CT, Jr., et al. J Neurol Sci 1995; 131 (1 ): 71 -73.) and ii) using the ProsperousPlus tool, we found a site of cleavage by cathepsin B with a very high score (0.95) at position 1840 (SEQ ID NO 79: TSFS|SIGG) of TAF1 (Fig. 47 and Table 2). Therefore, inhibition of TAF1 proteolysis with cathepsin inhibitors, particularly those affecting cathepsin B, may become a new candidate therapeutic strategy for MS. Of note, cathepsin B inhibition has recently been shown to alleviate the Th1 , Th17 and Th22 transcription factor signaling dysregulation in EAE (Ansari MA, et al. Exp Neurol 2022; 351 : 113997).
[0528] Table 2. Prediction of cleavage sites for cathepsin family. The predicted cleavage site at C-terminal TAF1 (from position 1709 to 1893) by different cathepsin proteases is indicated, together with its prediction score (range 0-1 , threshold 0.85).
[0529] Finally, an abnormal TAF1 L expression in MS brains may also account for the imbalance in C-terminal TAF1 detection. The gene encoding this protein arose in primates from retrotransposition of a TAF1 transcript and, due to a microduplication, a frameshift in the 3’ extreme of the gene was introduced, leading to a protein product that lacks the C-terminal TAF1 residues 1788 to 1893 (Fig. 48). Thus, an aberrant expression of this gene, which is expected to be significantly expressed only in male germ cells to replace TAF1 function during spermatogenesis, would lead to an infradetection of C-terminal TAF1 in MS brains respect to N-terminal region, as the antibodies recognizing the N-terminal sequences of TAF1 would also recognize TAF1 L.
[0530] In summary, this invention identifies that TAF1 , a protein that has not previously been linked to MS, is altered at the C-terminus in postmortem brains from individuals with MS, both in the NAG+WM and in the demyelinating plaques, and that TAF1 genetically modified mice develop a full, progressive, MS-like neuroinflammatory, demyelinating and disabling phenotype due to deficient transcriptional elongation that particularly affects oligodendrocytes.
[0531] Materials & Methods
[0532] Human brain tissue samples
[0533] Brain specimens from NAGM and NAWM cortex of individuals with PPMS or SPMS and control individuals used in immunoblotting, immunohistochemistry and RNA sequencing, were provided by the UK Multiple Sclerosis Tissue Bank at Imperial College, London. Written informed consent for brain removal was obtained from brain donors and / or next of kin. Procedures, information and consent forms have been approved by the Bioethics Subcommittee of Consejo Superior de Investigaciones Cientificas (CSIC, Madrid, Spain).
[0534] Animals
[0535] CRISPR mice lacking exon 38 of Taf1 (Taf1 d38) were generated (for details, see ‘Generation of Taf1d38 mice’ below) for this study and used in a C57BL / 6J background. All mice were housed in the CBM animal facility under ad libitum food and water availability and maintained in a temperature-controlled environment on 12 h light-dark cycles. Animal housing and maintenance protocols followed local authority guidelines approved by the CBM Institutional Animal Care and Utilization Committee (Comite de Etica de Experimentacion Animal del CBM, CEEA-CBM), and Comunidad de Madrid PROEX 293 / 15 and PROEX 247.1 / 20. Mice were euthanized with CO2.
[0536] Generation and genotyping of Taf1d38 mice
[0537] Taf1d38 mice were generated in a 50% C57BL / 6J and 50% CBA / CA background via CRISPR / Cas9-mediated gene editing, by the institutional core facility (CNB-CBM mouse transgenesis facility). We decided to emulate one of the alternative splicing events that in humans generate TAF1 isoforms lacking exon 38 encoded sequences (Fig. 8). More precisely, we mimicked the retention of intron 37 that happens in approximately 20% of TAF1 transcripts in human brain (httDs: / / vastdb.crg.eu / event / HsalNT 1043632@hg38), which results in a TAF1 protein truncated at the end of exon 37 with the only substitution of its last amino acid (S1820R), due to the highly evolutionary conserved stop codon at the beginning of intron 37 (Fig. 9). We designed a CRISPR-Cas9 strategy to eliminate the reported splicing acceptor sites at exon 38 and the canonical 3’UTR. For this, we generated a guide RNA mapping to intron 37 (SEQ ID NO: 102: CCAAGTTGTGAAGGACGTAAATG) and another mapping downstream of the canonical 3’UTR (SEQ ID NO: 103: GGGCTACAAGCCCCTAACTTAGG), the latter within an adjacent and longer cryptic alternative 3’UTR (Fig. 10). The CRISPR gRNAs were co-injected into fertilized oocytes with Cas9 protein. Confirmation of the correct 4,246 bp deletion without introducing any point mutations was performed through Sanger sequencing. RNA-seq analysis of both hemizygous males and homozygous females also showed the clean deletion and the constitution of a new 3’UTR downstream the canonical exon 37 (Fig. 11 ). Taf1d38 mice were genotyped using primers that anneal in intron 37 before the deletion and in the distal 3’ UTR after the deletion (SEQ ID NO: 104: GGCAATTCAGTCTTTGTGTATAGG; SEQ ID NO: 105: AGCTCTAGCTG ACCTTAAACTCC) .
[0538] Bioinformatic analysis of C-terminal TAF1
[0539] Genomic and protein sequences used in the alignments were downloaded from Ensembl and UniProt public databases. Sequence alignments were done with state-of-the-art aligner ClustalO (Madeira F, et al. Nucleic Acids Res 2019; 47(W1): W636-W641 ). JalView (Waterhouse AM, et al. Bioinformatics 2009; 25(9): 1189-1191 ) was used for sequence alignment editing, visualization and analysis. Chosen conservation measure for protein, among JalView’s repertoire, was Valdar’s conservation (Valdar WS. Proteins 2002; 48(2): 227-241 ). Protein Disorder prediction System (PrDOS) (Ishida T, Kinoshita K. Nucleic Acids Res 2007; 35(Web Server issue): W460-464) was used to study the intrinsic disorder of TAF1. ProsperousPlus (Li F, et al. Brief Bioinform 2023; 24(6)) was run to find putative protease cleavage sites in TAFI . Experimentally determined post-translational modifications were searched through PhosphoSitePlus (Hornbeck PV, et al. Nucleic Acids Res 2015; 43(Database issue): D512-520.), and predicted phosphorylation sites were found with NetPhos (Blom N, et al. J Mol Biol 1999; 294(5): 1351 -1362.) and ELM (Kumar M, et al. Nucleic Acids Res 2024; 52(D1 ): D442-D455.).
[0540] Western Blot
[0541] Samples from human brain were stored at -80 °C and ground with a mortar in a frozen environment to prevent thawing of the samples, resulting in tissue powder. Mouse brains were quickly dissected on an ice-cold plate and the different structures stored at -80 °C.
[0542] Mouse extracts were prepared by homogenizing the brain areas in ice-cold extraction buffer (20 mM HEPES pH 7.4, 100 mM NaCI, 5 mM EDTA, 20 mM NaF, 1% Triton X-100, 1 mM sodium orthovanadate, 1 pM okadaic acid, 5 mM sodium pyrophosphate, 30 mM p-glycerophosphate and protease inhibitors [Complete, Roche]). Homogenates were centrifuged at 15000g for 15 min at 4 °C, the supernatant was collected, and protein content was determined by Quick Start Bradford kit assay (Bio-Rad, 500-0203).
[0543] Nuclear and cytoplasmic extracts from human tissue were prepared by homogenizing the tissue powder in ice-cold extraction buffer A (10 mM HEPES pH 7.9, 10 mM KCI, 0.1 mM EDTA, 0.1 mM EGTA, 1% NP40, 2.5 mM DTT, 1 mM sodium orthovanadate, 1 pM okadaic acid, 5 mM sodium pyrophosphate, 30 mM p-glycerophosphate and protease inhibitors [Complete, Roche]) and centrifuging at 4000g for 1 min at 4 °C. The supernatant constituted the cytoplasmic fraction, and the pellet was washed in extraction buffer A and resuspended in ice-cold extraction buffer C (20 mM HEPES pH 7.9, 25% glycerol, 400 mM NaCI, 1 mM EDTA, 1 mM EGTA, 2.5 mM DTT, 1 mM sodium orthovanadate, 1 pM okadaic acid, 5 mM sodium pyrophosphate, 30 mM P-glycerophosphate and protease inhibitors [Complete, Roche]). After sonication (two pulses of 10 s and 100% amplitude, Labsonic M, Sartorius), 1 % NP40 was added, and the extracts were centrifuged at 15000g for 15 min at 4 °C. The supernatant constituted the nuclear extract, and protein content was determined by Quick Start Bradford kit assay (Bio-Rad, 500-0203).
[0544] Between 10 and 30 pg of total protein was electrophoresed on 6% SDS-polyacrylamide gel, transferred to a nitrocellulose blotting membrane (Amersham Protran 0.45 pm, GE Healthcare Life Sciences, 10600002) and blocked in TBS-T (150 mM NaCI, 20 mM Tris-HCI, pH 7.5, 0.1 % Tween 20) supplemented with 5% non-fat dry milk. Membranes were incubated overnight at 4 °C with either rabbit anti-TAF1 exon 3 (1 :500; Abeam, ab188452), rabbit anti-TAF1 exon 37 (1 :500; Bethyl, A303-504A), rabbit anti-TAF1 exon 38 (1 :500; Bethyl, A303-505A), rabbit anti-SMC1 (1 :2000; Bethyl, A300-055A), mouse anti-SREK1 (1 :500; Abnova, H00140890-B01 P), rabbit anti-SMC1 (1 :2000; Bethyl, A300-055A), rabbit anti-vinculin (1 :15,000; Abeam, ab129002) or mouse anti-p actin (1 :50,000; Sigma, A2228) in TBS-T supplemented with 5% non-fat dry milk. After washings with TBS-T, the membranes were incubated with secondary HRP-conjugated anti-rabbit IgG (1 :2,000, DAKO, P0448) or anti-mouse IgG (1 :2,000, DAKO, P0447), and developed using the ECL detection kit (Perkin Elmer, NEL105001 EA). Images were scanned with densitometer (Bio-Rad, GS-900) and quantified with Imaged software.
[0545] Immunohistochemical and myelin staining analyses
[0546] Human tissue. Postmortem control and MS cortical snap-frozen tissue was analysed. MS lesions were classified according to demyelination and cellular distribution as described (Clemente D, et al. J Neurosci 2011 ; 31 (42): 14899-14909; Ortega MC, et al. Acta Neuropathol 2023; 146(2): 263-282; Kuhlmann T, et al. Acta Neuropathol 2017; 133(1 ): 13-24; Luchetti S, et al. Acta Neuropathol 2018; 135(4): 51 1 -528).
[0547] Immunohistochemistry. The immunohistochemistry protocol was carried out: cryosections (10pm, Leica) were thawed and fixed in 4% paraformaldehyde (PFA) for 1 h at room temperature (RT). Antigen retrieval with citrate buffer 0.1 M pH 6 at 90 °C for 10 min was performed for TAF1 labelling. Endogenous peroxidases were quenched with 10% methanol and 3% H2O2and the slides were blocked with 5% normal goat serum and 0.2 % Triton X-100 in PBS. Slides were then incubated with rabbit anti-TAF1 exon 37+38 (1 :100; Sigma, HPA001075), mouse anti-TAF1 whole protein (1 :30; Santa Cruz, sc735) or anti-human leukocyte antigen (HLA)-DR (1 :200 Agilent, M0746, TAL.1 B5 clone) in blocking solution overnight. Anti-rabbit biotinylated (1 :200; Vector Labs) or anti-mouse biotinylated (1 :200; Vector Labs) secondary antibody was used and the reaction was amplified using the Vectastain Elite ABC reagent (Vector Labs). The peroxidase reaction product was developed with 0.05% 3,3'-diaminobenzidine (DAB, Sigma-Aldrich) and 0.003% H2O2in 0.1 M Tris-HCI, pH 7.6. The reaction was monitored under the microscope and terminated by rinsing the slides with PB.
[0548] Eriochrome cyanine for myelin staining. To visualize myelin after HLA-DR immunostaining, eriochrome cyanine (EC) staining was carried out as described: the sections were air-dried and placed in fresh acetone; then, they were stained in 0.5% EC for 1 h and differentiated in 5% iron alum and borax -ferricyanide for 10 and 5 min, respectively.
[0549] Quantification of TAF1 staining in human samples. To quantify the density of TAF1+cells and the intensity of TAF1 labelling, photomicrographs of the GM and WM from controls and MS samples were captured on BX61 microscope (Olympus) equipped with a color camera CX9000 (MBF). The quantification was performed using the QuPath software. A mean of 5 or 7 ROIs / sample were used respectively for GM or WM, and the regions were defined throughout all the cortical layers in GM (1 .78 ± 0.18 mm2 / ROI) or WM (0.67 ± 0.05 mm2 / ROI). A threshold for TAF1 staining was assessed and the number of TAF1+cells (density: mean value of TAF1+cells / mm2per sample) and TAF1 intensity (mean value of the optical density obtained from all TAF1+cells / sample) were measured.
[0550] Mouse tissue. Mice were deeply anaesthetized with an intraperitoneal injection of Dolethal (2 mg / g) and perfused transcardially with 50 ml of NaCI followed by 4% PFA. Brains and spinal cords were fixed overnight in 4% PFA, cryopreserved for 72 h in 30% sucrose in PBS and added to optimum cutting temperature (OCT) compound (Tissue-Tek, Sakura Finetek Europe, ref. 4583), frozen and stored at -80 °C until use. 30 pm sagittal (brain) or 20 pm transversal (spinal cord) sections were cut on a cryostat (Thermo-Fisher Scientific), collected and stored free floating at -20 °C in glycol containing buffer (30% glycerol, 30% ethylene glycol in 0.02 M PB).
[0551] Immunohistochemistry. Mouse sections were washed in PBS, immersed in 0.3% H2O2 in PBS for 45 min to quench endogenous peroxidase activity and incubated for 1 h in blocking solution (PBS containing 0.5% fetal bovine serum, 0.3% Triton X-100 and 1 % BSA). The sections were incubated overnight at 4 °C with either rat anti-CD68 (1 :200; Abeam ab53444), rabbit anti-lba1 (1 :1000; Wako 019-19741 ) or rabbit anti-GFAP (1 :20,000; Novus NB300-141 ) diluted in blocking solution. Sections were incubated with biotinylated anti-rat or anti-rabbit secondary antibody and then with avidin-biotin complex (Elite Vectastain kit, Vector Laboratories PK-6101 , PK-6104). Chromogen reactions were carried with 3,3'-diaminobenzidine (SIGMAFAST DAB, Sigma-Aldrich, D4293) for 10 min. Finally, the sections were mounted on glass slides and coverslipped with Mowiol (Calbiochem, Cat. 475904). A DMC6200 CMOS camera was used to capture the images.
[0552] Immunofluorescence. Sections were blocked with 4% of serum of the species in which secondary antibody was obtained, diluted in 0.1 M PBS and 0.1% Triton X-100 for 1 h at RT. Sections were incubated overnight at 4 °C with either mouse anti-SPP1 (1 :150; Biotechne MAB808), rat anti-CD3 (1 :50; BioRad) or rat-anti-B220 (1 :200; BD Pharmingen) in 1% serum and 0.1% Triton X-100. Afterwards, the sections were incubated with the secondary fluorochrome-conjugated antibodies (donkey anti-mouse Alexa 488 1 :500 [Thermo-Fisher Scientific, A-21202]; goat anti-rat Alexa 594 1 :400 [Invitrogen]) for 1 h at RT, and nuclei were labelled with Hoechst 33258 (1 .5ug / ml) for 10 min. The sections were mounted on glass slides using Glycergel mounting medium (Dako). A Leica Stellaris 5 confocal microscope was used to capture the images.
[0553] FluoroMyelin. Brain or spinal cord sections were mounted on glass slides and let dry. After hydration in PBS, staining in FluoroMyelin™ (Green Fluorescent Myelin Stain F34651 , Molecular Probes) was performed for 20 min at RT according to the manufacturer’s protocol. Sections were coverslipped with Prolong™ Gold Antifade Mountant (Invitrogen). A fluorescence microscope (Axiovision, Zeiss) was used to capture the images, and Imaged software was used to measure the mean intensity gray value in each region. One level per animal was analysed.
[0554] Black gold. Mouse sections were mounted on SuperFrost slides (Thermo Fisher) and let dry overnight. The staining was performed according to manufacturer’s instructions (Black Gold II Myelin Staining Kit, Avantor BSENTR-100-BG): after rehydration in distilled water, the slides were incubated with Black-Gold II reagent at 65 °C for 12 min in a water bath, and fixed with sodium thiosulfate solution for 3 min. The sections were coverslipped with DePeX (SERVA). A DMC6200 CMOS camera was used to capture the images.
[0555] Transmission electron microscopy
[0556] 7-month-old mice (3 wt and 3 Taf1d38) were deeply anaesthetized with an intraperitoneal injection of Dolethal (2 mg / g) and perfused transcardially with 50 ml of NaCI followed by a solution of 4% PFA and 2 % glutaraldehyde in 0.1 M PB pH 7.4. Spinal cords were dissected and 2 mm long fragments from the lumbar level were collected in the same solution for 2 h at RT and overnight at 4 °C. Fragments were postfixed in 1% osmium tetroxide in 0.1 M cacodylate buffer, dehydrated in ethanol and embedded in Epon-Araldite. Serial ultrathin transversal sections were collected on pioloform-coated, single-hole grids, and stained with uranyl acetate and lead citrate. The sections were analysed with a JEM-1010 transmission electron microscope (Jeol, Japan) equipped with a side-mounted CCD camera Mega View III from Olympus Soft Imaging System GmBH (Muenster, Germany). 20 photographs per white matter region (dorsal, lateral and anterior tracts) and per animal were taken at a magnification of 10,000X. Blind analysis of the g-ratio was performed using ImageJ, from circular areas equivalent to the measured areas of axons and myelin sheath including the axon.
[0557] Electrophysiology
[0558] Propagated compound action potentials (CAPs) were measured in the fornix of 14-month-old wt (n=4) and Taf1 d38 (n=4) mice. Animals were anesthetized with isofluorane and brains were cut in 400 pm-thick horizontal sections (Pelco 3000 vibrating blade microtome, Ted Pella Inc., USA) in an ice-cold solution (215 mM sucrose, 2.5 mM KCI, 26 mM NaHCO3, 1.6 mM NaH2PO4, 1 mM CaCI2, 4 mM MgCI2, 4 mM MgSO4and 20 mM glucose). Sections were incubated in artificial cerebrospinal fluid (aCFS) (124 mM NaCI, 2.5 mM KCI, 10 mM glucose, 25 mM NaHCO3, 1.25 mM NaH2PO4, 2.5 mM CaCI2and 1.3 mM MgCI2) for 1 h. CAPs were recorded with a pulled borosilicate glass pipette (== 1 MQ resistance) filled with aCFS by stimulation of the fornix with a bipolar electrode (CE2C55 FHC, USA) with up to 2 mA, 100 ps pulses (Master-8, AMPI, Israel). The distance between the recording and stimulating electrodes was 1 mm. Using pCIamp 10.0 (Molecular Devices, USA), latencies between the onset of the stimulus artifact and the two negative waves of the CAP’S (N1 and N2) were measured (ms).
[0559] Behavioural testing
[0560] Open field test. Locomotor activity was measured in clear Plexiglas® boxes of 27.5cm x 27.5cm, outfitted with photo-beam detectors for monitoring horizontal and vertical activity. Mouse movements were recorded with a MED Associates’ Activity Monitor and analysed with the MED Associates’ Activity Monitor Data Analysis v.5.93.773 software. Mice were placed in the center of the open-field apparatus and left to move freely for 15 min. Total traveled ambulatory distance and number of vertical episodes when the mouse stands on its hindlimbs (vertical counts) were measured. Limb clasping. Assessment of corticospinal function was performed by the limb clasping test. Mice were held vertically by the tail for 10 s and recorded. Videos were scored blind to the genotype as follows: 0 - normal extension of both hindlimbs;
[0561] 1 - unilateral hindlimb partial retraction during more than 50% of the recorded time;
[0562] 2 - bilateral hindlimb partial retraction during more than 50% of the recorded time;
[0563] 3 - bilateral hindlimb total retraction during more than 50% of the recorded time.
[0564] Rotarod. Motor coordination was assessed in an accelerating rotarod apparatus (Ugo Basile). The first time that the mice performed this test (2-month-old), they were trained at constant speed over two consecutive days: on the first day, 4 trials of 1 min with angular velocity fixed at 4 rpm were carried; and on the second day, 4 trials of 2 min (1 min at 4 rpm and 1 min at 8 rpm) were completed. The test was performed on the third consecutive day with the rotarod set to accelerate from 4 to 40 rpm over 5 min in 4 independent trials. For the following tests (6 and 10-month-old), the training was performed in a single session (2 trials of 1 min at 4 rpm and 2 trials of 2 min with acceleration from 4 to 8 rpm over 1 min and the second minute at 8 rpm). The test was performed on the second day as described above. The latency to fall from the rod was measured as the mean of the four accelerating trials.
[0565] Bulk RNA sequencing
[0566] RNA was isolated from striatum of 2-month-old wild type (wt) mice (n=4) and Taf1d38 mice (n=4) and from normal appearing gray and white matter (NAG+WM) from cortex of control (CT) subjects (n = 4) and individuals with PPMS (n = 3). Mice were subjected to daily manipulation for one week prior to sacrifice in their environment the day of the experiment, to avoid any stress that could modify the brain basal transcriptomic state.
[0567] Total RNA was extracted using the Maxwell 16 LEV simplyRNA Tissue Kit (Promega, AS1280). Quantification and quality determination of RNA was performed on a NanoDrop One spectrophotometer (Thermo Fisher) and an Agilent 2100 Bioanalyzer. For mouse samples an RNA integrity number (RIN) over 8.5 was required, while for human samples the RIN cut-off was set to 6.
[0568] Two independent libraries were prepared from the same mouse samples. The first (n = 4) was polyA+, had a read length of 1 x 50 bp, and was sequenced in HiSeq2500, generating >45 million single-end reads per sample. The second (n = 3) was ribo-depleted, had a read length of 2 x 150 bp and was sequenced in NovaSeq6000, generating >180 million paired-end reads per sample. For human samples, libraries were polyA+, with a read length of 2 x 150 bp and were sequenced in NovaSeq6000, generating >210 million paired-end reads per sample.
[0569] Bulk RNA-seq data processing and analysis
[0570] Quality analyses over the reads were performed using FastQC software. Salmon 1.9.0 alignment mapping algorithm was used, employing indexes constructed with either Mus musculus reference genome (Ensembl GRCm39) or Homo sapiens reference genome (Ensembl GRCh38). Differential expression analysis at gene level was carried with DESeq2 (Love Ml, et al. Genome Biol 2014; 15(12): 550). Genes with less than 20 normalized counts on average in both wt and Taf1d38 groups or CT and PPMS groups were filtered out. Significant differentially expressed genes (DEGs) were defined for an adjusted p value < 0.05.
[0571] DEGs were validated in silico via gene-based association test via MAGMA (de Leeuw CA, et al. PLoS Comput Biol 2015; 11 (4): e1004219), using the GW AS summary statistics from the International Multiple Sclerosis Genetics Consortium MS (47,429 MS cases; 68,374 controls). For MAGMA analysis the model applied was ‘multi=snp-wise’, which aggregates the results of the ‘snp-wise=top’ and ‘snp-wise=mean’ analyses models, improving the power and accuracy of gene-based association tests especially when the genetic model of the disease under study is unknown. For gene-based analysis, gene boundaries were retrieved from RefSeq (GRCh38.p14), lifted over to GRCh37 coordinates, and expanded to 5 kb flanking regions, upstream and downstream respectively, to encompass potential regulatory elements.
[0572] Hierarchical clustering based on sample distances and principal component analysis on human samples revealed that one of the controls (CT4) grouped together with the 3 PPMS samples (Fig. 28) and was removed it from the final analysis. Clustering of the remaining six samples resulted as expected (Fig. 28) and differential expression analysis was performed in this subset. Alternative splicing was also explored looking for TAF1 differential splicing events. We used three complementary software: 1 ) vast-tools v. 2.5.1 (Tapial J, et al. Genome Res 2017; 27(10): 1759-1768) using human junction libraries for the align and combine modules, and a |dPSI| (absolute difference in per cent spliced in) > 10 in the compare module; 2) rMATS v.4.1 .2 (Shen S, et al. Nucleic Acids Res 2012; 40(8): e61 ) using STAR v. 2.7.10 (Dobin A, et al. Bioinformatics 2013; 29(1 ): 15-21 ) to align fastq reads together with a custom splice-junction library, and a FDR < 5% for differential splicing analysis; and 3) Majiq v. 2.5 (Vaquero-Garcia J, et al. Elife 2016; 5: e11752) after alignment with STAR and default parameters and using deltapsi, with a probability of changing >95%.
[0573] We also reanalyzed raw data of ribo-depleted RNA-seq from Elkjaer et al. (Elkjaer ML, et al. Acta Neuropathol Common 2019; 7(1 ): 205) to find DEGs between control white matter samples and white matter with active MS lesions. Fastq files corresponding to samples from the same individual and the same type (i.e all the samples taken from the same control individual, or all the active lesions from a given subject with MS) were concatenated. Then, Salmon alignment mapping, Salmon quantification and DESeq2 differential expression analysis were carried as described earlier.
[0574] Overlap analyses. Overlap between MS and Taf1d38 DEG signatures was analysed using http: / / nemates.org / MA / progs / overlap_stats.html, which runs a hypergeometric test and provides a representation factor (RF), which is the number of overlapping genes divided by the expected number of overlapping genes drawn from two independent groups. PolyA+ datasets were compared to extract common DEG signature between Taf1d38 mice and PPMS NAG+WM; whereas ribo-depleted datasets were compared to find common DEG signature between Taf1d38 mice and MS active lesions.
[0575] Single nuclei RNA sequencing
[0576] 2-month-old wt (n = 3) and Taf1d38 (n = 3) mice were manipulated as described for bulk RNA-seq experiment, and the striata were quickly dissected on an ice-cold plate and stored at -80 °C. The six samples were processed in two batches (first batch n = 2, second batch n = 1 ). The whole striatum was homogenized in nuclei isolation buffer (10 mM Tris-HCI pH 8.0, 0.25 M sucrose, 5 mM MgCI2, 25 mM KCI, 0.1% Triton X-100, 0.1 mM dithiothreitol (DTT) and 0.2 ll / pl RNase inhibitors [RNAsin Plus 40 U / pl]). The suspension was sequentially filtered through a 70-pm strainer followed by a 30-pm strainer and centrifuged for 10 min at 900g at 4 °C. The pellet was resuspended in 150 pl of nuclei isolation buffer and was mixed with 250 pl of 50% iodixanol for debris removal. 500 pl of 29% iodixanol were overlaid with this solution, and centrifuged for 30 min at 13,500g at 4 °C. The supernatant was removed and the pellet was resuspended in 100 pl of phosphate-buffered saline (PBS) with 1% BSA with 0.2 ll / pl RNase inhibitors.
[0577] The nuclei were stained with Trypan Blue 0.4% and counted in a haemocytometer. A total of 10,000 estimated nuclei for each sample were loaded on the Chromium Next GEM Chip G, although a much lower number of nuclei was recovered after sequencing. cDNA libraries were prepared following the Chromium Next GEM Single Cell 3' Reagents Kits v3.1 User Guide and sequenced in NovaSeq6000.
[0578] Single nuclei RNA-seq data processing and analysis
[0579] The 6 samples were aligned with Cellranger v.7.0.1 with reference genome GRCm38-mm10. Filtered count matrices and barcodes were uploaded into R 4.3.2 to be processed following Seurat v4 (Hao Y, et al. Cell 2021 ; 184(13): 3573-3587 e3529). As samples were sequenced in two batches (replicates 1 and 2 run in the first and replicate 3 in the second for both conditions), they were processed separately for single initial QC to avoid differences related to sequencing batches. Cells were filtered for number UMIs per nuclei, number of genes per nuclei and mitochondrial percentage (percent_mito < 2 & nFeature_RNA > 50 & nFeature_RNA < 5000 & nCount_RNA > 5 & nCount_RNA < 10000). The final selected nuclei had the following features: median number of UMI, median number of genes and median mitochondrial percentage for WT samples (8056, 3376, 0.1804958) and for KO samples (8404, 3454, 0.1160056) with a total final number of 10,351 nuclei.
[0580] Log-normalization, feature selection and SCT normalization regressing mitochondrial percentage, was performed in the whole object with the 6 samples. The first 30 PCs were selected using variable features determined by SCTransform (2000 genes). Sample integration was performed with Harmony (RunHarmony) (70 Korsunsky I, et al. Nat Methods 2019; 16(12): 1289-1296), on the selected PCs using as grouping variables the replicates and the batch information. Dimension reduction was performed using runUMAP. KNN graph was constructed using Findneighbors on the harmony reduction, with annoy method, 50 trees and a K of 20. Clusters were identified by running Findclusters to different resolutions from 0.2 to 1 .2. First exploratory view of the clusters showed expression of a specific set of hormonal genes ("Prl", "Gh", "Pou1f1", "Tshb") from nuclei corresponding only to one of the replicates. These genes are not expected to be expressed in the striatum, and probably came from contamination from the hypothalamus or the bed nucleus of the stria terminalis in this specific sample. An “Hormonal_gene_score” was calculated using AddModuleScore with 4 controls and 50 bins for all the nuclei, and we filtered out the possibly contaminated nuclei by setting a “Hormonal_gene_score”>0. All the processing starting from the QC to cluster identification was repeated without these spurious cells, using the parameters and procedures as explained before.
[0581] Clusters were annotated using FindAIIMarkers in resolution 1 and 0.8. Top markers per cluster showed classical markers for main cell types in the striatum allowing the annotation of broad cell types, which was complemented with label transfer using as reference the single cell striatum dataset from Munoz-Manchado et al. (Munoz-Manchado AB, et al. Cell Rep 2018; 24(8): 2179-2190 e2177). The GSE97478 annotated expression matrix was downloaded from GEO repository, and common anchors between query and reference were calculated (FindTransferAnchors, with parameters cca reduction and SCT normalized query and reference). Annotations were transferred with TransferData on the first 30 dimensions of the cca reduction to the query dataset. The predicted cell types corresponded with our broad cell types and identified specific known striatum cells types, such as MSN cell types. Differential expression analysis was performed splitting the object according to cell type and running FindAIIMarkers between wt and Taf1d38 cells in each cluster using the default Wilcoxon Rank Sum test.
[0582] Identification of perturbed cell types
[0583] AUGUR R package (Squair JW, et al. Nat Protoc 2021 ; 16(8): 3836-3873) (https: / / github.com / neurorestore / Augur) was used to identify cell types that showed changes in Taf1d38 compared to wt, defined as perturbed. AUGUR trains a machine-learning model specific to the cell types to predict their condition (wt or Taf1d38) and, as a result, for each cell type AUGUR reports and AUC score of the accuracy of the model on that cell type. AUC score > 0.5 was considered as a significant perturbation. Input data was loaded after normalization and batch correction, as explained in the section above. AUGUR scores are not affected by cluster size; however, different iterations were run by downsampling to the same number of nuclei per comparison.
[0584] Functional enrichment analysis
[0585] The Ingenuity Pathway Analysis (IPA) software (QIAGEN Inc., http: / / www.ingenuity.com) was used for analysis of enriched canonical pathways within Taf1d38 DEG signatures, TAF1 FL>TAF1d38 interactors (after eliminating potential unspecific interactors such as cytoskeletal proteins), genes with higher total RNAPII occupancy at the promoter and genes with lower Ser2P RNAPII occupancy in Taf1d38 mice.
[0586] DEGs were also examined for enriched categories via gene-set enrichment analysis (GSEA) using clusterProfiler v.4.8.1 package in R (72 Wu T, et al. Innovation (Camb) 2021 ; 2(3): 100141 ), using a modified Kolmogorov-Smirnov statistics to assess significant categories. Categories found enriched (adjusted p-value <0.05) were additionally validated for MS with a gene-set level analysis via MAGMA using Z scores from gene p-values.
[0587] Plasmid generation
[0588] Human TAF1 cDNA (TAF1 FL, ENST00000373790.4 transcript) was cloned into a pcDNA3 vector plasmid and EGFP coding sequence was added N-terminally by restriction free cloning (RFC) to generate the pcDNA3-EGFP-TAF1 FL plasmid. Exon 38 sequence was deleted and substituted by intron 37 retention sequence until the stop codon (Fig. 9) by RFC to generate the pcDNA3-EGFP-TAF1d38 plasmid. pEGFP-C1 plasmid was used as an EGFP only expressing plasmid.
[0589] Cell culture and plasmid transfection
[0590] N2a cells (mouse neuroblastoma cell line) were grown in Dulbecco’s modified Eagle medium (DMEM) supplemented with 10% fetal bovine serum (FBS) (Gibco), 2 mM glutamine and penicillin-streptomycin (Sigma) at 37 °C with 5% CO2. Oli-neu cells were grown in DMEM supplemented with 1% horse serum (Gibco), 10 pg / mL insulin (Sigma I9278), 10 pg / mL apotransferrin (Sigma T4382), 100 pM putrescine dihydrochloride (Sigma P5780), 200 nM progesterone (Sigma P0130), 220 nM sodium selenite (Sigma S5261 ) and penicillin-streptomycin (Sigma) at 37 °C with 10% CO2. N2a cells were transfected with either pEGFP-C1 , pcDNA3-EGFP-TAF1 FL or pcDNA3-EGFP-TAF1d38 using Lipofectamine™ 2000 (Invitrogen; 4 pL / pg of plasmid), while jetOPTIMUS transfection agent (Polyplus; 1 .5 pL / pg of plasmid) was used for Oli-neu cells. Technical triplicates were carried for each condition.
[0591] GFP-Trap
[0592] 24 h after transfection, cells were harvested in PBS with protease inhibitors (1 mM PMSF, Phosphatase Inhibitor Cocktail 3 [Sigma-Aldrich, P004] and Complete [Roche]) and centrifuged for 3 min at 500g at 4°C. After washing the pellet in PBS with protease inhibitors, the samples were incubated in lysis buffer (50 mM NaCI, 20 mM TrisHCI pH 8, 0.5% Triton-X100, 2 mM MgCI2) with 0.1 LI / pL of benzonase for 30 min in ice. 250 mM NaCI, 0.5 mM EDTA and 0.1 mM DTT were added to the lysed sample and after 5 min centrifugation at 16,000g at 4°C, the supernatant was collected and 10% of the volume was saved as the input. GFP-Trap dynabeads (Chromotek) were washed with IP buffer (150 mM NaCI, 0.5 mM EDTA, 20 mM TrisHCI pH 7.5) supplemented with protease inhibitors, and each sample was incubated with 25 pL of beads in IP buffer for 1 h at 4 °C. After 3 washes with IP buffer, beads were stored at -20 °C until usage.
[0593] Mass spectrometry and proteomic data processing
[0594] N2a experiment. Proteins were eluted twice in 100 pL of 8 M urea in 100 mM Tris-HCI pH 8, and then reduced, alkylated (15 mM TCEP, 50 mM CAA, 30 min in the dark at RT) and digested with 200 ng of Lys-C (Wako) and 200 ng of trypsin (Promega) for 16 h at 37 °C. Peptides were desalted (C18 stage-tips, speed-vac dried) and re-dissolved in 21 pl of 0.5% formic acid (FA).
[0595] Liquid chromatography-mass spectrometry (LC-MS / MS) was performed by coupling an UltiMate 3000 RSLCnano LC system to an Orbitrap Exploris 480 mass spectrometer (Thermo Fisher Scientific). 5 pL of peptides were loaded into a trap column (Acclaim™ PepMap™ 100 C18 LC Columns 5 pm, 20 mm length) for 3 min at a flow rate of 10 pl / min in 0.1% FA. Then, peptides were transferred to an EASY-Spray PepMap RSLC C18 column (Thermo) (2 pm, 75 pm x 50 cm) operated at 45°C and separated using a 60 min effective gradient (buffer A: 0.1% FA; buffer B: 100% acetonitrile [ACN], 0.1% FA) at a flow rate of 250 nL / min (4% to 6% B the first 2 min, 6% to 33% B the following 58 minutes, plus 10 additional minutes at 98% B). Peptides were sprayed at 1 .5 kV into the mass spectrometer via the EASY-Spray source, with a capillary temperature of 300°C. The mass spectrometer was operated in a data-dependent mode, with an automatic switch between MS and MS / MS scans using a top 18 method (intensity threshold > 3.6e5, dynamic exclusion of 20 s and excluding charges, +1 and > +6). MS spectra were acquired from 350 to 1500 mass-to-charge ratio (m / z) with a resolution of 60,000 FWHM (200 m / z). Ion peptides were isolated using a 1.0 Th window and fragmented using higher-energy collisional dissociation (HCD) with a normalized collision energy of 27. MS / MS spectra resolution was set to 15,000 (200 m / z). The ion target values were 3e6 for MS (maximum IT of 25 ms) and 1 e5 for MS / MS (maximum IT set to auto).
[0596] Raw files were processed with MaxQuant (v 2.1.2.0) against a mouse protein database (UniProtKB / Swiss-Prot, 21 ,990 sequences) supplemented with contaminants, with the following settings: carbamidomethylation of cysteines as a fixed modification; oxidation of methionines and protein N-term acetylation as variable modifications; minimal peptide length = 7 amino acids; maximum of two tryptic missed-cleavages. Results were filtered at 0.01 FDR and loaded in Prostar (v1 .8.0) (Wieczorek S, et al. Bioinformatics 2017; 33(1 ): 135-136) using the intensity values for further statistical analysis. Proteins with less than 4 measured valid values in at least one condition (n=6) were filtered out. A global normalization of Iog2-transformed intensities across samples was performed using the LOESS function. Partial observed missing values (POV) were imputed using the SLSA algorithm, whereas missing in entire condition values (MEC) were imputed with the DETquantile option. Differential analysis was performed using the empirical Bayes statistics Limma. Proteins with an adjusted p-value<0.05 and a Iog2(fold-change)>|3| respect to EGFP were defined as interactors. Proteins with an adjusted p-value<0.05 and a Iog2(fold-change)>|0.6| in the comparison between TAF1 FL and vs TAF1d38 were defined as differential interactors. The adjusted p-value was estimated to be below 5% by Benjamini-Hochberg.
[0597] Oli-neu experiment. Identical volumes of each immunoprecipitate were loaded on STRAP columns (PROTIFI, Farmingdale, NY) and individually digested with trypsin (250 ng) in reducing and alkylating (chloroacetamide) conditions. Tryptic peptides were desalted using C18 tips and subjected to independent LC-MS / MS analysis using a nano liquid chromatography system (Ultimate 3000 nano HPLC system, Thermo Fisher Scientific) coupled to an Orbitrap Exploris 240 mass spectrometer (Thermo Fisher Scientific). 5 pL of each sample were injected on a C18 PepMap trap column (5 pm, 300 pm I.D. x 2 cm, Thermo Scientific) at 20 pL / min, in 0.1 % FA, and the trap column was switched on-line to a C18 PepMap Easy-spray analytical column (2 pm, 100 A, 75 pm I.D. x 50 cm, Thermo Scientific). Equilibration was done in mobile phase A (0.1 % FA), and peptide elution was achieved in a 90 min gradient from 4% - 50% B (0.1 % FA in 80% acetonitrile) at 250 nL / min. Data acquisition was performed in DDA (Data dependent acquisition) mode, in full scan positive mode. Survey MS1 scans were acquired at a resolution of 60,000 at m / z 200, scan range was 375-1200 m / z, with Normalized Automatic Gain Control (AGC) target of 300% and a maximum injection time of 45 ms. The top 20 most intense ions from each MS1 scan were selected and fragmented by Higher-energy collisional dissociation (HCD) of 30%. Resolution for HCD spectra was set to 15,000 at m / z 200, with AGC target of 100 % and maximum ion injection time of 80 ms. Precursor ions with single, unassigned, or more than six charge states from fragmentation selection were excluded.
[0598] Raw files were processed using Proteome Discoverer v2.5 (Thermo Fisher Scientific). MS2 spectra were searched using four Mascot v2.7 as search engine, against a target / decoy database built from sequences corresponding to the Mus musculus reference proteome downloaded from Uniprot Knowledge database including contaminants. Search parameters considered fixed carbamidomethyl modification of cysteine, and the following variable modifications: methionine oxidation, possible pyroglutamic acid from glutamine at the peptide N-terminus and acetylation of the protein N-terminus. The peptide precursor mass tolerance was set to 10 ppm and MS / MS tolerance was 0.02 Da, allowing for up to two missed tryptic cleavage sites. Spectra recovered after FDR <= 0.01 filter were selected for quantitative analysis. Label-free protein quantification used unique + razor peptides. The same criteria as described in the N2a experiment was adopted to define the interactors. MAGMA gene-based analysis was also performed for genes encoding both N2a and Oli-neu the TAF1 FL>TAF1 d38 interactors. In this regard, it should be noted Gtf2h4 is located on the complement region on Chr6 and, therefore, its genetic association might be not as robust as those found for Gtf2b and Flii.
[0599] Bulk ChIP sequencing
[0600] Mice were manipulated as described for bulk RNA-seq experiment, and the striata were quickly dissected on an ice-cold plate. A pool of 4 striata were used for each sample, and a total of 2 wt and 2 Taf1d38 pooled samples were processed in 2 batches. Chromatin immunoprecipitation was performed: after homogenizing the tissue in ice-cold PBS with protease inhibitors (1 mM PMSF; Phosphatase Inhibitor Cocktail 3, Sigma-Aldrich; Complete, Roche), the extracts were crosslinked with 1% formaldehyde for 10 minutes at 37 °C, and the reaction was quenched with 125mM glycine for 5 minutes. After centrifugation at 5000g for 5 minutes at 4°C and washing in ice-cold PBS with protease inhibitors, the pellets were resuspended in 5 mL lysis buffer A (5 mM Pipes pH 8.0, 85 mM KCI, 0.5% NP40) supplemented with protease inhibitors and incubated for 10 minutes on ice. Cell nuclei were pelleted by centrifugation at 4000g for 5 minutes at 4°C and resuspended in 1 .5 mL of lysis buffer B (50 mM Tris HCI pH 8.1 , 1% SDS, 10 mM EDTA) supplemented with protease inhibitors. Chromatin was sonicated for 25 minutes using a Covaris ultraso nicator (E220 Evolution) with the following parameters: 200 cycles / burst, 5% duty factor and 140 W of peak incident power. After a 10 min centrifugation at 20,000g the supernatant was collected. 50 pL of sonicated chromatin was reverse-crosslinked using Proteinase K in lysis buffer B at 65°C overnight and, after phenol chloroform extraction, DNA fragmentation was checked (intended fragment size 200-600 bp). For each immunoprecipitation (IP), 40 pg of chromatin spiked-in with 5% of human chromatin from MCF7 cell line, and 4 pg of the specific antibody (rabbit anti-Rpb1 , Cell Signaling, Cat.14958; rabbit anti-RNA Pol II phospho S2, Abeam, ab5095) was incubated in IP buffer (0,1% SDS, 1% TX-100, 2mM EDTA, 20 mM TrisHCI pH8, 150 mM NaCI) at 4 °C overnight. Inputs consisting of 10% of the mix of chromatin and IP buffer were stored at 4 °C. Dynabeads protein A and G (ThermoFisher) were blocked overnight with IP buffer supplemented with 1 mg / ml BSA, and each IP was incubated with 25 pL of beads for 4 hours at 4 °C. Beads were sequentially washed with wash buffer 1 (20 mM Tris HCI pH 8, 0.1% SDS, 2 mM EDTA, 150 mM NaCI), wash buffer 2 (20 mM Tris HCI pH 8, 0.1% SDS, 2 mM EDTA, 500 mM NaCI), wash buffer s (20 mM Tris HCI pH 8, 1% NP40, 1% sodium deoxycholate, 1 mM EDTA, 250 mM LiCI) and T buffer (10 mM Tris HCI pH8), all supplemented with protease inhibitors. ChlPmentation on beads was performed: using Tn5A enzyme provided by the Proteomics Service of CABD (Centro Andaluz de Biologfa del Desarrollo) in tagmentation buffer (10 mM Tris HCI pH 8, 10% dimethylformamide, 5 mM MgCI2) for 10 min at 37°C. After serial washes in wash buffer 3 and TE buffer (10 mM Tris HCI pH8, 1 mM EDTA), tagmented DNA was eluted with 1% SDS and 100 mM NaHCO3in TE buffer at 50°C for 30 minutes. Overnight decrosslinking was performed using Proteinase K in 175 mM NaCI at 37 °C. After incubation for 10 min with RNase A (ThermoFisher), DNA was purified (QIAGEN, 28106) and inputs were also tagmented as described above and purified. Libraries were amplified for the optimum cycles determined by qPCR reaction using NEBNext High-Fidelity Polymerase (New England Biolabs, M0541 ). Libraries were purified with Sera-Mag Select Beads (GE Healthcare, 29343052) and sequenced after Qubit quantification and size profile analysis with Agilent 2100 Bioanalyzer, using Illumina NextSeq 500 and single-end configuration.
[0601] ChlP-seq data processing and analysis
[0602] The reads from the IP and input libraries were aligned to the mouse and human reference genomes, and signal tracks were obtained using the bamCoverage tool (deepTools) (Ramirez F, et al. Nucleic Acids Res 2014; 42(Web Server issue): W187-191 ).
[0603] To correct technical biases and global differences, spike-in normalization was applied using human chromatin as control. Signal tracks were obtained without normalization and with a bin size of 50 bp. To measure the enrichment ratio in human and mouse signal tracks, the ratio between the IP and input samples was calculated to avoid local biases and background noise, providing a more accurate representation of the occupancy ratio (OR). OR was obtained by dividing the enrichment ratio in mouse and human reads. For the genome tracks, the OR was used to normalize the signal tracks using bamCoverage. The normalization was performed using the RPGC (Reads Per Genomic Content) method with a bin size of 50 bps, a read extension of 300 bps, and the OR as a scale factor.
[0604] Spike-in quantifications were used to generate definitive signal profiles and the normalized RPGC counts tables. Signal profiles were generated by Seqplots (Stempor P, et al. Wellcome Open Res 2016; 1 : 14): signal values were shown by mean in bin tracks of 100 bps, using 15,000 bp as default gene length with 10,000 bp as upstream and downstream regions.
[0605] For signal distribution analysis using normalized RPGCs, two regions were defined: the transcription start site (TSS) region, which comprised the sequence 1 .5 kb up and downstream of the TSS; and the gene body region, which was defined between the TSS and the transcriptional termination site (TTS). With such definition of the TSS region, only genes larger than 4.5 kb were plotted in total RNAPII metaplots, as shorter genes may introduce noise in the promoter / gene body definition. Log2FC of Taf1d38 vs wt counts for each IP and each experiment were calculated. The top 15% genes with higher log2FC for total RNAPII or lower log2FC for Ser2P RNAPII in a given experiment were extracted. Then, the overlapping genes which resulted from crossing the 15% top lists obtained in batches 1 and 2, were taken as the top affected genes. Thus, we defined the 1285 genes with highest total RNAPII promoter region occupancy in Taf1d38 mice and the 2397 genes with lowest Ser2P RNAPII gene body occupancy in Taf1d38 mice.
[0606] To estimate the level of paused RNAPII, we used the pausing index (PI), which is defined as the ratio of total RNAPII enrichment within the promoter to that in the gene body. The PI was estimated as the ratio between the occupation of RNAPII within the promoter (computed as the sum of reads in 400 bp surrounding the TSS), and the occupation of RNAPII within the gene (calculated as the average number of reads in 400 bp windows throughout the gene body, starting 200 bp downstream the TSS). Again, only genes larger than 4.5 kb were considered for PI computation.
[0607] Statistical analysis
[0608] Statistics were performed with SPSS 21 .0 (SPSS Statistic IBM). The normality of the data was checked by Shapiro-Wilk and Kolmogorov-Smirnov tests. For two-group comparison, two-tailed unpaired Student’s t-test or Mann-Whitney U test were performed. For multiple comparisons, data were analysed by one-way ANOVA test followed by a Tukey’s post hoc test. A critical value for significance of p-value < 0.05 was used throughout the study. Benjamini-Hochberg correction was applied for multiple testing for DEG.
Claims
CLAIMS1 . An in vitro method for diagnosing multiple sclerosis in a subject that comprises: a) detecting an isoform of the TATA-box binding protein associated factor 1 (TAF1 ) as established in SEQ ID NO: 1 , or an isoform of the corresponding mRNA, in a biological sample from a subject, wherein i) the isoform of SEQ ID NO: 1 comprises at least one deletion from residues: 1709, 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893 of SEQ ID NO: 1 , or ii) the isoform of the mRNA encodes the isoform i) of SEQ ID NO: 1 , wherein the presence of the isoform i) or the presence of the isoform ii) indicates that the subject suffers from, or is at risk of suffering from, multiple sclerosis.
2. The method according to claim 1 , wherein the isoform i) comprises the deletion from 1775 to residue 1893 of the sequence SEQ ID NO: 1 .
3. The method according to claim 1 , wherein the isoform i) comprises the deletion from 1777 to residue 1893 of the sequence SEQ ID NO: 1.
4. The method according to claim 1 , wherein the isoform i) comprises the deletion from 1803 to residue 1893 of the sequence SEQ ID NO: 1.
5. The method according to claim 1 , wherein the isoform i) comprises the deletion from 1840 to residue 1893 of the sequence SEQ ID NO: 1.
6. The method according to claim 1 , wherein the isoform i) comprises the deletion from 1841 to residue 1893 of the sequence SEQ ID NO: 1 .
7. The method according to claim 1 , wherein the isoform i) comprises the deletion from 1883 to residue 1893 of the sequence SEQ ID NO: 1.
8. The method according to claim 1 , wherein the isoform comprises the sequence SEQ ID NO: 2.
9. The method according to any one of claims 1 to 8, where in the step a) further comprises quantifying the expression levels of at least one biomarker selected from the list consisting of Ctsb, Atp13a2, C1 qc, Ina, C1 qa, C1 qb, H2-K1 , B2m, C4b, H2-D1 , Oasl2, 1117rb, Ifit3, Oas2, Spp1 , Gfap, Oasl a, Cd52, Serpina3n, Ccl4, CxcH O, Cst7, Klk6, Naaladl2, Phex, Ptgds, Hapln2, Anin, Ninj2, Enpp6, II33, Padi2, Trim59, Tppp3, Galnt6 and Abca8b, wherein:- the increased expression levels of at least one biomarker selected from Ctsb, Atp13a2, C1 qc, Ina, C1qa, C1 qb, H2-K1 , B2m, C4b, H2-D1 , Oasl2, 1117rb, Ifit3, Oas2, Spp1 , Gfap, Oasl a, Cd52, Serpina3n, Ccl4, CxcH O and Cst7 with respect to a reference value, or- the decreased expression levels of at least one biomarker selected from Klk6, Naaladl2, Phex, Ptgds, Hapln2, Anin, Ninj2, Enpp6, 1133, Padi2, Trim59, Tppp3, Galnt6 and Abca8b with respect to a reference value, indicates that the subject suffers from that the subject suffers from, or is at risk of suffering from, multiple sclerosis.
10. An in vitro method for the prognosis of multiple sclerosis in a subject that suffers from primary progressive multiple sclerosis (PPMS) or secondary progressive multiple sclerosis (SPMS) that comprises: a) detecting an isoform of the TAF1 as established in SEQ ID NO: 1 , or an isoform of the corresponding mRNA, in a biological sample from a subject, wherein i) the isoform of SEQ ID NO: 1 comprises at least one deletion from residues: 1709, 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893 of SEQ ID NO: 1 , or ii) the isoform of the mRNA encodes the isoform i) of SEQ ID NO: 1 ; b) calculating the abundance or concentration of isoform at step a) with a predetermined reference value for the same biological marker; wherein the abundance or concentration of the isoform i) or the isoform ii) indicates that the subject suffers poor prognosis for primary progressive multiple sclerosis (PPMS) or secondary progressive multiple sclerosis (SPMS).1 1. The method according to claim 10, wherein the isoform i) comprises the deletion from 1788 to residue 1893 of the sequence SEQ ID NO: 1 .
12. The method according to claim 10, wherein the isoform i) comprises the deletion from 1821 to residue 1893 of the sequence SEQ ID NO: 1 .
13. The method according to claim 10, wherein the isoform comprises the sequence SEQ ID NO: 2.
14. A method for screening of an active compound for treating and / or preventing multiple sclerosis that comprises: a) administering a potentially active compound for the treatment and / or prevention of multiple sclerosis to a cell or a tissue that expresses an isoform of the TAF1 as established in SEQ ID NO: 1 that comprises at least one deletion from residues: 1709, 1724, 1760, 1764, 1765, 1774, 1775, 1776, 1777, 1787, 1788, 1793, 1800, 1802, 1803, 1804, 1820, 1821 , 1825, 1838, 1839, 1840, 1841 , 1845, 1871 , 1880, 1882, 1883, 1884, 1889 to residue 1893 of SEQ ID NO: 1 , b) evaluating the change in the gene expression, phenotype change and / or histological change of the cell, tissue or non-human animal model in response to the potentially active compound.
15. The method according to claim 14 wherein in step a) the isoform comprises the sequence SEQ ID NO: 2.
16. The method according to claim 14 wherein in step a) the isoform, comprises the sequence SEQ ID NO: 3.
Citation Information
Patent Citations
Diagnosis and treatment of multiple sclerosis
WO2002059604A2