IN VITRO DIAGNOSIS OF MULTIPLE SCLEROSIS
Patent Information
- Application Number
- DE602021038335
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-07-14
- Filing Date
- 2021-07-13
- Publication Date
- 2025-09-10
- Estimated Expiration
- 2041-07-13
AI Technical Summary
Current diagnostic methods for multiple sclerosis lack definitive biomarkers and are reliant on subjective evaluations and invasive procedures, while existing HERV-W expression studies suffer from cross-reactivity, non-detection of mutated sequences, and artifact formation, leading to inconclusive results.
A method using next-generation sequencing (NGS) and bioinformatics analysis to quantify HERV transcript expression levels in blood samples, identifying specific HERV sequences differentially expressed in multiple sclerosis patients, overcoming cross-reactivity and ambiguity by precise molecular analysis.
The method enables clear differentiation between healthy individuals and MS patients using 17 selected HERV transcripts, providing a definitive diagnostic biomarker for multiple sclerosis with high predictive power.
Description
Technical field
[0001] The present invention relates to a method based on the analysis of the expression of transcripts deriving from Human Endogenous Retrovirus (HERV) for the in vitro diagnosis of Multiple Sclerosis (MS), namely a transcript of SEQ ID NO: 1.Prior art
[0002] Multiple Sclerosis (MS) is a demyelinating neurodegenerative disease affecting approximately 2.3 million people worldwide. The exact etiology and causes of the disease are still partially unknown, but suggest that the disease originates from a combination of environmental and genetic factors.
[0003] To date, there is no single test available that can definitely and indisputably confirm the diagnosis of MS, just as there are no diagnostic or predictive biomarkers for the disease.
[0004] Currently, the diagnosis is made by the doctor on the basis of three elements: the symptoms reported by the patient, the neurological examination and the instrumental (magnetic resonance - evoked potentials) and biological (blood and cerebrospinal fluid) analyses. The set of results and prolonged clinical observation allow to confirm or exclude the pathology but, to date, a single test is not available that can definitely and indisputably confirm the diagnosis of MS, just as there are no diagnostic biomarkers or predictive for the disease. Diagnosis is based on parameters affected in part by subjective evaluation (patient-reported symptoms, doctor's evaluation) and in part requiring expensive (magnetic resonance imaging) or invasive practices for the patient (cerebrospinal fluid sampling) to exclude the presence of MS.
[0005] Within the human genome are contained sequences that represent remnants of ancient infections by human endogenous retroviruses (HERV). These sequences make up about 8% of our DNA and are highly similar to each other in terms of nucleotide sequence. While the genomic characterization of HERVs is still ongoing, an impressive amount of data has been obtained regarding their general expression in different tissues (Grandi and Tramontano 2017, 2018). Among the HERVs, one of the most studied and most expressed groups is the W group, which has been studied extensively for its putative role in various diseases, such as cancer, inflammation and autoimmunity. Despite the great interest in the link between HERV-W expression and human pathogenesis, no conclusive correlation has so far been demonstrated. In general, (i) the absence of a correct identification of the specific HERV-W sequences expressed in a given condition; and (ii) the lack of studies that try to combine the different observations under the same experimental conditions are the main problems that prevent the definitive evaluation of the impact of HERV-W on human pathophysiology.
[0006] Regarding the correlation between HERV and multiple sclerosis, several studies have been proposed over time in which a link between group members and the pathology has been sought. The MSRV (MS-associated retrovirus) presented in Dolei and Perron 2009; H Perron et al. 1997; Voisset et al. 1999 as a "proposed cofactor of multiple sclerosis" has been widely debated in the literature for years, there is currently little evidence not only of its actual involvement, but even of its real existence. It has been proposed as an exogenous member of the HERV-W group but its nature has not been clarified at the moment. MSRV has never been isolated or confirmed as an existing virus. Among the various hypotheses made in this regard, there is also the possibility that it is actually an artifact originated collaterally during HERV study experiments based on sequence amplification processes (Grandi et al., 2016; Laufer et al., 2009). In the past years the same inventors have developed a database characterizing in detail about 3200 HERV sequences integrated into the human genome and classified in a total of 70 taxonomic groups (Vargiu et al. 2016), in this work which focused solely on the classification of these sequences does not refer to the correlation with the pathology of multiple sclerosis. These sequences are present in various positions on all human chromosomes, and their identification and characterization was possible thanks to an innovative software, RetroTector, which scans the genome identifying individual HERV sequences by recognizing conserved retroviral motifs and their relationship within the same integrated sequence (Sperber et al. 2007). The identified sequences are then classified using a multiple approach that includes phylogenetic relationships and structural comparison to a collection of reference retroviral sequences (Jern et al. 2005; Sperber et al. 2007). The dataset obtained uniquely identifies each HERV sequence in terms of genomic coordinates and nucleotide sequence, allowing it to be distinguished from the others and to specifically associate the resulting expression products. Moreover, this dataset, being based on the identification of conserved retroviral domains, includes the most intact HERV sequences and therefore more suitable to be still able to produce expression transcripts. To maximize the level of detection of HERV sequences of possible interest, the starting database was then integrated into a second characterization phase, focused on all members (regardless of the degree of conservation of the individual sequences) of some individual groups. HERV of particular relevance for MS and / or probable transcriptional activity, such as the HERV-W, HERVK (HML10), HERV-K (HML7) and HERV-K (HML6) groups (Grandi et al., 2016, 2017, 2021; Pisano et al., 2019). In particular, the HERV-W group is currently the most studied in relation to MS, with the first studies dating back to the 1990s and reporting the isolation of viral particles attributable to these sequences in MS patients (Grandi and Tramontano 2017). In the following decades, the HERV-W group was repeatedly hypothesized as a probable contributor to the disease, based on the results of multiple studies reporting i) greater expression of the group in biological fluids and brains of MS patients, ii) the presence of HERV-W proteins in the brain of MS patients, especially at the level of active lesions, iii) an immune response of the patient, both in the form of specific antibodies and with non-specific inflammatory phenomena against these proteins (Grandi and Tramontano 2017). In recent years, the HERV-W Envelope protein has shown neuro-inflammatory properties capable of developing cerebral symptoms coinciding with the main characteristics of MS (Van Horssen et al. 2016). The same protein has been studied in mouse models, confirming the ability to induce neuro-inflammation and damage to oligodendrocytes and myelin, resulting in an MS-like allergic encephalomyelitis (Antony et al. 2004; Grandi and Tramontano 2017, 2018; Perron et al. 2013). In light of these findings, a monoclonal antibody against HERV-W Env has shown neutralizing effects in vitro and in MS mouse models, restoring myelin expression, and is currently in phase II of clinical development (Curtin et al. 2015). Despite a clear indication of the pathogenic potential and of its exploitation as a therapeutic target, to date the debated association between HERV-W and MS still lacks both the identification of the specific HERV-W sequences that produce this Env protein, and the characterization of its / their expression in MS patients compared to healthy individuals. As for the other characterized groups, the HERV-K group (HML10) has a localized insertion in the C4 gene of the major histocompatibility complex, determining a polymorphic variant that has been investigated for a possible role in autoimmune diseases, this being the region genomics associated with the largest number of disorders of this type (Grandi et al. 2017; Trowsdale and Knight 2013). The KERV-K group (HML6), on the other hand, is one of the most expressed among all the HERV groups: various transcripts and peptides attributable to it have been reported both in physiological conditions and in various pathologies (Pisano et al. 2019). There are several studies in which attempts have been made to associate HERVs and / or their expression with multiple sclerosis.
[0007] To date, however, no HERV group or sequence has been recognized as actually associated with a human disease in general, and more particularly with multiple sclerosis.
[0008] It is evident that one of the main problems encountered in the study of the expression of HERV in various pathologies, including sclerosis, is precisely the fact that - being highly identical sequences - the use of generic primers or probes (and therefore not selective for each individual member of the HERV group analyzed) leads on the one hand to amplify only an indiscriminate portion of the sequences of interest, and on the other hand it makes the in vitro formation of chimeric sequences in reality non-existent very likely. Very often then the primers are designed for a single gene of the group, so if that gene is mutated or even is missing due to a deletion (as often happens for these sequences, which have been in the primate genome for millions of years and have accumulated many mutations) a portion of group members will not be detected.
[0009] The vast majority of expression studies conducted so far in patients with multiple sclerosis have been based on this approach, in which, through the use of unique primers for an entire HERV group, an increased expression was observed in patients compared to controls, without but to know which specific members of the group were actually expressed. The vast majority of the works cited in (Grandi and Tramontano 2017, 2018; Grandi and Tramontano 2018) belong to this vast group of studies. The results in this sense are largely contradictory, with different works that on the one hand support the increased expression of a group in MS patients and on the other report the absence of differences between cases and controls (Grandi and Tramontano 2017 and 2018). Finally, many of the studies cited are aimed at demonstrating that a HERV group or protein expressed more in patients is involved in the development of the disease, something that has never been confirmed for any pathology so far. The need is therefore felt to investigate the expression of HERV sequences with a different technological approach from that on which traditional HERV expression studies are based, allowing to identify specific HERV sequences of interest in the different groups with high resolution instead of measuring specifically the expression of a set of variously classified elements.
[0010] Grandi Nicole: 5 November 2019 (2019-11-05), Third International Workshop on Human Endogenous Retroviruses and Diseases; URL:https: / / www.hervsanddisease.com / program-2019 / indicates the aim of clarifying the contribution of individual HERV loci to both healthy and MS transcriptome, to identify loci differentially modulated in MS patients. Using RNA-seq from peripheral blood cells, a total of 53 HERVs loci were found to be differentially expressed (DE-HERVs) in MS patients as compared to healthy controls.
[0011] MORANDI ELENA ET AL: "The association between human endogenous retroviruses and multiple sclerosis: A systematic review and meta-analysis", PLOS ONE, vol. 12, no. 2, 16 February 2017 (2017-02-16), is a literature review referring to different studies according to which RNA expression of HERVs as detected as being differentially expressed in multiple sclerosis patients as compared to healthy controls. In these studies, the analysis of HERV expression relies on methods that cannot distinguish highly similar HERV sequences, being based on their nucleotide sequence.
[0012] In NOWAK JERZY ET AL: "Multiple Sclerosis-Associated Virus-Related pol Sequences Found Both in Multiple Sclerosis and Healthy Donors are More Frequently Expressed in Multiple Sclerosis Patients", JOURNAL OF NEUROVIROLOGY, vol. 9, no. 1, 1 January 2003 (2003-01-01), pages 112-117, GB ISSN: 1355-0284, DOI: 10.1080 / 13550280390173355, higher frequency of expression of MSRV sequences was detected in lymphocyte RNA as well as in the serum of MS patients in comparison to control groups, including healthy subjects. On the basis of obtained results, NOWAK claims that expression of MSRV RNA is specific for MS. The actual existance of MSRV is debated in literature, since it is claimed as an exogenous counterpart of the HERV-W group that wasbut never isolated, and several studies suggest it could represent an artifact arose from the in vitro recombination of highly similar amplicons derived from identical HERV-W group members, confirming the biases associated to this technique when primer specificity is not carefully accounted.
[0013] US 2004 / 054133 discloses a method of diagnosing multiple sclerosis by detection of RNA encoding HERV-W superantigen, e.g. using amplification and sequencing.
[0014] In view of the above, the need to identify markers strongly and uniquely correlated to the presence of MS pathology and the need for an analysis method capable of overcoming the technical problems related to the methods commonly used for the analysis of HERV-W sequences remain unsatisfied, i.e. cross-reactivity (amplification of several similar sequences recognized by the same primers), nondetection (non-amplification of mutated sequences at the level of the sites recognized by the primers), and the formation of artifacts.Summary of the invention
[0015] The problems posed by the known art have been solved with the method according to the present invention.
[0016] The invention is set out in the appended set of claims and relates to SEQ ID NO: 1.
[0017] The inventors have now devised a method for selecting molecular markers for use in the diagnosis of multiple sclerosis; in particular, the inventors carried out a precise molecular analysis by quantifying the expression levels of HERV transcripts in samples of nucleic acids isolated from blood obtained from patients with multiple sclerosis, comparing said levels with those detected in samples obtained from healthy donors.
[0018] The inventors found and selected from a database of about 3200 HERV sequences (HERVdb), 43 sequences with a significant modulation of expression, which was found to be up-regulated or down-regulated, in patients with multiple sclerosis, thus correlating this de-regulation with the presence of the pathology. Finally, the inventors also selected 17 particularly preferred HERV transcripts, with particularly significant levels of deregulation in patients suffering from multiple sclerosis.
[0019] Therefore, the subject of the present disclosure are the HERVs shown in tables 2 and 3 and in particular the HERV of SEQ ID NO: 1 for use as biomarker in the in vitro diagnosis of multiple sclerosis.
[0020] The present disclosure also relates to a method for the diagnosis of multiple sclerosis, the method comprising the steps of: isolate total cellular RNA ex vivo from a sample obtained from an individual determine the expression levels of HERV present at the loci indicated in tables 2 and / or 3 compare said expression levels with the expression levels of the respective HERV detected in a sample obtained from a healthy subject, in which altered expression levels (both positive and negative) of at least one of said HERV sequences in the sample obtained from the individual compared to the levels detected in a sample derived from a healthy control are indicative of the presence of MS in the subject.
[0021] The subject of the present invention is an in vitro method for the diagnosis of multiple sclerosis in an individual, the method comprising the following basic steps: to isolate ex vivo an RNA sample from a biological sample obtained from a subject; determine the expression levels of the transcript of the HERV sequence SEQ ID No. 1; compare these expression levels with the expression levels of the same HERV detected in an RNA sample isolated ex vivo from a sample obtained from a healthy individual in which the decreased expression levels of these HERV in the test sample compared to the levels detected in the sample obtained from a healthy individual are indicative of the presence of multiple sclerosis in the subject.
[0022] In a preferred embodiment, the expression analysis step of the HERV transcripts is carried out using next generation sequencing techniques, in particular RNAseq and subsequent bioinformatic analysis of the results.
[0023] In particular, the HERV of SEQ ID No 1 is for the use in an in vitro diagnostic method of multiple sclerosis; in particular, the object of the present invention is the HERV of SEQ ID No 1 for use as biomarker in the in vitro diagnosis of multiple sclerosis, in which expression levels decreased by the HERV transcript of SEQ ID NO: 1 compared to the levels found in samples obtained from a healthy individual is indicative of the presence of multiple sclerosis in the subject.
[0024] Other objects will become apparent from the detailed description of the invention which follows.Brief description of the Figures
[0025] Figure 1. Bioinformatics pipeline used for expression analysis The main software and tools used in the various phases are shown in square brackets. MS = multiple sclerosis, HERV = human endogenous retroviruses, hg38 = reference sequence of the human genome GRCh38 / hg38, TPM = transcripts per million kilobases, deHERV = differentially expressed HERV sequences in patients with multiple sclerosis. Figure 2A. Quality control of the reads contained in the RNAseq GSE77598 dataset Performed using the FastQC software. Mean quality score = average quality for each nucleotide position of the reads of each sample, by sequence quality score = average quality of the reads of each sample, adapter contents t= residual presence of adapters used in the sequencing process, by sequence GC content = percentage content of GC bases (expected = 50%), by base N content = percentage of unassigned bases for each nucleotide position of the reads of each sample, sequence length distribution = distribution of the reads of each sample in terms of length (expected = 100 bp). Figure 2B. Quality control of the mapping of the reads to the reference sequence of the human genome (GRCh38 / hg38) Carried out using the FastQC software. STAR alignment score= average quality of the alignment of the reads of each sample in terms of the number of uniquely mapped reads. The graph instead shows the relationship between the number of mapped reads and the total number of reads per sample. Figure 3. Principal Component Analysis (PCA) Responsible for Variance in HERV Expression Between Samples The healthy control or multiple sclerosis patient condition is the major variance component of HERV expression in samples (PC1, responsible for 39 % of variance). Note the clear division of healthy controls and patients, indicative of the discriminating power of HERV expression for multiple sclerosis. The samples representing biological triplicates of the same individual are grouped together. Figure 4. Distance between samples based on global expression of HERV (A) and cellular genes (B) Note the clear division of healthy controls and patients based on HERV expression, indicative of the discriminating power for multiple sclerosis. The triplicate representative samples of the same individual are grouped together. Figure 5. Heatmap of the 400 HERV loci (A) and cellular genes (B) with the highest mean expression Note the clear division of healthy controls and patients, indicative of the discriminating power of HERV expression for multiple sclerosis (A). On the other hand, this division is not observable considering the cellular genes (B). The triplicate representative samples of the same individual are grouped together. Figure 6. Heatmap of 140 HERV loci with the highest mean expression This set of loci, extracted from the larger one represented in Figure 5, represents the minimal set of HERV loci with discriminating power for multiple sclerosis, as indicated by the clear division of healthy controls and patients. The triplicate representative samples of the same individual are grouped together. Figure 7. Heatmap of 51 deHERV discriminating loci based on mean (A) and variance (B) expression values This set of HERV loci was found to be differentially expressed in the multiple sclerosis condition based on the statistical comparison between the expression levels of the same locus in healthy controls and patients (adjusted p-value ≤ 0.01 - log2 fold change> 1 in case of expression upregulation or <-1 in case of expression downregulation). Note the clear division of healthy controls and patients, indicative of the discriminating power of HERV expression for multiple sclerosis. Figure 8. Heatmap of the 15 discriminating deHERV loci with highest mean expression (A) and variance (B) This subset of deHERV loci, extracted from the larger one represented in Figure 7, represents the minimum set of deHERV loci with discriminating power for the multiple sclerosis, as indicated by the clear division of healthy controls and patients. Detailed description
[0026] The following definitions are given within the scope of the present disclosure.
[0027] By primer we mean an oligonucleotide of variable length which acts as a starting point for the synthesis of a nucleic acid of interest, which is recognized by sequence complementarity, allowing its selective amplification by means of a PCR (Polymerase Chain Reaction) reaction.
[0028] A probe is defined as an oligonucleotide of variable length that hybridizes to a complementary target nucleic acid.
[0029] For NGS (Next Generation Sequencing) or Deep sequencing is a set of technologies that allow to sequence within a narrow time entire genomes or transcriptomes
[0030] For RNA-Seq means the application of a method to NGS sequencing of an RNA sample
[0031] For transcriptome is means the sequencing product of all the transcripts present in a certain cell or tissue population, using NGS techniques. Within the scope of the present invention, transcriptome sequencing was performed specifically on the polyadenylated portion of the transcripts (corresponding to messenger RNA or mRNA).
[0032] By "reads" we mean the sequences originating from the sequencing process, which represent the fragments of cellular transcripts and will be present in quantities directly proportional to the level of expression of the gene that originated them. Reads can be of variable length, and be single or paired at one end.
[0033] By mapping we mean the alignment of the transcriptome reads to a reference genomic sequence, in order to identify its genetic origin. This process is implemented using bioinformatics software and involves the elimination of reads that do not map to a unique genomic location.
[0034] For expression analysis quantitative means the bioinformatic quantification of the expression of a given gene or genomic element by counting the map that reads at its genomic coordinates and normalization based on factors such as the length of the gene itself and the sequencing depth of the transcriptome. In this way, an expression value quantified in Transcripts Per Million Kilobases (TPM) is obtained, which allows the comparison of the expression of different genes.
[0035] By differential expression analysis we mean the statistical comparison of the expression levels of specific cellular genes or HERV, to identify significantly modulated elements (over-expressed or under-expressed) in a certain condition (in this case, MS).
[0036] By HERV locus we mean the single insertions of endogenous retroviruses present in the human genome (plural = loci).
[0037] The inventors with the method of the invention described below have overcome the technical problems relating to the traditional analysis of HERV sequences and their expression caused by the use of primers and probes, i.e. cross reactivity (amplification of several similar sequences recognized by the same primers) and the lack of detection (lack of amplification of mutated sequences at the level of the sites recognized by the primers), which together with the frequent formation of recombinant artifacts makes it difficult to identify which specific HERV sequence is actually expressed in a certain condition. Using the method described below, the inventors were able to analyze the expression of a large number of HERV sequences in healthy subjects and patients with multiple sclerosis (MS) and were able to arrive at a selection, establishing a clear correlation between the expression of specific HERV sequences present in the human genome and the presence of multiple sclerosis (MS). In particular, the inventors starting from a database (HERVdb) containing about 3200 sequences attributable to the related HERV genomic characterization publications of the Proponents (Grandi et al. 2016, 2017; Pisano et al. 2019; Vargiu et al. 2016) they first identified a set of 43 significant sequences and from these they selected a group of 17 sequences whose most variable expression was found to be predictive of the presence of the disease, allowing to clearly distinguish healthy individuals from MS patients.
[0038] The inventors have developed a method capable of obtaining a precise analysis of the expression of HERV transcripts, overcoming the problems of the known art and satisfying the need to detect each specific HERV expressed using Next Generation Sequencing (NGS) techniques and using universal adapters for amplify all polyadenylated cell transcripts, which are then uniquely mapped by bioinformatic alignment of each read to the human genome sequence.
[0039] In this way, the specific quantification of the expression of each cellular gene and HERV sequence is based on the quantification of the reads at the level of their unique position in the genome, and is therefore not susceptible to cross-reactivity or ambiguous results. In the specific case of HERVs, the use of sequencing technical characteristics such as paired reads with a length greater than 75 bases each guarantee a high mapping specificity even of highly similar sequences, since they will most likely include a portion of unique sequence. Another guarantee of specificity is that reads that do not uniquely map a specific position in the genome are excluded from the quantitative analysis, avoiding erroneous estimates.
[0040] To achieve this sensitivity of discrimination, it is first of all necessary to know the precise position in the genome of the elements to be analyzed.
[0041] In particular, the inventors used data derived from NGS experiments in which the transcriptome of healthy individuals and individuals with multiple sclerosis was sequenced.
[0042] HERVdb was then used as a starting point to analyze the expression of these elements in the transcriptome of monocytes from peripheral blood of healthy controls (3) and MS patients (5) sequenced in triplicate (24 total samples, contained in a public record registered in the GEO repository (https: / / www.ncbi.nlm.nih.gov / geo / , accession number: GSE77598). The comparison of the expression levels of each HERV between the two conditions allowed to identify specific sequences that showed a differential expression between healthy controls and MS patients. The methodology used to conduct the analysis is based on the use of an optimized bioinformatics process for the identification, mapping and quantitative evaluation of the transcripts produced by the HERV loci included in the said HERVdb, so compararne statistically presence and abundance in the two conditions taken into consideration (healthy controls and patients with MS).
[0043] This process allowed a first phase, to analyze the global expression of HERV loci by selecting the 400 most expressed loci and obtaining a general overview of the expression of these transcripts in healthy and affected individuals (Figure 5A). The inventors then deepened the analysis of the data obtained by considering the transcripts that had the highest variance and standard deviation, as well as reducing the number of HERV loci considered, down to a minimum discriminant set of 140 sequences (Listed in Table 1). The set of 140 HERV loci described in Table 1 is endowed with discriminating power on the basis of the overall expression in the samples under analysis selected on the basis of the expression mean. In order to obtain more effective diagnostic markers, the inventors carried out a further selection step by searching for transcripts with a significant difference in expression between healthy individuals and MS patients. A differential expression analysis was therefore conducted verifying which HERV loci contained in the HERVdb - beyond their expression values - showed a statistically significant variation between healthy controls and MS patients.
[0044] The expression variation of the HERV loci between healthy controls and MS patients was measured and expressed in log2 fold change. The HERV loci whose variation between controls and patients reached a threshold of statistical significance, represented by an adjusted p-value (p-adj, according to Benjamini-Hochberg) ≤ 0.01 and by a expression variation between healthy individuals and individuals with MS measured in terms of log2 fold change; loci with log2 fold change value> 1 (in case of expression up-regulation) or <-1 (in case of expression down-regulation) were selected as significant. The HERV loci considered for their discriminating power have an average length of about 8000 nucleotides.
[0045] The differential expression analysis allowed to identify 43 HERV loci located on autosomal chromosomes and reported in table 2 which are differentially expressed (deHERV) in healthy and MS patients. Of these 43 HERVs, 41 are down-regulated in patients with multiple sclerosis and only 2 are up-regulated (N ° 4 and 29).
[0046] Starting from this set, the inventors finally selected an even smaller number of deHERV loci with the highest discriminating power, identifying as the final subset a total of 17 deHERV loci, ordered on the basis of decreasing mean and variance values of expression, in order to select the most significant ones (Figure 8). These 17 HERVs make it possible to clearly distinguish MS patients from healthy controls (Figure 8, Table 3). The deHERV loci included in this minimal subset have statistical significance levels between 7.08-14 and 0.005, and most loci are down-regulated, with a log2 fold change value between -1.75 and -1.01 in MS patients, except two up-regulated loci with log2 fold change of 1.86 and 2.62.
[0047] The inventors have therefore selected, starting from a large initial dataset, HERV sequences with a diagnostic and predictive character of multiple sclerosis. The inventors therefore surprisingly found that the analysis of the expression of HERV present in the initial dataset and selected for their mean or variance of expression in patients has the value of a highly diagnostic and predictive biomarker of multiple sclerosis. In particular, the inventors selected 43 HERV transcripts whose differential expression in healthy subjects and MS patients has a high diagnostic power and, among these, 17 particularly preferred transcripts for the diagnosis of the pathology.
[0048] The inventors have developed a method for the ex vivo diagnosis of this pathology which includes the following basic steps isolate total cellular RNA ex vivo from a biological sample obtained from an individual; determine the expression levels of HERV present at the loci indicated in tables 2 and / or 3; compare said expression levels with the expression levels of the respective HERVs detected in a biological sample obtained from a healthy subject, in which decreased expression levels of said sequence HERV in the sample obtained from the individual versus the levels found in a sample derived from a healthy control are indicative of the presence of MS in the subject.
[0049] In a preferred embodiment, the method according to the disclosure comprises the steps of: isolate total cellular RNA ex vivo from a biological sample obtained from an individual; selection and deep sequencing of polyadenylated RNA to generate sequencing reads and subsequent bioinformatic analysis performed through a computer system. Said analysis comprising align the sequencing reads obtained with a reference human genome; determining relative quantities of the reads corresponding to the sequences reported in table 2 and / or preferably in table 3; statistically compare said relative quantities with the relative quantities detected in a sample obtained from a healthy subject, preferably by calculating the Fold change on a logarithmic scale and the adjusted p value and in which they are selected as significant with an adjusted p value of ≤ 0.01 and log2 fold change> 1 (in case of expression up-regulation) or <-1 (in case of expression down-regulation). in which altered expression levels (both positive and negative) of at least one of said HERV sequences in the sample obtained from the individual compared to the levels detected in a sample derived from a healthy control are indicative of the presence of MS in the subject.
[0050] The results obtained using the method according to the invention are intended to be submitted to the attending physician for subsequent and appropriate diagnostic investigations in order to confirm the diagnosis and identify suitable treatments.
[0051] In detail, the inventors have used a method and a bioinformatic process represented in Figure 1 and schematized in the following phases:· Step 0: RNAseq Dataset Download and Quality Control
[0052] The GSE77598 dataset was downloaded from NCBI's Gene Expression Omnibus (GEO) (https: / / www.ncbi.nlm.nih.gov / geo / ) in the form of raw transcriptomic data, ie including the pairs of reads of each sample (read_1 and read_2) in two files of type fastq. The raw data was then subjected to Fastqc software quality control (https: / / www.bioinformatics.babraham.ac.uk / projects / fastqc / ) to evaluate parameters related to the sequencing process, including length and total number of reads, nucleotide content, presence of ambiguous bases, content of GC bases and residual presence of adapters used for sequencing. The overall quality was satisfactory (Figure 2A), allowing the use of the dataset for the subsequent phases.· Phase 1: Mapping the reads to the reference human genome
[0053] The reads_1 and _2 of each sample were aligned to the most up-to-date reference sequence of the human genome currently available (GRCh38 / hg 38) downloaded as a GTF file from NCBI Genome Browser (https: / / hgdownload.soe.ucsc.edu / downloads) using the STAR software, and all unmapped or non-uniquely mapped reads (therefore aligned in two or more genomic locations) have been eliminated. Even the obtained alignments were checked in terms of quality before proceeding to the subsequent stages, showing a good percentage of reads mapped uniquely (on average 86%) (Figure 2B)· Phase 2: quantification of the expression of each HERV and cellular gene
[0054] The alignments carried out were used to count the reads aligned to each of the genomic coordinates uniquely identifying the HERV loci (HERVdb) and cellular genes (Gencode, version 29) (Harrow et al. 2012) using the HtseqCount software (Anders, Pyl, and Huber 2015). This allowed to obtain the raw number of counts of the reads uniquely mapped to each HERV / gene. Furthermore, in order to have comparable values, the expression levels of each HERV / gene were also calculated in terms of Transcripts Per Million kilobases (TPM), by importing the raw counts on the Rstudio (Rstudio Team 2016) platform. normalization based on factors such as HERV / gene length and transcriptome sequencing depth.· Phase 3: analysis of the variability of the expression of HERV and cellular genes in MS patients compared to healthy controls
[0055] The raw results obtained from the count of the reads mapped in each HERV locus and cellular gene were imported into the Rstudio platform (Rstudio Team 2016) and subjected to Rlog normalization (regularized-logarithm transformation, transforms the raw counts into log2 logarithmic scale, minimizing the differences between samples and with respect to the size of the dataset). The first analysis carried out was to identify the main components (PC) determining the variability in the expression of HERV and genes among all the samples. This analysis was carried out using unsupervised PCA (Principal Component Analysis), i.e. without indicating a particular condition of interest but evaluating the PCs based on the distribution of the samples according to their variance values in the expression of HERV and cellular genes. The result of the PCA (Figure 3) identified as PC1 (ie the main variance component) the status of "healthy" or "affected by MS". In fact, this condition alone is responsible for 39% of variance, and determines the net division of the samples by PC1 into two distinct groups, which correspond precisely to controls and patients (Figure 3). Confirming the quality of the dataset and the identifying power of the pipeline, the samples derived from the triplicate sequencing of the same individual in turn form well appreciable subgroups, reflecting an inter-individual variability that in any case does not influence the clear discrimination of patients and controls.
[0056] In light of this characteristic global expression difference in the presence of MS, the variability of expression of individual HERV and cellular genes between MS patients and healthy controls was compared both in terms of distance between samples (Figure 4) and in terms of mean expression (Figure 5). Both analyses confirmed that the global variation of HERV expression has a discriminating power for the condition of MS, allowing to clearly divide the sick individuals from the healthy controls into two well-defined groups, including in turn the subgroups with the 3 replicates of sequencing of the same individual (Figures 4A and 5A). On the contrary, the same analysis conducted on the expression of cellular genes did not show any discriminating power (Figures 4B and 5B).· Phase 4: Analysis of HERV Loci Expression in MS Patients and Healthy Controls
[0057] The discriminating power of HERVs for the MS condition was initially observed by considering the global expression of the HERV loci (as depicted in Figure 5A for the 400 HERV loci with the highest expression mean in the analyzed samples). The same results were also obtained by conducting the analysis based on the highest variance and standard deviation, as well as reducing the number of HERV loci considered to a minimum discriminant set of 140 sequences (Table 1). These sequences, although overall able to divide healthy controls from MS patients, do not, however, have a statistically significant variation in expression between the two groups in all cases. Table 1. Description of the set of 140 HERV loci discriminating for MS ID HERV Coordinates hg38 (chromosome: beginning-end) Bases St. Group Avera ge TPM (C) Averag e TPM (SM) 11143:129168485-1291718343350-ERRANTI-like632.64729.4360951:169683482-1696913017820+HERVH174.03189.55471319:36149712-3616102311312-HERVH102.64114.80472019:38823108-3883743314326+HERV936.6036.60365611:121632566-12164349110926-HERVH28.6125.7761711:207632285-2076412528968-HML223.4826.54479619:58305729-583151169388+HML619.9625.2824537:30572445-305796577213-HERV424.6329.4360621:150628159-1506357767618-HML218.9419.00418414:73702886-7371616713282+HERVH10.1510.1216584:153690317-1536939203604-HML328.6438.76377612:42440799-4245151110713+HML58.7711.68484920:49281128-492873436216-HERVIP17.3417.0660691:155626674-1556357039030-HML28.6712.96437816:35797699-358045236825+HML513.9116.32433116:2660502-26699859484-HML47.8110.29392713:19626060-196359079848+HML37.808.36461819:11942587-119489856399-HERVE10.5011.3225217:64990455-649999369482-HERV36.817.775102:37124923-3713762712705+HERVFA3.497.88418514:75577964-755876489685-HERVH6.037.546992:142929044-1429372918248-HERVH488.867.4318845:76881216-768895608345-HERVH8.265.075108Y:13282562-132896057044+HARLEQUIN6.747.53477819:53433375-534431109736-HERV35.104.9826377:135174033-13518446610434+HERVIP4.125.1010453:101692114-1017012539140+HML24.755.66331710:89406288-8941747111184+HERV92.616.62628322:39052278-3906627013993-HERVH2.763.80351111:60215182-6022550910328+HERVL3.834.998643:9841750-985471012961-HML23.183.4423846:158611703-1586211789476-HERVFA4.704.3760741:156180078-1561878977820-HML56.025.73432915:101072781-1010793176537+HERVH6.496.17370112:9959546-99686569111-HERVH3.646.2320305:160116064-16013113815075+MLT2.472.555429X:63426518-634346938176-LTR464.174.23475819:51976955-5199222815274+HML62.292.16626222:16611312-166167825471+HERVH6.377.16435316:21629614-216375257912-HERV94.504.14354711:74869434-7488448715054+HUERSP32.261.9662281:235670261-2356755505290+HERVIP6.635.256732:127369875-1273778537979+HML55.572.8924767:43853008-4386675213745-HML31.992.27479519:57896518-579038577340+HML13.934.15394913:40874210-408811616952-HERVE4.314.05449017:80537010-8055202015011-HML21.861.9225347:77339937-773458805944+HERVH6.503.81390212:111816238-1118256469409-HML64.222.3823826:158156792-15816936012569-unclassifiable2.182.306242:97821110-978294568347+HERV93.733.0418925:82267546-822737066161-HARLEQUIN5.103.96476619:52649616-526531163501+HERVIP9.746.4427768:41590933-415983267394-HERVE4.113.2331079:92071888-920783986511+HERVH4.853.505702:71388088-713955577470+HML53.523.44448517:75252632-752582975666+HERVH4.974.2821526:57069287-5708099611710+HML52.202.1560721:155650288-1556596319344-HERV42.412.84479019:57487177-5750081313637+unclassifiable1.561.99433216:3079068-30813182251-HML38.7312.525582:64252675-642585655891+HERVH3.464.51322210:35403060-354101207061-HERVH5.232.54368712:8281117-829158210466-HERVE2.142.38474419:47047398-470553978000-HERVH3.562.60433516:8837971-88454727502-HERVH3.503.00340311:10004937-1002070615770+MER501.821.17405013:99488624-994967528129-HERV92.073.26435216:20674233-206819957763+HERVI2.513.0618855:76909228-769129763749+HERV96.785.35400313:76983510-7699363010121+MER42.052.25340911:17370180-173792939114+HUERSP32.482.319093:32458604-324677149111-HERVH2.882.122p16.22:53756639-537614264788+HERVW3.155.6531709:133031013-1330373856373-HERVIP3.193.20337010:120844913-1208515806668+HERVH2.453.4314954:88442633-8845388111249+HERV41.112.589343:44329728-4434087811151-MER841.641.9125187:64679995-646865616567+HERVH2.603.4217455:1577341-15853898049-HML31.463.60442617:11971744-119781026359+HERVH2.353.11461119:9626252-963808211831+HERVIP1.531.417p14.27:35696138-35697017880+HERVW17.8620.72421514:92622417-926296087192-HERVIP3.061.8929848:145055669-14506669611028+HML41.411.53461319:9741762-97517489987+HERVH2.111.39476519:52620842-526273026461+HERVH2.622.385047Y:7377455-738816710713-HERV96.540.02320310:20118010-201249346925-HERVH3.341.71360011:94650031-946578307800+HERV41.772.02343611:34897254-349070219768+HUERSP31.901.42473619:43321084-433293938310-HERVH1.053.54390112:110403279-1104110687790-HERVH1.791.8829098:97175022-971807535732+HERVH2.902.338573:5140823-51498188996-HERVADP2.041.1411123:128919296-1289292319936+HML11.611.1312783:193599956-19361333313378-HEPSI11.350.76476719:52627908-526361648257+HERVH2.081.2762451:246634727-2466428038077-HERVH481.521.57415314:53128745-531355076763+HERVH2.901.2661231:183613210-1836224439234+HERVH1.251.31461019:9634663-96416817019-HERVH1.701.60414714:50893357-509010347678+HERVH1.961.244932:26297921-263064018481+HERVI1.041.538382:233213615-2332178294215-HERVH3.492.17339711:7513515-75204746960-HERVH1.641.54395513:48056441-480622975857+HERVH2.191.70338711:4454469-446477610308+HERVIP1.280,919773:72023117-720289175801+HERVH2.901,31467819:23340355-2335393813584+HERVH0.910,679333:44340864-443503939530-HERVH1.121,07448817:77167827-771752017375-HERVIP2.060.9314074:54020048-540255595512+HERVH2.261.5021246:35561857-355678235967-HERVIP2.091.35477719:53469501-5348221212712-HERV30.720.79394313:36316325-363212644940+HERVH1.782.0360561:143764840-14377833113492-MER660.570.81477019:52898515-529054346920-HERVE1.231.4762151:228947563-2289518914329+HERVH2.162.1558611:37907865-379140946230+HERVH2.111.16476319:52408081-524147336653-HML61.261.50490621:42916753-429257599007-HERVH481.370.8510803:119978070-11999030912240+LTR191.090.558362:231091066-23110140010335-HML51.150.7216384:139442392-1394498177426+HERVL1.261.23408614:24009540-240157826243-HML21.201.65459319:5158650-51657417092-HERV91.591.045462:58113432-581191305699-HERVH1.631.5260791:160690812-1606999879176+HML20.751.2226777:152027505-1520359458441-HERV91.220.9229168:102979373-10299171812346-HEPSI10.880.602q22.22:142898679-1429078999221-HERVW1.200.7825747:97948903-9796034311441-HERVE0.840.714662:6838342-68468838542-HUERSP30.371.92476019:52300607-523079777371-HML31.441.0024497:29636375-296462319857-HERVT0.760.9431239:100139459-1001465447086-HERVH1.660.882q11.22:96217144-962233256182-HERVW1.581.1920736:11103051-111121919141-HERVFRD1.380.62421314:92041755-920498638109-HERVH1.001.02ID HERV from Vargiu et al. 2016 and Grandi et al. 2016 (HERV-W). St = strand, TPM = transcripts per million kilobases, C = healthy controls, MS = patients with multiple sclerosis · Phase 5: Identification and selection of a set of deHERV loci having discriminating power for MS
[0058] Although the set of 140 HERV loci described above (Table 1) was endowed with discriminating power based on the global expression in the samples being analyzed, this number of sequences is however quite high, making the potential for variability in the population large. Furthermore, the sequences are selected based on the mean of expression, and not by a significant difference of the latter in patients with MS. Then, in order to select the HERV sequences used for discrimination, a differential expression analysis was performed verifying which HERV loci contained in the HERVdb - beyond their expression values - showed a statistically significant variation between healthy controls and MS patients. Differential expression analysis was performed on the RStudio platform using the DEseq2 package (Love, Huber, and Anders 2014), which estimates variance from raw counts and evaluates differential expression using a negative binomial distribution model. In particular, the variation of expression of the HERV loci between healthy controls and MS patients was measured on a logarithmic scale (log2 fold changes: for example, a log2 fold change of 1.5 corresponds to an increase in expression of a multiplication factor of 2 1.5< , then to ≈ 2.82 times). HERV loci were then identified as being differentially expressed whose variation between controls and patients reached a threshold of statistical significance, represented by an adjusted p-value (p-adj, according to Benjamini-Hochberg) ≤ 0.01 (the use of the p-adj with respect to the unadjusted p-value decreases the false discovery rate value to 1%) and by a variation of expression between healthy individuals and individuals with MS measured in terms of log2 fold change> 1 (in case of expression up-regulation) or <-1 (in case of expression down-regulation).
[0059] Differential expression analysis allowed the identification of 43 differentially expressed HERV loci (deHERV) located on autosomal chromosomes that were evaluated for their discriminating power based on both mean expression and variance, confirming the ability to clearly divide patients with MS from healthy controls (Figure 7). The deHERV loci identified are shown in Table 2 together with the significance level of differential expression in MS (padj) and the TPM values (Transcribed per Million Kilobases, is a quantitative measure of the expression level) mean in healthy controls and patients with SM. The sequences of the HERVs indicated in the table are understood to be incorporated thanks to the reference, present in the table itself, of their exact genomic position, which also specifies the beginning and the end of the sequence on the reference chromosome. The deHERV loci have levels of statistical significance between 7:08 -14< and 0.008, and most of the loci is down-regulated in patients with MS (healthy controls: TPM values between 5.6 and 0.1, average = 1.3, median = 0.95 ; MS patients: TPM values between 2.9 and 0.04, mean = 0.7, median = 0.5) (Table 1). The deHERV loci considered for their discriminating power have an average length of about 8000 nucleotides (from a minimum of about 3350 nucleotides to a maximum of about 14000 nucleotides, median = about 7600 nucleotides) (Table 2).
[0060] Finally, a small number of deHERV loci with the best discriminating power were selected, identifying as the final subset a total of 17 deHERV loci ordered on the basis of decreasing mean and variance values of expression, in order to select the most significant ones (Figure 8). In both cases, the 17 loci (13 of which shared between the two analyzes based on mean and variance of expression) make it possible to clearly distinguish MS patients from healthy controls (Figure 8, Table 2). The loci deHERV included in this subset have minimum levels of statistical significance between 7:08 -14< and 0.005, and most of the loci is down-regulated in patients with MS (healthy controls: TPM values between 5.6 and 0.5, average = 2.1, median = 1.7; MS patients: TPM values between 2.9 and 0.4, mean = 1.16, median = 0.9) (Table 3). Table 2. Description of the set of deHERV loci discriminating for MS N° HER V ID Coordinates hg38 (chromosome: beginning-end) st. Bases Group DE in SM Log2 Fold Chan ge Padj Averag e TPM (C) Averag e TPM (SM) 16732:127369875-127377853+7979HML5DOWN-1.137.08E-145.572.892484120:41594586-41602330+7745HERVIPDOWN-1.755.52E-091.210.423448817:77167827-77175201-7375HERVIPDOWN-1.362.47E-082.060.9342q132:112039346-112044989-5644HERVWUP1.862.68E-080.451.8458573:5140823-5149818-8996HERVAD PDOWN-1.011.09E-072.041.14620736:11103051-11112191-9141HERVFR DDOWN-1.311.27E-071.380.62710803:119978070-119990309+12240LTR19DOWN-1.157.11E-071.090.5589773:72023117-72028917+5801HERVHDOWN-1.333.41E-062.901.3197502:174202252-174216415+14164HUERSP 2DOWN-1.202.24E-050.590.2910415314:53128745-53135507+6763HERVHDOWN-1.383.15E-052.901.261112783:193599956-193613333-13378HEPSI1DOWN-1.023.68E-051.350.76126q21 a6:106228137-106235814+7678HERVWDOWN-1.234.84E-050.870.4213322210:35403060-35410120-7061HERVHDOWN-1.195.39E-055.232.541458831:48437944-48446631+8688HERV9DOWN-1.828.25E-050.470.141531679:131680373-131686728+6356HML3DOWN-1.140.00011.390.72167152:152701399-152709042+7644HERVIPDOWN-1.020.00011.280.721758611:37907865-37914094+6230HERVHDOWN-1.020.00032.111.161831239:100139459-100146544-7086HERVHDOWN-1.080.00031.660.881929828:144858915-144868670-9756LTR46DOWN-1.100.00040.950.5020320310:20118010-20124934-6925HERVHDOWN-1.110.00043.341.712121966:70436656-70448522-11867HML3DOWN-1.220.00040.540.272219925:131571798-131582180+10383HERVHDOWN-1.120.00040.840.432323076:118617102-118626810+9709HERV9DOWN-1.230.00050.890.4324367812:1711131-1715463-4333HERV9DOWN-1.200.00051.440.702510813:120038402-120045736+7335HERVIPDOWN-1.260.00050.870.4126477919:53831404-53845084+13681HERVHDOWN-2.790.00060.240.042711103:128828222-128834858-6637HERVHDOWN-1.200.00071.280.64287492:172542091-172548758+6668HERVHDOWN-1.790.00090.430.142920145:147868278-147874526-6249HERVHUP2.620.00090.110.7330374112:26777173-26782083+4911HERVLDOWN-1.550.00090.820.31315502:62088710-62100497-11788MER4DOWN-1.210.00140.400.203220596:2310409-2317792-7384HarlequinDOWN-1.270.00170.610.293317565:10716140-10724937+8798HERVLDOWN-1.350.00170.440.203412864:3046337-3056570-10234MER57DOWN-1.740.00210.300.103521266:38616529-38624444+7916HERV9DOWN-1.220.00220.610.3036409714:30965849-30972548+6700HERVIPDOWN-1.100.00251.090.57377912:203960014-203967570-7557HERVHDOWN-2.060.00280.310.083811093:128836995-128842027-5033HERVHDOWN-1.250.00370.810.3839416514:59610765-59618496+7732HUERSP 1DOWN-1.080.00421.170.614059721:94445506-94448856+3351HML2DOWN-1.280.00641.060.4941455518:46868452-46877781+9330HERVIPDOWN-1.410.00680.380.1642368812:8319624-8324920-5297HERVTDOWN-1.030.00741.210.694327218:8171833-8181319+9487HERVEDOWN-1.380.00800.300.13HERV ID from Vargiu et al. 2016 and Grandi et al. 2016 (HERV-W). St = strand, DE = differentially expressed (down-regulated or up-regulated in SM), padj = adjusted p-value (significance threshold = 0.01, the change in expression of the HERV loci was measured in log2 fold changes), TPM = transcripts per million kilobases, C = healthy controls, MS = patients with multiple sclerosis Table 3. Description of the selected subset of deHERV loci with discriminating power for MS N° HER V ID Coordinates hg38 (chromosome: beginning-end) St Bases Grou p DE in SM Log2 Fold Chan ge Padj Mean TPM (HC) Mean TPM (SM) SEQ ID No. 16732:127369875-127377853+7979HML5DOW N-1.137.08E-145.572.89SEQ ID No.1 3448817:77167827-77175201-7375HER VIPDOW N-1.362.47E-082.060.93SEQ ID No.2 42q132:112039346-112044989-5644HER VWUP1.862.68E-080.451.84SEQ ID No.3 58573:5140823-5149818-8996HER VADPDOW N-1.011.09E-072.041.14SEQ ID No.4 620736:11103051-11112191-9141HER VFRDDOW N-1.311.27E-071.380.62SEQ ID No.5 710803:119978070-119990309+12240LTR1 9DOW N-1.157.11E-071.090.55SEQ ID No.6 89773:72023117-72028917+5801HER VHDOW N-1.333.41E-062.901.31SEQ ID No.7 10415314:53128745-53135507+6763HER VHDOW N-1.383.15E-052.901.26SEQ ID No.8 1112783:193599956-193613333-13378HEPS I1DOW N-1.023.68E-051.350.76SEQ ID No.9 13322210:35403060-35410120-7061HER VHDOW N-1.195.39E-055.232.54SEQ ID No.10 1758611:37907865-37914094+6230HER VHDOW N-1.020.00032.111.16SEQ ID No.11 1831239:100139459-100146544-7086HER VHDOW N-1.080.00031.660.88SEQ ID No.12 20320310:20118010-20124934-6925HER VHDOW N-1.110.00043.341.71SEQ ID No.13 16 (m)7152:152701399-152709042+7644HER VIPDOW N-1.020.00011.280.72SEQ ID No.14 2(v)484120:41594586-41602330+7745HER VIPDOW N-1.755.52E-091.210.42SEQ ID No.15 19 (m)29828:144858915-144868670-9756LTR4 6DOW N-1.10.00040.950.50SEQ ID No.16 23 (v)23076:118617102-118626810+9709HER V9DOW N-1.230.00050.890.43SEQ ID No.17 The loci with the words (m) and (v) have been identified on the basis of only the mean and variance of expression, respectively, while the others are common to the two selections. HERV ID from Vargiu et al. 2016 and Grandi et al. 2016 (HERV-W). St = strand, DE = differentially expressed (down-regulated or up-regulated in SM), padj = adjusted p-value (significance threshold = 0.01, the change in expression of the HERV loci was measured in log2 fold changes), TPM = transcripts per million kilobases, C = healthy controls, MS = patients with multiple sclerosis
[0061] For this minimum discriminant subset of 17 deHERV loci, specific primer pairs were designed, suitable for verifying their expression levels by Real-Time PCR technique. This analysis confirmed the changes in expression detected through the bioinformatic analysis of the data obtained from the sequencing.
[0062] The analysis described above was developed starting from peripheral blood monocytes, but it can be used on RNA derived from any sample obtained from a subject, by way of non-limiting example RNA derived from tissues and body fluids such as blood, plasma, biopsy, muscle biopsy, spinal fluid, cerebrospinal fluid, saliva, oral mucosal cells.
[0063] The analysis and comparison of the expression levels can be carried out as described above or by means of techniques known to those skilled in the art that are currently available or that can be developed in the future; by way of non-limiting example, this analysis can be performed by RT-PCR, Real Time PCR, microarray, Northern Blotting, in situ hybridization, molecular beacons, forced intercalation probes and other RNA detection techniques known to those skilled in the art.
[0064] The method described above and the choice of data analysis deriving from sequencing of total messenger RNA with this method is particularly preferred as it is able to evaluate the expression of thousands of sequences with a high degree of structural homology, which makes them highly similar to each other. The preferred method is based on the mapping of all the transcripts to the sequence of the human genome, and the subsequent univocal identification based on their precise chromosomal position. In this way, "doubtful" transcripts as they map to several positions are excluded in the preliminary phase, and the subsequent quantifications concern only the sequences actually and univocally deriving from the HERV taken into consideration. The use of more traditional approaches is however possible by providing for the necessary passage, within the reach of experts, of designing specific probes or pairs of primers for each HERV sequence.
[0065] The object of the present invention is therefore a method for the in vitro diagnosis of multiple sclerosis comprising an analysis step of the expression levels of the HERV of the HERV of sequence SEQ ID No. 1 in a biological sample obtained from a subject and a step comparing said levels with the levels found in samples obtained from healthy individuals, in which decreased levels of expression by minus one of said transcripts indicate the presence of multiple sclerosis in the subject.
[0066] In particular, the subject of the present invention is an in vitro method for the diagnosis of multiple sclerosis in an individual, the method comprising the following basic steps: to isolate ex vivo an RNA sample from a biological sample obtained from a subject: determine the expression levels of the transcript of the HERV of sequence SEQ ID No. 1, in which a decrease of the expression levels of said HERVs in the sample under examination, compared to the levels of the corresponding product in a control sample, is indicative of the presence of multiple sclerosis in the subject.
[0067] The object of the present invention is therefore a method for thediagnosis in vitro of multiple sclerosis in an individual in which the levels of expression analyzed is the HERV transcript SEQ ID No. 1.
[0068] The expression levels decreased by the HERV transcript of SEQ ID No 1 in the sample under examination compared to the levels detected in the sample obtained from a healthy individual are indicative of the presence of multiple sclerosis in the subject.
[0069] The object of the present invention is therefore a human endogenous retroviral sequence (HERV) selected and validated in a specific way among the tens of thousands of similar sequences present in the genome which has shown early diagnostic power for the pathology Multiple Sclerosis ( SM). More in detail, the association with the presence of multiple sclerosis is represented by the overall pattern of expression of the set of sequences indicated which, taken together, showed a significantly higher diagnostic-predictive power than the evaluation of a single sequence or a subset of them for this reason, also in order to reach the most sensitive method in terms of diagnostic reliability, it is preferable to consider for the purposes of the invention this set of specific sequences as a whole, meaning it as a single panel of biological indicators, rather than patenting each of the single sequences, since the diagnostic power observed on the basis of the expression of the 43 HERV sequences according to table 2 as a whole, and preferably on the basis of the minimum discriminant set of the 17 sequences reported in table 3, is undoubtedly more accurate than that which is would have considered the expression individually the single sequences.
[0070] Thus, the present invention relates to the use of the HERV of SEQ ID No 1 as biomarker in the in vitro diagnosis of multiple sclerosis.
[0071] With the findings of the invention, therefore, the need to identify markers strongly and uniquely correlated to the presence of MS pathology and the need for an analysis method capable of overcoming the technical problems related to the methods commonly used for the analysis of sequences are satisfied.
Claims
1. An in vitro method for diagnosing multiple sclerosis in an individual, the method comprising the following basic steps: - isolating an RNA sample ex vivo from a biological sample obtained from a subject; - determining the expression levels of the human endogenous retrovirus (HERV) transcripts of sequence SEQ ID No. 1; - compare these levels of expression with the levels of expression of the same HERV detected in a sample of RNA isolated ex vivo from a sample obtained from a healthy individual in which the decreased levels of expression of said HERV in the test sample compared to the levels detected in the sample obtained from a healthy individual are indicative of the presence of multiple sclerosis in the subject.
2. Method according to claim 1 wherein the analysis to determine the expression levels is carried out by deep sequencing, RT-PCR, Real Time PCR, microarray, Northern blotting, in situ hybridization, molecular beacons, or other forced intercalation probes RNA detection techniques, preferably RNAseq.
3. Method for the in vitro diagnosis of multiple sclerosis comprising the following basic steps: - deep sequencing of polyadenylated RNA, isolated ex vivo, from a biological sample of a subject to generate sequencing reads, - bioinformatic analysis carried out through a system of computer comprising: I. To align the sequencing reads obtained with a reference human genome; II. Determine the relative quantities of the sequence in the sample corresponding to the HERV specific transcripts of SEQ ID 1 ; - compare said relative quantities with the relative quantities detected with said passages in a sample of polyadenylated RNA isolated ex vivo from a sample of a healthy individual in which the decreased expression levels of said HERV in the sample under examination compared to the levels detected in the sample obtained from a healthy individual are indicative of the presence of multiple sclerosis in the subject.
4. Method according to the preceding claims, wherein said samples are RNA samples derived from in-vitro tissues and body fluids, preferably blood, plasma, biopsy, muscle biopsy, spinal fluid, cerebrospinal fluid, saliva, oral mucosa cells.
5. Use of HERV of SEQ ID No 1 as biomarker in in vitro diagnosis of multiple sclerosis.
6. Use of the HERV of SEQ ID No 1 according to the preceding claim in which expression levels decreased of the HERV transcripts of SEQ ID No 1 compared to the levels detected in samples obtained from a healthy individual are indicative of the presence of multiple sclerosis in the subject.