METHOD FOR CHARACTERIZING A TUMOR

The method analyzes nucleosides from RNA and metabolites using mass spectrometry and machine learning to objectively characterize gliomas, addressing the limitations of current diagnostic methods by accurately distinguishing tumor grades and predicting patient outcomes.

FR3125882B1Active Publication Date: 2025-07-18CENTRE HOSPITALIER UNIVERSITAIRE DE MONTPELLIER +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2021008279
Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-07-29
Publication Date
2025-07-18
Estimated Expiration
2041-07-29

AI Technical Summary

Technical Problem

Current diagnostic methods for gliomas, particularly in distinguishing between grades II and III glial tumors, are subjective, costly, time-consuming, and lack sufficient biomarkers for guiding therapeutic decisions, necessitating an objective, precise, and reproducible method for tumor characterization.

Method used

A method involving the quantitative analysis of modified and unmodified nucleosides derived from total cellular RNA, extracellular RNA, and metabolites using mass spectrometry, combined with supervised machine learning, to create a nucleoside profile for predicting tumor grade and patient survival.

Benefits of technology

Enables accurate differentiation between glioma grades II and III, and predicts patient survival status with high precision, providing a robust and efficient diagnostic tool for gliomas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000031_0000
    Figure 00000031_0000
  • Figure 00000031_0001
    Figure 00000031_0001
  • Figure 00000032_0000
    Figure 00000032_0000
Patent Text Reader

Abstract

The present invention relates to an in vitro method for characterizing a tumor, based on the quantitative analysis of modified and unmodified nucleosides derived from total cellular RNA, extracellular RNA and / or isolated nucleosides extracted from a biological sample. More particularly, the invention relates to a method for predicting the grade of a glial tumor. The present invention is therefore in the fields of cancerology and molecular biology, more particularly applied to medical diagnosis. Figure to be published with the abstract: Fig. 1
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: METHOD FOR CHARACTERIZING A TUMOR Field of invention

[0001] The present invention relates to an in vitro method for characterizing a tumor, based on the quantitative analysis of modified and unmodified nucleosides isolated from a biological sample. More particularly, the invention relates to a method for predicting the grade of a glial tumor.

[0002] The present invention is therefore situated in the fields of cancerology and molecular biology, more particularly applied to medical diagnosis. State of the art

[0003] The characterization of a tumor is an essential prerequisite for choosing the most appropriate treatment for the patient. By "characterization of a tumor" is meant the characterization of the status or degree of evolution of a given tumor, it may involve, for example, the evaluation of the degree of evolution of a tumor of a known tissue, the attribution to a tumor of a previously defined grade, or any other characterization such as in particular the determination of the initial or metastatic character of a tumor.

[0004] Gliomas, or glial tumors, are the most common tumors of the central nervous system, characterized by significant variability in age of onset, classification, histological features, and ability to progress and eventually metastasize.

[0005] Gliomas are classified according to their morphology and degree of malignancy. The World Health Organization (WHO) consensus classification assigns a malignancy grade of I to IV to gliomas, with glioblastomas, or grade IV tumors, being the most aggressive and deadly form.

[0006] One of the main limitations in the management of gliomas and glioblastomas is related to the current lack of effective diagnostic strategies. The selection of personalized treatment requires accurate classification of tumors. Currently, the main diagnostic methods used clinically for the detection of gliomas are based on neurological tests and neuroimaging methods, performed when the disease is already at an advanced stage.

[0007] Tumor diagnosis requires analysis of the patient's tissues from a biopsy or surgical resection. From this sample, several molecular analyses are performed: candidate gene expression test, DNA copy number counting, methylation profile, phospho- pathway profiling protein and genetic sequencing. However, biopsy-based diagnostics have limitations regarding tumor grading and patient stratification. Indeed, with regard to glial tumors, for example, glioma grades are difficult to distinguish, especially grades II and III. Grading requires delicate pathological analysis, often performed independently by two specialists. Grade II designates a benign tumor, while grade III represents a transition to glioblastoma multiforme, which is the most aggressive condition.

[0008] Purely histological classifications are difficult to reproduce; they are based on visual expertise and require the intervention of two specialists. Anatomopathological coupling with image analysis by magnetic resonance imaging (MRI) is costly and time-consuming; it depends in particular on the time it takes to access the MRI. Currently, no biomarker is sufficient on its own to guide anti-cancer therapeutic decisions.

[0009] There is therefore a general need for an in vitro method for characterizing a tumor, said method being objective, precise, reproducible, easy and feasible at a stage if possible early in the disease. Said method would strengthen the diagnosis and facilitate the stratification of patients.

[0010] Janzer's publication ("Neuropathology and molecular pathology of gliomas" RC Janzer, Rev. Med. Suisse, 5, 1501-4, 2009) describes the WHO classification of gliomas, based on histological and immunohistochemical criteria, as well as on genetic profiles highlighting the alteration of the DNA of the cells: the determination of hypermethylation of the MGMT gene promoter (for glioblastomas) and the detection of losses of chromosomes Ip and 19q (for oligodendroglial tumors).

[0011] The publication by Relier et al (“FTO-mediated cytoplasmic m6 Am demethylation adjusts stem-like properties in colorectal cancer ce U.”, Nat. Commun 12, 1716, 2021) describes the regulation in the cytoplasm of the level of m6Am methylation by FTO (Fat mass and obesity associated protein) in cancer stem cell lines. The authors highlight the biological function of the m6Am modification and its potential side effects for the monitoring of colorectal cancer. This document mentions a step of mass spectrometry analysis (LC-MS / MS) of fragmented mRNA. Only the nucleosides m6A, A, m6Am and Am are detected and quantified.

[0012] International application WO 2007 / 008647 “Diagnosis and classification of gliomas using a proteomic approach” relates to a method for diagnosing and classifying gliomas using a proteomic approach. In this method, Tumor tissue is analyzed by mass spectrometry and a profile of expressed proteins is obtained.

[0013] There is therefore a particular need for an in vitro method for evaluating the degree of malignancy of a glial tumor, in particular its classification. In particular, there is a need for an objective method for distinguishing between grade II and grade III glial tumors, to strengthen diagnosis and facilitate patient stratification. Disclosure of the invention

[0014] The inventors have now developed a method for characterizing a tumor which exploits the quantitative data of the epitranscriptome.

[0015] The epitranscriptome encompasses all the chemical modifications carried by the bases of ribonucleic acids (RNA), a set which is also referred to as "RNA epigenetics". A method according to the invention comprises providing a biological sample from a subject suffering from a tumor and obtaining the quantities of modified and unmodified nucleosides from said sample, said quantities being grouped in a vector (in the mathematical sense of the term). According to a particular aspect, a method according to the invention comprises the subsequent computer analysis of said vector for the characterization of a tumor. Said characterization of a tumor makes it possible to predict clinical and medical information on the tumor from the sample subject to analysis. More particularly, a method according to the invention comprises the computer analysis of said vector for the prediction of the grade of said tumor.

[0016] For simplicity, for a given sample, we will call “epitranscriptomic profile” or simply “profile”, the vector which groups together the quantities of each nucleoside, modified or not modified.

[0017] The modified and unmodified nucleosides are derived from: i) total RNA extracted from cells of a biological sample of a patient, ii) extracellular RNA from a biological sample of a patient, and / or iii) an extract of metabolites from a biological sample isolated from a patient.

[0018] Nucleosides derived from total RNA extracted from cells of a biological sample of a patient and / or from extracellular RNA from a biological sample of a patient are obtained by fragmenting the RNA into nucleotides and then dephosphorylating them. Nucleosides derived from an extract of metabolites from a biological sample isolated from a patient are obtained by extracting the metabolites from a biological sample and then dephosphorylating said metabolites, according to appropriate methods well known to those skilled in the art. Said metabolites are in particular derived from the catabolism of RNA, the nucleosides present in the form monomeric nucleosides can also be referred to as so-called “free” nucleosides.

[0019] More particularly, the modified and unmodified nucleosides are the nucleosides present in the total RNA extracted from cells of a biopsy of said tumor.

[0020] By "nucleosides" is meant glycosamines consisting of a nucleic base linked to the anomeric carbon atom of a pentose residue by a glycosidic bond from the nitrogen atom N1 of a pyrimidine or the atom N9 of a purine. According to a particular aspect of a method according to the invention, when said pentose is ribose, the term "nucleosides" in this case means ribonucleosides.

[0021] Modified RNA nucleosides are also referred to as "epitranscriptomic marks" or "epitranscriptomic modifications". In addition to the usual RNA nucleosides (Table 1) which can be used for the characterization of tumors, modified nucleosides which can be used in a method according to the invention, in particular in the analysis of gliomas, are listed (Table 2).

[0022] [Tables 1] N ucKosiie D en oœma Osa chemical Adenoid C cytidise G guanosm U uridùte

[0023] [Tables2] BornJeosie No ml A iE66A m66Am :m6A w6Am ac4C Cm hm5C ai3-G m5C Gm ml G tll2i > G axo8G T ,1. Psi Q nGUm mcm5s2U Dice <Hnis&t»a chimHpe 2'4&-medi^adenosis the -networks of GSie N6rN6-dimethylM^exiosm. We^ÀZ-He-has^eth^adenosis N6-meîhyfadsK>sfae N6;2!-O-dime&yhdesi&8me N4-acehicytidme ' 2*-O-methylcytKJ iae 5^drox}'methyfcytîds3é 3-me&ylc ytidine 5 -methy Icytidine 2!-O-nie&ylguam&me 1 -inethyguaEosse N2SN 2,7-$mEethylguaED8he N2J-<&nethyiguanQsàie 7 -iæfeylguaosÈie S -h hydroxyguaosine inosiaé pseudonym Queuséie 3,2LO-dsEethytheideia 5 --jne&oxycarbonytaethyl - 2 - No troops A A At A; A A C C C c c G G G G G G I P Q u U

[0024] According to the embodiments of a method according to the invention, an epitranscriptomic profile may include, according to the needs of the application, a greater number of modified nucleosides, to be determined among the known nucleosides (Jonkhout et al, “The RNA modification landscape in human disease”, RNA, Dec; 23 (12): 1754-1769, 2017). A complete list of all modified nucleosides, which may, according to the needs of the analysis, be included in the transcriptomic profiles is publicly accessible.

[0025] The epitranscriptomic profile of a sample characterizes said sample. Said epitranscriptomic profile can be obtained by any technique known in the state of the art, and in particular by mass spectrometry, in particular by mass spectrometry coupled with chromatography.

[0026] To designate the medical information to be predicted, we will also speak of “clinical variables” or “clinical characteristics”.

[0027] In a method according to the invention, the step of analyzing the epitranscriptomic profile for clinical prediction purposes is based on a supervised machine learning method. The learning is carried out on the profiles from a cohort, i.e. cell samples for which the clinical characteristic variable to be predicted is known in advance. The “computer model” thus created by the learning can then be used (in prediction mode) to predict the clinical variable for any new sample.

[0028] The inventors have also developed a method for normalizing raw quantitative data to obtain an epitranscriptomic profile containing comparable relative quantities.

[0029] Prior to the learning method, the exploratory analysis of the profiles of the samples of the cohort revealed variations in the profiles, said variations being correlated with the grade of the tumors of the subjects from which the samples were extracted. From the profiles of the signature cohort and by means of a machine learning prediction tool, a method according to the invention makes it possible to predict the grade of a glioma from a biological sample of a patient suffering from a tumor, in particular from a sample comprising tumor cells. More particularly, a method according to the invention makes it possible to distinguish grades II and III of a glioma from a sample comprising tumor cells.

[0030] Finally, by combining the standardized epitranscriptomic profiles and patient survival data, and by means of a machine learning prediction tool, a method according to the invention makes it possible to predict the survival of a patient from a biological sample isolated from said patient, in particular from a tumor sample. Detailed description of the invention

[0031] According to a first aspect, the invention relates to an in vitro method for characterizing a tumor of an individual, from a biological sample isolated from this individual, comprising the steps of: a. isolation of nucleosides from said biological sample, by extraction of: i) total cellular RNA and its fragmentation into nucleosides, ii) extracellular RNA and its fragmentation into nucleosides, and / or iii) nucleosides from monomeric catabolites, b. isolation and determination of a respective quantity of at least 3, preferably at least 5, preferably at least 10, preferably at least 20, different nucleosides obtained during step a), and c. establishing, for said biological sample, a nucleoside profile from the respective quantities of each of the nucleosides obtained during step b), said profile being characteristic of said tumor.

[0032] According to a first embodiment, a method according to the invention is based on the simultaneous analysis of the quantity of different nucleosides derived from the total cellular RNA of a biological sample, and / or derived from extracellular RNA and its fragmentation into nucleosides and / or derived from nucleosides obtained from the monomeric catabolites present in said sample, a method according to the invention therefore comprises the simultaneous analysis of multiple variables, and not on the quantitative detection of a single marker.

[0033] By “profile” or “nucleoside profile” is meant a vector of quantities of nucleosides.

[0034] By "total cellular RNA" is meant all cellular RNA extracted according to well-known and accessible methods. Total cellular RNA includes transfer RNA (tRNA), messenger RNA (mRNA), ribosomal RNA (rRNA) and other non-coding RNAs. Said total cellular RNA is therefore present here in a polymeric form.

[0035] By "extracellular RNA" is meant all of the extracellular RNA present in polymeric form, extracted according to well-known and accessible methods. This polymeric form of extracellular RNA is also notably designated by the expression "circulating RNA". Said extracellular RNA is derived from the in vivo enzymatic degradation of transport RNA (tRNA), messenger RNA (mRNA) and / or ribosomal RNA (rRNA) and other types of RNA, notably non-coding RNA.

[0036] By "nucleosides derived from monomeric catabolites" is meant the nucleosides obtained, according to well-known and accessible methods, from the catabolites present in a monomeric form in the sample. These catabolites Monomeric RNAs are derived from the in vivo enzymatic degradation of transport RNA (tRNA), messenger RNA (mRNA) and / or ribosomal RNA (rRNA) and other types of RNA, including non-coding RNAs.

[0037] By "isolation and determination of a respective quantity of at least 3 different nucleosides" is meant the isolation and determination of a quantity of each of the "at least 3" nucleosides taken individually.

[0038] According to this first embodiment, the invention therefore relates to an in vitro method for characterizing a tumor of an individual, from a biological sample isolated from this individual, comprising the steps of: a. isolation of nucleosides from said biological sample by extraction of total cellular RNA and its fragmentation into nucleosides, b. isolation and determination of a respective quantity of at least 3, preferably at least 5, preferably at least 10, preferably at least 20, different nucleosides obtained during step a), and c. establishing, for said biological sample, a nucleoside profile from the respective quantities of each of the nucleosides obtained during step b), said profile being characteristic of said tumor.

[0039] Nucleosides may also be present in the biological sample in an extracellular polymeric form, in particular also referred to as "circulating RNA". Nucleosides may also be present in a monomeric form (metabolites) in the biological sample. Said extracellular RNAs and monomeric nucleosides are derived from the in vivo enzymatic degradation of transport RNA (tRNA), messenger RNA (mRNA) and / or ribosomal RNA (rRNA) and other types of RNA, in particular non-coding RNAs.

[0040] According to a second embodiment, a method according to the invention is based on the simultaneous analysis of the quantity of different nucleosides derived from the extracellular RNA of a biological sample.

[0041] According to this second embodiment, the subject of the invention is an in vitro method for characterizing a tumor of an individual, from a biological sample isolated from this individual, comprising the steps of: a. isolation of nucleosides from said biological sample, by extraction of extracellular RNA and its fragmentation into nucleosides, b. isolation and determination of a respective quantity of at least 3, preferably at least 5, preferably at least 10, preferably at least 20, different nucleosides obtained during step a), and c. establishing, for said biological sample, a nucleoside profile from the respective quantities of each of the nucleosides obtained during step b), said profile being characteristic of said tumor.

[0042] According to a third embodiment, a method according to the invention is based on the simultaneous analysis of the quantity of different nucleosides derived from the monomeric catabolites present in a biological sample.

[0043] According to this third embodiment, the invention relates to an in vitro method for characterizing a tumor of an individual, from a biological sample isolated from this individual, comprising the steps of: a. isolation of nucleosides from said biological sample, by extraction of monomeric catabolites present in said sample, b. isolation and determination of a respective quantity of at least 3, preferably at least 5, preferably at least 10, preferably at least 20, different nucleosides obtained during step a), and c. establishing, for said biological sample, a nucleoside profile from the respective quantities of each of the nucleosides obtained during step b), said profile being characteristic of said tumor.

[0044] In an in vitro method for characterizing a tumor of an individual, from a biological sample isolated from said individual, said biological sample is in particular chosen from: - a solid biological sample, particularly a biopsy, and more particularly a biopsy of said tumor, and - a liquid biological sample, in particular a sample of a bodily fluid from said individual, more particularly a sample of blood, plasma, serum or urine.

[0045] By "biopsy" is meant the removal of a very small part of an organ or tissue. When the biological sample is a biopsy, the first embodiment of the method according to the invention, in which the total cellular RNA is extracted and then fragmented, is preferred.

[0046] When the biological sample is a liquid biological sample, the second and third embodiments of the method according to the invention, in which respectively the extracellular RNA is extracted then fragmented, or in which the RNA in the form of isolated nucleosides is extracted, are preferred.

[0047] In a method according to the invention, said biological sample is of sufficient volume, or comprises a sufficient number of cells, to allow reliable quantitative determination of at least 3 nucleosides resulting from the fragmentation of an extract of the total cellular RNA of said sample.

[0048] In the case of a biopsy, the total cellular RNA is extracted according to a method chosen from among the methods accessible to those skilled in the art, in particular a method as described in the present example. In the case of a liquid sample, such as blood or urine, said sample is pretreated if necessary, in order in particular to eliminate any interfering compounds, to concentrate said sample and / or to determine a standard value of concentration of a reference element, such as creatinine in urine, this standard value being used to calibrate the concentration of the sample from which the nucleoside profile is established.

[0049] In a method according to the invention, the total cellular RNA, the extracellular RNA and the isolated nucleosides are obtained from a biological sample by any method known to a person skilled in the art, said method comprising in particular an extraction step, optionally a fragmentation step, and a dephosphorylation step.

[0050] According to a particular aspect, the invention therefore relates to an in vitro method for characterizing a tumor of an individual from a biopsy of said tumor, said method comprising the preparation, from said biopsy, of an extract of total cellular RNA and fragmentation of said RNA into nucleosides.

[0051] According to one embodiment, in an in vitro method for characterizing a tumor of an individual according to the invention, at least 3 isolated nucleosides from the biological sample, obtained by i) the preparation of an extract of the total cellular RNA and its fragmentation into nucleosides, ii) by the preparation of an extract of the extracellular RNA and its fragmentation into nucleosides, and / or iii) by the extraction of the isolated nucleosides, are isolated and their respective quantity determined, said at least 3 nucleosides are chosen from: - unmodified nucleosides: adenosine (A), cytidine (C), guanosine (G), uridine (U), and - modified nucleosides (see Table 2).

[0052] The modified nucleosides result from the action of a large number of highly specific enzymes, the nucleosides undergo in particular methylation and rearrangement of carbon-nitrogen bonds. Said modified nucleosides are all the modified nucleosides known at the date of the present application, these nucleosides are notably cited in the publication of Jonkhout et al (“The RNA modification landscape in human disease”, RNA, Dec; 23 (12): 1754-1769, 2017), and in Table 2 of the present application.

[0053] According to different embodiments of a method which is the subject of the invention, said at least 3 nucleosides are chosen from the groups consisting of: - unmodified nucleosides: adenosine (A), cytidine (C), guanosine (G), uridine (U), - 2'-O-methyladenosine (Am), 1-methyladenosine (mlA), N6,N6-dimethyladenosine (m66A), N6,N6,2'-O-lrimcthyladcnosinc (m66Am), N6-methyladenosine (m6A), N6,2'-O-dimethyladenosine (m6Am), N4-acetylcytidine (ac4C), 2'-O-methylcytidine (Cm), 5-hydroxymethylcytidine (hm5C), 3-methylcytidine (m3C), 5-methylcytidine (m5C), 2'-O-methylguanosine (Gm), 1-methylguanosine (mlG), N2,N2,7-trimethylguanosine (m227G), N2,7-dimethylguanosine (m27G), 7-methylguanosine (m7G), 8-hydroxyguanosine (oxo8G), inosine (I), pseudouridine (Psi), queuosine (Q), 3,2'-O-dimethyluridine (m3Um), 5-methoxycarbonylmethyl-2-thiouridine (mcm5s2U), 5-methoxycarbonylmethyluridine (mcm5U), 5-carbamoylmethyluridine (ncm5U), 2'-O-methyluridine (Um), and / or - 3-(3-amino-3-carboxypropyl)uridine (acp3U), 2'-O-ribosyladenosine (phosphat) (Ar(p)), 5-carboxymethylaminomethyl-2-thiouridine (cmnm5s2U), 5-carboxymethylaminomethyluridine (cmnm5U), 5-carboxymethylaminomethyl-2'-O-methyluridine (cmnm5Um), dihydrouridine (D), 5-formylcytidin (f5C), galactosyl-queuosine (galQ), 2'-O-methyl-5-hydroxymethylcytidine (hm5Cm), 5-hydroxyuridine (ho5U), 5-hydroxyadenosine (ho8A), 8-hydroxyguanosine (ho8G), N6-isopentenyladenosine (i6A), N6-(cis-hydroxyisopentenyl)adenosine (io6A), 1-methyl lino sine (mil), 1-methylpseudouridine (mlpsi), N2,N2-dimethylguanosine (m22G), 2-methyladenosine (m2A), N2-methylguanosine (m2G), 5-methyluridine (m5U), 5, 2'-O-dimethyluridine (m5Um), N6-methyl-N6-threonylcarbamoyladenosine (m6t6A), mannosyl-queuosine (manQ), 5-(carboxyhydroxymethyl)uridinemethyl ester (mchm5U), 5-methylaminomethyl-2-thiouridine (mnm5s2U), 2-methylthio-N6-isopentenyladenosine (ms2i6A), 2-methylthio-N6-threonylcarbamoyladenosine (ms2t6A),peroxywybutosine (o2yW), 2'-O-methylpseudouridine (psi m), 2-thiouridine (s2U), N6-threonylcarbamoyladenosine (t6A), wybutosine (yW). ,

[0054] According to one embodiment of a method which is the subject of the invention, said at least 3 nucleosides are chosen from the groups consisting of: - unmodified nucleosides: adenosine (A), cytidine (C), guanosine (G), uridine (U), and - 2'-O-methyladenosine (Am), 1-methyladenosine (mlA), N6,N6-dimethyladenosine (m66A), N6,N6,2'-O-lrimcthyladcnosinc (m66Am), N6-methyladenosine (m6A), N6,2'-O-dimethyladenosine (m6Am), N4-acetylcytidine (ac4C), 2'-O-methylcytidine (Cm), 5-hydroxymethylcytidine (hm5C), 3-methylcytidine (m3C), 5-methylcytidine (m5C), 2'-O-methylguanosine (Gm), 1-methylguanosine (mlG), N2,N2,7-trimethylguanosine (m227G), N2,7-dimethylguanosine (m27G), 7-methylguanosine (m7G), 8-hydroxyguanosine (oxo8G), inosine (I), pseudouridine (Psi), queuosine (Q), 3,2'-O-dimethyluridine (m3Um), 5-methoxycarbonylmethyl-2-thiouridine (mcm5s2U), 5-methoxycarbonylmethyluridine (mcm5U), 5-carbamoylmethyluridine (ncm5U), 2'-O-methyluridine (Um).

[0055] According to a more particular embodiment, a method which is the subject of the invention comprises the isolation and quantitative determination of at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28 or 29 different nucleosides resulting from the fragmentation of the total RNA of said biological sample.

[0056] According to another more particular embodiment, a method which is the subject of the invention comprises the isolation and quantitative determination of at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28 or 29 different nucleosides resulting from the fragmentation of the extracellular RNA of said biological sample.

[0057] According to another more particular embodiment, a method which is the subject of the invention comprises the isolation and quantitative determination of at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28 or 29 different nucleosides resulting from the extraction of nucleosides from said biological sample.

[0058] According to a more particular embodiment, a method which is the subject of the invention comprises the isolation and quantitative determination of at least 3 different nucleosides resulting from the fragmentation of the total RNA of said biological sample and / or from the fragmentation of the extracellular RNA and / or from the extraction of the isolated nucleosides, said nucleosides being chosen from the following: adenosine (A), cytidine (C), guanosine (G), uridine (U), 2'-O-methyladenosine (Am), 1-methyladenosine (mlA), N6,N6-dimethyladenosine (m66A), N6,N6,2'-O-trimethyladenosine (m66Am), N6-methyladenosine (m6A), N6,2'-O-dimethyladenosine (m6Am), N4-acetylcytidine (ac4C), 2'-O-methylcytidine (Cm), 5-hydroxymethylcytidine (hm5C), 3-methylcytidine (m3C), 5-methylcytidine (m5C), 2'-O-methylguanosine (Gm), 1-methylguanosine (mlG), N2,N2,7-trimethylguanosine (m227G), N2,7-dimethylguanosine (m27G), 7-methylguanosine (m7G), 8-hydroxyguanosine (oxo8G), inosine (I), pseudouridine (Psi), queuosine (Q), 3,2'-O-dimethyluridine (m3Um), 5-methoxycarbonylmethyl-2-thiouridine (mcm5s2U), 5-methoxycarbonylmethyluridine (mcm5U), 5-carbamoylmethyluridine (ncm5U), 2'-O-methyluridine (Um). ,

[0059] In a method according to the invention, the isolation and determination of a respective quantity of at least 3 nucleosides are carried out by any means analysis known to those skilled in the art. These means include in particular chromatography, in particular high-performance reversed-phase liquid chromatography (RP-HPLC) or capillary electrophoresis (CE).

[0060] These means also include spectrometry means, in particular mass spectrometry. More particularly, these means include liquid chromatography-tandem mass spectrometry (LC-MS / MS), an analytical technique that combines the separation power of liquid chromatography with the highly sensitive and selective mass analysis capability of triple quadrupole mass spectrometry. The strength of this technique lies in the separation power of liquid chromatography for a wide range of compounds, combined with the ability of mass spectrometry to quantify compounds with a high degree of sensitivity and selectivity, based on the unique mass / charge (m / z) transitions of each compound of interest.

[0061] According to a particular aspect, in a method according to the invention, the mixture of nucleosides obtained by fragmentation is analyzed using high-performance liquid chromatography coupled with triple quadrupole tandem mass spectrometry (LC-MS / MS) in the Multiple Reaction Monitoring (MRM) mode. The MRM mode is a highly sensitive and specific technique which allows the quantification of molecules by mass spectrometry. This scan mode is dependent on tandem mass spectrometry and more particularly on triple quadrupole or hybrid trap mass spectrometry systems. The MRM scan mode is based on the selection of ions of specific mass and charge of a molecule, ions called precursor ions or parent ions, as well as on the corresponding fragment ions after fragmentation in the collision cell.The first quadrupole will allow the precise selection of precursor ions specific to the molecules of interest which will then be fragmented in the second quadrupole. The resulting fragment ions are then selected in the third quadrupole. The two ions (mass / charge) then correspond to a highly specific transition of the molecule of interest.

[0062] According to a more particular embodiment, the subject of the invention is an in vitro method for characterizing a tumor of an individual, from a biopsy of this individual, comprising the steps of: a. preparation, from said biological sample, of an extract of total cellular RNA and fragmentation of the polymeric RNA into nucleosides, b. isolation and determination of a respective quantity of at least 3, preferably at least 5, preferably at least 10, preferably at least 20, different nucleosides from step a), c. establishing, for said biological sample, a nucleoside profile from the respective quantities of each of the nucleosides obtained during step b), said profile being characteristic of said tumor.

[0063] According to an even more particular embodiment, the invention relates to an in vitro method for characterizing a tumor of an individual, said tumor being a tumor located in one of the following organs: colon, breast, pancreas, kidney, lung, or a hematological tumor, in particular leukemia.

[0064] According to a more particular embodiment, the subject of the invention is an in vitro method for characterizing a glial tumor of an individual, from a biological sample isolated from this individual, comprising the steps of: a. preparation, from said biological sample, of an extract of total cellular RNA and fragmentation of said RNA into nucleosides, b. isolation and determination of a respective quantity of at least 3, preferably at least 5, preferably at least 10, preferably at least 20, different nucleosides from step a), c. establishing, for said biological sample, a nucleoside profile from the respective quantities of each of the nucleosides obtained during step b), said profile being characteristic of said tumor, and d. prediction of a grade of said glial tumor by a first classification model previously trained, from the profile established during step c).

[0065] The terms "glial tumor" or "glioma" group together various brain tumors that develop from normal glial cells in the brain. The grade of a glial tumor represents the most important determinant for the survival of an individual bearing such a tumor. Non-tumor brain tissue is characterized by many cells with normal characteristics and some mitotic characteristics, without endothelial proliferation. Grade II tumors, also called "astroblastomas," comprise a larger number of cells comprising polymorphic nuclei undergoing mitosis. Grade III tumors are also called "anaplastic astroblastomas." Grade IV tumors correspond to glioblastoma multiforme.

[0066] The term "classification model" means a previously trained machine learning algorithm, in particular during supervised learning, as well as a training data set allowing the training of the aforementioned algorithm, and an evaluation data set.

[0067] According to embodiments, said first classification model may comprise: - a machine learning algorithm, - more particularly a supervised learning neural network, or - a multi-class probabilistic classification algorithm, previously trained with a training data set.

[0068] The training data set may comprise a multitude of data pairs, each of the data pairs comprising a first data item representing a nucleoside profile and a second data item representing the tumor grade for that profile.

[0069] The training data set may comprise a training set and a test set of the model. The model may thus be tested on the training set and the test set may be used to determine whether the training of the model is satisfactory or not.

[0070] The training set and the test set may be different. Alternatively, the test set may correspond to a portion of the training set.

[0071] The training data set may be previously constituted from data obtained in the laboratory by analysis of samples obtained from individuals suffering from cancer and whose tumor grade has been previously determined.

[0072] It is estimated that the classification model has reached a satisfactory level of learning on all the profiles of the test set if the classification reaches, for example, 85% accuracy; in other words, it is estimated that the classification model has reached a satisfactory level of learning on all the profiles of the test set if the classification reaches, for example, at most 15% error.

[0073] The classification model may consist of a computer program. The computer program may be written in any computer language such as, for example, C, C++, JAVA, Python, etc.

[0074] According to exemplary embodiments, the classification model may comprise a support vector machine, a random forest, a linear discriminant analysis; these methods are called in English respectively / 'Support Vector Machines", "Random Forests" and "Linear Discriminant Analysis".

[0075] These three families of machine learning algorithms are conceptually described in the literature (Comuejols and Miclet, “Apprentissage Artificiel: Concepts et Algorithmes” Eyrolles, 2012; Hastie et al. “The Elements of Statistical Learning: Data Mining, Inference, and Prediction”, 2nd Edition. Springer Sériés in Statistics, Springer 2009, ISBN 9780387848570) and are perfectly suited to multi-class classification.

[0076] According to embodiments of a method according to the invention, the prediction of a grade of a glial tumor can comprise: - prediction of a grade II glial tumor, - prediction of a grade III glial tumor or - prediction of a grade IV glial tumor.

[0077] The invention more particularly relates to an in vitro method for predicting a grade of a glial tumor of an individual, from a biological sample of said individual, and in particular a biopsy of said glial tumor, in which the prediction of a grade of said glial tumor by a previously trained classification model, comprises: the prediction of a grade II glial tumor, the prediction of a grade III glial tumor and the prediction of a grade IV glial tumor. More particularly, a method according to the invention for predicting a grade of a glial tumor of an individual comprises the distinction between a grade II glial tumor and a grade III or IV glial tumor; the distinction between a grade III glial tumor and a grade II or IV glial tumor; the distinction between a grade IV glial tumor and a grade II or III glial tumor.

[0078] The invention even more particularly relates to an in vitro method for characterizing a glial tumor of an individual, from a biopsy of said tumor, comprising: a. the preparation, from said biopsy, of an extract of total cellular RNA and fragmentation of said RNA into nucleosides, b. the isolation and quantitative determination of at least 3 nucleosides resulting from said fragmentation, chosen from: adenosine (A), cytidine (C), guanosine (G), uridine (U), 2'-O-methyladenosine (Am), 1-methyladenosine (mlA), N6,N6-dimethyladenosine (m66A), N6,N6,2'-O-trimethyladenosine (m66Am), N6-methyladenosine (m6A), N6,2'-O-dimethyladenosine (m6Am), N4-acetylcytidine (ac4C), 2'-O-methylcytidine (Cm), 5-hydroxymethylcytidine (hm5C), 3-methylcytidine (m3C), 5-methylcytidine (m5C), 2'-O-methylguanosine (Gm), 1-methylguanosine (mlG), N2,N2,7-trimethylguanosine (m227G), N2,7-dimethylguanosine (m27G), 7-methylguanosine (m7G), 8-hydroxyguanosine (oxo8G), inosine (I), pseudouridine (Psi), queuosine (Q), 3,2'-O-dimethyluridine (m3Um), 5-methoxycarbonylmethyl-2-thiouridine (mcm5s2U), 5-methoxycarbonylmethyluridine (mcm5U), 5-carbamoylmethyluridine (ncm5U), 2'-O-methyluridine (Um). c. establishing, for said tumor, a profile from the respective quantitative values of the nucleosides obtained during step b), said profile being characteristic of said tumor, and d. the prediction of a grade of said glial tumor by a previously trained classification model, from the profile established during step c), in which the prediction of a grade of a glial tumor is chosen from: the prediction of a grade II glial tumor, the prediction of a grade III glial tumor and the prediction of a grade IV glial tumor.

[0079] According to another aspect, the invention relates to an in vitro method for characterizing a glial tumor of an individual, said method comprising the steps of: a. preparation, from said biological sample, of an extract of total cellular RNA and fragmentation of said RNA into nucleosides, b. isolation and determination of a respective quantity of at least 3, preferably at least 5, preferably at least 10, preferably at least 20, different nucleosides from step a), c. establishing, for said biological sample, a nucleoside profile from the respective quantities of each of the nucleosides obtained during step b), said profile being characteristic of said tumor, and d. prediction of a survival status of said individual, by a second previously trained classification model, from the profile established during step c).

[0080] According to embodiments, said second classification model may comprise: - a machine learning algorithm, - more particularly a supervised learning neural network, or - a probabilistic classification algorithm, previously trained with a second training dataset.

[0081] Said second training data set may comprise a multitude of data pairs, each of the data pairs comprising a first data item representing a nucleoside profile and a second data item representing the survival state for this profile.

[0082] This training data set may comprise a training set and a model evaluation set. The model may thus be tested on the training set and the evaluation set may be used to determine whether the model training is satisfactory or not. The training set and the evaluation set may be different. Alternatively, the evaluation set may correspond to a part of the training set. The training data set may be previously constituted from data obtained in the laboratory by analyzing samples obtained from individuals suffering from cancer and whose survival status has been previously determined.

[0083] The classification model is considered to have reached a satisfactory level of learning on all the profiles in the evaluation set if the classification reaches 85% accuracy; in other words, the classification model is considered to have reached a satisfactory level of learning on all the profiles in the evaluation set if the classification reaches at most 15% error.

[0084] Just like the first classification model, the second classification model may consist of a computer program. The program computer can be written in any computer language such as for example C, C++, JAVA, Python, etc.

[0085] According to exemplary embodiments, the second classification model may comprise a support vector machine, a random forest, a linear discriminant analysis; these methods are called in English respectively / 'Support Vector Machines", "Random Forests" and "Linear Discriminant Analysis".

[0086] According to another aspect, the invention relates to a classification model, previously trained on a learning data set, to predict a grade of a glial tumor of an individual suffering from a tumor, from a nucleoside profile obtained by implementing a method according to the invention.

[0087] According to another particular aspect, the subject of the invention is a classification model, previously trained on a learning data set, to predict a survival status of an individual suffering from a tumor, from a nucleoside profile obtained by implementing a method according to the invention.

[0088] According to another aspect, the present invention relates to the use of a classification model according to the invention for predicting a grade of a glial tumor.

[0089] According to one aspect, the present invention relates to the use of a classification model according to the invention for the stratification of a patient suffering from a glial tumor, in combination with at least one other biological marker characteristic of said patient.

[0090] According to another aspect, the present invention relates to the use of a classification model according to the invention for predicting a survival state of an individual.

[0091] According to another particular aspect, the present invention finally relates to a diagnostic method comprising the implementation of a method according to the invention for characterizing a tumor. The present invention also relates to a diagnostic method comprising the implementation of a method according to the invention for predicting a grade of a glial tumor. The present invention also relates to a diagnostic method comprising the implementation of a method according to the invention for predicting the survival status of a patient. Said diagnostic method may further comprise a histological analysis of the tissues. Description of figures and embodiments

[0092] Other advantages and characteristics will appear on examining the detailed description of a non-limiting embodiment, and the attached drawings, in which:

[0093] [Fig.l] represents the overall scheme of the experiment, where LC-MS / MS designates liquid chromatography associated with mass spectrometry and the raw data (data) are the epitranscriptomic profiles obtained by LC-MS / MS.

[0094] [Fig.2] represents the overall diagram of the bioinformatics process, the raw data are the epitranscriptomic profiles obtained by LC-MS / MS, the normalized data are the epitranscriptomic profiles after normalization, MS designates mass spectrometry (combined with liquid chromatography).

[0095] Figures 3A, 3B and 3C represent, in the form of a box plot, six graphs representing respectively the relative quantity (in percentage) of six modified nucleosides according to the glial tumor grade. For each of the graphs, said grade is designated on the abscissa by: "Normal", "Grade-II", "Grade-III" or "Grade IV" indicating respectively a sample of non-tumor glial tissue or a sample of glial tumor of grade II, III or IV. [Fig.3A] shows two examples of nucleosides whose quantity decreases with increasing glial tumor grade: (from left to right) oxo8G and mlG. [Fig.3B] represents two examples of nucleosides whose quantity increases with increasing glial tumor grade: (from left to right) m6Am and Gm. [Fig.3C] shows two examples of nucleosides whose quantity varies slightly with increasing glial tumor grade: (from left to right) mlA and m7G.The scales are different depending on the graphs.

[0096] [Fig.4] represents the percentage of variance explained by the first components of the Principal Component Analysis (PCA) of the epitranscriptomic profiles of the cohort. On the abscissa, the components are numbered from 0 to 9. On the ordinate, the percentages of variance explained by these components.

[0097] [Fig.5] represents the three-dimensional visualization of the cohort profiles according to the said first three components of the Principal Component Analysis (PCA), i.e. the three components that hold 39.2 + 23.3 + 8.6 = 71.1% of the variance of the epitranscriptomic profiles of the cohort. Each of the axes represents, respectively, principal component 0 (39.24%), principal component 1 (23.27%) and principal component 2 (8.58%). The symbols "star" represent the "normal" guard, "triangle" the grade II, "square" the grade III and "cross" the grade IV, respectively.

[0098] It is understood that the embodiments which will be described below are in no way limiting. In particular, it will be possible to imagine variants of the invention comprising only a selection of characteristics described below isolated from the other characteristics described, if this selection of characteristics is sufficient to confer a technical advantage or to differentiate the invention from the state of the prior art. The present invention will be better understood by reading the following example, which is given to illustrate the invention and not to limit its scope.

[0099] EXAMPLE: Analysis of transcriptomic data from glial cell samples

[0100] This section presents the cohort used, sample preparation, method for obtaining epitranscriptomic profiles and the computer analysis program. This section then presents the results of exploratory analysis of the cohort profiles, prediction of tumor grades and survival prediction.

[0101] Sample preparation and mass spectrometry profile generation were performed as follows: Fifty-eight samples from surgically resected tumors in adult patients diagnosed with glioma, none of the patients having received chemical treatment or radiotherapy prior to surgery, were used in accordance with French bioethics laws regarding patient information and consent. At the time of resection, for each tumor, an aliquot was immediately frozen and stored at -80°C and the remaining tissue was fixed in 4% formalin, embedded in paraffin and 3 micron sections were cut and stained with hematoxylin eosin. The histopathological type of the tumor was determined according to the revised classification of the World Health Organization (Wesseling & Capper, “WHO 2016 Classification of gliomas.” Neuropathol Appl Neurobiol. 44, 139-150, 2018).The tumor group consisted of grade II (n = 20), grade III (n = 20), and grade IV (n = 18) gliomas. In addition, 19 “control” non-tumor glial cell samples (n = 19) were prepared according to the same protocol (described below) as the tumor samples.

[0102] Total RNA was extracted from tumor samples using the acid-phenol guanidium method. The quality of the RNA samples was determined by electrophoresis on agarose gels and ethidium bromide staining, and the 18S and 28S RNA bands were visualized under UV light. The processing of the biological sample begins with RNA extraction by phase separation to obtain an RNA sample of at least 100 ng. The processing continues with enzymatic hydrolysis of the polymeric RNA and dephosphorylation of the nucleosides.

[0103] Enzymatic digestion of RNA is performed as follows: 400 ng of RNA is diluted in a total volume of 20 μL of milliQ water, to which 3 μl of ammonium acetate (0.1 M pH 5.3) and 0.001 enzyme unit (U) of Nuclease PI (Sigma, N8630) are added. Incubation at 42°C is performed for 2 hours. Then, 3 μl of 1 M ammonium acetate and 0.001 U of alkaline phosphatase (Sigma, P4252) are added. The mixture is then incubated at 37°C for 2 hours. Finally, the nucleoside solution is diluted twice and filtered with 0.22 μm filters (Millex®-GV, Millipore, SLGVR04NL). Finally, 5 pL of each sample is injected and all samples are analyzed in triplicate by LC-MSMS.

[0104] Liquid chromatography (LC) is performed as follows: Nucleosides are separated by Nexera LC-40 systems (Shimadzu) using a Synergi™ Fusion-RP Cl8 column (4 μm particle size, 250 mm x 2 mm, 80 Å) (Phenomenex, 00G-4424-B0). The mobile phase consists of 5 mM ammonium acetate adjusted to pH 5.3 with acetic acid (solvent A) and pure acetonitrile (solvent B). The 30-minute elution gradient starts with 100% phase A followed by a linear gradient to 8% solvent B at 13 minutes. Solvent B is further increased to 40% in 10 minutes. After 2 minutes, solvent B is reduced to 0% at 25.5 minutes. The initial conditions are regenerated by rinsing with 100% solvent A for an additional 4.5 minutes. The flow rate is 0.4 ml / min and the column temperature is 35 °C.

[0105] Mass spectrometry in “Multiple Reaction Monitoring” (MRM) mode is performed as follows: detection is performed by Shimadzu TripleQuad 8060 in positive ion mode. Mass spectrometry operates in dynamic MRM mode with a retention time window of 3 min and a maximum cycle time set at 258 ms. Peak areas are determined using Skyline 4.1 software (Pino LK et al, “The Skyline ecosystem: Informatics for quantitative mass spectrometry proteomics.” Mass Spectrom Rev. 2020 May;39(3):229-244. 2020).

[0106] The mass spectrometer was calibrated to accurately identify and quantify 25 modified nucleosides (Table 2) and 4 unmodified nucleosides (A, U, G, T) (Table 1). The mass spectrometry device used is a Shimadzu TripleQuad 8060 in Multiple Reaction Monitoring mode.

[0107] Each sample was injected three times, thus providing three technical replicates. For each nucleoside, the homogeneity of the retention time given by the mass spectrometer is checked. Measurements showing a divergence of more than 6% were discarded. This results in a data table containing the quantity measurements of each nucleoside, in each replicate, for all samples. This table is then analyzed using our computer programs.

[0108] All bioinformatics analyses are performed with in-house developed python programs. For this, the authors used well-known (open source) modules: “Pandas” for tabular data management (Reback et al, Pandas-dev / pandas: Pandas 1.0.3 (Version v 1.0.3). Zenodo Tuesday, 18, 2020), “scikit-learn” for exploratory statistical data analyses and machine learning (Pedregosa et al, Journal of Machine Learning Research, vol. 12, pp. 2825-2830, 2011), “Matplotlib” for visualization (J.D. Hunter, Computing in Science & Engineering, vol. 9, no. 3, pp. 90-95, 2007).

[0109] The necessary characteristics of the programs are as follows: a) they take as input mass spectrometry data in a file in tabular format (CSV format); b) the quantifications of area and retention time from the spectrometer must be given as real values with an accuracy of at least 10-1; c) they implement a multi-class supervised machine learning algorithm among those mentioned above; d) they implement the training phase, the evaluation phase, and the prediction mode; e) they use the classification model in prediction mode to classify the epitranscriptome profile of a patient sample in order to predict the tumor grade.

[0110] For computer preprocessing and normalization, the raw quantity table is loaded into memory and its format is checked. Then, the average quantity of each nucleoside is calculated, and the table is reformatted to obtain all measurements on one line for each biological sample. Mass spectroscopy does not produce absolute counts of molecules but relative measurements. The inventors propose a new normalization formula, in which the quantities of the unmodified nucleosides A, C, G and U are summed. This sum serves as a reference. Then all the reference measurements are divided by this sum. Thus, relative measurements are obtained, all included in the interval [0, 1]. As an example, an extract of such a data table is provided in Tables 3, 4 and 5. [YES] [Tables3] -Grade and A Am C Cm G Gæ I Psi Quen&sl U ne Grade 1.455 5.507 5.185 1.219 3.258 7194 1.774 4.391 5.045E- 1.021 -H E-20 E-Ol E-0 E-04 E-03 05 E-02 Grade 7,084- 7,919 4,303 1,670 4,915 1,006 1,632 4,090 3,43Œ- 7,414 -III E-02 E-03 E-01 E-02 E-02 Grade 1,360 7,150 3,581 1,776 4,894 9,499 2,988 3,922 X500E- 6,571 -IV E-The E-02 E-The E-02 E-01 E-03 E-05 Non E-03 6,030 5,014 1,406. 3,286 7,311 1,992 2,928 9,334E- 1,007 al E-01 E-02 E-01 E-02 E-01 E-03 E-04 E-03 05 E-02

[0112] [Table 4] Grade lia sc4C 1üb5C mlA mlG m227 G mSC Grade- 9,863E 1,255E l j ’o s J fri LD 'ld LU. DÛ LO73E L.300E 3.744E 3O17E .H-'L* w 2728E II -05 -04 -06 -02 -03 -04 -06 -03 -06 -02 Grade- L4S9E 8378E 3,249E ' 2,7405 3J93E 3J49E 3,73æ 2,587e 4.434E 2,273E IH -04 -05 -07 -02 -04 -04 -06 -03 -06 -02 Grade- 1.760E LÎ12E 1 56GE 6559E L593E 6.544E 5.398E 4, 159E IV -04 -04 -06 ■02 -04 -04 -06 -03 -00 -02 : Nonna 1,24^- l,780E 3.208E 1,9HE Î.553E' 9,01 SE 4.470E L799E 3.654E I -04 -04 -06 -02 -03 -04 -06 -03 -06 -02

[0113] [T ableaux5 ] Grade nid 6 A mâ6A Kl m6A us 6Am æ7G nu ns 5 tr mm5s2 L nonSU oxoSG Grade- 3.035E E746E- 2.785E L303E 6.834E L599E- 2.199E- 5 9 96 F 2.454E n -03 06 -03 -04 -03 06 05 -05- -06 ■Grade- 4.077E S.60OE- 4.1I5E 2.53 SE 6.673E 9.258E- 3.282E- L757E 2.357E m -03 07 -03 -04 -03 07 05 -05 -07 ■Grade- 3.899E 9.45 TE- 4.1IGE 3.98 IE L185E 2162E- 7 748F - 7G54E L06QÊ IV -03 07 -03 -04 -02 06 05 -05 -06 Nonna 3368E 6 77TE- 3.183E L283E 9.999E 8J95E- LÎ18E-05 5.544E 5.862E 1 -03 07 -03 -04 -03 06 -05 -06

[0114] Tables 3, 4 and 5 indicate, for each of the nucleosides analyzed, the normalized data value for each of the grades II, III and IV of glioma, and for healthy tissues (“normal”).

[0115] The joint analysis of epitranscriptomic profiles and clinical variables of interest ([Fig.l]) is performed, in particular regarding the grade in the case of gliomas. This process can be adapted to any type of clinical variable. In this example, we seek to distinguish cancer grades, which can be difficult to establish by means of anatomopathological examination.

[0116] Preprocessing of the cohort profiles resulted in a table of 77 rows, with one row per sample, and 29 columns, with one column per measurement. For each sample, the mention of tumor grades or the mention "normal" for healthy samples was added. An exploratory statistical analysis of this table was carried out to assess the relevance of the signal contained in the profiles on the grade information.

[0117] First, the variations in the amounts of each nucleoside are studied in samples of the same grade, and we compare these variations between grades. As shown in the box-and-whisker plots of Figures 3A, 3B and 3C, the experimental results suggest a grouping of nucleosides into four groups: i) those whose amount increases with grade, i.e. between non-tumor brain tissue (designated for simplification as "normal" on the ordinates of the graphs) and grades II, III and IV, including the nucleosides oxo8G, ml G, queuosine and Ac4C (as shown for example in [Fig.3A]); ii) those whose quantity decreases with the grade (as shown for example in [Fig.3B]), iii) those which vary weakly with the grades (as shown for example in [Fig.3C]) and iv) the remaining nucleosides, which do not satisfy the conditions of belonging to the first three groups.

[0118] At first glance, none of these groups is linked to a specific known feature of its constituents (e.g., modified nucleoside edge). Nevertheless, it is worth noting that 2'-O-methylations (Am, Um, Cm, Gm), mainly found in ribosomal RNA (rRNA) and small nuclear RNA (snRNA), behave similarly in a central cluster containing m6Am, a specific modification of rRNA.

[0119] Then, a Principal Component Analysis (PCA) of these data was carried out, in order to perform a dimension reduction, not to be confused with a selection of "features", in other words nucleosides, to see if the variations in quantity could be combined into a small number of components. [Fig.4] shows the percentage of variance explained for the first 10 components of the PCA: clearly the first three components group together a large majority of the variations in the profiles. We can see that the first three components alone group together: 39.2 + 23.3 + 8.6 = 71.1% of the variance in the epitranscriptomic profiles of the cohort.

[0120] Each epitranscriptomic profile comprising measurements for x nucleosides is seen mathematically as a point in an x-dimensional space. PCA is an exploratory multivariate analysis method that allows reducing the dimensions of the data while capturing their variability. Components are new variables that combine the data from the initial observations in order to best capture their variability while reducing the number of variables to be analyzed. The components result from the projection of the initial data onto other axes of the multidimensional space. The components are ordered in descending order of percentage of variance explained. This percentage associated with each component indicates its importance in describing the initial data. [Fig.4] shows the graph of the percentage of variance explained for the first 10 components. PCA is a classic technique of data analysis.

[0121] The 3-dimensional visualization of the projected profiles on the first three components is shown in [Fig.5]. First, non-tumor and grade II tissue samples clearly separate from those of grades III and IV. Furthermore, grade III samples occupy a relatively separate volume from those of grade IV. These exploratory results suggest that supervised machine learning algorithms should be able to learn a boundary between groups of samples of different grades.

[0122] Machine learning method for accurate prediction of tumor grade and healthy samples

[0123] A machine learning method was tested to determine whether the grade of the samples could be predicted from the epitranscriptomic profiles alone, i.e. without the help of any other information than the quantities of nucleosides ([Fig. 1]). To do this, the profiles were partitioned into two distinct subsets: the first was used only to train the machine learning model (n=60, or 78%), the second was used to evaluate the model (n=17, or 22%).

[0124] As the variable to be predicted (here the grade) is categorical data, the learning method must belong to the classification category. Among the major types of learning algorithms, a Support Vector Machine (SVM) classification algorithm was chosen for the possibility it offers to adapt the boundary formulas by changing the kernel type, as is the norm in learning. The prediction accuracy of the SVM algorithm equipped with a linear kernel on the profiles of the tested subset is 0.90, out of a maximum of 1, which is remarkable. The level of prediction accuracy is maintained when the learning and then the tests are repeated with new random partitionings of the data set, which shows the robustness of the developed learning tool.

[0125] Furthermore, the results of the evaluation allow us to compare our normalization method (designated SUM, for sum) with the formulas used in the literature. Indeed, the classic normalization which consists of dividing the measurement of a modified nucleoside, for example ml A, by that of the corresponding unmodified nucleoside, here the measurement of A. In Table 6, the precision according to the use of different formulas is between 0.8 and 0.9, and is therefore always less than or equal to (but never greater than) the precision of the SUM normalization formula.

[0126] [Tableauxô] ACGU SUM Standardization Accuracy of 0.85 0.90 0.85 0.80 0.90 prediction

[0127] Furthermore, the grade prediction is robust to changing the classification algorithm. Instead of an SVM algorithm, if an algorithm based on a Linear Discriminant Analysis approach is used, an accuracy of 92% is obtained, with a recall (or sensitivity) of 90% and an Fl-score of 90%. The details of the predictions for each grade are given in Table 7.

[0128] [Tables7] Grade precision recall score-fl Normal 1.00 0.80 0.89 Grade-II 0.67 1.00 0.80 Grade-in 1.00 0.83 0.91 Grade-IV 0.88 1.00 0.93 Weighted mean 0.92 0.90 0 90

[0129] In conclusion, the quality of grade prediction is not particularly linked to the optimization of a learning method on a given cohort, since two very different learning methods obtain similar results. The quality of the prediction is therefore linked to the power of the signal contained in the transcriptomic profiles.

[0130] Furthermore, the learning models whose results are reported here have deliberately not been optimized with respect to their parameters, in order to avoid a risk of over-learning which would affect the generalization capacity of the models. Prediction of the survival status of patients

[0131] The same supervised learning approach was used to predict the clinical variable indicating survival status, i.e., the status “alive” or “dead” at the end of the cohort follow-up, i.e., in 2020. Here, the classification is binary: “alive” or “dead”. The SVM learning algorithm gives an 80% correct prediction, which is convincing given the size of the cohort considered (Table 8).

[0132] [Tables8] Class precision recall seore-fl False (alive) 0.75 0.86 0.80 True (deceased) 0.89 0.80 0.84 Weighted mean 0.83 0.82 0.82 Conclusion

[0133] Differences in the relative quantities of certain epigenetic modifications of RNAs have been highlighted according to different samples, whether healthy or tumorous. These differences make it possible in particular to separate the different tumor grades. A supervised machine learning algorithm applied to the nucleoside quantity vectors makes it possible to effectively distinguish the grades of gliomas, and in particular to distinguish grades II and III, with remarkable precision given the relatively small size of the cohort. Furthermore, this method also makes it possible, from the same data, to estimate patient survival using a supervised machine learning method.

Claims

Claims

1. An in vitro method for characterizing a tumor of an individual, from a solid biological sample isolated from this individual, said biological sample being a biopsy, comprising the steps of: a. isolation of the nucleosides of said biological sample, by extraction of: i) total cellular RNA and its fragmentation into nucleosides, ii) extracellular RNA and its fragmentation into nucleosides and / or iii) nucleosides derived from monomeric catabolites, b. isolation by chromatography and determination of a respective quantity of at least 3, preferably at least 5, preferably at least 10, preferably at least 20, different nucleosides derived from step a), and c. establishment, for said biological sample, of a nucleoside profile from the respective quantities of each of the nucleosides obtained during step b), said profile being characteristic of said tumor.

2. Method according to the preceding claim, wherein said biological sample is a biopsy of said tumor.

3. A method according to any one of claims 1 or 2, wherein said nucleosides are selected from the following: • unmodified nucleosides: adenosine (A), cytidine (C), guanosine (G), uridine (U), and • modified nucleosides: 2'-O-methyladenosine (Am), 1-methyladenosine (mlA), N6,N6-dimethyladenosine (m66A), N6,N6,2'-O-trimethyladenosine (m66Am), N6-methyladenosine (m6A), N6,2'-O-dimethyladenosine (m6Am), N4-acetylcytidine (ac4C), 2'-O-methylcytidine (Cm), 5-hydroxymethylcytidine (hm5C), 3-methylcytidine (m3C), 5-methylcytidine (m5C), 2'-O-methylguanosine (Gm), 1-methylguanosine (mlG), N2,N2,7-trimethylguanosine (m227G), N2,7-dimethylguanosine (m27G), 7-methylguanosine (m7G), 8-hydroxyguanosine (oxo8G), inosine (I), pseudouridine (Psi), queuosine (Q), 3,2'-O-dimethyluridine (m3Um), 5-methoxycarbonylmethyl-2-thiouridine (mcm5s2U), 5-methoxycarbonylmethyluridine (mcm5U), 5-carbamoylmethyluridine (ncm5U), 2'-O-methyluridine (Um).

4. Method according to any one of the preceding claims, characterized in that said tumor is a glial tumor and in that it comprises a step of predicting a grade of said glial tumor by a first classification model previously trained, from the profile established during step c), and in that the first classification model comprises: • a machine learning algorithm, • a supervised learning neural network, or • a multi-class probabilistic classification algorithm, previously trained with a training data set.

5. A method according to any one of claims 4 or 5, wherein the prediction of a grade of a glial tumor is selected from: the prediction of a grade II glial tumor, the prediction of a grade III glial tumor and the prediction of a grade IV glial tumor.

6. Method according to any one of the preceding claims, characterized in that it comprises a step of predicting a survival state of said individual, by a second previously trained classification model, from the profile established during step c),

7. and in that the second classification model comprises: • a machine learning algorithm, • a supervised learning neural network, or • a probabilistic classification algorithm, previously trained with a training data set. Computer program product comprising instructions, which when executed by a computer, implement the prediction step of the method according to any one of claims 4 to 6.