A computational method for small RNA biomarker modification classification
By training deep learning models on nanopore sequencing platforms with structured adapters, the method addresses the challenge of sequencing and classifying small non-coding RNA modifications, enhancing diagnostic potential through precise RNA modification detection.
Patent Information
- Application Number
- PCT/EP2025/060457
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-16
- Filing Date
- 2025-04-15
- Publication Date
- 2025-11-20
AI Technical Summary
Current methods for detecting small non-coding RNA modifications, such as 2'-O-methylation, are limited by technical challenges, particularly in nanopore direct RNA sequencing, which has been optimized for longer nucleic acids and prohibits sequencing of shorter molecules under 100 nucleotides, hindering diagnostic applications.
A method for training deep learning classification models on a nanopore direct sequencing platform to classify target RNA modifications, using structured adapters and a neural network system to analyze small non-coding RNAs, enabling accurate sequencing and modification detection of RNA molecules less than 50 nucleotides.
Enables accurate classification and detection of RNA modifications in small non-coding RNAs, improving diagnostic capabilities by overcoming limitations in nanopore sequencing technology for shorter RNA molecules.
Smart Images

Figure IMGF000041_0001 
Figure IMGF000042_0001 
Figure IMGF000042_0002
Abstract
Description
[0001] A COMPUTATIONAL METHOD FOR SMALL RNA BIOMARKER MODIFICATION CLASSIFICATION
[0002] The present invention relates to a method classifying target RNA modifications in a neural network system, to a method for training a neural network system that is configured to classify target RNA modifications, to a neural network system, to a computer system adapted to classify target RNA modifications, to a computer program and to a non-transitory readable medium.
[0003] BACKGROUND OF THE INVENTION
[0004] Small non-coding RNAs can be used due to their diversity and characteristic expression for diagnostic purposes.
[0005] The canonical RNA bases are A, C, U, and G. However, in contrast to DNA, which contains only a dozen epigenetic modifications, RNA is decorated with almost 170 different modifications, collectively referred to as the ‘epitranscriptome’, which has emerged as an important regulatory layer of RNA metabolism and general cell homoeostasis. These RNA modifications encompass 2'-O-methylation (Nm), N6-methyladenosine (m6A), N6,2'-O- dimethyladenosine (m6Am), 8-oxo-7,8-dihydroguanosine (8-oxoG), pseudouridine ( ), 5- methylcytidine (m5C), and N4-acetylcytidine (ac4C).
[0006] Covalent modifications are also pervasive of small non-coding RNAs and have fundamental roles in the regulation of several cellular processes. These RNA modifications are dynamic and their frequency at certain variable sites has been shown to be correlated with certain diseases such as cancer.
[0007] Thus, detecting said small non-coding RNA modifications would improve diagnostics, but technical limitations have hampered these efforts so far. To date, for example, no algorithm has been described for the basecalling and / or single molecule classification of 2'-O-methylated (Nm) RNA bases.
[0008] Nanopore direct RNA sequencing is a third-generation sequencing technology that allows the sequencing of native RNA molecules, thus, providing a direct way to detect modifications at single-molecule resolution. Despite recent advances, the analysis of nanopore sequencing data for RNA modification detection is still a complex task that presents many challenges. The Oxford Nanopore Technology (ONT) platform, for example, has traditionally been optimized for the sequencing of long nucleic acids, and technical limitations have prohibited the sequencing of shorter molecules under 100 nucleotides.
[0009] Recently, the present inventors have developed a method allowing targeted sequencing of small non-coding RNAs on a nanopore direct sequencing platform (e.g. the ONT platform) in order to simultaneously measure both RNA abundance and modification patterns. This was a milestone on the way to direct sequencing and analysing of small non-coding RNAs.
[0010] Here, the present inventors have developed a method for the training of deep learning classification models for targeted small non-coding RNA biomarker sequencing and methylation classification on a nanopore direct sequencing platform (e.g. the ONT platform).
[0011] SUMMARY OF THE INVENTION
[0012] In a first aspect, the present invention relates to a method for classifying target RNA modifications in a neural network system, the method comprising: a) obtaining direct RNA sequencing data relating to at least one target RNA in at least one sample; b) selecting a first portion of the direct RNA sequencing data, wherein the first portion of the direct RNA sequencing data comprises first direct RNA sequencing information being indicative for the presence or absence and / or being indicative of the pattern of at least one modification in the at least one target RNA, c) providing the first portion of the direct RNA sequencing data to a first neural network of the neural network system, the first neural network being trained to determine the presence or absence and / or the pattern of at least one modification in said target RNA from direct RNA sequencing data, d) receiving a RNA modification prediction from the first neural network, the RNA modification prediction being indicative for the presence or absence and / or being indicative of the pattern of the at least one modification in the at least one target RNA in the at least one sample.
[0013] In a second aspect, the present invention relates to a method for training a neural network system that is configured to classify target RNA modifications, the method comprising the steps of a) selecting direct RNA sequencing training data, the direct RNA sequencing training data comprising information on target training RNA; b) selecting a first portion of the direct RNA sequencing training data, the first portion of the direct RNA sequencing training data comprising first target training RNA information being indicative of a presence or absence and / or being indicative of the pattern of at least one modification in the target training RNA; c) training a first neural network of the neural network system using the first portion of the direct RNA sequencing training data as input and the determination of the presence or absence and / or the determination of the pattern of at least one modification in the target training RNA as target.
[0014] In a third aspect, the present invention relates to a neural network system implemented on a computer, the system comprising a first neural network and / or a second neural network, wherein the first neural network and / or the second neural network are trained according to the method of the second aspect.
[0015] In a fourth aspect, the present invention relates to a computer system adapted to classify target RNA modifications comprising: at least one memory configured to store computer program code; at least one processor configured to access the computer program code and operate as instructed by the computer program code to perform the method according to the first or second aspect.
[0016] Ina fifth aspect, the present invention relates to a computer program comprising instructions which, when the program is executed by a computer system, causes the computer system to carry out the method according to the first or second aspect.
[0017] In a sixth aspect, the present invention relates to a non-transitory computer readable medium storing the computer programme according to the fifth aspect.
[0018] This summary of the invention does not necessarily describe all features of the present invention. Other embodiments will become apparent from a review of the ensuing detailed description.
[0019] DETAILED DESCRIPTION OF THE INVENTION
[0020] Definitions
[0021] Before the present invention is described in detail below, it is to be understood that this invention is not limited to the particular methodology, protocols and reagents described herein as these may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the present invention which will be limited only by the appended claims. Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art.
[0022] Preferably, the terms used herein are defined as described in “A multilingual glossary of biotechnological terms: (IUPAC Recommendations)”, Leuenberger, H.G.W, Nagel, B. and Kolbl, H. eds. (1995), Helvetica Chimica Acta, CH-4010 Basel, Switzerland).
[0023] Several documents are cited throughout the text of this specification. Each of the documents cited herein (including all patents, patent applications, scientific publications, manufacturer's specifications, instructions, GenBank Accession Number sequence submissions etc.), whether supra or infra, is hereby incorporated by reference in its entirety. Nothing herein is to be construed as an admission that the invention is not entitled to antedate such disclosure by virtue of prior invention. In the event of a conflict between the definitions or teachings of such incorporated references and definitions or teachings recited in the present specification, the text of the present specification takes precedence.
[0024] The term “comprise” or variations such as “comprises” or “comprising” according to the present invention means the inclusion of a stated integer or group of integers but not the exclusion of any other integer or group of integers. The term “consisting essentially of’ according to the present invention means the inclusion of a stated integer or group of integers, while excluding modifications or other integers which would materially affect or alter the stated integer. The term “consisting of’ or variations such as “consists of’ according to the present invention means the inclusion of a stated integer or group of integers and the exclusion of any other integer or group of integers.
[0025] The terms “a” and “an” and “the” and similar reference used in the context of describing the invention (especially in the context of the claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context.
[0026] The term “nucleotide”, as used herein, refers to an organic molecule consisting of a nucleoside and a phosphate. In particular, a nucleotide is composed of three subunit molecules: a nucleobase, a five-carbon sugar (ribose or deoxyribose), and a phosphate group consisting of one to three phosphates. The four nucleobases in DNA are guanine, adenine, cytosine and thymine; in RNA, uracil is used in place of thymine. The nucleotide serves as monomeric unit of nucleic acid polymers, such as deoxyribonucleotide acid (DNA) or ribonucleotide acid (RNA). Thus, the nucleotide is a molecular building-block of DNA and RNA.
[0027] The term “nucleoside”, as used herein, refers to a glycosylamine that can be thought of a nucleotide without a phosphate group. A nucleoside consists simply of a nucleobase (also termed a nitrogenous base) and a five-carbon sugar (ribose or 2'-deoxyribose) whereas a nucleotide is composed of a nucleobase, a five-carbon sugar, and one or more phosphate groups. In a nucleoside, the anomeric carbon is linked through a glycosidic bond to the N9 of a purine or the N1 of a pyrimindine.
[0028] The terms “nucleotide sequence” or “polynucleotide” are interchangeably used herein and refer to single-stranded and double-stranded polymers of nucleotide monomers, including without limitation, 2'-deoxyribonucleotides (DNA) and ribonucleotides (RNA) linked by internucleotide phosphodiester bond linkages, or internucleotide analogs, and associated counter ions, e.g., H+, NH4+, trialkylammonium, Mg2+, Na+, and the like. A nucleotide sequence or polynucleotide may be composed entirely of deoxyribonucleotides, entirely of ribonucleotides, or chimeric mixtures thereof and may include nucleotide analogs. The nucleotide monomer units may comprise any of the nucleotides described herein, including, but not limited to, nucleotides and / or nucleotide analogs.
[0029] The term “RNA molecule”, as used herein, refers to a polymeric form of ribonucleotides of any length. Like DNA, RNA is assembled as a chain of nucleotides, but unlike DNA, RNA is found in nature as a single strand folded onto itself, rather than a paired double strand.
[0030] The RNA molecule may be derived from any number of sources, including without limitation, humans and animals. These sources may include, but are not limited to, whole blood, a tissue biopsy, lymph, bone marrow, amniotic fluid, hair, skin, semen, biowarfare agents, anal secretions, vaginal secretions, perspiration, saliva, or buccal swabs. However, various environmental samples (for example, agricultural, water, and soil), research samples generally, purified samples generally, cultured cells and lysed cells may also be used as samples.
[0031] (All) RNA molecules comprised in a sample, e.g. biological sample, may be analysed such as sequenced. However, it is also possible to sequence and analyse a specific RNA molecule (i.e. the target).
[0032] The term “small RNA molecule”, as described herein, refers to a polymeric RNA molecule that is less than 200 ribonucleotides, preferably < 50 ribonucleotides, more preferably < 30 ribonucleotides, in length. Specifically, small RNA molecules have a length of between 10 and < 50 ribonucleotides. More specifically, small RNA molecules have a length of between 10 and < 30 ribonucleotides. Small RNA molecules are usually non-coding RNA molecules. The small (non-coding) RNA molecules may be miRNAs, piRNAs, siRNAs, or fragments of longer RNAs, preferably rRNA or tRNA fragments.
[0033] The term “target RNA”, as used herein, refers to a ribonucleotide sequence that is sought to be detected. The target RNA may be derived from any number of sources, including without limitation, humans and animals. These sources may include, but are not limited to, whole blood, a tissue biopsy, lymph, bone marrow, amniotic fluid, hair, skin, semen, biowarfare agents, anal secretions, vaginal secretions, perspiration, saliva, or buccal swabs. However, various environmental samples (for example, agricultural, water, and soil), research samples generally, purified samples generally, cultured cells and lysed cells may also be used as samples.
[0034] It is preferred that the target RNA is non-coding RNA, specifically non-coding small RNA. It is more preferred that the (non-coding) target RNA has a length of < 50 ribonucleotides, particularly a length of between 10 and < 50 ribonucleotides, e.g. a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, or 49 ribonucleotides. It is even more preferred that the (noncoding) target RNA has a length of < 30 ribonucleotides, particularly a length of between 10 and < 30 ribonucleotides, e.g. a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, or 29 ribonucleotides.
[0035] The (non-coding) target RNA, specifically (non-coding) small RNA, may be a miRNA, piRNA, siRNA, or a fragment of longer RNAs, preferably a rRNA or tRNA fragment.
[0036] It will be appreciated that the term “target RNA” may refer to the target molecule itself as well as to surrogates thereof, for example, products obtained by ligation / hybridization (RNA / adapter structures) or reverse transcription (e.g. cDNAs or RNA / cDNA duplexes).
[0037] Specifically, the (non-coding) target RNA is a miRNA or a miRNA isoform (an isomiR).
[0038] The term “miRNA” (the designation “microRNA” is also possible), as used herein, refers to a single-stranded RNA molecule. The miRNA may be a molecule of 10 to 45 nucleotides in length, e.g. 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 nucleotides in length, not including optionally labels and / or elongated sequences (e.g. biotin stretches).
[0039] The miRNAs regulate gene expression and are encoded by genes from whose DNA they are transcribed but miRNAs are not translated into protein (i.e. miRNAs are non-coding RNAs).
[0040] The term “isomiR” (or “miRNA isoform”), as used herein, refers to a miRNA that varies slightly in sequence, which results from variations in the cleavage site during miRNA biogenesis or by processes which affect the mature miRNA after the biogenesis has occurred, such as oligouridylation. IsomiRs (miRNA isoforms) can be divided into three main categories: 3' isomiRs (trimmed or addition of one or more nucleotides at the 3' position), 5' isomiRs (trimmed or addition of one or more nucleotides at the 5' position), and polymorphic isomiRs (some nucleotides within the sequence are different from the wild-type mature miRNA sequence.
[0041] The term “modified target RNA”, as used herein, refers to RNA which has been post- transcriptionally edited or modified. This process is also designated as RNA editing. RNA editing influences diverse RNA processes, including generation, transportation, function, and metabolization. Thus, RNA editing is a critical regulator of cell biology. Typical RNA modifications are selected from the group consisting of 2'-O-methylation, N6-methyladenosine (m6A), N6,2'-O-dimethyladenosine (m6Am), 8-oxo-7,8-dihydroguanosine (8-oxoG), pseudouridine ( ), 5 -methyl cytidine (m5C), and N4-acetylcytidine (ac4C). Preferably, the RNA modification is an RNA methylation.
[0042] The term “methylated target RNA”, as used herein, refers to RNA which has been post- transcriptionally edited or modified by methylation. The methylation can occur at a base (e.g. methyl-6-adenine, pseudouridine) and / or ribose ring (2'-ortho-methylated nucleotide (2'-O- m)). The methylation of RNA occurring at a base is preferably selected from the group consisting of 6-methyladenosine (m6A), 5 -methyl cytidine (m5C), 5-methyluridine (m5U), 3- methyluridine (m3U), 1 -methyladenosine (mfA), and 1 -methylguanosine (mfG), or is a combination thereof.
[0043] The 2'-O-methylation of the backbone ribose is the most common and conserved type of RNA modification. The methylation of RNA occurring at a ribose ring is selected from the group consisting of 3 '-end 2'-O-methyladenosine (Am), 2'-O-methyluridine (Um), 2'-O- methylguanosine (Gm), and 2'-O-methylcytidine (Cm), or is a combination thereof.
[0044] The term “target RNA modification proportion”, as used herein, refers to the modification percentage of the target RNA in a sample. The sample may be from a patient to be tested or from a control subject.
[0045] Specifically, the modification percentage / proportion of the target RNA ranges between 100% (i.e. fully modified, or all detected RNA entities are fully modified) and 0% (i.e. not modified, or all detected RNA entities are not modified).
[0046] More specifically, the modification percentage / proportion of the target RNA is 0, 1, 2, 3, 3, 4,
[0047] 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31,
[0048] 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56,
[0049] 57, 58, 59, 60, 61 ,62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81,
[0050] 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100%.
[0051] Preferably, the target RNA modification proportion is a target RNA methylation proportion. If the diagnosis of a disease or condition is desired, the target RNA modification proportion identified in a patient is compared to a target RNA modification proportion in a control subject (e.g. a healthy individual). This comparison then allows the diagnosis of a disease or condition.
[0052] The term “nanopore sequencing”, as used herein, is a third-generation approach used in the sequencing of biopolymers, specifically, polynucleotides in the form of DNA or RNA. Using nanopore sequencing, a single molecule of DNA or RNA can be sequenced without the need for Polymerase Chain Reaction (PCR) amplification or chemical labelling of the sample. Nanopore sequencing has the potential to offer relatively low-cost genotyping, high mobility for testing, and rapid processing of samples. Nanopore sequencing is the only sequencing technology to enable real-time analysis in fully scalable formats.
[0053] The principle of nanopore sequencing is as follows: The biological or solid-state membrane, where the nanopore is found, is surrounded by electrolyte solution. The membrane splits the solution into two chambers. A bias voltage is applied across the membrane inducing an electric field that drives charged particles, in this case the ions, into motion. This effect is known as electrophoresis. For high enough concentrations, the electrolyte solution is well distributed and all the voltage drop concentrates near and inside the nanopore. This means charged particles in the solution only feel a force from the electric field when they are near the pore region. This region is often referred as the capture region. Inside the capture region, ions have a directed motion that can be recorded as a steady ionic current by placing electrodes near the membrane. Imagine now a nano-sized polymer such as DNA or protein placed in one of the chambers. This molecule also has a net charge that feels a force from the electric field when it is found in the capture region. The molecule approaches this capture region aided by brownian motion and any attraction it might have to the surface of the membrane. Once inside the nanopore, the molecule translocates through via a combination of electro-phoretic, electroosmotic and sometimes thermo-phoretic forces. Inside the pore the molecule occupies a volume that partially restricts the flow of ions, observed as an ionic current drop. Based on various factors such as geometry, size and chemical composition, the change in magnitude of the ionic current and the duration of the translocation will vary. The magnitude of the electric current density across a nanopore surface depends on the nanopores dimensions and the composition of DNA or RNA that is occupying the nanopore.
[0054] Sequencing was made possible because, passing through the channel of the nanopore, the samples cause characteristic changes in the density of the electric current flowing through the nanopore. The total charge flowing through a nanopore channel is equal to the surface integral of electric current density flux across the nanopore unit normal surfaces between times ti and t2.
[0055] In other words, different molecules can be sensed and potentially identified based on the modulation in ionic current. A strand of DNA or RNA is made up of a sequence of different combinations of four nucleotide bases: A, T (or U for RNA), G and C. Each base that passes through the nanopore can be identified through the characteristic disruption it causes to the current in real-time. Not only the DNA or RNA sequence can be detected, but also DNA or RNA modifications can be quantified in this way.
[0056] Two types of nanopore sequencing exist: biological and solid state nanopore sequencing.
[0057] The term “biological nanopore sequencing”, as used herein, refers to the use of transmembrane proteins, called protein nanopores, in particular, formed by protein toxins, that are embedded in lipid membranes so as to create size dependent porous surfaces - with nanometer scale "holes" distributed across the membranes. Sufficiently low translocation velocity can be attained through the incorporation of various proteins that facilitate the movement of DNA or RNA through the pores of the lipid membranes
[0058] Alpha hemolysin, which is a nanopore from bacteria that causes lysis of red blood cells or Mycobacterium smegamatis porin A (MspA) may be used for nanopore sequencing.
[0059] The term “solid state nanopore sequencing”, as used herein, refers to a sequencing approach which, unlike biological nanopore sequencing, does not incorporate proteins into its system. Instead, solid state nanopore technology uses various metal or metal alloy substrates with nanometer sized pores that allow DNA or RNA to pass through. These substrates most often serve integral roles in the sequence recognition of nucleic acids as they translocate through the channels along the substrates.
[0060] Nanopore sequencing platforms are offered, for example, by Oxford Nanopore Technologies Ltd. All Oxford Nanopore sequencing devices use flow cells which contain an array of tiny holes - nanopores - embedded in an electro-resistant membrane. Each nanopore corresponds to its own electrode connected to a channel and sensor chip, which measures the electric current that flows through the nanopore. When a molecule passes through a nanopore, the current is disrupted to produce a characteristic ‘squiggle’. The squiggle is then decoded using basecalling algorithms to determine the DNA or RNA sequence in real time.
[0061] Specifically, a strand of DNA or RNA is made up of a sequence of different combinations of four nucleotide bases: A, T (or U for RNA), G and C. Each base that passes through the nanopore can be identified through the characteristic disruption it causes to the current in realtime. This makes nanopore sequencing unique, in that it is the only sequencing technology that enables direct, real-time analysis of DNA / RNA in fully scalable formats. Advantages of realtime sequencing include rapid access to time critical information (e.g. pathogen identification), the generation of early sample insights and more control over the sequencing experiment. Nanopore sequencing is limited only by the length of the DNA / RNA fragment presented to the pore and can, therefore, span entire repetitive regions, resolve structural variants, and differentiate between different isoforms. The ability to sequence native DNA and RNA without the requirement for amplification, eliminates PCR bias and allows for the identification of base modifications, such as methylation, alongside nucleotide sequence.
[0062] Particularly, the nanopore sequencing system, e.g. the Oxford Nanopore sequencing system, uses, in addition to flow cells, which contain an array of tiny holes - nanopores - embedded in an electro-resistant membrane, two more elements: a nanopore adapter / motor protein complex and a tether oligonucleotide. The nanopore adapter is required for attaching the DNA or RNA molecule to be sequenced to the nanopore. In addition, the motor protein is required for directing the DNA or RNA molecule to be sequence through the nanopore in order to allow sequencing. Specifically, the motor protein controls translocation of the RNA or DNA strand through the nanopore. Once the DNA or RNA has passed through, the motor protein detaches and the nanopore is ready to accept the next DNA or RNA molecule. The motor protein is often an enzyme. The tether oligonucleotide has the function of concentrating the RNA or DNA target which is to be sequenced at the membrane surface of a nanopore flow cell. An electrically resistant membrane is further used, which means that all current must pass through the nanopore to ensure a clean signal. The subsequent sequencing of the DNA or RNA molecule is possible as, when the DNA or RNA molecule passes through the nano-scale hole, the current changes / fluctuates. This signal can be detected and is converted to the nucleotide sequence by a basecalling.
[0063] The standard preparation of a direct RNA sequencing library for use on the Oxford Nanopore platform involves the ligation of a 3 ’adapter only. This however, prevents the sequencing of the 20-50 nt at the 5 ’end of RNAs, constituting a small region of sequence that is lost on long RNA inserts however a much greater proportion of sequence that is lost from smaller RNA inserts, such that the sequencing of RNA under 50 nt is prohibited.
[0064] Recently, the present inventors have developed a method allowing sequencing of RNA molecules having a length of < 50 ribonucleotides. Therefore, the present inventors have used a structured dumbbell 5 ’adapter that is ligated to the 5 ’end of the small RNA molecule and enables the extension of the direct RNA sequencing library so as to enable accurate nanopore sequencing of the full-length small RNA insert. In addition, the present inventors have developed a specific 3’adapter combination. Both, 5’adpater and 3’adapter combination can be designed for the capture of specific small RNAs via engineering of their overhangs with reverse complementary sequence to their target.
[0065] The term “adapter”, as used herein, refers to a polynucleotide that can be ligated to the 5’end of a target RNA (i.e. “5’adapter”) or to the 3’end of a target RNA (i.e. “3’adapter”). The nucleotides of the 5’adapter and 3’adapter may be standard or natural (i.e. adenosine, guanosine, cytidine, thymidine, and uridine) as well as non-standard nucleotides. Non-limiting examples of non-standard nucleotides include inosine, xanthosine, isoguanosine, isocytidine, diaminopyrimidine and deoxyuridine. The adapters may comprise modified or derivatized nucleotides. Non-limiting examples of modifications in the ribose or base moieties include the addition, or removal, of acetyl groups, amino groups, carboxyl groups, carboxymethyl groups, hydroxyl groups, methyl groups, phosphoryl groups and thiol groups. In particular, included are 2’-O-methyl and locked nucleic acids (LNA) nucleotides. Suitable examples of derivatized nucleotides include those with covalently attached dyes, such as fluorescent dyes or quenching dyes, or other molecules such as biotin, digoxygenin, or magnetic particles or microspheres. The adapters may also comprise synthetic nucleotide analogs such as morpholinos or peptide nucleic acids (PNA). Phosphodiester bonds or phosphothioate bonds may link the nucleotides or nucleotide analogs of the linkers.
[0066] In the context of the present invention, the term “5’adapter” refers to a polynucleotide that can be ligated to the 5’end of a target RNA (i.e. “5’adapter”). The 5’adapter, as described herein, is composed of ribonucleotides. The length of the 5’adapter, as described herein, can vary depending upon, for example, the desired length of the ligation product and the desired features of the adapter. In general, the 5’adapter may range from 15 to 70, e.g. 15, 16, 17, 18,
[0067] 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43,
[0068] 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68,
[0069] 69, or 70, ribonucleotides in length. Preferably, the 5’adapter comprises a 5 ’terminal ribonucleotide sequence comprising 6 to 15, e.g. 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15, ribonucleotides, wherein said 6 to 15 ribonucleotides are reverse complementary to a 5’- terminal sequence of a target RNA molecule.
[0070] The 5’adapter, as described herein, can be present as linear polynucleotide, e.g. after denaturation / when denatured. In this form, the 5’adapter is single-stranded. This primary structure may be converted into a secondary structure. Specifically, the 5’adapter, as described herein, is further capable of forming a stem-loop structure. Thus, the 5’adapter can also have a stem-loop structure, e.g. after re-naturation / when re-natured. As used herein, the term “stem-loop structure” refers to a pattern that can occur in singlestranded RNA. The structure is also known as a “hairpin” or “hairpin loop”. It occurs when two regions of the same strand, usually complementary in nucleotide sequence when read in opposite directions, base-pair to form a double helix that ends in an unpaired loop.
[0071] The 5 ’adapter, which is capable of forming a stem-loop structure, comprises a 5 ’positioned first stem sequence and a 3 ’positioned second stem sequence that are reverse complementary to each other. The first stem sequence and the second stem sequence form the “double-stranded region” or “double-stranded stem” of the stem-loop adapter.
[0072] In one embodiment, the stem of the 5’adapter is between 5 and 20, e.g. 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20, ribonucleotides in length. In one preferred embodiment, the stem of the 5’adapter is between 5 and 10, e.g. 5, 6, 7, 8, 9, or 10, ribonucleotides in length.
[0073] As used herein, the term “loop” refers to the single-stranded region of the stem-loop structure. In particular, the loop is located between the 5 ’positioned first stem sequence and the 3 ’positioned second stem sequence. In other words, the loop is located between the two reverse complementary strands of the stem and typically the loop comprises single-stranded ribonucleotides. In one embodiment, the loop sequence of the 5’adpater comprises between 10 and 40 ribonucleotides, e.g. 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 ribonucleotides. In one preferred embodiment, the loop sequence of the 5’adapter comprises between 12 and 22 ribonucleotides, e.g. 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, or 22 ribonucleotides.
[0074] In the context of the present invention, the term “3 ’adapter combination” refers to a combination of two single-stranded oligonucleotides, i.e. a first oligonucleotide and a second oligonucleotide. Both oligonucleotides comprise sections which are reverse complementary to each other so that they can form a hybrid / double-stranded structure. The length of the oligonucleotides of the 3 ’adapter combination, as described herein, can vary depending upon, for example, the desired length of the ligation product and the desired features of the adapter combination. In general, the oligonucleotides of the 3 ’adapter combination may range from 15 to 80, e.g. 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80, nucleotides in length. One of said two oligonucleotides can be ligated to the 3 ’end of a target RNA. This is possible as it is reverse complementary to the 3 ’-terminal sequence of a target RNA. Preferably, one of said two oligonucleotides of the 3 ’adapter combination comprises a 3 ’terminal deoxynucleotide sequence comprising 6 to 15, e.g. 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15, deoxynucleotide, wherein said 6 to 15 deoxynucleotide are reverse complementary to a 3 ’-terminal sequence of a target RNA molecule.
[0075] The two oligonucleotides of the 3 ’adapter combination, as described herein, can be present as linear oligonucleotides, e.g. after denaturation / when denatured. In this form, the 3 ’adapter combination is single-stranded. However, this primary structure may be converted into a secondary structure. Specifically, the two oligonucleotides of the 3 ’adapter combination, as described herein, are further capable of forming a hybrid (double-stranded) structure. Thus, the 3 ’adapter combination can also have a hybrid (double- stranded) structure, e.g. after re- naturation / when re-natured.
[0076] The 5’adapter and / or 3’adapter combination, as described herein, may comprise locked nucleic acids (LNAs).
[0077] The term “locked nucleic acids (LNAs)”, as used herein, refers to modified nucleotides, specifically deoxynucleotides or ribonucleotides. In case of locked ribonucleotides, the 2’-0 and 4’-C atoms of the ribose are joined through a methylene bridge. This additional bridge limits the flexibility normally associated with the ring, essentially locking the structure into a rigid conformation. These nucleic acid analogs are also referred to in some circles as “inaccessible ribonucleotides”. LNA nucleotides can be mixed with DNA or RNA residues in the polynucleotide, in effect hybridizing with DNA or RNA according to Watson-Crick basepairing rules. The inflexible nature of these molecules greatly enhances hybridization stability. Further, polynucleotides containing LNAs offer tremendous discriminatory power, allowing these molecules to distinguish between exact match and mismatched complementary target sequences with very little difficulty. In one embodiment, the 5’adapter and / or the 3’adapter comprise(s) locked nucleotides, in particular ribonucleotides or deoxynucleotides. In one preferred embodiment, the 5’positioned first stem sequence and / or the 3’positioned second stem sequence of the 5’adapter is (are) LNA enhanced. In one another preferred embodiment, the 5’positioned first stem sequence and / or the 3’positioned second stem sequence of the 3’adapter is (are) LNA enhanced.
[0078] The 5’adapter and 3’adapter combination, as described herein, allow to specifically capture small non-coding RNAs in a single ligation step reaction, thus, enabling direct RNA sequencing and quantification of RNA modifications on a nanopore sequencing platform.
[0079] The term “sample multiplexing (also known as multiplex sequencing)”, as used herein, refers to a technique in which nucleotide fragments from different samples are pooled and sequenced all together. The main reason is to increase sample throughput. The result is a mixture of sequencing reads from different samples. In a follow-up de-multiplexing step, the reads needs to be separated by using the attached barcode (sample marker) sequences.
[0080] The term “barcoding”, as used herein, refers to a method of specimen identification using short, standardized segments of nucleotides such as deoxynucleotides or ribonucleotides. Every species has its own barcode, just as every person has their own fingerprint. These DNA or RNA can be compared to a reference library to provide an ID.
[0081] The term “sample”, as used herein, refers to any sample comprising target RNA, particularly coding and / or non-coding, more particularly small non-coding, target RNA. Said sample is specifically derived from the body of a patient / subject. Said sample specifically comprises target RNA, in particular small non-coding target RNA, isolated from organisms, tissues, cells, or bodily fluids such as blood. Thus, the sample is particularly a biological sample.
[0082] The term “biological sample”, as used herein, refers to any sample having a biological origin and / or comprises biological material. The biological sample may be a body fluid sample, e.g. a blood sample or urine sample, or a tissue sample, e.g. a tissue biopsy sample. Biological samples may be mixed or pooled, e.g. a sample may be a mixture of a blood sample and a urine sample.
[0083] The term “body fluid sample”, as used herein, refers to any liquid sample comprising target RNAs. Said sample is specifically derived from the body of a patient / subject. Said body fluid sample may be a urine sample, blood sample, sputum sample, breast milk sample, cerebrospinal fluid (CSF) sample, cerumen (earwax) sample, gastric juice sample, mucus sample, lymph sample, endolymph fluid sample, perilymph fluid sample, peritoneal fluid sample, pleural fluid sample, saliva sample, sebum (skin oil) sample, semen sample, sweat sample, tears sample, cheek swab, vaginal secretion sample, liquid biopsy, or vomit sample including components or fractions thereof. The term “body fluid sample” also encompasses body fluid fractions, e.g. blood fractions, urine fractions or sputum fractions. Body fluid samples may be mixed or pooled. Thus, a body fluid sample may be a mixture of a blood and a urine sample or a mixture of a blood and cerebrospinal fluid sample.
[0084] The term “blood sample”, as used herein, encompasses whole blood or a blood fraction. Preferably, the blood fraction is selected from the group consisting of a blood cell fraction, plasma, and serum. In particular the blood fraction is selected from the group consisting of a blood cell fraction and plasma or serum. For example, the blood cell fraction encompasses erythrocytes, leukocytes, and / or thrombocytes.
[0085] The whole blood sample may be collected by means of a blood collection tube. It is, for example, collected in a PAXgene Blood RNA tube, in a Tempus Blood RNA tube, in an EDTA- tube, in a Na-citrate tube, Heparin-tube, or in a ACD-tube (Acid citrate dextrose). Alternatively, the whole blood sample may be collected in a blood collection tube containing cell-free nucleic acid stabilizing chemical agents, such as glutaraldehyde, formaldehyde, or similar (e.g. Streck cfRNA BCT tube, Streck cfDNA BCT tube), and others, or cellular crowding agents, such as polyethyleneglycol (PEG) (e.g. Norgen cfDNA / cfRNA preservation tube), and others.
[0086] The whole blood sample may also be collected by means of a bloodspot technique, e.g. using a Mitra Microsampling Device. This technique requires smaller sample volumes, typically 45-60 pl for humans or less. For example, the whole blood may be extracted from the patient via a finger prick with a needle or lancet. Thus, the whole blood sample may have the form of a blood drop. Said blood drop is then placed on an absorbent probe, e.g. a hydrophilic polymeric material such as cellulose, which is capable of absorbing the whole blood. Once sampling is complete, the blood spot is dried in air before transferring or mailing to labs for processing. Because the blood is dried, it is not considered hazardous. Thus, no special precautions need be taken in handling or shipping. Once at the analysis site, the desired components, e.g. miRNAs, are extracted from the dried blood spots into a supernatant which is then further analyzed.
[0087] The term “neural network”, as used herein, generally refers to a mathematical model (artificial neural network, ANN) composed of artificial neurons (nodes) interconnected by artificial synapses. The neural network may have one or more layers. A neural network may be a non-hierarchical type in apart thereof. A layer may be defined as a set of nodes. A layer may be defined to consist of a tensor-in tensor-out computation function (the layer's call method) and some state, held in TensorFlow variables (the layer's weights).
[0088] Embodiments of the invention
[0089] The present invention will now be further described. In the following passages, different aspects of the invention are defined in more detail. Each aspect so defined may be combined with any other aspect or aspects unless clearly indicated to the contrary. In particular, any feature indicated as being preferred or advantageous may be combined with any other feature or features indicated as being preferred or advantageous, unless clearly indicated to the contrary.
[0090] Covalent modifications of RNA molecules are pervasive of small non-coding RNAs and have fundamental roles in the regulation of several cellular processes. These RNA modifications are dynamic and their frequency at certain variable sites has been shown to be correlated with certain diseases such as cancer. Thus, detecting said RNA modifications would improve diagnostics. Nanopore direct RNA sequencing is a third-generation sequencing technology that allows the sequencing of native RNA molecules, thus, providing a direct way to detect modifications at single-molecule resolution. Despite recent advances, the analysis of nanopore sequencing data for RNA modification detection is still a complex task that presents many challenges. The Oxford Nanopore Technology (ONT) platform, for example, has traditionally been optimized for the sequencing of long nucleic acids, and technical limitations have prohibited the sequencing of shorter molecules under 100 nucleotides.
[0091] Recently, the present inventors have developed a method allowing targeted sequencing of small non-coding RNAs on a nanopore direct sequencing platform (e.g. the ONT platform) in order to simultaneously measure both RNA abundance and modification patterns. This was a milestone on the way to direct sequencing and analysing of small non-coding RNAs. However, technical limitations have hampered data analysis so far. To date, for example, no algorithm has been described for the basecalling and / or single molecule classification of 2'-O- methylated (Nm) RNA bases. The present inventors have now developed a method for the training of deep learning classification models for targeted small non-coding RNA biomarker sequencing and methylation classification on the Oxford Nanopore Technology (ONT) platform.
[0092] Thus, in a first aspect, the present invention relates to a method for classifying target RNA modifications in a neural network system. The method comprises the steps: a) obtaining direct RNA sequencing data relating to at least one target RNA (e.g. at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 target RNA(s)) in at least one sample (e.g. at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 50, 100, 150, 200, 300, 400, 500, or 1000 sample(s)), b) selecting a first portion of the direct RNA sequencing data, wherein the first portion of the direct RNA sequencing data comprises first direct RNA sequencing information being indicative for the presence or absence and / or being indicative of the pattern of at least one modification in the at least one target RNA, c) providing the first portion of the direct RNA sequencing data to a first neural network of the neural network system, the first neural network being trained to determine the presence or absence and / or the pattern of at least one modification in said target RNA from direct RNA sequencing data, and d) receiving a RNA modification prediction from the first neural network, the RNA modification prediction being indicative for the presence or absence and / or being indicative of the pattern of the at least one modification in the at least one target RNA in the at least one sample. The method of the first aspect therefore may lead to a RNA modification prediction, e.g. a prediction on the presence or absence of the at least one modification and / or on the pattern of the modification.
[0093] Modified RNA molecules are compounds which have been post-transcriptionally edited or modified. This process is also designated as RNA editing. RNA editing influences diverse RNA processes, including generation, transportation, function, and metabolization. Thus, RNA editing is a critical regulator of cell biology. Typical RNA modifications are selected from the group consisting of 2'-O-methylation, N6-methyladenosine (m6A), N6,2'-O- dimethyladenosine (m6Am), 8-oxo-7,8-dihydroguanosine (8-oxoG), pseudouridine ( ), 5- methylcytidine (m5C), and N4-acetylcytidine (ac4C). Preferably, the RNA modification is an RNA methylation.
[0094] Methylated RNA molecules are compounds which have been post-transcriptionally edited or modified by methylation. The methylation can occur at a base (e.g. methyl-6-adenine, pseudouridine) and / or ribose ring (2'-ortho-methylated nucleotide (2'-O-m)). The methylation of RNA occurring at a base is preferably selected from the group consisting of 6- methyladenosine (m6A), 5-methylcytidine (m5C), 5-methyluridine (m5U), 3 -methyluridine (m3U), 1 -methyladenosine (mfA), and 1 -methylguanosine (mfG), or is a combination thereof. The 2'-O-methylation of the backbone ribose is the most common and conserved type of RNA modification. The methylation of RNA occurring at a ribose ring is selected from the group consisting of 3 '-end 2'-O-methyladenosine (Am), 2'-O-methyluridine (Um), 2'-O- methylguanosine (Gm), and 2'-O-methylcytidine (Cm), or is a combination thereof.
[0095] The method of the first aspect has proven to be particularly advantageous in providing RNA modification predictions, as the first neural network is provided not with the entire direct RNA sequencing data relating to at least one target RNA in at least one sample, but only with a first portion of data, which comprises the first direct RNA sequencing information being relevant for providing RNA modifications prediction, namely being indicative for the presence or absence and / or being indicative of the pattern of at least one modification in the at least one target RNA.
[0096] The first neural network is therefore supplied with a smaller amount of data, which however is more relevant for the computation of the RNA modification prediction. This leads to quicker results, e.g. to a lower latency, and / or to an higher throughput without losing accuracy. It further minimizes the risk of errors in the computation of the neural networks due to irrelevant data in the data set, e.g. in the first portion of the direct RNA sequencing data. The method therefore has proven to be reliable by outputting highly accurate data (as will described with reference to the figures).
[0097] According to at least one embodiment, the target RNA comprises a 5'adapter ligated to the target RNA. The 5'adapater and its ligation to the target RNA will be described more in detail further below. The first direct RNA sequencing information may comprise information of the target RNA and of the 5'adapter, the information being indicative for the presence or absence and / or being indicative of the pattern of the at least one modification in the at least one target RNA in the at least one sample.
[0098] According to at least one embodiment, after step a) of the first aspect, e.g. after obtaining direct RNA sequencing data relating to at least one target RNA in at least one sample, the method further comprises the steps: i) segmenting the direct RNA sequencing data, ii-a) obtaining the first portion of the direct RNA sequencing data from the segmented direct RNA sequencing data.
[0099] The above steps i) and ii-a) may be performed after step a), e.g. after obtaining direct RNA sequencing data relating to at least one target RNA in at least one sample, and before step b) of the first aspect, e.g. after selecting a first portion of the direct RNA sequencing data.
[0100] The segmentation of the direct RNA sequencing may result in a segmentation of the RNA sequencing data into the first portion and a second portion. The first portion may comprise all the information which may be relevant for the RNA modification prediction. Through segmentation information which does not relate to the RNA modification is not comprised in the first portion of direct RNA sequencing data and therefore not fed to the first neural network.
[0101] According to at least one embodiment, after step a) of the first aspect, e.g. after obtaining direct RNA sequencing data relating to at least one target RNA in at least one sample, the method further comprises the steps of:
[0102] A) selecting a second portion of the direct RNA sequencing data, the second portion of the direct RNA sequencing data comprising second direct RNA sequencing information being indicative of a DNA barcode comprised in an adapter combination ligated to the at least one target RNA,
[0103] B) providing the second portion of the direct RNA sequencing data to a second neural network of the neural network system, the second neural network being trained to determine the DNA barcode,
[0104] C) receiving a barcode prediction from the second neural network, the barcode prediction being indicative of the DNA barcode. According to at least one embodiment the target RNA is ligated to a 3 'adapter combination. More details on the 3'adapter combination will be described further below. The second direct RNA sequencing information may comprise information of the 3'adapter, the information being indicative of a DNA barcode comprised in the 3'adapter combination.
[0105] The DNA barcode may be useful for sample multiplexing (also known as multiplex sequencing), which allows large numbers of libraries to be pooled and sequenced simultaneously during a single run on sequencing instruments. In other words in case of multiple samples, the method allows RNA modification prediction and DNA barcode prediction, with which the RNA modification prediction can be uniquely identified and mapped to the respective sample. The disclosure, however, is not limited to DNA barcodes and also RNA barcodes are envisaged. Thus, nucleotide barcodes in general such as DNA or RNA barcodes are possible.
[0106] The method according to the above embodiments has proven to be very advantageous for a variety of reasons.
[0107] Firstly using a first neural network and a second neural network specifically trained in a different manner for providing different predictions, e.g. the modification prediction and the barcode prediction, results in ameliorated overall predictions, as will become clear from the figures below showing prediction accuracy.
[0108] Secondly, the use of partitioned direct sequencing data, partitioned in its first portion and its second portion, the first portion used for the first neural network and the second portion used for the third neural network minimizes the latency of the respective first and second neural network as well as their respective accuracy. Using the whole data in a single neural network trained to provide both predictions, e.g. RNA modification predictions and DNA barcode predictions would results in a highly complex and expensive system with high latency and non- accurate results. Training two different neural networks in a neural network system and feeding the neural network only with the information which is important in the specific case, as done in the method according to the first aspect avoids these problems.
[0109] The segmentation step described above may comprise also the step ii-b) obtaining the second portion of the direct RNA sequencing data from the segmented direct RNA sequencing data.
[0110] With a single segmentation step, hence, both the first and the second portion of direct RNA sequencing data may be provided. The first portion may comprise information relative to the RNA modification prediction while the second portion may comprise information relative to the DNA barcode. In particular, the first portion of direct RNA sequencing data may comprise the first direct RNA sequencing information comprising information of the target RNA and of the 5'adapter while the second portion of direct sequencing data may comprise second direct RNA sequencing information comprising information of the 3'adapter combination. The direct RNA sequencing data may therefore be segmented, e.g. partitioned, in just two portion, e.g. the first portion comprising information of the target RNA and of the 5'adapter and the second portion comprising information on the 3'adapter combination.
[0111] The segmentation of the direct RNA sequencing data into the first portion of the direct sequencing data and the second portion of the direct sequencing data may comprise the step of determining the end of the 3'adapter combination (e.g. of the DNA barcode part). The end of the 3'adapter combination may be determined with a dynamic programming point detection approach according to Troung et al, 2020. This way of segmenting the data is advantageous as it is based on mostly mathematical considerations for segmenting the data and not directly to microbiological considerations, e.g. on microbiological information of the 3'adapter and / or the 5'adapter and / or on the target RNA sequence. Such a mathematical method has proven to provide information which are particularly compatible with the neural networks, e.g. with the first and / or the second neural networks, thereby improving the overall accuracy and latency.
[0112] The method of the first aspect may also comprise the step of associating the modification prediction and the barcode prediction to determine the at least one modification in the at least one target RNA in the at least one sample. This step permits to associate RNA modification prediction to the at least one target RNA in the at least one sample, thereby permitting the RNA modification prediction on multiple RNA targets in multiple samples. The method is therefore very well scalable to big data set of multiple samples (e.g. of multiple patients).
[0113] Once the data is associated, e.g. once the barcode prediction is associated to the RNA modification predictions the method may comprise the step of calculating a modification proportion for the at least one target RNA in the at least one sample.
[0114] The modification proportion is an indication on how many target RNA molecules exisit in a specific modification form as a percentage of all RNA target molecules in the at least one target RNA in the at least one target sample.
[0115] Especially, the RNA modification proportion refers to the modification percentage of the target RNA in a sample. The sample may be from a patient to be tested or from a control subject.
[0116] Specifically, the modification percentage / proportion of the target RNA ranges between 100% (i.e. fully modified, or all detected RNA entities are fully modified) and 0% (i.e. not modified, or all detected RNA entities are not modified). More specifically, the modification percentage / proportion of the target RNA is 0, 1, 2, 3, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61 ,62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100%.
[0117] Preferably, the target RNA modification proportion is a target RNA methylation proportion.
[0118] If the diagnosis of a disease or condition is desired, the target RNA modification proportion identified in a patient is compared to a target RNA modification proportion in a control subject (e.g. a healthy individual). This comparison then allows the diagnosis of a disease or condition.
[0119] Prior to selecting a first portion of the direct RNA sequencing data (and eventually the second portion of direct sequencing data), e.g. prior to segmenting the direct RNA sequencing data, the direct RNA sequencing data may be preprocessed in various way. Preprocessing the data (also called sometimes "cleaning the data") results in the elimination of outliers and anomalies which may have occurred either during the processes of the steps explained here above or e.g. during sample collection.. Preprocessing data furthermore may render the data more compatible for the neural networks in which the data is fed.
[0120] According to at least one embodiment the method of the first aspect comprises, after step a), e.g. after obtaining direct RNA sequencing data relating to at least one target RNA in at least one sample, and before step b) of the first aspect, e.g. after selecting a first portion of the direct RNA sequencing data, following (preprocessing) steps: basecalling the direct RNA sequencing data to obtain a predicted target RNA and to obtain quality control information, the quality control information comprising a quality control value indicative of the probability of a wrong basecalling, and filtering the data by disregarding data with a quality control value being below a quality control threshold value. The quality control threshold value may be an average quality score (PHRED) of seven. The disclosure is not limited to seven as a quality control threshold value, however it was proven, that by eliminating, e.g. disregarding data which has a PHRED score of lower 7, the accuracy of the prediction in the neural network system is very high (e.g. over 85%). A PHRED quality score may be a measure of the quality of the identification of the nucleobases generated by automated DNA sequencing.
[0121] Basecalling the direct RNA sequencing data may comprise processing direct RNA sequencing data with the guppy basecaller provided by Oxford Nanopore and the Nova lab rna_r9.4.1_70bps_sup model (Cruciani et al, 2023). The above described basecalling and filtering step are advantageous as they represent a first data clean-up step, which may facilitate the subsequent segmentation and processing in the neural networks.
[0122] According to at least one embodiment the method of the first aspect comprises, after step a), e.g. after obtaining direct RNA sequencing data relating to at least one target RNA in at least one sample, and before step b) of the first aspect, e.g. after selecting a first portion of the direct RNA sequencing data, following further (preprocessing) steps: transforming the direct RNA sequencing data into nucleotide sequence data of the target RNA, mapping the nucleotide sequence data of the target RNA against a reference target RNA, and disregarding the nucleotide sequence data which does not map to the reference target RNA.
[0123] The reference target RNA may also comprise a 5'adapter and a 3 'adapter combination. Mapping may be done using the Burrows-Wheeler Aligner. Disregarding nucleotide sequence data which does not map to the reference target RNA minimizes the provision of unrelated data in the method, e.g. in the neural networks, e.g. in the first and / or second neural network, and therefore results in more accurate predictions.
[0124] According to at least one embodiment, prior to providing the first portion of the direct RNA sequencing data and / or the second portion of the direct RNA sequencing data to the first neural network and / or to the second neural network, respectively, the method comprises: normalizing the first portion of the direct RNA sequencing data and / or the second portion of the direct RNA sequencing data, and / or scaling the first portion of the direct RNA sequencing data and / or the second portion of the direct RNA sequencing data to a fixed length, and / or removing outliers.
[0125] The three above steps may further "clean" and format the data for its ameliorated use in the neural network system. The normalization may for example be performed by Median Absolute Deviation (MAD). MAD measures the median deviation from the median value. It is a more robust measure of variability compared to the standard deviation. Normalizing with the MAD is a consistent estimation of the standard deviation.
[0126] MAD = median(\Xi — median X)\) xt— median xi
[0127] X~ 1.4826 * MAD Scaling to fix length may be performed through linear interpolation. The data may have different temporal granularity for each read but it may not be guaranteed to have measured the exact beginning and end of the RNA strand. The interpolation transforms the data into similar scale, thereby improving the use of the data in the neural network system.
[0128] The removal of outliers may be performed by clipping the data at three standard deviations from the median. The values at those positions may be replaced with an interpolation from a linear rolling window.
[0129] The first neural network and / or the second neural network may be a one dimensional (ID) residual neural network (ResNet). Residual Networks have proven to work very good in biological, e.g. microbiological, models.
[0130] The residual neural network might comprise eight residual blocks, each residual block comprising ID convolutional layers. Each ID convolutional layer may be followed by a batch normalization operation and ReLu activation (rectified linear unit activation). The channel width of the convolution may be 20 in a first block and increases in three steps to 67 channels in a last block of the eight residual blocks.
[0131] Before the first block one additional ID convolutional layer with 20 channels is arranged, the additional layer being followed by batch normalization operation, ReLU (rectified linear unit) activation and a pooling layer being configured to reduce the temporal size of the input of the neural network. After the last block an averaging operation may be performed wherein the averaging is performed over time.
[0132] The above architecture of the first and / or second neural network has proven to be advantageous with respect to the prediction of RNA modifications and barcode prediction as described above. It is also advantageous that both the first and the second neural network may be built with the same architecture. The disclosure, however, is neither limited to residual networks, nor to residual networks with this specific architecture. The first and the second neural networks may for example have different architecture with respect to each other.
[0133] Other Neural networks that may be considered are, e.g. Recurrent Neural Network (RNN), Long / Short Term Memory (LSTM), Bi-directional LSTM, WaveNet, Gated Recurrent Unit (GRU), Connectionist temporal classification (CTC), Auto Encoder (AE), Variational AE (VAE), Denoising Auto Encoder (DAE), Sparse AE (SAE), Temporal Convolutional Networks (TCN), Dilated Convolutional Neural Networks (DCNN), Spiking Neural Networks (SNN) have also proven to result in accurate predictions and are encompassed in this disclosure.
[0134] According to at least one embodiment, the method is a computer implemented method. According to at least one embodiment, the direct RNA sequencing data comprises ionic current data obtained by an ONT platform sequencers output of the target RNA. Different molecules can be sensed and potentially identified based on the modulation in ionic current as explained above.
[0135] In a second aspect, the present invention relates to a method for training a neural network system that is configured to classify target RNA modifications. The method comprising the steps of: a) selecting direct RNA sequencing training data, the direct RNA sequencing training comprising information on target training RNA; b) selecting a first portion of the direct RNA sequencing training data, the first portion of the direct RNA sequencing training data comprising first target training RNA information being indicative of a presence or absence and / or being indicative of the pattern of at least one modification in the target training RNA; c) training a first neural network, e.g. the first neural network of the first aspect, of the neural network system using the first portion of the direct RNA sequencing training data as input and the determination of the presence or absence and / or the determination of the pattern of at least one modification in the target training RNA as target.
[0136] According to at least one embodiment, after step a) the method further comprises the steps
[0137] A) selecting a second portion of the direct RNA sequencing training data, the second portion of the direct RNA sequencing training data comprising second target training RNA information being indicative for a DNA barcode comprised in an adapter combination ligated to the target training RNA;
[0138] B) training a second neural network of the neural network system using the second portion of the direct RNA sequencing training data as input and the determination of the DNA barcode as target.
[0139] Embodiments relating to the segmentation of the direct RNA sequencing data described for the first aspect also hold for the second aspect and for the direct RNA sequencing training data. Also the embodiments relating to preprocessing of the data, such as basecalling, filtering, mapping, normalizing, scaling, removing outliers, etc. described with respect to the first aspect may also hold for the second aspect, e.g. for training the neural networks, e.g. the first and / or the second neural network.
[0140] The direct RNA sequencing training data for training the first and / or second neural network comprises ground truth datasets based on synthetic small RNA biomarker targets and barcodes. These targets include in a specifc example described herein the ribosomal RNA fragment miLung#l (containing 2 modification sites, and four modification patterns; unmodified, double modified, and either site singly modified) and the miRNA cancer biomarker miR-17-5p (comprising unmodified and m6a modified miR-17-5p; miR-17-5p is known to be differentially m6A modified at position 13 in cancer (Konno et al, 2019)) for generation of ground truth data to train the first neural network. Each small RNA target synthetic library also contains one of up to 6 barcodes for the purpose of generation of ground truth data to train the second neural network.
[0141] In a third aspect, the present invention relates to a neural network system implemented on a computer, the system comprising a first neural network, e.g. the first neural network of the first aspect, and / or a second neural network, e.g. the second neural network of the first aspect, wherein the first neural network and / or the second neural network are trained according to the embodiments of the second aspect of the disclosure as described above.
[0142] In a fourth aspect, the present invention relates to a computer system adapted to classify target RNA modifications is provided. The computer system may comprise: at least one memory configured to store computer program code; at least one processor configured to access the computer program code and operate as instructed by the computer program code to perform the method according to the first and / or second aspect of the disclosure.
[0143] In a fifth aspect a computer program comprising instructions which, when the program is executed by a computer system, e.g. the computer system of the fourth aspect, causes the computer to carry out the method of according to the first and / or second aspect of the disclosure, is provided.
[0144] In a sixth aspect a non-transitory computer readable medium storing the computer programme of the fifth aspect is provided.
[0145] In the methods as described above, it is preferred that a 5’adapter comprising in the following order from 5’ to 3’ :
[0146] (i) a 5’terminal ribonucleotide sequence comprising 6 to 15 (e.g. 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15) ribonucleotides, wherein said 6 to 15 ribonucleotides are reverse complementary to a 5 ’-terminal sequence of a target RNA, and wherein optionally one or more (e.g. 1, 2, 4, or 5) of said 6 to 15 ribonucleotides are locked nucleotide- (LNA- ) enhanced, and
[0147] (ii) a ribonucleotide sequence capable of forming a stem-loop structure containing a loop and a double stranded stem is used.
[0148] The 5 ’adapter allows direct RNA sequencing of RNA targets, specifically non-coding small RNA targets. Particularly, the 5 ’adapter enables target RNA, specifically non-coding small RNA, sequencing via extension of the 5 ’end. In addition, the 5 ’adapter is sequence and structure specific as it only ligates to a specific target RNA sequence with a defined 5 ’end.
[0149] The locked nucleotides are specifically locked ribonucleotides. Examples of locked ribonucleotides are LNA-guanine, LNA-adenosine, or LNA-cytosine. For example, every, every second, every third, or every fourth ribonucleotide of the 5 ’terminal ribonucleotide sequence is a locked ribonucleotide. The LNAs increase the affinity of the 5 ’adapter.
[0150] The 5’adapter may range from 15 to 70, e.g. 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, or 70, nucleotides in length.
[0151] The 5’adapter may be present as linear polynucleotide, particularly in single-stranded form, e.g. after denaturation / when denatured. The 5’adapter is a polynucleotide that can be attached / ligated to the 5 ’end of a target RNA. When attached / ligated to the 5 ’end of a target RNA, the 5’adapter has a stem-loop structure. The attachment / ligation is possible as the 5’adapter comprises between 6 to 15, e.g. 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15, ribonucleotides which are reverse complementary to a 5 ’-terminal sequence of a target RNA.
[0152] In one embodiment, the ribonucleotide sequence capable of forming a stem-loop structure of the 5’adapter comprises a 5 ’positioned first stem sequence and a 3 ’positioned second stem sequence that are reverse complementary to each other. Thus, the 5 ’positioned first stem sequence and the 3 ’positioned second stem sequence can form the double stranded stem. The double stranded stem may have a length of between 5 and 20, e.g. 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20, ribonucleotides.
[0153] Preferably, each one of the 5 ’positioned first stem sequence and the 3 ’positioned second stem sequence has a length of between 5 to 10, e.g. 5, 6, 7, 8, 9, or 10, ribonucleotides. Particularly, the 5 ’positioned first stem sequence and the 3 ’positioned second stem sequence have the same length, e.g. a length of 5, 6, 7, 8, 9, or 10 ribonucleotides.
[0154] More preferably, the 5 ’positioned first stem sequence and / or the 3 ’positioned second stem sequence is (are) LNA enhanced. Particularly, the LNA enhanced sequence comprises between 2 to 5, e.g. 2, 3, 4, or 5, more particularly 3, locked nucleotides, specifically ribonucleotides. Examples of locked ribonucleotides are LNA-guanine, LNA-adenosine or LNA-cytosine. Even more preferably, the 5 ’positioned first stem sequence is LNA enhanced. Particularly, the LNA enhanced sequence comprises between 2 to 5, e.g. 2, 3, 4, or 5, more particularly 3, locked nucleotides, specifically ribonucleotides. Examples of locked ribonucleotides are LNA- guanine, LNA-adenosine or LNA-cytosine. Specifically, every, every second, or every third nucleotide may be LNA enhanced in the 5’positioned first stem sequence and / or the 3 ’positioned second stem sequence.
[0155] In one further embodiment, the nucleotide sequence capable of forming a stem-loop structure comprises a loop sequence which is located between the 5’positioned first stem sequence and the 3 ’positioned second stem sequence. The loop sequence may comprise between 10 and 40, e.g. 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40, ribonucleotides. Preferably, the loop sequence comprises between 12 and 22, e.g. 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, or 22, ribonucleotides.
[0156] In one another embodiment, the 5 ’-terminal sequence is configured such that it forms a single stranded 5 ’protrusion after formation of the stem-loop structure.
[0157] In one preferred embodiment, the 5 ’adapter comprises the following sequence from 5’ to 3’: (6-
[0158] 15x)rNrCrGrUrGrGrCrGrUrGrGrArGrUrGrUrUrArArUrUrArArUrGrUrGrCrUrUrUrGrCrCr ArUrG (SEQ ID NO: 1), wherein “r” stands for ribonucleotide, and wherein “(6-15x)rN” designates the ribonucleotide sequence reverse complementary to a 5 ’terminal sequence of a target RNA, or is a variant of this sequence.
[0159] The 5 ’adapter variant as described above has a sequence having at least 80%, preferably 85%, more preferably 90%, even more preferably 95%, and still even more preferably 99%, e.g. 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%, sequence identity to the sequence according to SEQ ID NO: 1. Such a 5’adapter variant still comprises ribonucleotides. In addition, such a 5’adapter variant is still capable of forming a stem-loop structure containing a loop and a double stranded stem. The skilled person can readily assess whether a 5’adapter variant is still capable of forming a stem-loop structure containing a loop and a double stranded stem. For example, the experimental section provides sufficient information in this respect.
[0160] Specifically, the 5’adapter allows (direct) nanopore sequencing of target RNA via the Oxford Nanopore Technology (ONT) platform.
[0161] In the methods as described above, it is preferred that a 3 ’adapter combination comprising (i) a first oligonucleotide and (ii) a second oligonucleotide, wherein (i) the first oligonucleotide comprises in the following order from 5’ to 3’: (a) a ribonucleotide sequence comprising between 4 and 8 (e.g. 4, 5, 6, 7, or 8) ribonucleotides, wherein the 5 ’terminal ribonucleotide is phosphorylated, which sequence is reverse complementary to a corresponding sequence in the second oligonucleotide,
[0162] (b) a poly(A) segment comprising between 8 and 12 (e.g. 8, 9, 10, 11, or 12) adenosine ribonucleotides,
[0163] (c) a deoxynucleotide spacer sequence comprising between 2 and 5 (e.g. 2, 3, 4, or 5) fixed deoxynucleotides, followed by optional 10 variable barcode deoxynucleotides, and further between 8 and 12 (e.g. 8, 9, 10, 11, or 12) fixed deoxynucleotides, all of which are reverse complementary to a corresponding sequence in the second oligonucleotide, and
[0164] (d) a 3 ’overhanging deoxynucleotide sequence comprising between 8 and 12 (e.g. 8, 9, 10, 11, or 12) deoxynucleotides which is capable of complexing with / ligating to a first element allowing nanopore sequencing of target RNA, and,
[0165] (ii) the second oligonucleotide comprises in the following order from 5’ to 3’ :
[0166] (a) a 5’overhanging deoxynucleotide sequence comprising between 15 and 25 (e.g. 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25) deoxynucleotides which is capable of complexing with / ligating to a second element allowing nanopore sequencing of target RNA,
[0167] (b) a deoxy nucleotide spacer sequence comprising between 8 and 12 (e.g. 8, 9, 10, 11, or 12) fixed deoxynucleotides, followed by optional 10 variable barcode deoxynucleotides, and further between 2 and 5 (e.g. 2, 3, 4, or 5) fixed deoxynucleotides, all of which are reverse complementary to a corresponding sequence in the first oligonucleotide,
[0168] (c) a poly(T) segment comprising between 8 and 12 (e.g. 8, 9, 10, 11, or 12) thymidine deoxynucleotides,
[0169] (d) a deoxynucleotide sequence comprising between 4 and 8 (e.g. 4, 5, 6, 7, or 8) deoxynucleotides which is reverse complementary to a corresponding sequence in the first oligonucleotide, and
[0170] (e) a 3 ’overhanging deoxynucleotide sequence comprising 6 to 15 (e.g. 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15) deoxynucleotides, wherein said 6 to 15 deoxynucleotides are reverse complementary to a 3 ’-terminal sequence of a target RNA, and wherein optionally one or more (e.g. 1, 2, 3, 4, or 5) of said 6 to 15 deoxynucleotides are locked nucleotide- (LNA-) enhanced is used.
[0171] The 3 ’adapter combination allows target RNA, specifically non-coding small RNA, sequencing. The 3’adpater combination is, in its renatured / hybridized state, a RNA / DNA hybrid polynucleotide. The inclusion of the RNA poly(A) / RNA poly(T) allows improved data processing, including segmentation and basecalling. The optional barcode deoxynucleotides enable multiplexed sequencing. They can also be designated as DNA barcode sequence.
[0172] The locked nucleotides are specifically locked deoxynucleotides. Examples of locked deoxynucleotides are LNA-guanine, LNA-thymidine, LNA-adenosine, or LNA-cytosine. For example, every, every second, every third, or every fourth nucleotide of the 3 ’overhanging nucleotide sequence is a locked deoxynucleotide. The LNAs increase annealing and ligation efficiency of the 3 ’adapter combination to the RNA target.
[0173] Thus, the 3 ’adapter combination comprises two oligonucleotides, a first oligonucleotide and a second oligonucleotide. In general, the oligonucleotides of the 3 ’adapter combination may range from 15 to 80, e.g. 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80, nucleotides in length. The second oligonucleotide of the two oligonucleotides can be ligated to the 3 ’end of a target RNA. This is possible as it is reverse complementary to the 3 ’-terminal sequence of a target RNA. Specifically, the second oligonucleotide of the two oligonucleotides of the 3’adapter combination comprises a 3 ’overhanging deoxynucleotide sequence comprising 6 to 15, e.g. 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15, deoxynucleotide, wherein said 6 to 15 deoxynucleotide are reverse complementary to a 3 ’-terminal sequence of a target RNA.
[0174] The two oligonucleotides of the 3’adapter combination, can be present as linear oligonucleotides, e.g. after denaturation / when denatured. In this form, the 3’adapter combination is single-stranded. This primary structure may be converted into a secondary structure. Specifically, the first and second oligonucleotides of the 3’adapter combination are capable of forming a hybrid (double-stranded) structure via their reverse complementary sequences. Thus, the 3’adapter combination can also have a hybrid (double-stranded) structure, e.g. after re-naturation / when re-natured. Particularly, when the second oligonucleotide of the 3’adapter combination is attached / ligated to the 3 ’end of a target RNA, the 3’adapter combination has a hybrid (double-stranded) structure.
[0175] In one preferred embodiment, the 3’overhanging deoxynucleotide sequence comprising between 8 and 12 (e.g. 8, 9, 10, 11, or 12) deoxynucleotides in (i)(d) is reverse complement to a corresponding 5 ’overhang sequence on a nanopore sequencing compatible motor protein / adapter complex, and / or the 5 ’overhanging deoxynucleotide sequence comprising between 15 and 25 deoxynucleotides in (ii)(a) is reverse complement to a tether oligonucleotide that concentrates the target RNA, specifically a ligation product or a RNA / cDNA duplex thereof (see below), at the membrane surface of a nanopore flow cell.
[0176] In one particularly preferred embodiment, the first element allowing nanopore sequencing of target RNA is a 5 ’overhang sequence on a nanopore sequencing compatible motor protein / adapter complex, and / or the second element allowing nanopore sequencing of target RNA is a tether oligonucleotide that concentrates the target RNA, specifically a ligation product or a RNA / cDNA duplex thereof (see below), at the membrane surface of a nanopore flow cell.
[0177] Specifically, the 3 ’adapter combination allows (direct) nanopore sequencing of target RNA via the Oxford Nanopore Technology (ONT) platform.
[0178] In one more preferred embodiment, the first oligonucleotide of the 3 ’adapter combination comprises the following sequence from 5’ to 3’: / 5Phos / rUrArGrCrUrC(8- / 2.N / GGGNNNNNNNNNNCTTGCTCTTAGGTAGTAGGTTC (SEQ ID NO: 2), wherein “r” stands for ribonucleotide, “ / 5Phos / ” indicates that the 5’-terminal ribonucleotide is phosphorylated, “(8-12x)rA” stands for the poly(A) segment, the deoxynucleotide spacer sequence is underlined, and the optional barcode (N) is highlighted in bold, or is a variant of this sequence.
[0179] The first oligonucleotide variant as described above has a sequence having at least 80%, preferably 85%, more preferably 90%, even more preferably 95%, and still even more preferably 99%, e.g. 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%, sequence identity to the sequence according to SEQ ID NO: 2. Such a first oligonucleotide variant still comprises the phosphorylated 5’-terminal ribonucleotide. Further, ribonucleotides in the first oligonucleotide are also ribonucleotides in the first oligonucleotide variant. Furthermore, the number of the fixed nucleotides in the first oligonucleotide is not amended in the first oligonucleotide variant. In addition, the poly(A) segment and / or the 3’overhanging deoxynucleotide sequence comprising between 8 and 12 (e.g. 8, 9, 10, 11, or 12) deoxynucleotides in (i)(d) which is reverse complement to a corresponding 5 ’overhang sequence on a nanopore sequencing compatible motor protein / adapter complex remain unchanged in the first oligonucleotide variant. Moreover, such a first oligonucleotide variant is still capable of forming a hybrid structure with the second oligonucleotide. Thus, amendments in the first oligonucleotide should also be included (in reverse complementary way) in the second oligonucleotide. The skilled person can readily assess whether the formation of a hybrid structure is still possible. For example, the experimental section provides sufficient information in this respect.
[0180] In one particularly more preferred embodiment, the first oligonucleotide of the 3 ’adapter combination comprises the following sequence from 5’ to 3’: / 5Phos / rUrArGrCrUrC GGNNNNNNNNNNCTTGCTCTTAGGTA GTAGGTTC (SEQ ID NO: 3), or is a variant of this sequence.
[0181] As to the variant language, it is referred to the above explanations.
[0182] In one another more preferred embodiment, the second oligonucleotide of the 3 ’adapter combination comprises the following sequence from 5’ to 3’: GAGGCGAGCGGTCAATTTTCCTAAGAGCAAGNNNNNNNNNNCC / 8- 72x)7GAGCTA(6-15x)N (SEQ ID NO: 4), wherein the deoxynucleotide spacer sequence is underlined, the optional barcode (N) is highlighted in bold, “(8-12x)T” stands for the poly(T) segment, and “(6-15x)N” designates the deoxynucleotide sequence reverse complementary to a 3 ’terminal sequence of a target RNA, or is a variant of this sequence.
[0183] The second oligonucleotide variant as described above has a sequence having at least 80%, preferably 85%, more preferably 90%, even more preferably 95%, and still even more preferably 99%, e.g. 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%, sequence identity to the sequence according to SEQ ID NO: 4.
[0184] Such a second oligonucleotide variant still comprises deoxynucleotides at positions where deoxynucleotides were previously present. Further, the number of the fixed nucleotides in the second oligonucleotide is not amended in the second oligonucleotide variant. In addition, the poly(T) segment and / or the 5’overhanging deoxynucleotide sequence comprising between 15 and 25 (e.g. 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25) deoxynucleotides in (ii)(a) which is reverse complement to a tether oligonucleotide that concentrates the target RNA, specifically a ligation product or a RNA / cDNA duplex thereof, at the membrane surface of a nanopore flow cell remain unchanged in the second oligonucleotide variant. Moreover, such a second oligonucleotide variant is still capable of forming a hybrid structure with the first oligonucleotide. Thus, amendments in the second oligonucleotide should also be included (in reverse complementary way) in the first oligonucleotide. The skilled person can readily assess whether the formation of a hybrid structure is still possible. For example, the experimental section provides sufficient information in this respect.
[0185] In one particularly more preferred embodiment, the second oligonucleotide of the 3 ’adapter combination comprises the following sequence from 5’ to 3’: GAGGCGAGCGGTCAATTTTCCTAAGAGCAAGNNNNNNNNNNCC7777777777GAG CTA(6-15x)N (SEQ ID NO: 5), or is a variant of this sequence.
[0186] As to the variant language it is referred to the above explanations.
[0187] The 5 ’adapter as described above and the 3 ’adapter combination as described above can be part of an adapter system.
[0188] The 5’adapter as described above and the 3’adapter combination as described above or the adapter system as described above allow the sequencing of target RNA present in a sample. The target RNA is preferably a small (non-coding) RNA.
[0189] Before the target RNA can be sequenced on a nanopore sequencing platform (e.g. the ONT platform), the 5’adapter as described above and the 3’adapter combination as described above have to be ligated to the target RNA, thereby producing a ligation product.
[0190] It is preferred that a method of ligating a 5’adapter and a 3’adapter combination to a target RNA (in a sample) comprises the steps of:
[0191] (i) providing a composition comprising a denatured target RNA (in a sample), the renatured 5’adapter as described above, and the renatured 3’adapter combination as described above, wherein the 5’adapter and the 3’adapter combination are annealed to the target RNA, and
[0192] (ii) ligating the 5’adapter and the 3’adapter combination to the target RNA using / with a double stranded RNA ligase, thereby producing a ligation product.
[0193] The annealing of the 5’adapter and 3’adapter combination to the target RNA requires that the target RNA is present in denatured form. In one embodiment, the denatured target RNA is produced by heating the target RNA at between 65°C and 75°C, e.g. 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, or 75°C, preferably at 70°C, for between 1 to 3 minutes, e.g. 1, 2, or 3, minutes, preferably for 2 minutes.
[0194] It is preferred that the target RNA is immediately placed on ice after denaturation. For the denaturation step, the target RNA is preferably given to an aqueous solution, e.g. water, or to a buffer solution.
[0195] The target RNA may, after its denaturation, treated with a Polynucleotide Kinase (PNK) which 5 ’phosphorylates RNA to enable ligation to the 3’ OH group at the 3’ end of the 5’ adapter. This PNK treatment is required for small RNAs that may not contain a 5’ phosphate such as for example fragments of rRNA or tRNA.
[0196] As to the 5 ’adapter, a denaturation and a renaturation step is required so that the adapter can form a stem-loop structure which allows annealing to the target RNA. As to the 3 ’adapter combination, a denaturation and a renaturation step is required so that the adapter combination can form a hybrid structure which allows annealing to the target RNA.
[0197] Annealing is a process of heating and cooling adapter / adapter combination with complementary sequences. Heat breaks all hydrogen bonds and cooling allows new bonds to form between the sequences. During this process, the adapter / adapter combination attaches to the denatured target RNA and forms their characteristic stem-loop structure / hybrid structure. In particular, the 5’adapter attaches to the 5 ’end / 5’ terminal sequence of the target RNA and the second oligonucleotide of the 3’adapter combination attaches to the 3 ’end / 3 ’terminal sequence of the target RNA. It is preferred that the adapter and adapter combination are denatured and renatured together, i.e. in a common reaction vessel. It is further preferred that the denaturation / renaturation of the adapter and adapter combination takes place separately and in the absence of the target RNA.
[0198] In one embodiment, the renatured 5’adapter is produced by denaturing the 5’adapter at between 75°C and 85°C, e.g. 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, or 85°C, preferably at 82°C, for between 1 to 3 minutes, e.g. 1, 2, or 3 minutes, preferably for 2 minutes, and renaturing the 5’adapter by cooling down to 4°C, preferably at a rate of 0. l°C / s.
[0199] In one additional or alternative embodiment, the renatured 3’adapter combination is produced by denaturing the 3’adapter combination at between 75°C and 85°C, e.g. 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, or 85°C, preferably at 82°C, for between 1 to 3 minutes, e.g. 1, 2, or 3 minutes, preferably for 2 minutes, and renaturing the 3’adapter combination by cooling down to 4°C, preferably at a rate of 0. l°C / s. For the denaturation and renaturation step, the 5’adapter and 3’adapter combination are preferably given to an aqueous buffer comprising 10 mM TRIS HC1 pH 7.5, 50 mM NaCl. In other words, the denaturing and renaturing of the 5’adapter and 3’adapter combination is preferably carried out in an aqueous buffer comprising 10 mM TRIS HC1 pH 7.5, 50 mM NaCl. The composition provided in step (i) of the above method is specifically produced by mixing the denatured target RNA, the renatured 5 ’adapter as described above, and the renatured 3 ’adapter combination as described above with each other, thereby annealing the 5 ’adapter and the 3 ’adapter combination to the target RNA.
[0200] The annealing of the 5 ’adapter with the target RNA particularly generates a doublestranded (DNA / RNA) hybrid containing a nick of RNA-OH-375’-P-RNA between the 3 ’end of the adapter and the 5’end of the target RNA. This is an efficient substrate for ligation by a double stranded RNA ligase. In addition, the annealing of the second oligonucleotide of the 3’adpater combination with the target RNA particularly generates a double-stranded (DNA / RNA) hybrid containing a nick of RNA-OH-375’-P-RNA between the 3’end of the target RNA and the 5’end of the second oligonucleotide of the 3 ’adapter combination. This is a substrate for ligation by a double stranded RNA ligase.
[0201] The ligation is usually carried out in a ligation buffer. An exemplarily ligation buffers is described in the experimental section of the present patent application. In one preferred embodiment, the ligation buffer comprises polyethylene glycol (PEG), e.g. PEG 8000 (5%), and / or adenosine triphosphate (ATP), e.g. 1 mM ATP.
[0202] In one embodiment, the ligation is carried out between 36°C and 38°C, e.g. 36, 37, or 38°C, preferably at 37°C, for between 30 minutes and 1.5 hours, e.g. 30, 35, 40, 45, 50, 55 minutes, 1, 1.25, or 1.5 hour(s), preferably for 1 hour, then at between 14°C and 18°C, e.g. 14, 15, 16, 17, or 18°C, preferably at 16°C, for between 1.5 hours and 2.5 hours, e.g. 1.5, 2, or 2.5 hours, preferably for 2 hours, and at 12°C overnight.
[0203] The double stranded RNA ligase can be any ligase capable of ligating double stranded RNA nicks / RNA structures. Preferably, the double stranded RNA ligase is a T4 RNA ligase 2 (Rnl2) or a Kodl ligase.
[0204] In this respect, it should be noted that only a perfectly hybridized molecule provides a substrate for the double stranded RNA ligase, in particular Rnl2. Also, in case of protrusion of either strand or a gap that is 2 nucleotides or longer, the Rnl2 will ligate the molecule with much lower efficiency.
[0205] By ligating the 5 ’adapter and 3 ’adapter combination to the target RNA using / with a double stranded RNA ligase, a ligation product is produced. The ligation product can be described as a hybrid molecule comprising at least one adapter / adapter combination and a target RNA. For example, the ligation product may comprise a 5 ’adapter and a target RNA such as miRNA or isomiR. The ligation product may comprise a 3 ’adapter combination and a target RNA such as miRNA or isomiR. In addition, the ligation produced may comprise a 5 ’adapter, a 3 ’adapter combination and a target RNA such as miRNA or isomiR.
[0206] Optionally, the ligation product obtained with the above method is reverse transcribed. By the reverse transcription, a RNA / cDNA duplex is obtained. Particularly, the reverse transcription of the ligation product is carried out by reverse transcribing the ligation product using a reverse transcriptase (RT). Usually, the reverse transcriptase (RT) uses a RT primer. Herein, however, the reverse transcriptase (RT) uses the second oligonucleotide of the 3 ’adapter combination as a self-primer to extend. The reverse transcriptase (RT) may be a SuperScript III RT, Maxima H-RT or Tth polymerase. Preferably, said reverse transcribing is carried out at between 45°C and 60°C, e.g. 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60°C, preferably at 50°C, for between 40 and 60 minutes, e.g. 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 minutes, preferably for 50 minutes, then at between 60°C and 80°C, e.g. 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80°C, preferably at 70°C, for between 8 and 15 minutes, e.g. 8, 9, 10, 11, 12, 13, 14, or 15 minutes, preferably for 10 minutes, before finally cooling down to 4°C. The result of the reverse transcription reaction is the formation of RNA / cDNA duplex. This duplex is composed of a top strand comprising the 5’ adapter, small RNA insert(s), and the first oligonucleotide of the 3’ adapter, and a bottom strand comprising the second oligonucleotide of the 3’ adapter that has been extended as cDNA by reverse transcription.
[0207] In a subsequent method the target RNA (in a sample) is determined and / or quantified. A method of determining and / or quantifying a target RNA (in a sample) preferably comprises the step of:
[0208] (i) subjecting the ligation product obtained as described above or the RNA / cDNA duplex obtained as described above to a nanopore sequencing reaction, thereby determining and / or quantifying the target RNA.
[0209] In this method, the sequencing can be determined as targeted sequencing.
[0210] Nanopore sequencing requires flow cells. Flow cells contain an array of tiny holes - nanopores - embedded in an electro-resistant membrane. The nanopore sequencing reaction particularly further requires / comprises the addition of a structure allowing nanopore sequencing.
[0211] Specifically, the structure allowing nanopore sequencing is composed of
[0212] (i) a first element, and
[0213] (ii) a second element.
[0214] More specifically, (i) the first element is a 5 ’overhang sequence on a nanopore sequencing compatible motor protein / adapter complex, and
[0215] (ii) the second element is a tether oligonucleotide that concentrates the ligation product (in case no reverse transcription is carried out) or the RNA / cDNA duplex (in case reverse transcription is carried out) at the membrane surface of a nanopore flow cell.
[0216] The 5 ’overhang sequence on a nanopore sequencing compatible motor protein / adapter complex, thus, established the connection / contact of the 5 ’adapter / 3’ adapter combination complex carrying the target RNA to be sequenced and the motor protein / adapter complex. Even more specifically,
[0217] (i) the first element recognizes the reverse complementary 3 ’overhanging deoxynucleotide sequence of the first oligonucleotide of the 3 ’adapter combination as described above, to which it anneals and is ligated to, to enable nanopore sequencing, and
[0218] (ii) the second element recognizes the reverse complementary 5 ’overhanging deoxynucleotide sequence of the second oligonucleotide of the 3 ’adapter combination as described above, to which it anneals and is ligated to, to enable nanopore sequencing. Thus, the nanopore sequencing system uses, in addition to flow cells including nanopores and an electro-resistant membrane, a nanopore adapter / motor protein complex and a tether oligonucleotide. The nanopore adapter is required for attaching the target RNA to be sequenced, specifically the ligation product or the RNA / cDNA duplex, to the nanopore. In addition, the motor protein is required for directing the target RNA to be sequence through the nanopore in order to allow sequencing. Specifically, the motor protein controls translocation of the target RNA strand through the nanopore. Once the target RNA has passed thought, the motor protein detaches and the nanopore is ready to accept the next target RNA. The tether oligonucleotide has the function of concentrating the RNA target which is to be sequenced at the membrane surface of a nanopore flow cell. An electrically resistant membrane means that all current must pass through the nanopore to ensure a clean signal. The final sequencing of the target RNA is possible as, when the target RNA passes through the nano-scale hole, the current changes / fluctuates. This signal can be detected and is converted to a nucleotide sequence by basecalling algorithms.
[0219] Especially, the added structure allows (direct) nanopore sequencing of a target RNA via the Oxford Nanopore Technology (ONT) platform.
[0220] The nanopore sequencing reaction preferably further allows the detection, characterization, and quantification of nucleotide modifications such as methylations. The nucleotide modifications are preferably selected from the group consisting of 2'-O-methylation, N6-methyladenosine (m6A), N6,2'-O-dimethyladenosine (m6Am), 8-oxo-7,8- dihydroguanosine (8-oxoG), pseudouridine ( ), 5 -methyl cytidine (m5C), and N4- acetylcytidine (ac4C). The detection, characterization, and quantification of nucleotide modifications is specifically carried out through analysis of the raw nanopore signal. Specific modifications will characteristically alter the current associated with a given nucleotide and enable their detection and quantification.
[0221] It is preferred that the target RNA as described above is non-coding RNA, specifically non-coding small RNA. It is more preferred that the (non-coding) target RNA has a length of < 50 ribonucleotides, particularly a length of between 10 and < 50 ribonucleotides, e.g. a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, or 49 ribonucleotides. It is even more preferred that the (non-coding) target RNA has a length of < 30 ribonucleotides, particularly a length of between 10 and < 30 ribonucleotides, e.g. a length of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, or 29 ribonucleotides.
[0222] The (non-coding) target RNA, specifically (non-coding) small RNA, may be a miRNA, piRNA, siRNA, or a fragment of longer RNAs, preferably a rRNA or tRNA fragment.
[0223] It is further preferred that the sample as described above is a biological sample. The biological sample may be any sample having a biological origin. For example, the biological sample may be a body fluid sample, e.g. a blood sample or urine sample, or a tissue sample, e.g. a tissue biopsy sample. Biological samples may be mixed or pooled, e.g. a sample may be a mixture of a blood sample and a urine sample.
[0224] The body fluid sample may be a urine sample, blood sample, sputum sample, breast milk sample, cerebrospinal fluid (CSF) sample, cerumen (earwax) sample, gastric juice sample, mucus sample, lymph sample, endolymph fluid sample, perilymph fluid sample, peritoneal fluid sample, pleural fluid sample, saliva sample, sebum (skin oil) sample, semen sample, sweat sample, tears sample, cheek swab, vaginal secretion sample, liquid biopsy, or vomit sample including components or fractions thereof. The term “body fluid sample” also encompasses body fluid fractions, e.g. blood fractions, urine fractions or sputum fractions. Body fluid samples may be mixed or pooled. Thus, a body fluid sample may be a mixture of a blood and a urine sample or a mixture of a blood and cerebrospinal fluid sample.
[0225] More preferably, the biological sample is a blood sample. Even more preferably, the blood sample is a whole blood or a blood fraction, preferably blood cells (e.g. erythrocytes, leukocytes, and / or thrombocytes), serum, or plasma. For example, the blood cell fraction encompasses erythrocytes, leukocytes, and / or thrombocytes. The whole blood sample may be collected by means of a blood collection tube. It is, for example, collected in a PAXgene Blood RNA tube, in a Tempus Blood RNA tube, in an EDTA-tube, in a Na-citrate tube, Heparin-tube, or in a ACD-tube (Acid citrate dextrose). The whole blood sample may also be collected by means of a bloodspot technique, e.g. using a Mitra Microsampling Device. This technique requires smaller sample volumes, typically 45-60 pl for humans or less. For example, the whole blood may be extracted from the patient via a finger prick with a needle or lancet. Thus, the whole blood sample may have the form of a blood drop. Said blood drop is then placed on an absorbent probe, e.g. a hydrophilic polymeric material such as cellulose, which is capable of absorbing the whole blood. Once sampling is complete, the blood spot is dried in air before transferring or mailing to labs for processing. Because the blood is dried, it is not considered hazardous. Thus, no special precautions need be taken in handling or shipping. Once at the analysis site, the desired components, e.g. miRNAs, are extracted from the dried blood spots into a supernatant which is then further analyzed.
[0226] Various modifications and variations of the invention will be apparent to those skilled in the art without departing from the scope of invention. Although the invention has been described in connection with specific preferred embodiments, it should be understood that the invention as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the invention which are obvious to those skilled in the art in the relevant fields are intended to be covered by the present invention.
[0227] BRIEF DESCRIPTION OF THE FIGURES
[0228] The following Figures are merely illustrative of the present invention and should not be construed to limit the scope of the invention as indicated by the appended claims in any way.
[0229] Figure 1: Shows the principle of nanopore sequencing on the platform of Oxford Nanopore Technologies (ONT). A MinlON flow cell contains 512 channels with 4 nanopores in each channel, for a total of 2,048 nanopores used to sequence DNA or RNA. The wells are inserted into an electrically resistant polymer membrane supported by an array of microscaffolds connected to a sensor chip. Each channel associates with a separate electrode in the sensor chip and is controlled and measured individually by the application-specific integration circuit (ASIC). Ionic current passes through the nanopore because a constant voltage is applied across the membrane, where the trans side is positively charged. Under the control of a motor protein, a double-stranded DNA (dsDNA) molecule (or an RNA-DNA hybrid duplex) is first unwound, then single-stranded DNA or RNA with negative charge is ratcheted through the nanopore, driven by the voltage. As nucleotides pass through the nanopore, a characteristic current change is measured and is used to determine the corresponding nucleotide type at -450 bases per s (R9.4 nanopore).
[0230] Figure 2: Shows a schematic diagram of the library preparation for small RNA direct RNA sequencing on the ONT platform. Ligation of both a structured 5’ adapter, and 3’ adapter in the 1ststep, enclose the small RNA in adapter sequence and enable accurate sequencing and RNA modification detection and quantification.
[0231] Figure 3: Shows a schematic diagram of the training and inference process for barcode demultiplexing and RNA modification quantification with Nano-smallRNAs.
[0232] Figure 4: Shows a schematic structure of a 5’adapter as used herein.
[0233] Figure 5: Shows a schematic structure of a 3 ’adapter combination as used herein.
[0234] Figure 6: Shows confusion matrices for trained deep learning RNA modification and barcode demultiplexing models, a. miLung#l modification status, b. miR-17-5p modification status, c. 3’ adapter barcode classification.
[0235] Figure 7: Shows the architecture of the networks used for predicting barcodes and methylation status. The two models have the exact same architecture and number of parameters, except the number of output neurons of the linear classifier. The network consists of four bottleneck blocks each consisting of three convolutions at the same channel depth before a channel deepening convolution which again is followed by three convolutions of the same depth. One extra convolution after the input, sets up the initial 20 channels, and the final classification is performed by a fully connected layer. All convolutions are ID and followed by batch normalisation.
[0236] EXAMPLES
[0237] The examples given below are for illustrative purposes only and do not limit the invention described above in any way.
[0238] 1. 5’adapter and 3 ’adapter combination
[0239] Nanopore sequencing is usually performed as shown in Figure 1. The Nano- smallRNAseq method as described herein allows sequencing of small non-coding target RNA (see Figure 2). Specifically, a structured dumbbell 5 ’adapter is used that is ligated to the 5 ’end of the small RNA and enables the extension of the 5 ’end of the direct RNA sequencing library so as to enable accurate nanopore sequencing of the full-length small RNA insert (see Figure 4). The structured dumbbell 5 ’ adapter is engineered with a 3 ’ overhang complementary to a given target for targeted sequencing. The overhang region may also contain LNA modified nucleotides to improve affinity to small RNAs and, therefore, ligation efficiency. The structured stem is also essential to impart structural specificity such that the adapter will only ligate to the free 5’end of a small RNA and not internally within a longer sequence.
[0240] The 3 ’adapter combination differs from the original ONT 3 ’adapter (Figure 5). The modifications include the addition of LNA modified nucleotides to the overhang of the second oligonucleotide of the 3 ’adapter combination, designed to capture the insert RNA with greater affinity. A poly(A) RNA tract has also been inserted into the first oligonucleotide and a poly(T) RNA tract has also been inserted into the second oligonucleotide of the 3 ’adapter combination to improve the data segmentation and basecalling for insert RNAs that lack a poly(A) tail, such as small RNAs. Finally, a DNA barcode is encoded within the 3 ’adapter combination to enable multiplexed direct RNA sequencing on the ONT platform.
[0241] 2, Direct RNA sequencing protocol with the 5 ’adapter and 3 ’adapter combination
[0242] Follow ONT instructions for Direct RNA sequencing protocol with the following key modifications:
[0243] 1. RNA PNK pre-treatment.
[0244] 1.1.In a 0.2 ml PCR tube, prepare a maximum of 1 mg of total RNA and adjust the volume to 16 pl with Nuclease-free Water.
[0245] 1.2.Place the tube in a thermal cycler and incubate at 70°C for 2 min. Immediately place on ice.
[0246] 1.3. In a new 0.2 ml PCR tube add 15.6 pl denatured RNA from the previous step and the following reagents: Place the tube(s) in a thermal cycler and incubate at 37°C for 30 min. Heat inactivate by incubating at 65°C for 20 minutes. Immediately place on ice.ing 5’ and 3’ adapters for RNA ligation: Prepare the Annealing buffer (10 mM Tris-HCl pH 7.5, 50 mM NaCl). Prepare the 3’ adapter by annealing 1.4 pM Oligo A and Oligo B in a 1 : 1 ratio in Annealing buffer. Resuspend and dilute the 5’ adapter in Annealing buffer to result in a 2.5 pM working concentration. Place the 3’ adapter and the 5’ adapter in the thermal cycler and incubate at 82°C for 2 min followed by a ramp down 0. l°C / sec to 4°C. Store in -20°C. NA Ligation: Mix by pipetting the following reagents together to make the RNA ligation master mix: If preparing more than one sample (for multiplexing), mix in different 0.2 ml PCR tubes 10 pl of the previous Master mix, 9 pl PNK treated RNA and 1 pl of the corresponding barcoded 3’ adapter (1.4 pM). Mix by pipetting. Place the tube(s) in a thermal cycler and incubate at 37°C for 60 min and the lid heated at 45°C, then at 16°C for 2 h and 12°C overnight. e Transcription: 4.1.Use the 20 pl of the previous adapter-ligated RNA to continue with the Reverse transcription:
[0247] 4.2. Place the tube(s) in a thermal cycler and incubate at 50°C for 50 min, then at 70°C for 10 min . eat inactivate the T4 Rnl2 at 80°C for 5 minutes and 4°C until the next step beads clean-up with Agencourt RNAClean XP Transfer the product to a 1.5 ml Eppendorf DNA LoBind tube. If preparing more than one sample for multiplexing, in this step, all the samples can be mixed homogeneously in a 1.5 ml Eppendorf DNA LoBind tube. Resuspend the stock of beads by vortexing. Adjust the ratio beads:RNA to 2x (example: for 40 pl of reverse transcription reaction, add 80 pl of resuspended beads) and mix by pipetting. Incubate on a Hula mixer (rotator mixer) for 5 minutes at room temperature. Prepare 500 pl of fresh 70%. Spin down the sample and pellet on a magnet. Wait until the supernatant is transparent (approx. 7-10 min) and pipette it off. Keep the tube on magnet and add 400 pl of 70% ethanol without disturbing the pellet. Keep the magnetic rack on the bench, rotate the bead-containing tube by 180°. Wait for the beads to migrate and then rotate the tube back to the starting position. Wait for the beads to migrate back and remove the ethanol. Spin down and place the tubes back on the magnet. Pipette off any residual ethanol. Remove the tube from the magnetic rack and resuspend the pellet in 21 pl of Nuclease-free water. Incubate 5 minutes at room temperature. Pellet the beads on a magnet until eluate is clear and colorless. Pipette 20 pl of eluate into a clean 1.5 ml Eppendorf DNA LoBind tube. cond RNA Ligation: RNA Adapters (RMX): dd the next reagents in the following order: cubate the reaction for 10 to 20 minutes at room temperature. d beads clean-up with Agencourt RNAClean XP Resuspend the stock of beads by vortexing Adjust the ratio beads:RNA to 0.4x (example: for 40 pl of ligation product, add 16 pl of resuspended beads) and mix by pipetting. Incubate on a Hula mixer (rotator mixer) for 5 minutes at room temperature. Spin down the sample and pellet on a magnet. Wait until the supernatant is transparent (approx. 7-10 min) and pipette it off Keep the tube on magnet and add 150 pl of Wash Buffer (WSB) to the beads. Close the tube lid and resuspend by flicking the tube. Short spin and return the tube to the magnetic rack. Wait for the beads to migrate and pipette off the supernatant. Repeat the previous step. . Remove the tube from the magnetic rack. . Resuspend the pellet in 23 pl of Elution Buffer (EB) by the gentle flicking the tube. . Incubate 10 minutes at room temperature. . Pellet the beads on a magnet until eluate is clear and colorless. . Pipette 20 pl of eluate into a clean 1.5 ml Eppendorf DNA LoBind tube. ing and loading the SpotON flow cell Prior to loading the reverse-transcribed and adapted RNA, the flow cell needs to be prepared. Add 30 pl of Flash Teher (FLT) in a full aliquote of Flush Buffer (FB). Mix by vortexing and spin down. This solution is the priming mix. Open the MinlON device lid and slide the flow cell under the clip. Press down firmly on the flow cell to ensure correct thermal and electrical contact. Slide the flow cell priming port cover clockwise to open the priming port. After opening the priming port, set a Pl 000 pipette to 200 pl and insert the tip into the priming port. Turn the wheel of the pipette until the dial shows 220-230 pl or until you see a small volume of buffer entering the pipette tip. Note: Viually check that there is continuous buffer from the priming port across the sensor array. Load 800 pl of the priming mix into the flow cell via the priming port, avoid the introduction of air bubbles. Wait for five minutes. During this time, prepare the library for loading by following the next steps: Complete the flow cell priming by gently lifting the SpotON sample port cover and loading 200 pl of the priming mix into the flow cell priming port (not the SpotON sample port), avoiding the introduction of air bubbles. Mix the prepared library gently by pipetting up and down and directly add the 75 pl of library to the Flow Cell vi the SpotON sample port in a dropwise fashion (not touching the port with the pipette tip). Ensure each drop flows into the port before adding the next. Genlty replace the SpotON sample port cover, making sure the bung enters the SpotON port, close the priming port and replace the MinlON device lid. Data analysis and bioinformatics: Following the flowchart detailed in Figure 3.
[0248] The left column of Figure 3 shows an exemplary embodiment of a method for training the neural network system. The method may comprise following steps:
[0249] Step 10: obtaining direct RNA sequencing training data relating to at least one target RNA in at least one sample, e.g. "raw data".
[0250] Step 12: basecalling the direct RNA sequencing training data to obtain a predicted target RNA and to obtain quality control (QC) information, the quality control information comprising a quality control value indicative of the probability of a wrong basecalling.
[0251] Step 14: filtering the data by disregarding data with a quality control value being below a quality control threshold value, the quality control threshold being PHRED lower 7.
[0252] Step 16: "filtering with alignment", which comprise: transforming the direct RNA sequencing training data into nucleotide sequence data of the target RNA; and mapping the nucleotide sequence data of the target RNA against a reference target RNA; and disregarding the nucleotide sequence data which does not map to the reference target RNA.
[0253] Step 18: segmenting the direct RNA sequencing training data.
[0254] Step 20: obtaining the first portion of the direct RNA sequencing training data from the segmented direct RNA sequencing training data (marked with "RNA" in the Figure), e.g. derived from RNA with the 5 'adapter, and obtaining the second portion of the direct RNA sequencing training data from the segmented direct RNA sequencing training data (marked with" DNA" in the Figure), e.g. derived from the 3'adapter combination.
[0255] Step 22: normalizing the first portion of the direct RNA sequencing training data and the second portion of the direct RNA sequencing training data, through Median Absolute Deviation (MAD). Step 24: scaling the first portion of the direct RNA sequencing data and the second portion of the direct RNA sequencing data to a fixed length.
[0256] Step 26: removing outliers, wherein the removal of outliers is performed by clipping the data at three standard deviations from the median and wherein the values at those positions are replaced with an interpolation from a linear rolling window.
[0257] Step 28: selecting the first portion of the direct RNA sequencing training data, the first portion of the direct RNA sequencing training data comprising first target training RNA information being indicative of a presence or absence and / or being indicative of the pattern of at least one modification in the target training RNA; and selecting the second portion of the direct RNA sequencing training data, the second portion of the direct RNA sequencing training data comprising second target training RNA information being indicative for a DNA barcode comprised in an adapter combination, e.g. in the 3'adapter combination, ligated to the target training RNA;
[0258] Step 30: training a first neural network of the neural network system using the first portion of the direct RNA sequencing training data as input and the determination of the presence or absence and / or being indicative of the pattern of at least one modification in the target training RNA as target, and training a second neural network of the neural network system using the second portion of the direct RNA sequencing training data as input and the determination of the barcode as target.
[0259] The trained first and second neural network may then be used for the classification of RNA modifications.
[0260] The right column of Figure 3 shows an exemplary embodiment of a method for classifying target RNA modifications in the neural network system. The method may comprise following steps:
[0261] Step 110: obtaining direct RNA sequencing data relating to at least one target RNA in at least one sample, e.g. "raw data.
[0262] Step 112: basecalling the direct RNA sequencing data to obtain a predicted target RNA and to obtain quality control (QC) information, the quality control information comprising a quality control value indicative of the probability of a wrong basecalling. Step 114: filtering the data by disregarding data with a quality control value being below a quality control threshold value, the quality control threshold being PHRED lower 7.
[0263] Step 116: "filtering with alignment", which comprise: transforming the direct RNA sequencing data into nucleotide sequence data of the target RNA; and mapping the nucleotide sequence data of the target RNA against a reference target RNA; and disregarding the nucleotide sequence data which does not map to the reference target RNA.
[0264] Step 118: segmenting the direct RNA sequencing data,
[0265] Step 120: obtaining the first portion of the direct RNA sequencing data from the segmented direct RNA sequencing data (marked with "RNA" in Figure 3), and obtaining the second portion of the direct RNA sequencing data from the segmented direct RNA sequencing data (marked with" DNA" in Figure 3).
[0266] Step 122: normalizing the first portion of the direct RNA sequencing data and the second portion of the direct RNA sequencing data, through Median Absolute Deviation (MAD).
[0267] Step 124: scaling the first portion of the direct RNA sequencing data and the second portion of the direct RNA sequencing data to a fixed length.
[0268] Step 126: removing outliers, wherein the removal of outliers is performed by clipping the data at three standard deviations from the median and wherein the values at those positions are replaced with an interpolation from a linear rolling window.
[0269] Step 128: selecting the first portion of the direct RNA sequencing data, wherein the first portion of the direct RNA sequencing data comprises first direct RNA sequencing information being indicative for the presence or absence and / or being indicative of the pattern of at least one modification in the at least one target RNA, and providing the first portion of the direct RNA sequencing data to a first neural network of the neural network system, the first neural network being trained to determine the presence or absence and / or the pattern of at least one modification in said target RNA from direct RNA sequencing data, e.g. according to the method of training described in the left column (e.g. steps 10 to 30), and receiving a RNA modification prediction from the first neural network, the RNA modification prediction being indicative for the presence or absence and / or being indicative of the pattern of the at least one modification in the at least one target RNA in the at least one sample; and selecting the second portion of the direct RNA sequencing data, the second portion of the direct RNA sequencing data comprising second direct RNA sequencing information being indicative of a DNA barcode comprised in an adapter combination ligated to the at least one target RNA, providing the second portion of the direct RNA sequencing data to a second neural network of the neural network system, the second neural network being trained to determine the DNA barcode, e.g. according to the method of training described in the left column (e.g. steps 10 to 30), receiving a barcode prediction from the second neural network, the barcode prediction being indicative of the DNA barcode.
[0270] Step 130: if samples have been multiplexed, the inferred barcode is used to divide target RNA reads into their respective samples. The modification proportion is then calculated per target RNA per sample by dividing the number of reads inferred to contain a specific modification pattern (eg. Double methylated miLung#l, GmCm) by the total number of miLung#l reads from that sample * 100. This is performed per target RNA form (eg. for miLung#l, methylation proportions are calculated for GC, GmC, GCm, GmCm).
[0271] 10. Established confusion matrices for trained deep learning RNA modification:
[0272] Figure 6 further shows confusion matrices for trained deep learning RNA modification and barcode demultiplexing models on held out ground truth test datasets, a. miLung#l modification status, b. miR-17-5p modification status, c. 3’ adapter barcode classification.
[0273] 11. Architecture of the networks used for predicting barcodes and methylation status.
[0274] Figure 7 further shows the architecture of the networks used for predicting barcodes and methylation status. The two models have the exact same architecture and number of parameters, except the number of output neurons of the linear classifier. The network consists of four bottleneck blocks each consisting of three convolutions at the same channel depth before a channel deepening convolution which again is followed by three convolutions of the same depth. One extra convolution after the input, sets up the initial 20 channels, and the final classification is performed by a fully connected layer. All convolutions are ID and followed by batch normalisation. Summary of the sequences used:
[0275] Table 1: Oligo sequences for targeted Nano-smallRNAseq. Bold underlined sequence indicates the barcode. ONT compatible sequence to enable ligation to the ONT RMX adapter is in italic. Bold black are the overhang regions designed to capture small RNAs into the Nano- smallRNAseq library, r denotes that the proceeding nucleotide is RNA. A + denotes that the proceeding base is LNA modified.
[0276] The ribosomal RNA fragment miLung#l exemplarily tested herein has the following (basic) sequence (irrespective of methylation) from 5’ to 3’: GCCGCCGGUGAAAUACCACUAC (SEQ ID NO: 23).
[0277] The ribosomal RNA fragment miLung#l having the sequence GCCGCCGGUGAAAUACCACUAC (SEQ ID NO: 23) has two methylations sites, namely G (if methylated Gm) and C (if methylated Cm). That means that either one methylation site, i.e. 5’ GCCGCCGmGUGAAAUACCACUAC 3’ (SEQ ID NO: 24) or 5’ GCCGCCGGUGAAAUACCACmUAC 3’ (SEQ ID NO: 25), or both methylation sites, i.e. 5’ GCCGCCGmGUGAAAUACCACmUAC 3’ (SEQ ID NO: 26), in the ribosomal RNA fragment miLung#l can be methylated.
[0278] The miR-17-5p exemplarily tested herein has the following basis sequence (irrespective of methylation) from 5’ to 3’ :
[0279] CAAAGUGCUUACAGUGCAGGUAG (SEQ ID NO: 27). This miRNA can be present as unmodified miR-17-5p or as m6a modified miR-17-5p. Said miR-17-5p is known to be differentially m6A modified at position 13 in cancer (Konno et al, 2019).
Claims
CLAIMS1. A method for classifying target RNA modifications in a neural network system, the method comprising: a) obtaining direct RNA sequencing data relating to at least one target RNA in at least one sample; b) selecting a first portion of the direct RNA sequencing data, wherein the first portion of the direct RNA sequencing data comprises first direct RNA sequencing information being indicative for the presence or absence and / or being indicative of the pattern of at least one modification in the at least one target RNA, c) providing the first portion of the direct RNA sequencing data to a first neural network of the neural network system, the first neural network being trained to determine the presence or absence and / or the pattern of at least one modification in said target RNA from direct RNA sequencing data, d) receiving a RNA modification prediction from the first neural network, the RNA modification prediction being indicative for the presence or absence and / or being indicative of the pattern of the at least one modification in the at least one target RNA in the at least one sample.
2. The method according to claim 1, wherein the target RNA comprises a 5'adapter ligated to the target RNA and wherein the first direct RNA sequencing information comprises information of the target RNA and of the 5'adapter, the information being indicative for the presence or absence and / or being indicative of the pattern of the at least one modification in the at least one target RNA in the at least one sample.
3. The method according to any one of the preceding claims, wherein after step a) of claim 1, the method further comprises the steps: i) segmenting the direct RNA sequencing data, ii-a) obtaining the first portion of the direct RNA sequencing data from the segmented direct RNA sequencing data, wherein the steps are performed after step a) and before step b).
4. The method according to any one of the preceding claims, wherein after step a) of claim1, the method further comprises the steps of:A) selecting a second portion of the direct RNA sequencing data, the second portion of the direct RNA sequencing data comprising second direct RNA sequencing information being indicative of a DNA barcode comprised in an adapter combination ligated to the at least one target RNA,B) providing the second portion of the direct RNA sequencing data to a second neural network of the neural network system, the second neural network being trained to determine the DNA barcode,C) receiving a barcode prediction from the second neural network, the barcode prediction being indicative of the DNA barcode.
5. The method according to claim 4, wherein the target RNA is ligated to a 3 'adapter combination and wherein the second direct RNA sequencing information comprises information of the 3 'adapter, the information being indicative of a DNA barcode comprised in the 3'adapter combination.
6. The method according to claim 4 or 5, when dependent on claim 3, further comprising the step of: ii-b) obtaining the second portion of the direct RNA sequencing data from the segmented direct RNA sequencing data.
7. The method according to any one of claims 4 to 6, comprising the step of: associating the modification prediction and the barcode prediction to determine the at least one modification in the at least one target RNA in the at least one sample.
8. The method according to claim 7, comprising the calculation of a modification proportion for the at least one target RNA in the at least one sample.
9. The method according to any one of the preceding claims, wherein after step a) and before step b) of claim 1, following step is performed:- basecalling the direct RNA sequencing data to obtain a predicted target RNA and to obtain quality control information, the quality control information comprising a quality control value indicative of the probability of a wrong basecalling, and- filtering the data by disregarding data with a quality control value being below a quality control threshold value.
10. The method according to claim 9, wherein the quality control threshold value is an average quality score (PHRED) of seven.
11. The method according to claim any one of the preceding claims, wherein after step a) and before step b) of claim 1, following steps are performed:- transforming the direct RNA sequencing data into nucleotide sequence data of the target RNA,- mapping the nucleotide sequence data of the target RNA against a reference target RNA, and disregarding the nucleotide sequence data which does not map to the reference target RNA.
12. The method according to any one of the preceding claims, wherein prior to providing the first portion of the direct RNA sequencing data and / or the second portion of the direct RNA sequencing data to the first neural network and / or to the second neural network, the method comprises:- normalizing the first portion of the direct RNA sequencing data and / or the second portion of the direct RNA sequencing data, and / or scaling the first portion of the direct RNA sequencing data and / or the second portion of the direct RNA sequencing data to a fixed length, and / or- removing outliers.
13. The method according to claim 12, wherein the normalization is performed by Median Absolute Deviation.
14. The method according to claims 12 or 13, wherein the scaling to fix length is performed through linear interpolation.
15. The method according to any one of claims 12 to 14, wherein the removal of outliers is performed by clipping the data at three standard deviations from the median and whereinthe values at those positions are replaced with an interpolation from a linear rolling window.
16. The method according to any one of the preceding claims, wherein the first neural network and / or the second neural network are ID residual neural networks (ResNet).
17. The method according to any one of the preceding claims, wherein the direct RNA sequencing data comprises ionic current data obtained by an ONT platform sequencers output of the target RNA.
18. The method according to any one of the preceding claims, wherein the method is a computer implemented method.
19. The method according to any one of claims 16 to 18, wherein the residual neural network comprises eight residual blocks, each residual block comprising ID convolutional layers, wherein each ID convolutional layer is followed by a batch normalization operation and Rectified linear unit (ReLU) activation, wherein the channel width of the convolution is 20 in a first block and increases in three steps to 67 channels in a last block of the eight residual blocks, wherein before the first block one additional ID convolutional layer with 20 channels is arranged, the additional layer being followed by batch normalization operation, ReLU activation and a pooling layer being configured to reduce the temporal size of the input of the neural network, wherein after the last block an averaging operation is performed, and wherein the averaging is performed over time.
20. A method for training a neural network system that is configured to classify target RNA modifications, the method comprising the steps of: a) selecting direct RNA sequencing training data, the direct RNA sequencing training data comprising information on target training RNA; b) selecting a first portion of the direct RNA sequencing training data, the first portion of the direct RNA sequencing training data comprising first target training RNA information being indicative of a presence or absence and / or being indicative of the pattern of at least one modification in the target training RNA;c) training a first neural network of the neural network system using the first portion of the direct RNA sequencing training data as input and the determination of the presence or absence and / or the determination of the pattern of at least one modification in the target training RNA as target.
21. The method for training a neural network system according to claim 20, wherein after step a) the method further comprises the steps:A) selecting a second portion of the direct RNA sequencing training data, the second portion of the direct RNA sequencing training data comprising second target training RNA information being indicative for a DNA barcode comprised in an adapter combination ligated to the target training RNA;B) training a second neural network of the neural network system using the second portion of the direct RNA sequencing training data as input and the determination of the DNA barcode as target.
22. A neural network system implemented on a computer, the system comprising a first neural network and / or a second neural network, wherein the first neural network and / or the second neural network are trained according to any one of claim 20 to 21.
23. A computer system adapted to classify target RNA modifications comprising: at least one memory configured to store computer program code; at least one processor configured to access the computer program code and operate as instructed by the computer program code to perform the method according to any one of claims 1 to 19 or 20 to 21.
24. A computer program comprising instructions which, when the program is executed by a computer system, causes the computer system to carry out the method of claims 1 to 19 or 20 to 21.
25. A non-transitory computer readable medium storing the computer programme of claim