Generalised framework for identification of RNA modification on specific sequences

The modNet-site framework optimizes model parameters for site-specific RNA modification detection in synthetic sequences, addressing the lack of specificity in existing methods by achieving high accuracy and sensitivity in RNA modification quantification.

WO2026089664A1PCT designated stage Publication Date: 2026-04-30AGENCY FOR SCI TECH & RES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
AGENCY FOR SCI TECH & RES
Filing Date
2025-10-24
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Current methods for detecting RNA modifications in synthetic RNAs, such as those used in RNA therapeutics, lack specificity and accuracy due to generalization from cellular RNA models, leading to suboptimal quantification and safety concerns.

Method used

A machine-learning framework, modNet-site, is developed to optimize model parameters for site-specific RNA modification detection in synthetic sequences by training on a known sequence context, using k-mers from before and after the site of interest, and updating models based on a loss function.

Benefits of technology

Achieves high accuracy and sensitivity in detecting multiple RNA modifications at single-nucleotide resolution, enabling precise quantification of RNA modifications in synthetic RNAs with minimal training data, suitable for commercial RNA products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SG2025050694_30042026_PF_FP_ABST
    Figure SG2025050694_30042026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein is a method for training a framework comprising one or more models for site-specific ribonucleic acid (RNA) modification detection on an RNA sequence of interest, comprising receiving a training dataset of sequences, each sequence being either an unmodified RNA sequence or a modified RNA sequence and, for each sequence, annotations indicating whether the respective sequence contains a modification at one or more sites; extracting a sequence context read j from each sequence, and extracting, for each site of interest i in read j, i being a site at which it is desired to detect a modification, a site-specific annotation from the annotations, the site-specific annotation indicating whether a specific modification m occurs at site i; training at least one said model to determine a probability of modification m at site i, using a plurality of k-mers, extracted from each read j, including one or more k-mers from before site i, at site i, and after site i; updating the at least one model based on a loss between the probability and the site-specific annotation; and outputting the framework, after a predefined number of epochs for training has been completed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] GENERALISED FRAMEWORK FOR IDENTIFICATION OF RNA MODIFICATION ON SPECIFIC SEQUENCES

[0002] Technical Field

[0003] This present invention relates, in general terms, to methods and systems for detection of ribonucleic acid (RNA) modification on specific sequences, and more specifically, to methods and systems for optimising model parameters for RNA sequences.

[0004] Background

[0005] This background is provided for generally presenting the context of the disclosure. Contents of this background section are neither expressly nor implied admitted as prior art against the present disclosure.

[0006] Post-transcriptional modifications are a hallmark of cellular RNA. Modifications occur at any of the four nucleobases, impacting ribonucleic acid (RNA) structure, recognition, and function. The ability to change RNA properties without impacting its coding sequence has made RNA modifications an integral part of RNA therapeutics that use in vitro transcribed, highly modified synthetic RNAs to reduce immunogenicity, increase stability, and translational efficiency.

[0007] Direct RNA-sequencing-based RNA modification detection methods have made epitranscriptome profiling accessible, revealing the roles of RNA modifications in cellular biology. Most methods for high throughput profiling of RNA modifications have been optimised for cellular RNA but tend to generalise across cellular RNAs in various organisms. However, modified synthetic RNAs can differ substantially from cellular RNA due to artificial sequence context, the use of non-natural modifications, or a higher density of modified bases. Even though RNA modifications are an essential component of RIMA therapeutics, dedicated methods for profiling RNA modifications for synthetic RNAs have not been developed. This makes them suboptimal for quantifying RNA modifications in commercial synthetic RNAs, which lack the sequence context learned in the cellular-RNA-trained models and are often more heavily modified than cellular RNAs.

[0008] RNA modifications are conventionally detected through laborious wet-lab assays coupled with next-generation sequencing (for e.g., m6ACE-Seq, miCLIP-Seq, bisulfite sequencing). There remains, however, a desire to detect RNA modifications via direct RNA-sequencing current signals.

[0009] It is, however, challenging to detect RNA modifications incorporated into specific synthetic mRNAs in RNA therapeutics such as RNA vaccines. Currently, there are limited solutions tailored for detecting RNA modifications in a specific sequence context to ensure accuracy, safety and efficacy. There is hence a need for development of approaches for detection of RNA modifications in a specific synthetic RNA sequence.

[0010] Summary

[0011] Disclosed is a method for training a framework comprising one or more models for site-specific ribonucleic acid (RNA) modification detection on an RNA sequence of interest, comprising:

[0012] receiving a training dataset of sequences, each sequence being either an unmodified RNA sequence or a modified RNA sequence and, for each sequence, annotations indicating whether the respective sequence contains a modification at one or more sites;

[0013] extracting a sequence context read j from each sequence, and extracting, for each site of interest / in read j, i being a site at which it is desired to detect a modification, a site-specific annotation from the annotations, the site-specific annotation indicating whether a specific modification m occurs at site / ; training at least one said model to determine a probability of modification m at site / , using a plurality of -mers, extracted from each read j, including one or more k-mers from before site / , at site i, and after site / ;

[0014] updating the at least one model based on a loss between the probability and the site-specific annotation; and

[0015] outputting the framework, after a predefined number of epochs for training has been completed.

[0016] Also disclosed is a framework for detecting site-specific ribonucleic acid (RNA) modifications, trained according to the method described above.

[0017] Also disclosed is a method for detecting site-specific ribonucleic acid (RNA) modifications on a new RNA sequence of interest with a known, specific sequence context, comprising:

[0018] subjecting the new RNA sequence of interest to direct RNA sequencing, to produce RNA sequence data comprising basecalled reads;

[0019] extracting a newly sequenced read with the known, specific sequence context from the basecalled reads;

[0020] applying the framework of claim 8 to the newly sequenced read with the known, specific sequence context, thereby generating a probability that the RNA sequence of interest comprises a specific modification or modifications at each site in the new sequence context read; and labelling each site in the newly sequenced read with the known, specific sequence context as comprising a said specific modification if the probability of the said specific modification at the respective site is above a predetermined threshold.

[0021] Also disclosed is a system for training a framework comprising one or more models for site-specific ribonucleic acid (RNA) modification detection on an RNA sequence of interest, comprising:

[0022] a receiver module, for receiving a training dataset of sequences, each sequence being either an unmodified RNA sequence or a modified RNA sequence and, for each sequence, annotations indicating whether the respective sequence contains a modification at one or more sites;

[0023] an extraction module, for extracting a sequence context read j from each sequence, and extracting, for each site of interest i in read j, i being a site at which it is desired to detect a modification, a site-specific annotation from the annotations, the site-specific annotation indicating whether a specific modification m occurs at site i;

[0024] a training module, for training at least one said model to determine a probability of modification m at site i, using a plurality of k-mers, extracted from each read j, including one or more k-mers from before site i, at site i, and after site i; and

[0025] an output module, for updating the at least one model based on a loss between the probability and the site-specific annotation and outputting the framework, after a predefined number of epochs for training has been completed.

[0026] Brief description of the drawings

[0027] Embodiments of the present invention will now be described, by way of nonlimiting example, with reference to the drawings in which:

[0028] Figure 1 illustrates identification of multiple RNA modifications from direct RNA-Seq with modNet-site, a framework optimised for predicting RNA modifications from new sequencing data from a known single RNA sequence. Top: model training; bottom: application scenarios, illustrated for the A-modification models. modNet-site directly predicts read-level modification probabilities.

[0029] Figure 2, comprising Figures 2(a) to 2(f), illustrates modNet-site effect in maximizing the RNA modification detection accuracy in known sequence context as disclosed in the present invention. Figure 2(a): ROC and Precision-Recall curves of the single-modification modNet-site models of nine modifications (m6A, mlA, m5C, f5C, TheinoG, Ψ, m1Ψ, mo5U, and s2U). Figure 2(b): ROC and precision-recall curves of the single-modification Ψ, m1Ψ, mo5U, and s2U modNet-site models in the new batch of Ψ, m1Ψ, mo5U, and s2U samples for validation. Figure 2(c): ROC AUCs of the multimodification modNet-site models across unmodified and different modifications. Figure 2(d): ROC AUCs of the single-modification modNet-site models between unmodified and modified. Figure 2(e): ROC curves of the multi-modification A-modification and U-modification modNet models and dorado's m6A and Ψ models. Figure 2(f): ROC curves of the single-modification m6A and Ψ modNet models and dorado's m6A and Ψ models.

[0030] Figure 3, comprising Figures 3(a) to 3(e), depicts identification of differentially modified positions as disclosed in the present invention. Figure 3(a): Absolute and relative stoichiometry estimation at various ratios of modified sites (0%, 25%, 50%, 75%, and 100%) of single-modification and multi-modification modNet models, and Figure 3(b): modNet-site models of the U modifications (MJ, mlUJ, mo5U, and s2U). Figure 3(c): modified sample estimation at various ratios of modified reads (0%, 25%, 50%, 75%, and 100%) based on the single-modification modNet-site model prediction for Ψ, m1Ψ, mo5U, and s2U. Figure 3(d): ROC and precision-recall curves of the single-modification mo5U modNet model for the TriLink eGFP mRNA. Figure 3(e): Sample estimation of Tri Link eGFP mRNA at various ratios of modified reads (0%, 25%, 50%, 75%, and 100%) based on the single modification mo5U modNet model.

[0031] Figure 4 illustrates the modNet-site model architecture, according to a preferred embodiment of the present invention.

[0032] Figure 5 provides a schematic of the generalised framework architecture for the identification of RNA modification, according to a preferred embodiment of the present invention.

[0033] Detailed Description The present invention relates to systems and methods for a machine-learning framework which detects and distinguishes multiple RNA modifications on all four bases using direct RNA-Seq data on specific sequences. Disclosed herein is a sequence-specific model (hereinafter referred to as "modNet-site"). By optimising model parameters for a known sequence context, substantially higher accuracy and sensitivity on synthetic RNAs may be achieved, while using minimal training data. modNet-site models are applied to detect nine different RNA modifications, at single molecule, single nucleotide resolution. Advantageously, this demonstrates the feasibility to sequence an extended alphabet of modified synthetic RNA bases using direct RNA-Seq.

[0034] While the systems and methods according to the presently claimed invention are described in detail with reference to preferred embodiments, it shall be understood that this is merely to facilitate discussion and not to limit the scope of the invention.

[0035] Throughout the specification, the term "RNA-Seq" may be used to refer to a sequencing technique used to quantify and identify RNA molecules in a biological sample, providing a snapshot of the transcriptome at a specific time. The term "sequence context" refers to a specific nucleotide (A, U, C or G) composition and structure surrounding a particular sequence within the RNA molecule. The term "modification", where used in the context of RNA modification detection, may refer to one of the nine modifications, i.e., N6-methyladenosine (m6A), Nl-methyladenosine (mlA), 5-formylcytidine (f5C), 5-methylcytidine (m5C), Pseudouridine (pU or UJ), 1-methylpseudouridine (mlpU or mlQJ), 5-methoxyuridine (mo5U), 2-thiouridine (s2U) and thienoguanosine (thG), but shall not be limited to these.

[0036] Figure 5 presents a schematic of the generalised architecture for a method for training a framework comprising one or more models for site-specific ribonucleic acid (RNA) modification detection on an RNA sequence of interest. The method comprises receiving a training dataset of sequences (102), each sequence being an unmodified RNA sequence or a modified RNA sequence and, for each sequence, annotations indicating whether the respective sequence contains a modification at one of the sites, extracting a sequence context read j from each sequence, and extracting, for each site of interest i in read j, / being a site at which it is desired to detect a modification, a sitespecific annotation from the annotations (104), the site-specific annotation indicating whether a specific modification m occurs at site / , training at least one said model to determine a probability of modification m at site / (106), using a plurality of k-mers, extracted from each readj, including one or more -mers from before site / , at site / , and after site / , updating the at least one model based on a loss between the probability and the site-specific annotation (108), and outputting the framework, after a predefined number of epochs for training has been completed (110). An epoch is a single complete pass through the entire training dataset - e.g., 30 epochs are 30 complete passes through the entire training dataset. The number of epochs is defined based on the model's performance plateaus. Beyond the predefined number of epochs for training, any further increase may be found to offer little to no improvement in the model's performance. In each of the training datasets, the sequence context associated with the dataset is fixed and a modification status associated with the modified RNA sequence of interest data may be known.

[0037] The models may be trained using a generalisable workflow, and with nine synthetic RNA samples with various modifications. The modifications may be, but are not limited to: N6-methyladenosine (m6A), Nl-methyladenosine (mlA), 5-formylcytidine (f5C), 5-methylcytidine (m5C), Pseudouridine (pU or UJ), 1-methylpseudouridine (mlpU or mlUJ), 5-methoxyuridine (mo5U), 2-thiouridine (s2U) and thienoguanosine (thG), incorporated through in vitro transcription (IVT). The single modification model trained with a site-based approach in detecting RNA modifications in unseen synthetic RNA achieves a precision-recall AUC 0.99. The multi-modification models achieve simultaneous same-base-modification detection for U modifications (Ψ, m1Ψ, mo5U, and s2U), C modifications (m5C and f5C), and A modifications (m6A and mlA), suggesting they can separate one same-base modification from another. These observations suggest that the generalised workflow for training the models, according to the present invention, enables highly accurate multiple RNA modifications optimised for specifically synthetic RNA.

[0038] Synthetic RNA modification model training using modNet

[0039] Figure 1(a) provides a schematic showing modNet-site optimised for predicting RNA modifications from new sequencing data, from new sequencing data from a single RNA sequence. modNet expands a multiple-instance-learning (MIL) framework of m6Anet to a multi-modification classifier.

[0040] The present site-specific modNet-site model is trained on each site of a single sequence of interest without any generalisability to unseen sequence context. Each model can either be trained to infer modification probabilities for a single modification or for multiple modifications. These models are implemented as part of the modNet package, which provides modules for training, inference, and arguments to use pre-trained models for N6-methyladenosine (m6A), Nl-methyladenosine (mlA), 5-formylcytidine (f5C), 5-methylcytidine (m5C), Pseudouridine (pU or Ψ), 1-methylpseudouridine (m1pU or m1Ψ), 5-methoxyuridine (mo5U), 2-thiouridine (s2U) and thienoguanosine (thG).

[0041] Increase in accuracy of modNet models, by optimisation for specific RNA sequences

[0042] While modNet maintains an ability to identify RNA modifications in a new sequence context, it typically requires training data which includes all relevant k-mers. However, if the sequence context of interest is limited and known in advance, the inclusion of non-relevant k-mers in model training may reduce the performance. According to an aspect of the invention, there is provided, a model, referred to as modNet-site in this disclosure, which optimises the model for a specific synthetic RNA of interest to maximize the modification prediction accuracy for a limited sequence context. Single-modification and multi-modification modNet-site models are trained with four selected sequences (See Example 1 - Methods'), each having between 31 and 39 modified sites, wherein the modNet-site approach learns one model per site (Figure 2(a)). modNet-site was trained on 50% of reads from a single sequence of interest for each site that contained a modified base and evaluated on the remaining 50% unseen reads (test data) and on the independent validation data (see Example 1 - Methods). Across all modifications, the single-modification modNet-site models achieve an accuracy of ROC AUC of 0.999 both in the test and validation data (Figure 2(b) and 2(c)) while the multi-modification modNet-site models have an average ROC AUC of 0.907 (Figure 2(d)).

[0043] A comparison with dorado's m6A, m5C, and J models shows that modNet-site achieves the same or better accuracy on these sequences (Figures 2(e) to 2(g)). While the Dorado models maintain generalisability to any sequence context, they require substantial training data that limits the number of modifications that can currently be detected. In contrast, modNet-site achieves similar performance, but with minimal training data, allowing the detection of potentially any modification within a limited sequence context.

[0044] Accurate estimation of fraction of modified RNA molecules

[0045] Advantageously, the development of a direct RNA-Seq technology is an ability to profile single RNA molecules. To evaluate the ability of modNet to identify modified bases for individual RNA molecules, in silico mixtures of modified and unmodified RNAs at specific ratios (0%, 25%, 50%, 75%, 100%) are generated using the validation data set, consisting of U modifications. For each modNet-site model, a read-level probability threshold is selected, which was then used to calculate the fraction of modified reads at each site (modification rate, or stoichiometry).

[0046] The normalised modification rate from single-modification and multimodification modNet-site models closely resembles the expected proportion from samples with 0%, 25%, 50%, 75%, and 100% modified reads for all modifications in the validation data set (Figure 3(a)). Further, the modNet-site model was found to directly capture the absolute stoichiometry of the different proportions of modified reads (Figure 3(b)).

[0047] The modNet-site models were also tested for the ability to estimate the fraction of modified reads in a sample is tested. Here, the probability of a read being modified as the average modification probability from all sites is calculated, with a read probability threshold of 0.5 to identify modified reads. This approach is applied to the validation data that contains both modified reads and unmodified reads at ratios of 0%, 25%, 50%, 75%, and 100%. On these data, the single-modification modNet-site models identified modified reads with 100% accuracy, leading to perfect estimates of the number of modified reads in each sample (Figure 3(c)).

[0048] Quantification of modified bases with single-molecule resolution for commercially available RNAs

[0049] Synthetic RNA products, which rely on modifications, such as in GFP mRNA, CAS RNA or mRNA vaccines, benefit from a high sensitivity and accuracy to identify modified bases with minimal training data. This motivates a test for feasibility of applying the modNet-site model to detect modifications present on a commercially available RNA, which may be done by generating direct RNA-Seq data for mo5U-modified and unmodified CleanCap® Enhanced Green Fluorescent Protein (eGFP) mRNA (TriLink). On these data, the pre-trained modNet-mo5U model classifies modified and unmodified bases with a ROC AUC of 0.901 (Figure 3(d)). In contrast, modNet-site-mo5U, which was trained specifically for the GFP sequence context (using independent training data, see Example 1 - Methods'), achieves a ROC and Precision-Recall AUC of 1, perfectly classifying modified and unmodified reads (Figure 3(d)). Similarly, modNet-site was able to perfectly estimate the fraction of modified reads in a mixed sample (Figure 3(e)), providing not just an estimate of the sample purity, but also of the modification status for each individual RNA molecule at single base resolution (Figure 3(f)). Collective analysis of the figures demonstrates that modNet-site may be trained using a single RIMA sequence, achieving high accuracy for profiling modified synthetic RNAs and commercial RNA products using direct RNA-Seq. Advantageously, learning the sequence context of a set of synthetic RNAs well to quantify their RNA modification stoichiometry accurately addresses a need for quantifying RNA modifications in commercial synthetic RNAs, via the modNet and modNet-site frameworks. Further, these models have also been extended to predict multiple same-base modifications, enabling simultaneous detection of multiple modifications from one model.

[0050] Discussion

[0051] Despite their good performance in cellular RNA, generalised RNA modification detection models require large amount of data and compute resources for training. Still, they may not perform well in synthetic RNAs as we have demonstrated that characteristics of RNA modifications differ significantly between cellular and synthetic RNA, making RNA detection in cellular RNAs and synthetic RNAs two different tasks. As synthetic-RNA-trained models do not learn the cellular sequence context, they may overall not perform as well as cellular-RNA-trained models, yet as we showed synthetic-RNA- and cellular-RNA-trained models perform similarly in some motifs.

[0052] To illustrate the generalised framework as described herein, models for nine modifications have been trained. While users may use these pre-trained models for predictions, uses may also train models for their modifications of interest with a limited amount of data. In spite of a small number of aligned reads from some IVT samples, it has been shown that accurate modNet-site models may be trained. Training of these models may not be limited to singlemodification prediction, and the modNet-site model training may be extended to multiple same-base modifications. Combing m5C and s2U in mRNAs increases the expression of the antigen, the mRNA product. With the multimodification detection approach presented herein, the exploration of incorporating an optimal amount of RIMA modifications in commercial RNA products, including therapeutics, to maximise their efficacy and safety may be enabled.

[0053] Examples

[0054] Example 1 - Methods for design and generation of synthetic sequences

[0055] A pooled library of synthetic sequences (synLibVl) which consists of the following two sub-libraries have been designed as follows:

[0056] Sub-library 1: Including all possible k-mers with curlcake

[0057] Curlcake (https: / / cb.csail.mit.edu / cb / curlcake / ), which is a synthetic sequence designing software that includes all possible k-mers (k=5) and minimizes secondary RNA structure with RNAshapes version 2.1.6 (https: ZZbteteer.cebjtec.unj--biete ete.de / rnasha eg) was used. Curlcake was run with the following parameters: k=5, length=50, multiplicity= 14 to 20, and attempts=100. A total of seven curlcake constructs were produced, one for each multiplicity parameter specified from 14 to 20. The lengths of the seven constructs are 29,070, 30,630, 33,390, 34,950, 37,230, 37,590, and 43,110 bp, respectively. The curlcake constructs were split into shorter sequences of 150-bp length, each overlapping by 30 nucleotides with the previous 150-bp sequence, resulting in 2,048 individual sequences.

[0058] Sub-library 2: Using human transcriptome fragments

[0059] To synthesise synthetic sequences resembling the human transcriptome, both cell-type-specific and constitutively expressed transcripts using the data from the Singapore Nanopore Expression Project (SG-NEx), which included the following 8 cell lines: A549, MCF7, HCT116, HepG2, K562, H9, HEYA8, and HEK293T, were selected. Here a transcript was considered to be expressed in a cell line if it is supported by at least five read counts in each replicate of the direct RNA-Seq samples from that cell line. Read counts were estimated based on the number of aligned reads for each transcript using the transcriptome alignment bam files downloaded from the SG-NEx project. Transcripts that are expressed in at least 5 out of the 8 SGNEx cell lines ("constitutively expressed"), or which are exclusively expressed in the HEK293T ("HEK293T-specific") or H9 cell line ("H9-specific") were then selected. We also filtered for transcripts with available 5'UTR, 3'UTR, start and stop codon annotation information. We then selected the transcript with the highest read count of each gene and split the transcript into 150-bp sequences, each overlapping by 30 nucleotides with the previous 150-bp sequence. As a final filter, we removed all 150 bp sequences that contain the same nucleotide in 15 consecutive positions. All 150 bp sequences generated were annotated by their locations on the transcript: TSS, 5'UTR, CDS, Exon Junction, and 3' UTR.

[0060] Final library design

[0061] A T7 promoter sequence was added to the 5' end of the 150-bp synthetic sequences to facilitate in vitro transcription. In addition, the entire construct (T7 promoter and 150-bp synthetic sequences) was also flanked by two primer binding sequences (5'-TTGGACCCTCGTACAGAAGC-3' and 5'-TGGTCTTTGAATAAAGCCTGAGTAGGAAG-3' at the 5' end and 3' end respectively) to facilitate PCR amplification. The designed DNA sequences were purchased from Twist Biosciences as two separate oligo pools.

[0062] Example 2 - PCR, in vitro transcription (IVT) and sequencing of synthetic sequences

[0063] Both sub-libraries were subjected to PCR amplification (KAPA HiFi HotStart PCR, #07958897001) and DNA purification (Zymo Research DNA Clean & Concentrator, D4013), before being pooled in equimolar concentrations.

[0064] The pooled synLibVl DNA templates were transcribed in vitro (Thermo Fisher Scientific MEGAscript™ T7 Transcription Kit, AM1334) with unmodified and modified nucleotides. For the unmodified RNA, only unmodified nucleotides were provided during in vitro transcription. For the m6A and mlA RNA samples, the ATP was substituted with either m6A or mlA (TriLink Biotechnologies) while the other 3 nucleotides remained unmodified. The transcribed RNAs were then purified (New England Biolabs Monarch RNA Cleanup Kit, T2040), poly(A)-tailed (New England Biolabs E. coli Poly(A) Polymerase, M0276) and purified again, before they were subjected to nanopore direct RNA sequencing (Oxford Nanopore Technologies, SQK-RNA002 or SQK-RNA004).

[0065] For the 4J, ml4J, mo5U, s2U, m5C, and f5C modified RNA samples, a 3' primer carrying a 120-nt poly(T) sequence was used for the PCR amplification of the IVT templates, allowing us to bypass the subsequent polyA-tailing reaction and second RNA purification before direct RNA sequencing. The modified nucleotides were also obtained from TriLink Biotechnologies.

[0066] The complete set of sequencing experiments is listed in Table 1. The RNA source was synLibVl - IVT for entries 1-24, and Trilink_CleanCap_EGFP for entries 25-26.

[0067] Table 1. Set of sequencing experiments

[0068]

[0069]

[0070]

[0071] Example 3 - Pre-processing of direct RNA-Seq samples

[0072] Alignment

[0073] The basecalled reads from the sequencing run were aligned to the synLibVl reference with minimap2 version 2.2211with the options -ax map-ont --secondary=no -k5 -w5 -g2k -Al -B2 -02,32 -El,0 -z200 --cap-sw-mem -s 20. The alignment files were converted from sam file format to bam file format, which were then indexed, with samtools version 1.18.

[0074] Segmentation

[0075] The fast5 files were converted to blow5 files with slowStools12version 1.1.0 with the / 2s command, and the pod5 files were converted to blow5 files with blue-crab13version 0.3.0 with the p2s command. The blow5 files were merged into a blow5 file with the slowStools merge command. Segmentation was performed with f5c14version 1.2. The reads from the fastq file and the raw signals from blow5 file were indexed with the f5c index command. Then, f5c eventalign was run with the options --rna — min-mapq 0 --min-recalib-events 100 --signal-index --scale-event. The --kmer_model argument was used to specify the modification-specific k-mer model for the RNA002 modified samples when comparing the performance between the original k-mer model to the modification-specific k-mer model only.

[0076] Data preparation

[0077] The output from f5c eventalign was summarised for each read and each site using modnet dataprep and site-dataprep, which outputs a json file containing sites of the base of the modification (e.g. for m5C, only NNCNN sites were outputted). For the unmodified sample, sites of all bases were outputted separately into four json files (one for NNANN, NNCNN, NNUNN, and NNGNN) using modnet dataprep and site-dataprep.

[0078] Example 4 - RNA modification models

[0079] Example 4.1 - Site-specific single-modification detection models

[0080] (1) Model architecture

[0081] In contrast to the other models, here we assume that the sequence context is fixed, and that the modification status for individual reads at each base is known. The site-specific single-modification models are then trained with read-level labels for each position / of a sequence S = {s1,...,sn} of length n. For simplicity, i = 1... n. However, in practice, only positions that have training data for modified and unmodified reads will be included in model training.

[0082] For each read j at position i, the modification status Yifjis defined as

[0083] Yi:j = {lifreadjismodifiedatpositioni, Ootherwise]

[0084] Here the modNet model architecture is modified to reflect the change in training data. Firstly, as each site contains three possible 5-mers from positions i, i - 1, and i + 1, we changed the k-mer-embedding layer to have 3 input channels. The raw signal deaggregation layer, the read-level feature encoder layer, and the read-level probability encoder layer are retained. As the model is trained with read-level training data, the site-level pooling layer was removed. We refer to this model as modNet-site, which predicts the read level probability for read j to be modified at position / site i:

[0085]

[0086] (2) Estimation of model parameters (training)

[0087] Similar to m6Anet, we trained the network by minimizing the binary cross entropy L between

[0088]

[0089] and Ytj for all reads j - 1.. N at site i:

[0090]

[0091] Example 4.2 - Site-specific multi-modification detection models

[0092] (1) Model architecture

[0093] Similar to the site-specific, single modification modNet-site models, the sitespecific multi-modification modNet-site models are trained with read-level labels for each position of a sequence of interest. For each read j, the modification status at position i, Yi is defined as

[0094]

[0095] with the index 0 representing unmodified base

[0096] ytjfi = {lifreadjisunmodifiedatpositioni, ^otherwise]

[0097] and with m — 1,..., M representing all possible M base modifications included in the model: yij,m={lifreadjhasmodificationmatpositioni, Ootherwise}

[0098] Here we assume that

[0099]

[0100] each read is either unmodified or has a single modification.

[0101] The multi-modification model then extends the site-specific single modification model described above by returning modification probabilities for each modification type (please also refer to the multi-modification model description above).

[0102] (2) Estimation of model parameters (training)

[0103] The model parameters are updated subject to minimising the cross entropy loss L betweenpij mand Yi^m for all reads j = 1... N and all modifications m = where m = 0 represents the unmodified base as described for the multi-modification modNet model:

[0104]

[0105] Example 5 - Training and evaluation of the RNA modification detection models

[0106] We used the direct RNA-Seq data from the synthetic RNAs to train and test RNA modification models.

[0107] The site-specific The site-specific modNet-site models are trained and evaluated to identify RNA modifications with high accuracy for a specific sequence context. The models are not generalizable to other sequence contexts. For each nucleotide (A, U, C, and G), we selected one representative synthetic sequence from the human transcriptome fragment library with more than 50 reads in all sites of the specific nucleotide (including sites containing multiple modified bases in the k-mer) in the unmodified sample and modified sample, respectively.

[0108] For the site-specific single modification models, reads from the unmodified sample were labelled as unmodified for all positions (yi7— 0), and reads from the modified samples with the modified base were labeled as modified (yi7= 1 ). For the site-specific multi-modification models, all reads from the unmodified sample were labeled as unmodified for all positions I (yij o= l,yi,7,m= Oforallm = 1... M) with 0 corresponding to the unmodified base and m representing all possible M base modifications included in the models. Reads from the modified samples with the modified bases were labeled as modified yLijim= lifreadjhasmodificationmatpositiom, Ootherwi.se] for m = 0,..., M). For each position, we generated position-specific training and testing sets and trained a separate model for each position / of the selected sequence.

[0109] Example 5.1 - Model training

[0110] Site-specific single-modification modNet-site models (m6A, mlA, f5C, m5C, pU, mlpU, mo5U, and s2U, TheinoG) were trained on all modified positions on four sequences with modNet site-train (one sequence for each base). We further trained site-specific modNet-site multi-modification models on the U modifications (unmodified, pU, mlpU, mo5U, and s2U), the C modifications (unmodified, f5C and m5C), and the A modifications (unmodified, m6A and mlA) on all modified positions on the same sequences with modNet site-train.

[0111] For each site-specific model, the site-specific training set was further split into training and validation (for model parameter optimization), and testing (for model selection) randomly at a ratio of 63:27:10, corresponding to the ratios of 7:3 (training:validation) and 9:1 (training+validation: testing). We trained each site-specific model independently with a fixed learning rate of 0.0001, 30 epochs, a seed set at 25, and 5 rounds of iterations using modNet site-train — Ir 0.0001 —seed 25 —epochs 30 — save_per_epoch 1 — numjterations 5. The number of epochs was set at 30, where the model's performance plateaus. Any further increase to the number of epochs was found to have little to no improvement in the model's performance.

[0112] Example 5.2 - Evaluation of model performance

[0113] For each position i, the site-specific modNet-site model was run on the test set using the modnet site-inference command, with the — model_state_dict argument specified to select the site-specific model. For the site-specific single modification modNet-site models, reads in the test data from the modified sample were labeled as 1 while reads from the unmodified sample were labeled as 0 (corresponding to Yi:j ). The site-specific multi-modification modNet-site models were evaluated separately for each modification, where only reads with the specific modifications were labeled as 1, whereas reads with any other modification and unmodified reads were labeled as 0 (corresponding

[0114]

[0115] Weaveraged the read-level probability for reads from each site to obtain the site-level probability and ranked the site-level probability across all sites in the same sequence. For comparison, we ran the general single- and multi-modification models on the site-specific test set and ranked the site-level probability across all sites in the same sequence.

[0116] The python packages scikit-learn and matplotlib were used to calculate the ROC curve, Precision Recall curve, and the Area under the ROC and PR curves.

[0117] Example 6 - Methods for specific analysis

[0118] Example 6.1 - Read threshold optimisation The read-probability threshold was defined as the value where the probability distributions from the modified and unmodified reads intersect, using a beta distribution estimate based on the training data.

[0119] For the U-modification modNet-site model, we defined the 90th percentile MJ, mlUJ, mo5U, and s2U-read-level probabilities across all reads as the readlevel probability thresholds. Using these read-level probability thresholds, we calculated the number of reads with read-level probabilities above the thresholds to obtain modification rate estimates.

[0120] Example 6.2 - Generation of in silico mixture data for evaluation

[0121] The modification rates were estimated for samples from the validation set, consisting of the U modifications (*4J, mlUJ, mo5U, and s2U), with the single-and multi-modification modNet and modNet-site models. To compare the modNet and modNet-site models, we evaluated results only on the 31 U sites on the human3356 synthetic sequence. For each site, 100 reads were randomly picked from the unmodified and modified samples at different modified-to-unmodified ratios (0:100, 25;75, 50:50, 75:25, and 100:0). The single-modification and multi-modification modNet inference and modNet-site inference commands were run on each simulated mixture sample. We performed the sampling and inferencing for five rounds and combined the results to obtain final modification-rate estimates.

[0122] Example 6.3 - Estimation of absolute and relative modification rates

[0123] We used the modification ratio estimates from modNet-site as the absolute modification rate. To obtain relative modification rates, we normalised the predicted modification rates to be relative to the 100:0 and the 0:100 modified-to-unmodified ratio with the following equation:

[0124]

[0125] The python package seaborn was used to plot the box plots of the absolute and relative modification rate across the samples.

[0126] Example 6.4 - Estimation of the fraction of modified reads in mixed samples The fraction of modified reads with the single-modification modNet-site models on the validation data set was estimated, consisting of U modifications (4J, ml, mo5U, and s2U). We randomly selected reads from the unmodified and modified samples with at least 20 U sites at different modified-to-unmodified read ratios (0:100, 25;75, 50:50, 75:25, and 100:0) and ran modnet-site inference on each sample. For each read, we calculated the probability being of being modified (read probability) by averaging the modification probability from all sites from the read. We use a read probability threshold of 0.5 to identify modified reads. We used the python package matplotlib to plot the bar plots.

[0127] Example 6.5 - Evaluation on commercial mRNA products

[0128] The unmodified and mo5U-incorporated TriLink CleanCap® Enhanced Green Fluorescent Protein (eGFP) mRNA was sequenced with the RNA004 sequencing chemistry. We trained a single-modification mo5U modNet-site model with 10,000 reads from each the unmodified and the mo5U-incorporated eGFP mRNA samples. We excluded three U sites from the 101 U sites for training as the three sites' read coverages are only 5% of the average U-site read coverage in the mo5U-incorperated eGFP mRNA sample. We evaluated the eGFP single-modification mo5U modNet-site model with 10,000 unseen reads from each of the unmodified and modified samples. We calculated the sitelevel probabilities by averaging read-level probabilities for each site. We evaluated the models' performance based on the ROC curve using the sitelevel probabilities. We used the python packages scikit-learn and matplotlib to calculate the ROC curve and the Area under the ROC curve. We also estimated the fraction of modified reads in the eGFP mRNA samples with the analysis described above. Read-level probabilities for 1,000 randomly chosen reads from the unmodified and mo5U-modified eGFP samples are visualised as a heatmap with the seaborn python package. Example 7 - Code Availability

[0129] modNet is available through GitHub: https: / / github.com / GoekeLab / modNet The code to generate figures in this manuscript is available through GitHub: https: / / github.com / GoekeLab / modNet_documentation

[0130] Example 8 - Data Availability

[0131] All data generated are available on the European Nucleotide Archive under the study accession PRJEB80152. The eGFP direct RNA sequencing fast5 files from the Grunter et al. study were downloaded from SRA SRP389850.

[0132] It will be appreciated that many further modifications and permutations of various aspects of the described embodiments are possible. Accordingly, the described aspects are intended to embrace all such alterations, modifications, and variations that fall within the spirit and scope of the appended claims.

[0133] Throughout this specification and the claims which follow, unless the context requires otherwise, the word "comprise", and variations such as "comprises" and "comprising", will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps.

[0134] The reference in this specification to any prior publication (or information derived from it), or to any matter which is known, is not, and should not be taken as an acknowledgment or admission or any form of suggestion that that prior publication (or information derived from it) or known matter forms part of the common general knowledge in the field of endeavour to which this specification relates.

Claims

CLAIMS1. A method for training a framework comprising one or more models for site-specific ribonucleic acid (RIMA) modification detection on an RNA sequence of interest, comprising:receiving a training dataset of sequences, each sequence being either an unmodified RNA sequence or a modified RNA sequence and, for each sequence, annotations indicating whether the respective sequence contains a modification at one or more sites; extracting a sequence context read j from each sequence, and extracting, for each site of interest / in read j, i being a site at which it is desired to detect a modification, a site-specific annotation from the annotations, the site-specific annotation indicating whether a specific modification m occurs at site / ; training at least one said model to determine a probability of modification m at siteusing a plurality of k-mers, extracted from each readj, including one or more k-mers from before site / , at site / ', and after siteupdating the at least one model based on a loss between the probability and the site-specific annotation; andoutputting the framework, after a predefined number of epochs for training has been completed.

2. The method according to claim 1, wherein the framework comprises at least one site-specific single-modification detection model (SMDM), the plurality of k-mers used to train each said SMDM comprising a respectively different specific modification m, and at least one sitespecific multi-modification detection model (MMDM), the plurality of k- mers used to train the MMDM comprising multiple specific modifications m.

3. The method according to claim 1 or 2, wherein the one or more k-mers extracted from each read j, including three k-mers being at sites / -l, iand 7+1.

4. The method according to any one of claims 1 to 3, wherein the loss between the probability and the site-specific annotation is a binary cross entropy L function, across all reads j at site 7, L being defined as follows:where Yij,m is an annotation relating to modification m at site 7 in read j, N is the number of sites and M is the number of modifications.

5. The method according to any one of claims 1 to 4, further comprising:pre-processing, before extracting the sequence context, the dataset by aligning the RNA sequence data to the modified and unmodified synthetic RNA sequences to obtain files of a first file format;segmenting the files to locate the sequence context read j; and annotating each read and each site.

6. The method according to any one of the preceding claims, further comprising validating the framework and testing the framework with a distinct test dataset.

7. The method according to any one of the preceding claims, wherein each of the models in the framework is trained independently of one another, at a pre-defined learning rate and number of iterations.

8. The method according to any one of the preceding claims, wherein the RNA modification type comprises one or more of: N6-methyladenosine (m6A), Nl-methyladenosine (mlA), 5-formylcytidine (f5C), 5- methylcytidine (m5C), Pseudouridine (pU or MJ), 1-methylpseudouridine (mlpl) or mlQJ), 5-methoxyuridine (mo5U), 2-thiouridine (s2U) andthienoguanosine (thG).

9. A framework for detecting site-specific ribonucleic acid (RNA) modifications, trained according to the method of any one of claims 1 to 8.

10. A method for detecting site-specific ribonucleic acid (RNA) modifications on a new RNA sequence of interest with a known, specific sequence context, comprising:subjecting the new RNA sequence of interest to direct RNA sequencing, to produce RNA sequence data comprising basecalled reads;extracting a newly sequenced read with the known, specific sequence context from the basecalled reads;applying the framework of claim 8 to the newly sequenced read with the known, specific sequence context, thereby generating a probability that the RNA sequence of interest comprises a specific modification or modifications at each site in the new sequence context read; andlabelling each site in the newly sequenced read with the known, specific sequence context as comprising a said specific modification if the probability of the said specific modification at the respective site is above a predetermined threshold.

11. The method according to claim 10, wherein the direct RNA sequencing method is nanopore sequencing.

12. The method according to any one of claims 1 to 8, wherein training the at least one said model within the framework comprises extracting the sequence context read j from each sequence.

13. A system for training a framework comprising one or more models for site-specific ribonucleic acid (RNA) modification detection on an RNA sequence of interest, comprising:a receiver module, for receiving a training dataset of sequences, each sequence being either an unmodified RIMA sequence or a modified RNA sequence and, for each sequence, annotations indicating whether the respective sequence contains a modification at one or more sites;an extraction module, for extracting a sequence context read j from each sequence, and extracting, for each site of interest i in read j, i being a site at which it is desired to detect a modification, a site-specific annotation from the annotations, the site-specific annotation indicating whether a specific modification m occurs at site i;a training module, for training at least one said model to determine a probability of modification m at site i, using a plurality of k-mers, extracted from each read j, including one or more k-mers from before site i, at site i, and after site i; and an output module, for updating the at least one model based on a loss between the probability and the site-specific annotation and outputting the framework, after a predefined number of epochs for training has been completed.