Methods for designing and utilizing extensive libraries of rna modifications comprising rna modifications all around a sequence and uses thereof

By constructing an RNA oligomer library containing RNA modifications and all surrounding random sequences, the limitations of existing RNA modification detection technologies are overcome, enabling high-precision detection of multiple RNA modifications.

CN122095100APending Publication Date: 2026-05-26SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
Filing Date
2024-10-25
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies cannot accurately and broadly detect RNA modifications, especially in predicting a variety of naturally occurring RNA modification sites.

Method used

RNA oligomers containing multiple RNA modifications and all surrounding random sequences are synthesized using chemical methods or template-independent enzymatic synthesis. An RNA modification library is constructed, and a high-quality reference dataset is generated by single-molecule RNA sequencing for use in deep learning and machine learning training.

Benefits of technology

It achieves high-precision detection of multiple RNA modifications, overcomes the limitations of existing methods, and can detect multiple RNA modifications simultaneously, thus improving the accuracy of the detection software.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122095100A_ABST
    Figure CN122095100A_ABST
Patent Text Reader

Abstract

This invention relates to a method for designing and utilizing a broad RNA modification library that includes all surrounding sequences of the RNA modification. By introducing multiple types of RNA modifications, this invention enables the construction of RNA modification libraries that encompass all possible combinations of surrounding sequences flanking the modified RNA. Therefore, it differs from existing reported methods for constructing RNA modification libraries. Furthermore, using the RNA modification library of this invention for single-molecule RNA sequencing generates a high-quality reference dataset containing all possible combinatorial motifs. If this dataset is used as training data for deep learning (or machine learning), it overcomes the limitations of existing software that is limited to specific RNA modifications or has low prediction accuracy for other modifications. This allows for the development of novel, more accurate RNA modification detection software and the simultaneous detection of multiple RNA modifications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for designing and utilizing a broad RNA modification library that includes all surrounding sequences of the RNA modification. Background Technology

[0002] RNA modification refers to the chemical changes in specific bases of RNA. Approximately 170 diverse RNA modifications exist in the RNA of all organisms, including m6A (N6-methyladenosine) and m5C (5-methylcytidine). RNA modifications are distributed at multiple locations in various RNA molecules and have a variety of biological functions. For example, they can alter base pairing or secondary structure and the binding ability with RNA-binding proteins. Furthermore, they participate in the overall RNA metabolism process, including RNA splicing and stability regulation, thus playing a crucial role in gene expression regulation. Literature such as RNA23.12 (2017):1754-1769 and Nature Cell Biology 21.5 (2019):552-559 have reported that abnormal regulation of RNA modification may lead to various diseases, including cancer. In addition, RNA modification, as a regulatory factor, plays an important role in biological processes such as fertilization capacity, development, and cell fate determination. This has been reported in publications such as Cell Research 27 (2017): 1100-1114 and Cell Stem Cell 15.6 (2014): 707-719.

[0003] As the biological functions and importance of RNA modifications are increasingly reported, there is a growing number of attempts to detect specific RNA modifications. However, due to the limitations of traditional antibody-based next-generation sequencing methods in experimental and practical applications, no method has yet been reported that can accurately and broadly detect multiple RNA modifications at once.

[0004] In recent years, in various models used to detect RNA modifications, the RNA modification data used for training has typically been synthetic RNA sequences or RNA sequences extracted from cells.

[0005] Synthetic RNA sequences refer to RNA generated through in vitro transcription (IVT). To construct RNA containing RNA modifications using IVT, specific types of RNA-modified NTPs (e.g., m6ATP, m5CTP) are typically used to replace NTPs (ATP, CTP, GTP, UTP). For example, when using m6ATP, all A atoms on the DNA are transcribed into m6A in the RNA. Therefore, the RNA synthesized via IVT will differ from actual biological sequences. That is, RNA sequences containing m6A at all A sites are constructed, unlike sequences actually existing in nature. The model in Nucleic Acids Research 49.2 (2021): e7, which used this sequence as training data, was confirmed by Nature Communications 14 (2023): 1906 to have very low accuracy in detecting RNA modifications present in actual cells. To overcome this shortcoming, in Nature Communications 10.1 (2019): 4079 and Nature Biotechnology 39 (2021): 1278-1291, 5-mer sequences containing other A sequences besides the central m6A were excluded from the training data, and the model was trained using only sequences that did not contain other A sequences. However, this method still suffers from the drawback of not being able to account for all RNA sequences with canonical A sequences surrounding naturally occurring m6A, and its accuracy remains low. When trained using data generated from synthetic RNA that differs from naturally occurring RNA or cannot account for all naturally occurring sequences, the model has the limitation of failing to successfully predict m6A sites with multiple surrounding sequences.

[0006] On the other hand, when using RNA extracted from cells as training data, m6A present in specific motifs is employed. For example, m6A is known to be predominantly located at the middle A site of the DRACH (D=A / G / U, R=A / G, H=A / C / U) motif. In models in RNA 26.1 (2020): 19-28 or Nature Methods 19.12 (2022): 1590-1598, when only this DR(m6A)CH sequence is screened and used as training data, only restricted predictions of m6A present in this motif are possible, thus limiting the ability to broadly detect naturally occurring m6A. Furthermore, for other RNA modifications besides m6A, since motifs predominantly present like DRACH are rarely reported, the strategy of using RNA extracted from cells or using RNA modification data on specific motifs based on information obtained from them is not applicable to other RNA modifications.

[0007] As mentioned above, existing literature and models suggest that using RNA modification sequences synthesized via IVT, extracted from cells, or restricted to specific motifs cannot account for all RNA modifications present throughout the organism, and may have limitations in predicting overall RNA modification sites when used as model training data. Based on this, the inventors, through in-depth research, aimed to design an RNA modification library capable of reflecting all RNA modifications present throughout the organism. As a result, a design method was developed to synthesize RNA oligomers containing multiple RNA modifications and all surrounding random sequences at desired positions via chemical methods or template-independent enzymatic synthesis, and these oligomers were used to construct an RNA modification library, thus completing this invention. Summary of the Invention

[0008] The technical problem that the invention aims to solve

[0009] The purpose of this invention is to provide a method for constructing an RNA modification library based on a design method for a broad RNA modification library containing all surrounding sequence combinations of RNA modification, and its applications.

[0010] means for solving problems

[0011] To achieve the aforementioned objective, the present invention provides a method for constructing an RNA modification library, comprising the following steps: (a) Construct one or more modified RNA blocks, and randomly link 4 to 50 random nucleotides in the 5′ and 3′ directions of the modified RNA, respectively; (b) Construct modified RNA oligomers such that 1 to 9 modified RNA blocks are randomly selected from the one or more modified RNA blocks; (c) The modified RNA oligomers are ligated to construct the modified RNA oligomer ligation product; (d) Screen the modified RNA oligomer ligation products, wherein the modified RNA oligomers are ligated from 2 to 20; (e) Attaching an RNA tail consisting of 5 to 150 arbitrary ribonucleotides to the selected modified RNA oligomer ligation product; and (f) Recover the modified RNA oligomer ligation product with the RNA tail attached.

[0012] In this invention, the modified RNA can be any one of the following modified RNAs from (1) to (4): (1) Modification A shown in Table 1: Table 1

[0013] (2) Modification C shown in Table 2: Table 2

[0014] (3) Modification G shown in Table 3: Table 3

[0015] as well as

[0016] (4) Modifications shown in Table 4.

[0017] Table 4

[0018] In this invention, step (b) may be to construct a modified RNA oligomer, wherein the modified RNA oligomer further includes an anchoring sequence at the following position: (i) Between the modified RNA blocks that make up the modified RNA oligomer; (ii) the 5' end of the modified RNA block located at the 5' end in the modified RNA block constituting the modified RNA oligomer; and / or (iii) The 3′ end of the modified RNA block located at the 3′ end in the modified RNA block that constitutes the modified RNA oligomer.

[0019] In this invention, the anchoring sequence has a sequence of 4 or more nucleotides. If the modified RNA oligomer contains 2 or fewer modified RNA blocks, the anchoring sequence has a length of 40 to 90% of the length of the modified RNA blocks. If the modified RNA oligomer contains 3 to 9 modified RNA blocks, the anchoring sequence can have a length of 20 to 30% of the length of the modified RNA blocks.

[0020] In this invention, step (c) can be ligation using RNA ligase.

[0021] In this invention, step (b) can be used to construct the modified RNA oligomer by introducing 2 to 4 specific nucleotide combinations at the 5′ and 3′ ends of the modified RNA oligomer, respectively.

[0022] In this invention, the RNA ligase can be T4 RNA ligase 1, the two specific nucleotide combinations are any of the combinations listed in Table 5 below, the three specific nucleotide combinations are any of the combinations listed in Table 6 below, and the four specific nucleotide combinations are any of the combinations listed in Table 7 below. Table 5

[0023] Table 6

[0024] Table 7

[0025] Alternatively, the RNA ligase may be a truncated form of T4 RNA ligase 2, the two specific nucleotide combinations may be any of the combinations listed in Table 8 below, the three specific nucleotide combinations may be any of the combinations listed in Table 9 below, and the four specific nucleotide combinations may be any of the combinations listed in Table 10 below.

[0026] Table 8

[0027] Table 9

[0028] Table 10

[0029] In this invention, steps (a) and (b) can be constructed using chemical synthesis or template-independent enzymatic synthesis. Step (b) can be constructed as an independent process after step (a), or steps (a) and (b) can be constructed as a single process.

[0030] In this invention, step (c) can be performed at 29°C to 37°C for 8 to 40 hours.

[0031] In this invention, step (c) can be performed in a reaction solution with PEG8000 at a concentration of 13% to 21% (v / v).

[0032] In this invention, the RNA tail in step (e) can be a poly(A) tail, a poly(U) tail, or a poly(I) tail, which consists of 5 to 40 ribonucleotides.

[0033] In this invention, the poly(A) tail, poly(U) tail, or poly(I) tail in step (e) can be reacted in a solution containing 5 to 50 μM of adenosine triphosphate (ATP), uridine triphosphate (UTP), or inosine triphosphate (ITP).

[0034] In this invention, an RNA modification library is provided, comprising a plurality of modified RNA oligomer ligation products, wherein the modified RNA oligomer ligation products are in the form of ligating 2 to 20 modified RNA oligomers, characterized in that... The modified RNA oligomer has modified RNA blocks with 4 to 50 random nucleotides randomly linked in the 5′ and 3′ directions, respectively, arranged in 1 to 9 random consecutive sequences, with the modified RNA at the center.

[0035] In this invention, the modified RNA is any one of the modified RNAs selected from the group consisting of (1) to (4) below: (1) Any modified RNA in the group consisting of modification A shown in Table 1; (2) Any modified RNA in the group consisting of C modification shown in Table 2; (3) Any modified RNA in the group consisting of the G modifications shown in Table 3; (4) Any modified RNA in the group consisting of the U modification shown in Table 4.

[0036] In this invention, the modified RNA is any one of the modified RNAs selected from the group consisting of (1) to (4) below: (i) Between the modified RNA blocks that make up the modified RNA oligomer; (ii) the 5' end of the modified RNA block located at the 5' end in the modified RNA block constituting the modified RNA oligomer; and / or (iii) The 3′ end of the modified RNA block located at the 3′ end in the modified RNA block that constitutes the modified RNA oligomer.

[0037] In this invention, the anchoring sequence has a sequence of 4 or more nucleotides. If the modified RNA oligomer contains 2 or fewer modified RNA blocks, the anchoring sequence has a length of 40 to 90% of the length of the modified RNA blocks. If the modified RNA oligomer contains 3 to 9 modified RNA blocks, the anchoring sequence may have a length of 20 to 30% of the length of the modified RNA blocks.

[0038] In this invention, the 5′ and 3′ ends of the modified RNA oligomer may each be additionally introduced with 2 to 4 specific nucleotide combinations.

[0039] In this invention, the 3′ end of the modified RNA oligomer ligation product may be further attached with an RNA tail consisting of 5 to 150 arbitrary ribonucleotides.

[0040] This invention provides an RNA library collection for training an RNA modification detection model, comprising: an RNA modification library according to any one of claims 13 to 18; and

[0041] A general RNA library is provided, the general RNA library comprising multiple general RNA oligomer ligation products, wherein the multiple general RNA oligomer ligation products have the same structure as the multiple modified RNA oligomer ligation products contained in the RNA modification library, except that the modified RNA is replaced by the corresponding general RNA.

[0042] In this invention, a sequencing dataset is provided, which is generated from an RNA library collection used for training the RNA modification detection model described above.

[0043] In this invention, the sequencing can be single-molecule RNA sequencing.

[0044] In this invention, the sequencing dataset can be used for deep learning and / or machine learning training.

[0045] The effects of the invention

[0046] This invention, by introducing multiple types of RNA modifications, enables the construction of RNA modification libraries that encompass all possible combinations of surrounding sequences flanking the modified RNA. Therefore, it differs from existing reported methods for constructing RNA modification libraries. Furthermore, using the RNA modification libraries of this invention for single-molecule RNA sequencing generates a high-quality reference dataset containing all possible combinatorial motifs. Moreover, using this dataset as training data for deep learning (or machine learning) overcomes the limitations of existing software, which is restricted to specific RNA modifications or has low prediction accuracy for other modifications. This allows for the development of novel, more accurate RNA modification detection software and enables the simultaneous detection of multiple RNA modifications. Attached Figure Description

[0047] Figure 1 The diagram illustrates a modified RNA oligomer constructed according to an embodiment of the present invention and its corresponding general RNA oligomer.

[0048] Figure 2 This is a schematic diagram illustrating the construction experiment of the RNA modified library and its corresponding general RNA library according to the present invention.

[0049] Figure 3a The diagram illustrates the ranking of ligation efficiency based on the combination of sequences at both ends of the RNA oligomers during the step of ligating RNA oligomers according to the present invention using T4 RNAligase 1.

[0050] Figure 3b The diagram illustrates the ranking of ligation efficiency based on the combination of sequences at both ends of the RNA oligomers during the ligation of RNA oligomers according to the present invention using T4 RNAligase 2 truncated KQ.

[0051] Figure 4aThis is a schematic diagram illustrating an experiment performed using various RNA ligases, namely RNA ligase 1, RNA ligase 2 truncated KQ, and RNA ligase 2, to ligate RNA oligomers according to the present invention.

[0052] Figure 4b The figure shows the results of polyacrylamide gel electrophoresis (PAGE) of the RNA oligomers according to the present invention after ligation using various RNA ligases, namely RNA ligase 1, RNA ligase 2 truncated KQ, and RNA ligase 2.

[0053] Figure 5a The figure shows the results of confirming the ligation efficiency based on the reaction temperature in the step of ligating RNA oligomers according to the present invention using T4 RNAligase 1.

[0054] Figure 5b The figure shows the results of confirming the ligation efficiency based on the PEG8000 concentration in the step of ligating RNA oligomers according to the present invention using T4 RNAligase 1.

[0055] Figure 5c The figure shows the results of confirming the ligation efficiency based on the DMSO concentration in the step of ligating RNA oligomers according to the present invention using T4 RNAligase 1.

[0056] Figure 5d The figure shows the results of confirming the ligation efficiency based on reaction time in the step of ligating RNA oligomers according to the present invention using T4 RNAligase 1.

[0057] Figure 5e The graph shows the results of comparing the ligation efficiency of RNA oligomers under optimized conditions established for ligating RNA oligomers according to the present invention with that under conventional conditions using T4 RNA ligase 1.

[0058] Figure 6The figure shows the confirmation results of the optimized poly(A) tailing length when performing nanopore sequencing using the RNA modified library constructed according to the present invention.

[0059] Figure 7a The figure illustrates the confirmation results of ATP concentration for achieving optimized length poly(A)tailing in the RNA modification library constructed according to the present invention.

[0060] Figure 7b The figure shows the results of confirming the amount of E-PAP added and the reaction time for achieving optimized length poly(A)tailing in the RNA modification library constructed according to the present invention.

[0061] Figure 7c This figure illustrates the confirmation results of achieving accurate RNA tail attachment by ligation in the construction of the RNA modification library of the present invention.

[0062] Figure 7d The figure illustrates the result of constructing an RNA modification library by App-RNA ligation of the modified RNA oligomer ligation product constructed according to the present invention to attach an RNA tail.

[0063] Figure 8a The figure illustrates the m6A library constructed according to the present invention, a general A library, and a practical example of single-molecule sequencing using them, along with the results (number of reads, number of bases).

[0064] Figure 8b The figure illustrates the m5C library constructed according to the present invention, a general C library, and a practical example of single-molecule sequencing using them, along with the results (number of reads, number of bases).

[0065] Figure 9a The figure illustrates the results of base quality confirmation based on the modified RNA sites in the motifs extracted from sequencing data of the m6A library constructed according to the present invention and a general A library.

[0066] Figure 9b The figure illustrates the results of confirming the motifs extracted from sequencing data of the m6A library constructed according to the present invention and a general A library, based on the average current signal of the modified RNA site.

[0067] Figure 9c The figure illustrates the results of confirming the motifs based on differences in current signals extracted from sequencing data of an m6A library constructed according to the present invention and a general A library.

[0068] Figure 10a This figure illustrates the results of comparing the degree of improvement in RNA oligomer design based on the anchor sequence insertion method.

[0069] Figure 10b This figure illustrates the results of comparing motif extraction accuracy based on the number of RNA segments and the length of the anchor sequence.

[0070] Figure 10c This diagram illustrates an example of the anchoring sequence obtained by inserting between blocks in the design of modified RNA oligomers and their corresponding general RNA oligomers. Detailed Implementation

[0071] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Generally, the nomenclature used in this specification and the experimental methods described below are methods well-known and commonly used in the art.

[0072] In this invention, terms such as "comprising" or "having" are intended to indicate the presence of features, values, steps, actions, constituent elements, components, or combinations thereof described in the specification, and should be understood as not precluding the possibility of the presence or addition of one or more other features, values, steps, actions, constituent elements, components, or combinations thereof.

[0073] This invention develops a method for designing a broad RNA modification library containing all combinations of surrounding sequences of the RNA modification. Specifically, this invention designs and chemically synthesizes RNA oligomers containing multiple RNA modifications and all random sequences surrounding them at desired positions, and uses these oligomers to construct RNA modification libraries.

[0074] Therefore, one aspect of the present invention relates to a method for constructing an RNA modification library, which includes the following steps: (a) Construct one or more modified RNA blocks, and randomly link 4 to 50 random nucleotides in the 5′ and 3′ directions of the modified RNA, respectively; (b) Constructing modified RNA oligomers in which 1 to 9 modified RNA blocks are randomly arranged in one or more modified RNA blocks; (c) The modified RNA oligomers are ligated to construct the modified RNA oligomer ligation product; (d) Screen the modified RNA oligomer ligation products, wherein the modified RNA oligomers are ligated from 2 to 20; (e) Attaching an RNA tail consisting of 5 to 150 arbitrary ribonucleotides to the selected modified RNA oligomer ligation product; and (f) Recover the modified RNA oligomer ligation product with the RNA tail attached.

[0075] As one approach, the construction of an RNA modification library may include the following steps: (a) Construct one or more modified RNA blocks, and randomly link 4 to 50 random nucleotides in the 5′ and 3′ directions of the modified RNA, respectively; (b) Constructing modified RNA oligomers in which 1 to 9 modified RNA blocks are randomly arranged in one or more modified RNA blocks; (c) The modified RNA oligomers are ligated to construct the modified RNA oligomer ligation product; (d) Screen the modified RNA oligomer ligation products, wherein the modified RNA oligomers are ligated from 2 to 20; (e) Attaching an RNA tail consisting of 5 to 150 arbitrary ribonucleotides to the selected modified RNA oligomer ligation product; and (f) Recover the modified RNA oligomer ligation product with the RNA tail attached.

[0076] In this invention, the modified RNA is any one of the modified RNAs selected from the group consisting of (1) to (4) below: (1) Modification A shown in Table 1: Table 1

[0077] (2) Modification C shown in Table 2: Table 2

[0078] (3) Modification G shown in Table 3: Table 3

[0079] as well as

[0080] (4) Modifications shown in Table 4.

[0081] Table 4

[0082] That is, in this invention, the modified RNA can be any of the various modified RNAs listed in Tables 1 to 4.

[0083] In step (a), random numbers of the same or different nucleotides can be randomly linked to both sides of the modified RNA. These nucleotides can be any of the A, G, C, and U types of common RNA, and the number of nucleotides linked to each side can be from 4 to 50. Through this process, modified RNA blocks with various combinations can be constructed centered around RNA modification.

[0084] In step (b), one to nine modified RNA blocks can be randomly selected from the modified RNA blocks with multiple combinations constructed in step (a) and arranged to construct modified RNA oligomers. This arrangement can further improve the diversity of the library according to the present invention.

[0085] As one implementation, various combinations of the following structures exist: In the modified RNA oligomer, modified RNA blocks with 4 to 50 random nucleotides randomly linked in the 5′ and 3′ directions, respectively, are arranged in 2 to 9 random consecutive sequences, centered on the modified RNA.

[0086] In this invention, step (b) may be to construct a modified RNA oligomer, wherein the modified RNA oligomer further includes an anchoring sequence at the following position: (i) Between the modified RNA blocks that make up the modified RNA oligomer; (ii) the 5' end of the modified RNA block located at the 5' end in the modified RNA block constituting the modified RNA oligomer; and / or (iii) The 3′ end of the modified RNA block located at the 3′ end in the modified RNA block that constitutes the modified RNA oligomer.

[0087] In this invention, the anchoring sequence is characterized by a sequence of four or more nucleotides. If the modified RNA oligomer contains two or fewer modified RNA blocks, the anchoring sequence has a length of 40 to 90% of the length of the modified RNA blocks. If the modified RNA oligomer contains three to nine modified RNA blocks, the anchoring sequence may have a length of 20 to 30% of the length of the modified RNA blocks.

[0088] As the length of the modified RNA block increases, the accuracy of the motif decreases. In this case, the length of the anchoring sequence can be adjusted as described above to maintain a high level of motif accuracy.

[0089] For example, when the modified RNA oligomer contains 3 or more modified RNA blocks and the length of the modified RNA block is 21, the anchoring length is suitable to be about 6; when the length of the modified RNA block is 31, the anchoring length is suitable to be about 9.

[0090] However, when the number of modified RNA blocks is less than two, the accuracy of motif extraction decreases. In this case, the length of the anchor sequence can be further increased to maintain a high motif accuracy.

[0091] For example, when the modified RNA block length is 21 and the number of modified RNA blocks is 2, the appropriate anchoring length is approximately 12; when the number of modified RNA blocks is 1, the appropriate anchoring length is approximately 18. When the modified RNA block length is 31 and the number of modified RNA blocks is 2, the appropriate anchoring length is approximately 18; when the number of modified RNA blocks is 1, the appropriate anchoring length is approximately 27.

[0092] The preferred anchor sequence is selected based on the following factors.

[0093] (1) The Levenshtein distance between each anchoring sequence should be at least 50% of the anchoring length.

[0094] (2) When all anchor sequences are combined, the proportion of each nucleotide should be equal, and the proportion of nucleotides within each anchor sequence should also be equal. An equal proportion means dividing the length of each anchor sequence into four integer ratios that are closest to 1:1:1:1.

[0095] (3) The average alignment error of all 5-mers that make up each anchor sequence during nanopore sequencing should be below the 10th percentile.

[0096] (4) When constructing partially random oligomers using all anchored sequences and sampled random blocks, the average value of the minimum free energy (MFE) should be above the 90th percentile to minimize the formation of RNA secondary structures. This is to minimize the formation of RNA secondary structures.

[0097] (5) When constructing partially random oligomers as described in (4), the average value of the minimum free energy (dimer MFE) of the dimers produced in all possible pairings should be above the 90th percentile, thereby minimizing the formation of RNA dimers. This is to minimize the formation of RNA dimers.

[0098] (6) When extracting all subsequences with 50% of the anchor length from all anchor sequences, there should be no reverse complementary sequence pairs among all possible pairings formed by them. This is to minimize the formation of RNA dimers.

[0099] Therefore, the anchoring sequence preferably satisfies one or more of the following conditions: (1) The Levenshtein distance between each anchoring sequence shall be at least 50% of the anchoring length; (2) The ratio of each nucleotide (A, G, C, U) within each anchored sequence does not deviate by more than 10% from the 1:1:1:1 ratio; (3) The average alignment error of all 5-mers constituting each anchor sequence during nanopore sequencing is below the 10th percentile; (4) When using all anchor sequences and sampled random blocks to construct partially random oligomers, the average value of their minimum free energy (MFE) is above the 90th percentile; (5) When constructing partially random oligomers using all anchored sequences and sampled random blocks, the average minimum free energy (dimer MFE) of the dimers produced across all possible pairings is above the 90th percentile; and (6) When all subsequences with an anchor length of 50% are extracted from all anchor sequences, there are no reverse complement sequences among all possible pairings formed by them.

[0100] For example, a set of four anchor sequences with a length of 6 nucleotides that meet the above conditions includes {CGACAU,CAGUUA, GUCCAG, GUAGUC}, {CGACAU, AGUCCG, CAGUUA, GUAGUC}, {CGACAU, GUAUCC,UUGACG, AGAGUC}, etc.

[0101] In this invention, the multiple anchoring sequences in the modified RNA oligomer may have the same length or different lengths, and may have the same sequence or different sequences.

[0102] In this invention, step (c) can be the ligation of modified RNA oligomers using an RNA ligase to construct modified RNA oligomer ligates, and the RNA ligase can be used without limitation on its type.

[0103] In this case, step (b) involves randomly arranging 1 to 9 arbitrary modified RNA blocks from one or more modified RNA blocks constructed in step (a) to construct a modified RNA oligomer, and 2 to 4 specific nucleotide combinations may be further introduced at the 5′ end and 3′ end of the modified RNA oligomer thus constructed.

[0104] For example, step (b) involves randomly arranging 2 to 9 arbitrary modified RNA blocks from the plurality of modified RNA blocks constructed in step (a) to construct a modified RNA oligomer, and 2 to 4 specific nucleotide combinations may be further introduced at the 5′ end and 3′ end of the modified RNA oligomer thus constructed.

[0105] When the specific nucleotide combination described above is introduced, the ligation efficiency can be improved in the subsequent step (c) of ligating the modified RNA oligomers to construct the modified RNA oligomer ligation product.

[0106] In one embodiment, the RNA ligase can be T4 RNA ligase 1, the two specific nucleotide combinations can be any of the combinations shown in Table 5, the three specific nucleotide combinations can be any of the combinations shown in Table 6, and the four specific nucleotide combinations can be any of the combinations shown in Table 7, but are not limited thereto: Table 5

[0107] Table 6

[0108] Table 7

[0109] In another embodiment, the RNA ligase may be T4 RNA ligase 2 truncated KQ, the two specific nucleotide combinations may be any of the combinations listed in Table 8 below, the three specific nucleotide combinations may be any of the combinations listed in Table 9 below, and the four specific nucleotide combinations may be any of the combinations listed in Table 10 below.

[0110] Table 8

[0111] Table 9

[0112] Table 10

[0113] This invention enables the construction of an RNA modification library through a single ligation reaction.

[0114] In this invention, steps (a) and (b) can be constructed using chemical methods or template-independent enzymatic synthesis known in the art. In this case, step (b) can be constructed as a separate process after step (a), or steps (a) and (b) can be constructed as a single process.

[0115] That is, in this invention, steps (a) and (b) can be constructed as a process as described below: (a') Construct modified RNA oligomers, in which 1 to 9 modified RNA blocks are randomly arranged in one or more modified RNA blocks with 4 to 50 random nucleotides randomly linked in the 5′ and 3′ directions of the modified RNA.

[0116] In one embodiment of the present invention, steps (a) and (b) are constructed by chemical synthesis. However, those skilled in the art will understand that they can also be constructed by template-independent enzymatic synthesis. The template-independent enzymatic synthesis method is described in detail, for example, in Wiegand, Daniel J., et al. “Template-independent enzymatic synthesis of RNA oligonucleotides.” Nature Biotechnology (2024). Those skilled in the art can refer to this literature to replace the chemical synthesis method with the template-independent enzymatic synthesis method to implement the present invention.

[0117] As another embodiment, in this invention, steps (a) and (b) can be constructed as a process as described below: (a') constructing modified RNA oligomers, such that 2 to 9 modified RNA blocks are randomly arranged in a plurality of modified RNA blocks in which 4 to 50 random nucleotides are randomly linked in the 5' and 3' directions of the modified RNA.

[0118] In this case, the additional process in step (b) can be directly applied to step (a') without contradiction.

[0119] In this invention, step (c) can be a connection performed for about 4 hours or more, for example about 8 hours to 40 hours, preferably about 10 hours to 24 hours, at a temperature of about 29°C to about 37°C to improve connection efficiency, for example, at a temperature of about 31°C to about 34°C, preferably about 32°C to about 33°C.

[0120] In this invention, step (c) may also be performed in a reaction solution containing about 11 to about 23% (v / v) of PEG8000, for example about 13 to about 21% (v / v), preferably about 15 to 19% (v / v), in order to improve the connection efficiency.

[0121] In this invention, step (c) may also involve, in order to improve the connection efficiency, the connection is carried out in a reaction solution containing DMSO, in which case the DMSO may be less than about 20% (v / v), for example about 5 to about 15% (v / v), preferably about 8 to about 12% (v / v).

[0122] Subsequently, in step (d), modified RNA oligomer ligation products with 2 to 20 modified RNA oligomers can be screened for use in RNA sequencing.

[0123] The step of screening modified RNA oligomer ligation products with a specific length is a well-known method in the art. As one implementation, it can be carried out by separating the modified RNA oligomer ligation products by electrophoresis on a PAGE gel and extracting the modified RNA oligomer ligation products with the desired size, but is not limited thereto.

[0124] Subsequently, in step (e), an RNA tail can be attached to the 3′ end of the modified RNA oligomer ligation product of a specific length selected.

[0125] In this context, the RNA tail is attached to improve RNA sequencing efficiency. It represents a well-defined random sequence whose exact sequence is known to the builder of the RNA modified library. The RNA tail can be a homonucleotide or heteronucleotide composed of about 5 to 150 arbitrary ribonucleotides (A, G, C, U, I).

[0126] The RNA tail can be composed of ribonucleotides selected from any ribonucleotide (A, G, C, U, I), such as 25 to 100 homonucleotides or heteronucleotides, preferably about 25 to 50 homonucleotides or heteronucleotides, more preferably about 25 to 40 homonucleotides or heteronucleotides.

[0127] In one embodiment, the RNA tail can be a poly(A) tail, which can be composed of 5 to 150 adenosines, for example 25 to 100 adenosines, preferably about 25 to 50 adenosines, and more preferably about 25 to 40 adenosines.

[0128] As another embodiment, the RNA tail can be a poly(U) tail, which can be composed of 5 to 150 uridines, for example 25 to 100 uridines, preferably about 25 to 50 uridines, and more preferably about 25 to 40 uridines.

[0129] As another embodiment, the RNA tail can be a poly(I) tail, which can be composed of 5 to 150 inosines, for example 25 to 100 inosines, preferably about 25 to 50 inosines, and more preferably about 25 to 40 inosines.

[0130] In this invention, the RNA tail may be introduced in the form of a poly(A) tail, a poly(U) tail, or a poly(I) tail for ease of construction. However, as long as the builder can clearly determine its sequence and can easily construct aptamers or primers for sequencing based on it, its type is not particularly limited. It can be constructed by linking ribonucleotides selected from any ribonucleotides (A, G, C, U, I) into a length range of about 5 to 150.

[0131] In this invention, for the formation of a poly(A) tail of about 25 to 40 adenosines, the poly(A) tail reaction solution may contain about 5 to 50 μM, for example about 5 to 25 μM, preferably about 5 to 15 μM of adenosine triphosphate (ATP).

[0132] In this case, the poly(A) tail reaction can be carried out using about 2.5 U to 4 U of Escherichia coli polyadenylate polymerase (E-PAP), for example about 3 to 3.5 U of E-PAP, for about 30 minutes to about 1 hour, for example about 40 minutes to about 50 minutes.

[0133] In one embodiment of the present invention, poly(A) tailing is performed using E. coli poly(A) polymerase (E-PAP). However, those skilled in the art will understand that poly(U) tailing or poly(I) tailing can also be performed using poly(U) polymerase. Poly(U) tailing or poly(I) tailing is described in detail, for example, in RNA 27 (2021): 1497-1511 and Molecular and Cellular Biology 27.10 (2007): 3612-3624.

[0134] On the other hand, in this invention, when the RNA tail is a homonucleotide or heteronucleotide other than a poly(A) tail, a poly(U) tail, or a poly(I) tail, the RNA tail can be constructed to have 25 to 100 homonucleotides or heteronucleotides, preferably about 25 to 50 homonucleotides or heteronucleotides, more preferably about 25 to 40 homonucleotides or heteronucleotides, and has phosphate groups (5′p-RNA-3′p) at its 5′ and 3′ ends. After the 5′-p is pre-adenosylated (to generate 5′ App-RNA-3′p), it is attached to the modified RNA oligomer linker product by ligation.

[0135] On the other hand, a method for constructing a general RNA library includes: comprising a plurality of general RNA oligomer ligation products, wherein, except for replacing the modified RNA with a corresponding general RNA, the plurality of general RNA oligomer ligation products have the same structure as the plurality of modified RNA oligomer ligation products contained in the RNA modification library. This method can be readily understood by those skilled in the art from the above-described method for constructing an RNA modification library.

[0136] In this invention, having the same structure means that, except for replacing the modified RNA with ordinary RNA, an ordinary RNA library is constructed through the same steps and conditions, thereby having a structure corresponding to the RNA modified library.

[0137] In another aspect, the present invention relates to RNA modification libraries constructed by the above-described methods.

[0138] In another aspect, the present invention relates to an RNA modification library comprising a plurality of modified RNA oligomer ligation products, wherein the modified RNA oligomer ligation products are in the form of ligating 2 to 20 modified RNA oligomers, wherein the modified RNA oligomers are centered on the modified RNA and the modified RNA blocks of 4 to 50 random nucleotides are randomly ligated in the 5′ and 3′ directions respectively, arranged in a random continuous sequence of 1 to 9.

[0139] For example, the modified RNA oligomer can be: a modified RNA block with 4 to 50 random nucleotides randomly linked in the 5′ and 3′ directions, centered on the modified RNA, arranged in 2 to 9 random consecutive blocks.

[0140] In this invention, the modified RNA may be any modified RNA selected from the group consisting of a variety of modified RNAs listed in Tables 1 to 4, but is not limited thereto.

[0141] In this invention, the modified RNA oligomer may include an anchoring sequence located at the following position: (i) Between the modified RNA blocks that make up the modified RNA oligomer, (ii) the 5' end of the modified RNA block located at the 5' end in the modified RNA block constituting the modified RNA oligomer; and / or (iii) The 3′ end of the modified RNA block that constitutes the modified RNA oligomer.

[0142] In this invention, the anchoring sequence has a sequence of 4 or more nucleotides. When the modified RNA oligomer contains 2 or fewer modified RNA blocks, the anchoring sequence has a length of 40 to 90% of the length of the modified RNA blocks. When the modified RNA oligomer contains 3 to 9 modified RNA blocks, the anchoring sequence has a length of 20 to 30% of the length of the modified RNA blocks.

[0143] In this invention, the 5′ and 3′ ends of the modified RNA oligomer can be further introduced with 2 to 4 specific nucleotide combinations, respectively.

[0144] In this invention, the modified RNA oligomer ligation product may further include an RNA tail consisting of 5 to 150 arbitrary ribonucleotides at its 3′ end.

[0145] The RNA tail has been described in detail above and will not be repeated here.

[0146] In another aspect, the present invention relates to an RNA library collection for training an RNA modification detection model, comprising: The RNA modification library; and A general RNA library comprising multiple general RNA oligomer ligation products, wherein the multiple general RNA oligomer ligation products have the same structure as the multiple modified RNA oligomer ligation products contained in the RNA modification library, except that the modified RNA is replaced by the corresponding general RNA.

[0147] In another aspect, the present invention relates to a method for training an RNA modification detection model using the aforementioned RNA library set, wherein the RNA library set includes: the RNA modification library; and

[0148] A general RNA library comprising multiple general RNA oligomer ligation products, wherein the multiple general RNA oligomer ligation products have the same structure as the multiple modified RNA oligomer ligation products contained in the RNA modification library, except that the modified RNA is replaced by the corresponding general RNA.

[0149] As one implementation, the RNA modification library and the corresponding general RNA library can be provided as components of the kit.

[0150] In this case, the kit may include the RNA modification library and a general RNA library corresponding to the RNA modification library, and may include instructions explaining the principles and experimental methods for generating sequencing data using the library set.

[0151] In another aspect, the present invention relates to sequencing datasets generated using RNA library sets trained with the said RNA modification detection model.

[0152] In this invention, the sequencing can be single-molecule RNA sequencing, but is not limited thereto.

[0153] The single-molecule RNA sequencing method can be any method known in the art, for example, see Nature Methods 15(2018): 201-206, but is not limited thereto.

[0154] For example, the sequencing dataset can be used for deep learning and / or machine learning training.

[0155] Therefore, in another aspect, the present invention relates to a training method for deep learning or machine learning using the sequencing dataset.

[0156] In another aspect, the present invention relates to the sequencing dataset used for training deep learning or machine learning.

[0157] In this invention, the oligomers in the RNA-modified library according to the invention have the following advantages: (1) They overcome the disadvantages of RNA constructed by IVT (the same RNA modification exists at all base sites), and can construct libraries with RNA modifications only at desired sites; (2) The RNA modification is centered, and the sequences on both sides are composed of random sequences of desired lengths that are not limited to all combinations of specific motifs. This makes it possible to construct libraries containing specific types of RNA modifications and their corresponding general bases. Thus, the disadvantages of libraries constructed using RNA extracted from cells or RNA synthesized by inserting only a few specific motif sequences can be overcome; (3) This invention synthesizes RNA by replacing only the middle specific base with other types of RNA modifications, so the same strategy can also be applied to the construction of libraries with multiple types of RNA modifications.

[0158] This invention differs from existing reported RNA modification libraries because it is an RNA modification library capable of reflecting various types of RNA modifications and all their surrounding sequences.

[0159] That is, the RNA modification library according to the present invention is not limited to specific RNA modifications and their known surrounding motifs, but has the advantage of including a variety of RNA modifications and all possible surrounding sequences in nature. Furthermore, by overcoming the limitations of RNA constructed by IVT, it is possible to design the introduction of RNA modifications only at the desired sites, and the surrounding sequences can also be designed as random sequences or specific sequences. Moreover, since only one RNA modification is introduced at the central position, it is not necessary to exclude specific combinations of surrounding sequences to avoid including more than two RNA modifications in the library constructed by IVT, as is done in existing literature, thus offering this advantage.

[0160] That is, the RNA modification library designed according to the present invention contains a wide combination of modified RNA oligomer ligation products, and therefore can be used as training data for RNA modification prediction models after generating sequencing data (e.g., single-molecule RNA sequencing data).

[0161] Furthermore, the RNA modification library according to the present invention, because it can introduce RNA modifications at desired locations, is closer to the transcriptome of actual organisms than existing known RNA modification libraries (e.g., those constructed via IVT), and has the advantage of being able to learn knowledge that can be generalized to actual organisms (e.g., deep learning or machine learning). That is, the RNA modification library according to the present invention can generate a dataset containing all possible motifs combined around the modified RNA, thus eliminating motif bias during training.

[0162] On the other hand, software trained using sequencing data generated from the RNA modification library according to the present invention as training data can overcome the limitations of software trained using existing RNA modification library data, which is limited to specific RNA modifications and has low accuracy in predicting other modifications, thereby achieving novel RNA modification detection with very high accuracy.

[0163] Furthermore, the software trained using various types of RNA modification library data designed according to the present invention can not only accurately detect RNA modifications present in all surrounding sequences, but also simultaneously detect multiple RNA modifications.

[0164] RNA modification is known to play a regulatory role in biological processes such as development, fertilization, and cell fate determination. Furthermore, there are ongoing reports indicating that abnormal RNA modification regulation is associated with various diseases, including cancer. Therefore, software utilizing the RNA modification library according to the present invention can also be applied to molecular diagnostics and medical fields related to various diseases. For example, by detecting RNA modifications in RNA obtained from various patient samples, disease-specific RNA modification sites can be identified, and novel RNA modification proteins can be identified by comparing them with RNA-binding protein sites. Moreover, based on the identified RNA modification protein data, candidate biomarkers for the diagnosis of RNA modification-related diseases can be discovered, and by identifying disease-specific drug targets, it can be actively applied in the diagnosis and treatment of RNA modification-related diseases.

[0165] Furthermore, software that uses various datasets derived from the RNA modification library according to the present invention to detect or predict modified RNAs can also be used to identify new associations between various RNA modifications and specific biological phenomena and to analyze their mechanisms. For example, if software trained using the RNA modification library according to the present invention is applied to cells during the maternal-zygotic transition process in embryonic development, it can identify new associations between these changes over time and various RNA modifications, and provide new directions for understanding their mechanisms.

[0166] Example

[0167] The present invention will be described in more detail below through embodiments. These embodiments are for illustrative purposes only and should not be construed as limiting the scope of the invention, as will be apparent to those skilled in the art.

[0168] Example 1: Construction of RNA Modification Library

[0169] In this invention, the aim is to design libraries that take into account all surrounding sequences for various RNA modifications. Figure 1 The specific construction process is as follows ( Figure 2 ).

[0170] (1) Construct modified RNA oligomers with modified RNA block repeats and their corresponding general RNA oligomers.

[0171] A block is formed by combining a specific type of RNA modification (or its corresponding general base) with a sequence of 4 to 50 nucleotides of random flanking it (N=A / G / C / U). Modified RNA oligomers (manufacturer: IDT) containing this modified RNA block repeated 1 to 9 times are constructed via chemical RNA synthesis. However, the modified RNA block can also be constructed using template-independent enzymatic synthesis.

[0172] The modified RNA oligomers were constructed with specific combinations of 2 to 4 nucleotides linked to their 5′ and 3′ ends, respectively. In this case, the optimal nucleotides for the 5′ and 3′ ends of the modified RNA oligomers were determined through additional experiments.

[0173] Specifically, after designing an RNA adapter with a random nucleotide at its 3′ end, T4 RNAligase 1 (NEB) is used to ligate RNA at the 5′ end of the modified RNA oligomer. Figure 3a NGS sequencing was performed using only the ligated RNA to synthesize cDNA. By analyzing the sequence at the ligation site, the 5′ end of the most common RNA-modified RNA oligomer was selected as the sequence combination with the 3′ end of the adapter. Since these combinations exhibit higher reactivity with T4 RNAligase 1, it was considered that designing the ends of the RNA oligomers with this combination could maximize ligation efficiency and yield during library construction using this enzyme. Furthermore, it can also serve as a marker for locating specific intermediate bases (X) after single-molecule sequencing.

[0174] The number of reads for the first 10 combinations was measured, and the enrichment (the increase relative to the random combination) of each combination was analyzed to confirm statistical significance. For 2-nucleotide combinations, all 10 combinations showed significant differences (p-value, q-value ≒0), and in replicates 1 and 2, the increase was confirmed to be more than 5.7-fold compared to the random combination. For 3-nucleotide combinations, all 10 combinations showed significant differences (p-value, q-value ≒0), and in replicates 1 and 2, the increase was confirmed to be more than 68-fold compared to the random combination. For 4-nucleotide combinations, all 10 combinations showed significant differences (p-value, q-value ≒0), and in replicates 1 and 2, the increase was confirmed to be more than 814-fold compared to the random combination. Figure 3a ).

[0175] Similarly, after designing an App-RNA adapter with a 5′ end composed of random nucleotides, the 3′ end of the modified RNA oligomer was ligated using T4 RNAligase 2 truncated KQ (NEB). Figure 3b By analyzing the sequences at the ligation sites, the sequences of the 3′ end of the most frequently occurring modified RNA oligomer and the 5′ end of the adapter were selected. Since these combinations exhibit higher reactivity with T4 RNAligase 2 truncated KQ, they were determined to maximize ligation efficiency and yield in the library construction step using this enzyme.

[0176] The number of reads for the first 10 combinations was measured, and the enrichment (the increase relative to random combinations) of each combination was analyzed to confirm statistical significance. For 2-nucleotide combinations, all 10 combinations showed significant differences (p-value, q-value ≒0), and in replicates 1 and 2, the increase was confirmed to be more than 2.2-fold compared to random combinations. For 3-nucleotide combinations, all 10 combinations showed significant differences (p-value, q-value ≒0), and in replicates 1 and 2, the increase was confirmed to be more than 4.47-fold compared to random combinations. For 4-nucleotide combinations, all 10 combinations showed significant differences (p-value, q-value ≒0), and in replicates 1 and 2, the increase was confirmed to be more than 12-fold compared to random combinations. Figure 3b ).

[0177] Using the above method, based on the type of enzyme used for ligation, 10 end combinations that maximize ligation efficiency were identified, and for subsequent processes, the 2 to 4 specific nucleotide combinations were inserted into the ends of the modified RNA oligomers.

[0178] (2) Linkage of RNA oligomers

[0179] The modified RNA oligomers and general RNA oligomers constructed above were ligated using T4 RNA ligase 1 (NEB), T4 RNA ligase 2 (NEB), or T4 RNA ligase 2 truncated KQ (NEB) according to the manufacturer's instructions to form different lengths.

[0180] (3) Screening of RNA oligomer ligates

[0181] The modified RNA oligomers and regular RNA oligomers were loaded into an 8% denaturing urea gel, and RNA was separated by length using PAGE (polyacrylamide gel electrophoresis). Bands corresponding to the ligation product range of at least 2 to 20 RNA oligomers were selected, excised, collected, and eluted to recover RNA from the gel.

[0182] (4) RNA tailing, including poly(A) tailing

[0183] The RNA oligomer ligation products with specific lengths obtained from the screening were subjected to poly(A) tailing using a poly(A) tailing kit (Invitrogen) according to the manufacturer's instructions. Subsequently, the poly(A) tailed RNA oligomer ligation products were recovered using Dynabeads Oligo(dT) (Invitrogen). This process selectively recovers only linear RNA with attached poly(A) tails.

[0184] Using the methods described above, an RNA modification library containing multiple modified RNA ligation products of all surrounding sequences was constructed.

[0185] Example 2: Optimization of RNA Modification Library Construction Process

[0186] 2-1. Selection of Ligase Types

[0187] Depending on the type of RNA ligase, the ligation mechanism and the required substrate form vary. Therefore, multiple RNA ligases are used for ligation, and the results are used to select the most suitable RNA ligase. Three enzymes were tested: T4 RNAligase 1 (NEB), T4 RNAligase 2 (NEB), and T4 RNAligase 2 truncated KQ (NEB). Figure 4a ).

[0188] T4 RNAligase 1 is a single-stranded RNA ligase that catalyzes the reaction of the 5′ phosphate group and 3′ hydroxyl group of RNA. Therefore, the modified RNA oligomer is used for the reaction without pretreatment. Conversely, T4 RNAligase 2 truncated KQ catalyzes the reaction of the pre-adenylated 5′ and 3′ ends of RNA. Therefore, the modified RNA oligomer is first 5′ adenylated using Mth RNAligase (NEB) before the reaction. Finally, T4 RNAligase 2 is a double-stranded RNA ligase that catalyzes the reaction of the 5′ phosphate group and 3′ hydroxyl group in double-stranded RNA or RNA-DNA hybrids. Therefore, after synthesizing single-stranded DNA complementary to the terminal sequence of the modified RNA oligomer, this DNA is added to the ligation reaction. All reactions were performed according to the manufacturer's instructions.

[0189] In the reaction used to compare T4 RNAligase 1 with T4 RNAligase 2 truncated KQ, the following modified RNA oligomers were used: [NNNN(m6A)NNNNNNNN(m6A)NNNNNNNN(m6A)NNNNNNNN(m6A)NNNNNN] (where N represents a random RNA sequence.) Results from confirming ligation using this oligomer showed that when using T4 RNA ligase 1, PAGE analysis confirmed the generation of more long-chain RNA oligomer ligation products. Conversely, experiments using T4 RNA ligase 2 truncated KQ confirmed the generation of more short-chain ligation products. Figure 4b ).

[0190] Furthermore, in the reaction used to compare T4 RNA ligase 1 and T4 RNA ligase 2, the following modified RNA oligomers were used: [CGACAGAUGAGUUCCNNNNNNNNNN(Um)NNNNNNNNNNCCUUGAUAGACAGUC] In addition, to perform the RNA ligase 2 reaction, the following single-stranded DNA, complementary to the terminal sequence of the modified RNA oligomer, was synthesized and experimentally tested: [TCATCTGTCGGACTGTCTAT] Results confirming ligation using these oligomers showed that using T4 RNA ligase 1 resulted in a greater production of long-chain RNA oligomer ligation products, as confirmed by PAGE. Compared to RNA ligase 2 experiments, bands were more prominent in longer regions, and less residual input (51-mer) was observed. Figure 4b ).

[0191] 2-2. Connection Optimization

[0192] Through the above experiments, T4 RNA ligase 1 (NEB) was selected as the RNA ligase. However, it was confirmed that the ligation efficiency and yield remained low when ligating according to the manufacturer's instructions. Therefore, the aim was to establish novel ligation conditions suitable for RNA-modified library design.

[0193] Therefore, by changing the reaction temperature, reaction time, PEG (polyethylene glycol) 8000 concentration, and DMSO (dimethyl sulfoxide) concentration, the amount of poly(A)+RNA relative to the amount of input RNA oligomers was calculated, and the conditions with the highest yield were screened.

[0194] (Condition 1) Reaction temperature

[0195] For reactants of the same composition, the linkage reaction was carried out at 16℃, 25℃, 29℃, 33℃, 37℃, and 41℃. Yield calculations showed that the highest yield was observed at 33℃. Figure 5a ).

[0196] (Condition 2) PEG8000 concentration

[0197] Experiments were conducted in the reaction system with PEG8000 concentrations of 10%, 15%, 19%, 23%, and 30%. Yield calculations showed that the highest yields were obtained at 15% or 19%. Figure 5b ).

[0198] (Condition 3) DMSO concentration

[0199] Experiments were conducted in the reaction system with DMSO concentrations of 0%, 10%, and 20%. The yield calculations showed that the highest yield was obtained at 10% DMSO. Figure 5c ).

[0200] (Condition 4) Reaction time

[0201] To determine the optimal reaction time at an optimal temperature of 33°C, experiments were conducted at 4 hours, 16 hours, and 40 hours. Yield calculations showed that the highest yield was obtained at a reaction time of 16 hours. Figure 5d ).

[0202] (Condition 5) Final Condition Comparison

[0203] Experiments were conducted using the conditions provided by NEB and the optimized conditions of this invention. The yield calculations showed that the yield increased from 3% to 11.5%, an increase of 3.8 times. Figure 5e ).

[0204] 2-3. Optimization of RNA tailing including poly(A) tailing

[0205] To stabilize the RNA oligomer ligation products and construct a sequenceable library, RNA tailing (adding a non-template nucleotide to the 3′ end of RNA) is required, which involves poly(A) tailing. Since the modified RNA oligomer ligation products constructed in this invention differ from mRNA and contain more shorter RNA molecules, it is considered necessary to first find an appropriate poly(A) tail length suitable for this library. Therefore, the process begins with 1) finding the optimal poly(A) tail length, followed by 2) optimizing the experimental conditions to achieve this length.

[0206] In the process of finding an appropriate poly(A) tail length, poly(A) tails of 20, 30, 40, 100, and 300 nucleotides were generated at the 3′ ends of experimental RNA molecules of the same length but with different indices. Subsequently, samples in which all RNA molecules were mixed in equal moles were sequenced. From the sequencing results, the poly(A) tail length corresponding to the RNA with the highest base quality and mapping accuracy was selected as the optimal poly(A) tail length.

[0207] The results showed that the base quality was higher under the poly(A) tail conditions of 30 and 100 nucleotides in length, while the proportion of mapped reads was highest under the poly(A) tail condition of 30 nucleotides in length. Figure 6 The condition that yielded the highest value in both results was a poly(A) tail length of 30 nucleotides, therefore this length was set as the most suitable length for the library of this invention.

[0208] Subsequently, an experiment was conducted to attach a poly(A) tail of approximately 30 nucleotides in length using E-PAP (E. coli poly(A) polymerase). When performed according to the manufacturer's instructions for the Invitrogen poly(A) tailing kit, it was confirmed that a longer poly(A) tail was attached. Figure 7a Therefore, by varying the ATP concentration, E-PAP enzyme dosage, and reaction time, conditions for attaching a 30-nucleotide-long poly(A) tail were established. The results were confirmed by PAGE on a denaturing gel.

[0209] The results showed that when using 1 mM ATP, no significant difference was observed even with changes in the E-PAP enzyme dosage. However, when ATP was reduced to 1 mM, 100 μM, and 10 μM, the length of the poly(A) tail decreased sharply. Under 10 μM ATP conditions, using 3.2 U of E-PAP and shortening the reaction time to 45 minutes yielded a poly(A) tail closest to approximately 30 nucleotides in length. Figure 7b Based on the above results, 10 uM ATP, 3.2 U E-PAP, and reaction at 37°C for 45 minutes were determined to be the optimal conditions for constructing the poly(A) tailing of the RNA modification library of this invention.

[0210] This process can also be performed using poly(U) polymerase (NEB) for poly(U) tailing or poly(I) tailing. This can be achieved by replacing adenosine triphosphate (ATP) with the same concentration of uridine triphosphate (UTP) or inosine triphosphate (ITP). Subsequently, the RNA-tailed oligomer ligation products can be recovered using oligodeoxyadenosine (Oligo(dA)) or oligodeoxycytidine (Oligo(dC)).

[0211] 2-4. Attachment of RNA tails via ligation

[0212] Poly(A) tailing using E-PAP enzyme is used to construct libraries of RNA in a form similar to mRNA, because adapters or RT primers used in library preparation typically contain oligo(dT) oligonucleotides. However, ligation may be preferred when attaching RNA tails of precise length, and it may be more suitable for testing various specific experimental conditions (e.g., subtle differences in RNA tail length). Furthermore, ligation is the only method available for attaching tails composed of heteronucleotides (rather than homonucleotides). In this case, adapters or RT primers with complementary sequences are directly designed for the sequencing preparation process.

[0213] RNA oligomers consisting of 20 to 40 adenosine molecules, with phosphate groups attached to both the 5′ and 3′ ends, were designed as RNA tails. During the ligation reaction, to prevent the tails from joining to form a circular form of RNA, or to attach multiple tails to the oligomer ligation product, a phosphate group was also attached to the 3′ end. Subsequently, pre-adenylated App-RNA oligomers were constructed using Mth RNA ligase (NEB), and the reaction was performed. To confirm the precise length of the tail attachment, an RNA tail was attached to the 3′ end of the oligomer with a length of 240 nucleotides via ligation.

[0214] In the Mth RNA ligase reaction, following the manufacturer's instructions, 100 pmol of enzyme and 100 uM adenosine triphosphate (ATP) were added to every 100 pmol of RNA oligomers, and the reaction was carried out at 65°C for 1 hour, followed by a reaction at 85°C for 5 minutes. Subsequently, in the reaction of ligating the oligomer ligation product with App-RNA oligomers, following the manufacturer's instructions, 40 pmol of App-RNA oligomers, 10% PEG8000, and 200 U T4 RNA ligase 2 truncated KQ were added to every 20 pmol of RNA substrate, and the reaction was carried out at 25°C for 2 hours. Experimental results showed that the length of the oligomer ligation products of various lengths increased precisely by 20, 30, and 40 nucleotides, respectively, and the degree of length increase could be distinguished. Therefore, it is considered that the attachment of RNA tails of precise length achieved through ligation has been successfully completed. Figure 7c ).

[0215] Subsequently, in order to attach the tail of the heteronucleotide instead of the non-homonucleotide form, the RNA tail is attached by ligation.

[0216] [GGUACCCGGGCGAAUUCCAAGCUUGAUCGC]

[0217] RNA oligomers with the above sequence and phosphate groups attached to both the 5′ and 3′ ends were designed. Subsequently, pre-adenylated App-RNA oligomers were constructed using Mth RNA ligase (NEB), and RNA tailing was performed by ligation at the 3′ end of oligomers of different lengths, following the same procedure as described above. Experimental results showed that the ligation products of oligomers of different lengths all increased by 30 nucleotides, therefore, the attachment of heteronucleotide RNA tails via ligation was considered successful. Figure 7d ).

[0218] The RNA tails obtained through ligation were recovered via PAGE. Recovery was achieved by selectively cleaving the oligomeric ligation product bands with increased length and then eluting. The yield was approximately 10%, which is lower than the approximately 30% yield of the poly(A) tailing method in Examples 2-3. However, this experiment confirmed that heteronucleotide RNA tails can be attached. After tailing, the 3′ phosphate group was removed and converted to the 3′ hydroxyl group by PNK treatment, thus completing the form that can be used for subsequent sequencing preparation.

[0219] Example 3: Sequencing using an RNA-modified library

[0220] As a first example of the present invention, RNA blocks with various combinations of 12 random nucleotides (A, G, C, U) randomly linked on both sides of the modified RNA's m6A and its corresponding general base A were designed. RNA oligomers with various combinations of RNA blocks arranged in a continuous random arrangement twice were also designed. In this experiment, 5′-CGAC and 3′-AGUC were further introduced at the 5′ and 3′ ends of each RNA oligomer, respectively, thereby designing RNA oligomers with a total length of 58 nucleotides. These RNA oligomers were then constructed using a chemical synthesis method.

[0221] The general A oligomer and m6A oligomer were ligated using T4 RNAligase 1 (NEB), respectively. In this case, the ligation reaction was carried out under the following conditions, which were changed from the manufacturer's (NEB) instructions: the reaction was carried out at 33°C for 16 hours in a reaction solution containing 15% PEG8000 and 10% DMSO.

[0222] RNA oligomers with more than four RNA oligomers and a length of more than 232 nucleotides were selectively recovered from the gel by PAGE. Subsequently, poly(A) tailing was performed using a poly(A) tailing kit (Invitrogen), where the reaction conditions were changed from the manufacturer's (Invitrogen) instructions: the reaction was carried out at 37°C for 45 minutes in a reaction solution containing 10 μM adenosine triphosphate (ATP) and 3.2 U E-PAP.

[0223] Subsequently, linear RNA with a poly(A) tail of approximately 30 nucleotides was selectively obtained using Dynabeads oligo(dT) (Invitrogen).

[0224] Sequencing was performed using the constructed m6A library and a general A library. Nanopore sequencing was selected for single-molecule RNA sequencing. Sequencing was performed according to the manufacturer's (Oxford Nanopore) instructions. For the m6A library, 1.32 million reads were obtained from a single sequencing run, yielding a total of 0.59 Gb of sequencing data. For the general A library, 2.00 million reads were obtained, yielding a total of 0.98 Gb of sequencing data. Figure 8a ).

[0225] As a second example of the present invention, RNA blocks with multiple combinations of five random nucleotides (A, G, C, U) randomly linked to each other on both sides of the m5C base and its corresponding general base C in modified RNA were designed. RNA oligomers with multiple combinations of RNA blocks arranged in a continuous random arrangement five times were also designed. 5′-GG and 3′-GG were further introduced into the 5′ and 3′ ends of each RNA oligomer, respectively, and RNA oligomers with a total length of 59 nucleotides were constructed using a chemical synthesis method. Experiments were conducted using these RNA oligomers under the same conditions described above, thereby constructing m5C libraries and general C libraries.

[0226] Sequencing was performed using both types of libraries described above, employing nanopore sequencing according to the manufacturer's instructions. For the m5C library, 1.30 million reads were obtained from a single sequencing run, yielding a total of 0.59 Gb of sequencing data. For the general C library, 2.15 million reads were obtained, yielding a total of 0.93 Gb of sequencing data. Figure 8b ).

[0227] Example 4: Application of Sequencing Datasets as Training Data

[0228] The RNA modification library designed according to the present invention is characterized by containing all surrounding sequences of various modified RNAs and their wide combinations. Therefore, sequencing data (e.g., single-molecule RNA sequencing data) of this library is used as training data for an RNA modification prediction model.

[0229] Motifs (modified RNA or general bases and their surrounding sequences, RNA blocks) are extracted from sequencing reads obtained from various modified RNA and general base libraries. This process utilizes a motif extraction algorithm based on a Directed Acyclic Graph (DAG).

[0230] The algorithm consists of the following steps: (1) k-mer index and matching First, k-mers (sequences of length k, where k is the length of a specific nucleotide at the end of the oligomer) are indexed in the sequencing reads. At this stage, sequence mutations, deletions, and insertions are allowed, with penalties assigned to each case. The cumulative penalty for each k-mer is defined as the specific sequence penalty. Subsequently, the indexed k-mers are matched against each specific sequence.

[0231] (2) Construct a Directed Acyclic Graph (DAG)

[0232] Construct a Directed Acyclic Graph (DAG) where nodes represent specific sequences and edges represent valid transitions between these sequences. The validity of each transition is determined by checking if the actual transition interval is within a certain error range and matches the closest expected transition interval from the given sequence and motif length. The penalty calculated based on the difference between the actual and expected transition intervals is called the transition penalty.

[0233] (3) Assigning edge weights

[0234] Each edge of the Directed Acyclic Graph (DAG) constructed in (2) is calculated and weighted. The weights are calculated by summing the specific sequence penalty calculated in (1), the transition penalty calculated in (2), and the central base penalty. The central base penalty is the minimum number of deletions and insertions required to find a specified central base (one of A, C, G, or U) at the motif center.

[0235] (4) Find the longest path

[0236] The longest path is found in the weighted directed acyclic graph (DAG) constructed in (3). The path length is defined as the sum of the weights of all edges. The process of finding the longest path is implemented using topological sorting and dynamic programming. After finding the longest path, motifs are extracted from the sequencing reads along that path and stored.

[0237] Using the aforementioned algorithm, motif extraction was performed on the sequencing reads of the general A library and the m6A library in Example 3, and the presence of motifs was confirmed. Sufficient sequencing depth was confirmed for combinations of 5-mer motifs (256 possibilities), 7-mer motifs (1,024 possibilities), and 9-mer motifs (65,536 possibilities). The sequencing data from both libraries showed a sequencing depth of at least 10× for all motifs.

[0238] To use the above data, which includes a wide range of motif combinations, as training data for the RNA modification prediction model, the impact of the location of modified RNA within the motif on the results was further analyzed. Compared to normal RNA, it was confirmed that the base quality of modified RNA was reduced, especially around the modified RNA, and showed a gradually decreasing trend in the surrounding sequence. Figure 9a Furthermore, it was confirmed that the electrical signals in general RNA and modified RNA, as well as in their surrounding sequences, also exhibited different patterns. Figure 9b In particular, the pattern of the confirmed current signal exhibits different characteristics depending on the surrounding motifs, that is, it shows a specific pattern according to the type of motif. Figure 9c ).

[0239] This analysis confirms that the library and sequencing data constructed in this invention contain data with multiple motifs (surrounding sequences) and motif-specific patterns, and therefore can be used as training data.

[0240] Example 5: Improving RNA Oligomer Design by Inserting Anchor Sequences

[0241] To improve the accuracy of motifs (modified RNA or general bases and their surrounding sequences) obtained from sequencing data of various RNA modification and general base libraries, experiments were conducted to improve the design of modified RNA oligomers.

[0242] First, to evaluate the improvement in RNA oligomer design when inserting anchor sequences, we hypothesized different RNA oligomer scenarios with varying anchor sequence forms and conducted simulations to compare motif extraction capabilities. Figure 10a ).

[0243] As an example, sequencing data from previous samples were used to test the case of specific anchoring sequences with 4 to 6 nucleotides inserted between blocks.

[0244] The data used in this test were sequencing data of an IVT RNA with a length of 274 nucleotides, the sequence of which is as follows.

[0245]

[0246] The data described is sequencing data used to simulate the following scenario: modified RNA oligomer ligation products containing three modified RNA oligomers (i.e., modified RNA oligomers ligated three times), each containing a modified RNA block of 21 nucleotides repeated three times. Motif extraction capability was simulated using specific sequence portions from the data. In each condition, the specific sequence portions used for extraction are shown in bold below.

[0247] (Condition 1) There is no anchoring sequence between the modified RNA blocks (in order to maintain the same number of nucleotides as (Condition 2), a specific sequence of 4 nucleotides is appended to the 5′ and 3′ ends).

[0248] (Condition 2) The case where there is a 4-nucleotide anchoring sequence between modified RNA blocks.

[0249] (Condition 3) The case where there is a 6-nucleotide anchoring sequence between modified RNA blocks.

[0250] To calculate the accuracy of motif selection under the three conditions described above, the precision-recall curve (PR curve) was calculated using the following method. The sequencing data of the in vitro transcribed (IVT) RNA was aligned with the original IVT reference sequence, and a motif extraction algorithm was performed on the IVT sequencing data. The motif extraction results were compared with the alignment results to confirm the consistency between the motif extraction results at each anchorage location and the anchor coordinates in the alignment results. Subsequently, the extraction score of each anchorage obtained in the motif extraction algorithm was used as the predicted value, and the alignment consistency of the anchor coordinates was used as the true value to calculate the precision-recall curve.

[0251] The precision-recall curves confirmed that motif extraction performance improved sequentially in the order of Condition 3 > Condition 2 > Condition 1. For example, at the same recall of 0.2, Condition 1 had a precision below 0.65, while Condition 3 had a precision above 0.90. Based on these results, it was confirmed that inserting a 6-nucleotide anchor sequence yielded the greatest improvement in identifying specific bases and motifs at the block center in sequencing data.

[0252] Subsequently, the need to adjust the anchor sequence length based on the number of modified RNA blocks was assessed. When the number of modified RNA blocks decreased to less than two, the accuracy of motif extraction may decrease. Therefore, to confirm whether a higher motif accuracy could be maintained by further increasing the anchor sequence length, relevant analyses were performed. As mentioned earlier, by assuming RNA oligomers with different numbers of RNA blocks and anchor sequence lengths, and comparing motif extraction capabilities through simulation (…),… Figure 10b Similar to the previous experiments, we assumed the case of modified RNA blocks containing 21 nucleotides in length and confirmed the motif extraction performance under conditions of reduced RNA block number.

[0253] (Condition 1) The case where there are three modified RNA blocks and a 6-nucleotide anchoring sequence.

[0254] (Condition 2) The case where there are two modified RNA blocks and a 6-nucleotide anchoring sequence.

[0255] (Condition 3) The case where there are two modified RNA blocks and an 8-nucleotide anchoring sequence.

[0256] (Condition 4) The existence of a modified RNA block and a 12-nucleotide anchoring sequence.

[0257] (Condition 5) The existence of a modified RNA block and a 15-nucleotide anchoring sequence.

[0258] The results of precision-recall curves calculated for the five conditions showed that when the number of RNA blocks decreased, motif extraction performance significantly decreased when using anchor sequences of the same length (Condition 1 > Condition 2), while performance improved by increasing the anchor sequence length (Condition 2 > Condition 3, Condition 4 > Condition 5). Furthermore, it was confirmed that motif extraction performance could be maintained to some extent by increasing the anchor sequence length as the number of RNA blocks decreased (Condition 3 ≒ Condition 4). Finally, it was confirmed that even when the number of RNA blocks decreased to one-third, the accuracy of motif extraction could be maintained by increasing the anchor sequence length by approximately three times (Condition 1 ≒ Condition 5).

[0259] Therefore, the anchoring length is preferably 20 to 30% of the length of the modified RNA block. However, when the number of RNA blocks is less than two, it has been confirmed that the accuracy of motif extraction can be maintained by increasing the length of the anchoring sequence inversely to the number of RNA blocks.

[0260] Subsequently, the sequence composition of the anchor sequence is selected considering the following factors.

[0261] (1) The Lewinstein distance between each anchoring sequence should be at least 50% of the anchoring length.

[0262] (2) When all anchor sequences are combined, the proportion of each nucleotide should be equal, and the proportion of nucleotides within each anchor sequence should also be equal. The equal proportion refers to the proportion in which the lengths of each anchor sequence are divided according to four integer ratios closest to 1:1:1:1.

[0263] (3) The average alignment error of all 5-mers that make up each anchor sequence in nanopore sequencing should not exceed the 10th percentile.

[0264] (4) When constructing partially random oligomers using all anchored sequences and sampled random blocks, the average value of the minimum free energy (MFE) should be above the 90th percentile. This is to minimize the formation of RNA secondary structures.

[0265] (5) When constructing partially random oligomers as described in (4), the average value of the minimum free energy (dimer MFE) of the dimers produced in all possible pairings should be above the 90th percentile, thereby minimizing the formation of RNA dimers. This is to minimize the formation of RNA dimers.

[0266] (6) When extracting all subsequences with a length of 50% of the anchor length from all anchor sequences, there should be no reverse complements among all possible pairs formed by them. This is to minimize the formation of RNA dimers.

[0267] As an example of an anchored sequence of 6 nucleotides that meets the above conditions, the following set is obtained.

[0268] Using the obtained anchoring sequence, in the modified RNA oligomer design of Example 1, an anchoring sequence with a specific sequence was inserted between the modified RNA blocks. Figure 10c After adopting the design that inserts the anchor sequence, the results show that the motif extraction accuracy improved from 66.5% to 99.7%.

[0269] The foregoing has described specific aspects of this invention in detail. It will be apparent to those skilled in the art that the above description is merely a preferred embodiment and is not intended to limit the scope of the invention. Therefore, the essential scope of this invention should be defined by the appended claims and their equivalents.

[0270] [National research and development projects supporting this invention]

[0271] [Project Number] 1711186079

[0272] [Project Number] 2020R1A2C3007032

[0273] [Name of the competent authority] Ministry of Science, Technology and Information

[0274] [Name of the organization responsible for project management (professional)] Korea Research Foundation

[0275] [Research Project Title] Individual Basic Research (Ministry of Science, Technology and Information) - Mid-Level Researcher Project

[0276] [Research Topic Title] Discovering Novel Oncogenic Non-coding Mutations by Elucidating the MicroRNA Targeting Regulation Mechanism Mediated by RNA-Binding Proteins

[0277] [Name of the Institution Undertaking the Project] Seoul National University

[0278] [Research Period] June 1, 2020 ~ February 28, 2025

[0279] [National research and development projects supporting this invention]

[0280] [Project Number] 1711200977

[0281] [Project Number] 2022M3A9I2082294

[0282] [Name of the competent authority] Ministry of Science, Technology and Information

[0283] [Name of the organization responsible for project management (professional)] Korea Research Foundation

[0284] [Research Project Title] Biomedical Technology Development Project

[0285] [Research Topic Title] (Common 3) Establishment of a Predictive Model for the Occurrence of Novel Coronavirus (COVID-19) Variant Strains and Development of a Prevention Platform Technology Based on Comprehensive Mechanism Analysis

[0286] [Name of the Institution Undertaking the Project] Seoul National University

[0287] [Research Period] April 1, 2022 ~ December 31, 2026

[0288] [National research and development projects supporting this invention]

[0289] [Project Number] 1711200846

[0290] [Project Number] 2019M3E5D3073104

[0291] [Name of the competent authority] Ministry of Science, Technology and Information

[0292] [Name of the organization responsible for project management (professional)] Korea Research Foundation

[0293] [Research Project Title] Omics-Based Precision Medicine Technology Development Project - Biomedical Technology Development Project

[0294] [Research Project Title] Development of an Exosome Multi-omics Analysis Platform for Precision Medicine of Diabetic Complications

[0295] [Name of the Institution Undertaking the Project] Seoul National University

[0296] [Research Period] July 1, 2019 ~ December 31, 2024

[0297] [National research and development projects supporting this invention]

[0298] [Project Number] 1711188895

[0299] [Project Number] 2020R1A5A1018081

[0300] [Name of the competent authority] Ministry of Science, Technology and Information

[0301] [Name of the organization responsible for project management (professional)] Korea Research Foundation

[0302] [Research Project Title] Group Research Support

[0303] [Research Topic Title] Research Center for Systemic Aging Mechanisms

[0304] [Name of the Institution Undertaking the Project] Seoul National University

[0305] [Research Period] July 1, 2020 ~ February 28, 2027

Claims

1. A method for constructing an RNA modification library, comprising the following steps: (a) Construct one or more modified RNA blocks, and randomly link 4 to 50 random nucleotides in the 5′ and 3′ directions of the modified RNA, respectively; (b) Constructing modified RNA oligomers in which 1 to 9 modified RNA blocks are randomly arranged in one or more modified RNA blocks; (c) The modified RNA oligomers are ligated to construct the modified RNA oligomer ligation product; (d) Screening the modified RNA oligomer ligation products, wherein the modified RNA oligomers are ligated from 2 to 20; (e) Attaching an RNA tail consisting of 5 to 150 arbitrary ribonucleotides to the selected modified RNA oligomer ligation product; and (f) Recover the modified RNA oligomer ligation product with the RNA tail attached.

2. The method for constructing an RNA modification library according to claim 1, wherein, The modified RNA is any one of the modified RNAs selected from the group consisting of (1) to (4) below: (1) Modification A shown in Table 1: Table 1 (2) Modification C shown in Table 2: Table 2 (3) Modification G shown in Table 3: Table 3 as well as (4) Modifications U shown in Table 4: Table 4 。 3. The method for constructing an RNA modification library according to claim 1, wherein, Step (b) involves constructing a modified RNA oligomer, which further includes an anchoring sequence at the following location: (i) Between the modified RNA blocks that make up the modified RNA oligomer; (ii) the 5' end of the modified RNA block located at the 5' end in the modified RNA block constituting the modified RNA oligomer; and / or (iii) The 3′ end of the modified RNA block located at the 3′ end in the modified RNA block that constitutes the modified RNA oligomer.

4. The method for constructing an RNA modification library according to claim 3, wherein, The anchoring sequence is a sequence of 4 or more nucleotides. If the modified RNA oligomer contains 2 or fewer modified RNA blocks, the anchoring sequence has a length of 40 to 90% of the length of the modified RNA blocks. If the modified RNA oligomer contains 3 to 9 modified RNA blocks, the anchoring sequence has a length of 20 to 30% of the length of the modified RNA blocks.

5. The method for constructing an RNA modification library according to claim 1, wherein, Step (c) involves ligating RNA using RNA ligase.

6. The method for constructing an RNA modification library according to claim 1, wherein, Step (b) involves constructing the modified RNA oligomer by introducing 2 to 4 additional specific nucleotide combinations at the 5′ and 3′ ends of the modified RNA oligomer, respectively.

7. The method for constructing an RNA modification library according to claim 6, wherein, The RNA ligase is T4 RNA ligase 1, the two specific nucleotide combinations are any of the combinations listed in Table 5 below, the three specific nucleotide combinations are any of the combinations listed in Table 6 below, and the four specific nucleotide combinations are any of the combinations listed in Table 7 below. Table 5 Table 6 Table 7 Alternatively, the RNA ligase is T4 RNA ligase 2 truncated KQ, the two specific nucleotide combinations are any of the combinations listed in Table 8 below, the three specific nucleotide combinations are any of the combinations listed in Table 9 below, and the four specific nucleotide combinations are any of the combinations listed in Table 10 below. Table 8 Table 9 Table 10 。 8. The method for constructing an RNA modification library according to claim 1, wherein, Steps (a) and (b) are constructed by chemical synthesis or template-independent enzymatic synthesis. Step (b) is constructed as an independent process after step (a), or steps (a) and (b) are constructed as a single process.

9. The method for constructing an RNA modification library according to claim 1, wherein, Step (c) involves connecting the components at 29°C to 37°C for 8 to 40 hours.

10. The method for constructing an RNA modification library according to claim 1, wherein, Step (c) involves bonding in a reaction solution containing 13 to 21% (v / v) PEG8000.

11. The method for constructing an RNA modification library according to claim 1, characterized in that, The RNA tail in step (e) is a poly(A) tail, a poly(U) tail, or a poly(I) tail, consisting of 5 to 40 ribonucleotides.

12. The method for constructing an RNA modification library according to claim 11, characterized in that, In step (e), the poly(A) tail, poly(U) tail, or poly(I) tail is reacted in a reaction solution containing 5 to 50 μM of adenosine triphosphate (ATP), uridine triphosphate (UTP), or inosine triphosphate (ITP).

13. An RNA modification library comprising multiple modified RNA oligomer ligation products, wherein, The modified RNA oligomer ligation product is in the form of ligating 2 to 20 modified RNA oligomers, characterized in that... The modified RNA oligomer has modified RNA blocks with 4 to 50 random nucleotides randomly linked in the 5′ and 3′ directions, respectively, arranged in 1 to 9 random consecutive sequences, with the modified RNA at the center.

14. The RNA modification library according to claim 13, wherein, The modified RNA is any one of the modified RNAs selected from the group consisting of (1) to (4) below: (1) Modification A shown in Table 1: Table 1 (2) Modification C shown in Table 2: Table 2 (3) Modification G shown in Table 3: Table 3 as well as (4) Modifications U shown in Table 4: Table 4 。 15. The RNA modification library according to claim 13, characterized in that, The modified RNA oligomer further includes an anchoring sequence at the following location: (i) Between the modified RNA blocks that make up the modified RNA oligomer; (ii) the 5′ end of the modified RNA block located at the 5′ end in the modified RNA block constituting the modified RNA oligomer; and / or (iii) The 3′ end of the modified RNA block located at the 3′ end in the modified RNA block that constitutes the modified RNA oligomer.

16. The RNA modification library according to claim 15, characterized in that, The anchoring sequence is a sequence of 4 or more nucleotides. If the modified RNA oligomer contains 2 or fewer modified RNA blocks, the anchoring sequence has a length of 40 to 90% of the length of the modified RNA blocks. If the modified RNA oligomer contains 3 to 9 modified RNA blocks, the anchoring sequence has a length of 20 to 30% of the length of the modified RNA blocks.

17. The RNA modification library according to claim 13, wherein, The modified RNA oligomer has an additional 2 to 4 specific nucleotide combinations introduced at its 5′ and 3′ ends, respectively.

18. The RNA modification library according to claim 13, wherein, The 3′ end of the modified RNA oligomer ligation product is further attached with an RNA tail consisting of 5 to 150 arbitrary ribonucleotides.

19. An RNA library collection for training an RNA modification detection model, comprising: RNA modification library according to any one of claims 13 to 18; and A general RNA library comprising multiple general RNA oligomer ligation products, wherein the multiple general RNA oligomer ligation products have the same structure as the multiple modified RNA oligomer ligation products contained in the RNA modification library, except that the modified RNA is replaced by the corresponding general RNA.

20. A sequencing dataset, wherein, The sequencing dataset is generated using the RNA library collection used for training the RNA modification detection model as described in claim 19.

21. The sequencing dataset according to claim 20, wherein, The sequencing was single-molecule RNA sequencing.

22. The sequencing dataset according to claim 20, wherein, The sequencing dataset is used for deep learning and / or machine learning training.