De novo sequencing method of RNA ribose and / or base modification
By using ribonuclease digestion and mass spectrometry analysis, the problem of identifying unknown ribose and base modification sites in RNA sequences in existing technologies has been solved, enabling precise analysis of nucleic acid drugs and improving the accuracy of nucleic acid drug design.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING GENSCRIPT BIOTECH CO LTD
- Filing Date
- 2025-10-30
- Publication Date
- 2026-05-01
AI Technical Summary
Existing sequencing technologies are insufficient to directly identify unknown ribose and/or base modification sites in RNA sequences, especially methoxy and fluorinated modifications, thus failing to meet the analytical requirements for nucleic acid drugs.
The method of ribonuclease digestion combined with mass spectrometry analysis was adopted. By calculating the difference between the measured molecular weight and the theoretical molecular weight of the RNA sequence to be tested, the types and quantities of ribose and/or base modifications were determined. The modification pattern was determined by matching analysis, and the modification information was further confirmed by secondary mass spectrometry detection.
It enables accurate identification of ribose and/or base modifications in RNA sequences, meeting the analytical needs of nucleic acid drugs and improving the accuracy of nucleic acid drug design and research.
Smart Images

Figure CN121963867A_ABST
Abstract
Description
[0001] Cross-citation of related applications This application claims priority to Chinese patent application CN202411546457.2, filed on October 31, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This invention relates to nucleic acid detection methods, specifically to de novo sequencing methods for RNA ribose and / or base modifications. Background Technology
[0003] In recent years, small nucleic acid drugs have become a research hotspot in the biopharmaceutical field due to their advantages such as high specificity, simple design, short development cycle, and abundant targets. Small nucleic acid drugs specifically refer to a class of oligonucleotide molecules that target RNA or proteins, including antisense oligonucleotides (ASO), siRNA, and aptamers, typically composed of single-stranded or double-stranded molecules of fewer than 30 nucleotides. With the continuous development of various nucleic acid drugs, in addition to specific sequence design, modification design is also crucial. Methoxy and fluorination modifications, as common and widely used modification methods, are extensively used in nucleic acid drug design. Whether it's short-chain small nucleic acid drugs or the recently popular sgRNA technology, methoxy and fluorination modifications are widely employed. It can be said that methoxy and fluorination modifications as ribose modifications are recognized and widely used, and these modifications play an irreplaceable role in the stability and unique targeting of nucleic acid drugs.
[0004] Most nucleic acid drugs are RNA-based, and with the rise of nucleic acid drugs, the demand for RNA analysis, especially structural analysis, is increasing. Existing next-generation sequencing (NGS) can identify unmodified sequences, such as verifying known bare sequences or identifying unknown bare sequences, but it struggles to identify modification sites on sequences. The recently emerging LC_MSMS sequence identification can identify modified sequences, but this method requires a reference theoretical sequence and modification site information before identification. Software uses this information to obtain theoretical daughter ion information, which is then matched with the measured daughter ion information. A match proves that the detected sequence is consistent with the reference theoretical sequence, thus achieving identification. In other words, LC_MSMS sequence identification requires prior information on the sequence and modifications, which can be seen as a form of verification. The verification function of LC_MSMS can falsify a limited number of true and false sequences. For example, given several sequences, if one sequence has a correct modification site, this technology can filter it out from a false positive database, thus achieving identification. However, without new analytical methods, a few unknown modification sites can lead to an astronomical number of possible sequences. Neither NGS sequencing nor LC-MSMS can directly identify unknown ribose and / or base modification sites. Therefore, methods for identifying ribose and / or base modification sites need to be developed to meet the increasing analytical demands. Methoxy and fluoropolymers, as the most common types of modifications on RNA ribose, are the first areas that need to be addressed. Summary of the Invention
[0005] A first aspect of the present invention provides a method for identifying ribose and / or base modification patterns in a test RNA sequence, wherein the naked sequence of the test RNA sequence is known. When the type and / or quantity of ribose and / or base modifications in the RNA sequence to be tested are unknown, the method includes the following steps: Step 1: Determine the type and / or quantity of ribose and / or base modifications in the RNA sequence to be tested based on the difference between the measured molecular weight of the RNA sequence to be tested and the theoretical molecular weight of the naked sequence of the RNA sequence to be tested. Step 2: Obtain the first measured molecular weight of the RNA sequence to be tested after digestion with a first ribonuclease. Match the first measured molecular weight with the first theoretical molecular weight of the first theoretical digestion fragment obtained by theoretically digesting the naked sequence of the RNA sequence to be tested with a first ribonuclease. This determines the nucleic acid sequence and partial ribose and / or base modification information of each digestion fragment obtained by specific digestion with a first ribonuclease of the RNA sequence to be tested. Optionally, obtain the second measured molecular weight of the RNA sequence to be tested after double digestion with a first ribonuclease and a second ribonuclease. Match the second measured molecular weight with the second theoretical molecular weight. The second theoretical digestion fragment is obtained by double digestion with a first ribonuclease and a second ribonuclease on a naked sequence of the RNA sequence to be tested with a known modification. This determines the nucleic acid sequence and the ribose and / or base modification information of each digestion fragment obtained by specific double digestion with a first ribonuclease and a second ribonuclease of the RNA sequence to be tested. Step 3: Based on the information obtained in Steps 1 and 2, determine all possible ribose and / or base modification patterns in the RNA sequence to be tested; and Step 4: When there is only one possible combination of ribose and / or base modification patterns, the modification patterns of all ribose and / or bases in the RNA sequence to be tested are determined; when there is more than one possible ribose and / or base modification pattern, the RNA sequence to be tested is further subjected to secondary mass spectrometry detection to determine the modification patterns of all ribose and / or bases in the RNA sequence to be tested. When the type and / or quantity of ribose and / or base modifications in the RNA sequence to be tested are known, step one is not required in this method.
[0006] In some embodiments, the type, quantity, and / or location of ribose and / or base modifications in the RNA sequence to be tested are unknown, and the method includes the following steps: Step 1: Obtain the measured molecular weight of the RNA sequence to be tested. Based on the difference between the measured molecular weight and the theoretical molecular weight of the bare sequence of the RNA sequence to be tested, determine the type and quantity of ribose and / or base modifications in the RNA sequence to be tested, which is constraint condition one. Step 2: Obtain the molecular weight of the first actual digested fragment obtained by digestion of the RNA sequence to be tested with the first ribonuclease, and obtain the first measured molecular weight, wherein each first measured molecular weight represents a first actual digested fragment; using the naked sequence of the RNA sequence to be tested as the first theoretical sequence, digest it with the theoretical first ribonuclease to obtain the first theoretical digested fragment; based on the constraint condition 1, perform the first matching analysis on the first theoretical digested fragment and the digested fragment represented by the first measured molecular weight, thereby determining the nucleic acid sequence and partial ribose and / or base modification information of each digested fragment obtained by specific digestion of the RNA sequence to be tested with the first ribonuclease, which is constraint condition 2; The first matching analysis includes comparing the theoretical molecular weight of the first theoretical digestion fragment with the first measured molecular weight to determine the first theoretical digestion fragment that matches each digestion fragment represented by the first measured molecular weight, as well as the additional ribose and / or base modifications contained in each digestion fragment represented by the first measured molecular weight relative to the matching first theoretical digestion fragment. The partial ribose and / or base modification information determined by the first matching analysis includes the type and quantity of ribose and / or base modifications contained in each digestion fragment obtained by specific digestion of the test RNA sequence with a first ribonuclease, and also includes precise information on whether a specific ribonucleotide has ribose and / or base modifications. The specific ribonucleotide refers to the characteristic nucleotide of each first ribonuclease contained in the test RNA sequence, excluding the nucleotide at the 3' end of the test RNA sequence; and / or The molecular weight of the second actual digested fragment obtained by double digestion of the RNA sequence to be tested with the first and second ribonucleases is obtained, and each second measured molecular weight represents a second actual digested fragment. The sequence of the modified fragment formed by adding ribose and / or base modifications at specific positions determined in the first matching analysis to the naked sequence is the second theoretical sequence. Double digestion with the theoretical first and second ribonucleases is performed to obtain the second theoretical digested fragment. Based on the satisfaction of constraint one and constraint two, a second matching analysis is performed on the digested fragments represented by the second theoretical digested fragment and the second measured molecular weight to determine the nucleic acid sequence and the types and numbers of ribose and / or base modifications contained in each digested fragment obtained by specific double digestion of the RNA sequence to be tested with the first and second ribonucleases, which is constraint three. The second matching analysis includes comparing the theoretical molecular weight of the second theoretical digestion fragment with the second measured molecular weight to determine the second theoretical digestion fragment that matches each second measured molecular weight and the additional ribose and / or base modifications contained in each second measured molecular weight fragment relative to the matching second theoretical digestion fragment. The partial ribose and / or base modification information determined by the second matching analysis includes the types and quantities of ribose and / or base modifications contained in each digestion fragment obtained by specific double digestion of the RNA sequence to be tested by the first and second ribonucleases, and also includes the exact information on whether a specific ribonucleotide has ribose and / or base modifications. The specific ribonucleotide refers to the characteristic nucleotide of each first or second ribonuclease contained in the RNA sequence to be tested, excluding the nucleotide at the 3' end of the RNA sequence to be tested. Step 3: Determine all possible ribose and / or base modification patterns in the RNA sequence to be tested; and Step 4: When there is only one type of ribose and / or base modification pattern, the types, positions, and numbers of all ribose and / or base modification patterns in the RNA sequence to be tested are determined; when there is more than one type of ribose and / or base modification pattern, the RNA sequence to be tested is further subjected to secondary mass spectrometry detection to determine the types, positions, and numbers of all ribose and / or base modification patterns in the RNA sequence to be tested. The first and second ribonucleases are site-specific ribonucleases and are not the same.
[0007] In some implementations, when the molecular weight of the RNA to be tested is known, specifically when the type and number of ribose and / or base modifications in the sequence are known, the method does not include step one.
[0008] In some embodiments, one of the first and second ribonucleases is RNase T1, whose characteristic nucleotide is G, and the other is RNase A, whose characteristic nucleotides are C and U; preferably, the first ribonuclease is RNase T1 and the second ribonuclease is RNase A.
[0009] In some embodiments, the ribose and / or base modification is selected from one or more of methoxy modification, fluorination modification, methoxyethyl modification, and locked nucleic acid modification.
[0010] In some implementations, in step one, the type and amount of ribose and / or base modifications in the RNA sequence to be tested are determined according to the following method: First, calculate the total molecular weight of all ribose and / or base modifications in the RNA sequence to be tested using the following formula: The total molecular weight of ribose and / or base modifications = molecular weight 实测 -Molecular weight 裸序列理论 Among them, "molecular weight" 实测 "Molecular weight" refers to the measured molecular weight of the obtained RNA sequence to be tested. 裸序列理论 "" refers to the theoretical molecular weight of the naked sequence of the RNA to be tested.
[0011] Subsequently, the theoretical molecular weight of different combinations of ribose and / or base modifications is compared with the total molecular weight of ribose and / or base modifications in the RNA sequence to be tested, as calculated above. When the theoretical molecular weight of a specific combination of ribose and / or base modifications is equal to the total molecular weight of ribose and / or base modifications in the RNA sequence to be tested, the types and quantities of ribose and / or base modifications contained in the RNA sequence to be tested can be determined as the types and quantities of ribose and / or base modifications contained in that specific combination of ribose and / or base modifications.
[0012] In some implementations, in step two, the matching of the theoretical digestion fragment with the digestion fragment represented by the measured molecular weight includes one or more of the following: (a) The molecular weight of a single theoretical digestion fragment is equal to the measured molecular weight; in this case, the digestion fragment represented by the measured molecular weight has the same sequence as the single theoretical digestion fragment and does not contain any additional ribose and / or base modifications relative to the single theoretical digestion fragment. (b) A single theoretical restriction fragment has a molecular weight smaller than a measured molecular weight, but contains a specific combination of ribose and / or base modifications such that the molecular weight of the modified fragment formed by adding the specific ribose and / or base modification combination to the single theoretical restriction fragment is equal to the measured molecular weight, provided that the addition of the specific ribose and / or base modification combination satisfies the established ribose and / or base modification constraints; in this case, the restriction fragment represented by the measured molecular weight is the modified fragment formed by adding the specific ribose and / or base modification combination to the theoretical restriction fragment, and when the restriction fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the restriction fragment represented by the measured molecular weight does not have ribose and / or base modifications; and (c) An extended theoretical restriction fragment, formed by covalently linking two or more consecutive theoretical restriction fragments, has a molecular weight smaller than that of a restriction fragment represented by a measured molecular weight. However, it contains a specific combination of ribose and / or base modifications such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modifications to the extended theoretical restriction fragment is equal to that of the measured molecular weight, provided that the addition of the specific combination of ribose and / or base modifications satisfies pre-defined ribose and / or base modification constraints. In this case, the restriction fragment represented by the measured molecular weight is the modified fragment formed by adding the specific combination of ribose and / or base modifications to the extended theoretical restriction fragment. Specifically, the ribonucleotide at the 3' end of the preceding theoretical restriction fragment in any two adjacent theoretical restriction fragments has a ribose and / or base modification, and when the restriction fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the restriction fragment represented by the measured molecular weight does not have a ribose and / or base modification.
[0013] In some implementations, the matching analysis in step two includes performing a matching analysis on each theoretical digestion fragment, such that each theoretical digestion fragment is matched to a digestion fragment represented by a measured molecular weight.
[0014] In some implementations, the matching analysis in step two is performed by sequentially executing the following steps: S1: Select a theoretical enzyme digestion fragment and proceed with step S2; S2: Check whether the molecular weight of the theoretical enzyme digestion fragment is equal to a certain measured molecular weight, and make the judgment in step S3. S3: If the molecular weight of the theoretical enzyme digestion fragment is equal to a measured molecular weight, then the enzyme digestion fragment represented by the measured molecular weight is determined to have the same sequence as the theoretical enzyme digestion fragment and does not contain any additional ribose and / or base modifications relative to the theoretical enzyme digestion fragment. Then, check whether there are other theoretical enzyme digestion fragments that have not yet been matched. If so, select one of the unmatched theoretical enzyme digestion fragments and perform step S2 on it. If not, complete all matching. If the molecular weight of the theoretical enzyme digestion fragment is not equal to any of the measured molecular weights, then perform step S4 on the theoretical enzyme digestion fragment. S4: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the theoretical enzyme digestion fragment is equal to a certain measured molecular weight (i.e., the sum of the molecular weight of the theoretical enzyme digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to the measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S5. S5: If such a combination of ribose and / or base modifications exists, the enzyme fragment represented by the measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the theoretical enzyme fragment. When the enzyme fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the enzyme fragment represented by the measured molecular weight does not have ribose and / or base modifications. Then, it is checked whether there are other theoretical enzyme fragments that have not been matched. If so, one unmatched theoretical enzyme fragment is selected and subjected to step S2. If not, all matching is completed. If such a combination of ribose and / or base modifications does not exist, a theoretical enzyme fragment adjacent to it before and / or after it is added to the theoretical enzyme fragment to form an extended theoretical enzyme fragment and subjected to step S6. S6: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the extended theoretical digestion fragment is equal to a certain measured molecular weight (i.e., the sum of the molecular weight of the extended theoretical digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to a certain measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S7. S7: If such a combination of ribose and / or base modifications exists, the enzyme digestion fragment represented by the measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the extended theoretical enzyme digestion fragment. In any two adjacent theoretical enzyme digestion fragments, the ribonucleotide at the 3' end of the preceding theoretical enzyme digestion fragment has ribose and / or base modifications. When the enzyme digestion fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the enzyme digestion fragment represented by the measured molecular weight does not have ribose and / or base modifications. Then, it is checked whether there are other theoretical enzyme digestion fragments that have not yet been matched. If so, one unmatched theoretical enzyme digestion fragment is selected and subjected to step S2. If not, all matching is completed. If such a combination of ribose and / or base modifications does not exist, the extended theoretical enzyme digestion fragment is added to a theoretical enzyme digestion fragment adjacent to it before and / or after it to form a further extended theoretical enzyme digestion fragment, which is then subjected to step S6. In some implementations, the matching analysis in step two includes arranging all theoretical restriction fragments sequentially in a 5' to 3' direction or in a 3' to 5' direction, and performing matching analysis on each theoretical restriction fragment sequentially, starting with the first theoretical restriction fragment in the sequence.
[0015] In some implementations, the matching analysis in step two includes matching each theoretical digestion fragment having two or more nucleotides, such that each theoretical digestion fragment having two or more nucleotides matches a digestion fragment represented by a measured molecular weight.
[0016] In some implementations, the matching analysis in step two is performed by sequentially executing the following steps: S1: Select a theoretically digestible fragment containing two or more nucleotides and proceed to step S2; S2: Check whether the molecular weight of the theoretical enzyme digestion fragment is equal to a certain measured molecular weight, and make the judgment in step S3. S3: If the molecular weight of the theoretical digestion fragment is equal to a measured molecular weight, then the digestion fragment represented by the measured molecular weight is determined to have the same sequence as the theoretical digestion fragment and does not contain any additional ribose and / or base modifications relative to the theoretical digestion fragment. Then, it is checked whether there are other theoretical digestion fragments containing two or more nucleotides that have not yet been matched. If so, one of the unmatched theoretical digestion fragments containing two or more nucleotides is selected and subjected to step S2. If not, all matching is completed. If the molecular weight of the theoretical digestion fragment is not equal to any of the measured molecular weights, then step S4 is performed on the theoretical digestion fragment. S4: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the theoretical enzyme digestion fragment is equal to a certain measured molecular weight (i.e., the sum of the molecular weight of the theoretical enzyme digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to the measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S5. S5: If such a combination of ribose and / or base modifications exists, the enzyme fragment represented by the measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the theoretical enzyme fragment. When the enzyme fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the enzyme fragment represented by the measured molecular weight does not have ribose and / or base modifications. Then, it is checked whether there are other theoretical enzyme fragments containing two or more nucleotides that have not yet been matched. If so, one unmatched theoretical enzyme fragment containing two or more nucleotides is selected and subjected to step S2. If not, all matching is completed. If such a combination of ribose and / or base modifications does not exist, a theoretical enzyme fragment adjacent to it before and / or after it is added to the theoretical enzyme fragment to form an extended theoretical enzyme fragment and subjected to step S6. S6: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the extended theoretical digestion fragment is equal to a certain measured molecular weight (i.e., the sum of the molecular weight of the extended theoretical digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to a certain measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S7. S7: If such a combination of ribose and / or base modifications exists, then the enzyme fragment represented by the measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the extended theoretical enzyme fragment, wherein the ribonucleotide at the 3' end of the preceding theoretical enzyme fragment in any two adjacent individual theoretical enzyme fragments has ribose and / or base modifications, and when the enzyme fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the 3' end of the enzyme fragment represented by the measured molecular weight... The ribonucleotide does not have ribose and / or base modifications. Then, it is checked whether there are other theoretical cleavage fragments containing two or more nucleotides that have not been matched. If they exist, one of the unmatched theoretical cleavage fragments containing two or more nucleotides is selected and subjected to step S2. If they do not exist, all matching is completed. If there is no such combination of ribose and / or base modifications, the extended theoretical cleavage fragment is added to a theoretical cleavage fragment adjacent to it before and / or after it to form a further extended theoretical cleavage fragment and subjected to step S6.
[0017] In some implementations, the matching analysis in step two includes arranging all theoretical restriction fragments with two or more nucleotides sequentially in a 5' to 3' direction or in a 3' to 5' direction, and performing matching analysis on each theoretical restriction fragment with two or more nucleotides sequentially, starting with the first theoretical restriction fragment with two or more nucleotides in sequence.
[0018] In some implementations, the first and second measured molecular weights are determined by mass spectrometry.
[0019] A second aspect of the invention provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the foregoing methods.
[0020] A third aspect of the invention provides a computer program product for storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the foregoing methods.
[0021] A fourth aspect of the invention provides a computer device including a processor and a memory, the memory storing a computer program configured to be executed by the processor, the computer program, when executed by the processor, implementing the steps of any of the aforementioned methods. Attached Figure Description
[0022] Figure 1 Flowchart for identifying methoxy / fluorinated modification sites.
[0023] Figure 2The result of first-level mass spectrometry deconvolution of the m_seq_1 modified sequence.
[0024] Figure 3 UV spectrum of RNAase T1 enzyme digestion product of m_seq_1 modified sequence.
[0025] Figure 4 Mass spectrometry results of RNAase T1 enzyme digestion products of m_seq_1 modified sequences.
[0026] Figure 5 Matching results of RNAase T1 enzyme digestion products of m_seq_1 modified sequences.
[0027] Figure 6 TIC map of the combined digestion products of m_seq_1 modified sequence RNase T1 and RNase A enzymes.
[0028] Figure 7 Mass spectrometry results of m_seq_1 modified sequence digested with RNase T1 and RNase A enzymes.
[0029] Figure 8 Summary of measured molecular weights of RNase T1 and RNase A enzyme digestion products with m_seq_1 modified sequences.
[0030] Figure 9 : m_seq_1 modified sequence RNase T1 enzyme and RNase A enzyme co-digestion expected sequence information.
[0031] Figure 10 The result of matching the expected sequence and the measured sequence information of the m_seq_1 modified sequence RNase T1 enzyme and RNase A enzyme combined digestion.
[0032] Figure 11 : Secondary mass spectrometry verification results of m_seq_1 modified sequences.
[0033] Figure 12 The result of first-level mass spectrometry deconvolution of the f_seq_1 modified sequence.
[0034] Figure 13 UV spectrum of RNAase T1 enzyme digestion product of f_seq_1 modified sequence.
[0035] Figure 14 Mass spectrometry results of RNAase T1 enzyme digestion products of f_seq_1 modified sequences.
[0036] Figure 15 Matching results of RNAase T1 enzyme digestion products of f_seq_1 modified sequences.
[0037] Figure 16 Mass spectrometry results of RNase T1 and RNase A enzyme co-digestion products of f_seq_1 modified sequence.
[0038] Figure 17 Summary of measured molecular weights of RNase T1 and RNase A enzyme digestion products with f_seq_1 modified sequences.
[0039] Figure 18 : f_seq_1 modified sequence RNase T1 enzyme and RNase A enzyme co-digestion expected sequence information.
[0040] Figure 19 The result of matching the expected sequence and the measured sequence information of the f_seq_1 modified sequence RNase T1 enzyme and RNase A enzyme combined digestion.
[0041] Figure 20 : Secondary mass spectrometry verification results of f_seq_1 modified sequences. Detailed Implementation
[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0043] All publications, patent applications, patents, and other references mentioned herein are incorporated herein by reference in their entirety. In case of conflict, this specification (including definitions) shall prevail. Furthermore, the materials, methods, and examples described herein are illustrative only and not intended to be restrictive.
[0044] When the terms “about” and “approximately” are used with numerical variables, they generally mean that the value of the variable and all values of the variable are within the measurement or experimental error (e.g., the 95% confidence interval of the mean) or within a wider range of specified values (e.g., ±5% or ±10%).
[0045] The term "comprising," or its variations such as "containing," "having," or "including," means to include the stated steps or elements, but does not exclude any other steps or elements. "Constitutes of," means to exclude steps or elements not listed. "Substantially constitutes of," means to include steps or elements that do not substantially affect the fundamental and novel features of the protected invention. The term "comprising" and its variations also include the cases of "consisting of specific steps or elements" and "substantially constitutes specific steps or elements."
[0046] When referring to a numerical range, it should be understood that the specific values of its upper and lower limits are disclosed, as well as all intermediate ranges included therein, such as the intermediate range between its upper or lower limit and any intermediate value, or the intermediate range between any two intermediate values. Furthermore, any intermediate ranges, subranges, and all individual numerical values described in the numerical range can be excluded from the numerical range.
[0047] Unless otherwise expressly stated in the context, the terms "first" and "second" are used merely to distinguish different elements and are not intended to imply, expressly or imply, that the elements have a specific spatial, temporal, or logical relationship.
[0048] In this document, when referring to molecular weight (or atomic weight), such as measured molecular weight (or atomic weight) or theoretical molecular weight (or atomic weight), it can be an exact molecular weight (or atomic weight). In some embodiments, the exact molecular weight (or atomic weight) is accurate to two decimal places.
[0049] The term “and / or” should be understood as any one of the multiple elements connected by the term, or a combination of any number of elements.
[0050] In this invention, "nucleotide" and "base" are used interchangeably and are usually represented by conventional single letters, where A is adenine deoxyribonucleotide or adenine ribonucleotide, C is cytosine deoxyribonucleotide or adenine ribonucleotide, G is guanine deoxyribonucleotide or adenine ribonucleotide, T is thymine deoxyribonucleotide, and U is uracil ribonucleotide.
[0051] The term "RNA sequence" in this invention has the same meaning as "RNA" and "RNA molecule," referring to a linear polymer composed of two or more ribonucleotide monomers polymerized by 3',5'-phosphodiester bonds. The term "ribonucleotide" refers to a nucleotide having a hydroxyl group at the 2' position of the β-D-furanose moiety. The ribonucleotide monomers included in the RNA sequence can be natural ribonucleotides having the following bases: adenine, cytosine, guanine, and uracil, abbreviated as A, C, G, and U, respectively, or can be ribonucleotides modified relative to these natural ribonucleotides, for example, modified at the base, ribose, and / or phosphate (phosphate backbone) moieties. Modifications to the phosphate (phosphate backbone) can include, for example, thiophosphate modifications. Modifications to the ribose can include, for example, the introduction of groups of different sizes and polarities at the 2' position, such as 2'-methoxy, 2'-methoxyethoxy, and / or 2'-fluoro (i.e., 2'-deoxy-2'-fluoro) modifications. They can also include modifications present simultaneously at the 2' and other ribose sites, such as modifications at the 2' position, the 4' position, or the entire sugar ring, such as locked nucleic acid modifications, unlocked nucleic acid modifications, restriction ethyl-bridged nucleic acid modifications, tricyclic DNA modifications, ethylene glycol nucleic acid modifications, phosphodiesteramide morpholino oligonucleotide (PMO) modifications, etc. Base modifications can include, for example, N6-methyladenosine (m6A), N1-methyladenosine (m1A), 5-methylcytidine (m5C), pseudouridine, N1-methylpseudouridine, 5-methoxyuridine, etc. In some embodiments, all ribonucleotide monomers in the RNA sequence do not have base and / or phosphate (phosphate backbone) modifications relative to the native ribonucleotides A, C, G, and U. In some embodiments, the RNA sequence contains only modifications to the ribose. In some embodiments, the RNA sequence contains only base modifications. In some embodiments, the RNA sequence contains both ribose modifications and base modifications. In some embodiments, the ribose and / or base modifications in the RNA sequence are selected from one, two, three, or four of the following: methoxy modifications, methoxyethoxy modifications, fluorinated modifications, and locked nucleic acid modifications. In some embodiments, the ribose and / or base modifications in the RNA sequence are methoxy modifications and / or fluorinated modifications.
[0052] Unless otherwise stated, nucleic acids are written from left to right in the 5' to 3' direction in this document, and amino acid sequences are written from left to right in the direction from the amino terminus to the carboxyl terminus.
[0053] The RNA sequence to be tested in this invention can be any type or any function of RNA, such as siRNA, shRNA, miRNA, gRNA (e.g., for use in CRISPR / Cas systems), antisense RNA, mRNA, tRNA, rRNA, or fragments thereof.
[0054] The length of the RNA sequence to be identified by the method of the present invention (i.e., the number of ribonucleotide monomers contained therein) can be 10-200 ribonucleotides, for example 100-180, 10-160, 10-140, 10-120, 10-100, 10-80, 10-60, 10-50, 20-160, 20-140, 20-120, 20-100, 20-80, 20-60, 20-50, 30-160, 30-140, 30-120, 30-100, 30-80, 30-60 or 30-50 ribonucleotides.
[0055] The term "ribose and / or base modification" refers to modifications on the ribose or base of one or more nucleotides in an RNA sequence. Ribose modifications can include, for example, the introduction of groups of different sizes and polarities at the 2' position of the ribose, such as 2'-methoxy, 2'-methoxyethoxy, and / or 2'-fluoro modifications (i.e., 2'-deoxy-2'-fluoro modifications). Modifications can also include those present simultaneously at the 2' position and other ribose sites, such as modifications at the 2' and 4' positions or throughout the sugar ring, such as locked nucleic acid modifications, unlocked nucleic acid modifications, restriction ethyl-bridged nucleic acid modifications, tricyclic DNA modifications, ethylene glycol nucleic acid modifications, phosphodiesteramide morpholino oligonucleotide (PMO) modifications, etc. Base modifications refer to the presence of 2'-methoxy, 2'-methoxyethoxy, and / or 2'-fluoro modifications (i.e., 2'-deoxy-2'-fluoro modifications) on the bases. In some embodiments, the RNA sequence to be tested has one or more ribose and / or base modifications, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, or more ribose and / or base modifications, such as 1-200, 1-100, 1-50, 1-20, 1-10, or 1-7 ribose and / or base modifications. In some embodiments, the RNA sequence to be tested has consecutive ribose and / or base modifications or does not have consecutive ribose and / or base modifications. The term "consecutive ribose and / or base modification" refers to modifications on the ribose of two or more consecutive nucleotides in a nucleic acid sequence (e.g., RNA). In some embodiments, the two or more ribose and / or base modifications contained in the RNA sequence to be tested can be the same or different. In some embodiments, the RNA sequence to be tested can simultaneously contain two or more different ribose and / or base modifications, such as simultaneously containing two or more of the following: methoxy modification, fluorine modification, methoxyethoxy modification, and locked nucleic acid modification, such as simultaneously containing methoxy modification and fluorine modification.
[0056] The term "methoxy modification" refers to the methylation of the hydroxyl group (-OH) at the 2' position of the ribose in a ribonucleotide to form a methoxy group.
[0057] The term "fluorine modification" refers to the substitution of the hydroxyl group (-OH) at the 2' position of the ribose in a ribonucleotide by fluorine.
[0058] The term "ribose and / or base modification site" refers to the location of a ribonucleotide containing a ribose and / or base modification in an RNA sequence, i.e., on which ribonucleotide monomers of the RNA sequence to be tested are the ribose and / or base modification located.
[0059] The term "base modification site" refers to the location of a ribonucleotide containing a base modification in an RNA sequence; that is, which ribonucleotide monomers in the RNA sequence to be tested are the base modification located on.
[0060] The term "mother ion" refers to the charged phase ion formed after a compound is ionized, while the term "daughter ion" refers to the ion formed by the chemical bond breaking of a specific mother ion after collision with gas molecules.
[0061] In some embodiments, the RNA sequence to be tested has no more than a specific number of consecutive ribose and / or base modifications. Consecutive ribose and / or base modifications mean that consecutive ribonucleotides in the RNA sequence have ribose and / or base modifications. In some embodiments, the specific number refers to 10, for example, 9, 8, 7, 6, 5, 4, 3, or 2.
[0062] The term "RNA sequence to be tested" refers to an RNA sequence containing the ribose and / or base modifications to be tested, wherein the type, number, and / or location of the ribose and / or base modifications to be tested are unknown. The term "RNA sequence to be tested" can include RNA sequences in which the type, number, and / or location of all ribose and / or base modifications are unknown, and also includes RNA sequences in which the type, number, and location of some ribose and / or base modifications are known, but the type, number, and / or location of some ribose and / or base modifications are unknown. In this invention, RNA sequences whose ribose and / or base modifications have not yet been detected by the methods described herein can be referred to as raw RNA sequences to be tested. Once the type, number, and / or location of some ribose and / or base modifications have been determined by one or more steps of the methods described herein, such sequences can be referred to as intermediate RNA sequences to be tested, which are subsequently used to determine the type, number, and / or location of other ribose and / or base modifications through subsequent steps.
[0063] The term "naked sequence" is used to describe the ribonucleotide composition and sequence of a test RNA sequence that does not contain information about ribose and / or base modifications (such as methoxy or fluorinated modifications). The term "naked sequence" can also be understood as a sequence that does not contain ribose and / or base modifications (such as methoxy or fluorinated modifications), has the same nucleic acid sequence as the test RNA sequence, but does not have any ribose and / or base modifications.
[0064] The terms "fragment," "polynucleotide fragment," "nucleic acid fragment," or "RNA fragment" are used interchangeably in this invention and refer to a portion of a nucleic acid sequence (such as the RNA sequence to be tested, naked sequence, or theoretical sequence described herein). The term "fragment" can be a separately cut fragment or a fragment covalently linked to other fragments within a longer sequence. The term "fragment" can be an actual enzyme-digested fragment or a theoretical enzyme-digested fragment.
[0065] The term "experimental fragment" refers to a fragment obtained through real biological experiments (e.g., cutting with ribonuclease) and / or chemical experiments (e.g., enzyme digestion).
[0066] The terms “detected sequence” or “detected fragment” refer to the nucleic acid sequence or fragment in a sample that is actually detected (e.g., detected using a detection device, such as mass spectrometry) and gives a detection result (e.g., determining its molecular weight by detection).
[0067] The term “theoretical sequence” or “theoretical fragment” refers to a hypothetical sequence that does not actually exist and has not been actually digested by enzymes and / or detected (e.g., detected using detection equipment, such as mass spectrometry).
[0068] The term "measured molecular weight" refers to the molecular weight of a component or components in a sample, determined by actual detection (e.g., detection using detection equipment, such as mass spectrometry).
[0069] The term "theoretical molecular weight" refers to the molecular weight calculated based on the elements contained in a nucleic acid sequence, nucleic acid fragment, molecule, or atom.
[0070] The term "ribonuclease" refers to an enzyme that hydrolyzes the phosphodiester bonds in the middle of an RNA molecule. Ribonucleases suitable for the present invention include site-specific ribonucleases, which can specifically cleave RNA at specific sites within the RNA molecule. Site-specific ribonucleases, for example, are nucleases that can specifically cleave the 5' or 3' side of an RNA sequence immediately adjacent to A, C, G, or U, or specifically cleave a specific number of nucleotides adjacent to A, C, G, or U, or specifically cleave a specific position within a specific ribonucleotide fragment; these include, but are not limited to, RNase A and RNase T1.
[0071] The term "RNase T1" may also be referred to as "T1 enzyme" in this invention, which specifically cleaves the phosphodiester bond on the 3' side of a guanine ribonucleotide (G) in an RNA sequence. The T1 enzyme can cleave the phosphodiester bond on the 3' side of a G without ribose and / or base modification, but cannot cleave the phosphodiester bond on the 3' side of a G with ribose and / or base modification. The fragment obtained by T1 enzyme cleavage of RNA may have a 3' end without ribose and / or base modification; when it contains G at positions other than the 3' end, it should be a G with ribose and / or base modification, not a G without ribose and / or base modification.
[0072] The term "RNase A" may also be referred to as "A enzyme" in this invention. It specifically cleaves the phosphodiester bonds on the 3' side of cytosine ribonucleotides (C) and uracil ribonucleotides (U) in an RNA sequence. A enzyme can cleave the phosphodiester bonds on the 3' side of C and U without ribose and / or base modification, but cannot cleave the phosphodiester bonds on the 3' side of C and U with ribose and / or base modification. The fragment obtained by A enzyme cleavage of RNA may have a 3' end without ribose and / or base modification (C or U). When it contains C and / or U at positions other than the 3' end, it should be C and / or U with ribose and / or base modification, not C and / or U without ribose and / or base modification.
[0073] The term "characteristic nucleotide" refers to a ribonuclease that determines the cleavage site of an RNA sequence. For a ribonuclease that specifically cleaves the RNA sequence at the phosphodiester bond immediately adjacent to the 3' side of a particular ribonuclease, that particular ribonuclease is its characteristic nucleotide. Modifications to the characteristic nucleotide (e.g., ribose and / or base modifications) can prevent the ribonuclease from cleaving at its specific cleavage site. For example, for the T1 enzyme, its characteristic nucleotide is G. For the A enzyme, its characteristic nucleotides are C and U.
[0074] The term "enzyme-digested fragment" refers to a fragment obtained by cleaving an RNA sequence with a ribonuclease. Those skilled in the art will understand that enzyme-digested fragments can include single nucleotide, dinucleotide, and / or polynucleotide fragments (e.g., fragments containing three or more nucleotides). In this document, when referring to the sequence of an enzyme-digested fragment, "single nucleotide" means the single nucleotide itself.
[0075] The term "type" of ribose and / or base modification refers to different types of ribose and / or base modifications, such as methoxy modification, fluorination modification, locked nucleic acid modification, methoxyethoxy modification, etc., which belong to different types of ribose and / or base modifications.
[0076] The term "theoretical enzyme digestion" or a similar term refers to a virtual enzyme digestion that does not actually occur.
[0077] The term "theoretical restriction fragment" is a hypothetical fragment, referring to the expected restriction fragment that will be obtained by digesting a specific RNA sequence with a specific ribonuclease.
[0078] The term "match" refers to two fragments that have the same nucleic acid sequence, but may differ only in chemical modifications (such as ribose and / or base modifications).
[0079] The term "ribose and / or base modification pattern" refers to the types, locations, and quantities of all ribose and / or base modifications contained in a modified RNA sequence.
[0080] The term "ribonucleotide sequence" refers to the order in which each ribonucleotide that makes up an RNA sequence is arranged in a specific direction (e.g., the 5' to 3' direction) within that RNA sequence.
[0081] In this article, when referring to the same ribonucleotide, it can be understood that the ribonucleotides have the same phosphate and base moieties and the same sequence, but the ribose moieties may differ due to modifications; when referring to the same nucleic acid sequence, it can be understood that two or more nucleic acid sequences have the same ribonucleotide sequence, and the ribonucleotides at the same position in these nucleic acid sequences have the same phosphate and base moieties, but the ribose moieties may differ due to modifications.
[0082] In this article, when comparing molecular weights, it should be understood that due to experimental errors, the precise values may not be exactly the same. Therefore, when two molecular weights are very close, they can be considered equal, and it is not required that the two molecular weights be exactly the same in value. There can be slight differences between them, for example, a difference of no more than 0.2 Da, or a difference of no more than 0.1 Da.
[0083] The terms "mass spectrometry," "mass spectrometry detection," or "MS" refer to the technique of detecting analytes based on their mass-to-charge ratio. Mass spectrometry detection typically involves ionizing the sample and then accelerating the ions through an electric or magnetic field. Ions with the same mass-to-charge ratio will undergo the same deflection and be detected by mechanisms that detect charged particles, such as electron multipliers.
[0084] Mass spectrometry is typically used to detect analytes. A mass spectrometer usually consists of an ionization source and a mass analyzer. Based on the working principle of the mass analyzer, it can be classified as a dual-focusing mass spectrometer, a quadrupole mass spectrometer (QMS), an ion trap mass spectrometer, a Fourier transform mass spectrometer (FT-MS), a time-of-flight mass spectrometer (TOF-MS), etc. Based on the ionization source, mass spectrometers can be classified as electron impact ionization mass spectrometers, chemical ionization mass spectrometers, matrix-assisted laser desorption / ionization mass spectrometers, electrospray ionization mass spectrometers, etc. It should be understood that different ionization sources and different mass analyzers can be combined. Different mass spectrometers are commercially available. Mass spectrometry can also be coupled with other detection methods, such as gas chromatography-mass spectrometry (GC-MS), liquid chromatography-mass spectrometry (LC-MS), and capillary electrophoresis-mass spectrometry (CES-MS).
[0085] The mass spectrometry detection described in this invention can be performed using the different types of mass spectrometers mentioned above. In some embodiments, mass spectrometry detection is performed using a detection method that couples mass spectrometry with other detection methods. In some embodiments, mass spectrometry detection is performed using liquid chromatography-mass spectrometry (LC-MS).
[0086] Mass spectrometry is a technique well-known to those skilled in the art. They understand how to select different mass spectrometers or coupled detection methods, and can learn, through routine experiments, how to set detection parameters in mass spectrometry. The data obtained from mass spectrometry includes the mass-to-charge ratio (m / z) and relative intensity (intensity is proportional to the amount of analyte) of the detected ions; ions with different mass-to-charge ratios appear as different peaks. In liquid chromatography-mass spectrometry (LC-MS) coupled detection methods, the obtained detection data also includes the retention time (RT) of the ions. In mass spectrometry, peaks (ions) with specific mass-to-charge ratios, specific intensities, and / or specific retention times can be extracted for subsequent analysis as needed.
[0087] The term "primary mass spectrometry" or "primary mass spectrometry detection" refers to a mass spectrometry method for detecting the mass-to-charge ratio and intensity of charged ions contained in an ionized sample.
[0088] The terms "secondary mass spectrometry" or "secondary mass spectrometry detection" may also be referred to as "tandem mass spectrometry" or "tandem mass spectrometry detection" in this invention. It comprises two mass spectrometry detection stages. Typically, in the first stage (primary mass spectrometry), charged ions contained in the ionized sample are detected and screened. Subsequently, one or more of these charged ions are used as parent ions to fragment them into daughter ions, and the daughter ions are then detected by mass spectrometry in the second stage (secondary mass spectrometry). Tandem mass spectrometry can detect the mass-to-charge ratio and intensity of charged ions contained in the ionized sample through primary mass spectrometry, and detect the mass-to-charge ratio and intensity of daughter ions through secondary mass spectrometry.
[0089] As those skilled in the art know, secondary mass spectrometry can be used to determine whether a tested nucleic acid sequence (e.g., a DNA sequence or an RNA sequence) or amino acid sequence is identical to a reference theoretical sequence. This can be achieved by performing secondary mass spectrometry on the tested sequence, comparing the measured daughter ion information with the theoretical secondary mass spectrometry daughter ion information obtained from the reference theoretical sequence, and if they match perfectly, it indicates that the tested sequence is identical to the reference theoretical sequence.
[0090] In this invention, the primary mass spectrometry detection can be a standalone primary mass spectrometry detection, i.e., the sample undergoes only one mass spectrometry detection, or it can be the primary mass spectrometry in a tandem mass spectrometry system. In some embodiments, the primary mass spectrometry detection is a detection method that combines primary mass spectrometry with other detection methods, such as liquid chromatography-mass spectrometry (LC-MS) detection.
[0091] In some implementations, a secondary mass spectrometry detection method is employed, combining tandem mass spectrometry with other detection methods, such as liquid chromatography-tandem mass spectrometry (LC-MSMS). In some implementations, the secondary mass spectrometry is the FullMS / ddms2 method (full scan of primary mass spectrometry plus data-dependent secondary mass spectrometry scan), which is a mass spectrometry detection method that fragments the ion while acquiring primary mass spectrometry data to generate secondary mass spectrometry fragments, thus simultaneously acquiring primary and secondary mass spectrometry data.
[0092] This invention provides a method for identifying ribose and / or base modifications in a test RNA sequence. This method involves molecular weight analysis of the complete test RNA sequence and its digested fragments (single and double enzyme digestions), comparing the results with theoretically digested fragments obtained from theoretical sequences containing or without ribose and / or base modifications. This comparison reveals the type, quantity, and location of ribose and / or base modifications in the test RNA sequence. Based on this information, the specific type, quantity, and location of ribose and / or base modifications in the test RNA sequence can be uniquely determined, or the possible ribose and / or base modification patterns in the test RNA sequence can be limited to a finite range. This facilitates subsequent analysis to determine the specific type, quantity, and location of ribose and / or base modifications in the test RNA sequence.
[0093] In some implementations, the raw nucleic acid sequence of the RNA sequence to be tested is known, but the type, quantity, and / or location of ribose and / or base modifications are unknown.
[0094] In some implementations, the raw nucleic acid sequence of the RNA to be tested is known, and the type, quantity, and location of ribose and / or base modifications are known.
[0095] In some implementations, the naked nucleic acid sequence of the RNA sequence to be tested is known, and the number, type, and location of ribose and / or base modifications are known.
[0096] In some implementations, the naked nucleic acid sequence of the RNA sequence to be tested is known, the type and number of ribose and / or base modifications are known, and the location is unknown.
[0097] The method includes the following steps: Step 1: Based on the difference between the measured molecular weight of the RNA sequence to be tested and the theoretical molecular weight of the naked sequence of the RNA sequence to be tested, determine the type and / or quantity of ribose and / or base modifications in the RNA sequence to be tested, which is constraint condition one; Step 2: Obtain the first measured molecular weight of the RNA sequence to be tested after digestion with the first ribonuclease. Match the first measured molecular weight with the first theoretical molecular weight of the first theoretical digestion fragment obtained by theoretically digesting the naked sequence of the RNA sequence to be tested with the first ribonuclease. This determines the nucleic acid sequence and partial ribose and / or base modification information of each digestion fragment obtained by specific digestion of the RNA sequence to be tested with the first ribonuclease. This is constraint condition 2. Optionally, obtain the second measured molecular weight of the RNA sequence to be tested after double digestion with the first and second ribonucleases. Match the second measured molecular weight with the second theoretical molecular weight of the second theoretical digestion fragment. The second theoretical digestion fragment is obtained by double digestion of the naked sequence of the RNA sequence to be tested with the first and second ribonucleases, which has been modified. This determines the nucleic acid sequence and ribose and / or base modification information of each digestion fragment obtained by specific double digestion of the RNA sequence to be tested with the first and second ribonucleases. This is constraint condition 3. Step 3: Using the information obtained in Steps 1 and 2 (which may include constraint 1, constraint 2, and optionally constraint 3), determine all possible ribose and / or base modification patterns in the RNA sequence to be tested; and Step 4: When there is only one possible combination of ribose and / or base modification patterns, the modification patterns of all ribose and / or bases in the RNA sequence to be tested are determined; when there is more than one possible combination of ribose and / or base modification patterns, the RNA sequence to be tested is further subjected to secondary mass spectrometry detection to determine the modification patterns of all ribose and / or bases in the RNA sequence to be tested.
[0098] In some implementations, the phrase "addition of a pre-determined modified naked RNA sequence" in step two refers to a modified sequence formed by adding a modification previously determined in a previous step to the naked sequence of the RNA to be tested. Specifically, the pre-determined modification may refer to the modification information of ribose and / or bases obtained in step two by matching the first measured molecular weight with the first theoretical molecular weight. For example, if in step two, matching the first measured molecular weight with the first theoretical molecular weight reveals a methoxy or fluorinated modification at one or more specific positions, then "addition of a pre-determined modified naked RNA sequence" refers to a modified sequence formed by adding the methoxy or fluorinated modification at the corresponding position in the naked sequence, while leaving other positions unmodified.
[0099] In some implementations, step two includes: obtaining the molecular weight of the first actual digested fragment obtained by digestion of the RNA sequence to be tested with a first ribonuclease, and obtaining the first measured molecular weight, wherein each first measured molecular weight represents a first actual digested fragment; using the naked sequence of the RNA sequence to be tested as the first theoretical sequence, digesting it with a theoretical first ribonuclease to obtain the first theoretical digested fragment; and, based on satisfying constraint one, performing a first matching analysis on the first theoretical digested fragment and the digested fragment represented by the first measured molecular weight, thereby determining the nucleic acid sequence and partial ribose and / or base modification information of each digested fragment obtained by specific digestion of the RNA sequence to be tested with a first ribonuclease, which is constraint two; The first matching analysis includes comparing the theoretical molecular weight of the first theoretical digestion fragment with the first measured molecular weight to determine the first theoretical digestion fragment that matches each digestion fragment represented by the first measured molecular weight, as well as the additional ribose and / or base modifications contained in each digestion fragment represented by the first measured molecular weight relative to the first theoretical digestion fragment that matches it. The partial ribose and / or base modification information determined by the first matching analysis includes the type and number of ribose and / or base modifications contained in each digestion fragment obtained by specific digestion of the test RNA sequence with a first ribonuclease, and also includes precise information on whether a specific ribonucleotide has ribose and / or base modifications (specifically, the specific ribonucleotide refers to the characteristic nucleotide of each first ribonuclease contained in the test RNA sequence, excluding the nucleotide at the 3' end of the test RNA sequence); and / or The molecular weight of the second actual digested fragment obtained by double digestion of the RNA sequence to be tested with the first and second ribonucleases is obtained, and each second measured molecular weight represents a second actual digested fragment. The sequence of the modified fragment formed by adding ribose and / or base modifications at specific positions determined in the first matching analysis to the naked sequence is the second theoretical sequence. Double digestion with the theoretical first and second ribonucleases is performed to obtain the second theoretical digested fragment. Based on the satisfaction of constraint one and constraint two, a second matching analysis is performed on the digested fragments represented by the second theoretical digested fragment and the second measured molecular weight to determine the nucleic acid sequence and the types and numbers of ribose and / or base modifications contained in each digested fragment obtained by specific double digestion of the RNA sequence to be tested with the first and second ribonucleases, which is constraint three. The second matching analysis includes comparing the theoretical molecular weight of the second theoretical digestion fragment with the second measured molecular weight to determine the second theoretical digestion fragment that matches each digestion fragment represented by the second measured molecular weight, as well as the additional ribose and / or base modifications contained in each digestion fragment represented by the second measured molecular weight relative to the matching second theoretical digestion fragment. The partial ribose and / or base modification information determined according to the second matching analysis includes the type and number of ribose and / or base modifications contained in each digestion fragment obtained by specific double digestion of the RNA sequence to be tested by the first ribonuclease and the second ribonuclease, and also includes the exact information on whether a specific ribonucleotide has ribose and / or base modifications (specifically, the specific ribonucleotide refers to the characteristic nucleotide of each of the first or second ribonuclease contained in the RNA sequence to be tested, excluding the nucleotide at the 3' end of the RNA sequence to be tested). Through steps one and two above, the range of types, quantities, and positions of ribose and / or base modifications can be gradually narrowed down, ultimately obtaining a unique or limited range of possible ribose and / or base modification patterns.
[0100] In some implementations, when only one possible ribose and / or base modification pattern is identified in step three, the type, quantity, and location of the ribose and / or base modification in the RNA sequence to be tested can be directly determined, or further confirmation can be made by secondary mass spectrometry. When more than one possible ribose and / or base modification pattern is identified in step four, the RNA sequence to be tested can be further subjected to secondary mass spectrometry to determine the type, quantity, and location of the ribose and / or base modification in the RNA sequence to be tested. In some implementations, when more than one possible ribose and / or base modification pattern is identified in step three, a modified RNA sequence having the possible ribose and / or base modification pattern is used as a reference theoretical sequence, and the RNA sequence to be tested is subjected to secondary mass spectrometry to determine the type, quantity, and location of the ribose and / or base modification in the RNA sequence to be tested.
[0101] The molecular weight of a nucleic acid sequence or fragment can be obtained by any molecular weight determination method. In some embodiments, the molecular weight determination method can measure the molecular weight to two decimal places; preferably, to three decimal places; more preferably, to four decimal places. In some embodiments, the molecular weight of a nucleic acid sequence or fragment can be determined by mass spectrometry, such as by primary mass spectrometry.
[0102] The first and second ribonucleases are different ribonucleases and can be site-specific ribonucleases. In some embodiments, one of the first and second ribonucleases is RNase T1 and the other is RNase A. In some embodiments, the first ribonuclease is RNase T1 and the second ribonuclease is RNase A.
[0103] The enzyme digestion fragments mentioned above, including both actual and theoretical digestion fragments, may contain one or more nucleic acid fragments; that is, they may be a mixture of multiple nucleic acid fragments. Correspondingly, the measured molecular weight of the actual digestion fragment and the theoretical molecular weight of the theoretical digestion fragment may also contain multiple different molecular weights.
[0104] This invention identifies the type, location, and / or quantity of additional ribose and / or base modifications in the actual sequence relative to the corresponding theoretical sequence by comparing the molecular weights of the enzyme-digested fragments obtained from the actual sequence and the corresponding theoretical sequence by digestion with the same enzyme. This allows for the determination of the type, location, and / or quantity of ribose and / or base modifications in the actual sequence.
[0105] In this article, when referring to the order of two or more theoretical digestion fragments (e.g., mentioning that one theoretical digestion fragment comes before or after another, or that they are adjacent to each other), the order refers to their arrangement in the corresponding theoretical sequence, where "before" refers to the 5' side and "after" refers to the 3' side.
[0106] In some implementations, step one can be omitted when the type and amount of ribose and / or base modifications in the RNA sequence to be tested are known.
[0107] In some implementations, the measured molecular weight of the RNA sequence to be tested is known, and step one may include determining the type of ribose and / or base modifications in the RNA sequence to be tested and the amount of each ribose and / or base modification based on the difference between the known measured molecular weight of the RNA sequence to be tested and the theoretical molecular weight of the bare sequence of the RNA sequence to be tested.
[0108] In some implementations, step one may further include obtaining the measured molecular weight of the RNA sequence to be tested, specifically the position of the measured molecular weight of the RNA sequence to be tested.
[0109] In some implementations, step one may further include performing primary mass spectrometry detection on the RNA sequence to be tested (which has not been treated with any nucleases, such as endonucleases and / or exonucleases) to obtain the measured molecular weight of the RNA sequence to be tested.
[0110] In some implementations, in step one, the type and amount of ribose and / or base modifications in the RNA sequence to be tested can be determined according to the following method: First, calculate the total molecular weight of all ribose and / or base modifications in the RNA sequence to be tested using the following formula: The total molecular weight of ribose and / or base modifications = molecular weight 实测 -Molecular weight 裸序列理论 Among them, "molecular weight" 实测 "Molecular weight" refers to the measured molecular weight of the RNA sequence to be tested, obtained (e.g., determined by primary mass spectrometry). 裸序列理论 "" refers to the theoretical molecular weight of the naked sequence of the RNA to be tested.
[0111] Subsequently, the theoretical molecular weight of different combinations of ribose and / or base modifications is compared with the total molecular weight of ribose and / or base modifications in the RNA sequence to be tested calculated above. When the theoretical molecular weight of a specific combination of ribose and / or base modifications is equal to the total molecular weight of ribose and / or base modifications in the RNA sequence to be tested calculated above, it can be determined that the RNA sequence to be tested contains the specific combination of ribose and / or base modifications.
[0112] The term "ribose and / or base modification combination" refers to a combination of different numbers of different types of ribose and / or base modifications. A ribose and / or base modification combination may contain one or more types of ribose and / or base modifications, and the number of each type of ribose and / or base modification may be one or more. The theoretical molecular weight of a ribose and / or base modification combination is the sum of the theoretical molecular weights of all ribose and / or base modifications contained in the combination; for example, the theoretical molecular weight M of a ribose and / or base modification combination containing n types of ribose and / or base modifications. c It can be calculated using the following formula.
[0113]
[0114] Where M i N is the theoretical molecular weight of the i-th ribose and / or base modification. i It represents the number of the i-th type of ribose and / or base modification.
[0115] For each ribose and / or base modification, the theoretical molecular weight of a single ribose and / or base modification refers to the difference in molecular weight between the ribose with the modification and the ribose without the modification. The theoretical molecular weights of different ribose and / or base modifications are well known to those skilled in the art; for example, the molecular weight of a methoxy modification is about 14.02 Da, and the molecular weight of a fluorinated modification is about 1.996 Da.
[0116] When only one type of ribose and / or base modification exists in the RNA sequence to be tested, the total number of such ribose and / or base modifications can also be determined according to the following formula: Total number of ribose and / or base modifications = (molecular weight) 实测 -Molecular weight 裸序列理论 ) / Molecular weight 单个核糖和 / 或碱基修饰 Among them, "molecular weight" 单个核糖和 / 或碱基修饰 "" refers to the theoretical molecular weight of a single ribose and / or base modification.
[0117] It should be understood that the value calculated based on this formula may not be an integer, but those skilled in the art will understand that the calculated value will be close to an integer, for example, the difference between it and a certain integer value is not greater than 0.2, such as not greater than 0.1, not greater than 0.09, not greater than 0.08, not greater than 0.07, not greater than 0.06, not greater than 0.05, not greater than 0.04, not greater than 0.03, not greater than 0.02 or not greater than 0.01. In this case, the calculated value is rounded to the nearest integer, which is the number of ribose and / or base modifications.
[0118] The information regarding ribose and / or base modifications of the RNA sequence to be tested, determined in Step One, is referred to as Constraint One in this invention. Further identification of ribose and / or base modifications in subsequent steps must satisfy Constraint One.
[0119] In some implementations, when the types and quantities of ribose and / or base modifications in the RNA sequence to be tested are known, step one can be omitted, and the known types and quantities of ribose and / or base modifications in the RNA sequence to be tested can be used directly as constraint one.
[0120] In some implementations, step two includes obtaining the first measured molecular weight of the RNA sequence to be tested after digestion with a first ribonuclease, and performing a matching analysis between the first measured molecular weight and the first theoretical molecular weight of the first theoretical digested fragment obtained by theoretical digestion of the naked sequence of the RNA sequence to be tested with the first ribonuclease, thereby determining the nucleic acid sequence and partial ribose and / or base modification information of each digested fragment obtained by specific digestion of the RNA sequence to be tested with the first ribonuclease.
[0121] In some implementations, this step alone can provide all information on the ribose and / or base modifications of the RNA sequence to be tested, eliminating the need for a second ribonuclease. In other implementations, this step alone may not provide all information on the ribose and / or base modifications of the RNA sequence to be tested. In such cases, step two may further include obtaining the second measured molecular weight of the RNA sequence to be tested after double digestion with a first and a second ribonuclease. This second measured molecular weight is then matched with the second theoretical molecular weight of the second theoretical digested fragment obtained from the double digestion with the first and second ribonucleases. This process determines the nucleic acid sequence and the ribose and / or base modifications of each digested fragment obtained from the specific double digestion with the first and second ribonucleases of the RNA sequence to be tested.
[0122] In some embodiments, step two may further include digesting the RNA sequence to be tested with a first ribonuclease to obtain a first actual digested fragment. In some embodiments, step two may further include performing primary mass spectrometry detection on the first actual digested fragment to obtain a first measured molecular weight.
[0123] In some embodiments, step two may further include sequentially or simultaneously digesting the RNA sequence to be tested with a first ribonuclease and a second ribonuclease to obtain a second actual digested fragment. For example, the first actual digested fragment can be digested with a second ribonuclease to obtain the second actual digested fragment. In some embodiments, step two may further include performing primary mass spectrometry detection on the second actual digested fragment to obtain a second measured molecular weight.
[0124] Step two involves fragmenting the RNA sequence. The actual digested RNA sequence containing ribose and / or base modifications is compared with corresponding theoretical digested RNA sequences that either do not contain ribose and / or base modifications or contain partially identified ribose and / or base modifications. Differences in molecular weight are used to identify theoretical digested RNA sequences that match the actual digested RNA sequences (i.e., have the same nucleic acid sequence, differing only in ribose and / or base modifications). This allows for the determination of the actual digested RNA sequence. Furthermore, the types and amounts of ribose and / or base modifications contained in the actual digested RNA sequence can be determined, as can the presence of ribose and / or base modifications on certain specific ribonucleotides within the actual digested RNA sequence. Additionally, by considering the positions of ribonucleotides permitted to contain ribose and / or base modifications within the actual digested RNA sequence, the positions of ribose and / or base modifications can be further determined.
[0125] The sequences of the first / second theoretical enzyme digestion fragments described in step two are known, and their molecular weights can be calculated. Calculating the molecular weight of a nucleic acid fragment with a given sequence is a conventional technique known to those skilled in the art, for example, based on whether each ribonucleotide constituting the fragment has a phosphate group attached to its 5' and 3' ends. For example, those skilled in the art know that chemically synthesized or transcribed RNA molecules typically have a phosphate group at their 5' end and not at their 3' end; while ribonucleases cleave the phosphodiester bond between the 3'-phosphate group of a ribonucleotide and the 5'-hydroxyl group of the adjacent ribonucleotide, thus the ribonucleotides formed by ribonuclease cleavage do not have a phosphate group at their 5' end, but do have a phosphate group at their 3' end.
[0126] In this article, the theoretical enzyme digestion fragment that matches the enzyme digestion fragment represented by a certain measured molecular weight refers to the theoretical enzyme digestion fragment that has the same nucleic acid sequence as the enzyme digestion fragment represented by the measured molecular weight, and may differ only in ribose and / or base modifications. Those skilled in the art will understand that a digestion fragment represented by a measured molecular weight can be matched with a single theoretical digestion fragment or with two or more consecutive theoretical digestion fragments, depending on whether there is ribose and / or base modification on the characteristic ribonucleotide of the ribonuclease contained in the digestion fragment represented by the measured molecular weight. When there is ribose and / or base modification on the characteristic nucleotide, the ribonuclease used cannot cleave at the cleavage position determined by the characteristic nucleotide. However, for the theoretical digestion fragment, since the theoretical sequence to be digested does not have the ribose and / or base modification, the theoretical digestion fragment is cleaved at the cleavage position. Therefore, the digestion fragment represented by the measured molecular weight will have the same nucleic acid sequence as an extended theoretical digestion fragment formed by covalently linking two or more consecutive theoretical digestion fragments, and the characteristic nucleotide contained in the previous theoretical digestion fragment of any two adjacent theoretical digestion fragments has ribose and / or base modification.
[0127] In this document, when theoretical fragments are referred to as "continuous" or "adjacent," it means that they are sequentially adjacent in their respective theoretical sequences. Those skilled in the art will understand that in step two, if a second measured molecular weight-represented enzyme fragment matches two or more consecutive theoretical enzyme fragments, the consecutive theoretical enzyme fragments cannot exceed the range of the first actual enzyme fragment. That is, two or more theoretical enzyme fragments from different first actual enzyme fragments cannot connect to form an extended theoretical enzyme fragment because the second actual enzyme fragment is confined within the first actual enzyme fragment; each second actual enzyme fragment can only be a product of a certain first actual enzyme fragment digested by a second ribonuclease.
[0128] In some implementations, matching a theoretical digestion fragment with a digestion fragment represented by a measured molecular weight means that the two have either of the following relationships (i) and (ii): (i) A theoretical digestion fragment can be matched as a single fragment with the digestion fragment represented by the measured molecular weight, including either of the following two cases (a) and (b): (a) The molecular weight of a single theoretical digestion fragment is equal to the measured molecular weight; in this case, the digestion fragment represented by the measured molecular weight has the same sequence as the single theoretical digestion fragment and does not contain any additional ribose and / or base modifications relative to the single theoretical digestion fragment; or (b) A single theoretical restriction fragment has a molecular weight smaller than a certain measured molecular weight, but contains a specific combination of ribose and / or base modifications such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modifications to the single theoretical restriction fragment is equal to the measured molecular weight (i.e., the sum of the molecular weight of the theoretical restriction fragment and the theoretical molecular weight of the specific combination of ribose and / or base modifications is equal to the measured molecular weight), provided that the addition of the specific combination of ribose and / or base modifications satisfies the established ribose and / or base modification constraints; in this case, the restriction fragment represented by the measured molecular weight is the modified fragment formed by adding the specific combination of ribose and / or base modifications to the theoretical restriction fragment, and when the restriction fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the restriction fragment represented by the measured molecular weight does not have ribose and / or base modifications; and (ii) An extended theoretical restriction fragment, formed by covalently linking two or more consecutive theoretical restriction fragments, has a molecular weight smaller than that of a restriction fragment represented by a measured molecular weight, but contains a specific combination of ribose and / or base modifications such that the molecular weight of the modified fragment formed by adding the specific ribose and / or base modification combination to the extended theoretical restriction fragment is equal to the measured molecular weight (i.e., the sum of the molecular weight of the extended theoretical restriction fragment and the theoretical molecular weight of the specific ribose and / or base modification combination is equal to the measured molecular weight), provided that the addition of the specific ribose and / or base modification combination satisfies the condition that... The determined ribose and / or base modification constraints; in this case, the enzyme digestion fragment represented by the measured molecular weight is the modified fragment formed by adding the specific combination of ribose and / or base modifications to the extended theoretical enzyme digestion fragment, wherein the ribonucleotide at the 3' end of the preceding theoretical enzyme digestion fragment in any two adjacent theoretical enzyme digestion fragments has ribose and / or base modifications, and when the enzyme digestion fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the enzyme digestion fragment represented by the measured molecular weight does not have ribose and / or base modifications.
[0129] For situations (a) and (b) above, when performing a matching analysis, you can first determine whether a theoretical enzyme digestion fragment matches situation (a). If it does not match, then determine whether it matches situation (b).
[0130] In cases (b) and (ii) of relationship (i) above, the added ribose and / or base modification combination should satisfy the already determined ribose and / or base modification constraints. "The already determined ribose and / or base modification constraints" include constraints determined in the previous step and constraints determined in earlier steps within the same step. For example, when performing matching analysis in step two, the added ribose and / or base modification combination should satisfy the ribose and / or base modification constraint I determined in step one and the constraints determined in steps already performed in step two. In some implementations, "determined ribose and / or base modification constraints" include constraints determined by the constraints determined in the previous step and the types, locations, and / or quantities of ribose and / or base modifications determined in the same step. For example, when performing matching analysis in step two, if the ribose and / or base modification information contained in one or more enzyme digestion fragments represented by the first / second measured molecular weight has been determined, when continuing to perform matching analysis on the remaining first / second theoretical enzyme digestion fragments and enzyme digestion fragments represented by the first / second measured molecular weight, "determined ribose and / or base modification constraints" may include constraints determined by the first constraint determined in step one and the ribose and / or base modification information determined in step two.
[0131] Those skilled in the art will understand that when RNA sequences are cleaved using certain ribonucleases (e.g., RNase A), both specific and non-specific cleavage may occur simultaneously. Therefore, the actual digested fragments obtained may include both specific and non-specific fragments. During matching analysis, the theoretically matched digested fragment can only match specific digested fragments, not non-specific ones. However, those skilled in the art will understand that identifying only the ribose and / or base modifications contained in the specific digested fragments that the theoretically matched fragment can match is sufficient to identify the ribose and / or base modifications in the entire RNA sequence being tested. This is because the ribose and / or base modifications contained in these specific digested fragments are sufficient to cover the ribose and / or base modifications in the entire RNA sequence being tested. Analysis of non-specific digested fragments is neither feasible nor necessary. Based on the above analysis, it is clear that the theoretically matched digested fragment must be a specific digested fragment.
[0132] In some implementations, the matching analysis in step two can refer to the matching analysis of the first / second theoretical enzyme digestion fragment and the specific enzyme digestion fragment represented by the first / second measured molecular weight. This can be achieved, for example, by performing matching analysis on the first / second theoretical enzyme digestion fragment one by one.
[0133] In some implementations, in step two, the first / second measured molecular weight may include the measured molecular weight representing each actual digestion fragment in the first / second actual digestion fragment. In this case, a matching analysis can be performed on all theoretical digestion fragments and the digestion fragments represented by the measured molecular weights. That is, a matching analysis is performed on each theoretical digestion fragment so that each theoretical digestion fragment is matched to a digestion fragment represented by a certain measured molecular weight (or, each specific digestion fragment represented by the measured molecular weight finds a matching theoretical digestion fragment). At this point, it can be considered that all specific digestion fragments have been matched, which is sufficient to determine the nucleic acid sequence and its ribose and / or base modification information of all digestion fragments obtained from the specific digestion of the RNA sequence to be tested. At this time, some digestion fragments represented by the measured molecular weights may not be matched. These digestion fragments may be non-specific digestion fragments and can be ignored, as they will not affect the identification of ribose and / or base modifications in the RNA sequence to be tested.
[0134] In this document, "matching a theoretical restriction enzyme fragment to a restriction enzyme fragment represented by a measured molecular weight" means that the theoretical restriction enzyme fragment matches the restriction enzyme fragment represented by the measured molecular weight as a single fragment, or that the theoretical restriction enzyme fragment is contained within an extended theoretical restriction enzyme fragment formed by covalently linking two or more consecutive theoretical restriction enzyme fragments that match the restriction enzyme fragment represented by the measured molecular weight. When an extended theoretical restriction enzyme fragment formed by covalently linking two or more consecutive theoretical restriction enzyme fragments matches a restriction enzyme fragment represented by a measured molecular weight, all two or more consecutive theoretical restriction enzyme fragments are considered to have matched to the restriction enzyme fragment represented by the measured molecular weight (also referred to as "matched" in this document). It can be understood that during the matching analysis, when an extended theoretical restriction enzyme fragment formed by linking a theoretical restriction enzyme fragment with one or more subsequent consecutive theoretical restriction enzyme fragments matches a restriction enzyme fragment represented by a measured molecular weight, the two or more theoretical restriction enzyme fragments contained within it are considered to have matched, and there is no need to repeat the matching analysis. Matching analysis can continue for other unmatched theoretical restriction enzyme fragments.
[0135] It should be noted that each theoretical enzyme digestion fragment can only be matched to one enzyme digestion fragment represented by the measured molecular weight, and cannot be matched to two enzyme digestion fragments represented by two measured molecular weights at the same time.
[0136] In some implementations, the matching analysis of all theoretical digestion fragments with the digestion fragments represented by the measured molecular weight in step two can be achieved by sequentially performing the following steps: S1: Select a theoretical enzyme digestion fragment and proceed with step S2; S2: Check whether the molecular weight of the theoretical enzyme digestion fragment is equal to a certain measured molecular weight, and make the judgment in step S3. S3: If the molecular weight of the theoretical enzyme digestion fragment is equal to a measured molecular weight, then the enzyme digestion fragment represented by the measured molecular weight is determined to have the same sequence as the theoretical enzyme digestion fragment and does not contain any additional ribose and / or base modifications relative to the theoretical enzyme digestion fragment. Then, check whether there are other theoretical enzyme digestion fragments that have not yet been matched. If so, select one of the unmatched theoretical enzyme digestion fragments and perform step S2 on it. If not, complete all matching. If the molecular weight of the theoretical enzyme digestion fragment is not equal to any of the measured molecular weights, then perform step S4 on the theoretical enzyme digestion fragment. S4: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the theoretical enzyme digestion fragment is equal to a certain measured molecular weight (i.e., the sum of the molecular weight of the theoretical enzyme digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to the measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S5. S5: If such a combination of ribose and / or base modifications exists, the enzyme fragment represented by the measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the theoretical enzyme fragment. When the enzyme fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the enzyme fragment represented by the measured molecular weight does not have ribose and / or base modifications. Then, it is checked whether there are other theoretical enzyme fragments that have not been matched. If so, one unmatched theoretical enzyme fragment is selected and subjected to step S2. If not, all matching is completed. If such a combination of ribose and / or base modifications does not exist, a theoretical enzyme fragment adjacent to it before and / or after it is added to the theoretical enzyme fragment to form an extended theoretical enzyme fragment and subjected to step S6. S6: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the extended theoretical digestion fragment is equal to a certain measured molecular weight (i.e., the sum of the molecular weight of the extended theoretical digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to a certain measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S7. S7: If such a combination of ribose and / or base modifications exists, the enzyme digestion fragment represented by the measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the extended theoretical enzyme digestion fragment. In any two adjacent theoretical enzyme digestion fragments, the ribonucleotide at the 3' end of the preceding theoretical enzyme digestion fragment has ribose and / or base modifications. When the enzyme digestion fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the enzyme digestion fragment represented by the measured molecular weight does not have ribose and / or base modifications. Then, it is checked whether there are other theoretical enzyme digestion fragments that have not yet been matched. If so, one unmatched theoretical enzyme digestion fragment is selected and subjected to step S2. If not, all matching is completed. If such a combination of ribose and / or base modifications does not exist, the extended theoretical enzyme digestion fragment is added to a theoretical enzyme digestion fragment adjacent to it before and / or after it to form a further extended theoretical enzyme digestion fragment, which is then subjected to step S6. In some implementations, in step two, each theoretical restriction enzyme fragment can be matched sequentially. For example, all theoretical restriction enzyme fragments can be arranged sequentially (e.g., in a 5' to 3' or 3' to 5' order). Starting with the first theoretical restriction enzyme fragment in the sequence, each theoretical restriction enzyme fragment is matched sequentially to a restriction enzyme fragment represented by a measured molecular weight, until each theoretical restriction enzyme fragment is matched to a restriction enzyme fragment represented by the measured molecular weight. It is understood that during the matching analysis, when an extended theoretical restriction enzyme fragment formed by connecting a theoretical restriction enzyme fragment with one or more subsequent consecutive theoretical restriction enzyme fragments matches a restriction enzyme fragment represented by a measured molecular weight, the two or more theoretical restriction enzyme fragments contained therein are considered matched, and there is no need to repeat the matching analysis. Matching can be directly performed on the first unmatched theoretical restriction enzyme fragment following the extended theoretical restriction enzyme fragment.
[0137] In some implementations, the matching analysis performed sequentially on each theoretical restriction fragment in step two can be achieved by sequentially performing the following steps: S1: Perform step S2 on the theoretical enzyme digestion fragment that is number 1 in sequence; S2: Check whether the molecular weight of the theoretical enzyme digestion fragment is equal to a certain measured molecular weight, and make the judgment in step S3. S3: If the molecular weight of the theoretical enzyme digestion fragment is equal to a measured molecular weight, then the enzyme digestion fragment represented by the measured molecular weight is determined to have the same sequence as the theoretical enzyme digestion fragment and does not contain any additional ribose and / or base modifications relative to the theoretical enzyme digestion fragment. Subsequently, it is checked whether there are other theoretical enzyme digestion fragments that have not yet been matched after the theoretical enzyme digestion fragment. If so, step S2 is performed on the first unmatched theoretical enzyme digestion fragment after the theoretical enzyme digestion fragment. If not, all matching is completed. If the molecular weight of the theoretical enzyme digestion fragment is not equal to any of the measured molecular weights, step S4 is performed on the theoretical enzyme digestion fragment. S4: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the theoretical enzyme digestion fragment is equal to a certain measured molecular weight (i.e., the sum of the molecular weight of the theoretical enzyme digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to the measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S5. S5: If such a combination of ribose and / or base modifications exists, the enzyme fragment represented by the measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the theoretical enzyme fragment. When the enzyme fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the enzyme fragment represented by the measured molecular weight does not have ribose and / or base modifications. Then, it is checked whether there are other theoretical enzyme fragments that have not been matched after the theoretical enzyme fragment. If so, step S2 is performed on the first unmatched theoretical enzyme fragment after the theoretical enzyme fragment. If not, all matching is completed. If such a combination of ribose and / or base modifications does not exist, the first theoretical enzyme fragment after the theoretical enzyme fragment is added to the theoretical enzyme fragment to form an extended theoretical enzyme fragment, and step S6 is performed on it. S6: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the extended theoretical digestion fragment is equal to a certain measured molecular weight (i.e., the sum of the molecular weight of the extended theoretical digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to a certain measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S7. S7: If such a combination of ribose and / or base modifications exists, then the enzyme fragment represented by the measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the extended theoretical enzyme fragment, wherein the ribonucleotide at the 3' end of the preceding theoretical enzyme fragment in any two adjacent individual theoretical enzyme fragments has ribose and / or base modifications, and when the enzyme fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the enzyme fragment represented by the measured molecular weight is... The glyconucleotide does not have ribose and / or base modifications. Then, it is checked whether there are other theoretical nucleotide fragments that have not been matched after the extended theoretical nucleotide fragment. If so, step S2 is performed on the first unmatched theoretical nucleotide fragment after the extended theoretical nucleotide fragment. If not, all matching is completed. If there is no such combination of ribose and / or base modifications, the extended theoretical nucleotide fragment is combined with a theoretical nucleotide fragment adjacent to it before and / or after it to form a further extended theoretical nucleotide fragment, and step S6 is performed on it. In some embodiments, in step two, the first / second measured molecular weight may also include only the measured molecular weight representing a portion of the actual enzyme digestion fragment in the first / second actual digestion fragment. In some embodiments, to improve detection efficiency, in step two, the first / second measured molecular weight may include only the measured molecular weight representing a longer actual enzyme digestion fragment (e.g., an actual enzyme digestion fragment with two or more ribonucleotides), excluding the measured molecular weight representing a single nucleotide.
[0138] Accordingly, in this case, if the corresponding theoretical digestion fragment contains one or more single nucleotides, when performing matching analysis on each theoretical digestion fragment, matching analysis can be performed on theoretical digestion fragments containing two or more nucleotides. In this case, for the above sequential analysis, when encountering a single nucleotide theoretical digestion fragment, its matching analysis can be skipped, and matching analysis can be directly performed on the first non-single nucleotide theoretical digestion fragment following the single nucleotide theoretical digestion fragment. However, when performing matching analysis on the first non-single nucleotide theoretical digestion fragment following the single nucleotide theoretical digestion fragment, if it is necessary to form an extended theoretical digestion fragment based on the theoretical digestion fragment, it is necessary to consider the case where the single nucleotide theoretical digestion fragment is added to the extended theoretical digestion fragment to form the extended theoretical digestion fragment.
[0139] In the case of matching analysis for theoretical restriction fragments containing two or more nucleotides, when all theoretical restriction fragments containing two or more nucleotides have been matched to a restriction fragment represented by a certain measured molecular weight, it is also possible to determine whether the ribonucleotide of the corresponding single nucleotide theoretical restriction fragment in the RNA sequence to be tested has ribose and / or base modification based on the matching results. Specifically, in this case, for any ribonucleotide in the RNA sequence to be tested that corresponds to a single nucleotide theoretical digestion fragment, if it has ribose and / or base modification, it will covalently link with other ribonucleotides in the digestion fragment represented by the measured molecular weight to form a digestion fragment of at least two nucleotides represented by the measured molecular weight, and will not be cleaved into a single nucleotide. In this case, there must be an extended theoretical digestion fragment formed by the single nucleotide theoretical digestion fragment and its adjacent theoretical digestion fragment that matches the digestion fragment represented by the measured molecular weight. If it does not have ribose and / or base modification, it will be cleaved into a single nucleotide. When the matching is complete, the digestion fragments represented by the measured molecular weight matched by the theoretical digestion fragment will not contain the single nucleotide. Based on the matching situation of the theoretical digestion fragment and the digestion fragment represented by the measured molecular weight, it is possible to infer whether each ribonucleotide in the RNA sequence to be tested that corresponds to a single nucleotide theoretical digestion fragment in the theoretical digestion fragment has ribose and / or base modification. In other words, for any ribonucleotide in the RNA sequence to be tested that corresponds to a single nucleotide theoretical digestion fragment, if, upon completion of the matching analysis, the extended theoretical digestion fragment formed by the single nucleotide theoretical digestion fragment and its adjacent theoretical digestion fragment matches a digestion fragment represented by a measured molecular weight, then the ribonucleotide has ribose and / or base modification; if none of the digestion fragments represented by the measured molecular weight matched by the theoretical digestion fragment contain the ribonucleotide, then the ribonucleotide does not have ribose and / or base modification.
[0140] In some embodiments, in step two, when determining the molecular weight of the enzyme digestion fragment represented by the measured molecular weight using mass spectrometry, mass spectrometry detection data containing only longer actual enzyme digestion fragment components can be screened out, and the measured molecular weight of these components can be determined based on the mass spectrometry detection data. This screening can be achieved by setting a molecular weight range (or mass-to-charge ratio range) and obtaining mass spectrometry peak data of fragments that conform to the molecular weight range. In some embodiments, the molecular weight range may be, for example, greater than a certain lower molecular weight limit. The lower molecular weight limit may be, for example, between 450 Da and 600 Da, such as 450 Da, 500 Da, 550 Da, or 600 Da. In some embodiments, the molecular weight range may be, for example, greater than a certain lower mass-to-charge ratio limit. The lower mass-to-charge ratio limit may be, for example, between 450 m / z and 600 m / z, such as 450 m / z, 500 m / z, 550 m / z, or 600 m / z.
[0141] In some implementations, in step two, the matching analysis for theoretical cleavage fragments containing two or more nucleotides can be performed by sequentially executing the following steps: S1: Select a theoretically digestible fragment containing two or more nucleotides and proceed to step S2; S2: Check whether the molecular weight of the theoretical enzyme digestion fragment is equal to a certain measured molecular weight, and make the judgment in step S3. S3: If the molecular weight of the theoretical digestion fragment is equal to a measured molecular weight, then the digestion fragment represented by the measured molecular weight is determined to have the same sequence as the theoretical digestion fragment and does not contain any additional ribose and / or base modifications relative to the theoretical digestion fragment. Then, it is checked whether there are other theoretical digestion fragments containing two or more nucleotides that have not yet been matched. If so, one of the unmatched theoretical digestion fragments containing two or more nucleotides is selected and subjected to step S2. If not, all matching is completed. If the molecular weight of the theoretical digestion fragment is not equal to any of the measured molecular weights, then step S4 is performed on the theoretical digestion fragment. S4: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the theoretical enzyme digestion fragment is equal to a certain measured molecular weight (i.e., the sum of the molecular weight of the theoretical enzyme digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to the measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S5. S5: If such a combination of ribose and / or base modifications exists, the enzyme fragment represented by the measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the theoretical enzyme fragment. When the enzyme fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the enzyme fragment represented by the measured molecular weight does not have ribose and / or base modifications. Then, it is checked whether there are other theoretical enzyme fragments containing two or more nucleotides that have not yet been matched. If so, one unmatched theoretical enzyme fragment containing two or more nucleotides is selected and subjected to step S2. If not, all matching is completed. If such a combination of ribose and / or base modifications does not exist, a theoretical enzyme fragment adjacent to it before and / or after it is added to the theoretical enzyme fragment to form an extended theoretical enzyme fragment and subjected to step S6. S6: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the extended theoretical digestion fragment is equal to a certain measured molecular weight (i.e., the sum of the molecular weight of the extended theoretical digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to a certain measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S7. S7: If such a combination of ribose and / or base modifications exists, then the enzyme fragment represented by the measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the extended theoretical enzyme fragment, wherein the ribonucleotide at the 3' end of the preceding theoretical enzyme fragment in any two adjacent individual theoretical enzyme fragments has ribose and / or base modifications, and when the enzyme fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the 3' end of the enzyme fragment represented by the measured molecular weight... The ribonucleotide does not have ribose and / or base modifications. Then, it is checked whether there are other theoretical cleavage fragments containing two or more nucleotides that have not been matched. If they exist, one of the unmatched theoretical cleavage fragments containing two or more nucleotides is selected and subjected to step S2. If they do not exist, all matching is completed. If there is no such combination of ribose and / or base modifications, the extended theoretical cleavage fragment is added to a theoretical cleavage fragment adjacent to it before and / or after it to form a further extended theoretical cleavage fragment and subjected to step S6.
[0142] In some implementations, when performing matching analysis on theoretical cleavage fragments containing two or more nucleotides, if the matching analysis is performed sequentially, it can be achieved by sequentially performing the following steps: S1: Perform step S2 on the theoretically digestible fragment containing two or more nucleotides that is ranked first. S2: Check whether the molecular weight of the theoretical enzyme digestion fragment is equal to a certain measured molecular weight, and make the judgment in step S3. S3: If the molecular weight of the theoretical digestion fragment is equal to a measured molecular weight, then the digestion fragment represented by the measured molecular weight is determined to have the same sequence as the theoretical digestion fragment and does not contain any additional ribose and / or base modifications relative to the theoretical digestion fragment. Then, it is checked whether there are other theoretical digestion fragments containing two or more nucleotides that have not yet been matched after the theoretical digestion fragment. If so, the first unmatched theoretical digestion fragment containing two or more nucleotides after the theoretical digestion fragment is selected and subjected to step S2. If not, all matching is completed. If the molecular weight of the theoretical digestion fragment is not equal to any of the measured molecular weights, then step S4 is performed on the theoretical digestion fragment. S4: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the theoretical enzyme digestion fragment is equal to a certain measured molecular weight (i.e., the sum of the molecular weight of the theoretical enzyme digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to the measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S5. S5: If such a combination of ribose and / or base modifications exists, the enzyme fragment represented by the measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the theoretical enzyme fragment. When the enzyme fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the enzyme fragment represented by the measured molecular weight does not have ribose and / or base modifications. Then, it is checked whether there are other theoretical enzyme fragments containing two or more nucleotides that have not yet been matched after the theoretical enzyme fragment. If so, one is selected from them. The first unmatched theoretical enzyme digestion fragment containing two or more nucleotides following the theoretical enzyme digestion fragment is subjected to step S2. If it does not exist, all matching is completed. If no such ribose and / or base modification combination exists, it is checked whether there is a single nucleotide theoretical enzyme digestion fragment adjacent to the theoretical enzyme digestion fragment. If it exists, the theoretical enzyme digestion fragment is added to the previous single nucleotide theoretical enzyme digestion fragment to form an extended theoretical enzyme digestion fragment, and step S6 is performed. If it does not exist, the theoretical enzyme digestion fragment is added to the first theoretical enzyme digestion fragment following it to form an extended theoretical enzyme digestion fragment, and step S6 is performed. S6: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the extended theoretical digestion fragment is equal to a certain measured molecular weight (i.e., the sum of the molecular weight of the extended theoretical digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to a certain measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S7. S7: If such a combination of ribose and / or base modifications exists, then the enzyme fragment represented by the measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the extended theoretical enzyme fragment. Specifically, the ribonucleotide at the 3' end of the preceding theoretical enzyme fragment in any two adjacent individual theoretical enzyme fragments has ribose and / or base modifications. Furthermore, when the enzyme fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence being tested, the ribonucleotide at the 3' end of the enzyme fragment represented by the measured molecular weight does not have ribose and / or base modifications. Subsequently, it is checked whether other theoretical enzymes containing two or more nucleotides exist. If a theoretical cleavage fragment containing two or more nucleotides has not yet been matched, select one such fragment and perform step S2. If no matching is found, complete all matching steps. If no such ribose and / or base modification combination exists, check if a single nucleotide theoretical cleavage fragment exists adjacent to the extended theoretical cleavage fragment. If it exists, add the extended theoretical cleavage fragment to the previous single nucleotide theoretical cleavage fragment to form a further extended theoretical cleavage fragment and perform step S6. If no matching is found, add the extended theoretical cleavage fragment to the first theoretical cleavage fragment following it to form a further extended theoretical cleavage fragment and perform step S6.
[0143] The following example illustrates step two using T1 as the first ribonuclease and A as the second ribonuclease. Those skilled in the art will understand that this method can be used similarly when using other ribonucleases.
[0144] In step two, the RNA sequence to be tested is digested with the T1 enzyme to obtain the first actual digested fragment. The molecular weight of the first actual digested fragment is then detected to obtain one or more first measured molecular weights. Unknown ribose and / or base modifications are identified by comparing the theoretical digested fragment with the actual digested fragment represented by the measured molecular weight.
[0145] The T1 enzyme cleaves on the 3' side of each G that does not have ribose and / or base modification.
[0146] The first actual enzyme digestion fragment may contain one or more of the following fragments: (a) A fragment with the 5' end of the RNA sequence to be tested as the 5' end and the 3' end as a G without ribose and / or base modification; (b) One or more fragments, the 5' end of which is a ribonucleotide adjacent to the 3' side of a G in the RNA sequence to be tested that does not have ribose and / or base modification, and the 3' end of which does not have ribose and / or base modification; (c) A fragment with the 3' end of the RNA sequence to be tested, defined by the ribonucleotide adjacent to the 3' side of the G nucleotide that does not have ribose and / or base modification; and (d) One or more individual Gs without ribose and / or base modification (when the 5' ribonucleotide of the RNA sequence to be tested is a G without ribose and / or base modification, and / or there are two or more adjacent Gs without ribose and / or base modification in the RNA sequence to be tested).
[0147] The first actual enzyme digestion fragments mentioned above are all products of T1 enzyme-specific digestion.
[0148] In each of the first actual enzyme digestion fragments mentioned above, if there is a G at a position other than the 3' end, then the G at the non-3' end position has ribose and / or base modification.
[0149] If all Gs in the RNA sequence to be tested, located at positions other than the 3' end, have ribose and / or base modifications, then the first actual digestion fragment can be the same as the RNA sequence to be tested.
[0150] Using the naked sequence of the RNA to be tested as the theoretical sequence, it is digested with the theoretical T1 enzyme to obtain the first theoretical digestion fragment, which may contain one or more of the following fragments: (a) A fragment with the 5' end of the naked sequence as the 5' end and the first G in the naked sequence from 5' to 3' as the 3' end; (b) One or more fragments whose 5' end is a G in a naked sequence and whose 3' adjacent ribonucleotides are G at their 3' ends; (c) A fragment with the 5' end of the ribonucleotide adjacent to the 3' side of G in the naked sequence and the 3' end of the ribonucleotide at the 3' end of the naked sequence; and (d) One or more individual Gs (when the 5' nucleotide of the RNA sequence is G, and / or there are two or more adjacent Gs in the RNA sequence).
[0151] Each of the above first theoretical enzyme digestion fragments does not contain G at a position other than the 3' end.
[0152] If the naked sequence does not contain G, its theoretical digestion fragment can be the same as the naked sequence mentioned above.
[0153] Subsequently, the first theoretical restriction fragment and the restriction fragment represented by the first measured molecular weight were matched and analyzed using the following method: S1: Select a first theoretical enzyme digestion fragment and proceed to step S2; S2: Check whether the molecular weight of the first theoretical enzyme digestion fragment is equal to a certain first measured molecular weight, and make the judgment in step S3. S3: If the molecular weight of the first theoretical enzyme digestion fragment is equal to a certain first measured molecular weight, then it is determined that the enzyme digestion fragment represented by the first measured molecular weight has the same sequence as the first theoretical enzyme digestion fragment, and does not contain any additional ribose and / or base modifications relative to the first theoretical enzyme digestion fragment. Then, it is checked whether there are other first theoretical enzyme digestion fragments that have not been matched. If there are, then one of the unmatched first theoretical enzyme digestion fragments is selected and performed on it in step S2. If there are no, then all matching is completed. If the molecular weight of the first theoretical enzyme digestion fragment is not equal to any of the first measured molecular weights, then step S4 is performed on the first theoretical enzyme digestion fragment. S4: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the first theoretical enzyme digestion fragment is equal to a certain first measured molecular weight (i.e., the sum of the molecular weight of the first theoretical enzyme digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to the first measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S5. S5: If such a combination of ribose and / or base modifications exists, the enzyme digestion fragment represented by the first measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the first theoretical enzyme digestion fragment. When the enzyme digestion fragment represented by the first measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the G at the 3' end of the enzyme digestion fragment represented by the first measured molecular weight does not have ribose and / or base modifications. Then, it is checked whether there are other first theoretical enzyme digestion fragments that have not been matched. If there are, one first theoretical enzyme digestion fragment that has not been matched is selected and performed on it in step S2. If there are no, all matching is completed. If such a combination of ribose and / or base modifications does not exist, the first theoretical enzyme digestion fragment is added to a first theoretical enzyme digestion fragment adjacent to it before and / or after it to form an extended first theoretical enzyme digestion fragment and performed on it in step S6. S6: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the extended first theoretical digestion fragment is equal to a certain first measured molecular weight (i.e., the sum of the molecular weight of the extended first theoretical digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to a certain first measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S7. S7: If such a combination of ribose and / or base modifications exists, then the enzyme digestion fragment represented by the first measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the extended first theoretical enzyme digestion fragment, wherein the G at the 3' end of the preceding theoretical enzyme digestion fragment in any two adjacent individual theoretical enzyme digestion fragments has a ribose and / or base modification, and when the enzyme digestion fragment represented by the first measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the G at the 3' end of the enzyme digestion fragment represented by the first measured molecular weight does not have a ribose and / or base modification. / or base modification, and determine that all individual theoretical digestion fragments contained in the extended first theoretical digestion fragment have been matched, then check whether there are other first theoretical digestion fragments that have not been matched. If there are, select one of the unmatched first theoretical digestion fragments and perform step S2 on it. If there are no such combinations of ribose and / or base modification, add a first theoretical digestion fragment adjacent to it before and / or after it to form a further extended first theoretical digestion fragment and perform step S6 on it.
[0154] In some implementations, all first theoretical digestion fragments can be arranged sequentially (e.g., in a 5' to 3' or 3' to 5' order), and each first theoretical digestion fragment can be matched and analyzed sequentially, starting with the first theoretical digestion fragment in the first position.
[0155] In some implementations, when performing a matching analysis between the first theoretical digestion fragment and the digestion fragment represented by the first measured molecular weight, the first theoretical digestion fragment containing two or more ribonucleotides can be matched. When the first theoretical digestion fragment contains one or more individual Gs, the matching analysis results can be used to determine whether the ribonucleotides in the first actual digestion fragment corresponding to each individual G in the first theoretical digestion fragment have ribose and / or base modifications.
[0156] Through the above matching analysis, the nucleic acid sequence of each digested fragment obtained by T1 enzyme digestion of the RNA sequence to be tested can be obtained. It can also be obtained the type and number of ribose and / or base modifications contained in each digested fragment obtained by T1 enzyme digestion of the RNA sequence to be tested, as well as whether each G in the RNA sequence located at a position other than the 3' end has ribose and / or base modifications, which is constraint condition two.
[0157] Subsequently, optionally, in step two, the first actual enzyme digestion fragment obtained above is digested with enzyme A to obtain a second actual enzyme digestion fragment, and the molecular weight of the second actual enzyme digestion fragment is detected to obtain a second measured molecular weight. Unknown ribose and / or base modifications are then identified by comparing the theoretically digested fragment with the enzyme digestion fragment represented by the measured molecular weight.
[0158] Enzyme A cleaves at the 3' sides of each C molecule without ribose and / or base modification and each U molecule without ribose and / or base modification. The second actual digested fragment may contain one or more of the following fragments: (a) A fragment with the 5' end of the RNA sequence to be tested as the 5' end and the 3' end as G without ribose and / or base modification, C without ribose and / or base modification, or U without ribose and / or base modification; (b) One or more fragments, the 5' end of which is a ribonucleotide located adjacent to the 3' side of a G without ribose and / or base modification, a C without ribose and / or base modification, or a U without ribose and / or base modification in the RNA sequence to be tested, and the 3' end of which is a G without ribose and / or base modification, a C without ribose and / or base modification, or a U without ribose and / or base modification; (c) A fragment with the 5' end of the RNA sequence to be tested, using the ribonucleotide adjacent to the 3' side of a G sequence without ribose and / or base modification, a C sequence without ribose and / or base modification, or a U sequence without ribose and / or base modification, and the 3' end of the RNA sequence to be tested; and (d) One or more individual Gs without ribose and / or base modification, Cs without ribose and / or base modification, and / or Us without ribose and / or base modification (when the 5' ribonucleotide of the RNA sequence to be tested is G without ribose and / or base modification, C without ribose and / or base modification, or U without ribose and / or base modification, and / or there are two or more adjacent ribonucleotides selected from Gs without ribose and / or base modification, Cs without ribose and / or base modification, and Us without ribose and / or base modification in the RNA sequence to be tested).
[0159] The second actual enzyme digestion fragments mentioned above are all products of specific enzyme digestion by enzyme A.
[0160] In each of the above-mentioned second actual enzyme digestion fragments, if there is G, C and / or U at a position other than the 3' end, then the G, C and / or U at the position other than the 3' end are all modified with ribose and / or bases.
[0161] If the C and U positions at positions other than the 3' end of the RNA sequence to be tested both have ribose and / or base modifications, then the second actual digestion fragment can be the same as the first actual digestion fragment.
[0162] In addition to specific cleavage, enzyme A may also produce non-specific cleavage. During identification, it is sufficient to focus on the enzyme fragment represented by the second measured molecular weight that the second theoretical enzyme fragment can match to complete the matching analysis of the first theoretical enzyme fragment and the enzyme fragment represented by the first measured molecular weight. It is not required to find a single second theoretical enzyme fragment or an extended second theoretical enzyme fragment that matches the enzyme fragment represented by each second measured molecular weight.
[0163] The second theoretical sequence used to obtain the second theoretical digestion fragment is a naked sequence plus a modified fragment sequence formed by adding ribose and / or base modifications whose specific positions have been determined in the matching analysis of the first theoretical digestion fragment and the digestion fragment represented by the first measured molecular weight. The second theoretical digestion fragment obtained by theoretical double digestion of the second theoretical sequence using T1 enzyme and A enzyme may include one or more of the following fragments: (a) A fragment with the 5' end of the second theoretical sequence as the 5' end, and with G without ribose and / or base modification or with C or U as the 3' end; (b) One or more fragments, wherein the 5' ribonucleotide is a ribonucleotide located adjacent to the 3' side of a G without ribose and / or base modification in the second theoretical sequence or a ribonucleotide located adjacent to the 3' side of a C or U, wherein the 3' ribonucleotide is a G without ribose and / or base modification, or is a C or U; (c) A fragment with a 5' end at a ribonucleotide adjacent to the 3' side of a G in the second theoretical sequence that is not ribose and / or has base modification, or adjacent to the 3' side of a C or U, and a 3' end at the 3' end of the second theoretical sequence; and (d) One or more individual G, C and / or U without ribose and / or base modification (when the 5' ribonucleotide of the second theoretical sequence is G or C or U without ribose and / or base modification, and / or there are two or more adjacent ribonucleotides selected from C, U and G without ribose and / or base modification in the second theoretical sequence). In each of the above second theoretical digestion fragments: (i) if there is a G at a position other than the 3' end, then all Gs at the non-3' end positions have ribose and / or base modifications; and (ii) all other ribonucleotides except for the Gs at the non-3' end positions do not have ribose and / or base modifications.
[0164] If the naked sequence does not contain C and U, then the second theoretical digestion fragment can be the same as the second theoretical sequence.
[0165] Subsequently, the second theoretical restriction fragment and the restriction fragment represented by the second measured molecular weight were matched and analyzed using the following method: S1: Select a second theoretical enzyme digestion fragment and proceed to step S2; S2: Check whether the molecular weight of the second theoretical enzyme digestion fragment is equal to a certain second measured molecular weight, and make the judgment in step S3. S3: If the molecular weight of the second theoretical digestion fragment is equal to a certain second measured molecular weight, then the digestion fragment represented by the second measured molecular weight is determined to have the same sequence as the second theoretical digestion fragment, and does not contain any additional ribose and / or base modifications relative to the second theoretical digestion fragment. Then, check whether there are other second theoretical digestion fragments that have not been matched. If there are, select one of the unmatched second theoretical digestion fragments and perform step S2 on it. If there are no, complete all matching. If the molecular weight of the second theoretical digestion fragment is not equal to any of the second measured molecular weights, then perform step S4 on the second theoretical digestion fragment. S4: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the second theoretical enzyme digestion fragment is equal to a certain second measured molecular weight (i.e., the sum of the molecular weight of the second theoretical enzyme digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to the second measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S5. S5: If such a combination of ribose and / or base modifications exists, the enzyme fragment represented by the second measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the second theoretical enzyme fragment. When the enzyme fragment represented by the second measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the G, C, or U at the 3' end of the enzyme fragment represented by the second measured molecular weight does not have ribose and / or base modifications. Then, it is checked whether there are other second theoretical enzyme fragments that have not been matched. If there are, one unmatched second theoretical enzyme fragment is selected and performed on it in step S2. If there are no such matching, all matching is completed. If such a combination of ribose and / or base modifications does not exist, a second theoretical enzyme fragment adjacent to it before and / or after it is added to the second theoretical enzyme fragment to form an extended second theoretical enzyme fragment and performed on it in step S6. S6: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the extended second theoretical digestion fragment is equal to a certain second measured molecular weight (i.e., the sum of the molecular weight of the extended second theoretical digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to a certain second measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S7; S7: If such a combination of ribose and / or base modifications exists, then the enzyme fragment represented by the second measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the extended second theoretical enzyme fragment, wherein the 3' end of the first theoretical enzyme fragment in any two adjacent individual theoretical enzyme fragments has ribose and / or base modifications at the G, C, or U position, and when the enzyme fragment represented by the second measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the 3' end of the enzyme fragment represented by the second measured molecular weight does not have ribose and / or base modifications at the G, C, or U position. Having ribose and / or base modifications, and determining that all individual theoretical digestion fragments contained in the extended second theoretical digestion fragment have been matched, then checking whether there are any other second theoretical digestion fragments that have not been matched. If so, a second theoretical digestion fragment that has not been matched is selected and subjected to step S2; if not, all matching is completed. If no such combination of ribose and / or base modifications exists, the extended second theoretical digestion fragment is combined with a second theoretical digestion fragment adjacent to it before and / or after it to form a further extended second theoretical digestion fragment and subjected to step S6.
[0166] In some implementations, all second theoretical digestion fragments can be arranged sequentially (e.g., in a 5' to 3' or 3' to 5' order), and each second theoretical digestion fragment can be matched and analyzed sequentially, starting with the first second theoretical digestion fragment in sequence.
[0167] In some implementations, when performing matching analysis on the second theoretical cleavage fragment and the cleavage fragment represented by the second measured molecular weight, the matching analysis can be performed on the second theoretical cleavage fragment containing two or more ribonucleotides, as described above. When the second theoretical cleavage fragment contains one or more individual G, C, and / or U, the matching analysis results can be used to determine whether each individual G, C, and / or U has ribose and / or base modification.
[0168] Through the above matching analysis, the nucleic acid sequence of each digested fragment obtained by double digestion of the RNA sequence to be tested with T1 enzyme and A enzyme can be obtained. It can also be obtained that each digested fragment obtained by double digestion of the RNA sequence to be tested with T1 enzyme and A enzyme contains the type and number of ribose and / or base modifications, as well as whether each G, C and U located at non-3' end positions in the RNA sequence to be tested has ribose and / or base modifications, which is constraint condition three.
[0169] The specific method for determining all possible ribose and / or base modification patterns according to constraint three is within the scope of technology mastered by those skilled in the art, and can be achieved, for example, through logical reasoning.
[0170] As an example, and not to limit the scope of the invention, the following example illustrates how possible ribose and / or base modification patterns can be determined according to constraint three: (1) For the restriction fragments obtained by double digestion of the RNA sequence to be tested with T1 and A enzymes, the restriction fragment located at the 3' end of the RNA sequence to be tested: If the number of ribose and / or base modifications contained in the restriction fragment is equal to the number of G, C, and U with ribose and / or base modifications already identified on the restriction fragment, then the 3' terminal ribonucleotide of the restriction fragment and its contained A (if present) do not have ribose and / or base modifications; if the number of ribose and / or base modifications contained in the restriction fragment is greater than the number of G, C, and U with ribose and / or base modifications already identified on the restriction fragment, and the difference between the two is equal to the number of 3' terminal ribonucleotides of the restriction fragment and its contained A (if present) nucleotides, then the restriction fragment is considered unrestricted. The 3' terminal ribonucleotide of the digestion fragment and all A's (if present) therein have ribose and / or base modifications. If the number of ribose and / or base modifications in the digestion fragment is greater than the number of G, C, and U's already identified as having ribose and / or base modifications in the digestion fragment, and the difference between the two is less than the number of 3' terminal ribonucleotides and A's (if present) in the digestion fragment, then a portion of the 3' terminal ribonucleotides and A's (if present) in the digestion fragment have ribose and / or base modifications. In this case, it is impossible to determine which positions of the 3' terminal ribonucleotides and A's (if present) in the digestion fragment have ribose and / or base modifications, and there are multiple possibilities. (2) For the restriction fragments obtained by double digestion of the RNA sequence to be tested with T1 and A enzymes, other than the restriction fragment located at the 3' end of the RNA sequence to be tested: if the number of ribose and / or base modifications contained in the restriction fragment is equal to the number of G, C, and U with ribose and / or base modifications already identified on the restriction fragment, then the A (if present) contained therein does not have ribose and / or base modifications; if the number of ribose and / or base modifications contained in the restriction fragment is greater than the number of G, C, and U with ribose and / or base modifications already identified on the restriction fragment, and the difference between the two is equal to the 3' terminal ribonucleoside of the restriction fragment. If the number of nucleotides containing nucleotides and the number of A nucleotides in the digestion fragment are large enough, then the 3' terminal ribonucleotide of the digestion fragment and all the A nucleotides it contains are modified with ribose and / or bases. If the number of ribose and / or base modifications in the digestion fragment is greater than the number of G, C, and U nucleotides with ribose and / or base modifications already identified in the digestion fragment, and the difference between the two is less than the number of nucleotides containing A nucleotides in the 3' terminal ribonucleotide of the digestion fragment, then a portion of the 3' terminal ribonucleotide of the digestion fragment and the A nucleotides it contains are modified with ribose and / or bases. In this case, it is impossible to determine which A nucleotides are modified with ribose and / or bases, and there are multiple possibilities. (3) For each digested fragment obtained by double digestion of the RNA sequence to be tested with T1 enzyme and A enzyme; if there is only one type of ribose and / or base modification on each digested fragment (the number of which can be one or more), then all ribose and / or base modifications on the digested fragment are that type of ribose and / or base modification; if there are two or more types of ribose and / or base modifications on each digested fragment, then it is impossible to determine which position of ribose and / or base modification is which type of ribose and / or base modification, and there are multiple possibilities.
[0171] The second actual digested fragments obtained by double digestion of the RNA sequence to be tested with T1 and A enzymes are usually short. After completing step two, the number and types of ribose and / or base modifications in these digested fragments have been determined, and it has been determined whether most positions (i.e., G, C, U) in these digested fragments have ribose and / or base modifications. At this point, only the A in these fragments and the 3' terminal ribonucleotide of the RNA sequence to be tested are allowed to have additional ribose and / or base modifications. Subsequently, for each second actual digested fragment, the number and types of ribose and / or base modifications can be determined based on the determined number and types of ribose and / or base modifications. Each enzyme fragment obtained by double digestion of the RNA sequence to be tested with T1 and A enzymes Because enzyme A cleaves at both the C and U positions, the resulting fragments are typically short, containing a limited number of ribose and / or base modifications. Therefore, in most cases, the ribose and / or base modification patterns of the RNA sequence being tested can be uniquely identified. In rare cases, the ribose and / or base modification patterns can be limited to a very narrow range of possibilities, such as reducing the number of possible patterns to at least 10 (e.g., fewer than 9, 8, 7, 6, 5, 4, 3, or even as few as 2). This allows for convenient determination of ribose and / or base modification patterns in the RNA sequence using secondary mass spectrometry. Specifically, it significantly reduces the number of theoretical reference sequences required for secondary mass spectrometry, thereby improving detection efficiency.
[0172] In some implementation schemes, any one or more of steps one through four above can be implemented by executing a computer program.
[0173] In some implementations, the method may further include step five: further verifying or determining the type, quantity, and location of all ribose and / or base modifications in the RNA to be tested using secondary mass spectrometry.
[0174] Secondary mass spectrometry detection can use a modified RNA sequence, formed by adding the ribose and / or base modification pattern determined in step four, to the bare sequence as a reference theoretical sequence.
[0175] If steps one through four can uniquely determine the type, quantity, and location of all ribose and / or base modifications in the RNA to be tested, then step five is unnecessary. Alternatively, step five can be used to further verify whether the type, quantity, and location of all ribose and / or base modifications in the RNA to be tested determined through steps one through four are correct. In this case, a modified RNA sequence formed by adding the ribose and / or base modification pattern determined in step four to a bare sequence can be used as a reference theoretical sequence. The RNA to be tested can be subjected to secondary mass spectrometry to determine whether the RNA to be tested is completely identical to the reference theoretical sequence in terms of ribonucleotide arrangement and ribose and / or base modifications.
[0176] If steps one through four cannot uniquely determine the types, quantities, and locations of all ribose and / or base modifications in the RNA to be tested, step five can be used to further verify which of the possible ribose and / or base modification patterns identified in steps one through four is present in the RNA sequence to be tested. In this case, different modified RNA sequences formed by adding different possible ribose and / or base modification patterns identified in step four to the bare sequence can be used as reference theoretical sequences. Secondary mass spectrometry analysis of the RNA sequence to be tested can then be performed to determine which reference theoretical sequence is identical to the RNA sequence to be tested in terms of ribonucleotide arrangement and ribose and / or base modifications.
[0177] Furthermore, those skilled in the art will recognize that the method of the present invention can be implemented by a computer program, for example, by performing any one or more steps of the above method by one or more computer programs, wherein the instructions in the program cause a computer or processor to perform the steps of the above method.
[0178] These programs can be stored and provided to a computer or processor using various types of non-transitory computer-readable media, which include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (such as floppy disks, magnetic tapes, and hard disk drives), magneto-optical recording media (such as magneto-optical disks), CD-ROMs (Compact Disc Read-Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (such as ROMs, PROMs (Programmable ROMs), EPROMs (Erasable and Writable PROMs), flash memory ROMs, and RAMs (Random Access Memory)).
[0179] These programs can also be included in various types of computer program products, which can be transmitted, distributed, and downloaded in the form of signals via wired or wireless communication paths such as wires and fiber optics, and are available, for example, via the Internet. Examples of computer program products include computer software, software packages, operating systems, applications, etc.
[0180] The present invention also provides a computer device including a processor and a memory, the memory storing a computer program configured to be executed by the processor, the computer program being configured to be executable by the processor and, when executed by the processor, to implement the steps in the above-described method.
[0181] The present invention is further described through the following embodiments, which should not be construed as limiting the invention. Unless otherwise specified, all reagents used in the following embodiments are commercially available products. Molecular biology experimental methods not specifically described in the embodiments were performed according to the specific methods listed in J. Sambrook, Molecular Cloning: A Laboratory Manual, Third Edition, or according to the kit and product instructions.
[0182] Example 1 1. Laboratory supplies 1.1 Experimental Samples
[0183] Note: m represents methoxyl modification. 1.2 Experimental Instruments and Equipment
[0184] 1.3 Experimental reagents and consumables
[0185] 1.4 Self-prepared reagents
[0186] 2. Experiment Content 2.1 Experiment Overview This study primarily utilizes various analytical techniques to narrow down the range of methoxy and / or fluorinated modification sites in RNA, reducing the originally infinite number of possible sequences to a finite number (e.g., single-digit possible sequences), and even directly identifying methoxy and / or fluorinated sites. Furthermore, secondary mass spectrometry can be used for true and false positive differentiation or verification, thereby achieving the identification of methoxy and / or fluorinated modification sites.
[0187] This study primarily utilizes various analytical techniques to narrow down the range of methoxy and / or fluorinated modification sites in RNA, reducing the originally infinite number of possible sequences to a finite number (e.g., single-digit possible sequences), and even directly identifying thiolated sites. Furthermore, secondary mass spectrometry can be used for true and false positive differentiation or verification, thereby achieving the identification of methoxy and / or fluorinated modification sites.
[0188] 2.2 Experimental Design First, high-resolution mass spectrometry is used to accurately determine the molecular weight of the sample, followed by... 实测 -Molecular weight 裸序列理论 ) / atomic weight 甲氧基 / 氟代 The first step involves determining the number of methoxy / fluorine modification sites. The second step uses RNase T1 digestion, which specifically cleaves the 3' end of G, breaking the RNA sequence into several fragments ending with G at their 3' ends. If the ribose of G is modified (e.g., methoxy or fluorine), G will resist digestion. By comparing the molecular weight of the actual digested fragments, we can infer whether there is a modification on the ribose of G and, more importantly, the number of methoxy / fluorine modifications in the RNA. The third step involves adding RNase A to the RNase T1 digestion product for further digestion. This enzyme specifically cleaves the 3' end of C and / or U (C / U). If the ribose of C and / or U has methoxy or fluorinated modifications, it will have natural resistance to digestion. The actual digestion fragments obtained through this treatment can be used to infer the presence or absence of methoxy or fluorinated modifications on C and / or U. Furthermore, by cutting the RNase T1 digestion product into smaller units and matching the molecular weight of the theoretical digestion sequence fragments with the measured digestion fragments, the modification composition of RNA can be largely determined. Combining the results of the first three steps with ensemble analysis, the methoxy and / or fluorinated modification sites can be almost certainly identified. Even if not, their probability is greatly reduced, generally to single digits. Finally, MSMS can be used to distinguish between true and false positive sequences or to verify individual sequences, thus ultimately achieving the goal of identifying methoxy and / or fluorinated modification sites. For the detection principle, please refer to [link to relevant documentation]. Figure 1 .
[0189] 3. Experimental Procedure 3.1 Preparation of LC MS mobile phase LC MS buffer A: Add 10 mL HFIPA + 500 μL DIEA to 1 L of water, mix well and sonicate to obtain mobile phase A.
[0190] LC MS buffer B: Add 10 mL HFIPA + 500 μL DIEA + 200 mL water to 800 mL acetonitrile, mix well and sonicate to obtain mobile phase B.
[0191] 3.2 Perform primary mass spectrometry (LC-MS) analysis on the samples. Take 1 nmol of the sample containing the modified sequence to be tested, insert it into the inner tube, add enzyme-free water to make up to 100 μL, and inject 10 μL for LS_MS analysis.
[0192] 3.3 RNase T1 enzyme digestion mass spectrometry analysis Take 2 nmol of the sample to be tested and place it in a PCR tube. Add 5 μL of buffer (200 mM Tris-HCl (pH 7.8), 400 mM KCl, 80 mM MgCl2, 10 mM DTT) and 1 μL of T1 enzyme. Add enzyme-free water to make up to 50 μL. Incubate at room temperature for 10 minutes, and then inject 10 μL for LC-MS analysis.
[0193] 3.4 RNase A enzyme digestion mass spectrometry analysis After digestion with T1 enzyme, add 1 μL of enzyme A (concentration of 100,000 U / μl diluted 1000 times) and incubate at room temperature for 10 minutes. Then inject 10 μL for LC-MS analysis.
[0194] 3.5 Perform secondary mass spectrometry (LS-MSMS) analysis on the samples. Take 1 nmol of the sample of the modified sequence to be tested (the sample that has not been treated with enzyme), put it into the inner tube, add enzyme-free water to make up to 100 μL, and inject 10 μL for LS_MSMS analysis.
[0195] 4. Experimental Results and Analysis 4.1 m_seq2_1 First-order mass spectrometry deconvolution The first-level mass spectrometry deconvolution result of the m_seq_1 modified sequence is as follows: Figure 2 As shown, the molecular weight was measured to be 8630.254. The number of methoxy sites was calculated as (8630.254-8532.1775) / 14.02=6.995=7 methoxy modification sites, where 8532.1775 is the theoretical molecular weight of the m_seq_1 bare sequence and 14.02 is the molecular weight of the methoxy modification.
[0196] 4.2 RNase T1 enzyme digestion results data The primary mass spectrometry (LC-MS) UV spectrum of the product obtained by digesting the m_seq_1 modified sequence with RNase T1 is shown below. Figure 3 As shown, further deconvolution yields... Figure 4 The mass spectrometry results shown indicate the measured molecular weights of the two main products as 2616.355 Da and 5686.856 Da, respectively.
[0197] The measured molecular weight of each product is matched with the theoretical enzyme digestion fragment, and different numbers of methoxy modification sites are assigned to the theoretical fragment. When the measured molecular weight of the product matches the molecular weight of the theoretical fragment with a specific number of methoxy modification sites added, it can be determined that the theoretical fragment contains that specific number of methoxy modification sites.
[0198] Matching results are as follows Figure 5 As shown. Figure 5 The left-hand table shows the theoretical RNase T1 enzyme digestion results obtained from the naked sequence and the theoretical molecular weight of the corresponding theoretical digested fragments. The right-hand table shows the measured molecular weight obtained from the digestion products, the matching fragment combinations, and the inferred number of methoxy groups in the fragments. For example, the theoretical digested fragment combination corresponding to the measured molecular weight of 2616.355 in the first row of the right-hand table should be theoretical fragment 1 + theoretical fragment 2, i.e., 1246.238 (theoretical fragment 1) + 1200.1836 (theoretical fragment 2) + 61.95 (phosphate group connecting the two fragments) + 79.95 (tri-terminal phosphate modification after digestion) + 28.04 (two methoxy groups) = 2616.3616.
[0199] The matching process starts from the 5' end of the naked sequence and proceeds towards the 3' end. When the matched result is very close to the measured molecular weight, the number of methoxy groups is increased. If the molecular weight can be matched when an integer number of methoxy groups is given, then the enzyme digestion fragment is considered to have the corresponding number of methoxy groups. The number of methoxy groups given here needs to be less than the total number of methoxy groups. The corresponding molecular weight can be obtained using relevant software or calculated manually.
[0200] Data analysis: The following results can be obtained from the RNase T1 digestion data: The G between theoretical fragment 1 and theoretical fragment 2 is modified with ribose and / or base, i.e., modified with methoxy group; Theoretical fragment 3 has no ribose and / or base modification in G; The G-ribose between theoretical fragment 4 and theoretical fragment 5 is modified with a methoxy group.
[0201] Based on the above analysis, we can tentatively derive the following sequence: {AACmGUUCG}G {CCACUUAAACmGACCAUCU} (SEQ ID NO: 3), and the fragments inside the corresponding curly braces have 2 and 5 methoxy groups respectively.
[0202] 4.3 RNase A enzyme digestion data analysis RNase A was added to the RNase T1 digestion product obtained in the previous step for further digestion. The TIC spectrum of the obtained product is shown below. Figure 6 As shown, the mass spectrometry results are extracted from it. Figure 7As shown, the extracted mass spectrometry signals are 625.06 (z=1); 641.10 (z=1); 643.07 (z=1); 648.08 (z=1); 651.09 (z=1); 657.09 (z=1); 666.09 (z=1); 980.15 (z=1); 1002.13 (z=1); 664.10 (z=2); 1157.68 (z=2). Mass spectrometry was performed in negative mode, resulting in the loss of protons. Therefore, molecular weight = mass-to-charge ratio × number of protons lost + number of protons. The measured molecular weights were 626.1 Da; 642.1 Da; 644.1 Da; 649.1 Da; 652.1 Da; 658.1 Da; 667.1 Da; 981.1 Da; 1003.1 Da; 1330.2 Da; and 2317.3 Da. Statistical results are shown below. Figure 8 .
[0203] like Figure 8 As shown, the measured molecular weights of the products after co-digestion by RNase T1 and RNase A were statistically summarized. Due to the slightly lower nonspecificity of enzyme A, there will be nonspecific cleavage products. That is, the digestion products include both specifically cleaved products and some nonspecific cleavage products. Therefore, the resulting result is greater than the actual expected result. In the subsequent data analysis, the expected digestion fragments will be matched with the measured digestion fragments. After matching, the specific modifications on the expected digestion fragments can be obtained.
[0204] 4.4 Data Summary and Analysis The sequence information AACmGUUCGGCCACUAAACmGACCAUCU (SEQ ID NO: 4) obtained in section 4.2 was subjected to theoretical RNase A specific digestion to obtain the theoretical digested fragment as follows: Figure 9 As shown.
[0205] Figure 9This is the theoretical result of specific RNase A enzyme digestion. Since RNase T1 enzyme has already digested the unmodified G, RNase A enzyme then digests the C / U sites in the sequence. Theoretically, this will form several expected fragments of dinucleotides or polynucleotides ending with an A base or a known modified G at the 5' end and C / U / G (unmodified G) at the 3' end. The theoretical molecular weights of these fragments are labeled. When the theoretical molecular weight corresponds exactly to the measured molecular weight, the theoretical fragment is the actual fragment. If there is no direct correspondence, it indicates an unknown modification on a certain ribose. An integer number of methoxy groups needs to be added. When adding a specific number of methoxy groups results in a perfect correspondence with the measured molecular weight, the actual fragment contains that specific number of methoxy groups on the theoretical fragment. Note: When modifying ribose and / or bases, adding a 3' terminal base requires adding another nucleotide to the expected fragment.
[0206] According to the above rules, we can obtain the following: Figure 10 The corresponding results are shown.
[0207] like Figure 10 As shown, the directly matched fragments AAC and AC have no modifications.
[0208] Since the G in mGU already has a methoxy group, after giving U a modification, the molecular weight of mGmUU will completely correspond to 1003.1.
[0209] Neither AAAC nor mGAC has a corresponding measured molecular weight. When no more than the remaining number of methoxy groups are given, it is also impossible to obtain a corresponding measured molecular weight. Therefore, it is speculated that AAAC and mGAC are linked in the measured fragment. Thus, if the C in AAAC is modified, AAAmCmGAC is obtained, and the molecular weight corresponds exactly to 2317.3.
[0210] When AU is modified with a methoxy group to become mAU, it can completely correspond to 667.1.
[0211] Based on the information confirmed above, the current sequence can be easily obtained as {AACmGmUUCG} G{CCACUUAAAmCmGACCmAUCU} (SEQ ID NO: 5). The bolded parts are all identified (including those identified as modified and / or unmodified). Except for the five marked methoxy groups, the remaining sites are unmodified. Therefore, the only possible combinations of the remaining modification sites are mCC / mUU / mCU / CmU. The molecular weight of mCC is 642.1, which matches the measured result. The molecular weight of mUU is 644.1, which matches the measured result perfectly. mCU and CmU are the 3' ends of this sequence, so there is no phosphate modification at the 3' end. Therefore, its molecular weight is 563.1, which does not match the measured molecular weight.
[0212] The final identification result is: AACmGmUUCGGmCCACmUUAAAmCmGACCmAUCU (SEQ ID NO: 2).
[0213] 4.5. Confirmation of m_seq_2 LC_MSMS results The sample containing the modified sequence to be tested was validated by secondary mass spectrometry, and the results are as follows: Figure 11 As shown, the sequences derived from 4.2 to 4.4 are completely correct after verification by secondary mass spectrometry, indicating that the modification sites obtained in the previous analysis are completely correct.
[0214] Example 2 The experimental design and the instruments and reagents used in this embodiment are the same as those in Example 1.
[0215] 1. Sample Information
[0216] Note: f represents fluorination modification. 2. Experimental Procedure 2.1 Preparation of LC MS mobile phase LC MS buffer A: Add 10 mL HFIPA + 500 μL DIEA to 1 L of water, mix well and sonicate to obtain mobile phase A.
[0217] LC MS buffer B: Add 10 mL HFIPA + 500 μL DIEA + 200 mL water to 800 mL acetonitrile, mix well and sonicate to obtain mobile phase B.
[0218] 2.2 Perform primary mass spectrometry (LC-MS) analysis on the samples. Take 1 nmol of the sample containing the modified sequence to be tested, put it into the insertion tube, add enzyme-free water to make up to 100 μL, and inject 10 μL for LS_MS analysis.
[0219] 2.3 RNase T1 enzyme digestion mass spectrometry analysis Take 2 nmol of the sample to be tested and place it in a PCR tube. Add 5 μL of buffer (200 mM Tris-HCl (pH 7.8), 400 mM KCl, 80 mM MgCl2, 10 mM DTT) and 1 μL of T1 enzyme. Add enzyme-free water to make up to 50 μL. Incubate at room temperature for 10 minutes, and then inject 10 μL for LC-MS analysis.
[0220] 2.4 RNase A enzyme digestion mass spectrometry analysis Add 1 μL of enzyme A (100,000 U / μl diluted 1000 times) to the RNase T1 enzyme digestion product and incubate at room temperature for 10 minutes. Then inject 10 μL for LC-MS analysis.
[0221] 2.5 Perform secondary mass spectrometry (LS-MSMS) analysis on the samples. Take 1 nmol of the sample containing the modified sequence to be tested, insert it into the insertion tube, add enzyme-free water to make up to 100 μL, and inject 10 μL for LS_MSMS analysis.
[0222] 3. Experimental Results and Analysis 3.1 f_seq2_1 First-order mass spectrometry deconvolution The results of the first-level mass spectrometry deconvolution of the f_seq_1 modified sequence are as follows: Figure 12 As shown, the molecular weight was measured to be 8546.114. The number of fluorinated modification sites was calculated as (8546.114-8532.1775) / 1.996=6.984=7 fluorinated modification sites, where 8532.1775 is the theoretical molecular weight of the f_seq_1 bare sequence and 1.996 is the molecular weight of the fluorinated modification.
[0223] 3.2 RNase T1 enzyme digestion results data The primary mass spectrometry (LC-MS) UV spectrum of the product obtained by digestion of the f_seq_1 modified sequence with RNase T1 is shown below. Figure 13 As shown, further deconvolution yields... Figure 14 The mass spectrometry results shown indicate the measured molecular weights of the two main products, which are 2592.315 and 5626.755, respectively.
[0224] The measured molecular weight of each product is matched with the theoretical enzyme digestion fragment, and different numbers of fluorinated modification sites are assigned to the theoretical fragment. When the measured molecular weight of the product matches the molecular weight of the theoretical fragment with a specific number of fluorinated modification sites, it can be determined that the theoretical fragment contains that specific number of fluorinated modification sites.
[0225] Matching results are as follows Figure 15 As shown. Figure 15The left-hand table shows the theoretical RNase T1 enzyme digestion results obtained from the naked sequence and the theoretical molecular weight of the corresponding theoretical digested fragments. The right-hand table shows the measured molecular weight obtained from the digestion products, the matching fragment combinations, and the number of fluorinated modifications in the inferred fragments. For example, the theoretical digested fragment combination corresponding to the measured molecular weight of 2592.315 in the first row of the right-hand table should be theoretical fragment 1 + theoretical fragment 2, i.e., 1246.238 (theoretical fragment 1) + 1200.1836 (theoretical fragment 2) + 61.95 (phosphate linking the two fragments) + 79.95 (tri-terminal phosphate modification after digestion) + 3.992 (two fluorinated modifications) = 2592.3136, which corresponds perfectly to the measured result of 2592.315.
[0226] The matching process starts from the 5' end of the naked sequence and proceeds towards the 3' end. When the matched result is very close to the measured molecular weight, the number of fluorinated molecules is increased. This ensures that when an integer number of fluorinated molecules is added, the molecular weight can be matched. In this case, the enzyme fragment is considered to have the corresponding number of fluorinated molecules. The number of fluorinated molecules added here must be less than the total number of fluorinated molecules. Data analysis: The following results can be obtained from the RNase T1 digestion data: The G between theoretical fragment 1 and theoretical fragment 2 is modified with ribose and / or bases, i.e., fluorinated. Theoretical fragment 3 has no ribose and / or base modification in G; The G-ribose between theoretical fragment 4 and theoretical fragment 5 is modified, namely fluorinated.
[0227] Based on the above analysis, we can tentatively derive the following sequence: {AACfGUUCG}G {CCACUUAAACfGACCAUCU} (SEQ ID NO: 8), and the segments inside the corresponding curly braces have 2 and 5 fluorinated modifications, respectively.
[0228] 3.3 RNase A enzyme digestion data analysis RNase A was added to the RNase T1 digestion product obtained in the previous steps for further digestion. Mass spectrometry results were extracted from the T1 chromatogram of the product, as shown below. Figure 16As shown, the mass spectrum signals are 613.03(z=1);629.08(z=1);631.05(z=1);636.06(z=1);651.09( z=1);654.07(z=1);657.09(z=1);960.08(z=1);978.09(z=1);980.14(z=1); 980.14(z=1); 986.14(z=1); 652.08(z=2); 984.13(z=2); 1145.66(z=2). Mass spectrometry was performed in negative mode, resulting in the loss of protons. Therefore, molecular weight = mass-to-charge ratio × number of protons lost + number of protons. The measured molecular weights were 614.0 Da; 630.1 Da; 632.1 Da; 637.1 Da; 652.1 Da; 655.1 Da; 658.1 Da; 961.1 Da; 979.1 Da; 981.1 Da; 987.1 Da; 1306.2 Da; 1970.3 Da; and 2293.3 Da. Statistical results are shown below. Figure 17 .
[0229] like Figure 17 As shown, the measured molecular weights of the products after co-digestion by RNase T1 and RNase A were statistically summarized. Due to the slightly lower nonspecificity of RNase A, there will be nonspecific cleavage products. That is, the digestion products include both specifically cleaved products and some nonspecific cleavage products. Therefore, the results are greater than the actual expected results. In the subsequent data analysis, the expected digestion fragments will be matched with the measured digestion fragments. After matching, the specific modifications on the expected digestion fragments can be obtained.
[0230] 3.4 Data Summary and Analysis Based on the sequence information AACfGUUCGGCCACUAAACfGACCAUCU (SEQ ID NO: 9) obtained in section 3.2, further RNase A-specific digestion was performed to obtain the theoretical digested fragment as follows: Figure 18 As shown.
[0231] Figure 18This is the theoretical result of specific RNase A enzyme digestion. Since RNase T1 enzyme has already digested the unmodified G, RNase A enzyme then digests the C / U sites in the sequence. Theoretically, this will form several expected fragments of dinucleotides or polynucleotides ending with an A base or a known modified G at the 5' end and C / U / G at the 3' end. Their theoretical molecular weights are labeled. When the theoretical molecular weight corresponds exactly to the measured molecular weight, the theoretical fragment is the actual fragment. If there is no direct correspondence, it indicates an unknown modification at a certain ribose. An integer number of fluorinated modifications needs to be applied. When the measured molecular weight corresponds exactly to the theoretical molecular weight after applying a specific number of fluorinated modifications, it means the actual fragment contains that specific number of fluorinated modifications on the theoretical fragment. Note: When modifying ribose and / or bases, if a 3' terminal base modification is applied, the expected fragment needs to be ligated with another nucleotide.
[0232] According to the above rules, we can obtain the following: Figure 19 The corresponding results are shown.
[0233] like Figure 19 As shown, the directly matched fragments AAC and AC have no modifications.
[0234] Since the G in fGU already has a fluorinated modification, after giving U a modification, the molecular weight of fGfUU will completely correspond to 979.1.
[0235] Neither AAAC nor fGAC has a corresponding measured molecular weight. Even when fluorination is performed on no more than the remaining number of molecules, a corresponding measured molecular weight cannot be obtained. Therefore, it is speculated that AAAC and fGAC are connected in the measured fragment. Thus, the C in AAAC is fluorinated, resulting in AAAfCfGAC, at which point the molecular weight corresponds exactly to 2293.3.
[0236] When AU is modified with a fluorine atom to become fAU, it can completely correspond to 655.1.
[0237] Based on the information confirmed above, the current sequence can be easily obtained as {AACfGfUUCG} G{CCACUUAAAfCfGACCfAUCU} (SEQ ID NO: 10). The parts shown in bold are all identified (including those identified as modified and / or those identified as unmodified). Except for the 5 fluorinated modifications that have been marked, the other sites are unmodified. Therefore, the combination caused by the remaining modification sites can only be fCC / fUU / fCU / CfU.
[0238] The molecular weight of fCC is 630.1, which matches the measured result. The molecular weight of mUU is 632.1, which matches the measured result perfectly. mCU and CmU are the 3' ends of this sequence, so there is no phosphate modification at the 3' end. Therefore, its molecular weight is 551.1, which does not match the measured molecular weight.
[0239] The final identification result is: AACfGfUUCGGfCCACfUUAAAfCfGACCfAUCU (SEQ ID NO: 7).
[0240] 3.5 Confirmation of f_seq_2 LC_MSMS results The sample containing the modified sequence to be tested was validated by secondary mass spectrometry, and the results are as follows: Figure 20 As shown, the sequences derived from 3.2-3.4 are completely correct after verification by secondary mass spectrometry, indicating that the modification sites obtained in the previous analysis are completely correct.
[0241] 4. Experimental Conclusions For any RNA sequence with ribose and / or base modifications (e.g., methoxy, fluorinated, locked nucleic acid modifications, or any combination thereof), if the modification sites are unknown, the number of possible sequences is astronomical. For example, in the above embodiment, a 27nt RNA sequence has 27 ribose groups, meaning there are 27 possible ribose and / or base modification sites. If there are 7 ribose and / or base modifications (e.g., methoxy and / or fluorinated modifications) in the sequence, the number of possible sequences would reach 27. 7Given the astronomical number of possibilities, even the increasingly popular LC-MSMS mass spectrometry sequencing method cannot achieve the purpose of identification, especially since LC-MSMS is merely a confirmatory (or falsification) characterization method. Therefore, there is currently no effective method for identifying such ribose and / or base modification sites. This paper employs multiple analytical steps to continuously narrow down the possible modification sites, ultimately identifying one or several possible sequences. Then, the verification or falsification function of LC-MSMS is used to achieve the purpose of identifying ribose and / or base modification sites. This scheme first uses the precise molecular weight of the primary mass spectrometer to compare with the precise molecular weight of the theoretical bare sequence to obtain the number of ribose modifications. Then, the sample sequence is digested with RNase T1 enzyme. T1 enzyme specifically cleaves G, forming a small fragment with G as the 3' end, such as AAAUCGUUUA (SEQ ID NO: 11) The enzyme will be digested into AAAUCGp and UUUA. Mass spectrometry will then provide the precise molecular weight of the digested fragments. This precise molecular weight will be matched with the theoretical precise molecular weight of the naked sequence digested fragment (this theoretical molecular weight can be modified by adding several ribose and / or base modifications, such as methoxy (14.02) or fluorinated (1.996). When a match is found, the number of added ribose and / or base modifications represents the number of ribose and / or base modifications in the fragment. Simultaneously, ribose and / or base modifications on G will resist enzyme digestion, thus confirming the presence or absence of ribose modifications on G in the sequence. This step confirms the distribution of ribose and / or base modifications and the presence or absence of ribose modifications on G in the sequence. The T1 enzyme digestion product will then be further digested with RNase A. RNase A specifically cleaves C / U, further dividing the sequence into smaller units. The theoretically expected fragments will be matched with the measured results (this theoretically expected fragment can be determined based on RNase A). Enzyme A specifically adds a particular ribose modification (e.g., methoxy and / or fluorinated modification). When a match is found, the added modification site becomes the actual modification site. Finally, by summarizing the above analytical data to ensure that the ribose and / or base modification sites meet all the above conditions, one confirmed possible modification sequence or several possible modification sequences will be obtained. The final step can use LC-MSMS analysis, utilizing its verification or falsification function for final identification, thereby achieving the purpose of identifying ribose and / or base modification sites in the RNA sequence. The principle of the verification or falsification function of LC-MSMS is consistent; here, the LC-MSMS analysis method specifically refers to Full MS / ddms2.
[0242] Full MS / ddms2, this method includes primary and secondary mass spectrometry scans. After primary mass spectrometry analysis, based on signal intensity, the 10 precursor ions with the highest signal intensity are selected using a quadrupole detector and sent to an HCD collision cell for fragmentation. Fragmentation generally occurs on the phosphate backbone. The fragmented daughter ions are then analyzed again to obtain the mass-to-charge ratio information, i.e., secondary mass spectrometry information. The theoretically formed daughter ion information for this fragment is calculated using BioPharma Finder and matched with the actually detected daughter ion information. If a corresponding daughter ion can be found at every fragmentation site, the sequence matching degree is 100%, indicating that the finally confirmed theoretical sequence is correct. This ultimately achieves the purpose of identifying ribose and / or base modification sites in RNA sequences.
[0243] The embodiments of the present invention are not limited to those described above. Without departing from the spirit and scope of the present invention, those skilled in the art can make various changes and improvements to the present invention in form and detail, and these are all considered to fall within the protection scope of the present invention.
Claims
1. A method for identifying ribose and / or base modification patterns in a test RNA sequence, wherein the naked sequence of the test RNA sequence is known. When the type and / or quantity of ribose and / or base modifications in the RNA sequence to be tested are unknown, the method includes the following steps: Step 1: Determine the type and / or quantity of ribose and / or base modifications in the RNA sequence to be tested based on the difference between the measured molecular weight of the RNA sequence to be tested and the theoretical molecular weight of the naked sequence of the RNA sequence to be tested. Step 2: Obtain the first measured molecular weight of the RNA sequence to be tested after digestion with the first ribonuclease. Match the first measured molecular weight with the first theoretical molecular weight of the first theoretical digestion fragment obtained by theoretical digestion of the naked sequence of the RNA sequence to be tested with the first ribonuclease. This determines the nucleic acid sequence and some ribose and / or base modification information of each digestion fragment obtained by specific digestion of the RNA sequence to be tested with the first ribonuclease. Optionally, the second measured molecular weight of the RNA sequence to be tested after double digestion with a first ribonuclease and a second ribonuclease is obtained. The second measured molecular weight and the second theoretical molecular weight of the theoretical digestion fragment are matched and analyzed. The second theoretical digestion fragment is obtained by double digestion with a first ribonuclease and a second ribonuclease on a naked sequence of the RNA to be tested with a known modification. Thus, the nucleic acid sequence and the modification information of the ribose and / or bases contained in each digestion fragment obtained by specific double digestion with a first ribonuclease and a second ribonuclease of the RNA sequence to be tested are determined. Step 3: Based on the information obtained in Steps 1 and 2, determine all possible ribose and / or base modification patterns in the RNA sequence to be tested; and Step 4: When there is only one possible combination of ribose and / or base modification patterns, the modification patterns of all ribose and / or bases in the RNA sequence to be tested are determined; when there is more than one possible ribose and / or base modification pattern, the RNA sequence to be tested is further subjected to secondary mass spectrometry detection to determine the modification patterns of all ribose and / or bases in the RNA sequence to be tested. When the type and / or quantity of ribose and / or base modifications in the RNA sequence to be tested are known, step one is not required in this method.
2. The method of claim 1, wherein the type, quantity, and / or location of ribose and / or base modifications in the RNA sequence to be tested are unknown, the method comprising the following steps: Step 1: Obtain the measured molecular weight of the RNA sequence to be tested. Based on the difference between the measured molecular weight and the theoretical molecular weight of the bare sequence of the RNA sequence to be tested, determine the type and quantity of ribose and / or base modifications in the RNA sequence to be tested, which is constraint condition one. Step 2: Obtain the molecular weight of the first actual digested fragment obtained by digestion of the RNA sequence to be tested with the first ribonuclease, and obtain the first measured molecular weight, wherein each first measured molecular weight represents a first actual digested fragment; using the naked sequence of the RNA sequence to be tested as the first theoretical sequence, digest it with the theoretical first ribonuclease to obtain the first theoretical digested fragment; based on the constraint condition 1, perform the first matching analysis on the first theoretical digested fragment and the digested fragment represented by the first measured molecular weight, thereby determining the nucleic acid sequence and partial ribose and / or base modification information of each digested fragment obtained by specific digestion of the RNA sequence to be tested with the first ribonuclease, which is constraint condition 2; The first matching analysis includes comparing the theoretical molecular weight of the first theoretical digestion fragment with the first measured molecular weight to determine the first theoretical digestion fragment that matches each digestion fragment represented by the first measured molecular weight, as well as the additional ribose and / or base modifications contained in each digestion fragment represented by the first measured molecular weight relative to the matching first theoretical digestion fragment. The partial ribose and / or base modification information determined by the first matching analysis includes the type and quantity of ribose and / or base modifications contained in each digestion fragment obtained by specific digestion of the test RNA sequence with a first ribonuclease, and also includes precise information on whether a specific ribonucleotide has ribose and / or base modifications. The specific ribonucleotide refers to the characteristic nucleotide of each first ribonuclease contained in the test RNA sequence, excluding the nucleotide at the 3' end of the test RNA sequence; and / or The molecular weight of the second actual digested fragment obtained by double digestion of the RNA sequence to be tested with the first and second ribonucleases is obtained, and the second measured molecular weight is obtained, wherein each second measured molecular weight represents a second actual digested fragment. The modified sequence formed by adding ribose and / or base modifications at specific positions determined in the first matching analysis to the bare sequence is the second theoretical sequence. This modified sequence is then double-digested with the theoretical first and second ribonucleases to obtain the second theoretical digested fragment. Based on constraints one and two, a second matching analysis is performed on the second theoretical digested fragment and the digested fragment represented by the second measured molecular weight. This determines the nucleic acid sequence and the types and quantities of ribose and / or base modifications contained in each digested fragment obtained by specific double digestion of the RNA sequence by the first and second ribonucleases, which constitutes constraint three. The second matching analysis includes comparing the theoretical molecular weight of the second theoretical digestion fragment with the second measured molecular weight to determine the second theoretical digestion fragment that matches each second measured molecular weight and the additional ribose and / or base modifications contained in each second measured molecular weight fragment relative to the matching second theoretical digestion fragment. The partial ribose and / or base modification information determined by the second matching analysis includes the types and quantities of ribose and / or base modifications contained in each digestion fragment obtained by specific double digestion of the RNA sequence to be tested by the first and second ribonucleases, and also includes the exact information on whether a specific ribonucleotide has ribose and / or base modifications. The specific ribonucleotide refers to the characteristic nucleotide of each first or second ribonuclease contained in the RNA sequence to be tested, excluding the nucleotide at the 3' end of the RNA sequence to be tested. Step 3: Determine all possible ribose and / or base modification patterns in the RNA sequence to be tested; and Step 4: When there is only one type of ribose and / or base modification pattern, the types, positions, and numbers of all ribose and / or base modification patterns in the RNA sequence to be tested are determined; when there is more than one type of ribose and / or base modification pattern, the RNA sequence to be tested is further subjected to secondary mass spectrometry detection to determine the types, positions, and numbers of all ribose and / or base modification patterns in the RNA sequence to be tested. The first and second ribonucleases are site-specific ribonucleases and are not the same.
3. The method according to claim 1 or 2, wherein one of the first and second ribonucleases is RNase T1, whose characteristic nucleotide is G, and the other is RNase A, whose characteristic nucleotides are C and U; preferably, the first ribonuclease is RNase T1, and the second ribonuclease is RNase A.
4. The method according to any one of claims 1-3, wherein the ribose and / or base modification is selected from one or more of methoxy modification, fluorination modification, methoxyethyl modification and locked nucleic acid modification.
5. The method according to any one of claims 1-4, wherein in step one, the type and quantity of ribose and / or base modifications in the RNA sequence to be tested are determined according to the following method: First, calculate the total molecular weight of all ribose and / or base modifications in the RNA sequence to be tested using the following formula: The total molecular weight of ribose and / or base modifications = molecular weight 实测 -Molecular weight 裸序列理论 Among them, "molecular weight" 实测 "Molecular weight" refers to the measured molecular weight of the obtained RNA sequence to be tested. 裸序列理论 "" refers to the theoretical molecular weight of the naked sequence of the RNA to be tested; Subsequently, the theoretical molecular weight of different combinations of ribose and / or base modifications is compared with the total molecular weight of ribose and / or base modifications in the RNA sequence to be tested, as calculated above. When the theoretical molecular weight of a specific combination of ribose and / or base modifications is equal to the total molecular weight of ribose and / or base modifications in the RNA sequence to be tested, the types and quantities of ribose and / or base modifications contained in the RNA sequence to be tested can be determined as the types and quantities of ribose and / or base modifications contained in that specific combination of ribose and / or base modifications.
6. The method according to any one of claims 1-5, wherein in step two, the matching of the theoretical digestion fragment with the digestion fragment represented by the measured molecular weight includes one or more of the following: (a) The molecular weight of a single theoretical digestion fragment is equal to the measured molecular weight; in this case, the digestion fragment represented by the measured molecular weight has the same sequence as the single theoretical digestion fragment and does not contain any additional ribose and / or base modifications relative to the single theoretical digestion fragment. (b) A single theoretical restriction fragment has a molecular weight smaller than a measured molecular weight, but contains a specific combination of ribose and / or base modifications such that the molecular weight of the modified fragment formed by adding the specific ribose and / or base modification combination to the single theoretical restriction fragment is equal to the measured molecular weight, provided that the addition of the specific ribose and / or base modification combination satisfies the established ribose and / or base modification constraints; in this case, the restriction fragment represented by the measured molecular weight is the modified fragment formed by adding the specific ribose and / or base modification combination to the theoretical restriction fragment, and when the restriction fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the restriction fragment represented by the measured molecular weight does not have ribose and / or base modifications; and (c) An extended theoretical restriction fragment, formed by covalently linking two or more consecutive theoretical restriction fragments, has a molecular weight smaller than that of a restriction fragment represented by a measured molecular weight. However, it contains a specific combination of ribose and / or base modifications such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modifications to the extended theoretical restriction fragment is equal to that of the measured molecular weight, provided that the addition of the specific combination of ribose and / or base modifications satisfies pre-defined ribose and / or base modification constraints. In this case, the restriction fragment represented by the measured molecular weight is the modified fragment formed by adding the specific combination of ribose and / or base modifications to the extended theoretical restriction fragment. Specifically, the ribonucleotide at the 3' end of the preceding theoretical restriction fragment in any two adjacent theoretical restriction fragments has a ribose and / or base modification, and when the restriction fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the restriction fragment represented by the measured molecular weight does not have a ribose and / or base modification.
7. The method according to any one of claims 1-6, wherein the matching analysis in step two includes performing a matching analysis on each theoretical enzyme digestion fragment, such that each theoretical enzyme digestion fragment is matched to an enzyme digestion fragment represented by a certain measured molecular weight.
8. The method of any one of claims 1-7, wherein the matching analysis in step two is performed by sequentially executing the following steps: S1: Select a theoretical enzyme digestion fragment and proceed with step S2; S2: Check whether the molecular weight of the theoretical enzyme digestion fragment is equal to a certain measured molecular weight, and make the judgment in step S3. S3: If the molecular weight of the theoretical enzyme digestion fragment is equal to a measured molecular weight, then the enzyme digestion fragment represented by the measured molecular weight is determined to have the same sequence as the theoretical enzyme digestion fragment and does not contain any additional ribose and / or base modifications relative to the theoretical enzyme digestion fragment. Then, check whether there are other theoretical enzyme digestion fragments that have not yet been matched. If so, select one of the unmatched theoretical enzyme digestion fragments and perform step S2 on it. If not, complete all matching. If the molecular weight of the theoretical enzyme digestion fragment is not equal to any of the measured molecular weights, then perform step S4 on the theoretical enzyme digestion fragment. S4: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the theoretical enzyme digestion fragment is equal to a certain measured molecular weight (i.e., the sum of the molecular weight of the theoretical enzyme digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to the measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S5. S5: If such a combination of ribose and / or base modifications exists, the enzyme fragment represented by the measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the theoretical enzyme fragment. When the enzyme fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the enzyme fragment represented by the measured molecular weight does not have ribose and / or base modifications. Then, it is checked whether there are other theoretical enzyme fragments that have not been matched. If so, one unmatched theoretical enzyme fragment is selected and subjected to step S2. If not, all matching is completed. If such a combination of ribose and / or base modifications does not exist, a theoretical enzyme fragment adjacent to it before and / or after it is added to the theoretical enzyme fragment to form an extended theoretical enzyme fragment and subjected to step S6. S6: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the extended theoretical digestion fragment is equal to a certain measured molecular weight (i.e., the sum of the molecular weight of the extended theoretical digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to a certain measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S7. S7: If such a combination of ribose and / or base modifications exists, the enzyme digestion fragment represented by the measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the extended theoretical enzyme digestion fragment. In any two adjacent theoretical enzyme digestion fragments, the ribonucleotide at the 3' end of the preceding theoretical enzyme digestion fragment has ribose and / or base modifications. When the enzyme digestion fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the enzyme digestion fragment represented by the measured molecular weight does not have ribose and / or base modifications. Then, it is checked whether there are other theoretical enzyme digestion fragments that have not yet been matched. If so, one unmatched theoretical enzyme digestion fragment is selected and subjected to step S2. If not, all matching is completed. If such a combination of ribose and / or base modifications does not exist, the extended theoretical enzyme digestion fragment is added to a theoretical enzyme digestion fragment adjacent to it before and / or after it to form a further extended theoretical enzyme digestion fragment, which is then subjected to step S6.
9. The method according to any one of claims 1-8, wherein the matching analysis in step two includes arranging all theoretical enzyme digestion fragments sequentially in a 5' to 3' direction or in a 3' to 5' direction, and performing matching analysis on each theoretical enzyme digestion fragment sequentially, starting from the theoretical enzyme digestion fragment that is ranked first in sequence.
10. The method of any one of claims 1-6, wherein the matching analysis in step two comprises performing a matching analysis on each theoretical digestion fragment having two or more nucleotides, such that each theoretical digestion fragment having two or more nucleotides is matched to a digestion fragment represented by a measured molecular weight.
11. The method of any one of claims 1-6 and 10, wherein the matching analysis in step two is performed by sequentially executing the following steps: S1: Select a theoretically digestible fragment containing two or more nucleotides and proceed to step S2; S2: Check whether the molecular weight of the theoretical enzyme digestion fragment is equal to a certain measured molecular weight, and make the judgment in step S3. S3: If the molecular weight of the theoretical digestion fragment is equal to a measured molecular weight, then the digestion fragment represented by the measured molecular weight is determined to have the same sequence as the theoretical digestion fragment and does not contain any additional ribose and / or base modifications relative to the theoretical digestion fragment. Then, it is checked whether there are other theoretical digestion fragments containing two or more nucleotides that have not yet been matched. If so, one of the unmatched theoretical digestion fragments containing two or more nucleotides is selected and subjected to step S2. If not, all matching is completed. If the molecular weight of the theoretical digestion fragment is not equal to any of the measured molecular weights, then step S4 is performed on the theoretical digestion fragment. S4: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the theoretical enzyme digestion fragment is equal to a certain measured molecular weight (i.e., the sum of the molecular weight of the theoretical enzyme digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to the measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S5. S5: If such a combination of ribose and / or base modifications exists, the enzyme fragment represented by the measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the theoretical enzyme fragment. When the enzyme fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the ribonucleotide at the 3' end of the enzyme fragment represented by the measured molecular weight does not have ribose and / or base modifications. Then, it is checked whether there are other theoretical enzyme fragments containing two or more nucleotides that have not yet been matched. If so, one unmatched theoretical enzyme fragment containing two or more nucleotides is selected and subjected to step S2. If not, all matching is completed. If such a combination of ribose and / or base modifications does not exist, a theoretical enzyme fragment adjacent to it before and / or after it is added to the theoretical enzyme fragment to form an extended theoretical enzyme fragment and subjected to step S6. S6: Check whether there is a specific combination of ribose and / or base modification such that the molecular weight of the modified fragment formed by adding the specific combination of ribose and / or base modification to the extended theoretical digestion fragment is equal to a certain measured molecular weight (i.e., the sum of the molecular weight of the extended theoretical digestion fragment and the theoretical molecular weight of the specific combination of ribose and / or base modification is equal to a certain measured molecular weight), provided that the addition of the specific combination of ribose and / or base modification satisfies the determined ribose and / or base modification constraints, and perform the judgment in step S7. S7: If such a combination of ribose and / or base modifications exists, then the enzyme fragment represented by the measured molecular weight is determined to be the modified fragment formed by adding the specific combination of ribose and / or base modifications to the extended theoretical enzyme fragment, wherein the ribonucleotide at the 3' end of the preceding theoretical enzyme fragment in any two adjacent individual theoretical enzyme fragments has ribose and / or base modifications, and when the enzyme fragment represented by the measured molecular weight is not located at the 3' end of the RNA sequence to be tested, the 3' end of the enzyme fragment represented by the measured molecular weight... The ribonucleotide does not have ribose and / or base modifications. Then, it is checked whether there are other theoretical cleavage fragments containing two or more nucleotides that have not been matched. If they exist, one of the unmatched theoretical cleavage fragments containing two or more nucleotides is selected and subjected to step S2. If they do not exist, all matching is completed. If there is no such combination of ribose and / or base modifications, the extended theoretical cleavage fragment is added to a theoretical cleavage fragment adjacent to it before and / or after it to form a further extended theoretical cleavage fragment and subjected to step S6.
12. The method according to any one of claims 1-6 and 10-11, wherein the matching analysis in step two comprises arranging all theoretical digestion fragments having two or more nucleotides sequentially in a 5' to 3' direction or in a 3' to 5' direction, and performing matching analysis sequentially on each theoretical digestion fragment having two or more nucleotides, starting from the theoretical digestion fragment having two or more nucleotides that is sequentially arranged in the first position.
13. The method according to any one of claims 1-12, wherein the first measured molecular weight and the second measured molecular weight are determined by mass spectrometry.
14. A computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the method of any one of claims 1-12.
15. A computer program product for storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1-12.
16. A computer device comprising a processor and a memory, the memory storing a computer program configured to be executed by the processor, the computer program, when executed by the processor, implementing the steps of the method according to any one of claims 1-12.