A method for the identification of modifications of a phosphorothioated modified nucleic acid sequence

By performing enzymatic digestion and secondary mass spectrometry analysis on nucleic acid sequences, combined with preset modification rules, the problem that LC-MS methods cannot identify thiophosphorylated modified nucleic acid sequences has been solved, achieving accurate and efficient modification identification, and simultaneously identifying pentose sugar and base modifications.

CN116973467BActive Publication Date: 2026-03-27NANJING GENSCRIPT BIOTECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-28
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing LC-MS methods cannot accurately identify the nucleotide sequence of nucleic acid sequences with thiophosphorylation modifications, leading to inaccurate identification results.

Method used

The nucleic acid sequence to be identified is digested using a mixed enzyme to generate digestion products. The target sequence fragment in the digestion products is determined to be consistent with the theoretical sequence fragment by primary and secondary mass spectrometry analysis combined with preset modification rules.

Benefits of technology

It enables accurate identification of thiophosphorylated modified nucleic acid sequences, and can simultaneously identify other chemical modifications of pentose sugars and bases, improving the accuracy and efficiency of identification results and eliminating the purification step of enzymatic digestion products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116973467B_ABST
    Figure CN116973467B_ABST
Patent Text Reader

Abstract

The embodiment of the specification provides a modification identification method of a thio-phosphorylated nucleic acid sequence, which comprises: performing enzymolysis on a to-be-identified nucleic acid sequence by using a nucleic acid mixed enzyme to obtain an enzymolysis product; performing secondary mass spectrum analysis on the enzymolysis product based on first mass spectrum information; and determining whether a target sequence fragment exists in the to-be-identified nucleic acid sequence based on the first mass spectrum information and the second mass spectrum information. At least the to-be-identified nucleic acid sequence modified by thio-phosphorylation is actually obtained based on a preset modification rule. The first mass spectrum information comprises a theoretical analysis result generated by analyzing a theoretical sequence fragment, and the theoretical sequence fragment is determined based on the preset modification rule, the to-be-identified nucleic acid sequence and the nucleic acid mixed enzyme. The modification identification method can accurately and efficiently identify the chemical modification of the thio-phosphorylated nucleic acid sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing

[0002] This application claims priority to Chinese application 202210467473.7, filed on April 29, 2022, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This specification relates to the field of nucleic acid detection technology, and in particular to a method for the modification and identification of nucleic acid sequences modified by thiophosphorylation. Background Technology

[0004] Nucleic acid sequence modification identification methods are mainly used to determine whether chemical modifications of nucleic acid sequences conform to the intended design, and they have wide applications in biological, medical, and pharmaceutical research and production activities. For example, the CRISPR / Cas9 system, widely used for gene editing, consists of the Cas9 protein and sgRNA (single guide RNA). sgRNA is designed based on the higher-order structure formed by crRNA and tracrRNA, and binds to the Cas9 nuclease protein, guiding it to recognize and cut target sequences. Because CRISPR / Cas9 originates from the prokaryotic acquired immune system's defense system against foreign genetic material, the Cas9 nuclease may inherit the characteristic of low sequence specificity, increasing the probability of non-specific cleavage and resulting in more off-target effects. Therefore, chemical modification or editing of Cas9 and sgRNA is usually required to reduce off-target effects. As a key raw material for CRISPR gene editing technology, sgRNA requires quality studies during clinical application, specifically the identification of modifications at the 5' and 3' ends of the sgRNA sequence. For sgRNAs with thiophosphorylation modifications at both the 5' and 3' ends, direct LC-MS analysis of the full-length sequence is inaccurate because it fails to reflect the nucleotide sequence of the chemically modified nucleic acid. This makes it impossible to determine whether the chemical modification is based on the correct sequence arrangement. Therefore, it is necessary to provide an accurate and efficient method for identifying thiophosphorylated nucleic acid sequences. Summary of the Invention

[0005] The one or more embodiments of the specification provide a modification identification method of a phosphorothioated nucleic acid sequence. The method comprises: performing enzymolysis on a nucleic acid sequence to be identified by using a nucleic acid mixed enzyme to obtain an enzymolysis product; wherein the nucleic acid sequence to be identified is actually obtained based on a preset modification rule, and the modification to be identified of the nucleic acid sequence to be identified at least includes phosphorothioation modification of an inter-nucleotide linkage; performing secondary mass spectrometry analysis on the enzymolysis product based on first mass spectrometry information to obtain second mass spectrometry information actually generated by the enzymolysis product, wherein the first mass spectrometry information includes a theoretical analysis result generated by analyzing a theoretical sequence fragment, and the theoretical sequence fragment is determined based on the preset modification rule, the nucleic acid sequence to be identified, and the nucleic acid mixed enzyme; and determining whether a target sequence fragment consistent with the theoretical sequence fragment exists in the nucleic acid sequence to be identified based on the first mass spectrometry information and the second mass spectrometry information.

[0006] In some embodiments, the nucleic acid mixed enzyme is capable of breaking a phosphodiester bond and exempt from breaking a phosphorothioate bond, so as to make the nucleic acid sequence to be identified with phosphorothioation modification generate at least two sequence fragments with different molecular weights.

[0007] In some embodiments, the nucleic acid mixed enzyme comprises a snake venom phosphodiesterase and / or a bovine spleen phosphodiesterase.

[0008] In some embodiments, the enzymolysis time of the nucleic acid mixed enzyme is 1 h to 8 h.

[0009] In some embodiments, the enzymolysis time of the nucleic acid mixed enzyme is 3 h to 5 h.

[0010] In some embodiments, the modification to be identified of the nucleic acid sequence to be identified further includes modification of a pentose and / or modification of a base.

[0011] In some embodiments, the modification of the pentose includes at least one of the following modifications: 2'-modification, 5'-modification, 3'-modification, locked nucleic acid modification, unlocked nucleic acid modification, and peptide nucleic acid modification.

[0012] In some embodiments, the 2'-modification of the pentose includes at least one of the following modifications: 2'-O-methyl modification, 2'-fluoro modification, 2'-O-methoxyethyl modification, 2'-O-methylation modification, and 2'-O-allyl modification.

[0013] In some embodiments, the modification of the base includes at least one of the following modifications: methylation modification and hydroxymethylation modification.

[0014] In some embodiments, the preset modification rule comprises that the nucleotide corresponding to the modified pentose is connected with at least one adjacent nucleotide by a phosphorothioate bond; and / or, the preset modification rule comprises that the nucleotide corresponding to the modified base is connected with at least one adjacent nucleotide by a phosphorothioate bond.

[0015] In some embodiments, the preset modification rule comprises that the 5' end and / or 3' end of the to-be-identified nucleic acid sequence has a modified sequence fragment, the length of the modified sequence fragment is 2-5 nucleotides, and adjacent nucleotides in the modified sequence fragment are connected by a phosphorothioate bond.

[0016] In some embodiments, adjacent nucleotides in the theoretical sequence fragment are connected by a phosphorothioate bond.

[0017] In some embodiments, the first mass spectrum information comprises first parent ion information and first daughter ion information of the theoretical sequence fragment; wherein the first parent ion information at least comprises the mass-to-charge ratio of the first parent ion, and the first daughter ion information at least comprises the complementary pairing information and mass-to-charge ratio of the first daughter ion.

[0018] In some embodiments, the second mass spectrum information comprises second parent ion information and second daughter ion information; the obtaining of the second mass spectrum information actually produced by the enzymatic product comprises: determining the second parent ion information based on the first parent ion information, the second parent ion information at least comprising the mass-to-charge ratio of the second parent ion; performing secondary mass spectrum analysis of the enzymatic product based on the second parent ion information to obtain second daughter ion information, the second daughter ion information at least comprising the complementary pairing information and mass-to-charge ratio of the second daughter ion.

[0019] In some embodiments, the determination of whether the target sequence fragment consistent with the theoretical sequence fragment exists in the to-be-identified nucleic acid sequence comprises: determining whether the first daughter ion information and the second daughter ion information satisfy a preset matching condition; if the preset matching condition is satisfied, it is determined that the target sequence fragment exists in the to-be-identified nucleic acid sequence.

[0020] In some embodiments, the preset matching condition is that there is at least one pair of first daughter ions and at least one pair of second daughter ions, and the complementary pairing information and mass-to-charge ratio of the at least one pair of first daughter ions and the at least one pair of second daughter ions are the same.

[0021] In some embodiments, the secondary mass spectrum analysis of the enzymatic product is performed by a linear ion trap mass spectrometer. BRIEF DESCRIPTION OF DRAWINGS

[0022] The present specification will be further demonstrated in the way of exemplary embodiments, which will be described in detail by the accompanying drawings. These embodiments are not restrictive, in which the same numbers represent the same structures, wherein:

[0023] Figure 1 is an exemplary flow chart of the modification identification method shown according to some embodiments of the present specification;

[0024] Figure 2 is the LC-MS mass spectrum chart of control sample 1 of embodiment 1 of the present specification;

[0025] Figure 3 is the LC-MS mass spectrum chart of control sample 2 of embodiment 1 of the present specification;

[0026] Figure 4 is the LC-MS total ion current chart of mixed enzyme reaction buffer of embodiment 1 of the present specification;

[0027] Figure 5 is the LC-MS mass spectrum chart of mixed enzyme reaction buffer of embodiment 1 of the present specification;

[0028] Figure 6 is the LC-MS mass spectrum chart of the enzymatic product of sample 1 enzymolysis for 2 min of embodiment 2 of the present specification;

[0029] Figure 7 is the LC-MS mass spectrum chart of the enzymatic product of sample 1 enzymolysis for 4 h of embodiment 2 of the present specification;

[0030] Figure 8 is the LC-MS mass spectrum chart of the enzymatic product of sample 1 enzymolysis for 24 h of embodiment 2 of the present specification;

[0031] Figure 9 is the LC-MS MS mass spectrum chart of the 5’ end modified fragment mG*mA*mG*rU corresponding to the enzymatic product of sample 1 of embodiment 3 of the present specification;

[0032] Figure 10 is the LC-MS MS mass spectrum chart of the 3’ end modified fragment mU*mU*mU*rU corresponding to the enzymatic product of sample 1 of embodiment 3 of the present specification;

[0033] Figure 11 is the LC-MS MS mass spectrum chart of the 5’ end modified fragment mA*mU*mC*rA corresponding to the enzymatic product of sample 2 of embodiment 3 of the present specification;

[0034] Figure 12 is the LC-MS MS mass spectrum chart of the 3’ end modified fragment mU*mU*mU*rU corresponding to the enzymatic product of sample 2 of embodiment 3 of the present specification;

[0035] wherein, Figures 2-3 and Figures 5-8 wherein the abscissa is the mass-to-charge ratio m / z and the ordinate is the absolute intensity of the ions; Figure 4 wherein the abscissa is the time Time and the ordinate is the total ion current intensity; Figures 10-12 wherein the abscissa is the mass-to-charge ratio m / z and the ordinate is the absolute intensity of the ions. DETAILED DESCRIPTION

[0036] As used in the specification and claims, the words "a," "an," and / or "the" do not exclude the plural. Generally, the term "includes" or "comprising" means "including, but not limited to."

[0037] The following are definitions of some terms used in this application.

[0038] As used in the specification and claims, the term "mass spectrometry" refers to an analytical technique that ionizes chemical entities and sorts them according to their mass-to-charge ratio. In a first mass spectrometry, the process typically includes ionization of a chemical entity to produce charged ions (ionization process) and measurement of the mass-to-charge ratio of the charged ions. In a second mass spectrometry, the process typically includes ionization of a chemical entity to produce charged ions (ionization process), collision-induced dissociation of a selected parent ion to produce daughter ions (dissociation process), and measurement of the mass-to-charge ratio of the daughter ions.

[0039] The term "parent ion" is an ion of interest produced by the ionization process in mass spectrometry.

[0040] The term "daughter ion" is a fragment ion produced by the dissociation process in mass spectrometry. In a second mass spectrometry, there is a clear relationship between the parent ion produced by the ionization process and the daughter ions produced by the dissociation process and their measured ion spectral peaks.

[0041] The embodiments of the present specification relate to a modification identification method of a thio-phosphorylated modified nucleic acid sequence. In some embodiments, the modification identification method can be used to identify the chemical modification of a thio-phosphorylated modified sgRNA sequence. For example, the 3' end and 5' end of the sgRNA sequence are modified by thio-phosphorylation and methoxy modification, and the modification identification method can identify the thio-phosphorylation modification and the methoxy modification of the sgRNA sequence at the same time. In some embodiments, the modification identification method can be used to identify the chemical modification of an antisense oligonucleotide (ASO) which is thio-phosphorylated. For example, the ASO therapeutic agent is modified by various chemical modifications, including phosphorothioate (PS) backbone modification and modification of pentose and base, to improve the pharmacological properties, so that the modified ASO therapeutic agent shows better binding affinity to the target RNA and enhances the binding to the protein. The modification identification method can identify the phosphorothioate (PS) backbone modification and the modification of pentose and base of the ASO sequence at the same time. In some embodiments, the modification identification method can be used to identify the chemical modification of other thio-phosphorylated modified nucleic acid sequences, such as siRNA. The identification of one or more chemical modifications of the thio-phosphorylated modified nucleic acid sequence can be accurately and efficiently achieved by the modification identification method.

[0042] It should be understood that the application scenarios of the modification identification method of the thio-phosphorylated modified nucleic acid sequence of the present specification are only some examples or embodiments of the present specification, and for those skilled in the art, the present specification can also be applied to other similar scenarios without creative labor.

[0043] The following will be combined with Figure 1 The modification identification method of the thio-phosphorylated modified nucleic acid sequence related to the embodiments of the present specification will be described in detail. It is worth noting that the following embodiments are only used to explain the present specification and do not constitute a limitation on the present specification.

[0044] Figure 1 is an exemplary flowchart of the modification identification method according to some embodiments of the present specification. In some embodiments, the modification identification flow 100 at least includes steps 110 to 130.

[0045] In step 110, the nucleic acid mixed enzyme is used to enzymatically digest the nucleic acid sequence to be identified to obtain the enzymatic digestion product.

[0046] In some embodiments, the nucleic acid sequence to be identified is actually obtained based on a preset modification rule. Specifically, whether the one or more chemical modifications of the nucleic acid sequence to be identified is correct or not needs to be identified.

[0047] In some embodiments, the modification to be identified of the nucleic acid sequence to be identified comprises at least a phosphorothioate modification (PS) of the internucleotide linkage. For example, a phosphorothioate linkage is introduced between the last 3-5 nucleotides of the 3’ end and the 5’ end of the nucleic acid sequence to inhibit degradation of the nucleic acid sequence by exonucleases.

[0048] In some embodiments, the modification to be identified of the nucleic acid sequence to be identified further comprises a modification of the pentose sugar, e.g., a modification at position 2 of the pentose sugar of any one or more nucleotides of the nucleic acid sequence, a modification of the 3’ end of the backbone, or a modification of the 5’ end of the backbone. In some embodiments, the pentose sugar comprises ribose and deoxyribose.

[0049] In some embodiments, the modification of the pentose sugar comprises a 2’- modification, i.e., a modification at position 2 of the pentose sugar. In some embodiments, the 2’- modification comprises one or more of the following modifications: 2’-O-methyl modification (2’-OMe), 2’-fluoro modification (2’-F), 2’-O-methoxyethyl modification (2’-O-MOE), 2’-O-methylated modification (2’-O-methyl), 2’-O-allyl modification (2’-O-Allyl).

[0050] In some embodiments, the modification of the pentose sugar comprises a 5’- modification, i.e., a modification of the 5’ end of the backbone. In some embodiments, the 5’- modification comprises one or more of the following modifications: 5’-HEX, 5’-FAM, 5’-CY3, 5’-CY5, 5’-VIC, 5’-ROX, 5’ C6 amino modification (5’-Amino-C6), 5’ C12 amino modification (5’-Amino-C12).

[0051] In some embodiments, the modification of the pentose sugar comprises a 3’- modification, i.e., a modification of the 3’ end of the backbone. In some embodiments, the 3’- modification comprises one or more of the following modifications: 3’-HEX, 3’-FAM, 3’-CY3, 3’-CY5, 3’-VIC, 3’-ROX, 3’ C6 amino modification (3’-Amino-C6), 3’ C12 amino modification (3’-Amino-C12).

[0052] In some embodiments, the modification of the pentose sugar further comprises other modifications, e.g., a locked nucleic acid modification (LNA), an unlocked nucleic acid modification (UNA), a peptide nucleic acid modification (PNA), etc.

[0053] In some embodiments, the modification to be identified of the nucleic acid sequence to be identified further comprises a modification of the base. In some embodiments, the modification of the base comprises a methylation modification and / or a hydroxymethylation modification.

[0054] In some embodiments, the preset modification rule refers to a rule of chemically modifying the initial nucleic acid sequence according to the intended modification position and modification type. In some embodiments, the preset modification rule at least includes: at least one phosphorothioate bond between the nucleotides in the nucleic acid sequence to be identified. By virtue of the anti-enzyme cutting property of the phosphorothioate bond, the nucleic acid sequence to be identified can be cleaved to generate a target sequence fragment modified by phosphorothioation and a non-target sequence fragment not modified by phosphorothioation after treatment with a functional enzyme, and the differences in molecular weight and structure between the two fragments make the nucleotide arrangement and chemical modification of the target sequence fragment easy to identify.

[0055] In some embodiments, the modification to be identified of the nucleic acid sequence to be identified further includes modification of the pentose; and the preset modification rule further includes that the nucleotide corresponding to the modified pentose is connected to at least one adjacent nucleotide by a phosphorothioate bond. In the case that the nucleic acid sequence to be identified can be cleaved to generate a target sequence fragment modified by phosphorothioation, modification of the pentose at the position corresponding to the target sequence fragment in the nucleic acid sequence to be identified can make the modification of the pentose be detected synchronously with the modification of the phosphorothioation between the nucleotides. For example, a phosphorothioate bond is introduced between the 1st to 4th nucleotides from the 5' end of the nucleic acid sequence, a 2'-O-methyl modification is introduced at the 2nd position of the pentose of the 1st to 3rd nucleotides from the 5' end, and by virtue of the anti-enzyme cutting property of the phosphorothioate bond, the 5' end modification fragment composed of the 4 nucleotides remains intact in the subsequent enzymatic treatment, so that the 2'-O-methyl modification and the phosphorothioation modification can be identified synchronously.

[0056] In some embodiments, the modification to be identified of the nucleic acid sequence to be identified further includes modification of the pentose; and the preset modification rule further includes that the nucleotide corresponding to the modified pentose is connected to at least one adjacent nucleotide by a phosphorothioate bond. In the case that the nucleic acid sequence to be identified can be cleaved to generate a target sequence fragment modified by phosphorothioation, modification of the pentose at the position corresponding to the target sequence fragment in the nucleic acid sequence to be identified can make the modification of the pentose be detected synchronously with the modification of the phosphorothioation between the nucleotides. For example, a phosphorothioate bond is introduced between the 1st to 4th nucleotides from the 5' end of the nucleic acid sequence, a 2'-O-methyl modification is introduced at the 2nd position of the pentose of the 1st to 3rd nucleotides from the 5' end, and by virtue of the anti-enzyme cutting property of the phosphorothioate bond, the 5' end modification fragment composed of the 4 nucleotides remains intact in the subsequent enzymatic treatment, so that the 2'-O-methyl modification and the phosphorothioation modification can be identified synchronously.

[0057] In some embodiments, the nuclease mixture can break phosphodiester bonds while avoiding the breakage of thiophosphate bonds, thereby generating at least two sequence fragments of different molecular weights from the thiophosphorylated nucleic acid sequence to be identified. Specifically, after treatment with the thiophosphorylated nucleic acid sequence to be identified, the phosphodiester bonds are broken due to the action of the nuclease mixture, forming several non-target sequence fragments (single nucleotide sequence fragments); the thiophosphate bonds of the thiophosphorylated nucleic acid sequence to be identified have enzyme-resistant characteristics and are not broken by the enzymatic digestion of the nuclease mixture, thus forming a complete target sequence fragment. Adjacent nucleotides of the target sequence fragment are linked by thiophosphate bonds. The target sequence fragment and the non-target sequence fragments differ significantly in length, molecular weight, structure, etc., thus allowing the target sequence fragment to be separated and identified from the enzymatic digestion products by mass spectrometry analysis.

[0058] In some embodiments, the nucleic acid mixed enzyme may include snake venom phosphodiesterase and / or bovine spleen phosphodiesterase. In other embodiments, the nucleic acid mixed enzyme may also include one or more other enzymes capable of breaking the phosphodiester bonds of the nucleic acid sequence to be identified while avoiding enzymatic digestion by breaking the thiophosphate bonds; this embodiment is not limited thereto.

[0059] In some embodiments, the nucleic acid mixed enzyme has a suitable digestion time, enabling it to fully digest the nucleic acid sequence to be identified while ensuring digestion efficiency. In some embodiments, the digestion time of the nucleic acid mixed enzyme is 1 h to 8 h. For example, the nucleic acid sequence to be identified can be treated with the nucleic acid mixed enzyme for about 1 h, 2 h, 3 h, 4 h, 5 h, 6 h, 7 h, or 8 h to ensure sufficient digestion. In some preferred embodiments, the digestion time of the nucleic acid mixed enzyme is 3 h to 5 h. In some even more preferred embodiments, the digestion time of the nucleic acid mixed enzyme is about 4 h.

[0060] In step 120, the enzymatic hydrolysis product is subjected to secondary mass spectrometry analysis based on the first mass spectrometry information to obtain the second mass spectrometry information of the actual enzymatic hydrolysis product.

[0061] In some embodiments, the first mass spectrum information comprises theoretical analysis results of theoretical sequence fragments generated by analysis. In some embodiments, the analysis of the theoretical sequence fragments comprises theoretical analysis of the first mass spectrum. In some embodiments, the analysis of the theoretical sequence fragments comprises theoretical analysis of the second mass spectrum. Specifically, for a nucleic acid sequence fragment with a certain molecular weight and structure, the results of the first mass spectrum analysis and the second mass spectrum analysis can be obtained by theoretical analysis. For example, for a theoretical sequence fragment A*A*A*A (the "*" represents a phosphorothioate modification) with sequence information determined, the charge states (e.g., the number of charges) and mass-to-charge ratios of a series of charged ions generated after ionization of the theoretical sequence fragment A*A*A*A can be obtained by theoretical analysis; after selecting any one of the series of charged ions as a parent ion, the charge states and mass-to-charge ratios of daughter ions generated by collision-induced dissociation of the parent ion can also be obtained by theoretical analysis.

[0062] In some embodiments, the theoretical sequence fragment is determined based on the preset modification rule, the nucleic acid sequence to be identified, and the nucleic acid mixed enzyme. In some embodiments, based on the preset modification rule and the nucleic acid sequence to be identified, a theoretical design sequence that can be theoretically obtained after correct modification of the nucleic acid sequence to be identified according to the preset modification rule can be determined. For example, the initial nucleic acid sequence is 5'-AATTCCTT-3', and the preset modification rule is to introduce a phosphorothioate bond between the 1st to 3rd nucleotides from the 5' end of the initial nucleic acid sequence. Then, the theoretical design sequence obtained after correct modification is 5'-A*A*TTCCTT-3', wherein the "*" represents a phosphorothioate modification. In some embodiments, based on the theoretical design sequence and the nucleic acid mixed enzyme, a theoretical sequence fragment can be determined, and the theoretical sequence fragment exists in the theoretical enzyme digestion product of the theoretical design sequence. For example, the theoretical design sequence is 5'-A*A*TTCCTT-3', and after treatment with the nucleic acid mixed enzyme, the theoretical design sequence can be cleaved into a modified sequence fragment A*A*T and a single nucleotide sequence fragment without modification, and the modified sequence fragment A*A*T is the theoretical sequence fragment.

[0063] In some embodiments, the adjacent nucleotides of the theoretical sequence fragment are connected by a phosphorothioate bond. The theoretical sequence fragment can be a chemically modified sequence fragment in the theoretical design sequence. Further, the first mass spectrum information (e.g., the molecular weight, the charge state of the charged ion generated after ionization, the mass-to-charge ratio of the charged ion, etc.) can be used to characterize the nucleotide arrangement order, the chemical modification, etc. of the theoretical sequence fragment, and the second mass spectrum information can be used to characterize the nucleotide arrangement order, the chemical modification, etc. of the chemically modified target sequence fragment in the nucleic acid sequence to be identified, so that the first mass spectrum information and the second mass spectrum information have direct correspondence, facilitating comparison in subsequent steps.

[0064] In some embodiments, there can be one or more theoretical sequence fragments based on different preset modification rules, and correspondingly, there can be one or more target sequence fragments chemically modified in the nucleic acid sequence to be identified, without limitation of the present embodiments. In some embodiments, each theoretical sequence fragment consists of at least two nucleotides.

[0065] In some embodiments, the first mass spectrum information can comprise first parent ion information of the theoretical sequence fragment. As used herein, the first parent ion is an ion of interest generated from the theoretical sequence fragment in an ionization process of mass spectrum analysis. In some embodiments, the ion of interest can be selected from a series of charged ions theoretically generated from the theoretical sequence fragment in the ionization process as the first parent ion. For example, the theoretical sequence fragment generates a series of charged ions including single-charged state ions, double-charged state ions and triple-charged state ions in the ionization process, and any one of the single-charged state ions, the double-charged state ions and the triple-charged state ions of the theoretical sequence fragment can be selected as the first parent ion.

[0066] In some embodiments, the first parent ion information at least comprises the mass-to-charge ratio of the first parent ion. In some embodiments, the first parent ion information further comprises the charge state of the first parent ion, such as the number of charges.

[0067] In some embodiments, the first parent ion information can be determined by theoretical analysis, such as determining the mass-to-charge ratio, the charge state and the like of the first parent ion by theoretical analysis. In some embodiments, the mass-to-charge ratio of the first parent ion can be determined based on the molecular weight of the theoretical sequence fragment and the charge state of the first parent ion. For example, the molecular weight of the theoretical sequence fragment is 300, the first parent ion is a single-charged state ion with a charge number of 1, and according to the relationship formula between the mass-to-charge ratio and the molecular weight: molecular weight = mass-to-charge ratio * charge number + charge number, the mass-to-charge ratio of the first parent ion is determined to be 299.

[0068] In some embodiments, the first mass spectrum information further comprises first daughter ion information of the theoretical sequence fragment.

[0069] As used herein, the first daughter ion is a fragment ion generated from the first parent ion in a dissociation step of mass spectrum analysis. Specifically, the first parent ion collides with a gas to cause specific chemical bonds (such as P-O bonds on the phosphodiester bonds between adjacent nucleotides) to break, thereby theoretically forming a series of fragment ions, and the first daughter ion can comprise one or more pairs of first daughter ions capable of complementary pairing. For example, a divalent ion generated from the ionization of the theoretical sequence fragment T*A*C*A is selected as the first parent ion. In the fragment ions generated by the dissociation of the first parent ion, there can be a pair of complementary paired daughter ions corresponding to the fragmentation fragments T*A* and C*A, and there can also be a pair of complementary paired daughter ions corresponding to the fragmentation fragments T*A and *C*A.

[0070] In some embodiments, the first sub-ion information comprises at least the complementary pairing information and the mass-to-charge ratio of the first sub-ion. In some embodiments, the complementary pairing information of the first sub-ion comprises the charge state of the first sub-ion and the complementary paired first sub-ion. In some embodiments, the complementary pairing information of the first sub-ion comprises the molecular weight of the fragmentation fragment corresponding to the first sub-ion and the complementary paired fragmentation fragment.

[0071] In some embodiments, the first sub-ion information can be determined by theoretical analysis, such as determining the charge state, mass-to-charge ratio, etc. of the first sub-ion by theoretical analysis. In some embodiments, the mass-to-charge ratio of the first sub-ion can be determined based on the molecular weight of the fragmentation fragment corresponding to the first sub-ion and the charge state of the first sub-ion. For more information about determining the mass-to-charge ratio of the first sub-ion, please refer to the relevant description of determining the mass-to-charge ratio of the first parent ion.

[0072] In some embodiments, the second mass spectrum information can comprise second parent ion information. In some embodiments, step 120 can further comprise a step of determining the second parent ion information. Specifically, the step of determining the second parent ion information comprises determining the second parent ion information based on the first parent ion information.

[0073] As used herein, the second parent ion is an ion of interest generated in the ionization process of the enzymatic product of the nucleic acid sequence to be identified in mass spectrometric analysis. In some embodiments, the ion of interest can be screened from a series of charged ions actually generated in the ionization process of the enzymatic product of the nucleic acid sequence to be identified as the second parent ion. In the case where the nucleic acid sequence to be identified is correctly chemically modified based on the preset modification rule, the charged ions actually generated in the ionization process of the enzymatic product of the nucleic acid sequence to be identified can include the charged ions theoretically generated in the ionization process of the theoretical sequence fragment. Determining the second parent ion information based on the first parent ion information can make the fragment ions generated by the second parent ion correspond to the fragment ions generated by the first parent ion, so that the comparison results of the two can reflect the modification of the nucleic acid sequence to be modified.

[0074] In some embodiments, the second parent ion information comprises at least the mass-to-charge ratio of the second parent ion. In some embodiments, the second parent ion information further comprises the charge state of the second parent ion, such as the number of charges, etc. For example, in the secondary mass spectrometric analysis of the nucleic acid sequence to be modified, the second parent ion with the mass-to-charge ratio and the charge state matching the first parent ion is selected for the dissociation process and detection. Wherein, the mass-to-charge ratio of the second parent ion is the same as that of the first parent ion, or the difference between the mass-to-charge ratios of the two is within the allowable range (such as the difference is less than 0.3, 0.5, 0.8 or 1.3), and the charge state of the second parent ion is the same as that of the first parent ion.

[0075] In some embodiments, the second mass spectrum information can comprise second sub-ion information. In some embodiments, step 120 can further comprise a step of determining the second sub-ion information. Specifically, the step of determining the second sub-ion information comprises performing a secondary mass spectrum analysis of the enzymatic product based on the second parent ion information to obtain the second sub-ion information.

[0076] As used herein, a second sub-ion is a fragment ion of a second parent ion produced in a dissociation process of a mass spectrum analysis. Specifically, a second parent ion collides with a gas to cause specific chemical bonds to break and actually form a series of fragment ions, among which one or more pairs of second sub-ions can be able to complementarily pair.

[0077] In some embodiments, the second sub-ion information at least comprises complementary pairing information and mass-to-charge ratio of the second sub-ion.

[0078] In some embodiments, the mass-to-charge ratio of the second sub-ion can be directly obtained from the result of actual secondary mass spectrum analysis, such as directly reading the mass-to-charge ratio data of the second sub-ion from a secondary mass spectrum graph.

[0079] In some embodiments, the complementary pairing information of the second sub-ion comprises charge states of the second sub-ion and the second sub-ion that complementarily pairs therewith. In some embodiments, the complementary pairing information of the second sub-ion comprises molecular weights of the fragmentation fragment corresponding to the second sub-ion and the fragmentation fragment that complementarily pairs therewith.

[0080] In step 130, based on the first mass spectrum information and the second mass spectrum information, it is determined whether a target sequence fragment exists in the nucleic acid sequence to be identified. The target sequence fragment is consistent with the theoretical sequence fragment.

[0081] In some embodiments, the determination of whether the target sequence fragment consistent with the theoretical sequence fragment exists in the nucleic acid sequence to be identified can further comprise a step of determining whether the first sub-ion information and the second sub-ion information satisfy a preset matching condition; if the preset matching condition is satisfied, it is determined that the target sequence fragment exists in the nucleic acid sequence to be identified; if the preset matching condition is not satisfied, it is determined that the target sequence fragment does not exist in the nucleic acid sequence to be identified.

[0082] Specifically, it is determined that the target sequence fragment exists in the to-be-identified nucleic acid sequence, i.e., the nucleotide arrangement order and the type, number, and position of chemical modifications of the target sequence fragment existing in the to-be-identified nucleic acid sequence meet the requirements of the preset modification rule, and the chemical modification of the to-be-identified nucleic acid sequence is correct. It should be understood that the identification result of correct modification means that at least part of the chemical modification of the to-be-identified nucleic acid sequence in the sample of the to-be-identified nucleic acid sequence is correct. Due to the characteristics of mass spectrometry, if the to-be-identified nucleic acid sequence has changes in the nucleotide arrangement order and / or changes in the type, number, and position of chemical modifications, the mass-to-charge ratio of the daughter ion produced by the enzymatic product will change significantly. Therefore, the comparison result of the actual daughter ion information and the theoretical daughter ion information of the to-be-identified nucleic acid sequence can accurately reflect whether the nucleotide arrangement order and the chemical modification of the to-be-identified nucleic acid sequence meet the expected target sequence fragment.

[0083] In some embodiments, the preset matching condition is set based on the complementary pairing information and the mass-to-charge ratio of the first daughter ion and the second daughter ion. In some embodiments, the preset matching condition is that there are at least one pair of first daughter ions and at least one pair of second daughter ions, and the complementary pairing information and the mass-to-charge ratio of the at least one pair of first daughter ions and the at least one pair of second daughter ions are the same.

[0084] For example, the series of fragment ions formed by the cleavage of the divalent parent ion of the theoretical sequence fragment A*C*A*G includes the complementary paired first daughter ion A-1 and the first daughter ion A-2, and the complementary paired first daughter ion B-1 and the first daughter ion B-2. Among them, the first daughter ion A-1 and the first daughter ion A-2 correspond to the fragmentation fragments A*C* and A*G respectively, and the first daughter ion B-1 and the first daughter ion B-2 correspond to the fragmentation fragments A*C and *A*G respectively. If the fragment ions produced by the enzymatic product of the to-be-identified nucleic acid sequence include the second daughter ion A’-1 and the second daughter ion A’-2, wherein the mass-to-charge ratio of the second daughter ion A’-1 is the same as that of the first daughter ion A-1, the mass-to-charge ratio of the second daughter ion A’-2 is the same as that of the first daughter ion A-2, and the second daughter ion A’-1 and the second daughter ion A’-2 are in a complementary pairing relationship, it can be determined that the first daughter ion information and the second daughter ion information meet the preset matching condition. If the second daughter ion obtained by the enzymatic product of the to-be-identified nucleic acid sequence only includes the second daughter ion A’-1, and the second daughter ion B’-1 with the same mass-to-charge ratio as the first daughter ion B-1, the mass-to-charge ratio of the second daughter ion A’-1 is the same as that of the first daughter ion A-1, the mass-to-charge ratio of the second daughter ion B’-1 is the same as that of the first daughter ion B-1, and the second daughter ion A’-1 and the second daughter ion B’-1 have no complementary pairing relationship, it can be determined that the first daughter ion information and the second daughter ion information do not meet the preset matching condition.

[0085] In some embodiments, the secondary mass spectrometry analysis of the enzymatic products can be performed by a linear ion trap mass spectrometer. In other embodiments, the secondary mass spectrometry analysis of the enzymatic products can also be performed by other functionally equivalent or similar instruments, and the embodiments are not limited in this respect.

[0086] The modification identification method of the thio-phosphorylated modified nucleic acid sequence disclosed in the specification can bring beneficial effects including but not limited to: (1) the modification identification method of the embodiments of the specification determines whether the chemical modification of the nucleic acid sequence to be identified is correct based on the comparison between the actual sub-ion information generated by the enzymatic and secondary mass spectrometry analysis of the nucleic acid sequence to be identified and the theoretically generated sub-ion information, and the identification result is accurate and intuitive; (2) in addition to the identification of thio-phosphorylation modification, the modification identification method of the embodiments of the specification can also identify pentose and / or other chemical modifications of bases simultaneously, and the method is efficient; (3) the modification identification method of the embodiments of the specification directly characterizes the nucleic acid sequence to be identified after enzymatic hydrolysis by secondary mass spectrometry analysis, omitting the cumbersome steps such as purification of the enzymatic products, saving time and samples. It should be noted that different embodiments can produce different beneficial effects, and in different embodiments, the beneficial effects that can be produced can be any one or a combination of several of the above, or any other beneficial effects that can be obtained.

[0087] The experimental methods in the following examples are all conventional methods unless otherwise specified. The experimental materials used in the following examples are all purchased from conventional biochemical reagent companies unless otherwise specified. The quantitative tests in the following examples all set three repeated experiments, and the results are averaged.

[0088] Materials and instruments used in the following examples include

[0089]

[0090]

[0091] LC-MS mobile phase configuration

[0092] LC-MS-buffer A: 2 mL HFIPA + 250 μL TEA + 10 mL EDTA is added to 1 L water, mixed and ultrasonicated to serve as mobile phase A;

[0093] LC-MS-buffer B: 2 mL HFIPA + 250 μL TEA + 10 mL EDTA + 190 mL water is added to 800 mL acetonitrile, mixed and ultrasonicated to serve as mobile phase B.

[0094] LC-MSMS mobile phase configuration:

[0095] ​​LC-MSMS-buffer A: 2 mL HFIPA + 800 μL DIEA + 8 mL LC-MS EDTA added to 792 mL water, mixed and sonicated to be mobile phase A;

[0096] LC-MSMS-buffer B: 600 μL HFIPA + 300 μL DIEA + 8 mL LC-MS EDTA + 160 mL water added to 640 mL acetonitrile, mixed and sonicated to be mobile phase B.

[0097] Example 1, Background interference of nucleic acid mixing enzyme

[0098] 1.1, LC-MS analysis of control sample 1

[0099] 1.1.1, Enzymatic treatment

[0100] Control sample 1 is a single-stranded RNA fragment 5'-sgRNA-4nt of enterprise control product, and the relevant information of the RNA fragment 5'-sgRNA-4nt is shown in Table 1.1.

[0101] Table 1.1 - Sequence information of 5'-sgRNA-4nt

[0102] Name Sequence Molecular weight 5'-sgRNA-4nt 5'mG*mA*mG*rU 3' 1353.84

[0103] Note: m represents 2'-OMe modification, * represents thio-phosphorylation modification, and r represents RNA.

[0104] Using control sample 1 as the substrate, the enzymatic reaction solution of control sample 1 was prepared according to Table 1.2, mixed and centrifuged, and then incubated in a digital temperature-controlled metal bath at 37°C to obtain the enzymatic product.

[0105] Table 1.2 - Components of enzymatic reaction solution

[0106]

[0107]

[0108] 1.1.2, LC-MS sample analysis

[0109] 5 μL of the enzymatic product of step 1.1.1 was diluted with 45 μL of water, and LC-MS analysis of the diluted enzymatic product was performed using the aforementioned LC-MS mobile phase configuration and a time-of-flight mass spectrometer.

[0110] 1.1.3, LC-MS mass spectrum results and analysis

[0111] As shown in Table 1.1, the theoretical molecular weight of the RNA fragment 5'-sgRNA-4nt is 1353.84, and through theoretical analysis, it can be known that the charged ions of the RNA fragment 5'-sgRNA-4nt include m / z = 1352.84 ([M-H] - ), m / z = 675.92 ([M-2H] 2- ) and m / z = 450.28 ([M-3H] 3- ).

[0112] Figure 2 is the LC-MS mass spectrum of the control sample 1. As shown in Figure 2 , the characteristic peaks of m / z = 1352.1973, m / z = 675.5928 and m / z = 450.0583 can correspond to [M-H] - , [M-2H] 2- and [M-3H] 3- of the RNA fragment 5'-sgRNA-4nt, respectively. It can be known that the actual parent ion information (species and mass-to-charge ratio) of the control sample 1 is consistent with the theoretical parent ion information. The nucleic acid mixed enzyme does not cut the phosphorothioate bond between adjacent nucleotides in the RNA fragment of the control sample 1, and the enzyme digestion effect is consistent with the expectation.

[0113] 1.2, LC-MS analysis of control sample 2

[0114] 1.2.1, enzyme digestion treatment

[0115] The control sample 2 is an enterprise control sample of the single-stranded RNA fragment 3'-sgRNA-4nt, and the related information of the RNA fragment 3'-sgRNA-4nt is shown in Table 1.3. The enzyme digestion reaction solution of the control sample 2 was prepared according to Table 1.2 with the control sample 2 as the substrate, and after mixing and centrifugation, it was placed in a digital display controlled temperature metal bath at 37°C for incubation to obtain the enzyme digestion product.

[0116] Table 1.3 - Sequence information of 3'-sgRNA-4nt

[0117] Name Sequence Molecular weight 3'-sgRNA-4nt 5'mU*mU*mU*rU 3' 1252.72

[0118] 1.2.2, LC-MS sample analysis

[0119] 5 μL of the enzyme digestion product of step 1.2.1 was diluted with 45 μL of water, and LC-MS analysis of the diluted enzyme digestion product was performed using the aforementioned LC-MS mobile phase configuration and time-of-flight mass spectrometer.

[0120] 1.2.3, LC-MS mass spectrum results and analysis

[0121] As shown in Table 1.3, the molecular weight of the RNA fragment 3'-sgRNA-4nt is 1252.72, and it can be known through theoretical analysis that the charged ions of the RNA fragment 3'-sgRNA-4nt include m / z = 1251.72 ([M-H] - ) and m / z = 625.36 ([M-2H] 2- ).

[0122] Figure 3 is the LC-MS mass spectrum of the control sample 2. As shown in Figure 3 , the characteristic peaks of m / z = 1251.1244 and m / z = 625.0575 can correspond to [M-H] - and [M-2H] 2- of the RNA fragment 3'-sgRNA-4nt, respectively. It can be known that the actual charged ion information of the control sample 2 is consistent with the theoretical charged ion information. The nucleic acid mixmer does not cut the phosphorothioate bond between adjacent nucleotides in the RNA fragment of the control sample 2, and the enzyme digestion effect is as expected.

[0123] 1.3, LC-MS analysis of the mixed enzyme reaction buffer

[0124] 1.3.1, Preparation of the mixed enzyme reaction buffer: prepare the mixed enzyme reaction buffer without substrate according to Table 1.2, and reserve.

[0125] 1.3.2, LC-MS sample analysis

[0126] Dilute 5 μL of the mixed enzyme reaction buffer of step 1.3.1 with 45 μL of water, use the aforementioned LC-MS mobile phase configuration, and use a time-of-flight mass spectrometer to perform LC-MS analysis on the diluted mixed enzyme reaction buffer.

[0127] 1.3.3, LC-MS mass spectrum results and analysis

[0128] Figure 4 is the total ion current chromatogram of the mixed enzyme reaction buffer, Figure 5 is the mass spectrum of the mixed enzyme reaction buffer LC-MS data at 5.925 min. As shown in Figure 4 and Figure 5 , in addition to the background peak at a retention time of 0.5 min, no characteristic peaks of the RNA fragment 5'-sgRNA-4nt and the RNA fragment 3'-sgRNA-4nt appear within the gradient of 3-10 min. It can be seen that the mixed enzyme reaction buffer does not interfere with the mass spectrum results of the enzyme digestion products.

[0129] Example 2, Mass spectrum analysis of sample 1 enzyme digestion products under different enzyme digestion times

[0130] 2.1, Enzyme digestion reaction

[0131] Sample 1 is a sample to be identified of single-stranded RNA fragment ET-02sgRNA. The relevant information of RNA fragment ET-02sgRNA is shown in Table 2.1. In the table, the sequence of the 5' end modification fragment of RNA fragment ET-02sgRNA is 5'-mG*mA*mG*rU-3', which is consistent with the sequence of RNA fragment 5'-sgRNA-4nt; the sequence of the 3' end modification fragment of RNA fragment ET-02sgRNA is 5'-mU*mU*mU*rU-3', which is consistent with the sequence of RNA fragment 3'-sgRNA-4nt. The enzyme hydrolysis reaction solution of sample 1 was prepared according to Table 1.2 of Example 1 using sample 1 as the substrate, mixed and centrifuged, and then incubated in a 37℃ digital controlled temperature metal bath to obtain the enzyme hydrolysis product.

[0132] Table 2.1 - Sequence information of ET-02sgRNA

[0133]

[0134] Note: N represents A, G, C, T or U.

[0135] 2.2, LC-MS sample analysis

[0136] After 2min, 4h and 24h of enzyme hydrolysis, 5μL of the enzyme hydrolysis product of step 2.1 was diluted with 45μL of water, and LC-MS analysis of the diluted enzyme hydrolysis product was performed using the aforementioned LC-MS mobile phase configuration and a time-of-flight mass spectrometer.

[0137] 2.3, LC-MS mass spectrum results and analysis

[0138] Figure 6 is the LC-MS mass spectrum of the enzyme hydrolysis product of sample 1 after 2min of enzyme hydrolysis. In the figure, Figure 6 The characteristic peak of m / z = 1251.1328 shown in A can correspond to [M-H] - of the 3' end modification fragment of ET-02sgRNA. Figure 6 The characteristic peaks of m / z = 1352.2057 and m / z = 675.6022 shown in B can correspond to [M-H] - , [M-H] 2- of the 5' end modification fragment of ET-02sgRNA, respectively. The appearance of the corresponding parent ion signals of the 5' end modification fragment and the 3' end modification fragment of ET-02sgRNA in the mass spectrum results of the enzyme hydrolysis product indicates that nucleic acid mixed enzyme can produce part of the enzyme hydrolysis product after a short time of enzyme hydrolysis treatment of sample 1.

[0139] Figure 7 is the LC-MS mass spectrum of the enzyme hydrolysis product of sample 1 after 4h of enzyme hydrolysis. In the figure,Figure 7 The characteristic peaks of m / z = 1251.1318 and m / z = 625.0617 shown in A can correspond to [M-H] of the 3' end modified fragment respectively - , [M-H] 2- . Figure 7 The characteristic peaks of m / z = 1352.2027 and m / z = 675.5978 shown in B can correspond to [M-H] of the 5' end modified fragment respectively - , [M-H] 2- . The corresponding parent ion signal of the complete sequence fragment of ET-02 sgRNA was not found in the LC-MS mass spectrum result. Compared with the enzymatic product of sample 1 enzymolysis 2 min, the corresponding parent ion signal response (such as absolute intensity) of the 5' end modified fragment and the 3' end modified fragment of ET-02 sgRNA in the LC-MS mass spectrum result of the enzymatic product of sample 1 enzymolysis 4 h increased significantly.

[0140] Figure 8 is the LC-MS mass spectrum of the enzymatic product of sample 1 enzymolysis 24 h. Among them, Figure 8 The characteristic peaks of m / z = 1251.1223 and m / z = 625.0563 shown in A can correspond to [M-H] of the 3' end modified fragment respectively - , [M-H] 2- . Figure 8 The characteristic peaks of m / z = 1352.1922 and m / z = 675.5920 shown in B can correspond to [M-H] of the 5' end modified fragment respectively - , [M-H] 2- . Compared with the enzymatic product of sample 1 enzymolysis 4 h, the corresponding parent ion signal response of the 5' end modified fragment and the 3' end modified fragment of ET-02 sgRNA in the LC-MS mass spectrum result of the enzymatic product of sample 1 enzymolysis 24 h did not increase significantly.

[0141] In summary, after 4 h of enzymolysis treatment by nucleic acid mixing enzyme, the enzymolysis of the substrate is relatively sufficient, so that the 5' end modified fragment and the 3' end modified fragment of ET-02 sgRNA are accurately obtained.

[0142] Example 3, modification identification of sample 1

[0143] 3.1, enzymolysis treatment

[0144] Take sample 1 as the substrate, prepare the enzymolysis reaction solution of sample 1 according to Table 1.2 of Example 1, mix well, centrifuge, and place in a 37℃ digital display temperature metal bath for incubation for 4 h to obtain the enzymatic product. The relevant information of sample 1 is shown in Table 2.1 of Example 2.

[0145] 3.2, LC-MSMS sample analysis

[0146] Take 5 μL of the enzymatic hydrolysis product from step 3.1 and add water to make up to 10 μL. Use the mobile phase prepared for LC-MSMS analysis as described above. Analyze the diluted enzymatic hydrolysis product using a linear ion well mass spectrometer. LC-MSMS is performed with a single injection of 10 μL.

[0147] 3.3 LC-MSMS Mass Spectrometry Results and Analysis

[0148] 3.3.1 Identification of 5' end modifications in the RNA sequence of Sample 1

[0149] The study in Example 1 confirmed that, due to the anti-enzymatic cleavage effect of the thiophosphate bond, the 5' end modified fragment mG*mA*mG*rU of ET-02sgRNA remains intact after digestion with a mixed nucleic acid enzyme. Theoretical analysis by LC-MSMS showed that the 5' end modified fragment mG*mA*mG*rU forms a series of charged ions after ionization. One of these charged ions is selected as the parent ion (e.g., a divalent parent ion with m / z = 675.92), and dissociation is induced by a gas collision with it at a certain energy level. This causes the specific chemical bond of the parent ion to break (generally the PO bond on the phosphodiester bond between adjacent nucleotides), forming several daughter ions. These sub-ions contain multiple pairs of sub-ions that can complement each other. For example, the monovalent sub-ion corresponding to the fragment mG*mA* (m / z = 733.6) and the monovalent sub-ion corresponding to the fragment mG*rU (m / z = 618.5) are a pair of complementary sub-ions.

[0150] Divalent precursor ions with m / z = 675.92 were screened in negative mode for LC-MS / MS analysis of the enzymatic digestion products of sample 1. The mass spectrometry results are shown in [Figure number missing]. Figure 9 . Figure 9 The label includes information on the fragmentation segments corresponding to the characteristic peaks and the charge number of the daughter ions (Z represents the charge number). For example... Figure 9 As shown, the enzymatic digestion product of sample 1 was cleaved into several daughter ions in LC-MS analysis. The characteristic peaks at m / z = 618.12 and m / z = 733.08 correspond to a pair of complementary daughter ions, namely, a monovalent daughter ion corresponding to fragment fragment mG*rU and fragment fragment mG*mA*; the characteristic peaks at m / z = 655.13 and m / z = 733.08 also correspond to a pair of complementary daughter ions, namely, a monovalent daughter ion corresponding to fragment fragment mG*mA and fragment fragment *mG*rU.

[0151] Therefore, the actual daughter ion information generated by the enzymatic digestion product of Sample 1 matches the theoretically generated daughter ion information of the 5' modified fragment mG*mA*mG*rU. The RNA sequence of Sample 1 contains a target sequence fragment identical to the 5' modified fragment mG*mA*mG*rU. This target sequence fragment has the same nucleotide sequence and chemical modification as the 5' modified fragment mG*mA*mG*rU, thus verifying that the 5' end of the RNA sequence of Sample 1 has the correct 2'-O-methylation and thiophosphorylation modifications.

[0152] 3.3.2 Identification of 3' end modifications in the RNA sequence of Sample 1

[0153] The study in Example 1 confirmed that, due to the anti-enzymatic cleavage effect of the thiophosphate bond, after cleavage of ET-02sgRNA by a mixed nuclease, the thiophosphate bond of the 3' modified fragment mU*mU*mU*rU of ET-02sgRNA is not destroyed and remains an intact fragment. In the theoretical analysis results of LC-MSMS mass spectrometry, after screening for a suitable precursor ion (such as a divalent precursor ion with m / z = 625.36), the 3' modified fragment mU*mU*mU*rU can form several daughter ions. These daughter ions contain multiple pairs of complementary daughter ions; for example, the monovalent daughter ion corresponding to the fragmented fragment mU*mU* (m / z = 671.5) and the monovalent daughter ion corresponding to the fragmented fragment mU*rU (m / z = 579.5) are a complementary pair of daughter ions.

[0154] Divalent precursor ions with m / z = 625.36 were screened in negative mode for LC-MS / MS analysis of the enzymatic digestion products of sample 1. The mass spectrometry results are shown in [Figure number missing]. Figure 10 . Figure 10 The label includes information on the fragmentation segments corresponding to the characteristic peaks and the charge number of the daughter ions. For example... Figure 10 As shown, the enzymatic digestion product of sample 1 was cleaved into several daughter ions in LC-MS analysis. Specifically, the characteristic peaks at m / z = 579.09 and m / z = 671.01 correspond to a pair of complementary daughter ions, namely, a monovalent daughter ion corresponding to fragment mU*rU and fragment mU*mU*; the characteristic peaks at m / z = 593.10 and m / z = 656.99 correspond to a pair of complementary daughter ions, namely, a monovalent daughter ion corresponding to fragment mU*mU and fragment *mU*rU; and the characteristic peaks at m / z = 321.00 and m / z = 929.08 correspond to a pair of complementary daughter ions, namely, a monovalent daughter ion corresponding to fragment *rU and fragment mU*mU*mU.

[0155] It can be seen that the actual sub-ion information of the sample 1 enzyme product can match the theoretical sub-ion information of the 3' end modified fragment mU*mU*mU*rU. The sample 1 RNA sequence has a target sequence fragment consistent with the 3' end modified fragment mU*mU*mU*rU, which has the same nucleotide arrangement order and chemical modification as the 3' end modified fragment mU*mU*mU*rU, thereby verifying that the sample 1 RNA sequence 3' end has correct 2'-O-methylation modification and phosphorothioation modification.

[0156] Example 4, modification identification of sample 2

[0157] 4.1, enzymatic treatment

[0158] The sample 2 is a single-stranded RNA fragment ET-01 sgRNA to be identified. The relevant information of the RNA fragment ET-01 sgRNA is shown in Table 4.1. Among them, the sequence of the 5' end modified fragment of the RNA fragment ET-01 sgRNA is 5'-mA*mU*mC*rA-3'; the sequence of the 3' end modified fragment of the RNA fragment ET-01 sgRNA is 5'-mU*mU*mU*rU-3', which is consistent with the sequence of the RNA fragment 3'-sgRNA-4nt. The sample 2 was used as the substrate, and the enzyme reaction solution of the sample 2 was prepared according to Table 1.2 of Example 1. After mixing and centrifugation, it was placed in a 37℃ digital display temperature metal bath for incubation for 4h, and the enzyme product was obtained.

[0159] Table 4.1-Sequence information of ET-01 sgRNA

[0160]

[0161]

[0162] 4.2, LC-MSMS sample analysis

[0163] 5μL of the enzyme product of step 4.1 was diluted to 10μL, and LC-MSMS analysis of the diluted enzyme product was performed using the aforementioned LC-MSMS mobile phase configuration using a linear ion well mass spectrometer. LC-MSMS was used for 10μL of sample.

[0164] 4.3, LC-MSMS mass spectrum results and analysis

[0165] 4.3.1, identification of the 5' end modification of the RNA sequence of sample 2

[0166] Due to the anti-enzyme cutting effect of the phosphorothioate bond, the 5' end modified fragment mA*mU*mC*rA of the ET-01 sgRNA will not be cut off by the enzyme and remains as a complete fragment. The theoretical molecular weight of the 5' end modified fragment mA*mU*mC*rA is 1298.1. In the theoretical analysis results of LC-MSMS mass spectrum analysis, after screening the appropriate parent ions (such as the divalent parent ion of m / z = 648), the 5' end modified fragment mA*mU*mC*rA can form several daughter ions. Among the several daughter ions, there are several pairs of complementary pairing daughter ions, for example, the monovalent daughter ion (m / z = 694.5) corresponding to the fragmentation fragment mA*mU and the monovalent daughter ion (m / z = 601.5) corresponding to the fragmentation fragment mC*rA are a pair of complementary pairing daughter ions.

[0167] LC-MSMS analysis of sample 2 enzyme product was carried out by screening the divalent parent ion of m / z = 648 in negative mode, and the mass spectrum results are shown in Figure 11 . Figure 11 The fragmentation fragment information corresponding to the characteristic peaks and the charge number of the daughter ions are marked in the figure. As shown in Figure 11 , the sample 2 enzyme product produces several daughter ions in LC-MSMS analysis. Among them, the characteristic peaks of m / z = 616.00 and m / z = 678.95 can correspond to a pair of complementary pairing daughter ions, that is, a pair of monovalent daughter ions corresponding to the fragmentation fragments mA*mU and *mC*rA; the characteristic peaks of m / z = 601.02 and m / z = 693.93 can correspond to a pair of complementary pairing daughter ions, that is, a pair of monovalent daughter ions corresponding to the fragmentation fragments mC*rA and *mA*mU*; the characteristic peaks of m / z = 343.98 and m / z = 950.99 can correspond to a pair of complementary pairing daughter ions, that is, a pair of monovalent daughter ions corresponding to the fragmentation fragments *rA and mA*mU*mC.

[0168] As can be seen, the actual daughter ion information produced by the sample 2 enzyme product can match the theoretical daughter ion information produced by the 5' end modified fragment mA*mU*mC*rA. The RNA sequence of sample 2 contains a target sequence fragment consistent with the 5' end modified fragment mA*mU*mC*rA, which has the same nucleotide arrangement order and chemical modification as the 5' end modified fragment mA*mU*mC*rA, thereby verifying that the 5' end of the RNA sequence of sample 2 has correct 2'-O-methylation modification and phosphorothioation modification.

[0169] 4.3.2, identification of the 3' end modification of the RNA sequence of sample 2

[0170] Referring to 3.3.2 of the embodiment 3, LC-MSMS analysis of sample 2 enzyme product was carried out by screening the divalent parent ion of m / z = 625.36 in negative mode, and the mass spectrum results are shown in Figure 12 .Figure 12 The label includes information on the fragmentation segments corresponding to the characteristic peaks and the charge number of the daughter ions. For example... Figure 12 As shown, the enzymatic digestion product of sample 2 was cleaved into several daughter ions in LC-MS analysis. Specifically, the characteristic peaks at m / z = 578.98 and m / z = 670.89 correspond to a pair of complementary daughter ions, namely, a monovalent daughter ion corresponding to fragment mU*rU and fragment mU*mU*; the characteristic peaks at m / z = 592.98 and m / z = 656.88 correspond to a pair of complementary daughter ions, namely, a monovalent daughter ion corresponding to fragment mU*mU and fragment *mU*rU; and the characteristic peaks at m / z = 320.94 and m / z = 928.93 correspond to a pair of complementary daughter ions, namely, a monovalent daughter ion corresponding to fragment *rU and fragment mU*mU*mU.

[0171] Therefore, the actual daughter ion information generated by the enzymatic digestion product of Sample 2 matches the theoretically generated daughter ion information of the 3' end modified fragment mU*mU*mU*rU. The RNA sequence of Sample 2 contains a target sequence fragment identical to the 3' end modified fragment mU*mU*mU*rU. This target sequence fragment has the same nucleotide sequence and chemical modification as the 3' end modified fragment mU*mU*mU*rU, thus verifying that the 3' end of the RNA sequence of Sample 2 has the correct 2'-O-methylation and thiophosphorylation modifications.

[0172] Those skilled in the art should understand that the above embodiments are merely illustrative of the present invention and do not constitute a limitation thereof. Any modifications, equivalent substitutions, and variations made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for identifying a modification of a phosphorothioated nucleic acid sequence, characterized in that, The method comprises: performing enzymolysis on the to-be-identified nucleic acid sequence by using a nucleic acid mixed enzyme to obtain an enzymolysis product; wherein the to-be-identified nucleic acid sequence is actually obtained based on a preset modification rule, and the to-be-identified modification of the to-be-identified nucleic acid sequence at least comprises a phosphorothioation modification of a connection between nucleotides; the nucleic acid mixed enzyme can break a phosphodiester bond and is immune to breaking a phosphorothioate bond, so that the to-be-identified nucleic acid sequence modified by phosphorothioation generates sequence fragments of at least two different molecular weights; the performing enzymolysis on the to-be-identified nucleic acid sequence by using the nucleic acid mixed enzyme to obtain the enzymolysis product comprises: using the anti-enzymolysis characteristic of the phosphorothioate bond, so that the to-be-identified nucleic acid sequence can be cleaved to generate a target sequence fragment modified by phosphorothioation and a non-target sequence fragment not modified by phosphorothioation after being processed by the nucleic acid mixed enzyme; performing secondary mass spectrometry analysis on the enzymolysis product based on first mass spectrometry information to obtain second mass spectrometry information actually generated by the enzymolysis product, wherein the first mass spectrometry information comprises a theoretical analysis result generated by analyzing a theoretical sequence fragment, and the theoretical sequence fragment is determined based on the preset modification rule, the to-be-identified nucleic acid sequence and the nucleic acid mixed enzyme; and determining, based on the first mass spectrometry information and the second mass spectrometry information, whether the target sequence fragment consistent with the theoretical sequence fragment exists in the to-be-identified nucleic acid sequence.

2. The method of claim 1, wherein, The nucleic acid mixed enzyme comprises a snake venom phosphodiesterase and / or a bovine spleen phosphodiesterase.

3. The method of claim 1, wherein, The enzymolysis time of the nucleic acid mixed enzyme is 1 h to 8 h.

4. The method of claim 1, wherein, The enzymolysis time of the nucleic acid mixed enzyme is 3 h to 5 h.

5. The method of any one of claims 1-4, wherein, The to-be-identified modification of the to-be-identified nucleic acid sequence can further comprise a modification of a pentose and / or a modification of a base.

6. The method of claim 5, wherein, The modification of the pentose comprises at least one of the following modifications: a 2'-modification, a 5'-modification, a 3'-modification, a locked nucleic acid modification, an unlocked nucleic acid modification and a peptide nucleic acid modification.

7. The method of claim 6, wherein, The 2'-modification of the pentose is selected from a 2'-O-methyl modification, a 2'-fluoro modification, a 2'-O-methoxyethyl modification, a 2'-O-methylation modification and a 2'-O-allyl modification.

8. The method of claim 5, wherein, The modification of the base comprises at least one of the following modifications: a methylation modification and a hydroxymethylation modification.

9. The method of claim 5, wherein, The preset modification rule comprises: a nucleotide corresponding to a modified pentose is connected with at least one adjacent nucleotide by a phosphorothioate bond; and / or the preset modification rule comprises: a nucleotide corresponding to a modified base is connected with at least one adjacent nucleotide by a phosphorothioate bond.

10. The method of any one of claims 1-4, wherein, The preset modification rule comprises: the 5' end and / or the 3' end of the to-be-identified nucleic acid sequence has a modified sequence fragment, the length of the modified sequence fragment is 2 to 5 nucleotides, and adjacent nucleotides in the modified sequence fragment are connected by a phosphorothioate bond.

11. The method of any one of claims 1-4, wherein, Adjacent nucleotides of the theoretical sequence fragment are connected by a phosphorothioate bond.

12. The method of any one of claims 1-4, wherein, The first mass spectrum information comprises first parent ion information and first daughter ion information of the theoretical sequence fragment; wherein the first parent ion information at least comprises mass-to-charge ratio of the first parent ion, and the first daughter ion information at least comprises complementary pairing information and mass-to-charge ratio of the first daughter ion.

13. The method of claim 12, wherein, The second mass spectrum information comprises second parent ion information and second daughter ion information; the obtaining of the second mass spectrum information actually produced by the enzymolysis product comprises: determining the second parent ion information based on the first parent ion information, the second parent ion information at least comprising mass-to-charge ratio of the second parent ion; performing secondary mass spectrum analysis of the enzymolysis product based on the second parent ion information to obtain second daughter ion information, the second daughter ion information at least comprising complementary pairing information and mass-to-charge ratio of the second daughter ion.

14. The method of claim 13, wherein, The determining of whether the target sequence fragment consistent with the theoretical sequence fragment exists in the nucleic acid sequence to be identified comprises: determining whether the first daughter ion information and the second daughter ion information satisfy a preset matching condition; if the preset matching condition is satisfied, determining that the target sequence fragment exists in the nucleic acid sequence to be identified.

15. The method of claim 14, wherein, The preset matching condition is that there are at least one pair of first daughter ions and at least one pair of second daughter ions, and the complementary pairing information and mass-to-charge ratio of the at least one pair of first daughter ions and the at least one pair of second daughter ions are the same.

16. The method of any one of claims 1-4, wherein, The secondary mass spectrum analysis of the enzymolysis product is performed by a linear ion trap mass spectrometer.

Citation Information

Patent Citations

  • Methods for inhibition of apolipoprotein h

    CA2746981A1

  • Method of detecting protein palmitoylation modification locus

    CN109541222A