Method for analyzing the end of a plant organelle rna and use thereof

By performing in vitro circularization and paired-end sequencing of plant organelle RNA, combined with the MeCi algorithm, the problems of low throughput, high cost, and high complexity in the analysis of plant organelle RNA ends in existing technologies have been solved. This has enabled efficient screening of RNase and RBP substrate RNA, improving the accuracy and efficiency of the analysis.

CN120700117BActive Publication Date: 2026-02-24SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510815072.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2026-02-24
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

Existing technologies are insufficient for efficiently analyzing the molecular mechanisms of RNA termination and stabilization in plant organelles. In particular, methods for analyzing mitochondrial and chloroplast RNA terminology suffer from low throughput, high cost, high error rate, and high complexity, making it difficult to effectively screen for RNase and RBP substrate RNAs.

Method used

RNA was treated with 5′ polyphosphatase and then circularized in vitro. High-reliability RNA circularization junctions were screened using paired-end sequencing and the MeCi algorithm, and the positions of the 5′ and 3′ ends of the RNA were determined. The length and origin of poly(A) were also analyzed.

Benefits of technology

This method enables high-throughput, low-cost analysis of plant organelle RNA terminal sequences, and can screen for RNase and RBP substrate RNAs involved in RNA terminal maturation and stabilization, thus improving the accuracy and efficiency of the analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120700117B_ABST
    Figure CN120700117B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of molecular biology experiment, especially to a kind of plant organelle RNA end analysis method and its application.The present application provides a kind of plant organelle RNA end analysis method, comprising the following steps: RNA is treated with 5' polyphosphatase, in vitro cyclization, then double-end sequencing is carried out, filtering, to obtain clean read length;MeCi algorithm is used to identify clean read length;According to the identification result and the number of clean read length, the 5' and 3' end position of RNA, the end type, the end position of mRNA and ncRNA and the 5' end type of mRNA and ncRNA are judged.The method not only analyzes organelle RNA end modification and organelle RNA poly(A) length and source, but also can be used for screening the substrate RNA of RNase and RBP involved in plant organelle RNA end processing and stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of molecular biology experimental technology, and in particular to an analytical method for the ends of RNA in plant organelles and its application. Background Technology

[0002] Mitochondria and chloroplasts are important organelles in higher plants. The former is the energy factory for cellular life activities, while the latter is the main site of photosynthesis. Both mitochondria and chloroplasts contain their own genomes, and the normal expression of these genes is crucial for organelle function and plant growth and development.

[0003] Unlike nuclear genes, the transcription process of chloroplast and mitochondrial genes is relatively loose, and their expression regulation mainly occurs in post-transcriptional processing, including RNA terminal maturation, RNA stabilization, RNA editing, and intron splicing (Holec et al., 2008; Stern et al., 2010; Eckardt, 2012; Castandet et al., 2019; Zhang et al., 2019). Research methods for RNA editing and intron splicing are relatively simple and well-established. In contrast, research on the molecular mechanisms of organelle RNA terminal maturation and stabilization is more challenging and progresses more slowly.

[0004] Plant mitochondria and chloroplasts contain two different types of transcripts: primary transcripts and secondary transcripts. Primary transcripts have a triphosphate group at their 5′ end, originating from transcription initiation; secondary transcripts have a monophosphate group at their 5′ end, primarily originating from post-transcriptional processing (Binder et al., 2016). In contrast to their 5′ ends, the 3′ ends of plant organelle RNAs primarily originate from post-transcriptional processing, rather than transcription termination (Holec et al., 2008).

[0005] It is generally believed that the maturation and stabilization of RNA ends in plant organelles is accomplished through the synergistic action of ribonuclease (RNase) and RNA secondary structures or RNA-binding proteins (RBPs), mainly in two forms: (1) endonucleases directly cleave RNA secondary structures located at the 3′ or 5′ ends, mediating the direct formation of mature ends (Canino et al., 2009; Zhang et al., 2019; MacIntosh and Castandet, 2020); (2) RBPs or RNA secondary structures located upstream of the 3′ end and downstream of the secondary 5′ end form mature ends by preventing cleavage by 3′→5′ and 5′→3′ exonucleases (Hammani et al., 2016; Jung et al., 2023). If RBPs are absent, exonucleases can cleave the 5′ or 3′ ends without restriction, thereby degrading RNA; if the relevant RNases are absent, it will lead to the lengthening of RNA ends. However, unlike the two models mentioned above, plant mitochondria lack 5′→3′ exonucleases, but they can still complete the 5′ end maturation and stabilization of RNA, indicating the existence of other mechanisms for RNA end maturation and stabilization (Ruwe et al., 2016; Zhang et al., 2019).

[0006] Research on the molecular mechanisms of RNA terminal maturation and stabilization in plant organelles hinges on identifying the substrate RNAs of RNases and RBPs. The primary method is analyzing changes in organelle RNA terminals in mutants. Currently, the main methods for determining plant organelle RNA terminal sequences include rapid amplification of cDNA ends (RACE), circularized RT-PCR (cRT-PCR), terminome-seq (Castandet et al., 2019; Zhang et al., 2019b; Jung et al., 2023), and long read sequencing (LRS). Additionally, there is circularized RNA-seq (Kuznetsova et al., 2017) for analyzing mammalian mitochondrial RNA terminals. RACE and cRT-PCR, based on PCR technology, suffer from low throughput and high error rates in large-scale substrate RNA screening, significantly limiting their application for RNA substrate identification. Terminome-seq combines adapter ligation and transcriptome sequencing, which can analyze chloroplast RNA terminal sequences in high throughput, but library construction and data analysis are relatively complex, and it cannot simultaneously determine the 5′ and 3′ ends of RNA molecules. Although LRS can simultaneously determine the 5′ and 3′ ends of RNA, it has high cost and error rate, and it is still uncertain whether it is suitable for screening RNase and RBP substrate RNA. Circularized RNA-seq combines in vitro RNA circularization and circular RNA sequencing (circRNA-seq), and its main problems are: (1) there is no suitable circRNA recognition algorithm, which leads to a complex process of identifying circularization connection points; (2) this system is established using mice as material, while the mammalian mitochondrial genome is very different from the plant mitochondrial and chloroplast genomes (Neupert, 2016). RNA terminal sequence analysis has become a technical bottleneck restricting the study of molecular mechanisms of RNA terminal maturation and stability in plant organelles, and there is an urgent need to establish a technical system for efficiently analyzing the terminal sequences of plant mitochondrial and chloroplast RNA. Summary of the Invention

[0007] The purpose of this invention is to provide a method for analyzing the ends of plant organelle RNA. This method can be combined with RNA 5′ end uncapping enzyme treatment to analyze the 5′ end modification of organelle RNA, as well as the length and origin of plant organelle RNA poly(A). It can also be used to screen substrate RNAs of RNase and RBP involved in the maturation and stabilization of plant organelle RNA ends.

[0008] To achieve the above-mentioned objectives, the present invention provides the following technical solution:

[0009] This invention provides a method for analyzing the ends of RNA in plant organelles, comprising the following steps:

[0010] (1) RNA was treated with RNA 5′ polyphosphatase to obtain the treated RNA;

[0011] (2) The treated RNA was circularized in vitro to obtain circularized RNA.

[0012] (3) Sequencing of in vitro circular RNA to obtain initial read lengths; filtering of initial read lengths to obtain clean read lengths;

[0013] (4) Concatenate the clean read lengths to obtain concatenated read lengths and unconcatenated read lengths; merge the concatenated read lengths and unconcatenated read lengths to obtain merged read lengths;

[0014] (5) Determine the reconstructed genome based on the reference genome;

[0015] (6) Use MeCi to detect in vitro circularized RNA linker sites with merged reads;

[0016] (7) Align the in vitro circular RNA linker sites with the reference genome and the reconstructed genome respectively to obtain the linker site positions aligned with the reference genome and the linker site positions aligned with the reconstructed genome; then convert the linker site positions aligned with the reconstructed genome into the positions of the reference genome and merge them with the data of the reference genome to obtain the positions of the in vitro circular RNA linker sites.

[0017] (8) Screening of in vitro circularized RNA linker sites to obtain high-confidence RNA in vitro circularized linker sites and determine their positions; based on the positions of the high-confidence RNA in vitro circularized linker sites, determine the positions of the 5′ and 3′ ends of the corresponding RNA molecules and infer the full length of the RNA molecules;

[0018] (9) Clustering RNA molecules containing high-confidence RNA in vitro circularization junctions yields the abundance of 5′ and 3′ ends and the number of reads;

[0019] (10) Determine the 5′ end type of RNA based on the number of clean reads and the number of 5′ and 3′ end reads;

[0020] (11) Determine the terminal positions of mRNA and ncRNA based on the number of clean reads, the number of 5′ and 3′ end reads, and the annotation information of the reference genome;

[0021] (12) Cluster the 5′ end types of RNA to obtain the cumulative abundance of the 5′ end types of RNA; determine the end types of mRNA and ncRNA based on the cumulative abundance of the 5′ end types of RNA.

[0022] Preferably, in step (3), the sequencing method is paired-end sequencing.

[0023] Preferably, in step (3), the filtering method is to remove sequencing adapters and low-quality reads from the initial read length;

[0024] The low-quality reads are those with a proportion of bases that cannot be determined exceeding 10% and those with a length less than 30 bp.

[0025] Preferably, in step (4), the unjoined read length includes the reverse complementary sequence of the unjoined 3′ end read length and the unjoined 5′ end read length.

[0026] Preferably, in step (5), the method for determining the reconstructed genome is as follows: for genes in the reference genome containing backsplicing introns, their exon sequences, as well as the sequences 2000-nt upstream of the first intron and 1000-nt downstream of the last exon, are spliced ​​together to form the reconstructed genome.

[0027] Preferably, the screening criteria in step (8) are as follows:

[0028] A. Retrieve join points that have been aligned to the reference genome and remove join points containing ≥3 alignment positions;

[0029] B. For join points with two alignment positions, if both contain the same overlapping bases or non-coding bases, both are retained. If the number of non-coding bases is different and the difference is >10bp, the join point with the fewer non-coding bases is retained. If the number of non-coding bases is different and the difference is ≤10bp, both are removed.

[0030] C. For single or multiple alignment sites containing 1 to 3 non-coding bases, all are retained. For alignment sites containing 4 to 100 non-coding bases, only the portion of the non-coding bases with an A base ratio ≥ 70% is retained.

[0031] D. All single alignment join points containing 1 to 3 repeating bases or neither repeating bases nor non-coding bases are retained;

[0032] E. Based on the above filtering of different types of linker points, only in vitro circularized linker points that appear in at least two sequencing samples are retained, i.e., those containing high-confidence in vitro circularized linker points.

[0033] Preferably, in step (10), the method for judging the 5'-end type of the RNA is as follows: calculate the RPTM value according to the read length and clean read length at the 5' and 3' ends; calculate the FCV value according to the RPTM value; judge the 5'-end type of the RNA according to the FCV value.

[0034] The calculation formula of the RPTM value is the read length at the end / the clean read length.

[0035] The judgment standard for the 5'-end type of the RNA is as follows: if FCV≥5, it is considered that the 5'-end contains a triphosphate group, marked as "primary", from transcription initiation; if FCV≤1, it is considered that the 5'-end contains a monophosphate group, marked as "secondary", from post-transcriptional processing; if 1 < FCV < 5, it is impossible to determine whether the 5'-end contains a monophosphate group or a triphosphate group, marked as "uncertain".

[0036] Preferably, in step (11), the judgment standard for the 5'-end position of the mRNA is as follows: the 5'-end needs to be located upstream of the translation initiation codon and RPTM≥10.

[0037] The judgment standard for the 5'-end position of the ncRNA is as follows: RPTM≥50 and the distance to its corresponding 3' is between 300 and 3000 bases.

[0038] The RPTM is calculated according to the read length and clean read length at the 5' and 3' ends.

[0039] The calculation formula of the RPTM value is the read length at the end / the clean read length.

[0040] Preferably, in step (12), the judgment standard for the 5'-end type of the mRNA and ncRNA is as follows: if the cumulative abundance of the 5'-end type of the RNA marked as "primary" in the 5'-end interval of the mRNA or ncRNA≥70%, it is considered primary; if the cumulative abundance of the 5'-end type of the RNA marked as "secondary" in the 5'-end interval of the mRNA or ncRNA≥70%, it is considered secondary; otherwise, it is "uncertain".

[0041] The present invention also provides the application of the above analysis method in screening substrate RNAs of RNase and RBP involved in the processing and stabilization of plant organelle RNA ends.

[0042] The beneficial effects of the present invention:

[0043] This invention establishes a high-throughput transcript end mapping system (Hiten) for analyzing the terminal sequences of plant organelle RNA. It is based on the circular RNA research approach for linear RNA end analysis. First, organelle RNA is circularized in vitro. Then, the circularized organelle RNA is sequenced using next-generation sequencing technology. Finally, the circular RNA recognition algorithm MeCi is used to screen organelle RNAs containing in vitro circularization junctions from the high-throughput sequencing data. Hiten can be combined with RNA 5′ end uncapping enzyme treatment to analyze 5′ end modifications of organelle RNA. Hiten can be used to screen substrate RNAs of RNases and RBPs involved in the maturation and stabilization of plant organelle RNA ends. Hiten can also analyze the poly(A) length and origin of plant organelle RNA. Attached Figure Description

[0044] Figure 1 A flowchart of a method for analyzing the RNA ends of plant organelles;

[0045] Figure 2 The results of PCR-PCR and RNA hybridization were used to verify the terminal positions of maize organelle mRNAs and ncRNAs located by Hiten; Figure A shows the chloroplast psbE-psbF-psbL-psbJ mRNAs and mitochondrial nad4L, nad9, and cox3 identified by Hiten. The distribution of the 5′ and 3′ ends of mRNA is shown in Figure B. The blue and red peaks represent the 5′ and 3′ ends identified by Hiten, respectively. The peak height is expressed as the average RPTM (number of reads per 10 million reads). Figure B shows the agarose gel electrophoresis results of cRT-PCR products (chloroplasts). + and - represent the 5′ polyphosphatase treatment group and the control group, respectively. The maize chloroplast 16S rRNA equilibration treatment group and the control group were amplified. The bands indicated by the arrows were recovered, cloned, and sequenced by Sanger sequencing. Figure C shows the agarose gel electrophoresis results of cRT-PCR products (mitochondria). Figure D shows the cumulative mRNA of chloroplast psbE-psbF-psbL-psbJ and psbE-psbF-psbL through RNA hybridization analysis. Figure E shows the cumulative mRNA of mitochondrial nad4L, nad9, and cox3 through RNA hybridization analysis.

[0046] Figure 3 A comparison of Hiten and cRT-PCR at the omics level to identify the number of 5′ and 3′ ends of mitochondrial mRNA;

[0047] Figure 4Figure 1 shows the subcellular localization of the ZmPPR67 protein and the phenotype of its deletion mutant. Figure 2 shows the ZmPPR67 gene structure and the Mu transposon mutation site. Figure 3 shows the ZmPPR67 protein structure SP: the mitochondrial signal peptide predicted by TargetP, and P: the PPR motif predicted by PROSITE. Figure 4 shows the ZmPPR67 protein subcellular localization, ZmPPR67-GFP is the fusion protein of ZmPPR67 and GFP, mitoTracker is the mitochondrial marker, Merged is the superimposed image of GFP and mitoTracker signals, White field indicates the field image, scale bar = 10 μm. Figure 5 shows the self-pollinated ears of ZmPPR67 heterozygotes, with black arrows indicating mutant grains, scale bar = 1 cm.

[0048] Figure 5 The results of Hiten's identification of substrate RNA for the maize mitochondrial localization protein ZmPPR67 are shown in Figure A. Figure A shows the distribution of the 5′ and 3′ ends of atp9 mRNA detected by Hiten, with blue and red peaks representing the 5′ and 3′ ends identified by Hiten, respectively. The peak height is expressed as the average RPTM (number of reads per 10 million reads). Figure B shows the formation of atp9 mRNA ends in the zmppr67 mutant using cRT-PCR analysis. Mitochondrial 26S rRNA, as well as nad9 and atp8 mRNA, were used as controls. Figure C shows the accumulation of atp9 transcripts in the zmppr67 mutant using RNA hybridization. Different types of atp9 transcripts were detected using P1, P2, and P3 probes. 26S rRNA was used as an internal control, and the loading amounts of mutant and wild-type RNA were balanced with cytoplasmic rRNA. Detailed Implementation

[0049] The technical solutions provided by the present invention will be described in detail below with reference to the embodiments, but they should not be construed as limiting the scope of protection of the present invention.

[0050] The materials used in the embodiments of this invention are sourced from the following sources:

[0051] RNA extraction reagents: TRIzol™ reagent (Thermo Fisher Scientific, USA); RNA ligase: T4RNAligase 1 (NEB, USA); RNA 5′ polyphosphatase purchased from Epicentre (Epicentre, USA); RNA purification and recovery kit: RNA recovery kit (Tiangen Biotech, China); RiboMinus kit purchased from Thermo Fisher Scientific (USA); NEBNext Ultra II RNA strand-specific library preparation kit purchased from NEB (USA);

[0052] DNA polymerase: 2×PhantaMax Master Mix (NovaZeneca Biotechnology Co., Ltd., China); Reverse transcriptase: PrimeScript TM II Reverse Transcriptase (Baori Biotechnology Co., Ltd., China); RNA Hybridization Kit: Roche DIG Northern Starter Kit (Sigma-Aldrich Chemical Co., Ltd., USA); DNA Purification and Recovery Kit (Tiangen Biotech Co., Ltd., China); Probe Vector: pEASY-Blunt (Beijing TransGen Biotechnology Co., Ltd., China); Cloning Vector: pMD19-T (Baori Biotechnology Co., Ltd., China); Competent Cells: Escherichia coli DH5α competent cells and Agrobacterium EHA105 competent cells (Shanghai Weidi Biotechnology Co., Ltd., China); DEPC purchased from Sigma-Aldrich Chemical Co., Ltd. (USA); RNA Hybridization Membrane: Positively Charged N+ Nylon Membrane: HybondN+ (GE, USA); TRIzol Reagent purchased from Invitrogen (USA);

[0053] DEPC treatment of water uses deionized water as a solvent and includes a final concentration of 1% DEPC.

[0054] The 10×MOPS electrophoresis buffer uses DEPC-treated water as a solvent and comprises the following components at final concentrations: 0.2M MOPS, 0.05M sodium acetate (NaOAc), 10mM EDTA, pH 7.0;

[0055] 20×SSC buffer uses deionized water as a solvent and includes the following components at final concentrations: 3M sodium chloride, 0.3M sodium citrate, 1% DEPC, pH 7.0;

[0056] The formaldehyde agarose gel with DEPC-treated water as solvent comprises the following components at final concentrations: 1.5% agarose, 2.2M 37% formaldehyde solution, and 1×10×MOPS.

[0057] The washing solution, using DEPC-treated water as a solvent, comprises the following components at final concentrations: 0.1M maleic acid, 0.15M sodium chloride, and 0.3% Tween 20.

[0058] The test solution, using DEPC-treated water as a solvent, comprises the following components at final concentrations: 0.1M Tris-HCl, 0.15M sodium chloride, and 1×10×MOPS.

[0059] Example 1: Analysis method for RNA ends of plant organelles (flowchart shown) Figure 1As shown in the figure, the red "P" represents the primary 5' end, and the blue "P" represents the secondary 5' end.

[0060] The genetic background of the maize material is the B73 inbred line, which comes from the fresh maize breeding laboratory of South China Agricultural University. The maize seedlings were planted in an artificial climate chamber and cultured under the following conditions: light culture: 25℃, 12h light / 12h dark; dark culture: 25℃, 0h light / 24h dark.

[0061] (1) Using TRIzol TM Total RNA was extracted from leaves of maize seedlings cultured under light and dark conditions using reagents. Three biological replicates were set up for seedlings cultured under light and two biological replicates were set up for seedlings cultured under dark conditions.

[0062] The total RNA of leaves was treated with RNA 5′ polyphosphatase (5′ polyphosphatase treatment group), which converted the triphosphate group at the 5′ end of organelle RNA into a monophosphate group to obtain the treated RNA. At the same time, the total RNA of untreated leaves was used as the control group.

[0063] The reaction system for treating total RNA from leaves with RNA 5′ polyphosphatase is shown in Table 1. The reaction conditions were 37℃ and incubation for 30 min. The resulting reaction solution was purified and recovered using an RNA purification kit for subsequent in vitro RNA circularization.

[0064] Table 1. Reaction system for treating total RNA in leaves with RNA 5′ polyphosphatase.

[0065]

[0066]

[0067] (2) In vitro cyclization

[0068] T4 RNA ligase was used to circularize total RNA from treated and untreated leaves in vitro to obtain circularized RNA.

[0069] The in vitro circularization reaction system is shown in Table 2. The reaction conditions were 16℃ and incubation for 15h. The obtained reaction solution was purified and recovered using an RNA purification kit for subsequent sequencing and cRT-PCR verification.

[0070] Table 2 Reaction system for in vitro RNA circularization

[0071] Components Dosage (μL) reagent concentration Final concentration T4RNAligase1 1.6 30U / μL 1.2 U / μL 10×lightbuffer 4 10× 1× RNase inhibitor 1 40 U / μL 1U / μL PEG8000 1.6 50% 20% TemplateRNA -- 8μg 0.2 μg / L Total Add to 40 μL -- --

[0072] (3) Sequencing

[0073] 1) Use the RiboMinus kit to remove rRNA from the in vitro circularized RNA, and use the RNA purification kit to purify and recover the RNA to obtain the in vitro circularized RNA after removing rRNA;

[0074] 2) Under ultrasound at a power of 150W, a time of 5 minutes, and a temperature of 4℃, the in vitro circularized RNA after removing rRNA was fragmented, and the fragmented RNA fragments were separated by agar electrophoresis, and the portion with a length of 250-350 bases was recovered.

[0075] 3) The fragmented RNA was constructed into a chain-specific library using the NEBNext Ultra II RNA chain-specific library preparation kit; the constructed RNA was sequenced at both ends using an Illumina Hiseq 4000 sequencing platform (Illumina, San Diego, CA, USA), with each read being 150 bp (PE150) long to obtain the initial read length (Tables 3-4).

[0076] 4) Use FastQC to assess the quality of the initial reads (without filtering); use Cutadapt to remove sequencing adapters and low-quality reads (low-quality reads are reads with more than 10% of unidentifiable bases and reads with a length of less than 30bp) from the initial reads to obtain clean reads (Table 3-4). When using Cutadapt for filtering, m is set to 50.

[0077] (4) Merge read lengths

[0078] 1) Use FLASH to concatenate the clean read length to obtain the concatenated read length and the unconcatenated read length: set FLASH's m to the default value;

[0079] 2) Use Seqkit to extract the reverse complementary sequence from the unassembled 3′ end sequencing read;

[0080] 3) Integrate the spliced ​​read, the unspliced ​​5′ end sequencing read, and the reverse-complementary 3′ end sequencing read to obtain the merged read;

[0081] (5) Determine the reconstructed genome based on the reference genome

[0082] Because some organelle genes contain trans-splicing introns, read lengths from these genes are lost. Therefore, it is necessary to reconstruct the exon sequences of genes containing trans-splicing introns. MeCi is used to compare these reconstructed genome sequences separately to ensure the integrity of the results. For genes in the reference genome containing reverse-splicing introns, their exon sequences, as well as the sequences 2000-nt upstream of the first intron and 1000-nt downstream of the last exon, are spliced ​​together to form a reconstructed genome.

[0083] (6) MeCi algorithm

[0084] The MeCi algorithm was used to detect reads containing in vitro circularized RNA linker sites in the merged reads. The OP parameters for MeCi RNA molecule detection were set to -100 to +3, where "+" indicates that there are overlapping bases at the linker site, and "-" indicates that the detected linker site contains bases that cannot be aligned to the gene, i.e., non-encoded bases.

[0085] (7) Align the in vitro circular RNA linker sites with the reference genome and the reconstructed genome respectively to obtain the linker site positions aligned with the reference genome and the linker site positions aligned with the reconstructed genome; then convert the linker site positions aligned with the reconstructed genome into the positions of the reference genome and merge them with the data of the reference genome to obtain the positions of the in vitro circular RNA linker sites.

[0086] (8) Screening of in vitro circularized RNA linker sites to obtain high-confidence RNA in vitro circularized linker sites and determine their positions (Tables 3-4); based on the positions of the high-confidence RNA in vitro circularized linker sites, determine the positions of the 5′ and 3′ ends of the corresponding RNA molecules (Tables 5-6), and infer the full length of the RNA molecules;

[0087] The specific steps of filtering are as follows:

[0088] A. Retrieve join sites that have been aligned to the genome and remove join sites containing ≥3 alignment positions;

[0089] B. For join points with two alignment positions, if both contain the same overlapping bases or non-coding bases, both are retained. If the number of non-coding bases is different and the difference is >10bp, the join point with the fewer non-coding bases is retained. If the number of non-coding bases is different and the difference is ≤10bp, both are removed.

[0090] C. For single or multiple alignment linkage sites containing 1 to 3 non-coding bases (i.e., OP = -3 to -1), all are retained. For linkage sites containing 4 to 100 non-coding bases (i.e., OP = -100 to -4), only the portion of non-coding bases with an A base ratio ≥ 70% is retained.

[0091] D. All single alignment join points containing 1 to 3 repeating bases (OP = +1 to +3) or neither repeating bases nor non-coding bases (OP = 0) are retained;

[0092] E. On the basis of filtering different types of junction points as described above, only in vitro circularized junction points that appear in at least two sequencing samples are retained, that is, high-confidence junction points;

[0093] Table 3 Statistical data of light-grown seedling identification

[0094]

[0095] Table 4 Statistical data of dark-grown seedling identification <s

[0096]

[0097] The results showed that using Hiten, in the three biological replicates of the control group of maize light-grown seedlings, 1,033,055, 939,610 and 946,096 chloroplast reads containing high-confidence in vitro circularized junction points were identified respectively, while in the three biological replicates of the treatment group, 1,367,816, 829,699 and 814,473 chloroplast reads containing high-confidence in vitro circularized junction points were identified respectively; in the two biological replicates of the control group of maize dark-grown seedlings, 231,287 and 237,709 mitochondrial reads containing high-confidence in vitro circularized junction points were identified respectively, while in the three biological replicates of the treatment group, 218,887 and 286,241 mitochondrial reads containing high-confidence in vitro circularized junction points were identified respectively.

[0098] (9) Cluster RNA molecules containing high-confidence RNA in vitro circularized junction points to obtain the number of 5′ and 3′ end reads;

[0099] (10) Judgment of the 5′ end type of RNA

[0100] According to the number of clean reads, balance the number of reads corresponding to the 5′ and 3′ ends to obtain the RPTM value of each 5′ and 3′ end, where the RPTM value = the number of reads at the end / the number of clean reads;

[0101] Compare the RPTM values of the corresponding 5′ ends of the 5′ polyphosphate treatment group and the control group (i.e., treatment group RPTM / control group RPTM) to obtain the FCV value; if FCV ≥ 5, it is considered that the 5′ contains a triphosphate group, labeled as "primary", from transcription initiation; if FCV ≤ 1, it is considered that it contains a monophosphate group, labeled as "secondary", from post-transcriptional processing; if 1 < FCV < 5, it is impossible to determine whether it contains a monophosphate group or a triphosphate group, labeled as "uncertain";

[0102] (11) Judgment of the terminal positions of mRNA and ncRNA:

[0103] The terminal positions of organelle mRNAs and ncRNAs are often distributed within a certain range, with the corresponding 5′ and 3′ ends typically representing a range of values. The terminal positions of mRNAs and ncRNAs can be determined based on the number of clean reads, the number of 5′ and 3′ end reads, and reference genome annotation information.

[0104] The criteria for determining the 5′ position of the mRNA are: the 5′ end must be located upstream of the translation start codon and RPTM≥10;

[0105] The criteria for determining the 5′ end position of the ncRNA are: RPTM ≥ 50 and the distance between it and its corresponding 3′ end is between 300 and 3000 bases; the RPTM is calculated based on the number of reads at the 5′ and 3′ ends and the number of clean reads; the formula for calculating the RPTM value is the number of reads at the end / the number of clean reads.

[0106] (12) Determining the 5′ end type of mRNA and ncRNA:

[0107] Because the 5′ ends of many organelle mRNAs and ncRNAs are distributed within a single region, the 5′ ends within this region may simultaneously be of primary, secondary, and indeterminate types. For the 5′ ends of these mRNAs and ncRNAs, it is impossible to distinguish whether they originate from transcription initiation or post-transcriptional processing; it can only be stated whether they are predominantly primary or secondary transcripts. In contrast, for a 5′ end located at a fixed position, its type can be determined.

[0108] Clustering of mRNA and ncRNA terminal types yields the cumulative abundance of RNA terminal types; the terminal types of mRNA and ncRNA are then determined based on the abundance of RNA terminal types.

[0109] The criteria for determination are as follows: if the cumulative abundance of the 5′ end region of the mRNA or ncRNA containing RNA labeled "primary" is ≥70%, it is considered primary; if the cumulative abundance of the 5′ end region of the mRNA or ncRNA containing RNA labeled "secondary" is ≥70%, it is considered secondary; otherwise, it is "uncertain".

[0110] The terminal positions and 5′ end types of the identified mRNA and ncRNA are shown in Tables 5-6.

[0111] Table 5. Terminal positions and types of mRNA and ncRNA in seedlings exposed to light.

[0112]

[0113]

[0114]

[0115]

[0116]

[0117]

[0118]

[0119]

[0120]

[0121] Table 6. Terminal positions and types of mRNA and ncRNA in dark seedlings.

[0122]

[0123]

[0124]

[0125]

[0126]

[0127] The results showed that under light culture, 108 mRNAs and 13 ncRNAs were detected in maize chloroplasts; under dark culture, 37 mRNAs and 23 ncRNAs were detected in maize mitochondria. These mRNAs included both precursor and mature mRNAs, while the ncRNAs included tRNAs, mRNA processing products, and transcripts from intergenic regions. The 5′ ends of chloroplast mRNAs and ncRNAs were distributed in a relatively small region (1–10 nts); in contrast, the 5′ ends of mitochondrial mRNAs and ncRNAs were mostly distributed in a larger region, some even exceeding 100 nts.

[0128] Therefore, in chloroplasts and mitochondria, 10 ± 0.26% and 7 ± 0.01% of RNA molecules, respectively, contain poly(A) sequences. The length of poly(A) in maize organelles ranges from 1 to 100 nts, with poly(A) tails of 1 to 7 nts accounting for over 95% of all poly(A) sequences; their proportion gradually decreases as the length of the poly(A) increases.

[0129] Example 2 Result Verification

[0130] (1) cDNA first-strand synthesis

[0131] Using PrimeScript TM II. Reverse transcriptase and random primers were used to reverse transcribe first-strand cDNA using in vitro circular RNA from the 5′ polyphosphatase-treated group and the control group, respectively. The reverse transcription reaction system is shown in Table 7.

[0132] Table 7 Reverse Transcription Reaction System

[0133] Components Final concentration volume random primers 50uM 1.25μl dNTP mixture 10mM 1.25μl template RNA 200ng 1μl RNAse inhibitors 40μ / μl 0.5μl reverse transcriptase 200μ / μl 1.25μl 5x Reverse Transcription Buffer 1x 5.0μl RNAse-deoxidizing water / 14.75 total / 25μl

[0134] Reaction conditions: incubate at 30℃ for 10 min, at 42℃ for 45 min, and at 70℃ for 15 min. The reaction solution was used directly for PCR amplification.

[0135] (2) Design of divergent PCR primers

[0136] Based on the terminal position located by Hiten, divergent primers were designed upstream and downstream of the in vitro looping junction (Table 8).

[0137] Table 8. Dispersion Primers

[0138]

[0139]

[0140] (4) PCR amplification

[0141] Using cDNA as a template and the above-mentioned divergent primers, fragments containing in vitro circularization junctions were amplified by PCR; the PCR reaction system and reaction conditions are shown in Tables 9 and 10, respectively.

[0142] Table 9 PCR Reaction System

[0143]

[0144]

[0145] Table 10 PCR Reaction Conditions

[0146]

[0147] (5) Isolation, recovery and cloning of PCR products

[0148] PCR products were separated by agarose gel electrophoresis; the expected bands were purified and recovered, ligated into the pMD19-T vector, and transformed into E. coli DH5α competent cells; transformed bacteria were screened under suitable culture conditions, and single clones were selected for PCR identification; clones carrying the target fragment were identified and sent to a commercial sequencing company for sequencing to determine the inserted fragment sequence.

[0149] (6) Determination of cyclization sites in PCR amplification products

[0150] The NCBI BLASTn (https: / / blast.ncbi.nlm.nih.gov) online alignment tool was used to compare and analyze the sequenced sequences with the reference genome sequence. Based on the alignment results, it was determined whether the insert fragment of the monoclonal molecule contained an in vitro circularization junction, and thus the end position of the corresponding linear RNA molecule was determined (Table 11).

[0151] Table 11. Single-clone sequencing results of CT-PCR products

[0152]

[0153]

[0154]

[0155]

[0156]

[0157]

[0158]

[0159]

[0160]

[0161]

[0162] (7) RNA hybridization experiment

[0163] RNA hybridization experiments were performed using the DIG Northern Starter Kit. The RNA samples used in the experiments were total RNA from leaves of maize seedlings cultured under light or in the dark. The amplification primers for the DNA fragments used to prepare RNA probes are shown in Table 12.

[0164] Table 12 cRT-PCR primer information.

[0165]

[0166] The PCR product was cloned into the pEASY-Blunt vector, and an RNA probe labeled with DIG-11-UTP was used in vitro. The main experimental steps are as follows:

[0167] 1) Dissolve 1.5g of agarose in 62ml of DEPC-treated water by heating, then add 20ml of 10×MOPS electrophoresis buffer and 18ml of formaldehyde to final concentrations of 1× and 2.2M, respectively. In a fume hood, pour the melted gel into a gel casting mold, ensuring a gel thickness of 3–5mm.

[0168] 2) Remove the gel, place it in the electrophoresis tank, and add 1×MOPS electrophoresis buffer until the liquid level is above the gel.

[0169] 3) Electrophoresis for 4 hours to allow the RNA samples to be fully separated on a 1.5% (w / v) denaturing formaldehyde agarose gel.

[0170] 4) After electrophoresis, remove the gel and rinse it to remove formaldehyde.

[0171] 5) Use capillary blotting to transfer RNA from the gel to a positively charged nylon membrane. Specifically: Select a plate larger than the gel, place a filter paper on it, and soak both sides of the filter paper in 20X SSC buffer; place the gel on the filter paper and remove any air bubbles between the two layers; lay the pre-wetted nylon membrane flat on the gel and cover it with two sheets of filter paper soaked in 2×SSC buffer; place a stack of dry absorbent paper on the filter paper, apply a weight, and transfer the membrane for 16 hours.

[0172] 6) After transfer, the nylon membrane is baked at 120℃ for 30 min to fix the RNA onto the membrane.

[0173] 7) The nylon membrane containing RNA was placed in a hybridization buffer and pre-hybridized at 68°C for 4 hours.

[0174] 8) Replace the hybridization buffer, add the RNA probe, and hybridize at 68°C for 16 hours.

[0175] 9) Remove the hybridization buffer, add 2×SSC+0.1%SDS buffer and wash twice, shaking for 5 minutes each time at room temperature.

[0176] 10) Wash twice with 1×SSC + 0.1% SDS buffer, shaking for 15 min at 68°C each time.

[0177] 11) Briefly rinse the HybondN+ membrane with washing solution for 5 min.

[0178] 12) Incubate in the sealing solution for 2 hours.

[0179] 13) Replace the blocking solution, add 1 / 1000 of the anti-DIG antibody, and incubate for 40 minutes.

[0180] 14) The nylon membrane was washed twice in the washing solution for 15 minutes each time.

[0181] 15) Take a petri dish, add the detection solution, place the nylon membrane inside, and soak for 5 minutes.

[0182] 16) Place the nylon film in the developing chamber, add CDP-Star, and incubate at room temperature for 5 minutes.

[0183] 17) Protein Blot Imaging System (Amersham ImageQuant) TM (800, Amersham, USA) Take a picture and observe the DIG signal.

[0184] (8) Fifteen organelle mRNAs (chloroplast: 12; mitochondria: 3) and four ncRNAs (chloroplast: 1; mitochondria: 3) were randomly selected to verify the terminal position and / or 5′ end type of Hiten (results are shown in the figure). Figure 2 (B, C); To differentiate between primary and secondary RNA 5′ end types, chloroplast 16S rRNA with secondary 5′ ends was used to balance the treatment and control groups. Based on this, PCR amplification of petN, psaJ, psbA, psbM, ccsA, and tncRNA1 was performed using divergent primers. The results showed that the abundance of PCR amplification products from these mRNAs or ncRNAs was significantly higher in the treatment group than in the control group, indicating that they have primary 5′ ends. In contrast, the PCR amplification levels of atpH-atpF, atpH-atpF-atpA, ycf3, rpl16-rpl14, ndhB, and ndhA in the control group were comparable to or higher than those in the treatment group, indicating that these mRNAs have secondary 5′ ends. The RNA 5′ end types determined by cRT-PCR were consistent with Hiten's results.

[0185] The results of PCR amplification product recovery, cloning, and Sanger sequencing (Table 11) showed that the terminal positions of these organelle mRNAs and ncRNAs located by PCR were all located in the terminal position region of Hiten localization.

[0186] Based on this, five mRNAs were randomly selected (chloroplasts: psbE-psbF-psbL-psbJ and psbE-psbF-psbL; mitochondria: nad4L, nad9, and cox3), and RNA hybridization was used to validate the Hiten results. Figure 2 A, D, E). Hiten results show ( Figure 2A) The 5′ ends of chloroplasts psbE-psbF-psbL-psbJ and psbE-psbF-psbL are located at 64256–64252, while the 5′ ends of the three mitochondrial mRNAs are distributed in a larger region and contain multiple enriched regions (nad4L: 226686–26709, 26788–26799 and 26817–26829; nad9: 100730–100735, 100979–100981, 100982 and 100992–101000; cox3: 443278, 442732–442729, 442688–442662 and 442643–442635). RNA hybridization experiments using the psbE probe detected two bands, the sizes of which were close to Hiten's predicted psbE-psbF-psbL-psbJ and psbE-psbF-psbL, at 991–1008 nts and 708–718 nts, respectively. Figure 2 D). Probes from all three mitochondrial genes detected diffuse bands, the size and band pattern of which were consistent with Hiten's results. Figure 2 E).

[0187] The above cRT-PCR and RNA hybridization results verified the accuracy of the terminal positions of organelle mRNA and ncRNA located by Hiten.

[0188] Example 3: Comparison of the application effects of Hiten and existing technologies

[0189] The terminal positions of maize mitochondrial mRNAs (including mature mRNAs and precursor mRNAs) were systematically analyzed at the omics level using the method described in Example 1 and cRT-PCR (Zhang et al., 2019). Hiten and cRT-PCR identified the number of 5′ and 3′ terminals of mitochondrial mRNAs at the omics level, as shown below. Figure 3 As shown;

[0190] The results showed that cRT-PCR identified 37 mitochondrial mRNAs, including 61 specific 5′ ends and 41 specific 3′ ends (Zhang et al., 2019); the Hiten method identified 38 mitochondrial mRNAs, including 71 specific 5′ ends and 38 specific 3′ ends. Of the 109 mRNA ends identified by the Hiten method, 87 (5′ ends: 53 / 71, 3′ ends: 34 / 38) were consistent with previously reported results, and 22 (5′ ends: 18 / 71, 3′ ends: 4 / 38) were newly identified ends; among them, nad1T2 was a newly identified precursor mRNA by the Hiten method.

[0191] cRT-PCR is based on PCR technology to locate the ends of RNA. However, because PCR technology is prone to errors, even a very small amount of template in the substrate can be amplified. Therefore, when cRT-PCR is used to locate the ends of mRNA at the omics level, it needs to be combined with RNA hybridization to determine the size of the target mRNA. In contrast, the Hiten method is simpler and more comprehensive in locating the ends of mRNA.

[0192] In summary, for systematically locating the RNA terminus of plant organelles at the omics level, the Hiten method is more comprehensive and accurate than the cRT-PCR method, and the operation process is relatively simple.

[0193] Example 4: Application of Hiten to identify substrate RNA of maize mitochondrial localization protein ZmPPR67

[0194] (1) RNA extraction, cDNA synthesis, PCR amplification

[0195] Total RNA from maize Zmppr67 mutant and wild-type kernels was extracted from embryonic and endosperm materials of 15DAP kernels. Before extraction, the seed coat was carefully peeled off, leaving only the embryonic and endosperm tissues. RNA extraction, cDNA synthesis, and PCR amplification were performed as described above.

[0196] (2) Subcellular localization

[0197] 1) Construction of a p1300 vector containing the ZmPPR67-GFP fusion gene.

[0198] Using cDNA as a template, convergent primers were designed based on the ZmPPR67 gene sequence, starting from the start codon ATG (excluding the stop codon). The forward primer sequence is: CTGCAGGGGCCCGGGaagcttATGCCCCCGCCTCGACTCCG, SEQ ID NO. 61; the reverse primer sequence is: CTCGGTACCGGATCCgtcgacTCCAGTTGGTGGGAGTAAAG, SEQ ID NO. 61. (ID NO. 62) Gene fragments were amplified by PCR. The PCR products were separated by agarose gel electrophoresis, and the target-sized fragments were recovered and purified. The p1300-GFP vector was double-digested with restriction endonucleases HindIII and SpeI. After agarose gel electrophoresis, the linearized vector fragments were recovered using a DNA recovery kit. The ZmPPR67 gene fragment was ligated to the linearized p1300-GFP vector using homologous recombinase. The homologous recombinant product was transformed into competent E. coli cells, and the bacterial culture was plated on LB solid medium containing kanamycin and incubated at 37°C for 16 h. Single clones were identified by colony PCR, and positive clones were sent to a commercial company for sequencing. Plasmids were extracted from positive clones containing the correct sequence.

[0199] 2) Agrobacterium transformation and culture

[0200] The constructed ZmPPR67-GFP recombinant plasmid was transformed into Agrobacterium competent cells, and the bacterial culture was plated on LB solid medium containing kanamycin and rifampin and cultured at 28°C for 3 days. Single-clone sequencing was performed to identify positive clones. A single colony of positive Agrobacterium was picked and inoculated into 5 mL of LB liquid medium containing kanamycin and rifampin and cultured overnight at 28°C with shaking at 250 rpm. The overnight culture was then transferred to 25 mL of fresh LB liquid medium at a ratio of 1:25 and cultured until OD600 = 0.6–0.8.

[0201] 3) Tobacco leaf contamination

[0202] Transfer the cultured Agrobacterium to centrifuge tubes and centrifuge at 5000 rpm for 10 min at 4°C to collect the bacterial cells. Discard the supernatant, resuspend the bacterial cells in infection buffer, and adjust the OD600 to 0.5. Incubate the resuspended Agrobacterium at room temperature for 3 h to fully activate the Agrobacterium. Using a 1 mL syringe (without the needle), slowly inject the bacterial solution into the lower epidermis of the tobacco leaf, ensuring even distribution of the solution within the leaf. Three different injection sites can be selected on each leaf.

[0203] 4) Post-infection treatment and result observation

[0204] After injection, tobacco plants are placed in a light incubator and cultured under suitable temperature (22-25℃), humidity (60-80%), and light conditions (16h light / 8h darkness) for 5 days. Leaf tissue around the injection site is selected to prepare temporary sections or epidermal cells are directly excised, placed on a glass slide, and an appropriate amount of distilled water or buffer solution is added and covered with a coverslip. Fluorescence signals at different time points over 36 hours are observed using a laser confocal microscope, and the time with the best effect is selected to take pictures and record the results.

[0205] (3) Experimental Results and Verification

[0206] 1) Using the method described in Example 1 (Hiten) to screen candidate substrates for ZmPPR67.

[0207] The ZmPPR67 gene encodes a P-type PPR protein located in mitochondria. Figure 4 AC); the Zmpp67 null mutation leads to maize kernel development arrest and embryo lethality. Figure 4 D);

[0208] Hiten results showed that in the wild type, atp9 mRNA had three major 5′ ends, enriched on chromosomes at locations 300098–300081, 299892–299867, and 299722–299694, respectively. Figure 5 A). In the Zmppr67 mutant, the abundance of these 5′ ends was significantly reduced, and atp9 transcripts lacking the 5′ end accumulated abnormally. Notably, the changes in the mitochondrial RNA ends were limited to the 5′ end of atp9, and Hiten analysis did not detect significant changes in the 5′ ends of other mitochondrial mRNAs.

[0209] 2) Verify Hiten results using cRT-PCR and RNA hybridization.

[0210] The ZmPPR67 substrate RNA identified by Hiten was verified using cRT-PCR and RNA hybridization.

[0211] PCR and agarose gel electrophoresis results showed that the accumulation level of mitochondrial atp9 mRNA was significantly reduced in the Zmppr67 mutant. Figure 5 B); as a negative control, mitochondrial nad9, atp8, and 26S rRNA showed no significant changes in either the mutant or wild-type; RNA hybridization analysis further confirmed that the three major banding patterns observed in the wild-type were significantly reduced in the mutant, while atp9 transcripts lacking the 5′ end showed a significant accumulation. Figure 5C); These results indicate that the stability of atp9 mRNA is significantly reduced in the Zmppr67 mutant, thus verifying the accuracy of Hiten's results.

[0212] Based on the above results, the stability of atp9 mRNA is impaired in the Zmppr67 mutant, and the ZmPPR67 protein may play a role in stabilizing RNA by protecting its 5′ end.

[0213] As can be seen from the above embodiments, the present invention provides a method for analyzing the ends of plant organelle RNA and its applications. This method not only analyzes organelle RNA end modifications and the length and origin of organelle RNA poly(A), but can also be used to screen substrate RNAs of RNases and RBPs involved in the processing and stabilization of plant organelle RNA ends.

[0214] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for analyzing the ends of RNA in plant organelles, characterized in that, Includes the following steps: (1) RNA was treated with RNA 5′ polyphosphatase to obtain the treated RNA; (2) The treated RNA was circularized in vitro to obtain circularized RNA; (3) Sequencing of in vitro circular RNA to obtain initial read lengths; filtering of initial read lengths to obtain clean read lengths; (4) Concatenate the clean read lengths to obtain concatenated read lengths and unconcatenated read lengths; merge the concatenated read lengths and unconcatenated read lengths to obtain merged read lengths; (5) Determine the reconstructed genome based on the reference genome; (6) Use MeCi to detect in vitro circularized RNA linker sites with merged reads; (7) Align the in vitro circular RNA linker sites with the reference genome and the reconstructed genome respectively to obtain the linker site positions aligned with the reference genome and the linker site positions aligned with the reconstructed genome; then convert the linker site positions aligned with the reconstructed genome into the positions of the reference genome and merge them with the data of the reference genome to obtain the positions of the in vitro circular RNA linker sites. (8) Screen the in vitro circularized RNA linker sites to obtain high-confidence RNA in vitro circularized linker sites and determine their positions; determine the positions of the 5′ and 3′ ends of the corresponding RNA molecules based on the positions of the high-confidence RNA in vitro circularized linker sites, and infer the full length of the RNA molecules; (9) Clustering RNA molecules containing high-confidence RNA in vitro circularization junctions yields the abundance and number of reads at the 5′ and 3′ ends; (10) Determine the 5′ end type of RNA based on the number of clean reads and the number of 5′ and 3′ end reads; (11) Determine the terminal positions of mRNA and ncRNA based on the number of clean reads, the number of 5′ and 3′ end reads, and the annotation information of the reference genome; (12) Cluster the 5′ end types of RNA to obtain the cumulative abundance of the 5′ end types of RNA; determine the end types of mRNA and ncRNA based on the cumulative abundance of the 5′ end types of RNA; In step (4), the unjoined read length includes the reverse complementary sequence of the unjoined 3′ end read length and the unjoined 5′ end read length.

2. The analytical method according to claim 1, characterized in that, In step (3), the sequencing method is paired-end sequencing.

3. The analytical method according to claim 2, characterized in that, In step (3), the filtering method is to remove sequencing adapters and low-quality reads from the initial read length; The low-quality reads are those where the proportion of bases that cannot be determined exceeds 10% and those with a length of less than 30 bp.

4. The analytical method according to claim 1, characterized in that, In step (5), the method for determining the reconstructed genome is as follows: for genes in the reference genome containing backsplicing introns, their exon sequences, as well as the sequences 2000-nt upstream of the first intron and 1000-nt downstream of the last exon, are spliced ​​together to form the reconstructed genome.

5. The analytical method according to claim 4, characterized in that, In step (8), the screening criteria are as follows: A. Retrieve join points that have been aligned to the reference genome and remove join points containing ≥3 alignment positions; B. For junction points with two alignment positions, if both contain the same overlapping bases or non-coding bases, both are retained. If the number of non-coding bases in both is different and the difference in the number of non-coding bases between them is > 10 bp, the junction point with the smaller number of non-coding bases is retained. If the number of non-coding bases in both is different and the difference in the number of non-coding bases between them is ≤ 10 bp, both are removed; C. Single or multiple alignment junction points containing 1 - 3 non-coding bases are all retained. For junction points containing 4 - 100 non-coding bases, only the part with an A base proportion ≥ 70% in the non-coding bases is retained; D. Single alignment junction points containing 1 - 3 repeat bases or neither repeat bases nor non-coding bases are all retained; E. On the basis of filtering different types of junction points above, only in vitro circularization junction points that appear in at least two sequencing samples are retained, that is, in vitro circularization junction points with high confidence are contained.

6. The analytical method according to claim 5, characterized in that, In step (10), the method for judging the 5′-end type of the RNA is as follows: Calculate the RPTM value according to the read length and clean read length at the 5′ and 3′ ends; Calculate the FCV value according to the RPTM value; Judge the 5′-end type of the RNA according to the FCV value; The calculation formula for the RPTM value is the read length at the end / the clean read length; The judgment criteria for the 5′-end type of the RNA are as follows: If FCV ≥ 5, it is considered that the 5′ end contains a triphosphate group, marked as "primary", from transcription initiation; If FCV ≤ 1, it is considered that the 5′ end contains a monophosphate group, marked as "secondary", from post-transcriptional processing; If 1 < FCV < 5, it is impossible to determine whether the 5′ end contains a monophosphate group or a triphosphate group, marked as "uncertain".

7. The analytical method according to claim 6, characterized in that, In step (11), the judgment criteria for the 5′-end position of the mRNA are as follows: The 5′ end needs to be located upstream of the translation initiation codon and RPTM ≥ 10; The judgment criteria for the 5′-end position of the ncRNA are as follows: RPTM ≥ 50 and the distance to its corresponding 3′ is between 300 - 3000 bases; The RPTM is calculated according to the read length and clean read length at the 5′ and 3′ ends; The calculation formula for the RPTM value is the read length at the end / the clean read length.

8. The analytical method according to claim 7, characterized in that, In step (12), the judgment criteria for the 5′-end type of the mRNA and ncRNA are as follows: If the cumulative abundance of the 5′-end type of the RNA marked as "primary" in the 5′-end interval of the mRNA or ncRNA ≥ 70%, it is considered primary; If the cumulative abundance of the 5′-end type of the RNA marked as "secondary" in the 5′-end interval of the mRNA or ncRNA ≥ 70%, it is considered secondary; Otherwise, it is "uncertain".

9. Use of the analysis method according to any one of claims 1 - 8 in screening substrate RNAs of RNase and RBP involved in the processing and stabilization of plant organelle RNA ends.