A method and system for detecting RNA multi-dimensional structure based on second-generation sequencing and application thereof

By performing paired-end sequencing on cDNA libraries used for RNA sequencing, dividing them into cap-related and non-cap-related reads, and analyzing information related to the 5' m7G cap structure and the 3' poly(A) tail of RNA, the problem of simultaneously detecting RNA multidimensional structures in existing technologies is solved, and efficient multidimensional information analysis is achieved in single-cell or low-input samples is realized.

CN122484264APending Publication Date: 2026-07-31SHENZHEN GENEDO MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN GENEDO MEDICAL TECH CO LTD
Filing Date
2026-05-20
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing next-generation sequencing methods have difficulty obtaining the 5' cap structure, 3' Poly(A) tail structure, tail length information, and gene expression information of RNA molecules in the same experiment, especially in single-cell or low-input samples. Furthermore, it is difficult to accurately identify m6A and the m6Am modification and Poly(A) sites adjacent to the cap structure.

Method used

By performing paired-end sequencing on cDNA libraries used for RNA sequencing, cap-related and non-cap-related reads are divided. Combining the characteristics of reverse transcriptase, information related to the 5' m7G cap structure and the 3' poly(A) tail is analyzed. Characteristic sequences are used to identify the m7G cap site, poly(A) tail site, APA window, and modification state, thus generating a multidimensional structure detection method.

Benefits of technology

It enables simultaneous analysis of the 5' end, 3' end, and genome information of RNA in single-cell or low-input samples, improving the accuracy of m7G cap recognition. It is suitable for rare sample detection and applicable to oocyte maturation quality assessment, assisted reproduction, and RNA metabolism mechanism research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

This invention belongs to the field of biotechnology, specifically relating to a method, system, and application of joint detection of RNA multidimensional structures based on next-generation sequencing. The joint detection method includes: S1, sequencing a cDNA library for sequencing and preprocessing the raw sequencing reads to obtain R1 and R2 reads; S2, dividing the R1 read into cap-related reads and non-cap reads based on the number of non-template G bases in the R1 read's initiation region; concatenating the R1 and R2 reads to obtain a full-length concatenated read, and extracting tail-related reads from the R2 read and the full-length concatenated read; S3, analyzing the cap-related and non-cap reads to obtain m7G cap structure information; and analyzing the tail-related reads to obtain poly(A) tail information. This invention's joint detection method simultaneously obtains multidimensional information such as the RNA m7G cap structure and poly(A) tail using the same library and analysis process, solving the problem of existing technologies requiring multiple independent experiments to obtain information in different dimensions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biotechnology, specifically to a method, system, and application for the joint detection of RNA multidimensional structures based on next-generation sequencing. Background Technology

[0002] RNA molecules typically possess a 5' cap and a 3' poly(A) tail. The 5' m7G cap is closely related to transcription initiation, RNA stability, and translation initiation, while the 3' poly(A) tail is closely related to transcription termination, RNA stability, translation efficiency, and degradation processes. In oocyte maturation, early embryonic development, and other biological systems, the cap, tail, and internal modifications of RNA molecules work together to determine RNA metabolic fate.

[0003] Current next-generation sequencing methods typically focus on a single dimension, such as gene expression, 5' end information, or 3' end information. For example, CAGE-seq emphasizes transcription start site detection, Tail-seq emphasizes poly(A) tail structure detection, and conventional single-cell RNA sequencing usually only covers the 3' or 5' end, making it difficult to simultaneously obtain the 5' cap structure, 3' poly(A) tail structure, tail length information, and gene expression information of the same RNA molecule in the same experiment. For precious samples such as oocytes and early embryos, existing methods also generally suffer from drawbacks such as high initial sample size requirements, fragmented information dimensions, and insufficient sensitivity to low-abundance transcripts.

[0004] Furthermore, existing m6A-related detection methods often struggle to reliably distinguish m6A from adjacent m6Am modifications when cap information is lacking; existing Poly(A) analysis methods also frequently fail to accurately identify Poly(A) sites, APA windows, and tail lengths simultaneously in single-cell or low-input RNA samples. Therefore, there is an urgent need for a technical solution that can simultaneously preserve RNA 5', 3', and genome information in single-cell or low-input samples, and achieve joint analysis of cap structure, tail structure, and modification status. Summary of the Invention

[0005] The purpose of this invention is to provide a method, system, and application for the joint detection of RNA multidimensional structures based on next-generation sequencing.

[0006] The first aspect of this invention provides a method for joint detection of RNA multidimensional structures based on next-generation sequencing, comprising the following steps: S1: Perform paired-end sequencing on the constructed RNA sequencing cDNA library to obtain R1 and R2 raw sequencing reads. Preprocess the R1 and R2 raw sequencing reads to obtain R1 and R2 reads. S2: Based on the number of non-template G bases in the starting region of the R1 read segment, the R1 read segment is divided into cap-related read segments and non-cap-related read segments; the R1 read segment is spliced ​​with the R2 read segment to obtain a full-length spliced ​​read segment, and tail-related read segments are extracted from the R2 read segment and the full-length spliced ​​read segment; S3: Analyze the m7G cap structure at the 5' end of RNA using cap-related and non-cap-related reads to obtain information related to the m7G cap structure of RNA; perform sequence analysis on tail-related reads to obtain information related to the poly(A) tail at the 3' end of RNA, thus completing the joint detection of the multidimensional structural features of RNA molecules.

[0007] According to the above-mentioned RNA multidimensional structure joint detection method, preferably, in step S2, the method for dividing cap-related reads and non-cap reads is as follows: reads with 4 or more non-template G bases in the R1 start region are cap-related reads, and reads with 3 or more non-template G bases in the R1 start region are non-cap reads.

[0008] According to the above-mentioned RNA multidimensional structure joint detection method, preferably, in step S2, the specific operation of extracting tail-related reads from the R2 read and the full-length spliced ​​read is as follows: extracting reads containing continuous A / T feature sequences from the R2 read and the full-length spliced ​​read, thus obtaining tail-related reads; wherein, the continuous A / T feature sequence refers to a sequence with a length of continuous A or continuous T bases at the end of the sequence greater than or equal to 8.

[0009] According to the above-mentioned RNA multidimensional structure joint detection method, preferably, in step S3, the m7G cap structure related information includes m7G cap site information and m7G cap modification level information; the poly(A) tail related information includes poly(A) tail site information, poly(A) tail length information, and APA window information.

[0010] According to the above-mentioned method for joint detection of RNA multidimensional structure, preferably, in step S3, the specific operation of analyzing the m7G cap structure at the 5' end of RNA using cap-related reads and non-cap reads includes: independently aligning the cap-related reads and non-cap reads to the reference genome of the corresponding species to obtain the location information of the cap-related reads and non-cap reads on the reference genome; based on the location information, statistically analyzing the genomic distribution enrichment differences of the cap-related reads and non-cap reads, identifying the m7G cap start site, and obtaining m7G cap site information; statistically analyzing the number and enrichment abundance of cap-related reads corresponding to each gene and each transcript, quantifying the m7G cap modification level of each transcript, and obtaining m7G cap modification level information.

[0011] According to the above-mentioned method for joint detection of RNA multidimensional structure, preferably, the method for obtaining the poly(A) tail site information is as follows: the tail-related read is aligned to the reference genome of the corresponding species to obtain the location information of the tail-related read on the reference genome; based on the enrichment information of the tail-related read on the reference genome, combined with the end coordinates of the continuous A feature sequences in the tail-related read, the poly(A) tail site of each transcript is located to obtain the poly(A) tail site information.

[0012] According to the above-mentioned RNA multidimensional structure joint detection method, preferably, the method for obtaining the poly(A) tail length information is as follows: calculate the poly(A) tail length based on the length of the continuous A / T feature sequence in the R2 read segment to obtain the poly(A) tail length information.

[0013] According to the above-mentioned RNA multidimensional structure joint detection method, preferably, the method for obtaining the APA window information is as follows: analyze all poly(A) tail sites of the same gene, screen out multiple adjacent poly(A) tail sites located in the same peak region, and calculate the maximum distance between these adjacent poly(A) tail sites, which is the APA window.

[0014] According to the above-mentioned RNA multidimensional structure joint detection method, preferably, the m7G cap structure related information also includes the m7G cap structure quantitative expression matrix. The method for generating the m7G cap structure quantitative expression matrix is ​​as follows: the cap reads and non-cap reads covering the selected reliable m7G cap sites (starting from the m7G site at the 5' end) are separately constructed into BAM files, and then the read coverage depth of each m7G cap site is counted using feature counts software as the expression level information of the m7G cap, thus obtaining the expression level matrix of a single m7G site; the modification level of the m7G cap is the proportion of the cap reads covering a specific m7G site to all reads of that m7G ​​site, thus obtaining the modification level matrix of a single m7G site.

[0015] According to the above-mentioned RNA multidimensional structure joint detection method, preferably, the poly(A) tail related information also includes poly(A) tail base modification information. The method for obtaining the poly(A) tail base modification information is as follows: scan the base composition at the end of the poly(A) tail, and determine the poly(A) tail base modification information based on the enrichment type of non-A bases at the end of the poly(A) tail; wherein, if the poly(A) tail end is enriched with U bases, it is determined to be uridine acidification modification; if the poly(A) tail end is enriched with G bases, it is determined to be guanylic acidification modification.

[0016] According to the above-described method for joint detection of RNA multidimensional structures, preferably, in step S1, the method for constructing the cDNA library for RNA sequencing includes: (1) Obtain RNA from the sample to be tested, and randomly break the RNA to obtain fragmented RNA; (2) Repair the 3' end of the fragmented RNA so that the 3' end of the fragmented RNA is suitable for connecting RNA adapters; (3) Connect the RNA adapter to the 3' end of the fragmented RNA after step (2) to obtain fragmented RNA with RNA adapter; use the fragmented RNA with RNA adapter as a template, perform reverse transcription using primers complementary to the RNA adapter, and add template conversion oligonucleotides to the reverse transcription system. After the reverse transcription is completed, cDNA is obtained. (4) Perform PCR amplification on the cDNA obtained in step (3), and purify the PCR amplification product to obtain the cDNA library for RNA sequencing.

[0017] Further, in step S1, the method for constructing the cDNA library for RNA sequencing includes: (1) Obtain RNA from the sample to be tested, and randomly break the RNA to obtain fragmented RNA; (2) Repair the 3' end of the fragmented RNA so that the 3' end of the fragmented RNA is suitable for connecting RNA adapters; (3) Connect the RNA adapter to the 3' end of the fragmented RNA after step (2) to obtain fragmented RNA with RNA adapter; use the fragmented RNA with RNA adapter as a template, perform reverse transcription using primers complementary to the RNA adapter, and add template conversion oligonucleotides to the reverse transcription system. After the reverse transcription is completed, cDNA is obtained. (4) Perform PCR amplification on the cDNA obtained in step (3) to obtain double-stranded cDNA product; (5) Perform in vitro reverse transcription amplification on the double-stranded cDNA product obtained in step (4) to obtain RNA product; enrich the poly(A) positive mRNA in the RNA product; process the poly(A) positive mRNA according to the above steps (1)-(4) to obtain the double-stranded cDNA product, which is the cDNA library for RNA sequencing.

[0018] According to the above-mentioned RNA multidimensional structure joint detection method, preferably, in step (1), the sample to be tested includes a cell sample or a tissue sample, wherein the cell sample is an oocyte, stem cell or tumor cell; the tissue sample is a clinical biopsy sample; in step (3), the template-converting oligonucleotide contains a template-converting anchor sequence and a blocking modification, and the reverse transcriptase used in the reverse transcription process has reverse transcription activity and terminal transferase activity.

[0019] A second aspect of the present invention provides an RNA multidimensional structure analysis system, comprising: a library construction module, a computational analysis module, and an output module; wherein, the library construction module is used to construct a cDNA library for RNA sequencing; the computational analysis module is used to perform paired-end sequencing on the constructed RNA sequencing cDNA library according to the RNA multidimensional structure joint detection method described in the first aspect, and to perform computational analysis based on the sequencing data to obtain information related to the 5' m7G cap structure and the 3' poly(A) tail of RNA; the output module is used to output the joint analysis results of the information related to the 5' m7G cap structure and the 3' poly(A) tail of RNA.

[0020] According to the above-mentioned RNA multidimensional structure analysis system, preferably, the output module is further used to generate oocyte quality assessment results, in vitro maturation culture monitoring results, or RNA metabolic remodeling maps between different cell states.

[0021] The third aspect of this invention provides the application of the RNA multidimensional structure joint detection method described in the first aspect or the RNA multidimensional structure analysis system described in the second aspect in oocyte maturation quality assessment, oocyte screening in assisted reproduction, in vitro maturation culture optimization, RNA metabolism mechanism research, or low input sample RNA structure detection.

[0022] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The RNA multidimensional structure joint detection method based on second-generation sequencing of the present invention obtains multidimensional information such as m7G cap, m6Am, Poly(A) site, APA window, Poly(A) tail length and expression level through the same library and analysis process, avoiding the problem of needing multiple independent experiments to obtain different dimensions of information in the prior art; it can simultaneously analyze the 5' end, 3' end and gene body information of RNA in single cells or low input samples. Therefore, it can be used for oocyte maturation quality assessment, oocyte screening in assisted reproduction, in vitro maturation culture optimization, RNA metabolism mechanism research or low input sample RNA structure detection and other application scenarios.

[0023] (2) The RNA multidimensional structure joint detection method based on second-generation sequencing of the present invention can improve the accuracy of m7G cap identification by comparing and analyzing cap-related reads and non-cap reads.

[0024] (3) The RNA multidimensional structure joint detection method based on second-generation sequencing of this invention locates the Poly(A) site and calculates the tail length by tail-related reads, which is suitable for the detection of rare and precious samples.

[0025] (4) The RNA multidimensional structure analysis system of the present invention can realize the construction of cDNA library for RNA sequencing, library sequencing and computational analysis of sequencing data, obtain the joint analysis results of RNA 5' end m7G cap structure related information and 3' end poly(A) tail related information and output them. It is easy to operate and quick to detect, and is particularly suitable for application scenarios such as oocytes and early embryos that have strict requirements for sample size and require joint analysis of multidimensional information.

[0026] (5) The RNA multidimensional structure joint detection method based on next-generation sequencing of this invention has been proven to identify highly repetitive m7G sites, accurately extract cap reads, tail reads, and gene body reads, and reveal the correlation between translation efficiency and m7G / m6Am and Poly(A) tail structure in oocytes. Therefore, the RNA multidimensional structure joint detection method of this invention has good application prospects in single-cell RNA metabolism research and assisted reproductive testing. Attached Figure Description

[0027] Figure 1 This is a schematic diagram of the cDNA library construction and targeted enrichment process for RNA sequencing in this invention. The diagram illustrates that after random fragmentation and 3' end repair of the RNA sample, an RNA adapter is ligated to the 3' end of the fragmented RNA, followed by reverse transcription and template conversion. This preserves the 5' cap-related features and 3' poly(A) tail-related features of the RNA in the cDNA. Subsequently, after cDNA amplification and in vitro transcription amplification, targeted enrichment can be further performed through rRNA removal or poly(dT) selection to obtain a sequencing library suitable for the joint detection of RNA cap structure, tail structure, and genome information. Figure 2 This is a schematic diagram illustrating the identification principle of RNA cap-related reads, non-cap-related reads, and poly(A)-tail-related reads in this invention; wherein, Figure 2 A is a schematic diagram showing the different non-template G / C characteristics generated by capped RNA and non-capped RNA during reverse transcription and template conversion. RNA with m7G cap structure generates cap-related non-template G / C signals during template conversion, while non-capped or broken RNA generates background signals. Figure 2 B is a schematic diagram of paired-end sequencing read classification. Cap-related reads and ordinary reads are identified based on the number of non-template G bases in the R1 read start region, and poly(A) tail-related reads are identified based on the continuous A / T feature sequence at the end of the R2 read or full-length spliced ​​read. Figure 2 C is a schematic diagram showing the distribution of cap-related reads, non-cap background reads, and tail-related reads in different regions of the transcript. Cap-related signals are mainly enriched at the 5' end of the transcript, tail-related signals are mainly enriched at the 3' end of the transcript, and non-cap background reads are distributed in the gene body region. Figure 3This is a schematic diagram illustrating the analysis results of the poly(A) tail site, APA window, poly(A) tail length, and tail-terminal base modification in this invention; wherein, Figure 3 A is a schematic diagram of the distribution of m7G cap-related reads, non-cap-related reads, positive tail-related reads, and negative tail-related reads on the genome at representative gene loci, and shows multiple poly(A) sites and APA peak regions; Figure 3 B is a schematic diagram of the poly(A) tail sequence and the composition of non-A bases at the tail end in the reads corresponding to different poly(A) sites, which is used to identify poly(A) tail end uridine acidification, guanylic acidification or other non-A base modification features; Figure 3 C is a schematic diagram of the clustering of adjacent poly(A) sites and the APA window within the same APA peak region, where the APA window represents the distribution range between multiple adjacent poly(A) tail sites within the same peak region; Figure 3 D is a poly(A) tail length distribution curve, used to show the overall distribution characteristics of the poly(A) tail length in the detected sample; Figure 3 E is a comparison graph of the length of the poly(A) tail of all poly(A) tail reads and the poly(A) tail of the U-tailed reads, used to show the difference in poly(A) tail length between the uridine-treated RNA at the tail and the total RNA. Figure 4 This is a schematic diagram illustrating the analysis of distinguishing between m6Am modification at the cap-adjacent site and internal m6A modification according to the present invention; wherein, Figure 4 A is a flowchart illustrating the process of distinguishing m6Am from m6A based on cap-related reads and methylation enrichment data. By extracting reads containing m7G cap signals and combining them with methylation enrichment signals, candidate m6Am sites located near the m7G cap can be identified and distinguished from m6A signals within the non-cap region. Figure 4 B is a distribution curve of non-cap GpppB related reads near the transcription start site; Figure 4 C is a distribution curve of cap-related GpppA reads near the transcription start site, used to show the enrichment characteristics of cap-related reads in the 5' end region of the transcript and to help determine the methylation signal at the cap-adjacent site; Figure 5 This is a schematic diagram illustrating the application of the RNA multidimensional structure joint detection method of the present invention in oocyte maturation quality assessment; wherein, Figure 5 A shows the morphological diagrams of different oocyte samples and the signal distribution diagrams of the input group, m6A immunoenrichment group, and supernatant group in representative genes, which are used to demonstrate that the method of the present invention can detect gene body methylation signals and cap-related methylation signals in single or low input oocyte samples. Figure 5 B is a schematic diagram of the poly(A) tail-related signal distribution of representative genes in oocytes at different maturity stages, used to show the differences in RNA tail structure characteristics among GV stage, MII stage and abnormal oocytes; Figure 5 C is a schematic diagram of the poly(A) tail base composition and poly(A) tail length distribution in oocytes during the GV and MII stages; Figure 5 D is a comparison diagram of the poly(A) tail length of oocytes in the GV stage and MII stage, used to show the differences in poly(A) tail length in oocytes at different maturation stages; Figure 5 E is a principal component analysis diagram constructed based on RNA cap structure, tail structure, methylation signal, and expression characteristics, used to distinguish oocytes in different states such as GV stage, MII stage, degeneration, and developmental arrest. Detailed Implementation

[0028] To enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention will be described in detail below with reference to specific embodiments.

[0029] The following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the exemplary embodiments of the invention. Without departing from the spirit and essence of the invention, those skilled in the art can make adjustments to the sequence of steps, reagent combinations, enrichment methods, and calculation parameters, and such equivalent substitutions fall within the scope of protection of this invention.

[0030] Unless otherwise stated, the routine molecular biology operations, nucleic acid purification, PCR amplification, magnetic bead purification, sequencing and bioinformatics analysis used in the following examples can all be performed using conventional techniques in the field. Unless otherwise specified, the reagents or instruments used are all commercially available products.

[0031] Example 1: Construction of cDNA library for RNA sequencing The specific steps for constructing a cDNA library for RNA sequencing are as follows: 1. Total RNA extraction and RNA fragmentation: Single oocytes were collected and placed in a 5 μL lysis buffer containing heat-sensitive DNase and 0.5% Triton X-100. Lysis was performed on ice for 15 minutes to extract total RNA. Subsequently, the RNA was extracted using MgSO4. 2+ The extracted total RNA was subjected to thermal fragmentation, ultrasonic fragmentation, or enzymatic fragmentation in the presence of the present environment. The reaction conditions were heating at 94°C for 4 minutes to obtain fragmented RNA of 100–500 bp.

[0032] 2. Repair of the 3' end of fragmented RNA: Add 1 μL T4 polynucleotide kinase, 0.5 μL RNase inhibitor, 1 μL 10× reaction buffer, and 2.5 μL sterile enzyme-free water to the fragmented RNA obtained in step 1. Incubate at 37°C for 30 min to repair the 3' end of the fragmented RNA, removing the 3' phosphate group and filling in any missing ends to form a smooth 3'-OH hydroxyl terminus suitable for subsequent adapter ligation. After repair, purify using magnetic beads to remove unreacted enzymes and impurities, collect the purified fragmented RNA, and store on ice for later use.

[0033] 3. Ligation of fragmented RNA 3' adapters: To the 3'-terminal-repaired fragmented RNA, add 1 μL of RNA 3' adapter, 0.5 μL of T4 RNA ligase, 0.5 μL of RNase inhibitor, and 12 μL of ligation buffer, and incubate overnight at 22°C (12-16 h) to covalently ligate the RNA 3' adapter to the 3' end of the fragmented RNA. After ligation, add an equal volume of magnetic beads for purification to remove free RNA 3' adapters and impurities, and collect the fragmented RNA after adapter ligation. The adapter sequence is: 5-App-NAGATCGGAAGAGCACACGTCT.

[0034] 4. Reverse transcription reaction: (1) Template and primer pretreatment: Take 17 μL of fragmented RNA obtained in step 2 after ligation of the adapter, add 1 μL of picoRT reverse transcription primer (20 μM) that is complementary to the 3' adapter, heat at 75°C for 3 min, and then quickly place on ice for 5 min to open the RNA secondary structure so that the primer and the RNA 3' adapter can bind complementaryly.

[0035] (2) Preparation of the reverse transcription reaction system (total volume 25 μL): Add 5 μL of 5× reverse transcription buffer, 1 μL of 10 mM dNTP mixture, 0.5 μL of RNase inhibitor, 0.5 μL of TSO (20 μM), and 1 μL of reverse transcriptase (the reverse transcriptase is a reverse transcriptase with terminal transferase activity, such as SMARTScribe reverse transcriptase, SuperScript IV reverse transcriptase, etc.) to the mixture after pretreatment in step (1) in sequence, and gently pipette to mix. The sequence of the TSO is: (dSpacer)3CACGACGCTCTTCCGATCTNNNNrGrG+G.

[0036] (3) Primer annealing: Set the PCR instrument program and incubate at 25°C for 15 min to ensure that the reverse transcription primers bind precisely to the RNA 3' end adapter and that TSO remains stable in the system.

[0037] (4) First-strand cDNA synthesis and template switching: The temperature of the PCR instrument was raised to 42℃ and incubated for 90 min. The reverse transcriptase extended along the RNA template from 3' to 5' to synthesize single-stranded cDNA. When it extended to the 5' end of the RNA, the reverse transcriptase relied on terminal transferase activity to add multiple consecutive CCC bases to the 3' end of the newly generated cDNA, which were complementary to the GGG bases at the 3' end of TSO. The reverse transcriptase switched templates and continued to extend with TSO as the new template, introducing the RNA 5' end characteristic sequence at the end of the cDNA.

[0038] (5) Reaction termination and single-stranded cDNA purification: The temperature of the PCR instrument was raised to 85℃ and incubated for 5 min to inactivate the reverse transcriptase and terminate the reverse transcription reaction. After the reaction, an equal volume of magnetic beads was used for purification to obtain pure single-stranded cDNA, which was then kept on ice for later use.

[0039] 5. cDNA PCR amplification and library construction: The single-stranded cDNA obtained in step 4 was dissolved in 30 μL of sterile, enzyme-free water. 4 μL of P7 primer and 4 μL of P5 primer (purchased from NEB) were added, along with 40 μL of 2× amplification premix and 2 μL of high-fidelity DNA polymerase. PCR amplification was then performed. The PCR amplification program was as follows: 95℃ pre-denaturation for 3 min; 98℃ pre-denaturation for 30 s, 55℃ annealing for 30 s, and 68℃ extension for 30 s, for a total of 14-18 cycles; final extension at 68℃ for 3 min; and incubation at 4℃. After the PCR amplification reaction, 88 μL of magnetic beads were added for purification to obtain the PCR amplification product, which is the cDNA library. This cDNA library can be directly used for paired-end sequencing and RNA cap and tail structure analysis. Furthermore, sequencing analysis of the cDNA library constructed in this example can improve the detection capability of total RNA information.

[0040] Example 2: Construction of cDNA libraries for targeted enrichment RNA sequencing The specific steps for constructing cDNA libraries for targeted enrichment RNA sequencing (e.g.) Figure 1 As shown below: 1. Total RNA extraction and RNA fragmentation: Single oocytes were collected and placed in a 5 μL lysis buffer containing heat-sensitive DNase and 0.5% Triton X-100. Lysis was performed on ice for 15 minutes to extract total RNA. Subsequently, the RNA was extracted using MgSO4. 2+ The extracted total RNA was subjected to thermal fragmentation, ultrasonic fragmentation, or enzymatic fragmentation in the presence of the present environment. The reaction conditions were heating at 94℃ for 4 minutes to obtain fragmented RNA of 100-500 bp.

[0041] 2. Repair of the 3' end of fragmented RNA: Add 1 μL T4 polynucleotide kinase, 0.5 μL RNase inhibitor, 1 μL 10× reaction buffer, and 2.5 μL sterile enzyme-free water to the fragmented RNA obtained in step 1. Incubate at 37°C for 30 min to repair the 3' end of the fragmented RNA, removing the 3' phosphate group and filling in any missing ends to form a smooth 3'-OH hydroxyl terminus suitable for subsequent adapter ligation. After repair, purify using magnetic beads to remove unreacted enzymes and impurities, collect the purified fragmented RNA, and store on ice for later use.

[0042] 3. Ligation of fragmented RNA 3' adapters: To the 3'-terminal-repaired fragmented RNA, add 1 μL of RNA 3' adapter, 0.5 μL of T4 RNA ligase, 0.5 μL of RNase inhibitor, and 12 μL of ligation buffer, and incubate overnight at 22°C (12-16 h) to covalently ligate the RNA 3' adapter to the 3' end of the fragmented RNA. After ligation, add an equal volume of magnetic beads for purification to remove free RNA 3' adapters and impurities, and collect the fragmented RNA after adapter ligation. The adapter sequence is: 5-App-NAGATCGGAAGAGCACACGTCT.

[0043] 4. Reverse transcription reaction: (1) Template and primer pretreatment: Take 17 μL of fragmented RNA obtained in step 2 after ligation of the adapter, add 1 μL of picoRT reverse transcription primer (20 μM) that is complementary to the 3' adapter, heat at 75°C for 3 min, and then quickly place on ice for 5 min to open the RNA secondary structure so that the primer and the RNA 3' adapter can bind complementaryly.

[0044] (2) Preparation of the reverse transcription reaction system (total volume 25 μL): Add 5 μL of 5× reverse transcription buffer, 1 μL of 10 mM dNTP mixture, 0.5 μL of RNase inhibitor, 0.5 μL of TSO (20 μM), and 1 μL of reverse transcriptase (the reverse transcriptase is a reverse transcriptase with terminal transferase activity, such as SMARTScribe reverse transcriptase, SuperScript IV reverse transcriptase, etc.) to the mixture after pretreatment in step (1) in sequence, and gently pipette to mix. The sequence of the TSO is: (dSpacer)3CACGACGCTCTTCCGATCTNNNNrGrG+G.

[0045] (3) Primer annealing: Set the PCR instrument program and incubate at 25°C for 15 min to ensure that the reverse transcription primers bind precisely to the RNA 3' end adapter and that TSO remains stable in the system.

[0046] (4) First-strand cDNA synthesis and template switching: The temperature of the PCR instrument was raised to 42℃ and incubated for 90 min. The reverse transcriptase extended along the RNA template from 3' to 5' to synthesize single-stranded cDNA. When it extended to the 5' end of the RNA, the reverse transcriptase relied on terminal transferase activity to add multiple consecutive CCC bases to the 3' end of the newly generated cDNA, which were complementary to the GGG bases at the 3' end of TSO. The reverse transcriptase switched templates and continued to extend with TSO as the new template, introducing the RNA 5' end characteristic sequence at the end of the cDNA.

[0047] (5) Reaction termination and single-stranded cDNA purification: The temperature of the PCR instrument was raised to 85℃ and incubated for 5 min to inactivate the reverse transcriptase and terminate the reverse transcription reaction; after the reaction was completed, the reverse transcriptase was purified by magnetic beads to obtain pure single-stranded cDNA, which was then kept on ice for later use.

[0048] 5. cDNA PCR amplification: The single-stranded cDNA obtained in step 4 was dissolved in 30 μL of sterile, enzyme-free water. 4 μL of T7-preUniversal primer (sequence: TAATACGACTCACTATAGGGAGACACGACGCTCTTCCGATC-sT) and 4 μL of picoRT reverse transcription primer (sequence: AGACGTGTGCTCTTCCGATCT), 40 μL of 2× amplification premix, and 2 μL of high-fidelity DNA polymerase were added, followed by PCR amplification. The PCR amplification program was as follows: 95℃ pre-denaturation for 3 min; 98℃ pre-denaturation for 30 s, 55℃ annealing for 30 s, and 68℃ extension for 30 s, for a total of 14-18 cycles; final extension at 68℃ for 3 min; and incubation at 4℃. After the PCR amplification reaction, 88 μL of magnetic beads were added for purification to obtain the PCR amplification product.

[0049] 6. In vitro transcription of PCR amplification products and enrichment of Poly(A) positive mRNA: The PCR amplification product obtained in step (5) was transcribed and amplified in vitro to obtain RNA. Then, the Poly(A) positive mRNA was enriched using Oligo(dT) magnetic beads to obtain poly(A) positive mRNA.

[0050] 7. Construction of cDNA libraries: The poly(A)-positive mRNA obtained in step 6 is processed according to steps 1 to 5 above. The resulting double-stranded cDNA product is the cDNA library for RNA sequencing. This cDNA library can be directly used for paired-end sequencing and RNA cap and tail structure analysis. Furthermore, sequencing analysis of the cDNA library constructed in this embodiment can improve the detection capability of RNA poly(A) information.

[0051] Example 3: Joint analysis of RNA cap and tail structures: 1. cDNA library paired-end sequencing and sequencing data preprocessing The cDNA libraries constructed in Example 1 or Example 2 were sequenced using the Illumina platform PE150 or PE250 paired-end sequencing mode. The raw FastQ data obtained from sequencing were processed by removing adapters, trimming low-quality bases, and removing PCR repeats to obtain high-quality clean R1 and R2 reads. The R1 and R2 reads were merged and spliced ​​according to the sequence overlap region to obtain full-length spliced ​​reads.

[0052] 2. Hat Structure Analysis: like Figure 2 As shown, the core scientific principle of this invention stems from the difference in strand conversion efficiency between the m7G cap and the ordinary phosphorylated end of reverse transcriptase at the RNA terminal. At the m7G end, reverse transcriptase can rely on the m7G cap as a template to add multiple guanine (G) bases. Figure 2 A) Therefore, this invention divides the R1 reads based on the number of non-template G bases in the R1 read start region from template-converted oligonucleotides (TSO). Reads with 4 or more non-template G bases in the R1 start region are cap-related reads, and reads with 3 or more non-template G bases in the R1 start region are non-cap reads.

[0053] The m7G cap structure at the 5' end of RNA was analyzed using cap-related and non-cap-related reads to obtain information related to the m7G cap structure of RNA. The specific procedures are as follows: Cap-related and non-cap-related reads were independently aligned to the reference genome of the corresponding species (Figure 2B) to obtain the location information of cap-related and non-cap-related reads on the genome. The read coverage of cap-related and non-cap-related reads at each position in the reference genome was statistically analyzed to obtain the coverage enrichment peak of cap-related reads and the distribution baseline of non-cap-related reads. Cap-related reads aligned to the genome showed significant enrichment at m7G ​​sites; these enriched sites are the m7G cap initiation sites. Figure 2C) Obtain m7G cap site information. Further, by comparing cap-related reads with the genome and identifying the starting positions of excess residues, the unit site information of m7G can be determined. For this unit site information, cap-related reads and non-cap reads covering the site are further screened. When cap-related reads account for less than or equal to 50%, it is determined to be a reliable m7G cap site, ensuring the accuracy of the identification results. This is determined by the efficiency of reverse transcriptase, which prevents all m7G sites from generating cap-related reads. Finally, the number of cap-related reads for each gene and transcript is further counted to obtain its m7G abundance. The m7G cap modification level of each transcript is quantified by calculating cap-related reads and non-cap reads covering m7G sites, thus obtaining m7G cap modification level information. Finally, the quantitative data of cap-related reads for all genes are summarized to generate a cap structure quantitative expression matrix for subsequent differential analysis, association analysis, and other downstream studies. The method for generating the m7G cap structure quantitative expression matrix is ​​as follows: cap reads and non-cap reads covering (starting from the m7G site at the 5' end) of the selected reliable m7G cap sites are separately constructed into BAM files, and then feature counts are used. The software calculates the read coverage depth of each m7G cap site as the expression level information of the m7G cap, obtaining an expression level matrix for a single m7G site. The modification level of the m7G cap is the proportion of the cap read covering a specific m7G site out of all reads at that m7G ​​site, obtaining a modification level matrix for a single m7G site. In summary, this invention efficiently captures the structural information of the RNA 5' m7G cap, enabling the localization and detection of the RNA cap.

[0054] 3. Poly(A) tail structure analysis: like Figure 3 As shown, this invention further extracts information near the polyA tail in the library to obtain multi-dimensional data such as RNA poly(A) tail processing sites, tail modifications, and length. First, reads containing consecutive A / T characteristic sequences (consecutive A / T characteristic sequences refer to sequences with consecutive A or T bases at the end of the sequence greater than or equal to 8 nt) are extracted from R2 reads and full-length spliced ​​reads to obtain tail-related reads, which are then aligned to the genome (e.g., ...). Figure 2(As shown in B). Further, sequence analysis is performed on the tail-related reads to obtain poly(A) tail-related information at the 3' end of the RNA. The specific operation is as follows: the tail-related reads are aligned to the reference genome of the corresponding species to obtain the location information of the tail-related reads on the reference genome. Based on the enrichment information of the tail-related reads on the reference genome, combined with the end coordinates of the continuous A characteristic sequences in the tail-related reads, the poly(A) tail sites of each transcript are located to obtain poly(A) tail site information. Specifically, the enrichment of the genomic location obtained from the tail-related reads can be analyzed using macs2 to obtain the enriched sites, which are the APA sites of the genes. Furthermore, these sites can be intersected with the reference APA sites of the species to obtain more reliable sites (Figure 3A). For a specific gene, the R1 reads are first aligned with the reference genome within 200 bp of the APA on the genome to obtain all poly(A) related reads associated with this gene. Then, the number of A bases in the consecutive A-characteristic sequences in the corresponding R2 reads is counted to obtain the poly(A) tail length (Figure 3B). The average and median of different tail-related reads of the same gene are calculated as the poly(A) length characteristic of a certain gene. Figure 3 C) Scan all poly(A) tail sites of the same gene, screen out multiple adjacent poly(A) tail sites located in the same peak region, and calculate the maximum distance between these adjacent poly(A) tail sites. This distance is the APA variable tail window (i.e., the APA window), and the APA window information is obtained (as shown in Figure 3C). In addition, scan the base composition at the end of the poly(A) tail. Based on the enrichment type of non-A bases at the end of the poly(A) tail, determine the poly(A) tail base modification information. Among them, the poly(A) tail enriched with U bases is determined to be uridine acidification modification; the poly(A) tail enriched with G bases is determined to be guanylate acidification modification (Figure 3C). Further, the poly(A) length difference information of RNA transcripts with different tail modifications can be obtained (Figures 3D, 3E). In summary, it can be seen that this invention patent efficiently captures the structure of RNA 3' poly(A) and realizes comprehensive analysis of RNA tail.

[0055] Example 4: m6Am Recognition Based on the previous example, further, when the library comes from sequencing data with methylation information, m7GpppA related reads can be extracted from the cap-related reads, and m6Am signals can be identified by combining them with immunoenrichment data. Specifically, data containing m7G signals are extracted to generate corresponding m7G_input reads, m7G_ip reads, non-Cap_input reads, and non-Cap_ip reads (Figure 4A). Then, the two sets of data are analyzed using macs2 software, and the regions with higher signals in the IP group are compared to determine whether they are m6A sites (as shown in Figure 4B) or m6Am sites (Figure 4C). This method can distinguish between m6Am and m6A by separating the methylation signals of sites adjacent to the cap from the m6A signals inside the non-cap. Preferably, the sequence motifs and signal intensities near the cap-related read enrichment regions are used as the criteria to obtain a set of m6Am candidate genes.

[0056] Example 5: Application in assessing oocyte maturation quality Following the procedures outlined in Examples 1 to 4, cDNA libraries were constructed, sequenced, and RNA cap and tail structures were jointly analyzed from human or mouse oocytes at different maturity stages or with varying morphological qualities. This allowed for the simultaneous acquisition of multidimensional characteristics such as m7G, m6Am, m6A, Poly(A) expression, Poly(A) tail length, and gene expression. Figure 5 As shown, individual oocytes were photographed before oocyte sequencing. Then, m6A enrichment and the aforementioned sequencing analysis were performed on individual oocytes. Input genomic data, m6A IP data, and supernatant data were displayed using IGV. m7G reads and non-cap reads were displayed using the methods described in the previous example, showing the distribution of m7G, m6Am, and m6A modifications in the transcriptome of each oocyte (Figure 5A). Further analysis of the poly(A) library yielded the expression of poly(A) modified transcripts in each oocyte (Figure 5B) and tail length information (Figures 5C and 5D). Furthermore, to integrate multi-dimensional omics data, the aforementioned RNA tail expression level, RNA cap level, gene expression level, m6A level, m6Am expression level, and tail length were all dimensionality-reduced to form a two-dimensional matrix. This two-dimensional matrix was then combined to obtain 12-dimensional data, which was further analyzed using clustering, principal component analysis, or classification model analysis. Further research combined with oocyte morphology revealed that different oocyte morphologies were successfully classified using multi-dimensional data (Figure 5E). Based on this, oocyte quality evaluation indicators can be further established based on the cap structure, tail length, and expression patterns of oocyte genes for oocyte screening and in vitro maturation culture monitoring in assisted reproduction.

[0057] The present invention has been described above by way of example. It should be noted that any simple modifications, alterations or other equivalent substitutions that can be made by those skilled in the art without creative effort without departing from the core of the present invention fall within the scope of protection of the present invention.

Claims

1. A method for joint detection of RNA multidimensional structure based on next-generation sequencing, characterized in that, Includes the following steps: S1: Perform paired-end sequencing on the constructed RNA sequencing cDNA library to obtain R1 and R2 raw sequencing reads. Preprocess the R1 and R2 raw sequencing reads to obtain R1 and R2 reads. S2: Based on the number of non-template G bases in the starting region of the R1 read segment, the R1 read segment is divided into cap-related read segments and non-cap-related read segments; the R1 read segment is spliced ​​with the R2 read segment to obtain a full-length spliced ​​read segment, and tail-related read segments are extracted from the R2 read segment and the full-length spliced ​​read segment; S3: Analyze the m7G cap structure at the 5' end of RNA using cap-related and non-cap-related reads to obtain information related to the m7G cap structure of RNA; perform sequence analysis on tail-related reads to obtain information related to the poly(A) tail at the 3' end of RNA.

2. The method for joint detection of RNA multidimensional structure according to claim 1, characterized in that, In step S2, the method for dividing cap-related reads and non-cap-related reads is as follows: reads with 4 or more non-template G bases in the R1 start region are cap-related reads, and reads with 3 or more non-template G bases in the R1 start region are non-cap-related reads. In step S2, the specific operation of extracting tail-related reads from the R2 read and the full-length spliced ​​read is as follows: extracting reads containing continuous A / T feature sequences from the R2 read and the full-length spliced ​​read, thus obtaining tail-related reads; wherein, the continuous A / T feature sequence refers to the sequence with a length of continuous A or continuous T bases at the end of the sequence greater than or equal to 8 nt.

3. The method for joint detection of RNA multidimensional structure according to claim 1, characterized in that, In step S3, the m7G hat structure-related information includes m7G hat location information and m7G hat modification level information; the poly(A) tail-related information includes poly(A) tail location information, poly(A) tail length information, and APA window information.

4. The method for joint detection of RNA multidimensional structure according to claim 3, characterized in that, In step S3, the specific operations for analyzing the m7G cap structure at the 5' end of RNA using cap-related reads and non-cap reads include: independently aligning cap-related reads and non-cap reads to the reference genome of the corresponding species to obtain the location information of cap-related reads and non-cap reads on the reference genome; based on the location information, statistically analyzing the genomic distribution enrichment differences of cap-related reads and non-cap reads, identifying m7G cap start sites, and obtaining m7G cap site information; statistically analyzing the number and enrichment abundance of cap-related reads corresponding to each gene and each transcript, quantifying the m7G cap modification level of each transcript, and obtaining m7G cap modification level information. The method for obtaining the poly(A) tail locus information is as follows: the tail-related reads are aligned to the reference genome of the corresponding species to obtain the location information of the tail-related reads on the reference genome; based on the enrichment information of the tail-related reads on the reference genome, combined with the end coordinates of the continuous A feature sequences in the tail-related reads, the poly(A) tail locus of each transcript is located to obtain the poly(A) tail locus information; the method for obtaining the poly(A) tail length information is as follows: the poly(A) tail length is calculated based on the length of the continuous A / T feature sequences in the R2 reads to obtain the poly(A) tail length information; The method for obtaining the APA window information is as follows: analyze all poly(A) tail sites of the same gene, screen out multiple adjacent poly(A) tail sites located in the same peak region, and calculate the maximum distance between these adjacent poly(A) tail sites, which is the APA window.

5. The method for joint detection of RNA multidimensional structure according to claim 3, characterized in that, The m7G cap structure related information also includes a quantitative expression matrix of the m7G cap structure. The method for generating the quantitative expression matrix of the m7G cap structure is as follows: the cap reads and non-cap reads covering the selected reliable m7G cap sites are separately constructed into a BAM file, and then feature counts are used. The software statistically analyzes the read coverage depth of each m7G cap site as the expression level information of the m7G cap, obtaining the expression level matrix of a single m7G site; the modification level of the m7G cap is the proportion of the cap read covering a specific m7G site to all reads of that m7G ​​site, obtaining the modification level matrix of a single m7G site; the poly(A) tail related information also includes poly(A) tail base modification information, which is obtained by scanning the base composition at the end of the poly(A) tail and determining the poly(A) tail base modification information based on the enrichment type of non-A bases at the end of the poly(A) tail; wherein, the poly(A) tail end is enriched with U bases, which is determined to be uridine acidification modification; the poly(A) tail end is enriched with G bases, which is determined to be guanylic acidification modification.

6. The method for joint detection of RNA multidimensional structures according to claim 1, characterized in that, In step S1, the method for constructing the cDNA library for RNA sequencing includes: (1) Obtain RNA from the sample to be tested, and randomly break the RNA to obtain fragmented RNA; (2) Repair the 3' end of the fragmented RNA so that the 3' end of the fragmented RNA is suitable for connecting RNA adapters; (3) Connect the RNA adapter to the 3' end of the fragmented RNA after step (2) to obtain fragmented RNA with RNA adapter; use the fragmented RNA with RNA adapter as a template, perform reverse transcription using primers complementary to the RNA adapter, and add template conversion oligonucleotides to the reverse transcription system. After the reverse transcription is completed, cDNA is obtained. (4) Perform PCR amplification on the cDNA obtained in step (3), and purify the PCR amplification product to obtain the cDNA library for RNA sequencing.

7. The method for joint detection of RNA multidimensional structures according to claim 1, characterized in that, In step S1, the method for constructing the cDNA library for RNA sequencing includes: (1) Obtain RNA from the sample to be tested, and randomly break the RNA to obtain fragmented RNA; (2) Repair the 3' end of the fragmented RNA so that the 3' end of the fragmented RNA is suitable for connecting RNA adapters; (3) Connect the RNA adapter to the 3' end of the fragmented RNA after step (2) to obtain fragmented RNA with RNA adapter; use the fragmented RNA with RNA adapter as a template, perform reverse transcription using primers complementary to the RNA adapter, and add template conversion oligonucleotides to the reverse transcription system. After the reverse transcription is completed, cDNA is obtained. (4) Perform PCR amplification on the cDNA obtained in step (3) to obtain double-stranded cDNA product; (5) Perform in vitro reverse transcription amplification on the double-stranded cDNA product obtained in step (4) to obtain RNA product; enrich the poly(A) positive mRNA in the RNA product; process the poly(A) positive mRNA according to the above steps (1)-(4) to obtain the double-stranded cDNA product, which is the cDNA library for RNA sequencing.

8. The method for joint detection of RNA multidimensional structures according to claim 6 or 7, characterized in that, In step (1), the sample to be tested includes a cell sample or a tissue sample, wherein the cell sample is an oocyte, stem cell or tumor cell; the tissue sample is a clinical biopsy tissue sample; in step (3), the template-converting oligonucleotide contains a template-converting anchor sequence and a blocking modification, and the reverse transcriptase used in the reverse transcription process has reverse transcription activity and terminal transferase activity.

9. A multidimensional RNA structure analysis system, characterized in that, include: The system comprises a library construction module, a computational analysis module, and an output module; wherein the library construction module is used to construct a cDNA library for RNA sequencing; the computational analysis module is used to perform paired-end sequencing on the constructed cDNA library for RNA sequencing using the RNA multidimensional structure joint detection method according to any one of claims 1-8, and to perform computational analysis based on the sequencing data to obtain information related to the 5' m7G cap structure and the 3' poly(A) tail of RNA; and the output module is used to output the joint analysis results of the information related to the 5' m7G cap structure and the 3' poly(A) tail of RNA.

10. The application of the RNA multidimensional structure combined detection method according to any one of claims 1-8 or the RNA multidimensional structure analysis system according to claim 9 in oocyte maturation quality assessment, oocyte screening in assisted reproduction, optimization of in vitro maturation culture, RNA metabolism mechanism research, or detection of RNA structure in low input samples.