A 3'race library sequencing method based on third-generation sequencing technology and application thereof

By constructing libraries using third-generation sequencing technology combined with specific primers, long-read sequencing can be performed directly, solving the problems of cumbersome procedures and PCR bias in RACE technology. This achieves efficient and accurate acquisition of cDNA 3' end sequences, which is suitable for molecular biology research.

CN122235273APending Publication Date: 2026-06-19WUHAN BIORUN BIO TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN BIORUN BIO TECH
Filing Date
2026-04-17
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing RACE technology has a cumbersome process and low throughput when obtaining cDNA 3' end sequences, making it difficult to meet the needs of large-scale, high-precision research. The PCR amplification and cloning steps have sequence bias, which leads to the loss or misjudgment of information on low-abundance transcripts and complex transcripts, affecting the accuracy of sequencing results.

Method used

The library was constructed using third-generation sequencing technology combined with specific primers, and long-read sequencing was performed directly, avoiding traditional cloning and Sanger sequencing. Through reverse transcription, specific amplification and third-generation sequencing adapter ligation, high-throughput, splice-free acquisition of complete 3' end sequences was achieved.

Benefits of technology

It improves the efficiency and accuracy of 3' Race assays, simplifies the operation process, reduces costs, and can obtain complete 3' end sequence information in one go. It is applicable to RNA molecules with complex structures and meets the needs of molecular biology research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122235273A_ABST
    Figure CN122235273A_ABST
Patent Text Reader

Abstract

This invention relates to a 3'RACE library construction and sequencing method and its application based on third-generation sequencing technology. The method includes: S1, using primers containing anchor sequences and Oligo d(T) to reverse transcribe sample RNA to obtain cDNA, which is then used as a template for specific amplification to obtain the target gene; S2, designing specific primers containing third-generation sequencing adapters and barcodes, using the target gene as a template for amplification, and constructing a high-throughput third-generation sequencing library; S3, performing third-generation sequencing and analysis on the library to obtain the complete 3' end sequence information of the target transcript. This invention abandons traditional vector construction and Sanger sequencing methods, leveraging the high throughput and long read length characteristics of third-generation sequencing to ensure that Race experiments obtain a large amount of long 3' end sequence information in a short period, improving the overall efficiency and data accuracy of the experiment, and laying the foundation for refined mRNA research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of cDNA end amplification, specifically relating to a 3'RACE library preparation and sequencing method based on third-generation sequencing technology and its application. Background Technology

[0002] RACE (Rapid amplification of cDNA ends) is a molecular biology technique based on reverse transcription that rapidly amplifies the 5' and 3' ends of cDNA from a sample. It was first proposed by Frohman et al. in 1988. This technique uses primers designed with known partial cDNA sequences to extend the cDNA to both ends, obtaining the complete 5' and 3' ends. Due to its unique advantages, RACE technology has been applied in various fields, including cDNA library construction, target gene cloning, 5' and 3' UTR function studies, and new sequence discovery.

[0003] As research into gene function and regulatory mechanisms deepens, the application of traditional race technology, while becoming increasingly widespread, has also revealed numerous problems, particularly in obtaining true and complete terminal sequences. Currently, most RACE methods still rely on traditional Sanger sequencing, which has the following limitations: First, the initial cloning and sequencing preparation is costly and time-consuming. Furthermore, Sanger sequencing has limited read length and low throughput, often requiring segmented cloning and splicing to obtain complete terminal sequences, resulting in limited throughput and overall low efficiency. Second, there is a significant template bias in PCR amplification; high-abundance transcripts can easily mask low-abundance target sequences, leading to the loss of important transcript information, particularly those with low abundance or complex structures. These problems often result in terminal sequence deletions, chimera formation, or sequence misinterpretation, affecting the accuracy and reliability of experimental results.

[0004] Although improved methods such as SMARTer RACE and nested RACE have been developed to enhance specificity and sensitivity, they still fundamentally rely on PCR amplification and Sanger sequencing. Therefore, their success rate remains unsatisfactory when dealing with low-expression genes, high-GC-content regions, or complex repetitive sequences. While existing optimized next-generation sequencing (NGS)-based schemes have improved throughput to some extent, their short read length characteristics still present challenges in resolving long 3' UTRs, identifying distal poly(A) sites, or accurately determining alternative splicing events, due to splicing difficulties and incomplete information.

[0005] Given that current RACE techniques involve cumbersome experimental procedures, long cycles, high technical requirements for operators, and results are greatly affected by PCR bias and sequencing length limitations, they are difficult to meet the current research needs for efficient and accurate analysis of full-length transcripts. Therefore, there is an urgent need in this field for a new generation of RACE technology that can simplify the operation process, fundamentally avoid PCR bias, improve the success rate of obtaining end sequences, and support the simultaneous detection of long fragment readings and epigenetic information. Summary of the Invention

[0006] This invention aims to address the technical bottlenecks of existing 3'Race technology in obtaining cDNA 3' end sequences. These bottlenecks include cumbersome procedures, low throughput, difficulty in meeting the needs of large-scale, high-precision research, and inherent sequence biases in PCR amplification and cloning steps. This leads to the inaccurate and complete capture of low-abundance transcripts, transcripts with complex secondary structures, or long 3'UTR fragments, often resulting in lost or misinterpreted end information and affecting the accuracy of sequencing results. This invention provides a method for constructing, sequencing, and analyzing cDNA end sequencing libraries based on a third-generation sequencing (long-read sequencing) platform, along with its applications. The core technical concept of this invention is to abandon traditional cloning construction and Sanger sequencing, as well as the PCR-biased second-generation sequencing library construction process, and cleverly combine 3'Race with third-generation sequencing technology to construct libraries that can be directly sequenced for long reads, thereby obtaining continuous, single-molecule sequence information from the known region to the poly(A) tail in a single step. This invention provides a high-fidelity solution that can directly obtain long-read, splice-free, complete 3' end sequences and reduce amplification bias from the source. This solution effectively improves the efficiency, accuracy, and comprehensiveness of 3' Race assays, meets the needs of various molecular biology experiments, and has broad application prospects.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: This invention provides a method for rapid 3' end amplification of cDNA ends for library construction and sequencing based on third-generation sequencing technology, comprising the following steps: S1. Based on the poly(A) structure of the transcript, the sample RNA is reverse transcribed using primers containing anchor sequences and Oligo d(T) to obtain cDNA, and the cDNA is used as a template for specific amplification to obtain the target gene; S2. The sequencing adapter and barcode are ligated to the target gene, and after purification, a high-throughput third-generation sequencing library containing the target gene fragment is constructed. S3. The third-generation sequencing library constructed in step S2 is subjected to high-throughput single-molecule sequencing to obtain raw reads of thousands of bases or even longer. By performing bioinformatics analysis on these long reads, including sequence alignment, end identification and quality correction, the complete 3' end sequence information of the target transcript can be obtained directly and accurately without complex splicing, including the precise poly(A) start site and tail length.

[0008] Preferably, step S1 specifically includes: (1) The sample RNA was reverse transcribed using primers containing the anchoring sequence and Oligo d(T) to obtain first-strand cDNA with a universal primer binding site complementary to the anchoring sequence at the 5' end; (2) Design a downstream primer that is complementary to the binding site of the universal primer and an upstream primer F1 that can specifically bind to the target transcript, and perform the first round of PCR amplification using the first strand cDNA as a template to enrich the target gene fragment. (3) Using the target gene fragment obtained in step (2) as a template, a second round of PCR amplification is performed using a downstream primer that is complementary to the binding site of the universal primer and an upstream primer F2 that can specifically bind to the target transcript to obtain the target gene.

[0009] Preferably, the primer containing the anchoring sequence and Oligo d(T) can be designed as a locking primer to improve the specificity of binding to the poly(A) junction sequence.

[0010] Preferably, the primer sequence containing the anchoring sequence and Oligo d(T) is shown in SEQ ID NO.5; and the downstream primer sequence that is complementary to the binding site of the universal primer is shown in SEQ ID NO.6.

[0011] Preferably, the binding site of the upstream primer F2 in step (3) is closer to the 3' end of the target transcript than that of primer F1.

[0012] Preferably, the first round of PCR amplification in step (2) and the second round of PCR amplification in step (3) can use high-fidelity enzymes and be optimized to minimize amplification errors and effectively amplify long fragments.

[0013] Preferably, the third-generation sequencing technology includes: single-molecule real-time sequencing technology or nanopore sequencing technology.

[0014] Preferably, the target transcript is derived from eukaryotes. The above method has specific biological system applicability, and is particularly suitable for eukaryotic mRNAs with typical poly(A) tails.

[0015] Preferably, the eukaryotes include: plants, animals, or fungi.

[0016] Preferably, when the target transcript is the AT3G013530 gene from Arabidopsis thaliana, primer F1 is as shown in SEQ ID NO.1 and primer F2 is as shown in SEQ ID NO.3.

[0017] Preferably, when the target transcript is the LOC_Os02g52650 gene derived from rice, primer F1 is as shown in SEQ ID NO.2, and primer F2 is as shown in SEQ ID NO.4. This invention also provides the application of any of the above-described methods in any of the following: A1) Application in obtaining or identifying the 3' end sequence of transcripts; A2) Application in obtaining information on gene transcription termination sites; A3) Application in obtaining alternative splicing of transcripts.

[0018] Preferably, obtaining or identifying the 3' end sequence of the transcript includes: obtaining or identifying the 3' UTR sequence, and / or, performing gene structure annotation.

[0019] Preferably, the method can be applied to accurately identify the 3' untranslated region sequence and transcription termination / polyadenylation site of a transcript.

[0020] Preferably, the method can be applied to obtain the complete 3' end sequence of a gene containing an ultra-long 3' UTR or a complex structural domain (such as one rich in repetitive sequences).

[0021] Preferably, the method can be applied to study the coupling events of variable polyadenylation and variable splicing at the full-length level.

[0022] Preferably, the method can be applied to the discovery and study of novel transcript 3' isoforms to improve genome annotation.

[0023] Beneficial effects: (1) Based on the traditional 3' Race technology, this invention abandons the traditional cloning construction and Sanger sequencing, and proposes an innovative sequencing strategy. By designing specific primers containing third-generation sequencing adapters to construct libraries, and using a third-generation sequencing platform, the sequence information of the 3' ends can be obtained in high throughput. This method solves the problems of low throughput and bias in traditional Sanger sequencing, and significantly improves the efficiency of Race experiments while simplifying the experimental process and reducing costs.

[0024] (2) The experimental results show that the 3'RACE technology (3'Race-Nano-seq) based on third-generation sequencing technology of the present invention can effectively improve the comprehensiveness of cDNA end amplification, especially when processing RNA molecules with complex structures. For example, the traditional 3'Race technology can only capture terminal sequences with high abundance, while the 3'Race-Nano-seq technology of the present invention can effectively detect terminal information of various morphologies of the gene, improve the comprehensiveness of 3'Race in obtaining 3'cDNA terminal information, and meet the needs of various molecular biology experiments. Attached Figure Description

[0025] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 These are the traditional 3'Race technology scheme and the 3'Race-Nano-seq optimization scheme of the present invention in the embodiments of the present invention, where a is the traditional 3'Race technology route and b is the 3'Race-Nano-seq technology route of the present invention.

[0027] Figure 2 This is a successful 3' Race-Nano-seq case based on Nano third-generation sequencing technology in an embodiment of the present invention. In the figure, a is the 3' Race-Nano-seq amplification electrophoresis image of the Arabidopsis a. AT3G03530 gene, b is the sequencing result of the AT3G03530 gene compared with the reference sequence, c is the 3' Race-Nano-seq amplification electrophoresis image of the rice LOC_Os02g52650 gene, and d is the sequencing result of the LOC_Os02g52650 gene compared with the reference sequence. Detailed Implementation

[0028] The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, and are therefore merely examples and should not be used to limit the scope of protection of the present invention. It should be noted that, unless otherwise stated, the technical or scientific terms used in this application should have the ordinary meaning understood by those skilled in the art. Unless specifically stated, the reagents, methods, and equipment used in this invention are conventional reagents, methods, and equipment in this technical field. Unless specifically stated, the reagents and materials used in the following embodiments are commercially available.

[0029] Example The traditional 3' Race technique involves: RNA extraction, reverse transcription to generate cDNA, specific amplification of the 3' fragment, gel extraction and recovery, ligation into a T vector, transformation of competent cells, selection of single clones for PCR detection, and then Sanger sequencing for alignment. The amplification process is often time-consuming and labor-intensive due to vector construction and Sanger sequencing, and the data output throughput is low. For long fragments, segmented cloning and splicing are required to obtain complete end sequences, resulting in low overall efficiency. To address this issue, this invention, based on traditional methods, omits cumbersome procedures such as vector construction, single-clone selection, and Sanger sequencing, proposing a novel detection method: After obtaining cDNA through reverse transcription, it is used as a template, and specific primers containing third-generation sequencing adapters are designed for PCR amplification. Barcodes are added to the target sequence using PCR, constructing a third-generation sequencing library of the target sequence. Sequencing is then performed on a third-generation sequencing platform to capture a large amount of target sequence information in a single pass. This leverages the high throughput and long read length advantages of third-generation sequencing to capture a large amount of ultra-long target sequence information in one go, effectively shortening experimental costs and timelines, and significantly increasing the capture efficiency of the cDNA 3' ends. For details of the technical route, please refer to [link to technical details]. Figure 1 .

[0030] This embodiment uses the target genes AT3G013530 (Arabidopsis thaliana phosphatase family protein NPC4) and LOC_Os02g52650 (rice light-harvesting chlorophyll a / b binding protein) as examples to illustrate the technical route and sequencing results of the 3'RACE (3'Race-Nano-seq) based on third-generation sequencing technology of this invention.

[0031] Specifically, amplifying the target genes AT3G013530 and LOC_Os02g52650 includes the following steps: 1. RNA extraction from Arabidopsis thaliana and rice leaves Leaves from Arabidopsis thaliana and rice were collected separately, rapidly frozen, and ground into powder. Total RNA was extracted using Trizol reagent or a commercially available RNA extraction kit (such as Vazyme's FastPure Universal Plant Total RNA Isolation Kit) following the instruction manual.

[0032] RNA quality testing: RNA samples are tested by 1% agarose gel electrophoresis. If the results show relatively complete 28S and 18S rRNA bands, the integrity of the total RNA is good and can be used for subsequent experiments.

[0033] 2. Primer design AT3G03530 is a gene derived from Arabidopsis thaliana, and LOC_Os02g52650 is a gene derived from rice. The complete sequences of these transcripts were obtained from reference databases. The cDNA reference sequence of the AT3G03530 gene is shown in SEQ ID NO.7, and the cDNA reference sequence of the LOC_Os02g52650 gene is shown in SEQ ID NO.8.

[0034] Based on the above reference sequences, specific primers NPC4-F1 and NPC4-F2 (specific sequences shown in SEQ ID NO.1 and SEQ ID NO.3 in Table 1) were designed for nested PCR amplification of the AT3G03530 gene, and specific primers 9-F1 and 9-F2 (specific sequences shown in SEQ ID NO.2 and SEQ ID NO.4 in Table 1) were designed for nested PCR amplification of the LOC_Os02g52650 gene.

[0035] Table 1 Primer Sequences 3. Synthesis of the first strand of cDNA (1) Preparation of the initial reaction system The reaction system was prepared in RNase-free test tubes, including Total RNA, OSYB01 (reverse transcription primer), dNTPMix, and RNase-free H2O. The reaction system is shown in Table 2. The control group was Oligo(dT). 20 Replace OSYB01 and prepare the reaction system.

[0036] Table 2 Initial reaction system First, heat the mixture to 72°C for 3 minutes to unwind the secondary structure of the RNA. Then immediately cool it on ice for 2 minutes to allow the primers to specifically bind to the poly(A) at the 3' end of the mRNA and inhibit non-specific binding.

[0037] (2) Preparation of reverse transcription reaction system The reaction system includes: the above reaction products, 5 × RT buffer, and M-MLV Reverse Transcriptase. The concentrations of each component are shown in the table below: Table 3 Reverse transcription reaction system The reaction conditions were as follows: reaction at 42℃ for 30 minutes to synthesize the first strand of cDNA via reverse transcription, followed by reaction at 85℃ for 5 minutes to shut down enzyme activity and stabilize the cDNA. After the reaction, the cDNA was stored at 4℃ for a short period. Primer OSYB01 contains an anchoring sequence and a poly(dT) band. The poly(dT) band specifically binds to the poly(A) band at the 3' end of eukaryotic mRNA, initiating the reverse transcription reaction to synthesize the first strand of cDNA. A universal primer binding site complementary to the OSYB01 anchoring sequence was introduced at the 5' end of the cDNA. Subsequent specific amplification and enrichment were performed using the complementary primer OSY364.

[0038] 4. Nested PCR amplification of the target gene fragment (1) First round of PCR amplification: After reverse transcription, the target gene fragment is amplified in the first round of PCR. The first strand cDNA is used as a template. PCR amplification is performed using primer OSY364 (as shown in SEQ ID NO.6 in Table 1), which is complementary to the binding sites of the above universal primers, and gene-specific primer F1 (NPC4-F1 or 9-F1) to specifically enrich the target gene fragment.

[0039] The PCR reaction system includes: 2 × PCR Mix, NPC4-F1, OSY364, template, and ddH2O. The concentrations of each component and the reaction procedures are shown in Tables 4 and 5 below: Table 4 PCR amplification system Table 5 PCR reaction procedure (2) Product detection and second round of PCR amplification: The first round of PCR products were analyzed by agarose gel electrophoresis. The electrophoretic fragments were recovered and purified and used as templates. The second round of PCR amplification was performed using primer OSY364 and another gene-specific primer F2 (NPC4-F2 or 9-F2) which is closer to the inner side of the 3' end of the target gene, so as to significantly improve the amplification specificity and sensitivity and effectively reduce non-specific background.

[0040] 5. Recovery and purification of the target fragment The second round of PCR was detected by electrophoresis, and the results are as follows: Figure 2 As shown in a and 2c, the results indicate that the target band was successfully amplified. Then, the target band was excised and purified according to the instructions of the DNA gel extraction kit (Nanjing Novizan Biotechnology Co., Ltd.).

[0041] 6. Perform rapid third-generation library construction on the target fragment (taking Nano sequencing as an example). A high-throughput third-generation sequencing library was constructed by adding third-generation sequencing adapters and barcodes to the purified PCR products using a high-speed barcode library construction kit (purchased from Beijing Qitan Technology Co., Ltd.).

[0042] 7. Sequencing and Alignment After concentration and quality testing of the library, it was mixed with sequencing reagents and loaded onto a nanopore sequencing chip for third-generation sequencing (using a QPursue-6k-hex sequencer from Beijing Qitan Technology Co., Ltd.). After obtaining the sequencing data, Snapgene was used to align the sequencing results with a reference sequence. The sequencing results are as follows: Figure 2 b and Figure 2 As shown in d, the results show that the 3' end sequence information of the target gene matches the reference sequence, indicating that the 3' Race-Nano-seq technology of the present invention can successfully sequence the target gene without any splicing, and obtain the full length in one read. This breaks through the low throughput limitation of traditional Sanger sequencing, and the results are reliable and the efficiency is effectively improved.

[0043] In summary, the 3' Race-Nano-seq technology based on third-generation sequencing exhibits significant advantages in experimental cycle, cost, and ease of operation, and can effectively improve the capture efficiency of cDNA 3' ends by leveraging the advantage of large data volume. In conclusion, this invention provides a more efficient, economical, and convenient tool for race experiments and downstream molecular biology research, possessing broad application prospects and significant scientific value.

[0044] The above detailed embodiments describe the implementation of the present invention; however, the present invention is not limited to the specific details described in the above embodiments. Within the scope of the claims and technical concept of the present invention, various simple modifications and changes can be made to the technical solution of the present invention, and these simple modifications all fall within the protection scope of the present invention.

Claims

1. A method for rapid 3' end amplification of cDNA ends for library construction and sequencing based on third-generation sequencing technology, characterized in that, Includes the following steps: S1. Based on the poly(A) structure of the transcript, the sample RNA is reverse transcribed using primers containing anchor sequences and Oligo d(T) to obtain cDNA, and the cDNA is used as a template for specific amplification to obtain the target gene; S2. The sequencing adapter and barcode are ligated to the target gene, and after purification, a high-throughput third-generation sequencing library containing the target gene fragment is constructed. S3. Perform high-throughput third-generation sequencing and analysis on the library to obtain the complete 3' end sequence information of the target transcript.

2. The method according to claim 1, characterized in that, Step S1 specifically includes: (1) The sample RNA was reverse transcribed using primers containing the anchoring sequence and Oligo d(T) to obtain first-strand cDNA with a universal primer binding site complementary to the anchoring sequence at the 5' end; (2) Design a downstream primer that is complementary to the binding site of the universal primer and an upstream primer F1 that can specifically bind to the target transcript, and perform the first round of PCR amplification using the first strand cDNA as a template to enrich the target gene fragment. (3) Using the target gene fragment obtained in step (2) as a template, a second round of PCR amplification is performed using a downstream primer that is complementary to the binding site of the universal primer and an upstream primer F2 that can specifically bind to the target transcript to obtain the target gene.

3. The method according to claim 2, characterized in that, The primer sequence containing the anchoring sequence and Oligo d(T) is shown in SEQ ID NO.5; the downstream primer sequence that is complementary to the binding site of the universal primer is shown in SEQ ID NO.

6.

4. The method according to claim 2, characterized in that, The binding site of upstream primer F2 in step (3) is closer to the 3' end of the target transcript than that of primer F1.

5. The method according to claim 1, characterized in that, The third-generation sequencing technologies include: single-molecule real-time sequencing technology or nanopore sequencing technology.

6. The method according to claim 1, characterized in that, The target transcript is derived from eukaryotes.

7. The method according to claim 6, characterized in that, The eukaryotes include: plants, animals, or fungi.

8. The application of the method according to any one of claims 1-7 in any of the following: A1) Application in obtaining or identifying the 3' end sequence of transcripts; A2) Application in obtaining information on gene transcription termination sites; A3) Application in obtaining alternative splicing of transcripts.

9. The application according to claim 8, characterized in that, The acquisition or identification of the 3' end sequence of the transcript includes: acquiring or identifying the 3' UTR sequence, and / or, performing gene structure annotation.