High-throughput double-end RNA-seq rapid library building method
By using one-step RT-PCR and transposase integration technology, high-throughput paired-end RNA-seq library construction has been achieved, solving the problems of cumbersome procedures, insufficient high throughput, and loss of transcript information in RNA-seq library construction technology. It is suitable for rapid and efficient library construction of low starting sample and meets the high-throughput RNA sequencing needs of scenarios such as multi-drug screening and plant breeding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-05
- Publication Date
- 2026-03-27
AI Technical Summary
Existing RNA-seq library preparation technology has a cumbersome process, insufficient high throughput, loss of transcript information, and weak adaptability to low starting sample volume, making it difficult to meet the high-throughput RNA sequencing needs of multi-drug screening and plant breeding.
A one-step RT-PCR and transposase integration technique was adopted, using a 3' end specific oligo(dT)-UMI-barcode-adaptor sequence C primer and a 5' end pore specific/universal TSO-barcode primer to simultaneously perform reverse transcription and amplification, and then perform transposition reaction in combination with the transposase complex to achieve high-throughput library construction of paired-end labeled cDNA.
It achieves high-throughput sample discrimination capability, simplified library construction process, high efficiency, and accurate quantification, significantly reducing costs. It is suitable for rapid high-throughput RNA sequencing of low starting sample volume, meeting the needs of scenarios such as multi-drug screening and plant breeding.
Smart Images

Figure CN121737270A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biotechnology. More specifically, it relates to a high-throughput method for rapid library construction using paired-terminal RNA-seq. Background Technology
[0002] In recent years, RNA-seq technology has been continuously developing due to its applications in gene expression, translation, regulatory RNA, and RNA epigenetics. Among these applications, the most common analysis is identifying differential gene expression (DGE), which provides guidance for the early initiation and direction determination of RNA research projects.
[0003] Today, various RNA-seq technologies have been derived from standard RNA-seq methods. For example, the 3' RNA-seq method using mRNA as a template, published in Nature Methods, is Lexogen's Quantseq library construction kit. During Quantseq library construction, only one fragment is generated per transcript, resulting in only one-tenth the data volume of conventional RNA-seq. It can also be used with UMI to perform single-molecule tagging of second-strand cDNA, making gene expression quantification more precise. Quantseq uses oligo(dT) primers for specific reverse transcription of mRNA, while the second strand is synthesized using random primers. The distance between the random primer binding site and the poly(A) determines the length of the inserted fragment, eliminating the need for fragmentation. PCR is performed directly after cDNA synthesis, significantly shortening the library construction process. This method can achieve the same sensitivity level as standard RNA-seq at low sequencing depths and can enable simultaneous sequencing of multiple libraries.
[0004] Due to the advantages of low cost and accurate gene expression quantification, 3' RNA-seq technology has been actively explored and developed by scientists in recent years. Major improvements include simplifying the library preparation process to further reduce costs, and improving the accuracy of gene expression quantification through the use of tags and UMI. However, the library preparation process is often quite cumbersome, requiring not only a separate two-strand synthesis of cDNA, but also end repair and adapter ligation after cDNA fragmentation before amplification and library construction. Furthermore, transcriptome libraries constructed only from the 3' end lose a significant portion of the 5' end information, resulting in only one end read being valid data during bioinformatics analysis.
[0005] While conventional transcriptome sequencing can obtain relatively comprehensive transcript information, the library construction process is lengthy and cumbersome, often requiring steps such as fragmentation, first-strand synthesis, second-strand synthesis, adapter ligation, purification, amplification, and further purification. Even with relatively rapid whole-transcriptome library construction, due to sample differentiation, a set of i5 and i7 tags can only correspond to one library, making it difficult to achieve high-throughput library construction.
[0006] Therefore, a paired-end RNA-seq library preparation method is needed to solve the technical problems of RNA-seq library preparation technology, such as cumbersome process, insufficient high throughput, loss of transcript information, and weak ability to adapt to low starting sample volume. Summary of the Invention
[0007] The purpose of this invention is to provide a high-throughput, rapid, and efficient method for paired-end RNA-seq library construction. This method is fast, efficient, and streamlined, enabling high-throughput sample differentiation. It can simultaneously capture the 5' and 3' ends of transcripts, and can also flexibly select single-end library construction to simultaneously construct libraries from the 5' and 3' ends of cDNA. It can perform high-throughput library construction on mRNA from hundreds of different cell types (e.g., drug screening) in a single run. It is suitable for low-volume samples (small number of cells / trace amounts of RNA), meeting the high-throughput RNA sequencing needs of a wide range of screening scenarios such as multi-drug screening and plant breeding, and providing technical support for the early and rapid acquisition of target genes.
[0008] To achieve the above objectives, the present invention adopts the following technical solution: This invention first provides a high-throughput method for rapid library construction using paired-terminal RNA-seq, comprising the following steps: S1. Obtaining total RNA by cell lysis: In an RNase-free environment, treated eukaryotic cells were lysed with Triton lysis buffer containing RNase inhibitor, and the lysis products were collected. S2. Obtaining paired-end labeled cDNA by one-step RT-PCR: Using the lysis product as a template, a 3' end pore-specific oligo(dT)-UMI-barcode-adaptor sequence C primer and a 5' end pore-specific / universal TSO-barcode primer were used to simultaneously complete reverse transcription and amplification through a one-step RT-PCR reaction to obtain paired-end labeled cDNA with the 3' end carrying the UMI-barcode-adaptor sequence C and the 5' end carrying the TSO-barcode. S3. Transposition reaction to obtain DNA fragments with partial transposable adapters, paired-end pore tags and target cDNA fragments: After purifying the paired-end labeled cDNA mixture, a purified cDNA mixture product is obtained. A transposase complex is added to carry out a transposition reaction to obtain DNA fragments with partial transposable adapters, paired-end pore tags and target cDNA fragments. S4. Indexed PCR amplification to introduce sequencing index: DNA fragments are amplified by PCR using indexed primer pairs to introduce sequencing index and obtain RNA-seq sequencing libraries.
[0009] In a specific embodiment of the present invention, the 3' end well-specific oligo(dT)-UMI-barcode-adaptor sequence C primer comprises a barcode, UMI, oligo(dT) sequence, and adapter sequence C, with each well corresponding to a unique barcode sequence, used to initiate the first strand synthesis of cDNA and introduce the 3' end carrying the UMI-barcode-adaptor sequence C sequence; the 5' end well-specific or universal TSO-barcode primer comprises a barcode and a TSO sequence, wherein the barcode is one sequence adapted to 96 wells, or 96 sequences adapted to 9216 samples, and the 5' end carrying the TSO-barcode sequence is introduced through a template conversion reaction; wherein the barcode sequence is any one of SEQ ID NO. 1 to 96.
[0010] In a specific embodiment of the present invention, the one-step RT-PCR reaction system includes: lysis products, 3' end pore-specific oligo(dT)-UMI-barcode-adaptor sequence C primer, 5' end pore-specific / universal TSO-barcode primer, RT-PCR buffer, and an enzyme mixture; wherein the enzyme mixture contains reverse transcriptase and polymerase.
[0011] In a preferred embodiment of the present invention, the one-step RT-PCR reaction system comprises: 2 μL of lysis product, 2 μL of 5 μM 3' end pore-specific oligo(dT)-UMI-barcode-adaptor sequence C primer, 0.5 μL of 10 μM 5' end pore-specific / universal TSO-barcode primer, 14 μL of RT-PCR buffer, and 1.5 μL of enzyme mixture.
[0012] In a specific embodiment of the present invention, the one-step RT-PCR reaction program is as follows: 42℃, 60 min; 95℃, 3 min; (98℃, 20 sec; 65℃, 45 sec; 72℃, 3 min) × 4 cycles; (98℃, 20 sec; 67℃, 20 sec; 72℃, 3 min) × 11 cycles; 72℃, 5 min; hold at 4℃.
[0013] In a specific embodiment of the present invention, the transposase complex includes transposase complex A and / or transposase complex B; transposase complex A is composed of a transposase and a transposon of coupled linker sequence A, and a library is constructed at the 5' end; transposase complex B is composed of a transposase and a transposon of coupled linker sequence B, and a library is constructed at the 3' end; transposase complex A and transposase complex B are used in a 1:1 ratio to achieve dual-end library construction.
[0014] In a specific embodiment of the present invention, the amount of the purified cDNA mixture is 50-200 ng; the transposition reaction system includes the cDNA mixture, transposition reaction buffer, transposase complex and RNase-free water.
[0015] In a preferred embodiment of the present invention, the transposition reaction system comprises 200 ng of cDNA mixture, 6 μl of 5× transposition reaction buffer, 6 μl of transposase complex, and RNase-free water to a final volume of 30 μl.
[0016] In a specific embodiment of the present invention, the transposable reaction procedure is an incubation at 55°C for 10 min.
[0017] In a specific embodiment of the present invention, the index primer pair includes index primer pair A and / or index primer pair B; wherein, index primer pair A includes P5-index(i5)_primer binding region T and P7-index(i7)_primer binding region A, the primer binding region T contained in the P5-index(i5)_primer binding region T is complementary to the TSO sequence, and the primer binding region A contained in the P7-index(i7)_primer binding region A is complementary to the transposon adapter sequence A. Sequencing indexes i5 and i7 are introduced through index PCR reaction to realize cDNA sequencing. 5' end library construction: The index primer pair B includes a P5-index(i5) primer binding region B and a P7-index(i7) primer binding region C. The primer binding region B in the P5-index(i5) primer binding region B is complementary to the transposon adapter sequence B, and the primer binding region C in the P7-index(i7) primer binding region C is complementary to the adapter sequence C. Sequencing indexes i5 and i7 are introduced through an index PCR reaction to achieve cDNA 3' end library construction. If index primer pair A and index primer pair B are used simultaneously, sequencing indexes i5 and i7 can be introduced through a single index PCR reaction to achieve cDNA paired-end library construction. Finally, a qualified RNA-seq sequencing library carrying barcode, UMI, sequencing index, and complete target sequence is obtained. This library is compatible with high-throughput sequencing platforms and allows for sample splitting and accurate transcript information analysis through barcode, UMI, and sequencing index (i5 / i7).
[0018] In a specific embodiment of the present invention, the indexed PCR reaction system includes a DNA fragment, a PCR Mix, index primer pair A and / or index primer pair B.
[0019] In a preferred embodiment of the present invention, the indexed PCR reaction system comprises 20 μl DNA fragment, 25 μl 2× high-fidelity PCR Mix, 2.5 μl index primer pair A and / or 2.5 μl index primer pair B.
[0020] In a specific embodiment of the present invention, the indexed PCR reaction program is 72°C, 3 min; 98°C, 3 min; (98°C, 10 sec; 62°C, 30 sec) × 7 cycles; 72°C, 3 min; 4°C hold.
[0021] The beneficial effects of this invention are as follows: This invention successfully establishes a novel high-throughput rapid library construction method for paired-terminal RNA-seq. Through one-step RT-PCR barcode synchronous labeling and transposase integration technology, rapid high-throughput paired-terminal RNA-seq library construction is achieved, as detailed below: 1. High-throughput sample discrimination capability: By introducing a TSO sequence with a well barcode at the 5' end of the cDNA and Oligo(dT) and UMI sequences with a well barcode at the 3' end, the barcode can accurately distinguish transcript information from different treatments. Furthermore, the UMI can be used to single-molecule label the cDNA fragments in each well, so that each cDNA fragment in the mixed library has a pairwise specific tag. Since each cDNA fragment from a sub-library in the mixed library carries 5' and 3' specific tag sequences, up to 9216 different cell treatments can be used to build libraries at once, fully meeting the needs of high-throughput library construction.
[0022] 2. Simplified and efficient library construction process: Reverse transcription (RT) and PCR amplification are combined into a one-step method, and one-step amplification is performed in a high-throughput PCR plate; the cDNA from each well is mixed and purified, and the mixed cDNA product is fragmented with one or two transposase complexes according to the single-end or double-end library construction requirements before amplification, so as to efficiently construct a mixed RNA-seq library, eliminating the cumbersome steps such as adapter ligation and purification in traditional library construction, with short library construction time, simple operation, significantly shortening the experimental cycle and significantly improving library construction efficiency.
[0023] 3. High library specificity and accurate quantification: One-step RT-PCR amplification of mRNA can retain the specific information of the original RNA ends, which is convenient for subsequent mixed library construction; it can also effectively reduce non-specific amplification, concentrate reverse transcription resources on the amplification of cDNA corresponding to mRNA, and improve library specificity; and from the perspective of the number of effective genes obtained from library construction, the information between different treatment groups can be distinguished by different bars, and the number of effective genes in each well corresponding to different bars is evenly distributed, and the number of genes captured will not be affected by the base composition of the barcode itself. Combined with UMI, it can achieve accurate single-molecule quantification of transcripts.
[0024] 4. High cost-effectiveness and significantly reduced costs: A single mixed RNA-seq library is constructed using one-step RT-PCR and transposase integration technology, replacing the traditional construction mode of multiple independent libraries. This greatly reduces the amount of library construction reagents used, significantly reduces the cost of library construction and sequencing, and balances the needs of high-throughput detection with economic efficiency.
[0025] 5. Wide range of applications and strong practicality: It is suitable for high-throughput rapid RNA library construction needs in a wide range of screening scenarios such as multi-drug screening and plant breeding. It is especially important for early rapid acquisition of target genes and clarification of research directions. It can efficiently adapt to different starting amounts of samples (including small amounts of cells, trace amounts of RNA and precious samples). Attached Figure Description
[0026] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0027] Figure 1 This is a schematic diagram illustrating the principle of the high-throughput paired-terminal RNA-seq library preparation method in this invention.
[0028] Figure 2 This is a peak diagram of cDNA with paired ends obtained by RT-PCR in 293T cells using the high-throughput paired-end RNA-seq library preparation method of this invention.
[0029] Figure 3 This is a peak diagram of DNA fragments obtained after transposition reaction following cDNA mixing and purification using the high-throughput paired-terminal RNA-seq library construction method of this invention. Detailed Implementation
[0030] To more clearly illustrate the present invention, the following description, in conjunction with preferred embodiments and accompanying drawings, further explains the invention. Similar components in the drawings are indicated by the same reference numerals. Those skilled in the art should understand that the specific description below is illustrative rather than restrictive and should not be construed as limiting the scope of protection of the present invention.
[0031] This invention provides a high-throughput paired-terminal RNA-seq library preparation method, the process of which is as follows: Figure 1 As shown, the main steps include: obtaining total RNA through cell lysis; obtaining paired-end labeled cDNA (Oligo(dT)-UMI sequence with a well tag at the 3' end and a TSO transformation sequence with a well tag at the 5' end) via one-step RT-PCR; purifying the cDNA and then performing a transposition reaction to obtain DNA fragments containing partial transposon adapters, paired-end well tags, and the target cDNA fragment; indexing PCR amplification to introduce a sequencing index; and finally obtaining an RNA-seq sequencing library. The specific operation methods are as follows: I. Obtaining Total RNA from Cell Lysis Cell lysis was performed in an RNase-free environment. Before the operation, the clean bench was sprayed with alcohol, wiped with RNase removal reagent, and irradiated with ultraviolet light for more than half an hour.
[0032] Eukaryotic cells treated with drugs were lysed using Triton lysis buffer containing RNase inhibitor. 2 μL of lysis product was taken from each group and transferred to the corresponding wells of a 96-well plate, with each well corresponding to one treatment sample.
[0033] II. One-step RT-PCR for obtaining paired-end labeled cDNA 1. Prepare the RT-PCR reaction system 2 μL of lysis product, 2 μL of 5 μM 3' end well-specific oligo(dT)-UMI-barcode-adaptor sequence C primer (one unique sequence per well, as shown in Table 1, and the sequence is as shown in SEQ ID NO.1-96), 0.5 μL of 10 μM 5' end well-specific or universal TSO-barcode primer (one adapter for 96 wells can be used for the same barcode, or 96 adapters can be used for 9216 sample expansion), 14 μL of RT-PCR buffer, and 1.5 μL of enzyme mixture (containing reverse transcriptase and polymerase). The 3' oligo(dT)-UMI-barcode-adaptor sequence C primer contains oligo(dT) primers for both barcode and UMI, used to initiate first-strand cDNA synthesis and introduce the barcode and UMI into the 5' end of the cDNA. The adapter sequence C is introduced into the sequencing index during subsequent indexed PCR amplification. The 5' well-specific or universal TSO-barcode primer is a TSO primer containing barcode, used to introduce the barcode into the 3' end of the same cDNA molecule via template conversion reaction. The 96 bars of the 3' oligo(dT)-UMI-barcode-adaptor sequence C primer, together with the UMI, achieve the purpose of high-throughput differentiation of different wells and different transcripts. In a single RT-PCR reaction, the barcode of the 5' end universal TSO-barcode primer is a single sequence as shown in any of SEQ ID NO.1-96. Each of the 96 wells uses the same barcode. During library construction, samples are distinguished only by the 3' end specific barcode. This is suitable for scenarios with ≤96 samples, such as small-scale drug screening. The barcode of the 5' end specific TSO-barcode primer is a sequence as shown in SEQ ID NO.1-96. Each of the 96 wells has a unique barcode. Combined with the 96 3' end barcodes, this forms 9216 unique tags. This is suitable for scenarios with >96 samples, such as large-scale plant breeding and high-throughput cell screening. Therefore, a single library can simultaneously contain up to 9216 seed libraries.
[0034] Table 1. Barcode sequence (applicable to both 5' and 3' ends)
[0035] 2. Run the RT-PCR reaction program (2-3 hours in total, achieving reverse transcription and amplification in one step). cDNA One-Strand Synthesis and Double-End Labeling: Incubation at 42℃ for 60 minutes (reverse transcriptase synthesizes cDNA from the 3' end of RNA by oligo(dT) binding to the poly(A) tail labeling. When extending to the 5' end of RNA, the TSO-barcode binds to the cDNA end through GC pairing, completing the 5' end labeling, and finally forming a product of 3' end-adaptor sequence C-barcode-UMI-full-length cDNA-barcode-TSO-5' end).
[0036] cDNA amplification double-stranded synthesis: 95℃, 3 min; (98℃, 20 sec; 65℃, 45 sec; 72℃, 3 min) × 4 cycles; (98℃, 20 sec; 67℃, 20 sec; 72℃, 3 min) × 11 cycles; 72℃, 5 min; hold at 4℃.
[0037] The final RT-PCR product is cDNA with paired-end labels: the 3' end carries a pore-specific UMI-barcode-adaptor sequence C sequence, and the 5' end carries a pore-specific or universal TSO-barcode sequence.
[0038] III. Transposition library construction reaction yields DNA fragments containing partial transposition adapters, paired-end tags, and the target cDNA fragment. 1. Purification 10 μl of cDNA product was aspirated from each well of a 96-well plate and mixed thoroughly to obtain 960 μl of cDNA. 384 μl of 0.4× purification beads were added to the cDNA mixture for purification, and the purified product was finally eluted with 50 μl of RNase-free water. 200 ng of the purified cDNA mixture was then used for subsequent transposase reactions.
[0039] 2. Preparation of transposase reaction system The transposase reaction system is 30 μl, comprising 200 ng of cDNA mixture, 6 μl of 5× transposase reaction buffer, 6 μl of transposase complex, and RNase-free water to a final volume of 30 μl. The transposase complex includes transposase complex A, composed of transposase and transposon DNA conjugated with adapter sequence A. After digestion with transposase A, DNA fragments containing adapter sequence A are obtained. A is amplified using index primers specifically matching the TSO-barcode conjugate sequence of adapter sequence A and cDNA 5' end, allowing selective enrichment and construction of a cDNA 5' end library. Transposase complex B, composed of transposase and transposon DNA conjugated with adapter sequence B, is also included. After digestion with transposase B, DNA fragments containing adapter sequence B are obtained. B is amplified using index primers specifically matching the Oligo(dT)-barcode-adaptor sequence C conjugate sequence of adapter sequence B and cDNA 3' end, allowing selective enrichment and construction of a cDNA library. 3' end library; transposase complex A+B (1:1 mixture): Both act together on the full-length cDNA, randomly cutting it to obtain DNA fragments containing adapter sequences A and B. By simultaneously using the above index primer pair A and index primer pair B for amplification, a mixed library containing both 5' and 3' end fragments can be obtained, achieving double-end coverage of the cDNA.
[0040] 3. Run the transposable reaction program Transposation procedure: Incubate at 55°C for 10 minutes. After the reaction, place the tube on ice and immediately add 30 μl of Stop Buffer. Mix well by pipetting and incubate at 55°C for 5 minutes to inactivate the transposase. Add 90 μl of purification magnetic beads (1.5×) to the final product for purification, and elute with 22 μl of RNase-free water to obtain the purified transposon product, which is a DNA fragment containing part of the transposon adapter, paired-end well tag, and target cDNA fragment.
[0041] IV. Indexing PCR Amplification and Sequencing Index The DNA fragment containing a partial transposon adapter, a paired-end well tag, and the target cDNA fragment was amplified using index primers to obtain a mixed library containing well tags, sequencing index, and the target cDNA fragment, which was then sent for sequencing.
[0042] 1. Prepare the indexed PCR amplification reaction system The index PCR reaction system consisted of 50 μl of 20 μl transposon product, 25 μl of 2× high-fidelity PCR Mix, 2.5 μl of index primer pair A, and 2.5 μl of index primer pair B, which were then mixed thoroughly.
[0043] Among them, index primer pair A includes P5-index(i5)_primer binding region T and P7-index(i7)_primer binding region A, primer binding region T is paired with the TSO sequence, and primer binding region A is paired with the transposon adapter sequence A; index primer pair B includes P5-index(i5)_primer binding region B and P7-index(i7)_primer binding region C, primer binding region B is paired with the transposon adapter sequence B, and primer binding region C is paired with the adapter sequence C.
[0044] 2. Run the indexed PCR amplification reaction program. The index PCR amplification reaction program was as follows: 72℃, 3 min; 98℃, 3 min; (98℃, 10 sec; 62℃, 30 sec) × 7 cycles; 72℃, 3 min; 4℃ hold.
[0045] The final qualified RNA-seq sequencing library contains well tags (barcodes) and UMI sequences, sequencing indexes (i5 / i7), and target cDNA fragments, which can be directly sent to the Illumina platform for sequencing.
[0046] V. High-throughput sequencing and data analysis 1. Sequencing: Sequencing the RNA-seq sequencing library (sequencing depth selected according to requirements) and outputting the data.
[0047] 2. Data Analysis: The data from the unloaded RNA sample were split and analyzed according to the labels of each well to obtain the RNA source information for each well, and further distinguished each transcript individually using UMI.
[0048] Data splitting: The data of the 96 wells is split into sub-library data according to the 5' end / 3' end barcode sequence, so as to distinguish between a minimum of 96 and a maximum of 9216 processed samples; Deduplication and quantification: PCR repetitive sequences are removed using UMI sequences to accurately quantify gene expression levels in each sample (FPKM>1 is considered a valid gene). Information analysis: Dual-end library preparation samples can be analyzed simultaneously for information at the 5' and 3' ends of the original RNA, while single-end library preparation samples only focus on information at the 3' end.
[0049] Compared to existing conventional RNA-seq and common 3' RNA-seq methods, the library construction method of this invention significantly increases the number of sample types (cells or RNA) that can be distinguished at one time because each transcript has a specific tag sequence at both ends. Furthermore, the application of a transposase complex allows for direct amplification of fragmented products, simplifying the library construction process and improving efficiency. This invention is suitable for small numbers of cells and trace amounts of RNA, effectively reverse transcribing mRNA information, and can simultaneously and rapidly process valuable samples and samples with varying starting quantities. The following example of 293T cells demonstrates high-throughput paired-end RNA-seq library construction and sequencing verification.
[0050] Example 1: High-throughput paired-terminal RNA-seq library construction and sequencing verification based on 293T cells I. Preparation and Lysis of 293T Cells (Obtaining Total RNA) Take an RNase-free 96-well PCR plate and add 10 μL of 293T cell suspension (containing 2000 cells in total) to each well, and seed 96 wells (well positions A1-H12).
[0051] Add 10 μL of Triton lysis buffer containing RNase inhibitor to each well, gently aspirate three times (avoiding air bubbles), and incubate at room temperature for 10 minutes (to ensure complete cell lysis and release of total RNA) to obtain the lysis products.
[0052] II. One-step RT-PCR for obtaining paired-end labeled cDNA 1. Prepare RT-PCR reaction system (20 μL per well, batch preparation for 96-well plates). Add 14 μL RT-PCR buffer, 2 μL 5 μM 3' end well-specific oligo(dT)-UMI-barcode-adaptor sequence C primer (one unique primer per well, added sequentially according to well position; the UMI sequence is a random sequence, the barcode sequence is shown in Table 1, and the adapter sequence C sequence is 5'-CAGACGAGCATCAG-3', SEQ ID NO. 97), 0.5 μL 10 μM 5' end universal TSO-barcode primer (sequence 5'-GCAGCATACGA-barcode-rGrGrG-3', barcode sequence is shown in SEQ ID NO. 1-96), 1.5 μL enzyme mixture (containing reverse transcriptase and high-fidelity polymerase, with final reaction concentrations of 10 μM and 8 μM respectively), and 2 μL of lysis product from the corresponding well. Mix well and start the RT-PCR reaction system. 2. RT-PCR reaction procedure (2.5 hours in total) cDNA One-Strand Synthesis and Double-End Labeling: Incubation at 42℃ for 60 minutes (reverse transcriptase synthesizes cDNA from the 3' end of RNA by oligo(dT) binding to the poly(A) tail labeling. When extending to the 5' end of RNA, the TSO-barcode binds to the cDNA end through GC pairing, completing the 5' end labeling, and finally forming a product of 3' end-adaptor sequence C-barcode-UMI-full-length cDNA-barcode-TSO-5' end).
[0053] Next, cDNA amplification and double-strand synthesis were performed: 95℃, 3 min; (98℃, 20 sec; 65℃, 45 sec; 72℃, 3 min) × 4 cycles; (98℃, 20 sec; 67℃, 20 sec; 72℃, 3 min) × 11 cycles; 72℃, 5 min; 4℃ hold.
[0054] The final RT-PCR product is cDNA, which carries a well-specific UMI-barcode adapter sequence C at the 3' end and a well-universal TSO-barcode sequence at the 5' end.
[0055] 3. Product Validation Use Qsep100 TM The fully automated nucleic acid and protein analysis system detects RT-PCR products, and the peak diagram is as follows: Figure 2 As shown, the cDNA fragments exhibited diffuse peaks with lengths concentrated between 500-5000 bp, without obvious short impurities or primer dimers, indicating that the one-step RT-PCR method successfully synthesized full-length cDNA with paired-end labels. The product showed no degradation, no non-specific amplification, and good uniformity, meeting the requirements for subsequent library construction.
[0056] III. Transposition and silo-building reaction 1. cDNA mixing and purification 1) cDNA mixing: Use an RNase-free pipette tip to pipette 10 μL of RT-PCR product from each well of a 96-well plate and mix well to obtain 960 μL of cDNA mixture.
[0057] 2) Magnetic bead purification: Add 384 μl of purification magnetic beads (0.4×) to the mixed cDNA for purification, and finally wash the purified product with 50 μl of RNase-free water. Take 200 ng of the purified cDNA mixture for subsequent transposase reaction.
[0058] 2. Transposase reaction 1) Preparation of the transposable reaction system: The transposition reaction system consisted of 30 μl, including 200 ng of cDNA mixture, 6 μl of 5× transposition reaction buffer, 6 μl of transposase complex, and finally RNase-free water to make up to 30 μl. After mixing, the transposition reaction system was obtained. The transposase complex A consists of a transposase and transposon DNA coupled with adapter sequence A (adaptor sequence A is composed of two complementary single strands, with sequences of 5'P-CTGTCTCTTATACACATCT-3', SEQ ID NO. 98 and 5'-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG-3', SEQ ID NO. 99, respectively), and the transposase complex B consists of a transposase and transposon DNA coupled with adapter sequence B (adaptor sequence B is composed of two complementary single strands, with sequences of 5'P-CTGTCTCTTATACACATCT-3', SEQ ID NO. 100 and 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3', SEQ ID NO. 101, respectively).
[0059] 2) Run the transposable reaction program: Transposation procedure: Incubate at 55°C for 10 minutes. After the reaction, place the tube on ice and immediately add 30 μl of Stop Buffer. Mix well by pipetting and incubate at 55°C for 5 minutes to inactivate the transposase. Add 90 μl of purification magnetic beads (1.5×) to the final product for purification, and elute with 22 μl of RNase-free water to obtain the purified transposon product, which is a DNA fragment containing part of the transposon adapter, paired-end well tag, and target cDNA fragment.
[0060] IV. Indexing PCR Amplification and Sequencing Index 1. Prepare the indexed PCR reaction system The index PCR reaction system is 50 μl, including 20 μl transposon product, 25 μl 2× high-fidelity PCR Mix, 2.5 μl index primer pair A (10 μM each), and 2.5 μl index primer pair B (10 μM each). Mix well before use.
[0061] Among them, the index primer pair A includes P5-index(i5)_primer binding region T (sequence 5'-AATGATACGGCGACCACCGAGATCTACAC[i5]GCAGCATACGA-3') and P7-index(i7)_primer binding region A (sequence 5'-CAAGCAGAAGACGGCATACGAGAT[i7]GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG-3'); the index primer pair B includes P5-index(i5)_primer binding region B (sequence 5'-AATGATACGGCGACCACCGAGATCTACAC[i5]TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3') and P7-index(i7)_primer binding region C (sequence 5'-CAAGCAGAAGACGGCATACGAGAT[i7]CAGACGAGCATCAG-3').
[0062] 2. Run the index PCR reaction program The index PCR amplification reaction program was as follows: 72℃, 3 min; 98℃, 3 min; (98℃, 10 sec; 62℃, 30 sec) × 7 cycles; 72℃, 3 min; 4℃ hold.
[0063] The final result is a qualified RNA-seq sequencing library containing PCR products with well tags (barcodes) and UMI sequences, sequencing indexes (i5 / i7), and target cDNA fragments, which can be directly sent to the Illumina platform for sequencing.
[0064] 3. PCR product verification Take 2 μl of PCR product and use Qsep100 TM The fully automated nucleic acid and protein analysis system detected the peaks as shown in the figure. Figure 3 As shown, the DNA fragment exhibits diffuse peaks, with lengths concentrated between 200-1000 bp. The peak shape is uniform with no obvious impurities, indicating successful transposase cleavage, satisfactory fragmentation, and that the PCR product meets the requirements for a sequencing library. This library contains well tags (barcodes), sequencing indexes (i5 / i7), and the target cDNA fragment, and can be directly sent to the Illumina platform for sequencing.
[0065] V. Sequencing Library Quality Testing RNA-seq sequencing libraries were sequenced on the Illumina platform, with sequencing depths of 100G, 75G, 50G, and 25G selected.
[0066] Data splitting: The data of the 96 wells is split into sub-library data according to the 5' end / 3' end barcode sequence to distinguish the 96 processed samples; Deduplication and quantification: PCR repetitive sequences were removed using UMI sequencing to accurately quantify gene expression levels in each sample (FPKM>1 was considered a valid gene). The number of valid genes in each well was counted at four sequencing depths, and the results are shown in Table 2: At a sequencing depth of 100G, the number of valid genes in each well was 12222-14178, with an average of 13380; at a sequencing depth of 75G, the number of valid genes in each well was 11880-13945, with an average of 13120; at a sequencing depth of 50G, the number of valid genes in each well was 11266-13888, with an average of 12700; and at a sequencing depth of 25G, the number of valid genes in each well was 9877-13440, with an average of 11900. The results show that the RNA-seq sequencing library obtained using the high-throughput paired-end RNA-seq library preparation method of this invention can still detect approximately 11,900 effective genes even with extremely low starting amounts (200 cells / well) and low sequencing depths (25G), demonstrating extremely high sensitivity and stability. At any given sequencing depth (100G, 75G, 50G, 25G), the number of effective genes in each well corresponding to different bars is uniformly distributed, and the gene capture number is not affected by the base composition of the barcode itself, exhibiting excellent reproducibility. Furthermore, reducing the sequencing depth from 100G to 50G results in relatively small loss of effective genes (an average decrease of approximately 5.1%), significantly reducing sequencing costs.
[0067] Table 24: Statistical table of effective gene counts at each well for each sequencing depth.
[0068] In summary, the high-throughput paired-terminal RNA-seq library construction method described in this invention can efficiently screen the expression information of transcripts in each well and is suitable for low initial sample input. It not only speeds up the testing process but also significantly reduces the cost for those who need rapid library construction. It is suitable for high-throughput rapid RNA library construction needs under various broad screening applications, such as multi-drug screening and plant breeding, and is of great help for early and rapid acquisition of target genes.
[0069] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. For those skilled in the art, other variations or modifications can be made based on the above description. It is impossible to exhaustively list all the implementation methods here. All obvious variations or modifications derived from the technical solutions of the present invention are still within the protection scope of the present invention.
Claims
1. A high-throughput rapid library construction method using paired-terminal RNA-seq, characterized in that, The rapid library preparation method using paired-terminal RNA-seq includes the following steps: S1. Obtaining total RNA by cell lysis: Treated eukaryotic cells were lysed with Triton lysis buffer containing RNase inhibitor, and the lysis products were collected; S2. Obtaining paired-end labeled cDNA by one-step RT-PCR: Using the lysis product as a template, a 3' end pore-specific oligo(dT)-UMI-barcode-adaptor sequence C primer and a 5' end pore-specific / universal TSO-barcode primer were used to simultaneously complete reverse transcription and amplification through a one-step RT-PCR reaction to obtain paired-end labeled cDNA with the 3' end carrying the UMI-barcode-adaptor sequence C and the 5' end carrying the TSO-barcode. S3. Transposition reaction to obtain DNA fragments with partial transposable adapters, paired-end pore tags and target cDNA fragments: After purifying the paired-end labeled cDNA mixture, a purified cDNA mixture product is obtained. A transposase complex is added to carry out a transposition reaction to obtain DNA fragments with partial transposable adapters, paired-end pore tags and target cDNA fragments. S4. Indexing PCR amplification to introduce sequencing index: DNA fragments are amplified by PCR using index primers to introduce sequencing index and obtain RNA-seq sequencing libraries.
2. The rapid library construction method for paired-terminal RNA-seq according to claim 1, characterized in that, The 3' end well-specific oligo(dT)-UMI-barcode-adaptor sequence C primer contains barcode, UMI, oligo(dT) sequences and adapter sequence C. Each well corresponds to a unique barcode sequence, which is used to initiate the synthesis of the first strand of cDNA and introduce the 3' end carrying the UMI-barcode-adaptor sequence C sequence. Preferably, the 5' end well-specific or universal TSO-barcode primer contains a barcode and a TSO sequence. The barcode is one sequence adapted to 96 wells or 96 sequences adapted to 9216 samples. The TSO-barcode sequence is introduced at the 5' end through a template conversion reaction.
3. The rapid library construction method for paired-terminal RNA-seq according to claim 2, characterized in that, The sequence of the barcode is as shown in any of SEQ ID NO.1~96.
4. The rapid library construction method for paired-terminal RNA-seq according to claim 1, characterized in that, The one-step RT-PCR reaction system includes: lysis products, 3' end pore-specific oligo(dT)-UMI-barcode-adaptor sequence C primer, 5' end pore-specific / universal TSO-barcode primer, RT-PCR buffer, and enzyme mixture; wherein the enzyme mixture contains reverse transcriptase and polymerase.
5. The rapid library construction method for paired-terminal RNA-seq according to claim 1, characterized in that, The one-step RT-PCR reaction program is as follows: 42℃, 60 min; 95℃, 3 min; (98℃, 20 sec; 65℃, 45 sec; 72℃, 3 min) × 4 cycles; (98℃, 20 sec; 67℃, 20 sec; 72℃, 3 min) × 11 cycles; 72℃, 5 min; hold at 4℃.
6. The rapid library construction method for paired-terminal RNA-seq according to claim 1, characterized in that, The transposase complex includes transposase complex A and / or transposase complex B; transposase complex A consists of a transposase and a transposon coupled to the header sequence A, with the 5' end for library construction; transposase complex B consists of a transposase and a transposon coupled to the header sequence B, with the 3' end for library construction; transposase complex A and transposase complex B are used in a 1:1 ratio to achieve dual-end library construction.
7. The rapid library construction method for paired-terminal RNA-seq according to claim 1, characterized in that, The amount of the purified cDNA mixture is 50-200 ng; the transposition reaction system includes the cDNA mixture, transposition reaction buffer, transposase complex and RNase-free water.
8. The rapid library construction method for paired-terminal RNA-seq according to claim 1, characterized in that, The transposable reaction procedure is to incubate at 55°C for 10 min.
9. The rapid library construction method for paired-terminal RNA-seq according to claim 1, characterized in that, The index primer pair includes index primer pair A and / or index primer pair B; wherein, index primer pair A includes a P5-index(i5) primer binding region T and a P7-index(i7) primer binding region A, the primer binding region T in the P5-index(i5) primer binding region T is complementary to the TSO sequence, and the primer binding region A in the P7-index(i7) primer binding region A is complementary to the transposon adapter sequence A. Sequencing indexes i5 and i7 are introduced through index PCR reaction to realize cDNA sequencing. 5' end library construction; the index primer pair B includes a P5-index(i5) primer binding region B and a P7-index(i7) primer binding region C. The primer binding region B in the P5-index(i5) primer binding region B is complementary to the transposon adapter sequence B, and the primer binding region C in the P7-index(i7) primer binding region C is complementary to the adapter sequence C. The sequencing indexes i5 and i7 are introduced through an index PCR reaction to achieve cDNA 3' end library construction; if index primer pair A and index primer pair B are used simultaneously, the sequencing indexes i5 and i7 can be introduced through a single index PCR reaction to achieve cDNA paired-end library construction.
10. The rapid library construction method for paired-terminal RNA-seq according to claim 1, characterized in that, The indexed PCR reaction system includes a DNA fragment, a PCR mix, index primer pair A, and / or index primer pair B.