Construction method of trace RNA sequencing library and reagent used in construction method
By using the transposase Tn5 complex for fragmentation and adapter ligation in a single tube, the problems of low efficiency and high cost in constructing sequencing libraries for trace RNA were solved, and efficient and low-cost construction of trace RNA sequencing libraries was achieved.
Patent Information
- Application Number
- CN202511178624.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-08-22
AI Technical Summary
Existing technologies make it difficult to effectively construct sequencing libraries for trace amounts of RNA, especially in plasma or serum samples, where the RNA amount is insufficient and the efficiency of adaptor ligation is low, resulting in a cumbersome, time-consuming and costly library construction process.
A new RNA sequencing library construction method is used to obtain DNA-RNA hybrid chains by reverse transcription of RNA in biological samples, followed by fragmentation and adapter ligation using the transposase Tn5 complex, and finally PCR amplification and purification to obtain high-quality RNA sequencing libraries.
This method can complete the entire process from RNA to sequencing library in a single tube, reducing intermediate purification steps, reducing RNA loss and library construction costs, shortening library construction time, and effectively reducing the proportion of rRNA in the library.
Smart Images

Figure CN120665989A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of nucleic acid determination methods, and particularly relates to a method for constructing a micro RNA sequencing library and reagents used therein. Background Art
[0002] Cell-free RNA (cfRNA) in body fluids, such as plasma or serum, possesses greater tissue specificity and diversity than existing DNA and protein biomarkers. cfRNA originates from a variety of cell death pathways, including apoptosis, necrosis, or pyroptosis, as well as release from extracellular vesicles. These molecules can be emitted from a variety of human cell types in remote organ systems or non-human microorganisms. As they circulate within biological systems, cfRNA exhibits remarkable resilience against degradation due to its association with protective proteins, lipids, or extracellular vesicles. Therefore, the use of trace amounts of cfRNA in plasma or serum to diagnose disease has become a research area of considerable interest in diagnostic medicine. Currently, RNA in plasma or serum can be detected, and numerous studies are hoping to identify new diagnostic biomarkers.
[0003] However, a major challenge in practical applications is how to construct high-quality, high-throughput sequencing libraries from cfRNA. Existing methods require sufficient amounts of RNA (at least >10 ng), followed by adaptor ligation to construct sequencing libraries. These methods suffer from the following issues: 1) limited access to plasma or serum samples, making it impossible to extract more than 10 ng of RNA; 2) low efficiency due to the multiple necessary purification steps during library construction, resulting in significant losses; 3) cumbersome and time-consuming library construction, typically requiring more than three days, resulting in low timeliness; 4) high overall library construction costs (>1200 RMB per sample); and 5) a high proportion of rRNA reads in the library (>75%), requiring increased sequencing depth to obtain sufficient information, further increasing costs. These issues have hindered research progress and clinical application of cell-free extracellular RNA in plasma or serum. Summary of the Invention
[0004] The technical problem to be solved by the present invention is how to construct a sequencing library for trace RNA and / or how to construct a sequencing library for trace RNA of low abundance rRNA and / or how to construct a high-throughput sequencing library for trace extracellular free RNA and / or how to construct a sequencing library for extracellular free RNA in plasma or serum and / or how to construct a sequencing library for extracellular free RNA in body fluids and / or how to construct a sequencing library for low rRNA abundance of extracellular free RNA in body fluids.
[0005] In order to effectively solve the above technical problems, the present invention provides a method for constructing an RNA sequencing library, which may include the following steps: A1) Reverse transcription of RNA from biological samples to obtain DNA-RNA hybrid chains; A2) cleaving the DNA-RNA hybrid chain to obtain DNA-RNA hybrid chain fragments, and connecting a linker M1 and a linker M2 to both ends of the DNA-RNA hybrid chain fragments to obtain linker-added DNA-RNA hybrid chain fragments; A3) using the adapter-added DNA-RNA hybrid fragment as a template, performing (first round) PCR amplification using sequencing primers 1 and 2 specific to the sequencing platform to obtain a PCR product, and purifying the PCR product to obtain a full RNA sequencing library; The sequencing primer 1 may include a sequencing platform adapter 1 and an adapter N1. The sequencing primer 2 may include a sequencing platform adapter 2 and an adapter N2.
[0006] The linker N1 may be a single-stranded nucleotide fragment selected from the linker M1. The linker N2 may be a single-stranded nucleotide fragment selected from the linker M2.
[0007] The length of the linker N1 can be greater than or equal to 14 nt and less than or equal to the length of one chain of the linker M1. The length of the linker N2 can be greater than or equal to 14 nt and less than or equal to the length of one chain of the linker M2. In a specific embodiment of the present invention, the length of the linker N1 is 14 nt, and the length of the linker N2 is 21 nt.
[0008] The above method may further include the following step A4): A4) Genome editing is performed on the RNA sequencing library using a gRNA targeting an rRNA-derived PCR product and an RNA-guided nuclease to obtain an RNA sequencing library after target DNA shearing. The RNA sequencing library after target DNA shearing is subjected to (second-round) PCR amplification using the sequencing platform adapter 1 and the sequencing platform adapter 2 to obtain PCR product 2. The PCR product 2 is purified to obtain a purified RNA sequencing library.
[0009] In the above method, the adapter N1 in the sequencing primer 1 can be paired with the partially complementary sequence of the adapter M1 to simultaneously initiate the first round of PCR amplification reaction and introduce the sequencing platform adapter 1. The adapter N2 in the sequencing primer 2 can be paired with the partially complementary sequence of the adapter M2 to simultaneously initiate the first round of PCR amplification reaction and introduce the sequencing platform adapter 2.
[0010] In the above method, the sequencing primer 1 and / or the sequencing primer 2 may further include a tag sequence.
[0011] The sequences of the linker M1 and the linker M2 may be the same.
[0012] In the above method, the mass of the RNA in A1) may be greater than or equal to 20 pg.
[0013] In some specific embodiments of the present invention, the mass of the RNA in A1) is greater than or equal to 20 pg and less than or equal to 10 ng.
[0014] In some specific embodiments of the present invention, the amount greater than or equal to 20 pg and less than or equal to 10 ng is 20 pg-50 pg, 20 pg-100 pg, 50 pg-200 pg, 100 pg-500 pg, 200 pg-1 ng or 500 pg-10 ng.
[0015] In the above method, the biological sample may be animal body fluid.
[0016] In the above method, step A2) can be achieved by the transposase Tn5 complex.
[0017] In the above method, the sequencing platform adapter 1 sequence and the sequencing platform adapter 2 sequence can be specific adapter sequences used in sequencing platform library construction. In a specific embodiment of the present invention, the sequence of sequencing platform adapter 1 can be the P5 sequence of the Illumina sequencing platform. The sequence of sequencing platform adapter 2 can be the P7 sequence of the Illumina sequencing platform.
[0018] In a specific embodiment of the present invention, the sequence of the sequencing primer 1 is SEQ ID NO. 4 in the sequence listing, and the sequence of the sequencing primer 2 is SEQ ID NO. 5 in the sequence listing.
[0019] In the above method, the body fluid refers to liquid substances in the body, specifically including plasma, tissue fluid, cerebrospinal fluid, lymph fluid, intracellular fluid, etc. The volume of the body fluid can be greater than or equal to 200 μL.
[0020] In the above method, the mass of the RNA derived from the biological sample may be greater than or equal to 20 pg.
[0021] The animal mentioned above may be a mammal, and the mammal may be a human.
[0022] In the above method, the reverse transcriptase may be SuperScript IV.
[0023] In the above method, the polypeptide guided by the gRNA having nuclease activity can be a CRISPR-Cas protein. In a specific embodiment of the present invention, the CRISPR-Cas protein is Cas9.
[0024] The term "sgRNA (single-guide RNA) or gRNA" is a component of the CRISPR-Cas system that is responsible for guiding the Cas protein to recognize and cut the target nucleic acid molecule. In actual gene editing applications, sgRNA can be directly synthesized or obtained through plasmid expression or in vitro transcription. In this field, "gRNA" is often used interchangeably with "sgRNA." In this article, "gRNA" and "sgRNA" are also used interchangeably. sgRNA generally refers to a single RNA structure formed by artificially modifying the crRNA / tracrRNA complex (gRNA) with a dual RNA structure, directly connecting the crRNA and tracrRNA (or through a linker). sgRNA is a short RNA that contains a recognition region and a framework region.
[0025] The term "recognition region," also referred to as a guide sequence, typically refers to an RNA sequence contained within a guide RNA (sgRNA or gRNA) that is identical or complementary to a target sequence or target site (referred to herein as the "guide sequence"). The guide sequence typically exhibits sufficient complementarity with the target sequence to hybridize therewith and guide the CRISPR / Cas complex to specific binding to the target sequence. While perfect complementarity between the guide and target sequences is preferred, some mismatches (e.g., 1-6 nucleotide mismatches) are tolerated as long as they still result in gene knockout. The degree of complementarity between a guide sequence and its corresponding target sequence is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. Methods for determining the complementarity of two nucleic acid sequences are within the capabilities of those skilled in the art.
[0026] The term "scaffold" generally refers to the structural or scaffolding RNA sequence within a guide RNA (sgRNA or gRNA) required for binding or interacting with an RNA-guided nuclease and / or other RNA molecules (e.g., tracrRNA). It can also be referred to as the backbone sequence of an sgRNA. The scaffold can be selected routinely by those skilled in the art. For example, the backbone sequence of the sgRNA corresponding to Cas9 can be selected, or variants constructed based on this backbone sequence that still retain the ability to bind to the corresponding Cas9.
[0027] In this application, "RNA-guided nuclease" refers to an RNA-guided DNA endonuclease associated with the CRISPR system. Non-limiting examples of RNA-guided nucleases include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, and Cpf1, homologs thereof, or modified forms thereof. In one embodiment, the RNA-guided nuclease is Cas9.
[0028] Furthermore, the Cas9 protein described in this application is not limited to a specific protein, as long as it can be used in conjunction with the gRNA of the present invention. Furthermore, the Cas9 protein described herein is selected from Streptococcus pyogenes Cas9 (spCas9, subtype II-A), spCas9HF (high fidelity), nickase Cas9 (nCas9), Staphylococcus aureus Cas9 (saCas9, subtype II-A), Neisseria meningitidis Cas9 (NmCas9, subtype II-C), Francisella novicida Cas9 (FnCas9, subtype II-B), Streptococcus thermophilus Cas9 (St1Cas9, St3Cas9), Campylobacter jejuni Cas9 (CjCas9) and Treponema sp. Cas9, as well as Cas9 orthologs of other organisms but not limited thereto. The Cas9 protein may also include high-fidelity Cas9 mutants (such as SpCas9-HF1, eSpCas9-1.1 and TrueCut™HiFiCas9 protein), etc.
[0029] In order to solve the above technical problems, the present invention also provides reagents for constructing an RNA sequencing library, which may include reverse transcriptase, the adapter M1 mentioned above, the adapter M2, the sequencing primer 1, the sequencing primer 2, the sequencing platform adapter 1 and the sequencing platform adapter 2.
[0030] The above reagents may also include gRNA and RNA-guided nuclease targeting rRNA-derived PCR products.
[0031] The reagents described above may further include a transposase complex and a DNA polymerase.
[0032] In the above reagents, the transposase complex may be a transposase Tn5 complex.
[0033] The above reagents may also include SuperScript IV reverse transcriptase.
[0034] The sequence of the linker M1 can be SEQ ID NO. 1 in the sequence listing. The sequence of the linker M2 can be SEQ ID NO. 2 in the sequence listing.
[0035] In a specific embodiment of the present invention, the sequencing platform adapter 1 may be P5 of the Illumina sequencing platform, and the sequencing platform adapter 2 may be P7 of the Illumina sequencing platform.
[0036] In a specific embodiment of the present invention, the sequence of the sequencing primer 1 is SEQ ID NO. 4 in the sequence listing, and the sequence of the sequencing primer 2 is SEQ ID NO. 5 in the sequence listing.
[0037] In a specific embodiment of the present invention, the gRNA targeting the rRNA-derived PCR product in the present invention is SEQ ID NO.16-SEQ ID NO.73 in the sequence listing.
[0038] In order to solve the above technical problems, the present invention also provides the use of the reagents described above in constructing an RNA sequencing library.
[0039] In order to solve the above technical problems, the present invention also provides a method for constructing an RNA sequencing library, which may include steps A1) to A3) of the method described above.
[0040] By leveraging the properties of Tn5, this invention comprehensively addresses the challenges encountered in constructing high-throughput sequencing libraries for extremely low-level cfRNA. This allows the entire process from cfRNA to first-round PCR amplification (cfRNA to whole-RNA sequencing library) to be completed in a single tube, avoiding the sample loss caused by purification during other library construction methods while significantly reducing library construction time and costs. This invention also includes the removal of rRNA-derived DNA from the library, effectively reducing the proportion of rRNA-derived reads in the library and increasing the probability of capturing other RNAs. Using this technology, libraries can be constructed in a short period of time from RNA levels as low as 10 ng, or even pg. High-throughput sequencing can then generate high-quality data, which can then be analyzed to identify disease-related RNA biomarkers.
[0041] Beneficial effects of the present invention: The RNA sequencing library construction process provided by this invention reduces RNA purification times and RNA loss; it can also construct libraries from trace amounts of RNA (e.g., 20 pg of RNA or cfRNA). This method can effectively reduce the proportion of rRNA in RNA sequencing libraries from biological samples (to below 32%). The library construction process of this method is time-efficient and low-cost, making it suitable for large-scale deployment. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 Schematic diagram of the experiment of the method of the present invention.
[0043] Figure 2 It is a schematic diagram of the process of the present invention.
[0044] Figure 3 Schematic diagram of the transposase Tn5 complex in an embodiment of the present invention.
[0045] Figure 4 These are the analysis results of the ribosomal RNA proportion and the exon proportion in the alignment region in the RNA sequencing library constructed using the method of the present invention for 10 plasma samples in Example 2.
[0046] Figure 5 These are the analysis results of the ribosomal RNA ratio and the exon ratio in the alignment region in the RNA sequencing library constructed using the method of the present invention using 6 samples of total RNA (trace amount) from human brain cells in Example 3.
[0047] Figure 6 Relevant data were constructed for RNA sequencing libraries derived from a total of 16 biological samples in Examples 2 and 3.
[0048] Figure 7This is the sequence information of 29 Guide-rRNAs (Grs-1 to Grs-29) contained in the Guide-rRNA Mix.
[0049] Figure 8 This is the sequence information of 29 Guide-rRNAs (Grs-30 to Grs-58) contained in the Guide-rRNA Mix. DETAILED DESCRIPTION
[0050] The present invention will be further described in detail below in conjunction with specific embodiments. The examples provided are only for illustrating the present invention and are not intended to limit the scope of the present invention. The examples provided below can serve as a guide for further improvements by those skilled in the art and are not intended to limit the present invention in any way.
[0051] Unless otherwise specified, the experimental methods in the following examples are conventional methods and were performed according to the techniques or conditions described in the literature in the field or according to the product instructions. The materials and reagents used in the following examples, unless otherwise specified, were all commercially available.
[0052] The nucleic acid sequences used in the transposase Tn5 complex, the whole RNA sequencing library, and the rRNA low-abundance sequencing library used in the following examples are as follows, wherein: N represents A, T, C, or G.
[0053] Tn5-Ad1 (5'-3'):TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG (SEQ ID NO.1); Tn5-Ad2 (5'-3'): GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG (SEQ ID NO. 2); Tn5-ME (5'-3'): pCTGTCTCTTATACACATCT (SEQ ID NO. 3, 5'-terminal nucleotide C is phosphorylated); N5 (5'-3'): AATGATACGGCGACCACCGAGATCTACACNNNNNNNNTCGTCGGCAGCGTC (SEQ ID NO. 4); N7 (5'-3'): CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTCTCGTGGGCTCGGAGATGT (SEQ ID NO.5); P5 (5'-3'):AATGATACGGCGACCACCGA (SEQ ID NO. 14); P7 (5'-3'): CAAGCAGAAGACGGCATACGA (SEQ ID NO. 15).
[0054] Example 1. Extracting cfRNA from plasma samples and constructing a whole RNA sequencing library.
[0055] This example uses plasma as the sample to be tested and describes in detail the specific process of extracting cfRNA from plasma samples and constructing a whole RNA sequencing library.
[0056] 1. Extract and purify cfRNA from plasma.
[0057] 1.1 Extraction of extracellular free nucleic acids.
[0058] Extracellular RNA (cfRNA) was extracted and purified from plasma samples using a serum / plasma cell-free RNA extraction kit (Jianshi Biotechnology, TR159). The specific steps involved were as follows: 1) Pipette 1 mL of plasma sample into a 15 mL tube and add 1 mL of the free RNA digestion solution (mainly the enzymatic reaction buffer) provided in the kit and mix thoroughly.
[0059] 2) Add 50 μL of proteinase K solution (50 units, prepared by mixing the proteinase K and proteinase K storage solution in the kit), mix well, and incubate at 37°C for 2 hours to digest the protein in the sample.
[0060] 3) Add 2 mL of free RNA binding solution (mainly nucleic acid binding buffer) and mix thoroughly.
[0061] 4) Add 6 mL of 100% (volume fraction) isopropanol (to promote nucleic acid precipitation) and mix thoroughly to obtain a crude free nucleic acid extract.
[0062] 5) Place the 25 mL funnel onto filter column 1 (column 3 Y) in the kit to assemble the filtration apparatus.
[0063] 6) Transfer the crude free nucleic acid extract obtained in step 4) to the filtration apparatus tube in step 5). Turn on the vacuum pump to filter the crude free nucleic acid extract so that it completely flows through the matrix of filter column 1. Turn off the vacuum pump and discard the 25 mL funnel.
[0064] 7) Add 600 μL of RNA Wash Buffer 1 to the matrix of Filter Column 1, turn on the vacuum pump (400 mmHg) to allow the liquid to flow through, and then turn off the vacuum pump.
[0065] 8) Transfer filter column 1 to a collection tube and centrifuge at 12,000 g for 2 minutes to remove residual liquid.
[0066] 9) Add 700 µL RNA Wash Buffer 2 (75% ethanol, to remove salt ions) and centrifuge at 12,000 g for 30 seconds. Discard the waste solution.
[0067] 10) Add 400 μl RNA Wash Buffer 2 (75% ethanol, to remove salt ions) and centrifuge at 12,000 g for 2 minutes to remove all liquid.
[0068] 11) Transfer filter column 1 to a clean 1.5 mL centrifuge tube. Add 80 μL of nuclease-free water from the kit to the filter column 1, incubate for 2 minutes, and centrifuge at 12,000 g for 30 seconds to elute the purified free nucleic acid.
[0069] 1.2 Remove DNA from extracellular free nucleic acids to obtain cfRNA.
[0070] Use the RNA purification and concentration kit (Jianshi Biotechnology TR113) to continue purifying and concentrating the purified free nucleic acid obtained in step 1.1 to remove the DNA and obtain cfRNA. The specific steps are as follows: 12) Add the purified free nucleic acid obtained in step 1.1 to the DNA digestion solution (10 µL), DNaseI (10 µL), and RNase Inhibitor (Vazyme, R301) in the kit.
[0071] 13) Incubate the entire reaction system at room temperature for 15 minutes.
[0072] 14) Add 200µL of RNA binding buffer from the kit to the incubated system and mix thoroughly.
[0073] 15) Add 300 mL of 95-100% (volume fraction) ethanol, mix, and transfer to column 1 of the kit containing filter column 2. Place in a clean collection tube and centrifuge at 12,000 g for 30 seconds to allow the mixture to completely pass through the matrix of filter column 2. Discard the waste liquid.
[0074] 16) Add 400 µL of RNA pre-wash solution from the kit, centrifuge at 12,000 g for 30 seconds, and discard the waste solution.
[0075] 17) Add 700 µL of RNA Wash Buffer from the kit, centrifuge at 12,000 g for 30 seconds, and discard the waste solution.
[0076] 18) Add 400 µL of RNA wash buffer from the kit and centrifuge at 12,000 g for 1 minute to remove all liquid.
[0077] 19) Transfer filter column 2 to a clean 1.5 mL centrifuge tube. Add 10 µL of nuclease-free water from the kit (pre-heated in a 65-70°C water bath) to filter column 2. Incubate for 2 minutes. Centrifuge at 12,000 g for 30 seconds to elute the purified cfRNA.
[0078] 2. Construction of whole RNA sequencing library.
[0079] 2.1 cfRNA reverse transcription to obtain hybrid chains.
[0080] 1) First, add the purified cfRNA obtained in step 1.2 to Reagent 1 and incubate at 65°C for 10 minutes and 72°C for 3 minutes. Reagent 1 contains: 1 µL reverse transcription primer (NEB, S1330), 1.25 µL dNTPs (NEB, N0447), and 0.15 µL RNase Inhibitor (Vazyme, R301).
[0081] 2) The reaction mixture from step 1) was then mixed with Reagent 2 and incubated to complete reverse transcription, generating a DNA-RNA hybrid. Reaction conditions were 23°C for 10 minutes, 50°C for 50 minutes, and 80°C for 10 minutes. Reagent 2 consisted of: 5 µL of 5× SSIV buffer (Reverse Transcriptase Buffer, Thermo, 18090050), 1.25 µL of DTT (Thermo, 18090050), 0.6 µL of RNase Inhibitor (Vazyme, R301), 5 µL of Betaine (Merck, B0300), 0.5 µL of SuperScript IV Reverse Transcriptase (Thermo, 18090050), and 0.2 µL of nuclease-free water.
[0082] 2.2 Transposase Tn5 cuts the hybridized strand and adds adapters.
[0083] 3) The present invention uses the transposase Tn5 complex to fragment the DNA-RNA hybrid and perform adapter ligation, resulting in gapped DNA-RNA hybrid fragments to which Tn5-Ad1 and Tn5-Ad2 adapters are added. The added adapters (Tn5-Ad1 and Tn5-Ad2) serve as primer binding sites for PCR reactions.
[0084] Specifically, the DNA-RNA hybrid reaction system obtained in step 2.1 was mixed with Reagent 3 and incubated at 55°C for 20 min. Reagent 3 consisted of: 4.5 µL of 8× TD buffer (prepared with 8 µL of 1M Tris-HCl (Beyotime, ST780), 4 µL of 1M MgCl₂ (Aladdin, M299562), 80 µL of N,N-dimethylformamide DMF (Merck, D4551), and 80 µL of nuclease-free water), 5.25 µL of 50% PEG₁₀ (Beyotime, R0056), 0.8 µL of ATP (NEB, P0756), 0.4 µL of RNase Inhibitor (Vazyme, R301), and 0.05 µL of transposase Tn5 complex.
[0085] The transposase Tn5 complex was synthesized using primers containing the naked Tn5 enzyme (TransGen, LT201), the Tn5-Ad1 linker (SEQ ID NO. 1 in the sequence listing), the Tn5-Ad2 linker (SEQ ID NO. 2 in the sequence listing), and the Tn5-ME core sequence (SEQ ID NO. 3 in the sequence listing) in the following steps: (3-1) Prepare a mixed system I of Tn5-Ad1, Tn5-ME, and Annealing Buffer, then treat the mixed system I at 95℃ for 2 minutes, cool it down from 95℃ to 22℃, and react at a speed of 0.1℃ / s; then treat it at 22℃ for 5 minutes; (3-2) Prepare a mixed system II of Tn5-Ad2, Tn5-ME, and Annealing Buffer, then treat the mixed system II at 95℃ for 2 minutes, cool it down from 95℃ to 22℃, and react at a speed of 0.1℃ / s; then treat it at 22℃ for 5 minutes; (3-3) Treat the mixed system I, mixed system II, Tn5 naked enzyme, and Tn5 Storage Buffer at 35℃ for 2 hours to obtain the transposase Tn5 complex ( Figure 3 ).
[0086] 2.3 Filling the gap in the DNA-RNA hybrid chain.
[0087] 4) Mix the reaction mixture containing the gapped DNA-RNA hybrid obtained in step 2.2 with Reagent 4 and incubate to obtain intact DNA-RNA hybrid fragments with adapters after filling the gap. Reaction conditions are: 72°C for 15 minutes; 80°C for 5 minutes. Reagent 4 contains: 10.5 µL DNA Polymerase Buffer, 1 µL dNTPs (NEB, N0447), and 0.5 µL DNA Polymerase (NEB, M0491).
[0088] 2.4 PCR amplification and addition of sequencing adapters.
[0089] 5) Mix the reaction mixture containing the intact DNA-RNA hybrid fragment obtained in step 2.3 with a mixture containing N5 primers and N7 primers (containing 2.5 µL of 10 µM primer N5 and 2.5 µL of 10 µM primer N7). Then perform the first round of PCR reaction to obtain PCR products with the addition of N5 and N7 sequencing primer sequences.
[0090] The PCR reaction program was as follows: initial denaturation at 98°C for 30 s; 10 cycles of denaturation at 98°C for 10 s, annealing at 60°C for 20 s, and extension at 72°C for 30 s; final extension at 72°C for 2 min; and hold at 4°C.
[0091] The N5 sequencing primer (SEQ ID NO. 4 in the sequence listing) comprises, in order, the NP5 sequence (P5 sequence, nucleotides 1-20 of SEQ ID NO. 4), the index sequence I (nucleotides 30-37 of SEQ ID NO. 4), and the adapter sequence I (nucleotides 38-51 of SEQ ID NO. 4, identical to nucleotides 1-14 of Tn5-Ad1, i.e., SEQ ID NO. 1 in the sequence listing). The N5 sequencing primer pairs with a partially complementary sequence of Tn5-Ad1 via the adapter sequence I, thereby simultaneously initiating the first round of PCR amplification and introducing the P5 sequence.
[0092] The N7 sequencing primer (SEQ ID NO. 5 in the sequence listing) comprises, in order, the NP7 sequence (P7 sequence, nucleotides 1-21 of SEQ ID NO. 5), the index sequence II (nucleotides 25-32 of SEQ ID NO. 5), and the adapter sequence II (nucleotides 33-53 of SEQ ID NO. 5, identical to nucleotides 1-21 of Tn5-Ad2, i.e., SEQ ID NO. 2 in the sequence listing). The N7 sequencing primer pairs with a partially complementary sequence of Tn5-Ad2 via the adapter sequence II, thereby simultaneously initiating the first round of PCR amplification and introducing the P7 sequence.
[0093] 2.5 Purification of the whole RNA sequencing library.
[0094] Use 50 µL of DNA screening magnetic beads to purify the PCR product obtained in step 2.4 above to obtain the whole RNA sequencing library of the sample.
[0095] Example 2. Extracting cfRNA from plasma samples and constructing a cfRNA sequencing library of low-abundance rRNA.
[0096] The samples to be tested in this embodiment are plasma from 10 pregnant women (sample volume: 200-1000 μL, e.g. Figure 6 As shown in Figure 1, the volume of sample 1 is 200 µL, the volume of sample 2 is 200 µL, the volume of sample 3 is 700 µL, the volume of sample 4 is 800 µL, the volume of sample 5 is 850 µL, the volume of sample 6 is 900 µL, the volume of sample 7 is 1000 µL, the volume of sample 8 is 1000 µL, the volume of sample 9 is 700 µL, and the volume of sample 10 is 900 µL. This example describes in detail the method for extracting cfRNA from trace plasma samples and constructing a cfRNA sequencing library for low-abundance ribosomal RNA (rRNA). The quality of the resulting library is then evaluated to evaluate the method.
[0097] 1. Construction of cfRNA sequencing library for low-abundance rRNA.
[0098] First, using the same method as in Example 1, full cfRNA sequencing libraries were obtained for 10 plasma samples (samples 1 to 10); Then, based on the gRNA of the PCR product corresponding to the rRNA in the whole cfRNA sequencing library ( Figure 7 Grs-1~Grs-29 and Figure 8 The whole cfRNA sequencing library of 8 samples (samples 1 to 8) was gene-edited to obtain the whole cfRNA sequencing library after rRNA shearing; Finally, a primer mixture containing P5 and P7 was used to amplify the full cfRNA sequencing library after rRNA shearing, and the PCR product was purified using 25µL DNA screening magnetic beads to obtain cfRNA sequencing libraries of low-abundance rRNA for 8 samples (samples 1 to 8).
[0099] The PCR reaction program was as follows: initial denaturation at 98°C for 30 s; 10 cycles of denaturation at 98°C for 10 s, annealing at 60°C for 20 s, and extension at 72°C for 30 s; final extension at 72°C for 2 min; and hold at 4°C.
[0100] In the process of obtaining full-cfRNA sequencing libraries for 10 plasma samples (samples 1 to 10), four different N5 primers were obtained by changing the index sequence I of N5: N5-index1, N5-index2, N5-index3, and N5-index4; and four different N7 primers were obtained by changing the index sequence II of N7: N7-index1, N7-index2, N7-index3, and N7-index4. By combining different N5 primers and N7 primers, full-cfRNA sequencing libraries for multiple samples (10 samples in this example) can be constructed simultaneously. The sequences of the four N5 primers and four N7 primers are as follows: N5-index1 (5'-3'): AATGATACGGCGACCACCGAGATCTACACTTCTAGCTTCGTCGGCAGCGTC (SEQ ID NO. 6); N5-index2 (5'-3'): AATGATACGGCGACCACCGAGATCTACACCCTAGAGTTCGTCGGCAGCGTC (SEQ ID NO. 7); N5-index3 (5'-3'): AATGATACGGCGACCACCGAGATCTACACGCGTAAGATCGTCGGCAGCGTC (SEQ ID NO. 8); N5-index4 (5'-3'): AATGATACGGCGACCACCGAGATCTACACCTATTAAGTCGTCGGCAGCGTC (SEQ ID NO. 9); N7-index1 (5'-3'): CAAGCAGAAGACGGCATACGAGATACCCAGCAGTCTCGTGGGCTCGGAGATGT (SEQ ID NO. 10); N7-index2 (5'-3'): CAAGCAGAAGACGGCATACGAGATAACCCCTCGTCTCGTGGGCTCGGAGATGT (SEQ ID NO. 11); N7-index3 (5'-3'): CAAGCAGAAGACGGCATACGAGATCCCAACCTGTCTCGTGGGCTCGGAGATGT (SEQ ID NO. 12); N7-index4 (5'-3'): CAAGCAGAAGACGGCATACGAGATCACCACACGTCTCGTGGGCTCGGAGATGT (SEQ ID NO. 13).
[0101] In the process of constructing cfRNA sequencing libraries for low-abundance rRNA in 8 samples (samples 1 to 8), the specific steps for gene editing of the whole cfRNA sequencing library based on gRNA targeting the PCR products corresponding to rRNA in the whole cfRNA sequencing library are as follows: 1) Prepare Reagent 5 and incubate at 25°C for 15 minutes. Reagent 5 contains: 0.4 µL 10× Cas9 Buffer (NEB, M0386), 0.3 µL Guide-rRNA Mix, 0.7 µL Cas9 Nuclease (NEB, M0386), and 0.6 µL nuclease-free water. The Guide-rRNA Mix contains 58 Guide-rRNAs (SEQ ID NO. 16-SEQ ID NO. 73 in the sequence listing, corresponding to Figure 7 Grs-1~Grs-29 and Figure 8 Grs-30~Grs-58).
[0102] 2) 14.4 µL of the full cfRNA sequencing library obtained from the plasma sample using the same method as in Example 1 was mixed with 1.6 µL of 10× Cas9 Buffer, and then Reagent 5 was added and incubated. The reaction conditions were 37°C for 2.5 hours and 65°C for 5 minutes.
[0103] 3) Purify the incubation product from step 2) using 20 µL of DNA screening magnetic beads to obtain a 10 µL full-cfRNA sequencing library after target DNA fragmentation.
[0104] 4) Mix the target DNA fragmented whole-cfRNA sequencing library obtained in step 3) with Reagent 6 for PCR reaction to obtain PCR products. Reagent 6 contains the following: 5 µL DNA Polymerase Buffer (NEB, M0491), 0.5 µL dNTPs (NEB, N0447), 0.25 µL DNA Polymerase (NEB, M0491), 1.25 µL P5, 1.25 µL P7, and 6.75 µL nuclease-free water.
[0105] The PCR reaction program was as follows: initial denaturation at 98°C for 30 s; 8 cycles of denaturation at 98°C for 10 s, annealing at 60°C for 20 s, and extension at 72°C for 30 s; final extension at 72°C for 2 min; and hold at 4°C.
[0106] 5) Purify the PCR product obtained in step 4) above using 25 µL of DNA screening magnetic beads to obtain the final low-abundance rRNA cfRNA sequencing library.
[0107] 2. On-machine sequencing and rRNA ratio analysis.
[0108] The low-abundance rRNA cfRNA sequencing libraries of 8 samples (samples 1-8) and the whole cfRNA sequencing libraries of 2 samples (samples 9 and 10) obtained in step 1 were sequenced using the T7 sequencer of the BGI sequencing platform to obtain the raw sequencing data of the cfRNA sequencing libraries of 10 plasma samples.
[0109] The raw sequencing data were subjected to data quality control, including removal of adapter sequences and low-quality bases. The data were then aligned to the human genome reference sequence (GRCh38) using STAR alignment software, retaining unique alignments and multiple alignments with no more than 10 positions. Transcript quantification was performed using RSEM software to generate gene and transcript expression matrices. MultiQC was used to integrate analysis reports from various software packages. Alignment was assessed using RNA-SeQC (v2.3.5), primarily analyzing the proportion of sequencing reads corresponding to ribosomal RNA in the total sequencing data and the proportion of exonic reads in the remaining reads after removing RCR repeats and sequencing reads corresponding to ribosomes.
[0110] The results are as follows Figure 4 and Figure 6 As shown, 8 samples after gene editing of the whole cfRNA sequencing library by gRNA targeting the PCR products corresponding to rRNA in the whole cfRNA sequencing library ( Figure 4 and Figure 6 rRNA abundance in cfRNA sequencing libraries of samples 1 to 8 ( Figure 4 and Figure 6 The proportion of rRNA in the whole cfRNA sequencing library was less than 32%; however, the rRNA abundance in the whole cfRNA sequencing library of the two samples without rRNA corresponding PCR product removal was very high ( Figure 4 and Figure 6The rRNA proportions in samples 9 and 10 were 87.93% and 89.26%, respectively. Therefore, the method of the present invention can effectively reduce the proportion of rRNA-derived reads in RNA sequencing libraries and increase the probability of capturing other RNAs.
[0111] Example 3. Construction of RNA sequencing libraries of low-abundance rRNA using total RNA of different qualities.
[0112] In this example, total RNA from human brain cells (1 μg / μL, Takara, 634485) was diluted with nuclease-free water at different ratios to obtain six RNA concentrations: 1 μg / μL was diluted three times at a ratio of 1:10 to 1 ng / μL; 1 ng / μL was diluted at a ratio of 1:2 to 500 pg / μL; 1 ng / μL was diluted at a ratio of 1:5 to 200 pg / μL; 1 ng / μL was diluted at a ratio of 1:10 to 100 pg / μL; 500 pg / μL was diluted at a ratio of 1:10 to 50 pg / μL; and 200 pg / μL was diluted at a ratio of 1:10 to 20 pg / μL.
[0113] Take 1 μL of the above 6 concentrations of RNA respectively to obtain 6 RNA samples with different mass gradients (1 ng, 500 pg, 200 pg, 100 pg, 50 pg and 20 pg) (corresponding to Figure 5 and Figure 6 Samples 11-16 in the .
[0114] The purified cfRNA in step 2 of Example 1 was replaced with the above 6 RNA samples with different mass gradients, and the operation in step 2 of Example 1 was followed to obtain human brain cell whole RNA sequencing libraries of the 6 RNA samples with different mass gradients.
[0115] Using the same method as in step 1 of Example 2, low-abundance rRNA RNA sequencing libraries were constructed for six human brain cell whole RNA sequencing libraries with RNA samples of different quality gradients (referred to as six whole RNA sequencing libraries), and the quality of the obtained low-abundance rRNA RNA sequencing libraries was identified using the same method as in step 2 of Example 2.
[0116] The results showed that the RNA sample from sample 11 (1ng, Figure 6 The proportion of ribosomal RNA (rRNA) in the RNA sequencing libraries of low-abundance rRNA of sample 1 (1000 pg of RNA sample), sample 12 (500 pg of RNA sample), sample 13 (200 pg of RNA sample), sample 14 (100 pg of RNA sample), sample 15 (50 pg of RNA sample) and sample 16 (20 pg of RNA sample) was less than 30%.
[0117] Therefore, the method of the present invention can be used to construct low-abundance rRNA RNA sequencing libraries for extremely small amounts of RNA greater than or equal to 20 pg. The rRNA proportion of the six RNA sequencing library samples was less than 30% ( Figure 5 and Figure 6 Samples 11-16).
[0118] The present invention has been described in detail above. For those skilled in the art, without departing from the purpose and scope of the present invention, and without the need to carry out unnecessary experimental conditions, the present invention can be implemented in a wide range under equivalent parameters, concentrations and conditions. Although the present invention provides specific embodiments, it should be understood that further improvements can be made to the present invention. In short, according to the principles of the present invention, this application is intended to include any changes, uses or improvements to the present invention, including changes that depart from the disclosed scope in this application and are made using conventional techniques known in the art.
Claims
1. A method for constructing an RNA sequencing library, characterized in that: The method comprises the following steps: A1) Reverse transcription of RNA from biological samples to obtain DNA-RNA hybrid chains; A2) cleaving the DNA-RNA hybrid chain to obtain DNA-RNA hybrid chain fragments, and connecting a linker M1 and a linker M2 to both ends of the DNA-RNA hybrid chain fragments to obtain linker-added DNA-RNA hybrid chain fragments; A3) using the adapter-added DNA-RNA hybrid fragment as a template, performing PCR amplification using sequencing primer 1 and sequencing primer 2 specific to the sequencing platform to obtain a PCR product, and purifying the PCR product to obtain an RNA sequencing library; The sequencing primer 1 comprises a sequencing platform adapter 1 and an adapter N1; the sequencing primer 2 comprises a sequencing platform adapter 2 and an adapter N2; The linker N1 is a single-stranded nucleotide fragment selected from the linker M1; the linker N2 is a single-stranded nucleotide fragment selected from the linker M2.
2. The method according to claim 1, wherein: The method further comprises the following step A4): A4) Genome editing is performed on the RNA sequencing library using a gRNA targeting an rRNA-derived PCR product and an RNA-guided nuclease to obtain an RNA sequencing library after target DNA shearing. PCR amplification is performed on the RNA sequencing library after target DNA shearing using the sequencing platform adapter 1 and the sequencing platform adapter 2 to obtain PCR product 2. The PCR product 2 is purified to obtain a purified RNA sequencing library.
3. The method according to claim 1 or 2, characterized in that: The mass of the RNA described in A1) is greater than or equal to 20 pg.
4. The method according to claim 3, wherein: The mass of the RNA in A1) is greater than or equal to 20 pg and less than or equal to 10 ng.
5. The method according to claim 1 or 2, characterized in that: The biological sample is animal body fluid.
6. A reagent for constructing an RNA sequencing library, characterized in that: The reagents include reverse transcriptase, the adapter M1 according to any one of claims 1 to 5, the adapter M2, the sequencing primer 1, the sequencing primer 2, the sequencing platform adapter 1 and the sequencing platform adapter 2.
7. The reagent according to claim 6, characterized in that: The reagents also include gRNA and RNA-guided nuclease that targets rRNA-derived PCR products.
8. The reagent according to claim 6 or 7, characterized in that: The reagents also include a transposase complex and a DNA polymerase.
9. The reagent according to claim 8, characterized in that: The transposase complex is a transposase Tn5 complex.
10. Use of the reagent according to any one of claims 6 to 9 in constructing an RNA sequencing library.
Citation Information
Patent Citations
Method for removing connection by-products of 5' and 3' adapters during construction of sequencing library
CN107488655A
CRISPR assisted DNA target enrichment method and application thereof
CN109837273A
Method and kit for simultaneously constructing sequencing library by DNA and RNA
CN111139532A
Sequencing library construction method and kit for pathogenic microorganism detection
CN111188094A
Transposome enabled dna / rna-sequencing (ted rna-seq)
CN112689673A