A method for constructing a trace RNA sequencing library and reagents used thereby

By using the transposase Tn5 complex and gene editing technology targeting rRNA, the problems of sample loss, long time and high cost in the construction of micro-cfRNA sequencing libraries have been solved, realizing efficient and low-cost cfRNA sequencing library construction, which is suitable for high-throughput sequencing of micro-RNA samples.

CN120665989BActive Publication Date: 2025-12-12PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511178624.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-12-12
Estimated Expiration
2045-08-22

AI Technical Summary

Technical Problem

Existing technologies struggle to construct high-quality micro-volume cell-free RNA (cfRNA) sequencing libraries. Problems include limited sample RNA acquisition, low aptamer ligation efficiency, cumbersome and time-consuming library construction process, high cost, and a high proportion of rRNA, which restrict the research and application of cfRNA in diagnostic medicine.

Method used

The transposase Tn5 complex was used to complete the first round of PCR amplification of cfRNA in a single tube. Gene editing was performed by combining gRNA from PCR products targeting rRNA and RNA-directed nucleases, reducing purification steps and time, and lowering costs. Adapter fragments were ligated through adapters M1 and M2, and PCR amplification was performed using adapters 1 and 2 on a sequencing platform to construct a high-throughput sequencing library.

Benefits of technology

It enables the construction of high-quality micro-RNA sequencing libraries in a short time, reduces the proportion of rRNA, increases the capture probability of other RNAs, reduces library construction costs and time, and is suitable for RNA samples smaller than 10ng or even pg.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120665989B_ABST
    Figure CN120665989B_ABST
Patent Text Reader

Abstract

The application discloses a method for constructing a trace RNA sequencing library in the field of a nucleic acid determination method and reagents used in the method. The application aims to solve the technical problem of how to construct a trace RNA sequencing library. The method comprises the following steps: performing reverse transcription on RNA to obtain DNA-RNA hybrid chains, cutting the hybrid chains to obtain hybrid chain fragments, adding a linker M to both ends of the hybrid chain fragments to obtain DNA-RNA hybrid chain fragments with linkers, using the DNA-RNA hybrid chain fragments with linkers as templates, performing PCR amplification on the templates by using sequencing primers for a sequencing platform to obtain PCR products, and purifying the PCR products to obtain a RNA sequencing library. The sequencing primers comprise a sequencing platform linker and a linker N, and the linker N is a single-stranded nucleotide fragment selected from the linkers M. The application can be used for constructing a trace RNA sequencing library.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of nucleic acid assay methods, specifically relating to a method for constructing a micro-RNA sequencing library and the reagents used therein. Background Technology

[0002] Cell-free RNA (cfRNA) in body fluids, such as plasma or serum cfRNA, exhibits greater tissue specificity and diversity compared to existing DNA and protein biomarkers. cfRNA originates from various cell death pathways, including apoptosis, necrosis, pyroptosis, and the release of extracellular vesicles. These molecules can be emitted from various human cell types in remote organ systems or non-human microorganisms. As they circulate in biological systems, cfRNA exhibits significant resistance to degradation due to its binding to protective proteins, lipids, or extracellular vesicles. Therefore, using trace amounts of cfRNA in plasma or serum to diagnose diseases has become a research focus of considerable interest in diagnostic medicine. Currently, RNA in plasma or serum can be detected, and numerous studies hope to identify new diagnostic biomarkers from it.

[0003] However, the main problem encountered in practice is how to construct high-quality, high-throughput sequencing libraries from cfRNA. Existing methods require sufficient RNA (at least >10 ng) and then construct the sequencing library by ligating aptamers. The problems with these methods are: 1) Obtaining plasma or serum samples is limited, making it impossible to extract more than 10 ng of RNA; 2) The aptamer ligation method is inefficient, and multiple necessary purification steps during library construction cause significant losses; 3) The library construction process is cumbersome and time-consuming, generally requiring more than 3 days, resulting in low timeliness; 4) The overall library construction cost is high (>1200 RMB / sample); 5) The high proportion of rRNA-derived reads in the library (>75%) necessitates increasing sequencing depth to obtain sufficient information, further increasing costs. All these problems limit the research progress and clinical application of extracellular free RNA from plasma or serum. Summary of the Invention

[0004] The technical problems to be solved by this invention are how to construct sequencing libraries of trace RNA and / or how to construct sequencing libraries of trace RNA with low abundance of rRNA and / or how to construct high-throughput sequencing libraries of trace extracellular free RNA and / or how to construct sequencing libraries of extracellular free RNA in plasma or serum and / or how to construct sequencing libraries of extracellular free RNA in body fluids and / or how to construct sequencing libraries of low rRNA abundance of extracellular free RNA in body fluids.

[0005] To effectively solve the above-mentioned technical problems, the present invention provides a method for constructing an RNA sequencing library, the method comprising the following steps:

[0006] A1) Obtain DNA-RNA hybrid strands by reverse transcription of RNA from biological samples;

[0007] A2) The DNA-RNA hybrid chain is cut to obtain a DNA-RNA hybrid chain fragment, and adapters M1 and M2 are connected to both ends of the DNA-RNA hybrid chain fragment to obtain a DNA-RNA hybrid chain fragment with adapters added.

[0008] A3) Using the DNA-RNA hybrid chain fragment with the added adapter as a template, PCR amplification (first round) is performed using sequencing primers 1 and 2 for the sequencing platform to obtain PCR products. The PCR products are then purified to obtain a whole RNA sequencing library.

[0009] The sequencing primer 1 may include a sequencing platform adapter 1 and an adapter N1. The sequencing primer 2 may include a sequencing platform adapter 2 and an adapter N2.

[0010] The linker N1 may be a single-stranded nucleotide fragment selected from the linker M1. The linker N2 may be a single-stranded nucleotide fragment selected from the linker M2.

[0011] The length of connector N1 can be greater than or equal to 14nt and less than or equal to the length of one chain of connector M1. The length of connector N2 can be greater than or equal to 14nt and less than or equal to the length of one chain of connector M2. In one specific embodiment of the present invention, the length of connector N1 is 14nt and the length of connector N2 is 21nt.

[0012] The method described above may also include the following step (A4):

[0013] A4) Gene editing of the RNA sequencing library using gRNA from PCR products targeting rRNA and RNA-guided nucleases to obtain RNA sequencing library with target DNA fragmentation. The RNA sequencing library with target DNA fragmentation is then subjected to (second round) PCR amplification using sequencing platform adapter 1 and sequencing platform adapter 2 to obtain PCR product 2. The PCR product 2 is then purified to obtain purified RNA sequencing library.

[0014] In the above method, the adapter N1 in sequencing primer 1 can pair with a partially complementary sequence of adapter M1 to introduce the sequencing platform adapter 1 simultaneously with the initiation of the first round of PCR amplification. Similarly, the adapter N2 in sequencing primer 2 can pair with a partially complementary sequence of adapter M2 to introduce the sequencing platform adapter 2 simultaneously with the initiation of the first round of PCR amplification.

[0015] In the above method, sequencing primer 1 and / or sequencing primer 2 may further include a tag sequence.

[0016] The sequences of connector M1 and connector M2 may be the same.

[0017] In the above method, the mass of the RNA described in A1) can be greater than or equal to 20 pg.

[0018] In some specific embodiments of the present invention, the mass of the RNA described in A1) is greater than or equal to 20 pg and less than or equal to 10 ng.

[0019] In some specific embodiments of the present invention, the greater than or equal to 20pg and less than or equal to 10ng is 20pg-50pg, 20pg-100pg, 50pg-200pg, 100pg-500pg, 200pg-1ng or 500pg-10ng.

[0020] In the above method, the biological sample may be animal body fluid.

[0021] In the above method, step A2) can be achieved by transposase Tn5 complex.

[0022] In the above method, the sequencing platform adapter 1 sequence and the sequencing platform adapter 2 can be specific adapter sequences used in the construction of sequencing platform libraries. In one specific embodiment of the present invention, the sequence of sequencing platform adapter 1 can be the P5 sequence of the Illumina sequencing platform. The sequence of sequencing platform adapter 2 can be the P7 sequence of the Illumina sequencing platform.

[0023] In one specific embodiment of the present invention, the sequence of sequencing primer 1 is SEQ ID NO.4 in the sequence listing, and the sequence of sequencing primer 2 is SEQ ID NO.5 in the sequence listing.

[0024] In the above method, the body fluid refers to liquid substances in the body, specifically including blood plasma, tissue fluid, cerebrospinal fluid, lymph, intracellular fluid, etc. The volume of the body fluid can be greater than or equal to 200µL.

[0025] In the above method, the mass of the RNA from the biological sample can be greater than or equal to 20 pg.

[0026] The animals mentioned above can be mammals, and the mammals can be humans.

[0027] In the above method, the reverse transcriptase can be SuperScript IV.

[0028] In the above method, the gRNA-guided polypeptide with nuclease activity can be a CRISPR-Cas protein. In one specific embodiment of the present invention, the CRISPR-Cas protein is Cas9.

[0029] The term "sgRNA (single-guide RNA) or gRNA" is a component of the CRISPR-Cas system, responsible for guiding the Cas protein to recognize and cleave target nucleic acid molecules. In practical gene editing applications, sgRNA can be synthesized directly or obtained through plasmid expression or in vitro transcription. In this field, "gRNA" and "sgRNA" are often used interchangeably. In this document, "gRNA" and "sgRNA" are also used interchangeably. sgRNA generally refers to a single RNA structure formed by artificially modifying the crRNA / tracrRNA complex (gRNA) with a dual RNA structure, directly (or through a linker) linking the crRNA and tracrRNA. sgRNA is a short RNA containing a recognition region and a framework region.

[0030] The term "recognition region," also known as a guide sequence, is typically an RNA sequence (referred to herein as the "guide sequence") that is identical to or complementary to the target sequence or target site within the target RNA (sgRNA or gRNA). The guide sequence is generally sufficiently complementary to the target sequence to hybridize with it and guide the CRISPR / Cas complex to specifically bind to the target sequence. Perfect complementarity between the guide sequence and the target sequence is preferred, but some mismatch (e.g., a mismatch of 1-6 nucleotides) is permissible, as long as it still results in gene knockout. The complementarity between the guide sequence and its corresponding target sequence is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. Methods for determining the complementarity of two nucleic acid sequences are within the capabilities of those skilled in the art.

[0031] The term "scaffold" generally refers to the structural or scaffold RNA sequence required to guide the binding or interaction of RNA with RNA-directed nucleases and / or other RNA molecules (e.g., tracrRNA) within RNA (sgRNA or gRNA), and can also be called the backbone sequence of sgRNA. The scaffold can be conventionally selected by those skilled in the art; for example, it can be the backbone sequence of the sgRNA corresponding to Cas9, or it can be a mutant constructed based on this sequence that still retains the function of binding the corresponding Cas9.

[0032] In this application, "RNA-directed nuclease" refers to RNA-directed DNA endonucleases associated with the CRISPR system. Unrestricted examples of RNA-directed nucleases include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, and Cpf1, their homologs, or modified forms thereof. In one implementation, the RNA-directed nuclease is Cas9.

[0033] Furthermore, the Cas9 protein described in this application is not limited to a specific protein, as long as it can be used in conjunction with the gRNA of this invention. Furthermore, the Cas9 protein described herein is selected from Streptococcus pyogenes Cas9 (spCas9, subtype II-A), spCas9HF (high fidelity), nickase Cas9 (nCas9), Staphylococcus aureus Cas9 (saCas9, subtype II-A), Neisseria meningitidis Cas9 (NmCas9, subtype II-C), Francisella novicida Cas9 (FnCas9, subtype II-B), Streptococcus thermophilus Cas9 (St1Cas9, St3Cas9), Campylobacter jejuni Cas9 (CjCas9), and Treponema pallidum Cas9, as well as orthologs of Cas9 from other organisms, but not limited to these. The Cas9 protein may also include high-fidelity Cas9 mutants (such as SpCas9-HF1, eSpCas9-1.1, and TrueCut™ HiFiCas9 protein).

[0034] To address the aforementioned technical problems, the present invention also provides reagents for constructing RNA sequencing libraries, wherein the reagents may include reverse transcriptase, adapter M1 and adapter M2 as described above, sequencing primer 1, sequencing primer 2, sequencing platform adapter 1, and sequencing platform adapter 2.

[0035] The reagents mentioned above may also include gRNA targeting PCR products derived from rRNA and RNA-directed nucleases.

[0036] The reagents mentioned above may also include transposase complexes and DNA polymerases.

[0037] In the above reagents, the transposase complex can be the transposase Tn5 complex.

[0038] The above reagents may also include SuperScript IV reverse transcriptase.

[0039] The sequence of connector M1 described above may be SEQ ID NO.1 in the sequence list. The sequence of connector M2 described above may be SEQ ID NO.2 in the sequence list.

[0040] In one specific embodiment of the present invention, the sequencing platform adapter 1 may be P5 of the Illumina sequencing platform, and the sequencing platform adapter 2 may be P7 of the Illumina sequencing platform.

[0041] In one specific embodiment of the present invention, the sequence of sequencing primer 1 is SEQ ID NO.4 in the sequence listing, and the sequence of sequencing primer 2 is SEQ ID NO.5 in the sequence listing.

[0042] In one specific embodiment of the present invention, the gRNA targeting the PCR product derived from rRNA is SEQ ID NO.16-SEQ ID NO.73 in the sequence listing.

[0043] To address the aforementioned technical problems, this invention also provides the application of the reagents described above in constructing RNA sequencing libraries.

[0044] To address the aforementioned technical problems, the present invention also provides a method for constructing an RNA sequencing library, the method comprising steps A1)-A3) of the method described above.

[0045] This invention comprehensively improves upon the problems encountered in constructing high-throughput sequencing libraries using trace amounts of cfRNA by utilizing the properties of Tn5. It allows the entire process from cfRNA to the first round of PCR amplification (cfRNA to whole RNA sequencing library) to be completed in a single tube, avoiding sample loss caused by mid-process purification in other library construction methods, while significantly reducing library construction time and cost. This invention also includes the removal of rRNA-derived DNA from the library, effectively reducing the proportion of rRNA-derived reads and increasing the probability of capturing other RNAs. Using the technology of this invention, libraries can be constructed for trace amounts of RNA less than 10 ng, or even pg, in a short time. High-quality data can then be obtained through high-throughput sequencing, and analysis can identify disease-related RNA biomarkers.

[0046] The beneficial effects of this invention are:

[0047] The RNA sequencing library construction process provided by this invention reduces the number of RNA purification steps and minimizes RNA loss; it can achieve library construction with trace amounts of RNA (such as 20 pg RNA or cfRNA). The method of this invention can effectively reduce the proportion of rRNA in RNA sequencing libraries from biologically derived samples (below 32%). The library construction process of this invention is time-efficient and low-cost, and can be widely adopted. Attached Figure Description

[0048] Figure 1 This is an experimental schematic diagram of the method of the present invention.

[0049] Figure 2 This is a schematic diagram of the process of the present invention.

[0050] Figure 3 This is a schematic diagram of the transposase Tn5 complex in an embodiment of the present invention.

[0051] Figure 4 The results show the proportion of ribosomal RNA and the proportion of exons in the alignment region of the RNA sequencing library constructed using the method of the present invention from 10 plasma samples in Example 2.

[0052] Figure 5 The results show the ribosomal RNA percentage and alignment region exon percentage analysis results in the RNA sequencing library constructed using the method of the present invention from the total RNA (trace amount) derived from 6 human brain cells in Example 3.

[0053] Figure 6 Data related to the construction of RNA sequencing libraries from 16 biological samples in Examples 2 and 3.

[0054] Figure 7 This contains the sequence information of 29 Guide-rRNAs (Grs-1~Grs-29) contained in the Guide-rRNA Mix.

[0055] Figure 8 This contains the sequence information of 29 Guide-rRNAs (Grs-30~Grs-58) contained in the Guide-rRNA Mix. Detailed Implementation

[0056] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.

[0057] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.

[0058] The specific nucleic acid sequences used in the transposase Tn5 complex, whole RNA sequencing library, and low-abundance rRNA sequencing library in the following examples are as follows, where N represents A, T, C, or G.

[0059] Tn5-Ad1 (5'-3'):TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG (SEQ ID NO.1);

[0060] Tn5-Ad2 (5'-3'): GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG (SEQ ID NO. 2);

[0061] Tn5-ME (5'-3'): pCTGTCTCTTATACACATCT (SEQ ID NO.3, 5' terminal nucleotide C is phosphorylated);

[0062] N5 (5'-3'):

[0063] AATGATACGGCGACCACCGAGATCTACACNNNNNNNNTCGTCGGCAGCGTC (SEQ ID NO. 4);

[0064] N7 (5'-3'):

[0065] CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTCTCGTGGGCTCGGAGATGT (SEQ ID NO.5);

[0066] P5 (5'-3'):AATGATACGGCGACCACCGA (SEQ ID NO. 14);

[0067] P7 (5'-3'): CAAGCAGAAGACGGCATACGA (SEQ ID NO. 15).

[0068] Example 1. Extracting cfRNA from plasma samples and constructing a whole RNA sequencing library.

[0069] This embodiment uses plasma as the test sample and describes in detail the specific process of extracting cfRNA from plasma samples and constructing a whole RNA sequencing library.

[0070] 1. Extract and purify cfRNA from plasma.

[0071] 1.1 Extraction of extracellular free nucleic acids.

[0072] Extracellular free RNA (cfRNA) was extracted and purified from plasma samples using a serum / plasma cell-free RNA extraction kit (Jianshi Biosciences, TR159). The specific steps included:

[0073] 1) Pipette 1 mL of plasma sample into a 15 mL tube, add 1 mL of the free RNA digestion solution (mainly enzyme reaction buffer) from the kit, and mix well.

[0074] 2) Add 50 μL of proteinase K solution (50 units, prepared by mixing proteinase K and proteinase K preservation solution from the kit), mix well, and incubate at 37°C for 2 hours to digest the protein in the sample.

[0075] 3) Add 2 mL of free RNA binding buffer (the main component of which is nucleic acid binding buffer) and mix well.

[0076] 4) Add 6 mL of 100% (volume fraction) isopropanol (to promote nucleic acid precipitation), mix well, and obtain crude free nucleic acid extract.

[0077] 5) Place the 25mL funnel onto filter column 1 (column Y, No. 3) of the kit and assemble the filter device.

[0078] 6) Transfer the crude free nucleic acid extract obtained in step 4) to the filter tube in step 5), turn on the vacuum pump to filter the crude free nucleic acid extract completely through the filter column 1 matrix, turn off the vacuum pump, and discard the 25mL funnel.

[0079] 7) Add 600 μL of RNA washing buffer 1 to the matrix of filter column 1, turn on the vacuum pump (400 mmHg) to allow the liquid to flow through, and then turn off the vacuum pump.

[0080] 8) Transfer filter column 1 to a collection tube and centrifuge at 12000g for 2 minutes to remove residual liquid.

[0081] 9) Add 700µL RNA washing buffer 2 (75% ethanol by volume to remove salt ions), centrifuge at 12000g for 30s, and discard the waste liquid.

[0082] 10) Add 400 μl of RNA washing buffer 2 (75% ethanol by volume to remove salt ions), and centrifuge at 12000g for 2 mins to remove the liquid completely.

[0083] 11) Transfer filter column 1 to a clean 1.5 mL centrifuge tube, add 80 μL of nuclease-free water from the kit to filter column 1, incubate for 2 min, centrifuge at 12000 g for 30 s to elute and obtain purified free nucleic acid.

[0084] 1.2 DNA was removed from extracellular free nucleic acids to obtain cfRNA.

[0085] The purified free nucleic acids obtained in step 1.1 were further purified and concentrated using an RNA purification and concentration kit (Jianshi Biotech TR113) to remove DNA and obtain cfRNA. The specific steps are as follows:

[0086] 12) Add the purified free nucleic acid obtained in step 1.1 to the DNA digestion solution (10 µL), DNase I (10 µL), and RNase Inhibitor (Vazyme, R301) in the kit.

[0087] 13) Incubate the entire reaction system at room temperature for 15 minutes.

[0088] 14) Add 200µL of the RNA binding solution from the kit to the incubated system and mix well.

[0089] 15) Add 300 mL of 95-100% (volume fraction) ethanol, mix and transfer to column 1 of the kit containing filter column 2, place in a clean collection tube, centrifuge at 12000 g for 30 s to ensure the mixture passes completely through the filter column 2 matrix, and discard the waste liquid.

[0090] 16) Add 400 µL of RNA pre-wash buffer from the kit, centrifuge at 12000g for 30s, and discard the waste liquid.

[0091] 17) Add 700 µL of RNA washing buffer from the kit, centrifuge at 12000g for 30s, and discard the waste liquid.

[0092] 18) Add 400 µL of RNA washing buffer from the kit, centrifuge at 12000g for 1 min to remove the liquid completely.

[0093] 19) Transfer filter column 2 to a clean 1.5 mL centrifuge tube, add 10 µL of nuclease-free water from the kit (preheated in a 65-70°C water bath beforehand), incubate for 2 mins, centrifuge at 12000g for 30 s to elute and obtain purified cfRNA.

[0094] 2. Construction of whole RNA sequencing library.

[0095] 2.1 cfRNA reverse transcription to obtain hybrid strands.

[0096] 1) First, add the purified cfRNA obtained in step 1.2 to reagent 1 and incubate. The reaction conditions are 65℃ for 10 mins and 72℃ for 3 mins. The components of reagent 1 are: 1 µL reverse transcription primer (NEB, S1330), 1.25 µL dNTPs (NEB, N0447) and 0.15 µL RNase inhibitor (Vazyme, R301).

[0097] 2) Then, the reaction system obtained from step 1) was mixed with reagent 2 and incubated to complete reverse transcription, yielding a DNA-RNA hybrid strand. The reaction conditions were 23℃ for 10 mins; 50℃ for 50 mins; and 80℃ for 10 mins. Reagent 2 consisted of: 5 µL of 5×SSIV buffer (reverse transcriptase buffer, Thermo, 18090050), 1.25 µL of DTT (Thermo, 18090050), 0.6 µL of RNase Inhibitor (Vazyme, R301), 5 µL of Betaine (Merck, B0300), 0.5 µL of SuperScript IV reverse transcriptase (Thermo, 18090050), and 0.2 µL of nuclease-free water.

[0098] 2.2 Transposase Tn5 cleaves the hybrid strand and adds adapters.

[0099] 3) This invention uses the transposase Tn5 complex to fragment and ligate adapters into DNA-RNA hybrid chains, resulting in fragments of DNA-RNA hybrid chains with gaps added to Tn5-Ad1 and Tn5-Ad2 adapters. The added adapters (Tn5-Ad1 and Tn5-Ad2 adapters) are primer binding sites used to introduce PCR reactions.

[0100] Specifically, the reaction system containing the DNA-RNA hybrid strand obtained in step 2.1 was mixed with reagent 3 and incubated at 55°C for 20 mins. The components of reagent 3 were: 4.5 µL of 8×TDbuffer (prepared by: 8 µL of 1M Tris-HCl (Beyotime, ST780), 4 µL of 1M MgCl2 (Aladdin, M299562), 80 µL of N,N-dimethylformamide DMF (Merck, D4551) and 80 µL of nuclease-free water), 5.25 µL of 50% PEG8000 (Beyotime, R0056), 0.8 µL of ATP (NEB, P0756), 0.4 µL of RNase Inhibitor (Vazyme, R301) and 0.05 µL of transposase Tn5 complex.

[0101] The transposase Tn5 complex was synthesized using primers containing the naked Tn5 enzyme (TransGen, LT201), the Tn5-Ad1 adapter (SEQ ID NO.1 in the sequence listing), the Tn5-Ad2 adapter (SEQ ID NO.2 in the sequence listing), and the Tn5-ME core sequence (SEQ ID NO.3 in the sequence listing), as follows:

[0102] (3-1) Prepare a mixture of Tn5-Ad1, Tn5-ME, and Annealing Buffer, then incubate mixture I at 95℃ for 2 min, followed by a temperature change of 0.1℃ / s to 22℃; then incubate at 22℃ for 5 min. (3-2) Prepare a mixture of Tn5-Ad2, Tn5-ME, and Annealing Buffer, then incubate mixture II at 95℃ for 2 min, followed by a temperature change of 0.1℃ / s to 22℃; then incubate at 22℃ for 5 min. (3-3) Incubate mixtures I, II, Tn5 naked enzyme, and Tn5 Storage Buffer at 35℃ for 2 h to obtain the transposase Tn5 complex. Figure 3 ).

[0103] 2.3 Filling the gaps in DNA-RNA hybrid strands containing gaps.

[0104] 4) Mix the reaction system of the nicked DNA-RNA hybrid strand obtained in step 2.2 with reagent 4 and incubate to obtain a complete DNA-RNA hybrid strand fragment with the adapter after the nick is filled. The reaction conditions are: 72℃ for 15 mins; 80℃ for 5 mins. The components of reagent 4 are: 10.5 µL of DNA polymerase buffer, 1 µL of dNTPs (NEB, N0447) and 0.5 µL of DNA polymerase (NEB, M0491).

[0105] 2.4 PCR amplification with the addition of sequencing adapters.

[0106] 5) Mix the reaction system containing the complete DNA-RNA hybridization strand obtained in step 2.3 with a mixture containing N5 primer and N7 primer (containing 2.5 µL of 10 µM primer N5 and 2.5 µL of 10 µM primer N7), and then perform the first round of PCR reaction to obtain PCR products with the N5 and N7 sequencing primer sequences added.

[0107] The PCR reaction procedure is as follows: 98℃, 30s initial denaturation; 98℃, denaturation for 10 seconds, 60℃, annealing for 20 seconds, 72℃, extension for 30 seconds, 10 cycles; 72℃, final extension for 2 minutes; 4℃, hold.

[0108] The N5 sequencing primer (SEQ ID NO.4 in the sequence listing) contains, in sequence, the NP5 sequence (P5 sequence, nucleotides 1-20 of SEQ ID NO.4), index sequence I (nucleotides 30-37 of SEQ ID NO.4), and adapter sequence I (nucleotides 38-51 of SEQ ID NO.4, identical to nucleotides 1-14 of Tn5-Ad1, i.e., SEQ ID NO.1 in the sequence listing). The N5 sequencing primer pairs with a partially complementary sequence of Tn5-Ad1 via adapter sequence I, thereby introducing the P5 sequence simultaneously with initiating the first round of PCR amplification.

[0109] The N7 sequencing primer (SEQ ID NO. 5 in the sequence listing) contains, in sequence, the NP7 sequence (P7 sequence, nucleotides 1-21 of SEQ ID NO. 5), index sequence II (nucleotides 25-32 of SEQ ID NO. 5), and adapter sequence II (nucleotides 33-53 of SEQ ID NO. 5, identical to nucleotides 1-21 of Tn5-Ad2, i.e., SEQ ID NO. 2 in the sequence listing). The N7 sequencing primer pairs with a partially complementary sequence of Tn5-Ad2 via its adapter sequence II, thereby introducing the P7 sequence simultaneously with initiating the first round of PCR amplification.

[0110] 2.5 Purify to obtain a whole RNA sequencing library.

[0111] The PCR product obtained in step 2.4 above was purified using 50µL DNA screening magnetic beads to obtain the whole RNA sequencing library of the sample.

[0112] Example 2. Extracting cfRNA from plasma samples and constructing a cfRNA sequencing library with low abundance rRNA.

[0113] In this embodiment, the test samples were plasma from 10 pregnant women (sample volume: 200-1000µL, e.g., ...). Figure 6 As shown: Sample 1 volume 200µL, Sample 2 volume 200µL, Sample 3 volume 700µL, Sample 4 volume 800µL, Sample 5 volume 850µL, Sample 6 volume 900µL, Sample 7 volume 1000µL, Sample 8 volume 1000µL, Sample 9 volume 700µL, Sample 10 volume 900µL). This example will describe in detail the specific method for extracting cfRNA from trace plasma samples and constructing a cfRNA sequencing library of low-abundance ribosomal RNA (rRNA), and evaluate the method by identifying the quality of the obtained library.

[0114] 1. Construction of cfRNA sequencing libraries for low-abundance rRNA.

[0115] First, using the same method as in Example 1, whole cfRNA sequencing libraries were obtained from 10 plasma samples (sample 1-sample 10);

[0116] Then, based on the gRNA of the PCR product corresponding to the rRNA in the targeted whole cfRNA sequencing library ( Figure 7 Grs-1 to Grs-29 and Figure 8 Gene editing was performed on the whole cfRNA sequencing libraries of 8 samples (sample 1-sample 8) using Grs-30~Grs-58 to obtain whole cfRNA sequencing libraries with fragmented rRNA.

[0117] Finally, a primer mixture containing P5 and P7 was used to amplify the whole cfRNA sequencing library after rRNA fragmentation by PCR, and the PCR product was purified using 25µL DNA screening magnetic beads to obtain cfRNA sequencing libraries of low-abundance rRNA from 8 samples (sample 1-sample 8).

[0118] The PCR reaction procedure is as follows: 98℃, 30s initial denaturation; 98℃, denaturation for 10 seconds, 60℃, annealing for 20 seconds, 72℃, extension for 30 seconds, 10 cycles; 72℃, final extension for 2 minutes; 4℃, hold.

[0119] In obtaining whole cfRNA sequencing libraries from 10 plasma samples (samples 1-10), four different N5 primers were obtained by modifying the N5 tag sequence I: N5-index1, N5-index2, N5-index3, and N5-index4; and four different N7 primers were obtained by modifying the N7 tag sequence II: N7-index1, N7-index2, N7-index3, and N7-index4. By combining different N5 and N7 primers, whole cfRNA sequencing libraries from multiple samples (10 samples in this example) can be constructed simultaneously. The sequences of the four N5 primers and four N7 primers are as follows:

[0120] N5-index1 (5'-3'):

[0121] AATGATACGGCGACCACCGAGATCTACACTTCTAGCTTCGTCGGCAGCGTC (SEQ ID NO. 6);

[0122] N5-index2 (5'-3'):

[0123] AATGATACGGCGACCACCGAGATCTACACCCTAGAGTTCGTCGGCAGCGTC(SEQ ID NO.7);

[0124] N5-index3(5’-3’):

[0125] AATGATACGGCGACCACCGAGATCTACACGCGTAAGATCGTCGGCAGCGTC(SEQ ID NO.8);

[0126] N5-index4(5’-3’):

[0127] AATGATACGGCGACCACCGAGATCTACACCTATTAAGTCGTCGGCAGCGTC(SEQ ID NO.9);

[0128] N7-index1(5’-3’):

[0129] CAAGCAGAAGACGGCATACGAGATACCCAGCAGTCTCGTGGGCTCGGAGATGT(SEQ ID NO.10);

[0130] N7-index2(5’-3’):

[0131] CAAGCAGAAGACGGCATACGAGATAACCCCTCGTCTCGTGGGCTCGGAGATGT(SEQ ID NO.11);

[0132] N7-index3(5’-3’):

[0133] CAAGCAGAAGACGGCATACGAGATCCCAACCTGTCTCGTGGGCTCGGAGATGT(SEQ ID NO.12);

[0134] N7-index4(5’-3’):

[0135] CAAGCAGAAGACGGCATACGAGATCACCACACGTCTCGTGGGCTCGGAGATGT(SEQ ID NO.13)。

[0136] In constructing cfRNA sequencing libraries of low-abundance rRNA from 8 samples (samples 1-8), the specific steps for gene editing of the whole cfRNA sequencing libraries based on gRNA targeting the PCR products corresponding to the rRNA in the whole cfRNA sequencing libraries are as follows:

[0137] 1) Prepare reagent 5 and incubate at 25°C for 15 mins. Reagent 5 consists of: 0.4 µL 10×Cas9 Buffer (NEB, M0386), 0.3 µL Guide-rRNA Mix, 0.7 µL Cas9 Nuclease (NEB, M0386), and 0.6 µL nuclease-free water. The Guide-rRNA Mix contains 58 Guide-rRNAs (SEQ ID NO. 16-SEQ ID NO. 73 in the sequence listing, corresponding to...) Figure 7 Grs-1 to Grs-29 and Figure 8 Grs-30 to Grs-58 (among others).

[0138] 2) Take 14.4 µL of the whole cfRNA sequencing library obtained from the plasma sample using the same method as in Example 1, mix it with 1.6 µL of 10× Cas9 Buffer, add reagent 5, and incubate under the following conditions: 37 °C for 2.5 h; 65 °C for 5 mins.

[0139] 3) The incubation product from step 2) was purified using 20µL DNA screening magnetic beads to obtain 10µL of whole cfRNA sequencing library after target DNA fragmentation.

[0140] 4) Mix the whole cfRNA sequencing library obtained from the target DNA fragmentation in step 3) with reagent 6 and perform a PCR reaction to obtain the PCR product. The components of reagent 6 are as follows: 5 µL DNA polymerase buffer (NEB, M0491), 0.5 µL dNTPs (NEB, N0447), 0.25 µL DNA polymerase (NEB, M0491), 1.25 µL P5, 1.25 µL P7 and 6.75 µL nuclease-free water.

[0141] The PCR reaction procedure was as follows: 98℃, 30s initial denaturation; 98℃, denaturation for 10s, 60℃, annealing for 20s, 72℃, extension for 30s, 8 cycles; 72℃, final extension for 2 minutes; 4℃, hold.

[0142] 5) The PCR product obtained in step 4) was purified using 25µL DNA screening magnetic beads to obtain the final rRNA low-abundance cfRNA sequencing library.

[0143] 2. Sequencing and rRNA percentage analysis.

[0144] The T7 sequencer of the BGI sequencing platform was used to sequence the low-abundance rRNA cfRNA sequencing libraries of 8 samples (samples 1-8) and the whole cfRNA sequencing libraries of 2 samples (samples 9 and 10) obtained in step 1, and the raw sequencing data of cfRNA sequencing libraries of 10 plasma samples were obtained.

[0145] Raw sequencing data underwent quality control, removing adapter sequences and low-quality bases. The sequencing data was then aligned with the human genome reference sequence (GRCh38) using STAR alignment software, retaining unique alignments and multiple alignments with no more than 10 alignment positions. Transcript quantification was performed using RSEM software, generating gene and transcript expression matrices. MultiQC integrated the analysis reports from various software programs. The alignment results were evaluated using RNA-SeQC (v2.3.5), primarily analyzing the proportion of ribosomal RNA reads in the total sequencing data and the proportion of exon reads in the remaining reads after removing RCR repeats and ribosomal reads.

[0146] The results are as follows Figure 4 and Figure 6 As shown, eight samples were gene-edited using gRNA based on the PCR products corresponding to rRNA in the whole cfRNA sequencing library. Figure 4 and Figure 6 rRNA abundance in cfRNA sequencing libraries of samples 1-8 (in the middle samples) Figure 4 and Figure 6 The proportion of rRNA in both samples was less than 32%; while the rRNA abundance in the whole cfRNA sequencing libraries of the two samples for which PCR products corresponding to rRNA were not removed was very high. Figure 4 and Figure 6 The rRNA percentages in samples 9 and 10 were 87.93% and 89.26%, respectively. Therefore, the method of this invention can effectively reduce the proportion of rRNA-derived reads in RNA sequencing libraries and increase the probability of capturing other RNAs.

[0147] Example 3. Constructing RNA sequencing libraries with low abundance rRNA using total RNA of different qualities.

[0148] In this embodiment, total RNA (1 μg / μL, Takara, 634485) derived from human brain cells was diluted with nuclease-free water at different ratios to obtain six concentrations of RNA. Specifically, 1 μg / μL was diluted three times at a ratio of 1:10 to 1 ng / μL; 1 ng / μL was diluted at a ratio of 1:2 to 500 pg / μL; 1 ng / μL was diluted at a ratio of 1:5 to 200 pg / μL; 1 ng / μL was diluted at a ratio of 1:10 to 100 pg / μL; 500 pg / μL was diluted at a ratio of 1:10 to 50 pg / μL; and 200 pg / μL was diluted at a ratio of 1:10 to 20 pg / μL.

[0149] Take 1 μL of RNA at each of the six concentrations mentioned above to obtain six RNA samples with different mass gradients (1 ng, 500 pg, 200 pg, 100 pg, 50 pg, and 20 pg, respectively). Figure 5 and Figure 6 (Samples 11-16 in the sample).

[0150] The purified cfRNA from step 2 of Example 1 was replaced with the RNA samples of the above 6 different quality gradients, and human brain cell whole RNA sequencing libraries of RNA samples of different quality gradients were obtained by following the operation of step 2 of Example 1.

[0151] Using the same method as step 1 of Example 2, RNA sequencing libraries with low abundance rRNA were constructed from human brain cell whole RNA sequencing libraries (hereinafter referred to as 6 whole RNA sequencing libraries) of RNA samples with different quality gradients, and the quality of the obtained low abundance rRNA RNA sequencing libraries was identified using the same method as step 2 of Example 2.

[0152] The results showed that the RNA sample from sample 11 (1 ng) Figure 6 The proportion of ribosomal RNA (rRNA) in the RNA sequencing libraries with low abundance rRNA in samples 1000pg, 12 (500pg RNA sample), 13 (200pg RNA sample), 14 (100pg RNA sample), 15 (50pg RNA sample), and 16 (20pg RNA sample) was less than 30%.

[0153] Therefore, the method of this invention can be used to construct low-abundance rRNA RNA sequencing libraries from trace amounts of RNA greater than or equal to 20 pg. The rRNA content in all 6 RNA sequencing library samples was less than 30%. Figure 5 and Figure 6 (Samples 11-16).

[0154] The present invention has been described in detail above. Those skilled in the art will recognize that the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. While specific embodiments have been provided, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein.

Claims

1. A method of constructing a sequencing library of RNA, characterized in that: The method comprises the following steps: A1) reverse transcribing RNA from a biological sample to obtain a DNA-RNA hybrid chain; A2) cleaving the DNA-RNA hybrid chain to obtain a DNA-RNA hybrid chain fragment, and connecting a linker M1 and a linker M2 to both ends of the DNA-RNA hybrid chain fragment respectively to obtain a DNA-RNA hybrid chain fragment with added linkers; A3) using the DNA-RNA hybrid chain fragment with added linkers as a template, performing PCR amplification using sequencing primer 1 and sequencing primer 2 for a sequencing platform to obtain a PCR product, and purifying the PCR product to obtain an RNA sequencing library; The sequencing primer 1 comprises a sequencing platform linker 1 and a linker N1; and the sequencing primer 2 comprises a sequencing platform linker 2 sequence and a linker N2. The linker N1 is a single-stranded nucleotide fragment selected from the linker M1; and the linker N2 is a single-stranded nucleotide fragment selected from the linker M2. A4) performing gene editing on the RNA sequencing library using a gRNA targeting a rRNA-derived PCR product and an RNA-guided nuclease to obtain a target DNA-disrupted RNA sequencing library, performing PCR amplification on the target DNA-disrupted RNA sequencing library using the sequencing platform linker 1 and the sequencing platform linker 2 to obtain a PCR product 2, and purifying the PCR product 2 to obtain a purified RNA sequencing library.

2. The method of claim 1, wherein: The RNA in A1) has a mass of greater than or equal to 20 pg.

3. The method of claim 2, wherein: The RNA in A1) has a mass of greater than or equal to 20 pg and less than or equal to 10 ng.

4. The method of any one of claims 1-3, wherein: The biological sample is an animal body fluid.

5. A reagent for constructing a sequencing library of RNA, characterized by: The reagents comprise a reverse transcriptase, the linker M1, the linker M2, the sequencing primer 1, the sequencing primer 2, the sequencing platform linker 1, the sequencing platform linker 2, a gRNA targeting a rRNA-derived PCR product, and an RNA-guided nuclease.

6. The agent of claim 5, wherein: The reagents further comprise a transposase complex and a DNA polymerase.

7. The agent of claim 6, wherein: The transposase complex is a transposase Tn5 complex.

8. Use of the reagents in any one of claims 5-7 in constructing an RNA sequencing library.

Citation Information

Patent Citations

  • Sequencing library construction method and kit for pathogenic microorganism detection

    CN111188094A

  • Transposome enabled dna / rna-sequencing (ted rna-seq)

    CN112689673A