Preparation method of strand-specific library for rapidly detecting multiple types of RNA (Ribonucleic Acid) and high-throughput sequencing technology
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-25
- Publication Date
- 2026-03-27
AI Technical Summary
The prior art cannot detect multiple types of RNA quickly and effectively, especially in the case of low starting amounts or degraded samples. The target RNA type of commercial sequencing library preparation kit is single, the experimental steps are complex and the cycle is long, so it cannot meet the high Systematic research requirements for multiple RNA types throughput sequencing.
Reverse transcription of multiple RNA types is achieved by adding polyadenylate (polyA) tail to the end of the RNA sample and combining multiple reactions with polydeoxythymidine ribonucleotide (oligo dT) primers. Steps: Use DNA ligase to replace RNA ligase, simplify the library preparation process, improve efficiency and sequencing quality.
It realizes fast, high sensitivity and strong anti-interference ability of strand-specific library preparation for various types of RNA, shortens library preparation time, improves sequencing data volume and quality, and is suitable for high-throughput sequencing technology and low-rise The starting RNA sample is suitable for histocellular and bodily fluid samples.
Smart Images

Figure CN121752736A_ABST
Abstract
Description
A strand-specific library preparation method and high-throughput sequencing technology for rapid detection of multiple types of RNA Technical Field
[0001] The present invention relates to the field of biotechnology, and in particular to a method for preparing a strand-specific library for rapidly detecting multiple types of RNA and a high-throughput sequencing technology. Background Art
[0002] With the advancement of fundamental disciplines such as molecular biology and genomics, liquid biopsy has become a cutting-edge focus in the field of precision medicine. Cell-free RNA (cfRNA) is present in various bodily fluids, such as blood, urine, cerebrospinal fluid, saliva, breast milk, pleural effusion, and ascites, and can monitor the body's physiological status. Intracellular RNA types include ribosomal RNA (rRNA), messenger RNA (mRNA), transfer RNA (tRNA), micronon-coding RNA (miRNA), and long non-coding RNA (lncRNA). Cell-free RNA in bodily fluid samples carries real-time gene expression information from various tissues and organs, providing solutions for early disease screening, auxiliary diagnosis, recurrence monitoring, and prognosis. Due to its non-invasive or minimally invasive sampling method and the advantages of real-time and comprehensive monitoring, cell-free RNA has the potential to become a key diagnostic tool in future precision medicine.
[0003] In order to fully explore the biomarker potential of free RNA, efficient, rapid, low-cost and stable free RNA detection technology is needed. RNA detection technology based on high-throughput sequencing has the advantages of single-base resolution, wide detection range and high throughput. However, there are still many challenges in the preparation of high-throughput sequencing libraries for free RNA, including but not limited to low sample starting amount, easy RNA degradation, long experimental cycle, low efficiency, single or limited target RNA type and other problems. Currently, there are no commercial sequencing library preparation kits specifically for free RNA on the market. The scientific research field mainly uses kits suitable for the preparation of sequencing libraries for RNA from conventional cell tissue samples to study free RNA. The commercial RNA sequencing library preparation kits on the market have relatively single or limited target RNA types, and the information that can be obtained is also limited, making it impossible to conduct systematic research on all types of RNA.
[0004] Currently, there are no commercial library preparation kits on the market that can comprehensively capture multiple types of RNA and specifically target free RNA. The alternative solution is to use conventional RNA sequencing library preparation kits to construct sequencing libraries for free RNA.
[0005] Conventional RNA sequencing library preparation kits are mainly divided into kits suitable for mRNA and miRNA, which cannot capture multiple types of RNA in total RNA at the same time. Various types of RNA include mRNA, lncRNA, tRNA, miRNA, etc. Most commercial RNA sequencing library preparation kits have high requirements for the starting amount of total RNA. In addition, commercial RNA sequencing libraries have more experimental steps, more complicated processes, and longer experimental cycles, which are not conducive to rapid and stable detection in the clinical field. Currently, commercial sequencing library preparation kits suitable for mRNA detection include Illumina's TruSeq RNA Library Prep Kit v2, TruSeq Stranded mRNA and Total RNA Library prep kits, NEB's NEBNext Ultra TM II RNA Library Prep Kit for Illumina, TAKARA's SMARTer Stranded Total RNA-Seq Kit, etc. The scheme of the above kits uses polydeoxythymidine ribonucleotide (oligo dT) magnetic beads or primers to selectively enrich mRNA containing polyadenylic acid (polyA) tails to specifically prepare mRNA sequencing libraries. The above kits can also capture lncRNA while capturing mRNA by adapting the rRNA removal step and then using primers containing multiple random bases for reverse transcription, but this method is difficult to capture a large number of short fragments of miRNA and other RNAs. The above kits all need to first reverse transcribe into single-stranded cDNA, and then synthesize complementary double-stranded DNA to form double-stranded DNA, and then prepare it as a sequencing library on this basis. Therefore, the experimental steps are relatively complicated and the experimental cycle is relatively long.
[0006] Currently, there are numerous commercially available kits suitable for miRNA or small RNA. Examples include Illumina's TruSeq Small RNA Library Prep Kit, Bioo Scientific's Nextflex Small RNA-Seq Kit v3, NEB's NEBNext Small RNA Library Prep Set for Illumina, MGI's MGIEasy Small RNA Library Preparation Kit, and QIAGEN's QIAseq miRNA Library Kit. These kits all utilize RNA ligase to directly ligate linkers containing specific sequences to the 3' and 5' ends of RNA molecules, transcribe them into cDNA, and then amplify them through PCR to prepare sequencing libraries. These kits are not suitable for capturing long RNAs such as mRNA and lncRNA, and the experiments require RNA ligase, which is relatively costly, inefficient, time-consuming, and requires multiple steps, resulting in lengthy experimental cycles.
[0007] There is a commercial kit for small RNA, the SMARTer smRNA-Seq Kit for Illumina, which has a rather special principle. First, polyadenylic acid tailing is performed on the 3' end of the small RNA, and then reverse transcription is performed using a primer containing an adapter sequence and polydeoxythymidine ribonucleotide (oligo dT). The template switch method of reverse transcription is used to introduce the adapter sequence at the other end at the end of the cDNA. Finally, PCR is performed using the adapter sequences added at both ends to complete the preparation of the sequencing library. However, the above method will introduce artificially added polydeoxythymidine ribonucleotide sequences, which reduces the complexity of the library sequence and seriously affects the sequencing quality. In order to reduce the impact of low-complexity sequences on sequencing, it is necessary to add a base-balanced library during sequencing, but this solution will result in a reduction in the amount of effective data.
[0008] Commonly used RNA library preparation methods are not suitable for low-input or degraded samples, and have low adaptability to plasma samples. Scientific research plans generally require mL-level plasma inputs, and this higher plasma input reduces its wide application.
[0009] Summary of the Invention
[0010] In view of this, the technical problem to be solved by the present invention is to provide a method for preparing chain-specific libraries and a high-throughput sequencing technology for rapidly detecting multiple types of RNA. The present invention provides a method for preparing chain-specific libraries and a high-throughput sequencing technology for rapidly detecting multiple types of RNA, which have high sensitivity, a wide range of applications, strong anti-interference ability, a simple and quick preparation method, and are suitable for high-throughput sequencing technology.
[0011] The present invention provides a method for preparing various types of RNA libraries, comprising: adding polyA to the end of an RNA sample, performing reverse transcription and U base digestion to prepare a cDNA library;
[0012] The RNA sample contains at least one of total RNA, rRNA, mRNA, tRNA, miRNA and / or lncRNA.
[0013] The present invention adds polyadenylic acid (poly A) tails to the ends of various types of RNA molecules and combines them with a reverse transcription method using polydeoxythymidine ribonucleotide (oligo dT) primers. This allows various types of RNA molecules, including non-coding RNA (lncRNA, miRNA, etc.), to be reverse transcribed in subsequent steps, thereby enabling the construction of multiple types of RNA libraries.
[0014] Furthermore, the preparation method specifically includes: DNA digestion, RNA end modification, poly A tailing, reverse transcription, U base digestion, cDNA end modification, denaturation, linker addition and library construction.
[0015] The present invention combines two or more of the above-mentioned specific reaction steps by optimizing reaction reagents and other conditions. For example, in some embodiments, the steps of DNA digestion, RNA end modification, and polyA tailing are combined. In some embodiments, the steps of DNA digestion and RNA end modification are combined, and then the steps of polyA tailing and reverse transcription are combined. In some embodiments, the steps of U base digestion, cDNA end modification, and denaturation are combined. In some embodiments, the steps of U base digestion and cDNA end modification are combined. By combining reaction steps, the number of reaction steps is reduced and the required time is shortened, while ensuring reaction efficiency and the quality of the resulting library.
[0016] In some embodiments, the DNA digestion and RNA end modification are performed in a first system;
[0017] The reverse transcription is performed in a second system;
[0018] The step of adding the polyA tail is performed in the first system or the second system.
[0019] Furthermore, in a specific embodiment,
[0020] The step of adding the polyA tail is carried out in the first system, wherein:
[0021] The first system includes: RNA sample, PolyA polymerase reaction buffer, ATP, BSA, DNase I, T4 polynucleotide kinase, PolyA polymerase, RNase inhibitor and nuclease-free water;
[0022] The second system includes: the reaction product of the first system, reverse transcription primer, HiScript III reaction buffer, HiScript III reverse transcriptase, dNTP Mix, RNase inhibitor and nuclease-free water.
[0023] To ensure that RNA end modification, PloyA tailing, and reverse transcription can proceed smoothly in the same system, the present invention optimizes the reaction system. The components in the reaction system are properly coordinated to ensure the reaction proceeds. Furthermore, the present invention optimizes the concentrations of the components in each system.
[0024] The concentrations of the components of the first system are as follows: 14 μL RNA sample, 3 μL 10× Poly A polymerase reaction buffer, 1 μL 10 mM ATP, 4 μL 10 mg / mL BSA, 2 μL DNase I, 0.5 μL 10 U / μL T4 polynucleotide kinase, 1 μL 5 U / μL Poly A polymerase, 0.5 μL 40 U / μL RNase inhibitor, and 4 μL nuclease-free water;
[0025] The concentrations of the components of the second system include: 30 μL of the reaction product of the first system, 2 μL of a 5 μM reverse transcription primer, 4 μL of a 5×HiScript III reaction buffer, 1 μL of a 200 U / μl HiScript III reverse transcriptase, 2 μL of a 5 mM dNTP Mix, 0.5 μL of a 40 U / μl RNase inhibitor, and 10.5 μL of nuclease-free water.
[0026] Furthermore, in another specific embodiment,
[0027] The step of adding the polyA tail is carried out in the second system, wherein:
[0028] The first system includes: RNA sample, DNase I reaction buffer, ATP, DNase I, T4 polynucleotide kinase and RNase inhibitor;
[0029] The second system includes: the reaction product of the first system, HiScript III reaction buffer, BSA, PEG8000, dNTP Mix, reverse transcription primer, Poly A polymerase, RNase inhibitor, HiScript III reverse transcriptase and nuclease-free water.
[0030] To ensure smooth DNA digestion and RNA end modification reactions in a single system, and smooth PloyA tailing and reverse transcription reactions in the same system, the present invention optimizes the reaction system. The components in the reaction system are properly coordinated to ensure the reaction proceeds. Furthermore, the present invention optimizes the concentrations of the components in each system.
[0031] The concentrations of the components of the first system are as follows: 14 μL RNA sample, 2 μL 10× DNase I reaction buffer, 1 μL 10 mM ATP, 2 μL DNase I, 0.5 μL 10 U / μL T4 polynucleotide kinase, and 0.5 μL 40 U / μL RNase inhibitor.
[0032] The concentrations of the components of the second system are: 20 μL of the reaction product of the first system, 6 μL of 5×HiScript III reaction buffer, 1 μL of 10 mg / mL BSA, 10 μL of 50% PEG8000, 2 μL of 5 mM dNTP Mix, 2 μL of 5 μM reverse transcription primer, 1 μL of 5 U / μL PolyA polymerase, 0.5 μL of 40 U / μl RNase inhibitor, 1 μL of 200 U / μl HiScript III reverse transcriptase, and 6.5 μL of nuclease-free water.
[0033] In some embodiments, the cDNA end modification and denaturation are performed in a third system.
[0034] The third system comprises: the reaction product after U base digestion, polynucleotide kinase reaction buffer, T4 polynucleotide kinase, Tris buffer with pH 8.0, ultra-thermostable single-stranded binding protein and nuclease-free water.
[0035] To ensure smooth cDNA end modification and denaturation reactions in the same system, the present invention optimizes the reaction system. The components in the reaction system are properly coordinated to ensure the reaction proceeds. Furthermore, the present invention optimizes the concentrations of the components in each system.
[0036] The concentrations of the components of the third system are: the reaction product after U base digestion, 5 μL of 10× polynucleotide kinase reaction buffer, 1 μL of 10 U / μL T4 polynucleotide kinase, 2 μL of 220 mM Tris buffer (pH 8.0), 0.6 μL of 500 ng / μL ultra-thermostable single-stranded binding protein, and 21.4 μL of nuclease-free water.
[0037] In other embodiments, cDNA end modification, denaturation and U-base digestion are performed in the same system. Therefore, the third system also includes reagents for U-base digestion.
[0038] In this embodiment, the third system includes: the reaction product of polyA tailing and reverse transcription, polynucleotide kinase reaction buffer, T4 polynucleotide kinase, Tris buffer at pH 8.0, ultra-thermostable single-stranded binding protein, uracil-specific excision reagent USER enzyme and nuclease-free water.
[0039] To ensure smooth cDNA end modification and denaturation reactions in the same system, the present invention optimizes the reaction system. The components in the reaction system are properly coordinated to ensure the reaction proceeds. Furthermore, the present invention optimizes the concentrations of the components in each system.
[0040] The concentrations of the components of the third system are: the reaction product after polyA tailing and reverse transcription, 5 μL of 10× polynucleotide kinase reaction buffer, 1 μL of 10 U / μL T4 polynucleotide kinase, 2 μL of 220 mM Tris buffer, pH 8.0, 0.6 μL of 500 ng / μL ultra-thermostable single-stranded binding protein, 2 μL of 10 U / μL uracil-specific excision reagent USER enzyme, and 19.4 μL of nuclease-free water.
[0041] Compared with the existing technology, the present invention adopts a flexible method of combining multiple reaction steps into the same experimental operation step and optimizing the experimental conditions, thereby reducing the operation steps and time required for library preparation and ensuring the efficiency of the reaction, thereby achieving better experimental results.
[0042] Furthermore, the reverse transcription primer has the following nucleotide sequence: poly(T)n-UVNm;
[0043] Where n represents the number of bases T, and m represents the number of bases N;
[0044] n is an integer of 8 to 50, and m is an integer of 1 to 4
[0045] At least one T in the poly(T)n is replaced by U; V is selected from any one of base A, base C and base G; and N is selected from any one of base A, base T, base C and base G.
[0046] The present invention provides a polydeoxythymidine ribonucleotide (oligo dT) reverse transcription primer containing deoxyuracil (dU) for synthesizing single-stranded cDNA molecules through a reverse transcription reaction. The primer sequence comprises deoxyuracil ribonucleotide (dU), polydeoxythymidine ribonucleotide (oligo dT), and other random bases (V and N). Utilizing the reverse transcription primer provided by the present invention, a single-stranded cDNA molecule containing a polydeoxythymidine ribonucleotide sequence containing deoxyuracil (dU) can be obtained, so that a uracil-specific excision reagent can be used to perform a digestion reaction on the single-stranded cDNA in a subsequent U base digestion step, thereby excising the polydeoxythymidine ribonucleotide sequence fragment containing deoxyuracil (dU) in the single-stranded cDNA, eliminating the influence of low-complexity sequences artificially introduced in the previous step on subsequent sequencing and analysis, and better achieving cDNA library preparation. Compared with the existing technology that artificially introduces low-complexity sequences to capture miRNA or single-stranded DNA and then adds a base-balanced library during sequencing (such as the SMARTer smRNA-Seq Kit for Illumina), the solution provided by the present invention is more conducive to increasing the amount of effective data in sequencing.
[0047] In some specific embodiments, the reverse transcription primer has a nucleotide sequence as shown in SEQ ID NO: 1. Experiments have shown that, compared with other reverse transcription primers, for example, primers with N of 2, 3, or 4, the reverse transcription primer shown in SEQ ID NO: 1 can achieve a better reverse transcription reaction, thereby obtaining better experimental results.
[0048] Furthermore, in the linker adding step, the linker includes a 5' end linker and a 3' end linker, the 5' end linker sequence has the nucleotide sequence shown as SEQ ID NO 2 and SEQ ID NO 3, and the 3' end linker sequence has the nucleotide sequence shown as SEQ ID NO 4 and SEQ ID NO 5.
[0049] The present invention provides a connector for introducing a specific target sequence at the 3' and 5' ends of a single-stranded cDNA molecule and a corresponding DNA connection scheme. Both connectors contain a double-stranded region and at least one protruding single-stranded region, wherein the double-stranded region contains a universal structural sequence for PCR or sequencing, and the protruding single-stranded region contains one or more (1 to 10) random base sequences for complementary pairing with the end of the single-stranded DNA molecule, so that the connector and the single-stranded DNA can be splinted. Each connector is formed by the interaction of two polynucleotide sequences to form its special structure. The two sequences contain complementary pairing regions, which will be complementary paired to form a specific double-stranded structure and a protruding single-stranded region containing random bases after solution mixing and static treatment. After the single-stranded cDNA molecule is added with connectors at both ends, PCR amplification can be directly performed to obtain the final library, thereby better realizing the preparation of cDNA library. Compared with the existing solutions that require step-by-step ligation of two adapter sequences at the 3' and 5' ends of the DNA molecule (such as SPLAT (Splinted ligation adapter tagging)), the present invention can simultaneously connect the DNA adapters at the 3' and 5' ends in a single step, shortening the steps and time while ensuring the connection effect.
[0050] In the present invention, the reaction system for adding a linker comprises: a 5' end linker solution, a 3' end linker solution, a ligation reaction buffer, a ligation reaction enhancer and T4 DNA ligase.
[0051] In the linker-adding reaction system described in the present invention, DNA ligase is used instead of RNA ligase, significantly shortening the ligation time. (In the scheme of adding linkers to both ends of RNA using RNA ligase, due to the relatively low ligation efficiency of RNA ligase, the single-end linker ligation reaction time generally takes 1 to 2 hours. The experimental steps for ligating linkers at both ends are also relatively complex, and the total time for linker ligation at both ends is 2 to 3 hours.) To enable the DNA ligase to work better, the present invention optimizes other components in the reaction system and their concentrations, thereby further improving the linker ligation efficiency.
[0052] In some embodiments, the reaction system for adding adapters includes: a 5' end adapter solution, a 3' end adapter solution, a ligation reaction buffer, a hexaamminecobalt chloride solution, and T4 DNA ligase.
[0053] The present invention also provides a sequencing method for various types of RNA libraries, wherein the cDNA library prepared by the above preparation method is used as a sample for on-machine sequencing.
[0054] Furthermore, the cDNA library is amplified by PCR, purified, sample mixed, and single-stranded circularized before sequencing.
[0055] In some specific embodiments, the upstream primer for PCR amplification has a nucleotide sequence as shown in SEQ ID NO 6, and the downstream primer for PCR amplification has a nucleotide sequence as shown in SEQ ID NO 7.
[0056] The downstream primers can contain a tag (Barcode) sequence for sample identification. By using this primer sequence to amplify library samples, library samples with different Barcode sequence tags can be obtained. At the same time, the Barcode sequence of the primer band can be matched with the Barcode of the universal sequencing library of the sequencing platform. The structure of the PCR amplification product can be consistent with the universal library structure of the sequencing platform, so that the source of the sample library can be accurately separated by the Barcode.
[0057] The present invention provides library construction reagents, including reagent I, reagent II and reagent III;
[0058] The reagent I comprises: DNase I reaction buffer, ATP, DNase I, T4 polynucleotide kinase, and RNase inhibitor;
[0059] The reagent II includes: reverse transcription primer, HiScript III reaction buffer, HiScript III reverse transcriptase, dNTP Mix, and RNase inhibitor;
[0060] The reagent III includes: polynucleotide kinase reaction buffer, T4 polynucleotide kinase, Tris buffer with pH 8.0, and ultra-thermostable single-stranded binding protein.
[0061] In some embodiments, the reagent I further comprises: PolyA polymerase reaction buffer, BSA, and PolyA polymerase;
[0062] In some embodiments, the reagent II further includes: BSA, PEG8000, and PolyA polymerase.
[0063] In some embodiments, the reagent III further includes a U base excision reagent, and the U base excision reagent includes a uracil-specific excision reagent USER enzyme.
[0064] Furthermore, the method further comprises adding a linker reagent, wherein the linker adding reagent comprises: Tris-HCl buffer, sodium chloride, EDTA, a linker as shown in SEQ ID NO 2 to 5, a ligation reaction buffer, T4 ligase and hexaamminecobalt chloride.
[0065] Furthermore, the method further comprises a purification reagent, wherein the purification reagent comprises Agencourt AMPure XP magnetic beads.
[0066] Furthermore, it also includes PCR amplification reagents, which include PCR enzyme reaction solution, an upstream primer of the nucleotide sequence shown in SEQ ID NO 6, and a downstream primer of the nucleotide sequence shown in SEQ ID NO 7.
[0067] Furthermore, RNA extraction reagents are also included.
[0068] The present invention also provides the use of the construction reagent in preparing various types of RNA libraries.
[0069] The present invention provides a method for preparing strand-specific libraries and a high-throughput sequencing technology for rapid detection of multiple types of RNA. This method utilizes polyadenylic acid polymerase to artificially add polyadenylic acid (poly A) tails to the 3' ends of multiple RNAs. Simultaneously, polydeoxythymidine ribonucleotide primers containing deoxyuracil are used to synthesize single-stranded cDNA under the action of reverse transcriptase. The resulting single-stranded cDNA molecules are then subjected to a series of reactions and ultimately amplified by PCR to obtain strand-specific libraries of multiple types of RNA. Compared to existing technologies:
[0070] (1) The present invention uses polyadenylic acid polymerase to artificially add polyadenylic acid (poly A) tails to the 3' ends of various RNAs, so that various types of RNA molecules, including non-coding RNAs (lncRNA, miRNA, etc.), can be reverse transcribed in subsequent steps, thereby enabling the construction and sequencing of various types of RNA libraries;
[0071] (2) Single-stranded cDNA is synthesized using polydeoxythymidine ribonucleotide primers containing deoxyuracil under the action of reverse transcriptase. The obtained single-stranded cDNA molecules can be cleared of artificially added polynucleotide sequences through U base digestion in subsequent reactions, avoiding the addition of low-complexity nucleotide sequences to the sequencing library, thereby ensuring sequencing quality and data volume;
[0072] (3) The present invention utilizes a single-stranded cDNA library preparation method, eliminating the need to synthesize complementary double-stranded DNA to the cDNA molecule, thereby reducing the steps and time required for library preparation. Furthermore, since the present invention utilizes a single-stranded cDNA library preparation method, the strand-specific information of the RNA is retained, making it more conducive to RNA analysis such as gene annotation.
[0073] (4) The present invention combines some reaction steps by optimizing the reaction reagents and other conditions. For example, in Example 3, the three reaction steps of DNA digestion, RNA end modification and polyadenylation tailing reaction are combined into the same operation step; the two reaction steps of cDNA end modification and denaturation reaction are combined into one operation step. In Example 4, the two reaction steps of DNA digestion and RNA end modification reaction are combined into one operation step; the two reaction steps of polyadenylation tailing and reverse transcription reaction are combined into one operation step; the three reaction steps of U base digestion, cDNA end modification and denaturation reaction are combined into the same operation step, etc. The overall solution reduces the number of operation steps and time required for library preparation and ensures the efficiency of the reaction;
[0074] (5) The present invention adopts a one-step ligation method to add linkers to both ends of the single-stranded DNA, which can simultaneously connect the DNA linkers at the 3' and 5' ends, reducing the number of experimental steps and time. The present invention adds a ligation reaction enhancer, such as the chemical reagent hexaaminocobalt chloride, to the ligation reaction, which can significantly improve the reaction efficiency of the ligation reaction and further shorten the ligation reaction time;
[0075] (6) The present invention utilizes reverse-transcribed cDNA and adapters for ligation, adding sequencing structure sequences. Since DNA ligase is used instead of RNA ligase, DNA ligase is superior to RNA ligase in both cost and efficiency. The total ligation time is only 30 minutes or even shorter, and the overall experimental process has fewer steps and requires less time. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 shows a schematic diagram of RNA library preparation and detection principles;
[0077] Figure 2 shows a schematic diagram of linker preparation, where P represents a phosphorylation modification group and B represents a modification group that blocks the connection;
[0078] FIG3 is a schematic diagram showing a method for preparing an RNA library;
[0079] FIG4 shows the sequencing quality distribution of RNA library sequencing along with the sequencing cycle number;
[0080] Figure 5 shows the number of various RNA genes detected in the samples, where the RNA samples in Figure 5a include mRNA, IncRNA, and pseudogene RNA, and the RNA samples in Figure 5b include miRNA, tRNA, mt-tRNA, mt-rRNA, snoRNA, and snRNA;
[0081] Figure 6 shows the number and percentage of various RNA genes detected in the samples, where Figure 6a is a 200 μL starting plasma free RNA sample, Figure 6b is a 10 ng starting amount of UHRR sample, and Figure 6c is a 2 ng starting amount of UHRR sample;
[0082] Figure 7 shows a scatter plot of gene expression and consistency analysis results between samples, where the Log2(TPM+1) value of gene expression was used for analysis. Figure 7a shows two technical parallels of a 200 μL starting plasma free RNA sample, Figure 7b shows UHRR samples with 10 ng and 100 ng starting amounts, and Figure 7c shows UHRR samples with 2 ng and 10 ng starting amounts;
[0083] FIG8 shows the Pearson correlation coefficient results of gene expression levels between samples, wherein FIG8a is the Pearson correlation coefficient calculated using the TPM value of the gene expression level, and FIG8b is the Pearson correlation coefficient calculated using the Log2(TPM+1) value after the gene expression level is converted. DETAILED DESCRIPTION
[0084] The present invention provides a method for preparing a strand-specific library and a high-throughput sequencing technology for rapid detection of multiple types of RNA. Those skilled in the art can refer to the contents of this article and appropriately improve the process parameters to achieve the desired results. It should be noted that all similar substitutions and modifications are obvious to those skilled in the art and are considered to be included in the present invention. The methods and applications of the present invention have been described through preferred embodiments, and relevant personnel can obviously modify or appropriately change and combine the methods and applications herein without departing from the content, spirit and scope of the present invention to implement and apply the technology of the present invention.
[0085] This invention provides strand-specific library preparation and high-throughput gene sequencing technologies for rapid detection of various types of RNA molecules. This technology can read multiple types of RNA information with a relatively low input amount of RNA and is suitable for a variety of scenarios, including tissue cells and body fluid samples. Through special reverse transcription primer design and digestion treatment, interference with sequencing caused by artificially added low-complexity sequences can be removed. Single-stranded library preparation technology can achieve strand-specific RNA library preparation. By introducing barcode sequences for sample identification and adapter structure sequences for sequencing reactions, high-throughput sequencing technology can be used to read and analyze multiple types of RNA information.
[0086] The present invention provides a chain-specific library preparation method and high-throughput sequencing technology for rapid detection of multiple types of RNA, which can simultaneously detect multiple RNA types including mRNA, lncRNA, tRNA, miRNA, etc. The method provided by the present invention is suitable for high-throughput sequencing detection of multiple sample types such as tissue cell RNA and cell-free RNA (abbreviated as cfRNA), and can be applied to low-starting RNA samples at the ng level or even the pg level. The present invention can be used in the fields of liquid biopsy, free RNA molecular diagnosis, RNA molecular detection of tissue cells, etc., and has broad application prospects in the fields of early disease screening, auxiliary diagnosis, recurrence monitoring, disease prognosis, etc. Among them, potential clinical application scenarios include but are not limited to complex diseases such as pregnancy diseases, tumors, cardiovascular and cerebrovascular diseases, infectious diseases, genetic diseases, neurological diseases, and mental illnesses. The present invention can be expanded to research in the fields of animals, plants, and microorganisms, and can detect various types of RNA at the same time; it can be used to study the growth and development characteristics, tissue and cell functions, gene functions, environmental adaptability, species interactions, etc. of different species, and it also has application prospects in the fields of agriculture, food, and biosafety. The technical principles of this invention can also be extended to single-cell or spatiotemporal omics technologies, enabling the simultaneous capture of different RNA types, including non-coding RNA, and characterization of the expression of multiple RNA types in single cells or at different spatial locations. Combining the principles of this invention with long-chain PCR and single-molecule sequencing can also potentially be applied to full-length transcriptome research.
[0087] First, an optional step involves DNA digestion of the RNA sample to remove any residual DNA molecules that might interfere with subsequent steps. This step is followed by the use of a terminal modification enzyme to modify the RNA molecule's ends, hydroxylating the 3' termini of the RNA molecule and allowing more natural RNA molecules to undergo polyadenylation tailing at the 3' termini. Polyadenylate polymerase is then used to complete polyadenylation tailing at the 3' termini of the RNA molecule. A polydeoxythymidine ribonucleotide (oligo dT) primer containing deoxyuracil (dU) is used to complement the RNA molecule, and single-stranded cDNA is synthesized using the RNA as a template under the action of a reverse transcriptase, yielding a single-stranded cDNA molecule containing deoxyuracil (dU) and polydeoxythymidine ribonucleotide (oligo dT) sequences. Next, the cDNA is digested with a uracil-specific excision reagent to remove polydeoxythymidine ribonucleotide sequence fragments containing deoxyuracil (dU) from the cDNA, removing the low-complexity sequences artificially introduced in the previous step that could affect subsequent sequencing and analysis. The cDNA library is then prepared.
[0088] The cDNA library preparation method provided by the present invention first denatures the cDNA at high temperature to unwind any double-stranded structures that may exist in localized regions of the cDNA. Single-strand binding proteins can be added to the denaturation reaction system to help maintain the linear structure of the single-stranded cDNA molecules, preventing annealing and renaturation into complex hairpin-like structures after denaturation. The high-temperature denatured cDNA is then treated with special double-stranded adapters and a supporting ligation technique to complete the cDNA library preparation.
[0089] This invention provides a single-stranded DNA library preparation technology. Using a special double-stranded adapter and one-step ligation technique, adapters can be added to both ends of a single-stranded cDNA molecule, introducing a universal structural sequence for PCR amplification. PCR amplification then generates a sequencing library. The sequencing library products are subsequently processed and sequenced according to the requirements of a gene sequencer. RNA detection is achieved through reading and analysis of the sequence information.
[0090] The schematic diagrams of the above library preparation are shown in Figures 1 and 3.
[0091] During the PCR amplification step of library preparation, PCR primers can be used to introduce a tag (barcode) sequence for sample identification. Library samples with different barcode sequences can be mixed into a single sequencing sample according to the required ratio of sequencing data. The sequencing sample is then processed and sequenced according to the requirements of the gene sequencer. Each sequencing result is precisely mapped to each sample after barcode sequence alignment and splitting. The sequencing results of each sample are then analyzed to achieve the goal of detecting RNA in each sample. Using barcode sequences allows multiple library samples to be mixed and sequenced together, improving detection throughput and reducing detection costs.
[0092] The present invention adds polyadenylic acid tails to the 3' ends of various RNA molecules. Combined with reverse transcription using a polydeoxythymidine (oligo dT) primer, this method enables efficient reverse transcription of various RNA types, thereby enabling the capture and construction of sequencing libraries for these diverse RNA types. These RNA types include, but are not limited to, mRNA, lncRNA, miRNA, and tRNA. The polyadenylic acid tail is poly(A)n (where n represents the number of adenines A and is an integer between 8 and 200).
[0093] The present invention designs a polydeoxythymidine ribonucleotide (oligo dT) primer containing deoxyuracil (dU) for synthesizing single-stranded cDNA molecules by reverse transcription reaction. This primer sequence is designated as the first nucleotide sequence. It should be noted that this primer sequence comprises deoxyuracil ribonucleotide (dU), polydeoxythymidine ribonucleotide (oligo dT), and other random bases (V and N), and its sequence has the nucleotide composition shown in Table 1 below. Using the primer sequence of the present invention for reverse transcription reaction, a cDNA molecule containing a polydeoxythymidine ribonucleotide sequence containing deoxyuracil (dU) can be obtained.
[0094] Table 1 Sequence structure of poly(deoxythymidine ribonucleotide) primers with deoxyuracil (5'-3' direction)
[0095] By optimizing the experimental conditions, the present invention can combine the multiple reaction steps mentioned in the above experimental principles into the same experimental operation step, reducing the operation steps and time required for library preparation and ensuring the efficiency of the reaction.
[0096] The preparation method of the library of the present invention specifically includes: DNA digestion, RNA end modification, poly A tail addition, reverse transcription, U base digestion, cDNA end modification, denaturation, linker addition and library construction.
[0097] In some embodiments of the present invention, the DNA digestion and RNA end modification are performed in a first system; the reverse transcription is performed in a second system; and the cDNA end modification and denaturation are performed in a third system.
[0098] Wherein, the step of adding the polyA tail is performed in the first system or the second system.
[0099] In some specific embodiments, the step of adding polyA tail is performed in a first system, wherein the first system comprises: 14 μL of RNA sample, 3 μL of 10× PolyA polymerase reaction buffer, 1 μL of 10 mM ATP, 4 μL of 10 mg / mL BSA, 2 μL of DNase I, 0.5 μL of 10 U / μL T4 polynucleotide kinase, 1 μL of 5 U / μL PolyA polymerase, 0.5 μL of 40 U / μL RNase inhibitor, and 4 μL of nuclease-free water;
[0100] The concentrations of the components of the second system were as follows: 30 μL of the reaction product of the first system, 2 μL of a 5 μM reverse transcription primer, 4 μL of a 5×HiScript III reaction buffer, 1 μL of a 200 U / μl HiScript III reverse transcriptase, 2 μL of a 5 mM dNTP Mix, 0.5 μL of a 40 U / μl RNase inhibitor, and 10.5 μL of nuclease-free water.
[0101] In some other specific embodiments, the poly A tailing step is performed in a second system, wherein the first system comprises: 14 μL of RNA sample, 2 μL of 10× DNase I reaction buffer, 1 μL of 10 mM ATP, 2 μL of DNase I, 0.5 μL of 10 U / μL T4 polynucleotide kinase, and 0.5 μL of 40 U / μL RNase inhibitor;
[0102] The second system consists of: 20 μL of the reaction product of the first system, 6 μL of 5×HiScript III reaction buffer, 1 μL of 10 mg / mL BSA, 10 μL of 50% PEG8000, 2 μL of 5 mM dNTP Mix, 2 μL of 5 μM reverse transcription primer, 1 μL of 5 U / μL PolyA polymerase, 0.5 μL of 40 U / μl RNase inhibitor, 1 μL of 200 U / μl HiScript III reverse transcriptase, and 6.5 μL of nuclease-free water.
[0103] In some specific embodiments, the third system comprises: the reaction product after U base digestion, 5 μL of 10× polynucleotide kinase reaction buffer, 1 μL of 10 U / μL T4 polynucleotide kinase, 2 μL of 220 mM Tris buffer (pH 8.0), 0.6 μL of 500 ng / μL ultra-thermostable single-stranded binding protein, and 21.4 μL of nuclease-free water.
[0104] In some other specific embodiments, the third system further includes a reagent for U base digestion. The third system comprises: the reaction product after poly A tailing and reverse transcription, 5 μL of 10× polynucleotide kinase reaction buffer, 1 μL of 10 U / μL T4 polynucleotide kinase, 2 μL of 220 mM Tris buffer (pH 8.0), 0.6 μL of 500 ng / μL ultrathermostable single-stranded binding protein, 2 μL of 10 U / μL uracil-specific excision reagent USER enzyme, and 19.4 μL of nuclease-free water.
[0105] In the above library preparation method, while ensuring the library construction effect, different reaction steps can be optionally combined into the same experimental operation step, thereby flexibly modifying the experimental steps and controlling the experimental time. The combination of reaction steps can be adjusted according to the specific experiment, and the present invention is not limited to this.
[0106] The present invention provides adapters for introducing specific target sequences at the 5' and 3' ends of single-stranded cDNA molecules, respectively, and a corresponding DNA ligation scheme. This scheme involves two adapters, a first adapter and a second adapter. Both adapters contain a double-stranded region and at least one protruding single-stranded region. The double-stranded region contains a universal structural sequence used for PCR or sequencing, and the protruding single-stranded region contains one or more (1-10) random base sequences for complementary pairing with the ends of the single-stranded DNA molecule, enabling the adapter and the single-stranded DNA to be splinted. The protruding single-stranded region of the first adapter is at the 5' end, and the protruding single-stranded region of the second adapter is at the 3' end. The first adapter is used for complementary pairing at the 5' end of the single-stranded DNA molecule and, the second adapter is used for complementary pairing and splinting at the 3' end of the single-stranded DNA molecule. One embodiment of the first adapter is that its special structure is formed by the interaction of two polynucleotide sequences. The two sequences contain complementary pairing regions. After solution mixing and static treatment, they complementarily pair to form a specific double-stranded structure and a protruding single-stranded region containing random bases. One implementation of the second linker is that its unique structure is formed by the interaction of two polynucleotide sequences. The two sequences contain complementary regions that, after solution mixing and static treatment, pair with each other to form a specific double-stranded structure with protruding single-stranded regions containing random bases. A schematic diagram of the preparation principles of the first and second linker solutions is shown in Figure 2.
[0107] The present invention provides a sequence scheme of a first linker and a second linker, which consists of four nucleotide sequences, and the sequences are shown in Table 2. The first linker is composed of two sequences, namely the second nucleotide sequence and the third nucleotide sequence in Table 2. The 5' end of the third nucleotide sequence contains a random base sequence, the number of random bases is 1-10, and the random base is base A, base T, base C or base G; the random base sequence can be complementary to the 5' end of the single-stranded cDNA molecule in the first linker structure, and is used for splint connection to improve the connection efficiency. The two sequences of the second linker are the fourth nucleotide sequence and the fifth nucleotide sequence in Table 2. The 3' end of the fifth nucleotide sequence contains a random base sequence, the number of random bases is 1-10, and the random base is base A, base T, base C or base G; the random base sequence can be complementary to the 3' end of the single-stranded cDNA molecule in the second linker structure, and is used for splint connection to improve the connection efficiency. The 5' end of the second nucleotide sequence needs to be specially modified to block it from undergoing a connection reaction. The 5' end of the fourth nucleotide sequence is phosphorylated, and the 3' end is specially modified to block ligation. Both the 5' and 3' ends of the third and fifth nucleotide sequences are specially modified to block ligation. With DNA ligase and appropriate reaction conditions, both the 5' and 3' ends of the single-stranded DNA molecule can be directly ligated to adapters, which can then be used in PCR amplification reactions using primers designed based on the adapter sequences.
[0108] The present invention uses a DNA ligation reaction enhancing reagent, for example, a chemical reagent hexaamminecobalt chloride of appropriate concentration, to improve the DNA ligation efficiency, realize the connection of the two end adapters of the 5' end and the 3' end of the single-stranded cDNA molecule in a one-step experimental operation, and shorten the reaction time of the ligation.
[0109] After adding adapters to the single-stranded cDNA molecules, PCR amplification can be performed directly to generate the final library. Through PCR primers, a complete sequencing structure sequence can be introduced, including barcode sequences that can distinguish samples, to adapt to the needs of the sequencing platform.
[0110] The present invention eliminates the need for a separate, independent step to synthesize the complementary double-stranded DNA of the cDNA molecule, thereby reducing the number of steps and time required for library preparation. Furthermore, because the present invention utilizes a single-stranded cDNA library preparation method, it preserves the strand-specific information of the RNA, further facilitating RNA analysis such as gene annotation.
[0111] Table 2 Specially modified double-stranded DNA linker sequences containing random bases (5'-3' direction)
[0112] The present invention provides a set of universal primer sequences for PCR amplification of library samples. The forward universal primer is the sixth nucleic acid sequence, and the reverse primer is the seventh nucleic acid sequence. The reverse primer may contain a tag (barcode) sequence for sample identification, designated as the seventh nucleic acid sequence-N, where N represents the barcode number. Different barcode sequences have different numbers. The universal primer sequence consists of the nucleotides listed in Table 3 below. Using this primer sequence to amplify library samples, library samples tagged with different barcode sequences can be obtained.
[0113] By using the principles of the present invention, the barcode sequence in the seventh nucleic acid sequence -N can be matched with the barcode of the universal sequencing library of the sequencing platform. The structure of the PCR amplification product can be consistent with the universal library structure of the sequencing platform, and the source of the sample library can be accurately separated by the barcode.
[0114] Table 3 Universal primer sequences (5'-3' direction)
[0115] The adapter and PCR primer sequence schemes provided in this invention (Tables 2 and 3) are suitable for preparing high-throughput sequencing libraries for the DNBSEQ and MGISEQ series of MGI sequencing platforms. However, the design principles provided in this invention can also be used to design sequence schemes suitable for other sequencing platforms to prepare compatible sequencing libraries.
[0116] The test materials used in the present invention are all common commercial products and can be purchased on the market. The present invention is further described below with reference to the following examples:
[0117] Example 1 Extraction and preparation of RNA samples
[0118] Universal Human Reference RNA (UHRR, Agilent, 740000) was extracted. One tube of commercially available standard (200 μg RNA, 70% ethanol, and 0.1 M sodium acetate solution) was centrifuged at 4°C, 12,000 × g for 15 minutes. The supernatant was discarded, and the pellet was washed with 70% ethanol. The pellet was centrifuged again at 4°C, 12,000 × g for 15 minutes, and the supernatant was carefully discarded. The pellet was dried at room temperature for 30 minutes and resolubilized in 1 mL of nuclease-free HO to a resolubilized RNA concentration of approximately 200 ng / μL. The exact concentration of the resolubilized UHRR solution was determined using a Qubit 3.0 fluorometer (Invitrogen, Q33216). According to the measured UHRR solution concentration, the standard solution was serially diluted with nuclease-free HO to prepare four samples with different RNA starting amounts: 100 ng, 10 ng, 2 ng, and 0.2 ng, with a total volume of 14 μL per tube.
[0119] A 400 μL human plasma sample was used and aliquoted into two 200 μL portions. Serum / plasma miRNA extraction and isolation kits (spin column, TIANGEN, DP503) were used for each extraction, strictly following the manufacturer's instructions.
[0120] Example 2 Preparation of Joint
[0121] 2.1. Preparation of 5× STE buffer
[0122] According to the reaction system in Table 4, prepare 5× STE buffer and shake thoroughly to mix.
[0123] Table 4 5×STE buffer
[0124] 2.2. Preparation of the linker
[0125] 5 × STE buffer was used to prepare the joint. According to the systems in Tables 5 and 6, the first joint solution and the second joint solution at a concentration of 25 μM were prepared respectively. The joint solution was fully shaken and mixed, and allowed to stand at room temperature for 30 minutes to allow the joint nucleotide sequence to fully complement each other. After the reaction was completed, 25 μM of the first joint solution and 25 μM of the second joint solution were diluted to 1 μM of the first joint solution and 1 μM of the second joint solution, respectively, and placed under -18 ° C to -22 ° C conditions for storage. When using the joint, the joint solution was placed at room temperature to melt, fully shaken and mixed, and then placed at 4 ° C for standby use. The schematic diagram of the preparation principle of the above-mentioned first joint and second joint solutions is shown in Figure 2.
[0126] Table 5 First linker (25 μM) solution configuration system
[0127] Table 6 Second linker (25 μM) solution preparation system
[0128] Example 3 RNA library preparation and high-throughput sequencing
[0129] The schematic diagram of library preparation is shown in Figures 1 and 3. The following is an explanation of the specific experimental steps.
[0130] 3.1 RNA sample description
[0131] According to the method described in Example 1, the starting amounts of 2 ng, 10 ng, and 100 ng of human universal RNA standard UHRR solution were prepared, and the samples were named UHRR-2 ng, UHRR-10 ng, and UHRR-100 ng.
[0132] 3.2 DNA digestion, RNA end modification and polyadenylation tailing reaction
[0133] DNA was digested using DNase I (RNase Free) (NEB, M0303L) to remove any residual DNA from the RNA preparation. RNA was end-modified using T4 polynucleotide kinase (T4 PNK, NEB, M0201L). Polyadenylic acid tails were added to the 3' end of the RNA using polyadenylate polymerase (NEB, E. coli Poly(A) Polymerase, M0276L). The reaction system is shown in Table 7. The reaction system was incubated at 37°C for 15 minutes, inactivated at 95°C for 5 minutes, and placed on ice for 2 minutes after the reaction.
[0134] Table 7 DNA digestion, end modification and polyadenylation tailing reaction
[0135] 3.3 Reverse transcription reaction
[0136] Add 2 μL of a 5 μM first nucleic acid sequence primer (i.e., reverse transcription primer) to the reaction product from the previous step. In this example, the first nucleotide sequence is 5'-TTTTTTTUTTTTTTTUVN-3' (SEQ ID NO: 1). After adding the primer solution, vortex to mix thoroughly, centrifuge briefly, denature at 65°C for 5 minutes, and incubate at 30°C for 1 minute.
[0137] Reverse transcription was performed using 200 U / μl HiScript III Reverse Transcriptase (Novozymes, R302-01). 18 μl of reverse transcription mixture (composition shown in Table 8) was added to the reaction system. The reaction system was placed in a PCR instrument (Bio-Tech, TC-96) and the program in Table 9 was run.
[0138] Table 8. Reverse transcription reaction mixture composition.
[0139] Table 9 Reverse transcription reaction procedure
[0140] 3.4 Purification of reverse transcription products
[0141] Add 100 μL of Agencourt AMPure XP magnetic beads (Beckman Coulter, A63881) to the reaction tube to purify the reverse transcription product. Elute with 22 μL of pH 8.0 TE buffer (AMBION, AM9849). Transfer 20 μL of the purified ligation product to a new PCR reaction tube and reserve it for library preparation.
[0142] 3.5 U base digestion reaction
[0143] Uracil-specific excision reagent (USER) enzyme (NEB, USER Enzyme, M5505L) was used for U base digestion, yielding a digested single-stranded cDNA product. 4 μL of the U base digestion mixture (composition shown in Table 10) was added to the reaction system. The reaction system was placed in a PCR instrument (Bio-Tech, TC-96) and the program in Table 11 was run.
[0144] Table 10. Composition of the U-base digestion reaction mixture.
[0145] Table 11 U base digestion reaction program
[0146] 3.6 cDNA end modification and denaturation reaction
[0147] The end modification and denaturation reaction mixture was added to the reaction system. Its composition is shown in Table 12. In the reaction system, T4 polynucleotide kinase (T4 PNK, NEB, M0201L) was used to phosphorylate the 5' end of the single-stranded cDNA molecules. The cDNA molecules were then denatured at high temperature to unwind the double-stranded structures that may appear in local regions of the cDNA. ET SSB (NEB, M2401S) was added to the denaturation reaction system to maintain the linear structure of the single-stranded cDNA molecules and prevent annealing and renaturation into complex hairpin-like structures after denaturation. The reaction system was placed in a PCR instrument (Bori, TC-96) and the program in Table 13 was run.
[0148] Table 12. Composition of the terminal modification and denaturation reaction mixture.
[0149] Table 13 Terminal modification and denaturation reaction procedures
[0150] 3.7 Ligation reaction
[0151] 2 μL of the first linker solution (1 μM) and 2 μL of the second linker solution (1 μM) prepared in Example 2 were added to the reaction system, followed by 28 μL of the ligation reaction mixture, the composition of which is shown in Table 14. T4 DNA Ligase (NEB, M0202L) was used for linker ligation, and the ligation reaction enhancer was a 0.5 mM hexaamminecobalt chloride solution. The reaction system was placed in a PCR instrument (Bori, TC-96) and the program in Table 15 was run.
[0152] The present invention uses a DNA ligation reaction enhancing reagent. In this embodiment, a chemical reagent hexaamminecobalt chloride of appropriate concentration is used to improve the DNA ligation efficiency, realize the connection of the two end adapters of the 5' end and the 3' end of the single-stranded cDNA molecule in a one-step experimental operation, and shorten the ligation reaction time.
[0153] Table 14 Ligation reaction mixture composition
[0154] Table 15 Ligation reaction program
[0155] 3.8 Purification of ligation product
[0156] Add 160 μL of Agencourt AMPure XP magnetic beads (Beckman Coulter, A63881) to the reaction tube and purify the ligation product. Elute with 23 μL of TE buffer, pH 8.0 (Invitrogen, AM9849). Transfer 21 μL of the purified ligation product to a new PCR reaction tube and reserve it for universal PCR amplification.
[0157] 3.9 Universal PCR Amplification
[0158] The primer of the sixth nucleic acid sequence (40 μM) was used as a forward primer, and the primer of the seventh nucleic acid sequence-N (40 μM) was used as a reverse primer for a universal PCR amplification reaction. The reverse primer contained a barcode sequence for sample identification. Each sample used a unique barcode sequence, and the different barcode sequences in the reverse primers of different samples were used to distinguish different samples in the offline data. The PCR enzyme reaction solution used in this embodiment was KAPA HiFi HotStart ReadyMix (2X) (Kapa Biosystems, KK2602), and the reaction composition is shown in Table 16. The reaction system was placed on a PCR instrument (Bori, TC-96) and the program in Table 17 was run.
[0159] Table 16 General PCR amplification reaction solution components
[0160] Table 17 General PCR amplification reaction program
[0161] 3.10 Universal PCR Amplification Product Purification and Pooling
[0162] 50 μL of PCR product obtained by universal PCR amplification was purified using 50 μL of Agencourt AMPure XP magnetic beads (Beckman Coulter, A63881). The concentration of 35 μL of purified DNA was measured using a Qubit3.0 fluorescence quantitative instrument (Invitrogen, Q33216). At the same time, the library samples were mixed into sequencing library samples at the same final concentration and shaken for use.
[0163] 3.11 Single-stranded circularization and sequencing reaction
[0164] Single-stranded DNA was circularized using the MGIEasy Circularization Module V2.0 (MGI, 1000005260) from Shenzhen MGI Intelligent Manufacturing Technology Co., Ltd., and sequenced using the MGISEQ-2000RS High-Throughput Sequencing Kit (FCL PE100) (MGI, 1000012554) from Shenzhen MGI Intelligent Manufacturing Technology Co., Ltd. All procedures were performed strictly according to the kit instructions. The PE100+10 (Paired End 100+10) sequencing method was used to obtain reliable base sequence information. Off-line data were split and filtered according to barcode sequence to obtain sequencing data for each sample.
[0165] Example 4 Rapid RNA library preparation and high-throughput sequencing
[0166] 4.1 RNA sample description
[0167] According to the method described in Example 1, the human universal RNA standard UHRR was prepared with starting amounts of 0.2 ng, 2 ng, and 10 ng, respectively. The samples were named UHRR-F-0.2 ng, UHRR-F-2 ng, and UHRR-F-10 ng.
[0168] According to the method described in Example 1, plasma free RNA samples were extracted and the samples were named Plasma-F-200uL-1 and Plasma-F-200uL-2.
[0169] 4.2 DNA digestion and RNA end modification reactions
[0170] DNA digestion was performed using DNase I (RNase Free) (NEB, M0303L) to remove any residual DNA from RNA preparation. RNA molecules were 3'-terminated using T4 polynucleotide kinase (T4 PNK, NEB, M0201L). The reaction system is shown in Table 18. The reaction system was incubated at 37°C for 15 minutes and inactivated at high temperature (95°C) for 5 minutes. After the reaction, the sample was placed on ice for 2 minutes.
[0171] Table 18 DNA digestion and RNA end modification reactions
[0172] 4.3 Polyadenylation and Reverse Transcription
[0173] Polyadenylic acid tailing was performed at the 3' end of the RNA molecule using polyadenylic acid polymerase (NEB, E. coli Poly (A) Polymerase, M0276L); reverse transcription was performed using a primer for the first nucleic acid sequence and 200 U / μl HiScript III Reverse transcriptase (Novozymes, R302-01). The first nucleic acid sequence in this embodiment is specifically 5'-TTTTTTTUTTTTTTTUVN-3'. 30 μL of polyadenylic acid tailing and reverse transcription reaction mixture was added to the reaction system, the composition of which is shown in Table 19. The reaction system was placed in a PCR instrument (Bori, TC-96) and the program in Table 20 was run.
[0174] Table 19. Composition of the poly(A) tailing and reverse transcription reaction mixture.
[0175] Table 20 Polyadenylation and reverse transcription reaction procedures
[0176] 4.4 Purification of reverse transcription products
[0177] Add 100 μL of Agencourt AMPure XP magnetic beads (Beckman Coulter, A63881) to the reaction tube to purify the reverse transcription product. Elute with 22 μL of TE buffer, pH 8.0 (Invitrogen, AM9849). Transfer 20 μL of the purified ligation product to a new PCR reaction tube and reserve it for library preparation.
[0178] 4.5 U base digestion, cDNA end modification and denaturation reaction
[0179] A U-base digestion reaction was performed using uracil-specific excision reagent (USER) enzyme (NEB, USER Enzyme, M5505L), yielding a digested single-stranded cDNA molecule. The 5' end of the digested single-stranded cDNA molecule was phosphorylated using T4 polynucleotide kinase (T4 PNK, NEB, M0201L). The cDNA molecule was then denatured at high temperature to unwind any double-stranded structures that may have appeared in localized regions of the cDNA. ET SSB (NEB, M2401S) was added to the denaturation reaction system to help maintain the linear structure of the single-stranded cDNA molecule and prevent annealing and renaturation into complex hairpin-like structures after denaturation. 30 μL of the U-base digestion reaction mixture was added to the reaction system. The composition of the mixture is shown in Table 21. The reaction system was placed in a PCR instrument (Bori, TC-96) and the program in Table 22 was run.
[0180] Table 21. Composition of the U-base digestion, end modification, and denaturation reaction mixture.
[0181] Table 22 U base digestion, end modification and denaturation reaction procedures
[0182] 4.6 Ligation reaction
[0183] 2 μL of the first linker solution (1 μM) and 2 μL of the second linker solution (1 μM) prepared in Example 2 were added to the reaction system, followed by 28 μL of the ligation reaction mixture, the composition of which is shown in Table 14. T4 DNA Ligase (NEB, M0202L) was used for linker ligation, and the ligation reaction enhancer was a 0.5 mM hexaamminecobalt chloride solution. The reaction system was placed in a PCR instrument (Bori, TC-96) and the program in Table 15 was run.
[0184] 4.7 Purification of ligation product
[0185] Add 80 μL of Agencourt AMPure XP magnetic beads (Beckman Coulter, A63881) to the reaction tube to purify the ligation product. Elute with 23 μL of TE buffer (Invitrogen, AM9849). Transfer 21 μL of the purified ligation product to a new PCR reaction tube and reserve it for universal PCR amplification.
[0186] 4.8 Follow steps 3.9, 3.10, and 3.11 to complete universal PCR amplification, universal PCR product purification and mixing, single-stranded circularization, and sequencing, respectively.
[0187] Example 5 RNA sequencing data analysis and result presentation
[0188] Sequencing quality analysis of the raw data from Examples 3 and 4 revealed that over 90% of the sequences had a quality of Q30 or higher. Analysis of the sequencing chip quality revealed no significant degradation throughout the sequencing cycle, maintaining a consistently high level of quality, as shown in Figure 4. These results demonstrate that the method of the present invention can generate high-quality sequencing data.
[0189] The original sequencing data is split and screened by the barcode sequence to obtain the sequencing data of each sample. The present invention can realize the simultaneous detection of multiple samples on the same sequencing chip.
[0190] The raw sequencing data (raw reads) of each sample were filtered for low-quality bases, short sequences were removed, adapters were removed, and microbial sequences were removed. Then, ribosomal RNA (rRNA) was removed. Using the human rRNA sequence as a reference, the rRNA sequence in the raw data was removed to obtain the filtered sequence (clean reads).
[0191] All filtered clean reads were aligned to the human genome (GRCh38) as a reference. The number of clean reads, total alignment rate, unique alignment rate, and proportion of alignments to exonic regions were calculated, along with the positive and negative strand information of successfully aligned clean reads to the reference gene. The strand-specificity of the library was then calculated. The results are shown in Table 23. For most UHRR samples, the total alignment rate reached 93%-97%, and the unique alignment rate reached 60%-73%. Alignment rates were lower for low-input UHRR samples. For example, UHRR-F-0.2ng had a total alignment rate of 43.63% and a unique alignment rate of 35.27%, while UHRR-F-2ng had a total alignment rate of 86.36% and a unique alignment rate of 61.15%. For a 200μL plasma cell-free RNA sample, the total alignment rate was approximately 35%, and the unique alignment rate was approximately 28%. The chain specificity ratios calculated for UHRR samples were all above 80%, with most reaching around 90% or above; the chain specificity ratios calculated for plasma samples were around 70%.
[0192] The total number of genes detected in the sample and the number of genes of different types of RNA were counted. The RNA types analyzed in this example include mRNA (messenger RNA), lncRNA (long non-coding RNA), pseudogene RNA, miRNA (microRNA), tRNA (transfer RNA), mt-tRNA (mitochondrial transfer RNA), mt-rRNA (mitochondrial ribosomal RNA), snoRNA (small neclear RNA), and snRNA (small cytoplasmic RNA). The results are shown in Figure 5. All UHRR samples, regardless of starting amount (0.2ng, 2ng, 10ng, or 100ng), were able to detect over 10,000 protein-coding genes; approximately 4,000 lncRNA genes; approximately 3,000 to 10,000 pseudogene RNAs; approximately 200-900 miRNA genes; approximately 200-300 tRNA genes; approximately 300-800 snoRNA genes and approximately 40-90 snRNA genes; and all 22 mitochondrial tRNAs and two rRNAs, namely mt-tRNA and mt-rRNA. For a starting plasma sample of 200μL, the number of genes detected in free RNA was relatively small compared to UHRR. However, nearly 10,000 protein-coding genes were detected, along with other types of RNA. Figure 6 of this example also shows the percentage of RNA genes detected for some samples.
[0193] The quantitative analysis of RNA expression levels used the TPM (Transcript per million) calculation method, based on the whole genome gtf (gene transfer format) file, to achieve qualitative and quantitative analysis of different types of RNA and obtain RNA expression profiles. The consistency of sample detection was analyzed using the RNA expression profile. As shown in the expression levels in Figure 7, both UHRR samples with different starting amounts and technical replicate samples of plasma-free RNA showed good consistency. The Pearson correlation coefficient was calculated for the expression levels between all samples according to the original TPM value and the converted Log2 (TPM+1) value, and the results are shown in Figure 8. According to the Pearson correlation coefficient calculated based on the original TPM value, the correlation coefficient between UHRR samples was at least above 0.85; the correlation coefficient of the technical replicate samples of plasma-free RNA was 0.99.
[0194] Quantitative analysis of RNA expression was performed according to Log2 Fluctuations in low-expression genes negatively impacted the results, resulting in lower correlation coefficients than those calculated using the original TPM values. The correlation coefficient for UHRR samples was at least 0.67, with most samples exceeding 0.70. The correlation coefficient for two technical replicates of plasma cell-free RNA was 0.83.
[0195] It is worth mentioning that the UHRR sample starting amount range implemented in the present invention is relatively wide, including 4 different starting amount ranges of 0.2ng, 2ng, 10ng, and 100ng. Among them, the lowest UHRR starting amount is as low as 0.2ng. The starting amount of plasma samples is also as low as 200μL. The above starting amount is much lower than the RNA starting amount used in most current commercial RNA kits and literature. In the case of the low starting amount, under the experimental conditions of two batches of two different library preparation processes (Example 3 and Example 4), a high correlation coefficient can still be shown, indicating that the method of the present invention has higher stability and low starting amount advantages.
[0196] Table 23 Basic information of sample sequencing data and alignment rate
[0197] In summary, the present invention provides a strand-specific library preparation method and high-throughput sequencing technology for rapid detection of multiple RNA types, including but not limited to mRNA, lncRNA, tRNA, and miRNA. This overcomes the problems of traditional RNA library preparation methods, which often capture a limited or limited number of RNA types, have complex experimental procedures, and are time-consuming.
[0198] The RNA sequencing library prepared by the present invention does not contain artificially added low-complexity sequences, thereby ensuring the sequencing quality and data volume.
[0199] The sequencing library prepared by the present invention contains strand-specific information of RNA, which is beneficial for RNA analysis such as gene annotation.
[0200] By optimizing reaction conditions, the present invention combines multiple reaction steps mentioned in the experimental principle into a single experimental procedure. For example, polyadenylation and reverse transcription reactions are combined into a single step, and the ligation of adapters at both ends of the cDNA is combined into a single step. This reduces the number of steps and time required for library preparation, making the experimental procedure simpler and faster.
[0201] Through experimental design, the present invention uses a DNA ligase scheme to replace the RNA ligase scheme in the miRNA library preparation method to add sequencing adapter sequences, thereby improving the efficiency of ligation, shortening the time required for ligation, and reducing costs.
[0202] The present invention adopts PCR amplification technology to amplify the library sequence, introduces the barcode sequence for library identification and the structural sequence required for cyclization reaction and sequencing through PCR technology, and can realize the mixing of multiple library samples and sequencing on the machine together, thereby improving the detection throughput and reducing the detection cost.
[0203] The method of the present invention is applicable to high-throughput sequencing detection of various sample types such as tissue cell RNA and free RNA, and can be applied to RNA samples with low starting amounts at the ng level or even the pg level.
[0204] The present invention can be used for free RNA detection in the field of liquid biopsies. Currently, there are no dedicated RNA library preparation kits for free RNA samples on the market, and this method has broad application prospects. The above are only preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications are also within the scope of protection of the present invention.
Claims
1. A method for preparing multiple types of RNA libraries, characterized in that: include: PolyA was added to the ends of RNA samples, and cDNA libraries were prepared by reverse transcription and U base digestion; The RNA sample contains at least one of total RNA, rRNA, mRNA, tRNA, miRNA and / or lncRNA.
2. The preparation method according to claim 1, characterized in that: The preparation method specifically includes: DNA digestion, RNA end modification, adding polyA tail, reverse transcription, U base digestion, cDNA end modification, denaturation, adding linkers and library construction.
3. The preparation method according to claim 2, characterized in that: The DNA digestion and RNA end modification are performed in a first system; The reverse transcription is performed in a second system; The step of adding the polyA tail is performed in the first system or the second system.
4. The preparation method according to claim 2 or 3, characterized in that: The cDNA terminal modification and denaturation are performed in a third system.
5. The preparation method according to claim 4, characterized in that: The third system also includes reagents for U base digestion.
6. The preparation method according to claim 3, characterized in that: The step of adding polyA tail is carried out in a first system, wherein: The first system comprises: RNA sample, PolyA polymerase reaction buffer, ATP, BSA, DNase I, T4 polynucleotide kinase, PolyA polymerase, RNase inhibitor and nuclease-free water; The second system includes: the reaction product of the first system, reverse transcription primer, HiScript III reaction buffer, HiScript III reverse transcriptase, dNTP Mix, RNase inhibitor and nuclease-free water.
7. The preparation method according to claim 3, characterized in that: The step of adding polyA tail is carried out in the second system, wherein: The first system comprises: RNA sample, DNase I reaction buffer, ATP, DNase I, T4 polynucleotide kinase and RNase inhibitor; The second system includes: the reaction product of the first system, HiScript III reaction buffer, BSA, PEG8000, dNTP Mix, reverse transcription primer, PolyA polymerase, RNase inhibitor, HiScript III reverse transcriptase and nuclease-free water.
8. The preparation method according to claim 6 or 7, characterized in that: The reverse transcription primer has the following nucleotide sequence: poly(T)n-UVNm; Where n represents the number of bases T, and m represents the number of bases N; n is an integer of 8 to 50, and m is an integer of 1 to 4 At least one T in the poly(T)n is replaced by U; V is selected from any one of base A, base C and base G; and N is selected from any one of base A, base T, base C and base G.
9. The preparation method according to claim 8, characterized in that: The reverse transcription primer has a nucleotide sequence as shown in SEQ ID NO:
1.
10. The preparation method according to claim 4, characterized in that: The third system comprises: the reaction product after U base digestion, polynucleotide kinase reaction buffer, T4 polynucleotide kinase, Tris buffer with pH 8.0, ultra-thermostable single-stranded binding protein and nuclease-free water.
11. The preparation method according to claim 5, characterized in that: The third system comprises: the reaction product of polyA tailing and reverse transcription, polynucleotide kinase reaction buffer, T4 polynucleotide kinase, Tris buffer of pH 8.0, ultra-thermostable single-strand binding protein, uracil specific excision reagent USER enzyme and nuclease-free water.
12. The preparation method according to claim 2, characterized in that: In the step of adding a linker, the linker includes a 5' end linker and a 3' end linker, the 5' end linker sequence has a nucleotide sequence as shown in SEQ ID NO 2 and SEQ ID NO 3, and the 3' end linker sequence has a nucleotide sequence as shown in SEQ ID NO 4 and SEQ ID NO 5.
13. A method for sequencing multiple types of RNA libraries, characterized in that: The cDNA library prepared by the preparation method according to any one of claims 1 to 12 is used as a sample for sequencing.
14. The sequencing method according to claim 13, characterized in that: The cDNA library is sequenced after PCR amplification, purification, sample mixing, and single-stranded circularization.
15. The sequencing method according to claim 14, characterized in that: The upstream primer for PCR amplification has a nucleotide sequence as shown in SEQ ID NO 6, and the downstream primer for PCR amplification has a nucleotide sequence as shown in SEQ ID NO 7.
16. A library construction reagent, characterized in that: It includes reagent I, reagent II and reagent III; The reagent I comprises: DNase I reaction buffer, ATP, DNase I, T4 polynucleotide kinase, and RNase inhibitor; The reagent II includes: reverse transcription primer, HiScript III reaction buffer, HiScript III reverse transcriptase, dNTP Mix, and RNase inhibitor; The reagent III comprises: polynucleotide kinase reaction buffer, T4 polynucleotide kinase, Tris buffer with pH 8.0, and super-thermostable single-stranded binding protein.
17. The construction reagent according to claim 16, characterized in that The reagent I also includes: PolyA polymerase reaction buffer, BSA, and PolyA polymerase.
18. The construction reagent according to claim 16, characterized in that The reagent II also includes: BSA, PEG8000, and PolyA polymerase.
19. The construction reagent according to claim 16, characterized in that The reagent III also includes a U base excision reagent, and the U base excision reagent includes a uracil-specific excision reagent USER enzyme.
20. The construction reagent according to claim 16, characterized in that The method also comprises a linker adding reagent, which comprises: Tris-Hcl buffer, sodium chloride, EDTA, a linker as shown in SEQ ID NO 2-5, a linker reaction buffer, T4 ligase and hexaamminecobalt chloride.
21. The construction reagent according to claim 16, characterized in that Also included is a purification reagent comprising Agencourt AMPure XP magnetic beads.
22. The construction reagent according to claim 16, characterized in that It also includes PCR amplification reagents, which include PCR enzyme reaction solution, an upstream primer of the nucleotide sequence shown in SEQ ID NO 6, and a downstream primer of the nucleotide sequence shown in SEQ ID NO 7.
23. The construction reagent according to any one of claims 16 to 22, characterized in that Also included are reagents for RNA extraction.
24. Use of the construction reagent according to any one of claims 16 to 23 in preparing multiple types of RNA libraries.