A method for constructing a single-cell whole transcriptome library and its application
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-02
- Publication Date
- 2026-08-14
AI Technical Summary
虽然VASA-seq提升了单细胞RNA测序技术的覆盖率和灵敏度,但是需要通过复杂的微流控操作对液滴进行多次精准的试剂添加,并且在片段化RNA后末端修复的过程中,会造成一定的转录片段损失
[0031]1)本发明的方法适用样本类型更广泛,不仅能够分析新鲜组织细胞,也能够分析新鲜细胞核、冰冻细胞、细胞核 、固定细胞、固定细胞核、石蜡包埋样本等;
Smart Images

Figure CN121951004B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of gene sequencing, and in particular relates to a method for constructing a single-cell whole transcriptome library and its application. Background Technology
[0002] Single-cell RNA sequencing, as a technology for characterizing intercellular heterogeneity, has made significant strides in obtaining more comprehensive cell maps, discovering new cell types, identifying cell type-specific genes, describing more refined gene regulatory networks, revealing potential gene functions, and reconstructing the trajectory of cell differentiation and disease development—achievements that bulk RNA sequencing cannot. Initially, single-cell RNA sequencing could only be used with a small number of individual cells. The introduction of droplet microfluidics increased the throughput, allowing a single experiment to study thousands to millions of cells. While current single-cell RNA sequencing technology offers a certain level of sensitivity and accuracy, providing higher resolution and more comprehensive information on single-cell gene expression and regulation, most techniques rely on hybridization between oligonucleotide polyT primers and polyA sequences of RNA to achieve RNA capture and complementary DNA synthesis. This means that current single-cell RNA sequencing technologies can only detect short fragments immediately adjacent to the polyA tail or 5' end; non-polyA transcribed fragments cannot be detected.
[0003] In recent years, Alexander van Oudenaarden et al. developed a single-cell whole transcriptome sequencing method called VASA-seq. This method fragments RNA, then performs end repair and poly-A tail hybridization to give each RNA fragment a poly-A tail, and finally uses poly-T primers to capture the RNA fragments and synthesize complementary DNA. Although VASA-seq improves the coverage and sensitivity of single-cell RNA sequencing technology, it requires complex microfluidic operations to precisely add reagents to the droplets multiple times, and the end repair process after fragmenting RNA causes some loss of transcription fragments.
[0004] Therefore, it is necessary to develop a method for constructing a single-cell whole transcriptome library and its application. Summary of the Invention
[0005] This invention provides a method for constructing a single-cell whole transcriptome library and its application, utilizing universal primers with specially modified bases and specifically encoded microsphere barcodes to achieve single-cell whole transcriptome sequencing. The universal primers described in this invention can capture all RNA in a single cell, improving the efficiency and sensitivity of capturing the whole transcriptome in a single cell.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] This invention provides a method for constructing a single-cell whole transcriptome library, the method comprising the following steps:
[0008] 1) Prepare single-cell or single-cell nucleus suspensions;
[0009] 2) Add universal primers and a reaction solution containing reverse transcriptase to a single cell or single cell nucleus suspension. The single cell or cell nucleus undergoes reverse transcription to synthesize cDNA in an independent reaction vessel. The universal primers also serve as templates to convert oligonucleotides to contact the 3' end of the cDNA and continue to synthesize cDNA through template conversion, resulting in a double-stranded molecule containing an RNA molecule and a cDNA product chain.
[0010] 3) Inactivate reverse transcriptase, add coding reaction reagents and coding microspheres, and use ribonuclease to digest RNA molecules, retaining only one cDNA product chain;
[0011] 4) Degrade the universal primers and release the single-cell coding primers on the coding microspheres, which then anneal and hybridize with the 3' end of the cDNA, and undergo a coding addition reaction under the action of DNA polymerase;
[0012] 5) Purify the cDNA product containing single-cell coding; add ddCTP to the ends of the primers using terminal transferase; extend the cDNA product containing single-cell coding using Read2 extension primers with specially modified bases to obtain the cDNA product with the complete Read2 sequencing primer sequence.
[0013] 6) After amplifying the obtained cDNA product using the Read1 and Read2 sequences, the cDNA product is then amplified again using the P5 primer sequence with index information and the P7 primer sequence with index information to obtain a high-throughput single-cell whole transcriptome library.
[0014] Optionally, in step 2), the independent reaction containers can be water-in-oil droplets, microfluidic chips, microwells, PCR tubes, 96-well plates, or 384-well plates.
[0015] Furthermore, when using water-in-oil droplets as the reaction vessel, the method of adding the coded reaction reagent and coded microspheres in step 3) is droplet fusion.
[0016] Furthermore, when using water-in-oil droplets as the reaction vessel, after obtaining the encoded cDNA product, the fused droplets are broken up using a demulsifier to release the cDNA product and other sequences from the droplets, and then purified using magnetic beads.
[0017] Further, in step 2), the universal primer is composed of a fixed sequence of three bases, G, A, and T, with a length of 10-30 nt; at least one thymine in the fixed sequence of the universal primer is replaced with deoxyuridine.
[0018] Further, in step 2), the reverse transcriptase includes Moloney mouse leukemia virus reverse transcriptase (MMLV) and its mutants, or TGIRT-III and its mutants.
[0019] Further, in step 2), the specific steps of the reverse transcription reaction and template conversion are as follows:
[0020] Reverse transcription reaction: After the universal primers are annealed and hybridized with RNA from a single cell or cell nucleus at a temperature of 8℃-35℃, cDNA is synthesized using reverse transcriptase at a temperature of 20℃-65℃.
[0021] Template conversion: After cDNA synthesis, the universal primers themselves are used as templates, and the template conversion activity of the reverse transcriptase is used to extend the cDNA to the 3' end to obtain a double-stranded molecule.
[0022] The reverse transcription reaction and template transformation were repeated at least twice, and the reaction temperature did not exceed 70°C.
[0023] Furthermore, the degradation universal primers use uracil DNA glycosylase and endonuclease, and the fixed sequence length of the degradation universal primers is 5-15 nt.
[0024] Further, in step 3), the ribonuclease is one or more of ribonuclease H, ribonuclease I, ribonuclease A, ribonuclease T1, and ribonuclease R.
[0025] Further, in step 3), the coding microsphere carries a cell-coding primer for the cell-coding addition of cDNA. The single-cell coding primer is linked to the microsphere by deoxyuridine nucleotides. The single-cell coding primer can be released under the combined action of uracil DNA glycosylase and endonuclease. The endonuclease can be one or more of endonuclease VIII, endonuclease III, or endonuclease IV.
[0026] Further, in step 4), the single-cell coding primer sequence consists of a fixed sequence A, a cell coding sequence, a unique molecular marker, and a fixed sequence B from the 5' end to the 3' end; wherein, the fixed sequence A is the complete Read1 sequencing primer sequence, used for subsequent PCR and sequencing; the cell coding sequence is different in different coding microspheres and is used to label cells; the unique molecular marker is different in the same coding microsphere and is used to label different cDNA products in a single cell; the sequence of the fixed sequence B is consistent with the forward sequence of the universal primer and is used to capture cDNA products and perform coding addition reactions.
[0027] Further, in step 4), the DNA polymerase is Taq and its mutants, including one of Phusion, Bst, Bst2.0, Bst2.0 WarmStart, Bst3.0, DeepVent, and Phi29.
[0028] Further, in step 5), the Read2 extension primer sequence with specially modified bases consists of the complete Read2 sequencing primer sequence and all or part of the universal primer sequence from the 5' end to the 3' end; the portion of the Read2 extension primer that is identical to the universal primer sequence contains at least 5 nt of the degraded universal primer sequence; the portion of the Read2 extension primer that is identical to the universal primer sequence contains at least 1 nt of specially modified bases; the special modification is locked nucleic acid or 2-fluoroRNA modification.
[0029] The present invention also provides an application of the method for constructing the single-cell whole transcriptome library in the whole transcriptome sequencing of single cells, single cell nuclei, and single microorganisms. The application includes applications in the fields of developmental biology, microbiology, basic medicine, clinical medicine, oncology, immunology, neuroscience, stem cell biology, regenerative medicine, genetics, hematology, agronomy, ecology, drug discovery and toxicology, pathology, spatial transcriptome integration analysis, aging research, and quality control of cell therapy products.
[0030] Compared with existing technologies, the method of the present invention has the following advantages:
[0031] 1) The method of the present invention is applicable to a wider range of sample types, and can analyze not only fresh tissue cells, but also fresh cell nuclei, frozen cells, cell nuclei, fixed cells, fixed cell nuclei, paraffin-embedded samples, etc.
[0032] 2) The method of the present invention can capture non-polyadenylated tail RNA and analyze important functions such as non-coding RNA and histone RNA;
[0033] 3) The method of this invention can achieve single-cell sequencing of microorganisms;
[0034] 4) The method of the present invention can detect the full length of transcripts, obtain a complete total RNA map, and analyze differential gene expression, alternative splicing, and RNA rate;
[0035] 5) Compared with existing microfluidic single-cell transcriptome sequencing technologies, the method of this invention has higher RNA detection sensitivity and can detect more gene expression at the same sequencing depth. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0037] Figure 1 This is a schematic diagram illustrating the process of constructing a single-cell whole transcriptome library according to the present invention;
[0038] Figure 2 This is a schematic diagram of the reverse transcription process in the construction of a single-cell whole transcriptome library;
[0039] Figure 3 This is a schematic diagram of the RNA on cDNA synthesized after reverse transcription degradation and the universal primers.
[0040] Figure 4 This is a schematic diagram of the reaction process that encodes cDNA products.
[0041] Figure 5 A schematic diagram illustrating the process of adding the complete Read2 sequence to the cDNA product after encoding;
[0042] Figure 6 This is a diagram showing the sequencing results of a mixed human-mouse cell sample; among them... Figure 6 In the figure, A represents the amplification curve of the qPCR experiment after reverse transcription. Figure 6 In the image, B represents a nucleic acid gel electrophoresis image, where pore 1 contains the reverse transcription product and pore 2 contains the sequencing library. Figure 6 The C in the scatter plot represents the number of UMIs detected for each cell. Figure 6 In the diagram, D represents a scatter plot of the number of UMIs mapped to the human and mouse genomes, with each point representing one cell. Figure 6 E in the figure represents the gene detection results obtained by sequencing a mixed human and mouse cell sample using the method of this invention;
[0043] Figure 7 This is a comparison of the sequencing results of the present invention method and the 10X Genomics method on mixed human and mouse cell samples; wherein, Figure 7In the figure, A is a statistical diagram showing the distribution of sequencing reads in the human genome from the 5' end to the 3' end of the reference gene, compared using the method of this invention and the 10X Genomics method. Figure 7 B in the figure represents a statistical graph of the total number of genes detected in a single human cell using this method and the 10X Genomics method. Figure 7 In the figure, C represents a statistical graph of the number of human protein-coding RNA genes detected using this method and the 10X Genomics method. Figure 7 In the figure, D represents a statistical graph of the number of human long non-coding RNA genes detected using this method and the 10X Genomics method. Figure 7 E in the figure represents the statistical graph of the number of human short noncoding RNA genes detected using this method and the 10X Genomics method;
[0044] Figure 8 This is a diagram showing the sequencing results of a frozen human glioma sample; in which, Figure 8 In the figure, A represents the amplification curve of the qPCR experiment after reverse transcription. Figure 8 In the image, B represents a nucleic acid gel electrophoresis image, where pore 1 contains the reverse transcription product and pore 2 contains the sequencing library. Figure 8 C in the diagram is a violin plot showing the number of UMIs aligned to the human genome. Figure 8 In this context, D represents the number of genes detected by sequencing frozen human glioma samples using the method of this invention.
[0045] Figure 9 A diagram showing the results of sequencing preparation for Klebsiella pneumoniae samples; in which, Figure 9 In the figure, A represents the amplification curve of the qPCR experiment after reverse transcription. Figure 9 In the image, B represents a nucleic acid gel electrophoresis image, where channel 1 is the sequencing library.
[0046] In the figure: 1. RNA molecule; 2. Universal primer; 21. Universal primer modified bases; 22. 3' end portion after universal primer degradation; 3. First cDNA product; 301. cDNA portion; 302. Protruding bases; 303. Reverse complementary sequence of universal primer; 31. Second cDNA product after degradation; 4. Cell-coding primer; 41. Fixed sequence A; 42. Cell-coding sequence; 43. Unique molecular marker; 44. Fixed sequence B; 5. Third cDNA product after encoding; 6. Read2 extension primer; 61. Read2 sequence; 62. Read2 extension primer modified bases; 7. Fourth cDNA product after encoding containing Read1 and Read2 sequences. Detailed Implementation
[0047] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the implementation methods of this invention, all other implementation methods obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0048] This invention provides a method for constructing a single-cell whole transcriptome library, the flowchart of which is shown below. Figure 1 As shown, it includes the following steps:
[0049] Step S10: Prepare a single-cell or single-cell nucleus suspension;
[0050] Step S20: Mix a single cell or single cell nucleus suspension with a reverse transcription reagent. The single cell or nucleus undergoes a reverse transcription reaction in an independent reaction vessel to obtain a single cell cDNA product.
[0051] Step S30: Add the coding reaction reagent and coding microspheres to the single-cell cDNA product to carry out the coding addition reaction;
[0052] Step S40: Construct a library from the encoded cDNA to obtain a single-cell whole transcriptome library.
[0053] In some embodiments, the reverse transcription reaction described in step S20 is as follows: Figure 2 As shown. In this invention, a specially designed universal primer 2 with a fixed sequence is used to randomly anneal and hybridize with RNA molecule 1 at low temperature. The resulting double-stranded molecule contains one RNA molecule and one cDNA product strand. The cDNA product strand includes universal primer 2, cDNA portion 301, prominent base 302, and the reverse complementary sequence 303 of the universal primer.
[0054] Specifically, universal primer 2 contains a universal primer-modified base 21. Universal primer 2 extends forward under the action of reverse transcriptase, generating cDNA portion 301. Due to the characteristics of MMLV and TGIRT family reverse transcriptases, one or more overhanging bases 302 are synthesized at the 3' end of cDNA portion 301. After universal primer 2 anneals with the overhanging base 302, the reverse transcriptase undergoes template switching, continuing to synthesize the remaining universal primer's reverse complementary sequence 303 using universal primer 2 as a template after the overhanging base 302. Universal primer 2, cDNA portion 301, overhanging base 302, and the universal primer's reverse complementary sequence 303 together constitute the first cDNA product 3.
[0055] The universal primer 2 consists of a fixed sequence composed of three bases: G, A, and T, with a length of 10-30 nt; at least one thymine in the fixed sequence of the universal primer is replaced with deoxyuridine.
[0056] In some embodiments, step S30 includes:
[0057] 1) Degrade the unreacted universal primer 2, the unreacted RNA molecule 1, and the RNA molecule 1 and universal primer 2 portions in the hybridized double strand of the first cDNA product 3 produced by the reaction, such as... Figure 3 As shown, RNA molecule 1 was degraded using ribonuclease; the universal primer modified base 21 of universal primer 2 was degraded using uracil DNA glycosylase and endonuclease. The second cDNA product 31 after degradation consisted of the 3' end portion 22 after degradation of the universal primer, cDNA portion 301, protruding base 302, and the reverse complementary sequence 303 of the universal primer.
[0058] 2) Encoding is added to the degraded second cDNA product 31, such as... Figure 4 As shown, the sequence of cell-coding primer 4 consists of fixed sequence A 41, cell-coding sequence 42, unique molecular marker 43, and fixed sequence B 44. Cell-coding primer 4 anneals to the degraded second cDNA product 31 through fixed sequence B 44, and is extended under the action of DNA polymerase to obtain the encoded third cDNA product 5.
[0059] The fixed sequence A 41 is the complete Read1 sequencing primer sequence, used for subsequent PCR and sequencing; the cell coding sequence 42 is a variable sequence composed of 20 oligonucleotides, unique to each cell, and the sequence is different between two cells; the unique molecular marker sequence 43 is a random sequence composed of 6 oligonucleotides, different for each transcript, used to mark different cDNA products in a single cell; the fixed sequence B 44 is consistent with the forward sequence of the universal primer, used to capture cDNA products and perform coding addition reactions.
[0060] In some embodiments, step S40 includes:
[0061] 1) Use Read2 extension primer 6 to add Read2 sequence 61 to the encoded third cDNA product 5, as follows: Figure 5As shown, Read2 extension primer 6 consists of Read2 sequence 61, the 3' end portion 22 after degradation of the universal primer, and Read2 extension primer modified base 62. After annealing with the encoded third cDNA product 5, Read2 extension primer 6 is extended under the action of DNA polymerase, thus completing the addition of the Read2 sequence. Then, using the Read1 sequence, i.e., the fixed sequence A 41, and Read2 sequence 61 as amplification primers, PCR amplification is performed to obtain the fourth cDNA product 7 containing the encoded Read1 and Read2 sequences.
[0062] 2) The fourth cDNA product 7, which contains the Read1 and Read2 sequences, is amplified by PCR using P5 and P7 primers with index information to obtain a single-cell transcriptome library.
[0063] The Read1 sequence is shown in SEQ ID NO.1, SEQ ID NO.1: 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3'; The Read2 sequence is shown in SEQ ID NO.2, SEQ ID NO.2: 5'-ATCTCGTATGCCGTCTTCTGCTTG-3';
[0064] The P5 primer sequence with index information is shown in SEQ ID NO.3, SEQ ID NO.3: 5'-AATGATACGGCGACCACCGAGATCTACAC[i5]TCGTCGGCAGCGTC-3';
[0065] The P7 primer sequence with index information is shown in SEQ ID NO.4; SEQ ID NO.4: 5'-CAAGCAGAAGACGGCATACGAGAT[i7]ATCTCGTATGCCGTCTTCTGCTTG-3';
[0066] [i5] and [i7] are index sequences, which are variable sequences of length 6-10 nt.
[0067] Example 1: Sequencing of human-mouse cell pooled samples
[0068] Cell culture: HEK293T cells and NIH / 3T3 cells were cultured in DMEM / high glucose medium containing 10% v / v FBS and 1% v / v penicillin-streptomycin, and passaged every 2-3 days.
[0069] Single-cell isolation and pretreatment: Cells were digested with trypsin until monodisperse, and washed twice with PBS. Aggregated cells were removed by filtration through a 40 μm cell strainer. Cells were diluted to 1 x 10⁻⁶ cells with PBS. 6 Cells / mL. Add the same volume of 30% v / v OptiPrep (PBS containing 30% v / v OptiPrep) to the cell suspension and mix well.
[0070] Droplet generation and reverse transcription: Prepare a 25 μL reverse transcription-lysis reagent containing 12.5 μL first-chain buffer (4x, containing 4% v / v Triton X-100, 0.04% v / v Tween-20), 300 U reverse transcriptase, 100 U RNaseOUT recombinant ribonuclease inhibitor, 1.25 μL dNTPs (10 mM each), 2.5 μL 50 μM universal primers, and 1 μL 25 mg / mL bovine serum albumin. Single-cell suspension, reverse transcription-lysis reagent, and oil phase (electronic fluorinated solution 7500 containing 2% v / v surfactant) were added to syringes and connected to the corresponding inlets of the microfluidic chip through tubing. The single-cell suspension, reverse transcription-lysis reagent, and oil phase were injected into the microfluidic chip at flow rates of 300 μL / h, 300 μL / h, and 800 μL / h, respectively, to obtain water-in-oil droplets containing single cells and reverse transcription-lysis reagent, which were collected in 0.2 mL PCR tubes.
[0071] The universal primer sequence used is shown in SEQ ID NO.5, SEQ ID NO.5: 5'-GTGAGTGGTGUGTGTAGTTGGAT-3'.
[0072] The collected droplets were placed on a thermal cycler and incubated at 10°C for 10 minutes to lyse the cells. This was followed by eight cycles of annealing, with the following program for each cycle: 12°C for 30 seconds, 30°C for 1 minute, and 42°C for 5 minutes. Finally, the cells were incubated at 75°C for 20 minutes to inactivate the reverse transcriptase. A portion of the reverse transcription product was then used for qPCR to determine the total amount of captured cDNA. The results showed a CT value of approximately 13 (see [link to qPCR]). Figure 6 (A) The amplification products were subjected to gel electrophoresis, and the results showed that the size distribution of the diffuse bands was mainly 100~1000 bp (see A). Figure 6 (B) indicates that the method of the present invention has high efficiency in capturing cDNA from cell line samples and covers long fragments, and can be used for high-throughput sequencing.
[0073] Single-cell coding addition: Prepare a solution containing 10 μL PCR buffer (10x), 2.5 μL dNTPs (10 mM each), 2% v / v Tween-20, 5 U USER enzyme, 5 U ribonuclease H, and 37.5 U ribonuclease I. f A total of 50 μL of 5 U DNA polymerase and 1.5 mg / mL bovine serum albumin-encoded extension reaction reagent were added. Droplets, the extension reaction reagent, the encoded microspheres, and the oil phase were simultaneously introduced into a microfluidic chip at flow rates of 100 μL / h, 200 μL / h, 100 μL / h, and 500 μL / h, respectively, to fuse single-cell droplets with the extension reaction reagent droplets, resulting in water-in-oil droplets containing single-cell reverse transcription products, the extension reaction reagent, and the microspheres. These droplets were collected in 0.2 mL PCR tubes. The collected droplets were placed on a thermal cycler and incubated at 37°C for 1 hour; 80°C for 10 seconds; for 10 cycles (64°C for 30 seconds, 70°C for 20 seconds); and finally at 70°C for 1 minute. After the reaction was complete, 20% v / v PFO (electronic fluorination solution 7500 containing 20% 1H, 1H, 2H, 2H-Perfluorooctanol) was added to the tube to break the droplets. The encoded cDNA was extracted using magnetic beads and eluted with 20 μL of DEPC water.
[0074] The coding microspheres contain cell-coding primers, the general structural formula of which is shown in SEQ ID NO. 6. SEQ ID NO. 6: 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG[NNNNNNNNNNNNNNNNNNNN][NNNNNN]GTGAGTGGTGTGTGTAGTTGGAT-3'.
[0075] Residual single-cell coding primer blocking: Prepare a primer blocking reagent containing 20 μL purified cDNA, 3 μL terminal transferase buffer (10x), 3 μL 2.5 mM cobalt chloride, 3 μL 10 mM ddCTP, and 5 U terminal transferase. Incubate at 37°C for 1 hour on a thermal cycler. Purify and extract cDNA using magnetic beads, and elute with 20 μL DEPC water.
[0076] Extension reaction: Prepare an extension reaction reagent containing 20 μL of the product from the previous step, 2.5 μL of PCR buffer (10x), 0.5 μL of dNTPs (10 mM each), 1.25 μL of 10 μM Read2 extension primers, and 1 U of DNA polymerase. Incubate on a thermal cycler at 95°C for 1 min, 62°C for 30 sec, and 72°C for 2 min. Purify and extract the reaction product using magnetic beads, and elute with 20 μL of DEPC water.
[0077] The Read2 extension primer sequence is shown in SEQ ID NO.7: 5'-ATCTCGTATGCCGTCTTCTGCTTGGTGTAGTTGGAT-3'.
[0078] Pre-amplification of cDNA product: An extension reaction reagent containing 20 μL of the product from the previous step, 4 μL of PCR buffer (10x), 1 μL of dNTPs (10 mM each), 2 μL of 10 μM Read1 primer, 2 μL of 10 μM TruSeq Read2 primer, and 1.2 U of DNA polymerase was prepared. The sequence of the Read1 primer is shown in SEQ ID NO.1; the sequence of the Read2 primer is shown in SEQ ID NO.2. The mixture was placed on a thermal cycler and incubated at 95°C for 1 minute; 15 PCR cycles were performed (95°C for 15 seconds, 58°C for 20 seconds, 72°C for 2 minutes); followed by incubation at 72°C for 5 minutes. The pre-amplified product was purified using magnetic beads and eluted with 20 μL of DEPC water.
[0079] Addition of sequencing library tags and sequencing adapters: Prepare library construction reaction reagents containing 10 ng of the product from the previous step, 25 μL of NEBNext Ultra II Q5 premix (2x), 2 μL of 10 μM P5 primer, and 2 μL of 10 μM P7 primer.
[0080] The P5 primer sequence is shown in SEQ ID NO.8; SEQ ID NO.8: 5'-AATGATACGGCGACCACCGAGATCTACACTAGATCGCTCGTCGGCAGCGTC-3'; The P7 primer sequence is shown in SEQ ID NO.9; SEQ ID NO.9: 5'-CAAGCAGAAGACGGCATACGAGATCGTACTAGATCTCGTATGCCGTCTTCTGCTTG-3'. The sample was placed on a thermal cycler and incubated at 98°C for 30 seconds, followed by 6 PCR cycles (98°C for 15 seconds, 70°C for 30 seconds, 72°C for 60 seconds), and then incubated at 72°C for 5 minutes. The amplified cDNA was purified using magnetic beads and eluted with 40 μL of DEPC water. Next-generation sequencing was performed after library quantification and sequencing.
[0081] Sequencing results showed that cells containing single-cell products encoded detected UFI numbers significantly higher than the background (see [link to data]). Figure 6 (C in the original text). A total of 117 cells were detected in the human-mouse mixed sample, of which 2 cells contained both human and mouse cDNA, indicating that the cell twinning rate in this method was 1.7% (see [link to original text]). Figure 6 (D in the original text). This indicates that the method can accurately distinguish single cells in mixed samples, has high transcriptome sequencing accuracy, and low cross-contamination between species. The method of this invention can detect a median of 13,809 expressed genes in human 293T cells and a median of 8,829 expressed genes in mouse 3T3 cells. This demonstrates the high sensitivity of the method (see D). Figure 6 (E in the original text). Further comparison of the data from this method with data from 10X Genomics shows that the sequencing reads from 293T cells obtained using this method are more evenly distributed across the reference gene sequence (see [link to original text]). Figure 7 (A) The proportion of non-coding RNA detected by this method is significantly higher than that of 10X Genomics, indicating that this method can detect more types of RNA, especially non-coding RNAs without polyadenylated tails (see A). Figure 7 (BE in the middle).
[0082] Example 2: Sequencing of frozen human glioma samples
[0083] Frozen sample lysis: Place the glioma sample in a tissue lysis buffer containing 10 mM Tris-HCl buffer (pH 8.0), 0.1% w / v Triton X-100, 250 mM sucrose, 25 mM potassium chloride, 5 mM magnesium chloride, 0.5 mM calcium chloride, 2 mg / mL collagenase IV, and 0.4 U / μL SUPERaseIN. Mince the sample with scissors and place it in a 2-3 mm size. 3Small pieces were then transferred to a Dunns homogenizer and crushed 5 times with a type A pestle and 10 times with a type B pestle. The tissue suspension was then incubated at 37°C for 2 minutes for further dissociation. The resulting tissue suspension was centrifuged at 800 g for 8 minutes at 4°C, resuspended, and flow cytometry was used to obtain a single-cell nucleus suspension. The flow-cytometry-sorted single-cell nucleus suspension was centrifuged at 800 g for 8 minutes at 4°C and resuspended in PBS solution containing 0.04% (w / v) BSA and 9% (w / v) Ficoll 400 (density approximately 1 x 10⁻⁶). 6 Mix well (each 1 / mL) and set aside.
[0084] Droplet generation and reverse transcription: Prepare a 25 μL reverse transcription-lysis reagent containing 12.5 μL first-chain buffer (4x, containing 4% v / v Triton X-100 and 0.04% v / v Tween-20), 300 U reverse transcriptase, 100 U RNaseOUT recombinant ribonuclease inhibitor, 1.25 μL dNTPs (10 mM each), 2 μL 50 μM universal primers, 1 μL 25 mg / mL bovine serum albumin, and 2 μL 50% w / v PEG8000. Single-cell suspension, reverse transcription-lysis reagent, and oil phase (electronic fluorinated solution 7500 containing 2% v / v surfactant) were added to syringes and connected to the corresponding inlets of the microfluidic chip via tubing. The single-cell suspension, reverse transcription-lysis reagent, and oil phase were injected into the microfluidic chip at flow rates of 300 μL / h, 300 μL / h, and 800 μL / h, respectively, to obtain water-in-oil droplets containing single cells and reverse transcription-lysis reagent, which were collected in 0.2 mL PCR tubes. The collected droplets were placed on a thermal cycler and incubated at 10°C for 10 minutes to lyse the cells, followed by 8 cycles of annealing (12°C to 42°C), and then incubated at 75°C for 20 minutes to inactivate the reverse transcriptase. A portion of the reverse transcription product was used for qPCR to detect the total amount of captured cDNA. The results showed a CT value of approximately 17 (see [link to qPCR]). Figure 8 (A) The amplification products were subjected to gel electrophoresis, and the results showed that the size distribution of the diffuse bands was mainly 150~500 bp (see A). Figure 8 (B) indicates that the method of the present invention has high efficiency in capturing cDNA from frozen tissue samples, covers a wide range of long fragments, and can be used for high-throughput sequencing.
[0085] The universal primer sequence used is shown in SEQ ID NO.10: 5'-GGTGTAGTGAGTUGATGTGTAGTTGG-3'.
[0086] Single-cell coding addition: Prepare a solution containing 10 μL PCR buffer (10x), 2.5 μL dNTPs (10 mM each), 2% v / v Tween-20, 5 U USER enzyme, 5 U ribonuclease H, and 37.5 U ribonuclease I. f 5 U DNA polymerase and 50 μL of extension reaction reagent encoded with 1.5 mg / mL bovine serum albumin were added. Using a microfluidic chip, droplets, extension reaction reagent, microspheres, and oil phase were simultaneously introduced into the microfluidic chip at flow rates of 100 μL / h, 200 μL / h, 100 μL / h, and 500 μL / h, respectively, to fuse single-cell droplets with the extension reaction reagent droplets, resulting in water-in-oil droplets containing single-cell reverse transcription products, extension reaction reagent, and microspheres. These droplets were collected in 0.2 mL PCR tubes. The collected droplets were placed on a thermal cycler and incubated at 37°C for 1 hour; 80°C for 10 seconds; for 10 cycles (64°C for 30 seconds, 70°C for 20 seconds); and 70°C for 1 minute. After the reaction was complete, 20% v / v PFO (electronic fluorination solution 7500 containing 20% 1H, 1H, 2H, 2H-Perfluorooctanol) was added to the tube to break the droplets. The encoded cDNA was extracted using magnetic beads and eluted with 20 μL of DEPC water.
[0087] The coding microspheres contain cell-coding primers, the general structural formula of which is shown in SEQ ID NO.11. SEQ ID NO.11: 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG[NNNNNNNNNNNNNNNNNNNN][NNNNNN]GGTGTAGTGAGTTGATGTGTAGTTGG-3'.
[0088] Residual single-cell coding primer blocking: Prepare a primer blocking reagent containing 20 μL purified cDNA, 3 μL terminal transferase buffer (10x), 3 μL 2.5 mM cobalt chloride, 3 μL 10 mM ddCTP, and 10 U terminal transferase. Incubate at 37°C for 1 hour on a thermal cycler. Purify and extract cDNA using magnetic beads, and elute with 20 μL DEPC water.
[0089] Extension reaction: Prepare an extension reaction reagent containing 20 μL of the product from the previous step, 2.5 μL of PCR buffer (10x), 0.5 μL of dNTPs (10 mM each), 1.25 μL of 10 μM Read2 extension primers, and 1 U of DNA polymerase. Incubate on a thermal cycler at 95°C for 1 min, 62°C for 30 sec, and 72°C for 2 min. Purify and extract the reaction product using magnetic beads, and elute with 20 μL of DEPC water.
[0090] The Read2 extension primer sequence is shown in SEQ ID NO.12: 5'-ATCTCGTATGCCGTCTTCTGCTTGGATGTGTAGTTGG-3'.
[0091] Pre-amplification of cDNA product: An extension reaction reagent containing 20 μL of the product from the previous step, 4 μL of PCR buffer (10x), 1 μL of dNTPs (10 mM each), 2 μL of 10 μM Read1 primer, 2 μL of 10 μM TruSeq Read2 primer, and 1.2 U of DNA polymerase was prepared. The sequence of the Read1 primer is shown in SEQ ID NO.1; the sequence of the Read2 primer is shown in SEQ ID NO.2. The mixture was placed on a thermal cycler and incubated at 95°C for 1 minute; 15 PCR cycles were performed (95°C for 15 seconds, 58°C for 20 seconds, 72°C for 2 minutes); followed by incubation at 72°C for 5 minutes. The pre-amplified product was purified using magnetic beads and eluted with 20 μL of DEPC water.
[0092] Addition of sequencing library tags and sequencing adapters: Prepare library construction reaction reagents containing 10 ng of the product from the previous step, 25 μL of NEBNext Ultra II Q5 premix (2x), 2 μL of 10 μM P5 primer, and 2 μL of 10 μM P7 primer.
[0093] The P5 primer sequence is shown in SEQ ID NO.13; SEQ ID NO.13: 5'-AATGATACGGCGACCACCGAGATCTACACTAGATCGCTCGTCGGCAGCGTC-3', and the P7 primer sequence is shown in SEQ ID NO.14; SEQ ID NO.14: 5'-CAAGCAGAAGACGGCATACGAGATCGTACTAGATCTCGTATGCCGTCTTCTGCTTG-3' was placed on a thermal cycler and incubated at 98°C for 30 seconds, followed by 6 PCR cycles (98°C for 15 seconds, 70°C for 30 seconds, 72°C for 60 seconds), and then incubated at 72°C for 5 minutes. The amplified cDNA was purified using magnetic beads and eluted with 40 μL of DEPC water. Next-generation sequencing was performed after library quantification and sequencing.
[0094] Sequencing results showed a median of 5805 UMIs and a median of 3142 expressed genes detected in frozen human glioma samples. This indicates that the method of this invention is also applicable to frozen human tissue samples and has high sensitivity (see [link to original text]). Figure 8 (C and D in the text).
[0095] Example 3: Sequencing of microbial samples
[0096] Microbial sample preparation: After culture, the bacterial suspension was centrifuged at 14,000 rpm for 1 minute at 4°C. The bacteria were then resuspended twice in 1x PBS and centrifuged again. The bacteria were then resuspended in 7 mL of PBS containing 4% w / v PFA. The suspension was incubated at 4°C for 16 h under rotation. After incubation, the bacterial suspension was centrifuged at 14,000 rpm for 1 minute at 4°C. The suspension was then resuspended in 1 mL of cold PBS containing 0.01 U / μL SUPERase-IN. The suspension was washed twice with 700 μL of cold PBS containing 0.01 U / μL SUPERase-IN and centrifuged again at 14,000 rpm for 1 minute at 4°C. The bacteria were then resuspended in 250 μL of PBS containing 0.04% v / v Tween-20 and incubated on ice for 3 minutes. (The number of bacteria in this step determines the permeation effect; generally, no more than 10 bacteria permeate per 250 μL.) 6 (For larger numbers of bacteria, it is recommended to divide them into multiple reaction tubes). After incubation, add 1 mL of cold PBS containing 0.01 U / μL SUPERase-IN and centrifuge. Resuspend the bacteria in a solution containing 100 mM Tris-HCl buffer (pH 7.0), 50 mM EDTA, 2.5 mg / mL lysozyme, and 0.25 U / μL SUPERase-IN. Incubate at 37°C for 15 minutes, then add 1 mL of solution containing 0.01 U / μL SUPERase-IN and wash the bacteria twice with PBS. Resuspend the bacteria in PBS to achieve a bacterial density of approximately 2 x 10⁻⁶. 6 per mL.
[0097] Droplet generation and reverse transcription: Prepare a 25 μL reverse transcription-lysis reagent containing 12.5 μL first-chain buffer (4x, containing 4% Triton X-100, 0.04% Tween-20), 300 U reverse transcriptase, 100 U RNaseOUT recombinant ribonuclease inhibitor, 1.25 μL dNTPs (10 mM each), 2.5 μL 50 μM universal primers, and 1 μL 25 mg / mL bovine serum albumin. Single-cell suspension, reverse transcription-lysis reagent, and oil phase (electronic fluorinated solution 7500 containing 2% v / v surfactant) were added separately to syringes and connected to the corresponding inlets of the microfluidic chip via tubing. The single-cell suspension, reverse transcription-lysis reagent, and oil phase were injected into the microfluidic chip at flow rates of 300 μL / h, 300 μL / h, and 800 μL / h, respectively, to obtain water-in-oil droplets containing single cells and reverse transcription-lysis reagent, which were collected in 0.2 mL PCR tubes. The collected droplets were placed on a thermal cycler and incubated at 10°C for 10 minutes to lyse the cells, followed by 8 cycles of annealing (12°C to 42°C), and then incubated at 75°C for 20 minutes to inactivate the reverse transcriptase. A portion of the reverse transcription product was used for qPCR to detect the total amount of captured cDNA. The results showed a CT value of approximately 19.5 (see [link to qPCR]). Figure 9 (A) The amplification products were subjected to gel electrophoresis, and the results showed that the size distribution of the diffuse bands was mainly 250~800bp (see A). Figure 9 (B) indicates that the method of the present invention has high efficiency and wide coverage in capturing cDNA from microbial samples and can be used for high-throughput sequencing.
[0098] The universal primer sequence used is shown in SEQ ID NO.15, SEQ ID NO.15: 5'-GGTGTGAAGTGUUGTAGTGATGTTGGG-3'.
[0099] Single-cell coding addition: Prepare a solution containing 10 μL PCR buffer (10x), 2.5 μL dNTPs (10 mM each), 2% v / v Tween-20, 5 U USER enzyme, 5 U ribonuclease H, and 37.5 U ribonuclease I. f50 μL of a combination of 5 U DNA polymerase and a 1.5 mg / mL bovine serum albumin-encoded extension reaction reagent were used. Using a microfluidic chip, droplets, the coded extension reaction reagent, coded microspheres, and the oil phase were simultaneously introduced at flow rates of 100 μL / h, 200 μL / h, 100 μL / h, and 500 μL / h to fuse single-cell droplets with the coded extension reaction reagent droplets, resulting in water-in-oil droplets containing single-cell reverse transcription products, the coded extension reaction reagent, and the coded microspheres. These droplets were collected in 0.2 mL PCR tubes. The collected droplets were placed on a thermal cycler and incubated at 37°C for 1 hour; 80°C for 10 seconds; for 10 cycles (64°C for 30 seconds, 70°C for 20 seconds); and finally at 70°C for 1 minute. After the reaction was complete, 20% v / v PFO (electronic fluorination solution 7500 containing 20% 1H, 1H, 2H, 2H-Perfluorooctanol) was added to the tube to break the droplets. The encoded cDNA was extracted using magnetic beads and eluted with 20 μL of DEPC water.
[0100] The coding microspheres contain cell-coding primers, the general structural formula of which is shown in SEQ ID NO.16: 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG[NNNNNNNNNNNNNNNNNNNN][NNNNNN]GGTGTGAAGTGTTGTAGTGATGTTGGG-3'.
[0101] Residual single-cell coding primer blocking: Prepare a primer blocking reagent containing 20 μL purified cDNA, 3 μL terminal transferase buffer (10x), 3 μL 2.5 mM cobalt chloride, 3 μL 10 mM ddCTP, and 10 U terminal transferase. Incubate on a thermal cycler at 37°C for 1 hour. Purify and extract cDNA using magnetic beads, and elute with 20 μL DEPC water.
[0102] Extension reaction: Prepare an extension reaction reagent containing 20 μL of the product from the previous step, 2.5 μL of PCR buffer (10x), 0.5 μL of dNTPs (10 mM each), 1.25 μL of 10 μM Read2 extension primers, and 1 U of DNA polymerase. Incubate on a thermal cycler at 95°C for 1 min, 62°C for 30 sec, and 72°C for 2 min. Purify and extract the reaction product using magnetic beads, and elute with 20 μL of DEPC water.
[0103] The Read2 extension primer sequence is shown in SEQ ID NO.17: 5'-ATCTCGTATGCCGTCTTCTGCTTG-3'.
[0104] Pre-amplification of cDNA product: An extension reaction reagent containing 20 μL of the product from the previous step, 4 μL of PCR buffer (10x), 1 μL of dNTPs (10 mM each), 2 μL of 10 μM Read1 primer, 2 μL of 10 μM TruSeq Read2 primer, and 1.2 U of DNA polymerase was prepared. The sequence of the Read1 primer is shown in SEQ ID NO.1; the sequence of the Read2 primer is shown in SEQ ID NO.2. The mixture was placed on a thermal cycler and incubated at 95°C for 1 minute; 15 PCR cycles were performed (95°C for 15 seconds, 58°C for 20 seconds, 72°C for 2 minutes); followed by incubation at 72°C for 5 minutes. The pre-amplified product was purified using magnetic beads and eluted with 20 μL of DEPC water.
[0105] Addition of sequencing library tags and sequencing adapters: Prepare library construction reaction reagents containing 10 ng of the product from the previous step, 25 μL of NEBNext Ultra II Q5 premix (2x), 2 μL of 10 μM P5 primer, and 2 μL of 10 μM P7 primer.
[0106] The P5 primer sequence is shown in SEQ ID NO.18; SEQ ID NO.18: 5'-AATGATACGGCGACCACCGAGATCTACACTAGATCGCTCGTCGGCAGCGTC-3'; the P7 primer sequence is shown in SEQ ID NO.19; SEQ ID NO.19: 5'-CAAGCAGAAGACGGCATACGAGATCGTACTAGATCTCGTATGCCGTCTTCTGCTTG-3'. The primers were placed on a thermal cycler and incubated at 98°C for 30 seconds, followed by 6 PCR cycles (98°C for 15 seconds, 70°C for 30 seconds, 72°C for 60 seconds), and then incubated at 72°C for 5 minutes. The amplified cDNA was purified using magnetic beads and eluted with 40 μL of DEPC water. Next-generation sequencing was performed after library quantification and analysis.
[0107] The above are merely preferred embodiments of the present invention and do not limit the patent scope of the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the patent protection scope of the present invention.
Claims
1. A method for constructing a single-cell whole transcriptome library, characterized in that, Includes the following steps: 1) Prepare single-cell or single-cell nucleus suspensions; 2) Universal primers and a reaction solution containing reverse transcriptase are added to a single cell or single cell nucleus suspension. The single cell or nucleus undergoes reverse transcription to synthesize cDNA in an independent reaction vessel. The universal primers also act as template-changing oligonucleotides that contact the 3' end of the cDNA and continue to synthesize cDNA through template conversion, resulting in a double-stranded molecule containing one RNA molecule and one cDNA product chain. The reverse transcriptase is a reverse transcriptase MMLV or TGIRT-III with template conversion activity. The specific steps of the reverse transcription reaction and template conversion are as follows: Reverse transcription reaction: After the universal primers are annealed and hybridized with RNA from a single cell or cell nucleus at a temperature of 8℃-35℃, cDNA is synthesized using reverse transcriptase at a temperature of 20℃-65℃. Template conversion: After cDNA synthesis, the universal primers themselves are used as templates, and the template conversion activity of the reverse transcriptase is used to extend the cDNA to the 3' end to obtain a double-stranded molecule. The reverse transcription reaction and template transformation were repeated at least twice, and the reaction temperature did not exceed 70°C. 3) Inactivate reverse transcriptase, add coding reaction reagents and coding microspheres, and use ribonuclease to digest RNA molecules, retaining only one cDNA product chain; 4) The universal primers are degraded using uracil DNA glycosylase and endonuclease, and the deoxyuracil sites inserted in the universal primers are specifically cut to release the single-cell coding primers on the coding microspheres. The degraded universal primers have a fixed sequence length of 5-15 nt, which are annealed and hybridized with the 3' end of cDNA, and then a coding addition reaction is carried out under the action of DNA polymerase. 5) Purify the cDNA product containing single-cell encoding; add ddCTP to the ends of the primers using terminal transferase; The cDNA product encoding a single cell was extended using Read2 extension primers with specially modified bases to obtain a cDNA product with the complete Read2 sequencing primer sequence. 6) After amplifying the obtained cDNA product using the Read1 and Read2 sequences, the cDNA product is then amplified again using the P5 primer sequence with index information and the P7 primer sequence with index information to obtain a high-throughput single-cell whole transcriptome library. The universal primer consists of a fixed sequence composed of three bases: G, A, and T, with a length of 10-30 nt; at least one thymine in the fixed sequence of the universal primer is replaced with deoxyuridine.
2. The method for constructing a single-cell whole transcriptome library according to claim 1, characterized in that, In step 2), the independent reaction containers include: water-in-oil droplets, microfluidic chips, microwells, PCR tubes, 96-well plates or 384-well plates; when water-in-oil droplets are used as reaction containers, the method of adding coded reaction reagents and coded microspheres in step 3) is droplet fusion.
3. The method for constructing a single-cell whole transcriptome library according to claim 1, characterized in that, In step 4), the single-cell coding primer sequence consists of a fixed sequence A, a cell coding sequence, a unique molecular marker, and a fixed sequence B from the 5' end to the 3' end. The fixed sequence A is the complete Read1 sequencing primer sequence, used for subsequent PCR and sequencing. The cell coding sequence differs in different coding microspheres and is used to label cells. The unique molecular marker differs in the same coding microsphere and is used to label different cDNA products in a single cell. The fixed sequence B is consistent with the forward sequence of the universal primer and is used to capture cDNA products and perform coding addition reactions.
4. The method for constructing a single-cell whole transcriptome library according to claim 1, characterized in that, In step 4), the DNA polymerase is Taq and its mutants, including one of Q5, Phusion, Bst, Bst2.0, Bst2.0 WarmStart, Bst3.0, DeepVent, and Phi29.
5. The method for constructing a single-cell whole transcriptome library according to claim 1, characterized in that, In step 5), the Read2 extension primer sequence with special modified bases consists of the complete Read2 sequencing primer sequence and all or part of the universal primer sequence from the 5' end to the 3' end; the portion of the Read2 extension primer that is identical to the universal primer sequence contains at least 5 nt of the degraded universal primer sequence; the portion of the Read2 extension primer that is identical to the universal primer sequence contains at least 1 nt of special modified bases. The special modification is locked nucleic acid or 2-fluoroRNA modification.
6. The method for constructing a single-cell whole transcriptome library according to any one of claims 1-5 is applied to the whole transcriptome sequencing of single cells, single cell nuclei, and single microorganisms.