A method and kit for constructing a high-throughput single-cell full-length transcriptome sequencing library and a sequencing method
Through droplet microfluidic control technology and combined coding technology, combined with specific tag methods, a high-throughput single-cell full-length transcriptome sequencing library is constructed, solving the problems of high error rate, low throughput and high cost in the existing technology, and achieving the efficient, economical and high-throughput characteristics of large-scale single-cell full-length transcriptome sequencing.
Patent Information
- Application Number
- CN202210481013.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-05
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-05-05
AI Technical Summary
The existing single-cell full-length transcriptome sequencing technology has the problems of high error rate, low throughput and high cost, and it is difficult to apply to large-scale single-cell full-length transcriptome sequencing.
Using droplet microfluidic technology combined with combined coding technology, high-throughput sequencing analysis of 10 to 107 single-cell full-length transcriptomes was achieved through specific tag methods, and a high-throughput single-cell full-length transcriptome sequencing library was constructed.
The technical problem of large-scale single-cell full-length transcriptome sequencing is solved, the sequencing cost is reduced, the detection throughput is improved, and efficient single-cell full-length transcriptome sequencing is achieved.
Smart Images

Figure CN114737258B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gene sequencing. Specifically, it relates to a method and kit for constructing a high-throughput single-cell full-length transcriptome sequencing library and a sequencing method. Background Art
[0002] Single-cell transcriptome sequencing (scRNA-seq) technology is a new focus in the global life science field and a new tool for biomedical research. Single-cell transcriptome sequencing can sequence all RNA molecules at the single-cell level, revealing the heterogeneity between cells. It can not only identify cell subtypes but also reveal their functional states, developmental processes, lineage origins, and interactions, etc., and has extremely important application prospects in the fields of oncology, developmental biology, immunology, microbiology, neuroscience, etc. Therefore, various countries have launched a series of major single-cell funding programs, such as the Single Cell Analysis Program, The Human Cell Atlas, Atlas of Blood Cells, The Human BioMolecular Atlas Program, etc., to develop special innovative tools and technologies for single-cell sequencing, accelerate the integration and transformation of technologies, and provide powerful technologies and tools for basic research, early screening, clinical diagnosis and treatment, and drug development of major diseases such as cancer.
[0003] Currently, the key points and difficulties of single-cell transcriptome sequencing mainly lie in aspects such as technological breakthroughs and cost control. Technologically, most of the current single-cell high-throughput transcriptome platforms are based on next-generation sequencing platforms. For example, the SMART-Seq2 technology is used to amplify full-length transcripts, covering the 5' end to the 3' end of the transcript, and then the next-generation sequencing is used to read the sequence information. However, when constructing a next-generation sequencing library, the cDNA needs to be fragmented, and there are problems such as transcript splicing errors and incomplete information in the assembly and splicing of sequencing data.
[0004] In addition, researchers have combined single-cell isolation systems with third-generation sequencing technologies to develop single-cell full-length transcriptome sequencing technologies such as SCAN-seq, and successfully obtained the full-length transcriptome of single cells for the quantification of transcripts at the single-cell level, the identification of isoforms such as alternative splicing and fusion genes, etc. However, the third-generation sequencing technology is still in its infancy so far, with defects such as high error rate (10-30%), low throughput, and high cost, and cannot be applied to large-scale single-cell full-length transcriptome sequencing. Existing method examples of full-length transcriptome sequencing: third-generation full-length transcriptome sequencing (Iso-Seq).
[0005] The existing full-length transcriptome is based on the third-generation sequencing platforms of PacBio and Nanopore. Without fragmentation and splicing, it can directly obtain the full-length mRNA sequences containing 5'UTR, 3'UTR, and polyA tails, as well as complete structural information, so as to accurately analyze the structural information such as alternative splicing and fusion genes in species with a reference genome, and overcome the problems of short and incomplete transcript splicing in species without a reference genome. At the same time, with the help of second-generation sequencing data, transcript-specific expression analysis can be carried out to obtain more comprehensive annotation information. For reference, see Figure 1 . During the full-length transcriptome sequencing process, in order to avoid the bias of short-fragment libraries and ensure the coverage of transcripts of different lengths, three or more libraries will be constructed: 1-2 kb, 2-3 kb, and ≥3 kb libraries. The experimental procedure can be referred to Figure 2 , and the analysis procedure can be referred to Figure 3 .
[0006] At present, some manual operation links in single-cell transcriptome sequencing have been replaced by automated equipment, but the overall cost is still relatively high. Taking the Chromium system of 10X Genomics as an example, the preparation cost of a single cell is about $0.15 - $1, which is still not suitable for large-scale clinical applications.
[0007] Therefore, large-scale single-cell full-length transcriptome sequencing remains a major challenge in existing single-cell sequencing.
[0008] In view of this, the present invention is specifically proposed. Summary of the Invention
[0009] The object of the present invention is to provide a method and kit for constructing a high-throughput single-cell full-length transcriptome sequencing library, as well as a sequencing method.
[0010] The present invention is implemented as follows:
[0011] In a first aspect, an embodiment of the present invention provides a method for constructing a high-throughput single-cell full-length transcriptome sequencing library, which includes: using droplet microfluidics technology to capture and encapsulate n cells to be analyzed, where n is a positive integer and ≥ 10, generating hydrogel microspheres containing single-cell nuclear mRNA capture probes, lysing and reverse transcribing the single cells in the hydrogel microspheres to obtain reverse transcription products of mRNA - full-length cDNA; dividing the n hydrogel microspheres into x parts, where x is a positive integer and < n, using random primers to extend the full-length cDNA into double-stranded cDNA fragments for each part of the hydrogel microspheres to form a hybrid chimera of full-length cDNA and double-stranded cDNA fragments, and different specific tags are labeled on the random primers of each part of the hydrogel microspheres; performing k steps of labeling new specific tags on the tags of all hydrogel microspheres to form multi-level tags, and stopping labeling when the types of the multi-level tags ≥ n; using a polymerase with strand displacement activity to extend and displace the double-stranded cDNA fragments in all hydrogel microspheres, and amplifying the obtained products after displacement to obtain a sequencing library.
[0012] In a second aspect, an embodiment of the present invention provides a method for sequencing a high-throughput single-cell full-length transcriptome, which includes: constructing a high-throughput single-cell full-length transcriptome sequencing library using the method for constructing a high-throughput single-cell full-length transcriptome sequencing library as described in the foregoing embodiment, and sequencing the sequencing library.
[0013] In a third aspect, an embodiment of the present invention provides an application of a composition in preparing a kit for high-throughput single-cell full-length transcriptome sequencing, where the composition includes: reagents for implementing the method for constructing a high-throughput single-cell full-length transcriptome sequencing library as described in the foregoing embodiment or the method for sequencing a high-throughput single-cell full-length transcriptome as described in the foregoing embodiment.
[0014] In a fourth aspect, an embodiment of the present invention provides a kit for high-throughput single-cell full-length transcriptome sequencing, which includes: reagents for implementing the method for constructing a high-throughput single-cell full-length transcriptome sequencing library as described in the foregoing embodiment or the method for sequencing a high-throughput single-cell full-length transcriptome as described in the foregoing embodiment.
[0015] The present invention has the following beneficial effects:
[0016] The present invention combines droplet microfluidics technology, combinatorial encoding technology, and high-throughput sequencing technology, and creatively uses a method of specific tags to achieve high-throughput sequencing analysis of the full-length transcriptomes of 10 to 10 7 single cells simultaneously, solves the technical problems of large-scale single-cell full-length transcriptome sequencing, modifies and improves the existing single-cell transcriptome sequencing technology, reduces the sequencing cost, and improves the detection throughput. Description of the Drawings
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0018] Figure 1 Schematic diagram for transcript-specific expression analysis of existing second-generation sequencing data;
[0019] Figure 2 Schematic diagram of the experimental process of the existing full-length transcriptome;
[0020] Figure 3 Schematic diagram of the bioinformatics analysis process of the existing full-length transcriptome;
[0021] Figure 4 Structural and micrograph of the droplet chip; wherein, A1, A2, and A3 are aqueous phase inlets, B1 and B2 are oil phase inlets, and C is the outlet. A1 is introduced with cell lysate and initiator: 0.6% (vol / vol) Triton X-100, 6U μl –1 RNase inhibitor, 0.5% APS (ammonium persulfate); A2 is introduced with cell suspension: Hela cells at a concentration of 3x10 5 cells / μl, and the buffer solution is PBS; A3 is introduced with 30% acrylamide solution and 60 μM Oligo-dT30VN; the oil phase inlets B1 and B2 are introduced with fluorinated oil, and its composition is FC-40 containing 3% (w / w) PFPE-PEG and 3% (v / v) tetramethylethylenediamine;
[0022] Figure 5 Schematic diagram of the sorting of single cells and the principle of mRNA capture;
[0023] Figure 6 Schematic diagram of the present invention for labeling single-cell full-length transcriptome;
[0024] Figure 7 Schematic diagram of the extension reaction without ddNTPs and the addition of the first barcode tag;
[0025] Figure 8 Schematic diagram of the extension reaction with ddNTPs and the addition of the first barcode tag;
[0026] Figure 9 Schematic diagram of the connection and addition of the second new DNA barcode tag;
[0027] Figure 10Schematic diagram for the ligation of the third new DNA barcode label;
[0028] Figure 11 Schematic diagram for the composition and structure of the complete coding;
[0029] Figure 12 Schematic diagram for the analysis process of the high-throughput single-cell full-length transcriptome provided by the present invention;
[0030] Figure 13 Schematic diagram for the splicing and assembly principle of the full-length transcriptome of the present invention. Detailed implementation manners
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. For those not specified in the embodiments, they shall be carried out according to conventional conditions or conditions recommended by the manufacturer. For the reagents or instruments whose manufacturers are not specified, they are all conventional products that can be obtained through commercial purchase.
[0032] The embodiments of the present invention provide a method for constructing a high-throughput single-cell full-length transcriptome sequencing library, which includes: using droplet microfluidics technology to capture and encapsulate n cells to be analyzed, where n is a positive integer and ≥ 10, to generate hydrogel microspheres containing single cells and mRNA capture probes, lysing and reverse transcribing the single cells in the hydrogel microspheres to obtain the reverse transcription product of mRNA - full-length cDNA; then, dividing the n hydrogel microspheres into x parts, where x is a positive integer and < n, and using random primers to extend the full-length cDNA into double-stranded cDNA fragments for each part of the hydrogel microspheres to form a hybrid chimera of full-length cDNA and double-stranded cDNA fragments, and different specific tags are labeled on the random primers of each part of the hydrogel microspheres; performing k steps of labeling new specific tags on the tags of all the hydrogel microspheres to form multi-level tags, and stopping the step of labeling new specific tags when the types of the multi-level tags ≥ n; using a polymerase with strand displacement activity to extend and displace the double-stranded cDNA fragments in all the hydrogel microspheres, and amplifying the obtained product after displacement to obtain a sequencing library.
[0033] It should be noted that the "double-stranded" in the double-stranded cDNA fragment does not refer to 2 strands, but is a substitute name for the cDNA fragment relative to the full-length cDNA (single-strand).
[0034] The present invention combines droplet microfluidics technology and high-throughput sequencing (such as next-generation sequencing), and creatively uses a specific tagging method to achieve 10 - 10 7Sequencing analysis of single-cell full-length transcriptomes has solved the technical problems of large-scale single-cell full-length transcriptome sequencing, modified and improved existing single-cell transcriptome sequencing technologies, reduced the sequencing cost, and increased the detection throughput.
[0035] In a preferred embodiment, each step of labeling new specific tags includes: mixing x hydrogel microspheres and then re-dividing them into y portions, where y is a positive integer and < n, and then labeling different specific tags on the tags of random primers in each portion of hydrogel microspheres to form multi-level specific tags. Preferably, the distribution method of the y portions of hydrogel microspheres is an even distribution. Optionally, x = y.
[0036] In a preferred embodiment, the step of labeling different specific tags on the tags of random primers in hydrogel microspheres to form multi-level specific tags includes: adding a ligase and a ligation probe containing a tag sequence to the hydrogel microspheres.
[0037] Preferably, the ligation probe is a nucleic acid sequence, which includes three sequentially connected parts: a first sequence, a tag sequence, and a second sequence. The first sequence is a sequence for hybridizing with an existing ligation probe on the double-stranded cDNA fragment, and the second sequence is a sequence for hybridizing with a subsequent ligation probe. It should be noted that the subsequent ligation probe refers to the ligation probe used in the next tagging, and it can hybridize and bind with the first sequence of the next ligation probe.
[0038] Preferably, the step of labeling different specific tags on the tags of random primers in hydrogel microspheres to form multi-level specific tags further includes: adding a linker to the hydrogel microspheres. The first sequence hybridizes with the existing ligation probe on the double-stranded cDNA fragment through the linker, and the second sequence hybridizes with the first sequence on the subsequent ligation probe through the linker.
[0039] In a preferred embodiment, the reaction conditions for labeling the new specific tag are as follows: the working concentration of the ligase is any value in the range of 1 to 5 unit / μL, such as any one or the range between any two of 1 unit / μL, 2 unit / μL, 3 unit / μL, 4 unit / μL, and 5 unit / μL, preferably 2 unit / μL; the working concentration of the ligation probe is any value in the range of 0.5 to 15 μM, and the reaction time is any value in the range of 0.5 to 2.5 h. Specifically, the working concentration of the probe can be any one or the range between any two of 0.5 μM, 1 μM, 2 μM, 3 μM, 4 μM, 5 μM, 6 μM, 7 μM, 8 μM, 9 μM, 10 μM, 11 μM, 12 μM, 13 μM, 14 μM, and 15 μM, and the reaction time can be any one or the range between any two of 0.5 h, 0.6 h, 0.7 h, 0.8 h, 0.9 h, 1.0 h, 1.1 h, 1.2 h, 1.3 h, 1.4 h, 1.5 h, 1.6 h, 1.7 h, 1.8 h, 1.9 h, 2.0 h, 2.1 h, 2.2 h, 2.3 h, 2.4 h, and 2.5 h, preferably 2 h.
[0040] In a preferred embodiment, when performing the k-th labeling of the new specific tag, the ligation probe further includes molecular tags UMIs (Unique Molecular Identifiers, random barcode sequences, with a length of 4 to 20 bp) to ensure that each mRNA carries a unique molecular tag. UMIs add a unique tag sequence to each fragment after the transcriptome of the original sample is fragmented, which is used to distinguish thousands of different fragments in the same sample. In subsequent data analysis, these tag sequences can be used to exclude errors introduced during DNA polymerase, amplification, and sequencing processes. The principle is as follows: for DNA fragments of the same sample, each fragment carries a unique tag sequence, which will go through library construction, PCR amplification together with the target sequence, and then be sequenced together. Among the finally sequenced sequences, sequences with different tags represent that they come from different original DNA fragment molecules; sequences with the same molecular tag represent that these sequences are all amplified from the same original DNA fragment. Since errors during PCR and sequencing occur randomly, based on these molecular tags, systematic mutations introduced during PCR, sequencing, etc. can be excluded during the redundancy removal process, which can greatly reduce the false positive rate of low-frequency mutations.
[0041] Preferably, in the extension system for extending full-length cDNA into multiple double-stranded cDNA fragments, the effective concentration of the random primer is such that the length of the double-stranded cDNA fragments is within the read length range of next-generation sequencing. Preferably, the read length of next-generation sequencing is 200-300 bp. Preferably, the effective concentration of the random primer is any value in the range of 0.2-30 μM, and this effective concentration can effectively control the length of the double-stranded cDNA fragments within the read length range of next-generation sequencing. Specifically, it can be any value or the range between any two of 0.2 μM, 0.5 μM, 1 μM, 2 μM, 4 μM, 8 μM, 10 μM, 12 μM, 14 μM, 16 μM, 18 μM, 20 μM, 22 μM, 24 μM, 26 μM, 28 μM, and 30 μM. Preferably, it is 20 μM; the length of the random primer is any value in the range of 6-10 nt, specifically, it can be any value or the range between any two of 6 nt, 7 nt, 8 nt, 9 nt, and 10 nt.
[0042] A linking sequence for hybridizing with a subsequent tagging linking probe is further linked to the random primer or the tag labeled on the random primer. When performing k times of tagging, the new linking probe hybridizes with this linking sequence to achieve the purpose of tagging. Preferably, the linking sequence is linked to the end of the random primer or the sequence end of the tag labeled on the random primer, and its structure can be: random primer-tag-linking sequence.
[0043] In addition to limiting the effective concentration of the random primer, there is another method for controlling the length of the double-stranded cDNA fragments within the next-generation sequencing length range, that is, by adding ddNTPs to the extension system and controlling the effective concentration of ddNTPs, the length of the double-stranded cDNA fragments can also be controlled within the next-generation sequencing length range. Preferably, when using the random primer to extend the full-length cDNA, the construction method further includes adding dNTPs and ddNTPs to the extension system.
[0044] Preferably, the effective concentrations of dNTPs and ddNTPs are such that the length of the double-stranded cDNA fragments is within the read length range of next-generation sequencing.
[0045] Preferably, the working concentration of the dNTPs is any value in the range of 0.05 to 0.5 mM, specifically, it can be any one of 0.05 mM, 0.1 mM, 0.2 mM, 0.3 mM, 0.4 mM, and 0.5 mM or the range between any two of them; the working concentration of the ddNTPs is any value in the range of 0.05 to 0.5 mM, specifically, it can be any one of 0.05 mM, 0.1 mM, 0.2 mM, 0.3 mM, 0.4 mM, and 0.5 mM or the range between any two of them. Within this range of working concentration, the length of the double-stranded cDNA fragments can be more effectively limited within the range of the second-generation sequencing length.
[0046] In some embodiments, the working concentrations of the random primers and the ddNTPs can also be simultaneously selected to make the length of the double-stranded cDNA fragments obtained by extension within the read length range of the second-generation sequencing.
[0047] In this article, "the type of multi-level tags" = x * y *... * z. Preferably, when the type of multi-level tags ≥ (n × 10), stop the step of marking new specific tags, which can ensure that the full-length transcriptomes of single cells in all hydrogel microspheres carry different tags respectively, and each single cell can be distinguished based on the tags.
[0048] In a preferred embodiment, n is any value in the range of 10 to 10 7 For example, n can be 10, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 and 10 7 or the range between any two of them. Using the construction method provided by the present invention, the sequencing analysis of the full-length transcriptomes of up to 10 7 single cells can be completed simultaneously, and the sequencing accuracy is high and the cost is low.
[0049] The embodiment of the present invention also provides a high-throughput single-cell full-length transcriptome sequencing method, which includes: constructing a high-throughput single-cell full-length transcriptome sequencing library by using the construction method of the high-throughput single-cell full-length transcriptome sequencing library described in any of the foregoing embodiments, and sequencing the sequencing library.
[0050] In a preferred embodiment, the sequencing includes the splicing of sequencing data, and the splicing steps are as follows: in the obtained sequencing library, there is a part of overlapping sequence between two adjacent cDNA fragments amplified from each full-length transcriptome. Use the overlapping sequence to splice the Reads to obtain the sequencing analysis result. Preferably, the sequencing is second-generation sequencing.
[0051] The short reads of existing second-generation sequencing have deficiencies in understanding certain information about cell-to-cell heterogeneity, such as alternative splicing, gene fusion, and copy number variation. However, the single-cell SmartSeq2 technology that can measure full-length transcripts has bottleneck problems of low throughput and high cost. The sequencing method provided by the present invention can detect the full-length transcriptome information of multiple single cells with high throughput, and has outstanding substantial effects and remarkable progress compared with the existing technology.
[0052] The embodiments of the present invention also provide the use of a composition in the preparation of a kit for high-throughput single-cell full-length transcriptome sequencing, and the composition includes: reagents and methods for constructing a high-throughput single-cell transcriptome sequencing library as described in any of the foregoing embodiments or reagents for the sequencing method of the high-throughput single-cell full-length transcriptome as described in any of the foregoing embodiments.
[0053] In some embodiments, the reagents include reagents for implementing droplet microfluidics technology, ligase, ligation probes, random primers, dNTPs, ddNTPs, etc. It can be understood that these reagents in this application and subsequent kits are the same as those described in any of the foregoing embodiments and will not be elaborated further.
[0054] In addition, the embodiments of the present invention also provide a kit for high-throughput single-cell full-length transcriptome sequencing, which includes: reagents and methods for constructing a high-throughput single-cell transcriptome sequencing library as described in any of the foregoing embodiments or reagents for the sequencing method of the high-throughput single-cell full-length transcriptome as described in any of the foregoing embodiments.
[0055] The features and properties of the present invention will be further described in detail below in conjunction with embodiments.
[0056] Example 1
[0057] Materials:
[0058] Table 1 Reagents and Materials
[0059]
[0060] 1. High-efficiency and high-throughput single-cell processing
[0061] Design and fabricate a cross-shaped polydimethylsiloxane (PDMS) microfluidic chip, and the PDMS chip is fabricated using standard soft lithography technology. The width of the aqueous phase at the cross is 200 microns, the width of the oil phase is 100 microns, and the height of the chip is 50 microns. When generating droplets, the total flow rate of the aqueous phase is 0.3 mL / h, and the total flow rate of the oil phase is 1.0 mL / h, and the diameter of the obtained droplets is approximately 65 microns ( Figure 4 ).
[0062] Using a microfluidic droplet chip, single cells are encapsulated in acrylamide droplets of about 65 microns. The acrydite-modified mRNA capture probe (Oligo-T30: acrydite-AAGCAGTGGTATCAACGCAGAGTACTTTTTTTTTTTTTTTTTTTTTTTTTTTTTT) and acrylamide copolymerize in the presence of an initiator, covalently coupling the mRNA capture probe inside the polyacrylamide hydrogel microspheres. At the same time, the cells are lysed in the presence of cell lysis buffer, and the released mRNA is captured by the capture probe. All the mRNA of a single cell is trapped in a single hydrogel microsphere. The oil phase is removed by centrifugation, and the hydrogel microspheres are dispersed in PBS buffer solution. The sorting of single cells and the principle of mRNA capture are as Figure 5 shown.
[0063] 2. High-throughput single-cell full-length transcriptome encoding
[0064] Using random primers, the full-length first-strand cDNA is extended into several second-strand cDNA fragments; combined with combinatorial chemistry methods and ligase reactions, encoding labels are added to the ends of the second-strand cDNA fragments; using a DNA polymerase with strand displacement enzyme activity, the encoded second-strand cDNA fragments are extended and replaced to form second-strand cDNA products with overlapping sequences, and then amplified, library constructed, and second-generation sequencing is performed. The principle is as Figure 6 shown.
[0065] First, these hydrogel microspheres are evenly divided into 96 parts (x), and the full-length cDNA is extended into short second-strand cDNA fragments using 96 specific DNA barcode tags (1stBC) random primers (Round1_XX) to form a hybrid chimera of first-strand full-length cDNA and second-strand cDNA fragments. At the same time, the 96 parts of microspheres are respectively labeled with different barcode tags. The schematic diagram of the extension reaction without ddNTPs and the addition of the first barcode tag is shown in Figure 7 , and the schematic diagram of the extension reaction with ddNTPs and the addition of the first barcode tag is shown in Figure 8 .
[0066] Composition of the random primer (Round1_XX) containing the DNA barcode tag (1st BC): (The bold 15 nt hybridizes with part of the 2nd barcodelinker (linking sequence), which is the same for all 96 wells; the underlined 8 nt part is the first barcode tag of the 1st ExtensionBarcode, which is different for each of the 96 wells; the last 6 nt is the random primer 6N).
[0067] The recipe of the extension reaction solution for each reaction well is as follows: 20 μM Round1_XX (1st BC), 0.2 mM dNTPs, 0.2 mM ddNTPs, 0.1 U / μL DNA polymerase, 1x PCR Reaction buffer.
[0068] Secondly, mark k new tags on the 1st BC, where k = 2: Mix the above chimeric hydrogel microspheres evenly, then divide them into 96 equal parts (y), and connect them to the ligation probes (Round2_XX) of 96 specific tags (2nd BC) with DNA barcodes respectively to obtain hydrogel microspheres with 96 2 specific tags.
[0069] Composition of the ligation probe of the specific tag (2nd BC) with DNA barcode (Round2_XX): (The bold 15 nt hybridizes with the 3rd barcode linker part, which is the same for all 96 wells; the underlined 8 nt part is the second barcode tag of the 2nd Ligation Barcode, which is different for each of the 96 wells; the last 15 nt hybridizes with the 2nd barcode linker part, which is the same for all 96 wells).
[0070] Among them, the sequence of the 3rd barcode linker is: 5’-AGTCGTACGCCGATGCGAAACATCGGCCAC-3’. The sequence of the 2nd Ligation Barcode is: 5’-CGAATGCTCTGGCCTCTCAAGCACGTGGAT-3’.
[0071] The schematic diagram of the addition of the second new DNA barcode tag is shown in Figure 9 .
[0072] The composition of the 2nd ligation reaction solution is shown in Table 2.
[0073] Table 2 Ligation reaction solution
[0074] Composition 20 μl reaction system T4 DNA Ligase Buffer(10X) 2 μl 2nd Ligation Barcodes, Round2_xx 10 μM 2nd barcode linker 15 μM Hydrogel microspheres --- Nuclease-free water to 20 μl T4 DNA Ligase 1 μl
[0075] Mark 3 new tags on the 2nd BC: Mix the above chimeric hydrogel microspheres evenly, then divide them into 96 equal parts, and connect them to the ligation probes (3rd LigationBarcodes, Round3_xx) of 96 specific tags (3rd BC) with DNA barcodes respectively to obtain hydrogel microspheres with 96 3 specific tags.
[0076] Composition of the ligation probe (3rd Ligation Barcode, Round3_XX) for the specific tag (3rd BC) of DNA barcode: The bold 22nt is the PCR primer sequence, which is the same for all 96 wells; NNNNNNNNNN (10nt) is the random sequence, which is the UMI; the underlined 8nt part is the third segment barcode tag of the 3rd Ligation Barcode, which is different for each of the 96 wells; the last 15nt hybridizes with the 3rd barcode linker part, which is the same for all 96 wells).
[0077] Schematic diagram of the connection addition of the third new DNA barcode tag is shown in Figure 10 .
[0078] The composition of the 3rd ligation reaction solution is shown in Table 3.
[0079] Table 3 Ligation reaction solution
[0080] Composition 20 μl reaction system T4 DNA Ligase Buffer(10X) 2 μl 3rd Ligation Barcodes, Round3_xx 10 μM 3rd barcode linker 15 μM Hydrogel microspheres --- Nuclease-free water to 20 μl T4 DNA Ligase 1 μl
[0081] At this time, hydrogel microspheres with nearly one million (96 3 ) independent barcode tags can be obtained. When the number of microspheres is less than one hundred thousand, it can ensure that each cell is labeled with a unique cell tag, and at most 10,000 single-cell full-length transcriptomes can be cell-coded and labeled. At the same time, UMIs are used to molecularly label each transcriptome to ensure that each mRNA has a unique molecular tag. Schematic diagrams of the composition and structure of the complete coding are shown in Figure 11 .
[0082] Finally, under the action of a DNA polymerase with strand displacement activity (such as Bst DNA polymerase, etc.), these double-stranded cDNA short fragments are extended and displaced, and then amplified, 3'-end enriched library construction and sequencing are carried out. After the strand displacement reaction, there will be a part of overlapping sequences between adjacent two cDNA fragments, and the overlapping part can be used for the accurate splicing and assembly of full-length transcripts.
[0083] The recipe for the strand displacement reaction is as follows: 0.5 mM dNTPs, 0.2 U / mL Bst DNA polymerase, 1X ThermoPol Buffer.
[0084] 3. High-throughput single-cell full-length transcriptome data analysis
[0085] After the raw sequencing data is downloaded, data quality control and filtering, insert fragment and coding information extraction, cell clustering, sequence splicing and alignment correction, and single-cell full-length transcript analysis are carried out. The process is as shown in Figure 12 .
[0086] First, software such as FastQC and Fastp is used to analyze the quality of the raw sequencing data, and low-quality data and adapter sequences are removed. Next, software such as Dropseq-tools, UMI-tools, and Cell Ranger is used to extract information (cell encoding, molecular tags, and transcript fragment information) and classify the screened data, and pair the DNA coding sequences with the mRNA sequences. Subsequently, the sequencing data of each single cell is assembled. The principle of the assembly is as Figure 13 shown. There is an overlapping sequence (Overlap sequence) between adjacent fragments (Cn) of each full-length transcriptome amplification. Using this overlapping sequence, these sequencing Reads (Rm) can be assembled. Software such as STAR, TopHat2, and Bowtie2 is used to align and correct the assembly to form a highly accurate full-length transcriptome sequence. Finally, differential expression, transcript structure, and variation of full-length transcripts are analyzed at the single-cell level.
[0087] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for constructing a high-throughput single-cell full-length transcriptome sequencing library, characterized in that, It includes: Using droplet microfluidics technology to capture and encapsulate n cells to be analyzed, where n is a positive integer and ≥ 10, generating hydrogel microspheres containing single cells and mRNA capture probes, lysing the single cells in the hydrogel microspheres and performing mRNA reverse transcription to obtain the reverse transcription product of mRNA: full-length cDNA; Dividing the n hydrogel microspheres into x portions, where x is a positive integer and < n, using random primers to extend the full-length cDNA into multiple double-stranded cDNA fragments in each portion of the hydrogel microspheres, forming a hybridization chimera of full-length cDNA and double-stranded cDNA fragments, with different specific tags labeled on the random primers of each portion of the hydrogel microspheres, and then combining the x portions of hydrogel microspheres into one portion; Performing k steps of labeling new specific tags on the tags of all hydrogel microspheres to form multi-level tags, and stopping labeling when the types of the multi-level tags ≥ n; Using a polymerase with strand displacement activity to extend and displace the double-stranded cDNA fragments in all hydrogel microspheres, and amplifying the product obtained after displacement to obtain a sequencing library; Each step of labeling new specific tags includes: mixing the x portions of hydrogel microspheres and then re-dividing them into y portions, where y is a positive integer and < n, and then respectively labeling different specific tags on the tags of the random primers in each portion of the hydrogel microspheres to form multi-level specific tags; the distribution method of the y portions of hydrogel microspheres is average distribution; y is equal to x; The step of labeling different specific tags on the tags of the random primers in the hydrogel microspheres to form multi-level specific tags includes: adding a ligase and a ligation probe containing a tag sequence to the hydrogel microspheres; The ligation probe is a nucleic acid sequence, including 3 connected parts: a first sequence, a tag sequence, and a second sequence. The first sequence is a sequence for hybridizing with the existing ligation probe on the double-stranded cDNA fragment, and the second sequence is a sequence for hybridizing with the subsequent ligation probe; The step of labeling different specific tags on the tags of the random primers in the hydrogel microspheres to form multi-level specific tags further includes adding a linker to the hydrogel microspheres. The first sequence hybridizes with the existing ligation probe on the double-stranded cDNA fragment through the linker, and the second sequence hybridizes with the first sequence on the subsequent ligation probe through the linker; The reaction conditions for labeling new specific tags are as follows: the working concentration of the ligase is 1 - 5 unit / μL; the working concentration of the ligation probe is 0.5 - 15 μM, and the reaction time is 0.5 - 2.5 h; When performing the kth step of labeling new specific tags, the ligation probe also includes molecular tags UMIs.
2. The method for constructing a high-throughput single-cell full-length transcriptome sequencing library according to claim 1, wherein The working concentration of the random primer is such that the length of the extended double-stranded cDNA fragment is within the read length range of next-generation sequencing.
3. The method for constructing a high-throughput single-cell full-length transcriptome sequencing library according to claim 2, wherein The read length of next-generation sequencing is 300 - 500 bp.
4. The method for constructing a high-throughput single-cell full-length transcriptome sequencing library according to claim 2, wherein The working concentration of the random primer is 0.2 - 30 μM, and the length of the random primer is 6 - 10 nt.
5. The method for constructing a high-throughput single-cell full-length transcriptome sequencing library according to claim 1, wherein, When extending the full-length cDNA with the random primers, the construction method further includes adding dNTPs and ddNTPs to the extension system.
6. The method for constructing a high-throughput single-cell full-length transcriptome sequencing library according to claim 5, wherein The working concentrations of dNTPs and ddNTPs are such that they can be used to make the length of the double-stranded cDNA fragments within the read length range of next-generation sequencing.
7. The method for constructing a high-throughput single-cell full-length transcriptome sequencing library according to claim 5, wherein The working concentration of the dNTPs is 0.05 - 0.5 mM, and the working concentration of the ddNTPs is 0.05 - 0.5 mM.
8. The method for constructing a high-throughput single-cell full-length transcriptome sequencing library according to any one of claims 1 to 7, characterized in that, n is from 10 to 10 7 .
9. A sequencing method for high-throughput single-cell full-length transcriptome, characterized in that, It includes: Constructing a high-throughput single-cell full-length transcriptome sequencing library using the construction method of the high-throughput single-cell full-length transcriptome sequencing library according to any one of claims 1 - 8, and sequencing the sequencing library.
10. The sequencing method for high-throughput single-cell full-length transcriptome according to claim 9, wherein The sequencing includes splicing of the sequencing data, and the steps of the splicing are as follows: in the obtained sequencing library, there is a partial overlapping sequence between two adjacent double-stranded cDNA fragments amplified from the full-length transcriptome. The Reads are spliced using the overlapping sequence to obtain the sequencing analysis result.
11. The sequencing method of the high-throughput single-cell full-length transcriptome according to claim 9 or 10, characterized in that, The sequencing is next-generation sequencing.
12. Use of the composition in the preparation of a kit for high-throughput single-cell full-length transcriptome sequencing, characterized in that, The composition includes: reagents for implementing the construction method of the high-throughput single-cell full-length transcriptome sequencing library according to any one of claims 1 - 8 or the sequencing method of the high-throughput single-cell full-length transcriptome according to any one of claims 9 - 11.
13. A kit for high-throughput single-cell full-length transcriptome sequencing, characterized in that, It includes: Reagents for implementing the construction method of the high-throughput single-cell full-length transcriptome sequencing library according to any one of claims 1 - 8 or the sequencing method of the high-throughput single-cell full-length transcriptome according to any one of claims 9 - 11.
Citation Information
Patent Citations
DNA encoding microsphere and synthetic method thereof
CN105925572A
Kit for constructing human single-cell BCR sequencing library and application of kit
CN113026112A
Phenotypic and molecular characterisation of single cells
US20210317522A1