A single-cell whole-genome amplification sequencing method
By assembling the Tn5 transposase complex using specific adapter sequences and then performing ribonuclease cleavage, the problem of uneven amplification and high false positive rate in single-cell whole-genome amplification has been solved, achieving genome detection with high uniformity, high coverage, and high accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2026-02-13
- Publication Date
- 2026-05-29
Smart Images

Figure CN122104878A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of single-cell sequencing technology, and in particular to a method for single-cell whole-genome amplification and sequencing library construction. Technical Background
[0002] The cell is the basic structural and functional unit of living organisms. From invisible single-celled microorganisms to the diverse array of plants and animals, and even the most complex human systems, all are built upon the blueprint of life, constructed from cells. Multicellular organisms are composed of numerous single cells, whose morphology and function are complex and diverse, working together to maintain the normal activities of the organism. Disease or functional errors in a single or partial cell often lead to disease, disorder, or even death throughout the organism. For example, tumors often originate from gene mutations in a single cell, eventually developing into a highly heterogeneous and abnormally proliferating tumor tissue, which can then metastasize throughout the body, leading to death. The immense heterogeneity and complexity of single cells is a major bottleneck and significant challenge for research in many other fields, including developmental biology, stem cell biology, and biomedicine.
[0003] Single-cell omics technology has become the most powerful tool for comprehensively analyzing cellular functional states and molecular mechanisms. It has demonstrated enormous potential in elucidating the complex systems and functions of multicellular organisms and in understanding and preventing complex diseases such as tumors and Alzheimer's disease. The genome contains the genetic code of a species, and single-cell whole-genome sequencing technology can map the entire genome landscape of a single cell, directly analyzing variations at the genomic level. It has significant application value in fields such as reproductive health, tumor evolution, and microorganisms. However, the main bottleneck restricting the development of this technology is that the genomic DNA content in a single cell is only at the picogram level, requiring whole-genome amplification to obtain sufficient DNA to meet the input requirements for sequencing. Ensuring uniform and accurate amplification of the single-cell genome from the picogram level to the nanogram level suitable for sequencing is extremely challenging. The high uniformity and fidelity of whole-genome amplification (WGA) are necessary conditions for the accurate detection of genomic variations such as copy number variation (CNV) and single nucleotide variation (SNV). Three main types of errors can occur during the amplification process: First, single-cell lysis can cause DNA damage, leading to false positives in SNV detection. Second, polymerase during DNA amplification can introduce amplification errors, also resulting in false positives in SNV detection. Third, whole-genome amplification reactions are biased, resulting in uneven amplification, partial genome deletion, low genome coverage, and inaccurate CNV detection.
[0004] Currently, various single-cell whole-genome sequencing technologies have been developed, each with its own advantages and disadvantages. Degenerate oligonucleotide primer PCR (DOP-PCR) uses oligonucleotide primers containing degenerate sequences for whole-genome amplification based on PCR. This method suffers from uneven amplification, low genome coverage, and a high false-negative rate. Multiple substitution amplification (MDA) uses random primers to induce strand substitution reactions under the action of phi29 DNA polymerase, ultimately forming multi-branched amplification products, exhibiting significant amplification bias. Multiple annealing circular amplification (MALBAC) introduces quasi-linear pre-amplification to reduce bias associated with nonlinear amplification, but it still cannot overcome problems such as false positives introduced by amplification errors, amplification bias, and allele loss. Transposon insertion linear amplification (LIANTI) further reduces amplification bias by using linear amplification, improving the resolution of single-cell CNV detection to 10 kb, but the false positive rate for SNV detection remains at 5.4 × 10⁻⁶. -6 In single-cell whole-genome sequencing, approximately 30,000 false-positive SNVs are generated, while the number of de novo mutations in SNVs within a single cell typically ranges from several hundred to several thousand. Therefore, accurate detection of SNVs within single cells remains impossible. Multiple-terminal labeling (META-CS) technology specifically labels the ends of both DNA strands, enabling specific identification of each strand. False-positive SNVs can then be filtered out by checking the complementarity of the two strands. Although this method can accurately detect SNVs, its amplification uniformity is poor, making accurate detection of CNVs impossible. In summary, existing single-cell whole-genome sequencing methods still cannot achieve simultaneous and accurate detection of both SNVs and CNVs, significantly hindering the application of single-cell whole-genome sequencing technology.
[0005] To meet the evolving demands of precision medicine, it is necessary to develop single-cell whole-genome sequencing technologies with high uniformity, high fidelity, and accuracy. Improving the uniformity of single-cell whole-genome amplification, increasing genome coverage, and reducing false positive rates are crucial issues that must be addressed. Summary of the Invention
[0006] This invention addresses the shortcomings of existing technologies by providing a single-cell whole-genome amplification and sequencing method with high uniformity, high coverage, and high accuracy, which can simultaneously and accurately detect CNVs and SNVs in the single-cell whole genome.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] This invention provides a single-cell whole-genome amplification and sequencing method, which includes the following steps:
[0009] 1) Assemble a Tn5 transposase complex loaded with a specific adapter sequence, including the Tn5 transposase and the specific adapter sequence, wherein the specific adapter sequence is an incomplete DNA double strand, including long single strands and short single strands, and the long single strands include the following sequence arranged from the 5' end to the 3' end:
[0010] T1 sequence;
[0011] The Mosaic End sequence is used to bind to the Tn5 transposase.
[0012] The short single-stranded sequence is the inverse complementary sequence of the Mosaic End sequence;
[0013] 2) Separate single cells, add single-cell lysis buffer to fully lyse the single cells and digest the proteins bound to the genomic DNA; 3) Add a Tn5 transposase complex loaded with a specific adapter sequence to the single cells obtained in step (2) to fragment the DNA and add adapter sequences to both ends of it.
[0014] 4) Complete the fragmented DNA obtained in step (3) at the end;
[0015] 5) Add ribonuclease to recognize and cleave ribonucleotides on specific linker sequences;
[0016] 6) Add linear amplification primers and DNA polymerase to perform at least three single-primer PCR linear amplifications of the DNA fragment;
[0017] 7) Purify the product obtained in step (6);
[0018] 8) Add the chain extension primer and DNA polymerase to the product of step (7) to carry out the chain extension reaction, and add the T2 sequence to one end of the linear copy;
[0019] The chain extension primers comprise primers ordered from the 5' end to the 3' end:
[0020] The T2 sequence;
[0021] Mosaic end sequence.
[0022] 9) Purify the product obtained in step (8).
[0023] 10) Add library primer 1 and library primer 2 to the product obtained in step (9), and DNA polymerase to perform PCR reaction, and connect the adapters required for sequencing to both ends of the DNA fragment to be sequenced.
[0024] 11) Sequencing the product obtained in step (10) to obtain the whole genome sequence information of a single cell.
[0025] Furthermore, the assembly steps of the Tn5 transposase complex loaded with a specific adapter sequence include:
[0026] (a) Mix the long single chain and the short single chain in equal molar amounts and add them to a buffer solution. Heat the mixture to ≥90°C and then slowly cool it to room temperature to anneal and bond the long single chain and the short single chain.
[0027] (b) Take a certain volume of glycerol and add it to a centrifuge tube, then take the same volume of the specific connector sequence and add it to the centrifuge tube, and mix the glycerol and the specific connector sequence thoroughly;
[0028] (c) Take the specific adapter-glycerol mixture and add it to a new centrifuge tube. Take the same volume of Tn5 transposase and add it to the centrifuge tube. Mix thoroughly.
[0029] Further, in step (1), the Mosaic End sequence is as shown in SEQ ID NO.2, wherein at least one of the nucleotides at positions 1 to 7 of the 5' end of the sequence is a ribonucleotide, and the rest are deoxyribonucleotides;
[0030] The short single-stranded sequence is selected from any of the following:
[0031] (a) The sequence is shown in SEQ ID NO.3;
[0032] (b) Based on the 3' truncated variant shown in SEQ ID NO.3, the sequence length is 15-18 nucleotides, and the first 15 nucleotides of the 5' end shown in SEQ ID NO.3 are retained;
[0033] (c) When the 3' terminal nucleotide of the truncated variant shown in sequence SEQ ID NO.3 is C, it is replaced with dideoxycytidine.
[0034] Furthermore, in step (2), the single-cell lysate contains one or more of the following components: protease, proteinase K, SDS, guanidine thiocyanate, and guanidine hydrochloride, and the single-cell lysis time is preferably 30 h or more.
[0035] Furthermore, in step (5), the ribonuclease used is one or more of ribonuclease H, ribonuclease HII, ribonuclease I, ribonuclease If, ribonuclease A, ribonuclease T1, and ribonuclease R.
[0036] Further, in step (6), the linear amplification primer sequence consists of a T1 sequence and a MosaicEnd sequence from the 5' end to the 3' end; the DNA polymerase is Taq and its mutants, including one of Taq, Q5, Phusion, DeepVent, Pfu, KAPAHiFi, and Platinum SuperFi II.
[0037] Further, in step (8), the strand extension primer consists of a T2 sequence and a Mosaic End sequence from the 5' end to the 3' end, and the 1st to 12th nucleotide residues at the 3' end contain at least 1 nt of specially modified bases; the special modification is locked nucleic acid or 2-fluoroRNA modification; the DNA polymerase is Taq and its mutants, including one of Taq, Q5, Phusion, DeepVent, Pfu, KAPA HiFi, and Platinum SuperFi II.
[0038] Furthermore, the DNA product is purified using magnetic beads in steps (7) and (9).
[0039] Further, in step (10), the library construction primers are the PCR primers provided in the Illumina Nextera® Index Kit.
[0040] The library construction primers consist of primers ordered from the 5' end to the 3' end:
[0041] The T1 sequence is shown in SEQ ID NO.1;
[0042] I5 sequence (Illumina, Nextera® Index Kit)
[0043] The P5 sequence is shown in SEQ ID NO.7 (illumina, Nextera® Index Kit).
[0044] The second primer for library construction includes primers ordered from the 5' end to the 3' end:
[0045] The T2 sequence is as shown in SEQ ID NO.5;
[0046] I7 sequence (Illumina, Nextera® Index Kit)
[0047] The P7 sequence is shown in SEQ ID NO.8 (illumina, Nextera® Index Kit).
[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0049] 1) This invention uses a single adapter to assemble Tn5 transposase, fragmenting genomic DNA and ligating identical adapters to both ends. Each adapter is specifically modified with a ribonucleotide residue site for subsequent ribonuclease cleavage. Ribonuclease cleavage of the adapter sequence breaks the symmetry of the adapters at both ends, allowing the DNA fragment product to be amplified by single-primer PCR. This is a linear amplification method that greatly improves amplification uniformity and reduces amplification errors.
[0050] 2) Another advantage of using a single adapter to assemble Tn5 transposase in this invention is that all DNA fragments with adapter sequences after ribonuclease cleavage can be uniformly amplified. Compared to existing techniques that use multiple adapters to assemble Tn5 transposase for DNA fragmentation, this method does not produce DNA fragment products that cannot be amplified due to perfect complementarity at both ends. Therefore, this method can recover more DNA fragments and greatly improve whole-genome coverage.
[0051] 3) The present invention has a unique directionality when adding adapter sequences at both ends of the DNA fragment. Therefore, during sequencing analysis, the directionality of the adapter can be used to distinguish whether the product comes from one of the complementary double strands of the original DNA template. Thus, by comparing the double strands, false positive errors in SNV detection can be completely filtered out, achieving high accuracy in SNV detection. Attached Figure Description
[0052] Figure 1 This is a schematic diagram of the process for constructing a single-cell whole-genome amplification and sequencing library as described in this invention.
[0053] Figure 2 A schematic diagram of the structure of a Tn5 transposase complex loaded with a specific adapter sequence.
[0054] Figure 3 This diagram illustrates the process of fragmenting the genomic DNA of a single cell and adding adapters to both ends for transposition using the Tn5 transposase complex.
[0055] Figure 4 This diagram illustrates the process of recognizing and cleaving adapter sequences using ribonuclease and then performing single-primer PCR amplification of the DNA fragment using linear amplification primers.
[0056] Figure 5 This is a schematic diagram of the process of using extension primers to perform chain extension reactions and construct a complete sequencing library.
[0057] Figure 6 This is a sequence of human skin fibroblasts, yielding copy number variation profiles of the entire genome at different resolutions; among them, Figure 6In this context, 'a' represents the copy number variation spectrum of the entire genome at a resolution of 1 M bin size. Figure 6 In this context, b represents the copy number variation spectrum of the entire genome at a resolution of 5 K bin size. Figure 6 In this context, c represents the copy number variation spectrum of the entire genome at a resolution of 1K bin size.
[0058] Figure 7 To sequence human skin fibroblasts, the coefficient of variation of the detection copy number was obtained using the method proposed in this invention (labeled "ours") and three typical methods in the prior art ("lianti", "malbac", "meta-cs") at various bin sizes.
[0059] Figure 8 This is a specific SNV site obtained from sequencing human skin fibroblasts. Detailed Implementation
[0060] The present invention will be further described below with reference to embodiments. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not intended to limit the present invention. Any simple improvements to the preparation method of the present invention under the premise of the concept of the present invention are within the protection scope of the present invention. Experimental methods in the following embodiments that do not specify specific conditions are generally carried out according to well-known means in the art. Unless otherwise specified, the experimental materials used in the following embodiments were all purchased from conventional biochemical reagent stores.
[0061] This invention provides a single-cell whole-genome amplification and sequencing method, the process of which is as follows: Figure 1 As shown, it includes the following steps:
[0062] Step S10: Assemble the Tn5 transposase complex loaded with a specific adapter sequence;
[0063] Step S20: Isolate single cells, add single-cell lysis buffer to fully lyse single cells and digest proteins bound to genomic DNA;
[0064] Step S30: Add a Tn5 transposase complex loaded with a specific adapter sequence to a single cell to perform a transposition reaction; after transposition, complete the fragmented DNA ends;
[0065] Step S40: Add ribonuclease to recognize and cleave the specific adapter sequence;
[0066] Step S50: Add linear amplification primers and DNA polymerase to perform single-primer PCR amplification of the DNA fragment;
[0067] Step S60: Add strand extension primers and DNA polymerase to carry out the strand extension reaction;
[0068] Step S70: Connect the adapters required for sequencing to both ends of the DNA fragment to be sequenced, and construct the sequencing library;
[0069] In some embodiments, step S10 includes:
[0070] 1.1 Take a certain amount of the long single strand and the same molar amount of the short single strand from the artificially synthesized specific adapter sequence, mix them and add them to TE buffer solution. After mixing evenly, put them into a PCR instrument, heat them to a high temperature of 90°C, and then gradually cool them to room temperature to allow the long single strand and the short single strand to anneal and bind together to form a complete adapter sequence.
[0071] 1.2 Heat glycerol at 90°C, add a certain volume of the heated glycerol to a centrifuge tube, and then cool it on ice. Add the same volume of the above-mentioned connector sequence to the glycerol and mix thoroughly to obtain a mixed solution.
[0072] 1.3 Take a certain volume of the above mixed solution and add it to a new centrifuge tube. Take the same volume of Tn5 transposase and add it to the centrifuge tube. Mix thoroughly and assemble to obtain a Tn5 transposase complex loaded with a specific adapter sequence.
[0073] The structure of the Tn5 transposase complex loaded with a specific adapter sequence is as follows: Figure 2 As shown. The Tn5 transposase complex includes a Tn5 transposase and a specific adapter sequence, wherein the specific adapter sequence is an incomplete DNA double strand, including a long single strand and a short single strand; the long single strand includes a T1 sequence-Mosaic End sequence (for binding to the Tn5 transposase) arranged from the 5' end to the 3' end, and the short strand is the reverse complementary sequence of the Mosaic End sequence.
[0074] Among them, the T1 sequence is: 5'-TCGTCGGCAGCGTC-3' (SEQ ID NO.1);
[0075] Mosaic End sequence: 5'-AGATGTGTATAAGAGACAG-3' (SEQ ID NO.2); wherein, at least one of the nucleotides at positions 1 to 7 of the 5' end of the sequence is a ribonucleotide, and the rest are deoxyribonucleotides;
[0076] The nucleotide sequence of the Mosaic End sequence containing ribonucleotides used in the following embodiments of the present invention is shown in SEQ ID NO.3:
[0077] SEQ ID NO. 3: 5'-rArGATGTGTATAAGAGACAG-3'.
[0078] The inverse complementary sequence of the Mosaic End sequence is a 5' phosphorylated oligonucleotide, the nucleotide sequence of which is shown in SEQ ID NO.4:
[0079] SEQ ID NO. 4: 5'-pCTGTCTCTTATACACATCT-3'.
[0080] In some embodiments, the transposition reaction and the process of finishing the fragmented DNA ends described in step S30 are as follows: Figure 3 As shown, it includes:
[0081] 3.1 A Tn5 transposase complex loaded with a specific adapter sequence is added to a single cell to initiate a transposition reaction, which is used to fragment DNA and simultaneously add adapter sequences to both ends of it.
[0082] 3.2 Complete the ends of the fragmented DNA obtained in step 3.
[0083] In some embodiments, the process of performing single-primer PCR amplification of the DNA fragment using linear amplification primers in step S50 is as follows: Figure 4 As shown. The linear amplification primers are arranged from the 5' end to the 3' end as follows: T1 sequence - Mosaic End sequence; the T1 sequence is shown in SEQ ID NO.1, and the Mosaic End sequence is shown in SEQ ID NO.2.
[0084] In some embodiments, the chain extension reaction in step S60 is performed using chain extension primers as follows: Figure 5 As shown. The chain extension primers are arranged from the 5' end to the 3' end as follows: T2 sequence - Mosaic End sequence.
[0085] T2 sequence (SEQ ID NO.5): 5'-GTCTCGTGGGCTCGG-3';
[0086] The Mosaic End sequence is shown in SEQ ID NO.2; wherein the 3' end of the sequence contains at least 1 nt of a specially modified base in the 1st-12th nucleotide residues; the special modification is a locked nucleic acid or a 2-fluoroRNA modification;
[0087] The nucleotide sequences containing specially modified Mosaic End sequences used in the following embodiments of the present invention are shown in SEQ ID NO. 6:
[0088] SEQ ID NO.6: 5'-AGATGTGTATAAGAG / LNA-A / / LNA-C / AG-3'.
[0089] In some embodiments, step S70 involves connecting the adapters required for sequencing to both ends of the DNA fragment to be sequenced, thus constructing a complete sequencing library. Figure 5 As shown. Among them, the library construction primer one is ordered from the 5' end to the 3' end as: T1 sequence - I5 Index sequence - P5 sequence, and the library construction primer two is ordered from the 5' end to the 3' end as: T2 sequence - I7 Index sequence - P7 sequence.
[0090] T1 sequence: as shown in SEQ ID NO.1;
[0091] I5 Index sequence (illumina, Nextera® Index Kit);
[0092] P5 sequence (SEQ ID NO.7): 5'-AATGATACGGCGACCACCGAGATCTACAC-3' (illumina, Nextera® Index Kit);
[0093] T2 sequence: as shown in SEQ ID NO.5;
[0094] I7 Index sequence (illumina, Nextera® Index Kit);
[0095] P7 sequence (SEQ ID NO. 8): 5'-CAAGCAGAAGACGGCATACGAGAT-3' (illumina, Nextera® Index Kit).
[0096] Example 1: Amplification and sequencing of the whole genome of a single cultured cell.
[0097] 1) Reagent preparation:
[0098] Assemble a Tn5 transposase complex loaded with a specific adapter sequence, the structure of which is as follows: Figure 2As shown: (a) Take 10 μL of 100 μM of a long single strand from a specific synthetic adapter sequence (nucleotide sequence as shown in SEQ ID NO. 9), 10 μL of 100 μM of a short single strand from a specific synthetic adapter sequence (as shown in SEQ ID NO. 10), and 80 μL of TE buffer (10 mM Tris, 1 mM EDTA, pH 8.0) and put them into a 0.2 ml EP tube. Mix well and then put it into a PCR instrument. Set the following reaction program: 95℃ for 3 min, 70℃ for 3 min, (70℃ for 30 s, -1 ℃, per cycle) × 45, 25℃ forever, to anneal and bind the long and short single strands.
[0099] SEQ ID NO.9: 5'-TCGTCGGCAGCGTCArGrATGTGTATAAGAGACAG -3'
[0100] SEQ ID NO.10: 5'-pCTGTCTCTTATACACATddC-3'
[0101] (b) Heat glycerol to 95 degrees Celsius, add 10 μL of glycerol to a centrifuge tube, and then cool it on ice. Add 10 μL of the solution from step (a) of the PCR reaction to a centrifuge tube, and repeatedly pipette the solution to thoroughly mix the glycerol with the adapter sequence in the solution.
[0102] (c) Add 10 μL of the adapter-glycerol mixture to a new centrifuge tube, add 10 μL of Tn5 transposase to the centrifuge tube, and repeatedly pipette the solution to mix the enzyme and the adapter-glycerol mixture thoroughly.
[0103] 2) Isolation of individual human skin fibroblasts:
[0104] (a) Preparation of single-cell suspension: Human skin fibroblast cell line was purchased from Zhihe Tech. For adherent cells, the culture medium was aspirated from the culture dish, 2 mL of PBS solution was added to wash the cells, and the PBS was aspirated; 750 µL of trypsin was added, and the culture dish was placed in an incubator at 37 °C for 4 min to digest; the culture dish was removed, and 750 µL of culture medium was added to neutralize the trypsin; the digested cell suspension was collected into a 15 mL centrifuge tube and centrifuged at 300 g for 2 min; the supernatant was discarded, 5 mL of PBS was added, and the solution was pipetted to resuspend the cells; the cells were centrifuged at 300 g for 2 min, the supernatant was discarded, and 1 mL of fresh PBS was added to resuspend the cells.
[0105] (b) Picking single cells: Prepare 0.2 mL EP and add 1.5 µL of single cell lysis buffer to each tube; under a microscope, use a pipette to pick up single cells and place them into the EP tube containing the lysis buffer.
[0106] The formulation of single-cell lysis buffer is as follows:
[0107] 20mM Tris-Cl, 20mM NaCl, 2mM EDTA, 10mM DTT, 0.1% Triton X-100, 1.5mg / ml protease.
[0108] 3) Digestion: EP tubes containing lysis buffer were reacted at 37°C for 10 h to fully lyse single cells and digest proteins bound to genomic DNA.
[0109] 4) Transposition: Add 3 µL of 2×TD buffer, 0.4 µL of Tn5 transposase complex, and 1.1 µL of H2O to the single-cell lysis product obtained in step 3). Incubate at 37°C for 10 min; then add 1 µL of 0.2 M EDTA and 1 µL of 1% SDS to terminate the reaction, and incubate at 37°C for 30 min to release Tn5 transposase; then add 1 µL of 0.2 M MgCl2 and 1 µL of 20% Triton X-100.
[0110] 5) End-completion: Add 1 µL of 10 mM dNTP and 0.5 μL of Bst DNA polymerase to the product obtained in step 4), react at 37 °C for 10 min, and then react at 72 °C for 30 min.
[0111] 6) Add 2 μL of RNase H to the product obtained in step 5) and react at 37°C for 2 h.
[0112] 7) Add 1 μL of linear amplification primer (as shown in SEQ ID NO.11) and 15 μL of 2 × Rapid Taq Master Mix (Vazyme) to the product obtained in step 6). Set the reaction program as follows: 95℃ for 1 min, (95℃ for 10 s, 72℃ for 1 min) x6, 4℃ forever to perform single primer PCR amplification and amplify the DNA fragment by linear amplification.
[0113] SEQ ID NO.11: 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG -3'
[0114] 8) Purify the amplification product obtained in step 7) using AMPure XP magnetic beads, and wash the nucleic acid product with 20 μL of nuclease-free water.
[0115] 9) Add 5 μL of 10X TP, 1 μL of 10 mM dNTP, 1 μL of 50 μM strand extension primer, 0.5 μL of Deep Vent DNA polymerase and 22.5 μL of H2O to the product obtained in step 8). Set the reaction program as follows: 95℃ for 1 min, 52℃ for 30 s, 72℃ for 30 s, 72℃ for 5 min, 4℃ forever to perform the strand extension reaction and add the T2 sequence (nucleotide sequence as shown in SEQ ID NO.5) to the 3' end of the DNA fragment.
[0116] 10) Purify the product obtained in step 9) using AMPure XP magnetic beads, and wash the nucleic acid product with 20 μL of nuclease-free water.
[0117] 11) Add 5 μL of 10X TP, 1 μL of 10 mM dNTP, 0.5 μL of 10 μM library primer 1 (nucleotide sequence as shown in SEQ ID NO. 12) and 0.5 μL of 10 μM library primer 2 (nucleotide sequence as shown in SEQ ID NO. 13), 0.5 μL of Deep Vent DNA polymerase and 22.5 μL of H2O to the product obtained in step 10). Set the reaction program as follows: 95℃ for 1 min, (95℃ for 10 s, 55℃ for 30 s, 72℃ for 30 s) X12, 72℃ for 5 min, 10℃ forever, and perform PCR reaction to ligate the adapters required for sequencing to both ends of the DNA fragment to be sequenced.
[0118] SEQ ID NO.12: 5'-AATGATACGGCGACCACCGAGATCTACACTAGATCGCTCGTCGGCAGCGTC-3'
[0119] SEQ ID NO.13: 5'-CAAGCAGAAGACGGCATACGAGATTAAGGCGAGTCTCGTGGGCTCGG-3'
[0120] 12) Sequencing the product obtained in step 11) to obtain the whole genome sequence information of a single cell.
[0121] Example 2: Amplification and sequencing of the whole genome of a single cell in a tumor tissue sample.
[0122] 1) Reagent preparation:
[0123] Assemble a Tn5 transposase complex loaded with a specific adapter sequence, the structure of which is as follows: Figure 2 As shown: (a) Take 5 μL of 100 μM artificially synthesized long single strand (nucleotide sequence as shown in SEQ ID NO.14), 5 μL of 100 μM artificially synthesized short single strand (as shown in SEQ ID NO.15), and 80 μL of TE buffer into a 0.2 ml EP tube, mix well, and then place it into a PCR instrument. Set the following reaction program: 95℃ for 3 min, 70℃ for 3 min, (70℃ for 30 s, -1 ℃, per cycle) × 45, 25℃ forever, to anneal and bind the long and short single strands.
[0124] SEQ ID NO.14: 5'-TCGTCGGCAGCGTCAGrATGTGTATAAGAGACAG-3'
[0125] SEQ ID NO.15: 5'-pCTGTCTCTTATACACA-3'
[0126] (b) Heat glycerol to 95 degrees Celsius, add 5 μL of glycerol to a centrifuge tube, and then cool it on ice. Add 10 μL of the solution from step (a) of the PCR reaction to a centrifuge tube, and repeatedly pipette the solution to thoroughly mix the glycerol with the adapter sequence in the solution.
[0127] (c) Add 5 μL of the adapter-glycerol mixture to a new centrifuge tube, then add 5 μL of Tn5 transposase to the centrifuge tube. Use a pipette tip to repeatedly pipette and agitate the solution to ensure thorough mixing of the enzyme and the adapter-glycerol mixture. 2) Isolate single cells:
[0128] (a) Preparation of single-cell suspension: Cryopreserved human glioma tumor tissue samples were obtained from the hospital. The glioma tumor tissue blocks were removed and placed in a culture dish containing pre-chilled PBS. The tissue blocks were repeatedly rinsed with PBS to remove the fascia and blood clots on the surface of the tissue. The culture dish containing the tissue was kept on ice during the operation. The washed tissue blocks were transferred to 1.5 mL centrifuge tubes, 400 μL of pre-chilled PBS was added, and the tissue blocks were quickly cut into 2-4 mm pieces with small scissors. 3The tissue fragments were removed; 800 μL of pre-cooled PBS was added, and the mixture was centrifuged at 4°C (300 g, 5 min), and the supernatant was discarded; 3 mL of pre-warmed trypsin hydrolysate at 37°C was added, and the tissue fragments and hydrolysate were transferred to a new 15 mL centrifuge tube using a wide-mouth pipette tip. The tube was sealed with sealing film and incubated at 37°C on a horizontal rotating shaker with gentle shaking at 90 rpm for 60 min. (Due to the significant differences in digestion conditions for different tissues, it is recommended to set different digestive enzyme formulations and digestion time gradients in the preliminary experiments. Commonly used digestive enzymes include various collagenases, hyaluronidases, elastases, pepsin, trypsin, DNases, etc. This example only uses glioma tissue samples as an example.) The centrifuge tube was removed, and complete cell culture medium was added at a volume ratio of 1:5 to terminate the digestion. The tube was gently inverted to mix and allowed to stand for 1-2 min.
[0129] According to experimental requirements, a 100 μm / 70 μm / 40 μm cell sieve was used to filter the digested cell suspension. The centrifuge tubes and cell sieve were rinsed with PBS, and the filtrate was collected. The cells were centrifuged at 500 g for 8 min at room temperature, and the supernatant was removed. 1 mL of PBS was added, and the cells were gently resuspended. The single-cell suspension separation effect and the proportion of red blood cells were evaluated by microscopic examination. Finally, the cell number and cell viability were measured.
[0130] (b) Picking single cells: Same as step (2b) in Example 1.
[0131] 3) Digestion: EP tubes containing lysis buffer were reacted at 37°C for 5 h to fully lyse single cells and digest proteins bound to genomic DNA.
[0132] 4) Transposition: Add 3 µL of 2×TD buffer, 0.4 µL of Tn5 transposase complex, and 1.1 µL of H2O to the single-cell lysate obtained in step 3). Incubate at 37°C for 30 min; then add 1 µL of 0.2 M EDTA and 1 µL of 1% SDS to terminate the reaction, and incubate at 37°C for 30 min to release Tn5 transposase; then add 1 µL of 0.2 M MgCl2 and 1 µL of 20% Triton X-100.
[0133] 5) End-completion: Add 1 µL of 10 mM dNTP and 0.5 μL of Bst DNA polymerase to the product obtained in step 4), react at 37 °C for 5 min, and then react at 72 °C for 30 min.
[0134] 6) Add 2 μL of RNase H to the product obtained in step 5) and react at 37°C for 5 h.
[0135] 7) Add 1 μL of linear amplification primer (as shown in SEQ ID NO.11) and 15 μL of 2 × Universal Taq Master Mix (iCloning) to the product obtained in step 6). Set the reaction program as follows: 95℃ for 1 min, (95℃ for 10 s, 72℃ for 1 min) ×6, 4℃ forever to perform single primer PCR amplification and amplify the DNA fragment by linear amplification.
[0136] 8) Purify the amplification product obtained in step 7) using AMPure XP magnetic beads, and wash the nucleic acid product with 20 μL of nuclease-free water.
[0137] 9) Add 5 μL of 10X TP, 1 μL of 10 mM dNTP, 1 μL of 50 μM strand extension primer, 0.5 μL of Deep Vent DNA polymerase and 22.5 μL of H2O to the product obtained in step 8). Set the reaction program as follows: 95℃ for 1 min, 52℃ for 30 s, 72℃ for 30 s, 72℃ for 5 min, 4℃ forever to perform the strand extension reaction and add the T2 sequence (nucleotide sequence as shown in SEQ ID NO.5) to the 3' end of the DNA fragment.
[0138] 10) Purify the product obtained in step 9) using AMPure XP magnetic beads, and wash the nucleic acid product with 20 μL of nuclease-free water.
[0139] 11) Add 5 μL of 10X TP, 1 μL of 10 mM dNTP, 0.5 μL of 10 μM library primer 1 (nucleotide sequence as shown in SEQ ID NO.16) and 0.5 μL of 10 μM library primer 2 (nucleotide sequence as shown in SEQ ID NO.17), 0.5 μL of Deep Vent DNA polymerase and 22.5 μL of H2O to the product obtained in step 10). Set the reaction program as follows: 95℃ for 1 min, (95℃ for 10 s, 55℃ for 30 s, 72℃ for 30 s) ×12, 72℃ for 5 min, 10℃ forever, and perform PCR reaction to ligate the adapters required for sequencing to both ends of the DNA fragment to be sequenced.
[0140] SEQ ID NO.16: 5'-AATGATACGGCGACCACCGAGATCTACACGTAAGGAGTCGTCGGCAGCGTC-3'
[0141] SEQ ID NO.17: 5'-CAAGCAGAAGACGGCATACGAGATTAGGCATGGTCTCGTGGGCTCGG-3'
[0142] 12) Sequencing the product obtained in step 11) to obtain the whole genome sequence information of a single cell.
[0143] Example 3: Analysis of single-cell whole-genome amplification and sequencing results.
[0144] The sequencing results of human skin fibroblasts in Example 1 were analyzed, and the analysis steps included...
[0145] 1) Adapter cutting: The sequencing adapter sequence is cut at the 3' end of R1 and R2 using cutadapt.
[0146] 2) Merging paired-end sequences: Use FLASH2 to merge paired-end sequences with an overlap greater than 8bp into single-end sequences.
[0147] 3) Alignment with reference genome: Use bwa mem to align the results of (1) with the reference genome hs37d5 and the results of (2) with the reference genome hg38.
[0148] 4) Quality control filtering: Use tools such as samtools to compare and filter the results, retaining reads with MAPQ=60, removing reads with excessive hard or soft shearing, and removing reads in the Blacklist region.
[0149] 5) Copy number detection and CV plotting: After deduplication of reads in the bam file aligned to hg38 in (3), the copy number of each bin region was obtained using cnvkit. Then, the DNAcopy CBS algorithm was used to re-segment the genome to obtain the copy number variation spectrum of the entire genome at 1 M bin size resolution, 5 K bin size resolution and 1 K bin size resolution respectively (see [link to documentation]). Figure 6 The chromosome copy number of human skin fibroblasts obtained by this method is mostly 2, which is consistent with the copy number characteristics of normal human diploid cells, except for a few special regions of the genome and sex chromosomes. This indicates that our method accurately detects the true copy number of genomic DNA in single cells, and that the amplification is uniform.
[0150] 6) Compare the method proposed in this invention (labeled "ours") with three typical methods in the prior art ("lianti", "malbac", "meta-cs"), and adjust the bin size parameter, with a value range covering 1 to 10. 8 For each bin size, the coefficient of variation (CV) of the detected copy number was calculated to characterize the dispersion of the data distribution. A lower CV value indicates higher data consistency, i.e., higher amplification uniformity. See the comparison results below. Figure 7 This indicates that our method amplifies uniformity similarly to the best meta-CS method.
[0151] 7) Single nucleotide variant detection: The BAM file aligned to hs37d5 in (3) is detected using the algorithm for detecting single nucleotide variants in meta-cs. False positives are identified by chain specificity. Figure 8 This demonstrates a definitive SNV site obtained after removing false-positive SNV sites.
[0152] The above are merely preferred embodiments of the present invention and do not limit the patent scope of the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the patent protection scope of the present invention.
Claims
1. A method for single-cell whole-genome amplification and sequencing, characterized in that, Includes the following steps: (1) Assemble a Tn5 transposase complex loaded with a specific adapter sequence, including a Tn5 transposase and a specific adapter sequence, wherein the specific adapter sequence is an incomplete DNA double strand, including a long single strand and a short single strand; (2) Isolate single cells, add single cell lysis buffer to fully lyse the single cells and digest the proteins bound to the genomic DNA; (3) Add a Tn5 transposase complex loaded with a specific adapter sequence to a single cell to perform a transposition reaction, which is used to fragment DNA and add adapter sequences to both ends of it at the same time. (4) Complete the ends of the fragmented DNA obtained in step (3); (5) Add ribonuclease to recognize and cleave the ribonucleotide sequence in the specific adapter sequence; (6) Add linear amplification primers and DNA polymerase to perform at least three single-primer PCR linear amplifications of the DNA fragment; (7) Purify the PCR product from step (6); (8) Add the chain extension primer and DNA polymerase to the product of step (7) to carry out the chain extension reaction and add the T2 sequence to one end of the linear copy; (9) Purify the product obtained in step (8); (10) Add library primer 1 and library primer 2 to the product obtained in step (9), and DNA polymerase to perform PCR reaction, and connect the adapters required for sequencing to both ends of the DNA fragment to be sequenced. (11) Sequencing the product obtained in step (10) to obtain the whole genome sequence information of a single cell.
2. The single-cell whole-genome amplification and sequencing method according to claim 1, characterized in that, Step (1) includes the following sub-steps: (1.1) Mix the long single chain and the short single chain in equal molar amounts and add them to a buffer solution. First heat to ≥90°C, then slowly cool to room temperature to allow the long single chain and the short single chain to anneal and bond together to form a complete linker sequence. (1.2) Take the same volume of glycerol and the connector sequence and mix them to obtain a mixed solution; (1.3) Take the mixed solution described in step (1.2) and mix it thoroughly with an equal volume of Tn5 transposase to assemble a Tn5 transposase complex loaded with a specific adapter sequence.
3. The single-cell whole-genome amplification and sequencing method according to claim 1, characterized in that, In step (1), the long single-stranded sequence from the 5' end to the 3' end is composed of a T1 sequence and a Mosaic End sequence; the short single-stranded sequence is the reverse complementary sequence of the Mosaic End sequence.
4. The single-cell whole-genome amplification and sequencing method according to claim 3, characterized in that, The Mosaic End sequence is shown in SEQ ID NO.2, wherein at least one of the nucleotides at positions 1 to 7 of the 5' end of the sequence is a ribonucleotide, and the rest are deoxyribonucleotides; The short single-stranded sequence is selected from any of the following: (a) The sequence is shown in SEQ ID NO.3; (b) Based on the 3' truncated variant shown in SEQ ID NO.3, the sequence length is 15-18 nucleotides, and the first 15 nucleotides of the 5' end shown in SEQ ID NO.3 are retained; (c) When the 3' terminal nucleotide of the truncated variant shown in sequence SEQ ID NO.3 is C, it is replaced with dideoxycytidine.
5. The single-cell whole-genome amplification and sequencing method according to claim 1, characterized in that, In step (2), the single-cell lysate contains one or more of the following components: protease, proteinase K, SDS, guanidine thiocyanate, and guanidine hydrochloride.
6. The single-cell whole-genome amplification and sequencing method according to claim 1, characterized in that, In step (5), the ribonuclease includes one or more of ribonuclease H, ribonuclease HII, ribonuclease I, ribonuclease If, ribonuclease A, ribonuclease T1, and ribonuclease R.
7. The single-cell whole-genome amplification and sequencing method according to claim 1, characterized in that, In step (6), the linear amplification primer sequence consists of a T1 sequence and a Mosaic End sequence from the 5' end to the 3' end; the DNA polymerase is Taq and its mutants, including one of Taq, Q5, Phusion, DeepVent, Pfu, KAPA HiFi, and Platinum SuperFi II.
8. The single-cell whole-genome amplification and sequencing method according to claim 1, characterized in that, In step (8), the strand extension primer consists of a T2 sequence and a Mosaic End sequence from the 5' end to the 3' end, and the 1st to 12th nucleotide residues at the 3' end contain at least 1 nt of a specially modified base; the special modification is locked nucleic acid or 2-fluoroRNA modification; the DNA polymerase is Taq and its mutants, including one of Taq, Q5, Phusion, DeepVent, Pfu, KAPA HiFi, and Platinum SuperFi II.
9. The single-cell whole-genome amplification and sequencing method according to claim 1, characterized in that, In step (10), the first primer for library construction consists of the P5 sequence, the I5 sequence, and the T1 sequence arranged from the 5' end to the 3' end; the second primer for library construction consists of the P7 sequence, the I7 sequence, and the T2 sequence arranged from the 5' end to the 3' end.
10. The application of the single-cell whole-genome amplification and sequencing method according to any one of claims 1-9 in the sequencing of single cells, single bacteria, single viruses, etc.