Primer group, kit and sequencing method for constructing single cell whole genome sequencing library
By using random primers and amplified primers, combined with pre-amplification and directional amplification techniques, a single-cell whole genome sequencing library is directly constructed, solving the problems of deviation, time and cost of sequencing results in the prior art, and achieving efficient and accurate single-cell whole genome sequencing.
Patent Information
- Application Number
- CN202311504074.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-10
- Publication Date
- 2025-05-13
AI Technical Summary
The existing ONT single-cell whole-genome sequencing technology relies on single-cell amplification methods, resulting in high deviations in sequencing results, time and cost, and an increased risk of sample contamination.
Provide a primer set for constructing a single-cell whole genome sequencing library, including random primers and amplification primers. By combining pre-amplification and directional amplification, the sequencing library is directly constructed, reducing the library construction steps and time.
It improves the accuracy and efficiency of single-cell whole genome sequencing, reduces costs, simplifies experimental steps, reduces the risk of sample contamination, and improves the coverage and uniformity of sequencing data.
Smart Images

Figure CN119979528A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of gene detection technology, and in particular, to a primer set, a kit and a sequencing method for constructing a single-cell whole genome sequencing library. Background Art
[0002] Single cells are the basic units of multicellular organisms. The continuous differentiation of mature organisms during development leads to great differences between cells, especially tumor tissues. The high heterogeneity of tumor tissues requires that related research be carried out at the single-cell level.
[0003] Single-cell sequencing refers to the amplification and sequencing of transcriptomes or genomes at the level of individual cells to detect data from multiple omics such as genomics (structural variations-SVs; copy number variations-CNVs; single nucleotide variants-SNVs, etc.), transcriptomics (RNA expression levels; alternative splicing of transcripts), epigenomics (DNA methylation, etc.), and proteomics. The outstanding advantage of single-cell sequencing technology is that it can detect cell specificity and differences between cells from the perspective of cell maps, explore the collaborative operation mode between cells, and study tissue heterogeneity. Therefore, based on the characteristics at the cellular level, it can deeply explore the cell state under healthy physiology and disease abnormalities, laying the foundation for understanding the functional mechanism of organisms and discovering effective disease treatment methods.
[0004] ONT (Oxford Nanopore Technologies) sequencing is the third generation of single-molecule real-time electrical signal sequencing technology based on nanopores. It does not require PCR amplification, avoiding the errors introduced by PCR amplification. For single-cell genome sequencing, it has the advantages of high throughput, high accuracy, and long read length. Based on these technical characteristics, third-generation sequencing plays an increasingly important role in genetic disease screening, tumor screening, microbial sequencing, and drug target design, especially in genome methylation sequencing. The conventional ONT sequencing process mainly includes nucleic acid extraction, library construction, on-machine sequencing, and bioinformatics analysis. Among them, library construction is a key step. The quality of library construction is directly related to the success or failure of sequencing and the accuracy of the results. At the same time, the speed of library construction affects the efficiency of sequencing.
[0005] When running single-cell sequencing technology on the ONT platform, the conventional method is to first amplify the whole genome of a single cell to obtain a uniform and complete genome, and then use a third-party library construction kit to construct a library of the single-cell whole genome amplification product. The obtained library is sequenced on the ONT platform to reveal the differences between individual cells or groups.
[0006] Traditional ONT single-cell sequencing technology is highly dependent on single-cell whole genome amplification technology. The uniformity and stability of products obtained by different single-cell amplification methods are different, resulting in large deviations in downstream sequencing results. In addition, traditional ONT single-cell sequencing technology separates single-cell amplification from library construction, which is a lengthy process, resulting in a significant increase in single-cell sequencing time and cost. The complicated experimental operations also increase the risk of sample contamination.
[0007] In summary, how to simultaneously reduce sequencing costs and improve sequencing accuracy is the difficulty of ONT single-cell whole genome sequencing. Summary of the invention
[0008] In order to solve the above problems, reduce the cost of ONT single-cell sequencing, and improve the versatility of amplification of the whole genome or trace genome of a single cell, the present application provides a primer set for constructing a single-cell whole genome sequencing library, the primer set comprising a random primer and an amplification primer, the sequence of the random primer is shown in SEQ ID No. 1, the sequence of the amplification primer is partially identical to the sequence of the random primer, the random primer is used to amplify the whole genome DNA, and the amplification primer is used to amplify the amplification product of the random primer.
[0009] In one embodiment, the primer set satisfies at least one of the following characteristics:
[0010] The sequence of the amplification primer includes the sequence shown in SEQ ID No. 2;
[0011] The sequence of the amplification primer also includes an adapter connection sequence, and the adapter connection sequence is complementary to at least a portion of the sequence of the ONT adapter;
[0012] The sequence of the amplification primer also includes a tag sequence, and the tag sequence is a sequence composed of random bases. Optionally, the number of random bases in the tag sequence is 22 to 26;
[0013] The sequence of the amplification primer is shown in SEQ ID No. 3;
[0014] The working concentration of the random primers is 1 μM to 5 μM; and the working concentration of the amplification primers is 1 μM to 5 μM.
[0015] The present application also provides a kit for constructing a single-cell whole genome sequencing library, the kit comprising a pre-amplification reagent and an amplification reagent, the pre-amplification reagent comprising the above-mentioned random primers, and the amplification reagent comprising the above-mentioned amplification primers.
[0016] In one embodiment, the pH value of the pre-amplification reagent is 8-9, and the pre-amplification reagent also includes a pre-amplification enzyme, and the working concentration of the pre-amplification enzyme is 0.01U-0.5U; optionally, the pre-amplification enzyme includes Taq DNA polymerase.
[0017] In one embodiment, the preamplification reagent further contains at least one of Tris-HCl, MgCl2, (NH4)2SO4, Triton-X100, DMSO, BSA and dNTP. Optionally, the preamplification reagent includes one or more of the following features: the concentration of Tris-HCl in the preamplification reagent is 10mM to 50mM; the concentration of MgCl2 in the preamplification reagent is 1mM to 4mM; the concentration of (NH4)2SO4 in the preamplification reagent is 10mM to 50mM; the volume fraction of Triton-X100 in the preamplification reagent is 0.1%-0.5%; the volume fraction of DMSO in the preamplification reagent is 1% to 10%; the concentration of BSA in the preamplification reagent is 0.01mg / mL to 0.1mg / mL; the sum of the base concentrations of dNTPs in the preamplification reagent is 0.1mM to 0.5mM.
[0018] In one embodiment, the pH value of the amplification reagent is 8-9, and the amplification reagent also includes an amplification enzyme, the amplification enzyme has a polymerization center and an enzyme cleavage center, and its working concentration is 1U-5U; optionally, the amplification enzyme includes a high-fidelity DNA polymerase.
[0019] In one embodiment, the amplification reagent further contains at least one of Tris-HCl, MgCl2, KCl, (NH4)2SO4, Triton-X100, DMSO, BSA and dNTP. Optionally, the amplification reagent includes one or more of the following features: the concentration of Tris-HCl in the amplification reagent is 10mM to 50mM;
[0020] The concentration of MgCl2 in the amplification reagent is 1mM to 4mM; the concentration of KCl in the amplification reagent is 10mM to 50mM; the concentration of (NH4)2SO4 in the amplification reagent is 10mM to 50mM; the volume fraction of Triton-X100 in the amplification reagent is 0.1%-0.5%; the volume fraction of DMSO in the amplification reagent is 1% to 10%; the concentration of BSA in the amplification reagent is 0.01mg / mL to 0.1mg / mL; the sum of the base concentrations of dNTPs in the amplification reagent is 0.1mM to 0.5mM.
[0021] In one embodiment, the kit further comprises a cell lysis reagent, the pH value of the cell lysis reagent is 8 to 9, and comprises at least one of a cell lysis enzyme, Tris-HCl, NP-40 and Triton X-100. Optionally, the cell lysis reagent satisfies at least one of the following characteristics: the concentration of Tris-HCl in the cell lysis reagent is 10mM to 50mM; the volume fraction of NP-40 in the cell lysis reagent is 0.1% to 0.5%; the volume fraction of Triton X-100 in the cell lysis reagent is 0.1% to 0.5%; the cell lysis enzyme comprises proteinase K, and the working concentration of the cell lysis enzyme is 10μg / mL to 30μg / mL.
[0022] In one embodiment, the kit further comprises at least one of an end repair reagent and an adapter ligation reagent; the end repair reagent is used to add an adenine fragment to the end of the DNA fragment; and / or the adapter ligation reagent contains an ONT adapter.
[0023] The present application also provides a single-cell whole genome sequencing method, comprising the following steps: using the above-mentioned kit to construct a sequencing library of a single-cell whole genome DNA sample; and performing ONT sequencing on the sequencing library.
[0024] In one embodiment, constructing a sequencing library of a single-cell whole genome DNA sample includes the following steps: pre-amplifying the single-cell whole genome DNA sample using a pre-amplification reagent to obtain a pre-amplification product; amplifying the pre-amplification product using an amplification reagent to obtain an amplification product; and optionally, performing end repair on the amplification product and connecting it to an ONT connector to obtain a single-cell whole genome sequencing library.
[0025] In one embodiment, before pre-amplifying the single-cell whole genome DNA sample using a pre-amplification reagent, the method further includes: lysing the single-cell sample using a cell lysis reagent to obtain the single-cell whole genome DNA sample. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0027] Figure 1 A schematic diagram of a method for constructing a single gene whole genome sequencing library provided in an embodiment of the present application;
[0028] Figure 2Schematic diagram of GC distribution of three libraries provided in Example 1 of the present application;
[0029] Figure 3 The three library Lorentz diagrams provided in Example 1 of the present application;
[0030] Figure 4 A comparison chart of chromosome dispersion of the S1 library sequencing results provided in Example 2 of the present application;
[0031] Figure 5 This is a comparison chart of chromosome dispersion of the S2 library sequencing results provided in Example 2 of the present application. DETAILED DESCRIPTION
[0032] References to embodiments of the present application will now be provided in detail, one or more examples of which are described below. Each example is provided as an explanation rather than a limitation of the present application. In fact, it will be apparent to those skilled in the art that various modifications and variations may be made to the present application without departing from the scope or spirit of the present application. For example, a feature described or described as part of one embodiment may be used in another embodiment to produce a further embodiment.
[0033] Therefore, it is intended that the present application covers such modifications and variations that fall within the scope of the appended claims and their equivalents. Other objects, features and aspects of the present application are disclosed in or are apparent from the following detailed description. It will be appreciated by those of ordinary skill in the art that this discussion is merely a description of exemplary embodiments and is not intended to limit the broader aspects of the present application.
[0034] As used herein, the term "primer" refers to an oligonucleotide, whether naturally present in a purified restriction digest or produced synthetically, which can act as a starting point for synthesis when the oligonucleotide is placed under conditions that induce the synthesis of a primer extension product complementary to a nucleic acid chain (e.g., in the presence of nucleotides and an inducing agent such as a DNA polymerase and at a suitable temperature and pH). The primer is preferably single-stranded to increase the maximum efficiency of amplification, but may alternatively be double-stranded. If double-stranded, the primer is first treated to separate its double strands before being used to prepare an extension product. Preferably, the primer is an oligodeoxyribonucleotide. The primer should be long enough, or the number of nucleotides contained in the primer is sufficient to initiate the synthesis of an extension product in the presence of an inducing agent. The exact length of the primer depends on many factors, including temperature, primer source, and method use. For example, in some embodiments, primers range from 10 to 100 or more nucleotides (e.g., 10-300, 15-250, 15-200, 15-150, 15-100, 15-90, 20-80, 20-70, 20-60, 20-50 nucleotides, etc.).
[0035] As mentioned above, traditional single-cell sequencing technology is highly dependent on single-cell whole genome amplification technology, and the sequencing cost is high. The uniformity and stability of its amplification determine the quality of downstream sequencing results. Therefore, traditional single-cell genome sequencing methods are expensive and time-consuming, with large deviations in sequencing results and poor versatility and flexibility.
[0036] To this end, the present application performs random amplification of a trace amount of single-cell genome to obtain a high-concentration product, and then performs ONT library construction and sequencing, thereby organically combining the single-cell amplification method and the ONT library construction method, and directly constructs a library for the whole genome of a single cell obtained by random primer amplification. The obtained amplified product can be used for third-generation sequencing, which greatly shortens the library construction and sequencing time, thereby quickly realizing the whole genome sequence analysis of a single cell and meeting the requirements for rapid conversion of single-cell sequencing products.
[0037] Specifically, the present application provides a primer set, which includes a random primer and an amplification primer. The sequence of the random primer is shown in SEQ ID No. 1, the sequence of the amplification primer is partially identical to the sequence of the random primer, the random primer is used to amplify genomic DNA, and the amplification primer is used to amplify the amplification product of the random primer to achieve enrichment of trace genome or single-cell genome DNA and meet the library construction concentration requirements for constructing a sequencing library.
[0038] Among them, SEQ ID No.1:
[0039] TGTGTTGGGTGTGTTTGGNNNNNKKKK,
[0040] In the sequence, N represents A, G, C or T, and K represents T or G;
[0041] The single cell may be a single cell from any tissue in an animal or human body, such as a single cell extracted from a tissue, a circulating tumor cell, a circulating fetal cell, etc.
[0042] The primer set of the present application performs pre-amplification and directional amplification on the whole genome of a single cell, and completes part of the work of pre-amplification product library construction during directional amplification, that is, library DNA fragments are enriched to meet the library construction concentration requirements, and the amplification step is omitted in the library construction process of the directional amplification product. It is only necessary to connect the ONT connector at the end of the directional amplification product to complete the entire ONT library construction process, thereby reducing the ONT library construction time, so that the ONT sequencing analysis of the whole genome of a single cell can be quickly realized, meeting the requirements for rapid conversion of single-cell sequencing products.
[0043] Compared with the sequencing libraries constructed by combining the traditional single-cell whole genome amplification methods such as the multiple displacement amplification MDA method and the multiple annealing replacement amplification MALBAC method with the library construction method, the method for constructing a single-cell whole genome sequencing library in the present application organically combines the single-cell amplification method with the ONT library construction method, omitting the step of amplifying the obtained single cell whole genome amplification product in the traditional library construction process, simplifying the ONT single-cell whole genome sequencing steps, avoiding the dependence of the sequencing results on the whole genome amplification method, greatly shortening the library construction and sequencing time, and having the characteristics of extremely simplified library construction process, good library construction quality, and relatively low cost, while improving the coverage and uniformity of single-cell whole genome sequencing data.
[0044] In some specific embodiments, the amplification sequence of the amplification primer is as shown in SEQ ID No. 2, and is used to complementarily pair with at least a portion of the sequence of the amplicon of the random primer, thereby amplifying the pre-amplification product;
[0045] Among them, SEQ ID No.2:
[0046] TGTGTTGGGTGTGTTTGG.
[0047] In some embodiments, the amplification primer also includes a label sequence, which is composed of random bases and can be used to distinguish whole genome samples from different sample sources. It can realize linear amplification of the pre-amplification primer while completing the labeling of the sample source in the library construction step.
[0048] In some embodiments, the sequence of the amplification primer further includes an adapter connection sequence, and the adapter connection sequence is complementary to at least a portion of the sequence of the ONT adapter for connecting to the ONT adapter.
[0049] In some specific embodiments, the sequence of the amplification primer is shown in SEQ ID No.3:
[0050] AAGGTTAANNNNNNNNTGTGTTGGGTGTGTTTGG;
[0051] NNNNNNNN in the sequence represents a sequence composed of several random bases N. The number of random bases in the sequence depends on the detection requirements, and can generally be 22 to 26, and specifically, can be 24. AAGGTTAA in the sequence is an adapter connection sequence, which is used to directly connect the amplification product of the amplification primer to the ONT adapter.
[0052] The second aspect of the present application is to provide a kit for constructing a single-cell sequencing library, including the above-mentioned primer set. Specifically, the kit includes a pre-amplification reagent and an amplification reagent, the pre-amplification reagent includes the above-mentioned random primers, and the amplification reagent includes the above-mentioned amplification primers. The present application configures the pre-amplification reagent and the amplification reagent according to the random primers and the amplification primers, respectively, and successfully achieves the enrichment of the micro-primer set and the construction of the micro-primer set ONT library, while simplifying the library construction process, improving the coverage and uniformity of the single-cell whole genome sequencing data.
[0053] Pre-amplification reagent: The pre-amplification reagent is used to achieve pre-amplification of the whole genome DNA of a single cell. The concentration of the random primer in the pre-amplification reagent is 1 μM to 5 μM. Further, it is 2 μM to 5 μM. Further, it is 3 μM to 5 μM.
[0054] In some embodiments, the preamplification reagent contains a preamplification enzyme, which is used in conjunction with a random primer to preamplify the whole genome of a single cell. For example, the preamplification enzyme includes Taq DNA polymerase. In order to achieve a better preamplification effect, the working concentration of the preamplification enzyme in the preamplification reagent is 0.01U-0.5U. Further, 0.1U-0.4U. Further, 0.2U-0.3U.
[0055] In some embodiments, the preamplification reagent contains Tris hydrochloride (Tris-HCl) to provide a suitable pH environment for the preamplification reaction of random primers. In some embodiments, the pH value of the preamplification reagent is 8 to 9. It can further be 8.5. In some embodiments, the concentration of Tris-HCl in the preamplification reagent is 10mM to 50mM. It is further 20mM to 40mM. It is further 30mM.
[0056] In some embodiments, the preamplification reagent contains magnesium chloride (KCl), which is beneficial to stabilize the amplification system and improve the activity of tapDNA polymerase. Specifically, the concentration of MgCl2 in the preamplification reagent is 1mM to 4mM. Further, 1mM to 3mM. Further, 2mM.
[0057] In some embodiments, the preamplification reagent contains ammonium salt (NH4)2SO4, which is beneficial to the annealing of the primers. Specifically, the concentration of (NH4)2SO4 in the preamplification reagent is 10mM to 50mM. Further, 20mM to 40mM. Further, 30mM.
[0058] In some embodiments, the preamplification reagent contains a nonionic surfactant Triton-X100 to reduce the secondary structure of DNA. Specifically, the volume fraction of Triton-X100 in the preamplification reagent is 0.1%-0.5%. Further, it is 0.2% to 0.4%. Further, it is 0.3%.
[0059] In some embodiments, the preamplification reagent contains dimethyl sulfoxide (DMSO) to reduce the secondary structure of DNA. Specifically, the volume fraction of DMSO in the preamplification reagent is 1% to 10%. Further, 3% to 7%. Further, 5%.
[0060] In some embodiments, the preamplification reagent contains bovine serum albumin (BSA) to protect the activity of Taq DNA polymerase, reduce the influence of phenolic compounds, and reduce the adhesion of reactants to the wall, thereby enhancing the efficiency of PCR amplification. Specifically, the concentration of BSA in the preamplification reagent is 0.01 mg / mL to 0.1 mg / mL. Further, it is 0.03 μg / mL to 0.08 μg / mL. Further, it is 0.05 μg / mL.
[0061] In some embodiments, the pre-amplification reagent contains dNTP, which is used to provide reaction raw materials for the pre-amplification reaction. Wherein, dNTP is the abbreviation of deoxy-ribonucleoside triphosphate, which is a general term including dATP, dGTP, dTTP, dCTP, etc., N refers to a nitrogenous base, and the representative variable refers to one of A, T, G, C, U, etc. Specifically, the sum of the base concentrations of dNTP in the pre-amplification reagent is 0.1mM to 0.5mM. Further, it is 0.2mM to 0.4mM. Further, it is 0.3mM.
[0062] Amplification reagent: used to achieve directional amplification of pre-amplification products and complete part of the sequencing library construction. The library construction can be completed after the amplification products are connected to the adapter. Specifically, the pH value of the amplification reagent is 8 to 9, further 8.3 to 8.7, and further 8.5. In the amplification reagent, the concentration of the amplification primer is 1 μM to 5 μM, further 2 μM to 5 μM, and further 1 μM to 5 μM.
[0063] In some embodiments, the amplification reagent contains an amplification enzyme, which is used in combination with an amplification primer to complete the amplification of the pre-amplification product and simultaneously complete part of the work of constructing a sequencing library.
[0064] In some specific embodiments, the amplification enzyme includes a high-fidelity DNA polymerase. Specifically, the high-fidelity DNA polymerase contains a polymerization center and an enzyme cleavage center. The polymerization center has 5'-3' polymerase activity, which is responsible for polymerization; the enzyme cleavage center has 3'-5' exonuclease activity (also called proofreading activity), which is responsible for removing unpaired nucleotides. In some specific embodiments, in order to achieve a better library construction effect, the working concentration of the amplification enzyme is 1U to 5U. Further, 2U-4U. Further, 3U.
[0065] In some embodiments, the amplification reagent contains Tris-HCl to provide a stable pH environment for the directional amplification reaction. The concentration of Tris-HCl in the amplification reagent is 10mM to 50mM. Further, 20mM to 40mM. Further, 30mM. The pH is 7.5 to 9.5. Further, 8.5.
[0066] In some embodiments, the amplification reagent contains magnesium chloride (MgCl2), which is an indispensable cofactor for polymerase. Specifically, the concentration of MgCl2 in the amplification reagent is 1 mM to 4 mM, further 1 mM to 3 mM, and further 2 mM.
[0067] In some embodiments, the amplification reagent contains potassium chloride (KCl), which is beneficial for the annealing of the primers. Specifically, the concentration of KCl in the amplification reagent is 10 mM to 50 mM, further 20 mM to 40 mM, and further 30 mM.
[0068] In some embodiments, the amplification reagent contains ammonium salt (NH4)2SO4, which is beneficial to the annealing of the primers. Specifically, the concentration of (NH4)2SO4 in the amplification reagent is 10mM to 50mM. Further, 20mM to 40mM. Further, 30mM.
[0069] In some embodiments, the amplification reagent contains a nonionic surfactant Triton-X100 to reduce the secondary structure of DNA. Specifically, the volume fraction of Triton-X100 in the amplification reagent is 0.1%-0.5%. Further, it is 0.2% to 0.4%. Further, it is 0.3%.
[0070] In some embodiments, the amplification reagent contains dimethyl sulfoxide (DMSO) to reduce the secondary structure of DNA. Specifically, the volume fraction of DMSO in the amplification reagent is 1% to 10%, further 3% to 7%, and further 5%.
[0071] In some embodiments, the amplification reagent contains bovine serum albumin (BSA) to protect the activity of TaqDNA polymerase, reduce the influence of phenolic compounds, and reduce the adhesion of reactants to the wall, thereby enhancing the efficiency of PCR amplification. Specifically, the concentration of BSA in the amplification reagent is 0.01 mg / mL to 0.1 mg / mL. Further, it is 0.03 μg / mL to 0.08 μg / mL. Further, it is 0.05 μg / mL.
[0072] In some embodiments, the amplification reagent contains dNTPs to provide reaction raw materials for the amplification reaction. Specifically, the sum of the base concentrations of the dNTPs in the amplification reagent is 0.1 mM to 0.5 mM, further 0.2 mM to 0.4 mM, and further 0.3 mM.
[0073] End repair reagent: In order to add adapters to both ends of the amplified product of the single-cell whole genome, the kit also includes end repair reagents to facilitate the subsequent completion of the adapter ligation reaction.
[0074] In some specific embodiments, the end repair reagent can use conventional end repair reagents in the art to perform end repair treatment on the amplified product and add adenine (A) fragments to the ends of the DNA fragments.
[0075] Adapter ligation reagent: In some embodiments, the kit further includes an adapter ligation reagent for implementing an adapter ligation reaction to complete the construction of a sequencing library.
[0076] In some specific embodiments, the adapter ligation reagent contains an ONT adapter, which contains a motor protein and an optional DNA helicase. The motor protein provides a driving force for DNA to pass through the nanopore; the DNA helicase is only required in the presence of double-stranded DNA, and it can unwind the double-stranded DNA into single strands for sequencing through the nanopore.
[0077] Cell lysis reagent: In some embodiments, in order to obtain a single cell whole genome DNA sample, the kit also includes a cell lysis reagent to release the genomic DNA in the cell. In some specific embodiments, the pH value of the cell lysis reagent is 8 to 9; further 8.3 to 8.8; further 8.5.
[0078] In some embodiments, the cell lysis reagent contains a cell lysis enzyme for enzymatically hydrolyzing histones and releasing DNA nucleic acids in cells. Specifically, the cell lysis enzyme includes proteinase K, which is a powerful proteinase with high specific activity. It is active in a wide pH range (4 to 12.5) and high temperature (50°C to 70°C). In genomic DNA extraction, it is used to hydrolyze histones bound to nucleic acids, freeing DNA in the solution to facilitate subsequent pre-amplification reactions. In some specific embodiments, the working concentration of the cell lysis enzyme is 10μg / mL to 30μg / mL, and further 20μg / mL to 30μg / mL.
[0079] In some embodiments, the cell lysis reagent contains Tris-HCl to provide a stable pH environment for the cell lysis enzyme. Specifically, the concentration of Tris-HCl in the cell lysis reagent is 10mM to 50mM; further 20mM to 40Mm, and further 30mM.
[0080] In some embodiments, the cell lysis reagent contains nonionic surfactant ethylphenyl polyethylene glycol (NP-40), which can reduce DNA secondary binding. Specifically, the volume fraction of NP-40 in the cell lysis reagent is 0.1%-0.5%, further 0.2%-0.4%, and further 0.3%.
[0081] In some embodiments, the cell lysis reagent contains Triton X-100. Specifically, the volume fraction of Triton X-100 in the cell lysis reagent is 0.1%-0.5%, further 0.2%-0.4%, and further 0.3%.
[0082] The third aspect of the present application provides a method for single-cell whole genome sequencing, comprising: using the above-mentioned kit to construct a sequencing library of a single-cell whole genome DNA sample; and performing ONT sequencing on the sequencing library.
[0083] Construct a sequencing library of a single-cell whole-genome DNA sample. Specifically, a whole-genome pre-amplification reaction is performed on the genomic DNA sample using a pre-amplification primer as shown in SEQ ID No. 1. Complementary sequences are introduced at both ends of the complete pre-amplification product, so that the newly amplified DNA molecule forms a circular structure and cannot be amplified as a template. Further, the newly amplified DNA molecule can be amplified as a template by the corresponding amplification primer, and a label sequence is introduced during the amplification process to obtain an amplification primer containing a label sequence. The amplification primer can be directly used to connect to the ONT connector after end repair, thereby completing the construction of a single-cell whole-genome sequencing library.
[0084] Compared with the traditional single-cell whole genome library construction method, the sequencing library construction method of the present application only needs one-step amplification treatment of the pre-amplification product to complete the library construction, which simplifies the library construction process, saves library construction time, and improves the coverage and uniformity of single-cell whole genome sequencing data. It has the characteristics of extremely simplified library construction process, good library construction quality, and relatively low cost.
[0085] In some specific embodiments, such as Figure 1 As shown in the figure, the specific steps for constructing a sequencing library for a single-cell whole genome DNA sample include:
[0086] Construction of a sequencing library for a single-cell whole-genome DNA sample includes:
[0087] Pre-amplifying the single-cell whole genome DNA sample to obtain a pre-amplification product;
[0088] The pre-amplification product is amplified to obtain an amplification product, and the amplification product is end-repaired and connected to an ONT connector to obtain a single-cell whole genome sequencing library.
[0089] Specifically, end repair refers to adding an adenine (A) base to the 3' end of the amplified product. The purpose of end repair is to repair the damaged or incomplete DNA end into a 5' phosphorylated blunt-end DNA for blunt-end DNA ligation with the ONT adapter.
[0090] In some embodiments, in order to improve the ligation efficiency, the amplified product is purified before the end repair is performed on the amplified product. The purification of the amplified product can be performed by conventional DNA purification methods in the art.
[0091] The single-cell whole genome sequencing library constructed by the above method has higher coverage and uniformity. Subsequently, ONT sequencing is performed on the above single-cell whole genome sequencing library, so that the sequencing data has higher coverage and uniformity.
[0092] Specifically, the above-mentioned single-cell whole genome sequencing library is purified and the concentration is tested, and the sequencing library that passes the concentration test is sequenced on the ONT platform. Compared with the traditional single-cell whole genome sequencing method, the sequencing data obtained by sequencing in this application has higher coverage and uniformity.
[0093] The implementation scheme of the present application will be described in detail below with reference to examples.
[0094] The primers used in the examples of this application were synthesized by Shanghai Bio-Tech Co., Ltd. The NEBNext ULtra II End Repair / dA Tailing Kit (#E7546L) and NEBNext Quick Ligation Module Kit (#E6056L) used in the examples of this application were purchased from NEW ENGLANDBioLabs; Single Cell WGA Kit (Takara, #112219) was purchased from TakaraBio; Ligation Sequencing Kit (SQK-LSK109) was purchased from Oxford Nanopore Technologies. Other reagents used in the examples of this application were obtained through conventional commercial channels.
[0095] Example 1
[0096] 1. Obtain cell line samples or tissue samples, prepare them into cell suspensions, and then use micromanipulation, flow cytometry, etc. to prepare the cell suspensions into single cells and dispense them into 0.2 ml PCR tubes. The solution volume is 2.5 μL.
[0097] 2. Cell lysis reaction
[0098] (1) Prepare cell lysis buffer: 25 mM Tris-HCl, 0.25% Triton-X100, 0.25% NP-40.
[0099] (2) Add 2.5 μL of cell lysis buffer and 0.2 μL of proteinase K to every 2.5 μL of single cell sample, centrifuge briefly, and place in a PCR instrument for reaction. Reaction procedure: 70°C for 10 min; 95°C for 4 min; store at 25°C.
[0100] 3. Pre-amplification reaction
[0101] (1) Prepare the pre-amplification buffer according to Table 1 below;
[0102] Table 1
[0103] Element Reaction volume μL 250 mM Tris-HCl (pH 8.5) 0.5 <![CDATA[20mM MgCl2]]> 0.5 100uM random primers 0.3 10mM dNTP 0.3 <![CDATA[250mM(NH4)2SO4]]> 0.5 50% DMSO 0.5 2.5% (v / v) Triton-X100 0.5 5mg / mL BSA 0.5 Nuclease-free water 1.4 total: 5
[0104] The random primer sequences are shown in Table 2 below, wherein NGP, NTP and NKP contain 5 N bases, NKP contains 4 K bases, N represents a random base, and K represents T or G.
[0105] Table 2
[0106] Primer number Primer sequences NGP TGTGTTGGGTGTGTTTGGNNNNNGTTT NTP TGTGTTGGGTGTGTTTGGNNNNNTGGG NKP TGTGTTGGGTGTGTTTGGNNNNNKKKK
[0107] 5 μL of pre-amplification buffer and 0.2 μL of pre-amplification enzyme were added to 5 μL of lysate, and the mixture was centrifuged briefly and then placed in a PCR instrument for reaction.
[0108] (2) The pre-amplification reaction procedure is shown in Table 3;
[0109] Table 3
[0110]
[0111]
[0112] 4. Amplification reaction
[0113] (1) Prepare the amplification reaction system according to Table 4 below;
[0114] Table 4
[0115] Element Reaction volume μL 1M Tris-HCl (pH 8.5) 1 <![CDATA[100mM MgCl2]]> 0.8 1M KCl 1 <![CDATA[1M(NH4)2SO4]]> 2 10%(v / v)Tritom-X100 1 100%(v / v)DMSO 2 10mg / ml BSA 2 100U high-fidelity DNA polymerase 1.2 Pre-amplification product 10 Amplification primers 10 Nuclease-free water 9 total: 40
[0116] (2) After the preparation is completed, centrifuge it instantly and place it in a PCR instrument to perform the reaction. The reaction procedure is shown in Table 5.
[0117] Table 5
[0118]
[0119]
[0120] Purification of amplified products:
[0121] (1) Add GCS Reagent (M) DNA Clean Beads at a ratio of 1:1. Add 60 μL of purified magnetic beads to every 60 μL of reaction product, pipette to mix, and adsorb at room temperature for 5 min.
[0122] (2) Wash twice with 100 μL of 80% ethanol and let dry for 2 min.
[0123] (3) Add 66 μL ddH2O and mix thoroughly. Let stand at room temperature for 2 min to elute.
[0124] 4. End repair
[0125] The ligation reaction system was prepared using the enhanced version of the second-generation NEB End Repair / dA Tailing Kit (NEB Next μLtra II End Repair / dA Tailing Kit (#E7546L)), as shown in Table 6.
[0126] Table 6
[0127] Element Reaction volume μL Amplification products 100fmol End Repair Enhancement Buffer 3.5 Super strong end repair enzyme mixture 3 total: Rehydrate to 60
[0128] 5. Ligation reaction
[0129] The ligation reaction system was prepared using the second generation NEB Quick Ligation Module Kit (#E6056L), as shown in Table 7.
[0130] Table 7
[0131] Element Reaction volume μL Unfinished product 65 NEB Quick Ligation Reaction Buffer 20 Joint Mixture II 5 Fast T4 DNA Ligase 10 total: 100
[0132] Using the above steps, 4 sets of parallel samples were constructed at the same time. Note: NEB is the manufacturer's name.
[0133] After the product is purified, it is sequenced:
[0134] (1) Add glycogen synthase reagent (GCS Reagent) (M) DNA Clean Beads (purified magnetic beads, the surface of which has been specifically modified to have a strong affinity with DNA molecules and adsorb DNA molecules) at a ratio of 2:1. Add 50 μL of purified magnetic beads to every 100 μL of reaction product, pipette to mix, and adsorb at room temperature for 5 min.
[0135] (2) Wash twice with 200 μL of SFB (Short Fragment Buffer, an ONT reagent derived from Oxford Nanopore Technologies' kit #SQK-109, which is used to clean DNA fragments below 3 kb) and let dry for 30 seconds.
[0136] (3) 15 μL Elution Buffer (EB, elution buffer, derived from Oxford Nanopore Technologies kit #SQK-109, which can dissolve DNA well and provide a more suitable pH environment) was pipetted to mix well and allowed to stand at room temperature for 10 min for elution.
[0137] (4) The purity of the obtained product was measured using an ultra-micro spectrophotometer (ThermoFisher, NanoDrop One).
[0138] (5) The obtained product is quantified using a Qubit fluorometer.
[0139] (6) After passing quality control, the samples were sequenced on the Oxford nanopore MinlON platform with an estimated sequencing depth of about 1x. The sequencing data were compared and analyzed to obtain basic information about the whole genome DNA.
[0140] The schematic diagram of GC distribution of the three libraries is as follows Figure 2 shown. Figure 2In the figure, the horizontal axis is the GC content of each sequencing statistical window, and the vertical axis is the number of sequencing Reads in each window. That is, when the GC content of the statistical window is x, the shorter the length of the box graph, the smaller the fluctuation range y of the number of Reads, and the GC content within the fluctuation range y of the number of Reads is x, indicating that the distribution of sequencing GC content is more concentrated.
[0141] from Figure 2 It can be seen from the test results that when NKP (Barcode60) is used as a random primer, the GC distribution in the sequencing data is the most concentrated and uniform.
[0142] The Lorenz diagram is drawn with the percentage of the cumulative number of reads in the first x windows in each statistical window of the sequencing data obtained from the sequencing library to the total number of windows as the horizontal axis and the percentage of the actual number of genome reads obtained in the first x windows to the total number of reads as the vertical axis. The Lorenz diagrams of the three libraries are shown in the figure. Figure 3 The part between the Lorenz curve and the 45° line (absolute equality line) is called the "inequality area". The smaller the area, the closer the number of genome reads obtained from each sequencing window is, that is, the wider the genome area covered, the better the amplification uniformity.
[0143] from Figure 3 It can be seen from the Lorentz plot that the curve represented by Barcode60 is closest to the absolute average line, indicating that the uniformity of amplification is better when NKP is used as a random primer.
[0144] Example 2
[0145] 1. Single cell isolation: The method is the same as in Example 1, and two groups of single cell samples from the same patient are obtained.
[0146] 2. Obtain the whole genome of a single cell and sequence it
[0147] Two sets of libraries were constructed, in which the first set of single cells was constructed using the system described in Example 1, in which the pre-amplification primer used was NKP, and the resulting library was named S1. Single Cell WGA Kit (Takara, #112219) was used for single cell whole genome amplification, and 150 ng of amplified product was used for library construction, end repair, label ligation, and adapter ligation through Ligation Sequencing Kit (SQK-LSK109), and finally a purified library was obtained and named S2.
[0148] The obtained products S1 and S2 were tested for purity using an ultra-micro spectrophotometer (ThermoFisher, NanoDrop One). The concentration was quantified using a Qubit fluorometer. After the two products passed the quality control, they were sequenced on the Oxford nanopore MinlON platform. The sequencing depth was expected to be about 1x. The sequencing data was compared and analyzed to obtain basic information about the amplified products.
[0149] The data obtained by sequencing were compared with the reference genome sequence, and the sequencing alignment rate, coverage rate and SD value were calculated, as shown in Table 8. The sequencing data were analyzed using the human cell chromosome detection data analysis system. The data analysis results of library S1 are shown in Figure 4 As shown, the data analysis results of library S2 are as follows Figure 5 shown.
[0150] Table 8
[0151]
[0152]
[0153] As can be seen from Table 8, the amount of sequencing data of library S1 and library S2 is roughly the same, and there is no significant difference in alignment rate and average sequencing depth; in terms of coverage, library S1 is slightly higher than S2; in terms of Figure 4 and Figure 5 It can be seen that the sequencing results of S1 and S2 show that the scatter distribution concentration of S1 is more concentrated, and the chromosome ploidy is equally clear.
[0154] The above results show that compared with the traditional method of combining single-cell whole genome amplification and library construction kits for single-cell whole genome sequencing, the single-cell whole genome sequencing effect of the ONT single-cell library construction method provided in this application has higher coverage and uniformity.
[0155] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0156] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A primer set for constructing a single-cell whole genome sequencing library, characterized in that: The primer set includes a random primer and an amplification primer, the sequence of the random primer is shown in SEQ ID No. 1, the sequence of the amplification primer is partially identical to the sequence of the random primer, the random primer is used to amplify genomic DNA, and the amplification primer is used to amplify the amplification product of the random primer.
2. The primer set according to claim 1, characterized in that The primer set meets at least one of the following characteristics: The sequence of the amplification primer includes the sequence shown in SEQ ID No.2; The sequence of the amplification primer also includes a linker connection sequence, and the linker connection sequence is complementary to at least a portion of the sequence of the ONT linker; The sequence of the amplification primer also includes a tag sequence, and the tag sequence is a sequence composed of random bases; optionally, the number of random bases in the tag sequence is 22 to 26; The sequence of the amplification primer includes the sequence shown in SEQ ID No.3; The working concentration of the random primers is 1 μM to 5 μM; and, The working concentration of the amplification primer is 1 μM to 5 μM.
3. A kit for constructing a single-cell whole genome sequencing library, characterized in that: The kit comprises a pre-amplification reagent and an amplification reagent, wherein the pre-amplification reagent comprises the random primer as described in claim 1 or 2; and the amplification reagent comprises the amplification primer as described in claim 1 or 2.
4. The kit according to claim 3, characterized in that The pH value of the pre-amplification reagent is 8-9, and the pre-amplification reagent also includes a pre-amplification enzyme, and the working concentration of the pre-amplification enzyme is 0.01U-0.5U; optionally, the pre-amplification enzyme includes Taq DNA polymerase.
5. The kit according to claim 3 or 4, characterized in that The pre-amplification reagent further contains at least one of Tris-HCl, MgCl2, (NH4)2SO4, Triton-X100, DMSO, BSA and dNTP; optionally, the pre-amplification reagent satisfies one or more of the following characteristics: The concentration of Tris-HCl in the pre-amplification reagent is 10mM to 50mM; The concentration of MgCl2 in the pre-amplification reagent is 1mM to 4mM; The concentration of (NH4)2SO4 in the pre-amplification reagent is 10mM to 50mM; The volume fraction of Triton-X100 in the pre-amplification reagent is 0.1%-0.5%; The volume fraction of DMSO in the pre-amplification reagent is 1% to 10%; The concentration of BSA in the pre-amplification reagent is 0.01 mg / mL to 0.1 mg / mL; The sum of the base concentrations of the dNTPs in the pre-amplification reagent is 0.1 mM to 0.5 mM.
6. The kit according to claim 3, characterized in that The pH value of the amplification reagent is 8-9. The amplification reagent also includes an amplification enzyme. The amplification enzyme has a polymerization center and an enzyme cleavage center, and its working concentration is 1U-5U; optionally, the amplification enzyme includes a high-fidelity DNA polymerase.
7. The kit according to claim 3 or 6, characterized in that The amplification reagent further contains at least one of Tris-HCl, MgCl2, KCl, (NH4)2SO4, Triton-X100, DMSO, BSA and dNTP; optionally, the amplification reagent satisfies one or more of the following characteristics: The concentration of Tris-HCl in the amplification reagent is 10mM to 50mM; The concentration of MgCl2 in the amplification reagent is 1mM to 4mM; The concentration of KCl in the amplification reagent is 10mM to 50mM; The concentration of (NH4)2SO4 in the amplification reagent is 10mM to 50mM; The volume fraction of Triton-X100 in the amplification reagent is 0.1%-0.5%; The volume fraction of DMSO in the amplification reagent is 1% to 10%; The concentration of BSA in the amplification reagent is 0.01 mg / mL to 0.1 mg / mL; The sum of the base concentrations of the dNTPs in the pre-amplification reagent is 0.1 mM to 0.5 mM.
8. The kit according to claim 3, characterized in that The kit further comprises a cell lysis reagent, the pH value of the cell lysis reagent is 8-9, and the cell lysis reagent comprises at least one of a cell lysis enzyme, Tris-HCl, NP-40 and Triton X-100; optionally, the cell lysis reagent satisfies at least one of the following characteristics: The concentration of Tris-HCl in the cell lysis reagent is 10mM to 50mM; The volume fraction of NP-40 in the cell lysis reagent is 0.1%-0.5%; The volume fraction of Triton X-100 in the cell lysis reagent is 0.1%-0.5%; The cell lytic enzyme includes proteinase K, and the working concentration of the cell lytic enzyme is 10 μg / mL to 30 μg / mL.
9. The kit according to claim 3, characterized in that The kit further comprises at least one of an end repair reagent and a linker ligation reagent; The end repair reagent is used to add an adenine fragment to the end of the DNA fragment; and / or the linker connection reagent contains an ONT linker.
10. A single-cell whole genome sequencing method, characterized in that: The method comprises the following steps: constructing a sequencing library of a single-cell whole genome DNA sample using the kit according to any one of claims 3 to 8; and The sequencing library is subjected to ONT sequencing.
11. The method according to claim 10, characterized in that The construction of a sequencing library of a single-cell whole genome DNA sample comprises the following steps: Pre-amplifying the single-cell whole genome DNA sample using a pre-amplification reagent to obtain a pre-amplification product; amplifying the pre-amplification product using an amplification reagent to obtain an amplification product; as well as Optionally, the amplified product is end-repaired and then connected to an ONT adapter to obtain a single-cell whole genome sequencing library.
12. The method according to claim 11, characterized in that Before pre-amplifying the single-cell whole genome DNA sample using a pre-amplification reagent, the method further includes: The single cell sample is lysed using a cell lysis reagent to obtain the single cell whole genome DNA sample.