Single cell four-recombination library construction and sequencing method
Through mtCOOL-seq technology, the problem of single-cell multi-omic sequencing in the existing technology is solved, and the problem of high cost of single-cell multi-omic sequencing is not possible at the same time is achieved, efficient and economical single-cell quadruplecommunication analysis is achieved, and more comprehensive epigenomic information is provided.
Patent Information
- Application Number
- CN202411964601.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-13
AI Technical Summary
The existing single-cell multi-omic sequencing technology is difficult to achieve sample mixing operations in the DNA part, resulting in high sequencing costs; the single-cell methylation mixing technology cannot perform multi-omics detection at the same time; the existing technology can only perform whole-genome detection on DNA, and cannot achieve simplified sequencing.
A medium-throughput single-cell quadruplecommunics library construction and sequencing technology mtCOOL-seq is developed to enrich key genomic sites by simultaneously sequencing the transcriptome, methylation group, chromatin openness and copy number variation of single cells, combining labeling technology and enzyme cutting to enrich key genomic sites, simplifying operations and reducing costs.
Efficient and economical single-cell multiomic analysis is achieved, providing more comprehensive epigenomic information, reducing sequencing costs and complexity, and improving experimental throughput and data quality.
Smart Images

Figure HDA0005218038720000011 
Figure HDA0005218038720000012 
Figure HDA0005218038720000021
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biotechnology, and in particular relates to a single-cell quadruple chemistry library construction and sequencing method. Background Art
[0002] Methylation and DNA methylation research and its significance: Methylation, especially DNA methylation, is an important area in current disease research because it is closely related to gene expression and phenotypic changes. DNA methylation refers to the process of transferring a methyl group from S-adenosylmethionine (SAM) to a specific base under the action of DNA methyltransferase (DMT). This process can occur at positions such as N-6 of adenine, N-7 of guanine, and C-5 of cytosine.
[0003] In mammals, DNA methylation is mainly concentrated on the cytosine of the 5'-CpG-3' sequence to form 5-methylcytosine (5mC). CpG sequences exist in two forms in mammals: one is a dispersed CpG dinucleotide, and the other is a high-density clustered CpG island. In the genome, 70% to 90% of the dispersed CpGs will be methylated, while CpG islands usually remain unmethylated, except in certain specific regions or genes. CpG islands are usually located in transcriptional regulatory regions and are associated with 56% of the human genome coding genes, so it is particularly important to study the methylation status of these regions. At the same time, DNA methylation is closely related to human development, cell differentiation, aging and various diseases. In particular, methylation of CpG islands may affect the activation of oncogenes, while hypomethylation of genomic repetitive sequences may lead to decreased genome stability. DNA methylation has become an important research direction in genetics and epigenomics.
[0004] Analysis of the human genome shows that there are approximately 28,890 CpG islands distributed in the genome, usually 5 to 15 CpG islands per 1Mb of chromosome, with an average of about 10.5 per Mb. In recent years, DNA methylation features have been used for the diagnosis and prognosis assessment of a variety of tumors. Studies have shown that DNA methylation patterns reveal the mechanisms of cancer occurrence and development and cellular heterogeneity in cancer tissues, providing new possibilities for early detection, prognosis assessment, and treatment. In addition, analyzing the methylation status of CpG islands in DNA sequences can also help explain the mechanisms of a variety of human diseases from an epigenetic level.
[0005] There are three main types of traditional methods for DNA methylation sequencing, each of which is based on different biochemical principles to detect DNA methylation status. (1) Bisulfite treatment is one of the most widely used methods. The principle is that bisulfite can selectively convert unmethylated cytosine into uracil, while methylated cytosine remains unchanged. In the subsequent PCR amplification and sequencing process, the originally unmethylated cytosine will be read as thymine, while the methylated cytosine will still be read as cytosine, thereby achieving the identification of methylation sites. The advantage of this method is that it can provide methylation information with single-base resolution, but there are also potential problems of DNA damage and incomplete conversion. (2) Methods based on the specific binding of methylated or unmethylated C or CpG DNA mainly include methylated DNA immunoprecipitation sequencing (MeDIP-seq) and methylated CpG binding domain protein sequencing (MBD-seq). MeDIP-seq uses antibodies that specifically recognize 5-methylcytosine to enrich methylated DNA fragments, while MBD-seq uses methyl CpG binding domain proteins to capture methylated DNA. Both methods can provide genome-wide methylation profiles, but the resolution is relatively low, usually at the level of hundreds to kilobases. (3) Methods based on the blocking effect of methylated DNA on methylation-sensitive restriction endonucleases, such as methylation-sensitive restriction endonuclease sequencing (MRE-seq). This method uses the property that certain restriction endonucleases can only cut unmethylated recognition sites but not methylated sites. By comparing the DNA fragments after enzyme digestion, the location of the methylation site can be inferred. The advantage of MRE-seq is that it can quickly and economically obtain genome-wide methylation information, but its resolution is limited by the distribution of recognition sites of the enzyme used.
[0006] However, these traditional methods have high requirements for the amount of DNA input, usually requiring hundreds of nanograms to micrograms of DNA samples, which limits their application in rare cell types or trace samples. Despite this, bisulfite sequencing is still considered the gold standard for DNA methylation analysis because it can provide single-base resolution, allowing researchers to accurately map methylation. This high resolution enables researchers to study the methylation status of individual CpG sites, which is crucial for understanding gene regulation and disease-related epigenetic variations.
[0007] In mammalian DNA methylation research, the methylation detection of CpG and CpG islands is most widely used by methods such as whole genome bisulfite sequencing (WGBS) and reduced representation bisulfite sequencing (RRBS). Whole genome BS (WGBS) technology can be used to study the DNA methylation status of population cells, providing the most comprehensive methylation information across the genome. However, since the entire genome needs to be sequenced, the cost of library construction and sequencing is relatively high, especially when dealing with a large number of samples. This high cost limits its application in large-scale studies. Reduced representation BS (RRBS) technology is a selective enrichment method based on restriction endonuclease digestion, which can provide researchers with an economical and efficient method for DNA methylation research. The main advantage of RRBS technology is that it can enrich CpG-dense regions such as promoters and CpG islands, while significantly reducing sequencing depth and cost, making it possible to analyze more samples within a limited budget. However, the common disadvantage of WGBS and RRBS technologies is that they both require complex sample processing, including DNA extraction, bisulfite conversion, and library construction. These tedious steps not only increase the experimental time, but also may introduce technical biases, making them inconvenient for medium- and high-throughput sequencing of multiple cell samples from scratch. In addition, these methods face challenges in processing trace samples, especially at the single-cell level, which may lead to uneven genome coverage and reduced data quality due to problems such as DNA loss and incomplete conversion.
[0008] With the rapid development of single-cell technology, researchers have gradually turned their attention to DNA methylation analysis at the single-cell level. This shift stems from the deepening understanding of the importance of cellular heterogeneity and the increasing demand for high-resolution epigenetic information. To meet this challenge, scientists have developed technologies such as single-cell whole-genome bisulfite sequencing (scWGBS) and single-cell reduced representation bisulfite sequencing (scRRBS). These methods are innovative improvements to traditional WGBS and RRBS technologies, designed to overcome the unique challenges faced in performing DNA methylation analysis at the single-cell level.
[0009] Main methods of single-cell DNA methylation sequencing: scWGBS technology is one of the first methods to achieve whole-genome methylation analysis at the single-cell level. It makes it possible to obtain whole-genome methylation information from a single cell by optimizing steps such as DNA extraction, bisulfite conversion, and library construction. The main improvements of scWGBS include the use of micro-sample processing technology, improving DNA conversion and amplification efficiency, and adopting special sequencing strategies to reduce data bias. However, scWGBS still faces some limitations, such as high sequencing costs, uneven genome coverage, and large technical noise.
[0010] To address these challenges, researchers developed the scRRBS technology. scRRBS has improved the original RRBS method. The most significant innovation is that all experimental steps of sample processing are integrated into a single-tube reaction. This improvement greatly simplifies the operation process and reduces the risk of sample loss and contamination. scRRBS can provide digital methylation information of about 1 million CpG sites in a single diploid mouse or human cell at single-base resolution. Although scRRBS covers fewer CpG sites than scWGBS (about 3.7 million CpG sites), it better covers CpG islands, which are considered to be the elements with the richest DNA methylation information. The core principle of scRRBS is to enrich CpG island sites in genomic DNA using specific restriction enzymes such as MspI. The DNA fragments after enzyme digestion are treated with bisulfite to convert unmethylated cytosine to uracil, while retaining methylated cytosine. Subsequently, PCR amplification is performed to make the DNA fragments reach the concentration required for sequencing. The main advantage of scRRBS is that it significantly reduces the cost and complexity of sequencing while maintaining high resolution. This allows researchers to analyze more single-cell samples within a limited budget, thereby more comprehensively revealing epigenetic heterogeneity in cell populations. In addition, the enrichment of CpG islands by scRRBS also makes it particularly suitable for studying the methylation status of promoter regions that are closely related to gene regulation.
[0011] With the rapid development of biotechnology, single-cell multi-omics technology has emerged and become an important tool for studying complex biological systems. This advanced method can simultaneously analyze multiple omics dimensions at the single-cell level, including but not limited to gene expression (transcriptome), DNA chemical modification (methylome) and chromatin structure (openness). Unlike traditional data analysis that relies on large-scale cell populations, single-cell multi-omics technology can more accurately capture the heterogeneity between cells, providing unprecedented opportunities for in-depth understanding of the functions of different cells in complex tissues and their roles in health and disease.
[0012] In epigenetic research, transcriptome, methylome and chromatin accessibility are three key omics fields. Transcriptome analysis reveals all RNA transcription products of cells under specific conditions, mainly including messenger RNA (mRNA), reflecting the gene expression pattern and cell functional status. This is of great significance for understanding cell development processes, environmental response mechanisms and disease mechanisms. Methylome research focuses on methylation modifications on DNA molecules, especially the methylation status of CpG dinucleotides. As an important epigenetic modification, DNA methylation plays a key role in gene expression regulation, cell differentiation and development. Abnormal DNA methylation patterns are closely related to a variety of diseases (such as cancer), so methylome research is crucial for disease diagnosis and the development of treatment strategies. Chromatin accessibility analysis focuses on the looseness of chromatin structure, which determines the accessibility of transcription factors and other regulatory proteins to DNA. Open chromatin regions are usually associated with active gene expression and play a key role in revealing cell type-specific gene expression and transcriptional regulation mechanisms.
[0013] At present, a variety of advanced technical methods have been developed in the field of single-cell multi-omics research, among which the following three technologies are particularly eye-catching: (1) scNOMe-seq (single-cell Nucleosome Occupancy and Methylome sequencing) is a method that can simultaneously measure DNA methylation and chromatin accessibility in single cells. This technology uses GpC methyltransferase to mark open chromatin regions in vitro, thereby distinguishing nucleosome-occupied and nucleosome-free regions, while retaining endogenous CpG methylation information. (2) scTrio-seq (single-cell triple omics sequencing) is another powerful single-cell multi-omics technology that can simultaneously analyze the genome, transcriptome, and DNA methylome of a single cell. This method provides multi-level information on cellular heterogeneity by combining single-cell RNA sequencing with genomic and epigenomic analysis, which helps to reveal the relationship between gene expression, genetic variation, and epigenetic modifications. (3)iscCOOL-seq (single-cell Chromatin Overall Omic-scale Landscape Sequencing) is an innovative single-cell multi-omics sequencing technology that can simultaneously analyze chromatin state / nucleosome positioning, DNA methylation, copy number variation, and ploidy from the same mammalian cell. The uniqueness of this technology lies in its ability to provide comprehensive epigenomic information at the single-cell level, including chromatin accessibility, DNA methylation patterns, and genomic structural variation. The development of iscCOOL-seq provides a powerful tool for studying epigenetic regulation in complex biological processes, allowing researchers to fully understand the dynamic changes of the epigenome at single-cell resolution.
[0014] The development of single-cell multi-omics sequencing technology is closely related to the progress of second-generation sequencing and third-generation sequencing technology. Second-generation sequencing, also known as high-throughput sequencing or next-generation sequencing, is mainly based on the principle of sequencing by synthesis, and achieves high-throughput data output through large-scale parallel sequencing. Its advantages include high throughput, relatively low cost and high accuracy, and it has been widely used in genomics and transcriptomics research. The main technical principles of third-generation sequencing include single-molecule real-time sequencing (SMRT) and nanopore sequencing. These technologies can directly read single DNA molecules, produce longer read lengths (up to thousands to tens of thousands of base pairs), and handle difficulties in complex genome structures and long-range information. Third-generation sequencing is particularly suitable for applications such as whole-genome de novo assembly, structural variation detection, and full-length transcript analysis. However, third-generation sequencing still faces some challenges, including high sequencing error rate, relatively low throughput and high cost. In addition, the data analysis methods and bioinformatics tools of third-generation sequencing are still under development and improvement. Therefore, despite the continuous progress of third-generation sequencing technology, second-generation sequencing is still the optimal solution for both cost and technical popularity. In single-cell multi-omics research, next-generation sequencing technology remains the mainstream choice, especially in application scenarios that require high throughput and high accuracy.
[0015] Current single-cell sequencing technologies are mainly divided into two types: one is to simultaneously sequence multiple epigenomics such as transcriptome, chromatin accessibility and methylome at the single cell level. Although this method provides comprehensive multi-omics data, it requires each cell to be processed individually, resulting in complex operations and high costs. The other is to achieve medium-throughput single-omics sequencing through labeling technology, which simplifies operations and reduces costs, but cannot obtain other omics data at the same time.
[0016] The first type of technical method is represented by iscCOOL-seq. It can simultaneously analyze the methylome, chromatin accessibility, and transcriptome in a single cell. Through optimized processes, iscCOOL-seq simplifies experimental operations while providing high-quality multi-omics data. It is an effective tool for exploring single-cell multi-omics. The main steps are to use the methylation library construction scheme and the RNA transcriptome library construction scheme in combination with the nuclear separation step. The main technical schemes of iscCOOL-seq technology are as follows: (1) In vitro methylation treatment: Use GpC methylase to methylate a single cell in vitro. (2) Separation of cytoplasm and nucleus: Use magnetic beads to capture the nucleus, thereby separating the nucleus for single-cell DNA methylation library construction, and the cytoplasm for single-cell RNA-seq library construction. (3) Single-cell DNA methylation library construction steps: Cell treatment: Perform protease digestion and bisulfite conversion to distinguish methylated cytosines. (4) DNA purification: Extract and purify single-cell genomic DNA. (5) Library construction (TAILS method): Random amplification: Random amplification was performed using the P5-N6-oligo1 primer, and the reaction system included 1× Blue buffer, 600 nM dNTPs, 400 nM P5-N6-oligo1, and 50 units of Klenow exo enzyme, and incubated at 37°C for 30 minutes. Deactivation: The remaining primers and dNTPs were deactivated using Exo-SAP IT reagent. dC tailing: dC tailing was performed under the conditions of 1× Blue buffer, 2.5 mM dCTPs, and 1 U / μL TdT enzyme. Second-strand synthesis: Second-strand synthesis was performed using the P7-G6-oligo2 primer, and the reaction system included 1× Blue buffer, 600 nM dNTPs, 400 nM P7-G6-oligo2, and 100 units of Klenow exo enzyme, and incubated at 37°C for 1 hour. Purification: Purification was performed using Agencourt AMPure XP magnetic beads. (6) Final library preparation: Universal primers and index primers were introduced through 18 rounds of PCR. Agencourt AMPureXP magnetic beads were used for final purification. The libraries were merged, quality tested, and 150-bp paired-end sequencing was performed on the Illumina HiSeq X-Ten platform.
[0017] Library construction steps for single-cell RNA-seq: Cell preparation: Single cells were manually picked and lysed. cDNA synthesis: First-strand cDNA synthesis was performed. Second-strand cDNA was synthesized, amplified, and fragmented. Library preparation: RNA-seq libraries were prepared using the KAPA Hyper Prep Kit. Libraries were pooled, quality tested, and 150-bp paired-end sequenced on the Illumina HiSeqX-Ten platform.
[0018] In summary, current sequencing technologies still have different technical defects. For example, the existing single-cell multi-omics technology cannot achieve sample mixing operations in the DNA part, resulting in excessively high DNA sequencing costs; single-cell methylation mixed sample technology cannot perform multi-omics detection at the same time; single-cell multi-omics technology can only perform whole-genome detection on DNA and cannot achieve simplified sequencing. In view of the limitations of existing single-cell multi-omics sequencing technologies and the urgent need for efficient, economical and comprehensive single-cell epigenomic analysis methods, this application is hereby filed. Summary of the invention
[0019] The present invention aims to develop a medium-throughput single-cell multi-omics library construction and sequencing technology mtCOOL-seq (middle throughput single-cell multi-omics sequencing technology), which comprehensively analyzes cell states and identifies cell heterogeneity in complex tissues by simultaneously sequencing the transcriptome, methylome, chromatin accessibility, and copy number variation of single cells. Combining labeling technology with enzyme digestion to enrich key genomic sites simplifies operations and reduces costs. This technology can be used to develop kits or testing services and is widely used in multi-omics analysis, especially for precision medicine and disease research.
[0020] The purpose of the first aspect of the present invention is to provide a method for constructing a medium-throughput single-cell quadruple cytometry library.
[0021] The second aspect of the present invention aims to provide a single cell quadruple biological library.
[0022] The third aspect of the present invention aims to provide a sequencing method.
[0023] The purpose of the fourth aspect of the present invention is to provide applications of the construction method of the first aspect of the present invention and the sequencing method of the third aspect of the present invention.
[0024] In order to achieve the above object, the technical solution adopted by the present invention is:
[0025] The first aspect of the present invention provides a method for constructing a medium-throughput single-cell quadruple chemistry library, comprising the following steps:
[0026] (1) Lysing cells and separating to obtain total RNA solution and cell nucleus solution;
[0027] (2) constructing a transcriptome library for the total RNA solution in step (1), and at the same time, constructing a genome library for the cell nucleus solution in step (1) to obtain the single cell quadruple chemistry library.
[0028] In some embodiments of the present invention, the cells are subjected to in vitro methylation treatment using GpC methyltransferase before single cell lysis.
[0029] In some embodiments of the present invention, the in vitro methylation treatment comprises mixing the cells with a solution containing GpC methyltransferase, reacting, and inactivating the mixture.
[0030] In some embodiments of the present invention, the solution containing GpC methyltransferase further comprises at least one of an RNase inhibitor, a surfactant, Lambda DNA and S-adenosylmethionine.
[0031] In some embodiments of the present invention, the reaction system of the in vitro methylation treatment contains 0.5-20 U / μL GpC methyltransferase.
[0032] In some embodiments of the present invention, the reaction conditions are vortexing for 5 to 15 seconds and then incubating at 35 to 40° C. for 5 to 15 minutes.
[0033] In some embodiments of the present invention, the RNase inhibitor includes at least one of Murine RNase inhibitor, diethyl pyrocarbonate, guanidine isothiocyanate, vanadyl ribonucleoside complex, SDS, urea, diatomaceous earth or RNasin.
[0034] In some embodiments of the present invention, the surfactant includes IGEPAL CA-630.
[0035] Before separating the total RNA solution and the cell nucleus of the single cell, oscillation and centrifugation can be performed. Oscillation can help permeabilize the cell membrane, and centrifugation can prevent a small amount of liquid from adhering to the tube wall and causing RNA and DNA loss. The cell lysate with carboxylic acid magnetic beads is placed on a magnetic rack to separate the cell nucleus.
[0036] The introduction of cationic components in the nuclear fractionation step optimized the experimental conditions, increased the DNA retention rate and RNA conversion efficiency, and further improved the quality and reliability of single-cell multi-omics sequencing.
[0037] In some embodiments of the present invention, the transcriptome library construction comprises:
[0038] 1) RNA is reverse transcribed to obtain cDNA, pre-amplified, and then amplified by biotin PCR;
[0039] 2) fragmenting the product of biotin PCR amplification to obtain fragmented DNA;
[0040] 3) Perform end repair on the fragmented DNA and add A to the 3' end of the nucleic acid fragment;
[0041] 4) adding a double-stranded DNA adapter and a ligation reaction reagent to carry out a ligation reaction to obtain a library containing adapter ligations;
[0042] 5) Amplifying the library containing the adapter ligation to obtain a transcriptome library.
[0043] In some embodiments of the present invention, the RNA reverse transcription system includes TSO primers, RNA tags, reverse transcriptase, RNase inhibitors, enhancers and buffer solutions.
[0044] In some embodiments of the present invention, the reverse transcriptase comprises M-MLV reverse transcriptase, AMV reverse transcriptase, Hiscript III reverse transcriptase or a combination thereof.
[0045] In some embodiments of the present invention, the reverse transcriptase includes at least one of SmartScribe reverse transcriptase, MaximaHMinus reverse transcriptase, SuperscriptII reverse transcriptase, SuperscriptIII reverse transcriptase or AIpha reverse transcriptase.
[0046] In some embodiments of the present invention, the buffer comprises at least one of dNTPs and metal ions.
[0047] In some embodiments of the present invention, the RNA tag contains an 8-bit barcode sequence.
[0048] In some embodiments of the present invention, the nucleotide sequence of the RNA tag is as follows: TCAGAC GTG TGC TCTTCC GAT CTX XXX XXX XDD DDD DDD, wherein X represents an 8-bit barcode sequence.
[0049] In some embodiments of the present invention, the length of the fragmented DNA is 100 to 600 bp.
[0050] In some embodiments of the present invention, the RNase inhibitor includes at least one of Murine RNase inhibitor, diethyl pyrocarbonate, guanidine isothiocyanate, vanadyl ribonucleoside complex, SDS, urea, diatomaceous earth or RNasin.
[0051] In some embodiments of the present invention, the enhancer includes at least one of betaine, trehalose, glycerol, DMSO, polyethylene glycol, formamide, ammonium sulfate, tetramethylammonium chloride, gelatin, BSA, Triton X-100 or Tween.
[0052] In some embodiments of the present invention, the stabilizer includes at least one of BSA, sucrose, trehalose, polyethyleneimine, dithiothreitol or DMSO.
[0053] In some embodiments of the present invention, the metal ions include at least one of magnesium ions, manganese ions or calcium ions.
[0054] In some embodiments of the present invention, the pre-amplification reaction system includes IS PCR primers, 3'P2 primers, DNA polymerase, dNTPs and metal ions (such as Mg 2+ ) and buffer.
[0055] In some embodiments of the present invention, the nucleotide sequence of the 3'P2 primer is 5'-GTG ACT GGA GTTCAG ACG TGT GCT CTT CCG ATC-3' (SEQ ID NO: 2), and the final concentration in the reaction system is 0.05-0.2 μM.
[0056] In some embodiments of the present invention, the final concentration of the IS PCR primer in the reaction system is 0.05-0.2 μM.
[0057] In some embodiments of the present invention, the pre-amplification reaction is 78-82°C for 2-3 minutes; 78-82°C for 15-20 seconds, 49-53°C for 25-35 seconds, and 3-5 cycles; 59-61°C for extension for 4-6 minutes; 89-92°C for 18-25 seconds, 59-61°C for 15-20 seconds, and 10-15 cycles; 62-65°C for extension for 4-6 minutes; and maintained at 4°C.
[0058] In some embodiments of the present invention, the reaction system of biotin PCR amplification includes IS PCR primers and Biotin primers.
[0059] In some embodiments of the present invention, the nucleotide sequence of the IS PCR primer is 5'-AAG CAG TGG TATCAA CGC AGA GT-3' (SEQ ID NO: 2), and the final concentration in the reaction system is 0.2-0.5 μM.
[0060] In some embodiments of the present invention, the nucleotide sequence of the Biotin primer is 5'- / 5Biotin / CAAGCA GAA GAC GGC ATA CGA GAT CGT GAT GTG ACT GGA GTT CAG ACG TGT GCT CTT CCGATC-3' (SEQ ID NO: 4), and the final concentration in the reaction system is 0.8-1.5 μM.
[0061] In some embodiments of the present invention, the reaction procedure of the biotin PCR amplification is preheating at 83-86°C for 2-4 minutes; 96-98°C for 15-20 seconds, 58-62°C for 13-17 seconds, 62-67°C for 4-6 minutes, and cycled 3-5 times; extension at 63-67°C for 4-6 minutes; and holding at 4°C.
[0062] In some embodiments of the present invention, the fragmentation treatment includes fragmentation by ultrasound.
[0063] In some embodiments of the present invention, the length of the fragmented DNA is 100 to 600 bp.
[0064] In some embodiments of the present invention, the fragmented DNA is enriched before the end repair and the addition of A to the 3' end of the nucleic acid fragment. The enrichment method includes mixing the fragmented DNA with magnetic beads (such as C1 streptavidin beads), incubating at room temperature for 15 to 30 minutes, and performing magnetic separation.
[0065] In some embodiments of the present invention, KAPA HyperPrep Kit is used for end repair and A addition, and the reaction conditions are incubation at 24-26° C. for 28-35 minutes, incubation at 62-66° C. for 18-23 minutes, and maintenance at 4° C.
[0066] In some embodiments of the present invention, the ligation reaction reagents include KAPA HyperPrep Kit.
[0067] In some embodiments of the present invention, the amplification reaction system described in 5) includes KAPA HiFi HotStartReadyMix, QP2 primers and short universal primers; the reaction program is 89-93°C for 25-35 seconds; 94-96°C for 13-18 seconds, 58-63°C for 25-35 seconds, 64-66°C for 28-23 seconds, and cycled 5-15 times; extension at 64-66°C for 1-2 minutes and maintained at 4°C.
[0068] In some embodiments of the present invention, the nucleotide sequence of the QP2 primer is 5'-CAA GCA GAA GACGGC ATA CGA-3' (SEQ ID NO: 5), and the final concentration in the reaction system is 0.1-0.5 μM.
[0069] In some embodiments of the present invention, the nucleotide sequence of the short universal primer is 5'-AAT GAT ACG GCGACC ACC GAG ATC TAC ACT CTT TCC CTA CAC GAC-3' (SEQ ID NO: 6), and the final concentration in the reaction system is 0.1-0.5 μM.
[0070] In some embodiments of the present invention, the genomic library construction comprises:
[0071] a) Protease digestion and enzyme cleavage of cell nucleus solution;
[0072] b) ligation of methylated adapters, purification by Biotin, and then bisulfite conversion to obtain single-stranded DNA;
[0073] c) performing linear amplification on single-stranded DNA;
[0074] d) performing library amplification on the amplified product of c) so that two different oligonucleotide sequences are added to the positive and reverse strands of the DNA fragment.
[0075] The DNA partial method achieves high sensitivity and sample multiplexing through an early labeling step, which is suitable for single cells and small sample amounts, enhancing the scalability of the technology.
[0076] In some embodiments of the present invention, the enzyme used for the enzymatic cleavage comprises at least one of AluI, TaqαI, SphI, MspI, BamHI, ApeKI, BanII, HpaII, HpyCH4V, BstNI, HaeIII, HpyCH4III, BglII, BssSI and KpnI.
[0077] In some embodiments of the present invention, the linker sequence connected by the methylated linker in b) is a Y-shaped double-stranded structure formed by annealing two single-stranded nucleic acids.
[0078] In some embodiments of the present invention, all cytosine nucleotides in the two single-stranded nucleic acids are methylated; and the 5' end of the bottom nucleic acid strand of the Y-shaped double-stranded structure formed after the two single-stranded nucleic acids are annealed is phosphorylated.
[0079] In some embodiments of the present invention, the 3' end of the Y-shaped double-stranded structure is a tag sequence having 8 nucleotides.
[0080] In some embodiments of the present invention, the nucleotide sequence of the upper chain of the linker sequence is as shown in SEQ ID NO:7, specifically 5'- / 5SpC3 / GGAGTT / iMe-dC / AGA / iMe-dC / GTGTG / iMe-dC / T / iMe-dC / TT / iMe-dC / / iMe-dC / GAT / iMe-dC / TTGGTATAG-3', and the nucleotide sequence of the lower chain of the linker sequence is as shown in SEQ ID NO:8, specifically 5'CG CTATACCA AG ATC GGA AGA GCA CAC GTC TGA ACT CC / 3Bio / -3'.
[0081] In some embodiments of the present invention, the reaction system for methylated linker ligation contains ATP, ligase, restriction endonuclease and buffer.
[0082] In some embodiments of the present invention, the ligase can catalyze the formation of a phosphodiester bond between the 5'-P end and the 3'-OH end between or within a single-stranded oligonucleotide or a single nucleotide, and can be selected from at least one of E. coli DNA Ligase, T4 RNA ligase 1, TS2126 RNA ligase, single-stranded DNA / RNA circular ligase (ssDNA / RNA CircLigase) or a truncated version of T4 RNA ligase 2.
[0083] In some embodiments of the present invention, the restriction endonuclease includes HpaII and / or MspI; preferably MspI.
[0084] The present invention applies the MspI enzyme to copy number variation (CNV) detection for the first time, which not only expands the application scope of the technology but also significantly improves the overall performance.
[0085] In some embodiments of the present invention, the buffer comprises NEB CutSmart buffer.
[0086] In some embodiments of the present invention, the reaction procedure for the methylated linker connection is incubation at 35-38° C. for 1-1.5 hours, incubation at 218-22° C. for 1-1.5 hours, and maintenance at 4° C.
[0087] In some embodiments of the present invention, the Biotin purification includes purification using C1 streptavidin beads.
[0088] In some embodiments of the present invention, the bisulfite conversion can be performed using conventional methods in the art, such as using a Zymo CT conversion kit to perform bisulfite conversion on a DNA sample.
[0089] By utilizing the representative bisulfite sequencing technology, we can more accurately focus on specific gene regions at the same sequencing amount, thereby improving the sensitivity and specificity of detection.
[0090] In some embodiments of the present invention, the linear amplification described in c) comprises the following steps: single-stranded DNA is mixed with NEBBuffer, dNTP and random primers, reacted at 90-98°C for 30-52 seconds, an enzyme with strand displacement activity is added on ice, the temperature is raised from 4°C to 37°C to promote base pairing, and incubated at 37°C for 60-100 minutes, wherein the temperature is raised by 1°C every 5-15 seconds.
[0091] In some embodiments of the present invention, the nucleotide sequence of the random primer is shown in SEQ ID NO:9.
[0092] In some embodiments of the present invention, the enzyme with strand displacement activity includes at least one of klenow polymerase (3'→5'exo-), klenow polymerase, bst DNA polymerase, vent DNA polymerase (3'→5'exo-), vent DNA polymerase, Phi 29 DNA polymerase, deep vent DNA polymerase (3'→5'exo-), and deep vent DNA polymerase.
[0093] In some embodiments of the present invention, the oligonucleotide sequence in d) is P5 or P7, and the nucleotide sequence is shown in SEQ ID NO:11 (5'-AAT GAT ACG GCG ACC ACC GAG ATC TAC ACT CTT TCC CTA CAC GAC GCTCTT CCG ATC T-3') or SEQ ID NO:10 (5'-CAA GCA GAA GAC GGC ATA CGA GAT GTG ACTGGA GTT CAG ACG TGT GC TCT T-3').
[0094] In some embodiments of the present invention, the cells are single cells and / or multi-cell samples.
[0095] In some embodiments of the present invention, when the cells are a multi-cell sample, during the construction of the transcriptome library, the products of the pre-amplification of each single cell are mixed before biotin PCR amplification; during the construction of the genomic library, before bisulfite conversion, the products of the biotin purification of each single cell are mixed.
[0096] By performing sample mixing operations on the DNA methylation part, the cost is significantly reduced and the process is simplified, which improves the economy and efficiency of sequencing.
[0097] The second aspect of the present invention provides a single-cell quartet biochemical library constructed by the construction method of the first aspect of the present invention.
[0098] The third aspect of the present invention provides a sequencing method, comprising the step of sequencing the single-cell quadruple cytokine library of the second aspect of the present invention.
[0099] In some embodiments of the present invention, the sequencing method includes any one of illumina, DNB, ABI, 454FLX, and SOLiD sequencing.
[0100] In some embodiments of the present invention, the sequencing method further comprises analyzing the data obtained by sequencing.
[0101] In some embodiments of the present invention, the analysis includes analyzing the data to obtain at least one of the following data: number of reads of offline data, Q30 of offline data, Barcode splitting rate, GC content, CV of single sample data, average sequencing depth of single sample, alignment rate, genome coverage, methylation C conversion rate, number of CpG sites, number of WCG and GCH sites, CNV and chromatin development. Of course, those skilled in the art can also increase or decrease the obtained analysis results according to the use requirements and purposes.
[0102] In some embodiments of the present invention, before sequencing, the sequencing method may further include performing a quality check on the library.
[0103] Compared with the scRRBS technology, the sequencing method provided by the present invention has the following advantages: First, in terms of coverage, although scRRBS improves the detection efficiency of CpG sites through enzyme enrichment, the genomic region covered is still limited, mainly concentrated in CpG islands and gene regulatory regions. In contrast, the present invention provides a wider genome coverage through mspl enzymes and multiple purification methods, and can capture more diverse DNA methylation information. Secondly, in terms of sensitivity, the scRRBS technology may face problems of signal loss and data quality degradation when processing very small amounts of DNA. The present invention significantly improves the detection sensitivity by optimizing the experimental process and introducing innovative data analysis strategies, and reliable methylation information can be obtained even in the case of extremely low starting amounts. In addition, the scRRBS technology is mainly limited to DNA methylation analysis and is difficult to provide other epigenetic information. In contrast, the present invention, as a comprehensive single-cell multi-omics sequencing technology, can not only analyze DNA methylation, but also simultaneously obtain multi-dimensional data such as chromatin state, nucleosome positioning, copy number variation and ploidy, providing a more comprehensive perspective for in-depth understanding of cell heterogeneity and epigenetic regulatory mechanisms. In terms of throughput, the present invention significantly improves sample processing capabilities by introducing innovative labeling strategies and optimized experimental processes, making it possible to analyze a large number of single cells simultaneously. Finally, in terms of cost, scRRBS technology still faces economic pressure when conducting large-scale studies at the single-cell level, and the present invention further reduces the cost of single-cell sequencing through multiple technical innovations and process optimization, making large-scale epigenomic research more economically feasible.
[0104] Compared with iscCOOL-seq, the sequencing method provided by the present invention has the following advantages: First, the experimental cost of iscCOOL-seq is high, which limits its application in large-scale research. In contrast, the present invention significantly reduces the sequencing cost of each cell by optimizing the experimental process and introducing innovative strategies, making large-scale single-cell epigenomic research more economical and feasible. Secondly, the sample processing throughput of iscCOOL-seq is relatively limited, which is difficult to meet the needs of high-throughput research. The present invention overcomes this shortcoming, and by introducing a labeling strategy, mixed processing of samples is achieved, the experimental throughput is greatly improved, and a large number of single-cell samples can be processed simultaneously. In addition, the present invention further improves the sequencing efficiency by introducing a strategy of enzyme digestion enrichment of CCGG sites. This method not only increases the sequencing depth of the target region, but also reduces the proportion of non-specific sequencing, thereby further reducing the sequencing cost while ensuring data quality. This enrichment strategy enables the present invention to more efficiently capture the epigenetic information of key regulatory regions, providing a more accurate and economical single-cell epigenomic analysis solution.
[0105] The fourth aspect of the present invention provides the use of the construction method of the first aspect of the present invention and the sequencing method of the third aspect of the present invention in any one of (1) to (3):
[0106] (1) Analysis of copy number variation and / or chromatin accessibility of single-cell genomes from multiple samples;
[0107] (2) DNA methylation and / or single-site mutation analysis of multiple samples and single cells;
[0108] (3)Multi-sample single-cell gene expression detection.
[0109] The beneficial effects of the present invention are:
[0110] The present invention provides a middle throughput single-cell multi-omics library construction and sequencing technology mtCOOL-seq (middle throughput single-cell multi-omics sequencing technology). The four-omics include: transcriptome (analyzing the expression of all RNA molecules in cells), DNA methylome (studying the influence of DNA methylation modification on gene expression), copy number variation (CNV, detecting the change of the copy number of DNA fragments in the genome) and chromatin openness (evaluating the relationship between the looseness of chromatin structure and gene regulation). The combination of these omics can fully reveal the gene expression regulation mechanism and functional state of a single cell. The mtCOOL-seq technology overcomes the shortcomings of existing technologies such as scRRBS and iscCOOL-seq, and integrates their advantages to provide a more powerful and practical tool for single-cell multi-omics research.
[0111] Specifically, efficient medium-throughput operation process: The construction method of the present invention realizes an efficient medium-throughput operation process, enabling operators to process up to 96 cells simultaneously in a single reaction system, greatly improving the experimental throughput. This method supports the use of different cell-specific markers, which is convenient for comparing batch effects, technical replicates, biological replicates and other experimental conditions, and is also convenient for measuring more single cells in the same sample. Through the second-generation sequencing technology, a large amount of single-cell methylation data can be obtained, and combined with bioinformatics analysis, the DNA methylation map of each cell can be accurately depicted.
[0112] Better data quality: In terms of data quality, the present invention reduces the sample operation steps and increases the total DNA amount in processes such as DNA conversion by optimizing the technical process, effectively reducing the risk of DNA damage and loss. The designed label connector improves the consistency of sample processing and effectively reduces the difference in coverage between samples. It is particularly worth mentioning that the present invention introduces a number of innovative modifications in the DNA processing link, among which the MspI enzyme is applied to CNV detection for the first time, which not only expands the scope of technical application, but also significantly improves the overall performance. At the same time, the latest cationic component application is adopted to further optimize the experimental conditions and improve the retention rate and conversion efficiency of DNA.
[0113] Low cost and high efficiency: The present invention also significantly reduces experimental costs and improves efficiency. By using mixed sample operations in the DNA methylation part, the sequencing cost is significantly reduced and the process is simplified. Using representative bisulfite sequencing technology, it is possible to focus more accurately on the promoter and enhancer regions of the gene of interest, avoiding the average sequencing of the whole genome, and providing an efficient and economical single-cell multi-omics detection solution. Compared with the prior art, the present invention has achieved significant improvements in both throughput and cost, making large-scale single-cell research more feasible and economical.
[0114] Multi-omics integration: Single-cell DNA methylome sequencing technology has achieved mid-throughput effect improvement, while also enabling single-cell multi-omics sequencing, and invented a low-cost, high-performance mid-throughput single-cell quadruple-omics sequencing solution. This integrated approach not only improves the comprehensiveness of the data, but also provides a powerful tool for revealing cellular heterogeneity and complex epigenetic regulatory mechanisms.
[0115] Unique advantages of the RNA part: In terms of RNA sequencing, the present invention uses 3' end sequencing technology, which has many advantages in principle. It can quantify gene expression levels more accurately, avoids deviations caused by RNA degradation or partial transcripts, and is particularly suitable for detecting low-abundance transcripts. At the same time, 3' end sequencing reduces the sequencing depth requirements and further reduces costs. In addition, this method can also effectively identify the diversity of the 3' end of the gene, providing a new perspective for studying gene expression regulation. BRIEF DESCRIPTION OF THE DRAWINGS
[0116] Figure 1 Schematic diagram of the process for constructing a medium-throughput single-cell quadruple chemistry library of the present invention.
[0117] Figure 2 This is the clustering heat map of K562 cells in Example 3.
[0118] Figure 3 This is the correlation heat map of K562 cells in Example 3. DETAILED DESCRIPTION
[0119] The present invention is further described in detail below through specific examples.
[0120] It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0121] In order to make the purpose, technical scheme and advantages of the embodiments of the present invention clearer, the technical scheme in the embodiments of the present invention will be described clearly and completely below. If the specific conditions are not specified in the embodiments, they are carried out according to conventional conditions or conditions recommended by the manufacturer. If the manufacturer of the reagents or instruments used is not specified, they are all conventional products that can be purchased commercially.
[0122] Terminology explanation:
[0123] Transcriptome: refers to the sum of all RNA molecules in a cell or a group of cells at a specific time point and under specific conditions. It mainly includes messenger RNA (mRNA), as well as other non-coding RNAs such as ribosomal RNA (rRNA), transfer RNA (tRNA) and small RNAs (such as miRNA). Transcriptome analysis can reveal gene expression levels and help understand cell functions and response mechanisms.
[0124] Methylome: refers to the collection of all DNA methylation modifications in an organism or a specific cell. DNA methylation usually occurs on cytosine nucleotides, especially in CpG dinucleotides, affecting gene expression and genome stability. Methylome analysis helps to study gene regulation, developmental processes and disease mechanisms such as cancer.
[0125] Chromatin accessibility: refers to the looseness of chromatin structure, which affects the accessibility of transcription factors and other regulatory proteins to DNA. Open chromatin regions are usually associated with active gene expression, while tightly packed chromatin may inhibit gene expression. Studying chromatin accessibility can help understand gene regulation and cell differentiation processes.
[0126] Cellular heterogeneity: refers to the differences in gene expression, morphology, function, etc. between different cells in the same tissue or sample. This heterogeneity is important in both normal physiological and disease states, especially in cancer, where the heterogeneity of tumor cells can affect treatment efficacy and prognosis.
[0127] Tagging Technology: Tagging technology involves attaching markers to molecules or cells to facilitate identification, separation, or analysis. For example, fluorescent tags can be used to visualize specific proteins or nucleic acid sequences. This technology is widely used in biological research, including molecular imaging, flow cytometry, and high-throughput sequencing.
[0128] Enzymatic Digestion: refers to the process of cutting DNA using specific restriction endonucleases. These enzymes recognize specific DNA sequences and cut at these sites, thereby enriching the genomic regions of interest. This process is used to analyze and study genome structure, helping to improve sequencing efficiency and reduce costs.
[0129] Copy Number Variation (CNV): refers to the phenomenon that the number of copies of DNA fragments in the genome changes. CNV can lead to gene duplication or deletion, affecting gene expression and function, and is an important source of genetic diversity and disease. CNV analysis helps to study genetic diseases, cancer mechanisms, and individual responses to drugs.
[0130] The features and performance of the present invention are further described in detail below in conjunction with the embodiments.
[0131] Example 1
[0132] A middle-throughput single-cell multi-omics library construction and sequencing technology mtCOOL-seq (middlethroughput single-cell multi-omics sequencing technology) includes the following steps (see flowchart for details) Figure 1 ):
[0133] (1) Preparation of cell lysis buffer
[0134] Add 2.5 μL of lysis mixture (containing 0.05 U / μL RNase inhibitor (Cat. No.: Takara2313A), 1× GC reaction buffer (M.CviPI enzyme supporting reagent), 0.1 mM SAM, 1% IGEPAL CA-630, 15 pg / μL Lambda DNA, and 3 U / μL M.CviPI) to each cell. Place a single cell in the lysis buffer, vortex for 10 seconds, then incubate at 37°C for 10 minutes, and then inactivate at 75°C for 5 minutes. After treatment, the sample directly proceeds to the next step.
[0135] (2) Single cell lysis and RNA release
[0136] Wash with Dynabeads Myone Carboxylic Acid beads (hereinafter referred to as CA beads), vortex for 10 seconds and then tilt and rotate for 2 minutes. Transfer the beads to a test tube and add 200μL of 4×SuperScript buffer for washing. Subsequently, resuspend the beads with a mixture containing 2.11U / μL RNase inhibitor, 0.1% Tween 20, SuperScript buffer, 2mM DTT, 0.1% Triton X-100, 1mM MgCl2 and water. Add 1μL of CA bead mixture to each cell. Finally, lyse single cells and vortex 10-15 times to release RNA. The sample was incubated at room temperature for 2 minutes, then centrifuged for 2 minutes and placed on a magnetic rack for 2 minutes to complete the treatment.
[0137] RNA part:
[0138] (3) RNA reverse transcription
[0139] For the RNA portion of the supernatant, add the following mixed solution: it contains 1.5M betaine, 1mM dNTP, Hiscript III Reverse Transcriptase, 2U / μL RNase inhibitor, 1μM TSO and different RNA tags (5'-TCA GAC GTG TGC TCT TCC GAT CTX XXX XXX XDD DDD DDD-3'(SEQ ID NO:1), where X represents the 8-bit barcode sequence). Gently pipette up and down to mix, avoiding bubbles. Then perform the following temperature-controlled reaction: incubate at 25℃ for 2 minutes, incubate at 55℃ for 10 minutes, incubate at 60℃ for 10 minutes, incubate at 65℃ for 10 minutes, and then maintain at 4℃. Finally, store the sample below -4℃ to pause the reaction.
[0140] (4) PCR amplification
[0141] First, add PCR mixture containing 1x KAPA HiFi HotStart ReadyMix, 0.1μM ISPCR primer, 0.1μM 3'P2 and water in a total volume of 30μL. IS PCR primer sequence: 5'-AAG CAG TGG TAT CAA CGCAGA GT-3' (SEQ ID NO: 2); 3'P2 primer sequence: 5'-GTG ACT GGA GTT CAG ACG TGT GCT CTTCCG ATC-3' (SEQ ID NO: 3). The temperature and time settings are as follows: heating at 80°C for 2 minutes, then 80°C for 20 seconds, 50°C for 30 seconds, and 4 cycles; followed by extension at 60°C for 5 minutes. Again at 90°C for 20 seconds, 60°C for 15 seconds, and 12 cycles; finally, extend at 65°C for 5 minutes, again at 65°C for 5 minutes, and finally hold at 4°C.
[0142] (5) Sample mixing and purification
[0143] First, the PCR products were combined and mixed. Purified using a DNA Clean&Concentrator kit and then eluted in 100 μL of water. Then, purified once using 0.8×Agencourt AMPure XP beads (hereinafter referred to as XP beads) and eluted in 10 μL of water.
[0144] (6) Biotin PCR amplification
[0145] In biotin PCR amplification, 10 to 70 ng of DNA was used as a template, and the PCR reaction was performed using IS PCR primers (SEQ ID NO: 2) and Biotin primers. Biotin primer sequence: 5'- / 5Biotin / CAA GCA GAA GAC GGC ATACGA GAT CGT GAT GTG ACT GGA GTT CAG ACG TGT GCT CTT CCG ATC-3' (SEQ ID NO: 4). The reaction mixture contained 1x KAPA HiFi HotStart ReadyMix, 0.3 μM IS PCR primer, 1 μM Biotin primer and DNA template in a total volume of 80 μL. The temperature was set as follows: preheating at 85°C for 3 minutes, followed by 98°C for 15 seconds, 60°C for 15 seconds and 65°C for 5 minutes, 4 cycles, and finally extension at 65°C for 5 minutes and holding at 4°C.
[0146] After PCR was completed, the DNA was purified once with 0.8× XP magnetic beads and eluted in 50 μL elution buffer.
[0147] (7) Ultrasonic interruption
[0148] 10 μL of DNA was mixed with 100 μL of elution buffer and sonicated using the DNA300 program of Covaris at room temperature of 22° C. for 1 minute to obtain fragmented DNA of 100 to 400 bp.
[0149] After ultrasonic shearing, the DNA was purified using a DNA Clean&Concentrator kit and then eluted in 25 μL of water. It was then purified once more using 1×XP magnetic beads and eluted in 25 μL of elution buffer.
[0150] (8) Biotin enrichment
[0151] During DNA biotin enrichment, 30 μL of C1 streptavidin beads (hereinafter referred to as C1 beads) were used for each sample. The beads were first washed once with 100 μL of 1× B&W buffer, and then resuspended with 15 μL of 1× B&W buffer. 25 μL of resuspended beads were mixed with 15 μL of DNA fragments and incubated with rotation at room temperature for 15 to 30 minutes. After incubation, each tube was placed on a magnetic rack and the supernatant was discarded. 100 μL of 1× B&W buffer was added to wash the beads, and then 100 μL of elution buffer was used to wash the beads. Finally, the beads were resuspended in 25 μL of water.
[0152] (9) End repair and A-tailing
[0153] During the end repair and A-tailing process, each reaction is assembled in a tube or well of a PCR plate to generate double-stranded DNA fragments that are end-repaired, 5' phosphorylated, and 3'-dA-tailed. End repair and A-tailing are performed using the KAPA HyperPrep Kit and follow its instructions in a total volume of 50μL. The reaction conditions are incubation at 25℃ for 30 minutes, then incubation at 65℃ for 20 minutes, and finally hold at 4℃. The buffer and enzyme mixture can be pre-mixed and added in a single pipetting step. After thorough mixing, centrifuge lightly to ensure that the liquid is at the bottom of the tube. Next, incubate according to the temperature control program of incubation at 20℃ for 30 minutes, incubation at 65℃ for 20 minutes, and hold at 4℃, and immediately proceed to the next step.
[0154] (10) Connector connection
[0155] During the adapter ligation process, double-stranded DNA adapters (provided in the KAPA HyperPrep Kit, Cat. No. E7335S) are ligated to the 3'-dA-tailed molecules. The reaction mixture contains 10 μL of end-repaired and A-tailed products and adapters, and is prepared according to the KAPA HyperPrep Kit instructions, with a total volume of 40 μL. After thorough mixing, centrifuge lightly to ensure that the liquid is at the bottom of the tube. Then, incubate at 20°C for 20 minutes and pause at 4°C. Next, add 1U of Enzyme, incubate at 37°C for 20 min, pause at 4°C.
[0156] During the post-ligation cleanup, place the product on a magnetic rack and aspirate the supernatant. Then resuspend the beads in 50 μL of EB buffer, place the resuspended beads on a magnetic rack again and aspirate the supernatant. Finally, add 15 μL of water to the beads, vortex and centrifuge.
[0157] (11) Library amplification
[0158] During library amplification, the reaction mixture contained 22 μL of adapter-ligated library, 1× KAPA HiFiHotStart ReadyMix, 0.3 μM QP2 primer and 0.3 μM short universal primer in a total volume of 50 μL. QP2 primer sequence: 5'-CAA GCA GAAGAC GGC ATA CGA-3' (SEQ ID NO: 5), short universal primer sequence: 5'-AAT GAT ACG GCGACC ACC GAG ATC TAC ACT CTT TCC CTA CAC GAC-3' (SEQ ID NO: 6). After thorough mixing, centrifugation was performed to ensure that the magnetic beads were in suspension. Amplification was performed according to the following cycle program: initial denaturation at 90°C for 30 seconds, followed by denaturation at 95°C for 15 seconds, annealing at 60°C for 30 seconds, extension at 65°C for 20 seconds, 5-15 cycles, and a final extension at 65°C for 1 minute, and a hold at 4°C.
[0159] After amplification, place the reaction on a magnetic stand and transfer the supernatant to a new tube. Then perform a post-amplification cleanup, purify twice using 40 μL 0.8x XP magnetic beads, and finally elute in 50 μL water to obtain the RNA library.
[0160] DNA part:
[0161] (3) Protease digestion
[0162] For the DNA portion, 4.1 μL of protease digestion buffer (composed of 3 μL M-digestion buffer, 0.1 μL 10 mg / mL Qiagen Protease, and 1 μL water) was added to the magnetic beads with captured DNA, followed by incubation at 55°C for 2 hours and then at 70°C for 30 minutes.
[0163] (4) Enzyme digestion and label ligation
[0164] In the experiment, the adapter consisted of a methylated upper strand and a biotinylated lower strand, which were resuspended in Tris EDTA (TE) buffer at a concentration of 50 μM and annealed before use. All adapters and primers were from Integrated DNA Technologies (IDT). The annealing process was briefly as follows: equimolar volumes of each adapter were mixed, heated to 95°C for 3 minutes, and then slowly cooled to 4°C at a rate of 0.5°C / second. The methylated upper strand contained a partial SBS12 sequence followed by an 8-base C-deficient tag (upper strand sequence 5'- / 5SpC3 / GGAGT T / iMe-dC / A GA / iMe-dC / GTG TG / iMe-dC / T / iMe-dC / TT / iMe-dC / / iMe-dC / GAT / iMe-dC / TTGGTATAG-3' (SEQ ID NO:7)). The biotinylated lower strand is complementary to the methylated upper strand and has two more bases (5'-CG-3') at the 5' end, which is complementary to the sticky end left by the MspI enzyme (lower strand sequence 5'-CG CTATACCA AG ATC GGA AGA GCA CAC GTC TGA ACT CC / 3Bio / -3' (SEQ ID NO: 8)).
[0165] Add 3 μL of digestion and ligation reagents to each well containing 3 μL of purified DNA or lysis buffer containing sorted cells. The final reaction contains 3mM ATP, 1× CutSmart buffer, 0.13 μL water, 500U / mL T4 DNA ligase, 2U / mL MspI enzyme, and 10nM annealed tag adapter. The reaction is incubated at 37°C for 1 hour, then at 20°C for 1 hour, and finally kept at 4°C.
[0166] (5) Biotin purification (biotin enrichment)
[0167] During the binding and sample combination process, the C1 magnetic beads were first washed three times with 200 μL 1× Bind&Wash (B&W) buffer according to the kit instructions, and then resuspended with 1× B&W buffer. Then, 16 μL of C1 magnetic bead mixture was added to the corresponding sample and incubated at room temperature on a rotator for 15 minutes.
[0168] (6) Mixed samples
[0169] For multiplexed single-cell libraries, combine the C1 beads from wells with different labels and place them in a centrifuge tube. Use a magnetic stand to separate the beads from the enzyme solution and resuspend them in 30 μL of water. Heat the resuspended beads to 95°C for 8 minutes and then immediately place them on ice to cool. Then centrifuge at 10,000 relative centrifugal force (rcf) at 4°C for 10 minutes. Finally, place the tube on the magnetic stand for 3 to 4 minutes and aspirate the supernatant into a new tube.
[0170] (7) Bisulfite conversion
[0171] According to the instructions, the DNA samples were bisulfite converted using the Zymo CT conversion kit. Specifically, the reaction conditions during the conversion were: 95°C for 8 minutes, 64°C for 3.5 hours, and then stored at 4°C (within 20 hours).
[0172] After bisulfite purification, the DNA was purified using a DNA Clean&Concentrator kit and eluted in 26.5 μL of elution buffer (EB).
[0173] (8) One-chain synthesis and two-chain synthesis
[0174] During the one-strand synthesis and purification process, the bisulfite-converted single-stranded DNA was mixed with 1x NEB Buffer, 0.4mM dNTP mixture and 1μM random primers in a total reaction volume of 31μL, including 26μL of converted DNA and 5.12μL of reaction mixture (containing 0.64μL of 100μM random primers, 3.2μL of 10x NEB buffer and 1.28μL of 10mM dNTPs). Random primer sequence: TAC ACG ACG CTC TTC CGA TCT NNN NN (SEQ ID NO:9), where N is a random base. The solution was heated at 95°C for 45 seconds and immediately transferred to ice. Then, 1μL of Klenowexo– with strand displacement activity was added. Base pairing was promoted by gradually increasing the temperature from 4°C to 37°C (1°C increase every 15 seconds) and incubated at 37°C for 90 minutes.
[0175] In the four-cycle system, the next step is to incubate at 95°C for 1 minute to denature the DNA, followed by a pause at 4°C. After that, 2.5 μl of a mixture containing 1 μl of 25 pmoles of random primers (made by diluting 10 μmoles of random primers to 25 nanomoles and then diluting to 1:1000 with 1× NEB buffer 2), 0.5 μl of Klenow exo-enzyme (50 units / μl, from NEB), and 1 μl of 2.5 nanomoles of each dNTP was added to each sample. This mixture was incubated at 4°C for 5 minutes, then the temperature was increased stepwise by 1°C every 15 seconds until it reached 25°C, and then incubated at 25°C for 5 minutes, and then gradually increased to 37°C and incubated for 30 minutes. This cycle was repeated 3 more times.
[0176] Next, XP magnetic bead purification was performed by adding 32 μL of XP magnetic beads and eluting with 12 μL of elution buffer, of which 10.5 μL of the eluate was used for the final library PCR.
[0177] (9) Library amplification
[0178] During the library PCR process, 13.5 μL of the mixture was added to the reaction tube and quickly centrifuged. The mixture included 0.1 μL of 10 μM i5-index primer, 0.3 μL of 15 μM i7-index primer, and 12.5 μL of 2×KAPA HiFiHotStart ReadyMix. 6+16 cycles of PCR were performed, with the specific program being: pre-denaturation at 95°C for 2 minutes, denaturation at 95°C for 1 minute; 3 to 8 cycles (95°C for 20 seconds, 58°C for 30 seconds, 65°C for 1 minute); followed by 10 to 18 cycles (95°C for 20 seconds, 65°C for 30 seconds, 65°C for 1 minute); and finally extension at 65°C for 3 minutes and hold at 4°C. The primer sequences used are: P7 primer with i7 index: 5'-CAA GCA GAA GAC GGC ATA CGA GAT-i7-GTG ACT GGA GTT CAG ACG TGT GC TCT T-3' (SEQ ID NO: 10); P5 primer without i5 index: 5'-AAT GAT ACG GCG ACC ACC GAG ATC TAC ACTCTT TCC CTA CAC GAC GCT CTT CCG ATC T-3' (SEQ ID NO: 11); P5 primer with i5 index: 5'-AAT GAT ACG GCG ACC ACC GAG ATC TCAC-i5-AC ACT CTT TCC CTACAC GAC GCT CTT CCGATC T-3'.
[0179] The excess library primers were removed using 1×XP magnetic beads to obtain a DNA library.
[0180] (10) Sequencing
[0181] Both DNA and RNA libraries were isolated using Illumina's NovaSeq TM X Plus sequencer for paired-end 150 bp sequencing.
[0182] Example 2
[0183] In order to verify the effectiveness and reliability of the mtCOOL-seq technology in Example 1 in single-cell DNA sequencing, this example selected the human chronic myeloid leukemia cell line K562 as an experimental model. K562 cells are a cell line widely used in molecular biology and hematology research, with a stable gene expression profile, and therefore become an ideal object for evaluating new single-cell sequencing technologies.
[0184] During the experiment, K562 cell lines were obtained from ATCC and cultured according to the standard method provided by ATCC to ensure that the cells were in the best growth state. A mouth pipette was used to pick and separate single cells into lysate. Subsequently, the experiment was carried out according to the DNA processing process of the mtCOOL-seq technology in Example 1. Following the mtCOOL-seq process, single-cell DNA was subjected to bisulfite conversion treatment, followed by amplification and library construction. Finally, NovaSeq TM The X Plus sequencer high-throughput sequencing platform performs deep sequencing on the constructed library to obtain high-quality DNA methylation data.
[0185] During the data processing phase, a standard bioinformatics analysis process was used. First, the sequencing data was quality controlled to remove low-quality reads. Then, the high-quality data was aligned to the reference genome to identify methylation sites. Next, a specialized software tool was used to quantitatively analyze the methylation levels, and the results were statistically validated to ensure the reliability and accuracy of the data.
[0186] This example used 5 to 48 cells and conducted multiple experiments using the DNA portion of the mtCOOL-seq technical process in Example 1, and obtained the following results:
[0187] (1) In terms of data matching rate, the mtCOOL-seq technology of Example 1 achieved a significant improvement, reaching 62.8%, which means that the sequencing method of Example 1 of the present invention can more accurately and comprehensively capture and analyze single-cell DNA information, thereby improving the reliability and utilization of the data.
[0188] (2) In terms of genome coverage, the genome coverage per Gb of data reached 6.4%. Higher coverage can not only provide more comprehensive genome information, but also significantly improve the detection ability of rare variations, providing strong support for in-depth research on genome structure and function. At the same time, the number of CpG sites detected per Gb of data reached 690,962; due to technical improvements and the addition of in vitro methylation treatment, the present invention can detect an average of 19,549,766 WCG sites and 1,838,811 GCH sites per Gb, which is something that scRRBS technology cannot do; the average data efficiency (data efficiency = data volume after cleaning / total data volume) is 89%.
[0189] (3) In terms of economy, the mtCOOL-seq technology of Example 1 significantly reduces the cost of single-cell DNA sequencing. This breakthrough not only improves the feasibility of the technology, but also creates conditions for the development of large-scale single-cell research projects, which is expected to accelerate the progress of biomedical research. Compared with the iscCOOL-seq technology, the present invention reduces the cost of the single-cell DNA part from 56 yuan / cell to 13.95-15.57 yuan / cell (16-96 cell mixed samples). The larger the number of cells in a single mixed sample, the lower the cost of a single cell.
[0190] In summary, the mtCOOL-seq technology of Example 1 has excellent performance in single-cell DNA sequencing and methylation analysis. This technology shows obvious advantages in detecting CpG, WCG and GCH sites, greatly expanding the breadth and depth of analysis, and providing a powerful tool for epigenetic research. Through innovative improvements, mtCOOL-seq solves the PolyG problem and significantly improves data quality and reliability. This lays a solid foundation for subsequent analysis. Overall, mtCOOL-seq shows significant advantages in data quality, genome coverage, methylation detection, economy and reliability. It is an efficient, economical and reliable technology that promotes single-cell omics research and provides strong support for biomedicine, disease diagnosis and personalized medicine.
[0191] Example 3
[0192] In order to verify the effectiveness and reliability of the mtCOOL-seq technology in Example 1 in single-cell RNA sequencing, this example also selected the K562 cell line as an experimental model. The inventors designed a series of gradually expanded experiments, and conducted four independent tests (Test1, Test2, Test3, and Test4), using 8, 11, 16, and 48 K562 single cell samples, respectively. This gradient increase experimental design is intended to evaluate the performance and stability of the mtCOOL-seq technology under different sample sizes.
[0193] During the experiment, K562 cell lines were first obtained from ATCC and cultured according to the standard method provided by ATCC to ensure that the cells were in the best growth state. A mouth pipette was used to pick individual cells into the lysate, and then the RNA part of the mtCOOL-seq technology was processed according to Example 1. Following the mtCOOL-seq experimental process, the mixed cDNA samples were amplified and the library was prepared, including a series of steps such as mtCOOL-seq specific fragmentation method, dedicated adapter connection and optimized PCR amplification steps. Finally, NovaSeq was used. TM The X Plus sequencer high-throughput sequencing platform performs deep sequencing on the constructed library to obtain high-quality transcriptome data.
[0194] After sequencing is completed, the obtained data is fully processed according to the bioinformatics analysis process. First, the raw data is strictly quality controlled and filtered using quality control tools to ensure the high quality of the data. Subsequently, the screened high-quality sequencing reads are aligned to the human reference genome (GRCh38). The algorithm is used to accurately calculate the expression level of each gene in each cell, and the raw count data is standardized to eliminate technical bias. During the analysis process, the inventors also focused on evaluating the key performance indicators of the mtCOOL-seq technology, including sequencing depth, gene detection rate, and cell-to-cell correlation.
[0195] In order to further verify the accuracy and reliability of the mtCOOL-seq technology of Example 1, the inventor introduced two bulk RNA sequencing data sets of K562 cell lines as references. These two data sets are from GSM5687482 and GSM5834052, respectively, and are referred to as Reference 1 (abbreviated as Ref1) and Reference 2 (abbreviated as Ref2) in their analysis. Although these reference data sets are based on the RNA sequencing results of the entire cell population, they still provide a benchmark for evaluating the performance of the mtCOOL-seq technology of Example 1 in capturing the transcriptome characteristics of K562 cells. The single-cell data obtained by the mtCOOL-seq technology were compared and analyzed in depth with these two reference data sets. This comparison not only helps to evaluate the ability of the mtCOOL-seq technology to capture gene expression patterns at the single-cell level, but also reveals the advantages and characteristics of the technology relative to traditional overall sequencing methods.
[0196] The analysis results are visualized through two heatmaps: cluster pheatmap of K562 ( Figure 2 ) and “cor pheatmap of K562” ( Figure 3). These two heat maps provide a visual representation of the performance of the mtCOOL-seq technology.
[0197] Figure 2 The "cluster pheatmap of K562" shows the sample clustering results under the gene expression pattern. This heatmap shows the similarities and differences between the single-cell data obtained by the mtCOOL-seq technology and the reference data, and reveals the heterogeneity between different single-cell samples. By observing the clustering pattern, it is possible to identify which single-cell samples are closer to the overall data and the degree of variation between samples. Figure 2 It can be seen that the single-cell data obtained by mtCOOL-seq are highly consistent with the reference data, proving that this technology can accurately capture the cellular gene expression profile.
[0198] Figure 3 The "correlation pheatmap of K562" visually shows the correlation between the single-cell data obtained by the mtCOOL-seq technology and the reference data set. This heatmap quantifies the consistency of the mtCOOL-seq technology and traditional methods in gene expression quantification, further verifying the reliability of the mtCOOL-seq technology. By analyzing the strength of the correlation, the consistency of the single-cell data and the reference data can be evaluated, confirming the accuracy and stability of mtCOOL-seq in gene expression detection. Figure 3 It can be seen that different single-cell samples show highly similar gene expression patterns, further demonstrating the stability and reliability of this technology.
[0199] By comparing with the authoritative K562 cell whole RNA sequencing data, the mtCOOL-seq technology of Example 1 showed significant advantages. The single-cell data obtained by mtCOOL-seq are highly consistent with the reference data, proving that the technology can accurately capture the gene expression profile of the cell. Different single-cell samples showed highly similar gene expression patterns, further demonstrating the stability and reliability of the technology. In addition, mtCOOL-seq can detect as many or even more genes as whole-cell sequencing, indicating that it has obvious advantages in capturing low-abundance transcripts. Compared with whole-cell sequencing, mtCOOL-seq can also reveal differences between individual cells, providing a powerful tool for studying the heterogeneity of cell populations. The optimized technical process reduces sample processing steps, reduces technical deviations, and improves experimental efficiency. The successful application in K562 cells shows that mtCOOL-seq technology is expected to play an important role in various cell types and biological problems. In summary, the results of Example 2 demonstrate the leading position of mtCOOL-seq technology in the field of single-cell RNA sequencing. Its high accuracy, high sensitivity and excellent repeatability provide researchers with a powerful and reliable tool to help deeply explore the complexity and diversity of the cellular transcriptome.
[0200] In summary, the present invention introduces an innovative labeling strategy, and the mtCOOL-seq technology realizes mixed processing of samples, greatly improves the experimental throughput, and makes it possible to analyze the multi-omics information of a large number of single cells at the same time. Secondly, the optimized experimental process and data analysis strategy significantly reduce the cost of single-cell sequencing, making large-scale epigenomic research more economical and feasible. mtCOOL-seq also has the ability to simultaneously analyze DNA methylation, chromatin state, nucleosome positioning, copy number variation and ploidy, and provides comprehensive single-cell epigenomic information. This is the first medium-throughput single-cell multi-omics technology on the market. It is particularly worth mentioning that the technology innovatively uses the MspI enzyme for CNV detection, which improves the accuracy and efficiency of copy number variation analysis. By enriching CCGG sites with enzyme digestion, the technology increases the sequencing depth of the target area, while reducing the proportion of non-specific sequencing, further improving sequencing efficiency and data quality. In addition, mtCOOL-seq cleverly draws on the library construction ideas of the second-generation sequencing, optimizes the technical route, not only significantly reduces costs, but also achieves enrichment of specific genomic regions. In the single-cell nuclear fractionation step, the technology innovatively introduced the use of specific metal cations, which greatly improved the retention rate of DNA and the conversion efficiency of RNA, thereby further improving the efficiency of nuclear fractionation. mtCOOL-seq has high flexibility and is suitable for various types of single-cell samples, including freshly isolated cells, frozen cells, and fixed cells. In the DNA analysis part, the technology preferably adopts a double-stranded DNA synthesis strategy, which significantly improves the genome coverage. Finally, through the optimized data analysis process, the technology improves the coverage and accuracy of DNA methylation detection. The main advantage of mtCOOL-seq technology is that it combines high throughput, low cost, multi-omics and high efficiency, which can provide researchers with a powerful tool to deeply explore cellular heterogeneity, epigenetic regulatory mechanisms and their role in development, disease and treatment response.
[0201] The embodiments of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above embodiments. Various changes can be made within the knowledge of ordinary technicians in the relevant technical field without departing from the purpose of the present invention. In addition, the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.
Claims
1. A method for constructing a medium-throughput single-cell quadruple chemistry library, comprising the following steps: (1) Lysing cells and separating to obtain total RNA solution and cell nucleus solution; (2) constructing a transcriptome library for the total RNA solution in step (1), and at the same time, constructing a genome library for the cell nucleus solution in step (1) to obtain the single cell quadruple chemistry library.
2. The construction method according to claim 1, characterized in that: The transcriptome library construction includes: 1) RNA is reverse transcribed to obtain cDNA, pre-amplified, and biotin PCR amplified; 2) fragmenting the product of biotin PCR amplification to obtain fragmented DNA; 3) Perform end repair on the fragmented DNA and add A to the 3' end of the nucleic acid fragment; 4) adding a double-stranded DNA adapter and a ligation reaction reagent to carry out a ligation reaction to obtain a library containing adapter ligations; 5) Amplifying the library containing the adapter ligation to obtain a transcriptome library.
3. The construction method according to claim 1, characterized in that: Before the single cell lysis, the cells are subjected to in vitro methylation treatment using GpC methyltransferase; Preferably, the in vitro methylation treatment comprises mixing the cells with a solution containing GpC methyltransferase, reacting, and inactivating the cells; Preferably, the solution containing GpC methyltransferase further comprises at least one of RNase inhibitor, surfactant, LambdaDNA and S-adenosylmethionine; Preferably, the reaction system of the in vitro methylation treatment contains 0.5-20 U / μL GpC methyltransferase; Preferably, the reaction conditions are vortexing for 5 to 15 seconds and then incubating at 35 to 40° C. for 5 to 15 minutes.
4. The construction method according to claim 2, characterized in that: The RNA reverse transcription system includes TSO primers, RNA tags, reverse transcriptase, RNase inhibitors, enhancers and buffer solutions; Preferably, the reverse transcriptase comprises M-MLV reverse transcriptase, AMV reverse transcriptase, Hiscript III reverse transcriptase or a combination thereof; Preferably, the buffer comprises at least one of dNTP and metal ions; Preferably, the RNA tag contains an 8-bit barcode sequence; Preferably, the length of the fragmented DNA is 100 to 600 bp.
5. The construction method according to any one of claims 1 to 4, characterized in that: The genomic library construction includes: a) Protease digestion and enzyme cleavage of cell nucleus solution; b) ligation of methylated adapters, purification by Biotin, and then bisulfite conversion to obtain single-stranded DNA; c) performing linear amplification on single-stranded DNA; d) performing library amplification on the amplified product of c) so that the positive strand and the reverse strand of the DNA fragment are added with two different oligonucleotide sequences; Preferably, the enzyme used for the enzyme digestion comprises at least one of AluI, TaqαI, SphI, MspI, BamHI, ApeKI, BanII, HpaII, HpyCH4V, BstNI, HaeIII, HpyCH4III, BglII, BssSI and KpnI; Preferably, the oligonucleotide sequence in d) is P5 or P7, and the nucleotide sequence is shown in SEQ ID NO:10 or SEQ ID NO:
11.
6. The construction method according to claim 5, characterized in that: b) The linker sequence connected by the methylated linker is a Y-shaped double-stranded structure formed by annealing two single-stranded nucleic acids; Preferably, all cytosine nucleotides in the two single-stranded nucleic acids are methylated; the 5' end of the bottom nucleic acid strand of the Y-shaped double-stranded structure formed after the two single-stranded nucleic acids are annealed is phosphorylated; Preferably, the 3' end of the Y-shaped double-stranded structure is a tag sequence having 8 nucleotides; Preferably, the nucleotide sequence of the upper chain of the linker sequence is as shown in SEQ ID NO:7, the nucleotide sequence of the lower chain of the linker sequence is as shown in SEQ ID NO:8, and the Y-shaped double-stranded structure is formed after the nucleic acids of the upper chain and the lower chain are annealed; Preferably, the linear amplification in c) comprises the following steps: single-stranded DNA is mixed with NEB Buffer, dNTP and random primers, reacted at 90-98°C for 30-52 seconds, an enzyme with strand displacement activity is added on ice, the temperature is increased from 4°C to 37°C to promote base pairing, and incubated at 37°C for 65-100 minutes, and the temperature is increased by 1°C every 5-15 seconds; Preferably, the nucleotide sequence of the random primer is as shown in SEQ ID NO:9; Preferably, the enzyme with strand displacement activity includes at least one of klenow polymerase (3'→5'exo-), klenow polymerase, bstDNA polymerase, vent DNA polymerase (3'→5'exo-), vent DNA polymerase, Phi 29DNA polymerase, deep ventDNA polymerase (3'→5'exo-), and deep vent DNA polymerase.
7. The construction method according to claim 6, characterized in that: The cells are single cells and / or multi-cell samples; Preferably, when the cells are a multi-cell sample, in the process of constructing the transcriptome library, the products of the pre-amplification of each single cell are mixed before the biotin PCR amplification; During the construction of the genomic library, before bisulfite conversion, the products of biotin purification of each single cell were mixed.
8. A single cell quartet chemical library constructed by the construction method according to any one of claims 1 to 7.
9. A sequencing method, comprising the step of sequencing the single cell quadruple chemical library of claim 8; Preferably, the sequencing method includes any one of illumina, DNB, ABI, 454FLX, and SOLiD sequencing.
10. Use of the construction method according to any one of claims 1 to 7 or the sequencing method according to claim 9 in any one of (1) to (3): (1) Analysis of copy number variation and / or chromatin accessibility of single-cell genomes from multiple samples; (2) DNA methylation and / or single-site mutation analysis of multiple samples and single cells; (3)Multi-sample single-cell gene expression detection.
Citation Information
Cited By
Construction method of single-cell multi-modal sequencing library
CN122146852A