A high-throughput single-cell level detection method for multiple protein-genome interactions and its application
By performing multiple ligation reactions of antibody tags and adapter sequences at the single-cell level, and combining cell tag information to construct sequencing libraries, the limitations of existing technologies in detecting limited protein types, cumbersome procedures, and high costs are solved, achieving efficient and low-cost detection of multiple protein-genome interactions.
Patent Information
- Application Number
- CN202510750018.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-06-06
AI Technical Summary
Existing high-throughput single-cell multiplex detection technologies suffer from limitations in the types of proteins that can be detected, cumbersome experimental procedures, high costs, and low efficiency.
By performing multiple ligation reactions, antibody tags and adapter sequences are labeled with nucleic acids in the cell nucleus. Combined with cell tag information, multiple protein-genome interaction detection at the single-cell level is achieved, including antibody tag-nucleic acid conjugation, the use of adapter sequences, and ligation reactions in microdroplets. Sequencing libraries are then constructed and high-throughput sequencing is performed.
It enables the simultaneous batch detection of genomic DNA binding sites of multiple histone modifications or transcription factors in tens of thousands of single cells, simplifying the operation process, reducing costs, and making it suitable for tumor heterogeneity and developmental biology research.
Smart Images

Figure CN120249444B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of gene sequencing and tissue cell sample analysis technology, and in particular to a high-throughput single-cell level detection method for multiple protein-genome interactions and its application. Background Technology
[0002] In complex biological systems, significant heterogeneity exists in the epigenetic states and protein-DNA interactions among single cells. For example, the differences in drug responses among different cell subpopulations in the tumor microenvironment may stem from dynamic changes in chromatin-binding proteins. High-throughput single-cell multiplex assays can reveal this heterogeneity, providing crucial information for understanding cell differentiation, disease development, and treatment resistance. Traditional single-target detection methods (such as ChIP-seq and CUT&Tag techniques) require multiple experiments, while multiplex assays can simultaneously capture genomic binding information of multiple proteins or histone modifications, significantly reducing experimental time and cost, and revealing multidimensional regulatory networks.
[0003] Currently, existing single-cell multiplex detection technologies, such as uCoTargetX, MULTI-CUT&Tag, and Nano-CUT&Tag, can detect a maximum of 2-5 proteins simultaneously. They require sequential labeling or antibody reuse and have limitations such as tag interference from transposases, such as Tn5 transposase (abbreviated as Tn5), and cumbersome experimental procedures. Alternatively, they rely on special reagents, such as the need to prepare new pA / pG-Tn5 antibody complexes or Tn5-Nanobody fusion proteins, resulting in high costs and reduced efficiency. Summary of the Invention
[0004] The main objective of this invention is to propose a high-throughput single-cell level detection method for multiple protein-genome interactions and its application, aiming to solve the problems of limited protein detection, cumbersome experimental procedures, high cost, and low efficiency in the process of using existing high-throughput single-cell technology.
[0005] To achieve the above objectives, this invention proposes a high-throughput single-cell level detection method for multiple protein-genome interactions, comprising the following steps:
[0006] S10 provides the cell nucleus;
[0007] S20. Mix the cell nucleus with the antibody tag to label the target protein in the cell nucleus;
[0008] S30. Provide a transposase with a linker sequence, and use the transposase with the linker sequence to fragment the nucleic acid in the cell nucleus to obtain multiple nucleic acid fragments containing the linker sequence.
[0009] S40. Perform a first ligation reaction between the plurality of nucleic acid fragments containing the adapter sequence and the labeled target protein, so as to label the plurality of nucleic acid fragments containing the adapter sequence with information of the target protein;
[0010] S50. Provide microdroplets with cell tags, and perform a second ligation reaction between the microdroplets with cell tags and the plurality of nucleic acid fragments containing the adapter sequence to label the plurality of nucleic acid fragments with the cell tag information;
[0011] S60. Combining the information of the cell tag and the information of the target protein, a library is constructed and sequenced to obtain information on protein-genome interactions in single-cell genomics of the sample to be tested.
[0012] In one embodiment, in step S20, the nucleic acid sequence length in the antibody tag is 20-150 nt.
[0013] In one embodiment, in step S20, the sequence of the antibody tag includes a 5' end modification group, a Read2 sequencing primer, a first random sequence, an antibody identification code index, a second random sequence, and a linker sequence.
[0014] The sequence of the antibody tag includes any one of S1 to S6, wherein the specific sequence of S1 is shown below:
[0015] NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNNNN-index-NNNNNNNNNGCTTTAAGGCGTTAGGTGATTA;
[0016] The specific sequence of S2 is shown below:
[0017] NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNN-index-NNGTTAGGTGAT*T*A;
[0018] The specific sequence of S3 is shown below:
[0019] NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNN-index-GGTGAT*T*A;
[0020] The specific sequence of S4 is shown below:
[0021] NH2-TTCCTTGGCACCCGAGAATTCCANN-index1-GTTAGGTGAT*T*A;
[0022] The specific sequence of S5 is shown below:
[0023] NH2-TTCCTTGGCACCCGAGAATTCCANN-index1-GGTGAT*T*A;
[0024] The specific sequence of S6 is shown below:
[0025] 5Biotin-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNNNN-TGACCAGTTCCGCAT-NNNNNNNNNGCTTTAAGGCCGGTCCTAGC*A*A, where * indicates thiophosphate bond modification.
[0026] In one embodiment, in step S20, the antibody tag is formed by conjugating an antibody and a nucleic acid: wherein the antibody is used to target the target protein, and the nucleic acid is used for labeling; and / or,
[0027] The coupling method includes streptavidin-biotin coupling; or,
[0028] The coupling methods include amino-crosslinker-mediated covalent coupling.
[0029] In one embodiment, in step S30, the adapter sequence includes a first adapter sequence and a second adapter sequence located at both ends of the transposase, wherein the first adapter sequence is PrimerB, and the sequence of PrimerB is shown in SEQ ID NO.6: where,
[0030] The second connector sequence is the sequence of primer C-1 as shown in SEQ ID NO.7; and / or,
[0031] The second connector sequence is the sequence of primerC-2 as shown in SEQ ID NO.8.
[0032] In one embodiment, the connector sequence further includes primer A, the sequence of which is shown in SEQ ID NO. 9.
[0033] In one embodiment, step S10 includes:
[0034] Provide a microfluidic chip system:
[0035] Multiple nucleic acid fragments containing the adapter sequence and the microdroplets with cell tags are introduced into a microfluidic system, where a second ligation reaction is performed inside the microdroplets to label the multiple nucleic acid fragments with the cell tag information.
[0036] The present invention also provides a sequencing library comprising a sequencing library constructed using the high-throughput single-cell level multiple protein-genome interaction detection method described in any one of the preceding claims.
[0037] The present invention also provides an application of a sequencing library constructed by the high-throughput single-cell level multiple protein-genome interaction detection method as described in any of the preceding claims in single-cell multiplex detection technology.
[0038] In this invention, through multiple ligation reactions, the antibody tag and adapter sequence are respectively ligated to the cell nucleus with a cell tag, simultaneously achieving dual labeling of genomic DNA of target proteins or protein modifications by single-cell tags and antibody tags. This method can simultaneously detect genomic DNA binding sites of multiple histone modifications or transcription factors in thousands of single cells. Compared with traditional single-cell multiplex CUT&Tag technology, which can only target 2 to 5 proteins or modifications at a time, this invention does not require custom protein design, so the number of targeted proteins is unlimited. Moreover, the operation method is simple, the cost is low, and the process is short, making it particularly suitable for tumor heterogeneity and developmental biology research. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0040] Figure 1 This is a flowchart of a high-throughput single-cell level multiple protein-genome interaction detection method according to an embodiment of the present invention;
[0041] Figure 2 This is a structural diagram of an antibody nucleic acid tag sequence according to an embodiment of the present invention;
[0042] Figure 3 This is a schematic diagram of the adapter sequence of the Tn5 transposase in one embodiment of the present invention.
[0043] Figure 4 This is a diagram illustrating the sequencing read structure of different adapter sequences of the Tn5 transposase in one embodiment of the present invention;
[0044] Figure 5 This is a comparison chart showing the proportion of high-quality DNA fragments that can overlap with peaks for antibody tags of different lengths in one embodiment of the present invention.
[0045] Figure 6Peaks of antibody tags of different lengths and traditional CUT&Tags in one embodiment of the present invention Figure 1 Comparison chart of consistency indicators.
[0046] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Where specific conditions are not specified in the embodiments, conventional conditions or conditions recommended by the manufacturer shall apply. Where the manufacturers of reagents or instruments are not specified, they are all conventional products that can be purchased commercially. Furthermore, the meaning of "and / or" throughout the text includes three parallel solutions; for example, "A and / or B" includes solution A, or solution B, or a solution where both A and B are satisfied simultaneously. In addition, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] Currently, existing single-cell multiplex detection technologies, such as uCoTargetX, MULTI-CUT&Tag, and Nano-CUT&Tag, can detect a maximum of 2-5 proteins simultaneously. They require sequential labeling or antibody reuse and have limitations such as tag interference from transposases, such as Tn5 transposase (abbreviated as Tn5), and cumbersome experimental procedures. Alternatively, they rely on special reagents, such as the need to prepare new pA / pG-Tn5 antibody complexes or Tn5-Nanobody fusion proteins, resulting in problems such as high cost and reduced efficiency.
[0049] In view of this, the present invention provides a high-throughput single-cell level detection method for multiple protein-genome interactions, comprising the following steps:
[0050] S10 provides the cell nucleus;
[0051] S20. Mix the cell nucleus with the antibody tag to label the target protein in the cell nucleus;
[0052] S30. Provide a transposase with a linker sequence, and use the transposase with the linker sequence to fragment the nucleic acid in the cell nucleus to obtain multiple nucleic acid fragments containing the linker sequence.
[0053] S40. Perform a first ligation reaction between the plurality of nucleic acid fragments containing the adapter sequence and the labeled target protein, so as to label the plurality of nucleic acid fragments containing the adapter sequence with information of the target protein;
[0054] S50. Provide microdroplets with cell tags, and perform a second ligation reaction between the microdroplets with cell tags and the plurality of nucleic acid fragments containing the adapter sequence to label the plurality of nucleic acid fragments with the cell tag information;
[0055] S60. Combining the information of the cell tag and the information of the target protein, a library is constructed and sequenced to obtain information on protein-genome interactions in single-cell genomics of the sample to be tested.
[0056] In this invention, through multiple ligation reactions, the antibody tag and adapter sequence are respectively ligated to the cell nucleus with a cell tag, simultaneously achieving dual labeling of genomic DNA of target proteins or protein modifications by single-cell tags and antibody tags. This method can simultaneously detect genomic DNA binding sites of multiple histone modifications or transcription factors in thousands of single cells. Compared with traditional single-cell multiplex CUT&Tag technology, which can only target 2 to 5 proteins or modifications at a time, this invention does not require custom protein design, so the number of targeted proteins is unlimited. Moreover, the operation method is simple, the cost is low, and the process is short, making it particularly suitable for tumor heterogeneity and developmental biology research.
[0057] It should be noted that the two ligation reactions in steps S40 and S50 occur in microdroplets and are not related to the order of time. They can be completed simultaneously or in steps. After the two ligation reactions are completed, the resulting nucleic acid fragments carry both the target protein tag and the cell tag. Based on these dual-labeled nucleic acid fragments, libraries can be constructed and sequenced.
[0058] Specifically, the cell nucleus needs to be permeabilized beforehand to allow the antibody to enter the nucleus and bind to the target site on the chromatin. The permeabilized cell nucleus is then incubated with an antibody carrying a nucleic acid tag to obtain an antibody-tagged cell nucleus, wherein the cell allows the antibody to target a specific protein or other modifications. Next, the antibody-treated cell nucleus is used to indiscriminately break down DNA using a Tn5 transposase complex. Then, an oil-in-water microfluidic device is used to ligate the antibody nucleic acid tag to the 5' end of one of the adapters of the adjacent Tn5-broken DNA fragment in a droplet environment, and to ligate the cell tag carried by the hydrogel microbead to the 5' end of another Tn5 adapter of the DNA fragment. After that, a sequencing library is constructed and high-throughput sequencing is performed. Finally, bioinformatics analysis is used to determine the protein-genome interaction sites.
[0059] It should be noted that only nucleic acid fragments located near the target protein or that can be indirectly linked to the target protein through some mechanism (such as cross-linking in chromatin immunoprecipitation ChIP) can obtain information about the target protein.
[0060] Specifically, such as Figure 1 As shown, it includes the following steps:
[0061] Step 1: Extract the cell nucleus from the cell to ensure that the experiment is performed only on the genomic DNA within the cell nucleus and to avoid interference from other cellular components. Then, incubate the cell nucleus with a specific antibody (oligo-antibody), i.e., an antibody tag, so that the antibody tag binds to the target protein, thereby marking a specific genomic region for the identification and localization of protein-DNA interaction sites of interest.
[0062] Step 2: Treat the cell nuclei after antibody incubation with Tn5 transposase to break down the genomic DNA and simultaneously insert adapter sequences (Primer A, Primer B, and Primer C) at both ends of the DNA fragments. These adapters contain universal sequences and possible barcodes or UMIs required for subsequent PCR amplification and sequencing.
[0063] Step 3: Tn5-treated cell nuclei are encapsulated together with gel beads containing cell barcodes (i.e., cell tags) in microdroplets. Each droplet contains a cell nucleus and a gel bead. The cell barcode on the gel bead will be linked to the DNA fragment in a subsequent step to generate a unique identifier for each cell, facilitating the differentiation of signals from different cells during subsequent data analysis, thereby enabling single-cell level operations.
[0064] Step 4: Within the microdroplet, the gel bead releases the cell barcode it carries and attaches it to the DNA fragment, adding a cell-specific barcode to each DNA fragment to ensure that all DNA fragments from the same cell can be correctly classified in subsequent analyses.
[0065] Step 5: Amplify the barcoded DNA fragments using PCR to increase the number of DNA fragments needed for high-throughput sequencing. Simultaneously, purify the PCR products to remove excess primers and other impurities, improving the quality of the sequencing library, reducing non-specific background, and ensuring the accuracy and reliability of the sequencing results. During the amplification process, additional sequences (such as i7 indexes) can be introduced for further sample differentiation and sequencing platform compatibility.
[0066] Step 6: In summary, a library suitable for high-throughput sequencing can be obtained. This library can then be directly used for sequencing. The sequencing data can be used to analyze the interaction sites between the target protein and genomic DNA, as well as the distribution of these sites at the single-cell level.
[0067] In one embodiment, in step S20, the sequence length of the nucleic acid in the antibody tag is 20-150 nt, for example, it can be 20 nt, 30 nt, 50 nt or 150 nt. Within this range, the data quality of the test library can be improved. If the detection rate of the obtained sequencing library is less than 20 nt or greater than 150 nt, the target nucleic acid may not be sequenced.
[0068] In one embodiment, in step S20, the sequence of the antibody tag includes: a 5' end modification group, a Read2 sequencing primer, a first random sequence, an antibody recognition code index, a second random sequence, and a linker sequence, specifically as follows: Figure 2 As shown, the antibody tag sequence includes a 5' end modification group, a Read2 sequencing primer, a random sequence, an antibody recognition code index, another random sequence, and a linker sequence. The 5' end modification group is used to stably couple the nucleic acid tag and the antibody to achieve antibody encoding. The antibody recognition code corresponds one-to-one with the antibody type, which can accurately identify and detect proteins or epigenetic modifications that interact with DNA during sequencing analysis. Since the base-encoded sequence has many possible combinations, for example, an 8nt nucleic acid sequence can theoretically recognize 4 to the power of 8, a total of 65,538 antibodies. Therefore, a single experiment can simultaneously detect multiple unrestricted protein or epigenetic modification interactions with the genome, as long as there is a corresponding specific antibody. Optionally, the addition of two random base sequences can improve the sequencing quality of the antibody recognition code. The linker sequence is used to connect to the 5' end of the primer C inserter of the transposase-broken genomic DNA in the microdroplet. This design not only improves the binding stability of the antibody and nucleic acid tag but also facilitates subsequent bioinformatics analysis.
[0069] In one embodiment, the sequence of the antibody tag includes any one of S1 to S6, wherein the sequences of S1 to S6 are respectively shown as SEQ ID NO.1, SEQ ID NO.2, SEQ ID NO.3, SEQ ID NO.4, SEQ ID NO.5 and SEQ ID NO.6.
[0070] The specific sequence of S1 is shown below:
[0071] NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNNNN-index-NNNNNNNNNGCTTTAAGGCGTTAGGTGATTA;
[0072] The specific sequence of S2 is shown below:
[0073] NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNN-index-NNGTTAGGTGAT*T*A;
[0074] The specific sequence of S3 is shown below:
[0075] NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNN-index-GGTGAT*T*A;
[0076] The specific sequence of S4 is shown below:
[0077] NH2-TTCCTTGGCACCCGAGAATTCCANN-index1-GTTAGGTGAT*T*A;
[0078] The specific sequence of S5 is shown below:
[0079] NH2-TTCCTTGGCACCCGAGAATTCCANN-index1-GGTGAT*T*A;
[0080] The specific sequence of S6 is shown below:
[0081] 5Biotin-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNNNN-TGACCAGTTCCGCAT-NNNNNNNNNGCTTTTAAGGCCGGTCCTAGC*A*A;
[0082] It should be noted that * indicates thiophosphate bond modification.
[0083] In one embodiment, in step S20, the antibody tag is formed by conjugating an antibody and a nucleic acid: wherein the antibody is used to target the target protein, and the nucleic acid is used to label the cell nucleus to achieve cell nucleus labeling. These nucleic acid tags can be further conjugated with fluorescent labels or other reporter molecules for visualization.
[0084] In one embodiment, the coupling method includes streptavidin-biotin coupling or amino-crosslinker-mediated covalent coupling. In the implementation of the present invention, streptavidin-coupled antibody is used for direct binding to biotin-modified nucleic acid, or amino-modified antibody is used for covalent binding mediated by amino-reactive crosslinker. Both methods have high coupling efficiency and stability and can meet different experimental requirements.
[0085] Specifically, the target protein is conjugated with an antibody, streptavidin, and then a biotin-modified DNA oligonucleotide is synthesized. Utilizing the extremely strong specific binding ability of the antibody streptavidin, the DNA oligonucleotide is linked to the target antibody for recognition of the target protein labeled by the DNA oligonucleotide antibody during sequencing.
[0086] In one embodiment, in step S30, the adapter sequence includes a first adapter sequence and a second adapter sequence located at both ends of the transposase, wherein the first adapter sequence is PrimerB, and the sequence of PrimerB is shown in SEQ ID NO.6: where,
[0087] The second connector sequence is the sequence of primer C-1 as shown in SEQ ID NO.7; and / or,
[0088] The second connector sequence is the sequence of primerC-2 as shown in SEQ ID NO.8.
[0089] Unlike the adapter sequences in the default sequencing primers of traditional Nextera library preparation kits and sequencers, the adapter sequences of this invention enable sequencing to begin from the designated nucleic acid tag Read2 sequencing primer. This design ensures that the sequencing read length can accurately cover the antibody tag region while detecting the genomic DNA sequence, thereby improving the accuracy and sensitivity of the detection.
[0090] It should be noted that using the original PrimerC sequence will cause Reads to fail to read the antibody nucleic acid tag sequence. In other words, there is a compatibility issue between the original PrimerC sequence and subsequent PCR primers or sequencing primers, resulting in an unsatisfactory binding position of the Read2 sequencing primer. The sequencing process skips the target region and directly starts sequencing from the Tn5-cut genomic DNA, thus ignoring the antibody tag attached to its 5' end and PrimerC itself. By introducing PrimerC-1 and PrimerC-2 sequences to replace the original PrimerC sequence, the Read2 sequencing primer can bind to the expected site more effectively instead of directly binding to the primerC for sequencing extension, thereby correctly reading the antibody tag sequence. The replacement principle is that the replaced primers are not complementary to the primers in the traditional Nextera library preparation kit and the sequencer's default sequencing primers, avoiding competition between the two sequencing primers. Therefore, the second adapter sequence of this invention can be either PrimerC-1 or PrimerC-2.
[0091] Specifically, the steps for replacing the original PrimerC sequence with primerC-1 or primerC-2 are as follows:
[0092] To prevent competition between the Rd2 Seq Primer and the Nextera N7 primer during sequencing, the T7 sequence needs to be modified. To ensure the modified sequence does not interfere with the Read2 sequencing primers and to reduce non-specific elongation during library construction, the following modification principles apply:
[0093] 1. The first two bases are selected from AC, GC, or GT;
[0094] 2. The sequence of the three-base sequence formed by GT and the first base of the generated sequence in sequential order cannot appear in AGATCGGAAGAGCGTCGTGTAG;
[0095] ACGAGCAACGACGGACGACAGCAA;
[0096] TTGCTGTCGTCCGTCGTTGCT;
[0097] CGAAATTCCGGCCAGGATCGTT;
[0098] TTTAAGGCCGGTCCTAGCAA;
[0099] In the 7 sequences ATGCGGAACTGGTCA;AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC, note that the order refers to the GT base being placed before the first base of the generated sequence.
[0100] The tribase sequence formed by T and the first two bases of the generated sequence in sequential order cannot appear in AGATCGGAAGAGCGTCGTGTAG;ACGAGCAACGACGGACGACAGCAA;
[0101] TTGCTGTCGTCCGTCGTTGCT;
[0102] CGAAATTCCGGCCAGGATCGTT;
[0103] TTTAAGGCCGGTCCTAGCAA;
[0104] In the 7 sequences ATGCGGAACTGGTCA;AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC, note that the order refers to the T base being before the first two bases of the generated sequence.
[0105] 4. If no sequence meets conditions 2 and 3, the conditions can be relaxed to the following: the 4-base sequence formed by GT and the first two bases of the generated sequence in sequential order cannot contain the following sequences:
[0106] AGATCGGAAGAGCGTCGTGTAG;
[0107] ACGAGCAACGACGGACGACAGCAA;
[0108] TTGCTGTCGTCCGTCGTTGCT;
[0109] CGAAATTCCGGCCAGGATCGTT;
[0110] TTTAAGGCCGGTCCTAGCAA;
[0111] In the 7 sequences ATGCGGAACTGGTCA;AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC, the first two bases in conditions 1-5 are AC;
[0112] 5. The generated sequence does not contain GTT, CGT, GTA, or TAC sequences;
[0113] 6. The generated sequence cannot contain three consecutive identical bases;
[0114] Finally, the specific sequences obtained for Primer C-1 are: ACGATGCATT; and for Primer C-2, ACGGCATGAT. Other sequences obtained based on the above principles should also be included in the protection scope, such as ACGCTACGAT, ACGATCGTCA, or ACGATCGTCA.
[0115] In one embodiment, the adapter sequence further includes primer A, the sequence of which is shown in SEQ ID NO.9. Primer A is mainly used to bind to the Tn5 protein to form a stable transposase complex. It does not usually participate directly in insertion into genomic DNA, but rather serves as structural support.
[0116] In one embodiment, step S10 includes: providing a microfluidic chip system, introducing multiple nucleic acid fragments containing the adapter sequence and the microdroplets with cell tags into the microfluidic system, and performing a second ligation reaction inside the microdroplets to label the multiple nucleic acid fragments with the cell tag information.
[0117] It should be noted that in the experiment, the microfluidic device can generate hundreds of thousands of microdroplets. Each droplet contains an informatically distinguishable single cell and a unique cell tag for each droplet. Genomic DNA is labeled in the microdroplets using antibody tags and cell tags. DNA from the same cell is labeled with the same cell tag, and DNA from different cells is labeled with different cell tags.
[0118] This invention provides a sequencing library comprising a sequencing library constructed using the high-throughput single-cell level multiple protein-genome interaction detection method described above. This sequencing library possesses all the technical solutions of the high-throughput single-cell level multiple protein-genome interaction detection method and thus has all the beneficial effects of the high-throughput single-cell level multiple protein-genome interaction detection method, which will not be elaborated upon here.
[0119] This invention also provides an application of the sequencing library constructed by the high-throughput single-cell level multiple protein-genome interaction detection method described above in single-cell multiplex detection technology.
[0120] The raw materials used in the following embodiments are sourced from:
[0121] K562 cell line: A passaged cell line owned by the laboratory.
[0122] H3K27Ac Specific Antibody: Recombinant Anti-Histone H3 (acetyl K27) Antibody [EP16602] - ChIP Grade - BSA and Azide free, Abcam (Abcam (Shanghai) Trading Co., Ltd.), catalog number ab302877;
[0123] Streptavidin Conjugation Kit - Lighting-link®: Abcam (Abcam (Shanghai) Trading Co., Ltd.), catalog number ab102921;
[0124] Commercial SeekOne® DD Single-Cell ATAC+RNA Dual-Omics Kit: Beijing SeekOne Biotechnology Co., Ltd., catalog numbers K02901-02 and K02901-08.
[0125] Example: Application of high-throughput single-cell genomic DNA and RNA co-detection
[0126] Example 1: The impact of different Tn5 coating sequences on sequencing results
[0127] 1. Preparation of streptavidin-labeled primary antibody: Histone-modified H3K27Ac specific antibody was conjugated with streptavidin SA using the Abcamlighting-link kit;
[0128] 2. Synthesis of the nucleic acid tag in the antibody tag: A 5' biotin-modified nucleic acid antibody tag was synthesized, with the specific sequence as follows:
[0129] 5Biotin-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNNNN-TGACCAGTTCCGCAT-NNNNNNNNNGCTTTTAAGGCCGGTCCTAGC*A*A;
[0130] 3. Preparation of antibodies with nucleic acid tags, i.e. antibody tags: The streptavidin antibody product conjugated in step 1 is combined with an excess of the biotin-modified nucleic acid product in step 2 to obtain antibody tags with nucleic acid tags.
[0131] 4. Take fresh K562 cells, add NIB nucleus extraction buffer and treat on ice for 5 min, centrifuge at 500g and resuspend in PBS buffer, adjust the cell density to 1×10^6 / mL to obtain K562 cell nuclei;
[0132] 5. Incubate permeabilized K562 cell nuclei overnight with an antibody containing a nucleic acid tag at a 1:100 volume ratio. Add 1 mL of PBS buffer, centrifuge at 500g for 5 min, and resuspend in PBS buffer. Repeat the washing process three times. 6. The simplified steps of the commercial SeekOne® DD single-cell ATAC+RNA dual-omics kit are as follows:
[0133] 6.1 The cell nuclei after antibody incubation were treated with Tn5 transposase to perform indiscriminate genomic DNA fragmentation. The original Tn5 adapter sequence PrimerC-0 was modified to PrimerC-1 and PrimerC-2 for the experimental group. PrimerB is also known as PrimerB-1. The specific adapter sequences are as follows: The PrimerA sequence is shown below:
[0134] The sequence of 5'-phos-CTGTCTCTTATACACATCT-NH2-3'PrimerB-1 is shown below:
[0135] 5′p-CGTCCGTCGTTGCTCGT-AGATGTGTATAAGAGACAG-3′; The sequence of Primer C-0 is shown below:
[0136] 5′P-GTCTCGTGGGCTCGG-AGATGTGTATAAGAGACAG-3′; The sequence of Primer C-1 is shown below:
[0137] 5'p-ACGATGCATT-AGATGTGTATAAGAGACAG;
[0138] The sequence of Primer C-2 is shown below:
[0139] 5'p-ACGGCATGAT-AGATGTGTATAAGAGACAG, the above sequence is arranged according to Figure 3 The adapter flow diagram completes the adapter process, which involves mixing primer A and primer B-1 in equal molar amounts to form PrimerAB; mixing primer A and primer C-0 / C-1 / C-2 in equal molar amounts to form PrimerAC-0 / AC-1 / AAC-2; then mixing primer AB and primer AC in equal volumes and coating them with Tn5 transposase to complete the naked enzyme coating of the transposase.
[0140] 6.2 Following the kit instructions, use a water-in-oil microfluidic device to ligate antibody tags in the cell nucleus to the Primer C5' end of nearby Tn5-broken DNA within the droplet. Simultaneously, ligate cell tags carried by hydrogel beads in the droplet to the Primer B5' end of the Tn5-broken DNA, thereby achieving cell labeling of each cell nucleus. Only ligation-related components are added to the reaction mixture; transcriptome-related components such as dNTPs are replaced with deionized water.
[0141] 6.3 Based on Illumina TruSeq Read1SeqPrimer and
[0142] microRNA Read2 Sequencing primer was used to directly amplify and obtain a DNA library, which was then sequenced.
[0143] 7. Analyze the sequencing results, as follows: Figure 4 As shown: The default Tn5 library buildup adapter PrimerC-0 starts its Read2 sequencing at the Tn5 break sequence, making it impossible to detect the antibody nucleic acid tag sequence and the PrimerC-0 sequence; while the modified PrimerC-1 and PrimerC-2 start their Read2 sequencing at the preset antibody tag, thus having the correct sequencing structure.
[0144] It should be noted that, Figure 4 The difference lies in the read2 sequence. When primer C-0 is used to coat the transposase, the final library structure lacks an antibody tag sequence. It is speculated that during sequencing, Rd2 Seq Primer and Nextera N7 primer compete, resulting in incorrect sequencing results for this library. When primer C-0 is replaced with primer C-1 or primer C-2, a normal library structure with an antibody tag appears, indicating that the library construction and sequencing results are normal (the antibody tag is the dark green and purple-red sequence in read2 of the figure, the bright blue and yellow are the transposase adapter sequences obtained from normal sequencing, and the green is the fixed sequence in the cell tag used to connect with the transposase adapter).
[0145] Example 2: The Influence of Antibody Tags of Different Lengths on Label Detection 1. Synthesis of Antibody Tags of Different Lengths: Antibody tags of different lengths were synthesized by Sangon Biotech (Shanghai) Co., Ltd., resulting in 5' amino-modified antibody tags, named AbB-20nt, AbB-50nt, AbB-58nt, AbB-81nt, and AbB-150nt, respectively. The specific sequence of AbB-20nt is as follows:
[0146] NH2-ACCCGAGAATTCCA-(Index)1bp-GAT*T*A
[0147] The specific sequence of AbB-81nt is as follows:
[0148] The specific sequence of NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNNNN-(Index)6bp-NNNNNNNNNGCTTTAAGGCGTTAGGTGAT*T*A;AbB-58nt is as follows:
[0149] NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNN-
[0150] The specific sequence of (Index)6bp-NNGTTAGGTGAT*T*A;AbB-50nt is as follows:
[0151] NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNN-(Index)
[0152] 6bp-GGTGAT*T*A;
[0153] The specific sequence of AbB-150nt is as follows:
[0154] NH2-CAAGCAGAAGACGGCATACGAGATGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNNNNNNNNNNNNNN-(Index)
[0155] 30bp-NNNNNNNNNNNNNNNNNNNNGCTTTAAGGCGTTAGGTGAT*T*A;
[0156] It should be noted that N here represents any base (A, T, C, G). The "NNNNNNNNN..." segment is an unknown or variable sequence preceding the "Index" region. Depending on the specific experimental design, it may be used to insert a specific recognition sequence or for other purposes. Antibody tag design must have read2 sequencing primers, which are approximately 20 bp in length. Sequencing is impossible without the sequencing primer sequence. Furthermore, the current mainstream read length for next-generation sequencing is 150 bp. If the tag is greater than or equal to 150 bp, only the tag sequence can be sequenced, not the genomic DNA sequence labeled with the tag.
[0157] 2. Using the Abcam Oligonucleotide Conjugation Kit (ab218260), the antibody tags synthesized in step 1 were conjugated to histone-modified H3K27Ac specific antibodies, respectively.
[0158] 3. Take fresh K562 cells, add NIB nucleus extraction buffer and treat on ice for 5 min, centrifuge at 500g and resuspend in PBS buffer, adjust the cell density to 1×10^6 / mL to obtain K562 cell nuclei;
[0159] 4. Incubate the permeabilized K562 cell nuclei overnight with the prepared nucleic acid tag antibody at a volume ratio of 1:100. Add 1 mL of PBS buffer, centrifuge at 500 g for 5 min, and resuspend in PBS buffer. Repeat the washing process 3 times.
[0160] 5. The simplified operating steps for the commercial SeekOne® DD single-cell ATAC+RNA dual-omics kit are as follows:
[0161] 5.1 The cell nuclei after antibody incubation were treated with Tn5 transposase to perform indiscriminate genomic DNA fragmentation; the adapter sequences of the Tn5 transposase include Primer A, Primer B-1, and Primer C-1, wherein the specific sequence of Primer A is as follows:
[0162] 5'-p-CTGTCTCTTATACACATCT-NH2-3'; The specific sequence of Primer B-1 is as follows:
[0163] 5′P-CGTCCGTCGTTGCTCGT-AGATGTGTATAAGAGACAG-3′;
[0164] The specific sequence of PrimerC-1 is as follows:
[0165] 5′P-ACGATGCATT-AGATGTGTATAAGAGACAG-3.
[0166] 5.2 Following the kit instructions, an oil-in-water microfluidic device was used to ligate antibody tags in the cell nucleus to the Primer C5' end of nearby Tn5-broken DNA within the droplet. Simultaneously, cell tags carried by hydrogel beads in the droplet were ligated to the Primer B5' end of the Tn5-broken DNA, thus achieving cell labeling of each cell nucleus. Only ligation-related components were added to the reaction mixture; transcriptome-related components such as dNTPs were replaced with deionized water.
[0167] 5.3 DNA libraries were directly amplified and sequenced using Illumina Truseq Read1Seq Primer and TruSeq Read2 Sequencing primer;
[0168] 6. Analyze the sequencing results, such as Figure 5 As shown, Figure 5The graph shows the proportion of high-quality DNA fragments overlapping peaks. The results indicate that shorter antibody-conjugated nucleic acid tags have a higher ratio of overlapping peaks. This ratio reflects the proportion of sequencing reads falling within known H3K27Ac enriched regions (peaks). A higher proportion generally means a higher signal-to-noise ratio and better targeting. Shorter antibody-conjugated nucleic acid tags (50 nt and 58 nt) have higher overlapping peaks than longer tags (81 nt). This is because shorter tags have advantages in ligation efficiency, PCR amplification efficiency, or sequencing efficiency, or less impact on Tn5 enzyme proximity. Figure 6 As shown, the peak signal of H3K27Ac obtained using the tag antibody method is similar to that of the conventional CUT&Tag method. CUT&Tag is a technique that directly binds antibodies and digests enzymes in the cell nucleus and is considered one of the gold standards for studying chromatin state. The similar peak signal indicates that the tag antibody-based method can accurately label and enrich the target chromatin region.
[0169] The above are merely preferred embodiments of the present invention and do not limit the patent scope of the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the patent protection scope of the present invention.
Claims
1. A high-throughput single-cell level method for detecting multiple protein-genome interactions, characterized in that, Includes the following steps: S10 provides the cell nucleus; S20. Mix the cell nucleus with the antibody tag to label the target protein in the cell nucleus; S30. Provide a Tn5 transposase with a linker sequence, and use the Tn5 transposase with the linker sequence to fragment the nucleic acid in the cell nucleus to obtain multiple nucleic acid fragments containing the linker sequence. S40. Perform a first ligation reaction between the plurality of nucleic acid fragments containing the adapter sequence and the labeled target protein, so as to label the plurality of nucleic acid fragments containing the adapter sequence with the information of the target protein; S50. Provide microdroplets with cell tags, and perform a second ligation reaction between the microdroplets with cell tags and the plurality of nucleic acid fragments containing the adapter sequence to label the plurality of nucleic acid fragments with the cell tag information; S60. Combining the information of the cell tag and the information of the target protein, a library is constructed and sequenced to obtain information on protein-genome interactions in single-cell genomics of the sample to be tested; in step S20, the antibody tag is composed of an antibody and a nucleic acid conjugated. The antibody is used to target the target protein, and the nucleic acid is used to label the target protein. The nucleic acid in the antibody tag has a sequence length of 20-150 nt, and from 5' to 3', it consists of a 5' end modification group, a Read2 sequencing primer, a first random sequence, an antibody recognition code index, a second random sequence, and a linker sequence connected in sequence.
2. The method for detecting multiple protein-genome interactions at the high-throughput single-cell level as described in claim 1, characterized in that, The nucleic acid sequence in the antibody tag is any one of S1 to S6, where the specific sequence of S1 is shown below: NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNNNN-index-NNNNNNNNNGCTTTAAGGCGTTAGGTGATTA; The specific sequence of S2 is shown below: NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNN-index-NNGTTAGGTGAT*T*A; The specific sequence of S3 is shown below: NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNN-index-GGTGAT*T*A; The specific sequence of S4 is shown below: NH2-TTCCTTGGCACCCGAGAATTCCANN-index1-GTTAGGTGAT*T*A; The specific sequence of S5 is shown below: NH2-TTCCTTGGCACCCGAGAATTCCANN-index1-GGTGAT*T*A; The specific sequence of S6 is shown below: 5Biotin-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNNNN-TGACCAGTTCCGCAT-NNNNNNNNNGCTTTAAGGCCGGTCCTAGC*A*A, where * indicates thiophosphate bond modification.
3. The method for detecting multiple protein-genome interactions at the high-throughput single-cell level as described in claim 1, characterized in that, The coupling method includes streptavidin-biotin coupling; or, The coupling methods include amino-crosslinker-mediated covalent coupling.
4. The method for detecting multiple protein-genome interactions at the high-throughput single-cell level as described in claim 1, characterized in that, In step S30, the adapter sequence includes a first adapter sequence and a second adapter sequence located at both ends of the Tn5 transposase. The first adapter sequence is PrimerB, and the sequence of PrimerB is shown in SEQ ID NO.6: where, The second connector sequence is primer C-1, the sequence of which is shown in SEQ ID NO.7; and / or, The second connector sequence is primerC-2, and the sequence of primerC-2 is shown in SEQ ID NO.
8.
5. The method for detecting multiple protein-genome interactions at the high-throughput single-cell level as described in claim 4, characterized in that, The connector sequence also includes primer A, the sequence of which is shown in SEQ ID NO.
9.
6. The method for detecting multiple protein-genome interactions at the high-throughput single-cell level as described in claim 1, characterized in that, Step S50 includes: A microfluidic chip system is provided; Multiple nucleic acid fragments containing the adapter sequence and the microdroplets with cell tags are introduced into a microfluidic system, where a second ligation reaction is performed inside the microdroplets to label the multiple nucleic acid fragments with the cell tag information.
7. A sequencing library, characterized in that, The sequencing library includes those constructed using the high-throughput single-cell level multiple protein-genome interaction detection method as described in any one of claims 1 to 6.
8. The application of the sequencing library as described in claim 7 in single-cell multiplex detection technology.
Citation Information
Patent Citations
Improvements in rotors for wind powered electric generators
EP0016602A1
Ultrahigh-flux single-cell chromatin transposase accessibility sequencing method
CN113604545A
Single cell DNA-protein interaction sequencing kit and method based on combinatorial index
CN118272509A