Detection method for high-throughput single cell level multiple protein-genome interaction and application thereof
By performing multiple ligation reactions at the single-cell level, combining antibody tags and linker sequences, high-throughput, low-cost multiprotein-genomic interaction detection is achieved, solving the problems of few proteins, high cost and low efficiency in the prior art detection, and is especially suitable for tumor heterogeneity and developmental biology research.
Patent Information
- Application Number
- CN202510750018.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The existing single-cell multiple detection technology can only detect up to 2-5 proteins at the same time, and requires sequential labeling or antibody multiplexing, which has problems such as transposase label interference and cumbersome experimental procedures, high cost and low efficiency.
Through multiple ligation reactions, the antibody tag and linker sequence are linked to the nucleic acid fragments in the cell nucleus, combined with cell tag information, library construction and sequencing is carried out, and the detection of genomic DNA binding sites of multiple histone modification or transcription factors in thousands of single cells is achieved, avoiding the need for customized proteins alone.
High-throughput, low-cost, and simplified single-cell multiprotein-genomic interaction detection is achieved, suitable for tumor heterogeneity and developmental biology research, and the number of targeted proteins is not limited.
Smart Images

Figure CN120249444A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of gene sequencing and tissue cell sample analysis, and particularly relates to a method for detecting high-throughput single-cell level multiplex protein-genome interactions and its application. Background Art
[0002] In complex biological systems, there are significant heterogeneities in the epigenetic states and protein-DNA interactions between single cells. For example, the differential responses of different cell subsets in the tumor microenvironment to drugs may stem from the dynamic changes of chromatin-binding proteins. High-throughput single-cell multiplex detection technologies can reveal such heterogeneities and provide key information for understanding cell differentiation, disease occurrence, and treatment resistance. Traditional single-target detection methods (such as ChIP-seq and CUT&Tag technologies) require separate experiments, while multiplex detection technologies can simultaneously capture the genomic binding information of multiple proteins or histone modifications, significantly reducing the experimental cycle and cost and revealing multi-dimensional regulatory networks.
[0003] Currently, existing single-cell multiplex detection technologies, such as uCoTargetX, MULTI-CUT&Tag, and Nano-CUT&Tag technologies, can detect at most 2-5 proteins simultaneously, require sequential labeling or antibody multiplexing, and have limitations such as tag interference of transposases, such as Tn5 transposase (abbreviated as Tn5), and cumbersome experimental procedures; or rely on special reagents, such as the need to re-prepare new pA / pG-Tn5 and antibody complexes or Tn5-Nanobody fusion proteins, resulting in problems such as high costs and low efficiency. Summary of the Invention
[0004] The main object of the present invention is to propose a method for detecting high-throughput single-cell level multiplex protein-genome interactions and its application, aiming to solve the problems of few detected proteins, cumbersome experimental procedures, high costs, and low efficiency in the process of using existing high-throughput single-cell technologies.
[0005] To achieve the above object, the present invention proposes a method for detecting high-throughput single-cell level multiplex protein-genome interactions, including the following steps: S10. Provide cell nuclei; S20. Mix the cell nuclei with antibody tags to label target proteins in the cell nuclei; S30. Provide a transposase with an adaptor sequence, and use the transposase with the adaptor sequence to fragment nucleic acids in the cell nuclei to obtain multiple nucleic acid fragments containing the adaptor sequence; S40. Perform a first ligation reaction on the multiple nucleic acid fragments containing the linker sequence and the labeled target protein, so as to label the information of the target protein on the multiple nucleic acid fragments containing the linker sequence; S50. Provide microdroplets with cell tags, and perform a second ligation reaction on the microdroplets with cell tags and the multiple nucleic acid fragments containing the linker sequence, so as to label the information of the cell tags on the multiple nucleic acid fragments; S60. Combine the information of the cell tags and the information of the target protein, perform library construction and sequencing, and obtain the information of the protein-genome interaction of single-cell genomics in the sample to be tested.
[0006] In one embodiment, in step S20, the sequence length of the nucleic acid in the antibody tag is 20 - 150 nt.
[0007] In one embodiment, in step S20, the sequence of the antibody tag includes a 5'-end modification group, a Read2 sequencing primer, a first random sequence, an antibody recognition code index, a second random sequence, and a linker sequence.
[0008] The sequence of the antibody tag includes S1 - S5, wherein the sequences of S1 - S5 are shown as SEQ ID NO.1, SEQ ID NO.2, SEQ ID NO.3, SEQ ID NO.4, and SEQ ID NO.5 respectively.
[0009] In one embodiment, in step S20, the antibody tag is formed by conjugating an antibody and a nucleic acid: wherein, the antibody is used to target the target protein, and the nucleic acid is used for labeling; and / or, The conjugation method includes a streptavidin-biotin conjugation method; or, The conjugation method includes an amino-crosslinker-mediated covalent conjugation method.
[0010] In one embodiment, in step S30, the linker sequence includes a first linker sequence and a second linker sequence respectively disposed at both ends of the transposase. The first linker sequence is PrimerB, and the sequence of PrimerB is shown as SEQ ID NO.6: wherein, The second linker sequence is the sequence of primerC-1 shown as SEQ ID NO.7; and / or, The second linker sequence is the sequence of primerC-2 shown as SEQ ID NO.8.
[0011] In one embodiment, the adapter sequence further includes primerA, and the sequence of primerA is as shown in SEQ ID NO. 9.
[0012] In one embodiment, step S10 includes: Providing a microfluidic chip system: Co-introducing a plurality of nucleic acid fragments containing the adapter sequence and the microdroplets with cell tags into the microfluidic system, and performing a second ligation reaction inside the microdroplets to label the information of the cell tags on the plurality of nucleic acid fragments.
[0013] The present invention also provides a sequencing library, including the sequencing library constructed by the detection method of high-throughput single-cell level multiplex protein-genome interaction described in any one of the above.
[0014] The present invention also provides an application of the sequencing library constructed by the detection method of high-throughput single-cell level multiplex protein-genome interaction described in any one of the above in the single-cell multiplex detection technology.
[0015] In the present invention, through multiple ligation reactions, the antibody tag and the adapter sequence are respectively ligated to the cell-tagged cell nuclei, synchronously realizing the dual labeling of the single-cell tag and the antibody tag on the genomic DNA of the target protein or protein modification interaction. This method can batch detect the genomic DNA binding sites of multiple histone modifications or transcription factors in thousands of single cells at one time. Compared with the traditional single-cell multiplex CUT&Tag technology that can only target 2-5 proteins or modifications at a time, this invention does not require separate customization of proteins, so the number of targeted proteins is not limited, and the operation method is simple, the cost is low, and the process is short, which is especially suitable for tumor heterogeneity and developmental biology research. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0017] Figure 1 It is a flowchart of the detection method of high-throughput single-cell level multiplex protein-genome interaction according to an embodiment of the present invention; Figure 2 It is a structural diagram of the antibody nucleic acid tag sequence according to an embodiment of the present invention; Figure 3 It is a schematic diagram of the adapter process of the adapter sequence of the Tn5 transposase in an embodiment of the present invention; Figure 4 It is a sequencing Read structure display diagram of different adapter sequences of Tn5 transposase in an embodiment of the present invention; Figure 5 It is a comparison diagram of the proportion of high-quality DNA fragments capable of overlapping peaks of different-length antibody tags in an embodiment of the present invention; Figure 6 It is about the peaks of different-length antibody tags and traditional CUT&Tag in an embodiment of the present invention Figure 1 Consistency index comparison diagram.
[0018] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0019] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. For those not specified in the embodiments, they are carried out according to conventional conditions or conditions recommended by the manufacturer. Those reagents or instruments without indicating the manufacturer can be obtained as conventional products through commercial purchase. In addition, the meaning of "and / or" appearing throughout the text includes three parallel solutions. Taking "A and / or B" as an example, it includes solution A, solution B, or a solution that satisfies both A and B at the same time. In addition, the technical solutions between the embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0020] Currently, existing single-cell multiplex detection technologies, such as uCoTargetX, MULTI-CUT&Tag, and Nano-CUT&Tag technologies, can detect at most 2-5 proteins simultaneously, require sequential labeling or antibody multiplexing, and have limitations such as tag interference of transposases, such as Tn5 transposase (abbreviated as Tn5), and cumbersome experimental procedures, or rely on special reagents, such as the need to re-prepare new pA / pG-Tn5 and antibody complexes or Tn5-Nanobody fusion proteins, resulting in problems such as high costs and reduced efficiency.
[0021] In view of this, the present invention provides a method for detecting high-throughput single-cell level multiplex protein-genome interactions, including the following steps: S10. Provide cell nuclei; S20. Mix the cell nuclei with antibody tags to label the target proteins in the cell nuclei; S30. Provide a transposase with an adapter sequence, and fragment the nucleic acids in the cell nucleus using the transposase with the adapter sequence to obtain a plurality of nucleic acid fragments containing the adapter sequence; S40. Perform a first ligation reaction on the plurality of nucleic acid fragments containing the adapter sequence and the labeled target protein to label the information of the target protein on the plurality of nucleic acid fragments containing the adapter sequence; S50. Provide microdroplets with cell tags, and perform a second ligation reaction on the microdroplets with cell tags and the plurality of nucleic acid fragments containing the adapter sequence to label the information of the cell tags on the plurality of nucleic acid fragments; S60. Combine the information of the cell tags and the information of the target protein, perform library construction and sequencing to obtain the information of the protein-genome interaction of single-cell genomics in the test sample.
[0022] In the present invention, through multiple ligation reactions, the antibody tag and the adapter sequence are respectively ligated to the cell nucleus with cell tags, synchronously realizing the dual labeling of the genomic DNA of the target protein or protein modification interaction by single-cell tags and antibody tags. This method can batch detect the genomic DNA binding sites of multiple histone modifications or transcription factors in thousands of single cells at one time. Compared with the traditional single-cell multiplex CUT&Tag technology that can only target 2-5 proteins or modifications at a time, this invention does not require customizing proteins separately, so the number of targeted proteins is not limited, and the operation method is simple, the cost is low, and the process is short, which is particularly suitable for tumor heterogeneity and developmental biology research.
[0023] It should be noted that the two ligation reactions in steps S40 and S50 occur in microdroplets, and there is no chronological order involved. They can be completed simultaneously or in a stepwise reaction. After the two ligation reactions are completed, the obtained nucleic acid fragments already carry the target protein tag and the cell tag at the same time. Library construction and sequencing can be performed based on these double-labeled nucleic acid fragments.
[0024] Specifically, the cell nuclei need to be permeabilized in advance so that antibodies can enter the cell nuclei and bind to the target sites on chromatin. The permeabilized cell nuclei are incubated with antibodies carrying nucleic acid tags to obtain cell nuclei with antibody tags, where the cells target specific proteins or other modifications. Then, the antibody-treated cell nuclei are randomly fragmented by the Tn5 transposase complex. Next, in a droplet environment through a water-in-oil microfluidic device, the 5'-end of one linker of the antibody nucleic acid tag is ligated to the Tn5-fragmented DNA fragment in the vicinity, and the 5'-end of the other Tn5 linker of the DNA fragment is ligated to the cell tag carried by the hydrogel bead. After that, a sequencing library is constructed and high-throughput sequencing is performed. Finally, the protein-genome interaction sites are determined through bioinformatics analysis.
[0025] It should be noted that only those nucleic acid fragments located near the target protein or indirectly connected to the target protein through a certain mechanism (such as cross-linking in chromatin immunoprecipitation ChIP) may obtain information about the target protein.
[0026] Specifically, as Figure 1 shown, it includes the following steps: Step 1: Extract the cell nuclei from the cells, ensuring that the experiment only operates on the genomic DNA within the cell nuclei to avoid interference from other cell components. Then, the cell nuclei are incubated with specific antibodies (oligo-Antibodies), i.e., antibody tags, so that the antibody tags bind to the target proteins, thereby marking specific genomic regions for identifying and localizing the protein-DNA interaction sites of interest.
[0027] Step 2: Treat the cell nuclei incubated with antibodies with Tn5 transposase to fragment the genomic DNA and simultaneously insert linker sequences (Primer A, Primer B, and Primer C) at both ends of the DNA fragments. These linkers contain the universal sequences required for subsequent PCR amplification and sequencing and possible barcodes or UMIs.
[0028] Step 3: Wrap the Tn5-treated cell nuclei together with gel beads containing cell barcodes (i.e., cell tags) in microdroplets. Each microdroplet contains one cell nucleus and one gel bead, and the cell barcode on the gel bead will be ligated to the DNA fragment in subsequent steps to generate a unique identifier for each cell, facilitating the differentiation of signals from different cells during subsequent data analysis and thus enabling single-cell level operations.
[0029] Step 4: Inside the microdroplet, the gel beads release the cell barcodes they carry and ligate them to the DNA fragments, adding cell-specific barcodes to each DNA fragment to ensure that all DNA fragments from the same cell can be correctly classified in subsequent analyses.
[0030] Step 5: Amplify the barcoded DNA fragments by PCR to increase the number of DNA fragments to meet the requirements of high-throughput sequencing. Meanwhile, purify the PCR products to remove excess primers and other impurities, improve the quality of the sequencing library, reduce non-specific background, and ensure the accuracy and reliability of the sequencing results. During the amplification process, additional sequences (such as i7 Index) can also be introduced for further sample differentiation and sequencing platform compatibility.
[0031] Step 6: In summary, a library suitable for high-throughput sequencing can be obtained. Then, directly use the library for sequencing. Through the sequencing data, the interaction sites between the target protein and genomic DNA, as well as the distribution of these sites at the single-cell level, can be analyzed.
[0032] In one embodiment, in step S20, the nucleic acid in the antibody tag has a sequence length of 20 - 150 nt. For example, it can be 20 nt, 30 nt, 50 nt, or 150 nt. Within this range, the data quality of the test library can be improved. If the detected rate of the obtained sequencing library is less than 20 nt or greater than 150 nt, it will lead to failure to sequence the target nucleic acid.
[0033] In one embodiment, in step S20, the sequence of the antibody tag includes: a 5'-end modification group, a Read2 sequencing primer, a first random sequence, an antibody recognition code index, a second random sequence, and a linker sequence, specifically as Figure 2As shown, the sequence of the antibody tag contains a 5'-end modification group, a Read2 sequencing primer, a random sequence, an antibody recognition code index, another random sequence, and a linker sequence. The 5'-end modification group is used to stably couple the nucleic acid tag and the antibody to achieve antibody coding; the antibody recognition code corresponds one-to-one with the antibody type, and can accurately identify and detect the protein or epigenetic modification that interacts with DNA during sequencing analysis. Since there are many variable combinations of base-encoded sequences, for example, a nucleic acid sequence with a length of 8 nt can theoretically recognize 4 to the 8th power, a total of 65,538 antibodies. Therefore, in a single experiment, the relationship between unrestricted multiple proteins or epigenetic modifications and genome interaction can be detected simultaneously, as long as there are corresponding specific antibodies; optionally, the addition of two random base sequences can improve the sequencing quality of the antibody recognition code; the linker sequence is used to connect to the 5' end of the adapter primerC in the microdroplet and the genomic DNA interrupted by the transposase; this design not only improves the binding stability between the antibody and the nucleic acid tag, but also facilitates subsequent bioinformatics analysis.
[0034] In one embodiment, the sequence of the antibody tag includes S1 to S5, wherein the sequences of S1 to S5 are shown in SEQ ID NO.1, SEQ ID NO.2, SEQ ID NO.3, SEQ ID NO.4, and SEQ ID NO.5, respectively.
[0035] Among them, the specific sequence of S1 is as follows: NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNNNN-index-NNNNNNNNNGCTTTAAGGCGTTAGGTGATTA; The specific sequence of S2 is as follows: NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNN-index-NNGTTAGGTGAT*T*A; The specific sequence of S3 is as follows: NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNN-index-GGTGAT*T*A; The specific sequence of S4 is as follows: NH2-TTCCTTGGCACCCGAGAATTCCANN-index1-GTTAGGTGAT*T*A; The specific sequence of S4 is as follows: NH2-TTCCTTGGCACCCGAGAATTCCANN-index1-GGTGAT*T*A; The specific sequence of S5 is shown as follows: 5Biotin-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNNNN-TGACCAGTTCCGCAT-NNNNNNNNNGCTTTAAGGCCGGTCCTAGC*A*A; It should be noted that * all represent phosphorothioate bond modifications.
[0036] In one embodiment, in step S20, the antibody tag is formed by coupling an antibody and a nucleic acid: wherein, the antibody is used to target the target protein, and the nucleic acid is used to label the cell nucleus so as to achieve the labeling of the cell nucleus. These nucleic acid tags can be further coupled with a fluorescent label or other reporter molecules for visualization.
[0037] In one embodiment, the coupling method includes the streptavidin-biotin coupling method or the amino-crosslinker-mediated covalent coupling method. In the implementation of the present invention, the direct binding of streptavidin-coupled antibody and biotin-modified nucleic acid is adopted, or the covalent binding mediated by amino-modified antibody and amino-reactive crosslinker is adopted. Both of these methods have high coupling efficiency and stability and can meet different experimental requirements.
[0038] Specifically, the antibody streptavidin of the target protein is coupled, and then a DNA oligonucleotide containing biotin modification is synthesized. Using the extremely strong specific binding ability of antibody streptavidin and biotin, the DNA oligonucleotide is connected to the target antibody to identify the target protein labeled by the DNA oligonucleotide antibody during sequencing.
[0039] In one embodiment, in step S30, the linker sequence includes a first linker sequence and a second linker sequence respectively disposed at both ends of the transposase. The first linker sequence is PrimerB, and the sequence of PrimerB is shown in SEQ ID NO.6: wherein, The second linker sequence is the sequence of primerC-1 as shown in SEQ ID NO.7; and / or, The second linker sequence is the sequence of primerC-2 as shown in SEQ ID NO.8.
[0040] Different from the adapter sequences in the traditional Nextera library construction kit and the default sequencing primers of the sequencer, the adapter sequence of the present invention is used to start sequencing from the set nucleic acid label Read2 sequencing primer. This design ensures that the sequencing read length can accurately cover the antibody label region while detecting the genomic DNA sequence, thereby improving the accuracy and sensitivity of the detection.
[0041] It should be noted that the use of the original PrimerC sequence will cause Reads to be unable to read the antibody nucleic acid tag sequence, that is, there is a certain compatibility problem between the original PrimerC sequence and the subsequent PCR primer or sequencing primer, resulting in an unsatisfactory binding position of the Read2 sequencing primer. The sequencing process skips the target area, causing it to start sequencing directly from the genomic DNA cut by Tn5, thereby ignoring the antibody tag and PrimerC itself connected to its 5' end. The PrimerC-1 and PrimerC-2 sequences are introduced to replace the original PrimerC sequence respectively, so that the Read2 sequencing primer can more effectively bind to the expected site instead of directly binding to primerC to extend the sequencing, thereby correctly reading the sequence of the antibody tag. The replacement principle is: the replaced primer is not complementary to the primer in the traditional Nextera library construction kit and the default sequencing primer of the sequencer, avoiding competition between the two sequencing primers, so that the second adapter sequence of the present invention can be primerC-1 or primerC-2.
[0042] Specifically, the steps for replacing the original PrimerC sequence with primerC-1 or primerC-2 are as follows: The T7 sequence needs to be modified to make the Rd2 SeqPrimer compete with the nextera N7 primer during sequencing. In order to prevent the modified sequence from competing with the Read2 sequencing primer and reduce nonspecific extension during library construction, the modification principles are as follows: 1. The first two bases are selected from AC, GC or GT; 2. The three-base sequence consisting of GT and the first base of the generated sequence in chronological order cannot appear in AGATCGGAAGAGCGTCGTGTAG; ACGAGCAACGACGGACGACAGCAA; TTGCTGTCGTCCGTCGTTGCT; CGAAATTCCGGCCAGGATCGTT; TTTAAGGCCGGTCCTAGCAA; Among the 7 sequences of ATGCGGAACTGGTCA; AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC, note that the front-back order means that the GT bases are arranged in front of the first base of the generated sequence; The three-base sequence formed by T and the first two bases of the generated sequence in the front-back order cannot appear in AGATCGGAAGAGCGTCGTGTAG; ACGAGCAACGACGGACGACAGCAA; TTGCTGTCGTCCGTCGTTGCT; CGAAATTCCGGCCAGGATCGTT; TTTAAGGCCGGTCCTAGCAA; Among the 7 sequences of ATGCGGAACTGGTCA; AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC, note that the front-back order means that the T base is in front of the first two bases of the generated sequence; 4. If there is no sequence that meets the requirements of condition 2 and condition 3, the condition can be relaxed to that the 4-base sequence formed by GT and the first two bases of the generated sequence in the front-back order cannot appear in the following sequences: AGATCGGAAGAGCGTCGTGTAG; ACGAGCAACGACGGACGACAGCAA; TTGCTGTCGTCCGTCGTTGCT; CGAAATTCCGGCCAGGATCGTT; TTTAAGGCCGGTCCTAGCAA; Among the 7 sequences of ATGCGGAACTGGTCA; AGATCGGAAGAGCACACGTCTGAACTCCAGTCAC, select AC for the first two bases in conditions 1-5; 5. The generated sequence does not contain the sequences GTT, CGT, GTA, and TAC; 6. There cannot be 3 consecutive identical bases in the generated sequence; Finally, obtain the specific sequence of primerC-1: ACGATGCATT; the specific sequence of PrimerC-2: ACGGCATGAT. Other sequences based on the above principles should also be within the scope of protection, such as ACGCTACGAT, ACGATCGTCA, or ACGATCGTCA, etc.
[0043] In one embodiment, the adapter sequence further includes primer A, and the sequence of primer A is shown in SEQ ID NO.9. Primer A is mainly used to bind to the Tn5 protein to form a stable transposase complex. It generally does not directly participate in the insertion into genomic DNA but serves as a structural support.
[0044] In one embodiment, step S10 includes: providing a microfluidic chip system, co-introducing a plurality of nucleic acid fragments containing the adapter sequence and the microdroplets with cell tags into the microfluidic system, and performing a second ligation reaction inside the microdroplets to label the information of the cell tags on the plurality of nucleic acid fragments.
[0045] It should be noted that in the experiment, the microfluidic device can generate hundreds of thousands of microdroplets. Each droplet contains a single cell that can be resolved informatically and a unique cell tag for each droplet. The genomic DNA is labeled in the microdroplets by the antibody tag and the cell tag through the microfluidic device. The DNA derived from the same cell is labeled with the same cell tag, and the DNA derived from different cells is labeled with different cell tags.
[0046] The present invention provides a sequencing library, including the sequencing library constructed by the method for detecting high-throughput single-cell level multiplex protein-genome interaction as described above. This sequencing library has all the technical solutions of the method for detecting high-throughput single-cell level multiplex protein-genome interaction and the method for detecting high-throughput single-cell level multiplex protein-genome interaction. Therefore, it has all the beneficial effects of the method for detecting high-throughput single-cell level multiplex protein-genome interaction and the method for detecting high-throughput single-cell level multiplex protein-genome interaction. The present invention will not elaborate on them one by one here.
[0047] The present invention also provides an application of the sequencing library constructed by the method for detecting high-throughput single-cell level multiplex protein-genome interaction as described above in the single-cell multiplex detection technology.
[0048] The sources of raw materials for the following examples are as follows: K562 cell line: a self-propagating cell line in the laboratory H3K27Ac specific antibody: recombinant Anti-Histone H3 (acetyl K27) antibody [EP16602] - ChIP Grade - BSA and Azide free, Abcam (Abcam (Shanghai) Trading Co., Ltd.), product number ab302877; Streptavidin Conjugation Kit - Lighting-link®: Abcam (Abcam (Shanghai) Trading Co., Ltd.), product number ab102921; Commercially available SeekOne® DD Single Cell ATAC + RNA Dual Omics Kit: Beijing Xunyin Biotechnology Co., Ltd., product numbers K02901 - 02; K02901 - 08.
[0049] Example: High - throughput co - detection application of single - cell genomic DNA and RNA Example 1: Influence of different Tn5 coating sequences on sequencing results 1. Preparation of streptavidin - labeled primary antibody: Use the Abcam lighting - link kit to conjugate the histone modification H3K27Ac - specific antibody with streptavidin SA; 2. Synthesis of nucleic acid tags in the antibody tag: Synthesize a 5′ - Biotin - modified nucleic acid antibody tag, and the specific sequence is as follows: 5Biotin - GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNNNN - TGACCAGTTCCGCAT - NNNNNNNNNGCTTTAAGGCCGGTCCTAGC*A*A; 3. Preparation of antibody with nucleic acid tags, that is, antibody tag: Combine the conjugated streptavidin antibody product in step 1 and the excessive biotin - modified nucleic acid product in step 2 to obtain an antibody tag with nucleic acid tags; 4. Take fresh K562 cells, add NIB nuclear extraction buffer and incubate on ice for 5 min, centrifuge at 500 g and resuspend in PBS buffer, adjust the cell density to 1×10^6 / mL to obtain K562 cell nuclei; 5. Incubate the permeabilized K562 cell nuclei overnight with the antibody containing nucleic acid tags at a volume ratio of 1:100, add 1 mL PBS buffer, centrifuge at 500 g for 5 min and resuspend in PBS buffer, repeat the washing 3 times; 6. The operation steps of the commercially available SeekOne® DD single - cell ATAC + RNA dual omics kit are briefly as follows: 6.1 Use Tn5 transposase to treat the nuclei after antibody incubation to perform non - discriminatory genomic DNA fragmentation. Among them, the original adapter sequence PrimerC - 0 of Tn5 is modified to PrimerC - 1 and PrimerC - 2 as the experimental group, and PrimerB is the PrimerB - 1 sequence. The specific adapter sequences are as follows: The sequence of PrimerA is shown as follows: The sequence of 5'-phos-CTGTCTCTTATACACATCT-NH2-3' PrimerB-1 is shown as follows: 5′p-CGTCCGTCGTTGCTCGT-AGATGTGTATAAGAGACAG-3′; The sequence of PrimerC-0 is shown as follows: 5′P-GTCTCGTGGGCTCGG-AGATGTGTATAAGAGACAG-3′; The sequence of PrimerC-1 is shown as follows: 5'p-ACGATGCATT-AGATGTGTATAAGAGACAG; The sequence of PrimerC-2 is shown as follows: 5'p-ACGGCATGAT-AGATGTGTATAAGAGACAG. The above sequences are processed according to the Figure 3 adapter flow chart to complete the adapter process, that is, equimolar mixing of primerA and primerB-1 to form PrimerAB; equimolar mixing of primerA and primerC-0 / C-1 / C-2 to form PrimerAC-0 / AC-1 / AAC-2; then equi-volume mixing of primerAB and primerAC and coating with Tn5 transposase to complete the coating of the naked transposase.
[0050] 6.2 Use the water-in-oil microfluidic device according to the kit instructions to ligate the antibody tag in the cell nucleus with the 5' end of PrimerC of the DNA interrupted by Tn5 in the droplet, and at the same time ligate the cell tag carried by the hydrogel beads in the droplet with the 5' end of primerB of the DNA interrupted by Tn5 to achieve cell labeling of each cell nucleus; only the costs related to ligation are added to the reaction mixture, and the transcriptome-related components such as dNTP are replaced with deionized water; 6.3 Based on illumina Truseq Read1SeqPrimer and microRNA Read2 Sequencing primer, directly amplify the DNA library and sequence it; 7. Analyze the sequencing results, and the results are as Figure 4 shown: It shows that the starting position of Read2 sequencing of the default Tn5 library construction adapter PrimerC-0 directly starts from the Tn5 interrupted sequence, and the antibody nucleic acid tag sequence and the PrimerC-0 sequence cannot be detected; while the starting positions of read2 sequencing of the modified primerC-1 and PrimerC-2 both start from the preset antibody tag and have the correct sequencing structure.
[0051] It should be noted thatFigure 4 The difference lies in the read2 sequence. After using the primerC-0-coated transposase, there is no antibody tag sequence in the final library structure. It is speculated that during sequencing, Rd2 SeqPrimer and nextera N7 primer compete, resulting in incorrect sequencing results for this library. When primerC-0 is replaced with primerC-1 or primerC-2, a normal library structure with antibody tags can appear, indicating normal library construction and sequencing results (the antibody tags are the dark green and purple sequences in read2 in the figure, the bright blue and yellow are the transposase adapter sequences obtained by normal sequencing, and the green is the fixed sequence in the cell tag used to connect to the transposase adapter).
[0052] Example 2: Influence of antibody tags of different lengths on marker detection 1. Synthesis of antibody tags of different lengths: Antibody tags of different lengths were given to Sangon Biotech (Shanghai) Co., Ltd. for synthesis of 5'-amino-modified antibody tags, named AbB-20nt, AbB-50nt, AbB-58nt, AbB-81nt, and AbB-150nt respectively. The specific sequence of AbB-20nt is as follows: NH2-ACCCGAGAATTCCA-(Index)1bp-GAT*T*A The specific sequence of AbB-81nt is as follows: NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNNNN-(Index)6bp-NNNNNNNNNGCTTTAAGGCGTTAGGTGAT*T*A; The specific sequence of AbB-58nt is as follows: NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNN- (Index)6bp-NNGTTAGGTGAT*T*A; The specific sequence of AbB-50nt is as follows: NH2-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNN-(Index) 6bp-GGTGAT*T*A; The specific sequence of AbB-150nt is as follows: NH2-CAAGCAGAAGACGGCATACGAGATGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNNNNNNNNNNNNNN-(Index) 30bp-NNNNNNNNNNNNNNNNNNNNGCTTTAAGGCGTTAGGTGAT*T*A; It should be noted that N here represents any base (A, T, C, G). The segment "NNNNNNNNN..." is an unknown or variable sequence before the "Index" region. According to specific experimental designs, it may be used to insert specific recognition sequences or for other purposes. The antibody tag design must have a read2 sequencing primer, which is about 20bp in length. Without the sequencing primer sequence, sequencing cannot be performed. Currently, the mainstream sequencing read length of next-generation sequencing is 150bp. If the tag is greater than or equal to 150bp, only the tag sequence can be detected, and the genomic DNA sequence labeled by the tag cannot be detected.
[0053] 2. Use the Abcam oligonucleotide conjugation kit (ab218260) to conjugate the antibody tags synthesized in step 1 to the histone modification H3K27Ac-specific antibody respectively; 3. Take fresh K562 cells, add NIB nuclear extraction buffer and treat on ice for 5 min. After centrifugation at 500g, resuspend in PBS buffer, and adjust the cell density to 1×10^6 / mL to obtain K562 cell nuclei; 4. Incubate the permeabilized K562 cell nuclei overnight with the prepared nucleic acid tag antibodies at a volume ratio of 1:100. Add 1 mL of PBS buffer, centrifuge at 500g for 5 min, and then resuspend in PBS buffer. Repeat the washing 3 times; 5. The operating steps of the commercial SeekOne® DD single-cell ATAC + RNA dual-omics kit are briefly as follows: 5.1 Use Tn5 transposase to treat the nuclei after antibody incubation to perform non-discriminatory genomic DNA fragmentation; the adapter sequences of the Tn5 transposase include PrimerA, PrimerB-1, and PrimerC-1. Among them, the specific sequence of PrimerA is as follows: 5'-p-CTGTCTCTTATACACATCT-NH2-3'; The specific sequence of PrimerB-1 is as follows: 5′P-CGTCCGTCGTTGCTCGT-AGATGTGTATAAGAGACAG-3′; The specific sequence of PrimerC-1 is as follows: 5′P-ACGATGCATT-AGATGTGTATAAGAGACAG-3.
[0054] 5.2 Use the water-in-oil microfluidic device according to the kit instructions to ligate the antibody tag in the nucleus with the 5'-end of PrimerC of the DNA interrupted by Tn5 in the droplet, and at the same time ligate the cell tag carried by the hydrogel beads in the droplet with the 5'-end of primerB of the DNA interrupted by Tn5, so as to achieve cell labeling for each nucleus; only the costs related to ligation are added to the reaction mixture, and the transcriptome-related components such as dNTP are replaced with deionized water; 5.3 Directly amplify the DNA library based on the illumina Truseq Read1 Seq Primer and TruSeq Read2 Sequencing primer and sequence it; 6. Analyze the sequencing results, as Figure 5 shown, Figure 5 is a graph of the proportion of high-quality DNA fragments that can overlap with the peaks (Fraction of high-quality fragments overlapping peaks). The results show that the shorter the nucleic acid tag conjugated with the antibody, the higher the Fraction reads overlapping peaks index. This index reflects the proportion of sequencing reads that fall within the known H3K27Ac enrichment regions (peaks). A higher proportion usually means a higher signal-to-noise ratio and better targeting. Shorter antibody nucleic acid tags (50 nt and 58 nt) have a higher Fraction reads overlapping peaks compared to longer tags (81 nt). This is because shorter tags have more advantages in terms of ligation efficiency, PCR amplification efficiency, or sequencing efficiency, or have less impact on the accessibility to the Tn5 enzyme. As Figure 6 shown, compared with the results of ordinary CUT&Tag, it has a similar peak signal of specific H3K27Ac. The peak signal of H3K27Ac obtained by using the method of labeled antibody is similar to that of the traditional CUT&Tag method. CUT&Tag is a technique that directly performs antibody binding and enzymatic digestion in the nucleus and is considered one of the gold standards for studying chromatin states. The similar peak signals indicate that the method based on labeled antibody can accurately label and enrich the target chromatin region.
[0055] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the patent protection scope of the present invention.
Claims
1. A method for detecting high-throughput multi-protein-genome interactions at the single-cell level, characterized in that, It includes the following steps: S10. Provide a cell nucleus; S20. Mix the cell nucleus with an antibody tag to label the target protein in the cell nucleus; S30. Provide a transposase with an adapter sequence, and use the transposase with the adapter sequence to fragment the nucleic acid in the cell nucleus to obtain a plurality of nucleic acid fragments containing the adapter sequence; S40. Perform a first ligation reaction on the plurality of nucleic acid fragments containing the adapter sequence and the labeled target protein to label the information of the target protein on the plurality of nucleic acid fragments containing the adapter sequence; S50. Provide a microdroplet with a cell tag, and perform a second ligation reaction on the microdroplet with the cell tag and the plurality of nucleic acid fragments containing the adapter sequence to label the information of the cell tag on the plurality of nucleic acid fragments; S60. Combine the information of the cell tag and the information of the target protein, perform library construction and sequencing to obtain the information of the protein-genome interaction of single-cell genomics in the sample to be tested.
2. The detection method for multiplex protein-genome interaction at the high-throughput single-cell level according to claim 1, characterized in that In step S20, the sequence length of the nucleic acid in the antibody tag is 20-150 nt.
3. The detection method for multiplex protein-genome interaction at the high-throughput single-cell level according to claim 1, wherein In step S20, the sequence of the antibody tag includes: a 5'-end modification group, a Read2 sequencing primer, a first random sequence, an antibody recognition code index, a second random sequence, and a ligation sequence.
4. The detection method for multiplexed protein-genome interaction at the high-throughput single-cell level according to claim 3, wherein The sequence of the antibody tag includes S1-S5, wherein the sequences of S1-S5 are shown in SEQ ID NO.1, SEQ ID NO.2, SEQ ID NO.3, SEQ ID NO.4, and SEQ ID NO.5 respectively.
5. The detection method for high-throughput single-cell level multiplex protein-genome interaction according to claim 1, wherein, In step S20, the antibody tag is formed by conjugating an antibody and a nucleic acid: wherein, the antibody is used to target the target protein, and the nucleic acid is used to label the target protein; and / or, the conjugation method includes a streptavidin-biotin conjugation method; or, the conjugation method includes an amino-crosslinker-mediated covalent conjugation method.
6. The detection method of multiplex protein-genome interaction at the high-throughput single-cell level according to claim 1, characterized in that, In step S30, the adapter sequence includes a first adapter sequence and a second adapter sequence respectively arranged at both ends of the transposase. The first adapter sequence is PrimerB, and the sequence of PrimerB is shown in SEQ ID NO.6: wherein, the second adapter sequence is the sequence of primerC-1 shown in SEQ ID NO.7; and / or, the second adapter sequence is the sequence of primerC-2 shown in SEQ ID NO.
8.
7. The detection method of multiplex protein-genome interaction at the high-throughput single-cell level according to claim 6, wherein, The adapter sequence further includes primerA, and the sequence of primerA is shown in SEQ ID NO.
9.
8. The detection method for multiplex protein-genome interaction at the high-throughput single-cell level according to claim 1, characterized in that Step S50 includes: Provide a microfluidic chip system; Co-import a plurality of nucleic acid fragments containing the adapter sequence and the microdroplet with the cell tag into the microfluidic system, and perform a second ligation reaction inside the microdroplet to label the information of the cell tag on the plurality of nucleic acid fragments.
9. A sequencing library, characterized in that, A sequencing library constructed by the method for detecting high-throughput single-cell level multiplex protein-genome interaction according to any one of claims 1 to 8.
10. Application of a sequencing library constructed by the method for detecting high-throughput single-cell level multiplex protein-genome interaction according to claim 9 in single-cell multiplex detection technology.
Citation Information
Patent Citations
Improvements in rotors for wind powered electric generators
EP0016602A1
Ultrahigh-flux single-cell chromatin transposase accessibility sequencing method
CN113604545A
Sequencing method for simultaneously acquiring whole genome transcription and protein-DNA binding information
CN115851876A
Single cell DNA-protein interaction sequencing kit and method based on combinatorial index
CN118272509A
Methods for detecting protein binding sequences and tagging nucleic acids
US20180335424A1