Specific methylation sequence marking system for living cells
Through the activity of dCas9 and MBD protein binding to TurboID protein, specific markers of methylated regions are achieved in living cells, solving the problem that DNA methylation and sequence cannot be recognized simultaneously in the prior art, and efficient screening of methylation site-regulating proteins is achieved.
Patent Information
- Application Number
- CN202510356929.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
AI Technical Summary
Existing proximity marker and high-throughput sequencing techniques are difficult to simultaneously identify DNA methylation modifications and specific sequences in genomic regions in living cells, resulting in the inability to effectively screen regulatory proteins at specific methylation sites.
A specific methylation sequence marker system for living cells is used to target the target genomic region through dCas9 protein, bind to MBD protein to recognize methylated DNA, and use the N- and C-terminal catalytic activities of TurboID protein to form biotinylated markers to achieve specific markers of the methylated region.
It can directly image and observe and screen regulatory proteins at specific methylation sites in living cells, reduce catalytic background signals, and improve label specificity and accuracy.
Smart Images

Figure CN120249398A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of molecular biology, cell biology, and epigenetics. Specifically, the present invention relates to a labeling system for specific methylation sequences in living cells. Background Art
[0002] Breakthrough progress has been continuously made in the study of protein interactions in living cells by proximity labeling methods. This method makes up for the deficiencies of traditional co-immunoprecipitation methods and can capture transient, low-affinity, or membrane-dependent interacting proteins. TurboID is a recently discovered biotin ligase that can be used in proximity labeling methods. It has the characteristics of high catalytic efficiency, short catalytic time, and can catalyze in living cells. Fusing TurboID with the target protein and adding a biotin substrate for catalytic reaction can label the proteins near the target region with biotin, and then enriching the labeled proteins with avidin magnetic beads for mass spectrometry identification can obtain potential regulatory proteins near the target region.
[0003] The CRISPR / Cas9 system combined with proximity labeling proteins can be used to capture DNA-protein interactions in specific regions of the genome. Fusing a biotin ligase to the dCas9 protein and then introducing an sgRNA targeting the target sequence can proximally label the proteins near the genomic region of interest in living cells. However, these reported methods cannot identify the epigenetic modifications of genomic regions.
[0004] High-throughput sequencing technology has greatly promoted the research on DNA methylation and its dynamic changes. Many new high-throughput sequencing strategies use methylation recognition protein sequences, such as the MBD protein domain, to enrich and analyze methylated or unmethylated DNA. However, the binding regions of the epigenetically differential regulatory proteins captured by it usually cover the entire genome without sequence differences, rather than specifically targeting the chromatin regions of interest. Therefore, developing a labeling system that can simultaneously identify methylation modifications and sequences on the genome can solve this bottleneck. Summary of the Invention
[0005] To solve the problems in the prior art, the present invention provides a labeling system for specific methylation sequences in living cells
[0006] In the first aspect of the present invention, a labeling system for specific methylation sequences in living cells is provided, which includes nucleic acid construct I and nucleic acid construct II;
[0007] The nucleic acid construct I includes at least a eukaryotic cell promoter sequence, a dCas9 protein sequence, a fluorescent protein sequence, and a C-terminal sequence of the TurboID protein from the 5'-end to the 3'-end;
[0008] The nucleic acid construct II includes at least a eukaryotic cell promoter sequence, a methyl - binding protein MBD domain, a fluorescent protein sequence, an N - terminal sequence of the TurboID protein, a murine U6 promoter sequence, and an sgRNA sequence that can be recognized by the dCas9 protein and target a target site, from the 5'-end to the 3'-end.
[0009] In another preferred embodiment, the eukaryotic cell promoter sequence in the nucleic acid construct I and the nucleic acid construct II is the SFFV promoter sequence.
[0010] In the second aspect of the present invention, there is provided an expression vector for a live - cell specific methylation sequence proximity - labeling fusion protein, and the nucleic acid construct I and the nucleic acid construct II in the live - cell specific methylation sequence labeling system are included on the vector.
[0011] In the third aspect of the present invention, there is provided a genetically engineered cell, and an exogenous DNA sequence is integrated into the genome of the genetically engineered cell, and the exogenous DNA sequence includes the nucleic acid construct I and the nucleic acid construct II in the live - cell specific methylation sequence labeling system.
[0012] Compared with the traditional genome labeling system, the live - cell specific methylation sequence labeling system of the present invention can directly perform imaging observation in live cells, distinguish methylation while identifying specific sequences, and thus can be applied to the study of regulatory proteins at specific methylation sites. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The following drawings are used to illustrate the specific embodiments of the present invention, and are not used to limit the scope of the present invention defined by the claims.
[0014] Figure 1 Showing the expression pattern diagram of the nucleic acid construct described in the present invention;
[0015] Figure 2 Showing the cellular localization of the dCas9 protein and the residual biotin - catalyzing activity at the C - terminal of the TurboID protein in 293T cells expressing only the nucleic acid construct 1;
[0016] Figure 3 Showing the cellular localization of the MBD protein and the residual biotin - catalyzing activity at the N - terminal of the TurboID protein in 293T cells expressing only the nucleic acid construct 2;
[0017] Figure 4 Showing the localization of the dCas9 protein and the MBD protein on human α - satellite DNA in two different 293T cells expressing both the nucleic acid construct 1 and the nucleic acid construct 2, and the biotin - catalyzing activity after dimerization of the N - terminal and C - terminal of the TurboID protein. Sub - figure (A) and sub - figure (B) are representative images and fluorescence signal statistics of the two cells.
[0018] Figure 5 The western blot verification results of biotin affinity immunoprecipitation enriched proteins after targeting the human α-satellite DNA region in the genetically engineered cells of the present invention are shown. Detailed implementation manners
[0019] The present invention will be further described and explained below in conjunction with the detailed implementation manners. The said embodiments are only demonstrations of the present disclosure content and do not delimit the scope of limitation. Without conflict, the technical features of each implementation manner in the present invention can be combined correspondingly.
[0020] The present invention has established a specific methylation sequence labeling system for living cells. By targeting the target genomic region with dCas9, the MBD protein domain binds to methylated DNA, and the N-terminal and C-terminal of the fused TurboID protein are spatially close to generate a biotin-catalyzing active dimer, which can achieve biotinylation labeling of regulatory proteins near the specified genomic hypermethylated region in living cells, so as to screen regulatory factors in the target region without any prior knowledge. This method can significantly reduce the catalytic background signal generated by using a single full-length TurboID protein for labeling, and can simultaneously recognize DNA methylation on the basis of recognizing DNA sequences.
[0021] The experimental methods in the following embodiments are all conventional methods without special instructions, and are carried out according to the literature methods or product specifications in the art. The materials and reagents in the following embodiments can be obtained from commercial channels without special instructions.
[0022] Example 1: Construction of a specific methylation sequence labeling system for living cells
[0023] The expression cassettes in the nucleic acid constructs of the present invention can all be obtained by gene synthesis, and then the expression cassettes are integrated into the second-generation lentiviral vector in a homologous recombination manner. The specific operation methods are well known to those skilled in the art.
[0024] In this embodiment, the specific methylation sequence labeling system for living cells includes nucleic acid construct I shown in Formula I and nucleic acid construct II shown in Formula II from the 5'-end to the 3'-end:
[0025] Formula I: S-C9-G-CT;
[0026] Formula II: S-MB-MC-NT-U-R;
[0027] Among them, S is a eukaryotic cell promoter sequence, C9 is the dCas9 protein sequence, G is the fluorescent protein sequence, and CT is the C-terminal sequence of the TurboID protein; MB is the methylated binding protein MBD domain, MC is the fluorescent protein sequence, NT is the N-terminal sequence of the TurboID protein, U is the murine U6 promoter sequence, and R is the sgRNA sequence that can be recognized by the dCas9 protein and target the target site.
[0028] In a preferred embodiment, the fluorescent protein sequence G in the nucleic acid construct is the enhanced green fluorescent protein EGFP sequence, and the fluorescent protein sequence MC in the nucleic acid construct II is the red fluorescent protein mCherry sequence. The eukaryotic cell promoter sequence S is the SFFV promoter sequence. Among them, the enhanced green fluorescent protein EGFP sequence, the red fluorescent protein mCherry sequence, the SFFV promoter sequence, and the dCas9 protein sequence are all well-known sequences in the art, and their sequence information can be queried and downloaded through the Addgene website (https: / / www.addgene.org / ).
[0029] In a preferred embodiment, nuclear localization signals NLS are added before and after dCas9 in the nucleic acid construct I. The nuclear localization signal sequence before the dCas9 protein sequence is shown in SEQ ID No.1; the nuclear localization signal sequence after the dCas9 protein sequence is shown in SEQ ID No.2.
[0030] In a preferred embodiment, a GS linker sequence is added before the enhanced green fluorescent protein EGFP sequence in the nucleic acid construct I, and its sequence is shown in SEQ ID No.3. In a preferred embodiment, a GS linker sequence and a MycTag tag are added after the enhanced green fluorescent protein EGFP sequence in the nucleic acid construct I, and its sequence is shown in SEQ ID No.4.
[0031] In a preferred embodiment, the C-terminal TurboID in the nucleic acid construct I is truncated from the 74th amino acid to the last amino acid of the complete TurboID protein, and the sequence is shown in SEQ ID No.5. The sequence of the complete TurboID protein is from a previously reported literature (Cho, K.F. et al. Split-TurboID enables contact-dependent proximity labeling in cells. Proc Natl Acad Sci U S A 117, 12143-12154 (2020). https: / / doi.org:10.1073 / pnas.1919528117).
[0032] In a preferred embodiment, three tandem nuclear localization signals and a 3xFlag tag were added before the MBD protein domain sequence in the nucleic acid construct II, and the sequence is shown in SEQ ID No.6.
[0033] In a preferred embodiment, the MBD protein domain in the nucleic acid construct II is derived from the MBD1 protein, and the sequence is shown in SEQ ID No.7.
[0034] In a preferred embodiment, GS linker sequences were added before and after the red fluorescent protein mCherry sequence in the nucleic acid construct II. Among them, the GS sequence shown in SEQ ID No.8 was added before the red fluorescent protein mCherry sequence, and the GS sequence shown in SEQ ID No.9 was added after the red fluorescent protein mCherry sequence.
[0035] In a preferred embodiment, the N-terminal TurboID in the nucleic acid construct II is truncated from the 2nd amino acid to the 73rd amino acid of the complete TurboID protein, and the sequence is shown in SEQ ID No.10.
[0036] In a preferred embodiment, the sgRNA in the nucleic acid construct II targets the human alpha satellite DNA region, and the sequence is shown in SEQ ID No.11. The sgRNA sequence targeting the human alpha satellite DNA region is derived from a previously reported literature (Gao, X.D. et al. C-BERST: defining subnuclear proteomic landscapes at genomic elements with dCas9-APEX2. Nat Methods 15, 433-436 (2018). https: / / doi.org:10.1038 / s41592-018-0006-2).
[0037] In a preferred embodiment, the sgRNA in the nucleic acid construct II is a negative control sequence that has no target in the human genome, and the sequence is shown in SEQ ID No.12. The sgRNA sequence of the negative control is derived from a previously reported literature (Chen, B. et al. Dynamic imaging of genomic loci in living human cells by an optimized CRISPR / Cas system. Cell 155, 1479-1491 (2013). https: / / doi.org:10.1016 / j.cell.2013.12.001).
[0038] In a preferred embodiment, an RNA scaffold sequence is contained after the sgRNA in the nucleic acid construct II, and the sequence is as shown in SEQ ID No. 13. The scaffold sequence is derived from a previously reported literature (Chen, B. et al. Dynamic imaging of genomic loci in living human cells by an optimized CRISPR / Cas system. Cell 155, 1479-1491 (2013). https: / / doi.org:10.1016 / j.cell.2013.12.001).
[0039] The vector sequence pattern diagram of the most preferred example of this example is shown in the appendix Figure 1 。
[0040] Example 2: Construction of a genetically engineered cell line containing the specific methylation sequence labeling system for living cells
[0041] The present invention provides a genetically engineered cell, in which the expression cassettes of the nucleic acid construct I and the nucleic acid construct II of the present invention are integrated into the genome for realizing specific methylation sequence labeling of living cells. Among them, the expression pattern diagrams of the nucleic acid construct I and the nucleic acid construct II used in Example 2 are as Figure 1 shown. In this example, the nuclear localization signals introduced in Example 1 are respectively added before and after dCas9 of the nucleic acid construct I, the GS linker sequence introduced in Example 1 is added before the enhanced green fluorescent protein EGFP sequence, and the GS linker sequence and MycTag tag introduced in Example 1 are added after the EGFP sequence. Three tandem nuclear localization signals and 3xFlag tag introduced in Example 1 are added before the MBD protein domain sequence of the nucleic acid construct II, and the GS linker sequences introduced in Example 1 are added before and after the red fluorescent protein mCherry sequence.
[0042] In the process of preparing the genetically engineered cell, preferably the second-generation lentivirus packaging and infection are used, and the cell population expressing the fluorescent protein is sorted by flow cytometry. The specific operation method is well known to those skilled in the art. Preferably, the positive cell population sorted by flow cytometry is weakly positive for GFP to ensure that the copy number of dCas9 will not be too high to cause false positive labeling, and strongly positive for mCherry to ensure sufficient copy numbers of the MBD protein domain and sgRNA.
[0043] In a preferred embodiment, the sgRNA of the nucleic acid construct 2 expression cassette contained in the genetically engineered cell of the present invention targets human alpha satellite DNA, and the labeling situation and biotinylation level of the target region can be observed by confocal imaging. The specific implementation process is as follows:
[0044] The genetically engineered cells were seeded in a confocal dish with a glass bottom. When the confluence reached 90%, 50 μM biotin was added and incubated for 1 hour. After terminating the catalytic reaction with PBS, immunofluorescence staining was performed. An avidin antibody conjugated with Alexa Flour 647 fluorescent molecules and DAPI dye were used to label biotinylated proteins and DNA. The fluorescence in each channel was observed under a confocal microscope. The genetically engineered cells containing only nucleic acid construct I are shown in Appendix Figure 2 , and the genetically engineered cells containing only nucleic acid construct II are shown in Appendix Figure 3 . In both cases, no effective biotin labeling could be produced in the cells; the genetically engineered cells containing both nucleic acid constructs I and II, i.e., this example, are shown in Appendix Figure 4 . Effective biotin labeling was produced at the satellite sequence, as indicated by the arrow. Appendix Figure 4 's statistical chart is the statistical analysis of the fluorescence signals in each channel at the enrichment site of the fusion protein produced by nucleic acid construct I.
[0045] Example 3: Screening for regulatory factors on specific methylated sequences in living cells
[0046] After adding a biotin substrate for catalysis to the genetically engineered cells provided in Example 2 of the present invention, the biotinylated proteins can be enriched using avidin magnetic beads, and potential binding proteins at the target site can be obtained by mass spectrometry identification. Therefore, the present invention can be applied to the screening of regulatory factors for specific methylated sequences in living cells.
[0047] In a preferred embodiment, the genetically engineered cells are set as an experimental group and two control groups. Among them, the sgRNA in the nucleic acid construct II expression cassette of the experimental group cells targets a hypermethylated target region or element; the sgRNA in the nucleic acid construct II expression cassette of one of the control group cells is a control sequence without a target site in the genome; the MBD protein domain in the nucleic acid construct II expression cassette of the other control group cells is an inactivated protein with a 44th amino acid mutation, and the sequence is shown in SEQ ID No. 14.
[0048] After obtaining the above three groups of genetically engineered cells, the specific implementation process is as follows:
[0049] (1) After sorting GFP-positive and mCherry-positive cells, they were expanded to a 15-cm culture dish. When the confluence reached 90%, 50 μM biotin was added and incubated for 1 hour. The catalytic reaction was terminated with PBS, and the cells were digested with trypsin and collected by centrifugation.
[0050] (2) Resuspend the cell pellet in 3 mL of lysis buffer A (10 mM HEPES, 10 mM KCl, 1.5 mM MgCl2, 0.15% NP40, 1 mM PMSF, pH 7.5) on ice, vortex until fully lysed, and then centrifuge (3000 g, 4 °C, 5 minutes) to obtain a pellet.
[0051] (3) Add 1 mL of lysis buffer B (50 mM Tris-HCl, 500 mM NaCl, 1 mM EDTA, 1% NP40, 0.1% SDS, 0.5% sodium deoxycholate, 1 mM PMSF, protease inhibitor) to the pellet obtained in (2), vortex to lyse, and then sonicate the chromatin under appropriate conditions until the solution is no longer viscous.
[0052] (4) Add 200 μL of avidin magnetic beads to each tube of lysis solution, incubate with rotation at 4 °C overnight, and then wash the magnetic beads with the following solutions at room temperature in sequence: lysis buffer B, 2 minutes, twice; 1 M KCl, 2 minutes; 0.1 M sodium carbonate, 10 seconds; 2 M urea (dissolved in 10 mM Tris-HCl, pH 8.0), 10 seconds; lysis buffer B, 2 minutes, twice. The magnetic beads at this time can be boiled with Western blot loading buffer and then the supernatant can be taken for Western blot detection.
[0053] (5) The magnetic beads used for preparing mass spectrometry samples are continued to be washed with the following solutions at room temperature: 50 mM Tris-HCl (pH 7.5) for 10 seconds, 2 M urea (dissolved in 50 mM Tris-HCl, pH 7.5) for 10 seconds, twice. Mass spectrometry detection can be carried out subsequently. Preferably, the proteins on the magnetic beads are directly digested with enzymes and then non-labeled quantitative mass spectrometry detection is carried out.
[0054] As shown in the Figure 5 appendix, Western blot was used to detect the content of satellite sequence-related proteins and fusion proteins produced by nucleic acid constructs I and II in the experimental group and the control group in total proteins and proteins enriched by magnetic beads. The experimental group was able to successfully enrich satellite sequence-related protein C (CENP-C).
[0055] The above embodiments only represent several implementation manners of the present invention, and the description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. For those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A specific methylation sequence labeling system for living cells, characterized in that, Comprising nucleic acid construct I and nucleic acid construct II; The nucleic acid construct I comprises at least a eukaryotic cell promoter sequence, a dCas9 protein sequence, a fluorescent protein sequence, and the C-terminal sequence of TurboID protein from the 5'-end to the 3'-end; The nucleic acid construct II comprises at least a eukaryotic cell promoter sequence, a methyl-binding protein MBD domain, a fluorescent protein sequence, the N-terminal sequence of TurboID protein, a murine U6 promoter sequence, and an sgRNA sequence that can be recognized by the dCas9 protein and target a target site from the 5'-end to the 3'-end.
2. The specific methylation sequence labeling system for living cells according to claim 1, characterized in that, The eukaryotic cell promoter sequence in the nucleic acid construct I and the nucleic acid construct II is the SFFV promoter sequence.
3. The live cell specific methylation sequence labeling system according to claim 1, characterized in that, Nuclear localization signals NLS are added before and after the dCas9 protein sequence in the nucleic acid construct I. Among them, the nuclear localization signal sequence before the dCas9 protein sequence is shown as SEQ ID No.1; the nuclear localization signal sequence after the dCas9 protein sequence is shown as SEQ ID No.
2.
4. The specific methylation sequence labeling system for living cells according to claim 1, characterized in that, The fluorescent protein sequence in the nucleic acid construct I is the enhanced green fluorescent protein EGFP sequence. A GS linker sequence shown as SEQ ID No.3 is added before the enhanced green fluorescent protein EGFP sequence; a GS linker sequence and a MycTag tag sequence are added after the enhanced green fluorescent protein EGFP sequence. The sequences of the GS linker sequence and the MycTag tag are shown as SEQ ID No.
4.
5. The specific methylation sequence labeling system for living cells according to claim 1, wherein The C-terminal sequence of TurboID protein in the nucleic acid construct I is shown as SEQ ID No.
5.
6. The live cell specific methylation sequence labeling system according to claim 1, characterized in that, Three tandem nuclear localization signals and a 3xFlag tag, the sequences of which are shown as SEQID No.6, are added before the MBD protein domain sequence in the nucleic acid construct II.
7. The specific methylation sequence labeling system for living cells according to claim 1, characterized in that, The MBD protein domain sequence in the nucleic acid construct II is shown as SEQ ID No.
7.
8. The specific methylation sequence labeling system for living cells according to claim 1, wherein The fluorescent protein sequence in the nucleic acid construct II is the red fluorescent protein mCherry sequence. A GS sequence shown as SEQ ID No.8 is added before the red fluorescent protein mCherry sequence, and a GS sequence shown as SEQ ID No.9 is added after the red fluorescent protein mCherry sequence.
9. The specific methylation sequence labeling system for living cells according to claim 1, characterized in that, The N-terminal sequence of TurboID protein in the nucleic acid construct II is shown as SEQ ID No.
10.
10. A fusion protein expression vector for proximity labeling of specific methylation sequences in living cells, wherein the vector contains the nucleic acid construct I and the nucleic acid construct II in the specific methylation sequence labeling system for living cells according to any one of claims 1-9.