A method for identifying a protospacer adjacent motif
By using a short-chain DNA detection array and a method that combines Cpf1 protein with crRNA, the problem of long time and complex steps in the identification of Cpf1 protein PAM sequences in existing technologies has been solved, and efficient and simple PAM sequence identification and detection has been achieved.
Patent Information
- Application Number
- CN202310173577.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-17
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2043-02-17
AI Technical Summary
Existing methods for identifying the PAM sequence of Cpf1 protein are time-consuming, complex, and lack specificity, making it difficult to accurately identify the PAM sequence.
A short-chain DNA detection array and a ribonucleoprotein complex formed by Cpf1 protein binding to crRNA were used to detect the specific recognition of PAM sequences by Cpf1 by fluorescence signal. The process included synthesizing primer pairs for the detection array, annealing to obtain the short-chain DNA detection array, measuring fluorescence values after incubation, and normalizing the results.
It achieves efficient and convenient PAM sequence identification, clearly demonstrating the recognition preference of Cpf1 protein for different PAM sequences, simplifying the operation process and shortening the detection time.
Smart Images

Figure CN116083536B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of biotechnology, in particular to a method for identifying a protospacer adjacent motif. BACKGROUND
[0002] The CRISPR-Cas system is originally derived from the RNA-guided adaptive immune system of bacteria, which is used to resist nucleic acid components of invading bacteria and archaea. CRISPR refers to clustered regularly interspaced short palindromic repeats, which is derived from a bacteriophage DNA fragment that can infect prokaryotes. It can detect and destroy similar DNA in other bacteriophages that can cause similar infections, so it is essential for prokaryotes to resist bacteriophages. Therefore, this sequence is the immune system of many prokaryotes. Cas protein refers to CRISPR-associated protein, which is a related protein in the CRISPR system. The CRISPR locus is composed of an operon encoding Cas protein and a repeat spacer. The corresponding RNA of the spacer in the CRISPR sequence can be used to guide the recognition and cutting of the DNA strand with specific sequence complementarity. The CRISPR / Cas9 system is simple and efficient, and only needs to synthesize a new RNA to achieve gene editing. Its simple design and easy operation are the greatest advantages, making this technology a convenient and highly adaptable tool for regulating and observing the genome. This allows it to be widely used in biological research and biotechnology, including gene regulation, epigenetic modification, base editing, gene imaging, and other fields, and has shown wide applicability in microorganisms, plants, animals, and even humans. The CRISPR-Cas tool greatly accelerates the pace of scientific research, and the development of Cas-based biotechnology is also rapidly advancing. A number of Cas9-based clinical trials are underway or will be conducted soon. The results of these clinical trials will guide future applications of somatic cell gene editing in vitro and in patients.
[0003] The CRISPR / Cas system is a detection technology that can specifically recognize target sequences and has catalytic amplification signal function, and can directly detect target sequences. However, the current detection method needs to pre-amplify the target nucleic acid, and the detection sensitivity can meet the diagnostic requirements, which increases the operation steps and takes more time. The trans-cleavage activity of the CRISPR / Cas system is random, and it is difficult to realize quantitative and multiplex detection. Most of the existing Cas proteins rely on a short conserved sequence upstream / downstream of the target site to recognize the target DNA, that is, the protospacer adjacent motif (PAM) plays a role, and the selection of the target sequence is limited. As a gene editing technology, the existence of off-target effects will also affect the accuracy of the diagnostic results.
[0004] CRISPR / Cas system has different types, the most widely used is the famous CRISPR / Cas9 system. Because it only needs one protein, Cas9, to screen, stick and then cut the target gene. In addition to Cas9, the most concerned is Cpf1 in class V, renamed as Cas12a. In September 2015, Zhang Feng's team (Zetsche B, Gootenberg JS. Abudayyeh O O, et al. Cpf1 is a single RNA-guided endonuclease of a class 2 CRISPR-Cas system. Cell, 2015, 163(3): 759-771.) added a new member to the CRISPR family-Cpf1, which belongs to CRISPR-Cas system class 2 type V. CRISPR / Cpf1 system belongs to Class2 Type V, which is smaller in size, has the ability of RNA endonuclease and DNA endonuclease, and can recognize PAM sequences rich in thymine (T), which expands the editable range of CRISPR / Cas system. This system not only has good targeted DNA cutting activity, but also shows significant editing characteristics different from CRISPR-Cas9. Internationally, Acidaminococcus Cpf1 (AsCpf1) and Lachnospiraceae Cpf1 (LbCpf1) are widely used, but due to their lack of flexibility in recognizing 5'-TTTN-3' protospacer adjacent motif, their application in gene editing field is limited. Tu M, Lin L, Cheng Y, et al. A 'new lease of life': FnCpf1 possesses DNA cleavage activity for genome editing in human cells. Nucleic Acids Research, 2017, 45(19): 11295-11304. Team proposed that Francisella novicida Cpf1 (FnCpf1) has considerable cutting efficiency in human cells for the first time, and systematically determined the related parameters, and determined that the sequence of PAM is expanded to 5'-KYTV-3'(K refers to T, G; Y refers to C, T; V refers to A, C, G). Therefore, Cpf1 will be used as a "new genome editing tool" together with the Cas9 family to provide more broad development space for scientific research and disease treatment
[0005] At present, the methods for identifying the PAM sequence of Cpf1 protein (referring to crRNA-dependent endonuclease, also known as Cas12a protein, which is a type V enzyme in the classification of CRISPR system) include in vivo library screening and in vitro library screening. The in vivo library screening method identifies the functional PAM by deep sequencing of cells that remain resistant to Cpf1 cleavage; in this method, the Cpf1 concentration of different cells is different and the cleavage time is uncontrollable, which makes it difficult to strictly identify the PAM sequence, and moreover, the process of constructing the PAM library plasmid is complicated and complex. The in vitro library screening method generally connects a 7-8 bp random sequence to the same spacer sequence, and identifies the PAM preference by in vitro cleavage experiment and deep sequencing; this identification method can clearly show the PAM type with obvious proportion of single type of base, but when there are two or more balanced proportions in two base positions of the PAM sequence, it is difficult to determine whether the editing activity is higher for a certain PAM or several PAMs. The above PAM detection methods have long test time, complex steps and insufficient specificity. Therefore, it is necessary to develop an efficient, accurate and convenient PAM identification method. SUMMARY
[0006] The problem to be solved by the present application is to overcome the defects of the existing PAM sequence identification method, such as long test time, complex steps and insufficient specificity, and to provide a high-throughput Cpf1 protein PAM sequence identification method.
[0007] The present application provides a method for identifying a protospacer adjacent motif, comprising the following steps:
[0008] (1) Synthesizing a detection array primer pair:
[0009] The sequence of the detection array primer pair contains a continuous 4bp base N, wherein N is A, C, G or T; and each determined sequence is a detection point of the array;
[0010] (2) Synthesizing a short-chain DNA detection array:
[0011] The detection array primer pair synthesized in step (1) is annealed to obtain a short-chain DNA detection array;
[0012] (3) PAM identification:
[0013] The short-chain DNA detection array obtained in step (2), Cpf1 protein, single-stranded DNA reporter gene and crRNA constitute an identification system, and the fluorescence value of the identification system is measured after incubation reaction.
[0014] As a preferred, the detection array primer pair in step (1) is at least one of the following:
[0015] (1) EMX1-F: ggggcctcctgagNNNNtcatctgtgcccctccctccctggcccaggtgaaggtg,
[0016] EMX1-R: caccttcacctgggccagggagggaggggcacagatgaNNNNctcaggaggcccc;
[0017] (2) MerS_F: ccattagtctctcNNNNgacatatggaaaacgaactatgttaccctttgtccaag,
[0018] MerS_R: cttggacaaagggtaacatagttcgttttccatatgtcNNNNgagagactaatgg;
[0019] (3) DNMT1_F: cagtacgttaatgNNNNctgatggtccatgtctgttactcgcctgtcaagtggcg,
[0020] DNMT1_R: cgccacttgacaggcgagtaacagacatggaccatcagNNNNcattaacgtactg;
[0021] (4) eGFP-1_F: cttgtggccgNNNNcgtcgccgtccagctcgaccaggatgggcaccaccccggtg,
[0022] eGFP-1_R: caccggggtggtgcccatcctggtcgagctggacggcgacgNNNNcggccacaag;
[0023] (5) eGFP-3_F: cgttggggtcNNNNctcagggcggactgggtgctcaggtagtggttgtcgggcag,
[0024] eGFP-3_R: ctgcccgacaaccactacctgagcacccagtccgccctgagNNNNgaccccaacg;
[0025] (6) FANCF_F: ctacttccgcNNNNaccttggagacggcgactctctgcgtactgattggaacatc,
[0026] FANCF_R: gatgttccaatcagtacgcagagagtcgccgtctccaaggtNNNNgcggaagtag.
[0027] Specifically, the annealing system in step (2) is 20 μL, specifically: 10 μL of forward primer and 10 μL of reverse primer, and the concentrations of the forward primer and the reverse primer are both 100 μM. The annealing condition is to keep the temperature at 95°C for 2 minutes.
[0028] The identification system in step (3) is specifically: 200 ng of Cpf1, 25 pM of single-stranded DNA reporter, 1 μM of crRNA, 1 μL of 8.5 nM short-chain DNA detection array, 2 μL of 10×NEBuffer r3.1, and the final volume is 20 μL.
[0029] In the detection system, the crRNA combines with the Cas12a to form a ribonucleoprotein complex, and after the Cas12a-crRNA recognizes the target sequence, the nuclease activity of the Cas12a is activated to cut the double-stranded DNA in the reaction system. At the same time, the transcleavage activity of Cas12a is activated to cut the ssDNA in the reaction system, and the single-stranded DNA reporter (ssDNA-FQ) will be cleaved.
[0030] Specifically, the single-stranded DNA reporter in step (3) has a fluorescent group at the 5' end and a fluorescent quenching group at the 3' end. The fluorescent group is 6-FAM, TET, CY3, CY5 or ROX, the fluorescent quenching group is BHQ1, BHQ2 or BHQ3, and the sequence of the single-stranded DNA is TTTATTT.
[0031] The crRNA in step (3) is at least one of the following:
[0032] (a) EMX1-crRNA1: UAAUUUCUACUAAGUGUAGAUUCAUCUGUGCCCCUCCCUCCCUG;
[0033] (b) Mers-crRNA1: UAAUUUCUACUAAGUGUAGAUGACAUAUGGAAAACGAACUAUGU;
[0034] (c) DNMT1-crRNA1: UAAUUUCUACUAAGUGUAGAUCUGAUGGUCCAUGUCUGUUACUC;
[0035] (d) eGFP-crRNA1: UAAUUUCUACUAAGUGUAGAUCGUCGCCGUCCAGCUCGACCAGG;
[0036] (e) eGFP-crRNA3: UAAUUUCUACUAAGUGUAGAUCUCAGGGCGGACUGGGUGCUCAG;
[0037] (f) FANCF-crRNA1: UAAUUUCUACUAAGUGUAGAUACCUUGGAGACGGCGACUCUCUG.
[0038] The method for identifying the pre-interval sequence adjacent motif further comprises averaging and normalizing the short-chain DNA detection array measurement in step (3). Normalization is to limit the processed data to a certain range you need (through an algorithm). First, normalization is for the convenience of subsequent data processing, and second, to ensure program convergence. The specific role of normalization is to summarize the statistical distribution of samples. Normalization is a statistical probability distribution between 0 and 1, and normalization is a statistical coordinate distribution in a certain interval. Normalization has the meaning of unity, uniformity and unity.
[0039] There are two forms of normalization method, one is to change the number to a decimal number between (0, 1), and the other is to change the dimensional expression to a dimensionless expression. It is mainly put forward for the convenience of data processing, which is more convenient and fast to map the data to 0-1 range, and should be classified into the category of digital signal processing.
[0040] The application also provides a kit for identifying a pre-interval sequence adjacent motif, the kit comprising 200 ng Cpf1, 25 pM single-stranded DNA reporter gene, 1 μM crRNA, 1 μL of 8.5 nM short-chain DNA detection array, 2 μL of 10x NEBuffer r3.1, and a final volume of 20 μL.
[0041] The short-chain DNA detection array is obtained by annealing at least one detection array primer pair:
[0042] (1) EMX1-F: ggggcctcctgagNNNNtcatctgtgcccctccctccctggcccaggtgaaggtg,
[0043] EMX1-R: caccttcacctgggccagggagggaggggcacagatgaNNNNctcaggaggcccc;
[0044] (2) MerS_F: ccattagtctctcNNNNgacatatggaaaacgaactatgttaccctttgtccaag,
[0045] MerS_R: cttggacaaagggtaacatagttcgttttccatatgtcNNNNgagagactaatgg;
[0046] (3) DNMT1_F: cagtacgttaatgNNNNctgatggtccatgtctgttactcgcctgtcaagtggcg,
[0047] DNMT1_R: cgccacttgacaggcgagtaacagacatggaccatcagNNNNcattaacgtactg;
[0048] (4) eGFP-1_F: cttgtggccgNNNNcgtcgccgtccagctcgaccaggatgggcaccaccccggtg,
[0049] eGFP-1_R: caccggggtggtgcccatcctggtcgagctggacggcgacgNNNNcggccacaag;
[0050] (5) eGFP-3_F: cgttggggtcNNNNctcagggcggactgggtgctcaggtagtggttgtcgggcag,
[0051] eGFP-3_R: ctgcccgacaaccactacctgagcacccagtccgccctgagNNNNgaccccaacg;
[0052] (6) FANCF_F: ctacttccgcNNNNaccttggagacggcgactctctgcgtactgattggaacatc,
[0053] FANCF_R: gatgttccaatcagtacgcagagagtcgccgtctccaaggtNNNNgcggaagtag;
[0054] the crRNA is at least one of:
[0055] (a) EMX1-crRNA1: UAAUUUCUACUAAGUGUAGAUUCAUCUGUGCCCCUCCCUCCCUG;
[0056] (b) Mers-crRNA1: UAAUUUCUACUAAGUGUAGAUGACAUAUGGAAAACGAACUAUGU;
[0057] (c) DNMT1-crRNA1: UAAUUUCUACUAAGUGUAGAUCUGAUGGUCCAUGUCUGUUACUC;
[0058] (d) eGFP-crRNA1: UAAUUUCUACUAAGUGUAGAUCGUCGCCGUCCAGCUCGACCAGG;
[0059] (e) eGFP-crRNA3: UAAUUUCUACUAAGUGUAGAUCUCAGGGCGGACUGGGUGCUCAG;
[0060] (f) FANCF-crRNA1: UAAUUUCUACUAAGUGUAGAUACCUUGGAGACGGCGACUCUCUG.
[0061] In the present application, when the target DNA, crRNA and Cpf1 form a ternary complex, the complex will cut other single-stranded DNA molecules in the system. By designing and synthesizing detection arrays with 256 PAM sequences targeting six sites of EMX1, DNMT1, MerS, eGFP-1, eGFP-3 and FANCF and their corresponding crRNA; adding the Cpf1 to be tested, crRNA and short-chain target DNA to the identification system; when Cpf1 can recognize short-chain DNA with specific PAM sequences, Cpf1 forms a ternary complex with crRNA and short-chain DNA, and at the same time the complex exercises its trans-cleavage activity and cuts single-stranded DNA labeled with a fluorescent signal (with a luminescent group and a quenching group at both ends, which can emit light after being cut), thereby emitting fluorescence; the specificity of Cas protein recognizing PAM sequence is determined by detecting the fluorescence intensity of the reaction system.
[0062] In the present application, the nucleic acid probe can also be referred to as a fluorescence-quenching single-stranded DNA reporter system, which generally contains a fluorescent group and a quenching group. In the intact state, the fluorescence is quenched by the quenching group, and after being cut during the reaction, the fluorescence is emitted.
[0063] On the basis of common sense in the art, the above-mentioned preferred conditions can be combined arbitrarily, i.e. to obtain each preferred example of the present application.
[0064] The reagents and raw materials used in the present application are commercially available.
[0065] The present application has the following beneficial effects:
[0066] The present application utilizes the nucleic acid cleavage activity of CRISP-Cas protein, and detects the specificity of CRISP-Cas protein to different PAM sequences in vitro using a DNA array with 256 PAM sequences, which can clearly show the recognition preference of CRISP-Cas protein to each PAM sequence. At the same time, the present application is simple in steps, easy to operate, does not require sequencing, and has short detection time, and can be used as a general detection method for CRISP-Cas protein. BRIEF DESCRIPTION OF DRAWINGS
[0067] Figure 1 Schematic diagram of short-chain DNA array for rapid detection of Cas protein PAM preference.
[0068] Figure 2 Schematic diagram of time efficiency comparison between short-chain DNA array rapid detection and intracellular sequencing method.
[0069] Figure 3 Experimental verification diagram for short-chain DNA array detection of AsCpf1 PAM preference.
[0070] Figure 4 Normalized result diagram for short-chain DNA array detection of AsCpf1 PAM preference.
[0071] Figure 5 Experimental verification diagram for short-chain DNA array detection of LbCpf1 PAM preference.
[0072] Figure 6 Normalized result diagram for short-chain DNA array detection of LbCpf1 PAM preference. DETAILED DESCRIPTION
[0073] The present application will be further described below in conjunction with specific examples. It should be understood that these examples are only used to illustrate the present application and not used to limit the scope of the present application. In addition, it should be understood that those skilled in the art can make various modifications or changes to the present application after reading the content taught by the present application, and these equivalent forms also fall within the scope of the appended claims of the present application.
[0074] The experimental methods in the following examples are not specified, which are selected according to conventional methods and conditions, or according to the instructions of the goods.
[0075] In the present application, the conventional reagents such as Tris-Base, BSA, NaCl, Tris-HCl, MgSO4 and glycerol are purchased from Thermo Fisher; the nucleic acid fragments, ssDNA probes and RNA synthesis are synthesized by Beijing Qikexing Technology Co., Ltd.
[0076] The general technical schematic diagram of the present application is shown in Figure 1 The present application comprises the following three parts; the design and synthesis of primer pairs, the annealing synthesis of short-chain DNA array and PAM detection three steps. Figure 2 As shown in the figure, the short-chain DNA array rapid detection has obvious advantages in time efficiency compared with the cell sequencing method, which is shortened from three and a half months to more than a week, and the measurement can be directly read immediately after completion without complex sequencing and decoding.
[0077] The present application provides a method for identifying a protospacer adjacent motif, comprising the following steps:
[0078] (1) Synthesis of detection array primer pairs:
[0079] The sequence of the detection array primer pair contains a continuous 4bp base N, wherein N is A, C, G or T; and each determined sequence is a detection point of the array;
[0080] (2) Synthesis of short-chain DNA detection array:
[0081] The detection array primer pairs synthesized in step (1) are annealed to obtain a short-chain DNA detection array;
[0082] (3) PAM identification:
[0083] The short-chain DNA detection array obtained in step (2), Cpf1 protein, single-stranded DNA reporter gene, crRNA and reaction solution constitute an identification system, and the fluorescence value of the identification system is measured after incubation reaction.
[0084] The measurement results of the short-chain DNA detection array in step (3) are averaged and normalized. There are two forms of normalization methods, one is to change the number to a decimal number between (0, 1), and the other is to change the dimensional expression to a dimensionless expression. In the example of the embodiment of the present application, the normalization method of changing the number to a decimal number between (0, 1) is preferred.
[0085] In the detection system, crRNA combines with Cas12a to form a ribonucleoprotein complex, Cas12a-crRNA recognizes the target sequence and activates the nuclease activity of Cas12a, cuts the double-stranded DNA in the reaction system, and at the same time activates the transcleavage activity of Cas12a, cuts the ssDNA in the reaction system, and the single-stranded DNA reporter gene (ssDNA-FQ) will be cleaved.
[0086] The single-stranded DNA reporter has a fluorescent group at the 5' end and a fluorescent quenching group at the 3' end. The fluorescent group can be 6-FAM, TET, CY3, CY5 or ROX, but is not limited thereto; the fluorescent quenching group can be BHQ1, BHQ2 or BHQ3, but is not limited thereto. The single-stranded DNA has a light-emitting group and a quenching group at both ends, and the light-emitting group can emit light after being cut, thereby emitting fluorescence.
[0087] The specificity of the Cas protein recognizing the PAM sequence is determined by detecting the fluorescence intensity of the reaction system.
[0088] The application will be further described below in combination with specific examples.
[0089] Example 1: Synthesis of short-chain DNA detection array targeting EMX1, DNMT1, MerS, eGFP-1, eGFP-3 and FANCF sites
[0090] Six target sites of EMX1, DNMT1, MerS, eGFP and FANCF genes were used as detection targets to design and synthesize short-chain DNA arrays. The length of the forward and reverse primers was 55 bases (synthesized by Beijing Qikexing Biological Technology Co., Ltd.), and the synthesis was completed at a concentration of 100 μM. The primer sequences are shown in Table 1.
[0091] Short-chain DNA primer annealing: the annealing system was 20 μL, 10 μL forward primer and 10 μL reverse primer (both at a concentration of 100 μM).
[0092] Annealing: the short-chain DNA annealing system was incubated at 95°C for 2 minutes, and then naturally cooled to room temperature.
[0093] Table 1: DNA detection array primers
[0094] Primer name Primer sequence (5'→3') (NNNN is 256 PAM combinations) EMX1_F ggggcctcctgagNNNNtcatctgtgcccctccctccctggcccaggtgaaggtg EMX1_R caccttcacctgggccagggagggaggggcacagatgaNNNNctcaggaggcccc MerS_F ccattagtctctcNNNNgacatatggaaaacgaactatgttaccctttgtccaag MerS_R cttggacaaagggtaacatagttcgttttccatatgtcNNNNgagagactaatgg DNMT1_F cagtacgttaatgNNNNctgatggtccatgtctgttactcgcctgtcaagtggcg DNMT1_R cgccacttgacaggcgagtaacagacatggaccatcagNNNNcattaacgtactg eGFP-1_F cttgtggccgNNNNcgtcgccgtccagctcgaccaggatgggcaccaccccggtg eGFP-1_R caccggggtggtgcccatcctggtcgagctggacggcgacgNNNNcggccacaag eGFP-3_F cgttggggtcNNNNctcagggcggactgggtgctcaggtagtggttgtcgggcag eGFP-3_R ctgcccgacaaccactacctgagcacccagtccgccctgagNNNNgaccccaacg FANCF_F ctacttccgcNNNNaccttggagacggcgactctctgcgtactgattggaacatc FANCF_R gatgttccaatcagtacgcagagagtcgccgtctccaaggtNNNNgcggaagtag
[0095] Example 2: PAM identification of AsCpf1
[0096] The AsCpf1 gene in the example (the gene sequence is shown as SEQ ID NO. 19) was codon-optimized, then cloned into the pET28a plasmid (this step was completed by Nanjing Kings River Co., Ltd.), expressed in E. coli, and used for PAM identification experiments after purification.
[0097] PAM identification.
[0098] The crRNA consists of a 21 nt region (UAAUUUCUACUAAGUGUAGAU) that interacts with Cas12a and a 23 nt spacer sequence complementary to the target DNA. The crRNA sequence is shown in Table 2. A single-stranded DNA reporter (ssDNA-FQ) labeled with FAM and BHQ1 was synthesized by Beijing Genesee Biotechnology Co., Ltd. The annealed short DNA array was diluted to 8.5 nM.
[0099] The detection system was 20 μL, and the specific components were as follows: 200 ng of purified AsCpf1, 25 pM of ssDNA-FQ, 1 μM of crRNA, 1 μL of short DNA, 2 μL of 10x NEBuffer r3.1, and a final volume of 20 μL. The reaction was incubated at 37°C. In the detection system, the crRNA binds to the Cas12a to form a ribonucleoprotein complex. After the Cas12a-crRNA recognizes the target sequence, the nuclease activity of the Cas12a is activated, and the double-stranded DNA in the reaction system is cleaved. At the same time, the transcleavage activity of the Cas12a is activated, and the ssDNA in the reaction system is cleaved. The ssDNA-FQ is cleaved. A full-wavelength microplate reader was used to measure the fluorescence value of the reaction system, with an excitation wavelength of 485 nm and an emission wavelength of 520 nm. The signal was recorded every 30 seconds for 30 minutes. In each batch of reaction, the detection reaction of the short DNA with the PAM sequence TTTG was added. The fluorescence value of this reaction was used to correct the fluorescence value of the batch reaction, that is, the fluorescence value of the batch reaction was divided by the fluorescence value of the TTTG reaction. Then the corrected value was normalized (that is, the series of values were changed to decimals between 0 and 1).
[0100] The PAM preference of AsCpf1 was detected using six short DNA arrays of EMX1, DNMT1, MerS, eGFP-1, eGFP-3, and FANCF, and the results are shown in Figure 3 To avoid the influence of target specificity, the average value of the detection results of the six short DNA arrays of each PAM sequence was taken, and the normalized processing (that is, the series of values were changed to decimals between 0 and 1) was performed, and the results are shown in Figure 4
[0101] Table 2: crRNA sequence
[0102]
[0103] Conclusion: Figure 3 It is shown that the preference of AsCpf1 for different PAM sequences in the short DNA array shows obvious differences, confirming that the short DNA array can be used to detect the preference of the PAM sequence recognized by Cpf1; Figure 4 The PAM preference fingerprint of Cpf1 protein obtained from the average of the detection results of the six DNA arrays can avoid the influence of target specificity on the detection results.
[0104] Example 3: PAM identification of LbCpf1
[0105] The LbCpf1 gene in the example (the gene sequence is shown as SEQ ID NO. 20) was codon-optimized, cloned into the pET28a plasmid (this step was completed by Nanjing Kings River Company), expressed in E. coli, and used for PAM identification experiments after purification.
[0106] PAM identification.
[0107] The crRNA is composed of a 21 nt region (UAAUUUCUACUAAGUGUAGAU) interacting with Cas12a and a 23 nt spacer sequence complementary to the target DNA. The crRNA sequence is shown in Table 2. A single-stranded DNA reporter (ssDNA-FQ) labeled with FAM and BHQ1 was synthesized by Beijing Qikexin Biotechnology Co., Ltd. The annealed short-chain DNA array was diluted to 8.5 nM.
[0108] The detection system was 20 μL, and the specific components were as follows: 200 ng of purified LbCpf1, 25 pM of ssDNA-FQ, 1 μM of crRNA, 1 μL of short-chain DNA, 2 μL of 10×NEBuffer r3.1, and the final volume was 20 μL. The reaction was incubated at 37°C. In the detection system, crRNA binds to Cas12a to form a ribonucleoprotein complex. Cas12a-crRNA recognizes the target sequence and activates the nuclease activity of Cas12a, which cuts the double-stranded DNA in the reaction system. At the same time, the transcleavage activity of Cas12a is activated, which cuts the ssDNA in the reaction system. ssDNA-FQ is cleaved. A full-wavelength microplate reader was used to measure the fluorescence value of the reaction system, with an excitation wavelength of 485 nm and an emission wavelength of 520 nm. The signal was recorded every 30 seconds for 30 minutes. In each batch of reaction, the detection reaction of the short-chain DNA with the PAM sequence TTTG was added. The fluorescence value of this reaction was used to correct the fluorescence value of the batch reaction, i.e. the fluorescence value of the batch reaction was divided by the fluorescence value of the TTTG reaction. Then the corrected value was normalized (i.e. the series of values were converted to decimals between 0 and 1).
[0109] The PAM preference of LbCpf1 was detected using EMX1, DNMT1, MerS, eGFP-1, eGFP-3 and FANCF six short-chain DNA arrays, and the results are shown in Figure 5The results of six short-chain DNA array detection of each PAM sequence were averaged and normalized (i.e. the series of values were changed into decimals between 0 and 1), and the results are shown in Table 2. Figure 6
[0110] Conclusion: Figure 5 It is shown that the preference of LbCpf1 for different PAM sequences in short-chain DNA array also shows obvious differences, confirming that short-chain DNA array can be used to detect the preference of PAM sequence recognized by Cpf1; Figure 6 It is shown that the PAM preference fingerprint of Cpf1 protein obtained from the average of six DNA array detection results again proves that this method can avoid the influence of target specificity on detection results; the detection results show that LbCpf1 recognizes the 5'-TTTN-3' PAM of double-stranded DNA.
Claims
1. A method of identifying a pre-int intergenic sequence proximal motif, characterized by, Comprising the following steps: (1) Synthesis of detection array primer pairs: The detection array primer pairs contain 4 consecutive base N in the sequence, wherein N is A, C, G or T; each determined sequence is a detection point of the array; The detection array primer pairs are at least one of the following: (a) EMX1-F: ggggcctcctgagNNNNtcatctgtgcccctccctccctggcccaggtgaaggtg, EMX1-R: caccttcacctgggccagggagggaggggcacagatgaNNNNctcaggaggcccc; (b) MerS_F: ccattagtctctcNNNNgacatatggaaaacgaactatgttaccctttgtccaag, MerS_R: cttggacaaagggtaacatagttcgttttccatatgtcNNNNgagagactaatgg; (c) DNMT1_F: cagtacgttaatgNNNNctgatggtccatgtctgttactcgcctgtcaagtggcg, DNMT1_R: cgccacttgacaggcgagtaacagacatggaccatcagNNNNcattaacgtactg; (d) eGFP-1_F: cttgtggccgNNNNcgtcgccgtccagctcgaccaggatgggcaccaccccggtg, eGFP-1_R: caccggggtggtgcccatcctggtcgagctggacggcgacgNNNNcggccacaag; (e) eGFP-3_F: cgttggggtcNNNNctcagggcggactgggtgctcaggtagtggttgtcgggcag, eGFP-3_R: ctgcccgacaaccactacctgagcacccagtccgccctgagNNNNgaccccaacg; (f) FANCF_F: ctacttccgcNNNNaccttggagacggcgactctctgcgtactgattggaacatc, FANCF_R: gatgttccaatcagtacgcagagagtcgccgtctccaaggtNNNNgcggaagtag; (2) Synthesis of short-chain DNA detection array: The detection array primer pairs synthesized in step (1) are annealed to obtain a short-chain DNA detection array; (3) PAM identification: The short-chain DNA detection array obtained in step (2), Cpf1 protein, single-stranded DNA reporter gene and crRNA form an identification system, and the fluorescence value of the identification system is measured after incubation reaction; The step (3) further comprises averaging the measurement results of the short-chain DNA detection array and normalizing the measurement results.
2. The method for identifying a pre-int intergenic sequence proximal motif according to claim 1, wherein, The annealing system in the step (2) is 10 μL of the forward primer and 10 μL of the reverse primer, and the concentration of the forward primer and the reverse primer is 100 μM.
3. The method for identifying a pre-int intergenic sequence proximal motif according to claim 2, wherein, The annealing condition is 2 minutes of constant temperature annealing at 95 °C.
4. The method for identifying a pre-int intergenic sequence proximal motif according to claim 1, wherein, The identification system in the step (3) is 200 ng of Cpf1, 25 pM of the single-stranded DNA reporter, 1 μM of crRNA, 1 μL of 8.5 nM of the short-chain DNA detection array, and 2 μL of 10× NEBuffer r3.1, and the final volume is 20 μL.
5. The method of identifying a pre-int intergenic sequence proximal motif according to claim 1, wherein, The single-stranded DNA reporter in the step (3) has a fluorescent group at the 5' end and a fluorescent quenching group at the 3' end.
6. The method for identifying a pre-int intergenic sequence proximal motif according to claim 5, wherein, The fluorescent group is 6-FAM, TET, CY3, CY5 or ROX, the fluorescent quenching group is BHQ1, BHQ2 or BHQ3, and the sequence of the single-stranded DNA is TTTATTT.
7. The method of identifying a pre-int intergenic sequence proximal motif according to claim 1, wherein, The crRNA in the step (3) is at least one of the following: (a) EMX1-crRNA1: UAAUUUCUACUAAGUGUAGAUUCAUCUGUGCCCCUCCCUCCCUG; (b) Mers-crRNA1: UAAUUUCUACUAAGUGUAGAUGACAUAUGGAAAACGAACUAUGU; (c) DNMT1-crRNA1: UAAUUUCUACUAAGUGUAGAUCUGAUGGUCCAUGUCUGUUACUC; (d) eGFP-crRNA1: UAAUUUCUACUAAGUGUAGAUCGUCGCCGUCCAGCUCGACCAGG; (e) eGFP-crRNA3: UAAUUUCUACUAAGUGUAGAUCUCAGGGCGGACUGGGUGCUCAG; (f) FANCF-crRNA1: UAAUUUCUACUAAGUGUAGAUACCUUGGAGACGGCGACUCUCUG.
8. A kit for the identification of a pre-int intronic sequence adjacent motif, characterized in that, The kit comprises 200 ng of Cpf1, 25 pM of the single-stranded DNA reporter, 1 μM of crRNA, 1 μL of 8.5 nM of the short-chain DNA detection array, and 2 μL of 10× NEBuffer r3.1, and the final volume is 20 μL. The short-chain DNA detection array is obtained by annealing at least one of the following detection array primer pairs: (1) EMX1-F: ggggcctcctgagNNNNtcatctgtgcccctccctccctggcccaggtgaaggtg, EMX1-R: caccttcacctgggccagggagggaggggcacagatgaNNNNctcaggaggcccc; (2) MerS_F: ccattagtctctcNNNNgacatatggaaaacgaactatgttaccctttgtccaag, MerS_R: cttggacaaagggtaacatagttcgttttccatatgtcNNNNgagagactaatgg; (3) DNMT1_F: cagtacgttaatgNNNNctgatggtccatgtctgttactcgcctgtcaagtggcg, DNMT1_R: cgccacttgacaggcgagtaacagacatggaccatcagNNNNcattaacgtactg; (4) eGFP-1_F: cttgtggccgNNNNcgtcgccgtccagctcgaccaggatgggcaccaccccggtg, eGFP-1_R: caccggggtggtgcccatcctggtcgagctggacggcgacgNNNNcggccacaag; (5) eGFP-3_F: cgttggggtcNNNNctcagggcggactgggtgctcaggtagtggttgtcgggcag, eGFP-3_R: ctgcccgacaaccactacctgagcacccagtccgccctgagNNNNgaccccaacg; (6) FANCF_F: ctacttccgcNNNNaccttggagacggcgactctctgcgtactgattggaacatc, FANCF_R: gatgttccaatcagtacgcagagagtcgccgtctccaaggtNNNNgcggaagtag; the crRNA is at least one of: (a) EMX1-crRNA1: UAAUUUCUACUAAGUGUAGAUUCAUCUGUGCCCCUCCCUCCCUG; (b) Mers-crRNA1: UAAUUUCUACUAAGUGUAGAUGACAUAUGGAAAACGAACUAUGU; (c) DNMT1-crRNA1: UAAUUUCUACUAAGUGUAGAUCUGAUGGUCCAUGUCUGUUACUC; (d) eGFP-crRNA1: UAAUUUCUACUAAGUGUAGAUCGUCGCCGUCCAGCUCGACCAGG; (e) eGFP-crRNA3: UAAUUUCUACUAAGUGUAGAUCUCAGGGCGGACUGGGUGCUCAG; (f) FANCF-crRNA1: UAAUUUCUACUAAGUGUAGAUACCUUGGAGACGGCGACUCUCUG. (f) FANCF-crRNA1: UAAUUUCUACUAAGUGUAGAUACCUUGGAGACGGCGACUCUCUG. (f) FAN
Citation Information
Patent Citations
Class II V-type CRISPR protein Lb2Cas12a and application thereof in gene editing
CN111394337A
Detection method based on Cas protein
CN113174433A
Cited By
Cas protease PAM analysis method based on local geometric structure and machine learning
CN121483365A
Cas protein pam analysis method based on local geometry and machine learning
CN121483365B