A method for detecting sites of DNA double-strand breaks generated by CRISPR-Cas

CN117025670BActive Publication Date: 2026-09-04GENEWIZ INC SZ
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310236984.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-13
Publication Date
2026-09-04
Estimated Expiration
2043-03-13

AI Technical Summary

Technical Problem

但其仍存在以下几个缺点:第一,dsODN本身没有靶向性,除了会插入CRISPR造成的DNA双链断裂处,也会插入基因组上自发的、随机的断裂损伤处,该方法无法区分两种插入,导致检测到的一部分脱靶位点实际上为假阳性;第二,dsODN借助NHEJ的修复机制插入到DNA双链断裂处的效率有限,这就导致靶位点和高频发生的脱靶位点比较容易被检测到,而一些低频发生的脱靶位点由于插入dsODN的几率太低,而不容易被检测到

Benefits of technology

[0057](1)本发明设计一种全新的CRISPR-Cas产生的DNA双链断裂位点的检测方法,利用ScFv-GCN4-mSA融合蛋白和Cas-GCN4融合蛋白系统,显著提高了Cas蛋白切割DNA双链时局部的dsODN浓度,进而提高了检测低频脱靶位点的灵敏度以及特异性,且适用性广,能够对本领域通用的CRISPR-Cas系统进行检测;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117025670B_ABST
    Figure CN117025670B_ABST
Patent Text Reader

Abstract

The application discloses a method for detecting DNA double-strand break sites generated by CRISPR-Cas. The method comprises the following steps: transfecting a Cas expression vector, a gRNA expression vector and a dsODN into cells, culturing and extracting cell genomes; carrying out fragmentation treatment, end repair treatment and universal adapter treatment on the genomes in sequence, carrying out amplification reaction on the dsODN sequence and the universal adapter sequence, and collecting amplification products; sequencing the sequencing library, and finding the sequence at the joint of the dsODN and the genome through data analysis. The application uses ScFv-GCN4-mSA fusion protein and Cas-GCN4 fusion protein systems and further designs dsODN, thereby significantly improving the sensitivity and specificity of detecting low-frequency off-target sites, and multiple CRISPR-Cas systems can be simultaneously detected and analyzed, so that the labor cost is saved and the detection throughput is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of genetic engineering technology and relates to a method for detecting DNA double-strand break sites generated by CRISPR-Cas. Background Technology

[0002] CRISPR technology is currently the most widely used gene-editing technology in the field of biology. Compared to previous generations of gene-editing technologies, such as meganucleases, zinc finger nucleases (ZFNs), and transcription activator effector-like nucleases (TALENs), CRISPR's biggest advantage is that it uses base complementarity of guide RNA (gRNA) to target different DNA double strands without relying on changes in protein structure. However, CRISPR technology still has the drawback of off-target effects. The gRNA does not need to be perfectly complementary to the target site's base sequence to cleave the DNA double strand, leading to the appearance of off-target sites.

[0003] CRISPR technology is gradually being applied to clinical treatment. CRISPR can directly alter the genome sequence of cells in the human body. Once the genome sequence of cells is altered, it remains stable. Any side effects caused by CRISPR may accompany the treated individual throughout their life, making CRISPR safety a key focus in clinical trials. Off-target effects are one of the main sources of risk. If off-target editing occurs at unknown sites, it can lead to the inactivation of other normally functioning genes, or even the activation and expression of oncogenes, causing cancer. To reduce the off-target risks of CRISPR technology, one approach is to continuously optimize the technology to reduce the probability of systemic off-target effects. Another key focus for researchers is selecting gRNA sequences with lower off-target risks within the desired editing sequence range. Therefore, identifying and evaluating the off-target sites and probabilities of different gRNA sequences is one of the urgent needs for the development of CRISPR gene therapy.

[0004] GUIDE-seq is a method for non-discriminatory detection of CRISPR off-target sites at the whole-genome level within cells. It utilizes the cellular DNA repair mechanism of non-homologous endjoining (NHEJ) to insert a double-stranded oligodeoxynucleotide sequence (dsODN) into the DNA double-strand break caused by CRISPR. This is equivalent to adding a marker to the DNA double-strand break caused by CRISPR. Subsequently, by extracting genomic DNA from the cell, designing specific primers for dsODN, and performing multiple rounds of PCR amplification and high-throughput sequencing, the sequence information at the junction of dsODN and DNA double-strand break can be obtained, including the location information of CRISPR cleavage target sites and off-target sites. However, it still has the following drawbacks: First, dsODN itself is not targeted. In addition to inserting into DNA double-strand breaks caused by CRISPR, it can also insert into spontaneous and random breaks in the genome. This method cannot distinguish between the two types of insertion, resulting in some off-target sites being false positives. Second, the efficiency of dsODN inserting into DNA double-strand breaks using the NHEJ repair mechanism is limited. This makes it easier to detect target sites and frequently occurring off-target sites, while some low-frequency off-target sites are not easily detected because the probability of insertion into dsODN is too low.

[0005] In conclusion, there is an urgent need to develop new methods to improve the sensitivity and specificity of current off-target detection. Summary of the Invention

[0006] To address the shortcomings of existing technologies and practical needs, this invention provides a method for detecting DNA double-strand breaks generated by CRISPR-Cas. On the one hand, this method increases the probability of detecting low-frequency off-target sites, and on the other hand, it reduces the probability that spontaneous DNA breaks in the genome are mistaken for CRISPR off-target sites.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] In a first aspect, the present invention provides a method for detecting DNA double-strand break sites generated by CRISPR-Cas, the method comprising:

[0009] The Cas expression vector, gRNA expression vector, and dsODN were transfected into cells, cultured, and the cell genome was extracted. The Cas expression vector contained a fusion coding sequence of Cas protein and GCN4 protein, the gRNA expression vector contained a fusion coding sequence of ScFv-GCN4 antibody and mSA protein as well as a gRNA coding sequence, and the dsODN contained biotinylation modification.

[0010] The genome was sequentially fragmented, end-repaired, and ligated with universal adapters. Amplification reactions were performed on the dsODN sequence and the universal adapter sequence, and the amplification products were collected to form the sequencing library.

[0011] The sequencing library was sequenced, and the junction sequence between dsODN and the genome was found through data analysis. This sequence is the DNA double-strand break site generated by CRISPR-Cas in the cell.

[0012] In this invention, the schematic diagram is as follows: Figure 1 As shown, Cas expression vector, gRNA expression vector, and dsODN were simultaneously transfected into cells. The gRNA expression vector expressed the ScFv-GCN4-mSA fusion protein, which can bind to a biotin-modified dsODN. The Cas expression vector expressed the Cas-GCN4 fusion protein, in which GCN4 specifically binds to ScFv-GCN4. GCN4 is a short peptide, and the Cas-GCN4 fusion protein can contain multiple GCN4s (e.g., 10). Multiple GCN4s can recruit multiple ScFv-GCN4-mSAs, i.e., Cas... The protein can recruit multiple dsODNs, significantly increasing the local dsODN concentration during Cas protein cleavage of DNA double strands compared to the simple Cas-mSA fusion protein. On the one hand, the higher local dsODN concentration increases the probability of dsODN insertion into the genome via NHEJ, improving the sensitivity for detecting low-frequency off-target sites. On the other hand, compared to dsODN enriched at the Cas protein cleavage sites, the dsODN at spontaneous breaks in the genome is not increased, resulting in a relatively lower probability of insertion into spontaneous breaks, which is equivalent to a reduction in false positives.

[0013] It is understood that the method of the present invention can be applied to any CRISPR-Cas system in the art, without being limited to specific Cas proteins and gRNAs.

[0014] Optionally, the Cas protein includes any one of spCas9, AsCas12a, or LbCas12a.

[0015] Optionally, the gRNA includes any one of spCas9-sgRNA, AsCas12a-crRNA, or LbCas12a-crRNA.

[0016] Preferably, the number of GCN4 proteins is 8 to 15, including but not limited to 9, 10, 11, 12, 13 or 14.

[0017] In this invention, controlling the amount of GCN4 in the Cas-GCN4 fusion protein can further control the amount of dsODN recruited by the Cas protein.

[0018] In this invention, GCN4 refers to a polypeptide sequence that can bind to ScFv-GCN4; ScFv-GCN4 antibody refers to a single-chain variable region fragment antibody that targets and binds to the GCN4 sequence. Both GCN4 and ScFv-GCN4 commonly used in the art are applicable to this invention, and no special restrictions are placed on the corresponding sequences.

[0019] Optionally, the amino acid sequence of the GCN4 protein includes the sequence shown in SEQ ID NO.1.

[0020] Optionally, the amino acid sequence of the ScFv-GCN4 antibody includes the sequence shown in SEQ ID NO.2.

[0021] SEQ ID NO.1:

[0022] EELLSKNYHLENEVARLKK.

[0023] SEQ ID NO.2:

[0024] MGPDIVMTQSPSSSLSASVGDRVTITTCRSSTGAVTTSNYASWVQEKPGKLFKGLIGGTNNRAPGVPSRFSGSLIGDKATLTISSLQPEDFATYFCALWYSNHWVFGQGTKVELKRGGGGSGGGG SGGGGSSGGGSEVKLLESGGGLVQPGGSLKLSCAVSGFSLTDYGVNWVRQAPGRGLEWIGVIWGDGITDYNSALKDRFIISKDNGKNTVYLQMSKVRSDDTALYYCVTGLFDYWGQGTLVTV.

[0025] In this invention, mSA refers to Monovalent streptavidin, a monovalent protein obtained by modifying streptavidin protein, which can strongly bind to a single biotin. mSA with the same function in the art are applicable to this invention, and no special restrictions are placed on the specific sequence.

[0026] Optionally, the amino acid sequence of the mSA protein includes the sequence shown in SEQ ID NO.3.

[0027] SEQ ID NO.3:

[0028] AEAGITGTWYNQSGSTFTVTAGADGNLTGQYENRAQGTGCQNSPYTLTGRYNGTKL EWRVEWNNSTENCHSRTEWRGQYQGGAEARINTQWNLTYEGGSGPATEQGQDTFTKVK.

[0029] Preferably, the 3' and 5' ends of the dsODN contain thiophosphate linkers.

[0030] In this invention, the dsODN is composed of two completely complementary single-stranded DNAs, with phosphate thioester linkers at the 3' and 5' ends of each single-stranded DNA, which can stabilize the dsODN in the cell.

[0031] Optionally, the number of the thiophosphate linkers is 2 to 8, including but not limited to 3, 4, 5, 6 or 7.

[0032] Preferably, the dsODN further contains a barcode sequence.

[0033] In this invention, a barcode refers to a short sequence that can be artificially altered to correspond to different samples. This facilitates the mixing of different samples for library preparation. The resulting sequencing data is then split and assigned to each sample based on the barcode sequence. By introducing the barcode sequence into dsODN, experiments with different CRISPR-Cas systems can be mixed during the library preparation stage, enabling simultaneous detection and analysis of multiple CRISPR-Cas systems. This saves labor costs and increases detection throughput.

[0034] It is understood that the specific sequence and length of dsODN in this invention can be adjusted according to requirements and are not subject to special restrictions.

[0035] Optionally, the length of the dsODN is 30 to 50 bp, including but not limited to 31 bp, 32 bp, 35 bp, 36 bp, 38 bp, 40 bp, 45 bp, 46 bp, 48 bp or 49 bp.

[0036] Optionally, the nucleic acid sequence of the dsODN can be selected from any one of the sequences shown in SEQ ID NO.4 to SEQ ID NO.7.

[0037] SEQ ID NO.4:

[0038] ACCGTTATTAACATATGACAACTCAATTAA(30bp).

[0039] SEQ ID NO.5:

[0040] ATACCGTTATTAACATATGACAACTCAATTAAAC(34bp).

[0041] SEQ ID NO.6:

[0042] TCGCGTATACCGTTATTAACATATGACAACTCAATTAAACGCGAGC(46bp).

[0043] SEQ ID NO.7:

[0044] CGTCGCGTATACCGTTATTAACATATGACAACTCAATTAAACGCGAGCGC(50bp).

[0045] Preferably, the 5' end of the dsODN does not contain phosphorylation modification.

[0046] In this invention, controlling the 5' end of dsODN to not contain phosphorylation modification can further reduce the probability of dsODN insertion at the spontaneous breakage site.

[0047] It is understood that the Cas expression vector and gRNA expression vector described in this invention refer to expression vectors known in the art that can express specific target genes in cells. When a specific target gene is determined (such as the fusion coding sequence of the Cas protein and GCN4 protein in this application), the type and other structures of the expression vector can be adjusted as needed.

[0048] Optionally, the Cas expression vector and the gRNA expression vector are each independently selected from plasmid vectors or viral vectors.

[0049] Optionally, the Cas expression vector further contains a pBR322 replicon, a constitutively expressed eukaryotic type II promoter, and a eukaryotic polyA signal coding sequence.

[0050] Optionally, the gRNA expression vector further contains a human U6 promoter, a constitutively expressed eukaryotic type II promoter, a T2A sequence, an antibiotic resistance gene sequence, a P2A sequence, an EGFP coding sequence, and a eukaryotic polyA signal coding sequence.

[0051] It is understood that library construction and sequencing methods commonly used in this field are applicable to this invention and are not subject to special limitations. For example, in specific embodiments of this invention, Novozymes is used. The Universal Plus DNA Library Prep Kit for Illumina V2 fragments, repairs, and ligates genomic DNA using universal adapters. These adapters contain 10–14 N (A / G / C / T) bases and serve as UMIs (Unique Molecular Identifiers).

[0052] As a preferred technical solution, the method for detecting DNA double-strand breaks generated by CRISPR-Cas includes the following steps:

[0053] (1) Transfect the Cas expression vector, gRNA expression vector and dsODN into cells, culture them and extract the cell genome. The Cas expression vector contains the fusion coding sequence of Cas protein and GCN4 protein. The gRNA expression vector contains the fusion coding sequence of ScFv-GCN4 antibody and mSA protein as well as the gRNA coding sequence. The dsODN contains biotinylation modification, 3' and 5' ends contain phosphate thioester linkers, and 5' end does not contain phosphorylation modification.

[0054] (2) The genome is sequentially fragmented, end-repaired and connected to universal adapters. Amplification reaction is performed on the dsODN sequence and the universal adapter sequence. The amplification products are collected to form the sequencing library.

[0055] (3) Sequencing the sequencing library and finding the junction sequence between dsODN and the genome through data analysis, which is the DNA double-strand break site generated by CRISPR-Cas in the cell.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] (1) This invention designs a novel method for detecting DNA double-strand break sites generated by CRISPR-Cas. By utilizing the ScFv-GCN4-mSA fusion protein and the Cas-GCN4 fusion protein system, the local dsODN concentration is significantly increased when the Cas protein cuts the DNA double strand, thereby improving the sensitivity and specificity of detecting low-frequency off-target sites. It is also widely applicable and can be used to detect CRISPR-Cas systems commonly used in this field.

[0058] (2) The present invention further designs dsODN. Controlling the 5' end of dsODN to not contain phosphorylation modification can further reduce the probability of dsODN insertion into spontaneous breakage sites. Introducing barcode sequences into dsODN can enable the mixing of samples in the library construction stage of experiments with different CRISPR-Cas systems, so that multiple CRISPR-Cas systems can be detected and analyzed at the same time, saving manpower costs and increasing detection throughput. Attached Figure Description

[0059] Figure 1 This is a schematic diagram illustrating the principle of GUIDE-advance in this invention;

[0060] Figure 2A The image shows the results of spCas9-gRNA off-target detection of dsODN (dsODN1) modified with biotin at different positions in Example 1.

[0061] Figure 2BThe image shows the results of spCas9-gRNA off-target detection of dsODN (dsODN2) modified with biotin at different positions in Example 1.

[0062] Figure 2C The image shows the results of spCas9-gRNA off-target detection of dsODN (dsODN3) modified with biotin at different positions in Example 1.

[0063] Figure 2D The image shows the results of spCas9-gRNA off-target detection of dsODN (dsODN4) modified with biotin at different positions in Example 1;

[0064] Figure 3A This is a graph showing the off-target detection results of spCas9-gRNA using traditional GUIDE-seq in Example 2;

[0065] Figure 3B This is a graph showing the off-target detection results of spCas9-gRNA in GUIDE-advance in Example 2;

[0066] Figure 4A The image shows the detection results of GUIDE-advance targeting the off-target sites of 4 spCas9-gRNAs in Example 3 (HEK1).

[0067] Figure 4B The image shows the detection results of GUIDE-advance targeting the off-target sites of 4 spCas9-gRNAs in Example 3 (HEK2).

[0068] Figure 4C The image shows the detection results of GUIDE-advance targeting the off-target sites of 4 spCas9-gRNAs in Example 3 (HEK3).

[0069] Figure 4D The image shows the detection results of GUIDE-advance targeting the off-target sites of 4 spCas9-gRNAs in Example 3 (HEK4).

[0070] Figure 5A This is a graph showing the detection results of GUIDE-advance on off-target sites of AsCas12a-crRNA in Example 4;

[0071] Figure 5B This is a graph showing the detection results of off-target sites of LbCas12a-crRNA by GUIDE-advance in Example 4. Figure 5CThe image shows the detection results of off-target sites of GUIDE-advance against AsCas12a-crRNA in Example 4 (AsCas12a-MS6). Figure 5D The image shows the detection results (LbCas12a-MS6) of off-target sites of LbCas12a-crRNA by GUIDE-advance in Example 4. Detailed Implementation

[0072] To further illustrate the technical means and effects of this invention, the following description, in conjunction with embodiments and accompanying drawings, provides a further explanation of the invention. It is understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it.

[0073] Where specific techniques or conditions are not specified in the examples, they shall be performed in accordance with the techniques or conditions described in the literature in this field, or in accordance with the product instructions. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased through legitimate channels.

[0074] In specific embodiments of the present invention, the corresponding expression vectors, using Cas plasmids and gRNA plasmids as examples, are used to verify the method of detecting DNA double-strand break sites generated by CRISPR-Cas (the present invention is named GUIDE-advance).

[0075] The Cas plasmid includes the pBR322 replicon, a constitutively expressed eukaryotic type II promoter, the Cas protein (Cas9), a 10×GCN4 protein fusion coding sequence (or an 8×GCN4 protein fusion coding sequence or a 15×GCN4 protein fusion coding sequence), and a eukaryotic polyA signal coding sequence.

[0076] Cas9 encoded sequence:

[0077]

[0078] 10×GCN44 encoded sequence:

[0079] .

[0080] It is understandable that GCN4 proteins of different lengths, such as 8× or 15×, can be reduced or increased sequentially from the 10× base.

[0081] The gRNA plasmid includes a human U6 promoter, a gRNA coding sequence, a constitutively expressed eukaryotic type II promoter, a ScFv-GCN4-mSA protein fusion coding sequence, a T2A sequence, a puromycin resistance gene sequence, a P2A sequence, an EGFP coding sequence, and a eukaryotic polyA signal coding sequence.

[0082] Material:

[0083] Lipofectamine TM 3000 transfection reagent (Thermo);

[0084] Blood / Cell / Tissue / BacteriaDNA Isolation Mini Kit (Norwegian);

[0085] Universal Plus DNA Library Prep Kit for IlluminaV2 (Norwegian);

[0086] VAHTS TM DNA Clean Beads (Novizan).

[0087] The primers to be synthesized are shown in Table 1, where N is a degenerate base representing A, G, C or T, * indicates a thiophosphate link between bases, and / 5'Phos / indicates 5' phosphorylation.

[0088] Table 1

[0089] Adapter-F1 gcttgacagtcaaccgcattggtgagatccgacaNNNNNNNNNNcgactcgtcagtg*t Adapter-R1 / 5′Phos / cactgacgagtcgacttggactgac Tag-F1 ggatctcgacgctctccctataccgttattaacatatgaca Tag-R1 ggatctcgacgctctccctgtttaattgagttgtcatatgttaataac Primer-R1 gcttgacagtcaaccgcattg Tag-F2 cactctttccctacacgacgctcttccgatctacatatgacaactcaattaaac Tag-R2 cactctttccctacacgacgctcttccgatctttgagttgtcatatgttaataacggta Prime-R2 gactggagttcagacgtgtgctcttccgatctcgcattggtgagatccgaca i5-D501 aatgatacggcgaccaccgagatctacactatagcctacactctttccctacacgac i5-D502 aatgatacggcgaccaccgagatctacacatagaggcacactctttccctacacgac i5-D503 aatgatacggcgaccaccgagatctacaccctatcctacactctttccctacacgac i5-D504 aatgatacggcgaccaccgagatctacacggctctgaacactctttccctacacgac i7-D701 caagcagaagacggcatacgagatcgagtaatgtgactggagttcagacgtg i7-D702 caagcagaagacggcatacgagattctccggagtgactggagttcagacgtg i7-D703 caagcagaagacggcatacgagataatgagcggtgactggagttcagacgtg i7-D704 caagcagaagacggcatacgagatggaatctcgtgactggagttcagacgtg

[0090] Example 1

[0091] In this embodiment, dsODN modified with biotin at different locations was used to detect intracellular off-target sites using GUIDE-advance.

[0092] 1. Plasmid construction and dsODN synthesis

[0093] 1.1 The gRNA spacer sequence on the gRNA plasmid was designed as follows: spCas9-sgRNA spacer sequence (HEK1): GGGAAAGACCCAGCATCCGT.

[0094] 1.2 The Cas protein sequence on the Cas plasmid was designed to be the spCas9 sequence.

[0095] 1.3 Design of dsODN sequences, including:

[0096] dsODN1: cgaattataccgttattaacatatgacaactcaattaaactatcgc;

[0097] dsODN2:tgattaataccgttattaacatatgacaactcaattaaaccttatt;

[0098] dsODN3: gagtaaataccgttattaacatatgacaactcaattaaacaagcag;

[0099] dsODN4: atccgagtttaattgagttgtcatatgttaataacggtatccgcga; The four primer pairs required for the synthesis of dsODN are shown in Table 2, where * indicates the phosphate thioester linkage between bases, / iBiodT / indicates biotin modification of the T base, and Bio indicates terminal biotin modification.

[0100] Table 2

[0101]

[0102] 2. Cell transfection and genome extraction

[0103] 2.1 Four 8×10^4 HEK293T cell lines were cultured, and after 18 hours, Lipofectamine was used. TM Transfection was performed using 3000 transfection reagents, including 300 ng of the designed Cas plasmid, 150 ng of gRNA plasmid, and 50 ng of annealed dsODN (corresponding to ODN-F and ODN-R annealing, with dsODN1-4 transfecting one copy of each cell).

[0104] 2.2 Three days after transfection, use Novizan. The Blood / Cell / Tissue / BacteriaDNAIsolation Mini Kit extracts cellular genomic gDNA.

[0105] 3. Random fragmentation of gDNA, end repair, addition of dA tail to the 3' end, and ligation of adapters.

[0106] 3.1 Take 400 ng of gDNA corresponding to each dsODN and use Novizan respectively. The Universal PlusDNA Library Prep Kit for Illumina V2 performs random fragmentation, end repair, and 3' tail addition on gDNA, with the expected fragment size controlled at around 500bp.

[0107] 3.2 Take the product from the previous step and continue to perform adapter ligation according to the kit instructions. Use the adapter sequence (Adapter-F1 / R1 annealed, 10 μM). The ligation system is shown in Table 3. Ligate at 20℃ for 15 minutes.

[0108] Table 3

[0109] Product of previous step 50 Rapid Ligation Buffer V2 25 Rapid DNA Ligase 5 Adapter-F1 / R1 annealed, 10μM 5 <![CDATA[ddH2O]]> 15

[0110] 3.3 Add 60 μL of VAHTS to the ligation reaction solution. TM DNA was purified using DNA Clean Beads and then eluted with 25 μL of nuclease-free ddH2O.

[0111] 4. First round of PCR amplification and purification

[0112] 4.1 Continue to use the DNA polymerase VAHTS HiFi Amplification Mix in the kit for the first round of PCR amplification. The reaction system is shown in Table 4.

[0113] Table 4

[0114] DNA eluted in the previous step 23 Tag-F1 (10μM) 1 Primer-R1 (10μM) 1 VAHTS HiFi Amplification Mix 25

[0115] The PCR procedure is shown in Table 5.

[0116] Table 5

[0117]

[0118] 4.2 Add 60 μL of VAHTS to the PCR reaction solution. TM DNA Clean Beads were purified and then moistened with 23 μL of nuclease-free ddH2O.

[0119] 5. Second round of PCR amplification and purification

[0120] 5.1 Use the DNA polymerase VAHTS HiFi Amplification Mix from the kit to perform a second round of PCR amplification. The reaction system is shown in Table 6.

[0121] Table 6

[0122] The magnetic bead suspension from the previous step 23 Tag-F2 (10μM) 1 Primer-R2 (10μM) 1 VAHTS HiFi Amplification Mix 25

[0123] The PCR procedure is shown in Table 7.

[0124] Table 7

[0125]

[0126] 5.2 Add 50 μL of VAHTS to the PCR reaction solution. TM DNA Clean Beads were purified and then eluted with 25 μL of nuclease-free ddH2O.

[0127] 6. Third round of PCR amplification and purification

[0128] 6.1 Determine the concentration of the PCR product from the previous step by diluting 20 ng to 20 μL;

[0129] 6.2 Use the DNA polymerase VAHTS HiFi Amplification Mix in the kit to perform the third round of PCR amplification. The reaction system is shown in Table 8. Different i5 / i7 primer combinations were used for each sample to distinguish them.

[0130] Table 8

[0131] Diluted second-round PCR products 20 i5-D501 / 502(10μM) 2.5 i7-D701 / D702 (10μM) 2.5 VAHTS HiFi Amplification Mix 25

[0132] The PCR procedure is shown in Table 9.

[0133] Table 9

[0134]

[0135] Add 45 μL of VAHTS to the PCR reaction solution TM DNA Clean Beads were purified and then eluted with 25 μL 1×TE.

[0136] 7. High-throughput sequencing and analysis

[0137] 7.1 Samples were generated using second-generation sequencing with paired 150bp end bands.

[0138] The results of the data analysis are as follows Figures 2A-2D As shown, comparing the off-target site detection results of four different biotin-modified dsODNs on the same sgRNA sequence, it can be seen that in the GUIDE-advance system, with the same amount of sequencing data, dsODN2, dsODN3, and dsODN4 can find more off-target sites than dsODN1, and the number of reads for each off-target site is also greater than that of dsODN1.

[0139] Example 2

[0140] This embodiment compares the GUIDE-advance of the present invention with the traditional GUIDE-seq for the detection of intracellular off-target sites.

[0141] 1. Plasmid construction and dsODN synthesis

[0142] 1.1 The design of gRNA plasmid and Cas plasmid is as described in Example 1.

[0143] 1.2 The two primer pairs required for the synthesis of dsODN (Table 10) are as follows: dsODN1 sequence: tcgcgtataccgttattaacatatgacaactcaattaaacgcgagc; and dsODN2 sequence: gagtaaataccgttattaacatatgacaactcaattaaacaagcag.

[0144] Table 10

[0145]

[0146]

[0147] 2. Cell transfection and genome extraction

[0148] 2.1 Two 8×10^4 HEK293T cell lines were cultured, and after 18 hours, Lipofectamine was used. TM Transfection was performed using 3000 transfection reagents, including 300 ng of spCas9 plasmid from GUIDE-seq (see: GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases, Nature Biotechnology, 2015, volume 33, 187-197), 150 ng of sgRNA plasmid from GUIDE-seq, and 50 ng of annealed dsODN1, or 300 ng of the Cas plasmid designed in this invention, 150 ng of the gRNA plasmid designed in this application, and 50 ng of annealed dsODN2;

[0149] 2.2 Three days after transfection, use Novizan. Blood / Cell / Tissue / BacteriaDNAIsolation Mini Kit for extracting cellular genomic gDNA;

[0150] 3. Random fragmentation of gDNA, end repair, addition of dA tail to the 3' end, and ligation of adapters.

[0151] 3.1 Take 400 ng of gDNA corresponding to each dsODN and use Novizan respectively. The Universal PlusDNA Library Prep Kit for Illumina V2 performs random fragmentation, end repair, and 3' tail addition on gDNA, with the expected fragment size controlled at around 500bp.

[0152] 3.2 Take the product from the previous step and continue to perform adapter ligation according to the kit instructions. Use the adapter sequence you designed (Adapter-F1 / R1 annealed, 10 μM). The ligation system is shown in Table 11. Ligate at 20℃ for 15 minutes.

[0153] Table 11

[0154] Previous product 50 Rapid Ligation Buffer V2 25 Rapid DNA Ligase 5 Adapter-F1 / R1 annealed, 10μM 5 <![CDATA[ddH2O]]> 15

[0155] 3.3 Add 60 μL of VAHTS to the ligation reaction solution. TM DNA was purified using DNA Clean Beads and then eluted with 25 μL of nuclease-free ddH2O.

[0156] 4. First round of PCR amplification and purification

[0157] 4.1 Continue to use the DNA polymerase VAHTS HiFi Amplification Mix in the kit for the first round of PCR amplification. The reaction system is shown in Table 12.

[0158] Table 12

[0159] DNA eluted in the previous step 23 Tag-F1 (10μM) 1 Primer-R1 (10μM) 1 VAHTS HiFi Amplification Mix 25

[0160] The PCR procedure is shown in Table 13.

[0161] Table 13

[0162]

[0163]

[0164] 4.2 Add 60 μL of VAHTS to the PCR reaction solution. TM DNA Clean Beads were purified and then moistened with 23 μL of nuclease-free ddH2O.

[0165] 5. Second round of PCR amplification and purification

[0166] 5.1 Use the DNA polymerase VAHTS HiFi Amplification Mix from the kit to perform the second round of PCR amplification. The reaction system is shown in Table 14.

[0167] Table 14

[0168] The magnetic bead suspension from the previous step 23 Tag-F2 (10μM) 1 Primer-R2 (10μM) 1 VAHTS HiFi Amplification Mix 25

[0169] The PCR procedure is shown in Table 15.

[0170] Table 15

[0171]

[0172] 5.2 Add 50 μL of VAHTS to the PCR reaction solution. TM DNA Clean Beads were purified and then eluted with 25 μL of nuclease-free ddH2O.

[0173] 6. Third round of PCR amplification and purification

[0174] 6.1 Determine the concentration of the PCR product from the previous step by diluting 20 ng to 20 μL.

[0175] 6.2 Use the DNA polymerase VAHTS HiFi Amplification Mix in the kit to perform the third round of PCR amplification. The reaction system is shown in Table 16. Different i5 / i7 primer combinations were used for each sample to distinguish them.

[0176] Table 16

[0177] Diluted second-round PCR products 20 i5-D501 (10μM) 2.5 i7-D701 / D702 (10μM) 2.5 VAHTS HiFi Amplification Mix 25

[0178] The PCR procedure is shown in Table 17.

[0179] Table 17

[0180]

[0181]

[0182] 6.2 Add 45 μL of VAHTS to the PCR reaction solution. TM DNA Clean Beads were purified and then eluted with 25 μL 1×TE.

[0183] 7. High-throughput sequencing and analysis

[0184] 7.1 The samples were obtained using second-generation sequencing with paired 150bp end-to-end sequencing.

[0185] 7.2 The results of the data analysis are as follows Figure 3A and Figure 3B As shown, comparing the off-target site detection results of the GUIDE-advance method in this invention and the traditional GUIDE-seq method on the same sgRNA sequence, it can be seen that under the same sequencing data volume, the GUIDE-advance method of this invention can find more off-target sites, and the number of reads for each off-target site is significantly more than that of GUIDE-seq.

[0186] Example 3

[0187] This embodiment verifies that the GUIDE-advance of the present invention can simultaneously detect intracellular off-target sites of multiple Cas9 / sgRNAs.

[0188] 1. Plasmid construction and dsODN synthesis

[0189] 1.1 The gRNA spacer sequences on the gRNA plasmids were designed as follows: spCas9-sgRNA spacer sequences (Table 18).

[0190] Table 18

[0191] HEK1 GGGAAAGACCCAGCATCCGT HEK2 GAACACAAAGCATAGACTGC HEK3 GGCCCAGACTGAGCACGTGA HEK4 GGCACTGCGGCTGGAGGTGG

[0192] 1.2 The Cas protein sequence on the Cas plasmid was designed to be the spCas9 sequence.

[0193] 1.3 The four primer pairs required for the synthesis of dsODN (Table 19) are as follows:

[0194] dsODN1: gagtaaataccgttattaacatatgacaactcaattaaacaagcag;

[0195] dsODN2: cgaattataccgttattaacatatgacaactcaattaaactatcgc;

[0196] dsODN3:taatcgataccgttattaacatatgacaactcaattaaactctgaa;

[0197] dsODN4:tgattaataccgttattaacatatgacaactcaattaaaccttatt.

[0198] Table 19

[0199]

[0200]

[0201] 2. Cell transfection and genome extraction

[0202] 2.1 Four 8×10^4 HEK293T cell lines were cultured, and after 18 hours, Lipofectamine was used. TMTransfection was performed using 3000 transfection reagents, including Cas plasmid, gRNA plasmid, and 50 ng of annealed dsODN (barcode sequences were designed for dsODN1-4, and the barcode sequences were used to correspond dsODN1-4 to one sgRNA).

[0203] 2.2 Three days after transfection, use Novizan. The Blood / Cell / Tissue / BacteriaDNAIsolation Mini Kit extracts cellular genomic gDNA.

[0204] 3. Random fragmentation of gDNA, end repair, addition of dA tail to the 3' end, and ligation of adapters.

[0205] 3.1 Take 250 ng of gDNA corresponding to each dsODN, mix them, and then use Novizan. The UniversalPlus DNA Library Prep Kit for Illumina V2 performs random fragmentation, end repair, and 3' tail addition on gDNA, with the expected fragment size controlled at around 500bp.

[0206] 3.2 Take the product from the previous step and continue to perform adapter ligation according to the kit instructions, using the adapter sequence you designed (Adapter-F1 / R1 annealed, 10 μM). The ligation system is shown in Table 20. Ligate at 20℃ for 15 minutes.

[0207] Table 20

[0208] Previous product 50 Rapid Ligation Buffer V2 25 Rapid DNA Ligase 5 Adapter-F1 / R1 annealed, 10μM 5 <![CDATA[ddH2O]]> 15

[0209] 3.3 Add 60 μL of VAHTS to the ligation reaction solution. TM DNA was purified using DNA Clean Beads and then eluted with 25 μL of nuclease-free ddH2O.

[0210] 4. First round of PCR amplification and purification

[0211] Refer to Example 1.

[0212] 5. Second round of PCR amplification and purification

[0213] Refer to Example 1.

[0214] 6. Third round of PCR amplification and purification

[0215] Refer to Example 1.

[0216] 7. High-throughput sequencing and analysis

[0217] 7.1 The samples were obtained using second-generation sequencing with paired 150bp end-to-end sequencing.

[0218] 7.2 The results of the data analysis are as follows Figures 4A-4D As shown, the off-target site information corresponding to the four sgRNAs can be separated from the same original sequencing data based on the different barcode sequences on their respective dsODNs (Table 21). This indicates that the present invention introduces barcode sequences into dsODNs, which enables the mixing of samples during the library construction stage for experiments with different CRISPR-Cas systems. This allows for the simultaneous detection and analysis of multiple CRISPR-Cas systems, saving manpower costs and increasing detection throughput.

[0219] Table 21

[0220]

[0221]

[0222] Example 4

[0223] This embodiment utilizes GUIDE-advance to simultaneously detect intracellular off-target sites of AsCas12a / LbCas12a-crRNA.

[0224] 1. Plasmid construction and dsODN synthesis

[0225] 1.1 The gRNA spacer sequence on the gRNA plasmid was modified to the following AsCas12a / LbCas12a-crRNA spacer sequence (Table 22).

[0226] Table 22

[0227] AsCas12a-MS6 TTTGGGGTGATCAGACCCAACAGCAGG AsCas12a-MS8 TTTGGGGACGGGGAGAAGGAAAAGAGG LbCas12a-MS6 TTTGGGGTGATCAGACCCAACAGCAGG LbCas12a-MS8 TTTGGGGACGGGGAGAAGGAAAAGAGG

[0228] 1.2 The Cas protein sequence on the Cas plasmid is designed to be either AsCas12a or LbCas12a.

[0229] 1.3 The two primer pairs required for the synthesis of dsODN (Table 23) have the following specific dsODN sequences: GAGTAAATACCGTTATTAACATdTGACAACTCAATTAAACAAGCAG; and CGAATTATACCGTTATTAACATATGACAACTCAATTAAACTATCGC.

[0230] Table 23

[0231]

[0232] 2. Cell transfection and genome extraction

[0233] 2.1 Four 8×10^4 HEK293T cell lines were cultured, and after 18 hours, Lipofectamine was used. TM Transfection was performed using 3000 transfection reagents, including Cas plasmid, gRNA plasmid, and 50 ng of annealed dsODN (dsODN1 and 2 correspond to AsCas12a-crRNA-MS6 and LbCas12a-crRNA-MS6, respectively; these two genomes were mixed for library construction. MS8 was operated in the same manner as MS6).

[0234] 2.2 Three days after transfection, use Novizan. The Blood / Cell / Tissue / BacteriaDNAIsolation Mini Kit extracts cellular genomic gDNA.

[0235] 3. Random fragmentation of gDNA, end repair, addition of dA tail to the 3' end, and ligation of adapters.

[0236] 3.1 Take 400 ng of gDNA corresponding to each dsODN, mix them, and then use Novizan. The UniversalPlus DNA Library Prep Kit for Illumina V2 performs random fragmentation, end repair, and 3' tail addition on gDNA, with the expected fragment size controlled at around 500bp.

[0237] 3.2 Take the product from the previous step and continue to perform adapter ligation according to the kit instructions, using the adapter sequence you designed (Adapter-F1 / R1 annealed, 10 μM), the ligation system is shown in Table 24, and ligation is performed at 20℃ for 15 minutes.

[0238] Table 24

[0239]

[0240]

[0241] 3.3 Add 60 μL of VAHTS to the ligation reaction solution. TM DNA was purified using DNA Clean Beads and then eluted with 25 μL of nuclease-free ddH2O.

[0242] The three rounds of PCR amplification and purification were performed as described in Example 1.

[0243] 4. High-throughput sequencing and analysis

[0244] The samples were analyzed using second-generation sequencing with paired 150bp end samples. The data analysis results are as follows: Figures 5A-5DAs shown, the GUIDE-advance of this invention can also simultaneously detect intracellular off-target sites of AsCas12a / LbCas12a-crRNA.

[0245] In summary, this invention presents a novel method for generating DNA double-strand break sites using CRISPR-Cas. Utilizing the ScFv-GCN4-mSA fusion protein and the Cas-GCN4 fusion protein system, it significantly increases the local dsODN concentration during Cas protein cleavage of the DNA double strand, thereby improving the sensitivity and specificity for detecting low-frequency off-target sites. This method has broad applicability, enabling detection of commonly used CRISPR-Cas systems. Furthermore, by designing the dsODN and ensuring that its 5' end does not contain phosphorylation modification, the probability of dsODN insertion into spontaneous break sites can be further reduced. Introducing a barcode sequence into the dsODN allows for sample mixing during the library preparation stage for experiments with different CRISPR-Cas systems, enabling simultaneous detection and analysis of multiple CRISPR-Cas systems, saving labor costs and increasing detection throughput.

[0246] The applicant declares that the detailed method of the present invention is illustrated by the above embodiments, but the present invention is not limited to the above detailed method, that is, it does not mean that the present invention must rely on the above detailed method to be implemented. Those skilled in the art should understand that any improvements to the present invention, equivalent substitutions of the raw materials of the product of the present invention, addition of auxiliary components, selection of specific methods, etc., all fall within the protection scope and disclosure scope of the present invention.

Claims

1. A method for detecting DNA double-strand breaks generated by CRISPR-Cas, characterized in that, The method includes: The Cas expression vector, gRNA expression vector, and dsODN were transfected into cells, cultured, and the cell genome was extracted. The Cas expression vector contained a fusion coding sequence of Cas protein and GCN4 protein, the gRNA expression vector contained a fusion coding sequence of ScFv-GCN4 antibody and mSA protein as well as a gRNA coding sequence, and the dsODN contained biotinylation modification. The genome was sequentially fragmented, end-repaired, and ligated with universal adapters. Amplification reactions were performed on the dsODN sequence and the universal adapter sequence, and the amplification products were collected to form the sequencing library. The sequencing library was sequenced, and the junction sequence between dsODN and the genome was found through data analysis. This sequence is the DNA double-strand break site generated by CRISPR-Cas in the cell. The method described is not for disease treatment or diagnosis purposes.

2. The method for detecting DNA double-strand breaks generated by CRISPR-Cas according to claim 1, characterized in that, The Cas protein includes any one of spCas9, AsCas12a, or LbCas12a.

3. The method for detecting DNA double-strand breaks generated by CRISPR-Cas according to claim 1, characterized in that, The number of GCN4 proteins is 8 to 15.

4. The method for detecting DNA double-strand breaks generated by CRISPR-Cas according to claim 1, characterized in that, The amino acid sequence of the GCN4 protein includes the sequence shown in SEQ ID NO.

1.

5. The method for detecting DNA double-strand breaks generated by CRISPR-Cas according to claim 1, characterized in that, The gRNA includes any one of spCas9-sgRNA, AsCas12a-crRNA, or LbCas12a-crRNA.

6. The method for detecting DNA double-strand breaks generated by CRISPR-Cas according to claim 1, characterized in that, The amino acid sequence of the ScFv-GCN4 antibody includes the sequence shown in SEQ ID NO.2; The amino acid sequence of the mSA protein includes the sequence shown in SEQ ID NO.

3.

7. The method for detecting DNA double-strand breaks generated by CRISPR-Cas according to claim 1, characterized in that, The 3' and 5' ends of the dsODN contain thiophosphate linkers.

8. The method for detecting DNA double-strand breaks generated by CRISPR-Cas according to claim 7, characterized in that, The number of thiophosphate linkers is 2 to 8.

9. The method for detecting DNA double-strand breaks generated by CRISPR-Cas according to claim 1, characterized in that, The dsODN also contains a Barcode sequence.

10. The method for detecting DNA double-strand breaks generated by CRISPR-Cas according to claim 1, characterized in that, The length of the dsODN is 30~50 bp.

11. The method for detecting DNA double-strand breaks generated by CRISPR-Cas according to claim 1, characterized in that, The 5' end of the dsODN does not contain phosphorylation modification.

12. The method for detecting DNA double-strand breaks generated by CRISPR-Cas according to claim 1, characterized in that, The Cas expression vector and gRNA expression vector are each independently selected from plasmid vectors or viral vectors.

13. The method for detecting DNA double-strand breaks generated by CRISPR-Cas according to claim 1, characterized in that, The Cas expression vector also contains a pBR322 replicon, a constitutively expressed eukaryotic type II promoter, and a eukaryotic polyA signal coding sequence.

14. The method for detecting DNA double-strand breaks generated by CRISPR-Cas according to claim 1, characterized in that, The gRNA expression vector also contains a human U6 promoter, a constitutively expressed eukaryotic type II promoter, a T2A sequence, an antibiotic resistance gene sequence, a P2A sequence, an EGFP coding sequence, and a eukaryotic polyA signal coding sequence.

15. The method for detecting DNA double-strand breaks generated by CRISPR-Cas according to claim 1, characterized in that, The method includes the following steps: (1) Transfect the Cas expression vector, gRNA expression vector and dsODN into cells, culture them and extract the cell genome. The Cas expression vector contains the fusion coding sequence of Cas protein and GCN4 protein. The gRNA expression vector contains the fusion coding sequence of ScFv-GCN4 antibody and mSA protein and gRNA coding sequence. The dsODN contains biotinylation modification, 3' and 5' ends contain phosphate thioester linkers, and 5' end does not contain phosphorylation modification. (2) The genome is sequentially fragmented, end-repaired and ligated with universal adapters. Amplification reaction is performed on the dsODN sequence and the universal adapter sequence. The amplification products are collected to form the sequencing library. (3) Sequencing the sequencing library and finding the junction sequence between dsODN and the genome through data analysis, which is the DNA double-strand break site generated by CRISPR-Cas in the cell.

Citation Information

Patent Citations

  • Nucleic acid molecule for detecting NHEJ repair and HDR repair in CRISPR-Cas gene editing, cell line, construction method and detection method

    CN118480578A

  • CRISPR / dCas9-SunTag mediated RBM25 controllable activation system and application thereof in ischemic heart failure myocardial repair

    CN120154740A

  • Base editor system and application

    CN120366269A