Engineered proteins
By performing specific amino acid substitutions and sgRNA optimization on the AsCas12f protein, a compact complex was formed, overcoming the molecular size limitations of Cas9 and Cas12a in eukaryotic cells and enabling efficient genome editing and in vivo gene therapy applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE UNIV OF TOKYO
- Filing Date
- 2024-08-23
- Publication Date
- 2026-06-02
AI Technical Summary
The existing Cas9 and Cas12a proteins are difficult to load effectively into AAV vectors in eukaryotic cells due to their large molecular size, which limits their application in genome engineering technology tools.
By performing specific amino acid substitutions in the amino acid sequence of the AsCas12f protein, a mutant of AsCas12f with high cleavage activity was developed. This mutant, combined with a small sgRNA, forms a compact AsCas12f-sgRNA-target DNA ternary complex, thereby improving genome editing efficiency.
It achieves genome editing activity comparable to or higher than SpCas9 or AsCas12a in human cells, and can be efficiently packaged in AAV vectors, making it suitable for genome editing and in vivo gene therapy.
Smart Images

Figure CN122139028A_ABST
Abstract
Description
[Technical Field]
[0002] This invention relates to engineered substances, such as engineered proteins. [Background Technology]
[0004] CRISPR-Cas systems in bacteria and archaea are known to provide adaptive immunity against foreign nucleic acids. Patent Document 1 discloses that CRISPR-Cas systems are classified into two categories (Categories 1 and 2) and six types (Types I–VI). Category 2 systems include Types II, V, and VI, comprising single multi-domain effector Cas proteins such as Cas9 (Type II) or Cas12 (Type V).
[0005] As disclosed in Patent Document 1, Cas9 and Cas12a are widely used as versatile tools in genome engineering technology because they exhibit strong nuclease activity in eukaryotic cells.
[0006] [Existing Technical Documents]
[0007] [Patent Documents]
[0008] Patent Document 1: International Publication No. 2022 / 092317 [Summary of the Invention]
[0010] [The problem the invention aims to solve]
[0011] Proteins such as Cas9 and Cas12, used as tools in genome engineering, are required to possess a variety of properties.
[0012] For example, SpCas9 derived from Streptococcus pyogenes is suitable as a genome engineering tool due to its high activity. However, its molecular size of 1367 amino acid residues (4.1 kb) is relatively large. On the other hand, in AAV vectors, the space for foreign genes to enter is less than 4.5 kb, making it difficult to even load SpCas9 into AAV vectors and introduce it into animal cells. Its use as a genome engineering tool is limited by this size.
[0013] Therefore, there is a need for engineered proteins that can be used as tools in genome engineering.
[0014] The present invention is made on the basis of the above background, and aims to provide an engineered substance that can be used as a genome engineering tool.
[0015] [Solutions for solving the problem]
[0016] The protein of one embodiment of the present invention is characterized in that, in the amino acid sequence shown in SEQ ID NO: 1, the protein has: a substitution of histidine at amino acid position 188, and further has one substitution selected from the following: a substitution of tyrosine at amino acid position 2, a substitution of tyrosine at amino acid position 70, a substitution of arginine at amino acid position 80, a substitution of threonine at amino acid position 105, a substitution of histidine at amino acid position 123, a substitution of lysine at amino acid position 195, a substitution of arginine at amino acid position 208, a substitution of alanine at amino acid position 232, a substitution of methionine at amino acid position 246, a substitution of methionine at amino acid position 316, or a substitution of isoleucine at amino acid position 337, having a sequence identity of more than 90% with the amino acid sequence shown in SEQ ID NO: 1, and having cleavage activity against a target site of DNA.
[0017] [The effects of the invention]
[0018] According to the present invention, engineered substances that can be used as tools in genome engineering technology can be provided. [Attached Image Description]
[0020]
【 Figure 1-1 A is the structure of the AsCas12f domain, and B is the composite structure analysis using cryo-electron microscopy.
[0021]
【 Figure 1-2 C is the structural model of the AsCas12f-sgRNA-target DNA ternary complex.
[0022]
【 Figure 1-3 D is a schematic diagram of sgRNA and target DNA.
[0023]
【 Figure 2-1 A is a spectrum obtained by size exclusion chromatography, and B is an explanatory diagram of a typical cryo-electron microscope image recorded in TitanKrios.
[0024]
【 Figure 2-2 C is an illustration of the single-particle cryo-electron microscopy image processing workflow, and D is a graph representing the Fourier shell correlation (FSC) curves used for 3D reconstruction.
[0025]
【 Figure 2-3 E is a spectrum obtained by size exclusion chromatography, and F is an explanatory diagram of a typical cryo-electron microscopy image recorded in TitanKrios.
[0026]
【 Figure 2-4G is an illustration of the single-particle cryo-electron microscopy image processing workflow, and H is a graph representing the Fourier shell correlation (FSC) curves used for 3D reconstruction.
[0027]
【 Figure 2-5 I is a spectrum obtained by size exclusion chromatography, and J is an explanatory diagram of a typical cryo-electron microscopy image recorded in TitanKrios.
[0028]
【 Figure 2-6 K is an illustration of the single-particle cryo-electron microscopy image processing workflow, and L is a graph representing the Fourier shell correlation (FSC) curves used for 3D reconstruction.
[0029]
【 Figure 3-1 A is an illustrative diagram comparing the structures of AsCas12f and UnCas12f.
[0030]
【 Figure 3-2 B is an illustration of the dimer interface of AsCas12f.1 and AsCas12f.2 and the recognition site of the sgRNA scaffold. C is an illustration of the overlap of AsCas12f.1 and AsCas12f.2 based on REC leaves.
[0031]
【 Figure 3-3 D is the dimer interface of REC.1 and REC.2, E is the dimer interface of RuvC.1 and RuvC.2, F is the recognition of stem 2 by AsCas12f.2, and G is an illustrative diagram of the recognition of stem 2 by AsCas12f.1.
[0032]
【 Figure 4-1 A is the recognition site of guide RNA and target DNA, and B is an illustration of the electrostatic surface potential of AsCas12f.
[0033]
【 Figure 4-2 [C-E are illustrative diagrams illustrating the recognition of stem 1 and PK 1, stem 3, and guide RNA-target DNA heteroduplex, respectively.]
[0034]
【 Figure 4-3 [F] is an illustrative diagram comparing the structures of the TS, NTS, RuvC active sites of AsCas12f and UnCas12f from uncultured archaea.
[0035]
【 Figure 5-1 A is the DMS library design for AsCas12f, and B is an illustrative diagram outlining the DMS method for evaluating gene editing efficiency in the context of GFP gene deletion.
[0036]
【 Figure 5-2C is an illustration of the effect of a single mutation on genome editing activity; D is an illustration of the Split-GFP-recombination assay used to evaluate genome editing activity of the VEGF gene in HEK293T cells; E is an illustration of the genome editing activity of the AsCas12f mutation; and F is a graph showing the time elapsed for GFP deletion induced by AsCas12f-HKRA and AsCas12f in HEK293T cells expressing d2EGFP.
[0037]
【 Figure 6-1 A is a graph showing the correlation between the effect of targeting the GFP-deficient single mutant.
[0038]
【 Figure 6-2 B-1 is a list of single amino acid substitutions that offer more than 20% improved editing efficiency compared to WT.
[0039]
【 Figure 6-3 B-2 is a list of single amino acid substitutions that offer more than 20% improved editing efficiency compared to WT.
[0040]
【 Figure 6-4 C, D, and E are graphs representing the genome editing activity of double, triple, and quadruple AsCas12f mutants, respectively.
[0041]
【 Figure 7-1 A is an illustration of the sequencing of sgRNA mutants; B is an illustration of the genome editing activity of sgRNA mutants as determined by flow cytometry; C is a graph showing the expression of sgRNA quantified by qPCR at a common sequencing site (n=3 means ±SD); and D is a graph showing the time elapsed for GFP deletion.
[0042]
【 Figure 7-2 E is a schematic diagram of wild-type sgRNA, and F is a schematic diagram of sgRNA_ΔS3-5_v7.
[0043]
【 Figure 7-3 G represents the structure of the wild-type sgRNA scaffold, and H represents the structure of the sgRNA_ΔS3-5_v7 scaffold.
[0044]
【 Figure 7-4 I is the PK1 region of wild-type sgRNA, J is the PK1 region of sgRNA_ΔS3-5_v7, K is the PK2 region of wild-type sgRNA, and L is an enlarged view of the stem 3 region of sgRNA_ΔS3-5_v7.
[0045]
【 Figure 8-1 Figure A is an illustration of the results of target DNA and target DNA complexes obtained using cryo-electron microscopy.
[0046]
【 Figure 8-2 [B is the dimer interface of AsCas12f-YHAM, C is the target DNA recognition of AsCas12f-YHAM, D is the hydrophobic interaction of AsCas12f-YHAM, E is the recognition of heteroduplex target DNA by the guide RNA of AsCas12f-HKRA, F is the recognition of target DNA, and G is the recognition of sgRNA scaffold.]
[0047]
【 Figure 9 A diagram illustrating the results of data collection, processing, model improvement, and validation.
[0048]
【 Figure 10 A diagram illustrating the results of data collection, processing, model improvement, and validation.
[0049]
【 Figure 11-1 A is an explanatory diagram of the self-targeted library, and B is a graph representing the editing frequency.
[0050]
【 Figure 11-2 C to E are graphs representing the genome editing activity in various Cas.
[0051]
【 Figure 12-1 A is a graph showing the comparison results of target editing efficiency, and B is an illustration of the recognition results of PAM double strands.
[0052]
【 Figure 12-2 C and D are curves representing the insertion / deletion efficiency in other target sites, etc.
[0053]
【 Figure 12-3 E is a graph representing the insertion / deletion efficiency in other target sites, etc.
[0054]
【 Figure 13-1 A is a graph showing the impact of mismatch, and B is an explanatory graph comparing the number of off-target edits.
[0055]
【 Figure 13-2 C-1 is an illustration of off-target sites identified by the guide sequence.
[0056]
【 Figure 13-3 C-2 is an illustration of off-target sites identified through a guide sequence.
[0057]
【 Figure 13-4 C-3 is an illustration of off-target sites identified through a guide sequence.
[0058]
【 Figure 14-1[Illustrative diagram of the preparation of artificial pluripotent stem cells (iPSCs) in a muscular dystrophy model and their differentiation into cardiomyocytes (iPSC-CM). In the diagram, A is an illustration of the complete deletion of DMD exons confirmed by PCR, and B is an illustration of the genome sequence of the DMD exon 44 deletion-iPS cell line.]
[0059]
【 Figure 14-2 [Illustrative diagram of the creation of artificial pluripotent stem cells (iPSCs) in a muscular dystrophy model and their differentiation into cardiomyocytes (iPSC-CM). C is an illustration of the flow cytometry analysis of α-actin-mCherry on day 14 of cardiomyocyte differentiation, and D is a diagram representing Western blot analysis.]
[0060]
【 Figure 14-3 E is a graph showing the comparison of luciferase activity, F is a schematic diagram of the plasmid construct, and G is a graph showing the increase in luciferase activity in cells in which the plasmid vector has been introduced.
[0061]
【 Figure 15-1 A through F are illustrative diagrams illustrating the use of mutants for treatment and knock-in.
[0062]
【 Figure 15-2 G to J are illustrations of treatment and knock-in using mutants.
[0063]
【 Figure 16-1 Figures A through F illustrate the use of the AsCas12f mutant for FIX knock-in at the Alb locus (in mice).
[0064]
【 Figure 16-2 G-J are illustrative diagrams illustrating gene transcription activation using the AsCas12f mutant.
[0065]
【 Figure 17 A is an illustration of the expression of luciferase activated by gene transcription in mice, and B is a graph representing the ROI value.
[0066]
【 Figure 18 [Illustrative diagram of nucleic acid sequences used for structural analysis associated with the STAR method]
[0067]
【 Figure 19 [Illustrative diagram of nucleic acid sequences used for genome editing associated with the STAR method]
[0068]
【 Figure 20-1 [Illustrative diagram of primers and constructs used for genome editing associated with the STAR method]
[0069]
【 Figure 20-2 [Illustrative diagram of primers and constructs used for genome editing associated with the STAR method]
[0070]
【 Figure 20-3 [Illustrative diagram of primers and constructs used for genome editing associated with the STAR method]
[0071]
【 Figure 20-4 [Illustrative diagram of primers and constructs used for genome editing associated with the STAR method]
[0072]
【 Figure 21-1 A diagram illustrating the key resources in the STAR method.
[0073]
【 Figure 21-2 A diagram illustrating the key resources in the STAR method.
[0074]
【 Figure 21-3 A diagram illustrating the key resources in the STAR method.
[0075]
【 Figure 21-4 A diagram illustrating the key resources in the STAR method.
[0076]
【 Figure 21-5 A diagram illustrating the key resources in the STAR method.
[0077]
【 Figure 21-6 A diagram illustrating the key resources in the STAR method.
[0078]
【 Figure 21-7 A diagram illustrating the key resources in the STAR method.
[0079]
【 Figure 22 Tobacco Benzoviae using a viral vector expressing the AsCas12f mutant (Tobacco Benzoviae var. rubrum) Nicotiana benthamiana A schematic diagram illustrating the genome editing process.
[0080]
【 Figure 23 (a) and (b) are illustrations of the results of the visualization of the genome-edited cells.
[0081]
【 Figure 24 (a) is an illustration of the NGS analysis results using PVX-AsCas12f mutant and PVX-SpCas9 respectively, and (b) is an illustration of the CAPS analysis results.
[0082]
【 Figure 25 (a) is a schematic diagram illustrating genome editing of tomato using a viral vector expressing the AsCas12f mutant, and (b) is a schematic diagram illustrating the CAPS analysis results using the PVX-AsCas12f mutant.
Detailed Implementation Methods
[0084] In this specification, when representing substitution mutations in an amino acid sequence, it is sometimes represented by the letter of the original amino acid, followed by a position number of 1 to 3 digits, and then by the letter of the substituted amino acid. For example, in the case where a mutation occurs in which aspartic acid (D) at amino acid position 1022 is replaced by asparagine (N), it is represented as "D1022N", which is synonymous with "substitution of Asp to Asn at amino acid position 1022". The -CO-NH- bond connecting amino acids is called a peptide bond. The amino acid in which these bonds are formed is called an amino acid residue. In addition, "mutation" refers to a variation that occurs in the starting amino acid or nucleic acid sequence. This term is intended to include substitution, insertion, and deletion. In addition, in this embodiment, the engineered substance is a protein, such as engineered AsCas12f. The engineered substance may also be used as a composition with other substances such as guide RNA, vectors, etc.
[0085] The CRISPR-Cas system in bacteria and archaea provides adaptive immunity against foreign nucleic acids and is classified into two categories (Category 1 and 2) and six types (Types I to VI). The Category 2 system includes Types II, V, and VI, which include single multi-domain effector Cas proteins such as Cas9 (Type II) or Cas12 (Type V).
[0086] The applications of CRISPR-Cas are not limited to genome editing. Because Cas, which eliminates cleavage activity, functions as an RNA-dependent DNA-binding protein, it can be used for techniques such as transcriptional activity modulation and single-base editing by fusing with various effectors.
[0087] In genome editing enzymes, viral vectors such as AAV are used to introduce the gene into the target tissue, but the size of the gene that can be carried is limited. Due to the size of Cas9 and Cas12a, they are difficult to carry.
[0088] Cas9 binds to double-stranded RNA guides (CRISPR RNA [crRNA] and trans-activating crRNA [tracrRNA]) or single-stranded guide RNA (sgRNA), cleaving double-stranded DNA (dsDNA) targets in sequences that are complementary to the 20nt guide fragment of the RNA guide and adjacent to the NGG (N is any nucleotide) proto-spacer adjacent motif (PAM).
[0089] Among the various types of V Cas12, type VA Cas12a (also known as Cpf1) binds to crRNA and cleaves dsDNA targets via TTTV (V is A, G, or C) PAM. Cas9 contains two nuclease domains, HNH and RuvC, that cleave the target strand (TS) and non-target strand (NTS) of the dsDNA target, respectively.
[0090] In contrast, Cas12a cleaves both TS and NTS via a single RuvC nuclease domain. Because Cas9 and Cas12a exhibit potent nuclease activity in eukaryotic cells, they are widely used as versatile tools in genome engineering.
[0091] In addition, Streptococcus pyogenes ( Streptococcus pyogenes Cas9 (type II) (SpCas9) and in various types V Cas12, from acetamococci ( Acidaminococcus sp. Cas12a (AsCas12a) is widely used as a genome editing tool in human cells.
[0092] However, these relatively large sizes become a limitation in the delivery of adeno-associated virus vectors due to cargo size constraints. On the other hand, those from *Bacillus oxidans* (… Acidibacillus sulfuroxidans The VF type Cas12f (AsCas12f) is very compact (422 amino acids) and functions in human cells. Table 1 shows the size comparison results of AsCas12f with the Cas types reported to date. AsCas12f has the amino acid sequence shown in SEQ ID NO: 1, and is the smallest reported Cas type with 422 residues, indicating that its size is less than half that of Cas9, Cas12a, Cas12b, Cas12e, etc.
[0093] Table 1
[0094]
[0095] However, the nuclease activity of wild-type AsCas12f is relatively low. Therefore, in this embodiment, two AsCas12f activity-enhancing mutants (enAsCas12f) were developed by using cryo-electron microscopy structural analysis and mutant screening.
[0096] In particular, the enAsCas12f mutant exhibited genome editing activity equivalent to or exceeding that of SpCas9 or AsCas12a in human cells. Structural analysis revealed that the mutant stabilized dimer formation, enhanced nucleic acid interactions, and improved DNA cleavage activity. Furthermore, the integrated AAV vector, packaged with a partner gene, demonstrated highly efficient knock-in / knockout activity and transcriptional activation in mice. Taken together, enAsCas12f has the potential to provide a minimal genome editing platform for in vivo gene therapy.
[0097] The Cas (clustered regularly spaced short palindromic repeats and CRISPR-associated proteins) system, which provides adaptive immunity to mobile genetic elements in bacteria and archaea, is classified into two classes (classes 1 and 2) and six types (types I-VI). SpCas9, accompanied by double-stranded RNA guides (CRISPR RNA [crRNA] and trans-activating crRNA [tracrRNA], or artificially linked single-stranded guide RNA [sgRNA]), uses its HNH and RuvC nuclease domains to cleave double-stranded DNA (dsDNA) targets embedded in a protospacer adjacent motif (PAM) of NGG (where N is any nucleotide). In contrast, AsCas12a binds to crRNA and uses a single RuvC nuclease domain to cleave dsDNA targets with a TTTV (V is A, G, or C) PAM.
[0098] SpCas9 and AsCas12a exhibit robust nuclease activity in eukaryotic cells and are widely used as highly versatile genome engineering tools. However, the large gene sizes of both SpCas9 (1,368 amino acids) and AsCas12a (1,307 amino acids) make them difficult to package efficiently into a single adeno-associated virus (AAV) vector, hindering their clinical application in vivo gene therapy.
[0099] The VF-type Cas12 effector is known to be a very compact (400–700 amino acids) RNA-guided DNA endonuclease. Cas12f binds to double-stranded crRNA:tracrRNA guides and cleaves target DNA with T-rich PAM. Based on the structural studies to date of Cas12f (UnCas12f) from uncultured archaea with 529 amino acids, it is known that UnCas12f functions as a dimer to compensate for its small size. However, UnCas12f only cleaves DNA targets under low-salt conditions in vitro and lacks activity in human cells, thus limiting its application as a genome editing tool. On the other hand, as explained below, Cas12f from *Bacillus sulfadiazine* (… Acidibacillus sulfuroxidansThe smallest Cas12f (AsCas12f) consists of only 422 amino acids and can cleave DNA targets with TTR (R is A or G) PAM under in vitro physiological conditions. Furthermore, AsCas12f exhibits small but detectable genome editing activity in human cells. Therefore, AsCas12f is useful as a small genome editing tool that can be packaged into a single AAV vector.
[0100] In this embodiment, improvements to the AsCas12f system for genome editing were investigated. Through combinatorial structural analysis and DMS (deep mutational scanning), the molecular basis was elucidated, and a detailed landscape of amino acid substitutions that significantly enhance the nuclease activity of AsCas12f was identified. Through the synergistic effect of these mutations and guide RNA engineering, the genome editing efficiency of AsCas12f in human cells was significantly improved, reaching levels comparable to or higher than SpCas9 or engineered Cas12a proteins. The compact size of AsCas12f is highly valuable for use with partner genes such as AAV-deliverable gRNAs, base editors, and epigenome modifiers, making the AsCas12f system a promising genome editing platform.
[0101] <Cryo-electron microscopy structure of the AsCas12f-sgRNA-target DNA ternary complex>
[0102] Figure 1-1 Figures showing the structure of the AsCas12f domain (A) and the cryo-electron microscopy pattern of the AsCas12f-sgRNA-target DNA ternary complex (B). Figure 1-2 A diagram (C) showing the structural model of the AsCas12f-sgRNA-target DNA ternary complex. Figure 1-3 This is a schematic diagram of sgRNA and target DNA (D). The disordered region is enclosed in a dashed box. TS represents the target strand, NTS represents the non-target strand, and PK represents spurious binding. To distinguish NTS from TS, NTS is marked with an asterisk.
[0103] Figure 2-1 and Figure 2-2 Figures (A-D) show the cryo-electron microscopy analysis of AsCas12f-sgRNA-target DNA. Figure 2-3 and Figure 2-4 The figure shows the cryo-electron microscopy analysis (E-H) of the target DNA of AsCas12f-YHAM-sgRNA_ΔS3-5_v7-. Figure 2-5 and Figure 2-6Figures showing cryo-electron microscopy analysis (I-L) of the AsCas12f-HKRA-sgRNA_ΔS3-5_v7-target DNA complex.
[0104] In detail, Figures 2-1 to 2-6 In the diagram, A represents the size-exclusion chromatography (MEC) chromatogram of AsCas12f-sgRNA-target DNA, E represents the MEC chromatogram of EAsCas12f-YHAM-sgRNA_ΔS3-5_v7-target DNA, and I represents the MEC chromatogram of the AsCas12f-HKRA-sgRNA_ΔS3-5_v7-target DNA complex. Peak fractionation was used for the following cryo-electron microscopy analysis.
[0105] In addition, Figures 2-1 to 2-6 In the diagram, B, F, and J represent representative cryo-electron microscopy images recorded with a K3 camera in a TitanKrios at 300 kV. C, G, and K are illustrative diagrams illustrating the single-particle cryo-electron microscopy image processing workflow. D, H, and L are graphs showing the Fourier shell correlation (FSC) curves (FSC=0.143) used for 3D reconstruction, representing the gold-standard cut-off value indicated by the black dotted line.
[0106] like Figures 1-1 to 1-2 A to C Figures 2-1 to 2-2 A through D, and the results representing data collection, processing, model improvement, and validation. Figure 9 and Figure 10 As shown, to understand the molecular mechanism of AsCas12f, cryo-electron microscopy (cryo-EM) was used to resolve the complex structure of AsCas12f and a 38-nt dsDNA modified with TTA PAM phosphate thiophosphate, consisting of a 222 nucleotide (nt) sgRNA and a DNA backbone around the cleavage site. The structure was reconstructed at full resolution of 3.1 Å. The structure reveals that an asymmetric homodimer is formed through the assembly of two AsCas12f molecules (AsCas12f.1 and AsCas12f.2) and one sgRNA molecule.
[0107] The Cas12f dimer adopts a bilobal structure with a recognition (REC) leaflet and a nuclease (NUC) leaflet, with the guide RNA-target DNA heteroduplex binding in a central channel between the two leaves. The REC leaflet consists of the wedge-shaped (WED) domain and REC domain of AsCas12f.1 and AsCas12f.2 (WED.1 / WED.2 / REC.1 / REC.2), while the NUC leaflet contains the RuvC domain and target nucleic acid binding (TNB) domain of AsCas12f.1 and AsCas12f.2 (RuvC.1 / RuvC.2 / TNB.1 / TNB.2). Furthermore, the AsCas12f-sgRNA-DNA complex was found to function as a dimer, similar to UnCas12f.
[0108] like Figure 1-3 As shown in D, the sgRNA consists of a 20 nt guide fragment (G1 to C20) and a 202 nt sgRNA scaffold (A(-202) to C(-1)). The sgRNA scaffold has an unexpected architecture with two pseudoknots (PK1, PK2) and five stems (stem 1 to stem 5) that cannot be predicted from its primary sequence. Stem 1 has a 6-base-pair (bp) double strand (U(-192):A(-175) to G(-187):C(180)), and PK1 consists of a 4 bp double strand (U(-185):A(-2) to A(-182):U(-5)).
[0109] Stem 2 comprises a 16bp double strand containing two atypical base pairs (G(-171):U(-125) to G(-161):U(-135)). Stem 3 consists of three base pairs (G(-118):C(-109) to C(-116):G(-111)). PK2 has two base pairs (G(-122):C(-88) and G(-121):C(-87)) connecting stems 2 and 3. Stem 4 comprises a 16bp double strand coaxially stacked from two double strands (U(-102):A(-80) to C(-96):G(-86)). Figure 1-3 As shown in D, in particular, the C (-103) located between stems 3 and 4 is reversed (turned out) and inserted into stems 1 and PK1, forming a base pair with G (-181), thereby forming a continuous helix composed of stems 1 and PK1. Figure 1-3 As shown, the 5' region (A(-202):A(-194)) and stem 5 (U(-70):U(-15)) are structurally disordered, suggesting their flexibility.
[0110] exist Figures 3-1 to 3-3 middle, Figure 3-1A represents a structural comparison between AsCas12f and UnCas12f from uncultured archaea (PDBID: 7C7L). Figure 3-2 The B represents the dimer interface of AsCas12f.1 and AsCas12f.2 and the recognition site of the sgRNA scaffold. Figure 3-2 The C represents the overlap of AsCas12f.1 and AsCas12f.2 based on REC leaves. Figure 3-3 D represents the dimer interface between REC.1 and REC.2. Figure 3-3 E represents the dimer interface between RuvC.1 and RuvC.2. Figure 3-3 F represents the recognition of stem 2 by AsCas12f.2. Figure 3-3 G represents the recognition of stem 2 by AsCas12f.1.
[0111] AsCas12f (422 residues) is more than 100 residues smaller than UnCas12f (529 residues), and its homologous sgRNA is about 40 nucleotides longer than UnCas12f. Figure 3-1 As shown in Figure A, a structural comparison of AsCas12f and UnCas12f reveals that, although their domain compositions are similar, AsCas12f lacks the zinc finger (ZF) domain inserted between the WED and REC domains observed in UnCas12f. Instead, in the AsCas12f structure, the ZF domain is replaced by PK2 and stem 3, which are absent in UnCas12f. These insights explain the miniaturization of AsCas12f, which is compensated for by a longer sgRNA.
[0112] <Small AsCas12f scaffolds associated with large sgRNAs>
[0113] like Figures 3-2 to 3-3 As shown in Figures B to E, AsCas12f.1 and AsCas12f.2 interact with each other primarily via hydrophobic interactions at the REC.1 / REC.2 and RuvC.1 / RuvC.2 interfaces, thereby promoting the dimerization of AsCas12f, as observed in UnCas12f. Figure 3-3 As shown in D and E, REC.1 and REC.2 form symmetrical interfaces with W43, F48, and I116 acting as centers on each other, while RuvC.1 and RuvC.2 form asymmetrical interfaces where the protopolymers adopt different conformations. Specifically, as... Figure 3-3 As shown in E, the α1 helix and α1-α2 ring of RuvC.1 interact with the α1 helix and α2 helix of RuvC.2. Figure 3-3As shown in F and G, in addition to protein-protein interactions, the lower and upper regions of stem 2 interact extensively with AsCas12f.1 and AsCas12f.2, respectively, suggesting that protein-nucleic acid interactions also contribute to the dimerization of AsCas12f.
[0114] Figure 4-1 The A in the diagram represents the recognition site between the guide RNA and the target DNA. Figure 4-1 B represents the electrostatic surface potential of AsCas12f. The heteroduplexes of sgRNA and target DNA, as well as the stem 1, stem 2, and PK1 regions of the sgRNA scaffold, are housed within positively charged grooves of the AsCas12f dimer. Figure 4-2 C to E represent stem 1 and PK1, stem 3, and the recognition of guide RNA-target DNA heteroduplex, respectively. Figure 4-3 F indicates a structural comparison of the TS, NTS, and RuvC active sites of AsCas12f with that of UnCas12f from uncultured archaea (PDBID: 7C7L). The location of the RuvC.1 active site for the target DNA is similar in both structures, suggesting that AsCas12f also uses the RuvC.1 domain to cleave the target DNA.
[0115] like Figure 4-1 As shown in A and B, the assembly of the sgRNA scaffold and AsCas12f is facilitated through both base-specific and non-specific interactions. Figure 4-2 As shown in Figure C, the continuous spiral formed by stem 1 and PK1 is primarily identified and contained within the grooves formed by the WED.1 and RuvC.1 structural domains. Figure 4-3 As shown in F and G, it is believed that stem 2 interacts extensively with the two protopolymers through the gap between AsCas12f.1 and AsCas12f.2, thereby enhancing dimerization as described above. Figure 4-2 As shown in D, PK2 is recognized by the WED.1 and REC.1 domains via sugar-phosphate backbone interactions and is subsequently stabilized through coordination with metal ions. Conversely, stems 3 and 4, when exposed to the solvent, exhibit minimal interactions with AsCas12f.
[0116] like Figure 4-1 B, Figure 4-2 As shown in Figure E, the heteroduplex of guide RNA and target DNA is housed within a positively charged central channel and recognized by AsCas12f through non-specific base interactions. Figure 4-2As shown in E, the PAM-proximal region (G1:dC18~G12:dC7) of the heteroduplex is mainly recognized by Cas12f.1, but the PAM-distant region (C13:dG6~G18:dC1) of the heteroduplex is recognized by AsCas12f.2. Figure 4-2 As shown in E, in particular, the P240.2 of RuvC.2 is stacked with the G18:dC1 base pair of the heteroduplex, indicating that the 18 nucleotides of the spacer sequence function as a guide fragment. Figure 4-3 As shown in Figure F, a structural comparison with UnCas12f reveals that the position of RuvC.1 in AsCas12f is similar to that of RuvC.1 in UnCas12f, which cleaves both the target and non-target strands. These structural observations suggest that AsCas12f.1 is responsible for cleaving the target DNA, while AsCas12f.2 plays a crucial role in recognizing the PAM-distant region of heteroduplex DNA.
[0117] <AsCas12f Engineering for Improving Genome Editing Efficiency in Mammalian Cells>
[0118] AsCas12f is known to exhibit limited genome editing activity in human cells. To expand the usefulness of this ultracompact protein, enhanced AsCas12f mutants were created. Deep mutational scanning (DMS) was performed to investigate the impact of all amino acid substitutions on genome editing efficiency in HEK293T cells. Specifically, using DMS, the effects of mutations replacing all 422 residues of AsCas12f with 20 amino acids on genome editing activity targeting GFP (Green Fluorescent Protein) were evaluated (genome editing efficiency was determined by the disappearance of GFP fluorescence).
[0119] exist Figure 5-1 In the diagram, A is an illustration of the DMS library design for AsCas12f, and B is a summary diagram representing the DMS method used for genome-wide evaluation of gene editing efficiency in the context of GFP gene deletion. Figure 5-2 In the diagram, C illustrates how a single mutant affects all genome editing activities. D is an illustration of the Split-GFP-reconstruction assay used to evaluate genome editing activity against the VEGF gene in HEK293T cells. The plasmid-encoded mCherry is co-infected to recognize plasmid-infected cells. Additionally, in... Figure 5-2In the diagram, E represents the genome editing activity of the AsCas12f mutant carrying editing efficiency-enhancing mutations, as determined by flow cytometry (n=4). F is a graph showing the time elapsed (n=3, mean ± standard deviation) of GFP deletion induced by AsCas12f-HKRA and AsCas12f in HEK293T cells expressing d2EGFP (n=3). sgRNA_ΔS5 and sgRNA_ΔS3-5_v7 are transduced using a single copy of lentiviral infection. It should be noted that in... Figure 5-2 In the curve of F, (a) represents the result of HKRA+ΔS3-5_v7, (b) represents the result of HKRA+ΔS5, (c) represents the result of WT+ΔS3-5_v7, and (d) represents the result of WT+ΔS5.
[0120] like Figure 5-1 As shown in Figure A, plasmids expressing both EGFP and AsCas12f were designed, and an AsCas12f library containing all 20 amino acid substitutions of the AsCas12f sequence (M1-K422) was constructed. Figure 5-1 As shown in Figure B, GFP-targeting sgRNA was introduced into HEK293T cells, and then an AsCas12f library packaged in lentivirus was expressed under conditions of infection multiplicity (MOI) less than 0.2, ensuring that the number of mutated AsCas12f1 expressed in each cell did not exceed 1. Since the genome editing efficiency of full-length sgRNA is not optimal, a stem 5 deletion mutant (sgRNA_ΔS5) with enhanced genome editing in human cells was used for screening.
[0121] On day 2 post-infection, GFP-positive cells were selected for lentiviral infection and cultured for another 5 days. Figure 5-1 As shown in Figure B, mRNA was extracted from GFP-positive and GFP-negative cells on day 7 post-infection and deep sequencing analysis was performed to identify mutations. Figure 5-2 The C-value represents genome editing efficiency using bar graphs and concentration distributions. Compared to the wild type, as genome editing efficiency increases, the bars become longer and the color in the concentration distribution becomes darker. It should be noted that... Figure 5-2 In C, the concentration distribution is approximate, and the editing efficiency is mainly represented by a bar graph. As shown in the figure, the editing efficiency of each mutation is defined as the ratio of GFP-negative read counts to the total read counts, and is normalized using the value of wild-type AsCas12f (WT: wild-type).
[0122] Figure 6-1In the figure, A represents a graph showing the correlation between the effects of a single mutant with GFP deletion, derived from the replicas of independent mutant libraries. Figure 6-2 B-1 and Figure 6-3 B-2 in the list is a single amino acid substitution that increases editing efficiency by more than 20% compared to WT. Figure 6-4 In the diagram, C represents the genome editing activity of the AsCas12f mutant when set to dual, D represents the genome editing activity of the AsCas12f mutant when set to triple, and E represents the genome editing activity of the AsCas12f mutant when set to quadruple.
[0123] like Figure 6-1 As shown in A, the DMS experiment was conducted under dual conditions and obtained the same results [R]. 2 (Determination coefficient) = ~0.6]. The result is as follows: Figure 6-2 and 6-3 B-1, B-2 and Figure 6-4 As shown in C, over 200 amino acid substitutions with editing efficiency exceeding 20% compared to WT were identified. It should be noted that in B-1 and B-2, 208 mutations with genome editing efficiency increasing by more than 1.2 times compared to wild-type are shown on the horizontal axis, and the vertical axis represents the genome editing efficiency of each mutation when the wild-type genome editing efficiency is set to 1. Figures 6-2 to 6-4 In B-1 to E, the mutations (F48Y, I123H, S188H, D195K, D208R, V232A, E316M) contained in the final mutants are emphasized by adding slashes.
[0124] Figure 6-4 C, D, E and Figure 5-2 E and F represent the results of mutation combinations and genome editing efficiency assays. In these assays, [the results are related to...]. Figure 6-2 and Figure 6-3 In contrast to B-1 and B-2, for genes that cause GFP fluorescence to disappear, GFP fluorescence is revived by inducing genome editing. Attention is focused on S188H and D195K in these mutations. Located near nucleic acids in this structure, they have the potential to form additional interactions that enhance AsCas12f-mediated target DNA binding. Therefore, multiple mutations were combined for S188H and D195K respectively, ultimately developing two highly active mutants ( Figure 5-2 E).
[0125] Next, as Figure 5-2As shown in Figure D, a split-GFP reporter system ligated with a frame transfer adapter containing a VEGFA sequence was developed. The genome editing activity of WT and several mutants was evaluated by measuring the ratio of frame-transfer GFP-positive cells. Figure 5-2 As shown in Figure E, consistent with the DMS analysis, the S188H and D195K mutants exhibited higher genome editing activity than the WT mutant. To create mutants with higher genome editing activity than the S188H or D195K mutants alone, more than 20 mutants were created by combining the S188H or D195K substitution with other effective substitution combinations identified in the DMS analysis.
[0126] like Figure 5-2 E, Figure 6-4 As shown in Figure C, several double mutants exhibited decreased genome editing activity, but S188H / V232A and D195K / V232A showed increased activity. To further enhance activity, mutations that enhance activity when combined with S188H or D195K were added, creating seven triple mutants. Figure 5-2 E and Figure 6-4 As shown in Figure D, the genome editing activity of the triple mutants S188H / V232A / E316M and D195K / D208R / V232A was further enhanced. Finally, 29 quadruple mutants were constructed by combining the triple mutants S188H / V232A / E316M or D195K / D208R / V232A with other effective substituent combinations identified in the DMS analysis, and their genome editing activity was evaluated. As shown in Figures 5E and 6E, most of these mutants showed equivalent or lower activity than the triple mutants, but the genome editing activity was further enhanced in the quadruple mutants F48Y / S188H / V232A / E316M and I123H / D195K / D208R / V232A. In this specification, the F48Y / S188H / V232A / E316M and I123H / D195K / D208R / V232A mutants are named AsCas12f-YHAM and AsCas12f-HKRA, respectively.
[0127] In addition, such as Figure 5-2As shown in the graph of E, the proportion of frame-shifted GFP-positive cells in the example S188H (no mutation at amino acid number 195 but a mutation at amino acid number 188, and no mutations in other amino acid numbers) and the example D195K (no mutation at amino acid number 188 but a mutation at amino acid number 195, and no mutations in other amino acid numbers) were approximately 18%, respectively. Therefore, it is shown that the proportion of frame-shifted GFP-positive cells in mutants with a mutation at amino acid number 188 is significantly higher than that in the comparative example WT, particularly in S188H where the mutation at amino acid number 188 is histidine (H), the proportion of frame-shifted GFP-positive cells is significantly higher than that in WT. Furthermore, it is shown that the proportion of frame-shifted GFP-positive cells in mutants with a mutation at amino acid number 195 is higher than that in WT, particularly in D195K where the mutation at amino acid number 195 is lysine (K), the proportion of frame-shifted GFP-positive cells is significantly higher than that in WT.
[0128] In contrast, under the condition of no mutation at amino acid position 195, in the example S188A / V232A, a double mutant with a further mutation at amino acid position 232 on top of a mutation at amino acid position 188, the number of frameshifted GFP-positive cells was approximately 2-fold higher than that of S188H, a significant increase. Furthermore, in the example S188H / V232A / E316M, a triple mutant with a mutation at amino acid position 316 on top of mutations at amino acid positions 188 and 232, the number of frameshifted GFP-positive cells was approximately 2.5-fold higher than that of S188H, a further increase.
[0129] Furthermore, in the example of the quadruple mutant F48Y / S188H / V232A / E316M, which has a mutation at amino acid position 48 based on mutations at amino acid positions 188, 232, and 316, the number of frameshifted GFP-positive cells was approximately three-fold higher compared to S188H, indicating a further increase in the value. Moreover, these results strongly suggest that even in multiple mutants such as pentapex or hexaplex mutants containing the F48Y / S188H / V232A / E316M mutation, the proportion of frameshifted GFP-positive cells is higher compared to WT. This demonstrates the usefulness of mutants with mutations at amino acid position 188, particularly mutants containing only the S188H mutation, and any multiple mutants such as double, triple, and quadruple mutants containing at least the S188H mutation. Therefore, the important insight that the proportion of frameshifted GFP-positive cells is significantly increased in the case of mutants with mutations at amino acid position 188 is obtained.
[0130] In addition, Figure 6-4In C, the double mutant of S188H and D208R showed a value approximately 3-fold higher than WT, but lower than other double mutants. Furthermore, as... Figure 6-4 As shown in E, the quadruple mutant of the combination of S188H, V232A, E316M and D208R has a value that is about 6 times higher than that of WT, but it has the lowest value compared with other quadruple mutants.
[0131] Therefore, among the double, triple, quadruple, quintuple, hexaple...m-fold mutants (where m is a natural number greater than 2) containing the S188H mutation, it can be strongly inferred that mutants without the D208R mutation have a higher frameshift rate of GFP-positive cells and can obtain higher genome editing activity.
[0132] Furthermore, in the case of the double mutant D195K / V232A, which has a mutation at amino acid position 232 in addition to a mutation at amino acid position 195, the number of frameshifted GFP-positive cells was approximately 2-fold higher than that of D195K, representing a significant increase. Moreover, in the case of the triple mutant D195K / D208R / V232A, which has a mutation at amino acid position 208 in addition to mutations at amino acid positions 195 and 232, the number of frameshifted GFP-positive cells was approximately 2.5-fold higher than that of D195K, representing a further increase.
[0133] Furthermore, in the example of the quadruple mutant I123H / D195K / D208R / V232A, which has a mutation at amino acid position 123 based on mutations at amino acid positions 195, 232, and 208, the number of frameshifted GFP-positive cells was approximately three-fold higher compared to D195K, indicating a further increase in the value. Additionally, these results strongly suggest that even in multiple mutants such as pentapex or hexaplex mutants containing the I123H / D195K / D208R / V232A mutation, the proportion of frameshifted GFP-positive cells is higher compared to WT. This demonstrates the usefulness of mutants with mutations at amino acid position 195, particularly mutants containing only the D195K mutation, and any multiple mutants such as double, triple, and quadruple mutants containing at least the D195K mutation. Therefore, we obtained the important insight that the proportion of frameshift GFP-positive cells was significantly increased in mutants with a mutation at amino acid position 195.
[0134] Thus, mutants exhibiting at least the S188H mutation and mutants exhibiting at least the D195K mutation are particularly useful. Additionally, as... Figure 6-4As shown in the graph of C, the percentage of frameshifted GFP-positive cells in these double mutants (S188H / D195K) containing mutations in both S188H and D195K is approximately 25%, as... Figure 6-4 As shown in C, higher values were obtained than those of mutants containing only the S188H mutation or mutants containing only the D195K mutation.
[0135] On the other hand, such as Figure 6-4 As shown in D, in triple mutants (S188H / V232A / E316M, etc.) that do not have a mutation at amino acid position 195 and contain at least mutations of S188H and V232A, the percentage of frameshifted GFP-positive cells is approximately 40% to 50%. This value is significantly higher than that in mutants with only a mutation of S188H or in double mutants containing mutations of both S188H and D195K, demonstrating their particular usefulness. Furthermore, as Figure 6-4 As shown in E, the proportion of frameshifted GFP-positive cells in triple mutants (S188H / V232A / E316M) containing mutations of S188H, V232A, and E316M, and in quadruple mutants (S188H / V232A / E316M / F48Y, etc.) containing mutations of at least S188H, V232A, and E316M, demonstrates their usefulness. In particular, in S188H / V232A / E316M / I23K, S188H / V232A / E316M / F48Y, S188H / V232A / E316M / K80R, S188H / V232A / E316M / I123H, S188H / V232A / E316M / E246M, S188H / V232A The values in the quadruple mutants / E316M / I262N, S188H / V232A / E316M / E354T, S188H / V232A / E316M / D364Q, and S188H / V232A / E316M / I388N are higher than those in the double mutants containing the mutations S188H and D195K, demonstrating their particular usefulness.
[0136] In addition, such as Figure 6-4 As shown in Figure D, in triple mutants (D195K / D208R / V232A, etc.) that do not have a mutation at amino acid position 188 and contain at least mutations of D195K and V232A, the percentage of frameshifted GFP-positive cells is approximately 30% to approximately 50% or higher. This value is significantly higher than that in mutants with only a D195K mutation or in double mutants containing mutations of both S188H and D195K, demonstrating their particular usefulness. Furthermore, as... Figure 6-4As shown in E, the proportion of frameshifted GFP-positive cells in triple mutants (D195K / D208R / V232A) containing mutations of D195K, V232A, and D208R, and in quadruple mutants (D195K / D208R / V232A / I123H, etc.) containing mutations of at least D195K, V232A, and D208R, demonstrates their usefulness.
[0137] Furthermore, in Figure 6-4 In E, the proportion of frameshifted GFP-positive cells in the quadruple mutant D195K / D208R / V232A / S188R, which contains a mutation at amino acid position 188, is approximately 25%, a similar level to the double mutant where both S188H and D195K are mutated. However, in other quadruple mutants containing the D195K / D208R / V232A mutation, such as... Figure 6-4 As shown to the right of E, the percentage of frame-shifted GFP-positive cells is high, ranging from approximately 35% to approximately 50%. This demonstrates the usefulness of quadruple mutants with a mutation at amino acid position 195 but no mutation at amino acid position 188, particularly quadruple mutants with a mutation at least D195K / D208R / V232A without a mutation at amino acid position 188. Furthermore, not only quadruple mutants, but also quintuple and hexaple mutants with a mutation at least D195K / D208R / V232A without a mutation at amino acid position 188 are considered useful.
[0138] in addition, Figure 6-4 The triple mutants shown in D are all triple mutants with mutations occurring only at amino acid positions 188 and 195, and in these, the rate of frameshifted GFP-positive cells exceeds 30%. Therefore, it is considered that (1) triple mutants without mutation at amino acid position 195 and containing at least the S188H / V232A mutation, and (2) triple mutants without mutation at amino acid position 188 and containing at least the D195K / V232A mutation, are both useful. In addition, in (1) and (2), it is considered that not only triple mutants, but also quadruple mutants, quintuple mutants, and any other multiple mutants are useful.
[0139] Based on the above, it was shown that mutating the protein at amino acid position 188 in the AsCas12f amino acid sequence significantly increased the rate of GFP-positive cells. Furthermore, it was shown that mutating the protein at amino acid position 195 in the AsCas12f amino acid sequence also significantly increased the rate of GFP-positive cells.
[0140] More specifically, such as Figure 6-4As shown on the left side of C, in each double mutant from S188H / I2Y to S188H / L337I, the percentage of frameshifted GFP-positive cells was significantly higher than that in the comparative example WT (approximately 3-4%). In particular, while the percentage of frameshifted GFP-positive cells in S188 / D208R was the lowest at approximately 6%, it was still significantly higher than that in WT. Therefore, it can be said that double mutants that satisfy the condition of a histidine substitution at amino acid position 188 are useful.
[0141] In addition, Figure 6-4 In the graph of C regarding S188H, the proportion of frameshifted GFP-positive cells was significantly higher in S188H / K80R, S188H / Y105T, S188H / I123H, S188H / D195K, S188H / V232A, S188H / E246M, S188H / E316M, and S188H / L337I compared to the case where only S188H was mutated. Therefore, these demonstrate that they are particularly useful dual mutants.
[0142] Secondly, such as Figure 6-4 As shown on the right side of C, in each double mutant from D195K / I22Y to D195K / L337I, the percentage of frameshifted GFP-positive cells was significantly higher than that in the comparative example WT (approximately 3-4%). In particular, while the percentage of frameshifted GFP-positive cells in D195K / Y105T was the lowest at approximately 16%, it was still significantly higher than that in WT. Therefore, it can be said that double mutants that satisfy the condition of a lysine change at amino acid position 195 are useful.
[0143] In addition, Figure 6-4 In the graph of C regarding D195K, the proportion of frameshifted GFP-positive cells was significantly higher in D195K / I2Y, D195K / N70Y, D195K / K80R, D195K / I123H, D195K / D208R, D195K / V232A, D195K / E246M, D195K / E316M, and D195K / L337I compared to the case where only S188H was mutated. Therefore, these demonstrate that they are particularly useful dual mutants.
[0144] Secondly, such as Figure 6-4As shown in Figure D, the percentage of frameshifted GFP-positive cells in each of the triple mutants S188H / V232A / I123H, S188H / V232A / E246M, S188H / V232A / E316M, and S188H / V232A / L337I was approximately 35% or higher, significantly higher than the value in WT (as a comparative example), and at the same or higher level as the S188H / V232A double mutant. In particular, the triple mutants S188H / V232A / E246M and S188H / V232A / E316M yielded higher values than the S188H / V232A double mutant, indicating that these triple mutants are useful. Furthermore, since this consistently yields a high percentage of frameshifted GFP-positive cells, it can be said that triple mutants satisfying the condition of including the S188H / V232A variant are particularly useful. Additionally, as... Figure 6-4 As shown on the left side of E, the proportion of frameshifted GFP-positive cells is also high in quadruple mutants that meet the condition of including the S188H / V232A variant. This demonstrates that not only triple mutants, but also n-fold mutants (where n is a natural number of 3 or more) that include at least the D195K / V232A variant are useful.
[0145] Furthermore, the proportion of frameshifting GFP-positive cells in each of the triple mutants D195K / V232A / N70Y, D195K / V232A / D206R, and D195K / V232A / E316M was approximately 30% or higher, significantly higher than the value in WT (as a comparative example), and at the same or higher level as the D195K / V232A double mutant. In particular, the triple mutants S188H / V232A / E246M and S188H / V232A / E316M showed higher values than the S188H / V232A double mutant, indicating that these triple mutants are useful. Moreover, since such consistently high proportions of frameshifting GFP-positive cells were obtained, it can be said that triple mutants satisfying the condition of including the D195K / V232A variant are particularly useful. Additionally, as... Figure 6-4 As shown on the right side of E, the proportion of frameshifted GFP-positive cells is also high in quadruple mutants that meet the condition of including the D195K / V232A variant. This demonstrates that not only triple mutants, but also n-fold mutants (where n is a natural number of 3 or more) that include at least the D195K / V232A variant are useful.
[0146] exist Figure 6-4In the quadruple mutant containing the S188H / V232A / E316M variant shown on the left side of E, the proportion of frameshifted GFP-positive cells was approximately 20% to approximately 50%, which is significantly higher than the value in the WT used as a comparative example. Furthermore, in Figure 6-4 Among the E variants, the proportion of frameshifting GFP-positive cells in the quadruple mutant S188H / V232A / E316M / D206, while the lowest at around 20%, is significantly higher than that in the WT variant. This consistent high proportion of frameshifting GFP-positive cells demonstrates the particular usefulness of quadruple mutants containing the S188H / V232A / E316M variant. Furthermore, among the quadruple mutants of S188H / V232A / E316M / F48Y, S188H / V232A / E316M / E246M, S188H / V232A / E316M / I262N, and S188H / V232A / E316M / I388N, the proportion of frameshifted GFP-positive cells was the same or greater than that of the triple mutant of S188H / V232A / E316M, indicating that these are particularly useful quadruple mutants. In addition, based on these findings, it can be strongly inferred that not only quadruple mutants containing at least the S188H / V232A / E316M variant, but also m-multiple mutants (m being a natural number greater than 4) containing at least the S188H / V232A / E316M variant are also useful.
[0147] exist Figure 6-4 In the quadruple mutant containing the D195K / D208R / V232A mutation shown on the right side of E, the proportion of frameshifted GFP-positive cells was approximately 30% to 50%, which was significantly higher than the value in the WT as a comparative example. Furthermore, in Figure 6-4In the E group, the proportion of frameshifting GFP-positive cells in the quadruple mutant D195K / D208R / V232A / S188R, while the lowest at around 30%, is clearly shown to be significantly higher than that in the WT group. Thus, quadruple mutants containing the D195K / D208R / V232A mutation are particularly useful because they consistently yield a high proportion of frameshifting GFP-positive cells. Furthermore, in the quadruple mutants D195K / D208R / V232A / I23K, D195K / D208R / V232A / K80R, D195K / D208R / V232A / Q93M, and D195K / D208R / V232A / I123H, the proportion of frameshifted GFP-positive cells became equal to or higher than that of the D195K / D208R / V232A triple mutant, demonstrating that these are particularly useful quadruple mutants. Additionally, based on these findings, not only quadruple mutants containing D195K / D208R / V232A, but also m-multiple mutants (where m is a natural number greater than 4) containing at least D195K / D208R / V232A are also presumed to be useful.
[0148] exist Figure 6-4 In the E group, the quadruple mutant with the combination of D195K / D208R / V232A and S188R shows a very high value, approximately 10-fold higher than WT, but a lower value compared to other quadruple mutants. Therefore, among quadruple, quintuple, hexaple, heptapule…m-fold mutants (m being a natural number greater than 4) containing the D195K / D208R / V232A mutation, it is speculated that the high value is obtained in mutants without the S188R mutation.
[0149] In addition, Figure 6-4 In addition to the double, triple, and quadruple mutants shown in C-E, it is speculated that high values are also obtained in triple, quadruple, pentaple, hexaple, heptapride, octaple...m-fold mutants (where m is a natural number greater than 3) that further incorporate amino acid substitutions from these mutants. Regarding these mutants, if the sequence identity with the amino acid sequence shown in SEQ ID NO: 1 becomes very low, it is possible that cleavage activity against the target site is lost. However, if the sequence identity is high to some extent, it is speculated that a high rate of GFP-positive cell translocation, similar to these mutants, results in high cleavage activity.
[0150] exist Figure 6-4In the example of C, for each of the double mutants having at least a substitution of histidine (H) at position 188 in the amino acid sequence shown in SEQ ID NO: 1 and the double mutant having at least a substitution of lysine (K) at position 195 in the amino acid sequence shown in SEQ ID NO: 1, it is preferred to have cleavage activity against the target site and sequence identity with the amino acid sequence shown in SEQ ID NO: 1 of, for example, 90% or more, preferably 95% or more, more preferably 98% or more.
[0151] Similarly, for Figure 6-4 The triple mutant shown in D and Figure 6-4 For each of the quadruple mutants shown in E, the mutant is preferably a mutant that has cleavage activity against the target site and has sequence identity of 90% or more, preferably 95% or more, more preferably 98% or more, with the amino acid sequence shown in SEQ ID NO: 1.
[0152] Furthermore, as mentioned above, among the double, triple, quadruple, quintuple, hexaple…m-fold mutants (m being a natural number greater than 2) containing the S188H mutation, mutants without the D208R mutation consistently exhibit higher genome editing activity. Therefore, even within these mutants with lower sequence identity to the amino acid sequence shown in SEQ ID NO: 1, high cleavage activity against the target site is observed. Mutants with lower sequence identity compared to the aforementioned mutants, such as those with sequence identity to the amino acid sequence shown in SEQ ID NO: 1 of, for example, 88% or higher, and possessing cleavage activity against the target site, are also considered suitable. Similarly, among the quadruple, quintuple, hexaple, heptaple…m-fold mutants (m being a natural number greater than 4) containing the D195K / D208R / V232A mutation, it is speculated that high values are obtained in mutants without the S188R mutation.
[0153] <Optimization of sgRNA for AsCas12f>
[0154] Figure 7-1 Figure A is an illustration of the sequencing of the sgRNA mutant (sgRNA_ΔS3-5_v1-v7). To remove stems 3 and 4 and preserve the original structure of the remaining regions, the length and sequence of the adapter region were modified. The adapter region is enclosed in a red dashed box.
[0155] Figure 7-1 B is an illustration representing the genome editing activity of sgRNA variants determined by flow cytometry. Figure 7-1C is a graph representing the expression of sgRNA quantified by qPCR using a common sequencing site (n=3 means ±SD). Figure 7-1 D is a graph representing the time elapsed for GFP loss induced by various Cas enzymes in HEK293T cells expressing d2EGFP (n=3 means ±SD). Each sgRNA was introduced via single-copy lentiviral infection. GFP loss induced by AsCas12f-HKRA in sgRNA_ΔS3-5_v7 was slightly lower than that induced by SpCas9, but faster than the loss rates observed in AsCas12a, enAsCas12a, and Ultra. Furthermore, in Figure 7-1 In D, (a) represents the result of WT+ΔS5, (b) represents the result of WT+ΔS3-5_v7, (c) represents the result of HKRA+ΔS5, (d) represents the result of WHKRA+ΔS3-5_v7, (e) represents the result of AsCas12a, (f) represents the result of enAsCas12a, (g) represents the result of Ultra, and (h) represents the result of SpCas9.
[0156] Figure 7-2 The E symbol is a simplified illustration of wild-type sgRNA. Figure 7-2 F is a schematic diagram of sgRNA_ΔS3-5_v7. Disordered regions are shown in gray. PK1, PK2, stem 1, and stem 2 form extensive interactions with AsCas12 proteins, but stems 3 and 4 form almost no interactions with proteins, and stem 5 is disordered. Therefore, sgRNA_ΔS3-5_v7 was designed by removing stems 3–5, which is enclosed in dashed lines and shown. The remaining regions retain their original structure.
[0157] Figure 7-3 The G in the figure represents the structure of the wild-type sgRNA scaffold. Figure 7-3 The H symbol represents the structure of the sgRNA_ΔS3-5_v7 scaffold. Disordered regions are represented by dashed lines. Figure 7-4 The "I" represents a magnified view of the PK1 region of the wild-type sgRNA. Figure 7-4 J represents a magnified view of the PK1 region of sgRNA_ΔS3-5_v7. Figure 7-4 The K in the domain represents the PK2 region of the wild-type sgRNA. Figure 7-4 L represents a magnified view of the stem 3 region of sgRNA_ΔS3-5_v7.
[0158] In the structure of this embodiment, stem 5 is completely disordered. Furthermore, it is shown that PK1, PK2, stem 1, and stem 2 extensively interact with the AsCas12f protein, while stems 3 and 4, exposed to the solvent, interact almost no with the protein.
[0159] Therefore, as Figure 7-1 As shown in Figure A, the sgRNA was further engineered by truncating these domains. To remove stems 3 and 4, seven sgRNA variants (sgRNA_ΔS3-5_v1 to sgRNA_ΔS3-5_v7) were designed, while the remaining regions retained their original structure. The results are shown in Figure A. Figure 7-1 As shown in B, all sgRNA variants showed increased genome editing activity compared to sgRNA_ΔS5, with sgRNA_ΔS3-5_v7 being the most effective.
[0160] like Figure 7-1 As shown in Figure C, the expression level of sgRNA_ΔS3-5_v7 in HEK293 cells was 4.5 times that of sgRNA_ΔS5. This demonstrates that shortening sgRNA increases expression levels and enhances genome editing activity. To further clarify the effect of sgRNA modification, GFP-targeting sgRNA_ΔS5 and sgRNA_ΔS3-5_v7 were introduced via 1-copy lentiviral infection, and the GFP deletion rate was evaluated. Figure 5-2 F, Figure 7-1 As shown in D, since sgRNA_ΔS3-5_v7 significantly improved the GFP deletion efficiency of both AsCas12f-WT and AsCas12f-HKRA compared with sgRNA_ΔS5, it suggests that sgRNA_ΔS3-5_v7 is effective under conditions where expression levels are limited, such as in vivo gene delivery.
[0161] <Structure of the optimized sgRNA AsCas12f mutant determined by cryo-electron microscopy>
[0162] Figure 8-1 A represents the structural model of the AsCas12f-sgRNA-target DNA (left), AsCas12f-YHAMsgRNA_ΔS3-5_v7-target DNA (center), and AsCas12f-HKRAsgRNA_ΔS3-5_v7-target DNA (right). Mutated residues are shown as a CPK model.
[0163] Figure 8-2 B represents the dimer interface of AsCas12f-YHAM. Figure 8-2 The C indicates the target DNA recognition of AsCas12f-YHAM. Figure 8-2 The 'D' indicates the hydrophobic interaction of AsCas12f-YHAM. Mutated residues are highlighted in red, indicated by the darker color in the figure. Figure 8-2 The "E" indicates the recognition of heteroduplex DNA target DNA by the AsCas12f-HKRA guide RNA. Figure 8-2 F in Figure 8 indicates target DNA recognition, and G in Figure 8 indicates sgRNA scaffold recognition. Mutated residues are highlighted in red and indicated by the darker color in the figure.
[0164] like Figure 8-1 A, Figures 2-3 to 2-6 The EL and explanatory diagrams representing the results of data collection, processing, model improvement, and validation. Figure 9 and Figure 10 As shown, to gain mechanistic insights into the enhanced DNA cleavage activity exhibited by the AsCas12f mutant, the structures of the AsCas12f-YHAM-sgRNA_ΔS3-5_v7-target DNA and AsCas12f-HKRA-sgRNA_ΔS3-5_v7-target DNA complexes were determined using cryo-electron microscopy at a total resolution of 2.9 Å. The results showed that the overall structures of the AsCas12f-YHAM and AsCas12f-HKRA variants were similar to those of wild-type AsCas12f, and the mutations introduced by the DMS experiment had no significant impact on the overall structure of the complexes. Furthermore, in Figure 8A, the introduced mutations are represented by cpk (distinguished by color according to atom type).
[0165] exist Figures 7-2 to 7-4 The diagram showing E, G, I, and K represents the original wild-type sgRNA. Correspondingly, in... Figures 7-2 to 7-4 F, H, J, and L show an illustrative diagram of the modified sgRNA variant, sgRNA__ΔS3-5_v7. As shown, sgRNA_ΔS3-5_v7 consists of a 20 nt guide fragment (G1–C20) and a 96 nt sgRNA scaffold (from A(-96) to C(-1)). In sgRNA_ΔS3-5_v7, stems 1, 2, and PK1 are maintained, while stems 3, 4, and 5 are missing, and it interacts extensively with AsCas12f.
[0166] like Figure 7-4 As shown in I and J, in particular, C(-6), corresponding to the inverted C(-103) of wild-type sgRNA, forms a base pair with G(-83), linking PK1 / stem 1 to form a continuous helix. As Figure 7-4 As shown in K and L, the stem 3 region, which corresponds to the wild-type PK2 region, maintains the G(-24):C(-8) and G(-23):C(-9) base pairs on the one hand (G(-122):C(-88) and G(-121):C(-87) in the wild-type), and on the other hand, G(-12) is inverted and stacked with Trp17. This is similar to the interaction between G(-90) and Trp17 in the wild-type. Unexpectedly, as Figure 7-4As shown in L, A(-10) and U(-22) form a non-classical base pair. In stem 3, the base pairs of G(-24):C(-8)~G(-23):C(-9) and C(-21):G(-14)~C(-20):G(-15) are linked, indicating that sgRNA_ΔS3-5_v7 is the most suitable sgRNA variant.
[0167] Furthermore, stems 1 through 5 are all stem-loops. By adjusting the sequence and length of the linker region connecting the missing parts, the optimal sgRNA variant was created. This confirmed that deleting a portion of the sgRNA in different Cas molecules increased genome editing activity. One reason for this increased activity is believed to be the increased transcriptional rate of the sgRNA in the introduced mammalian cells due to shorter sgRNA length, which facilitates complex formation with Cas. Specifically, given the structure of AsCas12f, stems 4 and 5 show almost no interaction with AsCas12f; therefore, a short sgRNA variant deleting this portion was created.
[0168] Figure 8-2 The diagrams B through G represent the introduced mutations. In Figure 8-2 In B, the F48Y mutation, denoted by Y48.1, exists at the dimer interface. By becoming Y, it interacts with another molecule, thereby enhancing dimer assembly. Figure 8-2 In C, S188H, represented by H188.1, interacts with the newly formed nucleic acid to create a hydrophobic interaction. Figure 8-2 In the D, E316M, represented by M316.1 in AsCas12f.1 and M316.2 in AsCas12f.2, becomes M through E because it is surrounded by hydrophobic residues, which contributes to the stability of the protein. Figure 8-2 In E, I123H / D195K, represented by H123.1 and K195.1, interact with RNA neoformation. Figure 8-2 In F, DR208.1 is used as the denoting symbol and... Figure 8-2 In G, denoted by R208.2, 208R interacts with newly formed DNA / RNA molecules. Through the formation of these new interactions and the resulting increase in stability, the genome editing activity of the variants increases.
[0169] In detail, such as Figure 8-2As shown in Figure B, in the AsCas12f-YHAM structure, Tyr48.1 (F48Y.1, referred to as Y48.1 in the figure), located at the dimer interface, forms hydrogen bonds with the main chain of G57.2 in addition to the hydrophobic interactions observed in wild-type AsCas12f, which is thought to promote more stable dimerization. Figure 8-2 As shown in Figure C, the side chain of His188.1 (S188H.1, referred to as H188.1 in the figure) is stabilized by Ile2 and interacts with the backbone phosphate groups between dT19 and dC18, demonstrating that the S188H mutation promotes the unwinding of the target DNA. Figure 8-2 As shown in D, Met316.1 and Met316.2 (E316M.1 and E316M.2, referred to as M316.1 and M316.2 in the figure) further form hydrophobic interactions with Thr239 and Ala241, which may improve the overall stability.
[0170] like Figure 8-2 As shown in Figure E, in the AsCas12f-HKRA structure, His123.1 (referred to as H123.1 in the figure) and Lys195.1 (referred to as K195.1 in the figure) provide hydrophobic and electrostatic interactions to the ribose portion of A5 and the backbone phosphate groups of A3 and A5, respectively, stabilizing the guide RNA-target DNA heteroduplex. Figure 8-2 As shown in Figure F, Arg208.1 (referred to as R208.1 in the figure) forms hydrogen bonds and hydrophobic interactions with dT19 and Ile2, respectively. This is the same as His188.1 in AsCas12f-YHAM, thereby promoting the initial heteroduplex formation. Furthermore, as Figure 8-2 As shown in G, another subunit (referred to as R208.2 in the figure), Arg208.2, also forms an electrostatic interaction with the phosphate backbone of G(-43) in the sgRNA scaffold, stabilizing the complex structure.
[0171] Although both AsCas12f-YHAM and AsCas12f-HKRA mutants contain the V232A mutation, the contribution of this mutation to the increased DNA cleavage activity of both mutants was not clearly understood in this structural analysis. These findings suggest that DMS-based engineering methods have a significant potential to generate highly active mutants that cannot be predicted solely from structural information.
[0172] <Characteristic Evaluation of AsCas12f Mutants in Human Cells>
[0173] Figure 11-1Figure A is an illustration of the PAM specificity assay, and Figure B is a graph representing the results. Wild-type (WT) almost never induces genome editing in any PAM, while the two mutants in this embodiment induced genome editing in TTRPAM. It should be noted that... Figure 11-1 The vertical axis of B represents the editing frequency (%), and the bars from the left to the right of the graph represent the values in WT (wild type), AsCas12f-YHAM, and AsCas12f-HKRA, respectively. Figure 11-2 C to E are curves representing the genome editing activity of various Cas. SpCas9 is the most commonly used highly active Cas9, AsCas12a is the most commonly used highly active Cas12a (Cpf1), and Ultra and enAsCas12a represent modified versions of AsCas12a, respectively.
[0174] These results show that, depending on the target gene, the AsCas12f mutants (AsCas12f-YHAM, AsCas12f-HKRA) of this embodiment sometimes exhibit higher activity than these commonly used Cas mutants, or exhibit equal or higher activity.
[0175] Figure 12-1 A is a graph representing the comparison of target editing efficiency in AsCas12f-YHAM and AsCas12f-HKRA. For example... Figure 11-1 A and Figure 12-1 As shown in Figure A, to evaluate the PAM specificity of the enAsCas12f mutant against a broad range of targets, a self-targeting library containing both sgRNA and its target sequence was developed. Indel formations induced by wild-type AsCas12f (referred to as AsCas12f for simplicity) as well as AsCas12f-YHAM and AsCas12f-HKRA were determined for 750 different spacer sequences with 16 NTTN PAMs. Figure 12-1 As shown in Figure B, based on the results of deep sequencing analysis, the three enzymes induce insertions and deletions in NTTR PAM, but not in NTTY (Y for T or C) PAM, indicating that they collectively recognize NTTR sequences as PAM. That is, it can be said that AsCas12f recognizes TTR (R for A or G) as PAM.
[0176] like Figure 12-1 As shown in B, based on the structure obtained from cryo-electron microscopy, the dT(-3) of the PAM double strand... ) and dT(-2 The nucleic acid bases of ) form a hydrophobic interaction with Tyr76.1, dA(-1 N6 of dG forms hydrogen bonds with His72.1. Modeling suggests that dG(-1) forms hydrogen bonds with His72.1. The O6 in ) forms hydrogen bonds with His72.1, providing a clear indication of the preference for NTTR PAM. For example... Figure 11-1 As shown in B, in NTTR PAM, AsCas12f induces an average of 3.0% insertions and deletions, while AsCas12f-YHAM and AsCas12f-HKRA induce an average of 40.8% and 44.7% insertions and deletions, respectively.
[0177] These results show that the enAsCas12f mutant exhibits significantly higher genome editing activity than AsCas12f across various target sequences using NTTR PAM. Next, using HEK 293T cells, the genome editing efficiency of the AsCas12f and enAsCas12f mutants (variants) was compared with that of the Cas12a and Cas12a mutants (UltraCas12a and enAsCas12a) at five target sites using NTTR PAM.
[0178] exist Figure 11-2 In Figure C, the bar charts for VEGFA, HEXA, etc., from left to right, represent the insertion / deletion frequencies (%) in AsCas12f, AsCas12f-YHAM, AsCas12f-HKRA, AsCas12a, UltraCas12a, and enAsCas12a, respectively. Among these five sites, AsCas12f, AsCas12f-YHAM, AsCas12f-HKRA, AsCas12a, UltraCas12a, and enAsCas12a generated insertions / deletions at average frequencies of 8.5%, 22.3%, 14.3%, 9.0%, 11.0%, and 11.1%, respectively. These results indicate that the enAsCas12f mutant can induce insertions / deletions with an efficiency equal to or greater than that of the Cas12a and Cas12a mutants; however, such efficiency in induction of insertions / deletions did not occur in AsCas12f. Furthermore, in... Figure 11-2 In D and E, the bar charts for each item, from left to right, represent the measurement results in AsCas12f, AsCas12f-YHAM, AsCas12f-HKRA, and SpCas9, respectively.
[0179] like Figure 11-2As shown in Figure D, the genome editing efficiency at eight target sites in HEK293T cells was compared with that of SpCas9. These eight target sites have the same spacer sequences as the 5'-TTTG PAM targeting AsCas12f or the NGG-3' PAM targeting SpCas9. As shown in the figure, AsCas13f, AsCas12f-YHAM, AsCas12f-HKRA, and SpCas9 generated insertions and deletions at the eight target sites at average frequencies of 1.7%, 14.8%, 16.9%, and 22.9%, respectively. Furthermore, the formation of insertions and deletions targeting three therapeutic targets was also measured in HEK 293T cells: PCSK9 and ANGPTL3 for atherosclerosis, and TTR for transthyretin amyloidosis.
[0180] like Figure 11-2 As shown in Figure E, in two of the three target sites, AsCas12f-YHAM and AsCas12f-HKRA exhibited higher insertion / deletion frequencies than SpCas9. Figure 12-2 C, D and Figure 12-3 As shown in Figure E, the same results were obtained at other target sites, in other human cell lines, in Huh-7 liver tumor cells or HT-1080 sarcoma cells, indicating that AsCas12f-YHAM and AsCas12f-HKRA efficiently induce insertions and deletions regardless of the target site or cell line.
[0181] Next, mismatch tolerance was investigated to study the specificity of the enAsCas12f mutant. Figure 13-1 Figure A is a graph showing the effect of mismatches in the Split-GFP-reframing assay using TTR sequences on EGFP activation via AsCas12f-YHAM and AsCas12f-HKRA (n=3). The leftmost "-" in the graph indicates data when using a standard guide RNA that is perfectly complementary to the target DNA. "M01", "M02", ... indicate the positions where mismatches are introduced between the guide RNA and the target DNA, i.e., the positions of the incorrect guide sequences that prevent base pair formation. Therefore, "M01" means a mismatch is introduced at the first position of the guide RNA, and "M01 / 02" means a double mismatch guide RNA with mismatches at the first and second positions. Figure 13-1 In A, the vertical axis represents the insertion / missing efficiency (%). In the respective bar charts M01, M02, M03, etc., the left bar chart represents the value in AsCas12f-YHAM, and the right bar chart represents the value in AsCas12f-HKRA.
[0182] like Figure 13-1 As shown in A, similar to what has been observed in Cas12a or other Cas12fs, both enAsCas12f mutants exhibit broad tolerance to single mismatches except for the distal PAM region (positions 16-20), but tolerance to double mismatches is negligible. Furthermore, the genome-wide specificity of AsCas12f-HKRA and SpCas9 at nine target sites was investigated using GUIDE-seq (unbiased identification of double-strand breaks based on genome-wide sequencing). Figure 13-1 B is about in relation to Figure 11-2 A diagram illustrating the number of off-target sites of AsCas12f-HKRA and SpCas9 detected in the guide sequences of the same phenotypic sites in D. Figure 13-2 C-1, Figure 13-3 C-2, Figure 13-4 C-3 are illustrations of off-target sites identified by the guide sequence.
[0183] like Figure 13-1 B, Figure 13-2 C-1, Figure 13-3 C-2, Figure 13-4 As shown in C-3, AsCas12f-HKRA and SpCas9 have the same number of off-target sites, and AsCas12f-HKRA allows mismatches in the distal 5-nt region of PAM, consistent with mismatch experiments. These results indicate that the enAsCas12f mutant with sgRNA_ΔS3-5_v7, despite its extremely compact size, exhibits equivalent or higher genome editing activity and specificity compared to SpCas9 and Cas12a.
[0184] <Therapeutic Potential Through AsCas12f Genome Editing>
[0185] Figures 14-1 to 14-2 An illustrative diagram illustrating the creation of artificial pluripotent stem cells (iPSCs) for a muscular dystrophy model and their differentiation into cardiomyocytes (iPSC-CM). Figure 14-1 A represents the complete deletion of the DMD exon, confirmed by PCR. Figure 14-1 B is an illustration of the genome sequence of the DMD exon 44 deletion-iPS cell line. Figure 14-2 Figure C represents the flow cytometry analysis of α-actin-mCherry 14 days after myocardial differentiation, showing that all DMD exon 44 deletion strains differentiated into cardiomyocytes efficiently without purification. Figure 14-2Figure D represents the Western blot analysis of AR and DMD iPSC cardiomyocytes (#2-16, 33, 66, and 84). Arrows in the figure indicate immunoreactive bands of dystrophin, and asterisks indicate nonspecific bands. Figure 14-3 E is a graph representing a comparison of luciferase activity between dAsCas12f WT and dAsCas12f-HKRA (n=3, mean ± SD). Figure 14-3 F is a schematic diagram illustrating plasmid constructs of dAsCas12f with VP64 or VPR bound to different sites. HEXA-targeted gRNA and plasmids were transduced into Huh-7 cells stably expressing luciferase under the control of a minimal CMV promoter with a HEXA gRNA binding site. Figure 14-3 G represents the increase in luciferase activity in cells in which the plasmid vector was introduced (n=3, mean ± SD).
[0186] like Figures 14-1 to 14-2 As shown in Figures A through D, to evaluate the potential for therapeutic applications, artificial pluripotent stem cells (iPSCs) lacking exon 44 (hDMDΔEx44), a major variant of Duchenne muscular dystrophy (DMD), were created and differentiated into cardiomyocytes (iPSC-CM). The loss of dystrophin causes myopathy and cardiomyopathy, and investigations are underway to restore DMD proteins by hopping exon 45.
[0187] Figure 15-1 A to F and Figure 15-2 G to J are illustrative diagrams illustrating the use of mutants for treatment and knock-in. In detail, Figure 15-1 A through C are illustrations of treatments using the AsCas12f mutant targeting dystrophin (iPSCs). Figure 15-1 Figures D-F are illustrative diagrams illustrating treatment using the AsCas12f mutant targeting transthyretin (administered in mice using AAV). Figure 15-1 In E, the vertical axis represents plasma transthyretin (μg / ml), and the horizontal axis represents the number of days (weeks) elapsed. Additionally, (a) represents 3.0 × 10⁻⁶. 11 vg's AsCas12f, (b) represents 1.0 × 10 12 vg's AsCas12f, (c) represents 3.0 × 10 11 The HKRA and (d) values in vg represent 1.0 × 10⁻⁶. 12 The value of plasma transthyretin in the HKRA of vg. Additionally, in Figure 15-1In F, the vertical axis represents the insertion / deletion frequency (%), and (a) on the horizontal axis represents 3.0 × 10⁻⁶. 11 vg's AsCas12f, (b) represents 1.0 × 10 12 vg's AsCas12f, (c) represents 3.0 × 10 11 The HKRA and (d) values in vg represent 1.0 × 10⁻⁶. 12 vg's HKR. Figure 15-2 The G-I indicates the knock-in of GFP / FIX at the Alb locus using the AsCas12f mutant (mammalian cells). Furthermore, in Figure 15-2 In Figure I, the vertical axis represents the percentage of EGFP-positive cells (%), and in Figure J, the vertical axis represents the FIX:C percentage (%), indicating an increase in plasma factor IX activity (FIX:C).
[0188] like Figure 15-1 As shown in Figure A, the results of expressing AsCas12f-HKRA and sgRNA targeting the DMD gene using an integrated AAV serotype 6 vector in WT iPSC-CM demonstrate that AsCas12f-HKRA efficiently reduces the amount of dystrophin. Figure 15-1 As shown in B and C, AsCas12f-HKRA, in turn, partially restores the dystrophin of hDMDΔEx44-iPSC-CM in pairs with sgRNAs designed to cause exon 45 hopping. This confirms the possibility of DMD treatment based on AAV gene delivery.
[0189] Next, to investigate the therapeutic potential of AsCas12f in vivo, AsCas12f-HKRA was applied to genome editing in mouse livers. Figure 15-1 As shown in Figure D, an AAV vector encoding AsCas12f (or AsCas12f-HKRA) under the HCRhAAT promoter and a promoter-driven sgRNA targeting the TTR gene of transthyretin amyloidosis, operating under the control of the U6 promoter, were constructed. 3 × 10⁻⁶ AAV vectors were injected into 7-week-old mice. 11 vg or 1×10 12 The liver-oriented AAV serotype 8 of vg was used to evaluate the expression level of plasma transthyretin protein. For example... Figure 15-1 As shown in Figure E, AsCas12f failed to regulate plasma transthyretin levels, but AsCas12f-HKRA decreased in a dose-dependent manner. Eight weeks after AAV-based gene delivery, target loci were analyzed using target amplicon deep sequencing. Results showed that, through 1×10⁻⁶... 12vg's AsCas12f-HKRA achieved a high edit rate (66.3%). Figure 15-1 (F).
[0190] Next, the in vivo knock-in efficiency of the AsCas12f system, which inserts the EGFP gene into the mAlb 3'UTR, was evaluated. Knock-in at the Alb locus is a research platform for the ectopic production of therapeutic proteins, including coagulation factors and lysosomal enzymes, from the liver. Two AAV vectors were prepared: one expressing AsCas12f or AsCas12f-HKRA and an sgRNA targeting the mAlb 3'UTR, and the other providing a donor template for knocking in EGFP only in the DSB via homology-directed repair (HDR). Figure 15-2 The two AAV vectors were intraperitoneally injected into C57BL / 6 wild-type newborn mice, and EGFP expression in the liver was evaluated using immunofluorescence microscopy 4 weeks after vector injection.
[0191] The result, such as Figure 15-2 As shown in H and I, a significant increase in EGFP-positive hepatocytes was confirmed by injecting an AAV vector containing AsCas12f-HKRA. Next, the EGFP gene in the donor vector was replaced with coagulation factor IX (F9) cDNA with a Padua mutation, and this was injected into newborn hemophilia B mice (F9-deficient mice). Figure 15-2 As shown in Figure J, knocking in the F9 gene at the Alb locus of AsCas12f-HKRA significantly increased plasma coagulation factor IX (FIX) activity beyond the therapeutic range compared to donor-only cases, but such an increase was not observed in AsCas12f. These data suggest that AsCas12f mutants with optimal sgRNA can be used for in vivo gene therapy, such as for hemophilia.
[0192] <Applications of the compact enAsCas12f>
[0193] Due to its small size, the AsCas12f gene can be packaged together with multiple sgRNAs or large partner genes into a single AAV vector, making it possible to apply genome editing therapy that was previously impossible with genome editing tools. Figure 16-1 Figures A through F illustrate the use of FIX knock-in (in mice) at the Alb locus of the AsCas12f mutant. Figure 16-2Figures G through J illustrate gene transcriptional activation (CRISPR) using the AsCas12f mutant (mammalian cells). As shown, a significant increase in transcriptional activity was achieved by simultaneously introducing MS2.
[0194] like Figure 16-1 As shown in Figure A, a single AAV vector encoding AsCas12f (or AsCas12f-HKRA) was designed to insert F9 cDNA with the Padua mutation into the Alb3'UTR locus. This AsCas12f (or AsCas12f-HKRA) operates under the conditions of a liver-tropic Ttr promoter, a donor sequence, and an sgRNA that functions under the control of the U6 promoter. Figure 16-1 As shown in B and C, when a single AAV serotype 8 vector was injected into neonatal hemophilia B mice, the plasma FIX activity (FIX:C) and antigen (FIX:Ag) were significantly increased by the AAV vector with AsCas12f-HKRA, but not by the AAV vector with AsCas12f. Figure 16-1 As shown in Figures D and E, consistent with these results, coagulation time assessed by activated partial thrombin time (APTT) and F9 mRNA expression levels assessed by quantitative RT-PCR were significantly improved. Figure 16-1 As shown in F, amylose gel analysis of the mRNA PCR fragments revealed insertion via non-homologous end-joining (NHEJ), but cDNA insertion via HDR also occurred.
[0195] Finally, as Figure 16-2 As shown in G, to investigate the usefulness of enAsCas12f in epigenome editing, transcriptional activation assays were performed using Huh-7 cells stably expressing luciferase driven by a minimal CMV promoter with two HEXA gRNA recognition sites. To enhance the transcriptional activity of enAsCas12f, sgRNAs with MS2 aptamers inserted into their stem-loop were created. Figure 16-2 As shown in G, a plasmid expressing AsCas12f-HKRA (dead AsCas12f-HKRA, hereinafter referred to as dAsCas12f-HKRA), which is formed by conjugating VP64 and MS2 fusion activator (MS2-p65-HSF1) and eliminating DNA cleavage activity, and an engineered sgRNA targeting HEXA were introduced into Huh-7 cells.
[0196] like Figure 16-2 H and Figure 14-3 As shown in Figure E, dAsCs12f-HKRA combined with sgRNA containing the MS2 aptamer significantly enhanced luciferase expression. Figure 14-3 As shown in F and G, although direct binding of VPRs (VP64, p65, Rta) does not enhance transcription via MS2-p65-HSF1, binding at the terminal VP64 of dAsCas12f-HKRA is most effective for transcriptional activation. Furthermore, as... Figure 16-2 As shown in Figures I and J, the results of introducing a single AAV serotype 6 vector encoding JVP64, MS2-p65-HSF1, and dAsCas12f-HKRA conjugated with sgRNA into Huh-7 cells confirmed a significant, dose-dependent increase in luciferase expression. Based on these results, enAsCas12f can be used as a transcriptional activation tool.
[0197] Figure 17 A represents the simultaneous intravenous administration of an AAV8 vector expressing luciferase driven by a minimal CMV promoter with two HEXA gRNA recognition sites to wild-type mice. Figure 16-2 The diagram illustrates an example of an AAV8-type carrier with the structure shown in Figure I. This figure converts a color heatmap of increasing count values as it changes from blue (50 counts) to red (200 counts) into grayscale. Figure 17 In Figure A, although it is a grayscale, compared to the four mice in the AsCas12fWT (left), the first mouse from the left in the AsCas12fHKRA (right) shows the presence of a ring-shaped area represented by a count value of approximately 50-100 in dark gray, a ring-shaped area represented by a count value of approximately 100-170 in light gray within it, and a small area represented by a count value of approximately 170-200 in slightly darker gray within it. In the second mouse from the left in the figure, areas represented by a count value of approximately 50-100 in dark gray and areas represented by a count value of approximately 100-170 in light gray are scattered throughout. In the third and fourth mice from the left in the figure, a ring-shaped area with a count value of about 50 to 100 is shown, a ring-shaped area with a count value of about 100 to 170 is shown in light gray inside, and a wide closed area with a count value of about 170 to 200 is shown in slightly darker gray inside.
[0198] Figure 17B is a graph representing the ROI values (vertical axis) in WT and HKRA. Additionally, as shown in the figure, luciferase activity in Cas12f-HKRA increased approximately 10-fold compared to wild-type AsCas12f. The figure shows P = 0.0474 (since p < 0.05), indicating statistical significance.
[0199] Considering this result, it is believed that the transcription of endogenous genes can also be increased by about 10 times, for example, by increasing the transcription of myosin, a protein in fetal muscular dystrophy, which is also believed to be able to treat Duchenne muscular dystrophy.
[0200] It should be noted that, in Figure 18 The diagram illustrates the nucleic acid sequence used for structural analysis in this embodiment. Figure 19 The diagram illustrates nucleic acid sequences used for structural analysis in association with the STAR method. Figures 20-1 to 20-4 The diagram illustrates primers and constructs used for genome editing in association with the STAR method. Figures 21-1 to 21-7 The diagram illustrates key resources in the STAR method. These are examples and do not limit the invention. Additionally, in Figures 19 to 20-4 In the diagram, 3A~3B and 3C~3F correspond to respectively Figure 5-1 A~B and Figure 5-2 C to F, 5A to 5B, and 5C to 5E correspond to respectively Figure 11-1 A~B and Figure 11-2 C~E, 6A~6F and 6G~6J correspond to respectively Figure 15-1 A to F and Figure 15-2 G to J. Additionally, S4A to S4D, S4E to S4F, S4G to S4H, and S4I to S4L correspond to respectively Figure 7-1 A~D, E~F, G~H, I~L.
[0201] As described above, by applying Deep Mutation Scanning (DMS) technology to CRISPR-Cas effectors, a compact AsCas12f mutant with enhanced activity was successfully prepared. The DMS approach, by combining comprehensive protein mutation introduction and functional screening with deep sequencing, enables the evaluation of the effects of thousands of mutations in a single experiment. One typical application of DMS is the introduction of libraries into yeast surface display systems and the evaluation of binding affinity for ligands, including antibodies and viral glycoproteins. In this embodiment, a yeast screening system was applied to mammalian cell-based screening to evaluate genome editing activity. All libraries covering 20 single amino acid substitutions at various positions along the full-length (422 residues) AsCas12f sequence were constructed, and over 200 effective mutations at various positions were identified. A structural perspective is advantageous for efficiently exploring effective combinations of these mutants. As shown in this embodiment, in the absence of experimentally determined structures, structural prediction models such as AlphaFold can be useful for efficiently selecting mutations identified by DMS.
[0202] Next, an example of applying the AsCas12f mutant to genome editing in plants is shown. It should be noted that the I123Y / D195K / D208R / V232A mutant was used as the AsCas12f mutant.
[0203] exist Figure 22 The image shows tobacco Benzovia henryi (Benzovia henryi) using a viral vector expressing the AsCas12f mutant. Nicotiana benthamiana A summary of genome editing for SpCas9. As shown in the figure, in this example, the target site was set as the PDS gene, and the target gene was knocked out using a viral vector (PVX) expressing the AsCas12f mutant. For each of the SpCas9 and AsCas12f mutants, the viral vector was infected with Nicotiana benthamiana (Nicotiana benthamiana) using an agro-infection method. Nicotiana benthamiana The leaves of the leaves, for Figure 22 The inoculated leaf (IL) and its epiphyseal leaf are shown to compare genome editing efficiency. This comparison visualizes genome-edited cells when the PDS gene is knocked out in all alleles, resulting in cell whitening. In the figure, the IL represents the inoculated leaf, and its epiphyseal leaves are sequentially labeled L1, L2, L3, L4… from the side closest to the inoculated leaf IL.
[0204] The visualization results of the genome-edited cells as described above are shown below. Figure 23 (a), (b). Figure 23 (a) shows the results using the PVX-AsCas12f mutant. Figure 23(b) shows PVX-SpCas9. Regarding the epiphyseal L2, for the PVX-AsCas12f mutant, as Figure 23 As shown in (a), it is clear that the leaves' cells are leukoplakia. On the other hand, as... Figure 23 As shown in (b), albinism was barely observed to the naked eye in PVX-SpCas9. Furthermore, in the PVX-AsCas12f mutant, albinism was clearly observed in L3–L7, although not as pronounced as in L2. Additionally, a tendency was shown that the closer to the inoculated leaf IL, the stronger the albinism. Moreover, albinism was also observed at the leaf tip in L8. On the other hand, no clear albinism was observed in any of the L2–L8 mutants of PVX-SpCas9.
[0205] exist Figure 24 (a) shows the NGS analysis results using the PVX-AsCas12f mutant and PVX-SpCas9, respectively. As shown in the figure, in the NGS analysis results using the PVX-AsCas12f mutant, the mutation rate was approximately 70% in both the inoculated leaf IL and the epiphyseal leaf L2. In the epiphyseal leaf L2, although the mutation rate was small, it was still higher than that in the inoculated leaf IL. On the other hand, in the NGS analysis results using PVX-SpCas9, the mutation rate showed a high value of approximately 80% in the inoculated leaf IL, but in the epiphyseal leaf L2, the mutation rate became very low at 2%.
[0206] exist Figure 24 (b) shows the CAPS analysis results for the use of the PVX-AsCas12f mutant and PVX-SpCas9, respectively. As shown in the figure, when the PVX-AsCas12f mutant was used, a black or gray line appeared in the line indicated by the black triangle in either the inoculated leaf IL or the superior leaves L2–L6, indicating a mutation. On the other hand, when PVX-SpCas9 was used, a black line appeared in the inoculated leaf IL, but no line was detected in the superior leaves L2–L6, indicating no mutation.
[0207] This demonstrates the efficient genome editing in plants inoculated with a viral vector using the AsCas12f mutant. Furthermore, a significantly higher genome editing efficiency is shown compared to the use of PVX-SpCas9. Next, experimental results using the AsCas12f mutant in tomato genome editing are presented. Figures 22-24 In the example, Benedict's tobacco ( Nicotiana benthamiana ), but as shown below Figure 25 In examples (a) and (b), the same experiment was conducted using tomatoes. In this example, compared to Figure 22 Benedict's tobacco ( Nicotiana benthamiana Similarly, in the example above, the genome editing of tomatoes was carried out using a viral vector expressing the AsCas12f mutant, with the target site set as the PDS gene, and the target gene was knocked out using a viral vector (PVX) expressing the AsCas12f mutant.
[0208] exist Figure 25 (a) shows an outline of genome editing in tomatoes using a viral vector expressing the AsCas12f mutant. Figure 25 (b) shows the CAPS analysis results using the PVX-AsCas12f mutant. As shown in the figure, by using the PVX-AsCas12f mutant, a gray line appears in the black triangle in the inoculated leaf IL, indicating the presence of the mutation. In the superior leaves L1 and L2, although the gray line is lighter than that in the inoculated leaf IL, it indicates the introduction of the mutation.
[0209] Furthermore, the concentration of the gene was not significantly different between the upper leaves L1 and L2. Therefore, the genome editing efficiency was similar between the upper leaves L1 and L2, and no sharp decrease in genome editing efficiency was observed. Additionally, compared with... Figure 24 The shown is a type of tobacco called Benjamin Tobacco ( Nicotiana benthamiana Similarly, the example in the example suggests that in leaves higher than the upper leaf L2, genome editing efficiency is not significantly different from that in the upper leaves L1 and L2.
[0210] As mentioned above, the AsCas12f mutant is also useful for improving genome editing efficiency in plant cells. Furthermore, it can be seen that the AsCas12f mutant can not only improve genome editing efficiency in mammalian cells, but also be applied to any plant or animal (or organism) with cleavage activity to improve genome editing efficiency.
[0211] The present invention discloses the following embodiments.
[0212] The first embodiment is a protein having, in the amino acid sequence shown in SEQ ID NO: 1, a substitution of histidine at amino acid position 188, and further having one substitution selected from the following: a substitution of tyrosine at amino acid position 2, a substitution of tyrosine at amino acid position 70, a substitution of arginine at amino acid position 80, a substitution of threonine at amino acid position 105, a substitution of histidine at amino acid position 123, a substitution of lysine at amino acid position 195, a substitution of arginine at amino acid position 208, a substitution of alanine at amino acid position 232, a substitution of methionine at amino acid position 246, a substitution of methionine at amino acid position 316, or a substitution of isoleucine at amino acid position 337, having a sequence identity of more than 90% with the amino acid sequence shown in SEQ ID NO: 1, and having cleavage activity against a target site of DNA.
[0213] In the first embodiment, a more preferred embodiment is a double mutant comprising a substitution of histidine at amino acid position 188 and a substitution selected as described above.
[0214] The second embodiment is the same as the first embodiment, which is a protein whose amino acid number 208 has not been replaced.
[0215] The third embodiment is a protein having, in the amino acid sequence shown in SEQ ID NO: 1, a substitution of histidine at amino acid position 188 and a substitution of alanine at amino acid position 232, and further having one of the following substitutions: a substitution of histidine at amino acid position 123, a substitution of methionine at amino acid position 246, a substitution of methionine at amino acid position 316, or a substitution of isoleucine at amino acid position 337, having a sequence identity of more than 90% with the amino acid sequence shown in SEQ ID NO: 1, and having cleavage activity against a target site of DNA.
[0216] In the third embodiment, a triple mutant comprising a substitution of histidine at amino acid position 188, a substitution of alanine at amino acid position 232, and a substitution selected as described above is preferred.
[0217] The fourth embodiment is a protein, which is described in SEQ ID NO. The amino acid sequence shown in NO: 1 contains: a substitution at amino acid position 188 to histidine, a substitution at amino acid position 232 to alanine, and a substitution at amino acid position 316 to methionine, further comprising one of the following substitutions: a substitution at amino acid position 2 to tyrosine, a substitution at amino acid position 23 to lysine, a substitution at amino acid position 48 to tyrosine, a substitution at amino acid position 70 to tyrosine, a substitution at amino acid position 80 to arginine, a substitution at amino acid position 93 to methionine, a substitution at amino acid position 105 to threonine, a substitution at amino acid position 123 to histidine, a substitution at amino acid position 208 to arginine, a substitution at amino acid position 246 to methionine, a substitution at amino acid position 262 to asparagine, a substitution at amino acid position 354 to threonine, a substitution at amino acid position 364 to glutamine, or a substitution at amino acid position 388 to asparagine, in accordance with the SEQ ID. The amino acid sequence shown in NO:1 has a sequence identity of over 90% and exhibits cleavage activity targeting DNA sites.
[0218] In the fourth embodiment, a more preferred embodiment is a quadruple mutant comprising a substitution at amino acid position 188 to histidine, a substitution at amino acid position 232 to alanine, a substitution at amino acid position 316 to methionine, and a substitution selected as described above.
[0219] The fifth embodiment is a protein having, in the amino acid sequence shown in SEQ ID NO: 1, a substitution of lysine at amino acid position 195, and further having one of the following substitutions: a substitution of tyrosine at amino acid position 2, a substitution of tyrosine at amino acid position 70, a substitution of arginine at amino acid position 80, a substitution of threonine at amino acid position 105, a substitution of histidine at amino acid position 123, a substitution of arginine at amino acid position 208, a substitution of alanine at amino acid position 232, a substitution of methionine at amino acid position 246, a substitution of methionine at amino acid position 316, or a substitution of isoleucine at amino acid position 337, having a sequence identity of more than 90% with the amino acid sequence shown in SEQ ID NO: 1, and having cleavage activity against a target site of DNA.
[0220] In the fifth embodiment, a more preferred embodiment is a double mutant comprising a substitution of amino acid number 195 to lysine and a substitution selected as described above.
[0221] The sixth embodiment is a protein having, in the amino acid sequence shown in SEQ ID NO: 1, a substitution at amino acid position 195 to lysine and a substitution at amino acid position 232 to alanine, and further having one of the following substitutions: a substitution at amino acid position 70 to tyrosine, a substitution at amino acid position 208 to arginine, or a substitution at amino acid position 316 to methionine, having a sequence identity of more than 90% with the amino acid sequence shown in SEQ ID NO: 1, and having cleavage activity against a target site of DNA.
[0222] In the sixth embodiment, a triple mutant is more preferably comprising: a substitution at amino acid position 195 to lysine, a substitution at amino acid position 232 to alanine, and a substitution selected as described above.
[0223] The seventh embodiment is a protein, which is described in SEQ ID NO. The amino acid sequence shown in NO: 1 contains: a substitution at amino acid position 195 to lysine, a substitution at amino acid position 208 to arginine, and a substitution at amino acid position 232 to alanine, further comprising a substitution selected from: a substitution at amino acid position 2 to tyrosine, a substitution at amino acid position 23 to lysine, a substitution at amino acid position 48 to tyrosine, a substitution at amino acid position 70 to tyrosine, a substitution at amino acid position 80 to arginine, a substitution at amino acid position 93 to methionine, a substitution at amino acid position 105 to threonine, a substitution at amino acid position 123 to histidine, a substitution at amino acid position 188 to arginine, a substitution at amino acid position 246 to methionine, a substitution at amino acid position 262 to asparagine, a substitution at amino acid position 354 to threonine, a substitution at amino acid position 364 to glutamine, or a substitution at amino acid position 388 to asparagine, and is consistent with the SEQ ID. The amino acid sequence shown in NO:1 has a sequence identity of over 90% and exhibits cleavage activity targeting DNA sites.
[0224] In the seventh embodiment, a more preferred embodiment is a quadruple mutant comprising: a substitution at amino acid position 195 to lysine, a substitution at amino acid position 208 to arginine, a substitution at amino acid position 232 to alanine, and a substitution selected as described above.
[0225] The eighth embodiment is the protein described in any one of the first embodiment and the third to seventh embodiments, wherein the protein has a sequence identity of 95% or more.
[0226] The ninth embodiment is the protein described in any one of the first embodiment and the third to seventh embodiments, wherein the protein has a sequence identity of 98% or more.
[0227] It should be noted that the amino acid sequence of AsCas12f is shown in SEQ ID NO: 1 of the attached sequence listing, and is represented by single letters as follows.
[0228] MIKVYRYEIV KPLDLDWKEF GTILRQLQQE TRFALNKATQ LAWEWMGFSS DYKDNHGEYPKSKDILGYTN VHGYAYHTIK TKAYRLNSGN LSQTIKRATD RFKAYQKEIL RGDMSIPSYK RDIPLDLIKENISVNRMNHG DYIASLSLLS NPAKQEMNVK RKISVIIIVR GAGKTIMDRI LSGEYQVSAS QIIHDDRKNKWYLNISYDFE PQTRVLDLNK IMGIDLGVAV AVYMAFQHTP ARYKLEGGEI ENFRRQVESR RISMLRQGKYAGGARGGHGR DKRIKPIEQL RDKIANFRDT TNHRYSRYIV DMAIKEGCGT IQMEDLTNIR DIGSRFLQNWTYYDLQQKII YKAEEAGIKV IKIDPQYTSQ RCSECGNIDS GNRIGQAIFK CRACGYEANA DYNAARNIAIPNIDKIIAES IK
Claims
1. A protein, characterized in that, In the amino acid sequence shown in SEQ ID NO: 1, It has: a substitution at amino acid position 188 to histidine, and Furthermore, it has one of the following permutations: The substitution of amino acid number 2 towards tyrosine, The substitution of amino acid position 70 with tyrosine, The substitution of amino acid 80 with arginine, The substitution of amino acid 105 with threonine, The substitution of histidine at amino acid position 123, The substitution of amino acid 195 with lysine, The substitution of amino acid 208 with arginine, The substitution of amino acid 232 with alanine, The substitution of amino acid 246 with methionine, The substitution of amino acid 316 with methionine, or The substitution at amino acid position 337 to isoleucine. The sequence identity with the amino acid sequence shown in SEQ ID NO: 1 is more than 90%, and It has cleavage activity targeting specific sites of DNA.
2. The protein according to claim 1, wherein, The amino acid number at position 208 was not replaced.
3. A protein, characterized in that, In the amino acid sequence shown in SEQ ID NO: 1, have: The substitution of amino acid position 188 to histidine, and The substitution of amino acid 232 to alanine, and Furthermore, it has one of the following permutations: The substitution of histidine at amino acid position 123, The substitution of amino acid 246 with methionine, The substitution of amino acid 316 with methionine, or The substitution at amino acid position 337 to isoleucine. The sequence identity with the amino acid sequence shown in SEQ ID NO: 1 is more than 90%, and It has cleavage activity targeting specific sites of DNA.
4. A protein, characterized in that, In the amino acid sequence shown in SEQ ID NO: 1, have: The substitution of histidine at amino acid position 188, The substitution of amino acid 232 to alanine, and The substitution of amino acid 316 with methionine, and Furthermore, it has one of the following permutations: The substitution of amino acid number 2 towards tyrosine, The substitution of amino acid 23 with lysine, The substitution of amino acid position 48 with tyrosine, The substitution of amino acid position 70 with tyrosine, The substitution of amino acid 80 with arginine, The substitution of amino acid 93 with methionine, The substitution of amino acid 105 with threonine, The substitution of histidine at amino acid position 123, The substitution of amino acid 208 with arginine, The substitution of amino acid 246 with methionine, The substitution of asparagine at amino acid position 262, The substitution of amino acid 354 with threonine, The substitution of amino acid number 364 with glutamine, or The substitution at amino acid position 388 of asparagine The sequence identity with the amino acid sequence shown in SEQ ID NO: 1 is more than 90%, and It has cleavage activity targeting specific sites of DNA.
5. A protein, characterized in that, In the amino acid sequence shown in SEQ ID NO: 1, It has the following characteristics: a substitution at amino acid position 195 to lysine, and Furthermore, it has one of the following permutations: The substitution of amino acid number 2 towards tyrosine, The substitution of amino acid position 70 with tyrosine, The substitution of amino acid 80 with arginine, The substitution of amino acid 105 with threonine, The substitution of histidine at amino acid position 123, The substitution of amino acid 208 with arginine, The substitution of amino acid 232 with alanine, The substitution of amino acid 246 with methionine, The substitution of amino acid 316 with methionine, or The substitution at amino acid position 337 to isoleucine. The sequence identity with the amino acid sequence shown in SEQ ID NO: 1 is more than 90%, and It has cleavage activity targeting specific sites of DNA.
6. A protein, characterized in that, In the amino acid sequence shown in SEQ ID NO: 1, have: The substitution of amino acid 195 with lysine, The substitution of amino acid 232 to alanine, and Furthermore, it has one of the following permutations: The substitution of amino acid position 70 with tyrosine, The substitution of arginine at amino acid position 208, or The substitution of amino acid 316 with methionine. The sequence identity with the amino acid sequence shown in SEQ ID NO: 1 is more than 90%, and It has cleavage activity targeting specific sites of DNA.
7. A protein, characterized in that, In the amino acid sequence shown in SEQ ID NO: 1, have: The substitution of amino acid 195 with lysine, The substitution of amino acid 208 towards arginine, and The substitution of amino acid 232 to alanine, and Furthermore, it has one of the following permutations: The substitution of amino acid number 2 towards tyrosine, The substitution of amino acid 23 with lysine, The substitution of amino acid position 48 with tyrosine, The substitution of amino acid position 70 with tyrosine, The substitution of amino acid 80 with arginine, The substitution of amino acid 93 with methionine, The substitution of amino acid 105 with threonine, The substitution of histidine at amino acid position 123, The substitution of amino acid 188 with arginine, The substitution of amino acid 246 with methionine, The substitution of asparagine at amino acid position 262, The substitution of amino acid 354 with threonine, The substitution of amino acid number 364 with glutamine, or The substitution at amino acid position 388 of asparagine The sequence identity with the amino acid sequence shown in SEQ ID NO: 1 is more than 90%, and It has cleavage activity targeting specific sites of DNA.
8. The protein according to any one of claims 1, 3 to 7, wherein the sequence identity is 95% or more.
9. The protein according to any one of claims 1, 3 to 7, wherein the sequence identity is 98% or more.