Compositions and methods for delivering nucleic acid base editing systems

JP7899276B2Active Publication Date: 2026-08-03BEAM THERAPEUTICS INC
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
BEAM THERAPEUTICS INC
Filing Date
2024-11-01
Publication Date
2026-08-03

Smart Images

  • Figure 0007899276000034
    Figure 0007899276000034
  • Figure 0007899276000035
    Figure 0007899276000035
  • Figure 0007899276000036
    Figure 0007899276000036
Patent Text Reader

Abstract

To provide a composition and a method for helping achievement of delivery of CRISPR / Cas9 component to a cell.SOLUTION: Provided are a composition and a method for delivering first and second polynucleotide severally encoding fragments of A→G base editor fusion protein containing one or more deaminase (e.g., adenosine deaminase) and nCas9, where the first polynucleotide encodes an N-terminal fragment of nCas9 fused to intein -N of a divided intein pair, and the second polynucleotide encodes a C-terminal fragment of nCas9 fused to intein -C of a divided intein pair. A method for delivering (e.g., AAV delivery) these fragments together with sgRNA cell is also provided, where these fragments are connected together by a divided intein system, by which a functional base editing system is reconstituted in a cell.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Cross-reference of related applications This application is based on U.S. Provisional Patent Application No. 62 / 728,703, filed on September 7, 2018, and December 13, 2018. This asserts the interests of U.S. Provisional Patent Application No. 62 / 779,404, filed on [date], and all of these. The contents of this document are incorporated herein by reference.

[0002] The discovery of Clustered Regularly-Interspaced Short Palindromic Repeats (CRISPR) It revolutionized the field of molecular biology. Much of this enthusiasm stems from the fact that CRISPR / Cas9 is used to treat human diseases. The focus is on the clinical potential of treating and editing the human genome. Disease-causing mutations are This can potentially be repaired using CRISPR or a CRISPR-based system. One challenge to achieving this is delivering the elements necessary for genome editing. For example, regarding CRISPR / Cas9, SpCas9 and sgRNA are encoded in a DNA plasmid vector. It can be delivered via adeno-associated virus (AAV). However, AAV is packaged Due to its low delivery capacity, it helps to achieve the delivery of CRISPR / Cas9 components to cells, and / or or other elements (e.g., polypeptide domains) to satisfy the desired gene editing purpose. , promoter, reporter, fluorescent tag, multiple sgRNA, or DNA template for HDR It is difficult to include things like (e.g., '.'). [Overview of the project]

[0003] In some embodiments, (a) a fusion protein comprising a deaminase and the N-terminal fragment of Cas9 The first polynucleotide that codes for quality, where the N-terminal fragment of Cas9 is the N-terminal of Cas9. A continuous sequence that starts at the end and ends at positions A292-G364 in the numbering of Cas9 in sequence number 2. Thus, the N-terminal fragment of Cas9 is fused to the split intein-N, forming the first polynucleus. (b) a second polynucleotide encoding the C-terminal fragment of Cas9, Here, the C-terminal fragment of Cas9 begins at position A292-G364 in the numbering of Cas9 in Sequence ID No. 2. It is a continuous sequence ending at the C-terminus of Cas9, and the C-terminal fragment of Cas9 is fused to the split intein-C. A composition comprising a second polynucleotide is provided herein.

[0004] In some embodiments, (a) a first polynucleotide encoding the N-terminal fragment of Cas9 Here, the N-terminal fragment of Cas9 starts at the N-terminus of Cas9 and is numbered in Sequence ID No. 2. It is a continuous sequence that terminates at positions A292-G364 of Cas9, and the N-terminal fragment of Cas9 is a partitioned integer. (b) The first polynucleotide fused to in-N, and the C-terminal fragment and deamina of Cas9. A second polynucleotide encoding a fusion protein containing -ase, where Cas9 The C-terminal fragment of Cas9 begins at positions A292-G364 in the numbering in Sequence ID No. 2, and is the C-terminal fragment of Cas9. It is a continuous sequence that terminates at the end, and the C-terminal fragment of Cas9 is fused to the split intein-C, the second Compositions comprising polynucleotides are provided herein.

[0005] In some embodiments, (a) a fusion protein comprising a deaminase and the N-terminal fragment of Cas9 A first polynucleotide encoding a quality, wherein the N-terminal fragment of Cas9 starts at the N-terminus of Cas9 and ends at positions A292-G364, F445-K483, or E565-T637 of Cas9 according to the numbering in SEQ ID NO: 2, and the N-terminal fragment of Cas9 is fused to split intein-N, the first polynucleotide, and (b) a second polynucleotide encoding a C-terminal fragment of Cas9, wherein the C-terminal fragment of Cas9 starts at positions A292-G364, F445-K483, or E565-T637 of Cas9 according to the numbering in SEQ ID NO: 2 and ends at the C-terminus of Cas9, and the N-terminal residue of the C-terminal fragment of Cas9 is Cys substituted with Ala, Ser, or Thr, and the C-terminal fragment of Cas9 is fused to split intein-C, a composition comprising the second polynucleotide is provided herein. starting at the N-terminus and ending at the continuous sequence of A292 - G364, F445 - K483, or E565 - T637 of Cas9 according to the numbering in SEQ ID NO: 2, and the N-terminal fragment of Cas9 is fused to split intein-N, and (b) a second polynucleotide encoding a C-terminal fragment of Cas9, wherein the C-terminal fragment of Cas9 starts at positions A292 - G364, F445 - K483, or E565 - T637 of Cas9 according to the numbering in SEQ ID NO: 2 and ends at the C-terminus of Cas9, and the N-terminal residue of the C-terminal fragment of Cas9 is Cys substituted with Ala, Ser, or Thr, and the C-terminal fragment of Cas9 is fused to split intein-C, A composition comprising the second polynucleotide is provided herein. starting at positions A292 - G364, F445 - K483, or E565 - T637 of Cas9 according to the numbering in SEQ ID NO: 2 and ending at the C-terminus of Cas9, and the N-terminal residue of the C-terminal fragment of Cas9 is Cys substituted with Ala, Ser, or Thr, and the C-terminal fragment of Cas9 is fused to split intein-C, A composition comprising the second polynucleotide is provided herein. In some embodiments, (a) a first polynucleotide encoding an N-terminal fragment of Cas9, wherein the N-terminal fragment of Cas9 starts at the N-terminus of Cas9

[0006] and ends at the continuous sequence of A292 - G364, F445 - K483, or E565 - T637 of Cas9 according to the numbering in SEQ ID NO: 2, and the N-terminal fragment of Cas9 is fused to split intein-N, the first polynucleotide, and ( b) a second polynucleotide encoding a fusion protein comprising a C-terminal fragment of Cas9 and a deaminase, wherein the C-terminal fragment of Cas9 starts at positions A292 - G364, F445 - K483, or E565 - T637 of Cas9 according to the numbering in SEQ ID NO: 2 and ends at the C-terminus of Cas9, and the N-terminal fragment of Cas9 is fused to split intein-N, the first polynucleotide, and ( b) a second polynucleotide encoding a fusion protein comprising a C-terminal fragment of Cas9 and a deaminase, wherein the C-terminal fragment of Cas9 starts at positions A292 - G364, F445 - K483, or E565 - T637 of Cas9 according to the numbering in SEQ ID NO: 2 and ends at the C-terminus of Cas9, and the N-terminal fragment of Cas9 is fused to split intein-N, the first polynucleotide, and ( b) a second polynucleotide encoding a fusion protein comprising a C-terminal fragment of Cas9 and a deaminase, wherein the C-terminal fragment of Cas9 starts at positions A292 - G364, F445 - K483, or E565 - T637 of Cas9 according to the numbering in SEQ ID NO: 2 and ends at the C-terminus of Cas9, and the N-terminal fragment of Cas9 is fused to split intein-N, the first polynucleotide, and ( a column, wherein the N-terminal residue of the C-terminal fragment of Cas9 is Cys with Ala, Ser, or Thr substituted, A composition comprising a second polynucleotide, wherein the C-terminal fragment of Cas9 is fused to split intein-C, is provided herein.

[0007] In some embodiments, the N-terminal fragment of Cas9 comprises amino acids numbered 302, 309, 312, 354, 455, 459, 462, 465, 471, 473, 576, 588, or 589 in SEQ ID NO: 2. In some embodiments, the C-terminal fragment of Cas9 or the N-terminal fragment of Cas9 comprises an Ala / Cys, Ser / Cys, or Thr / Cys mutation at a residue corresponding to amino acid S303, T310, T313, S355, A456, S460, A463, T466, S469, T472, T474, C574, S577, A589, or S590 numbered in SEQ ID NO: 2. In some embodiments, the composition further comprises single guide RNA (sgRNA) or a polynucleotide encoding the same. In some embodiments, the first and second polynucleotides are ligated. In one aspect, the first and second polynucleotides are expressed separately. In one aspect, the deaminase is adenosine deaminase. In one aspect, the deaminase is wild-type TadA or TadA7.10. In one aspect, the deaminase is a TadA dimer. In some embodiments, the TadA dimer comprises wild-type TadA and TadA 7.10. In one aspect, the fusion protein comprises a nuclear localization signal (NLS). In some embodiments, the N-terminal fragment of Cas9 or the C-terminal fragment of Cas9 is ligated to the NLS. In some embodiments In some embodiments, the TadA dimer comprises wild-type TadA and TadA 7.10. In one aspect, the fusion protein comprises a nuclear localization signal (NLS). In some embodiments, the N-terminal fragment of Cas9 or the C-terminal fragment of Cas9 is ligated to the NLS. In some embodiments​ In some embodiments, the NLS is a bipartite NLS. The fusion protein is linked to a base editor containing deaminase and SpCas9. It forms an protein. In one embodiment, the C-terminal fragment of Cas9 and the fusion protein are linked. This then forms a base editor protein containing deaminase and SpCas9. In that embodiment, SpCas9 is either nickase-active or catalytically inactive.

[0008] In some embodiments, the fusion protein disclosed herein and the N-terminal fragment of Cas9 A composition comprising the above is provided herein. In some embodiments, the following is disclosed herein. Compositions comprising a fusion protein and a C-terminal fragment of Cas9 are provided herein. In one embodiment, the N-terminal fragment of Cas9 or the C-terminal fragment of Cas9 and deaminase are linked to a linker. Thus, they are linked. In one embodiment, the linker is a peptide linker.

[0009] In some embodiments, the first and second polynucleotides disclosed herein are included A vector is provided herein. In one embodiment, the vector includes a promoter. In one embodiment, the promoter is a constitutive promoter. The active promoter is a CMV or CAG promoter. In one embodiment, the vector is Retrovirus vectors, adenovirus vectors, lentivirus vectors, herpes The selection is made from a group consisting of cystic virus vectors and adeno-associated virus vectors. In one embodiment, the vector is an adeno-associated virus vector.

[0010] In some embodiments, the compositions or materials disclosed herein Cells, including those containing cells, are provided herein. In one embodiment, the cells are mammalian cells. be.

[0011] In some embodiments, the Ala / Cys, Ser / Cys, or Thr / Cys mutations are referred to herein. A reconstituted A→G base editor protein is provided, containing a Cas9 domain. In some embodiments, the mutations occur in the amino acids S303, T310, T313, S355 of SpCas9. Corresponding to A456, S460, A463, T466, S469, T472, T474, C574, S577, A589, or S590. It is located at the residue.

[0012] In some embodiments, (a) Cas9 begins at the N-terminus and is numbered in sequence number 2. It is a continuous sequence that terminates at positions A292~G364 of as9 and is fused into the partitioned intein-N, Cas (b) The N-terminal fragment of 9, and Cas9 starting at positions A292-G364 in the numbering in Sequence ID No. 2. A continuous sequence terminating at the C-terminus of Cas9, which fuses with the split intein-C; this is the C-terminal fragment of Cas9. A composition comprising one or more polynucleotides encoding the above is provided herein.

[0013] In some embodiments, (a) Cas9 begins at the N-terminus and is numbered in sequence number 2. The amino acids of as9 are 302, 309, 312, 354, 455, 459, 462, 465, 471, 473, 576, 588, or 5 (b) In the numbering of sequence number 2, Cas9's 303, 310, 313, 355, 456, 460, 463, 466, 472, 4 A contiguous sequence starting at 74, 577, 589, or 590 and ending at the C-terminus of Cas9, and a partitioned integer It contains one or more polynucleotides that encode the C-terminal fragment of Cas9, which is fused to in-C. Compositions are provided herein.

[0014] In some embodiments, the N-terminal fragment of Cas9 or the C-terminal fragment of Cas9 is a nuclear localization signature. It is linked to the NLS (N-terminal segment). In some embodiments, the N-terminal segment of Cas9 and Cas9 Both C-terminal fragments are ligated to the NLS. In some embodiments, the NLS is a bipartite NLS. In some embodiments, the N-terminal fragment of Cas9 and the C-terminal fragment of Cas9 are ligated. This is used to form SpCa9. In some embodiments, SpCa9 has nickase activity. Alternatively, it is catalytically inactivated.

[0015] In some embodiments, the N-terminal fragment of Cas9 in (a) disclosed herein is (b A composition comprising the C-terminal fragment of Cas9 is provided herein.

[0016] In some embodiments, a vector comprising one or more polynucleotides disclosed herein A vector is provided herein. In one embodiment, the vector includes a promoter. In one embodiment, the promoter is a constitutive promoter. The promoter is a CMV or CAG promoter. In one embodiment, the vector is a let Rovirus vectors, adenovirus vectors, lentivirus vectors, herpesvirus Selected from the group consisting of Rus vectors and adeno-associated virus vectors. In this case, the vector is an adeno-associated virus vector.

[0017] In some embodiments, the compositions or materials disclosed herein Cells, including those containing cells, are provided herein. In one embodiment, the cells are mammalian cells. be.

[0018] In some embodiments, Cas9 variants containing Ala / Cys, Ser / Cys, or Thr / Cys mutations Ant polypeptides are provided herein. In some embodiments, Cys residues in amino acids 303, 310, 313, 355, 456, 460, 463, 466, 472, or 474 A Cas9 variant polypeptide containing this polypeptide is provided.

[0019] In some embodiments, this specification describes how to deliver a base editor system to cells. A method is provided which involves (a) the deaminase and the N-terminal fragment of Cas9 in cells. A first polynucleotide encoding a fusion protein comprising, where Cas9's N The terminal fragments begin at the N-terminus of Cas9 and correspond to positions A292-G364 of Cas9 in the numbering scheme in Sequence ID No. 2. It is a continuous sequence that terminates in a position, and the N-terminal fragment of Cas9 is fused to the split intein-N, the first (b) a polynucleotide, a second polynucleotide encoding the C-terminal fragment of Cas9, Here, the C-terminal fragment of Cas9 starts at position A292-G364 in the numbering in Sequence ID No. 2. Furthermore, it is a continuous sequence that terminates at the C-terminus of Cas9, and the C-terminal fragment of Cas9 is fused into the split intein-C. (c) a second polynucleotide, and (c) a single guide RNA (sgRNA) or it This involves contacting the polynucleotide with the nucleotide.

[0020] In some embodiments, this specification describes how to deliver a base editor system to cells. A method is provided which involves (a) a fusion protein containing the N-terminal fragment of Cas9 in cells. The first polynucleotide encoding, where the N-terminal fragment of Cas9 is the N-terminal A continuous sequence that starts at the end and ends at positions A292-G364 in the numbering of Cas9 in sequence number 2. The N-terminal fragment of Cas9 is fused to the split intein-N, the first polynucleotide, ( b) A second polynucleotide encoding the C-terminal fragment of Cas9 and deaminase, Thus, the C-terminal fragment of Cas9 starts at position A292-G364 in the numbering in Sequence ID No. 2. It is a continuous sequence that terminates at the C-terminus of Cas9, and the C-terminal fragment of Cas9 is fused to the split intein-C. (c) a second polynucleotide, and a single guide RNA (sgRNA) or code it This includes contacting the polynucleotide with the nucleotide.

[0021] In some embodiments, methods for delivering a base editor system to cells are revealed. Provided in the details, this method involves (a) melting cells with deaminase and the N-terminal fragment of Cas9. The first polynucleotide encoding the composite protein, where the N-terminal fragment of Cas9 is Starting from the N-terminus of Cas9, and numbered in Sequence ID No. 2 as Cas9 A292-G364, F445-K483 Alternatively, it is a continuous sequence terminating at positions E565~T637, and the N-terminal fragment of Cas9 is a split intein-N (b) The first polynucleotide, which is fused to (b) the second polynucleotide encoding the C-terminal fragment of Cas9. It is a creotide, and here the C-terminal fragment of Cas9 is A2 of Cas9 in the numbering in Sequence ID No. 2. A continuous distribution starting at positions 92-G364, F445-K483, or E565-T637 and ending at the C-terminus of Cas9. It is a series, and the N-terminal residue of the C-terminal fragment of Cas9 is a Cys, which is substituted with Ala, Ser, or Thr, and Ca The C-terminal fragment of s9 is fused to the split intein-C, a second polynucleotide, and (c) mono This involves contacting a single guide RNA (sgRNA) or a polynucleotide encoding it. .

[0022] In some embodiments, methods for delivering a base editor system to cells are revealed. Provided in the details, this method involves (a) a fusion protein containing the N-terminal fragment of Cas9 in cells. A first polynucleotide that starts at the N-terminus of Cas9, where the N-terminal fragment of Cas9 starts at the N-terminus of Cas9. Also, in the numbering of sequence number 2, Cas9 corresponds to A292-G364, F445-K483, or E565-T637. It is a continuous sequence that terminates at the position, and the N-terminal fragment of Cas9 is fused to the split intein-N. (b) The first polynucleotide, (b) the C-terminal fragment of Cas9 and the second polynucleotide encoding deaminase It is a nucleotide, and here the C-terminal fragment of Cas9 is numbered in SEQ ID NO: 2 A continuum starting at positions A292-G364, F445-K483, or E565-T637 and ending at the C-terminus of Cas9. The sequence is such that the N-terminal residue of the C-terminal fragment of Cas9 is Cys substituted with Ala, Ser, or Thr. The C-terminal fragment of Cas9 is fused to the split intein-C, a second polynucleotide, and (c) This includes contact with a single guide RNA (sgRNA) or a polynucleotide encoding it. nothing.

[0023] In one embodiment, the sgRNA is complementary to the target polynucleotide. Therefore, the target polynucleotide is present in the genome of an organism. In one embodiment, the organism is It is an animal, plant, or bacterium. In some embodiments, the first polynucleotide , the second polynucleotide, and / or the polynucleotides encoding them are vectors It comes into contact with cells via the vector. In one embodiment, the vector is a retroviral vector. - Adenovirus vectors, lentivirus vectors, herpesvirus vectors, Selected from the group consisting of adeno-associated virus vectors. In one embodiment, vector - is an adeno-associated virus vector. In some embodiments, the C-terminus of Cas9 The fragment or the N-terminal fragment of Cas9 is numbered as amino acids S303, T310, T313 in SEQ ID NO: 2. , S355, A456, S460, A463, T466, S469, T472, T474, C574, S577, A589, or S590 The corresponding residue contains an Ala / Cys, Ser / Cys, or Thr / Cys mutation. In some embodiments, In some cases, deaminase is adenosine deaminase. Deaminase is TadA or a variant thereof. In some embodiments, deaminase is It is wild-type TadA or Tad7.10. In one aspect, the deaminase is a TadA dimer. In some embodiments, the TadA dimer comprises wild-type TadA and TadA7.10. In some embodiments, the N-terminal or C-terminal fragment of Cas9 includes an NLS. In several embodiments, both the N-terminal and C-terminal fragments of Cas9 contain NLS. In some embodiments, the NLS is a two-part NLS. In some embodiments, The N-terminal and C-terminal fragments of Cas9 are ligated together to form SpCa9. In this context, SpCas9 either possesses nickase activity or is catalytically inactive.

[0024] In some embodiments, this specification describes polynucleotides encoding fusion proteins. A fusion protein is provided, in which the fusion protein comprises a deaminase and the N-terminal fragment of Cas9, and Cas The N-terminal fragment of 9 begins at the N-terminus of Cas9 and corresponds to the numbering A292-G364 of Cas9 in sequence number 2. It is a continuous sequence that terminates at the position, and the N-terminal fragment of Cas9 is fused to the split intein-N. In some embodiments, the polynucleotide encoding the fusion protein is used herein. Provided, here the fusion protein comprises deaminase and the N-terminal fragment of Cas9, Cas9 The N-terminal fragment of Cas9 begins at the N-terminus and is numbered F445-K483 in sequence number 2 of Cas9. It is a continuous sequence that terminates at the position, and the N-terminal fragment of Cas9 is fused to the split intein-N. In some embodiments, the polynucleotide encoding the fusion protein is used herein. Provided, here the fusion protein comprises deaminase and the N-terminal fragment of Cas9, and Cas9 The N-terminal fragment of Cas9 begins at the N-terminus and corresponds to the numbering E565~T637 of Cas9 in sequence number 2. It is a continuous sequence that terminates at the position, and the N-terminal fragment of Cas9 is fused to the split intein-N. In some embodiments, the polynucleotide encoding the fusion protein is used herein. Provided, here the fusion protein comprises deaminase and the C-terminal fragment of Cas9, and Cas9 The C-terminal fragment of Cas9 begins at positions A292-G364 in the numbering in Sequence ID No. 2, and is the C-terminal fragment of Cas9. It is a continuous sequence that terminates at the end, and the C-terminal fragment of Cas9 is fused to the split intein-C. In some embodiments, the polynucleotide encoding the fusion protein is used herein. Provided, here the fusion protein comprises deaminase and the C-terminal fragment of Cas9, and Cas9 The C-terminal fragment of Cas9 begins at position F445-K483 in the numbering in Sequence ID No. 2, and is the C of Cas9. It is a continuous sequence that terminates at the end, and the C-terminal fragment of Cas9 is fused to the split intein-C. In one embodiment, a polynucleotide encoding a fusion protein is provided herein, Here, the fusion protein contains a deaminase and the C-terminal fragment of Cas9, and the C-terminal fragment of Cas9 In the numbering of sequence number 2, it starts at positions E565-T637 of Cas9 and ends at the C-terminus of Cas9. In a certain manner, the Cas9 C-terminal fragment is fused to the split intein-C. Therefore, polynucleotides encoding fusion proteins are provided herein, and here, fusion The protein contains deaminase and the C-terminal fragment of Cas9, and the C-terminal fragment of Cas9 is sequence number In the numbering in 2, it starts at positions A292-G364, F445-K483, or E565-T637 in Cas9. It is a continuous sequence that terminates at the C-terminus of Cas9, and the N-terminal residues of the C-terminal fragment of Cas9 are Ala, Ser, Alternatively, it is a Cys with Thr substituted, and the C-terminal fragment of Cas9 is fused to the split intein C.

[0025] In some embodiments, the C-terminal or N-terminal fragment of Cas9 is Ala / Cys, S Includes er / Cys or Thr / Cys mutations. In some embodiments, the mutations are distributed In row number 2, the amino acids are numbered as S303, T310, T313, S355, A456, S460, A463, and T466. , located at residues corresponding to S469, T472, T474, C574, S577, A589, or S590. In a certain embodiment In this context, deaminase is adenosine deaminase. In one aspect, deaminase The enzyme is TadA or a variant thereof. In some embodiments, deaminase These are wild-type TadA or Tad7.10. In one embodiment, the fusion proteins are linked to each other. It contains two deaminases. In one aspect, the fusion protein is wild-type TadA It contains both TadA7.10. In one aspect, the fusion protein contains NLS. In one embodiment, the NLS is a bipartite NLS. The C-terminal fragment contains the amino acid sequence of SpCas9. In some embodiments, the N-terminal fragment of Cas9 The terminal fragment or C-terminal fragment of Cas9 contains one or more amino acids associated with reduced nuclease activity. Includes substitution.

[0026] In some aspects, amino acids 302, 309, 312, 354, 455, 459, and 46 are referred to herein. The N-terminal fragment of the Cas9 protein, including up to 2, 465, 471, or 473, is fused to the split intein-N. A combined product is provided. In some embodiments, the Cas9 protein is described herein. A C-terminal protein fragment is provided, where the N-terminal amino acid of the C-terminal fragment is amino acid 303. Cys substitution in 310, 313, 355, 456, 460, 463, 466, 472, or 474, and splitting in It is fused to tein-C. In some embodiments, as specified herein, A→G base editor fusion A polynucleotide encoding a fragment of the fusion protein is provided, and the fusion protein is one The above deaminase and the N-terminal fragment of Cas9 are included, and the N-terminal fragment is fused to split intein-N. In some embodiments, A→G base editor fusion proteins are used. A polynucleotide encoding a quality fragment is provided, and the fusion protein is one or more dear It contains a minase and a C-terminal fragment of Cas9, the C-terminal fragment being fused to a split intein-C. In some embodiments, the protein of an A→G base editor fusion protein A protein fragment is provided, and the fusion protein combines one or more deaminases with the N-terminal fragment of Cas9. The N-terminal fragment is fused to the split intein-N. In some embodiments, The specification provides a protein fragment of an A→G base editor fusion protein, and the fusion protein The protein contains one or more deaminases and the C-terminal fragment of Cas9, the C-terminal fragment being divided into It is integrated into Tein-C.

[0027] In some embodiments, A→G bases each containing one or more deaminases and Cas9 A set containing the first and second polynucleotides encoding the fragments of the editor fusion protein. A product is provided herein, in which a first polynucleotide is fused to split intein-N. The combined Cas9 N-terminal fragment is encoded, and the second polynucleotide is fused to the split intein-C. Encodes the C-terminal fragment of the combined Cas9. In some embodiments, one or more deaminers A composition comprising the N- and C-terminal fragments of an A→G base editor fusion protein containing Ze and Cas9. This is provided herein, where the N-terminal fragment is a SpCas9 fused to a split intein-N It contains fragments, and the C-terminal fragment contains the remainder of SpCas9 fused to the split intein-C.

[0028] In some embodiments, methods for delivering a base editor system to cells are described. Provided in the details, this method involves converting cells into A→G bases containing one or more deaminases and Cas9. The first and second polynucleotides each encode fragments of the editor fusion protein. This includes contacting the first polynucleotide, where the first polynucleotide is fused to the split intein-N. The second polynucleotide encodes the N-terminal fragment of Cas9, and is fused to the split intein-C. It encodes the C-terminal fragment of Cas9, and either the first or second polynucleotide. However, it encodes a single guide RNA. In some embodiments, a base editor system A method for delivering to cells is provided herein, and this method delivers cells to one or more dea N- and C-terminal fragments of an A→G base editor fusion protein containing minase and SpCas9, This involves contacting the N-terminal fragment with guide RNA, where the N-terminal fragment is divided into intein-N. It contains a fused SpCas9 fragment, and the C-terminal fragment is a fragment of SpCas9 fused to the split intein-C. Includes the remainder. In some embodiments, for editing target polynucleotides in cells A method is provided herein, which involves a cell containing one or more deaminases and Cas9. The first and second polynuclei each encode fragments of the A→G Base Editor fusion protein. Contact with rheotide (where the first polynucleotide is fused to the split intein-N) The combined Cas9 N-terminal fragment is encoded, and the second polynucleotide is fused to the split intein-C. Encoding the C-terminal fragment of the combined Cas9, and either the first or second polynucleotide is single (Encodes a single guide RNA), and the encoded protein and single guide RNA are further processed. This includes expressing the expression within the cell.

[0029] Other features and advantages of the present invention will become apparent from the detailed description and claims. It is likely.

[0030] [Definition] Unless otherwise defined, all technical and scientific terms used herein are defined as follows: This has meanings that are generally understood by those skilled in the art in which the present invention pertains. See below for reference. The literature provides general definitions of many terms used in this invention for those skilled in the art: Sing leton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); T he Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossa ry of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). used in this specification. When used, the following terms shall apply to them below unless otherwise specified. It has the meaning of [something].

[0031] "Adenosine deaminase" refers to the hydrolytic deamination of adenine or adenosine. This means a polypeptide or fragment thereof that can catalyze a substance. In one embodiment, The aminase or deaminase domain hydrolyzes adenosine to inosine. It catalyzes the deamination or hydrolysis of deoxyadenosine to deoxyinosine. It is adenosine deaminase. In one aspect, adenosine deaminase is deoxy It catalyzes the hydrolytic deamination of adenine or adenosine in siribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., genetically modified adenosine deaminases) Aminase (evolved adenosine deaminase) can originate from any organism, such as bacteria. It may be a deaminase or deaminase domain. It is a variant of a naturally occurring deaminase derived from a substance. In some embodiments, deaminase Alternatively, the deaminase domain does not exist in nature. For example, several states In this context, the deaminase or deaminase domain is a naturally occurring deaminase. against at least 50%, at least 55%, at least 60%, at least 65%, and at least 70% %, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, small At least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% It has the same identity as E. coli. In one embodiment, adenosine deaminase is, for example, E. coli, S. Aureus, S. typhi, S. putrefaciens, H. influenzae, or C. crescentus, etc. It is derived from bacteria. In one aspect, adenosine deaminase is TadA deaminase. In one embodiment, TadA deaminase is E. coli TadA(ecTadA) deaminase or This is a fragment of it.

[0032] For example, truncated ecTadA has one or more N-terminal axolars compared to full-length ecTadA. Mino acids may be absent. In some embodiments, the truncated ecTadA is full-length ec Compared to TadA, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, Alternatively, 20 N-terminal amino acid residues may be deleted. In some embodiments, The packed ecTadA is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 1 compared to the full-length ecTadA. The C-terminal amino acid residues 4, 15, 6, 17, 18, 19, or 20 may be deleted. In some embodiments, ecTadA deaminase does not contain N-terminal methionine. In the embodiment, the TadA deaminase is an N-terminally truncated TadA. Specific Embodiments In this context, TadA is incorporated herein by reference in its entirety by PCT / US2017 / 045381. It is one of the TadA listed.

[0033] In certain embodiments, adenosine deaminase comprises the following amino acid sequence: MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPT AHAEIMALRQGGLVMQNYRLIDAT LYVTLEPCVMCAGAMIHSRIGRVVFGARDAKT GAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKA QKKAQSSTD This is called the "TadA reference array".

[0034] In one embodiment, TadA deaminase is full-length E. coli TadA deaminase. For example, In certain embodiments, adenosine deaminase comprises the following amino acid sequence: MRRAFITGVFFLSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEG WNRPIGRHDPTAHAEIMALRQGGL VMQNYRLIDATLYVTLEPCVMCAGAMIHSRIG RVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSD FFRMRRQEI KAQKKAQSSTD

[0035] However, further adenosine deaminases useful in this application are obvious to those skilled in the art. It should be understood that this is within the scope of this disclosure. For example, adenosine Aminase may be a homolog of adenosine deaminase (AD AT) that acts on tRNA. Exemplary AD AT homologs include, but are not limited to, the following:

[0036] Staphylococcus aureus TadA: MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAH AEHIAIERAAKVLGSWRLEGCT LYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCS GS LMNLLQQS NFNHRAIVDKG VLKE AC S TLLTTFFKNL RANKKS TN

[0037] Bacillus subtilis TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEML VIDEACKALGTWRLEGATLYVT LEPCPMCAGAVVLSRVEKVVFGAFDPKGGC S GTLMN LLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAAR KNLSE

[0038] Salmonella typhimurium (S. typhimurium) TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEG WNRPIGRHDPTAHAEIMALRQGGL VLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIG RVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSD FFRMRRQEIK ALKKADRAEGAGPAV

[0039] Shewanella putrefaciens (S. putrefaciens) TadA: MDE YWMQVAMQM AEKAEAAGE VPVGA VLVKDGQQIATGYNLS IS QHDPT AHAEI LCLRSAGKKLENYRLLDA TLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGT VVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKK ALKLAQRAQQGIE

[0040] Haemophilus influenzae F3031 (H. influenzae) TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQS DPT ΑΗ AEIIALRNG AKNI QN YRLLNS TLY VTLEPCTMC AG AILHS RIKRLVFG AS D YK TGAIGSRFHFFDDYKMNHTLEITSGVLAEE CSQKLSTFFQKRREEKKIEKALLKSLSD K

[0041] Caulobacter crescentus (C. crescentus) TadA: MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAH DPTAHAEIAAMRAAAAKLGNYRL TDLTLVVTLEPCAMCAGAISHARIGRVVFGADD PKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRK AKI

[0042] Geobacter sulfurreducens (G. sulfurreducens) TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSN DPSAHAEMIAIRQAARRSANWRL TGATLYVTLEPCLMCMGAIILARLERVVFGCYDP KGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRR RKKAKATPALF IDERKVPPEP

[0043] TadA7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD

[0044] "Agent" refers to any small molecule compound, antibody, nucleic acid molecule, or polypeptide. , or fragments thereof.

[0045] "Modification" means detection by known methods of standard art as described herein. This refers to changes in the structure, expression level, or activity of a gene or polypeptide. When used in the specification, modifications (e.g., increase or decrease) are considered a 10% change in expression level, 2 This includes changes of 5%, 40%, and 50% or more in expression levels.

[0046] "Analog" refers to a molecule that is not identical but possesses similar functional or structural characteristics. It tastes like... For example, polypeptide analogs have the same biological activity as the corresponding natural polypeptide. While retaining at least some of the properties, it enhances the functionality of its analog compared to natural polypeptides. It has specific sequence modifications that enhance the polynucleotide binding activity. Such modifications include, for example, polynucleotide binding activity. Without altering the analog's protease resistance, membrane permeability, or half-life, It can be made possible. In another example, polynucleotide analogs are natural polynucleotides. Compared to the analog version, it has certain modifications that enhance the analog functionality, while the corresponding natural polynucleotide The biological activity of the rheotide is preserved. Such modifications are made to the polynucleotide DNA. The analog may increase affinity, half-life, and / or nuclease resistance, and is non-natural. It may contain nucleotides or amino acids.

[0047] A "base editor (BE)" or "nucleic acid base editor (NBE)" is a polynucleotide editor. This refers to a drug that binds to otide and has nucleic acid base modification activity. In one embodiment, the drug The agent has a domain that has base editing activity, i.e., a base within a nucleic acid molecule (e.g., DNA) (e.g. For example, it is a fusion protein containing domains that can modify A, T, C, G, U. In some embodiments, the domain having base editing activity removes bases within nucleic acid molecules. It can be mino-modified. In one embodiment, a base editor removes bases within a DNA molecule. It can be aminated. In one embodiment, a base editor can aminate cytosine (C) in DNA. ) or adenosine can be deaminated. In one embodiment, a base editor It can deaminate cytosine (C) and adenosine (A) within DNA. In this embodiment, the base editor is a cytidine base editor (CBE). Morphologically, a base editor is an adenosine base editor (ABE). Several implementations In terms of form, base editors include adenosine base editors (ABEs) and cytidine salts. It is a base editor (CBE). In one embodiment, the base editor is an adenosine deamin. This is a nuclease-inactivated Cas9 (dCas9) fused to an enzyme. In some embodiments, Therefore, Cas9 is a circular permutant Cas9 (e.g., spCas9 or saCas9). Yes, the cyclic substitution Cas9 is well known in the field, for example, Oakes et al., Cell 176, 254- As described in 267, 2019. In some embodiments, a base editor is used for base removal repair. A restorative inhibitor, such as a UGI domain, is fused to it. In one embodiment, the fusion protein is Cas9 fused to deaminase and a base excision repair inhibitor such as the UGI domain It contains gase. In other embodiments, the base editor is a debasement base editor.

[0048] Nucleic acid base components and polynucleotides programmable nucleotides in a base editor system. Otido bond components can be bonded to each other covalently or noncovalently. For example, In some embodiments, the deaminase domain is a polynucleotide programme Targeted to the target nucleotide sequence by a nucleotide-binding domain Obtain. In one embodiment, a polynucleotide programmable nucleotide binding domain The in can be fused to or linked to the deaminase domain. In some embodiments, The polynucleotide programmable nucleotide-binding domain is deaminase By interacting with or binding to the main body non-covalently, the deaminase domain Targeting of a target nucleotide sequence is possible. For example, in several embodiments... In this context, nucleic acid base editing components, such as deaminase components, are used in polynucleotide programming. Further heterogeneous parts or domains that form part of a nucleotide-binding domain Further heterogeneous parts or dormant molecules that can interact, associate, or form complexes It may include an additional heterogeneous part, such as polypeptide. They can bind, interact, associate, or form complexes with each other. In the application form, the additional heterogeneous parts bind to, interact with, and associate with the polynucleotides. Alternatively, a complex can be formed. In some embodiments, additional heterogeneous parts , can bind to guide polynucleotides. In some embodiments, additional The heterogeneous portion can be bonded to a polypeptide linker. In some embodiments... The additional heterogeneous portion can then be bound to the polynucleotide linker. The species portion may be a protein domain. In some embodiments, additional species may be added. The seed portion consists of a K homology (KH) domain, an MS2 coat protein domain, and a PP7 coat protein domain. Main component, SfMu Com coated protein domain, steryl α motif, telomerase Ku binding. motif and Ku protein, telomerase Sm7 binding motif and Sm7 protein, This could be an RNA recognition motif.

[0049] The base editor system may further include guide polynucleotide components. The components of a base editor system are covalent bonds, non-covalent interactions, or related It should be understood that they can be bound to one another through any combination of binding and interaction. In one embodiment, the deaminase domain is provided by a guide polynucleotide. Targeting can be performed on a target nucleotide sequence. For example, in some embodiments, Nucleic acid base editing components of the base editor system, such as deaminase components, are guide polynucleotides. It interacts with and binds to a part or segment of the ocide (e.g., a polynucleotide motif). or further heterogeneous parts or domains that can form a complex (e.g., RNA or DNA bonds) It may contain polynucleotide-binding domains (such as those found in composite proteins). Several embodiments In this context, additional heterogeneous parts or domains (e.g., RNA or DNA-binding proteins) The nucleotide-binding domain can be fused to or ligated to the deaminase domain. In some embodiments, additional heterogeneous parts bind to and interact with the polypeptide. It can combine or form a complex with a polypeptide. In some embodiments, The additional heterogeneous portion then binds to, interacts with, associates with, or polynucleotides. It can form a complex with creotide. In some embodiments, additional heterogene The portion can be bound to a guide polynucleotide. In some embodiments, Additional heterogeneous parts can be bonded to the polypeptide linker. Several implementations In this state, additional heterogeneous parts can be bound to the polynucleotide linker. The additional heterogeneous portion may be a protein domain. In some embodiments, The heterogeneous parts of the compound are the K homology (KH) domain, the MS2 coat protein domain, and the PP7 coat protein. Citrate domain, SfMu Com coated protein domain, sterile alpha motif, telomeres ZeKu binding motif and Ku protein, telomerase Sm7 binding motif and Sm7 protein It could be a quality or RNA recognition motif.

[0050] In one embodiment, the base editor system includes one or more proteins, fusion proteins It may contain polypeptides or polynucleotides that encode several real In the application form, the base editor system uses deaminase and the N-terminal fragment of napDNAbp The first polynucleotide encoding the fusion protein, and the C-terminal fragment of napDNAbp. It may include a second polynucleotide encoding napDNAb. For example, in certain embodiments, napDNAb The N-terminal and C-terminal fragments of p are reassembled to form a base editor protein. The N-terminal fragment of napDNAbp is fused to intein-N, and the C-terminal fragment of napDNAbp is fused to intein-C It can be merged with it.

[0051] In one embodiment, the base editor system is a base removal and repair (BER) component The inhibitor of the base may further be included. In one embodiment, base editing Them may further contain inhibitors of base excision repair (BER) components. The components of the base editor system are covalent, non-covalent interactions, or their It is understood that they can be bound to each other through any combination of bonding and interaction. Inhibitors of the BER component may include base excision repair inhibitors. In one embodiment Inhibitors of base excision repair may be uracil DNA glycosylase inhibitors (UGIs). In one embodiment, inhibitors of base excision repair may be inosine base excision repair inhibitors. In one embodiment, a base excision repair inhibitor is polynucleotide programming-enabled. The nucleotide-binding domain allows for targeting of specific nucleotide sequences. In one embodiment, the polynucleotide-programmable nucleotide-binding domain is They can be fused to or linked to inhibitors of base excision repair. In one embodiment, polynucleotides The nucleotide-programmable nucleotide-binding domain is the deaminase domain and salt It can be fused to or linked to inhibitors of group removal repair. In some embodiments, polynu The cleotide-programmable nucleotide-binding domain is an inhibitor of base excision repair. It either interacts non-covalently with or associates with inhibitors of base excision repair. This allows us to target the base excision repair inhibitor to the target nucleotide sequence. For example, in some embodiments, the inhibitor of the base excision repair component is polynucleotide. Further heterogeneous parts that are part of the ocidoprogrammable nucleotide-binding domain or further heterogeneous parts or domains that may interact, associate with, or form complexes with the domain. It may contain. In one embodiment, the inhibitor of base excision repair is a guide polynucleotide. This allows for targeting of the target nucleotide sequence. For example, in some embodiments... In this context, inhibitors of base excision repair are a part or segment of the guide polynucleotide ( For example, further interactions, associations, or complex formations with polynucleotide motifs A heterogeneous portion or domain (for example, a polynucleo such as an RNA or DNA-binding protein) It may include a tide-binding domain. In some embodiments, a guide polynucleotide. Further heterogeneous parts or domains (e.g., polynucleotides such as RNA or DNA-binding proteins) The nucleotide-binding domain can be fused to or linked to inhibitors of base excision repair. In some embodiments, additional heterogenes bind to, interact with, and meet polynucleotides. They can be combined or form a composite. In some embodiments, additional heterogeneous parts It can bind to a guide polynucleotide. In some embodiments, The additional heterogeneous portion can be bonded to the polypeptide linker. In some embodiments In this case, additional heterogeneous parts can be bound to the polynucleotide linker. The heterogeneous portion may be a protein domain. In some embodiments, additional The heterogeneous parts are the K homology (KH) domain, the MS2 coat protein domain, and the PP7 coat protein. Domain, SfMu Com coated protein domain, sterile alpha motif, telomerase Ku Binding motif and Ku protein, telomerase Sm7 binding motif and Sm7 protein, Alternatively, it could be an RNA recognition motif.

[0052] "Base editing activity" refers to the ability to chemically alter the bases within a polynucleotide. This means that. In one embodiment, the first base is converted to the second base. Another embodiment Then, the base is cleaved from the polynucleotide. In one embodiment, the base editing activity is This refers to cytidine deaminase activity, for example, the activity of converting target C·G to T·A. In the embodiment, the base editing activity is adenosine deaminase activity, for example, A This is the activity that converts T to G and C.

[0053] The terms "Cas9" or "Cas9 domain" refer to the Cas9 protein or a fragment of it (e.g., Cas9 The active, inactive, or partially active DNA cleavage domain, and / or gRNA binding of Cas9. This refers to RNA-induced nucleases containing proteins (including the domain). Cas9 nuclease is c ASN1 nuclease or CRISPR (clustered regularly interspaced short palindromic It is sometimes called a repeat-related nuclease. CRISPR is a mobile genetic element (virus, It is an adaptive immune system that provides protection against transposable elements (conjugation plasmids). CRISPR cluster —Includes a spacer, a sequence complementary to the preceding movable element, and a target entry nucleic acid. CRIS PR clusters are transcribed and processed into CRISPR RNA (crRNA). Type II CRISPR In TEM, the correct processing of pre-crRNA is transcoding small RNA (tracrRNA), endogenous TracrRNA requires sex ribonuclease 3 (rnc) and Cas9 protein. This guides the processing of pre-crRNA by rease-3. Subsequently, Cas9 / crRNA / tracrRN A cleaves a linear or circular dsDNA target complementary to the spacer with an endonuclease. The target strand that is not complementary to the crRNA is first cleaved endonuclease-like, and then exo It is trimmed 3'-5' nuclease-like. Both sides of crRNA and tracrRNA are combined into a single RN. A single guide RNA ("sgRNA," or simply "gRNA") is created to be incorporated into type A. It is possible. For example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, See Charpentier E. Science 337:816-821 (2012) (the entire content is referenced by...) (as incorporated herein). Cas9 is a short motif (PAM or) in CRISPR repeat sequences. The protospacer (adjacent motif) helps in recognizing the self and non-self. The sequence and structure of the s9 nuclease are well known to those skilled in the art (e.g., "Complete gene ome sequence of an Ml strain of Streptococcus pyogenes.” Ferretti et al., JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001); “CRISPR RNA maturation by tr ans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., S harma CM., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J., Charpentie r E., Nature 471:602-607(2011); and “A programmable dual-RNA-guided DNA end nuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., See Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012). The body is incorporated herein by reference. The Cas9 ortholog is, to the extent limited, However, it has been described in various species, including S. pyogenes and S. thermophilus. Other suitable Cas9 nucleases and sequences will become apparent to those skilled in the art based on this disclosure. Such Cas9 nucleases and sequences are described in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA B This includes Cas9 sequences from organisms and loci disclosed in iology 10:5, 726-737. All content is incorporated herein by reference.

[0054] Nuclease-inactivating Cas9 protein is interchangeable with the "dCas9" protein (nuclease- It can also be called "dead" Cas9 (meaning Cas9) or catalytically inactive Cas9. Inactive DNA cleavage domain Methods for generating the Cas9 protein (or fragment thereof) containing the following are known (e.g., Jinek et al.) al, Science. 337:816-821(2012); Qi et al, “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression”(2013) Cell. 28; 152 (5): See 1173-83 (each content is incorporated herein by reference). For example, Cas9's DN The A cleavage domain consists of two subdomains: the HNH nuclease subdomain and the RuvC1 subdomain. It is known to contain the main. The HNH subdomain cleaves the strand complementary to the gRNA and RuvC1 Subdomains cleave non-complementary strands. Mutations within these subdomains lead to Cas9 nuclei. It can suppress ase activity. For example, mutants D10A and H840A suppress the nucleotides of S. pyogenes Cas9. Completely inactivates crease activity (Jinek et al, Science. 337:816-821(2012); Qi et al, Cell. 28;152(5): 1173-83 (2013). In one embodiment, Cas9 nuclease, It has an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is "nCas9" It is a protein called nickase (meaning Cas9). In one embodiment, C A protein containing a fragment of as9 is provided. For example, in some embodiments, The protein contains one of the following two Cas9 domains: (1) the gRNA-binding domain of Cas9; ( 2) DNA cleavage domain of Cas9. In one embodiment, a protein containing Cas9 or a fragment thereof is These are called "Cas9 variants." Cas9 variants share homology with Cas9 or its fragments. It possesses. For example, Cas9 variants are at least about 70% identical to wild-type Cas9, and at least about 8% identical. 0% identical, at least approximately 90% identical, at least approximately 95% identical, at least approximately 96% identical, at least Approximately 97% identical, at least approximately 98% identical, at least approximately 99% identical, at least approximately 99.5% identical, They are at least approximately 99.9% identical. In some embodiments, Cas9 mutants are wild Compared to type Cas9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, It has amino acid changes of 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more. It is possible. In some embodiments, the Cas9 variant is a fragment of Cas9 (e.g., gRNA binding). It includes a domain or DNA cleavage domain, and its fragments are less than the corresponding fragment of wild-type Cas9. At least 70% are identical, at least 80% are identical, at least 90% are identical, At the very least, they are approximately 95% identical, at least approximately 96% identical, and at least approximately 97% identical. They are at least approximately 98% identical, at least approximately 99% identical, and at least approximately 99.5% identical. They are, or at least about 99.9% identical. In one embodiment, the fragment is the corresponding wild-type C at least 30%, at least 35%, at least 40%, at least 45%, of the amino acid length of as9 At least 50%, at least 55%, at least 60%, at least 65%, at least 70%, less Both 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, less All are 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%. ru.

[0055] In one embodiment, the fragment has a length of at least 100 amino acids. In this case, the fragments are at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 60 0, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or It is at least 1300 amino acids long. In one embodiment, wild-type Cas9 is Streptococc Corresponding to Cas9 derived from us pyogenes (NCBI reference sequence: NC_17053.1, nucleotide sequence and The amino acid sequence is as follows. ATGGATAAGAAATACTCAATAGGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGATTATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG GCAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAATTGGCAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAATCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTAGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAGAAATGGCTTGTTTGGGAAT CTCATTGCTTTGTCATTGGGATTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAATTAACAGCTTTC AAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATAGTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAGCGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAAATATGTTACTGAGGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGCGCCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGGGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGATATTCAAAAAGCACAGGTGTCTGGACAAGGCCATAGTTTA CATGAACAGATTGCTAACTTAGCTGGCAGTCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAATTGTTGATGAACT GGTCAAAGTAATGGGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCCAGA AAAATTCGCGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCTGTT GAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTACAAAATGGAAGAGACATGTATGTGGACCAAGAATT AGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCATTAAAGACGATTCAATAGACAATA AGGTACTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAAAAC TATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGGTTT GAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAATTT TGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCTAAA TTAGTTTCTGACTTCCGAAAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGCGTA TCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATAAAG TTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCTAAT ATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGGGGA AACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATATTG TCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAAAGAAATTCGGACAAGCTTATTGCT CGTAAAAAGACTGGGATCCAAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGCTAA GGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCTTTG AAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATAAGGAAGTTAAAAGACTTAATCATTAAACTACCTAATAT AGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAATTACAAAAAGGAAATGAGCTGGCTCT GCCAAGCAAATGTGAATTTTTTATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAACAAA AACAATTGTTTGTGGAGCAGCATAAGCATTATTTGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTTATT TTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGAAAA TATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTAAAC GATATACGTCTACAAAAGAAGTTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATTGAT TTGAGTCAGCTAGGAGGTGACTGA (Sequence No. 1) JPEG0007899276000001.jpg161169 (Single underline: HNH domain; Double underline: RuvC domain)

[0056] In one embodiment, wild-type Cas9 has the following nucleotide sequence and / or amino Corresponding to or containing an acid sequence: ATGGATAAAAAGTATTCTATTGGTTTAGACATCGGCACTAATTCCGTTGGATGGGCTGTCATAACCGATGAATACAAAGT ACCTTCAAAGAAATTTAAGGTGTTGGGGAACACAGACCGTCATTCGATTAAAAAAGAATCTTATCGGTGCCCTCCTATTCG ATAGTGGCGAAACGGCAGAGGCGACTCGCCTGAAACGAACCGCTCGGAGAAGGTATACACGTCGCAAGAACCGAATATGT TACTTACAAGAAATTTTTAGCAATGAGATGGCCAAAGTTGACGATTCTTTCTTTCACCGTTTGGAAGAGTCCTTCCTTGT CGAAGAGGACAAGAAACATGAACGGCACCCCATCTTTGGAAACATAGTAGATGAGGTGGCATATCATGAAAAGTACCCAA CGATTTATCACCTCAGAAAAAAGCTAGTTGACTCAACTGATAAAGCGGACCTGAGGTTAATCTACTTGGCTCTTGCCCAT ATGATAAAGTTCCGTGGGCACTTTCTCATTGAGGGTGATCTAAATCCGGACAACTCGGATGTCGACAAACTGTTCATCCA GTTAGTACAAACCTATAATCAGTTGTTTGAAGAACCCTATAAATGCAAGTGGCGTGGATGCGAAGGCTATTCTTAGCG CCCGCCTCTCTAAATCCCGACGGCTAGAAAAACCTGATCGCACAATTACCCGGAGAGAAAAAAATGGTTGTTCGGTAAC CTTATAGCGCTCTCACTAGGCCTGACACCAAATTTAAGTCGAACTTCGACTTAGCTGAAGATGCCAAATTGCAGCTTAG TAAGGACACGTACGATGACGATCTCGCAATCTACTGGCACAAATTGGAGATCAGTATGCGGACTTATTTTTGGCTGCCA AAAACCTTAGCGATGCAATCCTCCTATCTGACATACTGAGAGTTAATACTGAGATTACCAAGGCGCCGTTATCCCGCTTCA ATGATCAAAAGGTACGATGAACATCACCAAGACTTGACACTTCTCAAGGCCCTAGTCCGTCAGCAACTGCCTGAGAAATA TAAGGAAATATTCTTTGATCAGTCGAAAACGGGTACGCAGGTTATATTGACGGGCGGAGCGAGTCAAGAGGAATTCTACA AGTTTATCAAACCCATATTAGAGAATGGATGGGACGGAAGAGTTGCTTGTAAAACTCAATCGCGAAGATCTACTGCGA AAGCAGCGGACTTTCGACAACGGTAGCATTCCACATCAAATCCACTTAGGCGAATTGCATGCTATACTTAGAAGGCAGGA GGATTTTTATCCGTTCCTCAAAGACAATCGTGAAAAGATTGAGAAAATCCTAACCTTTCGCATACCTACTATGTGGGAC CCCTGGCCCGAGGGAACTCTCGGTTCGCATGGATGACAAGAAAGTCCGAAGAAACGATTACTCCATGGAATTTTGAGGAA GTTGTCGATAAAGGTGCGTCAGCTCAATCGTTCATCGAGAGGATGACCAACTTTGACAAGAATTTACCGAACGAAAAAGT ATTGCCTAAGCACAGTTTACTTTACGAGTATTTCACAGTGTACAATGAACTCACGAAAGTTAAGTATGTCACTGAGGGCA TGCGTAAACCCGCCTTTCTAAGCGGAGAACAGAAGAAAGCAATAGTAGATCTGTTATTCAAGACCAACCGCAAAGTGACA GTTAAGCAATTGAAAGAGGACTACTTTAAGAAAATTGAATGCTTCGATTCTGTCGAGATCTCCGGGGTAGAAGATCGATT TAATGCGTCACTTGGTACGTATCATGACCTCCTAAAGATAATTAAAGATAAGGACTTCCTGGATAACGAAGAGAATGAAG ATATCTTAGAAGATATAGTGTTGACTCTTACCCTCTTTGAAGATCGGGAAATGATTGAGGAAAGACTAAAAAACATACGCT CACCTGTTCGACGATAAGGTTATGAAACAGTTAAAGAGGCGTCGCTATACGGGCTGGGGACGATTGTCGCGGAAACTTAT CAACGGGATAAGAGACAAGCAAAGTGGTAAAACTATTCTCGATTTTCTAAAGAGCGACGGCTTCGCCATAGGAACTTTA TGCAGCTGATCCATGATGACTCTTTAACCTTCAAAGAGGATATACAAAAGGCACAGGTTTCCGGACAAGGGGACTCATTG CACGAACATATTGCGAATCTTGCTGGTTCGCCAGCCATCAAAAAGGGCATACTCCAGACAGTCAAAGTAGTGGATGAGCT AGTTAAGGTCATGGGACGTCACAAACCGGAAAACATTGTAATCGAGATGGCACGCGAAAATCAAACGACTCAGAAGGGGC AAAAAAACAGTCGAGAGCGGATGAAGAGAATAGAAGAGGGTATTAAAGAACTGGGCAGCCAGATCTTAAAGGAGCATCCT GTGGAAAATACCCAATTGCAGAACGAGAAACTTTACCTCTATTACCTACAAAATGGAAGGGACATGTATGTTGATCAGGA ACTGGACATAAACCGTTTATCTGATTACGACGTCGATCACATTGTACCCCAATCCTTTTTGAAGGACGATTCAATCGACA ATAAAGTGCTTACACGCTCGGATAAGAACCGAGGGAAAAGTGACAATGTTCCAAGCGAGGAAGTCGTAAAGAAAATGAAG AACTATTGGCGGCAGCTCCTAAATGCGAAACTGATAACGCAAAGAAAGTTCGATAACTTAACTAAAGCTGAGAGGGGTGG CTTGTCTGAACTTGACAAGGCCGGATTTATTAAACGTCAGCTCGTGGAAACCCGCCAAATCACAAAGCATGTTGCACAGA TACTAGATTCCCGAATGAATACGAAATACGACGAGAACGATAAGCTGATTCGGGAAGTCAAAGTAATCACTTTAAAGTCA AAATTGGTGTCGGACTTCAGAAAGGATTTTCAATTCTATAAAGTTAGGGAGATAAATAACTACCACCATGCGCACGACGC TTATCTTAATGCCGTCGTAGGGACCGCACTCATTAAGAAATACCCGAAGCTAGAAAGTGAGTTTGTGTATGGTGATTACA AAGTTTATGACGTCCGTAAGATGATCGCGAAAGCGAACAGGAGATAGGCAAGGCTACAGCCAAATACTTCTTTTATTCT AACATTATGAATTTCTTTAAGACGGAAATCACTCTGGCAAACGGAGAGATACGCAAACGACCTTTAATTGAAACCAATGG GGAGACAGGTGAAATCGTATGGGATAAGGGCCGGGACTTCGCGACGGTGAGAAAAGTTTTGTCCATGCCCCAAGTCAACA TAGTAAAGAAAACTGAGGTGCAGACCGGAGGGTTTTCAAAGGAATCGATTCTTCCAAAAAGGAATAGTGATAAGCTCATC GCTCGTAAAAAGGACTGGGACCCGAAAAAGTACGGTGGCTTCGATAGCCCTACAGTTGCCTATTCTGTCCTAGTAGTGGC AAAAGTTGAGAAGGGAAATCCAAGAAACTGAAGTCAGTCAAAGAATTATTGGGGATAACGATTATGGAGCGCTCGTCTT TTGAAAAGACCCCATCGACTTCCTTGAGGCGAAAGGTTACAAGGAAGTAAAAAAGGATCTCATAATTAAACTACCAAAG TATAGTCTGTTTGAGTTAGAAATGGCCGAAAACGGATGTTGGCTAGCGCCGGAGAGCTTCCAAAGGGGAACGAACTCGC ACTACCGTCTAAATACGTGAATTTCCTGTATTTAGCGTCCCATTACGAGAAGTTGAAAGGTTCACCTGAAGATAACGAAC AGAAGCAACTTTTTGTTGAGCAGCACAAACATTATCTCGACGAAATCATAGAGCAAATTTCGGAATTCAGTAAGAGAGTC ATCCTAGCTGATGCCAATCTGGACAAAGTATTAAGCGCATACAACAAGCACAGGGATAAACCCATACGTGAGCAGGCGGA AAATATTATCCATTTGTTTACTCTTACCAACCTCGGCGCTCCAGCCGCATTCAAGTATTTTGACACAACGATAGATCGCA AACGATACACTTCTACCAAGGAGGTGCTAGACGCGACACTGATTCACCAATCCATCACGGGATTATATGAAACTCGGATA GATTTGTCACAGCTTGGGGGTGACGGATCCCCAAGAAGAAGAGGAAAGTCTCGAGCGACTACAAAGACCATGACGGTGA TTATAAAGATCATGACATCGATTACAAGGATGACGATGACAAGGCTGCAGGA JPEG0007899276000002.jpg157169 (Single underline: HNH domain; Double underline: RuvC domain)

[0057] In some embodiments, wild-type Cas9 is compared to Cas9 from Streptococcus pyogenes. Corresponding (NCBI reference sequence: NC_002737.2 (the following nucleotide sequence); and Uniprot reference sequence Column: Q99ZW2 (followed by the amino acid sequence below). ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGAATATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG ACAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGTAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAACCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTGGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAAAAATGGCTTATTTGGGAAT CTCATTGCTTTGTCATTGGGTTTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATACTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAACGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAAATATGTTACTGAAGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGTACCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGAGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGACATTCAAAAAGCACAAGTGTCTGGACAAGGCGATAGTTTA CATGAACATATTGCAAATTTAGCTGGTAGCCCTGCTATTAAAAAGGTATTTTACAGACTGTAAAAGTTGTTGATGAATT GGTCAAAGTAATGGGGCGGCATAAGCCAGAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCC AGAAAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAAGAGCATCCT GTTGAAAATACTCAAATTGCAAATGAAAAGCTCTATCTCTATTATCTCCAAAATGGAAGAGACATGTATGTGGACCAAGA ATTAGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCCTTAAAGACGATTCAATAGACA ATAAGGTCTTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAA AACTATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGG TTTGAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCAGTGGCACAAA TTTTGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCT AATTAGTTTCTGACTTCCGAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGC GTATCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATA AAGTTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCT AATATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGG GGAAACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATA TTGTCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTTACCAAAAGAAATTCGGACAAGCTTATT GCTCGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGC TAAGGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCT TTGAAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAA TATAGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGC TCTGCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAAC AAAAACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTT ATTTTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGA AAATATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTA AACGATATACGTCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATT GATTTGAGTCAGCTAGGAGGTGACTGA JPEG0007899276000003.jpg156168(Underlined once: HNH domain; Underlined twice: RuvC domain)

[0058] ​In some embodiments, Cas9 is from Corynebacterium ulcerans (NCBI Refs: NC_01 5683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016 786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Strept ococcus iniae(NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquisI (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCB I Ref: YP_820832.1), Listeria innocua (NCBI Ref: NP_472073.1), Campylobacter jej uni (NCBI Ref: YP_002344900.1) or Neisseria. meningitidis (NCBI Ref: YP_00 2342100.1), or Cas9 from any other organism.

[0059] In some embodiments, dCas9 corresponds to a Cas9 amino acid sequence having one or more mutations that inactivate Cas9 nuclease activity, or comprises a part or all thereof. For example, in some embodiments, the dCas9 domain has D10A and H840A mutations as well as includes the corresponding mutation in another Cas9. In some embodiments, dCas9 is comprises the following amino acid sequences of dCas9 (D10A and H840A): JPEG0007899276000004.jpg157168 (single underline: HNH domain; double underline: RuvC domain)

[0060] In some embodiments, the Cas9 domain contains the D10A mutation, while the residue at position 840 in the amino acid sequence provided above, or the residue at the corresponding position in any of the amino acid sequences provided herein, remains histidine. [[ID=IO]]

[0061] In other embodiments, dCas9 variants having mutations other than D10A and H840A are provided, which result in, for example, nuclease-inactivated Cas9 (dCas9). Such mutations include, for example, other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). In certain aspects, variants or homologs of dCas9 are provided that have at least about 70% identity, at least about SO% identity, at least about 90% identity, at least about 95% identity, at least about 98% identity, at least about 99% identity, at least about ۹۹.5% identity, or at least about 99.9% identity. In certain aspects, variants of dCas9 are provided that have an amino acid sequence that is shorter or longer by about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids or more. [[ID=UI]] ​​​​​​​​​​​

[0062] In some embodiments, the Cas9 fusion protein provided herein is a Cas9 fusion protein. The full-length amino acid sequence of the protein, for example, one of the Cas9 sequences provided herein. In other embodiments, the fusion protein provided herein is a full-length Cas9 compound. It does not include the column, but includes only one or more fragments of it.

[0063] Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein. Further preferred arrangements of domains and fragments will be obvious to those skilled in the art.

[0064] In some embodiments, Cas9 is Corynebacterium ulcerans (NCBI Refs: NC_01 5683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016 786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense(NCBI Ref: NC_021846.1); Strepto coccus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquisI (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCB I Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP_472073.1); Campylobacter jej uni (NCBI Ref: YP_002344900.1); or Neisseria. meningitidis (NCBI Ref: YP_0023 This refers to Cas9 originating from 42100.1).

[0065] Additional Cas9 proteins (e.g., nuclease-inactivated (dead) Cas9 (dCas9), Cas9 ni Casase (nCas9, or nuclease-active Cas9) is a variant and homolog of nCas9. It should be understood that this is included within the scope of this disclosure. Exemplary Cas9 protein This includes, but is not limited to, the following: In one embodiment, Cas9 The protein is nuclease-inactivated Cas9 (dCas9). In one embodiment, the Cas9 protein The protein is Cas9 nickase (nCas9). In some embodiments, the Cas9 protein The substance is Cas9 with nuclease activity.

[0066] Exemplary catalyst-inactive Cas9 (dCas9): DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0067] Exemplary catalyst of Cas9ニッカーゼ(nCas9): DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0068] Exemplary catalyst activityCas9: DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD.

[0069] In one embodiment, Cas9 is a domain and kingdom of archaea that constitute unicellular prokaryotic microorganisms. For example, it refers to Cas9 derived from nanoarchaea. In one embodiment, the Cas protein is, for example, Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res. This refers to CasX or CasY described in doi: 10.1038 / cr.2017.21 on February 21, 2017, and all of its contents. The contents of the body are incorporated herein by reference. Using genome-degrading metagenomics, Many CRISPR-Cas systems were identified, including Cas9, which was the first to be reported in the Archaea domain of life. This branched Cas9 protein was found in nanoarchaea, which have been little studied. It was discovered as part of the active CRISPR-Cas system. In bacteria, two previously unknown mechanisms were found. Two systems, CRISPR-CasX and CRISPR-CasY, were discovered, and they are among the most advanced systems discovered to date. It falls into the most compact system category. In some embodiments, Cas9 is a variant of CasX or CasX. Represents a riant. In some embodiments, Cas9 represents CasY or a variant of CasY. Represents: Nucleic acid programmable DNA-binding protein (napDNAbp). Other RNA-induced DNA-binding proteins may also be used as DNA-binding proteins in this disclosure. It should be understood that this is within the acceptable range.

[0070] In some embodiments, napDNAbp is a Cas9 domain, for example, a nuclease This can be active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Non-limiting examples of nucleic acid-programmable DNA-binding proteins include Cas9 (e.g., dCa Cas effector proteins (s9 and nCas9), type II Cas effector proteins, type V Cas effector proteins Proteins, type VI Cas effector proteins, CARF, DinG, their homologs, or Modified or genetically engineered versions of these are examples. Other nucleic acid programming is possible. DNA-binding proteins, even if not specifically listed in this disclosure, are also included in the scope of this disclosure. It is within the scope of this classification. For example, Makarova et al. “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct;1:325-336. doi: 10.1089 / crispr.20 18.0033; Yan et al., “Functionally diverse type V CRISPR-Cas systems” Science. See 2019 Jan 4;363(6422):88-91. doi: 10.1126 / science.aav7271 (for the full content of each section). (The body is incorporated herein by reference.)

[0071] In some embodiments, any nucleic acid of the fusion protein provided herein Programmable DNA-binding proteins (napDNAbp) can be CasX or CasY proteins. In some embodiments, napDNAbp is the CasX protein. In this state, napDNAbp is a CasY protein. In some embodiments, napDNAbp It contains at least 85% and at least 90% of the naturally occurring CasX or CasY proteins. At least 91%, at least 92%, at least 93%, at least 94%, at least 95%, less At least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% It contains an amino acid sequence that is identical. In some embodiments, napDNAbp is naturally It is a CasX or CasY protein that is present. In some embodiments, napDNAbp is At least 85% of any of the CasX or CasY proteins described herein, At least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or less It contains amino acid sequences with at least 99.5% identity. CasX and CasY from other bacterial species are also included. Furthermore, please understand that this disclosure may be used in accordance with the terms of this disclosure.

[0072] CasX (uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) >tr|F0NN87|F0NN87_SULIH CRISPR-associated Casx protein OS = Sulfolobus islandicu s (strain HVE10 / 4) GN = SiH_0402 PE=4 SV=1 MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAV GQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNL NSDDGKVRDLKLISAYVNGELIRGEG

[0073] >tr|F0NH53|F0NH53_SULIR CRISPR associated protein, Casx OS = Sulfolobus islandic us (strain REY15A) GN=SiRe_0771 PE=4 SV=1 MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAV GQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLN SDDGKVRDLKLISAYVNGELIRGEG

[0074] CasY (ncbi.nlm.nih.gov / protein / APG80656.1) >APG80656.1 CRISPR-associated protein CasY [uncultured Parcubacteria group bacte rium] MSKRHPRISGVKGYRLHAQRLEYTGKSGAMRTIKYPLYSSPSGGRTVPREIVSAINDDYVGLYGLSNFDDLYNAEKRNEE KVYSVLDFWYDCVQYGAVFSYTAPGLLKNVAEVRGGSYELTKTLKGSHLYDELQIDKVIKFLNKKEISRANGSLDKLKKD IIDCFKAEYRERHKDQCNKLADDIKNAKKDAGASLGERQKKLFRDFFGISEQSENDKPSFTNPLNLTCCLLPFDTVNNNR NRGEVLFNKLKEYAQKLDKNEGSLEMWEYIGIGNSGTAFSNFLGEGFLGRLRENKITELKKAMMDITDAWRGQEQEEELE KRLRILAALTIKLREPKFDNHWGGYRSDINGKLSSWLQNYINQTVKIKEDLKGHKKDLKKAKEMINRFGESDTKEEAVVS SLLESIEKIVPDDSADDEKPDIPAIAIYRRFLSDGRLTLNRFVQREDVQEALIKERLEAEKKKKPKKRKKKSDAEDEKET IDFKELFPHLAKPLKLVPNFYGDSKRELYKKYKNAAIYTDALWKAVEKIYKSAFSSSLKNSFFDTDFDKDFFIKRLQKIF SVYRRFNTDKWKPIVKNSFAPYCDIVSLAENEVLYKPKQSRSRKSAAIDKNRVRLPSTENIAKAGIALARELSVAGFDWK DLLKKEEHEEYIDLIELHKTALALLLAVTETQLDISALDFVENGTVKDFMKTRDGNLVLEGRFLEMFSQSIVFSELRGLA GLMSRKEFITRSAIQTMNGKQAELLYIPHEFQSAKITTPKEMSRAFLDLAPAEFATSLEPESLSEKSLLKLKQMRYYPHY FGYELTRTGQGIDGGVAENALRLEKSPVKKREIKCKQYKTLGRGQNKIVLYVRSSYYQTQFLEWFLHRPKNVQTDVAVSG SFLIDEKKVKTRWNYDALTVALEPVSGSERVFVSQPFTIFPEKSAEEEGQRYLGIDIGEYGIAYTALEITGDSAKILDQN FISDPQLKTLREEVKGLKLDQRRGTFAMPSTKIARIRESLVHSLRNRIHHLALKHKAKIVYELEVSRFEEGKQKIKKVYA TLKKADVYSEIDADKNLQTTVWGKLAVASEISASYTSQFCGACKKLWRAEMQVDETITTQELIGTVRVIKGGTLIDAIKD FMRPPIFDENDTPFPKYRDFCDKHHISKKMRGNSCLFICPFCRANADADIQASQTIALLRYVKEEKKVEDYFERFRKLKN IKVLGQMKKI

[0075] The terms "CRISPR-Cas domain" or "CRISPR-Cas DNA-binding domain" are related to CRISPR. RNA-induced proteins containing the Cas protein or a fragment thereof (e.g., Cas protein) Active, inactive, or partially active DNA cleavage domains, and / or g of the Cas protein This refers to proteins that contain an RNA-binding domain. CRISPR clusters are transcribed into CRISPR RNA. The CRISPR cluster is transcribed and converted into CRISPR RNA (crRNA). In some CRISPR systems, the correct processing of pre-crRNA is impaired. Squadled small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas protein This is required. This tracrRNA is a pre-crRNA that is assisted by ribonuclease 3 in processing. It acts as a guide for the spacer. Subsequently, Cas9 / crRNA and / or tracrRNA act as spacers. - Endonucleaciately cleaves complementary linear or circular dsDNA targets. It may require both RNAs for DNA binding and cleavage. However, crRNA and tracrRNA A single guide RNA ("sgRNA" or simply "gNR") incorporates both aspects into a single RNA species. A) can be genetically engineered. For example, Jinek M., Chylinski K., Fonfara I., Hau See er M., Doudna JA, Charpentier E. Science 337:816-821 (2012) (the entire content is available therein). (Incorporated herein by reference). Cas proteins are short motifs in CRISPR repeat sequences. Recognizing the PAM (or protospacer adjacent motif) and distinguishing between self and non-self It helps with [something]. CRISPR-Cas proteins include Cas9, CasX, CasY, Cpf1, C2c1, and C2 Examples include, but are not limited to, C3 or its active fragments. Further suitable CRIS The PR-Cas protein and sequence will be apparent to those skilled in the art based on this disclosure.

[0076] Nuclease-inactivated CRISPR-Cas proteins are interchangeable with "dCas" proteins. Nuclease (meaning "dead" Cas) or catalytically inactive Cas. Inactive DNA Methods for generating Cas proteins (or fragments thereof) with cleavage domains are known. For example, Jinek et al., Science. 337:816-821(2012); Qi et al., “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression” (20 13) See Cell. 28;152(5):1173-83. Its entire contents are incorporated herein by reference. . ). For example, the DNA cleavage domain of Cas9 is the HNH nuclease subdomain and the RuvC1 subdomain. It is known to contain two subdomains called the bud domain. The HNH subdomain is gRNA. While complementary strands are cleaved, the RuvC1 subdomain cleaved non-complementary strands. Mutations within a subdomain suppress the nuclease activity of Cas9. For example, mutation D10A and H840A completely inactivates the nuclease activity of S. pyogenes Cas9 (Jinek et al., S Science. 337:816-821 (2012); Qi et al., Cell. 28;152(5):1173-83 (2013). A certain implementation Morphologically, Cas nucleases have an inactive (e.g., deactivated) DNA cleavage domain. Cas is a nickase, also known as an nCas protein. (Meaning). Cas variants share homology with the CRISPR-Cas protein or a fragment thereof. For example, Cas variants are at least approximately 70% identical to the wild-type CRISPR-Cas protein. Sex, at least approximately 80% identity, at least approximately 90% identity, at least approximately 95% identity, At least approximately 96% identity, at least approximately 97% identity, at least approximately 98% identity, at least At least approximately 99% identity, at least approximately 99.5% identity, or at least approximately 99.9% identity It has. In some embodiments, the Cas variant is wild-type CRISPR-Cas protein In comparison to quality, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 ,20,21,22,21,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39 , having 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes Obtain. In some embodiments, the Cas variant is obtained by having its fragments as wild-type CRISPR-Cas. At least approximately 70% identity, at least approximately 80% identity, for the corresponding fragment of protein. At least approximately 90% identity, at least approximately 95% identity, at least approximately 96% identity, less At least approximately 97% identity, at least approximately 98% identity, at least approximately 99% identity, at least The CRISPR-Cas protein has approximately 99.5% identity, or at least approximately 99.9% identity. A fragment of DNA (e.g., a gRNA programmable DNA-binding domain or a DNA-cleaving domain) Includes. In one embodiment, the fragment is the amino acid length of the corresponding wild-type CRISPR-Cas protein. at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, At least 55%, at least 60%, at least 65%, at least 70%, at least 75%, less 80% each, at least 85%, at least 90%, at least 95%, similarly, at least 96%, It is at least 97%, at least 98%, at least 99%, or at least 99.5%. In one embodiment, the fragment is at least 100 amino acids in length. , lengths of at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 70 0, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least It contains 1300 amino acids.

[0077] In this disclosure, the terms "comprises," "comprising," and "contains" are used. The terms "taining" and "having" have the meanings defined in U.S. patent law. It can mean "includes," "includes," etc., and "essentially from ~" "Consisting essentially of" or "Consisting essentially of" Similarly, "y)" has the same meaning as defined in U.S. Patent Law, and the term is open-ended. ...and the basic or novel characteristics of the described item are found in other entities that are more significant than the described item. Unless otherwise specified, other entities are permitted, provided that prior art is not implemented. Exclude the aspect.

[0078] "Cytidine deaminase" is an enzyme that performs deamination reactions, which convert amino groups to carbonyl groups. This means a polypeptide or fragment thereof that can catalyze. In one embodiment Cytidine deaminase converts cytosine to uracil, or 5-methylcytosine to thymine. Convert. PmCDA1 (Petromyzon marinus cytosine deaminase) derived from Petromyzon marinus. 1. AIDs (active) derived from mammals (e.g., humans, pigs, cattle, horses, monkeys, etc.) (PmCDA1) Activation-inducing cytidine deaminase (AICDA), and APOBEC are exemplary cytidine deaminases. It is Ze.

[0079] The nucleotide sequence and amino acid sequence of PmCDA1, and the nucleotide sequence and amino acid sequence of human AID CDS. The acid sequence is shown below.

[0080] >tr|A5H718|A5H718_PETMA Cytosine deaminase OS=Petromyzon marinus OX=7757 PE=2 SV =1 MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSG TERGIHAEIFSIRKVEEYLRDNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLK IWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKT LKRAEKRRSELSIMIQVKILHTTKSPAV

[0081] >EF094822.1 Petromyzon marinus isolate PmCDA.21 cytosine deaminase mRNA, complet e cds TGACACGACACAGCCGTGTATATGAGGAAGGGTAGCTGGATGGGGGGGGGGGGAATACGTTCAGAGAGGA CATTAGCGAGCGTCTTGTTGGTGGCCTTGAGTCTAGACACCTGCAGACATGACCGACGCTGAGTACGTGA GAATCCATGAGAAGTTGGACATCTACACGTTTAAGAAACAGTTTTTCAACAACAAAAAATCCGTGTCGCA TAGATGCTACGTTCTCTTTGAATTAAAACGACGGGGTGAACGTAGAGCGTGTTTTTGGGGCTATGCTGTG AATAAACCACAGAGCGGGACAGAACGTGGAATTCACGCCGAAATCTTTAGCATTAGAAAAGTCGAAGAAT ACCTGCGCGACAACCCCGGACAATTCACGATAAATTGGTACTCATCCTGGAGTCCTTGTGCAGATTGCGC TGAAAAGATCTTAGAATGGTATAACCAGGAGCTGCGGGGGAACGGCCACACTTTGAAAATCTGGGCTTGC AAACTCTATTACGAGAAAAATGCGAGGAATCAAATTGGGCTGTGGAACCTCAGAGATAACGGGGTTGGGT TGAATGTAATGGTAAGTGAACACTACCAATGTTGCAGGAAAATATTCATCCAATCGTCGCACAATCAATT GAATGAGAATAGATGGCTTGAGAAGACTTTGAAGCGAGCTGAAAAACGACGGAGCGAGTTGTCCATTATG ATTCAGGTAAAAATACTCCACACCACTAAGAGTCCTGCTGTTTAAGAGGCTATGCGGATGGTTTTC

[0082] >tr|Q6QJ80|Q6QJ80_HUMAN Activation-induced cytidine deaminase OS=Homo sapiens OX =9606 GN=AICDA PE=2 SV=1 MDSLLMNRRKFLYQFKNVRWAKGRRETYLCYVVKRRDSATSFSLDFGYLRNKNGCHVELL FLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNLSLRIFTARLYFCEDRK AEPEGLRRLHRAGVQIAIMTFKAPV

[0083] >NG_011588.1:5001-15681 Homo sapiens activation induced cytidine deaminase (AICD A), RefSeqGene (LRG_17) on chromosome 12 AGAGAACCATCATTAATTGAAGTGAGATTTTTCTGGCCTGAGACTTGCAGGGAGGCAAGAAGACACTCTG GACACCACTATGGACAGGTAAAGAGGCAGTCTTCTCGTGGGTGATTGCACTGGCCTTCCTCTCAGAGCAA ATCTGAGTAATGAGACTGGTAGCTATCCCTTTCTCTCATGTAACTGTCTGACTGATAAGATCAGCTTGAT CAATATGCATATATTTTTTGATCTGTCTCCTTTTCTTCTATTCAGATCTTATACGCTGTCAGCCCAAT TCTTTCTGTTTCAGACTTCTCTTGATTTCCCTCTTTTCATGTGGCAAAAGAAGTAGTGCGTACAATGTA CTGATTCGTCCTGAGATTTGTACCATGGTTGAAACTAATTTATGGTAATAATATTAACATAGCAAATCTT TAGAGACTCAAATCATGAAAAGGTAATAGCAGTACTGTACTAAAAACGGTAGTGCTAATTTTCGTAATAA TTTTGTAAATATTCAAACAGTAAAACACTTGAAGACACACTTTCCTAGGGAGGCGTTTACTGAAATAATTT AGCTATAGTAAGAAAATTTGTAATTTTAGAAATGCCAAGCATTCTAAATTAATTGCTTGAAAGTCACTAT GATTGTGTCCATTATAAGGAGACAAATTCATTCAAGCAAGTTATTTAATGTTAAAGGCCCAATTGTTAGG CAGTTAATGGCACTTTTACTATTAACTAATCTTTCCATTTGTTCAGACGTAGCTTAACTTACCTCTTAGG TGTGAATTTGGTTAAGGTCCTCATAATGTCTTTATGTGCAGTTTTTGATAGGTTATTGTCATAGAACTTA TTCTATTCCTACATTTTATGATTACTATGGATGTATGAGAATAACACCTAATCCTTTACTTTACCTCAAT TTAACTCCTTTATAAAGAACTTACATTACAGAATAAAGATTTTTTAAAAATATATTTTTTTGTAGAGACA GGGTCTTAGCCCAGCCGAGGCTGGTCTCTAAGTCCTGGCCCAAGCGATCCTCCTGCCTGGGCCTCCTAAA GTGCTGGAATTATAGACATGAGCCATCACATCCAATATACAGAATAAAGATTTTTAATGGAGGATTTAAT GTTCTTCAGAAAATTTTCTTGAGGTCAGACAATGTCAAATGTCTCCTCAGTTTACACTGAGATTTTGAAA ACAAGTCTGAGCTATAGGTCCTTGTGAAGGGTCCATTGGAAATACTTGTTCAAAGTAAAATGGAAAGCAA AGGTAAAATCAGCAGTTGAAATTCAGAGAAAGACAGAAAAGGAGAAAAGATGAAATTCAACAGGACAGAA GGGAAATATATTATCATTAAGGAGGACAGTATCTGTAGAGCTCATTAGTGATGGCAAAATGACTTGGTCA GGATTATTTTTAACCCGCTTGTTTCTGGTTTGCACGGCTGGGGATGCAGCTAGGGTTCTGCCTCAGGGAG CACAGCTGTCCAGAGCAGCTGTCAGCCTGCAAGCCTGAAACACTCCCTCGGTAAAGTCCTTCCTACTCAG GACAGAAATGACGAGAACAGGGAGCTGGAAACAGGCCCCTAACCAGAGAAGGGAAGTAATGGATCAACAA AGTTAACTAGCAGGTCAGGATCACGCAATTCATTTCACTCTGACTGGTAACATGTGACAGAAACAGTGTA GGCTTATTGTATTTTCATGTAGAGTAGGACCCAAAAATCCACCCAAAGTCCTTTATCTATGCCACATCCT TCTTATCTATACTTCCAGGACACTTTTTCTTCCTTATGATAAGGCTCTCTCTCTCTCCACACACACACAC ACACACACACACACACACACACACACACACACAAACACACACCCCGCCAACCAAGGTGCATGTAAAAAGA TGTAGATTCCTCTGCCTTTCTCATCTACACAGCCCAGGAGGGTAAGTTAATATAAGAGGGATTTATTGGT AAGAGATGATGCTTAATCTGTTTAACACTGGGCCTCAAAGAGAGAATTTCTTTTCTTCTGTACTTATTAA GCACCTATTATGTGTTGAGCTTATATATACAAAGGGTTATTATATGCTAATATAGTAATAGTAATGGTGG TTGGTACTATGGTAATTACCATAAAAATTATTATCCTTTTAAAATAAAGCTAATTATTATTGGATCTTTT TTAGTATTCATTTTATGTTTTTTATGTTTTTGATTTTTTAAAAGACAATCTCACCCTGTTACCCAGGCTG GAGTGCAGTGGTGCAATCATAGCTTTCTGCAGTCTTGAACTCCTGGGCTCAAGCAATCCTCCTGCCTTGG CCTCCCAAAGTGTTGGGATACAGTCATGAGCCACTGCATCTGGCCTAGGATCCATTTAGATTAAAATATG CATTTTAAATTTTAAAATAATATGGCTAATTTTTACCTTATGTAATGTGTATACTGGCAATAAATCTAGT TTGCTGCCTAAAGTTTAAAGTGCTTTCCAGTAAGCTTCATGTACGTGAGGGGAGACATTTAAAGTGAAAC AGACAGCCAGGTGTGGTGGCTCACGCCTGTAATCCCAGCACTCTGGGAGGCTGAGGTGGGTGGATCGCTT GAGCCCTGGAGTTCAAGACCAGCCTGAGCAACATGGCAAAACGCTGTTTCTATAACAAAAATTAGCCGGG CATGGTGGCATGTGCCTGTGGTCCCAGCTACTAGGGGGCTGAGGCAGGAGAATCGTTGGAGCCCAGGAGG TCAAGGCTGCACTGAGCAGTGCTTGCGCCACTGCACTCCAGCCTGGGTGACAGGACCAGACCTTGCCTCA AAAAAATAAGAAGAAAAATTAAAAATAAATGGAAACAACTACAAAGAGCTGTTGTCCTAGATGAGCTACT TAGTTAGGCTGATATTTTGGTATTTAACTTTTAAAGTCAGGGTCTGTCACCTGCACTACATTATTAAAAT ATCAATTCTCAATGTATATCCACACAAAGACTGGTACGTGAATGTTCATAGTACCTTTATTCACAAAACC CCAAAGTAGAGACTATCCAAATATCCATCAACAAGTGAACAAATAAACAAAATGTGCTATATCCATGCAA TGGAATACCACCCTGCAGTACAAAGAAGCTACTTGGGGATGAATCCCAAAGTCATGACGCTAAATGAAAG AGTCAGACATGAAGGAGGAGATAATGTATGCCATACGAAATTCTAGAAAATGAAAGTAACTTATAGTTAC AGAAAGCAAATCAGGGCAGGCATAGAGGCTCACACCTGTAATCCCAGCACTTTGAGAGGCCACGTGGGAA GATTGCTAGAACTCAGGAGTTCAAGACCAGCCTGGGCAACACAGTGAAACTCCATTCTCCACAAAAATGG GAAAAAAAGAAAGCAAATCAGTGGTTGTCCTGTGGGGAGGGGAAGGACTGCAAAGAGGGAAGAAGCTCTG GTGGGGTGAGGGTGGTGATTCAGGTTCTGTATCCTGACTGTGGTAGCAGTTTGGGGTGTTTACATCCAAA AATATTCGTAGAATTATGCATCTTAAATGGGTGGAGTTTACTGTATGTAAATTATACCTCAATGTAAGAA AAAATAATGTGTAAGAAAACTTTCAATTCTCTTGCCAGCAAACGTTATTCAAATTCCTGAGCCCTTTACT TCGCAAATTCTCTGCACTTCTGCCCCGTACCATTAGGTGACAGCACTAGCTCCACAAATTGGATAAATGC ATTTCTGGAAAAGACTAGGGACAAAATCCAGGCATCACTTGTGCTTTCATATCAACCATGCTGTACAGCT TGTGTTGCTGTCTGCAGCTGCAATGGGGACTCTTGATTTCTTTAAGGAAACTTGGGTTACCAGAGTATTT CCACAAATGCTATTCAAATTAGTGCTTATGATATGCAAGACACTGTGCTAGGAGCCAGAAAACAAAGAGG AGGAGAAATCAGTCATTATGTGGGAACAACATAGCAAGATATTTAGATCATTTTGACTAGTTAAAAAAGC AGCAGAGTACAAAATCACACATGCAATCAGTATAATCCAAATCATGTAAATATGTGCCTGTAGAAAGACT AGAGGAATAAACACAAGAATCTTAACAGTCATTGTCATTAGACACTAAGTCTAATTATTATTATTAGACA CTATGATATTTGAGATTTAAAAAATCTTTAATATTTTAAAATTTAGAGCTCTTCTATTTTTCCATAGTAT TCAAGTTTGACAATGATCAAGTATTACTCTTTCTTTTTTTTTTTTTTTTTTTTTTTTTGAGATGGAGTTT TGGTCTTGTTGCCCATGCTGGAGTGGAATGGCATGACCATAGCTCACTGCAACCTCCACCTCCTGGGTTC AAGCAAAGCTGTCGCCTCAGCCTCCCGGGTAGATGGGATTACAGGCGCCCACCACCACACTCGGCTAATG TTTGTATTTTTAGTAGAGATGGGGTTTCACCATGTTGGCCAGGCTGGTCTCAAACTCCTGACCTCAGAGG ATCCACCTGCCTCAGCCTCCCAAAGTGCTGGGATTACAGATGTAGGCCACTGCGCCCGGCCAAGTATTGC TCTTATACATTAAAAAACAGGTGTGAGCCACTGCGCCCAGCCAGGTATTGCTCTTATACATTAAAAAATA GGCCGGTGCAGTGGCTCACGCCTGTAATCCCAGCACTTTGGGAAGCCAAGGCGGGCAGAACACCCGAGGT CAGGAGTCCAAGGCCAGCCTGGCCAAGATGGTGAAACCCCGTCTCTATTAAAAATACAAACATTACCTGG GCATGATGGTGGGCGCCTGTAATCCCAGCTACTCAGGAGGCTGAGGCAGGAGGATCCGCGGAGCCTGGCA GATCTGCCTGAGCCTGGGAGGTTGAGGCTACAGTAAGCCAAGATCATGCCAGTATACTTCAGCCTGGGCG ACAAAGTGAGACCGTAACAAAAAAAAAAAAATTTAAAAAAAGAAATTTAGATCAAGATCCAACTGTAAAA AGTGGCCTAAACACCACATTAAAGAGTTTGGAGTTTATTCTGCAGGCAGAAGAGAACCATCAGGGGGTCT TCAGCATGGGAATGGCATGGTGCACCTGGTTTTTGTGAGATCATGGTGGTGACAGTGTGGGGAATGTTAT TTTGGAGGGACTGGAGGCAGACAGACCGGTTAAAAGGCCAAGCAACAGATAAGGAGGAAGAAGATGAGG GCTTGGACCGAAGCAGAGAAGAGCAAACAGGGAAGGTACAAATTCAAGAAATATTGGGGGGTTGAATCA ACACATTTAGATGATTAATTAAATATGAGGACTGAGGAATAAGAAATGAGTCAAGGATGGTTCCAGGCTG CTAGGCTGCTTACCTGAGGTGGCAAAGTCGGGAGGAGTGGCAGTTTAGGACAGGGGGCAGTTGAGGAATA TTGTTTTGATCATTTTGAGTTTGAGGTACAAGTTGGACACTTAGGTAAAGACTGGAGGGGAAATCTGAAT ATACAATTATGGGGACTGAGGAACAAGTTTTTTATTTTTTGTTTCGTTTTCTTGTTGAAGAACAAATTT AATTGTAATCCCAAGTCATCAGCATCTAGAAGACAGTGGCAGGAGGTGACTGTCTTGTGGGTAAGGGTTT GGGGTCCTTGATGAGTATCTCTCAATTGGCCTTAAATATAAGCAGGAAAAGGAGTTTATGATGGATTCCA GGCTCAGCAGGGCTCAGGGAGGGCTCAGGCAGCCAGCAGAGGAAGTCAGAGCATCTTCTTTGGTTTAGCCCC AAGTAATGACTTCCTTAAAAAAGCTGAAGGAAAATCCAGAGTGACCAGATTATAAACTGTACTCTTGCATT TTCTCTCCCTCCTTCACCCACAGCCTCTTGATCGAACCGGAGGAAGTTTCTTTACCAATTCAAAAATGTC CGCTGGGCTAAGGGTCGGCGTGAGACCTACCTGTGCTACGTAGTGAAGAGGCGTGACAGTGCTACATCCT TTTCACTGGACTTTGGTTATCTTCGCAATAAGGTATCAATTAAAGTCGGCTTTGCAAGCAGTTTAATGGT CAACTGTGAGTGCTTTTAGAGCCACCTGCTGATGGTATTACTTCCATCCTTTTTTGGCATTTGTGTCTCT ATCACATTCCTCAAATCCTTTTTTTTATTTCTTTTTCCATGTCCATGCACCCATATTAGACATGGCCCAA AATATGTGATTTAATTCCTCCCCAGTAATGCTGGGCACCCTAATACCACTCCTTCCTTCAGTGCCAAGAA CAACTGCTCCCAAACTGTTTACCAGCTTTCCTCAGCATCTGAATTGCCTTTGAGATTAATTAAGCTAAAA GCATTTTTATATGGGAGAATATTATCAGCTTGTCCAAGCAAAAATTTTAAATGTGAAAAACAAATTGTGT CTTAAGCATTTTTGAAAATTAAGGAAGAAGAATTTGGGAAAAAATTAACGGTGGCTCAATTCTGTCTTCC AAATGATTTCTTTTCCCTCCTACTCACATGGGTCGTAGGCCAGTGAATACATTCAACATGGTGATCCCCA GAAAACTCAGAGAAGCCTCGGCTGATGATTAATTAAATTGATCTTTCGGCTACCCGAGAGAATTACATTT CCAAGAGACTTCTTCACCAAAATCCAGATGGGTTTACATAAACTTCTGCCCACGGGTATCTCCTCTCTCC TAACACGCTGTGACGTCTGGGCTTGGTGGAATCTCAGGGAAGCATCCGTGGGGTGGAAGGTCATCGTCTG GCTCGTTGTTTGATGGTTATATTACCATGCAATTTTCTTTGCCTACATTTGTATTGAATACATCCCAATC TCCTTCCTATTCGGTGACATGACACATTCTATTTCAGAAGGCTTTGATTTTCAAGCACTTTTCATTTAC TTCTCATGGCAGTGCCTATTACTTCTCTTACAATACCCATCTGTCTGCTTTACCAAAATCTATTTCCCCT TTTCAGATCCTCCCAAATGGTCCTCATAAACTGTCCTGCCTCCACCTAGTGGTCCAGGTATATTTCCACA ATGTTACATCAACAGGCACTTCTAGCCATTTTCCTTCCAAAAGGTGCAAAAAGCAACTTCATAAACACA AATTAAATCTTCGGTGAGGTAGTGTGATGCTGCTTCCTCCCAACTCAGCGCACTTCGTCTTCCTCATTCC ACAAAAACCCATAGCCTTCCTTCACTCTGCAGGACTAGTGCTGCCAAGGGTTCAGCTCTACCTACTGGTG TGCTCTTTTGAGCAAGTTGCTTAGCCTCTCTGTAACAAGGACAATAGCTGCAAGCATCCCCAAAGATC ATTGCAGGAGACAATGACTAAGGCTACCAGAGCCGCAATAAAAGTCAGTGAATTTTAGCGTGGTCCTCTC TGTCTCTCCAGAACGGCTGCCACGTGGAATTGCTCTTCCTCCGCTACATCTCGGACTGGGACCTAGACCC TGGCCGCTGCTACCGCGTCACCTGGTTCACCTCCTGGAGCCCCTGCTACGACTGTGCCCGACATGTGGCC GACTTTCTGCGAGGGAACCCCAACCTCAGTCTGAGGATCTTCACCGCGCGCCTCTACTTCTGTGAGGACC GCAAGGCTGAGCCCGAGGGGCTGCGGCGGCTGCACCGCGCCGGGTGCAAATAGCCATCATGACCTTCAA AGGTGCGAAAGGGCCTTCCGCGCAGGCGCAGTGCAGCAGCCCGCATTCGGGATTGCGATGCGGAATGAAT GAGTTAGTGGGGAAGCTCGAGGGGAAGAAGTGGGCGGGGATTCTGGTTCACCTCTGGAGCCGAAATTAAA GATTAGAAGCAGAGAAAAGAGTGAATGGCTCAGAGACAAGGCCCCGAGGAAATGAGAAAATGGGGCCAGG GTTGCTTCTTTCCCCTCGATTTGGAACCTGAACTGTCTTCTACCCCCATATCCCCGCCTTTTTTTCCTTT TTTTTTTTTTGAAGATTATTTTTACTGCTGGAATACTTTTGTAGAAAACCACGAAAGAACTTTCAAAGCC TGGGAAGGGCTGCATGAAAATTCAGTTCGTCTCTCCAGACAGCTTCGGCGCATCCTTTTGGTAAGGGGCT TCCTCGCTTTTTAAATTTTCTTTCTTTCTCTACAGTCTTTTTTGGAGTTTCGTATATTTCTTATATTTTC TTATTGTTCAATCACTCTCAGTTTTCATCTGATGAAAACTTTATTTCTCCTCCACATCAGCTTTTTCTTC TGCTGTTTCACCATTCAGAGCCCTCTGCTAAGGTTCCTTTTCCCTCCCTTTTCTTTCTTTTGTTGTTTCA CATCTTTAAATTTCTGTCTCTCCCCAGGGTTGCGTTTCCTTCCTGGTCAGAATTCTTTTCTCCTTTTTTT TTTTTTTTTTTTTTTTTTTTAAACAAACAAACAAAAAACCCAAAAAAACTCTTTCCCAATTTACTTTCTT CCAACATGTTACAAAGCCATCCACTCAGTTTAGAAGACTCTCCGGCCCCACCGACCCCCAACCTCGTTTT GAAGCCATTCACTCAATTTGCTTCTCTCTTTCTCTACAGCCCCTGTATGAGGTTGATGACTTACGAGACG CATTTCGTACTTTGGGACTTTGATAGCAACTTCCAGGAATGTCACACACGATGAAATATATCTCTGCTGAAG ACAGTGGATAAAAAACAGTCCTTCAAGTCTTCTCTGTTTTTATTCTTCAACTCTCACTTTCTTAGAGTTT ACAGAAAAAATATTTATATACGACTCTTTAAAAAGATCTATGTCTTGAAAATAGAGAAGGAACACAGGTC TGGCCAGGGACGTGCTGCAATTGGTGCAGTTTTGAATGCAACATTGTCCCCTACTGGGAATAACAGAACT GCAGGACCTGGGAGCATCCTAAAGTGTCAACGTTTTTCTATGACTTTTAGGTAGGATGAGAGCAGAAGGT AGATCCTAAAGCATGGTGAGAGGATCAAATGTTTTTATATCAAACATCCTTTATTATTTGATTCATTTG AGTTAACAGTGGTGTTAGTGATAGATTTTTCTATTCTTTTCCCTTGACGTTTACTTTCAAGTAACACAAA CTCTTCCATCAGGCCATGATCTATAGGACCTCCTAATGAGAGTATCTGGGTGATTGTGACCCCAAACCAT CTCTCCAAAGCATTAATATCCAATCATCGCCGTGTATGTTTTAATCAGCAGAAGCATGTTTTTATGTTTGT ACAAAAGAAGATTGTTATGGGTGGGGATGGAGGTATAGACCATGCATGGTCACCTTCAAGCTACTTTAAT AAAGGATCTTAAAATGGGCAGGAGGACTGTGAACAAGACACCCTAATAATGGGTTGATGTCTGAAGTAGC AAATCTTCTGGAAACGCAAACTCTTTTAAGGAAGTCCCTAATTTGAAACACCCACAAACTTCACATATC ATAATTAGCAAACAATTGGAAGGAAGTTGCTTGAATGTTGGGGAGAGGAAAATCTATTGGCTCTCGTGGG TCTCTTCATCTCAGAAATGCCAATCAGGTCAAGGTTTGCTACATTTTGTGTGTGTGATGCTTCTCCCA AAGGTATATTAACTATATAAGAGTTGTGACAAAACAGAATGATAAAGCTGCGAACCGTGGCACACGCT CATAGTTCTAGCTGCTTGGGAGGTTGAGGAGGGAGGATGGCTTGAACACAGGTGTTCAAGGCCAGCCTGG GCAACATAACAAGATCCTGTCTCTCAAAAAAAAAAAAAAAAAAGAAAGAGAGGGCCGGGCGTGGTG GCTCACGCCTGTAATCCCAGCACTTTGGGAGGCCGAGCCGGGCGGATCACCTGTGGTCAGGAGTTTGAGA CCAGCCTGGCCAACATGGCAAAACCCCGTCTGTACTCAAAATGCAAAAATTAGCCAGGCGTGGTAGCAGG CACCTGTAATCCCAGCTACTTGGGAGGCTGAGGCAGGAGAATCGCTTTGAACCCAGGAGGTGGAGGTTGCA GTAAGCTGAGATCGTGCCGTTGCACTCCAGCCTGGGCGACAAGAGCAAGACTCTGTCTCAGAAAAAAAAA AAAAAAAGAGAGAGAGAGAAAGAGAACAATTTGGGAGAGAAGGATGGGAAGCATTGCAAGGAAAT TGTGCTTTATCCAACAAAATGTAAGGAGCCAATAAGGGATCCCTATTTGTCTCTTTGGTGTCTATTTGT CCCTAACAACTGTCTTTGACAGTGAGAAAAATATTCAGAATAACCATATCCCTGTGCCGTTATTACCTAG CAACCCTTGCAATGAAGATGAGCAGATCCACAGGAAAACTTGAATGCACAACTGTCTTATTTTAATCTTA TTGTACATAAGTTTGTAAAAGAGTTAAAAATTGTTACTTCATGTATTCATTTATATTTTATATTATTTTG CGTCTAATGATTTTTTATTAACATGATTTCCTTTTCTGATATATTGAAATGGAGTCTCAAAGCTTCATAA ATTTATAACTTTAGAAATGATTCTAATAACAACGTATGTAATTGTAACATTGCAGTAATGGTGCTACGAA GCCATTTCTCTTGATTTTTAGTAAACTTTTATGACAGCAAATTTGCTTCTGGCTCACTTTCAATCAGTTA AATAAATGATAAATAATTTTGGAAGCTGTGAAGATAAAATACCAAATAAAATAATATAAAAGTGATTTAT ATGAAGTTAAAATAAAAAATCAGTATGATGGAATAAACTTG

[0084] Apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like (APOBEC: Apolipoprotein B mRNA editing enzyme (catalytic polypeptide-like) is evolutionarily conserved cytidine It belongs to the deaminase family. Members of this family are editing enzymes that convert C to U. The N-terminal domain of APOBEC-like proteins is the catalytic domain, and the C-terminal domain is the pseudocatalytic domain. This is the main component. More specifically, the catalytic domain is the zinc-dependent cytidine deaminase domain. It is important for cytidine deamination. APOBEC family members include AP OBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D (now "APOBEC3E" is this one) (referring to APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, and activation-inducible (cytidine)) This includes deaminases, but is not limited to SaBE3, SaKKH-BE3, VQR-BE3, and EQR-B. Numerous modifications including E3, VRER-BE3, VRER-BE3, YE1-BE3, EE-BE3, YE2-BE3, and YEE-BE3. Cytidine deaminase is commercially available and can be obtained from Addgene (Plastic Sumido 85169, 85170, 85171, 85172, 85173, 85174, 85175, 85176, 85177).

[0085] Other exemplary deaminases that can be fused to Cas9 according to aspects of this disclosure are provided below. In some embodiments, the active domain of each sequence, for example, localization sequence Domains without NULS (nuclear localization sequence, nuclear export signal, cytoplasmic localization signal) It should be understood that it may be used.

[0086] Human AID: JPEG0007899276000005.jpg30169 (Underlined: Nuclear localization sequence; Double underlined: Nuclear export signal)

[0087] Mouse AID: JPEG0007899276000006.jpg29169 (Underlined: Nuclear localization sequence; Double underlined: Nuclear export signal)

[0088] Canine AID: JPEG0007899276000007.jpg28169 (Underlined: Nuclear localization sequence; Double underlined: Nuclear export signal)

[0089] Bovine AID: JPEG0007899276000008.jpg29169 (Underlined: Nuclear localization sequence; Double underlined: Nuclear export signal)

[0090] Rat AID JPEG0007899276000009.jpg29169 (Underlined: Nuclear localization sequence; Double underlined: Nuclear export signal)

[0091] Mouse APOBEC-3 JPEG0007899276000010.jpg45167 (italic: nucleic acid editing domain)

[0092] Rat APOBEC-3: JPEG0007899276000011.jpg45167 (italic: nucleic acid editing domain)

[0093] Rhesus macaque APOBEC-3 G: JPEG0007899276000012.jpg41167 (italic: nucleic acid editing domain; underline: cytoplasmic localization signal)

[0094] Chimpanzee APOBEC-3 G: JPEG0007899276000013.jpg38167 (italic: nucleic acid editing domain; underline: cytoplasmic localization signal)

[0095] Green monkey APOBEC-3G: JPEG0007899276000014.jpg40167 (italic: nucleic acid editing domain; underline: cytoplasmic localization signal)

[0096] Human APOBEC-3G: JPEG0007899276000015.jpg38168 (italic: nucleic acid editing domain; underline: cytoplasmic localization signal)

[0097] Human APOBEC-3F: JPEG0007899276000016.jpg38169 (italic: nucleic acid editing domain)

[0098] Human APOBEC-3B: JPEG0007899276000017.jpg38169 (italic: nucleic acid editing domain)

[0099] Rat APOBEC-3B: MQPQGLGPNAGMGPVCLGCSHRRPYSPIRNPLKKLYQQTFYFHFKNVRYAWGRKNNFLCYEVNGMDCALPVPLRQGVFRK QGHIHAELCFIYWFHDKVLRVLSPMEEFKVTWYMSWSPCSKCAEQVARFLAAHRNLSLAIFSSRLYYYLRNPNYQQKLCR LIQEGVHVAAMDLPEFKKCWNKFVDNDGQPFRPWMRLRINFSFYDCKLQEIFSRMNLLREDVFYLQFNNSHRVKPVQNRY YRRKSYLCYQLERANGQEPLKGYLLYKKGEQHVEILFLEKMRSMELSQVRITCYLTWSPCPNCARQLAAFKKDHPDLILR IYTSRLYFWRKKFQKGLCTLWRSGIHVDVMDLPQFADCWTNFVNPQRPFRPWNELEKNSWRIQRRLRRIKESWGL

[0100] Bovine APOBEC-3B: DGWEVAFRSGTVLKAGVLGVSMTEGWAGSGHPGQGACVWTPGTRNTMNLLREVLFKQQFGNQPRVPAPYYRRKTYLCYQL KQRNDLTLDRGCFRNKKQRHAERFIDKINSLDLNPSQSYKIICYITWSPCPNCANELVNFITRNNHLKLEIFASRLYFHW IKSFKMGLQDLQNAGISVAVMTHTEFEDCWEQFVDNQSRPFQPWDKLEQYSASIRRRLRQRILTAPI

[0101] Chimpanzee APOBEC-3B: MNPQIRNPMEWMYQRTFYYNFENEPILYGRSYTWLCYEVKIRRGHSNLLWDTGVFRGQMYSQPEHHAEMCFLSWFCGNQL SAYKCFQITWFVSWTPCPDCVAKLAKFLAEHPNVTLTISAARLYYYWERDYRRALCRLSQAGARVKIMDDEEFAYCWENF VYNEGQPFMPWYKFDDNYAFLHRTLKEIIRHLMDPDTFTFNFNNDPLVLRRHQTYLCYEVERLDNGTWVLMDQHMGFLCN EAKNLLCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGQVRAFLQENTHVRLRIFAARIYDYDPLYK EALQMLRDAGAQVSIMTYDEFEYCWDTFVYRQGCPFQPWDGLEEHSQALSGRLRAILQVRASSLCMVPHRPPPPQSPGP CLPLCSEPPLGSLLPTGRPAPSLPFLLTASSFPPPASLPPLPSLSLSPGHLPVPSFHSLTSCSIQPPCSSRIRETEGWA SWEDISH

[0102] Human APOBEC-3C: JPEG0007899276000018.jpg 20169

[0103] Gorilla APOBEC3C JPEG0007899276000019.jpg20169 (italic: nucleic acid editing domain)

[0104] Human APOBEC-3A: JPEG0007899276000020.jpg20169 (italic: nucleic acid editing domain)

[0105] Rhesus macaque APOBEC-3A: JPEG0007899276000021.jpg20169 (italic: nucleic acid editing domain)

[0106] Bovine APOBEC-3A: JPEG0007899276000022.jpg20169 (italic: nucleic acid editing domain)

[0107] Human APOBEC-3H: JPEG0007899276000023.jpg20169 (italic: nucleic acid editing domain)

[0108] Rhesus macaque APOBEC-3H: MALLTAKTFSLQFNNKRRVNKPYYPRKALLCYQLTPQNGSTPTRGHLKNKKKDHAEIRFINKIKSMGLDETQCYQVTCYL TWSPCPSCAGELVDFIKAHRHLNLRIFASRLYYHWRPNYQEGLLLLCGSQVPVEVMGLPEFTDCWENFVDHKEPPSFNPS EKLEELDKNSQAIKRRLERIKSRSVDVLENGLRSLQLGPVTPSSSIRNSR

[0109] Human APOBEC-3D: JPEG0007899276000024.jpg37169 (italic: nucleic acid editing domain)

[0110] Human APOBEC-1: MTSEKGPSTGDPTLRRRIEPWEFDVFYDPRELRKEACLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERDFHPSM SCSITWFLSWSPCWECSQAIREFLSRHPGVTLVIYVARLFWHMDQQNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYP PGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLTFFRLHLQNCHYQTIPPHILLATGLIHPSVAWR

[0111] Mouse APOBEC-1 : MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSVWRHTSQNTSNHVEVNFLEKFTTERYFRPNT RCSITWFLSWSPCGECSRAITEFLSRHPYVTLFIYIARLYHHTDQRNRQGLRDLISSGVTIQIMTEQEYCYCWRNFVNYP PSNEAYWPRYPHLWVKLYVLELYCIILGLPPCLKILRRKQPQLTFFTITLQTCHYQRIPPHLLWATGLK

[0112] Rat APOBEC-1 : MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHVEVNFIEKFTTERYFCPNT RCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYS PSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLK

[0113] Human APOBEC-2: MAQKEEAAVATEAASQNGEDLENLDDPEKLKELIELPPFEIVTGERLPANFFKFQFRNVEYSSGRNKTFLCYVVEAQGKG GQVQASRGYLEDEHAAAHAEEAFFNTILPAFDPALRYNVTWYVSSSPCAACADRIIKTLSKTKNLRLLILVGRLFMWEEP EIQAALKKLKEAGCKLRIMKPQDFEYVWQNFVEQEEGESKAFQPWEDIQENFLYYEEKLADILK

[0114] Mouse APOBEC-2: MAQKEEAAEAAAPASQNGDDLENLEDPEKLKELIDLPPFEIVTGVRLPVNFFKFQFRNVEYSSGRNKTFLCYVVEVQSKG GQAQATQGYLEDEHAGAHAEEAFFNTILPAFDPALKYNVTWYVSSSPCAACADRILKTLSKTKNLRLLILVSRLFMWEEP EVQAALKKLKEAGCKLRIMKPQDFEYIWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK

[0115] Rat APOBEC-2: MAQKEEAAEAAAPASQNGDDLENLEDPEKLKELIDLPPFEIVTGVRLPVNFFKFQFRNVEYSSGRNKTFLCYVVEAQSKG GQVQATQGYLEDEHAGAHAEEAFFNTILPAFDPALKYNVTWYVSSSPCAACADRILKTLSKTKNLRLLILVSRLFMWEEP EVQAALKKLKEAGCKLRIMKPQDFEYLWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK

[0116] Bovine APOBEC-2: MAQKEEAAAAAEPASQNGEEVENLEDPEKLKELIELPPFEIVTGERLPAHYFKFQFRNVEYSSGRNKTFLCYVVEAQSKG GQVQASRGYLEDEHATNHAEEAFFNSIMPTFDPALRYMVTWYVSSSPCAACADRIVKTLNKTKNLRLLILVGRLFMWEEP EIQAALRKLKEAGCRLRIMKPQDFEYIWQNFVEQEEGESKAFEPWEDIQENFLYYEEKLADILK

[0117] Petromyzon marinus CDA1 (pmCDAl) MTDAEYVRIHEKLDIYTFKKQFFNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLR DNPGQFTINWYSSWSPCADCAEKILEWYNQELRGNGHTLKIWACKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCR KIFIQSSHNQLNENRWLEKTLKRAEKRRSELSFMIQVKILHTTKSPAV

[0118] Human APOBEC3G D316R D317R MKPHFRNTVERMYRDTFSYNFYNRPILSRRNTVWLCYEVKTKGPSRPPLDAKIFRGQVYSELKYHPEMRFFHWFSKWRKL HRDQEYEVTWYISWSPCTKCTRDMATFLAEDPKVTLTIFVARLYYFWDPDYQEALRSLCQKRDGPRATMKFNYDEFQHCW SKFVYSQRELFEPWNNLPKYYILLHFMLGEILRHSMDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGF LCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDYRVTC FTSWSPCFSCAQEMAKFISKKHVSLCIFTARIYRRQGRC QEGLRTLAEAGAKISFTYSEFKHCWDTFVDHQGCPFQPWDGLDEHSQDLSGRLRAILQNQEN

[0119] Human APOBEC3G chain A MDPPTFTFNFNNEPWWGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQDY RVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYDDQGRCQEGLRTLAEAGAKISF TYSEFKHCWDTFVDHQGCP FQPWDGLD EHSQDLSGRLRAILQ

[0120] Human APOBEC3G chain A D120R D121R MDPPTFTFNFNNEPWVRGRHETYLCYEVERMHNDTWVLLNQRRGFLCNQAPHKHGFLEGRHAELCFLDVIPFWKLDLDQD YRVTCFTSWSPCFSCAQEMAKFISKNKHVSLCIFTARIYRRQGRCQEGLRTLAEAGAKISFMTYSEFKHCWDTFVDHQGC PFQPWDGLDEHSQDLSGRLRAILQ

[0121] The term "deaminase" or "deaminase domain" refers to a substance that catalyzes deamination reactions. It refers to a protein or a fragment of a protein.

[0122] "Detecting" refers to identifying the presence, absence, or quantity of an analyte that should be detected. In one embodiment, a sequence change in a polynucleotide or polypeptide is detected. In another embodiment, the presence of an indel is detected.

[0123] A "detectable label" is one that, when attached to the target molecule, is detected spectroscopically, photochemically, or biochemically. This refers to a composition that makes it detectable by scientific, immunochemical, or chemical means. For example, useful labels include radioactive isotopes, magnetic beads, metal beads, colloidal particles, and fluorescence. Dyes, high electron density reagents, enzymes (e.g., those commonly used in ELISA), biotin, digonol It contains xygenin or hapten.

[0124] "Disease" is any condition that damages or interferes with the normal function of cells, tissues, or organs. "T" means a disorder. In certain embodiments, a disease suitable for treatment with the composition of the present invention. This includes point mutations, splicing events, immature stop codons, or misfolding events. It is related to elephants.

[0125] A "DNA-binding protein domain" refers to a polypeptide or a fragment thereof that binds to DNA. It tastes good. In some embodiments, the DNA-binding protein domain binds sequence-specific DNA. It is a zinc finger or TALE domain that has synactivation. In other embodiments, DN The A-binding protein domain is the domain of the CRISPR-Cas protein that binds to DNA (e.g.) For example, Cas9, including those that bond to the protospacer adjacent motif (PAM). In this embodiment, the DNA-binding protein domain is a polynucleotide (e.g., a single guide It forms a complex with RNA, and this complex is formed by gRNA and the protospacer adjacent motif. It binds to a specific DNA sequence. In one embodiment, the DNA-binding protein domain is It may contain case activity (e.g., nCas9) or be catalytically inactive (e.g., dCas9, di In other embodiments, the DNA-binding protein (TALE). The main component is a catalytically inactive variant of the homing endonuclease I-SceI, or This is the DNA-binding domain of the TALE protein AvrBs4. For example, Gabsalilow et al., Nucleic Ac See ids Research, Volume 41, Issue 7, 1 April 2013, Pages e83. In one aspect The DNA-binding protein domain fuses with a catalytically active domain (e.g., FokI, MutH). They are combined. In a particular embodiment, the zinc finger domain is an endonuclease. It is fused to the catalytic domain of FokI. In other embodiments, TALE is used to fuse with site-specific DNA. It is fused with Muth, which contains King activity.

[0126] As used herein, the term "effective dose" means a sufficient amount to induce a desired biological response. This refers to the amount of a biologically active drug. In certain embodiments, the effective amount is obtained by ingesting the plasmid. A base editor sufficient to express an active base editing system in lancet cells. —The amount of two or more plasmids, including parts of the system. The effective amount of the drug (e.g., fusion protein) is, for example, the desired biological response, for example, editing The specific allele, genome, or target site to be targeted, the cell or tissue to be targeted, and It can vary depending on various factors, such as the medications used.

[0127] A "fragment" refers to a part of a polypeptide or nucleic acid molecule. This part is the reference nucleic acid. At least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% of the total length of the molecule or polypeptide. , or containing 90%. Fragments are 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 3 It may contain 00, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids. ru.

[0128] "Hybridization" refers to hydrogen bonding between complementary nucleic acid bases, as defined by Watson- These can be click, Hougsteen, or reverse Hougsteen hydrogen bonds. For example, Adenine and thymine are complementary nucleic acid bases that form a pair by forming a hydrogen bond.

[0129] The term "base repair inhibitor" or "IBR" refers to nucleic acid repair enzymes, such as base excision repair enzymes. This refers to a protein that can inhibit the activity of an enzyme. In one embodiment, IBR is an ino It is an inhibitor of syn base excision repair. Examples of base repair inhibitors include APE1 and Endo III. , Endo IV, Endo V, Endo VIII, Fpg, hOGGl, hNEILl, T7 Endol, T4 PDG, UDG, hSMUGL And inhibitors of hAAG are also mentioned. In one embodiment, IBR inhibits Endo V or hAAG. It is a harmful factor. In one embodiment, the IBR is catalytically inactive EndoV or catalytically inactive It is hAAG.

[0130] "Intein" is a term that cuts itself out and then leaves behind fragments (extein). (extein) is involved in a process known as protein splicing, where peptide bonds are formed. Inteins are protein fragments that can be linked together. It is also called "intei". The intein cuts itself out and ligates the rest of the protein. The process is referred to herein as "protein splicing" or "intene-mediated". This is called "protein splicing." In one aspect, the precursor protein (integer) The intein of the intein-containing protein prior to in-mediated protein splicing is It originates from two genes. Such an intein is referred to herein as a split intein. These are called inteins (for example, split inteins-N and split inteins-C). For example, shear In bacteria, DnaE, ​​the catalytic subunit a of DNA polymerase III, is derived from two separate genes. Encoded by the genes dnaE-n and dnaE-c. Encoded by the dnaE-n gene. In this specification, intein may be referred to as "inteiin N". The intein being used may be referred to as "Intein C" in this specification.

[0131] Other intein systems can also be used. For example, the dnaE intein, i.e., Cfa-N (for example) Based on the intein pairs of (partitioned intein-N) and Cfa-C (e.g., partitioned intein-C). Synthetic inteins are described (for example, Steven, incorporated herein by reference) s et al., J Am Chem Soc. 2016 Feb. 24; 138(7):2162-5). Use in accordance with this disclosure Non-restrictive examples of intelligent pairs that can perform this include the Cfa DnaE intelligent and the Ssp GyrB intelligent. Intein, Ssp DnaX Intein, Ter DnaE3 Intein, Ter ThyX Intein, Rma DnaB Intein Tein, and Cne Prp8 Intein (for example, the U.S. Patent incorporated herein by reference) Examples include those listed in Permit No. 8,394,604.

[0132] This provides exemplary nucleotide and amino acid sequences of the intein.

[0133] DnaE intein-N DNA:TGCCTGTCATACGAACCGAGATACTGACAGTAGAATATGGCCTTCTGCC AATCGGG AAGATTGTGGAGAAACGGATAGAATGCACAGTTTACTCTGTCG ATAACAATGGTAACATTTATACTCAGCCAGTTGCCC AGTGGCACGACCGG GGAGAGCAGGAAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAG GGCCACTAAGGACC ACAAATTTATGACAGTCGATGGCCAGATGCTGCCTA TAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGAGTTGA CAACCTT CCTAAT

[0134] DNA Intein-N Protein:CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDR GEQEVFEYCLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRVDNL PN

[0135] DnaE intein-C DNA:ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGA TATTGGA GTCGAAAGAGATCACAACTTTGCTCTGAAGAACGGATTCATAG CTTCTAAT

[0136] Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN

[0137] Cfa-N DNA: TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCC TATTGGAAAGATTGTCGAAGAGAGAATTG AATGCACAGTATATACTGTAG ACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCGC GGCGAAC AAGAAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACG AGCAACTAAAGATCATAAATTCATGACCACTGACGG GCAGATGTTGCCAA TAGATGAGATATTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTG CCA

[0138] Cfa-N protein: CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNR GEQEVFEYCLEDGSIIRATKDHKFMTTDG QMLPIDEIFERGLDLKQVDGL P

[0139] Cfa-C DNA: ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAG GAAAGTAAAGATAATATCTCGAAAAAGTC TTGGTACCCAAAATGTCTATG ATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGTA GCCAGCA AC

[0140] Cfa-C protein:MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLV ASN

[0141] In order to join the N-terminus of the divided Cas9 and the C-terminus of the divided Cas9, intein N and intein The tein C can be fused to the N-terminal and C-terminal portions of the segmented Cas9, respectively. In some embodiments, intein-N is the C-terminus of the N-terminus portion of the divided Cas9. They fuse together, that is, to form the structure N--[N-terminal portion of partitioned Cas9]-[intei-N]--C. In some embodiments, intein-C is the N-terminus of the C-terminus portion of the divided Cas9. They fuse together, that is, to form the structure N-[intei-C]--[C-terminal portion of split Cas9]-C. The intein is used to link the protein (e.g., split Cas9) at the site where it is fused. The mechanism of thein-mediated protein splicing is incorporated herein by reference, for example. As described in Shah et al., Chem Sci. 2014; 5(1):446-461, the technique It is publicly known in the art. The methods for designing and using the inteins are publicly known in the art. For example, by WO2014004336, WO2017132580, US20150344549 and US20180127780 These are described, and each of them is incorporated herein by reference in its entirety.

[0142] The terms “isolated,” “purified,” or “biologically pure” refer to their natural state. When found in its natural state, the associated components are usually removed to varying degrees. Refers to the degree of separation from the original source or surrounding environment. "Isolation" refers to the degree of separation from the original source or surrounding environment. It shows a higher degree of separation. "Purified" or "biologically pure" proteins are impure. The substance does not substantially affect the biological properties of the protein or cause other harmful consequences. In this way, other substances are sufficiently removed. That is, the nucleic acid or peptide of the present invention is When produced by recombinant DNA technology, the cellular material, viral material, or culture medium is essentially the same. If not included in the original, or if chemically synthesized, chemical precursors or other chemical substances. If it substantially contains no impurities, it is purified. Purity and homogeneity are typically determined by analysis. Chemical technologies, such as polyacrylamide gel electrophoresis or high-performance liquid chromatography. Determined using [this method]. The term "purified" refers to nucleic acids or proteins that have been processed on an electrophoresis gel. This could essentially mean producing one band. Modifications, for example, phosphorylation or globulinization. For proteins that can undergo lycosylation, different modifications are purified separately. This can result in different isolated proteins.

[0143] "Isolated polynucleotides" are derived from the natural genome of organisms from which the nucleic acid molecules of the present invention originate. This refers to nucleic acids (e.g., DNA) that do not contain genes adjacent to the gene in question. Therefore, The term is used, for example, to describe a plasmid or virus that is incorporated into a vector or replicates autonomously. Integrated; integrated into the genomic DNA of prokaryotes or eukaryotes; or independently of other sequences. Another molecule is established (e.g., cDNA produced by PCR or restriction endonuclease digestion). This includes recombinant DNA that exists as a genome or cDNA fragment. Furthermore, this term means High-frequency transcription of RNA molecules from DNA molecules, as well as encoding further polypeptide sequences. It contains recombinant DNA, which is part of the hybrid gene.

[0144] "Isolated polypeptide" refers to the polypeptide of the present invention that has been isolated from naturally occurring components. It means "cydo". Typically, polypeptides are proteins that it associates with in its natural state. It is isolated when it is at least 60% free by weight from qualitative and natural organic molecules. Preferably, the preparation is at least 75% by weight, more preferably at least 90%, and most Preferably, at least 99% is the polypeptide of the present invention. Butids are recombinant nucleic acids that encode polypeptides, for example, extracted from natural sources. It can be obtained by expression; or by chemically synthesizing the protein. Purity is at the discretion of the individual. Appropriate methods for the purpose, such as column chromatography and polyacrylamide gel electrophoresis. It can be measured by HPLC analysis or other methods.

[0145] As used herein, the term "linker" refers to two molecules or parts (e.g., fusion protein). A bond (e.g., a covalent bond) that connects two domains of a molecule, a chemical group, or a molecule. In one embodiment, the linker contains an RNA progesterone containing a Cas9 nuclease domain. The gRNA-binding domain of a ramming nuclease and the catalytic domain of a nucleic acid editing protein To connect. In one embodiment, the linker connects dCas9 and the nucleic acid editing protein. In terms of structure, a linker is positioned between two groups, molecules, or other parts, or between them They are adjacent to each other, connected to each other via covalent bonds, thus the two of them Linking. In one embodiment, the linker is an amino acid or a plurality of amino acids (e.g., pe It is a peptide or protein. In one embodiment, the linker is an organic molecule, group, or protein. It is a rimer, or chemical part. In one embodiment, the linker has a length of 5 to 200 amino acids. It is an acid, for example, with lengths of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102 , 103, 104, 105, 110, 120, 130, 140, 150, 160, 175, 180, 190, or 200 amino acids Therefore, longer or shorter linkers are also being considered. In one embodiment, the linker is It contains the amino acid sequence SGSETPGTSESATPES, which may also be called the XTEN linker. The linker contains the amino acid sequence SGGS. In one embodiment, the linker is (SGGS) n , (GG GS) n (GGGGS) n , (G) n、 (EAAAK) n (GGS) n , SGSETPGTSESATPES, or (XP) n motif , or any combination thereof, where n is an independent integer from 1 to 30, and X n is any amino acid. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9 , 10, 11, 12, 13, 14, or 15.

[0146] In some embodiments, the nucleic acid base editor domain is SGGSSGSETPGTSESATP ESSGGS, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPG TSTEPSEGSAPGSPAGSPTSTEEGTSTE PSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS no Ami They are fused via a linker containing a no-acid sequence. In some embodiments, nucleic acid bases The Ditter domain contains the amino acid sequence SGSETPGTSESATPES, which can also be called the XTEN linker. They are fused via a linker. In one embodiment, the linker has a length of 24 amino acids. In one embodiment, the linker contains the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In one embodiment, the linker has a length of 40 amino acids. In another embodiment, the linker is , containing the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS. In one embodiment, The linker has a length of 64 amino acids. In one embodiment, the linker is the amino acid sequence SG Includes GSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGS SGGS. In one embodiment In this configuration, the linker has a length of 92 amino acids. In one embodiment, the linker is ami ∇ acid sequence PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAP GTSTEPSEGS Includes APGTSESATPESGPGSEPATS.

[0147] A "marker" is a term used to describe a condition that exhibits changes in expression levels or activity associated with a disease or disorder. It means a protein or polynucleotide.

[0148] As used herein, the term "mutation" means a change within a sequence, for example, in a nucleic acid or amino acid. Substitution of a residue within an acid sequence by another residue, or deletion of one or more residues within a sequence. Alternatively, it refers to an insertion. A mutation, as used herein, typically involves identifying the original residue and then... By identifying the position of residues within the sequence and then identifying newly substituted residues, Various methods for producing amino acid substitutions (mutations) are provided herein. This is well known in the field of the art, for example, Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Provided by Spring Harbor, NY (2012).

[0149] As used herein, the terms “nucleic acid” and “nucleic acid molecule” mean nucleic acid bases and Compounds containing an acidic moiety, such as nucleosides, nucleotides, or polynucleotides. This refers to polymeric nucleic acids, such as nucleic acid molecules containing three or more nucleotides. This is a linear chain in which adjacent nucleotides are linked to each other via phosphodiester bonds. It is a child. In one aspect, "nucleic acid" is an individual nucleic acid residue (e.g., nucleotide and / It refers to a nucleic acid (or nucleoside). In one aspect, "nucleic acid" refers to three or more individual nucleosides. This refers to an oligonucleotide chain containing an oligonucleotide residue. The term "oligo" as used herein refers to an oligonucleotide chain containing an oligonucleotide residue. "Nucleotide" and "polynucleotide" are polymers of nucleotides (for example, a small amount of... It can be used interchangeably to refer to a chain of at least three nucleotides. In one embodiment "Nucleic acids" include RNA and single-stranded and / or double-stranded DNA. Nucleic acids are, for example, genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, It may be naturally occurring in relation to chromatids or other naturally occurring nucleic acid molecules. On the other hand, nucleic acid molecules include, for example, molecules that exist unnaturally, recombinant DNA or RNA, artificial chromosomes, and manipulative molecules. The created genome, or fragments thereof, or synthetic DNA, RNA, DNA / RNA hybrids, This refers to unnaturally occurring molecules containing nucleotides or nucleosides. It is possible. Furthermore, the terms “nucleic acid,” “DNA,” “RNA,” and / or similar terms refer to the nucleus. Acid analogs, such as analogs having a skeleton other than a phosphodiester skeleton. Nucleic acids are of natural origin. Chemically synthesized, produced using a recombinant expression system and purified as needed, purified from It is possible to do, etc. In the case of chemically synthesized molecules, nucleic acids, in appropriate cases, for example For example, nuclei such as chemically modified bases or sugars, and analogs having skeletal modifications. May contain osid analogs. Nucleic acid sequences are shown in the 5'-3' direction unless otherwise specified. In one embodiment, nucleic acids are natural nucleosides (e.g., adenosine, thymidine, guanoside). Syn, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine , and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-cinnabar Othymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2 - Aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5 -Propynyluridine, C5-Propynylcytidine, C5-Methylcytidine, 2-Aminoadeno Syn, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguani (O(6)-methylguanine and 2-thiocytidine); chemically modified bases; biologically modified Modified bases (e.g., methylated bases); inserted bases; modified sugars (e.g., 2'-fluororibose, ribose); Bose, 2'-deoxyribose, arabinose, and hexose; and / or modified Is it a phosphate group (e.g., phosphorothioate and 5'-N-phosphoamidite bond)? , or including them.

[0150] The terms "nuclear localization sequence," "nuclear localization signal," or "NLS" refer to the structure of a protein. This refers to an amino acid sequence that promotes translocation into the cell nucleus. Nuclear localization sequences are used in this field. It is publicly known, for example, filed on November 23, 2000, and registered as WO / 2001 / 038547 on May 31, 2001. It is described in Plank et al. in the publicly published international PCT application PCT / EP 2000 / 011690, among which The contents are incorporated herein by reference for the disclosure of exemplary nuclear localization sequences. In the embodiment, NLS is, for example, Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4 This is an optimized NLS described by 172. In one embodiment, NLS is an amino acid The sequence KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR This includes RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.

[0151] The term "nucleic acid programmable DNA-binding protein" or "napDNAbp" is derived from napDN Abp associates with nucleic acids (e.g., DNA or RNA), such as guide nucleic acids, which guide Abp to a specific nucleic acid sequence. This represents a protein that interacts with guide RNA. For example, the Cas9 protein interacts with guide RNA. It can bind to guide RNA that leads to a complementary specific DNA sequence. In some embodiments, napDNAbp is a Cas9 domain, for example, nuclease activity Cas9, Cas9 niccas ( nCas9, or nuclease-inactivated Cas9 (dCas9). Nucleic acid programmable DNA. Examples of binding proteins include Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, and C2. Examples include, but are not limited to, c1 and C2c3. Other nucleic acid programmable DNs. A-binding proteins are also within the scope of this disclosure, even if they are not specifically listed. .

[0152] As used herein, "to obtain" as in "to obtain a drug" means This includes synthesizing, purchasing, or otherwise obtaining drugs.

[0153] As used herein, “patient” or “subject” refers to a person diagnosed with a disease or disorder. or subjects that have or are suspected of having such, such as mammals. In one aspect, the term "patient" refers to someone who is more likely than average to develop a disease or disability. This refers to mammals. Examples of patients include humans, non-human primates, cats, dogs, pigs, and cows. Cats, horses, goats, sheep, rodents (e.g., mice, rabbits, rats, guinea pigs) and other mammals that may benefit from the treatments disclosed herein. The exemplary human patient may be male and / or female. “Patients who need it” In this specification, "those who require it" refers to those diagnosed with a disease or disability. They are referred to as patients suspected of having this condition.

[0154] The terms "RNA programmable nuclease" and "RNA-inducible nuclease" refer to cleavage It is used in conjunction with one or more non-target RNAs (e.g., bound to or associated with them). In one embodiment, if the RNA programmable nuclease is in complex with RNA, the nuclease Rease: Can be called an RNA complex. Typically, the bound RNA is called guide RNA (gRNA). gRNA can exist as a complex of two or more RNA molecules, or as a single RNA molecule. They can also exist. A gRNA that exists as a single RNA molecule is called a single guide RNA (sgRNA). Although it can be detected, "gRNA" exists as a single molecule or as a complex of two or more molecules. It is used interchangeably to refer to a guide RNA that exists as a single RNA species. Typically, it is used as a single RNA species. The gRNAs present are (1) domains that share homology with the target nucleic acid (for example, Cas9 multiple domains to the target) (1) A domain that directs the binding of the fusion; and (2) a domain that binds to the Cas9 protein. Includes the main. In one embodiment, domain (2) corresponds to a sequence known as tracrRNA. Accordingly, it includes a stem-loop structure. For example, in some embodiments, domain (2) is Jinek et al., Science 337:816-821 (2012) (The entire contents of this document are incorporated herein by reference) It is identical or homologous to tracrRNA provided to (the domain). gRNA (e.g., domain) Another example (including 2) is titled "Switchable Cas9 Nucleases and Uses Thereof" and is 20 U.S. Provisional Patent Application USSN61 / 874,682 and "Delivery System Fo" filed on September 6, 2013. U.S. Provisional Patent Application USSN61, titled "r Functional Nucleases," was filed on September 6, 2013. These can be found in / 874,746, and the entire contents of each of them are incorporated herein by reference. In some embodiments, the gRNA comprises two or more domains (1) and (2), and "extension It may be referred to as "extended gRNA." For example, extended gRNA is as described herein. For example, it binds to two or more Cas9 proteins and targets nucleic acids in two or more different regions. It binds to the target site. The gRNA contains a nucleotide sequence that complements the target site, and this is the target site. It mediates the binding of the nuclease / RNA complex to and the sequence specificity of the nuclease:RNA complex. Provides. In one embodiment, RNA programmable nucleases are (CRISPR-related systems) ) Cas9 endonuclease, for example, Cas9 (Csnl) derived from Streptococcus pyogenes. (For example, "Complete genome sequence of an Ml strain of Streptococcus pyogenes.") Ferretti JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Na jar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, Mc Laughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001); "CRISPR RNA mat uration by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chy linski K., Sharma CM., Gonzales K., Chao Y., Pirzada ZA, Eckert MR, Vogel J. (See Charpentier E., Nature 471:602-607 (2011)).

[0155] As used herein in relation to proteins or nucleic acids, the term “recombinant” means self This refers to proteins or nucleic acids that do not exist in nature but are products of human engineering. For example, In some embodiments, recombinant proteins or nucleic acid molecules are any naturally occurring Compared to the array, at least one, at least two, at least three, at least four, and fewer amino acids or nuclei containing at least five, at least six, or at least seven mutations Includes the Otid sequence.

[0156] "Decrease" means a negative change of at least 10%, 25%, 50%, 75%, or 100%. .

[0157] "Reference" means a standard or comparative condition. In one embodiment, the reference is an Nucleic acid base editor expressed in intein-containing fragments for thein-dependent reconstruction - Under the same conditions and within the same cell, the full-length nucleic acid bases expressed in a single plasmid. This is the activity of the filter.

[0158] A "reference sequence" is a predefined sequence used as the basis for sequence comparison. A reference sequence is: This can be a subset or the whole of a particular sequence; for example, a full-length cDNA or gene sequence. A segment, or complete cDNA or gene sequence. For polypeptides, see reference poly The length of a peptide sequence is generally at least about 16 amino acids, at least about 20 amino acids, and less At least about 25 amino acids, more preferably about 35 amino acids, about 50 amino acids, or about 10 It is 0 amino acids. For nucleic acids, the length of the reference nucleic acid sequence is generally at least about 50 nucleotides. Cleotide, at least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides Otid or approximately 300 nucleotides or any integer around or between them be.

[0159] The term "single nucleotide polymorphism (SNP)" refers to a mutation in a single nucleotide that occurs at a specific location in the genome. Here, each mutation exists to a degree that is recognizable within the population (e.g., >1%). At specific base locations in the human genome, C nucleotides can appear in most individuals, In a small number of individuals, that position is occupied by A. This is because there is an SNP at this specific position, and C Alternatively, this allele is a variation of two nucleotides, A. P underlies differences in susceptibility to disease. The severity of the disease and the body's response to treatment are also related. This is an expression of genetic variation. SNPs are found in the coding region and non-coding region of a gene. or it may be located in the intergenetic region (the region between genes). In one embodiment, SNPs within a gene sequence affect the amino acid distribution of the protein produced due to the degenerate nature of the gene code. The columns do not necessarily change. There are two types of SNPs in the code region: synonymous SNPs and non-synonymous SNPs. Synonymous SNPs do not affect the protein sequence, but non-synonymous SNPs alter the amino acid sequence of the protein. To cause. There are two types of non-synonymous SNPs: missense and nonsense. Protein coding SNPs outside of a region are involved in gene splicing, transcription factor binding, and messenger RNA degradation. or may affect the sequence of non-coding RNA. Gene expression that occurs is called an eSNP (expression SNP) and can be located upstream or downstream of a gene. A riant (SNV) is a single-nucleotide variation with no frequency limit and can occur in somatic cells. Somatic single-nucleotide variations can also be called single-nucleotide modifications.

[0160] "Specifically binding" means recognizing the polypeptide and / or nucleic acid molecule of the present invention. It binds, but does not substantially recognize and bind to other molecules in the sample (e.g., a biological sample). nucleic acid molecules, polypeptides, or complexes thereof (e.g., nucleic acid programmable DN This refers to a compound or molecule (including the A-binding domain and guide nucleic acid).

[0161] Nucleic acid molecules useful in the method of the present invention are polypeptides or fragments thereof. It contains any nucleic acid molecule that is 100% identical to the endogenous nucleic acid sequence. It is not necessary, but it typically shows substantial identity. "Substantial identity" for endogenous sequences. Polynucleotides typically consist of at least one strand of a double-stranded nucleic acid molecule and a hive. It can be reduced. Nucleic acid molecules useful in the method of the present invention are the polypeptide of the present invention. It includes any nucleic acid molecule encoding a cytoplasm or a fragment thereof. Such nucleic acid molecules are endogenous It does not need to be 100% identical to the nucleic acid sequence, but typically shows substantial identity. In contrast, polynucleotides that possess "substantial identity" are typically composed of a small number of double-stranded nucleic acid molecules. Even without it, it can hybridize with one chain. "Hybridizing" means that a species Complementary polynucleotide sequences under various stringency conditions (e.g., as described herein) This means forming a pair that creates a double-stranded molecule between (the genes) or between (parts of) them. For example, Wahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, A. R (See Methods Enzymol. 152:507, 1987).

[0162] For example, the stringent salt concentration is typically less than approximately 750 mM for NaCl and 75 mM for citric acid. Trisodium, preferably less than about 500 mM NaCl, and 50 mM trisodium citrate, more preferably The main components are less than 250 mM NaCl and 25 mM trisodium citrate. Low stringing C-hybridization can be obtained in the absence of organic solvents, such as formamide. On the other hand, high stringency hybridization requires at least about 35% formaldehyde. It can be obtained in the presence of an amide, more preferably at least about 50% formamide. The triggering temperature conditions are typically at least about 30°C, more preferably at least about 37°C. The temperature will include °C, most preferably at least about 42°C. During hybridization The concentration of the surfactant (e.g., sodium dodecyl sulfate (SDS)) and the content of the carrier DNA are also important factors. Various additional parameters, such as input or exclusion, are well known to those skilled in the art. By combining these various conditions, various levels of stringency This is achieved. In one embodiment, hybridization is performed at 30°C with 750 mM NaCl, 75 This occurs in mM trisodium citrate and 1% SDS. In another embodiment, hybridize The solution is prepared at 37°C with 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, and 35% formaldehyde. This occurs in denatured salmon sperm DNA (ssDNA) at a concentration of 100 μg / ml. In another embodiment, Ebridization was performed at 42°C with 250 mM NaCl, 25 mM trisodium citrate, and 1% This occurs in SDS, 50% formamide, and 200 μg / ml ssDNA. Useful variations of these conditions The explanation will be readily apparent to those skilled in the art.

[0163] In most applications, the washing process following hybridization is also done by stringen The conditions differ. Washing stringency conditions are defined by salt concentration and temperature. This can be done. As mentioned above, washing stringency reduces the salt concentration or warm It can be increased by raising the degree. For example, strips for the washing process The ideal salt concentration is preferably less than approximately 30 mM NaCl and 3 mM trisodium citrate. Yes, most preferably less than 15 mM NaCl and 1.5 mM trisodium citrate. Washing The stringent temperature conditions for the process are typically at least about 25°C, more preferably The temperature includes at least about 42°C, and more preferably at least about 68°C. In one embodiment, The washing process is carried out at 25°C with 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. This is carried out inside. In a more preferred embodiment, the washing step is performed at 42°C with 15 mM NaCl, 1.5 This is carried out in mM trisodium citrate and 0.1% SDS. In a more preferred embodiment... The washing process is carried out at 68°C with 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Further variations of these conditions will be readily apparent to those skilled in the art. It is likely. Hybridization techniques are well known to those skilled in the art, for example, Benton a nd Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wil ey Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular C loning: Described in *A Laboratory Manual*, Cold Spring Harbor Laboratory Press, New York. It is being done.

[0164] "Split" means to be divided into two or more pieces.

[0165] "Split Cas9 protein" or "split Cas9" refers to two separate nucleotides. The Cas9 protein is provided as an N-terminal and C-terminal fragment encoded by the sequence. The polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein are splices. It can be remodeled to form a Cas9 protein. In certain embodiments, The Cas9 protein is described, for example, in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-9. As described in 49, 2014, or Jiang et al. (2016) Science 351: 867-871. PDB As described in file: 5F9R (each incorporated herein by reference), The protein is divided into two fragments within a disordered region. The disordered region is analyzed by X-ray crystallography and NMR. Spectroscopy, electron microscopy (e.g., cryo-EM), and / or in silico protein modeling One or more protein structure determination techniques known in the art, including but not limited to ng. This can be determined by the procedure. In some embodiments, the protein is approximately the same as SpCas9. Any C, T, or A within the region between amino acids A292-G364, F445-K483, or E565-T637, if In S, or any other Cas9, Cas9 variant (e.g., nCas9, dCas9), if In some cases, the napDNAbp is split into two fragments at corresponding locations. The protein is then divided into two fragments at SpCas9 T310, T313, A456, S469, or C574. The protein is split. In some embodiments, the process of splitting a protein into two fragments is performed. This process is called "splitting" the protein.

[0166] "Substantially identical" means that the reference amino acid sequence (for example, the amino acids listed herein) is substantially identical. Any of the sequences) or nucleic acid sequences (for example, any of the nucleic acid sequences described herein) This means a polypeptide or nucleic acid molecule that exhibits at least 50% identity with respect to [the specified substance]. In a form, such an array has at least 60%, 80%, 85%, 90%, 95% or 99% identity with the array used for comparison at the amino acid level or in nucleic acids. .

[0167] Sequence identity is typically determined using sequence analysis software (e.g., Genetics Computer Group , University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705's Sequence Analysis Software Package, BLAST, BESTFIT, GAP, or PILE UP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning a degree of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: gly cine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, a sparagine, glutamine; serine, threonine; lysine, arginine; phenylalanine , tyrosine. In an exemplary approach for determining the degree of identity, the BLAST program can be used, and sequences showing a probability score between e and e -3 and e -100 that are closely related are shown. .

[0168] "Subject" means a mammal, including but not limited to human or non - human mammals such as cows, horses, dogs, sheep or cats. Subjects include livestock, breeding animals raised to supply products such as labor - producing food . and include cows, goats, chickens, horses This includes, but is not limited to, pigs, rabbits, and sheep.

[0169] The term "target site" refers to a sequence within a nucleic acid molecule that has been modified by a nucleic acid base editor. This refers to the sequence that is targeted. In one embodiment, the target site is a deaminase or a fusion containing it. Deamination by synthetic proteins (e.g., cytidine or adenine deaminase) ru.

[0170] RNA-programmable nucleases (e.g., Cas9) use RNA:D to target DNA cleavage sites. Since NA hybridization is used, these proteins are, in principle, guide RN Any sequence specified by A can be targeted. For site-specific cleavage, C Methods using RNA programmable nucleases like as9 (e.g., modifying genomes) (for this purpose) is known in the art (for example, Cong, L. et al., Multiplex gen ome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013 ); Hwang, WY et al., Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programme d genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., G enome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic acids research (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes See *Using CRISPR-Cas Systems*, *Nature Biotechnology* 31, 233-239 (2013). (The full contents of each of these are incorporated herein by reference.)

[0171] The ranges provided herein are understood to be abbreviations for all values ​​within the range. For example, the range from 1 to 50 is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 , 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36 From the group consisting of 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 It is understood to include any number, combination of numbers, or subrange.

[0172] The terms "treat," "treating," and "cure" as used herein refer to the same terms as used in this specification. "Treatment" refers to reducing or improving the symptoms of a disorder and / or related symptoms. This refers to treating a disorder or condition, which means completely eliminating the associated disorder, condition, or symptoms. It will be understood that removal is not necessary (complete removal is excluded). (But not.)

[0173] The terms “uracilglycosylase inhibitor” or “UGI” are used herein in the context of In combination, a protein that can inhibit uracil-DNA glycosylase base excision repair enzyme To refer to. In one aspect, the UGI domain includes wild-type UGI or its modified versions. In one embodiment, the UGI protein provided herein is a UGI fragment, and UGI or contains a protein homologous to the UGI fragment. For example, in some embodiments, the UGI fragment The main component includes a fragment of the amino acid sequence provided below. In some embodiments, The UGI fragment is at least 60%, at least 65%, and less than the exemplary UGI sequence provided herein. At least 70%, at least 75%, at least 80%, at least 85%, at least 90%, Also 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% Includes an amino acid sequence. In one embodiment, UGI provides the following in this specification. Amino acid sequences homologous to the amino acid sequence, or amino acids provided below in this specification. It contains amino acid sequences homologous to the sequence fragment. In some embodiments, UGI or U Proteins containing GI fragments, or UGI or homologs of UGI fragments, are classified as "UGI Varian." It is called "UGI variant". UGI variants have homology to UGI or its fragments. For example, UGI variant The riant has at least 70% identity with wild-type UGI or UGI described herein, and less At least 75% identity, at least 80% identity, at least 85% identity, at least 90% Identity of, at least 95% identity, at least 96% identity, at least 97% identity, At least 98% identity, at least 99% identity, at least 99.5% identity, or less It has at least 99.9% identity. In some embodiments, the UGI variant is UGI It includes a fragment, and that fragment is for the wild-type UGI or the corresponding fragment of the UGI provided below. And, at least 70% identity, at least 80% identity, at least 90% identity, less Both have 95% identity, at least 96% identity, at least 97% identity, and at least 98% identity. Oneness, at least 99% identity, at least 99.5% identity, or at least 99.9% identity It has one nature. In one embodiment, UGI includes the following amino acid sequence: >splP14739IUNGI_BPPB2 Uracil-DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLT SD APE YKPW ALVIQDS NGENKIKML

[0174] Unless otherwise stated or evident from the context, the term "as used herein" The term "or" is understood to be inclusive. Unless otherwise stated, or from the context Unless otherwise stated, the terms “a,” “an,” and “the” as used herein are: It is understood to be singular or plural.

[0175] Unless otherwise stated or made clear from the context, the terms used herein are: The term "approximately" refers to the normal acceptable range in this technical field, for example, within two standard deviations of the mean. It is understood that the approximate values ​​are 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, and 0.5% of the stated value. This can be understood as %, 0.1%, 0.05%, or within 0.01%. Unless otherwise clearly indicated in the context, All figures provided in the specification are qualified with the term "approximately."

[0176] The list of chemical groups in the definition of variable factors in this specification refers to any single one. This includes the definition of its variable factor as a group or a combination of the listed groups. The description of the embodiments relating to the variable factors or aspects is not a single embodiment, This includes embodiments of the same that are combined with any other embodiment or part thereof.

[0177] The compositions or methods provided herein may be used with any other compositions provided herein. It can be combined with one or more calling methods. [Brief explanation of the drawing]

[0178] [Figure 1] Figure 1 is a schematic diagram of an A→G base editor (ABE) fusion protein nucleic acid base editor, including one in which two adenosine deaminase domains, wild-type (wt) TadA and evolved (evo) TadA, are fused to S. pyogenes (Sp) Cas9 nickase (nCas9) with a C-terminal bipartite nuclear localization signal (NLS). This schematic also identifies three regions of the unstructured Cas9 protein in which the fusion protein can be split into N- and C-terminal fragments that can be reconstituted using a split intein system (i.e., by fusing intein-N and intein-C to the N-terminal and C-terminal fragments, respectively).

[0179] [Figure 2]Figure 2 provides three graphs that reproduce the fusion protein described in Figure 1 and quantify the base editing activity of a base editing system containing a nucleic acid base editor fusion protein (i.e., ABE) containing spliced ​​nCas9. The ABE was split at the indicated amino acid positions (e.g., T310, T313, A456, S469, and C574 with respect to the SpCas9 amino acid sequence). The N- and C-terminal fragments of the ABE were fused to intein-N and intein-C, respectively. These fragments and the indicated guide RNA were expressed on separate plasmids in cultured HEK293 cells expressing a protein containing the ABCA4 gene with the 5882G>A mutation, respectively. The base editing activity of the reconstituted ABE against the ABCA4 5882A>G target was compared to the activity of the control ABE. Base editing activity depended on the presence of both the N-terminal and C-terminal fragments of the ABE. No base editing activity was observed when only one of the N-terminal or C-terminal fragments of the ABE was expressed.

[0180] [Figure 3] Figure 3 is a graph confirming the base editing activity mediated by the 21 nt guide described in Figure 2. Note that the experiment in Figure 3 was performed in a different format than the experiment in Figure 2. All reconstituted ABEs showed good base editing activity. This activity depended on the presence of both the N- and C-terminal fragments of nCas9. No activity was observed when only the N- or C-terminal fragment of nCas9 was expressed. The activity of the reconstituted ABEs was compared to that of the control ABE7.09 and ABE7.10 fusion proteins.

[0181] [Figure 4] Figure 4 is a graph quantifying the base editing activity described in Figure 2. In Figure 4, a 20-nucleotide (nt) guide RNA containing hammer ribozyme (HRz) was used. The activity of various reconstituted ABEs was compared with the activity of control ABE7.09 and ABE7.10 fusion proteins.

[0182] [Figure 5]Figures 5A-5D show the determination of the multiplicity of infection (MOI) for AAV2 co-infection of ARPE-19 cells. Figure 5A is a graph showing 1:1 co-infection with AAV2 / CMV-mCherry and AAV2 / CMV-EmGFP in various viral loads (vg / cell). Figure 5B shows fluorescence images detecting EmGFP (left) and mCherry (center), as well as an overlay image (right) showing the co-localization of EmGFP and mCherry. Figure 5C is a graph showing the percentage of cells expressing mCherry in various viral loads (vg / cell) when infected with AAV2 / CMV-mCherry. Figure 5D is a graph showing the percentage of cells expressing EmGFP in various viral loads (vg / cell) when infected with AAV2 / EmGFP.

[0183] [Figure 6] Figure 6 is a series of graphs showing that delivery of a split editor to ARPE-19 cells via dual AAV2 infection results in a high A>G conversion rate in ABCA4 5882A. The multiplicity of infection (MOI) for dual infections at 20,000 vg / cell (top left); 30,000 vg / cell (top right); 40,000 vg / cell (bottom left); and 60,000 vg / cell (bottom right) is shown. [Modes for carrying out the invention]

[0184] As described below, the present invention relates to a composition and method for delivering a base editing system. The present invention provides a method for a nucleic acid base editor (ABE) from A to G to divide the intaine. This is at least partially based on the discovery that it can be "divided" and reconstructed using the division intelligence. N- and C-terminal fragments of ABE fused to the intein-N and intein-C of the in-pair, respectively. The polynucleotides encoding it are delivered to the cell along with a single guide RNA on separate vectors. The target was reached. The encoded ABE fragments were spliced ​​together to target the nucleic acid sequence. We reconstructed functional nucleic acid base editor fusion proteins that are particularly useful for editing. Ta.

[0185] [Intein] Intein in tervening pro tein ) is a self-processing compound found in a wide variety of organisms. It is a domain that performs a process known as protein splicing. Protein splicing is a multi-step biochemical process consisting of both the cleavage and formation of peptide bonds. This is a scientific reaction. The endogenous substrate for protein splicing is found in organisms that contain inteins. Although it is a protein that is produced, intein also has virtually any polypeptide backbone. It can also be used for chemical manipulation.

[0186] In protein splicing, the intein cleaves two peptide bonds. Therefore, it cleaves itself from a precursor polypeptide, thereby separating the adjacent extaine The (external protein) sequence is linked via the formation of a new peptide bond. This rearrangement is reversed. Occurs after translation (may also occur simultaneously with translation). Intein-mediated protein splinting. Icing occurs spontaneously and only requires the folding of the intent domain.

[0187] Approximately 5% of inteins are split inteins, which are N-intenes and C-intenes. It is transcribed and translated as two separate polypeptides, each of which is then converted into a single extension. They are fused. During translation, the intein fragments spontaneously and non-covalently fuse with the standard intein. Assemble into the in structure and perform protein splicing in trans. The pricing mechanism involves a series of acyl transfer reactions, which are the intein-e Cleavage of two peptide bonds at the kistine junction, and N-exstein and C-exstein This process leads to the formation of new peptide bonds between the N-extin and It is initiated by the activation of the peptide bond connecting the intein to the N-terminus. Each intein has cysteine ​​or serine at its N-terminus, which is the C-terminus of N-extine. It attacks the carbonyl carbon of the terminal residue. This N-to-O / S acyl group transfer is a conserved process. Rheonine and histidine (called the TXXH motif), and commonly found asparagus This is facilitated by ginate, leading to the formation of a linear (thio)ester intermediate. The intermediate is the first residue of C-extin, which is cysteine, serine, or threonine. It is trans-(thio)esterified by nucleophilic attack of +1). The resulting branched (thio)ester The intermediate is a unique product due to the cyclization of the highly conserved C-terminal asparagine of intein. This is resolved by conversion. This process involves histidine (a highly conserved HNF motif). (What is found) is promoted by the second-to-last histidine, and aspartic acid is also involved. This succinimide-forming reaction removes intein from the reaction complex and eliminates the peptin. It leaves the extein bound via cytobond. This structure is intein-independent. It then rapidly rearranges to form a stable peptide bond.

[0188] [Adenosine deaminase] In some embodiments, the fusion protein of the present invention is adenosine deaminase drug. Contains n. In some embodiments, adenosine deaminase provided herein This can deaminate adenine. In some embodiments, as specified herein The provided adenosine deaminase removes adenine from deoxyadenosine residues in DNA. It can be aminated. Adenosine deaminase can be used in any suitable organism (e.g., the large intestine). It can be derived from bacteria. In one embodiment, adenine deaminase is used herein One or more mutations corresponding to any of the provided mutations (e.g., the mutation in ecTadA) These are naturally occurring adenosine deaminases, including mutants. Those skilled in the art can, for example, describe the sequence Alignment and homologous residue determination reveals the corresponding residues in any homologous protein. The residues can be identified. Therefore, those skilled in the art can identify any naturally occurring adenosine residues. In a minase (for example, one having homology to ecTadA), the protrusion described herein This corresponds to any of the natural mutations (for example, any of the mutations identified in ecTadA). It can generate mutations. In one aspect, adenosine deaminase is prokaryotic. It is of biological origin. In one aspect, adenosine deaminase is of bacterial origin. In this context, adenosine deaminase is found in Escherichia coli, Staphylococcus aureus, and S almonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter cr Derived from escentus, or Bacillus subtilis. In one embodiment, adenosine deami The enzyme is derived from E. coli.

[0189] In one embodiment, the fusion protein of the present invention is linked to wild-type TadA7.10. It contains and is linked to Cas9 nickase. In certain embodiments, the fusion protein is single Includes one TadA7.10 domain (for example, one provided as a monomer). Other embodiments So, the ABE7.10 editor can form heterodimers with TadA7.10 and Tad Includes A (wt). The related sequences are as follows:

[0190] TadA(wt): SEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHEIMALRQGGLVMQNYRLIDATLY VTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKK AQSSTD

[0191] TadA7.10: SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATLY VTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKK AQSSTD

[0192] In some embodiments, TadA (e.g., one having double-strand substrate activity) or Ta dA7.10 is provided as a homodimer or monomer.

[0193] In one embodiment, adenosine deaminase is provided herein for adenosine deaminase At least 60% of any of the amino acid sequences described in any of the minases, and at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90% , at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or contains an amino acid sequence that is at least 99.5% identical. Provided herein Adenosine deaminase is a mutant of one or more mutants (e.g., the mutants provided herein). It should be understood that this disclosure may include any of the following (different). A deaminase domain having unisexuality, plus any of the mutations described herein. or provides a combination thereof. In some embodiments, adenosine deami The enzyme is compared to either the reference sequence or the adenosine deaminase provided herein. Compare 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , an amino acid sequence having 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations Includes. In some embodiments, adenosine deaminase is used in the art. Compared to any of the amino acid sequences known or described herein, at least 5, at least 10, at least 15, at least 20, at least 25, and at least 30, at least 35, at least 40, at least 45, at least 50, and at least 60, at least 70, at least 80, at least 90, at least 100, fewer 110 each, at least 120, at least 130, at least 140, at least 150, amino acids having at least 160 or at least 170 identical consecutive amino acid residues Contains acid sequences.

[0194] In one aspect, adenosine deaminase is a D108X mutation in the TadA reference sequence. , or including the corresponding mutation in another adenosine deaminase, where X is wild-type adenosine deaminase. This indicates any amino acid other than the corresponding amino acid in nosine deaminase. In this context, adenosine deaminase is found in D108G, D108N, D108V, and D108A in the TadA reference sequence. , or D108Y mutation, or corresponding mutation in another adenosine deaminase This includes, however, further deaminases are similarly aligned and described herein. It is understood that homologous amino acid residues that can be mutated as provided can be identified. It should be done.

[0195] In one aspect, adenosine deaminase is a mutation of the A106X mutation in the TadA reference sequence. , or including the corresponding mutation in another adenosine deaminase, where X is wild-type adenosine deaminase. This indicates any amino acid other than the corresponding amino acid in nosine deaminase. In this context, adenosine deaminase is caused by the A106V mutation in the TadA reference sequence, or another mutation. This includes corresponding mutations in adenosine deaminase.

[0196] In one aspect, adenosine deaminase is a mutation of the E155X mutation in the TadA reference sequence. , or including the corresponding mutation in another adenosine deaminase, the presence of X is in the field This indicates any amino acid other than the corresponding amino acid in bioactive adenosine deaminase. In one embodiment, adenosine deaminase is E155D, E155G in the TadA reference sequence, This includes the E155V mutation, or the corresponding mutation in another adenosine deaminase. .

[0197] In one aspect, adenosine deaminase is a D147X mutation in the TadA reference sequence. , or including the corresponding mutation in another adenosine deaminase, the presence of X is in the field This indicates any amino acid other than the corresponding amino acid in bioactive adenosine deaminase. In one aspect, adenosine deaminase is a D147Y mutation in the TadA reference sequence, This includes the corresponding mutation in another adenosine deaminase.

[0198] Mutations provided herein (for example, based on the ecTadA amino acid sequence of the TadA reference sequence) (These are all other adenosine deaminases such as Staphylococcus aureus TadA (saTadA), Alternatively, introduce it into other adenosine deaminases (e.g., bacterial adenosine deaminase). It should be understood that it is possible to do so. How homologous the mutant residue in ecTadA is This will be obvious to those skilled in the art. Therefore, none of the mutations identified in ecTadA are phase It can also be produced using other adenosine deaminases that have the same amino acid residues. None of the mutations provided in the specification affect ecTadA or another adenosine deaminase. It should also be understood that they can be manufactured individually or in any combination. For example, Adenosine deaminase is D108N, A106V, E155V, and / or in the TadA reference sequence. This may include the D147Y mutation, or a corresponding mutation in another adenosine deaminase. In one aspect, adenosine deaminase undergoes the following mutation in the TadA reference sequence. In different groups (mutation groups are separated by a semicolon), or in different adenosine deaminases Includes corresponding mutations: D108N and A106V; D108N and E155V; D108N and D147Y; A106V and E155V; A106V and D147Y; E155V and D147Y; D108N, A106V, and E55V; D108N, A106V, and D147Y; D108N, E55V, and D147Y; A106V, E55V, and D147Y; as well as D108N, A106V, E55V, and D147Y; however, the corresponding sudden changes provided here Any combination of different substances can be produced in adenosine deaminase (e.g., ecTadA). I want to be understood.

[0199] In some embodiments, adenosine deaminase is H8X in the TadA reference sequence. , T17X, ​​L18X, W23X, L34X, W45X, R51X, A56X, E59X, E85X, M94X, I95X, V102X, F104X , A106X, R107X, D108X, K110X, M118X, N127X, A138X, F149X, M151X, R153X, Q154X, I One or more of the 156X and / or K157X mutations, or other adenosine deaminase mutations. This includes one or more corresponding mutations in, where the presence of X is that of wild-type adenosine deamix. Shows any amino acid other than the corresponding amino acid in the enzyme. Several embodiments In this context, adenosine deaminase is associated with the TadA reference sequence H8Y, T17S, L18E, W23L L34S, W45L, R51H, A56E, or A56S, E59G, E85K, or E85G, M94L, 1951, V102A F104L, A106V, R107C, or R107H, or R107P, D108G, or D108N, or D108A , D108Y, K110I, M118K, N127S, A138V, F149Y, M151V, R153C, Q154L, I156D, and / or one or more K157R mutations, or one in other adenosine deaminases This includes the corresponding mutations mentioned above.

[0200] In one embodiment, adenosine deaminase is H8X, D108X, in the TadA reference sequence. and / or one or more of the N127X mutations, or in another adenosine deaminase It contains one or more corresponding mutations, where X indicates the presence of any amino acid. Several implementations Morphologically, adenosine deaminase has H8Y, D108N, and / Or one or more N127S mutations, or one or more correspondences in another adenosine deaminase. This includes mutations.

[0201] In some embodiments, adenosine deaminase is H8X in the TadA reference sequence. , R26X, M61X, L68X, M70X, A106X, D108X, A109X, N127X, D147X, R152X, Q154X, E155X , one or more K161X, Q163X, and / or T166X mutations, or another adenosine deamine mutation - Contains one or more corresponding mutations in the enzyme, where X is wild-type adenosine deaminator. This indicates the presence of any amino acid other than the corresponding amino acid in the ze. Several implementation forms In this state, adenosine deaminase is H8Y, R26W, M61I, L68Q in the TadA reference sequence. , M70V, A106T, D108N, A109T, N127S, D147Y, R152C, Q154H or Q154R, E155G if or one or more of the following mutations: E155V or E155D, K161Q, Q163H, and / or T166P or include one or more corresponding mutations in other adenosine deaminases.

[0202] In some embodiments, adenosine deaminase is H8X in the TadA reference sequence. One, two, or three selected from the group consisting of D108X, N127X, D147X, R152X, and Q154X. Four, five, or six mutations, or corresponding mutations in another adenosine deaminase. This includes mutations (multiple mutations are possible), where X is the corresponding mutation in wild-type adenosine deaminase. This indicates the presence of any amino acid other than amino acids. In some embodiments, adeno Syndeaminase is H8X, M61X, M70X, D108X, N127X, Q154X, E1 in the TadA reference sequence. One, two, three, four, five, six, or seven are selected from the group consisting of 55X and Q163X, and This involves eight mutations, or corresponding mutations in other adenosine deaminases (multiple mutations). (Possible) includes, where X is a different amino acid from the corresponding amino acid in wild-type adenosine deaminase. It indicates the presence of any of the amino acids. In one aspect, adenosine deaminase is TadA One selected from the group consisting of H8X, D108X, N127X, E155X, and T166X in the reference sequence. , 2, 3, 4, or 5 mutations, or pairings in a different adenosine deaminase Includes corresponding mutations (multiple mutations possible), where X is the corresponding mutation in wild-type adenosine deaminase. This indicates the presence of any amino acid other than the corresponding amino acid. In one embodiment, adenosine Deaminase has mutations in H8X, A106X, D108X, and other adenosine deaminases. It includes 1, 2, 3, 4, 5, or 6 mutations selected from a group consisting of (multiple) such mutations, where X is any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. It shows that. In one embodiment, adenosine deaminase is H8X, R in the TadA reference sequence. One, two, or three selected from the group consisting of 126X, L68X, D108X, N127X, D147X, and E155X. 4, 5, 6, 7, or 8 mutations, or a different adenosine deaminase This includes the corresponding mutation(s) in wild-type adenosine deaminase, where X is the corresponding mutation(s) in wild-type adenosine deaminase. This indicates the presence of any amino acid other than the corresponding amino acid. In one embodiment, adenos Ndeaminase is derived from H8X, D108X, A109X, N127X, and E155X in the TadA reference sequence. One, two, three, four, or five mutations selected from the group, or another adenos Includes the corresponding mutation(s) in an adenosinase, where X is wild-type adenosine. This indicates the presence of an amino acid other than the corresponding amino acid in the minase.

[0203] In some embodiments, adenosine deaminase is H8Y in the TadA reference sequence. One, two, or three selected from the group consisting of D108N, N127S, D147Y, R152C, and Q154H. Four, five, or six mutations, or corresponding mutations in another adenosine deaminase. This includes mutations. In some embodiments, adenosine deaminase is a TadA reference molecule. Select from the group consisting of H8Y, M61I, M70V, D108N, N127S, Q154R, E155G, and Q163H in the column. One, two, three, four, five, six, seven, or eight mutations or other adenosynium are selected. This includes corresponding mutations in adenoaminase. In some embodiments, adeno Syndeaminase is H8Y, D108N, N127S, E155V, and T166P in the TadA reference sequence. One, two, three, four, or five mutations selected from the group, or another adeno Includes the corresponding mutation(s) in syndeaminase. In some embodiments... Adenosine deaminase is found in the TadA reference sequence H8Y, A106T, D108N, N127S, E1 One, two, three, four, five, or six suddens selected from the group consisting of 55D and K161Q This includes mutations, or corresponding mutations in other adenosine deaminases. In this embodiment, adenosine deaminase is H8Y, R126W, L68Q in the TadA reference sequence. , select one, two, three, four, or five from the group consisting of D108N, N127S, D147Y, and E155V. 6, 7, or 8 mutations, or corresponding mutations in other adenosine deaminases This includes mutations. In some embodiments, adenosine deaminase is TadA-related. One selected from the group consisting of H8Y, D108N, A109T, N127S, and E155G in the symmetric array. Two, three, four, or five mutations, or corresponding in a different adenosine deaminase. This includes mutations (multiple mutations are possible).

[0204] In one aspect, adenosine deaminase is used in another adenosine deaminase. It contains one or more corresponding mutations. In one embodiment, adenosine deaminase is Ta In the dA reference sequence, a D108N, D108G, or D108V mutation, or another adenosine dea This includes the corresponding mutation in the minase. In one embodiment, adenosine deaminase This is due to A106V and D108N mutations in the TadA reference sequence, or another adenosine deaminator. This includes the corresponding mutation in the enzyme. In one embodiment, adenosine deaminase is Ta R107C and D108N mutations in the dA reference sequence, or in another adenosine deaminase This includes the corresponding mutation. In one aspect, adenosine deaminase is TadA (see reference). H8Y, D108N, N127S, D147Y, and Q154H mutations in the sequence, or other adenosine mutations. This includes corresponding mutations in deaminase. In one aspect, adenosine deaminase -se is a mutation in the TadA reference sequence with H8Y, R24W, D108N, N127S, D147Y, and E155V mutations. , or including a corresponding mutation in another adenosine deaminase. Adenosine deaminase has a sudden mutation at D108N, D147Y, and E155V in the TadA reference sequence. This includes mutations, or corresponding mutations in another adenosine deaminase. In this context, adenosine deaminase is found in the TadA reference sequence at H8Y, D108N, and S127S. This includes natural mutations or corresponding mutations in other adenosine deaminases. In this context, adenosine deaminase is found in the TadA reference sequence at A106V, D108N, D147Y and This includes the E155V mutation, or the corresponding mutation in another adenosine deaminase. .

[0205] In some embodiments, adenosine deaminase is involved in the S2X sequence of the tadA reference sequence. , one or more of the H8X, I49X, L84X, H123X, N127X, I156X and / or K160X mutations, X contains one or more corresponding mutations in a different adenosine deaminase, and the presence of X indicates that This indicates any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, adenosine deaminase is located in the S2 of the TadA reference sequence. One or more of the following mutations: A, H8Y, I49F, L84F, H123Y, N127S, I156F, and / or K160S , or including one or more corresponding mutations in another adenosine deaminase.

[0206] In one aspect, adenosine deaminase is L84X mutant adenosine deaminase X includes any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. It shows amino acid. In one embodiment, adenosine deaminase is L8 in the TadA reference sequence. This includes 4F mutations, or corresponding mutations in other adenosine deaminases.

[0207] In one aspect, adenosine deaminase is a result of the H123X mutation in the TadA reference sequence. , or including a corresponding mutation in another adenosine deaminase, where X is wild-type A This shows any amino acid other than the corresponding amino acid in denosine deaminase. In this context, adenosine deaminase is caused by the H123Y mutation in the TadA reference sequence, or another mutation. This includes the corresponding mutation in adenosine deaminase.

[0208] In one aspect, adenosine deaminase is a result of the I157X mutation in the TadA reference sequence. , or including a corresponding mutation in another adenosine deaminase, where X is wild-type A This shows any amino acid other than the corresponding amino acid in denosine deaminase. In this context, adenosine deaminase is caused by the I157F mutation in the TadA reference sequence, or another mutation. This includes the corresponding mutation in adenosine deaminase.

[0209] In one embodiment, adenosine deaminase has L84X, A106X in the TadA reference sequence. One, two, or three elements are selected from the group consisting of D108X, H123X, D147X, E155X, and I156X. Four, five, six, or seven mutations, or corresponding mutations in a different adenosine deaminase. This includes mutations (multiple mutations) in wild-type adenosine deaminase, where X is the opposite mutation in wild-type adenosine deaminase. This indicates the presence of any amino acid other than the corresponding amino acid. In one embodiment, adenosine Deaminases are S2X, I49X, A106X, D108X, D147X, and E155X in the tadA reference sequence. One, two, three, four, five, or six mutations selected from the group consisting of, or another This includes the corresponding mutation(s) in adenosine deaminase, where X is the wild type. This indicates the presence of an amino acid other than the corresponding amino acid in adenosine deaminase. In one embodiment, adenosine deaminase has H8X, A106X, in the TadA reference sequence. One, two, three, four, or five mutations selected from the group consisting of D108X, N127X, and K160X. , or including the corresponding mutation(s) in another adenosine deaminase, here X is any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. This indicates the presence of an acid.

[0210] In one embodiment, adenosine deaminase is L84F, A106V in the TadA reference sequence. One, two, or three elements are selected from the group consisting of D108N, H123Y, D147Y, E155V, and I156F. Four, five, six, or seven mutations, or corresponding mutations in a different adenosine deaminase. This includes mutations in the TadA reference sequence. In one embodiment, adenosine deaminase is present in the TadA reference sequence. Then, one or two selected from the group consisting of S2A, I49F, A106V, D108N, D147Y, and E155V. It contains one, three, four, five, or six mutations.

[0211] In some embodiments, adenosine deaminase is H8Y in the TadA reference sequence. One, two, three, four, or more are selected from the group consisting of A106T, D108N, N127S, and K160S. Or five mutations, or corresponding mutations in another adenosine deaminase. This includes mutations.

[0212] In some embodiments, adenosine deaminase is E 25 in the TadA reference sequence. One or more of the following mutations: X, R26X, R107X, A142X, and / or A143X, or The presence of X includes one or more corresponding mutations in another adenosine deaminase, and the wild type This indicates any amino acid other than the corresponding amino acid in type adenosine deaminase. In one embodiment, adenosine deaminase is E25M, E25D, E25A in the TadA reference sequence. , E25R, E25V, E25S, E25Y, R26G, R26N, R26Q, R26C, R26L, R26K, R107P, R07K, R107A , R107N, R107W, R107H, R107S, A142N, A142D, A142G, A143D, A143G, A143E, A143L, A One or more of the 143W, A143M, A143S, A143Q, and / or A143R mutations, or other A Includes one or more corresponding mutations in denosine deaminase. In some embodiments, In this context, adenosine deaminase is a mutation described herein that corresponds to the TadA reference sequence. One or more of the following, or the corresponding mutations in another adenosine deaminase Includes one or more of the following.

[0213] In one aspect, adenosine deaminase is a mutation in the TadA reference sequence, E25X mutation. Or include the corresponding mutation in another adenosine deaminase, where X is wild-type adenosine deaminase. This indicates any amino acid other than the corresponding amino acid in nosine deaminase. In this context, adenosine deaminase is found in the TadA reference sequence E25M, E25D, E25A, E25R, E2 5V, E25S, or E25Y mutations, or corresponding mutations in other adenosine deaminases Includes natural variations.

[0214] In one aspect, adenosine deaminase is a result of the R26X mutation in the TadA reference sequence. Or include the corresponding mutation in another adenosine deaminase, where X is wild-type adenosine deaminase. This shows any amino acids other than the corresponding amino acid in nosine deaminase. In this embodiment, adenosine deaminase is R26G, R26N, R26Q in the TadA reference sequence. , R26C, R26L, or R26K mutations, or corresponding mutations in other adenosine deaminases This includes mutations.

[0215] In one aspect, adenosine deaminase is a result of the R107X mutation in the TadA reference sequence. , or including a corresponding mutation in another adenosine deaminase, where X is wild-type A This shows any amino acid other than the corresponding amino acid in denosine deaminase. In this context, adenosine deaminase is R107P, R07K, R107A, R107 in the TadA reference sequence. N, R107W, R107H, or R107S mutations, or paired in another adenosine deaminase Includes corresponding mutations.

[0216] In one aspect, adenosine deaminase is a mutation in the TadA reference sequence, specifically the A142X mutation. , or including a corresponding mutation in another adenosine deaminase, where X is wild-type A This shows any amino acid other than the corresponding amino acid in denosine deaminase. In this context, adenosine deaminase has a mutation at A142N, A142D, and A142G in the TadA reference sequence. This includes mutations, or corresponding mutations in other adenosine deaminases.

[0217] In one aspect, adenosine deaminase is a mutation of the A143X mutation in the TadA reference sequence. , or including a corresponding mutation in another adenosine deaminase, where X is wild-type A This shows any amino acid other than the corresponding amino acid in denosine deaminase. In this context, adenosine deaminase is found in A143D, A143G, A143E, and A14 in the TadA reference sequence. Mutations of 3L, A143W, A143M, A143S, A143Q and / or A143R, or other adenosine This includes corresponding mutations in deaminase.

[0218] In some embodiments, adenosine deaminase is H36X in the TadA reference sequence. , N37X, P48X, I49X, R51X, M70X, N72X, D77X, E134X, S146X, Q154X, K157X, and / or one or more K161X mutations, or one or more in another adenosine deaminase This includes the corresponding mutation, where the presence of X is the opposite in wild-type adenosine deaminase. This represents any amino acid other than the corresponding amino acid. In some embodiments, adeno Syndeaminase is H36L, N37T, N37S, P48T, P48L, I49V, R51H in the TadA reference sequence. R51L, M70L, N72S, D77G, E134G, S146R, S146C, Q154H, K157N, and / or K161T One or more mutations in the mutation, or one or more corresponding mutations in another adenosine deaminase. Includes.

[0219] In some embodiments, adenosine deaminase is H36X in the TadA reference sequence. This includes mutations, or corresponding mutations in other adenosine deaminases, where X is any amino acid other than the corresponding amino acid in wild-type adenosine deaminase As shown, in some embodiments, adenosine deaminase is present in the TadA reference sequence. This includes the H36L mutation, or a corresponding mutation in another adenosine deaminase.

[0220] In one aspect, adenosine deaminase is a mutation of the N37X mutation in the TadA reference sequence. Or include the corresponding mutation in another adenosine deaminase, where X is wild-type adenosine. This indicates any amino acid other than the corresponding amino acid in syndeaminase. Adenosine deaminase is involved in the N37T or N37S mutations in the TadA reference sequence, This includes a corresponding mutation in another adenosine deaminase.

[0221] In one aspect, adenosine deaminase is a P48X mutation in the TadA reference sequence. Or include a corresponding mutation in another adenosine deaminase, where X is wild. This indicates any amino acid other than the corresponding amino acid in type adenosine deaminase. In this embodiment, adenosine deaminase is a P48T or P48L mutation in the TadA reference sequence. This includes corresponding mutations in different or different adenosine deaminases.

[0222] In one aspect, adenosine deaminase is a mutation of the R51X mutation in the TadA reference sequence. Or include the corresponding mutation in another adenosine deaminase, where X is wild-type adenosine. This indicates any amino acid other than the corresponding amino acid in syndeaminase. Adenosine deaminase is involved in R51H or R51L mutations in the TadA reference sequence, This includes a corresponding mutation in another adenosine deaminase.

[0223] In one aspect, adenosine deaminase is a result of the S146X mutation in the TadA reference sequence. , or including the corresponding mutation in another adenosine deaminase, where X is wild-type adenosine deaminase. This indicates any amino acid other than the corresponding amino acid in nosine deaminase. In this context, adenosine deaminase is caused by S146R or S146C mutations in the TadA reference sequence. or include corresponding mutations in other adenosine deaminases.

[0224] In one aspect, adenosine deaminase is a K157X mutation in the TadA reference sequence. , or including the corresponding mutation in another adenosine deaminase, where X is wild-type adenosine deaminase. This indicates any amino acid other than the corresponding amino acid in nosine deaminase. In this context, adenosine deaminase is caused by a K157N mutation in the TadA reference sequence, or another mutation. This includes corresponding mutations in adenosine deaminase.

[0225] In one aspect, adenosine deaminase is a P48X mutation in the TadA reference sequence. Or include a corresponding mutation in another adenosine deaminase, where X is wild. This indicates any amino acid other than the corresponding amino acid in type adenosine deaminase. In this embodiment, adenosine deaminase is P48S, P48T, or P4 in the TadA reference sequence. This includes the 8A mutation, or a corresponding mutation in another adenosine deaminase.

[0226] In one aspect, adenosine deaminase is a mutation in the TadA reference sequence, specifically the A142X mutation. , or including the corresponding mutation in another adenosine deaminase, where X is wild-type adenosine deaminase. This indicates any amino acid other than the corresponding amino acid in nosine deaminase. In this context, adenosine deaminase is caused by the A142N mutation in the TadA reference sequence, or another mutation. This includes corresponding mutations in adenosine deaminase.

[0227] In one aspect, adenosine deaminase is a W23X mutation in the TadA reference sequence. Or include a corresponding mutation in another adenosine deaminase, where X is wild. This indicates any amino acid other than the corresponding amino acid in type adenosine deaminase. In this embodiment, adenosine deaminase is a W23R or W23L mutation in the TadA reference sequence. This includes corresponding mutations in different or different adenosine deaminases.

[0228] In one aspect, adenosine deaminase is a result of the R152X mutation in the TadA reference sequence. , or a corresponding mutation in another adenosine deaminase, where X is a field This indicates any amino acid other than the corresponding amino acid in bioactive adenosine deaminase. In one embodiment, adenosine deaminase is a receptor for R152P or R52H in the TadA reference sequence. This includes natural mutations or corresponding mutations in other adenosine deaminases.

[0229] In one embodiment, adenosine deaminase is mutant H36L, R51L, L84F, A106V, D10 This may include 8N, H123Y, S146C, D147Y, E155V, I156F, and K157N. Some implementations Morphologically, adenosine deaminase is a combination of the following mutations associated with the tadA reference sequence. This includes, where each mutation in a combination is separated by "_", and each combination of mutations is It is in parentheses: (A106V_D108N), (R107C_D108N), (H8Y_D108N_S 127S_D 147Y_Q154H), (H8Y_R24W_D108N_N127S_D147Y_E155V), (D108N_D147 Y_E155V), (H8Y_D108N_S 127S), (H8Y_D108N_N127S_D147Y_Q154H), (A106V D108N D147Y E155V) (D108Q D147Y E155V) (D108M_D147Y_E155V), (D108L_D147Y_E155V), (D108K_D147 Y_E155V), (D108I_D147Y_E155V), (D108F_D147Y_E155V), (A106V_D108N_D147Y), (A106V_D108M_D147Y_E155V), (E59A_A106V_D108N_D147Y_E155V), (E59A cat dead_A106V_D108N_D147Y_E155V), (L84F_A106V_D108N_H123Y_D147Y_E155V_I156Y), (L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (D103A_D014N), (G22P_D 103 A_D 104N), (G22P_D 103 A_D 104N_S 138 A) , (D 103 A_D 104N_S 138A), (R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F), (E25G_R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I15 6F), (E25D_R26G_L84F_A106V_R107K_D108N_H123Y_A142N_A143G_D147Y_E155V_I15 6F), (R 26Q_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F), (E25M_R26G_L84F_A106V_R107P_D108N_H123Y_A142N_A143D_D147Y_E155V_I15 6F), (R26C_L 84F_A106V_R107H_D108N_H123Y_A142N_D147Y_E155V_I156F), (L84F_A106V_D108N_H123Y_A1 42N_A143L_D147Y_E155V_I156F), (R26G_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F), (E25A_R26G_L84F_A106V_R107N_D108N_H123Y_A142N_A143E_D147Y_E155V_I15 6F), (R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F), (A106V_D108N_A142N_D147Y_E155V), (R26G_A106V_D108N_A142N_D147Y_E155V), (E25D_R26G_A106V_R107K_D108N_A142N_A143G_D147Y_E155V), (R26G_A106V_D108N_R107H_A142N_A143D_D147Y_E155V), (E25D_R26G_A106V_D108N_A142N_D147Y_E155V), (A106V_R107K_D108N_A142N_D147Y_E155V), (A106V_D108N_A142N_A143G_D147Y_E155V), (A106V_D108N_A142N_A143L_D147Y_E155V), (H36L_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (N37T_P48T_M70L_L84F_A106V_D108N_H123Y_D147Y_I49V_E155V_I156F), (N37S_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K161T), (H36L_L84F_A106V_D108N_H123Y_D147Y_Q154H_E155V_I156F), (N72S_L84F_A106V_D108N_H123Y_S 146R_D147Y_E155V_I156F), (H36L_P48L_L84F_A106V_D108N_H123Y_E134G_D147Y_E155V_I156F), 57N), (H36L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F), (L84F_A106V_D108N_H123Y_S 146R_D147Y_E155V_I156F_K161T), (N37S_R51H_D77G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (R51L_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K157N), (D24G_Q71R_L84F_H96L_A106V_D108N_H123Y_D147Y_E155V_I156F_K160E), (H36L_G67V_L84F_A106V_D108N_H123Y_S 146T_D147Y_E155V_I156F), (Q71L_L84F_A106V_D108N_H123Y_L137M_A143E_D147Y_E155V_I156F), (E25G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_Q159L), (L84F_A91T_F104I_A106V_D108N_H123Y_D147Y_E155V_I156F), (N72D_L84F_A106V_D108N_H123Y_G125A_D147Y_E155V_I156F), (P48S_L84F_S97C_A106V_D108N_H123Y_D147Y_E155V_I156F), (W23G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (D24G_P48L_Q71R_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_Q159L), (L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F), (H36L_R51L_L84F_A106V_D108N_H123Y_A142N_S 146C_D147Y_E155V_I156F _K157N), (N37S_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F_K161T), (L84F_A106V_D108N_D147Y_E155V_I156F), (R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F_K157N_K161T), (L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F_K161T), (L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F_K157N_K160E_K161T), (L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F_K157N_K160E), (R74Q L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (R74A_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (R74Q_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (L84F_R98Q_A106V_D108N_H123Y_D147Y_E155V_I156F), (L84F_A106V_D108N_H123Y_R129Q_D147Y_E155V_I156F), (P48S_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F), (P48S_A142N), (P48T_I49V_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F_L157N), (P48T_I49V_A142N), (H36L_P48S_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (H36L_P48S_R51L_L84F_A106V_D108N_H123Y_S 146C_A142N_D147Y_E155V_I156F (H36L_P48T _I49V_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (H36L_P48T_I49V_R51L_L84F_A106V_D108N_H123Y_A142N_S 146C_D147Y_E155V_ I156F _K15 7N), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142N_S 146C_D147Y_E155V_I156F _K157N), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_A142N_D147Y_E155V_I156F _K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (W23R_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_E155V_I156F _K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146R_D147Y_E155V_I156F _K161T), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_R152H_E155V_I156F _K157N), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_R152P_E155V_I156F _K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_R152P_E155V _I156F _K15 7N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142A_S 146C_D147Y_E155 V_I156F _K15 7N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142A_S 146C_D147Y_R152P _E155V_I156 F _K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146R_D147Y_E155V_I156F _K161T), (W23R_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S 146C_D147Y_R152P_E155V _I156F _K15 7N), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142N_S 146C_D147Y_R152P_E155 V_I156F _K1 57N).

[0230] [Cytidine deaminase] In one embodiment, the fusion protein of the present invention comprises cytidine deaminase. In one embodiment, the cytidine deaminase provided herein is cytosine or 5-methylcytosine can be deaminated to uracil or thymine. In one embodiment, the cytosine deaminase provided herein is used to remove cytosine from DNA. It can deaminate cytidine deaminase. Cytidine deaminase can be derived from any suitable organism. It is possible. In some embodiments, cytidine deaminase is naturally occurring. One or more cytidine deaminases corresponding to any of the mutations provided herein This includes mutations. Those skilled in the art will know, for example, sequence alignment and phase By determining identical residues, it is possible to identify corresponding residues in any homologous protein. Therefore, a person skilled in the art can determine the mutation corresponding to any of the mutations described herein. This can be produced in any naturally occurring cytidine deaminase. In one embodiment, cytidine deaminase is of prokaryotic origin. In another embodiment, cytidine Zindeaminase is of bacterial origin. In one embodiment, cytidinedeaminase is of bacterial origin. It originates from a dairy animal (for example, a human).

[0231] In some embodiments, cytidine deaminase is the cytidine described herein. At least 60%, at least 65%, and less than any of the endoaminase amino acid sequences. 70% each, at least 75%, at least 80%, at least 85%, at least 90%, at least 9 5%, at least 96%, at least 97%, at least 98%, at least 99%, or at least Contains an amino acid sequence with 99.5% identity. The cytidine deamisin provided herein The enzyme contains one or more mutations (e.g., any of the mutations provided herein). Please understand that this disclosure may be perceived as any deamination having a certain percentage identity. - The ze domain is modified to include any of the mutations or combinations described herein. The present invention provides a cytidine deaminase that is a reference sequence Or compared to any of the cytidine deaminases provided herein, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, Contains amino acid sequences with 47, 48, 49, 50 or more mutations. Several implementations In terms of morphology, cytidine deaminase is known in the art or is described herein. Compared to any of the amino acid sequences described in the book, at least 5, at least 10, less At least 15, at least 20, at least 25, at least 30, at least 35, at least 40, At least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, or at least 170 identical consecutive amino acid residues It contains an amino acid sequence having the following characteristics.

[0232] The fusion protein of the present invention includes a nucleic acid editing domain. In one embodiment, the nucleic acid editing domain The main function is to catalyze the base change from C to U. In one embodiment, nucleic acid editing The main component is the deaminase domain. In one aspect, deaminase is cytidine It is a deaminase or adenosine deaminase. In one embodiment, deaminase is It is an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In one embodiment, the deaminase is APOBECl deaminase. In one embodiment, The deaminase is APOBEC 2 deaminase. In one aspect, the deaminase is APOBEC It is a 3-deaminase. In one aspect, the deaminase is APOBEC3A deaminase. In one embodiment, the deaminase is APOBEC3B deaminase. Therefore, the deaminase is APOBEC3C deaminase. In one aspect, the deaminase is , APOBEC3D deaminase. In one aspect, deaminase is APOBEC3E deaminase. It is an enzyme. In one aspect, the deaminase is APOBEC3F deaminase. In this case, the deaminase is APOBEC3G deaminase. In one embodiment, deaminase The enzyme is APOBEC3H deaminase. In one aspect, the deaminase is APOBEC4 It is a deaminase. In one aspect, the deaminase is an activation-inducing deaminase (AI). D) In ​​one aspect, deaminase is a vertebrate deaminase. In this context, deaminase is an invertebrate deaminase. In one aspect, deaminase Minazes are found in humans, chimpanzees, gorillas, monkeys, cattle, dogs, rats, or mice. It is an aminase. In one aspect, deaminase is a human deaminase. In one embodiment, the deaminase is a rat deaminase, for example, rAPOBEC1. In this context, the deaminase is Petromyzon marinus cytidine deaminase 1 (pmCDAl). Yes. In one embodiment, the deaminase is human APOBEC3G. In one embodiment, The aminase is a fragment of human APOBEC3G. In one embodiment, the deaminase is D316R This is a human APOBEC3G variant containing the D317R mutation. In one embodiment, deaminase This is a fragment of human APOBEC3G and contains mutations corresponding to the D316R and D317R mutations. In this embodiment, the nucleic acid editing domain is a deaminase of any deaminase described herein. At least 80%, at least 85%, at least 90%, and at least 92% relative to the enzyme domain , at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, Or at least 99.5% identical.

[0233] In certain embodiments, the fusion protein provided herein is a fusion protein with a base It includes one or more features that improve editing activity. For example, the fusion protein provided herein The substance may contain a Cas9 domain with reduced nuclease activity. In some embodiments, The fusion protein provided herein has a Cas9 domain that does not possess nuclease activity. Cas9 niccasse (nCas9), also known as dCas9, cleaves one strand of a double-stranded DNA molecule. It may contain a Cas9 Domain.

[0234] [Other nucleic acid base editors] The present invention relates to cytidine deaminase or adenosine in the fusion protein according to the present invention. The domain of the ndeaminase is virtually any nucleic acid base editor known in the art. It provides a nucleic acid base editor fusion protein that has been replaced with -.

[0235] In some embodiments, nucleic acid-programmable DNA-binding proteins (napDNAbp) , a Cas9 domain. Non-exclusive exemplary Cas9 domains are provided herein. The main components are the nuclease-active Cas9 domain, the nuclease-inactive Cas9 domain, or Ca It may be a Cas9 nickas. In one embodiment, the Cas9 domain is a nuclease activity domain. For example, the Cas9 domain is located on both strands of a double-stranded nucleic acid (for example, both strands of a double-stranded DNA molecule). It may be a Cas9 domain that cleaves the (one-sided) chain. In some embodiments, the Cas9 domain The amino acid sequence comprises one of the amino acid sequences described herein. In some embodiments, Therefore, the Cas9 domain has less than one of the amino acid sequences described herein. 60% each, at least 65%, at least 70%, at least 75%, at least 80%, at least 8 5%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, Contains amino acid sequences that are at least 99% or at least 99.5% identical. In embodiments, the Cas9 domain is one of the amino acid sequences described herein. Compared to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 , 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40 , amino acids having mutations 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more It contains an acid sequence. In some embodiments, the Cas9 domain is as described herein. Compared to any one of the mino acid sequences, at least 10, at least 15, at least 20, less At least 30, at least 40, at least 50, at least 60, at least 70, at least 80 , at least 90, at least 100, at least 150, at least 200, at least 250, less At least 300, at least 350, at least 400, at least 500, at least 600, and at least 700, at least 800, at least 900, at least 1000, at least 1100, or fewer It contains an amino acid sequence having at least 1200 identical consecutive amino acid residues.

[0236] In one embodiment, the Cas9 domain is a nuclease-inactive Cas9 domain (dCas9). For example, the dCas9 domain does not cleave either strand of a double-stranded nucleic acid molecule, but rather... It can bind to nucleic acid molecules (e.g., via gRNA molecules). In some embodiments, The rease-inactive dCas9 domain is a D10X mutation of the amino acid sequence described herein. and the H840X mutation, or in any of the amino acid sequences provided herein. This includes the corresponding mutation, where X is any amino acid change. In some embodiments, The nuclease-inactive dCas9 domain is a D10A mutation in the amino acid sequence described herein. and the H840A mutation, or the corresponding amino acid sequence in any of the amino acid sequences described herein This includes mutations. For example, the nuclease-inactive Cas9 domain is a cloning vector. The following amino acid sequence is included in the pPlatTET-gRNA2 (access number BAV54124): MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD (Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-specifi See "c control of gene expression." Cell. 2013; 152(5):1173-83. See the full content of that article. (This is incorporated herein by law.)

[0237] Further suitable nuclease-inactive dCas9 domains are relevant to this disclosure and the art. Such further examples would be obvious to those skilled in the art and are within the scope of this disclosure. A suitable nuclease-inactivating Cas9 domain is, but is not limited to, D10A / H84. This includes the 0A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains. For example, Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotech See nology. 2013; 31(9): 833-838 (the full contents of which are incorporated herein by reference). (Included). In some embodiments, the dCas9 domain is provided herein. For any of the s9 domains, at least 60%, at least 65%, at least 70%, and less 75% each, at least 80%, at least 85%, at least 90%, at least 95%, at least 9 6%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity It contains an amino acid sequence having the following: In some embodiments, the Cas9 domain is as described herein Compared to any of the listed amino acid sequences, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 , 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32 , 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or so It contains amino acid sequences having more than this number of mutations. In some embodiments, Cas9 dome The amino acid sequence is at least 10, and less than any of the amino acid sequences described herein. At least 15, at least 20, at least 30, at least 40, at least 50, at least 60, At least 70, at least 80, at least 90, at least 100, at least 150, and at least 200, at least 250, at least 300, at least 350, at least 400, at least 500 , at least 600, at least 700, at least 800, at least 900, at least 1000, small Amino acid mixture having at least 1100, or at least 1200, identical consecutive amino acid residues Includes columns.

[0238] In one embodiment, the Cas9 domain is Cas9 nickaase. Cas9 nickaase has two The Cas9 protein can cleave only one strand of a double-stranded nucleic acid molecule (for example, a double-stranded DNA molecule). It can be of quality. In some embodiments, Cas9 nickase targets double-stranded nucleic acid molecules. The strand is cut, and this is how Cas9 nickase breaks the gRNA (e.g., sgRNA) that is bound to Cas9 and salt This means cutting the chains that form a base pair (they are complementary). In one embodiment, Cas9 nicasse contains the D10A mutation, which has histidine at position 840. In its application mode, Cas9 nickase cleaves the non-target, non-base-edited strands of double-stranded nucleic acid molecules. This is because Cas9 nickase forms a base pair with the gRNA (e.g., sgRNA) bound to Cas9. This means cutting the chain that is not connected. In one embodiment, Cas9 nickase is H840A A natural mutation is included, resulting in an aspartic acid residue at position 10, or a corresponding mutation. In that embodiment, Cas9 nickas is any of the Cas9 nickas provided herein. either at least 60%, at least 65%, at least 70%, at least 75%, or at least 80% , at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, small It contains an amino acid sequence that is at least 98%, at least 99%, or at least 99.5% identical. Further suitable Cas9 nickase can be found based on this disclosure and knowledge in the art. This is clear to the business operator and falls within the scope of this disclosure.

[0239] [Cas9 domain with reduced PAM exclusivity] In one particular embodiment, the present invention includes a Cas9 domain divided into two fragments. It is characterized by a nucleic acid base editor, where each of the two fragments has a terminal intein, i.e., N The terminal fragment is fused with one member of the intein system at its C-terminus, and the C-terminal fragment is It has an intein system member at its N-terminus.

[0240] Typically, Cas9 proteins, such as Cas9 derived from S. pyogenes (spCas9), are specific nucleic acids. A standard NGG PAM sequence is required to bind to the region, where the "N" in "NGG" is adeno Syn (A), thymidine (T), or cytosine (C), and G is guanosine. This is, The ability to edit desired bases within the genome may be limited. In some embodiments, the present The base-editing fusion proteins provided in the document are located at precise locations, for example, upstream of PAM, target sites. It may be necessary to place it in a region containing a base. For example, Komor, AC, et al., “Progr ammable editing of a target base in genomic DNA without double-stranded DNA clear See “vage” Nature 533, 420-424 (2016) (The entire content of these articles is referred to herein by reference). (to be incorporated). Accordingly, in some embodiments, the fusion tank provided herein One of the proteins binds to a nucleotide sequence that does not contain a standard (e.g., NGG) PAM sequence. It may contain a Cas9 domain that can bind to a non-standard PAM sequence. This is described in the technical field and will be obvious to those skilled in the art. For example, in non-standard PAM sequences The Cas9 domain to bind is described in Kleinstiver, BP, et al., “Engineered CRISPR-Cas9 nuc "Leases with altered PAM specificities" Nature 523, 481-485 (2015); and Kleins Tiber, BP, et al., “Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition” Nature Biotechnology 33, 1293-1298 (2 As described in 015), the full contents of each are incorporated here by reference. Table 1 below shows Several PAM variants are listed.

[0241] Table 1. Cas9 protein and corresponding PAM sequence [Table 1]

[0242] In some embodiments, PAM is NGC. In some embodiments, NGC PA M is recognized by the Cas9 variant. In some embodiments, the NGC PAM variant The models are D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E and T1337R (combined). It contains one or more amino acid substitutions selected from (and called "MQKFRAER").

[0243] In one aspect, the Cas9 domain is a Cas9 domain derived from Staphylococcus aureus (SaCa s9) In one embodiment, the SaCas9 domain is nuclease-active SaCas9, nuclease These are SaCas9 nickases (SaCas9d) or SaCas9 nickases (SaCas9n). In embodiments, SaCas9 is an N579A mutation, or an amino acid combination provided herein. Includes the corresponding mutation in any of the columns.

[0244] In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain The ion can bind to nucleic acid sequences having non-standard PAMs, and in some embodiments, The SaCas9 domain, SaCas9d domain, or SaCas9n domain contains the NNGRRT PAM sequence. It can bind to nucleic acid sequences. In some embodiments, the SaCas9 domain , one or more of the E781X, N967X, and R1014X mutations, or amino acids provided herein. The sequence includes a corresponding mutation in any of the sequences, where X is any amino acid. In some embodiments, the SaCas9 domain has E781K, N967K, and R1014H mutations. One or more of them, or one or more of any of the amino acid sequences provided herein. Includes corresponding mutations. In some embodiments, the SaCas9 domain is E781K, N96 In the 7K, or R1014H mutation, or any of the amino acid sequences provided herein Includes the corresponding mutation.

[0245] Exemplary SaCas9 sequence KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHS ELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKD GEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYEMLMGHCTYFP EELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKP EFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAIN LILDELWHTNDNQIAIFNRKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIQUINAIIKKYGLPNDIIIELARE KNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPR SVSFDNSFNNKVLVKQEE N SKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTTKKEYLLEERDINRFSVQKDF INRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKKERNKGYKHHAEDALIIANADFIFKEWKKL DKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLI VNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIK YYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENNYEVNSKCYEEAKKLKKISNQAE FIASFYNDDLIKINGLYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKYSTDILGNLYEV KSKKHPQIIKKG The above residue N579, which is underlined and in bold, is mutated (for example to A579) into SaCas9 nicker It can produce ze.

[0246] Exemplary SaCas9n sequence KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHS ELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKD GEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYEMLMGHCTYFP EELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKP EFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAIN LILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELARE KNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPR SVSFDNSFNNKVLVKQEE A SKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDF INRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKL DKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLI VNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIK YYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAE FIASFYNNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKYSTDILGNLYEV KSKKHPQIIKKG The above residue A579 can be mutated from N579 to produce SaCas9 nickase, and is underlined and bolded. It is shown in writing.

[0247] Exemplary SaKKH Cas9 KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHS ELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKD GEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYEMLMGHCTYFP EELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKP EFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAIN LILDELWHTNDNQIAIFNRLKLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELARE KNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPR SVSFDNSFNNKVLVKQEE A SKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDF INRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAEDALIIANADFIFKEWKKL DKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNR K LINDTLYSTRKDDKGNTLI VNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIK YYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQAE FIASFY K NDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPP H IIKTIASKTQSIKKYSTDILGNLYEV KSKKHPQIIKKG. The above residue A579 can be mutated from N579 to produce SaCas9 nickase, and is underlined and bolded. The above residues K781, K967, and H1014 are derived from E781, N967, and R1014. This mutation can produce SaKKH Cas9, which is indicated in underline and italics.

[0248] In one aspect, the Cas9 domain is a Cas9 domain derived from Streptococcus pyogenes. (SpCas9). In one embodiment, the SpCas9 domain is nuclease activity SpCas9, nuclease These are either ase-inactive SpCas9 (SpCas9d) or SpCas9 nickase (SpCas9n). In this embodiment, SpCas9 is a D9X mutation, or an amino acid combination provided herein. The column contains the corresponding mutation, where X is any amino acid other than D. In some embodiments, SpCas9 is a D9A mutation, or provided herein Includes a corresponding mutation in any of the amino acid sequences. In one embodiment, SpCas9 The main, SpCas9d, or SpCas9n domain binds to nucleic acid sequences with non-standard PAMs. This is possible. In one embodiment, a SpCas9 domain, a SpCas9d domain, or a SpCas9n domain. The domain can bind to nucleic acid sequences containing NGG, NGA, or NGCG PAM sequences. In some embodiments, the SpCas9 domain is affected by D1134X, R1334X, and T1336X mutations. One or more of the amino acid sequences provided herein, or the corresponding aggravating This includes mutations, where X is any amino acid. In some embodiments, S The pCas9 domain is one or more of the D1134E, R1334Q, and T1336R mutations, or as specified herein. Includes the corresponding mutation in any of the provided amino acid sequences. Several implementations In this state, the SpCas9 domain has D1134E, R1334Q, and T1336R mutations, or the details described above. Includes a corresponding mutation in any of the amino acid sequences provided in the book. In the application morphology, the SpCas9 domain is one of the D1134X, R1334X, and T1336X mutations. One or more, or corresponding vertices in any of the amino acid sequences provided herein. This includes natural mutations, where X is any amino acid. In some embodiments, SpCas9 The domain is one or more of the D1134V, R1334Q, and T1336R mutations, or provided herein. Includes a corresponding mutation in any of the amino acid sequences. In some embodiments, Furthermore, the SpCas9 domain is a mutation of D1134V, R1334Q, and T1336R, or as specified herein. Includes a corresponding mutation in any of the provided amino acid sequences. Several embodiments In this case, the SpCas9 domain has one or more D1134X, G1217X, R1334X, and T1336X mutations. or comprising the corresponding mutation in any of the amino acid sequences provided herein. Here, X is any amino acid. In some embodiments, the SpCas9 domain is D One or more of the 1134V, G1217R, R1334Q, and T1336R mutations, or provided herein. This includes a corresponding mutation in any of the amino acid sequences. In some embodiments, The SpCas9 domain is a mutation of D1134V, G1217R, R1334Q, and T1336R, or as specified herein. This includes a corresponding mutation in any of the amino acid sequences provided in [the document].

[0249] In some embodiments, any of the fusion proteins provided herein may contain Cas9 The domain is at least 60% of the Cas9 polypeptide described herein, and at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90% , at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or contains an amino acid sequence that is at least 99.5% identical. In some embodiments, The Cas9 domain of any of the fusion proteins provided herein is as described herein. The present invention includes an amino acid sequence of any Cas9 polypeptide. In some embodiments, Any Cas9 domain of the fusion protein provided in this document is as described herein. It consists of the amino acid sequence of the Cas9 polypeptide.

[0250] Exemplary SpCas9 DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0251] Exemplary SpCas9n DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0252] Exemplary SpEQR Cas9 DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGF E SPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK Q Y R STKEVLDATLIHQSITGLYETRID LSQLGGD The above residues E1134, Q1334, and R1336 are mutated from D1134, R1334, and T1336. SpEQR can generate Cas9, as indicated by the underlined and bolded text.

[0253] Exemplary SpVQR Cas9 DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGF V SPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK Q Y R STKEVLDATLIHQSITGLYETRID LSQLGGD The above residues V1134, Q1334, and R1336 are mutated from D1134, R1334, and T1336. SpVQR Cas9 can be generated, and is shown in underlined and bold.

[0254] Exemplary SpVRER Cas9 DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGF VSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASA R ELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK E Y R STKEVLDATLIHQSITGLYETRID LSQLGGD The above residues V1134, R1217, Q1334, and R1336 are D1134, G1217, R1334, and T1336 This can be mutated to produce SpVRER Cas9, as indicated by the underlined and bolded text.

[0255] In certain embodiments, the fusion protein of the present invention binds to a dCas9 domain that binds to a canonical PAM sequence. nCas9 domains that bind to n and non-canonical PAM sequences (e.g., non-canonical PAMs identified in Table 1). This includes. In another embodiment, the fusion protein of the present invention binds to a canonical PAM sequence. dCas9 domains that bind to main and non-canonical PAM sequences (e.g., non-canonical PAMs identified in Table 1) Includes "in".

[0256] [High-fidelity Cas9 domain] Some aspects of this disclosure provide high-fidelity Cas9 domains. In some embodiments... In this context, the high-fidelity Cas9 domain is compared to the corresponding wild-type Cas9 domain. Humans with one or more mutations that reduce the electrostatic interaction between the DNA and the sugar-phosphate backbone. This is the Cas9 domain. I don't want to be tied to any particular theory, but the sugar-phosphate backbone of DNA and The high-fidelity Cas9 domain, which reduces electrostatic interaction, exhibits fewer off-target effects. It may have. In one aspect, a Cas9 domain (e.g., a wild-type Cas9 domain) is a Cas9 domain It includes one or more mutations that reduce the binding between the DNA and the sugar-phosphate backbone. In this context, the Cas9 domain reduces the binding between the Cas9 domain and the sugar-phosphate backbone of DNA. Each 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, At least 15%, at least 20%, at least 25%, at least 30%, at least 35%, less At least 40%, at least 45%, at least 50%, at least 55%, at least 60%, and It also includes one or more mutations that reduce the risk by 65%, or at least 70%.

[0257] In some embodiments, any of the Cas9 fusion proteins provided herein is , one or more of the N497X, R661X, Q695X, and / or Q926X mutations, or provided herein This includes a corresponding mutation in any of the amino acid sequences, where X is any amino acid It is an acid. In some embodiments, the Cas9 fusion protein provided herein Any of the following is one or more of the N497A, R661A, Q695A, and / or Q926A mutations, or as specified herein. Includes a corresponding mutation in any of the amino acid sequences provided. Several implementations Morphologically, the Cas9 domain is a D10A mutation, or an amino acid combination provided herein. Includes the corresponding mutation in any of the columns. A Cas9 domain with high fidelity is used in this technique. This is publicly known in the field and will be obvious to those skilled in the art. For example, a high-fidelity Cas9 domestic The source is Kleinstiver, BP, et al. “High-fidelity CRISPR-Cas9 nucleases with no "Detectable genome-wide off-target effects." Nature 529, 490-495 (2016); and S laymaker, IM, et al. “Rationally engineered Cas9 nucleases with improved spec This is described in "Environment." Science 351, 84-88 (2015), and the full content of each is available in reference. This is incorporated herein.

[0258] High-fidelity Cas9 domain mutations are shown in bold and underlined. DKKYSIGL A IGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMT A FDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWG A LSRKLINGIRDKQSGKTILDFLKSDGFANRNFM A LIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETR A ITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD

[0259] Cas9 nuclease has two functional endonuclease regions: RuvC and HNH. Cas9 undergoes a conformational change when it binds to target DNA, which then becomes a nucleazed. The main component is positioned and the opposite strand of the target DNA is cut. The final result of Cas9-mediated DNA cleavage. This is a double-strand break (DSB) in the target DNA (approximately 3-4 nucleotides upstream of the PAM sequence). SB is repaired by one of the following two common repair pathways: (1) Efficient but error-free Incidentally, the non-homologous end joining (NHEJ) pathway; or (2) homology induction with low efficiency but high fidelity HDR (recovery) route.

[0260] The "efficiency" of non-homologous end joining (NHEJ) and / or homology-oriented repair (HDR) is determined by the ease of use. It can be calculated by a convenient method. For example, in some cases, the efficiency is the success of HDR It can be expressed as a percentage. For example, cleavage production using nuclease analysis for testing. It is possible to produce a substance and calculate the percentage using the ratio of the product to the substrate. For example, directly cutting DNA containing restriction sequences newly incorporated as a result of successful HDR. Nuclease enzymes for measurement can be used. The more substrate is cleaved, the higher the HDR percentage. This indicates that the HDR efficiency was higher. As an example, the percentage of HDR (percentage) is shown. Centage can be calculated using the following formula: [(cleavage product) / (substrate + cleavage product)] (For example, (b+c) / (a+b+c) where "a" is the band intensity of the DNA substrate and "b" is the band intensity of the DNA substrate. (where "c" is a cleavage product.)

[0261] In some cases, efficiency can be expressed as the success rate of NHEJ. For example, T7 endonuclea A cleavage product was generated using the -ase I assay, and the ratio of the product to the substrate was used to determine the parsing of NHEJ. The duration can be calculated. T7 endonuclease I is wild-type and mutant D High NA chain (NHEJ produces small random insertions or deletions (indels) at the initial cleavage site) It cleaves mismatched heterodouble-stranded DNA resulting from hybridization. The more cleavage there is, the better. This indicates a high proportion of NHEJ (high efficiency of NHEJ). An example of this is the proportion of NHEJ. (Percentage) is given by the formula (1-(1-(b+c) / (a+b+c)) 1 / 2 It can be calculated using ) × 100. Here, "a" is the band intensity of the DNA substrate, and "b" and "c" are the cleavage products. Ran et. al., 2013 Sep. 12; 154(6):1380-9; and Ran et. al., Nat Protoc. 2013 No v.; 8(11): 2281-2308).

[0262] The NHEJ repair pathway is the most active repair mechanism, involving small nucleotide insertions or deletions at the DSB site. It frequently causes loss (indel). The randomness of NHEJ-mediated DSB repair is due to Cas9 and gRNA Alternatively, a population of cells expressing guide polynucleotides may result in a diverse array of mutations. Therefore, it has important practical implications. In most cases, NHEJ produces small indices in the target DNA. This causes immaturity within the open reading frame (ORF) of the target gene. This results in amino acid deletions, insertions, or frameshift mutations that lead to stop codons. The ideal final outcome is a loss-of-function mutation within the target gene.

[0263] NHEJ-mediated DSB repair often disrupts the open reading frame of genes, Same-sex-oriented repair (HDR) involves the addition of fluorophores or tags from single nucleotide changes. It can be used to generate specific nucleotide changes ranging from large insertions to large ones.

[0264] To use HDR for gene editing, a DNA repair template containing the desired sequence is created using gRNA. It can be delivered to the target cell type together with Cas9 or Cas9 nickase. Repair Temp The rate is determined by the desired edit and the area immediately upstream and downstream of its target (referred to as left and right identical arms). Further homologous arrangements can be included in ). The length of each homologous arm is the large change introduced. Depending on the size, larger insertions require longer homologous arms. A single repair mold is used. They are single-stranded oligonucleotides, double-stranded oligonucleotides, or double-stranded DNA plasmids. The efficiency of HDR is obtained in cells expressing Cas9, gRNA, and exogenous repair templates. However, it is generally low (less than 10% modified allele). HDR occurs between the S phase and G2 phase of the cell cycle. By synchronizing cells, the efficiency of HDR can be improved. Genes involved in NHEJ Chemically or genetically inhibiting offspring may also increase the frequency of HDR.

[0265] In some embodiments, Cas9 is a modified Cas9. Given a gRNA target sequence, The entire mass may have further parts where partial homology exists. These parts are This is called a target, and it needs to be considered when designing gRNA, but it optimizes the design of gRNA. In addition to this, modifying Cas9 can also enhance the specificity of CRISPR. Cas9 performs double-strand cleavage (DS) via the combined activity of two nuclease domains, RuvC and HNH. B) is produced. Cas9 nickase, a D10A mutant of SpCas9, has one nuclease domain. It retains and generates DNA nicks instead of DSBs. HDR-mediated genes for specific gene editing. You can also combine editing with a nickers effect.

[0266] In some cases, Cas9 is a variant Cas9 protein. Variant Cas9 polypeptide. This differs from the amino acid sequence of the wild-type Cas9 protein at the single amino acid level (for example) It has an amino acid sequence (with deletions, insertions, substitutions, and fusions). In some examples, The rianto-Cas9 polypeptide reduces the nuclease activity of the Cas9 polypeptide. It undergoes acidic changes (e.g., deletion, insertion, or substitution). For example, in some cases, Ant-Cas9 polypeptide exhibits 50% less nuclease activity than the corresponding wild-type Cas9 protein. Having a percentage of 10%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1%. The variant Cas9 protein has virtually no nuclease activity. If the substance is a variant Cas9 protein that does not have substantial nuclease activity, This can be referred to as "dCas9".

[0267] In some cases, variant Cas9 proteins reduce nuclease activity. For example, The variant Cas9 protein is similar to the wild-type Cas9 protein (e.g., wild-type Cas9 protein). Endonuclease activity is less than approximately 20%, less than approximately 15%, less than approximately 10%, less than approximately 5%, less than approximately 1%, and This indicates a figure of less than approximately 0.1%.

[0268] In some cases, the variant Cas9 protein cleaves the complementary strand of the guide target sequence. This is possible, but the ability to cleave the non-complementary strand of the double-stranded guide target sequence is reduced. For example, The variant Cas9 protein is a mutation (amino acid substitution) that reduces the function of the RuvC domain. It may have a variant. In some non-limiting embodiments, The Cas9 protein has D10A (aspartic acid to alanine at amino acid position 10). Therefore, it is possible to cleave the complementary strand of the double-stranded guide target sequence, but the double-stranded guide The ability to cleave the non-complementary strand of the target sequence is reduced (therefore, this variant Cas9 When a protein cleaves a double-stranded target nucleic acid, it produces a single-strand break (S) instead of a double-strand break (DSB). (This can lead to SB) (see, for example, Jinek et al., Science. 2012 Aug. 17; 337(6096):816-21) .

[0269] In some cases, the variant Cas9 protein cleaves the non-complementary strand of the double-stranded guide target sequence. While this is possible, the ability to cleave the complementary strand of the guide target sequence is reduced. For example, The variant Cas9 protein has mutations (amino acid substitutions) that reduce the function of the HNH domain. It is possible (RuvC / HNH / RuvC domain motif). As a non-exclusive example, several In this embodiment, the variant Cas9 protein is H840A (his at amino acid position 840) It has a thidine-to-alanine mutation, and therefore a non-complementary strand of the guide target sequence. It can cleave the target strand, but its ability to cleave the complementary strand of the guide target sequence is reduced. (Therefore, this variant Cas9 protein cleaves the double-stranded guide target sequence.) (This results in SSBs instead of DSBs). Such Cas9 proteins have a guide target sequence (for example) The ability to cleave single-stranded guide target sequences (for example, single strands) is reduced, but the ability to cleave guide target sequences (for example, single strands) is reduced. It retains the ability to bind to the chain guide target sequence.

[0270] In some cases, the variant Cas9 protein interacts with the complementary and non-complementary strands of the double-stranded target DNA. The ability to cut both is reduced. As a non-limiting example, in some cases, variant The Cas9 protein has both D10A and H840A mutations, and as a result the polypeptide This reduces the ability to cleave both the complementary and non-complementary strands of double-stranded target DNA. The Cas9 protein has reduced ability to cleave target DNA (e.g., single-stranded target DNA). However, it retains the ability to bind to target DNA (e.g., single-stranded target DNA).

[0271] As another non-limiting example, in some cases, the variant Cas9 protein is W4 The polypeptide has 76A and W1126A mutations, and as a result the polypeptide targets the DNA (e.g., single-stranded target). Although its ability to cleave target DNA (e.g., single-stranded target DNA) is reduced, its ability to bind to target DNA (e.g., single-stranded target DNA) is reduced. The power is still there.

[0272] As another non-limiting example, in some cases, variant Cas9 proteins are P4 It has the 75A, W476A, N477A, ​​D1125A, W1126A, and D1127A mutations, and as a result, polyp The plutidops has reduced ability to cleave target DNA. Such Cas9 proteins are unable to cleave target DNA. It has a reduced ability to cleave NA (e.g., single-stranded target DNA), but it does not cleave target DNA (e.g., single-stranded target It retains the ability to bind to DNA.

[0273] As another non-limiting example, in some cases, variant Cas9 proteins are H8 The polypeptide has 40A, W476A, and W1126A mutations, and as a result the polypeptide targets the DNA (for example) Although the ability to cleave single-stranded target DNA is reduced, the ability to cleave target DNA (e.g., single-stranded target DNA) The ability to bond is retained. As another non-limiting example, in some cases, barrier AntCas9 protein has H840A, D10A, W476A, and W1126A mutations, and as a result The polypeptide has a reduced ability to cleave target DNA. It has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but it does not cleave target DNA (e.g., It retains the ability to bind to single-stranded target DNA. In some embodiments, Varian In Cas9, the catalytic His residue at position 840 of the Cas9 HNH domain has been restored (A840H). .

[0274] As another non-limiting example, in some cases, variant Cas9 proteins are H8 It has the 40A, P475A, W476A, N477A, ​​D1125A, W1126A, and D1127A mutations, and as a result, Polypeptides reduce the ability to cleave target DNA (e.g., single-stranded target DNA), but target D It retains the ability to bind to NA (e.g., single-stranded target DNA). As another non-limiting example, several In that case, the variant Cas9 protein is D10A, H840A, P475A, W476A, N477A, The polypeptide has D1125A, W1126A, and D1127A mutations, and as a result, the polypeptide targets the DNA. The ability to cleave is reduced. Such Cas9 proteins cannot cleave target DNA (e.g., single-stranded DNA). It has reduced ability to cleave target DNA (e.g., single-stranded target DNA), but its ability to bind to target DNA (e.g., single-stranded target DNA) is reduced. It retains its strength. When the variant Cas9 protein has the W476A and W1126A mutations, Alternatively, the Cas9 protein variants P475A, W476A, N477A, ​​D1125A, W1126A, and D1127 In cases with the A mutation, the variant Cas9 protein does not efficiently bind to the PAM sequence. In such cases, this variant Cas9 protein is used as the binding method. Therefore, this method does not require a PAM sequence. In other words, in some cases, such a barrier When using the Cas9 protein as the binding method, this method may include guide RNA, This method can be performed in the absence of the PAM sequence (therefore, the binding specificity is guided by the RN). (Brought about by the target segment of A). To achieve the above effect, other residues are modified. It can cause a difference (i.e., inactivate one or the other nuclease moiety). Non-restrictive example For example, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and The A987 can be modified (i.e., substituted). Also, mutations other than alanine substitution are possible. This is also preferable.

[0275] In one embodiment, a variant Cas9 protein having reduced catalytic activity (e.g., Ca The s9 protein is D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, or / or A987 mutations, e.g., D10A, G12A, G17A, E762A, H840A, N854A, N863A, H98 If it has 2A, H983A, A984A, and / or D986A, it interacts with guide RNA. As long as it retains the ability to do so, it can bind to target DNA in a site-specific manner (guide RNA). (This is because it is induced to the target DNA sequence.)

[0276] In some embodiments, the variant Cas protein is spCas9, spCas9-VRQR, spCas9 -VRER, xCas9 (sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas It could be 9-LRVSQL.

[0277] As an alternative to Cas9 in S. pyogenes, the Cpf1 family exhibits cleavage activity in mammalian cells. Possible examples include RNA-induced endonucleases derived from Prevotella and Francisella 1. The next generation of CRISPR (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. 1 is a class II CRISPR / Cas RNA-induced endonuclease. This adaptive immune mechanism is Pre It is found in bacteria such as votella and Francisella. The Cpf1 gene is associated with the CRISPR locus, Encoding an endonuclease that uses guide RNA to locate and cleave viral DNA. Yes, Cpf1 is a smaller and simpler endonuclease than Cas9, and is a limitation of the CRISPR / Cas9 system. Overcoming some limitations. Unlike Cas9 nucleases, the result of DNA cleavage via Cpf1 is short. This is a double-strand break with a 3' overhang. The alternating cleavage pattern of Cpf1 is similar to that of traditional restriction enzymes. This opens up the possibility of directional gene transfer, similar to cloning, and this is a legacy This can improve the efficiency of gene editing. Similar to the Cas9 variants and orthologues mentioned above, Cpf1 This means that the number of sites that CRISPR can target is limited to ATs that lack the NGG PAM sites preferred by SpCas9. It can also be expanded to AT-rich regions or AT-rich genomes. The Cpf1 locus is an α / β mixed domain. In, RuvC-I and the subsequent helical region, RuvC-II and zinc finger-like domains It contains. The Cpf1 protein is a RuvC-like endonuclease genome similar to the RuvC domain of Cas9. It has an ine molecule. Furthermore, Cpf1 does not have an HNH endonuclease region, and the N-terminus of Cpf1 is Cas9 It lacks an alpha-helix recognition lobe. The Cpf1 CRISPR-Cas domain configuration is such that Cpf1 is functional It was shown to be unique and classified as a Class 2, Type V CRISPR system. The pf1 locus is more similar to the Cas1, Cas2, and Cas4 proteins of the type I and type III systems than to the type II system. It coded quality. Functional Cpf1 does not require trans-activated CRISPR RNA (tracrRNA). Therefore, only CRISPR (crRNA) is required. Cpf1 is not only smaller than Cas9, but also Because it has a small sgRNA molecule (about half the number of nucleotides as Cas9), this is used for genome editing. It is beneficial. In contrast to the G-rich PAM targeted by Cas9, the Cpf1-crRNA complex is a motif. The target DNA or RNA is cleaved by identifying the protospacer adjacent to 5'-YTN-3'. After identification, Cpf1 has sticky-end-like DNA with 4 or 5 nucleotide overhangs. We will introduce a true strand break.

[0278] [Protospacer adjacent motif] The term "protospacer adjacent motif (PAM)" or PAM-like motif is a CRISPR-compatible bacterial motif. The 2-6 base pairs of DNA immediately following the DNA sequence targeted by the Cas9 nuclease in the immune response. This refers to the sequence. In some embodiments, PAM is 5'PAM (i.e., the 5' end of the protospacer). In other embodiments, the PAM may be 3'PAM (i.e., Protospex). It could be located downstream of the 5' end of the sulcus.

[0279] The PAM sequence is essential for target binding, but its exact sequence depends on the type of Cas protein. To exist.

[0280] The base editors provided herein are adjacent to standard or non-standard protospacers. CRISPR proteins that can bind to nucleotide sequences containing motif (PAM) sequences. It may include the origin domain. The PAM site is adjacent to the target polynucleotide sequence. It is a creotide sequence. Several aspects of this disclosure are CRISPR sequences with different PAM specificities. Provides a base editor containing all or part of an protein. For example, derived from S. pyogenes. Cas9 proteins, such as Cas9 (spCas9), typically bind to specific nucleic acid regions. A standard NGG PAM sequence is required, where "N" in "NGG" stands for adenine (A) and thymine (T). PAM is a CRISPR protein. It can be protein-specific and contains different base editors with domains derived from different CRISPR proteins. It can vary between the two. The PAM can be at 5' or 3' of the target sequence. The PAM can be upstream of the target sequence or It can be downstream. PAM is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotids. It can be a certain length. In most cases, PAMs are 2-6 nucleotides long. Some PAMs The variants are listed in Table 1.

[0281] In some embodiments, SpCas9 is used in relation to the PAM nucleic acid sequence 5'-NGC-3' or 5'-NGC-3'. It has the characteristic of being. In various embodiments of the above-described model, SpCas9 is Cas9 or Table 1 The listed Cas9 variants are: In various embodiments of the above configuration, the modified SpCas9 is: The variant Cas protein is spCas9-MQKFRAER. In some embodiments, the variant Cas protein is sp Cas9, spCas9-VRQR, spCas9-VRER, xCas9 (sp), saCas9, saCas9-KKH, SpCas9-MQKFRAER It could be spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL. One specific implementation In this state, amino acid substitutions include D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and A modified PAM 5'-NGC-3' with specificity for the modified PAM 5'-NGC-3' containing T1337R (SpCas9-MQKFRAER). A modified SpCas9 is used.

[0282] In some embodiments, PAM is NGT. In some embodiments, NGT PAM is a variant. Yes. In some embodiments, the NGT PAM variant is one or more residues 1335 , produced through targeted mutations in 1337, 1135, 1136, 1218, and / or 1219 In some embodiments, the NGT PAM variant is one or more residues 1219, 1335, 133 It is generated through targeted mutations in 7,1218. In some embodiments, NGT PAM The variant is a targeted mutation at one or more residues 1135, 1136, 1218, 1219, and 1335. It is produced through the following process. In some embodiments, the NGT PAM variant is shown in Table 2 below. Selected from the set of targeted mutations provided to 3.

[0283] Table 2: NGT PAM variant mutations at residues 1219, 1335, 1337, and 1218 [Table 2]

[0284] Table 3: NGT PAM variant mutations at residues 1135, 1136, 1218, 1219, and 1335 [Table 3]

[0285] In some embodiments, the NGT PAM variants are variants 5, 7, and 2 in Tables 2 and 3. Selected from 8, 31, or 36. In some embodiments, the variant is an improved NGT. It has PAM recognition.

[0286] In some embodiments, the NGT PAM variant is residues 1219, 1335, 1337, and / or has a mutation at 1218. In some embodiments, the NGT PAM variant is as follows: Select from the variants provided in Table 4, with mutations to improve recognition. .

[0287] Table 4: NGT PAM variant mutations at residues 1219, 1335, 1337, and 1218 [Table 4]

[0288] In some embodiments, the NGT PAM is selected from the variants provided in Table 5 below. .

[0289] Table 5: NGT PAM Variants [Table 5]

[0290] In one aspect, the Cas9 domain is a Cas9 domain derived from Streptococcus pyogenes. (SpCas9). In one embodiment, the SpCas9 domain is nuclease activity SpCas9, nuclease These are either ase-inactive SpCas9 (SpCas9d) or SpCas9 nickase (SpCas9n). In this embodiment, SpCas9 is a D9X mutation, or an amino acid combination provided herein. The column contains the corresponding mutation, where X is any amino acid other than D. In some embodiments, SpCas9 is a D9A mutation, or provided herein Includes a corresponding mutation in any of the amino acid sequences. In one embodiment, SpCas9 The main, SpCas9d, or SpCas9n domain binds to nucleic acid sequences with non-standard PAMs. This is possible. In one embodiment, a SpCas9 domain, a SpCas9d domain, or a SpCas9n domain. The domain can bind to nucleic acid sequences containing NGG, NGA, or NGCG PAM sequences.

[0291] In some embodiments, the SpCas9 domain is D1135X, R1335X, and T1336X autism. One or more mutations, or corresponding amino acid sequences in any of the amino acid sequences provided herein. This includes mutations, where X is any amino acid. In some embodiments, SpCas The 9 domains are one or more of the D1135E, R1335Q, and T1336R mutations, or provided herein. This includes a corresponding mutation in any of the amino acid sequences. In some embodiments, In this context, the SpCas9 domain is a mutation of D1135E, R1335Q, and T1336R, or as specified herein. Includes the corresponding mutation in any of the provided amino acid sequences. Several implementations In this state, the SpCas9 domain is one of the D1135X, R1335X, and T1336X mutations. The corresponding mutation in any of the amino acid sequences provided above or herein It contains different amino acids, where X is any amino acid. In some embodiments, SpCas9 dome In is one or more of the D1135V, R1335Q, and T1336R mutations, or provided herein. This includes a corresponding mutation in any of the amino acid sequences. In some embodiments, The SpCas9 domain is a mutation of D1135V, R1335Q, and T1336R, or provided herein. Includes a corresponding mutation in any of the amino acid sequences. In some embodiments, Furthermore, the SpCas9 domain has one or more D1135X, G1217X, R1335X, and T1336X mutations. or includes a corresponding mutation in any of the amino acid sequences provided herein, Here, X is any amino acid. In some embodiments, the SpCas9 domain is D1135 One or more of the V, G1217R, R1335Q, and T1336R mutations, or the ami provided herein. Includes a corresponding mutation in any of the no-acid sequences. In some embodiments, Sp The Cas9 domain is a mutation of D1135V, G1217R, R1335Q, and T1336R, or as specified herein. Includes a corresponding mutation in any of the amino acid sequences provided.

[0292] In some embodiments, any of the fusion proteins provided herein may contain Cas9 The domain is at least 60% of the Cas9 polypeptide described herein, and at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90% , at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or contains an amino acid sequence that is at least 99.5% identical. In some embodiments, The Cas9 domain of any of the fusion proteins provided herein is as described herein. The present invention includes an amino acid sequence of any Cas9 polypeptide. In some embodiments, Any Cas9 domain of the fusion protein provided in this document is as described herein. It consists of the amino acid sequence of the Cas9 polypeptide.

[0293] In some cases, the base editors disclosed herein are derived from CRISPR proteins. The PAM recognized by the domain is an insert that codes for a base editor (e.g., AAV). The insert can be delivered to cells on a separate oligonucleotide. Such implementations In this state, providing PAM on a separate oligonucleotide is otherwise the target sequence and Targets that cannot be cleaved because there are no adjacent PAMs on the same polynucleotide. Enables array splitting.

[0294] In one embodiment, S. pyogenes Cas9 (SpCas9) is used for CRISPR engineering for genome engineering. It can be used as a donuclease. However, other substances may also be used. In that embodiment, a specific genomic target is targeted using different endonucleases. This is possible. In some embodiments, a synthetic SpCas9-derived varian having a non-NGG PAM sequence is used. It can be used. Furthermore, other Cas9 orthologues from various species have been identified. These "non-SpCas9" sequences can bind to various PAM sequences, which may also be useful in this disclosure. For example, a relatively large SpCas9 (a coding sequence of about 4 kilobases (kb)) is found inside an cell. This may result in a SpCas9 cDNA plasmid that cannot be efficiently expressed. Conversely, the coding sequence of Staphylococcus aureus Cas9 (SaCas9) is approximately 1 kilosalt higher than that of SpCas9. Because it is short, it can be efficiently expressed within cells. Similar to SpCas9, SaCas9 endonuclea Ze has the ability to modify target genes in in vitro mammalian cells and in vivo in mice. In some embodiments, the Cas protein can target different PAM sequences. In some embodiments, the target gene may be adjacent to a Cas9 PAM, such as 5'-NGG. In other embodiments, other Cas9 orthologues may have different PAM requirements. For example, S. ther PAM like that found in Mophilus (5'-NNAGAA for CRISPR1, 5'-NGGNG for CRISPR3) Other PAMs, such as those from Neisseria meningiditis (5'-NNNNGATT), are also target genes. It can be found adjacent to it.

[0295] In some embodiments, for the S. pyogenes system, the target gene sequence is 5'-NGG P It may be located before the AM (i.e., on its 5' side), and a 20 nt guide RNA sequence may be present, with base pairs forming with the opposite strand. It can be formed to mediate adjacent Cas9 cuts to the PAM. In one embodiment, adjacent cuts The break may be approximately 3 base pairs upstream of the PAM. In some embodiments, the adjacent break may be approximately 3 base pairs upstream of the PAM. ) may be 10 base pairs upstream. In one embodiment, the adjacent break is (about) 0 to 20 base pairs above the PAM. It could be a flow. For example, adjacent cuts are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 upstream of PAM. , 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 salt It can be adjacent to the base pair. Adjacent breaks can also occur 1 to 30 base pairs downstream of the PAM. The following is the PAM sequence. The following are exemplary SpCas9 protein sequences that can bind:

[0296] An example amino acid sequence of PAM-binding SpCas9 is as follows: MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD.

[0297] The amino acid sequence of an example PAM-binding SpCas9n is as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD.

[0298] The amino acid sequence of an example PAM-binding SpEQR Cas9 is as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGF ESPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK Q Y R STKEVLDATLIHQSITGLYETRID LSQLGGD In this sequence, mutations from D1135, R1335, and T1337 result in SpEQR Cas9. The residues E1135, Q1335, and R1337, which can be formed, are shown in underlined and bold.

[0299] The amino acid sequence of an example PAM-bound SpVQR Cas9 is as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGF V SPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK Q Y R STKEVLDATLIHQSITGLYETRI DLSQLGGD In this sequence, mutations from D1135, R1335, and T1336 result in SpVQR Cas9. The residues V1135, Q1335, and R1336, which can be used, are shown in underlined and bold.

[0300] An example of a PAM-binding SpVRER Cas9 amino acid sequence is as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGF V SPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELINGRKRMLASA R ELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRK E Y R STKEVLDATLIHQSITGLYETRI DLSQLGGD

[0301] In one embodiment, the Cas9 domain is a recombinant Cas9 domain. In one embodiment, The recombinant Cas9 domain is the SpyMacCas9 domain. In some embodiments, SpyM The acCas9 domain is divided into nuclease-active SpyMacCas9 and nuclease-inactive SpyMacCas9 (SpyM This is acCas9d) or SpyMacCas9 nickas (SpyMacCas9n). In some embodiments, In this context, SaCas9 domains, SaCas9d domains, or SaCas9n domains have non-standard PAM. It can bind to nucleic acid sequences. In some embodiments, SpyMacCas9 domain The SpCas9d domain or SpCas9n domain binds to nucleic acid sequences containing the NAA PAM sequence. It is possible.

[0302] Exemplary SpyMacCas9 MDKKYSIGLDIGTNSVGWAVITDDYKVPSKKFKVLGNTDRHSIKKNLIGALLFGSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLADSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQIYNQLFEENPINASRVDAKAILSARLSKSRRLENLIAQLPGEKRNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNSEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGAYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRGMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGHSL HEQIANLAGSPAIKKGILQTVKIVDELVKVMGHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFIKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEIQTVGQNGGLFDDNPKSPLEVT PSKLVPLKKELNPKKYGGYQKPTTAYPVLLITDTKQLIPISVMNKKQFEQNPVKFLRDRGYQQVGKNDFIKLPKYTLVDI GDGIKRLWASSKEIHKGNQLVVSKKSQILLYHAHHLDSDLSNDYLQNHNQQFDVLFNEIISFSKKCKLGKEHIQKIENVY SNKKNSASIEELAESFIKLLGFTQLGATSPFNFLGVKLNQKQYKGKKDYILPCTEGTLIRQSITGLYETRVDLSKIGED.

[0303] In some cases, the variant Cas9 protein is H840A, P475A, W476A, N477A, ​​D1125A, W1 It has mutations 126A and D1218A, resulting in a reduced ability to cleave target DNA or RNA. Such Cas9 proteins have the ability to cleave target DNA (e.g., single-stranded target DNA). Although reduced, it retains the ability to bind to target DNA (e.g., single-stranded target DNA). As a non-limiting example, in some cases the variant Cas9 protein D10A, H8 It has the 40A, P475A, W476A, N477A, ​​D1125A, W1126A, and D1218A mutations, and as a result, Polypeptides have a reduced ability to cleave target DNA (e.g., single-stranded target DNA). The Cas9 protein has reduced ability to cleave target DNA (e.g., single-stranded target DNA). The variant Cas9 protein retains the ability to bind to target DNA (e.g., single-stranded target DNA). If the protein has W476A and W1126A mutations, or if the variant Cas9 protein is P475A, If you have the W476A, N477A, ​​D1125A, W1126A, and D1218A mutations, the variant Cas9 protein The protein does not efficiently bind to the PAM sequence. Therefore, in such cases, When the riant Cas9 protein is used as the binding method, this method does not require a PAM sequence. In other words, in some cases, such variant Cas9 proteins are used as a binding method. In some cases, this method may include guide RNA, but this method can be performed in the absence of the PAM sequence. (Therefore, binding specificity is provided by the target segment of the guide RNA.) To achieve the above effect, other residues may be mutated (i.e., one or the other nucleus). (Inactivates the crease moiety). Non-limiting examples include residues D10, G12, G17, E762, and H840. Modify (i.e., replace) N854, N863, H982, H983, A984, D986, and / or A987. This is possible. Furthermore, mutations other than alanine substitution are also preferred.

[0304] In one embodiment, the CRISPR protein-derived domain of the base editor is a standard PAM compound. It may include all or part of the Cas9 protein having a row (NGG). In other embodiments, the base The Cas9-derived domain of the editor can use non-standard PAM sequences. The sequences are described in the art and will be obvious to those skilled in the art. For example, non-standard PAM sequences The Cas9 domains that bind to the column are described in Kleinstiver, BP, et al., “Engineered CRISPR-Cas9 “Nucleases with altered PAM specificities” Nature, 523, 481-485 (2015); and Kleinstiver, BP, et al., “Broadening the targeting range of Staphylococcus a ureus CRISPR-Cas9 by modifying PAM recognition”Nature Biotechnology, 33, 1293-1 This is described in 298 (2015), and its entire contents are incorporated here by reference.

[0305] [Fusion proteins containing nuclear localization sequences (NLS)] In some embodiments, the fusion proteins provided herein may be one or more (e.g., 2, 3, 4, 5) further includes nuclear targeting sequences, e.g., nuclear localization sequences (NLS). In terms of form, bipartite NLS is used. In some embodiments, NLS is NLS Amino acid sequences that promote the import of proteins containing into the cell nucleus (e.g., by nuclear transport) Includes. In some embodiments, the fusion protein provided herein One of these further includes a nuclear localization sequence (NLS). In some embodiments, the NLS is fused It is fused to the N-terminus of the protein. In some embodiments, the NLS is the fusion protein It is fused to the C-terminus. In some embodiments, the NLS is fused to the N-terminus of the Cas9 domain. In some embodiments, the NLS is located at the C-terminus of the nCas9 domain or dCas9 domain. They are fused. In some embodiments, the NLS is fused to the N-terminus of the deaminase. In some embodiments, NLS is fused to the C-terminus of the deaminase. In one embodiment, the NLS is fused to the fusion protein via one or more linkers. In this configuration, the NLS is fused to the fusion protein without a linker. In some embodiments... In this context, NLS is any one amino of the NLS sequences provided or referenced herein. It contains an acid sequence. Further nuclear localization sequences are known in the art and will be obvious to those skilled in the art. For example, the NLS sequence is described in Plank et al., PCT / EP2000 / 011690, and among them The details are incorporated herein by reference with respect to the disclosure of exemplary nuclear localization sequences. In this context, NLS has the amino acid sequence PKKKRKVEGADKRTADGSEFES PKKKRKV, KRTADGSEFESPKKKRK V, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRKPKK Includes KRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC. In some embodiments, NL S is present in the linker, or NLS is a linker, for example, a linker as described herein. —is adjacent to. In some embodiments, the N-terminus or C-terminus NLS is bipartite N This is LS. The bipartite NLS is separated by a relatively short spacer sequence into two basic elements. Contains a monopartite cluster (hence called bipartite, two-partite, and monopartite) (NLS is different). The NLS of nucleoplasmin, KR[PAATKKAGQA]KKKK, is a ubiquitous bipartite. It is a prototype signal, with two clusters of basic amino acids forming a group of approximately 10 amino acids. It is separated by a pacer. An example of a bipartite NLS sequence is PKKKRKVEGADKRTA It is DGSEFES PKKKRKV.

[0306] In some embodiments, the fusion protein of the present invention does not contain a linker sequence. In this embodiment, a linker sequence exists between one or more domains or proteins.

[0307] It should be understood that the fusion proteins of this disclosure may include one or more further features. For example, in some embodiments, the fusion protein acts as an inhibitor, localizing to the cytoplasm. Sequences, export sequences such as nuclear export sequences, or other localization sequences, as well as fusion proteins It may include sequence tags useful for dissolution, purification, or detection. Provided herein Appropriate protein tags include, but are not limited to, biotin carboxylase. Carrier tags (BCCP), myc-tags, Calmodulin tags, FLAG-tags, Hemagglutinator Nin (HA)-tag, polyhistidine tag (also called histidine tag or His-tag) , maltose-binding protein (MBP)-tag, nus-tag, glutathione-S-transfer GST-tag, green fluorescent protein (GFP)-tag, thioredoxin tag, S-tag, so FTAGs (e.g., Softag1, Softag3), streptotags, biotin ligase tags, FLAsH tags This includes V5 tags and SBP- tags. Further suitable arrangements will be obvious to those skilled in the art. In some embodiments, the fusion protein includes one or more His tags.

[0308] [Linker] In one embodiment, either the peptide or peptide domain of the present invention is linked. A linker may be used for this purpose. A linker can be as simple as a covalent bond, or It could be a polymer linker with a length of multiple atoms. The linker is a peptide linker. Alternatively, it may be a non-peptide linker. In certain embodiments, the linker is UV-cleavable It can be a linker. In some embodiments, the linker is a polynucleotide linker, e.g. For example, it could be an RNA linker. In one embodiment, the linker is a polypeptide. or is amino acid-based. In other embodiments, the linker is peptide-like. No. In one embodiment, the linker is a covalent bond (e.g., carbon-carbon bond, di These include sulfide bonds, carbon-heteroatom bonds, etc. In one embodiment, a linker This is an amide-linked carbon-nitrogen bond. In certain embodiments, the linker is cyclic or non- It is a cyclic, substituted or unsubstituted, branched or unbranched aliphatic or heteroaliphatic linker. In the embodiment, the linker is a polymer (e.g., polyethylene, polyethylene glyco) (Materials include polyamide, polyester, and others). In certain embodiments, the linker is It comprises monomers, dimers, or polymers of minoalkanoic acid. In one embodiment, phosphorus Car is an aminoalkanic acid (e.g., glycine, ethane, alanine, beta-alanine). Includes 3-aminopropanoic acid, 4-aminobutanoic acid, 5-pentanoic acid, etc. (Specific Embodiments) Therefore, the linker contains a monomer, dimer, or polymer of aminohexanoic acid (Ahx). In one embodiment, the linker is a carbocyclic moiety (e.g., cyclopentane, cyclohexane). ) is based on. In other embodiments, the linker is polyethylene glycol moiety (PEG) It includes. In other embodiments, the linker includes an amino acid. In one embodiment, the linker - contains peptides. In one embodiment, the linker is aryl or heteroaryl. It includes a portion. In one embodiment, the linker is based on a phenyl ring. The linker is Promotes the binding of nucleophiles (e.g., thiols, aminos) from peptides to linkers. It may include a functionalized portion for this purpose. Any electrophile can be used as part of the linker. Examples of electrophiles include activated esters, activated amides, Michael acceptors, and Alkyl chlorogenates, aryl halides, acyl halides, and isothiocyanates These include, but are not limited to, the following:

[0309] In one embodiment, the linker is one amino acid or a plurality of amino acids (e.g., peptides) It is a protein or a dendritic fluid. In some embodiments, the linker binds (for example, It is a covalent bond, organic molecule, group, polymer, or chemical part. In some embodiments, it is a covalent bond, organic molecule, group, polymer, or chemical part. Linkers are approximately 3 to 104 in length (for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16) , 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36 ,37,38,39,40,41,42,43,44,45,46,47,48,49,50,55,60,61,62,63,64 These are amino acids (65, 70, 75, 80, 85, 90, 95, or 100).

[0310] [Cas9 complex with guide RNA] Several aspects of this disclosure include multiple fusion proteins, including any of the fusion proteins provided herein. To provide a fusion and achieve the optimal length for the activity of the nucleic acid base editor, guide RNA can be used (for example, (GGGS) n (GGGGS) n , and (G) n A very flexible phosphorus in shape From the car, (EAAAK) n (SGGS) n up to a more rigid linker in the form of SGSETPGTSESATPES (For example, Guilinger JP, Thompson DB, Liu DR. Fusion of catalytically inactive Cas9) to FokI nuclease improves the specificity of genome modification. Nat. Biotechn See ol. 2014; 32(6): 577-82 (the full content of which is incorporated here by reference), and (XP) n ). In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 , 13, 14 or 15. In one embodiment, the linker is (GGS) n Including the motif, Here, n is 1, 3, or 7. In some embodiments, provided herein The Cas9 domain of the fusion protein is accessed via a linker containing the amino acid sequence SGSETPGTSESATPES. They are fused together.

[0311] In one embodiment, the guide nucleic acid (e.g., guide RNA) is 15 to 100 nucleotides long. It contains a sequence of at least 10 consecutive nucleotides that are complementary to the target sequence. In this embodiment, the guide RNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 , 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46 The length is 47, 48, 49, or 50 nucleotides. In some embodiments, The doRNAs are complementary to the target sequence, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 2 8, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 consecutive nucleotides It includes the sequence of. In one aspect, the target sequence is a DNA sequence. In one aspect, the target A sequence is a sequence found in the genome of bacteria, yeast, fungi, insects, plants, or animals. In some embodiments, the target sequence is a sequence in the human genome. The 3' end of the target sequence is immediately adjacent to the standard PAM sequence (NGG). Several implementations In this state, the 3' end of the target sequence is a non-standard PAM sequence (for example, the sequences listed in Table 1). It is immediately adjacent to. In one embodiment, guide nucleic acids (e.g., guide RNA) are disease and It is complementary to the sequences associated with the disorder.

[0312] Some aspects of this disclosure utilize fusion proteins or complexes provided herein. This provides a method for doing so. For example, several aspects of this disclosure involve introducing DNA molecules within this specification. Contact one of the provided fusion proteins and at least one guide RNA. This method provides a method in which the guide RNA is approximately 15-100 nucleotides long and the target RNA The column contains a sequence of at least 10 consecutive nucleotides that are complementary to the column. Several implementations Morphologically, the 3' end of the target sequence is immediately adjacent to an AGC, GAG, TTT, GTG, or CAA sequence. In some embodiments, the 3' end of the target sequence is NGA, NGCG, NGN, NNGRRT. , NNNRRT, NGCG, NGCN, NGTN, NGTN, NGTN, or immediately adjacent to the 5' (TTTV) sequence .

[0313] In some embodiments, the fusion protein of the present invention induces mutations in the target of interest. These mutations are used for the purpose of affecting the function of the target. For example, nucleic acid bases When a regulator is targeted using a DITER, the function of the regulatory region is altered, affecting downstream proteins. The expression of quality decreases.

[0314] The numbering of specific positions or residues in each sequence is determined by the specific protein used. It will be understood that this depends on the quality and numbering scheme. For example, mature tan The numbering system can differ between protein precursors and the mature protein itself, depending on the species. Differences in arrangement can affect numbering. A person skilled in the art can use methods well known to those skilled in the art, for example, Sequence alignment and homologous residue determination can be used to identify any homologous protein and its Each residue in each coding nucleic acid can be identified.

[0315] Any of the fusion proteins disclosed herein can be used to target a site, for example, an edited aggravating mutation. To target the site containing the mutation, the fusion protein is co-developed with the guide RNA. It will be obvious to those skilled in the art that it is typically necessary to make it apparent. As will be explained in more detail by the way, guide RNA typically enables Cas9 binding. The crRNA framework and Cas9: nucleic acid editing enzyme / domain fusion protein are given sequence specificity. It includes a contributing guide sequence. Alternatively, the guide RNA and tracrRNA are two nucleic acid molecules. They may be provided separately. In some embodiments, the guide RNA is a guide sequence with a target sequence. It contains a structure that includes a complementary sequence. The guide sequence is typically 20 nucleotides long. Cas9: Nucleic acid editing enzyme / domain fusion protein targets specific genomic target sites. The appropriate guide RNA sequence for this purpose will be apparent to those skilled in the art based on this disclosure. A suitable guide RNA sequence, such as the one described above, is typically 50 nucleotides of the target nucleotide to be edited. The provided fusion includes a guide sequence complementary to the nucleic acid sequence within the upstream or downstream of the ocidal region. Some exemplary guides suitable for targeting any of the proteins to a specific target sequence A doRNA sequence is provided herein.

[0316] [Fusion protein containing cytidine deaminase, adenosine deaminase, and Cas9 domain] How to use the product Some aspects of this disclosure utilize fusion proteins or complexes provided herein. This provides a method for doing so. For example, several aspects of this disclosure involve introducing DNA molecules within this specification. Contact one of the provided fusion proteins and at least one guide RNA. This method provides a method in which the guide RNA is approximately 15-100 nucleotides long and the target RNA The column contains a sequence of at least 10 consecutive nucleotides that are complementary to the column. Several implementations Morphologically, the 3' end of the target sequence is directly adjacent to the canonical PAM sequence (NGG). In this embodiment, the 3' end of the target sequence is not directly adjacent to the canonical PAM sequence (NGG). In some embodiments, the 3' end of the target sequence is an AGC, GAG, TTT, GTG, or CAA sequence. It is immediately adjacent to it. In some embodiments, the 3' end of the target sequence is NGA, NGCG, NGN, NNGRRT, NNNRRT, NGCG, NGCN, NGTN, NGTN, NGTN, or immediately adjacent to the 5' (TTTV) sequence. They are in contact.

[0317] In some embodiments, the fusion protein of the present invention induces mutations in the target of interest. Used for the purpose of, in particular, the multi-effector nucleic acid base editor described herein. This can create multiple mutations within the target sequence. These mutations affect the function of the target. It can have an effect. For example, using a multi-effector nucleic acid base editor to control the regulatory region. Targeting this protein alters the function of the regulatory region and reduces the expression of downstream proteins.

[0318] The numbering of specific positions or residues in each sequence is determined by the specific protein used. It will be understood that this depends on the quality and numbering scheme. For example, mature tan The numbering system can differ between protein precursors and the mature protein itself, depending on the species. Differences in arrangement can affect numbering. A person skilled in the art can use methods well known to those skilled in the art, for example, Sequence alignment and homologous residue determination can be used to identify any homologous protein and its Each residue in each coding nucleic acid can be identified.

[0319] The Cas9 domain and cytidine deaminase or adenosine deaminase disclosed herein One of the fusion proteins containing the enzyme has a target site, for example, a mutation that is edited. To target a specific site, the fusion protein is combined with a guide RNA, such as sgRNA. It will be obvious to those skilled in the art that it is typically necessary to bring it into existence. As will be explained in more detail elsewhere, guide RNA typically enables Cas9 binding. The racrRNA framework and Cas9: nucleic acid editing enzyme / domain fusion protein with sequence specificity It includes the guide sequence to be conferred. Alternatively, the guide RNA and tracrRNA are two nucleic acid molecules and They may be provided separately. In some embodiments, the guide RNA is a guide sequence that is targeted It includes a structure in which the column contains a complementary sequence. The guide sequence is typically 20 nucleotides long. Yes. Cas9: Nucleic acid editing enzyme / domain fusion protein targets specific genomic target sites. A suitable guide RNA sequence for this purpose will be apparent to those skilled in the art based on this disclosure. Such a suitable guide RNA sequence is typically 50 nucleotides of the target nucleotide to be edited. The rheotide contains a guide sequence complementary to the nucleic acid sequence either upstream or downstream. Several exemplary metamorphic proteins suitable for targeting specific target sequences. An id RNA sequence is provided herein.

[0320] [Efficiency of base editors] The fusion protein of the present invention generates a specific nucleo without producing a significant proportion of indels. The efficiency of the base editor is improved by modifying the cydo base. The term "indel" as used here refers to the insertion or deletion of a nucleotide base within a nucleic acid. This refers to a deletion. Such insertions or deletions are frameshift mutations within the coding region of a gene. It can cause abnormalities. In some embodiments, multiple insertions in nucleic acids Alternatively, it can efficiently extract specific nucleotides within nucleic acids without creating deletions (i.e., indels). It is desirable to generate a base editor that modifies (e.g., mutates) the base in a particular embodiment. This means that any of the base editors provided herein is intended to modify the indel. A larger proportion of mutations (e.g., spontaneous mutations) can be generated. In some embodiments, In this specification, the base editors provided are intended mutations greater than 1:1. A ratio to indel can be generated. In some embodiments, provided herein The base editors used are at least 1.5:1, at least 2:1, at least 2.5:1, and less Both 3:1, at least 3.5:1, at least 4:1, at least 4.5:1, at least 5:1, less 5.5:1, at least 6:1, at least 6.5:1, at least 7:1, at least 7.5:1, less At least 8:1, at least 10:1, at least 12:1, at least 15:1, at least 20:1, less At least 25:1, at least 30:1, at least 40:1, at least 50:1, at least 100:1, At least 200:1, at least 300:1, at least 400:1, at least 500:1, at least 600: 1. At least 700:1, at least 800:1, at least 900:1, or at least 1000:1, Or a higher, intended mutation-to-indel ratio can be produced. The number of mutations and indels can be determined using any suitable method.

[0321] In some embodiments, the base editor provided herein is used to define the region of nucleic acids. The formation of indels in the region can be restricted. In one embodiment, the region is a base It is located at the nucleotide targeted by the editor, or at the base editor. Within 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides of the nucleotide targeted by the system. ...

Claims

1. A composition comprising an adeno-associated virus (AAV) vector for use in a method for delivering a base editor system to cells, (a) A first AAV vector comprising a polynucleotide encoding a fusion protein comprising a deaminase and an N-terminal fragment of Cas9 in the order from N-terminus to C-terminus, wherein the N-terminal fragment of Cas9 is a continuous sequence beginning at the N-terminus of Cas9 and ending at positions 302, 309, 312, or 354 of Cas9 in the numbering in SEQ ID NO: 16, wherein the N-terminal fragment of Cas9 is fused to a split intein-N at its C-terminus, and (b) A second AAV vector comprising a polynucleotide encoding the C-terminal fragment of Cas9, wherein the C-terminal fragment of Cas9 is a continuous sequence beginning at position S303 of Cas9 if the N-terminal fragment terminates at position 302, at position T310 of Cas9 if the N-terminal fragment terminates at position 309, at position T313 of Cas9 if the N-terminal fragment terminates at position 312, and at position S355 of Cas9 if the N-terminal fragment terminates at position 354, and terminates at the C-terminus of Cas9, wherein the C-terminal fragment of Cas9 is fused to a split intein-C at its N-terminus, the second AAV vector Includes, The method comprises contacting the cells with the first AAV vector and the second AAV vector and a single guide RNA (sgRNA) or a polynucleotide encoding it. composition.

2. A composition comprising an adeno-associated virus (AAV) vector for use in a method for delivering a base editor system to cells, (a) A first AAV vector comprising a polynucleotide encoding a fusion protein comprising a deaminase and the N-terminal fragment of Cas9 in the order from N-terminus to C-terminus, wherein the N-terminal fragment of Cas9 is a continuous sequence beginning at the N-terminus of Cas9 and ending at positions 302, 309, 312, 354, 465, 471, or 576 of Cas9 as numbered in SEQ ID NO: 16, wherein the N-terminal fragment of Cas9 is fused to split intein-N at its C-terminus, and (b) A second AAV vector comprising a polynucleotide encoding the C-terminal fragment of Cas9, wherein the C-terminal fragment of Cas9 is numbered in SEQ ID NO: 16, at position S303 of Cas9 if the N-terminal fragment terminates at position 302, at position T310 of Cas9 if the N-terminal fragment terminates at position 309, at position T313 of Cas9 if the N-terminal fragment terminates at position 312, and at position S355 of Cas9 if the N-terminal fragment terminates at position 354. A second AAV vector, where the terminal fragment terminates at position 465, the sequence begins at position T466 of Cas9 if the N-terminal fragment terminates at position 471, the sequence begins at position T472 of Cas9 if the N-terminal fragment terminates at position 576, the sequence begins at position S577 of Cas9 and terminates at the C-terminus of Cas9, the N-terminal residue of the C-terminal fragment of Cas9 is a Cys substituted with Ser or Thr, and the C-terminal fragment of Cas9 is fused to a split intein-C at its N-terminus. Includes, The method comprises contacting the cells with the first AAV vector and the second AAV vector and a single guide RNA (sgRNA) or a polynucleotide encoding it. composition.

3. a) The composition further comprises the sgRNA or a polynucleotide encoding it, b) The composition further comprises an AAV vector containing the sgRNA or a polynucleotide encoding it, or c) The first AAV vector or the second AAV vector contains the sgRNA or a polynucleotide encoding it, The composition according to claim 1 or 2.

4. The composition according to claim 1 or 2, wherein the deaminase is adenosine deaminase.

5. The composition according to claim 1 or 2, wherein the deaminase is TadA.

6. The composition according to claim 1 or 2, wherein the deaminase is wild-type TadA or TadA7.

10.

7. The deaminase is a TadA dimer; or The deaminase is a TadA dimer containing wild-type TadA and TadA7.

10. The composition according to claim 1 or 2.

8. The composition according to claim 1 or 2, wherein the fusion protein comprises a nuclear localization signal (NLS).

9. The composition according to claim 1 or 2, wherein the N-terminal fragment or the C-terminal fragment of Cas9 is linked to a nuclear localization signal (NLS).

10. The composition according to claim 1 or 2, wherein the N-terminal fragment of Cas9 or the C-terminal fragment of Cas9 is linked to two NLSs.

11. The composition according to claim 1 or 2, wherein the Cas9 has nickase activity or is catalytically inactive.

12. The first AAV vector and / or the second AAV vector, a) containing a promoter; or b) containing a constitutive promoter; or c) Constitutive promoters that are CMV or CAG promoters, The composition according to claim 1 or 2.

13. The composition according to claim 1 or 2, wherein the C-terminal fragment of Cas9 contains a Ser / Cys or Thr / Cys mutation in the residue corresponding to amino acid S303, T310, T313, or S355 in the numbering in SEQ ID NO:

16.

14. An in vitro or ex vivo method for delivering a base editor system to a cell, the method comprising contacting the cell with the first and second AAV vectors described in claim 1 or 2 and a single guide RNA (sgRNA) or a polynucleotide encoding it.