Nucleobase editors with reduced off-target deamination and methods of using same to modify nucleobase target sequences
Nucleobase editors with tailored editing profiles and multi-effector nucleobase editors address off-target deamination issues by enhancing specificity and efficiency in targeted nucleic acid sequence modifications.
Patent Information
- Application Number
- JP2024210915
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-11-27
- Filing Date
- 2024-12-04
- Publication Date
- 2025-12-03
- Estimated Expiration
- 2040-01-31
AI Technical Summary
Current base editors lack specificity and efficiency in targeted nucleic acid sequence modifications, leading to off-target deamination issues.
Development of nucleobase editors with tailored editing profiles and multi-effector nucleobase editors, incorporating polynucleotide-programmable DNA binding domains and cytidine deaminases with enhanced cis-to-trans activity ratios, to minimize off-target deamination.
The improved nucleobase editors achieve enhanced specificity and efficiency in targeted nucleic acid sequence modifications, reducing off-target effects and increasing the ratio of cis to trans activity.
Smart Images

Figure 0007779988000073 
Figure 0007779988000074 
Figure 0007779988000075
Abstract
Description
[Technical Field]
[0001] Cross Reference This application is a continuation of U.S. Provisional Application No. 62 / 799,702, filed January 31, 2019, and U.S. Provisional Application No. 62 / 799,702, filed April 17, 2019. No. 62 / 835,456 filed on November 27, 2019, and No. 62 / 941,569 filed on November 27, 2019. and international PCT applications claiming the benefit of the following patents, the contents of each of which are incorporated herein by reference in their entirety: be incorporated into the book. [Background technology]
[0002] Targeted editing of nucleic acid sequences, for example, targeted cleavage or targeted modification of genomic DNA, can be used to modify genes. This is a very promising approach for functional studies and may lead to new therapeutic options for human genetic diseases. Currently available base editors can convert a target C G base pair to a T A base pair. A cytidine base editor (e.g., BE4) converts A to G and an adenine to C. Includes a base editor (e.g., ABE7.10). Summary of the Invention [Problem to be solved by the invention]
[0003] The art has demonstrated the ability to induce modifications in target sequences with greater specificity and efficiency. There is a need for improved base editors that can [Means for solving the problem]
[0004] As described below, the present invention provides improved methods for minimizing off-target deamination. Nucleobase editors with tailored editing profiles and multi-effector nucleobase editors editors, compositions containing such editors, and methods for using the same to target nucleic acid sequences The method for producing the modification is featured.
[0005] In one aspect, the present invention provides: (i) polynucleotide-programmable DNA binding; (ii) a cytidine base editor comprising a cytidine deaminase domain; and (iii) a cytidine base editor comprising: Increased cis vs. trans activity (in cis:in tr) compared to standard cytidine base editors ans) are provided.
[0006] In some embodiments, the canonical cytidine base editor is (i) a polynucleotide (ii) a reprogrammable DNA-binding domain; and (iii) an APOBEC cytidine deaminase. In some embodiments, the canonical cytidine base editor APOBEC cytidine deaminase is: In one embodiment, the target is rat APOBEC-1 cytidine deaminase (rAPOBEC-1). Polynucleotide-programmable DNA binding domains of quasi-cytidine base editors In one embodiment, the canonical cytidine base editor is a Cas9 nickase. In one embodiment, the nucleotide sequence of ... The canonical cytidine base editor is BE3 or BE4. The ratio of cis to trans activity is at least 2, 2.5, 5, 10, 15, 20, 25, 30, 35, 40, 45 In one embodiment, the canonical cytidine base enzyme is increased by 50, 60, or more times. Compared to dieter, at least 50%, 60%, 70%, 80%, 90%, 95%, 100%, 105%, 110%, 11 5%, 120%, or more cis activity.
[0007] In some embodiments, the amino acid sequence of at least 2, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 39, 38, 39, 4 10, 15, 20, 25, 30, 35, 40, 45, 50, 60 or more times less transactivation do.
[0008] In one embodiment, the cytidine deaminase is selected from the group consisting of APOBEC1, APOBEC2, APOBEC3A, APOB EC3B, APOBEC3C, APOBEC3D, APOBEC3E, APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, activity Aminolysis-induced cytidine deaminase (AID), hAPOBEC1, rAPOBEC1, ppAPOBEC1, AmAPOBEC1 (BEM 3.31), ocAPOBEC1, SSAPOBEC2 (BEM3.39), hAPOBEC3A, maAPOBEC1, mdAPOBEC1, Cysteine deaminase 1 (CDA1), hA3A, RrA3F (BEM3.14), PmCDA1, AID (activation-induced kinase In one embodiment, the enzyme is selected from the group consisting of cysteine deaminase (AICDA), human AID, and FENRY. In one embodiment, the cytidine deaminase is APOBEC1. Deaminase from (a) Mesocricetus auratus (MaAPOBEC-1), Pongo pygmaeus (PpAPOBEC-1), from Oryctolagus cuniculus (OcAPOBEC-1), from Monodelphis domestica APOBEC-1 from the genus Alligator (MdAPOBEC-1) or Alligator mississippiensis (AmAPOBEC-1) (b) Pongo pygmaeus-derived (PpAPOBEC-2), Bos taurus-derived (BtAPOBEC-2), or Su (c) APOBEC-2 from Macaca scrofa (SsAPOBEC-2); (d) APOBEC-4 from Macaca fascicularis (MfA POBEC-4); (d) Canis lupus familiaris-derived (ClAID) or Bos taurus-derived (BtAID) AID from Saccharomyces cerevisiae; (e) yeast cytosine deaminase (yCD) from Saccharomyces cerevisiae; (f) Rh APOBEC-3F (RrA3F) from R. inopithecus roxellana; or (g) any one of (a) to (f). an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to It is a cytidine deaminase with the amino acid sequence
[0009] In one embodiment, the cytidine deaminase is from Mesocricetus auratus (MaAPOB EC-1), Pongo pygmaeus (PpAPOBEC-1), Oryctolagus cuniculus (OcAPOBEC-1) ), or APOBEC-1 from Monodelphis domestica (MdAPOBEC-1), or at least having an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the In one embodiment, the cytidine deaminase is rAPO. In one embodiment, the cytidine deaminase is hAPOBEC3A. In some embodiments, the cytidine deaminase is ppAPOBEC1. The cysteine deaminase was derived from Pongo pygmaeus (PpAPOBEC-2), Bos taurus (BtAPOBEC -2), or APOBEC-2 from Sus scrofa (SsAPOBEC-2), or at least 80% APOBEC-2 , 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical amino acid sequence. In one embodiment, the cytidine deaminase is derived from Macaca f ascicularis-derived APOBEC-4 (MfAPOBEC-4), or at least 80%, 85%, 90%, 95%, or cytidine deaminers having amino acid sequences that are 96%, 97%, 98%, or 99% identical to In one embodiment, the cytidine deaminase is derived from Canis lupus familiaris. AIDs derived from Bos taurus (ClAID) or Bos taurus (BtAID), or at least 80%, 85% or more of 90%, 95%, 96%, 97%, 98%, or 99% identical amino acid sequence to a cytidiol It is a deaminase.
[0010] In one embodiment, the cytidine deaminase is an enzyme derived from Saccharomyces cerevisiae. Mother cytosine deaminase (yCD), or at least 80%, 85%, 90%, 95%, 96%, 97% 98%, or 99% identical amino acid sequence to the cytidine deaminase In one embodiment, the cytidine deaminase is APO from Rhinopithecus roxellana. BEC-3F (RrA3F), or at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or is a cytidine deaminase having an amino acid sequence that is 99% identical to wherein the cytidine deaminase is any of the cytidine deaminases provided in Table 13. or at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical thereto In one embodiment, the cytidine deaminase has an amino acid sequence The deaminase was identified as APOBEC-3F (RrA3F) from Rhinopithecus roxellana, and that from Alligator missi. APOBEC-1 from Sus ssippiensis (AmAPOBEC-1), APOBEC-2 from Sus scrofa (SsAPOBEC-2) , or APOBEC-1 from Pongo pygmaeus (PpAPOBEC-1), or at least 80 85%, 90%, 95%, 96%, 97%, 98% or 99% identical amino acid sequence It is a gin deaminase.
[0011] In one embodiment, the cytidine deaminase is located at position 166666 according to the numbering in SEQ ID NO: 1. R15X, R16X, H21X, R30X, R33X, K34X, R52X, K60X, R118X, H121X, H122X, R126X, R128 One or more of X, R169X, R198X, T36X, H53X, V62X, L88X, W90X, Y120X or R132X or one or more corresponding modifications, where X is any amino acid.
[0012] In one embodiment, the cytidine deaminase is R15A according to the numbering in SEQ ID NO: 1. , R16A, H21A, R30A, R33A, K34A, R52A, K60A, R118A, H121A, H122A, H122L, R126A, R 128A, R169A, R198A, T36A, H53A, V62A, L88A, W90F, W90A, Y120F, Y120A, H121R, H12 2R, R126E, W90Y, and R132E, or one or more modifications thereof selected from the group consisting of In one embodiment, the cytidine deaminase comprises one or more modifications corresponding to SEQ ID NO: In the numbering in No. 1, K34A+R33A, K34A+H122A, K34A+Y120F, K34A+R52A, K34A+H122A, K34A+H121A, W90A+ R126E, W90Y+R126E, H121R+H122R, R126+R132E, W90Y+R132E, and W90Y+R126E+R132E or a combination of modifications corresponding thereto In one embodiment, the cytidine deaminase comprises a cytidine deaminase having the sequence numbering in SEQ ID NO: 1. A modification at position Y120F and a modification from the group consisting of R33A, W90F, K34A, R52A, H122A, and H121A. In some embodiments, the modified form includes one or more modifications selected from the following: wherein the cytidine deaminase is at position Y130X or R28X according to the numbering in SEQ ID NO: 1 or a corresponding modification thereof, where X is any amino acid.
[0013] 1, wherein the cytidine deaminase is at position Y130A or R28A in the numbering of SEQ ID NO: 1 In one embodiment, the cytidine deaminase The enzyme comprises modifications at positions Y130A and R28A in the numbering in SEQ ID NO: 1 or In one embodiment, the cytidine deaminase comprises a modification corresponding to SEQ ID NO: 1. one or more alterations at positions H122X, K34X, R33X, W90X, or R128X in the numbering or one or more corresponding modifications thereof, where X is any amino acid. In some embodiments, the cytidine deaminase has the sequence H122A, K34A, as numbered in SEQ ID NO: 1. , R33A, W90F, W90A, and R128A, or one or more modifications thereof In one embodiment, the cytidine deaminase comprises one or more corresponding modifications of the sequence Numbering in number 1: R33A+K34A, W90F+K34A, R33A+K34A+W90F, and R33A+K34A+H1 22A+W90F or a combination of modifications corresponding thereto .
[0014] In one embodiment, the cytidine deaminase has the following amino acid sequence: MTSEKGPSTGDPTLRRRIESWEFDVFYDPRELRKETCLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERRFHSSI SCSITWFLSWSPCWECSQAIREFLSQHPGVTLVIYVARLFWHMDQRNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYP PGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLAFFRLHLQNCHYQTIPPHILLATGLIHPSVTWR and an amino acid sequence having at least 80% identity with
[0015] In one embodiment, the cytidine deaminase has the following amino acid sequence: MKPQIRDHRPNPMEAMYPHIFYFHFENLEKAYGRNETWLCFTVEIIKQYLPVPWKKGVFRNQVDPETHCHAEKCFLSWFC NNTLSPKKNYQVTWYTSWSPCPECAGEVAEFLAEHSNVKLTIYTARLYYFWDTDYQEGLRSLSEEGASVEIMDYEDFQYC WENFVYDDGEPFKRWKGLKYNFQSLTRRLREILQ and an amino acid sequence having at least 80% identity with
[0016] In one embodiment, the cytidine deaminase has the following amino acid sequence: MADSSEKMRGQYISRDTFEKNYKPIDGTKEAHLLCEIKWGKYGKPWLHWCQNQRMNIHAEDYFMNNIFKAKKHPVHCYVT WYLSWSPCADCASKIVKFLEERPYLKLTIYVAQLYYHTEEENRKGLRLLRSKKVIIRVMDISDYNYCWKVFVSNQNGNED YWPLQFDPWVKENYSRLLDIFWESKCRSPNPW and an amino acid sequence having at least 80% identity with
[0017] In one embodiment, the cytidine deaminase has the following amino acid sequence: MDPQRLRQWPGPGPASRGGYGQRPRIRNPEEWFHELSPRTFSFHFRNLRFASGNRRSYICCQVEGKNCFFQGIFQNQVPP DPPCHAELCFLSWFQSWGLSPDEHYYVTWFISWSPCCECAAKVAQFLEENRNVSLSLSAARLYYFWKSESREGLRRLSDL GAQVGIMSFQDFQHCWNNFVHNLGMPFQPWKKLHKNYQRLVTELKQILREEPATYGSPQAQGKVRIGSTAAGLRHSHSHT RSEAHLRPNHSSRQHRILNPPREARARTCVLVDASWICYR and an amino acid sequence having at least 80% identity with
[0018] In one embodiment, the cytidine deaminase comprises an H122A modification. The cytidine base editor of any one of the above embodiments further comprises at least one adenosine base. In one embodiment, the adenovirus further comprises an aminase or a catalytically active fragment thereof. In one embodiment, the syndeaminase is TadA deaminase. is a modified adenosine deaminase that does not occur in nature. A cysteine base editor comprises two adenosine deaminases, which may be the same or different. In embodiments, the two adenosine deaminases form a heterodimer or a homodimer. In one embodiment, the adenosine deaminase domain is a wild-type TadA and TadA7.10.
[0019] In one embodiment, the adenosine deaminase is selected from the group consisting of 149, 150, 151, 152, 153, 154, and a C-terminal deletion beginning at a residue selected from the group consisting of 155, 156, and 157. In this embodiment, the adenosine deaminase has 1, 2, 3 amino acids with respect to the full-length adenosine deaminase. , 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acids In one embodiment, the adenosine deaminase is a full-length adenosine deaminase. Compared to aminase, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, In one embodiment, the at least one amino acid sequence lacks 18, 19, or 20 C-terminal amino acid residues. The nucleobase editor domain further comprises an abasic nucleobase editor. In some embodiments, the cytidine base editor of any one of the above aspects comprises one or more nucleic acids. In one embodiment, the cytidine base editor further comprises a NLS. In one embodiment, the NLS comprises a bipartite NLS and / or a C-terminal NLS. )NLS.
[0020] In one embodiment, the polynucleotide-programmable DNA binding domain is C In one embodiment, the polynucleotide-programmable DNA binding domain is as9. The main ones are Staphylococcus aureus Cas9 (SaCas9) and Streptococcus pyogenes Cas9 (SpC as9), or a variant thereof. In one embodiment, the polynucleotide A more programmable DNA-binding domain allows the nuclease-dead Cas9 (dCas) 9), Cas9 nickase (nCas9), or nuclease-active Cas9. wherein the polynucleotide-programmable DNA binding domain is a reverse complement of a nucleic acid sequence In one embodiment, the polynucleotide comprises a catalytic domain capable of chain cleavage. A programmable DNA binding domain is used to generate a catalytic domain capable of cleaving nucleic acid sequences. In one embodiment, the Cas9 is dCas9. The Cas9 is a Cas9 nickase (nCas9). In one embodiment, the nCas9 is It contains the modified D10A or its corresponding amino acid substitution.
[0021] In certain embodiments, the cytidine base editor of any one of the above aspects comprises one or more In one embodiment, the one or more The above UGIs are derived from the Bacillus subtilis bacteriophage PBS1 and are the human UD In one embodiment, two uracil DNA glycosylase inhibitors (UGIs) are The cytidine base editor of any one of the above embodiments may further comprise one or more linkers. This includes:
[0022] Provided herein is a cell comprising a cytidine base editor according to any one of the above embodiments. In one embodiment, the cell is a bacterial cell, a plant cell, an insect cell, or a mammalian cell. It is a cell.
[0023] A cytidine base editor according to any one of the above embodiments, a guide RNA sequence, and a tracrRNA sequence. or one or more of the target DNA sequences.
[0024] A method for editing nucleobases of a nucleic acid sequence, comprising: and contacting the first nucleobase of the DNA sequence with a cytidine base editor such as
[0023] A method is provided herein that includes converting
[0025] In one embodiment, the method further comprises: linking the nucleic acid sequence to a guide polypeptide to effect the conversion. In one embodiment, the first nucleobase is cytosine and said second nucleobase is thymidine.
[0026] In one embodiment, a polynucleotide-programmable DNA binding domain and a cytochrome P4500 protein are and at least one nucleobase editor domain comprising a nucleobase deaminase. (i) a protein derived from Mesocricetus auratus (MaAP) OBEC-1), derived from Pongo pygmaeus (PpAPOBEC-1), derived from Oryctolagus cuniculus (OcAPOBEC- 1), Monodelphis domestica (MdAPOBEC-1), or Alligator mississippiensis (ii) APOBEC-1 from Pongo pygmaeus (PpAPOBEC-2), Bos taurus (iii) APOBEC-2 derived from B. scrofa (BtAPOBEC-2) or Sus scrofa (SsAPOBEC-2); (iv) APOBEC-4 from Canis lupus familiaris (ClAID (v) AID from Saccharomyces cerevisiae (BtAID); tosine deaminase (yCD); (vi) APOBEC-3F (RrA3F) from Rhinopithecus roxellana or (vii) at least 80%, 85%, 90%, 95%, or 96% of any one of (i) through (viii). cytidine deaminase having an amino acid sequence that is 97%, 98%, or 99% identical to Fusion proteins are provided herein.
[0027] In one embodiment, a polynucleotide-programmable DNA binding domain and a cytochrome P4500 protein are and at least one nucleobase editor domain comprising a nucleobase deaminase. The cytidine deaminase is derived from Mesocricetus auratus (MaAPOBEC -1), derived from Pongo pygmaeus (PpAPOBEC-1), derived from Oryctolagus cuniculus (OcAPOBEC-1) or APOBEC-1 from Monodelphis domestica (MdAPOBEC-1), or an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to Provided herein is a fusion protein that is a cytidine deaminase having a sequence.
[0028] In one embodiment, a polynucleotide-programmable DNA binding domain and a cytochrome P4500 protein are and at least one nucleobase editor domain comprising a nucleobase deaminase. a protein, wherein the cytidine deaminase is derived from Pongo pygmaeus (PpAPOBEC-2); APOBEC-2 from Bos taurus (BtAPOBEC-2) or Sus scrofa (SsAPOBEC-2), or at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto The present invention provides a fusion protein that is a cytidine deaminase having an amino acid sequence of It is served.
[0029] In one embodiment, a polynucleotide-programmable DNA binding domain and a cytochrome P4500 protein are and at least one nucleobase editor domain comprising a nucleobase deaminase. The cytidine deaminase is an APOBEC-4 protein derived from Macaca fascicularis. (MfAPOBEC-4), or at least 80%, 85%, 90%, 95%, 96%, 97%, 98% thereto; or a cytidine deaminase having an amino acid sequence that is 99% identical to the fusion protein. Proteins are provided herein.
[0030] In one embodiment, a polynucleotide-programmable DNA binding domain and a cytochrome P4500 protein are and at least one nucleobase editor domain comprising a nucleobase deaminase. The cytidine deaminase is derived from Canis lupus familiaris (ClAID). , or AIDs derived from Bos Taurus (BtAID), or at least 80%, 85%, 90%, 9 Cytidine deaminers having an amino acid sequence that is 5%, 96%, 97%, 98%, or 99% identical to the amino acid sequence of the cytidine deaminers. Provided herein are fusion proteins that are enzymes.
[0031] In one embodiment, a polynucleotide-programmable DNA binding domain and a cytochrome P4500 protein are and at least one nucleobase editor domain comprising a nucleobase deaminase. The cytidine deaminase is a yeast protein derived from Saccharomyces cerevisiae. Cytosine deaminase (yCD), or at least 80%, 85%, 90%, 95%, 96%, 97% , 98%, or 99% identical amino acid sequence to the cytidine deaminase Fusion proteins are provided herein.
[0032] In one embodiment, a polynucleotide-programmable DNA binding domain and a cytochrome P4500 protein are and at least one nucleobase editor domain comprising a nucleobase deaminase. The cytidine deaminase is an APOBE protein derived from Rhinopithecus roxellana. C-3F (RrA3F), or at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% thereof % identity of the fusion protein to the cytidine deaminase Provided herein.
[0033] In one embodiment, a polynucleotide-programmable DNA binding domain and a cytochrome P4500 protein are and at least one nucleobase editor domain comprising a nucleobase deaminase. wherein the cytidine deaminase is a cytidine deaminase provided in Table 13. or at least 80%, 85%, 90%, 95%, 96%, 97%, or 98% of any one of the above. or a cytidine deaminase having an amino acid sequence that is 99% identical to the fusion protein. Proteins are provided herein.
[0034] In one embodiment, a polynucleotide-programmable DNA binding domain and a cytochrome P4500 protein are and at least one nucleobase editor domain comprising a nucleobase deaminase. The cytidine deaminase is an APOBE protein derived from Rhinopithecus roxellana. C-3F (RrA3F), APOBEC-1 from Alligator mississippiensis (AmAPOBEC-1), Sus scrof APOBEC-2 from S. a (SsAPOBEC-2), APOBEC-1 from Pongo pygmaeus (PpAPOBEC-1), and is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to it. Provided herein is a fusion protein that is a cytidine deaminase having an amino acid sequence do.
[0035] In one embodiment, the cytidine deaminase is located at position 166666 according to the numbering in SEQ ID NO: 1. R15X, R16X, H21X, R30X, R33X, K34X, R52X, K60X, R118X, H121X, H122X, R126X, R128 One or more of X, R169X, R198X, T36X, H53X, V62X, L88X, W90X, Y120X or R132X or one or more corresponding modifications, where X is any amino acid. In one embodiment, the cytidine deaminase is selected from the group consisting of R15A, R15B, R15C, R15D, R15E, R15F, R15G, R15H, R15I ... 6A, H21A, R30A, R33A, K34A, R52A, K60A, R118A, H121A, H122A, H122L, R126A, R128A , R169A, R198A, T36A, H53A, V62A, L88A, W90F, W90A, Y120F, Y120A, H121R, H122R, one or more modifications selected from the group consisting of R126E, W90Y, and R132E or corresponding modifications thereof In one embodiment, the cytidine deaminase comprises one or more modifications selected from the group consisting of SEQ ID NO:1 In the numbering, K34A+R33A, K34A+H122A, K34A+Y120F, K34A+R52A, K34A+H122A, K34A +H121A, W90A+ R126E, W90Y+R126E, H121R+H122R, R126+R132E, W90Y+R132E, and W90Y +R126E+R132E or a combination of modifications corresponding thereto In one embodiment, the cytidine deaminase is Y1 20F, and R33A, W90F, K34A, R52A, H122A, and H121A. It includes one or more modifications or a combination of one or more corresponding modifications.
[0036] In one embodiment, the cytidine deaminase is located at position 166666 according to the numbering in SEQ ID NO: 1. containing one or more modifications in Y130X or R28X or one or more modifications corresponding thereto, wherein X is any amino acid. In one embodiment, the cytidine deaminase is One or more modifications selected from the group consisting of Y130A and R28A in the numbering in column number 1 or one or more corresponding modifications thereof. The enzyme has the modifications Y130A and R28A, or the corresponding modifications according to the numbering in SEQ ID NO: 1. In one embodiment, the cytidine deaminase comprises a modification of the amino acid sequence of SEQ ID NO: 1. one or more alterations at positions H122X, K34X, R33X, W90X, or R128X in the In one embodiment, X is any amino acid. wherein the cytidine deaminase is H122A, K34A, R33A, W, as numbered in SEQ ID NO: 1 90F, W90A, and R128A, or one or more modifications corresponding thereto It includes one or more modifications that:
[0037] In one embodiment, the cytidine deaminase has R33A as numbered in SEQ ID NO: 1. +K34A, W90F+K34A, R33A+K34A+W90F, and R33A+K34A+H122A+W90F. In some embodiments, the above-described modifications may be combined with one or more corresponding modifications. The cytidine deaminase has an H122A modification in the numbering of SEQ ID NO: 1, or a corresponding modification thereto. In one embodiment, the cytidine deaminase is rAPOBEC1 and Numbering in column 1: R15A, R16A, H21A, R30A, R33A, K34A, R52A, K60A, R118A , H121A, H122A, H122L, R126A, R128A, R169A, R198A, T36A, H53A, V62A, L88A, W90F , W90A, Y120F, Y120A, H121R, H122R, R126E, W90Y, and R132E. In one embodiment, the above-described The cytidine deaminase may comprise, in accordance with the numbering in SEQ ID NO: 1, K34A+R33A, K34A+H122A, K34A +Y120F, K34A+R52A, K34A+H122A, K34A+H121A, W90A+ R126E, W90Y+R126E, H121R+H122R , R126+R132E, W90Y+R132E, and W90Y+R126E+R132E, or or a combination of one or more corresponding modifications thereof.
[0038] In one embodiment, a polynucleotide-programmable DNA binding domain and an APOBE C2 family members, APOBEC3 family members, APOBEC4 family members, Citi CDA1 family members, A3A family members, and RrA3F family members A group consisting of Lee members, PmCDA1 family members, and FENRY family members at least one nucleobase editor domain comprising a cytidine deaminase selected from Provided herein are fusion proteins comprising:
[0039] In one embodiment, the APOBEC3 family member is APOBEC3A, APOBEC3B, APOBEC3 C, APOBEC3D, APOBEC3E, APOBEC3F, APOBEC3G, and APOBEC3H. In one embodiment, the APOBEC2 family member is SsAPOBEC2.
[0040] Polynucleotide-programmable DNA-binding domains and ppAPOBEC1 and AmAPOBEC1 (BEM3.31), ocAPOBEC1, SsAPOBEC2 (BEM3.39), hAPOBEC3A, maAPOBEC1, and mdAPOBE C1. At least one nucleobase editor domain comprising an APOBEC1 selected from the group consisting of: and
[0041] In one embodiment, the cytidine deaminase is at position R in the numbering of SEQ ID NO:1. 15X, R16X, H21X, R30X, R33X, K34X, R52X, K60X, R118X, H121X, H122X, R126X, R128X , R169X, R198X, T36X, H53X, V62X, L88X, W90X, Y120X or R132X or one or more corresponding modifications thereof, where X is any amino acid In one embodiment, the one or more modifications are R15A, R16A, R17A, R18A, R19A, R20A, R21A, R22A, R23A, R24A, R25A, R26A, R27A, R28A, R29A, R30A, R31A, R32A, R33A, R34A, R35A, R36A, R37A, R38A, R39A, R40A, R41A, R42A, R43A, R44A, R45A, R46A, R47A, R48A, R49A, R50A, R51A, R52A, R53A, R54A, R55A, R56A, R57A, R58A, R59A, R60A, R61A, R62A, R63A, , H21A, R30A, R33A, K34A, R52A, K60A, R118A, H121A, H122A, H122L, R126A, R128A, R169A, R198A, T36A, H53A, V62A, L88A, W90F, W90A, Y120F, Y120A, H121R, H122R, R1 26E, W90Y, and R132E, or one or more modifications corresponding thereto. In one embodiment, the cytidine deaminase is located at K34 in the numbering of SEQ ID NO: 1. A+R33A, K34A+H122A, K34A+Y120F, K34A+R52A, K34A+H122A, K34A+H121A, W90A+ R126E, Consists of W90Y+R126E, H121R+H122R, R126+R132E, W90Y+R132E, and W90Y+R126E+R132E The present invention includes a combination of one or more modifications selected from the group consisting of: In embodiments, the cytidine deaminase has a sequence similar to that of SEQ ID NO: 1, Y120F, and Y120G. one or more alterations selected from the group consisting of R33A, W90F, K34A, R52A, H122A, and H121A or a combination of one or more corresponding modifications thereof.
[0042] In one embodiment, a polynucleotide-programmable DNA binding domain and a cytochrome P4500 protein are and at least one nucleobase editor domain comprising a nucleobase deaminase. wherein the cytidine deaminase is located at position R15 in the numbering in SEQ ID NO: 1. X, R16X, H21X, R30X, R33X, K34X, R52X, K60X, R118X, H121X, H122X, R126X, R128X, One or more of R169X, R198X, T36X, H53X, V62X, L88X, W90X, Y120X or R132X or one or more corresponding modifications, wherein X is any amino acid. Synthetic proteins are provided herein.
[0043] In one embodiment, the cytidine deaminase is R15A according to the numbering in SEQ ID NO: 1. , R16A, H21A, R30A, R33A, K34A, R52A, K60A, R118A, H121A, H122A, H122L, R126A, R 128A, R169A, R198A, T36A, H53A, V62A, L88A, W90F, W90A, Y120F, Y120A, H121R, H12 2R, R126E, W90Y, and R132E, or one or more modifications thereof selected from the group consisting of In one embodiment, the cytidine deaminase comprises one or more modifications corresponding to SEQ ID NO: Numbering in No. 1: K34A+R33A, K34A+H122A, K34A+Y120F, K34A+R52A, K34A+H122A, K 34A+H121A, W90A+ R126E, W90Y+R126E, H121R+H122R, R126+R132E, W90Y+R132E, and W 90Y+R126E+R132E or one or more modifications corresponding thereto In one embodiment, the cytidine deaminase comprises a combination of the cytidine deaminase shown in SEQ ID NO: 1. The numbering indicates the alterations at position Y120F and R33A, W90F, K34A, R52A, H122A, and H121A. A) or one or more modifications corresponding thereto. nothing.
[0044] In one embodiment, a polynucleotide-programmable DNA binding domain and a cytochrome P4500 protein are and at least one nucleobase editor domain comprising a nucleobase deaminase. 1, wherein the cytidine deaminase is located at position Y13 in the numbering of SEQ ID NO: 1. 0X and R28X, or one or more corresponding modifications, Fusion proteins are provided herein, wherein X is any amino acid.
[0045] In one embodiment, the cytidine deaminase has the sequence Y130A as numbered in SEQ ID NO: 1. and R28A, or one or more modifications corresponding thereto. In one embodiment, the cytidine deaminase comprises the modifications Y130A and R28A. .
[0046] In one embodiment, a polynucleotide-programmable DNA binding domain and a cytochrome P4500 protein are and at least one nucleobase editor domain comprising a nucleobase deaminase. a protein in which the cytidine deaminase is located at position H12 in the numbering in SEQ ID NO: 1; 2X, K34X, R33X, W90X, or R128X, or the corresponding 1 fusion proteins are provided herein that contain one or more modifications, where X is any amino acid. It is served.
[0047] In one embodiment, the cytidine deaminase has the sequence H122A as numbered in SEQ ID NO: 1. , K34A, R33A, W90F, W90A, and R128A, or In one embodiment, the cytidine deaminase comprises one or more corresponding modifications. , R33A+K34A, W90F+K34A, R33A+K34A+W90F, and R33A+K3 4A+H122A+W90F, or one or more alterations corresponding thereto In one embodiment, the cytidine deaminase is selected from the group consisting of APOBEC1, APOBEC2, APOBEC3, APOBEC4, APOBEC5, APOBEC6, APOBEC7, APOBEC8, APOBEC9, APOBEC10, APOBEC11, APOBEC12, APOBEC13, APOBEC14, APOBEC15, APOBEC16, APOBEC17, APOBEC18, APOBEC19, APOB OBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3E, APOBEC3F, APOBEC3G, APOBEC3H, AP OBEC4, activation-induced (cytidine) deaminase (AID), hAPOBEC1, rAPOBEC1, ppAPOBEC1 , AmAPOBEC1 (BEM3.31), ocAPOBEC1, ssAPOBEC2 (BEM3.39), hAPOBEC3A, maAPOBEC1, m dAPOBEC1, cytidine deaminase 1 (CDA1), hA3A, RrA3F (BEM3.14), PmCDA1, AID (activation and a cytidine deaminase (AICDA), human AID, and FENRY. In one embodiment, the cytidine deaminase is APOBEC1. In one embodiment, the cytidine deaminase is rAPOBEC1. In one embodiment, the cytidine deaminase is hAPOBEC3A. be.
[0048] In one embodiment, a polynucleotide programmable DNA binding domain and a cytidine deoxyribonucleotide and a cytidine deaminase, wherein the cytidine deaminase is Acid sequence: MTSEKGPSTGDPTLRRRIESWEFDVFYDPRELRKETCLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERRFHSSI SCSITWFLSWSPCWECSQAIREFLSQHPGVTLVIYVARLFWHMDQRNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYP PGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLAFFRLHLQNCHYQTIPPHILLATGLIHPSVTWR A fusion protein comprising an amino acid sequence having at least 80% identity to the provided in writing.
[0049] In one embodiment, a polynucleotide programmable DNA binding domain and a cytidine deoxyribonucleotide and a cytidine deaminase, wherein the cytidine deaminase is Acid sequence: MKPQIRDHRPNPMEAMYPHIFYFHFENLEKAYGRNETWLCFTVEIIKQYLPVPWKKGVFRNQVDPETHCHAEKCFLSWFC NNTLSPKKNYQVTWYTSWSPCPECAGEVAEFLAEHSNVKLTIYTARLYYFWDTDYQEGLRSLSEEGASVEIMDYEDFQYC WENFVYDDGEPFKRWKGLKYNFQSLTRRLREILQ A fusion protein comprising an amino acid sequence having at least 80% identity to the provided in writing.
[0050] In one embodiment, a polynucleotide programmable DNA binding domain and a cytidine deoxyribonucleotide and a cytidine deaminase, wherein the cytidine deaminase is Acid sequence: MADSSEKMRGQYISRDTFEKNYKPIDGTKEAHLLCEIKWGKYGKPWLHWCQNQRMNIHAEDYFMNNIFKAKKHPVHCYVT WYLSWSPCADCASKIVKFLEERPYLKLTIYVAQLYYHTEEENRKGLRLLRSKKVIIRVMDISDYNYCWKVFVSNQNGNED YWPLQFDPWVKENYSRLLDIFWESKCRSPNPW A fusion protein comprising an amino acid sequence having at least 80% identity to the provided in writing.
[0051] In one embodiment, a polynucleotide programmable DNA binding domain and a cytidine deoxyribonucleotide and a cytidine deaminase, wherein the cytidine deaminase is Acid sequence: MDPQRLRQWPGPGPASRGGYGQRPRIRNPEEWFHELSPRTFSFHFRNLRFASGNRRSYICCQVEGKNCFFQGIFQNQVPP DPPCHAELCFLSWFQSWGLSPDEHYYVTWFISWSPCCECAAKVAQFLEENRNVSLSLSAARLYYFWKSESREGLRRLSDL GAQVGIMSFQDFQHCWNNFVHNLGMPFQPWKKLHKNYQRLVTELKQILREEPATYGSPQAQGKVRIGSTAAGLRHSHSHT RSEAHLRPNHSSRQHRILNPPREARARTCVLVDASWICYR A fusion protein comprising an amino acid sequence having at least 80% identity to the provided in writing.
[0052] In one embodiment, the cytidine deaminase comprises an H122A mutation.
[0053] In one embodiment, a polynucleotide programmable DNA binding domain and a cytidine deoxyribonucleotide and a cytidine deaminase, wherein the cytidine deaminase is an APOBEC1 deaminase. Provided herein is a fusion protein that is a cytosine aminotransferase and includes the H122A modification.
[0054] In one embodiment, a polynucleotide-programmable DNA binding domain and a cytidine and a cytidine deaminase, wherein the cytidine deaminase is rAPOBE C1, R15A, R16A, H21A, R30A, R33A, K34A, R52A, K60A, R118A, H121A, H122A, H 122L, R126A, R128A, R169A, R198A, T36A, H53A, V62A, L88A, W90F, W90A, Y120F, Y12 one or more modifications selected from the group consisting of 0A, H121R, H122R, R126E, W90Y, and R132E In one embodiment, the cytidine deoxyribonucleotide is The amino acid sequence of the enzymes is K34A+R33A, K34A+H122A, K34A+Y120F, K34A+R52A, K34A+H122A, and K34A+H121A. , W90A+ R126E, W90Y+R126E, H121R+H122R, R126+R132E, W90Y+R132E, and W90Y+R126E +R132E +R132E.
[0055] In one embodiment, a polynucleotide-programmable DNA binding domain and ppAPO BEC1, AmAPOBEC1 (BEM3.31), ocAPOBEC1, SsAPOBEC2 (BEM3.39), hAPOBEC3A, maAPOBEC 1, and mdAPOBEC1, Provided herein are fusion proteins comprising a nucleotide sequence and a nucleotide sequence that encodes ...
[0056] In one embodiment, the APOBEC1 has a sequence identical to that of positions R15X, R16X, H2 1X, R30X, R33X, K34X, R52X, K60X, R118X, H121X, H122X, R126X, R128X, R169X, R198 one or more modifications in X, T36X, H53X, V62X, L88X, W90X, Y120X or R132X; or and one or more corresponding modifications, where X is any amino acid.
[0057] In one embodiment, the one or more modifications are R15A, R16A, R17A, R18A, R19A, R20A, R21A, R22A, R23A, R24A, R25A, R26A, R27A, R28A, R29A, R30A, R31A, R32A, R33A, R34A, R35A, R36A, R37A, R38A, R39A, R40A, R41A, R42A, R43A, R44A, R45A, R46A, R47A, R48A, R49A, R50A, R51A, R52A, R53A, R54A, R55A, R56A, R57A, R58A, R59A, R60A, R61A, R62A, R63A, , H21A, R30A, R33A, K34A, R52A, K60A, R118A, H121A, H122A, H122L, R126A, R128A, R169A, R198A, T36A, H53A, V62A, L88A, W90F, W90A, Y120F, Y120A, H121R, H122R, R1 26E, W90Y, and R132E, or one or more corresponding thereto. In one embodiment, the APOBEC1 has the following modifications: K34A+R33 according to the numbering in SEQ ID NO: 1. A, K34A+H122A, K34A+Y120F, K34A+R52A, K34A+H122A, K34A+H121A, W90A+ R126E, W90Y+ From the group consisting of R126E, H121R+H122R, R126+R132E, W90Y+R132E, and W90Y+R126E+R132E In some embodiments, the selected modification or a combination of one or more modifications corresponding thereto. wherein the APOBEC1 has a modification at Y120F, and a modification at R33A, according to the numbering in SEQ ID NO: 1. , W90F, K34A, R52A, H122A, and H121A, or one or more modifications selected from the group consisting of includes one or more corresponding modifications.
[0058] In one embodiment, the fusion protein of any one of the above aspects comprises at least one amino acid sequence. In one embodiment, the enzyme further comprises adenosine deaminase or a catalytically active fragment thereof. The adenosine deaminase is TadA deaminase. In some embodiments, the aminase is a non-naturally occurring modified adenosine deaminase. In some embodiments, the fusion protein contains two adenosine deaminases, which may be the same or different. In embodiments, the two adenosine deaminases form a heterodimer or a homodimer. In one embodiment, the two adenosine deaminase domains are wild-type TadA and TadA7.10.
[0059] In one embodiment, the adenosine deaminase is selected from the group consisting of 149, 150, 151, 152, 153, 154, and a C-terminal deletion beginning at a residue selected from the group consisting of 155, 156, and 157. In one embodiment, the adenosine deaminase has 1, 2, or 3 amino acids longer than the full-length adenosine deaminase. 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19 or 20 N-terminal amino acids In one embodiment, the adenosine deaminase is a full-length adenosine deaminase. Compared to aminase, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, In one embodiment, the at least one C-terminal amino acid residue is missing. The nucleobase editor domain of further comprises an abasic nucleobase editor.
[0060] In one embodiment, the fusion protein of any one of the above aspects comprises one or more nuclear localization sequences. In one embodiment, the fusion protein further comprises an N-terminal NLS and / or or a C-terminal NLS. In one embodiment, the NLS is a bipartite NLS.
[0061] In one embodiment, the polynucleotide-programmable DNA binding domain is C In one embodiment, the polynucleotide-programmable DNA binding domain is as9. The main ones are Staphylococcus aureus Cas9 (SaCas9) and Streptococcus pyogenes Cas9 (SpC as9), or a variant thereof. Reprogrammable DNA-binding domains are used to engineer nuclease-inactive Cas9 (dCas9), Cas9 nickers, and In one embodiment, the polynucleoside comprises a nuclease (nCas9), or a nuclease-active Cas9. The DNA-binding domain, programmable by a peptide, can cleave the reverse complementary strand of a nucleic acid sequence. In one embodiment, the polynucleotide comprises a catalytic domain capable of programming the Possible DNA binding domains do not include catalytic domains capable of cleaving nucleic acid sequences.
[0062] In one embodiment, the Cas9 is dCas9. In one embodiment, the Cas9 is Cas9 nicotinamide adenine dinucleotide 1 (CdS1)-1 (CdS2)-1 (CdS3)-1 (CdS4)-1 (CdS5)-1 (CdS6)-1 (CdS7)-1 (CdS8)-1 (CdS9 ... In one embodiment, the nCas9 is a nucleotide sequence encoding ... In one embodiment, one or more uracil DNA glycosylators In one embodiment, the one or more UGIs further comprise a Bacillus subtilis enzyme inhibitor (UGI). The antibody is derived from the B. ubtilis bacteriophage PBS1 and inhibits human UDG activity. The fusion protein contains two uracil DNA glycosylase inhibitors (UGIs). In one embodiment, the fusion protein further comprises one or more linkers. The conjugated protein deaminates a nucleic acid base in a target nucleotide sequence, and said deamination , increased cis-to-trans activity (in cis:i) compared to standard cytidine base editors n trans).
[0063] In some embodiments, the canonical cytidine base editor (i) adds a cytidine base to a polynucleotide containing a more programmable DNA-binding domain and (ii) an APOBEC cytidine deaminase .
[0064] In one embodiment, the canonical cytidine base editor is an APOBEC cytidine deaminase. is rat APOBEC-1 cytidine deaminase (rAPOBEC-1). Polynucleotide-programmable DNA binding of canonical cytidine base editors In one embodiment, the canonical cytidine base editor domain is a Cas9 nickase. In one embodiment, the nucleotide sequence comprises a uracil glycosylase inhibitor (UGI) domain. The canonical cytidine base editor is BE3 or BE4. The ratio of cis to trans activity measured is at least 2, 2.5, 5, 10, 15, 20, 25, 30, 35, 40, 4 In one embodiment, the cytidine base edited fragment is increased by 5, 50, 60 or more times. the cytidine base editor has a cytidine base sequence that is at least 50%, 60%, 70%, 80% higher than the standard cytidine base editor. , 90%, 95%, 100%, 105%, 110%, 115%, 120% or more cis activity. In embodiments, the cytidine base editor is At least 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60 or more times smaller than It has in trans activity.
[0065] In one embodiment, a polynucleotide encoding the fusion protein of any one of the above embodiments is provided. In one embodiment, the polynucleotide comprises a codon It is optimized.
[0066] In one aspect, provided herein is an expression vector comprising the polynucleotide molecule described above. In one embodiment, the expression vector is a mammalian expression vector. In some embodiments, the vector is an adeno-associated virus (AAV), a retroviral vector, an adenovirus vector, or a recombinant vector. Noviral vectors, lentiviral vectors, Sendai virus vectors, and hepatitis C virus vectors In one embodiment, the viral vector is selected from the group consisting of pesvirus vectors. In the above, the vector comprises a promoter.
[0067] In one embodiment, a cell comprising the polynucleotide or the vector is described herein. In one embodiment, the cell is a bacterial cell, a plant cell, an insect cell, or a human cell. , or mammalian cells.
[0068] In one embodiment, the fusion protein of any one of the above embodiments is a fusion protein comprising a guide RNA sequence, a tracr Provided herein are molecular complexes comprising one or more of a target RNA sequence, a target DNA sequence, or a target DNA sequence. can be.
[0069] In one embodiment, the fusion protein of any one of the above embodiments, the polynucleotide, Provided herein is a kit comprising the vector or the molecular complex.
[0070] In one embodiment, there is provided a method for editing the nucleobases of a nucleic acid sequence, comprising: a base editor comprising the fusion protein of any one of embodiments, Provided herein are methods that include converting the base to a second nucleobase. In one embodiment, the first nucleobase is cytosine and the second nucleobase is thymidine.
[0071] In one embodiment, there is provided a method for editing the nucleobases of a nucleic acid sequence, comprising: a base editor comprising the fusion protein of any one of embodiments, Provided herein are methods that include converting the base to a second nucleobase. In one embodiment, the first nucleobase is cytosine and the second nucleobase is thymidine, or Alternatively, the first nucleobase is adenine and the second nucleobase is guanine. In an embodiment, the method further comprises converting the third nucleobase to a fourth nucleobase. In one embodiment, the third nucleobase is guanine and the fourth nucleobase is adenine. or said third nucleobase is thymine and said fourth nucleobase is cytosine; do.
[0072] In one embodiment, there is provided a method for optimized base editing, comprising: (i) a polynucleotide-programmable DNA-binding domain; (ii) contacting the nucleic acid sequence with a cytidine base editor comprising a cytidine deaminase; wherein the cytidine base editor is a canonical cytidine base editor including rAPOBEC1. nucleotides with lower unintended deamination in the target nucleobase sequence compared to In one embodiment, the cytidine base is deaminated. the editor selectively targets the target nucleic acid with greater efficiency than the canonical cytidine base editor. In one embodiment, the canonical cytidine base editor is In one embodiment, the polypeptide further comprises a guanine glycosylase inhibitor (UGI) domain. The canonical cytidine base editor is BE3 or BE4. In one embodiment, the cytidine The base editor is a cis / trans deamination assay. At least 20%, 30%, 50%, 70%, or 90% lower objective compared to the cytogenetic base editor In some embodiments, the cytidine base editor generates an external deamination. at least 50%, 60%, 70%, 80%, 90%, 95%, or 100% higher than a standard cytidine base editor , 105%, 110%, 115%, 120% or more cis activity. the cytidine base editor has, relative to the canonical cytidine base editor, at least 2x, 5x, 10x, 15x, 20x, 25x, 30x, 35x, 40x, 45x, 50x, 60x or more In one embodiment, the cytidine deaminase has low in trans activity. socricetus auratus (MaAPOBEC-1), Pongo pygmaeus (PpAPOBEC-1), Oryctola gus cuniculus (OcAPOBEC-1), Monodelphis domestica (MdAPOBEC-1), or (a) APOBEC-1 from Alligator mississippiensis (AmAPOBEC-1); (b) APOBEC-1 from Pongo pygmaeus (PpAPOBEC-2), derived from Bos taurus (BtAPOBEC-2), or derived from Sus scrofa (SsAPOBEC-2 ) APOBEC-2; (c) APOBEC-4 from Macaca fascicularis (MfAPOBEC-4); (d) Canis lup AIDs derived from Bos familiaris (ClAID) or Bos Taurus (BtAID); (e) Saccharomyces Yeast cytosine deaminase (yCD) from Saccharomyces cerevisiae; (f) from Rhinopithecus roxellana or (g) at least 80%, 8% or more of any one of (a) to (f). cytidines with an amino acid sequence that is 5%, 90%, 95%, 96%, 97%, 98% or 99% identical It is a deaminase.
[0073] In one embodiment, the cytidine deaminase is derived from Canis lupus familiaris (ClAID ), or AIDs derived from Bos Taurus (BtAID), or at least 80%, 85%, 90%, 9 Cytidine deaminers having amino acid sequences that are 5%, 96%, 97%, 98%, or 99% identical In one embodiment, the cytidine deaminase is derived from Rhinopithecus roxellana APOBEC-3F (RrA3F) derived from, or at least 80%, 85%, 90%, 95%, 96%, 97%, 98% %, or 99% identical amino acid sequence to the cytidine deaminase.
[0074] In one embodiment, the cytidine deaminase has the structure R15X according to the numbering in SEQ ID NO: 1. , R16X, H21X, R30X, R33X, K34X, R52X, K60X, R118X, H121X, H122X, R126X, R128X, R Selected from the group consisting of 169X, R198X, T36X, H53X, V62X, L88X, W90X, Y120X, and R132X or a corresponding modification, where X is any amino acid. In embodiments, the cytidine deaminase is selected from the group consisting of R15A, R16A, R17A, R18A, R19A, R20A, R21A, R22A, R23A, R24A, R25A, R26A, R27A, R28A, R29A, R30A, R31A, R32A, R33A, R34A, R35A, R36 , H21A, R30A, R33A, K34A, R52A, K60A, R118A, H121A, H122A, H122L, R126A, R128A, R169A, R198A, T36A, H53A, V62A, L88A, W90F, W90A, Y120F, Y120A, H121R, H122R, R1 26E, W90Y, and R132E, or a corresponding modification thereof. In one embodiment, the cytidine deaminase has a sequence similar to K34A in the numbering of SEQ ID NO: 1. +R33A, K34A+H122A, K34A+Y120F, K34A+R52A, K34A+H122A, K34A+H121A, W90A+ R126E, W The group consisting of 90Y+R126E, H121R+H122R, R126+R132E, W90Y+R132E, and W90Y+R126E+R132E or a combination of modifications corresponding thereto.
[0075] In one embodiment, the cytidine deaminase is at position 1444, as numbered in SEQ ID NO: 1. A modification at position Y120F and from the group consisting of R33A, W90F, K34A, R52A, H122A, and H121A. In some embodiments, the selected one or more modifications may be included, or one or more modifications corresponding thereto. wherein the cytidine deaminase is at position Y130X or R28X in the numbering in SEQ ID NO: 1 or a corresponding modification in which X is any amino acid. In embodiments, the cytidine deaminase has a Y130A mutation as numbered in SEQ ID NO: 1. or R28A mutation or a corresponding mutation. The deaminase has the modifications Y130A and R28A according to the numbering in SEQ ID NO: 1, or corresponding thereto This includes modifications that:
[0076] In one embodiment, the cytidine deaminase is at position H in the numbering of SEQ ID NO:1. 122X, K34X, R33X, W90X, and R128X or the corresponding modifications; wherein X is any amino acid. In one embodiment, the cytidine deaminase is From the group consisting of H122A, K34A, R33A, W90F, W90A, and R128A in the numbering in column 1 In one embodiment, the cytidine derivative comprises a modification selected from the group consisting of: The aminase may have the following sequences according to the numbering in SEQ ID NO: 1: R33A+K34A, W90F+K34A, R33A+K34A+W90F, and R33A+K34A+H122A+W90F, or a combination of alterations selected from the group consisting of This includes combinations of modifications corresponding to:
[0077] In one embodiment, the cytidine deaminase has the following amino acid sequence: MTSEKGPSTGDPTLRRRIESWEFDVFYDPRELRKETCLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERRFHSSI SCSITWFLSWSPCWECSQAIREFLSQHPGVTLVIYVARLFWHMDQRNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYP PGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLAFFRLHLQNCHYQTIPPHILLATGLIHPSVTWR and an amino acid sequence having at least 80% identity with
[0078] In one embodiment, the cytidine deaminase has the following amino acid sequence: MKPQIRDHRPNPMEAMYPHIFYFHFENLEKAYGRNETWLCFTVEIIKQYLPVPWKKGVFRNQVDPETHCHAEKCFLSWFC NNTLSPKKNYQVTWYTSWSPCPECAGEVAEFLAEHSNVKLTIYTARLYYFWDTDYQEGLRSLSEEGASVEIMDYEDFQYC WENFVYDDGEPFKRWKGLKYNFQSLTRRLREILQ and an amino acid sequence having at least 80% identity with
[0079] In one embodiment, the cytidine deaminase has the following amino acid sequence: MADSSEKMRGQYISRDTFEKNYKPIDGTKEAHLLCEIKWGKYGKPWLHWCQNQRMNIHAEDYFMNNIFKAKKHPVHCYVT WYLSWSPCADCASKIVKFLEERPYLKLTIYVAQLYYHTEEENRKGLRLLRSKKVIIRVMDISDYNYCWKVFVSNQNGNED YWPLQFDPWVKENYSRLLDIFWESKCRSPNPW and an amino acid sequence having at least 80% identity with
[0080] In one embodiment, the cytidine deaminase has the following amino acid sequence: MDPQRLRQWPGPGPASRGGYGQRPRIRNPEEWFHELSPRTFSFHFRNLRFASGNRRSYICCQVEGKNCFFQGIFQNQVPP DPPCHAELCFLSWFQSWGLSPDEHYYVTWFISWSPCCECAAKVAQFLEENRNVSLSLSAARLYYFWKSESREGLRRLSDL GAQVGIMSFQDFQHCWNNFVHNLGMPFQPWKKLHKNYQRLVTELKQILREEPATYGSPQAQGKVRIGSTAAGLRHSHSHT RSEAHLRPNHSSRQHRILNPPREARARTCVLVDASWICYR and an amino acid sequence having at least 80% identity with
[0081] In one embodiment, the cytidine deaminase comprises an H122A modification. In one embodiment, the contacting is performed in a cell. In one embodiment, the cell is a human cell or a mammalian cell. In one embodiment, the contacting is in vivo or ex vivo.
[0082] In one embodiment, a sequence having at least 80% identity to an amino acid sequence selected from the following: Provided herein is a cytidine deaminase comprising an amino acid sequence having: MTSEKGPSTGDPTLRRRIESWEFDVFYDPRELRKETCLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERRFHS SISCSITWFLSWSPCWECSQAIREFLSQHPGVTLVIYVARLFWHMDQRNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVN YPPGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLAFFRLHLQNCHYQTIPPHILLATGLIHPSVTWR; MKPQIRDHRPNPMEAMYPHIFYFHFENLEKAYGRNETWLCFTVEIIKQYLPVPWKKGVFRNQVDPETHCHAEKCFLSW FCNNTLSPKKNYQVTWYTSWSPCPECAGEVAEFLAEHSNVKLTIYTARLYYFWDTDYQEGLRSLSEEGASVEIMDYEDFQ YCWENFVYDDGEPFKRWKGLKYNFQSLTRRLREILQ; MADSSEKMRGQYISRDTFEKNYKPIDGTKEAHLLCEIKWGKYGKPWLHWCQNQRMNIHAEDYFMNNIFKAKKHPVHCY VTWYLSWSPCADCASKIVKFLEERPYLKLTIYVAQLYYHTEEENRKGLRLLRSKKVIIRVMDISDYNYCWKVFVSNQNGN EDYWPLQFDPWVKENYSRLLDIFWESKCRSPNPW; and MDPQRLRQWPGPGPASRGGYGQRPRIRNPEEWFHELSPRTFSFHFRNLRFASGNRRSYICCQVEGKNCFFQGIFQNQV PPDPPCHAELCFLSWFQSWGLSPDEHYYVTWFISWSPCCECAAKVAQFLEENRNVSLSLSAARLYYFWKSESREGLRRLS DLGAQVGIMSFQDFQHCWNNFVHNLGMPFQPWKKLHKNYQRLVTELKQILREEPATYGSPQAQGKVRIGSTAAGLRHSHS HTRSEAHLRPNHSSRQHRILNPPREARARTCVLVDASWICYR.
[0083] The description and examples herein set forth in detail embodiments of the present disclosure. It should be understood that the present invention is not limited to the particular embodiments described, as such may vary. Those skilled in the art will recognize that there are numerous variations and modifications that fall within the scope of this disclosure. You will recognize it.
[0084] The practice of some embodiments disclosed herein may be carried out in any suitable manner, including, but not limited to, immunology, Conventional techniques in biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA techniques are used, which are within the skill of those in the art. See, e.g., Sambrook and Green, Molecules lar Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocol ls in Molecular Biology (FM Ausubel, et al. eds.); the series Methods In Enzy mology (Academic Press, Inc.), PCR 2: A Practical Approach (MJ MacPherson, BD Hames and GR Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique See nd Specialized Applications, 6th Edition (RI Freshney, ed. (2010)).
[0085] The section headings used herein are for organizational purposes only and It should not be construed as limiting the subject matter described.
[0086] Various features of the disclosure may be described in the context of a single embodiment, but these features may also be They may be provided separately or in any suitable combination. Although the specification may, for clarity, be described in the context of separate embodiments, the present disclosure also provides a single The section headings used herein are intended to be general terms that are not intended to be limiting. is for illustrative purposes only and should not be construed as limiting the subject matter described. do not have.
[0087] The features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the advantages will be obtained from the following detailed description which sets forth illustrative embodiments in which the principles of the present disclosure are utilized. By reference to the description and by considering the accompanying drawings described below, This can be obtained.
[0088] definition The following definitions supplement those in the art and are intended for this application: Related or unrelated matters, such as those resulting from commonly owned patents or applications Any methods and materials similar or equivalent to those described herein are not intended to be limiting. Although the materials and methods described herein may be used in carrying out the tests shown, preferred materials and methods are described herein. Therefore, the terminology used herein is for the purpose of describing particular embodiments only. It is for illustrative purposes only and is not intended to be limiting.
[0089] Unless otherwise defined, all technical and scientific terms used herein are defined by the It has the meaning commonly understood by one of ordinary skill in the art to which the invention pertains. , provides those skilled in the art with general definitions of many of the terms used in this invention: n et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The C ambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary o f Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991).
[0090] In this application, the use of the singular includes the plural unless specifically stated otherwise. As used herein, the singular forms "a," "an," and "the" are used unless the context clearly indicates otherwise. It should be noted that unless specifically indicated, plural referents are included. In this context, the use of "or" means "and / or" and is inclusive, unless otherwise stated. Furthermore, the terms "including" and "include" should be understood to mean The use of other forms such as "," "includes," and "included" is non-limiting. It is quantitative.
[0091] As used in this specification and claims, the terms "comprising" (and " "comprise" and "comprises" and all its forms), "having" g) (any of its forms, such as "have" and "has"), "include "including" ("include" and "includes" and any of its forms) or "including containing (any of its forms, such as "contains" and "contain") ) is inclusive or open-ended and does not exclude additional, unrecited elements or method steps. Any embodiment discussed herein may be used with respect to any method or composition of the present disclosure. It is believed that the same can be done, and vice versa. The method of the present disclosure can be achieved by
[0092] The terms "about" or "approximately" refer to a range of values as determined by one of ordinary skill in the art. This means that the value is within an acceptable margin of error for a particular value, which indicates how How it is measured or determined depends in part on the limitations of the measurement system. For example, "about" means, according to practice in the art, within 1 or more than 1 standard deviation. Alternatively, "about" can mean up to 20%, up to 10%, up to 5%, or Alternatively, it may refer to a range of up to 1% of the total mass of a biological system or process. Therefore, the term can mean values within the same order of magnitude, such as within 5 or within 2. Where specific values are recited in the application and claims, unless otherwise stated, The term "about" is used to indicate an acceptable error range for that particular value. It should be determined.
[0093] Ranges provided herein are understood to be shorthand for all values within the range. For example, the range 1 to 50 is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16. , 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36 , 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 It is understood to include any number, combination of numbers, or subrange.
[0094] In the specification, "some embodiments," "an embodiment," "one embodiment," or Reference to "another embodiment" may include any particular features, structures, or feature is included in at least some embodiments of the present disclosure, but not necessarily all. This means that the embodiments are not necessarily included.
[0095] "Abasic base editor" is a tool that extracts nucleobases and converts them into DNA nucleobases (A, T, Abasic base editors refer to agents that can insert nucleic acid groups (C, G, or G). In one embodiment, the nucleic acid glycosylase comprises a glycosylase polypeptide or a fragment thereof. at amino acid 204 of the following sequence or the corresponding position in uracil DNA glycosylase: containing Asp (e.g., substituting Asn at amino acid 204) and cytosine-DNA glycosylation A mutant human uracil DNA glycosylase or an active fragment thereof having enzyme activity. In one embodiment, the nucleic acid glycosylase comprises amino acid 147 or uracil of the following sequence: containing Ala, Gly, Cys, or Ser at the corresponding positions of a DNA glycosylase (e.g., a mutation in which Tyr is substituted at amino acid 147, and which has thymine-DNA glycosylase activity An exemplary human uracil-DNA glycosylase is a human uracil-DNA glycosylase or an active fragment thereof. The sequence of A glycosylase, isoform 1 is as follows: 1 mgvfclgpwg lgrklrtpgk gplqllsrlc gdhlqaipak kapagqeepg tppssplsae 61 qldriqrnka aallrlaarn vpvgfgeswk khlsgefgkp yfiklmgfva eerkhytvyp 121 pphqvftwtq mcdikdvkvv ilgqdp y hgp nqahglcfsv qrpvppppsl eniykelstd 181 iedfvhpghg dlsgwakqgv lll navltvr ahqanshker gweqftdavv swlnqnsngl 241 vfllwgsyaq kkgsaidrkr hhvlqtahps p l svy r gffg crhfsktnel lqksgkkpid 301 wkel
[0096] The sequence of human uracil-DNA glycosylase, isoform 2 is as follows: 1 migqktlysf fspsparkrh apspepavqg tgvagvpees gdaaaipakk apagqeepgt 61 ppssplsaeq ldriqrnkaa allrlaarnv pvgfgeswkk hlsgefgkpy fiklmgfvae 121 erkhytvypp phqvftwtqm cdikdvkvvi lgqdp y hgpn qahglcfsvq rpvppppsle 181 niykelstdi edfvhpghgd lsgwakqgvl ll n avltvra hqanshkerg weqftdavvs 241 wlnqnsnglv fllwgsyaqk kgsaidrkrh hvlqtahpsp l svy r gffgc rhfsktnell 301 qksgkkpidw kel
[0097] In other embodiments, the abasic editor is a base editor as described in PCT / JP205 / 080958 and US20170321210. The base editor may be any of the base editors listed in the In certain embodiments, the abasic editor is shown in bold and underlined in the sequence above. or any other abasic editor or uracil degumming known in the art. In one embodiment, the abasic enzyme comprises a mutation at the corresponding amino acid in the glycosylase. The diter contains mutations at Y147, N204, L272, and / or R276 or corresponding positions. In another embodiment, the abasic editor is a Y147A or Y147G mutation or the corresponding In another embodiment, the abasic editor comprises a N204D mutation or a corresponding In another embodiment, the abasic editor comprises a L272A mutation or a corresponding In another embodiment, the abasic editor comprises a R276E or R276C mutation. or corresponding mutations.
[0098] "Adenosine deaminase" refers to the enzyme that hydrolyzes the deamination of adenine or adenosine. In one embodiment, the term "de" refers to a polypeptide or fragment thereof that is capable of catalyzing the deactivation of a protein. The aminase or deaminase domain converts adenosine to inosine or deoxyribonucleic acid. Adenosine deamina catalyzes the hydrolytic deamination of adenosine to deoxyinosine. In some embodiments, the adenosine deaminase is a deoxyribonucleic acid (DNA) The compounds provided herein catalyze the hydrolytic deamination of adenine or adenosine in the ribozyme. Adenosine deaminases (e.g., engineered adenosine deaminases, evolved The adenosine deaminase (enzyme-activated adenosine deaminase) can be from any organism, such as a bacterium.
[0099] In one embodiment, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is a TadA variant. The dA variant is TadA*7.10. In one embodiment, the deaminase or deaminase The enzyme domains are derived from, for example, humans, chimpanzees, gorillas, monkeys, cows, dogs, rats, or is a variant of a naturally occurring deaminase from an organism such as a mouse. The aminase or deaminase domain is non-naturally occurring. For example, some In embodiments, the deaminase or deaminase domain is a deaminase specific to a naturally occurring deaminase. At least 50%, at least 55%, at least 60%, at least 65%, at least 70%, At least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least at least 92%, at least 93%, at least, at least 94%, at least 95%, at least 96 %, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2% , at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identity. In is, for example, International PCT Application Nos. PCT / 2017 / 045381 (WO2018 / 027078) and PCT / US2016 / 058344 ( WO2017 / 070632), each of which is incorporated herein by reference in its entirety. Also included is Komor, AC, et al., The entire contents of which are incorporated herein by reference. “Programmable editing of a target base in genomic DNA without double-stranded D NA cleavage” Nature 533, 420-424 (2016); Gaudelli, NM, et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, AC, et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher eff iciency and product purity” Science Advances 3:eaao4774 (2017), and Rees, H.A. ., et al., “Base editing: precision chemistry on the genome and transcriptome o f living cells.” Nat Rev Genet. 2018 Dec;19(12):770-788. doi: 10.1038 / s41576-0 See also 18-0059-1.
[0100] In certain embodiments, the adenosine deaminase comprises a modification in the following sequence: MSEVEFSHEY WMRHALTLAK RARDEREVPV GAVLVLNNRV IGEGWNRAIG LHDPTAHAEI MALRQGGLVM QNY RLIDATL YVTFEPCVMC AGAMIHSRIG RVVFGVRNAK TGAAGSLMDV LHYPGMNHRV EITEGILADE CAALLC YFFR MPRQVFNAQK KAQSSTD (also known as TadA*7.10).
[0101] In certain embodiments, the adenosine deaminase homodimer comprises the TadA*7.10 domain. and an adenosine deaminase domain selected from one of the following: Staphylococcus aureus (S. aureus) TadA: MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTL YVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN Bacillus subtilis (B. subtilis) TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTL EPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLS E Salmonella typhimurium (S. typhimurium) TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLV LQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFF RMRRQEIKALKKADRAEGAGPAV Shewanella putrefaciens (S. putrefaciens) TadA: MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEP CAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFKRRRDEKKALQRAQ QGIE Haemophilus influenzae F3031 (H. influenzae) TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗAEIIALRNGAKNIQNYR LLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRRE EKKIEKALLKSLSDK Caulobacter crescentus (C. crescentus) TadA: MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLT DLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAK I Geobacter sulfurreducens (G. sulfurreducens) TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLT GATLYVTLEPCLMCCMGAIILARLERVVFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRK KAKATPALFIDERKVPPEP TadA*7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD
[0102] "Administering" refers to providing one or more compositions described herein to a patient or subject. By way of example and not limitation, administration of a composition For example, injections can be intravenous (iv), subcutaneous (sc), intradermal (id), or intraperitoneal (i. p.) injection or intramuscular (i.m.) injection. Using more than one such route. Parenteral administration can be, for example, by bolus injection or by gradual infusion over time. In some embodiments, parenteral administration can be achieved by infusing the Intraductal, intravenous, intramuscular, intraarterial, intrathecal, intratumoral, intradermal, intraperitoneal, transtracheal, subcutaneous, subcorneal , including intra-articular, intracapsular, intrathecal and intrasternal infusion or injection; or Concurrently, administration can be by the oral route.
[0103] An "agent" is any small molecule compound, antibody, nucleic acid molecule, or polypeptide. , or fragments thereof.
[0104] "Alteration" refers to any alteration that can be detected by standard art known methods such as those described herein. Such a change in the structure, expression level or activity of a gene or polypeptide (e.g., an increase or As used herein, modification means a change in a polynucleotide or A change in the sequence of a polypeptide or a change in expression level, for example, a 10% change, a 25% change, This includes a 40% change, a 50% change, or a greater change in expression level.
[0105] "Ameliorate" means to reduce, inhibit, or attenuate the occurrence or progression of a disease. "To cause, reduce, stop, or stabilize" means to cause, reduce, stop, or stabilize.
[0106] "Analog" means a molecule that is not identical but has similar functional or structural characteristics. For example, a polynucleotide or polypeptide analog is a polynucleotide or polypeptide analog that is similar to the corresponding naturally occurring polynucleotide. A naturally occurring polynucleotide or polypeptide can be synthesized while retaining the biological activity of the nucleotide or polypeptide. or polypeptides, and have specific modifications that enhance the function of the analogs. Such modifications can improve the DNA affinity, efficiency, etc. of the analogs without altering, for example, ligand binding. affinity, specificity, protease or nuclease resistance, membrane permeability, and / or half-life Analogs can be used to increase the frequency of non-naturally occurring polynucleotides or amino acids. It may include.
[0107] A "base editor (BE)" or "nucleobase editor (NBE)" is a In various embodiments, the term "agent" refers to an agent that binds to an oxidase and has nucleobase modifying activity. Base editors include nucleobase-modifying polypeptides (e.g., deaminases) and nucleic acid promoters. The ramifiable nucleotide-binding domain is attached to a guide polynucleotide (e.g., a guide RNase A). In various embodiments, the agent comprises a protein having base editing activity. domains, i.e., bases (e.g., A, T, C, G, U) within a nucleic acid molecule (e.g., DNA), In some embodiments, the biomolecular complex comprises a domain capable of binding to a polypeptide. The oligonucleotide programmable DNA binding domain is a nucleotide-programmable DNA binding domain that binds to one or more deaminase domains. In one embodiment, the agent is fused or linked to a base editing agent. In another embodiment, the fusion protein has base editing activity. The protein domain is linked to the guide RNA (e.g., an RNA-binding motif on the guide RNA). In one embodiment, the RNA-binding domain is fused to a ribosomal deaminase (e.g., via an RNA-binding domain fused to a ribosomal deaminase). A domain with base editing activity can deaminate bases within a nucleic acid molecule. In some embodiments, the base editor is capable of deaminating one or more bases in a DNA molecule. In some embodiments, the base editor modifies a cytosine (C) or adenosine (A) in DNA. In some embodiments, the base editor can deaminate (A). Tosine (C) and adenosine (A) can be deaminated. In , base editors can deaminate cytosines (C) in DNA In some embodiments, the base editor is a cytidine base editor (CBE) ( In some embodiments, the base editor is an adenovirus in DNA. In some embodiments, the base editor can deaminate syn(A). - is a naturally occurring molecule with base editing activity and / or programmable DNA binding activity A standard base editor is a protein domain that contains a standard cytogenetic protein. The base editor can include a cytidine deaminase (e.g., an APOBEC cytidine deaminase). In some embodiments, a standard cytidine deaminase is used. The enzyme includes an APOBEC1 cytidine deaminase (e.g., rAPOBEC1). In this state, canonical cytidine base editors are associated or linked to cytidine deaminases. and / or further comprising an additional domain linked to the UGI domain, e.g., one or more UGI domains may be linked to a cytidine deoxyribonucleotide. In some embodiments, the base editor can be linked to an adenosine base editor (ABE) and cytidine base editor (CBE).
[0108] In some embodiments, the base editor is an adenosine deaminase and / or or nuclease-inactive Cas9 (dCas9) fused to cytidine deaminase. In some embodiments, the Cas9 is a circular permutant Cas9 (e.g., spCa Circularly permuted Cas9 is known in the art and is described, for example, in Oakes et al. I., Cell 176, 254-267, 2019. In some embodiments, the base enzyme Deter is fused to an inhibitor of base excision repair, e.g., a UGI domain or a dISN domain. In one embodiment, the fusion protein comprises one or more deaminases and a UGI domain. Cas9 nickases fused to inhibitors of base excision repair, such as dISN domains, or In other embodiments, the base editor is an abasic base editor.
[0109] In some embodiments, the adenosine base editor is an adenosine deaminase The variants are derived from a circularly permuted Cas9 (e.g., spCAS9 or saCAS9) and a bipartite nuclear localization sequence (BNCS). Circularly permuted Cas9s are generated by cloning into a scaffold comprising This is known and described, for example, in Oakes et al., Cell 176, 254-267, 2019. Suitable circular permutations are described below, where bolded sequences indicate Cas9-derived sequences and italicized sequences. The sequence indicates the linker sequence, and the underlined sequence indicates the bipartite nuclear localization sequence.
[0110] JPEG0007779988000001.jpg187163
[0111] In some embodiments, polynucleotide-programmable DNA binding The domain is a CRISPR-associated (e.g., Cas or Cpf1) enzyme. In this study, a base editor was developed that uses catalytically inactive (dead) Cas9 (dCas9) in one or more deaminases. In some embodiments, the base editor is fused to a ribozyme domain. , in which a Cas9 nickase (nCas9) is fused to one or more deaminase domains. In some embodiments, the base editor inhibits base excision repair (BER). In one embodiment, the inhibitor of base excision repair is a uracil-DNA glycoprotein. In one embodiment, the inhibitor of base excision repair is a guanosine triphosphate (UGI). , an inosine base excision repair inhibitor.
[0112] Further details of the base editors are available in International PCT Application Nos. PCT / 2017 / 045381 (WO2018 / 027078) and PCT / U S2016 / 058344 (WO2017 / 070632), each of which is incorporated herein by reference in its entirety. Also, Komor, A., the entire contents of which are incorporated herein by reference. .C., et al., “Programmable editing of a target base in genomic DNA without doub "le-stranded DNA cleavage" Nature 533, 420-424 (2016); Gaudelli, NM, et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage ” Nature 551, 464-471 (2017); Komor, AC, et al., “Improved base excision re pair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao4774 (2017), and Rees, HA, et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec;19(12):770-788. doi: 1 See also 0.1038 / s41576-018-0059-1.
[0113] By way of example, adenine dinucleotides used in the base editing compositions, systems, and methods described herein can be used. The ABE has the following nucleic acid sequence (8877 base pairs) (Addgene, Waters, WA) own, MA.; Gaudelli NM, et al., Nature. 2017 Nov 23;551(7681):464-471. doi: 10. 1038 / nature24644; Koblan LW, et al., Nat Biotechnol. 2018 Oct;36(9):843-846. d oi: 10.1038 / nbt.4172.) ABE nucleic acid sequence with at least 95% identity Also included are polynucleotide sequences that
[0114] ATATGCCAAGTACGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTATGCCCAGTACAT GACCTTATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGTGATGCGG TTTTGGCAGTACATCAATGGGCGTGGATAGCGGTTTGACTCACGGGATTTCCAAGTCTCCACCCCATTG ACGTCAATGGGAGTTTGTTTTGGCACCAAAATCAACGGGACTTTCCAAAATGTCGTAACAACTCCGCCCC ATTGACGCAAATGGGCGGTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCTGGTTTAGTGAACCGT CAGATCCGCTAGAGATCCGCGGCCGCTAATACGACTCACTATAGGGAGAGCCGCCACCATGAAACGGACA GCCGACGGAAGCGAGTTCGAGTCACCAAAGAAGAAGCGGAAAGTCTCTGAAGTCGAGTTTAGCCACGAGT ATTGGATGAGGCACGCACTGACCCTGGCAAAGCGAGCATGGGATGAAAGAGAAGTCCCCGTGGGCGCCGT GCTGGTGCACAACAATAGAGTGATCGGAGAGGGATGGAACAGGCCAATCGGCCGCCACGACCCTACCGCA CACGCAGAGATCATGGCACTGAGGCAGGGAGGCCTGGTCATGCAGAATTACCGCCTGATCGATGCCACCC TGTATGTGACACTGGAGCCATGCGTGATGTGCGCAGGAGCAATGATCCACAGCAGGATCGGAAGAGTGGT GTTCGGAGCACGGGACGCCAAGACCGGCGCAGCAGGCTCCCTGATGGATGTGCTGCACCACCCCGGCATG AACCACCGGGTGGAGATCACAGAGGGAATCCTGGCAGACGAGTGCGCCGCCCTGCTGAGCGATTTCTTTA GAATGCGGAGACAGGAGATCAAGGCCCAGAAGAAGGCACAGAGCTCCACCGACTCTGGAGGATCTAGCGG AGGATCCTCTGGAAGCGAGACACCAGGCACAAGCGAGTCCGCCACACCAGAGAGCTCCGGCGGCTCCTCC GGAGGATCCTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGGG CACGCGATGAGAGGGAGGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCGAGGGCTG GAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCTGAGACAGGGCGGCCTG GTCATGCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCG GCGCCATGATCCACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACGCAAAAACCGGCGCCGCAGG CTCCCTGATGGACGTGCTGCACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCA GATGAATGTGCCGCCCTGCTGTGCTATTTCTTTCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGG CCCAGAGCTCCACCGACTCCGGAGGATCTAGCGGAGGCTCCTCTGGCTCTGAGACACCTGGCACAAGCGA GAGCGCAACACCTGAAAGCAGCGGGGGCAGCAGCGGGGGGTCAGACAAGAAGTACAGCATCGGCCTGGCC ATCGGCACCAACTCTGTGGGCTGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGG TGCTGGGCAACACCGACCGGCACAGCATCAAGAAGAACCTGATCGGAGCCCTGCTGTTCGACAGCGGCGA AACAGCCGAGGCCACCCGGCTGAAGAGAACCGCCAGAAGAAGATACACCAGACGGAAGAACCGGATCTGC TATCTGCAAGAGATCTTCAGCAACGAGATGGCCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGT CCTTCCTGGTGGAAGAGGATAAGAAGCACGAGCGGCACCCCATCTTCGGCAACATCGTGGACGAGGTGGC CTACCACGAGAAGTACCCCACCATCTACCACCTGAGAAAGAAACTGGTGGACAGCACCGACAAGGCCGAC CTGCGGCTGATCTATCTGGCCCTGGCCCACATGATCAAGTTCCGGGGCCACTTCCTGATCGAGGGCGACC TGAACCCCGACAACAGCGACGTGGACAAGCTGTTCATCCAGCTGGTGCAGACCTACAACCAGCTGTTCGA GGAAAACCCCATCAACGCCAGCGGCGTGGACGCCAAGGCCATCCTGTCTGCCAGACTGAGCAAGAGCAGA CGGCTGGAAAATCTGATCGCCCAGCTGCCCGGCGGAGAAAGAATGGCCTGTTCGGAAACCTGATTGCCC TGAGCCTGGGCCTGACCCCAACTTCAAGAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCAGCTGAG CAAGGACACCTACGACGACGACGACCTGGACAACCTGCTGGCCCAGATCGGCGACCAGTACGGCCGACCTGTTTT CTGGCCGCCAAGAACCTGCTCGACGCCATCCTGCTGAGCGACATCCTGAGAGTGAACACCGAGATCACCA AGGCCCCCCTGAGCGCCTCTATGATCAAGAGATACGACGAGCACCACCAGGACCTGACCCTGCTGAAAGC TCTCGTGCGGCAGCAGCTGCCTGAGAAGTACAAAGAGATTTTCTTCGACCAGAGCAAGAACGGCTACGCC GGCTACATTGACGGCGGAGCCAGCCAGGAAGAGTTCTACAAGTTCATCAAGCCCATCCTGGAAAAGATGG ACGGCACCGAGGAACTGCTCGTGAAGCTGAACAGAGGACCTGCTGCGGAAGCAGCGGACCTTCGACAA CGGCAGCATCCCCCACCAGATCCACCTGGGAGAGCTGCACGCCATTCTGCGGCGGCAGGAAGATTTTTAC CCATTCCTGAAGGACAACCGGGAAAGATCGAGAAGATCCTGACCTTCCGCATCCCCTACTACGTGGGCC CTCTGGCCAGGGGAAACAGCAGATTCGCCTGGATGACCAGAAAGAGCGAGGAAACCATCACCCCCTGGAA CTTCGAGGAAGTGGTGGACAAGGGCGCTTCCGCCCAGAGCTTCATCGAGCGGATGACCAACTTCGATAAG AACCTGCCCAACGAGAAGGTGCTGCCCAAGCACAGCCTGCTGTACGAGTACTTCACCGTGTATAACGAGC TGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCCGCCTTCCTGAGCGGCGAGCAGAAAAAGGC CATCGTGGACCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGACTACTTCAAG AAAATCGAGTGCTTCGACTCCGTGGAAATCTCCGGCGTGGAAGATCGGTTCAACGCCTCCCTGGGCACAT ACCACGATCTGCTGAAAATTATCAAGGACAAGGACTTCCTGGACAATGAGGAAAACGAGGACATTCTGGA AGATATCGTGCTGACCCTGACACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACCTATGCC CACCTGTTCGACGACAAAGTGATGAAGCAGCTGAAGCGGCGGAGATACACCGGCTGGGGCAGGCTGAGCC GGAAGCTGATCAACGGCATCCGGGACAAGCAGTCCGGCAAGACAATCCTGGATTTCCTGAAGTCCGACGG CTTCGCCAACAGAAACTTCATGCAGCTGATCCACGACGACAGCCTGACCTTTAAAGAGGACATCCAGAAA GCCCAGGTGTCCGGCCAGGGCGATAGCCTGCACGAGCACATTGCCAATCTGGCCGGCAGCCCCGCCATTA AGAAGGGCATCCTGCAGACAGTGAAGGTGGTGGACGAGCTCGTGAAAGTGATGGGCCGGCACAAGCCCGA GAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACCCAGAAGGGACAGAAGAACAGCCGCGAGAGA ATGAAGCGGATCGAAGAGGGCATCAAAGAGCTGGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAAACA CCCAGCTGCAGAACGAGAAGCTGTACCTGTACTACCTGCAGAATGGGCGGGATATGTACGTGGACCAGGA ACTGGACATCAACCGGCTGTCCGACTACGATGTGGACCATATCGTGCCTCAGAGCTTTCTGAAGGACGAC TCCATCGACAACAAGGTGCTGACCAGAAGCGACAAGAACCGGGGCAAGAGCGACAACGTGCCCTCCGAAG AGGTCGTGAAGAAGATGAAGAACTACTGGCGGCAGCTGCTGAACGCCAAGCTGATTACCCAGAGAAAGTT CGACAATCTGACCAAGGCCGAGAGAGGCGGCCTGAGCGAACTGGATAAGGCCGGCTTCATCAAGAGACAG CTGGTGGAAACCCGGCAGATCACAAAGCACGTGGCACAGATCCTGGACTCCCGGATGAACACTAAGTACG ACGAGAATGACAAGCTGATCCGGGAAGTGAAAGTGATCACCCTGAAGTCCAAGCTGGTGTCCGATTTCCG GAAGGATTTCCAGTTTTACAAAGTGCGCGAGATCAACAACTACCACCACGCCCACGACGCCTACCTGAAC GCCGTCGTGGGAACCGCCCTGATCAAAAAGTACCCTAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACA AGGTGTACGACGTGCGGAAGATGATCGCCAAGAGCGAGCAGGAAATCGGCAAGGCTACCGCCAAGTACTT CTTCTACAGCAACATCATGAACTTTTTCAAGACCGAGATTACCCTGGCCAACGGCGAGATCCGGAAGCGG CCTCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGTGGGATAAGGGCCGGGATTTTGCCACCGTGC GGAAAGTGCTGAGCATGCCCCAAGTGAATATCGTGAAAAAGACCGAGGTGCAGACAGGCGGCTTCAGCAA AGAGTCTATCCTGCCCAAGAGGAACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGGGACCCTAAGAAG TACGGCGGCTTCGACAGCCCCACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGGGCAAGT CCAAGAAACTGAAGAGTGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAGAAGCAGCTTCGAGAAGAA TCCCATCGACTTTCTGGAAGCCAAGGGCTACAAAGAAGTGAAAAAGGACCTGATCATCAAGCTGCCTAAG TACTCCCTGTTCGAGCTGGAAAACGGCCGGAAGAGAATGCTGGCCTCTGCCGGCGAACTGCAGAAGGGAA ACGAACTGGCCCTGCCCTCCAAATATGTGAACTTCCTGTACCTGGCCAGCCACTATGAGAAGCTGAAGGG CTCCCCCGAGGATAATGAGCAGAAACAGCTGTTTGTGGAACAGCACAAGCACTACCTGGACGAGATCATC GAGCAGATCAGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAATCTGGACAAAGTGCTGTCCGCCT ACAACAAGCACCGGGATAAGCCCATCAGAGAGCAGGCCGAGAATATCATCCACCTGTTTACCCTGACCAA TCTGGGAGCCCCTGCCGCCTTCAAGTACTTTGACACCACCATCGACCGGAAGAGGTACACCAGCACCAAA GAGGTGCTGGACGCCACCCTGATCCACCAGAGCATCACCGGCCTGTACGAGACACGGATCGACCTGTCTC AGCTGGGAGGTGACTCTGGCGGCTCAAAAAGAACCGCCGACGGCAGCGAATTCGAGCCCAAGAAGAAGAG GAAAGTCTAACCGGTCATCATCACCATCACCATTGAGTTTAAACCCGCTGATCAGCCTCGACTGTGCCTT CTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCAC TGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGT GGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCT CTATGGCTTCTGAGGCGGAAAGAACCAGCTGGGGCTCGATACCGTCGACCTCTAGCTAGAGCTTGGCGTA ATCATGGTCATAGCTGTTTCCTGTGTGAAATTGTTATCCGCTCACAATTCCACACAACATACGAGCCGGA AGCATAAAGTGTAAAGCCTAGGGTGCCTAATGAGTGAGCTAACTCACATTAATTGCGTTGCGCTCACTGC CCGCTTTCCAGTCGGGAAACCTGTCGTGCCAGCTGCATTAATGAATCGGCCAACGCGCGGGGAGAGGCGG TTTGCGTATTGGGCGCTCTTCCGCTTCCTCGCTCACTGACTCGCTGCGCTCGGTCGTTCGGCTGCGGCGA GCGGTATCAGCTCACTCAAAGGCGGTAATACGGTTATCCACAGAATCAGGGGATAACGCAGGAAAGAACA TGTGAGCAAAAGGCCAGCAAAAGGCCAGGAACCGTAAAAAGGCCGCGTTGCTGGCGTTTTTCCATAGGCT CCGCCCCCCTGACGAGCATCACAAAAATCGACGCTCAAGTCAGAGGTGGCGAAACCCGACAGGACTATAA AGATACCAGGCGTTTCCCCCTGGAAGCTCCCTCGTGCGCTCTCCTGTTCCGACCCTGCCGCTTACCGGAT ACCTGTCCGCCTTTCTCCCTTCGGGAAGCGTGGCGCTTTCTCATAGCTCACGCTGTAGGTATCTCAGTTC GGTGTAGGTCGTTCGCTCCAAGCTGGGCTGTGTGCACGAACCCCCCGTTCAGCCCGACCGCTGCGCCTTA TCCGGTAACTATCGTCTTGAGTCCAACCCGGTAAGACACGACTTATCGCCACTGGCAGCAGCCACTGGTA ACAGGATTAGCAGAGCGAGGTATGTAGGCGGTGCTACAGAGTTCTTGAAGTGGTGGCCTAACTACGGCTA CACTAGAAGAACAGTATTTGGTATCTGCGCTCTGCTGAAGCCAGTTACCTTCGGAAAAAGAGTTGGTAGC TCTTGATCCGGCAAACAAACCACCGCTGGTAGCGGTGGTTTTTTTGTTTGCAAGCAGCAGATTACGCGCA GAAAAAAAGGATCTCAAGAAGATCCTTTGATCTTTTCTACGGGGTCTGACACTCAGTGGAACGAAAACTC ACGTTAAGGGATTTTGGTCATGAGATTATCAAAAAGGATCTTCACCTAGATCCTTTTAAATTAAAAATGA AGTTTTAAATCAATCTAAAGTATATATGAGTAAACTTGGTCTGACAGTTACCAATGCTTAATCAGTGAGG CACCTATCTCAGCGATCTGTCTATTTCGTTCATCCATAGTTGCCTGACTCCCCGTCGTGTAGATAACTAC GATACGGGAGGGCTTACCATCTGGCCCCAGTGCTGCAATGATACCGCGAGACCCACGCTCACCGGCTCCA GATTTATCAGCAATAAACCAGCCAGCCGGAAGGGCCGAGCGCAGAAGTGGTCCTGCAACTTTATCCGCCT CCATCCAGTCTATTAATTGTTGCCGGGAAGCTAGAGTAAGTAGTTCGCCAGTTAATAGTTTGCGCAACGT TGTTGCCATTGCTACAGGCATCGTGGTGTCACGCTCGTCGTTTGGTATGGCTTCATTCAGCTCCGGTTCC CAACGATCAAGGCGAGTTACATGATCCCCCATGTTGTGCAAAAAAGCGGTTAGCTCCTTCGGTCCTCCGA TCGTTGTCAGAAGTAAGTTGGCCGCAGTGTTATCACTCATGGTTATGGCAGCACTGCATAATTCTCTTAC TGTCATGCCATCCGTAAGATGCTTTTCTGTGACTGGTGAGTACTCAACCAAGTCATTCTGAGAATAGTGT ATGCGGCGACCGAGTTGCTCTTGCCCGGCGTCAATACGGGATAATACCGCGCCACATAGCAGAACTTTAA AAGTGCTCATCATTGGAAAACGTTCTTCGGGGCGAAAACTCTCAAGGATCTTACCGCTGTTGAGATCCAG TTCGATGTAACCCACTCGTGCACCCAACTGATCTTCAGCATCTTTTACTTTCACCAGCGTTCTGGGTGA GCAAAAACAGGAAGGCAAAATGCCGCAAAAAAGGGAATAAGGGCGACACGGAAATGTTGAATACTCATAC TCTTCCTTTTTCAATATTATTGAAGCATTTATCAGGGTTATTGTCTCATGAGCGGATACATATTTGAATG TATTTAGAAAAATAAACAAATAGGGGTTCCGCGCACATTTCCCCGAAAGTGCCACCTGACGTCGACGGA TCGGGAGATCGATCTCCCGATCCCCTAGGGTCGACTCTCAGTACATACTGCTCCTGATGCCGCATAGTTAA GCCAGTATCTGCTCCCTGCTTGTGTGTTGGAGGTCGCTGAGTAGTGCGCGAGCAAAATTTAAGCTACAAC AAGGCAAGGCTTGACCGACAATTGCATGAAGAATCTGCTTAGGGTTAGGCGTTTGCGCTGCTTCGCGGAT GTACGGGCCAGATATACGCGTTGACATTGATTATTGACTAGTTATTAATAGTAATCAAATTACGGGGTCAT TAGTTCATAGCCCATATATGGAGTTCCGCGTTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCC CAACGACCCCCGCCCATTGACGTCAATAATGACGTATGTTCCCATAGTAACGCCAATAGGGACTTTCCAT TGACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAGTACATCAAGTGTATC
[0115] As an example, the specification described in this article will be used as a base The CBE has the nucleic acid sequence (8877 base pairs) provided below (Add gene, Watertown, MA.; Komor AC, et al., 2017, Sci Adv., 30;3(8):eaao4774. doi: 1 0.1126 / sciadv.aao4774) with at least 95% identity to the BE4 nucleic acid sequence Polynucleotide sequences are also included.
[0116] 1 ATATGCCAAG TACGCCCCCT ATTGACGTCA ATGACGGTAA ATGGCCCGCC TGGCATTATG 61 CCCAGTACAT GACCTTATGG GACTTTCCTA CTTGGCAGTA CATCTACGTA TTAGTCATCG 121 CTATTACCAT GGTGATGCGG TTTTGGCAGT ACATCAATGG GCGTGGATAG CGGTTTGACT 181 CACGGGGATT TCCAAGTCTC CACCCCATTG ACGTCAATGG GAGTTTGTTT TGGCACCAAA 241 ATCAACGGGA CTTTCCAAAA TGTCGTAACA ACTCCGCCCC ATTGACGCAA ATGGGCGGTA 301 GGCGTGTACG GTGGGAGGTC TATATAAGCA GAGCTGGTTT AGTGAACCGT CAGATCCGCT 361 AGAGATCCGC GGCCGCTAAT ACGACTCACT ATAGGGAGAG CCGCCACCAT GAGCTCAGAG 421 ACTGGCCCAG TGGCTGTGGA CCCCACATTG AGACGGCGGA TCGAGCCCCA TGAGTTTGAG 481 GTATTCTTCG ATCCGAGAGA GCTCCGCAAG GAGACCTGCC TGCTTTACGA AATTAATTGG 541 GGGGGCCGGC ACTCCATTTG GCGACATACA TCACAGAACA CTAACAAGCA CGTCGAAGTC 601 AACTTCATCG AGAAGTTCAC GACAGAAAGA TATTTCTGTC CGAACACAAG GTGCAGCATT 661 ACCTGGTTTC TCAGCTGGAG CCCATGCGGC GAATGTAGTA GGGCCATCAC TGAATTCCTG 721 TCAAGGTATC CCCACGTCAC TCTGTTTATT TACATCGCAA GGCTGTACCA CCACGCTGAC 781 CCCCGCAATC GACAAGGCCT GCGGGATTTG ATCTCTTCAG GTGTGACTAT CCAAATTATG 841 ACTGAGCAGG AGTCAGGATA CTGCTGGAGA AACTTTGTGA ATTATAGCCC GAGTAATGAA 901 GCCCACTGGC CTAGGTATCC CCATCTGTGG GTACGACTGT ACGTTCTTGA ACTGTACTGC 961 ATCATACTGG GCCTGCCTCC TTGTCTCAAC ATTCTGAGAA GGAAGCAGCC ACAGCTGACA 1021 TTCTTTACCA TCGCTCTTCA GTCTTGTCAT TACCAGCGAC TGCCCCCACA CATTCTCTGG 1081 GCCACCGGGT TGAAATCTGG TGGTTCTTCT GGTGGTTCTA GCGGCAGCGA GACTCCCGGG 1141 ACCTCAGAGT CCGCCACACC CGAAAGTTCT GGTGGTTCTT CTGGTGGTTC TGATAAAAAG 1201 TATTCTATTG GTTTAGCCAT CGGCACTAAT TCCGTTGGAT GGGCTGTCAT AACCGATGAA 1261 TACAAAGTAC CTTCAAAGAA ATTTAAGGTG TTGGGGAACA CAGACCGTCA TTCGATTAAA 1321 AAGAATCTTA TCGGTGCCCT CCTATTCGAT AGTGGCGAAA CGGCAGAGGC GACTCGCCTG 1381 AAACGAACCG CTCGGAGAAG GTATACACGT CGCAAGAACC GAATATGTTA CTTACAAGAA 1441 ATTTTTAGCA ATGAGATGGC CAAAGTTGAC GATTCTTTCT TTCACCGTTT GGAAGAGTCC 1501 TTCCTTGTCG AAGAGGACAA GAAACATGAA CGGCACCCCA TCTTTGGAAA CATAGTAGAT 1561 GAGGTGGCAT ATCATGAAAA GTACCCAACG ATTTATCACC TCAGAAAAAA GCTAGTTGAC 1621 TCAACTGATA AAGCGGACCT GAGGTTAATC TACTTGGCTC TTGCCCATAT GATAAAGTTC 1681 CGTGGGCACT TTCTCATTGA GGGTGATCTA AATCCGGACA ACTCGGATGT CGACAAACTG 1741 TTCATCCAGT TAGTACAAAC CTATAATCAG TTGTTTGAAG AGAACCCTAT AAATGCAAGT 1801 GGCGTGGATG CGAAGGCTAT TCTTAGCGCC CGCCTCTCTA AATCCCGACG GCTAGAAAAAC 1861 CTGATCGCAC AATTACCCGG AGAGAAGAAA AATGGGTTGT TCGGTAACCT TATAGCGCTC 1921 TCACTAGGCC TGACACCAAA TTTTAAGTCG AACTTCGACT TAGCTGAAGA TGCCAAATTG 1981 CAGCTTAGTA AGGACACGTA CGATGACGAT CTCGACAATC TACTGGCACA AATTGGAGAT 2041 CAGTATGCGG ACTTATTTTT GGCTGCCAAA AACCTTAGCG ATGCAATCCT CCTATCTGAC 2101 ATACTGAGAG TTAATACTGA GATTACCAAG GCGCCGTTAT CCGCTTCAAT GATCAAAAGG 2161 TACGATGAAC ATCACCAAGA CTTGACACTT CTCAAGGCCC TAGTCCGTCA GCAACTGCCT 2221 GAGAAATATA AGGAAATATT CTTTGATCAG TCGAAAAACG GGTACGCAGG TTATATTGAC 2281 GGCGGAGCGA GTCAAGAGGA ATTCTACAAG TTTATCAAAC CCATATTAGA GAAGATGGAT 2341 GGGACGGAAG AGTTGCTTGT AAAACTCAAT CGGAAGATC TACTGCGAAA GCAGCGGACT 2401 TTCGACAACG GTAGCATTCC ACATCAAATC CACTTAGGCG AATTGCATGC TATACTTAGA 2461 AGGCAGGAGG ATTTTTATCC GTTCCTCAAA GACAATCGTG AAAAGATTGA GAAAATCCTA 2521 ACCTTTCGCA TACCTTACTA TGTGGGACCC CTGGCCCGAG GGAACTCTCG GTTCGCATGG 2581 REPEAT AGTCCREACT AACGATTACT CCATGGAATTT TTREPEAT TGTCREADAA 2641 GGTGCGTCAG CTCAATCGTT CATCGAGAGG ATGACCAACT TTGACAAA TTTACCGAAC 2701 GAAAAGTAT TGCCTAXASSOCIATTACTT CHAPTER TGACTACTACATCHACTC 2761 ACGAAAGTTA AGTATGTCAC TGAGGGCATG CGTAAACCCG CCTTTCTAAG CGGAGAACAG 2821 AAGAAAGCAA TAGTAGATCT GTTATTCAAG ACCAACCGCA AAGTGACAGT TAAGCAATTG 2881 AAAGAGGACT ACTTTAAGAA AATTGAATGC TTCGATTCTG TCGAGATCTC CGGGGTAGAA 2941 GATCGATTTA ATGCGTCACT TGGTACGTAT CATGACCCTCC WINDOW INTERFACE 3001 GACTTCCTGG ATTACK REPEAT ATCTCTGG ATTACKGTT GACTCTTACC 3061 CTCTTTGAAG ATCGGGAAAT GATTGAGGAA AGACTAAAAAA CATACGCTCA CCTGTTCGAC 3121 GATAAGGTTA TGAAACAGTT AAAGAGGCGT CGCTATACGG GCTGGGGACG ATTGTCGCGG 3181 AAACTTATCA ACGGGATAAG AGACAAGCAA AGTGGTAAAA CTATTCTCGA TTTTCTAAAG 3241 AGCGACGGCT TCGCCAATAG GAACTTTATG CAGCTGATCC ATGATGACTC TTTAACCTTC 3301 AAAGAGGATA TACAAAAGGC ACAGGTTTCC GGACAAGGGG ACTCATTGCA CGAACATATT 3361 GCGAATCTTG CTGGTTCGCC AGCCATCAAA AAGGGCATAC TCCAGACAGT CAAAGTAGTG 3421 GATGAGCTAG TTAAGGTCAT GGGACGTCAC AAACCGGAAA ACATTGTAAT CGAGATGGCA 3481 CGCGAAAATC AAACGACTCA GAAGGGGCAA AAAAAACAGTC GAGAGCGGAT GAAGAGAATA 3541 GAAGAGGGTA TTAAAGAACT GGGCAGCCAG ATCTTAAAGG AGCATCCTGT GGAAAATACC 3601 CAATTGCAGA ACGAGAAACT TTACCTCTAT TACCTACAAA ATGGAAGGGA CATGTATGTT 3661 GATCAGGAAC TGGACATAAA CCGTTTATCT GATTACGACG TCGATCACAT TGTACCCCAA 3721 TCCTTTTTGA AGGACGATTC AATCGACAAT AAAGTGCTTA CACGCTCGGA TAAGAACCGA 3781 GGGAAAAGTG ACAATGTTCC AAGCGAGGAA GTCGTAAAGA AAATGAGAA CTATTGGCGG 3841 CAGCTCCTAA ATGCGAAACT GATAACGCAA AGAAAGTTCG ATAACTTAAC TAAAGCTGAG 3901 AGGGGTGGCT TGTCTGAACT TGACAAGGCC GGATTTATTA AACGTCAGCT CGTGGAAACC 3961 CGCCAAATCA CAAAGCATGT TGCACAGATA CTAGATTCCC GAATGAATAC GAAATACGAC 4021 GAGAACGATA AGCTGATTCG GGAAGTCAAA GTAATCACTT TAAAGTCAAA ATTGGTGTCG 4081 GACTTCAGAA AGGATTTTCA ATTCTATAAA GTTAGGGAGA TAATAACTA CCACCATGCG 4141 CACGACGCTT ATCTTAATGC CGTCGTAGGG ACCGCACTCA TTAAGAAATA CCCGAAGCTA 4201 GAAAGTGAGT TTGTGTATGG TGATTACAAA GTTTATGACG TCCGTAAGAT GATCGCGAAA 4261 AGCGAACAGG AGATAGGCAA GGCTACAGCC AAATACTTCT TTTATTCTAA CATTATGAAT 4321 TTCTTTAAGA CGGAAATCAC TCTGGCAAAC GGAGAGATAC GCAAACGACC TTTAATTGAA 4381 ACCAATGGGG AGACAGGTGA AATCGTATGG GATAAGGGCC GGGACTTCGC GACGGTGAGA 4441 AAAGTTTTGT CCATGCCCCA AGTCAACATA GTAAAGAAAA CTGAGGTGCA GACCGGAGGG 4501 TTTTCAAAGG AATCGATTCT TCCAAAAAGG AATAGTGATA AGCTCATCGC TCGTAAAAAG 4561 GACTGGGACC CGAAAAAGTA CGGTGGCTTC GATAGCCCTA CAGTTGCCTA TTCTGTCCTA 4621 GTAGTGGCAA AAGTTGAGAA GGGAAAATCC AAGAAACTGA AGTCAGTCAA AGAATTATTG 4681 GGGATAACGA TTATGGAGCG CTCGTCTTTT GAAAAGAACC CCATCGACTT CCTTGAGGCG 4741 AAAGGTTACA AGGATAAA AAAGGATCTC ATTACK TACTTACA TAGTCTGTTT 4801 GAGTTAGAAA ATGGCCGAAA ACGGATGTTG GCTAGCGCCG GAGAGCTTCA AAAGGGGAAC 4861 GAACTCGCAC TACCGTCTAA ATACGTGAAT TTCCTGTATT TAGCGTCCCA TTACGAGAAG 4921 TTGAAAGGTT CACCTGAAGA TAACGAACAG AAGCAACTTT TTGTTGAGCA GCACAAACAT 4981 TATCTCGACG AAATCATAGA GCAAATTTCG GAATTCAGTA AGAGAGTCAT CCTAGCTGAT 5041 GCCAATCTGG ACAAAGTATT AAGCGCATAC AACAAGCACA GGGATAAACC CATACGTGAG 5101 CAGGCGGAAA ATATTATCCA TTTGTTTACT CTTACCAACC TCGGCGCTCC AGCCGCATTC 5161 AAGTATTTTG ACACAACGAT AGATCGCAAA CGATACACTT CTACCAAGGA GGTGCTAGAC 5221 GCGACACTGA TTCACCAATC CATCACGGGA TTATATGAAA CTCGGATAGA TTTGTCACAG 5281 CTTGGGGGTG ACTCTGGTGG TTCTGGAGGA TCTGGTGGTT CTACTAATCT GTCAGATATT 5341 ATTGAAAAGG AGACCGGTAA GCAACTGGTT ATCCAGGAAT CCATCCTCAT GCTCCCAGAG 5401 GAGGTGGAAG AAGTCATTGG GAACAAGCCG GAAAGCGATA TACTCGTGCA CACCGCCTAC 5461 GACGAGAGCA CCGACGAGAA TGTCATGCTT CTGACTAGCG ACGCCCCTGA ATACAAGCCT 5521 TGGGCTCTGG TCATACAGGA TAGCAACGGT GAGAACAAGA TTAAGATGCT CTCTGGTGGT 5581 TCTGGAGGAT CTGGTGGTTC TACTAATCTG TCAGATATTA TTGAAAAGGA GACCGGTAAG 5641 CAACTGGTTA TCCAGGAATC CATCCTCATG CTCCCAGAGG AGGTGGAAGA AGTCATTGGG 5701 AACAAGCCGG AAAGCGATAT ACTCGTGCAC ACCGCCTACG ACGAGAGCAC CGACGAGAAT 5761 GTCATGCTTC TGACTAGGCGA CGCCCCTGAA TACAAGCCTT GGGCTCTGGT CATACAGGAT 5821 AGCAACGGTG AGAACAAGAT TAAGATGCTC TCTGGTGGTT CTCCCAAGAA GAAGAGGAAA 5881 GTCTAACCGG TCATCATCAC CATCACCATT GAGTTTAAAC CCGCTGATCA GCCTCGACTG 5941 TGCCTTCTAG TTGCCAGCCA TCTGTTGTTT GCCCCTCCCC CGTGCCTTCC TTGACCCTGG 6001 AAGGTGCCAC TCCCACTGTC CTTTCCTAAT AAAATGAGGA AATTGCATCG CATTGTCTGA 6061 GTAGGTGTCA TTCTATTCTG GGGGGTGGGG TGGGGCAGGA CAGCAAGGGG GAGGATTGGG 6121 AAGACAATAG CAGGCATGCT GGGGATGCGG TGGGCTCTAT GGCTTCTGAG GCGGAAAGAA 6181 CCAGCTGGGG CTCGATACCG TCGACCTCTA GCTAGAGCTT GGCGTAATCA TGGTCATAGC 6241 TGTTTCCTGT GTGAAATTGT TATCCGCTCA CAATTCCACA CAACATACGA GCCGGAAGCA 6301 TAAAGTGTAA AGCCTAGGGT GCCTAATGAG TGAGCTAACT CACATTAATT GCGTTGCGCT 6361 CACTGCCCGC TTTCCAGTCG GGAAACCTGT CGTGCCAGCT GCATTAATGA ATCGGCCAAC 6421 GCGCGGGGAG AGGCGGTTTG CGTATTGGGC GCTCTTCCGC TTCCTCGCTC ACTGACTCGC 6481 TGCGCTCGGT CGTTCGGCTG CGGCGAGCGG TATCAGCTCA CTCAAAGGCG GTAATACGGT 6541 TATCCACAGA ATCAGGGGAT AACGCAGGAA AGAACATGTG AGCAAAAGGC CAGCAAAAGG 6601 CCAGGAACCG TAAAAAGGCC GCGTTGCTGG CGTTTTTCCA TAGGCTCCGC CCCCCTGACG 6661 AGCATCACAA AAATCGACGC TCAAGTCAGA GGTGGCGAAA CCCGACAGGA CTATAAAGAT 6721 ACCAGGCGTT TCCCCCTGGA AGCTCCCTCG TGCGCTCTCC TGTTCCGACC CTGCCGCTTA 6781 CCGGATACCT GTCCGCCTTT CTCCCTTCGG GAAGCGTGGC GCTTTCTCAT AGCTCACGCT 6841 GTAGGTATCT CAGTTCGGTG TAGGTCGTTC GCTCCAAGCT GGGCTGTGTG CACGAACCCC 6901 CCGTTCAGCC CGACCGCTGC GCCTTATCCG GTAACTATCG TCTTGAGTCC AACCCGGTAA 6961 GACACGACTT ATCGCCACTG GCAGCAGCCA CTGGTAACAG GATTAGCAGA GCGAGGTATG 7021 TAGGCGGTGC TACAGAGTTC TTGAAGTGGT GGCCTAACTA CGGCTACACT AGAAGAACAG 7081 TATTTGGTAT CTGCGCTCTG CTGAAGCCAG TTACCTTCGG AAAAAGAGTT GGTAGCTCTT 7141 GATCCGGCAA ACAAACCACC GCTGGTAGCG GTGGTTTTTT TGTTTGCAAG CAGCAGATTA 7201 CGCGCAGAAA AAAAGGATCT CAAGAAGATC CTTTGATCTT TTCTACGGGG TCTGACGCTC 7261 AGTGGAACGA AAACTCACGT TAAGGGATTT TGGTCATGAG ATTATCAAAA AGGATCTTCA 7321 CCTAGATCCT TTTAAATTAA AAATGAAGTT TTAAATCAAT CTAAAGTATA TATGAGTAAA 7381 CTTGGTCTGA CAGTTACCAA TGCTTAATCA GTGAGGCACC TATCTCAGCG ATCTGTCTAT 7441 TTCGTTCATC CATAGTTGCC TGACTCCCCG TCGTGTAGAT AACTACGATA CGGGAGGGCT 7501 TACCATCTGG CCCCAGTGCT GCAATGATAC CGCGAGACCC ACGCTCACCG GCTCCAGATT 7561 TATCAGCAAT AAACCAGCCA GCCGGAAGGG CCGAGCGCAG AAGTGGTCCT GCAACTTTAT 7621 CCGCCTCCAT CCAGTCTATT AATTGTTGCC GGGAAGCTAG AGTAAGTAGT TCGCCAGTTA 7681 ATAGTTTGCG CAACGTTGTT GCCATTGCTA CAGGCATCGT GGTGTCACGC TCGTCGTTTG 7741 GTATGGCTTC ATTCAGCTCC GGTTCCCAAC GATCAAGGCG AGTTACATGA TCCCCCATGT 7801 TGTGCAAAAA AGCGGTTAGC TCCTTCGGTC CTCCGATCGT TGTCAGAAGT AAGTTGGCCG 7861 CAGTGTTATC ACTCATGGTT ATGGCAGCAC TGCATAATTC TCTTACTGTC ATGCCATCCG 7921 TAAGATGCTT TTCTGTGACT GGTGAGTACT CAACCAAGTC ATTCTGAGAA TAGTGTATGC 7981 GGCGACCGAG TTGCTCTTGC CCGGCGTCAA TACGGGATAA TACCGCGCCA CATAGCAGAA 8041 CTTTAAAAGT GCTCATCATT GGAAAACGTT CTTCGGGGCG AAAACTCTCA AGGATCTTAC 8101 CGCTGTTGAG ATCCAGTTCG ATGTAACCCA CTCGTGCACC CAACTGATCT TCAGCATCTT 8161 TTACTTTCAC CAGCGTTTCT GGGTGAGCAA AAACAGGAAG GCAAAATGCC GCAAAAAAGG 8221 GAATAAGGGC GACACGGAAA TGTTGAATAC TCATACTCTT CCTTTTTCAA TATTATTGAA 8281 GCATTTATCA GGGTTATTGT CTCATGAGCG GATACATATT TGAATGTATT TAGAAAAATA 8341 AACAAATAGG GGTTCCGCGC ACATTTCCCC GAAAAGTGCC ACCTGACGTC GACGGATCGG 8401 GAGATCGATC TCCCGATCCC CTAGGGTCGA CTCTCAGTAC AATCTGCTCT GATGCCGCAT 8461 AGTTAAGCCA GTATCTGCTC CCTGCTTGTG TGTTGGAGGT CGCTGAGTAG TGCGCGAGCA 8521 AAATTTAAGC TACAACAAGG CAAGGCTTGA CCGACAATTG CATGAAGAAT CTGCTTAGGG 8581 TTAGGCGTTT TGCGCTGCTT CGCGATGTAC GGGCCAGATA TACGCGTTGA CATTGATTAT 8641 TGACTAGTTA TTAATAGTAA TCAATTACGG GGTCATTAGT TCATAGCCCA TATATGGAGT 8701 TCCGCGTTAC ATAACTTACG GTAAATGGCC CGCCTGGCTG ACCGCCCAAC GACCCCCGCC 8761 CATTGACGTC AATAATGACG TATGTTCCCA TAGTAACGCC AATAGGGACT TTCCATTGAC 8821 GTCAATGGGT GGAGTATTTA CGGTAAACTG CCCACTTGGC AGTACATCAA GTGTATC
[0117] In some embodiments, the cytidine base editor is selected from any one of the following: BE4 has the selected nucleic acid sequence.
[0118] Original BE4 nucleic acid sequence: ATGagctcagagactggcccagtggctgtggaccccacattgagacggcggatcgagccccatgagtttgaggtattctt cgatccgagagagctccgcaaggagacctgcctgctttacgaaattaattgggggggccggcactccatttggcgacata catcacagaacactaacaagcacgtcgaagtcaacttcatcgagaagttcacgacagaaagatatttctgtccgaacaca aggtgcagcattacctggtttctcagctggagccgcgaatgtagtagggccatcactgaattcctgtcaaggtatcccca cgtcactctgtttatttacatcgcaaggctgtaccaccacgctgacccccgcaatcgacaaggcctgcgggatttgatct cttcaggtgtgactatccaaattatgactgagcaggagtcaggatactgctggagaaactttgtgaattatagcccgagt aatgaagcccactggcctaggtatccccatctgtgggtacgactgtacgttcttgaactgtactgcatcatactgggcct gcctccttgtctcaacattctgagaaggaagcagccacagctgacattctttaccatcgctcttcagtcttgtcattacc agcgactgcccccacacattctctgggccaccgggttgaaatctggtggttcttctggtggttctagcggcagcgagact cccgggacctcagagtccgccacacccgaaagttctggtggttcttctggtggttctgataaaaagtattctattggttt agccatcggcactaattccgttggatgggctgtcataaccgatgaatacaaagtaccttcaaagaaatttaaggtgttgg ggaacacagaccgtcattcgattaaaaagaatcttatcggtgccctcctattcgatagtggcgaaacggcagaggcgact cgcctgaaacgaaccgctcggagaaggtatacacgtcgcaagaaccgaatatgttacttacaagaaatttttagcaatga gatggccaaagttgacgattctttctttcaccgtttggaagagtccttccttgtcgaagaggacaagaaacatgaacggc accccatctttggaaacatagtagatgaggtggcatatcatgaaaagtacccaacgatttatcacctcagaaaaaagcta gttgactcaactgataaagcggacctgaggttaatctacttggctcttgcccatatgataaagttccgtgggcactttct cattgagggtgatctaaatccggacaactcggatgtcgacaaactgttcatccagttagtacaaacctataatcagttgt ttgaagagaaccctataaatgcaagtggcgtggatgcgaaggctattcttagcgcccgcctctctaaatcccgacggcta gaaaacctgatcgcacaattacccggagagaagaaaaatgggttgttcggtaaccttatagcgctctcactaggcctgac accaaattttaagtcgaacttcgacttagctgaagatgccaaattgcagcttagtaaggacacgtacgatgacgatctcg acaatctactggcacaaattggagatcagtatgcggacttatttttggctgccaaaaaccttagcgatgcaatcctccta tctgacatactgagagttaatactgagattaccaaggcgccgttatccgcttcaatgatcaaaaggtacgatgaacatca ccaagacttgacacttctcaaggccctagtccgtcagcaactgcctgagaaaataaggaatattctttgatcagtcga aaaacgggtacgcaggttatattgacggcggagcgagtcaagaggaattctacaagtttatcaaacccatattagagaag atggatgggacgggaagagttgcttgtaaaactcaatcgcgaagatctactgcgaaagcagcggactttcgacaacggtag cattccacatcaaatccacttaggcgaattgcatgctatacttagaaggcaggaggattttatccgttcctcaaagaca atcgtgaaaagattgagaaaatcctaacctttcgcatacctactatgtgggacccctggcccgagggaactctcggttc gcatggatgacaagaaagtccgaagaacgattactccatggaattttgaggaagttgtcgataaaggtgcgtcagctca atcgttcatcgagaggatgaccaactttgacaagaatttaccgaacgaaaaagtattgcctaagcacagtttactttacg agtatttcacagtgtacaatgaactcacgaaagttaagtatgtcactgagggcatgcgtaaacccgcctttctaagcgga gaacagaaagcaatagtagatctgttattcaagaccaaccgcaaagtgacagttaagcaattgaaaggactactt tagaaaattgaatgcttcgattctgtcgagatctccggggtagagaatcgatttaatgcgtcacttggtacgtatcatg acctcctaaagataattaaagataaggacttcctggataacgaagagaatgaagatatcttagaagatatagtgttgact cttaccctctttgaagatcgggaaatgattgaggaaagactaaaaacatacgctcacctgttcgacgataaggttatgaa acagttaaagaggcgtcgctatacgggctggggacgattgtcgcggaaacttatcaacgggataagagacaagcaaagtg gtaaaactattctcgattttctaaagagcgacggcttcgccaataggaactttatgcagctgatccatgatgactcttta accttcaaagaggatatacaaaaggcacaggtttccggacaaggggactcattgcacgaacatattgcgaatcttgctgg ttcgccagccatcaaaaagggcatactccagacagtcaaagtagtggatgagctagttaaggtcatgggacgtcacaaac cggaaaacattgtaatcgagatggcacgcgaaaatcaaacgactcagaaggggcaaaaaaacagtcgagagcggatgaag agaatagaagagggtattaaagaactgggcagccagatcttaaaggagcatcctgtggaaaatacccaattgcagaacga gaaactttacctctattacctacaaaatggaagggacatgtatgttgatcaggaactggacataaaccgtttatctgatt acgacgtcgatcacattgtaccccaatcctttttgaaggacgattcaatcgacaataaagtgcttacacgctcggataag aaccgagggaaagtgacaatgttccaagcgaggaagtcgtaaagaaaatgaagaactattggcggcagctcctaaatgc gaaactgataacgcaaagaaagttcgataacttaactaaagctgagaggggtggcttgtctgaacttgacaaggccggat ttattaaacgtcagctcgtggaaacccgccaaatcacaaagcatgttgcacagatactagattcccgaatgaatacgaaa tacgacgagaacgataagctgattcgggaagtcaaagtaatcactttaaagtcaaaattggtgtcggacttcagaaagga ttttcaattctataaagttagggagataaataactaccaccatgcgcacgacgcttatcttaatgccgtcgtagggaccg cactcattaagaaatacccgaagctagaaagtgagtttgtgtatggtgattacaaagtttatgacgtccgtaagatgatc gcgaaaagcgaacaggagataggcaaggctacagccaaatacttcttttatctaacattatgaatttctttaagacgga aatcactctggcaaacggagagatacgcaaacgacctttaattgaaaccaatggggagacaggtgaaatcgtatgggata agggccgggacttcgcgacggtgagaaaagttttgtccatgccccaagtcaacatagtaaagaaaactgaggtgcagacc ggagggttttcaaaggaatcgattcttccaaaaaggaatagtgataagctcatcgctcgtaaaaaggactgggacccgaa aaagtacggtggcttcgatagccctacagttgcctattctgtcctagtagtggcaaaagttgagaagggaaaatccaaga aactgaagtcagtcaaagaattattggggataacgattatggagcgctcgtcttttgaaaagaaccccatcgacttcctt gaggcgaaaggttacaaggaagtaaaaaaggatctcataattaaactaccaaagtatagtctgtttgagttagaaaatgg ccgaaaacggatgttggctagcgccggagagcttcaaaaggggaacgaactcgcactaccgtctaaatacgtgaatttcc tgtatttagcgtcccattacgagaagttgaaaggttcacctgaagataacgaacagaagcaactttttgttgagcagcac aaacattatctcgacgaaatcatagagcaaatttcggaattcagtaagagagtcatcctagctgatgccaatctggacaa agtattaagcgcatacaacaagcacagggataaacccatacgtgagcaggcggaaaatattatccatttgtttactctta ccaacctcggcgctccagccgcattcaagtattttgacacaacgatagatcgcaaacgatacacttctaccaaggaggtg ctagacgcgacactgattcaccaatccatcacgggattatatgaaactcggatagatttgtcacagcttgggggtgactc tggtggttctggaggatctggtggttctactaatctgtcagatattattgaaaaggagaccggtaagcaactggttatcc aggaatccatcctcatgctcccagaggaggtggaagaagtcattgggaacaagccggaaagcgatatactcgtgcacacc gcctacgacgagagcaccgacgagaatgtcatgcttctgactagcgacgcccctgaatacaagccttgggctctggtcat acaggatagcaacggtgagaacaagattaagatgctctctggtggttctggaggatctggtggttctactaatctgtcag atattattgaaaaggagaccggtaagcaactggttatccaggaatccatcctcatgctcccagaggagtggaagaagtc attgggaacaagccggaaagcgatatactcgtgcacaccgcctacgacgagagcaccgacgagaatgtcatgcttctgac tagcgacgcccctgaatacaagccttgggctctggtcatacagtagcaacggtgagaacaagattaagatgctctctg gtggttctAAAAGGACGGCGGACGGATCAGAGTTCGAGAGTCCGAAAAAAAAACGAAAGGTCGAAtaa
[0119] BE4コドン optimization1 nucleic acid sequence: ATGTCATCCGAAACCGGGCCAGTGGCCGTAGACCAACACTCAGGAGGCGGATAGAACCCCATGAGTTTGAAGTGTTCTT CGACCCCAGAGAGCTGCGCAAAGAGACTTGCCTCCTGTATGAAATAAATTGGGGGGGTCGCCATTCAATTTGGAGGCACA CTAGCCAGAATACTAACAAACACGTGGAGGTAAATTTATCGAGAAGTTTACCACCGAAAGATACTTTTGCCCCAATACA CGGTGTTCAATTACCTGGTTTCTGTCATGGAGTCCATGTGGAGAATGTAGTAGAGCGATAACTGAGTTCCTGTCTCGATA TCCTCACGTCACGTTGTTTATATACATCGCTCGGCTTTATCACCATGCGGACCCGCGGAACAGGCAAGGTCTTCGGGACC TCATATCCTCTGGGGTGACCATCCAGATAATGACGGAGCAAGAGAGCGGATACTGCTGGCGAAACTTTGTTAACTACAGC CCAAGCAATGAGGCACACTGGCCTAGATATCCGCATCTCTGGGTTCGACTGTATGTCCTTGAACTGTACTGCATAATTCT GGGACTTCCGCCATGCTTGAACATTCTGCGGCGGAAACAACCACAGCTGACCTTTTTCACGATTGCTCTCCAAAGTTGTC ACTACCAGCGATTGCCACCCCACATCTTGTGGGCTACTGGACTCAAGTCTGGAGGAAGTTCAGGCGGAAGCAGCGGGTCT GAAACGCCCGGAACCTCAGAGAGCGCAACGCCCGAAAGCTCTGGAGGGTCAAGTGGTGGTAGTGATAAGAAATACTCCAT CGGCCTCGCCATCGGTACGAATTCTGTCGGTTGGGCCGTTATCACCGATGAGTACAAGGTCCCTTCTAAGAAATTCAAGG TTTTGGGCAATACAGACCGCCATTCTATAAAAAAAAACCTGATCGGCGCCCTTTTGTTTGACAGTGGTGAGACTGCTGAA GCGACTCGCCTGAAGCGAACTGCCAGGAGGCGGTATACGAGGCGAAAAAACCGAATTTGTTACCTCCAGGAGATTTTCTC AAATGAAATGGCCAAGGTAGATGATAGTTTTTTTCACCGCTTGGAAGAAAGTTTTCTCGTTGAGGAGGACAAAAAGCACG AGAGGCACCCAATCTTTGGCAACATAGTCGATGAGGTCGCATACCATGAGAAATATCCTACGATCTATCATCTCCGCAAG AAGCTGGTCGATAGCACGGATAAAGCTGACCTCCGGCTGATCTACCTTGCTCTTGCTCACATGATTAAATTCAGGGGCCA TTTCCTGATAGAAGGAGACCTCAATCCCGACAATTCTGATGTCGACAAACTGTTTATTCAGCTCGTTCAGACCTATAATC AACTCTTTGAGGAGAACCCCATCAATGCTTCAGGGGTGGACGCAAAGGCCATTTTGTCCGCGCGCTTGAGTAAATCACGA CGCCTCGAGAATTTGATAGCTCAACTGCCGGGTGAGAAGAAAAACGGGTTGTTTGGGAATCTCATAGCGTTGAGTTTGGG ACTTACGCCAAACTTTAAGTCTAACTTTGATTTGGCCGAAGATGCCAAATTGCAGCTGTCCAAAGATACCTATGATGACG ACTTGGATAACCTTCTTGCGCAGATTGGTGACCAATACGCGGATCTGTTTCTTGCCGCAAAAAATCTGTCCGACGCCATA CTCTTGTCCGATATACTGCGCGTCAATACTGAGATAACTAAGGCTCCCCTCAGCGCGTCCATGATTAAAAGATACGATGA GCACCACCAAGATCTCACTCTGTTGAAAGCCCTGGTTCGCCAGCAGCTTCCAGAGAAGTATAAGGAGATATTTTTCGACC AATCTAAAAACGGCTATGCGGGTTACATTGACGGTGGCGCCTCTCAAGAAGAATTCTACAAGTTTATAAAGCCGATACTT GAGAAAATGGACGGTACAGAGGAATTGTTGGTTAAGCTCAATCGCGAGGACTTGTTGAGAAAGCAGCGCACATTTGCAA TGGTAGTATTCCACACCAGATTCATCTGGGCGAGTTGCATGCCATTCTTAGAAGACAAGAAGATTTTTATCCGTTTCTGA AAGATAACAGAGAAAAAGATTGAAAAGATACTTACCTTTCGCATACCGTATTATGTAGGTCCCCTGGCTAGAGGGAACAGT CGCTTCGCTTGGATGACTCGAAAATCAGAAGAAAACAATAACCCCCTGGAATTTTGAAAGAAGTGGTAGATAAAGGTGCGAG TGCCCAATCTTTTATTGAGCGGATGACAAATTTTGACAAGAATCTGCCTAACGAAAAGGTGCTTCCCAAGCATTCCCTTT TGTATGAATACTTTACAGTATATAATGAACTGACTAAAGTGAAGTACGTTACCGAGGGGATGCGAAAGCCAGCTTTTTCTC AGTGGCGAGCAGAAAAAGCAATAGTTGACCTGCTGTTCAAGACGAATAGGAAGGTTACCGTCAAACAGCTCAAAGAAGA TTACTTTAAAAGATCGAATGTTTTGATTCAGTTGAGATAAGCGGAGTAGAGGATAGATTTAACCCAAGTCTTGGGAACTT ATCATGACCTTTTGAAGATCATCAAGGATAAAGATTTTTTGGGACAACGAGGAGAATGAAGATATCCTGGAAGATATAGTA CTTACCTTGACGCTTTTTGAAGATCGAGAGATGATCGAGGAGCGACTTAAGACGTACGCACATCTCTTTGACGATAAGGT TATGAAACAATTGAAACGCCGGCGGTATACTGGCTGGGGCAGGCTTTCTCGAAAGCTGATTAATGGTATCCGCGATAAGC AGTCTGGAAAGACAATCCTTGACTTTCTGAAAAGTGATGGATTTGCAAATAGAAACTTTATGCAGCTTATACATGATGAC TCTTTGACGTTCAAGGAAGACATCCAGAAGGCACAGGTATCCGGCCAAGGGGATAGCCTCCATGAACACATAGCCAACCT GGCCGGCTCACCAGCTATTAAAAAGGGAATATTGCAAACCGTTAAGGTTGTTGACGAACTCGTTAAGGTTATGGGCCGAC ACAAACCAGAGAATATCGTGATTGAGATGGCTAGGGAGAATCAGACCACTCAAAAAGGTCAGAAAAATTCTCGCGAAAGG ATGAAGCGAATTGAAGAGGGAATCAAAGAACTTGGCTCTCAAATTTTGAAAGAGCACCCGGTAGAAAACACTCAGCTGCA GAATGAAAAGCTGTATCTGTATTATCTGCAGAATGGTCGAGATATGTACGTTGATCAGGAGCTGGATATCAATAGGCTCA GTGACTACGATGTCGACCACATCGTTCCTCAATCTTTCCTGAAAGATGACTCTATCGACAACAAAGTGTTGACGCGATCA GATAAGAACCGGGGAAAATCCGACAATGTACCCTCAGAAGAAGTTGTCAAGAAGATGAAAAACTATTGGAGACAATTGCT GAACGCCAAGCTCATAACACAACGCAAGTTCGATAACTTGACGAAAGCCGAAAGAGGTGGGTTGTCAGAATTGGACAAAG CTGGCTTTATTAAGCGCCAATTGGTGGAGACCCGGCAGATTACGAAACACGTAGCACAAATTTTGGATTCACGAATGAAT ACCAAATACGACGAAAACGACAAATTGATACGCGAGGTGAAAGTGATTACGCTTAAGAGTAAGTTGGTTTCCGATTTCAG GAAGGATTTTCAGTTTTACAAAGTAAGAGAAATAAACAACTACCACCACGCCCATGATGCTTACCTCAACGCGGTAGTTG GCACAGCTCTTATCAAAAAATATCCAAAGCTGGAAAGCGAGTTCGTTTACGGTGACTATAAAGTATACGACGTTCGGAAG ATGATAGCCAAATCAGAGCAGGAAATTGGGAAGGCAACCGCAAAATACTTCTTCTATTCAAACATCATGAACTTCTTTAA GACGGAGATTACGCTCGCGAACGGCGAAATACGCAAGAGGCCCCTCATAGAGACTAACGGCGAAACCGGGGAGATCGTAT GGGACAAAGGACGGGACTTTGCGACCGTTAGAAAAGTACTTTCAATGCCACAAGTGAATATTGTTAAAAAGACAGAAGTA CAAACAGGGGGGTTCAGTAAGGAATCCATTTTGCCCAAGCGGAACAGTGATAAATTGATAGCAAGGAAAAAAGATTGGGA CCCTAAGAAGTACGGTGGTTTCGACTCTCCTACCGTTGCATATTCAGTCCTTGTAGTTGCGAAAGTGGAAAAGGGGAAAA GTAAGAAGCTTAAGAGTGTTAAAGAGCTTCTGGGCATAACCATAATGGAACGGTCTAGCTTCGAGAAAAATCCAATTGAC TTTCTCGAGGCTAAAGGTTACAAGGAGGTAAAAAAGGACCTGATAATTAAACTCCCAAAGTACAGTCTCTTCGAGTTGGA GAATGGGAGGAAGAGAATGTTGGCATCTGCAGGGGAGCTCCAAAAGGGGAACGAGCTGGCTCTGCCTTCAAAATACGTGA ACTTTCTGTACCTGGCCAGCCACTACGAGAAACTCAAGGGTTCTCCTGAGGATAACGAGCAGAAACAGCTGTTTGTAGAG CAGCACAAGCATTACCTGGACGAGATAATTGAGCAAATTAGTGAGTTCTCAAAAAGAGTAATCCTTGCAGACGCGAATCT GGATAAAGTTCTTTCCGCCTATAATAAGCACCGGGACAAGCCTATACGAGAACAAGCCGAGAACATCATTCACCTCTTTA CCCTTACTAATCTGGGCGCGCCGGCCGCCTTCAAATACTTCGACACCACGATAGACAGGAAAAGGTATACGAGTACCAAA GAAGTACTTGACGCCACTCTCATCCACCAGTCTATAACAGGGTTGTACGAAACGAGGATAGATTTGTCCCAGCTCGGCGG CGACTCAGGAGGGTCAGGCGGCTCCGGTGGATCAACGAATCTTTCCGACATAATCGAGAAAGAAACCGGCAAACAGTTGG TGATCCAAGAATCAATCCTGATGCTGCCTGAAGAAGTAGAAGAGGTGATTGGCAACAAACCTGAGTCTGACATTCTTGTC CACACCGCGTATGACGAGAGCACGGACGAGAACGTTATGCTTCTCACTAGCGACGCCCCTGAGTATAAACCATGGGCGCT GGTCATCCAAGATTCCAATGGGGAAAACAAGATTAAGATGCTTAGTGGTGGGTCTGGAGGGAGCGGTGGGTCCACGAACC TCAGCGACATTATTGAAAAAAGACTGGTAAACAACTTGTAATACAAGAGTCTATTCTGATGTTGCCTGAAGAGGTGGAG GAGGTGATTGGGAACAAACCGGAGTCTGATATACTTGTTCATACCGCCTATGACGAATCTACTGATGAGAATGTGATGCT TTTaACGTCAGACGCTCCCGAGTACAAACCCTGGGCTCTGGTGATTCAGGACAGCAATGGTGAGAATAAGATTAAATGT TGAGTGGGGGCTCAAAGCGCACGGCTGACGGTAGCGAATTTGAGAGCCCCAAAAAAAAACGAAAGGTCGAAtaa
[0120] BE4コドン optimization2 nucleic acid sequence: ATGAGCAGCGAGACAGGCCCTGTGGCTGTGGATCCTACACTGCGGAGAAGAATCGAGCCCCACGAGTTCGAGGTGTTCTT CGACCCCAGAGAGCTGCGGAAAGAGACATGCCTGCTGTACGAGATCAACTGGGGCGGCAGACACTCTATCTGGCGCACA CAAGCCAGAACACCAACAAGCACGTGGAAGTGAACTTTATCGAGAAGTTTACGACCGAGCGGTACTTCTGCCCCAACACC AGATGCAGCATCACCTGGTTCTGAGCTGGTCCCCTTGCGGCGAGTGCAGCAGAGCCATCACCGAGTTTCTGTCCAGATA TCCCCACGTGACCCTGTTCATCTATATCGCCCGGCTGTACCACCACGCCGATCCTAGAAATAGACAGGGACTGCGCGACC TGATCAGCAGCGGAGTGACCATCCAGATCATGACCGAGCAAGAGAGCGGCTACTGCTGGCGGAACTTCGTGAACTACAGC CCCAGCAACGAAGCCCACTGGCCTAGATATCCTCACCTGTGGGTCCGACTGTACGTGCTGGAACTGTACTGCATCATCCT GGGCCTGCCTCCATGCCTGAACATCCTGAGAAGAAAGCAGCCTCAGCTGACCTTCTTCACAATCGCCCTGCAGAGCTGCC ACTACCAGAGACTGCCTCCACACATCCTGTGGGCCACCGGACTTAAGAGCGGAGGATCTAGCGGCGGCTCTAGCGGATCT GAGACACCTGGCACAAGCGAGTCTGCCACACCTGAGAGTAGCGGCGGATCTTCTGGCGGCTCCGACAAGAAGTACTCTAT CGGACTGGCCATCGGCACCAACTCTGTTGGATGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGG TGCTGGGCAACACCGACCGGCACAGCATCAAGAAGAATCTGATCGGCGCCCTGCTGTTCGACTCTGGCGAAACAGCCGAA GCCACCAGACTGAAGAGAACCGCCAGGCGGAGATACACCCGGCGGAAGAACCGGATCTGCTACCTGCAAGAGATCTTCAG CAACGAGATGGCCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGTCCTTCCTGGTGGAAGAGGACAAGAAGCACG AGCGGCACCCCATCTTCGGCAACATCGTGGATGAGGTGGCCTACCACGAGAAGTACCCCACCATCTACCACCTGAGAAAG AAACTGGTGGACAGCACCGACAAGGCCGACCTGAGACTGATCTACCTGGCTCTGGCCCACATGATCAAGTTCCGGGGCCA CTTTCTGATCGAGGGCGATCTGAACCCCGACAACAGCGACGTGGACAAGCTGTTCATCCAGCTGGTGCAGACCTACAACC AGCTGTTCGAGGAAAACCCCATCAACGCCTCTGGCGTGGACGCCAAGGCTATCCTGTCTGCCAGACTGAGCAAGAGCAGA AGGCTGGAAAACCTGATCGCCCAGCTGCCTGGCGAGAAGAAGAATGGCCTGTTCGGCAACCTGATTGCCCTGAGCCTGGG ACTGACCCCTAACTTCAAGAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCAGCTGAGCAAGGACACCTACGACGACG ACCTGGACAATCTGCTGGCCCAGATCGGCGATCAGTACGCCGACTTGTTTCTGGCCGCCAAGAACCTGTCCGACGCCATC CTGCTGAGCGATATCCTGAGAGTGAACACCGAGATCACAAAGGCCCCTCTGAGCGCCTCTATGATCAAGAGATACGACGA GCACCACCAGGATCTGACCCTGCTGAAGGCCCTCGTTAGACAGCAGCTGCCAGAGAAGTACAAAGAGATTTTCTTCGATC AGTCCAAGAACGGCTACGCCGGCTACATTGATGGCGGAGCCAGCCAAGAGGAATTCTACAAGTTCATCAAGCCCATCCTG GAAAAGATGGACGGCACCGAGGAACTGCTGGTCAAGCTGAACAGAGAGGACCTGCTGCGGAAGCAGCGGACCTTCGACAA TGGCTCTATCCCTCACCAGATCCACCTGGGAGAGCTGCACGCCATTCTGCGGAGACAAGAGGACTTTTACCCATTCCTGA AGGACAACCGGGAAAAGATCGAGAAGATCCTGACCTTCAGGATCCCCTACTACGTGGGACCACTGGCCAGAGGCAATAGC AGATTCGCCTGGATGACCAGAAAGAGCGAGGAAACCATCACACCCTGGAACTTCGAGGAAGTGGTGGACAAGGGCGCCAG CGCTCAGTCCTTCATCGAGCGGATGACCAACTTCGATAAGAACCTGCCTAACGAGAAGGTGCTGCCCAAGCACTCCCTGC TGTATGAGTACTTCACCGTGTACAACGAGCTGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCCGCCTTTCTG AGCGGCGAGCAGAAAAAGGCCATTGTGGATCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGA CTACTTCAAGAAAATCGAGTGCTTCGACAGCGTGGAAATCAGCGGCGTGGAAGATCGGTTCAATGCCAGCCTGGGCACAT ACCACGACCTGCTGAAAATTATCAAGGACAAGGACTTCCTGGACAACGAAGAGAACGAGGACATTCTCGAGGACATCGTG CTGACCCTGACACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACATACGCCCACCTGTTCGACGACAAAGT GATGAAGCAACTGAAGCGGAGGCGGTACACAGGCTGGGGCAGACTGTCTCGGAAGCTGATCAACGGCATCCGGGATAAGC AGTCCGGCAAGACAATCCTGGATTTCCTGAAGTCCGACGGCTTCGCCAACAGAAACTTCATGCAGCTGATCCACGACGAC AGCCTGACCTTTAAAGAGGACATCCAGAAAGCCCAGGTGTCCGGCCAAGGCGATTCTCTGCACGAGCACATTGCCAACCT GGCCGGATCTCCCGCCATTAAGAAGGGCATCCTGCAGACAGTGAAGGTGGTGGACGAGCTTGTGAAAGTGATGGGCAGAC ACAAGCCCGAGAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACACAGAAGGGCCAGAAGAACAGCCGCGAGAGA ATGAAGCGGATCGAAGAGGGCATCAAAGAGCTGGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAAACACCCAGCTGCA GAACGAGAAGCTGTACCTGTACTACCTGCAGAATGGACGGGATATGTACGTGGACCAAGAGCTGGACATCAACCGGCTGA GCGACTACGATGTGGACCATATCGTGCCCCAGAGCTTTCTGAAGGACGACTCCATCGATAACAAGGTCCTGACCAGAAGC GACAAGAACCGGGGCAAGAGCGATAACGTGCCCTCCGAAGAGGTGGTCAAGAAGATGAAGAACTACTGGCGACAGCTGCT GAACGCCAAGCTGATTACCCAGCGGAAGTTCGATAACCTGACCAAGGCCGAGAGAGGCGGCCTGAGCGAACTTGATAAGG CCGGCTTCATTAAGCGGCAGCTGGTGGAAACCCGGCAGATCACCAAACACGTGGCACAGATTCTGGACTCCCGGATGAAC ACTAAGTACGACGAGAATGACAAGCTGATCCGGGAAGTGAAAGTCATCACCCTGAAGTCTAAGCTGGTGTCCGATTTCCG GAAGGATTTCCAGTTCTACAAAGTGCGGGAAATCAACAACTACCATCACGCCCACGACGCCTACCTGAATGCCGTTGTTG GAACAGCCCTGATCAAGAAGTATCCCAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACAAGGTGTACGACGTGCGGAAG ATGATCGCCAAGAGCGAACAAGAGATCGGCAAGGCTACCGCCAAGTACTTTTTCTACAGCAACATCATGAACTTTTTCAA GACAGAGATCACCCTGGCCAACGGCGAGATCCGGAAAAGACCCCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGT GGGATAAGGGCAGAGATTTTGCCACAGTGCGGAAAGTGCTGAGCATGCCCCAAGTGAATATCGTGAAGAAAACCGAGGTG CAGACAGGCGGCTTCAGCAAAGAGTCTATCCTGCCTAAGCGGAACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGGGA CCCTAAGAAGTACGGCGGCTTCGATAGCCCTACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGGGCAAGT CCAAAAAGCTCAAGAGCGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAGAAGCAGCTTTGAGAAGAACCCGATCGAC TTTCTGGAAGCCAAGGGCTACAAAGAAGTCAAGAAGGACCTCATCATCAAGCTCCCCAAGTACAGCCTGTTCGAGCTGGA AAATGGCCGGAAGCGGATGCTGGCCTCAGCAGGCGAACTGCAGAAAGGCAATGAACTGGCCCTGCCTAGCAAATACGTCA ACTTCCTGTACCTGGCCAGCCACTATGAGAAGCTGAAGGGCAGCCCCGAGGACAATGAGCAAAAGCAGCTGTTTGTGGAA CAGCACAAGCACTACCTGGACGAGATCATCGAGCAGATCAGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAACCT GGATAAGGTGCTGTCTGCCTATAACAAGCACCGGGACAAGCCTATCAGAGAGCAGGCCGAGAATATCATCCACCTGTTTA CCCTGACCAACCTGGGAGCCCCTGCCGCCTTCAAGTACTTCGACACCACCATCGACCGGAAGAGGTACACCAGCACCAAA GAGGTGCTGGACGCCACACTGATCCACCAGTCTATCACCGGCCTGTACGAAACCCGGATCGACCTGTCTCAGCTCGGCGG CGATTCTGGTGGTTCTGGCGGAAGTGGCGGATCCACCAATCTGAGCGACATCATCGAAAAAGAGACAGGCAAGCAGCTCG TGATCCAAGAATCCATCCTGATGCTGCCTGAAGAGGTTGAGGAAGTGATCGGCAACAAGCCTGAGTCCGACATCCTGGTG CACACCGCCTACGATGAGAGCACCGATGAGAACGTCATGCTGCTGACAAGCGACGCCCCTGAGTACAAGCCTTGGGCTCT CGTGATTCAGGACAGCAATGGGGAGAACAAGATCAAGATGCTGAGCGGAGGTAGCGGAGGCAGTGGCGGAAGCACAAACC TGTCTGATATCATTGAAAAAGAAACCGGGAAGCAACTGGTCATTCAAGAGTCCATTCTCATGCTCCCGGAAGAAGTCGAG GAAGTCATTGGAAACAAACCCGAGAGCGATATTCTGGTCCACACAGCCTATGACGAGTCTACAGACGAAAACGTGATGCT CCTGACCTCTGACGCTCCCGAGTATAAGCCCTGGGCACTTGTTATCCAGGACTCTAACGGGGAAAACAAAATCAAAATGT TGTCCGGCGGCAGCAAGCGGACAGCCGATGGATCTGAGTTCGAGAGCCCCAAGAAGAAACGGAAGGTgGAGtaa
[0121] "Base editing activity" refers to the ability to chemically change bases within a polynucleotide. In one embodiment, the first base is converted to the second base. In this study, the base editing activity is a cytidine deaminase activity, e.g., targeting C·G to T·A. In another embodiment, the base editing activity is an activity that converts adenosine or adenine. In another embodiment, the activity is a deaminase activity, for example, an activity that converts A·T to G·C. base editing activity is mediated by cytidine deaminase activity (e.g., converting target C·G to T·A) and and adenosine or adenine deaminase activity (e.g., converting A·T to G·C). do.
[0122] The term "base editing system" refers to a system for editing nucleic acid bases in a target nucleotide sequence. In various embodiments, the base editor system comprises: (1) a polynucleotide sequence; Nucleotide-programmable nucleotide-binding domains (e.g., Cas9); ) one or more deaminase domains for deaminating the nucleobases (e.g. , adenosine deaminase and / or cytidine deaminase); (3) one or more In some embodiments, the guide polynucleotide (e.g., guide RNA) The base editor (BE) system (1) deamidates nucleobases in a target nucleotide sequence. Polynucleotide-programmable nucleotide binding domains (e.g., Cas9), adenosine deaminase domain and cytidine deaminase domain; and and (2) combining a polynucleotide with a programmable nucleotide binding domain. The nucleic acid sequence includes one or more guide polynucleotides (e.g., guide RNAs). In an embodiment, the polynucleotide programmable nucleotide binding domain comprises: In some embodiments, the polynucleotide is a programmable DNA binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor system is BE4. The base editor is an adenine or adenosine base editor (ABE). In this state, base editors are adenine or adenosine base editors (ABEs) and In some embodiments, the base editor is a cytidine base editor (CBE). , an abasic editor.
[0123] In some embodiments, the base editor system comprises two or more base editing components. For example, the base editor system can include one or more deaminases (e.g., In some embodiments, the enzymes may include adenosine deaminase, adenosine deaminase, and cytidine deaminase. In this method, a single guide polynucleotide is used to target different deaminases to a target nucleic acid sequence. In some embodiments, the guide polynucleoside can be targeted to A single pair of deaminases can be used to target different deaminases to a target nucleic acid sequence. This can be done.
[0124] Deaminase domains of base editor systems and polynucleotide programmability The nucleotide-binding moieties may be covalently or non-covalently bound to one another or by any of the bonds. The molecules may be linked by any combination of binding and interaction. wherein one or more deaminase domains are linked to a polynucleotide programmable nucleic acid. The nucleotide binding domain can be targeted to a target nucleotide sequence. In embodiments, the polynucleotide programmable nucleotide binding domain comprises: It may be fused or linked to one or more deaminase domains. Thus, the polynucleotide programmable nucleotide binding domain binds to the deaminase By non-covalently interacting with or binding to the domain, one or more deaminators The nucleotide domain can be targeted to a target nucleotide sequence. In some embodiments, the deaminase domain is a polynucleotide programmable interacting with additional heterologous moieties or domains that are part of the nucleotide-binding domain. It may contain additional heterologous moieties or domains with which it is capable of forming complexes, associations, or complexes. In some embodiments, the additional heterologous moiety is conjugated to the polypeptide. In some embodiments, the phospholipids may interact, associate, or form complexes. wherein the additional heterologous moiety binds to, interacts with, associates with, or complexes with the polynucleotide. In some embodiments, the additional heterologous moiety can form a combination of a guide In some embodiments, the additional heterologous moiety can be linked to a polynucleotide. In some embodiments, the additional moiety can be attached to a polypeptide linker. Additional heterologous moieties can be attached to the polynucleotide linker. , protein domain. In some embodiments, the additional heterologous moiety , K homology (KH) domain, MS2 coat protein domain, PP7 coat protein domain, SfMu Com coat protein domain, steryl α motif, telomerase Ku binding motif and Ku protein, telomerase Sm7 binding motif and Sm7 protein, or RNA recognition It could be a recognition motif.
[0125] The base editor system can further comprise a guide polynucleotide component. The components of the base editor system may be linked by covalent bonds, non-covalent interactions, or both. It should be understood that the compounds may be coupled to one another via any combination of bonds and interactions. In some embodiments, the one or more deaminase domains are The target nucleotide sequence can be targeted by a In this embodiment, the deaminase domain is a portion or segment of the guide polynucleotide. capable of interacting with, binding to, or forming a complex with (e.g., a polynucleotide motif) Further heterologous moieties or domains (e.g., polynucleotides such as RNA or DNA binding proteins) may be In some embodiments, additional heterologous moieties or domains (e.g., polynucleotide-binding domains of RNA- or DNA-binding proteins) In some embodiments, the additional The heterologous moiety binds to, interacts with, associates with, or complexes with the polypeptide. In some embodiments, the additional heterologous moiety can form a polynucleotide. It binds to, interacts with, associates with, or forms a complex with a polynucleotide. In some embodiments, the additional heterologous moiety can be a guide polynucleotide. In some embodiments, the additional heterologous moiety can be linked to a polypeptide. In some embodiments, the additional heterologous moiety can be linked to a linker. The additional heterologous moiety can be attached to a polynucleotide linker. In some embodiments, the additional heterologous moiety may be a K homology (KH) domain. Main, MS2 coat protein domain, PP7 coat protein domain, SfMu Com coat Protein domains, sterile alpha motif, telomerase Ku binding motif and Ku tan Protein, telomerase Sm7 binding motif and Sm7 protein, or RNA recognition motif could be.
[0126] In one embodiment, the base editor system comprises an inhibitor of base excision repair (BER). The components of the base editor system can further include a shared Covalent bonds, non-covalent interactions, or any combination of these bonds and interactions. It should be understood that the BER components can be linked to each other via a BER inhibitor. In one embodiment, the inhibitor of BER is uracil DNA glycosylase In one embodiment, the inhibitor of BER can be an inosine BER inhibitor (UGI). In one embodiment, the inhibitor of BER is a polynucleotide protease inhibitor. Targeting to target nucleotide sequences via a rammable nucleotide-binding domain In some embodiments, the polynucleotide may be a programmable nucleotide. The binding domain may be fused or linked to an inhibitor of BER. The programmable nucleotide-binding domain comprises one or more deamina In some embodiments, the ribozyme domain may be fused or linked to an inhibitor of BER. In this study, polynucleotide-programmable nucleotide-binding domains inhibited BER. B by noncovalently interacting with factors or by associating with inhibitors of BER. Inhibitors of ER can be targeted to target nucleotide sequences. In some embodiments, the inhibitor of the BER component is a polynucleotide programmable interacting with an additional heterologous moiety or domain that is part of the nucleotide-binding domain The polypeptide may comprise additional heterologous moieties or domains that are capable of binding, associating, or complexing with the polypeptide.
[0127] In one embodiment, the inhibitor of BER is a nucleotide sequence that is mediated by a guide polynucleotide to target nucleotides. For example, in some embodiments, the inhibitor may be targeted to a nucleotide sequence that inhibits BER. The deleterious agent may be a portion or segment of a guide polynucleotide (e.g., a polynucleotide Further heterologous moieties or domains (e.g., nucleotides) that can interact with, associate with, or complex with the nucleotide motif (e.g., nucleotides) For example, a polynucleotide binding domain such as an RNA or DNA binding protein. In some embodiments, the guide polynucleotide may further comprise a heterologous moiety or a The main (e.g., polynucleotide-binding domain, such as an RNA- or DNA-binding protein) In some embodiments, the additional heterologous nucleotide may be fused or linked to an inhibitor of BER. The moiety is capable of binding, interacting, associating, or complexing with a polynucleotide In some embodiments, the additional heterologous moiety is attached to the guide polynucleotide. In some embodiments, the additional heterologous moiety can be a polypeptide linker. In some embodiments, the additional heterologous moiety can be linked to a polynucleotide. The additional heterologous moiety can be a protein domain or a nucleotide linker. In some embodiments, the additional heterologous moiety may be a K homology (KH) domain, MS2 Coat protein domain, PP7 coat protein domain, SfMu Com coat protein domain, sterile alpha motif, telomerase Ku binding motif and Ku protein, Te It may be a chromosome enzyme Sm7 binding motif and Sm7 protein, or an RNA recognition motif.
[0128] The term "Cas9" or "Cas9 domain" refers to a Cas9 protein or a fragment thereof (e.g., Cas9 Active, inactive, or partially active DNA cleavage domains of Cas9 and / or gRNA binding of Cas9 Cas9 nuclease refers to an RNA-guided nuclease containing a c domain. asnl nuclease or CRISPR (clustered regularly interspaced short palindromic CRISPR is also known as a repeat-associated nuclease. The adaptive immune system provides defense against transposable elements (transposable elements, conjugative plasmids). CRIS comprises a spacer, a sequence complementary to the preceding mobile element, and a target invader nucleic acid. The PR cluster is transcribed and processed into CRISPR RNA (crRNA). In the endothelial cell line, the correct processing of pre-crRNA is mediated by a small trans-coding RNA (tracrRNA), an endogenous RNA. tracrRNA requires ribonuclease 3 (rnc) and Cas9 protein. This guides the processing of pre-crRNA by Cas9 / crRNA / tracrRNA. A endonucleolytically cleaves a linear or circular dsDNA target complementary to the spacer. The target strand that is not complementary to the crRNA is first endonucleolytically cleaved and then exonucleolytically cleaved. In nature, DNA binding and cleavage are carried out by proteins. However, both crRNA and tracrRNA aspects are typically required. A single guide RNA ("sgRNA," or simply "gRNA") is engineered to incorporate into a single RNA species. See, for example, Jinek M. et al., Science 337:816-821 (2012). Cas9 is a CRISPR repeat sequence. It recognizes a short motif (PAM or protospacer adjacent motif) in the spacer to distinguish self from non-self. The sequence and structure of Cas9 nuclease are well known to those skilled in the art. (e.g., "Complete genome sequence of an M1 strain of Streptococcus pyog" Ferretti et al., Proc. Natl. Acad. Sci. USA 98:4658-4663(2001); SPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltc heva E. et al., Nature 471:602-607(2011); and “A programmable dual-RNA-guide d DNA endonuclease in adaptive bacterial immunity.” Jinek M. et al., Science 33 7:816-821 (2012), the entire contents of which are incorporated herein by reference. Orthologs include, but are not limited to, S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences are described in the present disclosure. Such Cas9 nucleases and sequences will be apparent to those skilled in the art based on the disclosures of Chylinsk i, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas Organisms and genes disclosed in “immunity systems” (2013) RNA Biology 10:5, 726-737 The Cas9 sequences from the locus are included, the entire contents of which are incorporated herein by reference.
[0129] An exemplary Cas9 is Streptococcus pyogenes Cas9 (spCas9), the amino acid sequence of which is shown below. The following information is provided. JPEG0007779988000002.jpg173168 (single underline: HNH domain; double underline: RuvC domain)
[0130] The nuclease-inactivated Cas9 protein is interchangeably referred to as the “dCas9” protein (nuclease-“ It may also be referred to as "dead" Cas9 or catalytically inactive Cas9. Methods for generating a Cas9 protein (or a fragment thereof) having the following structure are known (see, e.g., Jinek et al. al, Science. 337:816-821(2012); Qi et al, “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression”(2013) Cell. 28; 152 (5): 1173-83, the contents of each of which are incorporated herein by reference. For example, the DN of Cas9 The A cleavage domain is composed of two subdomains: the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA and binds to RuvC1. The subdomains cleave the non-complementary strand. Mutations within these subdomains allow Cas9 to For example, mutations D10A and H840A can inhibit the nuclease activity of S. pyogenes Cas9. completely inactivates the ATPase activity (Jinek et al., Science. 337:816-821(2012); Qi et al. l, Cell. 28;152(5): 1173-83 (2013)). In some embodiments, the Cas9 nuclease is inactivated. The Cas9 has an active (e.g., inactivated) DNA cleavage domain, i.e., the Cas9 is referred to as an "nCas9" protein. The nickase is called Cas9 (meaning "nickase"). In some embodiments, Cas9 For example, in some embodiments, proteins comprising fragments of the protein The gene contains one of two Cas9 domains: (1) the gRNA-binding domain of Cas9; (2) the C In some embodiments, a protein comprising Cas9 or a fragment thereof is Cas9 variants are those that share homology with Cas9 or its fragments. For example, a Cas9 variant may have at least about 70% identity to wild-type Cas9, at least about 80% identity to wild-type Cas9, or at least about 80% identity to wild-type Cas9. 0% identity, at least about 90% identity, at least about 95% identity, at least about 96% Identity, at least about 97% identity, at least about 98% identity, at least about 99% identity have at least about 99.5% identity, or at least about 99.9% identity. In some embodiments, the Cas9 mutant has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 2 9, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 4 In some embodiments, the Cas9 vector may have 9, 50, or more amino acid changes. The variant contains a fragment of Cas9 (e.g., the gRNA binding domain or the DNA cleavage domain) and The fragments have at least about 70% identity to the corresponding fragment of wild-type Cas9, and at least about 80% identity to the corresponding fragment of wild-type Cas9. has at least about 90% identity, has at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity have at least about 99% identity, have at least about 99.5% identity, and In some embodiments, the fragment has at least about 99.9% identity to the corresponding wild-type Cas At least 30%, at least 35%, at least 40%, at least 45%, at least 9 amino acids in length at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least At least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, is 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% .
[0131] In some embodiments, the fragment is at least 100 amino acids in length. In the above, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 60 0, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or It is at least 1300 amino acids in length.
[0132] In one embodiment, the wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes. (NCBI reference sequence: NC_17053.1, nucleotide and amino acid sequences are as follows):
[0133] ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGATTATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG GCAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGCAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAATCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTAGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAGAAATGGCTTGTTTGGGAAT CTCATTGCTTTGTCATTGGGATTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAATTAACAGCTTTC AAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATAGTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAGCGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAGGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGCGCCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGGGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGATATTCAAAAAGCACAGGTGTCTGGACAAGGCCATAGTTTA CATGAACAGATTGCTAACTTAGCTGGCAGTCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAATTGTTGATGAACT GGTCAAAGTAATGGGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCCAGA AAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCTGTT GAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTACAAAATGGAAGAGACATGTATGTGGACCAAGAATT AGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCATTAAAGACGATTCAATAGACAATA AGGTACTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAAAAC TATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGGTTT GAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAATTT TGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCTAAA TTAGTTTCTGACTTCCGAAAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGCGTA TCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATAAAG TTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCTAAT ATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGGGGA AACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATATTG TCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAAAGAAATTCGGACAAGCTTATTGCT CGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGCTAA GGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCTTTG AAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAATAT AGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGCTCT GCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAACAAA AACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTTATT TTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGAAAA TATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTAAAC GATATACGTCTACAAAAGAAGTTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATTGAT TTGAGTCAGCTAGGAGGTGACTGA JPEG0007779988000003.jpg172166 (single underline: HNH domain; double underline: RuvC domain)
[0134] In one embodiment, the wild-type Cas9 has the following nucleotide and / or amino acid sequence: Corresponding to or containing columns: ATGGATAAAAAGTATTCTATTGGTTTAGACATCGGCACTAATTCCGTTGGATGGGCTGTCATAACCGATGAATACAAAGT ACCTTCAAAGAAATTTAAGGTGTTGGGGAACACAGACCGTCATTCGATTAAAAAAGAATCTTATCGGTGCCCTCCTATTCG ATAGTGGCGAAACGGCAGAGGCGACTCGCCTGAAACGAACCGCTCGGAGAAGGTATACACGTCGCAAGAACCGAATATGT TACTTACAAGAAATTTTTAGCAATGAGATGGCCAAAGTTGACGATTCTTTCTTTCACCGTTTGGAAGAGTCCTTCCTTGT CGAAGAGGACAAGAAACATGAACGGCACCCCATCTTTGGAAACATAGTAGATGAGGTGGCATATCATGAAAAGTACCCAA CGATTTATCACCTCAGAAAAAAGCTAGTTGACTCAACTGATAAAGCGGACCTGAGGTTAATCTACTTGGCTCTTGCCCAT ATGATAAAGTTCCGTGGGCACTTTCCATTGAGGGTGATCTAAATCCGGACAACTCGGATGTCGACAAACTGTTCATCCA GTTAGTACAAACCTATAATCAGTTGTTTGAAGAACCCTATAAATGCAAGTGGCGTGGATGCGAAGGCTATTCTTAGCG CCCGCCTCTCTAAATCCCGACGGCTAGAAAAACCTGATCGCACAATTACCCGGAGAGAAAAAAATGGTTGTTCGGTAAC CTTATAGCGCTCTCACTAGGCCTGACACCAAATTTAAGTCGAACTTCGACTTAGCTGAAGATGCCAAATTGCAGCTTAG TAAGGACACGTACGATGACGATCTCGCAATCTACTGGCACAAATTGGAGATCAGTATGCGGACTTATTTTTGGCTGCCA AAAACCTTAGCGATGCAATCCTCCTATCTGACATACTGAGAGTTAATACTGAGATTACCAAGGCGCCGTTATCCCGCTTCA ATGATCAAAAGGTACGATGAACATCACCAAGACTTGACACTTCTCAAGGCCCTAGTCCGTCAGCAACTGCCTGAGAAATA TAAGGAAATATTCTTTGATCAGTCGAAAACGGGTACGCAGGTTATATTGACGGGCGGAGCGAGTCAAGAGGAATTCTACA AGTTTATCAAACCCATATTAGAGAATGGATGGGACGGAAGAGTTGCTTGTAAAACTCAATCGCGAAGATCTACTGCGA AAGCAGCGGACTTTCGACAACGGTAGCATTCCACATCAAATCCACTTAGGCGAATTGCATGCTATACTTAGAAGGCAGGA GGATTTTTATCCGTTCCTCAAAGACAATCGTGAAAAGATTGAGAAAATCCTAACCTTTCGCATACCTTACTATGTGGGAC CCCTGGCCCGAGGGAACTCTCGGTTCGCATGGATGACAAGAAAGTCCGAAGAAACGATTACTCCATGGAATTTTGAGGAA GTTGTCGATAAAGGTGCGTCAGCTCAATCGTTCATCGAGAGGATGACCAACTTTGACAAGAATTTACCGAACGAAAAAGT ATTGCCTAAGCACAGTTTACTTTACGAGTATTTCACAGTGTACAATGAACTCACGAAAGTTAAGTATGTCACTGAGGGCA TGCGTAAACCCGCCTTTCTAAGCGGAGAACAGAAGAAAGCAATAGTAGATCTGTTATTCAAGACCAACCGCAAAGTGACA GTTAAGCAATTGAAAGAGGACTACTTTAAGAAAATTGAATGCTTCGATTCTGTCGAGATCTCCGGGGTAGAAGATCGATT TAATGCGTCACTTGGTACGTATCATGACCTCCTAAAGATAATTAAAGATAAGGACTTCCTGGATAACGAAGAGAATGAAG ATATCTTAGAAGATATAGTGTTGACTCTTACCCTCTTTGAAGATCGGGAAATGATTGAGGAAAGACTAAAAACATACGCT CACCTGTTCGACGATAAGGTTATGAAACAGTTAAAGAGGCGTCGCTATACGGGCTGGGGACGATTGTCGCGGAAACTTAT CAACGGGATAAGAGACAAGCAAAGTGGTAAAACTATTCTCGATTTTCTAAAGAGCGACGGCTTCGCCATAGGAACTTTA TGCAGCTGATCCATGATGACTCTTTAACCTTCAAAGAGGATATACAAAAGGCACAGGTTTCCGGACAAGGGGACTCATTG CACGAACATATTGCGAATCTTGCTGGTTCGCCAGCCATCAAAAAGGGCATACTCCAGACAGTCAAAGTAGTGGATGAGCT AGTTAAGGTCATGGGACGTCACAAACCGGAAAACATTGTAATCGAGATGGCACGCGAAAATCAAACGACTCAGAAGGGGC AAAAAAACAGTCGAGAGCGGATGAAGAGAATAGAAGAGGGTATTAAAGAACTGGGCAGCCAGATCTTAAAGGAGCATCCT GTGGAAAATACCCAATTGCAGAACGAGAAACTTTACCTCTATTACCTACAAAATGGAAGGGACATGTATGTTGATCAGGA ACTGGACATAAACCGTTTATCTGATTACGACGTCGATCACATTGTACCCCAATCCTTTTTGAAGGACGATTCAATCGACA ATAAAGTGCTTACACGCTCGGATAAGAACCGAGGGAAAAGTGACAATGTTCCAAGCGAGGAAGTCGTAAAGAAAATGAAG AACTATTGGCGGCAGCTCCTAAATGCGAAACTGATAACGCAAAGAAAGTTCGATAACTTAACTAAAGCTGAGAGGGGTGG CTTGTCTGAACTTGACAAGGCCGGATTTATTAAACGTCAGCTCGTGGAAACCCGCCAAATCACAAAGCATGTTGCACAGA TACTAGATTCCCGAATGAATACGAAATACGACGAGAACGATAAGCTGATTCGGGAAGTCAAAGTAATCACTTTAAAGTCA AAATTGGTGTCGGACTTCAGAAAGGATTTTCAATTCTATAAAGTTAGGGAGATAAATAACTACCACCATGCGCACGACGC TTATCTTAATGCCGTCGTAGGGACCGCACTCATTAAGAAATACCCGAAGCTAGAAAGTGAGTTTGTGTATGGTGATTACA AAGTTTATGACGTCCGTAAGATGATCGCGAAAGCGAACAGGAGATAGGCAAGGCTACAGCCAAATACTTCTTTTATTCT AACATTATGAATTTCTTTAAGACGGAAATCACTCTGGCAAACGGAGAGATACGCAAACGACCTTTAATTGAAACCAATGG GGAGACAGGTGAAATCGTATGGGATAAGGGCCGGGACTTCGCGACGGTGAGAAAAGTTTTGTCCATGCCCCAAGTCAACA TAGTAAAGAAAACTGAGGTGCAGACCGGAGGGTTTTCAAAGGAATCGATTCTTCCAAAAAGGAATAGTGATAAGCTCATC GCTCGTAAAAGGACTGGGACCCGAAAAAGTACGGTGGCTTCGATAGCCCTACAGTTGCCTATTCTGTCCTAGTAGTGGC AAAAGTTGAGAAGGGAAATCCAAGAAACTGAAGTCAGTCAAAGAATTATTGGGGATAACGATTATGGAGCGCTCGTCTT TTGAAAAGACCCCATCGACTTCCTTGAGGCGAAAGGTTACAAGGAAGTAAAAAAGGATCTCATAATTAAACTACCAAAG TATAGTCTGTTTGAGTTAGAAATGGCCGAAAACGGATGTTGGCTAGCGCCGGAGAGCTTCCAAAGGGGAACGAACTCGC ACTACCGTCTAAATACGTGAATTTCCTGTATTTAGCGTCCCATTACGAGAAGTTGAAAGGTTCACCTGAAGATAACGAAC AGAAGCAACTTTTTGTTGAGCAGCACAAACATTATCTCGACGAAATCATAGAGCAAATTTCGGAATTCAGTAAGAGAGTC ATCCTAGCTGATGCCAATCTGGACAAAGTATTAAGCGCATACAACAAGCACAGGGATAAACCCATACGTGAGCAGGCGGA AAATATTATCCATTTGTTTACTCTTACCAACCTCGGCGCTCCAGCCGCATTCAAGTATTTTGACACAACGATAGATCGCA AACGATACACTTCTACCAAGGAGGTGCTAGACGCGACACTGATTCACCAATCCATCACGGGATTATATGAAACTCGGATA GATTTGTCACAGCTTGGGGGTGACGGATCCCCAAGAAGAAGAGGAAAGTCTCGAGCGACTACAAAGACCATGACGGTGA TTATAAAGATCATGACATCGATTACAAGGATGACGATGACAAGGCTGCAGGA JPEG0007779988000004.jpg172165 (single underline: HNH domain; double underline: RuvC domain)
[0135] In some embodiments, the wild-type Cas9 is Cas9 from Streptococcus pyogenes (NC BI reference sequence: NC_002737.2 (nucleotide sequence below) and Uniprot reference sequence: Q99ZW2 (nucleotide sequence below) The amino acid sequence corresponds to the following: ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGAATATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG ACAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGTAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAACCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTGGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAAAAATGGCTTATTTGGGAAT CTCATTGCTTTGTCATTGGGTTTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATACTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAACGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAAGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGTACCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGAGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGACATTCAAAGGCACAAGTGTCTGGACAAGGCGATAGTTTA CATGAACATATTGCAAATTTAGCTGGTAGCCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAGTTGTTGATGAATT GGTCAAAGTAATGGGGCGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCC AGAAAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTCTTAAAGAGCATCCT GTTGAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTCCAAAATGGAAGAGACATGTATGTGGACCAAGA ATTAGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCCTTAAAGACGATTCAATAGACA ATAAGGTCTTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAA AACTATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGG TTTGAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAA TTTTGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCT AAATTAGTTTCTGACTTCCGAAAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGC GTATCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATA AAGTTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCT AATATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGG GGAAACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATA TTGTCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAAAGAAATTCGGACAAGCTTATT GCTCGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGC TAAGGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCT TTGAAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAA TATAGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGC TCTGCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAAC AAAAACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTT ATTTTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGA AAATATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTA AACGATATACGTCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATT GATTTGAGTCAGCTAGGAGGTGACTGA JPEG0007779988000005.jpg168163 (single underline: HNH domain; double underline: RuvC domain)
[0136] In one embodiment, Cas9 is isolated from Corynebacterium ulcerans (NCBI Refs: NC_015683.1 , NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1) ; Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcu s iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psyc hroflexus torquisI (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref : YP_820832.1), Listeria innocua (NCBI Ref: NP_472073.1), Campylobacter jejuni ( NCBI Ref: YP_002344900.1) or Neisseria meningitidis (NCBI Ref: YP_002342100 .1), or Cas9 from any other organism.
[0137] In some embodiments, dCas9 contains one or more nucleotides that inactivate Cas9 nuclease activity. It corresponds to, or contains part or all of, a mutated Cas9 amino acid sequence. For example, in some embodiments, the dCas9 domain contains the D10A and H840A mutations or another Ca mutation. In some embodiments, the dCas9 comprises a corresponding mutation in dCas9 (D10A and and H840A) contains the amino acid sequence: JPEG0007779988000006.jpg171165 (single underline: HNH domain; double underline: RuvC domain)
[0138] In some embodiments, the Cas9 domain comprises a D10A mutation, while the the residue at position 840 in the amino acid sequence provided herein, or The residue at the corresponding position in either sequence remains a histidine.
[0139] In other embodiments, D10A, e.g., resulting in nuclease-inactivated Cas9 (dCas9), and dCas9 variants having mutations other than H840A. For example, other amino acid substitutions at D10 and H840, or the nuclease domain of Cas9 may be used. Other substitutions within the domain (e.g., HNH nuclease subdomain and / or RuvC1 subdomain) In some embodiments, a variant or homolog of dCas9 is at least about 70% identity, at least about 80% identity, at least about 90% identity , at least about 95% identity, at least about 98% identity, at least about 99% identity, Those having at least about 99.5% identity, or at least about 99.9% identity are provided. In some embodiments, the amino acid sequence is about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, or Amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids Variants of dCas9 having amino acid sequences shorter or longer than 1000 or 10000 are provided. do.
[0140] In some embodiments, the Cas9 fusion proteins provided herein comprise a Cas9 protein. The full-length amino acid sequence of the protein, for example, one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein contain a full-length Cas9 sequence. Examples of Suitable Cas9 Domains and Cas9 Fragments Suitable amino acid sequences are provided herein, and further suitable sequences for Cas9 domains and fragments are , as will be apparent to those skilled in the art.
[0141] Additional Cas9 proteins (e.g., nuclease-dead Cas9 (dCas9), Cas9 nCas9, or nuclease-active Cas9, its variants and homologs It should be understood that within the scope of this disclosure are any and all Cas9 proteins, including: In some embodiments, Cas9 includes, but is not limited to, those provided below. The protein is a nuclease-inactive Cas9 (dCas9). In some embodiments, the Cas9 protein In some embodiments, the Cas9 protein is a Cas9 nickase (nCas9). The quality is nuclease-active Cas9.
[0142] Exemplary catalytically inactive Cas9 (dCas9): DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKLVDSTDKADLRIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD
[0143] Exemplary catalyst of Cas9ニッカーゼ (nCas9): DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD
[0144] Exemplary catalyst activityCas9: DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICY LQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHM IKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNL IALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASM IKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRK QRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEV VDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDRIEMIEERLKTYAH LFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLH EHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHPV ENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSK LVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSN IMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIA RKKDWDPCKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKY SLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVI LADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRID LSQLGGD.
[0145] In one embodiment, Cas9 is directed against archaeal organisms that comprise the domain and kingdom of unicellular prokaryotic microorganisms. In one embodiment, the Cas9 protein is derived from, for example, a fungus (e.g., nanoarchaea). , Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res . 2017 Feb 21. doi: 10.1038 / cr.2017.21 refers to CasX or CasY, and The entire contents of which are incorporated herein by reference. Many CRISPR-Cas systems, including Cas9, which was first reported in the archaeal domain of life, have been identified. This branched Cas9 protein was identified in the little-studied nanoarchaea. , discovered as part of an active CRISPR-Cas system, previously unknown in bacteria. Two systems, CRISPR-CasX and CRISPR-CasY, have been discovered that are the most efficient known to date. In some embodiments, Cas9 is a CasX or CasX-like protein. In some embodiments, Cas9 represents a variant of CasY or a variant of CasY. It represents nucleic acid programmable DNA binding protein (napDNAbp). Other RNA-guided DNA binding proteins may also be used, such as RNA-guided DNA binding proteins (RNA-guided DNA binding proteins), as disclosed herein. It should be understood that the range is
[0146] In certain embodiments, napDNAbp useful in the methods of the invention contain circular substitutions. , which is known and described, for example, in Oakes et al., Cell 176, 254-267, 2019 Below are exemplary circular permutations, where bolded sequences indicate Cas9-derived sequences and italicized sequences: The sequence represents the linker sequence and the underlined sequence represents the bipartite nuclear localization sequence. JPEG0007779988000007.jpg187163
[0147] Polynucleotide programmable nucleotides that can be incorporated into base editors Non-limiting examples of binding domains include domains derived from CRISPR proteins, restriction nucleases, and the like. Enzymes, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases Examples include ZFNs.
[0148] In some embodiments, the nucleic acid of any of the fusion proteins provided herein The programmable DNA binding protein (napDNAbp) can be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. is at least 85%, at least 90%, or At least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% In some embodiments, the napDNAbp comprises a naturally occurring amino acid sequence having identity to the napDNAbp. In some embodiments, the napDNAbp is a CasX or CasY protein present. At least 85%, at least at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least Contains amino acid sequences with at least 99.5% identity to Cas12b / C2c1 and CasX from other bacterial species. It should be understood that CasY and CasC may also be used in accordance with the present disclosure.
[0149] Cas12b / C2c1 (uniprot.org / uniprot / T0D7A2#2) sp|T0D7A2|C2C1_ALIAG CRISPR-associated endonuclease C2c1 OS = Alicyclobac illus acido- terrestris (strain ATCC 49025 / DSM 3922 / CIP 106132 / NCIMB 1313 7 / GD3B) GN=c2c1 PE=1 SV=1 MAVKSIKVKLRLDDMPEIRAGLWKLHKEVNAGVRYYTEWLSLLRQENLYRRSPNGDGEQECDKTAEECKAELLERLRARQ VENGHRGPAGSDDELLQLARQLYELLVPQAIGAKGDAQQIARKFLSPLADKDAVGGLGIAKAAGNKPRWVRMREAGEPGWE EEKEKAETRKSADRTADVLRALADFGLKPLMRVYTDSEMSSVEWKPLRKGQAVRTWDRDMFQQAIERMMSWESWNQRVGQ EYAKLVEQKNRFEQKNFVGQEHLVHLVNQLQQDMKEASPGLESKEQTAHYVTGRALRGSDKVFEKWGKLAPDAPFDLYDA EIKNVQRRNTRRFGSHDLFAKLAEPEYQALWREDASFLTRYAVYNSILRKLNHAKMFATTFTLPDATAHPIWTRFDKLGGN LHQYTFLFNEFGERRHAIRFHKLLKVENGVAREVDDVTVPISMSEQLDNLLPRDPNEPIALYFRDYGAEQHFTGEFGGAK IQCRRDQLAHMHRRRGARDVYLNVSVRVQSQSEARGERRPPYAAVFRLVGDNHRAFVHFDKLSDYLAEHPDDGKLGSEGL LSGLRVMSVDLGLRTSASISVFRVARKDELKPNSKGRVPFFFPIKGNDNLVAVHERSQLLKLPGETESKDLRAIREERQR TLRQLRTQLAYLRLLVRCGSEDVGRRERSWAKLIEQPVDAANHMTPDWREAFENELQKLKSLHGICSDKEWMDAVYESVR RVWRHMGKQVRDWRKDVRSGERPKIRGYAKDVVGGNSIEQIEYLERQYKFLKSWSFFGKVSGQVIRAEKGSRFAITLREH IDHAKEDRLKKLADRIIMEALGYVYALDERGKGKWVAKYPPCQLILLEELSEYQFNNDRPPSENNQLMQWSHRGVFQELI NQAQVHDLLVGTMYAAFSSRFDARTGAPGIRCRRVPARCTQEHNPEPFPWWLNKFVVEHTLDACPLRADDLIPTGEGEIF VSPFSAEEGDFHQIHADLNAAQNLQQRLWSDFDISQIRLRCDWGEVDGELVLIPRLTGKRTADSYSNKVFYTNTGVTYYE RERGKKRRKVFAQEKLSEEEAELLVEADEAREKSVVLMRDPSGIINRGNWTRQKEFWSMV NQRIEGYLVKQIRSRVPLQ DSACENTGDI
[0150] CasX (uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) >tr|F0NN87|F0NN87_SULIH CRISPR-associated Casx protein OS = Sulfolobus isla ndicus (strain HVE10 / 4) GN = SiH_0402 PE=4 SV=1 MEVPLYNIFGDNYIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAA AGWVLTRLGKAKVSEGDYVGVNVFTRPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVSSVTNPNVSVVRIYTISDAV GQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSSNMRERYIVLANIIYEYLTG SKRLEDLLYFANRDLIMNL NSDDGKVRDLKLISAIVNGELIRGEG
[0151] >tr|F0NH53|F0NH53_SULIR CRISPR associated protein, Casx OS = Sulfolobus is landicus (strain REY15A) GN=SiRe_0771 PE=4 SV=1 MEVPLYNIFGDNYIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAA AGWVLTRLGKAKVSEGDYVGVNVFTRPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVSSVTNPNVSVVSIYTISDAV GQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLN SDDGKVRDLKLISAYVNGELIRGEG
[0152] Deltaproteobacteria CasX MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMDTDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAIL QVYWQEFKDDHVGLMCKFAQPASKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFWYKLEQVSEKGKAITNYFGRCNVA EHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESHTPVKPLAQIAGNRYASGPVGKALSDACMGTIASFL SKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDAYNEVIARVRMWVNLNLWQKLKLSRDDAK PLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENPK KPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMD EKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTD GTDIKKSGKWQGLLYGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIG RDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQA AKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKL AYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELS AELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYK SGKQPFVGAWQAFYKRRLKEVWKPNA
[0153] CasY (ncbi.nlm.nih.gov / protein / APG80656.1) >APG80656.1 CRISPR-associated protein CasY (uncultured Parcubacteria group bacterium] MSKRHPRISGVKGYRLHAQRLEYTGKSGAMRTIKYPLYSSPSGGRTVPREIVSAINDDYVGLYGLSNFDDLYNAEKRNEE KVYSVLDFWYDCVQYGAVFSYTAPGLLKNVAEVRGGSYELTKTLKGSHLYDELQIDKVIKFLNKKEISRANGSLDKLKKD IIDCFKAEYRERHKDQCNKLADDIKNAKKDAGASLGERQKKLFRDFFGISEQSENDKPSFTNPLNLTCCLLPFDTVNNNR NRGEVLFNKLKEYAQKLDKNEGSLEMWEYIGIGNSGTAFSNFLGEGFLGRLRENKITELKKAMMDITDAWRGQEQEEELE KRLRILAALTIKLREPKFDNHWGGYRSDINGKLSSWLQNYINQTVKIKEDLKGHKKDLKKAKEMINRFGESDTKEEAVVS SLLESIEKIVPDDSADDEKPDIPAIAIYRRFLSDGRLTLNRFVQREDVQEALIKERLEAEKKKKPKKRKKKSDAEDEKET IDFKELFPHLAKPLKLVPNFYGDSKRELYKKYKNAAIYTDALWKAVEKIYKSAFSSSLKNSFFDTDFDKDFFIKRLQKIF SVYRRFNTDKWKPIVKNSFAPYCDIVSLAENEVLYKPKQSRSRKSAAIDKNRVRLPSTENIAKAGIALARELSVAGFDWK DLLKKEEHEEYIDLIELHKTALALLLAVTETQLDISALDFVENGTVKDFMKTRDGNLVLEGRFLEMFSQSIVFSELRGLA GLMSRKEFITRSAIQTMNGKQAELLYIPHEFQSAKITTPKEMSRAFLDLAPAEFATSLEPESLSEKSLLKLKQMRYYPHY FGYELTRTGQGIDGGVAENALRLEKSPVKKREIKCKQYKTLGRGQNKIVLYVRSSYYQTQFLEWFLHRPKNVQTDVAVSG SFLIDEKKVKTRWNYDALTVALEPVSGSERVFVSQPFTIFPEKSAEEEGQRYLGIDIGEYGIAYTALEITGDSAKILDQN FISDPQLKTLREEVKGLKLDQRRGTFAMPSTKIARIRESLVHSLRRNRIHHLALKHKAKIVYELEVSRFEEGKQKIKKVYA TLKKADVYSEIDADKNLQTTVWGKLAVASEISASYTSQFCGACKKLWRAEMQVDETITTQELIGTVRVIKGGTLIDAIKD FMRPPIFDENDTPFPKYRDFCDKHHISKKMRGNSCLFICPFCRANADADIQASQTIALLRYVKEEKKVEDYFERFRKLKN IKVLGQMKKI
[0154] The term "conservative amino acid substitution" or "conservative variation" refers to a mutation in which an amino acid is replaced by another amino acid that shares a common characteristic. It refers to the substitution of an amino acid with another amino acid that has the same properties. It defines the common properties between individual amino acids. A functional method for determining the normality of amino acid changes between corresponding proteins of homologous organisms is The goal is to analyze the frequency of the results (Schulz, GE and Schirmer, RH, Principles of f Protein Structure, Springer-Verlag, New York (1979). According to such analysis, Amino acids within a group are preferentially exchanged with each other, thus affecting the overall protein structure. Define the group of amino acids that are most similar to each other in their effect on the (Schulz, GE and Schirmer, RH, supra). Non-limiting examples of conservative mutations include: For example, the amino acids arginine to lysine and the like that can maintain a positive charge are used. Reverse; aspartic acid to glutamic acid and vice versa, which can maintain a negative charge; free threonine to serine, which maintains the -OH; and asparagine to asparagine, which maintains the free NH Examples include amino acid substitutions such as glutamine.
[0155] The terms "coding sequence" or "protein-coding sequence" are used interchangeably herein. " refers to a segment of a polynucleotide that encodes a protein. The sequence is bounded by a start codon near the 5' end and a stop codon near the 3' end. The coding sequence is also called an open reading frame.
[0156] "Cytidine deaminase" is a enzyme that converts amino groups into carbonyl groups through a deamination reaction. In one embodiment, the term "catalyzed polypeptide" refers to a polypeptide or fragment thereof that is capable of catalyzing , cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine The cytidine deaminases provided herein (e.g., engineered cytidine deaminases) convert Cytidine deaminases (evolved cytidine deaminases) can be derived from any organism, e.g., bacteria. It is possible that.
[0157] In some embodiments, the base editor cytidine deaminase is apolipoprotein B1. It may contain all or part of the protein B mRNA editing complex (APOBEC) family deaminase. APOBEC is an evolutionarily conserved family of cytidine deaminases. In some embodiments, a member of the family is a C→U editing enzyme. The amino acids include, but are not limited to, APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOB EC3D (now referred to as "APOBEC3E"), APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, Activation-induced (cytidine) deaminase (AID), hAPOBEC1 from Homo sapiens, Rattus no rAPOBEC1 from Pongo pygmaeus, rAPOBEC1 from Pongo pygmaeus, and Alligator mississippiensi AmAPOBEC1 (BEM3.31) from s, ocAPOBEC1 from Oryctolagus cuniculus, Sus scrof a, SsAPOBEC2 (BEM3.39), hAPOBEC3A, Mesocricetus a maAPOBEC1 from Monodelphis uratus, mdAPOBEC1 from Monodelphis domestica; aminase 1 (CDA1), hA3A, an APOBEC3A from Homo sapiens, and Rhinopithecus roxella APOBEC3F RrA3F (BEM3.14) from Petromyzon marinus, and PmCDA1 (Petr omyzon marinus cytosine deaminase 1, "PmCDA1"); mammalian (e.g., human, porcine) AID (activation-induced cytidine deaminase; AICDA) derived from mammals such as cattle, horses, and monkeys; hAID from Homo sapiens; and the APOBEC family, including but not limited to FENRY Including members.
[0158] As used herein, the terms "deaminase" or "deaminase domain" and "deaminase domain" refer to refers to a protein or enzyme that catalyzes a deamination reaction. The enzyme or deaminase domain binds to uridine or deoxyuridine, respectively. Cytidine deaminase catalyzing the hydrolytic deamination of cytidine or deoxycytidine In some embodiments, the deaminase or deaminase domain is a cytosine deaminase. enzyme, which catalyzes the hydrolytic deamination of cytosine to uracil. In this case, the deaminase catalyzes the hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase is adenosine or Adenosine deamination catalyzes the hydrolytic deamination of adenine (A) to inosine (I). In some embodiments, the deaminase or deaminase domain is an adenosine deaminase. adenosine or deoxyinosine to inosine or deoxyinosine, respectively. catalyzes the hydrolytic deamination of deoxyadenosine. Adenosine deaminase hydrolyzes the adenosine in deoxyribonucleic acid (DNA). The deaminases provided herein (e.g., engineered deaminases) catalyze The deaminase (evolved deaminase) can be from any organism, such as a bacterium. Adenosine deaminase is also found in Escherichia coli, Staphylococcus aureus, and Salmonella. lla typhimurium, Shewanella putrefaciens, Haemophilus influenzae, or Caulobac It is derived from bacteria such as ter crescentus.
[0159] "Detection" refers to identifying the presence, absence, or amount of an analyte to be detected. In an embodiment, sequence variations in a polynucleotide or polypeptide are detected. In another embodiment, the presence of indels is detected.
[0160] A "detectable label" means a label that, when attached to a molecule of interest, is detectable by spectroscopic, photochemical, or biochemical means. "detectable" refers to a composition that renders the latter detectable through biological, immunochemical, or chemical means. For example, useful labels include radioisotopes, magnetic beads, metal beads, colloidal particles, fluorescent Photochromic dyes, electron-dense reagents, enzymes (e.g., commonly used in enzyme-linked immunosorbent assays (ELISAs)) These include hydroxybenzoates, biotin, digoxigenin, or haptens.
[0161] "Disease" means any condition that damages or interferes with the normal function of a cell, tissue, or organ. Or it means disability.
[0162] As used herein, the term "effective amount" refers to an amount sufficient to induce a desired biological response. The term "amount of biologically active agent" refers to the amount of biologically active agent used to practice the present invention for the therapeutic treatment of disease. The effective amount of active agent(s) administered will depend on the mode of administration, the age, weight, and general health of the subject. Ultimately, your doctor or veterinarian will determine the appropriate amount and dosage. Such an amount is referred to as an "effective" amount. In one embodiment, an effective amount is a dose that is administered to a cell (e.g., an iPSC). A gene of interest in a cell (in vitro or in vivo) is prepared by subjecting the gene to a gene encoding the gene of interest to a gene encoding the gene of interest. Base editors (e.g., programmable DNA binding proteins, nucleobase editors, and In some embodiments, the fusion proteins provided herein are An effective amount of a protein, e.g., an nCas9 domain and one or more deaminase domains (e.g., Multi-effector nucleobases including ATPase, adenosine deaminase, and cytidine deaminase An effective amount of an editor is a compound that specifically binds to a nucleic acid sequence that is specifically bound by the multi-effector nucleobase editor. This may refer to the amount of fusion protein sufficient to induce editing of the target site to be edited. In one embodiment, an effective amount is an amount that provides a therapeutic effect (e.g., alleviates a disease or a symptom or condition thereof). The amount of base editor required to achieve such a treatment is the amount of base editor required to achieve a desired effect (reduction or control). The therapeutic effect is sufficient to alter the gene of interest in all cells of a subject, tissue, or organ. It is not necessary that the number of cells in the subject, tissue, or organ is about 1%, 5%, 10%, 25%, 50%, 7%, or 8% of the cells present in the subject, tissue, or organ. The gene of interest only needs to be altered by 5% or more.
[0163] In some embodiments, the fusion proteins provided herein (e.g., nCas9 domains) are A domain and one or more deaminase domains (e.g., adenosine deaminase, cytogenes deaminase, An effective amount of a nucleobase editor, including a nucleobase deaminase, is a nucleobase deaminase described herein. a fusion protein sufficient to induce editing of the target site that is specifically bound and edited by the editor; As will be appreciated by those skilled in the art, the amount of an agent (e.g., a fusion protein) that is present in a given amount may be increased by a factor of 10. proteins, nucleases, hybrid proteins, protein dimers, proteins (or an effective amount of a complex of a protein dimer and a polynucleotide, or a polynucleotide; may be used to determine, for example, the desired biological response, e.g., the specific allele to be edited, the genome, if or target site, cells or tissues to be targeted, and / or agents used, etc. The amount of time that the temperature can change can vary depending on a variety of factors.
[0164] By "fragment" is meant a portion of a polypeptide or nucleic acid molecule, which portion is identical to that of a reference nucleic acid At least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% of the entire length of the molecule or polypeptide , or 90%. The fragments may be 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 3 nucleotides or amino acids. do.
[0165] "Guide RNA" or "gRNA" is a polynucleotide that can be specific for a target sequence. a programmable nucleotide-binding domain protein (e.g., Cas9 or Cpf1) In one embodiment, the term "polynucleotide" refers to a polynucleotide that can form a complex with a The guide polynucleotide is a guide RNA (gRNA). gRNA is a complex of two or more RNAs. It can exist as a single RNA molecule or as a single RNA molecule. The gRNAs present are sometimes called single guide RNAs (sgRNAs), but "gRNA" refers to a single Interchangeable to refer to guide RNAs that exist as a single molecule or as a complex of two or more molecules. Typically, gRNAs exist as a single RNA species, and are used in a variety of ways. (1) They are homologous to the target nucleic acid. (2) a domain that shares a common property (e.g., directs binding of the Cas9 complex to the target); and In one embodiment, the domains are the s9 protein-binding domains. In (2), the sequence corresponds to the sequence known as tracrRNA and contains a stem-loop structure. In some embodiments, domain (2) is selected from the group consisting of the amino acid sequence of Jinek et al., Science 337:816-821(201 2) tracrRNA as provided in (the entire contents of which are incorporated herein by reference) Other examples of gRNAs (e.g., those containing domain 2) are listed on September 6, 2013. In a filed U.S. provisional patent application entitled "Switchable Cas9 Nucleases and Uses Thereof," U.S. SSN 61 / 874,682 and "Delivery System For Functional Numerical System" filed September 6, 2013 The entire disclosure of each of these patent applications can be found in U.S. Provisional Patent Application No. USSN 61 / 874,746 entitled "Patent Application No. 61 / 874,746," which is hereby incorporated by reference. The contents of which are incorporated herein by reference. In some embodiments, the gRNA A gRNA containing two or more of the main (1) and (2) sequences may be referred to as an "extended gRNA." The gRNA binds to two or more Cas9 proteins and expresses two or more Cas9 proteins as described herein. The gRNA binds to the target nucleic acid at a different region on the target site. sequence, which mediates binding of the nuclease / RNA complex to the target site, :Provides sequence specificity of the RNA complex.
[0166] "Hybridization" means hydrogen bonding between complementary nucleobases, as defined by Watson- It can be a Crick, Hoogsteen or reversed Hoogsteen hydrogen bond. For example, Adenine and thymine are complementary nucleobases that form hydrogen bonds to form pairs.
[0167] The term "inhibitor of base repair" or "IBR" refers to a nucleic acid Inhibiting the activity of repair enzymes, such as base excision repair (BER) enzymes In one embodiment, IBR refers to a protein capable of inhibiting inosine base excision repair. Examples of inhibitors of base repair include APE1, Endo III, Endo IV, Endo V, and Endo o VIII, Fpg, hOGGl, hNEILl, T7 Endol, T4 PDG, UDG, hSMUGL, and hAAG inhibitors In one embodiment, the IBR is an inhibitor of Endo V or hAAG. In some embodiments, the IBR is catalytically inactive EndoV or catalytically inactive hAAG. In one embodiment, the base repair inhibitor is an inhibitor of Endo V or hAAG. , the base repair inhibitor is catalytically inactive EndoV or catalytically inactive hAAG.
[0168] In some embodiments, the base repair inhibitor is a uracil glycosylase inhibitor (UGI). UGI is a gene that can inhibit the base excision repair enzyme uracil-DNA glycosylase. In some embodiments, the UGI domain refers to a protein that is a wild-type UGI or In some embodiments, the UGI proteins provided herein comprise a fragment of U In one embodiment, the present invention includes a fragment of GI and a protein homologous to UGI or a UGI fragment. The base repair inhibitor is an inhibitor of inosine base excision repair. In the present study, base repair inhibitors were identified as "catalytically inactive inosine-specific nucleases" or "dead" nucleases. Without wishing to be bound by any particular theory, However, catalytically inactive inosine glycosylases (e.g., alkyladenine glycosylases) Abasic amino acid sequence (AAG) can bind to inosine but also creates an abasic site. It is also unable to remove the inosine moiety, thereby blocking the newly formed inosine moiety from DNA damage / repair. In some embodiments, catalytically inactive inositols are sterically blocked from the cleavage mechanism. Inosine-specific nucleases can bind to inosine in nucleic acids but do not cleave the nucleic acid. Representative, non-limiting examples of catalytically inactive inosine-specific nucleases include: Catalytically inactive alkyl adenosine glycosylase (AAG nuclease) (e.g., from humans) and catalytically inactive endonuclease V (EndoV nuclease) (e.g., from E. coli). In some embodiments, catalytically inactive AAG nucleases are used. The AAG nuclease contains an E125Q mutation or a corresponding mutation in another AAG nuclease.
[0169] By "increase" is meant a positive change of at least 10%, 25%, 50%, 75%, or 100%.
[0170] An "intein" excises itself and releases the remaining fragment (the extein) (extein)) to form peptide bonds in a process known as protein splicing Inteins are fragments of proteins that can be linked together using a "protein intron." The intein excises itself and links the rest of the protein. The process is referred to herein as "protein splicing" or "intein-mediated In some embodiments, precursor proteins (integrins) are spliced. Inteins (intein-containing proteins before intein-mediated protein splicing) Such inteins are referred to herein as split inteins. These are called split inteins (e.g., split intein-N and split intein-C). In Bacteria, DnaE, the catalytic subunit a of DNA polymerase III, is expressed in two separate genes. Encoded by the genes dnaE-n and dnaE-c. Encoded by the dnaE-n gene The intein may be referred to herein as "intein N." The loaded intein may be referred to herein as "intein C."
[0171] Other intein systems can also be used. For example, the dnaE intein, i.e., Cfa-N (e.g., Based on the intein pair Cfa-C (e.g., split intein-N) and Cfa-C (e.g., split intein-C) Synthetic inteins have been described (see, e.g., Stevens, J. Med. Chem. Soc. 1999, 144:111-112, which is incorporated herein by reference). s et al., J Am Chem Soc. 2016 Feb. 24; 138(7):2162-5). Use in accordance with this disclosure Non-limiting examples of intein pairs that can be used include Cfa DnaE intein, Ssp GyrB intein, and Intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein intein, and Cne Prp8 intein (see, e.g., U.S. Pat. No. 6,223,629, incorporated herein by reference). Examples include those described in Patent No. 8,394,604.
[0172] Exemplary nucleotide and amino acid sequences of inteins are provided. DnaE Intein-N DNA: TGCCTGTCATACGAAACCGAGATACTGACAGTAGAATATGGCCTTCTGCCAATCGGG AAGATTGTGGAGAAACGGATAGAATGCACAGTTTACTCTGTCGATAACAATGGTAACATTTATACTCAGCCAGTTGCCCA GTGGCACGACCGGGGAGAGCAGGAAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAGGGCCACTAAGGACCACA AATTTATGACAGTCGATGGCCAGATGCTGCCTATAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGAGTTGACAAC CTTCCTAAT DnaE intein-N protein: CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDR GEQEVFEYCLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRVDNLPN DnaE intein-C DNA: ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGATATTGG AGTCGAAAGAGATCACAACTTTGCTCTGAAGAACGGATTCATAGCTTCTAAT Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN Cfa-N DNA: TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCCTATTGGAAAGATTGTCGAA GAGAGAATTGAATGCACAGTATATACTGTAGACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCG CGGCGAACAAGAAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACGAGCAACTAAAGATCATAAATTCATGACCA CTGACGGGCAGATGTTGCCAATAGATGAGATATTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTGCCA Cfa-N protein: CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQ MLPIDEIFERGLDLKQVDGLP Cfa-C DNA: ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAGGAAAGTAAAGATAATATCT CGAAAAAGCTTGGTACCCAAAATGTCTATGATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGT AGCCAGCAAC Cfa-C protein: MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN
[0173] To join the N-terminal part of the split Cas9 with the C-terminal part of the split Cas9, intein N and intein B were used. The tein C can be fused to the N-terminal end of the split-Cas9 and the C-terminal end of the split-Cas9, respectively. For example, In some embodiments, intein-N is C-terminal to the N-terminal portion of the split Cas9. The split-Cas9 is fused to the intein-N domain, forming the structure N--[N-terminal portion of split-Cas9]-[intein-N]--C. In some embodiments, the intein-C is located at the N-terminus of the C-terminal portion of the split Cas9. The intein is fused to the C-terminal portion of the split Cas9, forming the structure N-[intein-C]-[C-terminal portion of the split Cas9]-C. Intein for linking the protein (e.g., split-Cas9) to which the intein is fused The mechanism of tein-mediated protein splicing is described, for example, in the publications As described in Shah et al., Chem Sci. 2014; 5(1):446-461, the present invention Methods for designing and using inteins are known in the art. and are disclosed, for example, in WO2014004336, WO2017132580, US20150344549 and US20180127780. and US Pat. No. 6,263,999, each of which is incorporated herein by reference in its entirety.
[0174] The terms "isolated," "purified," or "biologically pure" refer to a substance that is in its native state. Substances that have been removed to varying degrees from components normally associated with them when found in the natural state. "Isolated" refers to the degree of separation from the original source or surrounding environment. "Purified" refers to the degree of separation from the original source or surrounding environment. A "purified" or "biologically pure" protein is one that is free from impurities. that the substance does not materially affect the biological properties of the protein or cause other adverse consequences. In other words, the nucleic acid or peptide of the present invention is or, if produced by recombinant DNA technology, cellular material, viral material, or culture medium. If the substance is not naturally contained in the substance, or if it is chemically synthesized, chemical precursors or other chemical substances Purity and homogeneity are typically determined by analytical methods. Chemical techniques, such as polyacrylamide gel electrophoresis or high performance liquid chromatography The term "purified" refers to the degree to which a nucleic acid or protein is purified by electrophoresis. This can mean that the resulting protein essentially produces one band. For proteins that can undergo glycosylation, the different modifications are purified separately. This can result in different isolated proteins that can be isolated.
[0175] An "isolated polynucleotide" is a nucleic acid molecule that does not occur in the naturally occurring genome of the organism from which the nucleic acid molecule of the invention is derived. It means nucleic acid (e.g., DNA) that does not contain the genes adjacent to the gene. The term refers to, for example, those incorporated into vectors; autonomously replicating plasmids or viruses. integrated into the genomic DNA of prokaryotes or eukaryotes; or independently of other sequences another molecule created (e.g., cDNA generated by PCR or restriction endonuclease digestion) or genomic or cDNA fragments). RNA molecules transcribed from DNA molecules, as well as hybrids encoding additional polypeptide sequences. It contains recombinant DNA that is part of the hybrid gene.
[0176] An "isolated polypeptide" is a polypeptide of the invention separated from components that naturally accompany it. Typically, a polypeptide is a polypeptide that is a protein with which it is naturally associated. A substance is isolated if it is at least 60% by weight free from substances and naturally occurring organic molecules. Preferably, the preparation comprises at least 75% by weight of soluble fiber, more preferably at least 90% by weight of soluble fiber, most preferably at least 10% by weight of soluble fiber. or at least 99% of the isolated polypeptide of the present invention. can be prepared by, for example, extraction from a natural source, expression of a recombinant nucleic acid encoding such a polypeptide, or the like. or by chemically synthesizing the protein. Purity can be achieved by any suitable method. Suitable methods, such as column chromatography, polyacrylamide gel electrophoresis, or can be measured by HPLC analysis.
[0177] As used herein, the term "linker" refers to a molecule that binds two molecules or moieties (e.g., a Two components of a protein or ribonucleocomplex, or two domains of a fusion protein In one embodiment, a polynucleotide programmable DNA binding domain (e.g., dCas9) and one or more The deaminase domain (e.g., adenosine deaminase and / or cytidine deaminase) Covalent linkers (e.g., covalent bonds), non-covalent linkers, chemical linkers, A linker can refer to a group, group, or molecule that connects different components of a base editor system. You can connect different parts of a component together. For example, In this embodiment, the linker is a polynucleotide programmable nucleotide linker. The guide polynucleotide binding domain of the binding domain and the catalytic domain of the deaminase In some embodiments, the linker can connect the CRISPR polypeptide and the deamid In some embodiments, the linker can link the Cas9 and the deaminase. In some embodiments, the linker can connect the dCas9 and the deaminase. In some embodiments, the linker can connect the nCas9 and the deaminase. In some embodiments, the linker can be a guide polynucleotide and a deamidating polynucleotide. In some embodiments, the linker can link a base editor Deamination component of the polynucleotide-programmable nucleotide binding component In some embodiments, the linker can link the base editor sequence. The RNA-binding portion of the stem deamination component and the polynucleotide programmable nucleotide In some embodiments, the linker can be: The RNA-binding moiety of the deaminating component of the base editor system and polynucleotide programming The linker can link two nucleotide-binding moieties to the RNA-binding portion of the nucleotide-binding moiety. Located between or sandwiched by groups, molecules, or other moieties and bonded covalently or are linked to each other through non-covalent interactions, thus allowing the two to be linked together. In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker can be a polynucleotide. The linker can be a DNA linker. In some embodiments, the linker can be an RNA linker. In some embodiments, the linker comprises an aptamer capable of binding to the ligand. In some embodiments, the ligand may be a carbohydrate, a peptide, a protein, or can be a nucleic acid. In some embodiments, the linker is derived from a riboswitch The riboswitch from which the aptamer is derived can include theophylline riboswitch. riboswitch, thiamine pyrophosphate (TPP) riboswitch, adenosine cobalamin (AdoCbl) Riboswitch, S-adenosylmethionine (SAM) riboswitch, SAH riboswitch, Fra FMN riboswitch, tetrahydrofolate riboswitch, lysine riboswitch switch, glycine riboswitch, purine riboswitch, GlmS riboswitch, or pre In some embodiments, the linker may be selected from the group consisting of: Aptamers bound to protein domains such as polypeptides or polypeptide ligands In some embodiments, the polypeptide ligand may comprise a K homology (KH) domain, an MS2 domain, Coat protein domain, PP7 coat protein domain, SfMu Com coat protein domain, sterile α motif, telomerase Ku binding motif and Ku protein, telomerase The motifs may be Sm7 enzyme binding motifs and Sm7 protein, or RNA recognition motifs. In embodiments, the polypeptide ligand can be part of a base editor system component. For example, the nucleobase editing component may comprise one or more deaminase domains and an RNA recognition motif. It can be seen.
[0178] In some embodiments, the linker is an amino acid or a plurality of amino acids (e.g., a peptide or In some embodiments, the linker may be about 5 to 100 amino acids in length. For example, lengths of approximately 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20-30 , 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, or 90-100 amino acids. In some embodiments, the linker has a length of about 100-150, 150-200, 200-250, 250-300, It can be 300-350, 350-400, 400-450, or 450-500 amino acids. Shorter linkers are also contemplated.
[0179] In some embodiments, the linker comprises an RNA programmable The gRNA-binding domain of a nuclease and a nucleic acid editing protein (e.g., cytidine and / or In some embodiments, the linker The linker connects the dCas9 and the nucleic acid editing protein. For example, the linker can be a group, molecule, or a combination of two groups. or other moieties, or adjacent by two groups, molecules, or other moieties and are linked to each other via a covalent bond, thus linking the two. wherein the linker is an amino acid or a plurality of amino acids (e.g., a peptide or protein) In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5 to 200 amino acids in length, e.g., 5, 6, 7, or ,8,9,10,11,12,13,14,15,16,17,18,19,20,25,35,45,50,55,60,60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130 , 140, 150, 160, 175, 180, 190, or 200 amino acids. Shorter linkers are also contemplated.
[0180] In some embodiments, a nucleobase editor (e.g., a multi-effector nucleobase editor) The domains for SGGSSGSETPGTSESATPESSGGS, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, or orGGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTE PSEGSAPGTSTEPSE The fusion occurs via a linker containing the amino acid sequence GSAPGTSESATPESGPGSEPATSGGSGGS. In some embodiments, a nucleobase editor (e.g., a multi-effector nucleobase editor) The domains are linked via a linker containing the amino acid sequence SGSETPGTSESATPES, which may also be referred to as an XTEN linker. In some embodiments, the linker comprises the amino acid sequence SGGS. In embodiments, the linker is (SGGS) n , (GGGS) n , (GGGGS) n , (G) n、 (EAAAK) n , (GGS) n , S GSETPGTSESATPES, or (XP) n motif, or any combination thereof, wherein n is independently an integer from 1 to 30, and X is any amino acid. wherein n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15.
[0181] In some embodiments, the linker is 24 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In some embodiments, the linker has the amino acid sequence SGGSSGGSSSGSETP In some embodiments, the linker comprises GTSESATPESSGGSSGGSSGGSSGGS. In some embodiments, the linker has the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSG In some embodiments, the linker comprises GSSGGSSGGSSGSETPGTSESATPESSGGS SGGS. In some embodiments, the linker has the amino acid sequence PGSPAGSPTSTEEGTSESATP. Contains ESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS.
[0182] A "marker" is any molecule that has an altered expression level or activity that is associated with a disease or disorder. The term "protein" refers to any protein or polynucleotide.
[0183] As used herein, the term "mutation" refers to a change in a sequence, e.g., a nucleic acid or amino acid sequence. The substitution of a residue in a sequence of amino acids by another residue, or the deletion of one or more residues in the sequence. Mutations, as used herein, are typically made by identifying the original residue and then and identifying the newly substituted residue, Various methods for making amino acid substitutions (mutations) are provided herein. are well known in the art and are described, for example, in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012). In some embodiments, the present disclosure The proposed base editor generates a significant number of unintended mutations, e.g., unintended point mutations. without introducing a "purposeful mutation," e.g., a point mutation, into a nucleic acid (e.g., a nucleic acid in a subject's genome) In some embodiments, the intended mutation can be efficiently generated. The mutation is a guide polynucleotide specifically designed to produce the intended mutation. specific base editors (e.g., cytidine base editors and These mutations are caused by nucleotides derived from nucleotides (e.g., nucleotides containing nucleotides of the nucleotide sequence or adenosine base editors).
[0184] Generally, the amino acid sequence of the present invention is made or identified in a sequence (e.g., an amino acid sequence described herein). The mutations detected are numbered relative to the reference (or wild-type) sequence, i.e., the sequence that does not contain the mutation. Those skilled in the art will recognize mutations in amino acid and nucleic acid sequences relative to a reference sequence. It will be easy to understand how to determine the position.
[0185] The term "non-conservative mutation" refers to amino acid substitutions between different groups, e.g. For example, tryptophan to lysine, or serine to phenylalanine. The non-conservative amino acid substitutions do not disrupt or inhibit the biological activity of the functional variant. Non-conservative amino acid substitutions are preferred because they improve the biological activity of the functional variant compared to the wild-type protein. The biological activity of the functional variant can be enhanced so that it is increased compared to the protein. do.
[0186] The term "nuclear localization sequence," "nuclear localization signal," or "NLS" refers to the localization of a protein. The term "nuclear localization sequence" refers to an amino acid sequence that promotes import into the cell nucleus. For example, WO / 2001 / 038547 filed on November 23, 2000 and published on May 31, 2001 is incorporated herein by reference. and Plank et al. in published international PCT application PCT / EP 2000 / 011690, in which No. 6,239,493, which is incorporated herein by reference for its disclosure of exemplary nuclear localization sequences. In embodiments, the NLS may be any of the NLSs described, for example, in Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4 172. In some embodiments, the NLS is Sequence KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR , RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.
[0187] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to a nucleic acid molecule comprising a nucleobase and Compounds containing an acidic moiety, such as nucleosides, nucleotides, or polynucleotides Typically, a polymeric nucleic acid, e.g., a nucleic acid molecule containing three or more nucleotides is a linear sequence in which adjacent nucleotides are linked to each other via phosphodiester bonds. In some embodiments, a "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or or nucleosides). In some embodiments, a "nucleic acid" refers to three or more individual nucleosides. As used herein, the term "oligonucleotide" refers to an oligonucleotide chain containing nucleotide residues. "Nucleotide" and "polynucleotide" refer to a polymer of nucleotides (e.g., may be used interchangeably to refer to a chain of at least three nucleotides. In this context, "nucleic acid" encompasses RNA and single- and / or double-stranded DNA. For example, genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome It may naturally occur in the context of a chromatid, or other naturally occurring nucleic acid molecule. On the other hand, nucleic acid molecules can be, for example, non-naturally occurring molecules, recombinant DNA or RNA, artificial chromosomes, Engineered genomes, or fragments thereof, or synthetic DNA, RNA, DNA / RNA hybrids, or or non-naturally occurring molecules containing non-naturally occurring nucleotides or nucleosides Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms may be used interchangeably. Nucleic acids include nucleic acid analogs, e.g., analogs having other than a phosphodiester backbone. purified from the source, produced using a recombinant expression system and optionally purified, or chemically synthesized In the case of chemically synthesized molecules, the nucleic acid may, where appropriate, be Nucleotides such as analogs with chemically modified bases or sugars, and backbone modifications, are also suitable. Nucleic acid sequences are shown in the 5' to 3' direction unless otherwise indicated. In some embodiments, nucleic acids contain natural nucleosides (e.g., adenosine, thymidine, guanine, adenosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine cytidine, and deoxycytidine; nucleoside analogs (e.g., 2-aminoadenosine, 2-thiocytidine); Othymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, 2- Aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5- Propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanine , O(6)-methylguanine, and 2-thiocytidine); chemically modified bases; biologically modified bases (e.g., methylated bases); inserted bases; modified sugars (e.g., 2'-fluororibose, ribose , 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages) or Includes these.
[0188] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" refers to Used interchangeably with "polynucleotide programmable nucleotide binding domain" A guide nucleic acid or guide polynucleotide that guides the napDNAbp to a specific nucleic acid sequence. It refers to a protein that associates with a nucleic acid (e.g., DNA or RNA) such as a nucleic acid (e.g., gRNA). In embodiments, a polynucleotide-programmable nucleotide binding domain is a polynucleotide-programmable DNA binding domain. In the present invention, the polynucleotide-programmable nucleotide binding domain is In one embodiment, the RNA-binding domain is programmable by oligonucleotides. The polynucleotide-programmable nucleotide binding domain is a Cas9 protein. The Cas9 protein binds to a specific DNA sequence complementary to the guide RNA. In some embodiments, the napDNAbp can bind to a guide RNA that guides the Cas9 domain. Main, e.g., nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease Inactive Cas9 (dCas9). Non-limiting examples of nucleic acid programmable DNA binding proteins. Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, and Cas12c / C2c3 Cas enzymes include Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i. Non-limiting examples include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12) ), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / Ca sX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5 e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3 , Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Cs x1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa 1, Csa2, Csa3, Csa4, Csa5, type II Cas effector protein, type V Cas effector protein Cas effector proteins, type VI Cas effector proteins, CARF, DinG, and their homologs Other nucleic acid programming may be used. Also included are DNA binding proteins that may not be specifically listed in this disclosure. It is within the scope of this disclosure, e.g., Makarova et al. CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct;1:325-336. doi: 10.10 89 / crispr.2018.0033; Yan et al., “Functionally diverse type V CRISPR-Cas system s” Science. 2019 Jan 4;363(6422):88-91. See doi: 10.1126 / science.aav7271 (see the entire contents of each of which are incorporated herein by reference).
[0189] The terms "nucleobase," "nitrogenous base," or "base" are used interchangeably herein. refers to nitrogen-containing biological compounds that form nucleosides, which are nucleotides The ability of nucleobases to form base pairs and stack with each other is directly related to the synthesis of ribonucleic acid ( Adenine (A) is a nucleotide that gives rise to long helical structures such as RNA and deoxyribonucleic acid (DNA). The five nucleic acid bases are cytosine (C), guanine (G), thymine (T), and uracil (U). They are called primary or standard bases. Adenine and guanine are derived from purines, while cytosine, Uracil and thymine are derived from pyrimidines. DNA and RNA contain modified other (non-primary) bases. Non-limiting exemplary modified nucleobases include hypoxanthine, xanthine, , 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydroxybenzoates. Hypoxanthine and xanthine are produced in the presence of mutagens. Both can be produced by deamination (replacement of an amine group with a carbonyl group). Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can be generated by deamination of cytosine. A "nucleoside" is a nucleic acid base and a five-carbon atom. They consist of a sugar (ribose or deoxyribose). Examples of nucleosides include adenosine , guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, These include deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine. Nucleosides with modified nucleobases include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), psoriasis Nucleotides consist of a nucleic acid base, a pentose sugar (ribose or is deoxyribose), and at least one phosphate group.
[0190] The terms "nucleobase editing domain" or "nucleobase editing protein" are used herein. When used in the formula, cytosine (or cytidine) to uracil (or uridine) or hypoxia to thymine (or thymidine) and from adenine (or adenosine) Deamination to sansine (or inosine), and non-templated nucleotide addition and insertion Proteins or enzymes that can catalyze nucleobase modifications in RNA or DNA, such as refers to an enzyme. In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., For example, adenine deaminase or adenosine deaminase; or cytidine deaminase In some embodiments, the nucleobase editing domain is a nucleobase editing domain (e.g., a nucleobase editing domain ... The enzymes may contain multiple deaminase domains (e.g., adenine deaminase or adenosine deaminase). aminase and cytidine or cytosine deaminase). The acid-base editing domain may be a naturally occurring nucleobase-editing domain. In embodiments, the nucleobase-editing domain is engineered from a naturally occurring nucleobase-editing domain. The nucleobase editing domain may be engineered or evolved in bacteria, Any living organism, such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse It can be of biological origin.
[0191] As used herein, "obtaining," as in "obtaining a drug," means This includes synthesizing, purchasing, or otherwise obtaining the drug.
[0192] As used herein, a "patient" or "subject" refers to a person who has been diagnosed with a disease or disorder. Mammals at risk of developing them or suspected of having or developing them In certain embodiments, the term "patient" refers to a subject or individual suffering from a disease or disorder. refers to a mammalian subject with a higher than average likelihood of developing a disease. Exemplary patients include humans, non-human primates, cats, dogs, pigs, cattle, horses, camels, llamas, goats, sheep, rodents (e.g. mice, rabbits, rats, guinea pigs) and will benefit from the treatments disclosed herein. Exemplary human patients include males and / or females. could be.
[0193] A "patient in need thereof" or a "subject in need thereof" is used herein to refer to a person with a disease or have been diagnosed with, are at risk of having, or may have a disability or disorder It is referred to as a patient who has been predetermined or is suspected of having it.
[0194] "Pathogenic mutation," "pathogenic variant," "disease-causing mutation," "disease-causing variant" The terms "deleterious mutation" or "predisposing mutation" refer to a mutation that is associated with a particular disease or disorder. Refers to a genetic change or mutation that increases an individual's susceptibility or predisposition. A pathogenic variant is a mutation that results in at least one wild-type mutation in the protein encoded by the gene. This includes those in which an amino acid has been replaced with at least one pathogenic amino acid.
[0195] The term "pharmaceutically acceptable carrier" refers to a liquid or solid filler, diluent, excipient, manufacturing agent, or Auxiliaries (e.g. lubricants, magnesium talc, calcium stearate or stearic acid zinc, or stearic acid), or solvent encapsulating materials, etc., in certain parts of the body (e.g. transport or transport of a compound from one site (e.g., a delivery site) to another site (e.g., an organ, tissue, or part of the body) Pharmaceutically acceptable materials, compositions, or vehicles involved in the delivery or transport of pharmaceuticals. A commercially acceptable carrier is one that is "compatible" with the other ingredients of the formulation and is not deleterious to the tissues of the subject. "Acceptable" (e.g., physiological compatibility, sterility, physiological pH, etc.). Terms such as "pharmaceutically acceptable carrier," "vehicle," and the like are used interchangeably herein. do.
[0196] The term "pharmaceutical composition" means a composition formulated for pharmaceutical use.
[0197] The terms "protein," "peptide," "polypeptide," and their grammatical equivalents are used interchangeably herein and are linked to each other by a peptide (amide) bond. The term refers to a polymer of amino acid residues bound together in a single chain. It typically refers to a protein, peptide, or polypeptide. A protein, peptide, or polypeptide is at least three amino acids in length. A polypeptide can refer to an individual protein or a group of proteins. One or more amino acids in a protein, peptide, or polypeptide may have, for example, a carbohydrate group, a hydrophobic group, xyl group, phosphate group, farnesyl group, isofarnesyl group, fatty acid group, linkage They may be modified by the addition of chemical entities such as anchors, functionalization, or other modifications. Proteins, peptides, or polypeptides may also be single molecules or multi-molecular. The protein, peptide, or polypeptide may be a naturally occurring It may be merely a fragment of a protein or peptide. Peptides may be naturally occurring, recombinant, or synthetic, or any of the foregoing. As used herein, the term "fusion protein" refers to a combination of: Hybrid polypeptides containing protein domains from at least two different proteins One protein is the amino-terminal (N-terminal) portion of the fusion protein or the capsid. can be located at the carboxy-terminus (C-terminus) of a protein, thus The protein forms a terminal or carboxy-terminal fusion protein. a nucleic acid binding domain (e.g., a nucleic acid binding domain that induces binding of a protein to a target site) Cas9 gRNA binding domain and nucleic acid cleavage domain, or nucleic acid editing protein In some embodiments, the protein may comprise a proteinaceous moiety, e.g., a catalytic domain of For example, an amino acid sequence constituting a nucleic acid binding domain and an organic compound, such as a nucleic acid cleaving agent In some embodiments, proteins include compounds that can act as nucleic acids (e.g., RNA). or DNA) are complexed with or associated with nucleic acids. Any protein can be produced by any method known in the art. For example, the proteins provided herein can be prepared via recombinant protein expression and purification. This is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and are described in Green and Samb rook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Labora and those described by the University of California Press, Cold Spring Harbor, NY (2012). the entire contents of which are incorporated herein by reference.
[0198] The polypeptides and proteins disclosed herein (including functional portions thereof and their functional variants) (including riboflavin) contains synthetic amino acids in place of one or more naturally occurring amino acids Such synthetic amino acids are known in the art and include, for example, aminocyclo Hexanecarboxylic acid, norleucine, α-amino n-decanoic acid, homoserine, S-acetylacetone aminomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenyl phenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine phenylalanine, β-phenylserine, β-hydroxyphenylalanine, phenylglycine, α-Naphthylalanine, cyclohexylalanine, cyclohexylglycine, indoline 2-carboxylic acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, Aminomalonic acid monoamide, N'-benzyl-N'-methyllysine, N',N'-dibenzyl-lysine, 6- Hydroxylysine, ornithine, α-aminocyclopentanecarboxylic acid, aminocyclohexyl cyclohexanecarboxylic acid, aminocyclohexanecarboxylic acid, α-aminocycloheptanecarboxylic acid carboxylic acid, α-(2-amino-2-norbornane)-carboxylic acid, α,γ-diaminobutyric acid, α,β-diamino These include monopropionic acid, homophenylalanine, and α-tert-butylglycine. Polypeptides and proteins are derived from post-translational modifications of one or more amino acids of the polypeptide construct. Non-limiting examples of post-translational modifications include phosphorylation, acetylation, and the like. and acylation, including formylation, glycosylation (including N-linked and O-linked), amino Derivatization, hydroxylation, alkylation including methylation and ethylation, ubiquitination, pyrolysis, Addition of lidocaine carboxylic acid, formation of disulfide bridges, sulfation, myristoylation, palmitoylation ylation, isoprenylation, farnesylation, geranylation, glypiation, lipoylation and iodination.
[0199] The term "recombinant" as used herein with respect to a protein or nucleic acid means that the protein or nucleic acid is It refers to proteins or nucleic acids that do not occur in nature but are the product of human engineering. For example, In some embodiments, the recombinant protein or nucleic acid molecule is any naturally occurring At least one, at least two, at least three, at least four, or at least at least five, at least six, or at least seven amino acid or nucleotide mutations Contains an octide sequence.
[0200] "Decrease" means a negative change of at least 10%, 25%, 50%, 75%, or 100% .
[0201] "Reference" means a standard or control condition. , the reference is a wild-type or healthy cell. In other embodiments, including but not limited to, the reference are not exposed to the test conditions or are exposed to placebo or normal saline, medium, or buffer. and / or exposed to a control vector that does not carry the polynucleotide of interest, It is a physiological cell.
[0202] A "reference sequence" is a defined sequence used as a basis for sequence comparison. It can be a subset or the entirety of a particular sequence; for example, a full-length cDNA or gene sequence. A segment of a gene, or the complete cDNA or gene sequence. For polypeptides, see the reference polypeptide. The length of the peptide sequence is generally at least about 16 amino acids, at least about 20 amino acids, At least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids. For nucleic acids, the length of a reference nucleic acid sequence is generally at least about 50 nucleotides, At least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides, or about 300 nucleotides thiazolinone, ... In one embodiment, the reference sequence is the wild-type sequence of the protein of interest. The sequence is a polynucleotide sequence that encodes the wild-type protein.
[0203] The terms "RNA programmable nuclease" and "RNA-guided nuclease" refer to Used in conjunction with (e.g., bound to or associated with) one or more non-target RNAs In some embodiments, the RNA-programmable nuclease, when complexed with RNA, Typically, the bound RNA is called a guide RNA (gRNA). can be.
[0204] In some embodiments, the RNA programmable nuclease is a Cas9 enzyme (CRISPR-associated system). endonuclease, such as Cas9 (Casnl) from Streptococcus pyogenes (e.g., For example, "Complete genome sequence of an Ml strain of Streptococcus pyogenes." Ferret ti JJ et al., Proc. Natl. Acad. Sci. USA 98:4658-4663(2001); "CRISPR RNA ma turation by trans-encoded small RNA and host factor RNase III." Deltcheva E. et al., Nature 471:602-607(2011)).
[0205] RNA-programmable nucleases (e.g., Cas9) use RNA:D to target DNA cleavage sites. Since NA hybridization is used, these proteins can, in principle, be used as guide RNAs. Any sequence specified by A can be targeted. For site-specific cleavage, C Methods using RNA-programmable nucleases such as as9 (e.g., to modify genomes) For this purpose, methods for generating multiplexed genomic DNA (e.g., Cong, L. et al., Multiplex gene expression) are known in the art (see, e.g., Cong, L. et al., Multiplex gene expression). ome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013 ); Hwang, WY et al., Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programme d genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al., G enome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic acids research (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes using CRISPR-Cas systems. See Nature biotechnology 31, 233-239 (2013). the entire contents of each of which are incorporated herein by reference).
[0206] The term "single nucleotide polymorphism (SNP)" refers to a single nucleotide variation that occurs at a specific position in the genome. where each mutation is present in the population to a noticeable extent (e.g., >1%). For example, At certain base positions in the human genome, C nucleotides can occur in most individuals, In a small number of individuals, the position is occupied by A. This means that there is a SNP at this particular position, and C or A, meaning that two nucleotide variations are alleles at this position. SNPs underlie differences in susceptibility to disease, affecting the severity of the disease and the body's response to treatment. SNPs are found in the coding region of genes, non-coding regions of genes, and so on. It may be present in a gene region, or in an intergenic region (the region between genes). In this case, SNPs within the coding sequence may affect the identity of the protein produced due to the degeneracy of the genetic code. SNPs in the coding region are of two types: synonymous SNPs and non-synonymous SNPs. Synonymous SNPs do not affect the protein sequence, whereas non-synonymous SNPs affect the amino acid sequence of a protein. There are two types of nonsynonymous SNPs: missense and nonsense. SNPs not in the coding region affect gene splicing, transcription factor binding, and messenger receptor (MR) activity. These types of SNPs can affect the degradation of RNA or the sequence of non-coding RNA. The gene expression that is affected is called an eSNP (expressed SNP) and can be upstream or downstream of the gene. Single nucleotide variants (SNVs) are single nucleotide variations with unlimited frequency and occur in somatic Somatic single base variations may also be referred to as single base modifications.
[0207] "Specifically bind" means to recognize and bind to the polypeptide and / or nucleic acid molecule of the present invention. and bind to other molecules in the sample (e.g., biological sample) but do not substantially recognize or bind to other molecules in the sample. Nucleic acid molecules, polypeptides, or complexes thereof (e.g., nucleic acid programmable DNA binding domain and guide nucleic acid), compound, or molecule.
[0208] Nucleic acid molecules useful in the methods of the present invention may encode a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules include any nucleic acid molecule that is 100% identical to the endogenous nucleic acid sequence. Typically, but not necessarily, substantial identity is shown. A polynucleotide having a nucleotide sequence typically hybridizes with at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention can be modified to The present invention also includes any nucleic acid molecule encoding an endogenous nucleotide sequence or a fragment thereof. It need not be 100% identical to the nucleic acid sequence, but typically will show substantial identity. Polynucleotides having "substantial identity" to a double-stranded nucleic acid molecule typically have a small number of double-stranded nucleic acid molecules. "Hybridize" means to hybridize with at least one of the strands of a given molecule. Complementary polynucleotide sequences (e.g., those described herein) are synthesized under various stringency conditions. This means that a double-stranded molecule is formed between the two genes (or a part of them). For example, Wahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, A. R (1987) Methods Enzymol. 152:507).
[0209] For example, a stringent salt concentration is typically less than about 750 mM NaCl and 75 mM citrate. trisodium, preferably less than about 500 mM NaCl and 50 mM trisodium citrate, more preferably Preferably, the concentration is less than about 250 mM NaCl and 25 mM trisodium citrate. See hybridization can be obtained in the absence of organic solvents, e.g., formamide. whereas high stringency hybridization requires at least about 35% formaldehyde. It can be obtained in the presence of at least about 50% formamide. Stringent temperature conditions are typically at least about 30°C, more preferably at least about 37°C. 0 C, and most preferably at least about 42°C. The concentration of detergent (e.g., sodium dodecyl sulfate (SDS)) and carrier DNA content Various additional parameters, such as inclusion or exclusion, are well known to those skilled in the art. By combining these various conditions, various levels of stringency can be achieved. In one embodiment, hybridization is achieved at 30° C. in 750 mM NaCl, 75 In another embodiment, hybridization occurs in 10 mM trisodium citrate and 1% SDS. The solution was incubated at 37°C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide. In another embodiment, the hybridization occurs in 100 μg / ml denatured salmon sperm DNA (ssDNA). Hybridization was performed at 42°C in 250 mM NaCl, 25 mM trisodium citrate, 1% S The reaction takes place in 500 μg / ml ssDNA, 50% formamide, and 200 μg / ml ssDNA. Useful variations of these conditions are The alternatives will be readily apparent to those skilled in the art.
[0210] For most applications, the washing steps that follow hybridization are also stringent. Wash stringency conditions are defined by salt concentration and temperature. As mentioned above, wash stringency can be increased by decreasing the salt concentration or increasing the temperature. This can be increased by increasing the thickness of the strip for the cleaning process. The appropriate salt concentration is preferably less than about 30 mM NaCl and 3 mM trisodium citrate. and most preferably less than about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for the process are typically at least about 25°C, more preferably In one embodiment, the temperature is at least about 42°C, and even more preferably at least about 68°C. Washing steps were performed at 25°C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the wash steps are carried out in 15 mM NaCl, 1.5 mM NaCl, at 42°C. 10 mM trisodium citrate, and 0.1% SDS. Washing steps were performed at 68°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Further variations of these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton et al. nd Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wil ey Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular C loning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York It has been done.
[0211] "Split" means divided into two or more pieces.
[0212] A "split Cas9 protein" or "split-Cas9" is a protein that is composed of two separate nucleotides. Cas9 protein provided as N-terminal and C-terminal fragments encoded by the sequences The polypeptides corresponding to the N-terminal and C-terminal parts of the Cas9 protein are spliced together. In certain embodiments, the Cas9 protein can be reconstituted to form a "reconstituted" Cas9 protein. The Cas9 protein is described, for example, in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-9 49, 2014, or Jiang et al. (2016) Science 351: 867-871. PDB As described in file: 5F9R (each of which is incorporated herein by reference), In some embodiments, the protein is split into two fragments within a disordered region of the protein. The protein binds to the SpCas9 protein between approximately amino acids A292-G364, F445-K483, or E565-T637. At any C, T, A, or S within the region, or any other Cas9, Cas9 variant ( For example, nCas9, dCas9), or other nap DNA fragments at corresponding positions in the two fragments. In some embodiments, the protein is split into SpCas9 T310, T313, A456, S469, or is split into two fragments at C574. In some embodiments, the protein is split into two fragments. The process of dividing a protein into multiple fragments is called "splitting" the protein. do.
[0213] In other embodiments, the N-terminal portion of the Cas9 protein is S. pyogenes Cas9 wild type (SpC as9) (NCBI Reference Sequence: NC_002737.2, Uniprot Reference Sequence: Q99ZW2) amino acids 1 to 573 The C-terminal part of the Cas9 protein contains amino acids 1 to 637, and the C-terminal part of the Cas9 protein contains amino acids 574 to 1368 of the wild-type SpCas9. includes the part 638 to 1368.
[0214] The C-terminal part of the split Cas9 is ligated with the N-terminal part of the split Cas9 to form the complete Cas9 tag. In some embodiments, the C-terminus of the Cas9 protein can form a protein. The end portion begins where the N-terminal portion of the Cas9 protein ends. In some embodiments, the C-terminal portion of the split Cas9 is amino acids (551-651)-1368 of spCas9. "(551-651) -1368" refers to the amino acids between 551 and 651 (inclusive). This means that the C-terminal portion of the split Cas9 begins with amino acid 1368 and ends with amino acid 1368. The amino acids 551-1368, 552-1368, 553-1368, 554-1368, 555-1368, and 556-1368 of spCas9 are , 557-1368, 558-1368, 559-1368, 560-1368, 561-1368, 562-1368, 563-1368, 564-1368 , 565-1368, 566-1368, 567-1368, 568-1368, 569-1368, 570-1368, 571-1368, 572-1368 , 573-1368, 574-1368, 575-1368, 576-1368, 577-1368, 578-1368, 579-1368, 580-1368 , 581-1368, 582-1368, 583-1368, 584-1368, 585-1368, 586-1368, 587-1368, 588-1368 , 589-1368, 590-1368, 591-1368, 592-1368, 593-1368, 594-1368, 595-1368, 596-1368 , 597-1368, 598-1368, 599-1368, 600-1368, 601-1368, 602-1368, 603-1368, 604-1368 , 605-1368, 606-1368, 607-1368, 608-1368, 609-1368, 610-1368, 611-1368, 612-1368 , 613-1368, 614-1368, 615-1368, 616-1368, 617-1368, 618-1368, 619-1368, 620-1368 , 621-1368, 622-1368, 623-1368, 624-1368, 625-1368, 626-1368, 627-1368, 628-1368 , 629-1368, 630-1368, 631-1368, 632-1368, 633-1368, 634-1368, 635-1368, 636-1368 , 637-1368, 638-1368, 639-1368, 640-1368, 641-1368, 642-1368, 643-1368, 644-1368 , 645-1368, 646-1368, 647-1368, 648-1368, 649-1368, 650-1368, or 651-1368 In some embodiments, the split Cas9 protein may comprise either one of the following portions: The C-terminal portion of SpCas9 includes amino acids 574-1368 or 638-1368 of SpCas9.
[0215] "Subject" means a mammal, including a human, or a bovine, equine, canine, ovine, or Subjects include, but are not limited to, non-human mammals such as cats. Subjects include livestock, labor Domestic animals (cows, goats, chickens) that are kept to produce power and provide goods such as food , horses, pigs, rabbits, and sheep).
[0216] "Substantially identical" means that the amino acid sequence of a reference amino acid sequence (e.g., an amino acid sequence described herein) is substantially identical to the amino acid sequence of a reference amino acid sequence (e.g., an amino acid sequence described herein). any one of the sequences) or nucleic acid sequence (e.g., any one of the nucleic acid sequences described herein) It means a polypeptide or nucleic acid molecule that exhibits at least 50% identity to In this form, such sequences are identified at the amino acid or nucleic acid level as sequences used for comparison. The sequence may have at least 60%, 80%, or 85%, 90%, 95% or even 99% identity in the sequence.
[0217] Sequence identity is typically determined using sequence analysis software (e.g., Genetics Computer Group , University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705 Sequence Analysis Software Package, BLAST, BESTFIT, GAP, or PILE UP / PRETTYBOX program). Such software is available with various permutations. By assigning degrees of homology to the sequences, deletions, and / or other modifications, it is possible to determine whether the sequences are identical or Similar sequences are matched. Conservative substitutions typically include substitutions within the following groups: leucine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, Paragine, glutamine; serine, threonine; lysine, arginine; phenylalanine In an exemplary approach to determining the degree of identity, the BLAST program You can use the RAM, -3 and e -100 Probability scores between indicate closely related sequences COBALT is used, for example, with the following parameters: a) Alignment parameters: Gap penalties -11, -1 and End-Gap penalties -5, -1 b) CDD parameters: Use RPS BLAST on; Blast E-value 0.003; Find Conserved column s and Recompute on c) Query clustering parameters: Use query clusters on; Word Size 4; M ax cluster distance 0.8; Alphabet Regular. The EMBOSS Needle is used, for example, with the following parameters: a) Matrix: BLOSUM62; b) GAP OPEN: 10; c) GAP EXTEND: 0.5; d) OUTPUT FORMAT: pair; e) END GAP PENALTY: false; f) END GAP OPEN: 10; and g) END GAP EXTEND: 0.5.
[0218] The term "target site" refers to a sequence within a nucleic acid molecule that is modified by a nucleobase editor. In one embodiment, the target site is a sequence that is targeted by a deaminase (e.g., a cytidine or or adenine deaminase) or a fusion protein containing it .
[0219] As used herein, the terms "treat," "treating," and "treatment" refer to "Treatment" and the like are intended to alleviate or improve a disorder and / or its associated symptoms. refers to the process of achieving a desired pharmacological and / or physiological effect. Treating a condition requires the complete elimination of the associated disorder, condition, or symptom. It will be understood that some aspects of the In some instances, the effect is therapeutic, i.e., the effect is, but is not limited to, the effect of treating a disease. Partially or completely reduce, diminish, eliminate or alleviate the disease and / or adverse symptoms resulting therefrom. In some embodiments, the effect is preventative, i.e., The effect is to protect or prevent the occurrence or recurrence of a disease or condition. The disclosed methods comprise administering a therapeutically effective amount of a composition as described herein. include.
[0220] "Uracil glycosylase inhibitor," or "UGI," is a protein that inhibits the uracil excision repair system. In one embodiment, the agent inhibits host uracil-DNA glycosylation. A protein or fragment thereof that binds to uracil and prevents the removal of uracil residues from DNA. In one embodiment, the UGI inhibits uracil-DNA glycosylase base excision repair enzyme. In some embodiments, the target protein is a protein, fragment, or domain thereof that can inhibit the target protein. In some embodiments, the UGI domain comprises wild-type UGI or a modified version thereof. In the present invention, the UGI domain comprises a fragment of the exemplary amino acid sequence provided below. In some embodiments, the UGI fragment comprises at least 60% of the exemplary UGI sequences provided below: At least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least In some embodiments, the UGI comprises an amino acid sequence that is at least 99%, or 100%, of the amino acid sequence of As described below, amino acids homologous to exemplary UGI amino acid sequences or fragments thereof may be used. In some embodiments, the UGI or a portion thereof comprises a sequence as described below. For example, at least 70%, at least 75% of the wild-type UGI or UGI sequence or a portion thereof , at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9% or have 100% identity. Exemplary UGIs include the following amino acid sequences: >splP14739IUNGI_BPPB2 Uracil-DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSD APEYKPWALVIQDSNGEN KIKML.
[0221] The term "vector" refers to a nucleic acid sequence that is introduced into a cell, resulting in a transformed cell. Vectors include plasmids, transposons, phages, viruses, and ribosomal An "expression vector" includes a vector that is expressed in a recipient cell. An expression vector is a nucleic acid sequence containing a nucleotide sequence that encodes a gene encoding a gene. Enhances expression of introduced sequences, such as promoter, secretory, and other sequences. and / or may contain additional nucleic acid sequences for facilitating
[0222] The compositions and methods provided herein may be used in combination with other compositions and methods provided herein. It can be combined with any one or more of the methods.
[0223] DNA editing is a method to correct disease states by correcting pathogenic mutations at the genetic level. Until recently, all DNA editing platforms induces DNA double-strand breaks (DSBs) at specific genomic sites and is partially repaired depending on endogenous DNA repair pathways. It functions by determining product outcomes in a probabilistic manner, resulting in a complex population of genetic products. Achieving accurate and user-defined repair outcomes via the homology-directed repair (HDR) pathway Although it is possible to achieve high resolution imaging using HDR in therapeutically relevant cell types, many challenges remain. Efficient repair is hindered. In reality, this pathway involves competing, error-prone, and non-correlated pathways. HDR is less efficient than the homo-end joining pathway. Furthermore, HDR is strictly confined to the G1 and S phases of the cell cycle. This limits the ability of DSBs to be repaired accurately in post-mitotic cells. Highly efficient genome sequencing in these populations in a user-defined and programmable manner It has proven difficult or impossible to alter [Brief explanation of the drawings]
[0224] [Figure 1] Figures 1A-C show the cis-trans activity of free deaminase. Figure 1A is a schematic showing the experimental design of the cis-trans assay for SpCas9 and deaminase in either the base editor complex format or the unlinked format. Figure 1B is a graph showing the cis-trans activity of rAPOBEC. Figure 1C is a graph showing the cis-trans activity of TadA7.10 and TadA-TadA7.10. [Figure 2]Figures 2A-2F show the cis-trans assay for base editors, an example of a deaminase similarity network, and the screening of 153 deaminases. Figure 2A is a schematic diagram illustrating the experimental design for the cis-trans assay. HEK293T cells were transfected with separate plasmids encoding SaCas9, gRNA for SaCas9, and the target base editor. Figure 2B is a schematic diagram illustrating the similarity network for APOBEC-like deaminases. Dots represent cytidine deaminases screened as next-generation CBEs, and the core next-generation CBE is shown. The shading of the dots represents the average trans / cis ratio, and the size of the dots represents the average cis activity. The method for creating the similarity network for cytidine deaminases shown in Figure 2B is as follows: To focus the search space within the APOBEC1-like protein family, a protein BLAST search was performed against the NCBI nonredundant protein sequence database (nr_v5) using human APOBEC1 as the query sequence. A sequence similarity network (SSN) was generated using the top 1,000 sequences using a protein BLAST-log(E-value) edge threshold of 115. A set of 43 deaminases was selected to sample the sequence space within the SSN. To identify deaminases from other families that may act as base editing enzymes, 80 sequences from the SSN constructed from all deaminases were sampled using the following InterPro annotations: IPR002125 (cytidine and deoxycytidylate deaminase domain), IPR016192 (APOBEC / CMP deaminase, zinc-binding), and IPR016193 (cytidine deaminase-like). This set of 82,043 sequences was first clustered at 55% identity using Cd-HIT3, and then an SSN network was generated using protein BLAST with a -log(E-value) edge threshold of 50. Sequences were selected based on the centrality of the sequences within the cluster in the network. Figure 2C is a graph showing the cis-trans activity of ppBE4 and its mutants. Figure 2D is a graph showing the cis-trans activity of selected editors.Separately, cis-trans activity data were generated based on in cis / in trans assays for three target sites, namely, site 1, site 4, and site 6, as shown in Figure 2E and Figure 2F. Figure 2E is a bar graph showing the cis- and trans-editing activities of the identified CBEs. Shown is a comparison of the cis- and trans-editing frequencies in mammalian cells treated with candidate CBEs. Editor numbers 1 to 36 are the base editors pYY-BEM3.8, pYY-BEM3.9, pYY-BEM3.10, pYY-BEM3.11, pYY-BEM3.12, pYY-BEM3.13, pYY-BEM3.14, pYY-BEM3.15, pYY-BEM3.16, pYY-BEM3.17, pYY-BEM3.18, pYY-BEM3.19, pYY-BEM3.20, pYY-BEM3.21, pYY-BEM3.22, pYY-BEM3.23, pYY-BEM3.24, and pYY-BEM3.25, respectively. The base editing efficiencies are reported for the most edited base at the target site. Figure 2F is a bar graph showing the cis- and trans-editing activities of the identified CBEs. A comparison of cis- and trans-editing frequencies in mammalian cells treated with candidate CBEs is shown.Editor numbers 1 to 37 are rBE4max, mAPOBEC-1, MaAPOBEC-1, hAPOBEC-1, ppAPOBEC-1, OcAPOBEC1, MdAPOBEC-1, mAPOBEC-2, hAPOBEC-2, ppAPOBEC-2, BtAPOBEC-2, mAPOBEC-3, hAPOBEC-3A, hAPOBEC-3B, hAPOBEC-3C, hAPOBEC-3D, and hAPO These are BEC-3F, hAPOBEC-3G, hAPOBEC-4, mAPOBEC-4, rAPOBEC-4, MfAPOBEC-4, hAID, negative control, btAID, mAID, pmCDA-1, pmCDA-2, pmCDA-5, yCD, pYY-BEM3.1, pYY-BEM3.2, pYY-BEM3.3, pYY-BEM3.4, pYY-BEM3.5, pYY-BEM3.6, and pYY-BEM3.7. Base editing efficiencies are reported for the most edited base at the target site. [Figure 3] Figures 3A and 3B show cis-trans activity. Figure 3A is a graph showing the cis-trans activity of ABE7.10. Figure 3B is a graph showing the cis-trans activity of BE4 max. [Figure 4] Figures 4A and 4B show the rAPOBEC1 homology model generated by SWISSMODEL using the hAPOBEC3C structure (PDB ID 3VM8). ssDNA from the hAPOBEC3A structure (PDB ID 5SWW) has been manually docked. Figure 4A is a schematic showing mutations that potentially affect ssDNA binding. Figure 4B is a schematic showing mutations that potentially affect catalytic activity. [Figure 5] 5A-5C show the cis-trans activity of rAPOBEC1 mutants. [Figure 6]Figures 6A to 6E show the cis-trans activity of the rAPOBEC1 double mutant. Figure 6A is a graph showing the in cis and in trans activity of the rAPOBEC1 double mutant. Figure 6B is a graph showing the in cis activity at six sites. Figure 6C is a graph showing the cis / trans ratio. Figure 6D is a graph showing the in cis activity at six sites. Figure 6E is a graph showing the cis / trans ratio. [Figure 7] Figures 7A and 7B show the cis-trans activity of the deaminases in the first round of screening. [Figure 8] 8A to 8C are graphs showing the targeting activity of ppAPOBEC1 against rAPOBEC1. [Figure 9] FIG. 9 is a schematic diagram showing the similarity network of APOBEC-like proteins. [Figure 10] Figures 10A and 10B are graphs showing dose-dependent studies of in cis and in trans activity in TadA-TadA7.10 and rAPOBEC1, respectively.
[0225] [Figure 11] Figure 11 is a graph showing off-target editing of selected CBEs. SNVs were identified by exome sequencing. [Figure 12] Figures 12A and 12B are graphs showing quantification of base editor mRNA and protein, respectively, from HEK293T cells transfected with base editor plasmids. [Figure 13] Figure 13 is a graph showing targeted RNA sequencing for selected editors. Three regions of 200-300 bp were sequenced. [Figure 14] FIG. 14 is a graph showing selected CBE-guided off-target editing. [Figure 15] 15A to 15E show the editing window of the selected editor. [Figure 16]FIG. 16 is a graph showing the indel rates of selected CBEs at 10 target sites. [Figure 17] Figures 17A-17D show diagrams and graphs for unguided ssDNA deamination and cis / trans assays. Figure 17A shows potential ssDNA formation in the genome during transcription or translation. Figure 17B shows the experimental design for the in cis / in trans assay. HEK293T cells were transfected with separate constructs encoding SaCas9, a gRNA for SaCas9, and a base editor. In cis and in trans activity was measured in different transfectants but at the target site defined by the NGGRRT PAM sequence. Figure 17C shows the in cis / in trans activity of BE4 with rAPOBEC1. Figure 17D shows the ABE7.10 variant at 34 genomic sites. The leftmost bar of each genomic site on the x-axis indicates on-target editing in cis. The rightmost bar of each genomic site on the x-axis indicates trans editing. Base editing efficiency is reported for the most edited base at the target site. Values and error bars reflect the mean and standard deviation (s.d.) of independent biological replicates. [Figure 18] Figure 18 is a bar graph showing identified next-generation CBEs with enhanced cis activity and reduced trans activity compared to BE4 with rAPOBEC1. Shown is a comparison of in cis and in trans editing frequencies at 10 genomic sites in mammalian cells treated with next-generation CBEs (BE4 with PpAPOBEC1 [wt, H122], RrA3F [wt, F130L], AmAPOBEC1, and SsAPOBEC2 [wt, R54Q]). Base editing efficiencies are reported for the most edited base at the target site. Values and error bars reflect the mean and standard deviation (SD) of four independent biological replicates. [Figure 19]Figures 19A-19E show allele frequencies and graphs for next-generation CBEs with reduced DNA and RNA off-target editing compared to BE4 in mammalian cells. Figure 19A shows whole transcriptome sequencing and targeted RNA sequencing (Figure 19B) of Hek293T cells expressing an off-target deamination-minimizing cytosine base editor. Figure 19C shows the percentage of C→T editing at known guiding off-target sites. Figure 19D shows the percentage of C→T editing in an in vitro enzymatic assay on a single-stranded DNA substrate. C→U editing of core next-generation CBEs on ssDNA substrates. Dots represent the NC local sequence context of editing. The black line indicates the average editing efficiency of the target cytosine in the substrate. Figure 29E shows the time course of product formation in an in vitro enzymatic assay from cell lysates containing selected CBEs. The sequences of the oligos used in Figures 19D and 19E are listed in the table provided in Example 5 below. Values and error bars reflect the mean and sd of independent triplicate biological replicates (Fig. 19 A, B, C) or duplicate replicates (Fig. 19 D, E). [Figure 20] Figure 20 graphically depicts the cis / trans editing activity of BE4 with the rAPOBEC1 mutants shown in Figures 4A and 4B at site 1. Base editing efficiencies are reported for the most edited base at the target site. In trans efficiencies are shown at the left edge of each target site on the x-axis; in cis efficiencies are shown in the right bar for each target on the x-axis. Values and error bars reflect the mean and SD of independent biological replicates.
[0226] [Figure 21] Figure 21 shows the cis / trans editing activity of BE4-rAPOBEC1 with HiFi mutations at 10 target sites. Values and error bars reflect the mean and SD of four independent biological replicates. [Figure 22]Figures 22A and 22B show graphs and sequence alignments relating to the cis / trans editing activity and sequence alignment of CBEs tested in the first round of screening. In cis / in trans editing activity (Figure 22A) and sequence alignment (Figure 22B) at site 10 of selected CBEs. Amino acid residues flanking the HiFi mutations in rAPOBEC1 are highlighted. Values and error bars reflect the mean and SD of independent biological replicates. [Figure 23] Figure 23 shows the in cis / in trans activity of BE4-PpAPOBEC1 and BE4-PpAPOBEC with HiFi mutations at 10 target sites. Base editing efficiencies are reported for the most edited base at the target site. Values and error bars reflect the mean and SD of four independent biological replicates. [Figure 24] Figure 24 shows a heatmap illustrating the leading base preference of the CBE shown in Figure 18 B. The values used to generate the heatmap reflect the average of four independent biological replicates. [Figure 25] Figure 25 shows the editing window of the CBE shown in Figure 18B at 10 target sites. Values reflect the average of four independent biological replicates. In cis and in trans editing are displayed in the heatmaps in the leftmost and rightmost panels, respectively. [Figure 26] Figure 26 is a table showing the indel rates of the CBEs shown in Figure 18B at the 10 target sites. The values used to generate the heatmap reflect the average of four independent biological replicates. [Figure 27]Figures 27A-27D show homology models of four cytidine deaminases selected based on existing crystal structures. Figure 27A: The homology model of PpAPOBEC1 is based on the predicted APOBEC3G structure (PDB ID 5K81). Figure 27B: RrA3F is based on the Vif-binding domain of hAPOBEC3F (PDB ID 3WUS). Figure 27C: AmAPOBEC1 is based on the hAPOBEC3B N-terminal domain (PDB ID 5TKM). Figure 27D: SsAPOBEC2 is based on the Vif-binding domain of hAPOBEC3F (PDB ID 3WUS). [Figure 28] Figures 28A-28D are graphs showing the guided off-target editing of selected next-generation CBEs. Figure 28A: Editing efficiency of next-generation CBEs at HEK2, HEK3, and HEK4 sites, and Figure 28B: Editing efficiency of next-generation CBEs at reported guided off-target editing sites for HEK2 sgRNA, HEK3 sgRNA, and Figure 28D: HEK4 sgRNA. Base editing efficiency is reported for the most edited base at the target site. Values and error bars reflect the mean and SD of three independent biological replicates. [Figure 29] Figure 29 is a graph showing the C→T editing efficiency of selected CBEs on ssDNA substrates in an in vitro enzymatic assay. Editing efficiency was measured for all 25 cytidines in two ssDNA substrates and grouped by NC sequence context. The sequences of the two substrates used are listed in Table 18 herein. Values and error bars reflect the mean and SD of data obtained from two independent biological replicates. [Figure 30] Figure 30 shows a graph showing quantification of CBE protein concentration in HEK293T cells transfected with base editor expression plasmids. Base editor protein concentration was quantified by measuring total Cas9 protein concentration and total protein amount in cell lysates. BE protein concentration was normalized to BE4-rAPOBEC1. Values and error bars reflect the mean and SD of data obtained from two or more independent biological replicates. [Figure 31] Figure 31 presents a graph showing the off-target deamination activity of CBEs as determined by whole genome sequencing (WGS). Relative mutation rates are shown as odds ratios. DETAILED DESCRIPTION OF THE INVENTION
[0227] The present invention provides an improved editing profile with minimized off-target deamination. Nucleic acid base editors and multi-effector nucleobase editors, such Compositions comprising the editor, as well as compositions comprising the editor to generate modifications in target nucleobase sequences. Provides a method for using
[0228] [Nucleobase Editor] A base editor for editing, modifying or altering a target nucleotide sequence of a polynucleotide. The present invention relates to a nucleic acid base editor or a multi-effector nucleic acid base editor. Described herein are polynucleotide-programmable a suitable nucleotide-binding domain (e.g., Cas9) and at least one nucleobase-editing domain Nucleobase edta containing enzymes (adenosine deaminase and / or cytidine deaminase) a nucleic acid base editor or a multi-effector nucleic acid base editor. Nucleotide-programmable nucleotide-binding domains (e.g., Cas9) are When combined with a guide polynucleotide (e.g., gRNA), The target nucleic acid is synthesized by the complementary base pairing between the bases of the target polynucleotide sequence and the bases of the target nucleic acid. capable of specifically binding to a target polynucleotide sequence and thereby being edited The base editor can be localized to a target nucleic acid sequence where modification is desired.
[0229] Polynucleotide-programmable nucleotide-binding domains Polynucleotide programmable nucleotide binding domains also bind to nuclear RNA. It is understood that the present invention may include an acid programmable protein. The polynucleotide-programmable nucleotide binding domain is The nucleotide-binding domain can be linked to a nucleic acid that guides the RNA. Other DNA-binding proteins are also within the scope of this disclosure, although they are not specifically listed in this disclosure. It has not been done.
[0230] The polynucleotide-programmable nucleotide-binding domain of the base editor is The polynucleotide programmable vector itself can contain one or more domains. The nucleotide-binding domain capable of binding to the nuclease may include one or more nuclease domains. In some embodiments, the nucleic acid of a polynucleotide-programmable nucleotide binding domain The nuclease domain can comprise an endonuclease or an exonuclease. As used herein, the term "exonuclease" refers to an enzyme that liberates nucleic acids (e.g., RNA or DNA). The term "endonuclease" refers to a protein or polypeptide that can be digested from its termini. A "clease" is a nucleic acid that can catalyze (e.g., cleave) an internal region of a nucleic acid (e.g., DNA or RNA). In some embodiments, an endonuclease refers to a protein or polypeptide that It is capable of cleaving a single strand of double-stranded nucleic acid. Both strands of a double-stranded nucleic acid molecule can be cleaved. The programmable nucleotide binding domain can be a deoxyribonuclease. In some embodiments, the polynucleotide programmable nucleotide binding domain is a ribonucleotide. It may be a nuclease.
[0231] In one embodiment, the nucleotide of the polynucleotide-programmable nucleotide binding domain The cleavage domain can cleave zero, one, or two strands of a target polynucleotide. In one embodiment, the polynucleotide-programmable nucleotide binding domain The protein may comprise a nickase domain. " refers to a nucleic acid that can cleave only one of the two strands in a double-stranded nucleic acid molecule (e.g., DNA). a polynucleotide-programmable nucleotide-binding domain containing a cleavage domain; In some embodiments, the nickase is an active polynucleotide programmable nuclease. By introducing one or more mutations into the nucleotide binding domain, polynucleotide protease activity can be increased. Derived from a fully catalytically active (e.g., native) form of a programmable nucleotide-binding domain For example, polynucleotide programmable nucleotide binding domains can be used. If the gene contains a nickase domain derived from Cas9, the nickase domain derived from Cas9 The protein may contain a D10A mutation and a histidine at position 840. In embodiments, residue H840 retains catalytic activity, thereby cleaving a single strand of a nucleic acid duplex. In another example, the nickase domain from Cas9 contains an H840A mutation. while the amino acid residue at position 10 remains D. In some embodiments, the nickase may comprise all of the nuclease domain that is not required for nickase activity. or by removing a portion of the polynucleotide, programmable nucleotide binding It can be derived from a fully catalytically active (e.g., native) form of the domain. A nucleotide-programmable nucleotide-binding domain derived from Cas9 is used to identify nickases If the domain is included, the nickase domain from Cas9 is a RuvC domain or an HNH domain. It may contain a deletion of all or part of the domain.
[0232] The amino acid sequence of an exemplary catalytically active Cas9 is as follows: MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD.
[0233] Thus, polynucleotides containing nickase domains can be used as programmable nucleotides. Base editors containing a base-binding domain can be used to bind to specific polynucleotide target sequences (e.g., binding domains). The DNA fragments are then ligated to the target nucleic acid, which then generates a single-stranded DNA break (a nick) at the target site (determined by the complementary sequence of the guide nucleic acid). In some embodiments, a nickase domain (e.g., a nickase domain derived from Cas9) can be used. Nucleic acid double-stranded target polynucleotides cleaved by base editors containing a base editor domain (a cleavage enzyme domain) The strand of the base sequence is the strand that is not edited by the base editor (i.e., (The strand cleaved by the cleavage is the opposite strand to the strand containing the base to be edited.) and base editors containing a nickase domain (e.g., a nickase domain derived from Cas9). can cleave the strand of a DNA molecule targeted for editing. In this state, the non-target strand is not cleaved.
[0234] Catalytically dead (i.e., unable to cleave target polynucleotide sequence) polynucleotides Base editors comprising nucleotide-programmable nucleotide-binding domains are also described herein. As used herein, the terms "catalytically dead" and "nuclease inactive" has one or more mutations and / or deletions that result in the inability to cleave a strand of nucleic acid. Polynucleotides with deletions are replaced to refer to programmable nucleotide binding domains. In some embodiments, catalytically dead polynucleotide proteases are used. A gram-capable nucleotide-binding domain base editor binds one or more nuclease domains Nuclease activity can be lost as a result of specific point mutations in the In the case of base editors containing the Cas9 domain, Cas9 is able to reverse the D10A and H840A mutations. Such a mutation would inactivate both nuclease domains. In another embodiment, catalytically dead polynucleotides are The nucleotide-programmable nucleotide-binding domain is a catalytic domain (e.g., RuvC1 and and / or the HNH domain). In embodiments, catalytically dead polynucleotide programmable nucleotide binding domains The mutants contain point mutations (e.g., D10A or H840A) as well as all or part of the nuclease domain. Contains some deletions.
[0235] Also provided herein are the following polynucleotide-programmable nucleotide binding domains: Catalytically dead polynucleotide programmable nucleic acid from a previously functioning version Mutations that can generate peptide-binding domains are also contemplated. In the case of dead Cas9 ("dCas9"), mutations other than D10A and H840A are present, leading to the nuclease Variants that result in inactive Cas9 are provided. Such mutations include, for example, D10 and other amino acid substitutions at H840 or other substitutions within the nuclease domain of Cas9. (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain) Additional suitable nuclease-inactive dCas9 domains are described in this disclosure and in the art. Such modifications would be apparent to those skilled in the art based on their knowledge and are within the scope of the present disclosure. Exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to: Contains D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (e.g., Prashant et al., CAS9 transcriptional activators for target specific ficity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013; 31(9): 833-838, the entire contents of which are incorporated herein by reference. incorporated into the book).
[0236] Polynucleotide programmable nucleotides that can be incorporated into base editors Non-limiting examples of binding domains include domains derived from CRISPR proteins, restriction nucleases, and the like. Enzymes, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases In some embodiments, the base editor is a CRISP gene for nucleic acids. R (i.e., Clustered Regularly Interspaced Short Palindromic Repeats)-mediated modifications In this case, natural or modified nucleic acids that can bind to a nucleic acid sequence via a binding guide nucleic acid are used. Polynucleotide-programmable nucleotide-binding domains containing proteins or portions thereof Such proteins are referred to herein as "CRISPR proteins." Thus, disclosed herein are methods for detecting CRISPR proteins, including all or part of the proteins. Polynucleotide base editors containing programmable nucleotide binding domains (i.e. That is, a base editor that contains all or part of the CRISPR protein as a domain (which is It is also called the "CRISPR protein-derived domain" of the base editor. The CRISPR protein-derived domain integrated into the CRISPR receptor is expressed as a wild-type or naturally occurring CRISPR protein. For example, as described below, CRISPR proteins can be modified relative to the CRISPR protein. The CRISPR domain may have one or more mutations, as compared to a wild-type or naturally occurring CRISPR protein, It may involve insertions, deletions, rearrangements and / or recombinations.
[0237] CRISPR provides defense against mobile genetic elements (viruses, transposable elements, conjugative plasmids) CRISPR clusters are composed of spacers, sequences that are complementary to the preceding mobile element. The CRISPR cluster contains a sequence of target nucleic acids, and the target invader nucleic acid. The CRISPR cluster is transcribed into CRISPR RNA (crRNA). In type II CRISPR systems, correct processing of pre-crRNA is essential for transcription. transcoding small RNA (tracrRNA), endogenous ribonuclease 3 (rnc), and Cas9 protein tracrRNA guides the processing of pre-crRNA by ribonuclease 3. Cas9 / crRNA / tracrRNA then binds to a linear or circular dsDNA target complementary to the spacer. The target strand that is not complementary to the crRNA is first cleaved by the endonuclease. It is cleaved protease-wise and then exonucleolytically trimmed to 3'-5'. In the case of mitochondrial DNA, both the protein and the RNA are required for DNA binding and cleavage. A single guide RNA ("sgRNA", For example, Jinek M. et al., Science 337: 816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 targets a short motif (PAM or protospacer adjacent motif) in the CRISPR repeats. It helps us to recognize and distinguish between "self" and "non-self."
[0238] In some embodiments, the methods described herein involve recombinantly engineered The guide RNA (gRNA) is required for Cas binding. The required scaffold sequence and a user-defined approximately 20-base spacer that defines the genomic target to be modified. Therefore, those skilled in the art can identify the genomic target of Cas protein specificity. The gRNA targeting sequence can be varied to target the genome relative to the rest of the genome. This is determined in part by how specific it is to the target.
[0239] In some embodiments, the gRNA scaffold sequence is: GUUUUAGAGC UAGAAAU AGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAAGU GGCACCGAGU CGGUGCUUUU.
[0240] In some embodiments, the base editor comprises a CRISPR protein-derived domain. The main component is capable of binding to a target polynucleotide when combined with a binding guide nucleic acid. Endonucleases (e.g., deoxyribonucleases or ribonucleases) that can In some embodiments, the CRISPR protein incorporated into the base editor The protein-derived domain binds to the target polynucleotide when combined with the binding guide nucleic acid. In some embodiments, the base editor is a nickase that can bind to The CRISPR protein-derived domains integrated into the CRISPR protein bind to the CRISPR protein when combined with the guide nucleic acid. A catalytically dead domain is capable of binding to a target polynucleotide when In embodiments, a target polynucleotide that binds to a CRISPR protein-derived domain of a base editor is The nucleic acid is DNA, and in some embodiments, the nucleic acid comprises a CRISPR protein-derived domain of a base editor. The target polynucleotide that binds to the in is RNA.
[0241] The CAs proteins that can be used herein include class 1 and class 2. Non-limiting examples of proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5d, Cas5t, as5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also called Csn1 or Csx12), Cas10, Csy1, C sy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5, Csn1, Csn2, Csm2, Csm3, Csm4, Csm 5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Cssx16, Cx16, Cx, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1 , Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i, CARF, DinG, and their homologs Unmodified CRISPR enzymes, like Cas9, are composed of two It has functional endonuclease domains, RuvC and HNH, and can have DNA cleavage activity. CRISPR enzymes target sequences, such as within the target sequence and / or the complementary strand of the target sequence. For example, CRISPR enzymes can induce cleavage of one or both strands of a target sequence. Approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 25 nucleotides from the first or last nucleotide of the string Induce breaks in one or both strands at 50, 100, 200, 500 base pairs or more. It is possible.
[0242] Loss of ability to cleave one or both strands of a target polynucleotide containing the target sequence To achieve this, vectors were used to encode CRISPR enzymes that were mutated relative to the corresponding wild-type enzyme. Cas9 can be a wild-type exemplary Cas9 polypeptide (e.g., Cas9 from S. pyogenes). Cas9) and at least approximately 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93% %, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology Cas9 can refer to a polypeptide having a wild-type exemplary Cas9 polypeptide ( For example, from S. pyogenes), at most, or at most, approximately, about 50%, 60%, 70% , 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or polypeptides having sequence homology. or deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any of these It can refer to modified forms of the Cas9 protein that may contain amino acid changes, such as combinations.
[0243] In some embodiments, the CRISPR protein-derived domain of the base editor is Corynebacterium diphtheria (NCBI Refs: NC_015683.1, NC_017317.1); NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021 284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (N CBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella ba ltica (NCBI Ref: NC_018010.1); Psychroflexus torquis (NCBI Ref: NC_018721.1); St reptococcus thermophilus (NCBI Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP _472073.1); Campylobacter jejuni (NCBI Ref: YP_002344900.1); Neisseria meningiti dis (NCBI Ref: YP_002342100.1), Streptococcus pyogenes, or Staphylococcus au The Cas9 fragment may comprise all or part of the Cas9 fragment derived from Cas9.
[0244] [Cas9 domain, a nucleobase editor] The sequence and structure of Cas9 nuclease are well known to those skilled in the art (see, e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes.” Ferretti et al., Pr oc. Natl. Acad. Sci. USA 98:4658-4663(2001); “CRISPR RNA maturation by trans -encoded small RNA and host factor RNase III.” Deltcheva E. et al., Nature 471: 602-607(2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptation See Jinek M. et al., Science 337:816-821 (2012). The entire contents of which are incorporated herein by reference.) Cas9 orthologs include, but are not limited to: It has been described in various species, including S. pyogenes and S. thermophilus, although not exclusively in the Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on the present disclosure. Such Cas9 nucleases and sequences are described in Chylinski, Rhun, and Charpentier, “ The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) R NA Biology 10:5, 726-737. The entire contents of which are incorporated herein by reference.
[0245] In some embodiments, nucleic acid programmable DNA binding proteins (napDNAbp) is a Cas9 domain. Non-limiting exemplary Cas9 domains are provided herein. The 9 domains are divided into nuclease-active Cas9 domains, nuclease-inactive Cas9 domains (dCas9 ), or Cas9 nickase (nCas9). In some embodiments, the Cas9 domain can be Nuclease activity domains. For example, the Cas9 domain binds both strands of a double-stranded nucleic acid (e.g., For example, it may be a Cas9 domain that cleaves both strands of a double-stranded DNA molecule. In some embodiments, the Cas9 domain comprises any one of the amino acid sequences described herein. In some embodiments, the Cas9 domain comprises any of the amino acid sequences described herein. At least 60%, at least 65%, at least 70%, at least 75%, or at least at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least or at least 97%, at least 98%, at least 99%, or at least 99.5% identical amino acid sequence. In some embodiments, the Cas9 domain comprises an amino acid sequence described herein. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 5, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 3 5, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more In some embodiments, the Cas9 domain comprises an amino acid sequence having a mutation. At least 10, at ... At least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 20 0, at least 250, at least 300, at least 350, at least 400, at least 500, a few At least 600, at least 700, at least 800, at least 900, at least 1000, at least amino acid sequences that have at least 1100 or at least 1200 identical stretches of amino acid residues Includes.
[0246] In some embodiments, proteins comprising fragments of Cas9 are provided. In embodiments, the protein includes one of the following two Cas9 domains: (1) Cas9 (2) a gRNA binding domain of Cas9; and (3) a DNA cleavage domain of Cas9. In some embodiments, Cas9 or a fragment thereof Proteins containing the fragments are referred to as "Cas9 variants." Cas9 variants are Cas9 or For example, a Cas9 variant may share at least about 70% homology with a wild-type Cas9. % identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, In some embodiments, the sequences are about 99.5% identical, or at least about 99.9% identical to each other. Cas9 mutants have the following advantages compared to wild-type Cas9: 4, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 3 4, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more In some embodiments, the Cas9 variant may have an amino acid change in the fragment of Cas9. The fragment contains a fragment of wild-type Cas9 (e.g., a gRNA binding domain or a DNA cleavage domain), At least about 70% identical to the corresponding fragment, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least In some embodiments, the fragment is at least about 99.5% identical, or at least about 99.9% identical. , at least 30%, at least 35%, at least 40% of the amino acid length of the corresponding wild-type Cas9; At least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least or at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least In some embodiments, the fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 50 0, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 12 50, or at least 1300 amino acids in length.
[0247] In some embodiments, the Cas9 fusion proteins provided herein comprise a Cas9 protein. The full-length amino acid sequence of the protein, for example, one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein contain a full-length Cas9 sequence. Examples of Suitable Cas9 Domains and Cas9 Fragments Suitable amino acid sequences are provided herein, and further suitable sequences for Cas9 domains and fragments are , as will be apparent to those skilled in the art.
[0248] The Cas9 protein guides the protein to a specific DNA sequence complementary to its guide RNA. In one embodiment, the polynucleotide protease is capable of binding to a guide RNA. The activatable nucleotide binding domain can be a Cas9 domain, e.g., a nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Examples of tunable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), CasX, C These include, but are not limited to, asY, Cpf1, Cas12b / C2C1, and Cas12c / C2C3. In one embodiment, the wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes ( NCBI reference sequence: NC_017053.1, nucleotide and amino acid sequences are as follows: ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGATTATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG GCAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGCAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAATCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTAGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAGAAATGGCTTGTTTGGGAAT CTCATTGCTTTGTCATTGGGATTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATAGTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAGCGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAGGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGCGCCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGGGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGATATTCAAAAAGCACAGGTGTCTGGACAAGGCCATAGTTTA CATGAACAGATTGCTAACTTAGCTGGCAGTCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAATTGTTGATGAACT GGTCAAAGTAATGGGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCCAGA AAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCTGTT GAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTACAAAATGGAAGAGACATGTATGTGGACCAAGAATT AGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCATTAAAGACGATTCAATAGACAATA AGGTACTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAAAAC TATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGGTTT GAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAATTT TGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCTAAA TTAGTTTCTGACTTCCGAAAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGCGTA TCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATAAAG TTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCTAAT ATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGGGGA AACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATATTG TCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAAAGAAATTCGGACAAGCTTATTGCT CGTAAAAAGACTGGGATCCAAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGCTAA GGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCTTTG AAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATAAGGAAGTTAAAAGACTTAATCATTAAACTACCTAATAT AGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAATTACAAAAAGGAAATGAGCTGGCTCT GCCAAGCAAATGTGAATTTTTTATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAACAAA AACAATTGTTTGTGGAGCAGCATAAGCATTATTTGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTTATT TTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGAAAA TATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTAAAC GATATACGTCTACAAAGAAGTTTTAGATGCCACTCTTATCCATCCATCCATGGTCTTTGAAACACGCATTGAT TTGAGTCAGCTAGGAGGTGACTGA JPEG0007779988000008.jpg172166 (single underline: HNH domain; double underline: RuvC domain)
[0249] In some embodiments, wild-type Cas9 contains the following nucleotides and / or amino acids: corresponding to or including the amino acid sequence: ATGGATAAAAAGTATTCTATTGGTTTAGACATCGGCACTAATTCCGTTGGATGGGCTGTCATAACCGATGAATACAAAGT ACCTTCAAAGAAATTTAAGGTGTTGGGGAACACAGACCGTCATTCGATTAAAAAAGAATCTTATCGGTGCCCTCCTATTCG ATAGTGGCGAAACGGCAGAGGCGACTCGCCTGAAACGAACCGCTCGGAGAAGGTATACACGTCGCAAGAACCGAATATGT TACTTACAAGAAATTTTTAGCAATGAGATGGCCAAAGTTGACGATTCTTTCTTTCACCGTTTGGAAGAGTCCTTCCTTGT CGAAGAGGACAAGAAACATGAACGGCACCCCATCTTTGGAAACATAGTAGATGAGGTGGCATATCATGAAAAGTACCCAA CGATTTATCACCTCAGAAAAAAGCTAGTTGACTCAACTGATAAAGCGGACCTGAGGTTAATCTACTTGGCTCTTGCCCAT ATGATAAAGTTCCGTGGGCACTTTCTCATTGAGGGTGATCTAAATCCGGACAACTCGGATGTCGACAAACTGTTCATCCA GTTAGTACAAACCTATAATCAGTTGTTTGAAGAGAACCCTATAAAATGCAAGTGGCGTGGATGCGAAGGCTATTCTTAGCG CCCGCCTCTCTAAATCCCGACGGCTAGAAAAACCTGATCGCACAATTACCCGGAGAGAAAAAAATGGTTGTTCGGTAAC CTTATAGCGCTCTCACTAGGCCTGACACCAAATTTAAGTCGAACTTCGACTTAGCTGAAGATGCCAAATTGCAGCTTAG TAAGGACACGTACGATGACGATCTCGCAATCTACTGGCACAAATTGGAGATCAGTATGCGGACTTATTTTTGGCTGCCA AAAACCTTAGCGATGCAATCCTCCTATCTGACATACTGAGAGTTAATACTGAGATTACCAAGGCGCCGTTATCCCGCTTCA ATGATCAAAAGGTACGATGAACATCACCAAGACTTGACACTTCTCAAGGCCCTAGTCCGTCAGCAACTGCCTGAGAAATA TAAGGAAATATTCTTTGATCAGTCGAAAACGGGTACGCAGGTTATATTGACGGGCGGAGCGAGTCAAGAGGAATTCTACA AGTTTATCAAACCCATATTAGAGAATGGATGGGACGGAAGAGTTGCTTGTAAAACTCAATCGCGAAGATCTACTGCGA AAGCAGCGGACTTTCGACAACGGTAGCATTCCACATCAAATCCACTTAGGCGAATTGCATGCTATACTTAGAAGGCAGGA GGATTTTTATCCGTTCCTCAAAGACAATCGTGAAAAGATTGAGAAAATCCTAACCTTTCGCATACCTACTATGTGGGAC CCCTGGCCGAGGGAACTCTCGGTTCGCATGGATGACAAGAAAGTCCGAAGAAACGATTACTCCATGGAATTTTGAGGAA GTTGTCGATAAAGGTGCGTCAGCTCAATCGTTCATCGAGAGGATGACCAACTTTGACAAGAATTTACCGAACGAAAAAGT ATTGCCTAAGCACAGTTTACTTTACGAGTATTTCACAGTGTACAATGAACTCACGAAAGTTAAGTATGTCACTGAGGGCA TGCGTAAACCCGCCTTTCTAAGCGGAGAACAGAAGAAAGCAATAGTAGATCTGTTATTCAAGACCAACCGCAAAGTGACA GTTAAGCAATTGAAAGAGGACTACTTTAAGAAAATTGAATGCTTCGATTCTGTCGAGATCTCCGGGGTAGAAGATCGATT TAATGCGTCACTTGGTACGTATCATGACCTCCTAAAGATAATTAAAGATAAGGACTTCCTGGATAACGAAGAGAATGAAG ATATCTTAGAAGATATAGTGTTGACTCTTACCCTCTTTGAAGATCGGGAAATGATTGAGGAAAGACTAAAAACATACGCT CACCTGTTCGACGATAAGGTTATGAAACAGTTAAAGAGGCGTCGCTATACGGGCTGGGGACGATTGTCGCGGAAACTTAT CAACGGGATAAGAGACAAGCAAAGTGGTAAAACTATTCTCGATTTTCTAAAGAGCGACGGCTTCGCCATAGGAACTTTA TGCAGCTGATCCATGATGACTCTTTAACCTTCAAAGAGGATATACAAAAGGCACAGGTTTCCGGACAAGGGGACTCATTG CACGAACATATTGCGAATCTTGCTGGGTTCGCCAGCCATCAAAGGGCATACTCCAGACAGTCAAAGTAGTGGATGAGCT AGTTAAGGTCATGGGACGTCACAAACCGGAAAACATTGTAATCGAGATGGCACGCGAAAATCAAACGACTCAGAAGGGGC AAAAAAACAGTCGAGAGCGGATGAAGAGAATAGAAGAGGGTATTAAAGAACTGGGCAGCCAGATCTTAAAGGAGCATCCT GTGGAAAATACCCAATTGCAGAACGAGAAACTTTACCTCTATTACCTACAAAATGGAAGGGACATGTATGTTGATCAGGA ACTGGACATAAACCGTTTATCTGATTACGACGTCGATCACATTGTACCCCAATCCTTTTTGAAGGACGATTCAATCGACA ATAAAGTGCTTACACGCTCGGATAAGAACCGAGGGAAAAGTGACAATGTTCCAAGCGAGGAAGTCGTAAAGAAAATGAAG AACTATTGGCGGCAGCTCCTAAATGCGAAACTGATAACGCAAAGAAAGTTCGATAACTTAACTAAAGCTGAGAGGGGTGG CTTGTCTGAACTTGACAAGGCCGGATTTATTAAACGTCAGCTCGTGGAAACCCGCCAAATCACAAAGCATGTTGCACAGA TACTAGATTCCCGAATGAATACGAAATACGACGAGAACGATAAGCTGATTCGGGAAGTCAAAGTAATCACTTTAAAGTCA AAATTGGTGTCGGACTTCAGAAAGGATTTTCAATTCTATAAAGTTAGGGAGATAAATAACTACCACCATGCGCACGACGC TTATCTTAATGCCGTCGTAGGGACCGCACTCATTAAGAAATACCCGAAGCTAGAAAGTGAGTTTGTGTATGGTGATTACA AAGTTTATGACGTCCGTAAGATGATCGCGAAAGCGAACAGGAGATAGGCAAGGCTACAGCCAAATACTTCTTTTATTCT AACATTATGAATTTCTTTAAGACGGAAATCACTCTGGCAAACGGAGAGATACGCAAACGACCTTTAATTGAAACCAATGG GGAGACAGGTGAAATCGTATGGGATAAGGGCCGGGACTTCGCGACGGTGAGAAAAGTTTTGTCCATGCCCCAAGTCAACA TAGTAAAGAAAACTGAGGTGCAGACCGGAGGGTTTTCAAAGGAATCGATTCTTCCAAAAAGGAATAGTGATAAGCTCATC GCTCGTAAAAGGACTGGGACCCGAAAAAGTACGGTGGCTTCGATAGCCCTACAGTTGCCTATTCTGTCCTAGTAGTGGC AAAAGTTGAGAAGGGAAATCCAAGAAACTGAAGTCAGTCAAAGAATTATTGGGGATAACGATTATGGAGCGCTCGTCTT TTGAAAAGACCCCATCGACTTCCTTGAGGCGAAAGGTTACAAGGAAGTAAAAAAGGATCTCATAATTAAACTACCAAAG TATAGTCTGTTTGAGTTAGAAATGGCCGAAAACGGATGTTGGCTAGCGCCGGAGAGCTTCCAAAGGGGAACGAACTCGC ACTACCGTCTAAATACGTGAATTTCCTGTATTTAGCGTCCCATTACGAGAAGTTGAAAGGTTCACCTGAAGATAACGAAC AGAAGCAACTTTTTGTTGAGCAGCACAAACATTATCTCGACGAAATCATAGAGCAAATTTCGGAATTCAGTAAGAGAGTC ATCCTAGCTGATGCCAATCTGGACAAAGTATTAAGCGCATACAACAAGCACAGGGATAAACCCATACGTGAGCAGGCGGA AAATATTATCCATTTGTTTACTCTTACCAACCTCGGCGCTCCAGCCGCATTCAAGTATTTTGACACAACGATAGATCGCA AACGATACACTTCTACCAAGGAGGTGCTAGACGCGACACTGATTCACCAATCCATCACGGGATTATATGAAACTCGGATA GATTTGTCACAGCTTGGGGGTGACGGATCCCCAAGAAGAAGAGGAAAGTCTCGAGCGACTACAAAGACCATGACGGTGA TTATAAAGATCATGACATCGATTACAAGGATGACGATGACAAGGCTGCAGGA JPEG0007779988000009.jpg172165 (single underline: HNH domain; double underline: RuvC domain)
[0250] In some embodiments, the wild-type Cas9 is Cas9 from Streptococcus pyogenes (NCBI Reference ce Sequence: NC_002737.2 (nucleotide sequence is as follows) and Uniprot Reference Sequence: Corresponding to Q99ZW2 (amino acid sequence is as follows): ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGAATATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG ACAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGTAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAACCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTGGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAAAAATGGCTTATTTGGGAAT CTCATTGCTTTGTCATTGGGTTTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATACTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAACGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAAGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGTACCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGAGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGACATTCAAAAAGCACAAGTGTCTGGACAAGGCGATAGTTTA CATGAACATATTGCAAATTTAGCTGGTAGCCCTGCTATTAAAAAGGTATTTTACAGACTGTAAAAGTTGTTGATGAATT GGTCAAAGTAATGGGGCGGCATAAGCCAGAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCC AGAAAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAAGAGCATCCT GTTGAAAATACTCAAATTGCAAATGAAAAGCTCTATCTCTATTATCTCCAAAATGGAAGAGACATGTATGTGGACCAAGA ATTAGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCCTTAAAGACGATTCAATAGACA ATAAGGTCTTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAA AACTATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGG TTTGAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAA TTTTGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCT AATTAGTTTCTGACTTCCGAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGC GTATCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATA AAGTTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCT AATATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGG GGAAACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATA TTGTCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTTACCAAAAGAAATTCGGACAAGCTTATT GCTCGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGC TAAGGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCT TTGAAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAA TATAGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGC TCTGCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAAC AAAAACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTT ATTTTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGA AAATATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTA AACGATATACGTCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATT GATTTGAGTCAGCTAGGAGGTGACTGA JPEG0007779988000010.jpg168163(Underlined once: HNH domain; Underlined twice: RuvC domain)
[0251] In some embodiments, Cas9 is isolated from Corynebacterium ulcerans (NCBI Refs: NC_01 5683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016 786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Strept ococcus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1) ; Psychroflexus torquisI (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NC BI Ref: YP_820832.1), Listeria innocua (NCBI Ref: NP_472073.1), Campylobacter je juni (NCBI Ref: YP_002344900.1) or Neisseria meningitidis (NCBI Ref: YP_00 2342100.1), or represents Cas9 from any other organism.
[0252] Additional Cas9 proteins (e.g., nuclease-dead Cas9 (dCas9), Cas9 nCas9, or nuclease-active Cas9, and its variants and homologs It is understood that within the scope of this disclosure are all such Cas9 proteins, including but not limited to: In some embodiments, the Cas9 protein is a nucleotide sequence that is nucleotides that are nucleotides of interest. is a nuclease-inactive Cas9 (dCas9). In some embodiments, the Cas9 protein is In some embodiments, the Cas9 protein is a nuclear nucleotide sequence encoding a nucleotide sequence of 9 (nCas9). It is an enzyme-active Cas9.
[0253] In some embodiments, the Cas9 domain is a nuclease-inactive Cas9 domain (dCas9). For example, the dCas9 domain can cleave a double-stranded nucleic acid molecule without cleaving either strand. In some embodiments, the nucleic acid molecule may be linked to a gRNA molecule (e.g., via a gRNA molecule). The enzyme-inactive dCas9 domain contains the D10X mutation of the amino acid sequence described herein. and H840X mutation, or in any of the amino acid sequences provided herein and the corresponding mutation, and X is any amino acid change. The nuclease-inactive dCas9 domain contains the D10A mutation of the amino acid sequence described herein. and H840A mutations, or the corresponding mutations in any of the amino acid sequences described herein. As an example, a nuclease-inactive Cas9 domain can be used in a cloning vector. pPlatTET-gRNA 2 (accession number BAV54124) contains the following amino acid sequence: :
[0254] The amino acid sequence of an exemplary catalytically inactive Cas9 (dCas9) is as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD (For example, Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-s specific control of gene expression.” Cell. 2013; 152(5):1173-83 (the entire contents of which are omitted). The bodies are incorporated herein by reference).
[0255] Additional suitable nuclease-inactive dCas9 domains are described in this disclosure and in the art. Such further examples will be apparent to those skilled in the art based on their knowledge and are within the scope of the present disclosure. Suitable nuclease-inactive Cas9 domains include, but are not limited to, D10A / H84 0A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains ( For example, Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotech 2013; 31(9): 833-838, the entire contents of which are incorporated herein by reference. be included).
[0256] In some embodiments, the Cas9 nuclease is an inactive (e.g., inactivated) DNA cleavage domain. Cas9 has a domain called the "nCas9" protein (for "nickase" Cas9). Nuclease-inactivated Cas9 proteins are interchangeably referred to as "dCas9" proteins. It may also be referred to as nuclease-"dead" Cas9 or catalytically inactive Cas9. Methods for generating a Cas9 protein (or a fragment thereof) having an inactive DNA cleavage domain are known. (e.g., Jinek et al., Science. 337:816-821(2012); Qi et al., “Repurposing C RISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression (2013) Cell. 28; 152(5): 1173-83 (the contents of each of which are incorporated herein by reference) For example, the DNA cleavage domain of Cas9 consists of an HNH nuclease subdomain and a RuvC1 subdomain. It is known that the HNH subdomain contains two subdomains: the HNH subdomain and the gRNA subdomain. The RuvC1 subdomain cleaves the complementary strand, and the RuvC2 subdomain cleaves the non-complementary strand. Mutations within the domain can suppress the nuclease activity of Cas9. For example, mutations D10A and H840A completely inactivates the nuclease activity of S. pyogenes Cas9 (Jinek et al., Science e. 337:816-821(2012); Qi et al, Cell. 28;152(5): 1173-83 (2013)).
[0257] In some embodiments, the dCas9 domain is a dCas9 domain provided herein. At least 60%, at least 65%, at least 70%, at least 75%, or at least at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least and amines having 97%, at least 98%, at least 99%, or at least 99.5% identity with each other. In some embodiments, the Cas9 domain comprises a nucleotide sequence described herein. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 5, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 3 5, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more In some embodiments, the Cas9 domain comprises an amino acid sequence having a natural mutation. at least 10, at least 15, At least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, At least 250, at least 300, at least 350, at least 400, at least 500, at least At least 600, at least 700, at least 800, at least 900, at least 1000, at least Contains an amino acid sequence having 1100, or at least 1200 identical contiguous amino acid residues .
[0258] In some embodiments, the dCas9 contains one or more mutations that inactivate Cas9 nuclease activity. For example, In some embodiments, the dCas9 domain contains the D10A and H840A mutations or another Cas9 including the corresponding mutation in
[0259] In some embodiments, the dCas9 comprises the amino acid sequence of dCas9 (D10A and H840A). nothing: JPEG0007779988000011.jpg171165 (single underline: HNH domain; double underline: RuvC domain)
[0260] In some embodiments, the Cas9 domain comprises a D10A mutation, while the the residue at position 840 in the amino acid sequence provided herein, or The residue at the corresponding position in either sequence remains a histidine.
[0261] In other embodiments, D10A, e.g., resulting in nuclease-inactivated Cas9 (dCas9), and dCas9 variants having mutations other than H840A. For example, other amino acid substitutions at D10 and H840, or the nuclease domain of Cas9 may be used. Other substitutions within the domain (e.g., HNH nuclease subdomain and / or RuvC1 subdomain) In some embodiments, a variant or homolog of dCas9 is , at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least In some embodiments, those having at least about 99.9% identity are provided. About 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids or more, or longer. Variants of dCas9 having amino acid sequences are provided.
[0262] In some embodiments, the Cas9 domain is a Cas9 nickase. Cas9 protein, which can cut only one strand of a nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase can target the double-stranded nucleic acid molecule. This allows the Cas9 nickase to separate the gRNA (e.g., sgRNA) bound to the Cas9 from the nucleotide sequence. It means cleaving paired (complementary) strands. The Cas9 nickase contains a D10A mutation and has a histidine at position 840. In embodiments, the Cas9 nickase cleaves the non-target, non-base-edited strand of the double-stranded nucleic acid molecule; This is because the Cas9 nickase base pairs with the gRNA (e.g., sgRNA) bound to Cas9. In some embodiments, the Cas9 nickase cleaves the H840A end strand. containing a natural mutation with an aspartic acid residue at position 10, or the corresponding mutation. In some embodiments, the Cas9 nickase is any of the Cas9 nickases provided herein. or at least 60%, at least 65%, at least 70%, at least 75%, at least 80% , at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, containing an amino acid sequence that is at least 98%, at least 99%, or at least 99.5% identical to the Additional suitable Cas9 nickases may be identified based on this disclosure and knowledge in the art. It is obvious to one skilled in the art and is within the scope of this disclosure.
[0263] The amino acid sequence of an exemplary catalytic Cas9 nickase (nCas9) is as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD
[0264] In some embodiments, Cas9 is expressed in the Archaea (Archaea), which comprise the domain and kingdom of unicellular prokaryotic microorganisms. In some embodiments, the programmable nucleic acid Cas9 is a Cas9 derived from a cell. The nucleotide-binding proteins are described, for example, in Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res. 2017 Feb 21. doi: 10.1038 / cr.2017.21 The CasX or CasY protein may be a CasX or CasY protein, which is incorporated herein by reference in its entirety. This is the first reported finding in the archaeal domain of life using genome-resolved metagenomics. Many CRISPR-Cas systems have been identified, including the previously reported Cas9 protein. Found as part of an active CRISPR-Cas system in little-studied nanoarchaea In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, They are among the most compact systems ever discovered. In some embodiments, in the base editor systems described herein, Cas9 is In some embodiments, the CasX gene is replaced by CasX or a variant of CasX. In the base editor systems described herein, Cas9 is a nucleotide sequence that is linked to CasY or a CasY barrier. Nucleic acid programmable DNA binding protein (napDNAbp) It is understood that other RNA-guided DNA binding proteins may also be used and are within the scope of the present disclosure. It should be.
[0265] In some embodiments, the nucleic acid of any of the fusion proteins provided herein The programmable DNA binding protein (napDNAbp) can be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. is at least 85%, at least 90%, or At least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% In some embodiments, the programmable Suitable nucleotide-binding proteins are the naturally occurring CasX or CasY proteins. In some embodiments, the programmable nucleotide binding protein is At least 85% of any CasX or CasY protein described in the specification, At least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 9 5%, at least 96%, at least 97%, at least 98%, at least 99%, or at least CasX and CasY from other bacterial species also contain amino acid sequences with 99.5% identity. It should be understood that any of the above may be used in accordance with the present disclosure.
[0266] Exemplary CasX ((uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) tr|F0NN 87|F0NN87_SULIHCRISPR-associatedCasx protein OS = Sulfolobus islandicus (strain The amino acid sequence of HVE10 / 4) GN = SiH_0402 PE=4 SV=1) is as follows: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAV GQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNL NSDDGKVRDLKLISAYVNGELIRGEG.
[0267] Exemplary CasX (>tr|F0NH53|F0NH53_SULIR CRISPR associated protein, Casx OS = Sul The amino acid sequence of folobus islandicus (strain REY15A) GN=SiRe_0771 PE=4 SV=1) is as follows: As follows: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAV GQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLN SDDGKVRDLKLISAYVNGELIRGEG.
[0268] Deltaproteobacteria CasX MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMDTDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAIL QVYWQEFKDDHVGLMCKFAQPASKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFWYKLEQVSEKGKAITNYFGRCNVA EHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESHTPVKPLAQIAGNRYASGPVGKALSDACMGTIASFL SKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDfAYNEVIARVRMWVNLNLWQKLKLSRDDA KPLRLLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENP KKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEM DEKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFT DGTDIKKSGKWQGLLYGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKI GRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQ AAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAK LAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKEL SAELDRLSEESGN...
Claims
1. (i) a nucleic acid programmable DNA binding protein (napDNAbp) domain; and (ii) a cytidine deaminase comprising an amino acid sequence having at least 90% sequence identity to any of the following amino acid sequences: a) MTSEKGPSTGDPTLRRRIESWEFDVFYDPRELRKETCLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERRFHSSISCSITWFLSWSPCWECSQAIREFLSQHPGVTLVIYVARLFWHMDQRNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLAFFRLHLQNCHYQTIPPHILLATGLIHPSVTWR, or a variant of the above amino acid sequence comprising an alteration selected from the group consisting of R33A, K34A, R52A, W90F, and H122A; b) MKPQIRDHRPNPMEAMYPHIFYFHFENLEKAYGRNETWLCFTVEIIKQYLPVPWKKGVFRNQVDPETHCHAEKCFLSWFCNNTLSPKKNYQVTWYTSWSPCPECAGEVAEFLAEHSNVKLTIYTARLYYFWDTDYQEGLRSLSEEGASVEIMDYEDFQYCWENFVYDDGEPFKRWKGLKYNFQSLTRRLREILQ, or a variant of the above amino acid sequence containing the F130L modification; c) MADSSEKMRGQYISRDTFEKNYKPIDGTKEAHLLCEIKWGKYGKPWLHWCQNQRMNIHAEDYFMNNIFKAKKHPVHCYVTWYLSWSPCADCASKIVKFLEERPYLKLTIYVAQLYYHTEEENRKGLRLLRSKKVIIRVMDISDYNYCWKVFVSNQNGNEDYWPLQFDPWVKENYSRLLDIFWESKCRSPNPW, and d) MDPQRLRQWPGPGPASRGGYGQRPRIRNPEEWFHELSPRTFSFHFRNLRFASGRNRSYICCQVEGKNCFFQGIFQNQVPPDPPCHAELCFLSWFQSWGLSPDEHYYVTWFISWSPCCECAAKVAQFLEENRNVSLSLSAARLYYFWKSESREGLRRLSDLGAQVGIMSFQDFQHCWNNFVHNLGMPFQPWKKLHKNYQRLVTELKQILREEPATYGSPQAQGKVRIGSTAAGLRHSHSHTRSEAHLRPNHSSRQHRILNPPREARARTCVLVDASWICYR, or a variant of the above amino acid sequence containing the R54Q modification; and a cytidine base editor fusion protein comprising:
2. 2. The cytidine base editor fusion protein of claim 1, wherein the canonical cytidine base editor fusion protein is BE3 or BE4.
3. 3. The cytidine base editor fusion protein of Claim 1 or 2, further comprising at least one adenosine deaminase or a catalytically active fragment thereof.
4. the adenosine deaminase is TadA deaminase; and / or The adenosine deaminase is TadA deaminase, a non-naturally occurring modified adenosine deaminase. The cytidine base editor fusion protein of claim 3.
5. the cytidine base editor fusion protein comprises two adenosine deaminases, which may be the same or different; and / or the cytidine base editor fusion protein comprises two adenosine deaminases capable of forming heterodimers or homodimers; and / or the cytidine base editor fusion protein comprises two adenosine deaminase domains, wild-type TadA and TadA7.
10. The cytidine base editor fusion protein of claim 3.
6. (i) the cytidine base editor fusion protein further comprises one or more nuclear localization signals (NLSs); and / or (ii) the cytidine base editor fusion protein comprises an N-terminal NLS and / or a C-terminal NLS; and / or (iii) the cytidine base editor fusion protein further comprises a bipartite NLS; and / or (iv) the cytidine base editor fusion protein further comprises an N-terminal bipartite NLS and / or a C-terminal bipartite NLS. The cytidine base editor fusion protein of any one of claims 1 to 5.
7. (i) the napDNAbp domain is Cas9; or (ii) the napDNAbp domain is Staphylococcus aureus Cas9 (SaCas9), Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof; or (iii) the napDNAbp domain comprises a nuclease-inactive Cas9 (dCas9), a Cas9 nickase (nCas9), an nCas9 containing the amino acid substitution D10A or a corresponding amino acid substitution, or a nuclease-active Cas9; or (iv) the napDNAbp domain comprises a catalytic domain capable of cleaving the reverse complement of a nucleic acid sequence, or does not comprise a catalytic domain capable of cleaving a nucleic acid sequence; The cytidine base editor fusion protein of any one of claims 1 to 6.
8. (i) further comprising one or more uracil DNA glycosylase inhibitors (UGIs); and / or (ii) further comprising one or more UGIs derived from Bacillus subtilis bacteriophage PBS1 and inhibiting human UDG activity; The cytidine base editor fusion protein of any one of claims 1 to 7.
9. A cell comprising the cytidine base editor fusion protein of any one of claims 1 to 8.
10. 9. A molecular complex comprising the cytidine base editor fusion protein of any one of claims 1 to 8 and one or more of a guide RNA, a tracrRNA, and a target DNA molecule.
11. 10. An in vitro or ex vivo method of editing a nucleobase of a nucleic acid molecule, comprising contacting the nucleic acid molecule with a cytidine base editor fusion protein of any one of claims 1-8 to convert a first nucleobase of the nucleic acid molecule to a second nucleobase, wherein the first nucleobase is cytosine and the second nucleobase is thymidine, and the nucleic acid molecule is a DNA molecule.
12. A polynucleotide molecule encoding the cytidine base editor fusion protein of any one of claims 1 to 8.
13. 13. The polynucleotide molecule of claim 12, which is codon-optimized.
14. An expression vector comprising the polynucleotide molecule of claim 12 or 13.
15. the expression vector is a mammalian expression vector; or The vector is a viral vector selected from the group consisting of an adeno-associated virus (AAV), a retroviral vector, an adenoviral vector, a lentiviral vector, a Sendai virus vector, and a herpes virus vector. The expression vector of claim 14.
16. A cell comprising a polynucleotide molecule according to claim 12 or 13 or a vector according to claim 14 or 15.
17. 16. A kit comprising the cytidine base editor fusion protein of any one of claims 1 to 8, the polynucleotide molecule of claim 12 or 13, the vector of claim 14 or 15, or the molecular complex of claim 10.
Citation Information
Patent Citations
Nucleobase editors comprising nucleic acid programmable DNA binding proteins
WO2018176009A1
Systems, methods, and compositions for targeted nucleic acid editing
WO2018213726A1
Base editors with improved precision and specificity
WO2018218188A2