Nucleobase editor with reduced deamination reaction off-target and method for modifying nucleobase target sequence using same
By designing a fusion protein, combining the polynucleotide programmable DNA binding domain and cytidine deaminase, the cis activity of the nucleobase editor is enhanced, the problem of off-target effect in the prior art is solved, and efficient and specific targeted modification is achieved.
Patent Information
- Application Number
- CN202510365522.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-27
- Filing Date
- 2020-01-31
- Publication Date
- 2025-08-01
AI Technical Summary
Existing nucleobase editors have off-target effects during targeted editing, resulting in nonspecific modifications and lack of high efficiency and specificity.
A fusion protein was designed containing a polynucleotide programmable DNA binding domain and cytidine deaminase, which enhanced the ratio of cis activity to trans activity, reduced off-target deaminescence response, and improved the specificity and efficiency of targeted modifications.
The modification specificity and efficiency of nucleobase editors in the target sequence are significantly improved, and non-specific off-target effects are reduced.
Smart Images

Figure BDA0005329899340000241 
Figure BDA0005329899340000242 
Figure BDA0005329899340000291
Abstract
Description
[0001] This application is a divisional application of a PCT patent application entering China, with the Chinese patent application number 2020800263104, the invention title "Nucleobase editors with reduced off-target deamination reactions and methods of using the same to modify nucleobase target sequences", and the filing date of January 31, 2020.
[0002] Cross-reference to related applications
[0003] This application is an international PCT application. This application claims the benefit of U.S. Provisional Patent Application No. 62 / 799,702 filed on January 31, 2019; U.S. Provisional Patent Application No. 62 / 835,456 filed on April 17, 2019; and U.S. Provisional Patent Application No. 62 / 941,569 filed on November 27, 2019, the entire contents of each of which are hereby incorporated by reference herein. Technical field
[0004] This application relates to nucleobase editors and multi-effector nucleobase editors with improved editing settings and minimal off-target deamination reactions, compositions comprising such editors, and methods of using the editors and / or compositions to generate modifications in target nucleobase sequences. Background art
[0005] Targeted editing of nucleic acid sequences, such as targeted cleavage or targeted modification of genomic DNA, is a promising method for gene function research and also has the potential to provide new therapies for human genetic diseases. Currently available base editors include cytidine base editors (such as BE4) that convert target C·G base pairs to T·A and adenine base editors (such as ABE7.10) that convert A·T to G·C. There is a need in the art for improved base editors that can induce modifications within target sequences with higher specificity and efficiency. Summary of the invention
[0006] As described below, the features of the present invention are: nucleobase editors and multi-effector nucleobase editors with improved editing settings (i.e., minimal off-target deamination reactions), compositions comprising such editors, and methods of using them to generate modifications in target nucleobase sequences.
[0007] In one aspect, the present disclosure provides a cytidine base editor comprising (i) a polynucleotide programmable DNA binding domain and (ii) a cytidine deaminase, wherein the cytidine base editor has an increased ratio of cis activity to trans activity (cis:trans) compared to a standard cytidine base editor.
[0008] In some embodiments, the standard cytidine base editor comprises (i) a polynucleotide programmable DNA binding domain and (ii) an APOBEC cytidine deaminase. In some embodiments, the APOBEC cytidine deaminase of the standard cytidine base editor is the rat APOBEC-1 cytidine deaminase (rAPOBEC-1). In some embodiments, the polynucleotide programmable DNA binding domain of the standard cytidine base editor is a Cas9 nickase. In some embodiments, the standard cytidine base editor comprises a uracil glycosylase inhibitor (UGI) domain. In some embodiments, the standard cytidine base editor is BE3 or BE4. In some embodiments, the increased cis-activity to trans-activity ratio is increased by at least 2, 2.5, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60-fold or more. In some embodiments, the cytidine base editor has at least 50%, 60%, 70%, 80%, 90%, 95%, 100%, 105%, 110%, 115%, 120%, or higher cis-activity compared to the standard cytidine base editor.
[0009] In some embodiments, the cytidine base editor has at least 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60-fold or more lower trans-activity compared to the standard cytidine base editor.
[0010] In some embodiments, the cytidine deaminase is selected from the group consisting of: APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3E, APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, activation-induced (cytidine) deaminase (AID), hAPOBEC1, rAPOBEC1, ppAPOBEC1, AmAPOBEC1 (BEM3.31), ocAPOBEC1, SsAPOBEC2 (BEM3.39), hAPOBEC3A, maAPOBEC1, mdAPOBEC1, cytidine deaminase 1 (CDA1), hA3A, RrA3F (BEM3.14), PmCDA1, AID (activation-induced cytidine deaminase; AICDA), hAID, and FENRY. In some embodiments, the cytidine deaminase is APOBEC1. In some embodiments, the cytidine deaminase is (a) APOBEC-1 from golden hamster (MaAPOBEC-1), Bornean orangutan (PpAPOBEC-1), European rabbit (OcAPOBEC-1), gray short-tailed opossum (MdAPOBEC-1), or American alligator (AmAPOBEC-1), (b) APOBEC-2 from Bornean orangutan (PpAPOBEC-2), domestic cattle (BtAPOBEC-2), or European pig (SsAPOBEC-2), (c) APOBEC-4 from cynomolgus macaque (MfAPOBEC-4), (d) AID from dog (ClAID) or domestic cattle (BtAID), (e) yeast cytosine deaminase (yCD) from Saccharomyces cerevisiae, (f) APOBEC-3F from Rhinopithecus roxellana (RrA3F), or (g) a cytidine deaminase having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity to any one of the proteins in (a)-(f).
[0011] In some embodiments, the cytidine deaminase is APOBEC-1, which is from golden hamster (MaAPOBEC-1), Bornean orangutan (PpAPOBEC-1), European rabbit (OcAPOBEC-1), gray short-tailed opossum (MdAPOBEC-1), or a cytidine deaminase having an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the foregoing proteins. In some embodiments, the cytidine deaminase is rAPOBEC1. In some embodiments, the cytidine deaminase is hAPOBEC3A. In some embodiments, the cytidine deaminase is ppAPOBEC1. In some embodiments, the cytidine deaminase is APOBEC-2, which is derived from Bornean orangutan (PpAPOBEC-2), domestic cattle (BtAPOBEC-2), or European pig (SsAPOBEC-2), or a cytidine deaminase having an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the foregoing proteins. In some embodiments, the cytidine deaminase is APOBEC-4, which is derived from cynomolgus macaque (MfAPOBEC-4), or a cytidine deaminase having an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the foregoing proteins. In some embodiments, the cytidine deaminase is AID, which is from dog (ClAID), domestic cattle (BtAID), or a cytidine deaminase having an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the foregoing proteins.
[0012] In some embodiments, the cytidine deaminase is yeast cytosine deaminase (yCD), which is from Saccharomyces cerevisiae, or a cytidine deaminase having an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the foregoing protein. In some embodiments, the cytidine deaminase is APOBEC-3F, which is from Rhinopithecus roxellana (RrA3F), or a cytidine deaminase having an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the foregoing protein. In some embodiments, the cytidine deaminase is any of the cytidine deaminases provided in Table 13, or a cytidine deaminase having an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the foregoing protein. In some embodiments, the cytidine deaminase is APOBEC-3F from Rhinopithecus roxellana (RrA3F), APOBEC-1 from Alligator mississippiensis (AmAPOBEC-1), APOBEC-2 from Sus scrofa (SsAPOBEC-2), APOBEC-1 from Pongo pygmaeus (PpAPOBEC-1), or a cytidine deaminase having an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the foregoing protein.
[0013] In some embodiments, the cytidine deaminase comprises one or more alterations located at position R15X, R16X, H21X, R30X, R33X, K34X, R52X, K60X, R118X, H121X, H122X, R126X, R128X, R169X, R198X, T36X, H53X, V62X, L88X, W90X, Y120X or R132X (as numbered in SEQ ID NO:1) or corresponding alterations of one or more of the foregoing alterations, where X is any amino acid.
[0014] In some embodiments, the cytidine deaminase comprises one or more alterations selected from the group consisting of: R15A, R16A, H21A, R30A, R33A, K34A, R52A, K60A, R118A, H121A, H122A, H122L, R126A, R128A, R169A, R198A, T36A, H53A, V62A, L88A, W90F, W90A, Y120F, Y120A, H121R, H122R, R126E, W90Y, and R132E (numbered as in SEQ ID NO:1) or corresponding alterations of one or more of said alterations. In some embodiments, the cytidine deaminase comprises a combination of alterations selected from the group consisting of: K34A+R33A, K34A+H122A, K34A+Y120F, K34A+R52A, K34A+H122A, K34A+H121A, W90A+R126E, W90Y+R126E, H121R+H122R, R126+R132E, W90Y+R132E, and W90Y+R126E+R132E (numbered as in SEQ ID NO:1) or corresponding alterations of said alterations. In some embodiments, the cytidine deaminase comprises the alteration at position Y120F and one or more alterations selected from the group consisting of: R33A, W90F, K34A, R52A, H122A, and H121A (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations. In some embodiments, the cytidine deaminase comprises an alteration at position Y130X or R28X (numbered as in SEQ ID NO:1) or corresponding alterations of said alterations, wherein X is any amino acid.
[0015] In some embodiments, the cytidine deaminase comprises a change at position Y130A or R28A (numbered as in SEQ ID NO:1) or a corresponding change of said change. In some embodiments, the cytidine deaminase comprises changes at positions Y130A and R28A (numbered as in SEQ ID NO:1) or a corresponding change of said change. In some embodiments, the cytidine deaminase comprises one or more changes at positions H122X, K34X, R33X, W90X, or R128X (numbered as in SEQ ID NO:1), or a corresponding change of one or more of said changes, wherein X is any amino acid. In some embodiments, the cytidine deaminase comprises one or more changes selected from the group consisting of: H122A, K34A, R33A, W90F, W90A, and R128A (numbered as in SEQ ID NO:1), or a corresponding change of one or more of said changes. In some embodiments, the cytidine deaminase comprises a combination of changes selected from the group consisting of: R33A+K34A, W90F+K34A, R33A+K34A+W90F, and R33A+K34A+H122A+W90F (numbered as in SEQ ID NO:1) or a corresponding change of said change.
[0016] In some embodiments, the cytidine deaminase comprises an amino acid sequence having at least 80% identity to the following amino acid sequence:
[0017] MTSEKGPSTGDPTLRRRIESWEFDVFYDPRELRKETCLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERRFHSSISCSITWFLSWSPCWECSQAIREFLSQHPGVTLVIYVARLFWHMDQRNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLAFFRLHLQNCHYQTIPPHILLATGLIHPSVTWR.
[0018] In some embodiments, the cytidine deaminase comprises an amino acid sequence having at least 80% identity to the following amino acid sequence:
[0019] MKPQIRDHRPNPMEAMYPHIFYFHFENLEKAYGRNETWLCFTVEIIKQYLPVPWKKGVFRNQVDPETHCHAEKCFLSWFCNNTLSPKKNYQVTWYTSWSPCPECAGEVAEFLAEHSNVKLTIYTARLYYFWDTDYQEGLRSLSEEGASVEIMDYEDFQYCWENFVYDDGEPFKRWKGLKYNFQSLTRRLREILQ.
[0020] In some embodiments, the cytidine deaminase comprises an amino acid sequence that has at least 80% identity to the following amino acid sequence:
[0021] MADSSEKMRGQYISRDTFEKNYKPIDGTKEAHLLCEIKWGKYGKPWLHWCQNQRMNIHAEDYFMNNIFKAKKHPVHCYVTWYLSWSPCADCASKIVKFLEERPYLKLTIYVAQLYYHTEEENRKGLRLLRSKKVIIRVMDISDYNYCWKVFVSNQNGNEDYWPLQFDPWVKENYSRLLDIFWESKCRSPNPW.
[0022] In some embodiments, the cytidine deaminase comprises an amino acid sequence that has at least 80% identity to the following amino acid sequence:
[0023] MDPQRLRQWPGPGPASRGGYGQRPRIRNPEEWFHELSPRTFSFHFRNLRFASGRNRSYICCQVEGKNCFFQGIFQNQVPPDPPCHAELCFLSWFQSWGLSPDEHYYVTWFISWSPCCECAAKVAQFLEENRNVSLSLSAARLYYFWKSESREGLRRLSDLGAQVGIMSFQDFQHCWNNFVHNLGMPFQPWKKLHKNYQRLVTELKQILREEPATYGSPQAQGKVRIGSTAAGLRHSHSHTRSEAHLRPNHSSRQHRILNPPREARARTCVLVDASWICYR.
[0024] In some embodiments, the cytidine deaminase comprises an H122A alteration. In some embodiments, the cytidine base editor of any of the foregoing aspects further comprises at least one adenosine deaminase or a catalytically active fragment thereof. In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the TadA deaminase is a modified adenosine deaminase that does not exist in nature. In some embodiments, the cytidine base editor comprises two adenosine deaminases, which are the same or different. In some embodiments, the two adenosine deaminases are capable of forming a heterodimer or a homodimer. In some embodiments, the adenosine deaminase domains are wild-type TadA and TadA7.10.
[0025] In some embodiments, the adenosine deaminase comprises a C-terminal deletion starting from a residue selected from the group consisting of: 149, 150, 151, 152, 153, 154, 155, 156, and 157. In some embodiments, the adenosine deaminase lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to the full-length adenosine deaminase. In some embodiments, the adenosine deaminase lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to the full-length adenosine deaminase. In some embodiments, at least one nucleobase editor domain further comprises a abasic nucleobase editor. In some embodiments, the cytidine base editor of any of the foregoing aspects further comprises one or more nuclear localization signals (NLS). In some embodiments, the cytidine base editor comprises an N-terminal NLS and / or a C-terminal NLS. In some embodiments, the NLS is a bipartite NLS.
[0026] In some embodiments, the polynucleotide programmable DNA binding domain is Cas9. In some embodiments, the polynucleotide programmable DNA binding domain is Staphylococcus aureus Cas9 (SaCas9), Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof. In some embodiments, the polynucleotide programmable DNA binding domain comprises nuclease-inactivated Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9. In some embodiments, the polynucleotide programmable DNA binding domain comprises a catalytic domain capable of cleaving the reverse complementary strand of a nucleic acid sequence. In some embodiments, the polynucleotide programmable DNA binding domain does not comprise a catalytic domain capable of cleaving a nucleic acid sequence. In some embodiments, the Cas9 is dCas9. In some embodiments, the Cas9 is Cas9 nickase (nCas9). In some embodiments, the nCas9 comprises the amino acid substitution D10A or its corresponding amino acid substitution.
[0027] In some embodiments, the cytidine base editor of any of the above aspects further comprises one or more uracil DNA glycosylase inhibitors (UGI). In some embodiments, the one or more UGI are derived from Bacillus subtilis bacteriophage PBS1 and inhibit human UDG activity. In some embodiments, the cytidine base editor comprises two uracil DNA glycosylase inhibitors (UGI). In some embodiments, the cytidine base editor of any of the above aspects further comprises one or more linkers.
[0028] Cells comprising the cytidine base editor of any of the above aspects are provided herein. In some embodiments, the cell is a bacterial cell, a plant cell, an insect cell, or a mammalian cell.
[0029] A molecular complex is provided herein that comprises the cytidine base editor of any of the above aspects and one or more guide RNA sequences, tracrRNA sequences, or target DNA sequences.
[0030] A method of editing a nucleobase of a nucleic acid sequence is provided herein, the method comprising contacting the nucleic acid sequence with the cytidine base editor of any of the above aspects and converting a first nucleobase of the DNA sequence to a second nucleobase.
[0031] In some embodiments, the method further comprises contacting the nucleic acid sequence with a guide polynucleotide to effect the above conversion. In some embodiments, the first nucleobase is cytosine and the second nucleobase is thymidine.
[0032] In one aspect, the present invention provides a fusion protein comprising a polynucleotide programmable DNA binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase is (i) APOBEC-1 from golden hamster (MaAPOBEC-1), Bornean orangutan (PpAPOBEC-1), burrowing rabbit (OcAPOBEC-1), gray short-tailed opossum (MdAPOBEC-1), or American alligator (AmAPOBEC-1), (ii) APOBEC-2 from Bornean orangutan (PpAPOBEC-2), domestic cattle (BtAPOBEC-2), or (iii) APOBEC-3 from swinhoe's brood (S ... BEC-2), or European pig (SsAPOBEC-2), (iii) APOBEC-4 from cynomolgus macaque (MfAPOBEC-4), (iv) AID from dog (ClAID) or cattle (BtAID), (v) yeast cytosine deaminase (yCD) from Saccharomyces cerevisiae, (vi) APOBEC-3F from golden snub-nosed monkey (RrA3F), or (vii) a cytidine deaminase having an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to any one of (i) to (viii).
[0033] In one aspect, the present invention provides a fusion protein comprising a polynucleotide-programmable DNA binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase is APOBEC-1 from golden hamster (MaAPOBEC-1), Bornean orangutan (PpAPOBEC-1), cave rabbit (OcAPOBEC-1), gray short-tailed opossum (MdAPOBEC-1), or a cytidine deaminase having an amino acid sequence at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the foregoing protein.
[0034] In one aspect, the present invention provides a fusion protein comprising a polynucleotide-programmable DNA binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase is APOBEC-2 from Bornean orangutan (PpAPOBEC-2), domestic cattle (BtAPOBEC-2), or European pigs (SsAPOBEC-2), or a cytidine deaminase having an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the foregoing proteins.
[0035] In one aspect, the present disclosure provides a fusion protein comprising a polynucleotide-programmable DNA-binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase is APOBEC-4, which is from cynomolgus macaque (MfAPOBEC-4), or a cytidine deaminase having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity to the foregoing protein.
[0036] In one aspect, the present disclosure provides a fusion protein comprising a polynucleotide-programmable DNA-binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase is AID, which is from dog (ClAID), bovine (BtAID), or a cytidine deaminase having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity to the foregoing protein.
[0037] In one aspect, the present disclosure provides a fusion protein comprising a polynucleotide-programmable DNA-binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase is yeast cytosine deaminase (yCD), which is from Saccharomyces cerevisiae, or a cytidine deaminase having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity to the foregoing protein.
[0038] In one aspect, the present disclosure provides a fusion protein comprising a polynucleotide-programmable DNA-binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase is APOBEC-3F, which is from Rhinopithecus roxellana (RrA3F), or a cytidine deaminase having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity to the foregoing protein.
[0039] In one aspect, the present disclosure provides a fusion protein comprising a polynucleotide-programmable DNA-binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase is any of the cytidine deaminases provided in Table 13, or a cytidine deaminase having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity to the foregoing protein.
[0040] On the one hand, the present invention provides a fusion protein comprising a polynucleotide-programmable DNA binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase is APOBEC-3F (RrA3F) from Sichuan golden snub-nosed monkey, APOBEC-1 (AmAPOBEC-1) from American alligator, APOBEC-2 from European pig (SsAPOBEC-2), APOBEC-1 from Bornean orangutan (PpAPOBEC-1), or a cytidine deaminase having an amino acid sequence at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the foregoing protein.
[0041] In some embodiments, the cytidine deaminase comprises one or more alterations located at positions R15X, R16X, H21X, R30X, R33X, K34X, R52X, K60X, R118X, H121X, H122X, R126X, R128X, R169X, R198X, T36X, H53X, V62X, L88X, W90X, Y120X, or R132X (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations, wherein X is any amino acid. In some embodiments, the cytidine deaminase comprises one or more alterations selected from the group consisting of: R15A, R16A, H21A, R30A, R33A, K34A, R52A, K60A, R118A, H121A, H122A, H122L, R126A, R128A, R169A, R198A, T36A, H53A, V62A, L88A, W90F, W90A, Y120F, Y120A, H121R, H122R, R126E, W90Y, and R132E (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations. In some embodiments, the cytidine deaminase comprises a combination of alterations selected from the group consisting of: K34A+R33A, K34A+H122A, K34A+Y120F, K34A+R52A, K34A+H122A, K34A+H121A, W90A+R126E, W90Y+R126E, H121R+H122R, R126+R132E, W90Y+R132E, and W90Y+R126E+R132E (numbered as in SEQ ID NO:1) or corresponding alterations of said alterations. In some embodiments, the cytidine deaminase comprises a combination of alterations, the alterations of the combination being selected from the group consisting of Y120F and one or more alterations selected from the group consisting of: R33A, W90F, K34A, R52A, H122A, and H121A (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations.
[0042] In some embodiments, the cytidine deaminase comprises one or more alterations located at position Y130X or R28X (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations, wherein X is any amino acid. In some embodiments, the cytidine deaminase comprises one or more alterations selected from the group consisting of: Y130A and R28A (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations. In some embodiments, the cytidine deaminase comprises the alterations Y130A and R28A (numbered as in SEQ ID NO:1) or corresponding alterations of said alterations. In some embodiments, the cytidine deaminase comprises one or more alterations located at position H122X, K34X, R33X, W90X, or R128X (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations, wherein X is any amino acid. In some embodiments, the cytidine deaminase comprises one or more alterations selected from the group consisting of: H122A, K34A, R33A, W90F, W90A, and R128A (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations.
[0043] In some embodiments, the cytidine deaminase comprises an altered combination of alterations selected from the group consisting of: R33A+K34A, W90F+K34A, R33A+K34A+W90F, and R33A+K34A+H122A+W90F (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations. In some embodiments, the cytidine deaminase comprises the H122A alteration (numbered as in SEQ ID NO:1), or a corresponding alteration of said alteration. In some embodiments, the cytidine deaminase is rAPOBEC1 and comprises one or more alterations selected from the group consisting of: R15A, R16A, H21A, R30A, R33A, K34A, R52A, K60A, R118A, H121A, H122A, H122L, R126A, R128A, R169A, R198A, T36A, H53A, V62A, L88A, W90F, W90A, Y120F, Y120A, H121R, H122R, R126E, W90Y, and R132E (numbered as in SEQ ID NO:1) or corresponding alterations of one or more of said alterations. In some embodiments, the cytidine deaminase comprises an altered combination of alterations selected from the group consisting of: K34A+R33A, K34A+H122A, K34A+Y120F, K34A+R52A, K34A+H122A, K34A+H121A, W90A+R126E, W90Y+R126E, H121R+H122R, R126+R132E, W90Y+R132E, and W90Y+R126E+R132E (numbered as in SEQ ID NO:1) or corresponding alterations of one or more of said alterations.
[0044] In one aspect, the present disclosure provides a fusion protein comprising a polynucleotide-programmable DNA binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase is selected from the group consisting of: APOBEC2 family members, APOBEC3 family members, APOBEC4 family members, cytidine deaminase 1 family members (CDA1), A3A family members, RrA3F family members, PmCDA1 family members, and FENRY family members.
[0045] In some embodiments, the APOBEC3 family member is selected from the group consisting of: APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3E, APOBEC3F, APOBEC3G, and APOBEC3H. In some embodiments, the APOBEC2 family member is SsAPOBEC2.
[0046] Provided herein is a fusion protein comprising a polynucleotide programmable DNA binding domain and at least one cytidine deaminase domain comprising APOBEC1, wherein the APOBEC1 is selected from the group consisting of: ppAPOBEC1, AmAPOBEC1 (BEM3.31), ocAPOBEC1, SsAPOBEC2 (BEM3.39), hAPOBEC3A, maAPOBEC1, and mdAPOBEC1.
[0047] In some embodiments, the cytidine deaminase comprises one or more alterations located at positions R15X, R16X, H21X, R30X, R33X, K34X, R52X, K60X, R118X, H121X, H122X, R126X, R128X, R169X, R198X, T36X, H53X, V62X, L88X, W90X, Y120X or R132X (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations, wherein X is any amino acid. In some embodiments, the one or more alterations are selected from the group consisting of: R15A, R16A, H21A, R30A, R33A, K34A, R52A, K60A, R118A, H121A, H122A, H122L, R126A, R128A, R169A, R198A, T36A, H53A, V62A, L88A, W90F, W90A, Y120F, Y120A, H121R, H122R, R126E, W90Y, and R132E (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations. In some embodiments, the cytidine deaminase comprises a combination of alterations selected from the group consisting of: K34A+R33A, K34A+H122A, K34A+Y120F, K34A+R52A, K34A+H122A, K34A+H121A, W90A+R126E, W90Y+R126E, H121R+H122R, R126+R132E, W90Y+R132E, and W90Y+R126E+R132E (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations. In some embodiments, the cytidine deaminase comprises a combination of alterations, the alterations of the combination being selected from the group consisting of Y120F and one or more alterations selected from the group consisting of: R33A, W90F, K34A, R52A, H122A, and H121A, (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations.
[0048] On the one hand, the present disclosure provides a fusion protein, which comprises a polynucleotide-programmable DNA-binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase comprises one or more alterations located at positions R15X, R16X, H21X, R30X, R33X, K34X, R52X, K60X, R118X, H121X, H122X, R126X, R128X, R169X, R198X, T36X, H53X, V62X, L88X, W90X, Y120X or R132X (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations, wherein X is any amino acid.
[0049] In some embodiments, the cytidine deaminase comprises one or more alterations selected from the group consisting of: R15A, R16A, H21A, R30A, R33A, K34A, R52A, K60A, R118A, H121A, H122A, H122L, R126A, R128A, R169A, R198A, T36A, H53A, V62A, L88A, W90F, W90A, Y120F, Y120A, H121R, H122R, R126E, W90Y, and R132E (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations. In some embodiments, the cytidine deaminase comprises a combination of alterations selected from the group consisting of: K34A+R33A, K34A+H122A, K34A+Y120F, K34A+R52A, K34A+H122A, K34A+H121A, W90A+R126E, W90Y+R126E, H121R+H122R, R126+R132E, W90Y+R132E, and W90Y+R126E+R132E (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations. In some embodiments, the cytidine deaminase comprises the alteration at position Y120F and one or more alterations selected from the group consisting of: R33A, W90F, K34A, R52A, H122A, and H121A (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations.
[0050] In one aspect, provided herein is a fusion protein comprising a polynucleotide-programmable DNA binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase comprises one or more alterations at positions Y130X and R28X (as numbered in SEQ ID NO: 1), or corresponding alterations thereof, wherein X is any amino acid.
[0051] In some embodiments, the cytidine deaminase comprises one or more alterations selected from the group consisting of Y130A and R28A (as numbered in SEQ ID NO: 1), or corresponding alterations of one or more of said alterations. In some embodiments, the cytidine deaminase comprises alterations Y130A and R28A.
[0052] In one aspect, provided herein is a fusion protein comprising a polynucleotide-programmable DNA binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase comprises one or more alterations at position H122X, K34X, R33X, W90X, or R128X (as numbered in SEQ ID NO: 1), or corresponding alterations of one or more of said alterations, wherein X is any amino acid.
[0053] In some embodiments, the cytidine deaminase comprises one or more alterations selected from the group consisting of H122A, K34A, R33A, W90F, W90A, and R128A (as numbered in SEQ ID NO: 1), or corresponding alterations to one or more of said alterations. In some embodiments, the cytidine deaminase comprises a combination of alterations selected from the group consisting of R33A+K34A, W90F+K34A, R33A+K34A+W90F, and R33A+K34A+H122A+W90F (as numbered in SEQ ID NO: 1), or corresponding alterations to one or more of said alterations. In some embodiments, the cytidine deaminase is selected from the group consisting of:
[0054] APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3E, APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, activation-induced (cytidine) deaminase (AID), hAPOBEC1, rAPOBEC1, ppAPOBEC1, AmAPOBEC1 (BEM3.31), ocAPOBEC1, SsAPOBEC2 (BEM3.39), hAPOBEC3A, maAPOBEC1, mdAPOBEC1, cytidine deaminase 1 (CDA1), hA3A, RrA3F (BEM3.14), PmCDA1, AID (activation-induced cytidine deaminase; AICDA), hAID, and FENRY. In some embodiments, the cytidine deaminase is APOBEC1. In some embodiments, the cytidine deaminase is rAPOBEC1. In some embodiments, the cytidine deaminase is hAPOBEC3A. In some embodiments, the cytidine deaminase is ppAPOBEC1.
[0055] In one aspect, the present disclosure provides a fusion protein comprising a polynucleotide programmable DNA binding domain and a cytidine deaminase, wherein the cytidine deaminase comprises an amino acid sequence having at least 80% identity to the following amino acid sequence:
[0056] MTSEKGPSTGDPTLRRRIESWEFDVFYDPRELRKETCLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERRFHSSISCSITWFLSWSPCWECSQAIREFLSQHPGVTLVIYVARLFWHMDQRNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLAFFRLHLQNCHYQTIPPHILLATGLIHPSVTWR.
[0057] In one aspect, the present disclosure provides a fusion protein comprising a polynucleotide programmable DNA binding domain and a cytidine deaminase, wherein the cytidine deaminase comprises an amino acid sequence having at least 80% identity to the following amino acid sequence:
[0058] MKPQIRDHRPNPMEAMYPHIFYFHFENLEKAYGRNETWLCFTVEIIKQYLPVPWKKGVFRNQVDPETHCHAEKCFLSWFCNNTLSPKKNYQVTWYTSWSPCPECAGEVAEFLAEHSNVKLTIYTARLYYFWDTDYQEGLRSLSEEGASVEIMDYEDFQYCWENFVYDDGEPFKRWKGLKYNFQSLTRRLREILQ.
[0059] On the one hand, the present disclosure provides a fusion protein comprising a polynucleotide programmable DNA binding domain and a cytidine deaminase, wherein the cytidine deaminase comprises an amino acid sequence that has at least 80% identity to the following amino acid sequence:
[0060] MADSSEKMRGQYISRDTFEKNYKPIDGTKEAHLLCEIKWGKYGKPWLHWCQNQRMNIHAEDYFMNNIFKAKKHPVHCYVTWYLSWSPCADCASKIVKFLEERPYLKLTIYVAQLYYHTEEENRKGLRLLRSKKVIIRVMDISDYNYCWKVFVSNQNGNEDYWPLQFDPWVKENYSRLLDIFWESKCRSPNPW.
[0061] On the one hand, the present disclosure provides a fusion protein comprising a polynucleotide programmable DNA binding domain and a cytidine deaminase, wherein the cytidine deaminase comprises an amino acid sequence that has at least 80% identity to the following amino acid sequence:
[0062] MDPQRLRQWPGPGPASRGGYGQRPRIRNPEEWFHELSPRTFSFHFRNLRFASGRNRSYICCQVEGKNCFFQGIFQNQVPPDPPCHAELCFLSWFQSWGLSPDEHYYVTWFISWSPCCECAAKVAQFLEENRNVSLSLSAARLYYFWKSESREGLRRLSDLGAQVGIMSFQDFQHCWNNFVHNLGMPFQPWKKLHKNYQRLVTELKQILREEPATYGSPQAQGKVRIGSTAAGLRHSHSHTRSEAHLRPNHSSRQHRILNPPREARARTCVLVDASWICYR.
[0063] In some embodiments, the cytidine deaminase comprises an H122A alteration.
[0064] In one aspect, the present disclosure provides a fusion protein comprising a polynucleotide programmable DNA binding domain and a cytidine deaminase, wherein the cytidine deaminase is APOBEC1 deaminase and comprises an H122A alteration.
[0065] In one aspect, the present disclosure provides a fusion protein comprising a polynucleotide programmable DNA binding domain and a cytidine deaminase, wherein the cytidine deaminase is rAPOBEC1 and comprises one or more alterations selected from the group consisting of: R15A, R16A, H21A, R30A, R33A, K34A, R52A, K60A, R118A, H121A, H122A, H122L, R126A, R128A, R169A, R198A, T36A, H53A, V62A, L88A, W90F, W90A, Y120F, Y120A, H121R, H122R, R126E, W90Y, and R132E. In some embodiments, the cytidine deaminase comprises a combination of alterations selected from the group consisting of: K34A+R33A, K34A+H122A, K34A+Y120F, K34A+R52A, K34A+H122A, K34A+H121A, W90A+R126E, W90Y+R126E, H121R+H122R, R126+R132E, W90Y+R132E, and W90Y+R126E+R132E.
[0066] In one aspect, the present disclosure provides a fusion protein comprising a polynucleotide programmable DNA binding domain and at least one APOBEC1-containing nucleobase editor domain, wherein the APOBEC1 is selected from the group consisting of: ppAPOBEC1, AmAPOBEC1 (BEM3.31), ocAPOBEC1, SsAPOBEC2 (BEM3.39), hAPOBEC3A, maAPOBEC1, and mdAPOBEC1.
[0067] In some embodiments, the APOBEC1 comprises one or more alterations at positions R15X, R16X, H21X, R30X, R33X, K34X, R52X, K60X, R118X, H121X, H122X, R126X, R128X, R169X, R198X, T36X, H53X, V62X, L88X, W90X, Y120X or R132X (numbered as in SEQ ID NO:1) or corresponding alterations of one or more of said alterations, wherein X is any amino acid.
[0068] In some embodiments, the one or more alterations are selected from the group consisting of: R15A, R16A, H21A, R30A, R33A, K34A, R52A, K60A, R118A, H121A, H122A, H122L, R126A, R128A, R169A, R198A, T36A, H53A, V62A, L88A, W90F, W90A, Y120F, Y120A, H121R, H122R, R126E, W90Y, and R132E (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations. In some embodiments, the APOBEC1 comprises a combination of alterations selected from the group consisting of: K34A+R33A, K34A+H122A, K34A+Y120F, K34A+R52A, K34A+H122A, K34A+H121A, W90A+R126E, W90Y+R126E, H121R+H122R, R126+R132E, W90Y+R132E, and W90Y+R126E+R132E (numbered as in SEQ ID NO:1) or corresponding alterations of one or more of said alterations. In some embodiments, the APOBEC1 comprises an alteration at position Y120F and one or more alterations selected from the group consisting of: R33A, W90F, K34A, R52A, H122A, and H121A (numbered as in SEQ ID NO:1), or corresponding alterations of one or more of said alterations.
[0069] In some embodiments, the fusion protein of any of the above aspects further comprises at least one adenosine deaminase or a catalytically active fragment thereof. In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is a modified adenosine deaminase that does not exist in nature. In some embodiments, the fusion protein comprises two adenosine deaminases, which are the same or different. In some embodiments, the two adenosine deaminases are capable of forming a heterodimer or a homodimer. In some embodiments, the two adenosine deaminase domains are wild-type TadA and TadA7.10.
[0070] In some embodiments, the adenosine deaminase comprises a C-terminal deletion starting from a residue selected from the group consisting of: 149, 150, 151, 152, 153, 154, 155, 156, and 157. In some embodiments, the adenosine deaminase lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 N-terminal amino acid residues relative to the full-length adenosine deaminase. In some embodiments, the adenosine deaminase lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 C-terminal amino acid residues relative to the full-length adenosine deaminase. In some embodiments, at least one nucleobase editor domain further comprises a base-free nucleobase editor.
[0071] In some embodiments, the fusion protein of any of the above aspects further comprises one or more nuclear localization signals (NLSs). In some embodiments, the fusion protein comprises an N-terminal NLS and / or a C-terminal NLS. In some embodiments, the NLS is a bipartite NLS.
[0072] In some embodiments, the polynucleotide programmable DNA binding domain is Cas9. In some embodiments, the polynucleotide programmable DNA binding domain is Staphylococcus aureus Cas9 (SaCas9), Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof. In some embodiments, the polynucleotide programmable DNA binding domain comprises nuclease-inactivated Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9.
[0073] In some embodiments, the polynucleotide programmable DNA binding domain comprises a catalytic domain capable of cleaving the reverse complementary strand of a nucleic acid sequence. In some embodiments, the polynucleotide programmable DNA binding domain does not comprise a catalytic domain capable of cleaving a nucleic acid sequence.
[0074] In some embodiments, the Cas9 is dCas9. In some embodiments, the Cas9 is Cas9 nickase (nCas9). In some embodiments, the nCas9 includes the amino acid substitution D10A or its corresponding amino acid substitution. In some embodiments, the fusion protein of any of the above aspects further includes one or more uracil DNA glycosylase inhibitors (UGI). In some embodiments, the one or more UGI are derived from Bacillus subtilis phage PBS1 and inhibit human UDG activity. In some embodiments, the fusion protein includes two uracil DNA glycosylase inhibitors (UGI). In some embodiments, the fusion protein of any of the above aspects further includes one or more linkers. In some embodiments, the fusion protein deaminates nucleobases in a target nucleotide sequence, and wherein the deamination reaction has an increased cis-activity to trans-activity ratio (cis:trans) compared to a standard cytidine base editor.
[0075] In some embodiments, the standard cytidine base editor includes (i) a polynucleotide programmable DNA binding domain and (ii) an APOBEC cytidine deaminase.
[0076] In some embodiments, the APOBEC cytidine deaminase of the standard cytidine base editor is rat APOBEC-1 cytidine deaminase (rAPOBEC-1). In some embodiments, the polynucleotide programmable DNA binding domain of the standard cytidine base editor is Cas9 nickase. In some embodiments, the standard cytidine base editor includes a uracil glycosylase inhibitor (UGI) domain. In some embodiments, the standard cytidine base editor is BE3 or BE4. In some embodiments, the increased cis-activity to trans-activity ratio is increased by at least 2, 2.5, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60-fold or more. In some embodiments, the cytidine base editor has at least 50%, 60%, 70%, 80%, 90%, 95%, 100%, 105%, 110%, 115%, 120%, or higher cis-activity compared to the standard cytidine base editor. In some embodiments, the cytidine base editor has at least 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, or more-fold lower trans-activity compared to the standard cytidine base editor.
[0077] In one aspect, the present disclosure provides a polynucleotide molecule encoding the fusion protein of any of the above aspects. In some embodiments, the polynucleotide molecule is codon-optimized.
[0078] The present disclosure provides an expression vector comprising the above-described polynucleotide molecule. In some embodiments, the expression vector is a mammalian expression vector. In some embodiments, the vector is a viral vector selected from the group consisting of adeno-associated virus (AAV) vectors, retroviral vectors, adenoviral vectors, lentiviral vectors, Sendai virus vectors, and herpesvirus vectors. In some embodiments, the vector comprises a promoter.
[0079] The present disclosure provides a cell comprising the above-described polynucleotide or the above-described vector. In some embodiments, the cell is a bacterial cell, a plant cell, an insect cell, a human cell, or a mammalian cell.
[0080] The present disclosure provides a molecular complex comprising a fusion protein of any of the above aspects and one or more guide RNA sequences, tracrRNA sequences, or target DNA sequences.
[0081] The present disclosure provides a kit comprising a fusion protein of any of the above aspects, the above-described polynucleotide, the above-described vector, or the above-described molecular complex.
[0082] The present disclosure provides a method of editing a nucleobase of a nucleic acid sequence, the method comprising contacting the nucleic acid sequence with a base editor comprising a fusion protein of any of the above aspects and converting a first nucleobase of the DNA sequence to a second nucleobase. In some embodiments, the first nucleobase is cytosine and the second nucleobase is thymidine.
[0083] The present disclosure provides a method of editing a nucleobase of a nucleic acid sequence, the method comprising contacting the nucleic acid sequence with a base editor comprising a fusion protein of any of the above aspects and converting a first nucleobase of the DNA sequence to a second nucleobase. In some embodiments, the first nucleobase is cytosine and the second nucleobase is thymidine or the first nucleobase is adenine and the second nucleobase is guanine. In some embodiments, the method further comprises converting a third nucleobase to a fourth nucleobase. In some embodiments, the third nucleobase is guanine and the fourth nucleobase is adenine or the third nucleobase is thymine and the fourth nucleobase is cytosine.
[0084] The present disclosure provides a method for optimal base editing, the method comprising contacting a target nucleobase in a target nucleotide sequence with a cytidine base editor comprising (i) a polynucleotide programmable DNA binding domain and (ii) a cytidine deaminase, wherein the cytidine base editor deaminates the target nucleobase in the target nucleotide sequence with a lower false deamination reaction as compared to a canonical cytidine base editor comprising rAPOBEC1. In some embodiments, the cytidine base editor deaminates the target nucleobase with a higher efficiency as compared to the canonical cytidine base editor. In some embodiments, the canonical cytidine base editor further comprises a uracil glycosylase inhibitor (UGI) domain. In some embodiments, the canonical cytidine base editor is BE3 or BE4. In some embodiments, as measured by a cis / trans deamination assay, the cytidine base editor generates at least 20%, 30%, 50%, 70%, or 90% less false deamination reaction as compared to the canonical cytidine base editor. In some embodiments, the cytidine base editor has at least 50%, 60%, 70%, 80%, 90%, 95%, 100%, 105%, 110%, 115%, 120%, or more cis activity as compared to the canonical cytidine base editor. In some embodiments, the cytidine base editor has at least 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, or more times lower trans activity as compared to the canonical cytidine base editor. In some embodiments, the cytidine deaminase is (a) APOBEC-1 from Mesocricetus auratus (MaAPOBEC-1), Pongo pygmaeus (PpAPOBEC-1), Oryctolagus cuniculus (OcAPOBEC-1), Monodelphis domestica (MdAPOBEC-1), or Alligator mississippiensis (AmAPOBEC-1), (b) APOBEC-2 from Pongo pygmaeus (PpAPOBEC-2), Bos taurus (BtAPOBEC-2), or Sus scrofa (SsAPOBEC-2), (c) APOBEC-4 from Macaca fascicularis (MfAPOBEC-4), (d) AID from Canis lupus familiaris (ClAID) or Bos taurus (BtAID), (e) yeast cytosine deaminase (yCD) from Saccharomyces cerevisiae, (f) APOBEC-3F from Rhinopithecus roxellana (RrA3F), or (g) a cytidine deaminase having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical amino acid sequence to any one of the proteins in (a)-(f).
[0085] In some embodiments, the cytidine deaminase is AID, which is from dog (ClAID), cattle (BtAID), or a cytidine deaminase having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity to the foregoing proteins. In some embodiments, the cytidine deaminase is APOBEC-3F, which is from Rhinopithecus roxellana (RrA3F), or a cytidine deaminase having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity to the foregoing proteins.
[0086] In some embodiments, the cytidine deaminase comprises a change selected from the group consisting of: R15X, R16X, H21X, R30X, R33X, K34X, R52X, K60X, R118X, H121X, H122X, R126X, R128X, R169X, R198X, T36X, H53X, V62X, L88X, W90X, Y120X, and R132X (numbered as in SEQ ID NO:1) or a corresponding change of said change, wherein X is any amino acid. In some embodiments, the cytidine deaminase comprises a change selected from the group consisting of: R15A, R16A, H21A, R30A, R33A, K34A, R52A, K60A, R118A, H121A, H122A, H122L, R126A, R128A, R169A, R198A, T36A, H53A, V62A, L88A, W90F, W90A, Y120F, Y120A, H121R, H122R, R126E, W90Y, and R132E (numbered as in SEQ ID NO:1) or a corresponding change of said change.
[0087] In some embodiments, the cytidine deaminase comprises a combination of changes selected from the group consisting of: K34A+R33A, K34A+H122A, K34A+Y120F, K34A+R52A, K34A+H122A, K34A+H121A, W90A+R126E, W90Y+R126E, H121R+H122R, R126+R132E, W90Y+R132E, and W90Y+R126E+R132E (numbered as in SEQID NO:1) or a combination of corresponding changes of said combination of changes.
[0088] In some embodiments, the cytidine deaminase comprises a change at position Y120F and one or more changes selected from the group consisting of: R33A, W90F, K34A, R52A, H122A, and H121A (numbered as in SEQ ID NO:1), or corresponding changes of one or more of said changes. In some embodiments, the cytidine deaminase comprises a change at position Y130X or R28X (numbered as in SEQ ID NO:1) or a corresponding change of said change, where X is any amino acid. In some embodiments, the cytidine deaminase comprises a Y130A change or an R28A change (numbered as in SEQ ID NO:1) or a corresponding change of said change. In some embodiments, the cytidine deaminase comprises Y130A and R28A changes (numbered as in SEQ ID NO:1) or a corresponding change of said change.
[0089] In some embodiments, the cytidine deaminase comprises changes at positions H122X, K34X, R33X, W90X, and R128X (numbered as in SEQ ID NO:1) or corresponding changes of said changes, where X is any amino acid. In some embodiments, the cytidine deaminase comprises changes selected from the group consisting of: H122A, K34A, R33A, W90F, W90A, and R128A (numbered as in SEQ ID NO:1), or corresponding changes of said changes. In some embodiments, the cytidine deaminase comprises a combination of changes selected from the group consisting of: R33A+K34A, W90F+K34A, R33A+K34A+W90F, and R33A+K34A+H122A+W90F (numbered as in SEQ ID NO:1) or a corresponding combination of changes of said combination of changes.
[0090] In some embodiments, the cytidine deaminase comprises an amino acid sequence having at least 80% identity to the following amino acid sequence:
[0091] MTSEKGPSTGDPTLRRRIESWEFDVFYDPRELRKETCLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERRFHSSISCSITWFLSWSPCWECSQAIREFLSQHPGVTLVIYVARLFWHMDQRNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLAFFRLHLQNCHYQTIPPHILLATGLIHPSVTWR.
[0092] In some embodiments, the cytidine deaminase comprises an amino acid sequence that has at least 80% identity to the following amino acid sequence:
[0093] MKPQIRDHRPNPMEAMYPHIFYFHFENLEKAYGRNETWLCFTVEIIKQYLPVPWKKGVFRNQVDPETHCHAEKCFLSWFCNNTLSPKKNYQVTWYTSWSPCPECAGEVAEFLAEHSNVKLTIYTARLYYFWDTDYQEGLRSLSEEGASVEIMDYEDFQYCWENFVYDDGEPFKRWKGLKYNFQSLTRRLREILQ.
[0094] In some embodiments, the cytidine deaminase comprises an amino acid sequence that has at least 80% identity to the following amino acid sequence:
[0095] MADSSEKMRGQYISRDTFEKNYKPIDGTKEAHLLCEIKWGKYGKPWLHWCQNQRMNIHAEDYFMNNIFKAKKHPVHCYVTWYLSWSPCADCASKIVKFLEERPYLKLTIYVAQLYYHTEEENRKGLRLLRSKKVIIRVMDISDYNYCWKVFVSNQNGNEDYWPLQFDPWVKENYSRLLDIFWESKCRSPNPW.
[0096] In some embodiments, the cytidine deaminase comprises an amino acid sequence that has at least 80% identity to the following amino acid sequence:
[0097] MDPQRLRQWPGPGPASRGGYGQRPRIRNPEEWFHELSPRTFSFHFRNLRFASGRNRSYICCQVEGKNCFFQGIFQNQVPPDPPCHAELCFLSWFQSWGLSPDEHYYVTWFISWSPCCECAAKVAQFLEENRNVSLSLSAARLYYFWKSESREGLRRLSDLGAQVGIMSFQDFQHCWNNFVHNLGMPFQPWKKLHKNYQRLVTELKQILREEPATYGSPQAQGKVRIGSTAAGLRHSHSHTRSEAHLRPNHSSRQHRILNPPREARARTCVLVDASWICYR.
[0098] In some embodiments, the cytidine deaminase comprises an H122A alteration. In some embodiments, the contacting is performed in a cell. In some embodiments, the cell is a human cell or a mammalian cell. In some embodiments, the contacting is in vivo or ex vivo.
[0099] In one aspect, provided herein is a cytidine deaminase comprising an amino acid sequence that has at least 80% identity to an amino acid sequence selected from:
[0100] MTSEKGPSTGDPTLRRRIESWEFDVFYDPRELRKETCLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERRFHSSISCSITWFLSWSPCWECSQAIREFLSQHPGVTLVIYVARLFWHMDQRNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLAFFRLHLQNCHYQTIPPHILLATGLIHPSVTWR;
[0101] MKPQIRDHRPNPMEAMYPHIFYFHFENLEKAYGRNETWLCFTVEIIKQYLPVPWKKGVFRNQVDPETHCHAEKCFLSWFCNNTLSPKKNYQVTWYTSWSPCPECAGEVAEFLAEHSNVKLTIYTARLYYFWDTDYQEGLRSLSEEGASVEIMDYEDFQYCWENFVYDDGEPFKRWKGLKYNFQSLTRRLREILQ;
[0102] MADSSEKMRGQYISRDTFEKNYKPIDGTKEAHLLCEIKWGKYGKPWLHWCQNQRMNIHAEDYFMNNIFKAKKHPVHCYVTWYLSWSPCADCASKIVKFLEERPYLKLTIYVAQLYYHTEEENRKGLRLLRSKKVIIRVMDISDYNYCWKVFVSNQNGNEDYWPLQFDPWVKENYSRLLDIFWESKCRSPNPW; and
[0103] MDPQRLRQWPGPGPASRGGYGQRPRIRNPEEWFHELSPRTFSFHFRNLRFASGRNRSYICCQVEGKNCFFQGIFQNQVPPDPPCHAELCFLSWFQSWGLSPDEHYYVTWFISWSPCCECAAKVAQFLEENRNVSLSLSAARLYYFWKSESREGLRRLSDLGAQVGIMSFQDFQHCWNNFVHNLGMPFQPWKKLHKNYQRLVTELKQILREEPATYGSPQAQGKVRIGSTAAGLRHSHSHTRSEAHLRPNHSSRQHRILNPPREARARTCVLVDASWICYR.
[0104] The description and examples herein illustrate in detail the embodiments disclosed in the present disclosure. It should be understood that the present disclosure is not limited to the specific embodiments described herein and can thus vary. Those skilled in the art will recognize the existence of many variations and modifications of the present disclosure, which are covered by the scope of the present disclosure.
[0105] Unless otherwise indicated, the practice of some embodiments disclosed herein uses conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, which are within the skill of the art. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); Current Protocols in Molecular Biology Series (edited by F.M. Ausubel et al.); Methods in Enzymology Series (Academic Press); PCR 2: A Practical Approach (edited by M.J. MacPherson, B.D. Hames, and G.R. Taylor (1995)); Antibodies, A Laboratory Manual (edited by Harlow and Lane (1988)); and Animal Cell Culture: A Manual of Basic Technique and Specialized Applications, 6th Edition (edited by R.I. Freshney (2010)).
[0106] The paragraph headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described.
[0107] Although various features of the present disclosure may be described in the context of a single embodiment, these features may also be provided separately or in any suitable combination. Conversely, although the present disclosure may be described herein for clarity in the context of a single embodiment, the present disclosure may also be practiced in a single embodiment. The paragraph headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described.
[0108] The features disclosed in this disclosure are set forth in detail in the appended claims. The features and advantages of the present invention will be better understood by reference to the following detailed description which sets forth illustrative embodiments, in which the principles of the present disclosure are utilized, and by reference to the drawings as described hereinafter.
[0109] Definitions
[0110] The following definitions supplement the original definitions in the art and are directed to the current application and are not to be attributed to any related or unrelated case, e.g., any co-owned patent or application. Although any methods and materials similar or equivalent to those described herein may be used in testing the practice of the present disclosure, the preferred materials and methods are described herein. Accordingly, the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.
[0111] Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. The following references provide one of ordinary skill in the art with the ordinary definitions of many of the terms used in this invention: Dictionary of Microbiology and Molecular Biology (2nd ed., Singleton et al. (1994)); Cambridge Dictionary of Science and Technology (Walker (ed.), 1988); Glossary of Genetics (5th ed., R. Rieger et al. (eds.), Springer-Verlag (1991)); and HarperCollins Dictionary of Biology (1991, Hale and Marham).
[0112] In this application, unless otherwise specifically stated, the use of the singular includes the plural. It must be noted that, as used in the specification, unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" include plural referents. In this application, unless otherwise stated, the use of "or" means "and / or" and is understood to be inclusive. Further, when the term "including" and other forms thereof, such as "include", "includes", and "included", are used, they are not limiting.
[0113] As used in this specification and the claims, the words "comprising" (and any form of comprising, such as "comprise" and "comprises"), "having" (and any form of having, such as "have" and "has"), "including" (and any form of including, such as "includes" and "include") or "containing" (and any form of containing, such as "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. It should be expected that any embodiment discussed in this specification can be practiced in conjunction with any method or composition disclosed in this disclosure, and vice versa. In addition, the compositions disclosed in this disclosure can be used to implement the methods disclosed in this disclosure.
[0114] The terms "about" or "approximately" mean within an acceptable error range of a particular numerical value (as determined by a person of ordinary skill in the art), which will depend in part on how the numerical value is measured or determined, i.e., the limitations of the measuring system. For example, in accordance with practice in the art, "about" can mean within one standard deviation or more than one standard deviation. Alternatively, "about" can mean within a range of 20%, 10%, 5%, or 1% of a given numerical value. Or, particularly for biological systems or processes, the term can mean within an order of magnitude, such as within 5-fold or 2-fold of a numerical value. When describing a particular numerical value in the application and claims, unless otherwise stated, the term "about" should be assumed to mean within an acceptable error range of the particular numerical value.
[0115] The ranges provided herein should be understood as a shorthand representation of all the numerical values within that range. For example, a range of 1 to 50 should be understood to include any number, combination of numbers, or any sub-range that comes from the group consisting of: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.
[0116] References in the specification to "some embodiments", "an embodiment", "one embodiment", or "other embodiments" mean that the particular features, structures, or characteristics described in connection with that embodiment are included in at least some embodiments of the disclosure, but not necessarily in all embodiments of the disclosure.
[0117] "Base-free base editor" means an agent capable of excising a nucleobase and inserting a DNA nucleobase (A, T, C, or G). The base-free base editor includes a nucleic acid glycosylase polypeptide or a fragment thereof. In one embodiment, the nucleic acid glycosylase is a mutant of human uracil DNA glycosylase, which includes Asp at amino acid 204 (or the corresponding position in uracil DNA glycosylase) in the following sequence (e.g., substituting Asn at amino acid 204), and has cytosine-DNA glycosylase activity or an active fragment thereof. In one embodiment, the nucleic acid glycosylase is a mutant of human uracil DNA glycosylase, which includes Ala, Gly, Cys, or Ser at amino acid 147 (or the corresponding position in uracil DNA glycosylase) in the following sequence (e.g., substituting Tyr at amino acid 147), and has thymine-DNA glycosylase activity, or an active fragment thereof. The sequence of an exemplary human uracil-DNA glycosylase, isoform 1, is as follows:
[0118]
[0119] The sequence of human uracil-DNA glycosylase, isoform 2, is as follows:
[0120]
[0121] In other embodiments, the base-free (base) editor is any of the base-free (base) editors described in PCT / JP2015 / 080958 and US20170321210 (which are incorporated herein by reference). In a particular embodiment, the base-free (base) editor includes a mutation that is at the position bolded and underlined in the above sequence, or at the corresponding amino acid position in any other base-free (base) editor or uracil deglycosylase known in the art. In one embodiment, the base-free (base) editor includes mutations at Y147, N204, L272, and / or R276, or the corresponding positions. In another embodiment, the base-free (base) editor includes a Y147A or Y147G mutation, or the corresponding mutation. In another embodiment, the base-free (base) editor includes an N204D mutation, or the corresponding mutation. In another embodiment, the base-free (base) editor includes an L272A mutation, or the corresponding mutation. In another embodiment, the base-free (base) editor includes an R276E or R276C mutation, or the corresponding mutation.
[0122] "Adenosine deaminase" means a polypeptide or a fragment thereof that can catalyze the hydrolytic deamination reaction of adenine or adenosine. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination reaction of adenosine to inosine or deoxyadenosine to deoxyinosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination reaction of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) can be from any organism, such as bacteria.
[0123] In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the TadA deaminase is a TadA variant. In some embodiments, the TadA variant is TadA*7.10. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain does not exist in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to a naturally occurring deaminase. For example, deaminase domains are described in International PCT Application Nos. PCT / 2017 / 045381 (WO 2018 / 027078) and PCT / US2016 / 058344 (WO 2017 / 070632), the entire contents of each of which are incorporated herein by reference. See also Komor, A.C. et al. "Programmable editing of a target base in genomic DNA without double-strand DNA cleavage" Nature 533, 420-424 (2016); Gaudelli, N.M. et al. "Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage" Nature 551, 464-471 (2017); Komor, A.C. et al. "Improved base excision repair inhibition enables C:G-to-T:A base editors with higher efficiency and higher product purity" Science Advances 3: eaao4774 (2017), and Rees, H.A. et al. "Base editing: precision chemistry on the genome and transcriptome of living cells" Nat Rev Genet. 2018 Dec; 19(12):770-788. doi:10.1038 / s41576-018-0059-1, the entire contents of which are incorporated herein by reference.
[0124] In some embodiments, the adenosine deaminase comprises alterations in the following sequences:
[0125] MSEVEFSHEY WMRHALTLAK RARDEREVPVGAVLVLNNRV IGEGWNRAIG LHDPTAHAEIMALRQGGLVM QNYRLIDATL YVTFEPCVMC AGAMIHSRIG RVVFGVRNAK TGAAGSLMDV LHYPGMNHRVEITEGILADE CAALLCYFFR MPRQVFNAQK KAQSSTD
[0126] (also known as TadA*7.10).
[0127] In certain embodiments, the adenosine deaminase heterodimer comprises a TadA*7.10 domain and an adenosine deaminase domain selected from one of the following:
[0128] Staphylococcus aureus (S. aureus) TadA:
[0129] MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN
[0130] Bacillus subtilis (B. subtilis) TadA:
[0131] MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLSE
[0132] Salmonella typhimurium (S. typhimurium) TadA:
[0133] MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIKALKKADRAEGAGPAV
[0134] Shewanella putrefaciens TadA:
[0135] MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE
[0136] Haemophilus influenzae F3031 TadA:
[0137] MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRREEKKIEKALLKSLSDK
[0138] Caulobacter crescentus TadA:
[0139] MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAKI
[0140] Geobacter sulfurreducens TadA:
[0141] MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKAKATPALFIDERKVPPEP
[0142] TadA*7.10
[0143] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD
[0144] "Administering" as used herein refers to providing a patient or subject with one or more of the compositions described herein. By way of example and not limitation, administration of a composition, such as an injection, can be carried out by intravenous (i.v.) injection, subcutaneous (s.c.) injection, intradermal (i.d.) injection, intraperitoneal (i.p.) injection, or intramuscular (i.m.) injection. One or more of the foregoing routes can be employed. Parenteral administration can be, for example, by bolus injection or by infusion over time. In some embodiments, parenteral administration includes intravascular, intravenous, intramuscular, intraarterial, intrathecal, intratumoral, intradermal, intraperitoneal, transtracheal, subcutaneous, transdermal, intraarticular, subcapsular, subarachnoid, and intrasternal infusion or injection. Alternatively, or concomitantly, administration can be by an oral route.
[0145] "Agent" means any small molecule chemical compound, antibody, nucleic acid molecule, or polypeptide, or fragment thereof.
[0146] "Alteration" means a change (e.g., an increase or decrease) in the structure, expression level, or activity of a gene or polypeptide detected by standard methods known in the art, such as those described herein. As used herein, alteration includes a change in the polynucleotide or polypeptide sequence or a change in the expression level, such as a 10% change, 25% change, 40% change, 50% or greater change.
[0147] "Ameliorate" means to reduce, inhibit, attenuate, decrease, arrest, or stabilize the development or progression of a disease.
[0148] "Analog" means a molecule that is not identical but has similar functional or structural characteristics. For example, a polynucleotide analog or a polypeptide analog retains the biological activity of the corresponding naturally occurring polynucleotide or polypeptide while having certain modifications that enhance the function of the analog relative to the naturally occurring polynucleotide or polypeptide. Such modifications can increase the analog's affinity for DNA, potency, specificity, protease or nuclease resistance, membrane permeability, and / or half-life without altering, for example, ligand binding. An analog may contain unnatural nucleotides or amino acids.
[0149] "Base editor (BE)" or "nucleobase editor (NBE)" means an agent that binds to a polynucleotide and has nucleobase modification activity. In several embodiments, the base editor comprises a nucleobase-modifying polypeptide (e.g., a deaminase) and a nucleic acid-programmable nucleotide-binding domain that is linked to a guide polynucleotide (e.g., guide RNA). In several embodiments, the agent is a biomolecular complex that comprises a protein domain having base editing activity, i.e., a domain capable of modifying a base (e.g., A, T, C, G, or U) within a nucleic acid molecule (e.g., DNA). In some embodiments, the polynucleotide-programmable DNA-binding domain is fused to or linked to one or more deaminase domains. In one embodiment, the agent is a fusion protein that comprises one or more domains having base editing activity. In another embodiment, the protein domain having base editing activity is linked to the aforementioned guide RNA (e.g., via an RNA-binding motif on the guide RNA and an RNA-binding domain fused to the deaminase). In some embodiments, the domain having base editing activity is capable of deaminating a base within a nucleic acid molecule. In some embodiments, the base editor is capable of deaminating one or more bases within a DNA molecule. In some embodiments, the base editor is capable of deaminating cytosine (C) or adenosine (A) in DNA. In some embodiments, the base editor is capable of deaminating both cytosine (C) and adenosine (A) in DNA. In some embodiments, the base editor is capable of deaminating cytosine (C) in DNA. In some embodiments, the base editor is a cytidine base editor (CBE) (e.g., BE4). In some embodiments, the base editor is capable of deaminating adenosine (A) in DNA. In some embodiments, the base editor is a standard base editor that comprises naturally occurring protein domains having base editing activity and / or programmable DNA-binding activity. For example, a standard cytidine base editor may contain a cytidine deaminase, such as an APOBEC cytidine deaminase or an AID deaminase. In some embodiments, the standard cytidine deaminase contains APOBEC1 cytidine deaminase, such as rAPOBEC1. In some embodiments, the standard cytidine base editor further comprises additional domains that are associated with or linked to the cytidine deaminase, e.g., one or more UGI domains that may be linked to the cytidine deaminase. In some embodiments, the base editor is an adenosine base editor (ABE) and a cytidine base editor (CBE).
[0150] In some embodiments, the base editor is a nuclease-inactivated Cas9 (dCas9) fused to an adenosine deaminase and / or a cytidine deaminase. In some embodiments, the Cas9 is a circularly permuted Cas9 (e.g., spCas9 or saCas9). Circularly permuted Cas9 is known in the art and is described, for example, in Oakes et al., Cell 176, 254–267, 2019. In some embodiments, the base editor is fused to an inhibitor of base excision repair, e.g., the UGI domain or the dISN domain. In some embodiments, the fusion protein comprises a Cas9 nickase fused to one or more deaminases and an inhibitor of base excision repair, such as the UGI or dISN domain). In other embodiments, the base editor is a base-free base editor.
[0151] In some embodiments, an adenosine base editor is generated by cloning an adenosine deaminase variant into a scaffold that comprises a circularly permuted Cas9 (e.g., spCAS9 or saCAS9) and a bipartite nuclear localization sequence. Circularly permuted Cas9 is known in the art and is described, for example, in Oakes et al., Cell 176, 254–267, 2019. Exemplary circularly permuted Cas9s are as follows, where the bolded sequences represent sequences derived from Cas9, the italicized sequences represent linker sequences, and the underlined sequences represent bipartite nuclear localization sequences.
[0152] CP5 (with MSP “NGC = Pam variant with mutations, normal Cas9 likes NGG” PID = protein interaction domain and “D10A” nickase):
[0153]
[0154]
[0155] In some embodiments, the polynucleotide programmable DNA binding domain is a CRISPR-associated (e.g., Cas or Cpf1) enzyme. In some embodiments, the base editor is a catalytically inactivated Cas9 (dCas9) fused to one or more deaminase domains. In some embodiments, the base editor is a Cas9 nickase (nCas9) fused to one or more deaminase domains. In some embodiments, the base editor is fused to an inhibitor of base excision repair (BER). In some embodiments, the inhibitor of base excision repair is a uracil DNA glycosylase inhibitor (UGI). In some embodiments, the inhibitor of base excision repair is an inosine base excision repair inhibitor.
[0156] Details of base editors are described in International PCT Application Nos. PCT / 2017 / 045381 (WO 2018 / 027078) and PCT / US2016 / 058344 (WO 2017 / 070632), the entire contents of each of which are incorporated herein by reference. See also Komor, A.C. et al., "Programmable editing of a target base in genomic DNA without double-strand DNA cleavage" Nature 533, 420-424 (2016); Gaudelli, N.M. et al., "Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage" Nature 551, 464-471 (2017); Komor, A.C. et al., "Improved base excision repair inhibition enables C:G-to-T:A base editors with higher efficiency and higher product purity" Science Advances 3: eaao4774 (2017), and Rees, H.A. et al., "Base editing: precision chemistry on the genomes and transcriptomes of living cells" Nat Rev Genet. 2018 Dec;19(12):770-788. doi:10.1038 / s41576-018-0059-1, the entire contents of which are incorporated herein by reference.
[0157] For example, an adenine base editor (ABE) as used in the base editing compositions, systems, and methods described herein has the nucleic acid sequence (8877 base pairs) provided below (Addgene, Watertown, MA.; Gaudelli NM et al., Nature. 2017 Nov 23;551(7681):464-471. doi:10.1038 / nature24644; Koblan LW et al., Nat Biotechnol. 2018 Oct;36(9):843-846. doi:10.1038 / nbt.4172.). Polynucleotide sequences having at least 95% or higher identity to the ABE nucleic acid sequence are also encompassed.
[0158]
[0159] For example, the cytidine base editor (CBE) used in the base editing compositions, systems, and methods as described herein has the following nucleic acid sequence (8,877 base pairs) provided below (Addgene, Watertown, MA.; Komor AC et al., 2017, Sci Adv., 30; 3(8): eaa04774. doi: 10.1126 / sciadv.aao4774). Polynucleotide sequences having at least 95% or higher identity to this BE4 nucleic acid sequence are also encompassed.
[0160]
[0161]
[0162]
[0163]
[0164] In some embodiments, the cytidine base editor is BE4 and has a nucleic acid sequence selected from one of the following:
[0165] The original BE4 nucleic acid sequence:
[0166]
[0167]
[0168] BE4 codon-optimized 1 nucleic acid sequence:
[0169]
[0170]
[0171]
[0172] BE4 codon-optimized 2 nucleic acid sequence:
[0173] ATGAGCAGCGAGACAGGCCCTGTGGCTGTGGATCCTACACTGCGGAGAAGAATCGAGCCCCACGAGTTCGAGGTGTTCTTCGACCCCAGAGAGCTGCGGAAAGAGACATGCCTGCTGTACGAGATCAACTGGGGCGGCAGACACTCTATCTGGCGGCACACAAGCCAGAACACCAACAAGCACGTGGAAGTGAACTTTATCGAGAAGTTTACGACCGAGCGGTACTTCTGCCCCAACACCAGATGCAGCATCACCTGGTTTCTGAGCTGGTCCCCTTGCGGCGAGTGCAGCAGAGCCATCACCGAGTTTCTGTCCAGATATCCCCACGTGACCCTGTTCATCTATATCGCCCGGCTGTACCACCACGCCGATCCTAGAAATAGACAGGGACTGCGCGACCTGATCAGCAGCGGAGTGACCATCCAGATCATGACCGAGCAAGAGAGCGGCTACTGCTGGCGGAACTTCGTGAACTACAGCCCCAGCAACGAAGCCCACTGGCCTAGATATCCTCACCTGTGGGTCCGACTGTACGTGCTGGAACTGTACTGCATCATCCTGGGCCTGCCTCCATGCCTGAACATCCTGAGAAGAAAGCAGCCTCAGCTGACCTTCTTCACAATCGCCCTGCAGAGCTGCCACTACCAGAGACTGCCTCCACACATCCTGTGGGCCACCGGACTTAAGAGCGGAGGATCTAGCGGCGGCTCTAGCGGATCTGAGACACCTGGCACAAGCGAGT
[0174]
[0175]
[0176] "Base editing activity" means the action of chemically altering a base within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is cytidine deaminase activity, such as converting a target C·G to T·A. In another embodiment, the base editing activity is adenosine or adenine deaminase activity, such as converting an A·T to G·C. In another embodiment, the base editing activity is cytidine deaminase activity (such as converting a target C·G to T·A) and adenosine or adenine deaminase activity (such as converting an A·T to G·C).
[0177] The term "base editor system" refers to a system for editing nucleobases of a target nucleotide sequence. In several embodiments, the base editor system includes (1) a polynucleotide-programmable nucleotide binding domain (such as Cas9); (2) one or more deaminase domains for deaminating the nucleobases (such as adenosine deaminase and / or cytidine deaminase); and (3) one or more guide polynucleotides (such as guide RNA). In some embodiments, the base editor (BE) system includes (1) a polynucleotide-programmable nucleotide binding domain (such as Cas9), an adenosine deaminase domain and a cytidine deaminase domain, which are used for deaminating the nucleobases in the target nucleotide sequence; and (2) one or more guide polynucleotides (such as guide RNA), which are linked to the polynucleotide-programmable nucleotide binding domain. In some embodiments, the polynucleotide-programmable nucleotide binding domain is a polynucleotide-programmable DNA binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor system is BE4. In some embodiments, the base editor is an adenine or adenosine base editor (ABE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE) and a cytidine base editor (CBE). In some embodiments, the base editor is a base-free (base) editor.
[0178] In some embodiments, the base editor system may include more than one base editing component. For example, the base editor system may contain one or more deaminases (such as adenosine deaminase, cytidine deaminase). In some embodiments, a single guide polynucleotide can be used to target different deaminases to the target nucleic acid sequence. In some embodiments, a single pair of guide polynucleotides can be used to target different deaminases to the target nucleic acid sequence.
[0179] The deaminase domain of the base editor system can be covalently or non-covalently linked to the polynucleotide-programmable nucleotide-binding component, or linked to each other in any combination of their linkage and interaction. For example, in some embodiments, one or more deaminase domains can be targeted to a target nucleotide sequence by a polynucleotide-programmable nucleotide-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can be fused to or linked to one or more deaminase domains. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can target one or more deaminase domains to a target nucleotide sequence through non-covalent interaction with or linkage to the deaminase domain. For example, in some embodiments, the deaminase domain can include additional heterologous portions or domains capable of interacting with, linking to, or forming a complex with an additional heterologous portion or domain that is part of the polynucleotide-programmable nucleotide-binding domain. In some embodiments, the additional heterologous portion may be capable of binding to, interacting with, linking to, or forming a complex with a polypeptide. In some embodiments, the additional heterologous portion may be capable of binding to, interacting with, linking to, or forming a complex with a polynucleotide. In some embodiments, the additional heterologous portion may be capable of binding to a guide polynucleotide. In some embodiments, the additional heterologous portion may be capable of binding to a polypeptide linker. In some embodiments, the additional heterologous portion may be capable of binding to a polynucleotide linker. The additional heterologous portion may be a protein domain. In some embodiments, the additional heterologous portion may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.
[0180] The base editor system may further include a guide polynucleotide component. It should be understood that the components of the base editor system can be linked to each other via covalent bonds, non-covalent interactions, or any combination of their associations and interactions. In some embodiments, one or more deaminase domains can be targeted to a target nucleotide sequence by a guide polynucleotide. For example, in some embodiments, the deaminase domain can include additional heterologous portions or domains (e.g., polynucleotide-binding domains such as RNA or DNA-binding domains) that are capable of interacting with, associating with, or forming a complex with a portion or segment of the guide polynucleotide (e.g., a polynucleotide motif). In some embodiments, the additional heterologous portion or domain (e.g., a polynucleotide-binding domain such as an RNA or DNA-binding protein) can be fused to or linked to the deaminase domain. In some embodiments, the additional heterologous portion may be capable of binding to, interacting with, associating with, or forming a complex with a polypeptide. In some embodiments, the additional heterologous portion may be capable of binding to, interacting with, associating with, or forming a complex with a polynucleotide. In some embodiments, the additional heterologous portion may be capable of binding to a guide polynucleotide. In some embodiments, the additional heterologous portion may be capable of binding to a polypeptide linker. In some embodiments, the additional heterologous portion may be capable of binding to a polynucleotide linker. The additional heterologous portion may be a protein domain. In some embodiments, the additional heterologous portion may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.
[0181] In some embodiments, the base editor system can further include an inhibitor of the base excision repair (BER) component. It should be understood that the components of the base editor system can be linked to each other via covalent bonds, non-covalent interactions, or any combination of their associations and interactions. The inhibitor of the BER component may include a BER inhibitor. In some embodiments, the inhibitor of the BER can be a uracil DNA glycosylase inhibitor (UGI). In some embodiments, the inhibitor of the BER can be an inosine BER inhibitor. In some embodiments, the inhibitor of the BER can be targeted to the target nucleotide sequence by a polynucleotide-programmable nucleotide-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can be fused to or linked to the inhibitor of the BER. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can be fused to or linked to one or more deaminase domains and the inhibitor of the BER. In some embodiments, the polynucleotide-programmable nucleotide-binding domain can target the inhibitor of the BER to the target nucleotide sequence by non-covalent interaction or association with the inhibitor of the BER. For example, in some embodiments, the inhibitor of the BER component can include an additional heterologous moiety or domain that can interact with, associate with, or form a complex with an additional heterologous moiety or domain that is part of the polynucleotide-programmable nucleotide-binding domain.
[0182] In some embodiments, an inhibitor of BER can be targeted to a target nucleotide sequence by a guide polynucleotide. For example, in some embodiments, the inhibitor of BER can include an additional heterologous moiety or domain (e.g., a polynucleotide-binding domain such as an RNA or DNA-binding domain) that is capable of interacting, associating, or forming a complex with a portion or segment (e.g., a polynucleotide motif) of the guide polynucleotide. In some embodiments, an additional heterologous moiety or domain of the guide polynucleotide (e.g., a polynucleotide-binding domain such as an RNA or DNA-binding domain) can be fused to or linked to the inhibitor of BER. In some embodiments, the additional heterologous moiety may be capable of binding, interacting, associating with, or forming a complex with a polynucleotide. In some embodiments, the additional heterologous moiety may be capable of binding to the guide polynucleotide. In some embodiments, the additional heterologous moiety may be capable of binding to a polypeptide linker. In some embodiments, the additional heterologous moiety may be capable of binding to a polynucleotide linker. The additional heterologous moiety may be a protein domain. In some embodiments, the additional heterologous moiety may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.
[0183] The term "Cas9" or "Cas9 domain" refers to an RNA-guided nuclease comprising a Cas9 protein or a fragment thereof (e.g., a protein that includes the DNA cleavage domain of active, inactive, or partially active Cas9, and / or the gRNA binding domain of Cas9). Cas9 nuclease is sometimes also referred to as Casnl nuclease or CRISPR (clustered regularly interspaced short palindromic repeats)-associated nuclease. CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain spacer sequences, i.e., sequences complementary to previous mobile elements, and target invading nucleic acids. The CRISPR clusters are transcribed and processed into CRISPR RNAs (crRNAs). In type II CRISPR systems, proper processing of pre-crRNA requires trans-encoded small RNAs (tracrRNAs), endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3-assisted processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA cleaves linear or circular dsDNA targets complementary to the aforementioned spacer sequences by endonucleolytic cleavage. The target strand that is not complementary to the crRNA is first cleaved by endonucleolytic cleavage and then trimmed 3'-5' by exonucleolytic cleavage. In nature, DNA-binding and cleavage generally require a protein and two RNAs. However, single guide RNAs ("sgRNAs", or simply "gRNAs") can be engineered to incorporate aspects of both crRNAs and tracrRNAs into a single RNA species. See, e.g., Jinek M. et al., Science 337:816-821 (2012), the entire content of which is incorporated herein by reference. Cas9 recognizes short motifs in the CRISPR repeats (i.e., PAM or protospacer adjacent motif) to help distinguish self from non-self.The Cas9 nuclease sequences and structures are well-known to those skilled in the art (see, for example, “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E. et al., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M. et al., Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference). Orthologs of Cas9 have been described in several species, including but not limited to Streptococcus pyogenes and Streptococcus thermophilus. Based on the present disclosure, additional suitable Cas9 nucleases and sequences will be apparent to those skilled in the art, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference).
[0184] Exemplary Cas9, Streptococcus pyogenes Cas9 (spCas9), has the amino acid sequence provided below:
[0185]
[0186]
[0187] (Single underline: HNH domain; double underline: RuvC domain)
[0188] The nuclease-inactivated Cas9 protein can be interchangeably referred to as the "dCas9" protein (for nuclease-"dead" Cas9) or catalytically inactivated Cas9. Methods for generating Cas9 proteins (or fragments thereof) with an inactivated DNA cleavage domain are known (see, e.g., Jinek et al., Science. 337:816-821 (2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell. 28; 152(5):1173-83, the entire contents of each of which are incorporated herein by reference). For example, it is known that the DNA cleavage domain of Cas9 contains two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of Streptococcus pyogenes Cas9 (Jinek et al., Science. 337:816-821 (2012); Qi et al., Cell. 28; 152(5):1173-83 (2013)). In some embodiments, the Cas9 nuclease has an inactivated (e.g., inactivated) DNA cleavage domain, that is, the Cas9 is a nickase, referred to as the "nCas9" protein (for "nickase" Cas9). In some embodiments, a protein comprising a fragment of Cas9 is provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) the gRNA binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, a protein comprising Cas9 or a fragment thereof is referred to as a "Cas9 variant". Cas9 variants share homology with Cas9 or a fragment thereof. For example, a Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to wild-type Cas9.In some embodiments, compared to wild-type Cas9, the Cas9 variant may have a change of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acids. In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., the gRNA-binding domain or the DNA-cleaving domain) such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding wild-type Cas9 fragment. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of the corresponding wild-type Cas9.
[0189] In some embodiments, the length of the fragment is at least 100 amino acids. In some embodiments, the length of the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids.
[0190] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1, nucleotide and amino acid sequences are as follows).
[0191]
[0192] AAAAAGGAAATGAGCTGGCTCTGCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAACAAAAACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTTATTTTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGAAAATATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTAAACGATATACGTCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATTGATTTGAGTCAGCTAGGAGGTGACTGA
[0193]
[0194]
[0195] (Single underline: HNH domain; double underline: RuvC domain)
[0196] In some embodiments, wild-type Cas9 corresponds to or comprises the following nucleotide and / or amino acid sequences:
[0197]
[0198]
[0199] CTTGGGGGTGACGGATCCCCCAAGAAGAAGAGGAAAGTCTCGAGCGACTACAAAGACCATGACGGTGATTATAAAGATCATGACATCGATTACAAGGATGACGATGACAAGGCTGCAGGA
[0200]
[0201]
[0202] (Single underline: HNH domain; double underline: RuvC domain)
[0203] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_002737.2) (nucleotide sequence is as follows); and Uniprot reference sequence: Q99ZW2 (amino acid sequence is as follows).
[0204]
[0205]
[0206]
[0207] (Single underline: HNH domain; double underline: RuvC domain)
[0208] In some embodiments, Cas9 refers to Cas9 from the following: Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheriae (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma aphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychrobacter torquis I (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref: YP_820832.1), Listeria innocua (NCBI Ref: NP_472073.1), Campylobacter jejuni (NCBI Ref: YP_002344900.1) or Neisseria meningitidis (NCBI Ref: YP_002342100.1) or Cas9 from any other organism.
[0209] In some embodiments, dCas9 corresponds to or comprises a part or all of the Cas9 amino acid sequence that has one or more mutations that inactivate Cas9 nuclease activity. For example, in some embodiments, the dCas9 domain comprises the D10A and H840A mutations, or corresponding mutations in another Cas9. In some embodiments, the dCas9 comprises the amino acid sequence of dCas9 (D10A and H840A):
[0210]
[0211]
[0212] (Single underline: HNH domain; double underline: RuvC domain).
[0213] In some embodiments, the Cas9 domain includes a D10A mutation, and the residue at position 840 in the amino acid sequence provided above, or the residue at the corresponding position in any amino acid sequence provided herein, remains histidine.
[0214] In other embodiments, dCas9 variants with mutations other than D10A and H840A are provided, such mutations resulting in, for example, nuclease-inactivated Cas9 (dCas9). Such mutations include, for example, other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions within the HNH nuclease subdomain and / or the RuvC1 subdomain). In some embodiments, variants or homologs of dCas9 are provided that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical. In some embodiments, variants of dCas9 are provided that are shorter or longer in amino acid sequence by about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids or more.
[0215] In some embodiments, the Cas9 fusion proteins provided herein include the full-length amino acid sequence of the Cas9 protein, e.g., one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein do not include the full-length Cas9 sequence, but rather contain only one or more fragments thereof. Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and additional suitable Cas9 domain and fragment sequences will be apparent to those skilled in the art.
[0216] It should be understood that additional Cas9 proteins (e.g., nuclease-inactivated Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9), including variants and homologs thereof, are within the scope of the present disclosure. Exemplary Cas9 proteins include, but are not limited to, those provided below. In some embodiments, the Cas9 protein is nuclease-inactivated Cas9 (dCas9). In some embodiments, the Cas9 protein is Cas9 nickase (nCas9). In some embodiments, the Cas9 protein is nuclease-active Cas9.
[0217] Exemplary catalytically inactivated Cas9 (dCas9):
[0218] DKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHV
[0219] AQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD
[0220] Exemplary catalytic Cas9 nickase (nCas9):
[0221]
[0222] Exemplary catalytically active Cas9:
[0223]
[0224] In some embodiments, Cas9 refers to Cas9 from archaea (e.g., Nanoarchaeota), which constitute a domain and kingdom of single-celled prokaryotic microorganisms. In some embodiments, Cas refers to CasX or CasY, which have been described, for example, in Burstein et al., “New CRISPR-Cas systems from uncultivated microbes.” Cell Res. 2017 Feb 21. doi:10.1038 / cr.2017.21, the entire content of which is hereby incorporated by reference herein. Using genome-resolved metagenomics, several CRISPR-Cas systems were identified, including Cas9, which was reported for the first time in the archaeal domain of life. This divergent Cas9 protein was found as part of an active CRISPR-Cas in the little-studied genus Nanoarchaeum. In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, were found, which are among the most compact systems discovered to date. In some embodiments, Cas9 refers to CasX or a variant of CasX. In some embodiments, Cas9 refers to CasY or a variant of CasY. It should be understood that other RNA-guided DNA-binding proteins may be used as nucleic acid programmable DNA-binding proteins (napDNAbps) and are within the scope of the present disclosure.
[0225] In certain embodiments, the napDNAbps useful in the methods of the present invention include circularly permuted variants, which are known in the art and are described, for example, in Oakes et al., Cell 176, 254–267, 2019. Exemplary circularly permuted variants are as follows, where the bold sequences represent sequences derived from Cas9, the italic sequences represent linker sequences, and the underlined sequences represent bipartite nuclear localization sequences.
[0226] CP5 (with MSP “NGC = Pam variant with mutations, normal Cas9 likes NGG” PID = protein interaction domain and “D10A” nickase):
[0227]
[0228]
[0229] Non-limiting examples of polynucleotide programmable nucleotide binding domains that can be incorporated into base editors include CRISPR protein-derived domains, restriction endonucleases, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs).
[0230] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) of any fusion protein provided herein can be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the amino acid sequence included in the napDNAbp is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring CasX or CasY protein. In some embodiments, the napDNAbp is a naturally occurring CasX or CasY protein. In some embodiments, the amino acid sequence included in the napDNAbp is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any CasX or CasY protein described herein. It should be understood that Cas12b / C2c1, CasX, and CasY, which are from other bacterial species, can also be used according to the disclosure of the present application.
[0231] Cas12b / C2c1(uniprot.org / uniprot / T0D7A2#2)
[0232] sp|T0D7A2|C2C1_ALIAG CRISPR-associated endo-nuclease C2c1 OS= Alicyclobacillus acido-terrestris (strain ATCC 49025 / DSM 3922 / CIP 106132 / NCIMB 13137 / GD3B) GN=c2c1 PE=1 SV=1
[0233]
[0234] CasX
[0235] (uniprot.org / uniprot / F0NN87;uniprot.org / uniprot / F0NH53)>tr|F0NN87|F0NN87_SULIH CRISPR-associated Casx protein OS=Sulfolobus islandicus (strain HVE10 / 4) GN=SiH_0402 PE=4 SV=1
[0236] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAVGQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG
[0237] >tr|F0NH53|F0NH53_SULIR CRISPR-associated protein, Casx OS=Sulfolobus islandicus (strain REY15A) GN=SiRe_0771 PE=4 SV=1
[0238] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAVGQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG
[0239] CasX of Deltaproteobacteria
[0240] MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAILQVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVAEHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFLSKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDAYNEVIARVRMWVNLNLWQKLKLSRDDAKPLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEK
[0241] RNTILEGYNYLPNENDHKKREGSLENPKKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMDEKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTDGTDIKKSGKWQGLLYGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIGRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQAAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKLAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELSAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYKSGKQPFVGAWQAFYKRRLKEVWKPNA
[0242] CasY (ncbi.nlm.nih.gov / protein / APG80656.1) > APG80656.1 CRISPR-associated protein CasY [uncultured Parcubacteria bacterium]
[0243]
[0244] The term "conservative amino acid substitution" or "conservative mutation" refers to the replacement of one amino acid by another amino acid with common properties. One functional way to define the common properties between individual amino acids is to analyze the normalized frequency of amino acid changes between the corresponding proteins of homologous organisms (Schulz, G.E. and Schirmer, R.H., Principles of Protein Structure, Springer-Verlag, New York (1979)). Based on such analysis, groups of amino acids can be defined where amino acids within a group are preferentially interchangeable and thus their effects on the overall protein structure are also most similar to each other (Schulz, G.E. and Schirmer, R.H., ibid.). Non-limiting examples of conservative mutations include amino acid substitutions of amino acids such as arginine for lysine and vice versa, while retaining the positive charge; glutamic acid for aspartic acid and vice versa, while retaining the negative charge; threonine for serine, while retaining the free –OH; and glutamine for asparagine, while retaining the free –NH2.
[0245] The term "coding sequence" or "protein-coding sequence", as used interchangeably herein, refers to a segment of a polynucleotide that encodes a protein. This region or sequence is bounded at the 5' end by a start codon and at the 3' end by a stop codon. The coding sequence may also be referred to as an open reading frame.
[0246] "Cytidine deaminase" means a polypeptide or a fragment thereof capable of catalyzing a deamination reaction that converts an amino group to a carbonyl group. In one embodiment, the cytidine deaminase converts cytosine to uracil, or 5-methylcytosine to thymine. The cytidine deaminases provided herein (e.g., engineered cytidine deaminases, evolved cytidine deaminases) can be from any organism, such as bacteria.
[0247] In some embodiments, the cytidine deaminase of the base editor can include all or part of the apolipoprotein B mRNA editing complex (APOBEC) family of deaminases. APOBEC is a family of evolutionarily conserved cytidine deaminases. Members of this family are C-to-U editing enzymes. In some embodiments, the cytidine deaminase includes, but is not limited to: APOBEC family members, including but not limited to: APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D (now referred to as "APOBEC3E"), APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, activation-induced (cytidine) deaminase (AID), hAPOBEC1, which is derived from Homo sapiens, rAPOBEC1, which is derived from Rattus norvegicus, ppAPOBEC1, which is derived from Pongo pygmaeus, AmAPOBEC1 (BEM3.31), which is derived from Alligator mississippiensis, ocAPOBEC1, which is derived from Oryctolagus cuniculus, SsAPOBEC2 (BEM3.39), which is derived from Sus scrofa, hAPOBEC3A, which is derived from Homo sapiens, maAPOBEC1, which is derived from Mesocricetus auratus, mdAPOBEC1, which is derived from Monodelphis domestica; cytidine deaminase 1 (CDA1), hA3A, which is APOBEC3A derived from Homo sapiens, RrA3F (BEM3.14), which is APOBEC3F derived from Rhinopithecus roxellana; PmCDA1, which is derived from Petromyzon marinus (Petromyzon marinus cytidine deaminase 1, "PmCDA1"); AID (activation-induced cytidine deaminase; AICDA), which is derived from mammals (e.g., human, pig, cow, horse, monkey, etc.); hAID, which is derived from Homo sapiens; and FENRY.
[0248] The term "deaminase" or "deaminase domain", as used herein, refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes the hydrolytic deamination of cytidine to uridine or deoxycytidine to deoxyuridine. In some embodiments, the deaminase or deaminase domain is a cytosine deaminase that catalyzes the hydrolytic deamination of cytosine to uracil. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or adenine (A) to inosine (I). In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine to inosine or deoxyadenosine to deoxyinosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenosine in deoxyribonucleic acid (DNA). The deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) can be from any organism, such as bacteria. In some embodiments, the deaminase is from bacteria such as Escherichia coli, Staphylococcus aureus, Salmonella typhimurium, Shewanella putrefaciens, Haemophilus influenzae, or Caulobacter crescentus.
[0249] "Detecting" means identifying the presence, absence, or amount of an analyte to be detected. In one embodiment, a sequence alteration in a polynucleotide or polypeptide is detected. In another embodiment, the presence of an insertion / deletion is detected.
[0250] "Detectable label" means a composition that, when linked to a target molecule, renders the target molecule detectable by spectroscopic, photochemical, biochemical, immunochemical, or chemical means. For example, useful labels include radioisotopes, magnetic beads, metal beads, colloidal particles, fluorescent dyes, electron-dense reagents, enzymes (e.g., as commonly used in enzyme-linked immunosorbent assays (ELISAs)), biotin, digoxin, or haptens.
[0251] "Disease" means any condition or disorder that impairs or interferes with the normal function of cells, tissues, or organs.
[0252] The term "effective amount", as used herein, refers to the amount of a bioactive agent sufficient to elicit a desired biological response. The effective amount of one (or more) active agents for practicing the present invention to treat a disease varies depending on the mode of administration, the age, weight, and general health of the subject. Ultimately, the attending physician or veterinarian will determine the appropriate amount and dosage regimen. This amount is referred to as the "effective" amount. In one embodiment, the effective amount is the amount of a base editor of the present invention (e.g., a fusion protein comprising a programmable DNA-binding protein, a nucleobase editor, and a gRNA) sufficient to introduce a change in a target gene in a cell (e.g., a cell in vivo or in vitro). In some embodiments, the effective amount of a fusion protein provided herein (e.g., a multi-effector nucleobase editor comprising an nCas9 domain and one or more deaminase domains (e.g., adenosine deaminase, cytidine deaminase)) can refer to the amount of the fusion protein sufficient to induce an editing process at a target site (which is specifically bound and edited by the multi-effector nucleobase editor). In one embodiment, the effective amount is the amount of a base editor required to achieve a therapeutic effect (e.g., reducing or controlling a disease or its symptoms or condition). Such a therapeutic effect does not require changing the target gene in all cells of a subject, tissue, or organ, but only requires changing about 1%, 5%, 10%, 25%, 50%, 75%, or more of the target genes in the existing cells in a subject, tissue, or organ.
[0253] In some embodiments, the effective amount of a fusion protein provided herein (e.g., a nucleobase editor comprising an nCas9 domain and one or more deaminase domains (e.g., adenosine deaminase, cytidine deaminase)) refers to the amount of the fusion protein sufficient to induce an editing process at a target site (which is specifically bound and edited by the multi-effector nucleobase editor described herein). Those skilled in the art will understand that the effective amount of an agent (e.g., a fusion protein, nuclease, hybrid protein, protein dimer, complex of a protein (or protein dimer) and a polynucleotide, or a polynucleotide) can vary depending on various factors, such as, by way of example, depending on the desired biological response, e.g., depending on a particular allele, genome, or target site to be edited, depending on the cell or tissue being targeted, and / or depending on the agent being used.
[0254] "Fragment" means a portion of a polypeptide or nucleic acid molecule. This portion contains at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the full length of the reference nucleic acid molecule or polypeptide. A fragment may contain 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids.
[0255] "Guide RNA" or "gRNA" means a polynucleotide that can be specific for a target sequence and can form a complex with a polynucleotide programmable nucleotide-binding domain protein (such as Cas9 or Cpf1). In one embodiment, the guide polynucleotide is guide RNA (gRNA). The gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule can be referred to as a single-guide RNA (sgRNA), although "gRNA" can be used interchangeably to refer to a guide RNA that exists as a single molecule or as a complex of two or more molecules. Generally, a gRNA that exists as a single RNA species includes two domains: (1) a domain that shares homology with the target nucleic acid (e.g., and guides the Cas9 complex to bind to the target); (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence called tracrRNA and includes a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to the tracrRNA provided in Jinek et al., Science 337:816-821 (2012), the entire content of which is incorporated herein by reference. Other examples of gRNAs (e.g., those that include domain 2) can be found in U.S. Provisional Patent Application U.S.S.N. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof," and in U.S. Provisional Patent Application U.S.S.N. 61 / 874,746, filed September 6, 2013, entitled "Delivery System For Functional Nucleases," the entire content of each of which is incorporated herein by reference. In some embodiments, the gRNA includes two or more of domains (1) and (2) and can be referred to as an "extended gRNA." As described herein, the extended gRNA will bind two or more Cas9 proteins and bind to the target nucleic acid at two or more distinct regions. The gRNA includes a nucleotide sequence complementary to the target site, which mediates the binding of the nuclease / RNA complex to the target site and provides sequence specificity for the nuclease:RNA complex.
[0256] "Hybridization" means hydrogen bonding between complementary nucleobases, which can be Watson-Crick, Hoogsteen or reverse Hoogsteen hydrogen bonding. For example, adenine and thymine are complementary nucleobases that pair by the formation of hydrogen bonds.
[0257] The term "inhibitor of base repair" or "IBR" refers to a protein that can inhibit the activity of nucleic acid repair enzymes (such as base excision repair (BER) enzymes). In some embodiments, the IBR is an inhibitor of inosine base excision repair. Exemplary inhibitors of base repair include inhibitors of APE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, hOGGl, hNEILl, T7 Endol, T4PDG, UDG, hSMUGl, and hAAG. In some embodiments, the IBR is an inhibitor of EndoV or hAAG. In some embodiments, the IBR is a catalytically inactivated EndoV or a catalytically inactivated hAAG. In some embodiments, the base repair inhibitor is an inhibitor of Endo V or hAAG. In some embodiments, the base repair inhibitor is a catalytically inactivated EndoV or a catalytically inactivated hAAG.
[0258] In some embodiments, the base repair inhibitor is a uracil glycosylase inhibitor (UGI). UGI refers to a protein that can inhibit the uracil-DNA glycosylase base-excision repair enzyme. In some embodiments, the UGI domain includes wild-type UGI or a fragment of wild-type UGI. In some embodiments, the UGI protein provided herein includes a fragment of UGI and a protein homologous to UGI or a UGI fragment. In some embodiments, the base repair inhibitor is an inhibitor of inosine base excision repair. In some embodiments, the base repair inhibitor is a "catalytically inactivated inosine-specific nuclease" or an "inactivated inosine-specific nuclease". Without wishing to be bound by any particular theory, a catalytically inactivated inosine glycosylase (such as alkyladenine glycosylase (AAG)) can bind inosine but cannot create an abasic site or remove the inosine, thereby spatially blocking the newly formed inosine moiety from the DNA damage / repair mechanism. In some embodiments, the catalytically inactivated inosine-specific nuclease can be capable of binding inosine in a nucleic acid but not cleaving the nucleic acid. Non-limiting exemplary catalytically inactivated inosine-specific nucleases include catalytically inactivated alkyladenosine glycosylase (AAG nuclease), for example, from humans, and catalytically inactivated endonuclease V (EndoV nuclease), for example, from Escherichia coli. In some embodiments, the catalytically inactivated AAG nuclease includes the E125Q mutation or a corresponding mutation in another AAG nuclease.
[0259] "Increase" means a positive change of at least 10%, 25%, 50%, 75%, or 100%.
[0260] An "intein" is a segment of a protein that is capable of excising itself and joining the remaining segments (exteins) together with peptide bonds in a process called protein splicing. Inteins are also referred to as "protein introns". The process by which an intein excises itself and joins the remaining parts of the protein is referred to herein as "protein splicing" or "intein-mediated protein splicing". In some embodiments, the intein of a precursor protein (a protein containing an intein prior to intein-mediated protein splicing) is from two genes. Such an intein is referred to herein as a split intein (e.g., split intein-N and split intein-C). For example, in cyanobacteria, DnaE, the catalytic subunit of DNA polymerase III, is encoded by two separate genes, dnaE-n and dnaE-c. The intein encoded by the dnaE-n gene may be referred to herein as "intein-N". The intein encoded by the dnaE-c gene may be referred to herein as "intein-C".
[0261] Other intein systems may also be used. For example, synthetic inteins based on the dnaE intein, namely the Cfa-N (e.g., split intein-N) and Cfa-C (e.g., split intein-C) intein pairs, have been described (e.g., in Stevens et al., J Am Chem Soc. 2016 Feb. 24;138(7):2162-5, which is incorporated herein by reference). Non-limiting examples of intein pairs that can be used according to the present disclosure include: Cfa DnaE intein, Ssp GyrB intein, Ssp DnaX intein, TerDnaE3 intein, Ter ThyX intein, Rma DnaB intein, and Cne Prp8 intein (e.g., as described in U.S. Patent No. 8,394,604, which is incorporated herein by reference).
[0262] Exemplary nucleotide and amino acid sequences of inteins are provided.
[0263] DnaE intein-N DNA:
[0264] TGCCTGTCATACGAAACCGAGATACTGACAGTAGAATATGGCCTTCTGCCAATCGGGAAGATTGTGGAGAAACGGATAGAATGCACAGTTTACTCTGTCGATAACAATGGTAACATTTATACTCAGCCAGTTGCCCAGTGGCACGACCGGGGAGAGCAGGAAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAGGGCCACTAAGGACCACAAATTTATGACAGTCGATGGCCAGATGCTGCCTATAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGAGTTGACAACCTTCCTAAT
[0265] DnaE intein-N protein:
[0266] CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDRGEQEVFEYCLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRVDNLPN
[0267] DnaE intein-C DNA:
[0268] ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGATATTGGAGTCGAAAGAGATCACAACTTTGCTCTGAAGAACGGATTCATAGCTTCTAAT
[0269] Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN
[0270] Cfa-N DNA:
[0271] TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCCTATTGGAAAGATTGTCGAAGAGAGAATTGAATGCACAGTATATACTGTAGACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCGCGGCGAACAAGAAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACGAGCAACTAAAGATCATAAATTCATGACCACTGACGGGCAGATGTTGCCAATAGATGAGATATTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTGCCA
[0272] Cfa-N protein:
[0273] CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQMLPIDEIFERGLDLKQVDGLP
[0274] Cfa-C DNA:
[0275] ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAGGAAAGTAAAGATAATATCTCGAAAAAGTCTTGGTACCCAAAATGTCTATGATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGTAGCCAGCAAC
[0276] Cfa-C protein:
[0277] MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN
[0278] Intein-N and intein-C can be fused to the N-terminal portion of split Cas9 and the C-terminal portion of split Cas9, respectively, to link the N-terminal portion of split Cas9 and the C-terminal portion of split Cas9. For example, in some embodiments, intein-N is fused to the C-terminus of the N-terminal portion of split Cas9, i.e., a structure of N--[N-terminal portion of split Cas9]-[intein-N]--C is formed. In some embodiments, intein-C is fused to the N-terminus of the C-terminal portion of split Cas9, i.e., a structure of N-[intein-C]--[C-terminal portion of split Cas9]-C is formed. The mechanism of intein-mediated protein splicing for linking the proteins (e.g., split Cas9) to which the intein is fused is known in the art, for example, as described in Shah et al., Chem Sci. 2014; 5(1):446-461, which is incorporated herein by reference. Methods for designing and using inteins are known in the art and are described, for example, in WO2014004336, WO2017132580, US20150344549, and US20180127780, the entire contents of each of which are incorporated herein by reference.
[0279] The terms "isolated", "purified", or "biologically pure" refer to different degrees of separation of a material from the components that are normally associated with it in its natural state. "Isolated" represents a degree of separation from the original source or the surrounding environment. "Purified" represents a degree of separation higher than isolation. A "purified" or "biologically pure" protein is sufficiently free of other materials such that any impurities do not substantially affect the biological properties of the protein or cause other adverse consequences. That is, a nucleic acid or peptide of the present invention is considered purified if it is produced by recombinant DNA techniques and is substantially free of cellular material, viral material, or culture medium; or, if it is chemically synthesized and is substantially free of chemical precursors or other chemicals. Purity and homogeneity are typically determined by analytical chemistry techniques, for example, polyacrylamide gel electrophoresis or high performance liquid chromatography. The term "purified" may indicate that a nucleic acid or protein produces substantially a single band in an electrophoretic gel. For proteins that can be modified, such as phosphorylation or glycosylation, different modifications may result in different isolated proteins, which can be purified separately.
[0280] "Isolated nucleic acid" means a nucleic acid (e.g., DNA) that does not contain the genes flanking that nucleic acid in the naturally occurring genome of the organism from which the nucleic acid molecule of the present invention is derived. The term thus encompasses, for example, recombinant DNA incorporated into a vector; recombinant DNA incorporated into an autonomously replicating plasmid or virus; or recombinant DNA incorporated into the genomic DNA of a prokaryote or eukaryote; or exists as a separate molecule independent of other sequences (e.g., cDNA or genomic fragments or cDNA fragments produced by PCR or restriction endonuclease digestion). Additionally, the term encompasses RNA molecules transcribed from DNA molecules, as well as recombinant DNA that is part of a hybrid gene encoding an additional polypeptide sequence.
[0281] "Isolated polypeptide" means a polypeptide of the present invention that has been separated from the components with which it is naturally associated. Generally, a polypeptide is considered isolated when it is at least 60% free, by weight, of the proteins and naturally occurring organic molecules with which it is naturally associated. Preferably, the polypeptide of the present invention is prepared to be at least 75%, more preferably at least 90%, and most preferably at least 99% pure, by weight. Isolated polypeptides of the present invention can be obtained, for example, by extraction from natural sources, by expression of recombinant nucleic acids encoding such polypeptides; or by chemical synthesis of the protein. Purity can be measured by any suitable method, such as column chromatography, polyacrylamide gel electrophoresis, or by HPLC (high performance liquid chromatography) analysis.
[0282] The term "linker", as used herein, can refer to a covalent linker (e.g., a covalent bond), a non-covalent linker, a chemical group, or a molecule that links two molecules or moieties, e.g., two components of a protein complex or ribonucleoprotein complex, or two domains of a fusion protein, such as, by way of example, a polynucleotide-programmable DNA-binding domain (e.g., dCas9) and one or more deaminase domains (e.g., adenosine deaminase and / or cytidine deaminase). The linker can connect different components or different parts of a base editor system. For example, in some embodiments, the linker can connect the guide polynucleotide-binding domain of a polynucleotide-programmable nucleotide-binding domain and the catalytic domain of a deaminase. In some embodiments, the linker can connect a CRISPR polypeptide and a deaminase. In some embodiments, the linker can connect Cas9 and a deaminase. In some embodiments, the linker can connect dCas9 and a deaminase. In some embodiments, the linker can connect nCas9 and a deaminase. In some embodiments, the linker can connect a guide polynucleotide and a deaminase. In some embodiments, the linker can connect the deaminating component and the polynucleotide-programmable nucleotide-binding component of a base editor system. In some embodiments, the linker can connect the RNA-binding portion of the deaminating component of a base editor system and the polynucleotide-programmable nucleotide-binding component. In some embodiments, the linker can connect the RNA-binding portion of the deaminating component of a base editor system and the RNA-binding portion of the polynucleotide-programmable nucleotide-binding component. The linker can be located between or on both sides of two groups, molecules, or other moieties and is linked to each of the two via a covalent bond or non-covalent interaction, thereby linking the two. In some embodiments, the linker can be an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker can be a polynucleotide. In some embodiments, the linker can be a DNA linker. In some embodiments, the linker can be an RNA linker. In some embodiments, the linker can include an aptamer capable of binding to a ligand. In some embodiments, the ligand can be a carbohydrate, peptide, protein, or nucleic acid. In some embodiments, the linker can include an aptamer derivable from a riboswitch. The riboswitch from which the aptamer is derived can be selected from the theophylline riboswitch, thiamine pyrophosphate (TPP) riboswitch, adenosylcobalamin riboswitch, S-adenosylmethionine (SAM) riboswitch, SAH riboswitch, flavin mononucleotide (FMN) riboswitch, tetrahydrofolate riboswitch, lysine riboswitch, glycine riboswitch, purine riboswitch, GlmS riboswitch, or precursor-Q nucleoside (PreQ1) riboswitch. In some embodiments, the linker can include an aptamer that binds to a polypeptide or protein domain, such as a polypeptide ligand).In some embodiments, the polypeptide ligand may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand may be part of a base editor system component. For example, a nucleobase editing component may include one or more deaminase domains and an RNA recognition motif.
[0283] In some embodiments, the linker may be one amino acid or a plurality of amino acids (e.g., a peptide or protein). In some embodiments, the linker may be about 5 - 100 amino acids in length, e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20 - 30, 30 - 40, 40 - 50, 50 - 60, 60 - 70, 70 - 80, 80 - 90, or 90 - 100 amino acids in length. In some embodiments, the linker may be about 100 - 150, 150 - 200, 200 - 250, 250 - 300, 300 - 350, 350 - 400, 400 - 450, or 450 - 500 amino acids in length. Longer or shorter linkers are also contemplated.
[0284] In some embodiments, the linker connects the gRNA-binding domain of an RNA-programmable nuclease (which includes a Cas9 nuclease domain) and the catalytic domain of a nucleic acid editing protein (such as a cytidine and / or adenosine deaminase). In some embodiments, the linker connects dCas9 and a nucleic acid editing protein. For example, the linker is located between or on both sides of two groups, molecules, or other moieties and is covalently linked to each of the two, thereby linking the two. In some embodiments, the linker is an amino acid or a plurality of amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5 to 200 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140, 150, 160, 175, 180, 190, or 200 amino acids in length. Longer or shorter linkers are also contemplated.
[0285] In some embodiments, the domains of a base editor (e.g., a multi-effector base editor) are fused via a linker comprising the amino acid sequence: SGGSSGSETPGTSESATPESSGGS, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS. In some embodiments, the domains of a base editor (e.g., a multi-effector base editor) are fused via a linker comprising the amino acid sequence: SGSETPGTSESATPES, which may also be referred to as the XTEN linker. In some embodiments, the linker comprises the amino acid sequence SGGS. In some embodiments, the linker comprises (SGGS) n , (GGGS) n , (GGGGS) n , (G) n、 (EAAAK) n , (GGS) n , SGSETPGTSESATPES, or (XP) n motif, or any combination of the foregoing, where n is independently an integer between 1 and 30, and where X is any amino acid. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15.
[0286] In some embodiments, the linker is 24 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In some embodiments, the linker is 40 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS. In some embodiments, the linker is 64 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGSSGGS. In some embodiments, the linker is 92 amino acids in length. In some embodiments, the linker comprises the amino acid sequence PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS.
[0287] "Marker" means any protein or polynucleotide having an altered expression level or activity associated with a disease or disorder.
[0288] The term "mutation", as used herein, refers to the replacement of one residue in a sequence (e.g., a nucleic acid or amino acid sequence) with another residue, or the deletion or insertion of one or more residues in the sequence. In this document, the description of a mutation is typically made by first indicating the original residue, then the position of that residue in the sequence, and then the identity of the newly substituted residue. The various methods for making the amino acid substitutions (mutations) provided herein are well known in the art, such as those provided by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)). In some embodiments, the base editors disclosed herein are capable of efficiently generating "desired mutations" (such as point mutations) within nucleic acids (e.g., nucleic acids within the genome of a subject) without generating a significant number of undesired mutations (such as undesired point mutations). In some embodiments, a desired mutation is a mutation generated by a specific base editor that binds to a guide polynucleotide (e.g., gRNA), and the guide polynucleotide is specifically designed to produce the desired mutation.
[0289] Typically, mutations created or identified in a sequence (e.g., an amino acid sequence as described herein) are numbered relative to a reference (or wild-type) sequence (i.e., a sequence that does not contain the mutation). One of ordinary skill in the art will readily understand how to determine the location of mutations in amino acid and nucleic acid sequences relative to a reference sequence.
[0290] The term "non-conservative mutation" refers to an amino acid substitution between different groups, e.g., tryptophan for lysine, serine for phenylalanine, etc. In such cases, the non-conservative amino acid substitution preferably does not interfere with or inhibit the biological activity of the functional variant. The non-conservative amino acid substitution can enhance the biological activity of the functional variant such that the biological activity of the functional variant is increased compared to the wild-type protein.
[0291] The term "nuclear localization sequence", "nuclear localization signal", or "NLS" refers to an amino acid sequence that promotes the import of a protein into the nucleus. Nuclear localization sequences are known in the art and are described, for example, in Plank et al., International PCT Application, PCT / EP2000 / 011690, filed on November 23, 2000 and published as WO / 2001 / 038547 on May 31, 2001, the disclosure of which regarding exemplary nuclear localization sequences is hereby incorporated by reference herein. In other embodiments, the NLS is an optimized NLS, which is described, for example, in Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises the amino acid sequence KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC.
[0292] The terms "nucleic acid" and "nucleic acid molecule", as used herein, refer to a compound that includes a nucleobase and an acidic moiety, such as a nucleoside, nucleotide, or polymer of nucleotides. Generally, polymeric nucleic acids, such as nucleic acid molecules that include three or more nucleotides, are linear molecules in which adjacent nucleotides are linked to each other via phosphodiester bonds. In some embodiments, "nucleic acid" refers to a single nucleic acid residue (e.g., nucleotide and / or nucleoside). In some embodiments, "nucleic acid" refers to an oligonucleotide chain that includes three or more single nucleotide residues. As used herein, the terms "oligonucleotide" and "polynucleotide" may be used interchangeably to refer to a polymer of nucleotides (e.g., a string of at least three nucleotides). In some embodiments, "nucleic acid" encompasses RNA as well as single-stranded and / or double-stranded DNA. Nucleic acids can be naturally occurring, such as in the context of a genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, chromatid, or other naturally occurring nucleic acid molecule. On the other hand, nucleic acid molecules can be non-naturally occurring molecules, such as, for example, recombinant DNA or RNA, artificial chromosomes, engineered genomes, or fragments thereof, or synthetic DNA, RNA, DNA / RNA hybrids, or include non-naturally occurring nucleotides or nucleosides. In addition, the terms "nucleic acid", "DNA", "RNA", and / or similar terms include nucleic acid analogs, such as, for example, analogs having a non-phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems, and optionally purified, chemically synthesized, etc. In appropriate circumstances, such as in the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, such as those having chemically modified bases or sugars, as well as backbone modifications. Unless otherwise indicated, nucleic acid sequences are presented in the 5′ to 3′ direction. In some embodiments, the nucleic acid is or includes natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine); chemically modified bases; biologically modified bases (e.g., methylated bases); chimeric bases; modified sugars (e.g., 2′-fluororibose, ribose, 2′-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioate and 5′-N-phosphoramidite linkages)
[0293] The term "nucleic acid programmable DNA-binding protein" or "napDNAbp" may be used interchangeably with "polynucleotide programmable nucleotide-binding domain" to refer to a protein associated with a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid that directs the napDNAbp to a specific nucleic acid sequence. In some embodiments, the polynucleotide programmable nucleotide-binding domain is a polynucleotide programmable DNA-binding domain. In some embodiments, the polynucleotide programmable nucleotide-binding domain is a polynucleotide programmable RNA-binding domain. In some embodiments, the polynucleotide programmable nucleotide-binding domain is a Cas9 protein. The Cas9 protein may be associated with a guide RNA that directs the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the napDNAbp is a Cas9 domain, such as nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactivated Cas9 (dCas9). Non-limiting examples of nucleic acid programmable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i.Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins, CARF, DinG, their homologs, or their modified or engineered versions. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, although they may not be explicitly listed in the present disclosure. See, for example, Makarova et al., "Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?" CRISPR J. 2018 Oct; 1: 325-336. doi: 10.1089 / crispr.2018.0033; Yan et al., "Functionally diverse type V CRISPR-Cas systems" Science. 2019 Jan 4; 363(6422): 88-91. doi: 10.1126 / science.aav7271, the entire contents of each of which are hereby incorporated by reference.
[0294] The terms "nucleobase", "nitrogenous base", or "base" are used interchangeably herein to refer to nitrogen-containing biological compounds that form nucleosides, which are components of nucleotides. The ability of nucleobases to form base pairs and stack on one another directly leads to long-chain helical structures such as ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). The five nucleobases – adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U) – are referred to as primary or canonical. Adenine and guanine are derived from purine, while cytosine, uracil, and thymine are derived from pyrimidine. DNA and RNA can also contain other (non-primary) modified bases. Non-limiting exemplary modified bases can include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydroxymethylcytosine. Hypoxanthine and xanthine can be produced by the presence of mutagens, both through deamination reactions (substituting an amine group with a carbonyl group). Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can be produced by the deamination of cytosine. A "nucleoside" consists of a nucleobase and a five-carbon sugar (either ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine. Examples of nucleosides with modified nucleobases include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), and pseudouridine (Ψ). A "nucleotide" consists of a nucleobase, a five-carbon sugar (either ribose or deoxyribose), and at least one phosphate group.
[0295] The term "nucleobase editing domain" or "nucleobase editing protein", as used herein, refers to a protein or enzyme that can catalyze nucleobase modification reactions in RNA or DNA, such as the deamination reactions of cytosine (or cytidine) to uracil (or uridine) or thymine (or thymidine), and adenine (or adenosine) to hypoxanthine (or inosine), as well as non-template nucleotide addition and insertion. In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., adenine deaminase or adenosine deaminase; or cytidine deaminase or cytosine deaminase). In some embodiments, the nucleobase editing domain is more than one deaminase domain (e.g., adenine deaminase or adenosine deaminase and cytidine or cytosine deaminase). In some embodiments, the nucleobase editing domain can be a naturally occurring nucleobase editing domain. In some embodiments, the nucleobase editing domain can be an engineered or evolved version of a naturally occurring nucleobase editing domain. The nucleobase editing domain can be from any organism, such as bacteria, human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse.
[0296] As used herein, "obtaining", as in "obtaining a(n) agent", includes synthesizing, purchasing, or otherwise acquiring the agent.
[0297] "Patient" or "subject", as used herein, refers to a mammalian subject or individual diagnosed with, at risk of developing, or in the process of developing, or suspected of having or developing a disease or disorder. In some embodiments, the term "patient" refers to a mammalian subject having a higher than average likelihood of developing a disease or disorder. Exemplary patients can be humans, non-human primates, cats, dogs, pigs, cows, cats, horses, camels, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, or guinea pigs), and other mammals that can benefit from the therapies disclosed herein. Exemplary human patients can be male and / or female.
[0298] "Patient in need" or "subject in need" herein refers to a patient diagnosed with, at risk of having, or having, predisposed to, or suspected of having a disease or disorder.
[0299] The terms "pathogenic mutation", "pathogenic variant", "disease-causing mutation", "disease-causing variant", "harmful mutation", or "susceptibility mutation" refer to a genetic alteration or mutation that increases an individual's susceptibility or predisposition to a disease or disorder. In some embodiments, the pathogenic mutation includes the replacement of at least one wild-type amino acid in a protein encoded by a gene with at least one pathogenic amino acid.
[0300] The term "pharmaceutically acceptable carrier" refers to a pharmaceutically acceptable material, composition, or vehicle, such as a liquid or solid filler, diluent, excipient, manufacturing aid (e.g., lubricant, talc, magnesium stearate, calcium stearate, or zinc stearate, or stearic acid), or solvent encapsulation material involved in carrying or transporting the compound from one location in the body (e.g., the delivery site) to another location (e.g., an organ, tissue, or body part of a human). A pharmaceutically acceptable carrier is "acceptable" in the sense of being compatible with the other ingredients of the formulation and not harmful to the tissues of the subject (e.g., physiologically compatible, sterile, physiological pH, etc.). Terms such as "excipient", "carrier", "pharmaceutically acceptable carrier", "vehicle", etc. are used interchangeably herein.
[0301] The term "pharmaceutical composition" means a formulated composition for pharmaceutical use.
[0302] The terms "protein", "peptide", "polypeptide" and their grammatical equivalents are used interchangeably herein and refer to polymers of amino acid residues linked together by peptide (amide) bonds. These terms refer to proteins, peptides or polypeptides of any size, structure or function. Generally, the length of a protein, peptide or polypeptide will be at least three amino acids. A protein, peptide or polypeptide can refer to a single protein or a collection of proteins. One or more of the amino acids in a protein, peptide or polypeptide can be modified, for example, by the addition of chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, geranylgeranyl groups, fatty acid groups, linkers for conjugation, functionalization or other modifications. A protein, peptide or polypeptide can also be a single molecule or can be a multimolecular complex. A protein, peptide or polypeptide can be just a fragment of a naturally occurring protein or peptide. A protein, peptide or polypeptide can be naturally occurring, recombinant, or synthetic, or any combination thereof. As used herein, the term "fusion protein" refers to a hybrid polypeptide that includes protein domains from at least two different proteins. One protein can be located in the amino-terminal (N-terminal) portion or the carboxyl-terminal (C-terminal) portion of the fusion protein, thus forming an amino-terminal fusion protein or a carboxyl-terminal fusion protein, respectively. A protein can include different domains, for example, a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9, which guides the binding of the protein to a target site) and a nucleic acid cleavage domain, or the catalytic domain of a nucleic acid editing protein. In some embodiments, a protein includes a portion of a protein, such as an amino acid sequence that constitutes a nucleic acid binding domain, and an organic compound, such as a compound that can act as a nucleic acid cleavage agent. In some embodiments, a protein forms a complex or associates with a nucleic acid (e.g., RNA or DNA). Any protein provided herein can be produced by methods known in the art. For example, the proteins provided herein can be produced via recombinant protein expression and purification, which is particularly suitable for fusion proteins that include peptide linkers. Methods for recombinant protein expression and purification are well known and include those described below: Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), the entire contents of which are incorporated herein by reference.
[0303] The polypeptides and proteins disclosed herein (including their functional portions and functional variants) may include synthetic amino acids in place of one or more naturally occurring amino acids. Such synthetic amino acids are known in the art and include, for example, aminocyclohexanecarboxylic acid, norleucine, α-amino-n-decanoic acid, homoserine, S-acetamidomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine, β-phenylserine, β-hydroxyphenylalanine, phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylglycine, indoline-2-carboxylic acid, 1,2,3,4-tetrahydroisoquinoline-3-carboxylic acid, aminomalonic acid, aminomalonic acid monoamide, N'-benzyl-N'-methyl-lysine, N',N'-dibenzyl-lysine, 6-hydroxylysine, ornithine, α-aminocyclopentanecarboxylic acid, α-aminocyclohexanecarboxylic acid, α-aminocycloheptanecarboxylic acid, α-(2-amino-2-norbornane)-carboxylic acid, α,γ-diaminobutyric acid, α,β-diaminopropionic acid, homophenylalanine, and α-tert-butylglycine. The polypeptides and proteins may be associated with post-translational modifications of one or more amino acids of the polypeptide construct. Non-limiting examples of post-translational modifications include: phosphorylation, acylation (including acetylation and formylation), glycosylation (including N-linked and O-linked), amidation, hydroxylation, alkylation (including methylation and ethylylation), ubiquitination, addition of pyrrolidonecarboxylic acid, formation of disulfide bonds, sulfation, myristoylation, palmitoylation, isoprenylation, farnesylation, geranylation, glypiation, lipoylation, and iodination.
[0304] The term "recombinant," as used herein in the context of a protein or nucleic acid, refers to a protein or nucleic acid that does not exist in nature but is the product of human engineering. For example, in some embodiments, a recombinant protein or recombinant nucleic acid molecule includes an amino acid or nucleotide sequence that includes at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.
[0305] "Reduce" means a negative change of at least 10%, 25%, 50%, 75%, or 100%.
[0306] "Reference" means a standard or control condition. In one embodiment, the reference is a wild-type or healthy cell. In other embodiments and without limitation, the reference is an untreated cell that has not been subjected to the test condition, or has been subjected to a placebo or a conventional saline solution, culture medium, buffer, and / or a control vector that does not have the target polynucleotide.
[0307] A "reference sequence" is a defined sequence that serves as a benchmark for sequence alignment. The reference sequence can be a subset or all of a designated sequence; for example, a segment of a full-length cDNA or gene sequence, or a complete cDNA or gene sequence. For a polypeptide, the length of the reference polypeptide sequence will generally be at least about 16 amino acids, at least about 20 amino acids, at least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids. For a nucleic acid, the length of the reference nucleic acid sequence will generally be at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides, or about 300 nucleotides, or any integer approximately equal to or between them. In some embodiments, the reference sequence is the wild-type sequence of the target protein. In other embodiments, the reference sequence is the polynucleotide sequence encoding the wild-type protein.
[0308] The terms "RNA-programmable nuclease" and "RNA-guided nuclease" are used (e.g., in combination or association) with one or more RNAs that do not cleave the target. In some embodiments, when an RNA-programmable nuclease forms a complex with an RNA, it may be referred to as a nuclease:RNA complex. Generally, the bound RNA is referred to as a guide RNA (gRNA).
[0309] In some embodiments, the RNA-programmable nuclease is a (CRISPR-associated system) Cas9 endonuclease, e.g., Cas9 from Streptococcus pyogenes (Csnl) (see, e.g., "Complete genome sequence of an Mlstrain of Streptococcus pyogenes." Ferretti J.J. et al., Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and hostfactor RNase III." Deltcheva E. et al., Nature 471:602-607 (2011)).
[0310] Because RNA-programmable nucleases (such as Cas9) use RNA:DNA hybridization to target DNA cleavage sites, these proteins are, in principle, capable of being targeted to any sequence specified by a guide RNA. Methods for using RNA-programmable nucleases (such as Cas9) for site-specific cleavage (e.g., to modify the genome) are known in the art (see, e.g., Cong, L. et al., Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, W.Y. et al., Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); DiCarlo, J.E. et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic acids research (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature biotechnology 31, 233-239 (2013); the entire contents of each of which are incorporated herein by reference).
[0311] The term "single nucleotide polymorphism (SNP)" refers to a variation in a single nucleotide that occurs at a specific location in the genome, where each variation exists in a population to a certain appreciable extent (e.g., >1%). For example, at a specific base position in the human genome, the C nucleotide may be present in most individuals, but in a small number of individuals, this position is occupied by an A. This means that there is an SNP at this specific position, and the two possible nucleotide variations, C or A, are considered to be alleles at this position. SNPs form the basis for differences in disease susceptibility. The severity of a disease and the way the human body responds to treatment are also clinical manifestations of genetic variation. SNPs can fall within the coding region of a gene, the non-coding region of a gene, or be located in the intergenic region (the region between genes). In some embodiments, due to the degeneracy of the genetic code, SNPs within the coding sequence do not necessarily change the amino acid sequence of the protein produced. There are two types of SNPs in the coding region: synonymous and non-synonymous SNPs. Synonymous SNPs do not affect the protein sequence, while non-synonymous SNPs change the amino acid sequence of the protein. Non-synonymous SNPs are of two types: missense and nonsense. SNPs that are not within the protein-coding region can still affect gene splicing, transcription factor binding, messenger RNA degradation, or the sequence of non-coding RNA. Gene expression affected by this type of SNP is called eSNP (expression SNP) and can be upstream or downstream of the gene. A single nucleotide variant (SNV) is a variation in a single nucleotide that is not restricted by frequency and can occur in somatic cells. Somatic single nucleotide variations can also be referred to as single-nucleotide alterations.
[0312] "Specifically binds" means recognizing and binding to the polypeptide and / or nucleic acid molecule of the present invention, but substantially not recognizing and binding to other molecules in the sample (e.g., biological sample), such as nucleic acid molecules, polypeptides, or their complexes (e.g., nucleic acid programmable DNA binding domain and guide nucleic acid), compounds, or molecules.
[0313] Nucleic acid molecules useful in the methods of the invention include any nucleic acid molecule encoding a polypeptide of the invention or a fragment of such polypeptide. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but will generally exhibit substantial identity. Polynucleotides having "substantial identity" to an endogenous sequence will generally be capable of hybridizing to at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the invention include any nucleic acid molecule encoding a polypeptide of the invention or a fragment of such polypeptide. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but will generally exhibit substantial identity. Polynucleotides having "substantial identity" to an endogenous sequence will generally be capable of hybridizing to at least one strand of a double-stranded nucleic acid molecule. "Hybridization" means the pairing and formation of a double-stranded molecule between complementary polynucleotide sequences (such as the genes described herein) or portions thereof under various stringent conditions (see, e.g., Wahl, G.M. and S.L.Berger (1987) Methods Enzymol. 152:399; Kimmel, A.R. (1987) Methods Enzymol. 152:507).
[0314] For example, stringent salt concentrations will generally be less than about 750 mM sodium chloride and 75 mM sodium citrate, preferably less than about 500 mM sodium chloride and 50 mM sodium citrate, and more preferably less than about 250 mM sodium chloride and 25 mM sodium citrate. Low stringency hybridization can be achieved in the absence of organic solvents (such as formamide), while high stringency hybridization can be achieved in the presence of at least about 35% formamide, and more preferably at least about 50% formamide. Stringent temperature conditions will generally include a temperature of at least about 30°C, more preferably at least about 37°C, and most preferably at least about 42°C. Variations in other parameters, such as hybridization time, concentration of detergent (such as sodium dodecyl sulfate (SDS)), and inclusion or exclusion of carrier DNA, are well known to those skilled in the art. Various degrees of stringency can be achieved by combining these different conditions as needed. In one embodiment, hybridization will occur at 30°C, 750 mM sodium chloride, 75 mM sodium citrate, and 1% SDS. In another embodiment, hybridization will occur at 37°C, 500 mM sodium chloride, 50 mM sodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA). In another embodiment, hybridization will occur at 42°C, 250 mM sodium chloride, 25 mM sodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA. Useful variations on these conditions will be apparent to those skilled in the art.
[0315] For most applications, the post-hybridization wash steps will also vary in stringency. The stringent conditions for washing can be defined by salt concentration and by temperature. As noted above, the stringency of the wash can be increased by lowering the salt concentration or by raising the temperature. For example, the stringent salt concentration for a wash step will preferably be less than about 30 mM sodium chloride and 3 mM sodium citrate, and most preferably less than about 15 mM sodium chloride and 1.5 mM sodium citrate. The stringent temperature conditions for a wash step will generally include a temperature of at least about 25° C., more preferably at least about 42° C., and even more preferably at least about 68° C. In one embodiment, the wash step will occur at 25° C. in 30 mM sodium chloride, 3 mM sodium citrate, and 0.1% SDS. In a more preferred embodiment, the wash step will occur at 42° C. in 15 mM sodium chloride, 1.5 mM sodium citrate, and 0.1% SDS. In a more preferred embodiment, the wash step will occur at 68° C. in 15 mM sodium chloride, 1.5 mM sodium citrate, and 0.1% SDS. Additional variations on these conditions will be apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described in, for example: Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.
[0316] "Cleavage" means to be divided into two or more fragments.
[0317] "Cleaved Cas9 protein" or "cleaved Cas9" refers to a Cas9 protein that is provided as an N-terminal fragment and a C-terminal fragment encoded by two separate nucleotide sequences. The polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein can be spliced to form a "reconstituted" Cas9 protein. In certain embodiments, the Cas9 protein is split into two fragments within a disordered region of the protein, e.g., as described in Nishimasu et al., Cell, Vol. 156, No. 5, pp. 935-949, 2014, or as described in Jiang et al. (2016) Science 351:867-871. PDB file: 5F9R, the entire contents of each of which are incorporated herein by reference. In some embodiments, within a region of SpCas9 between approximately amino acids A292-G364, F445-K483, or E565-T637, or at the corresponding positions of any other Cas9, Cas9 variant (e.g., nCas9, dCas9) or other napDNAbp, the protein is split into two fragments at any of C, T, A, or S. In some embodiments, the protein is split into two fragments at T310, T313, A456, S469, or C574 of SpCas9. In some embodiments, the process of splitting the protein into two fragments is referred to as "cleavage process" of the protein.
[0318] In other embodiments, the N-terminal portion of the Cas9 protein comprises amino acids 1-573 or 1-637 of wild-type Streptococcus pyogenes Cas9 (SpCas9) (NCBI reference sequence: NC_002737.2, Uniprot reference sequence: Q99ZW2) and the C-terminal portion of the Cas9 protein comprises a portion of amino acids 574-1368 or 638-1368 of wild-type SpCas9.
[0319] The C-terminal portion of the split Cas9 can be linked to the N-terminal portion of the split Cas9 to form a complete Cas9 protein. In some embodiments, the C-terminal portion of the Cas9 protein starts where the N-terminal portion of the Cas9 protein ends. Thus, in some embodiments, the C-terminal portion of the split Cas9 comprises a portion of amino acids (551-651)-1368 of spCas9. "(551-651)-1368" means starting at an amino acid between amino acids 551-651 (inclusive) and ending at amino acid 1368.For example, the C-terminal portion of the cleaved Cas9 may include a portion of any one of the following amino acids: amino acids 551-1368, 552-1368, 553-1368, 554-1368, 555-1368, 556-1368, 557-1368, 558-1368, 559-1368, 560-1368, 561-1368, 562-1368, 563-1368, 564-1368, 565-1368, 566-1368, 567-1368, 568-1368, 569-1368, 570-1368, 571-1368, 572-1368, 573-1368, 574-1368, 575-1368, 576-1368, 577-1368, 578-1368, 579-1368, 580-1368, 581-1368, 582-1368, 583-1368, 584-1368, 585-1368, 586-1368, 587-1368, 588-1368, 589-1368, 590-1368, 591-1368, 592-1368, 593-1368, 594-1368, 595-1368, 596-1368, 597-1368, 598-1368, 599-1368, 600-1368, 601-1368, 602-1368, 603-1368, 604-1368, 605-1368, 606-1368, 607-1368, 608-1368, 609-1368, 610-1368, 611-1368, 612-1368, 613-1368, 614-1368, 615-1368, 616-1368, 617-1368, 618-1368, 619-1368, 620-1368, 621-1368, 622-1368, 623-1368, 624-1368, 625-1368, 626-1368, 627-1368, 628-1368, 629-1368, 630-1368, 631-1368, 632-1368, 633-1368, 634-1368, 635-1368, 636-1368, 637-1368, 638-1368, 639-1368, 640-1368, 641-1368, 642-1368, 643-1368, 644-1368, 645-1368, 646-1368, 647-1368, 648-1368, 649-1368, 650-1368, or 651-1368 of spCas9.In some embodiments, the C-terminal portion of the cleaved Cas9 protein comprises a portion of amino acids 574-1368 or 638-1368 of SpCas9.
[0320] "Subject" means a mammal, including but not limited to a human or non-human mammal such as a cow, horse, dog, sheep, or cat. Subjects include domestic animals, i.e., domesticated animals raised to produce labor and provide goods (such as food), including but not limited to cows, goats, chickens, horses, pigs, rabbits, and sheep.
[0321] "Substantially identical" means a polypeptide or nucleic acid molecule that exhibits at least 50% identity to a reference amino acid sequence (e.g., any of the amino acid sequences described herein) or nucleic acid sequence (e.g., any of the nucleic acid sequences described herein). In one embodiment, such a sequence is at least 60%, 80%, or 85%, 90%, 95%, or even 99% identical to the sequence being compared at the amino acid level or nucleic acid level.
[0322] Sequence identity is typically measured using sequence analysis software (e.g., the Sequence Analysis Software Package of the Genetics Computer Group, Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning a degree of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include within-group substitutions from the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. In an exemplary method for determining the degree of identity, the BLAST program can be used, where the probability score between e -3 and e -100 indicates closely related sequences.
[0323] For example, COBALT was used with the following parameters:
[0324] a) alignment parameters: Gap penalties -11, -1 and End-Gap penalties -5, -1,
[0325] b) CDD parameters: Use RPS BLAST on; Blast E-value 0.003; FindConserved columns and Recompute on, and
[0326] c) Query Clustering parameters: Use query clusters on; Word Size 4; Max cluster distance 0.8; Alphabet Regular.
[0327] For example, EMBOSS Needle was used with the following parameters:
[0328] a) Matrix: BLOSUM62;
[0329] b) GAP OPEN: 10;
[0330] c) GAP EXTEND: 0.5;
[0331] d) OUTPUT FORMAT: pair;
[0332] e) END GAP PENALTY: false;
[0333] f) END GAP OPEN: 10; and
[0334] g) END GAP EXTEND: 0.5.
[0335] The term "(target) site" refers to a sequence within a nucleic acid molecule that is modified by a base editor. In one embodiment, the target site is deaminated by a deaminase or a fusion protein comprising a deaminase (such as a cytidine or adenine deaminase).
[0336] As used herein, the terms “treat,” “treating,” “treatment,” etc. refer to reducing or ameliorating a disorder and / or symptoms associated therewith or obtaining a desired pharmacological and / or physiological effect. It should be understood that, although not excluded, treatment of a disorder or condition does not require that the disorder, condition, or symptoms associated therewith be completely eliminated. In some embodiments, the foregoing effect is therapeutic, i.e., but not limited to, the effect partially or completely reduces, lessens, suppresses, alleviates, mitigates, or decreases the intensity of the disease and / or the adverse symptoms attributable to the disease, or cures the disease. In some embodiments, the effect is prophylactic, i.e., the effect protects or prevents the occurrence or recurrence of a disease or condition. To this end, the methods disclosed herein include administering a therapeutically effective amount of a composition as described herein.
[0337] “Uracil glycosylase inhibitor” or “UGI” means an agent that inhibits the uracil-excision repair system. In one embodiment, the agent is a protein or fragment thereof that binds to the host uracil-DNA glycosylase and prevents the removal of uracil from DNA. In one embodiment, UGI is a protein, fragment, or domain thereof that is capable of inhibiting the uracil-DNA glycosylase base-excision repair enzyme. In some embodiments, the UGI domain includes wild-type UGI or a modified version thereof. In some embodiments, the UGI domain includes a fragment of the exemplary amino acid sequences listed below. In some embodiments, the amino acid sequence included in the UGI fragment includes at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the exemplary UGI sequence provided below. In some embodiments, UGI includes an amino acid sequence homologous to the exemplary UGI amino acid sequence or a fragment thereof (as listed below). In some embodiments, the UGI or a portion thereof is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% or 100% identical to wild-type UGI or the UGI sequence or a portion thereof as listed below. Exemplary UGIs include the following amino acid sequences: >splP14739IUNGI_BPPB2 uracil-DNA glycosylase inhibitor
[0338] MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML.
[0339] The term "vector" refers to a tool for introducing a nucleic acid sequence into a cell to obtain a transformed cell. Vectors include plasmids, transposons, bacteriophages, viruses, liposomes, and episomes. An "expression vector" is a nucleic acid sequence that includes a nucleotide sequence to be expressed in a recipient cell. An expression vector may contain additional nucleic acid sequences to enhance and / or facilitate the expression of the introduced sequence, such as initiation, termination, enhancer, promoter, and secretion sequences.
[0340] Any composition or method provided herein can be combined with one or more of any other compositions and methods provided herein.
[0341] By correcting pathogenic mutations at the gene level, DNA editing has become a viable means of modulating disease states. Until recently, the utility of all DNA editing platforms has been through inducing DNA double-strand breaks (DSBs) at specific genomic loci and relying on endogenous DNA repair pathways to determine product outcomes in a semi-random manner, resulting in complex populations of gene products. Although precise, user-defined repair outcomes can be achieved through the homologous directed repair (HDR) pathway, several challenges have hindered the efficient use of HDR for repair in therapeutically relevant cell types. In practice, this pathway is inefficient relative to its competitor, the error-prone non-homologous end joining pathway. Additionally, HDR is strictly restricted to the G1 and S phases of the cell cycle, precluding precise repair of DSBs in post-mitotic cells. Thus, it has proven difficult or impossible to efficiently alter genomic sequences in a user-defined, programmable manner in these populations. BRIEF DESCRIPTION OF THE DRAWINGS
[0342] Figure 1A - 1C Depicts the cis-trans activity of a free deaminase. Figure 1A The schematic diagram of... describes the experimental design of the cis-trans assay of a base editor complex or SpCas9 and deaminase in an untethered form. Figure 1B Depicts the cis-trans activity of rAPOBEC. Figure 1C Depicts the cis-trans activity of TadA7.10 and TadA-TadA7.10.
[0343] Figure 2A - 2F Depicts the cis-trans assay of a base editor, the illustration of a deaminase similarity network, and the screening of 153 deaminases. Figure 2A The schematic diagram of... depicts the experimental design of the cis-trans assay. HEK293T cells were transfected with separate plasmids encoding SaCas9, gRNA for SaCas9, and a targeted base editor. Figure 2BThe schematic diagram depicts the similarity network of APOBEC-like deaminases. Each point represents a cytidine deaminase screened as a next-generation CBE, and the core next-generation CBEs are indicated. The shading of the points represents the average trans / cis ratio; the size of the points represents the average cis activity. Creation Figure 2B The method for creating the cytidine deaminase similarity network shown above is as follows: To focus the search space within the APOBEC1-like protein family, human APOBEC1 was used as the query sequence for a protein BLAST search against the NCBI non-redundant protein sequence database (nr_v5). The top 1000 sequences were used to generate a sequence similarity network (SSN) with a protein BLAST -log(E-value) edge threshold of 115. A set of 43 deaminases was selected to sample the sequence space within the SSN. To identify deaminases from other families that could act as base-editing enzymes, 80 sequences were sampled from the SSN constructed with all deaminases and the following InterPro annotations were used: IPR002125 (cytidine and deoxycytidylate deaminase domain), IPR016192 (APOBEC / CMP deaminase, zinc-binding), and IPR016193 (cytidine deaminase-like). The 82,043 sequences of this set were first clustered using Cd-HIT 3 with 55% identity, and then an SSN network was generated by protein BLAST with a -log(E-value) edge threshold of 50. Sequences were selected based on their centrality within the cluster. Figure 2C is an S-chart depicting the cis-trans activities of ppBE4 and its mutants. Figure 2D is a chart depicting the cis-trans activities of the selected editors. Individually, cis-trans-activity data were generated based on cis / trans assays at three target sites, namely site 1, site 4, and site 6, as Figure 2E and Figure 2F shown. Figure 2EA bar graph is presented that shows the cis and trans editing activities of the identified CBEs. Shown is a comparison of the cis and trans editing frequencies in mammalian cells treated with candidate CBEs. Editors numbered 1-36 are base editors: pYY-BEM3.8, pYY-BEM3.9, pYY-BEM3.10, pYY-BEM3.11, pYY-BEM3.12, pYY-BEM3.13, pYY-BEM3.14, pYY-BEM3.15, pYY-BEM3.16, pYY-BEM3.17, pYY-BEM3.18, pYY-BEM3.19, pYY-BEM3.20, pYY-BEM3.21, pYY-BEM3.22, pYY-BEM3.23, pYY-BEM3.24, pYY-BEM3.25, pYY-BEM3.26, pYY-BEM3.27, pYY-BEM3.28, pYY-BEM3.29, pYY-BEM3.30, pYY-BEM3.31, pYY-BEM3.32, pYY-BEM3.33, pYY-BEM3.34, pYY-BEM3.35, pYY-BEM3.36, pYY-BEM3.37, pYY-BEM3.38, pYY-BEM3.39, pYY-BEM3.40, pYY-BEM3.41, pYY-BEM3.42, pYY-BEM3.43. The reported base editing efficiency is for the most edited base in the target site. Figure 2FPresents a bar graph that shows the cis- and trans-editing activities of the identified CBEs. Shown is a comparison of the cis- and trans-editing frequencies in mammalian cells treated with candidate CBEs. Editors numbered 1-37 are: rBE4max, mAPOBEC-1, MaAPOBEC-1, hAPOBEC-1, ppAPOBEC-1, OcAPOBEC1, MdAPOBEC-1, mAPOBEC-2, hAPOBEC-2, ppAPOBEC-2, BtAPOBEC-2, mAPOBEC-3, hAPOBEC-3A, hAPOBEC-3B, hAPOBEC-3C, hAPOBEC-3D, hAPOBEC-3F, hAPOBEC-3G, hAPOBEC-4, mAPOBEC-4, rAPOBEC-4, MfAPOBEC-4, hAID, negative control, btAID, mAID, pmCDA-1, pmCDA-2, pmCDA-5, yCD, pYY-BEM3.1, pYY-BEM3.2, pYY-BEM3.3, pYY-BEM3.4, pYY-BEM3.5, pYY-BEM3.6, pYY-BEM3.7. The reported base editing efficiency is for the most edited base in the target site.
[0344] Figure 3A and 3B Depicts cis-trans activity. Figure 3A Is a graph depicting the cis-trans activity of ABE7.10. Figure 3B Is a graph depicting the cis-trans activity of BE4max.
[0345] Figure 4A and 4B Depicts a homology model of rAPOBEC1 generated by SWISSMODEL using the hAPOBEC3C structure (PDB ID 3VM8). The ssDNA from the hAPOBEC3A structure (PDB ID 5SWW) was manually docked. Figure 4A Is a schematic diagram depicting mutations that have the potential to affect ssDNA binding. Figure 4B Is a schematic diagram depicting mutations that have the potential to affect catalytic activity.
[0346] Figure 5A - 5C Depicts the cis-trans activity of rAPOBEC1 mutants.
[0347] Figure 6A - 6E Depicts the cis-trans activity of rAPOBEC1 double mutants. Figure 6A Is a graph depicting the cis and trans activities of rAPOBEC1 double mutants. Figure 6BIt is a graph depicting the cis-activity at 6 sites. Figure 6C It is a graph depicting the cis / trans activity. Figure 6D It is a graph depicting the cis-activity at 5 sites. Figure 6E It is a graph depicting the cis / trans activity.
[0348] Figure 7A and 7B Depicts the cis-trans activity of the deaminase in the first round of screening.
[0349] Figure 8A - 8C It is a graph depicting the on-target activity of ppAPOBEC1 relative to rAPOBEC1.
[0350] Figure 9 It is a schematic diagram depicting the similarity network of APOBEC-like proteins.
[0351] Figure 10A and 10B They are graphs respectively depicting the dose-dependent studies of cis-activity and trans-activity in TadA-TadA7.10 and rAPOBEC1.
[0352] Figure 11 It is a graph depicting the off-target editing of the selected CBE. SNVs were identified by whole-exome sequencing.
[0353] Figure 12A and 12B They are graphs respectively depicting the quantification of base editor mRNA and protein from HEK293T cells transfected with base editor plasmids.
[0354] Figure 13 It is a graph depicting the targeted RNA sequencing of the selected editors. Three regions of 200 - 300 bp were sequenced.
[0355] Figure 14 It is a graph depicting the guide-dependent off-target editing of the selected CBE.
[0356] Figure 15A - 15E Depicts the editing window of the selected editor.
[0357] Figure 16 It is a graph depicting the insertion / deletion rates of the selected CBE at 10 target sites.
[0358] Figure 17A - 17D Shows the schematic diagrams and graphs related to the non-guide ssDNA deamination reaction and cis / trans assays. Figure 17A Illustrates the formation of potential ssDNA in the genome during transcription or translation. Figure 17BIllustrated the experimental design for cis / trans determination. HEK293T cells were transfected with separate constructs encoding SaCas9, gRNA for SaCas9, and base editors. Cis and trans activities were measured at the target site (with NGGRRT PAM sequence) in different transfections. Figure 17C Showed the cis / trans activities of BE4 with rAPOBEC1. Figure 17D Showed ABE7.10 variants at 34 genomic loci. The leftmost bar for each genomic locus on the x-axis represents cis on-target editing. The rightmost bar for each genomic locus on the x-axis represents trans editing. The reported base editing efficiency is for the most edited base in the target site. Values and error bars reflect the mean and standard deviation (s.d.) of independent biological replicates.
[0359] Figure 18 The presented bar graph showed next-generation CBEs identified with high cis activity and reduced trans activity compared to BE4 with rAPOBEC1. Shown was the comparison of cis and trans editing frequencies at 10 genomic loci in mammalian cells treated with the next-generation CBEs (BE4 with PpAPOBEC1[wt,H122], RrA3F[wt,F130L], AmAPOBEC1, SsAPOBEC2[wt,R54Q]). The reported base editing efficiency is for the most edited base in the target site. Values and error bars reflect the mean and standard deviation of 4 independent biological replicates.
[0360] Figure 19A - 19E Showed allele frequencies and graphs related to next-generation CBEs that have reduced DNA and RNA off-target editing in mammalian cells relative to BE4. Figure 19A Showed whole-transcriptome sequencing and target RNA sequencing of Hek293T cells expressing cytosine base editors with minimized pseudodeamination ( Figure 19B ). Figure 19C Showed the percentage of C-to-T editing at known guide off-target sites. Figure 19D Showed the percentage of C-to-T editing in in vitro enzymatic assays on single-stranded DNA substrates. C-to-U editing was assayed for the core next-generation CBEs on ssDNA substrates. Each point represents the edited N C Local sequence context. The black line represents the average editing efficiency of the target cytosine in the substrate. Figure 19E Presented the time course of product formation in in vitro enzymatic assays with cell lysates containing the selected CBEs. Figure 19D and 19EThe oligonucleotide sequences used are listed in the table in Example 5 below. The values and error bars reflect the means and standard deviations of 3 independent biological replicates ( Figure 19A , B, C) and 2 independent biological replicates ( Figure 19D , E).
[0361] Figure 20 Graphically depicts the cis / trans editing activities of BE4 with rAPOBEC1 mutants as shown in Figure 4A and 4B at target site 1. The reported base editing efficiencies are for the most edited base in the target site. The trans efficiency is represented by the leftmost bar for each target site on the x-axis; the cis efficiency is represented by the bar on the right of each target site on the x-axis. The values and error bars reflect the means and standard deviations of independent biological replicates.
[0362] Figure 21 Depicts the cis / trans editing activities of BE4-rAPOBEC1 with HiFi mutations at 10 target sites. The values and error bars reflect the means and standard deviations of 4 independent biological replicates.
[0363] Figure 22A and 22B Show graphs and sequence alignments related to the cis / trans editing activities and sequence alignments of the CBEs tested in the previous first-round screening. Shown are the cis / trans editing activities ( Figure 22A ) and sequence alignments ( Figure 22B ) of the selected CBEs at site 10. The amino acid residues aligned to the HiFi mutations in rAPOBEC1 are highlighted. The values and error bars reflect the means and standard deviations of independent biological replicates.
[0364] Figure 23 Demonstrate the cis / trans activities of BE4-PpAPOBEC1 and BE4-PpAPOBEC with HiFi mutations at 10 target sites. The reported base editing efficiencies are for the most edited base in the target site. The values and error bars reflect the means and standard deviations of 4 independent biological replicates.
[0365] Figure 24 Show a heatmap that indicates the Figure 18 previous base preferences of the CBEs shown in
[0366] Figure 25 The values used to generate the heatmap reflect the means of 4 independent biological replicates. Figure 18 at 10 target sitesEditing windows of the CBEs shown in []. The values represent the average of 4 independent biological replicates. Cis- and trans-editing are presented in the leftmost and rightmost panel heatmaps, respectively.
[0367] Figure 26 A table is presented, which shows the Figure 18 insertion / deletion rates of the CBEs shown in []. The values used to generate this heatmap represent the average of 4 independent biological replicates.
[0368] Figure 27A - 27D Homology models of four selected cytidine deaminases based on existing crystal structures are depicted. Figure 27A : The homology model of PpAPOBEC1 is based on the putative APOBEC3G structure (PDB ID 5K81). Figure 27B : RrA3F is based on the Vif-binding domain of hAPOBEC3F (PDB ID 3WUS). Figure 27C : AmAPOBEC1 is based on the N-terminal domain of hAPOBEC3B (PDB ID 5TKM). Figure 27D : SsAPOBEC2 is based on the Vif-binding domain of hAPOBEC3F (PDB ID 3WUS).
[0369] Figure 28A - 28D A chart is presented, which illustrates the off-target editing of the selected next-generation CBEs. Figure 28A : The editing efficiency of the next-generation CBEs at the HEK2, HEK3, and HEK4 sites, and Figure 28B and 28C : The reported off-target sites of the HEK2 sgRNA and HEK3 sgRNA, and Figure 28D : The HEK4 sgRNA. The reported base editing efficiency is for the most edited base in the target site. The values and error bars represent the average and standard deviation of 3 independent biological replicates.
[0370] Figure 29 The presented chart shows the C-to-T editing efficiency of the selected CBEs on ssDNA substrates in in vitro enzymatic assays. The editing efficiency was measured at all 25 cytidines on 2 DNA substrates and grouped by N C sequence context. The sequences of the two substrates used are listed in Table 18 herein. The values and error bars represent the average and standard deviation of data from independent biological replicates.
[0371] Figure 30The presented graph shows the quantitative analysis of CBE protein concentration in HEK293T cells transfected with base editor expression plasmids. The base editor protein concentration was quantified by measuring the total Cas9 protein concentration and the total protein amount in cell lysates. The BE protein concentration was normalized relative to BE4-rAPOBEC1. The values and error bars reflect the mean and standard deviation from two or more independent biological replicates.
[0372] Figure 31 The presented graph shows the pseudo-deamination reaction activity of CBEs examined by whole genome sequencing (WGS). The relative mutation rate is shown as odds ratio. Detailed Description
[0373] The present invention provides nucleobase editors and multi-effector nucleobase editors with improved editing settings (i.e., minimal off-target deamination reaction), compositions comprising such editors, and methods of using them to generate modifications in target nucleobase sequences.
[0374] Nucleobase Editor
[0375] Disclosed herein are base editors or nucleobase editors or multi-effector nucleobase editors for editing, modifying or altering a target nucleotide sequence of a polynucleotide. Described herein are nucleobase editors or base editors or multi-effector nucleobase editors comprising a polynucleotide-programmable nucleotide binding domain (e.g., Cas9) and at least one nucleobase editing domain (e.g., adenosine deaminase and / or cytidine deaminase). The polynucleotide-programmable nucleotide binding domain (e.g., Cas9), when linked to a bound guide polynucleotide (e.g., gRNA), can specifically bind to a target nucleotide sequence (i.e., via complementary base pairing between the bound guide nucleic acid and the target nucleotide sequence bases), thereby localizing the base editor to the target nucleic acid sequence to be edited.
[0376] Polynucleotide-Programmable Nucleotide Binding Domain
[0377] It should be understood that the polynucleotide-programmable nucleotide binding domain can also comprise an RNA-binding nucleic acid-programmable protein. For example, the polynucleotide-programmable nucleotide binding domain can be linked to a nucleic acid that directs the polynucleotide-programmable nucleotide binding domain to RNA. Other nucleic acid-programmable DNA-binding proteins are also within the scope of the present disclosure, although they are not specifically listed in the present disclosure.
[0378] The polynucleotide programmable nucleotide binding domain of a base editor can itself include one or more domains. For example, the polynucleotide programmable nucleotide binding domain can include one or more nuclease domains. In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain can include an endonuclease or an exonuclease. The term "exonuclease" as used herein refers to a protein or polypeptide capable of digesting nucleic acids (e.g., RNA or DNA) from the ends, while the term "endonuclease" refers to a protein or polypeptide capable of catalyzing (e.g., cleaving) an internal region in a nucleic acid (e.g., DNA or RNA). In some embodiments, the endonuclease can cleave a single strand of a double-stranded nucleic acid. In some embodiments, the endonuclease can cleave both strands of a double-stranded nucleic acid molecule. In some embodiments, the polynucleotide programmable nucleotide binding domain can be a deoxyribonuclease. In some embodiments, the polynucleotide programmable nucleotide binding domain can be a ribonuclease.
[0379] In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain can cleave zero, one, or both strands of a target polynucleotide. In some embodiments, the polynucleotide programmable nucleotide binding domain can include a nickase domain. The term "nickase" as used herein refers to a polynucleotide programmable nucleotide binding domain that includes a nuclease domain that is capable of cleaving only one of the two strands in a double-stranded helical nucleic acid molecule (e.g., DNA). In some embodiments, the nickase can be derived from a fully catalytically active (e.g., native) form of the polynucleotide programmable nucleotide binding domain by introducing one or more mutations into the active polynucleotide programmable nucleotide binding domain. For example, when the polynucleotide programmable nucleotide binding domain includes a nickase domain derived from Cas9, the Cas9-derived nickase domain can contain a D10A mutation and a histidine at position 840. In this embodiment, the residue H840 retains catalytic activity and can thus cleave a single strand of the nucleic acid double helix. In another example, the Cas9-derived nickase domain can include an H840A mutation, while the amino acid residue at position 10 remains D. In some embodiments, the nickase can be derived from a fully catalytically active (e.g., native) form of the polynucleotide programmable nucleotide binding domain by removing all or part of the nuclease domain that is not essential for nickase activity. For example, when the polynucleotide programmable nucleotide binding domain includes a nickase domain derived from Cas9, the Cas9-derived nickase domain can include a deletion of all or part of the RuvC domain or the HNH domain.
[0380] An exemplary amino acid sequence of catalytically active Cas9 is as follows:
[0381] MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMT
[0382] NFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD.
[0383] Base editors that include a polynucleotide programmable nucleotide binding domain (which includes a nickase domain) are thus capable of generating a single-stranded DNA break (nick) at a specific polynucleotide target sequence (e.g., as determined by the complementary sequence of the bound guide nucleic acid). In some embodiments, the strand of the nucleic acid double helix target polynucleotide sequence that is nicked by the base editor that includes a nickase domain (e.g., a Cas9-derived nickase domain) is the strand that is not edited by the base editor (i.e., the strand that is nicked by the base editor is on the opposite side of the strand that includes the base to be edited). In other embodiments, a base editor that includes a nickase domain (e.g., a Cas9-derived nickase domain) can nick the strand of DNA that is to be targeted for editing. In such embodiments, the non-targeted strand is not nicked.
[0384] Also provided herein are base editors that include a catalytically inactive (i.e., unable to nick a target polynucleotide sequence) polynucleotide programmable nucleotide binding domain. The terms “catalytically inactive” and “nuclease-inactive” are used interchangeably herein to refer to a polynucleotide programmable nucleotide binding domain that has one or more mutations and / or deletions that result in its inability to nick a nucleic acid strand. In some embodiments, a catalytically inactive polynucleotide programmable nucleotide binding domain base editor may lack nuclease activity due to specific point mutations in one or more nuclease domains. For example, in the case of a base editor that includes a Cas9 domain, the Cas9 can include two mutations, namely, the D10A mutation and the H840A mutation. Such mutations inactivate the two nuclease domains, resulting in the loss of nuclease activity. In other embodiments, a catalytically inactive polynucleotide programmable nucleotide binding domain can include one or more deletions of all or part of a catalytic domain (e.g., the RuvC1 and / or HNH domains). In further embodiments, a catalytically inactive polynucleotide programmable nucleotide binding domain includes a point mutation (e.g., D10A or H840A) and a deletion of all or part of a nuclease domain.
[0385] This document also examines mutations that can generate a catalytically inactive polynucleotide-programmable nucleotide-binding domain from a previous functional version of the polynucleotide-programmable nucleotide-binding domain. For example, in the case of catalytically inactive Cas9 (“dCas9”), variants are provided that have mutations other than D10A and H840A (which result in inactivated Cas9). Such mutations include, for example, other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions within the HNH nuclease subdomain and / or the RuvC1 subdomain). Based on the disclosure herein and knowledge in the art, additional suitable nuclease-inactivated dCas9 domains may be apparent to those skilled in the art and are within the scope of the present disclosure. Such additional exemplary suitable nuclease-inactivated Cas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, e.g., Prashant et al., CAS9 transcriptional activators for targetspecificity screening and paired nickases for cooperative genomeengineering. Nature Biotechnology. 2013;31(9):833-838, the entire contents of which are incorporated herein by reference).
[0386] Non-limiting examples of polynucleotide programmable nucleotide binding domains that can be incorporated into base editors include CRISPR protein-derived domains, restriction endonucleases, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs). In some embodiments, a base editor includes a polynucleotide programmable nucleotide binding domain that includes a native or modified protein or portion thereof that, via a bound guide nucleic acid, is capable of binding to a nucleic acid sequence during CRISPR (i.e., clustered regularly interspaced short palindromic repeats)-mediated nucleic acid modification. This protein is referred to herein as a "CRISPR protein." Accordingly, base editors that include a polynucleotide programmable nucleotide binding domain are disclosed herein, where the polynucleotide programmable nucleotide binding domain includes all or a portion of a CRISPR protein (i.e., a base editor that includes all or a portion of a CRISPR protein as a domain, the domain is also referred to as the "CRISPR protein-derived domain" of the base editor). The CRISPR protein-derived domain incorporated into the base editor can be modified compared to the wild-type or native version of the CRISPR protein. For example, the CRISPR protein-derived domain can include one or more mutations, insertions, deletions, rearrangements, and / or recombinations relative to the wild-type or native version of the CRISPR protein as described below.
[0387] CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposons, and conjugative plasmids). CRISPR clusters contain spacer sequences, i.e., sequences complementary to previous mobile elements, and target invading nucleic acids. The CRISPR clusters are transcribed and processed into CRISPR RNAs (crRNAs). In type II CRISPR systems, correct processing of pre-crRNA requires trans-encoded small RNAs (tracrRNAs), endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA serves as a guide for the ribonuclease 3-assisted pre-crRNA processing. Subsequently, Cas9 / crRNA / tracrRNA cleaves linear or circular dsDNA targets complementary to the spacer sequence by endonucleolytic cleavage. The target strand that is not complementary to the crRNA is first cleaved by endonucleolytic cleavage and then trimmed 3′-5′ by exonucleolytic cleavage. In nature, DNA-binding and cleavage generally require a protein and two RNAs. However, single guide RNAs (“sgRNAs”, or simply “gRNAs”) can be engineered to incorporate aspects of both crRNA and tracrRNA into a single RNA species. See, e.g., Jinek M. et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes short motifs (PAMs or protospacer adjacent motifs) in the CRISPR repeat sequences to help distinguish self from non-self.
[0388] In some embodiments, the methods described herein can utilize engineered Cas proteins. A guide RNA (gRNA) is a short synthetic RNA that consists of a scaffold sequence required for Cas-binding and a user-defined ~20 nucleotide spacer sequence that defines the genomic target to be modified. Thus, one of skill in the art can alter the genomic target specificity of the Cas protein, which depends in part on how specific the gRNA targeting sequence is for the genomic target compared to the rest of the genome.
[0389] In some embodiments, the gRNA scaffold sequence is as follows:
[0390] In some embodiments, the gRNA scaffold sequence is as follows: GUUUUAGAGC UAGAAAUAGC AAGUUAAAAUAAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU.
[0391] In some embodiments, the CRISPR protein-derived domain incorporated within the base editor is an endonuclease (e.g., a deoxyribonuclease or ribonuclease) that is capable of binding to a target polynucleotide sequence when linked to the bound guide nucleic acid. In some embodiments, the CRISPR protein-derived domain incorporated within the base editor is a nickase that is capable of binding to a target polynucleotide sequence when linked to the bound guide nucleic acid. In some embodiments, the CRISPR protein-derived domain incorporated within the base editor is a catalytically inactivated domain that is capable of binding to a target polynucleotide sequence when linked to the bound guide nucleic acid. In some embodiments, the target polynucleotide bound by the CRISPR protein-derived domain of the base editor is DNA. In some embodiments, the target polynucleotide bound by the CRISPR protein-derived domain of the base editor is RNA.
[0392] Cas proteins useful in the present disclosure include Class 1 and Class 2. Non-limiting examples of Cas proteins include: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i, CARF, DinG, their homologs, or modified versions thereof. Unmodified CRISPR enzymes can have DNA cleavage activity, such as Cas9, which has two functional endonuclease domains: RuvC and HNH. The CRISPR enzyme can direct cleavage of one or both strands at the target sequence, such as within the target sequence and / or within the complementary (sequence) of the target sequence. For example, the CRISPR enzyme can direct cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence.
[0393] Vectors encoding CRISPR enzymes can be used, where the CRISPR enzyme (relative to the corresponding wild-type enzyme) is mutated such that the mutated CRISPR enzyme lacks the ability to cleave one or both strands of a target polynucleotide containing the target sequence. Cas9 can refer to a polypeptide having at least or at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology relative to a wild-type exemplary Cas9 polypeptide (e.g., Cas9 from Streptococcus pyogenes). Cas9 can refer to a polypeptide having at most or at most about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology relative to a wild-type exemplary Cas9 polypeptide (e.g., Cas9 from Streptococcus pyogenes). Cas9 can refer to the wild-type or modified form of the Cas9 protein, which includes amino acid changes such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof.
[0394] In some embodiments, the CRISPR protein-derived domain of the base editor can include all or part of Cas9, which is from: Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense, China (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquis (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref: YP_820832.1); Listeria innocua (NCBI Ref: NP_472073.1); Campylobacter jejuni (NCBI Ref: YP_002344900.1); Neisseria meningitidis (NCBI Ref: YP_002342100.1), Streptococcus pyogenes, or Staphylococcus aureus.
[0395] Cas9 domain of the nucleobase editor
[0396] The Cas9 nuclease sequence and structure are well-known to those skilled in the art (see, for example, “Complete genome sequence of an Ml strain of Streptococcus pyogenes.” Ferretti et al., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E. et al., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M. et al., Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference). Orthologs of Cas9 have been described in various species, including but not limited to Streptococcus pyogenes and Streptococcus thermophilus. Based on the present disclosure, additional suitable Cas9 nucleases and sequences will be apparent to those skilled in the art, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference).
[0397] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is a Cas9 domain. Non-limiting exemplary Cas9 domains are provided herein. The Cas9 domain may be a nuclease-active Cas9 domain, a nuclease-inactivated Cas9 domain (dCas9), or a Cas9 nickase (nCas9). In some embodiments, the Cas9 domain is a nuclease-active domain. For example, the Cas9 domain may be a Cas9 domain that cuts both strands of a double-stranded nucleic acid (e.g., both strands of a double-stranded DNA molecule). In some embodiments, the Cas9 domain comprises any of the amino acid sequences listed herein. In some embodiments, the amino acid sequence comprised by the Cas9 domain is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the amino acid sequences listed herein. In some embodiments, the amino acid sequence comprised by the Cas9 domain has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences listed herein. In some embodiments, the amino acid sequence comprised by the Cas9 domain has at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical consecutive amino acid residues compared to any of the amino acid sequences listed herein.
[0398] In some embodiments, a protein comprising a fragment of Cas9 is provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) the gRNA-binding domain of Cas9; or (2) the DNA-cleaving domain of Cas9. In some embodiments, a protein comprising Cas9 or a fragment thereof is referred to as a "Cas9 variant". The Cas9 variant shares homology with Cas9 or a fragment thereof. For example, the Cas9 variant is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to wild-type Cas9. In some embodiments, compared to wild-type Cas9, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes. In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., the gRNA-binding domain or the DNA-cleaving domain) such that the fragment is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to the corresponding wild-type Cas9 fragment. In some embodiments, the fragment is at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 99.5% of the amino acid length of the corresponding wild-type Cas9. In some embodiments, the length of the fragment is at least 100 amino acids. In some embodiments, the length of the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids.
[0399] In some embodiments, the Cas9 fusion proteins provided herein include the full-length amino acid sequence of the Cas9 protein, such as one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein do not include the full-length Cas9 sequence, but only include one or more of its fragments. Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and additional suitable sequences of Cas9 domains and fragments will be apparent to those skilled in the art.
[0400] The Cas9 protein can be associated with a guide RNA that directs the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the polynucleotide programmable nucleotide binding domain is a Cas9 domain, such as, for example, nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Examples of nucleic acid programmable DNA binding proteins include, but are not limited to, Cas9 (such as dCas9 and nCas9), CasX, CasY, Cpf1, Cas12b / C2C1, and Cas12c / C2C3.
[0401] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1, nucleotide and amino acid sequences are as follows).
[0402]
[0403]
[0404] AAGCTTATTGCTCGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGCTAAGGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCTTTGAAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAATATAGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGCTCTGCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAACAAAAACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTTATTTTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGAAAATATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTAAACGATATACGTCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATTGATTTGAGTCAGCTAGGAGGTGACTGA
[0405]
[0406] (Single underline: HNH domain; double underline: RuvC domain)
[0407] In some embodiments, wild-type Cas9 corresponds to or comprises the following nucleotide and / or amino acid sequences:
[0408]
[0409]
[0410] TAGCTGATGCCAATCTGGACAAAGTATTAAGCGCATACAACAAGCACAGGGATAAACCCATACGTGAGCAGGCGGAAAATATTATCCATTTGTTTACTCTTACCAACCTCGGCGCTCCAGCCGCATTCAAGTATTTTGACACAACGATAGATCGCAAACGATACACTTCTACCAAGGAGGTGCTAGACGCGACACTGATTCACCAATCCATCACGGGATTATATGAAACTCGGATAGATTTGTCACAGCTTGGGGGTGACGGATCCCCCAAGAAGAAGAGGAAAGTCTCGAGCGACTACAAAGACCATGACGGTGATTATAAAGATCATGACATCGATTACAAGGATGACGATGACAAGGCTGCAGGA
[0411]
[0412] (Single underline: HNH domain; double underline: RuvC domain).
[0413] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_002737.2) (nucleotide sequence as follows); and Uniprot reference sequence: Q99ZW2 (amino acid sequence as follows)
[0414]
[0415]
[0416]
[0417]
[0418] (Single underline: HNH domain; double underline: RuvC domain)
[0419] In some embodiments, Cas9 refers to Cas9 from the following: Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheriae (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma aphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychrobacter arcticus 273-4 (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Ref: YP_820832.1), Listeria innocua (NCBI Ref: NP_472073.1), Campylobacter jejuni (NCBI Ref: YP_002344900.1) or Neisseria meningitidis (NCBI Ref: YP_002342100.1) or Cas9 from any other organism.
[0420] It should be understood that additional Cas9 proteins (e.g., nuclease-inactivated Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9), including variants and homologs thereof, are within the scope of the present disclosure. Exemplary Cas9 proteins include but are not limited to those provided below. In some embodiments, the Cas9 protein is nuclease-inactivated Cas9 (dCas9). In some embodiments, the Cas9 protein is Cas9 nickase (nCas9). In some embodiments, the Cas9 protein is nuclease-active Cas9.
[0421] In some embodiments, the Cas9 domain is a nuclease-inactivated Cas9 domain (dCas9). For example, the dCas9 domain can bind to a double-stranded nucleic acid molecule (e.g., via a gRNA molecule) without cleaving either strand of the double-stranded nucleic acid molecule. In some embodiments, the nuclease-inactivated dCas9 domain comprises the D10X mutation and the H840X mutation of the amino acid sequences listed herein, or the corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid change. In some embodiments, the nuclease-inactivated dCas9 domain comprises the D10A mutation and the H840A mutation of the amino acid sequences listed herein, or the corresponding mutations in any of the amino acid sequences provided herein. As an example, the nuclease-inactivated Cas9 domain comprises the amino acid sequence listed in the cloning vector pPlatTET-gRNA2 (accession number BAV54124)
[0422] The amino acid sequence of an exemplary catalytically inactive Cas9 (dCas9) is as follows:
[0423]
[0424] SEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD
[0425] See, for example, Qi, et al., “Repurposing CRISPR as an RNA - guided platform for sequence - specific control of gene expression,” Cell. 2013; 152(5):1173 - 83, the entire content of which is incorporated herein by reference).
[0426] Based on the present disclosure and knowledge in the art, additional suitable nuclease - inactivated dCas9 domains will be apparent to those skilled in the art and are within the scope of the present disclosure. Such additional exemplary suitable nuclease - inactivated Cas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, for example, Prashant, et al., “CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering.” Nature Biotechnology. 2013; 31(9):833 - 838, the entire content of which is incorporated herein by reference)
[0427] In some embodiments, the Cas9 nuclease has an inactivated (e.g., dead) DNA cleavage domain, i.e., the Cas9 is a nickase and is referred to as the "nCas9" protein (for "nickase" Cas9). The nuclease-inactivated Cas9 protein may be interchangeably referred to as the "dCas9" protein (for nuclease-"dead" Cas9) or catalytically inactivated Cas9. Methods for generating Cas9 proteins (or fragments thereof) with inactivated DNA cleavage domains are known (see, e.g., Jinek et al., Science. 337:816-821 (2012); Qi et al., "Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression" (2013) Cell. 28; 152(5):1173-83, the entire contents of each of which are incorporated herein by reference). For example, it is known that the DNA cleavage domain of Cas9 contains two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of Streptococcus pyogenes Cas9 (Jinek et al., Science. 337:816-821 (2012); Qi et al., Cell. 28; 152(5):1173-83 (2013)).
[0428] In some embodiments, the amino acid sequence included in the dCas9 domain is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical compared to any of the dCas9 domains provided herein. In some embodiments, the amino acid sequence included in the Cas9 domain has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences listed herein. In some embodiments, the amino acid sequence included in the Cas9 domain has at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, or at least 1200 identical consecutive amino acid residues compared to any of the amino acid sequences listed herein.
[0429] In some embodiments, dCas9 corresponds to or includes a part or all of the Cas9 amino acid sequence that has one or more mutations that inactivate Cas9 nuclease activity. For example, in some embodiments, the dCas9 domain includes the D10A and H840A mutations, or the corresponding mutations in another Cas9.
[0430] In some embodiments, the dCas9 includes the amino acid sequence of dCas9(D10A and H840A):
[0431]
[0432] (Single underline: HNH domain; double underline: RuvC domain).
[0433] In some embodiments, the Cas9 domain includes the D10A mutation, and the residue at position 840 in the amino acid sequence provided above, or the residue at the corresponding position in any of the amino acid sequences provided herein, remains histidine.
[0434] In other embodiments, dCas9 variants having mutations other than D10A and H840A are provided, which mutations, for example, result in nuclease-deactivated Cas9 (dCas9). Such mutations include, for example, other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., substitutions within the HNH nuclease subdomain and / or the RuvC1 subdomain). In some embodiments, variants or homologs of dCas9 are provided that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical. In some embodiments, variants of dCas9 are provided that are shorter or longer in amino acid sequence by about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids or more.
[0435] In some embodiments, the Cas9 domain is a Cas9 nickase. The Cas9 nickase can be a Cas9 protein that is capable of cleaving only one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase cleaves the target strand of the double-stranded nucleic acid molecule, meaning that the Cas9 nickase cleaves the strand that is base-paired (complementary) to the gRNA (e.g., sgRNA) that is bound to the Cas9. In some embodiments, the Cas9 nickase includes the D10A mutation and has a histidine at position 840. In some embodiments, the Cas9 nickase cleaves the non-target, non-base-edited strand of the double-stranded nucleic acid molecule, meaning that the Cas9 nickase cleaves the strand that is not base-paired to the gRNA (e.g., sgRNA) that is bound to the Cas9. In some embodiments, the Cas9 nickase includes the H840A mutation and has an aspartic acid residue at position 10, or a corresponding mutation. In some embodiments, the Cas9 nickase includes an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical compared to any of the Cas9 nickases provided herein. Based on the disclosure herein and knowledge in the art, additional suitable Cas9 nickases will be apparent to those of skill in the art and are within the scope of the disclosure herein.
[0436] The amino acid sequence of an exemplary catalytic Cas9 nickase (nCas9) is as follows:
[0437]
[0438] In some embodiments, Cas9 refers to Cas9 from archaea (e.g., Nanoarchaeota), which constitute a domain and kingdom of single-celled prokaryotic microorganisms. In some embodiments, the programmable nucleotide-binding protein refers to CasX or CasY proteins, which have been described, for example, in Burstein et al., “New CRISPR-Cas systems from uncultivated microbes.” Cell Res. 2017 Feb 21. doi:10.1038 / cr.2017.21, the entire content of which is hereby incorporated by reference herein. Using genome-resolved metagenomics, several CRISPR-Cas systems were identified, including Cas9, which was reported for the first time in the archaeal domain of life. This divergent Cas9 protein was found as part of an active CRISPR-Cas in the little-studied genus Nanoarchaeum. In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, were discovered, which are among the most compact systems found to date. In some embodiments, in the base editor systems described herein, Cas9 is replaced by CasX or a variant of CasX. In some embodiments, in the base editor systems described herein, Cas9 is replaced by CasY or a variant of CasY. It should be understood that other RNA-guided DNA-binding proteins may be used as nucleic acid programmable DNA-binding proteins (napDNAbp) and are within the scope of the present disclosure.
[0439] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) of any of the fusion proteins provided herein may be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the amino acid sequence included in the napDNAbp has at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to a naturally occurring CasX or CasY protein. In some embodiments, the napDNAbp is a naturally occurring CasX or CasY protein. In some embodiments, the programmable nucleotide binding protein includes an amino acid sequence having at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to any of the CasX or CasY proteins described herein. It should be understood that CasX and CasY from other bacterial species may also be used in accordance with the disclosure herein.
[0440] An exemplary CasX ((uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) tr|F0NN87|F0NN87_SULIHCRISPR-associated Casx protein OS = Sulfolobus islandicus (strain HVE10 / 4) GN = SiH_0402 PE = 4 SV = 1) amino acid sequence is as follows:
[0441] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAVGQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG.
[0442] The amino acid sequence of exemplary CasX (>tr|F0NH53|F0NH53_SULIR CRISPR-associated protein, Casx OS=Sulfolobus islandicus (strain REY15A) GN=SiRe_0771 PE=4 SV=1) is as follows:
[0443] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAVGQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG.
[0444] CasX of Deltaproteobacteria
[0445] MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAILQVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVAEHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFLSKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDfAYNEVIARVRMWVNLNLWQKLKLSRDDAKPLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENPKKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMDEKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTDGTDIKKSGKWQGLLYGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIGRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQAAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKLAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELSAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYKSGKQPFVGAWQAFYKRRLKEVWKPNA
[0446] The amino acid sequence of exemplary CasY ((ncbi.nlm.nih.gov / protein / APG80656.1)>APG80656.1 CRISPR-associated protein CasY (uncultured Candidatus bacteria)) is as follows:
[0447] MSKRHPRISGVKGYRLHAQRLEYTGKSGAMRTIKYPLYSSPSGGRTVPREIVSAINDDYVGLYGLSNFDDLYNAEKRNEEKVYSVLDFWYDCVQYGAVFSYTAPGLLKNVAEVRGGSYELTKTLKGSHLYDELQIDKVIKFLNKKEISRANGSLDKLKKDIIDCFKAEYRERHKDQCNKLADDIKNAKKDAGASLGERQKKLFRDFFGISEQSENDKPSFTNPLNLTCCLLPFDTVNNNRNRGEVLFNKLKEYAQKLDKNEGSLEMWEYIGIGNSGTAFSNFLGEGFLGRLRENKITELKKAMMDITDAWRGQEQEEELEKRLRILAALTIKLREPKFDNHWGGYRSDINGKLSSWLQNYINQTVKIKEDLKGHKKDLKKAKEMINRFGESDTKEEAVVSSLLESIEKIVPDDSADDEKPDIPAIAIYRRFLSDGRLTLNRFVQREDVQEALIKERLEAEKKKKPKKRKKKSDAEDEKETIDFKELFPHLAKPLKL
[0448] VPNFYGDSKRELYKKYKNAAIYTDALWKAVEKIYKSAFSSSLKNSFFDTDFDKDFFIKRLQKIFSVYRRFNTDKWKPIVKNSFAPYCDIVSLAENEVLYKPKQSRSRKSAAIDKNRVRLPSTENIAKAGIALARELSVAGFDWKDLLKKEEHEEYIDLIELHKTALALLLAVTETQLDISALDFVENGTVKDFMKTRDGNLVLEGRFLEMFSQSIVFSELRGLAGLMSRKEFITRSAIQTMNGKQAELLYIPHEFQSAKITTPKEMSRAFLDLAPAEFATSLEPESLSEKSLLKLKQMRYYPHYFGYELTRTGQGIDGGVAENALRLEKSPVKKREIKCKQYKTLGRGQNKIVLYVRSSYYQTQFLEWFLHRPKNVQTDVAVSGSFLIDEKKVKTRWNYDALTVALEPVSGSERVFVSQPFTIFPEKSAEEEGQRYLGIDIGEYGIAYTALEITGDSAKILDQNFISDPQLKTLREEVKGLKLDQRRGTFAMPSTKIARIRESLVHSLRNRIHHLALKHKAKIVYELEVSRFEEGKQKIKKVYATLKKADVYSEIDADKNLQTTVWGKLAVASEISASYTSQFCGACKKLWRAEMQVDETITTQELIGTVRVIKGGTLIDAIKDFMRPPIFDENDTPFPKYRDFCDKHHISKKMRGNSCLFICPFCRANADADIQASQTIALLRYVKEEKKVEDYFERFRKLKNIKVLGQMKKI.
[0449] The Cas9 nuclease has two functional endonuclease domains: RuvC and HNH. Once bound to the target, Cas9 undergoes a conformational change that positions the nuclease domains to cleave opposite strands of the target DNA. The ultimate outcome of Cas9-mediated DNA cleavage is a double-strand break (DSB) within the target DNA (approximately 3-4 nucleotides upstream of the PAM sequence). The resulting DSB is then repaired by one of two general repair pathways: (1) the efficient but error-prone non-homologous end joining (NHEJ) pathway; or (2) the less efficient but high-fidelity homologous directed repair (HDR) pathway.
[0450] The "efficiency" of non - homologous end joining (NHEJ) and / or homology - directed repair (HDR) can be calculated by any convenient method. For example, in some embodiments, the efficiency can be expressed as the percentage of successful HDR. For example, the Surveyor nuclease assay can be used to generate cleaved products, and the ratio of products to substrate can be used to calculate this percentage. For example, the Surveyor nuclease can be used to directly cleave DNA containing a newly integrated restriction endonuclease recognition sequence (as a result of successful HDR). The more substrate that is cleaved, the higher the percentage of HDR (the higher the efficiency of HDR). As an illustrative example, the following equation can be used: [(cleaved products) / (substrate plus cleaved products)] (e.g., (b + c) / (a + b + c), where "a" is the band intensity of the DNA substrate and "b" and "c" are the cleaved products).
[0451] In some embodiments, the efficiency can be expressed as the percentage of successful NHEJ. For example, the T7 endonuclease I assay can be used to generate cleaved products, and the ratio of products to substrate can be used to calculate the percentage of NHEJ. T7 endonuclease I cleaves mismatched heteroduplex DNA, which occurs upon hybridization of wild - type and mutant DNA strands (NHEJ generates small random insertions or deletions at the original break site). The more cleavage, the higher the percentage of NHEJ (the higher the efficiency of NHEJ). As an illustrative example, the following equation can be used to calculate the fraction (percentage) of NHEJ: (1-(1-(b + c) / (a + b + c)) 1 / 2 )×100, where "a" is the band intensity of the DNA substrate and "b" and "c" are the cleaved products (Ran et al., Cell. 2013 Sep. 12; 154(6):1380 - 9; and Ran et al., Nat Protoc. 2013 Nov.; 8(11):2281–2308).
[0452] The NHEJ repair pathway is the most active repair mechanism, and it often causes small nucleotide insertions or deletions at the DSB site. The randomness of NHEJ - mediated DSB repair has important practical implications because a population of cells expressing Cas9 and gRNA or guide polynucleotide will result in a diverse array of mutations. In most embodiments, NHEJ causes small insertions / deletions in the target DNA, which lead to amino - acid deletions, insertions, or frameshift mutations, and these mutations result in premature stop codons within the open reading frame (ORF) of the targeted gene. The ideal end result is a loss - of - function mutation within the targeted gene.
[0453] NHEJ-mediated DSB repair often disrupts the open reading frame of genes, while homology-directed repair (HDR) can be used to generate specific nucleotide changes, ranging from single nucleotide changes to large insertions such as the addition of fluorophores or tags. To use HDR for gene editing, a DNA repair template containing the desired sequence can be delivered into the target cell type together with (one or more) gRNAs and Cas9 or Cas9 nickase. The repair template can contain the desired arrangement, as well as additional homologous sequences (referred to as left & right homology arms) upstream and downstream of the target. The length of each homology arm can depend on the size of the change to be introduced, with larger insertions requiring longer homology arms. The repair template can be a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a double-stranded DNA plasmid. Even in cells expressing Cas9, gRNA, and the exogenous repair template, the efficiency of HDR is usually low (modified alleles < 10%). Since HDR occurs in the S and G2 phases of the cell cycle, the efficiency of HDR can be increased by synchronizing the cells. Chemical or genetic inhibitors of genes involved in NHEJ can also increase HDR frequency.
[0454] In some embodiments, Cas9 is a modified Cas9. A given gRNA target sequence can have additional sites with partial homology throughout the genome. These sites are referred to as off-target sites and need to be taken into account when designing the gRNA. In addition to optimizing gRNA design, CRISPR specificity can also be improved by modifying Cas9. Cas9 generates a double-stranded break (DSB) through the combined activity of two nuclease domains, RuvC and HNH. Cas9 nickase, a D10A mutant of SpCas9, retains one nuclease domain and generates a DNA nick instead of a DSB. This nickase system can also be combined with HDR-mediated gene editing for specific gene arrangements.
[0455] In some embodiments, Cas9 is a variant Cas9 protein. When compared to the amino acid sequence of the wild-type Cas9 protein, the variant Cas9 polypeptide has an amino acid sequence that is different by one amino acid (e.g., having a deletion, insertion, substitution, fusion). In some instances, the variant Cas9 polypeptide has an amino acid change (e.g., deletion, insertion, or substitution) that reduces the nuclease activity of the Cas9 polypeptide. For example, in some instances, the variant Cas9 polypeptide has less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nuclease activity of the corresponding wild-type Cas9 protein. In some embodiments, the variant Cas9 protein has no substantial nuclease activity. When a test Cas9 protein is a variant Cas9 protein that does not have substantial nuclease activity, it can be referred to as "dCas9".
[0456] In some embodiments, the variant Cas9 protein has reduced nuclease activity. For example, the variant Cas9 protein exhibits less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% of the endonuclease activity of the wild-type Cas9 protein (e.g., wild-type Streptococcus pyogenes Cas9).
[0457] In some embodiments, the variant Cas9 protein can cleave the complementary strand of the guide target sequence, but its ability to cleave the non-complementary strand of the double-stranded guide target sequence is reduced. For example, the variant Cas9 protein may have a mutation (amino acid substitution) that reduces the function of the RuvC domain. As a non-limiting example, in some embodiments, the Cas9 variant protein has a D10A (aspartic acid at amino acid position 10 is changed to alanine) mutation and can thus cleave the complementary strand of the double-stranded guide target sequence, but its ability to cleave the non-complementary strand of the double-stranded guide target sequence is reduced (thus, when the variant Cas9 protein cleaves the double-stranded target nucleic acid, a single-strand break (SSB) rather than a double-strand break (DSB) is obtained) (see, e.g., Jinek et al., Science. 2012 Aug. 17; 337(6096):816-21).
[0458] In some embodiments, the variant Cas9 protein can cleave the non-complementary strand of the double-stranded guide target sequence, but its ability to cleave the complementary strand of the guide target sequence is reduced. For example, the variant Cas9 protein may have a mutation (amino acid substitution) that reduces the function of the HNH domain (RuvC / HNH / RuvC domain motif). As a non-limiting example, in some embodiments, the variant Cas9 protein has an H840A (histidine at amino acid position 840 is changed to alanine) mutation and can thus cleave the non-complementary strand of the guide target sequence, but its ability to cleave the complementary strand of the guide target sequence is reduced (thus, when the variant Cas9 protein cleaves the double-stranded target nucleic acid, an SSB rather than a DSB is obtained). The ability of such a Cas9 protein to cleave the guide target sequence (e.g., a single-stranded guide target sequence) is reduced, but it retains the ability to bind to the guide target sequence (e.g., a single-stranded guide target sequence).
[0459] In some embodiments, the variant Cas9 protein has a reduced ability to cleave both the complementary and non-complementary strands of double-stranded target DNA. As a non-limiting example, in some embodiments, the variant Cas9 protein bears two mutations, D10A and H840A, such that the polypeptide has a reduced ability to cleave both the complementary and non-complementary strands of double-stranded target DNA. The ability of such a Cas9 protein to cleave target DNA (e.g., single-stranded target DNA) is reduced, but it retains the ability to bind to target DNA (e.g., single-stranded target DNA).
[0460] As another non-limiting example, in some embodiments, the variant Cas9 protein bears two mutations, W476A and W1126A, such that the ability of the polypeptide to cleave target DNA is reduced. The ability of such a Cas9 protein to cleave target DNA (e.g., single-stranded target DNA) is reduced, but it retains the ability to bind to target DNA (e.g., single-stranded target DNA).
[0461] As another non-limiting example, in some embodiments, the variant Cas9 protein bears mutations P475A, W476A, N477A, D1125A, W1126A, and D1127A, such that the ability of the polypeptide to cleave target DNA is reduced. The ability of such a Cas9 protein to cleave target DNA (e.g., single-stranded target DNA) is reduced, but it retains the ability to bind to target DNA (e.g., single-stranded target DNA).
[0462] As another non-limiting example, in some embodiments, the variant Cas9 protein bears mutations H840A, W476A, and W1126A, such that the ability of the polypeptide to cleave target DNA is reduced. The ability of such a Cas9 protein to cleave target DNA (e.g., single-stranded target DNA) is reduced, but it retains the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some embodiments, the variant Cas9 protein bears mutations H840A, D10A, W476A, and W1126A, such that the ability of the polypeptide to cleave target DNA is reduced. The ability of such a Cas9 protein to cleave target DNA (e.g., single-stranded target DNA) is reduced, but it retains the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, the catalytic His residue (A840H) has been restored at position 840 in the Cas9 HNH domain of the variant Cas9.
[0463] As another non-limiting example, in some embodiments, the variant Cas9 protein bears the H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations such that the ability of the polypeptide to cleave the target DNA is reduced. The ability of such a Cas9 protein to cleave the target DNA (e.g., single-stranded target DNA) is reduced, but it retains the ability to bind to the target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some embodiments, the variant Cas9 protein bears the D10A, H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations such that the ability of the polypeptide to cleave the target DNA is reduced. The ability of such a Cas9 protein to cleave the target DNA (e.g., single-stranded target DNA) is reduced, but it retains the ability to bind to the target DNA (e.g., single-stranded target DNA). In some embodiments, when the variant Cas9 protein bears both the W476A and W1126A mutations, or when the variant Cas9 protein bears the P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, the variant Cas9 protein cannot effectively bind to the PAM sequence. Thus, in some such embodiments, when such a variant Cas9 protein is used in a binding method, the method does not require a PAM sequence. In other words, in some embodiments, when such a variant Cas9 protein is used in a binding method, the method can include a guide RNA, but the method can be performed in the absence of a PAM sequence (and the binding specificity is thus provided by the targeting segment of the guide RNA). Other residues can be mutated to achieve the above effect (i.e., to inactivate one or the other nuclease moiety). As non-limiting examples, the residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 can be altered (i.e., replaced). Similarly, mutations other than alanine substitutions are suitable.
[0464] In some embodiments, a variant Cas9 protein with reduced catalytic activity (e.g., when the Cas9 protein has D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 mutations, such as D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A), as long as it retains the ability to interact with the guide RNA, the variant Cas9 protein can still bind to the target DNA in a site-specific manner (because it can still be directed to the target DNA sequence by the guide RNA).
[0465] In some embodiments, the variant Cas protein can be spCas9, spCas9-VRQR, spCas9-VRER, xCas9(sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL.
[0466] In some embodiments, a modified SpCas9 is used, which contains the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (SpCas9-MQKFRAER) and is specific for the altered PAM 5'-NGC-3'.
[0467] Alternatives to Streptococcus pyogenes Cas9 can include RNA-guided endonucleases from the Cpf1 family, which show cleavage activity in mammalian cells. CRISPR from Prevotella and Francisella 1 (CRISPR / Cpf1) is a DNA-editing technology similar to the CRISPR / Cas9 system. Cpf1 is an RNA-guided endonuclease of the type II CRISPR / Cas system. This acquired immune mechanism has been found in Prevotella and Francisella bacteria. The Cpf1 gene is associated with the CRISPR locus and encodes an endonuclease that uses a guide RNA to find and cleave viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9, overcoming some limitations of the CRISPR / Cas9 system. Different from the Cas9 nuclease, the result of Cpf1-mediated DNA cleavage is a double-strand break with short 3' overhangs. The staggered cleavage pattern of Cpf1 opens up the possibility of directional gene transfer (similar to traditional restriction enzyme cloning), which can improve the efficiency of gene editing. Like the Cas9 variants and xenologs described above, Cpf1 can also expand the number of sites targetable by CRISPR to AT-rich regions or AT-rich genomes lacking the NGG PAM sites preferred by SpCas9. The Cpf1 locus contains a mixed α / β domain, RuvC-I, followed by a helical region, RuvC-II, and a zinc finger-like domain. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9. In addition, Cpf1 does not have an HNH endonuclease domain, and the N-terminus of Cpf1 does not have the α-helical recognition lobe of Cas9. The structure of the Cpf1 CRISPR-Cas domain shows that Cpf1 has unique functions and is classified as a type 2 class V CRISPR system. The Cas1, Cas2, and Cas4 proteins encoded by the Cpf1 locus are more similar to type I and type III than to type II systems. Functional Cpf1 does not require trans-activating CRISPR RNA (tracrRNA) and thus only requires CRISPR (crRNA). This is beneficial for genome editing not only because Cpf1 is smaller than Cas9, but also because it has a smaller sgRNA molecule (about half the number of nucleotides of Cas9). Instead of the G-rich PAM targeted by Cas9, the Cpf1-crRNA complex cleaves the target DNA or RNA by recognizing the protospacer adjacent motif 5'-YTN-3'. After recognizing the PAM, Cpf1 introduces a sticky-end-like DNA double-strand break with overhangs of 4 or 5 nucleotides.
[0468] Nucleic acid programmable DNA-binding protein
[0469] Some aspects disclosed herein provide fusion proteins that include a domain that acts as a nucleic acid-programmable DNA-binding protein, which can be used to direct a protein, such as a base editor, to a specific nucleic acid (e.g., DNA or RNA) sequence. In certain embodiments, the fusion protein includes a nucleic acid-programmable DNA-binding protein domain and one or more deaminase domains. Non-limiting examples of nucleic acid-programmable DNA-binding proteins include: Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i. Non-limiting examples of Cas enzymes include: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Type IICas effector proteins, Type V Cas effector proteins, Type VI Cas effector proteins, CARF, DinG, their homologs, or their modified or engineered versions. Other nucleic acid-programmable DNA-binding proteins are also within the scope of the disclosure herein, even though they may not be explicitly listed in this disclosure.See, for example, Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?”, CRISPR J. 2018 Oct; 1: 325-336. doi: 10.1089 / crispr.2018.0033; Yan et al., “Functionally diverse type V CRISPR-Cas systems”, Science. 2019 Jan 4; 363(6422): 88-91. doi: 10.1126 / science.aav7271, the entire contents of each of which are hereby incorporated by reference.
[0470] An example of a nucleic acid programmable DNA-binding protein having a PAM specificity different from Cas9 is the clustered regularly interspaced short palindromic repeats (Cpf1) from Prevotella and Francisella 1. Similar to Cas9, Cpf1 is also a type II CRISPR effector. It has been shown that Cpf1 mediates robust DNA interference, and its characteristics are different from Cas9. Cpf1 is a single RNA-guided endonuclease lacking tracrRNA, which utilizes a T-rich protospacer-adjacent motif (TTN, TTTN, or YTN). In addition, Cpf1 cleaves DNA via staggered DNA double-strand breaks. Among 16 Cpf1-family proteins, two enzymes from Acidaminococcus and Lachnospiraceae have been shown to have efficient genome-editing activity in human cells. Cpf1 proteins are known in the art and have been previously described, for example, Yamano et al., “Crystal structure of Cpf1 in complex with guide RNA and target DNA.” Cell (165) 2016, p. 949-96; the entire contents of which are hereby incorporated by reference.
[0471] The nuclease-inactivated Cpf1 (dCpf1) variants useful in the compositions and methods of the present invention can be used as guide nucleotide sequence-programmable DNA-binding protein domains. The Cpf1 protein has a RuvC-like nuclease domain similar to the RuvC domain of Cas9, but does not have an HNH endonuclease domain, and the N-terminus of Cpf1 does not have the α-helical recognition lobe of Cas9. It has been shown that in Zetsche et al., Cell, 163, 759-771, 2015 (which is incorporated herein by reference), this RuvC-like domain of Cpf1 is responsible for cleaving both DNA strands, and deactivation of this RuvC-like domain deactivates the activity of the Cpf1 nuclease. For example, mutations corresponding to D917A, E1006A, or D1255A in Francisella novicida Cpf1 deactivate Cpf1 nuclease activity. In some embodiments, the dCpf1 disclosed in the present disclosure includes mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A. It should be understood that any mutation that deactivates the RuvC domain of Cpf1 can be used according to the present disclosure, such as substitution mutations, deletions, or insertions.
[0472] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) of any fusion protein provided herein may be a Cpf1 protein. In some embodiments, the Cpf1 protein is a Cpf1 nickase (nCpf1). In some embodiments, the Cpf1 protein is a nuclease-inactivated Cpf1 (dCpf1). In some embodiments, the amino acid sequence included in the Cpf1, the nCpf1, or the dCpf1 has at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to the Cpf1 sequence described herein. In some embodiments, the amino acid sequence included in the dCpf1 has at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to the Cpf1 sequence described herein, and includes mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A. It should be understood that Cpf1 from other bacterial species can be used according to the present disclosure.
[0473] Wild-type Francisella novicida Cpf1 (D917, E1006, and D1255 are bold and underlined)
[0474] MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNK
[0475]
[0476] Francisella novicida Cpf1 D917A (A917, E1006, and D1255 are bold and underlined)
[0477]
[0478]
[0479] Francisella novicida Cpf1 E1006A (D917, A1006, and D1255 are bold and underlined)
[0480]
[0481] Francisella novicida Cpf1 D1255A (D917, E1006, and A1255 are bold and underlined)
[0482] MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYI TQQIAPKNLDNPS
[0483]
[0484] Francisella novicida Cpf1 D917A / E1006A (A917, A1006, and D1255 are bolded and underlined)
[0485]
[0486]
[0487] Francisella novicida Cpf1 D917A / D1255A (A917, E1006, and A1255 are bolded and underlined)
[0488]
[0489] Francisella novicida Cpf1 E1006A / D1255A (D917, A1006, and A1255 are bolded and underlined)
[0490] MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDY
[0491]
[0492]
[0493] Francisella novicida Cpf1 D917A / E1006A / D1255A (A917, A1006, and A1255 are bold and underlined)
[0494] MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYSDVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDITDIDEALEIIKSFKGWTTYFKGFHENRKNVYSSNDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAEELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYKMSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLTDLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEQELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILANFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNNLLHKLKIFHISQSEDKANILDKDEHFYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKIFDDKAIKENKGEGYKKIVYKLLPGANKMLPKVFFSAKSIKFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKFIDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISESYIDSVVNQGKLYLFQIYNKDFSAYSKGRPNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPKKITHPAKEAIANKNKDNPKKESVFEYDLI
[0495]
[0496] In some embodiments, one of the Cas9 domains present in the fusion protein can be replaced with a guide nucleotide sequence-programmable DNA-binding protein domain that has no requirement for a PAM sequence.
[0497] In some embodiments, the Cas9 domain is a Cas9 domain from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is nuclease-active SaCas9, nuclease-inactivated SaCas9 (SaCas9d), or SaCas9 nickase (SaCas9n). In some embodiments, the SaCas9 comprises the N579A mutation, or a corresponding mutation in any of the amino acid sequences provided herein.
[0498] In some embodiments, the SaCas9 domain, the SaCas9d domain, or the SaCas9n domain can bind to a nucleic acid sequence having a non-canonical PAM. In some embodiments, the SaCas9 domain, the SaCas9d domain, or the SaCas9n domain can bind to a nucleic acid sequence having an NNGRRT or NNGRRT PAM sequence. In some embodiments, the SaCas9 domain comprises one or more of the E781X, N967X, and R1014X mutations, or a corresponding mutation in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SaCas9 domain comprises one or more of the E781K, N967K, and R1014H mutations, or one or more corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SaCas9 domain comprises the E781K, N967K, or R1014H mutation, or a corresponding mutation in any of the amino acid sequences provided herein.
[0499] Exemplary SaCas9 sequences
[0500] KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYEMLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNS
[0501]
[0502] The residue N579 above, which is underlined and shown in bold, can be mutated (e.g., to A579) to produce SaCas9 nickase.
[0503] Exemplary SaCas9n sequence
[0504]
[0505] The residue A579 above, which can be mutated from N579 to produce SaCas9 nickase, is underlined and shown in bold.
[0506] Exemplary SaKKH Cas9
[0507]
[0508] The residue A579 above, which can be mutated from N579 to produce SaCas9 nickase, is underlined and shown in bold. The residues K781, K967, and H1014 above, which can be mutated from E781, N967, and R1014 to produce SaKKH Cas9, are underlined and shown in italics.
[0509] In some embodiments, the napDNAbp is a circular permutant. In the following sequences, plain text represents the adenosine deaminase sequence, bold sequences represent sequences derived from Cas9, italic sequences represent linker sequences, and underlined sequences represent bipartite nuclear localization sequences.
[0510] CP5 (with MSP “NGC” PID and “D10A” nickase):
[0511]
[0512] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is a single effector of a microbial CRISPR-Cas system. Single effectors of microbial CRISPR-Cas systems include, but are not limited to: Cas9, Cpf1, Cas12b / C2c1, and Cas12c / C2c3. Generally, microbial CRISPR-Cas systems are divided into class 1 and class 2 systems. Class 1 systems have multi-subunit effector complexes, while class 2 systems have single protein effectors. For example, Cas9 and Cpf1 are class 2 effectors. In addition to Cas9 and Cpf1, three different class 2 CRISPR-Cas systems (Cas12b / C2c1, and Cas12c / C2c3) have been described in: Shmakov et al., “Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems”, Mol. Cell, 2015 Nov. 5; 60(3):385-397, the entire content of which is incorporated herein by reference. Effectors of two of these systems, Cas12b / C2c1 and Cas12c / C2c3, contain an RuvC-like endonuclease domain related to Cpf1. The effector of the third system contains two predicted HEPN RNase domains. The production of mature CRISPR RNA does not depend on tracrRNA, which is different from the CRISPR RNA produced by Cas12b / C2c1. The cleavage of DNA by Cas12b / C2c1 depends on both CRISPR RNA and tracrRNA.
[0513] The crystal structure of the complex of Alicyclobacillus acidoterrestris Cas12b / C2c1 (AacC2c1) with a chimeric single-guide RNA (sgRNA) has been reported, see, e.g., Liu et al., "Structure of the C2c1-sgRNA Complex Reveals the RNA-Guided DNA Cleavage Mechanism", Mol. Cell, Jan. 19, 2017; 65(2):310-322, the entire content of which is incorporated herein by reference. The crystal structure of Alicyclobacillus acidoterrestris C2c1 bound to target DNA in a ternary complex has also been reported. See, e.g., Yang et al., "PAM-dependent Target DNA Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease", Cell, Dec. 15, 2016; 167(7):1814-1828, the entire content of which is incorporated herein by reference. Two catalytically competent conformations of AacC2c1 (with target and non-target DNA strands) have been independently captured, which are positioned within a single RuvC catalytic pocket, with Cas12b / C2c1-mediated cleavage resulting in staggered seven-nucleotide breaks in the target DNA. Structural comparisons between the Cas12b / C2c1 ternary complex and previously identified Cas9 and Cpf1 counterparts demonstrate the diversity of mechanisms used by the CRISPR-Cas9 system.
[0514] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) of any fusion protein provided herein may be Cas12b / C2c1, or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is Cas12b / C2c1 protein. In some embodiments, the napDNAbp is Cas12c / C2c3 protein. In some embodiments, the amino acid sequence included in the napDNAbp is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the amino acid sequence included in the napDNAbp is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any napDNAbp sequence provided herein. It should be understood that Cas12b / C2c1 or Cas12c / C2c3 from other bacterial species may also be used in accordance with the disclosure herein.
[0515] Cas12b / C2c1((uniprot.org / uniprot / T0D7A2#2)sp|T0D7A2|
[0516] C2C1_ALIAG CRISPR-associated endonuclease C2c1 OS= Alicyclobacillus acidoterrestris (strain ATCC 49025 / DSM 3922 / CIP 106132 / NCIMB 13137 / GD3B) GN=c2c1 PE=1 SV=1) Amino acid sequence is as follows:
[0517] MAVKSIKVKLRLDDMPEIRAGLWKLHKEVNAGVRYYTEWLSLLRQENLYRRSPNGDGEQECDKTAEECKAELLERLRARQVENGHRGPAGSDDELLQLARQLYELLVPQAIGAKGDAQQIARKFLSPLADKDAVGGLGIAKAGNKPRWVRMREAGEPGWEEEKEKAETRKSADRTADVLRALADFGLKPLMRVYTDSEMSSVEWKPLRKGQAVRTWDRDMFQQAIERMMSWESWNQRVGQEYAKLVEQ
[0518] KNRFEQKNFVGQEHLVHLVNQLQQDMKEASPGLESKEQTAHYVTGRALRGSDKVFEKWGKLAPDAPFDLYDAEIKNVQRRNTRRFGSHDLFAKLAEPEYQALWREDASFLTRYAVYNSILRKLNHAKMFATFTLPDATAHPIWTRFDKLGGNLHQYTFLFNEFGERRHAIRFHKLLKVENGVAREVDDVTVPISMSEQLDNLLPRDPNEPIALYFRDYGAEQHFTGEFGGAKIQCRRDQLAHMHRRRGARDVYLNVSVRVQSQSEARGERRPPYAAVFRLVGDNHRAFVHFDKLSDYLAEHPDDGKLGSEGLLSGLRVMSVDLGLRTSASISVFRVARKDELKPNSKGRVPFFFPIKGNDNLVAVHERSQLLKLPGETESKDLRAIREERQRTLRQLRTQLAYLRLLVRCGSEDVGRRERSWAKLIEQPVDAANHMTPDWREAFENELQKLKSLHGICSDKEWMDAVYESVRRVWRHMGKQVRDWRKDVRSGERPKIRGYAKDVVGGNSIEQIEYLERQYKFLKSWSFFGKVSGQVIRAEKGSRFAITLREHIDHAKEDRLKKLADRIIMEALGYVYALDERGKGKWVAKYPPCQLILLEELSEYQFNNDRPPSENNQLMQWSHRGVFQELINQAQVHDLLVGTMYAAFSSRFDARTGAPGIRCRRVPARCTQEHNPEPFPWWLNKFVVEHTLDACPLRADDLIPTGEGEIFVSPFSAEEGDFHQIHADLNAAQNLQQRLWSDFDISQIRLRCDWGEVDGELVLIPRLTGKRTADSYSNKVFYTNTGVTYYERERGKKRRKVFAQEKLSEEEAELLVEADEAREKSVVLMRDPSGIINRGNWTRQKEFWSMV NQRIEGYLVKQIRSRVPLQDSACENTGDI.
[0519] AacCas12b (Alicyclobacillus acidiphilus) - WP_067623834
[0520]
[0521] ELSEEEAELLVEADEAREKSVVLMRDPSGIINRGDWTRQKEFWSMVNQRIEGYLVKQIRSRVRLQESACENTGDI
[0522] BhCas12b (Bacillus hisashii) NCBI reference sequence: WP_095142515
[0523]
[0524] including a variant designated as BvCas12b V4 (S893R / K846R / E837G change, relative to the wild type above)
[0525] BhCas12b(V4) is represented as follows: 5’mRNA cap---5’UTR---bhCas12b---termination (STOP) sequence---3’UTR---120 polyA tail
[0526] 5’UTR:
[0527] GGGAAATAAGAGAGAAAAGAAGAGTAAGAAGAAATATAAGAGCCACC
[0528] 3’UTR (UTR of TriLink standard)
[0529] GCTGGAGCCTCGGTGGCCATGCTTCTTGCCCCTTGGGCCTCCCCCCAGCCCCTCCTCCCCTTCCTGCACCCGTACCCCCGTGGTCTTTGAATAAAGTCTGA
[0530] Nucleic acid sequence of bhCas12b(V4)
[0531]
[0532]
[0533] In some embodiments, the Cas12b is BvCas12B, which is a variant of BhCas12b and includes the following changes relative to BhCas12B: S893R, K846R, and E837G.
[0534] BvCas12b (Bacillus sp. V3-13) NCBI reference sequence: WP_101661451.1
[0535] MAIRSIKLKMKTNSGTDSIYLRKALWRTHQLINEGIAYYMNLLTLYRQEAIGDKTKEAYQAELINIIRNQQRNNGSSEEHGSDQEILALLRQLYELIIPSSIGESGDANQLGNKFLYPLVDPNSQSGKGTSNAGRKPRWKRLKEEGNPDWELEKKKDEERKAKDPTVKIFDNLNKYGLLPLFPLFTNIQKDIEWLPLGKRQSVRKWDKDMFIQAIERLLSWESWNRRVADEYKQLKEKTESYYKEHLTGGEEWIEKIRKFEKERNMELEKNAFAPNDGYFITSRQIRGWDRVYEKWSKLPESASPEELWKVVAEQQNKMSEGFGDPKVFSFLANRENRDIWRGHSERIYHIAAYNGLQKKLSRTKEQATFTLPDAIEHPLWIRYESPGGTNLNLFKLEEKQKKNYYVTLSKIIWPSEEKWIEKENIEIPLAPSIQFNRQIKLKQHVKGKQEISFSDYSSRISLDGVLGGSRIQFNRKYIKNHKELLGEGDIGPVFFNLVVDVAPLQETRNGRLQSPIGKALKVISSDFSKVIDYKPKELMDWMNTGSASNSFGVASLL
[0536] EGMRVMSIDMGQRTSASVSIFEVVKELPKDQEQKLFYSINDTELFAIHKRSFLLNLPGEVVTKNNKQQRQERRKKRQFVRSQIRMLANVLRLETKKTPDERKKAIHKLMEIVQSYDSWTASQKEVWEKELNLLTNMAAFNDEIWKESLVELHHRIEPYVGQIVSKWRKGLSEGRKNLAGISMWNIDELEDTRRLLISWSKRSRTPGEANRIETDEPFGSSLLQHIQNVKDDRLKQMANLIIMTALGFKYDKEEKDRYKRWKETYPACQIILFENLNRYLFNLDRSRRENSRLMKWAHRSIPRTVSMQGEMFGLQVGDVRSEYSSRFHAKTGAPGIRCHALTEEDLKAGSNTLKRLIEDGFINESELAYLKKGDIIPSQGGELFVTLSKRYKKDSDNNELTVIHADINAAQNLQKRFWQQNSEVYRVPCQLARMGEDKLYIPKSQTETIKKYFGKGSFVKNNTEQEVYKWEKSEKMKIKTDTTFDLQDLDGFEDISKTIELAQEQQKKYLTMFRDPSGYFFNNETWRPQKEYWSIVNNIIKSCLKKKILSNKVEL
[0537] Guide polynucleotide
[0538] In one embodiment, the guide polynucleotide is a guide RNA. The RNA / Cas complex can help "guide" the Cas protein to the target DNA. Cas9 / crRNA / tracrRNA cleaves linear or circular dsDNA targets complementary to the spacer sequence by endonucleolytic cleavage. The target strand that is not complementary to the crRNA is first cleaved by endonucleolytic cleavage and then trimmed 3'-5' by exonucleolytic cleavage. In nature, DNA-binding and cleavage generally require a protein and two RNAs. However, single guide RNAs ("sgRNAs", or simply "gNRAs") can be engineered to incorporate aspects of both crRNA and tracrRNA into a single RNA species. See, e.g., Jinek M. et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif (PAM or protospacer adjacent motif) in the CRISPR repeat to help distinguish self from non-self. The Cas9 nuclease sequence and structure are well known to those of skill in the art (see, e.g., "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Ferretti, J.J. et al., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E. et al., Nature 471:602-607 (2011); and "Programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M. et al., Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference). Orthologs of Cas9 have been described in various species, including but not limited to Streptococcus pyogenes and Streptococcus thermophilus.Based on the present disclosure, additional suitable Cas9 nucleases and sequences will be apparent to those skilled in the art, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II C RISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737; the entire content of which is incorporated herein by reference. In some embodiments, the Cas9 nuclease has an inactivated (e.g., deactivated) DNA cleavage domain, i.e., the Cas9 is a nickase.
[0539] In some embodiments, the guide polynucleotide is at least one single guide RNA ("sgRNA" or "gNRA"). In some embodiments, the guide polynucleotide is at least one tracrRNA. In some embodiments, the guide polynucleotide does not require a PAM sequence to direct the polynucleotide-programmable DNA-binding domain (e.g., Cas9 or Cpf1) to the target nucleotide sequence.
[0540] The polynucleotide-programmable nucleotide-binding domain (e.g., a CRISPR-derived domain) of the base editors disclosed herein can recognize a target polynucleotide by association with a guide polynucleotide. The guide polynucleotide (e.g., gRNA) is typically single-stranded and can be programmed to site-specifically bind (i.e., via complementary base pairing) to the target sequence of the polynucleotide, thereby guiding the base editor linked to the guide nucleic acid to the target sequence. The guide polynucleotide can be DNA. The guide polynucleotide can be RNA. In some embodiments, the guide polynucleotide includes natural nucleotides (e.g., adenosine). In some embodiments, the guide polynucleotide includes non-natural (or unnatural) nucleotides (e.g., peptide nucleic acids or nucleotide analogs). In some embodiments, the targeting region of the guide nucleic acid sequence can be at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. The targeting region of the guide nucleic acid can be between 10-30 nucleotides in length, or between 15-25 nucleotides in length, or between 15-20 nucleotides in length.
[0541] In some embodiments, the guide polynucleotide comprises two or more individual polynucleotides that can interact with each other via, for example, complementary base pairing (e.g., dual guide polynucleotides). For example, the guide polynucleotide can comprise a CRISPR RNA (crRNA) and a trans-activating CRISPR RNA (tracrRNA). For example, the guide polynucleotide can comprise one or more trans-activating CRISPR RNAs (tracrRNAs).
[0542] In type II CRISPR systems, targeting of nucleic acids by a CRISPR protein (e.g., Cas9) generally requires complementary base pairing between a first RNA molecule (crRNA) that comprises a sequence that recognizes the target sequence, and a second RNA molecule (trRNA) that comprises a repeat sequence that forms a scaffolding region that stabilizes the guide RNA-CRISPR protein complex. Such dual guide RNA systems can be used as guide polynucleotides to direct the base editors disclosed herein to a target polynucleotide sequence.
[0543] In some embodiments, the base editors provided herein utilize a single guide polynucleotide (e.g., gRNA). In some embodiments, the base editors provided herein utilize a dual guide polynucleotide (e.g., dual gRNA). In some embodiments, the base editors provided herein utilize one or more guide polynucleotides (e.g., multiplex gRNA). In some embodiments, a single guide polynucleotide is used with different base editors described herein. For example, a single guide polynucleotide can be used with a cytidine base editor and an adenosine base editor.
[0544] In other embodiments, the guide polynucleotide can comprise the polynucleotide targeting portion of the nucleic acid and the scaffolding portion of the nucleic acid in a single molecule (i.e., single-molecule guide nucleic acid). For example, the single-molecule guide polynucleotide can be a single guide RNA (sgRNA or gRNA). As used herein, the term guide polynucleotide sequence is intended to encompass any single, dual, or multi-molecule nucleic acid capable of interacting with a base editor and directing it to a target polynucleotide sequence.
[0545] Generally, a guide polynucleotide (e.g., a crRNA / trRNA complex or a gRNA) includes a "polynucleotide-targeting segment" that contains a sequence capable of recognizing and binding to a target polynucleotide sequence, and a "protein-binding segment" that stabilizes the guide polynucleotide within the polynucleotide programmable nucleotide-binding domain component of a base editor. In some embodiments, the polynucleotide-targeting segment of the guide polynucleotide recognizes and binds to a DNA polynucleotide, thereby facilitating editing of bases within the DNA. In other embodiments, the polynucleotide-targeting segment of the guide polynucleotide recognizes and binds to a polynucleotide, thereby facilitating editing of bases within the RNA. As used herein, a "segment" refers to a section or region of a molecule, e.g., a contiguous stretch of nucleotides within a guide polynucleotide. A segment can also refer to a region / section of a complex such that a segment can include regions of more than one molecule. For example, when a guide polynucleotide includes multiple nucleic acid molecules, the protein-binding segment can comprise all or part of multiple individual molecules (e.g., hybridized along complementary regions). In some embodiments, the protein-binding segment of a DNA-targeting RNA that includes two separate molecules can include (i) base pairs 40-75 of a first RNA molecule that is 100 base pairs in length; and (ii) base pairs 10-25 of a second RNA molecule that is 50 base pairs in length. The definition of a "segment", unless otherwise specifically defined in a particular context, is not limited to a particular total number of base pairs, is not limited to any particular number of base pairs from a given RNA molecule, is not limited to a particular number of individual molecules within a complex, and can comprise regions of RNA molecules of any total length and can comprise regions that are complementary to other molecules.
[0546] A guide RNA or guide polynucleotide can include two or more RNAs, such as CRISPR RNA (crRNA) and trans-activating crRNA (tracrRNA). A guide RNA or guide polynucleotide can sometimes include a single-stranded RNA, or a single guide RNA (sgRNA), which is formed by fusing a tracrRNA and a portion (e.g., the functional portion) of a crRNA. A guide RNA or guide polynucleotide can also be a dual RNA that includes a crRNA and a tracrRNA. Additionally, a crRNA can hybridize to a target DNA.
[0547] As described above, a guide RNA or guide polynucleotide can be a product of expression. For example, the DNA encoding the guide RNA can be a vector that includes a sequence encoding the guide RNA. A guide RNA or guide polynucleotide can be transferred into a cell by transfection with an isolated guide RNA or plasmid DNA (including a sequence encoding the guide RNA and a promoter). A guide RNA or guide polynucleotide can also be transferred into a cell by other means, such as using virus-mediated gene delivery.
[0548] The guide RNA or guide polynucleotide can be isolated. For example, the guide RNA can be transfected into cells or an organism in the form of isolated RNA. The guide RNA can be prepared by in vitro transcription using any in vitro transcription system known in the art. The guide RNA can be transferred into cells in the form of isolated RNA rather than in the form of a plasmid comprising the guide RNA coding sequence.
[0549] The guide RNA or guide polynucleotide can comprise three regions: a first region at the 5' end that can be complementary to a target site in a chromosomal sequence, a second internal region that can form a stem-loop structure, and a third region that can be single-stranded at the 3' end. The first region of each guide RNA can also be different such that each guide RNA directs the fusion protein to a specific target site. In addition, the second and third regions of each guide RNA can be the same in all guide RNAs.
[0550] The first region of the guide RNA or guide polynucleotide can be complementary at a target site in the chromosomal sequence such that the first region of the guide RNA can base pair with the target site. In some embodiments, the first region of the guide RNA can comprise from or from about 10 nucleotides to 25 nucleotides (i.e., from 10 nucleotides to 25 nucleotides; or from about 10 nucleotides to about 25 nucleotides; or from 10 nucleotides to about 25 nucleotides; or from about 10 nucleotides to 25 nucleotides) or more. For example, the base pairing region between the first region of the guide RNA and the target site in the chromosomal sequence can be or can be about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25, or more nucleotides in length. Sometimes, the first region of the guide RNA can be or can be about 19, 20, or 21 nucleotides in length.
[0551] The guide RNA or guide polynucleotide can also comprise a second region that forms a secondary structure. For example, the secondary structure formed by the guide RNA can comprise a stem (or hairpin) and a loop. The lengths of the loop and the stem can vary. For example, the length of the loop can range from or from about 3 to 10 nucleotides, while the length can range from or from about 6 to 20 base pairs. The stem can comprise one or more bulges of 1 to 10 or about 10 nucleotides. The total length of the second region can be in the range of about 16 to 60 nucleotides in length. For example, the loop can be or can be about 4 nucleotides in length, while the stem can be or can be about 12 base pairs.
[0552] The guide RNA or guide polynucleotide can also include a third region at the 3' end, which can be substantially single-stranded. For example, the third region sometimes is not complementary to any chromosomal sequence in the target cell, and sometimes is not complementary to the rest of the guide RNA. In addition, the length of the third region can vary. The length of the third region can be more than or more than about 4 nucleotides. For example, the length of the third region can be in the range of about 5 to 60 nucleotides.
[0553] The guide RNA or guide polynucleotide can target any exon or intron of a gene target. In some embodiments, the guide can target exon 1 or 2 of a gene; in other embodiments, the guide can target exon 3 or 4 of a gene. The composition can include multiple guide RNAs that all target the same exon, or in some embodiments, can include multiple guide RNAs that target different exons. Exons and introns of a gene can be targeted.
[0554] The guide RNA or guide polynucleotide can target a nucleic acid sequence that has or has about 20 nucleotides. The target nucleic acid can be less than or less than about 20 nucleotides. The target nucleic acid can be at least or at least about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, or any number of nucleotides in length between 1-100. The target nucleic acid can be at most or at most about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, or any number of nucleotides in length between 1-100. The distance of the target nucleic acid sequence from the 5' end of the first nucleotide of the PAM can be or can be about 20 bases. The guide RNA can target a nucleic acid sequence. The target nucleic acid can be at least or at least about 1-10, 1-20, 1-30, 1-40, 1-50, 1-60, 1-70, 1-80, 1-90, or 1-100 nucleotides.
[0555] A guide polynucleotide, such as a guide RNA, can refer to a nucleic acid that can hybridize to another nucleic acid, such as a target nucleic acid or protospacer sequence in the genome of a cell. The guide polynucleotide can be RNA. The guide polynucleotide can be DNA. The guide polynucleotide can be programmed or designed to site-specifically bind to a nucleic acid sequence. The guide polynucleotide can include a polynucleotide chain and can be referred to as a single guide polynucleotide. The guide polynucleotide can include two polynucleotide chains and can be referred to as a dual guide polynucleotide. The guide RNA can be introduced into a cell as an RNA molecule. For example, the RNA molecule can be transcribed in vitro and / or chemically synthesized. The RNA can be from a synthetic DNA molecule (such as It is transcribed from a gene fragment). Then, the guide RNA can be introduced into the cell as an RNA molecule. The guide RNA can also be introduced into the cell in the form of a non-RNA nucleic acid molecule (such as a DNA molecule). For example, the DNA encoding the guide RNA can be operably linked to a promoter control sequence for expressing the guide RNA in the target cell. The RNA coding sequence can be operably linked to a promoter sequence recognized by RNA polymerase subunit III (PolIII). Plasmid vectors that can be used to express the guide RNA include, but are not limited to, the px330 vector and the px333 vector. In some embodiments, the plasmid vector (such as the px333 vector) can include at least two DNA sequences encoding the guide RNA.
[0556] Methods for selecting, designing, and validating guide polynucleotides (such as guide RNAs) and target sequences are described herein and are known to those of skill in the art. For example, to minimize the effects of potential substrate promiscuity of the deaminase domain (such as the AID domain) in a base editor system, the number of residues that might be inadvertently targeted for deamination reactions (such as off-target C residues on ssDNA that might potentially reside within the target nucleic acid locus) (the location within the target nucleic acid) can be minimized. Additionally, software tools can be used to optimize the gRNA corresponding to the target nucleic acid sequence, for example, to minimize the overall off-target activity across the genome. For example, when using Streptococcus pyogenes Cas9, for each possible selection of a targeting domain, all off-target sequences (preceding the selected PAM, such as NAG or NGG) can be identified across the genome that contain at most a specific number (such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) of mismatched base-pairs. A first region of the gRNA that is complementary to the target site can be identified, and all first regions (such as crRNAs) can be ranked according to their overall predicted off-target scores; the top-ranked targeting domains represent those that are likely to have the greatest on-target and the least off-target activity. Functional evaluation of candidate targeting gRNAs can be performed using methods known in the art and / or as listed herein.
[0557] As a non-limiting example, a DNA sequence search algorithm can be used to identify the target DNA hybridization sequences in the crRNA of the guide RNA used with Cas9. Custom gRNA design software based on the public tool cas-offinder can be used for gRNA design. The public tool cas-offinder is described in the following reference: Bae S., Park J., & Kim J.-S., “Cas-OFFinder: A fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases.” Bioinformatics 30, 1473-1475 (2014). This software first calculates the genome-wide off-target propensity for each guide and then scores each guide. Generally, matches in the range from perfect match to 7 mismatches are considered for guides with lengths in the range from 17 to 24. Once the off-target sites are determined by computer calculations, a total score is calculated for each guide and summarized in a tabular output using a web interface. In addition to identifying potential target sites adjacent to the PAM sequence, this software can also identify all PAM-adjacent sequences that differ from the selected target site by 1, 2, 3, or more than 3 nucleotides. The genomic DNA sequence of the target nucleic acid sequence (e.g., the target gene) can be obtained, and publicly available tools such as the RepeatMasker program can be used to screen for repetitive elements. RepeatMasker searches the input DNA sequence for repetitive elements and low-complexity regions. The output is a detailed annotation of the repeats (elements) present in the given query sequence.
[0558] After identification, the first regions of the guide RNA (e.g., crRNA) can be ranked into different levels based on their distance from the target site, their orthogonality, and the presence of 5' nucleotides with near matches to the relevant PAM sequence (e.g., 5'G, based on identifying its near match in the human genome containing the relevant PAM (e.g., NGG PAM for Streptococcus pyogenes, or NNGRRT or NNGRRV PAM for Staphylococcus aureus)). As used herein, orthogonality refers to the number of sequences in the human genome with the minimum number of mismatches to the target sequence. For example, “high-level orthogonality” or “good orthogonality” can refer to a 20-base (20-mer) targeting domain that has no identical sequences in the human genome other than the expected target, and also has no sequences with one or two mismatches in that target sequence. Targeting domains with good orthogonality can be selected to minimize off-target DNA cleavage.
[0559] In some embodiments, a reporter system can be used to detect base-editing activity and test candidate guide polynucleotides. In some embodiments, the reporter system can include a reporter gene-based assay, wherein base-editing activity results in expression of the reporter gene. For example, the reporter system can comprise a reporter gene that includes a deactivated start codon, e.g., a mutation on the template strand from 3'-TAC-5' to 3'-CAC-5'. Once the deamination reaction of the target C is successful, the corresponding mRNA will be transcribed as 5'-AUG-3' instead of 5'-GUG-3', thereby allowing translation of the reporter gene. Suitable reporter genes will be apparent to those skilled in the art. Non-limiting examples of reporter genes include: genes encoding green fluorescent protein (GFP), red fluorescent protein (RFP), luciferase, secreted alkaline phosphatase (SEAP), or any other gene whose expression is detectable and apparent to those skilled in the art. The reporter system can be used to test many different gRNAs, e.g., to determine which residues of the target DNA sequence will be targeted by the corresponding deaminase. sgRNAs targeting the non-template strand can also be tested to evaluate off-target effects of a particular base-editing protein (e.g., a Cas9 deaminase fusion protein). In some embodiments, such gRNAs can be designed such that the mutated start codon does not base pair with the gRNA. The guide polynucleotide can include standard ribonucleotides, modified ribonucleotides (e.g., pseudouridine), ribonucleotide isomers, and / or ribonucleotide analogs. In some embodiments, the guide polynucleotide can include at least one detectable label. The detectable label can be a fluorophore (e.g., FAM, TMR, Cy3, Cy5, Texas Red, Oregon Green, Alexa Fluors, Halo tag, or a suitable fluorescent dye), a detection tag (e.g., biotin, digoxin, etc.), a tagging agent, a quantum dot, or a gold particle.
[0560] The guide polynucleotide can be chemically synthesized, enzymatically synthesized, or synthesized as a combination thereof. For example, the guide RNA can be synthesized using standard phosphoramidite-based solid-phase synthesis methods. Alternatively, the guide RNA can be synthesized in vitro by operably linking DNA encoding the guide RNA to a promoter sequence recognized by a phage RNA polymerase. Examples of suitable phage promoter sequences include T7, T3, SP6 promoter sequences, or variants thereof. In embodiments where the guide RNA comprises two separate molecules (e.g., crRNA and tracrRNA), the crRNA can be chemically synthesized while the tracrRNA can be enzymatically synthesized.
[0561] In some embodiments, the base editor system may include multiple guide polynucleotides, such as multiple gRNAs. For example, the multiple gRNAs may target one or more target loci included in the base editor system (e.g., at least 1 gRNA, at least 2 gRNAs, at least 5 gRNAs, at least 10 gRNAs, at least 20 gRNAs, at least 30 gRNAs, at least 50 gRNAs). The multiple gRNA sequences may be arranged in tandem and are preferably separated by direct repeats.
[0562] The DNA sequence encoding the guide RNA or guide polynucleotide may also be part of a vector. In addition, the vector may include additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), optional marker sequences (e.g., GFP or antibiotic resistance genes such as puromycin), an origin of replication, and the like. The DNA molecule encoding the guide RNA may also be linear. The DNA molecule encoding the guide RNA or guide polynucleotide may also be circular.
[0563] In some embodiments, one or more components of the base editor system may be encoded by several DNA sequences. Such several DNA sequences may be introduced into an expression system, such as a cell, together or separately. For example, two DNA sequences encoding the polynucleotide programmable nucleotide binding domain and the guide RNA may be introduced into a cell, where each DNA sequence may be part of a separate molecule (e.g., one vector containing the polynucleotide programmable nucleotide binding domain encoding sequence and a second vector containing the guide RNA encoding sequence) or the two may be part of the same molecule (e.g., one vector containing the encoding (and regulatory) sequences of both the polynucleotide programmable nucleotide binding domain and the guide RNA).
[0564] The guide polynucleotide may include one or more modifications to provide a nucleic acid with new or enhanced properties. The guide polynucleotide may include a nucleic acid affinity tag. The guide polynucleotide may include synthetic nucleotides, synthetic nucleotide analogs, nucleotide derivatives, and / or modified nucleotides.
[0565] In some embodiments, the gRNA or guide polynucleotide may include modifications. Modifications may be made at any location of the gRNA or guide polynucleotide. More than one modification may be made to a single gRNA or guide polynucleotide. The gRNA or guide polynucleotide may be subject to quality control after modification. In some embodiments, quality control may include PAGE, HPLC, MS, or any combination thereof.
[0566] Modifications of the gRNA or guide polynucleotide can be substitutions, insertions, deletions, chemical modifications, physical modifications, stabilization, purification, or any combination thereof.
[0567] The gRNA or guide polynucleotide can also be modified with: 5'-adenylation, 5'-guanosine-triphosphate cap, 5'-N7-methylguanosine-triphosphate cap, 5'-triphosphate cap, 3'-phosphate, 3'-thiophosphate, 5'-phosphate e, 5'-thiophosphate, cis-syn thymidine dimer, trimer, C12 spacer, C3 spacer, C6 spacer, dSpacer, PC spacer r, rSpacer, spacer 18, spacer 9, 3'-3' modification, 5'-5' modification, abasic / apurinic, acridine, azobenzene, biotin, biotin BB, biotin TEG, cholesterol TEG, desthiobiotin TEG, DNP TEG, DNP-X, DOTA, dT-biotin, dual biotin, PC biotin, psoralen C2, psoralen C6, TINA, 3'-DABCYL, Black Hole Quencher 1, Black Hole Quencher 2, DABCYL SE, dT-DABCYL, IRDye QC-1, QSY-21, QSY-35, QSY-7, QSY-9, carboxyl linker, thiol linker s, 2'-deoxyribonucleoside analog purine, 2'-deoxyribonucleoside analog pyrimidine, ribonucleoside analog, 2'-O-methyl ribonucleoside analog, sugar-modified analog, wobble / universal base, fluorescent dye label, 2'-fluoro RNA, 2'-O-methyl RNA, methylphosphonate, phosphodiester DNA, phosphodiester RNA, phosphorothioate DNA, phosphorothioate RNA, UNA, pseudouridine-5'-triphosphate, 5'-methylcytidine-5'-triphosphate, or any combination thereof.
[0568] In some embodiments, the modification is permanent. In other embodiments, the modification is temporary. In some embodiments, multiple modifications are made to the gRNA or guide polynucleotide. Modifications of the gRNA or guide polynucleotide can change the physicochemical properties of the nucleotides, such as their conformation, polarity, hydrophobicity, chemical reactivity, base-pairing interactions, or any combination thereof.
[0569] The PAM sequence can be any PAM sequence known in the art. Suitable PAM sequences include but are not limited to: NGG, NGA, NGC, NGN, NGT, NGCG, NGAG, NGAN, NGNG, NGCN, NGCG, NGTN, NNNRRT, NNNRRT, NNGRR(N), TTTV, TYCV, TYCV, TATV, NNNNGATT, NNAGAAW, or NAAAAC. Y is a pyrimidine; N is any nucleobase; W is A or T.
[0570] The modification can also be a phosphorothioate substitute. In some embodiments, the native phosphodiester bond can be readily and rapidly degraded by cellular nucleases; modification of the internucleotide linkages with phosphorothioate (PS) bond substitutes can be more stable to hydrolysis (by cellular degradation). The modification can increase the stability of the gRNA or guide polynucleotide. The modification can also enhance biological activity. In some embodiments, the gRNA whose RNA is enhanced with phosphorothioate can inhibit RNase A, RNase T1, calf serum nuclease, or any combination thereof. These features allow the PS-RNA gRNA to be used in applications with a high likelihood of exposure to nucleases in vivo or in vitro. For example, phosphorothioate (PS) bonds can be introduced between the last 3 to 5 nucleotides at the 5'- or 3'-end of the gRNA, which can inhibit exonuclease degradation. In some embodiments, phosphorothioate bonds can be added throughout the gRNA to reduce endonuclease attack.
[0571] Protospacer adjacent motif
[0572] The term "protospacer adjacent motif (PAM)" or PAM-like motif refers to a 2- to 6-base pair DNA sequence immediately following the DNA sequence targeted by the Cas9 nuclease in the CRISPR bacterial adaptive immune system. In some embodiments, the PAM can be a 5' PAM (i.e., upstream of the 5'-end of the protospacer). In other embodiments, the PAM can be a 3' PAM (i.e., downstream of the 5'-end of the protospacer).
[0573] The PAM sequence is required for target binding, but the exact sequence depends on the type of Cas protein.
[0574] The base editors provided herein can include CRISPR protein-derived domains that are capable of binding to nucleotide sequences containing a canonical or non-canonical protospacer adjacent motif (PAM) sequence. A PAM site is a nucleotide sequence adjacent to a target polynucleotide sequence. Some aspects of the present disclosure provide base editors that include all or part of a CRISPR protein having different PAM specificities. For example, the Cas9 protein, such as Cas9 from Streptococcus pyogenes (spCas9), typically requires a canonical NGG PAM sequence to bind to a specific nucleic acid region, where "N" in "NGG" is adenine (A), thymine (T), guanine (G), or cytosine (C), and G is guanine. The PAM can be CRISPR protein-specific and can vary between different base editors (including different CRISPR protein-derived domains). The PAM can be at the 5' or 3' of the target sequence. The PAM can be upstream or downstream of the target sequence. The length of the PAM can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides. Generally, the length of the PAM is between 2 and 6 nucleotides. Several PAM variants are described in Table 1 below.
[0575] Table 1. Cas9 proteins and corresponding PAM sequences
[0576] Variant PAM spCas9 NGG spCas9 - VRQR NGA spCas9 - VRER NGCG xCas9(sp) NGN saCas9 NNGRRT saCas9 - KKH NNNRRT spCas9 - MQKSER NGCG spCas9 - MQKSER NGCN spCas9 - LRKIQK NGTN spCas9 - LRVSQK NGTN spCas9 - LRVSQL NGTN spCas9 - MQKFRAER NGC Cpf1 5’(TTTV) SpyMac 5’ - NAA - 3’
[0577] In some embodiments, the PAM is NGC. In some embodiments, the NGC PAM is recognized by a Cas9 variant. In some embodiments, the NGC PAM variant contains one or more amino acid substitutions selected from: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (collectively referred to as "MQKFRAER").
[0578] In some embodiments, the PAM is NGT. In some embodiments, the NGT PAM is recognized by a Cas9 variant. In some embodiments, the NGT PAM variant is generated by targeted mutagenesis at one or more of residues 1335, 1337, 1135, 1136, 1218, and / or 1219. In some embodiments, the NGT PAM variant is generated by targeted mutagenesis at one or more of residues 1219, 1335, 1337, 1218. In some embodiments, the NGT PAM variant is generated by targeted mutagenesis at one or more of residues 1135, 1136, 1218, 1219, and 1335. In some embodiments, the NGT PAM variant is selected from the set of targeted mutations provided in Tables 2 and 3 below.
[0579] Table 2: NGT PAM variant mutations at residues 1219, 1335, 1337, 1218
[0580] Variant E1219V R1335Q T1337 G1218 1 F V T 2 F V R 3 F V Q 4 F V L 5 F V T R 6 F V R R 7 F V Q R 8 F V L R 9 L L T 10 L L R 11 L L Q 12 L L L 13 F I T 14 F I R 15 F I Q 16 F I L 17 F G C 18 H L N 19 F G C A 20 H L N V 21 L A W 22 L A F 23 L A Y 24 I A W 25 I A F 26 I A Y
[0581] Table 3: NGT PAM variant mutations at residues 1135, 1136, 1218, 1219, and 1335
[0582] Variant D1135L S1136R G1218S E1219V R1335Q 27 G 28 V 29 I 30 A 31 W 32 H 33 K 34 K 35 R 36 Q 37 T 38 N 39 I 40 A 41 N 42 Q 43 G 44 L 45 S 46 T 47 L 48 I 49 V 50 N 51 S 52 T 53 F 54 Y 55 N1286Q I1331F
[0583] In some embodiments, the NGT PAM variant is selected from variant 5, 7, 28, 31, or 36 in Tables 2 and 3. In some embodiments, these variants have improved NGT PAM recognition ability.
[0584] In some embodiments, these NGT PAM variants have mutations at residues 1219, 1335, 1337, and / or 1218. In some embodiments, the NGT PAM variant is selected from the variants provided in Table 4 below to obtain mutations with improved recognition ability.
[0585] Table 4: NGT PAM variant mutations at residues 1219, 1335, 1337, and 1218
[0586] Variant E1219V R1335Q T1337 G1218 1 F V T 2 F V R 3 F V Q 4 F V L 5 F V T R 6 F V R R 7 F V Q R 8 F V L R
[0587] In some embodiments, a base editor specific for NGT PAM can be generated as provided in Table 5 below.
[0588] Table 5. NGT PAM variants
[0589]
[0590] In some embodiments, the NGTN variant is variant 1. In some embodiments, the NGTN variant is variant 2. In some embodiments, the NGTN variant is variant 3. In some embodiments, the NGTN variant is variant 4. In some embodiments, the NGTN variant is variant 5. In some embodiments, the NGTN variant is variant 6.
[0591] In some embodiments, the Cas9 domain is the Cas9 domain of Streptococcus pyogenes (SpCas9). In some embodiments, the SpCas9 domain is nuclease-active SpCas9, nuclease-inactivated SpCas9 (SpCas9d), or SpCas9 nickase (SpCas9n). In some embodiments, the SpCas9 comprises a D9X mutation, or a corresponding mutation in any of the amino acid sequences provided herein, where X is any amino acid other than D. In some embodiments, the SpCas9 comprises a D9A mutation, or a corresponding mutation in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain, the SpCas9d domain, or the SpCas9n domain can bind to a nucleic acid sequence having a non-canonical PAM. In some embodiments, the SpCas9 domain, the SpCas9d domain, or the SpCas9n domain can bind to a nucleic acid sequence domain having an NGG, NGA, or NGCG PAM sequence. In some embodiments, the SpCas9 domain comprises one or more of the D1134X, R1334X, and T1336X mutations, or a corresponding mutation in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of the D1134E, R1334Q, and T1336R mutations, or a corresponding mutation in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises the D1134E, R1334Q, and T1336R mutations, or a corresponding multiple mutation in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises one or more of the D1134X, R1334X, and T1336X mutations, or a corresponding mutation in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of the D1134V, R1334Q, and T1336R mutations, or a corresponding mutation in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises the D1134V, R1334Q, and T1336R mutations, or a corresponding multiple mutation in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises one or more of the D1134X, G1217X, R1334X, and T1336X mutations, or a corresponding mutation in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of the D1134V, G1217R, R1334Q, and T1336R mutations, or a corresponding mutation in any of the amino acid sequences provided herein.In some embodiments, the SpCas9 domain comprises the D1134V, G1217R, R1334Q, and T1336R mutations, or a corresponding plurality of mutations in any of the amino acid sequences provided herein.
[0592] In some embodiments, the amino acid sequence comprised by the Cas9 domain of any of the fusion proteins provided herein is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the Cas9 polypeptide described herein. In some embodiments, the Cas9 domain of any of the fusion proteins provided herein comprises the amino acid sequence of any of the Cas9 polypeptides described herein. In some embodiments, the Cas9 domain of any of the fusion proteins provided herein consists of the amino acid sequence of any of the Cas9 polypeptides described herein.
[0593] In some examples, the PAM recognized by the CRISPR protein-derived domain of the base editor disclosed herein can be provided to the cell, and the PAM is on an oligonucleotide separate from the insert encoding the foregoing base editor (e.g., an AAV insert). In such embodiments, providing the PAM on a separate oligonucleotide can allow cleavage of a target sequence that otherwise would not be cleaved because there is no adjacent PAM on the same polynucleotide as the target sequence.
[0594] In one embodiment, Streptococcus pyogenes Cas9 (SpCas9) can be used as a CRISPR endonuclease suitable for genome engineering. However, other proteins can be used. In some embodiments, different endonucleases can be used to target certain genomic targets. In some embodiments, synthetic SpCas9-derived variants with non-NGG PAM sequences can be used. Additionally, other Cas9 orthologs from various species have been identified, and these "non-SpCas9s" can bind a variety of PAM sequences that can also be used in the present disclosure. For example, the relatively large size of SpCas9 (approximately 4 kilobases (kb) coding sequence) may result in inefficient expression of plasmids carrying the SpCas9 cDNA in cells. In contrast, the coding sequence of Staphylococcus aureus Cas9 (SaCas9) is approximately 1 kb shorter than SpCas9, which may allow for its efficient expression in cells. Similar to SpCas9, the SaCas9 endonuclease is capable of modifying target genes in vitro (in mammalian cells) and in vivo (in mice). In some embodiments, the Cas protein can target different PAM sequences. In some embodiments, for example, the target gene can be adjacent to a Cas9 PAM (5'-NGG, for example). In other embodiments, other Cas9 orthologs can have different PAM requirements. For example, other PAMs, such as those of Streptococcus thermophilus (5'-NNAGAA for CRISPR1 and 5'-NGGNG for CRISPR3) and Neisseria meningitidis (5'-NNNNGATT), can also be found adjacent to the target gene.
[0595] In some embodiments, for the Streptococcus pyogenes system, the target gene sequence can precede (i.e., be 5' to) the 5'-NGG PAM, and the 20-nt guide RNA sequence can base pair with the opposite strand to mediate Cas9 cleavage adjacent to the PAM. In some embodiments, the adjacent cleavage can be 3 base pairs upstream of the PAM or can be approximately 3 base pairs upstream of the PAM. In some embodiments, the adjacent cleavage can be 10 base pairs upstream of the PAM or can be approximately 10 base pairs upstream of the PAM. In some embodiments, the adjacent cleavage can be 0-20 base pairs upstream of the PAM or can be approximately 0-20 base pairs upstream of the PAM. For example, the adjacent cleavage can be immediately adjacent, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 base pairs upstream of the PAM. The adjacent cleavage can also be 1 to 30 base pairs downstream of the PAM. Sequences of exemplary SpCas9 proteins that can bind to the PAM sequence are as follows:
[0596] The amino acid sequences of exemplary PAM-binding SpCas9 are as follows:
[0597] MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFI
[0598]
[0599] The amino acid sequence o...
Claims
1. A cytidine base editor comprising (i) a polynucleotide-programmable DNA-binding domain and (ii) a cytidine deaminase, wherein the cytidine base editor has an increased cis-activity to trans-activity ratio (cis:trans) compared to a standard cytidine base editor.
2. A molecular complex comprising the cytidine base editor according to claim 1 and one or more guide RNA sequences, tracrRNA sequences, or target DNA sequences.
3. A method of editing a nucleobase of a nucleic acid sequence, comprising contacting the nucleic acid sequence with the cytidine base editor according to claim 1 and converting a first nucleobase of the DNA sequence to a second nucleobase.
4. A fusion protein comprising a polynucleotide-programmable DNA-binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase is (i) APOBEC-1 from Mesocricetus auratus (MaAPOBEC-1), Pongo pygmaeus (PpAPOBEC-1), Oryctolagus cuniculus (OcAPOBEC-1), Monodelphis domestica (MdAPOBEC-1), or Alligator mississippiensis (AmAPOBEC-1); (ii) APOBEC-2 from Pongo pygmaeus (PpAPOBEC-2), Bos taurus (BtAPOBEC-2), or Sus scrofa (SsAPOBEC-2); (iii) APOBEC-4 from Macaca fascicularis (MfAPOBEC-4); (iv) AID from Canis lupus familiaris (ClAID) or Bos taurus (BtAID); (v) yeast cytosine deaminase (yCD) from Saccharomyces cerevisiae; (vi) APOBEC-3F from Rhinopithecus roxellana (RrA3F); or (vii) a cytidine deaminase having an amino acid sequence with at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to any one of the proteins in (i) to (viii).
5. A fusion protein comprising a polynucleotide-programmable DNA-binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase is selected from the group consisting of members of the APOBEC2 family, members of the APOBEC3 family, members of the APOBEC4 family, members of the cytidine deaminase 1 family (CDA1), members of the A3A family, members of the RrA3F family, members of the PmCDA1 family, and members of the FENRY family.
6. A fusion protein comprising a polynucleotide-programmable DNA-binding domain and at least one nucleobase editor domain comprising APOBEC1, wherein the APOBEC1 is selected from the group consisting of ppAPOBEC1, AmAPOBEC1 (BEM3.31), ocAPOBEC1, SsAPOBEC2 (BEM3.39), hAPOBEC3A, maAPOBEC1, and mdAPOBEC1.
7. A fusion protein comprising a polynucleotide programmable DNA binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase comprises one or more alterations located at positions R15X, R16X, H21X, R30X, R33X, K34X, R52X, K60X, R118X, H121X, H122X, R126X, R128X, R169X, R198X, T36X, H53X, V62X, L88X, W90X, Y120X or R132X numbered as in SEQ ID NO: 1, or corresponding alterations of one or more of said alterations, wherein X is any amino acid.
8. A fusion protein comprising a polynucleotide programmable DNA binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase comprises one or more alterations located at positions Y130X and R28X numbered as in SEQ ID NO: 1, or corresponding alterations of one or more of said alterations, wherein X is any amino acid.
9. A fusion protein comprising a polynucleotide programmable DNA binding domain and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase comprises one or more alterations located at positions H122X, K34X, R33X, W90X or R128X numbered as in SEQ ID NO: 1, or corresponding alterations of one or more of said alterations, wherein X is any amino acid.
10. A fusion protein comprising a polynucleotide programmable DNA binding domain and a cytidine deaminase, wherein the cytidine deaminase comprises an amino acid sequence having at least 80% identity to the following amino acid sequence: MTSEKGPSTGDPTLRRRIESWEFDVFYDPRELRKETCLLYEIKWGMSRKIWRSSGKNTTNHVEVNFIKKFTSERRFHSSISCSITWFLSWSPCWECSQAIREFLSQHPGVTLVIYVARLFWHMDQRNRQGLRDLVNSGVTIQIMRASEYYHCWRNFVNYPPGDEAHWPQYPPLWMMLYALELHCIILSLPPCLKISRRWQNHLAFFRLHLQNCHYQTIPPHILLATGLIHPSVTWR.
Citation Information
Patent Citations
Towel rings
CN3315821D
cell phone
CN3329834D
Split inteins, conjugates and uses thereof
US20150344549A1
Method for modifying genome sequence to introduce specific mutation to targeted DNA sequence by base-removal reaction, and molecular complex used therein
US20170321210A1
AAV delivery of nucleobase editors
US20180127780A1