Compositions and methods for self-inactivation of base editors
Patent Information
- Application Number
- JP2023572544
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-05-28
- Filing Date
- 2022-05-27
- Publication Date
- 2025-06-03
AI Technical Summary
Existing gene editing technologies, such as CRISPR-Cas systems and base editors, face challenges with transient expression leading to off-target editing events, particularly when used for long-term delivery methods like adeno-associated virus (AAV) transduction, necessitating methods to inhibit or stop editing activity after successful on-target editing.
Incorporation of introns with splice acceptor or donor site modifications into polynucleotides encoding deaminase or nucleic acid programmable DNA binding protein (napDNAbp) domains within the open reading frame of base editors, which reduces or eliminates splicing of base editor mRNA, thereby reducing or eliminating expression.
The self-inactivating base editors achieve efficient and targeted genome editing with reduced off-target effects by minimizing expression over time, maintaining high editing efficiency while preventing unwanted editing activities.
Smart Images

Figure 00000241_0000 
Figure 00000241_0001 
Figure 00000241_0002
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 194,431, filed May 28, 2021, the entire contents of which are incorporated herein by reference.
[0002] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in ASCII format and is incorporated herein by reference in its entirety. This ASCII copy, created on May 27, 2022, is designated 180802_049001_PCT_SL.txt and has a size of 2,089,884 bytes. [Background technology]
[0003] The development of gene editing technology, such as the application of CRISPR-Cas system in eukaryotes and the emergence of base editing, allows genomes to be efficiently edited in a wide variety of cell types and organisms, and is rapidly developing the available methods for treating genetic diseases in humans. Although CRISPR-Cas system and base editors can be highly specific for the genome target of interest, the transient expression of genome modification tools in cells is preferred to mitigate potential off-target editing events that are more likely to occur if expression is sustained for a longer period of time. Therefore, methods for subsequently inhibiting or stopping editing activity after successful on-target editing are of widespread interest, especially when delivery methods that can result in long-term expression are used, such as adeno-associated virus (AAV) transduction, DNA transfection or other methods. Summary of the Invention
[0004] As described below, the invention features self-inactivating base editors, as well as related compositions and methods.
[0005] In one aspect, the invention of this disclosure features a polynucleotide encoding a deaminase domain or a nucleic acid programmable DNA binding protein (napDNAbp) domain, or a fragment thereof. The polynucleotide contains an intron. The intron is inserted within the open reading frame encoding the deaminase or napDNAbp, or a fragment thereof.
[0006] In another aspect, the invention of the disclosure features a polynucleotide encoding a deaminase domain or a nucleic acid programmable DNA binding protein (napDNAbp) domain open reading frame that contains an intron. The intron contains a modification at a splice acceptor or splice donor site. The modification reduces or eliminates splicing of the base editor mRNA, thereby reducing or eliminating expression of the base editor polypeptide.
[0007] In another aspect, the invention of the disclosure features a polynucleotide encoding a base editor polypeptide or a fragment thereof, the polynucleotide containing an intron, the intron being inserted within an open reading frame encoding the base editor polypeptide or a fragment thereof.
[0008] In another aspect, the invention of the disclosure features a polynucleotide containing a base editor open reading frame that contains an intron. The intron contains a modification at a splice acceptor or splice donor site that reduces or eliminates splicing of the base editor mRNA, thereby reducing or eliminating expression of the base editor polypeptide.
[0009] In another aspect, the invention of the disclosure features a polynucleotide encoding a base editor containing a nucleic acid programmable DNA binding protein (napDNAbp) domain or a deaminase domain. The polynucleotide contains an intron. The intron is inserted within the open reading frame encoding the napDNAbp domain or the deaminase domain.
[0010] In another aspect, the invention of the disclosure features a polynucleotide encoding a base editor or a fragment thereof that contains a nucleic acid programmable DNA binding protein (napDNAbp) domain and a deaminase domain. The polynucleotide contains a base editor open reading frame that contains an intron. The intron contains a modification at a splice acceptor or splice donor site. The modification reduces splicing of the base editor mRNA.
[0011] In another aspect, the invention of the disclosure features a composition that contains (i) a first polynucleotide encoding a deaminase domain and an N-terminal fragment of a nucleic acid programmable DNA binding protein (napDNAbp) domain, where the N-terminal fragment of the napDNAbp domain is fused to a split intein-N. The composition also contains (ii) a second polynucleotide encoding a C-terminal fragment of the napDNAbp domain, where the C-terminal fragment of the napDNAbp domain is fused to a split intein-C. The first polynucleotide or the second polynucleotide contains an intron, where the intron is inserted within the open reading frame of the polynucleotide.
[0012] In another aspect, the invention of the disclosure features a composition that contains (i) a first polynucleotide encoding an N-terminal fragment of a deaminase domain, where the N-terminal fragment of the deaminase domain is fused to a split intein-N. The composition also contains (ii) a second polynucleotide encoding a C-terminal fragment of a deaminase domain and a nucleic acid programmable DNA binding protein (napDNAbp) domain, where the C-terminal fragment of the deaminase domain is fused to a split intein-C. The first polynucleotide or the second polynucleotide contains an intron, where the intron is inserted within the open reading frame of the polynucleotide.
[0013] In another aspect, the invention of the disclosure features a base editor system that contains (i) a polynucleotide encoding a base editor containing a deaminase domain or a fragment thereof. The base editor system also contains (ii) one or more guide RNAs that direct the base editor to edit a site in the genome of a cell. The base editor system further contains (iii) one or more guide RNAs that direct the base editor to edit a polynucleotide encoding the base editor. The editing results in a decrease in activity and / or expression of the encoded base editor.
[0014] In another aspect, the invention of the disclosure features a base editor system containing (i) a polynucleotide encoding a self-inactivating base editor or a fragment thereof, where the polynucleotide contains an intron inserted within an open reading frame of the self-inactivating base editor or fragment thereof. The base editor system further contains (ii) one or more guide RNAs that direct the self-inactivating base editor to edit a site within the genome of a cell. The base editor system also contains (iii) one or more guide RNAs that direct the self-inactivating base editor to edit a splice acceptor or splice donor site present within an intron of the polynucleotide encoding the self-inactivating base editor.
[0015] In another aspect, the invention of the disclosure features a base editor system that contains (i) a polynucleotide of any one of the above aspects that encodes a base editor. The base editor system also contains (ii) one or more guide RNAs that direct the base editor to edit a site in the genome of a cell. The base editor system further contains (iii) one or more guide RNAs that direct the base editor to edit a splice acceptor or splice donor site present in an intron of the polynucleotide encoding the base editor.
[0016] In another aspect, the invention of the present disclosure features a base editor system that contains (i) a composition of any of the above aspects encoding a base editor. The base editor system further contains (ii) one or more guide RNAs that direct the base editor to edit a site in the genome of a cell. The base editor system also contains (iii) one or more guide RNAs that direct the base editor to edit a splice acceptor or splice donor site present in an intron of the composition of (i).
[0017] In another aspect, the invention of the disclosure features a base editor system that contains (i) a first polynucleotide encoding a deaminase domain and an N-terminal fragment of a nucleic acid programmable DNA binding protein (napDNAbp) domain, where the N-terminal fragment of the napDNAbp domain is fused to a split intein-N. The base editor system also contains (ii) a second polynucleotide encoding a C-terminal fragment of the napDNAbp domain, where the C-terminal fragment of the napDNAbp domain is fused to a split intein-C. The first polynucleotide or the second polynucleotide contains an intron, where the intron is inserted within the open reading frame, and the first polynucleotide and the second polynucleotide encode a base editor. The base editor system further contains (iii) one or more guide RNAs that direct the base editor to edit a site in the genome of a cell. The base editor system also contains (iv) one or more guide RNAs that direct the base editor to edit a splice acceptor or splice donor site present within an intron of the polynucleotide of (i) or (ii).
[0018] In another aspect, the invention of the disclosure features a base editor system that contains (i) a first polynucleotide encoding an N-terminal fragment of a deaminase domain, where the N-terminal fragment of the deaminase domain is fused to a split intein-N. The base editor system also contains (ii) a second polynucleotide encoding a C-terminal fragment of a deaminase domain and a nucleic acid programmable DNA binding protein (napDNAbp) domain, where the C-terminal fragment of the deaminase domain is fused to a split intein-C. The first polynucleotide or the second polynucleotide contains an intron, where the intron is inserted within an open reading frame, and the first polynucleotide and the second polynucleotide encode a base editor. The base editor system also contains (iii) one or more guide RNAs that direct the base editor to edit a site in the genome of a cell. The base editor system also contains (iv) one or more guide RNAs that direct the base editor to edit a splice acceptor or splice donor site present within an intron of the polynucleotide of (i) or (ii).
[0019] In another aspect, the invention of this disclosure features a vector containing a polynucleotide encoding a self-inactivating base editor or a fragment thereof, the polynucleotide containing an intron inserted within the open reading frame of the self-inactivating base editor or fragment thereof.
[0020] In another aspect, the invention of this disclosure features a vector containing a polynucleotide of any of the above aspects, or an embodiment thereof, or a base editor system of any of the above aspects, or an embodiment thereof.
[0021] In another aspect, the invention of this disclosure features a vector containing the first polynucleotide and / or the second polynucleotide of the composition of any one of the above aspects.
[0022] In another aspect, the invention of this disclosure features a cell containing a vector containing a polynucleotide encoding a self-inactivating base editor or a fragment thereof, the polynucleotide containing an intron inserted within the open reading frame of the self-inactivating base editor or fragment thereof.
[0023] In another aspect, the invention of the disclosure features a cell containing a polynucleotide of any of the above aspects, or an embodiment thereof, a composition of any of the above aspects, or an embodiment thereof, a base editor system of any of the above aspects, or an embodiment thereof, or a vector of any of the above aspects, or an embodiment thereof.
[0024] In another aspect, the invention of the present disclosure features a pharmaceutical composition containing a polynucleotide of any of the above aspects, or an embodiment thereof, a base editor system of any of the above aspects, or an embodiment thereof, a vector of any of the above aspects, or an embodiment thereof, or a cell of any of the above aspects, or an embodiment thereof.
[0025] In another aspect, the invention of this disclosure features a kit containing a polynucleotide, composition, base editor system, vector, cell, or pharmaceutical composition of any of the above aspects, or an embodiment thereof.
[0026] In another aspect, the invention of the disclosure features a method for reducing or eliminating expression of a self-inactivating base editor. The method includes (a) providing a polynucleotide encoding a self-inactivating base editor or a fragment thereof, where the polynucleotide contains an intron inserted within an open reading frame of the self-inactivating base editor or fragment thereof. The method also includes (b) contacting the polynucleotide with a guide RNA and a self-inactivating base editor polypeptide, where the guide RNA directs the base editor to edit a splice acceptor or splice donor site of the intron, thereby generating a modification that reduces or eliminates expression of the self-inactivating base editor.
[0027] In another aspect, the invention of the disclosure features a method of self-inactivating base editing. The method includes (a) expressing in a cell a polynucleotide encoding a base editor or a fragment thereof that contains a deaminase domain. The method also includes (b) contacting the cell with a first guide RNA, where the first guide RNA directs the base editor to edit a site in the genome of the cell, thereby generating a modification in the genome of the cell. The method further includes (c) contacting the cell with a second guide RNA, where the second guide RNA directs the base editor to edit a polynucleotide encoding the base editor, where the editing results in a decrease in activity and / or expression of the encoded base editor, thereby generating a modification that reduces or eliminates expression of the base editor.
[0028] In another aspect, the invention of the disclosure features a method of self-inactivating base editing. The method includes (a) expressing in a cell a polynucleotide encoding a self-inactivating base editor or a fragment thereof, where the polynucleotide contains an intron inserted within an open reading frame of the self-inactivating base editor or fragment thereof. The method also includes (b) contacting the cell with a first guide RNA, where the first guide RNA directs the self-inactivating base editor to edit a site in the genome of the cell, thereby generating a modification in the genome of the cell. The method further includes (c) contacting the cell with a second guide RNA, where the second guide RNA directs the self-inactivating base editor to edit a splice acceptor or splice donor site present within an intron of the polynucleotide of (a), thereby generating a modification that reduces or eliminates expression of the self-inactivating base editor.
[0029] In another aspect, the invention of the disclosure features a method of editing the genome of an organism. The method includes (a) expressing in a cell of the organism a polynucleotide encoding a self-inactivating base editor or a fragment thereof, where the polynucleotide contains an intron inserted within an open reading frame of the self-inactivating base editor or fragment thereof. The method also includes (b) contacting the cell with a first guide RNA, where the first guide RNA directs the self-inactivating base editor to edit a site in the genome of the cell, thereby generating a modification in the genome of the cell. The method further includes (c) contacting the cell with a second guide RNA, where the second guide RNA directs the self-inactivating base editor to edit a splice acceptor or splice donor site present within an intron of the polynucleotide of (a), thereby generating a modification that reduces or eliminates expression of the self-inactivating base editor.
[0030] In another aspect, the invention of the disclosure features a method of treating a subject. The method includes (a) expressing in a cell of the subject a polynucleotide encoding a self-inactivating base editor or a fragment thereof, where the polynucleotide contains an intron inserted within an open reading frame of the self-inactivating base editor or fragment thereof. The method further includes (b) contacting the cell with a first guide RNA, where the first guide RNA directs the self-inactivating base editor to edit a site in the genome of the cell, thereby generating a modification in the genome of the cell to treat the subject. The method also includes (c) contacting the cell with a second guide RNA, where the second guide RNA directs the self-inactivating base editor to edit a splice acceptor or splice donor site present in an intron of the polynucleotide of (a), thereby generating a modification that reduces or eliminates expression of the self-inactivating base editor.
[0031] In another aspect, the invention of this disclosure features a method of treating a subject, the method involving administering to the subject a base editor system, vector, cell, or pharmaceutical composition of any of the above aspects, or an embodiment thereof, thereby treating the subject.
[0032] In another aspect, the invention of the disclosure features a method of editing the genome of an organism. The method includes (a) expressing in a cell of the organism a first polynucleotide encoding a deaminase domain and an N-terminal fragment of a nucleic acid programmable DNA binding protein (napDNAbp) domain, where the N-terminal fragment of the napDNAbp domain is fused to a split intein-N, and a second polynucleotide encoding a C-terminal fragment of the napDNAbp domain, where the C-terminal fragment of the napDNAbp domain is fused to a split intein-C. The first polynucleotide or the second polynucleotide contains an intron. The intron is inserted within an open reading frame. Expression of the first polynucleotide and the second polynucleotide in the cell results in the formation of a self-inactivating base editor. The method also includes (b) contacting the cell with a first guide RNA, where the first guide RNA directs the self-inactivating base editor to edit a site within the genome of the cell, thereby generating a modification in the genome of the cell. The method also includes (c) contacting the cell with a second guide RNA, where the second guide RNA directs the self-inactivating base editor to edit a splice acceptor or splice donor site present within an intron of the polynucleotide of (a), thereby generating a modification that reduces or eliminates expression of the self-inactivating base editor.
[0033] In another aspect, the invention of the disclosure features a method for editing the genome of an organism. The method includes (a) expressing in a cell of the organism a first polynucleotide encoding an N-terminal fragment of a deaminase domain, the N-terminal fragment of the deaminase domain being fused to a split intein-N, and a second polynucleotide encoding a C-terminal fragment of the deaminase domain and a nucleic acid programmable DNA binding protein (napDNAbp) domain, the C-terminal fragment of the deaminase domain being fused to a split intein-C. The first polynucleotide or the second polynucleotide contains an intron, the intron being inserted within the open reading frame. Expression of the first polynucleotide and the second polynucleotide in the cell results in the formation of a self-inactivating base editor. The method also includes (b) contacting the cell with a first guide RNA, where the first guide RNA directs the self-inactivating base editor to edit a site within the genome of the cell, thereby generating a modification in the genome of the cell. The method further includes (c) contacting the cell with a second guide RNA, where the second guide RNA directs the self-inactivating base editor to edit a splice acceptor or splice donor site present within an intron of the polynucleotide of (a), thereby generating a modification that reduces or eliminates expression of the self-inactivating base editor.
[0034] In any of the above aspects, or embodiments thereof, the base editor has high editing efficiency in genomic DNA. In any of the above aspects, or embodiments thereof, the base editor contains a nucleic acid programmable DNA binding protein (napDNAbp) domain or a deaminase domain.
[0035] In any of the above aspects, or embodiments thereof, the deaminase domain is a cytidine deaminase domain or an adenosine deaminase domain. In any of the above aspects, or embodiments thereof, the deaminase domain is a TadA domain.
[0036] In any of the above aspects or embodiments thereof, the napDNAbp domain is a Cas domain selected from one or more of Cas9, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ domains.
[0037] In any of the above aspects or embodiments thereof, the intron is derived from a sequence selected from one or more of NF1, PAX2, EEF1A1, HBB, IGHG1, SLC50A1, ABCB11, BRSK2, PLXNB3, TMPRSS6, IL32, ANTXRL, PKHD1L1, PADI1, KRT6C, and HMCN2. In any of the above aspects or embodiments thereof, the intron is derived from NF1. In any of the above aspects or embodiments thereof, the intron is derived from PAX2. In any of the above aspects or embodiments thereof, the intron is derived from EEF1A1. In any of the above aspects or embodiments thereof, the intron is derived from HBB. In any of the above aspects or embodiments thereof, the intron is derived from IGHG1. In any of the above aspects or embodiments thereof, the intron is derived from SLC50A1. In any of the above aspects or embodiments thereof, the intron is derived from ABCB11. In any of the above aspects or embodiments thereof, the intron is derived from BRSK2. In any of the above aspects or embodiments thereof, the intron is derived from PLXNB3. In any of the above aspects or embodiments thereof, the intron is derived from TMPRSS6. In any of the above aspects or embodiments thereof, the intron is derived from IL32. In any of the above aspects or embodiments thereof, the intron is derived from PKHD1L1. In any of the above aspects or embodiments thereof, the intron is derived from PADI1. In any of the above aspects or embodiments thereof, the intron is derived from KRT6C. In any of the above aspects or embodiments thereof, the intron is derived from HMCN2. In any of the above aspects or embodiments thereof, the intron has at least about 85% nucleic acid sequence identity to an intron naturally occurring in a mammalian gene.In any of the above aspects, or embodiments thereof, the intron has at least about 85% nucleic acid sequence identity to an intron naturally occurring in the non-mammalian gene. In any of the above aspects, or embodiments thereof, the intron is a synthetic intron. In any of the above aspects, or embodiments thereof, the intron contains a sequence having at least about 85% nucleic acid sequence identity to one of the following: a) GTGAGATCAAATGAAAGTTTCATATAGAAATACAAAACCTAGAGAACTGGCATGTAAGAGAAGCAAAAATTACTTCAGCAAGGCCATGTTAGTAAATTTGCATCTGTTTGTCCACATTAG (SEQ ID NO: 226), b) GTAGGTGACAATGCTGCAGCTGCCTAATCTAGGTGGGGGGAACTAAATTGTGGGTGAGCTGCTGAATGGTCTGTAGTCTGAGGCTGGGGTGGGGGGAGACACAACGTCCCCTCCCTGCAAACCACTGCTATTCTGTCCCTCTCTCCTTAG (SEQ ID NO: 227), c) GTAAGTGGCTTTCAAGACCATTGTTAAAAAGCTCTGGGAATGGCGATTTCATGCTTACATAAATTGGCATGCTTGTGTTTCAG (SEQ ID NO: 228); d) GTAAGTATCAAGGTTACAAGACAGGTTTAAGGAGACCAATAGAAACTGGGCTTGTCTAGACAGAGAAGACTCTTGCGTTTCTGATAGGCACCTATTGGTCTTACTGACATCCACTTTGCCTTTCTCTCCACAG (SEQ ID NO: 229), e) GTAAGCACAACTGGGATGGGGTGACAGGGGTGCAAGATTGAAAACTGGCTCCTCTCCTCATAGCAGTTCTTGTGATTTCAG (SEQ ID NO: 230), f) GTAAGAAATGTTATTTTTCAGTAAGTGATTTAGTTATTTTTCCTTTTTTCTCATTAAAATTTCTCTAACATCTCCCTCTTCATGTTTTAG (SEQ ID NO: 231), g) GTGAGACCCTAGCCCCCTCAACCCTGCCCTGGCCTCTCCCCAAACCTGCCCCCCACGCTGACCCCCACACCCGGCCGCCCGCAG (SEQ ID NO: 232), h) GTGGGTGTCAGAGGCATCGGGCTGCGGGGTAGGGGGCTGCCCCACCCCTAACGAAGTCTGCTCCTCCAG (SEQ ID NO: 233), i) GCAGGGAAGTCCTGCTTCCGTGCCCCACCGGTGCTCAGCTGAGGCTCCCTTGAAAATGCGAGGCTGTTTCCAACTTTGGTCTGTTTCCCTGGCAG (SEQ ID NO: 234), j) GTGGGAGGTTGGGGTCCCCGAAGGTGAGGACCCTCTGGGGATGAGGGTGCTTCTCTGAGACACTTTCTTTTCCTACACCTGTTCCTCGCCAGCAG (SEQ ID NO: 235), k) GTATAGACCCCTTGATCTCCTAACCCTAACCCTAACCCTAACCCTAACCCTAACCTACAAAATCTTAGAGCATCAGTGGGAGCATCTCACTGTCCAGGCTCAATATTTCTTCATTTTCTTGCAG (SEQ ID NO: 236), l) GTAATTATGATAAAGATGGTGATTGTTTATTTTCTTTTATGATTGTCCTTAGTATTATGTAACCTGCAAATTCTATTGCAG (SEQ ID NO: 237), m) GTGAGTGACACAAGGTGTTGTCTGGGGAGTGGGGAAGGGGGATGGAAGTGAATCCTGTTGGTGGGGTGGAGAAAGGGCGATCTCAAGAGGGCCACTCTCTCCAG (SEQ ID NO: 238), n) GTAAGCATCTCCACCATCCTTCTGTTTACTCTGATGGGGTCTGCAAAGGGGAGATGATGTATAGGGTTGGGTATCCTGTAAATGTCAGATGTGAAGTTGATCTTATGACCTTCTGTTCTGCAG (SEQ ID NO: 239), o) GTGAGGGTCTCCCAGGCTGGGCAGGGGGAGGGGGCTGCTGCCTTGATTGCGTCCCAGGACACAGCCCTCCTCCAGCTGCCCTCGCCTTGCTCATCCCCTCCCCATCTCAGCCCCCCCCACTAACTCTCTCTCTGCTCTGACTCAG (SEQ ID NO: 240), p) GTAATGATGATTGCAATGTATGATTACAATAATCTCAGTATAAGTTCAGTAATAATAACCTTCCACTGCTGTCCTCTGTGTGCACCCAG (SEQ ID NO: 241), or q) GTAAATATATACAACAGTTTTTCATTTAAATAAGTGCACGGCACAAATAAGAAAAATATGTCAAAAATGTAACCAATAGTTTTTTTCAAATTTAG (SEQ ID NO: 242). In any of the above aspects, or embodiments thereof, the intron contains a nucleic acid sequence from one of the following: a) GTGAGATCAAATGAAAGTTTCATATAGAAATACAAAACCTAGAGAACTGGCATGTAAGAGAAGCAAAAATTACTTCAGCAAGGCCATGTTAGTAAATTTGCATCTGTTTGTCCACATTAG (SEQ ID NO: 226), b) GTAGGTGACAATGCTGCAGCTGCCTAATCTAGGTGGGGGGAACTAAATTGTGGGTGAGCTGCTGAATGGTCTGTAGTCTGAGGCTGGGGTGGGGGGAGACACAACGTCCCCTCCCTGCAAACCACTGCTATTCTGTCCCTCTCTCCTTAG (SEQ ID NO: 227), c) GTAAGTGGCTTTCAAGACCATTGTTAAAAAGCTCTGGGAATGGCGATTTCATGCTTACATAAATTGGCATGCTTGTGTTTCAG (SEQ ID NO: 228), d) GTAAGTATCAAGGTTACAAGACAGGTTTAAGGAGACCAATAGAAACTGGGCTTGTCTAGACAGAGAAGACTCTTGCGTTTCTGATAGGCACCTATTGGTCTTACTGACATCCACTTTGCCTTTCTCTCCACAG (SEQ ID NO: 229), e) GTAAGCACAACTGGGATGGGGTGACAGGGGTGCAAGATTGAAAACTGGCTCCTCTCCTCATAGCAGTTCTTGTGATTTCAG (SEQ ID NO: 230), f) GTAAGAAATGTTATTTTTCAGTAAGTGATTTAGTTATTTTTCCTTTTTTCTCATTAAAATTTCTCTAACATCTCCCTCTTCATGTTTTAG (SEQ ID NO: 231), g) GTGAGACCCTAGCCCCCTCAACCCTGCCCTGGCCTCTCCCCAAACCTGCCCCCCCACGCTGACCCCCACACCCGGCCGCCCGCAG (SEQ ID NO: 232), h) GTGGGTGTCAGAGGCATCGGGGCTGCGGGGTAGGGGGCTGCCCCACCCCTAACGAAGTCTGCTCCTCCAG (SEQ ID NO: 233), i) GCAGGGAAGTCCTGCTTCCGTGCCCCACCGGTGCTCAGCTGAGGCTCCCTTGAAAATGCGAGGCTGTTTCCAACTTTGGTCTGTTTCCCTGGCAG (SEQ ID NO: 234), j) GTGGGGAGTTGGGGTCCCCGAAGGTGAGGACCCTCTGGGGATGAGGGTGCTTCTCTGAGACACTTTCTTTTCCTCACACCTGTTCCTCGCCAGCAG (SEQ ID NO: 235), k) GTATAGACCCCTTGATCTCCTAACCCTAACCCTAACCCTAACCCTAACCCTAACCTACAAAATCTTAGAGCATCAGTGGGAGCATCTCACTGTCCAGGCTCAATATTTCTTCATTTTCTTGCAG (SEQ ID NO: 236), l) GTAATTATGATAAAGATGGTGATTGTTTATTTTCTTTTATGATTGTCCTTAGTATTATGTAACCTGCAAATTCTATTGCAG (SEQ ID NO: 237), m) GTGAGTGACACAAGGTGTTGTCTGGGGAGTGGGGAAGGGGGATGGAAGTGAATCCTGTTGGTGGGGTGGAGAAAGGGCGATCTCAAGAGGGCCACTCTCTCCAG (SEQ ID NO: 238), n) GTAAGCATCTCCACCATCCTTCTGTTTACTCTGATGGGGTCTGCAAAGGGGAGATGATGTATAGGGTTGGGTATCCTGTAAATGTCAGATGTGAAGTTGATCTTATGACCTTCTGTTCTGCAG (SEQ ID NO: 239), o) GTGAGGGTCTCCCAGGCTGGGCAGGGGGAGGGGGCTGCTGCCTTGATTGCGTCCCAGGACACAGCCCTCCTCCAGCTGCCCTCGCCTTGCTCATCCCCTCCCCATCTCAGCCCCCCCCACTAACTCTCTCTCTGCTCTGACTCAG (SEQ ID NO: 240), p) GTAATGATGATTGCAATGTATGATTACAATAATCTCAGTATAAGTTCAGTAATAATAACCTTCCACTGCTGTCCTCTGTGTGCACCCAG (SEQ ID NO: 241), or q) GTAAATATATACAACAGTTTTTCATTTAAATAAGTGCACGGCACAAATAAGAAAAATATGTCAAAAATGTAACCAATAGTTTTTTTCAAATTTAG (SEQ ID NO: 242).
[0038] In any of the above aspects, or embodiments thereof, the intron contains about 10 base pairs to about 500 base pairs. In any of the above aspects, or embodiments thereof, the intron contains about 70 base pairs to 150 base pairs. In any of the above aspects, or embodiments thereof, the intron contains about 100 base pairs to 200 base pairs. In any of the above aspects, or embodiments thereof, the intron is inserted adjacent to the protospacer sequence. In any of the above aspects, or embodiments thereof, the intron is inserted within about 10 to 30 base pairs of the protospacer sequence. In any of the above aspects, or embodiments thereof, the protospacer sequence is NGG or NNGRRT.
[0039] In any of the above aspects or embodiments thereof, the deaminase domain contains a TadA domain.
[0040] In any of the above aspects, or embodiments thereof, the intron is inserted within or immediately after codon 18, 23, 59, 62, 87, or 129 of TadA. In any of the above aspects, or embodiments thereof, the intron is inserted immediately after codon 87 of TadA. In any of the above aspects, or embodiments thereof, the modification is a single base edit. In any of the above aspects, or embodiments thereof, the single base edit is an A-to-G base edit. In any of the above aspects, or embodiments thereof, the single base edit is a C-to-T base edit.
[0041] In any of the above aspects, or embodiments thereof, the polynucleotide further contains a polynucleotide sequence encoding a linker. In any of the above aspects, or embodiments thereof, an intron is inserted within the polynucleotide sequence encoding the linker.
[0042] In any of the above aspects, or embodiments thereof, the programmable DNA binding protein domain is a Cas9 domain. In any of the above aspects, or embodiments thereof, the Cas9 domain is split into an amino acid residue corresponding to Asn309 and an amino acid residue corresponding to Thr310 of Cas9, and residue 310 is mutated to Thr310Cys.
[0043] In any of the above aspects, or embodiments thereof, the intron contains a modification at a splice acceptor or splice donor site that reduces or eliminates splicing of the base editor mRNA.
[0044] In any of the above aspects, or embodiments thereof, the napDNAbp domain is a Cas9 domain. In any of the above aspects, or embodiments thereof, the N-terminal and C-terminal domains of the Cas9 domain are divided into amino acid residues Asn309 and Thr310. In any of the above aspects, or embodiments thereof, the Cas9 domain contains the mutation Thr310Cys.
[0045] In any of the above aspects, or embodiments thereof, the composition further contains a linker polynucleotide sequence. In any of the above aspects, or embodiments thereof, the intron is inserted within the linker polynucleotide sequence.
[0046] In any of the above aspects, or embodiments thereof, the editing alters a catalytic residue of the deaminase domain. In any of the above aspects, or embodiments thereof, the deaminase domain is an adenosine deaminase domain. In any of the above aspects, or embodiments thereof, the deaminase domain is a cytidine deaminase domain. In any of the above aspects, or embodiments thereof, the modified catalytic residue of the deaminase domain is His57 (H57), Glu59 (E59), Cys87 (C87), or Cys90 (C90) of the following reference sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 1), or the corresponding position in another adenosine deaminase. In any of the above aspects, or embodiments thereof, the modified catalytic residue is E59. In any of the above aspects, or embodiments thereof, the modification to the catalytic residue is E59G. In any of the above aspects, or embodiments thereof, the modified catalytic residue is H57. In any of the above aspects, or embodiments thereof, the modification to the catalytic residue is H57R. In any of the above aspects, or embodiments thereof, the modified catalytic residue is C87. In any of the above aspects, or embodiments thereof, the modification to the catalytic residue is C87R. In any of the above aspects, or embodiments thereof, the modified catalytic residue is C90. In any of the above aspects, or embodiments thereof, the modification to the catalytic residue is C90R.
[0047] In any of the above aspects, or embodiments thereof, the base editor system contains a polynucleotide sequence selected from: a) gGUUUUAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 191), b) gUUUCUUACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 192), c) gGUUUCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 193), d) GCCACUUACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 194), e) gACAUUAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 195), f) gGAUCUCACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 196), g) gUCCUUAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 197), h) GUCACCUACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 198), i) GAUUUCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 190), j) gGUGCUUACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 200), k) gUCCACAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 201), l) GAUACUUACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 202), m) gUGUUUUAGCUGCGGCAAGGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 203), n) gUUUCUUACAGCCAUAAUUUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 204), o) gCUCCACAGCUGCGGCAAGGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 205), p) GAUACUUACAGCCAUAAUUUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 206), q) gUGUUUUAGGGACGAAAGAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 207), r) gUUACCUGGCUCUCUUAGCCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 208), s) gCUCCACAGGGACGAAAGAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 209), t) gCUUGCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 210), u) gAUUGCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 211), v) gUCUCCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 212), w) gUCUGCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 213), x) gGACUCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 214), y) GCACCCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 215), z) gAAUUUAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 216), aa) gCAUUAGGUCGAGAUCACAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 217), bb) gCCUUAGGUCGAGAUCACAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 218), cc) GUUUCAGGUCGAGAUCACAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 219), dd) gACAUUAGGCUAAGAGAGCCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 220), ee) gUCCUUAGGCUAAGAGAGCCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 221), ff) gGUUUCAGGCUAAGAGAGCCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 222), gg) gACAUUAGAUUAUGGCUCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 223), hh) gUCCUUAGAUUAUGGCUCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 224), ii) gGUUUCAGAUUAUGGCUCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 225), jj) gCACCAUGAGCGAGGUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 524), kk) gGCCACCAUGAGCGAGGUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 525), ll) GUGUCGAAGUUCGCCCUGGAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 526), mm) gAUGCCGAGAUAAUGGCCCUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 527), nn) gAUGCCGAGAUAAUGGCCCUUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 528), oo) gAUGCCGAGAUCAUGGCACUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 529), pp) gAUGCCGAGAUCAUGGCACUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 530), qq) gAUGCCGAGAUCAUGGCACUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 531), rr) gAUGCCGAGAUCAUGGCGCUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 532), ss) gAUGCCGAGAUCAUGGCGCUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 533), tt) gAUGCCGAGAUCAUGGCGUUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 534), uu) gAUGCCGAGAUUAUGGCACUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 535), vv) gAUGCCGAGAUUAUGGCACUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 536), ww) gAUGCCGAGAUUAUGGCACUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 537), xx) gAUGCCGAGAUUAUGGCACUUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 538), yy) gAUGCCGAGAUUAUGGCGCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 539), zz) gAUGCCGAGAUUAUGGCUCUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 540), aaa) gAUGCGGAGAUCAUGGCGCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 541), bbb) gAUGCUGAGAUAAUGGCCCUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 542), ccc) gAACCGCACAUGCCGAAAUUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 543), ddd) gGCAGGUGUCGACAUAUCUAUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 544), eee) gAUGCCGAAAUUAUGGCUCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 545), fff) gACACAUGACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 546), or ggg)gGCCCCAGCACACAUGACACAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (sequence number 547).
[0048] In any of the above aspects, or embodiments thereof, the expression vector is a mammalian expression vector. In any of the above aspects, or embodiments thereof, the vector is a lipid nanoparticle. In any of the above aspects, or embodiments thereof, the vector is a viral vector selected from one or more of adeno-associated virus (AAV), retroviral vector, adenoviral vector, lentiviral vector, Sendai virus vector, and herpes virus vector. In any of the above aspects, or embodiments thereof, the vector is an AAV vector. In any of the above aspects, or embodiments thereof, the AAV vector is AAV2 or AAV8. In any of the above aspects, or embodiments thereof, the vector contains a promoter. In any of the above aspects, or embodiments thereof, the promoter is a CMV promoter.
[0049] In any of the above aspects or embodiments thereof, the cell is in vitro or in vivo.
[0050] In any of the above aspects, or embodiments thereof, the composition or pharmaceutical composition further comprises a pharma- ceutically acceptable excipient, diluent, or carrier.
[0051] In any of the above aspects, or embodiments thereof, the kit contains instructions for use in the methods according to any of the above aspects, or embodiments thereof.
[0052] In any of the above aspects or embodiments thereof, the method is performed in vivo.
[0053] In any of the above aspects, or embodiments thereof, the first polynucleotide and / or the second polynucleotide are expressed in the cell by a vector. In any of the above aspects, or embodiments thereof, the first polynucleotide and the second polynucleotide are expressed in the cell by separate vectors. In any of the above aspects, or embodiments thereof, the first guide RNA and / or the second guide RNA are delivered to the cell by a vector. In any of the above aspects, or embodiments thereof, the first guide RNA and / or the second guide RNA are delivered to the cell in the same vector as the first polynucleotide and / or the second polynucleotide. In any of the above aspects, or embodiments thereof, the first guide RNA and / or the second guide RNA are delivered to the cell in a different vector than the first polynucleotide and / or the second polynucleotide. In any of the above aspects, or embodiments thereof, the vector is a viral vector.
[0054] In any of the above aspects, or embodiments thereof, the base editor contains a nucleic acid programmable DNA binding protein (napDNAbp) domain and a deaminase domain. In any of the above aspects, or embodiments thereof, the open reading frame containing an intron is in the napDNAbp domain or the deaminase domain.
[0055] In any of the above aspects, or embodiments thereof, the self-inactivating base editor polypeptide maintains high editing efficiency in genomic DNA. In any of the above aspects, or embodiments thereof, the deaminase domain is a cytidine deaminase domain or an adenosine deaminase domain. In any of the above aspects, or embodiments thereof, the modification is in a consensus splice donor site at the 5' end of the intron or a consensus splice acceptor sequence at the 3' end of the intron.
[0056] In any of the above aspects, or embodiments thereof, the intron contains a sequence having at least about 85%, 90%, 95%, or 99% nucleic acid sequence identity to one of the following: a) GTGAGATCAAATGAAAGTTTCATATAGAAATACAAAACCTAGAGAACTGGCATGTAAGAGAAGCAAAAATTACTTCAGCAAGGCCATGTTAGTAAATTTGCATCTGTTTGTCCACATTAG (SEQ ID NO: 226), b) GTAGGTGACAATGCTGCAGCTGCCTAATCTAGGTGGGGGGAACTAAATTGTGGGTGAGCTGCTGAATGGTCTGTAGTCTGAGGCTGGGGTGGGGGGAGACACAACGTCCCCTCCCTGCAAACCACTGCTATTCTGTCCCTCTCTCCTTAG (SEQ ID NO: 227), c) GTAAGTGGCTTTCAAGACCATTGTTAAAAAGCTCTGGGAATGGCGATTTCATGCTTACATAAATTGGCATGCTTGTGTTTCAG (SEQ ID NO: 228); d) GTAAGTATCAAGGTTACAAGACAGGTTTAAGGAGACCAATAGAAACTGGGCTTGTCTAGACAGAGAAGACTCTTGCGTTTCTGATAGGCACCTATTGGTCTTACTGACATCCACTTTGCCTTTCTCTCCACAG (SEQ ID NO: 229), e) GTAAGCACAACTGGGATGGGGTGACAGGGGTGCAAGATTGAAAACTGGCTCCTCTCCTCATAGCAGTTCTTGTGATTTCAG (SEQ ID NO: 230), f) GTAAGAAATGTTATTTTTCAGTAAGTGATTTAGTTATTTTTCCTTTTTTCTCATTAAAATTTCTCTAACATCTCCCTCTTCATGTTTTAG (SEQ ID NO: 231), g) GTGAGACCCTAGCCCCCTCAACCCTGCCCTGGCCTCTCCCCAAACCTGCCCCCCACGCTGACCCCCACACCCGGCCGCCCGCAG (SEQ ID NO: 232), h) GTGGGTGTCAGAGGCATCGGGCTGCGGGGTAGGGGGCTGCCCCACCCCTAACGAAGTCTGCTCCTCCAG (SEQ ID NO: 233), i) GCAGGGAAGTCCTGCTTCCGTGCCCCACCGGTGCTCAGCTGAGGCTCCCTTGAAAATGCGAGGCTGTTTCCAACTTTGGTCTGTTTCCCTGGCAG (SEQ ID NO: 234), j) GTGGGAGGTTGGGGTCCCCGAAGGTGAGGACCCTCTGGGGATGAGGGTGCTTCTCTGAGACACTTTCTTTTCCTACACCTGTTCCTCGCCAGCAG (SEQ ID NO: 235), k) GTATAGACCCCTTGATCTCCTAACCCTAACCCTAACCCTAACCCTAACCCTAACCTACAAAATCTTAGAGCATCAGTGGGAGCATCTCACTGTCCAGGCTCAATATTTCTTCATTTTCTTGCAG (SEQ ID NO: 236), l) GTAATTATGATAAAGATGGTGATTGTTTATTTTCTTTTATGATTGTCCTTAGTATTATGTAACCTGCAAATTCTATTGCAG (SEQ ID NO: 237), m) GTGAGTGACACAAGGTGTTGTCTGGGGAGTGGGGAAGGGGGATGGAAGTGAATCCTGTTGGTGGGGTGGAGAAAGGGCGATCTCAAGAGGGCCACTCTCTCCAG (SEQ ID NO: 238), n) GTAAGCATCTCCACCATCCTTCTGTTTACTCTGATGGGGTCTGCAAAGGGGAGATGATGTATAGGGTTGGGTATCCTGTAAATGTCAGATGTGAAGTTGATCTTATGACCTTCTGTTCTGCAG (SEQ ID NO: 239), o) GTGAGGGTCTCCCAGGCTGGGCAGGGGGAGGGGGCTGCTGCCTTGATTGCGTCCCAGGACACAGCCCTCCTCCAGCTGCCCTCGCCTTGCTCATCCCCTCCCCATCTCAGCCCCCCCCACTAACTCTCTCTCTGCTCTGACTCAG (SEQ ID NO: 240), p) GTAATGATGATTGCAATGTATGATTACAATAATCTCAGTATAAGTTCAGTAATAATAACCTTCCACTGCTGTCCTCTGTGTGCACCCAG (SEQ ID NO: 241), or q) GTAAATATATACAACAGTTTTTCATTTAAATAAGTGCACGGCACAAATAAGAAAAATATGTCAAAAATGTAACCAATAGTTTTTTTCAAATTTAG (SEQ ID NO: 242). In any of the above aspects, or embodiments thereof, the second guide RNA contains a polynucleotide sequence selected from: a) gGUUUUAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 191), b) gUUUCUUACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 192), c) gGUUUCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 193), d) GCCACUUACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 194), e) gACAUUAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 195), f) gGAUCUCACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 196), g) gUCCUUAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 197), h) GUCACCUACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 198), i) GAUUUCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 190), j) gGUGCUUACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 200), k) gUCCACAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 201), l) GAUACUUACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 202), m) gUGUUUUAGCUGCGGCAAGGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 203), n) gUUUCUUACAGCCAUAAUUUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 204), o) gCUCCACAGCUGCGGCAAGGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 205), p) GAUACUUACAGCCAUAAUUUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 206), q) gUGUUUUAGGGACGAAAGAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 207), r) gUUACCUGGCUCUCUUAGCCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 208), s) gCUCCACAGGGACGAAAGAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 209), t) gCUUGCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 210), u) gAUUGCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 211), v) gUCUCCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 212), w) gUCUGCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 213), x) gGACUCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 214), y) GCACCCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 215), z) gAAUUUAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 216), aa) gCAUUAGGUCGAGAUCACAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 217), bb) gCCUUAGGUCGAGAUCACAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 218), cc) GUUUCAGGUCGAGAUCACAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 219), dd) gACAUUAGGCUAAGAGAGCCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 220), ee) gUCCUUAGGCUAAGAGAGCCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 221), ff) gGUUUCAGGCUAAGAGAGCCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 222), gg) gACAUUAGAUUAUGGCUCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 223), hh) gUCCUUAGAUUAUGGCUCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 224), ii) gGUUUCAGAUUAUGGCUCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 225), jj) gCACCAUGAGCGAGGUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 524), kk) gGCCACCAUGAGCGAGGUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 525), ll) GUGUCGAAGUUCGCCCUGGAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 526), mm) gAUGCCGAGAUAAUGGCCCUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 527), nn) gAUGCCGAGAUAAUGGCCCUUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 528), oo) gAUGCCGAGAUCAUGGCACUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 529), pp) gAUGCCGAGAUCAUGGCACUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 530), qq) gAUGCCGAGAUCAUGGCACUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 531), rr) gAUGCCGAGAUCAUGGCGCUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 532), ss) gAUGCCGAGAUCAUGGCGCUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 533), tt) gAUGCCGAGAUCAUGGCGUUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 534), uu) gAUGCCGAGAUUAUGGCACUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 535), vv) gAUGCCGAGAUUAUGGCACUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 536), ww) gAUGCCGAGAUUAUGGCACUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 537), xx) gAUGCCGAGAUUAUGGCACUUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 538), yy) gAUGCCGAGAUUAUGGCGCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 539), zz) gAUGCCGAGAUUAUGGCUCUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 540), aaa) gAUGCGGAGAUCAUGGCGCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 541), bbb) gAUGCUGAGAUAAUGGCCCUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 542), ccc) gAACCGCACAUGCCGAAAUUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 543), ddd) gGCAGGUGUCGACAUAUCUAUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 544), eee) gAUGCCGAAAUUAUGGCUCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 545), fff) gACACAUGACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 546), or ggg)gGCCCCAGCACACAUGACACAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (sequence number 547).
[0057] In any of the above aspects, or embodiments thereof, the polynucleotide further contains a linker polynucleotide sequence. In any of the above aspects, or embodiments thereof, an intron is inserted within the linker polynucleotide sequence.
[0058] In any of the above aspects, or embodiments thereof, the subject or organism is a human. In any of the above aspects, or embodiments thereof, the subject or organism is a mammal. In any of the above aspects, or embodiments thereof, the mammal is a human.
[0059] definition Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by those skilled in the art to which this invention belongs. The following references provide those skilled in the art with general definitions of many terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed.1994), The Cambridge Dictionary of Science and Technology (Walker ed.,1988), The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991), and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the following meanings unless otherwise specified.
[0060] "Adenine" or "9H-purin-6-amine" has the molecular formula C 5 H 5 N 5 The structure [ka] and means the purine nucleobase corresponding to CAS number 73-24-5.
[0061] "Adenosine" or "4-amino-1-[(2R,3R,4S,5R)-3,4-dihydroxy-5-(hydroxymethyl)oxolan-2-yl]pyrimidin-2(1H)-one" is a glycosidic linkage attached to a ribose sugar and has the structure [ka] It means the adenine molecule having the formula C, which corresponds to the CAS number 65-46-3. 10 H 13 N 5 O 4It is.
[0062] "Adenosine deaminase" or "adenine deaminase" refers to a polypeptide or functional fragment thereof that can catalyze the hydrolytic deamination of adenine or adenosine. The terms "adenine deaminase" and "adenosine deaminase" are used interchangeably throughout this application. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine to inosine or deoxyadenosine to deoxyinosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminase (e.g., modified adenosine deaminase, evolved adenosine deaminase) provided herein can be from any organism (e.g., eukaryote, prokaryote), including, but not limited to, algae, bacteria, fungi, plants, invertebrates (e.g., insects), vertebrates (e.g., amphibians, mammals). In some embodiments, the adenosine deaminase is an adenosine deaminase variant having one or more modifications and capable of deaminating both adenine and cytosine in a target polynucleotide (e.g., DNA, RNA). In some embodiments, the target polynucleotide is single-stranded or double-stranded. In some embodiments, the adenosine deaminase variant is capable of deaminating both adenine and cytosine in DNA. In some embodiments, the adenosine deaminase variant is capable of deaminating both adenine and cytosine in single-stranded DNA. In some embodiments, the adenosine deaminase variant is capable of deaminating both adenine and cytosine in RNA.
[0063] "Adenosine deaminase activity" refers to catalyzing the deamination of adenine or adenosine to guanine in a polynucleotide. In some embodiments, the adenosine deaminase variants provided herein maintain adenosine deaminase activity (e.g., at least about 30%, 40%, 50%, 60%, 70%, 80%, 90% or more of the activity of a reference adenosine deaminase (e.g., TadA*8.20 or TadA*8.19)).
[0064] By "adenosine base editor (ABE)" is meant a base editor that comprises an adenosine deaminase.
[0065] By "adenosine base editor (ABE) polynucleotide" is meant a polynucleotide that encodes an ABE. By "Adenosine Base Editor 8 (ABE8) polypeptide" or "ABE8" is meant a base editor, as defined herein, including an adenosine deaminase or an adenosine deaminase variant that comprises one or more of the modifications listed in Table 14, one of the combinations of modifications listed in Table 14, or a modification at one or more of the amino acid positions listed in Table 14, where such modifications are relative to the following reference sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 1), or a corresponding position in another adenosine deaminase. In embodiments, ABE8 comprises modifications at amino acids 82 and / or 166 of SEQ ID NO:1.
[0066] In some embodiments, ABE8 comprises further modifications, as described herein, relative to the reference sequence.
[0067] By "Adenosine Base Editor 8 (ABE8) polynucleotide" is meant a polynucleotide that encodes an ABE8 polypeptide.
[0068] "Administering" is referred to herein as providing one or more compositions described herein to a patient or subject.
[0069] By "agent" is meant any small molecule chemical compound, antibody, nucleic acid molecule, or polypeptide, or fragments thereof.
[0070] "Alteration" refers to a change (increase or decrease) in the level, structure or activity of an analyte, gene or polypeptide, as detected by standard methods known in the art, such as the methods described herein. As used herein, alteration includes a 10% change in expression levels, a 25% change, a 40% change, and a 50% or greater change in expression levels. In some embodiments, alteration includes an insertion, deletion, or substitution of a nucleic acid base or amino acid.
[0071] "Ameliorate" means to reduce, suppress, attenuate, decrease, arrest, or stabilize the onset or progression of a disease.
[0072] "Analog" refers to a molecule that has similar, but not identical, functional or structural characteristics. For example, a polypeptide analog retains the biological activity of the corresponding native polypeptide while having certain biochemical modifications that enhance the function of the analog relative to the native polypeptide. Such biochemical modifications may increase the analog's protease resistance, membrane permeability, or half-life, for example, without altering ligand binding. Analogs may also contain unnatural amino acids.
[0073] "Base editor (BE)" or "nucleobase editor polypeptide (NBE)" refers to an agent that binds to a polynucleotide and has nucleobase modifying activity. In various embodiments, a base editor comprises a polynucleotide programmable nucleotide binding domain (e.g., Cas9 or Cpf1) in combination with a nucleobase modifying polypeptide (e.g., a deaminase) and a guide polynucleotide (e.g., a guide RNA (gRNA)). Representative nucleic acid and protein sequences of base editors are provided as SEQ ID NOs: 2-11 in the Sequence Listing.
[0074] "Base editing activity" means acting to chemically modify a base within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is a cytidine deaminase activity, e.g., converting a targeted C·G to T·A. In another embodiment, the base editing activity is an adenosine or adenine deaminase activity, e.g., converting A·T to G·C.
[0075] The term "base editor system" refers to an intermolecular complex for editing a nucleobase of a target nucleotide sequence. In various embodiments, the base editor (BE) system includes (1) a polynucleotide programmable nucleotide binding domain, a deaminase domain (e.g., cytidine deaminase or adenosine deaminase) for deaminating a nucleobase in a target nucleotide sequence, and (2) one or more guide polynucleotides (e.g., guide RNAs) in combination with the polynucleotide programmable nucleotide binding domain. In various embodiments, the base editor (BE) system includes a nucleobase editor domain selected from adenosine deaminase or cytidine deaminase, and a domain having a nucleic acid sequence-specific binding activity. In some embodiments, the base editor system includes (1) a base editor (BE) including a polynucleotide programmable DNA binding domain and a deaminase domain for deaminating one or more nucleobases in a target nucleotide sequence, and (2) one or more guide RNAs in combination with the polynucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable DNA binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE), or a cytidine or cytosine base editor (CBE).
[0076] The term "Cas9" or "Cas9 domain" refers to an RNA-guided nuclease that includes a Cas9 protein or a fragment thereof (e.g., a protein that includes an active, inactive, or partially active DNA cleavage domain of Cas9 and / or a gRNA binding domain of Cas9). Cas9 nuclease is also sometimes referred to as a casnl nuclease or a CRISPR (clustered regularly interspaced short palindromic repeats)-associated nuclease.
[0077] The term "conservative amino acid substitution" or "conservative mutation" refers to the replacement of one amino acid with another amino acid that has common properties. A functional method for defining common properties between individual amino acids is to analyze the normalized frequency of amino acid changes between corresponding proteins of homologous organisms (Schulz, GE and Schirmer, RH, Principles of Protein Structure, Springer-Verlag, New York (1979)). According to such an analysis, groups of amino acids can be defined where the amino acids within the group are preferentially exchanged with each other and therefore are most similar to each other in their impact on the overall protein structure (Schulz, GE and Schirmer, RH, supra). Non-limiting examples of conservative mutations include amino acid substitutions, such as amino acid substitutions of arginine to lysine, which can maintain a positive charge, and vice versa, aspartic acid to glutamic acid, which can maintain a negative charge, and vice versa, threonine to serine, which can maintain a free -OH, and vice versa, which can maintain a free -NH 2 Examples of suitable substitutions include asparagine to glutamine, which can maintain the above-mentioned structure.
[0078] The terms "coding sequence" or "protein coding sequence," as used interchangeably herein, refer to a segment of a polynucleotide that encodes a protein. A coding sequence may also be referred to as an open reading frame. A region or sequence is bounded proximal to the 5' end by a start codon and proximal to the 3' end by a stop codon. Stop codons useful in the base editors described herein include:
[0079] TIFF2024521750000003.tif47165
[0080] "Complex" refers to a combination of two or more molecules whose interactions depend on intermolecular forces. Non-limiting examples of intermolecular forces include covalent and non-covalent interactions. Non-limiting examples of non-covalent interactions include hydrogen bonds, ionic bonds, halogen bonds, hydrophobic bonds, van der Waals interactions (e.g., dipole-dipole interactions, dipole-induced dipole interactions, and London dispersion forces), and π-effects. In one embodiment, the complex comprises a polypeptide, a polynucleotide, or a combination of one or more polypeptides and one or more polynucleotides. In one embodiment, the complex comprises one or more polypeptides that associate to form a base editor (e.g., a nucleic acid programmable DNA binding protein such as Cas9, and a base editor including a deaminase) and a polynucleotide (e.g., a guide RNA). In one embodiment, the complex is held together by hydrogen bonds. It is understood that one or more components of a base editor (e.g., a deaminase, or a nucleic acid programmable DNA binding protein) may be covalently or non-covalently associated. As an example, a base editor may include a deaminase covalently linked (e.g., by a peptide bond) to a nucleic acid-programmable DNA binding protein. Alternatively, a base editor may include a deaminase and a nucleic acid-programmable DNA binding protein that are non-covalently associated (e.g., where one or more components of the base editor are provided in trans and associated directly or via another molecule, such as a protein or nucleic acid). In one embodiment, one or more components of the complex are held together by hydrogen bonds. Throughout this disclosure, whenever an embodiment of a base editor is contemplated to contain a fusion protein, complexes comprising one or more domains of a base editor or fragments thereof are also contemplated.
[0081] "Cytosine" or "4-aminopyrimidin-2(1H)-one" has the molecular formula C 4 H 5 N 3 O, and the structure [ka] and means the purine nucleobase corresponding to CAS number 71-30-7.
[0082] "Cytidine" is linked to the ribose sugar via a glycosidic bond and has the structure [ka] It refers to the cytosine molecule having the formula C 9 H 13 N 3 O 5 It is.
[0083] By "cytidine base editor (CBE)" is meant a base editor that comprises a cytidine deaminase.
[0084] By "cytidine base editor (CBE) polynucleotide" is meant a polynucleotide that contains a CBE.
[0085] "Cytidine deaminase" or "cytosine deaminase" refers to a polypeptide or fragment thereof capable of deaminating cytidine or cytosine. In one embodiment, cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine. The terms "cytidine deaminase" and "cytosine deaminase" are used interchangeably throughout this application. Petromyzon marinus cytosine deaminase 1 (PmCDA1) (SEQ ID NOs: 12-13), activation-induced cytidine deaminase (AICDA) (SEQ ID NOs: 14-16 and 18-21), and APOBEC (SEQ ID NOs: 22-62) are exemplary cytidine deaminases. Further exemplary cytidine deaminase (CDA) sequences are set forth in the sequence listing as SEQ ID NOs: 63-67 and SEQ ID NOs: 68-190.
[0086] "Cytosine deaminase activity" means catalyzing the deamination of cytosine or cytidine. In one embodiment, a polypeptide having cytosine deaminase activity converts an amino group to a carbonyl group. In one embodiment, a cytosine deaminase converts cytosine to uracil (i.e., C to U) or 5-methylcytosine to thymine (i.e., 5mC to T). In some embodiments, the cytosine deaminase variants provided herein have high cytosine deaminase activity (e.g., at least 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, or more) relative to a reference cytosine deaminase. As used herein, the term "deaminase" or "deaminase domain" refers to a protein or fragment thereof that catalyzes a deamination reaction.
[0087] "Detection" refers to identifying the presence, absence, or amount of an analyte being detected. In one embodiment, a sequence alteration in a polynucleotide or polypeptide is detected. In another embodiment, the presence of an indel is detected.
[0088] "Detectable label" refers to a composition that, when linked to a molecule of interest, allows the molecule of interest to be detected through spectroscopic, photochemical, biochemical, immunochemical, or chemical means. For example, useful labels include radioisotopes, magnetic beads, metal beads, colloidal particles, fluorescent dyes, electron-dense reagents, enzymes (e.g., enzymes commonly used in enzyme-linked immunosorbent assays (ELISAs)), biotin, digoxigenin, or haptens.
[0089] By "disease" is meant any condition or disorder that damages or interferes with the normal function of a cell, tissue, or organ.
[0090] "Effective amount" refers to the amount of a drug or active compound, e.g., a base editor as described herein, required to ameliorate the symptoms of a disease in an untreated patient or an individual without a disease, i.e., a healthy individual, or is the amount of a drug or active compound that is sufficient to induce a desired biological response. The effective amount of the active compound(s) used to practice the present invention for the therapeutic treatment of a disease will vary depending on the mode of administration, the age, weight, and general health of the subject. Ultimately, the attending physician or veterinarian will determine the appropriate amount and dosing regimen. Such an amount is referred to as an "effective" amount. In one embodiment, an effective amount is an amount of a base editor of the present invention that is sufficient to introduce a modification in a gene of interest in a cell (e.g., a cell in vitro or in vivo). In one embodiment, an effective amount is the amount of a base editor required to achieve a therapeutic effect. Such a therapeutic effect need not be sufficient to modify pathogenic genes in all cells of a subject, tissue or organ, but need only be sufficient to modify pathogenic genes in about 1%, 5%, 10%, 25%, 50%, 75% or more of the cells present in the subject, tissue or organ. In one embodiment, an effective amount is sufficient to ameliorate one or more symptoms of a disease.
[0091] The term "exonuclease" refers to a protein or polypeptide capable of digesting nucleic acids (eg, RNA or DNA) from their free ends.
[0092] The term "endonuclease" refers to a protein or polypeptide that can catalyze (eg, cleave) an internal region in a nucleic acid (eg, DNA or RNA).
[0093] By "fragment" is meant a portion of a polypeptide or nucleic acid molecule. The portion contains at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the full length of the reference nucleic acid molecule or polypeptide. A fragment may contain 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids.
[0094] "Guide polynucleotide" refers to a polynucleotide or polynucleotide complex that is specific to a target sequence and can form a complex with a polynucleotide-programmable nucleotide-binding domain protein (e.g., Cas9 or Cpf1). In one embodiment, the guide polynucleotide is a guide RNA (gRNA). The gRNA can exist as a complex of two or more RNAs or as a single RNA molecule.
[0095] In some embodiments, the guide polynucleotide has a nucleotide sequence selected from the following, where a lower case "g" indicates a 5' mismatch to the target sequence:
[0096] TIFF2024521750000006.tif236165 TIFF2024521750000007.tif220165
[0097] "Heterologous" or "exogenous" refers to a polynucleotide or polypeptide that is 1) experimentally incorporated into a polynucleotide or polypeptide sequence not normally found in nature, or 2) experimentally placed into a cell that does not normally contain that polynucleotide or polypeptide. In some embodiments, "heterologous" means that the polynucleotide or polypeptide has been experimentally placed into a non-natural context. In some embodiments, the heterologous polynucleotide or polypeptide is derived from a first species or host organism and is incorporated into a polynucleotide or polypeptide derived from a second species or host organism. In some embodiments, the first species or host organism is different from the second species or host organism. In some embodiments, the heterologous polynucleotide is DNA. In some embodiments, the heterologous polynucleotide is RNA.
[0098] In some embodiments, the heterologous polynucleotide is a heterologous intron. In some embodiments, the heterologous intron is a synthetic intron. In some embodiments, the heterologous intron is derived from a mammalian gene (e.g., NF1, PAX2, EEF1A1, HBB, IGHG1, SLC50A1, ABCB11, BRSK2, PLXNB3, TMPRSS6, IL32, ANTXRL, PKHD1L1, PADI1, KRT6C, or HMCN2). In some embodiments, the heterologous intron is derived from a non-mammalian gene (e.g., HMCN2-Salmon, ENPEP-Gecko). In some embodiments, a polynucleotide encoding a base editor provided herein comprises a heterologous intron. In some embodiments, the base editor is an adenosine base editor (ABE). In some embodiments, the base editor is a cytidine base editor (CBE).
[0099] In some embodiments, the heterologous intron is incorporated into a polynucleotide encoding a polynucleotide programmable DNA binding protein or a fragment thereof. In some embodiments, the polynucleotide programmable DNA binding protein is Cas9, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ domain. In some embodiments, the polynucleotide programmable DNA binding domain is Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof.
[0100] In some embodiments, the heterologous intron is integrated into the polynucleotide that encodes the deaminase or a fragment thereof. In some embodiments, the heterologous intron is integrated into the polynucleotide that encodes the adenosine deaminase. In some embodiments, the adenosine deaminase is TadA. In some embodiments, the heterologous intron is integrated into the polynucleotide that encodes the cytidine deaminase.
[0101] "Hybridization" means hydrogen bonding, which may be Watson-Crick, Hoogsteen, or reversed Hoogsteen hydrogen bonding, between complementary nucleobases. For example, adenine and thymine are complementary nucleobases that pair through the formation of hydrogen bonds.
[0102] By "increase" is meant a positive change of at least 10%, 25%, 50%, 75%, or 100%.
[0103] The terms "inhibitor of base repair," "base repair inhibitor," "IBR," or grammatical equivalents thereof, refer to a protein capable of inhibiting the activity of a nucleic acid repair enzyme, e.g., a base excision repair enzyme.
[0104] An "intein" is a fragment of a protein that can excise itself and join the remaining fragment (the extein) with a peptide bond in a process known as protein splicing.
[0105] "Intron" refers to a non-coding nucleotide sequence that is removed by splicing prior to translation of a transcript. In some embodiments, the intron is removed by RNA splicing during the precursor messenger RNA stage of maturation of the mRNA. In some embodiments, the intron is derived from a gene of an organism. In some embodiments, the intron is synthetic. In some embodiments, the intron comprises a splice acceptor and a splice donor site. In some embodiments, the intron is about 10, 25, 50, 75, 100, 125, 150, 175, 200, 250, 300, 350, 400, 450, or 500 nucleotides in length. In some embodiments, the intron is about 50, 100, 125, 150, 175, or 200 nucleotides in length. In some embodiments, the intron is about 150 nucleotides in length.
[0106] In some embodiments, the intron is derived from a mammalian gene (e.g., NF1, PAX2, EEF1A1, HBB, IGHG1, SLC50A1, ABCB11, BRSK2, PLXNB3, TMPRSS6, IL32, ANTXRL, PKHD1L1, PADI1, KRT6C, or HMCN2). In some embodiments, the intron is derived from a non-mammalian gene (e.g., HMCN2-Salmon, ENPEP-Gecko). In some embodiments, the intron has a polynucleotide sequence selected from the following: TIFF2024521750000008.tif219165
[0107] In some embodiments, a polynucleotide encoding a base editor provided herein comprises a heterologous intron. In some embodiments, the base editor is an adenosine base editor (ABE). In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the intron is heterologously incorporated within a polynucleotide sequence. In some embodiments, the polynucleotide sequence is DNA. In some embodiments, the polynucleotide sequence is RNA. In some embodiments, the intron is heterologously incorporated within a polynucleotide encoding a polynucleotide programmable DNA binding protein. In some embodiments, the polynucleotide programmable DNA binding protein is a Cas9, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ domain. In some embodiments, the polynucleotide programmable DNA binding domain is Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof.
[0108] In some embodiments, the intron is heterologously integrated into the polynucleotide encoding the deaminase. In some embodiments, the intron is heterologously integrated into the polynucleotide encoding the adenosine deaminase. In some embodiments, the adenosine deaminase is TadA. In some embodiments, the intron is heterologously integrated into the polynucleotide encoding the cytidine deaminase. In some embodiments, the intron is heterologously integrated into the polynucleotide programmable DNA binding protein (e.g., Cas9). In some embodiments, the intron is heterologously integrated into the linker region.
[0109] The terms "isolated," "purified," or "biologically pure" refer to material that is free to different degrees from components that normally accompany it as found in the natural state. "Isolated" refers to a degree of separation from the original source or surroundings. "Purified" refers to a degree of separation that is greater than isolation. A "purified" or "biologically pure" protein is sufficiently free of other substances so that no impurities substantially affect the biological properties of the protein or cause other adverse events. That is, the nucleic acids or peptides of the invention are purified to be substantially free of cellular material, viral material, or culture medium if produced by recombinant DNA technology, or substantially free of chemical precursors or other chemicals if chemically synthesized. Purity and homogeneity are usually determined using analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high performance liquid chromatography. The term "purified" may indicate that the nucleic acid or protein gives rise to essentially one band in an electrophoretic gel. For example, in the case of proteins that can be subject to modifications such as phosphorylation or glycosylation, the different modifications can give rise to different isolated proteins that can be purified separately.
[0110] "Isolated polynucleotide" refers to a nucleic acid molecule that does not contain the genes adjacent to the gene in the natural genome of the organism from which the nucleic acid molecule of the present invention is derived. In an embodiment, the nucleic acid molecule contains DNA or is a DNA molecule. Thus, the term includes, for example, recombinant DNA that is incorporated into a vector, an autonomously replicating plasmid or virus, or an incorporated into the genomic DNA of a prokaryotic or eukaryotic organism, or exists as a separate molecule independent of other sequences (e.g., cDNA or genomic or cDNA fragments generated by PCR or restriction endonuclease digestion). In addition, the term includes RNA molecules that are transcribed from DNA molecules, as well as recombinant DNA that is part of a hybrid gene that codes for additional polypeptide sequences.
[0111] By "isolated polypeptide" is meant a polypeptide of the invention separated from components which naturally accompany it. Typically, a polypeptide is isolated when it is at least 60%, by weight, free from the proteins and naturally occurring organic molecules with which it is naturally associated. Preferably, a preparation is at least 75%, more preferably at least 90%, and most preferably at least 99%, by weight, the polypeptide of the invention. An isolated polypeptide of the invention can be obtained, for example, by extraction from a natural source, by expression of a recombinant nucleic acid encoding such a polypeptide, or by chemically synthesizing the protein. Purity can be measured by any appropriate method, for example, column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.
[0112] As used herein, the term "linker" refers to a molecule that connects two moieties. In one embodiment, the term "linker" refers to a covalent linker (e.g., a covalent bond) or a non-covalent linker.
[0113] As used herein, the term "mutation" refers to the substitution of a residue in a sequence, e.g., a nucleic acid or amino acid sequence, with another residue, or the deletion or insertion of one or more residues in a sequence. Mutations are usually described herein by identifying the original residue, followed by identifying the position of the residue in the sequence, and then identifying the newly substituted residue. Various methods for making amino acid substitutions (mutations) provided herein are well known in the art and are described, for example, in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4 th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012).
[0114] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to a compound that includes a nucleobase and an acidic moiety, e.g., a nucleoside, a nucleotide, or a polymer of nucleotides. Typically, polymeric nucleic acids, e.g., nucleic acid molecules that include three or more nucleotides, are linear molecules in which adjacent nucleotides are linked to each other via phosphodiester bonds. In some embodiments, "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides). In some embodiments, "nucleic acid" refers to an oligonucleotide chain that includes three or more individual nucleotide residues. As used herein, the terms "oligonucleotide" and "polynucleotide" may be used interchangeably to refer to a polymer of nucleotides (e.g., a string of at least three nucleotides). In some embodiments, "nucleic acid" encompasses RNA as well as single-stranded and / or double-stranded DNA. Nucleic acids may be naturally occurring, e.g., in the context of a genome, a transcript, an mRNA, a tRNA, an rRNA, an siRNA, an snRNA, a plasmid, a cosmid, a chromosome, a chromatid, or other naturally occurring nucleic acid molecule. On the other hand, the nucleic acid molecule may be a non-naturally occurring molecule, such as a recombinant DNA or RNA, an artificial chromosome, an engineered genome or a fragment thereof, or a synthetic DNA, RNA, DNA / RNA hybrid, or may contain non-naturally occurring nucleotides or nucleosides.
[0115] Additionally, "nucleic acid," "DNA," "RNA," and / or similar terms include nucleic acid analogs, e.g., analogs having other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems, and optionally purified or chemically synthesized. Optionally, for example, in the case of chemically synthesized molecules, nucleic acids can include nucleoside analogs, such as analogs having chemically modified bases or sugars, and backbone modifications. Nucleic acid sequences are presented in the 5' to 3' direction unless otherwise specified. In some embodiments, the nucleic acid may be any combination of naturally occurring nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), nucleotide analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C ...bromouridine, C5-bromouridine, C5-bromouridine, C5-bromouridine, C5-bromouridine, C5-bromouridine, C5-bromouridine, C5-bromouridine, C5-bromouridine, C5-bromouridine, C5-bromouridine, C5-bromouridine, C5-bromour -aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalated bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose), and / or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages).
[0116] The term "nuclear localization sequence", "nuclear localization signal", or "NLS" refers to an amino acid sequence that facilitates the import of a protein into a cell nucleus. Nuclear localization sequences are known in the art and are described, for example, in International PCT Application PCT / EP2000 / 011690, filed November 23, 2000 by Plank et al. (published May 31, 2001 as WO / 2001 / 038547), the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, for example, as described in Koblan et al., Nature Biotech.2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises the amino acid sequence KRTADGSEFESPKKKRKV (SEQ ID NO: 243), KRPAATKKAGQAKKKK (SEQ ID NO: 244), KKTELQTTNAENKTKKL (SEQ ID NO: 245), KRGINDRNFWRGENGRKTR (SEQ ID NO: 246), RKSGKIAAIVVKRPRK (SEQ ID NO: 247), PKKKRKV (SEQ ID NO: 248), or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 249).
[0117] The terms "nucleobase", "nitrogenous base", or "base" are used interchangeably herein to refer to the nitrogen-containing biological compounds that form nucleosides, which are components of nucleotides. The ability of nucleobases to base pair and stack with one another leads directly to long-chain helical structures such as ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). The five nucleobases, adenine (A), cytosine (C), guanine (G), thymine (T) and uracil (U), are referred to as primary or canonical. Adenine and guanine are derived from purines, while cytosine, uracil, and thymine are derived from pyrimidines. DNA and RNA can also contain other (non-primary) bases that are modified. Non-limiting exemplary modified nucleobases include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydromethylcytosine. Both hypoxanthine and xanthine can be generated through deamination (replacement of an amine group with a carbonyl group) in the presence of mutagens. Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can result from the deamination of cytosine. A "nucleoside" consists of a nucleobase and a five-carbon sugar (either ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine. Examples of nucleosides having modified nucleobases include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), and pseudouridine (Ψ). A "nucleotide" consists of a nucleobase, a five-carbon sugar (either ribose or deoxyribose), and at least one phosphate group.Non-limiting examples of modified nucleobases and / or chemical modifications that a modified nucleobase may include are: pseudouridine, 5-methyl-cytosine, 2'-O-methyl-3'-phosphonoacetate, 2'-O-methylthio PACE (MSP), 2'-O-methyl-PACE (MP), 2'-fluoro RNA (2'-F-RNA), constrained ethyl (S-cEt), 2'-O-methyl ("M"), 2'-O-methyl-3'-phosphorothioate ("MS"), 2'-O-methyl-3'-thiophosphonoacetate ("MSP"), 5-methoxyuridine, phosphorothioate, and N1-methylpseudouridine.
[0118] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" may be used interchangeably with "polynucleotide programmable nucleotide binding domain" and may refer to a protein that associates with a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid or guide polynucleotide (e.g., gRNA) that guides the napDNAbp to a specific nucleic acid sequence. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable RNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a Cas9 protein. The Cas9 protein can associate with a guide RNA that guides the Cas9 protein to a specific DNA sequence that is complementary to the guide RNA. In some embodiments, the napDNAbp is a Cas9 domain, such as a nuclease-active Cas9, a Cas9 nickase (nCas9), or a nuclease-inactive Cas9 (dCas9). Non-limiting examples of nucleic acid programmable DNA binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ (Cas12j / Casphi).Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Cas12j / CasΦ, Cpf1, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Css1, Css2, Css1 ... sn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins, CARF, DinG, homologs thereof, or modified or engineered versions thereof. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, but may not be specifically listed in the present disclosure. See, for example, Makarova et al., "Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?" CRISPR J. 2018 Oct; 1: 325-336. doi: 10.1089 / crispr.2018.0033, Yan et al., "Functionally diverse type V CRISPR-Cas systems" Science. 2019 Jan 4; 363(6422): 88-91. doi: 10.1126 / science.aav7271 (the entire contents of each are incorporated herein by reference).Exemplary nucleic acid programmable DNA binding proteins and nucleic acid sequences encoding nucleic acid programmable DNA binding proteins are provided in the Sequence Listing as SEQ ID NOs: 250-283 and 490.
[0119] As used herein, the term "nucleobase editing domain" or "nucleobase editing protein" refers to a protein or enzyme that can catalyze nucleobase modifications in RNA or DNA, such as the deamination of cytosine (or cytidine) to uracil (or uridine) or thymine (or thymidine), and the deamination of adenine (or adenosine) to hypoxanthine (or inosine), as well as non-templated nucleotide addition and insertion. In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., an adenine deaminase or adenosine deaminase, or a cytidine deaminase or cytosine deaminase).
[0120] As used herein, "obtaining" in "obtaining a drug" includes synthesizing, purchasing, or otherwise obtaining a drug.
[0121] As used herein, "patient" or "subject" refers to a mammalian subject or individual who has been diagnosed with, is at risk of having or developing, or is suspected of having or developing a disease or disorder. In some embodiments, the term "patient" refers to a mammalian subject who has a higher than average likelihood of developing a disease or disorder. Exemplary patients can be humans, non-human primates, cats, dogs, pigs, cows, cats, horses, camels, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, or guinea pigs), and other mammals that can benefit from the therapies disclosed herein. Exemplary human patients can be male and / or female.
[0122] A "patient in need thereof" or a "subject in need thereof" as used herein refers to a patient who has been diagnosed with, is at risk of or has, has been predetermined to have, or is suspected of having a disease or disorder.
[0123] The term "pathogenic mutation", "pathogenic variant", "causative mutation", "pathogenic variant", "harmful mutation" or "predisposing mutation" refers to a genetic modification or mutation associated with a disease or disorder, or that increases an individual's susceptibility or predisposition to a particular disease or disorder. In some embodiments, a pathogenic mutation comprises at least one wild-type amino acid replaced by at least one pathogenic amino acid in a protein encoded by a gene. In some embodiments, a pathogenic mutation is in a termination region (e.g., a stop codon). In some embodiments, a pathogenic mutation is in a non-coding region (e.g., an intron, a promoter, etc.).
[0124] The terms "protein," "peptide," "polypeptide," and their grammatical equivalents are used interchangeably herein to refer to a polymer of amino acid residues linked together by peptide (amide) bonds. A protein, peptide, or polypeptide may be natural, recombinant, or synthetic, or any combination thereof.
[0125] As used herein, the term "fusion protein" refers to a hybrid polypeptide that contains protein domains derived from at least two different proteins.
[0126] The term "recombinant" as used herein in the context of a protein or nucleic acid refers to a protein or nucleic acid that does not occur in nature but is the product of human engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that contains at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.
[0127] By "reduction" is meant a negative change of at least 10%, 25%, 50%, 75%, or 100%.
[0128] "Reference" refers to a standard or control condition. In one embodiment, the reference is the level of editing provided by a base editor encoded by an intron-free polynucleotide. In another embodiment, the reference is the level of editing provided by a base editor encoded by an intron-containing polynucleotide that does not include an alteration in the splice acceptor or splice donor site. In one embodiment, the reference is the level, structure, or activity of an analyte present in a wild-type or healthy cell. In other embodiments, without limitation, the reference is the level, structure, or activity of an analyte present in an untreated cell that is not subjected to the test condition, or in an untreated cell that is subjected to a placebo or normal saline, medium, buffer, and / or a control vector that does not carry the polynucleotide of interest.
[0129] A "reference sequence" is a defined sequence used as a basis for sequence comparison. A reference sequence may be a subset of a particular sequence or the entirety thereof, for example, a segment of a full-length cDNA or gene sequence, or a complete cDNA or gene sequence. For polypeptides, the length of a reference polypeptide sequence is generally at least about 16 amino acids, at least about 20 amino acids, at least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids. For nucleic acids, the length of a reference nucleic acid sequence is generally at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides, or about 300 nucleotides, or any integer thereabout or therebetween. In some embodiments, the reference sequence is a wild-type sequence of a protein of interest. In other embodiments, the reference sequence is a polynucleotide sequence that encodes a wild-type protein.
[0130] The terms "RNA programmable nuclease" and "RNA-guided nuclease" are used with (e.g., bound or associated with) one or more RNA(s) that are not targets for cleavage. In some embodiments, the RNA programmable nuclease may be referred to as a nuclease:RNA complex when in a complex with an RNA. Typically, the bound RNA(s) is referred to as a guide RNA (gRNA). In some embodiments, the RNA programmable nuclease is a (CRISPR-associated system) Cas9 endonuclease, such as Cas9 from Streptococcus pyogenes (Csnl) (e.g., SEQ ID NO: 250), Cas9 from Neisseria meningitidis (NmeCas9, SEQ ID NO: 261), Nme2Cas9 (SEQ ID NO: 262) or a derivative thereof (e.g., a sequence having at least about 85% sequence identity to Cas9, such as Nme2Cas9 or spCas9).
[0131] The term "single nucleotide polymorphism (SNP)" refers to a single nucleotide variation that occurs at a specific location in the genome, with each variation occurring to some degree (eg, >1%) in the population.
[0132] By "specifically binds" is meant that a nucleic acid molecule, polypeptide, polypeptide / polynucleotide complex, compound or molecule recognizes and binds to a polypeptide and / or nucleic acid molecule of the invention, but does not substantially recognize or bind to other molecules in a sample, e.g., a biological sample.
[0133] "Substantially identical" refers to a polypeptide or nucleic acid molecule that exhibits at least 50% identity to a reference amino acid sequence. In one embodiment, the reference sequence is a wild-type amino acid or nucleic acid sequence. In another embodiment, the reference sequence is any one of the amino acid or nucleic acid sequences described herein. In one embodiment, such a sequence is at least 60%, 80%, 85%, 90%, 95%, or even 99% identical at the amino acid or nucleic acid level to the sequence used for comparison.
[0134] Sequence identity is typically measured using sequence analysis software (e.g., the sequence analysis software packages of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, GAP or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. In an exemplary approach for determining the degree of identity, e, which indicates closely related sequences, is used. -3 ~e -100 The BLAST program may be used, using a probability score of
[0135] Use COBALT, for example, with the following parameters: a) alignment parameters: gap penalty -11, -1 and end gap penalty -5, -1; b) CDD parameters: Use RPS BLAST, Blast E value 0.003, find and recalculate conserved columns, and c) Query clustering parameters: Use query cluster, word size 4, max cluster distance 0.8, alphabet normal. EMBOSS Needle is used, for example, with the following parameters: a) Matrix: BLOSUM62; b) GAP OPEN: 10, c) GAP EXTEND: 0.5; d) OUTPUT FORMAT: vs. e) END GAP PENALTY: False; f) END GAP OPEN: 10, and g) END GAP EXTEND: 0.5.
[0136] Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule that encodes a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to an endogenous nucleic acid sequence, but will typically exhibit substantial identity. A polynucleotide with "substantial identity" to an endogenous sequence can typically hybridize with at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule that encodes a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to an endogenous nucleic acid sequence, but will typically exhibit substantial identity. A polynucleotide with "substantial identity" to an endogenous sequence can typically hybridize with at least one strand of a double-stranded nucleic acid molecule. By "hybridize" it is meant that the pair forms a double-stranded molecule between complementary polynucleotide sequences (e.g., genes described herein) or portions thereof under various stringency conditions. (See, e.g., Wahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, AR (1987) Methods Enzymol. 152:507).
[0137] For example, stringent salt concentrations are usually less than about 750 mM NaCl and less than 75 mM trisodium citrate, preferably less than about 500 mM NaCl and less than 50 mM trisodium citrate, more preferably less than about 250 mM NaCl and less than 25 mM trisodium citrate. Low stringency hybridization can be obtained in the absence of organic solvents, such as formamide, while high stringency hybridization can be obtained in the presence of at least about 35% formamide, more preferably at least about 50% formamide. Stringent temperature conditions usually include a temperature of at least about 30°C, more preferably at least about 37°C, and most preferably at least about 42°C. Various additional parameters, such as hybridization time, concentration of detergent, such as sodium dodecyl sulfate (SDS), and inclusion or exclusion of carrier DNA, are well known to those skilled in the art. Various levels of stringency can be achieved by combining these various conditions as needed. In a preferred embodiment, hybridization is carried out in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS at 30° C. In a more preferred embodiment, hybridization is carried out in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA) at 37° C. In a most preferred embodiment, hybridization is carried out in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA at 42° C. Useful variations of these conditions will be readily apparent to those of skill in the art.
[0138] In most applications, the washing steps following hybridization will also vary in stringency. Stringency conditions for washing can be defined by salt concentration and temperature. As mentioned above, washing stringency can be increased by lowering the salt concentration or increasing the temperature. For example, stringent salt concentrations for washing steps are preferably less than about 30 mM NaCl and less than 3 mM trisodium citrate, and most preferably less than about 15 mM NaCl and less than 1.5 mM trisodium citrate. Stringent temperature conditions for washing steps usually include a temperature of at least about 25°C, more preferably at least about 42°C, and even more preferably at least about 68°C. In one embodiment, the washing step is carried out at 25°C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In another embodiment, the washing step is carried out at 42°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the washing step is carried out in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS at 68°C. Additional variations in these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196:180, 1977), Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975), Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001), Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York), and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.
[0139] "Split" means divided into two or more pieces.
[0140] "Split Cas9 protein" or "split Cas9" refers to a Cas9 protein that is provided as an N-terminal fragment and a C-terminal fragment encoded by two separate nucleotide sequences. Polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein can be spliced to form a "reconstituted" Cas9 protein.
[0141] The term "target site" refers to a sequence within a nucleic acid molecule that is modified. In embodiments, the nucleic acid molecule is deaminated by a deaminase, a fusion protein or complex comprising a deaminase, or a base editor disclosed herein. In embodiments, the deaminase is a cytidine or adenine deaminase. In some cases, the deaminase is a dCas9-adenosine deaminase fusion protein. In some cases, the base editor is an adenine or adenosine base editor (ABE), or a cytidine or cytosine base editor (CBE).
[0142] As used herein, terms such as "treat", "treating" and "treatment" refer to reducing or ameliorating a disorder and / or symptoms associated therewith, or obtaining a desired pharmacological and / or physiological effect. It will be understood that treating a disorder or condition does not necessarily require, but does not preclude, that the disorder, condition, or symptoms associated therewith be completely eliminated. In some embodiments, the effect is therapeutic, i.e., without limitation, the effect reduces, decreases, inhibits, alleviates, alleviates, reduces the intensity, or cures, partially or completely, the disease and / or adverse symptoms resulting from the disease. In some embodiments, the effect is prophylactic, i.e., the effect prevents or prevents the occurrence or recurrence of the disease or condition. To this end, the methods disclosed herein include administering a therapeutically effective amount of a composition as described herein.
[0143] "Uracil glycosylase inhibitor" or "UGI" refers to an agent that inhibits the uracil excision repair system. Base editors that contain cytidine deaminase convert cytosine to uracil, which is then converted to thymine during DNA replication or repair. Inclusion of an inhibitor of uracil DNA glycosylase (UGI) in a base editor prevents base excision repair that changes U back to C. An exemplary UGI includes the following amino acid sequence: >splP14739IUNGI_BPPB2 Uracil DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (SEQ ID NO: 284).
[0144] It is understood that ranges provided herein are shorthand for all values within the range. For example, the range 1 to 50 is understood to include any number, combination of numbers, or subrange from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.
[0145] The recitation of a listing of chemical groups within any definition of a variable herein includes the definition of that variable as any single group or combination of listed groups. The recitation of an embodiment for a variable or aspect herein includes that embodiment as any single embodiment or in combination with any other embodiment or portion thereof.
[0146] All terms are intended to be understood as understood by one of ordinary skill in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. In this application, the use of the singular includes the plural unless specifically stated otherwise. It should be noted that as used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. In this application, the use of "or" means "and / or" unless otherwise specified. Furthermore, the use of the term "including," as well as other forms such as "include," "includes," and "included," is not limiting.
[0147] As used in the specification and claim(s), the words "comprising" (and any of its forms, e.g., "comprise" and "comprises"), "having" (and any of its forms, e.g., "have" and "has"), "including" (and any of its forms, e.g., "includes" and "include"), or "containing" (and any of its forms, e.g., "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. Any embodiment specified as "comprising" a particular component(s) or element(s) is also contemplated in some embodiments to "consist of" or "consist essentially of" the particular component(s) or element(s). It is contemplated that any embodiment discussed herein may be implemented with respect to any method or composition of the disclosure, and vice versa. Additionally, the compositions of the present disclosure can be used to achieve the methods of the present disclosure.
[0148] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which will depend, in part, on how the value is measured or determined (i.e., the limitations of the measurement system).
[0149] References herein to "some embodiments," "embodiments," "one embodiment," or "other embodiments" mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least some embodiments of the present disclosure, but not necessarily in all embodiments. [Brief description of the drawings]
[0150] [Figure 1A]A schematic diagram showing the mechanism of self-inactivation of base editors is provided. Two gRNAs direct base editing to occur simultaneously at the target site in the host genome and within the coding region of the base editor. If the base editor used is an adenine base editor (ABE), the catalytic residues of the deaminase domain (His57 (H57), Glu59 (E59), Cys87 (C87) or Cys90 (C90)) can be inactivated via a single A-to-G editor to install Arg, Gly, Arg or Arg at each site, respectively. If the base editor used is a cytosine base editor (CBE), a premature stop codon can be installed at any Arg, Gln or Trp residue within the editor via a single C-to-T edit. [Figure 1B] FIG. 13 provides a bar graph showing base editing activity in HEK293T cells after lipofection of ABE7.10-m and ABE7.10-m variants containing preinstalled TadA mutations (His57Arg, Glu59Gly, Cys87Arg or Cys90Arg) at a genomic site (ABCA4 c.5882G>A) and a self-inactivating site in TadA (His57, Glu59, Cys87 or Cys90) using two gRNAs. [Figure 1C] FIG. 1C provides a schematic showing the DNA sequence of the self-inactivation target sites His57 and Glu59 within the TadA coding region. The 3'PAM sequence is highlighted in grey, and the target nucleotide and its position within the protospacer in each sequence is in bold. The nucleotide sequences provided in FIG. 1C correspond, in order of appearance from top to bottom, to SEQ ID NOs: 458-459. The amino acid sequences provided in FIG. 1C correspond, in order of appearance from top to bottom, to SEQ ID NOs: 460-461. [Figure 1D]Graph showing base editing activity in HEK293T cells after lipofection of ABE8.5-m codon variants and two gRNAs targeting a genomic site (ABCA4 c.5882G>A) and the self-inactivating site Glu59 of TadA. The activity of the variants was compared to the activity of ABE8.5-m, which was not provided with a self-inactivating gRNA. [Figure 1E] 1 provides a bar graph showing base editing kinetics of AAV2-delivered ABE8.5-m codon variant and two gRNAs at the genomic site and the TadA catalytic residue of ABE in ARPE-19 cells. 2 provides a bar graph showing a 5-week time course of base editing at the genomic site (ABCA4 c.5882G>A) after AAV2 delivery of ABE8.5-m codon variant and two gRNAs. [Figure 1F] 1 provides bar graphs showing base editing kinetics of AAV2-delivered ABE8.5-m codon variants and two gRNAs at genomic sites and ABE TadA catalytic residues in ARPE-19 cells. 2 provides bar graphs showing editing at amino acid residue His57 or residue Glu59, the self-inactivating site of TadA, in the same samples from a 5-week time course. [Figure 1G] 1 provides a bar graph showing base editing kinetics of AAV2-delivered ABE8.5-m codon variant and two gRNAs at a genomic site and the TadA catalytic residue of ABE in ARPE-19 cells, where self-inactivating editing is assessed by two different methods. 2 weeks after AAV2 delivery of ABE8.5-m codon variant and two gRNAs, a bar graph showing base editing at a genomic site (ABCA4 c.5882G>A) is provided. [Figure 1H]1 provides bar graphs showing base editing kinetics of AAV2-delivered ABE8.5-m codon variants and two gRNAs at genomic sites and the TadA catalytic residue of ABE in ARPE-19 cells, where self-inactivating editing is assessed by two different methods. 2 provides bar graphs showing self-inactivation rates assessed by targeted sequencing of either DNA from cell lysates or cDNA generated from mRNA of technical replicate samples in the same experiment. [Figure 2A] FIG. 2A provides a diagram of mutations made to TadA to inactivate the editor via modification of the ABE start codon. Mutations in the DNA and protein sequences are highlighted in black. Alternative out-of-frame start codons are identified by grey boxes. The nucleotide sequences provided in FIG. 2A correspond, in order of appearance from top to bottom, to SEQ ID NOs: 462-466. The amino acid sequences provided in FIG. 2A correspond, in order of appearance from top to bottom, to SEQ ID NOs: 467-469. [Figure 2B] 1 provides a bar graph showing base editing activity at genomic site ABCA4 c.5882G>A in HEK293T cells following lipofection of ABE8.5-m variants containing pre-installed start codon mutations. No self-inactivating gRNA was provided in this experiment. [Figure 2C] FIG. 2C provides a diagram showing mutations made to ABE8.5-m to incorporate a PAM sequence (NGG) that allows base editing to occur at Met1 of TadA. The nucleotide sequences provided in FIG. 2C correspond, in order of appearance from top to bottom, to SEQ ID NOs: 470-476. The amino acid sequences provided in FIG. 2C correspond, in order of appearance from top to bottom, to SEQ ID NOs: 477-480. [Figure 2D] 1 provides a bar graph showing base editing activity at genomic site ABCA4 c.5882G>A in HEK293T cells after lipofection of ABE8.5-m variants containing an installed PAM sequence in TadA compared to unmutated controls. No self-inactivating gRNA was provided in the experiment. [Figure 2E]1 provides a bar graph showing base editing activity in HEK293T cells after lipofection of ABE8.5-m and ABE8.5-m variants at a genomic site (ABCA4 c.5882G>A) and a self-inactivating site Met1 in TadA using two gRNAs. [Figure 3A]
[0023] Figure 1 provides a schematic showing the mechanism of self-inactivation of adenine base editors (ABEs) through intron incorporation into DNA. [Figure 3B] 10 provides bar graphs showing base editing activity in HEK293T cells following lipofection of ABE variants containing introns in the coding sequence after or within specific codons (residues) of TadA. 11 provides bar graphs showing base editing activity following intron incorporation after residue 87 of TadA (NF1, PAX2, EEF1A1, Chimera, SLC50A1, ABCB11, BRSK2, PLXNB3, TMPRSS6, IL32), incorporation after residue 62 (Chimera, ABCB11, PLXNB3, IL32), or incorporation within residue 23 (Chimera, ABCB11, PLXNB3, IL32). [Figure 3C] A bar graph is provided showing base editing activity in HEK293T cells after lipofection of ABE variants containing introns in the coding sequence after or within specific codons (residues) of TadA. A bar graph is provided showing base editing activity after incorporation of several additional introns (ANTXRL, PKHD1L1, PADI1, KRT6C, HMCN2, HMCN2-Salmon, or ENPEP-Gecko) after residue 87 in addition to NF1, PAX2, and EEF1A1. No self-inactivating gRNA was provided in this experiment. [Figure 3D]1 provides a bar graph showing base editing activity in HEK293T cells after lipofection of ABE variants containing introns with pre-installed edits in either the splice acceptor or splice donor sites. The introns were located after TadA residue 87 (NF1 acceptor, PAX2 acceptor, EEF1A1 acceptor, Chimera acceptor, ANTXRL acceptor, PKHK1L1 acceptor, PADI1 acceptor, KRT6C acceptor, HMCN2 acceptor, ENPEP-Gecko acceptor, HMCN2-Salmon acceptor, NF1 donor, PAX2 donor, EEF1A1 donor, or Chimera donor). No self-inactivating gRNA was provided in this experiment. [Figure 3E] 1 provides a bar graph showing base editing activity in HEK293T cells after lipofection of ABE variants containing introns with preinstalled edits in either the splice acceptor or splice donor sites. The introns were located after TadA residue 129 (NF1 acceptor, PAX2 acceptor, EEF1A1 acceptor), after TadA residue 59 (NF1 acceptor, PAX2 acceptor, EEF1A1 acceptor), after TadA residue 18 (NF1 acceptor, PAX2 acceptor, EEF1A1 acceptor), after TadA residue 62 (ABCB11 acceptor) or within residue 23 (ABCB11 donor). No self-inactivating gRNA was provided in this experiment. [Figure 3F] 1 provides a bar graph showing base editing activity in lipofected HEK293T cells at the genomic site (ABCA4 c.5882G>A) and the intronic (NF1 or PAX2) acceptor site located after residue 87 within TadA. [Figure 3G]FIG. 1 provides bar graphs showing base editing activity in lipofected HEK293T cells at the genomic site (ABCA4 c.5882G>A) and acceptor sites in introns (NF1, PAX2, and EEF1A1) located after residue 87 and introns (ABCB11) located after residue 62 within TadA. [Figure 3H] A bar graph showing the base editing activity of ABE8.5-m variants containing introns (NF1, PAX2 or EEF1A1) at various positions within TadA (after residues 87, 129, 59 or 18) with or without pre-installed mutations at the splice acceptor site. No self-inactivating gRNA was provided in this experiment. [Figure 3I] FIG. 1 provides a bar graph showing base editing activity in lipofected HEK293T cells at the genomic site (ABCA4 c.5882G>A) and acceptor sites in introns (NF1, PAX2, and EEF1A1) located after residues 87, 129, 59, and 18 within TadA. [Figure 3J] 1 provides a bar graph showing base editing activity in lipofected HEK293T cells at the genomic site (ABCA4 c.5882G>A) and the acceptor sites of introns NF1, PAX2, EEF1A1, ANTXRL, PKHD1L1, PADI1, and ENPEP-Gecko located after residue 87 within TadA. [Figure 3K]Figure 3K provides bar graphs showing base editing activity in HEK293T cells after plasmid lipofection of plasmid DNA encoding a self-inactivating gRNA, a gRNA targeting a genomic site, and an ABE variant containing an intron in the coding sequence of TadA. Figure 3K provides bar graphs showing base editing activity at a genomic site (ABCA4 c.5882G>A) and an intronic NF1 or PAX2 acceptor site located after residue 87 in TadA, where editing was assessed by targeted sequencing of DNA from cell lysates. Figures 3L and 3M provide stacked bar graphs showing the proportion of splice variants in ABE8.5-m mRNA, assessed by RNA-seq of total mRNA. All analyses in Figures 3K, 3L, and 3M were performed in technical replicates in the same experiment. [Figure 3L] Figure 3K provides bar graphs showing base editing activity in HEK293T cells after plasmid lipofection of plasmid DNA encoding a self-inactivating gRNA, a gRNA targeting a genomic site, and an ABE variant containing an intron in the coding sequence of TadA. Figure 3K provides bar graphs showing base editing activity at a genomic site (ABCA4 c.5882G>A) and an intronic NF1 or PAX2 acceptor site located after residue 87 in TadA, where editing was assessed by targeted sequencing of DNA from cell lysates. Figures 3L and 3M provide stacked bar graphs showing the proportion of splice variants in ABE8.5-m mRNA, assessed by RNA-seq of total mRNA. All analyses in Figures 3K, 3L, and 3M were performed in technical replicates in the same experiment. [Figure 3M]Figure 3K provides bar graphs showing base editing activity in HEK293T cells after plasmid lipofection of plasmid DNA encoding a self-inactivating gRNA, a gRNA targeting a genomic site, and an ABE variant containing an intron in the coding sequence of TadA. Figure 3K provides bar graphs showing base editing activity at a genomic site (ABCA4 c.5882G>A) and an intronic NF1 or PAX2 acceptor site located after residue 87 in TadA, where editing was assessed by targeted sequencing of DNA from cell lysates. Figures 3L and 3M provide stacked bar graphs showing the proportion of splice variants in ABE8.5-m mRNA, assessed by RNA-seq of total mRNA. All analyses in Figures 3K, 3L, and 3M were performed in technical replicates in the same experiment. [Figure 3N] A bar graph is provided showing base editing activity in ARPE-19 cells 2 weeks after AAV2 delivery of a self-inactivating gRNA targeting a splice acceptor site, a gRNA targeting a genomic site, and an ABE variant containing the NF1 intron at residue 87 in the coding sequence of TadA. Editing was measured at the genomic site by targeted sequencing of genomic DNA. Editing at the self-inactivating site is measured by both targeted sequencing of the recovered AAV genome and RNAseq of total mRNA from the cells. All measurements were performed in technical replicates in the same experiment. [Figure 4A] A bar graph is provided showing a 5-week AAV2 transduction experiment in which A>G base conversion was measured at weeks 1, 3, and 5 (x-axis) in ARPE-19 cells, a cell line derived from retinal pigment epithelium. A bar graph is provided showing editing at genomic site (ABCA4 c.5882G>A). The term "_scrmbl" indicates that the self-inactivating guide sequence is scrambled. The NF1 and PAX2 splice acceptor sites were edited using guides g235 and g239, respectively (see Table 1C). [Figure 4B]A bar graph is provided showing a 5-week AAV2 transduction experiment in which A>G base conversion was measured at weeks 1, 3, and 5 (x-axis) in ARPE-19 cells, a cell line derived from retinal pigment epithelium. A bar graph is provided showing editing at TadA catalytic residues or intron splice acceptor sites as measured by DNA sequencing. The term "_scrmbl" indicates that the self-inactivating guide sequence is scrambled. NF1 and PAX2 splice acceptor sites were edited using guides g235 and g239, respectively (see Table 1C). [Figure 4C] A bar graph is provided showing a 5-week AAV2 transduction experiment in which A>G base conversion was measured at weeks 1, 3, and 5 (x-axis) in ARPE-19 cells, a cell line derived from retinal pigment epithelium. A bar graph is provided showing measurement of editing of the same locus via RNA amplicon sequencing. The term "_scrmbl" indicates that the self-inactivating guide sequence is scrambled. The NF1 and PAX2 splice acceptor sites were edited using guides g235 and g239, respectively (see Table 1C). [Figure 5A] A bar graph is provided showing a two-week AAV2 transduction experiment in ARPE-19 cells after the indicated days of transduction (x-axis). Each bar graph shows the number of viral genomes (high, medium or low) added to transduce the cells. The number of viral genomes added to transduce the cells was either high (89 kvg / cell), medium (17 kvg / cell) or low (9 kvg / cell). A bar graph is provided showing the editing rate of a genomic site (ABCA4 c.5882G>A) at days 3, 7 and 14 post-transduction for the amount of virus added. [Figure 5B]Provided are bar graphs showing a two-week AAV2 transduction experiment in ARPE-19 cells after the indicated days of transduction (x-axis). Each bar graph shows the number of viral genomes added to transduce the cells (high, medium or low). The number of viral genomes added to transduce the cells was either high (89 kvg / cell), medium (17 kvg / cell) or low (9 kvg / cell). Provided are bar graphs showing editing at TadA catalytic residues or intron splice acceptor sites as measured by DNA sequencing at the indicated time points. [Figure 6A] A bar graph is provided showing a two week AAV2 time course transduction experiment in ARPE-19 cells in which editing was measured on days 4, 7 and 14. A bar graph is provided showing the editing rate of a genomic site (ABCA4 c.5882G>A) as measured via next generation sequencing. [Figure 6B] 1 provides a bar graph showing a 2 week AAV2 time course transduction experiment in ARPE-19 cells in which editing was measured on days 4, 7 and 14. 2 provides a bar graph showing editing of TadA catalytic residues or intron splice acceptors as measured via RNA amplicon sequencing. [Figure 7A] Provided are bar graphs showing the results of plasmid lipofection in HEK293T cells and the editing rates measured at 2 and 7 days post-lipofection. Provided are bar graphs showing editing of a genomic site (ABCA4 c.5882G>A) as measured via next generation sequencing. The term "_scrmbl" indicates that the self-inactivating guide sequence is scrambled. [Figure 7B] 1 provides bar graphs showing the results of plasmid lipofection in HEK293T cells and the editing rates measured at 2 and 7 days post-lipofection. 2 provides bar graphs showing the editing of TadA catalytic residues or intron splice acceptor sites measured via RNA amplicon sequencing. 3 The term "_scrmbl" indicates that the self-inactivating guide sequence is scrambled. [Figure 8A]1 provides a bar graph showing editing data collected after IV tail vein injection of AAV8 in BALB / c mice. 2 provides graphs showing editing at a genomic site (ABCA4 c.5882G>A) and editing of TadA catalytic residues or intron splice acceptor sites measured one week after transduction via both DNA and RNA amplicon sequencing. Editing of genomic sites is shown on the left y-axis and editing of TadA catalytic residues or intron splice acceptor is shown on the right y-axis. The term "_scrmbl" indicates that the self-inactivating guide sequence is scrambled. [Figure 8B] A bar graph is provided showing editing data collected after IV tail vein injection of AAV8 in BALB / c mice. A graph is provided showing the same results 4 weeks later. Editing of the genomic site is shown on the left y-axis, and editing of the TadA catalytic residue or intron splice acceptor is shown on the right y-axis. The term "_scrmbl" indicates that the self-inactivating guide sequence is scrambled. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0151] The invention features compositions that include self-inactivating base editors and methods of using such editors. The invention also features polynucleotides that encode base editors with heterologous introns for self-inactivation, compositions that include such polynucleotides, and methods of inactivating base editors encoded by such polynucleotides.
[0152] DNA base editing techniques generally use engineered DNA binding domains, such as RNA-guided Cas9 nickase (nCas9), in protein fusions with either cytosine or adenine deaminase. Cytosine base editors (CBEs) catalyze the transversion of cytosine to thymine (C>T) via a uracil intermediate, and adenine base editors (ABEs) catalyze the transversion of adenine to guanine (A>G) via a hypoxanthine intermediate (Rees, HA, & Liu, DR (2018). Base editing: precision chemistry on the genome and transcriptome of living cells. Nat Rev Genet, 19 (12), 770-788). DNA base editing relies on the RNA-guided nCas9 domain binding at a region of interest in the genome, which displaces the non-targeted strand of genomic DNA, which is extruded from nCas9 as an R-loop, thus exposing these unpaired bases for deamination. The targeted strand of DNA bound to the gRNA is also nicked by nCas9, which biases cellular DNA mismatch repair toward incorporation of the installed mutation in the R-loop, rather than degrading it to the wild-type base pairs of the unedited targeted strand.
[0153] As with all genome modification tools, care should be taken to protect against unwanted off-target edits in DNA that are persistent and potentially harmful (Kim, D., et al. (2017). Genome-wide target specificities of CRISPR RNA-guided programmable deaminases. Nat Biotechnol, 35(5), 475-480; Liang, P., et al. (2019). Genome-wide profiling of adenine base editor specificity by EndoV-seq. Nature Communications, 10(1), 67; Zuo, E., et al. (2019). Cytosine base editor generates substantial off-target single-nucleotide variants in mouse embryos. Science, 364(6437), 289-292).In situations where the DNA editor is expressed indefinitely, for example when delivered by AAV (Colella, P., et al. (2018). Emerging Issues in AAV-Mediated In Vivo Gene Therapy. Molecular Therapy-Methods & Clinical Development, 8, 87-104, Nathwani, AC, et al. (2011). Long-term Safety and Efficacy Following Systemic Administration of a Self-complementary AAV Vector Encoding Human FIX Pseudotyped With Serotype 5 and 8 Capsid Proteins. Molecular Therapy, 19(5), 876-885, Nguyen, GN, et al. (2021). A long-term study of AAV gene therapy in dogs with hemophilia A identifies clonal expansions of transduced liver cells. Nature Biotechnology, 39(1), 47-55, Niemeyer, GP, et al. (2009). Long-term correction of inhibitor-prone hemophilia B dogs treated with liver-directed AAV2-mediated factor IX gene therapy. Blood, 113(4), 797-806) could potentially be problematic even if off-target activity is very low, because the risk of editing at these sites increases with time of exposure.
[0154] In addition, persistence of off-target RNA deamination by base editors, although non-persistent, can alter the transcriptome profile of affected cells (Grunewald, J., et al. (2019). Transcriptome-wide off-target RNA editing induced by CRISPR-guided DNA base editors. Nature, 569 (7756), 433-437, Rees, HA, et al. (2019). Analysis and minimization of cellular RNA editing by DNA adenine base editors. Sci Adv, 5 (5), eaax5717, Zhou, C., et al. (2019). Off-target RNA mutation induced by DNA base editing and its elimination by mutagenesis. Nature, 571 (7764), 275-278). A mechanism for programmed self-inactivation of AAV-delivered Cas9 nuclease has been described previously, in which a transgene expressing Cas9 is targeted for double-stranded DNA breaks in addition to the on-target site in the host genome (Epstein, BE, & Schaffer, DV (2016). Engineering a Self-Inactivating CRISPR System for AAV Vectors. Molecular Therapy, 24, S50; Li, A., et al. (2019). A Self-Deleting AAV-CRISPR System for In Vivo Genome Editing. Mol Ther Methods Clin Dev, 12, 111-122). Thus, the instructions for Cas9 expression are removed from the cells to which Cas9 was originally delivered.
[0155] To realize the broadest therapeutic utility of base editing technology, the present invention provides a method to attenuate the activity and expression of base editors after delivery methods that may otherwise result in long-term expression. In contrast to CRISPR-Cas nucleases, base editors use either nCas9 or catalytically inactive "dead" variants (dCas9) to avoid the formation of indels resulting from unmodified Cas9 nucleases (Gaudelli, NM, et al. (2017). Programmable base editing of A*T to G*C in genomic DNA without DNA cleavage. Nature, 551(7681), 464-471; Komor, AC, et al. (2016). Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature, 533(7603), 420-424). Self-inactivation of base editors via the generation of double-stranded breaks in the DNA encoding is possible, but has several considerations. The nickase Cas9 in the base editor can be used to generate nicks on both strands that code the base editor DNA. The sites for each nick can occur close enough to favor dissociation of base-paired nucleotides, including up to the nick on each strand, resulting in blunt-end double-stranded DNA breaks. In addition, such a method may require these nicks to be made simultaneously, not sequentially, to avoid their religation, and may include at least two additional gRNAs to target the nicks. The base editors that incorporate dCas9 are not capable of using this strategy.Thus, in one embodiment, the invention provides methods that rely on making single base edits in editor DNA to reduce or eliminate further editing activity or expression, with the goal of minimizing the potential for both guide-dependent and guide-independent (e.g., pseudo-deamination) activity (Yu, Y., et al. (2020). Cytosine base editors with minimized unguided DNA and RNA off-target events and high on-target activity. Nature Communications, 11(1), 2052). The present invention also provides that any of the four sense codons CAA, CAG, CGA, or TGG encoding Gln, Arg, and Trp residues in the CBE can be directly converted to a stop codon via a single C-to-T base edit (Billon, P., et al. (2017). CRISPR-Mediated Base Editing Enables Efficient Disruption of Eukaryotic Genes through Induction of STOP Codons. Molecular Cell, 67(6), 1068-1079. e1064). However, achieving self-inactivation in the ABE requires an alternative approach, as sense codons cannot be converted to nonsense codons via A-to-G base edits.
[0156] The invention described herein features compositions and methods for facilitating the self-inactivation of base editors following cellular delivery of their encoding genetic material. The methods of the invention for self-inactivation of ABEs do not rely on direct conversion of a sense codon to a stop codon and can be adapted to inactivate CBEs using C-to-T single base editing. These compositions and methods use base editing to programmatically install single base mutations into DNA encoding the editor, resulting in the elimination or alteration of expression of DNA editing activity.
[0157] In one embodiment, the invention is based, at least in part, on the discovery that a guide RNA can direct a base editor to mutate active site residues in the deaminase subunit of the base editor, resulting in a catalytically inactive enzyme and loss of base editing activity. In another embodiment, the invention is also based, at least in part, on the discovery that targeting the start codon of a base editor for a single base mutation prevents translation.
[0158] In another embodiment, the invention is based, at least in part, on the discovery that an intron can be inserted within a base editor coding sequence (e.g., an open reading frame). The intron provides a sequence that can be targeted for base editing to disrupt or alter productive splicing of the base editor transcript (e.g., mRNA), resulting in loss of expression of the base editor (e.g., ABE, CBE). In some embodiments, base editing is made at the 5' or 3' end of the intron sequence (e.g., within a splice donor or splice acceptor site). Editing of target polynucleotide
[0159] The compositions of the invention are used, for example, to effect gene editing for a defined period of time. Once a desired level of editing is reached, expression of the base editor is reduced or eliminated by disrupting an intronic splice acceptor or donor site present in the polynucleotide sequence encoding the base editor.
[0160] Generally, base editing is performed to induce therapeutic changes in the genome of a cell of a subject. In some embodiments of the invention, a cell (in vivo or in vitro) is contacted with two or more guide RNAs and contacted with a nucleobase editor polypeptide comprising a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9), a deaminase (e.g., cytidine deaminase or adenosine deaminase). In some embodiments, the cell to be edited is contacted with at least one nucleic acid molecule, the at least one nucleic acid molecule encoding two or more guide RNAs and a nucleobase editor polypeptide, the nucleobase editor polypeptide comprising a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9) domain, a deaminase (e.g., cytidine deaminase or adenosine deaminase) domain, and a portion of the nucleic acid molecule encoding the nucleobase editor polypeptide comprises an intron, the intron comprising a splice acceptor or splice donor site. In some embodiments, the cell to be edited is contacted with at least one nucleic acid molecule, the at least one nucleic acid molecule encoding two or more guide RNAs and a nucleobase editor polypeptide, the nucleobase editor polypeptide comprising a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9) domain, a cytidine deaminase domain, and a portion of the nucleic acid molecule encoding the nucleobase editor polypeptide comprises an intron, the intron comprising a splice acceptor or a splice donor site. In some embodiments, the cell to be edited is contacted with at least one nucleic acid molecule, the at least one nucleic acid molecule encoding two or more guide RNAs and a nucleobase editor polypeptide, the nucleobase editor polypeptide comprising a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9) domain, an adenosine deaminase domain, and a portion of the nucleic acid molecule encoding the nucleobase editor polypeptide comprises an intron, the intron comprising a splice acceptor or a splice donor site.In some embodiments, at least one nucleic acid molecule encoding two or more guide RNAs and a nucleobase editor polypeptide are delivered to a cell by one or more vectors (e.g., AAV vectors).
[0161] In some embodiments, a cell to be edited is contacted with at least one nucleic acid molecule encoding two or more guide RNAs and contacted with at least two nucleic acid molecules encoding split nucleobase editor polypeptides, where one nucleic acid molecule encodes an N-terminal fragment of a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9) domain and a deaminase (e.g., cytidine deaminase or adenosine deaminase) domain fused to a split intein-N, and a second nucleic acid molecule encodes a C-terminal fragment of a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9) domain fused to a split intein-C, and either the first nucleic acid molecule or the second nucleic acid molecule comprises an intron, wherein the intron comprises a splice acceptor or a splice donor site. In some embodiments, a cell to be edited is contacted with at least one nucleic acid molecule encoding two or more guide RNAs and contacted with at least two nucleic acid molecules encoding split nucleobase editor polypeptides, where one nucleic acid molecule encodes an N-terminal fragment of a deaminase (e.g., cytidine deaminase or adenosine deaminase) domain fused to a split intein-N and a second nucleic acid molecule encodes a C-terminal fragment of a deaminase (e.g., cytidine deaminase or adenosine deaminase) domain and a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9) domain fused to a split intein-C, and where either the first nucleic acid molecule or the second nucleic acid molecule comprises an intron, where the intron comprises a splice acceptor or a splice donor site.
[0162] In some embodiments, at least one nucleic acid molecule encoding two or more guide RNAs and the first and second nucleic acid molecules encoding the split nucleobase editor polypeptide are delivered to a cell by one or more vectors (e.g., AAV vectors). In some embodiments, at least one nucleic acid molecule encoding two or more guide RNAs and the first and second nucleic acid molecules encoding the split nucleobase editor polypeptide are each delivered to a cell by a separate vector (e.g., AAV vector). In some embodiments, at least one nucleic acid molecule encoding two or more guide RNAs and the first and second nucleic acid molecules encoding the split nucleobase editor polypeptide are delivered to a cell in the same vector (e.g., AAV vector).
[0163] In some embodiments, the nucleic acid molecule encoding the nucleobase editor polypeptide comprises a linker. In some embodiments, an intron is inserted within an open reading frame in a nucleic acid molecule encoding the nucleobase editor polypeptide. In some embodiments, the intron is inserted within a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9) domain, a deaminase (e.g., cytidine deaminase or adenosine deaminase) domain, or a linker. In some embodiments, the intron is inserted adjacent to a protospacer sequence. In some embodiments, the intron is inserted within about 10-30 base pairs of the protospacer sequence. In some embodiments, the protospacer sequence is NGG or NNGRRT. In some embodiments, the intron is about 10 base pairs to about 500 base pairs in length. In some embodiments, the intron is about 70 base pairs to 150 base pairs. In some embodiments, the intron is about 100 base pairs to 200 base pairs.
[0164] In some embodiments, the two or more guide RNAs comprise one or more guide RNAs that direct a nucleobase editor polypeptide to edit a site in the genome of a cell, and one or more guide RNAs that direct the nucleobase editor polypeptide to edit a splice acceptor or splice donor site (e.g., A-to-G or C-to-T base editing) present within an intron of a nucleic acid encoding the nucleobase editor polynucleotide. In some embodiments, the gRNA comprises a nucleotide analog. These nucleotide analogs can inhibit degradation of the gRNA by cellular processes.
[0165] In various cases, it is advantageous for the spacer sequence to include a 5' and / or 3' "G" nucleotide. In some cases, for example, any spacer sequence or guide polynucleotide provided herein includes or further includes a 5' "G", and in some embodiments, the 5' "G" is complementary or not complementary to the target sequence. In some embodiments, the 5' "G" is added to a spacer sequence that does not already contain a 5' "G". For example, when a guide RNA is expressed under the control of, such as, a U6 promoter, it may be advantageous for the guide RNA to include a 5' terminal "G" because the U6 promoter prefers a "G" at the transcription start site (see Cong, L. et al. "Multiplex genome engineering using CRISPR / Cas systems. Science 339:819-823 (2013) doi:10.1126 / science.1231143). In some cases, a 5' terminal "G" is added to a guide polynucleotide that is expressed under the control of a promoter, but optionally is not added to a guide polynucleotide if or when the guide polynucleotide is not expressed under the control of a promoter.
[0166] In some embodiments, the base editing of the present invention is performed in vivo in a subject. In some embodiments, one or more vectors (e.g., AAV vectors) comprising at least one nucleic acid molecule, wherein the at least one nucleic acid molecule encodes two or more guide RNAs and a nucleobase editor polypeptide, wherein the nucleobase editor polypeptide comprises a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9) domain, a deaminase (e.g., a cytidine deaminase or adenosine deaminase) domain, and wherein a portion of the nucleic acid molecule encoding the nucleobase editor polypeptide comprises an intron, wherein the intron comprises a splice acceptor or splice donor site, are delivered to cells in a subject in vivo.
[0167] In some embodiments, one or more vectors (e.g., AAV vectors) comprising at least one nucleic acid molecule encoding one or more guide RNAs, where the one or more guide RNAs direct a nucleobase editor polypeptide to edit a site within the genome of a cell, and at least one nucleic acid molecule encoding a nucleobase editor polypeptide, where the nucleobase editor polypeptide comprises a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9) domain, a deaminase (e.g., cytidine deaminase or adenosine deaminase) domain, and an intron, where the intron comprises a splice acceptor or a splice donor site, are delivered in vivo to a cell in a subject to edit the site within the genome of the cell. In some embodiments, once a desired level of base editing is achieved in a subject, one or more vectors (e.g., AAV vectors) comprising at least one nucleic acid molecule encoding one or more guide RNAs, where the one or more guide RNAs are targeted to edit a splice acceptor or splice donor site present within an intron of a nucleic acid molecule encoding a nucleobase editor polynucleotide, are delivered in vivo to cells in the subject to edit the splice acceptor or splice donor site within the intron of the nucleic acid molecule encoding the nucleobase editor polynucleotide (e.g., an A-to-G or C-to-T base editing), thereby self-inactivating the nucleobase editor polynucleotide and reducing or eliminating base editing activity.
[0168] In some embodiments, one or more vectors (e.g., AAV vectors) comprising at least one nucleic acid molecule encoding two or more guide RNAs and at least two nucleic acid molecules encoding split nucleobase editor polypeptides, where one nucleic acid molecule encodes an N-terminal fragment of a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9) domain and a deaminase (e.g., cytidine deaminase or adenosine deaminase) domain fused to a split intein-N and a second nucleic acid molecule encodes a C-terminal fragment of a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9) domain fused to a split intein-C, where either the first or second nucleic acid molecule comprises an intron, where the intron comprises a splice acceptor or splice donor site, are delivered in vivo to cells in a subject. In some embodiments, one or more vectors (e.g., AAV vectors) comprising at least one nucleic acid molecule encoding two or more guide RNAs and at least two nucleic acid molecules encoding split nucleobase editor polypeptides, where one nucleic acid molecule encodes an N-terminal fragment of a deaminase (e.g., cytidine deaminase or adenosine deaminase) domain fused to a split intein-N and a second nucleic acid molecule encodes a C-terminal fragment of a deaminase (e.g., cytidine deaminase or adenosine deaminase) domain and a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9) domain fused to a split intein-C, wherein either the first or second nucleic acid molecule comprises an intron, wherein the intron comprises a splice acceptor or a splice donor site, are delivered in vivo to cells in a subject.
[0169] In some embodiments, the present invention comprises at least one nucleic acid molecule encoding one or more guide RNAs, where the one or more guide RNAs direct a nucleobase editor polypeptide to edit a site in the genome of a cell, and at least two nucleic acid molecules encoding split nucleobase editor polypeptides, where one nucleic acid molecule comprises an N-terminal fragment of a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9) domain and a deaminase (e.g., a cytidine deaminase) fused to a split intein-N. One or more vectors (e.g., AAV vectors) comprising at least two nucleic acid molecules, where a first nucleic acid molecule encodes a C-terminal fragment of a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9) domain fused to a split intein-C, and a second nucleic acid molecule encodes a C-terminal fragment of a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9) domain fused to a split intein-C, where either the first nucleic acid molecule or the second nucleic acid molecule comprises an intron, the intron comprising a splice acceptor or a splice donor site, are delivered in vivo to a cell in a subject to edit a site in the genome of the cell.In some embodiments, the present invention comprises at least one nucleic acid molecule encoding one or more guide RNAs, where the one or more guide RNAs direct a nucleobase editor polypeptide to edit a site in the genome of a cell, and at least two nucleic acid molecules encoding split nucleobase editor polypeptides, where one nucleic acid molecule encodes an N-terminal fragment of a deaminase (e.g., cytidine deaminase or adenosine deaminase) domain fused to a split intein-N and a second nucleic acid molecule encodes an N-terminal fragment of a deaminase (e.g., cytidine deaminase or adenosine deaminase) domain fused to a split intein-N. One or more vectors (e.g., AAV vectors) comprising at least two nucleic acid molecules encoding a C-terminal fragment of a deaminase (e.g., cytidine deaminase or adenosine deaminase) domain and a nucleic acid programmable DNA binding protein (napDNAbp) (e.g., Cas9) domain fused to In-C, where either the first nucleic acid molecule or the second nucleic acid molecule comprises an intron, the intron comprising a splice acceptor or splice donor site, are delivered in vivo to a cell in a subject to edit a site in the genome of the cell. When the one or more vectors (e.g., AAV vectors) are delivered to the cell, the cell expresses the N-terminal and C-terminal fragments of the split nucleobase editor polypeptide, and the N-terminal and C-terminal fragments combine together to form the nucleobase editor polypeptide. In some embodiments, once a desired level of base editing is achieved in a subject, one or more vectors (e.g., AAV vectors) comprising at least one nucleic acid molecule encoding one or more guide RNAs, where the one or more guide RNAs are targeted to edit a splice acceptor or splice donor site present within an intron of a nucleic acid molecule encoding a nucleobase editor polynucleotide, are delivered in vivo to cells in the subject to edit the splice acceptor or splice donor site present within the nucleic acid molecule encoding an intron of the nucleobase editor polynucleotide (e.g., an A-to-G or C-to-T base edit), thereby self-inactivating the nucleobase editor polynucleotide and reducing or eliminating base editing activity.
[0170] The present invention provides a method of treating a patient, for example, with a disease and a SNP of interest, by administering two AAV vectors containing the split intein base editor system provided herein. In some embodiments, the AAV vectors each encode the following portions of a base editor: an N-terminal portion fused to intein-N and a C-terminal portion fused to intein-C. An intron sequence is encoded in the coding sequence of one or more of the two halves of the base editor. In some embodiments, a guide RNA targeting the SNP is also included in one of the AAV vectors. In some embodiments, the AAV vector has a tropism associated with a diseased cell, tissue, or organ (e.g., the AAV vector is of a single serotype). When a cell is infected with two AAV vectors of the base editing system, a transcript encoding the two halves of the base editor is expressed and the intron is spliced out and removed. Upon expression of the two half polypeptides, the base editor is reconstituted by intracellular protein splicing via the split intein tag. In some embodiments, after base editing is performed for a period of time to allow base editing to occur, a third AAV is provided, the third AAV encoding a guide RNA, which targets a donor or acceptor splice site in an intron with a base editor in the cell. When a cell expressing the base editor is infected with this AAV, the AAV modifies the splice site to prevent splicing from occurring. Because a portion of the base editor is not properly expressed, base editing is inactivated or attenuated at on-target and off-target sites in the cell.
[0171] The present invention also provides guide RNAs that target introns of polynucleotides encoding self-inactivating base editors. Table 1A provides target intron sequences that are used to target gRNAs to intron acceptor or donor sites. [Table 1A-1] [Table 1A-2]
[0172] Table 1B provides gRNA sequences for targeting intron acceptor or donor sites. In some embodiments, the gRNA sequences are expressed from a U6 promoter. The lowercase "g" in Table 1B below indicates a 5' mismatch to the target sequence. [Table 1B-1] [Table 1B-2] [Table 1B-3] [Table 1B-4] [Table 1B-5] [Table 1B-6] [Table 1B-7]
[0173] In some embodiments, the deaminase domain is a TadA domain. In some embodiments, the intron is inserted within or immediately after the codon of TadA. In some embodiments, the intron is inserted within or immediately after codon 18, 23, 59, 62, 87, or 129 of TadA. In some embodiments, the intron is inserted immediately after codon 87 of TadA.
[0174] Table 1C below provides target sequence coordinates for inserting an intron into the TadA open reading frame (e.g., c.100+1 indicates that the first base pair of the intron sequence was immediately following the 100th coding nucleotide of TadA). Thus, in some embodiments, the intron sequence is positioned immediately following the specified amino acid position. In other embodiments, the intron sequence is positioned immediately before the specified amino acid position. [Table 1C-1] [Table 1C-2] [Table 1C-3]
[0175] Nucleic acid base editor Nucleobase editors that edit, modify or alter a target nucleotide sequence of a polynucleotide (e.g., a self-inactivating nucleobase editor) are useful in the methods and compositions described herein. The nucleobase editors described herein generally comprise a polynucleotide programmable nucleotide binding domain and a nucleobase editing domain (e.g., an adenosine deaminase or a cytidine deaminase). The polynucleotide programmable nucleotide binding domain, when combined with a bound guide polynucleotide (e.g., a gRNA), can specifically bind to a target polynucleotide sequence, thereby allowing the base editor to localize to the target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence is present within an intron (e.g., a splice acceptor or splice donor site).
[0176] In certain embodiments, the nucleobase editors provided herein include one or more features that improve base editing activity. For example, any of the nucleobase editors provided herein may include a Cas9 domain with reduced nuclease activity. In some embodiments, any of the nucleobase editors provided herein may have a Cas9 domain that does not have nuclease activity (dCas9), or a Cas9 domain that cleaves one strand of a double-stranded DNA molecule, called Cas9 nickase (nCas9). Without wishing to be bound by any particular theory, the presence of a catalytic residue (e.g., H840) maintains the activity of Cas9 to cleave the non-edited (e.g., non-deaminated) strand opposite the target nucleobase. Mutation of the catalytic residue (e.g., D10 to A10) prevents cleavage of the edited (e.g., deaminated) strand that contains the target residue (e.g., A or C). Such Cas9 variants generate single-stranded DNA breaks (nicks) at specific locations based on the target sequence defined by the gRNA, triggering repair of the non-edited strand and ultimately altering nucleobases on the non-edited strand.
[0177] Polynucleotide Programmable Nucleotide Binding Domains The polynucleotide programmable nucleotide binding domain binds to a polynucleotide (e.g., RNA, DNA). In some embodiments, an intron is present in the open reading frame encoding the nucleotide programmable nucleotide binding domain of the base editor. The polynucleotide programmable nucleotide binding domain of the base editor may itself comprise one or more domains (e.g., one or more nuclease domains). In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain may comprise an endonuclease or an exonuclease. An endonuclease can cleave one strand of a double-stranded nucleic acid molecule, or both strands of a double-stranded nucleic acid molecule. In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain can cleave zero, one, or two strands of a target polynucleotide.
[0178] Non-limiting examples of polynucleotide programmable nucleotide binding domains that can be incorporated into base editors include CRISPR protein-derived domains, restriction nucleases, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs). In some embodiments, the base editor comprises a polynucleotide programmable nucleotide binding domain comprising a natural or modified protein or a portion thereof, and can bind to a nucleic acid sequence during CRISPR (i.e., clustered regularly interspaced short palindromic repeats)-mediated modification of the nucleic acid via a bound guide nucleic acid. Such proteins are referred to herein as "CRISPR proteins." Thus, disclosed herein are base editors that comprise a polynucleotide programmable nucleotide binding domain comprising all or a portion of a CRISPR protein (i.e., a base editor that comprises all or a portion of a CRISPR protein as a domain, also referred to as the "CRISPR protein-derived domain" of the base editor). The CRISPR protein-derived domain incorporated into the base editor can be modified compared to a wild-type or natural version of the CRISPR protein. For example, as described below, a domain derived from a CRISPR protein can contain one or more mutations, insertions, deletions, rearrangements, and / or modifications relative to a wild-type or naturally occurring version of the CRISPR protein.
[0179] Cas proteins that may be used herein include class 1 and class 2. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Cs x17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1 (e.g., SEQ ID NO: 320), Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ, CARF, DinG, homologs thereof, or modified versions thereof. CRISPR enzymes can direct cleavage of one or both strands at a target sequence, e.g., within the target sequence and / or within a complementary sequence of the target sequence. For example, CRISPR enzymes can direct cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence.
[0180] A vector can be used that encodes a CRISPR enzyme that is mutated relative to the corresponding wild-type enzyme, such that the mutant CRISPR enzyme lacks the ability to cleave one or both strands of a target polynucleotide that contains a target sequence. A Cas protein (e.g., Cas9, Cas12) or Cas domain (e.g., Cas9, Cas12) can refer to a polypeptide or domain that has at least or at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology to a wild-type exemplary Cas polypeptide or Cas domain. Cas (e.g., Cas9, Cas12) can refer to a wild-type or modified form of a Cas protein that can include amino acid changes such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof.
[0181] In some embodiments, the base editor CRISPR protein-derived domains are selected from the group consisting of Corynebacterium ulcerans (NCBI References: NC_015683.1, NC_017317.1), Corynebacterium diphtheria (NCBI References: NC_016782.1, NC_016786.1), Spiroplasma syrphidicola (NCBI Reference: NC_021284.1), Prevotella intermedia (NCBI Reference: NC_017861.1), Spiroplasma taiwanense (NCBI Reference: NC_021846.1), Streptococcus iniae (NCBI Reference: NC_021314.1), Belliella baltica (NCBI Reference: NC_018010.1), Psychroflexus The Cas9 may include all or a portion of Cas9 from C. torquis (NCBI Reference: NC_018721.1), Streptococcus thermophilus (NCBI Reference: YP_820832.1), Listeria innocua (NCBI Reference: NP_472073.1), Campylobacter jejuni (NCBI Reference: YP_002344900.1), Neisseria meningitidis (NCBI Reference: YP_002342100.1), Streptococcus pyogenes, or Staphylococcus aureus.
[0182] The sequence and structure of Cas9 nuclease are well known to those of skill in the art (see, e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes," Ferretti et al., Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III," Deltcheva E., et al., Nature 471:602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity," Jinek M., et al., Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including, but not limited to, S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to the skilled artisan based on this disclosure, and include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737, the entire contents of which are incorporated herein by reference.
[0183] High-fidelity Cas9 domain Some aspects of the present disclosure provide high-fidelity Cas9 domains. High-fidelity Cas9 domains are known in the art and are described, for example, in Kleinstiver, BP, et al. "High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects." Nature 529, 490-495 (2016) and Slaymaker, IM, et al. "Rationally engineered Cas9 nucleases with improved specificity." Science 351, 84-88 (2015), the entire contents of each of which are incorporated herein by reference. An exemplary high-fidelity Cas9 domain is shown in the sequence listing as SEQ ID NO: 321. In some embodiments, the high-fidelity Cas9 domain is a modified Cas9 domain that contains one or more mutations that reduce the electrostatic interaction between the Cas9 domain and the sugar-phosphate backbone of DNA relative to the corresponding wild-type Cas9 domain. High fidelity Cas9 domains with reduced electrostatic interactions with the sugar phosphate backbone of DNA have fewer off-target effects. In some embodiments, the Cas9 domain (e.g., wild-type Cas9 domain (SEQ ID NOs: 250 and 253)) comprises one or more mutations that reduce the association between the Cas9 domain and the sugar phosphate backbone of DNA. In some embodiments, the Cas9 domain comprises one or more mutations that reduce the association between the Cas9 domain and the sugar phosphate backbone of DNA by at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, or at least 70%.
[0184] In some embodiments, any of the Cas9 fusion proteins provided herein comprises one or more of D10A, N497X, R661X, Q695X and / or Q926X mutations, or corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the high-fidelity Cas9 enzyme is SpCas9(K855A), eSpCas9(1.1), SpCas9-HF1, or hyper-precise Cas9 variant (HypaCas9). In some embodiments, the modified Cas9, eSpCas9(1.1), contains an alanine substitution, which weakens the interaction between the HNH / RuvC groove and non-target DNA strands, preventing strand separation and cleavage at off-target sites. Similarly, SpCas9-HF1 reduces off-target editing via an alanine substitution that disrupts the interaction of Cas9 with the DNA phosphate backbone. HypaCas9 contains mutations in the REC3 domain (SpCas9 N692A / M694A / Q695A / H698A) that enhance Cas9 proofreading and target discrimination. All three high-fidelity enzymes generate fewer off-target edits than wild-type Cas9.
[0185] Reduced exclusivity of Cas9 domains Typically, Cas9 proteins, such as Cas9 from S. pyogenes (spCas9), require a "protospacer adjacent motif (PAM)" or PAM-like motif, which is a DNA sequence of 2-6 base pairs immediately following the DNA sequence targeted by the Cas9 nuclease in the CRISPR bacterial adaptive immune system. The presence of the NGG PAM sequence is required to bind to a specific nucleic acid region, where the "N" in "NGG" is adenosine (A), thymidine (T) or cytosine (C), and the G is guanosine. This may limit the ability to edit a desired base in a genome. In some embodiments, the base editing fusion proteins provided herein may need to be placed at a precise location (e.g., a region containing the target base upstream of the PAM). See, e.g., Komor, A.C., et al., "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage," Nature 533, 420-424 (2016), the entire contents of which are incorporated herein by reference. Exemplary polypeptide sequences of spCas9 proteins capable of binding to PAM sequences are set forth in the Sequence Listing as SEQ ID NOs: 250, 254, and 322-325. Thus, in some embodiments, any of the fusion proteins provided herein can include a Cas9 domain capable of binding to a nucleotide sequence that does not contain a canonical (e.g., NGG) PAM sequence. Cas9 domains that bind to non-canonical PAM sequences have been described in the art and would be apparent to one of ordinary skill in the art.For example, Cas9 domains that bind to non-canonical PAM sequences are described in Kleinstiver, BP, et al., "Engineered CRISPR-Cas9 nucleases with altered PAM specificities," Nature 523, 481-485 (2015), and Kleinstiver, BP, et al., "Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition," Nature Biotechnology 33, 1293-1298 (2015), the entire contents of each of which are incorporated herein by reference.
[0186] Nickase In some embodiments, the polynucleotide programmable nucleotide binding domain can include a nickase domain. As used herein, the term "nickases" refers to polynucleotide programmable nucleotide binding domains that include a nuclease domain that can cleave only one of the two strands of a double-stranded nucleic acid molecule (e.g., DNA). In some embodiments, a nickase can be derived from a fully catalytically active (e.g., native) form of a polynucleotide programmable nucleotide binding domain by introducing one or more mutations into the active polynucleotide programmable nucleotide binding domain. For example, when a polynucleotide programmable nucleotide binding domain includes a nickase domain derived from Cas9, the Cas9-derived nickase domain can include a D10A mutation and a histidine at position 840. In such an embodiment, residue H840 retains catalytic activity and can thereby cleave one strand of a nucleic acid duplex. In another example, a nickase domain derived from Cas9 can include a H840A mutation, but the amino acid residue at position 10 remains D. In some embodiments, the nickase may be derived from a fully catalytically active (e.g., native) form of a polynucleotide programmable nucleotide binding domain by removing all or part of a nuclease domain that is not required for nickase activity. For example, if the polynucleotide programmable nucleotide binding domain comprises a nickase domain derived from Cas9, the Cas9-derived nickase domain may comprise a deletion of all or part of the RuvC domain or the HNH domain.
[0187] In some embodiments, the wild-type Cas9 corresponds to or comprises the following amino acid sequence: TIFF2024521750000021.tif136165
[0188] In some embodiments, the strand of a nucleic acid duplex target polynucleotide sequence that is cleaved by a base editor that comprises a nickase domain (e.g., a Cas9-derived nickase domain, a Cas12-derived nickase domain) is the strand that is not edited by the base editor (i.e., the strand that is cleaved by the base editor is opposite the strand that contains the base to be edited). In other embodiments, a base editor that comprises a nickase domain (e.g., a Cas9-derived nickase domain, a Cas12-derived nickase domain) can cleave the strand of a DNA molecule that is targeted for editing. In such embodiments, the non-target strand is not cleaved.
[0189] In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., a nickase referred to as a "nCas9" protein ("nickase" Cas9). The Cas9 nickase can be a Cas9 protein that can cleave only one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase cleaves the target strand of the double-stranded nucleic acid molecule, meaning that the Cas9 nickase cleaves the strand that is base-paired (complementary) to the gRNA (e.g., sgRNA) that is bound to the Cas9. In some embodiments, the Cas9 nickase comprises a D10A mutation and has a histidine at position 840. In some embodiments, the Cas9 nickase cleaves the non-targeted, un-base-edited strand of the double-stranded nucleic acid molecule, meaning that the Cas9 nickase cleaves the strand that is not base-paired to the gRNA (e.g., sgRNA) that is bound to the Cas9. In some embodiments, the Cas9 nickase comprises an H840A mutation and has an aspartic acid residue at position 10, or a corresponding mutation. In some embodiments, the Cas9 nickase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the Cas9 nickases provided herein. Additional suitable Cas9 nickases will be apparent to one of skill in the art based on this disclosure and knowledge in the art, and are within the scope of this disclosure.
[0190] The amino acid sequence of an exemplary catalytic Cas9 nickase (nCas9) is as follows:
[0191] The Cas9 nuclease has two functional endonuclease domains (RuvC and HNH). Upon target binding, Cas9 undergoes a conformational change that positions the nuclease domains to cleave opposite strands of the target DNA. The end result of Cas9-mediated DNA cleavage is a double-strand break (DSB) in the target DNA (approximately 3-4 nucleotides upstream of the PAM sequence). The resulting DSB is then repaired by one of two general repair pathways: (1) the efficient but error-prone non-homologous end joining (NHEJ) pathway, or (2) the less efficient but high fidelity homology-directed repair (HDR) pathway.
[0192] The "efficiency" of non-homologous end joining (NHEJ) and / or homology directed repair (HDR) can be calculated by any convenient method. For example, in some embodiments, the efficiency can be expressed as a percentage of successful HDR. For example, Surveyor Nuclease assay can be used to generate cleavage products, and the ratio of products to substrate can be used to calculate the percentage. For example, Surveyor Nuclease enzyme can be used to directly cleave DNA containing the newly incorporated restriction sequence resulting from successful HDR. Cleavage of more substrates indicates a higher percentage of HDR (higher efficiency of HDR). As an illustrative example, the rate (percentage) of HDR can be calculated using the following formula [(cleavage product) / (substrate+cleavage product)] (e.g., (b+c) / (a+b+c), where "a" is the band intensity of the DNA substrate, and "b" and "c" are the band intensities of the cleavage products).
[0193] In some embodiments, efficiency may be expressed as a percentage of successful NHEJ. For example, a T7 endonuclease I assay can be used to generate cleavage products, and the ratio of products to substrates can be used to calculate the percentage of NHEJ. T7 endonuclease I cleaves mismatched heteroduplex DNA resulting from hybridization of wild-type and mutant DNA strands (NHEJ generates small random insertions or deletions (indels) at the original cleavage site). More cleavage indicates a higher percentage of NHEJ (higher efficiency of NHEJ). As an illustrative example, the percentage of NHEJ can be calculated using the following formula: (1-(1-(b+c) / (a+b+c)). 1 / 2 ) × 100, where "a" is the band intensity of the DNA substrate and "b" and "c" are the cleavage products (Ran et al., Cell. 2013 Sep. 12; 154(6): 1380-9 and Ran et al., Nat Protoc. 2013 Nov.; 8(11): 2281-2308).
[0194] The NHEJ repair pathway is the most active repair mechanism, frequently causing small nucleotide insertions or deletions (indels) at DSB sites. The randomness of NHEJ-mediated DSB repair has important practical implications, since a cell population expressing Cas9 and gRNA, or guide polynucleotide, can result in diverse mutations. In most embodiments, NHEJ generates small indels in the target DNA, resulting in amino acid deletions, insertions, or frameshift mutations that lead to premature stop codons in the open reading frame (ORF) of the target gene. The ideal end result is a loss-of-function mutation in the target gene.
[0195] While NHEJ-mediated DSB repair often disrupts the open reading frame of a gene, homology-directed repair (HDR) can be used to generate specific nucleotide changes ranging from single nucleotide changes to large insertions (e.g., addition of fluorophores or tags).
[0196] To utilize HDR for gene editing, a DNA repair template containing the desired sequence can be delivered to a cell type of interest using gRNA(s) and Cas9 or Cas9 nickase. The repair template can contain the desired edit as well as additional homologous sequences immediately upstream and downstream of the target (referred to as left and right homology arms). The length of each homology arm can depend on the size of the change to be introduced, with larger insertions requiring longer homology arms. The repair template can be a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a double-stranded DNA plasmid. Even in cells expressing Cas9, gRNA, and an exogenous repair template, the efficiency of HDR is generally low (less than 10% of modified alleles). Because HDR takes place during the S and G2 phases of the cell cycle, the efficiency of HDR can be enhanced by synchronizing cells. Chemical or genetic inhibition of genes involved in NHEJ can also increase HDR frequency.
[0197] In some embodiments, the Cas9 is a modified Cas9. A given gRNA targeting sequence may have additional sites throughout the genome where partial homology exists. These sites are called off-targets and need to be considered when designing the gRNA. In addition to optimizing gRNA design, the specificity of CRISPR can also be increased through modifications to Cas9. Cas9 generates double-strand breaks (DSBs) through the combined activity of two nuclease domains, RuvC and HNH. Cas9 nickase, a D10A mutant of SpCas9, retains one nuclease domain and generates DNA nicks instead of DSBs. The nickase system can also be combined with HDR-mediated gene editing for specific gene editing.
[0198] Catalytically inactive nucleases Also provided herein are base editors that include a polynucleotide programmable nucleotide binding domain that is catalytically inactive (i.e., unable to cleave a target polynucleotide sequence). As used herein, the terms "catalytically dead" and "nuclease dead" are used interchangeably to refer to a polynucleotide programmable nucleotide binding domain that has one or more mutations and / or deletions that result in an inability to cleave a strand of nucleic acid. In some embodiments, a catalytically inactive polynucleotide programmable nucleotide binding domain base editor may lack nuclease activity as a result of specific point mutations in one or more nuclease domains. For example, in the case of a base editor that includes a Cas9 domain, Cas9 may include both a D10A mutation and an H840A mutation. Such mutations inactivate both nuclease domains, thereby resulting in loss of nuclease activity. In other embodiments, a catalytically inactive polynucleotide programmable nucleotide binding domain may include one or more deletions of all or part of a catalytic domain (e.g., RuvC1 and / or HNH domain). In further embodiments, the catalytically inactive polynucleotide programmable nucleotide binding domain comprises a point mutation (e.g., D10A or H840A) and a deletion of all or part of the nuclease domain. dCas9 domains are known in the art and are described, for example, in Qi et al., "Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression." Cell. 2013;152(5):1173-83, the entire contents of which are incorporated herein by reference.
[0199] Additional suitable nuclease-inactive dCas9 domains will be apparent to those of skill in the art based on this disclosure and knowledge in the art and are within the scope of this disclosure. Additional exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013;31(9):833-838, the entire contents of which are incorporated herein by reference).
[0200] In some embodiments, the dCas9 corresponds to, or partially or entirely comprises, a Cas9 amino acid sequence having one or more mutations that inactivate Cas9 nuclease activity. In some embodiments, the nuclease-inactive dCas9 domain comprises a D10X mutation and a H840X mutation in the amino acid sequence described herein, or a corresponding mutation in any amino acid sequence provided herein, where X is any amino acid change. In some embodiments, the nuclease-inactive dCas9 domain comprises a D10A mutation and a H840A mutation in the amino acid sequence described herein, or a corresponding mutation in any amino acid sequence provided herein. In some embodiments, the nuclease-inactive Cas9 domain comprises the amino acid sequence described in the cloning vector pPlatTET-gRNA2 (Accession No. BAV54124).
[0201] In some embodiments, variant Cas9 protein can cleave the complementary strand of guide target sequence, but has a reduced ability to cleave the non-complementary strand of double-stranded guide target sequence. For example, variant Cas9 protein can have a mutation (amino acid substitution) that reduces the function of RuvC domain. As a non-limiting example, in some embodiments, variant Cas9 protein has D10A (aspartic acid to alanine at amino acid position 10), and thus can cleave the complementary strand of double-stranded guide target sequence, but has a reduced ability to cleave the non-complementary strand of double-stranded guide target sequence (thus, when variant Cas9 protein cleaves double-stranded target nucleic acid, it generates single-strand break (SSB) instead of double-strand break (DSB)) (see, for example, Jinek et al., Science. 2012 Aug. 17; 337 (6096): 816-21).
[0202] In some embodiments, the variant Cas9 protein can cleave the non-complementary strand of the double-stranded guide target sequence, but has a reduced ability to cleave the complementary strand of the guide target sequence. For example, the variant Cas9 protein can have a mutation (amino acid substitution) that reduces the function of the HNH domain (RuvC / HNH / RuvC domain motif). As a non-limiting example, in some embodiments, the variant Cas9 protein has H840A (histidine to alanine at amino acid position 840), and thus can cleave the non-complementary strand of the guide target sequence, but has a reduced ability to cleave the complementary strand of the guide target sequence (so that when the variant Cas9 protein cleaves the double-stranded guide target sequence, an SSB is generated instead of a DSB). Such a Cas9 protein has a reduced ability to cleave the guide target sequence (e.g., a single-stranded guide target sequence), but retains the ability to bind to the guide target sequence (e.g., a single-stranded guide target sequence).
[0203] As another non-limiting example, in some embodiments, a variant Cas9 protein has a W476A and a W1126A mutation, which reduces the ability of the polypeptide to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA).
[0204] As another non-limiting example, in some embodiments, a variant Cas9 protein has P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations that result in the polypeptide having a reduced ability to cleave target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA).
[0205] As another non-limiting example, in some embodiments, a variant Cas9 protein has H840A, W476A, and W1126A mutations, which reduces the ability of the polypeptide to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some embodiments, a variant Cas9 protein has H840A, D10A, W476A, and W1126A mutations, which reduces the ability of the polypeptide to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, a variant Cas9 has a catalytic His residue restored to position 840 of the Cas9 HNH domain (A840H).
[0206] As another non-limiting example, in some embodiments, a variant Cas9 protein has H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, which results in the polypeptide having a reduced ability to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some embodiments, a variant Cas9 protein has D10A, H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, which results in the polypeptide having a reduced ability to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, when the variant Cas9 protein has W476A and W1126A mutations, or when the variant Cas9 protein has P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, the variant Cas9 protein does not bind efficiently to the PAM sequence. Thus, in some such embodiments, when such a variant Cas9 protein is used in a binding method, the method does not require a PAM sequence. In other words, in some embodiments, when such a variant Cas9 protein is used in a binding method, the method can include a guide RNA, but the method can be performed in the absence of a PAM sequence (thus, the specificity of binding is provided by the targeting segment of the guide RNA). Other residues can be mutated to achieve the above effects (i.e., to inactivate one or other nuclease moieties). As non-limiting examples, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 can be altered (i.e., substituted). Mutations other than alanine substitutions are also suitable.
[0207] In some embodiments, a variant Cas9 protein with reduced catalytic activity (e.g., where the Cas9 protein has a D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 mutation (e.g., D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A)) can still bind to target DNA in a site-specific manner (as it is still guided to the target DNA sequence by the guide RNA) so long as it retains the ability to interact with the guide RNA.
[0208] In some embodiments, the variant Cas protein can be spCas9, spCas9-VRQR, spCas9-VRER, xCas9(sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL.
[0209] In some embodiments, the Cas9 domain is a Cas9 domain from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is a nuclease-active SaCas9, a nuclease-inactive SaCas9 (SaCas9d), or a SaCas9 nickase (SaCas9n). In some embodiments, the SaCas9 comprises a N579A mutation or a corresponding mutation in any of the amino acid sequences provided in the sequence listing submitted herewith.
[0210] In some embodiments, the SaCas9 domain, the SaCas9d domain, or the SaCas9n domain can bind to a nucleic acid sequence having a non-canonical PAM. In some embodiments, the SaCas9 domain, the SaCas9d domain, or the SaCas9n domain can bind to a nucleic acid sequence having an NNGRRT or NNGRRV PAM sequence. In some embodiments, the SaCas9 domain comprises one or more of E781X, N967X, and R1014X mutations, or corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SaCas9 domain comprises one or more of E781K, N967K, and R1014H mutations, or corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SaCas9 domain comprises E781K, N967K, or R1014H mutations, or corresponding mutations in any of the amino acid sequences provided herein.
[0211] In some embodiments, one of the Cas9 domains present in the fusion protein can be replaced with a guide nucleotide sequence programmable DNA binding protein domain that does not require a PAM sequence. In some embodiments, the Cas9 is SaCas9. Residue A579 of SaCas9 can be mutated from N579 to obtain SaCas9 nickase. Residues K781, K967, and H1014 can be mutated from E781, N967, and R1014 to obtain SaKKH Cas9.
[0212] In some embodiments, a modified SpCas9 was used that contains the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (SpCas9-MQKFRAER) and has specificity for the engineered PAM 5'-NGC-3'.
[0213] Alternatives to S. pyogenes Cas9 include RNA-guided endonucleases from the Cpf1 family that show cleavage activity in mammalian cells. CRISPR from Prevotella and Francisella 1 (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. Cpf1 is an RNA-guided endonuclease of the class II CRISPR / Cas system. This acquired immunity mechanism is found in Prevotella and Francisella bacteria. The Cpf1 gene is associated with the CRISPR locus and encodes an endonuclease that uses guide RNA to find and cleave viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9, overcoming some of the limitations of the CRISPR / Cas9 system. Unlike Cas9 nuclease, Cpf1-mediated DNA cleavage results in double-stranded breaks with short 3' overhangs. The staggered cleavage pattern of Cpf1 opens the possibility of directional gene transfer similar to conventional restriction enzyme cloning, and can increase the efficiency of gene editing. Similar to the above-mentioned Cas9 variants and orthologs, Cpf1 can also extend the number of sites that can be targeted by CRISPR to AT-rich regions or AT-rich genomes that lack NGG PAM sites suitable for SpCas9. The Cpf1 locus contains a mixed alpha / beta domain, RuvC-I, followed by a helical region, RuvC-II, and a zinc finger-like domain. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9.
[0214] Furthermore, unlike Cas9, Cpf1 does not have an HNH endonuclease domain, and the N-terminus of Cpf1 does not have the alpha-helical recognition lobe of Cas9. The Cpf1 CRISPR-Cas domain architecture indicates that Cpf1 is functionally unique and is classified as a class 2 type V CRISPR system. The Cpf1 locus encodes Cas1, Cas2, and Cas4 proteins that are more similar to type I and type III systems than to type II systems. Functional Cpf1 does not require trans-activating CRISPR RNA (tracrRNA), and therefore only CRISPR (crRNA) is required. This is beneficial for genome editing, as Cpf1 is not only smaller than Cas9, but also has a smaller sgRNA molecule (about half the number of nucleotides of Cas9). The Cpf1-crRNA complex cleaves the target DNA or RNA by identifying the protospacer adjacent motifs 5'-YTN-3' or 5'-TTN-3', in contrast to the G-rich PAM targeted by Cas9. After PAM recognition, Cpf1 introduces sticky-end-like DNA double-strand breaks with 4- or 5-nucleotide overhangs.
[0215] In some embodiments, the Cas9 is a Cas9 variant with specificity for modified PAM sequences. In some embodiments, additional Cas9 variants and PAM sequences are described in Miller, SM, et al. Continuous evolution of SpCas9 variants compatible with non-G PAMs, Nat. Biotechnol. (2020), the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 variant does not have a specific PAM requirement. In some embodiments, the Cas9 variant, e.g., the SpCas9 variant, has specificity for NRNH PAM, where R is A or G, and H is A, C, or T. In some embodiments, the SpCas9 variant has specificity for the PAM sequence AAA, TAA, CAA, GAA, TAT, GAT, or CAC. In some embodiments, the SpCas9 variant comprises an amino acid substitution at or corresponding to positions 1114, 1134, 1135, 1137, 1139, 1151, 1180, 1188, 1211, 1218, 1219, 1221, 1249, 1256, 1264, 1290, 1318, 1317, 1320, 1321, 1323, 1332, 1333, 1335, 1337, or 1339. In some embodiments, the SpCas9 variant comprises an amino acid substitution at or corresponding to positions 1114, 1135, 1218, 1219, 1221, 1249, 1320, 1321, 1323, 1332, 1333, 1335, or 1337. In some embodiments, the SpCas9 variant comprises an amino acid substitution at positions 1114, 1134, 1135, 1137, 1139, 1151, 1180, 1188, 1211, 1219, 1221, 1256, 1264, 1290, 1318, 1317, 1320, 1323, 1333, or their corresponding positions.In some embodiments, the SpCas9 variant comprises an amino acid substitution at positions 1114, 1131, 1135, 1150, 1156, 1180, 1191, 1218, 1219, 1221, 1227, 1249, 1253, 1286, 1293, 1320, 1321, 1332, 1335, 1339, or their corresponding positions. In some embodiments, the SpCas9 variant comprises an amino acid substitution at positions 1114, 1127, 1135, 1180, 1207, 1219, 1234, 1286, 1301, 1332, 1335, 1337, 1338, 1349, or their corresponding positions. Exemplary amino acid substitutions and PAM specificities of SpCas9 variants are shown in Tables 2A-2D.
[0216] [Table 2A]
[0217] [Table 2B]
[0218] [Table 2C]
[0219] [Table 2D]
[0220] Additional exemplary Cas9 (e.g., SaCas9) polypeptides with altered PAM recognition are described in Kleinstiver, et al., “Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition,” Nature Biotechnology, 33:1293-1298 (2015) DOI: 10.1038 / nbt.3404, the disclosure of which is incorporated herein by reference in its entirety for all purposes. In some embodiments, a Cas9 variant (e.g., a SaCas9 variant) comprising one or more of the modifications E782K, N929R, N968K and / or R1015H has or is associated with increased editing activity specificity relative to a reference polypeptide (e.g., SaCas9) at NNNRRT or NNHRRT PAM sequences, where N represents any nucleotide, H represents any nucleotide other than G (i.e., “non-G”), and R represents a purine. In embodiments, the Cas9 variant (e.g., a SaCas9 variant) includes the modifications E782K, N968K and R1015H, or the modifications E782K, K929R and R1015H.
[0221] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is the single effector of the microbial CRISPR-Cas system. The single effectors of the microbial CRISPR-Cas system include, but are not limited to, Cas9, Cpf1, Cas12b / C2c1, and Cas12c / C2c3. Generally, the microbial CRISPR-Cas system is divided into class 1 and class 2 systems. Class 1 systems have a multi-subunit effector complex, and class 2 systems have a single protein effector. For example, Cas9 and Cpf1 are class 2 effectors. In addition to Cas9 and Cpf1, three distinct class 2 CRISPR-Cas systems (Cas12b / C2c1 and Cas12c / C2c3) have been described by Shmakov et al., “Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems”, Mol. Cell, 2015 Nov. 5; 60(3):385-397, the entire contents of which are incorporated herein by reference. The effectors of two of the systems, Cas12b / C2c1 and Cas12c / C2c3, contain a RuvC-like endonuclease domain related to Cpf1. The third system contains an effector with two predicted HEPN RNase domains. The generation of mature CRISPR RNA is independent of tracrRNA, unlike the generation of CRISPR RNA by Cas12b / C2c1. Cas12b / C2c1 depends on both CRISPR RNA and tracrRNA for DNA cleavage.
[0222] In some embodiments, the napDNAbp is circularly permuted (eg, SEQ ID NO:326).
[0223] The crystal structure of Alicyclobacillus acidoterrastris Cas12b / C2c1 (AacC2c1) in complex with a chimeric single-molecule guide RNA (sgRNA) has been reported. See, e.g., Liu et al., “C2c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism”, Mol. Cell, 2017 Jan. 19; 65(2): 310-322, the entire contents of which are incorporated herein by reference. A crystal structure has also been reported of Alicyclobacillus acidoterrestris C2c1 bound to target DNA as a ternary complex. See, e.g., Yang et al., “PAM-dependent Target DNA Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease”, Cell, 2016 Dec.15;167(7):1814-1828, the entire contents of which are incorporated herein by reference. Both the target and non-target DNA strands have captured catalytically competent conformations of AacC2c1 that are independently positioned within a single RuvC catalytic pocket and allow Cas12b / C2c1-mediated cleavage to produce a staggered seven-nucleotide cleavage of the target DNA. Structural comparison of the Cas12b / C2c1 ternary complex with previously identified Cas9 and Cpf1 counterparts shows the diversity of mechanisms employed by the CRISPR-Cas9 system.
[0224] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) of any of the fusion proteins provided herein can be a Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a Cas12b / C2c1 protein. In some embodiments, the napDNAbp is a Cas12c / C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the napDNAbp sequences provided herein. It is understood that Cas12b / C2c1 or Cas12c / C2c3 from other bacterial species may also be used in accordance with the present disclosure.
[0225] In some embodiments, napDNAbp refers to Cas12c. In some embodiments, the Cas12c protein is Cas12c1 (SEQ ID NO: 327) or a variant of Cas12c1. In some embodiments, the Cas12 protein is Cas12c2 (SEQ ID NO: 328) or a variant of Cas12c2. In some embodiments, the Cas12 protein is a Cas12c protein from Oleiphilus species HI0009 (i.e., OspCas12c, SEQ ID NO: 329) or a variant of OspCas12c. These Cas12c molecules are described in Yan et al., "Functionally Diverse Type V CRISPR-Cas Systems," Science, 2019 Jan. 4; 363: 88-91, the entire contents of which are incorporated herein by reference. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring Cas12c1, Cas12c2, or OspCas12c protein. In some embodiments, the napDNAbp is a naturally occurring Cas12c1, Cas12c2, or OspCas12c protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any Cas12c1, Cas12c2, or OspCas12c protein described herein. It should be understood that Cas12c1, Cas12c2, or OspCas12c from other bacterial species may also be used in accordance with the present disclosure.
[0226] In some embodiments, napDNAbp refers to Cas12g, Cas12h, or Cas12i, e.g., as described in Yan et al., "Functionally Diverse Type V CRISPR-Cas Systems," Science, 2019 Jan. 4; 363: 88-91, the entire contents of each of which are incorporated herein by reference. Exemplary Cas12g, Cas12h, and Cas12i polypeptide sequences are set forth in the Sequence Listing as SEQ ID NOs: 330-333. By aggregating over 10 terabytes of sequence data, a new classification of type V Cas proteins, such as Cas12g, Cas12h, and Cas12i, that show weak similarity to previously characterized class V proteins, has been identified. In some embodiments, the Cas12 protein is Cas12g or a variant of Cas12g. In some embodiments, the Cas12 protein is Cas12h or a variant of Cas12h. In some embodiments, the Cas12 protein is Cas12i or a variant of Cas12i. It should be understood that other RNA-guided DNA binding proteins may be used as napDNAbp and are within the scope of the present disclosure. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a natural Cas12g, Cas12h, or Cas12i protein. In some embodiments, the napDNAbp is a natural Cas12g, Cas12h, or Cas12i protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any Cas12g, Cas12h, or Cas12i protein described herein.It should be understood that Cas12g, Cas12h, or Cas12i from other bacterial species may also be used according to the present disclosure. In some embodiments, Cas12i is Cas12i1 or Cas12i2.
[0227] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) of any of the fusion proteins provided herein can be a Cas12j / CasΦ protein. Cas12j / CasΦ is described in Pausch et al., "CRISPR-CasΦ from huge phages is a hypercompact genome editor," Science, 17 July 2020, Vol. 369, Issue 6501, pp. 333-337, which is incorporated herein by reference in its entirety. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a native Cas12j / CasΦ protein. In some embodiments, the napDNAbp is a native Cas12j / CasΦ protein. In some embodiments, the napDNAbp is a nuclease-inactive ("dead") Cas12j / CasΦ protein. It is understood that Cas12j / CasΦ from other species may also be used in accordance with the present disclosure.
[0228] Fusion proteins with internal insertions Provided herein is a fusion protein comprising a heterologous polypeptide fused to a nucleic acid programmable nucleic acid binding protein, such as a napDNAbp. As detailed below, the disclosure provides a polynucleotide encoding a fusion protein featuring a heterologous polypeptide, the polynucleotide comprising an intron within an open reading frame encoding all or a portion of the heterologous domain of the fusion protein. The heterologous polypeptide can be a polypeptide not found in a natural or wild-type napDNAbp polypeptide sequence. The heterologous polypeptide can be fused to the napDNAbp at the C-terminus of the napDNAbp, the N-terminus of the napDNAbp, or inserted at an internal position of the napDNAbp. In some embodiments, the heterologous polypeptide is a deaminase (e.g., cytidine adenosine deaminase) or a functional fragment thereof. For example, the fusion protein can include a deaminase adjacent to the N-terminal and C-terminal fragments of a Cas9 or Cas12 (e.g., Cas12b / C2c1) polypeptide. In some embodiments, the cytidine deaminase is an APOBEC deaminase (e.g., APOBEC1). In some embodiments, the adenosine deaminase is TadA (e.g., TadA*7.10 or TadA*8). In some embodiments, TadA is TadA*8 or TadA*9. The TadA sequences described herein (e.g., TadA7.10 or TadA*8) are suitable deaminases for the above fusion proteins.
[0229] In some embodiments, the fusion protein has the following structure: NH2-[N-terminal fragment of napDNAbp]-[deaminase]-[C-terminal fragment of napDNAbp]-COOH, NH2-[N-terminal fragment of Cas9]-[adenosine deaminase]-[C-terminal fragment of Cas9]-COOH, NH2-[N-terminal fragment of Cas12]-[adenosine deaminase]-[C-terminal fragment of Cas12]-COOH, NH2-[N-terminal fragment of Cas9]-[cytidine deaminase]-[C-terminal fragment of Cas9]-COOH, NH2-[N-terminal fragment of Cas12]-[cytidine deaminase]-[C-terminal fragment of Cas12]-COOH; where each instance of "]-[" is an optional linker.
[0230] The deaminase can be a circular permutant deaminase.For example, the deaminase can be a circular permutant adenosine deaminase.In some embodiments, the deaminase is a circular permutant TadA that is circularly permuted at amino acid residue 116, 136 or 65 numbered in the TadA reference sequence.
[0231] A fusion protein may contain multiple deaminases. A fusion protein may contain, for example, 1, 2, 3, 4, 5 or more deaminases. In some embodiments, a fusion protein contains one or two deaminases. The two or more deaminases of a fusion protein may be adenosine deaminase, cytidine deaminase, or a combination thereof. The two or more deaminases may be homodimers or heterodimers. The two or more deaminases may be inserted in tandem in napDNAbp. In some embodiments, the two or more deaminases may not be in tandem in napDNAbp.
[0232] In some embodiments, the napDNAbp in the fusion protein is a Cas9 polypeptide or a fragment thereof. The Cas9 polypeptide can be a variant Cas9 polypeptide. In some embodiments, the Cas9 polypeptide is a Cas9 nickase (nCas9) polypeptide or a fragment thereof. In some embodiments, the Cas9 polypeptide is a nuclease-inactive Cas9 (dCas9) polypeptide or a fragment thereof. The Cas9 polypeptide in the fusion protein can be a full-length Cas9 polypeptide. In some cases, the Cas9 polypeptide in the fusion protein may not be a full-length Cas9 polypeptide. The Cas9 polypeptide can be truncated, for example, at the N-terminus or C-terminus, compared to a native Cas9 protein. The Cas9 polypeptide can be a circularly permuted Cas9 protein. The Cas9 polypeptide can be a fragment, part, or domain of a Cas9 polypeptide, but it can still bind to a target polynucleotide and a guide nucleic acid sequence.
[0233] In some embodiments, the Cas9 polypeptide is Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), or a fragment or variant of any of the Cas9 polypeptides described herein.
[0234] In some embodiments, the fusion protein comprises an adenosine deaminase domain and a cytidine deaminase domain inserted into Cas9. In some embodiments, the adenosine deaminase is fused in Cas9 and the cytidine deaminase is fused to the C-terminus. In some embodiments, the adenosine deaminase is fused in Cas9 and the cytidine deaminase is fused to the N-terminus. In some embodiments, the cytidine deaminase is fused in Cas9 and the adenosine deaminase is fused to the C-terminus. In some embodiments, the cytidine deaminase is fused in Cas9 and the adenosine deaminase is fused to the N-terminus.
[0235] Exemplary structures of fusion proteins with adenosine deaminase and cytidine deaminase and Cas9 are provided below: NH2-[Cas9(adenosine deaminase)]-[cytidine deaminase]-COOH, NH2-[cytidine deaminase]-[Cas9(adenosine deaminase)]-COOH, NH2-[Cas9(cytidine deaminase)]-[adenosine deaminase]-COOH, or NH2-[adenosine deaminase]-[Cas9 (cytidine deaminase)]-COOH.
[0236] In some embodiments, the "-" used in the general architecture above indicates the presence of an optional linker.
[0237] In various embodiments, the catalytic domain has a DNA modifying activity (e.g., deaminase activity), such as adenosine deaminase activity. In some embodiments, the adenosine deaminase is TadA (e.g., TadA*7.10). In some embodiments, TadA is TadA*8. In some embodiments, TadA*8 is fused in Cas9 and a cytidine deaminase is fused to the C-terminus. In some embodiments, TadA*8 is fused in Cas9 and a cytidine deaminase is fused to the N-terminus. In some embodiments, a cytidine deaminase is fused in Cas9 and TadA*8 is fused to the C-terminus. In some embodiments, a cytidine deaminase is fused in Cas9 and TadA*8 is fused to the N-terminus. Exemplary structures of fusion proteins having TadA*8 and cytidine deaminase and Cas9 are provided below: NH2-[Cas9(TadA*8)]-[cytidine deaminase]-COOH, NH2-[cytidine deaminase]-[Cas9(TadA*8)]-COOH, NH2-[Cas9(cytidine deaminase)]-[TadA*8]-COOH or NH2-[TadA*8]-[Cas9(cytidine deaminase)]-COOH.
[0238] In some embodiments, the "-" used in the general architecture above indicates the presence of an optional linker.
[0239] A heterologous polypeptide (e.g., a deaminase) can be inserted into the napDNAbp (e.g., Cas9 or Cas12 (e.g., Cas12b / C2c1)) at a suitable position, for example, such that the napDNAbp retains the ability to bind to a target polynucleotide and a guide nucleic acid. A deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted into the napDNAbp without impairing the function of the deaminase (e.g., base editing activity) or the napDNAbp (e.g., the ability to bind to a target nucleic acid and a guide nucleic acid). A deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted into the napDNAbp, for example, in a disordered region, or in a region that contains a high temperature factor or B factor, as shown by crystallographic studies. Ordered, disordered, or unstructured regions of proteins (e.g., solvent-exposed regions and loops) can be used for insertion without compromising structure or function. Deaminases (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted into napDNAbp in flexible loop regions or solvent-exposed regions. In some embodiments, deaminases (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) are inserted into flexible loops of Cas9 or Cas12b / C2c1 polypeptides.
[0240] In some embodiments, the insertion location of the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is determined by B-factor analysis of the crystal structure of the Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted into a region of the Cas9 polypeptide that contains a higher than average B-factor (e.g., a higher B-factor compared to the whole protein or a protein domain that contains a disordered region). The B-factor or temperature factor may indicate the variation of atoms from their average position (e.g., as a result of temperature-dependent atomic vibrations or static disorder in the crystal lattice). A high B-factor (e.g., a higher than average B-factor) of backbone atoms may indicate a region with relatively high local mobility. Such a region can be used to insert the deaminase without compromising the structure or function. A deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted at a position that includes a residue having a Cα atom with a B factor that is 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, or greater than the average B factor of the entire protein. A deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted at a position that contains a residue that has a B factor that is 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, or greater than the average B factor of the Cas9 protein domain that contains the residue. Positions in the Cas9 polypeptide that contain higher than average B factors can include, for example, residues 768, 792, 1052, 1015, 1022, 1026, 1029, 1067, 1040, 1054, 1068, 1246, 1247, and 1248, as numbered in the Cas9 reference sequence above.Regions of a Cas9 polypeptide that contain a higher than average B factor can include, for example, residues 792-872, 792-906, and 2-791, as numbered in the Cas9 reference sequence above.
[0241] A heterologous polypeptide (e.g., a deaminase) may be inserted into the napDNAbp at an amino acid residue selected from the group consisting of: residues 768, 791, 792, 1015, 1016, 1022, 1023, 1026, 1029, 1040, 1052, 1054, 1067, 1068, 1069, 1246, 1247, and 1248, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues of another Cas9 polypeptide. In some embodiments, the heterologous polypeptide is inserted between amino acid positions 768-769, 791-792, 792-793, 1015-1016, 1022-1023, 1026-1027, 1029-1030, 1040-1041, 1052-1053, 1054-1055, 1067-1068, 1068-1069, 1247-1248, or 1248-1249, or their corresponding amino acid positions, as numbered in the above Cas9 reference sequences. In some embodiments, the heterologous polypeptide is inserted between amino acid positions 769-770, 792-793, 793-794, 1016-1017, 1023-1024, 1027-1028, 1030-1031, 1041-1042, 1053-1054, 1055-1056, 1068-1069, 1069-1070, 1248-1249, or 1249-1250, or their corresponding amino acid positions, as numbered in the above Cas9 reference sequences. In some embodiments, the heterologous polypeptide replaces an amino acid residue selected from the group consisting of residues 768, 791, 792, 1015, 1016, 1022, 1023, 1026, 1029, 1040, 1052, 1054, 1067, 1068, 1069, 1246, 1247, and 1248 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. It should be understood that reference to the above Cas9 reference sequence with respect to insertion positions is for exemplary purposes.Insertions discussed herein are not limited to the Cas9 polypeptide sequences of the above Cas9 reference sequences, but include insertions at corresponding locations in variant Cas9 polypeptides (e.g., Cas9 nickase (nCas9), nuclease-inactive Cas9 (dCas9), Cas9 variants lacking a nuclease domain, truncated Cas9, or a Cas9 domain partially or completely lacking the HNH domain).
[0242] A heterologous polypeptide (e.g., a deaminase) may be inserted into the napDNAbp at an amino acid residue selected from the group consisting of: amino acid residues 768, 792, 1022, 1026, 1040, 1068, and 1247 as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues of another Cas9 polypeptide. In some embodiments, the heterologous polypeptide is inserted between amino acid positions 768-769, 792-793, 1022-1023, 1026-1027, 1029-1030, 1040-1041, 1068-1069, or 1247-1248 as numbered in the above Cas9 reference sequence, or their corresponding amino acid positions. In some embodiments, the heterologous polypeptide is inserted between amino acid positions 769-770, 793-794, 1023-1024, 1027-1028, 1030-1031, 1041-1042, 1069-1070, or 1248-1249 numbered in the above Cas9 reference sequence, or their corresponding amino acid positions. In some embodiments, the heterologous polypeptide replaces an amino acid residue selected from the group consisting of amino acid residues 768, 792, 1022, 1026, 1040, 1068, and 1247 numbered in the above Cas9 reference sequence, or the corresponding amino acid residues of another Cas9 polypeptide.
[0243] A heterologous polypeptide (e.g., a deaminase) may be inserted into the napDNAbp at an amino acid residue described herein or a corresponding amino acid residue of another Cas9 polypeptide. In one embodiment, a heterologous polypeptide (e.g., a deaminase) may be inserted into the napDNAbp at an amino acid residue selected from the group consisting of: amino acid residues 1002, 1003, 1025, 1052-1056, 1242-1247, 1061-1077, 943-947, 686-691, 569-578, 530-539, and 1060-1077 as numbered in the Cas9 reference sequence above, or a corresponding amino acid residue of another Cas9 polypeptide. A deaminase (e.g., an adenosine deaminase, a cytidine deaminase, or an adenosine deaminase and a cytidine deaminase) may be inserted at the N-terminus or C-terminus of the residue or may replace the residue. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at the C-terminus of the residue.
[0244] In some embodiments, an adenosine deaminase (e.g., TadA) may be inserted at an amino acid residue selected from the group consisting of: amino acid residues 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246 numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, an adenosine deaminase (e.g., TadA) is inserted in place of residues 792-872, 792-906, or 2-791 numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the adenosine deaminase is inserted at the N-terminus of an amino acid selected from the group consisting of amino acid residues 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246 numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the adenosine deaminase is inserted at the C-terminus of an amino acid selected from the group consisting of amino acid residues 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246 numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the adenosine deaminase is inserted to replace an amino acid selected from the group consisting of amino acid residues 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246 numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide.
[0245] In some embodiments, the cytidine deaminase (e.g., APOBEC1) is inserted at an amino acid residue selected from the group consisting of: amino acid residues 1016, 1023, 1029, 1040, 1069, and 1247 numbered in the above Cas9 reference sequence, or the corresponding amino acid residues of another Cas9 polypeptide. In some embodiments, the cytidine deaminase is inserted at the N-terminus of an amino acid selected from the group consisting of: amino acid residues 1016, 1023, 1029, 1040, 1069, and 1247 numbered in the above Cas9 reference sequence, or the corresponding amino acid residues of another Cas9 polypeptide. In some embodiments, the cytidine deaminase is inserted at the C-terminus of an amino acid selected from the group consisting of: amino acid residues 1016, 1023, 1029, 1040, 1069, and 1247 numbered in the above Cas9 reference sequence, or the corresponding amino acid residues of another Cas9 polypeptide. In some embodiments, the cytidine deaminase is inserted to replace an amino acid selected from the group consisting of amino acid residues 1016, 1023, 1029, 1040, 1069, and 1247, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide.
[0246] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 768 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminally to amino acid residue 768 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminally to amino acid residue 768 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid residue 768, as numbered in the Cas9 reference sequence above, or the corresponding amino acid residue in another Cas9 polypeptide.
[0247] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 791 or amino acid residue 792 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminal to amino acid residue 791 or amino acid residue 792 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminally to amino acid 791 or N-terminally to amino acid 792 numbered in the above Cas9 reference sequence, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid 791 or amino acid 792 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.
[0248] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1016 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminally to amino acid residue 1016 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminally to amino acid residue 1016 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid residue 1016, numbered in the Cas9 reference sequence above, or the corresponding amino acid residue in another Cas9 polypeptide.
[0249] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1022 or amino acid residue 1023 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminally to amino acid residue 1022 or amino acid residue 1023 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminal to amino acid residue 1022 or amino acid residue 1023 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid residue 1022 or amino acid residue 1023 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.
[0250] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1026 or amino acid residue 1029, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminal to amino acid residue 1026 or amino acid residue 1029, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminal to amino acid residue 1026 or amino acid residue 1029 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid residue 1026 or amino acid residue 1029 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.
[0251] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1040 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminally to amino acid residue 1040 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminally to amino acid residue 1040 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid residue 1040, numbered in the Cas9 reference sequence above, or the corresponding amino acid residue in another Cas9 polypeptide.
[0252] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1052 or amino acid residue 1054, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminally to amino acid residue 1052 or amino acid residue 1054, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminal to amino acid residue 1052 or amino acid residue 1054 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid residue 1052 or amino acid residue 1054 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.
[0253] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1067 or amino acid residue 1068 or amino acid residue 1069 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminal to amino acid residue 1067 or amino acid residue 1068 or amino acid residue 1069 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminal to amino acid residue 1067 or amino acid residue 1068 or amino acid residue 1069 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue of another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid residue 1067 or amino acid residue 1068 or amino acid residue 1069 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue of another Cas9 polypeptide.
[0254] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1246 or amino acid residue 1247 or amino acid residue 1248 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminal to amino acid residue 1246 or amino acid residue 1247 or amino acid residue 1248 numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminal to amino acid residue 1246, or amino acid residue 1247, or amino acid residue 1248, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted to replace amino acid residue 1246, or amino acid residue 1247, or amino acid residue 1248, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.
[0255] In some embodiments, the heterologous polypeptide (e.g., a deaminase) is inserted into a mobile loop of a Cas9 polypeptide. The mobile loop portion can be selected from the group consisting of amino acid residues 530-537, 569-570, 686-691, 943-947, 1002-1025, 1052-1077, 1232-1247, or 1298-1300 as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues of another Cas9 polypeptide. The mobile loop portion can be selected from the group consisting of amino acid residues 1-529, 538-568, 580-685, 692-942, 948-1001, 1026-1051, 1078-1231, or 1248-1297 as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues of another Cas9 polypeptide.
[0256] A heterologous polypeptide (e.g., adenine deaminase) can be inserted into a region of a Cas9 polypeptide corresponding to the following amino acid residues: amino acid residues 1017-1069, 1242-1247, 1052-1056, 1060-1077, 1002-1003, 943-947, 530-537, 568-579, 686-691, 1242-1247, 1298-1300, 1066-1077, 1052-1056, or 1060-1077 numbered in the Cas9 reference sequence above, or the corresponding amino acid residues of another Cas9 polypeptide.
[0257] A heterologous polypeptide (e.g., adenine deaminase) may be inserted in place of the deleted region of the Cas9 polypeptide. The deleted region may correspond to the N-terminal or C-terminal portion of the Cas9 polypeptide. In some embodiments, the deleted region corresponds to residues 792-872 as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues of another Cas9 polypeptide. In some embodiments, the deleted region corresponds to residues 792-906 as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues of another Cas9 polypeptide. In some embodiments, the deleted region corresponds to residues 2-791 as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues of another Cas9 polypeptide. In some embodiments, the deleted region corresponds to residues 1017-1069 as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues thereof.
[0258] Exemplary internal fusion base editors are provided in Table 3 below. [Table 3]
[0259] A heterologous polypeptide (e.g., a deaminase) can be inserted within a structural or functional domain of a Cas9 polypeptide. A heterologous polypeptide (e.g., a deaminase) can be inserted between two structural or functional domains of a Cas9 polypeptide. A heterologous polypeptide (e.g., a deaminase) can be inserted in place of a structural or functional domain of a Cas9 polypeptide, for example, after deleting the domain from the Cas9 polypeptide. A structural or functional domain of a Cas9 polypeptide can include, for example, RuvCI, RuvCII, RuvCIII, Rec1, Rec2, PI, or HNH.
[0260] In some embodiments, the Cas9 polypeptide lacks one or more domains selected from the group consisting of: RuvCI, RuvCII, RuvCIII, Rec1, Rec2, PI, or HNH domain. In some embodiments, the Cas9 polypeptide lacks a nuclease domain. In some embodiments, the Cas9 polypeptide lacks an HNH domain. In some embodiments, the Cas9 polypeptide lacks a portion of the HNH domain, such that the Cas9 polypeptide has reduced or eliminated HNH activity. In some embodiments, the Cas9 polypeptide comprises a deletion of the nuclease domain, and a deaminase is inserted to replace the nuclease domain. In some embodiments, the HNH domain is deleted and a deaminase is inserted in its place. In some embodiments, one or more of the RuvC domains are deleted and a deaminase is inserted in its place.
[0261] A fusion protein comprising a heterologous polypeptide may be flanked by N- and C-terminal fragments of napDNAbp. In some embodiments, the fusion protein comprises a deaminase flanked by N- and C-terminal fragments of a Cas9 polypeptide. The N- or C-terminal fragment can bind to a target polynucleotide sequence. The C-terminus of the N- or C-terminal fragment can comprise a portion of a flexible loop of a Cas9 polypeptide. The C-terminus of the N- or C-terminal fragment can comprise a portion of an alpha-helical structure of a Cas9 polypeptide. The N- or C-terminal fragment can comprise a DNA-binding domain. The N- or C-terminal fragment can comprise a RuvC domain. The N- or C-terminal fragment can comprise an HNH domain. In some embodiments, either the N- or C-terminal fragment does not comprise an HNH domain.
[0262] In some embodiments, the C-terminus of the N-terminal Cas9 fragment comprises an amino acid that is adjacent to the target nucleobase when the fusion protein deaminates the target nucleobase. In some embodiments, the N-terminus of the C-terminal Cas9 fragment comprises an amino acid that is adjacent to the target nucleobase when the fusion protein deaminates the target nucleobase. The insertion location of the different deaminase can be different to allow for proximity between the target nucleobase and the amino acid at the C-terminus of the N-terminal Cas9 fragment or the N-terminus of the C-terminal Cas9 fragment. For example, the insertion location of the deaminase can be at an amino acid residue selected from the group consisting of: amino acid residues 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246 numbered in the above Cas9 reference sequence, or the corresponding amino acid residues of another Cas9 polypeptide.
[0263] The N-terminal Cas9 fragment of the fusion protein (i.e., the N-terminal Cas9 fragment adjacent to the deaminase of the fusion protein) can comprise the N-terminus of a Cas9 polypeptide. The N-terminal Cas9 fragment of the fusion protein can comprise a length of at least about: 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, or 1300 amino acids. The N-terminal Cas9 fragment of the fusion protein can comprise a sequence corresponding to the following amino acid residues: amino acid residues 1-56, 1-95, 1-200, 1-300, 1-400, 1-500, 1-600, 1-700, 1-718, 1-765, 1-780, 1-906, 1-918, or 1-1100 numbered in the Cas9 reference sequence above, or the corresponding amino acid residues of another Cas9 polypeptide. An N-terminal Cas9 fragment can comprise amino acid residues 1-56, 1-95, 1-200, 1-300, 1-400, 1-500, 1-600, 1-700, 1-718, 1-765, 1-780, 1-906, 1-918, or 1-1100 as numbered in the above Cas9 reference sequences, or a sequence that comprises at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% sequence identity to the corresponding amino acid residues of another Cas9 polypeptide.
[0264] The C-terminal Cas9 fragment of the fusion protein (i.e., the C-terminal Cas9 fragment adjacent to the deaminase of the fusion protein) can comprise the C-terminus of the Cas9 polypeptide. The C-terminal Cas9 fragment of the fusion protein can comprise a length of at least about: 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, or 1300 amino acids. The C-terminal Cas9 fragment of the fusion protein can comprise a sequence corresponding to the following amino acid residues: amino acid residues 1099-1368, 918-1368, 906-1368, 780-1368, 765-1368, 718-1368, 94-1368, or 56-1368 numbered in the Cas9 reference sequence above, or the corresponding amino acid residues in another Cas9 polypeptide. An N-terminal Cas9 fragment can comprise amino acid residues 1099-1368, 918-1368, 906-1368, 780-1368, 765-1368, 718-1368, 94-1368, or 56-1368 as numbered in the above Cas9 reference sequences, or a sequence that comprises at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to the corresponding amino acid residues of another Cas9 polypeptide.
[0265] The N-terminal Cas9 fragment and the C-terminal Cas9 fragment of the fusion protein taken together may not correspond to a native full-length Cas9 polypeptide sequence, e.g., as set forth in the Cas9 reference sequence above.
[0266] The fusion proteins described herein can provide targeted deamination and reduce deamination at non-target sites (e.g., off-target sites) (e.g., reduce genome-wide spurious deamination). The fusion proteins described herein can provide targeted deamination and reduce bystander deamination at non-target sites. Unwanted or off-target deamination can be reduced by at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%, for example, compared to an end terminus fusion protein comprising a deaminase fused to the N-terminus or C-terminus of a Cas9 polypeptide. Unwanted or off-target deamination can be reduced, for example, by at least 1-fold, at least 2-fold, at least 3-fold, at least 4-fold, at least 5-fold, at least 10-fold, at least 15-fold, at least 20-fold, at least 30-fold, at least 40-fold, at least 50-fold, at least 60-fold, at least 70-fold, at least 80-fold, at least 90-fold, or at least 100-fold, as compared to a terminal fusion protein comprising a deaminase fused to the N-terminus or C-terminus of a Cas9 polypeptide.
[0267] In some embodiments, the deaminase of the fusion protein (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) deaminates no more than two nucleobases within the R-loop. In some embodiments, the deaminase of the fusion protein deaminates no more than three nucleobases within the R-loop. In some embodiments, the deaminase of the fusion protein deaminates no more than 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleobases within the R-loop. An R-loop is a triple-stranded nucleic acid structure, including a DNA:RNA hybrid, a DNA:DNA, or an RNA:RNA complementary structure, associated with a single-stranded DNA. As used herein, an R-loop can be formed when a target polynucleotide contacts a CRISPR complex or a base editing complex, and a portion of the guide polynucleotide (e.g., a guide RNA) hybridizes with and replaces a portion of the target polynucleotide (e.g., a target DNA). In some embodiments, the R-loop comprises a spacer sequence and a hybridized region of the target DNA complementary sequence. The R-loop region can be about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleobase pairs in length. In some embodiments, the R-loop region is about 20 nucleobase pairs in length. It should be understood that as used herein, the R-loop region is not limited to the target DNA strand that hybridizes with the guide polynucleotide. For example, editing of the target nucleobase in the R-loop region may be on the DNA strand that contains the complementary strand to the guide RNA, or on the DNA strand that is the opposite strand to the complementary strand to the guide RNA. In some embodiments, editing in the region of the R-loop comprises editing the nucleobase on the non-complementary strand to the guide RNA (protospacer strand) in the target DNA sequence.
[0268] The fusion proteins described herein can provide targeted deamination in an editing window distinct from canonical base editing. In some embodiments, the targeted nucleobase is about 1 to about 20 bases upstream of the PAM sequence in the target polynucleotide sequence. In some embodiments, the targeted nucleobase is about 2 to about 12 bases upstream of the PAM sequence in the target polynucleotide sequence. In some embodiments, the targeted nucleobase is about 1 to 9 base pairs, about 2 to 10 base pairs, about 3 to 11 base pairs, about 4 to 12 base pairs, about 5 to 13 base pairs, about 6 to 14 base pairs, about 7 to 15 base pairs, about 8 to 16 base pairs, about 9 to 17 base pairs, about 10 to 18 base pairs, about 11 to 19 base pairs, about 12 to 20 base pairs, about 1 to 7 base pairs, about 2 to 8 base pairs, about 3 to 9 base pairs, about 4 to 5 base pairs, about 5 to 6 ... base pairs, about 4-10 base pairs, about 5-11 base pairs, about 6-12 base pairs, about 7-13 base pairs, about 8-14 base pairs, about 9-15 base pairs, about 10-16 base pairs, about 11-17 base pairs, about 12-18 base pairs, about 13-19 base pairs, about 14-20 base pairs, about 1-5 base pairs, about 2-6 base pairs, about 3-7 base pairs, about 4-8 base pairs, about 5-9 base pairs, about 6-10 base pairs, about 7-11 base pairs, about 8-12 base pairs, about 9-13 base pairs, about 10-14 base pairs, about 11-15 base pairs, about 12-16 base pairs, about 13-17 base pairs, about 14-18 base pairs, about 15-19 base pairs, about 16-20 base pairs, about 1-3 base pairs, about 2-4 base pairs, about 3-5 base pairs, about 4-6 base pairs, about 5-7 base pairs, about 6-8 base pairs, about 7-9 base pairs, about 8-10 base pairs, about 9-11 base pairs, about 10-12 base pairs, about 11-13 base pairs, about 12-14 base pairs, about 13-15 base pairs, about 14-16 base pairs, about 15-17 base pairs, about 16-18 base pairs, about 17-19 base pairs, about 18-20 base pairs away from the PAM sequence or upstream of the PAM sequence. In some embodiments, the target nucleobase is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more base pairs away from the PAM sequence or upstream of the PAM sequence. In some embodiments, the target nucleobase is about 1, 2, 3, 4, 5, 6, 7, 8, or 9 base pairs upstream of the PAM sequence. In some embodiments, the targeted nucleobase is about 2, 3, 4, or 6 base pairs upstream of the PAM sequence.
[0269] A fusion protein can contain multiple heterologous polypeptides. For example, a fusion protein can further contain one or more UGI domains and / or one or more nuclear localization signals. Two or more heterologous domains can be inserted in tandem. Two or more heterologous domains can be inserted in a position such that they are not in tandem in the NapDNAbp.
[0270] The fusion protein may include a linker between the deaminase and the napDNAbp polypeptide. The linker may be a peptide or a non-peptide linker. For example, the linker may be XTEN, (GGGS)n (SEQ ID NO: 334), (GGGGS)n (SEQ ID NO: 335), (G)n, (EAAAK)n (SEQ ID NO: 336), (GGS)n, SGSETPGTSESATPES (SEQ ID NO: 337). In some embodiments, the fusion protein includes a linker between the N-terminal Cas9 fragment and the deaminase. In some embodiments, the fusion protein includes a linker between the C-terminal Cas9 fragment and the deaminase. In some embodiments, the N-terminal fragment and the C-terminal fragment of the napDNAbp are connected to the deaminase with a linker. In some embodiments, the N-terminal fragment and the C-terminal fragment are linked to the deaminase domain without a linker. In some embodiments, the fusion protein includes a linker between the N-terminal Cas9 fragment and the deaminase, but does not include a linker between the C-terminal Cas9 fragment and the deaminase. In some embodiments, the fusion protein comprises a linker between the C-terminal Cas9 fragment and the deaminase, but does not comprise a linker between the N-terminal Cas9 fragment and the deaminase.
[0271] In some embodiments, the napDNAbp in the fusion protein is a Cas12 polypeptide (e.g., Cas12b / C2c1) or a fragment thereof. The Cas12 polypeptide can be a variant Cas12 polypeptide. In other embodiments, the N-terminal or C-terminal fragment of the Cas12 polypeptide comprises a nucleic acid programmable DNA binding domain or a RuvC domain. In other embodiments, the fusion protein comprises a linker between the Cas12 polypeptide and the catalytic domain. In other embodiments, the amino acid sequence of the linker is GGSGGS (SEQ ID NO: 338) or GSSGSETPGTSESATPESSG (SEQ ID NO: 339). In other embodiments, the linker is a rigid linker. In other embodiments of the above aspects, the linker is encoded by GGAGGCTCTGCAGGAGGAAGC (SEQ ID NO: 340) or GGCTCTTGCTGAAACACCTGGCACAAGCGAGAGCGCCACCCCTGAGAGCTCTGGC (SEQ ID NO: 341).
[0272] Fusion proteins comprising heterologous catalytic domains flanking the N- and C-terminal fragments of Cas12 polypeptide are also useful for base editing in the methods described herein. Fusion proteins comprising Cas12 and one or more deaminase domains (e.g., adenosine deaminase, or adenosine deaminase domain flanking Cas12 sequences) are also useful for highly specific and efficient base editing of target sequences. In one embodiment, a chimeric Cas12 fusion protein comprises a heterologous catalytic domain (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) inserted within a Cas12 polypeptide. In some embodiments, the fusion protein comprises an adenosine deaminase domain and a cytidine deaminase domain inserted within Cas12. In some embodiments, the adenosine deaminase is fused within Cas12 and the cytidine deaminase is fused to the C-terminus. In some embodiments, adenosine deaminase is fused within Cas12 and cytidine deaminase is fused to the N-terminus. In some embodiments, cytidine deaminase is fused within Cas12 and adenosine deaminase is fused to the C-terminus. In some embodiments, cytidine deaminase is fused within Cas12 and adenosine deaminase is fused to the N-terminus. Exemplary structures of fusion proteins having adenosine deaminase and cytidine deaminase and Cas12 are provided below: NH2-[Cas12(adenosine deaminase)]-[cytidine deaminase]-COOH, NH2-[cytidine deaminase]-[Cas12 (adenosine deaminase)]-COOH, NH2-[Cas12(cytidine deaminase)]-[adenosine deaminase]-COOH, or NH2-[adenosine deaminase]-[Cas12 (cytidine deaminase)]-COOH, In some embodiments, the "-" used in the general architecture above indicates the presence of an optional linker.
[0273] In various embodiments, the catalytic domain has a DNA modifying activity (e.g., deaminase activity), such as adenosine deaminase activity. In some embodiments, the adenosine deaminase is TadA (e.g., TadA*7.10). In some embodiments, TadA is TadA*8. In some embodiments, TadA*8 is fused in Cas12 and a cytidine deaminase is fused to the C-terminus. In some embodiments, TadA*8 is fused in Cas12 and a cytidine deaminase is fused to the N-terminus. In some embodiments, a cytidine deaminase is fused in Cas12 and TadA*8 is fused to the C-terminus. In some embodiments, a cytidine deaminase is fused in Cas12 and TadA*8 is fused to the N-terminus. Exemplary structures of fusion proteins having TadA*8 and cytidine deaminase and Cas12 are provided below: N-[Cas12(TadA*8)]-[cytidine deaminase]-C, N-[cytidine deaminase]-[Cas12(TadA*8)]-C, N-[Cas12(cytidine deaminase)]-[TadA*8]-C, or N-[TadA*8]-[Cas12(cytidine deaminase)]-C. In some embodiments, the "-" used in the general architecture above indicates the presence of an optional linker.
[0274] In other embodiments, the fusion protein comprises one or more catalytic domains. In other embodiments, at least one of the one or more catalytic domains is inserted into a Cas12 polypeptide or fused at the N-terminus or C-terminus of Cas12. In other embodiments, at least one of the one or more catalytic domains is inserted into a loop, an alpha-helical region, an unstructured portion, or a solvent-accessible portion of a Cas12 polypeptide. In other embodiments, the Cas12 polypeptide is Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ. In other embodiments, the Cas12 polypeptide has at least about 85% amino acid sequence identity to Bacillus hisashii Cas12b, Bacillus thermoamylovorans Cas12b, Bacillus species V3-13 Cas12b, or Alicyclobacillus acidiphilus Cas12b (SEQ ID NO: 342). In other embodiments, the Cas12 polypeptide has at least about 90% amino acid sequence identity to Bacillus hisashii Cas12b (SEQ ID NO: 343), Bacillus thermoamylovorans Cas12b, Bacillus species V3-13 Cas12b, or Alicyclobacillus acidiphilus Cas12b. In other embodiments, the Cas12 polypeptide has at least about 95% amino acid sequence identity to Bacillus hisashii Cas12b, Bacillus thermoamylovorans Cas12b (SEQ ID NO: 344), Bacillus species V3-13 Cas12b, or Alicyclobacillus acidiphilus Cas12b (SEQ ID NO: 345).In other embodiments, the Cas12 polypeptide comprises or consists essentially of a fragment of Bacillus hisashii Cas12b, Bacillus thermoamylovorans Cas12b, Bacillus species V3-13 Cas12b, or Alicyclobacillus acidiphilus Cas12b. In embodiments, the Cas12 polypeptide comprises BvCas12b (V4), which in some embodiments is represented as 5'mRNACap---5'UTR---bhCas12b---termination sequence---3'UTR---120 polyA tail (SEQ ID NOs:346-348).
[0275] In other embodiments, the catalytic domain is inserted between amino acid positions 153-154, 255-256, 306-307, 980-981, 1019-1020, 534-535, 604-605, or 344-345 of BhCas12b, or between the corresponding amino acid residues of Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ. In other embodiments, the catalytic domain is inserted between amino acids P153 and S154 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids K255 and E256 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids D980 and G981 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids K1019 and L1020 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids F534 and P535 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids K604 and G605 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids H344 and F345 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acid positions 147 and 148, 248 and 249, 299 and 300, 991 and 992, or 1031 and 1032 of BvCas12b, or between the corresponding amino acid residues of Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ. In other embodiments, the catalytic domain is inserted between amino acids P147 and D148 of BvCas12b. In other embodiments, the catalytic domain is inserted between amino acids G248 and G249 of BvCas12b. In other embodiments, the catalytic domain is inserted between amino acids P299 and E300 of BvCas12b. In other embodiments, the catalytic domain is inserted between amino acids G991 and E992 of BvCas12b. In other embodiments, the catalytic domain is inserted between amino acids K1031 and M1032 of BvCas12b.In other embodiments, the catalytic domain is inserted between amino acid positions 157 and 158, 258 and 259, 310 and 311, 1008 and 1009, or 1044 and 1045 of AaCas12b, or between the corresponding amino acid residues of Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ. In other embodiments, the catalytic domain is inserted between amino acids P157 and G158 of AaCas12b. In other embodiments, the catalytic domain is inserted between amino acids V258 and G259 of AaCas12b. In other embodiments, the catalytic domain is inserted between amino acids D310 and P311 of AaCas12b. In other embodiments, the catalytic domain is inserted between amino acids G1008 and E1009 of AaCas12b. In other embodiments, the catalytic domain is inserted between amino acids G1044 and K1045 of AaCas12b.
[0276] In other embodiments, the fusion protein comprises a nuclear localization signal (e.g., a bipartite nuclear localization signal). In other embodiments, the amino acid sequence of the nuclear localization signal is MAPKKKRKVGIHGVPAA (SEQ ID NO: 349). In other embodiments of the above aspects, the nuclear localization signal is encoded by the following sequence: ATGGCCCCAAAGAAGAAGCGGAAGGTCGGTATCCACGGAGTCCCAGCAGCC (SEQ ID NO: 350). In other embodiments, the Cas12b polypeptide comprises a mutation that silences the catalytic activity of the RuvC domain. In other embodiments, the Cas12b polypeptide comprises a D574A, D829A and / or D952A mutation. In other embodiments, the fusion protein further comprises a tag (e.g., an influenza hemagglutinin tag).
[0277] In some embodiments, the fusion protein comprises a napDNAbp domain (e.g., a Cas12-derived domain) and has an internally fused nucleobase editing domain (e.g., all or a portion of a deaminase domain (e.g., an adenosine deaminase domain)). In some embodiments, the napDNAbp is Cas12b. In some embodiments, the base editor comprises a BhCas12b domain and has an internally fused TadA*8 domain inserted at a locus provided in Table 4 below. [Table 4]
[0278] As a non-limiting example, an adenosine deaminase (e.g., TadA*8.13) can be inserted into BhCas12b to generate a fusion protein (e.g., TadA*8.13-BhCas12b) that effectively edits a nucleic acid sequence.
[0279] In some embodiments, the base editor system described herein is an ABE with TadA inserted into Cas9. The polypeptide sequences of relevant ABEs with TadA inserted into Cas9 are provided as SEQ ID NOs: 351-396 in the attached sequence listing.
[0280] In some embodiments, an adenosine base editor was generated to insert TadA, or a variant thereof, into the identified position of the Cas9 polypeptide.
[0281] Exemplary, but non-limiting, fusion proteins are described in International PCT Application No. PCT / US2020 / 016285 and U.S. Provisional Application Nos. 62 / 852,228 and 62 / 852,224, the contents of which are incorporated by reference in their entireties.
[0282] Editing A to G In some embodiments, the base editors described herein comprise an adenosine deaminase domain. Such an adenosine deaminase domain of a base editor can facilitate the editing of an adenine (A) nucleobase to a guanine (G) nucleobase by deaminating A to form inosine (I), which exhibits the base pairing properties of G. An adenosine deaminase can deaminate (i.e., remove an amine group) the adenine of a deoxyadenosine residue of a deoxyribonucleic acid (DNA). In some embodiments, an A-to-G base editor further comprises an inhibitor of inosine base excision repair, such as a uracil glycosylase inhibitor (UGI) domain or a catalytically inactive inosine-specific nuclease. Without wishing to be bound by any particular theory, the UGI domain or catalytically inactive inosine-specific nuclease can inhibit or prevent base excision repair of deaminated adenosine residues (e.g., inosine), which can improve the activity or efficiency of the base editor.
[0283] Base editors comprising adenosine deaminase can act on any polynucleotide, including DNA, RNA, and DNA-RNA hybrids. In certain embodiments, base editors comprising adenosine deaminase can deaminate target A of polynucleotides comprising RNA. For example, base editors can comprise an adenosine deaminase domain capable of deaminating target A of RNA polynucleotides and / or DNA-RNA hybrid polynucleotides. In one embodiment, the adenosine deaminase incorporated in the base editor comprises all or a portion of an adenosine deaminase acting on RNA (ADAR, e.g., ADAR1 or ADAR2) or tRNA (ADAT). Base editors comprising an adenosine deaminase domain can also deaminate A nucleobases of DNA polynucleotides. In one embodiment, the adenosine deaminase domain of the base editor comprises all or a portion of ADAT comprising one or more mutations that enable ADAT to deaminate target A in DNA. For example, a base editor can include all or a portion of ADAT from Escherichia coli (EcTadA) that contains one or more of the following mutations: D108N, A106V, D147Y, E155V, L84F, H123Y, I156F, or corresponding mutations in another adenosine deaminase. Exemplary ADAT homolog polypeptide sequences are provided in the Sequence Listing as SEQ ID NOs: 1, 397-403.
[0284] The adenosine deaminase may be from any suitable organism (e.g., E. coli). In some embodiments, the adenosine deaminase is from a prokaryote. In some embodiments, the adenosine deaminase is from a bacterium. In some embodiments, the adenosine deaminase is from Escherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis. In some embodiments, the adenosine deaminase is from E. coli. In some embodiments, the adenosine deaminase is a naturally occurring adenosine deaminase that includes one or more mutations corresponding to any of the mutations provided herein (e.g., mutations in ecTadA). Corresponding residues in any homologous protein can be identified, for example, by sequence alignment and determining homologous residues. Mutations in any naturally occurring adenosine deaminase (e.g., having homology to ecTadA) that correspond to any of the mutations described herein (e.g., any of the mutations identified in ecTadA) can be generated accordingly.
[0285] In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the amino acid sequences described in any of the adenosine deaminases provided herein. It should be understood that the adenosine deaminases provided herein can comprise one or more mutations (e.g., any of the mutations provided herein). The present disclosure provides any deaminase domain that has a particular percent identity, as well as any of the mutations or combinations thereof described herein. In some embodiments, the adenosine deaminase comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to a reference sequence or any of the adenosine deaminases provided herein. In some embodiments, the adenosine deaminase comprises an amino acid sequence having at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, or at least 170 identical contiguous amino acid residues compared to any one of the amino acid sequences known in the art or described herein. It should be understood that any of the mutations provided herein (e.g., based on the TadA reference sequence) can be introduced into other adenosine deaminases, such as E. coli TadA (ecTadA), S. aureus TadA (saTadA), or other adenosine deaminases (e.g., bacterial adenosine deaminases). It will be apparent to one of skill in the art that additional deaminases can be similarly aligned to identify homologous amino acid residues and mutated as provided herein. Thus, any of the mutations identified in the TadA reference sequence can be made in other adenosine deaminases (e.g., ecTada) that have homologous amino acid residues. It should also be understood that any of the mutations provided herein can be made in the TadA reference sequence or another adenosine deaminase, either individually or in any combination.
[0286] In some embodiments, the adenosine deaminase comprises a D108X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a D108G, D108N, D108V, D108A, or D108Y mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase. However, it should be understood that additional deaminases can be aligned similarly to identify homologous amino acid residues and can be mutated as provided herein.
[0287] In some embodiments, the adenosine deaminase comprises an A106X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in a wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A106V mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0288] In some embodiments, the adenosine deaminase comprises an E155X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where the presence of X indicates any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an E155D, E155G, or E155V mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0289] In some embodiments, the adenosine deaminase comprises a D147X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where the presence of X indicates any amino acid other than the corresponding amino acid in wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises D147Y, a mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0290] In some embodiments, the adenosine deaminase comprises A106X, E155X, or D147X, a mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA), where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises E155D, E155G, or E155V mutation. In some embodiments, the adenosine deaminase comprises D147Y.
[0291] It should also be understood that any of the mutations provided herein can be made in ecTadA or another adenosine deaminase, individually or in any combination. For example, an adenosine deaminase can include D108N, A106V, E155V, and / or D147Y mutations in the TadA reference sequence, or the corresponding mutations in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises the following mutations in the TadA reference sequence (mutations are separated by ";"), or corresponding mutations in another adenosine deaminase: D108N and A106V; D108N and E155V; D108N and D147Y; A106V and E155V; A106V and D147Y; E155V and D147Y; D108N, A106V, and E155V; D108N, A106V, and D147Y; D108N, E155V, and D147Y; A106V, E155V, and D147Y; and D108N, A106V, E155V, and D147Y. However, it should be understood that any combination of the corresponding mutations provided herein can be made in an adenosine deaminase (eg, ecTadA).
[0292] In some embodiments, the adenosine deaminase comprises the following combinations of mutations in the TadA reference sequence (e.g., TadA*7.10) or corresponding mutations in another adenosine deaminase: V82G+Y147T+Q154S, I76Y+V82G+Y147T+Q154S, L36H+V82G+Y147T+Q154S+N157K, V82G+Y147D+F149Y+Q154S+D167N, L36H+V82G+Y147D+F149Y+Q154S+N 157K+D167N, L36H+I76Y+V82G+Y147T+Q154S+N157K, I76Y+V82G+Y147D+F149Y+Q154S+D167N, or L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N.
[0293] In some embodiments, the adenosine deaminase comprises one or more of the H8X, T17X, L18X, W23X, L34X, W45X, R51X, A56X, E59X, E85X, M94X, I95X, V102X, F104X, A106X, R107X, D108X, K110X, M118X, N127X, A138X, F149X, M151X, R153X, Q154X, I156X and / or K157X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein the presence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the H8Y, T17S, L18E, W23L, L34S, W45L, R51H, A56E or A56S, E59G, E85K or E85G, M94L, I95L, V102A, F104L, A106V, R107C or R107H or R107P, D108G or D108N or D108V or D108A or D108Y, K110I, M118K, N127S, A138V, F149Y, M151V, R153C, Q154L, I156D, and / or K157R mutations in a TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.
[0294] In some embodiments, the adenosine deaminase comprises one or more of an H8X, D108X and / or N127X mutation in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, where X indicates the presence of any amino acid. In some embodiments, the adenosine deaminase comprises one or more of an H8Y, D108N and / or N127S mutation in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.
[0295] In some embodiments, the adenosine deaminase comprises one or more of the H8X, R26X, M61X, L68X, M70X, A106X, D108X, A109X, N127X, D147X, R152X, Q154X, E155X, K161X, Q163X and / or T166X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in a wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of H8Y, R26W, M61I, L68Q, M70V, A106T, D108N, A109T, N127S, D147Y, R152C, Q154H or Q154R, E155G or E155V or E155D, K161Q, Q163H, and / or T166P in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.
[0296] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8X, D108X, N127X, D147X, R152X, and Q154X in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA), where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8X, M61X, M70X, D108X, N127X, Q154X, E155X, and Q163X in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA), where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8X, D108X, N127X, E155X, and T166X in the TadA reference sequence, or the corresponding mutation(s) in another adenosine deaminase (e.g., ecTadA), where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase.
[0297] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8X, A106X, and D108X, or a corresponding mutation(s) in another adenosine deaminase, where X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8X, R26X, L68X, D108X, N127X, D147X, and E155X, or a corresponding mutation in another adenosine deaminase, where X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase.
[0298] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, or seven mutations selected from the group consisting of H8X, R126X, L68X, D108X, N127X, D147X, and E155X in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8X, D108X, A109X, N127X, and E155X in the TadA reference sequence, or a corresponding one or more mutations in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase.
[0299] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8Y, D108N, N127S, D147Y, R152C, and Q154H in the TadA reference sequence, or a corresponding mutation(s) in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8Y, M61I, M70V, D108N, N127S, Q154R, E155G, and Q163H in the TadA reference sequence, or a corresponding mutation(s) in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8Y, D108N, N127S, E155V, and T166P in the TadA reference sequence, or a corresponding mutation(s) in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8Y, A106T, D108N, N127S, E155D, and K161Q in the TadA reference sequence, or a corresponding mutation(s) in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8Y, R26W, L68Q, D108N, N127S, D147Y, and E155V in the TadA reference sequence, or a corresponding mutation(s) in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8Y, D108N, A109T, N127S, and E155G in the TadA reference sequence, or the corresponding one or more mutations in another adenosine deaminase (e.g., ecTadA).
[0300] In some embodiments, the adenosine deaminase comprises one or more of the above mutations or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises a D108N, D108G, or D108V mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A106V and a D108N mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R107C and a D108N mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises H8Y, D108N, N127S, D147Y, and Q154H mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises H8Y, D108N, N127S, D147Y, and E155V mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises D108N, D147Y, and E155V mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises H8Y, D108N, and N127S mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises the A106V, D108N, D147Y, and E155V mutations in the TadA reference sequence, or the corresponding mutations in another adenosine deaminase (e.g., ecTadA).
[0301] In some embodiments, the adenosine deaminase comprises one or more of the S2X, H8X, I49X, L84X, H123X, N127X, I156X and / or K160X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, where the presence of X indicates any amino acid other than the corresponding amino acid in a wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the S2A, H8Y, I49F, L84F, H123Y, N127S, I156F and / or K160S mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase (e.g., ecTadA).
[0302] In some embodiments, the adenosine deaminase comprises an L84X mutant adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in a wild-type adenosine deaminase, In some embodiments, the adenosine deaminase comprises an L84F mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0303] In some embodiments, the adenosine deaminase comprises an H123X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in a wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an H123Y mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0304] In some embodiments, the adenosine deaminase comprises an I156X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an I156F mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.
[0305] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, or seven mutations selected from the group consisting of L84X, A106X, D108X, H123X, D147X, E155X, and I156X in the TadA reference sequence, or a corresponding one or more mutations in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of S2X, I49X, A106X, D108X, D147X, and E155X in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8X, A106X, D108X, N127X, and K160X in the TadA reference sequence, or the corresponding mutation(s) in another adenosine deaminase (where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase).
[0306] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six or seven mutations selected from the group consisting of L84F, A106V, D108N, H123Y, D147Y, E155V and I156F in the TadA reference sequence, or the corresponding one or more mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five or six mutations selected from the group consisting of S2A, I49F, A106V, D108N, D147Y and E155V in the TadA reference sequence.
[0307] In some embodiments, the adenosine deaminase comprises one, two, three, four or five mutations selected from the group consisting of H8Y, A106T, D108N, N127S and K160S in the TadA reference sequence, or the corresponding one or more mutations in another adenosine deaminase.
[0308] In some embodiments, the adenosine deaminase comprises one or more of the E25X, R26X, R107X, A142X and / or A143X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein the occurrence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of E25M, E25D, E25A, E25R, E25V, E25S, E25Y, R26G, R26N, R26Q, R26C, R26L, R26K, R107P, R107K, R107A, R107N, R107W, R107H, R107S, A142N, A142D, A142G, A143D, A143G, A143E, A143L, A143W, A143M, A143S, A143Q, and / or A143R in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the mutations described herein corresponding to the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.
[0309] In some embodiments, the adenosine deaminase comprises an E25X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in a wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an E25M, E25D, E25A, E25R, E25V, E25S, or E25Y mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0310] In some embodiments, the adenosine deaminase comprises a R26X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in a wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a R26G, R26N, R26Q, R26C, R26L, or R26K mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0311] In some embodiments, the adenosine deaminase comprises a R107X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in a wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a R107P, R107K, R107A, R107N, R107W, R107H, or R107S mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0312] In some embodiments, the adenosine deaminase comprises an A142X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in a wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A142N, A142D, A142G mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0313] In some embodiments, the adenosine deaminase comprises an A143X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in a wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A143D, A143G, A143E, A143L, A143W, A143M, A143S, A143Q, and / or A143R mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0314] In some embodiments, the adenosine deaminase comprises one or more of the H36X, N37X, P48X, I49X, R51X, M70X, N72X, D77X, E134X, S146X, Q154X, K157X and / or K161X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein the occurrence of X indicates any amino acid other than the corresponding amino acid in a wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the following mutations in the TadA reference sequence: H36L, N37T, N37S, P48T, P48L, I49V, R51H, R51L, M70L, N72S, D77G, E134G, S146R, S146C, Q154H, K157N, and / or K161T, or one or more corresponding mutations in another adenosine deaminase (e.g., ecTadA).
[0315] In some embodiments, the adenosine deaminase comprises an H36X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an H36L mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.
[0316] In some embodiments, the adenosine deaminase comprises an N37X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an N37T or N37S mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.
[0317] In some embodiments, the adenosine deaminase comprises a P48X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a P48T or P48L mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.
[0318] In some embodiments, the adenosine deaminase comprises an R51X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R51H or R51L mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.
[0319] In some embodiments, the adenosine deaminase comprises a S146X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a S146R or S146C mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.
[0320] In some embodiments, the adenosine deaminase comprises a K157X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in a wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a K157N mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.
[0321] In some embodiments, the adenosine deaminase comprises a P48X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a P48S, P48T, or P48A mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.
[0322] In some embodiments, the adenosine deaminase comprises an A142X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A142N mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.
[0323] In some embodiments, the adenosine deaminase comprises a W23X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a W23R or W23L mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.
[0324] In some embodiments, the adenosine deaminase comprises a R152X mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a R152P or R52H mutation in the TadA reference sequence or a corresponding mutation in another adenosine deaminase.
[0325] In one embodiment, the adenosine deaminase may comprise the mutations H36L, R51L, L84F, A106V, D108N, H123Y, S146C, D147Y, E155V, I156F, and K157N. In some embodiments, the adenosine deaminase comprises the following combinations of mutations relative to the TadA reference sequence, where each mutation in the combination is separated by an "_" and each combination of mutations is between brackets: (A106V_D108N), (R107C_D108N), (H8Y_D108N_N127S_D147Y_Q154H), (H8Y_D108N_N127S_D147Y_E155V), (D108N_D147Y_E155V), (H8Y_D108N_N127S), (H8Y_D108N_N127S_D147Y_Q154H), (A106V_D108N_D147Y_E155V), (D108Q_D147Y_E155V), (D108M_D147Y_E155V), (D108L_D147Y_E155V), (D108K_D147Y_E155V), (D108I_D147Y_E155V), (D108F_D147Y_E155V), (A106V_D108N_D147Y), (A106V_D108M_D147Y_E155V), (E59A_A106V_D108N_D147Y_E155V)、 (E59A cat dead_A106V_D108N_D147Y_E155V)、 (L84F_A106V_D108N_H123Y_D147Y_E155V_I156Y)、 (L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (D103A_D104N)、 (G22P_D103A_D104N)、 (D103A_D104N_S138A)、 (R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F)、 (E25G_R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F)、 (E25D_R26G_L84F_A106V_R107K_D108N_H123Y_A142N_A143G_D147Y_E155V_I156F)、 (R26Q_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F)、 (E25M_R26G_L84F_A106V_R107P_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F)、 (R26C_L84F_A106V_R107H_D108N_H123Y_A142N_D147Y_E155V_I156F)、 (L84F_A106V_D108N_H123Y_A142N_A143L_D147Y_E155V_I156F)、 (R26G_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F)、 (E25A_R26G_L84F_A106V_R107N_D108N_H123Y_A142N_A143E_D147Y_E155V_I156F)、 (R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F)、 (A106V_D108N_A142N_D147Y_E155V)、 (R26G_A106V_D108N_A142N_D147Y_E155V)、 (E25D_R26G_A106V_R107K_D108N_A142N_A143G_D147Y_E155V)、 (R26G_A106V_D108N_R107H_A142N_A143D_D147Y_E155V)、 (E25D_R26G_A106V_D108N_A142N_D147Y_E155V)、 (A106V_R107K_D108N_A142N_D147Y_E155V)、 (A106V_D108N_A142N_A143G_D147Y_E155V)、 (A106V_D108N_A142N_A143L_D147Y_E155V)、 (H36L_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、 (N37T_P48T_M70L_L84F_A106V_D108N_H123Y_D147Y_I49V_E155V_I156F)、 (N37S_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K161T)、 (H36L_L84F_A106V_D108N_H123Y_D147Y_Q154H_E155V_I156F)、 (N72S_L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F)、 (H36L_P48L_L84F_A106V_D108N_H123Y_E134G_D147Y_E155V_I156F)、 (H36L_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K157N) (H36L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F)、 (L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F_K161T)、 (N37S_R51H_D77G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (R51L_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K157N)、 (D24G_Q71R_L84F_H96L_A106V_D108N_H123Y_D147Y_E155V_I156F_K160E)、 (H36L_G67V_L84F_A106V_D108N_H123Y_S146T_D147Y_E155V_I156F)、 (Q71L_L84F_A106V_D108N_H123Y_L137M_A143E_D147Y_E155V_I156F)、 (E25G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_Q159L)、 (L84F_A91T_F104I_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (N72D_L84F_A106V_D108N_H123Y_G125A_D147Y_E155V_I156F)、 (P48S_L84F_S97C_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (W23G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (D24G_P48L_Q71R_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_Q159L)、 (L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F)、 (H36L_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_E155V_I156F_K157N)、 (N37S_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F_K161T)、 (L84F_A106V_D108N_D147Y_E155V_I156F)、 (R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N_K161T)、 (L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K161T)、 (L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N_K160E_K161T)、 (L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N_K160E)、 (R74Q_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (R74A_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (R74Q_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (L84F_R98Q_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (L84F_A106V_D108N_H123Y_R129Q_D147Y_E155V_I156F)、 (P48S_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F)、 (P48S_A142N)、 (P48T_I49V_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F_L157N)、 (P48T_I49V_A142N)、 (H36L_P48S_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、 (H36L_P48S_R51L_L84F_A106V_D108N_H123Y_S146C_A142N_D147Y_E155V_I156F(H36L_P48T_I49V_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、 (H36L_P48T_I49V_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_E155V_I156F_K157N)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_E155V_I156F_K157N)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_A142N_D147Y_E155V_I156F_K157N)、 (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、 (W23R_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、 (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F_K161T)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152H_E155V_I156F_K157N)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152P_E155V_I156F_K157N)、 (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152P_E155V_I156F_K157N)、 (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142A_S146C_D147Y_E155V_I156F_K157N)、 (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142A_S146C_D147Y_R152P_E155V_I156F_K157N)、 (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F_K161T)、 (W23R_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152P_E155V_I156F_K157N)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_R152P_E155V_I156F_K157N)。
[0326] In some embodiments, the TadA deaminase is a TadA variant. In some embodiments, the TadA variant is TadA*7.10. In certain embodiments, the fusion protein comprises a single TadA*7.10 domain (e.g., provided as a monomer). In other embodiments, the fusion protein comprises TadA*7.10 and TadA(wt), which can form a heterodimer. In one embodiment, the fusion protein of the invention comprises wild-type TadA bound to TadA*7.10, which is bound to a Cas9 nickase.
[0327] In some embodiments, TadA*7.10 comprises at least one modification. In some embodiments, the adenosine deaminase comprises a modification in the following sequence: TadA*7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 1)
[0328] In some embodiments, TadA*7.10 comprises modifications at amino acids 82 and / or 166. In certain embodiments, TadA*7.10 comprises one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. In other embodiments, the variant of TadA*7.10 comprises a combination of modifications selected from the following group: Y147T+Q154R, Y147T+Q154S, Y147R+Q154S, V82S+Q154S, V82S+Y147R, V82S+Q154R, V82S+Y123H, I76Y+V82S, V82S+Y123H+Y147T, V82 S+Y123H+Y147R, V82S+Y123H+Q154R, Y147R+Q154R+Y123H, Y147R+Q154R+I76Y, Y147R+Q154R+T166R, Y123H+Y147R+Q154R+I76Y, V82S+Y123H+Y147R+Q154R, and I76Y+V82S+Y123H+Y147R+Q154R.
[0329] In some embodiments, the variant of TadA*7.10 comprises one or more modifications selected from the group of L36H, I76Y, V82G, Y147T, Y147D, F149Y, Q154S, N157K, and / or D167N. In some embodiments, the variant of TadA*7.10 comprises V82G, Y147T / D, Q154S, and one or more of L36H, I76Y, F149Y, N157K, and D167N. In other embodiments, the variant of TadA*7.10 comprises a combination of modifications selected from the following group: V82G+Y147T+Q154S, I76Y+V82G+Y147T+Q154S, L36H+V82G+Y147T+Q154S+N157K, V82G+Y147D+F149Y+Q154S+D167N, L36 H+V82G+Y147D+F149Y+Q154S+N157K+D167N, L36H+I76Y+V82G+Y147T+Q154S+N157K, I76Y +V82G+Y147D+F149Y+Q154S+D167N, L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N.
[0330] In some embodiments, the adenosine deaminase variant (e.g., TadA*8) comprises a deletion. In some embodiments, the adenosine deaminase variant comprises a C-terminal deletion. In certain embodiments, the adenosine deaminase variant comprises TadA*7.10, a C-terminal deletion starting at residues 149, 150, 151, 152, 153, 154, 155, 156, and 157 relative to the TadA reference sequence, or a corresponding mutation in another TadA.
[0331] In other embodiments, the adenosine deaminase variant (e.g., TadA*8) is a monomer that includes one or more of the following modifications: TadA*7.10, Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R relative to the TadA reference sequence, or corresponding mutations in another TadA. In another embodiment, the variant of adenosine deaminase (TadA*8) is a monomer that comprises a combination of modifications selected from the following group: TadA*7.10, Y147T+Q154R, Y147T+Q154S, Y147R+Q154S, V82S+Q154S, V82S+Y147R, V82S+Q154R, V82S+Y123H, I76Y+V82S, V82S+Y123H+ Y147T, V82S+Y123H+Y147R, V82S+Y123H+Q154R, Y147R+Q154R+Y123H, Y147R+Q154R+I76Y, Y147R+Q154R+T166R, Y123H+Y147R+Q154R+I76Y, V82S+Y123H+Y147R+Q154R, and I76Y+V82S+Y123H+Y147R+Q154R, or the corresponding mutations in another TadA.
[0332] In other embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase domains (e.g., TadA*8), each having one or more of the following modifications relative to the TadA reference sequence, TadA*7.10, Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R, or corresponding mutations in another TadA. In other embodiments, the adenosine deaminase variant is a monomer that includes two adenosine deaminase domains (e.g., TadA*8), each having a combination of modifications selected from the following group: TadA*7.10, Y147T+Q154R, Y147T+Q154S, Y147R+Q154S, V82S+Q154S, V82S+Y147R, V82S+Q154R, V82S+Y123H, I76Y+V 82S, V82S+Y123H+Y147T, V82S+Y123H+Y147R, V82S+Y123H+Q154R, Y147R+Q154R+Y123H, Y147R+Q154R+I76Y, Y147R+Q154R+T166R, Y123H+Y147R+Q154R+I76Y, V82S+Y123H+Y147R+Q154R, and I76Y+V82S+Y123H+Y147R+Q154R, or the corresponding mutations in another TadA.
[0333] In other embodiments, base editors of the disclosure comprise variant adenosine deaminase (e.g., TadA*8) monomers and include one or more of the following modifications: TadA*7.10, R26C, V88A, A109S, T111R, D119N, H122N, Y147D, F149Y, T166I and / or D167N relative to the TadA reference sequence, or corresponding mutations in another TadA. In other embodiments, the variant (TadA*8) monomer of adenosine deaminase comprises a combination of modifications selected from the following group: TadA*7.10, R26C+A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N, V88A ... 22N+F149Y+T166I+D167N, R26C+A109S+T111R+D119N+H122N+F149Y+T166I+D167N, V88A+T111R+D119N+F149Y, and A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N, or the corresponding mutations in another TadA.
[0334] In some embodiments, the adenosine deaminase variant (e.g., MSP828) is a monomer that includes one or more of the following modifications, L36H, I76Y, V82G, Y147T, Y147D, F149Y, Q154S, N157K, and / or D167N, relative to TadA*7.10, the TadA reference sequence, or the corresponding mutation in another TadA. In some embodiments, the adenosine deaminase variant (e.g., MSP828) is a monomer that includes one or more of the following modifications, V82G, Y147T / D, Q154S, and L36H, I76Y, F149Y, N157K, and D167N, relative to TadA*7.10, the TadA reference sequence, or the corresponding mutation in another TadA. In other embodiments, the adenosine deaminase variant (TadA variant) is a monomer that comprises a combination of modifications relative to the TadA reference sequence TadA*7.10 selected from the following group or corresponding mutations in another TadA: V82G+Y147T+Q154S, I76Y+V82G+Y147T+Q154S, L36H+V82G+Y147T+Q154S+N157K, V ... G+Y147D+F149Y+Q154S+D167N, L36H+V82G+Y147D+F149Y+Q154S+N157K+D167N, L36H+I76Y+V82G+Y147T+Q1 54S+N157K, I76Y+V82G+Y147D+F149Y+Q154S+D167N, L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N.
[0335] In other embodiments, the adenosine deaminase variant is a heterodimer of a wild-type adenosine deaminase domain and an adenosine deaminase variant domain that includes one or more of the following modifications Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R relative to the TadA reference sequence, TadA*7.10, or a corresponding mutation in another TadA (e.g., TadA*8). In other embodiments, the variant of adenosine deaminase is a heterodimer of a wild-type adenosine deaminase domain and an adenosine deaminase variant domain (e.g., TadA*8) that includes a combination of modifications selected from the following group: TadA*7.10, Y147T+Q154R, Y147T+Q154S, Y147R+Q154S, V82S+Q154S, V82S+Y147R, V82S+Q154R, V82S+Y12 relative to the TadA reference sequence. 3H, I76Y+V82S, V82S+Y123H+Y147T, V82S+Y123H+Y147R, V82S+Y123H+Q154R, Y147R+Q154R+Y123H, Y147R+Q154R+I76Y, Y147R+Q154R+T166R, Y123H+Y147R+Q154R+I76Y, V82S+Y123H+Y147R+Q154R, and I76Y+V82S+Y123H+Y147R+Q154R, or the corresponding mutations in another TadA.
[0336] In other embodiments, a base editor of the disclosure comprises an adenosine deaminase variant (e.g., TadA*8) homodimer that includes two adenosine deaminase domains (e.g., TadA*8), each having one or more of the following modifications, R26C, V88A, A109S, T111R, D119N, H122N, Y147D, F149Y, T166I and / or D167N, relative to TadA*7.10, a TadA reference sequence, or a corresponding mutation in another TadA. In other embodiments, the adenosine deaminase variant is a monomer that includes two adenosine deaminase domains (e.g., TadA*8), each having a combination of modifications selected from the following group: TadA*7.10, R26C+A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N, V88A+A10 relative to the TadA reference sequence. 9S+T111R+D119N+H122N+F149Y+T166I+D167N, R26C+A109S+T111R+D119N+H122N+F149Y+T166I+D167N, V88A+T111R+D119N+F149Y, and A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N, or the corresponding mutations in another TadA.
[0337] In some embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase domains (e.g., TadA*7.10), each having one or more of the following modifications relative to TadA*7.10, the TadA reference sequence, or a corresponding mutation in another TadA: L36H, I76Y, V82G, Y147T, Y147D, F149Y, Q154S, N157K, and / or D167N. In some embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase variant domains (e.g., MSP828), each having one or more of the following modifications, V82G, Y147T / D, Q154S, and L36H, I76Y, F149Y, N157K, and D167N, relative to TadA*7.10, the TadA reference sequence, or a corresponding mutation in another TadA. In other embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase domains (e.g., TadA*7.10), each of which has a combination of modifications relative to the TadA reference sequence TadA*7.10 selected from the following group or corresponding mutations in another TadA: V82G+Y147T+Q154S, I76Y+V82G+Y147T+Q154S, L36H+V 82G+Y147T+Q154S+N157K, V82G+Y147D+F149Y+Q154S+D167N, L36H+V82G+Y147D+F149Y+Q154S+N157K+D167N, L36H+I76Y+ V82G+Y147T+Q154S+N157K, I76Y+V82G+Y147D+F149Y+Q154S+D167N, L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N.
[0338] In other embodiments, the adenosine deaminase variant is a heterodimer of a TadA*7.10 domain and an adenosine deaminase variant domain that includes one or more of the following modifications Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R relative to the TadA*7.10, TadA reference sequence, or a corresponding mutation in another TadA (e.g., TadA*8). In other embodiments, the adenosine deaminase variant is a heterodimer of a TadA*7.10 domain and an adenosine deaminase variant domain (e.g., TadA*8), wherein the adenosine deaminase variant domain compris...
Claims
**Claim 1** A polynucleotide encoding a deaminase domain, a nucleic acid programmable DNA binding protein (napDNAbp) domain, a base editor polypeptide, or a fragment thereof, wherein the polynucleotide comprises an intron, and the intron is inserted into an open reading frame encoding the deaminase, napDNAbp, base editor polypeptide, or a fragment thereof, optionally, the intron comprises a modification at a splice acceptor or splice donor site, the modification reducing or eliminating splicing of the base editor mRNA, thereby reducing or eliminating expression of the deaminase, napDNAbp, base editor polypeptide, or a fragment thereof, said polynucleotide. **Claim 2** A polynucleotide encoding a base editor comprising a nucleic acid programmable DNA binding protein (napDNAbp) domain and a deaminase domain, wherein the polynucleotide comprises an intron, and the intron is inserted into an open reading frame encoding the napDNAbp domain or the deaminase domain, optionally, the intron comprises a modification at a splice acceptor or splice donor site, the modification reducing splicing of the base editor mRNA, said polynucleotide. **Claim 3** i) the intron is derived from a sequence selected from the group consisting of NF1, PAX2, EEF1A1, HBB, IGHG1, SLC50A1, ABCB11, BRSK2, PLXNB3, TMPRSS6, IL32, ANTXRL, PKHD1L1, PADI1, KRT6C, and HMCN2; and / or ii) the intron is as follows: a) GTGAGATCAAATGAAAGTTTCATATAGAAATACAAAACCTAGAGAACTGGCATGTAAGAGAAGCAAAAATTACTTCAGCAAGGCCATGTTAGTAAATTTGCATCTGTTTGTCCACATTAG (SEQ ID NO: 226), b) GTAGGTGACAATGCTGCAGCTGCCTAATCTAGGTGGGGGGAACTAAATTGTGGGTGAGCTGCTGAATGGTCTGTAGTCTGAGGCTGGGGTGGGGGGAGACACAACGTCCCCTCCCTGCAAACCACTGCTATTCTGTCCCTCTCTCTCCTTAG (SEQ ID NO: 227), c) GTAAGTGGCTTTCAAGACCATTGTTAAAAAGCTCTGGGAATGGCGATTTCATGCTTACATAAATTGGCATGCTTGTGTTTCAG (SEQ ID NO: 228), d) GTAAGTATCAAGGTTACAAGACAGGTTTAAGGAGACCAATAGAAACTGGGCTTGTCTAGACAGAGAAGACTCTTGCGTTTCTGATAGGCACCTATTGGTCTTACTGACATCCACTTTGCCTTTCTCTCCACAG (SEQ ID NO: 229), e) GTAAGCACAACTGGGATGGGGTGACAGGGGTGCAAGATTGAAAACTGGCTCCTCTCCTCATAGCAGTTCTTGTGATTTCAG (SEQ ID NO: 230), f) GTAAGAAATGTTATTTTTCAGTAAGTGATTTAGTTATTTTTCCTTTTTTCTCATTAAAATTTCTCTAACATCTCCCTCTTCATGTTTTAG (SEQ ID NO: 231), g) GTGAGACCCTAGCCCCCTCAACCCTGCCCTGGCCTCTCCCCAAACCTGCCCCCCCACGCTGACCCCCACACCCGGCCGCCCGCAG (SEQ ID NO: 232), h) GTGGGTGTCAGAGGCATCGGGGCTGCGGGGTAGGGGGCTGCCCCACCCCTAACGAAGTCTGCTCCTCCAG (SEQ ID NO: 233), i) GCAGGGAAGTCCTGCTTCCGTGCCCCACCGGTGCTCAGCTGAGGCTCCCTTGAAAATGCGAGGCTGTTTCCAACTTTGGTCTGTTTCCCTGGCAG (SEQ ID NO: 234), j) GTGGGGAGTTGGGGTCCCCGAAGGTGAGGACCCTCTGGGGATGAGGGTGCTTCTCTGAGACACTTTCTTTTCCTCACACCTGTTCCTCGCCAGCAG (SEQ ID NO: 235), k) GTATAGACCCCTTGATCTCCTAACCCTAACCCTAACCCTAACCCTAACCTACAAAATCTTAGAGCATCAGTGGGAGCATCTCACTGTCCAGGCTCAATATTTCTTCATTTTCTTGCAG (SEQ ID NO: 236), l) GTAATTATGATTAAAGATGGTGATTGTTTATTTTCTTTTATGATTGTCCTTAGTATTATGTAACCTGCAAATTCTATTGCAG (SEQ ID NO: 237), m) GTGAGTGACACAAGGTGTTGTCTGGGGAGTGGGGAAGGGGGATGGAAGTGAATCCTGTTGGTGGGGTGGAGAAAGGGCGATCTCAAGAGGGCCACTCTCTCCAG (SEQ ID NO: 238), n) GTAAGCATCTCCACCATCCTTCTGTTTACTCTGATGGGGTCTGCAAAGGGGAGATGATGTATAGGGTTGGGTATCTCTGTAAATGTCAGATGTGAAGTTGATCTTATGACCTTCTGTTCTGCAG (SEQ ID NO: 239), o) GTGAGGGTCTCCCAGGCTGGGCAGGGGGAGGGGGCTGCTGCCTTGATTGCGTCCCAGGACACAGCCCTCCTCCAGCCTGCCCTCGCCTTGCTCATCCCCTCCCCATCTCAGCCCCACCCCCACTAACTCTCTCTCTGCTCTGACTCAG (SEQ ID NO: 240), p) GTAATGATTGATTGCAATGTATGATTACAATAATCTCAGTATAAGTTCAGTAATAATAACCTTCCACTGCTGTCCTCTGTGTGCACCCAG (SEQ ID NO: 241), or q) GTAAATATATACAACAGTTTTTCATTTAAATAAGTGCACGGCACAAATAAGAAAAATATGTCAAAAATGTAACCAATAGTTTTTTTCAAATTTAG (SEQ ID NO: 242) comprises one of the sequences or a sequence having at least about 85% nucleic acid sequence identity to one of the sequences; and / or (iii) The intron contains from about 10 base pairs to about 500 base pairs, and in particular, the intron contains from about 70 base pairs to 150 base pairs, or the intron contains from about 100 base pairs to 200 base pairs. The polynucleotide according to claim 1 or 2. **Claim 4** The polynucleotide according to claim 1 or 2, wherein the intron is inserted within about 10 to 30 base pairs of the protospacer sequence, and optionally, the protospacer sequence is NGG or NNGRRT. **Claim 5** a) The deaminase domain contains a TadA domain, and optionally, the intron is inserted within or immediately after codons 18, 23, 59, 62, 87, or 129 of TadA, and further optionally, the intron is inserted immediately after codon 87 of TadA; and / or b) further comprises a polynucleotide sequence encoding a linker, and optionally, the intron is inserted within the polynucleotide sequence encoding the linker; and / or c) the programmable DNA-binding protein domain is a Cas9 domain, and optionally, the Cas9 domain is split between the amino acid residue corresponding to Asn309 of Cas9 and the amino acid residue corresponding to Thr310, and residue 310 is mutated to Cys. The polynucleotide according to claim 1 or 2. **Claim 6** A composition comprising: (i) a first polynucleotide encoding an N-terminal fragment of a deaminase domain and a nucleic acid programmable DNA-binding protein (napDNAbp) domain, wherein the N-terminal fragment of the napDNAbp domain is fused to split intein-N; and (ii) a second polynucleotide encoding a C-terminal fragment of the napDNAbp domain, wherein the C-terminal fragment of the napDNAbp domain is fused to split intein-C; or (i) a first polynucleotide encoding an N-terminal fragment of a deaminase domain, wherein the N-terminal fragment of the deaminase domain is fused to split intein-N; and (ii) a second polynucleotide encoding a C-terminal fragment of the deaminase domain and a nucleic acid programmable DNA-binding protein (napDNAbp) domain, wherein the C-terminal fragment of the deaminase domain is fused to split intein-C. The first polynucleotide or the second polynucleotide contains an intron, and the intron is inserted within the open reading frame of the polynucleotide, optionally, the intron contains modifications at splice acceptor or splice donor sites, and the modifications reduce or eliminate splicing of the base editor mRNA, the composition. **Claim 7**: a) the napDNAbp domain is a Cas9 domain, where: i) the N-terminal and C-terminal domains of the Cas9 domain are split between amino acid residues Asn309 and Thr310; and / or ii) the Cas9 domain contains the mutation Thr310Cys; and / or b) the intron is derived from a sequence selected from the group consisting of NF1, PAX2, EEF1A1, HBB, IGHG1, SLC50A1, ABCB11, BRSK2, PLXNB3, TMPRSS6, IL32, ANTXRL, PKHD1L1, PADI1, KRT6C, and HMCN2; and / or c) further comprises a linker polynucleotide sequence, and optionally the intron is inserted within the linker polynucleotide sequence, The composition according to claim 6. **Claim 8** (i) A polynucleotide encoding a base editor or a fragment thereof containing a deaminase domain, (ii) One or more guide RNAs that direct the base editor to edit a site within the genome of a cell, (iii) One or more guide RNAs that direct the base editor to edit the polynucleotide encoding the base editor, A base editor system comprising, wherein the editing in (iii) results in a decrease in the activity and / or expression of the encoded base editor, the base editor system. **Claim 9**: The editing in (iii) modifies the catalytic residue of the deaminase domain, Preferably, the deaminase domain is an adenosine deaminase domain, and the catalytic residue to be modified in the deaminase domain is one of His57 (H57), Glu59 (E59), Cys87 (C87) or Cys90 (C90) of the following reference sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 1), or the corresponding position in another adenosine deaminase. More preferably, the modification to the catalytic residue is E59G, H57R, C87R, or C90R. The base editor system according to claim 8. **Claim 10** A base editor system, comprising: (i) a polynucleotide encoding a self-inactivating base editor or a fragment thereof, and comprising an intron inserted within the open reading frame of the self-inactivating base editor or the fragment thereof; (ii) one or more guide RNAs that direct the self-inactivating base editor to edit a site within the genome of a cell; (iii) one or more guide RNAs that direct the self-inactivating base editor to edit a splice acceptor or splice donor site present within the intron of the polynucleotide encoding the self-inactivating base editor; or (i) the polynucleotide according to claim 1 or 2, encoding a base editor; (ii) one or more guide RNAs that direct the base editor to edit a site within the genome of a cell; (iii) one or more guide RNAs that direct the base editor to edit a splice acceptor site or splice donor site present within the intron of the polynucleotide encoding the base editor; or (i) the composition according to claim 6 or 7, encoding a base editor; (ii) one or more guide RNAs that direct the base editor to edit a site within the genome of a cell. (iii) One or more guide RNAs that induce the base editor to edit a splice acceptor or splice donor site present within the intron of the composition of (i). The base editor system. **Claim 11**: A base editor system, comprising: (i) A first polynucleotide encoding an N-terminal fragment of a deaminase domain and a nucleic acid programmable DNA-binding protein (napDNAbp) domain, wherein the N-terminal fragment of the napDNAbp domain is fused to split intein-N; (ii) A second polynucleotide encoding a C-terminal fragment of the napDNAbp domain, wherein the C-terminal fragment of the napDNAbp domain is fused to split intein-C; The first polynucleotide or the second polynucleotide contains an intron, the intron is inserted within an open reading frame, the first polynucleotide and the second polynucleotide encode a base editor, (iii) One or more guide RNAs that induce the base editor to edit a site within the genome of a cell; (iv) One or more guide RNAs that induce the base editor to edit a splice acceptor or splice donor site present within the intron of the polynucleotide of (i) or (ii); or (i) A first polynucleotide encoding an N-terminal fragment of a deaminase domain, wherein the N-terminal fragment of the deaminase domain is fused to split intein-N; (ii) A second polynucleotide encoding a C-terminal fragment of the deaminase domain and a nucleic acid programmable DNA-binding protein (napDNAbp) domain, wherein the C-terminal fragment of the deaminase domain is fused to split intein-C; The first polynucleotide or the second polynucleotide contains an intron, the intron is inserted within an open reading frame, the first polynucleotide and the second polynucleotide encode a base editor, (iii) One or more guide RNAs that induce the base editor to edit a site within the genome of a cell. (iv) one or more guide RNAs that induce the base editor to edit a splice acceptor or splice donor site present within the intron of the polynucleotide of (i) or (ii), the base editor system. **Claim 12** The base editor system is as follows: a) gGUUUUAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 191), b) gUUUCUUACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 192), c) gGUUUCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 193), d) GCCACUUACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 194), e) gACAUUAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 195), f) gGAUCUCACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 196), g) gUCCUUAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 197), h) GUCACCUACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 198), i) GAUUUCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 190), j) gGUGCUUACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 200), k) gUCCACAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 201), l) GAUACUUACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 202), m) gUGUUUUAGCUGCGGCAAGGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 203), n) gUUUCUUACAGCCAUAAUUUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 204), o) gCUCCACAGCUGCGGCAAGGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 205), p) GAUACUUACAGCCAUAAUUUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 206), q) gUGUUUUAGGGACGAAAGAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 207), r) gUUACCUGGCUCUCUUAGCCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 208), s) gCUCCACAGGGACGAAAGAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 209), t) gCUUGCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 210), u) gAUUGCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 211), v) gUCUCCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 212), w) gUCUGCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 213), x) gGACUCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 214), y) GCACCCAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 215), z) gAAUUUAGGUCAUGUGUGCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 216), aa) gCAUUAGGUCGAGAUCACAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 217), bb) gCCUUAGGUCGAGAUCACAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 218), cc) GUUUCAGGUCGAGAUCACAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 219), dd) gACAUUAGGCUAAGAGAGCCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 220), ee) gUCCUUAGGCUAAGAGAGCCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 221), ff) gGUUUCAGGCUAAGAGAGCCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 222), gg) gACAUUAGAUUAUGGCUCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 223), hh) gUCCUUAGAUUAUGGCUCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 224), ii) gGUUUCAGAUUAUGGCUCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 225), jj) gCACCAUGAGCGAGGUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 524), kk) gGCCACCAUGAGCGAGGUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 525), ll) GUGUCGAAGUUCGCCCUGGAGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 526), mm) gAUGCCGAGAUAAUGGCCCUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 527), nn) gAUGCCGAGAUAAUGGCCCUUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 528), oo) gAUGCCGAGAUCAUGGCACUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 529), pp) gAUGCCGAGAUCAUGGCACUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 530), qq) gAUGCCGAGAUCAUGGCACUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 531), rr) gAUGCCGAGAUCAUGGCGCUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 532), ss) gAUGCCGAGAUCAUGGCGCUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 533), tt) gAUGCCGAGAUCAUGGCGUUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 534), uu) gAUGCCGAGAUUAUGGCACUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 535), vv) gAUGCCGAGAUUAUGGCACUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 536), ww) gAUGCCGAGAUUAUGGCACUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 537), xx) gAUGCCGAGAUUAUGGCACUUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 538), yy) gAUGCCGAGAUUAUGGCGCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 539), zz) gAUGCCGAGAUUAUGGCUCUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 540), aaa) gAUGCGGAGAUCAUGGCGCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 541), bbb) gAUGCUGAGAUAAUGGCCCUCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 542), ccc) gAACCGCACAUGCCGAAAUUAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 543), ddd) gGCAGGUGUCGACAUAUCUAUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 544), eee) gAUGCCGAAAUUAUGGCUCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 545), fff) gACACAUGACACAGGGCUCGAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 546), or ggg) gGCCCCAGCACACAUGACACAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (SEQ ID NO: 547) The base editor system according to any one of claims 8, 9, and 11, comprising a polynucleotide sequence selected from **Claim 13** A vector, a) comprising a polynucleotide encoding a self-inactivating base editor or a fragment thereof, wherein the polynucleotide contains an intron inserted within the open reading frame of the self-inactivating base editor or a fragment thereof; or b) comprising the polynucleotide according to claim 1 or 2, or the base editor system according to any one of claims 8, 9, and 11; or c) comprising the first polynucleotide and / or the second polynucleotide of the composition according to claim 6 or 7, The said vector. **Claim 14** The said vector is a) a lipid nanoparticle; or b) a viral vector selected from the group consisting of adeno-associated virus (AAV), retroviral vector, adenoviral vector, lentiviral vector, Sendai virus vector, and herpesvirus vector, optionally, the AAV vector is AAV2 or AAV8, The vector according to claim 13. **Claim 15** A cell, a) a cell comprising a vector comprising a polynucleotide encoding a self-inactivating base editor or a fragment thereof, wherein the polynucleotide comprises an intron inserted within the open reading frame of the self-inactivating base editor or the fragment thereof; or b) a cell comprising the polynucleotide according to claim 1 or 2, the composition according to claim 6 or 7, or the base editor system according to any one of claims 8, 9, 11; or c) a vector, i) a polynucleotide encoding a self-inactivating base editor or a fragment thereof, wherein the polynucleotide comprises an intron inserted within the open reading frame of the self-inactivating base editor or the fragment thereof; or ii) the polynucleotide according to claim 1 or 2, or the base editor system according to any one of claims 8, 9, 11, or iii) the cell comprising the vector, comprising the first polynucleotide and / or the second polynucleotide of the composition according to claim 6 or 7, the cell.
16. A pharmaceutical composition comprising the polynucleotide according to claim 1 or 2, or the base editor system according to any one of claims 8, 9, 11, and a pharmaceutically acceptable excipient, diluent, or carrier.
17. A kit comprising the polynucleotide according to claim 1 or 2, the composition according to claim 6 or 7, or the base editor system according to any one of claims 8, 9, 11.
18. A composition for use in a method for reducing or eliminating the expression of a self-inactivating base editor, comprising a guide RNA, wherein the method comprises (a) providing a polynucleotide encoding a self-inactivating base editor or a fragment thereof, wherein the polynucleotide comprises an intron inserted within the open reading frame of the self-inactivating base editor or the fragment thereof, the providing; (b) contacting the polynucleotide with the guide RNA and the self-inactivating base editor polypeptide comprised in the composition, wherein the guide RNA induces the base editor to edit a splice acceptor or splice donor site of the intron, thereby generating a modification that reduces or eliminates the expression of the self-inactivating base editor, and the contacting, in the composition.
19. A composition for use in a method of self-inactivating base editing, comprising a second guide RNA, the method comprising: (a1) expressing, within a cell, a polynucleotide encoding a base editor or a fragment thereof comprising a deaminase domain; (b1) contacting the cell with a first guide RNA, wherein the first guide RNA induces the base editor to edit a site within the genome of the cell, thereby generating a modification within the genome of the cell, and the contacting; (c1) contacting the cell with the second guide RNA, wherein the second guide RNA induces the base editor to edit the polynucleotide encoding the base editor, the editing resulting in a decrease in the activity and / or expression of the encoded base editor, thereby generating a modification that reduces or eliminates the expression of the base editor, and the contacting; or (a2) expressing, within a cell, a polynucleotide encoding a self-inactivating base editor or a fragment thereof, wherein the polynucleotide comprises an intron inserted within the open reading frame of the self-inactivating base editor or the fragment thereof, and the expressing; (b2) contacting the cell with a first guide RNA, wherein the first guide RNA induces the self-inactivating base editor to edit a site within the genome of the cell, thereby generating a modification within the genome of the cell, and the contacting; (c2) contacting the cell with the second guide RNA, wherein the second guide RNA induces the self-inactivating base editor to edit a splice acceptor or splice donor site present within the intron of the polynucleotide of (a2), thereby generating a modification that reduces or eliminates the expression of the self-inactivating base editor, and said contacting, the composition. **Claim 20** The editing in (c1) modifies the catalytic residues of the deaminase domain, and optionally, the deaminase domain is an adenosine deaminase domain, preferably, the catalytic residue of the deaminase domain to be modified is one of His57 (H57), Glu59 (E59), Cys87 (C87) or Cys90 (C90) of the following reference sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 1), or the corresponding position in another adenosine deaminase, and more preferably, the modification to the catalytic residue is E59G, H57R, C87R, or C90R, The composition according to claim 19. **Claim 21** A composition for use in a method of editing the genome of an organism, comprising a second guide RNA, wherein the method comprises: (a1) expressing, in a cell of the organism, a polynucleotide encoding a self-inactivating base editor or a fragment thereof, wherein the polynucleotide comprises an intron inserted within the open reading frame of the self-inactivating base editor or the fragment thereof, and said expressing, (b1) contacting the cell with a first guide RNA, wherein the first guide RNA induces the self-inactivating base editor to edit a site within the genome of the cell, thereby generating a modification within the genome of the cell, and said contacting, (c1) contacting the cell with the second guide RNA, wherein the second guide RNA induces the self-inactivating base editor to edit a splice acceptor or splice donor site present within the intron of the polynucleotide of (a1), thereby generating a modification that reduces or eliminates the expression of the self-inactivating base editor, the contacting; or (a2) expressing, within a cell of the organism, a first polynucleotide encoding a deaminase domain and an N-terminal fragment of a nucleic acid programmable DNA binding protein (napDNAbp) domain, wherein the N-terminal fragment of the napDNAbp domain is fused to split intein-N, and a second polynucleotide encoding a C-terminal fragment of the napDNAbp domain, wherein the C-terminal fragment of the napDNAbp domain is fused to split intein-C, wherein the first polynucleotide or the second polynucleotide comprises an intron, the intron is inserted within an open reading frame, and the expression of the first polynucleotide and the second polynucleotide within the cell results in the formation of a self-inactivating base editor, the expressing; (b2) contacting the cell with a first guide RNA, wherein the first guide RNA induces the self-inactivating base editor to edit a site within the genome of the cell, thereby generating a modification within the genome of the cell, the contacting; (c2) contacting the cell with the second guide RNA, wherein the second guide RNA induces the self-inactivating base editor to edit a splice acceptor or splice donor site present within the intron of the polynucleotide of (a2), thereby generating a modification that reduces or eliminates the expression of the self-inactivating base editor, the contacting; or (a3)In the cells of the organism, expressing a first polynucleotide encoding an N-terminal fragment of the deaminase domain, wherein the N-terminal fragment of the deaminase domain is fused to Split Intein-N, and a second polynucleotide encoding a C-terminal fragment of the deaminase domain and a nucleic acid programmable DNA-binding protein (napDNAbp) domain, wherein the C-terminal fragment of the deaminase domain is fused to Split Intein-C, wherein the first polynucleotide or the second polynucleotide contains an intron, the intron is inserted within the open reading frame, and the expression of the first polynucleotide and the second polynucleotide in the cell results in the formation of a self-inactivating base editor, the step of expressing; (b3)Contacting the cell with a first guide RNA, wherein the first guide RNA induces the self-inactivating base editor to edit a site within the genome of the cell, thereby generating a modification within the genome of the cell, the step of contacting; (c3)Contacting the cell with a second guide RNA, wherein the second guide RNA induces the self-inactivating base editor to edit a splice acceptor or splice donor site present within the intron of the polynucleotide of (a3), thereby generating a modification that reduces or eliminates the expression of the self-inactivating base editor, the step of contacting; comprising The composition. A pharmaceutical composition for use in a method of treating a subject, comprising a polynucleotide encoding a self-inactivating base editor or a fragment thereof, wherein the method comprises (a)Delivering into the cells of the subject a polynucleotide encoding a self-inactivating base editor or a fragment thereof for expression, wherein the polynucleotide contains an intron inserted within the open reading frame of the self-inactivating base editor or a fragment thereof, the step of delivering; (b) contacting the cell with a first guide RNA, wherein the first guide RNA induces the self-inactivating base editor to edit a site within the genome of the cell, thereby generating a modification within the genome of the cell for treating the subject, said contacting; (c) contacting the cell with a second guide RNA, wherein the second guide RNA induces the self-inactivating base editor to edit a splice acceptor or splice donor site present within the intron of the polynucleotide of (a), thereby generating a modification that reduces or eliminates the expression of the self-inactivating base editor, said contacting; The pharmaceutical composition comprising. **Claim 23** A base editor system according to any one of claims 8, 9, 11 for use in a method of treating a subject, the method comprising administering the base editor system to the subject, thereby treating the subject. The system. **Claim 24** The composition according to claim 21, wherein the napDNAbp domain is a Cas9 domain, optionally the N-terminal domain and the C-terminal domain of the Cas9 domain are split between amino acid residues Asn309 and Thr310, and further optionally, the Cas9 domain comprises the mutation Thr310Cys. **Claim 25** The composition according to any one of claims 18-22, 24, wherein the base editor comprises a nucleic acid programmable DNA binding protein (napDNAbp) domain and a deaminase domain, and the open reading frame comprising the intron is within the napDNAbp domain or the deaminase domain. **Claim 26** i) the intron is derived from a sequence selected from the group consisting of NF1, PAX2, EEF1A1, HBB, IGHG1, SLC50A1, ABCB11, BRSK2, PLXNB3, TMPRSS6, IL32, ANTXRL, PKHD1L1, PADI1, KRT6C, and HMCN2, and / or ii) the intron is as follows: a) GTGAGATCAAATGAAAGTTTCATATAGAAATACAAAACCTAGAGAACTGGCATGTAAGAGAAGCAAAAATTACTTCAGCAAGGCCATGTTAGTAAATTTGCATCTGTTTGTCCACATTAG (SEQ ID NO: 226), b) GTAGGTGACAATGCTGCAGCTGCCTAATCTAGGTGGGGGGAACTAAATTGTGGGTGAGCTGCTGAATGGTCTGTAGTCTGAGGCTGGGGTGGGGGGAGACACAACGTCCCCTCCCTGCAAACCACTGCTATTCTGTCCCTCTCTCTCCTTAG (SEQ ID NO: 227), c) GTAAGTGGCTTTCAAGACCATTGTTAAAAAGCTCTGGGAATGGCGATTTCATGCTTACATAAATTGGCATGCTTGTGTTTCAG (SEQ ID NO: 228), d) GTAAGTATCAAGGTTACAAGACAGGTTTAAGGAGACCAATAGAAACTGGGCTTGTCTAGACAGAGAAGACTCTTGCGTTTCTGATAGGCACCTATTGGTCTTACTGACATCCACTTTGCCTTTCTCTCCACAG (SEQ ID NO: 229), e) GTAAGCACAACTGGGATGGGGTGACAGGGGTGCAAGATTGAAAACTGGCTCCTCTCCTCATAGCAGTTCTTGTGATTTCAG (SEQ ID NO: 230), f) GTAAGAAATGTTATTTTTCAGTAAGTGATTTAGTTATTTTTCCTTTTTTCTCATTAAAATTTCTCTAACATCTCCCTCTTCATGTTTTAG (SEQ ID NO: 231), g) GTGAGACCCTAGCCCCCTCAACCCTGCCCTGGCCTCTCCCCAAACCTGCCCCCCCACGCTGACCCCCACACCCGGCCGCCCGCAG (SEQ ID NO: 232), h) GTGGGTGTCAGAGGCATCGGGGCTGCGGGGTAGGGGGCTGCCCCACCCCTAACGAAGTCTGCTCCTCCAG (SEQ ID NO: 233), i) GCAGGGAAGTCCTGCTTCCGTGCCCCACCGGTGCTCAGCTGAGGCTCCCTTGAAAATGCGAGGCTGTTTCCAACTTTGGTCTGTTTCCCTGGCAG (SEQ ID NO: 234), j) GTGGGGAGTTGGGGTCCCCGAAGGTGAGGACCCTCTGGGGATGAGGGTGCTTCTCTGAGACACTTTCTTTTCCTCACACCTGTTCCTCGCCAGCAG (SEQ ID NO: 235), k) GTATAGACCCCTTGATCTCCTAACCCTAACCCTAACCCTAACCCTAACCTACAAAATCTTAGAGCATCAGTGGGAGCATCTCACTGTCCAGGCTCAATATTTCTTCATTTTCTTGCAG (SEQ ID NO: 236), l) GTAATTATGATTAAAGATGGTGATTGTTTATTTTCTTTTATGATTGTCCTTAGTATTATGTAACCTGCAAATTCTATTGCAG (SEQ ID NO: 237), m) GTGAGTGACACAAGGTGTTGTCTGGGGAGTGGGGAAGGGGGATGGAAGTGAATCCTGTTGGTGGGGTGGAGAAAGGGCGATCTCAAGAGGGCCACTCTCTCCAG (SEQ ID NO: 238), n) GTAAGCATCTCCACCATCCTTCTGTTTACTCTGATGGGGTCTGCAAAGGGGAGATGATGTATAGGGTTGGGTATCTCTGTAAATGTCAGATGTGAAGTTGATCTTATGACCTTCTGTTCTGCAG (SEQ ID NO: 239), o) GTGAGGGTCTCCCAGGCTGGGCAGGGGGAGGGGGCTGCTGCCTTGATTGCGTCCCAGGACACAGCCCTCCTCCAGCCTGCCCTCGCCTTGCTCATCCCCTCCCCATCTCAGCCCCACCCCCACTAACTCTCTCTCTGCTCTGACTCAG (SEQ ID NO: 240), p) GTAATGATTGATTGCAATGTATGATTACAATAATCTCAGTATAAGTTCAGTAATAATAACCTTCCACTGCTGTCCTCTGTGTGCACCCAG (SEQ ID NO: 241), or q) GTAAATATATACAACAGTTTTTCATTTAAATAAGTGCACGGCACAAATAAGAAAAATATGTCAAAAATGTAACCAATAGTTTTTTTCAAATTTAG (SEQ ID NO: 242) comprising one of the sequences or a sequence having at least about 85%, 90%, 95%, or 99% nucleic acid sequence identity to one of the sequences; The composition according to any one of claims 18 to 22, 24.