Compositions and methods for genome editing of neonatal FC receptors
Patent Information
- Application Number
- JP2024522490
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-10-13
- Filing Date
- 2022-10-13
- Publication Date
- 2025-10-16
AI Technical Summary
There is a need for improved compositions and methods to target FcRn-mediated autoimmune disorders, as existing treatments like IVIg and high-affinity antibody injections have limitations in effectively reducing IgG levels and inflammatory responses.
The development of compositions and methods that modify the neonatal Fc receptor (FcRn) protein in mammalian cells using guide RNAs and genome editors to alter the amino acid sequence, reducing its ability to bind IgG, and the use of lipid nanoparticles to target and silence FcRn, thereby limiting IgG half-life in circulation.
These approaches specifically target FcRn binding to IgG without affecting albumin half-life, providing a more effective treatment for autoimmune disorders by reducing IgG-mediated immune responses.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Related Applications
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 255,290, filed October 13, 2021, the entire contents of which are incorporated herein by reference.
[0002] Sequence Listing
[0001] This application contains a Sequence Listing that has been submitted electronically in XML format, and is hereby incorporated by reference in its entirety. The Sequence Listing XML file, created on October 9, 2022, is named 180802_049101_PCT_SL.xml and is 970,484 bytes in size. [Technical Field]
[0003] The present disclosure relates to the field of genome editing. Specifically, the present disclosure relates to compositions and methods for editing, altering the expression of, and / or silencing the neonatal Fc receptor (FcRn) gene, FCGRT. [Background technology]
[0004] Immunoglobulin G (IgG) is the most common type of antibody found in the blood circulation and extracellular fluids where it is used to control infection in body tissues. While IgG can bind directly to antigens, the fragment crystallizable (Fc) region of IgG also binds to receptors on cells, resulting in an immune response. The Fc gamma receptor (FcγR) family includes the atypical neonatal Fc receptor (FcRn), encoded by the FCGRT gene. FcRn functions to recycle and maintain IgG and albumin and transport them across polarized cell barriers, thereby increasing their half-life in the circulation. FcRn also interacts with peptides derived from IgG immune complexes (ICs) and promotes their antigen presentation.
[0005] FcRn was first identified as the receptor responsible for the transport of maternal IgG antibodies from mother to child. Initially, FcRn was thought to be present only in the placenta and intestinal tissues during fetal and neonatal life. However, it is now known that FcRn is expressed in many tissues throughout the body, including epithelial, endothelial, and hematopoietic cells. Specifically, epithelial FcRn expression has been detected in the intestine, placenta, kidney, and liver.
[0006] Several autoimmune disorders, including myasthenia gravis, warm autoimmune hemolytic anemia (wAIHA), idiopathic thrombocytopenic purpura (ITP), Graves' disease, chronic inflammatory demyelinating polyneuropathy (CIDP), pemphigus vulgaris, and hemolytic disease of the fetus and newborn (HDFN), are caused by IgG responses to autoantigens. Because FcRn functions to maintain circulating IgG levels, it also prolongs the half-life of antibodies that cause these autoimmune disorders. Intravenous immunoglobulin (IVIg) is a recently developed therapy that saturates the IgG recycling capacity of FcRn, reducing the levels of pathogenic IgG that bind to FcRn and thereby promoting a reduction in the levels of IgG autoantibodies. Other strategies for treating autoimmune disorders include injection of higher-affinity antibodies to reduce the inflammatory response to autoantigens. Summary of the Invention [Problem to be solved by the invention]
[0007] There is a need for improved compositions and methods for the targeted treatment of FcRn-mediated autoimmune disorders. [Means for solving the problem]
[0008] Summary of the Invention Provided herein are compositions and methods for modifying the neonatal Fc receptor (FcRn) protein for IgG and / or its expression or activity in mammalian cells. The compositions and methods disclosed herein result in the production of modified variant FcRn proteins with reduced ability to bind to the Fc region of IgG antibodies. Such compositions and methods are useful for improving IgG-mediated autoimmune disorders. Advantageously, the compositions and methods disclosed herein specifically target FcRn binding to IgG without interfering with albumin half-life in subjects.
[0009] Thus, in one embodiment, there is provided a method of modifying an FcRn protein in a mammalian cell, the method comprising contacting the cell with a guide RNA and a genome editor, wherein the guide RNA comprises a nucleotide sequence complementary to a portion of the FCGRT gene and targets the genome editor to effect modification of the FCGRT gene in the cell, wherein the modification changes the amino acid sequence of the FcRn protein encoded by the FCGRT gene.
[0010] In another embodiment, a method of treating an IgG-mediated autoimmune disorder in a subject in need thereof is provided, the method comprising modifying FcRn protein in mammalian cells of the subject.
[0011] In another embodiment, a composition is provided comprising a guide RNA and a genome editor, wherein the guide RNA comprises a nucleotide sequence complementary to a portion of the FCGRT gene and targets the genome editor to effect a modification in the FCGRT gene in a cell, wherein the modification changes the amino acid sequence of the FcRn protein encoded by the FCGRT gene.
[0012] In another embodiment, lipid nanoparticles (LNPs) are provided that are surface-functionalized to incorporate the Fc fragment of an IgG antibody or other targeting moiety. The disclosed LNPs can target the neonatal Fc receptor (FcRn) on epithelial surfaces, fuse with or internalize, and deliver their payload to targeted cells. The LNPs disclosed herein may also contain siRNA to silence FcRn, thereby limiting the half-life of IgG in circulation and treating IgG-mediated autoimmune disorders in subjects in need thereof.
[0013] In another embodiment, an LNP is provided that comprises a lipid monolayer membrane comprising at least one fragment crystallizable (Fc) region of an IgG antibody, or a functional fragment thereof, embedded in the lipid monolayer membrane; and a lipid core matrix encapsulated in the lipid monolayer membrane.
[0014] In another embodiment, an LNP is provided that comprises a lipid monolayer membrane comprising at least one fragment, Fc region, of an IgG antibody or a functional fragment thereof embedded in the lipid monolayer membrane; and a lipid core matrix encapsulated in the lipid monolayer membrane, wherein the lipid core matrix comprises at least one nucleic acid.
[0015] In another embodiment, an LNP is provided comprising: a lipid monolayer membrane comprising at least one Fc region of an IgG antibody or a functional fragment thereof embedded in the lipid monolayer membrane; and a lipid core matrix encapsulated in the lipid monolayer membrane, wherein the lipid core matrix comprises at least one siRNA or guide RNA that regulates expression of or silences the FCGRT gene.
[0016] In another embodiment, a pharmaceutical composition is provided comprising: at least one LNP comprising a lipid monolayer membrane comprising at least one Fc region of an IgG antibody or a functional fragment thereof embedded in the lipid monolayer membrane; and a lipid core matrix encapsulated in the lipid monolayer membrane, wherein the lipid core matrix comprises at least one nucleic acid; and at least one pharmaceutically acceptable excipient.
[0017] In another embodiment, there is provided a method of treating an IgG-mediated autoimmune disorder in a subject in need thereof, the method comprising administering to the subject LNPs comprising: a lipid monolayer membrane comprising at least one Fc region of an IgG antibody or other targeting moiety disclosed herein, or a functional fragment thereof, embedded in the lipid monolayer membrane; and a lipid core matrix encapsulated in the lipid monolayer membrane, wherein the lipid core matrix comprises at least one siRNA or guide RNA that regulates expression of or silences the FCGRT gene.
[0018] In another embodiment, a method for silencing FcRn expression in a cell is provided, the method comprising contacting the cell with an LNP comprising: a lipid monolayer membrane comprising at least one Fc region of an IgG antibody, or a functional fragment thereof, embedded in the lipid monolayer membrane; and a lipid core matrix encapsulated in the lipid monolayer membrane, wherein the lipid core matrix comprises at least one siRNA that silences the FCGRT gene. In one aspect, the disclosure features a method for altering a nucleobase of an IgG receptor and transporter Fc fragment (FcRn) polynucleotide. The method comprises contacting the FcRn polynucleotide with a base editor system comprising one or more guide polynucleotides and a base editor, or one or more polynucleotides encoding the base editor system, thereby altering the nucleobase of the FcRn polynucleotide. The base editor comprises a nucleic acid-programmable DNA-binding protein (napDNAbp) domain and a deaminase domain. The base editor system comprises: (a) one or more guide polynucleotides comprise a nucleic acid sequence comprising at least 10 to 23 contiguous nucleotides of a spacer nucleic acid sequence listed in Table 2B; or (b) the one or more guide polynucleotides comprise the following reference sequences: FcRn amino acid sequence AESHLSLLYHLTAVSSPAPGTPAFWVSGWLGPQQYLSYNSLRGEAEPCGAWVWENQVSWYWEKETTDLRIKEKLFLEAFKALGGKGPYTLQGLLGCELGPDNTSVPTAKFALNGEEFMNFDLKQGTWGGDWPEALAISQRWQQQDKAANKELTFLLFSCPHRLREHLERGRGNLEWKEPPSMRLKARPSSPGFSVLTCSAFSFYPPELQLRFLRNGLAAGTGQGDFGPNSDGSFHASSSLTVKSGDEHHYCCIVQHAGLAQPLRVELESPAKSSVLVVGIVIGVLLLTAAAVGGALLWRRMRSGLPAPWISLRGDDTGVLLPTPGEAQDADLKDVNVIPATA (SEQ ID NO: 530), or to the corresponding position in another FcRn polypeptide sequence Base editors are targeted to effect nucleobase changes in codons encoding amino acid residues selected from one or more of F110, L112, N113, E115, E116, F117, M118, N119, D121, L122, T126, W127, G128, D130, W131, P132, E133, A134, L135, and I137.
[0019] In another aspect, the disclosure features a cell produced by a method of any of the aspects of the disclosure, or embodiments thereof.
[0020] In another aspect, the disclosure features a base editor system for altering nucleobases in IgG receptor and transporter Fc fragment (FcRn) polynucleotides. The base editor system includes (i) one or more guide polynucleotides, or one or more polynucleotides encoding the one or more guide polynucleotides, and (ii) a base editor comprising a nucleic acid programmable DNA binding protein (napDNAbp) domain and a deaminase domain, or one or more polynucleotides encoding the base editor. The base editor system includes: (a) one or more guide polynucleotides comprise a nucleic acid sequence comprising at least 10 to 23 contiguous nucleotides of a spacer nucleic acid sequence listed in Table 2B; or (b) the one or more guide polynucleotides comprise the following reference sequences: FcRn amino acid sequence AESHLSLLYHLTAVSSPAPGTPAFWVSGWLGPQQYLSYNSLRGEAEPCGAWVWENQVSWYWEKETTDLRIKEKLFLEAFKALGGKGPYTLQGLLGCELGPDNTSVPTAKFALNGEEFMNFDLKQGTWGGDWPEALAISQRWQQQDKAANKELTFLLFSCPHRLREHLERGRGNLEWKEPPSMRLKARPSSPGFSVLTCSAFSFYPPELQLRFLRNGLAAGTGQGDFGPNSDGSFHASSSLTVKSGDEHHYCCIVQHAGLAQPLRVELESPAKSSVLVVGIVIGVLLLTAAAVGGALLWRRMRSGLPAPWISLRGDDTGVLLPTPGEAQDADLKDVNVIPATA (SEQ ID NO: 530), or to the corresponding position in another FcRn polypeptide sequence Base editors are targeted to effect nucleobase changes in codons encoding amino acid residues selected from one or more of F110, L112, N113, E115, E116, F117, M118, N119, D121, L122, T126, W127, G128, D130, W131, P132, E133, A134, L135, and I137.
[0021] In another aspect, the disclosure features a polynucleotide encoding a base editor system of any of the aspects of the disclosure, or embodiments thereof.
[0022] In another aspect, the disclosure features a vector including a polynucleotide of any of the aspects of the disclosure, or embodiments thereof.
[0023] In another aspect, the disclosure features a cell including a polynucleotide or vector of any of the aspects of the disclosure, or embodiments thereof.
[0024] In another aspect, the disclosure features a composition including a base editor system, a polynucleotide, a vector, or a cell of any of the aspects of the disclosure, or embodiments thereof.
[0025] In another aspect, the disclosure features a pharmaceutical composition including a composition of any of the aspects of the disclosure, or embodiments thereof, and a pharmaceutically acceptable excipient.
[0026] In another aspect, the disclosure features a method of treating an immunoglobulin G-mediated autoimmune disorder in a subject in need thereof. The method includes altering a nucleobase of an FcRn polynucleotide in the subject by administering to the subject a base editor system, or one or more polynucleotides encoding the base editor system, thereby treating the autoimmune disorder. The base editor system includes one or more guide polynucleotides and a base editor. The base editor includes a nucleic acid-programmable DNA-binding protein (napDNAbp) domain and a deaminase domain. The base editor system includes: (a) one or more guide polynucleotides comprise a nucleic acid sequence comprising at least 10 to 23 contiguous nucleotides of a spacer nucleic acid sequence listed in Table 2B; or (b) the one or more guide polynucleotides comprise the following reference sequences: FcRn amino acid sequence AESHLSLLYHLTAVSSPAPGTPAFWVSGWLGPQQYLSYNSLRGEAEPCGAWVWENQVSWYWEKETTDLRIKEKLFLEAFKALGGKGPYTLQGLLGCELGPDNTSVPTAKFALNGEEFMNFDLKQGTWGGDWPEALAISQRWQQQDKAANKELTFLLFSCPHRLREHLERGRGNLEWKEPPSMRLKARPSSPGFSVLTCSAFSFYPPELQLRFLRNGLAAGTGQGDFGPNSDGSFHASSSLTVKSGDEHHYCCIVQHAGLAQPLRVELESPAKSSVLVVGIVIGVLLLTAAAVGGALLWRRMRSGLPAPWISLRGDDTGVLLPTPGEAQDADLKDVNVIPATA (SEQ ID NO: 436), or to the corresponding position in another FcRn polypeptide sequence The base editor is targeted to effect a nucleobase change in a codon encoding an amino acid residue selected from one or more of F110, L112, N113, E115, E116, F117, M118, N119, D121, L122, T126, W127, G128, D130, W131, P132, E133, A134, L135, and I137.
[0027] In another aspect, the disclosure features a kit suitable for use in a method of any of the aspects of the disclosure, or embodiments thereof, that includes a guide polynucleotide comprising a sequence listed in Table 2A or Table 2B.
[0028] In another aspect, the disclosure features a method for altering a nucleobase of an IgG receptor and transporter Fc fragment (FcRn) polynucleotide. The method includes contacting the FcRn polynucleotide with a base editor system, thereby altering the nucleobase of the FcRn polynucleotide. The base editor system includes one or more guide polynucleotides selected from one or more of gRNA1583, gRNA1578, and gRNA3265, or one or more polynucleotides encoding them, and a base editor comprising a nucleic acid programmable DNA binding protein (napDNAbp) domain and an adenosine deaminase domain, or one or more polynucleotides encoding the base editor.
[0029] In another aspect, the disclosure features a base editor system including one or more guide polynucleotides, or one or more polynucleotides encoding them, selected from one or more of gRNA1583, gRNA1578, gRNA3265, and a base editor comprising a nucleic acid programmable DNA binding protein (napDNAbp) domain and an adenosine deaminase domain, or one or more polynucleotides encoding the base editor.
[0030] In another aspect, the disclosure features a guide polynucleotide comprising a sequence listed in Table 2A or Table 2B.
[0031] In aspects of the disclosure, or any of its embodiments, the nucleobase changes are the following amino acid changes in the FcRn polypeptide encoded by the FcRn polynucleotide relative to the reference sequence: F110L, F110S, F110P, L112P, N113S, N113D, E115G, E115K, E116G, E116K, E116Q, F117P, M118N, M118V, M118I, M118N, M118V, M118I, M118N, M118V, M118I, M118N, M118V, M118V, M118I, M118V ... In aspects of the disclosure, or any of its embodiments, the one or more guide polynucleotides target a base editor to effect a nucleobase change in a codon encoding amino acid M118 or W131 in a reference sequence. In this aspect of the disclosure, or any of its embodiments, the nucleobase change results in an amino acid change in the FcRn polypeptide encoded by the FcRn polynucleotide selected from one or more of M118V, M118V, M118I, M118T, W131R, and W131Q.
[0032] In an aspect of the present disclosure, or any of its embodiments, one or more amino acid changes in an FcRn polypeptide reduce or eliminate binding of the FcRn polypeptide to IgG1, IgG2, IgG3, and / or IgG4. In an aspect of the present disclosure, or any of its embodiments, one or more amino acid changes in an FcRn polypeptide reduce or eliminate binding of the FcRn polypeptide to the Fc region of IgG1, IgG2, IgG3, and / or IgG4. In an aspect of the present disclosure, or any of its embodiments, an FcRn polypeptide comprising one or more amino acid changes has a K in solution for binding to IgG1, IgG2, IgG3, and / or IgG4 of greater than 3000 nM. D It has.
[0033] In an aspect of the present disclosure, or any of its embodiments, an FcRn polypeptide encoded by an FcRn polynucleotide comprising an altered nucleobase is capable of binding to albumin. In an aspect of the present disclosure, or any of its embodiments, an FcRn polypeptide comprising one or more amino acid alterations has a K of less than 2000 nM in solution for binding to albumin. D In an aspect of the present disclosure, or any of its embodiments, the FcRn polypeptide comprising one or more amino acid changes has a K of less than 1000 nM in solution for binding to albumin. D In an aspect of the present disclosure, or any of its embodiments, the binding of an FcRn polypeptide comprising one or more amino acid changes has a K of less than 500 nM in solution for binding to albumin. D It has.
[0034] In an aspect of the present disclosure, or any of its embodiments, an FcRn polypeptide comprising one or more amino acid changes exhibits a K value in solution for albumin binding that is lower than that of a reference FcRn polypeptide having the same amino acid sequence except that it does not comprise the one or more amino acid changes. D K less than 1.5 times D In aspects of the present disclosure, or any of its embodiments, an FcRn polypeptide comprising one or more amino acid changes may have a K in solution for binding to albumin that is lower than the K of a reference FcRn polypeptide having the same amino acid sequence except that it does not comprise the one or more amino acid changes. D Between 0.5 and 1.5 times K D In aspects of the present disclosure, or any of its embodiments, an FcRn polypeptide comprising one or more amino acid changes may have a K in solution for binding to IgG1, IgG2, IgG3, and / or IgG4 of a reference FcRn polypeptide having the same amino acid sequence except that it does not comprise the one or more amino acid changes. D At least 5 times the K DIn aspects of the present disclosure, or any of its embodiments, an FcRn polypeptide comprising one or more amino acid changes may have a K in solution for binding to IgG1, IgG2, IgG3, and / or IgG4 of a reference FcRn polypeptide having the same amino acid sequence except that it does not comprise the one or more amino acid changes. D At least 10 times the K D In any of the aspects of the present disclosure, or embodiments thereof, the FcRn polypeptide comprising one or more amino acid changes does not bind to IgG1, IgG2, IgG3, and / or IgG4 at detectable levels, as measured in a suitable assay, e.g., an SPR assay described herein.
[0035] In aspects of the present disclosure, or any of its embodiments, the nucleobases of the FcRn polynucleotide are altered with a base editing efficiency of at least about 20%. In aspects of the present disclosure, or any of its embodiments, the nucleobases of the FcRn polynucleotide are altered with a base editing efficiency of at least about 40%. In aspects of the present disclosure, or any of its embodiments, the nucleobases of the FcRn polynucleotide are altered with a base editing efficiency of at least about 50%.
[0036] In an aspect of the present disclosure, or any of its embodiments, the deaminase domain is capable of deaminating cytidine or adenine in DNA. In an aspect of the present disclosure, or any of its embodiments, the deaminase domain is an adenosine deaminase domain or a cytidine deaminase domain. In an aspect of the present disclosure, or any of its embodiments, the adenosine deaminase converts a target A·T to G·C in an FcRn polynucleotide. In an aspect of the present disclosure, or any of its embodiments, the cytidine deaminase converts a target C·G to T·A in an FcRn polynucleotide. In an aspect of the present disclosure, or any of its embodiments, the cytidine deaminase domain is an APOBEC deaminase domain or a derivative thereof.
[0037] In aspects of the disclosure, or any of its embodiments, the base editor is a BE4 base editor.
[0038] In an aspect of the present disclosure, or any of its embodiments, the adenosine deaminase domain is a TadA deaminase domain. In an aspect of the present disclosure, or any of its embodiments, the deaminase domain is an adenosine deaminase domain. In an aspect of the present disclosure, or any of its embodiments, the adenosine deaminase is a TadA * 8 or Tad * In an aspect of the present disclosure, or any of its embodiments, the adenosine deaminase is TadA. * 8.1, TadA * 8.2, TadA * 8.3, TadA * 8.4, TadA * 8.5, TadA * 8.6, TadA * 8.7, TadA * 8.8, TadA * 8.9, TadA * 8.10, TadA * 8.11, TadA * 8.12, TadA * 8.13, TadA * 8.14, TadA * 8.15, TadA * 8.16, TadA * 8.17, TadA * 8.18, TadA * 8.19, TadA * 8.20, TadA * 8.21, TadA * 8.22, TadA * 8.23, or TadA * It's 8.24.
[0039] In any aspect of the disclosure, or embodiment thereof, the deaminase domain is a monomer or a heterodimer.
[0040] In aspects of the present disclosure, or any of its embodiments, the napDNAbp domain is Cas9 or Cas12. In aspects of the present disclosure, or any of its embodiments, the napDNAbp domain is a nuclease-inactive or nickase variant. In aspects of the present disclosure, or any of its embodiments, the napDNAbp domain comprises a Cas9, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ polynucleotide, or a functional portion thereof. In aspects of the present disclosure, or any of its embodiments, the napDNAbp domain comprises dead Cas9 (dCas9) or Cas9 nickase (nCas9). In aspects of the present disclosure, or any of its embodiments, the napDNAbp domain is Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1Cas9 (St1Cas9), Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof.
[0041] In some aspects of the present disclosure, or any of its embodiments, the napDNAbp domain comprises a variant of SpCas9 or SaCas9 with altered protospacer adjacent motif (PAM) specificity. In some aspects of the present disclosure, or any of its embodiments, the SpCas9 or SaCas9 has specificity for PAM sequences selected from one or more of NGG, NGA, NGC, NNGRRT, and NNNRRT, where N is any nucleotide and R is A or G.
[0042] In aspects of the disclosure, or any of its embodiments, the napDNAbp domain comprises a nuclease-active Cas9.
[0043] In any of the aspects of the disclosure, or embodiments thereof, the base editor further comprises one or more uracil glycosylase inhibitors (UGIs), or the method further comprises expressing a UGI in the cell in trans with the base editor.
[0044] In aspects of the present disclosure, or any of its embodiments, the base editor further comprises one or more nuclear localization signals (NLSs). In aspects of the present disclosure, or any of its embodiments, the NLS is a bisecting NLS.
[0045] In aspects of the present disclosure, or any of its embodiments, the one or more guide polynucleotides comprise a scaffold comprising one of the following nucleotide sequences: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (SpCas9 scaffold; SEQ ID NO: 317), or GUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGAUUUU (SaCas9 scaffold; SEQ ID NO: 436). In an aspect of the present disclosure, or any of its embodiments, the one or more guide polynucleotides comprise one or more modified nucleotides. In an aspect of the present disclosure, or any of its embodiments, the one or more modified polynucleotides are at the 5' and / or 3' ends of the one or more guide polynucleotides. In an aspect of the present disclosure, or any of its embodiments, the one or more modified nucleotides are 2'-O-methyl-3'-phosphorothioate nucleotides. In an aspect of the present disclosure, or any of its embodiments, the one or more guide polynucleotides comprise a spacer containing only 19 to 23 nucleotides. In an aspect of the present disclosure, or any of its embodiments, the one or more guide polynucleotides comprise a spacer containing only 19 or 20 nucleotides.
[0046] In aspects of the disclosure, or any of the embodiments thereof, the base editor comprises a complex comprising a deaminase domain, a napDNAbp domain, and a guide polynucleotide, or the base editor is a fusion protein comprising a napDNAbp domain fused to a deaminase domain.
[0047] In an aspect of the present disclosure, or any of its embodiments, the FcRn polynucleotide is in a cell. In an aspect of the present disclosure, or any of its embodiments, the cell is a hepatocyte, an endothelial cell, a myeloid cell, or an epithelial cell. In an aspect of the present disclosure, or any of its embodiments, the cell is in vivo or ex vivo. In an aspect of the present disclosure, or any of its embodiments, the cell is in a subject. In an aspect of the present disclosure, or any of its embodiments, the cell is a mammalian cell. In an aspect of the present disclosure, or any of its embodiments, the cell is a human cell.
[0048] In this aspect of the disclosure, or any of its embodiments, the subject is a mammal. In this aspect of the disclosure, or any of its embodiments, the mammal is a human.
[0049] In any of the aspects of the disclosure, or embodiments thereof, the base editor further comprises one or more uracil glycosylase inhibitors (UGIs), or the base editor system further comprises a UGI in trans with the base editor.
[0050] In one aspect of the present disclosure, or in one of its embodiments, the vector comprises lipid nanoparticles.In one aspect of the present disclosure, or in one of its embodiments, the lipid nanoparticles comprise a lipid monolayer comprising lipids selected from one or more of lecithin, phosphatidylcholine, phosphatidic acid, phosphatidylethanolamine, phosphatidylglycerol, phosphatidylserine, phosphatidylinositol, cardiolipin, lipid-polyethylene glycol conjugates, and combinations thereof.In one aspect of the present disclosure, or in one of its embodiments, the lipid monolayer comprises PEGylated lipids.In one aspect of the present disclosure, or in one of its embodiments, the lipid monolayer further comprises cholesterol. In aspects of the present disclosure, or any of its embodiments, the lipid nanoparticles may be selected from the group consisting of N-methyl-N-(2-(arginoylamino)ethyl)-N,N-dioctadecylaminium chloride or distearoylarginylammonium chloride] (DSAA); N,N-dimyristoyl-N-methyl-N-2[N'-(N6-guanidino-L-lysinyl)]aminoethylammonium chloride (DMGLA); N,N-dimyristoyl-N-methyl-N-2[N2-guanidino-L-lysinyl]aminoethylammonium chloride; N,N-dimyristoyl-N-methyl-N-2[N'-(N2,N6-di-guanidino-L-lysinyl)]aminoethylammonium chloride; N,N-di-stearoyl-N-methyl-N-2[N '-(N6-guanidino-L-lysinyl)]aminoethyl ammonium chloride; N,N-dioleyl-N,N-dimethylammonium chloride (DODAC); N-(2,3-dioleoyloxy)propyl)-N,N,N-trimethylammonium chloride (DOTAP); N-(2,3-dioleyloxy)propyl)-N,N,N-trimethylammonium chloride (DOTMA); N,N-distearyl-N,N-dimethylammonium bromide (DDAB); 3-(N-(N',N'-dimethylaminoethane)-carbamoyl)cholesterol (DC-Choi); N-(l,2-dimyristyloxyprop-3-yl)-N,N-dimethyl-N-hydroxyethylammonium bromide (DMRIE);l,3-Dioleoyl-3-trimethylammonium-propane, N-(l-(2,3-dioleyloxy)propyl)-N-(2-(sperminecarboxamido)ethyl)-N,N-dimethyl-1 ammonium trifluoroacetate (DOSPA); GAP-DLRIE; DMDHP; 3-p[4N-(H8N-diguanidinospermidine)-carbamoyl]cholesterol (BGSC); 3-p[N,N-diguanidinoethyl-aminoethane)-carbamoyl]cholesterol (BGTC); N,N\N2,N3 tetraoleoyl -methyltetrapalmitylspermine (Cellfectin); Nt-butyl-N'-tetradecyl-3-tetradecyl-aminopropionamidine (CLONfectin); dimethyldioctadecylammonium bromide (DDAB); l,3-dioleoyloxy-2-(6-carboxyspermyl)-propylamide (DOSPER); 4-(2,3-bis-palmitoyloxy-propyl)-1-methyl-lH-imidazole (DPIM) N,N,N',N'-tetramethyl-N,N'-bis(2-hydroxyethyl)-2 ,3-Dioleoyloxy-1,4-butanediammonium iodide) (Tfx-50); 1,2-Dioleoyl-3-(4'-trimethylammonio)butanol-sn-glycerol (DOBT); cholesteryl (4'-trimethylammonium) butanoate (ChOTB) in which the trimethylammonium group is attached to either the duplex (in the case of DOTB) or the cholesteryl group (in the case of ChOTB) via a butanol spacer arm; DL-l,2-Dioleoyl-3-dimethylaminopropyl-p-hydroxyethyl ammonium (DORI); DL-l,2-O-dioleoyl-3-dimethylaminopropyl-P-hydroxyethylammonium (DORIE); l,2-dioleoyl-3-succinyl-sn-glycerol choline ester (DOSC); cholesteryl hemisuccinate ester (ChOSC); dioctadecylamidoglycylspermine (DOGS); dipalmitoylphosphatidylethanolamylspermine (DPPES); cholesteryl-3P-carboxyl-amido-ethylenetrimethylammonium iodide;l-Dimethylamino-3-trimethylammonio-DL-2-propyl-cholesterylcarboxylate iodide;Cholesteryl-3-β-carboxyamidoethyleneamine;Cholesteryl-3-P-oxysuccinamido-ethylenetrimethylammonium iodide;l-Dimethylamino-3-trimethylammonio-DL-2-propyl-cholesteryl-3-P-oxysuccinamido-ethylenetrimethylammonium iodide;2-(2-Trimethylammonio)-ethylmethylaminoethyl-cholesteryl-3-P- The ionizable cationic lipid comprises one or more of: oxysuccinate iodide; 3-β-N-(polyethyleneimine)-carbamoylcholesterol, DC-cholesterol; N4-cholesteryl-spermine HCl salt (GL67); N1-[2-((1S)-1-[(3-aminopropyl)amino]-4-[di(3-amino-propyl)amino]butylcarboxamido)ethyl]-3,4-di[oleyloxy]-benzamide (MVL5); and combinations thereof.
[0051] In an aspect of the present disclosure, or any of its embodiments, the vector comprises a polymeric nanoparticle. In an aspect of the present disclosure, or any of its embodiments, the vector is a viral vector. In an aspect of the present disclosure, or any of its embodiments, the viral vector is a retroviral vector or an adeno-associated viral vector.
[0052] In aspects of the present disclosure, or any of its embodiments, the disorder is selected from one or more of myasthenia gravis (gMG), warm autoimmune hemolytic anemia (wAIHA), idiopathic thrombocytopenic purpura (ITP), Graves' disease, chronic inflammatory demyelinating polyneuropathy (CIDP), pemphigus vulgaris, and hemolytic disease of the fetus and newborn (HDFN).
[0053] In any of the aspects of the disclosure, or embodiments thereof, the base editor further comprises one or more uracil glycosylase inhibitors (UGIs), or the method further comprises expressing a UGI in the cell in trans with the base editor.
[0054] In this aspect of the disclosure, or any of its embodiments, the administration is local. In this aspect of the disclosure, or any of its embodiments, the administration is systemic.
[0055] In this aspect of the disclosure, or any of its embodiments, the base editor system is administered to a subject using a vector.
[0056] In an aspect of the present disclosure, or any of its embodiments, the vector is a lipid nanoparticle.In an aspect of the present disclosure, or any of its embodiments, the vector is targeted to the liver.
[0057] In this aspect of the disclosure, or any of its embodiments, the subject is a mammal. In this aspect of the disclosure, or any of its embodiments, the mammal is a human.
[0058] These and other objects, features, embodiments and advantages will become apparent to those skilled in the art upon reading the following detailed description and the appended claims.
[0059] definition Although the following terms are believed to be well understood in the art, definitions are provided to facilitate description of the presently disclosed subject matter. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the presently disclosed subject matter belongs.
[0060] The following references provide those skilled in the art with general definitions of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed., 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them unless otherwise specified.
[0061] Unless otherwise indicated, all numerical values expressing properties such as quantities of ingredients, reaction conditions, and the like used in the specification and claims are to be understood as being modified in all instances by the term "about." Accordingly, unless indicated to the contrary, the numerical parameters set forth in the specification and claims are approximations that may vary depending upon the desired properties sought to be obtained by the presently disclosed subject matter.
[0062] It should be understood that every maximum numerical limitation given throughout this specification includes every lower numerical limitation, as if such lower numerical limitations were expressly written herein. Every minimum numerical limitation given throughout this specification includes every higher numerical limitation, as if such higher numerical limitations were expressly written herein. Every numerical range given throughout this specification includes every narrower numerical range that is included in such broader numerical range, as if such narrower numerical ranges were all expressly written herein. "Adenine" or "9H-purin-6-amine" refers to a compound having the molecular formula CHN and the structure
[0063] [ka]
[0064] and refers to the purine nucleobase having the CAS number 73-24-5.
[0065] "Adenosine" or "4-amino-1-[(2R,3R,4S,5R)-3,4-dihydroxy-5-(hydroxymethyl)oxolan-2-yl]pyrimidin-2(1H)-one" means the compound of the structure
[0066] [ka]
[0067] and corresponds to the CAS number 65-46-3. Its molecular formula is C 10 H 13 It is N5O4.
[0068] "Adenosine deaminase" or "adenine deaminase" refers to a polypeptide or fragment thereof that can catalyze the hydrolytic deamination of adenine or adenosine. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine to inosine or deoxyadenosine to deoxyinosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) can be derived from any organism (e.g., eukaryotes, prokaryotes), including, but not limited to, algae, bacteria, fungi, plants, invertebrates (e.g., insects), and vertebrates (e.g., amphibians, mammals). In some embodiments, the adenosine deaminase is an adenosine deaminase variant having one or more changes and capable of deaminating both adenine and cytosine in a target polynucleotide (e.g., DNA, RNA), and may be referred to as a "dual deaminase." Non-limiting examples of dual deaminases include those described in PCT / US22 / 22050. In some embodiments, the target polynucleotide is single-stranded or double-stranded. In some embodiments, the adenosine deaminase variant is capable of deaminating both adenine and cytosine in DNA. In some embodiments, the adenosine deaminase variant is capable of deaminating both adenine and cytosine in single-stranded DNA. In some embodiments, the adenosine deaminase variant is capable of deaminating both adenine and cytosine in RNA. In embodiments, the adenosine deaminase variant is selected from those described in PCT / US2020 / 018192, PCT / US2020 / 049975, and PCT / US2017 / 045381.
[0069] "Adenosine deaminase activity" means catalyzing the deamination of adenine or adenosine to guanine in a polynucleotide. In some embodiments, the adenosine deaminase variants provided herein exhibit adenosine deaminase activity (e.g., a reference adenosine deaminase (e.g., TadA) * 8.20 or TadA * 8.19) maintains at least about 30%, 40%, 50%, 60%, 70%, 80%, 90%, or more) of the activity of
[0070] "Adenosine base editor (ABE)" means a base editor that includes an adenosine deaminase.
[0071] By "adenosine base editor (ABE) polynucleotide" is meant a polynucleotide that encodes an ABE.
[0072] By "Adenosine Base Editor 8 (ABE8) polypeptide" or "ABE8" is meant a base editor, as defined herein, including an adenosine deaminase or adenosine deaminase variant, that comprises one or more of the changes listed in Table 15, one of the combinations of changes listed in Table 15, or a change at one or more of the amino acid positions listed in Table 15, wherein such changes are consistent with the following reference sequence: ABE8 is related to the corresponding position in another adenosine deaminase, such as MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 1). In embodiments, ABE8 contains changes at amino acids 82 and / or 166 of SEQ ID NO: 1. In some embodiments, ABE8 contains additional changes relative to the reference sequence, as described herein.
[0073] By "adenosine base editor 8 (ABE8) polynucleotide" is meant a polynucleotide that encodes an ABE8 polypeptide.
[0074] "Administering," as used herein, refers to providing one or more compositions described herein to a patient or subject. By way of example and not limitation, administration (e.g., injection) of a composition can be performed by intravenous (iv), subcutaneous (sc), intradermal (id), intraperitoneal (ip), or intramuscular (im) injection. One or more such routes can be utilized. Parenteral administration can be, for example, by bolus injection or by gradual perfusion over time. In some embodiments, parenteral administration includes intravascular, intravenous, intramuscular, intraarterial, intrathecal, intratumoral, intradermal, intraperitoneal, transtracheal, subcutaneous, subcuticular, intraarticular, subcapsular, subarachnoid, and intrasternal infusion or injection. Alternatively, or concurrently, administration can be by oral route. In embodiments, one or more compositions described herein are administered by subretinal or subfoveal injection. In some cases, subretinal injection results in the formation of a bleb in the fovea.
[0075] By "agent" is meant any small molecule chemical compound, antibody, nucleic acid molecule, or polypeptide, or fragments thereof.
[0076] By "albumin polypeptide" is meant a protein having at least about 85% amino acid sequence identity with GenBank Accession No. CAA23754.1, provided below, or a fragment thereof capable of binding to an FcRn polypeptide. >CAA23754.1 Serum Albumin [Homo sapiens] (SEQ ID NO: 425).
[0077] By "albumin polynucleotide" is meant a nucleic acid molecule encoding an albumin polypeptide, as well as introns, exons, 3' untranslated regions, 5' untranslated regions, and regulatory sequences associated with its expression, or fragments thereof. In embodiments, an albumin polynucleotide is a genomic sequence, cDNA, mRNA, or gene associated with and / or required for albumin expression. An exemplary albumin nucleotide sequence from Homo sapiens is provided below (GenBank: V00495.1:76-1905):
[0078] "Alteration" or "modification" refers to a change in the expression level, structure, or activity of an analyte, gene, or polypeptide, as detected by standard art-known methods, such as those described herein. As used herein, alteration includes a change (e.g., an increase or decrease) in expression levels. In embodiments, the increase or decrease in expression levels is 10%, 25%, 40%, 50%, or more. In some embodiments, the alteration includes an insertion, deletion, or substitution of a nucleic acid base or amino acid (e.g., by genetic engineering).
[0079] By "ameliorate" is meant to reduce, inhibit, attenuate, alleviate, arrest, or stabilize the onset or progression of a disease.
[0080] "Analog" refers to a molecule that is not identical but has similar functional or structural characteristics. For example, a polypeptide analog retains the biological activity of the corresponding naturally occurring polypeptide while having certain biochemical modifications that enhance the function of the analog compared to the naturally occurring polypeptide. Such biochemical modifications can, for example, increase the analog's protease resistance, membrane permeability, or half-life without altering ligand binding. Analogs can also include unnatural amino acids.
[0081] As used herein, the term "antibody" refers to an immunoglobulin molecule that specifically binds to or immunologically reacts with a particular antigen, and includes, but is not limited to, chimeric antibodies, humanized antibodies, heteroconjugate antibodies (e.g., bi-, tri-, and tetra-specific antibodies, diabodies, triabodies, and tetrabodies), and polyclonal, monoclonal, genetically engineered, and other modified forms of antibodies, including antigen-binding fragments of antibodies, including, for example, Fab', F(ab')2, Fab, Fv, rgG, and scFv fragments. Unless otherwise indicated, the term "monoclonal antibody" (mAb) is meant to include both intact molecules and antibody fragments (e.g., including Fab and F(ab')2 fragments) that are capable of specifically binding to a target protein. As used herein, Fab and F(ab')2 fragments refer to antibody fragments that lack the Fc fragment of the intact antibody.
[0082] An antibody (immunoglobulin) contains two heavy chains linked together by disulfide bonds and two light chains, with each light chain linked to its respective heavy chain by a disulfide bond in a "Y"-shaped configuration. Each heavy chain has a variable domain (VH) at one end followed by multiple constant domains (CH). Each light chain has a variable domain (VL) at one end and a constant domain (CL) at the other end. The variable domain (VL) of the light chain is aligned with the variable domain (VL) of the heavy chain, and the light chain constant domain (CL) is aligned with the first constant domain (CH1) of the heavy chain. The variable domains of each pair of light and heavy chains form the antigen-binding site. The heavy chain isotype (gamma, alpha, delta, epsilon, or mu) determines the immunoglobulin class (IgG, IgA, IgD, IgE, or IgM, respectively). Light chains are of one of two isotypes, kappa (κ) or lambda (λ), found in all antibody classes. The term "antibody" or "antibody(s)" includes intact antibodies, such as polyclonal or monoclonal antibodies (mAbs), as well as proteolytic portions or fragments thereof, such as Fab or F(ab')2 fragments, that are capable of specific binding to a target protein. Antibodies can include chimeric antibodies; recombinant and engineered antibodies, and antigen-binding fragments thereof.Exemplary functional antibody fragments containing the entire or essentially the entire variable regions of both the light and heavy chains are defined as follows: (i) Fv, defined as a genetically engineered fragment consisting of the variable region of the light chain and the variable region of the heavy chain represented as two chains; (ii) single-chain Fv ("scFv"), which is a genetically engineered single-chain molecule comprising the variable region of the light chain and the variable region of the heavy chain linked by a suitable polypeptide linker; (iii) an intact antibody is treated with the enzyme papain to produce an Fd fragment of the heavy chain consisting of an intact light chain and the variable and CH1 domains of the heavy chain. (iv) Fab', the fragment of an antibody molecule containing a monovalent antigen-binding portion of an antibody molecule, can be obtained by treating an intact antibody with the enzyme pepsin, followed by reduction (to produce two Fab' fragments per antibody molecule); and (v) F(ab')2, the fragment of an antibody molecule containing a monovalent antigen-binding portion of an antibody molecule, can be obtained by treating an intact antibody with the enzyme pepsin (i.e., a dimer of Fab' fragments held together by two disulfide bonds).
[0083] "Base editor (BE)" or "nucleobase editor polypeptide (NBE)" refers to an agent that binds to a polynucleotide and has nucleobase-modifying activity. In various embodiments, a base editor comprises a nucleobase-modifying polypeptide (e.g., a deaminase) combined with a guide polynucleotide (e.g., a guide RNA (gRNA)) and a polynucleotide-programmable nucleotide-binding domain (e.g., Cas9 or Cpf1). Exemplary nucleic acid and protein sequences of base editors include sequences having about or at least about 85% sequence identity to any of the base editor sequences provided in the Sequence Listing, such as those corresponding to SEQ ID NOs: 2-11.
[0084] "BE4 cytidine deaminase (BE4) polypeptide" refers to a base editor comprising a nucleic acid programmable DNA binding protein (napDNAbp) domain, a cytidine deaminase domain, and two uracil glycosylase inhibitor domains (UGI). In embodiments, the napDNAbp is a Cas9n (D10A) polypeptide. Non-limiting examples of cytidine deaminase domains include rAPOBEC, ppAPOBEC, RrA3F, AmAPOBEC1, and SsAPOBEC3B.
[0085] By "BE4 cytidine deaminase (BE4) polynucleotide" is meant a polynucleotide that encodes a BE4 polypeptide.
[0086] "Base editing activity" refers to the act of chemically altering a base within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is cytidine deaminase activity, e.g., converting a target C·G to T·A. In another embodiment, the base editing activity is adenosine or adenine deaminase activity, e.g., converting A·T to G·C.
[0087] The term "base editor system" refers to an intermolecular complex for editing nucleobases of a target nucleotide sequence. In various embodiments, a base editor (BE) system comprises: (1) a polynucleotide-programmable nucleotide-binding domain, a deaminase domain (e.g., cytidine deaminase or adenosine deaminase) for deaminating nucleobases in a target nucleotide sequence; and (2) one or more guide polynucleotides (e.g., guide RNAs) combined with the polynucleotide-programmable nucleotide-binding domain. In various embodiments, the base editor (BE) system comprises a nucleobase editor domain selected from adenosine deaminase or cytidine deaminase, and a domain with nucleic acid sequence-specific binding activity. In some embodiments, the base editor system comprises: (1) a base editor (BE) comprising a polynucleotide-programmable DNA-binding domain and a deaminase domain for deaminating one or more nucleobases in a target nucleotide sequence; and (2) one or more guide RNAs combined with the polynucleotide-programmable DNA-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable DNA-binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE) or a cytidine or cytosine base editor (CBE). In some embodiments, the base editor system (e.g., a base editor system comprising cytidine deaminase) comprises a uracil glycosylase inhibitor or other agent or peptide that inhibits the inosine base excision repair system (e.g., a uracil stabilizing protein as provided in WO2022015969, the disclosure of which is incorporated herein by reference in its entirety for all purposes).
[0088] The term "Cas9" or "Cas9 domain" refers to an RNA-guided nuclease comprising the Cas9 protein, or a fragment thereof (e.g., a protein comprising an active, inactive, or partially active DNA-cleavage domain of Cas9, and / or the gRNA-binding domain of Cas9). Cas9 nucleases are also sometimes referred to as casnl nucleases or CRISPR (clustered regularly interspaced short palindromic repeats)-associated nucleases.
[0089] As used interchangeably herein, the terms "coding sequence" or "protein-coding sequence" refer to a segment of a polynucleotide that encodes a protein. A coding sequence may also be referred to as an open reading frame. A region or sequence is bounded proximal to the 5' end by a start codon and proximal to the 3' end by a stop codon. Stop codons useful in the base editors described herein include:
[0090] [Table A]
[0091] A "complex" refers to a combination of two or more molecules whose interaction relies on intermolecular forces. Non-limiting examples of intermolecular forces include covalent and non-covalent interactions. Non-limiting examples of non-covalent interactions include hydrogen bonds, ionic bonds, halogen bonds, hydrophobic bonds, van der Waals interactions (e.g., dipole-dipole interactions, dipole-induced dipole interactions, and London dispersion forces), and the π effect. In embodiments, a complex comprises a polypeptide, a polynucleotide, or a combination of one or more polypeptides and one or more polynucleotides. In one embodiment, a complex comprises one or more polypeptides that associate to form a base editor (e.g., a nucleic acid-programmable DNA-binding protein such as Cas9, and a base editor comprising a deaminase) and a polynucleotide (e.g., a guide RNA). In embodiments, the complex is held together by hydrogen bonds. It should be understood that one or more components of a base editor (e.g., a deaminase, or a nucleic acid-programmable DNA-binding protein) can be covalently or non-covalently associated. As an example, a base editor can include a deaminase covalently linked (e.g., by a peptide bond) to a nucleic acid-programmable DNA-binding protein. Alternatively, a base editor can include a deaminase and a nucleic acid-programmable DNA-binding protein that are non-covalently associated (e.g., when one or more components of the base editor are provided in trans and associated directly or via another molecule, such as a protein or nucleic acid). In embodiments, one or more components of the complex are held together by hydrogen bonds.
[0092] "Cytosine" or "4-aminopyrimidin-2(1H)-one" has the molecular formula C4H5N3O and the structure
[0093] [ka]
[0094] and corresponds to CAS number 71-30-7.
[0095] "Cytidine" has the structure
[0096] [ka]
[0097] and corresponds to CAS number 65-46-3. Its molecular formula is CH 13 It is N3O5.
[0098] "Cytidine base editor (CBE)" means a base editor that includes a cytidine deaminase.
[0099] By "cytidine base editor (CBE) polynucleotide" is meant a polynucleotide that encodes a CBE.
[0100] "Cytidine deaminase" or "cytosine deaminase" refers to a polypeptide or fragment thereof capable of deaminating cytidine or cytosine. In embodiments, the cytidine or cytosine is present in a polynucleotide. In one embodiment, the cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine. The terms "cytidine deaminase" and "cytosine deaminase" are used interchangeably throughout this application. Sea lamprey (Petromyzon marinus) cytosine deaminase 1 (PmCDA1) (SEQ ID NOS: 13-14), activation-induced cytidine deaminase (AICDA) (SEQ ID NOS: 15-21), and APOBEC (e.g., SEQ ID NOS: 12-61) are exemplary cytidine deaminases. Additional exemplary cytidine deaminase (CDA) sequences are provided in the Sequence Listing as SEQ ID NOS: 62-66 and 67-189. Non-limiting examples of cytidine deaminases include those described in PCT / US20 / 16288, PCT / US2018 / 021878, 180802-021804 / PCT, PCT / US2018 / 048969, and PCT / US2016 / 058344. "Cytosine deaminase activity" refers to catalyzing the deamination of cytosine or cytidine. In one embodiment, a polypeptide with cytosine deaminase activity converts an amino group to a carbonyl group. In an embodiment, cytosine deaminase converts cytosine to uracil (i.e., C to U) or 5-methylcytosine to thymine (i.e., 5mC to T). In some embodiments, the cytosine deaminases provided herein have increased cytosine deaminase activity relative to a reference cytosine deaminase (e.g., at least 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold or more).
[0101] As used herein, the term "deaminase" or "deaminase domain" refers to a protein or fragment thereof that catalyzes a deamination reaction.
[0102] "Detecting" refers to determining the presence, absence, or amount of the analyte being detected. In one embodiment, a sequence variation in a polynucleotide or polypeptide is detected. In another embodiment, the presence of an indel is detected.
[0103] "Disease" refers to any condition or disorder that damages or interferes with the normal function of cells, tissues, or organs. Exemplary diseases include autoimmune disorders, such as IgG-mediated autoimmune disorders. Non-limiting examples of autoimmune disorders include myasthenia gravis (gMG), warm autoimmune hemolytic anemia (wAIHA), idiopathic thrombocytopenic purpura (ITP), Graves' disease, chronic inflammatory demyelinating polyneuropathy (CIDP), pemphigus vulgaris, and hemolytic disease of the fetus and newborn (HDFN).
[0104] An "effective amount" refers to the amount of a drug or active compound, e.g., a base editor described herein, required to ameliorate the symptoms of a disease compared to an untreated (treated) patient or an individual not suffering from the disease, i.e., a healthy individual, or the amount of a drug or active compound sufficient to elicit a desired biological response. The effective amount of an active compound used to practice the present invention for the therapeutic treatment of a disease will vary depending on the method of administration, the age, weight, and general health of the subject. Ultimately, the attending physician or veterinarian will determine the appropriate amount and dosage regimen. Such an amount is referred to as an "effective" amount. In one embodiment, an effective amount is the amount of a base editor of the present invention sufficient to introduce an alteration into a gene of interest in a cell (e.g., a cell in vitro or in vivo). In one embodiment, an effective amount is the amount of a base editor required to achieve a therapeutic effect. Such a therapeutic effect need not be sufficient to alter the gene of interest in all cells of a subject, tissue, or organ; altering the gene of interest in about 1%, 5%, 10%, 25%, 50%, 75%, or more of the cells present in the subject, tissue, or organ is sufficient. In one embodiment, the effective amount is sufficient to ameliorate one or more symptoms of the disease.
[0105] "Neonatal Fc receptor for IgG (FcRn) polypeptide" or "Fc fragment of IgG receptor and transporter (FCGRT) polypeptide" refers to a protein having at least about 85% amino acid sequence identity to the NCBI reference sequence NP_001129491 or a fragment thereof that is capable of binding to albumin. Exemplary FcRn polypeptide sequences are provided below. Throughout this disclosure, references are made to amino acid positions within the FcRn polypeptide sequence (e.g., E115(138) or E115). Unless otherwise indicated, such references are made with reference to the sequence below, where the position number outside parentheses corresponds to the position within the FcRn sequence below, excluding the first 23 amino acids corresponding to the signal peptide, and the position within parentheses corresponds to the position within the FcRn sequence below, including the first 23 amino acids.
[0106] [Table B]
[0107] "IgG receptor and transporter Fc fragment (FcRn;FCGRT) polynucleotide" or "IgG receptor and transporter Fc fragment (FCGRT) polynucleotide" refers to a nucleic acid molecule encoding an FcRn polypeptide, as well as introns, exons, 3' untranslated region, 5' untranslated region, and regulatory sequences associated with its expression, or fragments thereof. In embodiments, an FcRn polynucleotide is a genomic sequence, cDNA, mRNA, or gene associated with and / or required for FcRn expression. Exemplary FcRn nucleotide sequences from Homo sapiens are provided below. Further exemplary FcRn nucleotide sequences from Homo sapiens are provided in Ensembl Accession No. ENSG00000211893. 1 aggatgtgag agaggaactg gggtctccag tcacgggagc caggagccgg ccagggccgc 61 121 ctcagccctg ggcgctgggg ctcctgctct ttctccttcc tgggagcctg ggcgcagaaa 181 gccacctctc cctcctgtac caccttaccg cggtgtcctc gcctgccccg gggactcctg 241 ccttctgggt gtccggctgg ctgggcccgc agcagtacct gagctacaat agcctgcggg 301 gcgaggcgga gccctgtgga gcttgggtct gggaaaacca ggtgtcctgg tattgggaga 361 aagagaccac agatctgagg atcaaggaga agctctttct ggaagctttc aaagctttgg 421 ggggaaaagg tccctacact ctgcagggcc tgctggggctg tgaactgggc cctgacaaca 481 cctcggtgcc caccgccaag ttcgccctga acggcgagga gttcatgaat ttcgacctca 541 agcagggcac ctggggtggg gactggcccg aggccctggc tatcagtcag cggtggcagc 601 agcaggacaa ggcggccaac aaggagctca ccttcctgct attctcctgc ccgcaccgcc 661 tgcgggagca cctggagagg ggccgcggaa acctggagtg gaaggagccc ccctccatgc 721 gcctgaaggc ccgacccagc agccctggct tttccgtgct tacctgcagc gccttctcct 781 tctaccctcc ggagctgcaa cttcggttcc tgcggaatgg gctggccgct ggcaccggcc 841 agggtgactt cggccccaac agggtgactt ccttccacgc ctcgtcgtca ctaacagtca 901 aaagtggcga tgagcaccac tactgctgca ttgtgcagca cgcggggctg gcgcagcccc 961 tcagggtgga gctggaatct ccagccaagt cctccgtgct cgtggtggga atcgtcatcg 1021 gtgtcttgct actcacggca gcggctgtag rich gttgtggaga rich 1081 gtgggctgcc agccccttgg atctcccttc gtggagacga caccggggtc ctcctgccca 1141 ccccagggga ggcccaggat gctgatttga aggatgtaaa tgtgattcca gccaccgcct 1201 gaccatccgc cattccgact gctaaaagcg aatgtagtca ggccctttc atgctgtgag 1261 acctcctgga acactggcat ctctgagcct ccagaagggg ttctgggcct agttgtcctc 1321 cctctggagc cccgtcctgt ggtctgcctc agtttcccct cctaatacat atggctgttt 1381 tccacctcga father cgagtttggg cccgaatcag tgtgttctca tcattttca 1441 ggcaggggag gtaagggaat aagtcgggggg actgaatggc ggctgggcct cggatctctc 1501 ctacaggtaa c (SEQ ID NO: 428) "Fragment" refers to a portion of a polypeptide or nucleic acid molecule. This portion comprises at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the entire length of the reference nucleic acid molecule or polypeptide. A fragment may comprise 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids. In some embodiments, a fragment is a functional fragment. "Guide polynucleotide" refers to a polynucleotide or polynucleotide complex that is specific to a target sequence and can form a complex with a polynucleotide-programmable nucleotide-binding domain protein (e.g., Cas9 or Cpf1). In embodiments, the guide polynucleotide is a guide RNA (gRNA). A gRNA can exist as a complex of two or more RNAs or as a single RNA molecule.
[0108] "Hybridization" refers to hydrogen bonding between complementary nucleobases, which may be Watson-Crick, Hoogsteen, or reversed Hoogsteen hydrogen bonds. For example, adenine and thymine are complementary nucleobases that pair through the formation of hydrogen bonds.
[0109] By "immunoglobulin gamma 1 (IgG1) polypeptide" is meant a protein having at least about 85% amino acid sequence identity to GenBank Accession No. CAA75030.1, provided below, or a fragment thereof that has immunomodulatory activity. An exemplary IgG1 amino acid sequence from Homo sapiens is provided in Figure 2A. >CAA75030.1 Immunoglobulin kappa heavy chain [Homo sapiens] MEFGLRWVFLVAILKDVQCDVQLVESGGGLVQPGGSLRLSCAASGFAYSSFWMHWVRQAPGRGLVWVSRINPDGRITVYADAVKGRFTISRDNAKNTLYLQMNNLRAEDTAVYYCARGTR FLELTSRGQMDQWGQGTLVTVSSASTKGPSVFPLAPSSKSTSGGTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFPAVLQSSGLYSLSSVVTVPSSSLGTQTYICNVNHKPSNTKVDKKV EPKSCDKTHTCPPCPAPELLGGPSVFLFPPKPKDTLMISRTPEVTCVVVDVSHEDPEVKFNWYVDGVEVHNAKTKPREEQYNSTYRVVSVLTVLHQDWLNGKEYKCKVSNKALPAPIEKTISKAKGQPREPQVYTLPPSRDELTKNQVSLTCLVKGFYPSDIAVEWESNGQPENNYKTTPPVLDSDGSFFLYSKLTVDKSRWQQGNVFSCSVMHEALHNHYTQKSLSLSPGK (SEQ ID NO: 429).
[0110] "Immunoglobulin gamma 1 (IgG1) polynucleotide" refers to a nucleic acid molecule encoding an IgG1 polypeptide, as well as introns, exons, 3' untranslated regions, 5' untranslated regions, and regulatory sequences associated with its expression, or fragments thereof. In embodiments, an IgG1 polynucleotide is a genomic sequence, cDNA, mRNA, or gene associated with and / or required for IgG1 expression. An exemplary IgG1 nucleotide sequence from Homo sapiens is provided below (GenBank: Y14735.1:36-1457): >Y14735.1:36-1457 Homo sapiens mRNA for immunoglobulin kappa heavy chain
[0111] By "immunoglobulin gamma 2 (IgG2) polypeptide" is meant a protein having at least about 85% amino acid sequence identity to GenBank Accession No. AAB59393.1, provided below, or a fragment thereof having immunomodulatory activity. Exemplary IgG2 amino acid sequences from Homo sapiens include GenBank Accession Nos. AH005273.2:216-509, 902-937, 1056-1382, 1480-1802, and are provided below and in Figure 2A: >AAB59393.1 Immunoglobulin gamma-2 heavy chain, partial [Homo sapiens] STKGPSVFPLAPCSRSTSESTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFPAVLQSSGLYSLSSVVTVPSSNFGTQTYTCNVDHKPSNTKVDKTVERKCCVECPPCPAPPVAGPSVFLFPPKPKDTLMISRTPEVTCVVVDVSHEDPEVQFNWYVDGVEVHNAKTKPREEQFNSTFRVVSVLTVVHQDWLNGKEYKCKVSNKGLPAPIEKTISKTKGQPREPQVYTLPPSREEMTKNQVSLTCLVKGFYPSDIAVEWESNGQPENNYKTTPPMLDSDGSFFLYSKLTVDKSRWQQGNVFSCSVMHEALHNHYTQKSLSLSPGK (SEQ ID NO: 431).
[0112] Exemplary IgG2 Amino Acid Sequences ASTKGPSVFPLAPCSRSTSESTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFPAVLQSSGLYSLSSVVTVPSSNFGTQTYTCNVDHKPSNTKVDKTVERKCCVECPPCPAPPVAGPSVFLFPPKPKDTLMISRTPEVTCVVVDVSHEDPEVQFNWYVDGVEVHNAKTKPREEQFNSTFRVVSVLTVVHQDWLNGKEYKCKVSNKGLPAPIEKTISKTKGQPREPQVYTLPPSREEMTKNQVSLTCLVKGFYPSDISVEWESNGQPENNYKTTPPMLDSDGSFFLYSKLTVDKSRWQQGNVFSCSVMHEALHNHYTQKSLSLSPGK (SEQ ID NO: 432).
[0113] "Immunoglobulin gamma 2 (IgG2) polynucleotide" refers to a nucleic acid molecule encoding an IgG2 polypeptide, as well as introns, exons, 3' untranslated regions, 5' untranslated regions, and regulatory sequences associated with its expression, or fragments thereof. In embodiments, the IgG2 polynucleotide is a genomic sequence, cDNA, mRNA, or gene associated with and / or required for IgG2 expression. Exemplary IgG2 nucleotide sequences from Homo sapiens are provided below (GenBank: AH005273.2:216-509, 902-937, 1056-1382, 1480-1802): >AH005273.2: 216–509, 902–937, 1056–1382, 1480–1802 Homo sapiens immunoglobulin gamma-2 heavy chain (IgH), immunoglobulin gamma-4 heavy chain (IgH), immunoglobulin epsilon chain constant region (IgH), and immunoglobulin alpha-2 heavy chain (IgH) genes, partial cds CCTCCACCAAGGGCCCATCGGTCTTCCCCCTGGCGCCCTGCTCCAGGAGCACCTCCGAGAGCACAGCCGCCCTGGGCTGCCTGGTCAAGGACTACTTCCCCGAACCGGTGACGGTGTCGTGGAACTCAGGCGCTCTGACCAGCGGCGTGCACACCTTCCCAGCTGTCCTACAGTCCTCAGGACTCTACTCCCTCAGCAGCGTGGTGACCGTGCCCTCCAGCAACTTCGGCACCCAGACCTACACCTGCAACGTAGATCACAAGCCCAGCAACACCAAGGTGGACAAGACAGTTGAGCGCAAATGTTGTGTCGAGTGCCCACCGTGCCCAGCACCACCTGTGGCAGGACCGTCAGTCTTCCTCTTCCCCCCAAAACCCAAGGACACCCTCATGATCTCCCGGACCCCTGAGGTCACGTGCGTGGTGGTGGACGTGAGCCACGAAGACCCCGAGGTCCAGTTCAACTGGTACGTGGACGGCGTGGAGGTGCATAATGCCAAGACAAAGCCACGGGAGGAGCAGTTCAACAGCACGTTCCGTGTGGTCAGCGTCCTCACCGTTGTGCACCAGGACTGGCTGAACGGCAAGGAGTACAAGTGCAAGGTCTCCAACAAAGGCCTCCCAGCCCCCATCGAGAAAACCATCTCCAAAACCAAAGGGCAGCCCCGAGAACCACAGGTGTACACCCTGCCCCCATCCCGGGAGGAGATGACCAAGAACCAGGTCAGCCTGACCTGCCTGGTCAAAGGCTTCTACCCCAGCGACATCGCCGTGGAGTGGGAGAGCAATGGGCAGCCGGAGAACAACTACAAGACCACACCTCCCATGCTGGACTCCGACGGCTCCTTCTTCCTCTACAGCAAGCTCACCGTGGACAAGAGCAGGTGGCAGCAGGGGAACGTCTTCTCATGCTCCGTGATGCATGAGGCTCTGCACAACCACTACACGCAGAAGAGCCTCTCCCTGTCTCCGGGTAAATGA (SEQ ID NO: 433).
[0114] By "increase" is meant a positive change of at least 10%, 25%, 50%, 75%, or 100%, or about 1.5-fold, about 2-fold, about 3-fold, about 4-fold, about 5-fold, about 6-fold, about 7-fold, about 8-fold, about 9-fold, about 10-fold, about 15-fold, about 20-fold, about 25-fold, about 30-fold, about 35-fold, about 40-fold, about 45-fold, about 50-fold, or about 100-fold.
[0115] The terms "inhibitor of base repair," "base repair inhibitor," "IBR," or grammatical equivalents thereof, refer to a protein capable of inhibiting the activity of a nucleic acid repair enzyme, e.g., a base excision repair enzyme.
[0116] An "intein" is a fragment of a protein that can excise itself and join the remaining fragment (an extein) via a peptide bond in a process known as protein splicing.
[0117] The terms "isolated," "purified," or "biologically pure" refer to material that is free, to varying degrees, from components that normally accompany it as found in its natural state. "Isolated" refers to a degree of separation from the original source or surrounding environment. "Purified" refers to a degree of separation greater than isolation. A "purified" or "biologically pure" protein is sufficiently free from other substances so that impurities do not significantly affect or otherwise adversely affect the biological properties of the protein. That is, a nucleic acid or peptide of the invention is purified when it is substantially free of cellular material, viral material, or culture medium, if produced by recombinant DNA technology, or when it is chemically synthesized, when it is substantially free of chemical precursors or other chemicals. Purity and homogeneity are typically determined using analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high-performance liquid chromatography. The term "purified" can mean that the nucleic acid or protein gives rise to essentially one band in an electrophoretic gel. In the case of proteins that may undergo modifications, such as phosphorylation or glycosylation, different modifications may result in different isolated proteins, which can be individually purified.
[0118] "Isolated polynucleotide" refers to a nucleic acid molecule that does not contain the genes adjacent to it in the naturally occurring genome of the organism from which the nucleic acid molecule of the present invention is derived. Thus, this term includes, for example, a vector; an autonomously replicating plasmid or virus; or a recombinant DNA that is integrated into the genomic DNA of a prokaryotic or eukaryotic organism; or exists as a separate molecule independent of other sequences (e.g., cDNA or genomic or cDNA fragments generated by PCR or restriction endonuclease digestion). Furthermore, this term also includes RNA molecules transcribed from DNA molecules, and recombinant DNA that is part of a hybrid gene encoding an additional polypeptide sequence.
[0119] By "isolated polypeptide" is meant a polypeptide of the invention that has been separated from components that naturally accompany it. Typically, a polypeptide is isolated when it is at least 60%, by weight, free from the proteins and naturally-occurring organic molecules with which it is naturally associated. Preferably, the preparation is at least 75%, more preferably at least 90%, and most preferably at least 99%, by weight, a polypeptide of the invention. Isolated polypeptides of the invention can be obtained, for example, by extraction from a natural source, by expressing a recombinant nucleic acid encoding such a polypeptide; or by chemically synthesizing the protein. Purity can be measured by any appropriate method, for example, column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.
[0120] As used herein, the term "linker" refers to a molecule that connects two moieties. In one embodiment, the term "linker" refers to a covalent linker (e.g., a covalent bond) or a non-covalent linker.
[0121] "Marker" refers to any analyte, protein, or polynucleotide whose expression, level, structure, or activity is altered in association with a disease or disorder. In embodiments, the marker is an IgG polypeptide or an FcRn polypeptide capable of binding to an autoantigen and / or associated with an autoimmune disease. As used herein, the term "mutation" or "alteration" refers to the substitution of a residue in a polynucleotide or polypeptide sequence with another nucleotide or residue, or the deletion or insertion of one or more nucleotides or residues in the sequence. Mutations are typically described herein by identifying the original residue, followed by the position of the residue in the sequence and the identity of the newly substituted residue. Various methods for making the amino acid substitutions (mutations) provided herein are well known in the art and are provided, for example, in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)).
[0122] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to a compound comprising a nucleic acid base and an acidic moiety, e.g., a nucleoside, a nucleotide, or a polymer of nucleotides. Typically, polymeric nucleic acids, e.g., nucleic acid molecules comprising three or more nucleotides, are linear molecules in which adjacent nucleotides are linked to each other via phosphodiester bonds. In some embodiments, "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides). In some embodiments, "nucleic acid" refers to an oligonucleotide chain comprising three or more individual nucleotide residues. As used herein, the terms "oligonucleotide" and "polynucleotide" can be used interchangeably to refer to a polymer of nucleotides (e.g., a string of at least three nucleotides). In some embodiments, "nucleic acid" encompasses RNA and single- and / or double-stranded DNA. Nucleic acids can occur naturally, e.g., in the context of a genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, chromatid, or other naturally occurring nucleic acid molecule. On the other hand, a nucleic acid molecule may be a non-naturally occurring molecule, e.g., recombinant DNA or RNA, an artificial chromosome, an engineered genome, or a fragment thereof, or synthetic DNA, RNA, or DNA / RNA hybrid, or may contain non-naturally occurring nucleotides or nucleosides. Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms include nucleic acid analogs, e.g., analogs having other than a phosphodiester backbone. Nucleic acids may be purified from natural sources, produced using recombinant expression systems, and, if desired, purified, chemically synthesized, etc. Where appropriate, e.g., in the case of chemically synthesized molecules, nucleic acids include nucleoside analogs, e.g., analogs having chemically modified bases or sugars and backbone modifications. Nucleic acid sequences are presented in the 5' to 3' direction unless otherwise indicated.In some embodiments, nucleic acids are selected from natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, The bases may be or contain: 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine; chemically modified bases; biologically modified bases (e.g., methylated bases); intercalated bases; modified sugars (2'-e.g., fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages).
[0123] The term "nuclear localization sequence," "nuclear localization signal," or "NLS" refers to an amino acid sequence that promotes the import of a protein into the cell nucleus. Nuclear localization sequences are known in the art and are described, for example, in International PCT Application PCT / EP2000 / 011690 by Plank et al., filed November 23, 2000, and published May 31, 2001 as WO / 2001 / 038547, the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, for example, as described in Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises the amino acid sequence KRTADGSEFESPKKKRKV (SEQ ID NO: 190), KRPAATKKAGQAKKKK (SEQ ID NO: 191), KKTELQTTNAENKTKKL (SEQ ID NO: 192), KRGINDRNFWRGENGRKTR (SEQ ID NO: 193), RKSGKIAAIVVKRPRK (SEQ ID NO: 194), PKKKRKV (SEQ ID NO: 195), or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 196).
[0124] The terms "nucleobase," "nitrogenous base," or "base," as used interchangeably herein, refer to nitrogen-containing biological compounds that form nucleosides, which then become components of nucleotides. The ability of nucleobases to base pair and stack with each other directly gives rise to long-chain helical structures such as ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). The five nucleobases, adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), are referred to as primary or standard. Adenine and guanine are derived from purines, while cytosine, uracil, and thymine are derived from pyrimidines. DNA and RNA can also contain other (non-primary) modified bases. Non-limiting exemplary modified nucleobases include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrolauracil, 5-methylcytosine (m5C), and 5-hydromethylcytosine. Hypoxanthine and xanthine can be produced in the presence of mutagens, but both can also be produced by deamination (substitution of an amine group with a carbonyl group). Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can be produced by deamination of cytosine. A "nucleoside" consists of a nucleic acid base and a five-carbon sugar (either ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine. Examples of nucleosides having modified nucleobases include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), and pseudouridine (Ψ). A "nucleotide" consists of a nucleobase, a five-carbon sugar (either ribose or deoxyribose), and at least one phosphate group.Non-limiting examples of modified nucleobases and / or chemical modifications that the modified nucleobases may comprise are as follows: pseudo-uridine, 5-methyl-cytosine, 2'-O-methyl-3'-phosphonoacetate, 2'-O-methylthioPACE (MSP), 2'-O-methyl-PACE (MP), 2'-fluoroRNA (2'-F-RNA), constrained ethyl (S-cEt), 2'-O-methyl ("M"), 2'-O-methyl-3'-phosphorothioate ("MS"), 2'-O-methyl-3'-thiophosphonoacetate ("MSP"), 5-methoxyuridine, phosphorothioate, and N1-methylpseudouridine.
[0125] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" is used interchangeably with "polynucleotide programmable nucleotide binding domain" and can refer to a protein that associates with a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid or guide polynucleotide (e.g., gRNA), that guides the napDNAbp to a specific nucleic acid sequence. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable RNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a Cas9 protein. The Cas9 protein can associate with a guide RNA that guides the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the napDNAbp is a Cas9 domain, such as a nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Non-limiting examples of nucleic acid programmable DNA binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ (Cas12j / Casphi).Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Cas12j / CasΦ, Cpf1, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Examples include Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins, CARF, DinG, homologs thereof, or modified or engineered versions thereof. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, but may not be specifically listed in the present disclosure. See, for example, Makarova et al., "Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?" CRISPR J. 2018 October; 1:325-336. doi:10.1089 / crispr.2018.0033; Yan et al., "Functionally diverse type V CRISPR-Cas systems," Science. 2019 January 4; 363(6422):88-91. doi:10.1126 / science.aav7271, the entire contents of each of which are hereby incorporated by reference herein.Exemplary nucleic acid programmable DNA binding proteins and nucleic acid sequences encoding nucleic acid programmable DNA binding proteins are provided in the Sequence Listing as SEQ ID NOs: 197-230, and 378.
[0126] As used herein, the term "nucleobase editing domain" or "nucleobase editing protein" refers to a protein or enzyme that can catalyze nucleobase modifications in RNA or DNA, such as the deamination of cytosine (or cytidine) to uracil (or uridine) or thymine (or thymidine), and adenine (or adenosine) to hypoxanthine (or inosine), and the addition and insertion of non-templated nucleotides. In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., adenine deaminase or adenosine deaminase; or cytidine deaminase or cytosine deaminase).
[0127] As used herein, "obtaining," as in "obtaining an agent," includes synthesizing, purchasing, or otherwise acquiring an agent.
[0128] "Subject" or "patient" means a mammal, including, but not limited to, a human or non-human mammal. In embodiments, the mammal is a cow, horse, dog, sheep, rabbit, rodent, non-human primate, or cat. In embodiments, "patient" refers to a mammalian subject who has a higher than average likelihood of developing a disease or disorder. Exemplary patients can be humans, non-human primates, cats, dogs, pigs, cows, cats, horses, camels, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, or guinea pigs), and other mammals that can benefit from the therapies disclosed herein. Exemplary human patients can be male and / or female.
[0129] A "patient in need thereof" or "subject in need thereof," as used herein, refers to a patient who has been diagnosed with, is at risk of having, has, has been pre-determined to have, or is suspected of having a disease or disorder.
[0130] The terms "pathogenic mutation," "pathogenic variant," "causing mutation," "pathogenic variant," "deleterious mutation," or "predisposing mutation" refer to a genetic change or mutation that is associated with a disease or disorder or that increases an individual's susceptibility or predisposition to a particular disease or disorder. In some embodiments, a pathogenic mutation comprises at least one wild-type amino acid substituted with at least one pathogenic amino acid in a protein encoded by a gene.
[0131] The terms "protein," "peptide," "polypeptide," and their grammatical equivalents are used interchangeably herein to refer to a polymer of amino acid residues linked together by peptide (amide) bonds. A protein, peptide, or polypeptide can be naturally occurring, recombinant, or synthetic, or any combination thereof.
[0132] As used herein, the term "fusion protein" refers to a hybrid polypeptide comprising protein domains derived from at least two different proteins.
[0133] The term "recombinant," as used herein with respect to a protein or nucleic acid, refers to a protein or nucleic acid that does not occur in nature but is the product of human manipulation. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that contains at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.
[0134] By "reduce" is meant a negative change of at least 10%, 25%, 50%, 75%, or 100%.
[0135] "Reference" refers to a standard or control condition. In one embodiment, the reference is a wild-type or healthy cell. In other embodiments, without limitation, the reference is an untreated cell that is not subjected to the test condition or is subjected to a placebo or normal saline, medium, buffer, and / or a control vector that does not have the polynucleotide of interest. In embodiments, the reference is a cell or a subject that has not been contacted with the base editor system provided herein, or a component thereof. In some cases, the reference is a cell or a subject that has been administered an agent (e.g., a small molecule drug) that interferes with the activity of FcRn in the subject. In some cases, the reference is an FcRn polypeptide (i.e., a wild-type FcRn polypeptide sequence) that does not contain a change in the amino acid residue of interest or does not contain any of the changes provided herein. In various examples, the reference is a cell that has not been altered according to the methods provided herein.
[0136] A "reference sequence" is a defined sequence used as a basis for sequence comparison. A reference sequence may be a subset or the entirety of a specified sequence; for example, a segment of a full-length cDNA or gene sequence, or the complete cDNA or gene sequence. For polypeptides, the length of a reference polypeptide sequence is generally at least about 16 amino acids, at least about 20 amino acids, at least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids. For nucleic acids, the length of a reference nucleic acid sequence is generally at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides, or about 300 nucleotides, or any integer thereabout or therebetween. In some embodiments, the reference sequence is the wild-type sequence of the protein of interest. In other embodiments, the reference sequence is a polynucleotide sequence encoding a wild-type protein.
[0137] The terms "RNA-programmable nuclease" and "RNA-guided nuclease" refer to a nuclease that forms a complex with one or more RNAs that are not the target of cleavage. In some embodiments, when complexed with RNA, the RNA-programmable nuclease can be referred to as a nuclease-RNA complex. Typically, the bound RNA is referred to as a guide RNA (gRNA). In some embodiments, the RNA-programmable nuclease is a (CRISPR-associated) Cas9 endonuclease, such as Cas9 from Streptococcus pyogenes (Csn1) (e.g., SEQ ID NO: 197), Cas9 from Neisseria meningitidis (NmeCas9; SEQ ID NO: 208), Nme2Cas9 (SEQ ID NO: 209), Streptococcus constellatus (ScoCas9), or a derivative thereof (e.g., a sequence having at least about 85% sequence identity to Cas9, such as Nme2Cas9 or spCas9).
[0138] As used herein, the term "scFv" refers to a single-chain Fv antibody in which the variable domains of the heavy and light chains from an antibody are linked to form a single chain. An scFv fragment comprises a single polypeptide chain comprising the variable region of an antibody light chain (VL) (e.g., CDR-L1, CDR-L2, and / or CDR-L3) and the variable region of an antibody heavy chain (VH) (e.g., CDR-H1, CDR-H2, and / or CDR-H3), separated by a linker. The linker connecting the VL and VH regions of an scFv fragment can be a peptide linker composed of proteinogenic amino acids. Alternative linkers may be used to increase the resistance of scFv fragments to proteolysis (e.g., linkers containing D-amino acids), enhance the solubility of scFv fragments (e.g., hydrophilic linkers such as polyethylene glycol-containing linkers or polypeptides containing repeated glycine and serine residues), improve the biophysical stability of the molecules (e.g., linkers containing cysteine residues that form intramolecular or intermolecular disulfide bonds), or attenuate the immunogenicity of scFv fragments (e.g., linkers containing glycosylation sites). It will also be understood by those skilled in the art that the variable regions of the scFv molecules described herein may be modified to differ in amino acid sequence from the antibody molecule from which they are derived. For example, nucleotide or amino acid substitutions resulting in conservative substitutions or changes in amino acid residues may be made (e.g., in CDR and / or framework residues) so as to maintain or enhance the ability of the scFv to bind to the antigen recognized by the corresponding antibody.
[0139] By "specifically binds" is meant a nucleic acid molecule, polypeptide, polypeptide / polynucleotide complex, compound, or molecule that recognizes and binds to a polypeptide and / or nucleic acid molecule of the invention, but does not substantially recognize and bind to other molecules in a sample, e.g., a biological sample.
[0140] "Substantially identical" refers to a polypeptide or nucleic acid molecule that exhibits at least 50% identity to a reference amino acid sequence. In one embodiment, the reference sequence is a wild-type amino acid or nucleic acid sequence. In another embodiment, the reference sequence is any one of the amino acid or nucleic acid sequences described herein. In one embodiment, such a sequence is at least about 60%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or even 99.99% identical at the amino acid or nucleic acid level to the sequence used for comparison.
[0141] Sequence identity is typically measured using sequence analysis software (e.g., the BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs in the sequence analysis software package of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705). Such software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. In an exemplary approach to determining the degree of identity, e, which indicates closely related sequences, is used. -3 From e -100 A BLAST program using a probability score between 0 and 1 may be used.
[0142] COBALT is used, for example, with the following parameters: a) Alignment parameters: Gap penalty -11, -1 and End-Gap penalty -5, -1; b) CDD parameters: Use RPS BLAST; Blast E-value 0.003; Find and recalculate conserved columns; and c) Query clustering parameters: Use query cluster; word size 4; max cluster distance 0.8; Alphabet Regular.
[0143] The EMBOSS Needle is used, for example, with the following parameters: a) Matrix: BLOSUM62; b) GAP OPEN: 10; c) GAP EXTEND: 0.5; d)OUTPUT FORMAT:pair; e)END GAP PENALTY:false; f) END GAP OPEN: 10; and g)END GAP EXTEND:0.5.
[0144] Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding a polypeptide of the present invention or a functional fragment thereof. Such nucleic acid molecules need not be 100% identical to an endogenous nucleic acid sequence, but will typically exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence will typically be able to hybridize with at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding a polypeptide of the present invention or a functional fragment thereof. Such nucleic acid molecules need not be 100% identical to an endogenous nucleic acid sequence, but will typically exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence will typically be able to hybridize with at least one strand of a double-stranded nucleic acid molecule. "Hybridizing" means that the pair forms a double-stranded molecule between complementary polynucleotide sequences (e.g., genes described herein) or portions thereof under various stringency conditions. (See, e.g., Wahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, AR (1987) Methods Enzymol. 152:507).
[0145] For example, a stringent salt concentration is usually less than about 750 mM NaCl and 75 mM trisodium citrate, preferably less than about 500 mM NaCl and 50 mM trisodium citrate, and more preferably less than about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be achieved in the absence of organic solvents, such as formamide, while high stringency hybridization can be achieved in the presence of at least about 35% formamide, more preferably at least about 50% formamide. Stringent temperature conditions usually include a temperature of at least about 30°C, more preferably at least about 37°C, and most preferably at least about 42°C. Various additional parameters, such as hybridization time, the concentration of detergents, such as sodium dodecyl sulfate (SDS), and the inclusion or exclusion of carrier DNA, are well known to those skilled in the art. Various levels of stringency can be achieved by combining these various conditions as needed. In a preferred embodiment, hybridization is carried out in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS at 30° C. In a more preferred embodiment, hybridization is carried out in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA) at 37° C. In a most preferred embodiment, hybridization is carried out in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA at 42° C. Useful variations on these conditions will be readily apparent to those of skill in the art.
[0146] In most applications, the stringency of the washing step following hybridization also varies. Wash stringency conditions can be defined by salt concentration and temperature. As described above, washing stringency can be increased by decreasing the salt concentration or increasing the temperature. For example, stringent salt concentrations for washing steps are preferably less than about 30 mM NaCl and 3 mM trisodium citrate, and most preferably less than about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for washing steps typically include temperatures of at least about 25°C, more preferably at least about 42°C, and even more preferably at least about 68°C. In one embodiment, the washing step is performed at 25°C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In another embodiment, the washing step is performed at 42°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the washing step is carried out in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS at 68°C. Additional variations of these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.
[0147] "Split" means divided into two or more pieces.
[0148] "Split Cas9 protein" or "split Cas9" refers to a Cas9 protein that is provided as an N-terminal fragment and a C-terminal fragment encoded by two separate nucleotide sequences. Polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein can be spliced to form a "reconstituted" Cas9 protein.
[0149] The term "target site" refers to the sequence within a nucleic acid molecule that is being modified. In embodiments, the modification is base deamination. The deaminase can be a cytidine or adenine deaminase. The fusion protein or base editing complex comprising the deaminase can include a dCas9-adenosine deaminase fusion protein, a Cas12b-adenosine deaminase fusion, or a base editor disclosed herein.
[0150] As used herein, the terms "treat," "treating," "treatment," and the like refer to reducing and / or ameliorating a disorder and / or its associated symptoms, or achieving a desired pharmacological and / or physiological effect. It will be understood that treating a disorder or condition does not necessarily require, although does not exclude, that the disorder, condition, or its associated symptoms be completely eliminated. In some embodiments, the effect is therapeutic, i.e., the effect includes, but is not limited to, partially or completely reducing, diminishing, suppressing, alleviating, ameliorating, reducing the intensity of, or curing, the disease and / or adverse symptoms resulting from the disease. In some embodiments, the effect is prophylactic, i.e., the effect protects against or prevents the onset or recurrence of the disease or condition. To this end, the methods disclosed herein comprise administering a therapeutically effective amount of a composition described herein.
[0151] "Uracil glycosylase inhibitor" or "UGI" refers to an agent that inhibits the uracil excision repair system. Base editors, including cytidine deaminase, convert cytosine to uracil, which is then converted to thymine during DNA replication or repair. In various embodiments, uracil DNA glycosylase (UGI) prevents base excision repair that changes U back to C. In some examples, contacting a cell and / or polynucleotide with a UGI and a base editor prevents base excision repair that changes U back to C. Exemplary UGIs include amino acid sequences such as: >splP14739IUNGI_BPPB2 uracil-DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (SEQ ID NO: 231).
[0152] In some embodiments, the agent that inhibits the uracil-excision repair system is a uracil-stabilizing protein (USP). See, e.g., WO2022015969 A1, which is incorporated herein by reference.
[0153] As used herein, the term "vector" refers to a means for introducing a nucleic acid sequence into a cell. Vectors include plasmids, transposons, phages, viruses, liposomes, lipid nanoparticles, and episomes. An "expression vector" is a nucleic acid sequence containing a nucleotide sequence to be expressed in a recipient cell. An expression vector contains a polynucleotide sequence and additional nucleic acid sequences to promote and / or facilitate the expression of the introduced sequence, such as initiation, termination, enhancer, promoter, and secretion sequences, into the genome of a mammalian cell. Examples of vectors include nucleic acid vectors, such as DNA vectors such as plasmids, RNA vectors, viruses, or other suitable replicons (e.g., viral vectors). Various vectors have been developed for delivering polynucleotides encoding foreign proteins into prokaryotic or eukaryotic cells. Examples of such expression vectors are disclosed, for example, in WO1994 / 11026, which is incorporated herein by reference. Specific vectors that can be used to express the editors, e.g., base editors or prime editors, and / or guide polynucleotides of some aspects and embodiments herein include plasmids containing regulatory sequences, such as promoter and enhancer regions, that direct gene transcription. Other useful vectors for expressing antibodies and antibody fragments contain polynucleotide sequences that enhance the translation rate of these genes or improve the stability or nuclear export of the mRNA resulting from gene transcription. These sequence elements include, for example, 5' and 3' untranslated regions, internal ribosome entry sites (IRES), and polyadenylation signal sites to direct efficient transcription of genes carried on the expression vector. Expression vectors of some aspects and embodiments herein may also contain a polynucleotide encoding a marker for selecting cells containing such a vector. Examples of suitable markers include genes encoding resistance to antibiotics such as ampicillin, chloramphenicol, kanamycin, or nourseothricin.
[0154] It is understood that ranges provided herein are shorthand notations for all values within that range. For example, a range of 1 to 50 is understood to include any number, combination of numbers, or subrange from the group including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.
[0155] The recitation of a list of chemical groups within any definition of a variable herein includes definitions of that variable as any single group or combination of listed groups. The recitation of an embodiment for a variable or aspect herein includes that embodiment as any single embodiment or in combination with any other embodiment or portion thereof.
[0156] All terms are intended to be understood as understood by one of ordinary skill in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0157] In this application, the use of the singular includes the plural unless clearly stated otherwise. It must be noted that as used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. In this application, the use of "or" means "and / or" unless stated otherwise. Furthermore, the use of the term "including," as well as other forms such as "include," "includes," and "including," is non-limiting.
[0158] As used in this specification and claims, the words "comprising" (and any form of comprising, such as "comprise" and "comprises"), "having" (and any form of having, such as "have" and "has"), "including" (and any form of including, such as "includes" and "include"), or "containing" (and any form of containing, such as "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. Any embodiment designated as "comprising" a particular component or element is also contemplated in some embodiments as "consisting of" or "consisting essentially of" the particular component or element. It is contemplated that any embodiment discussed herein can be implemented with respect to any method or composition of the disclosure, and vice versa. Additionally, the compositions of the disclosure can be used to achieve the methods of the disclosure.
[0159] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system.
[0160] References herein to "some embodiments," "embodiments," "one embodiment," or "other embodiments" mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least some embodiments, but need not be included in all embodiments of the present disclosure. [Brief explanation of the drawings]
[0161] [Figure 1A]Figures 1A and 1B provide 3D stick structures and plots taken from European Journal of Immunology, 29:2819-2825 (1999), the disclosure of which is incorporated herein by reference in its entirety for all purposes. Figure 1A provides a 3D stick structure of the Fc region of human IgG1. The figure was generated using the RASMOL program (Roger Sayle, Bioinformatics Institute, University of Edinburgh, UK). Figure 1B provides a plot showing the elimination curves of recombinant human Fc-hinge derivatives and Fc-papain fragments in mice, demonstrating the interaction of FcRn with IgG. [Figure 1B] Same as description for Figure 1A. [Figure 2A] Figures 2A and 2B provide a multiple sequence alignment and ribbon structure of IgG2 bound to FcRn. Figure 2A provides an alignment of the amino acid sequences of IgG1 and IgG2, with important binding residues underlined. Figure 2B provides a ribbon structure showing IgG2 binding to FcRn, with key residues indicated. The following sequence is depicted from top to bottom in Figure 2A: LGGPSVFLFPPKPKDTLMISRTPEVTCVVVDVSHEDPEVKFNWYVDGVEVHNAKTKPREEQYNSTYRVVSVLTVLHQDWLNGKEYKCKVSNKALPAPIEKTISKAKGQPREPQVYTLPPSRDELTKNQVSLTCLVKGFYPSDIAVEWESNGQPENNYKTTPPVLDSDGSFFLYSKLTVDKSRWQQGNVFSCSVMHEALHNHYTQKSLSLS PGK (SEQ ID NO: 434) and VAGPSVFLFPPKPKDTLMISRTPEVTCVVVDVSHEDPEVQFNWYVDGVEVHNAKTKPREEQFNSTFRVVSVLTVVHQDWLNGKEYKCKVSNKGLPAPIEKTISKTKGQPREPQVYTLPPSREEMTKNQVSLTCLVKGFYPSDIAVEWESNGQPENNYKTTPPMLDSDGSFFLYSKLTVDKSRWQQGNVFSCSVMHEALHNHYTQKSLSLSPGK (SEQ ID NO: 435). [Figure 2B]Same as description for Figure 2A. [Figure 3A] 3A and 3B provide a ribbon structure for the FcRn:IgG interface. [Figure 3B] Same as description for Figure 3A. [Figure 4] FIG. 4 provides a ribbon structure for the FcRn:IgG binding site, with key residues indicated. [Figure 5] Figure 5 provides a ribbon structure of FcRn bound to IgG, and the structure of FcRn amino acids important for forming a complex with IgG are depicted using spheres. In Figure 5, amino acids that form part of the hydrophobic pocket that helps define W131 (a critical residue for IgG binding) are shown as a cluster of amino acids depicted using spheres in the lightest shade of gray. In Figure 5, amino acids corresponding to the pH-dependent FcRn IgG-binding site are depicted using a cluster of spheres in the darkest shade of gray. In Figure 5, amino acids associated with stabilization and reduced binding affinity of the complex between IgG and FcRn at neutral pH are depicted using spheres in an intermediate shade of gray. Changes to the amino acids depicted in Figure 5 using spheres can, in various embodiments, reduce binding and recycling to IgG1, IgG2, IgG3, and / or IgG4 while advantageously preserving albumin recycling and FcRn expression. In some cases, changes to amino acid residues in FcRn are associated with a >50% reduction in circulating IgG in vivo. [Figure 6A-1]Figures 6A and 6B provide bar graphs showing the base editing rates achieved when HEK293T cells were contacted with a base editing system containing the guide polynucleotide and base editor (i.e., ABE or CBE) indicated on the x-axis. The base editors used were SpCas9-ABE8.8, spCas9-BE4, VRQR spCas9-ABE8.8, VRQR spCas9-BE4, KKH-saCas9-ABE8.8, KKH-saCas9-BE4, SaABE8.8, SaBE4, and spCas9-ABE. Base editing rates are shown for each specific FcRn modification or combination of modifications observed in the base-edited cells. In Figure 6A, bars corresponding to base editing systems that achieved base editing efficiencies greater than 40% are outlined by shaded boxes in Figure 6A. A base editing system containing an adenosine base editor (ABE) and guide RNA gRNA1583 achieved a base editing efficiency of over 70% in HEK293T cells, introducing the W131R modification into FcRn. Figure 6B shows a subset of the data presented in Figure 6A. The arrows in Figure 6B indicate bars corresponding to the base editing efficiencies measured for combined amino acid changes, including E116K and M118I. In Figure 6A, the four rightmost bars correspond to the positive control base editor system. In Figure 6B, the two rightmost bars correspond to the positive control base editor system. In Figure 6A, the amino acid positions listed along the x-axis are numbered from the first amino acid of the FcRn 23-amino acid signal peptide. [Figure 6A-2] Continued from Figure 6A. [Figure 6B] Same as the description for Figure 6A-1. [Figure 7A]Figures 7A-7D provide results from surface plasmon resonance (SPR) measurements of albumin or IgG1 binding to FcRn polypeptides. Figures 7A and 7B provide bar graphs showing results from surface plasmon resonance (SPR) measurements of albumin or IgG1 binding to FcRn polypeptides containing the 10 modifications indicated on the x-axis. All 10 FcRn variants maintained albumin binding. Figure 7A provides a bar graph showing surface plasmon resonance measurements of albumin binding to FcRn polypeptides containing the modifications indicated on the x-axis. [Figure 7B] Figure 7B provides a bar graph showing surface plasmon resonance measurements of IgG1 binding to FcRn polypeptides containing the changes indicated on the X-axis. In Figure 7B, arrows indicate amino acid changes associated with significant reductions in IgG1 binding by FcRn. Four of the FcRn variants evaluated showed reduced IgG binding. [Figure 7C] Figure 7C shows a comparison of the binding of wild-type, M118I, and W131R FcRn to IgG. Measurements were performed with FcRn-biotin on the surface. IgG was injected at the indicated concentrations. [Figure 7D] Figure 7D shows a comparison of the binding of wild-type, M118I, and W131R to albumin. Measurements were performed with FcRn-biotin on the surface. Albumin was injected at the indicated concentrations. [Figure 8A]Figures 8A-8C provide schematic diagrams and bar graphs. Figure 8A provides a schematic diagram describing the experimental scheme used to evaluate base editing in primary human hepatocyte (PHH) coculture. In Figure 8A, "MC" indicates medium change, "TF" indicates transfection with the base editing system, "NGS" indicates next-generation sequencing, and "RT-qPCR" indicates reverse transcriptase quantitative polymerase chain reaction. Samples were collected for next-generation sequencing on day 10 post-transfection, and samples were taken for RT-qPCR measurements on day 13 post-transfection. Cells were transfected with a subsaturating dose of the base editing system (600 ng total, including 160 ng of the terminally modified guide polynucleotide + 450 ng of mRNA encoding the base editor). The gRNA1583 guide, which promoted the creation of the W131(154)R modification, functioned well in the PHH coculture system. [Figure 8B] Figure 8B provides a bar graph showing the base editing efficiency achieved using the base editor system indicated on the x-axis in relation to the specific FcRn modification indicated on the x-axis. Figure 8C provides a bar graph showing the levels of FCGRT exons 5-6 and exons 4-5 detected in mRNA isolated from transfected cells. Transcript levels were normalized to the transcript levels measured for ACTB. Cells edited using a base editor system containing gRNA1583 and an adenosine base editor showed an approximately 30% reduction in FcRn mRNA expression compared to untreated cells and cells edited using guide sg23. In Figures 8B and 8C, a base editor system containing guide g23 (alternatively referred to as gRNA23) and the ABE base editor was used as a positive control. [Figure 8C]Figure 8C provides a bar graph showing the levels of exons 5-6 and exons 4-5 of FCGRT detected in mRNA isolated from transfected cells. Transcript levels were normalized to the transcript levels measured for ACTB. Cells edited with a base editor system containing gRNA1583 and an adenosine base editor showed an approximately 30% reduction in FcRn mRNA expression compared to untreated cells and cells edited with guide sg23. In Figures 8B and 8C, a base editor system containing guide g23 (alternatively referred to as gRNA23) and the ABE base editor was used as a positive control. [Figure 9A] Figures 9A and 9B provide a schematic and bar graph of spacer length optimization in HEK293T cells. HEK293T cells were transfected with mRNA encoding an adenosine base editor and the guide RNA indicated on the x-axis. This included spacers of varying lengths, ranging from 19 to 23 nucleotides. Figure 9A provides a schematic outlining the experimental design for evaluating the effect of spacer length on base editing efficiency. HEK293T cells were seeded on day 0 and transfected with the base editor system on day 1. Medium was changed on day 2, and genomic DNA from the cells was sequenced 72 hours post-transfection using next-generation sequencing. Figure 9B shows the base editing efficiency associated with the indicated FcRn modifications made with the indicated base editing system. Cells were transfected with a subsaturating dose of base editing system (600 ng total, including 160 ng of terminally modified guide polynucleotide + 450 ng of mRNA encoding the base editor). All spacer lengths evaluated showed similar base editing efficiencies for the primary modifications achieved. [Figure 9B] Same as description for Figure 9A. DETAILED DESCRIPTION OF THE INVENTION
[0162] The present invention features compositions and methods for editing, altering the expression of, and / or silencing the neonatal Fc receptor (FcRn) gene, FCGRT.
[0163] The present invention is based, at least in part, on the discovery that base editing can be used to alter an FcRn polypeptide encoded by a cell, such that the polypeptide exhibits reduced binding to IgG while maintaining binding to albumin. Thus, in various embodiments, the methods and base editing systems provided herein can be used to treat IgG-mediated autoimmune disorders by introducing alterations into FcRn that reduce binding to IgG, thereby advantageously reducing IgG half-life in a subject in need of treatment while maintaining the beneficial function of FcRn in albumin cycling.
[0164] Thus, the present disclosure provides improved compositions and methods for the treatment of FcRn-mediated autoimmune disorders.
[0165] Detailed descriptions of embodiments of the subject matter of the present disclosure are provided herein. Modifications to the embodiments described in this document, as well as other embodiments, will become apparent to those skilled in the art after reviewing the information provided in this document.
[0166] Genome editing involves molecular manipulation of genetic material by deleting, substituting, or inserting nucleotide sequences in target genes, and optionally correcting genetic mutations in genes. In embodiments, genome editing includes CRISPR systems, base editing, prime editing, etc.
[0167] The clustered regularly interspaced short palindromic repeat (CRISPR) system is a naturally occurring bacterial and archaeal defense mechanism against viruses. CRISPR systems have been adapted for genome editing in living cells by introducing double-stranded DNA breaks (DSBs) or RNA cleavage at user-defined loci. Porto et al., Base editing: advances and therapeutic opportunities, Nature Reviews 19:839-59 (2020). CRISPR methods involve the use of guide RNAs and nucleic acid-programmable DNA-binding domain Cas proteins, which together introduce cleavage at the target nucleotide sequence. Cas proteins include Cas9, catalytically inactivated (dead) dCas9, nCas9 (nickase), Cas12, and Cas13. Repair of the cleavage by nonhomologous end joining (NHEJ) or homology-directed repair (HDR) introduces insertions, deletions, or point mutations at the cleavage site. The non-specific nature of the mutation may introduce a frameshift into the target nucleotide sequence.
[0168] Base editing allows for the direct conversion of target residues at specific genetic loci without introducing DSBs. Base editing directly introduces single nucleotide modifications into DNA or RNA in living cells. Base editors include those that target DNA and RNA. A DNA base editor comprises a nucleic acid-programmable DNA-binding domain and a cytidine deaminase domain that converts a target CG to TA or a target GC to AT in a target region of DNA (e.g., the FCGRT gene), or an adenosine deaminase domain that converts a target AT to GC or a target TA to CG in a target region of DNA (e.g., the FCGRT gene). In some embodiments, a base editor comprising a cytidine deaminase domain further comprises a uracil glycosylase inhibitor (UGI). Base editing techniques are described in detail, for example, in Porto et al. (2020), the entire contents of which are incorporated herein by reference. In embodiments, the nucleic acid-programmable DNA-binding domain comprises a catalytically inactivated (dead) Cas9 (dCas9) or Cas9 nickase (nCas9).
[0169] Prime editing preserves the target specificity of CRISPR by incorporating an edited RNA template extending from a guide RNA (prime editing guide RNA, or "pegRNA") and a reverse transcriptase fused to nCas9. See, for example, Scholefield et al., "Prime editing: an update on the field," Gene Therapy 28:396-401 (2021). nCas9 does not introduce a DSB, but instead nicks the non-complementary strand of DNA upstream of the PAM site. This nickase exposes a DNA overhang with a 3' OH that binds to the primer binding site (PBS) of the pegRNA. This acts as a primer for reverse transcriptase, which fills in the 3' overhang by copying the sequence of the edited pegRNA. The 5' overhang is excised and the strand is ligated to complete the editing. The prime editing technique is described in detail in Scholefield et al. (2021), which is incorporated herein by reference in its entirety. In embodiments, the prime editor comprises a nucleic acid programmable DNA-binding domain and a reverse transcriptase, the guide RNA is a prime editing guide RNA (pegRNA), and the prime editor replaces one or more nucleotides in the FCGRT gene with different nucleotides. In embodiments, the nucleic acid programmable DNA-binding domain comprises a catalytically inactivated (dead) Cas9 (dCas9) or Cas9 nickase (nCas9).
[0170] In another embodiment, a method for modifying an FcRn protein in a mammalian cell is provided, comprising contacting the cell with a guide RNA and a genome editor, wherein the guide RNA comprises a nucleotide sequence complementary to a portion of the FCGRT gene and targets the genome editor to effect a modification in the FCGRT gene in the cell, which modification changes the amino acid sequence of the FcRn protein encoded by the FCGRT gene. In an embodiment, the genome editor comprises a base editor or a prime editor.
[0171] In another embodiment, a method for treating an IgG-mediated autoimmune disorder in a subject in need thereof is provided, the method comprising modifying an FcRn protein in a mammalian cell of the subject. In certain embodiments, modifying the FcRn protein comprises genome editing of an FCGRT gene in the mammalian cell of the subject. Optionally, genome editing comprises contacting the mammalian cell with a guide RNA and a genome editor, wherein the guide RNA comprises a nucleotide sequence complementary to a portion of the FCGRT gene and targets the genome editor to effect a modification in the FCGRT gene in the cell, which modification changes the amino acid sequence of the FcRn protein encoded by the FCGRT gene. In certain embodiments, the genome editor comprises a base editor or a prime editor.
[0172] Genome editors can be delivered to mammalian cells of interest via various delivery techniques known in the art. In embodiments, genome editors are delivered to mammalian cells via nanoparticles, viral vectors, or electroporation. Suitable nanoparticles for use in the compositions and methods include inorganic nanoparticles (e.g., gold), lipid-based particles (e.g., lipid nanoparticles, liposomes, exosomes, cell-derived membrane-bound particles, etc.), peptide nanoparticles, polymer nanoparticles, etc.
[0173] Various viral vectors are known in the art and are suitable for use in delivering the compositions of the present disclosure.In embodiments, the viral vector is selected from the group consisting of retrovirus (e.g., HIV, lentivirus), adenovirus, adeno-associated virus (AAV), herpesvirus (e.g., HSV) and Sendai virus.
[0174] The compositions and methods disclosed herein modify nucleic acids encoding FcRn proteins by introducing one or more single nucleotide modifications into the FCGRT gene. In embodiments, the modified or variant FcRn proteins exhibit reduced ability to bind to the Fc region of IgG antibodies. In further embodiments, the modified or variant FcRn proteins contain at least one amino acid change compared to a reference FcRn protein, such as a wild-type FcRn protein.
[0175] The disclosed methods can be performed ex vivo, in vitro, or in vivo. That is, the compositions disclosed herein can be administered directly to a subject (e.g., intravenously or topically, by injection, inhalation, etc.) or can be administered to a cell, optionally obtained from the subject. In embodiments, the subject is a human.
[0176] Various modifications can be made to the FCGRT gene to provide the modified FcRn proteins disclosed herein. In embodiments, the modified FcRn protein differs from the reference FcRn protein in one or more amino acids selected from the group consisting of leucine (L) at position 112, glutamic acid (E) at position 115, glutamic acid (E) at position 116, tryptophan (W) at position 131, proline (P) at position 132, and glutamic acid (E) at position 133. In other embodiments, the modified FcRn protein contains one or more mutations as shown in Figure 4.
[0177] Optionally, the genome editor or delivery vehicle is conjugated or incorporated with a targeting moiety that binds to FcRn or albumin. In certain embodiments, the targeting moiety is selected from the group consisting of the Fc domain of IgG, an antibody that specifically binds to FcRn, an antibody that specifically binds to albumin, a peptide that binds to albumin, albumin, or a fragment or derivative thereof.
[0178] Additional targeting moieties include, but are not limited to, variant Fc domains that bind to the extracellular domain of FcRn; antibodies or other specific binding agents (e.g., engineered scaffold proteins such as affibodies, darpins, or peptides (which can be selected using display technologies such as phage display)); albumin or a fragment or variant thereof that retains the ability to bind to FcRn. In this approach, albumin (or a fragment / variant) binds to FcRn, and the delivery vehicle / active agent is internalized along with the albumin (or fragment / variant). Other targeting moieties include other specific binding agents (e.g., engineered scaffold proteins such as affibodies or darpins or peptides (which can be selected using display technologies such as phage display)) that bind to albumin but do not substantially interfere with albumin binding to FcRn. The delivery vehicle is internalized by cells along with albumin when albumin binds to FcRn.
[0179] Also provided herein is a composition comprising a guide RNA and a genome editor, wherein the guide RNA comprises a nucleotide sequence complementary to a portion of the FCGRT gene and targets the genome editor to cause a modification to the FCGRT gene in a cell, which modification changes the amino acid sequence of the FcRN protein encoded by the FCGRT gene. The disclosed composition may further comprise a delivery vehicle described herein and / or a targeting moiety that binds to FcRn and / or albumin.
[0180] In certain embodiments, the delivery vehicle disclosed herein comprises a guide RNA and a genome editor or a nucleic acid encoding the genome editor. In embodiments, the delivery vehicle comprises a targeting moiety that binds to FcRn and / or albumin.
[0181] Lipid nanoparticles (LNPs) are spherical, nanometer-scale particles containing an ionizable lipid monolayer shell and a lipid core matrix that can solubilize lipophilic molecules such as drugs or nucleic acids. Traditional LNPs are taken up by host cells via endocytosis, escape endosomes, and release their cargo into the host cell cytoplasm. LNPs are generally considered safe, effective, and suitable for industrial manufacturing and clinical use in drug delivery.
[0182] An embodiment of the LNP disclosed herein includes an Fc region or fragment of an Fc region of an IgG antibody or other targeting moiety embedded or incorporated into a lipid monolayer shell, and encapsulates within the core a nucleic acid for silencing or modulating the expression of FcRn (the FCGRT gene). When the LNP contacts FcRn on the surface of an epithelial cell, the Fc region or fragment thereof binds to FcRn, and the LNP fuses with the cell or otherwise internalizes, delivering its payload. The released nucleic acid then silences, modulates, or alleviates the expression of FcRn, which in turn results in a reduction in circulating IgG (but preferably not albumin) in the host and a reduction in the symptoms and pathology of autoimmune disorders.
[0183] In one embodiment, a solid LNP is provided, comprising a lipid monolayer membrane containing at least one Fc region of an IgG antibody or functional fragment thereof embedded in the lipid monolayer membrane; and a lipid core matrix encapsulated in the lipid monolayer membrane. In one embodiment, the lipid core of the LNP comprises at least one nucleic acid.
[0184] In one embodiment, the IgG or fragment thereof incorporated into the LNP is of the IgG1 subclass. In a specific embodiment, the IgG1 or fragment thereof has the following amino acid substitution: aspartic acid at position 265 is substituted with alanine, or proline at position 238 is substituted with alanine.
[0185] In another embodiment, the IgG incorporated into the LNP is of the IgG2 subclass or a fragment thereof.
[0186] In another embodiment, the IgG incorporated into the LNP is of the IgG3 subclass or a fragment thereof.
[0187] In another embodiment, the IgG incorporated into the LNP is of the IgG4 subclass or a fragment thereof.
[0188] In another embodiment, the IgG incorporated into the LNP recognizes the FcRn receptor. In a specific embodiment, the IgG1 or fragment thereof has the following amino acid substitution: aspartic acid at position 265 is substituted with alanine, or proline at position 238 is substituted with alanine.
[0189] In another embodiment, the IgG is not incorporated into LNP and recognizes the FcRn receptor. In a specific embodiment, the IgG1 or fragment thereof has the following amino acid substitution: aspartic acid at position 265 is substituted with alanine, or proline at position 238 is substituted with alanine. In a specific embodiment, the IgG can directly deliver a payload.
[0190] In some embodiments, the engineered Fc variant has increased affinity for FcRn at basic pH (e.g., a pH typical of blood, e.g., 7.35-7.45) compared to a naturally occurring Fc region.
[0191] In some embodiments, the nucleic acid incorporated into the lipid core of the LNP is DNA or RNA. In certain embodiments, the nucleic acid is small interfering RNA (siRNA), microRNA (miRNA), guide RNA, pegRNA, or short hairpin RNA (shRNA). In very specific embodiments, the nucleic acid is siRNA. In another specific embodiment, the nucleic acid is guide RNA or pegRNA. In another embodiment, the nucleic acid encodes a genome editor.
[0192] In embodiments, the siRNA is functional to modulate the expression of one or more genes. In a specific embodiment, the siRNA modulates the expression of the gene encoding FCGRT, the neonatal Fc receptor (FcRn).
[0193] In embodiments, the nucleic acid incorporated into the LNP is a guide RNA that is functional to target a genome editor to edit or modify the gene encoding FCGRT, the neonatal Fc receptor (FcRn). Suitable modifications of the FCGRT gene are shown, for example, in Figure 4 of the present disclosure.
[0194] In certain embodiments, the tryptophan residues at positions 51 or 61 and the histidine at position 166 are not modified as these amino acids are responsible for binding and extending the half-life of human serum albumin.
[0195] Various lipids are suitable for use in the lipid monolayer of the disclosed LNPs. In embodiments, the lipid monolayer is composed of lipids selected from the group consisting of lecithin, phosphatidylcholine, phosphatidic acid, phosphatidylethanolamine, phosphatidylglycerol, phosphatidylserine, phosphatidylinositol, cardiolipin, lipid-polyethylene glycol conjugates, and combinations thereof. In embodiments, the lipids of the lipid monolayer can be at least partially PEGylated to facilitate immune clearance of the LNPs. In embodiments, the lipid monolayer can further comprise cholesterol as a stabilizer.
[0196] The lipid core matrix of the disclosed LNP comprises a cationic lipid suitable for complexing with nucleic acid in the core. As used herein, the term "cationic lipid" encompasses any of several lipid species that carry a net positive charge at physiological pH, which can be determined using any method known to those skilled in the art. Such lipids include, but are not limited to, the cationic lipids of formula (I) disclosed in International Application No. PCT / US2009 / 042476, entitled "Methods and Compositions Comprising Novel Cationic Lipids," filed May 1, 2009, the entire contents of which are incorporated herein by reference. These include, but are not limited to, N-methyl-N-(2-(arginoylamino)ethyl)N,N-dioctadecylaminium chloride or distearoylarginylammonium chloride (DSAA), N,N-dimyristoyl-N-methyl-N-2[N'-(N6-guanidino-L-lysinyl)]aminoethylammonium chloride (DMGLA), N,N-dimyristoyl-N-methyl-N-2[N2-guanidino-L-lysinyl]aminoethylammonium chloride, N,N-dimyristoyl-N-methyl-N-2[N'-(N2,N6-di-guanidino-L-lysinyl)]aminoethylammonium chloride, and N,N-di-stearoyl-N-methyl-N-2[N'-(N6-guanidino-L-lysinyl)]aminoethylammonium chloride (DSGLA). Other non-limiting examples of cationic lipids that can be present in the liposomes or lipid bilayers of the lipid nanoparticles of the present disclosure include N,N-dioleyl-N,N-dimethylammonium chloride (DODAC); N-(2,3-dioleyloxy)propyl)-N,N,N-trimethylammonium chloride (DOTAP); N-(2,3-dioleyloxy)propyl)-N,N,N-trimethylammonium chloride (DOTMA) or other N(N,Nl-dialkoxy)-alkyl-N,N,N-trisubstituted ammonium surfactants; N,N-distearyl-N,N-dimethylammonium bromide (DAB);3-(N-(N',N'-dimethylaminoethane)carbamoyl)cholesterol (DC-Choi) and N-(1,2-dimyristyloxyprop-3-yl)-N,N-dimethyl-N-hydroxyethylammonium bromide (DMRIE); 1,3-dioleyl-3-trimethylammonium propane, N-(1-(2,3-dioleyloxy)propyl)-N-(2-(sperminecarboxamido)ethyl)-N,N-dimethyl-1-ammonium trifluoroacetate (DOSPA); GAP-DLRIE; DMDHP; 3-p[4N -(H8N-diguanidinospermidine)-carbamoyl]cholesterol (BGSC); 3-P[N,N-diguanidinoethyl-aminoethane)-carbamoyl]cholesterol (BGTC); N,N\N2,N3 tetra-methyltetrapalmitylspermine (Cellfectin); Nt-butyl-N'-tetradecyl-3-tetradecyl-aminopropionamidine (CLONfectin); dimethyldioctadecylammonium bromide (DDAB); 1,3-dioleoyloxy-2-(6-carboxspermyl)-propylamide (DOS PER); 4-(2,3-bis-palmitoyloxy-propyl)-1-methyl-1H-imidazole (DPIM); N,N,N',N'-tetramethyl-N,N'-bis(2-hydroxyethyl)-2,3 dioleoyloxy-1,4 butanediammonium iodide (Tfx-50); 1,2-dioleoyl-3-(4'-trimethylammonio)butanol-sn-glycerol (DOBT) or trimethylammonium groups are attached via a butanol spacer arm to the double chain (for DOTB) or cholesteryl group (for ChOTB). Cholesteryl (4'-trimethylammonium) butanoate (ChOTB) bound to either; DL-1,2-dioleoyl-3 dimethylaminopropyl-p-hydroxyethylammonium (DORI) or DL-1,2-0-dioleoyl-3-dimethylaminopropyl-p-hydroxyethylammonium (DORIE), or analogs thereof disclosed in International Application Publication No. WO 93 / 03709, the entire contents of which are incorporated herein by reference; 1,2-dioleoyl-3-succinyl-sn-glycerol choline ester (DOSC);Cholesteryl hemisuccinate ester (ChOSC); lipopolyamines, such as dioctadecylamidoglycylspermine (DOGS) and dipalmitoylphosphatidylethanolamylspermine (DPPES), or cationic lipids disclosed in U.S. Pat. No. 5,283,185, the entirety of which is incorporated herein by reference; cholesteryl-3P-carboxyl-amido-ethylenetrimethylammonium iodide; 1-dimethylamino-3-trimethylammonium-DL-2-propyl-cholesterylcarboxylate iodide; cholesteryl -3-β-carboxyamidoethyleneamine; cholesteryl-3-P-oxysuccinamido-ethylenetrimethylammonium iodide; 1-dimethylamino-3-trimethylammonio-DL-2-propyl-cholesteryl-3-P-oxysuccinate iodide; 2-(2-trimethylammonio)-ethylmethylaminoethyl-cholesteryl-3-P-oxysuccinate iodide; 3-β-N-(polyethyleneimine)-carbamoylcholesterol, DC-cholesterol; and N4-cholesteryl-spermine HCl salt (GL67).
[0197] In embodiments, the lipid core matrix further comprises cholesterol as a stabilizer.
[0198] In another embodiment, a pharmaceutical composition is provided comprising at least one LNP comprising a lipid monolayer membrane comprising at least one Fc region of an IgG antibody or a functional fragment thereof embedded in the lipid monolayer membrane, and a lipid core matrix encapsulated in the lipid monolayer membrane, wherein the lipid core matrix comprises at least one nucleic acid and at least one pharmaceutically acceptable excipient.
[0199] If necessary, pharmaceutical compositions are formulated for local administration or systemic administration to the subject.The administration for delivering the compound of combination therapy systemically or to desired surface or target can include but is not limited to injection, infusion, drip infusion and inhalation administration.Injection includes but is not limited to intravenous, intramuscular, intraarterial, intrathecal, intracerebroventricular, intracapsular, intraorbital, intracardiac, intradermal, intraperitoneal, transtracheal, subcutaneous, subcuticular and intraarticular injection and infusion.
[0200] Pharmaceutical compositions for injection include aqueous solutions or dispersions and sterile powders for the extemporaneous preparation of sterile injectable solutions or dispersions. For intravenous administration, suitable carriers include, but are not limited to, physiological saline, bacteriostatic water, Cremophor EL™ (BASF, Parsippany, NJ), or phosphate-buffered saline (PBS). The carrier can be a solvent or dispersion medium containing, for example, water, ethanol, polyols (e.g., glycerol, propylene glycol, liquid polyethylene glycol, and the like), and suitable mixtures thereof. Fluidity can be maintained, for example, by the use of a coating such as lecithin, by maintaining the required particle size in the case of dispersions, and by the use of surfactants. Isotonic agents, for example, sugars, polyalcohols such as mannitol, sorbitol, and sodium chloride, can be included in the composition. The resulting solution can be packaged for immediate use or lyophilized. The lyophilized preparation can then be combined with a sterile solution prior to administration.
[0201] In another embodiment, there is provided a method for treating an IgG-mediated autoimmune disorder in a subject in need thereof, the method comprising: administering to the subject LNPs comprising a lipid monolayer membrane comprising at least one Fc region of an IgG antibody, or a functional fragment thereof, embedded in the lipid monolayer membrane; and a lipid core matrix encapsulated in the lipid monolayer membrane, wherein the lipid core matrix comprises at least one siRNA or guide RNA that attenuates expression of or silences the FCGRT gene.
[0202] IgG-mediated autoimmune disorders include, but are not limited to, myasthenia gravis, warm autoimmune hemolytic anemia (wAIHA), idiopathic thrombocytopenic purpura (ITP), Graves' disease, chronic inflammatory demyelinating polyneuropathy (CIDP), pemphigus vulgaris, and hemolytic disease of the fetus and newborn (HDFN).
[0203] In another embodiment, a method for silencing FcRn expression in a cell is provided, comprising contacting the cell with LNPs, wherein the LNPs comprise a lipid monolayer membrane comprising at least one Fc region of an IgG antibody or a functional fragment thereof embedded in the lipid monolayer membrane, and a lipid core matrix encapsulated in the lipid monolayer membrane, wherein the lipid core matrix comprises at least one siRNA that silences the FCGRT gene. In embodiments, the method is ex vivo, in vivo, or in vitro.
[0204] FcRn Immunoglobulin G (IgG) (see, for example, Figures 1 and 2A) is the most common type of antibody found in the blood circulation and extracellular fluids, where it controls infection of living tissues. While IgG can bind directly to antigens, the neonatal Fc receptor for IgG (FcRn) also binds to receptors on cells to initiate an immune response. The Fc gamma receptor (FcγR) family includes the atypical neonatal Fc receptor (FcRn), encoded by the FCGRT gene. FcRn functions to recycle and maintain IgG and albumin and transport them across polarized cell barriers, thereby increasing the half-life of IgG and albumin in the circulation. FcRn also interacts with peptides derived from IgG immune complexes (ICs) and promotes their antigen presentation.
[0205] FcRn was first identified as the receptor responsible for transporting maternal IgG antibodies from mother to child and promoting passive humoral immunity in the mother-to-child relationship. FcRn binds to the Fc region of monomeric immunoglobulin gamma (see Figures 1B and 2B-5) and mediates its selective uptake from milk. IgG in milk binds to the apical surface of the intestinal epithelium. The resulting FcRn-IgG complex is transcytosed across the intestinal epithelium, and IgG is released from FcRn into the blood or tissue fluids. Throughout life, it contributes to effective humoral immunity by recycling IgG and extending its half-life in the circulation. Mechanistically, monomeric IgG binds to FcRn in the acidic endosomes of endothelial and hematopoietic cells, where it is recycled to the cell surface and released into the circulation.
[0206] Initially, FcRn was thought to be present only in fetal and neonatal placenta and intestinal tissues. However, it is now known that FcRn is expressed in many tissues throughout the body, including epithelial, endothelial, and hematopoietic tissues. Specifically, epithelial FcRn expression has been detected in the intestine, placenta, kidney, and liver.
[0207] Mechanistically, monomeric IgG that binds to FcRn in the acidic endosomes of endothelial and hematopoietic cells is recycled to the cell surface, where it is released into the circulation. In addition to IgG, FcRn regulates the homeostasis of the other most abundant circulating protein, albumin / ALB.
[0208] FcRn is expressed in many tissues. For example, it is expressed in the liver, hepatocytes, and Müller cells. FcRn is also highly expressed in epithelial, endothelial, and myeloid cells and plays multiple roles in adaptive immunity. On myeloid cells, FcRn, along with classical FcγRs and complement, is involved in both phagocytosis and antigen presentation. In podocytes (kidneys), FcRn reabsorbs IgG from the glomerular basement membrane, preventing the deposition of immune complexes that can cause glomerular disease.
[0209] For example, several autoimmune diseases, including myasthenia gravis (gMG), warm autoimmune hemolytic anemia (wAIHA), idiopathic thrombocytopenic purpura (ITP), Graves' disease, chronic inflammatory demyelinating polyneuropathy (CIDP), pemphigus vulgaris, and hemolytic disease of the fetus and newborn (HDFN), are caused by IgG responses to autoantigens. Because FcRn functions to maintain circulating IgG levels, it also prolongs the half-life of antibodies that cause these autoimmune diseases. Intravenous immunoglobulin (IVIg) is a recently developed treatment that promotes the reduction of IgG autoantibody levels by saturating the IgG recycling capacity of FcRn and reducing the level of pathogenic IgG binding to FcRn.
[0210] Efgartigimod (ARGX-113, VYVGART) is an IV / SC therapy originally developed by Argenx to treat myasthenia gravis (gMG). Efgartigimod is an IgG1 Fc fragment with increased affinity for FcRn. Efgartigimod blocks IgG access to FcRn and reduces its overall serum half-life. Administration of Efgartigimod (approximately 10 mg / kg / week administered using a single IV infusion) to subjects has been associated with a 50-70% reduction in IgG in subjects.
[0211] Various modifications can be made to the FCGRT gene to provide the modified FcRn proteins disclosed herein. The modifications affect the serum half-life of IgG in a subject comprising an FcRn protein modified according to the methods provided herein. In embodiments, the modified FcRn protein differs from a reference FcRn protein in one or more amino acids selected from the group consisting of leucine (L) at position 112, glutamic acid (E) at position 115, glutamic acid (E) at position 116, tryptophan (W) at position 131, proline (P) at position 132, and glutamic acid (E) at position 133. In other embodiments, the modified FcRn protein includes one or more modifications as shown in any one of Table 1, Figures 2B, 4-7B, 8B, 8C, and 9B, and / or a modification at position M118(141) (e.g., M118(141)I).
[0212] [Table 1-1]
[0213] [Table 1-2]
[0214] In some embodiments, the methods and compositions of the disclosure are used to introduce alterations into one or more of the amino acids underlined or bolded in the following FcRn amino acid sequence:
[0215] [Table C]
[0216] In embodiments, the methods provided herein are used to produce FcRn containing alterations that modify one or more of the following properties of FcRn: A) the stability (e.g., reduced or increased) of the complex formed between FcRn and IgG; B) the binding affinity (e.g., reduced or increased) for IgG at neutral pH; C) the binding affinity (e.g., reduced or increased) for IgG at a pH below or above neutral; D) the positioning of W131 (e.g., to reduce or increase binding to IgG).
[0217] In certain embodiments, the tryptophan residues at positions 51 or 61 and the histidine at position 166 are not modified as these amino acids are responsible for binding and extending the half-life of human serum albumin.
[0218] In another embodiment, a method for silencing FcRn expression in a cell is provided, comprising contacting the cell with LNP, wherein the LNP comprises a lipid monolayer membrane, the lipid monolayer membrane comprising at least one Fc region of an IgG antibody or a functional fragment thereof, and a lipid core matrix encapsulated in the lipid monolayer membrane, the lipid core matrix comprising at least one siRNA that silences the FCGRT gene. In an embodiment, the method is ex vivo, in vivo, or in vitro.
[0219] Targeted gene editing In some embodiments, to generate the gene edits described herein, cells (e.g., cells from a subject, such as hepatocytes, endothelial cells, epithelial cells, or myeloid cells) are contacted in vivo or in vitro with one or more guide RNAs and a nucleobase editor polypeptide comprising a nucleic acid-programmable DNA-binding protein (napDNAbp) and a cytidine deaminase or adenosine deaminase. In some embodiments, the cells to be edited are contacted with at least one polynucleotide, wherein the polynucleotide encodes one or more guide RNAs and a nucleobase editor polypeptide comprising a nucleic acid-programmable DNA-binding protein (napDNAbp) and a cytidine deaminase. In some embodiments, the gRNA comprises one or more nucleotide analogs. In some examples, the gRNA is added directly to the cell. In some embodiments, these nucleotide analogs can inhibit degradation of the gRNA from cellular processes.
[0220] In various examples, it is advantageous to include 5' and / or 3' "G" nucleotides in the spacer sequence. In some embodiments, for example, any spacer sequence or guide polynucleotide provided herein includes or further includes a 5' "G", where in some embodiments, the 5' "G" is complementary or not complementary to the target sequence. In some embodiments, a 5' "G" is added to a spacer sequence that does not already include a 5' "G". For example, because the U6 promoter prefers a "G" at the transcription start site, it may be advantageous to include a 5'-terminal "G" in the guide RNA when the guide RNA is expressed under the control of, for example, a U6 promoter (see Cong, L. et al., "Multiplex genome engineering using CRISPR / Cas systems." Science 339:819-823 (2013) doi:10.1126 / science.1231143). In some embodiments, a 5'-terminal "G" is added to a guide polynucleotide that is expressed under the control of a promoter, but is optionally not added to a guide polynucleotide if or when the guide polynucleotide is not expressed under the control of a promoter.
[0221] In embodiments, the guide polynucleotide comprises: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (SpCas9 scaffold; SEQ ID NO: 317) and GUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGAUUUU (SaCas9 scaffold; SEQ ID NO: 436) The scaffold comprises a nucleotide sequence selected from:
[0222] Tables 2A and 2B provide exemplary gRNA sequences (e.g., complete guide and spacer sequences) suitable for use in embodiments of the present disclosure.
[0223] [Table 2A-1]
[0224] [Table 2A-2]
[0225] [Table 2A-3]
[0226] [Table 2A-4]
[0227] [Table 2B-1]
[0228] [Table 2B-2]
[0229] Nucleic acid base editor Nucleic acid base editors that edit, modify, or change the target nucleotide sequence of a polynucleotide are useful in the methods and compositions described herein. The nucleic acid base editors described herein typically include a polynucleotide programmable nucleotide binding domain and a nucleic acid base editing domain (e.g., adenosine deaminase, cytidine deaminase). When combined with a bound guide polynucleotide (e.g., gRNA), the polynucleotide programmable nucleotide binding domain specifically binds to the target polynucleotide sequence, thereby allowing the base editor to localize to the target nucleic acid sequence that is desired to be edited.
[0230] In certain embodiments, the nucleobase editors provided herein comprise one or more features that improve base editing activity. For example, any of the nucleobase editors provided herein may comprise a Cas9 domain with reduced nuclease activity. In some embodiments, any of the nucleobase editors provided herein may comprise a Cas9 domain without nuclease activity (dCas9) or a Cas9 domain that cleaves one strand of a double-stranded DNA molecule, referred to as a Cas9 nickase (nCas9). Without wishing to be bound by any particular theory, the presence of a catalytic residue (e.g., H840) maintains the activity of Cas9 to cleave the non-edited (e.g., non-deaminated) strand opposite the targeted nucleobase. Mutation of the catalytic residue (e.g., D10 to A10) prevents cleavage of the edited (e.g., deaminated) strand containing the targeted residue (e.g., A or C). Such Cas9 variants can generate single-stranded DNA breaks (nicks) at specific locations based on the target sequence defined by the gRNA, leading to repair of the non-edited strand and ultimately resulting in nucleobase changes on the non-edited strand.
[0231] Polynucleotide Programmable Nucleotide Binding Domains The polynucleotide-programmable nucleotide binding domain binds to a polynucleotide (e.g., RNA, DNA). The polynucleotide-programmable nucleotide binding domain of a base editor may itself comprise one or more domains (e.g., one or more nuclease domains). In some embodiments, the nuclease domain of the polynucleotide-programmable nucleotide binding domain comprises an endonuclease or an exonuclease. An endonuclease can cleave one strand of a double-stranded nucleic acid molecule or both strands of a double-stranded nucleic acid molecule. In some embodiments, the nuclease domain of the polynucleotide-programmable nucleotide binding domain can cleave zero, one, or two strands of a target polynucleotide.
[0232] Non-limiting examples of polynucleotide-programmable nucleotide binding domains that can be incorporated into base editors include CRISPR protein-derived domains, restriction nucleases, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs). In some embodiments, the base editor comprises a polynucleotide-programmable nucleotide binding domain comprising a natural or engineered protein or a portion thereof, which can bind to a nucleic acid sequence during CRISPR (i.e., clustered regularly interspaced short palindromic repeats)-mediated modification of the nucleic acid via a bound guide nucleic acid. Such proteins are referred to herein as "CRISPR proteins." Thus, disclosed herein are base editors comprising a polynucleotide-programmable nucleotide binding domain comprising all or a portion (e.g., a functional portion) of a CRISPR protein (i.e., a base editor comprising all or a portion (e.g., a functional portion) of a CRISPR protein as a domain, also referred to as the "CRISPR protein-derived domain" of the base editor). The domain from the CRISPR protein incorporated into the base editor can be modified compared to a wild-type or naturally occurring version of the CRISPR protein. For example, as described below, the domain from the CRISPR protein can include one or more mutations, insertions, deletions, rearrangements, and / or recombinations compared to a wild-type or naturally occurring version of the CRISPR protein.
[0233] Cas proteins that may be used herein include class 1 and class 2. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csb4, Csb5, Csb6, Csb7, Csb8, Csb9, Csb10, Csb11, Csb12, Csb13, Csb14, Csb15, Csb16, Csb17, Csb18, Csb19, Csb19, Csb19, Csb11, Csb12, Csb13, Csb14, Csb15, Csb16 ...7, Csb18, Csb19, Csb19, Csb19, Csb19, Csb19, Csb19, Csb19, Csb19, Csb19, Csb19, Csb19, Csb19, Csb19, Csb Examples of such proteins include sx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1 (e.g., SEQ ID NO: 232), Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ, CARF, DinG, homologs thereof, or modified versions thereof. CRISPR enzymes can direct cleavage of one or both strands at a target sequence, for example, within the target sequence and / or within the complement of the target sequence. For example, CRISPR enzymes can direct cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence.
[0234] A vector can be used that encodes a CRISPR enzyme that is mutated relative to the corresponding wild-type enzyme, so that the mutant CRISPR enzyme lacks the ability to cleave one or both strands of the target polynucleotide containing the target sequence.Cas protein (e.g., Cas9, Cas12) or Cas domain (e.g., Cas9, Cas12) can refer to a polypeptide or domain that has at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity and / or sequence homology with the wild-type exemplary Cas polypeptide or Cas domain.Cas (e.g., Cas9, Cas12) can refer to wild-type or modified forms of Cas protein, which can include amino acid changes such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof.
[0235] In some embodiments, the base editor CRISPR protein-derived domains are derived from Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcus iniae (NCBI Ref: NC_021846.1); iniae (NCBI Ref:NC_021314.1); Belliella baltica (NCBI Ref:NC_018010.1); Psychroflexus torquis (NCBI Ref:NC_018721.1); Streptococcus thermophilus (NCBI Ref:YP_820832.1); Listeria innocua (NCBI Ref:NP_472073.1); Campylobacter jejuni (NCBI Ref:YP_002344900.1); Neisseria meningitidis (NCBI Ref:YP_002344900.1) Ref:YP_002342100.1); may include all or a portion (e.g., a functional portion) of Cas9 from Streptococcus pyogenes, or Staphylococcus aureus.
[0236] The sequence and structure of Cas9 nuclease are well known to those of skill in the art (see, e.g., "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Ferretti et al., Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E. et al., Nature 471:602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M. et al., Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including, but not limited to, S. pyogenes and S. thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those of skill in the art based on the present disclosure, including Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737, the entire contents of which are incorporated herein by reference.
[0237] High-fidelity Cas9 domain Some aspects of the present disclosure provide high-fidelity Cas9 domains. High-fidelity Cas9 domains are known in the art and are described, for example, in Kleinstiver, BP et al., "High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects." Nature 529, 490-495 (2016); and Slaymaker, IM et al., "Rationally engineered Cas9 nucleases with improved specificity." Science 351, 84-88 (2015); the entire contents of each are incorporated herein by reference. An exemplary high-fidelity Cas9 domain is provided in the Sequence Listing as SEQ ID NO: 233. In some embodiments, the high-fidelity Cas9 domain is an engineered Cas9 domain containing one or more mutations that reduce the electrostatic interaction between the Cas9 domain and the sugar-phosphate backbone of DNA compared to the corresponding wild-type Cas9 domain. High-fidelity Cas9 domains with reduced electrostatic interactions with the sugar-phosphate backbone of DNA result in fewer off-target effects. In some embodiments, the Cas9 domain (e.g., a wild-type Cas9 domain (SEQ ID NOs: 197 and 200)) comprises one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of DNA. In some embodiments, the Cas9 domain comprises one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of DNA by at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, or at least 70%.
[0238] In some embodiments, any of the Cas9 fusion proteins or complexes provided herein contain one or more of the following mutations: D10A, N497X, R661X, Q695X, and / or Q926X, or corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the high-fidelity Cas9 enzyme is SpCas9(K855A), eSpCas9(1.1), SpCas9-HF1, or a hyper-precision Cas9 variant (HypaCas9). In some embodiments, the modified Cas9 eSpCas9(1.1) contains an alanine substitution that weakens the interaction between the HNH / RuvC groove and non-target DNA strands, preventing strand separation and cleavage at off-target sites. Similarly, SpCas9-HF1 reduces off-target editing through an alanine substitution that disrupts the interaction between Cas9 and the DNA phosphate backbone. HypaCas9 contains mutations in the REC3 domain (SpCas9 N692A / M694A / Q695A / H698A) that improve Cas9 proofreading and target discrimination. All three high-fidelity enzymes generate fewer off-target edits than wild-type Cas9.
[0239] Reduced exclusivity of Cas9 domains Typically, Cas9 proteins, such as S. pyogenes Cas9 (SpCas9), require a "protospacer adjacent motif (PAM)" or PAM-like motif, a 2-6 base pair DNA sequence immediately following the DNA sequence targeted by the Cas9 nuclease in the CRISPR bacterial adaptive immune system. The presence of an NGG PAM sequence is required to bind to specific nucleic acid regions, where the "N" in "NGG" is adenosine (A), thymidine (T), or cytosine (C), and the "G" is guanosine. This can limit their ability to edit desired bases in the genome. In some embodiments, the base-editing fusion proteins or complexes provided herein may need to be placed at a precise location, e.g., in the region containing the target base, upstream of the PAM. See, e.g., Komor, AC et al., "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage," Nature 533, 420-424 (2016), the entire contents of which are hereby incorporated by reference. Exemplary polypeptide sequences of spCas9 proteins capable of binding to PAM sequences are provided in the Sequence Listing as SEQ ID NOS: 197, 201, and 234-237. Thus, in some embodiments, any of the fusion proteins or complexes provided herein can contain a Cas9 domain capable of binding to a nucleotide sequence that does not contain a standard (e.g., NGG) PAM sequence. Cas9 domains that bind to non-standard PAM sequences have been described in the art and would be apparent to one of skill in the art.For example, Cas9 domains that bind to non-canonical PAM sequences are described in Kleinstiver, B.P. et al., "Engineered CRISPR-Cas9 nucleases with altered PAM specificities," Nature 523, 481-485 (2015); and Kleinstiver, B.P. et al., "Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition," Nature Biotechnology 33, 1293-1298 (2015); the entire contents of each are hereby incorporated by reference herein.
[0240] Nickase In some embodiments, the polynucleotide-programmable nucleotide-binding domain comprises a nickase domain. As used herein, the term "nickase" refers to a polynucleotide-programmable nucleotide-binding domain, including a nuclease domain, that can cleave only one of the two strands in a double-stranded nucleic acid molecule (e.g., DNA). In some embodiments, a nickase can be derived from a fully catalytically active (e.g., native) form of the polynucleotide-programmable nucleotide-binding domain by introducing one or more mutations into the active form of the polynucleotide-programmable nucleotide-binding domain. For example, if the polynucleotide-programmable nucleotide-binding domain comprises a nickase domain derived from Cas9, the Cas9-derived nickase domain can comprise a D10A mutation and a histidine at position 840. In such an embodiment, residue H840 retains catalytic activity, thereby enabling cleavage of one strand of a nucleic acid duplex. In another example, the Cas9-derived nickase domain comprises an H840A mutation, but the amino acid residue at position 10 remains D. In some embodiments, a nickase can be derived from a fully catalytically active (e.g., native) form of a polynucleotide-programmable nucleotide-binding domain by removing all or a portion (e.g., a functional portion) of a nuclease domain that is not required for nickase activity. For example, if the polynucleotide-programmable nucleotide-binding domain comprises a nickase domain derived from Cas9, the nickase domain from Cas9 can comprise a deletion of all or a portion (e.g., a functional portion) of the RuvC domain or the HNH domain.
[0241] In some embodiments, the wild-type Cas9 corresponds to or comprises the following amino acid sequence:
[0242] [Table D]
[0243] In some embodiments, the strand of a nucleic acid duplex target polynucleotide sequence that is cleaved by a base editor comprising a nickase domain (e.g., a nickase domain from Cas9, a nickase domain from Cas12) is the strand that is not edited by the base editor (i.e., the strand cleaved by the base editor is opposite the strand containing the base to be edited). In other embodiments, a base editor comprising a nickase domain (e.g., a nickase domain from Cas9, a nickase domain from Cas12) can cleave the strand of a DNA molecule that is targeted for editing. In such embodiments, the non-target strand is not cleaved.
[0244] In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., the Cas9 is a nickase referred to as a "nCas9" protein (for "nickase" Cas9). The Cas9 nickase can be a Cas9 protein that can cleave only one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase cleaves the target strand of the double-stranded nucleic acid molecule, meaning that the Cas9 nickase cleaves the strand that is base-paired (complementary) to the gRNA (e.g., sgRNA) that binds to the Cas9. In some embodiments, the Cas9 nickase comprises a D10A mutation and has a histidine at position 840. In some embodiments, the Cas9 nickase cleaves the non-target, unbase-edited strand of the double-stranded nucleic acid molecule, meaning that the Cas9 nickase cleaves the strand that is not base-paired to the gRNA (e.g., sgRNA) that binds to the Cas9. In some embodiments, the Cas9 nickase comprises an H840A mutation and has an aspartic acid residue at position 10, or a corresponding mutation. In some embodiments, the Cas9 nickase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the Cas9 nickases provided herein. Additional suitable Cas9 nickases will be apparent to those of skill in the art based on this disclosure and knowledge in the art, and are within the scope of this disclosure.
[0245] The amino acid sequence of an exemplary catalytic Cas9 nickase (nCas9) is as follows: The Cas9 nuclease has two functional endonuclease domains: RuvC and HNH. Upon binding to its target, Cas9 undergoes a conformational change, positioning the nuclease domains to cleave opposite strands of the target DNA. The end result of Cas9-mediated DNA cleavage is a double-strand break (DSB) within the target DNA (approximately 3–4 nucleotides upstream of the PAM sequence). The resulting DSB is then repaired by one of two general repair pathways: (1) the efficient but error-prone non-homologous end joining (NHEJ) pathway; or (2) the less efficient but high-fidelity homology-directed repair (HDR) pathway.
[0246] In some embodiments, the Cas9 is a modified Cas9. A given gRNA targeting sequence may have additional sites of partial homology throughout the genome. These sites are called off-targets and must be considered when designing the gRNA. In addition to optimizing the gRNA design, Cas9 can also be modified to enhance CRISPR specificity. Cas9 combines the activities of two nuclease domains, RuvC and HNH, to generate double-strand breaks (DSBs). Cas9 nickase, a D10A mutant of SpCas9, retains one nuclease domain and generates DNA nicks instead of DSBs. The nickase system can also be combined with HDR-mediated gene editing for specific gene editing.
[0247] Catalytically inactive nucleases Also provided herein are base editors comprising a catalytically inactive (i.e., unable to cleave a target polynucleotide sequence) polynucleotide-programmable nucleotide-binding domain. As used herein, the terms "catalytically inactive" and "nuclease-inactive" are used interchangeably to refer to a polynucleotide-programmable nucleotide-binding domain having one or more mutations and / or deletions that result in an inability to cleave a strand of nucleic acid. In some embodiments, a catalytically inactive polynucleotide-programmable nucleotide-binding domain base editor may lack nuclease activity as a result of specific point mutations in one or more nuclease domains. For example, in the case of a base editor comprising a Cas9 domain, Cas9 may contain both the D10A and H840A mutations. Such mutations inactivate both nuclease domains, thereby resulting in loss of nuclease activity. In other embodiments, a catalytically inactive polynucleotide-programmable nucleotide-binding domain comprises one or more deletions of all or a portion (e.g., functional portions) of a catalytic domain (e.g., RuvC1 and / or HNH domain). In further embodiments, the catalytically inactive polynucleotide programmable nucleotide binding domain comprises a point mutation (e.g., D10A or H840A) and a deletion of all or a portion (e.g., a functional portion) of the nuclease domain. dCas9 domains are known in the art and are described, for example, in Qi et al., "Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression." Cell. 2013;152(5):1173-83, the entire contents of which are incorporated herein by reference.
[0248] Additional suitable nuclease-inactive dCas9 domains will be apparent to those skilled in the art based on this disclosure and knowledge in the art and are within the scope of this disclosure. Exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013;31(9):833-838, the entire contents of which are incorporated herein by reference).
[0249] In some embodiments, the dCas9 corresponds to, or partially or entirely comprises, a Cas9 amino acid sequence with one or more mutations that inactivate Cas9 nuclease activity. In some embodiments, the nuclease-inactive dCas9 domain comprises a D10X mutation and an H840X mutation in the amino acid sequence described herein, or a corresponding mutation in any of the amino acid sequences provided herein, where X is any amino acid change. In some embodiments, the nuclease-inactive dCas9 domain comprises a D10A mutation and an H840A mutation in the amino acid sequence described herein, or a corresponding mutation in any of the amino acid sequences provided herein. In some embodiments, the nuclease-inactive Cas9 domain comprises the amino acid sequence described in the cloning vector pPlatTET-gRNA2 (accession number BAV54124).
[0250] In some embodiments, the variant Cas9 protein is capable of cleaving the complementary strand of a guide target sequence but has a reduced ability to cleave the non-complementary strand of a double-stranded guide target sequence. For example, the variant Cas9 protein may have a mutation (amino acid substitution) that reduces the function of the RuvC domain. As a non-limiting example, in some embodiments, the variant Cas9 protein has D10A (aspartic acid to alanine at amino acid position 10), and thus is capable of cleaving the complementary strand of a double-stranded guide target sequence but has a reduced ability to cleave the non-complementary strand of a double-stranded guide target sequence (thus generating a single-strand break (SSB) instead of a double-strand break (DSB) when the variant Cas9 protein cleaves a double-stranded target nucleic acid) (see, e.g., Jinek et al., Science. 2012 Aug. 17;337(6096):816-21).
[0251] In some embodiments, the variant Cas9 protein can cleave the non-complementary strand of a double-stranded guide target sequence, but has a reduced ability to cleave the complementary strand of the guide target sequence. For example, the variant Cas9 protein can have a mutation (amino acid substitution) that reduces the function of the HNH domain (RuvC / HNH / RuvC domain motif). As a non-limiting example, in some embodiments, the variant Cas9 protein has an H840A (histidine to alanine at amino acid position 840) mutation, which allows it to cleave the non-complementary strand of a guide target sequence, but has a reduced ability to cleave the complementary strand of a guide target sequence (thus generating an SSB rather than a DSB when the variant Cas9 protein cleaves a double-stranded guide target sequence). Such a Cas9 protein has a reduced ability to cleave a guide target sequence (e.g., a single-stranded guide target sequence), but retains the ability to bind to the guide target sequence (e.g., a single-stranded guide target sequence).
[0252] As another non-limiting example, in some embodiments, a variant Cas9 protein has a W476A and a W1126A mutation, resulting in the polypeptide having a reduced ability to cleave target DNA. Such a Cas9 protein has a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retains the ability to bind to target DNA (e.g., single-stranded target DNA).
[0253] As another non-limiting example, in some embodiments, a variant Cas9 protein has P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, resulting in a polypeptide with a reduced ability to cleave target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA).
[0254] As another non-limiting example, in some embodiments, a variant Cas9 protein has H840A, W476A, and W1126A mutations, resulting in a polypeptide with reduced ability to cleave target DNA. Such Cas9 proteins have reduced ability to cleave target DNA (e.g., single-stranded target DNA) but retain the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some embodiments, a variant Cas9 protein has H840A, D10A, W476A, and W1126A mutations, resulting in a polypeptide with reduced ability to cleave target DNA. Such Cas9 proteins have reduced ability to cleave target DNA (e.g., single-stranded target DNA) but retain the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, a variant Cas9 restores a catalytic His residue at position 840 of the Cas9 HNH domain (A840H).
[0255] As another non-limiting example, in some embodiments, a variant Cas9 protein has H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, thereby resulting in a reduced ability of the polypeptide to cleave target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some embodiments, a variant Cas9 protein has D10A, H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, thereby resulting in a reduced ability of the polypeptide to cleave target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA) but retain the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, when a variant Cas9 protein contains the W476A and W1126A mutations, or when a variant Cas9 protein contains the P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, the variant Cas9 protein does not efficiently bind to PAM sequences. Thus, in some such embodiments, when such a variant Cas9 protein is used in a binding method, the method does not require a PAM sequence. In other words, in some embodiments, when such a variant Cas9 protein is used in a binding method, the method can include a guide RNA, but the method can be performed without a PAM sequence (and binding specificity is therefore provided by the targeting segment of the guide RNA). Other residues can be mutated to achieve the above effects (i.e., to inactivate one or other nuclease moieties). By way of non-limiting example, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 can be altered (i.e., substituted). Mutations other than alanine substitutions are also suitable.
[0256] In some embodiments, variant Cas9 proteins with reduced catalytic activity (e.g., when the Cas9 protein has a D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 mutation, e.g., D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A), the variant Cas9 protein can still bind to target DNA in a site-specific manner (because it is still guided to the target DNA sequence by the guide RNA), so long as it retains the ability to interact with the guide RNA.
[0257] In some embodiments, the variant Cas protein can be spCas9, spCas9-VRQR, spCas9-VRER, xCas9(sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL.
[0258] In some embodiments, the Cas9 domain is a Cas9 domain derived from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is a nuclease-active SaCas9, a nuclease-inactive SaCas9 (SaCas9d), or a SaCas9 nickase (SaCas9n). In some embodiments, the SaCas9 comprises an N579A mutation or a corresponding mutation in any of the amino acid sequences provided in the Sequence Listing submitted herewith.
[0259] In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to a nucleic acid sequence with a non-canonical PAM. In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to a nucleic acid sequence with an NNGRRT or NNGRRV PAM sequence. In some embodiments, the SaCas9 domain comprises one or more of the following mutations: E781X, N967X, and R1014X, or a corresponding mutation in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SaCas9 domain comprises one or more of the following mutations: E781K, N967K, and R1014H, or a corresponding mutation in any of the amino acid sequences provided herein. In some embodiments, the SaCas9 domain comprises the following mutations: E781K, N967K, or R1014H, or a corresponding mutation in any of the amino acid sequences provided herein.
[0260] In some embodiments, one of the Cas9 domains present in the fusion protein or complex may be replaced with a guide nucleotide sequence-programmable DNA binding protein domain that does not require a PAM sequence. In some embodiments, the Cas9 is SaCas9. Residue A579 of SaCas9 can be mutated from N579 to generate SaCas9 nickase. Residues K781, K967, and H1014 can be mutated from E781, N967, and R1014 to generate SaKKH Cas9.
[0261] In some embodiments, a modified SpCas9 was used containing the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (SpCas9-MQKFRAER) and with altered specificity for PAM5'-NGC-3'.
[0262] An alternative to S. pyogenes Cas9 would be RNA-guided endonucleases from the Cpf1 family, which exhibit cleavage activity in mammalian cells. CRISPR from Prevotella and Francisella 1 (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. Cpf1 is an RNA-guided endonuclease in the Class II CRISPR / Cas system. This adaptive immune mechanism is found in bacteria of the genera Prevotella and Francisella. The Cpf1 gene is associated with the CRISPR locus and encodes an endonuclease that uses guide RNA to locate and cleave viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9, overcoming some of the limitations of the CRISPR / Cas9 system. Unlike Cas9 nuclease, Cpf1-mediated DNA cleavage results in double-strand breaks with short 3' overhangs. The staggered cleavage pattern of Cpf1 may open up the possibility of directed gene transfer, which could increase the efficiency of gene editing, similar to conventional restriction enzyme cloning. Like the Cas9 variants and orthologs mentioned above, Cpf1 can also expand the number of sites that can be targeted by CRISPR, even to AT-rich regions or AT-rich genomes lacking the NGG PAM site preferred by SpCas9. The Cpf1 locus contains a mixed alpha / beta domain, RuvC-I, followed by a helical region, RuvC-II, and a zinc finger-like domain. The Cpf1 protein possesses a RuvC-like endonuclease domain similar to the RuvC domain of Cas9.
[0263] Furthermore, unlike Cas9, Cpf1 lacks the HNH endonuclease domain, and the N-terminus of Cpf1 lacks the alpha-helix recognition lobe of Cas9. The Cpf1 CRISPR-Cas domain architecture indicates that Cpf1 is functionally unique and falls within the class 2, type V CRISPR system. The Cpf1 locus encodes Cas1, Cas2, and Cas4 proteins, which are more similar to type I and type III CRISPR systems than type II systems. Functional Cpf1 does not require trans-activating CRISPR RNA (tracrRNA); only CRISPR (crRNA) is required. Cpf1 is not only smaller than Cas9, but also possesses a smaller sgRNA molecule (approximately half the nucleotides of Cas9), which is advantageous for genome editing. The Cpf1-crRNA complex cleaves target DNA or RNA by recognizing the protospacer adjacent motif 5'-YTN-3' or 5'-TTN-3', in contrast to the G-rich PAM targeted by Cas9. After recognizing the PAM, Cpf1 introduces sticky-end-like DNA double-strand breaks with 4- or 5-nucleotide overhangs.
[0264] In some embodiments, the Cas9 is a Cas9 variant with specificity for an altered PAM sequence. In some embodiments, additional Cas9 variants and PAM sequences are described in Miller, SM et al., "Continuous evolution of SpCas9 variants compatible with non-G PAMs," Nat. Biotechnol. (2020), the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 variant does not have a specific PAM requirement. In some embodiments, the Cas9 variant, e.g., the SpCas9 variant, has specificity for the NRNH PAM, where R is A or G and H is A, C, or T. In some embodiments, the SpCas9 variant has specificity for the PAM sequence AAA, TAA, CAA, GAA, TAT, GAT, or CAC. In some embodiments, the SpCas9 variant comprises an amino acid substitution at position 1114, 1134, 1135, 1137, 1139, 1151, 1180, 1188, 1211, 1218, 1219, 1221, 1249, 1256, 1264, 1290, 1318, 1317, 1320, 1321, 1323, 1332, 1333, 1335, 1337, or 1339, or a corresponding position thereof. In some embodiments, the SpCas9 variant comprises an amino acid substitution at position 1114, 1135, 1218, 1219, 1221, 1249, 1320, 1321, 1323, 1332, 1333, 1335, or 1337, or a corresponding position thereof. In some embodiments, the SpCas9 variant comprises an amino acid substitution at positions 1114, 1134, 1135, 1137, 1139, 1151, 1180, 1188, 1211, 1219, 1221, 1256, 1264, 1290, 1318, 1317, 1320, 1323, 1333, or their corresponding positions. In some embodiments, the SpCas9 variant comprises an amino acid substitution at positions 1114, 1131, 1135, 1150, 1156, 1180, 1191, 1218, 1219, 1221, 1227, 1249, 1253, 1286, 1293, 1320, 1321, 1332, 1335, 1339, or their corresponding positions.In some embodiments, the SpCas9 variant comprises an amino acid substitution at positions 1114, 1127, 1135, 1180, 1207, 1219, 1234, 1286, 1301, 1332, 1335, 1337, 1338, 1349, or their corresponding positions. Exemplary amino acid substitutions and PAM specificities of SpCas9 variants are shown in Tables 3A-3D.
[0265] [Table 3A]
[0266] [Table 3B]
[0267] [Table 3C]
[0268] [Table 3D]
[0269] Additional exemplary Cas9 (e.g., SaCas9) polypeptides with altered PAM recognition are described in Kleinstiver et al., "Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition," Nature Biotechnology, 33:1293-1298 (2015) DOI: 10.1038 / nbt.3404, the disclosure of which is incorporated herein by reference in its entirety for all purposes. In some embodiments, Cas9 variants (e.g., SaCas9 variants) containing one or more of the following changes are associated with specific or increased editing activity for NNNRRT or NNHRRT PAM sequences compared to a reference polypeptide (e.g., SaCas9), where N represents any nucleotide, H represents any nucleotide other than G (i.e., "non-G"), and R represents a purine. In embodiments, the Cas9 variant (e.g., a SaCas9 variant) comprises the following changes: E782K, N968K, and R1015H, or E782K, K929R, and R1015H.
[0270] In some embodiments, nucleic acid programmable DNA binding protein (napDNAbp) is the single effector of microbial CRISPR-Cas system.Single effectors of microbial CRISPR-Cas system include but are not limited to Cas9, Cpf1, Cas12b / C2c1 and Cas12c / C2c3.Typically, microbial CRISPR-Cas system is divided into class 1 system and class 2 system.Class 1 system has multi-subunit effector complex, while class 2 system has single protein effector.For example, Cas9 and Cpf1 are class 2 effectors. In addition to Cas9 and Cpf1, three distinct Class 2 CRISPR-Cas systems (Cas12b / C2c1 and Cas12c / C2c3) have been described by Shmakov et al., "Discovery and Functional Characterization of Diverse Class 2 CRISPR-Cas Systems," Mol. Cell, November 5, 2015;60(3):385-397, the entire contents of which are hereby incorporated by reference. Two effectors in the system, Cas12b / C2c1 and Cas12c / C2c3, contain a RuvC-like endonuclease domain related to Cpf1. The third system contains an effector with two predicted HEPN RNase domains. Unlike CRISPR RNA production by Cas12b / C2c1, mature CRISPR RNA production is independent of tracrRNA. Cas12b / C2c1 depends on both CRISPR RNA and tracrRNA for DNA cleavage.
[0271] In some embodiments, the napDNAbp is a circular permutation (eg, SEQ ID NO: 238).
[0272] The crystal structure of Alicyclobacillus acidoterrastris Cas12b / C2c1 (AacC2c1) in complex with a chimeric single-molecule guide RNA (sgRNA) has been reported. See, e.g., Liu et al., "C2c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism," Mol. Cell, 2017 Jan. 19;65(2):310-322, the entire contents of which are hereby incorporated by reference. A crystal structure has also been reported of Alicyclobacillus acidoterrastris C2c1 bound to target DNA as a ternary complex. See, e.g., Yang et al., "PAM-dependent Target DNA Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease," Cell, 2016 Dec. 15;167(7):1814-1828, the entire contents of which are hereby incorporated by reference. Both target and non-target DNA strands capture catalytically competent conformations of AacC2c1, which are independently positioned within a single RuvC catalytic pocket, resulting in a staggered seven-nucleotide cleavage of the target DNA upon Cas12b / C2c1-mediated cleavage. Structural comparison of the Cas12b / C2c1 ternary complex with previously identified Cas9 and Cpf1 counterparts demonstrates the diversity of mechanisms employed by the CRISPR-Cas9 system.
[0273] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) of any of the fusion proteins or complexes provided herein can be a Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a Cas12b / C2c1 protein. In some embodiments, the napDNAbp is a Cas12c / C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the napDNAbp sequences provided herein. It should be understood that Cas12b / C2c1 or Cas12c / C2c3 from other bacterial species may also be used in accordance with the present disclosure.
[0274] In some embodiments, napDNAbp refers to Cas12c. In some embodiments, the Cas12c protein is Cas12c1 (SEQ ID NO: 239) or a variant of Cas12c1. In some embodiments, the Cas12 protein is Cas12c2 (SEQ ID NO: 240) or a variant of Cas12c2. In some embodiments, the Cas12 protein is a Cas12c protein from Oleiphilus sp. HI0009 (i.e., OspCas12c; SEQ ID NO: 241) or a variant of OspCas12c. These Cas12c molecules are described in Yan et al., "Functionally Diverse Type V CRISPR-Cas Systems," Science, 2019 Jan. 4;363:88-91, the entire contents of which are hereby incorporated by reference herein. In some embodiments, the napDNAbp comprises an amino acid sequence at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring Cas12c1, Cas12c2, or OspCas12c protein. In some embodiments, the napDNAbp is a naturally occurring Cas12c1, Cas12c2, or OspCas12c protein. In some embodiments, the napDNAbp comprises an amino acid sequence at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any Cas12c1, Cas12c2, or OspCas12c protein described herein. It should be understood that Cas12c1, Cas12c2, or OspCas12c from other bacterial species may also be used in accordance with the present disclosure.
[0275] In some embodiments, napDNAbp refers to Cas12g, Cas12h, or Cas12i, e.g., as described in Yan et al., "Functionally Diverse Type V CRISPR-Cas Systems," Science, January 4, 2019; 363:88-91; the entire contents of each are hereby incorporated by reference. Exemplary Cas12g, Cas12h, and Cas12i polypeptide sequences are provided in the Sequence Listing as SEQ ID NOs: 242-245. By aggregating over 10 terabytes of sequence data, a new classification of type V Cas proteins, including Cas12g, Cas12h, and Cas12i, was identified that shows weak similarity to previously characterized class V proteins. In some embodiments, the Cas12 protein is Cas12g or a variant of Cas12g. In some embodiments, the Cas12 protein is Cas12h or a variant of Cas12h. In some embodiments, the Cas12 protein is Cas12i or a variant of Cas12i. It should be understood that other RNA-guided DNA-binding proteins may also be used as napDNAbp and are within the scope of the present disclosure. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring Cas12g, Cas12h, or Cas12i protein. In some embodiments, the napDNAbp is a naturally occurring Cas12g, Cas12h, or Cas12i protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any Cas12g, Cas12h, or Cas12i protein described herein.It should be understood that Cas12g, Cas12h, or Cas12i from other bacterial species may also be used in accordance with the present disclosure. In some embodiments, the Cas12i is Cas12i1 or Cas12i2.
[0276] In some embodiments, the nucleic acid-programmable DNA-binding protein (napDNAbp) of any of the fusion proteins or complexes provided herein can be a Cas12j / CasΦ protein. Cas12j / CasΦ is described in Pausch et al., "CRISPR-CasΦ from Huge Phages Is a Hypercompact Genome Editor," Science, July 17, 2020, Vol. 369, Issue 6501, pp. 333-337, which is incorporated herein by reference in its entirety. In some embodiments, the napDNAbp comprises an amino acid sequence at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring Cas12j / CasΦ protein. In some embodiments, the napDNAbp is a naturally occurring Cas12j / CasΦ protein. In some embodiments, the napDNAbp is a nuclease-inactive ("dead") Cas12j / CasΦ protein. It should be understood that Cas12j / CasΦ from other species may also be used in accordance with the present disclosure.
[0277] Fusion proteins or complexes with internal insertions Provided herein are fusion proteins or complexes comprising a heterologous polypeptide fused to a nucleic acid programmable nucleic acid binding protein, e.g., napDNAbp. The heterologous polypeptide can be a polypeptide not found in the native or wild-type napDNAbp polypeptide sequence. The heterologous polypeptide can be fused to the napDNAbp at the C-terminus of the napDNAbp, the N-terminus of the napDNAbp, or inserted at an internal position of the napDNAbp. In some embodiments, the heterologous polypeptide is a deaminase (e.g., cytidine or adenosine deaminase) or a functional fragment thereof. For example, the fusion protein can comprise a deaminase flanked by N-terminal and C-terminal fragments of a Cas9 or Cas12 (e.g., Cas12b / C2c1) polypeptide. In some embodiments, the cytidine deaminase is an APOBEC deaminase (e.g., APOBEC1). In some embodiments, the adenosine deaminase is TadA (e.g., TadA * 7.10 or TadA * 8). In some embodiments, TadA is TadA * 8 or TadA * 9. The TadA sequences described herein (e.g., TadA7.10 or TadA * 8) is a suitable deaminase for the above fusion protein or complex.
[0278] In some embodiments, the fusion protein has the following structure: NH2-[N-terminal fragment of napDNAbp]-[deaminase]-[C-terminal fragment of napDNAbp]-COOH; NH2-[N-terminal fragment of Cas9]-[adenosine deaminase]-[C-terminal fragment of Cas9]-COOH; NH2-[N-terminal fragment of Cas12]-[adenosine deaminase]-[C-terminal fragment of Cas12]-COOH; NH2-[N-terminal fragment of Cas9]-[cytidine deaminase]-[C-terminal fragment of Cas9]-COOH; NH2-[N-terminal fragment of Cas12]-[cytidine deaminase]-[C-terminal fragment of Cas12]-COOH where each instance of "]-[" is an optional presence of a linker (i.e., the linker is optionally present).
[0279] The deaminase can be a circularly permuted deaminase. For example, the deaminase can be a circularly permuted adenosine deaminase. In some embodiments, the deaminase is a circularly permuted TadA that is circularly permuted at amino acid residues 116, 136, or 65 numbered in the TadA reference sequence.
[0280] A fusion protein or complex may contain more than one deaminase. A fusion protein or complex may contain, for example, one, two, three, four, five, or more deaminases. In some embodiments, a fusion protein or complex contains one or two deaminases. The two or more deaminases in a fusion protein or complex may be adenosine deaminase, cytidine deaminase, or a combination thereof. The two or more deaminases may be homodimers or heterodimers. Two or more deaminases may be inserted in tandem into the napDNAbp. In some embodiments, the two or more deaminases may not be present in tandem in the napDNAbp.
[0281] In some embodiments, the napDNAbp in the fusion protein or complex is a Cas9 polypeptide or a fragment thereof. The Cas9 polypeptide can be a variant Cas9 polypeptide. In some embodiments, the Cas9 polypeptide is a Cas9 nickase (nCas9) polypeptide or a fragment thereof. In some embodiments, the Cas9 polypeptide is a nuclease-inactive Cas9 (dCas9) polypeptide or a fragment thereof. The Cas9 polypeptide in the fusion protein or complex can be a full-length Cas9 polypeptide. In some cases, the Cas9 polypeptide in the fusion protein or complex may not be a full-length Cas9 polypeptide. The Cas9 polypeptide can be truncated, for example, at the N-terminus or C-terminus compared to a naturally occurring Cas9 protein. The Cas9 polypeptide can be a circularly permuted Cas9 protein. The Cas9 polypeptide can be a fragment, portion, or domain of a Cas9 polypeptide, but it can still bind to a target polynucleotide and a guide nucleic acid sequence.
[0282] In some embodiments, the Cas9 polypeptide is Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), or a fragment or variant of any of the Cas9 polypeptides described herein.
[0283] In some embodiments, the fusion protein comprises an adenosine deaminase domain and a cytidine deaminase domain inserted into Cas9. In some embodiments, adenosine deaminase is fused into Cas9, and cytidine deaminase is fused to the C-terminus. In some embodiments, adenosine deaminase is fused into Cas9, and cytidine deaminase is fused to the N-terminus. In some embodiments, cytidine deaminase is fused into Cas9, and adenosine deaminase is fused to the C-terminus. In some embodiments, cytidine deaminase is fused into Cas9, and adenosine deaminase is fused to the N-terminus.
[0284] An exemplary structure of a fusion protein having adenosine deaminase and cytidine deaminase and Cas9 is as follows: NH2-[Cas9(adenosine deaminase)]-[cytidine deaminase]-COOH; NH2-[cytidine deaminase]-[Cas9(adenosine deaminase)]-COOH; NH2-[Cas9(cytidine deaminase)]-[adenosine deaminase]-COOH; or NH2-[adenosine deaminase]-[Cas9 (cytidine deaminase)]-COOH It is provided as follows.
[0285] In some embodiments, the "-" used in the general structures above indicates the presence of an optional linker.
[0286] In various embodiments, the catalytic domain has a DNA modifying activity (e.g., deaminase activity), such as adenosine deaminase activity. In some embodiments, the adenosine deaminase is TadA (e.g., TadA * 7.10). In some embodiments, TadA is TadA * 8. In some embodiments, TadA * In some embodiments, TadA is fused to Cas9 and cytidine deaminase is fused to the C-terminus. * In some embodiments, 8 is fused to Cas9 and cytidine deaminase is fused to the N-terminus. In some embodiments, cytidine deaminase is fused to Cas9 and TadA is fused to the N-terminus. * In some embodiments, cytidine deaminase is fused to Cas9 and TadA is fused to the C-terminus. * 8 is fused to the N-terminus. * An exemplary structure of a fusion protein having 8 and cytidine deaminase and Cas9 is as follows: NH2-[Cas9(TadA * 8)]-[cytidine deaminase]-COOH; NH2-[cytidine deaminase]-[Cas9(TadA * 8)]-COOH; NH2-[Cas9(cytidine deaminase)]-[TadA * 8]-COOH; or NH2-[TadA * 8]-[Cas9 (cytidine deaminase)]-COOH It is provided as follows.
[0287] In some embodiments, the "-" used in the general structures above indicates the presence of an optional linker.
[0288] A heterologous polypeptide (e.g., a deaminase) can be inserted into the napDNAbp (e.g., Cas9 or Cas12 (e.g., Cas12b / C2c1)) at an appropriate position, e.g., so that the napDNAbp retains the ability to bind to a target polynucleotide and a guide nucleic acid. The deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted into the napDNAbp without impairing the function of the deaminase (e.g., base editing activity) or the function of the napDNAbp (e.g., the ability to bind to a target nucleic acid and a guide nucleic acid). The deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted into the napDNAbp, e.g., in a denatured region or a region containing a high temperature factor or B factor, as shown in crystallographic studies. Regions of less ordered, denatured, or amorphous proteins, such as solvent-exposed regions and loops, can be used for insertion without impairing structure or function. Deaminases (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted into the napDNAbp of flexible loop regions or solvent-exposed regions. In some embodiments, deaminases (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) are inserted into the flexible loops of Cas9 or Cas12b / C2c1 polypeptides.
[0289] In some embodiments, the insertion location of the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is determined by B-factor analysis of the crystal structure of the Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted into a region of the Cas9 polypeptide that contains a higher-than-average B-factor (e.g., a higher B-factor compared to the entire protein or a protein domain containing the denatured region). The B-factor or temperature factor can indicate variation of atoms from their average positions (e.g., as a result of temperature-dependent atomic vibrations or static disorder in the crystal lattice). A high B-factor (e.g., a higher-than-average B-factor) of backbone atoms can indicate a region with relatively high local mobility. Such a region can be used to insert the deaminase without compromising structure or function. A deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted at a position having a residue with a Cα atom that has a B factor that is 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, or more than 200% higher than the average B factor for all proteins. A deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted at a position having a residue with a Cα atom that has a B-factor that is 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, or greater than 200% higher than the average B-factor for the Cas9 protein domain containing the residue. Positions in the Cas9 polypeptide containing higher than average B factors can include, for example, residues 768, 792, 1052, 1015, 1022, 1026, 1029, 1067, 1040, 1054, 1068, 1246, 1247, and 1248, as numbered in the above Cas9 reference sequence.Regions of the Cas9 polypeptide containing higher than average B factors can include, for example, residues 792-872, 792-906, and 2-791, as numbered in the Cas9 reference sequence above.
[0290] A heterologous polypeptide (e.g., a deaminase) can be inserted into the napDNAbp at an amino acid residue selected from the group consisting of 768, 791, 792, 1015, 1016, 1022, 1023, 1026, 1029, 1040, 1052, 1054, 1067, 1068, 1069, 1246, 1247, and 1248, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the heterologous polypeptide is inserted between amino acid positions 768-769, 791-792, 792-793, 1015-1016, 1022-1023, 1026-1027, 1029-1030, 1040-1041, 1052-1053, 1054-1055, 1067-1068, 1068-1069, 1247-1248, or 1248-1249, as numbered in the above Cas9 reference sequence, or their corresponding amino acid positions. In some embodiments, the heterologous polypeptide is inserted between amino acid positions 769-770, 792-793, 793-794, 1016-1017, 1023-1024, 1027-1028, 1030-1031, 1041-1042, 1053-1054, 1055-1056, 1068-1069, 1069-1070, 1248-1249, or 1249-1250, as numbered in the above Cas9 reference sequence, or their corresponding amino acid positions. In some embodiments, the heterologous polypeptide replaces an amino acid residue selected from the group consisting of 768, 791, 792, 1015, 1016, 1022, 1023, 1026, 1029, 1040, 1052, 1054, 1067, 1068, 1069, 1246, 1247, and 1248, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. It should be understood that reference to the above Cas9 reference sequence for insertion positions is for illustrative purposes.Insertions discussed herein are not limited to the Cas9 polypeptide sequences of the above Cas9 reference sequences, but include insertions at corresponding positions in variant Cas9 polypeptides, such as Cas9 nickase (nCas9), nuclease-inactive Cas9 (dCas9), Cas9 variants lacking a nuclease domain, truncated Cas9, or Cas9 domains partially or completely lacking the HNH domain.
[0291] A heterologous polypeptide (e.g., a deaminase) can be inserted into the napDNAbp at an amino acid residue selected from the group consisting of 768, 792, 1022, 1026, 1040, 1068, and 1247, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the heterologous polypeptide is inserted between amino acid positions 768-769, 792-793, 1022-1023, 1026-1027, 1029-1030, 1040-1041, 1068-1069, or 1247-1248, as numbered in the above Cas9 reference sequence, or their corresponding amino acid positions. In some embodiments, the heterologous polypeptide is inserted between amino acid positions 769-770, 793-794, 1023-1024, 1027-1028, 1030-1031, 1041-1042, 1069-1070, or 1248-1249, as numbered in the above Cas9 reference sequence, or their corresponding amino acid positions. In some embodiments, the heterologous polypeptide replaces an amino acid residue selected from the group consisting of 768, 792, 1022, 1026, 1040, 1068, and 1247, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.
[0292] A heterologous polypeptide (e.g., a deaminase) can be inserted into the napDNAbp at an amino acid residue described herein, or at a corresponding amino acid residue in another Cas9 polypeptide. In embodiments, a heterologous polypeptide (e.g., a deaminase) can be inserted into the napDNAbp at an amino acid residue selected from the group consisting of 1002, 1003, 1025, 1052-1056, 1242-1247, 1061-1077, 943-947, 686-691, 569-578, 530-539, and 1060-1077, as numbered in the above Cas9 reference sequence, or at a corresponding amino acid residue in another Cas9 polypeptide. The deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted N-terminally or C-terminally to, or replace, a residue. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminal to the residue.
[0293] In some embodiments, an adenosine deaminase (e.g., TadA) is inserted at an amino acid residue selected from the group consisting of 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, an adenosine deaminase (e.g., TadA) is inserted at residues 792-872, 792-906, or 2-791, as numbered in the above Cas9 reference sequence, or in place of the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, adenosine deaminase is inserted N-terminally to an amino acid selected from the group consisting of 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, adenosine deaminase is inserted C-terminally to an amino acid selected from the group consisting of 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, adenosine deaminase is inserted to replace an amino acid selected from the group consisting of 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.
[0294] In some embodiments, a cytidine deaminase (e.g., APOBEC1) is inserted at an amino acid residue selected from the group consisting of 1016, 1023, 1029, 1040, 1069, and 1247, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, a cytidine deaminase is inserted N-terminally to an amino acid selected from the group consisting of 1016, 1023, 1029, 1040, 1069, and 1247, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, a cytidine deaminase is inserted C-terminally to an amino acid selected from the group consisting of 1016, 1023, 1029, 1040, 1069, and 1247, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the cytidine deaminase is substituted by an insertion of an amino acid selected from the group consisting of 1016, 1023, 1029, 1040, 1069, and 1247, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.
[0295] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 768 numbered in the above Cas9 reference sequence, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminally to amino acid residue 768 numbered in the above Cas9 reference sequence, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminally to amino acid residue 768 numbered in the above Cas9 reference sequence, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, a deaminase (e.g., an adenosine deaminase, a cytidine deaminase, or an adenosine deaminase and a cytidine deaminase) is inserted into and replaces amino acid residue 768, numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.
[0296] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 791, or at amino acid residue 792, or at the corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminal to amino acid residue 791, or at amino acid residue 792, or at the corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminally to amino acid residue 791, or N-terminally to amino acid residue 792, or the corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at and replaces amino acid residue 791, or the corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence.
[0297] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1016 numbered in the above Cas9 reference sequence, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminally to amino acid residue 1016 numbered in the above Cas9 reference sequence, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminally to amino acid residue 1016 numbered in the above Cas9 reference sequence, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, a deaminase (e.g., an adenosine deaminase, a cytidine deaminase, or an adenosine deaminase and a cytidine deaminase) is inserted to replace amino acid residue 1016, numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.
[0298] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1022, or at amino acid residue 1023, or at the corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminal to amino acid residue 1022, or at amino acid residue 1023, or at the corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminal to amino acid residue 1022, or to amino acid residue 1023, or to the corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted and replaced by amino acid residue 1022, or to amino acid residue 1023, or to the corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence.
[0299] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1026, or at amino acid residue 1029, or at the corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminal to amino acid residue 1026, or at amino acid residue 1029, or at the corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminal to amino acid residue 1026, or to amino acid residue 1029, or to the corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted and replaced by amino acid residue 1026, or to amino acid residue 1029, or to the corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence.
[0300] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1040 numbered in the above Cas9 reference sequence, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminally to amino acid residue 1040 numbered in the above Cas9 reference sequence, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminally to amino acid residue 1040 numbered in the above Cas9 reference sequence, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, a deaminase (e.g., an adenosine deaminase, a cytidine deaminase, or an adenosine deaminase and a cytidine deaminase) is inserted into and replaces amino acid residue 1040, numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.
[0301] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1052, or at amino acid residue 1054, or at the corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminal to amino acid residue 1052, or at amino acid residue 1054, or at the corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminal to amino acid residue 1052, or to amino acid residue 1054, or to the corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted and replaced by amino acid residue 1052, or to amino acid residue 1054, or to the corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence.
[0302] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1067, or amino acid residue 1068, or amino acid residue 1069, as numbered in the above Cas9 reference sequence, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminally to amino acid residue 1067, or amino acid residue 1068, or amino acid residue 1069, or at the corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminally to amino acid residue 1067, or to amino acid residue 1068, or to amino acid residue 1069, or to a corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at and replaces amino acid residue 1067, or to and replaces amino acid residue 1068, or to and replaces amino acid residue 1069, or to a corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence.
[0303] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at amino acid residue 1246, or amino acid residue 1247, or amino acid residue 1248, as numbered in the above Cas9 reference sequence, or at the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted N-terminally to amino acid residue 1246, or amino acid residue 1247, or amino acid residue 1248, as numbered in the above Cas9 reference sequence, or at the N-terminal of the corresponding amino acid residue in another Cas9 polypeptide. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted C-terminally to amino acid residue 1246, or to amino acid residue 1247, or to amino acid residue 1248, or to a corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence. In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted at and replaces amino acid residue 1246, or to and replaces amino acid residue 1247, or to and replaces amino acid residue 1248, or to a corresponding amino acid residue in another Cas9 polypeptide, as numbered in the above Cas9 reference sequence.
[0304] In some embodiments, a heterologous polypeptide (e.g., a deaminase) is inserted into a flexible loop of a Cas9 polypeptide. The flexible loop portion may be selected from the group consisting of 530-537, 569-570, 686-691, 943-947, 1002-1025, 1052-1077, 1232-1247, or 1298-1300 amino acid residues numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. The flexible loop portion may be selected from the group consisting of 1-529, 538-568, 580-685, 692-942, 948-1001, 1026-1051, 1078-1231, or 1248-1297 amino acid residues numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide.
[0305] A heterologous polypeptide (e.g., adenine deaminase) can be inserted into a region of a Cas9 polypeptide corresponding to the following numbered amino acid residues in the above Cas9 reference sequence: 1017-1069, 1242-1247, 1052-1056, 1060-1077, 1002-1003, 943-947, 530-537, 568-579, 686-691, 1242-1247, 1298-1300, 1066-1077, 1052-1056, or 1060-1077, or the corresponding amino acid residues in another Cas9 polypeptide.
[0306] A heterologous polypeptide (e.g., adenine deaminase) can be inserted in place of the deleted region of the Cas9 polypeptide. The deleted region may correspond to the N-terminal or C-terminal portion of the Cas9 polypeptide. In some embodiments, the deleted region corresponds to residues 792-872, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the deleted region corresponds to residues 792-906, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the deleted region corresponds to residues 2-791, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. In some embodiments, the deleted region corresponds to residues 1017-1069, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues.
[0307] Exemplary internal fusion base editors are provided in Table 4 below:
[0308] [Table 4]
[0309] A heterologous polypeptide (e.g., a deaminase) can be inserted within a structural or functional domain of a Cas9 polypeptide. A heterologous polypeptide (e.g., a deaminase) can be inserted between two structural or functional domains of a Cas9 polypeptide. A heterologous polypeptide (e.g., a deaminase) can be inserted in place of a structural or functional domain of a Cas9 polypeptide, for example, after deleting the domain from the Cas9 polypeptide. The structural or functional domain of a Cas9 polypeptide can include, for example, RuvC I, RuvC II, RuvC III, Rec1, Rec2, PI, or HNH.
[0310] In some embodiments, the Cas9 polypeptide lacks one or more domains selected from the group consisting of the RuvC I, RuvC II, RuvC III, Rec1, Rec2, PI, or HNH domain. In some embodiments, the Cas9 polypeptide lacks the nuclease domain. In some embodiments, the Cas9 polypeptide lacks the HNH domain. In some embodiments, the Cas9 polypeptide lacks a portion of the HNH domain, thereby reducing or eliminating HNH activity of the Cas9 polypeptide. In some embodiments, the Cas9 polypeptide contains a deletion of the nuclease domain, and a deaminase is inserted to replace the nuclease domain. In some embodiments, the HNH domain is deleted and a deaminase is inserted in its place. In some embodiments, one or more of the RuvC domains are deleted and a deaminase is inserted in its place.
[0311] A fusion protein comprising a heterologous polypeptide may be flanked by N- and C-terminal fragments of napDNAbp. In some embodiments, the fusion protein comprises a deaminase flanked by N- and C-terminal fragments of a Cas9 polypeptide. The N- or C-terminal fragment can bind to a target polynucleotide sequence. The C-terminus of the N- or C-terminal fragment may comprise a portion of the flexible loop of the Cas9 polypeptide. The C-terminus of the N- or C-terminal fragment may comprise a portion of the alpha-helical structure of the Cas9 polypeptide. The N- or C-terminal fragment may comprise a DNA-binding domain. The N- or C-terminal fragment may comprise a RuvC domain. The N- or C-terminal fragment may comprise an HNH domain. In some embodiments, neither the N- or C-terminal fragment comprises an HNH domain.
[0312] In some embodiments, the C-terminus of the N-terminal Cas9 fragment comprises an amino acid that is adjacent to the target nucleobase when the fusion protein deaminates the target nucleobase.In some embodiments, the N-terminus of the C-terminal Cas9 fragment comprises an amino acid that is adjacent to the target nucleobase when the fusion protein deaminates the target nucleobase.In order to achieve proximity between the target nucleobase and the C-terminus of the N-terminal Cas9 fragment or the amino acid at the N-terminus of the C-terminal Cas9 fragment, the insertion position of different deaminase can be different.For example, the insertion position of deaminase can be selected from the group consisting of amino acid residues 1015, 1022, 1029, 1040, 1068, 1247, 1054, 1026, 768, 1067, 1248, 1052, and 1246, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residue in another Cas9 polypeptide.
[0313] The N-terminal Cas9 fragment of the fusion protein (i.e., the N-terminal Cas9 fragment adjacent to the deaminase in the fusion protein) can comprise the N-terminus of the Cas9 polypeptide. The N-terminal Cas9 fragment of the fusion protein can comprise at least about 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, or 1300 amino acids in length. The N-terminal Cas9 fragment of the fusion protein can comprise the following numbered amino acid residues in the above Cas9 reference sequence: 1-56, 1-95, 1-200, 1-300, 1-400, 1-500, 1-600, 1-700, 1-718, 1-765, 1-780, 1-906, 1-918, or 1-1100, or a sequence corresponding to the corresponding amino acid residues in another Cas9 polypeptide. An N-terminal Cas9 fragment can comprise the following numbered amino acid residues in the above Cas9 reference sequences: 1-56, 1-95, 1-200, 1-300, 1-400, 1-500, 1-600, 1-700, 1-718, 1-765, 1-780, 1-906, 1-918, or 1-1100, or a sequence comprising at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to the corresponding amino acid residues in another Cas9 polypeptide.
[0314] The C-terminal Cas9 fragment of the fusion protein (i.e., the C-terminal Cas9 fragment adjacent to the deaminase in the fusion protein) can comprise the C-terminus of the Cas9 polypeptide. The C-terminal Cas9 fragment of the fusion protein can comprise at least about 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, or 1300 amino acids in length. The C-terminal Cas9 fragment of the fusion protein can comprise the following numbered amino acid residues in the above Cas9 reference sequence: 1099-1368, 918-1368, 906-1368, 780-1368, 765-1368, 718-1368, 94-1368, or 56-1368, or a sequence corresponding to the corresponding amino acid residues in another Cas9 polypeptide. An N-terminal Cas9 fragment can comprise the following numbered amino acid residues in the above Cas9 reference sequence: 1099-1368, 918-1368, 906-1368, 780-1368, 765-1368, 718-1368, 94-1368, or 56-1368, or a sequence comprising at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to the corresponding amino acid residues in another Cas9 polypeptide.
[0315] The N-terminal Cas9 fragment and the C-terminal Cas9 fragment of the fusion protein together may not correspond to a full-length, naturally occurring Cas9 polypeptide sequence, e.g., as set forth in the Cas9 reference sequence above.
[0316] The fusion proteins or complexes described herein can provide targeted deamination accompanied by reduced deamination at non-target sites (e.g., off-target sites), such as reduced spurious deamination throughout the genome. The fusion proteins or complexes described herein can provide targeted deamination with reduced bystander deamination at non-target sites. Unwanted or off-target deamination can be reduced by at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%, for example, compared to a terminal fusion protein comprising a deaminase fused to the N-terminus or C-terminus of a Cas9 polypeptide. Unwanted or off-target deamination can be reduced by at least 1-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold, 15-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, or 100-fold, for example, compared to a terminal fusion protein comprising a deaminase fused to the N- or C-terminus of a Cas9 polypeptide.
[0317] In some embodiments, the deaminase of the fusion protein or complex (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) deaminates no more than two nucleobases within the R-loop. In some embodiments, the deaminase of the fusion protein or complex deaminates no more than three nucleobases within the R-loop. In some embodiments, the deaminase of the fusion protein or complex deaminates no more than 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleobases within the R-loop. An R-loop is a triple-stranded nucleic acid structure comprising a DNA-RNA hybrid, a DNA:DNA, or an RNA:RNA complementary structure, and associated with single-stranded DNA. As used herein, an R-loop can be formed when a target polynucleotide is contacted with a CRISPR complex or a base editing complex, and a portion of the guide polynucleotide, e.g., a guide RNA, hybridizes with and replaces a portion of the target polynucleotide, e.g., a target DNA. In some embodiments, R-loop comprises a spacer sequence and a hybridized region of target DNA complementary sequence.The R-loop region can be about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleic acid base pairs in length.In some embodiments, the R-loop region is about 20 nucleic acid base pairs in length.It should be understood that as used herein, R-loop region is not limited to the target DNA strand that hybridizes with guide polynucleotide. For example, editing of a target nucleobase in an R-loop region can be editing of a DNA strand that contains a complementary strand to the guide RNA, or editing of a DNA strand that is the opposite strand to the strand complementary to the guide RNA. In some embodiments, editing in the R-loop region involves editing a nucleobase on a non-complementary strand (protospacer strand) to the guide RNA in the target DNA sequence.
[0318] The fusion proteins or complexes described herein can effect targeted deamination in an editing window that differs from standard base editing. In some embodiments, the target nucleobase is about 1 to about 20 bases upstream of the PAM sequence in the target polynucleotide sequence. In some embodiments, the target nucleobase is about 2 to about 12 bases upstream of the PAM sequence in the target polynucleotide sequence. In some embodiments, the target nucleobase is about 1 to 9 base pairs, about 2 to 10 base pairs, about 3 to 11 base pairs, about 4 to 12 base pairs, about 5 to 13 base pairs, about 6 to 14 base pairs, about 7 to 15 base pairs, about 8 to 16 base pairs, about 9 to 17 base pairs, about 10 to 18 base pairs, about 11 to 19 base pairs, about 12 to 20 base pairs, about 1 to 7 base pairs, about 2 to 3 base pairs, about 3 to 4 base pairs, about 4 to 5 base pairs, about 5 to 6 base pairs, about 6 to 7 base pairs, about 7 to 8 base pairs, about 8 to 9 base pairs, about 9 to 10 ... 8 base pairs, about 3 to 9 base pairs, about 4 to 10 base pairs, about 5 to 11 base pairs, about 6 to 12 base pairs, about 7 to 13 base pairs, about 8 to 14 base pairs, about 9 to 15 base pairs, about 10 to 16 base pairs, about 11 to 17 base pairs, about 12 to 18 base pairs, about 13 to 19 base pairs, about 14 to 20 base pairs, about 1 to 5 base pairs, about 2 to 6 base pairs, about 3 to 7 base pairs, about 4 to 8 base pairs , about 5 to 9 base pairs, about 6 to 10 base pairs, about 7 to 11 base pairs, about 8 to 12 base pairs, about 9 to 13 base pairs, about 10 to 14 base pairs, about 11 to 15 base pairs, about 12 to 16 base pairs, about 13 to 17 base pairs, about 14 to 18 base pairs, about 15 to 19 base pairs, about 16 to 20 base pairs, about 1 to 3 base pairs, about 2 to 4 base pairs, about 3 to 5 base pairs, about 4 to 6 base pairs, The target nucleobase is separated from or upstream of the PAM sequence by about 5-7 base pairs, about 6-8 base pairs, about 7-9 base pairs, about 8-10 base pairs, about 9-11 base pairs, about 10-12 base pairs, about 11-13 base pairs, about 12-14 base pairs, about 13-15 base pairs, about 14-16 base pairs, about 15-17 base pairs, about 16-18 base pairs, about 17-19 base pairs, or about 18-20 base pairs. In some embodiments, the target nucleobase is separated from or upstream of the PAM sequence by about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more base pairs. In some embodiments, the target nucleobase is separated from or upstream of the PAM sequence by about 1, 2, 3, 4, 5, 6, 7, 8, or 9 base pairs. In some embodiments, the target nucleobase is about 2, 3, 4, or 6 base pairs upstream from the PAM sequence.
[0319] The fusion protein or complex may contain more than one heterologous polypeptide.For example, the fusion protein or complex may further contain one or more UGI domains and / or one or more nuclear localization signals.Two or more heterologous domains may be inserted in tandem.Two or more heterologous domains may be inserted in a position that is not tandem within the NapDNAbp.
[0320] The fusion protein may include a linker between the deaminase and the napDNAbp polypeptide. The linker may be a peptide linker or a non-peptide linker. For example, the linker may be XTEN, (GGGS)n (SEQ ID NO:246), (GGGGS)n (SEQ ID NO:247), (G)n, (EAAAK)n (SEQ ID NO:248), (GGS)n, or SGSETPGTSESATPES (SEQ ID NO:249). In some embodiments, the fusion protein includes a linker between the N-terminal Cas9 fragment and the deaminase. In some embodiments, the fusion protein includes a linker between the C-terminal Cas9 fragment and the deaminase. In some embodiments, the N- and C-terminal fragments of the napDNAbp are connected to the deaminase by a linker. In some embodiments, the N- and C-terminal fragments are linked to the deaminase domain without a linker. In some embodiments, the fusion protein includes a linker between the N-terminal Cas9 fragment and the deaminase, but does not include a linker between the C-terminal Cas9 fragment and the deaminase. In some embodiments, the fusion protein comprises a linker between the C-terminal Cas9 fragment and the deaminase, but does not comprise a linker between the N-terminal Cas9 fragment and the deaminase.
[0321] In some embodiments, the napDNAbp in the fusion protein or complex is a Cas12 polypeptide, such as Cas12b / C2c1, or a fragment thereof, capable of associating with a nucleic acid (e.g., gRNA) that directs Cas12 to a specific nucleic acid sequence. The Cas12 polypeptide may be a variant Cas12 polypeptide. In other embodiments, an N-terminal or C-terminal fragment of a Cas12 polypeptide comprises a nucleic acid programmable DNA-binding domain or a RuvC domain. In other embodiments, the fusion protein comprises a linker between the Cas12 polypeptide and the catalytic domain. In other embodiments, the amino acid sequence of the linker is GGSGGS (SEQ ID NO: 250) or GSSGSETPGTSESATPESSG (SEQ ID NO: 251) In other embodiments, the linker is a rigid linker. In other embodiments of the above aspect, the linker is GGAGGCTCTGGAGGAAGC (SEQ ID NO: 252) or GGCTCTTCTGGATCTGAAACACCTGGCACAAGCGAGAGCGCCACCCCTGAGAGCTCTGGC (SEQ ID NO: 253) is coded by
[0322] Fusion proteins containing heterologous catalytic domains flanked by N-terminal and C-terminal fragments of a Cas12 polypeptide are also useful for base editing in the methods described herein. Fusion proteins containing Cas12 and one or more deaminase domains, such as adenosine deaminase, or containing an adenosine deaminase domain flanked by Cas12 sequences, are also useful for highly specific and efficient base editing of target sequences. In embodiments, chimeric Cas12 fusion proteins contain heterologous catalytic domains (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) inserted into a Cas12 polypeptide. In some embodiments, the fusion protein contains an adenosine deaminase domain and a cytidine deaminase domain inserted into Cas12. In some embodiments, the adenosine deaminase is fused into Cas12, and the cytidine deaminase is fused to the C-terminus. In some embodiments, adenosine deaminase is fused within Cas12 and cytidine deaminase is fused to the N-terminus. In some embodiments, cytidine deaminase is fused within Cas12 and adenosine deaminase is fused to the C-terminus. In some embodiments, cytidine deaminase is fused within Cas12 and adenosine deaminase is fused to the N-terminus. An exemplary structure of a fusion protein having adenosine deaminase, cytidine deaminase, and Cas12 is as follows: NH2-[Cas12(adenosine deaminase)]-[cytidine deaminase]-COOH; NH2-[cytidine deaminase]-[Cas12(adenosine deaminase)]-COOH; NH2-[Cas12(cytidine deaminase)]-[adenosine deaminase]-COOH; or NH2-[adenosine deaminase]-[Cas12 (cytidine deaminase)]-COOH It is provided as follows.
[0323] In some embodiments, the "-" used in the general structures above indicates the presence of an optional linker.
[0324] In various embodiments, the catalytic domain has a DNA modifying activity (e.g., deaminase activity), such as adenosine deaminase activity. In some embodiments, the adenosine deaminase is TadA (e.g., TadA * 7.10). In some embodiments, TadA is TadA * 8. In some embodiments, TadA * In some embodiments, TadA is fused to Cas12 and cytidine deaminase is fused to the C-terminus. * In some embodiments, 8 is fused into Cas12 and cytidine deaminase is fused to the N-terminus. In some embodiments, cytidine deaminase is fused into Cas12 and TadA * In some embodiments, cytidine deaminase is fused to Cas12 and TadA is fused to the C-terminus. * 8 is fused to the N-terminus. * An exemplary structure of a fusion protein having 8 and cytidine deaminase and Cas12 is as follows: N-[Cas12(TadA * 8)]-[cytidine deaminase]-C; N-[cytidine deaminase]-[Cas12(TadA * 8)]-C; N-[Cas12 (cytidine deaminase)]-[TadA * 8]-C; or N-[TadA * 8]-[Cas12 (cytidine deaminase)]-C It is provided as follows.
[0325] In some embodiments, the "-" used in the general structures above indicates the presence of an optional linker.
[0326] In other embodiments, the fusion protein comprises one or more catalytic domains. In other embodiments, at least one of the one or more catalytic domains is inserted within a Cas12 polypeptide or fused to the N-terminus or C-terminus of Cas12. In other embodiments, at least one of the one or more catalytic domains is inserted within a loop, alpha-helical region, amorphous portion, or solvent-accessible portion of a Cas12 polypeptide. In other embodiments, the Cas12 polypeptide is Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ. In other embodiments, the Cas12 polypeptide has at least about 85% amino acid sequence identity to Bacillus hisashii Cas12b, Bacillus thermoamylovorans Cas12b, Bacillus sp. V3-13 Cas12b, or Alicyclobacillus acidiphilus Cas12b (SEQ ID NO: 254). In other embodiments, the Cas12 polypeptide has at least about 90% amino acid sequence identity to Bacillus hisashii Cas12b (SEQ ID NO: 255), Bacillus thermoamylovorans Cas12b, Bacillus sp. V3-13 Cas12b, or Alicyclobacillus acidiphilus Cas12b. In other embodiments, the Cas12 polypeptide has at least about 95% amino acid sequence identity to Bacillus hisashii Cas12b, Bacillus thermoamylovorans Cas12b (SEQ ID NO:256), Bacillus sp. V3-13 Cas12b (SEQ ID NO:257), or Alicyclobacillus acidiphilus Cas12b. In other embodiments, the Cas12 polypeptide comprises or consists essentially of a fragment of Bacillus hisashii Cas12b, Bacillus thermoamylovorans Cas12b, Bacillus sp. V3-13 Cas12b, or Alicyclobacillus acidiphilus Cas12b.In embodiments, the Cas12 polypeptide comprises BvCas12b (V4), which in some embodiments is represented as: 5' mRNA Cap---5' UTR---bhCas12b---termination sequence---3' UTR---120 polyA tail (SEQ ID NOs: 258-260).
[0327] In other embodiments, the catalytic domain is inserted between amino acid positions 153-154, 255-256, 306-307, 980-981, 1019-1020, 534-535, 604-605, or 344-345 of BhCas12b, or between the corresponding amino acid residues of Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ. In other embodiments, the catalytic domain is inserted between amino acids P153 and S154 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids K255 and E256 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids D980 and G981 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids K1019 and L1020 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids F534 and P535 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids K604 and G605 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acids H344 and F345 of BhCas12b. In other embodiments, the catalytic domain is inserted between amino acid positions 147 and 148, 248 and 249, 299 and 300, 991 and 992, or 1031 and 1032 of BvCas12b, or between the corresponding amino acid residues of Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ. In other embodiments, the catalytic domain is inserted between amino acids P147 and D148 of BvCas12b. In other embodiments, the catalytic domain is inserted between amino acids G248 and G249 of BvCas12b. In other embodiments, the catalytic domain is inserted between amino acids P299 and E300 of BvCas12b. In other embodiments, the catalytic domain is inserted between amino acids G991 and E992 of BvCas12b. In other embodiments, the catalytic domain is inserted between amino acids K1031 and M1032 of BvCas12b.In other embodiments, the catalytic domain is inserted between amino acid positions 157 and 158, 258 and 259, 310 and 311, 1008 and 1009, or 1044 and 1045 of AaCas12b, or between the corresponding amino acid residues of Cas12a, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ. In other embodiments, the catalytic domain is inserted between amino acids P157 and G158 of AaCas12b. In other embodiments, the catalytic domain is inserted between amino acids V258 and G259 of AaCas12b. In other embodiments, the catalytic domain is inserted between amino acids D310 and P311 of AaCas12b. In other embodiments, the catalytic domain is inserted between amino acids G1008 and E1009 of AaCas12b. In other embodiments, the catalytic domain is inserted between amino acids G1044 and K1045 of AaCas12b.
[0328] In other embodiments, the fusion protein or complex comprises a nuclear localization signal (e.g., a bipartite nuclear localization signal). In other embodiments, the amino acid sequence of the nuclear localization signal is MAPKKKRKVGIHGVPAA (SEQ ID NO: 261). In other embodiments of the above aspects, the nuclear localization signal is encoded by the following sequence: ATGGCCCCAAAGAAGAAGCGGAAGGTCGGTATCCACGGAGTCCCAGCAGCC (SEQ ID NO: 262). In other embodiments, the Cas12b polypeptide comprises a mutation that silences the catalytic activity of the RuvC domain. In other embodiments, the Cas12b polypeptide comprises a D574A, D829A, and / or D952A mutation. In other embodiments, the fusion protein or complex further comprises a tag (e.g., an influenza hemagglutinin tag).
[0329] In some embodiments, the fusion protein or complex comprises a napDNAbp domain (e.g., a Cas12-derived domain) with an internally fused nucleobase editing domain (e.g., all or a portion (e.g., a functional portion) of a deaminase domain, e.g., an adenosine deaminase domain). In some embodiments, the napDNAbp is Cas12b. In some embodiments, the base editor is an internally fused TadA inserted into a locus provided in Table 5 below. * Contains the BhCas12b domain, which has 8 domains.
[0330] [Table 5]
[0331] Non-limiting examples include adenosine deaminases (e.g., TadA * 8.13) into BhCas12b to generate a fusion protein (e.g., TadA) that effectively edits the nucleic acid sequence. * 8.13-BhCas12b) can be produced.
[0332] In some embodiments, the base editing system described herein is an ABE with TadA inserted into Cas9. The polypeptide sequences of relevant ABEs with TadA inserted into Cas9 are provided in the accompanying sequence listing as SEQ ID NOs: 263-308.
[0333] In some embodiments, an adenosine base editor was generated to insert TadA, or a variant thereof, into the Cas9 polypeptide at the identified position.
[0334] Exemplary, but non-limiting, fusion proteins are described in International PCT Application No. PCT / US2020 / 016285 and U.S. Provisional Application Nos. 62 / 852,228 and 62 / 852,224, the contents of which are incorporated by reference herein in their entireties.
[0335] Editing A to G In some embodiments, the base editors described herein comprise an adenosine deaminase domain. The adenosine deaminase domain of such base editors can facilitate editing of adenine (A) nucleobases to guanine (G) nucleobases by deaminating A to form inosine (I), which exhibits the base pairing properties of G. Adenosine deaminase can deaminate (i.e., remove the amine group from) the adenine of deoxyadenosine residues of deoxyribonucleic acid (DNA). In some embodiments, the A to G base editor further comprises an inhibitor of inosine base excision repair, such as a uracil glycosylase inhibitor (UGI) domain or a catalytically inactive inosine-specific nuclease. Without wishing to be bound by any particular theory, the UGI domain or catalytically inactive inosine-specific nuclease can inhibit or prevent base excision repair of deaminated adenosine residues (e.g., inosine), thereby improving the activity or efficiency of the base editor.
[0336] Base editors comprising adenosine deaminase can act on any polynucleotide, including DNA, RNA, and DNA-RNA hybrids. In certain embodiments, base editors comprising adenosine deaminase can deaminate target A in polynucleotides comprising RNA. For example, a base editor may comprise an adenosine deaminase domain capable of deaminating target A in RNA polynucleotides and / or DNA-RNA hybrid polynucleotides. In embodiments, the adenosine deaminase incorporated into the base editor comprises all or a portion (e.g., a functional portion) of an adenosine deaminase that acts on RNA (ADAR, e.g., ADAR1 or ADAR2) or tRNA (ADAT). Base editors comprising an adenosine deaminase domain can also deaminate A nucleobases in DNA polynucleotides. In embodiments, the adenosine deaminase domain of the base editor comprises all or a portion (e.g., a functional portion) of ADAT containing one or more mutations that enable ADAT to deaminate target A in DNA. For example, a base editor can include all or a portion (e.g., a functional portion) of ADAT from Escherichia coli (EcTadA) containing one or more of the following mutations: D108N, A106V, D147Y, E155V, L84F, H123Y, I156F, or corresponding mutations in another adenosine deaminase. Polypeptide sequences of exemplary ADAT homologs are provided in the Sequence Listing as SEQ ID NOs: 1 and 309-315.
[0337] The adenosine deaminase may be derived from any suitable organism (e.g., E. coli). In some embodiments, the adenosine deaminase is derived from a prokaryote. In some embodiments, the adenosine deaminase is derived from a bacterium. In some embodiments, the adenosine deaminase is derived from E. coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis. In some embodiments, the adenosine deaminase is derived from E. coli. In some embodiments, the adenine deaminase is a naturally occurring adenosine deaminase that contains one or more mutations (e.g., mutations in ecTadA) corresponding to any of the mutations provided herein. Corresponding residues in any homologous protein can be identified, for example, by sequence alignment and determining homologous residues. Mutations of any naturally occurring adenosine deaminase (e.g., having homology with ecTadA) that correspond to any of the mutations described herein (e.g., any of the mutations identified in ecTadA) can be generated accordingly.
[0338] In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the amino acid sequences described in any of the adenosine deaminases provided herein. It should be understood that the adenosine deaminases provided herein may contain one or more mutations (e.g., any of the mutations provided herein). The present disclosure provides any deaminase domain that has a specific percentage of identity as well as any of the mutations described herein, or a combination thereof. In some embodiments, the adenosine deaminase comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more mutations compared to a reference sequence or any of the adenosine deaminases provided herein. In some embodiments, the adenosine deaminase comprises an amino acid sequence having at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, or at least 170 identical contiguous amino acid residues compared to any one of the amino acid sequences known in the art or described herein.
[0339] The mutations provided herein (e.g., TadA *It should be understood that any of the TadA reference sequences (based on a TadA reference sequence such as SEQ ID NO: 1) can be introduced into other adenosine deaminases, such as E. coli TadA (ecTadA), S. aureus TadA (saTadA), or other adenosine deaminases (e.g., bacterial adenosine deaminases). In some embodiments, the TadA reference sequence is * 7.10 (SEQ ID NO: 1). It will be apparent to one skilled in the art that additional deaminases can be similarly aligned to identify homologous amino acid residues that can be mutated as provided herein. Thus, any of the mutations identified in the TadA reference sequence can be made in other adenosine deaminases (e.g., ecTada) that have homologous amino acid residues. It should also be understood that any of the mutations provided herein can be made individually or in any combination in the TadA reference sequence or another adenosine deaminase.
[0340] In some embodiments, the adenosine deaminase comprises a D108X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a D108G, D108N, D108V, D108A, or D108Y mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. However, it should be understood that additional deaminases can be similarly aligned to identify homologous amino acid residues that can be mutated as provided herein.
[0341] In some embodiments, the adenosine deaminase comprises an A106X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A106V mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0342] In some embodiments, the adenosine deaminase comprises an E155X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where the presence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an E155D, E155G, or E155V mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0343] In some embodiments, the adenosine deaminase comprises a D147X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where the presence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a D147Y mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0344] In some embodiments, the adenosine deaminase comprises an A106X, E155X, or D147X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA), where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an E155D, E155G, or E155V mutation. In some embodiments, the adenosine deaminase comprises D147Y.
[0345] It should also be understood that any of the mutations provided herein can be made individually or in any combination in ecTadA or another adenosine deaminase. For example, an adenosine deaminase can include D108N, A106V, E155V, and / or D147Y mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase (e.g., ecTadA). In some embodiments, an adenosine deaminase can include a TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1), or the corresponding mutations in another adenosine deaminase: D108N and A106V; D108N and E155V; D108N and D147Y; A106V and E155V; A106V and D147Y; E155V and D147Y; D108N, A106V, and E155V; D108N, A106V, and D147Y; D108N, E155V, and D147Y; A106V, E155V, and D147Y; and D108N, A106V, E155V, and D147Y. However, it should be understood that any combination of the corresponding mutations provided herein can be made in an adenosine deaminase (eg, ecTadA).
[0346] In some embodiments, the adenosine deaminase is a TadA reference sequence (e.g., TadA *7.10 (SEQ ID NO: 1)), or a combination of corresponding mutations in another adenosine deaminase: V82G + Y147T + Q154S; I76Y + V82G + Y147T + Q154S; L36H + V82G + Y147T + Q154S + N157K; V82G + Y147D + F149Y + Q154S + D167N; L36H+V82G+Y147D+F149Y+Q154S+N157K+D167N;L36H+I76Y+V82G+Y147T+Q154S+N157K;I76Y+V82G+Y147D+F149Y+Q154S+D167N;or L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N In some embodiments, the adenosine deaminase comprises one or more of the H8X, T17X, L18X, W23X, L34X, W45X, R51X, A56X, E59X, E85X, M94X, I95X, V102X, F104X, A106X, R107X, D108X, K110X, M118X, N127X, A138X, F149X, M151X, R153X, Q154X, I156X, and / or K157X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein the presence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase is selected from the group consisting of H8Y, T17S, L18E, W23L, L34S, W45L, R51H, A56E, or A56S, E59G, E85K, or E85G, M94L, I95L, V102A, F104L, A106V, R107C, or R107H, or R107P, D108G, or D108N, or D108V, or D108A, or one or more of D108Y, K110I, M118K, N127S, A138V, F149Y, M151V, R153C, Q154L, I156D, and / or K157R mutations, or one or more corresponding mutations in another adenosine deaminase.
[0347] In some embodiments, the adenosine deaminase comprises one or more of an H8X, D108X, and / or N127X mutation in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, where X indicates the presence of any amino acid. In some embodiments, the adenosine deaminase comprises one or more of an H8Y, D108N, and / or N127S mutation in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.
[0348] In some embodiments, the adenosine deaminase comprises one or more of the H8X, R26X, M61X, L68X, M70X, A106X, D108X, A109X, N127X, D147X, R152X, Q154X, E155X, K161X, Q163X, and / or T166X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the H8Y, R26W, M61I, L68Q, M70V, A106T, D108N, A109T, N127S, D147Y, R152C, Q154H or Q154R, E155G or E155V or E155D, K161Q, Q163H, and / or T166P mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.
[0349] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8X, D108X, N127X, D147X, R152X, and Q154X in the TadA reference sequence, or the corresponding mutation or mutations in another adenosine deaminase (e.g., ecTadA), where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8X, M61X, M70X, D108X, N127X, Q154X, E155X, and Q163X in the TadA reference sequence, or the corresponding mutation or mutations in another adenosine deaminase (e.g., ecTadA), where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8X, D108X, N127X, E155X, and T166X in the TadA reference sequence, or the corresponding mutation or mutations in another adenosine deaminase (e.g., ecTadA), where X indicates the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase.
[0350] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8X, A106X, and D108X, or a corresponding mutation or mutations in another adenosine deaminase, where X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8X, R26X, L68X, D108X, N127X, D147X, and E155X, or a corresponding mutation or mutations in another adenosine deaminase, where X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase.
[0351] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, or seven mutations selected from the group consisting of H8X, R126X, L68X, D108X, N127X, D147X, and E155X in the TadA reference sequence, or a corresponding mutation or mutations in another adenosine deaminase, where X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8X, D108X, A109X, N127X, and E155X in the TadA reference sequence, or a corresponding mutation or mutations in another adenosine deaminase, where X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase.
[0352] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8Y, D108N, N127S, D147Y, R152C, and Q154H in the TadA reference sequence, or a corresponding mutation or mutations in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8Y, M61I, M70V, D108N, N127S, Q154R, E155G, and Q163H in the TadA reference sequence, or a corresponding mutation or mutations in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8Y, D108N, N127S, E155V, and T166P in the TadA reference sequence, or a corresponding mutation or mutations in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of H8Y, A106T, D108N, N127S, E155D, and K161Q in the TadA reference sequence, or a corresponding mutation or mutations in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, seven, or eight mutations selected from the group consisting of H8Y, R26W, L68Q, D108N, N127S, D147Y, and E155V in the TadA reference sequence, or a corresponding mutation or mutations in another adenosine deaminase (e.g., ecTadA). In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8Y, D108N, A109T, N127S, and E155G in the TadA reference sequence, or a corresponding mutation or mutations in another adenosine deaminase (e.g., ecTadA).
[0353] In some embodiments, the adenosine deaminase comprises one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises a D108N, D108G, or D108V mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A106V and D108N mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R107C and D108N mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises an H8Y, D108N, N127S, D147Y, and Q154H mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises H8Y, D108N, N127S, D147Y, and E155V mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises D108N, D147Y, and E155V mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises H8Y, D108N, and N127S mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises A106V, D108N, D147Y, and E155V mutations in the TadA reference sequence, or corresponding mutations in another adenosine deaminase (e.g., ecTadA).
[0354] In some embodiments, the adenosine deaminase comprises one or more of the S2X, H8X, I49X, L84X, H123X, N127X, I156X, and / or K160X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, where the presence of X denotes any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the S2A, H8Y, I49F, L84F, H123Y, N127S, I156F, and / or K160S mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase (e.g., ecTadA).
[0355] In some embodiments, the adenosine deaminase comprises an L84X mutant adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an L84F mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0356] In some embodiments, the adenosine deaminase comprises an H123X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an H123Y mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0357] In some embodiments, the adenosine deaminase comprises an I156X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an I156F mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0358] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, or seven mutations selected from the group consisting of L84X, A106X, D108X, H123X, D147X, E155X, and I156X in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of S2X, I49X, A106X, D108X, D147X, and E155X in the TadA reference sequence, or a corresponding mutation or mutations in another adenosine deaminase, where X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8X, A106X, D108X, N127X, and K160X in the TadA reference sequence, or the corresponding mutation or mutations in another adenosine deaminase, where X represents the presence of any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase.
[0359] In some embodiments, the adenosine deaminase comprises one, two, three, four, five, six, or seven mutations selected from the group consisting of L84F, A106V, D108N, H123Y, D147Y, E155V, and I156F in the TadA reference sequence, or the corresponding mutation or mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one, two, three, four, five, or six mutations selected from the group consisting of S2A, I49F, A106V, D108N, D147Y, and E155V in the TadA reference sequence.
[0360] In some embodiments, the adenosine deaminase comprises one, two, three, four, or five mutations selected from the group consisting of H8Y, A106T, D108N, N127S, and K160S in the TadA reference sequence, or the corresponding mutation or mutations in another adenosine deaminase.
[0361] In some embodiments, the adenosine deaminase comprises one or more of the E25X, R26X, R107X, A142X, and / or A143X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein the occurrence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the E25M, E25D, E25A, E25R, E25V, E25S, E25Y, R26G, R26N, R26Q, R26C, R26L, R26K, R107P, R107K, R107A, R107N, R107W, R107H, R107S, A142N, A142D, A142G, A143D, A143G, A143E, A143L, A143W, A143M, A143S, A143Q, and / or A143R mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the mutations described herein corresponding to the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase.
[0362] In some embodiments, the adenosine deaminase comprises an E25X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an E25M, E25D, E25A, E25R, E25V, E25S, or E25Y mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0363] In some embodiments, the adenosine deaminase comprises an R26X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R26G, R26N, R26Q, R26C, R26L, or R26K mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0364] In some embodiments, the adenosine deaminase comprises an R107X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R107P, R107K, R107A, R107N, R107W, R107H, or R107S mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0365] In some embodiments, the adenosine deaminase comprises an A142X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A142N, A142D, A142G mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0366] In some embodiments, the adenosine deaminase comprises an A143X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A143D, A143G, A143E, A143L, A143W, A143M, A143S, A143Q, and / or A143R mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase (e.g., ecTadA).
[0367] In some embodiments, the adenosine deaminase comprises one or more of the H36X, N37X, P48X, I49X, R51X, M70X, N72X, D77X, E134X, S146X, Q154X, K157X, and / or K161X mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase, wherein the occurrence of X indicates any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises one or more of the H36L, N37T, N37S, P48T, P48L, I49V, R51H, R51L, M70L, N72S, D77G, E134G, S146R, S146C, Q154H, K157N, and / or K161T mutations in the TadA reference sequence, or one or more corresponding mutations in another adenosine deaminase (e.g., ecTadA).
[0368] In some embodiments, the adenosine deaminase comprises an H36X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an H36L mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0369] In some embodiments, the adenosine deaminase comprises an N37X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an N37T or N37S mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0370] In some embodiments, the adenosine deaminase comprises a P48X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a P48T or P48L mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0371] In some embodiments, the adenosine deaminase comprises an R51X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R51H or R51L mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0372] In some embodiments, the adenosine deaminase comprises a S146X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a S146R or S146C mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0373] In some embodiments, the adenosine deaminase comprises a K157X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a K157N mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0374] In some embodiments, the adenosine deaminase comprises a P48X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a P48S, P48T, or P48A mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0375] In some embodiments, the adenosine deaminase comprises an A142X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an A142N mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0376] In some embodiments, the adenosine deaminase comprises a W23X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises a W23R or W23L mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0377] In some embodiments, the adenosine deaminase comprises an R152X mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase, where X represents any amino acid other than the corresponding amino acid in the wild-type adenosine deaminase. In some embodiments, the adenosine deaminase comprises an R152P or R52H mutation in the TadA reference sequence, or a corresponding mutation in another adenosine deaminase.
[0378] In one embodiment, the adenosine deaminase may include the mutations H36L, R51L, L84F, A106V, D108N, H123Y, S146C, D147Y, E155V, I156F, and K157N. In some embodiments, the adenosine deaminase includes the following combinations of mutations relative to the TadA reference sequence, where each mutation in the combination is separated by an "_" and each combination of mutations is between parentheses: (A106V_D108N), (R107C_D108N), (H8Y_D108N_N127S_D147Y_Q154H), (H8Y_D108N_N127S_D147Y_E155V), (D108N_D147Y_E155V), (H8Y_D108N_N127S), (H8Y_D108N_N127S_D147Y_Q154H), (A106V_D108N_D147Y_E155V), (D108Q_D147Y_E155V), (D108M_D147Y_E155V), (D108L_D147Y_E155V), (D108K_D147Y_E155V), (D108I_D147Y_E155V), (D108F_D147Y_E155V), (A106V_D108N_D147Y), (A106V_D108M_D147Y_E155V), (E59A_A106V_D108N_D147Y_E155V), (E59A cat dead_A106V_D108N_D147Y_E155V), (L84F_A106V_D108N_H123Y_D147Y_E155V_I156Y), (L84F_A106V_D108N_H123Y_D147Y_E155V_I156F), (D103A_D104N), (G22P_D103A_D104N), (D103A_D104N_S138A), (R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F), (E25G_R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F), (E25D_R26G_L84F_A106V_R107K_D108N_H123Y_A142N_A143G_D147Y_E155 V_I156F), (R26Q_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F), (E25M_R26G_L84F_A106V_R107P_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F), (R26C_L84F_A106V_R107H _D108N_H123Y_A142N_D147Y_E155V_I156F), (L84F_A106V_D108N_H123Y_A142N_A143L_D147Y_E155V_I156F), (R26G_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F), (E25A_R26G_L84F_A106V_R107N_D108N_H123Y_A142N_A143E_D147Y_E155V_I156F)、(R26G_L84F_A106V_R107H_D108N_H123Y_A142N_A143D_D147Y_E155V_I156F)、 (A106V_D108N_A142N_D147Y_E155V)、 (R26G_A106V_D108N_A142N_D147Y_E155V)、 (E25D_R26G_A106V_R107K_D108N_A142N_A143G_D147Y_E155V)、(R26G_A106V_D108N_R107H_A142N_A143D_D147Y_E155V)、 (E25D_R26G_A106V_D108N_A142N_D147Y_E155V)、 (A106V_R107K_D108N_A142N_D147Y_E155V)、 (A106V_D108N_A142N_A143G_D147Y_E155V)、 (A106V_D108N_A142N_A143L_D147Y_E155V)、 (H36L_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、 (N37T_P48T_M70L_L84F_A106V_D108N_H123Y_D147Y_I49V_E155V_I156F)、 (N37S_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K161T)、 (H36L_L84F_A106V_D108N_H123Y_D147Y_Q154H_E155V_I156F)、 (N72S_L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F)、 (H36L_P48L_L84F_A106V_D108N_H123Y_E134G_D147Y_E155V_I156F)、 (H36L_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K157N)(H36L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F)、 (L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F_K161T)、 (N37S_R51H_D77G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (R51L_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_K157N)、 (D24G_Q71R_L84F_H96L_A106V_D108N_H123Y_D147Y_E155V_I156F_K160E)、 (H36L_G67V_L84F_A106V_D108N_H123Y_S146T_D147Y_E155V_I156F)、 (Q71L_L84F_A106V_D108N_H123Y_L137M_A143E_D147Y_E155V_I156F)、 (E25G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_Q159L)、 (L84F_A91T_F104I_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (N72D_L84F_A106V_D108N_H123Y_G125A_D147Y_E155V_I156F)、 (P48S_L84F_S97C_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (W23G_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (D24G_P48L_Q71R_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F_Q159L)、 (L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F)、 (H36L_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_E155V_I156F_K157N)、(N37S_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F_K161T)、 (L84F_A106V_D108N_D147Y_E155V_I156F)、 (R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N_K161T)、 (L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K161T)、 (L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N_K160E_K161T)、 (L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N_K160E)、(R74Q_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (R74A_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (R74Q_L84F_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (L84F_R98Q_A106V_D108N_H123Y_D147Y_E155V_I156F)、 (L84F_A106V_D108N_H123Y_R129Q_D147Y_E155V_I156F)、 (P48S_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F)、 (P48S_A142N)、 (P48T_I49V_L84F_A106V_D108N_H123Y_A142N_D147Y_E155V_I156F_L157N)、 (P48T_I49V_A142N)、 (H36L_P48S_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、(H36L_P48S_R51L_L84F_A106V_D108N_H123Y_S146C_A142N_D147Y_E155V_I156F(H36L_P48T_I49V_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、 (H36L_P48T_I49V_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_E155V_I156F_K157N)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_E155V_I156F_K157N)、(H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_A142N_D147Y_E155V_I156F_K157N)、 (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、(W23R_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_E155V_I156F_K157N)、(W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F_K161T)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152H_E155V_I156F_K157N)、 (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152P_E155V_I156F_K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152P_E155V_I156F_K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142A_S146C_D147Y_E155V_I156F_K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142A_S146C_D147Y_R152P_E155V_I156F_K157N), (W23L_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146R_D147Y_E155V_I156F_K161T), (W23R_H36L_P48A_R51L_L84F_A106V_D108N_H123Y_S146C_D147Y_R152P_E155V_I156F_K157N), (H36L_P48A_R51L_L84F_A106V_D108N_H123Y_A142N_S146C_D147Y_R152P_E155V_I156F_K157N).
[0379] In some embodiments, the TadA deaminase is a TadA variant. In some embodiments, the TadA variant is TadA * 7.10. In certain embodiments, the fusion protein or complex contains a single TadA * In other embodiments, the fusion protein comprises a TadA domain (e.g., provided as a monomer). * 7.10 and TadA(wt), which can form heterodimers. In one embodiment, the fusion protein of the invention comprises TadA linked to Cas9 nickase. * 7.10, containing wild-type TadA linked to TadA.
[0380] In some embodiments, TadA * 7.10 comprises at least one alteration. In some embodiments, the adenosine deaminase comprises the following sequence alteration: TadA * 7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 1).
[0381] In some embodiments, TadA * 7.10 contains changes at amino acids 82 and / or 166. In certain embodiments, TadA * 7.10 contains one or more of the following changes: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. * The 7.10 variants are Y147T+Q154R;Y147T+Q154S;Y147R+Q154S;V82S+Q154S;V82S+Y147R;V82S+Q154R;V82S+Y123H;I76Y+V82S;V82S+Y123H+Y147T;V82S+Y123H+Y147R;V82S+Y123H+Q15 Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R.
[0382] In some embodiments, TadA *The 7.10 variant comprises one or more of the changes selected from the group of L36H, I76Y, V82G, Y147T, Y147D, F149Y, Q154S, N157K, and / or D167N. * The 7.10 variant includes V82G, Y147T / D, Q154S, and one or more of L36H, I76Y, F149Y, N157K, and D167N. * The 7.10 variants are V82G+Y147T+Q154S;I76Y+V82G+Y147T+Q154S;L36H+V82G+Y147T+Q154S+N157K;V82G+Y147D+F149Y+Q154S+D167N;L36H+V82G+Y147D+F149Y+Q154S+N15 L36H+I76Y+V82G+Y147T+Q154S+N157K+D167N; L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N; L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N; L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N.
[0383] In some embodiments, an adenosine deaminase variant (e.g., TadA * 8) comprises a deletion. In some embodiments, the adenosine deaminase variant comprises a C-terminal deletion. In certain embodiments, the adenosine deaminase variant comprises a TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1)), or a C-terminal deletion starting at residues 149, 150, 151, 152, 153, 154, 155, 156, and 157 relative to the corresponding mutation in another TadA.
[0384] In other embodiments, adenosine deaminase variants (e.g., TadA * 8) is a TadA reference sequence (e.g., TadA *7.10 (SEQ ID NO: 1), or a corresponding mutation in another TadA, is a monomer that includes one or more of the following changes: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. In other embodiments, an adenosine deaminase variant (TadA * 8) is a TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1)), or for the corresponding mutations in another TadA, Y147T+Q154R;Y147T+Q154S;Y147R+Q154S;V82S+Q154S;V82S+Y147R;V82S+Q154R;V82S+Y123H;I76Y+V82S;V82S+Y123H+Y147T;V82S+Y123H+Y147R;V82S +Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R.
[0385] In other embodiments, the adenosine deaminase variant is a variant of the TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1)), or a corresponding mutation in another TadA, two adenosine deaminase domains each having one or more of the following changes: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R (e.g., TadA * 8). In other embodiments, the adenosine deaminase variant is a homodimer comprising a TadA reference sequence (e.g., TadA *7.10 (SEQ ID NO: 1)), or the corresponding mutations in another TadA, each of which is Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q two adenosine deaminase domains (e.g., TadA) having a combination of changes selected from the group consisting of: 154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R. * 8) is a homodimer containing
[0386] In other embodiments, the TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1)), or adenosine deaminase variants comprising one or more of the following changes: R26C, V88A, A109S, T111R, D119N, H122N, Y147D, F149Y, T166I and / or D167N relative to the corresponding mutation in another TadA (e.g., TadA * 8) A base editor of the present disclosure comprising a monomer. In other embodiments, adenosine deaminase variant (TadA * 8) The monomer is a TadA reference sequence (e.g., TadA *7.10 (SEQ ID NO: 1)), or for the corresponding mutations in another TadA, R26C+A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N; V88A+A109S+T111R+D119N+H122N+F149Y+T166I+D167N; R26C+A109S+T111R+D119N+H122N+F149Y+T166I+D167N; V88A+T111R+D119N+F149Y; and A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N.
[0387] In some embodiments, the adenosine deaminase variant (e.g., MSP828) is a variant of the TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1), or a corresponding mutation in another TadA, is a monomer that includes one or more of the following changes: L36H, I76Y, V82G, Y147T, Y147D, F149Y, Q154S, N157K, and / or D167N. In some embodiments, the adenosine deaminase variant (e.g., MSP828) is a monomer that includes one or more of the following changes relative to a TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1), or a corresponding mutation in another TadA. In other embodiments, the adenosine deaminase variant (TadA variant) is a monomer that includes V82G, Y147T / D, Q154S, and one or more of L36H, I76Y, F149Y, N157K, and D167N relative to a TadA reference sequence (e.g., TadA *7.10 (SEQ ID NO: 1)), or for the corresponding mutations in another TadA, V82G+Y147T+Q154S; I76Y+V82G+Y147T+Q154S; L36H+V82G+Y147T+Q154S+N157K; V82G+Y147D+F149Y+Q154S+D167N; L36H+V82G+Y147D+F149 Y+Q154S+N157K+D167N; L36H+I76Y+V82G+Y147T+Q154S+N157K; I76Y+V82G+Y147D+F149Y+Q154S+D167N; L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N.
[0388] In other embodiments, the adenosine deaminase variant comprises a wild-type adenosine deaminase domain and a TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1)), or a corresponding mutation in another TadA, an adenosine deaminase variant domain comprising one or more of the following changes: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. * 8). In other embodiments, the adenosine deaminase variant is a heterodimer of a wild-type adenosine deaminase domain and a TadA reference sequence (e.g., TadA *7.10 (SEQ ID NO: 1)), or for the corresponding mutations in another TadA, Y147T+Q154R;Y147T+Q154S;Y147R+Q154S;V82S+Q154S;V82S+Y147R;V82S+Q154R;V82S+Y123H;I76Y+V82S;V82S+Y123H+Y147T;V82S+Y123H+Y147R;V82S+Y123H+Q154 and I76Y+V82S+Y123H+Y147R+Q154R. An adenosine deaminase variant domain (e.g., TadA) comprising a combination of changes selected from the group consisting of: R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R. * 8) is a heterodimer.
[0389] In other embodiments, the TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1)), or a corresponding mutation in another TadA, two adenosine deaminase domains each having one or more of the following changes: R26C, V88A, A109S, T111R, D119N, H122N, Y147D, F149Y, T166I and / or D167N (e.g., TadA * 8) adenosine deaminase variants (e.g., TadA * 8) Base editors of the present disclosure that comprise homodimers. In other embodiments, the adenosine deaminase variant is a base editor that is a homodimer of a TadA reference sequence (e.g., TadA *7.10 (SEQ ID NO: 1)), or the corresponding mutations in another TadA, each of which is R26C+A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N; V88A+A109S+T111R+D119N+H122N+F149Y+T166I+D167N; R26C+A109S+ two adenosine deaminase domains (e.g., TadA) with a combination of changes selected from the group of: T111R + D119N + H122N + F149Y + T166I + D167N; V88A + T111R + D119N + F149Y; and A109S + T111R + D119N + H122N + Y147D + F149Y + T166I + D167N * 8) is a homodimer containing
[0390] In some embodiments, the adenosine deaminase variant is a variant of the TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1)), or a corresponding mutation in another TadA, two adenosine deaminase domains each having one or more of the following changes: L36H, I76Y, V82G, Y147T, Y147D, F149Y, Q154S, N157K, and / or D167N (e.g., TadA * 7.10). In some embodiments, the adenosine deaminase variant is a homodimer comprising a TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1), or a corresponding mutation in another TadA, is a homodimer comprising two adenosine deaminase variant domains (e.g., MSP828) each having one or more of the following changes: V82G, Y147T / D, Q154S, and L36H, I76Y, F149Y, N157K, and D167N. In other embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase variant domains (e.g., MSP828) each having one or more of the following changes: V82G, Y147T / D, Q154S, and L36H, I76Y, F149Y, N157K, and D167N, relative to a TadA reference sequence (e.g., TadA *7.10 (SEQ ID NO: 1)), or the corresponding mutations in another TadA, each of which is V82G+Y147T+Q154S; I76Y+V82G+Y147T+Q154S; L36H+V82G+Y147T+Q154S+N157K; V82G+Y147D+F149Y+Q154S+D167N; L36H+V82G+Y147D+F149Y+Q154S+N two adenosine deaminase domains (e.g., TadA) having a combination of changes selected from the group consisting of: L36H+I76Y+V82G+Y147T+Q154S+N157K; L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K; L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N; L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N * 7.10) is a homodimer containing
[0391] In other embodiments, the adenosine deaminase variant is TadA * 7.10 domain and the TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1)), or a corresponding mutation in another TadA, an adenosine deaminase variant domain comprising one or more of the following changes: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. * 8). In another embodiment, the adenosine deaminase variant is a heterodimer with TadA. * 7.10 domain and the TadA reference sequence (e.g., TadA *7.10 (SEQ ID NO: 1)), or for the corresponding mutations in another TadA, Y147T+Q154R;Y147T+Q154S;Y147R+Q154S;V82S+Q154S;V82S+Y147R;V82S+Q154R;V82S+Y123H;I76Y+V82S;V82S+Y123H+Y147T;V82S+Y123H+Y147R;V82S+Y123H+Q154 and I76Y+V82S+Y123H+Y147R+Q154R. An adenosine deaminase variant domain (e.g., TadA) comprising a combination of changes selected from the group consisting of: R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R. * 8) is a heterodimer.
[0392] In other embodiments, the base editor comprises a wild-type adenosine deaminase domain and a TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1)), or a corresponding mutation in another TadA, an adenosine deaminase variant domain comprising one or more of the following changes R26C, V88A, A109S, T111R, D119N, H122N, Y147D, F149Y, T166I and / or D167N (e.g., TadA * 8). In other embodiments, the base editor comprises a heterodimer of a wild-type adenosine deaminase domain and a TadA reference sequence (e.g., TadA *7.10 (SEQ ID NO: 1)), or for the corresponding mutations in another TadA, R26C+A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N; V88A+A109S+T111R+D119N+H122N+F149Y+T166I+D167N; R26C+A109S+T111R+D119N+H122N+F149Y+T166I+D167N; an adenosine deaminase variant domain (e.g., TadA) comprising a combination of changes selected from the group consisting of: V88A+T111R+D119N+H122N+F149Y+T166I+D167N; V88A+T111R+D119N+F149Y; and A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N * 8) and heterodimers.
[0393] In other embodiments, the adenosine deaminase variant comprises a wild-type adenosine deaminase domain and a TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1)), or a corresponding mutation in another TadA, an adenosine deaminase variant domain comprising one or more of the following changes L36H, I76Y, V82G, Y147T, Y147D, F149Y, Q154S, N157K, and / or D167N (e.g., TadA * 7.10). In some embodiments, the adenosine deaminase variant is a heterodimer of a wild-type adenosine deaminase domain and a TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1), or a corresponding mutation in another TadA, and an adenosine deaminase variant domain having one or more of the following changes: V82G, Y147T / D, Q154S, and L36H, I76Y, F149Y, N157K, and D167N (e.g., MSP828). In other embodiments, the adenosine deaminase variant is a heterodimer comprising a wild-type adenosine deaminase domain and a TadA reference sequence (e.g., TadA *7.10 (SEQ ID NO: 1)), or for the corresponding mutations in another TadA, V82G+Y147T+Q154S; I76Y+V82G+Y147T+Q154S; L36H+V82G+Y147T+Q154S+N157K; V82G+Y147D+F149Y+Q154S+D167N; L36H+V82G+Y147D+F149Y+Q154S+N15 an adenosine deaminase variant domain (e.g., TadA) comprising a combination of changes selected from the group consisting of: L36H+I76Y+V82G+Y147T+Q154S+N157K; L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K; L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N; L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N * 7.10) is a heterodimer.
[0394] In other embodiments, the adenosine deaminase variant is TadA * 7.10 domain and the TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1)), or a corresponding mutation in another TadA, an adenosine deaminase variant domain comprising one or more of the following changes: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. * 8). In another embodiment, the adenosine deaminase variant is a heterodimer with TadA. * 7.10 domain and the TadA reference sequence (e.g., TadA *7.10 (SEQ ID NO: 1)), or for the corresponding mutations in another TadA, Y147T+Q154R;Y147T+Q154S;Y147R+Q154S;V82S+Q154S;V82S+Y147R;V82S+Q154R;V82S+Y123H;I76Y+V82S;V82S+Y123H+Y147T;V82S+Y123H+Y147R;V82S+Y123H+Q154 and I76Y+V82S+Y123H+Y147R+Q154R. An adenosine deaminase variant domain (e.g., TadA) comprising a combination of changes selected from the group consisting of: R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R. * 8) is a heterodimer.
[0395] In certain embodiments, the adenosine deaminase heterodimer is TadA * 8 domains and Staphylococcus aureus (S. aureus) TadA, Bacillus subtilis (B. subtilis) TadA, Salmonella typhimurium (S. typhimurium) TadA, Shewanella putrefaciens (S. putrefaciens) TadA, Haemophilus influenzae F3031 (H. influenzae) TadA, Caulobacter crescentus (C. crescentus) TadA, Geobacter sulfurreducens (G. sulfurreducens) TadA, or TadA * 7.10.
[0396] In some embodiments, the adenosine deaminase is TadA. * 8. In one embodiment, the adenosine deaminase comprises or consists essentially of the following sequence or a fragment thereof having adenosine deaminase activity: TadA * 8 is: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 316).
[0397] In some embodiments, TadA * 8 is truncated. In some embodiments, the truncated TadA * 8 is the full length TadA * In some embodiments, the truncated TadA lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acid residues relative to 8. * 8 is the full length TadA * In some embodiments, the adenosine deaminase variant is a full-length TadA variant lacking 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 C-terminal amino acid residues relative to TadA. * It is 8.
[0398] In some embodiments, TadA * 8 is TadA * 8.1, TadA * 8.2, TadA * 8.3, TadA * 8.4, TadA * 8.5, TadA * 8.6, TadA * 8.7, TadA * 8.8, TadA * 8.9, TadA * 8.10, TadA * 8.11, TadA * 8.12, TadA * 8.13, TadA * 8.14, TadA * 8.15, TadA * 8.16, TadA * 8.17, TadA* 8.18, TadA * 8.19, TadA * 8.20, TadA * 8.21, TadA * 8.22, TadA * 8.23, or TadA * It's 8.24.
[0399] In other embodiments, the TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1)), or adenosine deaminase variants comprising one or more of the following changes: R26C, V88A, A109S, T111R, D119N, H122N, Y147D, F149Y, T166I and / or D167N relative to the corresponding mutation in another TadA (e.g., TadA * 8) A base editor of the present disclosure comprising a monomer. In other embodiments, adenosine deaminase variant (TadA * 8) The monomer is a TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1)), or for the corresponding mutations in another TadA, R26C+A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N; V88A+A109S+T111R+D119N+H122N+F149Y+T166I+D167N; R26C+A109S+T111R+D119N+H122N+F149Y+T166I+D167N; V88A+T111R+D119N+F149Y; and A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N.
[0400] In other embodiments, the base editor comprises a wild-type adenosine deaminase domain and a TadA reference sequence (e.g., TadA *7.10 (SEQ ID NO: 1)), or a corresponding mutation in another TadA, an adenosine deaminase variant domain comprising one or more of the following changes R26C, V88A, A109S, T111R, D119N, H122N, Y147D, F149Y, T166I and / or D167N (e.g., TadA * 8). In other embodiments, the base editor comprises a heterodimer of a wild-type adenosine deaminase domain and a TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1)), or for the corresponding mutations in another TadA, R26C+A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N; V88A+A109S+T111R+D119N+H122N+F149Y+T166I+D167N; R26C+A109S+T111R+D119N+H122N+F149Y+T166I+D167N; an adenosine deaminase variant domain (e.g., TadA) comprising a combination of changes selected from the group consisting of: V88A+T111R+D119N+H122N+F149Y+T166I+D167N; V88A+T111R+D119N+F149Y; and A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N * 8) and heterodimers.
[0401] In other embodiments, the base editor is TadA * 7.10 domain and the TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1)), or a corresponding mutation in another TadA, an adenosine deaminase variant domain comprising one or more of the following changes R26C, V88A, A109S, T111R, D119N, H122N, Y147D, F149Y, T166I and / or D167N (e.g., TadA * 8). In other embodiments, the base editor comprises a heterodimer with TadA. * 7.10 domain and the TadA reference sequence (e.g., TadA *7.10 (SEQ ID NO: 1)), or for the corresponding mutations in another TadA, R26C+A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N; V88A+A109S+T111R+D119N+H122N+F149Y+T166I+D167N; R26C+A109S+T111R+D119N+H122N+F149Y+T166I+D167N; an adenosine deaminase variant domain (e.g., TadA) comprising a combination of changes selected from the group consisting of: V88A+T111R+D119N+H122N+F149Y+T166I+D167N; V88A+T111R+D119N+F149Y; and A109S+T111R+D119N+H122N+Y147D+F149Y+T166I+D167N * 8) and heterodimers.
[0402] In other embodiments, the adenosine deaminase variant is TadA * 7.10 domain and the TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1)), or a corresponding mutation in another TadA, an adenosine deaminase variant domain comprising one or more of the following changes L36H, I76Y, V82G, Y147T, Y147D, F149Y, Q154S, N157K, and / or D167N (e.g., TadA * 7.10). In some embodiments, the adenosine deaminase variant is a heterodimer with TadA. * 7.10 domain and the TadA reference sequence (e.g., TadA * 7.10 (SEQ ID NO: 1), or a corresponding mutation in another TadA (e.g., MSP828) with an adenosine deaminase variant domain having one or more of the following changes: V82G, Y147T / D, Q154S, and L36H, I76Y, F149Y, N157K, and D167N. In other embodiments, the adenosine deaminase variant is a heterodimer comprising TadA * 7.10 domain and the TadA reference sequence (e.g., TadA *7.10 (SEQ ID NO: 1)), or for the corresponding mutations in another TadA, V82G+Y147T+Q154S; I76Y+V82G+Y147T+Q154S; L36H+V82G+Y147T+Q154S+N157K; V82G+Y147D+F149Y+Q154S+D167N; L36H+V82G+Y147D+F149Y+Q154S+N15 an adenosine deaminase variant domain (e.g., TadA) comprising a combination of changes selected from the group consisting of: L36H+I76Y+V82G+Y147T+Q154S+N157K; L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K; L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N; L36H+I76Y+V82G+Y147D+F149Y+Q154S+N157K+D167N * 7.10) is a heterodimer.
[0403] In some embodiments, TadA * 8 is a variant shown in Table 6. Table 6 shows specific amino acid position numbers in the TadA amino acid sequence and the amino acids present at those positions in TadA-7.10 adenosine deaminase. Table 6 also shows the amino acid changes of TadA variants relative to TadA-7.10 after phage-assisted discontinuous evolution (PANCE) and phage-assisted continuous evolution (PACE), as described in M. Richter et al., 2020, Nature Biotechnology, doi.org / 10.1038 / s41587-020-0453-z, the entire contents of which are incorporated herein by reference. In some embodiments, TadA * 8 is TadA * 8a, TadA * 8b, TadA * 8c, TadA * 8d, or TadA * 8e. In some embodiments, TadA * 8 is TadA * It's 8e.
[0404] [Table 6]
[0405] In some embodiments, the TadA variant is a variant shown in Table 6.1. Table 6.1 lists the specific amino acid position numbers in the TadA amino acid sequence and the TadA * The amino acids present at those positions in 7.10 adenosine deaminase are shown. In some embodiments, the TadA variant is MSP605, MSP680, MSP823, MSP824, MSP825, MSP827, MSP828, or MSP829. In some embodiments, the TadA variant is MSP828. In some embodiments, the TadA variant is MSP829.
[0406] [Table 6-1]
[0407] In one embodiment, the fusion protein or complex of the invention comprises an adenosine deaminase variant described herein (e.g., TadA) linked to a Cas9 nickase. * 8). In certain embodiments, the fusion protein or complex comprises a single TadA. * In other embodiments, the fusion protein or complex comprises TadA 8 domains (e.g., provided as a monomer). * 8 and TadA(wt), which can form heterodimers.
[0408] In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%,...
Claims
1. 1. A method of altering a nucleobase of an IgG receptor and transporter Fc fragment (FcRn) polynucleotide, the method comprising contacting the FcRn polynucleotide with a base editor system comprising one or more guide polynucleotides and a base editor comprising a nucleic acid programmable DNA binding protein (napDNAbp) domain and a deaminase domain, or one or more polynucleotides encoding said base editor system; (a) the one or more guide polynucleotides comprise a nucleic acid sequence comprising at least 10 to 23 contiguous nucleotides of a spacer nucleic acid sequence listed in Table 2B; or (b) the one or more guide polynucleotides comprise the following reference sequence: FcRn amino acid sequence AESHLSLLYHLTAVSSPAPGTPAFWVSGWLGPQQYLSYNSLRGEAEPCGAWVWENQVSWYWEKETTDLRIKE KLFLEAFKALGGKGPYTLQGLLGCELGPDNTSVPTAKFALNGEEFMNFDLKQGTWGGDWPEALAISQRWQQQD KAANKELTFLLFSCPHRLREHLERGRGNLEWKEPPSMRLKARPSSPGFSVLTCSAFSFYPPELQLRFLRNGL AAGTGQGDFGPNSDGSFHASSSLTVKSGDEHHYCCIVQHAGLAQPLRVELESPAKSSVLVVGIVIGVLLLTAA and targeting the base editor to effect a nucleobase change in a codon encoding an amino acid residue selected from the group consisting of: W131, F110, L112, N113, E115, E116, F117, M118, N119, D121, L122, T126, W127, G128, D130, P132, E133, A134, L135, and 1137 relative to AVGGALLWRRMRSGLPAPWISLRGDDTGVLLPTPGEAQDADLKDVNVIPATA (SEQ ID NO: 530), or a corresponding position in another FcRn polypeptide sequence, thereby altering the nucleobase of the FcRn polynucleotide.
2. The method of claim 1 , wherein the nucleobase changes result in one or more of the following amino acid changes in the FcRn polypeptide encoded by the FcRn polynucleotide relative to the reference sequence: W131R, W131Q, F110L, F110S, F110P, L112P, N113S, N113D, . E115G, E115K, E116G, E116K, E116Q, F117P, M118N, M118V, M118I, M118T, N119G, N119D, N119S, N119C, D121G, L122F, L122A, L12 2P, T126I, T126S, T126N, T126A, W127R, G128S, D130G, D130N, D130H, P132L, P132S, P132P, E133G, A134V, L135P, I137V, I137T.
3. 2. The method of Claim 1, wherein said one or more guide polynucleotides target said base editor to effect a nucleobase change in a codon encoding amino acid M118 or W131 in said reference sequence.
4. the one or more amino acid changes in the FcRn polypeptide reduce or eliminate binding of the FcRn polypeptide to IgG1, IgG2, IgG3, and / or IgG4; and the FcRn polypeptide comprising the one or more amino acid alterations has a K in solution for binding to IgG1, IgG2, IgG3, and / or IgG4 of greater than 3000 nM D 4. The method according to claim 2, wherein the
5. the FcRn polypeptide encoded by the FcRn polynucleotide comprising the altered nucleobase is capable of binding to albumin; and 3. The method of claim 2, wherein the FcRn polypeptide comprising the one or more amino acid alterations has a KD in solution for binding to albumin of less than 2000 nM, less than 1000 nM, or less than 500 nM.
6. 2. The method of Claim 1, wherein the nucleobases of the FcRn polynucleotide are altered with a base editing efficiency of at least about 20%, at least about 40%, or at least about 50%.
7. 2. The method of claim 1, wherein the deaminase domain is an adenosine deaminase domain or a cytidine deaminase domain, and the adenosine deaminase converts a target A-T in the FcRn polynucleotide to a G-C, and the cytidine deaminase converts a target C-G in the FcRn polynucleotide to a T-A.
8. The cytidine deaminase domain is an APOBEC deaminase domain or a derivative thereof; and the adenosine deaminase domain is a TadA deaminase domain; The method of claim 7.
9. The method described in claim 1, wherein the base editor is a BE4 base editor.
10. The adenosine deaminase is TadA * 8.1, TadA * 8.2, TadA * 8.
3. TadA * 8.4, TadA * 8.5, TadA * 8.6, TadA * 8.7, TadA * 8.8, TadA * 8.9, TadA * 8.10, TadA * 8.11, TadA * 8.12, TadA * 8.13, TadA * 8.14, TadA * 8.15, TadA * 8.16, TadA * 8.17, TadA * 8.18, TadA * 8.19, TadA * 8.20, TadA * 8.21, TadA * 8.22, TadA * 8.23, or TadA * 8.24, or 9. The method of claim 8, wherein the Tad*9 variant is a Tad*9 variant.
11. 2. The method of claim 1, wherein the napDNAbp domain comprises a Cas9, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ polynucleotide, or a functional portion thereof.
12. 2. The method of claim 1, wherein the napDNAbp domain comprises a nuclease-active Cas9, a dead Cas9 (dCas9), or a Cas9 nickase (nCas9).
13. 2. The method of claim 1, wherein the napDNAbp domain is Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof.
14. 14. The method of claim 13, wherein the napDNAbp domain comprises a variant of SpCas9 or SaCas9 with altered protospacer adjacent motif (PAM) specificity.
15. 15. The method of Claim 14, wherein the SpCas9 or SaCas9 has specificity for a PAM sequence selected from the group consisting of NGG, NGA, NGC, NNGRRT, and NNNRRT, wherein N is any nucleotide and R is A or G.
16. 2. The method of Claim 1, wherein the base editor further comprises one or more uracil glycosylase inhibitors (UGIs), or wherein the method further comprises expressing a UGI in the cell in trans with the base editor.
17. 2. The method of Claim 1, wherein the base editor further comprises one or more nuclear localization signals (NLS).
18. 10. The method of claim 1, wherein the one or more guide polynucleotides comprise a scaffold comprising one of the following nucleotide sequences: GUUUUAGAGCUAGAAAAUAGCAAGUUAAAAAAAAAGGCUAGUCCGUUAUCAACUUGAAAAAAGUGGCACCGAGUCGGUGCU UUU (SpCas9 scaffold; SEQ ID NO: 317), or GUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGAUUUU (SaCas9 scaffold; SEQ ID NO: 436).
19. the one or more guide polynucleotides comprise one or more modified polynucleotides; and the one or more modified polynucleotides are at the 5' and / or 3' ends of the one or more guide polynucleotides; and, 2. The method of claim 1, wherein the one or more modified nucleotides are 2'-O-methyl-3'-phosphorothioate nucleotides.
20. 10. The method of claim 1, wherein the one or more guide polynucleotides comprise a spacer consisting of 19 to 23 nucleotides.
21. The method of claim 1, wherein the FcRn polynucleotide is in a hepatic cell, an endothelial cell, a myeloid cell, or an epithelial cell.
22. 1. A base editor system for altering nucleobases in IgG receptor and transporter Fc fragment (FcRn) polynucleotides, the base editor system comprising: (i) one or more guide polynucleotides, or one or more polynucleotides encoding the one or more guide polynucleotides; and (ii) a base editor comprising a nucleic acid programmable DNA binding protein (napDNAbp) domain and a deaminase domain, or one or more polynucleotides encoding the base editor; (a) the one or more guide polynucleotides comprise a nucleic acid sequence comprising at least 10 to 23 contiguous nucleotides of a spacer nucleic acid sequence listed in Table 2B; or (b) the one or more guide polynucleotides comprise the following reference sequence: FcRn amino acid sequence AESHLSLLYHLTAVSSPAPGTPAFWVSGWLGPQQYLSYNSLRGEAEPCGAWVWENQVSWYWEKETTDLR IKEKLFLEAFKALGGKGPYTLQGLLGCELGPDNTSVPTAKFALNGEEFMNFDLKQGTWGGDWPEALAIS QRWQQQDKAANKELTFLLFSCPHRLREHLERGRGNLEWKEPPSMRLKARPSSPGFSVLTCSAFSFYPPE LQLRFLRNGLAAGTGQGDFGPNSDGSFHASSSLTVKSGDEHHYCCIVQHAGLAQPLRVELESPAKSSVLV VGIVIGVLLLTAAAVGGALLWRRMRSGLPAPWISLRGDDTGVLLPTPGEAQDADLKDVNVIPATA (SEQ ID NO: 530), or a corresponding position in another FcRn polypeptide sequence, wherein the base editor is targeted to effect a nucleobase change in a codon encoding an amino acid residue selected from the group consisting of W131, F110, L112, N113, E115, E116, F117, M118, N119, D121, L122, T126, W127, G128, D130, P132, E133, A134, L135, and I137.
23. The nucleobase changes result in the following amino acid changes in the FcRn polypeptide encoded by the FcRn polynucleotide relative to the reference sequence: W131R, W131Q, F110L, F110S, F110P, L112P, N113S, N113D, .
23. The base editor system of claim 22, wherein the base editor system results in one or more of: E115G, E115K, E116G, E116K, E116Q, F117P, M118N, M118V, M118I, M118T, N119G, N119D, N119S, N119C, D121G, L122F, L122A, L122P, T126I, T126S, T126N, T126A, W127R, G128S, D130G, D130N, D130H, P132L, P132S, P132P, E133G, A134V, L135P, I137V, I137T.
24. 23. The base editor system of Claim 22, wherein the one or more guide polynucleotides target the base editor to effect a nucleobase change in a codon encoding amino acid M118 or W131 in the reference sequence.
25. A vector comprising the polynucleotide of claim 22, The vector is a viral vector, and The viral vector is a retroviral vector or an adeno-associated viral vector.
26. The method of claim 1, wherein one or more guide polynucleotides comprise a sequence listed in Table 2A or Table 2B.
27. A method according to claim 1 or 26, wherein the alteration of FcRn does not interfere with the half-life of albumin.