Compositions and methods for editing mutation to permit transcription or expression
Patent Information
- Application Number
- JP2025135451
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-08-29
- Filing Date
- 2025-08-15
- Publication Date
- 2026-02-25
AI Technical Summary
There is currently no cure for Shwachman-Diamond syndrome (SDS), a rare autosomal recessive multisystem disorder characterized by exocrine pancreatic insufficiency, impaired hematopoiesis, and a predisposition to leukemia, with patients typically surviving only to about age 35 due to frequent hospitalizations and complications.
A method involving the use of a base editor, comprising a polynucleotide programmable DNA binding domain and a deaminase domain, to introduce targeted mutations in the SBDS gene, allowing for the production of a functional gene product by altering stop codons or splice sites, using adenosine or cytidine deaminases to correct aberrant splicing associated with SDS.
This approach aims to treat SDS by correcting gene mutations, potentially leading to the expression of a functional SBDS protein, thereby alleviating the symptoms and improving patient survival and quality of life.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Background technology]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application was filed on August 29, 2019, the contents of which are incorporated herein by reference in their entirety. International Patent Application No. 62 / 893,638, filed on 2004, which claims priority to and benefit of U.S. Provisional Patent Application No. 62 / 893,638, filed on 2004, This is a CT application.
[0002] Shwachman-Diamond syndrome (SDS) is a condition characterized by exocrine pancreatic insufficiency, impaired hematopoiesis, and SDS is a rare autosomal recessive multisystem disorder characterized by a predisposition to leukemia. Patients with HIV infection present with bone marrow failure. Other clinical manifestations include skeletal, immune, and liver damage. Approximately 90% of patients with the clinical picture of SDS have a genetic mutation on chromosome 7. The evolutionarily conserved Shwachman-Bodian-Diamond syndrome (S The SDBS protein has a biallelic mutation in the BDS gene. Although its molecular function remains unknown, it is thought to be involved in the biogenesis of ribosomes and the stability of the mitotic spindle. Currently, there is no cure for SDS, and patients with the disorder typically However, they are frequently hospitalized due to complications and, on average, only survive to about age 35. Therefore, there is an urgent need for improved methods and therapeutic agents for treating SDS. Summary of the Invention
[0003] As described below, the present invention provides a method for treating Shwachman-Diamond syndrome (SDS). The gene associated with the gene is programmed to be spliced to produce a functional gene product. Products, compositions, and methods for editing this gene using a MABUL nucleobase editor It is characterized by law.
[0004] In one aspect, a method for editing a polynucleotide to allow transcription is provided, The method further comprises: injecting a polynucleotide into a base editor in complex with one or more guide polynucleotides; and contacting the nucleotides, wherein the base editor comprises a polynucleotide programmer. a DNA binding domain and a deaminase domain, and one or more guide polypeptides Nucleotide modifications are targeted by base editors to introduce mutations that allow transcription. In one embodiment, the mutation that allows transcription is a mutation that changes a stop codon, Mutations that introduce splice acceptor or splice donor sites, or These are mutations that modify the splice acceptor or splice donor site.
[0005] In one embodiment, an S gene comprising a mutation associated with Shwachman-Diamond Syndrome (SDS) is A method for editing a BDS polynucleotide is provided, the method comprising: contacting an SBDS polynucleotide with a base editor in complex with the polynucleotide; wherein the base editor comprises a polynucleotide programmable DNA binding domain. and a deaminase domain, and one or more guide polynucleotides Targeting gene editors to treat Shwachman-Diamond syndrome (SDS)-associated In one embodiment of this method or an embodiment thereof, the method comprises: Mutations associated with Mann-Diamond syndrome (SDS) result from gene conversion. or in one of its embodiments, Shwachman-Diamond syndrome (SD) Mutations associated with S) induce stop codons in the gene or disrupt splicing. In one embodiment of the method or its embodiments, the Shwachman Diagram Mutations associated with SDS encode a truncated SBDS polypeptide. do.
[0006] In one embodiment of any of the above methods and embodiments thereof, the deaminase is a cytoplasmic deaminase. In one embodiment, the deaminase is adenosine deaminase or adenosine deaminase. In some embodiments, the adenosine deaminase is Select from ABE8 or ABE8 variants listed in Table 7A or Table 7B in the specification In another embodiment of the above method and its embodiments, the deaminase is a cytoplasmic deaminase. In one embodiment, the cytosine deaminase is BE4, rAPOB EC1, PpAPOBEC1, PpAPOBEC1 containing the H122A substitution, AmAP OBEC1, SsAPOBEC2, RrA3F, RrA3F containing the F130L substitution, A variant of BE4 in which APOBEC-1 is replaced by the sequence of rAPOBEC1, AP A variant of BE4 in which OBEC-1 is replaced by the sequence of AmAPOBEC1, APO A variant of BE4 in which BEC-1 is replaced by the sequence of SsAPOBEC2, APOB A variant of BE4 in which EC-1 is replaced by the sequence of PpAPOBEC1, or AP B, in which OBEC-1 was replaced with the sequence of PpAPOBEC1 containing the H122A substitution In one embodiment, the H122A variant is selected from one or more of the following variants of E4: PpAPOBEC1 containing the substitution, or APOBEC-1 containing the H122A substitution The BE4 variants, in which the PpAPOBEC1 sequence was replaced, were R33A, W90 One or more selected from F, K34A, R52A, H121A, or Y120F In some embodiments of the above method and its embodiments, Two or more guide polynucleotides are used to target base editors and It results in two or more mutational changes associated with Mann-Diamond syndrome (SDS).
[0007] In another embodiment, the gene includes a mutation associated with Shwachman-Diamond Syndrome (SDS). A method for editing an SBDS polynucleotide is provided, the method comprising: The SBDS polynucleotide is then attached to an adenosine base editor (ABE) in complex with the nucleotide polynucleotide. and contacting a nucleotide with the base editor, wherein the base editor is a polynucleotide program. A rammablable DNA binding domain and a deaminase domain, and one or more guide Polynucleotides target base editors to 183-184TA>CT Rs resulting in an A·T to G·C change at 113993991, generating a missense mutation In one embodiment, the one or more guide polynucleotides have one of the following sequences: Targets: TGTAAATGTTTCCTAAGGTC or AATGTTCCT In one embodiment, the one or more sgRNAs have the sequence: Contains one of: UGUAAAUGUUUCCUAAGGUC or AAUGUUUCC In one embodiment, the ABE is 5'-NGC-3' or 5'- It has PAM specificity of NGG-3'.
[0008] In another embodiment, the gene includes a mutation associated with Shwachman-Diamond Syndrome (SDS). A method for editing an SBDS polynucleotide, the method comprising: SBDS polynucleotide to a cytidine base editor in complex with the polynucleotide contacting the cytidine base editor (CBE) with the polynucleotide protease; a ramaable DNA binding domain and a cytidine deaminase domain, The guide polynucleotide targets the base editor to rs113993993 258+2T>C, resulting in a C·G to T·A change. In one embodiment, the CBE , specificity for 5'-NGC-3' PAM or PAM containing 5'-NGC-3' In one embodiment, the guide polynucleotide has the specificity GTAAGCAGGCG GGTAACAGCTGC,AGCAGGCGGGTAACAGCTGCAGC,GCG GGTAACAGCTGCAGCATAGC, GTAAGCAGGCGGGTAACAG C, AGCAGGCGGGTAACAGCTGC, GCGGGTAACAGCTGCAG CAT, GCAGGCGGGTAACAGCTGC, CAGGCGGGTAACAGCT GC, AGGCGGGTAACAGCTGC, or AAGCAGCGGGTAACA In one embodiment, the polynucleotide target sequence is selected from the group consisting of: GCTGC. The gRNA contains one of the following sequences: GUAAGCAGGCGGGUAACAGC ;AGCAGGCGGGUAACAGCUGC;GCGGGUAACAGCUGCAGC A;GCAGGCGGGUAACAGCUGC, CAGGCGGGUAACAGCUGC , AGGCGGGUAACAGCUGC, or AAGCAGGCGGGUAACAGC UGC.
[0009] In other embodiments of any of the above methods and embodiments thereof, the contacting is within a cell, The cell is a eukaryotic cell, a mammalian cell, or a human cell. In one embodiment, the cell is an in vivo or ex vivo. In this study, base editors introduced missense mutations and created new splice acceptors or Inserts a splice donor site and / or splice acceptor site containing mutations Any of the above methods and embodiments thereof may be used to modify the splice donor site or splice donor site. In one embodiment, the polynucleotide programmable DNA binding domain is a streptococcal Coccus pyrogenes Cas9 (SpCas9), Staphylococcus aureus Ca s9 (SaCas9), Streptococcus thermophilus 1Cas9 (St1Cas 9), Streptococcus canis Cas9 (ScCas9), or their variants In one embodiment, the polynucleotide programmable D The NA-binding domain is expressed in wild-type or engineered Streptococcus pyrogenes Cas9. (SpCas9), or a variant thereof. The programmable DNA-binding domain is a modified SpCas9 or SpCas9 variant. In one embodiment, the polynucleotide programmable DNA binding domain is Modified SpCas9 or SpCas9 with altered spacer adjacent motif (PAM) specificity In one embodiment, the Cas9 variant comprises a Cas9 variant 5′- of the PAM nucleic acid sequence. In one embodiment, Sp Cas9 is a PAM nucleic acid sequence 5'-NGC-3' or a PAM containing 5'-NGC-3'. A modified SpCas9 or SpCas9 variant with specificity for a nucleic acid sequence. In one embodiment, the modified SpCas9 or SpCas9 variant is selected from those listed in Table 1. In one embodiment, the modified SpCas9 comprises an amino acid sequence designated spCas9-M. In one embodiment, the modified SpCas9 or SpCas9 barrier The compounds include combinations of amino acid substitutions shown in Figures 3A-3C or Figure 10. In one embodiment, the modified SpCas9 or SpCas9 variant is selected from: Combinations of amino acid sequence substitutions include: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 , R1335E, and T1337R (224 SpCas9); D1135M, S11 36Q, G1218K, E1219F, A1322R, D1332A, R1335E, and and T1337R(225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 K, R1335E, and T1337R(226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 A, R1335E, and T1337Q(227 Cas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 A, R1335Q, and T1337Q(230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1 136, G1218T, E1219W, A1322R, D1332, R1335N, and and T1337(237 SpCas9); D1135H, S1136, G1218S, E 1219W, A1322R, D1332, R1335V, and T1337(242 S pCas9);D1135C, S1136W, G1218N, E1219W, A1322 R, D1332, R1335N, and T1337(244 SpCas9);D113 LM, S1136W, G1218R, E1219S, A1322R, D1332, R13 35E, and T1337(245 SpCas9); D1135G, S1136W, G 1218S, E1219M, A1322R, D1332, R1335Q, and T133 7R(259 SpCas9);L111R, D1135V, S1136Q, G1218 K, E1219F, A1322R, D1332, R1335A, and T1337R(N ureki SpCas9);D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337(NGCRd 1 SpCas9); or D1135G, S1136, S1216G, G1218, E1219, A1322R, D 1332A, R1335E, and T1337R (267 (NGCRd2SpCas9 ).
[0010] In other embodiments of any of the above methods and their embodiments, a polynucleotide program The MABBLE DNA-binding domain is a nuclease-inactive or nickase variant. In one embodiment, the nickase variant has the amino acid substitution D10A or its corresponding In one embodiment, the deaminase domain comprises an amino acid substitution that In one embodiment, adenosine or cytosine in DNA can be deaminated. adenosine deaminase or cytidine deaminase are modified adenosine deaminases that do not occur naturally. In one embodiment, the enzyme is adenosine deaminase or cytidine deaminase. The deaminase is a TadA deaminase. In one embodiment, the TadA deaminase , TadA*7.10, TadA*8.1, TadA*8.2, TadA*8.3, Ta dA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8 .8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.1 2, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.1 6, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.2 0, TadA*8.21, TadA*8.22, TadA*8.23, or TadA* In one embodiment, TadA*7.10 is one of the following modifications: Multiple including: Y147T, Y147R, Q154S, Y123H, V82S, T166R , Q154R.
[0011] In one embodiment, TadA*7.10 is Y147R+Q154R+Y123H;Y1 47R+Q154R+I76Y;Y147R+Q154R+T166R;Y147T+Q 154R; Y147T + Q154S; V82S + Q154S; and Y123H + Y14 It comprises a combination of alterations selected from the group consisting of: 7R+Q154R+I76Y.
[0012] In another embodiment of any of the above-described methods and embodiments thereof, one or more guides R NA is a gene encoding CRISPR RNA (crRNA) and transcoding small RNA (trancoding small RNA). crRNA), wherein the crRNA is an SBDS nucleic acid sequence containing SDS-related modifications. In another embodiment of any of the above methods and embodiments thereof, a salt The base editor contains a nucleic acid sequence complementary to the SBDS nucleic acid sequence containing the SDS-related modifications. It is in a complex with a single guide RNA (sgRNA) containing
[0013] In another embodiment, a polynucleotide programmable DNA binding domain and a deaminase domain a base editor comprising a domain; a polynucleotide encoding the base editor; and targeting base editors to produce changes associated with aberrant splicing. and introducing one or more guide polynucleotides into the cell or its precursor. In one embodiment, the cell or a precursor thereof is produced by: In one embodiment, the cells are SB stem cells, induced pluripotent stem cells, or hematopoietic stem cells. In one embodiment, the cells express Shwachman-Diamond disease (DS) protein. In one embodiment, the cells are derived from a subject suffering from Syndrome Disorders (SDS). The cell is a human cell. In one embodiment of the cell, the mutation or alteration is in aberrant splicing. In one embodiment, the mutation results from a gene conversion containing a stop codon and / or mutation resulting from In one embodiment, the cells are selected for gene conversion associated with SDS. The nucleotide-programmable DNA-binding domain is expressed in wild-type or engineered Streptococcus In one embodiment, the vector is Caspyrogenic Cas9 (SpCas9), or a variant thereof. In this study, the polynucleotide programmable DNA binding domain was constructed using a protospacer-adjacent domain. Wild-type SpCas9 or modified SpCas9 with altered chief (PAM) specificity In one embodiment, the modified SpCas9 comprises the nucleic acid sequence 5'-NGC-3' or or 5'-NGC-3'. In one embodiment, the modified SpCas9 is a Cas9 variant listed in Table 1. In one embodiment, the modified SpCas9 is spCas9-MQKFRAER. In embodiments, the modified SpCas9 contains the amino acid sequence shown in Figures 3A-3C or 10. In one embodiment of the cell, the SpC As9 variants include amino acid sequence / substitution combinations selected from the following: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 , R1335E, and T1337R (224 SpCas9); D1135M, S11 36Q, G1218K, E1219F, A1322R, D1332A, R1335E, and and T1337R(225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 K, R1335E, and T1337R(226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 A, R1335E, and T1337Q(227 Cas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 A, R1335Q, and T1337Q(230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1 136, G1218T, E1219W, A1322R, D1332, R1335N, and and T1337(237 SpCas9); D1135H, S1136, G1218S, E 1219W, A1322R, D1332, R1335V, and T1337(242 S pCas9);D1135C, S1136W, G1218N, E1219W, A1322 R, D1332, R1335N, and T1337(244 SpCas9);D113 LM, S1136W, G1218R, E1219S, A1322R, D1332, R13 35E, and T1337(245 SpCas9); D1135G, S1136W, G 1218S, E1219M, A1322R, D1332, R1335Q, and T133 7R(259 SpCas9);L111R, D1135V, S1136Q, G1218 K, E1219F, A1322R, D1332, R1335A, and T1337R(N ureki SpCas9);D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337(NGCRd 1 SpCas9); or D1135G, S1136, S1216G, G1218, E1219, A1322R, D 1332A, R1335E, and T1337R (267 (NGC Rd2 SpCa In embodiments of the above cells, the programmable polynucleotide binding domain is a nucleotide. In one embodiment, the nickase variant is a clease-inactive variant or a nickase variant. The enzyme variant comprises the amino acid substitution D10A or its corresponding amino acid substitution. In one embodiment of the cell, the deaminase domain is a cytidine deaminase domain capable of deaminating cytidine, or adenosine deaminase domain that can deaminate adenosine in DNA In one embodiment, the adenosine deaminase or cytidine deaminase is naturally occurring. The modified adenosine deaminase or cytidine deaminase does not generate the above-mentioned cell membrane. In another embodiment of the cell, the adenosine deaminase is TadA deaminase. In terms of structure, TadA deaminase is composed of TadA*7.10, TadA*8.1, and TadA *8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6 , TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, Ta dA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, Ta dA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, Ta dA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, Ta dA*8.23, or TadA*8.24. In one embodiment, TadA*7.1 0 contains one or more of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, Q154R. In one embodiment, TadA*7.10 is Y147R+Q154R+Y123H;Y147R+Q154R+I76Y;Y14 7R+Q154R+T166R;Y147T+Q154R;Y147T+Q154S;V 82S+Q154S. In the form of cytosine deaminase, BE4; rAPOBEC1, PpAPOBEC1, PpAPOBEC1, AmAPOBEC1, and SsAPOBEC containing the H122A substitution 2. RrA3F, containing the F130L substitution, which inhibits APOBEC-1 A variant of BE4 in which the sequence of EC1, APOBEC-1, is replaced by AmAPOBE A variant of BE4 in which the sequence of C1 was replaced by APOBEC-1, SsAPOBEC A variant of BE4 in which APOBEC-1 was replaced with the sequence of PpAPOBEC1 or a variant of BE4 in which the sequence of APOBEC-1 is replaced by the H122A substitution One of the variants of BE4 was replaced with a PpAPOBEC1 sequence containing or more than one. In one embodiment, a PpAPOBE containing an H122A substitution is C1, or the sequence of PpAPOBEC1, in which APOBEC-1 contains the H122A substitution The replaced BE4 variants are R33A, W90F, K34A, R52A, and H1 21A, or Y120F. In another embodiment of the above cell, the one or more guide RNAs are CRISPR RNA (crRNA) and trans-encoded small RNA (tracrRNA), The NA comprises a nucleic acid sequence complementary to the SBDS nucleic acid sequence, including the modifications associated with SDS. In one embodiment of the cell, a base editor and one or more guide polynucleotides In one embodiment, the base editor is associated with SDS. A single guide RNA ( It is in a complex with the sgRNA.
[0014] In another embodiment, Shwachman-Diamond syndrome (SDS) or aberrant splicing A method for treating a disease associated with steroid use in a subject in need thereof is provided, The method includes administering to a subject cells according to the above aspects and described embodiments thereof. In one embodiment of this method, the cells are autologous, allogeneic, or xenogeneic to the subject. do.
[0015] In another aspect, a method for growing or expanding a cell according to the above aspect and described embodiments thereof is provided. Enlarged, isolated cells or cell populations are provided.
[0016] In another aspect, a method for treating Shwachman-Diamond Syndrome (SDS) in a subject is provided. A method is provided, the method comprising combining a polynucleotide programmable DNA binding domain with a deamidating a base editor comprising a nucleotide sequence encoding a base editor, and targeting base editors to alter SDS-associated mutations. administering one or more guide polynucleotides to a subject in need thereof. .
[0017] In another aspect, a method for treating a genetic disorder associated with aberrant splicing in a subject is provided. The method comprises: a base editor comprising the domain, or a polynucleotide encoding the base editor and targeting base editors to alter pathogenic mutations that alter splicing. and administering one or more guide polynucleotides that result in the expression of the target gene to a subject in need thereof. Includes:
[0018] the aforementioned method of treating Shwachman-Diamond Syndrome (SDS) in a subject; or the above-mentioned method of treating a genetic disease associated with aberrant splicing in a subject. In one embodiment, the subject is a mammal or a human. a polynucleotide encoding a base editor or base editor, and one or more and a guide polynucleotide to a cell of the subject. In one embodiment of the above method, the alteration is to express a truncated polypeptide. In another embodiment of this method, the modification is to convert the TAA stop to TGG in the nucleotide. , altering K62X in the SBDS polypeptide relative to SDS. In this manner, gene conversion associated with SDS results in the expression of a truncated SBDS polypeptide. In another embodiment of the method, the base editor modification results in In another embodiment of the method, lysine (K) is replaced with tryptophan (W). Nucleotide-programmable DNA-binding domains are engineered to bind to engineered Streptococcus pyogenes. In another embodiment of the method, the method comprises: , the polynucleotide programmable DNA binding domain is a protospacer adjacent motif. In one embodiment, the modified SpCas9 has altered PAM specificity. pCas9 is a PAM nucleic acid sequence 5'-NGC-3' or a PA containing 5'-NGC-3'. In one embodiment, the modified SpCas9 has specificity for the M nucleic acid sequence. In one embodiment, the modified SpCas9 is a Cas9 variant that can be used to In another embodiment of these methods, the modified SpCas9 SpCas9 vectors containing the combinations of amino acid substitutions shown in Figures 3A to 3C or Figure 10 In one embodiment, the SpCas9 variant is an amino acid selected from the following: Combinations of amino acid sequence substitutions include: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 , R1335E, and T1337R (224 SpCas9); D1135M, S11 36Q, G1218K, E1219F, A1322R, D1332A, R1335E, and and T1337R(225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 K, R1335E, and T1337R(226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 A, R1335E, and T1337Q(227 Cas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 A, R1335Q, and T1337Q(230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1 136, G1218T, E1219W, A1322R, D1332, R1335N, and and T1337(237 SpCas9); D1135H, S1136, G1218S, E 1219W, A1322R, D1332, R1335V, and T1337(242 S pCas9);D1135C, S1136W, G1218N, E1219W, A1322 R, D1332, R1335N, and T1337(244 SpCas9);D113 LM, S1136W, G1218R, E1219S, A1322R, D1332, R13 35E, and T1337(245 SpCas9); D1135G, S1136W, G 1218S, E1219M, A1322R, D1332, R1335Q, and T133 7R(259 SpCas9);L111R, D1135V, S1136Q, G1218 K, E1219F, A1322R, D1332, R1335A, and T1337R(N ureki SpCas9);D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337(NGCRd 1 SpCas9); or D1135G, S1136, S1216G, G1218, E1219, A1322R, D 1332A, R1335E, and T1337R (267 (NGC Rd2 SpCa In other embodiments of the above method and its embodiments, the polynucleotide programmer The blue DNA binding domain is a nuclease-inactive variant. In embodiments, the polynucleotide programmable DNA binding domain comprises a nickase barrier In one embodiment, the nickase variant has the amino acid substitution D10A or In one embodiment of the above method, the deaminase domain comprises the corresponding amino acid substitution of: It can deaminate adenosine or cytidine in deoxyribonucleic acid (DNA). In one embodiment, the deaminase domain comprises a non-naturally occurring modified adenosine deaminase. aminase or cytidine deaminase. In one embodiment, adenosine deaminase In one embodiment, the TadA deaminase is TadA *7.10, TadA*8.1, TadA*8.2, TadA*8.3, TadA*8. 4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, Ta dA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, Tad A*8.13, TadA*8.14, TadA*8.15, TadA*8.16, Tad A*8.17, TadA*8.18, TadA*8.19, TadA*8.20, Tad A*8.21, TadA*8.22, TadA*8.23, or TadA*8.24 In one embodiment, TadA*7.10 contains one or more of the following modifications: Y147 Includes T, Y147R, Q154S, Y123H, V82S, T166R, and Q154R ;or TadA*7.10 is Y147R+Q154R+Y123H;Y147R+ Q154R+I76Y;Y147R+Q154R+T166R;Y147T+Q154R ;Y147T+Q154S;V82S+Q154S;and Y123H+Y147R+Q 154R+I76Y. In another embodiment of the present invention, the deaminase domain is selected from the group consisting of BE4, rAPOBEC1, PpAPOBEC1, PpAPOBEC1 containing the H122A substitution, AmAPOBEC 1, SsAPOBEC2, RrA3F, RrA3F containing F130L substitution, APOB A variant of BE4 in which EC-1 is replaced with the sequence of rAPOBEC1, APOBEC A variant of BE4 in which -1 was replaced with the sequence of AmAPOBEC1, APOBEC- A variant of BE4 in which 1 was replaced with the sequence of SsAPOBEC2, APOBEC-1 A variant of BE4 in which the sequence of PpAPOBEC1 has been replaced, or APOBEC The BE4 variant in which -1 was replaced with the sequence of PpAPOBEC1 containing the H122A substitution In one embodiment, the cytidine deaminase is selected from one or more of the following: In this study, PpAPOBEC1 containing the H122A substitution, or APOBEC-1, was found to be H12 A variant of BE4 replaced with the sequence of PpAPOBEC1 containing the 2A substitution was Selected from R33A, W90F, K34A, R52A, H121A, or Y120F and further comprising one or more amino acid mutations that result in the formation of a polypeptide of the present invention. In this embodiment, the base editor modifies the SNP rs11399 in the SBDS polynucleotide sequence. 3993 258+2T>C is targeted to restore correct splicing. In one embodiment of the method, the one or more guide polynucleotides are CRISPR RNase A (CRISPR RNase B) or A (crRNA) and transcoding small RNA (tracrRNA), The RNA comprises a nucleic acid sequence complementary to an SBDS nucleic acid sequence that comprises the gene conversion. The base editor is a gene encoding a SBDS nucleic acid sequence containing the gene conversion associated with the SBDS. It is in a complex with a single guide RNA (sgRNA) that contains a complementary nucleic acid sequence.
[0019] In another aspect, a method of producing a cell or a precursor thereof is provided, the method comprising: (a) A gene transformation-related gene associated with Shwachman-Diamond syndrome (SDS) Induced pluripotent stem cells Polynucleotide-programmable nucleotide-binding domain and cytidine deaminers and a base editor or base editor comprising an adenosine deaminase domain or an adenosine deaminase domain. a polynucleotide encoding a diluent; and One or more targeting base editors to alter the mutations associated with SDS Guide polynucleotide to introduce; and (b) differentiating the induced pluripotent stem cells or progenitors into a desired cell type. In one embodiment of this method, the mutation is a gene conversion associated with SDS. In one embodiment, the cells or precursors are obtained from a subject suffering from SDS. The cell or precursor is a mammalian cell or a human cell. A nucleotide-programmable DNA-binding domain is expressed in Streptococcus pyrogenes Cas9 (SpCas9), modified Streptococcus pyrogenes Cas9 (SpCa s9), or variants thereof. Ramablable DNA-binding domains alter protospacer adjacent motif (PAM) specificity In one embodiment of this method, the SpCas9 comprises a modified SpCas9 having the nucleic acid sequence The modified SpCas9 has specificity for the nucleic acid sequence 5'-NGG-3'. It has specificity for PAM nucleic acid sequences containing 5'-NGC-3' or 5'-NGC-3'. In one embodiment of the method, the modified SpCas9 is a Cas9 variant listed in Table 1. Alternatively, the modified SpCas9 is spCas9-MQKFRAER. In this embodiment, the modified SpCas9 has the amino acid sequence shown in Figures 3A-3C or 10. In one embodiment of the method, the SpCas9 variant comprises a combination of SpC As9 variants contain a combination of amino acid sequence substitutions selected from the following: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 , R1335E, and T1337R (224 SpCas9); D1135M, S11 36Q, G1218K, E1219F, A1322R, D1332A, R1335E, and and T1337R(225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 K, R1335E, and T1337R(226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 A, R1335E, and T1337Q(227 Cas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 A, R1335Q, and T1337Q(230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1 136, G1218T, E1219W, A1322R, D1332, R1335N, and and T1337(237 SpCas9); D1135H, S1136, G1218S, E 1219W, A1322R, D1332, R1335V, and T1337(242 S pCas9);D1135C, S1136W, G1218N, E1219W, A1322 R, D1332, R1335N, and T1337(244 SpCas9);D113 LM, S1136W, G1218R, E1219S, A1322R, D1332, R13 35E, and T1337(245 SpCas9); D1135G, S1136W, G 1218S, E1219M, A1322R, D1332, R1335Q, and T133 7R(259 SpCas9);L111R, D1135V, S1136Q, G1218 K, E1219F, A1322R, D1332, R1335A, and T1337R(N ureki SpCas9);D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337(NGCRd 1 SpCas9); or D1135G, S1136, S1216G, G1218, E1219, A1322R, D 1332A, R1335E, and T1337R (267 (NGC Rd2 SpCa s9). In one embodiment of this method, the polynucleotide programmable DNA binding domain is a nuclease-inactive or nickase variant. The case variant contains the amino acid substitution D10A or its corresponding amino acid substitution. In one embodiment of the method, the adenosine deaminase domain is a deoxyribonucleic acid (DNA ) can deaminate adenosine in cytidine deaminase, which deoxyribonucleotides It can deaminate cytosines in nucleic acids (DNA). The adenosine deaminase is a modified adenosine deaminase that does not occur in nature. In this state, adenosine deaminases TadA*7.10, TadA*8.1, and TadA *8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6 , TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, Ta dA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, Ta dA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, Ta dA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, Ta TadA deaminase selected from TadA*8.23, TadA*8.24, or TadA*8.25. In another embodiment of this method, the deaminase domain is selected from the group consisting of BE4, rAPOBEC1, P pAPOBEC1, PpAPOBEC1 containing the H122A substitution, AmAPOBEC1 , SsAPOBEC2, RrA3F, RrA3F containing F130L substitution, APOBE A variant of BE4 in which C-1 was replaced with the sequence of rAPOBEC1, APOBEC- A variant of BE4 in which 1 was replaced with the sequence of AmAPOBEC1, APOBEC-1 A variant of BE4 in which the sequence of APOBEC-1 was replaced by that of SsAPOBEC2. A variant of BE4 replaced with the sequence of PpAPOBEC1, or APOBEC- 1 was replaced with the sequence of PpAPOBEC1 containing the H122A substitution. In one embodiment, the cytidine deaminase is selected from one or more of the following: In this study, PpAPOBEC1 containing the H122A substitution, or APOBEC-1, was found to be H12 A variant of BE4 replaced with the sequence of PpAPOBEC1 containing the 2A substitution was Selected from R33A, W90F, K34A, R52A, H121A, or Y120F In one embodiment of this method, the one or more amino acid mutations The guide polynucleotide is transduced with CRISPR RNA (crRNA). The crRNA contains a small RNA (tracrRNA) that encodes a gene related to SDS. In one embodiment of the method, the nucleic acid sequence is complementary to the SBDS nucleic acid sequence containing the base The editor and one or more guide polynucleotides form a complex within the cell. In one embodiment of this method, the base editor comprises a SDS-associated gene conversion. A complex with a single guide RNA (sgRNA) containing a nucleic acid sequence complementary to the BDS nucleic acid sequence. It is in this state.
[0020] In another embodiment, the 5' to 3' nucleic acid sequence is selected from one or more of the following: or a 5' truncated fragment of the guide RNA by 1, 2, 3, 4, or 5 nucleotides. Provided are:GUAAGCAGGCGGGUAACAGC;AGCAGGCGGGUA ACAGCUGC;GCGGGUAACAGCUGCAGCAU;UGUAAAUGUU UCCUAAGGUC;AAUGUUUCCUAAGGUCAGGU, GCAGGCGG GUAACAGCUGC, CAGGCGGGUAACAGCUGC, AGGCGGGUA ACAGCUGC, and AAGCAGGCGGGUAACAGCUGC.
[0021] In another embodiment, a base editor system for editing pathogenic mutations in the SBDS gene is provided. The base editor system provides: (a)(i) a polynucleotide-programmable DNA-binding domain; (ii) the polynucleotide or its complementary nucleobase present in the SBDS gene conversion; a deaminase domain capable of deaminating base editors, including: (b) A guide polynucleotide in association with a polynucleotide programmable DNA binding domain. a peptide comprising at least a portion of the SBDS gene, an SBDS pseudogene, or a linker thereof; Directing the base editor to a target polynucleotide sequence located on the base complementary strand Guide polynucleotide Includes; SBDS by deaminating the polynucleotide or its complementary nucleobases This allows for gene transcription.
[0022] In another embodiment, a gene is prepared by editing a mutation in a gene that results in aberrant splicing. A base editor system is provided for: (a)(i) a polynucleotide-programmable DNA-binding domain; (ii) a mutation or its complementary nucleobase that results in aberrant splicing a deaminase domain capable of deaminating; and (b) A guide polynucleotide in association with a polynucleotide programmable DNA binding domain. a nucleic acid fragment, at least a portion of which is located in the gene or its reverse complement. A guide polynucleotide that directs a base editor to a target polynucleotide sequence. Includes; Mutation or deamination of the complementary nucleobase allows transcription.
[0023] In another embodiment, a pathogenic mutation in a gene that results in aberrant splicing is edited. A method for performing a method for detecting a stoichiometric amount of a compound is provided, the method comprising: A target nucleotide sequence at least partially located in the gene or its reverse complement. Column, (i) a target polynucleotide at least a portion of which is located in the gene or its reverse complement; and linking the base editor to a guide polynucleotide that targets the nucleic acid sequence. a polynucleotide programmable DNA binding domain; (ii) a pathogenic mutation or its complementary nucleic acid that results in aberrant splicing a deaminase domain capable of deaminating bases; contacting a base editor comprising: When targeting a base editor against a target nucleotide sequence, pathogenic mutations or Editing pathogenic mutations by deaminating complementary nucleobases Including, Deamination of the pathogenic mutation or its complementary nucleobase results in Conversion of the pathogenic mutation into a sequence that allows splicing occurs, thereby It will be corrected.
[0024] In another aspect, a method of editing a pathogenic mutation in an SBDS gene is provided, the method comprising: A target nucleotide sequence at least partially located in the gene or its reverse complement. Column, (i) a target polynucleotide at least a portion of which is located in the gene or its reverse complement; and linking the base editor to a guide polynucleotide that targets the nucleic acid sequence. a polynucleotide programmable DNA binding domain; (ii) a pathogenic mutation or a deamination capable of deaminating its complementary nucleobase Nase domain and contacting a base editor comprising: Targeting a base editor to a target nucleotide sequence results in the identification of a pathogenic mutation or its related mutations. Editing pathogenic mutations by deaminating complementary nucleobases Including, Deamination of the pathogenic mutation or its complementary nucleobase disrupts splicing. This allows for the editing of pathogenic mutations in the SBDS gene. In one embodiment of the described method, a pathogenic mutation in the SBDS results in a gene conversion. In one embodiment, the pathogenic mutation introduces or splices a stop codon in the gene. In one embodiment, the pathogenic mutation alters the coding sequence of a polypeptide having a truncation. In one embodiment, the base editor introduces a missense mutation or creates a new splice. Insert a splice acceptor site or splice donor site, or create a splice containing a mutation. In one embodiment, the salt modifies the splice acceptor site or the splice donor site. The base editor identified a splice containing the rs113993993 C→T mutation in the SBDS gene. Correct the rice donor SNP site.
[0025] In another embodiment, the method comprises the steps of: (a) generating a pathogenic mutation in the SBDS gene in a subject; A method of treating DS is provided, the method comprising: A base editor or a polynucleotide encoding a base editor is provided. administering to a subject a base editor comprising: (i) a polynucleotide-programmable DNA-binding domain; (ii) the pathogenic mutation or its complementary nucleobase is capable of deaminating the nucleobase deaminase domain that can administering; administering a guide polynucleotide to a subject, wherein the guide polynucleotide: a target nucleotide sequence at least a portion of which is located in said gene or its reverse complement; administering a base editor to target the Targeting a base editor to a target nucleotide sequence results in the generation of a disease within the SBDS gene. The pathogenic mutation is eliminated by deaminating the pathogenic mutation or its complementary nucleobase. Editing Including, Deamination of the pathogenic mutation or its complementary nucleobase allows transcription. or correct the pathogenic mutation.
[0026] In another embodiment, correcting a pathogenic mutation in the SBDS gene of a cell, tissue, or organ. and a cell, tissue, or A method for producing an organ is provided, the method comprising: Cells, tissues, or organs (i) a polynucleotide programmable DNA binding domain; (ii) a pathogenic mutation or a deamination capable of deaminating its complementary nucleobase Nase domain and contacting a base editor comprising: contacting a cell, tissue, or organ with a guide polynucleotide, The polynucleotide is located at least in part in the gene or its reverse complement. Directing or contacting a base editor to the target nucleotide sequence; and Targeting a base editor to a target nucleotide sequence results in the identification of a pathogenic mutation or its related mutations. Editing the mutation by deaminating the complementary nucleobase. Including, by deaminating the pathogenic mutation or its complementary nucleobase, resulting in splicin This allows for the production of cells, tissues, or organs for treating SDS. In one embodiment, the mutation results from a gene conversion. Mutations associated with Lancaster-Diamond syndrome either induce a stop codon in the gene or In another embodiment, the gene alters splicing. Mutations associated with (SDS) encode truncated SBDS polypeptides. In embodiments, the base editor introduces a missense mutation or creates a new splice access splice accessors that insert or contain mutations in the promoter or splice donor sites In another embodiment, the method comprises: In one embodiment of this method, the method comprises administering to the subject a cell, tissue, or organ. The tissue or organ may be autologous, allogeneic, or xenogeneic to the subject. In this embodiment, the deaminase domain is a cytidine deaminase domain or an adenosine deaminase domain. In one embodiment, the adenosine deaminase domain is It can deaminate adenine in deoxyribonucleic acid (DNA) and cytidine deamination Aminases can deaminate cytosines in DNA.
[0027] The base editor system, editing method, or treatment method described above, and In one embodiment of any of the embodiments, the guide polynucleotide is a ribonucleic acid (RNA ) or deoxyribonucleic acid (DNA). In one embodiment of the method of collecting or treating, and any of its embodiments, The CRISPR RNA (crRNA) sequence, the transactivating C RISPR RNA (tracrRNA) sequences, or combinations thereof, A comprises a nucleic acid sequence complementary to the SBDS nucleic acid sequence containing the modifications associated with SDS. Base editor systems, or editing or treatment methods, and embodiments thereof In one embodiment of any of the aspects, the base editor system or method comprises a second guidepoint In one embodiment, the second guide polynucleotide further comprises a ribonucleotide. In another embodiment, the second group comprises a nucleic acid (RNA) or a deoxyribonucleic acid (DNA). The CRISPR RNA (crRNA) sequence, the transactivating C RISPR RNA (tracrRNA) sequences, or combinations thereof. Base editor system, editing method, or treatment method, and embodiments thereof In one embodiment of any of the above, the polynucleotide programmable DNA binding domain comprises: The base editor system described above has no nuclease activity or is a nickase. In one embodiment of the present invention, a method for editing or treating a gene, and any of the embodiments thereof, In one embodiment, the polynucleotide programmable DNA binding domain comprises a Cas9 domain. In one embodiment, the Cas9 domain is a nuclease-free Cas9 (dCas9). , Cas9 nickase (nCas9), or nuclease-active Cas9. In embodiments, the Cas9 domain comprises a Cas9 nickase. A method for editing or a method for treating, and any of the embodiments thereof. In embodiments, the polynucleotide programmable DNA binding domain is engineered The polynucleotide programmable DNA binding domain is a polynucleotide programmable DNA binding domain that has been modified or altered. Base editor system, editing method, or treatment method, and embodiments thereof In one embodiment of any of the above, the editing results in less than 20% indel forms. formation, less than 15% indel formation, less than 10% indel formation, less than 5% indel formation, <4% indel formation, <3% indel formation, <2% indel formation, <1% 0.5% indel formation, or less than 0.1% indel formation The base editor system, editing method, or treatment method described above, and and in one embodiment of any of the embodiments thereof, the editing results in a rearrangement. The base editor system, editing method, or treatment method described above, and and any of the embodiments thereof, the base editor is s113993993 Corrects the splice donor SNP site containing the C→T mutation.
[0028] In another embodiment, Shwachman-Diamond syndrome (SDS) is a condition requiring A method of treatment in a subject is provided, the method comprising the above-described aspects and their described implementations. The method includes administering cells in the form of a medicament to a subject.
[0029] The above-mentioned method and embodiments thereof, the above-mentioned cell and embodiments thereof, or the above-mentioned base enzyme Editor systems and embodiments thereof, or for editing, treating, producing cells, tissues, etc. In one embodiment of any of the above-described methods and embodiments thereof, a base editor and and / or its components are encoded by mRNA. the above-mentioned cell and embodiments thereof, or the above-mentioned base editor system and embodiments thereof. or the above-described methods and embodiments thereof for editing, treating, producing cells, tissues, etc. In another embodiment of any of the above, the base enzyme according to any one of claims 126 to 157 is a base editor system or method, wherein the base editor is a nucleic acid sequence complementary to the SBDS nucleic acid sequence. In one embodiment, the nucleic acid sequence is in a complex with a single guide RNA (sgRNA). The sgRNA comprises at least 10 contiguous nucleotides complementary to the SBDS nucleic acid sequence. In another embodiment, the sgRNA comprises a nucleic acid sequence that is complementary to the SBDS nucleic acid sequence. 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 4 In another embodiment, the sgRNA comprises a nucleic acid sequence comprising: a nucleic acid sequence containing 18, 19, or 20 contiguous nucleotides complementary to the SBDS nucleic acid sequence Contains the acid sequence.
[0030] In another aspect, a composition is provided comprising a base editor bound to a guide RNA, Guide RNAs are essential for the transcription of the SBDS gene associated with Shwachman-Diamond syndrome (SDS). In one embodiment, the base editor comprises a nucleic acid sequence complementary to the gene. In one embodiment, the enzyme is adenosine deaminase or cytidine deaminase. , which can deaminate adenosine in deoxyribonucleic acid (DNA). In this state, adenosine deaminases TadA*7.10, TadA*8.1, and TadA *8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6 , TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, Ta dA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, Ta dA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, Ta dA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, Ta Tad selected from one or more of dA*8.23 or TadA*8.24 In one embodiment, the cytidine deaminase is a deoxyribonucleic acid ( In another embodiment, the cytidine deamination can be performed by deaminating cytidine residues in DNA. In one embodiment of the composition, the amino acid sequence is APOBEC, A3F, or a derivative thereof. , base editor, (i) containing Cas9 nickase; (ii) containing nuclease-inactive Cas9; (iii) SpC containing a combination of amino acid substitutions shown in Figures 3A to 3C or 10 Includes as9 variants; (iv) an SpCas9 variant comprising a combination of amino acid sequence substitutions selected from: Includes: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 , R1335E, and T1337R (224 SpCas9); D1135M, S11 36Q, G1218K, E1219F, A1322R, D1332A, R1335E, and and T1337R(225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 K, R1335E, and T1337R(226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 A, R1335E, and T1337Q(227 Cas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 A, R1335Q, and T1337Q(230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332 A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1 136, G1218T, E1219W, A1322R, D1332, R1335N, and and T1337(237 SpCas9); D1135H, S1136, G1218S, E 1219W, A1322R, D1332, R1335V, and T1337(242 S pCas9);D1135C, S1136W, G1218N, E1219W, A1322 R, D1332, R1335N, and T1337(244 SpCas9);D113 LM, S1136W, G1218R, E1219S, A1322R, D1332, R13 35E, and T1337(245 SpCas9); D1135G, S1136W, G 1218S, E1219M, A1322R, D1332, R1335Q, and T133 7R(259 SpCas9);L111R, D1135V, S1136Q, G1218 K, E1219F, A1322R, D1332, R1335A, and T1337R(N ureki SpCas9);D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337(NGCRd 1 SpCas9); or D1135G, S1136, S1216G, G1218, E1219, A1322R, D 1332A, R1335E, and T1337R (267 (NGC Rd2 SpCa s9). (v) does not contain a UGI domain; and / or (vi) BE4, rAPOBEC1, PpAPOBEC1, containing the H122A substitution PpAPOBEC1, AmAPOBEC1, SsAPOBEC2, RrA3F, F130 RrA3F, which contains the L substitution, is a mutant of APOBEC-1, where the sequence of APOBEC-1 is replaced by that of rAPOBEC1. A variant of BE4, in which APOBEC-1 is replaced with the sequence of AmAPOBEC1 A variant of BE4 in which APOBEC-1 was replaced with the sequence of SsAPOBEC2 A variant of BE4, in which APOBEC-1 is replaced by the sequence of PpAPOBEC1 Variants of E4 or PpAPOB in APOBEC-1 containing the H122A substitution A cytidine deaminator selected from variants of BE4 replaced with the sequence of EC1. In one embodiment of the composition, (vi) comprises a PpAP containing an H122A substitution. OBEC1, or PpAPOBEC1, where APOBEC-1 contains the H122A substitution The BE4 variants replaced by the sequence R33A, W90F, K34A, and R52A , H121A, or Y120F. In one embodiment, the composition further comprises a pharmaceutically acceptable excipient, diluent, or carrier. This includes:
[0031] In another aspect, a pharmaceutical composition for the treatment of Shwachman-Diamond Syndrome (SDS) is provided. and a pharmaceutical composition is provided, the pharmaceutical composition comprising the composition of the above aspects and embodiments, and In one embodiment of the pharmaceutical composition, the gRNA and and the base editor may be formulated together or separately. In one embodiment of the pharmaceutical composition, The gRNA may comprise a 5' to 3' nucleic acid sequence selected from one or more of the following: includes 5' truncated fragments of 1, 2, 3, 4, or 5 nucleotides thereof: GCGGGUAACAGCUGCAGCAU;UGUAAAUGUUCCUAAGGU C;AAUGUUUCCUAAGGUCAGGU, GCAGGGCGGUAACAGCU GC, CAGGCGGGUAACAGCUGC, AGGCGGUUAACAGCUGC, and AAGCAGGCGGGUAACAGCUGC. In one embodiment, the pharmaceutical composition comprises: Further included are vectors suitable for expression in mammalian cells, the vectors encoding a base editor. In one embodiment of the pharmaceutical composition, the pharmaceutical composition comprises a polynucleotide encoding a base editor. In one embodiment of the pharmaceutical composition, the vector is a viral vector. In one embodiment, the viral vector is a retroviral vector. adenoviral vector, lentiviral vector, herpesvirus vector, or In one embodiment, the pharmaceutical composition is an adeno-associated viral vector (AAV). Further included are ribonucleic particles suitable for expression in cells.
[0032] In one embodiment, (i) a nucleic acid encoding a base editor; and (ii) a guide according to any of the above embodiments. RNA, e.g., having a 5' to 3' nucleus selected from one or more of the following: a nucleic acid sequence or a 5' truncated fragment thereof of 1, 2, 3, 4, or 5 nucleotides and a pharmaceutical composition comprising: CAGC;AGCAGGCGGGUAACAGCUGC;GCGGGUAACAGCUG CAGCAU;UGUAAAUGUUUCCUAAGGUC;AAUGUUUCCUAA GGUCAGGU,GCAGGCGGGUAAACAGCUGC,CAGGCGGUAA CAGCUGC, AGGCGGGUAACAGCUGC, and AAGCAGGCGGG UAACAGCUGC. One embodiment of the pharmaceutical composition of any of the above aspects or embodiments thereof In some embodiments, the pharmaceutical composition further comprises a lipid.
[0033] In one aspect, a method of treating Shwachman-Diamond Syndrome (SDS) is provided. The method comprises administering the pharmaceutical composition of any of the above aspects and embodiments thereof to a patient in need thereof. This includes administering to a subject.
[0034] In one aspect, in treating Shwachman-Diamond Syndrome (SDS) in a subject. Use of the pharmaceutical composition of any of the above aspects and embodiments thereof is provided. In an embodiment, the subject is a human.
[0035] definition The following definitions are supplemental to those in the art and are directed to this application and any related or unrelated matters, e.g., not attributable to any commonly owned patents or applications. Any methods and materials similar or equivalent to those described herein may be incorporated by reference in their entirety in light of the present disclosure. Suitable materials and methods are described herein, although practice may be used to test. Therefore, the terminology used herein is intended to describe specific embodiments only. The present disclosure is illustrative and not intended to be limiting.
[0036] Unless otherwise defined, all technical and scientific terms used herein are defined by the It has the meaning commonly understood by one of ordinary skill in the art to which the specification pertains. provides those skilled in the art with general definitions of many of the terms used in this invention: Singleton et al. l., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994);The Cambrid ge Dictionary of Science and Technology (Walker ed., 1988);The Glossary of Gene tics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Mar ham, The Harper Collins Dictionary of Biology (1991). As used herein The following terms have the meanings ascribed to them unless otherwise specified.
[0037] In this application, the use of the singular includes the plural unless specifically stated otherwise. As used herein, the singular forms "a," "an," and "the" shall apply unless otherwise specified in the context. Unless the text clearly states otherwise, the use of "or" in this application includes plural references. Unless otherwise stated, "and / or" means "and / or." g)" and other forms such as "include," "includes," and " The use of "included" is not limiting.
[0038] As used in this specification and claims, the words "comprising" ( and any forms of including, e.g., "comprise" and "comprises" etc.), "having" (and any form of having, e.g., "ha "have" and "has," "including" (and the Any form, such as "includes" and "include" or "containing" containing" (and any form of containing, e.g., "containing" "ns" and "contain" are inclusive or open-ended, It does not exclude additional, unrecited elements or method steps. Any of the above embodiments may be implemented with respect to any method or composition of the disclosure, and vice versa. Furthermore, it is believed that the compositions of the present disclosure can be used to achieve the methods of the present disclosure. can be done.
[0039] The terms "about" or "approximately" refer to a particular value as determined by one of ordinary skill in the art. It means that the value is within an acceptable range of error, and that range depends on how the value was measured or determined. The accuracy of the measurement depends in part on the accuracy of the measurement system, i.e., the limitations of the measurement system. , may mean within 1 or more than 1 standard deviation per the practice of the art. About means a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, with specific reference to biological systems or processes, the term may be used , may mean within one digit of a certain value, for example, within 5 times or within 2 times. When used in the claims, unless otherwise stated, the term "about" means the amount of the specific This should be considered to mean that the value is within an acceptable error range.
[0040] As used herein, the terms "some embodiments," "embodiments," "one embodiment," or "other" may refer to Reference to an "embodiment" does not necessarily imply any particular feature, structure, or or features may be included in at least some embodiments of the present disclosure, but not necessarily in all embodiments. This means that it does not include.
[0041] "Adenosine deaminase" refers to the hydrolytic deamination of adenine or adenosine. In some embodiments, the present invention refers to a polypeptide or fragment thereof that is capable of catalyzing Deaminases or deaminase domains convert adenosine to inosine or deoxyribonucleic acid. Adenosine deamina catalyzes the hydrolytic deamination of adenosine to deoxyinosine. In some embodiments, the adenosine deaminase is a deoxyribonucleic acid (DN) deaminase. A) catalyzes the hydrolytic deamination of adenine or adenosine in Adenosine deaminases that can be used (e.g., engineered adenosine deaminases, The modified adenosine deaminase may be derived from any organism, such as a bacterium.
[0042] In some embodiments, the deaminase or deaminase domain is naturally occurring from an organism. In some embodiments, the deaminase or deaminase The deaminase domain is not naturally occurring. For example, in some embodiments, The aminase domain should be at least 50%, at least 55%, or at least at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, At least 97%, at least 98%, at least 99%, or at least 99.5% In some embodiments, the adenosine deaminase is the same as that of Escherichia coli, Staphylococcus aureus, Cass aureus, Salmonella typhimurium, Shewanella putrefaciens, of bacterial origin, such as Haemophilus influenzae or Caulobacter crescentus In some embodiments, the adenosine deaminase is TadA deaminase. In some embodiments, the TadA deaminase is an E. coli TadA (ecTadA) deaminase. It is a protein or a fragment thereof.
[0043] In some embodiments, the adenosine deaminase comprises an alteration in the following sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD (Also known as TadA*7.10).
[0044] In some embodiments, TadA*7.10 includes an alteration at amino acid 82 or 166. In specific embodiments, variants of the above reference sequences include one or more of the following changes: Includes: Y147T, Y147R, Q154S, Y123H, V82S, T166 R, and Q154R. The Y123H mutation reverted to Y123H TadA (wild type). This refers to the change H123Y in TadA*7.10. In another embodiment, TadA*7.10 The sequence variants are Y147R+Q154R+Y123H;Y147R+Q154R+ I76Y;Y147R+Q154R+T166R;Y147T+Q154R;Y147T +Q154S;V82S+Q154S; and Y123H+Y147R+Q154R+I 76Y.
[0045] In other embodiments, the present invention provides adenosine deaminase variants comprising deletions, e.g. , residues 149, 150, 151, 152, 153, 154, 155, 156, or 15 In another embodiment, TadA*8 is provided, which contains a C-terminal deletion starting at adenosine 7. Deaminase variants are TadA monomers containing one or more of the following modifications (e.g., For example, TadA*8): Y147T, Y147R, Q154S, Y123H, V82 In another embodiment, the adenosine deaminase variant is , a monomer containing the following modifications: Y147R+Q154R+Y123H; Y147R+ Q154R+I76Y;Y147R+Q154R+T166R;Y147T+Q154R ;Y147T+Q154S;V82S+Q154S;and Y123H+Y147R+Q 154R+I76Y. In yet other embodiments, the adenosine deaminase variant is A homodimer containing two adenosine deaminase domains, each domain having one or having more than one of the following modifications: Y147T, Y147R, Q154S, Y123H, V82S, T166R, Q154R. In another embodiment, the adenosine deaminase barrier The components consist of the wild-type adenosine deaminase domain or the TadA*7.10 domain and adenosine deaminase barrier of TadA*7.10 containing one or more of the following modifications: and a nucleotide domain (e.g., TadA*8): Y147T, Y 147R, Q154S, Y123H, V82S, T166R, Q154R. Other embodiments The adenosine deaminase variants consist of the TadA*7.10 domain and the following variants: Adenosine deaminase variants of TadA*7.10 containing TadA* 8) A heterodimer comprising: Y147R + Q154R + Y123H; Y147R +Q154R+I76Y;Y147R+Q154R+T166R;Y147T+Q154 R;Y147T+Q154S;V82S+Q154S;and Y123H+Y147R+ Q154R+I76Y.
[0046] In one embodiment, the adenosine deaminase is a compound having adenosine deaminase activity. TadA*8 comprising or consisting essentially of the following sequence or a fragment thereof: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQK KAQSSTD.
[0047] In some embodiments, TadA*8 is truncated. dA*8 has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, and 11 amino acids compared to full-length TadA*8. 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal amino acids In some embodiments, the truncated TadA*8 is missing a residue. Compare 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6 , 17, 18, 19, or 20 C-terminal amino acid residues. In TadA, the adenosine deaminase variant is full-length TadA*8.
[0048] In a specific embodiment, the adenosine deaminase heterodimer comprises a TadA*8 domain. and an adenosine deaminase domain selected from one of the following: Staphylococcus aureus (S. aureus) TadA: MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAH AEHIAIERAAKVLGSWRLEGCT LYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGS LMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKS TN Bacillus subtilis (B. subtilis) TadA: MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEML VIDEACKALGTWRLEGATLYVT LEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMN LLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKN LSE Salmonella Typhimurium (S. Typhimurium) TadA: MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEG WNRPIGRHDPTAHAEIMALRQGGL VLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIG RVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSD FFRMRRQEIK ALKKADRAEGAGPAV Shewanella putrefaciens (S. putrefaciens) TadA: MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEI LCLRSAGKKLENYRLLDATLYITLE PCAMCAGAMVHSRIARVVYGARDEKTGAAGT VVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQR AQQGIE Haemophilus influenzae F3031 (H. influenzae) TadA: MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗ AEIIALRNGAKNIQNY RLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYK TGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKR REEKKIEKALLKSLSDK Caulobacter crescentus (C. crescentus) TadA: MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAH DPTAHAEIAAMRAAAKLGNYRL TDLTLVVTLEPCAMCAGAISHARIGRVVFGADD PKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRK AKI Geobacter sulfreducens (G. sulfreducens) TadA: MSSLKKTPIRDDAYWMGKAIREAAKAAAARDEVPIGAVIVRDGAVIGRGHNLREGSN DPSAHAEMIAIRQAARRSANWRL TGATLYVTLEPCLMCMGAIILARLERVVFGCYDP KGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRR RKKAKATPALF IDERKVPPEP TadA*7.10 MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK KAQSSTD.
[0049] "Administering" as used herein refers to administering one or more of the compositions described herein. By way of example and not limitation, administering a composition to a patient or subject is referred to as administering a composition to a patient or subject. For example, injections may be intravenous (iv), subcutaneous (sc), or intradermal (id). It can be administered by intravenous, intraperitoneal (ip) injection, or intramuscular (im) injection. One or more such routes can be employed. Parenteral administration can be, for example, This can be done by bolus injection or by gradual perfusion over time. Alternatively, administration can be oral.
[0050] An "agent" is any small molecule compound, antibody, nucleic acid molecule, or polypeptide, or It means a fragment of these.
[0051] "Alteration" refers to any change that can be detected by standard art known methods, such as those described herein. Changes in the sequence, expression level, or activity of a gene or polypeptide when it is expressed (increases or decreases) As used herein, an alteration is defined as a change in expression level of 10%. , 25% change, 40% change, and 50% or more change in expression level.
[0052] "Ameliorate" means to reduce, suppress, attenuate, diminish, arrest, or stabilize the development or progression of a disease. It means to stabilize.
[0053] "Analog" means a molecule that is not identical but has similar functional or structural properties. For example, a polypeptide analog retains the biological activity of the corresponding naturally occurring polypeptide. and certain biochemical modifications that enhance the function of the analogue compared to the native polypeptide. Such biochemical modifications may improve the analog's protease resistance, membrane permeability, and The half-life or half-life may be increased, for example, without altering ligand binding. May contain amino acids.
[0054] "Base editor (BE)" or "nucleobase editor (NBE)" refers to a polynucleotide In various embodiments, the term "nucleobase modifying agent" refers to an agent that binds to a nucleic acid and has nucleobase modifying activity. The editors are designed to synthesize nucleobase-modifying polypeptides (e.g., deaminases) and guide polynucleotides. Polynucleotide programmable nucleotides linked to nucleotides (e.g., guide RNAs) In various embodiments, the agent comprises a protein having base editing activity and a binding domain. A structural domain, i.e., a sequence of bases (e.g., A, T, C, G, or In some embodiments, the biomolecular complex includes a domain capable of modifying a target molecule (e.g., a nucleotide or a U). In this embodiment, the polynucleotide programmable DNA binding domain is a deaminase domain. In one embodiment, the agent is fused or linked to one or more base editing molecules having base editing activity. In another embodiment, the base editing activity is a fusion protein comprising a plurality of domains. The protein domain that binds to the guide RNA (e.g., the RNA-binding motif on the guide RNA) Some of the proteins are linked to the RNA (through RNA-binding domains fused to the RNA deaminase and RNA-binding domains fused to the RNA deaminase). In embodiments, the domain having base editing activity can deaminate bases within a nucleic acid molecule. In some embodiments, the base editor can modify one or more In some embodiments, the base editor can deaminate bases within DNA. In some embodiments, the cytosine (C) or adenosine (A) of In this state, base editors deamidate cytosine (C) and adenosine (A) in DNA. In some embodiments, the base editor is a cytidine base editor. In some embodiments, the base editor is an adenosine base editor ( In some embodiments, the base editor is an adenosine base editor ( In some embodiments, the base editors are cytidine base editors (ABEs) and cytidine base editors (CBEs). The editor uses a nuclease-inactive Cas9 (dC) fused to adenosine deaminase. In some embodiments, the Cas9 is a circularly permuted Cas9 (e.g., sp Circularly permuted Cas9 is known in the art. , for example, Oakes et al., Cell 176, 254-267, 2019. In some embodiments, base editors contain inhibitors of base excision repair, e.g., UGI domains or dIS In some embodiments, the fusion protein is fused to a deaminase N domain. Fused Cas9 nickase and base excision repair domains such as UGI or dISN domains In other embodiments, the base editor is an abasic base editor. do.
[0055] In some embodiments, the adenosine deaminase is evolved from TadA. In embodiments, the polynucleotide programmable DNA binding domain is a CRISPR In some embodiments, the base editor is a related (e.g., Cas or Cpf1) enzyme. is a catalytically inactive Cas9 (dCas9) fused to a deaminase domain. In some embodiments, the base editor is a Ca fused to a deaminase domain. In some embodiments, the base editor is a base excision enzyme. In some embodiments, the inhibitor of base excision repair (BER) is fused to an inhibitor of base excision repair. In some embodiments, the toxic agent is a uracil DNA glycosylase inhibitor (UGI). is an inhibitor of base excision repair, and is an inhibitor of inosine base excision repair. For details, see International PCT Application No. PCT / 2017 / 045381 (WO2018 / 0270 78) and PCT / US2016 / 058344 (WO2017 / 070632) and U.S. Pat. No. 6,229,133, both of which are incorporated herein by reference in their entireties. Komor, AC, et al., “Programmable editing,” the entire contents of which are incorporated herein by reference. ing of a target base in genomic DNA without double-stranded DNA cleavage” Natur e 533, 420-424 (2016);Gaudelli, NM, et al., “Programmable base editing of A ·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017);K omor, AC, et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and produc t purity” Science Advances 3:eaao4774 (2017), and Rees, HA, et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec;19(12):770-788. See also doi: 10.1038 / s41576-018-0059-1 I want to be.
[0056] In some embodiments, the base editor (e.g., ABE8) is a circular permutation Cas9 (e.g., The adenosine deaminase is embedded within a scaffold containing a bipartite nuclear localization sequence (e.g., spCAS9). It is generated by cloning enzyme variants (e.g., TadA*8). Ring-substituted Cas9s are known in the art, see, for example, Oakes et al., Cell 176, 254- 267, 2019. Exemplary circularly permuted sequences are described below, with bolded The sequences in the base indicate the Cas9-derived sequences, the sequences in italics indicate the linker sequences, and the sequences in the bottom indicate the The lined arrangement indicates the bipartite nuclear orientation sequence.
[0057] CP5 (MSP) NGC = Pam variant containing the mutation. Regular Cas9 selects NGG. (preferably "PID" = protein interaction domain and "D10A" nickase): JPEG2025183226000002.jpg117169
[0058] In some embodiments, ABE8 is selected from a base editor from Table 7 below. In some embodiments, ABE8 is an adenosine deaminase that is evolved from TadA. In some embodiments, the adenosine deaminase variant of ABE8 comprises In some embodiments, the TadA*8 variant is a TadA*8 variant as described in Table 7 below. Nosine deaminase variants are Y147T, Y147R, Q154S, Y123H, V82S, T166R, Q154R. In various embodiments, ABE8 is Y147R+Q154 R+Y123H;Y147R+Q154R+I76Y;Y147R+Q154R+T16 6R;Y147T+Q154R;Y147T+Q154S;V82S+Q154S;and and Y123H+Y147R+Q154R+I76Y. In some embodiments, ABE8 is a monomeric construct.
[0059] In some embodiments, ABE8 is a heterodimeric construct. The BE8 base editor contains the following sequences: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHEIMALRQGGLVMQNYRLIDATL YVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQK KAQSSTD
[0060] For example, base editing compositions, systems, and methods described herein may be used The adenine base editor ABE was synthesized by the following steps: (Addgene, Watertown, MA.; Gaudelli NM, et al., Nature. 2017 Nov 23;551 (7681):464-471. doi: 10.1038 / nature24644;Koblan LW, et al., Nat Biotechnol. 2 018 Oct;36(9):843-846. doi: 10.1038 / nbt.4172). ABE nucleic acid sequence and at least 95 % or more identity to the sequence. ATATGCCAAGTACGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTATGCCCAGTACAT GACCTTATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGTGATGCGG TTTTGGCAGTACATCAATGGGCGTGGATAGCGGTTTGACTCACGGGATTTCCAAGTCTCCACCCCATTG ACGTCAATGGGAGTTTGTTTTGGCACCAAAATCAACGGGACTTTCCAAAATGTCGTAACAACTCCGCCCC ATTGACGCAAATGGGCGGTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCTGGTTTAGTGAACCGT CAGATCCGCTAGAGATCCGCGGCCGCTAATACGACTCACTATAGGGAGAGCCGCCACCATGAAACGGACA GCCGACGGAAGCGAGTTCGAGTCACCAAAGAAGAAGCGGAAAGTCTCTGAAGTCGAGTTTAGCCACGAGT ATTGGATGAGGCACGCACTGACCCTGGCAAAGCGAGCATGGGATGAAAGAGAAGTCCCCGTGGGCGCCGT GCTGGTGCACAACAATAGAGTGATCGGAGAGGGATGGAACAGGCCAATCGGCCGCCACGACCCTACCGCA CACGCAGAGATCATGGCACTGAGGCAGGGAGGCCTGGTCATGCAGAATTACCGCCTGATCGATGCCACCC TGTATGTGACACTGGAGCCATGCGTGATGTGCGCAGGAGCAATGATCCACAGCAGGATCGGAAGAGTGGT GTTCGGAGCACGGGACGCCAAGACCGGCGCAGCAGGCTCCCTGATGGATGTGCTGCACCACCCCGGCATG AACCACCGGGTGGAGATCACAGAGGGAATCCTGGCAGACGAGTGCGCCGCCCTGCTGAGCGATTTCTTTA GAATGCGGAGACAGGAGATCAAGGCCCAGAAGAAGGCACAGAGCTCCACCGACTCTGGAGGATCTAGCGG AGGATCCTCTGGAAGCGAGACACCAGGCACAAGCGAGTCCGCCACACCAGAGAGCTCCGGCGGCTCCTCC GGAGGATCCTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGGG CACGCGATGAGAGGGAGGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCGAGGGCTG GAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCTGAGACAGGGCGGCCTG GTCATGCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCG GCGCCATGATCCACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACGCAAAAACCGGCGCCGCAGG CTCCCTGATGGACGTGCTGCACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCA GATGAATGTGCCGCCCTGCTGTGCTATTTCTTTCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGG CCCAGAGCTCCACCGACTCCGGAGGATCTAGCGGAGGCTCCTCTGGCTCTGAGACACCTGGCACAAGCGA GAGCGCAACACCTGAAAGCAGCGGGGGCAGCAGCGGGGGGTCAGACAAGAAGTACAGCATCGGCCTGGCC ATCGGCACCAACTCTGTGGGCTGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGG TGCTGGGCAACACCGACCGGCACAGCATCAAGAAGAACCTGATCGGAGCCCTGCTGTTCGACAGCGGCGA AACAGCCGAGGCCACCCGGCTGAAGAGAACCGCCAGAAGAAGATACACCAGACGGAAGAACCGGATCTGC TATCTGCAAGAGATCTTCAGCAACGAGATGGCCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGT CCTTCCTGGTGGAAGAGGATAAGAAGCACGAGCGGCACCCCATCTTCGGCAACATCGTGGACGAGGTGGC CTACCACGAGAAGTACCCCACCATCTACCACCTGAGAAAGAAACTGGTGGACAGCACCGACAAGGCCGAC CTGCGGCTGATCTATCTGGCCCTGGCCCACATGATCAAGTTCCGGGGCCACTTCCTGATCGAGGGCGACC TGAACCCCGACAACAGCGACGTGGACAAGCTGTTCATCCAGCTGGTGCAGACCTACAACCAGCTGTTCGA GGAAAACCCCATCAACGCCAGCGGCGTGGACGCCAAGGCCATCCTGTCTGCCAGACTGAGCAAGAGCAGA CGGCTGGAAAATCTGATCGCCCAGCTGCCCGGCGGAGAAAGAATGGCCTGTTCGGAAACCTGATTGCCC TGAGCCTGGGCCTGACCCCAACTTCAAGAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCAGCTGAG CAAGGACACCTACGACGACGACGACCTGGACAACCTGCTGGCCCAGATCGGCGACCAGTACGGCCGACCTGTTTT CTGGCCGCCAAGAACCTGCTCGACGCCATCCTGCTGAGCGACATCCTGAGAGTGAACACCGAGATCACCA AGGCCCCCCTGAGCGCCTCTATGATCAAGAGATACGACGAGCACCACCAGGACCTGACCCTGCTGAAAGC TCTCGTGCGGCAGCAGCTGCCTGAGAAGTACAAAGAGATTTTCTTCGACCAGAGCAAGAACGGCTACGCC GGCTACATTGACGGCGGAGCCAGCCAGGAAGAGTTCTACAAGTTCATCAAGCCCATCCTGGAAAAGATGG ACGGCACCGAGGAACTGCTCGTGAAGCTGAACAGAGGACCTGCTGCGGAAGCAGCGGACCTTCGACAA CGGCAGCATCCCCCACCAGATCCACCTGGGAGAGCTGCACGCCATTCTGCGGCGGCAGGAAGATTTTTAC CCATTCCTGAAGGACAACCGGGAAAGATCGAGAAGATCCTGACCTTCCGCATCCCCTACTACGTGGGCC CTCTGGCCAGGGGAAACAGCAGATTCGCCTGGATCCAGAAAAGGCGAGGAAACCATCACCCCCTGGAA CTTCGAGGAAGTGGTGGACAAGGGCGCTTCCGCCCCAGAGCTTCATCGAGCGGATGACCACTTCGATAAG AACCTGCCCAACGAGAAGGTGCTGCCCAAGCACAGCCTGCTGTACGAGTACTTCACCGTGTATAACGAGC TGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCCGCCTTCCTGAGCGGCGAGCAGAAAAAGGC CATCGTGGACCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGACTACTTCAAG AAAATCGAGTGCTTCGACTCCGTGGAAATCTCCGGCGTGGAAGATCGGTTCAACGCCTCCCTGGGCACAT ACCACGATCTGCTGAAAATTATCAAGGACAAGGACTTCCTGGACAATGAGGAAAACGAGGACATTCTGGA AGATATCGTGCTGACCCTGACACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACCTATGCC CACCTGTTCGACGACAAAGTGATGAAGCAGCTGAAGCGGCGGAGATACACCGGCTGGGGCAGGCTGAGCC GGAAGCTGATCAACGGCATCCGGGACAAGCAGTCCGGCAAGACAATCCTGGATTTCCTGAAGTCCGACGG CTTCGCCAACAGAAACTTCATGCAGCTGATCCACGACGACAGCCTGACCTTTAAAGAGGACATCCAGAAA GCCCAGGTGTCCGGCCAGGGCGATAGCCTGCACGAGCACATTGCCAATCTGGCCGGCAGCCCCGCCATTA AGAAGGGCATCCTGCAGACAGTGAAGGTGGTGGACGAGCTCGTGAAAGTGATGGGCCGGCACAAGCCCGA GAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACCCAGAAGGGACAGAAGAACAGCCGCGAGAGA ATGAAGCGGATCGAAGAGGGCATCAAAGAGCTGGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAAACA CCCAGCTGCAGAACGAGAAGCTGTACCTGTACTACCTGCAGAATGGGCGGGATATGTACGTGGACCAGGA ACTGGACATCAACCGGCTGTCCGACTACGATGTGGACCATATCGTGCCTCAGAGCTTTCTGAAGGACGAC TCCATCGACAACAAGGTGCTGACCAGAAGCGACAAGAACCGGGGCAAGAGCGACAACGTGCCCTCCGAAG AGGTCGTGAAGAAGATGAAGAACTACTGGCGGCAGCTGCTGAACGCCAAGCTGATTACCCAGAGAAAGTT CGACAATCTGACCAAGGCCGAGAGAGGCGGCCTGAGCGAACTGGATAAGGCCGGCTTCATCAAGAGACAG CTGGTGGAAACCCGGCAGATCACAAAGCACGTGGCACAGATCCTGGACTCCCGGATGAACACTAAGTACG ACGAGAATGACAAGCTGATCCGGGAAGTGAAAGTGATCACCCTGAAGTCCAAGCTGGTGTCCGATTTCCG GAAGGATTTCCAGTTTTACAAAGTGCGCGAGATCAACAACTACCACCACGCCCACGACGCCTACCTGAAC GCCGTCGTGGGAACCGCCCTGATCAAAAAGTACCCTAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACA AGGTGTACGACGTGCGGAAGATGATCGCCAAGAGCGAGCAGGAAATCGGCAAGGCTACCGCCAAGTACTT CTTCTACAGCAACATCATGAACTTTTTCAAGACCGAGATTACCCTGGCCAACGGCGAGATCCGGAAGCGG CCTCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGTGGGATAAGGGCCGGGATTTTGCCACCGTGC GGAAAGTGCTGAGCATGCCCCAAGTGAATATCGTGAAAAAGACCGAGGTGCAGACAGGCGGCTTCAGCAA AGAGTCTATCCTGCCCAAGAGGAACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGGGACCCTAAGAAG TACGGCGGCTTCGACAGCCCCACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGGGCAAGT CCAAGAAACTGAAGAGTGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAGAAGCAGCTTCGAGAAGAA TCCCATCGACTTTCTGGAAGCCAAGGGCTACAAAGAAGTGAAAAAGGACCTGATCATCAAGCTGCCTAAG TACTCCCTGTTCGAGCTGGAAAACGGCCGGAAGAGAATGCTGGCCTCTGCCGGCGAACTGCAGAAGGGAA ACGAACTGGCCCTGCCCTCCAAATATGTGAACTTCCTGTACCTGGCCAGCCACTATGAGAAGCTGAAGGG CTCCCCCGAGGATAATGAGCAGAAACAGCTGTTTGTGGAACAGCACAAGCACTACCTGGACGAGATCATC GAGCAGATCAGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAATCTGGACAAAGTGCTGTCCGCCT ACAACAAGCACCGGGATAAGCCCATCAGAGAGCAGGCCGAGAATATCATCCACCTGTTTACCCTGACCAA TCTGGGAGCCCCTGCCGCCTTCAAGTACTTTGACACCACCATCGACCGGAAGAGGTACACCAGCACCAAA GAGGTGCTGGACGCCACCCTGATCCACCAGAGCATCACCGGCCTGTACGAGACACGGATCGACCTGTCTC AGCTGGGAGGTGACTCTGGCGGCTCAAAAAGAACCGCCGACGGCAGCGAATTCGAGCCCAAGAAGAAGAG GAAAGTCTAACCGGTCATCATCACCATCACCATTGAGTTTAAACCCGCTGATCAGCCTCGACTGTGCCTT CTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCAC TGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGT GGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCT CTATGGCTTCTGAGGCGGAAAGAACCAGCTGGGGCTCGATACCGTCGACCTCTAGCTAGAGCTTGGCGTA ATCATGGTCATAGCTGTTTCCTGTGTGAAATTGTTATCCGCTCACAATTCCACACAACATACGAGCCGGA AGCATAAAGTGTAAAGCCTAGGGTGCCTAATGAGTGAGCTAACTCACATTAATTGCGTTGCGCTCACTGC CCGCTTTCCAGTCGGGAAACCTGTCGTGCCAGCTGCATTAATGAATCGGCCAACGCGCGGGGAGAGGCGG TTTGCGTATTGGGCGCTCTTCCGCTTCCTCGCTCACTGACTCGCTGCGCTCGGTCGTTCGGCTGCGGCGA GCGGTATCAGCTCACTCAAAGGCGGTAATACGGTTATCCACAGAATCAGGGGATAACGCAGGAAAGAACA TGTGAGCAAAAGGCCAGCAAAAGGCCAGGAACCGTAAAAAGGCCGCGTTGCTGGCGTTTTTCCATAGGCT CCGCCCCCCTGACGAGCATCACAAAAATCGACGCTCAAGTCAGAGGTGGCGAAACCCGACAGGACTATAA AGATACCAGGCGTTTCCCCCTGGAAGCTCCCTCGTGCGCTCTCCTGTTCCGACCCTGCCGCTTACCGGAT ACCTGTCCGCCTTTCTCCCTTCGGGAAGCGTGGCGCTTTCTCATAGCTCACGCTGTAGGTATCTCAGTTC GGTGTAGGTCGTTCGCTCCAAGCTGGGCTGTGTGCACGAACCCCCCGTTCAGCCCGACCGCTGCGCCTTA TCCGGTAACTATCGTCTTGAGTCCAACCCGGTAAGACACGACTTATCGCCACTGGCAGCAGCCACTGGTA ACAGGATTAGCAGAGCGAGGTATGTAGGCGGTGCTACAGAGTTCTTGAAGTGGTGGCCTAACTACGGCTA CACTAGAAGAACAGTATTTGGTATCTGCGCTCTGCTGAAGCCAGTTACCTTCGGAAAAAGAGTTGGTAGC TCTTGATCCGGCAAACAAACCACCGCTGGTAGCGGTGGTTTTTTTGTTTGCAAGCAGCAGATTACGCGCA GAAAAAAAGGATCTCAAGAAGATCCTTTGATCTTTTCTACGGGGTCTGACACTCAGTGGAACGAAAACTC ACGTTAAGGGATTTTGGTCATGAGATTATCAAAAAGGATCTTCACCTAGATCCTTTTAAATTAAAAATGA AGTTTTAAATCAATCTAAAGTATATATGAGTAAACTTGGTCTGACAGTTACCAATGCTTAATCAGTGAGG CACCTATCTCAGCGATCTGTCTATTTCGTTCATCCATAGTTGCCTGACTCCCCGTCGTGTAGATAACTAC GATACGGGAGGGCTTACCATCTGGCCCCAGTGCTGCAATGATACCGCGAGACCCACGCTCACCGGCTCCA GATTTATCAGCAATAAACCAGCCAGCCGGAAGGGCCGAGCGCAGAAGTGGTCCTGCAACTTTATCCGCCT CCATCCAGTCTATTAATTGTTGCCGGGAAGCTAGAGTAAGTAGTTCGCCAGTTAATAGTTTGCGCAACGT TGTTGCCATTGCTACAGGCATCGTGGTGTCACGCTCGTCGTTTGGTATGGCTTCATTCAGCTCCGGTTCC CAACGATCAAGGCGAGTTACATGATCCCCCATGTTGTGCAAAAAAGCGGTTAGCTCCTTCGGTCCTCCGA TCGTTGTCAGAAGTAAGTTGGCCGCAGTGTTATCACTCATGGTTATGGCAGCACTGCATAATTCTCTTAC TGTCATGCCATCCGTAAGATGCTTTTCTGTGACTGGTGAGTACTCAACCAAGTCATTCTGAGAATAGTGT ATGCGGCGACCGAGTTGCTCTTGCCCGGCGTCAATACGGGATAATACCGCGCCACATAGCAGAACTTTAA AAGTGCTCATCATTGGAAAACGTTCTTCGGGGCGAAAACTCTCAAGGATCTTACCGCTGTTGAGATCCAG TTCGATGTAACCCACTCGTGCACCCAACTGATCTTCAGCATCTTTTACTTTCACCAGCGTTTCTGGGTGA GCAAAAACAGGAAGGCAAAATGCCGCAAAAAAGGGAATAAGGGCGACACGGAAATGTTGAATACTCATAC TCTTCCTTTTTCAATATTATTGAAGCATTTATCAGGGTTATTGTCTCATGAGCGGATACATATTTGAATG TATTTAGAAAAATAAACAAATAGGGGTTCCGCGCACATTTCCCCGAAAAGTGCCACCTGACGTCGACGGA TCGGGAGATCGATCTCCCGATCCCCTAGGGTCGACTCTCAGTACAATCTGCTCTGATGCCGCATAGTTAA GCCAGTATCTGCTCCCTGCTTGTGTGTTGGAGGTCGCTGAGTAGTGCGCGAGCAAAATTTAAGCTACAAC AAGGCAAGGCTTGACCGACAATTGCATGAAGAATCTGCTTAGGGTTAGGCGTTTTGCGCTGCTTCGCGAT GTACGGGCCAGATATACGCGTTGACATTGATTATTGACTAGTTATTAATAGTAATCAATTACGGGGTCAT TAGTTCATAGCCCATATATGGAGTTCCGCGTTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCC CAACGACCCCCGCCCATTGACGTCAATAATGACGTATGTTCCCATAGTAACGCCAATAGGGACTTTCCAT TGACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAGTACATCAAGTGTATC
[0061] For example, base editing compositions, systems, and methods described herein may be used The cytidine base editor (CBE) was synthesized using the following nucleic acid sequence (887 7 base pairs) (Addgene, Watertown, MA.; Komor AC, et al., 2017, Sci Adv., 3 0;3(8):eaao4774. doi: 10.1126 / sciadv.aao4774). BE4 nucleic acid sequence and at least 95% Polynucleotide sequences having the above identities are also encompassed. 1 ATATGCCAAG TACGCCCCCT ATTGACGTCA ATGACGGTAA ATGGCCCGCC TGGCATTATG 61 CCCAGTACAT GACCTTATGG GACTTTCCTA CTTGGCAGTA CATCTACGTA TTAGTCATCG 121 CTATTACCAT GGTGATGCGG TTTTGGCAGT ACATCAATGG GCGTGGATAG CGGTTTGACT 181 CACGGGGATT TCCAAGTCTC CACCCCATTG ACGTCAATGG GAGTTTGTTT TGGCACCAAA 241 ATCAACGGGA CTTTCCAAAA TGTCGTAACA ACTCCGCCCC ATTGACGCAA ATGGGCGGTA 301 GGCGTGTACG GTGGGAGGTC TATATAAGCA GAGCTGGTTT AGTGAACCGT CAGATCCGCT 361 AGAGATCCGC GGCCGCTAAT ACGACTCACT ATAGGGAGAG CCGCCACCAT GAGCTCAGAG 421 ACTGGCCCAG TGGCTGTGGA CCCCACATTG AGACGGCGGA TCGAGCCCCA TGAGTTTGAG 481 GTATTCTTCG ATCCGAGAGA GCTCCGCAAG GAGACCTGCC TGCTTTACGA AATTAATTGG 541 GGGGGCCGGC ACTCCATTTG GCGACATACA TCACAGAACA CTAACAAGCA CGTCGAAGTC 601 AACTTCATCG AGAAGTTCAC GACAGAAAGA TATTTCTGTC CGAACACAAG GTGCAGCATT 661 ACCTGGTTTC TCAGCTGGAG CCCATGCGGC GAATGTAGTA GGGCCATCAC TGAATTCCTG 721 TCAAGGTATC CCCACGTCAC TCTGTTTATT TACATCGCAA GGCTGTACCA CCACGCTGAC 781 CCCCGCAATC GACAAGGCCT GCGGGATTTG ATCTCTTCAG GTGTGACTAT CCAAATTATG 841 ACTGAGCAGG AGTCAGGATA CTGCTGGAGA AACTTTGTGA ATTATAGCCC GAGTAATGAA 901 GCCCACTGGC CTAGGTATCC CCATCTGTGG GTACGACTGT ACGTTCTTGA ACTGTACTGC 961 ATCATACTGG GCCTGCCTCC TTGTCTCAAC ATTCTGAGAA GGAAGCAGCC ACAGCTGACA 1021 TTCTTTACCA TCGCTCTTCA GTCTTGTCAT TACCAGCGAC TGCCCCCACA CATTCTCTGG 1081 GCCACCGGGT TGAAATCTGG TGGTTCTTCT GGTGGTTCTA GCGGCAGCGA GACTCCCGGG 1141 ACCTCAGAGT CCGCCACACC CGAAAGTTCT GGTGGTTCTT CTGGTGGTTC TGATAAAAAG 1201 TATTCTATTG GTTTAGCCAT CGGCACTAAT TCCGTTGGAT GGGCTGTCAT AACCGATGAA 1261 TACAAAGTAC CTTCAAAGAA ATTTAAGGTG TTGGGGAACA CAGACCGTCA TTCGATTAAA 1321 AAGAATCTTA TCGGTGCCCT CCTATTCGAT AGTGGCGAAA CGGCAGAGGC GACTCGCCTG 1381 AAACGAACCG CTCGGAGAAG GTATACACGT CGCAAGAACC GAATATGTTA CTTACAAGAA 1441 ATTTTTAGCA ATGAGATGGC CAAAGTTGAC GATTCTTTCT TTCACCGTTT GGAAGAGTCC 1501 TTCCTTGTCG AAGAGGACAA GAAACATGAA CGGCACCCCA TCTTTGGAAA CATAGTAGAT 1561 GAGGTGGCAT ATCATGAAAA GTACCCAACG ATTTATCACC TCAGAAAAAA GCTAGTTGAC 1621 TCAACTGATA AAGCGGACCT GAGGTTAATC TACTTGGCTC TTGCCCATAT GATAAAGTTC 1681 CGTGGGCACT TTCTCATTGA GGGTGATCTA AATCCGGACA ACTCGGATGT CGACAAACTG 1741 TTCATCCAGT TAGTACAAAC CTATAATCAG TTGTTTGAAG AGAACCCTAT AAATGCAAGT 1801 GGCGTGGATG CGAAGGCTAT TCTTAGCGCC CGCCTCTCTA AATCCCGACG GCTAGAAAAAC 1861 CTGATCGCAC AATTACCCGG AGAGAAGAAA AATGGGTTGT TCGGTAACCT TATAGCGCTC 1921 TCACTAGGCC TGACACCAAA TTTTAAGTCG AACTTCGACT TAGCTGAAGA TGCCAAATTG 1981 CAGCTTAGTA AGGACACGTA CGATGACGAT CTCGACAATC TACTGGCACA AATTGGAGAT 2041 CAGTATGCGG ACTTATTTTT GGCTGCCAAA AACCTTAGCG ATGCAATCCT CCTATCTGAC 2101 ATACTGAGAG TTAATACTGA GATTACCAAG GCGCCGTTAT CCGCTTCAAT GATCAAAAGG 2161 TACGATGAAC ATCACCAAGA CTTGACACTT CTCAAGGCCC TAGTCCGTCA GCAACTGCCT 2221 GAGAAATATA AGGAAATATT CTTTGATCAG TCGAAAAACG GGTACGCAGG TTATATTGAC 2281 GGCGGAGCGA GTCAAGAGGA ATTCTACAAG TTTATCAAAC CCATATTAGA GAAGATGGAT 2341 GGGACGGAAG AGTTGCTTGT AAAACTCAAT CGGAAGATC TACTGCGAAA GCAGCGGACT 2401 TTCGACAACG GTAGCATTCC ACATCAAATC CACTTAGGCG AATTGCATGC TATACTTAGA 2461 AGGCAGGAGG ATTTTTATCC GTTCCTCAAA GACAATCGTG AAAAGATTGA GAAAATCCTA 2521 ACCTTTCGCA TACCTTACTA TGTGGGACCC CTGGCCCGAG GGAACTCTCG GTTCGCATGG 2581 REPEAT AGTCCREACT AACGATTACT CCATGGAATTT TTREPEAT TGTCREADAA 2641 GGTGCGTCAG CTCAATCGTT CATCGAGAGG ATGACCAACT TTGACAAA TTTACCGAAC 2701 GAAAAGTAT TGCCTAXASSOCIATTACTT CHAPTER TGACTACTACATCHACTC 2761 ACGAAAGTTA AGTATGTCAC TGAGGGCATG CGTAAACCCG CCTTTCTAAG CGGAGAACAG 2821 AAGAAAGCAA TAGTAGATCT GTTATTCAAG ACCAACCGCA AAGTGACAGT TAAGCAATTG 2881 AAAGAGGACT ACTTTAAGAA AATTGAATGC TTCGATTCTG TCGAGATCTC CGGGGTAGAA 2941 GATCGATTTA ATGCGTCACT TGGTACGTAT CATGACCCTCC WINDOW INTERFACE 3001 GACTTCCTGG ATTACK REPEAT ATCTCTGG ATTACKGTT GACTCTTACC 3061 CTCTTTGAAG ATCGGGAAAT GATTGAGGAA AGACTAAAAAA CATACGCTCA CCTGTTCGAC 3121 GATAAGGTTA TGAAACAGTT AAAGAGGCGT CGCTATACGG GCTGGGGACG ATTGTCGCGG 3181 AAACTTATCA ACGGGATAAG AGACAAGCAA AGTGGTAAAA CTATTCTCGA TTTTCTAAAG 3241 AGCGACGGCT TCGCCAATAG GAACTTTATG CAGCTGATCC ATGATGACTC TTTAACCTTC 3301 AAAGAGGATA TACAAAAGGC ACAGGTTTCC GGACAAGGGG ACTCATTGCA CGAACATATT 3361 GCGAATCTTG CTGGTTCGCC AGCCATCAAA AAGGGCATAC TCCAGACAGT CAAAGTAGTG 3421 GATGAGCTAG TTAAGGTCAT GGGACGTCAC AAACCGGAAA ACATTGTAAT CGAGATGGCA 3481 CGCGAAAATC AAACGACTCA GAAGGGGCAA AAAAACAGTC GAGAGCGGAT REPEAT 3541 GAAGAGGGTA TTAAAGAACT GGGCAGCCAG ATCTTAAAGG AGCATCCTGT GGAAAATACC 3601 CAATTGCAGA ACGAGAAACT TTACCTCTAT TACCTACAAA ATGGAAGGGA CATGTATGTT 3661 GATCAGGAAC TGGACATAAA CCGTTTATCT GATTACGACG TCGATCACAT TGTACCCCAA 3721 TCCTTTTTGA AGGACGATTC AATCGACAAT AAAGTGCTTA CACGCTCGGA TAAGAACCGA 3781 GGGAAAAGTG ACAATGTTCC AAGCGAGGAA GTCGTAAAGA AAATGAGAA CTATTGGCGG 3841 CAGCTCCTAA ATGCGAAACT GATAACGCAA AGAAAGTTCG ATAACTTAAC TAAAGCTGAG 3901 AGGGGTGGCT TGTCTGAACT TGACAAGGCC GGATTTATTA AACGTCAGCT CGTGGAAACC 3961 CGCCAAATCA CAAAGCATGT TGCACAGATA CTAGATTCCC GAATGAATAC GAAATACGAC 4021 GAGAACGATA AGCTGATTCG GGAAGTCAAA GTAATCACTT TAAAGTCAAA ATTGGTGTCG 4081 GACTTCAGAA AGGATTTTCA ATTCTATAAA GTTAGGGAGA TAATAACTA CCACCATGCG 4141 CACGACGCTT ATCTTAATGC CGTCGTAGGG ACCGCACTCA TTAAGAAATA CCCGAAGCTA 4201 GAAAGTGAGT TTGTGTATGG TGATTACAAA GTTTATGACG TCCGTAAGAT GATCGCGAAA 4261 AGCGAACAGG AGATAGGCAA GGCTACAGCC AAATACTTCT TTTATTCTAA CATTATGAAT 4321 TTCTTTAAGA CGGAAATCAC TCTGGCAAAC GGAGAGATAC GCAAACGACC TTTAATTGAA 4381 ACCAATGGGG AGACAGGTGA AATCGTATGG GATAAGGGCC GGGACTTCGC GACGGTGAGA 4441 AAAGTTTTGT CCATGCCCCA AGTCAACATA GTAAAGAAAA CTGAGGTGCA GACCGGAGGG 4501 TTTTCAAAGG AATCGATTCT TCCAAAAAGG AATAGTGATA AGCTCATCGC TCGTAAAAAG 4561 GACTGGGACC CGAAAAAGTA CGGTGGCTTC GATAGCCCTA CAGTTGCCTA TTCTGTCCTA 4621 GTAGTGGCAA AAGTTGAGAA GGGAAAATCC AAGAAACTGA AGTCAGTCAA AGAATTATTG 4681 GGGATAACGA TTATGGAGCG CTCGTCTTTT GAAAAGAACC CCATCGACTT CCTTGAGGCG 4741 AAAGGTTACA AGGATAAA AAAGGATCTC ATTACK TACTTACA TAGTCTGTTT 4801 GAGTTAGAAA ATGGCCGAAA ACGGATGTTG GCTAGCGCCG GAGAGCTTCA AAAGGGGAAC 4861 GAACTCGCAC TACCGTCTAA ATACGTGAAT TTCCTGTATT TAGCGTCCCA TTACGAGAAG 4921 TTGAAAGGTT CACCTGAACAG AAGCAACTTT TTGTTGAGCA GCACAACAT 4981 TATCTCGACG AAATCATAGA GCAAATTTCG GAATTCAGTA AGAGAGTCAT CCTAGCTGAT 5041 GCCAATCTGG ACAAAGTATT AAGCGCATAC AACAAGCACA GGGATAAACC CATACGTGAG 5101 CAGGCGGAAA ATATTATCCA TTTGTTTACT CTTACCAACC TCGGCGCTCC AGCCGCATTC 5161 AAGTATTTTG ACACAACGAT AGATCGCAAA CGATACACTT CTACCAAGGA GGTGCTAGAC 5221 GCGACACTGA TTCACCAATC CATCACGGGA TTATATGAAA CTCGGATAGA TTTGTCACAG 5281 CTTGGGGGTG ACTCTGGTGG TTCTGGAGGA TCTGGTGGTT CTACTAATCT GTCAGATATT 5341 ATTGAAAAGG AGACCGGTAA GCAACTGGTT ATCCAGGAAT CCATCCTCAT GCTCCCAGAG 5401 GAGGTGGAAG AAGTCATTGG GAACAAGCCG GAAAGCGATA TACTCGTGCA CACCGCCTAC 5461 GACGAGAGCA CCGACGAGAA TGTCATGCTT CTGACTAGCG ACGCCCCTGA ATACAAGCCT 5521 TGGGCTCTGG TCATACAGGA TAGCAACGGT GAGAACAAGA TTAAGATGCT CTCTGGTGGT 5581 TCTGGAGGAT CTGGTGGTTC TACTAATCTG TCAGATATTA TTGAAAAGGA GACCGGTAAG 5641 CAACTGGTTA TCCAGGAATC CATCCTCATG CTCCCAGAGG AGGTGGAAGA AGTCATTGGG 5701 AACAAGCCGG AAAGCGATAT ACTCGTGCAC ACCGCCTACG ACGAGAGCAC CGACGAGAAT 5761 GTCATGCTTC TGACTAGGCGA CGCCCCTGAA TACAAGCCTT GGGCTCTGGT CATACAGGAT 5821 AGCAACGGTG AGAACAAGAT TAAGATGCTC TCTGGTGGTT CTCCCAAGAA GAAGAGGAAA 5881 GTCTAACCGG TCATCATCAC CATCACCATT GAGTTTAAAC CCGCTGATCA GCCTCGACTG 5941 TGCCTTCTAG TTGCCAGCCA TCTGTTGTTT GCCCCTCCCC CGTGCCTTCC TTGACCCTGG 6001 AAGGTGCCAC TCCCACTGTC CTTTCCTAAT AAAATGAGGA AATTGCATCG CATTGTCTGA 6061 GTAGGTGTCA TTCTATTCTG GGGGGTGGGG TGGGGCAGGA CAGCAAGGGG GAGGATTGGG 6121 AAGACAATAG CAGGCATGCT GGGGATGCGG TGGGCTCTAT GGCTTCTGAG GCGGAAAGAA 6181 CCAGCTGGGG CTCGATACCG TCGACCTCTA GCTAGAGCTT GGCGTAATCA TGGTCATAGC 6241 TGTTTCCTGT GTGAAATTGT TATCCGCTCA CAATTCCACA CAACATACGA GCCGGAAGCA 6301 TAAAGTGTAA AGCCTAGGGT GCCTAATGAG TGAGCTAACT CACATTAATT GCGTTGCGCT 6361 CACTGCCCGC TTTCCAGTCG GGAAACCTGT CGTGCCAGCT GCATTAATGA ATCGGCCAAC 6421 GCGCGGGGAG AGGCGGTTTG CGTATTGGGC GCTCTTCCGC TTCCTCGCTC ACTGACTCGC 6481 TGCGCTCGGT CGTTCGGCTG CGGCGAGCGG TATCAGCTCA CTCAAAGGCG GTAATACGGT 6541 TATCCACAGA ATCAGGGGAT AACGCAGGAA AGAACATGTG AGCAAAAGGC CAGCAAAAGG 6601 CCAGGAACCG TAAAAAGGCC GCGTTGCTGG CGTTTTTCCA TAGGCTCCGC CCCCCTGACG 6661 AGCATCACAA AAATCGACGC TCAAGTCAGA GGTGGCGAAA CCCGACAGGA CTATAAAGAT 6721 ACCAGGCGTT TCCCCCTGGA AGCTCCCTCG TGCGCTCTCC TGTTCCGACC CTGCCGCTTA 6781 CCGGATACCT GTCCGCCTTT CTCCCTTCGG GAAGCGTGGC GCTTTCTCAT AGCTCACGCT 6841 GTAGGTATCT CAGTTCGGTG TAGGTCGTTC GCTCCAAGCT GGGCTGTGTG CACGAACCCC 6901 CCGTTCAGCC CGACCGCTGC GCCTTATCCG GTAACTATCG TCTTGAGTCC AACCCGGTAA 6961 GACACGACTT ATCGCCACTG GCAGCAGCCA CTGGTAACAG GATTAGCAGA GCGAGGTATG 7021 TAGGCGGTGC TACAGAGTTC TTGAAGTGGT GGCCTAACTA CGGCTACACT AGAAGAACAG 7081 TATTTGGTAT CTGCGCTCTG CTGAAGCCAG TTACCTTCGG AAAAAGAGTT GGTAGCTCTT 7141 GATCCGGCAA ACAAACCACC GCTGGTAGCG GTGGTTTTTT TGTTTGCAAG CAGCAGATTA 7201 CGCGCAGAAA AAAAGGATCT CAAGAAGATC CTTTGATCTT TTCTACGGGG TCTGACGCTC 7261 AGTGGAACGA AAACTCACGT TAAGGGATTT TGGTCATGAG ATTATCAAAA AGGATCTTCA 7321 CCTAGATCCT TTTAAATTAA AAATGAAGTT TTAAATCAAT CTAAAGTATA TATGAGTAAA 7381 CTTGGTCTGA CAGTTACCAA TGCTTAATCA GTGAGGCACC TATCTCAGCG ATCTGTCTAT 7441 TTCGTTCATC CATAGTTGCC TGACTCCCCG TCGTGTAGAT AACTACGATA CGGGAGGCT 7501 TACCATCTGG CCCCAGTGCT GCAATGATAC CGCGAGACCC ACGCTCACCG GCTCCAGATT 7561 TATCAGCAAT AAACCAGCCA GCCGGAAGGG CCGAGCGCAG AAGTGGTCCT GCAACTTTAT 7621 CCGCCTCCAT CCAGTCTATT AATTGTTGCC GGGAAGCTAG AGTAAGTAGT TCGCCAGTTA 7681 ATAGTTTGCG CAACGTTGTT GCCATTGCTA CAGGCATCGT GGTGTCACGC TCGTCGTTTG 7741 GTATGGCTTC ATTCAGCTCC GGTTCCCAAC GATCAAGGCG AGTTACATGA TCCCCCATGT 7801 TGTGCAAAAA AGCGGTTAGC TCCTTCGGTC CTCCGATCGT TGTCAGAAGT AAGTTGGCCG 7861 CAGTGTTATC ACTCATGGTT ATGGCAGCAC TGCATAATTC TCTTACTGTC ATGCCATCCG 7921 TAAGATGCTT TTCTGTGACT GGTGAGTACT CAACCAAGTC ATTCTGAGAA TAGTGTATGC 7981 GGCGACCGAG TTGCTCTTGC CCGGCGTCAA TACGGGATAA TACCGCGCCA CATAGCAGAA 8041 CTTTAAAAGT GCTCATCATT GGAAAACGTT CTTCGGGGCG AAAACTCTCA AGGATCTTAC 8101 CGCTGTTGAG ATCCAGTTCG ATGTAACCCA CTCGTGCACC CAACTGATCT TCAGCATCTT 8161 TTACTTTCAC CAGCGTTTCT GGGTGAGCAA AAACAGGAAG GCAAAATGCC GCAAAAAAGG 8221 GAATAAGGGC GACACGGAAA TGTTGAATAC TCATACTCTT CCTTTTTCAA TATTATTGAA 8281 GCATTTATCA GGGTTATTGT CTCATGAGCG GATACATATT TGAATGTATT TAGAAAAATA 8341 AACAAATAGG GGTTCCGCGC ACATTTCCCC GAAAAGTGCC ACCTGACGTC GACGGATCGG 8401 GAGATCGATC TCCCGATCCC CTAGGGTCGA CTTCCAGTAC AATCTGCTCT GATGCCGCAT 8461 AGTTAAGCCA GTATCTGCTC CCTGCTTGTG TGTTGGAGGT CGCTGAGTAG TGCGCGAGCA 8521 AAATTTAAGC TACAACAAGG CAAGGCTTGA CCGACAATTG CATGAAGAAT CTGCTTAGGG 8581 TTAGGCGTTT TGCGCTGCTT CGCGATGTAC GGGCCAGATA TACGCGTTGA CATTGATTAT 8641 TGACTAGTTA TTAATAGTAA TCAATTACGG GGTCATTAGT TCATAGCCCA TATATGGAGT 8701 TCCGCGTTAC ATAACTTACG GTAAATGGCC CGCCTGGCTG ACCGCCCAAC GACCCCCGCC 8761 CATTGACGTC AATAATGACG TATGTTCCCA TAGTAACGCC AATAGGGACT TTCCATTGAC 8821 GTCAATGGGT GGAGTATTTA CGGTAAACTG CCCACTTGGC AGTACATCAA GTGTATC
[0062] In some embodiments, the cytidine base editor is a nucleic acid selected from one of the following: BE4 has the sequence: Native BE4 nucleic acid sequence: ATGagctcagagactggcccagtggctgtggaccccacattgagacggcggatcgagccccatgagtttgaggtattctt cgatccgagagagctccgcaaggagacctgcctgctttacgaaattaattgggggggccggcactccatttggcgacata catcacagaacactaacaagcacgtcgaagtcaacttcatcgagaagttcacgacagaaagatatttctgtccgaacaca aggtgcagcattacctggtttctcagctggagccgcgaatgtagtagggccatcactgaattcctgtcaaggtatcccca cgtcactctgtttatttacatcgcaaggctgtaccaccacgctgacccccgcaatcgacaaggcctgcgggatttgatct cttcaggtgtgactatccaaattatgactgagcaggagtcaggatactgctggagaaactttgtgaattatagcccgagt aatgaagcccactggcctaggtatccccatctgtgggtacgactgtacgttcttgaactgtactgcatcatactgggcct gcctccttgtctcaacattctgagaaggaagcagccacagctgacattctttaccatcgctcttcagtcttgtcattacc agcgactgcccccacacattctctgggccaccgggttgaaatctggtggttcttctggtggttctagcggcagcgagact cccgggacctcagagtccgccacacccgaaagttctggtggttcttctggtggttctgataaaaagtattctattggttt agccatcggcactaattccgttggatgggctgtcataaccgatgaatacaaagtaccttcaaagaaatttaaggtgttgg ggaacacagaccgtcattcgattaaaaagaatcttatcggtgccctcctattcgatagtggcgaaacggcagaggcgact cgcctgaaacgaaccgctcggagaaggtatacacgtcgcaagaaccgaatatgttacttacaagaaatttttagcaatga gatggccaaagttgacgattctttctttcaccgtttggaagagtccttccttgtcgaagaggacaagaaacatgaacggc accccatctttggaaacatagtagatgaggtggcatatcatgaaaagtacccaacgatttatcacctcagaaaaaagcta gttgactcaactgataaagcggacctgaggttaatctacttggctcttgcccatatgataaagttccgtgggcactttct cattgagggtgatctaaatccggacaactcggatgtcgacaaactgttcatccagttagtacaaacctataatcagttgt ttgaagagaaccctataaatgcaagtggcgtggatgcgaaggctattcttagcgcccgcctctctaaatcccgacggcta gaaaacctgatcgcacaattacccggagagaagaaaaatgggttgttcggtaaccttatagcgctctcactaggcctgac accaaattttaagtcgaacttcgacttagctgaagatgccaaattgcagcttagtaaggacacgtacgatgacgatctcg acaatctactggcacaaattggagatcagtatgcggacttatttttggctgccaaaaaccttagcgatgcaatcctccta tctgacatactgagagttaatactgagattaccaaggcgccgttatccgcttcaatgatcaaaaggtacgatgaacatca ccaagacttgacacttctcaaggccctagtccgtcagcaactgcctgagaaatataaggaaatattctttgatcagtcga aaaacgggtacgcaggttatattgacggcggagcgagtcaagaggaattctacaagtttatcaaacccatattagagaag atggatgggacggaagagttgcttgtaaaactcaatcgcgaagatctactgcgaaagcagcggactttcgacaacggtag cattccacatcaaatccacttaggcgaattgcatgctatacttagaaggcaggaggattttatccgttcctcaaagaca atcgtgaaaagattgagaaaatcctaacctttcgcatacctactatgtgggacccctggcccgagggaactctcggttc gcatggatgacaagaaagtccgaagaacgattactccatggaattttgaggaagttgtcgataaaggtgcgtcagctca atcgttcatcgagaggatgaccaactttgacaagaatttaccgaacgaaaaagtattgcctaagcacagtttactttacg agtatttcacagtgtacaatgaactcacgaaagttaagtatgtcactgagggcatgcgtaaacccgcctttctaagcgga gaacagaaagcaatagtagatctgttattcaagaccaaccgcaaagtgacagttaagcaattgaaaggactactt tagaaaattgaatgcttcgattctgtcgagatctccggggtagaagatcgatttaatgcgtcacttggtacgtatcatg acctcctaaagataattaaagataaggacttcctggataacgaagagaatgaagatatcttagaagatatagtgttgact cttaccctcttttgaagatcgggaaatgattgaggaaagaactaaaaacatacgctcacctgttcgacgataaggttatgaa acagttaaaagaggcgtcgctatacgggctggggacgattgtcgcgggaaacttatcaacgggataagaagacaagcaaagtg gtaaaactattctcgattttctaaagagcgacggcttcgccaataggaactttatgcagctgatccatgatgactcttta accttcaaagaggatatacaaaaggcacaggtttccggacaaggggactcattgcacgaacatattgcgaatcttgctgg ttcgccagccatcaaaaagggcatactccagacagtcaaagtagtggatgagctagttaaggtcatgggacgtcacaaac cggaaaacattgtaatcgagatggcacgcgaaaatcaaacgactcagaaggggcaaaaaaacagtcgagagcggatgaag agaatagaagagggtattaaagaactgggcagccagatcttaaaggagcatcctgtggaaaatacccaattgcagaacga gaaactttacctctattacctacaaaatggaagggacatgtatgttgatcaggaactggacataaaccgtttatctgatt acgacgtcgatcacattgtaccccaatcctttttgaaggacgattcaatcgacaataaagtgcttacacgctcggataag aaccgagggaaaagtgacaatgttccaagcgaggaagtcgtaaagaaaatgaagaactattggcggcagctcctaaatgc gaaactgataacgcaaagaaagttcgataacttaactaaagctgagaggggtggcttgtctgaacttgacaaggccggat ttattaaacgtcagctcgtggaaacccgccaaatcacaaagcatgttgcacagatactagattcccgaatgaatacgaaa tacgacgagaacgataagctgattcgggaagtcaaagtaatcactttaaagtcaaaattggtgtcggacttcagaaagga ttttcaattctataaagttagggagataaataactaccaccatgcgcacgacgcttatcttaatgccgtcgtagggaccg cactcattaagaaatacccgaagctagaaagtgagtttgtgtatggtgattacaaagtttatgacgtccgtaagatgatc gcgaaaagcgaacaggagataggcaaggctacagccaaatacttcttttatctaacattatgaatttctttaagacgga aatcactctggcaaacggagagatacgcaaacgacctttaattgaaaccaatggggagacaggtgaaatcgtatgggata agggccgggacttcgcgacggtgagaaaagttttgtccatgccccaagtcaacatagtaaagaaaactgaggtgcagacc ggagggttttcaaaggaatcgattcttccaaaaaggaatagtgataagctcatcgctcgtaaaaaggactgggacccgaa aaagtacggtggcttcgatagccctacagttgcctattctgtcctagtagtggcaaaagttgagaagggaaaatccaaga aactgaagtcagtcaaagaattattggggataacgattatggagcgctcgtcttttgaaaagaaccccatcgacttcctt gaggcgaaaggttacaaggaagtaaaaaaggatctcataattaaactaccaaagtatagtctgtttgagttagaaaatgg ccgaaaacggatgttggctagcgccggagagcttcaaaaggggaacgaactcgcactaccgtctaaatacgtgaatttcc tgtatttagcgtcccattacgagaagttgaaaggttcacctgaagataacgaacagaagcaactttttgttgagcagcac aaacattatctcgacgaaatcatagagcaaatttcggaattcagtaagagagtcatcctagctgatgccaatctggacaa agtattaagcgcatacaacaagcacagggataaacccatacgtgagcaggcggaaaatattatccatttgtttactctta ccaacctcggcgctccagccgcattcaagtattttgacacaacgatagatcgcaaacgatacacttctaccaaggaggtg ctagacgcgacactgattcaccaatccatcacgggattatatgaaactcggatagatttgtcacagcttgggggtgactc tggtggttctggaggatctggtggttctactaatctgtcagatattattgaaaaggagaccggtaagcaactggttatcc aggaatccatcctcatgctcccagaggaggtggaagaagtcattgggaacaagccggaaagcgatatactcgtgcacacc gcctacgacgagagcaccgacgagaatgtcatgcttctgactagcgacgcccctgaatacaagccttgggctctggtcat acaggatagcaacggtgagaacaagattaagatgctctctggtggttctggaggatctggtggttctactaatctgtcag atattattgaaaaggagaccggtaagcaactggttatccaggaatccatcctcatgctcccagaggagtggaagaagtc attgggaacaagccggaaagcgatatactcgtgcacaccgcctacgacgagagcaccgacgagaatgtcatgcttctgac tagcgacgcccctgaatacaagccttgggctctggtcatacagtagcaacggtgagaacaagattaagatgctctctg gtggttctAAAAGGACGGCGGACGGATCAGAGTTCGAGAGTCCGAAAAAAAAACGAAAGGTCGAAtaa BE4コドン optimization1 nucleic acid sequence: ATGTCATCCGAAACCGGGCCAGTGGCCGTAGACCAACACTCAGGAGGCGGATAGAACCCCATGAGTTTGAAGTGTTCTT CGACCCCAGAGAGCTGCGCAAAGAGACTTGCCTCCTGTATGAAATAAATTGGGGGGGTCGCCATTCAATTTGGAGGCACA CTAGCCAGAATACTAACAAACACGTGGAGGTAAATTTATCGAGAAGTTTACCACCGAAAGATACTTTTGCCCCAATACA CGGTGTTCAATTACCTGGTTTCTGTCATGGAGTCCATGTGGAGAATGTAGTAGAGCGATAACTGAGTTCCTGTCTCGATA TCCTCACGTCACGTTGTTTATATACATCGCTCGGCTTTATCACCATGCGGACCCGCGGAACAGGCAAGGTCTTCGGGACC TCATATCCTCTGGGGTGACCATCCAGATAATGACGGAGCAAGAGAGCGGATACTGCTGGCGAAACTTTGTTAACTACAGC CCAAGCAATGAGGCACACTGGCCTAGATATCCGCATCTCTGGGTTCGACTGTATGTCCTTGAACTGTACTGCATAATTCT GGGACTTCCGCCATGCTTGAACATTCTGCGGCGGAAACAACCACAGCTGACCTTTTTCACGATTGCTCTCCAAAGTTGTC ACTACCAGCGATTGCCACCCCACATCTTGTGGGCTACTGGACTCAAGTCTGGAGGAAGTTCAGGCGGAAGCAGCGGGTCT GAAACGCCCGGAACCTCAGAGAGCGCAACGCCCGAAAGCTCTGGAGGGTCAAGTGGTGGTAGTGATAAGAAATACTCCAT CGGCCTCGCCATCGGTACGAATTCTGTCGGTTGGGCCGTTATCACCGATGAGTACAAGGTCCCTTCTAAGAAATTCAAGG TTTTGGGCAATACAGACCGCCATTCTATAAAAAAAAACCTGATCGGCGCCCTTTTGTTTGACAGTGGTGAGACTGCTGAA GCGACTCGCCTGAAGCGAACTGCCAGGAGGCGGTATACGAGGCGAAAAAACCGAATTTGTTACCTCCAGGAGATTTTCTC AAATGAAATGGCCAAGGTAGATGATAGTTTTTTTCACCGCTTGGAAGAAAGTTTTCTCGTTGAGGAGGACAAAAAGCACG AGAGGCACCCAATCTTTGGCAACATAGTCGATGAGGTCGCATACCATGAGAAATATCCTACGATCTATCATCTCCGCAAG AAGCTGGTCGATAGCACGGATAAAGCTGACCTCCGGCTGATCTACCTTGCTCTTGCTCACATGATTAAATTCAGGGGCCA TTTCCTGATAGAAGGAGACCTCAATCCCGACAATTCTGATGTCGACAAACTGTTTATTCAGCTCGTTCAGACCTATAATC AACTCTTTGAGGAGAACCCCATCAATGCTTCAGGGGTGGACGCAAAGGCCATTTTGTCCGCGCGCTTGAGTAAATCACGA CGCCTCGAGAATTTGATAGCTCAACTGCCGGGTGAGAAGAAAAACGGGTTGTTTGGGAATCTCATAGCGTTGAGTTTGGG ACTTACGCCAAACTTTAAGTCTAACTTTGATTTGGCCGAAGATGCCAAATTGCAGCTGTCCAAAGATACCTATGATGACG ACTTGGATAACCTTCTTGCGCAGATTGGTGACCAATACGCGGATCTGTTTCTTGCCGCAAAAAATCTGTCCGACGCCATA CTCTTGTCCGATATACTGCGCGTCAATACTGAGATAACTAAGGCTCCCCTCAGCGCGTCCATGATTAAAAGATACGATGA GCACCACCAAGATCTCACTCTGTTGAAAGCCCTGGTTCGCCAGCAGCTTCCAGAGAAGTATAAGGAGATATTTTTCGACC AATCTAAAAACGGCTATGCGGGTTACATTGACGGTGGCGCCTCTCAAGAAGAATTCTACAAGTTTATAAAGCCGATACTT GAGAAAATGGACGGTACAGAGGAATTGTTGGTTAAGCTCAATCGCGAGGACTTGTTGAGAAAGCAGCGCACATTTGACAA TGGTAGTATTCCACACCAGATTCATCTGGGCGAGTTGCATGCCATTCTTAGAAGACAAGAAGATTTTTATCCGTTTCTGA AAGATAACAGAGAAAAAGATTGAAAAGATACTTACCTTTCGCATACCGTATTATGTAGGTCCCCTGGCTAGAGGGAACAGT CGCTTCGCTTGGATGACTCGAAAATCAGAAGAAAACAATAACCCCCTGGAATTTTGAAAGAAGTGGTAGATAAAGGTGCGAG TGCCCAATCTTTTATTGAGCGGATGACAAATTTTGACAAGAATCTGCCTAACGAAAAGGTGCTTCCCAAGCATTCCCTTT TGTATGAATACTTTACAGTATATAATGAACTGACTAAAGTGAAGTACGTTACCGAGGGGATGCGAAAGCCAGCTTTTTCTC AGTGGCGAGCAGAAAAAGCAATAGTTGACCTGCTGTTCAAGACGAATAGGAAGGTTACCGTCAAACAGCTCAAAGAAGA TTACTTTAAAAGATCGAATGTTTTGATTCAGTTGAGATAAGCGGAGTAGAGGATAGATTTAACCCAAGTCTTGGGAACTT ATCATGACCTTTTGAAGATCATCAAGGATAAAGATTTTTTGGGACAACGAGGAGAATGAAGATATCCTGGAAGATATAGTA CTTACCTTGACGCTTTTTGAAGATCGAGAGATGATCGAGGAGCGACTTAAGACGTACGCACATCTCTTTGACGATAAGGT TATGAAACAATTGAAACGCCGGCGGTATACTGGCTGGGGCAGGCTTTTCCGAAAGCTGATTAATGGTATCCGCGATAAGC AGTCTGGAAAGACAATCCTTGACTTTCTGAAAAGTGATGGATTTGCAAATAGAAACTTTATGCAGCTTATACATGATGAC TCTTTGACGTTCAAGGAAGACATCCAGAAGGCACAGGTATCCGGCCAAGGGGATAGCCTCCATGAACACATAGCCAACCT GGCCGGCTCACCAGCTATTAAAAAGGGAATATTGCAAACCGTTAAGGTTGTTGACGAACTCGTTAAGGTTATGGGCCGAC ACAAACCAGAGAATATCGTGATTGAGATGGCTAGGGAGAATCAGACCACTCAAAAAGGTCAGAAAAATTCTCGCGAAAGG ATGAAGCGAATTGAAGAGGGAATCAAAGAACTTGGCTCTCAAATTTTGAAAGAGCACCCGGTAGAAAACACTCAGCTGCA GAATGAAAAGCTGTATCTGTATTATCTGCAGAATGGTCGAGATATGTACGTTGATCAGGAGCTGGATATCAATAGGCTCA GTGACTACGATGTCGACCACATCGTTCCTCAATCTTTCCTGAAAGATGACTCTATCGACAACAAAGTGTTGACGCGATCA GATAAGAACCGGGGAAAATCCGACAATGTACCCTCAGAAGAAGTTGTCAAGAAGATGAAAAACTATTGGAGACAATTGCT GAACGCCAAGCTCATAACACAACGCAAGTTCGATAACTTGACGAAAGCCGAAAGAGGTGGGTTGTCAGAATTGGACAAAG CTGGCTTTATTAAGCGCCAATTGGTGGAGACCCGGCAGATTACGAAACACGTAGCACAAATTTTGGATTCACGAATGAAT ACCAAATACGACGAAAACGACAAATTGATACGCGAGGTGAAAGTGATTACGCTTAAGAGTAAGTTGGTTTCCGATTTCAG GAAGGATTTTCAGTTTTACAAAGTAAGAGAAATAAACAACTACCACCACGCCCATGATGCTTACCTCAACGCGGTAGTTG GCACAGCTCTTATCAAAAAATATCCAAAGCTGGAAAGCGAGTTCGTTTACGGTGACTATAAAGTATACGACGTTCGGAAG ATGATAGCCAAATCAGAGCAGGAAATTGGGAAGGCAACCGCAAAATACTTCTTCTATTCAAACATCATGAACTTCTTTAA GACGGAGATTACGCTCGCGAACGGCGAAATACGCAAGAGGCCCCTCATAGAGACTAACGGCGAAACCGGGGAGATCGTAT GGGACAAAGGACGGGACTTTGCGACCGTTAGAAAAGTACTTTCAATGCCACAAGTGAATATTGTTAAAAAGACAGAAGTA CAAACAGGGGGGTTCAGTAAGGAATCCATTTTGCCCAAGCGGAACAGTGATAAATTGATAGCAAGGAAAAAAGATTGGGA CCCTAAGAAGTACGGTGGTTTCGACTCTCCTACCGTTGCATATTCAGTCCTTGTAGTTGCGAAAGTGGAAAAGGGGAAAA GTAAGAAGCTTAAGAGTGTTAAAGAGCTTCTGGGCATAACCATAATGGAACGGTCTAGCTTCGAGAAAAATCCAATTGAC TTTCTCGAGGCTAAAGGTTACAAGGAGGTAAAAAAGGACCTGATAATTAAACTCCCAAAGTACAGTCTCTTCGAGTTGGA GAATGGGAGGAAGAGAATGTTGGCATCTGCAGGGGAGCTCCAAAAGGGGAACGAGCTGGCTCTGCCTTCAAAATACGTGA ACTTTCTGTACCTGGCCAGCCACTACGAGAAACTCAAGGGTTCTCCTGAGGATAACGAGCAGAAACAGCTGTTTGTAGAG CAGCACAAGCATTACCTGGACGAGATAATTGAGCAAATTAGTGAGTTCTCAAAAAGAGTAATCCTTGCAGACGCGAATCT GGATAAAGTTCTTTCCGCCTATAATAAGCACCGGGACAAGCCTATACGAGAACAAGCCGAGAACATCATTCACCTCTTTA CCCTTACTAATCTGGGCGCGCCGGCCGCCTTCAAATACTTCGACACCACGATAGACAGGAAAAGGTATACGAGTACCAAA GAAGTACTTGACGCCACTCTCATCCACCAGTCTATAACAGGGTTGTACGAAACGAGGATAGATTTGTCCCAGCTCGGCGG CGACTCAGGAGGGTCAGGCGGCTCCGGTGGATCAACGAATCTTTCCGACATAATCGAGAAAGAAACCGGCAAACAGTTGG TGATCCAAGAATCAATCCTGATGCTGCCTGAAGAAGTAGAAGAGGTGATTGGCAACAAACCTGAGTCTGACATTCTTGTC CACACCGCGTATGACGAGAGCACGGACGAGAACGTTATGCTTCTCACTAGCGACGCCCCTGAGTATAAACCATGGGCGCT GGTCATCCAAGATTCCAATGGGGAAAACAAGATTAAGATGCTTAGTGGTGGGTCTGGAGGGAGCGGTGGGTCCACGAACC TCAGCGACATTATTGAAAAAGAGACTGGTAAACAACTTGTAATACAAGAGTCTATTCTGATGTTGCCTGAAGAGGTGGAG GAGGTGATTGGGAACAAACCGGAGTCTGATATACTTGTTCATACCGCCTATGACGAATCTACTGATGAGAATGTGATGCT TTTaACGTCAGACGCTCCCGAGTACAAACCCTGGGCTCTGGTGATTCAGGACAGCAATGGTGAGAATAAGATTAAATGT TGAGTGGGGGCTCAAAGCGCACGGCTGACGGTAGCGAATTTGAGAGCCCCAAAAAAAAACGAAAGGTCGAAtaa BE4コドン optimization2 nucleic acid sequence: ATGAGCAGCGAGACAGGCCCTGTGGCTGTGGATCCTACACTGCGGAGAAGAATCGAGCCCCACGAGTTCGAGGTGTTCTT CGACCCCAGAGAGCTGCGGAAAGAGACATGCCTGCTGTACGAGATCAACTGGGGCGGCAGACACTCTATCTGGCGCACA CAAGCCAGAACACCAACAAGCACGTGGAAGTGAACTTTATCGAGAAGTTTACGACCGAGCGGTACTTCTGCCCCAACACC AGATGCAGCATCACCTGGTTCTGAGCTGGTCCCCTTGCGGCGAGTGCAGCAGAGCCATCACCGAGTTTCTGTCCAGATA TCCCCACGTGACCCTGTTCATCTATATCGCCCGGCTGTACCACCACGCCGATCCTAGAAATAGACAGGGACTGCGCGACC TGATCAGCAGCGGAGTGACCATCCAGATCATGACCGAGCAAGAGAGCGGCTACTGCTGGCGGAACTTCGTGAACTACAGC CCCAGCAACGAAGCCCACTGGCCTAGATATCCTCACCTGTGGGTCCGACTGTACGTGCTGGAACTGTACTGCATCATCCT GGGCCTGCCTCCATGCCTGAACATCCTGAGAAGAAAGCAGCCTCAGCTGACCTTCTTCACAATCGCCCTGCAGAGCTGCC ACTACCAGAGACTGCCTCCACACATCCTGTGGGCCACCGGACTTAAGAGCGGAGGATCTAGCGGCGGCTCTAGCGGATCT GAGACACCTGGCACAAGCGAGTCTGCCACACCTGAGAGTAGCGGCGGATCTTCTGGCGGCTCCGACAAGAAGTACTCTAT CGGACTGGCCATCGGCACCAACTCTGTTGGATGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGG TGCTGGGCAACACCGACCGGCACAGCATCAAGAAGAATCTGATCGGCGCCCTGCTGTTCGACTCTGGCGAAACAGCCGAA GCCACCAGACTGAAGAGAACCGCCAGGCGGAGATACACCCGGCGGAAGAACCGGATCTGCTACCTGCAAGAGATCTTCAG CAACGAGATGGCCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGTCCTTCCTGGTGGAAGAGGACAAGAAGCACG AGCGGCACCCCATCTTCGGCAACATCGTGGATGAGGTGGCCTACCACGAGAAGTACCCCACCATCTACCACCTGAGAAAG AAACTGGTGGACAGCACCGACAAGGCCGACCTGAGACTGATCTACCTGGCTCTGGCCCACATGATCAAGTTCCGGGGCCA CTTTCTGATCGAGGGCGATCTGAACCCCGACAACAGCGACGTGGACAAGCTGTTCATCCAGCTGGTGCAGACCTACAACC AGCTGTTCGAGGAAAACCCCATCAACGCCTCTGGCGTGGACGCCAAGGCTATCCTGTCTGCCAGACTGAGCAAGAGCAGA AGGCTGGAAAACCTGATCGCCCAGCTGCCTGGCGAGAAGAAGAATGGCCTGTTCGGCAACCTGATTGCCCTGAGCCTGGG ACTGACCCCTAACTTCAAGAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCAGCTGAGCAAGGACACCTACGACGACG ACCTGGACAATCTGCTGGCCCAGATCGGCGATCAGTACGCCGACTTGTTTCTGGCCGCCAAGAACCTGTCCGACGCCATC CTGCTGAGCGATATCCTGAGAGTGAACACCGAGATCACAAAGGCCCCTCTGAGCGCCTCTATGATCAAGAGATACGACGA GCACCACCAGGATCTGACCCTGCTGAAGGCCCTCGTTAGACAGCAGCTGCCAGAGAAGTACAAAGAGATTTTCTTCGATC AGTCCAAGAACGGCTACGCCGGCTACATTGATGGCGGAGCCAGCCAAGAGGAATTCTACAAGTTCATCAAGCCCATCCTG GAAAAGATGGACGGCACCGAGGAACTGCTGGTCAAGCTGAACAGAGAGGACCTGCTGCGGAAGCAGCGGACCTTCGACAA TGGCTCTATCCCTCACCAGATCCACCTGGGAGAGCTGCACGCCATTCTGCGGAGACAAGAGGACTTTTACCCATTCCTGA AGGACAACCGGGAAAAGATCGAGAAGATCCTGACCTTCAGGATCCCCTACTACGTGGGACCACTGGCCAGAGGCAATAGC AGATTCGCCTGGATGACCAGAAAGAGCGAGGAAACCATCACACCCTGGAACTTCGAGGAAGTGGTGGACAAGGGCGCCAG CGCTCAGTCCTTCATCGAGCGGATGACCAACTTCGATAAGAACCTGCCTAACGAGAAGGTGCTGCCCAAGCACTCCCTGC TGTATGAGTACTTCACCGTGTACAACGAGCTGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCCGCCTTTCTG AGCGGCGAGCAGAAAAAGGCCATTGTGGATCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGA CTACTTCAAGAAAATCGAGTGCTTCGACAGCGTGGAAATCAGCGGCGTGGAAGATCGGTTCAATGCCAGCCTGGGCACAT ACCACGACCTGCTGAAAATTATCAAGGACAAGGACTTCCTGGACAACGAAGAGAACGAGGACATTCTCGAGGACATCGTG CTGACCCTGACACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACATACGCCCACCTGTTCGACGACAAAGT GATGAAGCAACTGAAGCGGAGGCGGTACACAGGCTGGGGCAGACTGTCTCGGAAGCTGATCAACGGCATCCGGGATAAGC AGTCCGGCAAGACAATCCTGGATTTCCTGAAGTCCGACGGCTTCGCCAACAGAAACTTCATGCAGCTGATCCACGACGAC AGCCTGACCTTTAAAGAGGACATCCAGAAAGCCCAGGTGTCCGGCCAAGGCGATTCTCTGCACGAGCACATTGCCAACCT GGCCGGATCTCCCGCCATTAAGAAGGGCATCCTGCAGACAGTGAAGGTGGTGGACGAGCTTGTGAAAGTGATGGGCAGAC ACAAGCCCGAGAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACACAGAAGGGCCAGAAGAACAGCCGCGAGAGA ATGAAGCGGATCGAAGAGGGCATCAAAGAGCTGGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAAACACCCAGCTGCA GAACGAGAAGCTGTACCTGTACTACCTGCAGAATGGACGGGATATGTACGTGGACCAAGAGCTGGACATCAACCGGCTGA GCGACTACGATGTGGACCATATCGTGCCCCAGAGCTTTCTGAAGGACGACTCCATCGATAACAAGGTCCTGACCAGAAGC GACAAGAACCGGGGCAAGAGCGATAACGTGCCCTCCGAAGAGGTGGTCAAGAAGATGAAGAACTACTGGCGACAGCTGCT GAACGCCAAGCTGATTACCCAGCGGAAGTTCGATAACCTGACCAAGGCCGAGAGAGGCGGCCTGAGCGAACTTGATAAGG CCGGCTTCATTAAGCGGCAGCTGGTGGAAACCCGGCAGATCACCAAACACGTGGCACAGATTCTGGACTCCCGGATGAAC ACTAAGTACGACGAGAATGACAAGCTGATCCGGGAAGTGAAAGTCATCACCCTGAAGTCTAAGCTGGTGTCCGATTTCCG GAAGGATTTCCAGTTCTACAAAGTGCGGGAAATCAACAACTACCATCACGCCCACGACGCCTACCTGAATGCCGTTGTTG GAACAGCCCTGATCAAGAAGTATCCCAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACAAGGTGTACGACGTGCGGAAG ATGATCGCCAAGAGCGAACAAGAGATCGGCAAGGCTACCGCCAAGTACTTTTTCTACAGCAACATCATGAACTTTTTCAA GACAGAGATCACCCTGGCCAACGGCGAGATCCGGAAAAGACCCCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGT GGGATAAGGGCAGAGATTTTGCCACAGTGCGGAAAGTGCTGAGCATGCCCCAAGTGAATATCGTGAAGAAAACCGAGGTG CAGACAGGCGGCTTCAGCAAAGAGTCTATCCTGCCTAAGCGGAACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGGGA CCCTAAGAAGTACGGCGGCTTCGATAGCCCTACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGGGCAAGT CCAAAAAGCTCAAGAGCGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAGAAGCAGCTTTGAGAAGAACCCGATCGAC TTTCTGGAAGCCAAGGGCTACAAAGAAGTCAAGAAGGACCTCATCATCAAGCTCCCCAAGTACAGCCTGTTCGAGCTGGA AAATGGCCGGAAGCGGATGCTGGCCTCAGCAGGCGAACTGCAGAAAGGCAATGAACTGGCCCTGCCTAGCAAATACGTCA ACTTCCTGTACCTGGCCAGCCACTATGAGAAGCTGAAGGGCAGCCCCGAGGACAATGAGCAAAAGCAGCTGTTTGTGGAA CAGCACAAGCACTACCTGGACGAGATCATCGAGCAGATCAGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAACCT GGATAAGGTGCTGTCTGCCTATAACAAGCACCGGGACAAGCCTATCAGAGAGCAGGCCGAGAATATCATCCACCTGTTTA CCCTGACCAACCTGGGAGCCCCTGCCGCCTTCAAGTACTTCGACACCACCATCGACCGGAAGAGGTACACCAGCACCAAA GAGGTGCTGGACGCCACACTGATCCACCAGTCTATCACCGGCCTGTACGAAACCCGGATCGACCTGTCTCAGCTCGGCGG CGATTCTGGTGGTTCTGGCGGAAGTGGCGGATCCACCAATCTGAGCGACATCATCGAAAAAGAGACAGGCAAGCAGCTCG TGATCCAAGAATCCATCCTGATGCTGCCTGAAGAGGTTGAGGAAGTGATCGGCAACAAGCCTGAGTCCGACATCCTGGTG CACACCGCCTACGATGAGAGCACCGATGAGAACGTCATGCTGCTGACAAGCGACGCCCCTGAGTACAAGCCTTGGGCTCT CGTGATTCAGGACAGCAATGGGGAGAACAAGATCAAGATGCTGAGCGGAGGTAGCGGAGGCAGTGGCGGAAGCACAAACC TGTCTGATATCATTGAAAAAGAAACCGGGAAGCAACTGGTCATTCAAGAGTCCATTCTCATGCTCCCGGAAGAAGTCGAG GAAGTCATTGGAAACAAACCCGAGAGCGATATTCTGGTCCACACAGCCTATGACGAGTCTACAGACGAAAACGTGATGCT CCTGACCTCTGACGCTCCCGAGTATAAGCCCTGGGCACTTGTTATCCAGGACTCTAACGGGGAAAACAAAATCAAAATGT TGTCCGGCGGCAGCAAGCGGACAGCCGATGGATCTGAGTTCGAGAGCCCCAAGAAGAAACGGAAGGTgGAGtaa
[0063] "Base editing activity" refers to the ability to chemically modify bases within a polynucleotide. In one embodiment, the first base is converted to the second base. Base editing activity is, for example, cytidine deaminase activity that converts the target C·G to T·A. In another embodiment, the base editing activity is, for example, an adduct that converts A·T to G·C. In another embodiment, the base editing activity is adenine or adenine deaminase activity. For example, the cytidine deaminase activity converting the target C·G to T·A, e.g., A· It is an adenosine or adenine deaminase activity that converts T to G·C.
[0064] The term "base editor system" refers to a system that uses nucleic acid bases to edit a target nucleotide sequence. In various embodiments, the base editor (BE) system comprises: (1) Polynucleotide programmable nucleotide binding domain, target nucleotide sequence Deaminase domains and cytidine deaminase domains for deaminating nucleic acid bases (2) a polynucleotide linked to a programmable nucleotide binding domain; and one or more guide polynucleotides (e.g., guide RNAs). In this paper, the base editor (BE) system is a system that encodes adenosine deaminase or cytidine deaminase. and a nucleic acid base editor domain selected from a nucleotide sequence-specific binding domain of ... In some embodiments, the base editor system comprises: (1) a domain that encodes a polynucleotide; a nucleotide-programmable DNA binding domain and one or more target nucleotide sequences; A base editor ( BE); and (2) 1 linked to a polynucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmer comprises one or more guide RNAs. The programmable nucleotide binding domain is a polynucleotide-programmable DNA binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine or adenosine base editor ( In some embodiments, the base editor is adenine or adenosine base editor (ABE) or cytidine base editor (CBE).
[0065] The term "Cas9" or "Cas9 domain" refers to the Cas9 protein or a fragment thereof. fragments (e.g., active, inactive, or partially active DNA cleavage domains of Cas9) a protein containing the gRNA-binding domain of Cas9, and / or an RNA guide containing the gRNA-binding domain of Cas9 Cas9 nuclease is also known as casnl nuclease or C RISPR (clustered regularly interspaced short palindromic repeats)-related genes An exemplary Cas9 is a Cas9 enzyme found in Streptococcus pyrogenes. spCas9, the amino acid sequence of which is shown below: JPEG2025183226000003.jpg114169 (single underline: HNH domain; double underline: RuvC domain)
[0066] The term "conservative amino acid substitution" or "conservative mutation" refers to the substitution of one amino acid with another amino acid that shares a common characteristic. It refers to the substitution of a specific amino acid with another amino acid that has the same function as the amino acid. A practical method is to calculate the normalized frequencies of amino acid changes between corresponding proteins of homologous organisms. (Schulz, GE and Schirmer, RH, Principles of Protein Struc- ture Structure, Springer-Verlag, New York (1979). Following such analysis, the amino acid groups It is possible to define a group in which amino acids within a group preferentially exchange with each other, and thus are most similar to each other in terms of their effect on the overall protein structure (Schu Ilz, GE and Schirmer, RH, supra. Non-limiting examples of conservative mutations include amino Acid amino acid substitutions, e.g., arginine to lysine so that a positive charge can be maintained, and so on. vice versa; aspartic acid to glutamic acid and vice versa so that the negative charge can be maintained; Threonine to serine to maintain a free -OH; and One example is changing asparagine to glutamine to make it more durable.
[0067] The terms "coding sequence" or "protein-coding sequence" as used interchangeably herein " refers to a segment of a polynucleotide that encodes a protein. The sequence is bounded near the 5' end by a start codon and near the 3' end by a stop codon. Stop codons useful for the base editors described herein include: Can be: Glutamine CAG → TAG stop codon CAA → TAA Arginine CGA → TGA Tryptophan TGG → TGA TGG → TAG TGG→TAA A coding sequence can also be referred to as an open reading frame.
[0068] "Cytidine deaminase" is a enzyme that converts amino groups into carbonyl groups through a deamination reaction. In one embodiment, the term "cytochrome P450" refers to a polypeptide or fragment thereof that is capable of catalyzing a cytochrome P450 polypeptide. Cytosine deaminase converts cytosine to uracil or 5-methylcytosine to thymine. PmCDA1 (sea lamprey cytosinde) derived from sea lamprey (Petromyzon marinus) aminase 1, "PmCDA1"), mammals, or mammals of different species (e.g., humans) mammals, such as rats, pigs, cows, horses, monkeys, etc.), and non-mammals, such as alligators. AID (activation-induced cytidine deaminase; AICDA) and APOBEC An exemplary cytidine deaminase.
[0069] As used herein, the term "deaminase" or "deaminase domain" refers to a deaminase domain that Deaminase refers to a protein or enzyme that catalyzes a deaminization reaction. or the deaminase domain converts cytidine or deoxycytidine to uridine, respectively. or cytidine deaminase, which catalyzes the hydrolytic deamination to deoxyuridine. In some embodiments, the deaminase or deaminase domain converts cytosine to uracil. A cytosine deaminase that catalyzes the hydrolytic deamination of cytosine to cytosine. In this state, deaminases catalyze the hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the deaminase is an adenosine deaminase that or adenosine, which catalyzes the hydrolytic deamination of adenine (A) to inosine (I). In some embodiments, the deaminase or deaminase domain is a deaminase. converts adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of Adenosine deaminase hydrolyzes adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases described herein (e.g., engineered Adenosine deaminase (evolved adenosine deaminase) is a type of enzyme found in bacteria, In some embodiments, the adenosine deaminase is derived from E. coli, S. coli, or Taphylococcus aureus, Salmonella Typhimurium, Shewanella putrefa S. cereus, Haemophilus influenzae, or Caulobacter crescentus In some embodiments, the adenosine deaminase is of bacterial origin. In some embodiments, the deaminase or deaminase domain is human, chin Natural decoys derived from organisms such as pansies, gorillas, monkeys, cows, dogs, rats, or mice. In some embodiments, the deaminase is a variant of an aminase. The domain is not naturally occurring. For example, in some embodiments, a deaminase or deaminase domain is The deaminase domain has at least 50%, at least 55%, or at least 100% identity with the natural deaminase. At least 60%, at least 65%, at least 70%, at least 75%, at least 80 %, at least 85%, at least 90%, at least 91%, at least 92%, at least at least 93%, at least 94%, at least 95%, at least 96%, at least 9 7%, at least 98%, at least 99%, at least 99.1%, at least 99. 2%, at least 99.3%, at least 99.4%, at least 99.5%, at least Even 99.6%, at least 99.7%, at least 99.8%, or at least 99. 9% identical.
[0070] "Direct" refers to determining the presence, absence, or amount of an analyte to be detected. In embodiments, sequence alterations in polynucleotides or polypeptides are detected. In an embodiment, the presence of indels is detected.
[0071] A "detectable label" means a label that, when attached to a molecule of interest, is spectroscopically, photochemically, or biochemically detectable. "Detectable" refers to a composition that allows the latter to be detected through immunochemical, immunological, or chemical means. For example, useful labels include radioisotopes, magnetic beads, metal beads, colloidal particles, Fluorescent dyes, electron-dense reagents (e.g., commonly used in enzyme-linked immunosorbent assays (ELISAs)), Examples of suitable antibodies include enzymes (such as those used), biotin, digoxigenin, or haptens.
[0072] "Disease" means any condition that damages or interferes with the normal function of a cell, tissue, or organ. In a specific embodiment, the present invention relates to a method for treating a pulmonary artery disease or a disorder that is amenable to treatment using the compositions of the present invention. In a specific embodiment, the disease is Schwarzenegger's disease. This is Schuckman-Diamond syndrome (SDS).
[0073] "Diseases associated with abnormal splicing" are diseases that affect genes that affect splicing. Alterations in the sequence, e.g., alterations within the splice acceptor or splice donor sites The term "condition or disorder" refers to any condition or disorder associated with disruption of transcription due to, for example, a gene encoding a gene for ... encoding a gene for a gene for a gene encoding a gene for a gene encoding a gene for
[0074] An "effective amount" is an amount that reduces the severity of a disease compared to an untreated patient or a disease-free individual, i.e., a healthy individual. Agents or active compounds required to ameliorate the symptoms of the disease, such as those described herein or an amount of such a base editor sufficient to elicit a desired biological response. The amount of agent or active compound used in practicing the invention for the therapeutic treatment of disease. The effective amount of active compound to be administered will depend on the mode of administration, the age, weight, and general health of the subject. Ultimately, your doctor or veterinarian will determine the appropriate dosage and administration regimen. Such an amount is referred to as an "effective" amount. In one embodiment, an effective amount refers to the insertion of a gene of interest into a cell (e.g., a cell in vitro or in vivo). The amount of a base editor of the invention is sufficient to introduce the alteration. In one embodiment, an effective amount is the amount of base editor required to achieve a therapeutic effect. Such a therapeutic effect is It is not necessary that the gene be sufficient to alter the pathogenic gene in all cells of a subject, tissue, or organ. However, approximately 1%, 5%, 10%, 25%, 5% of the cells present in a subject, tissue, or organ In one embodiment, the effective amount is a dose that alters 0%, 75% or more of the pathogenic genes of the disease. Sufficient to relieve one or more symptoms.
[0075] In some embodiments, nCas9 may be in the form of a fusion protein as described herein. domain and deaminase domain (e.g., adenosine deaminase, cytidine deaminase) a nucleobase editor comprising a nCas9 domain and a deaminase domain, or a (e.g., adenosine deaminase, cytidine deaminase) and nucleic acid base editors, including An effective amount of an agent or composition comprising a nucleobase editor as described herein is It refers to an amount sufficient to induce editing of the heterologously linked and edited target site. As will be appreciated, the effective amount of an agent, e.g., a fusion protein, will depend on a variety of factors, e.g., For example, depending on the desired biological response, the specific alleles, genomes, or depending on the target site, depending on the cell or tissue being targeted, and / or may vary depending on the agent being used.
[0076] In some embodiments, the agent may be in the form of a fusion protein, e.g., an nCas9 domain. An effective amount of a fusion protein comprising a deaminase domain and a deaminase domain is an amount of a protein that is capable of being inhibited by the fusion protein. The target site is specifically bound and edited by the agent sufficient to induce editing, e.g., a fusion target. As will be understood by those skilled in the art, the amount of an agent, e.g., a fusion protein, , nucleases, hybrid proteins, protein dimers, protein complexes (or or protein dimer), and a polynucleotide, or an effective amount of the polynucleotide Depending on various factors, e.g., depending on the desired biological response, e.g., the specific Depending on the specific allele, genome, or target site, the targeted cell or tissue This may vary depending on the tissue and / or the agent being used.
[0077] By "fragment" is meant a portion of a polypeptide or nucleic acid molecule. at least about 10%, 20%, 30%, 40% of the total length of the nucleic acid molecule or polypeptide of Fragments containing 50%, 60%, 70%, 80%, or 90% of the total DNA fragments are 10, 20, 30, or 40% of the total DNA fragments. , 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 50 0, 600, 700, 800, 900, or 1000 nucleotides or amino acids may contain
[0078] "Guide RNA" or "gRNA" refers to a polynucleotide that is specific to a target sequence. a programmable nucleotide-binding domain protein (e.g., Cas9 or Cp In one embodiment, the term "polynucleotide" refers to a polynucleotide capable of forming a complex with f1). The guide polynucleotide is a guide RNA (gRNA). It can exist as a complex of the above RNAs or as a single RNA molecule. gRNAs that exist as NA molecules are also called single guide RNAs (sgRNAs). However, "gRNA" can be expressed as a single molecule or as a complex of two or more molecules. are used interchangeably to refer to the guide RNA present in a single RNA species. gRNAs present in the present study have two domains: (1) homology to the target nucleic acid (and e.g., Ca) (2) the Cas9 protein; and In some embodiments, domain (2) comprises a domain that binds to tracrRN. A and has a stem-loop structure. For example, in some embodiments, , domain (2) is described in Jinek et al., the entire contents of which are incorporated herein by reference. , Science 337:816-821 (2012) or Other examples of gRNAs (e.g., those containing domain 2) are "Switchable Cas9 U.S. Patent Application Publication No. 20160208288, entitled "Nucleases and Uses Thereof" and U.S. Patent No. 9,733 entitled "Delivery System For Functional Nucleases." No. 7,604, the entire contents of each of which are incorporated herein by reference in their entirety. In some embodiments, the gRNA comprises two or more domains (1) and ( 2), and may also be referred to as "extended gRNA." As described in, two or more Cas9 proteins can be linked together and expressed at two or more distinct regions. The gRNA binds to the target nucleic acid. The gRNA contains a nucleotide sequence complementary to the target site. This sequence mediates the binding of the nuclease / RNA complex to the target site and This results in sequence specificity of the enzyme:RNA complex.
[0079] "Hybridization" refers to hydrogen bonding between complementary nucleic acid bases, and this bonding are Watson-Crick hydrogen bonds, Hoogsteen hydrogen bonds, or reverse Hoogsteen hydrogen bonds. For example, adenine and thymine pair through the formation of hydrogen bonds. are complementary nucleobases.
[0080] "Increase" means at least a 10%, 25%, 50%, 75%, or 100% increase in This means a change in
[0081] The terms "inhibitor of base repair," "base repair inhibitor," "IBR," or their equivalents Legally equivalent terms include those capable of inhibiting the activity of nucleic acid repair enzymes, e.g., base excision repair enzymes. In some embodiments, IBR refers to a protein that inhibits inosine base excision repair. Exemplary inhibitors of base repair include APE1, EndoIII, EndoI, V, EndoV, EndoVIII, Fpg, hOGGl, hNEILl, T7Endo These include inhibitors of T4PDG, UDG, hSMUGl, and hAAG. In this embodiment, the base repair inhibitor is an inhibitor of EndoV or hAAG. In some embodiments, the IBR is an inhibitor of EndoV or hAAG. In this form, IBR binds catalytically inactive EndoV or catalytically inactive hAAG. In some embodiments, the base repair inhibitor is catalytically inactive EndoV or In some embodiments, the base repair inhibitor is uracil. UGI is a uracil-DNA glycosylase inhibitor. In some embodiments, U refers to a protein that can inhibit base excision repair enzymes. The GI domain comprises wild-type UGI or a fragment of wild-type UGI. The UGI proteins provided herein include fragments of UGI and UGI or UGI fragments. In some embodiments, the base repair inhibitor is an inosine salt In some embodiments, the base repair inhibitor is a "catalytic inhibitor of base excision repair." Inactive inosine-specific nuclease" or "Inactive inosine-specific nuclease" Without wishing to be bound by any particular theory, catalytically inactive inno- Adenine glycosylases (e.g., alkyladenine glycosylase (AAG)) are produced in wild boar. It can bind to inosine but cannot create abasic sites or remove inosine. damage, thereby protecting the newly formed inosine moiety from DNA damage / repair machinery. In some embodiments, the catalytically inactive inosine-specific nuclease is It can allow inosine binding within nucleic acids but does not cleave nucleic acids. A representative catalytically inactive inosine-specific nuclease is catalytically inactive arginine. Alkyl adenosine glycosylase (AAG nuclease), e.g., of human origin and catalytically inactive endonuclease V (EndoV nuclease), For example, from E. coli. In some embodiments, catalytically inactive AAG The nuclease contains the E125Q mutation or a corresponding mutation in another AAG nuclease .
[0082] An "intein" is a protein that binds itself to other proteins in a process known as protein splicing. Proteins that can excise the fragment and join the remaining fragment (extein) with peptide bonds. Inteins are also called "protein introns." The "process of an intein excising itself and splicing the remainder of the protein" is used herein to refer to is "protein splicing" or "intein-mediated protein splicing" In some embodiments, the intein of the precursor protein (intein-mediated transcription factor) is The intein-containing protein (pre-spliced protein) is derived from two genes Such inteins are referred to herein as split inteins (e.g., In cyanobacteria, for example, In this study, DnaE, the catalytic subunit of DNA polymerase III, is expressed in two separate regions. It is encoded by the genes dnaE-n and dnaE-c. The intein that is loaded is sometimes referred to herein as "intein-N." The intein encoded by the aE-c gene is referred to herein as "intein-C." This may be the case.
[0083] Other intein systems may also be used, for example, those based on the dnaE intein. The synthetic inteins Cfa-N (e.g., split intein-N) and Cfa-C Intein pairs (e.g., split intein-C) have been described (e.g., Stevens et al., J Am Chem Soc. 2016 Feb. 24; 138(7):21, incorporated herein by reference. 62-5). Non-limiting examples of intein pairs that can be used in accordance with the present disclosure include Cfa DnaE intein, Ssp GyrB intein, Ssp DnaX intein, T er DnaE3 intein, Ter ThyX intein, Rma DnaB intein and Cne Prp8 intein (U.S. Patent No. 5,629,393, which is incorporated herein by reference). Examples include those described in US Pat. No. 8,394,604.
[0084] Exemplary nucleotide and amino acid sequences of inteins are shown. DnaE intein-N DNA:TGCCTGTCATACGAACCGAGATACTGACAGTAGAATATGGCCTTCTG CCAATCGGGAAGATTGTGGAGAAACGGATAGAATGCACAGTTTACTCTGTCGATAACAATGGTAACATTTATACTCAGCC AGTTGCCCAGTGGCACGACCGGGGAGAGCAGGAAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAGGGCCACTA AGGACCACAAATTTATGACAGTCGATGGCCAGATGCTGCCTATAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGA GTTGACAACCTTCCTAAT DnaE intein-N protein: CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQP VAQWHDR GEQEVFEYCLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRVDNL PN DnaE intein-C DNA: ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTT ATGA TATTGGAGTCGAAAGAGATCACAACTTTGCTCTGAAGAACGGATTCATAG CTTCTAAT Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN Cfa-N DNA: TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCCTATTGGAAAG ATTGTCGAAGAGAGAATTGAATGCACAGTATATACTGTAGACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATG GCACAATCGCGGCGAACAAGAAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACGAGCAACTAAAGATCATAAAT TCATGACCACTGACGGGCAGATGTTGCCAATAGATGAGATATTTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTG CCA Cfa-N protein: CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQ MLPIDEIFERGLDLKQVDGLP Cfa-C DNA:ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAGGAAAGTAAAG ATAATATCTCGAAAAAGTCTTGGTACCCAAAATGTCTATGATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAA CGGTCTCGTAGCCAGCAAC Cfa-C protein:MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN
[0085] Intein-N and intein-C are the N-terminal and split portions of split Cas9. For the ligation of the C-terminal part of split Cas9, the N-terminal part of split Cas9 and the C-terminal part of split Cas9 are used. The split Cas9 fragments may be fused to the C-terminal end of the split Cas9 fragment, respectively. For example, in some embodiments, In this state, intein-N is fused to the C-terminus of the N-terminal part of split Cas9, i.e. That is, the structure N-[N-terminal part of split Cas9]-[intein-N]-C is formed. In some embodiments, intein-C is attached to the N-terminus of the C-terminal portion of the split Cas9. The C-terminal portion of the split Cas9 is fused to the N-[intein-C]-[C-terminal portion of the split Cas9]. ]-C structure. The protein to which the intein is fused (e.g., split Ca The mechanism of intein-mediated protein splicing to join s9) is well known in the art. These methods are well known in the art, see, for example, Shah et al., Chem Sci. 2002, incorporated herein by reference. 014; 5(1):446-461. Methods for designing and using inteins are well known in the art. These are known in the art, and are described in, for example, WO2014004336, WO2017132580, US Patent Application Publication No. 20150344549 and U.S. Patent Application Publication No. 20180127 780, both of which are incorporated herein by reference in their entireties.
[0086] The terms "isolated," "purified," or "biologically pure" mean that a material is in its pure state. It refers to being free, to varying degrees, from components normally associated with it when found in its native state. "Isolated" refers to the degree of separation from the original source or environment. refers to a degree of separation that is greater than isolation. The protein is free from any impurities that may materially affect the biological properties of the protein. In other words, the nucleic acid of the present invention is free from other substances to the extent that it does not cause adverse effects. Alternatively, the peptide may be derived from cell-derived material, virus-derived material, or produced by recombinant DNA techniques. The culture medium in which the compound is produced, or the chemical precursors or other chemicals in which it is chemically synthesized, Purity and homogeneity are typically determined by analytical chemistry techniques. For example, using polyacrylamide gel electrophoresis or high performance liquid chromatography The term "purified" refers to a nucleic acid or protein that is purified by electrophoresis. It can be said that the resulting band is essentially one band. For proteins that can be subjected to polymerization, different modifications result in different isolated proteins. These may be purified separately.
[0087] An "isolated polynucleotide" is a polynucleotide that is a naturally occurring genomic sequence of the organism from which the nucleic acid molecule of the present invention is derived. In the context of gene expression, it refers to the gene-free nucleic acid (e.g., DNA) that flanks the gene. Therefore, the term is used to refer to, for example, an autonomously replicating vector, a plasmid, or a virus. integrated into a virus or into the genomic DNA of a prokaryotic or eukaryotic organism, or Separate molecules (e.g., PCR or restriction endonuclease digestion) independent of other sequences cDNA or genomic DNA fragments or cDNA fragments resulting from In addition, the term includes RNA molecules transcribed from DNA molecules, as well as a recombinant DNA that is part of a hybrid gene that encodes an additional polypeptide sequence include.
[0088] An "isolated polypeptide" is a polypeptide of the present invention that has been separated from components that naturally accompany it. Typically, a polypeptide refers to a polypeptide that is a protein or polypeptide with which it is naturally associated. and is isolated when it is at least 60%, by weight, free from naturally occurring organic molecules. Preferably, the preparation is at least 75% by weight, more preferably at least 90% by weight, most preferably or at least 99% of the polypeptide of the present invention. A gene encoding such a polypeptide may be obtained, for example, by extraction from a natural source. It may be obtained by expression of a recombinant nucleic acid or by chemical synthesis of the protein. Purity can be determined by any suitable method, for example, column chromatography, polyacrylamide gel electrophoresis, or the like. The concentration can be measured by column electrophoresis or by HPLC analysis.
[0089] The term "linker," as used herein, refers to a molecule or moiety that connects two molecules or moieties, e.g., For example, two components of a protein complex or ribonucleocomplex, or a fusion protein Two domains, e.g., a polynucleotide programmable DNA binding domain (e.g., dCas9) and a deaminase domain (e.g., adenosine deaminase, cytogenes adenosine deaminase, or adenosine deaminase and cytidine deaminase) Covalent linkers (e.g., covalent bonds), non-covalent linkers, chemical groups, or A linker can refer to a molecule. A linker can refer to a molecule that connects different components or components of a base editor system. For example, in some embodiments, the linker can be a poly Programmable nucleotide-binding domains guide polynucleotide binding In some embodiments, the domain can be joined to the catalytic domain of a deaminase. The linker can join the CRISPR polypeptide and the deaminase. In some embodiments, a linker can join the Cas9 and the deaminase. In some embodiments, a linker can join the dCas9 and the deaminase. In some embodiments, the linker can join the nCas9 and the deaminase In some embodiments, a linker joins the guide polynucleotide and the deaminase. In some embodiments, the linker can be used to deaminase the base editor system. and joining the polynucleotide-programmable nucleotide binding component to the polynucleotide. In some embodiments, the linker can be a deaminated structure of the base editor system. RNA-binding portion of the component and polynucleotide-programmable nucleotide-binding component In some embodiments, the linker can be a RNA-binding moiety of deamination component and polynucleotide programmable nucleotide binding A linker can be a group, molecule, or a combination of two groups, molecules, or a linker that can connect the RNA-binding portion of a conjugated component to the RNA-binding portion of a conjugated component. or located between or flanked by other moieties, covalently or non-covalently bonded and thus the two can be connected. The linker can be an organic molecule, group, polymer, or chemical moiety. In embodiments, the linker may be a polynucleotide. The linker can be a DNA linker. In some embodiments, the linker is an RNA linker. In some embodiments, the linker is capable of binding to a ligand. In some embodiments, the ligand can be a carbohydrate, peptide, or other suitable aptamer. In some embodiments, the linker may be a ribonucleotide, a protein, or a nucleic acid. The aptamer may include an aptamer that can be derived from a riboswitch. The riboswitches are theophylline riboswitch, thiamine pyrophosphate (TPP) riboswitch, and adeno Cincobalamin (AdoCbl) riboswitch, S-adenosylmethionine (SAM) riboswitch riboswitch, SAH riboswitch, flavin mononucleotide (FMN) riboswitch, Tetrahydrofolate riboswitch, lysine riboswitch, glycine riboswitch, purine riboswitch riboswitch, GlmS riboswitch, or prequosin 1 (PreQ1) riboswitch In some embodiments, the linker may be selected from a polypeptide or protein domain. The antibody may comprise an aptamer bound to a domain, such as a polypeptide ligand. In some embodiments, the polypeptide ligand comprises a K homology (KH) domain, an MS2 Coat protein domain, PP7 coat protein domain, SfMu Com coat Protein domains, sterile alpha motif, telomerase Ku binding motif and Ku Protein, telomerase Sm7 binding motif and Sm7 protein, or RNA recognition In some embodiments, the polypeptide ligand may be a base editor motif. For example, the nucleobase editing component may be a deaminator. It may contain a nucleotide domain and an RNA recognition motif.
[0090] In some embodiments, the linker may be one or more amino acids (e.g., a peptide). In some embodiments, the linker may be about 5 to 1000 nucleotides long. 100 amino acids in length, e.g., about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 , 15, 16, 17, 18, 19, 20, 20-30, 30-40, 40-50, 50- Length of 60, 60-70, 70-80, 80-90, or 90-100 amino acids In some embodiments, the linker may be about 100-150, 150-200, 200~250, 250~300, 300~350, 350~400, 400~450, or 450-500 amino acids in length. Cars are also expected.
[0091] In some embodiments, the linker comprises an RNA promoter, including a Cas9 nuclease domain. The gRNA-binding domain of the gramabru nuclease and the catalytic domain of the nucleic acid editing protein (e.g., cytidine or adenosine deaminase). The linker connects the dCas9 and the nucleic acid editing protein. For example, the linker may be 2 Covalently bonded to or positioned between or flanking two groups, molecules, or other moieties In some embodiments, the linker is one or more amino acids (e.g., a peptide or protein) In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5 to 200 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 2 5, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90 ,90,95,100,101,102,103,104,105,110,120,1 30, 140, 150, 160, 175, 180, 190, or 200 amino acids in length is.
[0092] In some embodiments, the base editor domain is SGGSSGSETPGTSES ATPESSGGS, SGGSGGSSGSETPGTSESATPESSGGSSG GS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTS TEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPS Contains the amino acid sequence EGSAPGTSESATPESGPGSEPATSGGSGGS In some embodiments, the base editor domain is fused via a linker. The amino acid sequence SGSETPGTSESATPES is sometimes called the XTEN linker In some embodiments, the linker is fused via a linker comprising the 24 amino acid In some embodiments, the linker has a length of the amino acid sequence SGGSSGGSSGS In some embodiments, the linker comprises 40 amino acids in length. In some embodiments, the linker has the amino acid sequence SGGSSGGSSGSET In some embodiments, the compound comprises PGTSESATPESSGGSSGGSSGGSSGGS. In some embodiments, the linker is 64 amino acids in length. Column SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGG In some embodiments, the link In some embodiments, the linker is 92 amino acids in length. GSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPA GSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATP Includes ESGPGSEPATS.
[0093] A "marker" is any molecule whose altered expression level or activity is associated with a disease or disorder. It refers to a protein or polynucleotide of the
[0094] The term "mutation" as used herein refers to a mutation in a sequence, e.g., a nucleic acid sequence or Substitution of a residue in an amino acid sequence, substitution of a residue for another residue, or one or more residues in the sequence In some embodiments, an insertion refers to the deletion or insertion of a residue from the wild-type sequence. is a gene conversion that replaces a portion of the original residue, followed by a by identifying the position of that residue in the To make the amino acid substitutions (mutations) shown herein, Various methods for this purpose are well known in the art and are described, for example, in Green and Sambrook, Molecule r Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Co Spring Harbor, NY (2012).
[0095] In some embodiments, the base editors disclosed herein may reduce a significant number of unintended mutations. , nucleic acids (e.g., nucleic acids in a subject's genome) without generating unintended point mutations, etc. ) can be efficiently generated "target mutations," such as point mutations. In this form, the mutation of interest is a guide polypeptide specifically designed to generate the mutation of interest. A specific base editor (e.g., a citrine) conjugated to a nucleotide (e.g., a gRNA) These mutations are generated by adenosine base editors (DNA base editors or adenosine base editors).
[0096] Generally, mutations made or identified in a sequence (e.g., an amino acid sequence described herein) Variations are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain that mutation. Those skilled in the art will appreciate how to determine the location of amino acid and nucleic acid sequence variations relative to a reference sequence. It becomes easy to understand how decisions are made.
[0097] The term "non-conservative mutation" refers to a mutation between different groups, e.g., tryptophan to lysine, or a serine to phenylalanine amino acid substitution. The amino acid substitutions do not interfere with or inhibit the biological activity of the functional variant. Non-conservative amino acid substitutions are preferred because they improve the biological activity of the functional variant compared to the wild-type protein. The biological activity of the functional variant can be enhanced such that it is increased relative to the protein.
[0098] The term "nuclear localization sequence," "nuclear localization signal," or "NLS" refers to a sequence that directs a target protein into the cell nucleus. It refers to an amino acid sequence that promotes the import of a protein. Nuclear localization sequences are known in the art. See, for example, Plank et al., International PCT Application PCT / EP2 filed November 23, 2000. 000 / 011690, as described in WO / 2001 / 038547 published on May 31, 2001 No. 6,239,999, the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences. In another embodiment, the NLS is an optimized NLS, e.g. , Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. Optimized sequences useful in the methods of the invention are shown in Figures 8A-8E (Koblan et al., supra). In some embodiments, the NLS has the amino acid sequence KRTADGSEFESPK KKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL ,KRGINDRNFWRGENGRKTR,RKSGKIAAIVVKRPRK,PK KKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC include.
[0099] The terms "nucleobase," "nitrogen-containing base," or "base," used interchangeably herein, are Nitrogen-containing biological compounds that form the building blocks of nucleotides, nucleosides, and even nucleotides The ability of nucleic acid bases to form base pairs and stack one on top of another allows them to form ribonucleotides. Long chain helical structures such as nucleic acids (RNA) and deoxyribonucleic acid (DNA) are directly generated. There are five nucleobases: adenine (A), cytosine (C), guanine (G), and thymine (C). Adenine (T), and uracil (U) are called major or canonical. Nine is derived from purine, while cytosine, uracil, and thymine are derived from pyrimidine. NA and RNA can also contain other (non-major) bases that are modified. Non-limiting exemplary modified nucleobases include hypoxanthine, xanthine, 7-methyl guanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydroxybenzoates Hypoxanthine and xanthine are mutagens. through the presence of both, through deamination (replacement of an amine group with a carbonyl group), Hypoxanthine can be produced from adenine. Uracil can be modified from guanine. Uracil is the result of deamination of cytosine. A "nucleoside" is a nucleic acid consisting of a nucleic acid base and a five-carbon sugar (ribose or deoxyribose). Examples of nucleosides are adenosine, guanosine, and uridin. cytidine, 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine Modified nucleosides include thiazol-1, thymidine, deoxyuridine, and deoxycytidine. Examples of nucleosides containing an acid base include inosine (I), xanthosine (X), 7-methyl- Chirguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), and pseudouridine (Ψ). A "nucleotide" is a nucleic acid consisting of a nucleic acid base, a five-carbon sugar ( It consists of a nucleotide sequence consisting of either ribose or deoxyribonucleotides, and at least one phosphate group.
[0100] The terms "nucleic acid" and "nucleic acid molecule" as used herein refer to a nucleic acid molecule comprising a nucleic acid base and an acid Compounds containing moieties, e.g., nucleosides, nucleotides, or polymers of nucleotides Typically, a polymeric nucleic acid, e.g., a nucleic acid molecule containing three or more nucleotides, is a linear molecule in which adjacent nucleotides are joined by phosphodiester bonds In some embodiments, a "nucleic acid" refers to individual nucleic acid residues ( In some embodiments, "nucleic acid" refers to a nucleic acid. " refers to an oligonucleotide chain containing three or more individual nucleotide residues. As used herein, the terms "oligonucleotide" and "polynucleotide" refer to nucleic acids. To refer to a polymer of nucleotides (e.g., a stretch of at least three nucleotides) In some embodiments, "nucleic acid" refers to RNA as well as single stranded nucleic acids. Nucleic acids include, for example, genomes, transcripts, mRNAs, and the like. A, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, staining It may be naturally occurring in the context of a chromophore or other naturally occurring nucleic acid molecule. On the other hand, nucleic acid molecules can be non-naturally occurring molecules, such as recombinant DNA or RNA, artificial chromosomes, Engineered genomes, or fragments thereof, or synthetic DNA, RNA, or DNA / RNA hybrids or contain unnatural nucleotides or nucleosides. Additionally, the terms "nucleic acid," "DNA," "RNA," and / or similar The term includes nucleic acid analogs, e.g., analogs having other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems, and optionally Where appropriate, the molecule may be chemically synthesized, e.g. In some cases, nucleic acids may contain nucleoside analogs, e.g., nucleosides with chemically modified bases or sugars. Nucleic acid sequences may include analogs, etc., and backbone modifications. In some embodiments, the nucleic acid is present in a 3' to 3' orientation. Adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs ( For example, 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3 -methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromourizidine C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C 5-Propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deaza Adenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O (6)-methylguanine, and 2-thiocytidine; chemically modified bases; biologically modified modified bases (e.g., methylated bases); intercalated bases; modified sugars (e.g., 2'- Fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose and / or modified phosphate groups (e.g., phosphorothioates and 5'-N-phosphates) is or contains a phosphoamidite bond).
[0101] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" refers to A nucleic acid (e.g., DNA or RNA) that guides the apDNAbp to a specific nucleic acid sequence, e.g. For example, a protein associated with a guide nucleic acid or guide polynucleotide (e.g., gRNA). The term "polynucleotide programmable nucleotide binding domain" is used interchangeably to refer to a protein. In some embodiments, a polynucleotide programmable nucleic acid The nucleotide binding domain is a polynucleotide-programmable DNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain comprises: In some embodiments, the polynucleotide is a programmable RNA binding domain. The nucleotide-programmable nucleotide-binding domain is a Cas9 protein. The Cas9 protein binds to a specific DNA sequence complementary to the guide RNA. In some embodiments, napDNAbp may be associated with a guide RNA that guides the quality of the DNA. Cas9 domain, e.g., Cas9 with nuclease activity, Cas9 nickase (n Cas9), or nuclease-inactive Cas9 (dCas9). Non-limiting examples of matrix DNA binding proteins include Cas9 (e.g., dCas9 and and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas1 Non-limiting examples of Cas enzymes include Cas12i, Cas2i, and Cas3i. 1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, C as8c, Cas9 (also known as Csn1 or Csx12), Cas10, C as10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c 3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse 3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Cs m1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cm r4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, C sf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Cs h1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas e effector protein, type V Cas effector protein, type VI Cas effector Proteins, CARF, DinG, their homologues, or modified or engineered versions thereof Other nucleic acid programmable DNA binding proteins include engineered versions of the nucleic acid. are also within the scope of this disclosure even though they may not be specifically recited in this disclosure. See, for example, Makarova et al. "Class sification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct;1:325-336. doi: 10.1089 / crispr.2018.0033; Yan et al., “Functionally di verse type V CRISPR-Cas systems” Science. 2019 Jan 4;363(6422):88-91. doi: 10.1 Please refer to 126 / science.aav7271.
[0102] The term "nucleobase-editing domain" or "nucleobase-editing protein" as used herein When used, modifications of nucleic acid bases in RNA or DNA, such as cytosine (or thymine (or thymidine) to uracil (or uridine) or thymine (or thymidine), and and deamination of adenine (or adenosine) to hypoxanthine (or inosine) Proteins that can catalyze the synthesis of non-templated nucleotides, as well as the addition and insertion of non-templated nucleotides. In some embodiments, the nucleobase editing domain refers to a protein or enzyme. enzymes (e.g., adenine deaminase or adenosine deaminase; or cytidine In some embodiments, the nucleobase amino acid is a nucleobase deaminase or a cytosine deaminase. The clustering domain may comprise one or more deaminase domains (e.g., adenine deaminase or adenine deaminase). adenosine deaminase, and cytidine or cytosine deaminase). In embodiments, the nucleobase-editing domain may be a naturally occurring nucleobase-editing domain. In some embodiments, the nucleobase-editing domain is engineered from a naturally occurring nucleobase-editing domain. Nucleobase editing domains that have been modified or evolved. can be used in any organism, such as bacteria, humans, chimpanzees, gorillas, monkeys, cows, dogs, rats, etc. , or may be of murine origin.
[0103] As used herein, "obtaining" as in "obtaining an agent" means combining an agent. To produce, isolate, extract, purchase, or otherwise obtain include.
[0104] As used herein, a "patient" or "subject" refers to a person suffering from a disease or disorder. mammalian subjects diagnosed with, at risk for, or suspected of developing In some embodiments, the term refers to a subject or individual having a mutation in the gene encoding the SDSP. Elephants have or are at risk of developing Shwachman-Diamond Syndrome (SDS). In some embodiments, the term "patient" refers to a person who is at risk for a disease or disorder. An exemplary patient is a mammalian subject who has a higher than average likelihood of developing a disease described herein. Humans, non-human primates, cats, dogs, pigs, cattle, etc. that can benefit from the indicated therapy. Cats, horses, camels, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, gerbils, or guinea pigs), and other mammals. An exemplary human patient is a male. It can be gender and / or female.
[0105] A "patient in need thereof" or a "subject in need thereof" as used herein refers to a person who is suffering from a disease or or have been diagnosed with or are at risk of having a disorder, such as SDS, Patients are referred to as those who are either predetermined to be affected or suspected to be affected.
[0106] The terms "pathogenic mutation," "pathogenic variant," "disease-causing mutation," and "disease-causing mutation" are used interchangeably. A "causing variant," "deleterious mutation," or "predisposing mutation" is a specific Genetic alterations or mutations that increase an individual's susceptibility or predisposition to a disease or disorder In some embodiments, the pathogenic mutation refers to a polynucleotide that encodes an SBDS protein. This includes alterations within the splice acceptor or splice donor sites in the polypeptide. In some embodiments, the pathogenic variant is a variant of a polynucleotide encoding an SBDS protein. altering the splicing of the SBD, resulting in, for example, protein truncation or other This results in a negative effect on the expression or activity of the S protein.
[0107] The terms "protein," "peptide," "polypeptide," and their grammatical equivalents The terms "peptide" and "peptide-linked" are used interchangeably herein to refer to compounds linked together by peptide (amide) bonds. Refers to a polymer of amino acid residues. The term applies to proteins of any size, structure, or function. It typically refers to a protein, peptide, or polypeptide. A polypeptide is at least three amino acids long. A protein, peptide, or Polypeptide can refer to an individual protein or a collection of proteins One or more amino acids in a protein, peptide, or polypeptide may be conjugated. Chemical entities for conjugation, functionalization, or other modification, e.g., carbohydrate groups, hydrophobic groups, xyl group, phosphate group, farnesyl group, isofarnesyl group, fatty acid group, linker, etc. Thus, the protein, peptide, or polypeptide can be modified by a single They can be molecules or multi-molecular complexes. Proteins, peptides Alternatively, the polypeptide may be only a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide may be natural, recombinant, or synthetic, or The term "fusion protein" as used herein can be any combination thereof. A "protein" is a hybrid that contains protein domains from at least two different proteins. A hybrid polypeptide is a hybrid polypeptide in which one protein is attached to the amino terminus (N-terminus) of the fusion protein. ) portion or carboxy-terminal (C-terminal) protein, and therefore, the amino-terminal fusion A fusion protein or a carboxy-terminal fusion protein, respectively, can be formed. The proteins may contain distinct domains, such as the nucleic acid binding domain of a nucleic acid editing protein (e.g. , the gRNA-binding domain of Cas9, which directs the binding of the protein to the target site) and the nucleus. In some embodiments, the protein may comprise an acid cleavage domain or a catalytic domain. Proteins consist of a proteinaceous portion, e.g., an amino acid sequence constituting a nucleic acid binding domain, and an organic In some embodiments, the compound may act as a nucleic acid cleaving agent. In some cases, the protein is complexed with or binds to a nucleic acid, e.g., RNA or DNA. Any of the proteins described herein can be associated with any method known in the art. For example, the proteins described herein can be produced by peptide linkers. Produced via recombinant protein expression and purification, it is particularly suitable for fusion proteins containing anchors. Methods for the expression and purification of recombinant proteins are well known. , Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spr. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)) , the entire contents of which are incorporated herein by reference.
[0108] The polypeptides and proteins disclosed herein (including functional moieties and their functional barriers) The amino acids (including the amino acids) may contain synthetic amino acids in place of one or more naturally occurring amino acids. Such synthetic amino acids are known in the art and include, for example, aminocyclohexane. Xanthate, norleucine, α-amino n-decanoic acid, homoserine, S-acetyl Aminomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine 4-carboxyphenylalanine, β-phenylserine, β-hydroxyphenylalanine phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylalanine Cylglycine, Indoline-2-carboxylic acid, 1,2,3,4-tetrahydroisoquinolinol N'-Benzyl-N'-malonic acid, N'-benzyl-N'-malonic acid monoamide '-Methyl-lysine, N',N'-dibenzyl-lysine, 6-hydroxylysine, orni tin, α-aminocyclopentanecarboxylic acid, α-aminocyclohexanecarboxylic acid, α -Aminocycloheptanecarboxylic acid, α-(2-amino-2-norbornane)-carboxylic acid acid, α,γ-diaminobutyric acid, α,β-diaminopropionic acid, homophenylalanine, and α-tert-butylglycine. Polypeptides and proteins include poly The polypeptide construct may be accompanied by post-translational modifications of one or more amino acids. Non-limiting examples include phosphorylation, acylation, including acetylation and formylation, glycosylation, and the like. Sylation (including N-linked and O-linked), amidation, hydroxylation, methylation and alkylation, including ethylation, ubiquitination, pyrrolidone carboxylic acid addition, disulfide Formation of amide bridges, sulfation, myristoylation, palmitoylation, isoprenylation, farnesylation These include sylation, geranylation, glypiation, lipoylation, and iodination.
[0109] The term "recombinant" as used herein in the context of a protein or nucleic acid , refers to proteins or nucleic acids that do not occur in nature but are the product of human engineering. For example, some In this embodiment, the recombinant protein or nucleic acid molecule contains less than or equal to any naturally occurring sequence. At least one, at least two, at least three, at least four, at least five, at least an amino acid sequence or a nucleotide sequence containing at least six or at least seven mutations; include.
[0110] "Reduce" means at least a 10%, 25%, 50%, 75%, or 100% reduction This means a change in
[0111] "Reference" refers to a standard or control condition. In one embodiment, the reference is a wild-type For example, a wild-type or healthy cell is a cell that is healthy and and / or may be derived from or obtained from disease-free subjects. In a specific embodiment, wild-type or healthy cells contain wild-type SBDS protein (i.e., SBDS protein, which is the product of the wild-type SBDS gene exhibiting wild-type splicing. In another embodiment, and without limitation, the reference is to a cell expressing a gene (protein) that is in a test condition. or placebo or normal saline, medium, buffer, and / or are untreated cells that are subjected to a control vector that does not contain the polynucleotide of interest.
[0112] A "reference sequence" is a defined sequence used as a basis for sequence comparison. The reference sequence may be a subset of the specified sequence or the entire sequence, e.g., the full-length cDNA. A segment of a DNA or gene sequence, or the complete cDNA or gene sequence For polypeptides, the length of a reference polypeptide sequence is generally at least about 16 amino acids, at least about 20 amino acids, at least about 25 amino acids, about 35 amino acids, about For nucleic acids, the length of a reference nucleic acid sequence is: Generally, at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides, or about 300 nucleotides, or In some embodiments, the reference sequence is a sequence of interest. In other embodiments, the reference sequence is a sequence encoding a wild-type protein. It is a polynucleotide sequence that
[0113] The terms "RNA-programmable nuclease" and "RNA-guided nuclease" is used with one or more RNAs that are not targets for cleavage (e.g., In some embodiments, the RNA programmable nucleic acid When in a complex with RNA, the nuclease is sometimes called a nuclease:RNA complex. Typically, the RNA that is bound is called a guide RNA (gRNA). In embodiments, the RNA programmable nuclease (CRISPR-associated system) s9 endonuclease, e.g., Cas from Streptococcus pyrogenes 9 (Csnl) (see, for example, "Complete genome sequence of an Ml strain of Strept ococcus pyogenes." Ferretti JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C, Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qi an Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L., White J., Yuan X., Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(200 1);"CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III. " Deltcheva E., Chylinski K., Sharma CM., Gonzales K., Chao Y., Pirzada ZA, Ec See Kert MR, Vogel J., Charpentier E., Nature 471:602-607(2011).
[0114] What is the "Shwachman-Bodian-Diamond syndrome (SBDS) protein"? NCBI accession number NP_057122.2 and have at least about 85% amino acid sequence identity. It means a polypeptide or fragment thereof having the biological activity of SBDS. In embodiments, the biological activity of SBDS is an activity that inhibits RNA processing, ribosome biogenesis, or It refers to the role it plays in binding to antibodies that specifically bind to the SBDS protein.
[0115] An example of the amino acid sequence of an SBDS protein is shown below: MSIFTPTNQI RLTNVAVVRM KRAGKRFEIA CYKNKVVGWR SGVEKDLDEV LQTHSVFVNV SKGQVAKKED LIS AFGTDDQ TEICKQILTK GEVQVSDKER HTQLEQMFRD IATIVADKCV NPETKRPYTV ILIERAMKDI HYSVKT NKST KQQALEVIKQ LKEKMKIERA HMRLRFILPV NEGKKLKEKL KPLIKVIESE DYGQQLEIVC LIDPGCFRE I DELIKKETKG KGSLEVLNLK DVEEGDEKFE.
[0116] In a specific embodiment, the SBDS protein comprises a protein truncation.
[0117] "Shwachman-Bodian-Diamond Syndrome (SBDS) Polynucleotide" means a nucleic acid sequence encoding an SBDS protein. An example sequence is shown in NM_016038.2, reproduced below. SBDS The open reading frame (ORF) of the polynucleotide begins at nucleotide 185. It extends to 937 (shown in underlined part). GTAAGTAAGC CTGCCAGACA CACTGTGACG GCTGCCTGAA GCTAGTGAGT CGCGGCGCCG CGCACTGGTG GTT GGGTCAG TGCCGCGCGC CGATCGGTCG TTACCGCGAG GCGCTGGTGG CCTTCAGGCT GGACGGCGCG GGTCAG CCCT GGTTCGCCGG CTTCTGGGTC TTTGAACAGC CGCG ATGTCG ATCTTCACCC CCACCAACCA GATCCGCCT A ACCAATGTGG CCGTGGTACG GATGAAGCGT GCCGGGAAGC GCTTCGAAAT CGCCTGCTAC AAAAACAAGG T CGTCGGCTG GCGGAGCGGC GTGGAAAAAG ACCTCGATGA AGTTCTGCAG ACCCACTCAG TGTTTGTAAA TGTT TCTAAA GGTCAGGTTG CCAAAAAGGA AGATCTCATC AGTGCGTTTG GAACAGATGA CCAAACTGAA ATCTGTA AGC AGATTTTGAC TAAAGGAGAA GTTCAAGTAT CAGATAAAGA AAGACACACA CAACTGGAGC AGATGTTTAG GGACATTGCA ACTATTGTGG CAGACAAATG TGTGAATCCT GAAACAAAGA GACCATACAC CGTGATCCTT AT TGAGAGAG CCATGAAGGA CATCCACTAT TCGGTGAAAA CCAACAAGAG TACAAAACAG CAGGCTTTGG AAGTG ATAAA GCAGTTAAAA GAGAAAATGA AGATAGAACG TGCTCACATG AGGCTTCGGT TCATCCTTCC AGTCAATG AA GGCAAGAAGC TGAAAGAAAA GCTCAAGCCA CTGATCAAGG TCATAGAAAG TGAAGATTAT GGCCAACAGT TAGAAATCGT ATGTCTGATT GACCCGGGCT GCTTCCGAGA AATTGATGAG CTAATAAAAA AGGAAACTAA AGG CAAAGGT TCTTTGGAAG TACTCAATCT GAAAGATGTA GAAGAAGGAG ATGAGAAATT TGAATGA CAC CCATCA ATCT CTTCACCTCT AAAACACTAA AGTGTTTCCG TTTCCGACGG CACTGTTTCA TGTCTGTGGT CTGCCAAAT A CTTGCTTAAA CTATTTGACA TTTTCTATCT TTGTGTTAAC AGTGGACACA GCAAGGCTTT CCTACATAAG T FATHER TGGGAATGAT TTGGTTTTAA TTATAAACTG GGGTCTAAAT CCTAAAGCAA AATTGAAACT CCAA GATGCA AAGTCCAGAG TGGCATTTTG CTACTCTGTC TCATGCCTTG ATAGCTTTCC AAAATGAAAG TTACTTG AGG CAGCTCTTGT GGGTGAAAAG TTATTTGTAC TAKE GATTTATTAGG GGTATGTCTA TACAACAAAA GGGGGGGTCT TTCCTAAAAAA AGAAAACATA TGATGCTTCA TTTCTACTTA ATGGAACTTG TGTTCTGAGG GT CATTATGG TATCGTAATG TAAAGCTTGG ATGATGTTCC TGATTATCTG AGAACAGAT ATAGAAAAT TGTGC CGGAC TTACCTTTCA TTGAACATGC TGCCATAACT TAGATTATTC TTGGTTAAAA AATAAAAGTC ACTTATTT CT AATTCTTAAA GTTTATAATA TATATTAATA TAGCTAAAAT TGTATGTAAT CAATAAAACC ACTCTTATGT TTATT
[0118] In some embodiments, Shwachman-Bodian-Diamond syndrome (SBDS) The polynucleotides include polynucleotides derived from SBDS pseudogenes. In this form, the SBDS polynucleotide is a mutation resulting from gene conversion associated with SDS ( For example, 258+2T>C and / or 183–184TA>CT mutations) alone or or in combination with other alterations in the SBDS pseudogene.
[0119] "Shwachman-Bodian-Diamond syndrome (SBDS) pseudogene" refers to the S means a nucleic acid sequence having at least about 85% nucleic acid sequence identity to a BDS polynucleotide. In one embodiment, exemplary pseudogenes include the following fragments thereof: >NR_024109.1 Human (Homo sapiens) SBDS pseudogene 1 (SBDSP1) , transcript variant 4, non-coding RNA CCTTTTTGGGCGTGGAAAGATGGCGGTAAAAGCCACAATGCGCAGGCGTCATCGCTCACTTCTCCCCTCC CGGCTTCTG CTCCACCTGACGCCTGCGCAGTAAGTAAGCCTGCCAGACACGCTGTGGCGGCTGCCTGAAG CTAGTGAGTCGCGGCGCC GCGCACTTGTGGTTGGGTCAGTGCCGCGCGCCGCTCGGTCGTTACCGCGAGG CGCTGGTGGCCTTCAGGCTGGACGGCG CGGGTCAGCCCTGGTTTGCCGGCTTCTGGGTCTTTGAACAGCC GCGATGTCGATCTTCACCCCCACCAACCAGATCCGC CTAACCAATGTGGCCGTGGTACGGATGAAGCGCG CCAGGAAGCGCTTCGAAATCGCCTGCTACAGAAACAAGGTCGTCG GCTGGCGGAGCGGCTTATTTTGACT AAAGGAGAAGTTCAAGTATCAGATAAAGACACACACAACTGGAGCAGATGTTTA GGGACATTGCAATTAT TGTGGCAGACAAATGTGTGACTCCTGAAACAAAGAGACCATACACCGTGATCCTTATTGAGAG AGCCATG AAGGACATCCACTATTTGGTGAAAACCAACAGGAGTACAAAACAGCAGGCTTTGGAAGTGATAAAGCAGT T AAAAGAGAAAATGAAGATAGAACGTGCTCACATGAGGCTTCAGTTCATCCTTCCAGTGAATGAAGGCAA GAAGCTGAAA GAAAAGCTCAAGCCACTGATCAAGGTCATAGAAAGTAAAGATTATGGCCAACAGTTAGAA ATCGTAAGAGTCAAATATT TTCTTTGCTTCATGTTACCTAAATATTGTATTCTCTAGTAATAAATTTGTA GCAAACATTCAAAAAAAAAAAAAAAAAA AA >NR_024110.1 Homo sapiens SBDS pseudogene 1 (SBDSP1), transcript variant 1, non-coding RNA Transcript variant 1, non-coding RNA CCTTTTTGGGCGTGGAAAGATGGCGGTAAAAGCCACAATGCGCAGGCGTCATCGCTCACTTCTCCCCTCC CGGCTTCTG CTCCACCTGACGCCTGCGCAGTAAGTAAGCCTGCCAGACACGCTGTGGCGGCTGCCTGAAG CTAGTGAGTCGCGGCGCC GCGCACTTGTGGTTGGGTCAGTGCCGCGCGCCGCTCGGTCGTTACCGCGAGG CGCTGGTGGCCTTCAGGCTGGACGGCG CGGGTCAGCCCTGGTTTGCCGGCTTCTGGGTCTTTGAACAGCC GCGATGTCGATCTTCACCCCCACCAACCAGATCCGC CTAACCAATGTGGCCGTGGTACGGATGAAGCGCG CCAGGAAGCGCTTCGAAATCGCCTGCTACAGAAACAAGGTCGTCG GCTGGCGGAGCGGCTTGGAAAAAGA CCTTGATGAAGTTCTGCAGACCCACTCAGTGTTTGTAAATGTTTCCTAAGGTCA<( GGTTGCCAAGAAGGAA GATCTCATCAGTGCGTTTGGAACAGATGACCAAACTGAAATCTATTTTGACTAAAGGAGAAGT TCAAGTA TCAGATAAAGACACACACAACTGGAGCAGATGTTTAGGGACATTGCAATTATTGTGGCAGACAAATGTGT G It should be noted that there may be an error in ID=19 where the opening parenthesis in the tag is not in the correct format according to the original. It is presented as is based on the translation requirements.ACTCCTGAAACAAAGAGACCATACACCGTGATCCTTATTGAGAGAGCCATGAAGGACATCCACTATTTG GTGAAAACCA ACAGGAGTAACAACAGCAGGCTTTGGAAGTGATAAAGCAGTTAAAAGAGAAAATGAAGA TAGAACGTGCTCACATGAG GCTTCAGTTCATCCTTCCAGTGAAGGCAAGAAGCTGAAAGAAAGCT CAAGCCACTGATCAAGGTCATAGAAAGT AAAGATTATGGCCAACAGTTAGAAATCGTATGTCTGATTGAC CTGGGCTGCTTCCGAGAAATTGATGAGCTAATAAA AGGAAACCAAAGGCAAAGGTCTTTTGGAAGTAC TCAATCTGAAAGATTTGAAGAAGGAGATGAGAATATTTGAATGACAC CCATCAGTCTCTTCACCTCTAAAA CACTAAAGTGTTTTCGTTTCCAACAGCACTGTTTCATGTCTGTGGTCTGCCAAAT ACTTGCTCAAACTAT TTGACATTTTCTATCTTTGTGTTAACAGTGGACACAGCAAGGCTTCCTACATAAGTATAATAA TGTGGG AATGATTTGGTTTTAATTATAAACTGGGGTCTAAATCCTAAAGCAAAATTGAAACTCCAGGATGCAAAAT CC AGAGTGGCATTTTGCTACTCTGTCTCATGCCTTGATAGCTTTCCAAAATGAAAGTTACTTGAGGCAGC TCTTGTGGGTG AAAAGTTTTTTGTACAGTAGAGTAAGATTATTAGGGGTATGTCTATACGACAAAAGGG GGTCTTTCCTAAAAAAAGAAA ACATGATGCTTCATTTCTACTTAATGGAACTTGTGTTCTGAGGGTCATTA TGGTATCGTAATATAAAGCTTGGATGATG TTCCTGATTATCTGAGAAACAGATATAGAAAAATTGTGTCG GACTTAAATAATTTTCGTTGAACATGCTGCCATAACTT AGATTATTCTTGGTTAAAAAATAAAAGTCACT TATTTCTAATTCTTAAAGTTTATAATATATATTAATATAGCTAAAAT TGTATGTAATCAATAAAACCACT CTTATGTTTATTAAACTATGGCTTGTGTTTCTAGACAAAAAAAAAAAAAAAAAA >NR_024111.1 Homo sapiens SBDS pseudogene 1 (SBDSP1), transcript variant 2, non-coding RNA Transcript variant 2, non-coding RNA CCTTTTTGGGCGTGGAAAGATGGCGGTAAAAGCCACAATGCGCAGGCGTCATCGCTCACTTCTCCCCTCC CGGCTTCTG CTCCACCTGACGCCTGCGCAGTAAGTAAGCCTGCCAGACACGCTGTGGCGGCTGCCTGAAG CTAGTGAGTCGCGGCGCC GCGCACTTGTGGTTGGGTCAGTGCCGCGCGCCGCTCGGTCGTTACCGCGAGG CGCTGGTGGCCTTCAGGCTGGACGGCG CGGGTCAGCCCTGGTTTGCCGGCTTCTGGGTCTTTGAACAGCC GCGATGTCGATCTTCACCCCCACCAACCAGATCCGC CTAACCAATGTGGCCGTGGTACGGATGAAGCGCG CCAGGAAGCGCTTCGAAATCGCCTGCTACAGAAACAAGGTCGTCG GCTGGCGGAGCGGCTTATTTTGACT AAAGGAGAAGTTCAAGTATCAGATAAAGACACACACAACTGGAGCAGATGTTTA GGGACATTGCAATTAT TGTGGCAGACAAATGTGTGACTCCTGAAACAAAGAGACCATACACCGTGATCCTTTATTGAGAG AGCCATG AAGGACATCCACTATTGGTGAAAACCAACAGGAGTACAAAACAGCAGGCTTTGGAAGTGATAAAGCAGT T AAAAGAGAAAATGAAGATAGAACGTGCTCACATGAGGCTTCAGTTCATCCTTCCAGTGAATGAAGGCAA GAAGCTGAAA GAAAAGCTCAAGCCACTGATCAAGGTCATAGAAAGTAAAGATTATGGCCAACAGTTAGAA ATCGTATGTCTGATTGACC TGGGCTGCTTCCGAGAAATTGATGAGCTAATAAAAAGGAAACCAAAGGCA AAGGTTCTTTGGAAGTACTCAATCTGAA AGATTTGAAGAAGGAGATGAGAAATTGAATGACACCCATCA GTCTCTTCCACCTCTAAAACACTAAAGTGTTTCGTTT CCAACAGCACTGTTTCATGTCTGTGGTCTGCCA AATACTTGCTCAAACTATTTGACATTTTCTATCTTTGTGTTAACAG TGGACACAGCAAGGCTTTCCTACA TAAGTATAATAATGTGGGAATGATTTGGTTTTAATTATAAACTGGGGTCTAAATC CTAAAGCAAAATTGA AACTCCAGGATGCAAAATCCAGAGTGGCATTTTGCTACTCTGTCTCATGCCTTGATAGCTTTCC AAAATG AAAGTTACTTGAGGCAGCTCTTGTGGGTGAAAGTTTTTTGTACAGTAGAGTAAGATTATTAGGGGTATG TC TATACGACAAAAGGGGGGTCTTTCCTAAAAAAGAAAACATGATGCTTCATTTCTACTTAATGGAACTT GTGTTCTGAGG GTCATTATGGTATCGTAATATAAAGCTTGGATGATGTTCCTGATTATCTGAGAAACAGA TATAGAAAAATTGTGTCGGA CTTAAATAATTTTCGTTGAACATGCTGCCATAACTTAGATTATTCTTGGT TAAAAAATAAAAGTCACTTATTTCTAATT CTTAAAGTTTATAATATATATTAATATAGCTAAAATTGTAT GTAATCAATAAAACCACTCTTATGTTTATTAAACTATG GCTTGTGTTTCTAGACAAAAAAAAAAAAAAAA AA >NR_001588.2 Homo sapiens SBDS pseudogene 1 (SBDSP1), transcript variant 3, non-coding RNA Transcript variant 3, non-coding RNA CCTTTTTGGGCGTGGAAAGATGGCGGTAAAAGCCACAATGCGCAGGCGTCATCGCTCACTTCTCCCCTCC CGGCTTCTG CTCCACCTGACGCCTGCGCAGTAAGTAAGCCTGCCAGACACGCTGTGGCGGCTGCCTGAAG CTAGTGAGTCGCGGCGCC GCGCACTTGTGGTTGGGTCAGTGCCGCGCGCCGCTCGGTCGTTACCGCGAGG CGCTGGTGGCCTTCAGGCTGGACGGCG CGGGTCAGCCCTGGTTTGCCGGCTTCTGGGTCTTTGAACAGCC GCGATGTCGATCTTCACCCCCACCAACCAGATCCGC CTAACCAATGTGGCCGTGGTACGGATGAAGCGCG CCAGGAAGCGCTTCGAAATCGCCTGCTACAGAAACAAGGTCGTCG GCTGGCGGAGCGGCTTGGAAAAAGA CCTTGATGAAGTTCTGCAGACCCACTCAGTGTTTGTAAATGTTTCCTAAGGTCA GGTTGCCAAGAAGGAA GATCTCATCAGTGCGTTTGGAACAGATGACCAAACTGAAATCTATTTTGACTAAAGGAGAAGT TCAAGTA TCAGATAAAGACACACACAACTGGAGCAGATGTTTAGGGACATTGCAATTATTGTGGCAGACAAATGTGT G ACTCCTGAAACAAAGAGACCATACACCGTGATCCTTATTGAGAGAGCCATGAAGGACATCCACTATTTG GTGAAACCA ACAGGAGTACAAAACAGCAGGCTTTGGAAGTGATAAAGCAGTTAAAAAGAGAAAATGAAGA TAGAACGTGCTCACATGAG GCTTCAGTTCATCCTTCCAGTGAATGAAGGCAAGAAGCTGAAAGAAAAGCT CAAGCCACTGATCAAGGTCATAGAAAGT AAAGATTATGGCCAACAGTTAGAAATCGTAAGAGTCAAATAT TTTCTTTGCTTCATGTTACCTAAATATTGTATTCTCT AGTAATAAATTTGTAGCAAACATTCAAAAAAAA AAAAAAAAAAAA
[0120] The term "single nucleotide polymorphism (SNP)" refers to a single nucleotide that occurs at a specific location in the genome. Each variation is somewhat noticeable within the population. For example, at specific base positions in the human genome, C nuclei Ochid may occur in most individuals, but in a few individuals this position is occupied by A. This means that the SNP is at this specific location, and the two possible nucleotides The variants C or A of the octide are said to be alleles at this position. P underlies differences in susceptibility to disease and also influences the severity of illness and how well our bodies respond to treatment. The way we respond to stress is also a manifestation of genetic variation. SNPs are found in the coding region of genes. can range within a gene, within the non-coding region of a gene, or in an intergenic region (region between genes) In some embodiments, SNPs within coding sequences are generated due to the degeneracy of the genetic code. SNPs in the coding region do not necessarily change the amino acid sequence of the protein being produced. There are two types of SNPs: synonymous and non-synonymous. Synonymous SNPs affect the protein sequence. Nonsynonymous SNPs alter the amino acid sequence of a protein, whereas nonsynonymous SNPs do not affect the There are two types of P: missense and nonsense. SNPs that are not present in the genome also affect gene splicing, transcription factor binding, and messenger receptor (MR) signaling. This type of SNP can affect the degradation of RNA or the sequence of non-coding RNA. Gene expression that is affected by the gene is called an eSNP (expressed SNP), and is located upstream or downstream from the gene. Single nucleotide variants (SNVs) are single nucleotide variants with no frequency limitation. Somatic single nucleotide variations are also , may be a single nucleotide change.
[0121] "Specifically bind" means to recognize and bind to the polypeptide and / or nucleic acid molecule of the present invention. nucleic acid molecules that bind to the nucleic acid sequence but do not substantially recognize or bind to other molecules in a sample, e.g., a biological sample , polypeptides, or complexes thereof (e.g., nucleic acid programmable DNA binding proteins "protein and guide nucleic acid"), compound, or molecule.
[0122] Nucleic acid molecules useful in the methods of the present invention include polypeptides of the present invention or fragments thereof Such nucleic acid molecules may be any nucleic acid molecule that encodes an endogenous nucleic acid sequence. It does not have to be 100% identical, but typically it will exhibit substantial identity. Polynucleotides that have "substantial identity" to a given sequence are typically double-stranded nucleic acid molecules. Nucleic acids useful in the methods of the present invention are capable of hybridizing to at least one strand of the nucleic acid. The nucleic acid molecule may be any nucleic acid molecule that encodes a polypeptide of the invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but Typically, the endogenous sequence will have substantial identity. A polynucleotide having a double-stranded nucleic acid molecule typically has a hybrid in at least one strand of the double-stranded nucleic acid molecule. "Hybridize" can be performed under various stringencies. a polynucleotide sequence complementary to the gene described herein under the conditions This refers to the pairing between portions of a double-stranded molecule (e.g., Wahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, AR (1987) Methods Enzymol. Mol. 152:507).
[0123] For example, stringent salt concentrations are typically about 750 mM NaCl and 75 mM Not more than trisodium citrate, preferably about 500 mM NaCl and 50 mM citrate. Not more than trisodium enoate, preferably about 250 mM NaCl and 25 mM Not to exceed trisodium citrate. Low stringency hybridization The synthesis can be achieved in the absence of organic solvents, such as formamide, whereas the synthesis High-stringency hybridization requires at least about 35% formamide, More preferably, it can be obtained in the presence of at least about 50% formamide. The appropriate temperature conditions are typically at least about 30°C, more preferably at least about 3 7°C, and most preferably at least about 42°C. meters, e.g., hybridization time, detergents, e.g., sodium dodecyl sulfate, The concentration of sodium dodecyl sulfate (SDS), and the inclusion or exclusion of carrier DNA, etc. are well known in the art. Varying levels of stringency can be used to combine these different conditions as needed. In a preferred embodiment, hybridization is achieved by combining in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS at 30°C In a more preferred embodiment, hybridization occurs in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and This occurs in 100 μg / mL denatured salmon sperm DNA (ssDNA) at 37°C. In a more preferred embodiment, hybridization is carried out in 250 mM NaCl, 25 mM Cl. Trisodium enoate, 1% SDS, 50% formamide, and 200 μg / mL ss This occurs in DNA at 42°C. Useful variations on these conditions include: It will be readily apparent to one skilled in the art.
[0124] In most applications, the washing steps that follow hybridization are also stringent. The stringency of washing varies depending on the salt concentration and As mentioned above, the stringency of the wash can be determined by the salt concentration and temperature. It can be increased by decreasing or by increasing the temperature. For example, the stringent salt concentration for the wash step is preferably about 30 mM NaCl and no more than 3 mM trisodium citrate, most preferably about 15 mM aCl and 1.5 mM trisodium citrate. Generally, the most stringent temperature conditions are at least about 25°C, and more preferably at least about 25°C. and preferably at least about 42°C, and even more preferably at least about 68°C. In one embodiment, the wash step is performed in 30 mM NaCl, 3 mM trisodium citrate. and 0.1% SDS at 25° C. In another embodiment, the washing step The buffer was diluted in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the wash step occurs at 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS at 68°C Additional variations on these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Bent et al. on and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. S ci., USA 72:3961, 1975);Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001);Berger and Kimmel (Guide to Molecular Clone ing Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecula r Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York. It is written.
[0125] "Split" means to divide into two or more pieces.
[0126] A "split Cas9 protein" or "split Cas9" is a protein that is split into two separate Ca provided as N- and C-terminal fragments encoded by nucleotide sequences The Cas9 protein is a polypeptide that corresponds to the N-terminal and C-terminal parts of the Cas9 protein. The polypeptides may be separated to form a "reconstituted" Cas9 protein. In specific embodiments, the Cas9 protein is, for example, a Cas9 protein as described herein, each of which is incorporated by reference. Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, which is incorporated herein by reference. As described in Jiang et al. (2016) Science 351: 867-871. PDB file: 5F9R In some embodiments, the protein is split into two fragments within the disordered region of the protein, such that The protein can be roughly aligned with any C, T, A, or S within the SpCas9 region. Between the amino acids A292-G364, F445-K483, or E565-T637, or any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or at the corresponding position in another nap DNA fragment. In this state, the protein is SpCas9T310, T313, A456, S469, or In some embodiments, the protein is split into two fragments at C574. The process of dividing the protein into two is called splitting the protein.
[0127] In other embodiments, the N-terminal portion of the Cas9 protein is selected from the group consisting of S. pyrogenes Cas9 wild-type and S. pyrogenes Cas9 wild-type. Type (SpCas9) (NCBI reference sequence: NC_002737.2, Uniprot reference The Cas9 protein contains amino acids 1 to 573 or 1 to 637 of the sequence Q99ZW2. The C-terminal portion of the protein is located at amino acids 574–1368 or 638–1368 in the wild-type SpCas9. Contains 68 parts.
[0128] The C-terminal part of the split Cas9 is joined to the N-terminal part of the split Cas9 to complete the In some embodiments, the Cas9 protein can be transformed into a complete Cas9 protein. The C-terminal part of the protein starts where the N-terminal part of the Cas9 protein ends. As such, in some embodiments, the C-terminal portion of the split Cas9 is Contains the amino acid (551-651)-1368 part. "(551-651)-1368" means that the amino acid sequence begins with amino acids 551 to 651 (inclusive) and ends with amino acid 1368. For example, the C-terminal part of split Cas9 ends with spCas9. Amino acids 551-1368, 552-1368, 553-1368, 554-1368, 555~1368, 556~1368, 557~1368, 558~1368, 559~ 1368、560~1368、561~1368、562~1368、563~1368 、564~1368、565~1368、566~1368、567~1368、568 ~1368、569~1368、570~1368、571~1368、572~136 8、573~1368、574~1368、575~1368、576~1368、57 7~1368、578~1368、579~1368、580~1368、581~13 68、582~1368、583~1368、584~1368、585~1368、5 86~1368、587~1368、588~1368、589~1368、590~1 368、591~1368、592~1368、593~1368、594~1368、 595~1368、596~1368、597~1368、598~1368、599~ 1368、600~1368、601~1368、602~1368、603~1368 、604~1368、605~1368、606~1368、607~1368、608 ~1368、609~1368、610~1368、611~1368、612~136 8、613~1368、614~1368、615~1368、616~1368、61 7~1368、618~1368、619~1368、620~1368、621~13 68、622~1368、623~1368、624~1368、625~1368、6 26~1368、627~1368、628~1368、629~1368、630~1 368、631~1368、632~1368、633~1368、634~1368、 635~1368、636~1368、637~1368、638~1368、639~ 1368, 640~1368, 641~1368, 642~1368, 643~1368 , 644~1368, 645~1368, 646~1368, 647~1368, 648 ~1368, 649~1368, 650~1368, or 651~1368 In some embodiments, the C-terminus of the split Cas9 protein may comprise a portion of one of the C-terminus of the split Cas9 protein. The end portion contains part of amino acids 574 to 1368 or 638 to 1368 of SpCas9. nothing.
[0129] "Subject" means a mammal, including, but not limited to, Human or non-human mammals, such as non-human primates (monkeys), cows, horses, dogs, sheep, In some embodiments, the subject described herein is a cat or a mammal having SDS. or identifying a subject as having a predisposition to developing SDS. Contains a pathogenic mutation within the SDS polynucleotide sequence encoding the gene.
[0130] "Substantially identical" means that a polypeptide or nucleic acid molecule has a similar amino acid sequence ( For example, any one of the amino acid sequences described herein) or nucleic acid sequences (e.g., "Any one of the nucleic acid sequences described in the specification" means that the nucleic acid sequence exhibits at least 50% identity with the nucleic acid sequence of the present invention. In one embodiment, such sequences are used for comparison at the amino acid or nucleic acid level. 9. The sequence is at least 60%, 65%, 70%, 75%, 80%, or 85% identical to the sequence 0%, 95%, or even 99% identical.
[0131] Sequence identity is typically determined using sequence analysis software (e.g., Genetics Computing). Sequence Analysis Software Package from the Computer Group, University of Wisconsin Biotechnology The University of Wisconsin, 1710 University Avenue, Madison, Wisconsin 53705, BLAST, BESTFIT, GAP, or PILEUP / PRETTY BOX program). Such software can detect various substitutions, deletions, , and / or by associating the degree of homology with other modifications Conservative substitutions are typically the following: glycine, alanine; valine, Isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; Within the group of serine, threonine; lysine, arginine; and phenylalanine, tyrosine In one exemplary approach to determining the degree of identity, closely related sequences are compared using a indicates that there is -3 and e -100 BLAST program with probability scores between may also be used.
[0132] COBALT is used, for example, with the following parameters: a) Alignment parameters: Gap penalty -11, -1 and End-G AP penalty -5, -1, b) CDD parameters: RPS BLAST used; Blast E-value 0.003; Find and recalculate saved columns, and c) Query clustering parameters: Use query cluster; Word size 4; Maximum cluster distance 0.8; Alphabet Regular. Using EMBOSS Needle, for example, with the following parameters: a) Matrix: BLOSUM62; b) GAP OPEN: 10; c) GAP EXTEND: 0.5; d)OUTPUT FORMAT:pair; e)END GAP PENALTY:false; f) END GAP OPEN: 10; and g)END GAP EXTEND:0.5.
[0133] The term "target site" refers to a deaminase or a fusion protein containing a deaminase (e.g., For example, a dCas9-adenosine deaminase fusion protein or a base enzyme disclosed herein. Deamination of a nucleic acid molecule refers to a sequence within a nucleic acid molecule that is deaminated by a deamidator.
[0134] RNA-programmable nucleases (e.g., Cas9) are capable of cleaving RNA:DNA hybrids. These proteins target DNA cleavage sites using cleavage. Cas can target any sequence specified by the guide RNA. RNA programmable nucleases such as 9 for site-specific cleavage (e.g., genome Methods used to modify the respective overall Cong, L. et al., Multiplex genome engineering, the contents of which are incorporated herein by reference. ng using CRISPR / Cas systems. Science 339, 819-823 (2013);Mali, P. et ah, RNA-gu ided human genome engineering via Cas9. Science 339, 823-826 (2013);Hwang, WY et ah, Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013);Jinek, M. et ah, RNA-programmed genome editing in human cells. eLife 2, e00471 (2013);Dicarlo, JE et ah, Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic acids research (2 013);Jiang, W. et ah RNA-guided editing of bacterial genomes using CRISPR-Cas s systems. Nature biotechnology 31, 233-239 (2013).
[0135] As used herein, the terms "treat," "treating," "treatment," and the like mean Reducing, diminishing, or reducing a disease or disorder and / or its associated symptoms to reduce, taper, alleviate, or ameliorate, or to achieve the desired pharmacological and refers to the acquisition of certain physiological and / or physiological effects, including but not limited to, disorders or conditions Treating a disorder, condition, or associated symptoms does not require complete elimination of the disorder, condition, or associated symptoms. It will be understood that in some embodiments, the effect is therapeutic, i.e., Without limitation, this effect may be partially or completely attributable to the disease and / or the condition. Reduce, diminish, prevent, abate, ameliorate, or lessen the intensity of adverse symptoms In some embodiments, the effect is preventative, i.e., That is, this effect protects against or prevents the occurrence and recurrence of a disease or condition. The methods disclosed herein comprise administering a therapeutically effective amount of a composition as described herein. In one embodiment, the present invention provides a treatment for SDS.
[0136] "Uracil glycosylase inhibitor" or "UGI" refers to an inhibitor of the uracil excision repair system. In one embodiment, the agent inhibits uracil DNA glycosylation in a host. A protein or fragment thereof that binds to an enzyme and prevents the removal of uracil residues from DNA In one embodiment, the UGI is a uracil DNA glycosylase base excision repair enzyme. In some embodiments, the protein is a protein or a fragment or domain thereof that can be inhibited. In some embodiments, the UGI domain comprises wild-type UGI or a modified version thereof. In some embodiments, the UGI domain comprises a fragment of the exemplary amino acid sequence set forth below. In some embodiments, the UGI fragment is at least 60% of the exemplary UGI sequences shown below. %, at least 65%, at least 70%, at least 75%, at least 80%, at least at least 85%, at least 90%, at least 95%, at least 96%, at least 9 7%, at least 98%, at least 99%, or 100% of the amino acid sequence In some embodiments, the UGI is selected from the group consisting of the exemplary UGI amino acid sequences as defined below: or a fragment thereof. In some embodiments, the UGI or A portion of the UGI sequence may be at least partially identical to wild-type UGI or a UGI sequence or portion thereof as defined below. at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, At least 99%, at least 99.5%, at least 99.9%, or 100% identical Exemplary UGIs include the following amino acid sequences: >splP14739IUNGI_BPPB2 Uracil DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLT SD APE YKPW ALVIQDS NGENKIKML.
[0137] Ranges set forth herein are understood to be shorthand for all values within the range. For example, the range 1 to 50 is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 3 8, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 It is understood that the term "range" includes any number, combination of numbers, or subrange from the group consisting of:
[0138] The recitation of a chemical group listing in any definition of a variable herein constitutes a definition of that variable. The definitions include any single group or combination of the listed groups. A description of an embodiment of an element or aspect does not necessarily translate that embodiment into any single embodiment or aspect. or in combination with any other embodiment or portion thereof.
[0139] Any composition or method described herein may be used in combination with other compositions and methods described herein. It can be combined with any one or more of the methods.
[0140] The description and examples herein provide detailed descriptions of embodiments of the present disclosure. It should be understood that the present invention is not limited to the specific embodiments described, as such may vary. Those of ordinary skill in the art will recognize numerous variations and modifications of this disclosure that would fall within its scope. It will be something to recognize.
[0141] All terms are intended to be understood as would be understood by a person skilled in the art. Unless otherwise defined, all technical and scientific terms used herein are defined by the The terms have the same meaning as commonly understood by one of ordinary skill in the art to which they pertain.
[0142] In the practice of some embodiments disclosed herein, unless otherwise indicated, Immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, etc., all of which are within the scope of the Conventional techniques of chemistry and recombinant DNA are employed. See, e.g., Sambrook and Green, Molec ular Cloning: A Laboratory Manual, 4th Edition (2012);the series Current Protoc ols in Molecular Biology (FM Ausubel, et al. eds.);the series Methods In Enz ymology (Academic Press, Inc.), PCR 2: A Practical Approach (MJ MacPherson, B. D. Hames and GR Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique an d Specialized Applications, 6th Edition (RI Freshney, ed. (2010)) stomach.
[0143] Various features of the disclosure may be described in the context of a single embodiment, but this The features of may be presented separately or in any suitable combination. Although they may be described herein in the context of separate embodiments, The disclosure may also be implemented in a single embodiment. It is for organizational purposes only and is not to be construed as limiting the subject matter described.
[0144] The features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the present disclosure may be obtained in view of the accompanying drawings, as described herein below. By reference to the following detailed description, illustrative embodiments of the principles are set forth. This is what you will gain. [Brief explanation of the drawings]
[0145] [Figure 1A]Figures 1A and 1B show mutations within SBDS that cause SDS. Figure 1A shows a map of SBDS (coding regions in light shading, noncoding regions in dark shading) and a sequence alignment of the exon 2 region of SBDS with the SBDS protein, indicating gene-specific (gray; top) and pseudogene-specific (gray; bottom) sequences. Compared to SBDS, exon 2 of SBDSP, resulting from the conversion event, contains sequence changes predicted to result in protein truncation (underlined). These include an in-frame stop codon at position 184 and a T→C change at 250+10 (corresponding to the invariant T of the donor splice site at 258+2 in SBDS), the latter change resulting in the use of an alternative donor splice site at 250+1 (the position of the invariant splice site is boxed). Figure 1B shows sequence reads for cloned segments derived from the exon 2 region of SBDS, highlighting sequence changes in individuals with SBDS resulting from gene conversion events between SBDS and its pseudogene. Three converted alleles are shown. These include 183-184TA → CT, 258 + 2T → C, and an extended conversion mutation, 183-184TA → CT + 201A → G + 258 + 2T → C. In each case, informative flanking positions, including 141 and 258 + 124, were unconverted (green). [Figure 1B] Same as above. [Figure 2A] Figures 2A-2D are schematic diagrams illustrating strategies for restoring transcription in SBDS genes containing one or more pathogenic mutations. Figure 2A illustrates a strategy for introducing a mutation that eliminates the stop codon and results in expression of an SBDS protein containing an alternative amino acid (e.g., Trp(W)) at amino acid position 62 (e.g., (K62X)). Figures 2B and 2D illustrate strategies for correcting the splice site at nucleotide position 258 (target SNP rs113993993 C→T). Figure 2C illustrates the splice donor position where the canonical splice donor is restored to correct the SNP mutation. [Figure 2B] Same as above. [Figure 2C] Same as above. [Figure 2D] Same as above. [Figure 3-1] Figures 3A-3C present tables showing amino acid positions where substitutions occur in modified Cas9 proteins, e.g., modified SpCas9, to produce Cas9 variants with specificity for an altered PAM 5'-NGC-3' or a 5'-NGC-3' containing a PAM, and plasmid constructs encoding the SpCas9 variant sequences. Cytidine base editors (CBEs) comprising at least one cytidine deaminase and at least one Cas9 variant as described are used to correct mutations in the SBDS gene associated with SDS, as described in Example 3. Figure 3A presents the amino acid positions in the Cas9 protein that were altered from wild-type to produce Cas9 variants (designated by numbers in the left column) that were capable of binding to the NGC PAM. These Cas9 variants are components of the CBEs evaluated in the base editing studies described herein. Figure 3B presents a subset of Cas9 variants that produced particularly good and robust on-target editing with limited bystander effects in the studies. Also shown in Figure 3B is a schematic representation of the Cas9 protein domains and their location within the Cas9 protein sequence. Figure 3C illustrates the plasmid vector components encoding SpCas9 variants with specificity for the modified PAM 5'-NGC-3' as described herein, and the sequence mutations therein. [Figure 3-2] Same as above. [Figure 3-3] Same as above. [Figure 3-4] Same as above. [Figure 3B] Same as above. [Figure 3C-1] Same as above. [Figure 3C-2] Same as above. [Figure 4] FIG. 4 depicts a graph comparing the relative mutation rates of base editing achieved by CBEs with different cytidine deaminases indicated on the abscissa. [Figure 5] 5 is a table showing the guide RNAs (gRNAs) used with the CBEs evaluated in the studies described herein. In embodiments, the gRNA sequences were components of the plasmid constructs used in the base editing studies described in the Examples. [Figure 6A] Figures 6A-6C show the percentage of editing (e.g., on-target editing) versus the percentage of bystander editing achieved by NGC CBE variants as described herein and 19-mer and 20-mer gRNAs, such as G88 and G44. In the right-hand graph of Figure 6A and Figure 6B, "PV226" and "PV230" refer to the plasmids used in the study. The PV226 plasmid contains a polynucleotide encoding Cas9 variant #226, the sequence of which is shown in Figures 3A-3C, and the PV230 plasmid contains a polynucleotide encoding Cas9 variant #230, the sequence of which is shown in Figures 3A-3C. The percentage of editing achieved by another NGC CBE containing a different Cas9 variant, the sequence of which is shown in Figures 3A-3C, and the 20-mer gRNA, G44, is shown in Figure 6C. [Figure 6B] Same as above. [Figure 6C] Same as above. [Figure 7A] Figures 7A and 7B show graphs of percent editing by NGC CBEs containing cytidine deaminases and Cas9 variants shown in Table 13 used in conjunction with a 19-mer gRNA (G88) and a 20-mer gRNA (G44) as described in Example 4 herein. [Figure 7B] Same as above. [Figure 8A]Figures 8A-8J show graphs of the percent base editing (on-target editing and bystander editing) achieved by NGC CBEs containing different cytidine deaminases and Cas9 (e.g., SpCas9) variant polypeptides with specific combinations of mutations within the Cas9 amino acid sequence as presented in Figures 3A-3C or Table 13, in conjunction with either a 19mer or 20mer gRNA, as assessed in a cell-based (HEK293) assay correcting a splice site SNP within the SBDS polynucleotide sequence. Figure 8A shows the contrast between on-target editing and bystander editing exerted by an NGC CBE containing Cas9 variant 225 and PpAPOBEC1, and by NGC CBEs 454 and 459 (Table 13), containing PpAPOBEC1 and Cas9 variants 226 and 244, respectively (Figures 3A-3C), when used with a 19mer (guide 88) gRNA. Figure 8B shows the percentage of on-target versus bystander editing achieved by NGC CBEs 454 and 459 containing Cas9 variant 225 and PpAPOBEC1, respectively, with a 20-mer (Guide44) gRNA, and PpAPOBEC1 and Cas9 variants 226 and 244 (Table 13). Figures 8C and 8D show the percentage of on-target and bystander base editing achieved by NGC CBEs containing AmAPOBEC1 cytidine deaminase and Cas9 variants 225, 226, and 244 (Figures 3A-3C), with either a 19-mer (Guide88) or a 20-mer (Guide44) gRNA. Figures 8E and 8F show the percentage of on-target and bystander base edits of NGC CBEs containing PmCDA1 cytidine deaminase and Cas9 variants 225, 453, and 458 (Table 13), together with either a 19-mer (Guide88) or a 20-mer (Guide44) gRNA.Figures 8G and 8H show the percentage of on-target and bystander base edits for NGC CBEs containing RRA3F cytidine deaminase and Cas9 variants 225, 455, and 460 (Table 13) with either a 19-mer (Guide88) or 20-mer (Guide44) gRNA. Figures 8I and 8J show the percentage of on-target and bystander base edits for NGC CBEs containing SsAPOBEC2 cytidine deaminase and Cas9 variants 225, 456, and 461 (Table 13) with either a 19-mer (Guide88) or 20-mer (Guide44) gRNA. In Figures 8A-8J, Cas9 variant 225 (or PV225) is alternatively referred to as "Beam Shuffle." [Figure 8B] Same as above. [Figure 8C] Same as above. [Figure 8D] Same as above. [Figure 8E] Same as above. [Figure 8F] Same as above. [Figure 8G] Same as above. [Figure 8H] Same as above. [Figure 8I] Same as above. [Figure 8J] Same as above. [Figure 9A] Figures 9A-9D show graphs and dot plots of percent editing by NGC CBEs containing PpAPOBEC1 cytidine deaminase polypeptide sequences containing various mutations as described in Example 4, such as the H122A mutation alone and in combination with the amino acid mutations R33A, W90F, K34A, R52A, H121A, and Y120F, using either a 19-mer gRNA (Figure 9A) or a 20-mer gRNA (Figure 9B). The percentage of on-target versus bystander editing was assessed in an in vitro cell-based assay. Figures 9C and 9D present the data from Figures 9A and 9B, respectively, in dot plot format. [Figure 9B] Same as above. [Figure 9C] Same as above. [Figure 9D] Same as above. [Figure 10-1] Figure 10 presents a table illustrating the mutations and combinations of mutations made in the SpCas9 protein to create SpCas9 variants with the indicated combinations of mutations, including "NRCH" mutations as described in S. Miller et al., April 2020, "Continuous evolution of SpCas9 variants compatible with non-G PAMs," Nature Biotechnology, 38(4):471-481 (published online February 10, 2020, doi: 10.1038 / s41587-020-0412-8), the contents of which are incorporated herein by reference in their entirety. Combinations of NRCH mutations (amino acid substitutions) were included in several different SpCas9 variants to determine which would produce SpCas9 variant components of the NGC CBE for use in correcting a splice site SNP in the SBDS gene associated with SDS with enhanced on-target versus bystander base editing (Example 6). In Figure 10, darker shaded amino acids reflect amino acid substitutions in the Cas9 (SpCas9) amino acid sequence compared to the sequence of the wild-type, unmutated Cas9 (SpCas9) protein, and lighter shaded amino acids reflect amino acid residues in the wild-type, unmutated Cas9 (SpCas9) protein. [Figure 10-2] Same as above. [Figure 11A]11A and 11B show graphs illustrating the percent editing by an NGC CBE containing a cytidine deaminase (e.g., PpAPOBEC1) and an SpCas9 variant containing one or more NRCH mutations as defined in FIG. 10 and Example 5, used in conjunction with either a 19-mer or a 20-mer gRNA, in a cell-based assay evaluating the on-target and bystander editing efficiencies of these CBEs to correct a splice site SNP in the SBDS gene associated with SDS. NGC CBEs 468 and 469 (FIG. 10) showed high levels of on-target versus off-target base editing when used in conjunction with either a 19-mer or a 20-mer gRNA. [Figure 11B] Same as above. [Figure 12A] 12A-12C show graphs illustrating the results of in vitro cell-based assays performed to assess the base editing efficiency of NGC CBEs encoded by the mRNAs described in Example 6 in conjunction with gRNAs of different lengths (17mer, 18mer, 19mer, 20mer, or 21mer), and the percentage of on-target edits versus bystander edits. As observed, mRNA342 with 18mer and 20mer gRNAs has the fewest C-to-A or C-to-G transitions compared to mRNA340 or mRNA341. [Figure 12B] Same as above. [Figure 12C] Same as above. DETAILED DESCRIPTION OF THE INVENTION
[0146] The present invention uses programmable nucleic acid base editors to induce aberrant splicing events within genes. The aim of this study was to develop a recombinant protein that edits the pathogenic gene mutations that cause HIV-1 infection and enables transcription to achieve a therapeutic effect. In some embodiments, the editing comprises editing a sequence from a stop codon to a target sequence. In some embodiments, the editing comprises converting the codon to a codon that allows transcription. providing and modifying splice acceptor or splice donor sites; or This includes providing alternative splice acceptor or splice donor sites. In some embodiments, one or more mutations resulting from aberrant splicing are corrected.
[0147] The present invention is based, at least in part, on the discovery of adenosine base editors or cytidine base editors. Using ABE and CBE, we investigated Shwachman-Diamond syndrome (SDS). A strategy to edit pathogenic mutations (e.g., mutations resulting from gene conversion) within the relevant gene Thus, the present invention provides an ABE or a medicament useful for treating or preventing SDS. A base editor system comprising a CBE is provided.
[0148] Shwachman-Diamond Syndrome (SDS) Shwachman-Diamond syndrome (SDS) is an autosomal recessive disorder. Approximately 90% of patients who meet the clinical diagnostic criteria for Shwachman-Bodian-Dye syndrome The carrier frequency of this mutation is approximately It is estimated that 1 in 110 people have this condition. This highly conserved gene has five exons. They encompass 7.9 kb and map to the centromeric region of chromosome 7, 7q11. The SDBS gene is a novel gene lacking homology to known protein functional domains. It encodes a 250 amino acid protein. The adjacent pseudogene, SBDSP, encodes a 250 amino acid protein. It shares 97% homology with S but contains deletions and nucleosides that prevent the production of functional protein. Approximately 75% of patients with SDS have a gene conversion event involving this pseudogene. Gene conversion occurs when homologous sequences (parasites) located at different genomic loci are exchanged. This occurs when recombination occurs between homologous sequences. The existence of the gene may be due to past gene duplication. The protein is widely expressed throughout human tissues at both the mRNA and protein levels. The early truncating SBDS mutation 183TA>CT is common among patients with SDS. However, no patients homozygous for this mutation have been identified, suggesting that SBDS This suggests that complete lack of expression may be lethal in human patients.
[0149] A common sequence change associated with SDS is the TA→CT dinucleotide at positions 183–184. These include a nucleotide change or an 8-bp deletion at the end of exon 2. Analysis revealed that individuals with SDS who express the deletion transcript have a 183-184TA→CT The presence of the mutation was confirmed and the 258+2T→C mutation was identified. An 8-bp deletion predicted to disrupt the donor splice site in intron 2 This is consistent with the use of a potential upstream splice donor site at positions 251-252. Nucleotide change 183-184 TA → CT results in an in-frame stop codon (K62X). The 258+2T→C and resulting 8-bp deletion resulted in a frameshift. This results in premature truncation of the encoded protein (84Cfs3).
[0150] The present invention relates to one or more alterations (e.g., genetic alterations) that result in aberrant splicing. allowing transcription of a polynucleotide containing a functional SBDS A protein (e.g., a protein with sufficient activity to mitigate the effects of SBDS gene conversion) In a specific embodiment, the present invention provides compositions and methods for the expression of a gene encoding a gene encoding a protein. The present invention provides a method for inserting Trp from the TAA end into the SBDS gene containing 183-184TA→CT. This results in the introduction of a modification that converts the coding sequence to TGG, thereby allowing transcription. In one embodiment, the present invention provides a method for the preparation of a polypeptide encoding a biologically active protein. Polynucleotides that introduce splice donor or effector sites that allow cleavage In some embodiments, the present invention provides a method for introducing nucleotide sequence alterations into the SBDS gene. A site within exon 2 (e.g., at nucleotide position 1495 as shown in Figure 2B) (by editing cytosines) to correct.
[0151] Nucleic acid base editor Disclosed herein are methods for editing, modifying, or otherwise targeting nucleotide sequences of polynucleotides. or a base editor or nucleobase editor for modifying the base of a nucleic acid. The target is a polynucleotide programmable nucleotide binding domain and a nucleic acid base editing domains (e.g., adenosine deaminase, cytidine deaminase) and nucleobases A polynucleotide programmable nucleotide is a base editor or a base editor. The binding domain, when linked to a binding guide polynucleotide (e.g., gRNA), , specifically to the target polynucleotide sequence (i.e., the bases of the bound guide nucleic acid and the target (through complementary base pairing between the bases of the polynucleotide sequence) This allows the base editor to be localized to target the nucleic acid sequence targeted for editing. In some embodiments, the target polynucleotide sequence is single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid.
[0152] Polynucleotide Programmable Nucleotide Binding Domains Polynucleotide programmable nucleotide binding domains bind RNA to nucleic acids It should be understood that programmable proteins may also be included. For example, polynucleosides The polynucleotide programmable nucleotide binding domain The nucleic acid may be associated with a nucleic acid that guides the nucleic acid binding domain to the RNA. DNA binding proteins are also within the scope of this disclosure, even though they are not specifically listed in this disclosure. is located.
[0153] The polynucleotide-programmable nucleotide-binding domain of the base editor is It may itself contain one or more domains. For example, a polynucleotide programmable The nucleotide binding domain may comprise one or more nuclease domains. In an embodiment, the nuclease of the polynucleotide programmable nucleotide binding domain The enzyme domain may comprise an endonuclease or an exonuclease. The term "exonuclease" refers to an enzyme that breaks down nucleic acids (e.g., RNA or DNA) into free The term "endonuclease" refers to a protein or polypeptide that can be digested from its ends. "Protease" refers to a protease that catalyzes (e.g., catalyzes) an internal region of a nucleic acid (e.g., DNA or RNA). In some embodiments, an enzyme refers to a protein or polypeptide that can be cleaved. An endonuclease can cleave one strand of a double-stranded nucleic acid. In some cases, endonucleases can cleave both strands of a double-stranded nucleic acid molecule. In embodiments, the polynucleotide programmable nucleotide binding domain is a deoxyribonucleotide. In some embodiments, the polynucleotide programmable nuclease The tide binding domain may be a ribonuclease.
[0154] In some embodiments, the nucleotide of a polynucleotide programmable nucleotide binding domain The cleavage domain can cleave zero, one, or two strands of a target polynucleotide. In some cases, the polynucleotide programmable nucleotide binding domain comprises a As used herein, the term "nickase" refers to a double-stranded nucleic acid molecule that may contain a nickase domain. Nuclease domains that can cut only one strand of a double strand in a molecule (e.g., DNA) A polynucleotide-programmable nucleotide binding domain containing In embodiments, the nickase is capable of converting one or more mutations into an active polynucleotide program. By incorporating a nucleotide-binding domain into the polynucleotide programmer derived from a fully catalytically active (e.g., native) form of the nucleotide-binding domain For example, the polynucleotide programmable nucleotide binding domain can be When a Cas9-derived nickase domain is included, the Cas9-derived nickase domain may contain a D10A mutation and a histidine at position 840. In such a case, residue H840 retains catalytic activity and can thereby cleave a single strand of a nucleic acid duplex. , whereas the Cas9-derived nickase domain can contain the H840A mutation. The amino acid residue at position 0 remains D. In some embodiments, the nickase By removing all or part of the nuclease domain that is not required for activity, The programmable nucleotide binding domain is fully catalytically active (e.g. For example, the polynucleotide programmable nucleic acid may be obtained from a (naturally occurring) form. When the peptide-binding domain contains the nickase domain from Cas9, this Cas The nickase domain derived from 9 is a whole or part of the RuvC domain or HNH domain. It may include a deletion of
[0155] The amino acid sequence of an exemplary catalytically active Cas9 is as follows: MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD.
[0156] Thus, polynucleotides containing a nickase domain can programmably bind nucleotides. The base editor containing the domain can target a specific polynucleotide target sequence (e.g., a target region of the target polypeptide). The DNA fragments are then inserted into the DNA strands, which then generate single-stranded DNA breaks (nicks) at the target site (determined by the complementary sequence of the inserted guide nucleic acid). In some embodiments, a nickase domain (e.g., a Cas9-derived Nucleic acid duplex target polynucleotides cleaved by base editors containing a nickase domain The strand of the nucleotide sequence is the strand that is not edited by the base editor (i.e., (The strand cut by the editor is opposite to the strand containing the base to be edited.) Other Embodiments In this embodiment, a base sequence containing a nickase domain (e.g., a nickase domain derived from Cas9) is used. The editor can cut the strand of the DNA molecule that is targeted for editing. In this case, the non-targeted strand is not cleaved.
[0157] Also provided herein are catalytically inactive (i.e., target polynucleosides) a polynucleotide programmable nucleotide binding domain (which cannot cleave the nucleotide sequence) As used herein, the terms "catalytically inactive" and "nucleic acid" refer to base editors that "No cleavage activity" refers to one or more mutations and / or modifications that result in the inability to cleave nucleic acid strands. and / or deletions in the polynucleotide programmable nucleotide binding domain. In some embodiments, catalytically inactive polynucleotides are used interchangeably to refer to The nucleotide-programmable nucleotide-binding domain base editor comprises one or more Lacking nuclease activity as a result of specific point mutations within the nuclease domain For example, in the case of base editors containing a Cas9 domain, Cas9 may Such mutations may include both the H840A and H840A mutations. In another embodiment, the nuclease domain is inactivated, thereby resulting in a loss of nuclease activity. The catalytically inactive polynucleotide programmable nucleotide binding domain , all or part of a catalytic domain (e.g., RuvC1 and / or HNH domain) In yet another embodiment, the catalytically inactive polypeptide may comprise one or more deletions of the The programmable nucleotide binding domain may be modified by a point mutation (e.g., D10A or H840A) as well as deletion of all or part of the nuclease domain.
[0158] Also contemplated herein are polynucleotide programmable nucleotide binding domains. Catalytically inactive polynucleotide proteases are synthesized from the previously functional versions of the main Mutations that can generate mutable nucleotide-binding domains, e.g., catalytic and In the case of Cas9 ("dCas9"), which is inactive relative to the nuclease-inactivated Cas9, Variants having mutations other than D10A and H840A are provided. Mutations include, for example, other amino acid substitutions at D10 and H840, or substitutions of the Cas9 nucleotides. Other substitutions within the nuclease domain (e.g., the HNH nuclease subdomain and / or (substitutions within the RuvC1 subdomain). Additional suitable nuclease-inactive d The Cas9 domain will be apparent to those skilled in the art based on this disclosure and knowledge in the art. Such additional exemplary suitable nuclease-inactive Cas9 polypeptides are within the scope of the disclosure. Domains include, but are not limited to, the D10A / H840A mutant domain, D1 0A / D839A / H840A mutant domain, and D10A / D839A / H840 A / N863A mutant domain (see, e.g., the entire contents of the present invention). Prashant et al., CAS9 transcriptional activators for target sp ecificity screening and paired nickases for cooperative genome engineering. Natu re Biotechnology. 2013; 31(9): 833-838).
[0159] Polynucleotide programmable nucleotide binding domains that can be incorporated into base editors Main, non-limiting examples include domains from CRISPR proteins, restriction nucleases, Nucleases, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases In some cases, base editors can be used to manipulate nucleic acid CRIs. SPRs (i.e., clustered regularly interspaced short palindromic repeats) A gene that can bind to a nucleic acid sequence via a bound guide nucleic acid during mediated modification. Polynucleotide programmable nucleic acids containing natural or modified proteins or portions thereof Such proteins are referred to herein as "CRISPR targets." Thus, disclosed herein are CRISPR proteins. a polynucleotide containing all or a portion of the protein; base editors (i.e., the "CRISPR protein-derived domain" of the base editor) A base sequence containing all or part of the domain of a CRISPR protein, also known as a "CRISPR gene," is The CRISPR protein-derived domain incorporated into the base editor is CRISPR proteins can be modified compared to the wild-type or native version. For example, as described below, a domain from a CRISPR protein may be one or more mutations, insertions, or These may include insertions, deletions, rearrangements, and / or recombinations.
[0160] CRISPR targets mobile genetic elements (viruses, transposable elements, and conjugative plasmids). The CRISPR cluster is a spacer, a precursor, and a A CRISPR cluster contains a sequence complementary to an example mobile element and a target entry nucleic acid. is processed into transcribed CRISPR RNA (crRNA). Type II C In the RISPR system, correct processing of the crRNA precursor is required for transcription in trans. The transduced small RNA (tracrRNA) and endogenous ribonuclease 3 (rnc) The tracrRNA requires the ribonucleotides of the crRNA precursor and the Cas9 protein. It serves as a guide for nuclease 3-assisted processing. s9 / crRNA / tracrRNA contains a linear or circular dsD complementary to the spacer. The target strand that is not complementary to the crRNA is cleaved endonucleolytically. First, it is cleaved endonucleolytically, and then cleaved 3'-5' exonucleolytically. In nature, DNA binding and cleavage typically involves both proteins and RNA. However, both aspects of crRNA and tracrRNA are incorporated into a single RNA species. A single guide RNA ("sgRNA" or simply "gNRA") was engineered to target the target gene. See, for example, J.I. k M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 3 37:816-821 (2012). Cas9 targets short moieties within CRISPR repeats. recognizes the PAM or protospacer adjacent motif to distinguish self from non-self This will help.
[0161] In some embodiments, the methods described herein involve the use of engineered Cas proteins. Guide RNA (gRNA) is a scanning RNA required for Cas binding. The fold sequence and a user-defined sequence of approximately 20 nucleotides that define the genomic target to be modified. It is a short synthetic RNA consisting of a nucleic acid sequence and a spacer of nucleic acid fragments. By altering the target sequence present in the gRNA, the Cas protein can be targeted to the genome or The target of the polynucleotide can be varied. Cas targets RNA depending on how specific it is relative to the rest of the genome. The specificity of a protein is determined in part.
[0162] In some embodiments, the gRNA scaffold sequence is: GUUUUAGAGC UAGAAAUAGC AAGUUAAAAU AAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU.
[0163] In one embodiment, the RNA scaffold comprises a stem loop. The NA scaffold contains the following nucleic acid sequence: GUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAAC GAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAU GGCAGGGUG.
[0164] In one embodiment, the RNA scaffold comprises the following nucleic acid sequence: GUUUUAGAGCUAGAAA UAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU.
[0165] In one embodiment, the S. pyrogenes sgRNA scaffold polynucleotide sequence is as follows: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC.
[0166] In one embodiment, the S. aureus sgRNA scaffold polynucleotide sequence is as follows: GUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGC GAGA.
[0167] In one embodiment, the sgRNA scaffold for BhCas12b is composed of the following polynucleotides: It has the following sequence: GUUCUGTCUUUUGGUCAGGACAACCGUCUAGCUAUAAGUGCUGCAGGGUGUGAGAAACUCCUAUUGCUGGACGAUGUCUC UUACGAGGCAUUAGCAC.
[0168] In one embodiment, the sgRNA scaffold for BvCas12b is composed of the following polynucleotides: It has the nucleotide sequence: GACCUAUAGGGUCAAUGAAUCUGUGCGUGUGCCAUAAGUAAUUAAAAAUUACCCACCA CAGGAGCACCUGAAAACAGGUGCUUGGCAC.
[0169] In some embodiments, the domains from the CRISPR protein incorporated into the base editors The nucleic acid is capable of binding to a target polynucleotide when linked to a binding guide nucleic acid. capable endonucleases (e.g., deoxyribonucleases or ribonucleases) In some embodiments, the base editor is derived from a CRISPR protein incorporated therein. The binding domain binds to a target polynucleotide when linked to a binding guide nucleic acid. In some embodiments, a CR The ISPR protein-derived domain binds to the target polynucleotide when linked to the bound guide nucleic acid. A catalytically inactive domain capable of binding to ribonucleotides. In this state, the target bound by the CRISPR protein-derived domain of the base editor In some embodiments, the polynucleotide is DNA. The target polynucleotide bound by the R protein-derived domain is RNA.
[0170] Cas proteins that can be used herein include class 1 and class 2 Ca Non-limiting examples of Cas proteins include Cas1, Cas2, 1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h , Cas5a, Cas6, Cas7, Cas8, Cas9 (with Csn1 or Csx12) (also known as Cas10, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1 , Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1 , Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx1 7, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S , Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12 a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / Ca sY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i, C ARF, DinG, their homologs, or modified versions thereof. The modified CRISPR enzymes can have DNA cleavage activity, such as Cas9, and CRISPR enzymes have functional endonuclease domains: RuvC and HNH. The enzyme may induce cleavage of one or both of the target sequences, e.g., within and / or between the target sequences. For example, the CRISPR enzyme can guide one or both of the cuts The fragment is cut at about 1, 2, 3, 4, 5, 6, 7, 8 nucleotides from the first or last nucleotide of the target sequence. , 9, 10, 15, 20, 25, 50, 100, 200, 500 or more can be introduced into the base pair.
[0171] The ability to cleave one or both strands of a target polynucleotide containing the target sequence is determined by mutating the The CRISPR enzyme was mutated to eliminate the corresponding wild-type enzyme. A vector encoding the enzyme can be used. Cas9 is an exemplary wild-type Cas 9 polypeptide (e.g., Cas9 from S. pyrogenes), at least 50%, 6 0%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 9 7%, 98%, 99%, or 100% or at least about 50%, about 60%, or about 70% , about 80%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96% , about 97%, about 98%, about 99%, or about 100% sequence identity and / or sequence homology Cas9 can refer to a polypeptide having the same structure as the wild-type exemplary Cas9 polypeptide. up to 50%, 60%, 70%, 80%, 90% against thrombin (e.g., from S. pyrogenes) %, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, and is 100% or up to approximately 50%, approximately 60%, approximately 70%, approximately 80%, approximately 90%, approximately 91%, Approximately 92%, approximately 93%, approximately 94%, approximately 95%, approximately 96%, approximately 97%, approximately 98%, approximately 99%, or may refer to a polypeptide with about 100% sequence identity and / or sequence homology. Cas9 can be used to generate amino acid changes, such as deletions, insertions, substitutions, variants, mutations, fusions, and wild-type or wild-type Cas9 proteins, which may include, for example, a nucleotide sequence encoding a nucleotide sequence encoding a nucleotide sequence of ... It may refer to a variant.
[0172] In some embodiments, the base editor CRISPR protein-derived domain is Nebacterium ulcerans (NCBI reference numbers: NC_015683.1, NC_0 17317.1); Corynebacterium diphtheriae (NCBI reference number: NC_016 782.1, NC_016786.1); Spiroplasma sylphidicola (NCB Reference number: NC_021284. 1); Prevotella intermedia (NCBI reference No.: NC_017861.1); Spiroplasma taiwanense (NCBI Reference No.: NC_021846.1); Streptococcus iniae (NCBI reference number: NC_0 21314.1); Verriella baltica (NCBI reference number: NC_018010.1 ); Cycloflexus torquhis (NCBI reference number: NC_018721.1); Streptococcus thermophilus (NCBI reference number: YP_820832.1); Streptomyces innocua (NCBI reference number: NP_472073.1); Campylobacter Neisseria jejuni (NCBI reference number: YP_002344900.1); Neisseria ningitidis (NCBI reference number: YP_002342100.1), Streptococcus Cas9 from Staphylococcus pyrogenes or Staphylococcus aureus Or it may include a part of it.
[0173] The Cas9 Cas9 Castle Case The Cas9 range of casings is specifically designed to be specific (for example Its publication is titled “Complete genome sequence o f an Ml strain of Streptococcus pyogenes.” Ferretti et al , JJ , McShan WM . Ajdic DJ, Savic DJ, Savic G, Lyon K, Primeaux C, Sezate S, Suvorov AN, [ PMC free article ] [ PubMed ] Kenton S, Lai HS, Lin SP, Qian Y, Jia HG, Najar FZ, Ren Q, Zhu H, S. et al ong L, White J, Yuan X, Clifton SW, Roe BA, McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001);“CRISPR RNA maturation by trans-encoded s mall RNA and host factor RNase III.” Deltcheva E , Chylinski K , Sharma CM , G onzales K, Chao Y, Pirzada ZA, Eckert MR, Vogel J, Carpenter E, Nature 471:602-607(2011);および “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” . Jinek M. , Chylinski K. , Fonfara I. , Hauer M. , Do See Udna JA, Charpentier E. Science 337:816-821 (2012). Cas9 orthologs has been described in a variety of species, including but not limited to S. Additional suitable Cas9 nucleases include S. pyrogenes and S. thermophilus. The sequences and sequences will be apparent to those skilled in the art based on this disclosure, and the sequences and sequences of such Cas9 nucleic acids may be used in The enzymes and sequences are described in Ch, the entire contents of which are incorporated herein by reference. ylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRIS PR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737 The Cas9 sequence and locus are included.
[0174] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is C Non-limiting exemplary Cas9 domains are provided herein. The Cas9 domain is divided into two types: the nuclease-active Cas9 domain and the nuclease-inactive Cas9 domain. In some embodiments, the Cas9 domain may be a Cas9 nickase. The Cas9 domain is a nuclease active domain. For example, the Cas9 domain is a double-stranded nuclease. The Cas9 domain may cleave both strands of a DNA molecule (e.g., both strands of a double-stranded DNA molecule). In some embodiments, the Cas9 domain has the amino acid sequence as defined herein: In some embodiments, the Cas9 domain comprises any one of the following: At least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, At least 95%, at least 96%, at least 97%, at least 98%, at least In some embodiments, the amino acid sequence is at least 99%, or at least 99.5%, identical to the amino acid sequence of the target gene. In the present specification, the Cas9 domain is compared to any one of the amino acid sequences defined herein. All 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 , 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 4 3, 44, 45, 46, 47, 48, 49, 50, or more mutations In some embodiments, the Cas9 domain comprises an amino acid sequence as defined herein. at least 10, at least 15, at least At least 20, at least 30, at least 40, at least 50, at least 60, at least At least 70, at least 80, at least 90, at least 100, at least 150, at least At least 200, at least 250, at least 300, at least 350, at least 4 00, at least 500, at least 600, at least 700, at least 800, At least 900, at least 1000, at least 1100, or at least 1200 It includes an amino acid sequence having 10 identical adjacent amino acid residues.
[0175] In some embodiments, proteins comprising fragments of Cas9 are provided. For example, some In embodiments, the protein comprises one of the following two Cas9 domains: (1) (2) the gRNA-binding domain of Cas9; or (3) the DNA-cleavage domain of Cas9. In embodiments, proteins comprising Cas9 or fragments thereof are referred to as "Cas9 variants." Cas9 variants share homology with Cas9 or fragments thereof. For example, The Cas9 variant is at least about 70% homologous to wild-type Cas9, or at least or at least about 80% homologous, or at least about 90% homologous, or at least about 95% homologous. or at least about 96% homologous, or at least about 97% homologous, or at least or at least about 98% homologous, or at least about 99% homologous, or at least about 99.5% homologous. In some embodiments, the Cas9 The variants are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 1 compared to wild-type Cas9. 1, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 , 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, In some embodiments, the Cas9 Variants include fragments of Cas9 (e.g., the gRNA binding domain or the DNA cleavage domain). fragment) is at least about 70% homologous to the corresponding fragment of wild-type Cas9 or at least about 80% homologous, or at least about 90% homologous, or at least about 95% homologous or at least about 96% homologous, or at least about 97% homologous, or at least about 98% homologous, or at least about 99% homologous, or at least about 99.5% homologous or has this fragment so as to be at least about 99.9% homologous. In embodiments, the fragments are at least 30%, at least 35% of the corresponding wild-type Cas9. , at least 40%, at least 45%, at least 50%, at least 55%, less At least 60%, at least 65%, at least 70%, at least 75%, at least 80% %, at least 85%, at least 90%, at least 95% identical, at least 96% , at least 97%, at least 98%, at least 99%, or at least 99.5% In some embodiments, the fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 25 0, 300, 350, 400, 450, 500, 550, 600, 650, 700, 75 0, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 12 00, 1250, or at least 1300 amino acids in length.
[0176] In some embodiments, the Cas9 fusion proteins as provided herein comprise a Cas The full-length amino acid sequence of the Cas9 protein, e.g., one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins as provided herein comprise It does not contain the full-length Cas9 sequence, but only one or more fragments thereof. Exemplary amino acid sequences of s9 domains and Cas9 fragments are provided herein: Additional suitable sequences for Cas9 domains and fragments will be apparent to those skilled in the art.
[0177] The Cas9 protein binds to a specific DNA sequence that is complementary to the guide RNA. In some embodiments, the polynucleotide may be associated with a guide RNA that guides the protein. The programmable nucleotide binding domain is a Cas9 domain, e.g., For example, nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease It is a nucleic acid-programmable DNA-binding protein called dCas9 (dCas9). Examples of such proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), Ca sX, CasY, Cpf1, Cas12b / C2C1, and Cas12c / C2C3 Examples include:
[0178] In some embodiments, wild-type Cas9 is expressed in Streptococcus pyrogenes (NCBI). Reference sequence: NC_017053.1, nucleotide and amino acid sequences as shown below It is compatible with the upcoming Cas9. ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGATTATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG GCAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGCAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAATCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTAGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAGAAATGGCTTGTTTGGGAAT CTCATTGCTTTGTCATTGGGATTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATAGTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAGCGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAGGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGCGCCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGGGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGATATTCAAAAAGCACAGGTGTCTGGACAAGGCCATAGTTTA CATGAACAGATTGCTAACTTAGCTGGCAGTCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAATTGTTGATGAACT GGTCAAAGTAATGGGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCCAGA AAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCTGTT GAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTACAAAATGGAAGAGACATGTATGTGGACCAAGAATT AGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCATTAAAGACGATTCAATAGACAATA AGGTACTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAAAAC TATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGGTTT GAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAATTT TGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCTAAA TTAGTTTCTGACTTCCGAAAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGCGTA TCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATAAAG TTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCTAAT ATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGGGGA AACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATATTG TCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTACCAAAAAGAAATTCGGACAAGCTTATTGCT CGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGCTAA GGTGGAAAAAGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCTTTG AAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAAAGACTTAATCATTAAACTACCTAAATAT AGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAAAAGGAAATGAGCTGGCTCT GCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAACAAA AACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTTATT TTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGAAAA TATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTAAAC GATATACGTCTACAAAAGAAGTTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATTGAT TTGAGTCAGCTAGGAGGTGACTGA JPEG2025183226000004.jpg117168 (single underline: HNH domain; double underline: RuvC domain)
[0179] In some embodiments, the wild-type Cas9 comprises the following nucleotides and / or amino acids: Corresponding to or containing the sequence: ATGGATAAAAAGTATTCTATTGGTTTAGACATCGGCACTAATTCCGTTGGATGGGCTGTCATAACCGATGAATACAAAGT ACCTTCAAAGAAATTTAAGGTGTTGGGGAACACAGACCGTCATTCGATTAAAAAAGAATCTTATCGGTGCCCTCCTATTCG ATAGTGGCGAAACGGCAGAGGCGACTCGCCTGAAACGAACCGCTCGGAGAAGGTATACACGTCGCAAGAACCGAATATGT TACTTACAAGAAATTTTTAGCAATGAGATGGCCAAAGTTGACGATTCTTTCTTTCACCGTTTGGAAGAGTCCTTCCTTGT CGAAGAGGACAAAACATGAACGGCACCCCATCTTTGGAAACATAGGATGAGGTGGCATATCATGAAAAGTACCCAA CGATTTATCACCTCAGAAAAAAGCTAGTTGACTCAACTGATAAAGCGGACCTGAGGTTAATCTACTTGGCTCTTGCCCAT ATGATAAAGTTCCGTGGGCACTTTCCATTGAGGGTGATCTAAATCCGGACAACTCGGATGTCGACAAACTGTTCATCCA GTTAGTACAAACCTATAATCAGTTGTTTGAAGAACCCTATAAATGCAAGTGGCGTGGATGCGAAGGCTATTCTTAGCG CCCGCCTCTCTAAATCCCGACGGCTAGAAAAACCTGATCGCACAATTACCCGGAGAGAAAAAAATGGTTGTTCGGTAAC CTTATAGCGCTCTCACTAGGCCTGACACCAAATTTAAGTCGAACTTCGACTTAGCTGAAGATGCCAAATTGCAGCTTAG TAAGGACACGTACGATGACGATCTCGCAATCTACTGGCACAAATTGGAGATCAGTATGCGGACTTATTTTTGGCTGCCA AAAACCTTAGCGATGCAATCCTCCTATCTGACATACTGAGAGTTAATACTGAGATTACCAAGGCGCCGTTATCCCGCTTCA ATGATCAAAAGGTACGATGAACATCACCAAGACTTGACACTTCTCAAGGCCCTAGTCCGTCAGCAACTGCCTGAGAAATA TAAGGAAATATTCTTTGATCAGTCGAAAACGGGTACGCAGGTTATATTGACGGGCGGAGCGAGTCAAGAGGAATTCTACA AGTTTATCAAACCCATATTAGAGAATGGATGGGACGGAAGAGTTGCTTGTAAAACTCAATCGCGAAGATCTACTGCGA AAGCAGCGGACTTTCGACAACGGTAGCATTCCACATCAAATCCACTTAGGCGAATTGCATGCTATACTTAGAAGGCAGGA GGATTTTTATCCGTTCCTCAAAGACAATCGTGAAAAGATTGAGAAAATCCTAACCTTTCGCATACCTACTATGTGGGAC CCCTGGCCGAGGGAACTCTCGGTTCGCATGGATGACAAGAAAGTCCGAAGAAACGATTACTCCATGGAATTTTGAGGAA GTTGTCGATAAAGGTGCGTCAGCTCAATCGTTCATCGAGAGGATGACCACTTTGACAAGAATTTACCGAACGAAAAAGT ATTGCCTAAGCACAGTTTACTTTACGAGTATTTCACAGTGTACAATGAACTCACGAAAGTTAAGTATGTCACTGAGGGCA TGCGTAAACCCGCCTTTCTAAGCGGAGAACAGAAGAAGCAATAGTAGATCTGTTATTCAAGACCAACCGCAAAGTGACA GTTAAGCAATTGAAAGAGGACTACTTTAAGAAAATTGAATGCTTCGATTCTGTCGAGATCTCCGGGGTAGAAGATCGATT TAATGCGTCACTTGGTACGTATCATGACCTCCTAAAGATAATTAAAGATAAGGACTTCCTGGATAACGAAGGAATGAAG ATATCTTAGAAGATATAGTGTTGACTCTTACCCTCTTTGAAGATCGGGAAATGATTGAGGAAAGACTAAAAACATACGCT CACCTGTTCGACGATAAGGTTATGAAACAGTTAAAGAGGCGTCGCTATACGGGCTGGGGACGATTGTCGCGGAAACTTAT CAACGGGATAAGAGACAAGCAAAGTGGTAAAACTATTCTCGATTTTCTAAAGAGCGACGGCTTCGCCAATAGGAACTTTA TGCAGCTGATCCATGATGACTCTTTAACCTTCAAAGAGGATATACAAAAGGCACAGGTTTCCGGACAAGGGGACTCATTG CACGAACATATTGCGAATCTTGCTGGTTCGCCAGCCATCAAAAAGGGCATACTCCAGACAGTCAAAGTAGTGGATGAGCT AGTTAAGGTCATGGGACGTCACAAACCGGAAAACATTGTAATCGAGATGGCACGCGAAAATCAAACGACTCAGAAGGGGC AAAAAAACAGTCGAGAGCGGATGAAGAGAATAGAAGAGGGTATTAAAGAACTGGGCAGCCAGATCTTAAAGGAGCATCCT GTGGAAAATACCCAATTGCAGAACGAGAAACTTTACCTCTATTACCTACAAAATGGAAGGGACATGTATGTTGATCAGGA ACTGGACATAAACCGTTTATCTGATTACGACGTCGATCACATTGTACCCCAATCCTTTTTGAAGGACGATTCAATCGACA ATAAAGTGCTTACACGCTCGGATAAGAACCGAGGGAAAAGTGACAATGTTCCAAGCGAGGAAGTCGTAAAGAAAATGAAG AACTATTGGCGGCAGCTCCTAAATGCGAAACTGATAACGCAAAGAAAGTTCGATAACTTAACTAAAGCTGAGAGGGGTGG CTTGTCTGAACTTGACAAGGCCGGATTTATTAAACGTCAGCTCGTGGAAACCCGCCAAATCACAAAGCATGTTGCACAGA TACTAGATTCCCGAATGAATACGAATACGACGAGAACGATAAGCTGATTCGGGAAGTCAAAGTAATCACTTTAAAGTCA AAATTGGTGTCGGACTTCAGAAAGGATTTTCAATTCTATAAAGTTAGGGAGATAAATAACTACCACCATGCGCACGACGC TTATCTTAATGCCGTCGTAGGGACCGCACTCATTAAGAAATACCCGAAGCTAGAAAGTGAGTTTGTGTATGGTGATTACA AAGTTTATGACGTCCGTAAGATGATCGCGAAAGCGAACAGGAGATAGGCAAGGCTACAGCCAAATACTTCTTTTATTCT AACATTATGAATTTCTTTAAGACGGAAATCACTCTGGCAAACGGAGAGATACGCAAACGACCTTTAATTGAAACCAATGG GGAGACAGGTGAAATCGTATGGGATAAGGGCCGGGACTTCGCGACGGTGAGAAAAGTTTTGTCCATGCCCCAAGTCAACA TAGTAAAGAAAACTGAGGTGCAGACCGGAGGGTTTTCAAAGGAATCGATTCTTCCAAAAAGGAATAGTGATAAGCTCATC GCTCGTAAAAAGGACTGGGACCCGAAAAAGTACGGTGGCTTCGATAGCCCTACAGTTGCCTATTCTGTCCTAGTAGTGGC AAAAGTTGAGAAGGGAAAATCCAAGAAACTGAAGTCAGTCAAAGAATTATTGGGGATAACGATTATGGAGCGCTCGTCTT TTGAAAAGAACCCCATCGACTTCCTTGAGGCGAAAGGTTACAAGGAAGTAAAAAAGGATCTCATAATTAAACTACCAAAG TATAGTCTGTTTGAGTTAGAAAATGGCCGAAAACGGATGTTGGCTAGCGCCGGAGAGCTTCAAAAGGGGAACGAACTCGC ACTACCGTCTAAATACGTGAATTTCCTGTATTTAGCGTCCCATTACGAGAAGTTGAAAGGTTCACCTGAAGATAACGAAC AGAAGCAACTTTTTGTTGAGCAGCACAAACATTATCTCGACGAAATCATAGAGCAAATTTCGGAATTCAGTAAGAGAGTC ATCCTAGCTGATGCCAATCTGGACAAAGTATTAAGCGCATACAACAAGCACAGGGATAAACCCATACGTGAGCAGGCGGA AAATATTATCCATTTGTTTACTCTTACCAACCTCGGCGCTCCAGCCGCATTCAAGTATTTTGACACAACGATAGATCGCA AACGATACACTTCTACCAAGGAGGTGCTAGACGCGACACTGATTCACCAATCCATCACGGGATTATATGAAACTCGGATA GATTTGTCACAGCTTGGGGGTGACGGATCCCCCAAGAAGAAGAGGAAAGTCTCGAGCGACTACAAAGACCATGACGGTGA TTATAAAGATCATGACATCGATTACAAGGATGACGATGACAAGGCTGCAGGA JPEG2025183226000005.jpg113168(Underlined once: HNH domain; Underlined twice: RuvC domain)
[0180] In some embodiments, wild-type Cas9 is expressed in Streptococcus pyrogenes (NCBI). Reference sequence: NC_002737.2 (nucleotide sequence as below); and Unip rot reference sequence: Corresponding to Cas9 from Q99ZW2 (amino acid sequence as shown below) : ATGGATAAGAAATACTCAATGCTTAGATATCGGCACAAATAGCGTCGGATGGGCGGTGATCACTGATGAATATAAGGT TCCGTCTAAAAAGTTCAAGGTTCTGGGAAATACAGACCGCCACAGTATCAAAAAAAATCTTATAGGGGCTCTTTTATTTG ACAGTGGAGAGACAGCGGAAGCGACTCGTCTCAAACGGACAGCTCGTAGAAGGTATACACGTCGGAAGAATCGTATTTGT TATCTACAGGAGATTTTTTCAAATGAGATGGCGAAAGTAGATGATAGTTTCTTTCATCGACTTGAAGAGTCTTTTTTGGT GGAAGAAGACAAGAAGCATGAACGTCATCCTATTTTTGGAAATATAGTAGATGAAGTTGCTTATCATGAGAAATATCCAA CTATCTATCATCTGCGAAAAAAATTGGTAGATTCTACTGATAAAGCGGATTTGCGCTTAATCTATTTGGCCTTAGCGCAT ATGATTAAGTTTCGTGGTCATTTTTTGATTGAGGGAGATTTAAATCCTGATAATAGTGATGTGGACAAACTATTTATCCA GTTGGTACAAACCTACAATCAATTATTTGAAGAAAACCCTATTAACGCAAGTGGAGTAGATGCTAAAGCGATTCTTTCTG CACGATTGAGTAAATCAAGACGATTAGAAAATCTCATTGCTCAGCTCCCCGGTGAGAAGAAAAATGGCTTATTTGGGAAT CTCATTGCTTTGTCATTGGGTTTGACCCCTAATTTTAAATCAAATTTTGATTTGGCAGAAGATGCTAAATTACAGCTTTC AAAAGATACTTACGATGATGATTTAGATAATTTATTGGCGCAAATTGGAGATCAATATGCTGATTTGTTTTTGGCAGCTA AGAATTTATCAGATGCTATTTTACTTTCAGATATCCTAAGAGTAAATACTGAAATAACTAAGGCTCCCCTATCAGCTTCA ATGATTAAACGCTACGATGAACATCATCAAGACTTGACTCTTTTAAAAGCTTTAGTTCGACAACAACTTCCAGAAAAGTA TAAAGAAATCTTTTTTGATCAATCAAAAAACGGATATGCAGGTTATATTGATGGGGGAGCTAGCCAAGAAGAATTTTATA AATTTATCAAACCAATTTTAGAAAAAATGGATGGTACTGAGGAATTATTGGTGAAACTAAATCGTGAAGATTTGCTGCGC AAGCAACGGACCTTTGACAACGGCTCTATTCCCCATCAAATTCACTTGGGTGAGCTGCATGCTATTTTGAGAAGACAAGA AGACTTTTATCCATTTTTAAAAGACAATCGTGAGAAGATTGAAAAAATCTTGACTTTTCGAATTCCTTATTATGTTGGTC CATTGGCGCGTGGCAATAGTCGTTTTGCATGGATGACTCGGAAGTCTGAAGAAACAATTACCCCATGGAATTTTGAAGAA GTTGTCGATAAAGGTGCTTCAGCTCAATCATTTATTGAACGCATGACAAACTTTGATAAAAATCTTCCAAATGAAAAAGT ACTACCAAAACATAGTTTGCTTATGAGTATTTTACGGTTTATAACGAATTGACAAAGGTCAAATATGTTACTGAAGGAA TGCGAAAACCAGCATTTCTTTCAGGTGAACAGAAGAAAGCCATTGTTGATTTACTCTTCAAAACAAATCGAAAAGTAACC GTTAAGCAATTAAAAAGAAGATTATTTCAAAAAAATAGAATGTTTTGATAGTGTTGAAATTTCAGGAGTTGAAGATAGATT TAATGCTTCATTAGGTACCTACCATGATTTGCTAAAAATTATTAAAGATAAAGATTTTTTGGATAATGAAGAAAATGAAG ATATCTTAGAGGATATTGTTTTAACATTGACCTTATTTGAAGATAGGGAGATGATTGAGGAAAGACTTAAAACATATGCT CACCTCTTTGATGATAAGGTGATGAAACAGCTTAAACGTCGCCGTTATACTGGTTGGGGACGTTTGTCTCGAAAATTGAT TAATGGTATTAGGGATAAGCAATCTGGCAAAACAATATTAGATTTTTTGAAATCAGATGGTTTTGCCAATCGCAATTTTA TGCAGCTGATCCATGATGATAGTTTGACATTTAAAGAAGACATTCAAAGGCACAAGTGTCTGGACAAGGCGATAGTTTA CATGAACATATTGCAAATTTAGCTGGTAGCCCTGCTATTAAAAAAGGTATTTTACAGACTGTAAAAGTTGTTGATGAATT GGTCAAAGTAATGGGGCGGCATAAGCCAGAAAATATCGTTATTGAAATGGCACGTGAAAATCAGACAACTCAAAAGGGCC AGAAAAATTCGCGAGAGCGTATGAAACGAATCGAAGAAGGTATCAAAGAATTAGGAAGTCAGATTCTTAAAGAGCATCCT GTTGAAAATACTCAATTGCAAAATGAAAAGCTCTATCTCTATTATCTCCAAAATGGAAGAGACATGTATGTGGACCAAGA ATTAGATATTAATCGTTTAAGTGATTATGATGTCGATCACATTGTTCCACAAAGTTTCCTTAAAGACGATTCAATAGACA ATAAGGTCTTAACGCGTTCTGATAAAAATCGTGGTAAATCGGATAACGTTCCAAGTGAAGAAGTAGTCAAAAAGATGAAA AACTATTGGAGACAACTTCTAAACGCCAAGTTAATCACTCAACGTAAGTTTGATAATTTAACGAAAGCTGAACGTGGAGG TTTGAGTGAACTTGATAAAGCTGGTTTTATCAAACGCCAATTGGTTGAAACTCGCCAAATCACTAAGCATGTGGCACAAA TTTTGGATAGTCGCATGAATACTAAATACGATGAAAATGATAAACTTATTCGAGAGGTTAAAGTGATTACCTTAAAATCT AAATTAGTTTCTGACTTCCGAAAAGATTTCCAATTCTATAAAGTACGTGAGATTAACAATTACCATCATGCCCATGATGC GTATCTAAATGCCGTCGTTGGAACTGCTTTGATTAAGAAATATCCAAAACTTGAATCGGAGTTTGTCTATGGTGATTATA AAGTTTATGATGTTCGTAAAATGATTGCTAAGTCTGAGCAAGAAATAGGCAAAGCAACCGCAAAATATTTCTTTTACTCT AATATCATGAACTTCTTCAAAACAGAAATTACACTTGCAAATGGAGAGATTCGCAAACGCCCTCTAATCGAAACTAATGG GGAAACTGGAGAAATTGTCTGGGATAAAGGGCGAGATTTTGCCACAGTGCGCAAAGTATTGTCCATGCCCCAAGTCAATA TTGTCAAGAAAACAGAAGTACAGACAGGCGGATTCTCCAAGGAGTCAATTTTTACCAAAAGAAATTCGGACAAGCTTATT GCTCGTAAAAAAGACTGGGATCCAAAAAAATATGGTGGTTTTGATAGTCCAACGGTAGCTTATTCAGTCCTAGTGGTTGC TAAGGTGGAAAGGGGAAATCGAAGAAGTTAAAATCCGTTAAAGAGTTACTAGGGATCACAATTATGGAAAGAAGTTCCT TTGAAAAAAATCCGATTGACTTTTTAGAAGCTAAAGGATATAAGGAAGTTAAAAGACTTAATCATTAAACTACCTAAA TATAGTCTTTTTGAGTTAGAAAACGGTCGTAAACGGATGCTGGCTAGTGCCGGAGAATTACAAGGGAAATGAGCTGGC TCTGCCAAGCAAATATGTGAATTTTTTATATTTAGCTAGTCATTATGAAAAGTTGAAGGGTAGTCCAGAAGATAACGAAC AAAAACAATTGTTTGTGGAGCAGCATAAGCATTATTTAGATGAGATTATTGAGCAAATCAGTGAATTTTCTAAGCGTGTT ATTTTAGCAGATGCCAATTTAGATAAAGTTCTTAGTGCATATAACAAACATAGAGACAAACCAATACGTGAACAAGCAGA AAATATTATTCATTTATTTACGTTGACGAATCTTGGAGCTCCCGCTGCTTTTAAATATTTTGATACAACAATTGATCGTA AACGATATACGTCCTACAAAAGAAGTTTTAGATGCCACTCTTATCCATCAATCCATCACTGGTCTTTATGAAACACGCATT GATTTGAGTCAGCTAGGAGGTGACTGA JPEG2025183226000006.jpg115168 (single underline: HNH domain; double underline: RuvC domain)
[0181] In some embodiments, Cas9 is expressed in Corynebacterium ulcerans (NCBI reference Numbers: NC_015683.1, NC_017317.1); Corynebacterium zieff terrier (NCBI reference numbers: NC_016782.1, NC_016786.1); Loplasma silphidicola (NCBI reference number: NC_021284.1); Pre Botella intermedia (NCBI reference number: NC_017861.1); Spiropla Zuma taiwanense (NCBI reference number: NC_021846.1); Streptococcus · inie (NCBI reference number: NC_021314.1); Verriella baltica (NC NCBI Reference Number: NC_018010.1); Cycloflexus torquhis (NCBI Reference Streptococcus thermophilus (NCBI Reference Number: NC_018721.1); Streptococcus thermophilus (NCBI Reference Number: NC_018721.1); Reference number: YP_820832.1), Listeria innocua (NCBI reference number: NP _472073.1), Campylobacter jejuni (NCBI reference number: YP_00 2344900.1), or Neisseria meningitidis (NCBI reference number: Y P_002342100.1) or any other organism. This refers to the upcoming Cas9.
[0182] Additional Cas9 proteins (e.g., Cas9 without nuclease activity (dCas9) , Cas9 nickase (nCas9), or nuclease-active Cas9), It should be appreciated that variants and homologs are included within the scope of the present disclosure. Exemplary Cas9 proteins include, but are not limited to, those shown below. In some embodiments, the Cas9 protein is a Cas9 protein with no nuclease activity ( In some embodiments, the Cas9 protein is Cas9 nickase ( In some embodiments, the Cas9 protein is a C It is as9.
[0183] In some embodiments, the Cas9 domain is a nuclease-inactive Cas9 domain (d For example, the dCas9 domain cleaves both strands of a double-stranded nucleic acid molecule. Some of the nucleic acid molecules may bind to the double-stranded nucleic acid molecule (e.g., via a gRNA molecule) without any need for a specific nucleic acid molecule. In embodiments, the nuclease-inactive dCas9 domain is an amino acid sequence as defined herein. D10X and H840X mutations in the amino acid sequence, or any of the amino acid sequences shown herein. The corresponding mutations in the amino acid sequence are included, where X is any amino acid change. In embodiments, the nuclease-inactive dCas9 domain comprises an amino acid sequence as defined herein. The D10A and H840A mutations in the amino acid sequence, or any of the amino acid sequences shown herein, The corresponding mutations in the amino acid sequence are included. The defined amino acid sequence was inserted into the cloning vector pPlatTET-gRNA2( Accession number BAV54124).
[0184] The amino acid sequence of an exemplary catalytically inactive Cas9 (dCas9) is as follows: :MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRK NRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYL ALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNG LFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAP LSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNRE DLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPW NFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTN RKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERL KTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQ GDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQIL KEHPVENTQLQNEKLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVV KKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVI TLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKY FFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNS DKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLII KLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEF SKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLY ETRIDLSQLGGD (See, e.g., Qi et al., “Repurposing,” the entire contents of which are incorporated herein by reference. ng CRISPR as an RNA-guided platform for sequence-specific control of gene expres sion.” Cell. 2013; 152(5):1173-83).
[0185] In some embodiments, the Cas9 nuclease is an inactive (e.g., inactivated) D The "nCas9" tag has a NA cleavage domain, i.e., Cas9 is a nickase. Also called "nickase" Cas9 protein. Nuclease-inactivating Cas9 protein The protein is interchangeably referred to as "dCas9" protein (nuclease "dead" Cas9) or It is also sometimes called catalytically inactive Cas9. Methods for generating a Cas9 protein (or a fragment thereof) having the following structure are known (e.g., Jinek et al., Science. 337:816-821(2012); Qi et al., “Repurposing CRISPR as an RNA-Guided Platform fo r Sequence-Specific Control of Gene Expression” (2013) Cell. 28;152(5):1173-83 For example, the DNA cleavage domain of Cas9 is divided into two subdomains: H It is known to contain the NH nuclease subdomain and the RuvC1 subdomain. The H subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain Mutations within these subdomains can reduce the nuclease activity of Cas9. For example, mutations D10A and H840A can block sex in S. pyrogenes C. Completely inactivates the nuclease activity of as9 (Jinek et al., Science. 337:816-821(2 012); Qi et al., Cell. 28;152(5):1173-83 (2013)).
[0186] In some embodiments, the dCas9 domain is a dCas9 domain as described herein. At least 60%, at least 65%, at least 70%, or at least At least 75%, at least 80%, at least 85%, at least 90%, at least 95% %, at least 96%, at least 97%, at least 98%, at least 99%, or In some embodiments, Cas9 comprises an amino acid sequence that is at least 99.5% identical. A domain may have one, two, three or more amino acid sequences different from any one of the amino acid sequences defined herein. , 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 3 2, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45 , 46, 47, 48, 49, 50, or more mutations In some embodiments, the Cas9 domain comprises an amino acid sequence having the amino acid sequence as defined herein. At least 10, at least 15, or fewer amino acid sequences than any one of the amino acid sequences At least 20, at least 30, at least 40, at least 50, at least 60, at least At least 70, at least 80, at least 90, at least 100, at least 150, At least 200, at least 250, at least 300, at least 350, at least At least 400, at least 500, at least 600, at least 700, at least 800 , at least 900, at least 1000, at least 1100, or at least 12 It comprises an amino acid sequence having 00 identical contiguous amino acid residues.
[0187] In some embodiments, dCas9 partially or completely inhibits Cas9 nuclease activity. corresponding to a Cas9 amino acid sequence containing one or more mutations that inactivate the For example, in some embodiments, the dCas9 domain comprises D10A and H 840A mutation, or the corresponding mutation in another Cas9.
[0188] In some embodiments, dCas9 is an amino acid sequence of dCas9(D10A and H840A). The amino acid sequence includes: JPEG2025183226000007.jpg116168 (single underline: HNH domain; double underline: RuvC domain)
[0189] In some embodiments, the Cas9 domain comprises a D10A mutation and a CDR at residue 840. The group may be present in the amino acid sequences shown above or in the amino acid sequences shown herein. At the corresponding positions of either of them, histidine remains.
[0190] In other embodiments, dCas9 variants with mutations other than D10A and H840A are Examples of methods are provided that result in nuclease-inactivated Cas9 (dCas9). Such mutations include, for example, other amino acid substitutions at D10 and H840, or other substitutions within the nuclease domain of Cas9 (e.g., HNH nuclease subunits). In some embodiments, substitutions within the RuvC1 domain and / or RuvC1 subdomain are included. At least about 70% identical, at least about 80% identical, at least about 90% identical At least about 95% identical, at least about 98% identical, at least about 99% identical to , at least about 99.5% identical, or at least about 99.9% identical dCas9 variants In some embodiments, variants or homologs of the nucleotides are provided. about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids about 50 amino acids, about 75 amino acids, about 100 amino acids, or shorter Alternatively, variants of dCas9 having longer amino acid sequences are provided.
[0191] In some embodiments, the Cas9 domain is a Cas9 nickase. The enzyme cleaves only one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase may be a Cas9 protein capable of The target strand of the double-stranded nucleic acid molecule is cleaved, which allows the gRNA (e.g., The Cas9 nickase cleaves the strand that base pairs with (or is complementary to) the target RNA (e.g., sgRNA). In some embodiments, the Cas9 nickase comprises a D10A mutation, In some embodiments, the Cas9 nickase is capable of amplifying the double-stranded nucleic acid. The acid molecule cleaves the non-target strand of the gene that is not edited, which allows it to bind to Cas9. Cas9 nickase cleaves the strand that does not base-pair with the gRNA (e.g., sgRNA) that it is targeting. In some embodiments, the Cas9 nickase comprises an H840A mutation. and has an aspartic acid residue at position 10 or a corresponding mutation. In some embodiments, the Cas9 nickase can be any one of the Cas9 nickases provided herein. At least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 9 9.5% identical amino acid sequence. Additional suitable Cas9 nickases are described in this disclosure and This would be apparent to a person skilled in the art based on knowledge in the art and is within the scope of this disclosure.
[0192] The amino acid sequence of an exemplary catalytic Cas9 (nCas9) is as follows: MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRIC YLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAH MIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGN LIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSAS MIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLR KQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEE VVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYA HLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSL HEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIGELGSQILKEHP VENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMK NYWRQLLNAKLITQRKFDNLTKAEGGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKS KLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYS NIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLI ARKKDWDPKKYGGFDSPTVAYSVLVVAKVEGKKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPK YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRV ILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRI DLSQLGGD
[0193] In some embodiments, Cas9 is an archaic organism that comprises the domain and kingdom of unicellular prokaryotic microorganisms. It refers to a Cas9 derived from bacteria (e.g., nanoarchaea). In some embodiments, a programmable The nucleotide-binding protein may be a CasX or CasY protein, See, for example, Burstein et al., "New CR ISPR-Cas systems from uncultivated microbes." Cell Res. 2017 Feb 21. doi: 10.103 8 / cr.2017.21. Using genome sequencing metagenomics, multiple CRISPR The first known -Cas system has been identified in the archaeal domain of life. The first reported Cas9 protein is Cas9. This diverse Cas9 protein has been little studied. It was found as part of an active CRISPR-Cas system in uncharted nanoarchaea. In bacteria, two previously unknown systems, CRISPR-CasX and CRIS The PR-CasY system is the most compact system discovered to date. In some embodiments, in the base editor systems described herein, Ca s9 is replaced by CasX or a variant of CasX. In the base editor system described herein, Cas9 is CasY or C It is replaced by variants of asY. Other RNA-guided DNA-binding proteins is used as a nucleic acid programmable DNA binding protein (napDNAbp) and should be understood to be within the scope of this disclosure.
[0194] In some embodiments, the nucleic acid programmable fragment of any of the fusion proteins provided herein The napDNAbp protein binds to the CasX or CasY protein. In some embodiments, the napDNAbp is a CasX protein. In some embodiments, the napDNAbp is a CasY protein. napDNAbp is at least 85% identical to the native CasX or CasY protein. at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, Contains an amino acid sequence that is at least 99%, or comfortably 99.5%, identical. In this state, the programmable nucleotide-binding protein binds to the natural CasX or CasY In some embodiments, the programmable nucleotide binding protein is , at least 85% with any CasX or CasY protein described herein; At least 90%, at least 91%, at least 92%, at least 93%, at least At least 94%, at least 95%, at least 96%, at least 97%, at least 98% , containing amino acid sequences that are at least 99%, or comfortably 99.5%, identical to other bacterial species. It should be appreciated that derived CasX and CasY may also be used in accordance with the present disclosure.
[0195] Exemplary CasX((uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) tr| F0NN87|F0NN87_SULIH CRISPR-associated CasX protein OS = Sulfolobus The amino acid sequence of B. islandicus (HVE10 / 4 strain) GN = SiH_0402 PE=4 SV=1) is as follows: As shown below: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAV GQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLN SDDGKVRDLKLISAYVNGELIRGEG.
[0196] Exemplary CasX(>tr|F0NH53|F0NH53_SULIR CRISPR-associated proteins, Cas X OS = Sulfolobus islandicus (strain REY15A) GN = SiRe_0771 PE = 4 SV =1) has the following amino acid sequence: MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIIL PLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAA AGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAV GQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLN SDDGKVRDLKLISAYVNGELIRGEG.
[0197] Deltaproteobacteria CasX MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAIL QVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVA EHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFL SKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDfAYNEVIARVRMWVNLNLWQKLKLSRDDA KPLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENP KKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEM DEKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFT DGTDIKKSGKWQGLLYGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKI GRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQ AAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAK LAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKEL SAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSY KSGKQPFVGAWQAFYKRRLKEVWKPNA
[0198] An example of CasY ((ncbi.nlm.nih.gov / protein / APG80656.1) >APG80656.1 CRIS The amino acid sequence of the PR-related protein CasY (from an uncultured Percival bacterium) is as follows: As shown below: MSKRHPRISGVKGYRLHAQRLEYTGKSGAMRTIKYPLYSSPSGGRTVPREIVSAINDDYVGLYGLSNFDDLYNAEKRNEE KVYSVLDFWYDCVQYGAVFSYTAPGLLKNVAEVRGGSYELTKTLKGSHLYDELQIDKVIKFLNKKEISRANGSLDKLKKD IIDCFKAEYRERHKDQCNKLADDIKNAKKDAGASLGERQKKLFRDFFGISEQSENDKPSFTNPLNLTCCLLPFDTVNNNR NRGEVLFNKLKEYAQKLDKNEGSLEMWEYIGIGNSGTAFSNFLGEGFLGRLRENKITELKKAMMDITDAWRGQEQEEELE KRLRILAALTIKLREPKFDNHWGGYRSDINGKLSSWLQNYINQTVKIKEDLKGHKKDLKKAKEMINRFGESDTKEEAVVS SLLESIEKIVPDDSADDEKPDIPAIAIYRRFLSDGRLTLNRFVQREDVQEALIKERLEAEKKKKPKKRKKKSDAEDEKET IDFKELFPHLAKPLKLVPNFYGDSKRELYKKYKNAAIYTDALWKAVEKIYKSAFSSSLKNSFFDTDFDKDFFIKRLQKIF SVYRRFNTDKWKPIVKNSFAPYCDIVSLAENEVLYKPKQSRSRKSAAIDKNRVRLPSTENIAKAGIALARELSVAGFDWK DLLKKEEHEEYIDLIELHKTALALLLAVTETQLDISALDFVENGTVKDFMKTRDGNLVLEGRFLEMFSQSIVFSELRGLA GLMSRKEFITRSAIQTMNGKQAELLYIPHEFQSAKITTPKEMSRAFLDLAPAEFATSLEPESLSEKSLLKLKQMRYYPHY FGYELTRTGQGIDGGVAENALRLEKSPVKKREIKCKQYKTLGRGQNKIVLYVRSSYYQTQFLEWFLHRPKNVQTDVAVSG SFLIDEKKVKTRWNYDALTVALEPVSGSERVFVSQPFTIFPEKSAEEEGQRYLGIDIGEYGIAYTALEITGDSAKILDQN FISDPQLKTLREEVKGLKLDQRRGTFAMPSTKIARIRESLVHSLRNRIHHLALKHKAKIVYELEVSRFEEGKQKIKKVYA TLKKADVYSEIDADKNLQTTVWGKLAVASEISASYTSQFCGACKKLWRAEMQVDETITTQELIGTVRVIKGGTLIDAIKD FMRPPIFDENDTPFPKYRDFCDKHHISKKMRGNSCLFICPFCRANADADIQASQTIALLRYVKEEKKVEDYFERFRKLKN IKVLGQMKKI.
[0199] In some embodiments, a nucleic acid programmable DNA binding protein (napDNAbp) is the single effector of the microbial CRISPR-Cas system. Single effectors of the R-Cas system include, but are not limited to, Cas9, Cpf1, Cas12b / C2c1, and Cas12c / C2c3. Biological CRISPR-Cas systems are divided into class 1 and class 2 systems. Class 1 systems have multisubunit effector complexes, whereas class 2 systems have multisubunit effector complexes. The Cas2 system has a single protein effector, e.g., Cas9 and Cp f1 is a class 2 effector. In addition to Cas9 and Cpf1, three distinct Class 2 CRISPR-Cas systems (Cas12b / C2c1 and Cas12 c / C2c3), but Shmakov et al., “Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems”, Mol. Cell, 2015 Nov. 5; 60(3): 385-397 The entire contents of which are incorporated herein by reference. The effectors of C2c1 and C2c3 are expressed by Cp Contains a RuvC-like endonuclease domain related to f1. The system contains an effector with two putative HEPN RNase domains. Mature CRISPR RNA production is achieved by Cas12b / C2c1-mediated CRISPR RNA transcription. Unlike Cas12b / C2c1, it is tracrRNA-dependent. depends on both CRISPR RNA and tracrRNA.
[0200] Alicyclobacillus acidoterrestris Cas12b / C2c1 (AacC2c1) The crystal structure of α-glucan has been reported in complex with a chimeric single-stranded guide RNA (sgRNA). See, for example, Liu et al., "C2c1 -sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism”, Mol. Cell, See 2017 Jan. 19; 65(2):310-322. The crystal structure shows a tertiary complex with the target DNA. reported in Alicyclobacillus acidoterrestris C2c1 bound as a soluble form. See, for example, Yang et al., "PAM-d" (Parametric Assessment Method for Multi-Objective Learning), the entire contents of which are incorporated herein by reference. ependent Target DNA Recognition and Cleavage by C2C1 CRISPR-Cas endonuclease”, See Cell, 2016 Dec. 15; 167(7):1814-1828. Both targeted and non-targeted DNA The catalytically competent conformation of AacC2c1 with the A chain resides within a single RuvC catalytic pocket. are independently located and captured, and targeted as a result of Cas12b / C2c1-mediated cleavage. This results in a staggered cut of seven nucleotides in the target DNA. Structural comparison of this enzyme with previously identified Cas9 and Cpf1 counterparts revealed that C The diversity of mechanisms used by the RISPR-Cas9 system has been demonstrated.
[0201] In some embodiments, the nucleic acid programmable fragment of any of the fusion proteins provided herein The napDNAbp protein binds to Cas12b / C2c1 or Cas1 In some embodiments, napDNAbp may be a C12c / C2c3 protein. In some embodiments, the napDNAbp is a Ca In some embodiments, the napDNAbp is a naturally occurring s12c / C2c3 protein. Cas12b / C2c1 or Cas12c / C2c3 protein and at least 85% , at least 90%, at least 91%, at least 92%, at least 93%, at least At least 94%, at least 95%, at least 96%, at least 97%, at least 98% %, at least 99%, or comfortably 99.5% identical amino acid sequence. In embodiments, napDNAbp is a naturally occurring Cas12b / C2c1 or Cas12c / In some embodiments, the napDNAbp is a C2c3 protein. At least 85%, at least 90%, or at least at least 91%, at least 92%, at least 93%, at least 94%, at least 9 5%, at least 96%, at least 97%, at least 98%, at least 99%, or or easily 99.5% identical amino acid sequence to Cas12b / C from other bacterial species. It should be appreciated that Cas12c / C2c3 or Cas12c / C2c1 may also be used in accordance with the present disclosure. be.
[0202] Cas12b / C2c1((uniprot.org / uniprot / T0D7A2 #2)sp|T0D7A2|C2C1_ALIAG CRISPR-associated endonuclease ZeC2c1 OS = Alicyclobacillus acidoterrestris (ATCC49025 / D SM3922 / CIP106132 / NCIMB13137 / GD3B strain) GN=c2 The amino acid sequence of c1 (PE=1 SV=1) is as follows: MAVKSIKVKLRLDDMPEIRAGLWKLHKEVNAGVRYYTEWLSLLRQENLYRRSPNGDGEQECDKTAEECKAELLERLRARQ VENGHRGPAGSDDELLQLARQLYELLVPQAIGAKGDAQQIARKFLSPLADKDAVGGLGIAKAAGNKPRWVRMREAGEPGWE EEKEKAETRKSADRTADVLRALADFGLKPLMRVYTDSEMSSVEWKPLRKGQAVRTWDRDMFQQAIERMMSWESWNQRVGQ EYAKLVEQKNRFEQKNFVGQEHLVHLVNQLQQDMKEASPGLESKEQTAHYVTGRALRGSDKVFEKWGKLAPDAPFDLYDA EIKNVQRRNTRRFGSHDLFAKLAEPEYQALWREDASFLTRYAVYNSILRKLNHAKMFATTFTLPDATAHPIWTRFDKLGGN LHQYTFLFNEFGERRHAIRFHKLLKVENGVAREVDDVTVPISMSEQLDNLLPRDPNEPIALYFRDYGAEQHFTGEFGGAK IQCRRDQLAHMHRRRGARDVYLNVSVRVQSQSEARGERRPPYAAVFRLVGDNHRAFVHFDKLSDYLAEHPDDGKLGSEGL LSGLRVMSVDLLGLRTSASISVFRVARKDELKPNSKGRVPFFFPIKGNDNLVAVHERSQLLKLPGETESKDLRAIREERQR TLRQLRTQLAYLRLLVRCGSEDVGRRERSWAKLIEQPVDAANHMTPDWREAFENELQKLKSLHGICSDKEWMDAVYESVR RVWRHMGKQVRDWRKDVRSGERPKIRGYAKDVVGGNSIEQIEYLERQYKFLKSWSFFGKVSGQVIRAEKGSRFAITLREH IDHAKEDRLKKLADRIIMEALGYVYALDERGKGKWVAKYPPCQLILLEELSEYQFNNDRPPSENNQLMQWSHRGVFQELI NQAQVHDLLVGTMYAAFSSRFDARTGAPGIRCRRVPARCTQEHNPEPFPWWLNKFVVEHTLDACPLRADDLIPTGEGEIF VSPFSAEGDFHQIHADLNAAQNLQQRLWSDFDISQIRLRCDWGEVDGELVLIPRLTGKRTADSYSNKVFYTNTGVTYYE RERGKKRRKVFAQEKLSEEEAELLVEADEAREKSVVLMRDPSGIINRGNWTRQKEFWSMV NQRIEGYLVKQIRSRVPLQ DSACENTGDI.
[0203] BhCas12b (Bacillus hisashii) NCBI reference sequence: WP_09514251 5 MAPKKKRKVGIHGVPAAATRSFILKIEPNEEVKKGLWKTHEVLNHGIAYYMNILKLIRQEAIYEHHEQDPKNPKKVSKAE IQAELWDFVLKMQKCNSFTHEVDKDEVFNILRELYEELVPSSVEKKGEANQLSNKFLYPLVDPNSQSGKGTASSGRKPRW YNLKIAGDPSWEEEKKKWEEDKKKDPLAKILGKLAEYGLIPLFIPYTDSNEPIVKEIKWMEKSRNQSVRRLDKDMFIQAL ERFLSWESWNLKVKEEYEKVEKEYKTLEERIKEDIQALKALEQYEKERQEQLLRDTLNTNEYRLSKRGLRGWREIIQKWL KMDENEPSEKYLEVFKDYQRKHPREAGDYSVYEFLSKKENHFIWRNHPEYPYLYATFCEIDKKKKDAKQQATFTLADPIN HPLWVRFEERSGSNLNKYRILTEQLHTEKLKKKLTVQLDRLIYPTESGGWEKGKVDIVLLPSRQFYNQIFLDIEEEKGKH AFTYKDESIKFPLKTGTLGGARVQFDRDHLRRYPHKVESGNVGRIYFNMTVNIEPTESPVSKSLKIHRDDFPKVVNFKPKE LTEWIKDSKGKKLKSGIESLEIGLRVMSIDLGQRQAAAASIFEVVDQKPDIEGKLFFPIKGTELYAVHRASFNIKLPGET LVKSREVLRKAREDNLKLMNQKLNFLRNVLHFQQFEDITEREKRVTKWISRQENSDVPLVYQDELIQIRELMYKPYKDWV AFLKQLHKRLEVEIGKEVKHWRKSLSDGRKGLYGISLKNIDEIDRTRKFLLRWSLRPTEPGEVRRLEPGQRFAIDQLNHL NALKEDRLKKMANTIIMHALGYCYDVRKKKWQAKNPACQIILFEDLSNYNPYEERSRFENSKLMKWSRREIPRQVALQGE IYGLQVGEVGAQFSSRFHAKTGSPGIRCSVVTKEKLQDNRFFKNLQREGRLTLDKIAVLKEGDLYPDKGGEKFISLSKDR KCVTTHADINAAQNLQKRFWTRTHGFYKVYCKAYQVDGQTVYIPESKDQKQKIIEEFGEGYFILKDGVYEWVNAGKLKIK KGSSKQSSSELVDSDILKDSFDLASELKGEKLMLYRDPSGNVFPSDKWMAAGVFFGKLERILISKLTNQYSISTIEDDSS KQSMKRPAATKKAGQAKKKK
[0204] In some embodiments, the Cas12b is BvCas12B, and the BvCas12B is B A variant of hCas12b, containing the following changes relative to BhCas12B: S8 93R, K846R, and E837G. BvCas12b (Bacillus V3-13) NCBI reference sequence: WP_10166145 1.1 MAIRSIKLKMKTNSGTDSIYLRKALWRTHQLINEGIAYYMNLLTLYRQEAIGDKTKEAYQAELINIIRNQQRNNGSSEEH GSDQEILALLRQLYELIIPSSIGESGDANQLGNKFLYPLVDPNSQSGKGTSNAGRKPRWKRLKEEGNPDWELEKKKDEER KAKDPTVKIFDNLNKYGLLPLFPLFTNIQKDIEWLPLGKRQSVRKWDKDMFIQAIERLLSWESWNRRVADEYKQLKEKTE SYYKEHLTGGEEWIEKIRKFEKERNMELEKNAFAPNDGYFITSRQIRGWDRVYEKWSKLPESAPEELWKVVAEQQNKMS EGFGDPKVFSFLANRENRDIWRGHSERIYHIAAYNGLQKKLSRTKEQATFTLPDAIEHPLWIRYESPGGTNLNLFKLEEK QKKNYYVTLSKIIWPSEEKWIEKENIEIPLAPSIQFNRQIKLKQHVKGKQEISFSDYSSRISLDGVLGGSRIQFNRKYIK NHKELLGEGDIGPVFFNLVVDVAPLQETRNGRLQSPIGKALKVISSDFSKVIDYKPKELMDWMNTGSASNSFGVASLLEG MRVMSIDMGQRTSASVSIFEVVKELPKDQEQKLFYSINDTELFAIHKRSFLLNLPGEVVTKNNKQQRQERRKKRQFVRSQ IRMLANVLRLETKKTPDERKKAIHKLMEIVQSYDSWTASQKEVWEKELNLLTNMAAFNDEIWKESLVELHHRIEPYVGQI VSKWRKGLSEGRKNLAGISMWNIDELEDTRRLLISWSKRSRTPGEANRIETDEPFGSSLLQHIQNVKDDRLKQMANLIIM TALGFKYDKEEKDRYKRWKETYPACQIILFENLNRYLFNLDRSRRRENSRLMKWAHRSIPRTVSMQGEMFGLQVGDVRSEY SSRFHAKTGAPGIRCHALTEEDLKAGSNTLKRLIEDGFINESELAYLKKGDIIPSQGGELFVTLSKRYKKDSDNNELTVI HADINAAQNLQKRFWQQNSEVYRVPCQLARMGEDKLYIPKSQTETIKKYFGKGSFVKNNTEQEVYKWEKSEKMKIKTDTT FDLQDLDGFEDISKTIELAQEQQKKYLTMFRDPSGYFFNNETWRPQKEYWSIVNNIIKSCLKKKILSNKVEL
[0205] Cas9 nuclease contains two functional endonuclease domains: RuvC and Cas9 undergoes a conformational change upon target binding, and this conformational change Positioning the nuclease domain to cut opposite strands of DNA. Cas9-mediated DNA The net result of NA cleavage is a nucleotide sequence within the target DNA (approximately 3–4 nucleotides upstream of the PAM sequence). The resulting double-strand break (DSB) then splits into two general Repair pathways: (1) the efficient but error-prone non-homologous end joining (NHEJ) pathway; (2) by one of the less efficient but more fidelity homology-directed repair (HDR) pathways It will be repaired.
[0206] "Efficiency" of non-homologous end joining (NHEJ) and / or homology-directed repair (HDR) can be calculated by any convenient method. For example, in some cases the efficiency is , can be expressed in terms of the percentage of successful HDR. For example, A cleavage assay can be used to generate cleavage products, and the ratio of the products to the substrate can be Ratios can be used to calculate percentages. For example, surveyor nuclei The enzyme can be used, but this enzyme is newly incorporated as a result of successful HDR. The amount of cleaved substrate is determined by the HDR indicates a higher percentage of HDR (higher HDR efficiency). The percentage of HDR is calculated using the following formula: [(cleavage product) / (substrate + cleavage product)] ] (e.g., (b+c) / (a+b+c), where , "a" is the band intensity of the DNA substrate, and "b" and "c" are the cleavage products).
[0207] In some cases, efficiency may be expressed in terms of the percentage of successful NHEJ. For example, a T7 nuclease assay can be used to generate cleavage products. The ratio of product to substrate can be used to calculate the percentage NHEJ. 7 Endonuclease I acts by hybridizing wild-type DNA strands with mutant DNA strands. NHEJ is a process that breaks mismatched heteroduplex DNA resulting from small random fragments. (This creates insertions or deletions (indels) at the site of the original break.) The number of breaks is a function of the NHEJ. Higher percentages (higher NHEJ efficiency) are indicated. The percentage of NHEJ is calculated using the following formula: (1-(1-(b+c) / (a+b +c)) 1 / 2 ) × 100, where "a" is the balance of the DNA substrate. "b" and "c" are the cleavage products (Ran et al., Cell. 2013 Sep. 12; 154(6):1380-9; and Ran et al., Nat Protoc. 2013 Nov.; 8(11): 2281-2308).
[0208] The NHEJ repair pathway is the most active repair mechanism and frequently results in small nucleotide insertions. or deletions (indels) at the DSB site. This has important practical implications because Cas9 and gRNA or guide polypeptides A population of cells expressing nucleotides can result in a range of different mutations. In most cases, NHEJ generates small indels in the target DNA, resulting in , amino acid deletions, insertions, or mutations within the open reading frame (ORF) of the target gene This results in a frameshift mutation that results in a premature stop codon. The ideal end result is It is a loss-of-function mutation within the target gene.
[0209] NHEJ-mediated DSB repair often results in the loss of the open reading frame of a gene. Homology-directed repair (HDR) repairs single nucleotide changes to generate fluorescent fragments. to generate specific nucleotide changes, up to large insertions such as addition of nucleotides or tags. To utilize HDR for gene editing, DNA repair templates are introduced into cells of interest using gRNA and Cas9 or Cas9 nickase. The repair template can be delivered into the cell type containing the desired editing sequence as well as additional homologous sequences. The target sequence can be contained immediately upstream and downstream of the target (left & right homology arms). The length of each homology arm may depend on the size of the change being introduced. Larger insertions require longer homology arms. The repair template is a single-stranded oligonucleotide. It can be a double-stranded oligonucleotide, a double-stranded DNA plasmid, or a double-stranded DNA plasmid. The efficiency of DR is generally low (<10% for modified alleles), and Cas9, gRNA, HDR is low even in cells expressing exogenous repair templates. Because it occurs during G2, the efficiency of HDR can be increased by synchronizing cells. Chemical or genetic inhibition of genes involved in NHEJ can also enhance HDR. The efficiency of the
[0210] In some embodiments, the Cas9 is a modified Cas9. A given gRNA target sequence is There may be additional sites of partial homology throughout the genome. These are called off-targets and must be taken into consideration when designing gRNAs. In addition to optimizing the design of CRISPR, the specificity of CRISPR was also increased through modifications to Cas9. Cas9 contains two nuclease domains, RuvC and HNH. The D10A mutation of SpCas9 generates double-strand breaks (DSBs) through the combination of these activities. The variant Cas9 nickase retains one nuclease domain and is capable of repairing DSBs rather than DSBs. This nickase system generates DNA nicks, allowing for specific gene editing. It can also be combined with R-mediated gene editing.
[0211] In some cases, the Cas9 is a variant Cas9 protein. 9 polypeptide has one amino acid sequence different from that of the wild-type Cas9 protein. have different amino acid sequences (e.g., have deletions, insertions, substitutions, fusions). The variant Cas9 polypeptide reduces the nuclease activity of the Cas9 polypeptide. have amino acid changes (e.g., deletions, insertions, or substitutions) that reduce the In this study, the variant Cas9 polypeptides were synthesized from the corresponding wild-type Cas9 proteins. Less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5% of the enzyme activity In some cases, the variant Cas9 protein has substantially The target Cas9 protein has no substantial nuclease activity. If the variant Cas9 protein has no activity, it is called "dCas9." This sometimes happens.
[0212] In some cases, the variant Cas9 proteins have reduced nuclease activity. For example, a variant Cas9 protein can be a wild-type Cas9 protein, e.g., a wild-type Less than about 20%, less than about 15%, or about 10% of the endonuclease activity of the Cas9 protein less than about 5%, less than about 1%, or less than about 0.1%.
[0213] In some cases, the variant Cas9 protein cleaves the complementary strand of the guide target sequence. can cleave the non-complementary strand of the double-stranded guide target sequence, but with a reduced ability to cleave the non-complementary strand of the double-stranded guide target sequence. For example, variant Cas9 proteins contain mutations that reduce the function of the RuvC domain ( As a non-limiting example, in some embodiments, the variant Cas 9 protein has D10A (aspartic acid to alanine at amino acid position 10), and Therefore, it is possible to cleave the complementary strand of the double-stranded guide target sequence, but The variant Cas9 protein has a reduced ability to cleave the non-complementary strand of the sequence (thus When proteins cleave double-stranded target nucleic acids, they create single-strand breaks (SS) instead of double-strand breaks (DSB). B) (see, for example, Jinek et al., Science. 2012 Aug. 17; 337(6096):816-21) reference).
[0214] In some cases, the variant Cas9 protein binds to a non-complementary strand of the double-stranded guide target sequence. is capable of cleaving the complementary strand of the guide target sequence, but has a reduced ability to cleave the complementary strand of the guide target sequence. For example, variant Cas9 proteins may contain the HNH domain (RuvC / HNH / Ruv The C domain motif may have a mutation (amino acid substitution) that reduces the function of the C domain motif. By way of example, in some embodiments, the variant Cas9 protein is H840A (amino A mutation (histidine to alanine at position 840) is present, and therefore non-complementary to the guide target sequence. capable of cleaving the complementary strand of the guide target sequence, but has a reduced ability to cleave the complementary strand of the guide target sequence (As a result, when the variant Cas9 protein cleaves the double-stranded guide target sequence, Such a Cas9 protein generates a guide target sequence ( For example, it has a reduced ability to cleave a single-stranded guide target sequence, but not a guide target sequence ( For example, the single-stranded guide retains the ability to bind to the target sequence.
[0215] In some cases, the variant Cas9 protein binds to the complementary strand of the double-stranded target DNA and As a non-limiting example, in some cases The variant Cas9 protein carried both the D10A and H840A mutations. The polypeptide therefore cleaves both the complementary and non-complementary strands of double-stranded target DNA. Such Cas9 proteins have a reduced ability to target DNA (e.g., single-stranded DNA). a target DNA (e.g., single-stranded target DNA), but has reduced ability to cleave the target DNA (e.g., single-stranded target DNA) NA).
[0216] As another non-limiting example, in some cases, the variant Cas9 protein is 6A and W1126A mutations, and therefore the polypeptide is capable of cleaving the target DNA. Such Cas9 proteins have a reduced ability to target DNA (e.g., single-stranded DNA). a target DNA (e.g., single-stranded target DNA), but has reduced ability to cleave the target DNA (e.g., single-stranded target DNA) NA).
[0217] As another non-limiting example, in some cases, the variant Cas9 protein may be P47 5A, W476A, N477A, D1125A, W1126A, and D1127A mutations The polypeptide therefore has a reduced ability to cleave the target DNA. Cas9 proteins such as those shown in Figure 1 have the ability to cleave target DNA (e.g., single-stranded target DNA). Reduced force but retains the ability to bind to target DNA (e.g., single-stranded target DNA) do.
[0218] As another non-limiting example, in some cases, the variant Cas9 protein is H84 0A, W476A, and W1126A mutations, and therefore the polypeptide Such Cas9 proteins have a reduced ability to cleave target DNA. (e.g., single-stranded target DNA), but have reduced ability to cleave target DNA (e.g., As another non-limiting example, in some cases In this study, the variant Cas9 proteins were H840A, D10A, W476A, and W Carrying the 1126A mutation, the polypeptide therefore has a reduced ability to cleave target DNA Such Cas9 proteins are capable of targeting target DNA (e.g., single-stranded target DNA). ) but have reduced ability to cleave target DNA (e.g., single-stranded target DNA). In some embodiments, the variant Cas9 retains the ability to The catalytically active His residue was restored at main position 840 (A840H).
[0219] As another non-limiting example, in some cases, the variant Cas9 protein is H84 0A, P475A, W476A, N477A, D1125A, W1126A, and D1 127A mutation, and therefore the polypeptide has a reduced ability to cleave target DNA. Such Cas9 proteins are capable of cleaving target DNA (e.g., single-stranded target DNA). The ability of the target DNA to cleave the target DNA (e.g., single-stranded target DNA) is reduced. As another non-limiting example, in some cases, variant Cas9 proteins retain the ability to Proteins: D10A, H840A, P475A, W476A, N477A, D1125A , W1126A, and D1127A mutations, and therefore the polypeptide is Such Cas9 proteins have a reduced ability to cleave target DNA ( For example, they have a reduced ability to cleave single-stranded target DNA, but In some cases, variant Cas9 proteins retain the ability to bind to single-stranded target DNA. If the protein carries the W476A and W1126A mutations, or if it is a variant Cas 9 proteins were P475A, W476A, N477A, D1125A, W1126A, and When carrying the D1127A and D1127A mutations, the variant Cas9 protein efficiently transduced P It does not bind to AM sequences. Therefore, in some cases, such variants C When the as9 protein is used in the conjugation method, the method does not require a PAM sequence. In other words, in some cases, such variant Cas9 proteins may be used in a variety of binding methods. When used in a method, the method may include a guide RNA, but the method may be used in the absence of a PAM sequence. (and therefore the specificity of binding can be determined by the targeting site of the guide RNA. Other residues can be mutated to achieve the above effect (i.e. (i.e., inactivating one or the other nuclease moiety). and residues D10, G12, G17, E762, H840, N854, N863, and H98 2. Modify (i.e., replace) H983, A984, D986, and / or A987. Mutations other than alanine substitutions are also suitable.
[0220] In some embodiments, variant Cas9 proteins with reduced catalytic activity (e.g., Cas9 protein was found at D10, G12, G17, E762, H840, N854, and N86. 3, H982, H983, A984, D986, and / or A987 mutations, e.g., D10A, G12A, G17A, E762A, H840A, N854A, N863A, H If you have a 982A, H983A, A984A, and / or D986A, this The modified Cas9 protein can be used at the site as long as it retains the ability to interact with the guide RNA. It can still bind to the target DNA in a specific manner (because it (This is because the target DNA sequence is guided by the RNA.)
[0221] In some embodiments, the variant Cas protein is spCas9, spCas9- VRQR, spCas9-VRER, xCas9(sp), saCas9, saCas9 -KKH, SpCas9-MQKFRAER, spCas9-MQKSER, spCas 9-LRKIQK, or spCas9-LRVSQL.
[0222] In one specific embodiment, the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R(Sp Cas9-MQKFRAER) and specific for the modified PAM5'-NGC-3' A modified SpCas9 with the desired activity is used.
[0223] As an alternative to S. pyrogenes Cas9, Cpf1 fragments exhibit cleavage activity in mammalian cells. Examples of RNA-guided endonucleases include those derived from the Prevotella and and Francisella 1-derived CRISPR (CRISPR / Cpf1) It is a DNA editing technology similar to the as9 system. Cpf1 is a class II CRISPR system. The R / Cas system is an RNA-guided endonuclease. This adaptive immune mechanism The Cpf1 gene was found in the bacteria Prevotella and Francisella. It is associated with a genetic locus that uses guide RNA to look at viral DNA. Cpf1 encodes an endonuclease that extrudes and cleaves the target protein. This is a fast and simple endonuclease, which is one of the limitations of the CRISPR / Cas9 system. Unlike Cas9 nuclease, Cpf1-mediated DNA cleavage results in short The staggered cleavage pattern of Cpf1 is similar to that of the traditional This opens up the possibility of directed gene transfer similar to restriction enzyme cloning. This may increase the efficiency of gene editing. Thus, Cpf1 significantly reduces the number of sites that can be targeted by CRISPR and SpCas9. In AT-rich regions or AT-rich genomes that lack preferred NGG PAM sites, The Cpf1 locus has a mixed alpha / beta domain and R UVC-I and the following helical region, RuvC-II, and zinc finger-like The Cpf1 protein contains a domain similar to the RuvC domain of Cas9. Cpf1 has a RuvC-like endonuclease domain. The N-terminus of Cpf1 does not have a cleavage domain, and the alpha-helix recognition domain of Cas9 is The structure of the Cpf1 CRISPR-Cas domain is similar to that of the Cpf1 CRISPR-Cas domain. We demonstrate that this system is functionally unique and is classified as a Class 2 V-type CRISPR system. The Cpf1 locus is more similar to type I and III Cas systems than to type II systems. It encodes the 1, Cas2, and Cas4 proteins. Functional Cpf1 is expressed in trans It does not require activating CRISPR RNA (tracrRNA), hence CRISPR Only R(crRNA) is required, which is a boon for genome editing. This is because Cpf1 is not only smaller than Cas9, but also a smaller sgRN. This is because it contains the A molecule (almost half the number of nucleotides of Cas9). The NA complex is a protospacer complex, in contrast to the G-rich PAM targeted by Cas9. - Target DNA or RNA by identifying the adjacent motif 5'-YTN-3' After identifying the PAM, Cpf1 cleaves the 4- or 5-nucleotide fragment. It introduces sticky end-like DNA double-strand breaks at the proximal end.
[0224] Some embodiments of the present disclosure relate to nucleic acid programmable DNA binding protein domains and deamina Some embodiments of the present disclosure provide nucleic acid programmable DNA binding proteins. and a fusion protein comprising a domain that acts as a protein, the domain being a salt. Proteins such as gene editors are directed to specific nucleic acid (e.g., DNA or RNA) sequences. In a specific embodiment, the fusion protein can be used to encode a nucleic acid programmer. It contains a soluble DNA-binding protein domain and a deaminase domain. Proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), Ca s12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12 i.e., a programmable polynucleoside with PAM specificity different from that of Cas9. An example of a tide-binding protein forms a cluster from Prevotella and Francisella 1 It is a regularly interspaced short palindromic repeat (Cpf1) that encodes a nucleotide sequence similar to Cas9. Cpf1 is also a class 2 CRISPR effector. It has been shown to mediate robust DNA interference through different features. A single RNA-guided endonuclease lacking crRNA and a T-rich protos Utilizes pacer adjacent motifs (TTN, TTTN, or YTN). 1 cleaves DNA through staggered DNA double-strand breaks. In addition to the Lee protein, two enzymes from Acidaminococcus and Lachnospiraceae were It has been shown to have efficient genome editing activity in human cells. Quality is known in the art and has been previously described, e.g., in the entire contents of which are incorporated herein. Yamano et al., “Crystal structure of Cpf1 in complex with th guide RNA and target DNA.” Cell (165) 2016, pp. 949-962.
[0225] Also useful in the compositions and methods herein are guide nucleotide sequence programmable Nuclease-inactive Cp that can be used as a polynucleotide-binding protein domain The Cpf1 protein is a variant of the RuvC domain of Cas9. It has a RuvC-like endonuclease domain similar to that of ribozymes, but lacks the HNH endonuclease domain. The N-terminus of Cpf1 does not contain a cleavage domain, and the N-terminus of Cpf1 is the alpha-helical recognition locus of Cas9. Zetsche et al., Cell, 163, 759-771, 2015 (incorporated herein by reference). In the case of cleavage of the ribosomal DNA, the RuvC-like domain of Cpf1 is responsible for cleavage of both DNA strands. , and inactivation of the RuvC-like domain inactivates Cpf1 nuclease activity. For example, D917A, E1006A, and Mutations corresponding to D1255A or D1255A inactivate Cpf1 nuclease activity. In an embodiment, the dCpf1 of the present disclosure includes D917A, E1006A, D1255A, D91 7A / E1006A, D917A / D1255A, E1006A / D1255A, or It contains the mutations corresponding to D917A / E1006A / D1255A. Any mutation that inactivates the domain, e.g., a substitution mutation, deletion, or insertion, may be used in accordance with the present disclosure. It will be understood that it may also be used.
[0226] In some embodiments, the nucleic acid programmable The nucleotide binding protein may be a Cpf1 protein. The pf1 protein is Cpf1 nickase (nCpf1). The pf1 protein is nuclease-inactive Cpf1 (dCpf1). In the form, Cpf1, nCpf1, or dCpf1 is a Cpf1 disclosed herein. At least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical In some embodiments, dCpf1 comprises a Cpf1 sequence disclosed herein. Column and at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 9 7%, at least 98%, at least 99%, or comfortably 99.5% identical amino acid sequence Including D917A, E1006A, D1255A, D917A / E1006A, D91 7A / D1255A, E1006A / D1255A, or D917A / E1006A / Cpfl from other bacterial species may also be used in accordance with the present disclosure. It should be recognized that
[0227] The amino acid sequence of wild-type Francisella novicida Cpf1 is: D917, E1 006, and D1255 are in bold and underlined. MSIYQEFVNKYSLSKTLRFELIPQGKTLENIKARGLILDDEKRAKDYKKAKQIIDKYHQFFIEEILSSVCISEDLLQNYS DVYFKLKKSDDDNLQKDFKSAKDTIKKQISEYIKDSEKFKNLFNQNLIDAKKGQESDLILWLKQSKDNGIELFKANSDIT DIDEALEIIKSFKGWTTYFKGFHENRKNVYSSDIPTSIIYRIVDDNLPKFLENKAKYESLKDKAPEAINYEQIKKDLAE ELTFDIDYKTSEVNQRVFSLDEVFEIANFNNYLNQSGITKFNTIIGGKFVNGENTKRKGINEYINLYSQQINDKTLKKYK MSVLFKQILSDTESKSFVIDKLEDDSDVVTTMQSFYEQIAAFKTVEEKSIKETLSLLFDDLKAQKLDLSKIYFKNDKSLT DLSQQVFDDYSVIGTAVLEYITQQIAPKNLDNPSKKEKELIAKKTEKAKYLSLETIKLALEEFNKHRDIDKQCRFEEILA NFAAIPMIFDEIAQNKDNLAQISIKYQNQGKKDLLQASAEDDVKAIKDLLDQTNLLHKLKIFHISQSEDKANILDKDEH FYLVFEECYFELANIVPLYNKIRNYITQKPYSDEKFKLNFENSTLANGWDKNKEPDNTAILFIKDDKYYLGVMNKKNNKI FDDKAIKENKGGEGYKKIVYKLLPGANKMLPKVFFSASKIFYNPSEDILRIRNHSTHTKNGSPQKGYEKFEFNIEDCRKF IDFYKQSISKHPEWKDFGFRFSDTQRYNSIDEFYREVENQGYKLTFENISEYIDSVVNQGKLYLFQIYNKDFSAYSKGR PNLHTLYWKALFDERNLQDVVYKLNGEAELFYRKQSIPCKITHPAKEAIANKNKDNPKKESVFEYDLIKDKRFTEDKFFF HCPITINFKSSGANKFNDEINLLLKEKANDVHILSI DRGERHLAYYTLVDGKGNIIKQDTFNIIGNDRMKTNYHDKLAAI EKDRDSARKDWKKINNIKEMKEGYLSQVVHEIAKLVIEYNAIVVF E DLNFGFKRGRFKVEKQVYQKLEKMLIEKLNYLVF KDNEFDKTGGVLRAYQLTAPFETFKKMGKQTGIIYYVPAGFTSKICPVTGFVNQLYPKYESVSKSQEFFSKFDKICYNLD KGYFEFSFDYKNFGDKAAKGKWTIASFFGSRLINFRNSDKNHNWDTREVYPTKELEKLLKDYSIEYGHGECIKAAICGESD KKFFAKLTSVLNTILQMRNSKTGTELDYLISPVADV...
Claims
1. An in vitro or ex vivo method of editing a polynucleotide to alter a stop codon, introduce a splice acceptor or splice donor site, or modify a splice acceptor or splice donor site, comprising contacting the polynucleotide with a base editor in complex with one or more guide polynucleotides, wherein the base editor: i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising a combination of amino acid sequence substitutions with respect to SEQ ID NO:69 selected from the group consisting of: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R (224 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q (227 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237 SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242 SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R, D1332, R1335N, and T1337 (244 SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E, and T1337 (245 SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q, and T1337R (259 SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A, and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGC Rd1 SpCas9); and D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E, and T1337R (267 (NGC Rd2 SpCas9) an SpCas9 domain comprising: ii) a deaminase domain; wherein the one or more guide polynucleotides target the base editor to result in an alteration that alters a stop codon, introduces a splice acceptor or splice donor site, or modifies a splice acceptor or splice donor site, wherein the method does not include editing the genome of a human embryo.
2. 1. An in vitro or ex vivo method of editing an SBDS polynucleotide comprising a mutation associated with Shwachman-Diamond Syndrome (SDS), comprising contacting the SBDS polynucleotide with a base editor in complex with one or more guide polynucleotides, wherein the base editor comprises: i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising a combination of amino acid sequence substitutions with respect to SEQ ID NO:69 selected from the group consisting of: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R (224 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q (227 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237 SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242 SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R, D1332, R1335N, and T1337 (244 SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E, and T1337 (245 SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q, and T1337R (259 SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A, and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGC Rd1 SpCas9); and D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E, and T1337R (267 (NGC Rd2 SpCas9) an SpCas9 domain comprising: ii) a deaminase domain; wherein the one or more guide polynucleotides target the base editor to alter a mutation associated with Shwakman-Diamond Syndrome (SDS), and wherein the method does not include editing the genome of a human embryo.
3. The method described in claim 1 or 2, wherein the deaminase domain is a cytidine deaminase domain or an adenosine deaminase domain.
4. The cytidine deaminase a) an amino acid sequence having at least 90% identity to SEQ ID NO: 130; b) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 and further containing the amino acid modification H122A relative to SEQ ID NO: 130; c) an amino acid sequence having at least 90% identity to SEQ ID NO: 127; d) an amino acid sequence having at least 90% identity to SEQ ID NO: 136; e) an amino acid sequence having at least 90% identity to SEQ ID NO: 126; or f) an amino acid sequence having at least 90% identity to SEQ ID NO: 126 and further comprising the amino acid modification F130L. The method of claim 3, comprising:
5. The method described in claim 3, wherein the adenosine deaminase domain comprises an amino acid sequence having at least 90% identity to SEQ ID NO: 23 and further comprises one or more of the amino acid modifications I76Y, V82S, Y147T, Y147R, Q154S, and T166R.
6. The method of claim 5, wherein the adenosine deaminase domain comprises a combination of modifications selected from the group consisting of Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y147T+Q154R; Y147T+Q154S; V82S+Q154S; and Y123H+Y147R+Q154R+I76Y, based on SEQ ID NO:
23.
7. 5. The method of Claim 3 or 4, wherein the one or more guide polynucleotides target the base editor to result in a C.G to T.A change at rs113993993 258+2T>C.
8. 8. The method of claim 7, wherein the guide polynucleotide comprises one of the following sequences: GUAAGCAGGCGGGUAACAGC; AGCAGGCGGGUAACAGCUGC; GCGGGUAACAGCUGCAGCA; GCAGGCGGGUAACAGCUGC; CAGGCGGGUAACAGCUGC; AGGCGGGUAACAGCUGC; or AAGCAGGCGGGUAACAGCUGC.
9. The method of claim 3, 5, or 6, wherein the one or more guide polynucleotides target the base editor to result in an A.T to G.C change in the 183-184TA>CT Rs113993991 mutation, resulting in a missense mutation.
10. The method described in claim 9, wherein the one or more guide polynucleotides comprise one of the following sequences: UGUAAAUGUUUCCUAAGGUC or AAUGUUUCCUAAGGUCAGGU.
11. i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising a combination of amino acid sequence substitutions with respect to SEQ ID NO:69 selected from the group consisting of: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R (224 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q (227 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237 SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242 SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R, D1332, R1335N, and T1337 (244 SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E, and T1337 (245 SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q, and T1337R (259 SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A, and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGC Rd1 SpCas9); and D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E, and T1337R (267 (NGC Rd2 SpCas9) an SpCas9 domain comprising: ii) a deaminase domain; a base editor, or a polynucleotide encoding said base editor, comprising: one or more guide polynucleotides that target the base editor to effect a change associated with aberrant splicing or to alter a stop codon 10. A cell comprising: a human embryonic stem cell;
12. The cell described in claim 11, wherein the deaminase domain is a cytidine deaminase domain or an adenosine deaminase domain.
13. The cytidine deaminase a) an amino acid sequence having at least 90% identity to SEQ ID NO: 130; b) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 and further containing the amino acid modification H122A relative to SEQ ID NO: 130; c) an amino acid sequence having at least 90% identity to SEQ ID NO: 127; d) an amino acid sequence having at least 90% identity to SEQ ID NO: 136; e) an amino acid sequence having at least 90% identity to SEQ ID NO: 126; or f) an amino acid sequence having at least 90% identity to SEQ ID NO: 126 and further comprising the amino acid modification F130L. The cell of claim 12, comprising:
14. The cell described in claim 12, wherein the adenosine deaminase domain comprises an amino acid sequence having at least 90% identity to SEQ ID NO: 23 and further comprises one or more of the amino acid modifications I76Y, V82S, Y147T, Y147R, Q154S, and T166R.
15. The cell described in claim 14, wherein the adenosine deaminase domain comprises a combination of modifications selected from the group consisting of Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y147T+Q154R; Y147T+Q154S; V82S+Q154S; and Y123H+Y147R+Q154R+I76Y, based on SEQ ID NO:
23.
16. 12. A composition for use in a method for treating Shwakman-Diamond Syndrome (SDS) or a disease associated with abnormal splicing in a subject in need thereof, the composition comprising the cells described in claim 11.
17. 1. A composition for use in a method of treating Shwakman-Diamond Syndrome (SDS) in a subject, comprising: i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising a combination of amino acid sequence substitutions with respect to SEQ ID NO:69 selected from the group consisting of: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R (224 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q (227 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237 SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242 SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R, D1332, R1335N, and T1337 (244 SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E, and T1337 (245 SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q, and T1337R (259 SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A, and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGC Rd1 SpCas9); and D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E, and T1337R (267 (NGC Rd2 SpCas9) an SpCas9 domain comprising: ii) a deaminase domain; a base editor, or a polynucleotide encoding said base editor, comprising: one or more guide polynucleotides that target the base editor to effect mutational alterations associated with SDS; A composition comprising:
18. 1. A composition for use in a method of treating a genetic disease associated with aberrant splicing in a subject, comprising: i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising a combination of amino acid sequence substitutions with respect to SEQ ID NO:69 selected from the group consisting of: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R (224 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q (227 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237 SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242 SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R, D1332, R1335N, and T1337 (244 SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E, and T1337 (245 SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q, and T1337R (259 SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A, and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGC Rd1 SpCas9); and D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E, and T1337R (267 (NGC Rd2 SpCas9) an SpCas9 domain comprising: ii) a deaminase domain; a base editor, or a polynucleotide encoding said base editor, comprising: one or more guide polynucleotides that target the base editor to modify a pathogenic mutation that alters splicing; A composition comprising:
19. The composition described in claim 17 or 18, wherein the deaminase domain is a cytidine deaminase domain or an adenosine deaminase domain.
20. The cytidine deaminase a) an amino acid sequence having at least 90% identity to SEQ ID NO: 130; b) an amino acid sequence having at least 90% identity to SEQ ID NO: 130 and further containing the amino acid modification H122A relative to SEQ ID NO: 130; c) an amino acid sequence having at least 90% identity to SEQ ID NO: 127; d) an amino acid sequence having at least 90% identity to SEQ ID NO: 136; e) an amino acid sequence having at least 90% identity to SEQ ID NO: 126; or f) an amino acid sequence having at least 90% identity to SEQ ID NO: 126 and further comprising the amino acid modification F130L.
20. The composition of claim 19, comprising:
21. The composition described in claim 19, wherein the adenosine deaminase domain comprises an amino acid sequence having at least 90% identity to SEQ ID NO: 23 and further comprises one or more of the amino acid modifications I76Y, V82S, Y147T, Y147R, Q154S, and T166R.
22. The composition described in claim 21, wherein the adenosine deaminase domain comprises a combination of modifications selected from the group consisting of Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y147T+Q154R; Y147T+Q154S; V82S+Q154S; and Y123H+Y147R+Q154R+I76Y, based on SEQ ID NO:
23.
23. 1. An in vitro or ex vivo method for producing non-human mammalian cells or human cells or precursors thereof, comprising: (a) in induced pluripotent stem cells containing a gene conversion associated with Shwachman-Diamond syndrome (SDS), i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising a combination of amino acid sequence substitutions with respect to SEQ ID NO:69 selected from the group consisting of: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R (224 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q (227 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237 SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242 SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R, D1332, R1335N, and T1337 (244 SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E, and T1337 (245 SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q, and T1337R (259 SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A, and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGC Rd1 SpCas9); and D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E, and T1337R (267 (NGC Rd2 SpCas9) an SpCas9 domain comprising: ii) a deaminase domain; a base editor or a polynucleotide encoding the base editor, comprising: One or more guide polynucleotides that target base editors to alter mutations associated with SDS to introduce; and (b) differentiating the induced pluripotent stem cells or progenitors into a desired cell type. wherein the method does not include editing the genome of a human embryo.
24. 1. An in vitro or ex vivo method for editing a pathogenic mutation in a gene that results in aberrant splicing, comprising: a target nucleotide sequence, at least a portion of which is located in the gene or its reverse complement, (i) a Streptococcus pyogenes Cas9 (SpCas9) domain in association with a guide polynucleotide that directs the base editor to target a target polynucleotide sequence, at least a portion of which is located in the gene or its reverse complement, wherein the domain has specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5′-NGC-3′, and comprises an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprises a combination of amino acid sequence substitutions with respect to SEQ ID NO:69 selected from the group consisting of: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R (224 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q (227 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237 SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242 SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R, D1332, R1335N, and T1337 (244 SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E, and T1337 (245 SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q, and T1337R (259 SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A, and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGC Rd1 SpCas9); and D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E, and T1337R (267 (NGC Rd2 SpCas9) an SpCas9 domain comprising: (ii) a deaminase domain; contacting a base editor comprising: and editing the pathogenic mutation by directing a base editor at the target nucleotide sequence and deaminating the pathogenic mutation or its complementary nucleobase. Including, deaminating the pathogenic mutation or its complementary nucleobase, thereby converting the pathogenic mutation into a sequence that permits splicing, thereby correcting the pathogenic mutation, wherein the method does not include editing the genome of a human embryo.
25. 1. An in vitro or ex vivo method of producing a cell, tissue, or organ for treating SDS in a subject in need thereof by correcting a pathogenic mutation in the SDS gene of the cell, tissue, or organ, comprising: the cells, tissues, or organs, (i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', and comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising a combination of amino acid sequence substitutions with respect to SEQ ID NO:69 selected from the group consisting of: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R (224 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q (227 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237 SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242 SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R, D1332, R1335N, and T1337 (244 SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E, and T1337 (245 SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q, and T1337R (259 SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A, and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGC Rd1 SpCas9); and D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E, and T1337R (267 (NGC Rd2 SpCas9) an SpCas9 domain comprising: (ii) a deaminase domain; contacting a base editor comprising: contacting the cell, tissue, or organ with a guide polynucleotide, wherein the guide polynucleotide directs a base editor to a target nucleotide sequence located at least in part in the gene or its reverse complement; and editing the pathogenic mutation or its complementary nucleobase by deaminating the mutation when a base editor is directed against the target nucleotide sequence. Including, deaminating the pathogenic mutation or its complementary nucleobase to permit splicing, thereby producing the cell, tissue, or organ for treating SDS, wherein the method does not include editing the genome of a human embryo.
26. 1. A composition comprising a base editor bound to a guide RNA, wherein the guide RNA comprises a nucleic acid sequence complementary to an SBDS gene associated with Shwachman-Diamond Syndrome (SDS), and wherein the base editor i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising a combination of amino acid sequence substitutions with respect to SEQ ID NO:69 selected from the group consisting of: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R (224 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q (227 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237 SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242 SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R, D1332, R1335N, and T1337 (244 SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E, and T1337 (245 SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q, and T1337R (259 SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A, and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGC Rd1 SpCas9); and D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E, and T1337R (267 (NGC Rd2 SpCas9) an SpCas9 domain comprising: ii) a deaminase domain; A composition comprising:
27. 27. A pharmaceutical composition for the treatment of Shwakman-Diamond Syndrome (SDS), comprising the composition of claim 26 and further comprising a pharmaceutically acceptable excipient, diluent, or carrier.
28. The guide RNA is a nucleic acid sequence selected from the group consisting of: GUAAGCAGGCGGGUAACAGC; AGCAGGCGGGUAACAGCUGC; GCGGGUAACAGCUGC; GCAGGCGGGUAACAGCUGC; CAGGCGGGUAACAGCUGC; AGGCGGGUAACAGCUGC; AAGCAGGCGGGUAACAGCUGC; UGUAAAUGUUUCCUAAGGUC; and AAUGUUUCCUAAGGUCAGGU 28. The pharmaceutical composition of claim 27, comprising:
29. (i) a nucleic acid encoding a base editor; and (ii) a guide RNA comprising a nucleic acid sequence selected from the group consisting of GUAAGCAGGCGGGUAACAGC; AGCAGGCGGGUAACAGCUGC; GCGGGUAACAGCUGC; GCAGGCGGGUAACAGCUGC; CAGGCGGGUAACAGCUGC; AGGCGGGUAACAGCUGC; AAGCAGGCGGGUAACAGCUGC; UGUAAAUGUUUCCUAAGGUC; and AAUGUUUCCUAAGGUCAGGU and wherein the base editor is i) a Streptococcus pyogenes Cas9 (SpCas9) domain having specificity for a protospacer adjacent motif (PAM) comprising the nucleic acid sequence 5'-NGC-3', comprising an amino acid sequence having at least 90% identity to SEQ ID NO:69, and further comprising a combination of amino acid sequence substitutions with respect to SEQ ID NO:69 selected from the group consisting of: D1135M, S1136Q, G1218K, E1219F, A1322R, D1332, R1335E, and T1337R (224 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (225 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332K, R1335E, and T1337R (226 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337Q (227 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335Q, and T1337Q (230 SpCas9); D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335D, and T1337Q (235 SpCas9); D1135Q, S1136, G1218T, E1219W, A1322R, D1332, R1335N, and T1337 (237 SpCas9); D1135H, S1136, G1218S, E1219W, A1322R, D1332, R1335V, and T1337 (242 SpCas9); D1135C, S1136W, G1218N, E1219W, A1322R, D1332, R1335N, and T1337 (244 SpCas9); D113LM, S1136W, G1218R, E1219S, A1322R, D1332, R1335E, and T1337 (245 SpCas9); D1135G, S1136W, G1218S, E1219M, A1322R, D1332, R1335Q, and T1337R (259 SpCas9); L111R, D1135V, S1136Q, G1218K, E1219F, A1322R, D1332, R1335A, and T1337R (Nureki SpCas9); D1135M, S1136, S1216G, G1218, E1219, A1322, D1332A, R1335Q, and T1337 (NGC Rd1 SpCas9); and D1135G, S1136, S1216G, G1218, E1219, A1322R, D1332A, R1335E, and T1337R (267 (NGC Rd2 SpCas9) an SpCas9 domain comprising: ii) a deaminase domain; A pharmaceutical composition comprising: