Methods and compositions for non-genotoxic conditioning based on anti-cd45
By introducing missense mutations into hematopoietic stem cells and combining them with anti-CD45 antibodies, the genotoxicity risk in the busulfan conditioning method was resolved, achieving safe hematopoietic stem cell conditioning and reducing the risk of tumor and organ toxicity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEAM THERAPEUTICS INC
- Filing Date
- 2024-05-14
- Publication Date
- 2026-06-05
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Cross-reference to related applications
[0001] This application claims priority to U.S. Provisional Application No. 63 / 502,335, filed May 15, 2023, the entire contents of which are hereby incorporated by reference.
[0002] sequence list This application contains a sequence list, which has been submitted electronically in XML format and is hereby incorporated in its entirety by reference. The sequence list XML file, created on May 14, 2024, is named 180802-052002PCT_SL.xml and has a size of 952,822 bytes. Background Technology
[0003] Busulfan is a DNA alkylating agent that induces bone marrow immunosuppression and is widely used in conditioning prior to allogeneic hematopoietic stem cell transplantation and autologous cell therapy. While busulfan is currently the standard of care for patients requiring allogeneic or autologous transplantation and engrafted cell therapy, the use of this potent cytotoxic agent carries associated risks, including genotoxicity, primary or secondary malignancy, and organ toxicity (including infertility). These risks present a barrier for patients seeking treatment. Therefore, there is a need for improved methods of conditioning prior to allogeneic hematopoietic stem cell transplantation. Summary of the Invention
[0004] As described below, this disclosure provides compositions and methods for non-genotoxic conditioning with anti-CD45 antibodies, anti-CD45 antibody-drug conjugates, or anti-CD45 chimeric antigen receptor-expressing T cells (CAR-T), wherein the method uses a base editor to alter the differentiation cluster 45 (protein tyrosine phosphatase, receptor type C, also known as PTPRC) polynucleotide sequence in hematopoietic stem cells or progenitor cells (HSPCs) to encode a CD45 polypeptide with reduced antibody binding. In some embodiments, this disclosure provides compositions and methods for treating diseases and disorders by non-genotoxic conditioning with anti-CD45 antibodies, anti-CD45 antibody-drug conjugates, or anti-CD45 chimeric antigen receptor-expressing T cells (CAR-T) before, simultaneously with, or after transplantation of base-edited HSPCs. The compositions and methods disclosed herein can advantageously reduce or eliminate the use of genotoxic agents (such as busulfan) associated with primary or secondary malignancies and organ toxicity (including infertility). Furthermore, the compositions and methods disclosed herein provide edited HSPCs resistant to the disclosed anti-CD45 antibody, anti-CD45 antibody drug conjugate, and anti-CD45 chimeric antigen receptor-expressing T cells (CAR-T), and thus can be simultaneously exposed to opsonizing agents in the patient.
[0005] In one aspect, this disclosure provides a method for generating edited hematopoietic stem cells or progenitor cells (HSPCs) for treating diseases or ailments. The method includes (a) expressing a nucleobase editor polypeptide in the hematopoietic stem cells or progenitor cells. The nucleobase editor polypeptide contains a nucleic acid programmable DNA-binding protein (napDNAbp) domain and a deaminase domain. The method further includes (b) contacting the hematopoietic stem cells or progenitor cells with a guide RNA (gRNA) or a polynucleotide encoding the gRNA. The gRNA targets a polynucleotide encoding a differentiation cluster 45 (CD45) polypeptide. The method results in the introduction of a missense mutation in the CD45 polynucleotide in the cell. The missense mutation is located in a portion of the polynucleotide encoding an extracellular domain and / or in fibronectin domains 1 to 4 of the CD45 polypeptide. The missense mutation is associated with reduced binding of anti-CD45 antibodies to the CD45 polypeptide expressed by the edited HSPC.
[0006] In another aspect, this disclosure provides a method for conditioning a subject with a disease or disorder concurrently with or prior to hematopoietic stem cell transplantation (HSCT). The method includes (a) expressing a nucleobase editor polypeptide containing a nucleic acid-programmable DNA-binding protein (napDNAbp) and a deaminase in hematopoietic stem cells of the subject or donor. The method further includes (b) contacting hematopoietic stem cells or progenitor cells (HSPCs) with a guide RNA (gRNA) or a polynucleotide encoding the gRNA. The gRNA targets the polynucleotide encoding the differentiation cluster 45 (CD45) polypeptide. The method results in the introduction of a missense mutation in the polynucleotide and the generation of an edited HSPC. The missense mutation is in a portion of the polynucleotide encoding an extracellular domain and / or in fibronectin domains 1 to 4 of the CD45 polypeptide. The method further includes (c) administering the edited hematopoietic stem cells to the subject. The method further includes (d) administering to the subject an anti-CD45 antibody, an anti-CD45 antibody-drug conjugate, or anti-CD45 chimeric antigen receptor-expressing T cells (CAR-T) that selectively binds to wild-type CD45 protein. Step (d) is administered before, simultaneously with, or after step (c). Missense mutations are associated with reduced binding of anti-CD45 antibodies, anti-CD45 antibody conjugates, or anti-CD45 chimeric antigen receptor-expressing T cells (CAR-T) to the CD45 peptide expressed by edited HSPCs.
[0007] In another aspect, this disclosure provides a base editor system containing a nucleobase editor polypeptide or a polynucleotide encoding a nucleobase editor polypeptide. The nucleobase editor polypeptide contains a nucleic acid programmable DNA-binding protein (napDNAbp) domain and a deaminase domain. The base editor system also contains a guide RNA (gRNA) or a polynucleotide encoding gRNA. The gRNA targets the polynucleotide encoding the differentiation cluster 45 (CD45) polypeptide. The base editor system is capable of introducing missense mutations into the CD45 polynucleotide in cells. The missense mutations are located in a portion of the polynucleotide encoding the extracellular domain and / or in fibronectin domains 1 to 4 of the CD45 polypeptide. The missense mutations are associated with reduced binding of anti-CD45 antibodies to the CD45 polypeptide expressed by the edited HSPC.
[0008] In another aspect, this disclosure provides a polynucleotide or polynucleotide set that encodes a base editor system, or a component thereof, any aspect of this disclosure or an embodiment thereof.
[0009] In another respect, this disclosure provides a vector or set of vectors containing any aspect of this disclosure or embodiments thereof, or a set of polynucleotides.
[0010] In another embodiment, this disclosure provides an HSPC prepared by any aspect of this disclosure or an embodiment thereof.
[0011] In another aspect, this disclosure provides an HSPC containing any aspect of this disclosure or an embodiment thereof, a base editor system, a polynucleotide or polynucleotide group, or a vector or vector group.
[0012] In another aspect, this disclosure provides a pharmaceutical composition comprising a base editor system, a polynucleotide or group of polynucleotides, a carrier or group of carriers, or an HSPC of any aspect of this disclosure or an embodiment thereof, and a pharmaceutically acceptable excipient.
[0013] In another aspect, this disclosure provides a kit suitable for use in any aspect of this disclosure or in any embodiment thereof, wherein the kit contains a base editor system, a polynucleotide or polynucleotide group, a carrier or carrier group, an HSPC or pharmaceutical composition, and a container, for any aspect of this disclosure or any embodiment thereof.
[0014] In another aspect, this disclosure provides a method for generating edited hematopoietic stem cells or progenitor cells (HSPCs) for treating diseases or ailments. The method includes (a) expressing a nucleobase editor polypeptide in the hematopoietic stem cells or progenitor cells. The nucleobase editor polypeptide contains a SpCas9 polypeptide specific to adjacent motifs selected from the prototype spacer region of NRCH and NGC, wherein “N” is A, T, G, or C, “R” is A or G, and “H” is A, C, or T. The nucleobase editor polypeptide also contains an adenosine deaminase domain. The adenosine deaminase domain contains TadA*8.20, TadA-8e, or TadA*8.20 with amino acid alterations S82T, Y147D, F149Y, T166I, and D167N. The method further includes (b) contacting the hematopoietic stem cells or progenitor cells with gRNA4630 or gRNA2696. The method results in the introduction of a missense mutation in the CD45 polynucleotide in the cells. Missense mutations result in an alteration of E259G in the CD45 polypeptide encoded by the CD45 polynucleotide.
[0015] In another aspect, this disclosure provides an HSPC prepared according to any aspect of this disclosure or an embodiment thereof.
[0016] In another aspect, this disclosure provides a method for treating a subject suffering from a disease or ailment. The method includes a) administering HSPC cells of any aspect of this disclosure or an embodiment thereof to the subject. The method further includes b) administering the mAb030-9 peptide to the subject before, simultaneously with, or after step a).
[0017] In any aspect of this disclosure or in any implementation thereof, missense mutations do not alter the biological activity of CD45.
[0018] In any aspect of this disclosure or any embodiment thereof, the deaminase domain contains an adenosine deaminase domain or a cytidine deaminase domain. In any aspect of this disclosure or any embodiment thereof, the adenosine deaminase domain contains TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, and TadA*8.1. 4. TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, TadA*8.24, TadA-8e, or TadA*8.20 having amino acid modifications S82T, Y147D, F149Y, T166I, and D167N. In any aspect of this disclosure or any embodiment thereof, the adenosine deaminase domain contains TadA*8.20, TadA-8e, or TadA*8.20 having amino acid modifications S82T, Y147D, F149Y, T166I, and D167N. In any aspect of this disclosure or any embodiment thereof, the adenosine deaminase domain also contains wild-type TadA.
[0019] In any aspect of this disclosure or any embodiment thereof, the adenosine deaminase domain is inserted within the napDNAbp domain. In any aspect of this disclosure or any embodiment thereof, the napDNAbp domain is SpCas9, and the adenosine deaminase domain is inserted between amino acid positions 1247 and 1248, as numbered in SEQ ID NO: 197.
[0020] In any aspect of this disclosure or in any embodiment thereof, the nucleobase editor polypeptide contains an amino acid sequence that has at least about 85% identity with the sequences listed in Table 17.
[0021] In any aspect of this disclosure or its embodiments, the missense mutation is selected from one or more of the following: E256G; D238G, Y239C; E259G; E259K; E259G, N262D; I283M; I283M, H285R; I283M, H285R, N286G; K231R, Y232C; N253G, N255G; N255G, E256G; N255G, E256G, N257D; N255G E256G, N257G; N257G; N257S; N257D; E259R; N257G, E256G; N257G, E256G, N257D; N263G, T264A; N263S, T264A; N263S, T264A, T266A; N267G, N268G; N267S; N267G; N286G; T264A; T266A; T266A, N267G; and N286D.
[0022] In any aspect of this disclosure or any embodiment thereof, the gRNA contains a spacer sequence selected from the spacer sequences provided in Table 1 or Table 2. In any aspect of this disclosure or any embodiment thereof, the gRNA contains a gRNA sequence selected from those listed in Table 2.
[0023] In any aspect of this disclosure or in any embodiment thereof, the napDNAbp domain is SpCas9. In any aspect of this disclosure or in any embodiment thereof, SpCas9 is specific for the motif sequence adjacent to the prototype spacer region of NGC or NRCH, wherein “N” is A, T, G or C, “R” is A or G, and “H” is A, C or T.
[0024] In any aspect of this disclosure or its implementation, the disease or ailment is a human disease selected from one or more of the following: sickle cell disease, β-thalassemia, multiple sclerosis, systemic sclerosis (sSC), systemic lupus erythematosus (SLE), rheumatoid arthritis (RA), multiple myeloma, plasma cell disease, acute myeloid leukemia, non-Hodgkin lymphoma, myelodysplastic syndrome, myeloproliferative neoplasm, acute lymphoblastic leukemia, Hodgkin lymphoma, chronic myeloid leukemia, HIV, and Crohn's disease.
[0025] In any aspect of this disclosure or any embodiment thereof, the anti-CD45 antibody, anti-CD45 antibody-drug conjugate, or anti-CD45 chimeric antigen receptor comprises the following CDRs: VH CDR1: GFDFSRYW (SEQ ID NO: 430); VH CDR2: INPTSSTI (SEQ ID NO: 431); VH CDR3: ARGNYYRYGDAMDY (SEQ ID NO: 432); VL CDR1: KSVSTSGYSYL (SEQ ID NO: 433); VL CDR2: LAS; and VL CDR3: QHSRELPFT (SEQ ID NO: 434). In any aspect of this disclosure or any embodiment thereof, the anti-CD45 antibody or anti-CD45 antibody-drug conjugate is a humanized anti-CD45 antibody or anti-CD45 antibody-drug conjugate.
[0026] In any aspect of this disclosure or in any embodiment thereof, the vector or group of vectors is a lipid nanoparticle. In any aspect of this disclosure or in any embodiment thereof, the vector or group of vectors is a viral vector. In any aspect of this disclosure or in any embodiment thereof, the viral vector is an adeno-associated virus vector, a lentiviral vector, or a rabies virus vector.
[0027] In any aspect of this disclosure or in any embodiment thereof, the adenosine deaminase domain is inserted between amino acid positions 1247 and 1248 of the SpCas9 polypeptide, as numbered in SEQ ID NO: 197.
[0028] In any aspect of this disclosure or in any embodiment thereof, the nucleobase editor polypeptide contains an amino acid sequence having at least about 90% identity with a sequence selected from one or more of the following: 1570 2518 3626 3167
[0029] In no aspect or implementation thereof provided herein is the method intended to modify human phylogenetic identity.
[0030] definition Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by one of ordinary skill in the art to which this disclosure pertains. The following references provide general definitions for many of the terms used in this disclosure: Singleton et al. , Dictionary of Microbiology and Molecular Biology (2nd edition 1994); The Cambridge Dictionary of Science and Technology (Walker, 1988); TheGlossary of Genetics, 5th edition, R. Rieger et al. (Editor), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, unless otherwise specified, the following terms have the meanings assigned to them below.
[0031] "Differentiation Cluster 45 (CD45) polypeptide" or "Protein Tyrosine Phosphatase, Receptor Type C (PTPRC)" refers to a protein or functional fragment thereof that possesses tyrosine phosphatase activity, shares at least approximately 85% amino acid sequence identity with NCBI Reference Sequence Entrance No. NP_002829.3 provided below, and has CD45 biological activity. The following provides information from... Homo sapiens An exemplary CD45 peptide amino acid sequence (NCBI RefSeq accession number NP_002829.3).
[0032] CD45 amino acid sequence [Homo sapiens] NCBI reference sequence
[0033] "CD45 polynucleotide" refers to a nucleic acid molecule or fragment thereof encoding the CD45 polypeptide and its introns, exons, and regulatory sequences associated with its expression. In the implementation scheme, the CD45 polynucleotide is a genomic sequence corresponding to the sequence encoding the CD45 polypeptide. Homo sapiens Genomic sequences, mRNAs associated with and / or required for CD45 peptide expression. An exemplary CD45 gene sequence is provided in ENSEMBL accession number ENSG00000081237. The following are from... Homo sapiens An exemplary CD45 nucleotide sequence (NCBI RefSeq accession number NM_002838.5).
[0034] >NM_002838.5:142-4062 Homo sapiens protein tyrosine phosphatase receptor type C (PTPRC), transcript variant 1, mRNA
[0035] "CD45 polypeptide bioactivity" refers to tyrosine phosphatase activity.
[0036] The term "mAb U139AGL030-9 (mAb030-9) polypeptide" refers to an antibody having at least approximately 85% amino acid sequence identity with the antibody sequence of antibody mAb030-9, or an antibody comprising VH and / or VL CDRs 1-3 of mAb030-9 or an antigen-binding fragment thereof, wherein each antibody, CDR, and antigen-binding fragment specifically binds to the wild-type CD45 polypeptide, but cannot detectably bind to altered CD45 polypeptides or only exhibits reduced binding. In an embodiment, the antibody sequence is the variable light chain region amino acid sequence, the variable heavy chain region amino acid sequence, the heavy chain amino acid sequence, the light chain amino acid sequence of the mAb030-9 polypeptide, or an antigen-binding fragment thereof. In an embodiment, the antibody or its antigen-binding fragment has at least 90%, 93%, 95%, 98%, 99%, or 100% amino acid sequence identity with the antibody sequence of antibody mAb030-9. Below are exemplary heavy and light chain sequences of antibody mAb030-9, where variable domains are plain text, constant domains are bolded, and complementarity-determining regions (CDRs) (i.e., CDR1, CDR2, and CDR3) are underlined: mAb030-9 Heavy Chain EVKLLESGGGLVQPGGSLKLSCAAS GFDFSRYW MSWVRQAPGKGLEWIGE INPTSSTI NFTPSLKDKVFISRDNAKNTLYLQMSKVRSEDTALYYC ARGNYYRYGDAMDY WGQGTSVTVSSAKTTPPSVYPLAPGSAAQTNSMVTLGCLVKGYFPEPVTVTWNSGSLSSGVHTFPAVLQSDLYTLSSSVTVPSSTWPSETVTCNVAHPASSTKVDKKIVPRDCGCKPCICTVPEVSSVFIFPPKPKDVLTITLTPKVTCVVVDISKDDPEVQFSWFVDDV EVHTAQTQPREEQFNSTFRVSELPIMHQDWLNGKEFKCRVNSAAFPAPIEKTISKTKGRPKAPQVYTIPPPKEQMAKDKVSLTCMITDFFPEDITVEWQWNGQPAENYKNTQPIMDTDGSYFVYSKLNVQKSNWEAGNTFTCSVLHEGLHNHHTEKSLSHSPGK(SEQID NO: 428) mAb030-9 Light Chain DIALTQSPASLAVSLGQRATISCRAS KSVSTSGYSYL HWYQQKPGQPPKLLIY LAS NLESGVPARFSGSGSGTDFTLNIHPVEEEDAATYYC QHSRELPFT FGSGTKLEIKRADAAPTVSIFPPSSEQLTSGGASVVCFLNNFYPKDINVKWKIDGSERQNGVLNSWTDQDSKDSTYSMSSTLTLTKDEYERHNSYTCEATHKTSTSPIVKSFNRNEC (SEQ ID NO: 429).
[0037] The three CDRs in the VH region of the mAb030-9 antibody are as follows: VH CDR1: GDFFSRYW (SEQ ID NO: 430); VH CDR2: INPTSSTI (SEQ ID NO: 431); and VH CDR3: ARGNYYRYGDAMDY (SEQ ID NO: 432).
[0038] The three CDRs in the VL region of the mAb030-9 antibody are as follows: VL CDR1: KSVSTSGYSYL (SEQ ID NO: 433); VL CDR2: LAS; and VL CDR3: QHSRELPFT (SEQ ID NO: 434).
[0039] The four framework (FR) regions (FR1, FR2, FR3, and FR4) of the mAb030-9 antibody are located in The above text The CDRs in the VH and VL regions shown are flanked by the following. Specifically, the four FRs in the VH region of the mAb030-9 antibody are as follows: VH FR1: EVKLLESGGGLVQPGGSLKLSCAAS (SEQ ID NO: 435); VH FR2: MSWVRQAPGKGLEWIGE (SEQ ID NO: 436); VH FR3: NFTPSLKDKVFISRDNAKNTLYLQMSKVRSEDTALYYC (SEQ ID NO: 437); and VH FR4: WGQGTSVTVSS (SEQ ID NO: 438).
[0040] The four free radicals (FRs) in the VL region of the mAb030-9 antibody are as follows: VL FR1: DIALTQSPASLAVSLGQRATISCRAS (SEQ ID NO: 439); VL FR2: HWYQQKPGQPPKLLIY (SEQ ID NO: 440); VL FR3: NLESGVPARFSGSGSGTDFTLNIHPVEEEDAATYYC (SEQ ID NO: 441); and VL FR4: FGSGTKLEIKRA (SEQ ID NO: 442).
[0041] "mAb U139AGL030-9 (mAb030-9) polynucleotide" refers to a nucleic acid molecule (e.g., DNA) encoding at least one fragment of the mAb030-9 antibody. In one embodiment, the encoded fragment has antigen-binding activity.
[0042] "Adenine" or "9" H "-purine-6-amine" refers to a purine nucleobase with the molecular formula C5H5N5, which has a structural... And it corresponds to CAS number 73-24-5.
[0043] "Adenosine" or "4-amino-1-[(2 R ,3 R 4 S 5 R )-3,4-dihydroxy-5-(hydroxymethyl)oxacyclopentan-2-yl]pyrimidine-2(1 H "-Keto" refers to an adenine molecule linked to a ribose via a glycosidic bond, which has a structural... And it corresponds to CAS number 65-46-3. Its molecular formula is C 10 H 13 N5O4.
[0044] "Adenosine deaminase" or "adenine deaminase" refers to a polypeptide or fragment thereof capable of catalyzing the hydrolytic deamination of adenine or adenosine. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine to inosine or catalyzes the hydrolytic deamination of deoxyadenosine to deoxyinosine. In some embodiments, adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminase (…) described herein… For example,Engineered adenosine deaminases (evolved adenosine deaminases) can be derived from any organism (e.g., eukaryotes, prokaryotes), including but not limited to algae, bacteria, fungi, plants, invertebrates (e.g., insects), and vertebrates (e.g., amphibians, mammals). In some embodiments, the target polynucleotide is single-stranded or double-stranded. In some embodiments, the adenosine deaminase variant is capable of deaminating adenine and cytosine in DNA. In some embodiments, the adenosine deaminase variant is capable of deaminating adenine and cytosine in single-stranded DNA. In some embodiments, the adenosine deaminase variant is capable of deaminating adenine and cytosine in RNA. In embodiments, the adenosine deaminase variant is selected from those described in PCT / US2020 / 018192, PCT / US2020 / 049975, PCT / US2017 / 045381, and PCT / US2020 / 028568, the entire contents of which are incorporated herein by reference for all purposes.
[0045] "Adenosine deaminase activity" refers to the catalytic deamination of adenine or adenosine in polynucleotides into guanine.
[0046] "Adenosine base editor (ABE)" refers to a base editor that contains adenosine deaminase.
[0047] "Adenosine base editor (ABE) polynucleotide" means a polynucleotide encoding ABE. "Adenosine base editor 8 (ABE8) polypeptide" or "ABE8" means a base editor comprising an adenosine deaminase or a variant of an adenosine deaminase as defined herein, wherein the adenosine deaminase variant comprises one or more of the changes listed in Table 5B, one combination of the changes listed in Table 5B, or one or more of the amino acid positions listed in Table 5B, wherein such changes are relative to the following reference sequence: MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 1), or the corresponding position in another adenosine deaminase. In some embodiments, ABE8 comprises a change at amino acid 82 and / or 166 of SEQ ID NO: 1. In some embodiments, as described herein, ABE8 comprises other changes relative to the reference sequence.
[0048] "Adenosine base editor 8 (ABE8) polynucleotide" refers to a polynucleotide that encodes the ABE8 polypeptide.
[0049] "Administration" herein refers to providing a patient or subject with one or more of the compositions described herein. For example, but not limited to, the composition may be administered (e.g., by injection) via intravenous (iv), subcutaneous (sc), intradermal (id), intraperitoneal (ip), or intramuscular (im). One or more of these routes may be used. Parenteral administration may be performed, for example, by bolus injection or by gradual perfusion over time. In some embodiments, parenteral administration includes intravascular, intravenous, intramuscular, intra-arterial, intrathecal, intratumoral, intradermal, intraperitoneal, tracheal, subcutaneous, subepidermal, intra-articular, subcapsular, subarachnoid, and intrasternal infusion or injection. Alternatively or concurrently, it may be administered orally.
[0050] The term "agent" means any small molecule compound, antibody, nucleic acid molecule, or polypeptide or fragment thereof. Non-limiting examples of agents suitable for use in the methods of this disclosure include antibody-drug conjugates (ADCs).
[0051] "Change" means a change in the level, structure, or activity of an analyte, gene, or peptide, as detected by standard methods known in the art, such as those described herein. As used herein, change includes changes in expression levels (e.g., increases or decreases). In embodiments, expression levels are increased or decreased by 10%, 25%, 40%, 50%, or more. In some embodiments, change includes the insertion, deletion, or substitution of nucleobases or amino acids (e.g., through genetic engineering).
[0052] "Improvement" means to reduce, suppress, weaken, eliminate, prevent, or stabilize the development or progression of a disease.
[0053] "Analogous" refers to molecules that are different but have similar functions or structural features. For example, peptide analogs retain the biological activity of the corresponding naturally occurring peptide while having certain biochemical modifications that enhance the function of the analog compared to the naturally occurring peptide. Such biochemical modifications can increase the protease resistance, membrane permeability, or half-life of the analog without altering, for example, ligand binding. Analogs may include non-natural amino acids.
[0054] As used herein, the term "antibody" refers to an immunoglobulin molecule that specifically binds to or reacts with a specific antigen, and includes polyclonal, monoclonal, genetically engineered, and other modified forms of antibodies, including but not limited to chimeric antibodies, humanized antibodies, and heteroconjugated antibodies. For exampleAntibodies include bispecific, trispecific, and tetraspecific antibodies, as well as antigen-binding fragments of antibodies, including, for example, Fab', F(ab')2, Fab, Fv, rlgG, and scFv fragments. Other non-limiting examples of antibodies include the VHH domain. Unless otherwise stated, the term "monoclonal antibody" (mAb) means comprising the complete molecule and antibody fragments (including, for example, Fab and F(ab')2 fragments) capable of specifically binding to a target protein. As used herein, Fab and F(ab')2 fragments refer to antibody fragments lacking the Fc fragment of the complete antibody.
[0055] Antibodies (immunoglobulins) consist of two heavy chains linked together by disulfide bonds and two light chains, each light chain connected to the corresponding heavy chain in a "Y" configuration via disulfide bonds. Each heavy chain has a variable domain (VH) at one end, followed by multiple constant domains (CH). Each light chain has a variable domain (VL) at one end and a constant domain (CL) at the other end. The variable domains (VL) of the light chains are aligned with the variable domains (VL) of the heavy chains, and the constant domains (CL) of the light chains are aligned with the first constant domain (CH1) of the heavy chains. Each pair of variable domains of the light and heavy chains forms an antigen-binding site. The isotype of the heavy chain (γ, α, δ, ε, or μ) determines the immunoglobulin class (IgG, IgA, IgD, IgE, or IgM, respectively). The light chain is either one of two isotypes found in all antibody classes (kappa (κ) or lambda (λ)). The term "antibody" includes complete antibodies, such as polyclonal or monoclonal antibodies (mAbs), as well as proteolytic portions or fragments that enable them to specifically bind to target proteins, such as Fab or F(ab')2 fragments. Antibodies can include chimeric antibodies; recombinant antibodies and engineered antibodies and their antigen-binding fragments. Exemplary functional antibody fragments containing complete or substantially complete variable regions of both the light and heavy chains are defined as follows: (i) Fv, defined as a genetically engineered fragment consisting of a variable region of the light chain and a variable region of the heavy chain expressed as two chains; (ii) Single-chain Fv (“scFv”), a genetically engineered single-chain molecule comprising a variable region of the light chain and a variable region of the heavy chain linked by a suitable polypeptide linker; (iii) Fab, an antibody molecule fragment containing a monovalent antigen-binding portion of an antibody molecule, obtained by treating an intact antibody with the enzyme papain to produce an Fd fragment of both the light and heavy chains, the fragment consisting of a variable domain and a CH1 domain of the heavy chain; (iv) Fab', an antibody molecule fragment containing a monovalent antigen-binding portion of an antibody molecule, obtained by treating an intact antibody with the enzyme pepsin, followed by reduction (two Fab' fragments are generated per antibody molecule); and (v) F(ab')2, an antibody molecule fragment containing the monovalent antigen-binding portion of an antibody molecule, is obtained by treating an intact antibody with the enzyme pepsin (i.e., a dimer of the Fab' fragment held together by two disulfide bonds).
[0056] Antibody structures are well known in the art. In short, the variable regions (V) or domains of the antibody heavy chain (H) and light chain (L) contain complementarity-determining regions (CDRs) that bind to specific antigens or immunogens (e.g., protein antigens or immunogens). The CDRs are located in the antibody heavy chain (V). H ) and light chains (V LThe CDRs are located within the frame (FR) sequence of the V domain of an antibody. The CDRs are the most variable part of the antibody and a key component of the antigen-specific diversity of antibodies produced by B lymphocytes. Generally, three CDRs (CDR1, CDR2, and CDR3) are arranged consecutively within the V domain of an antibody. Because VHHs (such as Camelidae VHHs) are essentially single-chain antibody polypeptides, they contain three CDRs that bind to the antigen or target protein (such as CD45) against the background of four frame (FR) regions, as follows: FR1-CDR1-FR2-CDR2-FR3-CDR3-FR4. Because most sequence variations associated with immunoglobulin and antigen binding occur in the CDRs, these regions are sometimes referred to as hypervariable regions. Typically, CDR1, CDR2, and CDR3 of VHHs contribute to and / or do not interfere with antigen binding.
[0057] "Antigen" refers to a factor that an antibody or other polypeptide capture molecule specifically binds to. In one embodiment, the antigen is a tumor antigen. Exemplary antigens include small molecules, carbohydrates, proteins, and polynucleotides.
[0058] "Base editor (BE)" or "nucleobase editor peptide (NBE)" refers to an agent that binds to polynucleotides and has nucleobase modification activity. In various embodiments, the base editor comprises a nucleobase modification peptide (NBE). For example (deaminase) and polynucleotide programmable nucleotide binding domain ( For example (Cas9 or Cpf1). Representative nucleic acid and protein sequences of base editors include those sequences that have about or at least about 85% sequence identity with any base editor sequence provided in the sequence listing (such as those corresponding to SEQ ID NO: 2-11).
[0059] "BE4 cytidine deaminase (BE4) polypeptide" refers to a base editor comprising a nucleic acid programmable DNA-binding protein (napDNAbp) domain, a cytidine deaminase domain, and two uracil glycosylation inhibitor domains. In the embodiments, napDNAbp is a Cas9n (D10A) polypeptide. Non-limiting examples of the cytidine deaminase domain include rAPOBEC, ppAPOBEC, RrA3F, AmAPOBEC1, and SsAPOBEC3B.
[0060] "BE4 cytidine deaminase (BE4) polynucleotide" refers to a polynucleotide that encodes the BE4 polypeptide.
[0061] "Base editing activity" refers to the ability to chemically alter bases within a polynucleotide. In one embodiment, the first base is converted to the second base. In one embodiment, the base editing activity is cytidine deaminase activity. For exampleThe target C•G is converted to T•A. In another embodiment, the base editing activity is adenosine or adenine deaminase activity. For example Convert A • T to G • C.
[0062] "Base editing efficiency" refers to the total percentage of one or more target bases in a sample modified using a base editor. In some cases, base editing efficiency is calculated as the total percentage of target polynucleotides in a sample containing modified target bases. In other cases, base editing efficiency is calculated as the total percentage of target polynucleotides in a sample containing modifications to one or more of 2, 3, 4, 5, 6, 7, 8, 9, or 10 target bases (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10). Methods for measuring the base editing efficiency of a base editor are known in the art (see, for example, Gaudelli). et al. Nature 551:464-471 (2017), the contents of which are incorporated herein by reference in their entirety for all purposes. In some cases, base editing efficiency is the median base editing efficiency calculated at 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100 or more target sites.
[0063] The "base editing window" of a base editor refers to the bases within a target polynucleotide sequence that can be modified using the base editor. In some embodiments, the positions of nucleobases in the target polynucleotide sequence are numbered relative to their specific prototype spacer adjacent motifs (PAMs) within the base editor's nucleoprogrammable DNA-binding protein (napDNAbp) domain, where base 1 corresponds to the base immediately adjacent to the PAM. In some embodiments, the positions of nucleobases in the target polynucleotide sequence are numbered relative to the 5′ or 3′ end of the spacer region of the guide polynucleotide used to direct the base editor's nucleoprogrammable DNA-binding protein (napDNAbp) domain to the target site, where base 1 corresponds to the 5′ or 3′ terminal base of the spacer region.
[0064] The term "base editor system" refers to an intermolecular complex used to edit the nucleobases of a target nucleotide sequence. In various embodiments, the base editor (BE) system comprises (1) a multinucleotide programmable nucleotide-binding domain for deamination of nucleobases in the target nucleotide sequence, a deaminase domain, and ( For example (1) cytidine deaminase or adenosine deaminase; and (2) one or more guide polynucleotides that bind to a polynucleotide programmable nucleotide binding domain. For example(Guide RNA). In various embodiments, the base editor (BE) system comprises a nucleobase editor domain selected from adenosine deaminase or cytidine deaminase, and a domain having nucleic acid sequence-specific binding activity. In some embodiments, the base editor system comprises (1) a base editor (BE) comprising a polynucleotide programmable DNA binding domain and a deaminase domain for deaminating one or more nucleobases in a target nucleotide sequence; and (2) one or more guide RNAs that bind to the nucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a polynucleotide programmable DNA binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE) or a cytidine or cytosine base editor (CBE). In some embodiments, the base editor system (e.g., a base editor system comprising cytidine deaminase) contains a uracil glycosylase inhibitor or other agent or peptide that inhibits the inosine base excision repair system (e.g., a uracil stabilizing protein (such as that provided in WO2022015969, the disclosure of which is incorporated herein by reference in its entirety for all purposes)).
[0065] The term "Cas9" or "Cas9 domain" refers to the structure containing the Cas9 protein or a fragment thereof. For example Cas9 nucleases are RNA-directed nucleases that contain the active, inactive, or partially active DNA-cutting domain of Cas9 and / or the gRNA-binding domain of Cas9. Cas9 nucleases are sometimes also referred to as casnl nucleases or CRISPR (clustered regularly spaced short palindromic repeats)-associated nucleases.
[0066] "Chimeric antigen receptor" or "CAR" refers to a receptor that contains one or more intracellular signal transduction domains. For example A synthetic or engineered receptor that binds to an extracellular antigen-binding domain (T cell signaling domain), conferring antigen specificity to immune effector cells (e.g., T cells, NK cells, or macrophages). In embodiments, the CAR is a SUPRA CAR, an anti-tag CAR, a TCR CAR, or a TCR-like CAR (see, for example, Guedan). et al. “Engineering and Design of ChimericAntigen Receptors,” Methods and Clinical Development , 12:145-156 (2019); Poorebrahim et al., “TCR-like CARs and TCR-CARs targeting neoepitopes: anemerging potential,” Cancer Gene Therapy , 28:581-589 (2021); and Minutolo et al. “The Emergence of Universal Immune Receptor T Cell Therapy for Cancer,” Front Oncol. , 9:176 (2019), the contents of which are incorporated herein by reference in their entirety for all purposes.
[0067] "Chimeric antigen receptor (CAR) T cells" or "CAR-T cells" refers to T cells that express a CAR, which has antigen specificity defined by the antibody-derived targeting domain of the CAR. As used herein, "CAR-T cells" includes T cells, regulatory T cells (T cells, etc.). REG ), macrophages, or NK cells. As used herein, "CAR-T cells" include cells engineered to express CAR or T cell receptor (TCR, sometimes called TCR-CAR or TCR-like CAR). CAR production ( For example The methods used to treat cancer are publicly available. See, for example Park et al. , Trends Biotechnol., 29:550-557, 2011; Grupp et al. , N Engl J Med., 368:1509-1518, 2013; Han et al. , J. Hematol Oncol. 6:47,2013; Haso et al., (2013) Blood, 121, 1165-1174; Mohseni et al., (2020) Front.Immunol., 11, art. 1608, doi: 10.3389 / fimmu.2020.01608; Eggenhuizen et al., Int.J. Mol. Sci. (2020), 21:7015, doi: 10.3390 / ijms21197015; Poorebrahim et al., Cancer Gene Ther 28, 581–589 (2021), doi.org / 10.1038 / s41417-021-00307-7, PCTPubs. WO2012 / 079000, WO2013 / 059593; and U.S. Publication 2012 / 0213783, the contents of which are incorporated herein by reference in their entirety.
[0068] "Chimeric antigen receptor" or "CAR" refers to a synthetic or engineered receptor comprising an extracellular antigen-binding domain operatively bound to one or more intracellular signaling domains, wherein the CAR confers specificity on immune effector cells for antigens bound by the extracellular antigen-binding domain. In some cases, the intracellular signaling domain is a T cell signaling domain. In embodiments, the immune effector cells are T cells, NK cells, or macrophages. In embodiments, the CAR is a SUPRACAR, an anti-tag CAR, a TCR CAR, or a TCR-like CAR (see, for example, Guedan). et al. “Engineering and Design ofChimeric Antigen Receptors,” Methods and Clinical Development , 12:145-156(2019); Poorebrahim et al. , “TCR-like CARs and TCR-CARs targeting neoepitopes: anemerging potential,” Cancer Gene Therapy , 28:581-589 (2021); and Minutolo et al. “The Emergence of Universal Immune Receptor T Cell Therapy for Cancer,” Front Oncol. , 9:176 (2019), the contents of which are incorporated herein by reference in their entirety for all purposes.
[0069] "Chimeric antigen receptor (CAR) T cells" or "CAR-T cells" refers to T cells that express a CAR, which has antigen specificity defined by the antibody-derived targeting domain of the CAR. As used herein, "CAR-T cells" includes T cells, regulatory T cells (T cells, etc.). REG ), macrophages, or NK cells. As used herein, the term "CAR-T cells" includes cells engineered to express CAR or T-cell receptor (TCR, sometimes called TCR-CAR or TCR-like CAR). Methods for manufacturing CARs ( For example (For the treatment of cancer) is publicly available. See, for example Park et al. , Trends Biotechnol., 29:550-557, 2011; Grupp et al. , N Engl J Med., 368:1509-1518, 2013; Han et al. , J. Hematol Oncol. 6:47, 2013; Haso et al. , (2013) Blood, 121, 1165-1174; Mohseni et al., (2020) Front.Immunol., 11, art. 1608, doi: 10.3389 / fimmu.2020.01608; Eggenhuizen et al., Int.J. Mol. Sci. (2020), 21:7015, doi: 10.3390 / ijms21197015; Poorebrahim et al., Cancer Gene Ther 28, 581–589 (2021), doi.org / 10.1038 / s41417-021-00307-7, PCTPubs. WO2012 / 079000, WO2013 / 059593; and U.S. Publication 2012 / 0213783, the contents of which are incorporated herein by reference in their entirety.
[0070] The term "conserved amino acid substitution" or "conserved mutation" refers to the substitution of one amino acid for another that shares a common characteristic. One functional way to define the common characteristic between individual amino acids is to analyze the normalized frequencies of amino acid changes between corresponding proteins in homologous organisms (Schulz, GE, and Schirmer, RH, Principles of Protein Structure, Springer-Verlag, New York (1979)). Based on this type of analysis, amino acid groups can be defined where amino acids within a group preferentially exchange with each other and are therefore most similar in their effects on the overall protein structure (Schulz, GE, and Schirmer, RH, ibid.). Non-restrictive examples of conserved mutations include amino acid substitutions, such as lysine replacing arginine, and in turn, maintaining a positive charge; glutamic acid replacing aspartic acid, and in turn, maintaining a negative charge; serine replacing threonine, maintaining free –OH; and glutamine replacing asparagine, maintaining free –NH2.
[0071] Amino acids can generally be classified into several categories based on the following common side chain characteristics: (1) Hydrophobic: Leucine, Met, Ala, Val, Leu, He; (2) Neutral and hydrophilic: Cys, Ser, Thr, Asn, Gln; (3) Acidic: Asp, Glu; (4) Alkaline: His, Lys, Arg; (5) Residues that affect chain orientation: Gly, Pro; (6) Fang ethnic group: Trp, Tyr, Phe.
[0072] In some embodiments, conservative substitution may include exchanging a member of one of these classes for another member of the same class. In some embodiments, non-conservative amino acid substitution may include exchanging a member of one of these classes for another class.
[0073] In this document, the terms "coding sequence" or "protein-coding sequence," used interchangeably, refer to a segment of a polynucleotide that encodes a protein. A coding sequence may also be referred to as an open reading frame. The region or sequence is bounded by a start codon closer to the 5' end and by a stop codon closer to the 3' end. The stop codons available for the base editor described herein include the following: TAG, TAA, and TGA.
[0074] As used herein, the terms “conditioning” and “conditioning” refer to the process by which a patient is prepared to receive a transplant containing hematopoietic stem cells. Such procedures facilitate the engraftment of the hematopoietic stem cell transplant (e.g., as inferred from the sustained increase in the number of live hematopoietic stem cells in blood samples isolated from the patient after the conditioning procedure and subsequent hematopoietic stem cell transplantation). According to the methods described herein, a patient can be adapted for hematopoietic stem cell transplantation by administering an antibody or an antigen-binding fragment thereof capable of binding to antigens expressed by hematopoietic stem cells, such as CD45. Such antibodies are expected to exert their effects via complement-mediated cytotoxicity and antibody-dependent cell-mediated cytotoxicity. As described herein, the transplanted cells have been edited so that the anti-CD45 antibody no longer binds to the antigen (…). For example CD45). Administering drugs to patients requiring hematopoietic stem cell transplantation that can bind to one or more antigens (CD45). For example Antibodies against CD45, its antigen-binding fragments, antibody-drug conjugates (ADCs), or chimeric antigen receptor-expressing T cells (CAR-T) can promote the engraftment of hematopoietic stem cell transplants, for example, by selectively depleting endogenous hematopoietic stem cells expressing CD45, thereby creating vacancies to be filled by exogenous hematopoietic stem cell transplantation.
[0075] A “complex” refers to a combination of two or more molecules whose interactions depend on intermolecular forces. Non-limiting examples of intermolecular forces include covalent and non-covalent interactions. Non-limiting examples of non-covalent interactions include hydrogen bonds, ionic bonds, halogen bonds, hydrophobic bonds, van der Waals interactions (e.g., dipole-dipole interactions, dipole-induced dipole interactions, and London dispersion forces), and π-effects. In one embodiment, the complex comprises a polypeptide, a polynucleotide, or a combination of one or more polypeptides and one or more polynucleotides. In one embodiment, the complex comprises one or more polypeptides and polynucleotides (e.g., guide RNA) associated to form a base editor (e.g., a base editor comprising a nucleic acid programmable DNA-binding protein (such as Cas9) and a deaminase). In one embodiment, the complex is held together by hydrogen bonds. It should be understood that one or more components of the base editor (e.g., a deaminase or a nucleic acid programmable DNA-binding protein) may be covalently or non-covalently associated. As an example, the base editor may include a deaminase covalently linked (e.g., via a peptide bond) to a nucleic acid programmable DNA-binding protein. Alternatively, the base editor may include a non-covalently associated deaminase and a nucleic acid-programmable DNA-binding protein (e.g., wherein one or more components of the base editor are provided in trans form and associated directly or via another molecule such as a protein or nucleic acid). In one embodiment, one or more components of the complex are held together by hydrogen bonds.
[0076] "Cytosine" or "4-aminopyrimidine-2(1)" H "-ketone" refers to a purine nucleobase with the molecular formula C4H5N3O, which has a structural... And it corresponds to CAS number 71-30-7.
[0077] "Cytosine" refers to a cytosine molecule linked to ribose via a glycosidic bond, which has a structural... And it corresponds to CAS number 65-46-3. Its molecular formula is C9H. 13 N3O5.
[0078] "Cytidine base editor (CBE)" refers to a base editor that contains cytidine deaminase.
[0079] "Cytidine base editor (CBE) polynucleotide" refers to a polynucleotide that encodes CBE.
[0080] "Cytidine deaminase" or "cytosine deaminase" refers to a polypeptide or fragment thereof capable of deaminating cytidine or cytosine. In embodiments, cytidine or cytosine is present in a polynucleotide. In one embodiment, cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine. The terms "cytidine deaminase" and "cytosine deaminase" are used interchangeably throughout the application. Marine lamprey (Petromyzon marinus) Cytosine deaminase 1 (PmCDA1) (SEQ ID NO: 13-14), activation-induced cytidine deaminase (AICDA) (SEQ ID NO: 15-21), and APOBEC (SEQ ID NO: 12-61) are exemplary cytidine deaminases. Other exemplary cytidine deaminase (CDA) sequences are provided in the sequence listing as SEQ ID NO: 62-66 and SEQ ID NO: 67-189. Non-limiting examples of cytidine deaminases include those described in PCT / US20 / 16288, PCT / US2018 / 021878, 180802-021804 / PCT, PCT / US2018 / 048969, and PCT / US2016 / 058344.
[0081] "Cytosine deaminase activity" refers to the catalytic deamination of cytosine or cytidine. In one embodiment, a polypeptide having cytosine deaminase activity converts an amino group to a carbonyl group. In one embodiment, cytosine deaminase converts cytosine to uracil (…). Right now (C to U) or converting 5-methylcytosine to thymine ( Right now ,5mC to T). In some embodiments, the cytosine deaminases provided herein have increased cytosine deaminase activity relative to a reference cytosine deaminase ( , 5mC to T). example like(At least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 times or higher).
[0082] As used herein, the term "deaminase" or "deaminase domain" refers to a protein or fragment thereof that catalyzes a deamination reaction.
[0083] The term "detection" refers to identifying the presence, absence, or amount of an analyte to be detected. In one embodiment, it detects a sequence alteration in a polynucleotide or polypeptide. In another embodiment, it detects the presence of an insertion or deletion.
[0084] "Detectable label" means a composition that, when linked to a molecule of interest, enables the latter to be detected by spectroscopic, photochemical, biochemical, immunochemical, or chemical means. Useful labels include, for example, radioactive isotopes, magnetic beads, metal beads, colloidal particles, fluorescent dyes, electron-dense reagents, enzymes (e.g., commonly used in enzyme-linked immunosorbent assays (ELISA)), biotin, digoxigenin, or haptens.
[0085] "Disease" means any disorder or condition that impairs or interferes with the normal function of cells, tissues, or organs. Exemplary diseases include, but are not limited to, sickle cell disease, beta-thalassemia, multiple sclerosis, systemic sclerosis (sSC), systemic lupus erythematosus (SLE), rheumatoid arthritis (RA), multiple myeloma, plasma cell disease, acute myeloid leukemia, non-Hodgkin's lymphoma, myelodysplastic syndrome, myeloproliferative neoplasm, acute lymphoblastic leukemia, Hodgkin's lymphoma, chronic myeloid leukemia, HIV, or Crohn's disease. In some embodiments, the disease is an autoimmune disease. In some embodiments, the disease is a blood cancer.
[0086] "Effective dose" refers to the amount of medication available to an untreated patient or an individual who is not currently ill. Right now For healthy individuals, the agents (as described in this article) needed to improve disease symptoms are... For example The effective amount of an active compound used in practicing embodiments of this disclosure to therapeutically treat a disease varies depending on the method of administration, the subject's age, weight, and general health condition. Ultimately, the attending physician or veterinarian will determine the appropriate amount and dosage regimen. This amount is referred to as the "effective" amount. In one embodiment, the effective amount is sufficient to induce a biological response in cells (base editors, cells). For example , in vitro or in vivoThe present disclosure discloses a method for introducing altered base editors into genes of interest in cells. In one embodiment, the effective amount is the amount of base editor required to achieve a therapeutic effect. Such a therapeutic effect does not require alteration of the pathogenic gene in all cells of a subject, tissue, or organ, but only in approximately 1%, 5%, 10%, 25%, 50%, 75%, or more of the pathogenic gene present in the cells of the subject, tissue, or organ. In one embodiment, the effective amount is sufficient to improve one or more symptoms of the disease.
[0087] "Fragment" refers to a portion of a polypeptide or nucleic acid molecule. This portion contains at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the full length of a reference nucleic acid molecule or polypeptide. A fragment may contain 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids. In some embodiments, the fragment is a functional fragment.
[0088] The “FR zone” or “FR zone” includes areas related to V H and V L The FR region and the amino acid residues adjacent to the CDR in the VHH. For example, FR region residues can be found in antibodies as described herein, camelid antibodies (VHH), human antibodies, rodent-derived antibodies (e.g., mouse and rat antibodies), humanized antibodies, primate-derived antibodies, chimeric antibodies, antibody fragments (e.g., Fab fragments), VHH, single-chain antibody fragments (e.g., scFv fragments), antibody domains, and bispecific antibodies.
[0089] "Guiding polynucleotide" refers to a polynucleotide or polynucleotide complex that is specific to a target sequence and can bind to a polynucleotide programmable nucleotide-binding domain protein. For example The nucleotide (Cas9 or Cpf1) forms a complex. In one embodiment, the guide polynucleotide is a guide RNA (gRNA). The gRNA may exist as a complex of two or more RNAs or as a single RNA molecule.
[0090] As used herein, the term "hematopoietic stem cell" ("HSC") refers to immature blood cells capable of self-renewal and differentiation into mature blood cells, which include various lineages, including but not limited to granulocytes. For example Promyelocytes, neutrophils, eosinophils, basophils), and erythrocytes ( For example Reticulocytes, erythrocytes), thrombus cells ( For example Promegakaryocytes and platelets produce megakaryocytes and platelets), monocytes ( For exampleMonocytes, macrophages), dendritic cells, microglia, osteoclasts, and lymphocytes ( For example NK cells, B cells, and T cells). These cells can include CD34+ cells. CD34+ cells are immature cells that express the CD34 cell surface marker. In humans, CD34+ cells are considered to include a cell subpopulation with the stem cell characteristics defined above, while in mice, HSCs are CD34-. Furthermore, HSCs also refer to long-term refilled HSCs (LT-HSCs) and short-term refilled HSCs (ST-HSCs). LT-HSCs and ST-HSCs are distinguished based on functional potential and cell surface marker expression. For example, human HSCs are CD34+, CD38-, CD45RA-, CD90+, CD49F+, and lin- (negative for maturation lineage markers (including CD2, CD3, CD4, CD7, CD8, CD10, CD11B, CD19, CD20, CD56, and CD235A)). In mice, bone marrow LT-HSCs were CD34-, SCA-1+, C-kit+, CD135-, Slamfl / CD150+, CD48-, and lin- (negative for maturation lineage markers including Ter119, CD11b, Gr1, CD3, CD4, CD8, B220, and IL7ra), while ST-HSCs were CD34+, SCA-1+, C-kit+, CD135-, Slamfl / CD150+, and lin- (negative for maturation lineage markers including Ter119, CD11b, Gr1, CD3, CD4, CD8, B220, and IL7ra). Furthermore, under steady-state conditions, ST-HSCs exhibited lower quiescence and higher proliferative capacity than LT-HSCs. However, LT-HSCs possessed greater self-renewal potential. Right now They survive throughout adulthood and can be continuously transplanted through successive recipients, while ST-HSCs have limited self-renewal ( Right now (These HSCs can only survive for a limited time and do not have the potential for continuous transplantation.) Any of these HSCs can be used in the methods described herein. ST-HSCs are particularly useful because they are highly proliferative and therefore produce differentiated offspring more quickly.
[0091] As used herein, the term "functional potential of hematopoietic stem cells" refers to the functional characteristics of hematopoietic stem cells, including: 1) pluripotency (referring to the ability to differentiate into multiple different blood lineages (including but not limited to granulocytes)). For example Promyelocytes, neutrophils, eosinophils, basophils), and erythrocytes ( For example Reticulocytes, erythrocytes), thrombus cells ( For examplePromegakaryocytes and platelets produce megakaryocytes and platelets), monocytes ( For example Monocytes, macrophages), dendritic cells, microglia, osteoclasts, and lymphocytes ( For example The hematopoietic stem cells have the following functions: 1) the ability of NK cells, B cells, and T cells to produce daughter cells with the same potential as the mother cells, and further, this ability can be repeated throughout an individual's life without being depleted; 2) the ability of hematopoietic stem cells or their progeny to be reintroduced into the transplant recipient, and then they home to the hematopoietic stem cell niche and re-establish productive and continuous hematopoiesis.
[0092] "Heterologous" or "exogenous" means 1) a polynucleotide or polypeptide sequence that has been experimentally incorporated into the polynucleotide or polypeptide sequence, which is not normally found in nature; and / or 2) a polynucleotide or polypeptide that has been experimentally placed in cells that do not normally contain the polynucleotide or polypeptide. In some embodiments, "heterologous" means that the polynucleotide or polypeptide has been experimentally placed in a non-natural environment. In some embodiments, the heterologous polynucleotide or polypeptide is derived from a first species or host organism and is incorporated into a polynucleotide or polypeptide derived from a second species or host organism. In some embodiments, the first species or host organism is different from the second species or host organism. In some embodiments, the heterologous polynucleotide is DNA. In some embodiments, the heterologous polynucleotide is RNA.
[0093] The term "humanized" antibody refers to the form of non-human (e.g., mouse) antibodies, camel-derived single-domain antibody (sdAb) binding molecules, which consist of only heavy chain antibodies (Ab) or VHH with variable heavy chain (V). H Humanized antibodies consist of chimeric immunoglobulins, immunoglobulin chains, or fragments thereof (such as Fv, Fab, Fab', F(ab')2, or other target-binding subdomains of the antibody) containing a minimal sequence derived from a non-human immunoglobulin. Generally, a humanized antibody or VHH may contain substantially all of at least one variable domain (or two variable domains in the case of a non-VHH antibody), wherein all or substantially all of the CDR regions correspond to the CDR regions of a non-human immunoglobulin. All or substantially all of the FR regions of a humanized antibody may also be derived from a human immunoglobulin sequence. In the case of a non-VHH antibody, the VHH or humanized antibody may also contain at least a portion of an immunoglobulin constant region (Fc) (which may be a constant region of a common sequence of human immunoglobulins). Techniques and methods for humanized antibodies (and VHHs) are known and practiced in the art, such as, for example, in Riechmann et al. Nature 332:323-7, 1988; Kasmiri et al., MethodsThe contents of U.S. Patents 5,530,101, 5,585,089, 5,693,761, 5,693,762, and 6,180,370, EP239400, WO 1991 / 09967, 5,225,539, EP592106, and EP519596, all of which are incorporated herein by reference. Humanized antibodies or VHHs are molecularly engineered to contain even more human-like immunoglobulin domains, and only the CDR of the V region of the monoclonal antibody or VHH is incorporated by carefully examining the hypervariable loop sequence of the V region and adapting it to the structure of the human antibody chain. This process is routine and usually performed by those skilled in the art. See, for example, U.S. Patent No. 6,187,287, the contents of which are incorporated herein by reference.
[0094] "Hybridization" refers to hydrogen bonding between complementary nucleobases, which can be Watson-Crick, Hoogsteen, or reverse Hoogsteen hydrogen bonds. For example, adenine and thymine are paired by forming hydrogen bonds between their complementary nucleobases.
[0095] "Increase" means a positive change of at least 10%, 25%, 50%, 75%, or 100%, or approximately 1.5 times, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times, 15 times, 20 times, 25 times, 30 times, 35 times, 40 times, 45 times, 50 times, or 100 times.
[0096] The terms “base repair inhibitor”, “base repair inhibitor”, “IBR”, or their grammatical equivalents refer to proteins that can inhibit the activity of nucleic acid repair enzymes (such as base excision repair enzymes).
[0097] "Intrinates" are segments of proteins that can self-cut and bind to the remaining segments (expeptides) with peptide bonds in a process called protein splicing.
[0098] The terms "isolated," "purified," or "biopure" refer to substances released to varying degrees from components typically found in their native state. "Isolate" indicates the degree of separation from the original source or surrounding environment. "Purified" indicates a degree of separation beyond isolation. A "purified" or "biopure" protein is sufficiently free of other substances such that any impurities do not materially affect the protein's biological properties or cause other adverse consequences. That is, the nucleic acid or peptide disclosed herein is purified if it is substantially free of cellular material, viral material, or culture medium when produced by recombinant DNA technology, or substantially free of chemical precursors or other chemicals during chemical synthesis. Purity and homogeneity are typically determined using analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high-performance liquid chromatography. The term "purified" can mean that the nucleic acid or protein produces essentially one band in the electrophoretic gel. For proteins that can be modified (e.g., phosphorylation or glycosylation), different modifications can produce different isolated proteins that can be purified separately.
[0099] "Separated polynucleotide" means a nucleic acid molecule that does not contain a gene side-linked in the naturally occurring genome of the organism from which the nucleic acid molecule of this disclosure is derived. Therefore, this term includes, for example, recombinant DNA incorporated into a vector; incorporated into an autonomously replicating plasmid or virus; or incorporated into the genomic DNA of a prokaryote or eukaryote; or present as a separate molecule independent of other sequences (e.g., cDNA or genomic or cDNA fragments produced by PCR or restriction endonuclease digestion). Furthermore, this term includes RNA molecules transcribed from DNA molecules and recombinant DNA that is part of a hybrid gene encoding an additional polypeptide sequence.
[0100] "Isolated polypeptide" means a polypeptide of this disclosure that has been isolated from its naturally occurring associated components. Typically, a polypeptide is considered isolated when it contains at least 60% by weight of proteins and naturally occurring organic molecules that are naturally associated with it. In one embodiment, the formulation is at least 75%, at least 90%, and at least 99% by weight of the polypeptide of this disclosure. The isolated polypeptide of this disclosure can be obtained, for example, by extraction from a natural source, by expression of a recombinant nucleic acid encoding such polypeptide, or by chemical synthesis of the protein. Purity can be measured by any suitable method (e.g., column chromatography, polyacrylamide gel electrophoresis, or analysis by HPLC).
[0101] As used herein, the term "joint" refers to a molecule that connects two parts. In one embodiment, the term "joint" refers to a covalent joint ( For example (covalent bond) or non-covalent joint.
[0102] "Marker" refers to any protein or polynucleotide that has alterations in expression, level, structure, or activity associated with a disease or condition.
[0103] As used herein, the term "mutation" refers to the substitution of a residue within a sequence (e.g., a nucleic acid or amino acid sequence) by another residue, or the deletion or insertion of one or more residues within a sequence. Mutations are generally described herein by identifying the original residue, then identifying the position of said residue within the sequence, and then identifying the identity of the newly substituted residue. Various methods for performing the amino acid substitutions (mutations) described herein are well known in the art and have been employed by, for example, Green and Sambrook. Molecular Cloning: A Laboratory Manual (4th edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)) Available.
[0104] The terms "nick guide RNA (nickRNA; nsgRNA; ngRNA; nRNA)" or "nicking guide RNA (nickRNA; nsgRNA; ngRNA; nRNA)" refer to a guide polynucleotide that contains a spacer region and a scaffold sequence and is capable of guiding the target site in the nick target polynucleotide of a lead editor's nucleic acid programmable DNA-binding (napDNAbp) protein domain. The nick guide RNA does not contain the intended nucleotide editing for incorporation into the double-stranded target polynucleotide. In some cases, the target polynucleotide is double-stranded DNA.
[0105] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to molecules containing nucleobases and an acidic portion. For example Compounds that are polymers of nucleosides, nucleotides, or nucleotides. Typically, polymeric nucleic acids (...) For example Nucleic acid molecules (containing three or more nucleotides) are linear molecules in which adjacent nucleotides are linked together by phosphodiester bonds. In some embodiments, "nucleic acid" refers to a single nucleic acid residue (…). For example Nucleotides and / or nucleosides). In some embodiments, "nucleic acid" refers to an oligonucleotide chain containing three or more individual nucleotide residues. As used herein, the terms "oligonucleotide" and "polynucleotide" are used interchangeably and refer to polymers of nucleotides ( For example (A string of at least three nucleotides). In some implementations, "nucleic acid" encompasses RNA as well as single-stranded and / or double-stranded DNA. Nucleic acids can be naturally occurring, for example, in the case of genomes, transcripts, mRNA, tRNA, rRNA, siRNA, snRNA, plasmids, viscera, chromosomes, chromatids, or other naturally occurring nucleic acid molecules. On the other hand, nucleic acid molecules can be non-naturally occurring molecules. For example Recombinant DNA or RNA, artificial chromosomes, engineered genomes or fragments thereof, or synthetic DNA, RNA, DNA / RNA hybrids, or including non-naturally occurring nucleotides or nucleosides. Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms include nucleic acid analogues. For example Analogs containing components other than the phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems, and optionally purified or chemically synthesized. wait. Under appropriate circumstances, For example In the case of chemically synthesized molecules, nucleic acids contain nucleoside analogs, such as analogs with chemically modified bases or sugars and backbone modifications. Unless otherwise stated, nucleic acid sequences are presented in a 5' to 3' orientation. In some embodiments, the nucleic acid is or contains a natural nucleoside ( For example Adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine; nucleoside analogues ( For example 2-Aminoadenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyluridine, C5-propynylcytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazoadenosine, 7-deazoguanosine, 8-oxadenosine, 8-oxazoguanosine, O(6)-methylguanine and 2-thiocytidine); chemically modified bases; biologically modified bases ( example like methylated bases); inserted bases; modified sugars ( For example 2'-Fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups ( For example Thiophosphates and 5'- N -phosphamide bond).
[0106] The terms "nuclear localization sequence," "nuclear localization signal," or "NLS" refer to an amino acid sequence that facilitates the importation of proteins into the cell nucleus. Nuclear localization sequences are known in the art and described, for example, in Plank. et al. International PCT application PCT / EP2000 / 011690, filed on November 23, 2000 and published on May 31, 2001 as WO / 2001 / 038547, is incorporated herein by reference in connection with the disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, for example, via Koblan. et al.As described in Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises the amino acid sequence KRTADGSEFESPKKKRKV (SEQ ID NO: 190), KRPAATKKAGQAKKKK (SEQ ID NO: 191), KKTELQTTNAENKTKKL (SEQ ID NO: 192), KRGINDRNFWRGENGRKTR (SEQ ID NO: 193), RKSGKIAAIVVKRPRK (SEQ ID NO: 194), PKKKRKV (SEQ ID NO: 195), MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 196), PKKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 328), or RKSGKIAAIVVKRPRKPKKKRKV (SEQ ID NO: 329).
[0107] The terms “nucleobase,” “nitrogenous base,” or “base,” used interchangeably herein, refer to nitrogenous biological compounds that form nucleosides, which are components of nucleotides. The ability of nucleobases to form base pairs and stack on each other directly results in long-chain helical structures, such as ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). Five nucleobases—adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U)—are referred to as primary or canonical. Adenine and guanine are derived from purines, while cytosine, uracil, and thymine are derived from pyrimidines. DNA and RNA may also contain other modified (non-basic) bases. Modified nucleobases, in non-limiting examples, may include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydroxymethylcytosine. Hypoxanthine and xanthine can be produced in the presence of mutagens; both are produced through deamination (replacing an amine group with a carbonyl group). Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can be produced by deamination of cytosine. A nucleoside consists of a nucleobase and a pentose sugar (ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine. Examples of nucleosides with modified nucleobases include inosine (I), xanthine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), and pseudouridine (Ψ). A nucleotide consists of a nucleobase, a pentose sugar (ribose or deoxyribose), and at least one phosphate group. Non-limiting examples of chemical modifications that may be included in the modified nucleobases and / or modified nucleobases are as follows: pseudouridine, 5-methyl-cytosine, 2'- O 2'-methyl-3'-phosphonoacetate, 2'- O -MethylthioPACE (MSP), 2'- O 2'-methyl-PACE (MP), 2'-fluoroRNA (2'-F-RNA), restricted ethyl (S-cEt), 2'- O -Methyl ('M'), 2'-O-methyl-3'-thiophosphate ('MS'), 2'- O 3'-methyl-3'-thiophosphonate ('MSP'), 5-methoxyuridine, thiophosphate and N1-methylpseudouridine.
[0108] The terms "nucleic acid programmable DNA-binding protein" or "napDNAbp" are used interchangeably with "polynucleotide programmable nucleotide-binding domain" to refer to a protein that binds to nucleic acids (...). For example Proteins associated with DNA or RNA, such as guide nucleic acids or guide polynucleotides (RNAs) that direct the nap DNA bp to a specific nucleic acid sequence. For example,(gRNA). In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable DNA-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable RNA-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a Cas9 protein. The Cas9 protein can associate with guide RNA that directs the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, napDNAbp is a Cas9 domain, such as an active nuclease Cas9, a Cas9 nickase (nCas9), or an inactive nuclease Cas9 (dCas9). Non-limiting examples of nucleic acid-programmable DNA-binding proteins include Cas9 ( For exampleCas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ (Cas12j / Casphi). Non-restrictive examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, and Cas9. (Also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Cas12j / CasΦ, Cpf1, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm 6. Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins, CARF, DinG, their homologues, or modified or engineered forms thereof. Other nucleic acid-programmable DNA-binding proteins are also within the scope of this disclosure, although they may not be specifically listed herein. See also, For example Makarova et al. “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR-J October 2018; 1:325-336. doi:10.1089 / crispr.2018.0033; Yan et al. , “Functionally diverse type V CRISPR-Cassystems” ScienceJanuary 4, 2019; 363(6422):8891. doi: 10.1126 / science.aav7271, the entire contents of each of the aforementioned documents are incorporated herein by reference. Exemplary nucleic acid programmable DNA-binding proteins and nucleic acid sequences encoding nucleic acid programmable DNA-binding proteins are provided in the sequence listing as SEQ ID NO: 197-231, 232-245, 254-257, 260, and 378. In some embodiments, napDNAbp is a (CRISPR-related system) Cas9 endonuclease, for example, derived from Streptococcus pyogenes Cas9 (Csnl) (e.g., SEQ ID NO: 197), from Neisseria meningitidis meningitidis) Cas9 (NmeCas9; SEQ ID NO: 208), Nme2Cas9 (SEQ ID NO: 209), Constellation Chain Streptococcus constellatus (ScoCas9) or its derivatives (e.g., sequences having at least about 85% sequence identity with Cas9, such as Nme2Cas9 or spCas9). Other non-limiting examples of nucleic acid-programmable DNA-binding proteins include Rufflow. et al. , “Design of highly functional genome editors by modeling of the universe of CRISPR-Cas Sequences,” bioRxiv Those published or referenced in doi:10.1101 / 2024.04.22.590591 on April 22, 2024, whose disclosures are incorporated herein by reference in their entirety for all purposes, are designed using artificial intelligence. In some embodiments, napDNAbp is OpenCRISPR-1 or a variant thereof (e.g., a variant containing a D10A amino acid change and / or lacking an N-terminal methionine).
[0109] As used herein, the terms "nucleobase editing domain" or "nucleobase editing protein" refer to proteins or enzymes that catalyze nucleobase modifications in RNA or DNA, such as deamination from cytosine (or cytidine) to uracil (or uridine) or thymine (or thymidine), and deamination from adenine (or adenosine) to hypoxanthine (or inosine), as well as the addition and insertion of non-templated nucleotides. In some embodiments, the nucleobase editing domain is a deaminase domain ( For example (Adenine deaminase or adenosine deaminase; or cytidine deaminase or cytosine deaminase).
[0110] As used herein, “acquisition” in “acquisition agent” includes the synthesis, purchase or other acquisition of the agent.
[0111] "OpenCRISPR-1 peptide" refers to a protein having an amino acid sequence that is at least about 85% identical to that of SEQ ID NO: 443, or a fragment thereof associated with a nucleic acid (such as a guide nucleic acid or guide polynucleotide that directs nap DNAbp to a specific nucleic acid sequence). Further details relating to the OpenCRISPR-1 peptide are disclosed on Rufflow. et al. , “Design of highly functional genome editors by modeling of the universe ofCRISPR-Cas Sequences,” bioRxiv Published on April 22, 2024, doi: 10.1101 / 2024.04.22.590591, the contents of which are hereby incorporated herein by reference in their entirety for all purposes.
[0112] "OpenCRISPR-1 polynucleotide" refers to a nucleic acid molecule or fragment thereof encoding the OpenCRISPR-1 polypeptide and its introns, exons, 3' untranslated region, 5' untranslated region, and regulatory sequences associated with its expression. In embodiments, the OpenCRISPR-1 polynucleotide is a genomic sequence, cDNA, mRNA, or gene associated with and / or required for OpenCRISPR-1 expression. An exemplary OpenCRISPR-1 nucleotide sequence is provided in SEQ ID NO: 444.
[0113] In various implementations, the guide RNA suitable for use in combination with the OpenCRISPR-1 peptide contains a scaffold with at least 85% sequence identity to a nucleotide sequence selected from the following or a fragment thereof capable of binding to the OpenCRISPR-1 peptide: GUUUUAGAGCUGUGUUGAAAAACACAGCAAGUUAAAAUAAGGCUUUGUCCGUAUCCAACUUGAAAAAGUGAGCACCGAUUCGGUGC (SEQ ID NO: 445); GUUUUAGAGCUGGAAACAGCAAGUUAAAAUAAGGCUUUGUCCGUAUCCAACUUGAAAAAGUGAGCACCGAUUCGGUGC (SEQ ID NO: 446); and GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 447).
[0114] "Subject" or "patient" means mammal, including but not limited to humans or non-human mammals. In one implementation, mammal is a bovine, equine, canine, sheep, rabbit, rodent, non-human primate, or feline. In one implementation, "patient" refers to a mammalian subject with a higher-than-average likelihood of developing a disease or symptom. Exemplary patients may be humans, non-human primates, cats, dogs, pigs, cattle, horses, camels, llamas, goats, sheep, rodents (…). For example (mice, rabbits, rats, or guinea pigs) and other mammals that may benefit from the treatments disclosed herein. Exemplary human patients may be male and / or female.
[0115] "Patients in need" or "subjects in need" in this article refers to patients who are diagnosed with, at risk of or have, pre-determined to have or are suspected of having a disease or condition.
[0116] The terms "pathogenic mutation," "pathogenic variant," "disease-causing mutation," "deleterious mutation," or "predisposing mutation" refer to genetic alterations or mutations that are associated with or increase an individual's susceptibility to or predisposition to a disease or condition. In some embodiments, a pathogenic mutation includes the substitution of at least one wild-type amino acid for at least one pathogenic amino acid in a protein encoded by a gene. In some embodiments, the pathogenic mutation is located in a termination region (…). For example In some implementations, the pathogenic mutation is located in the non-coding region (stop codon). For example (including introns, promoters, etc.)
[0117] "Lead editor (PE)" refers to a polypeptide or polypeptide component involved in lead editing. In various embodiments, the lead editor contains a polypeptide domain with DNA-binding activity. For example DNA-binding domain and polypeptide domain with DNA polymerase activity. For example (DNA polymerase domain). In the implementation scheme, the pilot editor contains Streptococcus pyogenesThe Cas9 protein domain or its functional fragments or variants, and those derived from retroviruses ( For example The reverse transcriptase domain or a functional fragment or variant thereof of Moloney murine leukemia virus. A lead editor and its usage are described in International Patent Application Publication No. WO 2023 / 283092, the disclosure of which is incorporated herein by reference in its entirety for all purposes.
[0118] The term "lead editing" refers to the programmable editing of target DNA using a lead editor complexed with lead editing guide RNA (PEgRNA) to incorporate desired nucleotide changes into the target DNA through target-lead DNA synthesis. The target DNA polynucleotides in the lead edit... For example The target gene may comprise a double-stranded DNA molecule with two complementary strands: a first strand, which may be referred to as the “target strand” or “non-editing strand,” and a second strand, which may be referred to as the “non-target strand” or “editing strand.” In some embodiments, in the lead editing guide RNA (PEgRNA), the spacer sequence is complementary or substantially complementary to a specific sequence on the target strand (referred to as the “search target sequence”). In some embodiments, the spacer sequence anneals to the target strand at the search target sequence. The target strand may also be referred to as the “non-prototype spacer adjacent motif (non-PAM strand).” In some embodiments, the non-target strand may also be referred to as the “PAM strand.” In some embodiments, the PAM strand comprises the prototype spacer sequence and an optional prototype spacer adjacent motif (PAM) sequence. In lead editing using a Cas protein-based lead editor, the PAM sequence refers to a short DNA sequence on the PAM strand of the target gene immediately adjacent to the prototype spacer sequence. The PAM sequence can be programmed by a DNA-binding protein (…). For example Cas nickase or Cas nuclease) specific recognition. In some implementations, a specific PAM is a specific programmable DNA-binding protein (PAM). For example Cas cleavage enzyme or Cas nuclease For example The characteristics of Cas9 cleavage enzyme or Cas9 nuclease. The prototype spacer sequence refers to the double-stranded target DNA (Cas9 cleavage enzyme or Cas9 nuclease). For example A specific sequence in the PAM strand of the target gene that is complementary to the target sequence. In PEgRNA, the spacer region sequence may have a complementary sequence to the double-stranded target DNA (the target gene). For example The prototypical spacer sequence on the editing strand of the target gene is essentially the same as the sequence, except that the spacer sequence may contain uracil (U), while the prototypical spacer sequence may contain thymine (T).
[0119] The term "lead editing guidance RNA" or "PEgRNA" refers to a guide polynucleotide containing one or more intended nucleotide edits incorporated into a double-stranded target polynucleotide. In some cases, the target polynucleotide is double-stranded DNA. In some embodiments, the PEgRNA is associated with a lead editor and guides the lead editor to incorporate one or more intended nucleotide edits into the double-stranded target DNA via lead editing. For example , target genes) in.
[0120] The term "lead editing system" or "lead editor system" refers to an intermolecular complex used for editing the nucleobases of a target nucleotide sequence. In various embodiments, the lead editor (PE) system comprises (1) a domain containing DNA-binding activity ( For example Nucleic acid programmable DNA-binding proteins) and polymerases ( For example A fusion protein of DNA polymerase or reverse transcriptase; and (2) one or more guide polynucleotides (DNA polymerase or reverse transcriptase) that bind to a domain having DNA binding activity. For example , lead editing guide RNA and / or nickRNA (“nickRNA”).
[0121] The terms “protein,” “peptide,” “polypeptide,” and their grammatical equivalents are used interchangeably herein and refer to a polymer of amino acid residues linked together by peptide (amide) bonds. Proteins, peptides, or polypeptides can be naturally occurring, recombinant, or synthetic, or any combination thereof.
[0122] As used in this article, the term "fusion protein" refers to a hybrid polypeptide that contains protein domains from at least two different proteins.
[0123] As used herein in the context of proteins or nucleic acids, the term "recombinant" refers to a protein or nucleic acid that is not found in nature but is an engineered product. For example, in some embodiments, the recombinant protein or nucleic acid molecule contains at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutated amino acid or nucleotide sequences compared to any naturally occurring sequence.
[0124] The term "reduction" refers to a negative change of at least 10%, 25%, 50%, 75%, or 100%.
[0125] The term "reference" refers to a standard or control condition. In one embodiment, the reference is wild-type or healthy cells. In other embodiments, and not limitingly, the reference is untreated cells that have not been subjected to the test conditions, or to a placebo or saline, culture medium, buffer, and / or a control vector that does not contain the polynucleotide of interest. In some cases, the reference is unedited cells. In some cases, the reference is a subject who has not been given the edited cells of this disclosure.
[0126] A “reference sequence” is a defined sequence used as the basis for sequence comparison. The reference sequence can be a subset or the entirety of a specified sequence; for example, a segment of a full-length cDNA or gene sequence, or the complete cDNA or gene sequence. For peptides, the reference peptide sequence will typically be at least about 16 amino acids, at least about 20 amino acids, at least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids in length. For nucleic acids, the reference nucleic acid sequence will typically be at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides, or about 300 nucleotides in length, or any integer near or between these values. In some embodiments, the reference sequence is the wild-type sequence of the protein of interest. In other embodiments, the reference sequence is a polynucleotide sequence encoding a wild-type protein.
[0127] The terms "RNA-programmable nuclease" and "RNA-directed nuclease" refer to enzymes that form complexes with one or more RNA molecules that are not cleavage targets. For example RNA-programmable nucleases (which bind or associate with RNA). In some embodiments, when in a complex with RNA, the RNA-programmable nuclease may be referred to as a nuclease-RNA complex. Typically, the bound RNA is referred to as guide RNA (gRNA). In some embodiments, the RNA-programmable nuclease is a Cas9 endonuclease (from a CRISPR-related system), for example, derived from... Streptococcus pyogenes Cas9 (Csnl) (e.g., SEQ ID NO: 197), from Neisseria meningitidis Cas9 (NmeCas9; SEQ ID NO: 208), Nme2Cas9 (SEQ ID NO: 209), Constellation Streptococcus (ScoCas9) or its derivatives (e.g., sequences that have at least about 85% sequence identity with Cas9, such as Nme2Cas9 or spCas9).
[0128] As used herein, the term "scFv" refers to a single-chain Fv antibody in which the variable domains of the heavy and light chains of the antibody have joined together to form a single chain. An scFv fragment contains a single polypeptide chain that includes a variable region (VL) of the antibody light chain separated by a linker. For exampleCDR-L1, CDR-L2 and / or CDR-L3) and the variable region (VH) of the antibody heavy chain ( For example (CDR-H1, CDR-H2, and / or CDR-H3). The linkers that connect the VL and VH regions of the scFv fragment can be peptide linkers composed of proteogenic amino acids. Alternative linkers can be used to increase the resistance of the scFv fragment to proteolytic degradation (e.g., linkers containing D-amino acids), to enhance the solubility of the scFv fragment (e.g., hydrophilic linkers, such as linkers containing polyethylene glycol or peptides containing repeating glycine and serine residues), to improve the biophysical stability of the molecule (e.g., linkers containing cysteine residues that form intramolecular or intermolecular disulfide bonds), or to reduce the immunogenicity of the scFv fragment (e.g., linkers containing glycosylation sites). Those skilled in the art will also understand that the variable regions of the scFv molecules described herein can be modified such that they differ in amino acid sequence from the antibody molecules from which they are derived. For example, nucleotide or amino acid substitutions that result in conserved substitutions or alterations at amino acid residues can be made. For example (in CDR and / or framework residues) in order to preserve or enhance the ability of scFv to bind to antigens recognized by the corresponding antibody.
[0129] "Selective binding" means that the CD45 peptide specifically binds to the wild-type form, but exhibits reduced or no binding to the mutated CD45 peptide. In the embodiments, the mutations in the CD45 peptide are selected from those provided herein.
[0130] The term "single nucleotide polymorphism (SNP)" refers to a variation of a single nucleotide occurring at a specific location in the genome, where each variation is present to a considerable degree in the population. For example, >1%). SNPs can be located in the coding region of a gene, the non-coding region of a gene, or the intergenetic region (the region between genes). In some implementations, due to the degeneracy of the genetic code, SNPs within the coding sequence do not necessarily change the amino acid sequence of the resulting protein. There are two types of SNPs in coding regions: synonymous and non-synonymous SNPs. Synonymous SNPs do not affect the protein sequence, while non-synonymous SNPs change the amino acid sequence of the protein. There are two types of non-synonymous SNPs: missense and nonsense. SNPs not in protein-coding regions can still affect gene splicing, transcription factor binding, messenger RNA degradation, or non-coding RNA sequences. Gene expression affected by such SNPs is called an eSNP (expressed SNP) and can be located upstream or downstream of the gene. Single nucleotide variants (SNVs) are variations of a single nucleotide, have no frequency limit, and can occur in somatic cells. Somatic single nucleotide variants are also called single nucleotide alterations.
[0131] "Specific binding" means a nucleic acid molecule, polypeptide, polypeptide / polynucleotide complex, compound, or molecule that recognizes and binds to the polypeptides and / or nucleic acid molecules of this disclosure but substantially does not recognize and bind to other molecules in a sample (e.g., a biological sample).
[0132] "Substantially identical" means that the polypeptide or nucleic acid molecule exhibits at least 50% identity with a reference amino acid sequence. In one embodiment, the reference sequence is a wild-type amino acid or nucleic acid sequence. In another embodiment, the reference sequence is any of the amino acid or nucleic acid sequences described herein. In one embodiment, such a sequence is at least about 60%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or even 99.99% identical to the sequence used for comparison at the amino acid or nucleic acid level.
[0133] Sequence identity is typically measured using sequence analysis software (e.g., Sequence Analysis Software Package of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs). This type of software assigns homology levels to various substitutions, deletions, and / or other modifications. Conserved substitutions typically include those within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine.
[0134] Nucleic acid molecules that can be used in the methods of this disclosure include any nucleic acid molecule encoding the polypeptide or a functional fragment thereof of this disclosure. Such nucleic acid molecules do not need to be 100% identical to an endogenous nucleic acid sequence but generally exhibit substantially similarity. Polynucleotides having "substantially similarity" to an endogenous sequence are generally capable of hybridizing with at least one strand of a double-stranded nucleic acid molecule. "Hybridization" means pairing under various stringent conditions to achieve hybridization with complementary polynucleotide sequences (…). For example (The genes described in this article) or portions thereof form double-stranded molecules. (See also: For exampleWahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, AR (1987) Methods Enzymol. 152:507).
[0135] "Split" means to divide into two or more segments.
[0136] A “split polypeptide” or “split protein” refers to a protein provided as an N-terminal fragment and a C-terminal fragment, which translate into two separate polypeptides derived from a nucleotide sequence. The polypeptides corresponding to the N-terminal and C-terminal portions of the split protein can be spliced in some embodiments to form a “reconstructed” protein. In embodiments, the split polypeptide is a nucleic acid-programmable DNA-binding protein (e.g., Cas9) or a base editor.
[0137] The term "target site" refers to a nucleotide sequence or nucleobase of interest within a modified nucleic acid molecule. In embodiments, the modification is the deamination of a base. The deaminase may be cytidine or adenine deaminase. Fusion proteins or base editing complexes containing deaminases may include dCas9-adenosine deaminase fusion proteins, Cas12b-adenosine deaminase fusions, or base editors disclosed herein.
[0138] As used herein, the term "treatment" (such as "treat / treating / treatment") refers to reducing or improving a condition and / or related symptoms, or achieving the desired pharmacological and / or physiological effect. It will be understood that treating a condition or disorder does not require the complete elimination of the condition, disorder, or related symptoms, although this is not excluded. In some implementations, the effect is therapeutic. Right now (But not limited to) the effects partially or completely reduce, weaken, eliminate, alleviate, relieve, decrease the intensity of the disease and / or adverse symptoms attributable to the disease, or cure the disease and / or adverse symptoms attributable to the disease. In some embodiments, the effect is preventative. Right now The effect is to protect against or prevent the occurrence or recurrence of disease or ailment. To this end, the method disclosed in this invention includes administering a therapeutically effective amount of the composition as described herein.
[0139] The term "uracil glycosylase inhibitor" or "UGI" refers to an agent that inhibits the uracil excision repair system. A base editor containing cytidine deaminase converts cytosine to uracil, which is then converted to thymine via DNA replication or repair. In various embodiments, uracil DNA glycosylase (UGI) prevents the base excision repair that converts U back to C. In some cases, contacting cells and / or polynucleotides with the UGI and the base editor prevents the base excision repair that converts U back to C. An exemplary UGI contains the following amino acid sequence: >splP14739IUNGI_BPPB2 Uracil-DNA Glycosyltransferase Inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (SEQ ID NO: 231).
[0140] In some implementations, the agent that inhibits the uracil excision repair system is uracil stable protein (USP). See, for example, WO 2022015969 A1, which is incorporated herein by reference.
[0141] This refers to the method of introducing nucleic acid molecules into cells to obtain transformed cells. Vectors include plasmids, transposons, bacteriophages, viruses, liposomes, lipid nanoparticles, and episomes.
[0142] The ranges provided herein should be understood as abbreviations of all values within the range. For example, the range 1 to 50 is understood to include any number, combination of numbers, or subrange that comes from a group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.
[0143] Any definition of a variable in this document, including a statement of a list of chemical groups, includes defining this variable as any single group or a combination of the listed groups. References to embodiments of a variable or aspect herein include embodiments as a single embodiment or as a combination of any other embodiment or part thereof.
[0144] All terms should be understood as they would be understood by one of ordinary skill in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0145] In this application, the use of the singular includes the plural unless otherwise specifically stated. It must be noted that, as used in the specification, the singular forms “a” and “the” include the plural referent unless the context clearly specifies otherwise. In this application, unless otherwise stated, the use of “or” means “and / or”. Furthermore, the use of the term “including” and other forms such as “include,” “includes,” and “included” is not limited.
[0146] As used in this specification and claims, the terms “comprising” (and any form of inclusion, such as “comprise” and “comprises”), “having” (and any form of having, such as “have” and “has”), “including” (and any form of inclusion, such as “includes” and “include”), or “containing” (and any form of containing, such as “contains” and “contain”) are inclusive or open-ended. This wording indicates the presence of a specified element, feature, component, and / or method step, but does not exclude the presence of other elements, features, components, and / or method steps. In some embodiments, any embodiment designated as “comprising” a particular component or element is contemplated as also “consisting of a particular component or element” or “substantially consisting of a particular component or element.” It is contemplated that any embodiment discussed in this specification can be implemented with respect to any method or composition of this disclosure, and vice versa. Furthermore, the compositions of this disclosure can be used to implement the methods of this disclosure.
[0147] The terms "about" or "approximately" mean that a particular value, as determined by a person skilled in the art, is within an acceptable margin of error, which will depend in part on the method of measurement or determination of the value. Right now Limitations of the measurement system.
[0148] The use of terms such as "some implementation schemes," "implementation schemes," "one implementation scheme," or "other implementation schemes" in the specification means that a particular feature, structure, or characteristic described in connection with an implementation scheme is included in at least some implementation schemes, but not necessarily in all implementation schemes. Attached Figure Description
[0149] Figure 1A and 1B Schematic diagrams and bar graphs related to editing of mAb030-9 in cell lines are provided. Figure 1BIn the diagram, each left-hand bar corresponds to D3 (day 3 after electroporation (EP), and each right-hand bar corresponds to D7 (day 7 after EP). Figure 1A and 1B In Chinese, "NGS" stands for Next Generation Sequencing.
[0150] Figures 2A to 2C Charts and bar graphs related to the editing of mAb030-9 in MOLM13 are provided.
[0151] Figure 3 Scatter plots and schematic diagrams are provided illustrating mAb030-9 antibody escape observed based on Y232C and K231R / Y232C mutations. CD34-guided mutation frequency: Y232C = 56.02%, K231R / Y232C = 10.24%. Combinatorial escape editing = 66.26%. Both Y232C and K231R / Y232C showed escape from mAb030-9 and represent a high frequency of combinatorial editing in CD34+ cells.
[0152] Figure 4 Scatter plots and schematic diagrams are provided illustrating the escape of mAb030-9 antibody based on the E259G mutation. CD34-guided mutation frequency: E259G = 63.6%; silencing = 28.71%. 63.6% escape editing. E259G exhibits escape from the mAb030-9 antibody and represents a high editing frequency in CD34+ cells.
[0153] Figure 5 Scatter plots and schematic diagrams are provided showing the mAb030-9 antibody escape observed in the transient pool based on N255G / E256G. Guided mutation frequencies: N255G = 48.79%; N255D = 1.03%; N255G / E256G = 3.36%. 3.36% escape edit. N255G / E256G has some effect.
[0154] Figures 6A to 6C A graph illustrating the binding kinetics of mAb030-9 to a CD45 polypeptide containing specified amino acid substitutions is provided. Figures 6A to 6C In this context, the term "PBD" stands for phosphate-buffered saline, and "WT" indicates wild-type CD45 peptide. Figures 6A to 6CEach data point was generated by immobilizing 100 nM mAB030-9 onto a biosensor equipped with an anti-human IgG Fc capture (AHC) antibody, contacting the immobilized mAB030-9 with 300 nM of a specified mutant or wild-type CD45 protein, and then replacing the CD45 protein solution with PBS at the times indicated by the dashed vertical lines. CD45 peptides containing amino acid substitutions of E259R, E259K, or N257S, relative to WT CD45, all showed reduced binding kinetics with mAB030-9, indicating that editing resulted in a decreased binding affinity of the CD45 mutant to mAB030-9. All CD45 variants exhibited responses below 0.05 nm.
[0155] Figure 7 A graph showing the geometric mean fluorescence intensity of cells expressing a specified CD45 variant and immunostained with a labeled mAb030-9 peptide is provided. Cells were electroporated using a base editor system (see Tables 10–15) represented along the x-axis, and measurements were taken on day 7 post-electroporation. "Simulated" indicates cells that were simulated for editing and thus expressing WT CD45. EP9–EP15 showed a reduced MFI compared to the simulated version, indicating that the mutation reduced the binding affinity to mAb030-9 and was mutated with sufficient efficiency for a detectable percentage of cells to have been edited. Detailed Implementation
[0156] This disclosure provides a base editing strategy targeting CD45 cell surface proteins, which can be used for non-toxic conditioning of subjects.
[0157] Various aspects of this disclosure are based, at least in part, on the discovery that base editing can be used to edit CD45 polynucleotides, thereby resulting in alterations in the binding of anti-CD45 antibodies, antibody-drug conjugates (ADCs), or chimeric antigen receptor (CAR)-T cells to cells containing the edited CD45 polynucleotides. In embodiments, editing of the CD45 polynucleotides alters the binding of anti-CD45 antibodies to the CD45 peptide, for example, by altering epitopes present on the CD45 peptide. In embodiments, editing of the CD45 polynucleotides reduces or eliminates CD45 expression. In some embodiments, this alteration in the CD45 peptide does not alter the biological activity of the CD45 peptide.
[0158] In one instance, an anti-CD45 antibody is used to target and ablate CD45-expressing cells present in the subject. Ablation can be performed before, during, or after transplantation of HSCs containing edited CD45 polynucleotides into the subject, such that the transplanted HSCs are not targeted by the antibody. In one embodiment, the anti-CD45 antibody does not bind to or depletes the CD45-edited cells. Therefore, cells containing edited CD45 polynucleotides can be used in the presence of a CD45 antibody (e.g., a CD45 antibody-drug conjugate). In the body Amplification.
[0159] The conditioning method disclosed herein has several advantages. The method described herein provides selective targeting of endogenous HSCs while preserving the edited HSCs. Therefore, antibody or ADC therapy can be continued after HSCT to... In the body Expanding gene-edited cells. In the implementation scheme, the base-edited HSC cells are disease-independent, and therefore, it is advantageous to expand base-edited HSC cells in subjects to replace pathogenic cells with non-pathogenic base-edited HSC cells. The edited cells allow for the administration of antibodies or ADCs without toxicity risks, and there is no need to ensure that the ADC has a short serum half-life. This makes it possible to use antibodies with longer half-lives and simplifies ADC development. Non-limiting examples of ADCs for the treatment of multiple sclerosis include those provided herein. Clinical trial design is also simplified—HSCs can be infused before or concurrently with conditioning with little or no risk of depletion. The method provides benefit to all patients, regardless of their immune status.
[0160] In one implementation, CD45 is altered in the cells used for transplantation. For example(Base editing) prevents the binding of anti-CD45 antibodies without interfering with the normal function of CD45 in regulating signal transduction in hematopoiesis. Base editing generates nucleobase changes to create amino acid substitutions in CD45. Administration of CD45-targeting anti-CD45 antibody-drug conjugates to patients with native CD45-expressing cells depletes the patient's HSCs and progenitor cells (opsonization). Autologous gene-edited HSCs can be transplanted into the patient. The gene-edited cells compete with residual host HSCs for bone marrow refill (BM). Anti-CD45 antibody drugs target cells with unedited CD45 but cannot bind to HSCs with edited CD45. Therefore, anti-CD45 antibodies target the subject's native HSCs but not gene-edited HSCs. Thus, both cell types express CD45 peptides that play a role in signal transduction. However, the binding of the antibody-drug conjugate to native CD45-expressing cells leads to cell death. In contrast, anti-CD45 antibody-drug conjugates do not target gene-edited HSCs because the amino acid substitutions introduced into CD45 prevent the binding of anti-CD45 antibodies, but do not interfere with the normal function of CD45.
[0161] Methods for identifying candidate agents for the selective depletion or ablation of endogenous stem cell populations are also within the scope of this disclosure. Such methods may include the following steps: (a) mixing a sample containing a stem cell population with a test reagent ( For example (a) Contacting the sample with an antibody or antibody-drug conjugate; and (b) detecting whether one or more cells of a stem cell population have been depleted or ablated from the sample; wherein depletion or ablation of one or more cells of the stem cell population after the contacting step identifies the test agent as a candidate. In a further step, the edited cells are identified as not having been similarly depleted or ablated by the agent. In some embodiments, the cells are contacted with the test agent for at least about 2-24 hours.
[0162] Editing of target genes To produce the gene editing described herein, cells (e.g., hematopoietic stem cells (HSCs)) are collected from a subject and contacted with one or more guide RNAs and a nucleobase editor polypeptide comprising a nucleic acid-programmable DNA-binding protein (napDNAbp) and a cytidine deaminase or adenosine deaminase, or comprising one or more deaminases. In some embodiments, the cells to be edited are contacted with at least one nucleic acid, wherein said at least one nucleic acid encodes one or more guide RNAs and a nucleobase editor polypeptide comprising a nucleic acid-programmable DNA-binding protein (napDNAbp) and a cytidine deaminase. In some embodiments, the gRNA comprises a nucleotide analog. In some cases, the gRNA is added directly to the cells. These nucleotide analogs can inhibit the degradation of the gRNA during cellular processes.
[0163] In some implementations, lead editing is used to modify immune cells. Methods for using lead editing to edit polynucleotide sequences are well known in the art (see, for example, Petrova IO, Smirnikhina SA. The Development, Optimization and Future of Prime Editing). Int J Mol Sci . 1 December 2023;24(23):17045. doi: 10.3390 / ijms242317045, the contents of which are incorporated herein by reference in their entirety for all purposes. In the presence or absence of secondary cut guide RNA (nRNA), lead editor constructs and paired lead editing guide RNAs (pegRNAs) can induce targeted, programmable changes in genomic DNA.
[0164] In various cases, it is advantageous for the spacer sequence to contain a 5' and / or 3' "G" nucleotide. In some cases, for example, any spacer sequence or guide polynucleotide provided herein contains or also contains a 5' "G", wherein, in some embodiments, the 5' "G" is complementary to or not complementary to the target sequence. In some embodiments, a 5' "G" is added to a spacer sequence that does not already contain a 5' "G". For example, when the guide RNA is expressed under the control of the U6 promoter, etc., it may be advantageous for the guide RNA to contain a 5' terminal "G" because the U6 promoter prefers "G" at the transcription start site. See Cong, L. et al. “Multiplex genome engineering using CRISPR / Cas systems. Science 339:819-823 (2013) doi:10.1126 / science.1231143). In some cases, a 5' G is added to the guide polynucleotide that will be expressed under promoter control, but optionally, a 5' G is not added to the guide polynucleotide if or when the guide polynucleotide is not expressed under promoter control.
[0165] Exemplary guidance sequences are provided in Tables 1 and 2.
[0166] Consider the variants of the spacer sequence listed in Table 1 or 2 that contain 1, 2, 3, 4, or 5 nucleobase alterations. For example, variations in the target polynucleotide sequence within a population (e.g., single nucleotide polymorphisms) may require the aforementioned alterations to the spacer sequence to allow the spacer to better bind to the target sequence in the subject.
[0167] Table 1. Representative guide sequences associated with missense mutations identified by yeast display screening.
[0168] Table 1 (continued)
[0169] Table 2. Representative guide sequences. Guides tested on CD34, Jurkat, Molm13, and / or CD34+ hematopoietic stem cell and progenitor (HSPC) cells.
[0170] Table 2 (continued). Nucleotide base editor The nucleobase editors that can be used in the methods and compositions described herein are nucleobase editors that edit, modify, or alter the target nucleotide sequence of polynucleotides. The nucleobase editors described herein typically include a polynucleotide programmable nucleotide-binding domain and a nucleobase editing domain. For example Adenosine deaminase, cytidine deaminase). When bound to a guide polynucleotide ( For example When bound to gRNA, the polynucleotide programmable nucleotide binding domain can specifically bind to the target polynucleotide sequence, thereby positioning the base editor to the target nucleic acid sequence that needs to be edited.
[0171] Multinucleotide programmable nucleotide binding domain Multinucleotide programmable nucleotide binding domain binds to multinucleotides ( For example (RNA, DNA). The multinucleotide programmable nucleotide-binding domain of the base editor can itself contain one or more domains (RNA, DNA). For example (one or more nuclease domains). In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide-binding domain contains an endonuclease or an exonuclease.
[0172] This document discloses a base editor comprising a multinucleotide programmable nucleotide-binding domain containing all or part (e.g., the functional portion) of a CRISPR protein. Right nowA base editor that contains all or part of a CRISPR protein (e.g., a Cas protein) as a domain is also called a "CRISPR protein-derived domain" of the base editor. The CRISPR protein-derived domain incorporated into the base editor can be modified compared to the wild-type or natural form of the CRISPR protein. The CRISPR protein-derived domain may contain one or more mutations, insertions, deletions, rearrangements, and / or recombinations relative to the wild-type or natural form of the CRISPR protein.
[0173] The Cas proteins used in this article include classes 1 and 2. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, and Csb1. , Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, C sd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1 ( example like (SEQ ID NO: 232), Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i and Cas12j / CasΦ, CARF, DinG, their homologues or modified forms thereof. CRISPR enzymes can direct the cleavage of one or both strands at a target sequence, such as within the target sequence and / or within the complementary sequence of the target sequence. For example, a CRISPR enzyme can direct the cleavage of one or both strands within approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500 or more base pairs from the first or last nucleotide of the target sequence.
[0174] Vectors encoding a CRISPR enzyme mutated relative to the corresponding wild-type enzyme can be used, such that the mutated CRISPR enzyme lacks the ability to cleave one or both strands of a target polynucleotide containing the target sequence. Cas protein ( For exampleCas9, Cas12) or Cas structural domain ( For example Cas9, Cas12) can refer to a polypeptide or domain having at least or at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence homology with a wild-type exemplary Cas polypeptide or Cas domain. For example Cas9 and Cas12 can refer to the wild-type or modified form of the Cas protein, which may include amino acid changes such as deletion, insertion, substitution, variant, mutation, fusion, chimera or any combination thereof.
[0175] In some implementations, the CRISPR protein-derived domain of the base editor may include all or part (e.g., the functional portion) of Cas9: Corynebacterium ulcerans (NCBIRef: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBIRef: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref: NC_021284.1); Prevotella intermedia (NCBI Ref: NC_017861.1); Taiwan Spiral Spiroplasma taiwanense (NCBI Ref: NC_021846.1); Streptococcus dolphin iniae) (NCBI Ref: NC_021314.1); Belliella baltica (NCBI Ref: NC_018010.1); Psychroflexus torquis (NCBI Ref: NC_018721.1); Thermophilic chain bacteria (NCBI Ref: YP_820832.1); Harmless Listeria innocua (NCBI Ref: NP_472073.1); Campylobacter jejuni (NCBI Ref: YP_002344900.1); meningitis Neisseria (NCBI Ref: YP_002342100.1), Streptococcus pyogenes or Staphylococcus aureus .
[0176] Some aspects of this disclosure provide high-fidelity Cas9 domains. High-fidelity Cas9 domains are known in the art and described, for example, by Kleinstiver and BP. et al.“High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects.” Nature 529, 490-495 (2016); and Slaymaker, IM et al. “Rationally engineered Cas9 nucleases with improved specificity.” Science 351, 84-88 (2015); the entire contents of each of the above references are incorporated herein by reference. An exemplary high-fidelity Cas9 domain is provided in the sequence listing as SEQ ID NO: 233.
[0177] In some embodiments, any Cas9 fusion protein or complex provided herein contains one or more of the D10A, N497X, R661X, Q695X and / or Q926X mutations or corresponding mutations in any amino acid sequence provided herein, wherein X is any amino acid.
[0178] Typically, Cas9 proteins (such as those from...) Streptococcus pyogenes Cas9 (spCas9) requires a "prototype spacer adjacent motif (PAM)" or PAM-like motif, which is a 2-6 base pair DNA sequence immediately following the DNA sequence targeted by the Cas9 nuclease in the CRISPR bacterial adaptive immune system. The presence of the NGG PAM sequence is essential for binding to specific nucleic acid regions, where "N" in "NGG" is adenosine (A), thymidine (T), or cytosine (C), and "G" is guanosine. In some embodiments, any fusion protein or complex provided herein may contain a motif capable of binding to non-canonical ( For example The Cas9 domain of the nucleotide sequence of the PAM sequence (NGG). Cas9 domains that bind to non-canonical PAM sequences have been described in the art and will be apparent to those skilled in the art. For example, Cas9 domains that bind to non-canonical PAM sequences have been described in Kleinstiver, BP. et al. , “Engineered CRISPR-Cas9 nucleases with altered PAMspecificities” Nature 523, 481-485 (2015); and Kleinstiver, B. P. et al. , “Broadening the targeting range of Staphylococcus aureusCRISPR-Cas9 by modifying PAMrecognition” Nature Biotechnology In 33, 1293-1298 (2015); the entire contents of each are hereby incorporated by reference.
[0179] In some implementations, napDNAbp is a circular mutant ( For example (SEQ ID NO: 238).
[0180] In some implementations, the polynucleotide programmable nucleotide-binding domain includes a nicking enzyme domain. The term "nicking enzyme" herein refers to a polynucleotide programmable nucleotide-binding domain containing a nuclease domain capable of cleaving only double-stranded nucleic acid molecules (…). For example One of the two strands of DNA. For example, when a polynucleotide programmable nucleotide-binding domain contains a nickase domain derived from Cas9, the Cas9-derived nickase domain may contain a D10A mutation and a histidine at position 840. In another example, the Cas9-derived nickase domain contains an H840A mutation, while the amino acid residue at position 10 remains D.
[0181] In some implementations, the Cas9 nuclease is inactive ( For example The inactivated DNA cleavage domain, i.e., Cas9 is a cleavage enzyme, referred to as the "nCas9" protein (for "cleavage enzyme" Cas9; SEQ ID NO: 201). Cas9 cleavage enzymes can be capable of cleaving only double-stranded nucleic acid molecules (…). For example A Cas9 protein comprising one strand of a double-stranded DNA molecule. In some embodiments, the Cas9 nickase comprises at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the same amino acid sequence as any Cas9 nickase provided herein. Further suitable Cas9 nickases will be apparent to those skilled in the art based on this disclosure and knowledge of the art, and are within the scope of this disclosure.
[0182] This article also provides a base editor, which includes catalytic deactivation (…). Right now A multinucleotide-programmable nucleotide-binding domain that cannot cleave the target polynucleotide sequence. For example, in the case of a base editor containing a Cas9 domain, Cas9 may contain D10A and H840A mutations. In other embodiments, the catalytically inactivating multinucleotide-programmable nucleotide-binding domain contains a point mutation ( For example,The deletion of D10A or H840A) and all or part (e.g., the functional portion) of the nuclease domain. The dCas9 domain is known in the art and described in, for example, Qi et al. , “Repurposing CRISPRas an RNA-guided platform for sequence-specific control of gene expression.” Cell The entire contents of the article are incorporated herein by reference in 2013; 152(5):1173-83.
[0183] The term "prototype spacer adjacent motif (PAM)" or PAM-like motif refers to a 2-6 base pair DNA sequence immediately following a DNA sequence targeted by a nucleic acid-programmable DNA-binding protein. In some implementations, the PAM may be a 5' PAM ( Right now (Located upstream of the 5' end of the prototype spacing region). In other embodiments, the PAM can be a 3' PAM ( Right now (Located downstream of the 5' end of the prototype spacer region). The PAM sequence can be any PAM sequence known in the art. Suitable PAM sequences include, but are not limited to, NGG, NGA, NGC, NGN, NGT, NGTT, NGCG, NGAG, NGAN, NGNG, NGCN, NGCG, NGTN, NNGRRT, NNNRRT, NNGRR(N), TTTV, TYCV, TYCV, TATV, NNNGATT, NNAGAAW, or NAAMC. Y is pyrimidine; N is any nucleotide base; W is A or T.
[0184] The base editor provided in this article can contain CRISPR protein-derived domains that can bind nucleotide sequences containing canonical or non-canonical prototype spacer adjacent motif (PAM) sequences.
[0185] In some implementations, PAM is “NRN” PAM, where “N” in “NRN” is adenine (A), thymine (T), guanine (G), or cytosine (C), and R is adenine (A) or guanine (G); or PAM is “NYN” PAM, where “N” in NYN is adenine (A), thymine (T), guanine (G), or cytosine (C), and Y is cytidine (C) or thymine (T), for example, as in RT Walton. et al. , 2020, Science As described in , 10.1126 / science.aba8853 (2020), the entire contents of which are incorporated herein by reference.
[0186] Several PAM variants are described in Table 3 below.
[0187] Table 3. Cas9 protein and corresponding PAM sequence. N is A, C, T, or G; and V is A, C, or G.
[0188] In some embodiments, the PAM is an NGC. In some embodiments, the NGC PAM is recognized by a Cas9 variant. In some embodiments, the Cas9 variant contains one or more amino acid substitutions selected from spCas9 (SEQ ID No: 197) of D1135V, G1218R, R1335Q, and T1337R (collectively, VRQR) or corresponding mutations in another Cas9. In some embodiments, the Cas9 variant contains one or more amino acid substitutions selected from spCas9 (SEQ ID No: 197) of D1135V, G1218R, R1335E, and T1337R (collectively, VRER) or corresponding mutations in another Cas9. In some embodiments, the Cas9 variant contains one or more amino acid substitutions selected from saCas9 (SEQ ID NO: 218) of E782K, N968K, and R1015H (collectively, KHH).
[0189] In some embodiments, the CRISPR protein-derived domain of the base editor comprises all or a portion (e.g., the functional portion) of a Cas9 protein having a canonical PAM sequence (NGG). In other embodiments, the Cas9-derived domain of the base editor may employ a non-canonical PAM sequence. Such sequences have been described in the art and will be readily apparent to those skilled in the art. For example, Cas9 domains binding non-canonical PAM sequences have been described in Kleinstiver, BP. et al. , “Engineered CRISPR-Cas9 nucleases with altered PAM specificities” Nature 523, 481-485 (2015); and Kleinstiver, BP et al. , "Broadening thetargeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAMrecognition" Nature Biotechnology 33, 1293-1298 (2015); RT Walton et al.“Unconstrained genome targeting with near-PAMless engineered CRISPR-Cas9 variants” Science 10.1126 / science.aba8853 (2020); Hu et al. “Evolved Cas9 variants with broad PAM compatibility and high DNA specificity,” Nature April 5, 2018, 556(7699), 57-63; Miller et al. , “Continuous evolution of SpCas9 variantscompatible with non-G PAMs” Nat. Biotechnol ., 2020 Apr;38(4):471-481; The entire contents of each are incorporated herein by reference.
[0190] A fusion protein or complex containing NapDNAbp and cytidine deaminase and / or adenosine deaminase Some aspects of this disclosure provide a programmable DNA-binding protein containing a Cas9 domain or other nucleic acid. For example A fusion protein or complex of Cas9 domains and one or more cytidine deaminase, adenosine deaminase, or cytidine adenosine deaminase domains. It should be understood that the Cas9 domain can be any Cas9 domain or Cas9 protein (as described herein). For example dCas9 or nCas9). In some implementations, any Cas9 domain or Cas9 protein (as described herein) For example (dCas9 or nCas9) can be fused with any cytidine deaminase and / or adenosine deaminase provided herein. The domains of the base editors disclosed herein can be arranged in any order.
[0191] In some implementations, it includes cytidine deaminase or adenosine deaminase and napDNAbp ( For example Fusion proteins or complexes (containing Cas9 or Cas12 domains) do not contain adapter sequences. In some embodiments, the adapter is present between cytidine or adenosine deaminase and napDNAbp. In some embodiments, cytidine or adenosine deaminase and napDNAbp are fused via any adapter provided herein. For example, in some embodiments, cytidine or adenosine deaminase and napDNAbp are fused via any adapter provided herein.
[0192] It should be understood that the fusion proteins or complexes disclosed herein may include one or more additional features. For example, in some embodiments, the fusion protein or complex may include inhibitors, cytoplasmic localization sequences, export sequences (such as nuclear export sequences) or other localization sequences, and sequence tags that can be used to dissolve, purify, or detect the fusion protein or complex. Suitable protein tags provided herein include, but are not limited to, biotinylate carboxylase carrier protein (BCCP) tags, myc tags, calmodulin tags, FLAG tags, hemagglutinin (HA) tags, multihistidine tags (also known as histidine tags or His tags), maltose-binding protein (MBP) tags, nus tags, glutathione S-transferase (GST) tags, green fluorescent protein (GFP) tags, thioredoxin tags, S tags, Softag ( For example The fusion protein or complex contains one or more His tags, including Softag 1, Softag 3, streptococcal tags, biotin ligase tags, FLASH tags, V5 tags, and SBP tags. Other suitable sequences will be apparent to those skilled in the art. In some embodiments, the fusion protein or complex contains one or more His tags.
[0193] Exemplary but non-limiting fusion proteins are described in International PCT Applications PCT / US2017 / 045381, PCT / US2019 / 044935 and PCT / US2020 / 016288, each of which is incorporated herein by reference in its entirety.
[0194] Fusion proteins or complexes with internal insertions This document provides fusion proteins or complexes comprising a heterologous polypeptide fused to a nucleic acid-programmable nucleic acid-binding protein (e.g., napDNAbp). The heterologous polypeptide may be fused to the C-terminus or N-terminus of napDNAbp, or inserted internally into the napDNAbp. In some embodiments, the heterologous polypeptide is a deaminase (…). For example (cytidine or adenosine deaminase) or a functional fragment thereof. For example, fusion proteins may contain Cas9 or Cas12 (crystalpine or adenosine deaminase) side-mounted. For example Deaminases of the N-terminal and C-terminal fragments of the Cas12b / C2c1 polypeptide.
[0195] The deaminase can be a circularly arranged mutant deaminase. In some embodiments, the deaminase is a circularly arranged mutant TadA, with amino acid residues arranged in a ring at positions 116, 136, or 65 according to the TadA reference sequence number.
[0196] Fusion proteins or complexes may contain more than one deaminase. For example, 1, 2, 3, 4, 5, or more deaminases may be present. The deaminases in the fusion protein or complex may be adenosine deaminase, cytidine deaminase, or a combination thereof.
[0197] In some embodiments, the napDNAbp in the fusion protein or complex contains a Cas9 polypeptide or a fragment thereof. The Cas9 polypeptide can be a variant of the Cas9 polypeptide. The Cas9 polypeptide can be a cyclically arranged Cas9 protein.
[0198] Heteropeptides ( For example (deaminase) can insert into nap DNAbp ( For example Cas9 or Cas12 ( For example The appropriate position of Cas12b / C2c1 is crucial, for example, to allow napDNAbp to retain its ability to bind target polynucleotides and guide nucleic acids. Deaminases ( For example Adenosine deaminase or cytidine deaminase can be inserted into nap DNAbp without impairing the function of the deaminase. For example (base editing activity) or the function of nap DNAbp ( For example (The ability to bind to target nucleic acids and guide nucleic acids).
[0199] In some implementation schemes, deaminase ( For example Adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) insertions contain above-average factor B ( For example The Cas9 polypeptide region containing a higher B factor than the total protein or protein domains containing disordered regions. Cas9 polypeptide positions containing a higher-than-average B factor may include, for example, residues 768, 792, 1052, 1015, 1022, 1026, 1029, 1067, 1040, 1054, 1068, 1246, 1247, and 1248 as numbered in SEQ ID NO: 197. Cas9 polypeptide regions containing a higher-than-average B factor may include, for example, residues 792-872, 792-906, and 2-791 as numbered in SEQ ID NO: 197.
[0200] In some implementations, heterologous peptides ( For exampleThe deaminase is inserted into the flexible ring of the Cas9 polypeptide. The flexible ring portion can be selected from the group consisting of numbers 530-537, 569-570, 686-691, 943-947, 1002-1025, 1052-1077, 1232-1247, or 1298-1300 in SEQ ID NO: 197, or the corresponding amino acid residues in another Cas9 polypeptide. The flexible ring portion can be selected from the group consisting of numbers 1-529, 538-568, 580-685, 692-942, 948-1001, 1026-1051, 1078-1231, or 1248-1297 in SEQ ID NO: 197, or the corresponding amino acid residues in another Cas9 polypeptide.
[0201] Heteropeptides ( For example (adenine deaminase) can be inserted into the Cas9 polypeptide region corresponding to the following amino acid residues: such as 1017-1069, 1242-1247, 1052-1056, 1060-1077, 1002-1003, 943-947, 530-537, 568-579, 686-691, 1242-1247, 1298-1300, 1066-1077, 1052-1056 or 1060-1077 in SEQ ID NO: 197, or the corresponding amino acid residues in another Cas9 polypeptide.
[0202] A heterologous polypeptide (e.g., adenine deaminase) can be inserted to replace the missing region of the Cas9 polypeptide. The missing region may correspond to the N-terminal or C-terminal portion of the Cas9 polypeptide. Exemplary internal fusion base editors are provided in Table 4A below: Table 4A: Insertion loci in the Cas9 protein
[0203] A heterologous polypeptide (e.g., a deaminase) may be inserted into a structural or functional domain of the Cas9 polypeptide. A heterologous polypeptide (e.g., a deaminase) may be inserted between two structural or functional domains of the Cas9 polypeptide. A heterologous polypeptide (e.g., a deaminase) may be inserted in place of a structural or functional domain of the Cas9 polypeptide, for example, after the domain has been deleted from the Cas9 polypeptide. The structural or functional domains of the Cas9 polypeptide may include, for example, RuvC I, RuvC II, RuvC III, Rec1, Rec2, PI, or HNH.
[0204] Fusion proteins may contain a linker between the deaminase and the nap DNAbp polypeptide. The linker can be a peptide or a non-peptide linker. For example, the linker could be XTEN, (GGGS) n(SEQ ID NO: 246), SGGSSGGS (SEQ ID NO: 330), (GGGGS) n (SEQ ID NO: 247), (G) n , (EAAAK)n (SEQ ID NO: 248), (GGS) n SGSETPGTSESATPES (SEQ ID NO: 249). In some embodiments, the fusion protein includes a linker between the N-terminal Cas9 fragment and the deaminase. In some embodiments, the fusion protein includes a linker between the C-terminal Cas9 fragment and the deaminase. In some embodiments, the N-terminal and C-terminal fragments of napDNAbp are linked to the deaminase using linkers. In some embodiments, the N-terminal and C-terminal fragments are conjugated to the deaminase domain without a linker. In some embodiments, the fusion protein includes a linker between the N-terminal Cas9 fragment and the deaminase, but not between the C-terminal Cas9 fragment and the deaminase. In some embodiments, the fusion protein includes a linker between the C-terminal Cas9 fragment and the deaminase, but not between the N-terminal Cas9 fragment and the deaminase.
[0205] In some implementations, the napDNAbp in the fusion protein or complex is a Cas12 polypeptide ( For example The Cas12 peptide (Cas12b / C2c1) or a functional fragment thereof capable of associating with a nucleic acid (e.g., gRNA) that directs Cas12 to a specific nucleic acid sequence. The Cas12 peptide may be a variant of the Cas12 peptide. In other embodiments, the N-terminal or C-terminal fragment of the Cas12 peptide contains a nucleic acid-programmable DNA-binding domain or a RuvC domain. In other embodiments, the fusion protein contains a linker between the Cas12 peptide and the catalytic domain. In other embodiments, the amino acid sequence of the linker is GGSGGS (SEQ ID NO: 250) or GSSGSETPGTSESATPESSG (SEQ ID NO: 251). In other embodiments, the linker is a rigid linker. In other embodiments of the foregoing aspects, the linker is encoded by GGAGGCTCTGGAGGAAGC (SEQ ID NO: 252) or GGCTCTTCTGGATCTGAAACACCTGGCACAAGCGAGAGCGCCACCCCTGAGAGCTCTGGC (SEQ ID NO: 253).
[0206] In other embodiments, the fusion protein or complex contains a nuclear localization signal ( For example(Dichotomous nuclear localization signal). In other embodiments, the amino acid sequence of the nuclear localization signal is MAPKKKRKVGIHGVPAA (SEQ ID NO: 261). In other embodiments of the above aspects, the nuclear localization signal is encoded by the following sequence: ATGGCCCCAAAGAAGAAGCGGAAGGTCGGTATCCACGGAGTCCCAGCAGCC (SEQ ID NO: 262). In other embodiments, the Cas12b peptide contains a mutation that silences the catalytic activity of the RuvC domain. In other embodiments, the Cas12b peptide contains D574A, D829A, and / or D952A mutations.
[0207] In some implementations, the fusion protein or complex includes a nucleobase editing domain with internal fusion ( For example Deaminase domain ( For example The napDNAbp domain (including all or part of the adenosine deaminase domain, e.g., the functional portion) For example (Cas12-derived domains). In some embodiments, napDNAbp is Cas12b. In some embodiments, the base editor contains a BhCas12b domain having an internally fused TadA*8 domain inserted at the locus provided in Table 4B below.
[0208] Table 4B: Insertion loci in the Cas12b protein
[0209] In some embodiments, the base editing system described herein is an ABE having a TadA with an inserted Cas9. The polypeptide sequence of the relevant ABE having a TadA with an inserted Cas9 is provided in the appended sequence listing as SEQ ID NO: 263-308.
[0210] Exemplary but non-limiting fusion proteins are described in International PCT Application No. PCT / US2020 / 016285 and U.S. Provisional Applications Nos. 62 / 852,228 and 62 / 852,224, the contents of which are incorporated herein by reference in their entirety.
[0211] A to G editing In some embodiments, the base editor described herein includes an adenosine deaminase domain. This adenosine deaminase domain of the base editor facilitates the editing of adenine (A) nucleosides to guanine (G) nucleosides by deaminating A to form inosine (I), which exhibits the base-pairing properties of G. In some embodiments, the A-to-G base editor also includes an inosine base excision repair inhibitor, such as a uracil glycosylase inhibitor (UGI) domain or a non-catalytically active inosine-specific nuclease. Without being bound by any particular theory, the UGI domain or the non-catalytically active inosine-specific nuclease can inhibit or prevent the deamination of adenosine residues (…). For example Base excision repair of inosine can improve the activity or efficiency of base editors.
[0212] A base editor containing an adenosine deaminase can act on any polynucleotide, including DNA, RNA, and DNA-RNA hybrids. In one embodiment, the adenosine deaminase domain of the base editor contains all or part (e.g., a functional portion) of ADAT, which contains one or more mutations that allow ADAT to deaminate target A in DNA. For example, the base editor may contain all or part (e.g., a functional portion) of ADAT from *Escherichia coli* (EcTadA), which contains one or more of the following mutations: D108N, A106V, D147Y, E155V, L84F, H123Y, I156F, or a corresponding mutation in another adenosine deaminase. Exemplary ADAT homologous polypeptide sequences are provided in the sequence listing as SEQ ID NO: 1 and 309-315.
[0213] Adenosine deaminase can be derived from any suitable organism. For example , E. coli In some implementations, adenosine deaminase originates from... E. coli , Staphylococcus aureus, Salmonella typhi, and Sheva putrefactive bacteria Shewanella putrefaciens, Haemophilus influenzae, and Crescentella (Caulobacter crescentus) or Bacillus subtilis In some embodiments, adenine deaminase is a naturally occurring adenine deaminase containing any mutations corresponding to those provided herein. For example One or more mutations (mutations in ecTadA). Corresponding residues in any homologous protein can be detected through... For example Sequence alignment and identification of homologous residues are used for identification. Corresponding mutations can be generated accordingly for any mutations described herein. For example Any naturally occurring adenosine deaminase (any mutation identified in ecTadA) For example Mutations in (homogeneous with ecTadA)
[0214] In some embodiments, the adenosine deaminase comprises an amino acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with any amino acid sequence shown in any adenosine deaminase provided herein. It should be understood that the adenosine deaminases provided herein may include one or more mutations (…). For example (Any mutations provided herein). This disclosure provides any deaminase domain having a certain percentage of identity plus any mutations or combinations thereof described herein. In some embodiments, the adenosine deaminase comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to the reference sequence provided herein or any adenosine deaminase.
[0215] It should be understood that any mutations provided in this article ( For example Based on the TadA reference sequence, sequences such as TadA*7.10 (SEQ ID NO: 1) can be introduced into other adenosine deaminases, such as... E. coli TadA (ecTadA), Staphylococcus aureus TadA (saTadA) or other adenosine deaminases ( For example (Bacterial adenosine deaminase). In some embodiments, the TadA reference sequence is TadA*7.10 (SEQ ID NO: 1). It will be apparent to those skilled in the art that other deaminases can be similarly compared to identify homologous amino acid residues that can be mutated as provided herein. Therefore, any mutation identified in the TadA reference sequence can be identified in other adenosine deaminases having homologous amino acid residues (bacterial adenosine deaminase). For example It appears in (ecTada). It should also be understood that any mutations provided herein may appear alone or in any combination in the TadA reference sequence or in another adenosine deaminase.
[0216] In some implementations, the adenosine deaminase comprises altered or altered groups selected from those listed in Tables 5A to 5E below: Table 5A. Adenosine deaminase variants. Indicates... E. coli Residue positions in the TadA variant (TadA*).
[0217] Table 5B. TadA*8 adenosine deaminase variants. Indicates... E. coli The residue positions in the TadA variant (TadA*). See reference TadA*7.10 (first line).
[0218] Table 5C. TadA*9 adenosine deaminase variants. Changes refer to TadA*7.10. Further details of TadA*9 adenosine deaminase are described in International PCT Application No. PCT / US2020 / 049975, which is incorporated herein by reference in its entirety for all purposes.
[0219] In some embodiments, the adenosine deaminase comprises a TadA*8.20 adenosine deaminase variant, which further comprises the F149Y amino acid alteration. In some embodiments, the adenosine deaminase comprises a TadA*8.20 adenosine deaminase variant, which further comprises the amino acid alterations R147D, F149Y, T166I, and D167N (TadA*8.10+). In some embodiments, the adenosine deaminase comprises a TadA*8.20 adenosine deaminase variant, which further comprises the amino acid alterations S82T and F149Y (TadA*9v1). In some embodiments, the adenosine deaminase comprises a TadA*8.20 adenosine deaminase variant, which further comprises the amino acid alterations Y147D, F149Y, T166I, D167N, and S82T (TadA*9v2).
[0220] In some embodiments, the adenosine deaminase comprises the following sequences from the TadA reference sequence (e.g., TadA*7.10, ecTadA, or TadA8e): M1I, M1S, S2A, S2E, S2H, S2R, S2L, E3L, V4D, V4E, V4M, V4K, V4S, V4T, V4A, E5K, F6S, F6G, F6H, F6Y, F6I, F6E, S7K, H8E, H8Y, H8H, H8Q, H8E, H8G, H8S, E9Y, E9K, E9V, E9E, Y10F, Y10W, Y10Y, M12S, M12L, M12R, M12W, R13H, R13I, R13Y, R13R, R13G. R13S, H14N, A15D, A15V, A15L, A15H, T17T, T17A, T17W, T17L, T17F, T17R, T1 7S, L18A, L18E, L18N, L18L, L18S, A19N, A19H, A19K, A19A, A19D, A19G, A19M, R21N, K20K, K20A, K20R, K20E, K20G, K20C, K20Q, R21A, R21R, R21N, R21Y, R2 1C, G22P, A22W, A22R, W23D, R23H, W23G, W23Q, W23L, W23R, W23H, W23D, W23M, W23W, W23I, D24E, D24G, D24W, D24D, D24R, E25F, E25M, E25D, E25A, E25G, E2 5R, E25E, E25H, E25V, E25S, E25Y, R26D, R26E, R26G, R26N, R26Q, R26C, R26L, R26K, R26W, R26C, R26P, R26R, R26A, R26H, E27E, E27Q, E27H, E27C, E27G, E2 7K, E27S, E27P, E27R, E27L, E27V, E27D, V28V, V28A, V28C, V28G, V28P, V28S, V28T, P29V, P29P, P29A, P29G, P29K, P29L, V30V, V30I, V30L, V30F, V30G, V3 0A, V30M, L34S, L34V, L34L, L34M, L34W, L34G, H36E, H36V, L36H, H36L, H36N, N37N, N37H, N37R, N37T, N37S, N38G, N38R, N38N, N38E, V40I, W45A, W45W, W4 5R, W45L, W45N, N46N, N46M, N46P, N46G, N46L, N46R, N46V, R46W, R46F, R46Q,R46M、R47A、R47Q、R47F、R47K、R47P、R47W、R47M、R47R、R47G、R47S、R47V、R47H、P48T、P48L、P48A、P48I、P48S、P48R、P48K、P48D、P48E、P48H、P48G、P48P、P48N、I49G、I49H、I49V、I49F、I49H、I49I、I49M、I49N、I49K、I49Q、I49T、G50L、G50S、G50R、G50G、R51H、R51L、R51N、L51W、R51Y、R51G、R51V、R51R、H52D、H52Y、H52I、H52H、D53D、D53E、D53G、D53P、P54C、P54T、P54P、P54E、A55H、T55A、T55I、T55V、T55G、T55T、A56A、A56H、A56W、A56E、A56S、H57P、H57A、H57H、H57N、A58G、A58E、A58A、A58R、E59A、E59G、E59I、E59Q、E59W、E59E、E59T、E59H、E59P、M61A、M61I、M61L、M61V、M61P、M61G、M61I、L63S、L63V、L63T、L63R、L63H、L63A、R64A、R64Q、R64R、R64D、Q65V、Q65H、Q65G、Q65P、Q65F、Q65Q、Q65R、G66V、G66E、G66T、G66G、G66C、G67G、G67W、G67I、G67A、G67D、G67L、G67V、L68Q、L68M、L68V、L68H、L68L、L68G、V69A、V69M、V69V、M70V、M70L、E70A、M70A、M70M、M70E、M70T、M70v、Q71M、Q71N、Q71L、Q71R、Q71Q、Q71I、N72A、N72K、N72S、N72D、N72Y、N72N、N72H、N72G、N72M、Y73G、Y73I、Y73K、Y73R、Y73S、Y73Y、Y73H、Y73A、R74A、R74Q、R74G、R74K、R74L、R74N、R74G、R74K、R74R、I76H、I76R、I76W、I76Y、I76V、I76Q、I76L、I76D、I76F、I76I、I76N、I76T、I76Y、D77G、D77D、D77A、D77Q、A78Y、A78T、A78G、A78A、A78I、T79M、T79R、T79L、T79T、L80M、L80Y、L80I、L80V、L80L、Y81D、Y81V、Y81Y、Y81M、V82A、V82S、V82G、V82T、V82V、V82Q、V82Y、T83L、T83F、T83T、T83N、L84E、L84F、L84Y、L84I、L84L、L84M、L84A、L84T、L84S、E85K、E85G、E85P、E85S、E85E、E85F、E85V、E85R、P86T、P86C、P86P、P86L、P86N、P86K、P86H、C87M、C87I、C87S、C87N、C87P、S87C、S87L、S87V、V88A、V88M、V88V、V88T、V88E、V88D、V88S、C90S、C90P、C90A、C90T、C90M、A91A、A91G、A91S、A91V、A91T、A91C、A91L、G92T、G92M、G92A、G92Y、G92G、A93I、A93C、A93M、A93V、A93A、M94M、M94T、M94A、M94V、M94L、M94I、M94H、I95S、I95G、I95L、I95H、I95V、H96A、H96L、H96R、H96S、H96H、H96N、H96E、S97C、S97G、S97I、S97M、S97R、S97S、S97P、R98K、R98I、R98N、R98Q、R98G、R98H、R98C、R98L、R98R、G100R、G100V、G100K、G100A、G100S、G100M、G100I、R101V、R101R、R101S、R101C、V102A、V102F、V102I、V102V、D103A、V103A、V103G、V103F、V103V、F104G、D104N、F104V、F104I、F104L、F104A、F104F、F104R、G105V、G105W、G105G、G105M、G105A、A106T、V106Q、V106F、V106W、V106M、A106A、A106Q、A106F、A106G、A106W、A106M、A106V、A106R、A106L、A106S、A106B、A106I、R107C、R107G、R107P、R107K、R107A、R107N、R107W、R107H、R107S、R107R、R107F、D108N、D108F、D108G、D108V、D108A、D108Y、D108H、D108I、D108K、D108L、D108M、D108Q、N108Q、N108F、N108W、N108M、N108K、D108K、D108F、D108M、D108Q、D108R、D108W、D108S、D108E、D108T、D108R、D108D、A109H、A109K、A109R、A109S、A109T、A109V、A109A、A109D、K110G、K110H、K110I、K110R、K110T、K110K、K110A、K110l、T111A、T111G、T111H、T111R、T111T、T111K、G112A、G112G、G112H、G112T、G112R、A113N、A114G、A114H、A114V、A114C、A114S、A114A、G115S、G115G、G115M、G115L、G115A、G115F、L117M、L117L、L117W、L117A、L117S、L117N、L117V、M118D、M118G、M118K、M118N、M118V、M118M、M118L、M118R、D119L、D119N、D119S、D119V、D119D、V120H、V120L、V120V、V120T、V120A、V120E、V120G、V120D、L121D、L121M、L121N、L121K、L121L、H122H、H122N、H122P、H122R、H122S、H122Y、H122G、H122T、H122L、H123C、H123G、H123P、H123V、H123Y、Y123H、H123Y、H123H、P124P、P124H、P124A、P124Y、P124D、P124G、P124I、P124L、P124W、G125H、G125I、G125A、G125M、G125K、G125G、G125P、M126D、M126H、M126K、M126I、M126N、M126O、M126S、M126Y、M126M、M126G、N127H、N127S、N127D、N127K、N127R、N127N、N127I、N127P、N127M、H128R、H128N、H128L、H128H、R129H、R129Q、R129V、R129I、R129E、R129V、R129R、R129M、R129P、V130R、V130V、V130E、V130D、E131E、E131I、E131V、E131K、I132I、I132F、I132T、I132L、I132V、I132E、T133V、T133E、T133G、T133K、T133T、T133A、T133H、T133F、T133I、E134A、E134E、E134G、E134I、E134H、E134K、E134T、G135G、G135V、G135I、G135P、G135E、I136G、I136L、I136T、I136I、l137A、l137D、l137E、L137M、l137S、L137L、L137I、A138D、A138E、A138G、S138A、A138N、A138S、A138T、A138V、A138Y、A138A、A138M、A138L、D139E、D139I、D139C、D139L、D139M、D139D、D139G、D139H、D139A、E140A、E140C、E140L、E140R、E140K、E140E、E140D、C141S、C141A、C141C、C141V、C141E、A142N、A142D、A142G、A142A、A142L、A142S、A142T、A142N、A142S、A142V、A142E、A142C、A143D、A143E、A143G、A143D、A143G、A143E、A143L、A143W、A143M、A143S、A143Q、A143R、A143A、A143I、L144S、L144L、L144T、L144A、L145A、L145F、L145G、L145D、L145L、L145C、L145E、L145s、C146R、S146A、S146C、S146D、S146F、S146R、S146T、S146D、S146G、S146S、S146L、D147D、D147L、D147F、D147G、D147Y、Y147T、Y147R、Y147D、D147R、D147Y、D147A、D147T、D147H、D147F、D147U、D147V、D147I、D147C、F148L、F148F、F148R、F148Y、F148A、F148T、F149C、F149M、F149R、F149Y、F149N、F149F、F149A、F149T、F149V、R150R、R150M、R150D、R150F、M151F、M151P、M151R、M151V、M151M、M151E、R152C、R152F、R152H、R152P、R152R, R152P, R152Q, R152M, R152O, R153C, R153Q, R153R, R153V, R153E, R153A, R153P, Q154E, Q154H, Q154M, Q154R, Q154L, Q154S, Q1 54V, Q154Q, Q154F, Q154I, Q154A, Q154K, E155F, E155G, E155I, E155K, E155P, E155V, E155D, E155E, E155L, E155Q, I156V, I156A, I156 I, I156L, I156F, I156D, I156K, I156N, I156R, I156Y, E157A, E157F, E157I, E157P, E157T, E157V, N157K, K157N, K157V, K157P, K157I, K157F, K157F, K157T, K157A, K157S, K157R, A158Q, A158K, A158V, A158A, A158D, A158S, A158T, A158N, Q159S, Q159Q, Q159A, Q159F, Q1 59K, Q159L, Q159N, K160A, K160S, K160E, K160K, K160N, K160F, K160Q, K161T, K161K, K161R, K161I, K161A, K161N, K161Q, K161S, K161 T, A162D, A162Q, R162H, R162P, A162S, A162A, A162N, A162M, A162K, Q163G, Q163S, Q163Q, Q163A, Q163H, Q163N, Q163R, S164F, S164S, One or more of the following mutations: S164Q, S164I, S164R, S164Y, S165S, S165P, S165Q, S165A, S165D, S165I, S165T, S165Y, T166T, T166Q, T166E, T166S, T166D, T166K, T166I, T166N, T166P, T166R, D167S, D167D, D167I, D167G, D167T, D167A and / or D167N, and any alternative mutations at the corresponding positions, or one or more corresponding mutations in another adenosine deaminase. Other mutations are described in U.S. Patent Application Publication No. 2022 / 0307003 A1, U.S. Patent No. 11,155,803, and International Patent Application Publication Nos. WO 2023 / 288304 A2, PCT / CN2022 / 143408, and WO 2018 / 027078 A1.The disclosures of patents WO 2021 / 158921 A1 and WO 2023 / 034959 A2 are incorporated herein by reference in their entirety for all purposes.
[0221] In some embodiments, this disclosure provides TadA variants comprising V82T, Y147T, and / or Q154S mutations. In some embodiments, this disclosure provides TadA variants comprising V82T, Y147T, and / or Q154S mutations. In some embodiments, this disclosure provides TadA*8.8, which also comprises a V82T mutation. In some embodiments, this disclosure provides TadA*8.8, which also comprises V82T, Y147T, and Q154S mutations. In some embodiments, this disclosure provides TadA*8.17, which also comprises a V82T mutation. In some embodiments, this disclosure provides TadA*8.17, which also comprises V82T, Y147T, and Q154S mutations. In some embodiments, this disclosure provides TadA*8.20, which also comprises a V82T mutation. In some embodiments, this disclosure provides TadA*8.20, which also comprises V82T, Y147T, and Q154S mutations.
[0222] In the implementation, variations of TadA*7.10 include one or more changes selected from any of the changes provided herein.
[0223] In a specific implementation, the adenosine deaminase heterodimer comprises a TadA*8 domain and an adenosine deaminase domain, wherein the adenosine deaminase domain is selected from... Staphylococcus aureus ( S. aureus TadA、 Bacillus subtilis ( B. subtilis TadA、 Salmonella typhimurium ( S. typhimurium TadA、 rot Shewanella ( S. putrefaciens TadA, Haemophilus influenzae F3031 ( H. influenzae TadA, Crested Bacillus ( C. crescentus TadA、 Geobacter sulfurreducens ( G. sulfurreducens ) TadA, or TadA*7.10.
[0224] In some implementations, TadA*8 is a variant as shown in Table 5D. Table 5D shows the position numbers of certain amino acids in the TadA amino acid sequence and the amino acids present at these positions in the TadA-7.10 adenosine deaminase. Table 5D also shows the amino acid changes in the TadA variants relative to TadA-7.10 after phage-assisted discontinuous evolution (PANCE) and phage-assisted continuous evolution (PACE), as described by M. Richter. et al. The entire contents of the description in Nature Biotechnology, doi.org / 10.1038 / s41587-020-0453-z, 2020, are incorporated herein by reference. In some embodiments, TadA*8 is TadA*8a, TadA*8b, TadA*8c, TadA*8d, or TadA*8e. In some embodiments, TadA*8 is TadA*8e. In one embodiment, the adenosine deaminase is TadA*8, comprising, or consisting substantially of, SEQ ID NO: 316 or a fragment thereof having adenosine deaminase activity.
[0225] Table 5D. Selecting the TadA*8 variant
[0226] In some embodiments, the TadA variants are those shown in Table 5E. Table 5E shows the position numbers of certain amino acids in the TadA amino acid sequence and the amino acids present at those positions in TadA*7.10 adenosine deaminase. In some embodiments, the TadA variants are MSP605, MSP680, MSP823, MSP824, MSP825, MSP827, MSP828, or MSP829. In some embodiments, the TadA variant is MSP828. In some embodiments, the TadA variant is MSP829.
[0227] Table 5E. TadA variants
[0228] In a specific implementation scheme, the fusion protein or complex comprises a single ( For exampleProvided as a monomer, TadA* (e.g., TadA*8 or TadA*9). Throughout this disclosure, adenosine deaminase base editors comprising a single TadA* domain are referred to by the term ABEm or ABE#m, where "#" is an identification number (e.g., ABE8.20m) and "m" signifies "monomer". In some embodiments, TadA* is linked to a Cas9 nickase. In some embodiments, the fusion protein or complex of this disclosure comprises a heterodimer of wild-type TadA (TadA(wt)) linked to TadA*. Throughout this disclosure, adenosine deaminase base editors comprising a single TadA* domain and a TadA(wt) domain are referred to by the term ABEd or ABE#d, where "#" is an identification number (e.g., ABE8.20d) and "d" signifies "dimer". In other embodiments, the fusion protein or complex of this disclosure comprises a heterodimer of TadA*7.10 linked to TadA*. In some embodiments, the base editor is ABE8, which comprises a TadA* variant monomer. In some embodiments, the base editor is ABE, which comprises a heterodimer of TadA* and TadA (wt). In some embodiments, the base editor is ABE, which comprises a heterodimer of TadA* and TadA*7.10. In some embodiments, the base editor is ABE, which comprises a heterodimer of TadA*. In some embodiments, TadA* is selected from Tables 5A to 5E.
[0229] In some embodiments, adenosine deaminase is expressed as a monomer. In other embodiments, adenosine deaminase is expressed as a heterodimer. In some embodiments, the deaminase or other polypeptide sequence lacks a methionine, for example, when included as a component of a fusion protein. This may alter the positional numbering. However, those skilled in the art will understand that such corresponding mutations refer to the same mutation.
[0230] Any mutations described in this article and any other mutations ( For example Based on the ecTadA amino acid sequence, this can be introduced into any other adenosine deaminase. Any mutations provided herein can be introduced alone or in any combination into the TadA reference sequence or into another adenosine deaminase (…). For example It appears in ,ecTadA).
[0231] Detailed descriptions of the A-to-G nucleobase editing protein can be found in International PCT Application No. PCT / US2017 / 045381 (No. WO2018 / 027078) and Gaudelli, NM et al.“Programmable base editing of A • T to G • C in genomic DNA without DNA cleavage” Nature, 551, 464-471 (2017), the entire contents of the above reference are hereby incorporated by way of citation.
[0232] C to T editing In some embodiments, the base editor disclosed herein comprises a fusion protein or complex containing a cytidine deaminase capable of deaminating a target cytidine (C) base of a polynucleotide to produce uridine (U), said uridine having the base-pairing property of thymine. In some embodiments, for example when the polynucleotide is double-stranded (… For example When DNA is involved, uridine bases can then be replaced by thymidine bases. For example (Through cellular repair mechanisms) to achieve the C:G to T:A conversion. In other embodiments, the base editor deamination of C in nucleic acids to U cannot be accompanied by U to T substitution.
[0233] The deamination of a target C in a polynucleotide to produce a U is a non-limiting example of a type of base editing that can be performed by the base editor described herein. In another example, a base editor containing a cytidine deaminase domain can mediate the conversion of a cytosine (C) base to a guanine (G) base. For example, the U of a polynucleotide produced by deamination of cytidine via the cytidine deaminase domain of a base editor can be obtained through a base excision repair mechanism (…). For example The uracil DNA glycosylation enzyme (UDG) removes a polynucleotide from the uracil DNA glycosylation enzyme, creating a debasement site. The nucleobase opposite the debasement site can then be replaced by another base (such as C) via, for example, a trans-damage polymerase. For example (Through base repair mechanisms). Although the nucleobase opposite the debasement site is usually replaced by a C, other substitutions may also occur ( For example (A, G, or T).
[0234] Therefore, in some embodiments, the base editor described herein includes a deamination domain capable of deaminating a target C to U in a polynucleotide. For example(Cytidine deaminase domain). Furthermore, as described below, the base editor may include additional domains that, in some embodiments, facilitate the conversion of the deamination-derived U to T or G. For example, a base editor containing a cytidine deaminase domain may also include a uracil glycosylation inhibitor (UGI) domain to mediate the substitution of U for T, thereby completing the C-to-T base editing event. In another instance, the base editor may include a uracil-stabilizing protein as described herein. In yet another instance, the base editor may incorporate a trans-damage polymerase to improve the efficiency of C-to-G base editing, since the trans-damage polymerase facilitates the incorporation of the C opposite the deamination site (i.e., This leads to the incorporation of G at the debasement site, thereby completing C. (To the G base editing event).
[0235] Base editors containing cytidine deaminase as a domain can deaminate target C in any polynucleotide, including DNA, RNA, and DNA-RNA hybrids.
[0236] In some implementations, the base editor's cytidine deaminase comprises all or part (e.g., the functional portion) of apolipoprotein B mRNA editing complex (APOBEC) family of deaminases. APOBEC is an evolutionarily conserved family of cytidine deaminases. Members of this family are C-to-U editing enzymes. The N-terminal domain of APOBEC-like proteins is the catalytic domain, while the C-terminal domain is the pseudocatalytic domain. More specifically, the catalytic domain is the zinc-dependent cytidine deaminase domain and is important for cytidine deamination. APOBEC family members include APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D (“APOBEC3E” now refers to this), APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, and activation-induced (cytidine) deaminase.
[0237] Other exemplary deaminases that can be fused to Cas9 according to aspects of this disclosure are provided below. In embodiments, the deaminase is an activation-induced deaminase (AID). It should be understood that in some embodiments, the active domain of the corresponding sequence may be used. For example It has no domains with localization signals (nuclear localization sequences, no nuclear output signals, and no cytoplasmic localization signals).
[0238] Some aspects of this disclosure are based on the understanding that modulating the deaminase domain catalytic activity of any fusion protein or complex described herein (e.g., by point mutation of the deaminase domain) affects the fusion protein ( For exampleThe persistence of base editors or complexes. For example, mutations that reduce but do not eliminate the catalytic activity of the deaminase domain within a base-editing fusion protein or complex can make the deaminase domain less likely to catalyze deamination of residues adjacent to the target residue, thereby narrowing the deamination window. The ability to narrow the deamination window prevents unwanted deamination of residues adjacent to a specific target residue, which can reduce or prevent off-target effects.
[0239] In some embodiments, the APOBEC deaminase incorporated into the base editor may comprise one or more mutations selected from the group consisting of H121R, H122R, R126A, R126E, R118A, W90A, W90Y, and R132E of rAPOBEC1; D316R, D317R, R320A, R320E, R313A, W285A, W285Y, and R326E of hAPOBEC3G; and any alternative mutations at the corresponding positions, or one or more corresponding mutations in another APOBEC deaminase.
[0240] Many modified cytidine deaminases are commercially available, including but not limited to SaBE3, SaKKH-BE3, VQR-BE3, EQR-BE3, VRER-BE3, YE1-BE3, EE-BE3, YE2-BE3, and YEE-BE3, which are available from Addgene (plasmids 85169, 85170, 85171, 85172, 85173, 85174, 85175, 85176, and 85177). In some embodiments, the deaminase incorporated into the base editor comprises all or part (e.g., the functional portion) of the APOBEC1 deaminase.
[0241] In some embodiments, the fusion protein or complex of this disclosure comprises one or more cytidine deaminase domains. In some embodiments, the cytidine deaminases provided herein are capable of deaminating cytosine or 5-methylcytosine to uracil or thymine. In some embodiments, the cytidine deaminases provided herein are capable of deaminating cytosine in DNA. Cytidine deaminases can be derived from any suitable organism. In some embodiments, the cytidine deaminase is a naturally occurring cytidine deaminase comprising one or more mutations corresponding to any mutation provided herein. Those skilled in the art will be able to For example Corresponding residues in any homologous protein are identified by sequence alignment and determination of homologous residues. Therefore, those skilled in the art will be able to generate mutations corresponding to any mutations described herein in any naturally occurring cytidine deaminase. In some embodiments, the cytidine deaminase is derived from prokaryotes. In some embodiments, the cytidine deaminase is derived from bacteria. In some embodiments, the cytidine deaminase is derived from mammals (…). For example ,people).
[0242] In some embodiments, the cytidine deaminase comprises an amino acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with any of the cytidine deaminases shown herein. It should be understood that the cytidine deaminases provided herein may include one or more mutations (…). For example (Any mutations provided herein). Some embodiments provide multinucleotide molecules encoding any of the foregoing aspects or cytidine deaminase nucleobase editor polypeptides as described herein. In some embodiments, the multinucleotide is codon-optimized.
[0243] In the implementation scheme, the fusion protein of this disclosure comprises two or more nucleic acid editing domains.
[0244] Detailed descriptions of the C-to-T nucleobase editing protein can be found in International PCT Application No. PCT / US2016 / 058344 (No. WO2017 / 070632) and Komor, AC et al. , “Programmable editing of a target base ingenomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016), the entire contents of which are hereby incorporated by reference.
[0245] Guided polynucleotides The programmable nucleotide binding domain of a polynucleotide binds to a guide polynucleotide (PNP). For example When gRNA binds, it can specifically bind to the target polynucleotide sequence. Right now The base editor is directed to the target nucleic acid sequence to be edited via complementary base pairing between the bases of the binding guide nucleic acid and the bases of the target polynucleotide sequence. In some embodiments, the target polynucleotide sequence comprises single-stranded or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid.
[0246] In one embodiment, the guiding polynucleotide described herein can be RNA or DNA. In one embodiment, the guiding polynucleotide is gRNA.
[0247] In some embodiments, the guiding polynucleotide is at least one single guiding RNA (“sgRNA” or “gRNA”). In some embodiments, the guiding polynucleotide comprises two or more separate polynucleotides, which may be coupled via, for example, complementary base pairing (…). For example, Dual guide polynucleotides (dual gRNAs) interact with each other. For example, the guide polynucleotide may contain CRISPR RNA (crRNA) and trans-activating CRISPR RNA (tracrRNA), or may contain one or more trans-activating CRISPR RNAs (tracrRNA).
[0248] Guide polynucleotides can contain either natural or non-natural nucleotides. example like (peptide nucleic acid or nucleotide analogue). In some cases, the length of the target region guiding the nucleic acid sequence (e.g., spacer region) can be at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides.
[0249] In some embodiments, the methods described herein may utilize engineered Cas proteins. The guide RNA (gRNA) is a short synthetic RNA consisting of a scaffold sequence necessary for Cas binding and a user-defined spacer region of about 20 nucleotides defining a genomic target to be modified. Exemplary gRNA scaffold sequences are provided in the sequence listing as SEQ ID NOs: 317-327 and 425. Therefore, those skilled in the art can modify the Cas protein-specific genomic target, depending in part on how specific the gRNA targeting sequence is to the genomic target compared to the rest of the genome. In embodiments, the length of the spacer region is about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25 or more nucleotides. The length of the spacer region of the gRNA may be or may be about 19, 20 or 21 nucleotides.
[0250] gRNAs or guide polynucleotides can target any exon or intron of a gene target. In some embodiments, the composition comprises multiple gRNAs that all target the same exon or multiple gRNAs that target different exons. It can target exons and / or introns of a gene. The gRNA or guide polynucleotide can target about 20 nucleotides or fewer than about 20 nucleotides. For example At least about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 30 nucleotides) or any number of nucleotides between about 1 and 100. For exampleThe target nucleic acid sequence can be a nucleic acid sequence of at least 1-10, 1-20, 1-30, 1-40, 1-50, 1-60, 1-70, 1-80, 1-90, or 1-100 nucleotides. The target nucleic acid sequence can be the 5' 20 bases immediately adjacent to the first nucleotide of the PAM, or approximately the 5' 20 bases immediately adjacent to the first nucleotide of the PAM. gRNA can target nucleic acid sequences. The target nucleic acid can be at least or approximately 1-10, 1-20, 1-30, 1-40, 1-50, 1-60, 1-70, 1-80, 1-90, or 1-100 nucleotides.
[0251] Guided polynucleotides may include standard ribonucleotides, modified ribonucleotides ( For example (pseudouridine), ribonucleotide isomers and / or ribonucleotide analogs.
[0252] In some implementations, the base editor system may contain multiple guide polynucleotides. For example gRNA. For example, gRNA can target one or more target loci contained in a base editor system. For example (At least 1 gRNA, at least 2 gRNAs, at least 5 gRNAs, at least 10 gRNAs, at least 20 gRNAs, at least 30 gRNAs, at least 50 gRNAs). In one embodiment, the multiple gRNA sequences may be arranged in tandem and separated by unidirectional repeats.
[0253] Modified polynucleotides To enhance expression, stability, and / or genome / base editing efficiency, and / or reduce potential toxicity, base editor coding sequences (e.g., mRNA) and / or guide polynucleotides (e.g., gRNA) can be modified to include one or more modified nucleotides and / or chemical modifications. For example Using pseudouridine, 5-methylcytosine, 2'- O 2'-methyl-3'-phosphonoacetate, 2'- O -MethylthioPACE (MSP), 2'- O -Methyl-PACE (MP), 2'-FluoroRNA (2'-F-RNA), =Restricted Ethyl (S-cEt), 2'-O-Methyl ('M'), 2'- O -Methyl-3'-thiophosphate ('MS'), 2'- O 3'-Methyl-3'-thiophosphonoacetate ('MSP'), 5-methoxyuridine, thiophosphate, and N1-methylpseudouridine. Chemically protected gRNAs can enhance... in vivo and In vitroStability and editing efficiency. Methods using chemically modified mRNA and guide RNA are known in the art, and are illustrated, for example, by Jiang et al., Chemical modifications of adenine-based editor mRNA and guide RNA expand its application scope. Nat Commun 11, 1979 (2020). doi.org / 10.1038 / s41467-020-15892-8, and Callum et al. N 1-Methylpseudouridine substitution enhances the performance of synthetic mRNA switches in cells, Nucleic Acids Research Volume 48, Issue 6, April 6, 2020, page 35, and Andries et al., Journal of Controlled Release, Volume 217, November 10, 2015, pp. 337-344, each of which is incorporated herein by reference in its entirety.
[0254] In some embodiments, the guiding polynucleotide contains one or more modified nucleotides at the 5' and / or 3' end of the guide. In some embodiments, the guiding polynucleotide contains two, three, four, or more modified nucleotides at the 5' and / or 3' end of the guide. In some embodiments, the guiding polynucleotide contains two, three, four, or more modified nucleotides at the 5' and / or 3' end of the guide.
[0255] In some embodiments, the guide contains at least about 50% to 75% modified nucleotides. In some embodiments, the guide contains at least about 85% or more modified nucleotides. In some embodiments, at least about 1 to 5 nucleotides at the 5' end of the gRNA are modified, and at least about 1 to 5 nucleotides at the 3' end of the gRNA are modified. In some embodiments, at least about 3 to 5 consecutive nucleotides at both the 5' and 3' ends of the gRNA are modified. In some embodiments, at least about 20% of the nucleotides present in the direct or inverted repeats are modified. In some embodiments, at least about 50% of the nucleotides present in the direct or inverted repeat sequences are modified. In some embodiments, at least about 50% to 75% of the nucleotides present in the direct or inverted repeat sequences are modified. In some embodiments, at least about 100% of the nucleotides present in the direct or inverted repeat sequences are modified. In some embodiments, at least about 20% or more of the nucleotides present in the hairpins within the gRNA scaffold are modified. In some embodiments, at least about 50% or more of the nucleotides present in the hairpin of the gRNA scaffold are modified. In some embodiments, the guide comprises a variable-length spacer. In some embodiments, the guide comprises a 20-40 nucleotide spacer. In some embodiments, the guide comprises a spacer containing at least about 20-25 nucleotides or at least about 30-35 nucleotides. In some embodiments, the spacer comprises modified nucleotides. In some embodiments, the guide comprises two or more of the following: • At least about 1-5 nucleotides at the 5' end of the gRNA are modified and at least about 1-5 nucleotides at the 3' end of the gRNA are modified; • At least about 20% of the nucleotides present in either direct or inverted repeats are modified; • At least about 50-75% of the nucleotides present in the direct or inverted repeat sequence are modified; • At least about 20% or more of the nucleotides present in the hairpins of the gRNA scaffold are modified; • Variable-length spacers; and • Spacer regions containing modified nucleotides.
[0256] In the implementation scheme, the gRNA contains numerous modified nucleotides and / or chemical modifications. Such modifications can be... in vivo or in vitro This increases base editing by approximately two times. In the implementation method, the gRNA contains 2'- O -Methyl or thiophosphate modification. In one embodiment, the gRNA contains a 2'- O-Methyl and thiophosphate modifications. In one embodiment, the modifications increase base editing by at least about 2 times.
[0257] Directing polynucleotides may contain one or more modifications to provide new or enhanced features to nucleic acids. Directing polynucleotides may contain nucleic acid affinity tags. Directing polynucleotides may contain synthetic nucleotides, synthetic nucleotide analogs, nucleotide derivatives, and / or modified nucleotides.
[0258] gRNA or guide polynucleotides can also be modified by the following: 5' adenosine, 5' guanosine-triphosphate cap, 5' N7-methylguanosine-triphosphate cap, 5' triphosphate cap, 3' phosphate, 3' thiophosphate, 5' phosphate, 5' thiophosphate, Cis-Syn thymidine dimer, trimer, C12 spacer, C3 spacer, C6 spacer, dSpacer, PC spacer, rSpacer, spacer 18, spacer 9, 3'-3' modification, 2'- O -MethylthioPACE (MSP), 2'- O -Methyl-PACE (MP) and restricted ethyl (S-cEt), 5'-5' modification, debasing, acridine, azobenzene, biotin, biotin BB, biotin TEG, cholesterol TEG, desulfurized biotin TEG, DNP TEG, DNP-X, DOTA, dT-biotin, dibiotin, PC-biotin, psoralen C2, psoralen C6, TINA, 3' DABCYL, black hole quencher 1, black hole quencher 2, DABCYL SE, dT-DABCYL, IRDye QC-1, QSY-21, QSY-35, QSY-7, QSY-9, carboxyl linker, thiol linker, 2'-deoxyribonucleoside analog purine, 2'-deoxyribonucleoside analog pyrimidine, ribonucleoside analog, 2'- O -Methylribonucleoside analogs, sugar-modified analogs, swing / universal bases, fluorescent dye labeling, 2'-fluoroRNA, 2'- O -MethylRNA, methylphosphonate, phosphodiester DNA, phosphodiester RNA, thiophosphate DNA, thiophosphate RNA, UNA, pseudouridine-5'-triphosphate, 5'-methylcytidine-5'-triphosphate, or any combination thereof.
[0259] In some cases, phosphate-thiolated RNA gRNAs can inhibit RNase A, RNase T1, calf serum nuclease, or any combination thereof. These properties allow PS-RNA gRNAs to be used in... in vivo or in vitroUses in applications where there is a high likelihood of exposure to nucleases. For example, introducing a phosphate-thiocyanate (PS) bond between the last 3-5 nucleotides at the 5' or 3' end of the gRNA can inhibit exonuclease degradation. In some cases, a phosphate-thiocyanate bond can be added throughout the gRNA to reduce endonuclease attack.
[0260] Fusion proteins or complexes containing nuclear localization sequences (NLS) In some embodiments, the fusion protein or complex provided herein further comprises one or more ( For example 2, 3, 4, or 5) nuclear targeting sequences, such as nuclear localization sequences (NLS). In one embodiment, a bipartite NLS is used. In some embodiments, the NLS contains components that facilitate the entry of proteins (containing the NLS) into the cell nucleus. For example, The NLS is an amino acid sequence (via nuclear transport). In some embodiments, the NLS is fused to the N-terminus or C-terminus of the fusion protein. In some embodiments, the NLS is fused to the C-terminus or N-terminus of the nCas9 or dCas9 domain. In some embodiments, the NLS is fused to the N-terminus or C-terminus of the Cas12 domain. In some embodiments, the NLS is fused to the N-terminus or C-terminus of cytidine or adenosine deaminase. In some embodiments, the NLS is fused to the fusion protein via one or more linkers. In some embodiments, the NLS is fused to the fusion protein without a linker. In some embodiments, the NLS comprises the amino acid sequence of any NLS sequence provided or mentioned herein. Additional nuclear localization sequences are known in the art and will be apparent to those skilled in the art. For example, the NLS sequence is described in Plank. et al. The contents of PCT / EP2000 / 011690 are incorporated herein by reference because they disclose exemplary nuclear localization sequences.
[0261] In some embodiments, the NLS is present in the linker or a linker-side NLS, as described herein. A bipartite NLS comprises two basic amino acid clusters separated by a relatively short spacer sequence (thus bipartite into two parts, unlike a monopartite NLS). The nucleoplasmic protein NLS, KR[PAATKKAGQA]KKKK (SEQ ID NO: 191), is a prototype of the ubiquitous bipartite signal: a cluster of two basic amino acids separated by a spacer of approximately 10 amino acids. An exemplary bipartite NLS sequence is shown below: PKKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 328).
[0262] In some embodiments, any fusion protein or complex provided herein comprises an NLS containing the amino acid sequence EGADKRTADGSEFESPKKKRKV (amino acids 8 to 29 of SEQ ID NO: 328). In some embodiments, any adenosine base editor provided herein comprises an NLS containing the amino acid sequence EGADKRTADGSEFESPKKKRKV (amino acids 8 to 29 of SEQ ID NO: 328). In some embodiments, the NLS is located at the C-terminal portion of the adenosine base editor. In some embodiments, the NLS is located at the C-terminus of the adenosine base editor.
[0263] Other structural domains The base editor described herein may include any domain that facilitates the editing, modification, or alteration of the nucleobases of polynucleotides. In some embodiments, the base editor comprises a polynucleotide programmable nucleotide-binding domain (…). For example Cas9), nucleobase editing domain ( For example (a deaminase domain) and one or more additional domains. In some embodiments, the additional domains may contribute to the enzymatic or catalytic function of the base editor, the binding function of the base editor, or be inhibitors of cellular mechanisms that may interfere with the desired base editing outcome. For example (Enzymes). In some implementations, the base editor includes nuclease, nickase, recombinase, deaminase, methyltransferase, methyltransferase, acetyltransferase, transcription activator, or transcription repressor domains.
[0264] In some embodiments, the base editor comprises a uracil glycosylase inhibitor (UGI) domain. In some embodiments, cellular DNA repair responses in the presence of U:G heteroduplex DNA can lead to reduced nuclear base editing efficiency in the cell. In such embodiments, uracil DNA glycosylase (UDG) catalyzes the removal of U from the DNA in the cell, which can initiate base excision repair (BER), primarily resulting in the restoration of U:G pairs to C:G pairs. In such embodiments, BER can be suppressed in base editors comprising one or more domains that bind single strands, block edited bases, inhibit UGI, inhibit BER, protect edited bases, and / or promote the repair of non-edited strands. Therefore, this disclosure contemplates base editor fusion proteins or complexes comprising a UGI domain and / or a uracil stabilizing protein (USP) domain.
[0265] Base Editor System This document provides systems, compositions, and methods for editing nucleobases using a base editor system. In some embodiments, the base editor system comprises (1) a base editor (BE) that includes a multinucleotide programmable nucleotide binding domain and a nucleobase editing domain for editing nucleobases. For example (1) deaminase domain; and (2) guide polynucleotides that bind to the polynucleotide programmable nucleotide binding domain. For example (Guiding RNA). In some embodiments, the base editor system is a cytidine base editor (CBE) or an adenosine base editor (ABE). In some embodiments, the polynucleotide programmable nucleotide-binding domain is a polynucleotide programmable DNA or RNA-binding domain. In some embodiments, the nucleobase editing domain is a deaminase domain. In some embodiments, the deaminase domain may be cytidine deaminase or cytosine deaminase. In some embodiments, the deaminase domain may be adenine deaminase or adenosine deaminase. In some embodiments, the adenosine base editor can deaminate adenine in DNA. In some embodiments, the base editor is capable of deaminating cytidine in DNA.
[0266] The use of the base editor system provided in this article includes the following steps: (a) making the subject's polynucleotide ( example like The target nucleotide sequence of (double-stranded or single-stranded DNA or RNA) contains a nucleobase editor. For example (Adenosine base editor or cytidine base editor) and guide polynucleotides ( For example The system contacts a base editor system for gRNA, wherein the target nucleotide sequence contains a targeted nucleobase pair; (b) induces strand separation of the target region; (c) converts a first nucleobase of the target nucleobase pair in a single strand of the target region to a second nucleobase; and (d) cleaves no more than one strand of the target region, wherein a third nucleobase complementary to the first nucleobase is replaced by a fourth nucleobase complementary to the second nucleobase. It should be understood that in some embodiments, step (b) is omitted. In some embodiments, the targeted nucleobase pair is multiple nucleobase pairs in one or more genes. In some embodiments, the base editor system provided herein is capable of multiplex editing of multiple nucleobase pairs in one or more genes. In some embodiments, multiple nucleobase pairs are located in the same gene. In some embodiments, multiple nucleobase pairs are located in one or more genes, wherein at least one gene is located at a different locus.
[0267] Components of a base editor system (e.g., deaminase domains, guide RNA, and / or polynucleotide programmable nucleotide-binding domains) can be covalently or non-covalently associated with each other. For example, in some embodiments, the deaminase domain can be targeted to a target nucleotide sequence via a polynucleotide programmable nucleotide-binding domain, optionally wherein the polynucleotide programmable nucleotide-binding domain is complexed with a polynucleotide (e.g., guide RNA). In some embodiments, the polynucleotide programmable nucleotide-binding domain can be fused to or linked to the deaminase domain. In some embodiments, the polynucleotide programmable nucleotide-binding domain can target the deaminase domain to a target nucleotide sequence via non-covalent interactions or association with the deaminase domain. For example, in some embodiments, nucleobase editing components ( For example The deaminase component contains an additional heterologous portion or domain that is capable of interacting, associating, or forming a complex with a corresponding heterologous portion, antigen, or domain that is part of a polynucleotide programmable nucleotide binding domain and / or a guide polynucleotide (e.g., guide RNA) complexed therewith. In some embodiments, the polynucleotide programmable nucleotide binding domain and / or the guide polynucleotide (e.g., guide RNA) complexed therewith contains an additional heterologous portion or domain that is capable of interacting, associating, or forming a complex with a corresponding heterologous portion, antigen, or domain that is part of a nucleobase editing domain (e.g., deaminase component). In some embodiments, the additional heterologous portion may be capable of binding to, interacting with, associating with, or forming a complex with a peptide. In some embodiments, the additional heterologous portion may be capable of binding to, interacting with, associating with, or forming a complex with a polynucleotide. In some embodiments, the additional heterologous portion may be capable of binding to a guide polynucleotide. In some embodiments, the additional heterologous portion may be capable of binding to a peptide linker. In some embodiments, the additional heterologous portion is capable of binding to a polynucleotide linker. The additional heterologous portion can be a protein domain. In some embodiments, the additional heterologous portion includes peptides, such as the 22-amino acid RNA-binding domain (N22p) of λ phage anti-termination protein N, the 2G12 IgG homodimer domain, ABI, antibodies ( For example Antibodies or fragments thereof that bind to components of a base editor system or their heterologous portions For exampleHeavy chain domain 2 (CH2) (MHD2) of IgM or heavy chain domain 2 (EHD2) of IgE, immunoglobulin Fc region, heavy chain domain 3 (CH3) of IgG or IgA, heavy chain domain 4 (CH4) of IgM or IgE, Fab, Fab2, microantibodies and / or ZIP antibodies, barnase-barstar dimer domain, Bcl-xL domain, calcineurin A (CAN) domain, cardiac phosphatase transmembrane pentamer domain, collagen domain, ComRNA-binding protein domain ( For example SfMu Com outer shell protein domain and SfMu Com binding protein domain), Cypophilic protein-Fas fusion protein (CyP-Fas) domain, Fab domain, Fe domain, cellulose protein folding domain, FK506 binding protein (FKBP) domain, mTOR FKBP binding domain (FRB) domain, folding domain, fragment X domain, GAI domain, GID1 domain, blood group glycoprotein A transmembrane domain, GyrB domain, Halo tag, HIV Gp41 trimerization domain, HPV45 oncoprotein E7 C-terminal dimer domain, hydrophobic polypeptide, K homology (KH) domain, Ku protein domain (e.g., Ku heterodimer), leucine zipper, LOV domain, mitochondrial antiviral signaling protein CARD filament domain, MS2 capsid protein domain (MCP), non-natural RNA aptamer ligand binding corresponding RNA motif / aptamer, parathyroid hormone dimerization domain, PP7 capsid protein (PCP) domain, PSD95-Dlgl-zo-1 (PDZ) domain, PYL domain, SNAP tag, SpyCatcher portion, SpyTag portion, streptavidin domain, streptavidin-binding protein domain, streptavidin-binding protein (SBP) domain, telomerase Sm7 protein domain ( For example Sm7 homoheptamer or monomeric Sm-like protein) and / or fragments thereof. In embodiments, additional heterologous portions include polynucleotides (e.g., RNA motifs), such as MS2 phage operon stem loops (…). For example,MS2, MS2 C-5 mutant, or MS2 F-5 mutant), non-natural RNA motifs, PP7 operon stem loops, SfMu phate Com stem loops, sterile α motifs, telomerase Ku-binding motifs, telomerase Sm7-binding motifs, and / or fragments thereof. Other non-limiting examples of heterologous portions include polypeptides having at least about 85% sequence identity with any one or more of SEQ ID NO: 380, 382, 384, 386-388 or fragments thereof. Other non-limiting examples of heterologous portions include polynucleotides having at least about 85% sequence identity with any one or more of SEQ ID NO: 379, 381, 383, 385 or fragments thereof.
[0268] In some cases, the components of the base editing system associate with each other through the interaction of leucine zipper domains (e.g., SEQ ID NO: 387 and 388). In other cases, the components of the base editing system associate with each other through polypeptide domains (e.g., FokI domains) to form a protein complex containing about, at least about, or no more than about 1, 2 (i.e., dimerized), 3, 4, 5, 6, 7, 8, 9, or 10 polypeptide domain units, optionally including alterations that reduce or eliminate their activity.
[0269] In some cases, the components of the base editing system associate with each other through the interaction of polyantibodies or fragments thereof (e.g., heavy chain domain 2 (CH2) of IgG, IgD, IgA, IgM, IgE, IgM (MHD2) or IgE (EHD2), the Fc region of immunoglobulin, heavy chain domain 3 (CH3) of IgG or IgA, heavy chain domain 4 (CH4) of IgM or IgE, Fab and Fab2). In some cases, the antibodies are dimers, trimers or tetramers. In embodiments, the dimer antibody binds to the peptide or polynucleotide component of the base editing system.
[0270] In some cases, components of a base editing system associate with each other through interactions between polynucleotide-binding protein domains and polynucleotides. In other cases, components of a base editing system associate with each other through interactions between one or more polynucleotide-binding protein domains and self-complementary and / or mutually complementary polynucleotides, such that the complementary binding of polynucleotides to each other associates their respective polynucleotide-binding protein domains.
[0271] In some cases, components of a base editing system associate with each other through interactions between peptide domains and small molecules (e.g., chemically inducing dimerization agents (CIDs), also known as "dimerizing agents"). Non-limiting examples of CIDs include Amara. et al., “A versatile synthetic dimerizer for the regulation of protein-protein interactions,” PNAS , 94:10618-10623 (1997); and Voß et al. “Chemically induced dimerization: reversible and spatiotemporal control of proteinfunction in cells,” Current Opinion in Chemical Biology Those disclosed in , 28:194-201 (2015), the disclosures of each of which are incorporated herein by reference in their entirety for all purposes. In some embodiments, the base editor inhibits base excision repair (BER) of the edited strand. In some embodiments, the base editor protects or binds to the unedited strand. In some embodiments, the base editor includes UGI activity or USP activity. In some embodiments, the base editor comprises an inosine-specific nuclease with no catalytic activity.
[0272] The base editor disclosed herein may include any domain, feature, or amino acid sequence that facilitates editing of the target polynucleotide sequence. For example, in some embodiments, the base editor includes a nuclear localization sequence (NLS). In some embodiments, the NLS of the base editor is located between a deaminase domain and a polynucleotide-programmable nucleotide-binding domain. In some embodiments, the NLS of the base editor is located at the C-terminus of the polynucleotide-programmable nucleotide-binding domain.
[0273] Protein domains contained in fusion proteins can be heterologous functional domains. Non-limiting examples of protein domains that can be contained in fusion proteins include deaminase domains (e.g., cytidine deaminase or adenosine deaminase), uracil glycosylation inhibitor (UGI) domains, epitope tags, and reporter gene sequences.
[0274] In some implementations, an adenosine base editor (ABE) can deamination of adenine in DNA. In some implementations, ABE is performed using natural or engineered [materials / technology]. E. coliThe ABE3 is generated by replacing the APOBEC1 component with TadA, human ADAR2, mouse ADA, or human ADAT2. In some embodiments, the ABE comprises an evolved variant of TadA. In some embodiments, the base editor is ABE8.1, which comprises the following sequence or a fragment thereof having adenosine deaminase activity, or is substantially composed of the following sequence or a fragment thereof having adenosine deaminase activity: SEQ ID NO: 331. Other ABE8 sequences are provided in the appended sequence listing (SEQ ID NO: 332-354).
[0275] In some embodiments, the base editor comprises an adenosine deaminase variant containing an altered amino acid sequence relative to the ABE7*10 reference sequence as described herein. The term "monomer" as used in Table 7 refers to the monomeric form of TadA*7.10 containing the alteration. The term "heterodimer" as used in Table 7 refers to a designated wild-type fused with TadA*7.10 containing the alteration. E. coli TadA adenosine deaminase.
[0276] Table 7. Variants of adenosine deaminase base editor
[0277] In some implementations, the base editor comprises a domain containing all or part (e.g., the functional portion) of a uracil glycosylase inhibitor (UGI) or uracil stable protein (USP) domain.
[0278] connector In some embodiments, the linker can be used to connect any peptide or peptide domain of this disclosure. The linker can be as simple as a covalent bond, or it can be a polymeric linker with a length of many atoms. In some embodiments, the linker is a polypeptide or amino acid-based. In other embodiments, the linker is not peptide-like. In some embodiments, the linker is a covalent bond (…). example like Carbon-carbon bonds, disulfide bonds, carbon-heteroatom bonds wait ).
[0279] In some embodiments, any fusion protein provided herein comprises a cytidine or adenosine deaminase and a Cas9 domain fused together via a linker. Various linker lengths and flexibilityes may be employed between the cytidine or adenosine deaminase and the Cas9 domain. For example The range extends from highly flexible connector types (GGGS) n (SEQ ID NO: 246), (GGGGS) n (SEQ ID NO: 247) and (G)n to a more rigid joint form (EAAAK) n(SEQ ID NO: 248), (SGGS) n (SEQ ID NO: 355), SGSETPGTSESATPES (SEQ ID NO: 249) (See also, For example Guilinger JP et al. Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification. Nat. Biotechnol. 2014; 32(6): 577-82; the entire contents are incorporated herein by reference) and (XP)n) to obtain the optimal active length of the cytidine or adenosine deaminase nucleobase editor. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15. In some embodiments, the linker contains the (GGS)n motif, where n is 1, 3 or 7. In some embodiments, the cytidine deaminase or adenosine deaminase and the Cas9 domain of any fusion protein provided herein are fused via a linker (also referred to as the XTEN linker) containing the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 249).
[0280] In some implementations, the base editor's domain is fused via a linker comprising the following amino acid sequence: SGGSSGSETPGTSESATPESSGGS (SEQ ID NO: 356), SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 357) or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS (SEQ ID NO: 358).
[0281] In some embodiments, the base editor domain is fused via a linker (also referred to as an XTEN linker) comprising the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 249). In some embodiments, the linker comprises the amino acid sequence SGGS (SEQ ID NO: 355). In some embodiments, the linker is 24 amino acids long. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES (SEQ ID NO: 359). In some embodiments, the linker is 40 amino acids long. In some embodiments, the linker comprises the amino acid sequence: SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS (SEQ ID NO: 360). In some embodiments, the linker is 64 amino acids long. In some embodiments, the linker comprises the amino acid sequence: SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 361). In some embodiments, the linker is 92 amino acids long. In some embodiments, the linker comprises the amino acid sequence: PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS (SEQ ID NO: 362).
[0282] In some embodiments, the linker contains multiple proline residues and has a length of 5-21, 5-14, 5-9, or 5-7 amino acids. For example PAPAP (SEQ ID NO: 363), PAPAPA (SEQ ID NO: 364), PAPAPAP (SEQ ID NO: 365), PAPAPAPA (SEQ ID NO: 366), P(AP)4 (SEQ ID NO: 367), P(AP)7 (SEQ ID NO: 368), P(AP)10 (SEQ ID NO: 369) (see, For exampleTan J, Zhang F, Karcher D, Bock R. Engineering of high-precision base editors for site-specific single nucleotide replacement. Nat Commun. Jan 25, 2019; 10 (1): 439; all contents are incorporated herein by reference. Such proline-rich adapters are also called “rigid” adapters.
[0283] Nucleic acid programmable DNA-binding proteins with guiding RNA This article provides compositions and methods for base editing in cells. This article also provides compositions comprising a guiding polynucleotide sequence. For example Guide RNA sequences, or combinations of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more guide RNAs as provided herein. In some embodiments, the compositions for base editing as provided herein also comprise encoding a base editor (…). For example A polynucleotide (C-base editor or A-base editor). For example, a composition for base editing may comprise an mRNA sequence encoding a combination of BE, BE4, ABE, and one or more guide RNAs provided herein. A composition for base editing may comprise a base editor polypeptide and a combination of one or more of the guide RNAs provided herein. Such compositions can be used to achieve base editing in cells via various delivery methods, such as electroporation, nuclear transfection, viral transduction, or transfection. In some embodiments, a composition for base editing comprises an mRNA sequence encoding a base editor for electroporation provided herein and a combination of one or more guide RNA sequences.
[0284] Some aspects of this disclosure provide for any fusion protein or complex provided herein, as well as a nucleic acid programmable DNA-binding protein (napDNAbp) domain associated with the fusion protein or complex. For example Cas9 ( For example A system in which dCas9, nuclease activity (Cas9 or Cas9 cleavage enzyme), or Cas12 binds to guide RNA. These complexes are also known as ribonucleoproteins (RNPs). In some implementations, the guide nucleic acid ( For exampleThe guide RNA is a sequence 15-100 nucleotides long and contains at least 10 consecutive nucleotides complementary to the target sequence. In some embodiments, the target sequence is a DNA sequence. In some embodiments, the target sequence is an RNA sequence. In some embodiments, the target sequence is a sequence in the genome of bacteria, yeast, fungi, insects, plants, or animals. In some embodiments, the target sequence is a sequence in the human genome. In some embodiments, the 3' end of the target sequence is adjacent to a canonical PAM sequence (NGG). In some embodiments, the 3' end of the target sequence is adjacent to a non-canonical PAM sequence (NGG). For example The sequences listed in Table 3 or 5'-NAA-3' are used. In some implementations, the nucleic acid is guided ( For example (guide RNA) and genes of interest ( For example Sequence complementation in genes (related to diseases or symptoms).
[0285] Some aspects of this disclosure provide methods for using the fusion proteins or complexes provided herein. For example, some aspects of this disclosure provide methods including contacting a DNA molecule with any fusion protein or complex provided herein and at least one guide RNA, wherein the guide RNA is a sequence of about 15-100 nucleotides in length and contains at least 10 consecutive nucleotides complementary to a target sequence.
[0286] The domains of the base editor disclosed in this paper can be arranged in any order.
[0287] The defined target region can be a deamination window. A deamination window can be a defined region in which a base editor acts on the target nucleotide and causes it to deaminate. In some embodiments, the deamination window is within a 2, 3, 4, 5, 6, 7, 8, 9, or 10-base region. In some embodiments, the deamination window is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 bases upstream of the PAM.
[0288] The base editor disclosed herein may contain any domain, feature, or amino acid sequence that facilitates the editing of target polynucleotide sequences.
[0289] Methods using fusion proteins or complexes containing cytidine or adenosine deaminase and a Cas9 domain. Some aspects of this disclosure provide methods for using the fusion proteins or complexes provided herein. For example, some aspects of this disclosure provide methods including contacting a DNA molecule with any fusion protein or complex provided herein and at least one guide RNA described herein.
[0290] In some embodiments, the fusion proteins or complexes disclosed herein are used to edit target genes of interest. Specifically, the cytidine deaminase or adenosine deaminase nucleobase editors described herein are capable of making multiple mutations within the target sequence. These mutations can affect the function of the target. For example, when a regulatory region is targeted using a cytidine deaminase or adenosine deaminase nucleobase editor, the function of the regulatory region is altered and the expression of downstream proteins is reduced or eliminated.
[0291] Base editor efficiency In some implementations, the methods provided herein aim to alter genes and / or gene products via gene editing. The nucleobase editing proteins provided herein can be used for in vitro or in vivo gene-editing-based human therapies. Those skilled in the art will understand that the nucleobase editing proteins provided herein (… For example Contains a multinucleotide programmable nucleotide binding domain ( For example Cas9) and nucleobase editing domains ( For example Fusion proteins or complexes of adenosine deaminase domain or cytidine deaminase domain can be used to edit nucleotides from A to G or C to T.
[0292] Advantageously, the base editing systems provided herein offer genome editing without generating double-strand DNA breaks, do not require donor DNA templates, and do not induce excessive random insertions and deletions as CRISPR might. In some embodiments, this disclosure provides a base editor that, in nucleic acids ( For example The system effectively generates the desired mutations (such as stop codons) in the nucleic acids within the subject's genome, without generating a large number of undesired mutations (such as undesired point mutations).
[0293] The base editor disclosed herein advantageously modifies specific nucleotide bases encoding proteins without generating a significant proportion of insertions or deletions (i.e., insertions or deletions). Such insertions or deletions can lead to frameshift mutations within the coding region of a gene.
[0294] In some implementations, the base editor provided herein is able to generate a ratio greater than 1:1 of expected mutation to insertion / deletion. Right now(Expected point mutations: Unexpected point mutations). In some implementations, the base editor provided herein is capable of generating expected mutation to insertion / deletion ratios of at least 1.5:1, at least 2:1, at least 2.5:1, at least 3:1, at least 3.5:1, at least 4:1, at least 4.5:1, at least 5:1, at least 5.5:1, at least 6:1, at least 6.5:1, at least 7:1, at least 7.5:1, at least 8:1, at least 10:1, at least 12:1, at least 15:1, at least 20:1, at least 25:1, at least 30:1, at least 40:1, at least 50:1, at least 100:1, at least 200:1, at least 300:1, at least 400:1, at least 500:1, at least 600:1, at least 700:1, at least 800:1, at least 900:1, or at least 1000:1 or greater. The number of expected mutations and insertions / deletions can be determined using any suitable method.
[0295] In some embodiments, the base editors provided herein can restrict the formation of insertions and deletions in nucleic acid regions. In some embodiments, the regions are located at the nucleotide targeted by the base editor or within 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides of the nucleotide targeted by the base editor. In some embodiments, any base editor provided herein can restrict the formation of insertions and deletions at nucleic acid regions to less than 1%, less than 1.5%, less than 2%, less than 2.5%, less than 3%, less than 3.5%, less than 4%, less than 4.5%, less than 5%, less than 6%, less than 7%, less than 8%, less than 9%, less than 10%, less than 12%, less than 15%, or less than 20%.
[0296] Base editing is often referred to as “modification,” such as genetic modification, gene modification, and modification of nucleic acid sequences, and this is clearly understood based on the context that said modification is a base editing modification. Therefore, base editing modification is a modification at the nucleotide base level (e.g., due to deaminase activity discussed throughout this disclosure), which then leads to changes in the gene sequence and may affect the gene product.
[0297] In some implementation schemes, modifications ( For example Single base editing (SLA) results in a reduction of gene-targeted expression by approximately or at least approximately 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100%, or to an undetectable level.
[0298] This disclosure provides adenosine deaminase variants with increased efficiency and specificity. For example(ABE8 variant). Specifically, the adenosine deaminase variants described herein are more likely to edit desired bases within polynucleotides and less likely to edit bases with unintended alterations. For example (bystander).
[0299] In some implementations, it is associated with a base editor containing ABE7 ( For example Compared to base editor systems containing one of the ABE8 base editor variants, the bystander edits or mutations of any base editing system described herein that includes one of the ABE8 base editor variants are reduced by at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99%.
[0300] In some implementations, any of the ABE8 base editor variants described herein exhibit higher base editing efficiency compared to the ABE7 base editor. In some implementations, compared to the ABE7 base editor ( For example Compared to ABE7.10, any ABE8 base editor variant described herein has at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, 105%, 110%, 115%, 120%, 125%, 130%, 135%, 140%, 145% higher base editing ratio. Base editing efficiencies of %, 150%, 155%, 160%, 165%, 170%, 175%, 180%, 185%, 190%, 195%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 450%, or 500%.
[0301] The ABE8 base editor variants described herein can be delivered to host cells via plasmids, vectors, LNP complexes, or mRNA. In some embodiments, any of the ABE8 base editor variants described herein are delivered to host cells as mRNA.
[0302] In some embodiments, the methods described herein (e.g., base editing methods) have off-target effects minimized to none. In some embodiments, the methods described herein (e.g., base editing methods) have chromosomal translocations minimized to none.
[0303] In some implementations, the base editing methods described herein resulted in approximately 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the cell population being successfully edited.
[0304] In some embodiments, the percentage of viable cells in the cell population after base editing intervention is greater than at least 60%, 70%, 80%, or 90% of the initial cell population at the time of the base editing event. In some embodiments, the percentage of viable cells in the edited cell population is about 70%. In some embodiments, the percentage of viable cells in the edited cell population is about 75%. In some embodiments, the percentage of viable cells in the edited cell population is about 80%. In some embodiments, the percentage of viable cells in the cell population described above is about 85%. In some embodiments, the percentage of viable cells in the cell population described above is about 90% of the cells in the population at the time of the base editing event, or about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.
[0305] In the embodiments, the cell population is a group of cells that come into contact with the base editor, complex, or base editor system of this disclosure.
[0306] The number of expected mutations and insertions / deletions can be determined using any suitable method, such as those described in, for example, International PCT Applications PCT / US2017 / 045381 (WO2018 / 027078) and PCT / US2016 / 058344 (WO2017 / 070632); Komor, AC et al. , “Programmable editing of a target base ingenomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, NM et al. , “Programmable base editing of A • T to G • C in genomicDNA without DNA cleavage” Nature 551, 464-471 (2017); and Komor, AC et al."Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity" Science Advances 3:eaao4774 (2017); the entire contents of which are hereby incorporated by reference.
[0307] In some embodiments, to calculate the insertion / deletion frequency, sequencing reads are scanned to obtain exact matches with two 10-bp sequences flanking a window where insertions / deletions may occur. If no exact match is found, the read is excluded from the analysis. If the length of this insertion / deletion window exactly matches the reference sequence, the read is classified as having no insertions / deletions. If the insertion / deletion window is two or more bases longer or shorter than the reference sequence, the sequencing read is classified as an insertion or a deletion, respectively. In some embodiments, the base editor provided herein can restrict the formation of insertions / deletions in nucleic acid regions. In some embodiments, the region is located at the nucleotide targeted by the base editor or within 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides of the nucleotide targeted by the base editor.
[0308] Multiple editing In some embodiments, the base editor systems provided herein are capable of multiplexing multiple nucleobase pairs in one or more genes or polynucleotide sequences. In some embodiments, the multiple nucleobase pairs are located in the same gene or one or more genes, wherein at least one gene is located at a different locus. In some embodiments, multiplexing includes one or more guide polynucleotides. In some embodiments, multiplexing includes one or more base editor systems. In some embodiments, multiplexing includes one or more base editor systems having a single guide polynucleotide or multiple guide polynucleotides. In some embodiments, multiplexing includes one or more guide polynucleotides and a single base editor system. It should be understood that the features of multiplexing using any base editor as described herein can be applied to any combination of methods using any base editor provided herein. It should also be understood that multiplexing using any base editor as described herein can include sequential editing of multiple nucleobase pairs.
[0309] In some implementations, a base editor system capable of multiple editing of multiple nucleobase pairs in one or more genes includes one of the ABE7, ABE8, and / or ABE9 base editors.
[0310] Expression of fusion proteins or complexes in host cells Fusion proteins or complexes of this disclosure containing adenosine deaminase variants can be expressed in virtually any host cell of interest (including, but not limited to, bacterial, yeast, fungal, insect, plant, and animal cells) using conventional methods known to those skilled in the art. For example, DNA encoding an adenosine deaminase of this disclosure can be cloned by designing suitable primers for upstream and downstream of the CDS based on the cDNA sequence. The cloned DNA can be directly ligated to DNA encoding one or more additional components of the base editing system, or ligated after digestion with restriction enzymes if needed, or after the addition of suitable adapters and / or nuclear localization signals. The base editing system is translated in the host cell to form the complex.
[0311] The polynucleotides encoding the polypeptides described herein can be obtained through chemical synthesis of polynucleotides or by constructing full-length polynucleotides (e.g., DNA) encoding them by linking partially overlapping oligonucleotides synthesized using PCR and Gibson assembly methods. The advantage of constructing full-length polynucleotides through chemical synthesis or a combination of PCR or Gibson assembly methods is that the codons to be used can be selected based on the host in which the polynucleotide is introduced. In the expression of heterologous DNA molecules, it is expected that increasing protein expression levels will be achieved by converting their DNA sequence to codons frequently used in the host organism. Codon usage data from host cells (e.g., codon usage data available at kazusa.or.jp / codon / index.html) can be used to guide codon optimization of polynucleotide sequences encoding polypeptides. Codons used less frequently in the host can be converted to codons encoding the same amino acids that are used more frequently.
[0312] Expression vectors containing polynucleotides encoding nucleic acid sequence recognition modules and / or nucleic acid base converting enzymes can be generated, for example, by linking DNA downstream of a promoter in a suitable expression vector.
[0313] As a vehicle for expression, use E. coli plasmids from which the source ( For example , pBR322, pBR325, pUC12, pUC13); Bacillus subtilis plasmids from which the source ( For example pUB110, pTP5, pC194); yeast-derived plasmids ( For example pSH19, pSH15); insect cell expression plasmids ( For example pFast-Bac); animal cell expression plasmids ( For example pA1-11, pXT1, pRc / CMV, pRc / RSV, pcDNAI / Neo); bacteriophages such as .λ phage; insect viral vectors such as baculoviruses, etc. For example(e.g., BmNPV, AcNPV); animal viral vectors such as retroviruses, vaccinia viruses, adenoviruses, etc.
[0314] Regarding the promoter to be used, any promoter suitable for the host for gene expression can be used. In conventional methods using double-strand breaks, since host cell survival is sometimes significantly reduced due to toxicity, it is desirable to increase cell numbers at the start of induction by using inducible promoters. However, since sufficient cell proliferation can also be provided by expressing the nuclease complex disclosed herein, constitutive promoters can be used without restriction.
[0315] For example, when the host is an animal cell, the SR.α promoter, SV40 promoter, LTR promoter, cytomegalovirus (CMV) promoter, Rous sarcoma virus (RSV) promoter, Moloney mouse leukemia virus (MoMuLV), LTR, and herpes simplex virus thymidine kinase (HSV-TK) promoter can be used. Among these, the CMV promoter and SR.α promoter are particularly suitable.
[0316] When the host is E. coli At this time, the trp promoter, lac promoter, recA promoter, λ.P.sub.L promoter, lpp promoter, T7 promoter, etc. can be used.
[0317] When the host is Bacillus At this time, the SPO1 promoter, SPO2 promoter, penP promoter, etc. can be used.
[0318] When the host is yeast, promoters such as Gal1 / 10, PHO5, PGK, GAP, and ADH can be used.
[0319] When the host is an insect cell, polyhedrome promoters, P10 promoters, etc. can be used.
[0320] When the host is a plant cell, the CaMV35S promoter, CaMV19S promoter, NOS promoter, etc. can be used.
[0321] In addition to those mentioned above, the expression vectors used in this disclosure may include enhancers, splicing signals, terminators, polyA addition signals, selection markers such as drug resistance genes, auxotrophic complement genes, etc., and may use origin of replication, etc.
[0322] RNA encoding the protein domains described herein can be, for example... in vitro Prepared by transcribing the nucleic acid sequence encoding any fusion protein or complex disclosed herein.
[0323] The disclosed fusion proteins or complexes can be expressed intracellularly by introducing an expression vector containing a nucleic acid sequence encoding a fusion protein or complex into the cell.
[0324] Host cells of interest include, but are not limited to, bacterial, yeast, fungal, insect, plant, and animal cells. For example, host cells may include those from… E. coli Bacteria of the genus, such as E. coli K12.cndot.DH1[Proc. Natl.Acad. Sci. USA, 60, 160 (1968)], E. coli JM103 [Nucleic Acids Research, 9, 309(1981)]、 E. coli JA221 [Journal of Molecular Biology, 120, 517 (1978)], Escherichia coli bacteria HB101 [Journal of Molecular Biology, 41, 459 (1969)], E. coli C600 [Genetics, 39, 440 (1954)] etc.
[0325] Host cells may include Bacillus bacteria, such as Bacillus subtilis M1114 [Gene, 24, 255(1983)], Bacillus subtilis 207-21 [Journal of Biochemistry, 95, 87 (1984)] etc.
[0326] The host cell can be a yeast cell. Examples of yeast cells include... Saccharomyces cerevisiae cerevisiae) AH22, AH22R.sup.-, NA87-11A, DKD-5D, 20B-12, Saccharomyces cerevisiae (Schizosaccharomyces pombe) NCYC1913, NCYC2036 Pichia pastoris KM71, etc.
[0327] In cases where the viral delivery method utilizes the AcNPV virus, cells derived from a cabbage armyworm larvae can be used as the established line. Fall armyworm (Spodoptera frugiperda) Cells; Sf cells), derived from Powdered Noctuid Moth (Trichoplusia ni) MG1 cells in the midgut, HighFive cells derived from the ovaries of the moth *Spodoptera litura* TM cell, Cabbage cutworm (Mamestra brassicae) Derived cells, Salt marsh moth (Estigmena acrea)Derived cells, etc. When the virus is BmNPV, use... Domestic silkworm (Bombyx mori) The cells from which the lineage was established ( silkworm N cells; BmN cells, etc. As Sf cells, for example, Sf9 cells (ATCC CRL1711), Sf21 cells [all above, InVivo, 13, 213-217 (1977)], etc.
[0328] Insects can be any insect, such as silkworm larvae, fruit flies, crickets, etc. [Nature, 315, 592 (1985)].
[0329] The animal cells anticipated in this disclosure include, but are not limited to, cell lines such as monkey COS-7 cells, monkey Vero cells, Chinese hamster ovary (CHO) cells, dhfr gene-deficient CHO cells, mouse L cells, mouse AtT-20 cells, mouse myeloma cells, rat GH3 cells, human FL cells, etc., pluripotent stem cells (e.g., iPS cells), ES cells derived from humans and other mammals, and primary cultured cells prepared from various tissues. Furthermore, zebrafish embryos, African clawed toad Oocytes, etc.
[0330] This disclosure also considers plant cells. Plant cells are used, including but not limited to those made from various plants ( For example Grains such as rice, wheat, and corn; product crops such as tomatoes, cucumbers, and eggplants; garden plants such as carnations and lisianthus (Eustoma russellianum); and other plants such as tobacco, Arabidopsis thaliana Suspension cultured cells, callus, protoplasts, leaves, roots, etc., prepared by (etc.).
[0331] All of the above-mentioned host cells can be haploid (monoploid) or polyploid (monoploid). For example (Diploid, triploid, tetraploid, etc.). Using conventional methods, mutations introduced into only one homologous chromosome in principle produce heterologous cells. Therefore, unless the mutation is dominant, the desired phenotype will not be expressed. For recessive mutations, obtaining homozygous cells may be inconvenient due to labor and time requirements. Conversely, according to this disclosure, since mutations can be introduced into any allele on a homologous chromosome of the genome, even in the case of recessive mutations, the desired phenotype can be expressed in a single generation, thus solving the problems associated with conventional mutagenesis methods.
[0332] Expression vectors can be used according to the type of host through known methods ( For example Lysozyme method, competent cell method, PEG method, CaCl2 coprecipitation method, electroporation, microinjection, particle gun method, lipid transfection, spp. of Agrobacterium (Mediated delivery, etc.) are introduced.
[0333] E. coli Transformation can be performed according to methods described, for example, in Proc. Natl. Acad. Sci. USA, 69, 2110 (1972), Gene, 17, 107 (1982).
[0334] The method described, for example, in Molecular & General Genetics, 168, 111 (1979), can be used to... bud spores of Bacillus Introduce a carrier.
[0335] Yeast can be introduced into the vector according to methods described, for example, in Methods in Enzymology, 194, 182-187 (1991), Proc. Natl. Acad. Sci. USA, 75, 1929 (1978).
[0336] Insect cells and insects can be introduced into the vector according to methods described, for example, in Bio / Technology, 6, 47-55 (1988).
[0337] The vector can be introduced into animal cells according to methods described, for example, in Cell Engineering Supplement 8, New Cell Engineering Experiment Protocol, 263-267 (1995) (published by Shujunsha) and Virology, 52, 456 (1973).
[0338] Cells containing the vector can be cultured using known methods, depending on the host species. For example, in culturing... E. coli or Bacillus Liquid culture media can be used for cultivation. In one embodiment, the culture medium contains carbon sources, nitrogen sources, and inorganic substances necessary for the growth of the transformant. Examples of carbon sources include glucose, dextrin, soluble starch, and sucrose; examples of nitrogen sources include inorganic or organic substances such as ammonium salts, nitrates, corn steep liquor, peptone, casein, meat extracts, soybean meal, and potato extracts; and examples of inorganic substances include calcium chloride, sodium dihydrogen phosphate, and magnesium chloride. The culture medium may contain yeast extract, vitamins, and growth promoters. The pH of the culture medium is between approximately 5 and approximately 8.
[0339] As for cultivation E. coliThe culture medium can be, for example, M9 medium containing glucose and casein amino acids [Journal of Experiments in Molecular Genetics, 431-433, Cold SpringHarbor Laboratory, New York 1972]. For example, if necessary, additives such as 3β-indolylacrylic acid can be added to the medium to ensure effective promoter function. Culture is generally carried out at approximately 15 to approximately 43°C. E. coli If necessary, ventilation and stirring can be performed.
[0340] Generally, it is cultured at approximately 30 to 40°C. Bacillus If necessary, ventilation and stirring can be performed.
[0341] Examples of culture media used for culturing yeast include Burkholder Minimal Medium [Proc. Natl. Acad. Sci. USA, 77, 4505 (1980)] and SD Medium containing 0.5% casein amino acids [Proc. Natl. Acad. Sci. USA, 81, 5330 (1984)]. The pH of the medium can be between approximately 5 and approximately 8. Incubation is typically carried out at approximately 20°C to approximately 35°C. Aeration and stirring may be performed if necessary.
[0342] As a culture medium for culturing insect cells or insects, for example, Grace's insect medium [Nature, 195, 788 (1962)] containing appropriate additives such as inactivated 10% bovine serum is used. In one embodiment, the pH of the medium can be between about 6.2 and about 6.4. Culturing is generally carried out at about 27°C. Aeration and stirring may be performed if necessary.
[0343] As culture media for culturing animal cells, for example, minimal basal medium (MEM) containing approximately 5% to approximately 20% fetal bovine serum is used [Science, 122, 501 (1952)], Durbecco's modified Eagle medium (DMEM) [Virology, 8, 396 (1959)], RPMI 1640 medium [The Journal of the American Medical Association, 199, 519 (1967)], and 199 medium [Proceeding of the Society for the Biological Medicine, 73, 1 (1950)]. The pH of the medium is approximately 6 to approximately 8. Culture is usually carried out at approximately 30°C to approximately 40°C. Aeration and stirring may be performed if necessary.
[0344] As a culture medium for culturing plant cells, for example, MS medium, LS medium, B5 medium, etc. are used. The pH of the medium is from about 5 to about 8. Incubation is usually carried out at about 20°C to about 30°C. Aeration and stirring may be performed if necessary.
[0345] When higher eukaryotic cells, such as animal cells, insect cells, and plant cells, are used as host cells, inducible promoters ( For example The base editing system disclosed herein will be encoded under the regulation of metallothionein promoters (induced by heavy metal ions), heat shock protein promoters (induced by heat shock), Tet-ON / Tet-OFF system promoters (induced by the addition or removal of tetracycline or its derivatives), steroid-responsive promoters (induced by steroid hormones or their derivatives), etc. For example By introducing polynucleotides (including adenosine deaminase variants) into host cells, adding inducing substances to (or removing from) the culture medium at appropriate stages to induce the expression of nucleic acid-modifying enzyme complexes, culturing for a given time to perform base editing, and introducing mutations into target genes, transient expression of the base editing system can be achieved.
[0346] Prokaryotic cells such as E. coli Inducible promoters can be utilized. Examples of inducible promoters include, but are not limited to, the lac promoter (induced by IPTG), the cspA promoter (induced by cold shock), and the araBAD promoter (induced by arabinose).
[0347] Alternatively, when higher eukaryotic cells such as animal cells, insect cells, or plant cells are used as host cells, the aforementioned inducible promoters can also be used as vector removal mechanisms. In other words, the vector contains the origin of replication that functions in the host cell and nucleic acids encoding the proteins required for replication. For example For animal cells, the expression of nucleic acids encoding proteins (such as SV40 on and large T antigen, oriP and EBNA-1, etc.) is regulated by the aforementioned inducible promoters. Therefore, although the vector can replicate autonomously in the presence of the inducing substance, autonomous replication is not available when the inducing substance is removed, and the vector naturally detaches with cell division (autonomous replication is not possible by adding tetracycline and doxycycline to the vector in the Tet-OFF system).
[0348] Immunoconjugates In some embodiments, the anti-CD45 antibody of this disclosure is an immunoconjugate (“anti-CD45 antibody immunoconjugate”) or a portion thereof, wherein the anti-CD45 antibody is conjugated to one or more heterologous molecules (such as, but not limited to, cytotoxic agents or imaging agents). The fusion of a cytotoxic agent with an anti-CD45 antibody may have therapeutic value. Cytotoxic agents include, but are not limited to, radioisotopes (e.g., At...). 211 I 131 I 125 Y 90 、Rel 86 、Rel 88 、Sml 53 Bi 212 P 32 Pb 212 The antibody may be conjugated to one or more cytotoxic agents, such as chemotherapeutic agents or drugs, growth inhibitors, toxins (e.g., vincristine, vinblastine, etoposide), doxorubicin, melphalan, mitomycin C, chlorambucil, daunorubicin, or other intercalating agents); growth inhibitors; enzymes and fragments thereof such as lysozymes; antibiotics; toxins such as small molecule toxins or enzyme-active toxins. In some embodiments, the antibody is conjugated to one or more cytotoxic agents, such as chemotherapeutic agents or drugs, growth inhibitors, toxins (e.g., bacterial, fungal, plant or animal-derived protein toxins, enzyme-active toxins or fragments thereof), or radioactive isotopes.
[0349] Anti-CD45 antibody immunoconjugates include antibody-drug conjugates (ADCs) in which an anti-CD45 antibody is conjugated to one or more drugs, including but not limited to maytansine (see U.S. Patents 5,208,020 and 5,416,064 and European Patent EP 0 425 235 B 1); olistatin, such as monomethylolistatin drug portions DE and DF. (MMAE and MMAF) (see U.S. Patent Nos. 5,635,483, 5,780,588, and 7,498,298); dorlastatin; chazim or its derivatives (see U.S. Patent Nos. 5,712,374, 5,714,586, 5,739,116, 5,767,285, 5,770,701, 5,770,710, 5,773,001, and 5,877,296; Hinman et al., Cancer Res. 53: 3336-3342 (1993); and Lode et al., Cancer Res. 58: 2925-2928 (1998)); anthracyclines, such as doxorubicin or doxorubicin (see Kratz et al., Current Med.). Chem. 13: 477-523 (2006); Jeffrey et al., Bioorganic & Med. Chem. Letters 16: 358-362 (2006); Torgov et al., Bioconj. Chem. 16: 717-721 (2005); Nagy et al., Proc. Natl. Acad. Sci. USA 97: 829-834 (2000); Dubowchik et al., Bioorg. & Med. Chem. Letters 12: 1529-1532 (2002); King et al., J. Med. Chem. 45: 4336-4343 (2002); and U.S. Patent No. 6,630,579); methotrexate; vindesine; taxanes such as docetaxel, paclitaxel, larotaxel, tesetaxel, and ortataxel; trichothecene; and CC1065.
[0350] Anti-CD45 antibody immunoconjugates also include those conjugated with enzyme-active toxins or fragments thereof, including but not limited to diphtheria A chain, non-conjugated active fragments of diphtheria toxin, exotoxin A chain (from Pseudomonas aeruginosa), ricin A chain, abrin A chain, sucralose root toxin A chain, α-arbusculin, tung oil protein, caryophyllin protein, American pokeweed protein (PAPI, PAPII, and PAP-S), bitter melon inhibitor, jatropha toxin, croton toxin, soapwort inhibitor, white tree toxin, mitogellin, localized arbusculin, phenolmycin, enoxamin, and trichothecenes.
[0351] In some embodiments, the anti-CD45 antibody is conjugated with a protein degrading agent, such as Frere, G. Methods in Cell Biology , 167:1-26 (2022) and Sasso, J. et al. Biochemistry , 62:601-623 (2023), or those international patent applications described in WO 2021 / 053555, WO 2021 / 249517, WO 2020 / 006264, WO 2008 / 115516, WO 2021 / 126805, WO 2021 / 178920, WO 2021 / 127080, WO 2014 / 094138, WO 2015 / 200795, WO 2017 / 117118 and WO 2020 / 079103, the disclosures of which are incorporated herein by reference in their entirety for all purposes. In some implementations, the degrading agent is CC-122, CC-220, CC-99282, CFT7455, DKY709, CR8, Glue01, HQ005, FPFT-2216, TMX-4116, Ergidomide, BTX-1188, MG-277, ZHX-1-161, Indisulam, E7820, dCeMM1, CQS, NRX-252114, NRX-252262, BI-3802, CCT369260, cyclosporine A, luckynis, sanglifehrin A, auxin, methyl jasmonate, lenalidomide, pomalyst, or thalidomide. Non-limiting examples of protein degrading agents suitable for the compositions, conjugates and / or methods of this disclosure include heterobifunctional degrading agents and molecular gel degrading agents.
[0352] Anti-CD45 antibody immunoconjugates also include those in which the anti-CD45 antibody is conjugated with a radioactive atom to form a radioactive conjugate. Exemplary radioactive isotopes include At.211 I 131 I 125 Y 90 Re 186 Re 188 、Sm 153 Bi 212 P 32 Pb 212 And radioactive isotopes of Lu.
[0353] Conjugates of anti-CD45 antibodies and cytotoxic agents can be made using a variety of known protein conjugates (e.g., linkers) (see Vitetta et al., Science 238: 1098 (1987)). Linkers can be “cleavable linkers” that promote the release of cytotoxic drugs from cells, such as acid-labile linkers, peptidase-sensitive linkers, light-labile linkers, dimethyl linkers, and disulfide-containing linkers (Chari et al., Cancer Res. 52: 127-131 (1992); U.S. Patent No. 5,208,020).
[0354] Delivery system Nucleic acid-based delivery of base editor systems This can be done by methods known in the art or as described herein. in vitro or in vivo The nucleic acid molecule encoding the base editor system according to this disclosure is administered to a subject or delivered to cells. For example, it can be administered via a vector ( For example Viral or non-viral vectors), or delivered via naked DNA, DNA complexes, lipid nanoparticles, or a combination of the foregoing components, containing deaminases ( For example A base editor system (cytidine or adenine deaminase). The base editor system can be delivered to cells using any method available in the art, including but not limited to physical methods (e.g., electroporation, particle gun, calcium phosphate transfection), viral methods, non-viral methods (e.g., liposomes, cationic methods, lipid nanoparticles, polymeric nanoparticles), or biological non-viral methods (e.g., attenuated bacteria, engineered bacteriophages, mammalian virus-like particles, biological liposomes, erythrocyte shadows, exosomes).
[0355] Nanoparticles (which can be organic or inorganic) can be used to deliver base editor systems or components thereof. Nanoparticles are well known in the art, and any suitable nanoparticle can be used to deliver base editor systems or components thereof, or nucleic acid molecules encoding such components. In one example, organic ( For exampleLipid and / or polymer nanoparticles are suitable as delivery media in certain embodiments of this disclosure. Non-limiting examples of lipid nanoparticles suitable for the methods of this disclosure include those described in International Patent Application Publications Nos. WO2022140239, WO2022140252, WO2022140238, WO2022159421, WO2022159472, WO2022159475, WO2022159463, WO2021113365, WO2024019936 and WO2021141969, the disclosure of each of which is incorporated herein by reference in its entirety for all purposes.
[0356] Viral vector The base editor described herein can be delivered together with a viral vector. In some embodiments, the base editor disclosed herein can encode on nucleic acids contained in a viral vector. In some embodiments, one or more components of the base editor system can be encoded on one or more viral vectors.
[0357] Viral vectors may include lentiviruses ( For example (based on HIV and FIV vectors), adenovirus ( For example AD100), retroviruses ( For example Moloney murine leukemia virus (MML-V), herpesvirus vector ( For example The formulation and dosage of lentiviruses, including HSV-2, rabies virus (see, for example, U.S. Patent Application Publication No. 2022 / 0290164 A1, the disclosure of which is incorporated herein by reference in its entirety for all purposes), and adeno-associated virus (AAV), or other plasmid or viral vector types, particularly using formulations and dosages from, for example, U.S. Patent No. 8,454,972 (Formulations and Doses of Adenovirus), U.S. Patent No. 8,404,658 (Formulations and Doses of AAV), and U.S. Patent No. 5,846,946 (Formulations and Doses of DNA Plasmids), as well as from clinical trials involving lentiviruses, AAV, and adenoviruses, and publications relating to said clinical trials. For example, for AAV, the route of administration, formulation, and dosage may be as in U.S. Patent No. 8,454,972 and as in clinical trials involving AAV. For adenovirus, the route of administration, formulation, and dosage may be as in U.S. Patent No. 8,404,658 and as in clinical trials involving adenovirus. For plasmid delivery, the route of administration, formulation, and dosage can be as described in U.S. Patent No. 5,846,946 and in clinical studies involving plasmids. Dosage can be based on or extrapolated to an average individual of 70 kg. For example(Male adults), and can be adjusted for patients, subjects, and mammals of different weights and species. Administration frequency is recommended for medical or veterinary practitioners (…). For example Within the authority of physicians and veterinarians, and depending on common factors including the patient's or subject's age, sex, general health condition, other ailments, and the specific disease or symptom being addressed, viral vectors can be injected into the tissue of interest. For cell type-specific base editing, the base editor and, optionally, the expression of the guiding nucleic acid can be driven by a cell type-specific promoter.
[0358] Viral vectors can be selected based on the application. For example, for in vivo gene delivery, AAV may be superior to other viral vectors. In some embodiments, AAV allows for low toxicity, possibly because the purification method does not require ultracentrifugation of cell particles that can activate an immune response. In some embodiments, AAV allows for a low probability of inducing insertional mutagenesis because it does not integrate into the host genome. Adenoviruses are commonly used as vaccines because they induce strong immunogenic responses. The packaging capacity of a viral vector can limit the size of the base editor that can be encapsulated into the vector.
[0359] AAVs have a packaging capacity of approximately 4.5 Kb or 4.75 Kb, comprising two 145-base inverted terminal repeat (ITR) sequences. This means that the disclosed base editor, along with the promoter and transcription terminator, can be assembled into a single viral vector. Constructs larger than 4.5 or 4.75 Kb can result in a significant reduction in viral production. For example, SpCas9 is quite large, with the gene itself exceeding 4.1 Kb, making it difficult to package into AAVs. Therefore, embodiments of this disclosure include utilizing a base editor that is shorter than a conventional base editor. In some examples, the base editor is less than 4 kb. The disclosed base editor may be less than 4.5 kb, 4.4 kb, 4.3 kb, 4.2 kb, 4.1 kb, 4 kb, 3.9 kb, 3.8 kb, 3.7 kb, 3.6 kb, 3.5 kb, 3.4 kb, 3.3 kb, 3.2 kb, 3.1 kb, 3 kb, 2.9 kb, 2.8 kb, 2.7 kb, 2.6 kb, 2.5 kb, 2 kb, or 1.5 kb. In some embodiments, the disclosed base editor is 4.5 kb or shorter.
[0360] AAV may be AAV1, AAV2, AAV5, AAV6, AAV9, PHP.EB, PHP.B, AAV.CAP-B10, AAV, CAP-B22, AAV-rh10, PAL family AAVs, or any combination thereof. In implementations, AAVs are capable of crossing the blood-brain barrier (see, for example, Liu...). et al.“Crossing the blood-brain barrier with AAV vectors,” Metabolic Brain Disease The AAV vectors disclosed in , 36:45-52 (2021), the contents of which are incorporated herein by reference in their entirety for all purposes. The type of AAV can be selected based on the cells to be targeted; For example AAV serotypes 1, 2, 5, or hybrid capsids AAV1, AAV2, AAV5, or any combination thereof, can be selected for targeting brain or neuronal cells; and AAV4 can be selected for targeting cardiac tissue. AAV8 can be used for delivery to the liver. A list of certain AAV serotypes for these cells can be found in Grimm, D. et al., J. Virol. 82: 5887-5911 (2008)). In some embodiments, the AAV vector contains a PAL family AAV capsid (see, Stanton, A.). Med 4:31-50 (2023) (doi: doi.org / 10.1016 / j.medj.2022.11.002), the contents of which are hereby incorporated herein by reference in their entirety for all purposes.
[0361] In some implementations, lentiviral vectors are used to transduce cells of interest with polynucleotides encoding a base editor or base editor system as provided herein. Lentivirals are complex retroviruses with the ability to infect and express their genes in mitotic and postmitotic cells. The most common lentivirus is human immunodeficiency virus (HIV), which uses envelope glycoproteins of other viruses to target a wide range of cell types.
[0362] In another embodiment, minimal nonprimate lentiviral vectors based on equine infectious anemia virus (EIAV) are also considered. In another embodiment, RetinoStat® is an EIAV-based lentiviral gene therapy vector expressing the angiogenesis inhibitors endostatin and angiostatin, intended for subretinal injection. In yet another embodiment, self-inactivated lentiviral vectors are considered.
[0363] Any RNA in the system, such as guide RNA or base editor-encoded mRNA, can be delivered in RNA form. Base editor-encoded mRNA can be generated using in vitro transcription. For example, nuclease mRNA can be synthesized using a PCR cassette containing the following elements: a T7 promoter, an optional kozak sequence (GCCACC), a nuclease sequence, and a 3' UTR, such as a 3' UTR from a β-globin-polyA tail. This cassette can be used for transcription of T7 polymerase. Guide polynucleotides can also be transcribed from a cassette containing a T7 promoter, followed by the sequence "GG," and a guide polynucleotide sequence using in vitro transcription. For example (gRNA).
[0364] Non-viral platform for gene transfer Nonviral platforms for introducing heterologous polynucleotides into cells of interest are known in the art.
[0365] For example, this disclosure provides a method for inserting heterologous polynucleotides into a cellular genome using a Cas9 or Cas12 (e.g., Cas12b) ribonucleoprotein complex (RNP)-DNA template complex, wherein the RNP comprises a Cas9 or Cas12 nuclease domain and a guide RNA, wherein the guide RNA specifically hybridizes to a target region of the cellular genome, and wherein the Cas9 nuclease domain cleaves the target region to create an insertion site in the cellular genome. The heterologous polynucleotide is then introduced using a DNA template. In embodiments, the DNA template is a double-stranded or single-stranded DNA template, wherein the DNA template is about 200 nucleotides or larger, and wherein the 5' and 3' ends of the DNA template contain nucleotide sequences homologous to the genomic sequence flanking the insertion site. In some embodiments, the DNA template is a single-stranded circular DNA template. In embodiments, the molar ratio of RNP to DNA template in the complex is about 3:1 to about 100:1.
[0366] In some implementations, the DNA template is a linear DNA template. In some instances, the DNA template is a single-stranded DNA template. In some implementations, the single-stranded DNA template is a pure single-stranded DNA template. In some implementations, the single-stranded DNA template is a single-stranded oligodeoxynucleotide (ssODN).
[0367] In other embodiments, single-stranded DNA (ssDNA) can generate efficient homology-directed repair (HDR) with minimal off-target integration. In one embodiment, ssDNA phages are used to efficiently and inexpensively produce circular ssDNA (cssDNA) donors. When used with Cas9 or Cas12 (e.g., Cas12a, Cas12b), these cssDNA donors act as efficient HDR templates, exhibiting integration frequencies superior to linear ssDNA (lssDNA) donors.
[0368] In some implementations, transposable elements such as transposons can be used to insert heterologous polynucleotides into the genome of a cell, such as Tipanee. Human Gene Therapy As described in November 2017, 1087-1104, DOI:10.1089 / hum.2017.128. Transposable elements are divided into two categories: retrotransposons and DNA transposons. Transposable elements can alter the host cell genome through insertion, replication, deletion, and translocation. Retrotransposons are described as mobile elements employing an RNA intermediate, which is first reverse transcribed into a complementary single-stranded (c) DNA strand by a reverse transcriptase encoded by the retrotransposon. Subsequently, the single-stranded DNA is converted into double-stranded DNA, which is then integrated into the host genome. This so-called “replication mechanism” produces several new copies of the retrotransposon, which expand throughout the target genome in the course of evolution. Retrotransposons are classified into many subtypes based on their long terminal repeat sequences and the DNA sequence of their open reading frames. Retrotransposons are used to integrate transgenes into the DNA of target cells, and in some cases, rely on adenovirus delivery. Alternatively, DNA transposons translocate via a “non-replicative mechanism,” whereby the two terminal inverted repeat (TIR) sequences are recognized and cleaved by a transposase, releasing a homologous DNA transposon with free DNA ends. The excised DNA transposon is then integrated into a new genomic region where the target site is recognized and cleaved by the same transposase. This cleavage and pasting mechanism typically replicates the DNA target site upon insertion, leaving target site duplication (TSD). Non-limiting examples of transposons include… Sleeping Beauty (SB) transposable piggyBac (PB) transposons and Tol2 Rotatable element.
[0369] Intended peptides Intrin (intercalary proteins) are autoprocessing domains present in many different organisms, which perform a process called protein splicing.
[0370] Non-limiting examples of inteins include any inteins or intein pairs known in the art, including synthetic inteins based on dnaE inteins, Cfa-N (e.g., split intein-N) and Cfa-C (e.g., split intein-C) intein pairs that have been described (e.g., in Stevens et al., J Am Chem Soc. 2016 Feb. 24; 138(7):2162-5, which are incorporated herein by reference), and DnaE. Non-limiting examples of intein pairs that may be used under this disclosure include: Cfa DnaE inteins, Ssp GyrB inteins, Ssp DnaX inteins, Ter DnaE3 inteins, Ter ThyX inteins, Rma DnaB inteins, and Cne Prp8 inteins (e.g., as described in U.S. Patent No. 8,394,604, which is incorporated herein by reference). Exemplary nucleotide and amino acid sequences of the integrities are provided in the sequence listing as SEQ ID NO: 370-377 and 389-424. Integrities suitable for embodiments of this disclosure and methods of use thereof are described in U.S. Patent No. 10,526,401, International Patent Application Publications Nos. WO 2013 / 045632, WO 2024 / 073385 and WO 2020 / 051561, and U.S. Patent Application Publication No. US 2020 / 0055900, the entire disclosures of which are incorporated herein by reference for all purposes.
[0371] Inteptide-N and inteptide-C can be fused to the N-terminal and C-terminal portions of cleaved Cas9, respectively, for conjugation of the N-terminal and C-terminal portions of cleaved Cas9. For example, in some embodiments, inteptide-N is fused to the C-terminus of the N-terminal portion of cleaved Cas9. Right now This forms an N-[N-terminal portion of cleaved Cas9]-[intein-N]-C structure. In some embodiments, the intein-C is fused to the N-terminus of the C-terminal portion of cleaved Cas9. Right now A structure is formed of N-[intein-C]--[cleaved C-terminal portion of Cas9]-C. In an embodiment, the base editor is encoded by two polynucleotides, one of which encodes a base editor fragment fused to the intein-N, and the other polynucleotide encodes a base editor fragment fused to the intein-C. Methods for designing and using inteins are known in the art and are described, for example, by WO2014004336, WO2017132580, WO2013045632A1, US20150344549, and US20180127780, each of which is incorporated herein by reference in its entirety.
[0372] In some implementations, ABE splits into N-terminal and C-terminal fragments at Ala, Ser, Thr, or Cys residues within selected regions of SpCas9. These regions correspond to ring regions identified by Cas9 crystal structure analysis.
[0373] The N-terminal fragment is fused to the inteptide-N at the C-terminus, and the C-terminal fragment is fused to the inteptide-C at the N-terminal amino acid selected from the group consisting of: S303, T310, T313, S355, A456, S460, A463, T466, S469, T472, T474, C574, S577, A589, and S590 (see SEQ ID NO: 197). In various embodiments, SpCas9 is split between amino acid positions 302 and 303, 309 and 310, 312 and 313, 354 and 355, 455 and 456, 459 and 460, 462 and 463, 465 and 466, 468 and 469, 471 and 472, 473 and 474, 573 and 574, 576 and 577, 588 and 589, or 589 and 590 (see SEQ ID NO: 197) to produce an N-terminal fragment and a C-terminal fragment, wherein the N-terminal fragment is fused with an inteptide-N at the C-terminus, and wherein the C-terminal fragment is fused with an inteptide-C at the N-terminus.
[0374] Pharmaceutical Composition In some aspects, this disclosure provides a pharmaceutical composition comprising any of the cells, polynucleotides, carriers, base editors, base editor systems, guide polynucleotides, fusion proteins, complexes, or fusion protein-guide polynucleotide complexes described herein.
[0375] The pharmaceutical compositions disclosed herein can be prepared using known techniques. See also, For exampleRemington, *The Science and Practice of Pharmacy* (21st edition, 2005). Generally, cells or populations thereof are mixed with a suitable carrier before administration or storage, and in some embodiments, the pharmaceutical composition also includes a pharmaceutically acceptable carrier. A suitable pharmaceutically acceptable carrier typically comprises an inert substance that facilitates administration of the pharmaceutical composition to a subject, facilitates the formulation of the pharmaceutical composition into a deliverable formulation, or facilitates the storage of the pharmaceutical composition prior to administration. Pharmaceutically acceptable carriers can include agents that can stabilize, optimize, or otherwise alter the form, consistency, viscosity, pH, pharmacokinetics, or solubility of the formulation. Such agents include buffers, wetting agents, emulsifiers, diluents, encapsulating agents, and skin penetration enhancers. For example, carriers may include, but are not limited to, saline, buffered saline, dextran, arginine, sucrose, water, glycerol, ethanol, sorbitol, dextran, sodium carboxymethyl cellulose, and combinations thereof.
[0376] In some embodiments, the pharmaceutical composition is formulated for delivery to a subject. Suitable routes of administration of the pharmaceutical compositions described herein include, but are not limited to: topical, subcutaneous, transdermal, intradermal, intralesional, intra-articular, intraperitoneal, intrabladder, transmucosal, transgingival, intradental, intracochlear, transtympanic, intra-organ, epidural, intrathecal, intramuscular, intravenous, intravascular, intraosseous, periorbital, intratumoral, intracerebral, and intraventricular administration.
[0377] In some embodiments, the pharmaceutical composition described herein is applied topically to the affected area. In some embodiments, the pharmaceutical composition described herein is administered to a subject by injection, via a catheter, via a suppository, or via an implant, said implant being a porous, non-porous, or gel-like material, including membranes such as sialastic membranes or fibers.
[0378] In some embodiments, any fusion protein, gRNA, and / or complex described herein is provided as part of a pharmaceutical composition. In some embodiments, the pharmaceutical composition comprises any fusion protein or complex provided herein. In some embodiments, the pharmaceutical composition comprises gRNA, nucleic acid-programmable DNA-binding protein, cationic lipid, and pharmaceutically acceptable excipients. In some embodiments, the pharmaceutical composition comprises lipid nanoparticles and pharmaceutically acceptable excipients. In some embodiments, the lipid nanoparticles contain gRNA, base editor, complex, base editor system, or a component thereof disclosed herein, and / or one or more polynucleotides encoding them. The pharmaceutical composition may optionally comprise one or more additional therapeutically active substances.
[0379] The composition described above can be administered in an effective amount. The effective amount will depend on the method of administration, the specific disease being treated, and the desired outcome. It may also depend on the stage of the disease, the age and physical condition of the subject, the nature of concurrent treatments (if any), and similar factors well known to medical practitioners. For therapeutic applications, the amount is sufficient to achieve the medically desired outcome.
[0380] In some embodiments, the compositions according to this disclosure can be used to treat any of a variety of diseases, conditions and / or disorders.
[0381] Treatment Some aspects of this disclosure provide methods for treating a subject in need, the methods comprising administering to the subject in need an effective therapeutic amount of a pharmaceutical composition as described herein. More specifically, the treatment method comprises administering to the subject in need one or more pharmaceutical compositions comprising one or more cells having at least one editing gene. In other embodiments, the methods of this disclosure comprise expressing or introducing into cells a base editor polypeptide and one or more guide RNAs capable of targeting a nucleic acid molecule encoding at least one polypeptide.
[0382] Those skilled in the art will recognize that multiple administrations of the pharmaceutical composition considered in a particular implementation may be required to achieve the desired treatment. For example, the composition may be administered to a subject 1, 2, 3, 4, 5, 6, 7, 8, 9 or more times over a span of 1 week, 2 weeks, 3 weeks, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 1 year, 2 years, 5 years, 10 years or more.
[0383] The administration of the pharmaceutical compositions considered herein can be performed using conventional techniques, including but not limited to infusion, intravenous infusion, or parenteral administration. In some embodiments, parenteral administration includes intravascular, intravenous, intramuscular, intra-arterial, intrathecal, intratumoral, intradermal, intraperitoneal, tracheal, subcutaneous, subepidermal, intra-articular, subcapsular, subarachnoid, and intrasternal infusion or injection.
[0384] Reagent test kit This disclosure provides a kit for treating a disease in a subject. In some embodiments, the kit further includes a base editor system or a polynucleotide encoding a base editor system, wherein the base editor polypeptide system comprises a nucleic acid programmable DNA-binding protein (napDNAbp), a deaminase, and guide RNA. In some embodiments, napDNAbp is Cas9 or Cas12. In some embodiments, the polynucleotide encoding the base editor is an mRNA sequence. In some embodiments, the deaminase is cytidine deaminase or adenosine deaminase. In some embodiments, the kit includes edited cells and instructions for using such cells.
[0385] The kit may further include written instructions for use with a base editor, base editor system, and / or edited cells as described herein. In other embodiments, the instructions include at least one of the following: precautions; warnings; clinical studies; and / or references. The instructions may be printed directly on the container (where present), affixed as a label to the container, or included as a separate sheet of paper, brochure, card, or in a folder provided with the container. In a further embodiment, the kit includes instructions in the form of a label or separate insert (packaging insert) with appropriate operating parameters. In yet another embodiment, the kit includes one or more containers with appropriate positive and negative controls or control samples for use as standards for detection, calibration, or normalization. The kit also includes a second container containing a pharmaceutically acceptable buffer, such as (sterile) phosphate-buffered saline, Ringer's solution, or dextran solution. It may also include other materials required from a commercial and user perspective, including additional buffers, diluents, filters, needles, syringes, and packaging inserts with instructions for use.
[0386] Unless otherwise stated, the practice of embodiments of this disclosure employs conventional techniques of molecular biology (including recombinant technologies), microbiology, cell biology, biochemistry, and immunology, which are well within the skill of a person skilled in the art. Such techniques are well explained in the literature, such as "Molecular Cloning: A Laboratory Manual", 2nd edition (Sambrook, 1989); "Oligonucleotide Synthesis" (Gait, 1984); "Animal Cell Culture" (Freshney, 1987); "Methods in Enzymology"; "Handbook of Experimental Immunology" (Weir, 1996); "Gene Transfer Vectors for Mammalian Cells" (Miller and Calos, 1987); "Current Protocols in Molecular Biology" (Ausubel, 1987); "PCR: The Polymerase Chain Reaction" (Mullis, 1994); and "Current Protocols in Immunology" (Coligan, 1991). These techniques are suitable for producing the polynucleotides and peptides of this disclosure and are therefore contemplated in the manufacture and practice of embodiments of this disclosure. Techniques particularly useful for specific implementations will be discussed in the following sections.
[0387] The following embodiments are provided to provide a complete disclosure and description of how to manufacture and use the assays, screening and treatments disclosed herein for those skilled in the art, and are not intended to limit the scope of the invention as perceived by the inventors.
[0388] Example Example 1: Non-genotoxic conditioning Experiments were conducted to demonstrate a method for non-genotoxicity modulation. Yeast display screening was used to identify amino acid alterations in the differentiation cluster 45 (CD45) peptide that resulted in reduced binding of mAb039-9 to the peptide (see Table 1). To demonstrate the effect of the amino acid alterations on the binding of mAb030-9 to the CD45 peptide, as... Figure 1AThe base editing system shown was used to perform base editing on MOLM13 and JURKAT cells using a base editor containing mRNA encoding adenosine deaminase and a base editor selected from gRNA1610, gRNA1611, gRNA1613, gRNA1614 and gRNA2696 (see Table 2). Figure 1B The maximum percentage of A-to-G base editing observed in cells at 3 and 7 days post-electroporation is shown. The base editor system correlated with the maximum percentage of A-to-G base editing (up to approximately 80%) at days 3 and 7, with an increase in base editing rate observed between days 3 and 7. Edited MOLM13 cells showed reduced binding to mAb30-9 antibody (Table 8;). Figure 2A and 2B ). Figure 2C A stacked bar graph showing the allelic composition of the edited cells is provided. CD45 variants containing the following alterations show reduced mAB030-9 binding ( Figures 3 to 5 The bases used to introduce CD45 into CD34+ cells were: Y232C; K231R and Y232C; E259G; N255G; N255D; N255G and N256G. Base editing to introduce amino acid substitutions of Y232C, K231R+Y232C, or E259G showed a high editing frequency in CD34+ cells. These data demonstrate that base editing can be used to alter CD45 polynucleotides in cells to encode CD45 peptides with reduced binding to anti-CD45 antibodies (e.g., mAb030-9).
[0389] Table 8. Binding of mAbU139AGL030-9 to mutant CD45 D1-D2 protein compared to binding to WT CD45 D1-D2.
[0390] To allow for further evaluation, the CD45 variant shown to have reduced binding with mAb030-9 was overexpressed and purified (Table 9). Figures 6A to 6C The CD45 peptides listed are described. The binding kinetics of purified CD45 variants containing alterations selected from E259G, N257G, E256G, E259K, N286D, N267S, N257S, E259R, N257D, and N267G with mAb030-9 were measured using a biosensor. Figures 6A to 6C Alterations to E259R, E259K, and N257S all showed similar or lower responses to phosphate-buffered saline, and all mutants containing one of the alterations selected from E259G, N257G, E256G, E259K, N286D, N267S, N257S, E259R, N257D, and N267G exhibited responses below 0.05 nm.
[0391] Table 9. Representative CD45 peptides generated in CHO-S cells using the CD45R0 overexpression construct.
[0392] Experiments were conducted to demonstrate reduced binding of mAb030-9 to cells that had undergone base editing to express the altered CD45 peptide. CD34+ HSPC cells were base-edited using the base editor systems described in Table 15, and the base editing rate and amino acid changes associated with each base editor system are listed in Tables 10–14 and 16. Cells edited using base editor systems EP9–EP15 all showed reduced binding compared to unedited (“simulated”) cells. Figure 7 ).
[0393] Table 10. Editing rates measured for the evaluated base editor systems. Nucleotide changes are referenced to the spacer region sequence of the gRNA.
[0394] Table 11. Editing rates measured for the evaluated base editor systems. Nucleotide changes are referenced to the spacer region sequence of the gRNA.
[0395] Table 12. Editing rates measured for the evaluated base editor systems. Nucleotide changes are referenced to the spacer region sequence of the gRNA.
[0396] Table 13. Editing rates measured for the evaluated base editor systems. Nucleotide changes are referenced to the spacer region sequence of the gRNA.
[0397] Table 14. Editing rates measured for the evaluated base editor systems. Nucleotide changes are referenced to the spacer region sequence of the gRNA.
[0398] Table 15. Description of base editor systems E1 to E18 and the volume of specified stock solutions used for transfecting cells using electroporation.
[0399] Table 16. Base editing efficiency of E259G changes performed using base editor systems EP9 to EP15.
[0400] The following materials and methods were used in the above embodiments.
[0401] CD34+ cell preparation Mobilized peripheral blood was obtained and human CD34+ hematopoietic stem cell and progenitor (HSPC) cells (HemaCare, M001F-GCSF / MOZ-2) were enriched. CD34+ HSPC cells were thawed and placed in X-VIVO10 (Lonza) containing 1% Glutamax (Gibco), 100 ng / mL TPO (Peprotech), SCF (Peprotech), and Flt-3 (Peprotech) for 48 hours prior to electroporation.
[0402] Electroporation of CD34+ cells Forty-eight hours after thawing, cells were decanted to remove X-VIVO 10 medium and washed in MaxCyte buffer (HyClone) with 0.1% HSA (Akron Biotechnologies). Cells were then resuspended in cold MaxCyte buffer at 1,250,000 cells per mL and aliquoted into multiple 20 μL aliquots. ABE mRNA and guide polynucleotides were then aliquoted according to experimental conditions (see Table 15) and added to a total of 5 μL in MaxCyte buffer. 20 μL cells were added in triplicate to each of the 5 μL RNA mixture and loaded into each chamber of an OC25x3 MaxCyte cuvette for electroporation. After receiving charge, 25 μL was collected from the chamber and placed in the center of each well of a 24-well untreated culture plate. Cells were incubated in an incubator (37°C, 5% CO2) for 20 min. After 20 minutes of recovery, X-VIVO 10 (hematopoietic cell culture medium) containing 1% Glutamax, 100 ng / mL TPO, SCF, and Flt-3 was added to the cells at a concentration of 1,000,000 cells per mL. The cells were then placed in an incubator (37°C, 5% CO2) for a further 48 hours of recovery.
[0403] Genomic DNA extraction from CD34+ cells Following electroporation (after 48 h), aliquots of cells were cultured in X-VIVO 10 medium (Lonza) containing 1% Glutamax (Gibco), 100 ng / mL TPO (Peprotech), stem cell factor (SCF) (Peprotech), and Flt-3 (Peprotech). After 48 h and 144 h of culture, 100,000 cells were collected and vortexed. 50 μL of rapid extract (Lucigen) was added to the cell pellet, and the cell mixture was transferred to a 96-well PCR plate (Bio-Rad). The lysates were heated at 65 °C for 15 min, followed by heating at 98 °C for 10 min. Cell lysates were stored at -20 °C.
[0404] Combining dynamics 100 nM mAb030-9 was immobilized onto the biosensor using anti-human IgG Fc capture (AHC). 300 nM of purified CD45 protein was then contacted with the immobilized antibody. After 300 s, the biosensor was immersed in buffer containing no purified CD45 protein.
[0405] sequence Table 17 below provides the amino acid sequences of base editors encoded by specified mRNA sequences.
[0406] Table 17. Base editor amino acid sequences encoded by specified mRNA sequences.
[0407] Other implementation plans As will be apparent from the foregoing description, changes and modifications can be made to the various embodiments disclosed herein to adapt them to a variety of uses and conditions. Such embodiments are also within the scope of the following claims.
[0408] The description of the list of elements in any definition of a variable in this document includes the definition of the variable as any single element or a combination (or sub-combination) of the listed elements. The description of an implementation scheme in this document includes the implementation scheme as any single implementation scheme or in combination with any other implementation scheme or part thereof.
[0409] All patents and publications mentioned in this specification are incorporated herein by reference as if each individual patent and publication were specifically and separately indicated to be incorporated by reference. This application may relate to PCT / US2020 / 048586, filed August 28, 2020, the disclosure of which is incorporated herein by reference in its entirety for all purposes.
Citation Information
Patent Citations
towel ring
CN3315821D
cell phone
CN3329834D
Recombinant antibodies and methods for their production
EP0239400A2
A method for reducing the immunogenicity of antibody variable domains
EP0519596A1
Resurfacing of rodent antibodies
EP0592106A1