Optimized protein fusions and linkers
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INTEGRATED DNA TECHNOLOGIES INC
- Filing Date
- 2021-04-20
- Publication Date
- 2026-05-29
Smart Images

Figure CN115968301B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the priority benefit of U.S. Provisional Patent Application Serial No. 63 / 012658, entitled “Optimized Protein Fusion and Connector,” filed April 20, 2020, the contents of which are incorporated herein by reference in their entirety. Technical Field
[0003] This invention relates to chimeric proteins and methods of using them in living cells. Technical Field
[0005] This invention relates to optimized protein fusion linkers for generating multifunctional chimeric proteins and methods of using them. Furthermore, this invention relates to chimeric proteins for guided endonuclease systems. Background Technology
[0006] For many reasons, the fusion of guided endonucleases with one or more unrelated enzymes is desirable. In some cases, it is useful to observe and monitor the successful delivery of the guided endonuclease reagent to the target cell nucleus or nucleolus. In other cases, it may be useful to fuse DNA-modifying enzymes after guided endonuclease cleavage to influence specific repair outcomes.
[0007] The construction of recombinant fusion proteins requires the selection of appropriate linkers to connect protein domains. Direct fusion of functional domains without linkers can lead to a number of problems, including misfolding of the fusion protein, low protein yield, or impaired biological activity. For example, one obstacle to using engineered nucleases fused with fluorescent proteins is that such fusions often exhibit lower editing activity, possibly due to steric turbulence caused by the non-natural fusion of the two proteins.
[0008] Achieving high functional activity in fusion proteins may depend on the nature of the linkers between subunits, as direct protein fusion can inhibit folding, stability, and biological activity. Furthermore, the connected domains can interfere with each other's function; physically separating these domains can improve function. The nature of the linkers between domains can affect the average separation distance, depending on the rigidity and length of the linker.
[0009] Linkers are generally classified into three categories: flexible, rigid, and cleavable. Flexible linkers typically consist of glycine repeat sequences with periodically added polar serine residues to disrupt the interactions between linker proteins. Rigid linkers include A(EAAAK) linkers that form an α-helix structure. n A repeating sequence or a Pro-rich sequence (XP) forming a relatively rigid, extended rod-like structure. n .
[0010] Therefore, there is a need for methods and compositions that overcome the challenges of existing technologies. A novel set of universal peptide linkers, specifically optimized for guided endonuclease fusions, is desirable for improving the function of guided endonucleases and covalently fused protein chaperones. Summary of the Invention
[0011] In general, the present invention relates to methods and compositions for improving multifunctional chimeric proteins and improved linkers. In some embodiments, the chimeric protein comprises a guided endonuclease protein, such as an RNA-guided endonuclease (“RGEN”), including Cas / CRISPR proteins, which are covalently fused to a chaperone protein via a set of universal peptide linkers. In other embodiments, the fusion protein comprises a rigid linker system for fusing two proteins. In a further embodiment, the rigid linker can be used for a fusion protein comprising a Cas protein fused to a fluorescent protein.
[0012] For many reasons, the fusion of a guided endonuclease Cas protein with one or more unrelated protein chaperones is desirable. For example, one implementation includes the ability to generate a CRISPR / Cas9 protein fusion to cleave double-stranded DNA at a precise location within a living cell. Cas9 is an RNA-guided endonuclease derived from the bacterial adaptive immune system of *Steptococcus pyogenes*, characterized by clustered, regularly spaced short palindromic repeats ((CRISPR)-Cas (CRISPR-associated)). Cas9 is guided to a 23-nt DNA target sequence by a target-site-specific 20-nt complementary RNA (part of a 44-nt crRNA) and a universal 89-nt tracrRNA (collectively referred to as the guide RNA (gRNA) complex). The Cas9-gRNA ribonucleoprotein (RNP) complex mediates double-stranded DNA breaks (DSBs), which are then repaired via non-homologous end joining (NHEJ), typically introducing mutations or deletions at the cleavage site if a suitable template nucleic acid is present. This often leads to gene disruption via frameshift mutations or homology-directed repair (HDR) systems.
[0013] In one embodiment, a set of universal peptide linkers specifically optimized for CRISPR enzyme fusion is desirable for improving the function of Cas9 and the covalently fused chaperone protein or protein domain. The universal linkers are not specific to the Cas9 protein or any mutant variant thereof. In another embodiment, the CRISPR enzyme can be optimally linked to a fluorescent protein using universal linkers.
[0014] In one implementation, a rigid linker is used for the covalent fusion of two different proteins or protein domains. In one aspect, a rigid linker is used, which includes forming an α-helix structure to provide rigidity for A(EAAAK).n A repeating sequence. On the other hand, (XP) is used to form a relatively rigid, extended rod-like structure. n Repeating sequence.
[0015] In the first aspect, chimeric proteins are provided. Chimeric proteins include guided endonuclease proteins, rigid linkers, and secondary proteins.
[0016] In a second aspect, a method for enriching cells with chimeric proteins is provided. This method comprises several steps. The first step includes incubating the chimeric protein according to the first aspect with guide RNA to form an RNP complex. The second step includes contacting the RNP complex with multiple target cells to generate recipient cells containing the RNP complex. The third step includes sorting the recipient cells based on fluorescence signals.
[0017] In the third aspect, isolated nucleic acids encoding chimeric proteins are provided. Chimeric proteins include those described in the first aspect.
[0018] In the fourth aspect, chimeric proteins are provided. These chimeric proteins include guided endonuclease proteins, universal linkers, and secondary proteins.
[0019] In the fifth aspect, isolated nucleic acids are provided. These isolated nucleic acids encode chimeric proteins according to any aspect of the fourth aspect.
[0020] In the sixth aspect, a method for enriching cells with chimeric proteins is provided. This method comprises several steps. The first step includes incubating the chimeric protein according to any aspect of the fourth aspect with guide RNA to form an RNP complex. The second step includes contacting the RNP complex with multiple target cells to generate recipient cells containing the RNP complex. The third step includes sorting the recipient cells based on fluorescence signals. Attached Figure Description
[0021] Figure 1AThe percentage of cleavage endonuclease activities of the Cas9-eGFP adaptor variants (L1, Cas9-GSAGSAAGSGEF-eGFP (SEQ ID NO:38); L2, Cas9-(AP)7A-eGFP (SEQ ID NO:49); L3, Cas9-A(EAA AK)4A-eGFP (SEQ ID NO:40); L4, Cas9-GGGGSEAAAKGGGG-eGFP (SEQ ID NO:71) and L5, Cas9-(GGGGS)4A-eGFP (SEQ ID NO:39)) was depicted compared to the unlabeled Cas9 protein (denoted as "C" (SEQ ID NO:324)) delivered as a plasmid (400 ng), where sgRNA was delivered via reverse transfection targeting HPRT-38087 (SEQ ID NO:324) 24 hours after plasmid delivery. Editing activity at 48 hours was measured by the T7E1 assay.
[0022] Figure 1B The percentage of cleavage endonuclease activities of the Cas9-eGFP adaptor variants (L1, Cas9-GSAGSAAGSGEF-eGFP (SEQ ID NO:38); L2, Cas9-(AP)7A-eGFP (SEQ ID NO:49); L3, Cas9-A(EAAAK)4A-eGFP (SEQ ID NO:40); L4, Cas9-GGGGSEAAAKGGGG-eGFP (SEQ ID NO:71) and L5, Cas9-(GGGGS)4A-eGFP (SEQ ID NO:39)) was depicted compared to the unlabeled Cas9 protein (denoted as "C" (SEQ ID NO:326)) delivered as a plasmid (400 ng), where sgRNA was delivered via reverse transfection targeting HPRT-38285 (SEQ ID NO:326) 24 hours after plasmid delivery. Editing activity at 48 hours was measured by the T7E1 assay.
[0023] Figure 2The percentage editing efficiency of plasmid-expressed Cas9-eGFP adaptor variants (Cas9-GSAGSAAGSGEF-eGFP (SEQ ID NO:38); Cas9-(AP)7A-eGFP (SEQ ID NO:49); and Cas9-A(EAAAK)4A-eGFP (SEQ ID NO:40)) compared to the unlabeled Cas9 protein (SEQ ID NO:1), where the gRNAs were delivered to HEK293 cells via reverse transfection 24 hours post-plasmid delivery. Data are presented as box and whisker plots, showing the aggregation and cleavage efficiency of 12 gRNAs targeting a unique locus with HPRT (SEQ ID NOS:319-330). Cleavage efficiency was measured using the T7EI assay 48 hours post-transfection.
[0024] Figure 3A The editing efficiency of plasmid-expressed Cas9-eGFP adaptor variants (L1, Cas9-GSAGSAAGSGEF-eGFP (SEQ ID NO:38); L2, Cas9-(AP)7A-eGFP (SEQ ID NO:49); L3, Cas9-A(EAAAK)4A-eGFP (SEQ ID NO:40); L4, Cas9-A(EAAAK)4ALEA(EAAAK)4A-eGFP (SEQ ID NO:41); and L5, Cas9 LEA(EAAAK)4ALEA(EAAAK)4ALE-eGFP (SEQ ID NO:42)) was depicted compared with unlabeled Cas9 protein (denoted as "C", SEQ ID NO:1) co-delivered to HEK293 cells via electroporation with sgRNA expression plasmids, wherein the expressed sgRNA targets HPRT-38087 (SEQ ID NO:324). After transfecting the two plasmids into HEK293 cells for 72 hours, T7E1 cleavage assays were performed on each transfected cell population.
[0025] Figure 3BThe editing efficiency of plasmid-expressed Cas9-eGFP adaptor variants (L1, Cas9-GSAGSAAGSGEF-eGFP (SEQ ID NO:38); L2, Cas9-(AP)7A-eGFP (SEQ ID NO:49); L3, Cas9-A(EAAAK)4A-eGFP (SEQ ID NO:40); L4, Cas9-A(EAAAK)4ALEA(EAAAK)4A-eGFP (SEQ ID NO:41); and L5, Cas9-LEA(EAAAK)4ALEA(EAAAK)4ALE-eGFP (SEQ ID NO:42)) was depicted compared with unlabeled Cas9 protein (denoted as "C", SEQ ID NO:1) co-delivered to HEK293 cells via electroporation with sgRNA expression plasmids, wherein the expressed sgRNA targets HPRT-38285 (SEQ ID NO:326). After transfecting the two plasmids into HEK293 cells for 72 hours, T7E1 cleavage assays were performed.
[0026] Figure 4A The study depicts the expression of Cas9-eGFP adaptor variants (1, Cas9-GSAGSAAGSGEF-eGFP (SEQ ID NO:38); 2, Cas9-APA-eGFP (SEQ ID NO:43); 3, Cas9-(AP)2A-eGFP (SEQ ID NO:44); 4, Cas9-(AP)3A-eGFP (SEQ ID NO:45); 5, Cas9-(AP)4A-eGFP (SEQ ID NO:46); 6, Cas9-(AP)5A-eGFP (SEQ ID NO:47); 7, Cas9-(AP)6A-eGFP (SEQ ID NO:48); 8, Cas9-(AP)7A-eGFP (SEQ ID NO:49); 9, Cas9-(AP)8A-eGFP (SEQ ID NO:48)) expressed by plasmids compared to unlabeled Cas9 protein (denoted as "C", SEQ ID NO:1) co-delivered to HEK293 cells via electroporation with sgRNA expression plasmids. ID NO:50); 10, Cas9-(AP)9A-eGFP (SEQ ID NO:51) 11, Cas9-(AP) 10 A-eGFP (SEQ ID NO: 52); 12, Cas9-(AP) 11 A-eGFP (SEQ ID NO: 53); 13, Cas9-(AP) 12 A-eGFP (SEQ ID NO: 54); 14, Cas9-(AP) 13A-eGFP (SEQ IDNO:55); and 15, Cas9-(AP) 14 The editing efficiency of A-eGFP (SEQ ID NO:56) was measured, in which the expressed sgRNA targets HPRT-38087 (SEQ ID NO:324). T7E1 cleavage was measured 72 hours after transfecting both plasmids into HEK293 cells.
[0027] Figure 4B The image depicts the expression of Cas9-eGFP adaptor variants (1, Cas9-GSAGSAAGSGEF-eGFP (SEQ ID NO:38); 2, Cas9-APA-eGFP (SEQ ID NO:43); 3, Cas9-(AP)2A-eGFP (SEQ ID NO:44); 4, Cas9-(AP)3A-eGFP (SEQ ID NO:45); 5, Cas9-(AP)4A-eGFP (SEQ ID NO:46); 6, Cas9-(AP)5A-eGFP (SEQ ID NO:47); 7, Cas9-(AP)6A-eGFP (SEQ ID NO:48); 8, Cas9-(AP)7A-eGFP (SEQ ID NO:49); 9, Cas9-(AP)8A-eGFP (SEQ ID NO:48)) expressed by plasmids compared to unlabeled Cas9 protein (denoted as "C", SEQ ID NO:1) co-delivered to HEK293 cells via electroporation with sgRNA expression plasmids. ID NO:50); 10, Cas9-(AP)9A-eGFP (SEQ ID NO:51); 11, Cas9-(AP) 10 A-eGFP (SEQ ID NO: 52); 12, Cas9-(AP) 11 A-eGFP (SEQ ID NO: 53); 13, Cas9-(AP) 12 A-eGFP (SEQ ID NO: 54); 14, Cas9-(AP) 13 A-eGFP (SEQ IDNO:55); and 15, Cas9-(AP) 14 The editing efficiency of A-eGFP (SEQ ID NO:56) was measured, in which the expressed sgRNA targets HPRT-38285 (SEQ ID NO:326). T7E1 cleavage was performed 72 hours after transfecting both plasmids into HEK293 cells.
[0028] Figure 5AEditing efficiency of Cas9-GFP adaptor variants (1, Cas9-GSAGSAAGSGEF-eGFP (SEQ ID NO:38); 2, Cas9-(AP)7A-eGFP (SEQ ID NO:49); 3, Cas9-(AP)9A-eGFP (SEQ ID NO:51); 4, Cas9-GSAGSAAGSGEF-mCherry (SEQ ID NO:72); 5, Cas9-(AP)7A-mCherry (SEQ ID NO:83); and 6, Cas9-(AP)9A-mCherry (SEQ ID NO:85)) was depicted compared to unlabeled Cas9 protein (denoted as "C", SEQ ID NO:1) delivered to HEK293 cells at a concentration of 0.0625 μmol. Data are presented as box and whisker plots, showing the aggregation and cleavage efficiency of 12 gRNAs targeting a unique locus with HPRT (SEQ ID NOS:319-330). Cutting efficiency was measured using the T7EI assay 48 hours after transfection.
[0029] Figure 5B Editing efficiency of Cas9-GFP adapter variants (1, Cas9-GSAGSAAGSGEF-eGFP (SEQ ID NO:38); 2, Cas9-(AP)7A-eGFP (SEQ ID NO:49); 3, Cas9-(AP)9A-eGFP (SEQ ID NO:51); 4, Cas9-GSAGSAAGSGEF-mCherry (SEQ ID NO:72); 5, Cas9-(AP)7A-mCherry (SEQ ID NO:83); and 6, Cas9-(AP)9A-mCherry (SEQ ID NO:85)) was depicted compared to unlabeled Cas9 protein (denoted as "C", SEQ ID NO:1) delivered to HEK293 cells at a concentration of 0.25 μmol. Data are presented as box and whisker plots, showing the aggregation and cleavage efficiency of 12 gRNAs targeting a unique locus with HPRT (SEQ ID NOS:319-330). Cutting efficiency was measured using the T7EI assay 48 hours after transfection.
[0030] Figure 5CEditing efficiency of Cas9-GFP adapter variants (1, Cas9-GSAGSAAGSGEF-eGFP (SEQ ID NO:38); 2, Cas9-(AP)7A-eGFP (SEQ ID NO:49); 3, Cas9-(AP)9A-eGFP (SEQ ID NO:51); 4, Cas9-GSAGSAAGSGEF-mCherry (SEQ ID NO:72); 5, Cas9-(AP)7A-mCherry (SEQ ID NO:83); and 6, Cas9-(AP)9A-mCherry (SEQ ID NO:85)) was depicted compared to unlabeled Cas9 protein (denoted as "C", SEQ ID NO:1) delivered to HEK293 cells at a concentration of 2.0 μmol. Data are presented as box and whisker plots, showing the aggregation and cleavage efficiency of 12 gRNAs targeting a unique locus with HPRT (SEQ ID NOS:319-330). Cutting efficiency was measured using the T7EI assay 48 hours after transfection.
[0031] Figure 6 The effect of FACS selection on Cas9-eGFP fusion protein editing in HEK293 cells compared to unsorted cells was described. An RNP containing a chimeric Cas9-(AP)7A-eGFP protein (SEQ ID NO:49) was introduced into HEK293 cells at a final concentration of 10 nM via lipid transfection. This chimeric Cas9-(AP)7A-eGFP protein formed a double strand with guide RNA. Cells were divided into three populations based on GFP signals corresponding to high, medium, and low GFP signals (high GFP signal, top 20% of the signal; medium GFP signal, middle 80-60% of the signal; and low GFP signal, bottom 60% of the signal). Next-generation sequencing (NGS) of the target sites was performed on each population to measure the percentage of total editing activity 48-72 hours after transfection.
[0032] Figure 7The LbCas12a(E795L)-GFP adaptor variants (1, LbCas12a(E795L)-eGFP (SEQ ID NO:314); 2, LbCas12a(E795L)-A(EAAAK)4A-eGFP (SEQ ID NO:244); 3, LbCas12a(E795L)-(GGGS)4-eGFP (SEQ ID NO:243); 4, LbCas12a(E795L)-(AP)7A-eGFP (SEQ ID NO:253); 5, LbCas12a(E795L)-(AP)9A-eGFP (SEQ ID NO:314)) compared to unlabeled LbCas12a(E795L) protein (denoted as "C", SEQ ID NO:310) delivered to HEK293 cells at a concentration of 50 nM are depicted. Exemplary results of editing efficiency for 6, LbCas12a(E795L)-A(EAAAK)4ALEA(EAAAK)4A-eGFP(SEQ ID NO:245), wherein the selection targets a unique locus with HPRT. LbCas 12a crRNAs. The results shown are 38115:Cpf1HPRT 38115-AS (SEQ ID NO:333) and 38186:Cpf1 HPRT 38186-S (SEQ ID NO:334); the other two guide RNAs (CPf1 HPRT 38330-AS (SEQ ID NO:335) and Cpf1 HPRT 38486-S (SEQ ID NO:336)) yielded comparable results (data not shown). Detailed Implementation
[0033] The methods and compositions of the present invention described herein provide improved universal linkers for the covalent fusion of two or more proteins or protein domains. Furthermore, the invention describes methods using chimeric fusion. Specifically, the disclosed chimeric proteins provide a robust manner in which cells containing chimeric proteins can be identified and sorted based on their ability to generate fluorescent signals. Moreover, compared to unmodified endonucleases or fusion proteins lacking linkers, not only do chimeric proteins retain their enzymatic activity as guided endonucleases, but some chimeric proteins exhibit remarkably enhanced activity. These and other advantages of the invention, as well as additional inventive features, will become apparent from the description of the invention provided herein.
[0034] definition
[0035] In the context of describing the invention (especially in the context of the following claims), the terms “a” and “an” and “the” should be understood to cover both singular and plural forms, unless otherwise stated herein or the context clearly contradicts this. Unless otherwise stated, the terms “comprising,” “having,” “including,” and “containing” should be interpreted as open-ended terms (i.e., meaning “including, but not limited to”). Unless otherwise stated herein, the descriptions of numerical ranges herein are intended only as a shorthand method for individually referring to each individual value falling within that range, and each individual value is incorporated into the specification as if it were individually described herein. Unless otherwise stated herein or clearly contradicted by the context, all methods described herein can be performed in any suitable order. Unless otherwise stated, the use of any and all examples or exemplary language (e.g., “such as”) provided herein is merely for the purpose of better illustrating the invention and not for limiting the scope of the invention. No language in the specification should be construed as indicating that any unclaimed element is essential to the practice of the invention.
[0036] The term "codon-optimized," as a term used to describe a "nucleic acid," "gene," "DNA," or "RNA" that modifies a protein or polypeptide, refers to a nucleic acid, gene, DNA, or RNA that includes preferred codons for the efficient expression of a protein or polypeptide in a given host cell or organism, based on the abundance of naturally occurring codon-specific charged tRNAs in that host cell or organism. For example, when a nucleic acid is expressed in *E. coli*, a codon-optimized Cas9 nucleic acid for *E. coli* will be suitable for optimal expression of the Cas9 protein. Similarly, when a nucleic acid is expressed in human cells, a codon-optimized Cas9 nucleic acid for human cells will be suitable for optimal expression of the Cas9 protein. Exemplary codon-optimized nucleic acids are disclosed herein. Codon-optimized nucleic acids, genes, DNA, and RNA can be readily generated from naturally occurring endogenous DNA or RNA derived from the original host cell or organism, or from the reverse transcription of the amino acid sequence of a protein from the relevant protein or polypeptide sequence. Those skilled in the art will understand various codon preferences or biases for host cells or organisms from the literature. Furthermore, codon-optimized nucleic acid conversion software programs are readily available online or are generally known in the field.
[0037] The terms "RNA-guided endonuclease" and "RGEN" refer to ribonucleoprotein endonucleases, which include RNA components that have targeting enzyme activity in a given substrate. Exemplary RGENs include CRISPR-associated endonucleases.
[0038] The term "CRISPR" refers to the bacterial adaptive immune system consisting of clusters of regularly spaced short palindromic repeat sequences.
[0039] The terms “Cas” and “Cas endonuclease” usually refer to CRISPR-related endonucleases.
[0040] The term "Cas protein" generally refers to the wild-type protein of a CRISPR-associated endonuclease (including the interchangeable terms Cas and Cas endonuclease), including its variants.
[0041] The term "Cas nucleic acid" usually refers to the nucleic acid of a CRISPR-associated endonuclease, including guide RNA, sgRNA, crRNA, or tracrRNA.
[0042] The terms “Cas9” and “CRISPR / Cas9” refer to the CRISPR-associated bacterial adaptive immune system of Steptococcus pyogenes.
[0043] The terms "AsCas12a" and "CRISPR / AsCas12a" refer to the CRISPR-associated bacterial adaptive immune system of the genus *Acidaminococcus* sp.
[0044] The terms “LbCas12a” and “CRISPR / LbCas12a” refer to the CRISPR-associated bacterial adaptive immune system of Lachnospiraceae bacterium.
[0045] The term "Cas9 protein" refers to a protein of the Cas9 or CRISPR / Cas9 endonuclease system. For the purposes of this disclosure, the amino acid sequence of the wild-type Cas9 protein is SEQ ID NO:1.
[0046] The term "AsCas12a protein" refers to a protein of the AsCas12a or CRISPR / AsCas12a endonuclease system. For the purposes of this disclosure, the amino acid sequence of the wild-type AsCas12a protein is SEQ ID NO:2.
[0047] The term "LbCas12a protein" refers to the protein of LbCas12a or the CRISPR / LbCas12a endonuclease system. For the purposes of this disclosure, the amino acid sequence of the wild-type LbCas12a protein is SEQ ID NO:3.
[0048] The term "Cas9 nucleic acid" refers to a nucleic acid (e.g., DNA or RNA) encoding a Cas9 protein or polypeptide. Cas9 nucleic acids can be selected from naturally occurring nucleic acids derived from Streptococcus pyogenes or from codon-optimized nucleic acids for efficient expression in a given host cell or organism. For the purposes of this disclosure, exemplary codon-optimized versions of Cas9 nucleic acids are SEQ ID NO:337 and SEQ ID NO:338.
[0049] The term "AsCas12a nucleic acid" refers to a nucleic acid (e.g., DNA or RNA) encoding an AsCas12a protein or polypeptide. AsCas12a nucleic acid may be selected from native nucleic acids of the genus *Acidaminococcus* or from codon-optimized nucleic acids for efficient expression in a given host cell or organism. For the purposes of this disclosure, exemplary codon-optimized versions of AsCas12a nucleic acid are SEQ ID NO:339 and SEQ ID NO:340.
[0050] The term "LbCas12a nucleic acid" refers to a nucleic acid (e.g., DNA or RNA) encoding an LbCas12a protein or polypeptide. LbCas12a nucleic acid may be selected from naturally occurring nucleic acids from *Lachnospiraceae bacterium* or from codon-optimized nucleic acids for efficient expression in a given host cell or organism. For the purposes of this disclosure, exemplary codon-optimized versions of LbCas12a nucleic acid are SEQ ID NO:341 and SEQ ID NO:342.
[0051] The terms “guide RNA,” “guide RNA complex,” and “gRNA complex” refer to target-specific crRNA, universal tracrRNA, or a combination of both.
[0052] The term "sgRNA" refers to the guide RNA complex, in which crRNA is covalently linked to tracrRNA in a single molecule.
[0053] The term "variant," as used for modified proteins (e.g., Cas9, AsCas12a, or LbCas12a proteins), refers to a protein containing at least one amino group substituted with a reference wild-type protein amino acid sequence, an additional amino acid (e.g., an affinity tag or nuclear localization signal), or a combination thereof. An exemplary LbCas12a protein variant amino acid sequence is LbCas12a(E795L) (SEQ ID NO:310), and chimeric proteins containing this amino acid sequence are given in Tables V.7, V.8, and V.10.
[0054] the term The term "modified RNA" (e.g., crRNA, tracrRNA, guide RNA, or sgRNA) refers to isolated, chemically synthesized synthetic RNA.
[0055] the term The term "modified protein" (e.g., Cas9 protein, AsCas12a protein, or LbCas12a protein) refers to isolated recombinant proteins.
[0056] The terms "fusion protein," "protein fusion," "chimeric protein," "protein chimera," and "chimeric fusion protein" refer to a first protein or polypeptide covalently bonded to at least one or more proteins or polypeptides, wherein the at least one or more proteins or polypeptides differ in their primary sequence composition from the first protein or polypeptide. The terms fusion protein, protein fusion, chimeric protein, protein chimera, and chimeric fusion protein have the same meaning and are used interchangeably. Exemplary fusion proteins are disclosed herein.
[0057] The term "general" as used to modify "guide RNA," "connector," or "connector peptide" refers to a guide RNA, connector, or connector peptide that functions as a guide RNA, connector, or connector peptide among multiple RGENs (associated with guide RNA) or multiple fusion proteins (associated with connectors or connector peptides). Exemplary general connectors and general connector peptides include flexible connectors, rigid connectors, and hybrid connectors, including those identified in Table I and the references cited therein, the entire contents of which are incorporated herein by reference.
[0058] The term "editing activity assay" refers to a assay that determines the degree of editing by RGEN (e.g., a Cas endonuclease) at at least one locus targeting guide RNA. Exemplary editing activity assays include those selected from T7EI assays and next-generation sequencing (NGS) assays. An exemplary T7EI assay procedure is disclosed in U.S. Patent Application Serial No. 14 / 975,709 (Attorney General's File No. IDT01-008-US), filed December 18, 2015, the contents of which are incorporated herein by reference. An exemplary NGS assay is disclosed in U.S. Patent Application Serial No. 13 / 935,451 (Attorney General's File No. IDT01-001-US), filed July 3, 2013, the contents of which are incorporated herein by reference.
[0059] The terms "linker" and "linker peptide" refer to a polypeptide amino acid sequence that covalently links two or more amino acids of a first protein or polypeptide to a second protein or polypeptide. Exemplary linkers and linker peptides are disclosed herein.
[0060] The term "rigid connector" refers to a connector that restricts at least one degree of freedom of motion or conformation of the affected fusion protein, protein fusion, chimeric protein, or protein chimera, including the connector. Exemplary rigid connectors are disclosed herein.
[0061] In a first embodiment, a composition for an improved multifunctional chimeric protein and an improved adapter is provided. On the other hand, a universal peptide adapter is a rigid adapter system.
[0062] In another embodiment, a rigid linker is used to fuse two proteins, recombinant proteins, or protein domains into a single covalently fused protein construct. In one aspect, a (EAAAK) connector is used. n The α-helix of the sequence forms a linker. On the other hand, the EAAAK amino acid sequence is repeated n times, where n represents 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 repetitions.
[0063] In another embodiment, the rigid joint is a Pro-rich sequence. On the other hand, the Pro-rich sequence is (XP). n In this sequence, X represents any amino acid, preferably including alanine, lysine, or glutamic acid, more preferably alanine. In another aspect, the XP sequence is repeated n times, where n represents 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 repetitions. In another aspect, the number of repetitions is preferably 5, 6, 7, 8, or 9. In yet another aspect, the number of repetitions is more preferably 7 or 9.
[0064] In another embodiment, the rigid joint is a Pro-rich sequence. On the other hand, the Pro-rich sequence is (XP). n A, where X represents any amino acid, preferably including alanine, lysine, or glutamic acid, more preferably alanine, wherein the XP sequence is repeated n times in the range of 1-14 times, including n being 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 times. Alternatively, the number of repetitions is preferably 5, 6, 7, 8, or 9 times. Even more preferably, the number of repetitions is 7 or 9 times.
[0065] In another embodiment, a rigid adapter is used for covalently fusing a guided endonuclease and a second protein. In one aspect, the rigid adapter is used to covalently fuse the C-terminus of the guided endonuclease to the N-terminus of the second protein. Preferably, the rigid adapter is used to covalently fuse the C-terminus of RGEN to the N-terminus of the second protein. More preferably, the rigid adapter is used to covalently fuse the C-terminus of CRISPR-Cas to the N-terminus of the second protein.
[0066] In a further implementation, Pro-enriched (XP)n A rigid linker, wherein n is a repeating unit in the range of 1-14 repeats, including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 repeats, is used to covalently fuse the C-terminus of a guided endonuclease to the N-terminus of a second protein. Preferably, (XP) n A rigid linker is used to covalently fuse the C-terminus of RGEN to the N-terminus of a second protein. More preferably, (XP) n Rigid connectors are used to covalently fuse the C-terminus of a Cas protein to the N-terminus of a second protein.
[0067] In another embodiment, the Pro-rich sequence includes (PX). n P-rigid joint system, where n is a repeating unit in the range of 1-14 repetitions, including n being 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 repetitions. In one aspect, Pro-rich sequences include (PA). n P-rigid joint system, where n is a repeating unit in the range of 1-14 repetitions, including n being 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 repetitions. On the other hand, Pro-rich sequences include (PA)6P rigid joint system.
[0068] In another embodiment, a rigid adapter is used to covalently fuse the C-terminus of the first protein to the N-terminus of the guided endonuclease. In another embodiment, a rigid adapter is used to covalently fuse the C-terminus of the first protein to the N-terminus of RGEN. In yet another embodiment, a rigid adapter is used to covalently fuse the C-terminus of the first protein to the N-terminus of the Cas protein.
[0069] In another implementation, Pro-rich (XP) n Rigid linkers are used to covalently fuse the C-terminus of the first protein to the N-terminus of a guided endonuclease. On the other hand, (XP) n The linker is used to covalently fuse the C-terminus of the first protein to the N-terminus of the RGEN. On the other hand, (XP) n Rigid connectors are used to covalently fuse the C-terminus of the first protein to the N-terminus of the Cas protein.
[0070] In another embodiment, the CRISPR / Cas9 protein is one of the proteins covalently fused to a rigid linker. In one aspect, the Cas9 protein can be a wild-type Cas9 protein. On the other hand, the Cas9 protein can be a mutant or variant protein.
[0071] In another embodiment, the first protein is a guided endonuclease linked to a fluorescent protein via a rigid linker. Preferably, the guided endonuclease is RGEN, more preferably a CRISPR / Cas enzyme. In one aspect, the fluorescent protein is eGFP or mCherry. In another aspect, the CRISPR / Cas enzyme is covalently linked to the fluorescent protein via a rigid linker to produce a CRISPR / Cas fluorescent chimeric fusion protein. The CRISPR / Cas fluorescent fusion protein has the ability to visualize and monitor the successful delivery of CRISPR / Cas reagents to the target cell nucleus or nucleolus. In another aspect, cells can be sorted and enriched based on the detection of the fluorescent fusion protein.
[0072] In a first aspect, a chimeric protein is provided. The chimeric protein comprises a guided endonuclease protein, a rigid linker, and a second protein. In the first aspect, the guided endonuclease protein is a Cas protein. In the second aspect, the Cas protein is selected from the group consisting of Cas9 protein, AsCas12a protein, LbCas12a protein, Cas9 protein variants, AsCas12a protein variants, or LbCas12a protein variants. In the third aspect, the Cas protein is selected from the group comprising SEQ ID NO:1, SEQ ID NO:2, and SEQ ID NO:3. In the fourth aspect, the Cas protein is preferably SEQ ID NO:1. In the fifth aspect, the Cas protein is preferably SEQ ID NO:2. In the sixth aspect, the Cas protein is preferably SEQ ID NO:3. In the seventh aspect, the second protein is a fluorescent protein. In the eighth aspect, the fluorescent protein is selected from the group comprising SEQ ID NO:4 and SEQ ID NO:5. In the ninth aspect, the fluorescent protein is SEQ ID NO:4. In the tenth aspect, the fluorescent protein is SEQ ID NO:5. In the eleventh aspect, the rigid joint is selected from the group comprising XP, APA, SEQ ID NO:8, SEQ ID NO:11-23, and SEQ ID NO:24-36. In the twelfth aspect, the rigid joint is SEQ ID NO:8. In the thirteenth aspect, the rigid joint is selected from the group comprising XP and SEQ ID NO:24-36. In the fourteenth aspect, the rigid joint is selected from SEQ ID NO:29-31. In the fifteenth aspect, the rigid joint is selected from the group comprising APA and SEQ ID NO:11-23. In the sixteenth aspect, the rigid joint is selected from the group comprising APA and SEQ ID NO:17-18. In the eighteenth aspect, the chimeric protein is selected from the group comprising SEQ ID NO:40-70, SEQ ID NO:74-104, SEQ ID NO:108-138, SEQ ID NO:142-172, SEQ ID NO:176-206, SEQ ID NO:210-240, SEQ ID NO:244-274 and SEQ ID NO:278-308.
[0073] In a second aspect, a method for enriching cells containing chimeric proteins is provided. The method comprises several steps. The first step comprises incubating the chimeric protein according to the first aspect with guide RNA to form an RNP complex. The second step comprises contacting the RNP complex with a plurality of target cells to generate recipient cells containing the RNP complex. The third step comprises sorting the recipient cells based on a fluorescence signal. In the first aspect, the method includes an additional step of assaying editing activity at at least one locus targeted by the guide RNA. In the second aspect, the step of assaying editing activity at at least one locus targeted by the guide RNA is selected from T7EI assays and next-generation sequencing assays. In the third aspect, the method comprises chimeric proteins selected from the group consisting of SEQ ID NO:40-70, SEQ ID NO:74-104, SEQ ID NO:108-138, SEQ ID NO:142-172, SEQ ID NO:176-206, SEQ ID NO:210-240, SEQ ID NO:244-274, and SEQ ID NO:278-308.
[0074] In a third aspect, an isolated nucleic acid encoding a chimeric protein is provided. The chimeric protein includes the chimeric protein of the first aspect. In the first aspect, the chimeric protein comprises a member selected from the group consisting of SEQ ID NO:40-70, SEQ ID NO:74-104, SEQ ID NO:108-138, SEQ ID NO:142-172, SEQ ID NO:176-206, SEQ ID NO:210-240, SEQ ID NO:244-274, and 278-308. In a second aspect, the isolated nucleic acid is codon-optimized for expression in an organism or host cell. In a third aspect, the organism or host cell is selected from *Escherichia coli* or *Homo sapiens*.
[0075] In a fourth aspect, a chimeric protein is provided. The chimeric protein comprises a guided endonuclease protein, a universal adapter, and a second protein. In a first aspect, the guided endonuclease protein is a Cas protein. In a second aspect, the Cas protein is selected from the group consisting of Cas9 protein, AsCas12a protein, LbCas12a protein, Cas9 protein variants, AsCas12a protein variants, and LbCas12a protein variants. In a third aspect, the Cas protein is selected from the group comprising SEQ ID NO:1, SEQ ID NO:2, and SEQ ID NO:3. In a fourth aspect, the Cas protein is SEQ ID NO:1. In a fifth aspect, the Cas protein is SEQ ID NO:2. In a sixth aspect, the Cas protein is SEQ ID NO:3. In a seventh aspect, the second protein is a fluorescent protein. In an eighth aspect, the fluorescent protein is selected from the group comprising SEQ ID NO:4 and SEQ ID NO:5. In a ninth aspect, the fluorescent protein is SEQ ID NO:4. In a tenth aspect, the fluorescent protein is SEQ ID NO:5. In an eleventh aspect according to any of the preceding aspects of the fourth aspect, the universal adapter is selected from flexible adapters, rigid adapters, and hybrid adapters. In a twelfth aspect, the universal adapter comprises a flexible adapter selected from the group consisting of SEQ ID NO:6 and SEQ ID NO:7. In a thirteenth aspect, the universal adapter comprises a rigid adapter selected from the group consisting of XP, APA, and SEQ ID NO:8-36. In a fourteenth aspect, the universal adapter comprises a hybrid adapter of SEQ ID NO:37. In a fifteenth aspect, the chimeric protein is selected from the group consisting of SEQ ID NO:38-309.
[0076] In the fifth aspect, an isolated nucleic acid is provided. The isolated nucleic acid encodes a chimeric protein according to any of the fourth aspects. In the first aspect, the isolated nucleic acid is codon-optimized for expression in an organism or host cell. In the second aspect, the organism or host cell is selected from *Escherichia coli* or *Homo sapiens*. In the third aspect, the isolated nucleic acid is selected from the group consisting of SEQ ID NO:348-387 and SEQ ID NO:389-429.
[0077] In a sixth aspect, a method for enriching cells containing chimeric proteins is provided. This method comprises several steps. The first step comprises incubating the chimeric protein according to any aspect of the fourth aspect with guide RNA to form an RNP complex. The second step comprises contacting the RNP complex with multiple target cells to generate recipient cells containing the RNP complex. The third step comprises sorting the recipient cells based on a fluorescence signal. In a first aspect, the method includes an additional step of assessing editing activity at at least one locus targeted by the guide RNA. In a second aspect, the step of assessing editing activity at at least one locus targeted by the guide RNA is selected from T7EI assays and next-generation sequencing assays.
[0078] Example 1
[0079] The impact of connectors on the efficiency of Cas9 recombinant editing
[0080] This example identified a linker that improves Cas9-eGFP activity. Nineteen independent plasmids expressing recombinant Cas9 in mammalian cells were constructed, in which the encoded protein has eGFP downstream of the Cas9 carboxyl-terminal domain (CTD), and the peptide sequences between Cas9 and eGFP are distinct (Table I).
[0081] Table I. Connector Sequence List.
[0082]
[0083]
[0084]
[0085] 1 Waldo 1999; 2 Gergeron 2009; 3 Bae and Shen 2006; 4 McCormick 2001; 5 Chen 2017
[0086] The construct was first analyzed by expression in plasmids delivered to HEK293 cells, followed by T7EI digestion assays after 48–72 hours. The construct showing enhanced activity was then subcloned into a vector for protein expression in *E. coli*, and the recombinant Cas9-fluorescent protein fusion was purified by immobilized metal affinity chromatography followed by ion exchange chromatography. The Cas9 endonuclease activity of the purified protein was tested upon delivery to HEK293 cells as an RNP, followed by T7 endonuclease (T7EI) digestion assays after 48 hours.
[0087] The editing efficiency of 19 different recombinant forms of Cas9 protein was determined when expressed in HEK293 by transiently transfected plasmids. Furthermore, a subset of the constructs was delivered as RNPs. The methods and compositions disclosed herein identified adapters with improved editing efficiency relative to baseline flexible adapter designs. Rigid adapters rich in Pro sequence (XP)n and EAAAK repeat variants were found to lead to increased Cas9 activity upon plasmid delivery. Additionally, specific (XP)n repeat structures were found to exhibit enhanced editing activity upon RNP delivery.
[0088] Example 2
[0089] The expression of Cas9-eGFP protein with rigid linkers by plasmids demonstrated improved editing efficiency.
[0090] This example demonstrates that when human cells are expressed by a transfected plasmid, the introduction of a rigid linker between Cas9 and eGFP leads to enhanced Cas9 editing activity of the Cas9-eGFP fusion protein. The rigid and flexible linker sets generated in Example 1 were further tested in the Cas9-linker-eGFP context to evaluate their impact on Cas9 activity.
[0091] The editing activity of Cas9-eGFP fusion proteins containing rigid helical [A(EAAAK)4A] (linker SEQ ID NO:8), rigid alanine-proline [(AP)7A] (linker SEQ ID NO:16), flexible (GGGGS)4 (linker SEQ ID NO:7), or a mixture of flexible linkers GGGGSEAAAKGGGGS (linker SEQ ID NO:37) was compared with that of wild-type Cas9 protein. Cas9-eGFP protein with a flexible GSAGSAAGSGEF (abbreviated GSA; linker SEQ ID NO:6) linker was used as a baseline for comparing editing activity. The fusion protein was expressed by plasmid using a CMV promoter in HEK293 cells. Using a Lonza Nucleofector 96-well shuttle, the plasmid was delivered to a DS-150 with a background of 400 ng of protein expression plasmid and 350,000 cells per well using SF cell line solution. Cells were then divided into three wells after transfection. Twenty-four hours after plasmid delivery, sgRNAs targeting two sites within the HPRT1 gene region (HPRT-38087 (SEQ ID NO:324) and HPRT-38285 (SEQ ID NO:326)) were delivered at a final concentration of 30 nM using Lipofectamine RNAi Max reverse transcription. To analyze editing activity, Quick Extract was used 48 hours after reverse transcription. TMGenomic DNA was extracted using the extraction buffer. A 1kb region of the HPRT1 gene containing the target cleavage site was amplified by PCR. The PCR product was then unwound and rehybridized. The product was digested with a T7EI digester, and the ratio of the cleaved fragment to the full-length sequence was analyzed using a fragment analyzer (Agilent).
[0092] The results are shown in Figure 1, with error bars representing standard deviations among the three biological samples. While all four adapters showed improved activity compared to the baseline GSA adapter, both rigid adapters improved editing activity.
[0093] Example 3
[0094] To further test the rigid adapter, another experiment was performed using the same protocol as in Example 2. A set of 12 different gRNA targeting sites in the HPRTI gene (Table II) were tested. PCR primers are provided in Table III.
[0095] Table II. List of protospacer sequences for sgRNA
[0096]
[0097] Table III. Primer List
[0098]
[0099] The results are as follows Figure 2 As shown. Rigid connectors offer different levels of improvement depending on the target site relative to the original GSA connector design; both connectors provide improvement at each site.
[0100] Example 4
[0101] This example demonstrates the effect of adapter length by altering the alanine-proline repeat sequence (APA and adapter SEQ ID NO: 11-23) between 1 and 14 and testing two different extended helical adapter designs. The experiment followed the same protocol established in Example 2; however, the guide RNA was expressed by a co-delivered plasmid (100 ng per well), and cells were allowed to grow for three days prior to collecting genomic DNA.
[0102] The results are shown in Figures 3 and 4. Example 4 demonstrates that the rigid helical joint length or rigid AP joint length is functional in a wide range of repeating sequences.
[0103] Example 5
[0104] When delivered as RNP, the editing activity of Cas9-GFP and Cas9-mCherry linked by (AP)7A and (AP)9A is enhanced.
[0105] This example demonstrates that when delivered to human cells in the form of RNP, the introduction of the (AP)7A adapter (SEQ ID NO:16) or (AP)9A adapter (SEQ ID NO:18) between Cas9 and eGFP leads to an increase in the editing activity of the Cas9-eGFP fusion protein.
[0106] To determine whether rigid adapters increase Cas9-GFP editing activity when delivered as RNPs, Cas9-GFP adapter variants containing GSA adapters, helical rigid adapters, or alanine-proline rigid adapters were expressed and purified in *E. coli*. This was achieved by combining Cas9 with... sgRNA was incubated in PBS at a ratio of 1:1.2 for 10 min to form an RNP complex. Then, it was incubated with 4 μmol of [unspecified ingredient] at final RNP concentrations of 0.0625 μmol, 0.25 μmol, or 2.0 μmol. The Cas9 electroporation enhancer was delivered together with the Cas9 electroporation enhancer into HEK293 cells. Editing activity was determined by the T7EI assay as described above, and the results are shown in Figure 5.
[0107] At a 2 μmol dose of RNP, increased editing activity was observed in adapter variants containing (AP)7A adapters (SEQ ID NO:16; chimeric Cas9 SEQ ID NO:49) and (AP)9A adapters (SEQ ID NO:18; chimeric Cas9 SEQ ID NO:51) compared to flexible GSA adapters (Adapter SEQ ID NO:6; chimeric Cas9 SEQ ID NO:38) at the test site. Benefits from these adapters were also observed for a second distinct fluorophore. The second fluorophore tested was mCherry, the same (AP)... n Connector A is used to covalently attach mCherry to Cas9. As shown in Figure 5, the same rigid connector improves the editing activity of Cas9.
[0108] Example 6
[0109] Cas9-eGFP fusion cells were enriched using fluorescence-activated cell sorting (FACS).
[0110] This example demonstrates the use of (AP) between Cas9 and eGFP. n Connector A and use with Cas9-(AP) n A-eGFP compatible, using FACS to enrich edited cell populations.
[0111] To determine Cas9-(AP) nCan A-eGFP be used in conjunction with FACS to enrich cells with editing activity? As in Example 5, an RNP complex was formed using chimeric Cas9-(AP)7A-eGFP protein (SEQ ID NO:49) or wild-type Cas9 (SEQ ID NO:1), and delivered to cells using Lipofectamine RNAiMAX at a final concentration of 10 nM in three wells (1.2 million cells per well) of a 6-well dish. After approximately 16 hours, cells were digested with trypsin, washed with PBS, and resuspended in PBS containing 1% FBS. Cells were filtered through a 70 μM Flowmi tip filter (Bel-Art). Approximately 10% of the cells were transferred to collection tubes containing 500 μl PBS and 1% FBS as an unsorted control. Cells delivered with Cas9-eGFP were sorted based on GFP signal using a Becton Dickinson Aria II cell sorter. Cells were divided into three populations: those with apex ~20% signal, those with middle ~80-60% signal, and those with apex ~60% signal. Total editing activity in cells was measured by NGS 48-72 hours after delivery. Figure 6 The results show that the highest signal cell population exhibited significantly higher editing activity than the unsorted cell population, indicating that Cas9-eGFP fusion with a rigid adapter is suitable for this application.
[0112] Example 7
[0113] Evaluation of chimeric LbCas12a protein variants using eGFP protein
[0114] To determine whether rigid adapters increase editing efficiency in the presence of another RNA-guided endonuclease, LbCas12a protein variants (E795L) (SEQ ID NO:310) were fused with eGFP (SEQ ID NO:4) without the adapter (chimeric LbCas12a(E795L)-eGFP protein; SEQ ID NO:314) or using the subset of adapters expressed and purified in E. coli listed in Table 1. This was achieved by fusing unlabeled or chimeric LbCas12a(E795L) protein with... LbCas12a crRNA (see Table IV) was incubated in PBS at a ratio of 1:1.2 for 10 minutes to form an RNP complex, and then diluted to a final concentration of 50 nM along with 3 μmol of [unclear text - likely a typo, should be 3 μmol]. Cas12a electroporation enhancer was delivered to HEK293 cells. Editing activity was measured by NGS 48 hours after delivery, and the results showed... Figure 7The results show that, compared with other adapters without adapters or those tested, adapter variants containing (AP)7A and (AP)9A adapters exhibited increased editing activity. These are the same adapters that were beneficial for editing when chimeric Cas9-GFP protein was used.
[0115] Table IV. Guide RNA used in conjunction with the LbCas12a protein variant (E795L).
[0116] Serial Number: Original interval sequence Boot name 333 ACACACCCAAGGAAAGACTAT Cpf1 HPRT 38115-AS 334 TAATGCCCTGTAGTCTCTCTG Cpf1 HPRT 38186-S 335 GGTTAAAGATGGTTAAATGAT Cpf1 HPRT 38330-AS 336 TTGTAGGATATGCCCTTGACT Cpf1 HPRT 38486-S
[0117] Example 8
[0118] Exemplary sequences of Cas proteins and nucleic acids
[0119] The amino acid sequence of the wild-type Cas9 protein is shown in SEQ ID NO:1. Codon-optimized polynucleotides for expression in *E. coli* and human cells are given in SEQ ID NO:337 and SEQ ID NO:338, respectively. Exemplary variants of the Cas9 protein, the polynucleotides encoding them, and exemplary guide RNAs are disclosed in U.S. Patent Application Serial Nos. 15 / 729,491 and 15 / 964,041 (Attorney General's File Nos. IDT01-009-US and IDT01-009-US-CIP, respectively, filed October 10, 2017 and April 26, 2018, the contents of which are incorporated herein by reference.
[0120] SEQ ID NO:1 (Cas9 protein amino acid sequence)
[0121]
[0122] SEQ ID NO: 337 (Cas9 DNA sequence optimized for E. coli codons)
[0123]
[0124] SEQ ID NO: 338 (Cas9 DNA sequence optimized for Homo sapiens codons)
[0125]
[0126] The amino acid sequence of the wild-type AsCas12a protein is shown in SEQ ID NO:2. Codon-optimized polynucleotides for expression in *E. coli* and human cells are shown in SEQ ID NO:339 and SEQ ID NO:340, respectively. U.S. Patent Application Serial No. 16 / 536,256 (Attorney General's File No. IDT01-013-US), filed August 8, 2019, discloses exemplary variants of the AsCas12a protein, the polynucleotides encoding them, and exemplary guide RNAs, the contents of which are incorporated herein by reference.
[0127] SEQ ID NO:2 (AsCas12a protein amino acid sequence)
[0128]
[0129] SEQ ID NO:339 (AsCas12a DNA sequence optimized for E. coli codons)
[0130]
[0131]
[0132] The amino acid sequence of the wild-type LbCas12a protein is shown in SEQ ID NO:3. Codon-optimized polynucleotides for expression in *E. coli* and human cells are shown in SEQ ID NO:341 and SEQ ID NO:342, respectively. U.S. Patent Application Serial No. 63 / 0185,92 (Attorney’s File No. IDT01-017-PRO), filed May 1, 2020 (now U.S. Patent Application Serial No. 17 / 245,401, filed April 30, 2021), discloses exemplary variants of the LbCas12a protein, the polynucleotides encoding them, and exemplary guide RNAs, the contents of which are incorporated herein by reference.
[0133] SEQ ID NO:3 (LbCas12a protein amino acid sequence)
[0134]
[0135] SEQ ID NO:341 (LbCas12a DNA sequence optimized for E. coli codons)
[0136]
[0137] SEQ ID NO:342 (LbCas12a DNA sequence optimized for Homo sapiens codons)
[0138]
[0139] Example 9
[0140] Fluorescent protein sequence
[0141] SEQ ID NO:4 (eGFP protein amino acid sequence) (Cormack et al. (1996))
[0142] MVSKGEELFTGVVPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTTLTYGVQCFSRYPDHMKQHDFFKSAMPEGYVQERTIFFKDDGNYKT RAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYNYNSHNVYIMADKQKNGIKVNFKIRHNIEDGSVQLADHYQQNTPIGDGPVLLPDNHYLSTQSALSKDPNEKRDHM
[0143] VLLEFVTAAGITLGMDELYK
[0144] SEQ ID NO:343eGFP DNA sequence (codon optimized for E. coli)
[0145] ATGGTAAGTAAAGGTGAAGAGCTTTTTACAGGTGTGGTGCCCATCCTTGTGGAGTTGGACGGTGATGTCAATGGCCATAAATTTTCTGTGTCTGGTGAAGGGGAAGGGGACGCGACGTATGGAAAACTTACCCTGAAGTTTATCTGCACGACAGGTAAGTTGCCAGTACCGTGGCCTACCCTGGTCACCACATTAACATATGGTGTTCAATGCTTTTCACGCTACCCTGACCACATGAAACAACATGACTTTTTTAAAAGTGCCATGCCAGAGGGCTACGTGCAGGAACGCACAATCTTTTTCAAAGACGACGGCAATTATAAAACACGCGCGGAGGTAAAGTTTGAGGGAGACACACTGGTTAATCGCATCGAACTGAAAGGCATTGACTTTAAGGAGGATGGGAATATCTTAGGCCATAAACTGGAGTATAACTATAACTCTCACAACGTCTATATTATGGCGGACAAGCAAAAGAATGGTATCAAGGTAAACTTTAAGATTCGTCATAACATTGAGGATGGGAGCGTGCAGTTGGCTGACCACTATCAGCAGAATACTCCCATTGGCGACGGCCCCGTGCTTTTACCTGACAATCACTATCTTTCTACGCAGTCAGCTTTGTCCAAAGACCCCAACGAGAAGCGTGATCACATGGTTCTTCTGGAATTCGTCACAGCCGCCGGAATCACTTTAGGCATGGACGAACTTTATAAA
[0146] SEQ ID NO: 344 eGFP DNA sequence (codon-optimized for Homo sapiens)
[0147] ATGGTCTCAAAAGGCGAAGAATTGTTCACCGGTGTAGTCCCTATCTTGGTGGAGCTTGACGGCGATGTTAATGGCCATAAATTTAGTGTGTCCGGCGAAGGTGAGGGCGACGCCACATATGGTAAACTTACGCTTAAATTTATTTGCACGACGGGAAAGCTCCCCGTGCCGTGGCCAACCCTTGTAACTACCCTTACGTACGGGGTGCAGTGCTTCAGTAGGTATCCCGACCATATGAAGCAGCACGATTTTTTCAAAAGTGCTATGCCCGAGGGGTATGTGCAAGAGAGGACTATATTCTTTAAGGATGATGGGAATTACAAGACGCGAGCCGAGGTTAAGTTCGAAGGTGATACGCTTGTAAATCGAATCGAATTGAAAGGGATTGACTTTAAGGAAGACGGAAATATACTTGGACACAAATTGGAATATAACTACAACAGCCACAACGTCTATATAATGGCCGACAAGCAAAAGAACGGAATAAAAGTCAACTTTAAGATTCGACACAATATAGAGGATGGATCCGTGCAGCTTGCTGACCACTATCAGCAGAACACTCCGATAGGCGATGGACCAGTGCTGCTTCCCGACAATCACTACCTCTCCACGCAGTCTGCTCTGTCTAAAGACCCTAATGAAAAACGAGACCACATGGTGCTGCTCGAATTTGTCACTGCCGCCGGGATTACCTTGGGAATGGACGAACTGTATAAG
[0148] SEQ ID NO:5 (Amino acid sequence of mCherry) (Shaner et al. (2004))
[0149] MVSKGEEDNMAIIKEFMRFKVHMEGSVNGHEFEIEGEGEGRPYEGTQTAKLKVTKGGPLPFAWDILSPQFMYGSKAYVKHPADIPDYLKLSFPEGFKWERVMNFEDGGVVTVTQDSSL QDGEFIYKVKLRGTNFPSDGPVMQKKTMGWEASSERMYPEDGALKGEIKQRLKLKDGGHYDAEVKTTYKAKKPVQLPGAYNVNIKLDITSHNEDYTIVEQYERAEGRHSTGGMDELYK
[0150] SEQ ID NO:345 (mCherry DNA sequence optimized for E. coli codons)
[0151] ATGGTGTCCAAAGGAGAAGAGGATAACATGGCGATCATCAAGGAATTTATGCGTTTCAAAGTTCACATGGAGGGAAGCGTAAACGGCCATGAATTTGAGATTGAGGGGGAAGGCGAAGGCCGTCCATACGAAGGCACTCAAACCGCCAAGCTTAAGGTTACCAAGGGAGGTCCGCTTCCTTTCGCCTGGGATATCCTTAGCCCACAGTTCATGTATGGCTCTAAAGCATACGTCAAGCACCCTGCAGACATTCCTGATTATTTGAAACTTAGTTTTCCTGAGGGTTTTAAGTGGGAACGCGTTATGAATTTCGAGGATGGAGGAGTAGTGACAGTAACTCAAGATTCCAGTCTTCAAGACGGAGAGTTTATCTACAAGGTGAAATTACGTGGCACGAACTTTCCATCCGACGGGCCAGTGATGCAGAAGAAAACCATGGGTTGGGAGGCATCTTCTGAGCGTATGTACCCGGAAGACGGGGCGTTGAAAGGCGAAATCAAGCAGCGCCTGAAGTTGAAAGACGGGGGTCACTACGATGCGGAAGTAAAAACTACATATAAGGCCAAGAAGCCCGTCCAGCTTCCGGGCGCGTACAACGTCAACATCAAATTGGACATTACTTCCCACAACGAGGATTACACTATTGTTGAACAATATGAGCGTGCGGAGGGACGCCATTCAACCGGGGGGATGGATGAGCTTTACAAG
[0152] SEQ ID NO:346 (mCherry DNA sequence optimized for Homo sapiens codons)
[0153] ATGGTCAGTAAGGGGGGAAGAGGACAACATGGCGATAATCAAGGAATTTATGAGATTTAAGGTCCATATGGAAGGCCTGTCAACGGTCATGAGTTCGAAATTGAGGGGGAGGGGGAAGGCCGCCCATATGAGGGGACTCAAACAGCCAAACTTAAGGTCACAAAAGGAGGTCCTCTG CCCTTCGCGTGGGACATACTCAGCCCACAATTCATGTATGGAAGTAAGGCATATGTTAAACACCCGGCGGACATACCCGACTACCTCAAACTGTCATTTCCTGAGGGTTTTAAGTGGGAGAGGGTTATGAATTTCGAGGACGGTGGAGTAGTAACAGTGACACAAGACTCTTCTCTC CAGGATGGAGAATTTATATACAAAGTGAAACTGCGGGGCACTAACTTTCCGTCAGACGGCCCAGTAATGCAAAAGAAAACTATGGGGTGGGAGGCTTCCAGCGAGCGCATGTATCCCGAGGATGGGGCCCTTAAGGGAGAAATAAAACAACGCTTGAAGCTCAAAGACGGAGGGCAC TATGATGCGGAGGTCAAAACGACCTACAAAGCAAAAAAGCCAGTACAGCTTCCGGGAGCTTATAACGTGAATATAAAGCTCGATATAACCTCACACAACGAGGACTACACGATTGTAGAACAATACGAAAGAGCCGAGGGGACATAGTACGGGGGGCATGGACGAACTTTACAAG
[0154] Example 10
[0155] Chimeric protein sequence
[0156] Table V.1 provides information on chimeric fusion proteins between the Cas9 amino acid sequence, the linker amino acid sequence (underlined), and the eGFP amino acid sequence (bold, italic).
[0157] Table V.1. Cas9-linker-eGFP chimerism
[0158]
[0159]
[0160]
[0161]
[0162]
[0163]
[0164]
[0165]
[0166]
[0167]
[0168]
[0169]
[0170]
[0171]
[0172]
[0173]
[0174]
[0175]
[0176]
[0177]
[0178]
[0179]
[0180]
[0181]
[0182]
[0183] Table V.2 provides information on chimeric fusion proteins between the Cas9 amino acid sequence, the linker amino acid sequence (underlined), and the mCherry amino acid sequence (bold, italic).
[0184] Table V.Cas9 - Connector - mCherry Fitting
[0185]
[0186]
[0187]
[0188]
[0189]
[0190]
[0191]
[0192]
[0193]
[0194]
[0195]
[0196]
[0197]
[0198]
[0199]
[0200]
[0201]
[0202]
[0203]
[0204]
[0205]
[0206]
[0207]
[0208]
[0209]
[0210]
[0211] Table V.3 provides information on the chimeric fusion protein between the AsCas12a amino acid sequence, the linker amino acid sequence (underlined), and the eGFP amino acid sequence (bold, italic).
[0212] Table V.3. AsCas12a-linker-eGFP chimerism
[0213]
[0214]
[0215]
[0216]
[0217]
[0218]
[0219]
[0220]
[0221]
[0222]
[0223]
[0224]
[0225]
[0226]
[0227]
[0228]
[0229]
[0230]
[0231]
[0232]
[0233]
[0234]
[0235]
[0236]
[0237] Table V.4 provides information on the chimeric fusion protein between the AsCas12a amino acid sequence, the linker amino acid sequence (underlined), and the mCherry amino acid sequence (bold, italic).
[0238] Table V.4. AsCas12a - Connector - mCherry Fitting
[0239]
[0240]
[0241]
[0242]
[0243]
[0244]
[0245]
[0246]
[0247]
[0248]
[0249]
[0250]
[0251]
[0252]
[0253]
[0254]
[0255]
[0256]
[0257]
[0258]
[0259]
[0260]
[0261]
[0262]
[0263] Table V.5 provides information on the chimeric fusion protein between the LbCas12a amino acid sequence, the linker amino acid sequence (underlined), and the eGFP amino acid sequence (bold, italic).
[0264] Table V.5. LbCas12a-linker-eGFP chimerism
[0265]
[0266]
[0267]
[0268]
[0269]
[0270]
[0271]
[0272]
[0273]
[0274]
[0275]
[0276]
[0277]
[0278]
[0279]
[0280]
[0281]
[0282]
[0283]
[0284]
[0285]
[0286]
[0287]
[0288] Table V.6 provides information on the chimeric fusion protein between the LbCas12a amino acid sequence, the linker amino acid sequence (underlined), and the mCherry amino acid sequence (bold, italic).
[0289] Table V.6.LbCas 12a - Connector - mCherry Fitting
[0290]
[0291]
[0292]
[0293]
[0294]
[0295]
[0296]
[0297]
[0298]
[0299]
[0300]
[0301]
[0302]
[0303]
[0304]
[0305]
[0306]
[0307]
[0308]
[0309]
[0310]
[0311]
[0312]
[0313]
[0314] Table V.7 provides information on the chimeric fusion protein between the LbCas12a variant amino acid sequence (E795L), the linker amino acid sequence (underlined), and the eGFP amino acid sequence (bold, italic).
[0315] Table V.7. LbCas12a(E795L) - Linker - eGFP Chimerism
[0316]
[0317]
[0318]
[0319]
[0320]
[0321]
[0322]
[0323]
[0324]
[0325]
[0326]
[0327]
[0328]
[0329]
[0330]
[0331]
[0332]
[0333]
[0334]
[0335]
[0336]
[0337]
[0338]
[0339] Table V.8 provides information on the chimeric fusion protein between the LbCas12a variant amino acid sequence (E795L), the linker amino acid sequence (underlined), and the mCherry amino acid sequence (bold, italic).
[0340] Table V.8.LbCas12a(E795L) - Connector - mCherry Fitting
[0341]
[0342]
[0343]
[0344]
[0345]
[0346]
[0347]
[0348]
[0349]
[0350]
[0351]
[0352]
[0353]
[0354]
[0355]
[0356]
[0357]
[0358]
[0359]
[0360]
[0361]
[0362]
[0363]
[0364] Table V.9 provides the amino acid sequence of the LbCas12a protein variant (E795L).
[0365] Table V.9. Amino acid sequence of LbCas12a protein variant (E795L)
[0366]
[0367]
[0368] Table V.10 provides chimeric fusion proteins of Cas protein amino acid sequences with no inserting linker polypeptides between the amino acid sequences of eGFP or mCherry (bold, italic).
[0369] Table V.10. Cas protein-fluorescent protein chimerism
[0370]
[0371]
[0372]
[0373]
[0374]
[0375]
[0376] Exemplary DNA sequence encoding Cas9 protein chimerism
[0377] SEQ ID NO: 347 (Wild-type Cas9 protein DNA sequence)
[0378]
[0379] SEQ ID NO: 348 (Cas9-GSAGSAAGSGEF-eGFP interlaced DNA sequence)
[0380]
[0381] SEQ ID NO: 349 (Cas9-(GGGGS)4-eGFP intermigration DNA sequence)
[0382]
[0383] SEQ ID NO: 350 (Cas9-GGGGSEAAAKGGGGS-eGFP interlaced DNA sequence)
[0384]
[0385] SEQ ID NO: 351 (Cas9-A(EAAAK)4A-eGFP intermigration DNA sequence)
[0386]
[0387] SEQ ID NO: 352 (Cas9-A(EAAAK)4ALEA(EAAAK)4A-eGFP intermating DNA sequence)
[0388]
[0389] SEQ ID NO: 353 (Cas9-LEA(EAAAK)4ALEA(EAAAK)4ALE-eGFP interlock DNA sequence)
[0390]
[0391] SEQ ID NO: 354 (Cas9-APA-eGFP chimeric DNA sequence)
[0392]
[0393] SEQ ID NO: 355 (Cas9-APAPA-eGFP chimeric DNA sequence)
[0394]
[0395] SEQ ID NO: 356 (Cas9-APAPAPA-eGFP chimeric DNA sequence)
[0396]
[0397] SEQ ID NO: 357 (Cas9-APAPAPAPA-eGFP chimeric DNA sequence)
[0398]
[0399] SEQ ID NO: 358 (Cas9-APAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0400]
[0401] SEQ ID NO: 359 (Cas9-APAPAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0402]
[0403] SEQ ID NO: 360 (Cas9-APAPAPAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0404]
[0405] SEQ ID NO: 361 (Cas9-APAPAPAPAPAPAPAPAPA-eGFP chimera DNA sequence)
[0406]
[0407] SEQ ID NO:362 (Cas9-APAPAPAPAPAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0408]
[0409] SEQ ID NO:363 (Cas9-APAPAPAPAPAPAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0410]
[0411] SEQ ID NO:364 (Cas9-APAPAPAPAPAPAPAPAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0412]
[0413] SEQ ID NO:365 (Cas9-APAPAPAPAPAPAPAPAPAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0414]
[0415] SEQ ID NO:366 (Cas9-APAPAPAPAPAPAPAPAPAPAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0416]
[0417] SEQ ID NO:367 (Cas9-APAPAPAPAPAPAPAPAPAPAPAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0418]
[0419] SEQ ID NO:368 (Cas9-GSAGSAAGSGEF-mCherry intermotherapeutic DNA sequence)
[0420]
[0421] SEQ ID NO:369 (Cas9-(GGGGS)4-mCherry intermotherapeutic DNA sequence)
[0422]
[0423] SEQ ID NO:370 (Cas9-GGGGSEAAAKGGGGS-mCherry intermotherapeutic DNA sequence)
[0424]
[0425] SEQ ID NO:371 (Cas9-A(EAAAK)4A-mCherry intermating DNA sequence)
[0426]
[0427] SEQ ID NO: 372 (Cas9-A(EAAAK)4ALEA(EAAAK)4A-mCherry intermotherapeutic DNA sequence)
[0428]
[0429] SEQ ID NO: 373 (Cas9-LEA(EAAAK)4ALEA(EAAAK)4ALE-mCherry intermotherapeutic DNA sequence)
[0430]
[0431] SEQ ID NO:374 (Cas9-APA-mCherry intermotherapeutic DNA sequence)
[0432]
[0433] SEQ ID NO:375 (Cas9-APAPA-mCherry chimeric DNA sequence)
[0434]
[0435] SEQ ID NO:376 (Cas9-APAPAPA-mCherry chimeric DNA sequence)
[0436]
[0437] SEQ ID NO:377 (Cas9-APAPAPAPA-mCherry chimeric DNA sequence)
[0438]
[0439] SEQ ID NO:378 (Cas9-APAPAPAPAPAPA-mCherry chimeric DNA sequence)
[0440]
[0441] SEQ ID NO:379 (Cas9-APAPAPAPAPAPAPA-mCherry chimeric DNA sequence)
[0442]
[0443] SEQ ID NO:380 (Cas9-APAPAPAPAPAPAPAPA-mCherry chimeric DNA sequence)
[0444]
[0445] SEQ ID NO:381 (Cas9-APAPAPAPAPAPAPAPAPA-mCherry chimeric DNA sequence)
[0446]
[0447] SEQ ID NO:382 (Cas9-APAPAPAPAPAPAPAPAPAPAPA-mCherry chimeric DNA sequence)
[0448]
[0449] SEQ ID NO:383 (Cas9-APAPAPAPAPAPAPAPAPAPAPAPA-mCherry chimeric DNA sequence)
[0450]
[0451] SEQ ID NO:384 (Cas9-APAPAPAPAPAPAPAPAPAPAPAPAPA-mCherry chimeric DNA sequence)
[0452]
[0453] SEQ ID NO:385 (Cas9-APAPAPAPAPAPAPAPAPAPAPAPAPAPA-mCherry chimeric DNA sequence)
[0454]
[0455] SEQ ID NO:386 (Cas9-APAPAPAPAPAPAPAPAPAPAPAPAPAPAPA-mCherry chimeric DNA sequence)
[0456]
[0457] SEQ ID NO:387 (Cas9-APAPAPAPAPAPAPAPAPAPAPAPAPAPAPAPA-mCherry chimeric DNA sequence)
[0458]
[0459] An exemplary DNA sequence encoding an amino acid chimera of the LbCas12a (E795L) variant protein.
[0460] SEQ ID NO:388 (LbCas12a(E795L)-eGFP chimeric DNA sequence)
[0461]
[0462]
[0463] SEQ ID NO:390(LbCas12a(E795L)-(GGGGS)4-eGFP chimeric DNA sequence)
[0464]
[0465] SEQ ID NO:391(LbCas12a(E795L)-GGGGSEAAAKGGGGS-eGFP chimeric DNA sequence)
[0466]
[0467] SEQ ID NO:392(LbCas12a(E795L)-A(EAAAK)4A-eGFP chimeric DNA sequence)
[0468]
[0469] SEQ ID NO:393(LbCas12a(E795L)-A(EAAAK)4ALEA(EAAAK)4A-eGFP chimeric DNA sequence)
[0470]
[0471] SEQ ID NO:394(LbCas12a(E795L)-LEA(EAAAK)4ALEA(EAAAK)4ALE-eGFP chimeric DNA sequence)
[0472]
[0473] SEQ ID NO:395(LbCas12a(E795L)-APA-eGFP chimeric DNA sequence)
[0474]
[0475] SEQ ID NO:396(LbCas12a(E795L)-APAPA-eGFP chimeric DNA sequence)
[0476]
[0477] SEQ ID NO:397(LbCas12a(E795L)-APAPAPA-eGFP chimeric DNA sequence)
[0478]
[0479] SEQ ID NO:398(LbCas12a(E795L)-APAPAPAPA-eGFP chimeric DNA sequence)
[0480]
[0481] SEQ ID NO:399(LbCas12a(E795L)-APAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0482]
[0483]
[0484] SEQ ID NO:401(LbCas12a(E795L)-APAPAPAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0485]
[0486] SEQ ID NO:402(LbCas12a(E795L)-APAPAPAPAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0487]
[0488] SEQ ID NO:403(LbCas12a(E795L)-APAPAPAPAPAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0489]
[0490] SEQ ID NO:404(LbCas12a(E795L)-APAPAPAPAPAPAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0491]
[0492] SEQ ID NO:405(LbCas12a(E795L)-APAPAPAPAPAPAPAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0493]
[0494] SEQ ID NO:406(LbCas12a(E795L)-APAPAPAPAPAPAPAPAPAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0495]
[0496] SEQ ID NO:407(LbCas12a(E795L)-APAPAPAPAPAPAPAPAPAPAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0497]
[0498] SEQ ID NO:408(LbCas12a(E795L)-APAPAPAPAPAPAPAPAPAPAPAPAPAPAPAPA-eGFP chimeric DNA sequence)
[0499]
[0500] SEQ ID NO:409 (LbCas12a(E795L)-mCherry chimeric DNA sequence)
[0501]
[0502] SEQ ID NO: 410 (LbCas12a(E795L)-GSAGSAAGSGEF-mCherry intermotherapeutic DNA sequence)
[0503]
[0504] SEQ ID NO:411(LbCas12a(E795L)-(GGGGS)4-mCherry chimeric DNA sequence)
[0505]
[0506] SEQ ID NO:412(LbCas12a(E795L)-GGGGSEAAAKGGGGS-mCherry chimeric DNA sequence)
[0507]
[0508] SEQ ID NO:413(LbCas12a(E795L)-A(EAAAK)4A-mCherry chimeric DNA sequence)
[0509]
[0510] SEQ ID NO:414(LbCas12a(E795L)-A(EAAAK)4ALEA(EAAAK)4A-mCherry chimeric DNA sequence)
[0511]
[0512] SEQ ID NO:415(LbCas12a(E795L)-LEA(EAAAK)4ALEA(EAAAK)4ALE-mCherry chimeric DNA sequence)
[0513]
[0514] SEQ ID NO:416(LbCas12a(E795L)-APA-mCherry chimeric DNA sequence)
[0515]
[0516] SEQ ID NO:417(LbCas12a(E795L)-APAPA-mCherry chimeric DNA sequence)
[0517]
[0518] SEQ ID NO:418(LbCas12a(E795L)-APAPAPA-mCherry chimeric DNA sequence)
[0519]
[0520] SEQ ID NO:419(LbCas12a(E795L)-APAPAPAPA-mCherry chimeric DNA sequence)
[0521]
[0522] SEQ ID NO:420(LbCas12a(E795L)-APAPAPAPAPAPA-mCherry chimeric DNA sequence)
[0523]
[0524] SEQ ID NO:421(LbCas12a(E795L)-APAPAPAPAPAPAPA-mCherry chimeric DNA sequence)
[0525]
[0526] SEQ ID NO:422(LbCas12a(E795L)-APAPAPAPAPAPAPAPA-mCherry chimeric DNA sequence)
[0527]
[0528] SEQ ID NO:423(LbCas12a(E795L)-APAPAPAPAPAPAPAPAPA-mCherry chimeric DNA sequence)
[0529]
[0530] SEQ ID NO:424(LbCas12a(E795L)-APAPAPAPAPAPAPAPAPAPA-mCherry chimeric DNA sequence)
[0531]
[0532] SEQ ID NO:425(LbCas12a(E795L)-APAPAPAPAPAPAPAPAPAPAPAPA-mCherry chimeric DNA sequence)
[0533]
[0534]
[0535] SEQ ID NO:427(LbCas12a(E795L)-APAPAPAPAPAPAPAPAPAPAPAPAPAP-mCherry chimeric DNA sequence)
[0536]
[0537] SEQ ID NO:428(LbCas12a(E795L)-APAPAPAPAPAPAPAPAPAPAPAPAPAPAP-mCherry chimeric DNA sequence)
[0538]
[0539] SEQ ID NO:429(LbCas12a(E795L)-APAPAPAPAPAPAPAPAPAPAPAPAPAPAPAP-mCherry chimeric DNA sequence)
[0540]
[0541] All references cited in this article, including publications, patent applications and patents, are incorporated herein by reference as if each reference were individually and specifically indicated to be incorporated herein by reference and fully elaborated herein.
[0542] References
[0543] 1. Jinek, M., et al., A programmabledual-RNA-guided DNA endonuclease inadaptive bacterial immunity. Science, 2012.337(6096): p.816-21.
[0544] 2. Rees, HA and DRLiu, Base editing: precision chemistry on the genome and transcriptome of living cells. Nat Rev Genet, 2018.19(12):p.770-788.
[0545] 3.Komor,AC,et al.,Programmable editing of a target base in genomicDNA without double-stranded DNA cleavage.Nature,2016.533(7603):p.420-4.
[0546] 4.Gaudelli, NM, et al., Programmable base editing of A*T to G*C ingenomic DNA without DNA cleavage. Nature, 2017.551(7681): p.464-471.
[0547] 5.Nishida,K.,et al.,Targeted nucleotide editing using hybridprokaryotic and vertebrate adaptive immune systems.Science,2016.353(6305).
[0548] 6.Qi,L.S.,et al.,Repurposing CRISPR as an RNA-guided platform forsequence specific control of gene expression.Cell,2013.152(5):p.1173-83.
[0549] 7.Gilbert,L.A.,et al.,CRISPR-mediated modular RNA-guided regulationof transcription in eukaryotes.Cell,2013.154(2):p.442-51.
[0550] 8.Gilbert,L.A.,et al.,Genome-Scale CRISPR-Mediated Control of GeneRepression and Activation.Cell,2014.159(3):p.647-61.
[0551] 9.Anzalone,A.V.,et al.,Search-and-replace genome editing withoutdouble-strand breaks or donor DNA.Nature,2019.576(7785):p.149-157.
[0552] 10.Cong,L.,et al.,Multiplex genome engineering using CRISPR / Cassystems.Science,2013.339(6121):p.819-23.
[0553] 11.Vakulskas,C.A.,et al.,A high-fidelity Cas9 mutant delivered as aribonucleoprotein complex enables efficient gene editing in humanhematopoietic stem and progenitor cells.Nat Med,2018.24(8):p.1216-1224.
[0554] 12.Slaymaker,I.M.,et al.,Rationally engineered Cas9 nucleases withimproved specificity.Science,2016.351(6268):p.84-8.
[0555] 13.Kleinstiver,B.P.,et al.,High-fidelity CRISPR-Cas9 nucleases withno detectable genome-wide off-target effects.Nature,2016.529(7587):p.490-5.
[0556] 14.Chen,J.S.,et al.,Enhanced proofreading governs CRISPR-Cas9targeting accuracy.Nature,2017.550(7676):p.407-410.
[0557] 15.McCormick,A.L.,M.S.Thomas,and A.W.Heath,Immunization with aninterferongamma-gp120 fusion protein induces enhanced immune responses tohuman immunodeficiency virus gp120.J Infect Dis,2001.184(11):p.1423-30.
[0558] 16.Bai,Y.and W.C.Shen,Improving the oral efficacy of recombinantgranulocyte colony-stimulating factor and transferrin fusion protein byspacer optimization.Pharm Res,2006.23(9):p.2116-21.
[0559] 17.Wriggers,W.,S.Chakravarty,and P.A.Jennings,Control of proteinfunctional dynamics by peptide linkers.Biopolymers,2005.80(6):p.736-46.
[0560] 18.Arai,R.,et al.,Design of the linkers which effectively separatedomains of a bifunctional fusion protein.Protein Eng,2001.14(8):p.529-32.
[0561] 19.Marqusee,S.and R.L.Baldwin,Helix stabilization by Glu-...Lys+saltbridges in short peptides of de novo design.Proc Natl Acad Sci U S A,1987.84(24):p.8898-902.
[0562] 20.Bhandari,D.G.,et al.,1H-NMR study of mobility and conformationalconstraints within the proline-rich N-terminal of the LC1 alkali light chainof skeletal myosin.Correlation with similar segments in other proteinsystems.Eur J Biochem,1986.160(2):p.349-56.
[0563] 21.Chen,X.,J.L.Zaro,and W.C.Shen,Fusion protein linkers:property,design and functionality.Adv Drug Deliv Rev,2013.65(10):p.1357-69.
[0564] 22.Brendan P.Cormack,Raphael H.Valdivia,Stanley Falkow,FACS-optimizedmutants of the green fluorescent protein(GFP),Gene,1996.173(1):33-38.
[0565] 23.Shaner NC,Campbell RE,Steinbach PA,Giepmans BN,Palmer AE,TsienRY.Improved monomeric red,orange and yellow fluorescent proteins derived fromDiscosoma sp.red fluorescent protein.Nat Biotechnol.2004 Dec;22(12):1567-72.
[0566] 24.Peng,R.,Li,Z.,Xu,Y.,He,S.,Peng,Q.,Wu,L.A.,Wu,Y.,Qi,J.,Wang,P.,Shi,Y.and Gao,G.F.Structural insight into multistage inhibition of CRISPR-Cas12aby AcrVA4,Proc.Natl.Acad.Sci.U.S.A.2019 116(38):18928-18936.
Claims
1. A chimeric protein comprising a guided endonuclease protein, a rigid linker, and a second protein, wherein the chimeric protein is selected from the group consisting of SEQ ID NO: 49, SEQ ID NO: 51, SEQ ID NO: 83, and SEQ ID NO:
85.
2. A method for enriching cells containing chimeric proteins, comprising: a) Incubate the chimeric protein according to claim 1 with guide RNA to form an RNP complex; b) Contact the RNP complex with multiple target cells to generate recipient cells with the RNP complex; and c) Receptor cells are sorted based on fluorescence signals.
3. An isolated nucleic acid, wherein the isolated nucleic acid encodes the chimeric protein according to claim 1.
4. The isolated nucleic acid according to claim 3, wherein the isolated nucleic acid is codon-optimized for expression in an organism or host cell.
5. The nucleic acid isolation according to claim 4, wherein the organism or host cell is selected from Escherichia coli or Homo sapiens.
6. The method for enriching cells with chimeric proteins according to claim 2 further includes a step of determining editing activity by at least one locus targeted by guide RNA.
7. The method of claim 6, wherein the step of determining editing activity at at least one locus targeted by the guide RNA is selected from T7EI assay and next-generation sequencing assay.