Nuclease system for genome editing
Chimeric nucleases formed by combining Cpfl domains from multiple species enhance editing efficiency and specificity in mammalian cells, addressing the limitations of existing Cpfl-based systems.
Patent Information
- Application Number
- PCT/US2025/020754
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-21
- Filing Date
- 2025-03-20
- Publication Date
- 2025-09-25
AI Technical Summary
Existing CRISPR-Cas systems, particularly Cpfl-based systems, face challenges in achieving high editing efficiency and specificity, especially in mammalian cells, with existing chimeric nucleases like M44 showing inconsistent performance across different guide RNA sequences.
Development of chimeric nucleic acid-guided nucleases by combining domains from at least two distinct Cpfl species, such as Eubacterium rectale (ErCpfl) with domains from other species, creating a highly efficient and specific nucleic acid-guided nuclease system.
The chimeric nucleases demonstrate enhanced nuclease activity and specificity, rescuing non-functional guide RNAs and improving genome editing efficiency in mammalian cells.
Smart Images

Figure US2025020754_25092025_PF_FP_ABST
Abstract
Description
NUCLEASE SYSTEM FOR GENOME EDITINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority under 35 U.S.C. § 119(e) to U.S. Provisional Application No. 63 / 568,362, filed March 21, 2024. The content of the prior application is considered part of and is hereby incorporated by reference in its entirety.INCORPORATION BY REFERENCE OF SEQUENCE LISTING
[0002] The material in the accompanying sequence listing is hereby incorporated by reference into this application and is hereby submitted via patent center. The accompanying sequence listing xml file, name BAYS1110-lWO.xml, was created on February 24, 2025, and is 172,603 bytes.BACKGROUND OF THE INVENTIONFIELD OF THE INVENTION
[0003] The present disclosure relates generally to systems and methods for editing a genome and more specifically to chimeric nucleic acid-guided nucleases with improved editing efficiency, specificity and accuracy.BACKGROUND INFORMATION
[0004] The discovery of CRISPR adaptive immune systems in bacteria and archaea, which rely upon RNA-guided nucleases to target and cleave nucleic acids from invading viruses or mobile genetic elements, has revolutionized the field of molecular biology. Researchers are increasingly using these nucleic acid-guided nucleases to rapidly modify the genomes of many organisms.
[0005] CRISPR-associated (Cas) systems have evolved in many different species of bacteria and archaea. Each CRISPR-Cas system consists of a combination of Cas effector proteins and CRISPR RNAs (crRNAs). The defense activity of a typical CRISPR-Cas system covers three stages: (1) adaptation, which occurs when a complex of Cas proteins excises a segment of the target DNA (referred to as a “proto spacer”) and inserts it between the repeats at the 5’ end of a CRISPR array in the host genome, thus yielding a new spacer; (2) expression and processing of the precursor CRISPR RNA (pre-crRNA), resulting in the formation of mature small CRISPR RNAs (crRNAs); and (3) interference, when the effector module - either another Cas protein complex or a single multidomain protein - uses the crRNA as a guide to recognize and clear target DNA or RNA. Although the proteins involved in the adaptation stage (Casl and Cas2) are highly conserved among CRISPR-Cas systems, there is considerable diversity in the proteins involved in the pre-crRNA to crRNA processing step, as well as in the effector modules that mediate the recognition and cleavage of the target DNA or RNA.
[0006] As with other immune defense mechanisms, CRISPR-Cas systems evolved in the context of an ongoing arms race, with bacteria or archaea on one side, and viruses or mobile genetic elements on the other. Each bacteria or archaea must survive in a specific environment and defend itself against a specific set of invaders (e.g., a specific set of viruses). Thus, the specificity, accuracy and efficiency of CRISPR-Cas systems can vary greatly depending upon the species from which they are derived.
[0007] Researchers have separated CRISPR-Cas systems into two different classes based on the number of effector protein subunits they contain. These two classes have been further sorted into six types. The effector complexes of Class 1 CRISPR-Cas systems typically consist of 4-7 Cas protein subunits in an uneven stoichiometry; in contrast, Class 2 CRISPR-Cas systems only require a single, large, multidomain Cas protein. This relatively simple architecture has made Class 2 CRISPR-Cas systems an increasingly popular choice for use in genome editing.
[0008] Cas9 is a Class 2, Type II CRISPR-Cas nuclease often used in genome editing. When bound to a single guide RNA, or a combination of a short targeting crRNA and an accessory transactivating CRISPR RNA (tracrRNA), Cas9 will hybridize to a recognition sequence within the target genome. The Cas9 recognition sequence is typically about 20 nucleotides long and is located near a short protospacer adjacent motif (PAM). Cas9 discriminates strongly against mismatches within the first 10 or so base pairs of the RNA-DNA helix (also called the R-loop) that is proximal to the PAM. Once bound, the two nuclease domains (HNH and RuvC) will each cut one target DNA strand in a near-simultaneous manner, leaving blunt ends (no overhanging nucleotides). In mammalian cells, CRISPR-induced double-strand breaks (DSBs) are repaired by either of two mechanisms: (1) non-homologous end joining (NHEJ); or (2) homology-directed repair (HDR). NHEJ can result in the incorporation of small insertions / deletions near the cleavage site, which in turn can lead to gene inactivation due to frameshift mutations. When two DSBs occur on the same chromosome, a substantial segment can be deleted, whereas DSBs on different chromosomes can give rise to chromosomal rearrangements. In contrast, HDR of a CRISPR-induced DSB can result in more precise gene edits if an exogenous DNA sequence is used as a template for correction.
[0009] Cpfl is a Class 2, Type V CRISPR-Cas nuclease that was more recently discovered, and it differs from Cas9 in several important respects. For example, Cpfl is a single-RNA-guided nuclease that does not require a tracrRNA and can auto-processes its pre-crRNA into mature crRNA. It also uses shorter guide RNAs compared to Cas9, and has different PAM specificities than Cas9 (e.g., Cpfl recognizes T-rich PAM sequences instead of guanine-rich PAM sequenceslike Cas9). Although Cpfl was initially thought to contain only one endonuclease domain (RuvC), x-ray crystallographic studies of Cpfl in complex with crRNA and target DNA revealed a second nuclease domain with a unique fold (Nuc) which is functionally analogous to the HNH nuclease domain of Cas9. Cpfl makes staggered cuts in double-stranded DNA, which leaves short 4- nucleotide overhangs on the 5' end of each DNA strand. These “sticky ends” are thought to improve recombination repair efficiency. Further, Cpfl has both DNA and RNA nuclease activity, making it more attractive for use when trying to modify multiple genomic loci.
[0010] Given the widespread interest in developing Cpfl -based CRISPR-Cas systems as an alternative to CRISPR-Cas9 systems, researchers have sought to modify the editing characteristics of Cpfl -like nucleases by generating chimeras in which conserved functional domains from multiple Cpfl orthologs were exchanged. Although one Cpfl chimera has been created (M44- SEQ ID NO:7 in this application), its editing efficiency in E. coll bacteria cells, when compared to a positive control (MAD7 - SEQ ID NO:2 in this application, an engineered Cpfl variant originating from the bacterium Eubacterium rectale) could not be conclusively established, as it varied significantly when the same bacterial loci was targeted with different guide RNA sequences. When tested in yeast and HEK293T cells, M44 also did not work as intended, as it had significantly lower editing efficiency than MAD7. Thus, there remains a need for CRISPR-Cas nucleases and guide RNAs with improved specificity, accuracy, and efficiency, particularly in mammalian cells.SUMMARY OF THE INVENTION
[0011] The present disclosure is based on the discovery that chimeric proteins made by combining Cpfl domains from at least two distinct species result in high nuclease activity and specificity. The present invention replaces domains of the Eubacterium rectale Cpfl (ErCpfl) with domains from other species, generating a highly efficient chimeric nucleic acid-guided nuclease system. E. rectale is also known as Agathobacter rectalis (NCBI). The described nuclease systems unexpectedly rescue nonfunctional guide RNAs.
[0012] In one embodiment, the present disclosure provides a nucleic acid-guided nuclease system including: (a) a polypeptide sequence including a nucleic acid-guided nuclease, wherein the polypeptide sequence has at least 90% sequence identity to: SEQ ID NO:1, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:9, SEQ ID NQ:25, SEQ ID NO:^ SEQ ID NQ:28, SEQ ID NQ:29, SEQ ID NQ:30, or SEQ ID NO:31 ; or (b) a nucleic acid molecule encoding the polypeptide sequence of (a); and (c) an engineered guide nucleic acid including: (i) a region for complexing with the nucleic acid-guided nuclease, and (ii) a region for hybridizingwith a target sequence, wherein the engineered guide nucleic acid includes at least 95% identity to SEQ ID NO: 13.
[0013] In one embodiment, the present disclosure provides a nucleic acid-guided nuclease system including: (a) a polypeptide sequence including a nucleic acid-guided nuclease, wherein the polypeptide sequence has at least 90% sequence identity to: SEQ ID NO:3, SEQ ID NON, SEQ ID NON, SEQ ID NO:6, SEQ ID NO:9, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, or SEQ ID NO:31; and wherein the polypeptide sequence has less than 90% sequence identity to SEQ ID NO:2; or (b) a nucleic acid molecule encoding the polypeptide sequence of (a); and (c)an engineered guide nucleic acid including: (i) a region for complexing with the nucleic acid-guided nuclease; and (ii) a region for hybridizing with a target sequence, wherein the engineered guide nucleic acid includes at least 95% identity to SEQ ID NO: 12. In one aspect, the target sequence encodes a protein. In one aspect, the present disclosure provides a vector including at least one of: (a) a polynucleotide encoding the polypeptide sequence including the nucleic acid-guided nuclease with at least 90% sequence identity to SEQ ID NO:1, SEQ ID NON, SEQ ID NON, SEQ ID NON, SEQ ID NO:6 or SEQ ID NON; and (b) the engineered guide nucleic acid with at least 95% sequence identity to SEQ ID NO: 13.
[0014] In one aspect, the present disclosure provides a vector including at least one of: (a) a polynucleotide encoding the polypeptide sequence with at least 90% sequence identity to SEQ ID NON, SEQ ID NON, SEQ ID NON, SEQ ID NON or SEQ ID NON and less than 90% sequence identity to SEQ ID NON; and (b) the engineered guide nucleic acid with at least 95% sequence identity to SEQ ID NO: 12. In one aspect, the vector is a plasmid or a viral vector.
[0015] In one aspect, the present disclosure provides a delivery system configured to deliver one or more components of the nucleic acid-guided system or the vector including at least one of: (a) a polynucleotide encoding the polypeptide sequence with at least 90% sequence identity to SEQ ID NON, SEQ ID NON, SEQ ID NON, SEQ ID NON or SEQ ID NON and less than 90% sequence identity to SEQ ID NON; and (b) the engineered guide nucleic acid with at least 95% sequence identity to SEQ ID NO: 12.
[0016] In one aspect, the delivery vehicle is selected from a liposome, a particle, an exosome, a microvesicle or a viral vector. In one aspect, the delivery method is selected from electroporation, a gene-gun, calcium phosphate mediated transfer, nucleofection, sonoporation, heat shock, magneto fection, or micro injection.
[0017] In one embodiment, the present disclosure provides a method of modifying a target nucleic acid with a nucleic acid-guided nuclease, wherein the guide sequence directs sequencespecific binding to the target nucleic acid sequence, whereby the target nucleic acid sequence, theexpression of the target nucleic acid, or both are modified. In one aspect, modifying of the target nucleic acid occurs in vitro, ex vivo, or in vivo. In another aspect, modifying the target nucleic acid includes cleaving the target nucleic acid. In an additional aspect, modifying expression of the target nucleic acid includes increasing or decreasing transcription or translation of the target nucleic acid. In a further aspect, the target nucleic acid is in a prokaryotic cell. In an additional aspect, the target nucleic acid is in a eukaryotic cell. In one aspect, the eukaryotic cell is a mammalian cell. In another aspect, the mammalian cell is a human cell.
[0018] In one embodiment, the present disclosure provides an isolated cell including a modified target nucleic acid according to the described methods. In one aspect, modification of the target nucleic acid of interest results in the cell including altered expression of at least one gene product. In another aspect, expression of the at least one gene product is increased. In an additional aspect, expression of the at least one gene product is decreased. In a further aspect, at least one gene product is modified. In one aspect, the present disclosure provides a plant or animal model including one or more cells described in the present disclosure.
[0019] In certain embodiments. The present disclosure provides a nucleic acid-guided nuclease system including: (a) a nucleic acid-guided nuclease including at least 90% amino acid sequence identity to a chimera nuclease based on ErCpfl , wherein the chimera nuclease based on ErCpf 1 includes a nuclease (Nuc) domain substituted with a Nuc domain of Cpfl from another specie, or (b) a nucleic acid molecule encoding the nucleic acid-guided nuclease of (a); and (c) an engineered guide nucleic acid including: (i) a region for complexing with the nucleic acid-guided nuclease, and (ii) a region for hybridizing with a target sequence, wherein the engineered guide nucleic acid includes at least 95% sequence identity to SEQ ID NO: 13. In some aspects the Nuc domain substitution is at the crossover points shown in Figure 9.BRIEF DESCRIPTION OF THE DRAWINGS
[0020] FIG. 1 illustrates a schematic diagram showing the research design for detecting homology directed repairs (HDRs). Two plasmids, one expressing a nuclease and the other a guide RNA expression cassette and donor DNA (containing homology arms and a polynucleotide encoding a reporter flanked by the homology arms), were introduced by electroporation and the reporter was detected by fluorescence or luminescence after 3 days.
[0021] FIG. 2 illustrates a schematic diagram showing plasmid constructs used for HDR detection. The sequence of a representative nuclease expression plasmid is shown in SEQ ID NO: 19. The nuclease expression was driven by the hCMV-IE enhancer and CAG promoter. The nuclease expression plasmid also contained nuclear localization signal sequences at the N- and C-termini (SEQ ID NO:20 and SEQ ID NO:22 respectively). An internal linker (SEQ ID N0:21) was used for tethering the nuclear localization signal peptide to the nuclease. Guide RNA expression was driven by a human U6 promoter. The reporter was designed without an exogenous promoter using a T2A self-cleaving peptide or internal ribosome entry site (IRES) sequence so that it was expressed in conjunction with endogenous gene expression only when inserted at the targeted site (also known as a gene trap).
[0022] FIG. 3 illustrates a schematic diagram of a genome editing design: a fluorescent protein reporter was inserted into the ACTB gene used in Example 1 (SEQ ID NO: 48). The donor vector was designed so that a cDNA encoding glycine 268 to phenylalanine 375 of beta-actin and a cDNA encoding 2A self-cleaving peptide and tdTomato, a fluorescent protein, were inserted inframe downstream of the genomic sequence encoding leucine 267 of beta-actin. Because the donor vector does not have a promoter, tdTomato fluorescence is not observed during episomal expression or random integration, and only when targeted integration occurs is tdTomato fluorescence observed with the expression of the endogenous ACTB gene. In addition, the amino acid sequence encoded in exons 5 and 6 of the ACTB gene (amino acids 268 to 375 of beta-actin) is not translated in this system after the targeted integration, but instead the sequence supplied by the donor vector provides the lost sequence, so that targeted integration by the HDR does not alter the amino acid sequence of the beta-actin protein.
[0023] FIG. 4 illustrates a schematic diagram of Cpfl domains, including the origin of each chimeric Cpfl domain used in Example 1 , and the results of qualitative HDR detection using a fluorescent reporter.
[0024] FIGS. 5A-5B illustrate genome editing induced pluripotent stem cells (iPSCs) in human. FIG. 5A illustrates a microscopic image showing fluorescent (mCherry) reporter-positive cells detected in Example 3. FIG. 5B illustrates a schematic diagram showing the system used for single-cell cloning of reporter-positive cells, and an image of PCR results confirming the insertion of the donor DNA-derived sequence at the target sequence (SEQ ID NO: 49).
[0025] FIG. 6 illustrates a schematic diagram showing crossover points of the present disclosure.
[0026] FIG. 7 illustrates a schematic diagram showing an effector nuclease and associated guide RNA molecule.
[0027] FIG. 8 illustrates the template alignment of SEQ ID NOG with templates SEQ ID NO: 1.
[0028] FIG. 9 illustrates the template alignment of SEQ ID NOG with templates SEQ ID NO:1.
[0029] FIG. 10 illustrates the template alignment of SEQ ID NO:5 with templates SEQ ID NO: 1.
[0030] FIG. 11 illustrates the template alignment of SEQ ID NO:6 with templates SEQ ID NO: 1.
[0031] FIG. 12 illustrates the template alignment of SEQ ID NO:7 with templates SEQ ID NO: 1.
[0032] FIG. 13 illustrates the template alignment of SEQ ID NO:8 with templates SEQ ID NO: 1.
[0033] FIG. 14 illustrates the template alignment of SEQ ID NO:9 with templates SEQ ID NO: 1.
[0034] FIG. 15 illustrates the template alignment of SEQ ID NO: 10 with templates SEQ ID NO: 1.
[0035] FIG. 16 illustrates the template alignment of SEQ ID NO: 11 with templates SEQ ID NO: 1.
[0036] FIG. 17 illustrates a sequence showing detailed sequence of a guide expression and tdTomato reporter donor plasmid: (SEQ ID NO: 14) and (SEQ ID Nos: 50-53).
[0037] FIG. 18 illustrates a sequence showing detailed sequence of a guide expression (scaffold SEQ ID NO:13) and tdTomato reporter donor plasmid: (SEQ ID NO:15) and (SEQ ID Nos: 51- 54).
[0038] FIG. 19 illustrates a sequence showing detailed sequence of a guide expression (scaffold SEQ ID NO:12) and luciferase reporter donor plasmid: (SEQ ID NO:16) and (SEQ ID Nos: 51, 53 & 55-56).
[0039] FIG. 20 illustrates a sequence showing detailed sequence of a guide expression (scaffold SEQ ID NO: 13) and luciferase reporter donor plasmid: (SEQ ID NO: 17) and (SEQ ID Nos: 51, 53 & 56-57) and (SEQ ID Nos: 53 & 58-60).
[0040] FIG. 21 illustrates a sequence showing detailed sequence of a guide expression (scaffold SEQ ID NO:12) and mCherry-tCD19 reporter donor plasmid: (SEQ ID NO:18).
[0041] FIG. 22 illustrates a diagram showing plasmid construct for modified HDR quantification system.
[0042] FIG. 23 illustrates a schematic diagram showing the design of a nuclease chimera that exchanges the Nuc domain based on ErCpfl (SEQ ID NO: 66).
[0043] FIGS. 24A-24B illustrate improvement or decrease in HDR efficiency due to Nuc domain replacement. FIG. 24A illustrates a graph showing HDR activity quantified by luciferaseactivity relative to wild-type ErCpfl . FIG. 24B illustrates a table showing the origin of the Nuc domain used.
[0044] FIG. 25 illustrates a sequence showing detailed sequence of the expression plasmid for the EFla promoter-driven nuclease (SEQ ID NO:32) and (SEQ ID Nos: 61-63).
[0045] The figures described herein are for illustrative purposes only and are not drawn to scale.DETAILED DESCRIPTION OF THE INVENTION
[0046] The present disclosure is based on the discovery that chimeric proteins made by combining Cpfl domains from at least two distinct species result in high nuclease activity and specificity. The present invention replaces domains of the Eubacterium rectale Cpfl (ErCpfl) with domains from other species, generating a highly efficient chimeric nucleic acid-guided nuclease system.
[0047] Before the present systems and methods are described, it is to be understood that this invention is not limited to the particular systems, methods, and experimental conditions described, as such systems, methods, and conditions may vary. It is also to be understood that the terminology used herein is for the purposes of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only in the appended claims.
[0048] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Thus, for example, references to “the method” include one or more methods, and / or steps of the type described herein which will become apparent to those persons skilled in the art upon reading this disclosure and so forth.
[0049] As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items.
[0050] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the invention, it will be understood that modifications and variations are encompassed within the spirit and scope of the instant disclosure. The preferred methods and materials are now described.
[0051] The present disclosure provides chimeric nucleic acid-guided nucleases and their methods of use. Nucleic acid-guided nucleases of the present disclosure include systems including proteins, polynucleotides, vectors, and cells used in methods to target, edit, or modify a nucleic acid in a prokaryotic or eukaryotic cell, in vitro, ex vivo, or in vivo. In one embodiment, the present disclosure includes systems and methods for targeting multiple polynucleotide positions within a prokaryotic or eukaryotic genome. In one aspect, the present disclosure describes systems and methods for genome engineering of prokaryotic or eukaryotic cells. In another aspect, the present disclosure describes systems and methods for efficient genome engineering including nucleic acid-guided nucleases. In an additional aspect, the present disclosure describes chimeric nucleic acid-guided nucleases. In a further aspect, the nucleic acid-guided nuclease is a CRISPR- Cas system.
[0052] The term “nucleic acid-guided nuclease,” as used herein, refers to a protein-nucleic acid complex with nuclease activity. In one aspect, a nucleic acid-guided nuclease is a nucleic acid- guided system. In one aspect, the protein-nucleic acid nuclease is a CRISPR-Cas system. In one embodiment, a CRISPR-Cas system includes an effector complex and a guide nucleic acid. In another aspect, the effector complex includes at least one, at least two, at least three, or at least four proteins. In an additional aspect, the guide nucleic acid includes at least one, at least two, or at least three nucleic acids.
[0053] In one embodiment, the nucleic acid-guided nuclease of the present disclosure binds to a target nucleic acid sequence and cleaves, nicks, or otherwise modifies the target sequence. A nucleic acid of interest is a target gene on a plasmid, in a prokaryotic or eukaryotic genome, or any other location on a polynucleotide for which mutation is desired. In one aspect, the nucleic acid-guided nuclease does not cleave, nick, or otherwise modify the target sequence. In another aspect, the nucleic acid-guided nuclease alters or manipulates expression of one or more genes in prokaryotic or eukaryotic cells. In a further aspect, the nucleic acid-guided nuclease is fused to a transcriptional repressor to regulate translation of the target sequence. In an additional aspect, the nucleic acid-guided nuclease is fused to a histone methylase to regulate translation of the target sequence. In one aspect, the nucleic acid-guided nuclease is fused to a detectable marker so that a location within a cell is visualized. In an additional aspect, the nucleic acid-guided nuclease is fused to a detectable marker so that a location within an organism is visualized. In a further aspect, the nucleic acid-guided nuclease is fused to a detectable marker so that a location of a polynucleotide sequence is visualized.
[0054] In one embodiment, the target sequence is a nucleic acid. In one aspect, the target sequence is a double-stranded nucleic acid. In another aspect, the target sequence is a single-stranded nucleic acid. In various aspects, the nucleic acid-guided nuclease cleaves a doublestranded target sequence. In a further aspect, the nucleic acid-guided nuclease cleaves a singlestranded nucleic acid. In another aspect, the nucleic acid-guided nuclease cleaves one strand of a double-stranded target sequence. In an additional aspect, the nucleic acid-guided nuclease cleaves a double-stranded target sequence to generate a blunt end. In another aspect, the nucleic acid- guided nuclease cleaves a double-stranded target sequence in a staggered manner to generate sticky ends. In various aspects, the staggered cut has a 5' overhang. In a further aspect, the staggered cut has a 5' overhang of 1, 2, 3, 4 or 5 nucleotides. In a one aspects, the CRISPR-Cas cleaves one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more nucleotides from the first or last nucleotide of the target nucleotide sequence or its complementary sequence.
[0055] In one embodiment, the nucleic acid-guided nuclease includes an effector complex and at least one guide RNA. In one aspect, the effector complex includes a recognition region, a nuclease region, and a connection region. In another aspect, the recognition region includes a recognition lobe domain (Rec). In an additional aspect, the nuclease region includes a nuclease lobe (Nuc). In a further aspect, the Nuc includes a RuvC nuclease domain, a bridge helix domain (BH), and a Nuc domain. In an additional aspect, the connection region includes a Wedge domain (WED) and a PAM-interacting domain (PI). In another aspect, the effector complex includes at least one of (a) a recognition lobe domain (Rec), (b) a RuvC domain, (c) a bridge helix domain (BH), (d) a Nuc domain, (e) a Wedge domain (WED), and (f) a PAM-interacting domain (PI). In one aspect, the effector complex includes at least one, at least two, at least three, at least four, at least five, or at least six domains. In an additional aspect, the effector complex includes at least one domain from a first species and at least one domain from a second species. In another aspect, the effector complex includes at least one domain from a first species, at least one domain from a second species, and at least one domain from a third species.
[0056] In one embodiment, the nucleic acid-guided nuclease is a CRISPR-Cas. In one aspect, the CRISPR-Cas includes an effector complex and a guide RNA (FIG. 7). In an additional aspect, the guide RNA includes a scaffold region (also called a constant region) and a guide region. In another aspect, the scaffold region interacts with the effector complex, whereas the guide region is at least partially complimentary to the 3' sequence downstream of the PAM sequence in the target polynucleotide sequence. In a further aspect, the guide region includes a spacer sequence at least partially complementary to the target polynucleotide sequence. In an additional aspect, the guide RNA includes, from 5' to 3', an optional 5' sequence called a “tail”, a modulator stem sequence, a loop, a targeter stem sequence complementary to the modulator stem sequence, and aspacer sequence at least partially complementary to a target sequence (FIG. 7). In another aspect, the spacer sequence has at least 50%, at least 60%, at least 70%, at least 80%, at least 90, or at least 95% homology to a target sequence.
[0057] Design requirements of guide RNAs for nucleic acid-guided nucleases such as CRISPR- Cas include modification of GC content, length of the guide RNA, and mismatches between the spacer and the target site. Mismatches between the spacer and the target nucleotide sequence may affect the stability and / or conformation of the CRISPR-Cas complex, leading to modified binding of the CRISPR-Cas complex to the target polynucleotide sequence, such as increased binding to the target locus, decreased binding to the target locus, increased binding to off-target loci, or decreased binding to off-target loci.
[0058] In one embodiment, the PAM sequence is TTTN, where N is A, C, G, or T. In one aspect, the PAM sequence is TTTT. In another aspect, the PAM sequence is TTTC. In an additional aspect, the PAM sequence is TTTA. Optional PAM sequences include TTTA, TTTC, TTTG, TTTT, CTTA, CTTC, CTTG, or CTTT. In one aspect, the PAM in the non-target strand of the target polynucleotide sequence binds the effector complex.
[0059] In one embodiment, a PAM or PAM-like motif directs binding of a CRISPR-Cas to a target locus. In a further aspect, the sequence and length requirements for the PAM differ depending on the effector complex used, the guide RNA used, or a combination thereof. In another aspect, the PAM sequence is from 2-5 base pairs (bp) in length and is adjacent to the target polynucleotide strand. In an additional aspect, the PAM sequence is on a different strand of target DNA than the target nucleotide sequence.
[0060] In one embodiment, an CRISPR-Cas includes a modification that alters the targeting specificity and modifies the targeting range. In one aspect, an engineered CRISPR-Cas is designed to have increased target specificity as well as accommodating modifications in PAM recognition. In another aspect, target specificity of an engineered CRISPR-Cas is modified by including mutations that alter PAM specificity. In an additional aspect, a mutation that modifies target specificity is in the PI domain. In a further aspect, an engineered CRISPR-Cas is modified by combining mutations that modify target specificity with groove mutations that increase or decrease specificity for the on-target locus versus off-target loci. In another aspect, a CRISPR- Cas modification counters loss of specificity resulting from alteration of PAM recognition. In an additional aspect, a CRISPR-Cas modification enhances gain of specificity resulting from alteration of PAM recognition. In one aspect, a CRISPR-Cas modification counters gain of specificity resulting from alteration of PAM recognition. In another aspect, a CRISPR-Cas modification enhances loss of specificity resulting from alteration of PAM recognition.
[0061] In one embodiment, the CRISPR-Cas system edits a polynucleotide via homology- directed repair (HDR). In one aspect, the CRISPR-Cas system includes a donor sequence used to repair a target sequence via homology-directed repair. In an additional aspect, the donor sequence is a non-homologous sequence flanked by two regions homologous to the target sequence, called homology arms, such that homology-directed repair between the target DNA and the two flanking sequences results in insertion of the non-homologous sequence at the target DNA region. In another aspect, the donor sequence includes a non-homologous sequence from 10-100 nucleotides, from 50-500 nucleotides, from 100-1,000 nucleotides, from 200-2,000 nucleotides, or from 500-5,000 nucleotides in length positioned between two homology arms.
[0062] Homology arms are added to a donor sequence to allow incorporation of a non- homologous sequence into the desired genomic location via homologous recombination, or homology-driven repair. Homology arms are added to a donor sequence by synthesis, in vitro assembly, PCR, or other known methods in the art. In a further aspect, a homology arm is added to both ends of a barcode, recorder sequence, and / or editing sequence, thereby flanking the sequence with two distinct homology arms, for example, a 5' homology arm and a 3' homology arm.
[0063] The present disclosure is based on the finding that domains of a nucleic acid-guided nuclease are interchangeable with domains of a nucleic acid-guided nuclease from distinct species, and that substitution of the WED-I domain and the RuvC-III domain with domains from distinct species rescues function of non-functional guide RNAs. An engineered nucleic acid-guided nuclease including domains from at least one, at least two, or at least three distinct species is a chimeric nucleic acid-guided nuclease. In one embodiment, a guide RNA that is not functional in one CRISPR-Cas system is functional in a chimeric CRISPR-Cas system. In one embodiment, combinations of domains from one, two or three species produce a functional chimeric CRISPR- Cas nuclease.
[0064] Cpfl is a putative Class 2, Type V CRISPR-Cas nuclease. Cpfl has several advantages over Cas9, including functionality with a single RNA instead of two RNA strands, the ability to cleave double-stranded DNA in a staggered manner to produce sticky ends, and the ability cleave RNA and DNA. One feature of this invention is that when making chimeric proteins with effector complex domains from multiple distinct species, high activity is obtained by combining wild-type domains from two or more different species.
[0065] Cpfl is composed of the domains WED, REC, PI, RuvC, BH, and Nuc as shown in the upper part of FIG. 4. SEQ ID NO: 1 is a Cpfl nuclease from Eubacterium rectale (ErCpfl). SEQ ID NO:3 is a Cpfl nuclease with the RECI domain from ErCpfl substituted with a RECI domainfrom Thiomicrospira sp. XS5 (TxCpfl) (FIG. 8). SEQ ID NO:4 is a Cpfl nuclease with the Nuc domain from ErCpfl substituted with a Nuc domain from Succinivibrio dextrinosolvens (ScCpfl) (FIG. 9). SEQ ID NO:5 is an ErCpfl nuclease with the RECI domain substituted with a RECI domain from TxCpfl, and a Nuc domain substituted with a Nuc domain from ScCpfl (FIG. 10). SEQ ID NO:6 is an ErCpfl substituted with the RECI domain from TxCpfl (FIG. 11). SEQ ID NO:7 is a ErCpfl nuclease with the WED-I and RECI domains substituted with those of TxCpfl (FIG. 12). SEQ ID NO:8 is an ErCpfl nuclease with part of the RuvC-I, and the BH, RuvC-II, Nuc, and RuvC-III domains substituted with those of ScCpfl (FIG. 13). SEQ ID NO:9 is a Cpfl nuclease with part of the RuvC-I, and the BH, RuvC-II, Nuc, and RuvC-III domains substituted with those of Acidaminococcus sp. BV3L6 (FIG. 14).
[0066] The chimeric protein called M44 (SEQ ID NO:7) includes TxCpflfor 317 amino acids from the N-terminal side and ErCpfl for 966 amino acids from the C-terminal side. In contrast, SEQ ID NO:3 and SEQ ID NO:6, which include three domains from the N-terminal side, including ErCpfl, TxCpfl, and ErCpfl again, had surprisingly higher nuclease activity than M44 (SEQ ID NO:7, FIG. 4).
[0067] In one embodiment, an engineered effector complex includes at least one nuclear localization signal (NLS) motif. In one aspect, an engineered effector complex comprises at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 NLS motifs. In another aspect, the NLS motif includes: the NLS of SV40 large T-antigen, having the amino acid sequence of PKKKRKV (SEQ ID NO:20); the NLS from nucleoplasmin, e.g., the nucleoplasmin bipartite NLS having the amino acid sequence of KRPAATKKAGQAKKKK (SEQ ID NO:22); the c-myc NLS, having the amino acid sequence of PAAKRVKLD (SEQ ID NO:42) or RQRRNELKRSP (SEQ ID NO:43); the hRNPAl M9 NLS, having the amino acid sequence of NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO:44); the importin- a IBB domain NLS, having the amino acid sequence of RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO:45); the myoma T protein NLS, having the amino acid sequence of VSRKRPRP (SEQ ID NO:46) or PPKKARED (SEQ ID NO:47); the human p53 NLS, having the amino acid sequence of PQPKKKPL (SEQ ID NO:33); the mouse c-abl IV NLS, having the amino acid sequence of SALIKKKKKMAP (SEQ ID NO:34); the influenza virus NS1 NLS, having the amino acid sequence of DRLRR (SEQ ID NO:35) or PKQKKRK (SEQ ID NO:36); the hepatitis virus 6 antigen NLS, having the amino acid sequence of RKLKKKIKKL (SEQ ID NO:37); the mouse Mxl protein NLS, having the amino acid sequence of REKKKFLKRR (SEQ ID NO:38); the human poly(ADP-ribose) polymerase NLS, having the amino acid sequence ofKRKGDEVDGVDEVAKKKSKK (SEQ ID NO:39); the human glucocorticoid receptor NLS, having the amino acid sequence of RKCLQAGMNLEARKTKK (SEQ ID NO:40), and synthetic NLS motifs such as PAAKKKKLD (SEQ ID NO:41).
[0068] In one embodiment, the at least one NLS motif drives accumulation of the CRISPR-Cas in the nucleus of a eukaryotic cell. In one aspect, the amount of nuclear localization is correlated with the number of NLS motifs in the effector complex, the particular NLS motif used, the position of the NLS motif, or a combination thereof. In another aspect, an engineered effector complex includes at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least9, or at least 10 NLS motifs at or near the N-terminus. In an additional aspect, an engineered effector complex includes at least one NLS motif within about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, or more amino acids along the polypeptide chain from the N-terminus. In one aspect, an engineered effector complex includes at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 NLS motif(s) at or near the C -terminus. In another aspect, an engineered effector complex includes at least 1 NLS motif within about 1, 2, 3, 4, 5,10, 15, 20, 25, 30, 40, 50, or more amino acids along the polypeptide chain from the C-terminus. In a further aspect, an engineered effector complex includes at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 NLS motif(s) at or near the C- terminus and at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 NLS motif(s) at or near the N-terminus. In an additional aspect, the engineered effector complex includes one, two, or three NLS motifs at or near the C-terminus. In one aspect, the engineered Cas protein includes one NLS motif at or near the N-terminus and one, two, or three NLS motifs at or near the C-terminus. In another aspect, the engineered Cas protein comprises a nucleoplasmin NLS at or near the C-terminus.
[0069] In one embodiment, a protective nucleotide sequence is located at the 5' or 3' end of the guide nucleic acid. In one aspect, the guide nucleic acid includes a protective nucleotide sequence at the 5' end, at the 3' end, or at both ends. In another aspect, the guide nucleic acid includes a protective nucleotide sequence linked through a nucleotide linker.
[0070] In one embodiment, various nucleotide sequences are present in the 5' portion of a guide nucleic acid, including but not limited to a donor template-recruiting sequence, an editing enhancer sequence, a protective nucleotide sequence, and a linker connecting a sequence to a 5' tail, or to the 5' portion of the guide nucleic acid.
[0071] In one embodiment, the disclosure provides a vector including polynucleotides coding for an effector complex, a guide RNA, and optionally a donor sequence. In one aspect, the disclosure provides a vector including a polynucleotide coding for an effector complex. In anotheraspect, the disclosure provides a vector including a polynucleotide coding for a guide RNA. In an additional aspect, the disclosure provides a vector including a polynucleotide coding for a donor sequence. In a further aspect, the effector complex and guide RNA are expressed in a single vector. In an additional aspect, the effector complex is expressed in a first vector and the guide RNA is expressed in a second vector. In another aspect, the donor sequence and the effector complex are expressed in a first vector. In one aspect, the donor sequence and the guide RNA are expressed in a first vector. In one aspect, the donor sequence is conjugated covalently to a guide nucleic acid. In an additional aspect, the covalent linkages are described in U.S. Patent No. 9,982,278 and Savic et al. (2018) ELIFE 7:e33761. In another aspect, the donor sequence is covalently linked to the 5' end of the guide nucleic acid through an inter-nucleotide bond. In a further aspect, the donor sequence is covalently linked to a the 5' end of the guide nucleic acid through a linker.
[0072] In one aspect, the present disclosure provides a vector including at least one of: (a) a polynucleotide encoding the polypeptide sequence with at least 90% sequence identity to SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6 or SEQ ID NO:9 and less than 90% sequence identity to SEQ ID NO:2; and (b) the engineered guide nucleic acid with at least 95% sequence identity to SEQ ID NO: 12. In one aspect, the vector is a plasmid or a viral vector.
[0073] As used herein the term “engineered guide nucleic acid” refers to a synthetically modified nucleic acid sequence designed to specifically target a particular location on a genome. An engineered guide nucleic acid is a custom-made guide RNA with altered features to enhance its targeting ability or functionality within a specific biological context.
[0074] In one embodiment, the guide nucleic acid comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 modified nucleotides or inter-nucleotide linkages. In one aspect, the modified nucleic acids direct the CRISPR-Cas to the target nucleotide sequence. In another aspect, the modified nucleotide or inter-nucleotide linkage is selected based on the functionality of the nucleotide or inter-nucleotide linkage at a position on the guide nucleic acid. In an additional aspect, a specificity-enhancing modification is suitable for a nucleotide or inter-nucleotide linkage in the spacer sequence, the targeter stem sequence, or the modulator stem sequence. In a further aspect, a stability-enhancing modification includes one or more terminal nucleotides or inter- nucleotide linkages in the targeter nucleic acid or the modulator nucleic acid. In one aspect, at least 1, at least 2, at least 3, at least 4, or at least 5 terminal nucleotides or inter-nucleotide linkages and / or at least 1, at least 2, at least 3, at least 4, or at least 5 terminal nucleotides or inter-nucleotide linkages are modified.
[0075] Positions appropriate for modifications of nucleotides or inter-nucleotide linkages are described in U.S. Patent Nos. 10,900,034 and 10,767,175. In one aspect, when the targeter stem or modulator section of the guide nucleic acid is a combination of DNA and RNA, the nucleic acid as a whole is considered as an RNA, and the DNA nucleotide is considered a modification of the RNA, including a 2'-H modification of a ribose and a modification of the nucleobase.
[0076] In one embodiment, polynucleotides expressing the nucleic acid-guided nucleases of the present invention, included in a vector, are introduced into a cell, thus expressing the nucleic acid- guided nucleases within the cell. A variety of methods are suitable for introduction of nucleic acid into a cell, including viral and non- viral mediated delivery methods. In one embodiment, the non- viral mediated delivery methods include, but are not limited to, electroporation, calcium phosphate mediated transfer, nucleofection, sonoporation, heat shock, magnetofection, liposome mediated transfer, micro injection, microprojectile mediated transfer (nanoparticles), a gene-gun, cationic polymer mediated transfer (DEAE-dextran, polyethylenimine, polyethylene glycol (PEG) and the like) or cell fusion, and delivery vehicles selected from a liposome, a particle, an exosome, and a microvesicle.
[0077] In one embodiment, the viral vector is a recombinant expression vector. In one aspect, the recombinant expression vector is a viral construct. In another aspect, the viral vector is a recombinant adeno-associated virus construct, a recombinant adenoviral construct, a recombinant lentiviral construct, or a recombinant retroviral construct.
[0078] In one embodiment, expression vectors include, but are not limited to, viral vectors, for example, viral vectors based on vaccinia virus; poliovirus, adenovirus, adeno-associated virus, herpes simplex virus, human immunodeficiency virus, a retroviral vector, a vector derived from retroviruses such as Rous sarcoma virus, Harvey sarcoma virus, avian leukosis virus, a lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, or mammary tumor virus.
[0079] Numerous suitable expression vectors are known to those of skill in the art. In one aspect, the following vectors are used: pXTl, pSG5, pSVK3, pBPV, pMSG, and pSVLSV40. In one aspect, any vector that is compatible with the host cell is used. In one aspect, the present disclosure provides a delivery system configured to deliver one or more components of the nucleic acid- guided system or the vector including at least one of: (a) a polynucleotide encoding the polypeptide sequence with at least 90% sequence identity to SEQ ID NOG, SEQ ID NO:4, SEQ ID NOG, SEQ ID NO:6 or SEQ ID NO:9 and less than 90% sequence identity to SEQ ID NOG; and (b) the engineered guide nucleic acid with at least 95% sequence identity to SEQ ID NO: 12.
[0080] In one aspect, the delivery vehicle is selected from a liposome, a particle, an exosome, a microvesicle or a viral vector. In one aspect, the delivery method is selected from electroporation,a gene-gun, calcium phosphate mediated transfer, nucleofection, sonoporation, heat shock, magneto fection, or micro injection.
[0081] In one embodiment, the present disclosure provides a nucleic acid-guided nuclease system including: (a) a polypeptide sequence including a nucleic acid-guided nuclease, wherein the polypeptide sequence has at least 90% sequence identity to: SEQ ID NO:1, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6 or SEQ ID NO:9; or (b) a nucleic acid molecule encoding the polypeptide sequence of (a); and (c) an engineered guide nucleic acid including: (i) a region for complexing with the nucleic acid-guided nuclease, and (ii) a region for hybridizing with a target sequence, wherein the engineered guide nucleic acid includes at least 95% identity to SEQ ID NO: 13.
[0082] The “region for complexing with the nucleic acid-guided nuclease” or “recognition lobe” or “REC domain” specifically binds to the guide RNA, allowing the nuclease to target the complementary DNA sequence based on the guide’s sequence information.
[0083] In one embodiment, the present disclosure provides a nucleic acid-guided nuclease system including: (a) a polypeptide sequence including a nucleic acid-guided nuclease, wherein the polypeptide sequence has at least 90% sequence identity to: SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6 or SEQ ID NO:9; and wherein the polypeptide sequence has less than 90% sequence identity to SEQ ID NO:2; or (b) a nucleic acid molecule encoding the polypeptide sequence of (a); and (c)an engineered guide nucleic acid including: (i) a region for complexing with the nucleic acid-guided nuclease; and (ii) a region for hybridizing with a target sequence, wherein the engineered guide nucleic acid includes at least 95% identity to SEQ ID NO: 12. In one aspect, the target sequence encodes a protein. In one aspect, the present disclosure provides a vector including at least one of: (a) a polynucleotide encoding the polypeptide sequence including the nucleic acid-guided nuclease with at least 90% sequence identity to SEQ ID NO:1, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6 or SEQ ID NO:9; and (b) the engineered guide nucleic acid with at least 95% sequence identity to SEQ ID NO: 13.
[0084] The present disclosure describes systems and methods to modify the function of chimeric non-naturally occurring nucleic acid-guided nucleases by substituting at least one domain of an effector complex from a first species of bacteria with at least one domain of an effector complex from a second species of bacteria. In one embodiment, domains from at least one, at least two, at least three, at least four, or at least five species of bacteria are combined in a single effector complex. In one embodiment, the domains of an effector complex include at least one of WED-I, RECI, REC2, WED-II, PI, WED-III, RuvC-I, BH, RuvC-II, Nuc, and RuvC-III. In one aspect, the effector complex is from a first species and the WED-I domain is from a second species. Inanother aspect, the effector complex is from a first species and the RECI domain is from a second species. In an additional aspect, the effector complex is from a first species and the REC2 domain is from a second species. In a further aspect, the effector complex is from a first species and the WED-II domain is from a second species. In one aspect, the effector complex is from a first species and the PI domain is from a second species. In another aspect, the effector complex is from a first species and the WED-III domain is from a second species. In an additional aspect, the effector complex is from a first species and the RuvC-I domain is from a second species. In a further aspect, the effector complex is from a first species and the BH domain is from a second species. In another aspect, the effector complex is from a first species and the RuvC-II domain is from a second species. In one aspect, the effector complex is from a first species and the Nuc domain is from a second species. In an additional aspect, the effector complex is from a first species and the RuvC-III domain is from a second species.
[0085] In one embodiment, the effector complex is from a first species and at least one domain is substituted from a second species. In one aspect, the at least one domain is WED-I, RECI, REC2, WED-II, PI, WED-III, RuvC-I, BH, RuvC-II, Nuc, or RuvC-III. In another aspect, the effector complex is from a first species, at least one domain is substituted from a second species and at least one domain is substituted from a third species. In an additional aspect, the at least one domain is WED-I, RECI, REC2, WED-II, PI, WED-III, RuvC-I, BH, RuvC-II, Nuc, or RuvC-III. In a further aspect, the effector complex is from a first species and at least one domain is substituted from a second species, at least one domain is substituted from a third species, and at least one domain is substituted from a fourth species. In another aspect, the at least one domain is WED-I, RECI, REC2, WED-II, PI, WED-III, RuvC-I, BH, RuvC-II, Nuc, or RuvC-III.
[0086] In one embodiment, the effector complex includes crossover points at the beginning of at least one of WED-I, RECI, REC2, WED-II, PI, WED-III, RuvC-I, BH, RuvC-II, Nuc, or RuvC- III. In one aspect, the effector complex includes crossover points at the beginning of REC 1 , REC2, Nuc, and RuvC-III (see FIG. 6). In an additional aspect, the effector complex includes crossover points at the beginning of REC 1 and RuvC-III.
[0087] Examples of non-naturally occurring nucleic acid sequences which are described herein include nucleic acid sequences optimized for expression in multicellular eukaryotes, and plasmids including nucleic acid sequences operably linked to a heterologous promoter or nuclear localization signal or other heterologous elements, e.g., SEQ ID NO: 14-19. In one aspect, the nucleic acid sequences are operably linked to proteins generated from engineered or codon- optimized nucleic acid sequences (e.g., SEQ ID NO: 1-11). In one aspect, the nucleic acid sequences are operably linked to engineered guide nucleic acids including SEQ ID NO: 12 or SEQID NO: 13. The non-naturally occurring nucleic acid sequences described herein are amplified, cloned, assembled, synthesized, and generated from synthesized oligonucleotides or dNTPs, or otherwise obtained using methods known by those skilled in the art.
[0088] In one embodiment, the polynucleotide sequences described herein are optimized for expression in prokaryotes or eukaryotes. In one aspect, the polynucleotide sequences are optimized for expression in prokaryotes. In another aspect, the polynucleotide sequences are optimized for expression in eukaryotes. In an additional aspect, the polynucleotide sequences are optimized for expression in mammalian cells. In a further aspect, the polynucleotide sequences are optimized for expression in a human cell. In another aspect, the polynucleoide sequences are optimized for expression in cells from a human or other animal, including rodents, ungulates, or mammals, for example, horses, cattle, sheep, pigs, goats, llamas, camels, dogs, cats, birds, ferrets, rabbits, squirrels, mice, rats, or ferrets.
[0089] In one embodiment, the invention describes an isolated cell including a modified target nucleic acid according to the described systems and methods. In one aspect, modification of the target nucleic acid results in altered expression of at least one gene product in the cell. In an additional aspect, expression of the at least one gene product is increased. In another aspect, expression of the at least one gene product is decreased. In one aspect, expression of the at least one gene product is silenced. In another aspect, at least one gene product is modified. In an additional aspect, the present disclosure includes a plant or animal model including one or more cells described in the present disclosure. In certain other aspects, the animal models used for experimental trials include, but are not limited to, mouse, rat, cat, chimpanzee, dog, guinea pig and hamster. In still other aspects, the plant models used for experimental trials include, but are not limited to maize, rice, petunia, Arabidopsis, tomato, and snapdragon.
[0090] In one embodiment, a CRISPR-Cas system is expressed in a prokaryotic or eukaryotic cell. In another embodiment, a CRISPR-Cas nuclease is expressed in a prokaryotic or eukaryotic cell. In one aspect, the eukaryotic cell is a plant, insect, or animal cell. In another aspect, the eukaryotic cell is a mammalian cell. In a further aspect, the mammalian cell is a human cell. In an additional aspect, the human cell is an immune cell. In one aspect, the immune cell is a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, or a lymphocyte. In another aspect, the immune cell is a T cell. In a further aspect, the T cell is a CAR- T cell. In an additional aspect, the human cell is a stem cell. In another aspect, the stem cell is a human pluripotent, multipotent stem cell, embryonic stem cell, induced pluripotent stem cell, CD34+ stem cell, or hematopoietic stem cell. In one aspect, the human cell is allogeneic; in otherwords, a cell that provokes little or no immune response when introduced into an allogeneic host and produces little or no graft versus host response.
[0091] In one embodiment, the immune cell is a T cell. In one aspect, the T cell is a cultured T cell, a primary T cell, a T cell from a cultured T cell line, or a T cell obtained from a mammal. In a further aspect, the T cell is obtained from a subject to be treated. In another aspect, the T cell line is Jurkat or SupTi. In an additional aspect, the T cell is obtained from a tissue or fluid. In one aspect, the tissue or fluid is blood, bone marrow, lymph node, the thymus, or other tissue or fluid. In one aspect, the T cell is enriched or purified. In another aspect, the T cell is any type of T cell and is of any developmental stage. In an additional aspect, the T cell is a CD4+ / CD8+ double positive T cell, a CD4+ helper T cells such as a Thl or Th2 cell, a CD8+ T cell such as a cytotoxic T cell, a tumor infiltrating lymphocyte (TIL), a memory T cell such as a central memory T cell or effector memory T cell, a regulatory T cell, a naive T cell, or another kind of T cell.
[0092] In one embodiment, an immune cell or a T cell is engineered to express an exogenous gene. In one aspect, the CRISPR-Cas system disclosed herein cleaves DNA at a gene locus, integrating an exogenous gene via site-specific genetic modification at the gene locus by HDR.
[0093] In one embodiment, an immune cell or a T cell expresses a chimeric antigen receptor (CAR). In a further aspect, the T cell includes an exogenous nucleotide sequence encoding a CAR. As used herein, the term “chimeric antigen receptor” or “CAR” refers to any artificial receptor including an antigen-specific binding moiety and one or more signaling proteins derived from an immune receptor. As used herein, a T cell expressing a chimeric antigen receptor is referred to as a CAR T cell. In one aspect, a CAR comprises a single chain fragment variable (scFv) of an antibody specific for an antigen coupled via hinge and transmembrane regions to a cytoplasmic domain of a T cell signaling molecule. In another aspect, the scFv is coupled to a T cell costimulatory domain. In an additional aspect, the T cell costimulatory domain is from CD28, CD 137, 0X40, ICOS, or CD27. In a further aspect, the scFv is coupled to a T cell costimulatory domain in tandem with a T cell triggering domain. In another aspect, the T cell triggering domain is from CD3Q. In an additional aspect, a CAR T cells includes a CD 19 targeted CTL019 cell, a 19-28z cell, or a KTE-C19 cell. Additional CAR T cells are described in U.S. Patent Nos. 7,446,190, 8,399,645, 8,906,682, 9,181,527, 9,272,002, 9,266,960, 10,253,086, 10640569, and 10,808,035, and International (PCT) Publication Nos. WO 2013 / 142034, WO 2015 / 120180, WO 2015 / 188141, WO 2016 / 120220, and WO 2017 / 040945.
[0094] In one embodiment, an immune cell binds an antigen through an endogenous T cell receptor (TCR). In one aspect, the immune cell is a T cell. In another aspect, the antigen is a cancer antigen. In an additional aspect, the immune cell is engineered to express an exogenous TCR. Ina further aspect, the immune cell is engineered to express an exogenous naturally occurring TCR or an exogenous engineered TCR.
[0095] In one embodiment, the nucleic acid-guided nuclease includes a domain from at least one organism from a genus including but not limited to Thiomicrospira, Succinivibrio, Candidatus, Porphyromonas, Acidaminococcus, Acidomonococcus, Prevotella, Smithella, Moraxella, Synergistes, Francisella, Leptospira, Catenibacterium, Kandleria, Clostridium, Dorea, Coprococcus, Enterococcus, Fructobacillus, Weissella, Pediococcus, Corynebacter, Sutterella, Legionella, Treponema, Roseburia, Filifactor, Eubacterium, Streptococcus, Lactobacillus, Mycoplasma, Bacteroides, Flaviivola, Flavobacterium, Sphaerochaeta, Azospirillum, Gluconacetobacter, Neisseria, Roseburia, Parvibaculum, Staphylococcus, Nitratifractor, Mycoplasma, Alicyclobacillus, Brevibacilus, Bacillus, Bacteroidetes, Brevibacilus, Carnobacterium, Clostridiaridium, Clostridium, Desulfonatronum, Desulfovibrio, Helcococcus, Leptotrichia, Listeria, Methanomethyophilus, Methylobacterium, Opitutaceae, Paludibacter, Rhodobacter, Sphaerochaeta, Tuberibacillus, Oleiphilus, Omnitrophica, Parcubacteria, and Campylobacter. Suitable nucleic acid-guided nucleases are described in at least US Patent Application Publication No. US20160208243 filed Dec. 18, 2015, US Application Publication No. US20140068797 filed Mar. 15, 2013, U.S. Pat. No. 8,697,359 filed Oct. 15, 2013, and Zetsche et al., Cell 2015 Oct. 22; 163(3):759-71, each of which are incorporated herein by reference in their entirety.
[0096] Nucleic acid-guided nuclease systems for use in the described systems and methods include those derived from an organism, wherein the organism is Thiomicrospira sp. XS5, Eubacterium rectale, Succinivibrio dextrinosolvens, Candidatus Methanoplasma termitum, Candidatus Methanomethylophilus alvus, Porphyromonas crevioricanis, Flavobacterium branchiophilum, Acidaminococcus sp., Acidomonococcus sp., Lachnospiraceae bacterium COE1, Prevotella brevis AT CC 19188, Smithella sp. SCADC, Moraxella bovoculi, Synergistes jonesii, Bacteroidetes oral taxon 274, Francisella tularensis, Leptospira inadai serovar Lyme str. 10, A ci domonococc us sp. crystal structure (5B43), S. mutans, S. agalactiae, S. equisimilis, S. sanguinis, S. pneumonia, C. jejuni, C. coli, N. salsuginis, N. tergarcus, S. auricularis, S. carnosus, N. meningitides, N. gonorrhoeae, L. monocytogenes, L. ivanovii, C. botulinum, C. difficile, C. tetani, C. sordellii, Francisella tularensis 1, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Butyrivibrio proteoclasticus B316, Peregrinibacteria bacterium GW201 l_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp. SCADC, Acidaminococcus sp. BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237 , Leptospirainadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis 3, Prevotella disiens, Porphyromonas macacae, Catenibacterium sp. CAG:290, Kandleria vitulina, Clostridiales bacterium KA00274, Lachnospiraceae bacterium 3-2, Dorea longicatena, Coprococcus catus GD / 7, Enterococcus columbae DSM 7374, Fructobacillus sp. EFB-N1, Weissella halotolerans, Pediococcus acidilactici, Lactobacillus curvatus, Streptococcus pyogenes, Lactobacillus versmoldensis, Filifactor alocis ATCC 35896, Alicyclobacillus acidoterrestris, Alicyclobacillus acidoterrestris ATCC 49025, Desulfovibrio inopinatus, Desulfovibrio inopinatus DSM 10711, Oleiphilus sp. Oleiphilus sp. HI0009, Candidtus kefeldibacteria, Parcubacteria CasY.4, Omnitrophica WOR 2 bacterium GWF2, Bacillus sp. NSP2.1, o Bacillus thermoamylovorans.
[0097] In one embodiment, the nucleic acid-guided nuclease disclosed herein includes an amino acid sequence having at least 50% amino acid sequence identity to SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO: 10, or SEQ ID NO:11 (FIG. 15-16). In one aspect, the nucleic acid-guided nuclease includes an amino acid sequence having at least about 10%, 20%, 30%, 40%, 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, greater than 95%, or 100% amino acid sequence identity to SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO: 10, or SEQ ID NO:11. In some cases, the nucleic acid-guided nuclease has at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or greater than 95%, amino acid sequence identity to SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO: 10, or SEQ ID NO:11. In some cases, the nucleic acid-guided nuclease has at least about 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or greater than 95%, amino acid sequence identity to SEQ ID NO:4.
[0098] In one aspect, the present disclosure provides methods of modifying expression of a polynucleotide in a cell. In an additional aspect, the modifying of expression of a polynucleotide in a cell includes contacting a target polynucleotide with a nucleic acid-guided nuclease, wherein expression of the target polynucleotide is increased. In another aspect, the modifying of expression of a polynucleotide in a cell includes contacting a target polynucleotide with a nucleic acid-guided nuclease, wherein expression of the target polynucleotide is decreased. In a further aspect, the modifying of expression of a polynucleotide in a cell includes contacting a target polynucleotide with a nucleic acid-guided nuclease, wherein expression of the target polynucleotide is silenced.
[0099] In one embodiment, the disclosure includes systems and methods for modifying one or more polynucleotides for treating a disease or disorder. In an additional aspect, the disclosureincludes administering an effective amount of the compositions or systems to a subject in need thereof. In one embodiment, the disclosure includes methods of detecting a target polynucleotide in a sample from a subject. In another aspect, the method of detecting a target polynucleotide in a sample from a subject includes contacting the sample with the compositions or systems described herein, which then generates a detectable signal, indicating the presence or absence of the target polynucleotide. In one aspect, the presence or absence of the target polynucleotide is used to diagnose a disease. In an additional aspect, the subject is treated when the target polynucleotide is present. In another aspect, the subject is treated when the target polynucleotide is absent.
[0100] In one embodiment, the present disclosure provides a method of modifying a target nucleic acid with a nucleic acid-guided nuclease, wherein the guide sequence directs sequencespecific binding to the target nucleic acid sequence, whereby the target nucleic acid sequence, the expression of the target nucleic acid, or both are modified. In one aspect, modifying of the target nucleic acid occurs in vitro, ex vivo, or in vivo. In another aspect, modifying the target nucleic acid includes cleaving the target nucleic acid. In an additional aspect, modifying expression of the target nucleic acid includes increasing or decreasing transcription or translation of the target nucleic acid. In a further aspect, the target nucleic acid is in a prokaryotic cell. In an additional aspect, the target nucleic acid is in a eukaryotic cell. In one aspect, the eukaryotic cell is a mammalian cell. In another aspect, the mammalian cell is a human cell.
[0101] In one embodiment, the present disclosure provides an isolated cell including a modified target nucleic acid according to the described methods. In one aspect, modification of the target nucleic acid of interest results in the cell including altered expression of at least one gene product. In another aspect, expression of the at least one gene product is increased. In an additional aspect, expression of the at least one gene product is decreased. In a further aspect, at least one gene product is modified. In one aspect, the present disclosure provides a plant or animal model including one or more cells described in the present disclosure.
[0102] In one embodiment, a nucleic acid-guided nuclease is provided or expressed in an in vitro system, a cell, or a eukaryotic organism, transiently or stably. In one aspect, the nucleic acid- guided nuclease encompasses homologs or orthologs of nucleic acid-guided nucleases described herein. As used herein, the term “homolog” of a protein refers to a protein of the same species which performs the same or a similar function. As used herein, the term “ortholog” refers to a protein of a distinct species which performs the same or a similar function. In another aspect, the homologue or orthologue of a nucleic acid-guided nuclease as described herein has a sequence homology or identity of at least 80%, at least 85%, at least 90%, at least 95%, or greater than 95% to a nucleic acid-guided nuclease. In a further aspect, the homologue or orthologue of an effectorcomplex described herein has a sequence identity of at least 80%, at least 85%, at least 90%, at least 95%, or greater than 95% to a wild type nucleic acid-guided nuclease. Homologs and orthologs are identified by various methods homology modelling known in the art.
[0103] An unmodified gene or nucleic acid sequence present naturally in an organism is referred to as a natural, endogenous, or wild type sequence.
[0104] In one embodiment, the nucleic acid-guided nuclease described herein has less than 50%, less than 60%, less than 70%, less than 80%, or less than 90% amino acid sequence identity to a polypeptide as described in U.S. Patent Publication 10,011,849, incorporated herein by reference.
[0105] In one embodiment, the effector complex includes a chimeric effector complex including a first fragment from a first effector complex ortholog (e.g., a Cpfl) and a second fragment from a second effector (e.g., a Cpfl) complex ortholog, and the first and second effector complex orthologs are different. In one aspect, at least one of the first and second effector complex orthologs (e.g., a Cpfl) includes an effector complex from an organism (e.g., a Cpfl), wherein the organism is Eubacterium, Thiomicrospira, Succinivibrio, Acidaminococcus, Alicyclobacillus, Desulfovibrio, Desulfonatronum, Opitutaceae, Tuberibacillus, Bacillus, Brevibacillus, Candidatus, Desulfatirhabdium, Elusimicrobia, Citrobacter, Methylobacterium, Omnitrophicai, Phycisphaerae, Planctomycetes, Spirochaetes, or Verrucomicrobiaceae. In an additional aspect, the chimeric effector complex includes a first fragment and a second fragment wherein the first and second fragments are selected from a Cpfl of an organism, wherein the organism tis Eubacterium, Thiomicrospira, Succinivibrio, Acidaminococcus, Alicyclobacillus, Desulfovibrio, Desulfonatronum, Opitutaceae, Tuberibacillus, Bacillus, Brevibacillus, Candidatus, Desulfatirhabdium, Elusimicrobia, Citrobacter, Methylobacterium, Omnitrophicai, Phycisphaerae, Planctomycetes, Spirochaetes, or Verrucomicrobiaceae wherein the first and second fragments are not from the same species. In a further aspect, the chimeric effector complex includes a first fragment and a second fragment wherein the first and second fragments are selected from a Cpfl of Eubacterium rectale, Thiomicrospira sp. XS5, Sussinivibrio dextrinosolvens, Acidaminococcus sp. BV3L6, Alicyclobacillus acidoterrestris (e.g., ATCC 49025), Alicyclobacillus contaminans (e.g., DSM 17975), Alicyclobacillus macrosporangiidus (e.g. DSM 17980), Bacillus hisashii strain C4, Candidatus lindowbacteria hacterzwm RIFCSPLOWO2, Desulfovibrio inopinatus (e.g., DSM 10711), Desulfonatronum thiodismutans (e.g., strain MLF- 1 ), Elusimicrobia bacterium RIFOXYA I 2, Omnitrophica W0R 2 bacterium RIFCSPHIGHO2, Opitutaceae bacterium TAV5, Phycisphaerae bacterium ST-NAGAB-D I, Planctomycetes bacterium RBG I 3 46 I 0, Spirochaetes bacterium GWB1_27 _13, Verrucomicrobiaceae bacterium UBA2429,Tuberibacillus calidus (e.g., DSM 17572), Bacillus thermoamylovorans (e.g., strain B4166), Brevibacillus sp. CF112, Bacillus sp. NSP2.1, Desulf atirhabdium butyrativorans (e.g., DSM 18734), Alicyclobacillus herbarius (e.g., DSM 13609), Citrobacter freundii (e.g., ATCC 8090), Brevibacillus agri (e.g., BAB-2500), or Methylobacterium nodulans (e.g., ORS 2060), wherein the first and second fragments are not from the same species.
[0106] In one embodiment, the effector complex described herein is derived from a bacterial species selected from Eubacterium rectale, Thiomicrospira sp. XS5, Succinivibrio dextrinosolvens, Acidaminococcus sp. BV3L6, Alicyclobacillus acidoterrestris (e.g., ATCC 49025), Alicyclobacillus contaminans (e.g., DSM 17975), Alicyclobacillus macrosporangiidus (e.g. DSM 17980), Bacillus hisashii strain C4, Candidatus lindowbacteria hacterzwm RIFCSPLOWO2, Desulfovibrio inopinatus (e.g., DSM 10711), Desulf onatronum thiodismutans (e.g., strain MLF- 1 ), Elusimicrobia bacterium RIFOXYA I 2,Omnitrophica WOR 2 bacterium RIFCSPHIGHO2, Opitutaceae bacterium TAV5,Phycisphaerae bacterium ST-NAGAB-D 1, Planctomycetes bacterium RBG I 3 46 I 0,Spirochaetes bacterium GWB1_27_13, Verrucomicrobiaceae bacterium UBA2429,Tuberibacillus calidus (e.g., DSM 17572), Bacillus thermoamylovorans (e.g., strainB4166), Brevibacillus sp. CF112, Bacillus sp. NSP2.1, Desulf atirhabdium butyrativorans e.g., DSM 18734), Alicyclobacillus herbarius e.g., DSM 13609), Citrobacter freundii (e.g., ATCC 8090), Brevibacillus agri (e.g., BAB-2500), and Methylobacterium nodulans (e.g., ORS 2060). In one aspect, the Cpfl is derived from a bacterial species selected from Eubacterium rectale, Thiomicrospira sp. XS5, Succinivibrio dextrinosolvens, and Acidaminococcus sp. BV3L6.
[0107] In one embodiment, the homolog or ortholog of the effector complex described herein has an amino acid sequence homology or identity of at least 80%, at least 85%, at least 90%, at least 95%, or greater than 95% to Cpfl. In further embodiments, the homologue or orthologue of Cpfl described herein has an amino acid sequence identity of at least 80%, at least 85%, at least 90%, at least 95%, or greater than 95% to a wild type Cpfl. Where the effector complex has one or more mutations (mutated), the homologue or orthologue of said effector complex as described herein has an amino acid sequence identity of at least 80%, at least 85%, at least 90%, at least 95%, or greater than 95% with the mutated effector complex.
[0108] In one embodiment, the effector complex protein described herein has an amino acid sequence homology or identity of at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, or greater than 95% of the wild type effector complex. In one aspect, the effector complex protein has an amino acid sequence identity of at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, or greater than 95% to the wild type effectorcomplex. In another aspect, a homologous or orthologous effector complex includes truncated forms of the effector complex, wherein the sequence identity or homology is determined over the length of the truncated effector complex.
[0109] As used herein, the term “guide sequence,” “crRNA,” “guide RNA,” “guide nucleotide,” “single guide RNA,” or “gRNA” refers to a polynucleotide including any polynucleotide sequence having sufficient complementarity with a target nucleic acid sequence to hybridize with the target nucleic acid sequence and to direct sequence-specific binding of a nucleic acid-guided nuclease to the target nucleic acid sequence. In one aspect, the degree of complementarity is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more. The ability of a guide sequence (within a nucleic acid-targeting guide RNA) to direct sequence-specific binding of a nucleic acid-guided nuclease to a target nucleic acid sequence is assessed by any suitable assay. In one embodiment, the components of a nucleic acid-guided nuclease, including an effector complex and a guide sequence, are provided to a host cell having the corresponding target nucleic acid sequence. In one aspect, the nucleic acid-guided nuclease is provided to a host cell by transfection with at least one vector encoding the components of the nucleic acid-guided nuclease, followed by an assessment of preferential targeting of the target nucleic acid sequence. In an additional aspect, a guide sequence and an effector complex are selected to target any target nucleic acid sequence. In another aspect, the target sequence is DNA. In a further aspect, the target sequence is DNA encoding any RNA sequence, such as messenger RNA (mRNA), pre-mRNA, ribosomal RNA (rRNA), transfer RNA (tRNA), micro-RNA (miRNA), small interfering RNA (siRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), double stranded RNA (dsRNA), noncoding RNA (ncRNA), long non-coding RNA (IncRNA), and small cytoplasmic RNA (scRNA).
[0110] In one embodiment, the nucleic acid-guided nuclease includes a guide RNA and a tracrRNA. In one aspect, the guide RNA and tracrRNA are two separate polynucleotides. In an additional aspect, the guide RNA and tracrRNA are a single polynucleotide.
[0111] In one embodiment, the guide sequence or spacer length of the guide RNA is from 15 to 50 nt. In one aspect, the spacer length of the guide RNA is at least 15 nucleotides. In another aspect, the spacer length is from 15 to 17 nt, from 17 to 20 nt, from 20 to 24 nt, from 23 to 25 nt, from 24 to 27 nt, from 27-30 nt, from 30-35 nt or longer. In one aspect, the spacer length is 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, 26 nt, 27 nt, 28 nt, 29 nt, 30 nt, 31 nt, 32 nt, 33 nt, 34 nt, or 35 nt, or longer. In an additional aspect, the guide sequence is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39 40, 41, 42, 43, 44, 45, 46, 47 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69,70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nt.
[0112] In one embodiment, the guide sequence is derived from a different species than the effector complex. In one embodiment, the effector complex is a Cpfl from a Eubacterium rectale, Thiomicrospira sp. XS5, Succinivibrio dextrinosolvens, or Acidaminococcus sp. BV3L6. When a chimeric effector complex is utilized, a related guide RNA is used. In one aspect, when a chimeric effector complex is utilized, a guide sequence is from the same or another species. In an aspect, the guide sequence includes at least 95%, 96%, 97% or more sequence similarity to the guide sequence of another species.
[0113] In one embodiment, a target polynucleotide is inactivated to modify expression of the target polynucleotide in a cell. In one aspect, the target polynucleotide is inactivated such that the sequence is not transcribed, or the coded protein is not produced. In another aspect, a protein or microRNA coding sequence is inactivated, and the protein is not produced.
[0114] In one embodiment, an inactivated target sequence includes a deletion mutation, an insertion mutation, or a nonsense mutation. In one aspect, the inactivation of a target sequence results in “knockout” of the target sequence. In an additional aspect, knockout of the target sequence halts expression of a gene. In one aspect, a deletion means that at least part of the object nucleic acid sequence is deleted. In another aspect, a deletion means that the entire sequence is deleted.
[0115] A wide variety of labels suitable for detecting protein levels are known in the art. Nonlimiting examples include radioisotopes, enzymes, colloidal metals, fluorescent compounds, bioluminescent compounds, and chemiluminescent compounds. In one aspect, the detection system includes tandem dimer tomato (tdTomato), luciferase, or mCherry fluorescent reporters.
[0116] The amount of agent: polypeptide complexes formed during a binding reaction is quantified by standard quantitative assays. In another aspect, the formation of an agent: polypeptide complex is measured by quantifying the amount of label remaining at the site of binding.
[0117] In one embodiment, a suitable vector is introduced to a cell, tissue, organism, or embryo via one or more methods known in the art, including micro injection, electroporation, sonoporation, biolistics, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendrimer transfection, heat shock transfection, nucleofection transfection, magnetofection, lipofection, impalefection, optical transfection, a gene-gun, proprietary agent- enhanced uptake of nucleic acids, delivery via liposomes, immunoliposomes, virosomes, particles, exosomes, microvesicles, viral vectors, or artificial virions. In an additional aspect, the vector isintroduced into a cell, tissue, organism, or embryo by micro injection. The vector or vectors are micro injected into the nucleus or the cytoplasm of the cell, tissue, or embryo. In another aspect, the vector or vectors are introduced into a cell by nucleofection.
[0118] A target polynucleotide of a nucleic acid-guided nuclease is any polynucleotide endogenous or exogenous to the host cell. In one aspect, the target polynucleotide is a polynucleotide residing in the nucleus of the eukaryotic cell. In another aspect, the target polynucleotide is a polynucleotide residing in the genome of a prokaryotic cell. In an additional aspect, the target polynucleotide is a polynucleotide residing in an extrachromosomal vector of a host cell. In a further aspect, the target polynucleotide is a sequence coding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or a junk DNA).
[0119] Examples of target polynucleotides include a sequence associated with a signaling biochemical pathway. In one aspect, the sequence is a signaling biochemical pathway-associated gene or polynucleotide. Examples of target polynucleotides include a disease-associated gene or polynucleotide. As used herein, a “disease-associated” gene or polynucleotide refers to any gene or polynucleotide which produces transcription or translation products at an abnormal level or in an abnormal form in cells derived from affected tissues compared with tissues or cells of a nondiseased tissue. In one aspect, a disease-associated gene is a gene that becomes expressed at an abnormally high level or at an abnormally low level, wherein the altered expression correlates with a disease or disorder. In another aspect, a disease-associated gene refers to a gene possessing a mutation or genetic variation that is directly responsible or is in linkage disequilibrium with a gene that is responsible for the etiology of a disease or disorder.
[0120] The term “genetic modification” as used herein refers to a mutation, a disruption, a gene “knock out,” a deletion, an insertion, insertion of a stop codon, an inactivation, an attenuation, a rearrangement, one or more point mutations, a frameshift mutation, an inversion, a single nucleotide polymorphism (SNP), a truncation, or a point mutation that changes the activity or expression of one or more genes or nucleic acids. In one aspect, the change in expression is a reduction or inactivation of expression. The genetic modification is made in any polynucleotide sequence that affects expression of a gene or nucleic acid sequence, or that modifies the nature or quantity of a gene product. In another aspect, the genetic modification is to a coding or non-coding sequence, a promoter, a terminator, an exon, an intron, a 3' or 5' UTR, or another regulatory sequence. In one aspect, a genetic modification to a structure of the gene results in attenuation or elimination of the nucleic acid product. In an additional aspect, the genetic modification is made to the host cell’s native genome. In another aspect, a recombinant cell or organism having attenuated expression of a gene has one or more genetic mutations, which are one or morenucleobase changes, one or more nucleobase deletions, or one or more nucleobase insertions. In a further aspect, the one or more genetic mutations are in the region of a gene located 5' of the transcriptional start site. In one example, the one or more genetic mutations are within about 2 kb, within about 1.5 kb, within about 1 kb, or within about 0.5 kb of a known or putative transcriptional start site, or within about 3 kb, within about 2.5 kb, within about 2kb, within about 1.5 kb, within about 1 kb, or within about 0.5 kb of a translational start site.Editing Cassette
[0121] Disclosed herein are systems and methods for editing a target polynucleotide sequence. Such systems include polynucleotides containing one or more components of a targetable nuclease system. Polynucleotide sequences for use in these systems and methods are referred to as editing cassettes. In one aspect, an editing cassette includes one or more primer sites. Primer sites amplify an editing cassette by hybridizing to oligonucleotide primers including reverse complementary sequences that hybridize to the one or more primer sites. In another aspect, an editing cassette includes two or more primer sites. In an additional aspect, an editing cassette includes a primer site on each end of the editing cassette, said primer sites flanking one or more of the other components of the editing cassette. In one aspect, primer sites are approximately 8, 10, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 or more nucleotides in length.
[0122] In one aspect, an editing cassette includes an editing template as described herein. In an additional aspect, an editing cassette includes a donor sequence. A donor sequence is homologous to a target sequence. In a further aspect, a donor sequence includes at least one mutation relative to a target sequence. In one aspect, a donor sequence includes a homology region (or homology arms) flanking at least one mutation relative to a target sequence, such that the flanking homology regions facilitate homologous recombination of the donor sequence into a target sequence. In a further aspect, a donor sequence includes an editing template as described herein.
[0123] A vector is linear or circular and is generated by chemical synthesis, Gibson assembly, SLIC, CPEC, PCA, ligation-free cloning, overlapping oligo extension, in vitro assembly, in vitro oligo assembly, PCR, traditional ligation-based cloning, or other known methods in the art.
[0124] In one embodiment, a nucleic acid molecule is synthesized which includes one or more elements disclosed herein. In one aspect, a nucleic acid molecule is synthesized that includes an editing cassette. In a further aspect, a nucleic acid molecule is synthesized that includes a guide nucleic acid. In another aspect, a nucleic acid molecule is synthesized that includes a homology arm. In any of these aspects, the guide nucleic acid is optionally operably linked to a promoter.
[0125] In one embodiment, the methods described herein are carried out in any type of cell in which a targetable nuclease system functions, including prokaryotic and eukaryotic cells. In oneaspect the cell is a bacterial cell, such as Escherichia spp. (e.g., E. coli) or a fungal cell such as a yeast cell, e.g., Saccharomyces spp. In another aspect, the cell is a eukaryotic cell, a plant cell, an insect cell, a mammalian cell, or a human cell. In another aspect, the cell is a human cancer cell. In a further aspect, the cell is a stem cell. In one aspect, the cell is a mammalian stem cell, such as an embryonic or induced pluripotent stem cell.
[0126] Methods for genome editing include: (a) introducing a vector that encodes at least one editing cassette and at least one guide nucleic acid into a first population of cells, thereby producing a second population of cells including the vector; (b) maintaining the second population of cells under conditions in which a nucleic acid-guided nuclease is expressed or maintained, wherein the nucleic acid-guided nuclease is encoded on the vector, on a second vector, on the genome of cells of the second population of cells, or otherwise encoded in the cell, resulting in DNA cleavage and incorporation of the editing cassette; (c) obtaining viable cells; and (d) sequencing the target DNA molecule in at least one cell of the second population of cells to identify the mutation of at least one target sequence.
[0127] Methods for genome editing include: (a) introducing a vector that encodes at least one editing cassette and at least one guide nucleic acid into a first population of cells, thereby producing a second population of cells including the vector; (b) maintaining the second population of cells under conditions in which a nucleic acid-guided nuclease is expressed or maintained, wherein the nucleic acid-guided nuclease is encoded on the vector, on a second vector, on the genome of cells of the second population of cells, or otherwise encoded in the cell, resulting in regulation of transcription of a target sequence; (c) obtaining viable cells; and (d) sequencing the target DNA molecule in at least one cell of the second population of cells to identify the mutation of at least one target sequence. In one aspect, a target gene is a gene that is overexpressed, underexpressed, or not expressed, wherein the over-expression, under-expression, or non-expression causes a disease or disorder in a subject.
[0128] The terms “polynucleotide,” “nucleotide,” “nucleotide sequence,” “nucleic acid,” and “oligonucleotide” are used interchangeably. The term polynucleotide refers to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. Polynucleotides include coding or non-coding regions of a gene or gene fragment, a locus identified from linkage analysis, an exon, an intron, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), a ribozyme, cDNA, a recombinant polynucleotide, a branched polynucleotide, a plasmid, a vector, or isolated DNA of any sequence, isolated RNA of any sequence, a nucleic acid probe, or a primer. The term also encompasses nucleic acid-like structureswith synthetic backbones. A polynucleotide may include one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. In one aspect, the sequence of polynucleotide is interrupted by a non-nucleotide component. In another aspect, a polynucleotide is further modified after polymerization, such as by conjugation with a labeling component.
[0129] The terms “expressing” and “expression” in reference to a polypeptide are intended to be understood in the ordinary meaning as used in the art. A polypeptide is expressed by a cell via transcription of a nucleic acid into mRNA, followed by translation into an initial polypeptide, which is folded and possibly further processed to a mature polypeptide. The statement that a cell or an organism is expressing a polypeptide indicates that the polypeptide is found in or on the cell and implies that the polypeptide has been synthesized by the polypeptide expression machinery of the cell.
[0130] With regard to the respective biological process itself, the terms “expression”, “gene expression” or “expressing” refer to the pathways converting the information encoded in the nucleic acid sequence of a gene first into messenger RNA (mRNA) and then into a polypeptide. The expression of a gene includes its transcription into a primary pre-RNA, the processing of this pre-RNA into a mature RNA and the translation of the mRNA sequence into the corresponding amino acid sequence of the polypeptide. In this context, it is also noted that the term “gene product” refers not only to a polypeptide, including for example a mature polypeptide (including a splice variant thereof) encoded by that gene and a respective precursor protein where applicable, but also to the respective mRNA, which may be regarded as the first gene product during the course of gene expression.
[0131] The terms “subject,” “patient,” or “subjects” as used herein, refer to a human or other animal, including rodents, ungulates, or mammals, for example, horses, cattle, sheep, pigs, goats, llama, camel, dogs, cats, birds, ferrets, rabbits, squirrels, mice, rats, or ferrets. In one embodiment, the subject is a human subject.
[0132] The term “treatment” is used interchangeably herein with the term “therapeutic method” or “therapy” and refers to therapeutic treatments or measures that cure, slow down, lessen symptoms of, and / or halt progression of a diagnosed pathologic conditions or disorder, for example, inherited eye disease, neurodegenerative disease, cancer, or HIV. The term “treatment” also refers to prophylactic / preventative measures. A subject in need of treatment includes an individual already diagnosed with a disease or disorder as well as an individual who is at risk for acquiring a disease or disorder, for example, a subject in need of a preventive measure.
[0133] Presented below are examples discussing genomic editing to increase nuclease expression intensity contemplated for the discussed applications. The following examples areprovided to further illustrate the embodiments of the present invention but are not intended to limit the scope of the invention. While they are typical of those that might be used, other procedures, methodologies, or techniques known to those skilled in the art may alternatively be used.EXAMPLES
[0134] The following examples are provided to further illustrate the embodiments of the present invention but are not intended to limit the scope of the invention. While they are typical of those that might be used, other procedures, methodologies, or techniques known to those skilled in the art may alternatively be used.EXAMPLE 1MethodsCell Culture
[0135] K562 cells (ATCC catalog no. CCL-243) were cultured in RPMI 1640 supplemented with 10% FBS, lx Antibiotic-Antimycotic . 771-3G cells (human induced pluripotent stem cells established from endothelial progenitor cells derived from peripheral blood) (Passage number 14) were maintained in a serum-free and feeder-free maintenance medium containing stabilized FGF2 / bFGF (mTeSR Plus Basal Medium), and plated on feeder-free iMatrix 511 (Nippi)-coated plates. The medium was changed every day. Subculture was carried out every 4-6 days via the EDTA method. After plating, 10 pM Y-27632 (ROCK inhibitor) was added for 1 day.Plasmid Construction
[0136] Plasmid vectors were constructed by NEBuilder HiFi DNA Assembly (NEB) with DNA fragments prepared by PCR using genomic DNA extracted from induced pluripotent stem cells (iPSCs, 771-3G) as a template, or synthesized by Integrated DNA Technologies. For preparation of plasmids for transfection, the ZymoPURE Plasmid Miniprep Kit was used.Transfection
[0137] iPSCs (70-80% confluent) were harvested with Accutase™. 2 million of K562 cells or iPSCs were resuspended with 100 uL of OptiMEM containing plasmid vectors (donor & guide 2- 5 ug and nuclease 2-5 ug), transferred to EC-002S Electroporation Cuvette, 2mm gap, then subjected to electroporation using NEPA21 Super Electroporator. The parameters for K562 cells were as follows; voltage, 275 V; pulse length, 2.5 ms; pulse interval, 50 m; number of pulse, 2; decay rate, 10%; polarity + as poring pulse and 20 V; pulse length, 50 ms; pulse interval, 50 ms; number of pulses, 5; decay rate, 40%; and polarity + / - as transfer pulse. The parameters for iPSCs were as follows; voltage, 150 V; pulse length, 5 ms; pulse interval, 50 m; number of pulse, 2; decay rate, 10%; polarity + as poring pulse and 20 V; pulse length, 50 ms; pulse interval, 50 ms;number of pulses, 5; decay rate, 40%; and polarity + / - as transfer pulse. After transfection, K562 cells were cultured in 12 well plates in RPMI Medium supplemented with 10% FBS and 1 x Antibiotic-Antimycotic, and iPSCs were plated into three wells of an iMatrix511 -coated six-well plate and maintained with mTeSR Plus supplemented with 10 pM Y-27632 for 3 days. Then cells were maintained with mTeSR without Y-27632.Single Cell Cloning
[0138] After 7 days of transfection with plasmids (SEQ ID NO:18 and SEQ ID NO:19), truncated CD 19 (tCD19) positive iPSCs were enriched by magnetic assisted cell sorting (MACS) using CD 19 microbeads, human and MS columns according to manufacturer’s protocol with slight modification (10 pM Y-27632 was added to MACS buffer). Then, cells were plated at 200 cells / well in 6 well plate coated with iMatrix-511 and cultured with mTeSR Plus supplemented with 10 pM Y-27632 for 3 days, then cultured with mTeSR plus without Y-27632 for 5 days. After treatment of 0.5 pM EDTA / PBS for 10 min, mCherry positive colonies were manually picked under microscopic observation, and the cells were transferred to 24 well plate coated with iMatrix- 511, then the cells were expanded.Quantification of HDR efficiency by Bioluminescence Detection
[0139] After 3 days of transfection with luciferase reporting donor plus guide RNA expression vector and nuclease expression vector, 1 ml of K562 cells were harvested, resuspended in 100 uL of culture media supplemented with 200 ug / mL D-luciferin, transferred to a 96 well white plate, incubated for 5 min at room temperature, and bioluminescence was measured using GloMax Navigator. After the measurement of bioluminescence, relative cell number was determined using CellTiter-Glo 2.0 Cell Viability Assay according to manufacturer’s protocol.Genotyping
[0140] Selection marker integration into targeted sites was detected by PCR using PrimeSTAR GXL according to the manufacturer’s protocol, followed by agarose gel electrophoresis. DNA amplicons were visualized using Midori Green Advance and a blue light LED transilluminator. Sanger sequencing was outsourced.EXAMPLE 2Quantitative analysis of genome editing using HDR detection system in K562 cells
[0141] Genome editing was evaluated using the HDR detection system shown in FIG. 1 and FIG. 2.
[0142] Qualitative detection of genome editing through HDRs was performed using a tdTomato fluorescent protein reporter in K562 cells. Donor vector and guide RNA were designed as shown in FIG. 3. After 3 days of electroporation with a donor / guide expression vector and a nucleaseexpression vector, HDR was detected when the nucleases of SEQ ID NO: 1 (wild-type full-length E. rectale Casl2a nuclease), SEQ ID NO:3 (Chimera), SEQ ID NON, SEQ ID NO:5 or SEQ ID NO:6 were combined with guide RNAs with the backbone of SEQ ID NO: 12 or SEQ ID NO: 13 (FIG. 4).
[0143] The results obtained from the comparison of M44 (SEQ ID NO:7) with SEQ ID NO:3 and SEQ ID NO:6 indicated that when the REC 1 domain of ErCpfl was exchanged with the REC 1 domain of TxCpfl, the nuclease function of the resultant chimeric Cpfl was maintained. When the WED I domain was also exchanged, in SEQ ID NO:7, nuclease function was lost. When the Nuc domain of ErCpfl was replaced with the Nuc domain of ScCpfl (SEQ ID NON), nuclease function was maintained at a level comparable to that of wild-type ErCpfl (FIG. 4 and Table 1).
[0144] Furthermore, when the RECI and Nuc domains of ErCpfl were replaced with the RECI domain from TxCpfl and the Nuc domain of ScCpfl (SEQ ID NON), the chimeric Cpfl showed nuclease activity (FIG. 4). Thus, the invention disclosed here includes a new chimeric Cpfl nuclease including domains from multiple species by exchanging only REC 1 , only Nuc, or only RECI and Nuc with sequences of Cpfl from other species. The present invention replaces domains of one species of Cpfl with domains from other species, generating an efficient nuclease system.EXAMPLE 3Quantitative analysis of HDR using a luciferase reporter in K562 cells.
[0145] The activity of various chimeric nucleases was quantitatively evaluated in K562 cells by switching from tdTomato to Luciferase and measuring the amount of luminescence emitted.
[0146] The nucleases of SEQ ID NON and SEQ ID NON showed improved performance compared to other nucleases tested in this study (Table 1, Table la). Table 1 shows the results of the comparison of SEQ ID NO: 12 and SEQ ID NO: 13 guide RNA scaffolds with nucleases from SEQ ID NON, SEQ ID NON, and SEQ ID NON detecting luminescence based on Luciferase expression caused by HDR. The efficacy of these chimeric nucleases is unexpected due to the fact that the species-swapped chimeras do not exist in nature, and therefore the sequences were not subject to evolutionary pressures over time that increase structural stability and functional activity of wild type nucleases. A person of skill in the art would have assumed that chimeric nucleases, especially chimeric nucleases previously shown to be non-functional, would not be as stable as non-chimeric nucleases, and thus would not have improved activity.Table 1 : Luciferase Expression Caused by HDR
[0147] Table 1A shows the results of the comparison of several Cpfl chimeras, detecting luminescence based on Luciferase expression caused by HDR.Table 1 A: Luciferase Expression Caused by HDR
[0148] Table 2 shows the results of the comparison of SEQ ID NO: 12 and SEQ ID NO: 13 guide RNA scaffolds with nucleases from SEQ ID NO:1, SEQ ID NO:4, and SEQ ID NO:6, detecting luminescence using an ATP measurement kit. Consistent effects on cell viability assessed by ATP measurement were not detected for all nucleases, as indicated by the results shown in Tables 2 and 2A.Table 2: Luminescence Detection Using ATP Measurement Kit
[0149] Table 2A shows the results of the comparison of several Cpfl chimeras, detecting luminescence using an ATP measurement kit.Table 2A: Luminescence Detection Using ATP KitEXAMPLE 4Genome editing in human induced pluripotent stem cells (iPSCs)
[0150] SEQ ID NO:4 and RNAs with a backbone of SEQ ID NO: 12 were introduced by transfection of pCAG-GS plasmid vectors (SEQ ID NO: 18 and SEQ ID NO: 19). After 3 days of electroporation, mCherry positive cells appeared (FIG. 5A), then these positive cells were subjected to single cell cloning and genotyping using primers TTGGAAGGTTTCTGCTGTCACTC (SEQ ID NO:23) and GAATCCAGACCTCAGCCCATAG (SEQ ID NO:24). Polynucleotide sequence flanked by thehomology arms of the donor vector was inserted into the targeting region in all mCherry positive clones (FIG. 5B).EXAMPLE 5Quantitative analysis of genome editing using HDR detection system in conditions of Cpfl chimeras with Nuc domain replacements
[0151] Compared to the system described in FIG. 2, changing the nuclease expression promoter from the CAG promoter to EFla resulted in an increase in nuclease expression intensity, and stable experimental results were obtained. The detailed sequence of the expression plasmid for the EFla promoter-driven nuclease (SEQ ID NO: 4) is shown in SEQ ID NO: 32 (FIG. 25), and the various chimeric nucleases used in the example in FIG. 24A-24B were expressed using plasmids in which the cDNA fragment encoding the nuclease (SEQ ID NO:4) in the expression plasmid shown in SEQ ID NO:32 (FIG. 11) was replaced with the cDNA sequence encoding each chimeric nuclease. The same Nuc domain replacement as that of chimera Cpfl shown in SEQ ID NO:4 was performed. FIG. 23 shows the detailed crossover points.
[0152] Cpfl chimeras with Nuc domain replacements at the positions shown in FIG. 23 were evaluated using the HDR quantification system shown in FIG. 22 using K562 cells. The guide scaffold used here is SEQ ID NO: 13. The Cpfl chimera indicated by SEQ ID NO:26 and SEQ ID NO:30 showed higher HDR activity than wild-type ErCpfl (FIG. 24A-24B), indicating that the replacement of the Nuc domain can flexibly regulate HDR activity. Three independent evaluations were conducted, each data point shows the data from each experiment, and the bar graph shows the average of these.Sequences
[0153] Sequences below are written as DNA sequences from which RNA sequences can be synthesized.
[0154] SEQ ID NO: 1 is a Cpfl derived from Eubacterium rectal e
[0155] MNNGTNNFQNFIGISSLQKTLRNALIPTETTQQFIVKNGIIKEDELRGENRQILKD IMDDYYRGFISETLSSIDDIDWTSLFEKMEIQLKNGDNKDTLIKEQAEKRKAIYKKFADD DRFKNMFSAKLISDILPEFVIHNNNYSASEKEEKTQVIKLFSRFATSFKDYFKNRANCFSA DDISSSSCHRIVNDNAEIFFSNALVYRRIVKNLSNDDINKISGDIKDSLKEMSLEEIYSYEK YGEFITQEGISFYNDICGKVNSFMNLYCQKNKENKNLYKLRKLHKQILCIADTSYEVPY KFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNLDKIYIVSKFYESVSQKTYRD WETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSITEINELVSNYKLCPDDNIKA ETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVLDVIMNAFHWCSVFMTEELV DKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQKPYSTKKIKLNFGIPTLADGWSKSKEY31SNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDYKKMIYNLLPGPNKMIPKVFL SSKTGVETYKPSAYILEGYKQNKHLKSSKDFDITFCHDLIDYFKNCIAIHPEWKNFGFDF SDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKGQLYLFQIYNKDFSKKSSGND NLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPIIHKKGSILVNRTYEAEEKDQF GNIQIVRKTIPENIYQELYKYFNDKSDKELSDEAAKLKNWGHHEAATNIVKDYRYTYD KYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGIDRGERNLIYVSVIDTCGNIVEQK SFNIVNGYDYQIKLKQQEGARQIARKEWKEIGKIKEIKEGYLSLVIHEISKMVIKYNAIIA MEDLSYGFKKGRFKVERQVYQKFETMLINKLNYLVFKDISITENGGLLKGYQLTYIPEK LKNVGHQCGCIFYVPAAYTSKIDPTTGFVNIFKFKDLTVDAKREFIKKFDSIRYDSEKNL FCFTFDYNNFITQNTVMSKSSWSVYTYGVRIKRRFVNGRFSNESDTIDITKDMEKTLEM TDINWRDGHDLRQDIIDYEIVQHIFEIFKLTVQMRNSLSELEDRDYDRLISPVLNENNIFY DSAKAGDALPKDADANGAYCIALKGLYEIKQITENWKEDGKFSRDKLKISNKDWFDFI QNKRYL
[0156] SEQ ID NO:2 is a Cpfl derived from Eubacterium rectale (MAD7):
[0157] MNNGTNNFQNFIGISSLQKTLRNALIPTETTQQFIVKNGIIKEDELRGENRQILKD IMDDYYRGFISETLSSIDDIDWTSLFEKMEIQLKNGDNKDTLIKEQTEYRKAIHKKFAND DRFKNMFSAKLISDILPEFVIHNNNYSASEKEEKTQVIKLFSRFATSFKDYFKNRANCFSA DDISSSSCHRIVNDNAEIFFSNALVYRRIVKSLSNDDINKISGDMKDSLKEMSLEEIYSYE KYGEFITQEGISFYNDICGKVNSFMNLYCQKNKENKNLYKLQKLHKQILCIADTSYEVP YKFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNLDKIYIVSKFYESVSQKTYR DWETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSITEINELVSNYKLCSDDNIK AETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVLDVIMNAFHWCSVFMTEEL VDKDNNFYAELEEIYDEIYPVISLYNLVRNYVTQKPYSTKKIKLNFGIPTLADGWSKSKE YSNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDYKKMIYNLLPGPNKMIPKV FLSSKTGVETYKPSAYILEGYKQNKHIKSSKDFDITFCHDLIDYFKNCIAIHPEWKNFGFDFSDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKGQLYLFQIYNKDFSKKSTGN DNLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPIIHKKGSILVNRTYEAEEKDQ FGNIQIVRKNIPENIYQELYKYFNDKSDKELSDEAAKLKNWGHHEAATNIVKDYRYTY DKYFLHMPITINFKANKTGFINDRILQYIAKEKDLHVIGIDRGERNLIYVSVIDTCGNIVE QKSFNIVNGYDYQIKLKQQEGARQIARKEWKEIGKIKEIKEGYLSLVIHEISKMVIKYNAI IAMEDLSYGFKKGRFKVERQVYQKFETMLINKLNYLVFKDISITENGGLLKGYQLTYIP DKLKNVGHQCGCIFYVPAAYTSKIDPTTGFVNIFKFKDLTVDAKREFIKKFDSIRYDSEK NLFCFTFDYNNFITQNTVMSKSSWSVYTYGVRIKRRFVNGRFSNESDTIDITKDMEKTLE MTDINWRDGHDLRQDIIDYEIVQHIFEIFRLTVQMRNSLSELEDRDYDRLISPVLNENNIFYDSAKAGDALPKDADANGAYCIALKGLYEIKQITENWKEDGKFSRDKLKISNKDWFDF IQNKRYL
[0158] SEQ ID NO:3 is Chimera 1 :
[0159] MNNGTNNFQNFIGISSLQKTLRNALIPTETTQQFIVKNGIIKEDERRAVDYQKVK EIIDDYHRDFIEESLNYFPEQVSKDALEQAFHLYQKLKAAKVEEREKALKEWEALQKKL REKWKCFSDSNKARFSRIDKKELIKEDLINWLVAQNREDDIPTVETFNNFTTYFTGFHE NRKNIYSKDDHATAISFRLIHENLPKFFDNVISFNKLKEGFPELKFDKVKEDLEVDYDLK HAFEIEYFVNFVTQAGIDQYNYLLGGKTLEDGTKKQGMNEQINLFKQQQTRDKARQIP KLIPLFKQILSERTESQSFIPYKFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNL DKIYIVSKFYESVSQKTYRDWETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSI TEINELVSNYKLCPDDNIKAETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVL DVIMNAFHWCSVFMTEELVDKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQKPYSTKK IKLNFGIPTLADGWSKSKEYSNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDY KKMIYNLLPGPNKMIPKVFLSSKTGVETYKPSAYILEGYKQNKHLKSSKDFDITFCHDLI DYFKNCIAIHPEWKNFGFDFSDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKG QLYLFQIYNKDFSKKSSGNDNLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPII HKKGSILVNRTYEAEEKDQFGNIQIVRKTIPENIYQELYKYFNDKSDKELSDEAAKLKN WGHHEAATNIVKDYRYTYDKYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGID RGERNLIYVSVIDTCGNIVEQKSFNIVNGYDYQIKLKQQEGARQIARKEWKEIGKIKEIK EGYLSLVIHEISKMVIKYNAIIAMEDLSYGFKKGRFKVERQVYQKFETMLINKLNYLVF KDISITENGGLLKGYQLTYIPEKLKNVGHQCGCIFYVPAAYTSKIDPTTGFVNIFKFKDLT VDAKREFIKKFDSIRYDSEKNLFCFTFDYNNFITQNTVMSKSSWSVYTYGVRIKRRFVN GRFSNESDTIDITKDMEKTLEMTDINWRDGHDLRQDIIDYEIVQHIFEIFKLTVQMRNSLS ELEDRDYDRLISPVLNENNIFYDSAKAGDALPKDADANGAYCIALKGLYEIKQITENWK EDGI<FSRDI<LI<ISNI<DWFDFIQNI<RYL
[0160] SEQ ID NO:4 is Chimera 2:
[0161] MNNGTNNFQNFIGISSLQKTLRNALIPTETTQQFIVKNGIIKEDELRGENRQILKD IMDDYYRGFISETLSSIDDIDWTSLFEKMEIQLKNGDNKDTLIKEQAEKRKAIYKKFADD DRFKNMFSAKLISDILPEFVIHNNNYSASEKEEKTQVIKLFSRFATSFKDYFKNRANCFSA DDISSSSCHRIVNDNAEIFFSNALVYRRIVKNLSNDDINKISGDIKDSLKEMSLEEIYSYEK YGEFITQEGISFYNDICGKVNSFMNLYCQKNKENKNLYKLRKLHKQILCIADTSYEVPY KFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNLDKIYIVSKFYESVSQKTYRD WETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSITEINELVSNYKLCPDDNIKA ETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVLDVIMNAFHWCSVFMTEELVDKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQKPYSTKKIKLNFGIPTLADGWSKSKEY SNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDYKKMIYNLLPGPNKMIPKVFL SSKTGVETYKPSAYILEGYKQNKHLKSSKDFDITFCHDLIDYFKNCIAIHPEWKNFGFDF SDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKGQLYLFQIYNKDFSKKSSGND NLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPIIHKKGSILVNRTYEAEEKDQF GNIQIVRKTIPENIYQELYKYFNDKSDKELSDEAAKLKNWGHHEAATNIVKDYRYTYD KYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGIDRGERNLIYVSVIDTCGNIVEQK SFNIVNGYDYQIKLKQQEGARQIARKEWKEIGKIKEIKEGYLSLVIHEISKMVIKYNAIIA MEDLSYGFKKGRFKVERQVYQKFETMLINKLNYLVFKDISITENGGLLKGYQLTYIPEK LKNVGHQCGCIFYVPAAYTSKIDPTTGFANVLNLSKVRNVDAIKSFFSNFNEISYSKKEA LFKFSFDLDSLSKKGFSSFVKFSKSKWNVYTFGERIIKPKNKQGYREDKRINLTFEMKKL LNEYKVSFDLENNLIPNLTSANLKDTFWKELFFIFKTTLQLRNSVTNGKEDVLISPVKNA KGEFFVSGTHNKTLPQDCDANGAYCIALKGLYEIKQITENWKEDGKFSRDKLKISNKD WFDFIQNKRYL
[0162] SEQ ID NO:5 is Chimera 3:
[0163] MNNGTNNFQNFIGISSLQKTLRNALIPTETTQQFIVKNGIIKEDERRAVDYQKVK EIIDDYHRDFIEESLNYFPEQVSKDALEQAFHLYQKLKAAKVEEREKALKEWEALQKKL REKWKCFSDSNKARFSRIDKKELIKEDLINWLVAQNREDDIPTVETFNNFTTYFTGFHE NRKNIYSKDDHATAISFRLIHENLPKFFDNVISFNKLKEGFPELKFDKVKEDLEVDYDLK HAFEIEYFVNFVTQAGIDQYNYLLGGKTLEDGTKKQGMNEQINLFKQQQTRDKARQIP KLIPLFKQILSERTESQSFIPYKFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNL DKIYIVSKFYESVSQKTYRDWETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSI TEINELVSNYKLCPDDNIKAETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVL DVIMNAFHWCSVFMTEELVDKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQKPYSTKK IKLNFGIPTLADGWSKSKEYSNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDY KKMIYNLLPGPNKMIPKVFLSSKTGVETYKPSAYILEGYKQNKHLKSSKDFDITFCHDLI DYFKNCIAIHPEWKNFGFDFSDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKG QLYLFQIYNKDFSKKSSGNDNLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPII HKKGSILVNRTYEAEEKDQFGNIQIVRKTIPENIYQELYKYFNDKSDKELSDEAAKLKN WGHHEAATNIVKDYRYTYDKYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGID RGERNLIYVSVIDTCGNIVEQKSFNIVNGYDYQIKLKQQEGARQIARKEWKEIGKIKEIK EGYLSLVIHEISKMVIKYNAIIAMEDLSYGFKKGRFKVERQVYQKFETMLINKLNYLVF KDISITENGGLLKGYQLTYIPEKLKNVGHQCGCIFYVPAAYTSKIDPTTGFANVLNLSKV RNVDAIKSFFSNFNEISYSKKEALFKFSFDLDSLSKKGFSSFVKFSKSKWNVYTFGERIIKPKNKQGYREDKRINLTFEMKKLLNEYKVSFDLENNLIPNLTSANLKDTFWKELFFIFKTTLQLRNSVTNGKEDVLISPVKNAKGEFFVSGTHNKTLPQDCDANGAYCIALKGLYEIKQITENWKEDGKFSRDKLKISNKDWFDFIQNKRYL
[0164] SEQ ID NO:6 is Chimera 4:
[0165] MNNGTNNFQNFIGISSLQKTLRNALIPVGETASFVEDFKNEGLKRWSEDERRAVDYQKVKEIIDDYHRDFIEESLNYFPEQVSKDALEQAFHLYQKLKAAKVEEREKALKEWEALQKKLREKWKCFSDSNKARFSRIDKKELIKEDLINWLVAQNREDDIPTVETFNNFTTYFTGFHENRI<NIYSI<DDHATAISFRLIHENLPI<FFDNVISFNI<LI<EGFPELI<FDI<VI<EDLEVDYDLKHAFEIEYFVNFVTQAGIDQYNYLLGGKTLEDGTKKQGMNEQINLFKQQQTRDKARQIPKLIPLFKQILSERTESQSFIPYKFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNLDKIYIVSKFYESVSQKTYRDWETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSITEINELVSNYKLCPDDNIKAETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVLDVIMNAFHWCSVFMTEELVDKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQKPYSTKKIKLNFGIPTLADGWSKSKEYSNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDYKKMIYNLLPGPNKMIPKVFLSSKTGVETYKPSAYILEGYKQNKHLKSSKDFDITFCHDLIDYFKNCIAIHPEWKNFGFDFSDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKGQLYLFQIYNKDFSKKSSGNDNLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPIIHKKGSILVNRTYEAEEKDQFGNIQIVRKTIPENIYQELYKYFNDKSDKELSDEAAKLKNWGHHEAATNIVKDYRYTYDKYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGIDRGERNLIYVSVIDTCGNIVEQKSFNIVNGYDYQIKLKQQEGARQIARKEWKEIGKIKEIKEGYLSLVIHEISKMVIKYNAIIAMEDLSYGFKKGRFKVERQVYQKFETMLINKLNYLVFKDISITENGGLLKGYQLTYIPEKLKNVGHQCGCIFYVPAAYTSKIDPTTGFVNIFKFKDLTVDAKREFIKKFDSIRYDSEKNLFCFTFDYNNFITQNTVMSKSSWSVYTYGVRIKRRFVNGRFSNESDTIDITKDMEKTLEMTDINWRDGHDLRQDIIDYEIVQHIFEIFKLTVQMRNSLSELEDRDYDRLISPVLNENNIFYDSAKAGDALPKDADANGAYCIALKGLYEIKQITENWKEDGKFSRDKLKISNKDWFDFIQNKRYL
[0166] SEQ ID NO:7 is Chimera 5 (M44):
[0167] MTKTFDSEFFNLYSLQKTVRFELKPVGETASFVEDFKNEGLKRVVSEDERRAVDYQKVKEIIDDYHRDFIEESLNYFPEQVSKDALEQAFHLYQKLKAAKVEEREKALKEWEALQKKLREKVVKCFSDSNKARFSRIDKKELIKEDLINWLVAQNREDDIPTVETFNNFTTYFTGFHENRKNIYSKDDHATAISFRLIHENLPKFFDNVISFNKLKEGFPELKFDKVKEDLEVDYDLKHAFEIEYFVNFVTQAGIDQYNYLLGGKTLEDGTKKQGMNEQINLFKQQQTRDKARQIPKLIPLFKQILSERTESQSFIPYKFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNLDKIYIVSKFYESVSQKTYRDWETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSITEINELVSNYKLCPDDNIKAETYIHEISHILNNFEAQELKYNPEIHLVESELKAS ELKNVLDVIMNAFHWCSVFMTEELVDKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQ KPYSTKKIKLNFGIPTLADGWSKSKEYSNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDYKKMIYNLLPGPNKMIPKVFLSSKTGVETYKPSAYILEGYKQNKHLKSSKDFD ITFCHDLIDYFKNCIAIHPEWKNFGFDFSDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKGQLYLFQIYNKDFSKKSSGNDNLHTMYLKNLFSEENLKDIVLKLNGEAEIFFR KSSIKNPIIHKKGSILVNRTYEAEEKDQFGNIQIVRKTIPENIYQELYKYFNDKSDKELSDE AAKLKNVVGHHEAATNIVKDYRYTYDKYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGIDRGERNLIYVSVIDTCGNIVEQKSFNIVNGYDYQIKSamLKQQEGARQIARKEW KEIGKIKEIKEGYLSLVIHEISKMVIKYNAIIAMEDLSYGFKKGRFKVERQVYQKFETMLI NKLNYLVFKDISITENGGLLKGYQLTYIPEKLKNVGHQCGCIFYVPAAYTSKIDPTTGFV NIFKFKDLTVDAKREFIKKFDSIRYDSEKNLFCFTFDYNNFITQNTVMSKSSWSVYTYGV RIKRRFVNGRFSNESDTIDITKDMEKTLEMTDINWRDGHDLRQDIIDYEIVQHIFEIFKLT VQMRNSLSELEDRDYDRLISPVLNENNIFYDSAKAGDALPKDADANGAYCIALKGLYEI KQITENWKEDGKFSRDKLKISNKDWFDFIQNKRYL
[0168] SEQ ID NO:8 is Chimera 6:
[0169] MNNGTNNFQNFIGISSLQKTLRNALIPTETTQQFIVKNGIIKEDELRGENRQILKD IMDDYYRGFISETLSSIDDIDWTSLFEKMEIQLKNGDNKDTLIKEQAEKRKAIYKKFADD DRFKNMFSAKLISDILPEFVIHNNNYSASEKEEKTQVIKLFSRFATSFKDYFKNRANCFSA DDISSSSCHRIVNDNAEIFFSNALVYRRIVKNLSNDDINKISGDIKDSLKEMSLEEIYSYEK YGEFITQEGISFYNDICGKVNSFMNLYCQKNKENKNLYKLRKLHKQILCIADTSYEVPY KFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNLDKIYIVSKFYESVSQKTYRD WETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSITEINELVSNYKLCPDDNIKAETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVLDVIMNAFHWCSVFMTEELV DKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQKPYSTKKIKLNFGIPTLADGWSKSKEY SNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDYKKMIYNLLPGPNKMIPKVFL SSKTGVETYKPSAYILEGYKQNKHLKSSKDFDITFCHDLIDYFKNCIAIHPEWKNFGFDF SDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKGQLYLFQIYNKDFSKKSSGND NLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPIIHKKGSILVNRTYEAEEKDQF GNIQIVRKTIPENIYQELYKYFNDKSDKELSDEAAKLKNWGHHEAATNIVKDYRYTYDKYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGIDRGERNLIYVTVINQKGEILDSV SFNTVTNKSSKIEQTVDYEEKLAVREKERIEAKRSWDSISKIATLKEGYLSAIVHEICLLM IKHNAIWLENLNAGFKRIRGGLSEKSVYQKFEKMLINKLNYFVSKKESDWNKPSGLLN GLQLSDQFESFEKLGIQSGFIFYVPAAYTSKIDPTTGFANVLNLSKVRNVDAIKSFFSNFNEISYSKKEALFKFSFDLDSLSKKGFSSFVKFSKSKWNVYTFGERIIKPKNKQGYREDKRI NLTFEMKKLLNEYKVSFDLENNLIPNLTSANLKDTFWKELFFIFKTTLQLRNSVTNGKE DVLISPVKNAKGEFFVSGTHNKTLPQDCDANGAYHIALKGLMILERNNLVREEKDTKKI MAISNVDWFEYVQKRRGVL
[0170] SEQ ID NO:9 is Chimera 7:
[0171] MNNGTNNFQNFIGISSLQKTLRNALIPTETTQQFIVKNGIIKEDELRGENRQILKD IMDDYYRGFISETLSSIDDIDWTSLFEKMEIQLKNGDNKDTLIKEQAEKRKAIYKKFADD DRFKNMFSAKLISDILPEFVIHNNNYSASEKEEKTQVIKLFSRFATSFKDYFKNRANCFSA DDISSSSCHRIVNDNAEIFFSNALVYRRIVKNLSNDDINKISGDIKDSLKEMSLEEIYSYEK YGEFITQEGISFYNDICGKVNSFMNLYCQKNKENKNLYKLRKLHKQILCIADTSYEVPY KFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNLDKIYIVSKFYESVSQKTYRD WETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSITEINELVSNYKLCPDDNIKA ETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVLDVIMNAFHWCSVFMTEELV DKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQKPYSTKKIKLNFGIPTLADGWSKSKEY SNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDYKKMIYNLLPGPNKMIPKVFL SSKTGVETYKPSAYILEGYKQNKHLKSSKDFDITFCHDLIDYFKNCIAIHPEWKNFGFDF SDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKGQLYLFQIYNKDFSKKSSGND NLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPIIHKKGSILVNRTYEAEEKDQF GNIQIVRKTIPENIYQELYKYFNDKSDKELSDEAAKLKNWGHHEAATNIVKDYRYTYD KYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGIDRGERNLIYVSVIDTCGNIVEQK SFNIVNGYDYQIKLKQQEGARQIARKEWKEIGKIKEIKEGYLSLVIHEISKMVIKYNAIIA MEDLSYGFKKGRFKVERQVYQKFETMLINKLNYLVFKDISITENGGLLKGYQLTYIPEK LKNVGHQCGCIFYVPAAYTSKIDPLTGFVDPFVWKTIKNHESRKHFLEGFDFLHYDVKT GDFILHFKMNRNLSFQRGLPGFMPAWDIVFEKNETQFDAKGTPFIAGKRIVPVIENHRFT GRYRDLYPANELIALLEEKGIVFRDGSNILPKLLENDDSHAIDTMVALIRSVLQMRNSNA ATGEDYINSPVRDLNGVCFDSRFQNPEWPMDADANGAYHIALKGQLLLNHLKESKDLK LQNGISNQDWLAYIQELRN
[0172] SEQ ID NO:10 is Chimera 8:
[0173] MNNGTNNFQNFIGISSLQKTLRNALIPTETTQQFIVKNGIIKEDELRGENRQILKD IMDDYYRGFISETLSSIDDIDWTSLFEKMEIQLKNGDNKDTLIKEQAEKRKAIYKKFADD DRFKNMFSAKLISDILPEFVIHNNNYSASEKEEKTQVIKLFSRFATSFKDYFKNRANCFSA DDISSSSCHRIVNDNAEIFFSNALVYRRIVKNLSNDDINKISGDIKDSLKEMSLEEIYSYEK YGEFITQEGISFYNDICGKVNSFMNLYCQKNKENKNLYKLRKLHKQILCIADTSYEVPY KFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNLDKIYIVSKFYESVSQKTYRDWETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSITEINELVSNYKLCPDDNIKA ETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVLDVIMNAFHWCSVFMTEELV DKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQKPYSTKKIKLNFGIPTLADGWSKSKEY SNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDYKKMIYNLLPGPNKMIPKVFL SSKTGVETYKPSAYILEGYKQNKHLKSSKDFDITFCHDLIDYFKNCIAIHPEWKNFGFDF SDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKGQLYLFQIYNKDFSKKSSGND NLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPIIHKKGSILVNRTYEAEEKDQF GNIQIVRKTIPENIYQELYKYFNDKSDKELSDEAAKLKNWGHHEAATNIVKDYRYTYD KYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGIDRGERNLIYITVIDSTGKILEQRS LNTIQQFDYQKKLDNREKERVAARQAWSVVGTIKDLKQGYLSQVIHEIVDLMIHYQAV WLENLNFGFKSKRTGIAEKAVYQQFEKMLIDKLNCLVLKDYPAEKVGGVLNPYQLTD QFTSFAKMGTQSGFLFYVPAPYTSKIDPLTGFVDPFVWKTIKNHESRKHFLEGFDFLHY DVKTGDFILHFKMNRNLSFQRGLPGFMPAWDIVFEKNETQFDAKGTPFIAGKRIVPVIEN HRFTGRYRDLYPANELIALLEEKGIVFRDGSNILPKLLENDDSHAIDTMVALIRSVLQMR NSNAATGEDYINSPVRDLNGVCFDSRFQNPEWPMDADANGAYHIALKGQLLLNHLKES KDLKLQNGISNQDWLAYIQELRN
[0174] SEQ ID NO: 11 is Chimera 9:
[0175] MNNGTNNFQNFIGISSLQKTLRNALIPTETTQQFIVKNGIIKEDELRGENRQILKD IMDDYYRGFISETLSSIDDIDWTSLFEKMEIQLKNGDNKDTLIKEQAEKRKAIYKKFADD DRFKNMFSAKLISDILPEFVIHNNNYSASEKEEKTQVIKLFSRFATSFKDYFKNRANCFSA DDISSSSCHRIVNDNAEIFFSNALVYRRIVKNLSNDDINKISGDIKDSLKEMSLEEIYSYEK YGEFITQEGISFYNDICGKVNSFMNLYCQKNKENKNLYKLRKLHKQILCIADTSYEVPY KFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNLDKIYIVSKFYESVSQKTYRD WETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSITEINELVSNYKLCPDDNIKA ETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVLDVIMNAFHWCSVFMTEELV DKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQKPYSTKKIKLNFGIPTLADGWSKSKEY SNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDYKKMIYNLLPGPNKMIPKVFL SSKTGVETYKPSAYILEGYKQNKHLKSSKDFDITFCHDLIDYFKNCIAIHPEWKNFGFDF SDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKGQLYLFQIYNKDFSKKSSGND NLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPIIHKKGSILVNRTYEAEEKDQF GNIQIVRKTIPENIYQELYKYFNDKSDKELSDEAAKLKNWGHHEAATNIVKDYRYTYD KYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGIDRGERNLIYITVIDSTGKILEQRS LNTIQQFDYQKKLDNREKERVAARQAWSVVGTIKDLKQGYLSQVIHEIVDLMIHYQAV WLENLNFGFKSKRTGIAEKAVYQQFEKMLIDKLNCLVLKDYPAEKVGGVLNPYQLTDQFTSFAKMGTQSGFLFYVPAPYTSKIDPLTGFVDPFVWKTIKNHESRKHFLEGFDFLHY DVKTGDFILHFKMNRNLSFQRGLPGFMPAWDIVFEKNETQFDAKGTPFIAGKRIVPVIEN HRFTGRYRDLYPANELIALLEEKGIVFRDGSNILPKLLENDDSHAIDTMVALIRSVLQMR NSNAATGEDYINSPVRDLNGVCFDSRFQNPEWPMDADANGAYCIALKGLYEIKQITEN WKEDGKFSRDKLKISNKDWFDFIQNKRYL
[0176] SEQ ID NO: 12 is a guide scaffold derived from Eubacterium rectale
[0177] GTCTGGCCCCAAATTTTAATTTCTACTGTTGTAGAT
[0178] SEQ ID NO: 13 is a guide scaffold derived from Thiomicrospira sp. XS5:
[0179] CTCTAGCAGGCCTGGCAAATTTCTACTGTTGTAGAT
[0180] SEQ ID NO: 14 is a guide expression (scaffold SEQ ID NO: 12) and tdTomato reporter donor plasmid (FIG. 17):
[0181] tttctactttcccatgattccttcatatttgcatatacgatacaaggctgttagagagataattggaattaatttgactgtaaacacaa agatattagtacaaaatacgtgacgtagaaagtaataatttcttgggtagtttgcagttttaaaattatgttttaaaatggactatcatatgcttacc gtaacttgaaagtatttcgatttcttggctttatatatcttGTGGAAAGGACGAAACACCGTCTGGCCCCAAATT TTAATTTCTACTGTTGTAGATCCTCTCAGGCATGGAGTCCTTTTTTTggggtctgacgctcagtg gaacgaaaactcacgttaagggattttggtcatgagattatcaaaaaggatcttcacctagatccttttaaattaaaaatgaagttttaaatcaat ctaaagtatatatgagtaaacttggtctgacagttaccaatgcttaatcagtgaggcacctatctcagcgatctgtctatttcgttcatccatagtt gcctgactccccgtcgtgtagataactacgatacgggagggcttaccatctggccccagtgctgcaatgataccgcgagacccacgctca ccggctccagatttatcagcaataaaccagccagccggaagggccgagcgcagaagtggtcctgcaactttatccgcctccatccagtcta ttaattgttgccgggaagctagagtaagtagttcgccagttaatagtttgcgcaacgttgttgccattgctacaggcatcgtggtgtcacgctc gtcgtttggtatggcttcattcagctccggttcccaacgatcaaggcgagttacatgatcccccatgttgtgcaaaaaagcggttagctccttc ggtcctccgatcgttgtcagaagtaagttggccgcagtgttatcactcatggttatggcagcactgcataattctcttactgtcatgccatccgt aagatgcttttctgtgactggtgagtactcaaccaagtcattctgagaatagtgtatgcggcgaccgagttgctcttgcccggcgtcaatacg ggataataccgcgccacatagcagaactttaaaagtgctcatcattggaaaacgttcttcggggcgaaaactctcaaggatcttaccgctgtt gagatccagttcgatgtaacccactcgtgcacccaactgatcttcagcatcttttactttcaccagcgtttctgggtgagcaaaaacaggaag gcaaaatgccgcaaaaaagggaataagggcgacacggaaatgttgaatactcatactcttcctttttcaatattattgaagcatttatcagggtt attgtctcatgagcggatacatatttgaatgtatttagaaaaataaacaaataggggttccgcgcacatttccccgaaaagtgccactctagag ccacctgacgttaacccggggaCATGGAGAAAATCTGGCACCACACCTTCTACAATGAGCTGCG TGTGGCTCCCGAGGAGCACCCCGTGCTGCTGACCGAGGCCCCCCTGAACCCCAAGG CCAACCGCGAGAAGATGACCCAGGTGAGTGGCCCGCTACCTCTTCTGGTGGCCGCC TCCCTCCTTCCTGGCCTCCCGGAGCTGCGCCCTTTCTCACTGGTTCTCTCTTCTGCCG TTTTCCGTAGGACTCTCTTCTCTGACCTGAGTCTCCTTTGGAACTCTGCAGGTTCTAT TTGCTTTTTCCCAGATGAGCTCTTTTTCTGGTGTTTGTCTCTCTGACTAGGTGTCTAA GACAGTGTTGTGGGTGTAGGTACTAACACTGGCTCGTGTGACAAGGCCATGAGGCTGGTGTAAAGCGGCCTTGGAGTGTGTATTAAGTAGGcGCACAGTAGGTCTGAACAGA CTCCCCATCCCAAGACCCCAGCACACTTAGCCGTGTTCTTTGCACTTTCTGCATGTCC CCCGTCTGGCCTGGCTGTCCCCAGTGGCTTCCCCAGTGTGACATGGTGcATCTCTGCC TTACAGATCATGTTTGAGACCTTCAACACCCCAGCCATGTACGTTGCTATCCAGGCT GTGCTATCCCTGTACGCCTCTGGCCGTACCACTGGCATCGTGATGGACTCCGGTGAC GGGGTCACCCACACTGTGCCCATCTACGAGGGGTATGCCCTCCCCCATGCCATCCTG CGTCTGGACCTGGCTGGCCGGGACCTGACTGACTACCTCATGAAGATCCTCACCGA GCGCGGCTACAGCTTCACCACCACGGCCGAGCGGGAAATCGTGCGTGACATTAAGG AGAAGCTGTGCTACGTCGCCCTGGACTTCGAGCAAGAGATGGCCACGGCTGCTTCC AGCTCCTCCCTGGAGAAGAGCTACGAGCTGCCTGACGGCCAGGTCATCACCATTGG CAATGAGCGGTTCCGCTGCCCTGAGGCACTCTTCCAGCCTTCCTTCCTGGGCATGGA AAGCTGTGGCATACACGAGACCACTTTCAATTCAATCATGAAGTGCGACGTCGACA TTCGCAAAGATCTTTACGCTAACACAGTGCTCTCTGGGGGAACGACCATGTATCCGG GGATAGCCGATCGGATGCAGAAGGAGATTACAGCATTGGCTCCCAGTACTATGAAG ATTAAGATCATCGCGCCACCCGAGAGAAAATATAGTGTTTGGATCGGTGGCTCTAT CCTGGCCTCACTGAGCACCTTTCAACAGATGTGGATTTCCAAACAGGAATACGACG AATCTGGCCCTAGCATCGTGCATAGGAAGTGCTTCGGATCCGGtgaAggGagCggATCTc tCctAacAtgTggAgaTgtCgaAgagaaTccTggAccAgtgagtaagggcgaggaagtgatcaaagagttcatgcggtttaa ggtgagaatggaaggaagcatgaacggccacgagttcgaaattgagggagaaggagagggacggccctacgagggcacccagacag ccaagctgaaagtgacaaagggcgggcctctgccattcgcttgggacatcctgagcccacagtttatgtacggctccaaggcctatgtgaa acatccagctgacattcccgattataagaaactgagcttccccgaggggtttaagtgggaaagagtgatgaacttcgaggacggaggcctg gtgactgtgacccaggacagctccctgcaggatgggaccctgatctacaaggtgaaaatgagagggacaaattttccccctgatggacct gtgatgcagaagaaaactatgggatgggaggcctccaccgaaaggctgtatccacgcgacggggtgctgaaaggagaaatccaccagg ctctgaagctgaaagatgggggacattacctggtggagttcaagacaatctacatggccaagaaacctgtgcagctgccaggctactatta cgtggacacaaaactggatatcacttcacacaacgaggactacactattgtggagcagtatgaacggagcgaggggagacaccatctgtt cctgggccatgggactggaagtaccggctcagggtctagtggaaccgcctcaagcgaggataacaatatggctgtgatcaaagagttcat gaggtttaaggtgcgcatggagggcagcatgaatgggcacgaatttgagattgaaggagagggcgaagggaggccttacgagggcac acagactgccaagctgaaagtgaccaagggaggaccactgcctttcgcttgggatatcctgtctcctcagtttatgtacggaagtaaggcct atgtcaagcatcccgctgacattcctgattacaagaaactgtctttcccagagggctttaagtgggagagagtgatgaattttgaagatggag gcctggtgaccgtgacacaggactcctctctgcaggatggcactctgatctacaaagtcaaaatgcgcggcaccaattttccacccgatgg gcccgtgatgcagaagaaaacaatggggtgggaggccagcactgaacggctgtatcctagagacggagtgctgaagggcgaaatccac caggccctgaagctgaaagacggcggccactacctggtggagttcaaaaccatctacatggccaagaaaccagtgcagctgcccggcta ttactatgtggacaccaagctggatatcacatcccacaatgaagactacaccattgtggaacagtatgagaggtctgaaggacgccaccatc tgtttctgtacggcatggatgagctgtataagGGCATGGAGTCCTGTGGCATCCACGAAACTACCTTCAACTCCATCATGAAGTGTGACGTGGACATCCGCAAAGACCTGTACGCCAACACAGTG CTGTCTGGCGGCACCACCATGTACCCTGGCATTGCCGACAGGATGCAGAAGGAGAT CACTGCCCTGGCACCCAGCACAATGAAGATCAAGGTGGGTGTCTTTCCTGCCTGAG CTGACCTGGGCAGGTCaGCTGTGGGGTCCTGTGGTGTGTGGGGAGCTGTCACATCCA GGGTCCTCACTGCCTGTCCCCTTCCCTCCTCAGATCATTGCTCCTCCTGAGCGCAAGT ACTCCGTGTGGATCGGCGGCTCCATCCTGGCCTCGCTGTCCACCTTCCAGCAGATGT GGATCAGCAAGCAGGAGTATGACGAGTCCGGCCCCTCCATCGTCCACCGCAAATGC TTCTAGGCGGACTATGACTTAGTTGCGTTACACCCTTTCTTGACAAAACCTAACTTG CGCAGAAAACAAGATGAGATTGGCATGGCtttatttgttttttttgttttgttttggttttttttttttggCTTGACT CAGGATTTAAAAACTGGAACGGTGAAGGTGACAGCAGTCGGTTGGAGCGAGCATCC CCCAAAGTTCACAATGTGGCCGAGGACTTTGATTGCACATTGTTGTTTTTTTAATAG TCATTCCAAATATGAGATGCGTTGTTACAGGAAGTCCCTTGCCATCCTAAAAGCCAC CCCACTTCTCTCTAAGGAGAATGGCCCAGTCCTCTCCCAAGTCCACACAGGGGAGG TGATAGCATTGCTTTCGTGTAAATTATGTAATGCAAAAtttttttaatcttcgccttaatacttttttattttgttt tattttgAATGATGAGCCTTCGTGcccccccttcccccttttttgtcccccAACTTGAGATGTATGAAGGC TTTTGGTCTCCCTGGGAGTGGGTGGAGGCAGCCAGGGCTTACCTGTACACTGACTgga tccccgggttaacgtcaggtggccggccgttgctggcgtttttccataggctccgcccccctgacgagcatcacaaaaatcgacgctcaag tcagaggtggcgaaacccgacaggactataaagataccaggcgtttccccctggaagctccctcgtgcgctctcctgttccgaccctgccg cttaccggatacctgtccgcctttctcccttcgggaagcgtggcgctttctcatagctcacgctgtaggtatctcagttcggtgtaggtcgttcg ctccaagctgggctgtgtgcacgaaccccccgttcagcccgaccgctgcgccttatccggtaactatcgtcttgagtccaacccggtaaga cacgacttatcgccactggcagcagccactggtaacaggattagcagagcgaggtatgtaggcggtgctacagagttcttgaagtggtgg cctaactacggctacactagaagaacagtatttggtatctgcgctctgctgaagccagttaccttcggaaaaagagttggtagctcttgatccg gcaaacaaaccaccgctggtagcggtggtttttttgtttgcaagcagcagattacgcgcagaaaaaaaggatctcaagaagatcctttgatct
[0182] SEQ ID NO: 15 is a guide expression (scaffold SEQ ID NO: 13) and tdTomato reporter donor plasmid (FIG. 18):
[0183] tttctactttcccatgattccttcatatttgcatatacgatacaaggctgttagagagataattggaattaatttgactgtaaacacaa agatattagtacaaaatacgtgacgtagaaagtaataatttcttgggtagtttgcagttttaaaattatgttttaaaatggactatcatatgcttacc gtaacttgaaagtatttcgatttcttggctttatatatcttGTGGAAAGGACGAAACACCGCTCTAGCAGGCCT GGCAAATTTCTACTGTTGTAGATCCTCTCAGGCATGGAGTCCTTTTTTTggggtctgacgctc agtggaacgaaaactcacgttaagggattttggtcatgagattatcaaaaaggatcttcacctagatccttttaaattaaaaatgaagttttaaat caatctaaagtatatatgagtaaacttggtctgacagttaccaatgcttaatcagtgaggcacctatctcagcgatctgtctatttcgttcatccat agttgcctgactccccgtcgtgtagataactacgatacgggagggcttaccatctggccccagtgctgcaatgataccgcgagacccacgc tcaccggctccagatttatcagcaataaaccagccagccggaagggccgagcgcagaagtggtcctgcaactttatccgcctccatccagt ctattaattgttgccgggaagctagagtaagtagttcgccagttaatagtttgcgcaacgttgttgccattgctacaggcatcgtggtgtcacgctcgtcgtttggtatggcttcattcagctccggttcccaacgatcaaggcgagttacatgatcccccatgttgtgcaaaaaagcggttagctcc ttcggtcctccgatcgttgtcagaagtaagttggccgcagtgttatcactcatggttatggcagcactgcataattctcttactgtcatgccatcc gtaagatgcttttctgtgactggtgagtactcaaccaagtcattctgagaatagtgtatgcggcgaccgagttgctcttgcccggcgtcaatac gggataataccgcgccacatagcagaactttaaaagtgctcatcattggaaaacgttcttcggggcgaaaactctcaaggatcttaccgctgt tgagatccagttcgatgtaacccactcgtgcacccaactgatcttcagcatcttttactttcaccagcgtttctgggtgagcaaaaacaggaag gcaaaatgccgcaaaaaagggaataagggcgacacggaaatgttgaatactcatactcttcctttttcaatattattgaagcatttatcagggtt attgtctcatgagcggatacatatttgaatgtatttagaaaaataaacaaataggggttccgcgcacatttccccgaaaagtgccactctagag ccacctgacgttaacccggggaCATGGAGAAAATCTGGCACCACACCTTCTACAATGAGCTGCG TGTGGCTCCCGAGGAGCACCCCGTGCTGCTGACCGAGGCCCCCCTGAACCCCAAGG CCAACCGCGAGAAGATGACCCAGGTGAGTGGCCCGCTACCTCTTCTGGTGGCCGCC TCCCTCCTTCCTGGCCTCCCGGAGCTGCGCCCTTTCTCACTGGTTCTCTCTTCTGCCG TTTTCCGTAGGACTCTCTTCTCTGACCTGAGTCTCCTTTGGAACTCTGCAGGTTCTAT TTGCTTTTTCCCAGATGAGCTCTTTTTCTGGTGTTTGTCTCTCTGACTAGGTGTCTAA GACAGTGTTGTGGGTGTAGGTACTAACACTGGCTCGTGTGACAAGGCCATGAGGCT GGTGTAAAGCGGCCTTGGAGTGTGTATTAAGTAGGcGCACAGTAGGTCTGAACAGA CTCCCCATCCCAAGACCCCAGCACACTTAGCCGTGTTCTTTGCACTTTCTGCATGTCC CCCGTCTGGCCTGGCTGTCCCCAGTGGCTTCCCCAGTGTGACATGGTGcATCTCTGCC TTACAGATCATGTTTGAGACCTTCAACACCCCAGCCATGTACGTTGCTATCCAGGCT GTGCTATCCCTGTACGCCTCTGGCCGTACCACTGGCATCGTGATGGACTCCGGTGAC GGGGTCACCCACACTGTGCCCATCTACGAGGGGTATGCCCTCCCCCATGCCATCCTG CGTCTGGACCTGGCTGGCCGGGACCTGACTGACTACCTCATGAAGATCCTCACCGA GCGCGGCTACAGCTTCACCACCACGGCCGAGCGGGAAATCGTGCGTGACATTAAGG AGAAGCTGTGCTACGTCGCCCTGGACTTCGAGCAAGAGATGGCCACGGCTGCTTCC AGCTCCTCCCTGGAGAAGAGCTACGAGCTGCCTGACGGCCAGGTCATCACCATTGG CAATGAGCGGTTCCGCTGCCCTGAGGCACTCTTCCAGCCTTCCTTCCTGGGCATGGA AAGCTGTGGCATACACGAGACCACTTTCAATTCAATCATGAAGTGCGACGTCGACA TTCGCAAAGATCTTTACGCTAACACAGTGCTCTCTGGGGGAACGACCATGTATCCGG GGATAGCCGATCGGATGCAGAAGGAGATTACAGCATTGGCTCCCAGTACTATGAAG ATTAAGATCATCGCGCCACCCGAGAGAAAATATAGTGTTTGGATCGGTGGCTCTAT CCTGGCCTCACTGAGCACCTTTCAACAGATGTGGATTTCCAAACAGGAATACGACG AATCTGGCCCTAGCATCGTGCATAGGAAGTGCTTCGGATCCGGtgaAggGagCggATCTc tCctAacAtgTggAgaTgtCgaAgagaaTccTggAccAgtgagtaagggcgaggaagtgatcaaagagttcatgcggtttaa ggtgagaatggaaggaagcatgaacggccacgagttcgaaattgagggagaaggagagggacggccctacgagggcacccagacag ccaagctgaaagtgacaaagggcgggcctctgccattcgcttgggacatcctgagcccacagtttatgtacggctccaaggcctatgtgaaacatccagctgacattcccgattataagaaactgagcttccccgaggggtttaagtgggaaagagtgatgaacttcgaggacggaggcctg gtgactgtgacccaggacagctccctgcaggatgggaccctgatctacaaggtgaaaatgagagggacaaattttccccctgatggacct gtgatgcagaagaaaactatgggatgggaggcctccaccgaaaggctgtatccacgcgacggggtgctgaaaggagaaatccaccagg ctctgaagctgaaagatgggggacattacctggtggagttcaagacaatctacatggccaagaaacctgtgcagctgccaggctactatta cgtggacacaaaactggatatcacttcacacaacgaggactacactattgtggagcagtatgaacggagcgaggggagacaccatctgtt cctgggccatgggactggaagtaccggctcagggtctagtggaaccgcctcaagcgaggataacaatatggctgtgatcaaagagttcat gaggtttaaggtgcgcatggagggcagcatgaatgggcacgaatttgagattgaaggagagggcgaagggaggccttacgagggcac acagactgccaagctgaaagtgaccaagggaggaccactgcctttcgcttgggatatcctgtctcctcagtttatgtacggaagtaaggcct atgtcaagcatcccgctgacattcctgattacaagaaactgtctttcccagagggctttaagtgggagagagtgatgaattttgaagatggag gcctggtgaccgtgacacaggactcctctctgcaggatggcactctgatctacaaagtcaaaatgcgcggcaccaattttccacccgatgg gcccgtgatgcagaagaaaacaatggggtgggaggccagcactgaacggctgtatcctagagacggagtgctgaagggcgaaatccac caggccctgaagctgaaagacggcggccactacctggtggagttcaaaaccatctacatggccaagaaaccagtgcagctgcccggcta ttactatgtggacaccaagctggatatcacatcccacaatgaagactacaccattgtggaacagtatgagaggtctgaaggacgccaccatc tgtttctgtacggcatggatgagctgtataagGGCATGGAGTCCTGTGGCATCCACGAAACTACCTTCA ACTCCATCATGAAGTGTGACGTGGACATCCGCAAAGACCTGTACGCCAACACAGTG CTGTCTGGCGGCACCACCATGTACCCTGGCATTGCCGACAGGATGCAGAAGGAGAT CACTGCCCTGGCACCCAGCACAATGAAGATCAAGGTGGGTGTCTTTCCTGCCTGAG CTGACCTGGGCAGGTCaGCTGTGGGGTCCTGTGGTGTGTGGGGAGCTGTCACATCCA GGGTCCTCACTGCCTGTCCCCTTCCCTCCTCAGATCATTGCTCCTCCTGAGCGCAAGT ACTCCGTGTGGATCGGCGGCTCCATCCTGGCCTCGCTGTCCACCTTCCAGCAGATGT GGATCAGCAAGCAGGAGTATGACGAGTCCGGCCCCTCCATCGTCCACCGCAAATGC TTCTAGGCGGACTATGACTTAGTTGCGTTACACCCTTTCTTGACAAAACCTAACTTG CGCAGAAAACAAGATGAGATTGGCATGGCtttatttgttttttttgttttgttttggttttttttttttggCTTGACT CAGGATTTAAAAACTGGAACGGTGAAGGTGACAGCAGTCGGTTGGAGCGAGCATCC CCCAAAGTTCACAATGTGGCCGAGGACTTTGATTGCACATTGTTGTTTTTTTAATAG TCATTCCAAATATGAGATGCGTTGTTACAGGAAGTCCCTTGCCATCCTAAAAGCCAC CCCACTTCTCTCTAAGGAGAATGGCCCAGTCCTCTCCCAAGTCCACACAGGGGAGG TGATAGCATTGCTTTCGTGTAAATTATGTAATGCAAAAtttttttaatcttcgccttaatacttttttattttgttt tattttgAATGATGAGCCTTCGTGcccccccttcccccttttttgtcccccAACTTGAGATGTATGAAGGC TTTTGGTCTCCCTGGGAGTGGGTGGAGGCAGCCAGGGCTTACCTGTACACTGACTgga tccccgggttaacgtcaggtggccggccgttgctggcgtttttccataggctccgcccccctgacgagcatcacaaaaatcgacgctcaag tcagaggtggcgaaacccgacaggactataaagataccaggcgtttccccctggaagctccctcgtgcgctctcctgttccgaccctgccg cttaccggatacctgtccgcctttctcccttcgggaagcgtggcgctttctcatagctcacgctgtaggtatctcagttcggtgtaggtcgttcg ctccaagctgggctgtgtgcacgaaccccccgttcagcccgaccgctgcgccttatccggtaactatcgtcttgagtccaacccggtaagacacgacttatcgccactggcagcagccactggtaacaggattagcagagcgaggtatgtaggcggtgctacagagttcttgaagtggtgg cctaactacggctacactagaagaacagtatttggtatctgcgctctgctgaagccagttaccttcggaaaaagagttggtagctcttgatccg gcaaacaaaccaccgctggtagcggtggtttttttgtttgcaagcagcagattacgcgcagaaaaaaaggatctcaagaagatcctttgatct
[0184] SEQ ID NO: 16 is a guide expression (scaffold SEQ ID NO: 12) and luciferase reporter donor plasmid (FIG. 19):
[0185] tttctactttcccatgattccttcatatttgcatatacgatacaaggctgttagagagataattggaattaatttgactgtaaacacaa agatattagtacaaaatacgtgacgtagaaagtaataatttcttgggtagtttgcagttttaaaattatgttttaaaatggactatcatatgcttacc gtaacttgaaagtatttcgatttcttggctttatatatcttGTGGAAAGGACGAAACACCGTCTGGCCCCAAATT TTAATTTCTACTGTTGTAGATCCTCTCAGGCATGGAGTCCTTTTTTTggggtctgacgctcagtg gaacgaaaactcacgttaagggattttggtcatgagattatcaaaaaggatcttcacctagatccttttaaattaaaaatgaagttttaaatcaat ctaaagtatatatgagtaaacttggtctgacagttaccaatgcttaatcagtgaggcacctatctcagcgatctgtctatttcgttcatccatagtt gcctgactccccgtcgtgtagataactacgatacgggagggcttaccatctggccccagtgctgcaatgataccgcgagacccacgctca ccggctccagatttatcagcaataaaccagccagccggaagggccgagcgcagaagtggtcctgcaactttatccgcctccatccagtcta ttaattgttgccgggaagctagagtaagtagttcgccagttaatagtttgcgcaacgttgttgccattgctacaggcatcgtggtgtcacgctc gtcgtttggtatggcttcattcagctccggttcccaacgatcaaggcgagttacatgatcccccatgttgtgcaaaaaagcggttagctccttc ggtcctccgatcgttgtcagaagtaagttggccgcagtgttatcactcatggttatggcagcactgcataattctcttactgtcatgccatccgt aagatgcttttctgtgactggtgagtactcaaccaagtcattctgagaatagtgtatgcggcgaccgagttgctcttgcccggcgtcaatacg ggataataccgcgccacatagcagaactttaaaagtgctcatcattggaaaacgttcttcggggcgaaaactctcaaggatcttaccgctgtt gagatccagttcgatgtaacccactcgtgcacccaactgatcttcagcatcttttactttcaccagcgtttctgggtgagcaaaaacaggaag gcaaaatgccgcaaaaaagggaataagggcgacacggaaatgttgaatactcatactcttcctttttcaatattattgaagcatttatcagggtt attgtctcatgagcggatacatatttgaatgtatttagaaaaataaacaaataggggttccgcgcacatttccccgaaaagtgccactctagag ccacctgacgttaacccggggaCATGGAGAAAATCTGGCACCACACCTTCTACAATGAGCTGCG TGTGGCTCCCGAGGAGCACCCCGTGCTGCTGACCGAGGCCCCCCTGAACCCCAAGG CCAACCGCGAGAAGATGACCCAGGTGAGTGGCCCGCTACCTCTTCTGGTGGCCGCC TCCCTCCTTCCTGGCCTCCCGGAGCTGCGCCCTTTCTCACTGGTTCTCTCTTCTGCCG TTTTCCGTAGGACTCTCTTCTCTGACCTGAGTCTCCTTTGGAACTCTGCAGGTTCTAT TTGCTTTTTCCCAGATGAGCTCTTTTTCTGGTGTTTGTCTCTCTGACTAGGTGTCTAA GACAGTGTTGTGGGTGTAGGTACTAACACTGGCTCGTGTGACAAGGCCATGAGGCT GGTGTAAAGCGGCCTTGGAGTGTGTATTAAGTAGGcGCACAGTAGGTCTGAACAGA CTCCCCATCCCAAGACCCCAGCACACTTAGCCGTGTTCTTTGCACTTTCTGCATGTCC CCCGTCTGGCCTGGCTGTCCCCAGTGGCTTCCCCAGTGTGACATGGTGcATCTCTGCC TTACAGATCATGTTTGAGACCTTCAACACCCCAGCCATGTACGTTGCTATCCAGGCT GTGCTATCCCTGTACGCCTCTGGCCGTACCACTGGCATCGTGATGGACTCCGGTGAC GGGGTCACCCACACTGTGCCCATCTACGAGGGGTATGCCCTCCCCCATGCCATCCTGCGTCTGGACCTGGCTGGCCGGGACCTGACTGACTACCTCATGAAGATCCTCACCGAGCGCGGCTACAGCTTCACCACCACGGCCGAGCGGGAAATCGTGCGTGACATTAAGGAGAAGCTGTGCTACGTCGCCCTGGACTTCGAGCAAGAGATGGCCACGGCTGCTTCCAGCTCCTCCCTGGAGAAGAGCTACGAGCTGCCTGACGGCCAGGTCATCACCATTGGCAATGAGCGGTTCCGCTGCCCTGAGGCACTCTTCCAGCCTTCCTTCCTGGGCATGGAAAGCTGTGGCATACACGAGACCACTTTCAATTCAATCATGAAGTGCGACGTCGACATTCGCAAAGATCTTTACGCTAACACAGTGCTCTCTGGGGGAACGACCATGTATCCGGGGATAGCCGATCGGATGCAGAAGGAGATTACAGCATTGGCTCCCAGTACTATGAAGATTAAGATCATCGCGCCACCCGAGAGAAAATATAGTGTTTGGATCGGTGGCTCTATCCTGGCCTCACTGAGCACCTTTCAACAGATGTGGATTTCCAAACAGGAATACGACGAATCTGGCCCTAGCATCGTGCATAGGAAGTGCTTCGGATCCGGtgaAggGagCggATCTc tCctAacAtgTggAgaTgtCgaAgagaaTccTggAccAATGGAGATGGAGAAGGAGGAGAACGTGGTGTACGGACCTCTCCCCTTTTACCCTATTGAAGAAGGGTCCGCCGGTATCCAGTTGCACAAGTATATGCATCAGTACGCCAAACTGGGAGCAATAGCATTTTCAAATGCTCTAACAGGTGTCGATATCTCATATCAGGAGTATTTTGACATTACATGCAGGCTGGCCGAGGCAATGAAAAACTTCGGGATGAAACCTGAAGAACATATCGCTCTCTGCTCCGAGAACTGTGAAGAGTTCTTTATACCCGTTTTGGCTGGACTTTACATCGGTGTGGCGGTGGCCCCCACGAATGAGATATACACACTGAGGGAACTCAATCATAGTCTGGGAATCGCCCAGCCCACGATCGTGTTCAGTTCTCGGAAGGGACTCCCAAAGGTGCTCGAGGTCCAAAAAACTGTCACCTGTATTAAGAAGATTGTTATCCTGGACAGCAAAGTTAACTTCGGGGGTCACGATTGCATGGAGACTTTTATCAAGAAGCACGTAGAGCTCGGATTTCAGCCGTCCTCTTTCGTGCCAATTGATGTAAAGAACCGAAAGCAACACGTTGCTTTACTGATGAATTCTAGCGGATCTACCGGATTGCCGAAAGGGGTCAGGATAACCCACGAAGGCGCGGTCACAAGATTCTCACATGCAAAGGACCCAATATACGGCAATCAAGTCTCCCCTGGAACAGCTATTTTAACCGTGGTGCCCTTTCACCATGGTTTTGGCATGTTTACTACACTGGGCTATTTTGCGTGTGGCTACCGTGTCGTGATGCTGACAAAATTTGACGAAGAACTGTTCCTCCGGACCCTGCAGGACTACAAGTGCACCAGCGTCATTCTGGTGCCCACCTTATTTGCCATCCTGAACAAATCCGAACTGATTGATAAATTCGATCTCTCCAATCTTACTGAAATCGCCAGCGGTGGAGCTCCCCTCGCCAAGGAGGTGGGCGAGGCCGTCGCACGGAGATTCAATCTGCCTGGCGTTAGACAGGGGTATGGACTGACTGAGACAACCAGCGCCTTCATTATTACGCCTCGCGGGGATGACAAACCTGGGGCGTCGGGCAAGGTTGCCCCTCTGTTCAAAGTGAAGGTGATCGACTTGGACACGAAAAAAACCTTAGGGGTCAACCGCCGAGGCGAAATCTGTGTAAAAGGGCCAAGCCTTATGCTAGGTTATTCAAACAACCCTGAGGCAACCAGAGAAACTATTGACGAAGAGGGCTGGCTGCACACCGGGGACCTTGGGTATTACGATGAAGACGAACACTTCTTCATAGTGGATAGGCTGAAG TCTCTGATCAAGTACCAGGGCTACCAGGTGCCGCCCGCTGAGCTTGAGAGTGTCCTC CTACAACACCCAAATATTTTCGATGCTGGGGTAGCCGGCGTGCCAGATCCCGATGC AGGCGAGCTGCCCGGCGCCGTAGTGGTCATGGAAAAGGGTAAAACTATGACTGAG AAGGAGATCGTGGACTATGTTAACAGTCAGGTTGTTAATCATAAGCGCTTGCGTGG CGGAGTACGGTTCGTGGACGAGGTGCCAAAAGGCCTTACAGGCAAGATCGACGCA AAGGTGATCCGCGAGATTTTGAAAAAACCACAGGCTAAGATGGGCATGGAGTCCTG TGGCATCCACGAAACTACCTTCAACTCCATCATGAAGTGTGACGTGGACATCCGCAAAGACCTGTACGCCAACACAGTGCTGTCTGGCGGCACCACCATGTACCCTGGCATT GCCGACAGGATGCAGAAGGAGATCACTGCCCTGGCACCCAGCACAATGAAGATCA AGGTGGGTGTCTTTCCTGCCTGAGCTGACCTGGGCAGGTCaGCTGTGGGGTCCTGTG GTGTGTGGGGAGCTGTCACATCCAGGGTCCTCACTGCCTGTCCCCTTCCCTCCTCAG ATCATTGCTCCTCCTGAGCGCAAGTACTCCGTGTGGATCGGCGGCTCCATCCTGGCC TCGCTGTCCACCTTCCAGCAGATGTGGATCAGCAAGCAGGAGTATGACGAGTCCGG CCCCTCCATCGTCCACCGCAAATGCTTCTAGGCGGACTATGACTTAGTTGCGTTACA CCCTTTCTTGACAAAACCTAACTTGCGCAGAAAACAAGATGAGATTGGCATGGCtttatt tgttttttttgttttgttttggttttttttttttggCTTGACTCAGGATTTAAAAACTGGAACGGTGAAGGTGAC AGCAGTCGGTTGGAGCGAGCATCCCCCAAAGTTCACAATGTGGCCGAGGACTTTGA TTGCACATTGTTGTTTTTTTAATAGTCATTCCAAATATGAGATGCGTTGTTACAGGA AGTCCCTTGCCATCCTAAAAGCCACCCCACTTCTCTCTAAGGAGAATGGCCCAGTCC TCTCCCAAGTCCACACAGGGGAGGTGATAGCATTGCTTTCGTGTAAATTATGTAATG CAAAAtttttttaatcttcgccttaatacttttttattttgttttattttgAATGATGAGCCTTCGTGcccccccttcccccttttttg tcccccAACTTGAGATGTATGAAGGCTTTTGGTCTCCCTGGGAGTGGGTGGAGGCAGCCAGGGCTTACCTGTACACTGACTggatccccgggttaacgtcaggtggccggccgttgctggcgtttttccataggctcc gcccccctgacgagcatcacaaaaatcgacgctcaagtcagaggtggcgaaacccgacaggactataaagataccaggcgtttccccct ggaagctccctcgtgcgctctcctgttccgaccctgccgcttaccggatacctgtccgcctttctcccttcgggaagcgtggcgctttctcata gctcacgctgtaggtatctcagttcggtgtaggtcgttcgctccaagctgggctgtgtgcacgaaccccccgttcagcccgaccgctgcgc cttatccggtaactatcgtcttgagtccaacccggtaagacacgacttatcgccactggcagcagccactggtaacaggattagcagagcg aggtatgtaggcggtgctacagagttcttgaagtggtggcctaactacggctacactagaagaacagtatttggtatctgcgctctgctgaag ccagttaccttcggaaaaagagttggtagctcttgatccggcaaacaaaccaccgctggtagcggtggtttttttgtttgcaagcagcagatta cgcgcagaaaaaaaggatctcaagaagatcctttgatct
[0186] SEQ ID NO: 17 is a guide expression (scaffold SEQ ID NO: 13) and luciferase reporter donor plasmid (FIG. 20):
[0187] tttctactttcccatgattccttcatatttgcatatacgatacaaggctgttagagagataattggaattaatttgactgtaaacacaa agatattagtacaaaatacgtgacgtagaaagtaataatttcttgggtagtttgcagttttaaaattatgttttaaaatggactatcatatgcttacc gtaacttgaaagtatttcgatttcttggctttatatatcttGTGGAAAGGACGAAACACCGCTCTAGCAGGCCT GGCAAATTTCTACTGTTGTAGATCCTCTCAGGCATGGAGTCCTTTTTTTggggtctgacgctc agtggaacgaaaactcacgttaagggattttggtcatgagattatcaaaaaggatcttcacctagatccttttaaattaaaaatgaagttttaaat caatctaaagtatatatgagtaaacttggtctgacagttaccaatgcttaatcagtgaggcacctatctcagcgatctgtctatttcgttcatccat agttgcctgactccccgtcgtgtagataactacgatacgggagggcttaccatctggccccagtgctgcaatgataccgcgagacccacgc tcaccggctccagatttatcagcaataaaccagccagccggaagggccgagcgcagaagtggtcctgcaactttatccgcctccatccagt ctattaattgttgccgggaagctagagtaagtagttcgccagttaatagtttgcgcaacgttgttgccattgctacaggcatcgtggtgtcacg ctcgtcgtttggtatggcttcattcagctccggttcccaacgatcaaggcgagttacatgatcccccatgttgtgcaaaaaagcggttagctcc ttcggtcctccgatcgttgtcagaagtaagttggccgcagtgttatcactcatggttatggcagcactgcataattctcttactgtcatgccatcc gtaagatgcttttctgtgactggtgagtactcaaccaagtcattctgagaatagtgtatgcggcgaccgagttgctcttgcccggcgtcaatac gggataataccgcgccacatagcagaactttaaaagtgctcatcattggaaaacgttcttcggggcgaaaactctcaaggatcttaccgctgt tgagatccagttcgatgtaacccactcgtgcacccaactgatcttcagcatcttttactttcaccagcgtttctgggtgagcaaaaacaggaag gcaaaatgccgcaaaaaagggaataagggcgacacggaaatgttgaatactcatactcttcctttttcaatattattgaagcatttatcagggtt attgtctcatgagcggatacatatttgaatgtatttagaaaaataaacaaataggggttccgcgcacatttccccgaaaagtgccactctagag ccacctgacgttaacccggggaCATGGAGAAAATCTGGCACCACACCTTCTACAATGAGCTGCG TGTGGCTCCCGAGGAGCACCCCGTGCTGCTGACCGAGGCCCCCCTGAACCCCAAGG CCAACCGCGAGAAGATGACCCAGGTGAGTGGCCCGCTACCTCTTCTGGTGGCCGCC TCCCTCCTTCCTGGCCTCCCGGAGCTGCGCCCTTTCTCACTGGTTCTCTCTTCTGCCG TTTTCCGTAGGACTCTCTTCTCTGACCTGAGTCTCCTTTGGAACTCTGCAGGTTCTAT TTGCTTTTTCCCAGATGAGCTCTTTTTCTGGTGTTTGTCTCTCTGACTAGGTGTCTAA GACAGTGTTGTGGGTGTAGGTACTAACACTGGCTCGTGTGACAAGGCCATGAGGCT GGTGTAAAGCGGCCTTGGAGTGTGTATTAAGTAGGcGCACAGTAGGTCTGAACAGA CTCCCCATCCCAAGACCCCAGCACACTTAGCCGTGTTCTTTGCACTTTCTGCATGTCC CCCGTCTGGCCTGGCTGTCCCCAGTGGCTTCCCCAGTGTGACATGGTGcATCTCTGCC TTACAGATCATGTTTGAGACCTTCAACACCCCAGCCATGTACGTTGCTATCCAGGCT GTGCTATCCCTGTACGCCTCTGGCCGTACCACTGGCATCGTGATGGACTCCGGTGAC GGGGTCACCCACACTGTGCCCATCTACGAGGGGTATGCCCTCCCCCATGCCATCCTG CGTCTGGACCTGGCTGGCCGGGACCTGACTGACTACCTCATGAAGATCCTCACCGA GCGCGGCTACAGCTTCACCACCACGGCCGAGCGGGAAATCGTGCGTGACATTAAGG AGAAGCTGTGCTACGTCGCCCTGGACTTCGAGCAAGAGATGGCCACGGCTGCTTCC AGCTCCTCCCTGGAGAAGAGCTACGAGCTGCCTGACGGCCAGGTCATCACCATTGG CAATGAGCGGTTCCGCTGCCCTGAGGCACTCTTCCAGCCTTCCTTCCTGGGCATGGAAAGCTGTGGCATACACGAGACCACTTTCAATTCAATCATGAAGTGCGACGTCGACATTCGCAAAGATCTTTACGCTAACACAGTGCTCTCTGGGGGAACGACCATGTATCCGGGGATAGCCGATCGGATGCAGAAGGAGATTACAGCATTGGCTCCCAGTACTATGAAGATTAAGATCATCGCGCCACCCGAGAGAAAATATAGTGTTTGGATCGGTGGCTCTATCCTGGCCTCACTGAGCACCTTTCAACAGATGTGGATTTCCAAACAGGAATACGACGAATCTGGCCCTAGCATCGTGCATAGGAAGTGCTTCGGATCCGGtgaAggGagCggATCTc tCctAacAtgTggAgaTgtCgaAgagaaTccTggAccAATGGAGATGGAGAAGGAGGAGAACGTGGTGTACGGACCTCTCCCCTTTTACCCTATTGAAGAAGGGTCCGCCGGTATCCAGTTGCACAAGTATATGCATCAGTACGCCAAACTGGGAGCAATAGCATTTTCAAATGCTCTAACAGGTGTCGATATCTCATATCAGGAGTATTTTGACATTACATGCAGGCTGGCCGAGGCAATGAAAAACTTCGGGATGAAACCTGAAGAACATATCGCTCTCTGCTCCGAGAACTGTGAAGAGTTCTTTATACCCGTTTTGGCTGGACTTTACATCGGTGTGGCGGTGGCCCCCACGAATGAGATATACACACTGAGGGAACTCAATCATAGTCTGGGAATCGCCCAGCCCACGATCGTGTTCAGTTCTCGGAAGGGACTCCCAAAGGTGCTCGAGGTCCAAAAAACTGTCACCTGTATTAAGAAGATTGTTATCCTGGACAGCAAAGTTAACTTCGGGGGTCACGATTGCATGGAGACTTTTATCAAGAAGCACGTAGAGCTCGGATTTCAGCCGTCCTCTTTCGTGCCAATTGATGTAAAGAACCGAAAGCAACACGTTGCTTTACTGATGAATTCTAGCGGATCTACCGGATTGCCGAAAGGGGTCAGGATAACCCACGAAGGCGCGGTCACAAGATTCTCACATGCAAAGGACCCAATATACGGCAATCAAGTCTCCCCTGGAACAGCTATTTTAACCGTGGTGCCCTTTCACCATGGTTTTGGCATGTTTACTACACTGGGCTATTTTGCGTGTGGCTACCGTGTCGTGATGCTGACAAAATTTGACGAAGAACTGTTCCTCCGGACCCTGCAGGACTACAAGTGCACCAGCGTCATTCTGGTGCCCACCTTATTTGCCATCCTGAACAAATCCGAACTGATTGATAAATTCGATCTCTCCAATCTTACTGAAATCGCCAGCGGTGGAGCTCCCCTCGCCAAGGAGGTGGGCGAGGCCGTCGCACGGAGATTCAATCTGCCTGGCGTTAGACAGGGGTATGGACTGACTGAGACAACCAGCGCCTTCATTATTACGCCTCGCGGGGATGACAAACCTGGGGCGTCGGGCAAGGTTGCCCCTCTGTTCAAAGTGAAGGTGATCGACTTGGACACGAAAAAAACCTTAGGGGTCAACCGCCGAGGCGAAATCTGTGTAAAAGGGCCAAGCCTTATGCTAGGTTATTCAAACAACCCTGAGGCAACCAGAGAAACTATTGACGAAGAGGGCTGGCTGCACACCGGGGACCTTGGGTATTACGATGAAGACGAACACTTCTTCATAGTGGATAGGCTGAAGTCTCTGATCAAGTACCAGGGCTACCAGGTGCCGCCCGCTGAGCTTGAGAGTGTCCTCCTACAACACCCAAATATTTTCGATGCTGGGGTAGCCGGCGTGCCAGATCCCGATGCAGGCGAGCTGCCCGGCGCCGTAGTGGTCATGGAAAAGGGTAAAACTATGACTGAGAAGGAGATCGTGGACTATGTTAACAGTCAGGTTGTTAATCATAAGCGCTTGCGTGGCGGAGTACGGTTCGTGGACGAGGTGCCAAAAGGCCTTACAGGCAAGATCGACGCA AAGGTGATCCGCGAGATTTTGAAAAAACCACAGGCTAAGATGGGCATGGAGTCCTG TGGCATCCACGAAACTACCTTCAACTCCATCATGAAGTGTGACGTGGACATCCGCA AAGACCTGTACGCCAACACAGTGCTGTCTGGCGGCACCACCATGTACCCTGGCATT GCCGACAGGATGCAGAAGGAGATCACTGCCCTGGCACCCAGCACAATGAAGATCA AGGTGGGTGTCTTTCCTGCCTGAGCTGACCTGGGCAGGTCaGCTGTGGGGTCCTGTG GTGTGTGGGGAGCTGTCACATCCAGGGTCCTCACTGCCTGTCCCCTTCCCTCCTCAG ATCATTGCTCCTCCTGAGCGCAAGTACTCCGTGTGGATCGGCGGCTCCATCCTGGCC TCGCTGTCCACCTTCCAGCAGATGTGGATCAGCAAGCAGGAGTATGACGAGTCCGG CCCCTCCATCGTCCACCGCAAATGCTTCTAGGCGGACTATGACTTAGTTGCGTTACA CCCTTTCTTGACAAAACCTAACTTGCGCAGAAAACAAGATGAGATTGGCATGGCtttatt tgttttttttgttttgttttggttttttttttttggCTTGACTCAGGATTTAAAAACTGGAACGGTGAAGGTGAC AGCAGTCGGTTGGAGCGAGCATCCCCCAAAGTTCACAATGTGGCCGAGGACTTTGATTGCACATTGTTGTTTTTTTAATAGTCATTCCAAATATGAGATGCGTTGTTACAGGA AGTCCCTTGCCATCCTAAAAGCCACCCCACTTCTCTCTAAGGAGAATGGCCCAGTCC TCTCCCAAGTCCACACAGGGGAGGTGATAGCATTGCTTTCGTGTAAATTATGTAATG CAAAAtttttttaatcttcgccttaatacttttttattttgttttattttgAATGATGAGCCTTCGTGcccccccttcccccttttttg tcccccAACTTGAGATGTATGAAGGCTTTTGGTCTCCCTGGGAGTGGGTGGAGGCAGCC AGGGCTTACCTGTACACTGACTggatccccgggttaacgtcaggtggccggccgttgctggcgtttttccataggctcc gcccccctgacgagcatcacaaaaatcgacgctcaagtcagaggtggcgaaacccgacaggactataaagataccaggcgtttccccct ggaagctccctcgtgcgctctcctgttccgaccctgccgcttaccggatacctgtccgcctttctcccttcgggaagcgtggcgctttctcata gctcacgctgtaggtatctcagttcggtgtaggtcgttcgctccaagctgggctgtgtgcacgaaccccccgttcagcccgaccgctgcgc cttatccggtaactatcgtcttgagtccaacccggtaagacacgacttatcgccactggcagcagccactggtaacaggattagcagagcg aggtatgtaggcggtgctacagagttcttgaagtggtggcctaactacggctacactagaagaacagtatttggtatctgcgctctgctgaag ccagttaccttcggaaaaagagttggtagctcttgatccggcaaacaaaccaccgctggtagcggtggtttttttgtttgcaagcagcagatta cgcgcagaaaaaaaggatctcaagaagatcctttgatct
[0188] SEQ ID NO:18 is a guide expression (scaffold SEQ ID NO:12) and mCherry-tCD19 reporter donor plasmid (FIG. 21):
[0189] cttaatcagtgaggcacctatctcagcgatctgtctatttcgttcatccatagttgcctgactccccgtcgtgtagataactacgat acgggagggcttaccatctggccccagtgctgcaatgataccgcgagacccacgctcaccggctccagatttatcagcaataaaccagcc agccggaagggccgagcgcagaagtggtcctgcaactttatccgcctccatccagtctattaattgttgccgggaagctagagtaagtagtt cgccagttaatagtttgcgcaacgttgttgccattgctacaggcatcgtggtgtcacgctcgtcgtttggtatggcttcattcagctccggttcc caacgatcaaggcgagttacatgatcccccatgttgtgcaaaaaagcggttagctccttcggtcctccgatcgttgtcagaagtaagttggcc gcagtgttatcactcatggttatggcagcactgcataattctcttactgtcatgccatccgtaagatgcttttctgtgactggtgagtactcaaccaagtcattctgagaatagtgtatgcggcgaccgagttgctcttgcccggcgtcaatacgggataataccgcgccacatagcagaactttaaa agtgctcatcattggaaaacgttcttcggggcgaaaactctcaaggatcttaccgctgttgagatccagttcgatgtaacccactcgtgcacc caactgatcttcagcatcttttactttcaccagcgtttctgggtgagcaaaaacaggaaggcaaaatgccgcaaaaaagggaataagggcg acacggaaatgttgaatactcatactcttcctttttcaatattattgaagcatttatcagggttattgtctcatgagcggatacatatttgaatgtattt agaaaaataaacaaataggggttccgcgcacatttccccgaaaagtgccacctTCCGGTGACGGGGTCACCCACAC TGTGCCCATCTACGAGGGGTATGCCCTCCCCCATGCCATCCTGCGTCTGGACCTGGC TGGCCGGGACCTGACTGACTACCTCATGAAGATCCTCACCGAGCGCGGCTACAGCT TCACCACCACGGCCGAGCGGGAAATCGTGCGTGACATTAAGGAGAAGCTGTGCTAC GTCGCCCTGGACTTCGAGCAAGAGATGGCCACGGCTGCTTCCAGCTCCTCCCTGGA GAAGAGCTACGAGCTGCCTGACGGCCAGGTCATCACCATTGGCAATGAGCGGTTCC GCTGCCCTGAGGCACTCTTCCAGCCTTCCTTCCTGGGTGAGTGGAGACTGTCTCCCG GCTCTGCCTGACATGAGGGTTACCCCTCGGGGCTGTGCTGTGGAAGCTAAGTCCTGC CCTCATTTCCCTCTCAGGCATGGAGTCCTGTGGCATCCACGAAACTACCTTCAACTC CATCATGAAGTGTGACGTGGACATCCGCAAAGACCTGTACGCCAACACAGTGCTGT CTGGCGGCACCACCATGTACCCTGGCATTGCCGACAGGATGCAGAAGGAGATCACT GCCCTGGCACCCAGCACAATGAAGATCAAGGTGGGTGTCTTTCCTGCCTGAGCTGA CCTGGGCAGGTCGGCTGTGGGGTCCTGTGGTGTGTGGGGAGCTGTCACATCCAGGG TCCTCACTGCCTGTCCCCTTCCCTCCTCAGATCATTGCTCCTCCTGAGCGCAAGTACT CCGTGTGGATCGGCGGCTCCATCCTGGCCTCGCTGTCCACCTTCCAGCAGATGTGGA TCAGCAAGCAGGAGTATGACGAGTCCGGCCCCTCCATCGTCCACCGCAAATGCTTC TAGGCGGACTATGACTTAGTTGCGTTACACCCTTTCTTGACAAAACCTAACTTGCGC AGAAAACAAGATGAGATTGGCATGGCgacgagctgtacaagtaaggcgcgcccccccctaacgttactggccg aagccgcttggaataaggccggtgtgcgtttgtctatatgttattttccaccatattgccgtcttttggcaatgtgagggcccggaaacctggc cctgtcttcttgacgagcattcctaggggtctttcccctctcgccaaaggaatgcaaggtctgttgaatgtcgtgaaggaagcagttcctctgg aagcttcttgaagacaaacaacgtctgtagcgaccctttgcaggcagcggaaccccccacctggcgacaggtgcctctgcggccaaaag ccacgtgtataagatacacctgcaaaggcggcacaaccccagtgccacgttgtgagttggatagttgtggaaagagtcaaatggctctcct caagcgtattcaacaaggggctgaaggatgcccagaaggtaccccattgtatgggatctgatctggggcctcggtgcacatgctttacatgt gtttagtcgaggttaaaaaaacgtctaggccccccgaaccacggggacgtggttttcctttgaaaaacacgatgataatatggccacaacc ATGGTCTCAAAAGGCGAAGAAGACAACATGGCAATAATCAAGGAGTTTATGCGCTT TAAAGTACATATGGAAGGGTCAGTAAACGGACATGAATTTGAGATAGAAGGTGAA GGAGAAGGTAGACCCTACGAAGGCACACAAACTGCCAAGTTGAAAGTTACTAAAG GCGGCCCCCTCCCTTTTGCTTGGGACATTTTGTCACCCCAGTTCATGTATGGCAGTA AAGCATATGTTAAACACCCAGCCGACATCCCCGATTACCTCAAGCTCTCATTTCCCG AGGGTTTTAAGTGGGAGAGGGTTATGAACTTTGAAGATGGGGGTGTAGTAACAGTTACTCAGGATTCAAGCCTCCAGGATGGGGAGTTTATATATAAAGTAAAACTTCGAGGAACTAACTTCCCTAGTGATGGACCTGTCATGCAGAAAAAGACCATGGGCTGGGAAGCCTCAAGTGAACGGATGTACCCGGAGGACGGAGCCTTGAAAGGCGAGATAAAGCAGCGATTGAAATTGAAGGATGGAGGACATTACGATGCTGAGGTAAAGACTACCTATAAAGCGAAGAAACCCGTTCAGCTGCCCGGAGCTTATAACGTCAATATAAAACTTGACATCACGTCCCACAACGAAGACTACACGATAGTGGAGCAGTATGAACGAGCTGAAGGCCGGCACTCCACGGGAGGTATGGACGAGTTGTATAAGGAGGGCAGGGGCAGCCTGCTGACCTGCGGCGACGTGGAGGAGAACCCCGGCCCCATGCCGCCACCTCGGCTTCTGTTCTTCCTCCTCTTCCTTACACCTATGGAAGTAAGACCGGAGGAACCGTTGGTTGTGAAAGTCGAGGAAGGTGATAATGCAGTCCTTCAATGTCTCAAAGGGACCAGTGATGGTCCTACTCAGCAGCTCACCTGGTCCAGGGAATCTCCCCTGAAACCATTTTTGAAGTTGTCCCTTGGTCTGCCCGGACTGGGCATTCACATGAGACCCCTGGCAATCTGGCTTTTCATCTTTAATGTCTCCCAACAAATGGGTGGTTTTTATCTTTGCCAGCCTGGCCCACCTAGTGAGAAGGCCTGGCAACCGGGTTGGACGGTCAATGTGGAAGGGAGCGGTGAGCTCTTTCGGTGGAACGTCAGCGATTTGGGGGGCCTTGGTTGCGGACTGAAAAATCGATCATCAGAGGGACCGTCTAGCCCTAGCGGGAAACTGATGTCCCCGAAACTTTATGTATGGGCGAAAGACAGGCCTGAAATTTGGGAGGGCGAACCGCCGTGCCTCCCGCCCAGGGATTCACTGAATCAATCCTTGTCTCAGGATCTTACGATGGCTCCCGGCAGTACTCTGTGGCTGTCCTGTGGCGTCCCGCCAGACTCCGTAAGCAGAGGGCCACTCTCCTGGACCCATGTGCACCCGAAAGGCCCGAAGTCATTGTTGTCACTCGAATTGAAAGACGACCGCCCCGCACGAGATATGTGGGTGATGGAAACCGGTCTGCTTTTGCCGCGCGCGACTGCCCAAGATGCCGGCAAGTACTACTGTCATCGGGGGAATCTCACCATGTCTTTCCATTTGGAAATAACAGCAAGACCCGTATTGTGGCACTGGCTTTTGAGAACCGGGGGTTGGAAGGTTTCTGCTGTCACTCTGGCTTACTTGATCTTCTGTCTCTGTTCCCTTGTCGGAATTCTGCATCTCCAGCGGGCACTGGTTCTTCGAAGGAAACGAAAACGAATGACCGATCCTACTCGGCGATTTTGAcatcggccgcgctcccgatGAACGGTGAAGGTGACAGCAGTCGGTTGGAGCGAGCATCCCCCAAAGTTCACAATGTGGCCGAGGACTTTGATTGCACATTGTTGTTTTTTTAATAGTCATTCCAAATATGAGATGCGTTGTTACAGGAAGTCCCTTGCCATCCTAAAAGCCACCCCACTTCTCTCTAAGGAGAATGGCCCAGTCCTCTCCCAAGTCCACACAGGGGAGGTGATAGCATTGCTTTCGTGTAAATTATGTAATGCAAAAtttt tttaatcttcgccttaatacttttttattttgttttattttgAATGATGAGCCTTCGTGcccccccttcccccttttttgtcccccAACTTGAGATGTATGAAGGCTTTTGGTCTCCCTGGGAGTGGGTGGAGGCAGCCAGGGCTTACCTGTACACTGACTTGAGACCAGTTGAATAAAAGTGCACACCTTAAAAATGAGGCCAAGTGTGACTTTGTGGTGTGGCTGGGTTGGGGGCAGCAgagggtgaaccctgcaggagggtgaaccctgcaaaagggtgGGGCAGTGGGGGCCAACTTGTCCTTACCCAGAGTGCAGGTGTGTGG AGATCCCTCCTGCCTTGACATTGAGCAGCCTTAGAGGGTGGGGGAGGCTCAGGGGT CAGGTCTCTGTTCCTGCTTATTGGGGAGTTCCTGGCCTGGCCCTTCTATGTCTCCCCA GGTACCCCAGTTTTTCTGGGTTCACCCAGAGTGCAGATGCTTGAGGAGGTGGGAAG GGACTATTTGGGGGTGTCTGGCTCAGGTGCCATGCCTCACTGGGGCTGGTTGGCACC TGCATTTCCTGGGAGTGGGGCTGTCTCAGGGTAGCTGGGCACGGTGTTCCCTTGAGT GGGGGTGTAGTGGGTGTTCCTAGCTGCCACGCCTTTGCCTTCACCTATGGGATCGTG GCTGTCAGCCTTGAGGGTCAGCCTGGCCCAGGCTCCCATAGGCTTAGGAGAGGCCG CAATTCCTACCTGTTCATCCAGAATTCacgtggggctcacctcgaccacgctggtagcggtggtttttttgtttgca agcagcagattacgcgcagaaaaaaaggatctcaagaagatcctttgatcttttctactttcccatgattccttcatatttgcatatacgatacaa ggctgttagagagataattggaattaatttgactgtaaacacaaagatattagtacaaaatacgtgacgtagaaagtaataatttcttgggtagtt tgcagttttaaaattatgttttaaaatggactatcatatgcttaccgtaacttgaaagtatttcgatttcttggctttatatatcttGTGGAAAG GACGAAACACCGTCTGGCCCCAAATTTTAATTTCTACTGTTGTAGATtggCTTGACTC AGGATTTAAttttTtggtaatagcgatgactaatacgtagatgtactgccaagtaggaaagtcccataaggtcatgtactgggcata atgccaggcgggccatttaccgtcattgacgtcaatagggggcgtacttggcatatgatacacttgatgtactgccaagtgggcagtttacc gtaaatactccacccattgacgtcaatggaaagtGAATtcgttcggctgcggcgagcggtatcagctcactcaaaggcggtaatacggt tatccacagaatcaggggataacgcaggaaagaacatgtgagcaaaaggccagcaaaaggccaggaaccgtaaaaaggccgcgttgct ggcgtttttccataggctccgcccccctgacgagcatcacaaaaatcgacgctcaagtcagaggtggcgaaacccgacaggactataaag ataccaggcgtttccccctggaagctccctcgtgcgctctcctgttccgaccctgccgcttaccggatacctgtccgcctttctcccttcggga agcgtggcgctttctcatagctcacgctgtaggtatctcagttcggtgtaggtcgttcgctccaagctgggctgtgtgcacgaaccccccgtt cagcccgaccgctgcgccttatccggtaactatcgtcttgagtccaacccggtaagacacgacttatcgccactggcagcagccactggta acaggattagcagagcgaggtatgtaggcggtgctacagagttcttgaagtggtggcctaactacggctacactagaagaacagtatttgg tatctgcgctctgctgaagccagttaccttcggaaaaagagttggtagctcttgatccggcaaacaaaccaccgctggtagcggtggttttttt gtttgcaagcagcagattacgcgcagaaaaaaaggatctcaagaagatcctttgatcttttctacggggtctgacgctcagtggaacgaaaa ctcacgttaagggattttggtcatgagattatcaaaaaggatcttcacctagatccttttaaattaaaaatgaagttttaaatcaatctaaagtatat atgagtaaacttggtctgacagttaccaatg
[0190] SEQ ID NO: 19 is a chimeric nuclease (SEQ ID NO:4) expression vector:
[0191] ctggccccagtgctgcaatgataccgcgagacccacgctcaccggctccagatttatcagcaataaaccagccagccggaa gggccgagcgcagaagtggtcctgcaactttatccgcctccatccagtctattaattgttgccgggaagctagagtaagtagttcgccagtta atagtttgcgcaacgttgttgccattgctacaggcatcgtggtgtcacgctcgtcgtttggtatggcttcattcagctccggttcccaacgatca aggcgagttacatgatcccccatgttgtgcaaaaaagcggttagctccttcggtcctccgatcgttgtcagaagtaagttggccgcagtgtta tcactcatggttatggcagcactgcataattctcttactgtcatgccatccgtaagatgcttttctgtgactggtgagtactcaaccaagtcattct gagaatagtgtatgcggcgaccgagttgctcttgcccggcgtcaatacgggataataccgcgccacatagcagaactttaaaagtgctcat cattggaaaacgttcttcggggcgaaaactctcaaggatcttaccgctgttgagatccagttcgatgtaacccactcgtgcacccaactgatcttcagcatcttttactttcaccagcgtttctgggtgagcaaaaacaggaaggcaaaatgccgcaaaaaagggaataagggcgacacggaa atgttgaatactcatactcttcctttttcaatattattgaagcatttatcagggttattgtctcatgagcggatacatatttgaatgtatttagaaaaat aaacaaataggggttccgcgcacatttccccgaaaagtgccacctaaattgtaagcgttaatattttgttaaaattcgcgttaaatttttgttaaat cagctcattttttaaccaataggccgaaatcggcaaaatcccttataaatcaaaagaatagaccgagatagggttgagtgttgttccagtttgg aacaagagtccactattaaagaacgtggactccaacgtcaaagggcgaaaaaccgtctatcagggcgatggcccactacgtgaaccatca ccctaatcaagttttttggggtcgaggtgccgtaaagcactaaatcggaaccctaaagggagcccccgatttagagcttgacggggaaagc cggcgaacgtggcgagaaaggaagggaagaaagcgaaaggagcgggcgctagggcgctggcaagtgtagcggtcacgctgcgcgt aaccaccacacccgccgcgcttaatgcgccgctacagggcgcgtcccattcgccattcaggctgcgcaactgttgggaagggcgatcgg tgcgggcctcttcgctattacgccagctgcgcgctcgctcgctcactgaggccgcccgggcaaagcccgggcgtcgggcgacctttggt cgcccggcctcagtgagcgagcgagcgcgcagagagggagtggccaactccatcactaggggttccttgtagttaatgattaacccgcc atgctacttatctacgtagccatgctctaggaagagtaccattgacgtcaataatgacgtatgttcccatagtaacgccaatagggactttccat tgacgtcaatgggtggagtatttacggtaaactgcccacttggcagtacatcaagtgtatcatatgccaagtacgccccctattgacgtcaatg acggtaaatggcccgcctggcattatgcccagtacatgaccttatgggactttcctacttggcagtacatctacgtattagtcatcgctattacc atggtcgaggtgagccccacgttctgcttcactctccccatctcccccccctccccacccccaattttgtatttatttattttttaattattttgtgca gcgatgggggcggggggggggggggggsgcgcgccggggggggggggggggggggggggggggggsggggsgrggcggag aggtgcggcggcagccaatcagagcggcgcgctccgaaagtttccttttatggcgaggcggcggcggcggcggccctataaaaagcga agcgcgcggcgggcgggagtcgctgcgcgctgccttcgccccgtgccccgctccgccgccgcctcgcgccgcccgccccggctctga ctgaccgcgttactcccacaggtgagcgggcgggacggcccttctcctccgggctgtaattagcgcttggtttaatgacggcttgtttcttttc tgtggctgcgtgaaagccttgaggggctccgggagggccctttgtgcggggggagcggctcggggctgtccgcggggggacggctgc cttcgggggggacggggcagggcggggttcggcttctggcgtgtgaccggcggctctagagcctctgctaaccatgttcatgccttcttctt tttcctacagctcctgggcaacgtgctggttattgtgctgtctcatcattttggcaaagaattggatccgccaccATGcctaagaagaagcg gaaggttggtattcacggggtgcctgcggctATGAACAATGGTACAAATAATTTCCAGAACTTCATCG GCATTTCAAGTCTTCAGAAGACCCTCAGAAATGCACTTATCCCTACCGAGACTACAC AGCAGTTTATTGTAAAAAACGGGATCATTAAAGAAGACGAGCTTCGAGGAGAAAAT AGGCAGATTCTCAAGGATATTATGGACGACTACTACCGTGGGTTCATATCCGAAAC TTTGTCCTCCATAGACGACATCGACTGGACTAGCCTGTTCGAGAAAATGGAAATTCA ACTTAAGAATGGGGATAACAAAGACACACTAATCAAAGAGCAGGCAGAAAAAAGG AAAGCCATTTATAAGAAATTCGCCGATGACGATAGATTTAAGAATATGTTCAGTGC CAAGCTCATCTCAGATATCCTGCCGGAGTTTGTCATACATAATAATAACTACAGTGC CTCGGAAAAGGAAGAGAAGACTCAGGTTATCAAGCTGTTCTCTCGATTTGCCACGA GCTTCAAAGATTATTTCAAGAATAGAGCCAACTGTTTCAGTGCAGATGACATTTCTT CCTCCTCCTGTCACCGGATAGTGAATGACAACGCCGAGATATTTTTTTCAAACGCAC TTGTATACCGAAGAATTGTGAAGAATCTGAGCAATGACGACATTAACAAGATTTCA GGCGATATAAAGGACTCCCTTAAAGAGATGAGCCTCGAAGAAATTTATAGCTATGAGAAGTACGGCGAGTTCATAACGCAAGAAGGGATTAGCTTCTACAACGACATCTGCGGCAAAGTCAATTCCTTTATGAACCTGTACTGCCAGAAAAATAAAGAGAACAAAAATCTGTACAAATTGAGGAAACTCCACAAACAGATCCTGTGTATTGCAGATACCTCTTACGAAGTCCCCTATAAATTCGAGTCCGACGAAGAGGTATACCAATCCGTCAACGGTTTTCTCGACAATATCAGTTCCAAACACATTGTGGAGAGGCTGCGGAAGATCGGTGACAATTATAATGGTTACAACTTGGATAAAATTTATATAGTGTCAAAGTTCTATGAATCCGTAAGCCAAAAGACATATCGGGATTGGGAGACCATTAATACTGCACTCGAAATCCATTACAATAACATCCTGCCAGGTAACGGGAAAAGTAAGGCGGATAAGGTTAAAAAGGCTGTCAAGAACGACCTGCAAAAGTCAATTACAGAAATCAATGAGCTGGTGAGCAACTACAAGCTGTGCCCCGACGACAATATCAAAGCCGAAACCTACATACATGAAATCTCCCACATACTGAACAACTTTGAGGCCCAGGAGCTGAAATATAATCCCGAGATTCACCTGGTCGAAAGCGAACTGAAAGCAAGCGAGCTGAAGAACGTGCTGGACGTCATAATGAATGCATTCCATTGGTGTAGTGTATTTATGACCGAAGAACTAGTTGACAAAGATAACAACTTTTATGCCGAACTGGAGGAAATCTACGACGAGATCTATACTGTTATCAGCCTATATAACCTCGTGCGGAATTACGTCACTCAGAAACCGTACAGCACTAAAAAAATCAAGCTGAATTTTGGAATCCCCACGTTAGCAGATGGCTGGTCCAAGTCTAAAGAGTATAGCAACAACGCCATCATCCTGATGCGAGACAACTTATATTATCTCGGGATCTTCAACGCCAAGAATAAGCCCGATAAGAAAATTATTGAGGGCAATACCAGCGAGAATAAAGGAGACTACAAGAAGATGATTTACAACCTGCTCCCAGGACCTAACAAGATGATTCCTAAGGTGTTTCTGTCTTCCAAGACTGGTGTGGAAACCTATAAGCCATCAGCTTACATCCTGGAGGGATACAAACAAAACAAGCATCTCAAATCTAGCAAGGACTTCGACATTACCTTCTGTCACGACCTTATAGACTATTTTAAGAACTGCATTGCCATTCACCCAGAGTGGAAGAACTTTGGGTTCGACTTCTCTGACACATCGACATATGAAGATATATCAGGCTTTTACCGCGAGGTTGAGCTGCAGGGATACAAGATCGACTGGACCTATATTAGCGAAAAGGACATTGACCTCCTGCAGGAGAAGGGCCAGCTCTATCTATTCCAGATTTATAATAAGGATTTCAGCAAAAAGAGTTCTGGTAACGATAACTTGCATACCATGTACCTTAAGAACTTGTTCAGCGAAGAGAATTTGAAAGACATCGTGCTTAAGCTGAACGGGGAGGCTGAAATCTTTTTTCGGAAGTCATCTATCAAGAATCCAATCATCCACAAAAAAGGAAGCATCTTGGTGAACCGGACGTACGAGGCGGAAGAGAAGGATCAGTTTGGGAACATTCAGATTGTGCGCAAAACTATTCCTGAGAATATCTATCAGGAGCTTTACAAGTACTTTAATGACAAGTCAGACAAGGAGCTCAGCGATGAAGCTGCGAAGCTGAAAAATGTGGTCGGGCATCACGAAGCGGCCACAAACATCGTTAAAGATTACAGATATACTTACGATAAATATTTTCTCCACATGCCCATCACGATCAACTTTAAGGCTAACAAAACTAGTTTCATCAATGATCGCATCCTACAGTATATTGCAAAAGAAAAAGATTTACATGTGATTGGCATAGACAGGGGTGAGCGCAATTTGATTTACGTCTCTGTGATCGACACCTGTGGC AACATCGTTGAACAGAAGAGCTTTAACATCGTCAATGGGTACGATTACCAGATTAA ACTGAAACAACAGGAGGGAGCTAGACAAATTGCTAGAAAGGAGTGGAAAGAGATA GGTAAGATAAAGGAGATAAAGGAAGGCTACTTGTCATTAGTGATTCACGAGATATC GAAAATGGTGATTAAATACAACGCTATTATCGCTATGGAGGATCTGTCGTATGGGTT CAAGAAAGGAAGGTTCAAAGTGGAGCGCCAAGTTTATCAGAAGTTCGAAACAATG CTGATAAACAAGCTCAATTACCTCGTGTTCAAGGACATCTCTATTACCGAGAACGG GGGACTGTTGAAAGGCTATCAGCTCACCTACATCCCGGAGAAGCTTAAAAATGTGG GCCACCAGTGCGGATGCATCTTTTATGTGCCAGCCGCTTATACAAGTAAAATCGACC CTACCACGGGTTTCGCTAACGTTCTGAACCTCTCCAAAGTCCGAAATGTGGATGCCA TCAAGAGCTTCTTCTCTAATTTTAACGAGATATCATACTCAAAGAAGGAGGCCCTCT TCAAGTTCAGCTTCGACCTGGATAGTCTGAGTAAGAAGGGATTCAGTAGCTTTGTGA AGTTCAGTAAATCTAAATGGAATGTTTATACGTTCGGGGAGCGGATCATAAAACCT AAAAATAAACAAGGCTACCGGGAGGACAAGAGAATTAATCTGACATTTGAGATGAAGAAACTGCTGAACGAATATAAAGTCAGTTTTGACCTTGAGAATAACCTTATCCCA AACCTCACCTCAGCCAACCTCAAAGATACTTTTTGGAAGGAATTATTCTTTATCTTC AAGACAACCCTGCAGCTGCGGAATAGCGTGACCAACGGTAAGGAAGACGTGCTGA TCTCTCCAGTCAAGAACGCAAAGGGCGAATTTTTCGTGTCTGGGACACATAATAAG ACTTTGCCTCAGGACTGTGATGCGAATGGGGCATACTGCATCGCCCTGAAGGGCCT GTACGAAATCAAGCAGATCACTGAGAACTGGAAAGAAGATGGTAAGTTCTCCCGGGACAAGCTAAAAATCTCTAACAAGGATTGGTTTGACTTCATACAGAACAAACGCTAT CTGaagcgtcctgctgccaccaaaaaggccggacaggctaagaaaaagaagTGAgcggccgcttcgagcagacatgataagata cattgatgagtttggacaaaccacaactagaatgcagtgaaaaaaatgctttatttgtgaaatttgtgatgctattgctttatttgtaaccattataa gctgcaataaacaagttaacaacaacaattgcattcattttatgtttcaggttcagggggagatgtgggaggttttttaaagcaagtaaaacctc tacaaatgtggtaaaatcgataaggatcttcctagagcatggctacgtagataagtagcatggcgggttaatcattaactacaaggaacccct agtgatggagttggccactccctctctgcgcgctcgctcgctcactgaggccgggcgaccaaaggtcgcccgacgcccgggctttgccc gggcggcctcagtgagcgagcgagcgcgcagctgcattaatgaatcggccaacgcgcggggagaggcggtttgcgtattgggcgctct tccgcttcctcgctcactgactcgctgcgctcggtcgttcggctgcggcgagcggtatcagctcactcaaaggcggtaatacggttatccac agaatcaggggataacgcaggaaagaacatgtgagcaaaaggccagcaaaaggccaggaaccgtaaaaaggccgcgttgctggcgttt ttccataggctccgcccccctgacgagcatcacaaaaatcgacgctcaagtcagaggtggcgaaacccgacaggactataaagatacca ggcgtttccccctggaagctccctcgtgcgctctcctgttccgaccctgccgcttaccggatacctgtccgcctttctcccttcgggaagcgt ggcgctttctcatagctcacgctgtaggtatctcagttcggtgtaggtcgttcgctccaagctgggctgtgtgcacgaaccccccgttcagcc cgaccgctgcgccttatccggtaactatcgtcttgagtccaacccggtaagacacgacttatcgccactggcagcagccactggtaacagg attagcagagcgaggtatgtaggcggtgctacagagttcttgaagtggtggcctaactacggctacactagaagaacagtatttggtatctgcgctctgctgaagccagttaccttcggaaaaagagttggtagctcttgatccggcaaacaaaccaccgctggtagcggtggtttttttgtttgc aagcagcagattacgcgcagaaaaaaaggatctcaagaagatcctttgatcttttctacggggtctgacgctcagtggaacgaaaactcac gttaagggattttggtcatgagattatcaaaaaggatcttcacctagatccttttaaattaaaaatgaagttttaaatcaatctaaagtatatatgag taaacttggtctgacagttaccaatgcttaatcagtgaggcacctatctcagcgatctgtctatttcgttcatccatagttgcctgactccccgtc gtgtagataactacgatacgggagggcttaccat
[0192] SEQ ID NO:20 is a SV40 nuclear localization signal sequence:
[0193] PKKKRKV
[0194] SEQ ID NO:21 is an internal linker:
[0195] GIHGVPAA
[0196] SEQ ID NO:22 is a nucleoplasmin nuclear localization signal sequence:
[0197] I<RPAATI<I<AGQAI<I<I<I<
[0198] SEQ ID NO:23 is a genotyping primer:
[0199] TTGGAAGGTTTCTGCTGTCACTC
[0200] SEQ ID NO:24 is a genotyping primer:
[0201] GAATCCAGACCTCAGCCCATAG
[0202] SEQ ID NO:25 is Chimera 10:
[0203] MNNGTNNFQNFIGISSLQKTLRNALIPTETTQQFIVKNGIIKEDELRGENRQILKDIMDDYYRGFISETLSSIDDIDWTSLFEKMEIQLKNGDNKDTLIKEQAEKRKAIYKKFADD DRFKNMFSAKLISDILPEFVIHNNNYSASEKEEKTQVIKLFSRFATSFKDYFKNRANCFSA DDISSSSCHRIVNDNAEIFFSNALVYRRIVKNLSNDDINKISGDIKDSLKEMSLEEIYSYEK YGEFITQEGISFYNDICGKVNSFMNLYCQKNKENKNLYKLRKLHKQILCIADTSYEVPY KFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNLDKIYIVSKFYESVSQKTYRD WETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSITEINELVSNYKLCPDDNIKA ETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVLDVIMNAFHWCSVFMTEELVDKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQKPYSTKKIKLNFGIPTLADGWSKSKEY SNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDYKKMIYNLLPGPNKMIPKVFL SSKTGVETYKPSAYILEGYKQNKHLKSSKDFDITFCHDLIDYFKNCIAIHPEWKNFGFDF SDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKGQLYLFQIYNKDFSKKSSGND NLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPIIHKKGSILVNRTYEAEEKDQF GNIQIVRKTIPENIYQELYKYFNDKSDKELSDEAAKLKNWGHHEAATNIVKDYRYTYD KYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGIDRGERNLIYVSVIDTCGNIVEQKSFNIVNGYDYQIKLKQQEGARQIARKEWKEIGKIKEIKEGYLSLVIHEISKMVIKYNAIIA MEDLSYGFKKGRFKVERQVYQKFETMLINKLNYLVFKDISITENGGLLKGYQLTYIPEK LKNVGHQCGCIFYVPAAYTSKIDPTTGFVNLFNTSSKTNAQERKEFLQKFESISYSAKDGGIFAFAFDYRKFGTSKTDHKNVWTAYTNGERMRYIKEKKRNELFDPSKEIKEALTSSGI KYDGGQNILPDILRSNNNGLIYTMYSSFIAAIQMRVYDGKEDYIISPIKNSKGEFFRTDPK RRELPIDADANGAYCIALKGLYEIKQITENWKEDGKFSRDKLKISNKDWFDFIQNKRYL
[0204] SEQ ID NO:26 is Chimera 11 :
[0205] MNNGTNNFQNFIGISSLQKTLRNALIPTETTQQFIVKNGIIKEDELRGENRQILKD IMDDYYRGFISETLSSIDDIDWTSLFEKMEIQLKNGDNKDTLIKEQAEKRKAIYKKFADD DRFKNMFSAKLISDILPEFVIHNNNYSASEKEEKTQVIKLFSRFATSFKDYFKNRANCFSA DDISSSSCHRIVNDNAEIFFSNALVYRRIVKNLSNDDINKISGDIKDSLKEMSLEEIYSYEK YGEFITQEGISFYNDICGKVNSFMNLYCQKNKENKNLYKLRKLHKQILCIADTSYEVPY KFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNLDKIYIVSKFYESVSQKTYRD WETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSITEINELVSNYKLCPDDNIKA ETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVLDVIMNAFHWCSVFMTEELV DKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQKPYSTKKIKLNFGIPTLADGWSKSKEY SNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDYKKMIYNLLPGPNKMIPKVFL SSKTGVETYKPSAYILEGYKQNKHLKSSKDFDITFCHDLIDYFKNCIAIHPEWKNFGFDF SDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKGQLYLFQIYNKDFSKKSSGND NLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPIIHKKGSILVNRTYEAEEKDQF GNIQIVRKTIPENIYQELYKYFNDKSDKELSDEAAKLKNWGHHEAATNIVKDYRYTYD KYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGIDRGERNLIYVSVIDTCGNIVEQK SFNIVNGYDYQIKLKQQEGARQIARKEWKEIGKIKEIKEGYLSLVIHEISKMVIKYNAIIA MEDLSYGFKKGRFKVERQVYQKFETMLINKLNYLVFKDISITENGGLLKGYQLTYIPEKLKNVGHQCGCIFYVPAAYTSKIDPTTGFADLFALSNVKNVASMREFFSKMKSVIYDKAE GKFAFTFDYLDYNVKSECGRTLWTVYTVGERFTYSRVNREYVRKVPTDIIYDALQKAG ISVEGDLRDRIAESDGDTLKSIFYAFKYALDMRVENREEDYIQSPVKNASGEFFCSKNAG KSLPQDSDANGAYCIALKGLYEIKQITENWKEDGKFSRDKLKISNKDWFDFIQNKRYL
[0206] SEQ ID NO:27 is Chimera 12:
[0207] MNNGTNNFQNFIGISSLQKTLRNALIPTETTQQFIVKNGIIKEDELRGENRQILKD IMDDYYRGFISETLSSIDDIDWTSLFEKMEIQLKNGDNKDTLIKEQAEKRKAIYKKFADD DRFKNMFSAKLISDILPEFVIHNNNYSASEKEEKTQVIKLFSRFATSFKDYFKNRANCFSA DDISSSSCHRIVNDNAEIFFSNALVYRRIVKNLSNDDINKISGDIKDSLKEMSLEEIYSYEK YGEFITQEGISFYNDICGKVNSFMNLYCQKNKENKNLYKLRKLHKQILCIADTSYEVPY KFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNLDKIYIVSKFYESVSQKTYRD WETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSITEINELVSNYKLCPDDNIKA ETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVLDVIMNAFHWCSVFMTEELVDKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQKPYSTKKIKLNFGIPTLADGWSKSKEY SNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDYKKMIYNLLPGPNKMIPKVFL SSKTGVETYKPSAYILEGYKQNKHLKSSKDFDITFCHDLIDYFKNCIAIHPEWKNFGFDF SDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKGQLYLFQIYNKDFSKKSSGND NLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPIIHKKGSILVNRTYEAEEKDQF GNIQIVRKTIPENIYQELYKYFNDKSDKELSDEAAKLKNWGHHEAATNIVKDYRYTYD KYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGIDRGERNLIYVSVIDTCGNIVEQK SFNIVNGYDYQIKLKQQEGARQIARKEWKEIGKIKEIKEGYLSLVIHEISKMVIKYNAIIA MEDLSYGFKKGRFKVERQVYQKFETMLINKLNYLVFKDISITENGGLLKGYQLTYIPEK LKNVGHQCGCIFYVPAAYTSKIDPTTGFVNLFHAQYENVDKAKSFFQKFDSISYNPKKD WFEFAFDYKNFTKKAEGSRSMWILCTHGSRIKNFRNSQKNGQWDSEEFALTEAFKSLF VRYEIDYTADLKTAIVDEKQKDFFVDLLKLFKLTVQMRNSWKEKDLDYLISPVAGADG RFFDTREGNKSLPKDADANGAYCIALKGLYEIKQITENWKEDGKFSRDKLKISNKDWF DFIQNKRYL
[0208] SEQ ID NO:28 is Chimera 13:
[0209] MNNGTNNFQNFIGISSLQKTLRNALIPTETTQQFIVKNGIIKEDELRGENRQILKD IMDDYYRGFISETLSSIDDIDWTSLFEKMEIQLKNGDNKDTLIKEQAEKRKAIYKKFADD DRFKNMFSAKLISDILPEFVIHNNNYSASEKEEKTQVIKLFSRFATSFKDYFKNRANCFSA DDISSSSCHRIVNDNAEIFFSNALVYRRIVKNLSNDDINKISGDIKDSLKEMSLEEIYSYEK YGEFITQEGISFYNDICGKVNSFMNLYCQKNKENKNLYKLRKLHKQILCIADTSYEVPY KFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNLDKIYIVSKFYESVSQKTYRD WETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSITEINELVSNYKLCPDDNIKA ETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVLDVIMNAFHWCSVFMTEELV DKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQKPYSTKKIKLNFGIPTLADGWSKSKEY SNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDYKKMIYNLLPGPNKMIPKVFL SSKTGVETYKPSAYILEGYKQNKHLKSSKDFDITFCHDLIDYFKNCIAIHPEWKNFGFDFSDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKGQLYLFQIYNKDFSKKSSGND NLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPIIHKKGSILVNRTYEAEEKDQF GNIQIVRKTIPENIYQELYKYFNDKSDKELSDEAAKLKNWGHHEAATNIVKDYRYTYD KYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGIDRGERNLIYVSVIDTCGNIVEQK SFNIVNGYDYQIKLKQQEGARQIARKEWKEIGKIKEIKEGYLSLVIHEISKMVIKYNAIIA MEDLSYGFKKGRFKVERQVYQKFETMLINKLNYLVFKDISITENGGLLKGYQLTYIPEK LKNVGHQCGCIFYVPAAYTSKIDPTTGFVNYFYTKYENVDKAKAFFEKFEAIRFNAEKK YFEFEVKKYSDFNPKAEGTQQAWTICTYGERIETKRQKDQNNKFVSTPINLTEKIEDFLGKNQIVYGDGNCIKSQIASKDDKAFFETLLYWFKMTLQMRNSETRTDIDYLISPVMNDNG TFYNSRDYEKLENPTLPKDADANGAYCIALKGLYEIKQITENWKEDGKFSRDKLKISNK DWFDFIQNKRYLSEQ ID NO:29 is Chimera 14:
[0210] MNNGTNNFQNFIGISSLQKTLRNALIPTETTQQFIVKNGIIKEDELRGENRQILKD IMDDYYRGFISETLSSIDDIDWTSLFEKMEIQLKNGDNKDTLIKEQAEKRKAIYKKFADD DRFKNMFSAKLISDILPEFVIHNNNYSASEKEEKTQVIKLFSRFATSFKDYFKNRANCFSA DDISSSSCHRIVNDNAEIFFSNALVYRRIVKNLSNDDINKISGDIKDSLKEMSLEEIYSYEK YGEFITQEGISFYNDICGKVNSFMNLYCQKNKENKNLYKLRKLHKQILCIADTSYEVPY KFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNLDKIYIVSKFYESVSQKTYRD WETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSITEINELVSNYKLCPDDNIKA ETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVLDVIMNAFHWCSVFMTEELV DKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQKPYSTKKIKLNFGIPTLADGWSKSKEY SNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDYKKMIYNLLPGPNKMIPKVFL SSKTGVETYKPSAYILEGYKQNKHLKSSKDFDITFCHDLIDYFKNCIAIHPEWKNFGFDFSDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKGQLYLFQIYNKDFSKKSSGND NLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPIIHKKGSILVNRTYEAEEKDQF GNIQIVRKTIPENIYQELYKYFNDKSDKELSDEAAKLKNWGHHEAATNIVKDYRYTYD KYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGIDRGERNLIYVSVIDTCGNIVEQK SFNIVNGYDYQIKLKQQEGARQIARKEWKEIGKIKEIKEGYLSLVIHEISKMVIKYNAIIA MEDLSYGFKKGRFKVERQVYQKFETMLINKLNYLVFKDISITENGGLLKGYQLTYIPEK LKNVGHQCGCIFYVPAAYTSKIDPTTGFVNMLNLNYTNMKDAQTLLSGMDKISFNADA NYFEFELDYEKFKTNQTDHTNKWTICTVGEKRFTYNSATKETTTVNVTEDLKKLLDKF EVKYSNGDNIKDEICRQTDAKFFEIILWLLKLTMQMRNSNTKTEEDFILSPVKNSNGEFF RSNDDANGIWPADADANGAYCIALKGLYEIKQITENWKEDGKFSRDKLKISNKDWFDF IQNKRYL
[0211] SEQ ID NO:30 is Chimera 15:
[0212] MNNGTNNFQNFIGISSLQKTLRNALIPTETTQQFIVKNGIIKEDELRGENRQILKD IMDDYYRGFISETLSSIDDIDWTSLFEKMEIQLKNGDNKDTLIKEQAEKRKAIYKKFADD DRFKNMFSAKLISDILPEFVIHNNNYSASEKEEKTQVIKLFSRFATSFKDYFKNRANCFSA DDISSSSCHRIVNDNAEIFFSNALVYRRIVKNLSNDDINKISGDIKDSLKEMSLEEIYSYEK YGEFITQEGISFYNDICGKVNSFMNLYCQKNKENKNLYKLRKLHKQILCIADTSYEVPY KFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNLDKIYIVSKFYESVSQKTYRD WETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSITEINELVSNYKLCPDDNIKA ETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVLDVIMNAFHWCSVFMTEELVDKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQKPYSTKKIKLNFGIPTLADGWSKSKEY SNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDYKKMIYNLLPGPNKMIPKVFL SSKTGVETYKPSAYILEGYKQNKHLKSSKDFDITFCHDLIDYFKNCIAIHPEWKNFGFDF SDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKGQLYLFQIYNKDFSKKSSGND NLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPIIHKKGSILVNRTYEAEEKDQF GNIQIVRKTIPENIYQELYKYFNDKSDKELSDEAAKLKNWGHHEAATNIVKDYRYTYD KYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGIDRGERNLIYVSVIDTCGNIVEQK SFNIVNGYDYQIKLKQQEGARQIARKEWKEIGKIKEIKEGYLSLVIHEISKMVIKYNAIIA MEDLSYGFKKGRFKVERQVYQKFETMLINKLNYLVFKDISITENGGLLKGYQLTYIPEK LKNVGHQCGCIFYVPAAYTSKIDPTTGFVNLLYADDLKYKNQEQAKSFIQKIDNIYFEN GEFKFDIDFSKWNNRYSISKTKWTLTSYGTRIQTFRNPQKNNKWDSAEYDLTEEFKLIL NIDGTLKSQDVETYKKFMSLFKLMLQLRNSVTGTDIDYMISPVTDKTGTHFDSRENIKN LPADADANGAYCIALKGLYEIKQITENWKEDGKFSRDKLKISNKDWFDFIQNKRYL
[0213] SEQ ID NO:31 is Chimera 16:
[0214] MNNGTNNFQNFIGISSLQKTLRNALIPTETTQQFIVKNGIIKEDELRGENRQILKD IMDDYYRGFISETLSSIDDIDWTSLFEKMEIQLKNGDNKDTLIKEQAEKRKAIYKKFADD DRFKNMFSAKLISDILPEFVIHNNNYSASEKEEKTQVIKLFSRFATSFKDYFKNRANCFSA DDISSSSCHRIVNDNAEIFFSNALVYRRIVKNLSNDDINKISGDIKDSLKEMSLEEIYSYEK YGEFITQEGISFYNDICGKVNSFMNLYCQKNKENKNLYKLRKLHKQILCIADTSYEVPY KFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNLDKIYIVSKFYESVSQKTYRD WETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSITEINELVSNYKLCPDDNIKA ETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVLDVIMNAFHWCSVFMTEELV DKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQKPYSTKKIKLNFGIPTLADGWSKSKEY SNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDYKKMIYNLLPGPNKMIPKVFL SSKTGVETYKPSAYILEGYKQNKHLKSSKDFDITFCHDLIDYFKNCIAIHPEWKNFGFDFSDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKGQLYLFQIYNKDFSKKSSGND NLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPIIHKKGSILVNRTYEAEEKDQF GNIQIVRKTIPENIYQELYKYFNDKSDKELSDEAAKLKNWGHHEAATNIVKDYRYTYD KYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGIDRGERNLIYVSVIDTCGNIVEQK SFNIVNGYDYQIKLKQQEGARQIARKEWKEIGKIKEIKEGYLSLVIHEISKMVIKYNAIIA MEDLSYGFKKGRFKVERQVYQKFETMLINKLNYLVFKDISITENGGLLKGYQLTYIPEK LKNVGHQCGCIFYVPAAYTSKIDPTTGFVDPFVWKTIKNHESRKHFLEGFDFLHYDVKT GDFILHFKMNRNLSFQRGLPGFMPAWDIVFEKNETQFDAKGTPFIAGKRIVPVIENHRFT GRYRDLYPANELIALLEEKGIVFRDGSNILPKLLENDDSHAIDTMVALIRSVLQMRNSNAATGEDYINSPVRDLNGVCFDSRFQNPEWPMDADANGAYCIALKGLYEIKQITENWKED GKFSRDKLKISNKDWFDFIQNKRYL
[0215] SEQ ID NO:32 is Chimeric nuclease (Seq ID NO: 4) expression vector EFla promoter:
[0216] GCCATCTTCCTCTGGTAAGAACCTCGACCTCCCCACCCTGCCATTGCAGGAA AGGTTGGGTCTGGGTAATACCCCAGGGCATGGGATGTCTGAAGttaacgtcaggtggcacttttcg gggaaatgtgcgcggaacccctatttgtttatttttctaaatacattcaaatatgtatccgctcatgagacaataaccctgataaatgcttcaataa tattgaaaaaggaagagtatgagtattcaacatttccgtgtcgcccttattcccttttttgcggcattttgccttcctgtttttgctcacccagaaac gctggtgaaagtaaaagatgctgaagatcagttgggtgcacgagtgggttacatcgaactggatctCaacagcggtaagatccttgagagt tttcgccccgaagaacgttttccaatgatgagcacttttaaagttctgctatgtggcgcggtattatcccgtattgacgccgggcaagagcaac tcggtcgccgcatacactattctcagaatgacttggttgagtactcaccagtcacagaaaagcatcttacggatggcatgacagtaagagaat tatgcagtgctgccataaccatgagtgataacactgcggccaacttacttctgacaacgatcggaggaccgaaggagctaaccgcttttttg cacaacatgggggatcatgtaactcgccttgatcgttgggaaccggagctgaatgaagccataccaaacgacgagcgtgacaccacgat gcctgtagcaatggcaacaacgttgcgcaaactattaactggcgaactacttactctagcttcccggcaacaattaatagactggatggagg cggataaagttgcaggaccacttctgcgctcggcccttccggctggctggtttattgctgataaatctggagccggtgagcgtgggtctcgc ggtatcattgcagcactggggccagatggtaagccctcccgtatcgtagttatctacacgacggggagtcaggcaactatggatgaacgaa atagacagatcgctgagataggtgcctcactgattaagcattggtaactgtcagaccaagtttactcatatatactttagattgatttaaaacttc atttttaatttaaaaggatctaggtgaagatcctttttgataatctcatgaccaaaatcccttaacgtgagttttcgttccactgagcgtcagaccc cgtagaaaagatcaaaggatcttcttgagatcctttttttctgcgcgtaatctgctgcttgcaaacaaaaaaaccaccgctaccagcggtggttt gtttgccggatcaagagctaccaactctttttccgaaggtaactggcttcagcagagcgcagataccaaatactgttcttctagtgtagccgta gttaggccaccacttcaagaactctgtagcaccgcctacatacctcgctctgctaatcctgttaccagtggctgctgccagtggcgataagtc gtgtcttaccgggttggactcaagacgatagttaccggataaggcgcagcggtcgggctgaacggggggttcgtgcacacagcccagctt ggagcgaacgacctacaccgaactgagatacctacagcgtgagctatgagaaagcgccacgcttcccgaagggagaaaggcggacag gtatccggtaagcggcagggtcggaacaggagagcgcacgagggagcttccagggggaaacgcctggtatctttatagtcctgtcgggt ttcgccacctctgacttgagcgtcgatttttgtgatgctcgtcaggggggcggagcctatggaaaaacgccagcaacggccgaattaggct gaatttatttctgaattacccgggTGGGTGGATAAACTTGCTGATGCTGTAGATCTTAGCCTCTACA TGAGATCATGTGGAAAATCTGAAAGCATTTTAGGTTCCTTATGTTTGCAATCAAATA ACTGTACACCTTTTAATTTAAAAAGTACCATGAGGCACACACACACACTCGCAGGA ACTTTTTGGCGTAACAAAACTAGAATTAGATCTAAAGCTAACTGTAGGACTGAGTCT ATTCTAAACTGAAAGCCTGGACATCTGGAGTACCAGGGGGAGATGACGTGTTACGG GCTTCCATAAAAGCAGCTGGCTTTGAATGGAAGGAGCCAAGAGGCCAGCACAGGA GCGGATTCGTCGCTTTCACGGCCATCGAGCCGAACCTCTCGCAAGTCCGTGAGCCGT TAAGGAGGCCCCCAGTCCCGACCCTTCGCCCCAAGCCCCTCGGGGTCCCCGGGCCT GGTACTCCTTGCCACACGGGAGGGGCGCGGAAGCCGGGGCGGAGGAGGAGCCAACCCCGGGCTGGGCTGAGACCCGCAGAGGAAGACGCTCTAGGGATTTGTCCCGGACTAGCGAGATGGCAAGGCTGAGGACGGGAGGCTGATTGAGAGGCGAAGGTACACCCTA ATCTCAATACAACCTTTGGAGCTAAGCCAGCAATGGTAGAGGGAAGATTCTGCACGTCCCTTCCAGGCGGCCTccccgtcaccaccccccccaacccgccccGACCGGAGCTGAGAGTAATTC ATACAAAAGGACTCGCCCCTGCCTTGGGGAATCCCAGGGACCGTCGTTAAACTCCCACTAACGTAGAACCCAGAGATCGCTGCGTTCCCGCCCCCTCACCCGCCCGCTCTCGT CATCACTGAGGTGGAGAAGAGCATGCGTGAGGCTCCGGTGCCCGTCAGTGGGCAGA GCGCACATCGCCCACAGTCCCCGAGAAGTTGGGGGGAGGGGTCGGCAATTGAACCG GTGCCTAGAGAAAGTGGCGCGGGGTAAACTGGGAAAGTGATGTCGTGTACTGGCTC CGCCTTTTTCCCGAGGGTGGGGGAGAACCGTATATAAGTGCAGTAGTCGCCGTGAA CGTTCTTTTTCGCAACGGGTTTGCCGCCAGAACACAGGTAAGTGCCGTGTGTGGTTCCCGCGGGCCTGGCCTCTTTACGGGTTATGGCCCTTGCGTGCCTTGAATTACTTCCAC GCCCCTGGCTGCAGTACGTGATTCTTGATCCCGAGCTTCGGGTTGGAAGTGGGTGGGAGAGTTCGAGGCCTTGCGCTTAAGGAGCCCCTTCGCCTCGTGCTTGAGTTGAGGCCT GGCCTGGGCGCTGGGGCCGCCGCGTGCGAATCTGGTGGCACCTTCGCGCCTGTCTCGCTGCTTTCGATAAGTCTCTAGCCATTTAAAATTTTTGATGACCTGCTGCGACGCTTT TTTTCTGGCAAGATAGTCTTGTAAATGCGGGCCAAGATCTGCACACTGGTATTTCGG TTTTTGGGGCCGCGGGCGGCGACGGGGCCCGTGCGTCCCAGCGCACATGTTCGGCG AGGCGGGGCCTGCGAGCGCGGCCACCGAGAATCGGACGGGGGTAGTCTCAAGCTGGCCGGCCTGCTCTGGTGCCTGGCCTCGCGCCGCCGTGTATCGCCCCGCCCTGGGCGG CAAGGCTGGCCCGGTCGGCACCAGTTGCGTGAGCGGAAAGATGGCCGCTTCCCGGC CCTGCTGCAGGGAGCTCAAAATGGAGGACGCGGCGCTCGGGAGAGCGGGCGGGTGAGTCACCCACACAAAGGAAAAGGGCCTTTCCGTCCTCAGCCGTCGCTTCATGTGACT CCACGGAGTACCGGGCGCCGTCCAGGCACCTCGATTAGTTCTCGAGCTTTTGGAGTA CGTCGTCTTTAGGTTGGGGGGAGGGGTTTTATGCGATGGAGTTTCCCCACACTGAGT GGGTGGAGACTGAAGTTAGGCCAGCTTGGCACTTGATGTAATTCTCCTTGGAATTTG CCCTTTTTGAGTTTGGATCTTGGTTCATTCTCAAGCCTCAGACAGTGGTTCAAAGTTT TTTTCTTCCATTTCAGGTGTCGTGAAAACTACCCCTAAAAGCCAccATGGGCcctaagaag aagcggaaggttggtattcacggggtgcctgcggct AT G AAC AAT GGT AC AA AT AATTTC C AG AAC TT C A TCGGCATTTCAAGTCTTCAGAAGACCCTCAGAAATGCACTTATCCCTACCGAGACTACACAGCAGTTTATTGTAAAAAACGGGATCATTAAAGAAGACGAGCTTCGAGGAGAA AATAGGCAGATTCTCAAGGATATTATGGACGACTACTACCGTGGGTTCATATCCGA AACTTTGTCCTCCATAGACGACATCGACTGGACTAGCCTGTTCGAGAAAATGGAAA TTCAACTTAAGAATGGGGATAACAAAGACACACTAATCAAAGAGCAGGCAGAAAA AAGGAAAGCCATTTATAAGAAATTCGCCGATGACGATAGATTTAAGAATATGTTCAGTGCCAAGCTCATCTCAGATATCCTGCCGGAGTTTGTCATACATAATAATAACTACAGTGCCTCGGAAAAGGAAGAGAAGACTCAGGTTATCAAGCTGTTCTCTCGATTTGCCACGAGCTTCAAAGATTATTTCAAGAATAGAGCCAACTGTTTCAGTGCAGATGACATTTCTTCCTCCTCCTGTCACCGGATAGTGAATGACAACGCCGAGATATTTTTTTCAAACGCACTTGTATACCGAAGAATTGTGAAGAATCTGAGCAATGACGACATTAACAAGATTTCAGGCGATATAAAGGACTCCCTTAAAGAGATGAGCCTCGAAGAAATTTATAGCTATGAGAAGTACGGCGAGTTCATAACGCAAGAAGGGATTAGCTTCTACAACGACATCTGCGGCAAAGTCAATTCCTTTATGAACCTGTACTGCCAGAAAAATAAAGAGAACAAAAATCTGTACAAATTGAGGAAACTCCACAAACAGATCCTGTGTATTGCAGATACCTCTTACGAAGTCCCCTATAAATTCGAGTCCGACGAAGAGGTATACCAATCCGTCAACGGTTTTCTCGACAATATCAGTTCCAAACACATTGTGGAGAGGCTGCGGAAGATCGGTGACAATTATAATGGTTACAACTTGGATAAAATTTATATAGTGTCAAAGTTCTATGAATCCGTAAGCCAAAAGACATATCGGGATTGGGAGACCATTAATACTGCACTCGAAATCCATTACAATAACATCCTGCCAGGTAACGGGAAAAGTAAGGCGGATAAGGTTAAAAAGGCTGTCAAGAACGACCTGCAAAAGTCAATTACAGAAATCAATGAGCTGGTGAGCAACTACAAGCTGTGCCCCGACGACAATATCAAAGCCGAAACCTACATACATGAAATCTCCCACATACTGAACAACTTTGAGGCCCAGGAGCTGAAATATAATCCCGAGATTCACCTGGTCGAAAGCGAACTGAAAGCAAGCGAGCTGAAGAACGTGCTGGACGTCATAATGAATGCATTCCATTGGTGTAGTGTATTTATGACCGAAGAACTAGTTGACAAAGATAACAACTTTTATGCCGAACTGGAGGAAATCTACGACGAGATCTATACTGTTATCAGCCTATATAACCTCGTGCGGAATTACGTCACTCAGAAACCGTACAGCACTAAAAAAATCAAGCTGAATTTTGGAATCCCCACGTTAGCAGATGGCTGGTCCAAGTCTAAAGAGTATAGCAACAACGCCATCATCCTGATGCGAGACAACTTATATTATCTCGGGATCTTCAACGCCAAGAATAAGCCCGATAAGAAAATTATTGAGGGCAATACCAGCGAGAATAAAGGAGACTACAAGAAGATGATTTACAACCTGCTCCCAGGACCTAACAAGATGATTCCTAAGGTGTTTCTGTCTTCCAAGACTGGTGTGGAAACCTATAAGCCATCAGCTTACATCCTGGAGGGATACAAACAAAACAAGCATCTCAAATCTAGCAAGGACTTCGACATTACCTTCTGTCACGACCTTATAGACTATTTTAAGAACTGCATTGCCATTCACCCAGAGTGGAAGAACTTTGGGTTCGACTTCTCTGACACATCGACATATGAAGATATATCAGGCTTTTACCGCGAGGTTGAGCTGCAGGGATACAAGATCGACTGGACCTATATTAGCGAAAAGGACATTGACCTCCTGCAGGAGAAGGGCCAGCTCTATCTATTCCAGATTTATAATAAGGATTTCAGCAAAAAGAGTTCTGGTAACGATAACTTGCATACCATGTACCTTAAGAACTTGTTCAGCGAAGAGAATTTGAAAGACATCGTGCTTAAGCTGAACGGGGAGGCTGAAATCTTTTTTCGGAAGTCATCTATCAAGAATCCAATCATCCACAAAAAAGGAAGCATCTTGGTGAACCGGACGTACGAGGCGGAAGAGAAGGATCAGTTTGGGAACATTCAGATTGTGCGCAAAACTATTCCTGAGAATATCTATCAGGAGCTTTACAAGTACTTTAATGACAAGTCAGACAAGGAGCTCAGCGATGAAGCTGCGAAGCTGAAAAATGTGGTCGGGCATCACGAAGCGGCCACAAACATCGTTAAAGATTACAGATATACTTACGATAAATATTTTCTCCACATGCCCATCACGATCAACTTTAAGGCTAACAAAACTAGTTTCATCAATGATCGCATCCTACAGTATATTGCAAAAGAAAAAGATTTACATGTGATTGGCATAGACAGGGGTGAGCGCAATTTGATTTACGTCTCTGTGATCGACACCTGTGGCAACATCGTTGAACAGAAGAGCTTTAACATCGTCAATGGGTACGATTACCAGATTAAACTGAAACAACAGGAGGGAGCTAGACAAATTGCTAGAAAGGAGTGGAAAGAGATAGGTAAGATAAAGGAGATAAAGGAAGGCTACTTGTCATTAGTGATTCACGAGATATCGAAAATGGTGATTAAATACAACGCTATTATCGCTATGGAGGATCTGTCGTATGGGTTCAAGAAAGGAAGGTTCAAAGTGGAGCGCCAAGTTTATCAGAAGTTCGAAACAATGCTGATAAACAAGCTCAATTACCTCGTGTTCAAGGACATCTCTATTACCGAGAACGGGGGACTGTTGAAAGGCTATCAGCTCACCTACATCCCGGAGAAGCTTAAAAATGTGGGCCACCAGTGCGGATGCATCTTTTATGTGCCAGCCGCTTATACAAGTAAAATCGACCCTACCACGGGTTTCGCTAACGTTCTGAACCTCTCCAAAGTCCGAAATGTGGATGCCATCAAGAGCTTCTTCTCTAATTTTAACGAGATATCATACTCAAAGAAGGAGGCCCTCTTCAAGTTCAGCTTCGACCTGGATAGTCTGAGTAAGAAGGGATTCAGTAGCTTTGTGAAGTTCAGTAAATCTAAATGGAATGTTTATACGTTCGGGGAGCGGATCATAAAACCTAAAAATAAACAAGGCTACCGGGAGGACAAGAGAATTAATCTGACATTTGAGATGAAGAAACTGCTGAACGAATATAAAGTCAGTTTTGACCTTGAGAATAACCTTATCCCAAACCTCACCTCAGCCAACCTCAAAGATACTTTTTGGAAGGAATTATTCTTTATCTTCAAGACAACCCTGCAGCTGCGGAATAGCGTGACCAACGGTAAGGAAGACGTGCTGATCTCTCCAGTCAAGAACGCAAAGGGCGAATTTTTCGTGTCTGGGACACATAATAAGACTTTGCCTCAGGACTGTGATGCGAATGGGGCATACTGCATCGCCCTGAAGGGCCTGTACGAAATCAAGCAGATCACTGAGAACTGGAAAGAAGATGGTAAGTTCTCCCGGGACAAGCTAAAAATCTCTAACAAGGATTGGTTTGACTTCATACAGAACAAACGCTATCTGaagcgtcctgctgccaccaaaaaggccggacaggctaagaaaaagaagTGATCAGCCTCGACTGTGCCTTCTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGTGGATGCGGTGGGCTCTATGGCGTACGACGGTCTCCGATCGATGTCTCGGCACTGGTGTATACTCCTGAATCATTGGAATCTAGAGATTTGATGGAAGGGGGTGTCTCTATCCAAGAAGAGCCAATTTCCTCTTTCACTTTCTCCTTGGAGAGCTGATCTTCAGAGAACACAGTGGGGCTTGTTCTGGAGGTCCATGGTAGT CCAGCT
[0217] SEQ ID NO: 32:
[0218] gccatcttcctctggtaagaacctcgacctccccaccctgccattgcaggaaaggttgggtctgggtaataccccagggcatg ggatgtctgaagttaacgtcaggtggcacttttcggggaaatgtgcgcggaacccctatttgtttatttttctaaatacattcaaatatgtatccg ctcatgagacaataaccctgataaatgcttcaataatattgaaaaaggaagagtatgagtattcaacatttccgtgtcgcccttattcccttttttg cggcattttgccttcctgtttttgctcacccagaaacgctggtgaaagtaaaagatgctgaagatcagttgggtgcacgagtgggttacatcg aactggatctcaacagcggtaagatccttgagagttttcgccccgaagaacgttttccaatgatgagcacttttaaagttctgctatgtggcgc ggtattatcccgtattgacgccgggcaagagcaactcggtcgccgcatacactattctcagaatgacttggttgagtactcaccagtcacag aaaagcatcttacggatggcatgacagtaagagaattatgcagtgctgccataaccatgagtgataacactgcggccaacttacttctgaca acgatcggaggaccgaaggagctaaccgcttttttgcacaacatgggggatcatgtaactcgccttgatcgttgggaaccggagctgaatg aagccataccaaacgacgagcgtgacaccacgatgcctgtagcaatggcaacaacgttgcgcaaactattaactggcgaactacttactct agcttcccggcaacaattaatagactggatggaggcggataaagttgcaggaccacttctgcgctcggcccttccggctggctggtttattg ctgataaatctggagccggtgagcgtgggtctcgcggtatcattgcagcactggggccagatggtaagccctcccgtatcgtagttatctac acgacggggagtcaggcaactatggatgaacgaaatagacagatcgctgagataggtgcctcactgattaagcattggtaactgtcagac caagtttactcatatatactttagattgatttaaaacttcatttttaatttaaaaggatctaggtgaagatcctttttgataatctcatgaccaaaatcc cttaacgtgagttttcgttccactgagcgtcagaccccgtagaaaagatcaaaggatcttcttgagatcctttttttctgcgcgtaatctgctgctt gcaaacaaaaaaaccaccgctaccagcggtggtttgtttgccggatcaagagctaccaactctttttccgaaggtaactggcttcagcagag cgcagataccaaatactgttcttctagtgtagccgtagttaggccaccacttcaagaactctgtagcaccgcctacatacctcgctctgctaat cctgttaccagtggctgctgccagtggcgataagtcgtgtcttaccgggttggactcaagacgatagttaccggataaggcgcagcggtcg ggctgaacggggggttcgtgcacacagcccagcttggagcgaacgacctacaccgaactgagatacctacagcgtgagctatgagaaa gcgccacgcttcccgaagggagaaaggcggacaggtatccggtaagcggcagggtcggaacaggagagcgcacgagggagcttcca gggggaaacgcctggtatctttatagtcctgtcgggtttcgccacctctgacttgagcgtcgatttttgtgatgctcgtcaggggggcggagc ctatggaaaaacgccagcaacggccgaattaggctgaatttatttctgaattacccgggtgggtggataaacttgctgatgctgtagatctta gcctctacatgagatcatgtggaaaatctgaaagcattttaggttccttatgtttgcaatcaaataactgtacaccttttaatttaaaaagtaccat gaggcacacacacacactcgcaggaactttttggcgtaacaaaactagaattagatctaaagctaactgtaggactgagtctattctaaactg aaagcctggacatctggagtaccagggggagatgacgtgttacgggcttccataaaagcagctggctttgaatggaaggagccaagagg ccagcacaggagcggattcgtcgctttcacggccatcgagccgaacctctcgcaagtccgtgagccgttaaggaggcccccagtcccga cccttcgccccaagcccctcggggtccccgggcctggtactccttgccacacgggaggggcgcggaagccggggcggaggaggagc caaccccgggctgggctgagacccgcagaggaagacgctctagggatttgtcccggactagcgagatggcaaggctgaggacgggag gctgattgagaggcgaaggtacaccctaatctcaatacaacctttggagctaagccagcaatggtagagggaagattctgcacgtcccttcc aggcggcctccccgtcaccaccccccccaacccgccccgaccggagctgagagtaattcatacaaaaggactcgcccctgccttgggg aatcccagggaccgtcgttaaactcccactaacgtagaacccagagatcgctgcgttcccgccccctcacccgcccgctctcgtcatcact gaggtggagaagagcatgcgtgaggctccggtgcccgtcagtgggcagagcgcacatcgcccacagtccccgagaagttggggggaggggtcggcaattgaaccggtgcctagagaaagtggcgcggggtaaactgggaaagtgatgtcgtgtactggctccgcctttttcccgag ggtgggggagaaccgtatataagtgcagtagtcgccgtgaacgttctttttcgcaacgggtttgccgccagaacacaggtaagtgccgtgt gtggttcccgcgggcctggcctctttacgggttatggcccttgcgtgccttgaattacttccacgcccctggctgcagtacgtgattcttgatc ccgagcttcgggttggaagtgggtgggagagttcgaggccttgcgcttaaggagccccttcgcctcgtgcttgagttgaggcctggcctgg gcgctggggccgccgcgtgcgaatctggtggcaccttcgcgcctgtctcgctgctttcgataagtctctagccatttaaaatttttgatgacct gctgcgacgctttttttctggcaagatagtcttgtaaatgcgggccaagatctgcacactggtatttcggtttttggggccgcgggcggcgac ggggcccgtgcgtcccagcgcacatgttcggcgaggcggggcctgcgagcgcggccaccgagaatcggacgggggtagtctcaagc tggccggcctgctctggtgcctggcctcgcgccgccgtgtatcgccccgccctgggcggcaaggctggcccggtcggcaccagttgcgt gagcggaaagatggccgcttcccggccctgctgcagggagctcaaaatggaggacgcggcgctcgggagagcgggcgggtgagtca cccacacaaaggaaaagggcctttccgtcctcagccgtcgcttcatgtgactccacggagtaccgggcgccgtccaggcacctcgattag ttctcgagcttttggagtacgtcgtctttaggttggggggaggggttttatgcgatggagtttccccacactgagtgggtggagactgaagtta ggccagcttggcacttgatgtaattctccttggaatttgccctttttgagtttggatcttggttcattctcaagcctcagacagtggttcaaagttttt ttcttccatttcaggtgtcgtgaaaactacccctaaaagccaccatgggccctaagaagaagcggaaggttggtattcacggggtgcctgcg gctatgaacaatggtacaaataatttccagaacttcatcggcatttcaagtcttcagaagaccctcagaaatgcacttatccctaccgagacta cacagcagtttattgtaaaaaacgggatcattaaagaagacgagcttcgaggagaaaataggcagattctcaaggatattatggacgactac taccgtgggttcatatccgaaactttgtcctccatagacgacatcgactggactagcctgttcgagaaaatggaaattcaacttaagaatggg gataacaaagacacactaatcaaagagcaggcagaaaaaaggaaagccatttataagaaattcgccgatgacgatagatttaagaatatgt tcagtgccaagctcatctcagatatcctgccggagtttgtcatacataataataactacagtgcctcggaaaaggaagagaagactcaggtta tcaagctgttctctcgatttgccacgagcttcaaagattatttcaagaatagagccaactgtttcagtgcagatgacatttcttcctcctcctgtca ccggatagtgaatgacaacgccgagatatttttttcaaacgcacttgtataccgaagaattgtgaagaatctgagcaatgacgacattaacaa gatttcaggcgatataaaggactcccttaaagagatgagcctcgaagaaatttatagctatgagaagtacggcgagttcataacgcaagaag ggattagcttctacaacgacatctgcggcaaagtcaattcctttatgaacctgtactgccagaaaaataaagagaacaaaaatctgtacaaatt gaggaaactccacaaacagatcctgtgtattgcagatacctcttacgaagtcccctataaattcgagtccgacgaagaggtataccaatccgt caacggttttctcgacaatatcagttccaaacacattgtggagaggctgcggaagatcggtgacaattataatggttacaacttggataaaatt tatatagtgtcaaagttctatgaatccgtaagccaaaagacatatcgggattgggagaccattaatactgcactcgaaatccattacaataaca tcctgccaggtaacgggaaaagtaaggcggataaggttaaaaaggctgtcaagaacgacctgcaaaagtcaattacagaaatcaatgagc tggtgagcaactacaagctgtgccccgacgacaatatcaaagccgaaacctacatacatgaaatctcccacatactgaacaactttgaggc ccaggagctgaaatataatcccgagattcacctggtcgaaagcgaactgaaagcaagcgagctgaagaacgtgctggacgtcataatgaa tgcattccattggtgtagtgtatttatgaccgaagaactagttgacaaagataacaacttttatgccgaactggaggaaatctacgacgagatc tatactgttatcagcctatataacctcgtgcggaattacgtcactcagaaaccgtacagcactaaaaaaatcaagctgaattttggaatcccca cgttagcagatggctggtccaagtctaaagagtatagcaacaacgccatcatcctgatgcgagacaacttatattatctcgggatcttcaacg ccaagaataagcccgataagaaaattattgagggcaataccagcgagaataaaggagactacaagaagatgatttacaacctgctcccag gacctaacaagatgattcctaaggtgtttctgtcttccaagactggtgtggaaacctataagccatcagcttacatcctggagggatacaaac aaaacaagcatctcaaatctagcaaggacttcgacattaccttctgtcacgaccttatagactattttaagaactgcattgccattcacccagagtggaagaactttgggttcgacttctctgacacatcgacatatgaagatatatcaggcttttaccgcgaggttgagctgcagggatacaagatc gactggacctatattagcgaaaaggacattgacctcctgcaggagaagggccagctctatctattccagatttataataaggatttcagcaaa aagagttctggtaacgataacttgcataccatgtaccttaagaacttgttcagcgaagagaatttgaaagacatcgtgcttaagctgaacggg gaggctgaaatcttttttcggaagtcatctatcaagaatccaatcatccacaaaaaaggaagcatcttggtgaaccggacgtacgaggcgga agagaaggatcagtttgggaacattcagattgtgcgcaaaactattcctgagaatatctatcaggagctttacaagtactttaatgacaagtca gacaaggagctcagcgatgaagctgcgaagctgaaaaatgtggtcgggcatcacgaagcggccacaaacatcgttaaagattacagata tacttacgataaatattttctccacatgcccatcacgatcaactttaaggctaacaaaactagtttcatcaatgatcgcatcctacagtatattgca aaagaaaaagatttacatgtgattggcatagacaggggtgagcgcaatttgatttacgtctctgtgatcgacacctgtggcaacatcgttgaa cagaagagctttaacatcgtcaatgggtacgattaccagattaaactgaaacaacaggagggagctagacaaattgctagaaaggagtgg aaagagataggtaagataaaggagataaaggaaggctacttgtcattagtgattcacgagatatcgaaaatggtgattaaatacaacgctatt atcgctatggaggatctgtcgtatgggttcaagaaaggaaggttcaaagtggagcgccaagtttatcagaagttcgaaacaatgctgataaa caagctcaattacctcgtgttcaaggacatctctattaccgagaacgggggactgttgaaaggctatcagctcacctacatcccggagaagc ttaaaaatgtgggccaccagtgcggatgcatcttttatgtgccagccgcttatacaagtaaaatcgaccctaccacgggtttcgctaacgttct gaacctctccaaagtccgaaatgtggatgccatcaagagcttcttctctaattttaacgagatatcatactcaaagaaggaggccctcttcaag ttcagcttcgacctggatagtctgagtaagaagggattcagtagctttgtgaagttcagtaaatctaaatggaatgtttatacgttcggggagc ggatcataaaacctaaaaataaacaaggctaccgggaggacaagagaattaatctgacatttgagatgaagaaactgctgaacgaatataa agtcagttttgaccttgagaataaccttatcccaaacctcacctcagccaacctcaaagatactttttggaaggaattattctttatcttcaagaca accctgcagctgcggaatagcgtgaccaacggtaaggaagacgtgctgatctctccagtcaagaacgcaaagggcgaatttttcgtgtctg ggacacataataagactttgcctcaggactgtgatgcgaatggggcatactgcatcgccctgaagggcctgtacgaaatcaagcagatcac tgagaactggaaagaagatggtaagttctcccgggacaagctaaaaatctctaacaaggattggtttgacttcatacagaacaaacgctatct gaagcgtcctgctgccaccaaaaaggccggacaggctaagaaaaagaagtgatcagcctcgactgtgccttctagttgccagccatctgtt gtttgcccctcccccgtgccttccttgaccctggaaggtgccactcccactgtcctttcctaataaaatgaggaaattgcatcgcattgtctgag taggtgtcattctattctggggggtggggtggggcaggacagcaagggggaggattgggaagacaatagcaggcatgctgtggatgcg gtgggctctatggcgtacgacggtctccgatcgatgtctcggcactggtgtatactcctgaatcattggaatctagagatttgatggaagggg gtgtctctatccaagaagagccaatttcctctttcactttctccttggagagctgatcttcagagaacacagtggggcttgttctggaggtccat ggtagtccagct
[0219] SEQ ID NO: 33:PQPKKKPL
[0220] SEQ ID NO: 34:SALIKKKKKMAP
[0221] SEQ ID NO: 35:DRLRR
[0222] SEQ ID NO: 36:PKQKKRK1
[0223] SEQ ID NO: 37:RKLKKKIKKL
[0224] SEQ ID NO: 38:REKKKFLKRR
[0225] SEQ ID NO: 39:KRKGDEVDGVDEVAKKKSKK
[0226] SEQ ID NO: 40:RKCLQAGMNLEARKTKK
[0227] SEQ ID NO: 41:
[0228] PAAKKKKLD
[0229] SEQ ID NO: 42:
[0230] PAAKRVKLD
[0231] SEQ ID NO: 43:
[0232] RQRRNELKRSP
[0233] SEQ ID NO: 44:
[0234] NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY
[0235] SEQ ID NO: 45:
[0236] RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV
[0237] SEQ ID NO: 46:
[0238] VSRKRPRP
[0239] SEQ ID NO: 47
[0240] PPKKARED
[0241] SEQ ID NO: 48 cctctcaggcatggagtcct
[0242] SEQ ID NO: 49 tggcttgactcaggatttaa
[0243] SEQ ID NO: 50 aaagatgaaagggtactaaggaagtataaacgtatatgctatgttccgacaatctctctattaaccttaattaaactgacatttgtgtttctataat catgttttatgcactgcatctttcattattaaagaacccatcaaacgtcaaaattttaatacaaaattttacctgatagtatacgaatggcattgaac tttcataaagctaaagaaccgaaatatatagaacacctttcctgctttgtggcagaccggggtttaaaattaaagatgacaacatctaggagag tccgtacctcaggaaaaaaaccccagactgcgagtcaccttgcttttgagtgcaattccctaaaaccagtactctaatagtttttcctagaagtg gatctaggaaaatttaatttttacttcaaaatttagttagatttcatatatactcatttgaaccagactgtcaatggttacgaattagtcactccgtgg atagagtcgctagacagataaagcaagtaggtatcaacggactgaggggcagcacatctattgatgctatgccctcccgaatggtagaccg gggtcacgacgttactatggcgctctgggtgcgagtggccgaggtctaaatagtcgttatttggtcggtcggccttcccggctcgcgtcttcaccaggacgttgaaataggcggaggtaggtcagataattaacaacggcccttcgatctcattcatcaagcggtcaattatcaaacgcgttgca acaacggtaacgatgtccgtagcaccacagtgcgagcagcaaaccataccgaagtaagtcgaggccaagggttgctagttccgctcaatg tactagggggtacaacacgttttttcgccaatcgaggaagccaggaggctagcaacagtcttcattcaaccggcgtcacaatagtgagtacc aataccgtcgtgacgtattaagagaatgacagtacggtaggcattctacgaaaagacactgaccactcatgagttggttcagtaagactctta tcacatacgccgctggctcaacgagaacgggccgcagttatgccctattatggcgcggtgtatcgtcttgaaattttcacgagtagtaaccttt tgcaagaagccccgcttttgagagttcctagaatggcgacaactctaggtcaagctacattgggtgagcacgtgggttgactagaagtcgta gaaaatgaaagtggtcgcaaagacccactcgtttttgtccttccgttttacggcgttttttcccttattcccgctgtgcctttacaacttatgagtat gagaaggaaaaagttataataacttcgtaaatagtcccaataacagagtactcgcctatgtataaacttacataaatctttttatttgtttatcccc aaggcgcgtgtaaaggggcttttcacggtgagatctcggtggactgcaattgggcccctgtacctcttttagaccgtggtgtggaagatgtta ctcgacgcacaccgagggctcctcgtggggcacgacgactggctccggggggacttggggttccggttggcgctcttctactgggtcca ctcaccgggcgatggagaagaccaccggcggagggaggaaggaccggagggcctcgacgcgggaaagagtgaccaagagagaag acggcaaaaggcatcctgagagaagagactggactcagaggaaaccttgagacgtccaagataaacgaaaaagggtctactcgagaaa aagaccacaaacagagagactgatccacagattctgtcacaacacccacatccatgattgtgaccgagcacactgttccggtactccgacc acatttcgccggaacctcacacataattcatccgcgtgtcatccagacttgtctgaggggtagggttctggggtcgtgtgaatcggcacaag aaacgtgaaagacgtacagggggcagaccggaccgacaggggtcaccgaaggggtcacactgtaccacgtagagacggaatgtctag tacaaactctggaagttgtggggtcggtacatgcaacgataggtccgacacgatagggacatgcggagaccggcatggtgaccgtagca ctacctgaggccactgccccagtgggtgtgacacgggtagatgctccccatacgggagggggtacggtaggacgcagacctggaccga ccggccctggactgactgatggagtacttctaggagtggctcgcgccgatgtcgaagtggtggtgccggctcgccctttagcacgcactgt aattcctcttcgacacgatgcagcgggacctgaagctcgttctctaccggtgccgacgaaggtcgaggagggacctcttctcgatgctcga cggactgccggtccagtagtggtaaccgttactcgccaaggcgacgggactccgtgagaaggtcggaaggaaggacccgtacctttcga caccgtatgtgctctggtgaaagttaagttagtacttcacgctgcagctgtaagcgtttctagaaatgcgattgtgtcacgagagacccccttg ctggtacataggcccctatcggctagcctacgtcttcctctaatgtcgtaaccgagggtcatgatacttctaattctagtagcgcggtgggctc tcttttatatcacaaacctagccaccgagataggaccggagtgactcgtggaaagttgtctacacctaaaggtttgtccttatgctgcttagacc gggatcgtagcacgtatccttcacgaagcctaggccacttccctcgcctagagaggattgtacacctctacagcttctcttaggacctggtca ctcattcccgctccttcactagtttctcaagtacgccaaattccactcttaccttccttcgtacttgccggtgctcaagctttaactccctcttcctct ccctgccgggatgctcccgtgggtctgtcggttcgactttcactgtttcccgcccggagacggtaagcgaaccctgtaggactcgggtgtc aaatacatgccgaggttccggatacactttgtaggtcgactgtaagggctaatattctttgactcgaaggggctccccaaattcaccctttctca ctacttgaagctcctgcctccggaccactgacactgggtcctgtcgagggacgtcctaccctgggactagatgttccacttttactctccctgt ttaaaagggggactacctggacactacgtcttcttttgataccctaccctccggaggtggctttccgacataggtgcgctgccccacgactttc ctctttaggtggtccgagacttcgactttctaccccctgtaatggaccacctcaagttctgttagatgtaccggttctttggacacgtcgacggt ccgatgataatgcacctgtgttttgacctatagtgaagtgtgttgctcctgatgtgataacacctcgtcatacttgcctcgctcccctctgtggta gacaaggacccggtaccctgaccttcatggccgagtcccagatcaccttggcggagttcgctcctattgttataccgacactagtttctcaag tactccaaattccacgcgtacctcccgtcgtacttacccgtgcttaaactctaacttcctctcccgcttccctccggaatgctcccgtgtgtctga cggttcgactttcactggttccctcctggtgacggaaagcgaaccctataggacagaggagtcaaatacatgccttcattccggatacagttcgtagggcgactgtaaggactaatgttctttgacagaaagggtctcccgaaattcaccctctctcactacttaaaacttctacctccggaccact ggcactgtgtcctgaggagagacgtcctaccgtgagactagatgtttcagttttacgcgccgtggttaaaaggtgggctacccgggcacta cgtcttcttttgttaccccaccctccggtcgtgacttgccgacataggatctctgcctcacgacttcccgctttaggtggtccgggacttcgactt tctgccgccggtgatggaccacctcaagttttggtagatgtaccggttctttggtcacgtcgacgggccgataatgatacacctgtggttcga cctatagtgtagggtgttacttctgatgtggtaacaccttgtcatactctccagacttcctgcggtggtagacaaagacatgccgtacctactcg acatattcccgtacctcaggacaccgtaggtgctttgatggaagttgaggtagtacttcacactgcacctgtaggcgtttctggacatgcggtt gtgtcacgacagaccgccgtggtggtacatgggaccgtaacggctgtcctacgtcttcctctagtgacgggaccgtgggtcgtgttacttct agttccacccacagaaaggacggactcgactggacccgtccagtcgacaccccaggacaccacacacccctcgacagtgtaggtccca ggagtgacggacaggggaagggaggagtctagtaacgaggaggactcgcgttcatgaggcacacctagccgccgaggtaggaccgg agcgacaggtggaaggtcgtctacacctagtcgttcgtcctcatactgctcaggccggggaggtagcaggtggcgtttacgaagatccgc ctgatactgaatcaacgcaatgtgggaaagaactgttttggattgaacgcgtcttttgttctactctaaccgtaccgaaataaacaaaaaaaac aaaacaaaaccaaaaaaaaaaaaccgaactgagtcctaaatttttgaccttgccacttccactgtcgtcagccaacctcgctcgtagggggtt tcaagtgttacaccggctcctgaaactaacgtgtaacaacaaaaaaattatcagtaaggtttatactctacgcaacaatgtccttcagggaacg gtaggattttcggtggggtgaagagagattcctcttaccgggtcaggagagggttcaggtgtgtcccctccactatcgtaacgaaagcacat ttaatacattacgttttaaaaaaattagaagcggaattatgaaaaaataaaacaaaataaaacttactactcggaagcacgggggggaaggg ggaaaaaacagggggttgaactctacatacttccgaaaaccagagggaccctcacccacctccgtcggtcccgaatggacatgtgactga cctaggggcccaattgcagtccaccggccggcaacgaccgcaaaaaggtatccgaggcggggggactgctcgtagtgtttttagctgcg agttcagtctccaccgctttgggctgtcctgatatttctatggtccgcaaagggggaccttcgagggagcacgcgagaggacaaggctgg gacggcgaatggcctatggacaggcggaaagagggaagcccttcgcaccgcgaaagagtatcgagtgcgacatccatagagtcaagc cacatccagcaagcgaggttcgacccgacacacgtgcttggggggcaagtcgggctggcgacgcggaataggccattgatagcagaac tcaggttgggccattctgtgctgaatagcggtgaccgtcgtcggtgaccattgtcctaatcgtctcgctccatacatccgccacgatgtctcaa gaacttcaccaccggattgatgccgatgtgatcttcttgtcataaaccatagacgcgagacgacttcggtcaatggaagcctttttctcaacca tcgagaactaggccgtttgtttggtggcgaccatcgccaccaaaaaaacaaacgttcgtcgtctaatgcgcgtctttttttcctagagttcttcta ggaaactaga
[0244] SEQ ID NO: 51WHKILSAGIEAIQRNREDMTAQSGTTYIWIRSPKGDPGLAAIIGRSGREGAGSKDAIFW GAPLASRLLPGAVKDAEMWDILQQRSALTLLEGTLLKRLTTAMAVPMTTDREDNPIAE NLEPEWRDLRTVHDGMNHLFATLEKPGGITTLLLNAATNDSMTIAASCLERVTMGDTL HKETVPSYEVLDNQSYHIRRGLQEQGADIRSLVAGCLLVKFTSMMPFREEPRFSELIKGS NLDLEIYGVRAGLQDEADKVKVLTEPHAFVPLCFAAFFPILAVRFHQISM
[0245] SEQ ID NO: 52IMFETFNTPAMYVAIQAVLSLYASGRTTGIVMDSGDGVTHTVPIYEGYALPHAILRLDL AGRDLTDYLMKILTERGYSFTTTAEREIVRDIKEKLCYVALDFEQEMATAASSSSLEKSY ELPDGQVITIGNERFRCPEALFQPSFLGMESCGIHETTFNSIMKCDVDIRKDLYANTVLSGGTTMYPGIADRMQKEITALAPSTMKIKIIAPPERKYSVWIGGSILASLSTFQQMWISKQE YDESGPSIVHRKCFGSGEGSGSLLTCGDVEENPGPVSKGEEVIKEFMRFKVRMEGSMNGHEFEIEGEGEGRPYEGTQTAKLKVTKGGPLPFAWDILSPQFMYGSKAYVKHPADIPDYK KLSFPEGFKWERVMNFEDGGLVTVTQDSSLQDGTLIYKVKMRGTNFPPDGPVMQKKT MGWEASTERLYPRDGVLKGEIHQALKLKDGGHYLVEFKTIYMAKKPVQLPGYYYVDT KLDITSHNEDYTIVEQYERSEGRHHLFLGHGTGSTGSGSSGTASSEDNNMAVIKEFMRF KVRMEGSMNGHEFEIEGEGEGRPYEGTQTAKLKVTKGGPLPFAWDILSPQFMYGSKAY VKHPADIPDYKKLSFPEGFKWERVMNFEDGGLVTVTQDSSLQDGTLIYKVKMRGTNFP PDGPVMQKKTMGWEASTERLYPRDGVLKGEIHQALKLKDGGHYLVEFKTIYMAKKPV QLPGYYYVDTKLDITSHNEDYTIVEQYERSEGRHHLFLYGMDELYKGMESCGIHETTFN SIMKCDVDIRKDLYANTVLSGGTTMYPGIADRMQKEITALAPSTMKIK
[0246] SEQ ID NO: 53IIAPPERKYSVWIGGSILASLSTFQQMWISKQEYDESGPSIVHRKCF
[0247] SEQ ID NO: 54 aaagatgaaagggtactaaggaagtataaacgtatatgctatgttccgacaatctctctattaaccttaattaaactgacatttgtgtttctataat catgttttatgcactgcatctttcattattaaagaacccatcaaacgtcaaaattttaatacaaaattttacctgatagtatacgaatggcattgaac tttcataaagctaaagaaccgaaatatatagaacacctttcctgctttgtggcgagatcgtccggaccgtttaaagatgacaacatctaggaga gtccgtacctcaggaaaaaaaccccagactgcgagtcaccttgcttttgagtgcaattccctaaaaccagtactctaatagtttttcctagaagt ggatctaggaaaatttaatttttacttcaaaatttagttagatttcatatatactcatttgaaccagactgtcaatggttacgaattagtcactccgtg gatagagtcgctagacagataaagcaagtaggtatcaacggactgaggggcagcacatctattgatgctatgccctcccgaatggtagacc ggggtcacgacgttactatggcgctctgggtgcgagtggccgaggtctaaatagtcgttatttggtcggtcggccttcccggctcgcgtcttc accaggacgttgaaataggcggaggtaggtcagataattaacaacggcccttcgatctcattcatcaagcggtcaattatcaaacgcgttgc aacaacggtaacgatgtccgtagcaccacagtgcgagcagcaaaccataccgaagtaagtcgaggccaagggttgctagttccgctcaat gtactagggggtacaacacgttttttcgccaatcgaggaagccaggaggctagcaacagtcttcattcaaccggcgtcacaatagtgagta ccaataccgtcgtgacgtattaagagaatgacagtacggtaggcattctacgaaaagacactgaccactcatgagttggttcagtaagactct tatcacatacgccgctggctcaacgagaacgggccgcagttatgccctattatggcgcggtgtatcgtcttgaaattttcacgagtagtaacc ttttgcaagaagccccgcttttgagagttcctagaatggcgacaactctaggtcaagctacattgggtgagcacgtgggttgactagaagtc gtagaaaatgaaagtggtcgcaaagacccactcgtttttgtccttccgttttacggcgttttttcccttattcccgctgtgcctttacaacttatgag tatgagaaggaaaaagttataataacttcgtaaatagtcccaataacagagtactcgcctatgtataaacttacataaatctttttatttgtttatcc ccaaggcgcgtgtaaaggggcttttcacggtgagatctcggtggactgcaattgggcccctgtacctcttttagaccgtggtgtggaagatg ttactcgacgcacaccgagggctcctcgtggggcacgacgactggctccggggggacttggggttccggttggcgctcttctactgggtc cactcaccgggcgatggagaagaccaccggcggagggaggaaggaccggagggcctcgacgcgggaaagagtgaccaagagaga agacggcaaaaggcatcctgagagaagagactggactcagaggaaaccttgagacgtccaagataaacgaaaaagggtctactcgaga aaaagaccacaaacagagagactgatccacagattctgtcacaacacccacatccatgattgtgaccgagcacactgttccggtactccgaccacatttcgccggaacctcacacataattcatccgcgtgtcatccagacttgtctgaggggtagggttctggggtcgtgtgaatcggcaca agaaacgtgaaagacgtacagggggcagaccggaccgacaggggtcaccgaaggggtcacactgtaccacgtagagacggaatgtct agtacaaactctggaagttgtggggtcggtacatgcaacgataggtccgacacgatagggacatgcggagaccggcatggtgaccgtag cactacctgaggccactgccccagtgggtgtgacacgggtagatgctccccatacgggagggggtacggtaggacgcagacctggacc gaccggccctggactgactgatggagtacttctaggagtggctcgcgccgatgtcgaagtggtggtgccggctcgccctttagcacgcac tgtaattcctcttcgacacgatgcagcgggacctgaagctcgttctctaccggtgccgacgaaggtcgaggagggacctcttctcgatgctc gacggactgccggtccagtagtggtaaccgttactcgccaaggcgacgggactccgtgagaaggtcggaaggaaggacccgtacctttc gacaccgtatgtgctctggtgaaagttaagttagtacttcacgctgcagctgtaagcgtttctagaaatgcgattgtgtcacgagagaccccct tgctggtacataggcccctatcggctagcctacgtcttcctctaatgtcgtaaccgagggtcatgatacttctaattctagtagcgcggtgggc tctcttttatatcacaaacctagccaccgagataggaccggagtgactcgtggaaagttgtctacacctaaaggtttgtccttatgctgcttaga ccgggatcgtagcacgtatccttcacgaagcctaggccacttccctcgcctagagaggattgtacacctctacagcttctcttaggacctggt cactcattcccgctccttcactagtttctcaagtacgccaaattccactcttaccttccttcgtacttgccggtgctcaagctttaactccctcttcc tctccctgccgggatgctcccgtgggtctgtcggttcgactttcactgtttcccgcccggagacggtaagcgaaccctgtaggactcgggtg tcaaatacatgccgaggttccggatacactttgtaggtcgactgtaagggctaatattctttgactcgaaggggctccccaaattcaccctttct cactacttgaagctcctgcctccggaccactgacactgggtcctgtcgagggacgtcctaccctgggactagatgttccacttttactctccct gtttaaaagggggactacctggacactacgtcttcttttgataccctaccctccggaggtggctttccgacataggtgcgctgccccacgact ttcctctttaggtggtccgagacttcgactttctaccccctgtaatggaccacctcaagttctgttagatgtaccggttctttggacacgtcgacg gtccgatgataatgcacctgtgttttgacctatagtgaagtgtgttgctcctgatgtgataacacctcgtcatacttgcctcgctcccctctgtgg tagacaaggacccggtaccctgaccttcatggccgagtcccagatcaccttggcggagttcgctcctattgttataccgacactagtttctca agtactccaaattccacgcgtacctcccgtcgtacttacccgtgcttaaactctaacttcctctcccgcttccctccggaatgctcccgtgtgtct gacggttcgactttcactggttccctcctggtgacggaaagcgaaccctataggacagaggagtcaaatacatgccttcattccggatacag ttcgtagggcgactgtaaggactaatgttctttgacagaaagggtctcccgaaattcaccctctctcactacttaaaacttctacctccggacca ctggcactgtgtcctgaggagagacgtcctaccgtgagactagatgtttcagttttacgcgccgtggttaaaaggtgggctacccgggcact acgtcttcttttgttaccccaccctccggtcgtgacttgccgacataggatctctgcctcacgacttcccgctttaggtggtccgggacttcgac tttctgccgccggtgatggaccacctcaagttttggtagatgtaccggttctttggtcacgtcgacgggccgataatgatacacctgtggttcg acctatagtgtagggtgttacttctgatgtggtaacaccttgtcatactctccagacttcctgcggtggtagacaaagacatgccgtacctactc gacatattcccgtacctcaggacaccgtaggtgctttgatggaagttgaggtagtacttcacactgcacctgtaggcgtttctggacatgcgg ttgtgtcacgacagaccgccgtggtggtacatgggaccgtaacggctgtcctacgtcttcctctagtgacgggaccgtgggtcgtgttacttc tagttccacccacagaaaggacggactcgactggacccgtccagtcgacaccccaggacaccacacacccctcgacagtgtaggtccca ggagtgacggacaggggaagggaggagtctagtaacgaggaggactcgcgttcatgaggcacacctagccgccgaggtaggaccgg agcgacaggtggaaggtcgtctacacctagtcgttcgtcctcatactgctcaggccggggaggtagcaggtggcgtttacgaagatccgc ctgatactgaatcaacgcaatgtgggaaagaactgttttggattgaacgcgtcttttgttctactctaaccgtaccgaaataaacaaaaaaaac aaaacaaaaccaaaaaaaaaaaaccgaactgagtcctaaatttttgaccttgccacttccactgtcgtcagccaacctcgctcgtagggggtt tcaagtgttacaccggctcctgaaactaacgtgtaacaacaaaaaaattatcagtaaggtttatactctacgcaacaatgtccttcagggaacggtaggattttcggtggggtgaagagagattcctcttaccgggtcaggagagggttcaggtgtgtcccctccactatcgtaacgaaagcacat ttaatacattacgttttaaaaaaattagaagcggaattatgaaaaaataaaacaaaataaaacttactactcggaagcacgggggggaaggg ggaaaaaacagggggttgaactctacatacttccgaaaaccagagggaccctcacccacctccgtcggtcccgaatggacatgtgactga cctaggggcccaattgcagtccaccggccggcaacgaccgcaaaaaggtatccgaggcggggggactgctcgtagtgtttttagctgcg agttcagtctccaccgctttgggctgtcctgatatttctatggtccgcaaagggggaccttcgagggagcacgcgagaggacaaggctgg gacggcgaatggcctatggacaggcggaaagagggaagcccttcgcaccgcgaaagagtatcgagtgcgacatccatagagtcaagc cacatccagcaagcgaggttcgacccgacacacgtgcttggggggcaagtcgggctggcgacgcggaataggccattgatagcagaac tcaggttgggccattctgtgctgaatagcggtgaccgtcgtcggtgaccattgtcctaatcgtctcgctccatacatccgccacgatgtctcaa gaacttcaccaccggattgatgccgatgtgatcttcttgtcataaaccatagacgcgagacgacttcggtcaatggaagcctttttctcaacca tcgagaactaggccgtttgtttggtggcgaccatcgccaccaaaaaaacaaacgttcgtcgtctaatgcgcgtctttttttcctagagttcttcta ggaaactaga
[0248] SEQ ID NO: 55 aaagatgaaagggtactaaggaagtataaacgtatatgctatgttccgacaatctctctattaaccttaattaaactgacatttgtgtttctataat catgttttatgcactgcatctttcattattaaagaacccatcaaacgtcaaaattttaatacaaaattttacctgatagtatacgaatggcattgaac tttcataaagctaaagaaccgaaatatatagaacacctttcctgctttgtggcagaccggggtttaaaattaaagatgacaacatctaggagag tccgtacctcaggaaaaaaaccccagactgcgagtcaccttgcttttgagtgcaattccctaaaaccagtactctaatagtttttcctagaagtg gatctaggaaaatttaatttttacttcaaaatttagttagatttcatatatactcatttgaaccagactgtcaatggttacgaattagtcactccgtgg atagagtcgctagacagataaagcaagtaggtatcaacggactgaggggcagcacatctattgatgctatgccctcccgaatggtagaccg gggtcacgacgttactatggcgctctgggtgcgagtggccgaggtctaaatagtcgttatttggtcggtcggccttcccggctcgcgtcttca ccaggacgttgaaataggcggaggtaggtcagataattaacaacggcccttcgatctcattcatcaagcggtcaattatcaaacgcgttgca acaacggtaacgatgtccgtagcaccacagtgcgagcagcaaaccataccgaagtaagtcgaggccaagggttgctagttccgctcaatg tactagggggtacaacacgttttttcgccaatcgaggaagccaggaggctagcaacagtcttcattcaaccggcgtcacaatagtgagtacc aataccgtcgtgacgtattaagagaatgacagtacggtaggcattctacgaaaagacactgaccactcatgagttggttcagtaagactctta tcacatacgccgctggctcaacgagaacgggccgcagttatgccctattatggcgcggtgtatcgtcttgaaattttcacgagtagtaaccttt tgcaagaagccccgcttttgagagttcctagaatggcgacaactctaggtcaagctacattgggtgagcacgtgggttgactagaagtcgta gaaaatgaaagtggtcgcaaagacccactcgtttttgtccttccgttttacggcgttttttcccttattcccgctgtgcctttacaacttatgagtat gagaaggaaaaagttataataacttcgtaaatagtcccaataacagagtactcgcctatgtataaacttacataaatctttttatttgtttatcccc aaggcgcgtgtaaaggggcttttcacggtgagatctcggtggactgcaattgggcccctgtacctcttttagaccgtggtgtggaagatgtta ctcgacgcacaccgagggctcctcgtggggcacgacgactggctccggggggacttggggttccggttggcgctcttctactgggtcca ctcaccgggcgatggagaagaccaccggcggagggaggaaggaccggagggcctcgacgcgggaaagagtgaccaagagagaag acggcaaaaggcatcctgagagaagagactggactcagaggaaaccttgagacgtccaagataaacgaaaaagggtctactcgagaaa aagaccacaaacagagagactgatccacagattctgtcacaacacccacatccatgattgtgaccgagcacactgttccggtactccgacc acatttcgccggaacctcacacataattcatccgcgtgtcatccagacttgtctgaggggtagggttctggggtcgtgtgaatcggcacaag aaacgtgaaagacgtacagggggcagaccggaccgacaggggtcaccgaaggggtcacactgtaccacgtagagacggaatgtctagtacaaactctggaagttgtggggtcggtacatgcaacgataggtccgacacgatagggacatgcggagaccggcatggtgaccgtagca ctacctgaggccactgccccagtgggtgtgacacgggtagatgctccccatacgggagggggtacggtaggacgcagacctggaccga ccggccctggactgactgatggagtacttctaggagtggctcgcgccgatgtcgaagtggtggtgccggctcgccctttagcacgcactgt aattcctcttcgacacgatgcagcgggacctgaagctcgttctctaccggtgccgacgaaggtcgaggagggacctcttctcgatgctcga cggactgccggtccagtagtggtaaccgttactcgccaaggcgacgggactccgtgagaaggtcggaaggaaggacccgtacctttcga caccgtatgtgctctggtgaaagttaagttagtacttcacgctgcagctgtaagcgtttctagaaatgcgattgtgtcacgagagacccccttg ctggtacataggcccctatcggctagcctacgtcttcctctaatgtcgtaaccgagggtcatgatacttctaattctagtagcgcggtgggctc tcttttatatcacaaacctagccaccgagataggaccggagtgactcgtggaaagttgtctacacctaaaggtttgtccttatgctgcttagacc gggatcgtagcacgtatccttcacgaagcctaggccacttccctcgcctagagaggattgtacacctctacagcttctcttaggacctggtta cctctacctcttcctcctcttgcaccacatgcctggagaggggaaaatgggataacttcttcccaggcggccataggtcaacgtgttcatata cgtagtcatgcggtttgaccctcgttatcgtaaaagtttacgagattgtccacagctatagagtatagtcctcataaaactgtaatgtacgtccg accggctccgttactttttgaagccctactttggacttcttgtatagcgagagacgaggctcttgacacttctcaagaaatatgggcaaaaccg acctgaaatgtagccacaccgccaccgggggtgcttactctatatgtgtgactcccttgagttagtatcagacccttagcgggtcgggtgcta gcacaagtcaagagccttccctgagggtttccacgagctccaggttttttgacagtggacataattcttctaacaataggacctgtcgtttcaat tgaagcccccagtgctaacgtacctctgaaaatagttcttcgtgcatctcgagcctaaagtcggcaggagaaagcacggttaactacatttct tggctttcgttgtgcaacgaaatgactacttaagatcgcctagatggcctaacggctttccccagtcctattgggtgcttccgcgccagtgttct aagagtgtacgtttcctgggttatatgccgttagttcagaggggaccttgtcgataaaattggcaccacgggaaagtggtaccaaaaccgta caaatgatgtgacccgataaaacgcacaccgatggcacagcactacgactgttttaaactgcttcttgacaaggaggcctgggacgtcctg atgttcacgtggtcgcagtaagaccacgggtggaataaacggtaggacttgtttaggcttgactaactatttaagctagagaggttagaatga ctttagcggtcgccacctcgaggggagcggttcctccacccgctccggcagcgtgcctctaagttagacggaccgcaatctgtccccatac ctgactgactctgttggtcgcggaagtaataatgcggagcgcccctactgtttggaccccgcagcccgttccaacggggagacaagtttca cttccactagctgaacctgtgctttttttggaatccccagttggcggctccgctttagacacattttcccggttcggaatacgatccaataagttt gttgggactccgttggtctctttgataactgcttctcccgaccgacgtgtggcccctggaacccataatgctacttctgcttgtgaagaagtatc acctatccgacttcagagactagttcatggtcccgatggtccacggcgggcgactcgaactctcacaggaggatgttgtgggtttataaaag ctacgaccccatcggccgcacggtctagggctacgtccgctcgacgggccgcggcatcaccagtaccttttcccattttgatactgactcttc ctctagcacctgatacaattgtcagtccaacaattagtattcgcgaacgcaccgcctcatgccaagcacctgctccacggttttccggaatgt ccgttctagctgcgtttccactaggcgctctaaaacttttttggtgtccgattctacccgtacctcaggacaccgtaggtgctttgatggaagttg aggtagtacttcacactgcacctgtaggcgtttctggacatgcggttgtgtcacgacagaccgccgtggtggtacatgggaccgtaacggc tgtcctacgtcttcctctagtgacgggaccgtgggtcgtgttacttctagttccacccacagaaaggacggactcgactggacccgtccagtc gacaccccaggacaccacacacccctcgacagtgtaggtcccaggagtgacggacaggggaagggaggagtctagtaacgaggagg actcgcgttcatgaggcacacctagccgccgaggtaggaccggagcgacaggtggaaggtcgtctacacctagtcgttcgtcctcatact gctcaggccggggaggtagcaggtggcgtttacgaagatccgcctgatactgaatcaacgcaatgtgggaaagaactgttttggattgaac gcgtcttttgttctactctaaccgtaccgaaataaacaaaaaaaacaaaacaaaaccaaaaaaaaaaaaccgaactgagtcctaaatttttga ccttgccacttccactgtcgtcagccaacctcgctcgtagggggtttcaagtgttacaccggctcctgaaactaacgtgtaacaacaaaaaaattatcagtaaggtttatactctacgcaacaatgtccttcagggaacggtaggattttcggtggggtgaagagagattcctcttaccgggtcagg agagggttcaggtgtgtcccctccactatcgtaacgaaagcacatttaatacattacgttttaaaaaaattagaagcggaattatgaaaaaata aaacaaaataaaacttactactcggaagcacgggggggaagggggaaaaaacagggggttgaactctacatacttccgaaaaccagag ggaccctcacccacctccgtcggtcccgaatggacatgtgactgacctaggggcccaattgcagtccaccggccggcaacgaccgcaaa aaggtatccgaggcggggggactgctcgtagtgtttttagctgcgagttcagtctccaccgctttgggctgtcctgatatttctatggtccgca aagggggaccttcgagggagcacgcgagaggacaaggctgggacggcgaatggcctatggacaggcggaaagagggaagcccttc gcaccgcgaaagagtatcgagtgcgacatccatagagtcaagccacatccagcaagcgaggttcgacccgacacacgtgcttgggggg caagtcgggctggcgacgcggaataggccattgatagcagaactcaggttgggccattctgtgctgaatagcggtgaccgtcgtcggtga ccattgtcctaatcgtctcgctccatacatccgccacgatgtctcaagaacttcaccaccggattgatgccgatgtgatcttcttgtcataaacc atagacgcgagacgacttcggtcaatggaagcctttttctcaaccatcgagaactaggccgtttgtttggtggcgaccatcgccaccaaaaa aacaaacgttcgtcgtctaatgcgcgtctttttttcctagagttcttctaggaaactaga
[0249] SEQ ID NO: 56IMFETFNTPAMYVAIQAVLSLYASGRTTGIVMDSGDGVTHTVPIYEGYALPHAILRLDL AGRDLTDYLMKILTERGYSFTTTAEREIVRDIKEKLCYVALDFEQEMATAASSSSLEKSY ELPDGQVITIGNERFRCPEALFQPSFLGMESCGIHETTFNSIMKCDVDIRKDLYANTVLSGGTTMYPGIADRMQKEITALAPSTMKIKIIAPPERKYSVWIGGSILASLSTFQQMWISKQEYDESGPSIVHRKCFGSGEGSGSLLTCGDVEENPGPMEMEKEENWYGPLPFYPIEEGSA GIQLHKYMHQYAKLGAIAFSNALTGVDISYQEYFDITCRLAEAMKNFGMKPEEHIALCS ENCEEFFIPVLAGLYIGVAVAPTNEIYTLRELNHSLGIAQPTIVFSSRKGLPKVLEVQKTV TCIKKIVILDSKVNFGGHDCMETFIKKHVELGFQPSSFVPIDVKNRKQHVALLMNSSGST GLPKGVRITHEGAVTRFSHAKDPIYGNQVSPGTAILTWPFHHGFGMFTTLGYFACGYR WMLTKFDEELFLRTLQDYKCTSVILVPTLFAILNKSELIDKFDLSNLTEIASGGAPLAKE VGEAVARRFNLPGVRQGYGLTETTSAFIITPRGDDKPGASGKVAPLFKVKVIDLDTKKT LGVNRRGEICVKGPSLMLGYSNNPEATRETIDEEGWLHTGDLGYYDEDEHFFIVDRLKSLIKYQGYQVPPAELESVLLQHPNIFDAGVAGVPDPDAGELPGAVWMEKGKTMTEKEI VDYVNSQWNHKRLRGGVRFVDEVPKGLTGKIDAKVIREILKKPQAKMGMESCGIHET TFNSIMKCDVDIRKDLYANTVLSGGTTMYPGIADRMQKEITALAPSTMKIK
[0250] SEQ ID NO: 57
[0251] aaagatgaaagggtactaaggaagtataaacgtatatgctatgttccgacaatctctctattaaccttaattaaactgacatttgtg tttctataatcatgttttatgcactgcatctttcattattaaagaacccatcaaacgtcaaaattttaatacaaaattttacctgatagtatacgaatgg cattgaactttcataaagctaaagaaccgaaatatatagaacacctttcctgctttgtggcgagatcgtccggaccgtttaaagatgacaacat ctaggagagtccgtacctcaggaaaaaaaccccagactgcgagtcaccttgcttttgagtgcaattccctaaaaccagtactctaatagttttt cctagaagtggatctaggaaaatttaatttttacttcaaaatttagttagatttcatatatactcatttgaaccagactgtcaatggttacgaattagt cactccgtggatagagtcgctagacagataaagcaagtaggtatcaacggactgaggggcagcacatctattgatgctatgccctcccgaatggtagaccggggtcacgacgttactatggcgctctgggtgcgagtggccgaggtctaaatagtcgttatttggtcggtcggccttcccggc tcgcgtcttcaccaggacgttgaaataggcggaggtaggtcagataattaacaacggcccttcgatctcattcatcaagcggtcaattatcaa acgcgttgcaacaacggtaacgatgtccgtagcaccacagtgcgagcagcaaaccataccgaagtaagtcgaggccaagggttgctagt tccgctcaatgtactagggggtacaacacgttttttcgccaatcgaggaagccaggaggctagcaacagtcttcattcaaccggcgtcacaa tagtgagtaccaataccgtcgtgacgtattaagagaatgacagtacggtaggcattctacgaaaagacactgaccactcatgagttggttca gtaagactcttatcacatacgccgctggctcaacgagaacgggccgcagttatgccctattatggcgcggtgtatcgtcttgaaattttcacg agtagtaaccttttgcaagaagccccgcttttgagagttcctagaatggcgacaactctaggtcaagctacattgggtgagcacgtgggttga ctagaagtcgtagaaaatgaaagtggtcgcaaagacccactcgtttttgtccttccgttttacggcgttttttcccttattcccgctgtgcctttac aacttatgagtatgagaaggaaaaagttataataacttcgtaaatagtcccaataacagagtactcgcctatgtataaacttacataaatcttttta tttgtttatccccaaggcgcgtgtaaaggggcttttcacggtgagatctcggtggactgcaattgggcccctgtacctcttttagaccgtggtgt ggaagatgttactcgacgcacaccgagggctcctcgtggggcacgacgactggctccggggggacttggggttccggttggcgctcttct actgggtccactcaccgggcgatggagaagaccaccggcggagggaggaaggaccggagggcctcgacgcgggaaagagtgacca agagagaagacggcaaaaggcatcctgagagaagagactggactcagaggaaaccttgagacgtccaagataaacgaaaaagggtct actcgagaaaaagaccacaaacagagagactgatccacagattctgtcacaacacccacatccatgattgtgaccgagcacactgttccg gtactccgaccacatttcgccggaacctcacacataattcatccgcgtgtcatccagacttgtctgaggggtagggttctggggtcgtgtgaa tcggcacaagaaacgtgaaagacgtacagggggcagaccggaccgacaggggtcaccgaaggggtcacactgtaccacgtagagac ggaatgtctagtacaaactctggaagttgtggggtcggtacatgcaacgataggtccgacacgatagggacatgcggagaccggcatggt gaccgtagcactacctgaggccactgccccagtgggtgtgacacgggtagatgctccccatacgggagggggtacggtaggacgcaga cctggaccgaccggccctggactgactgatggagtacttctaggagtggctcgcgccgatgtcgaagtggtggtgccggctcgcccttta gcacgcactgtaattcctcttcgacacgatgcagcgggacctgaagctcgttctctaccggtgccgacgaaggtcgaggagggacctcttc tcgatgctcgacggactgccggtccagtagtggtaaccgttactcgccaaggcgacgggactccgtgagaaggtcggaaggaaggacc cgtacctttcgacaccgtatgtgctctggtgaaagttaagttagtacttcacgctgcagctgtaagcgtttctagaaatgcgattgtgtcacgag agacccccttgctggtacataggcccctatcggctagcctacgtcttcctctaatgtcgtaaccgagggtcatgatacttctaattctagtagcg cggtgggctctcttttatatcacaaacctagccaccgagataggaccggagtgactcgtggaaagttgtctacacctaaaggtttgtccttatg ctgcttagaccgggatcgtagcacgtatccttcacgaagcctaggccacttccctcgcctagagaggattgtacacctctacagcttctctta ggacctggttacctctacctcttcctcctcttgcaccacatgcctggagaggggaaaatgggataacttcttcccaggcggccataggtcaa cgtgttcatatacgtagtcatgcggtttgaccctcgttatcgtaaaagtttacgagattgtccacagctatagagtatagtcctcataaaactgta atgtacgtccgaccggctccgttactttttgaagccctactttggacttcttgtatagcgagagacgaggctcttgacacttctcaagaaatatg ggcaaaaccgacctgaaatgtagccacaccgccaccgggggtgcttactctatatgtgtgactcccttgagttagtatcagacccttagcgg gtcgggtgctagcacaagtcaagagccttccctgagggtttccacgagctccaggttttttgacagtggacataattcttctaacaataggac ctgtcgtttcaattgaagcccccagtgctaacgtacctctgaaaatagttcttcgtgcatctcgagcctaaagtcggcaggagaaagcacggt taactacatttcttggctttcgttgtgcaacgaaatgactacttaagatcgcctagatggcctaacggctttccccagtcctattgggtgcttccg cgccagtgttctaagagtgtacgtttcctgggttatatgccgttagttcagaggggaccttgtcgataaaattggcaccacgggaaagtggta ccaaaaccgtacaaatgatgtgacccgataaaacgcacaccgatggcacagcactacgactgttttaaactgcttcttgacaaggaggcctgggacgtcctgatgttcacgtggtcgcagtaagaccacgggtggaataaacggtaggacttgtttaggcttgactaactatttaagctagag aggttagaatgactttagcggtcgccacctcgaggggagcggttcctccacccgctccggcagcgtgcctctaagttagacggaccgcaa tctgtccccatacctgactgactctgttggtcgcggaagtaataatgcggagcgcccctactgtttggaccccgcagcccgttccaacgggg agacaagtttcacttccactagctgaacctgtgctttttttggaatccccagttggcggctccgctttagacacattttcccggttcggaatacga tccaataagtttgttgggactccgttggtctctttgataactgcttctcccgaccgacgtgtggcccctggaacccataatgctacttctgcttgt gaagaagtatcacctatccgacttcagagactagttcatggtcccgatggtccacggcgggcgactcgaactctcacaggaggatgttgtg ggtttataaaagctacgaccccatcggccgcacggtctagggctacgtccgctcgacgggccgcggcatcaccagtaccttttcccattttg atactgactcttcctctagcacctgatacaattgtcagtccaacaattagtattcgcgaacgcaccgcctcatgccaagcacctgctccacggt tttccggaatgtccgttctagctgcgtttccactaggcgctctaaaacttttttggtgtccgattctacccgtacctcaggacaccgtaggtgcttt gatggaagttgaggtagtacttcacactgcacctgtaggcgtttctggacatgcggttgtgtcacgacagaccgccgtggtggtacatggga ccgtaacggctgtcctacgtcttcctctagtgacgggaccgtgggtcgtgttacttctagttccacccacagaaaggacggactcgactgga cccgtccagtcgacaccccaggacaccacacacccctcgacagtgtaggtcccaggagtgacggacaggggaagggaggagtctagt aacgaggaggactcgcgttcatgaggcacacctagccgccgaggtaggaccggagcgacaggtggaaggtcgtctacacctagtcgtt cgtcctcatactgctcaggccggggaggtagcaggtggcgtttacgaagatccgcctgatactgaatcaacgcaatgtgggaaagaactgt tttggattgaacgcgtcttttgttctactctaaccgtaccgaaataaacaaaaaaaacaaaacaaaaccaaaaaaaaaaaaccgaactgagtc ctaaatttttgaccttgccacttccactgtcgtcagccaacctcgctcgtagggggtttcaagtgttacaccggctcctgaaactaacgtgtaac aacaaaaaaattatcagtaaggtttatactctacgcaacaatgtccttcagggaacggtaggattttcggtggggtgaagagagattcctctta ccgggtcaggagagggttcaggtgtgtcccctccactatcgtaacgaaagcacatttaatacattacgttttaaaaaaattagaagcggaatt atgaaaaaataaaacaaaataaaacttactactcggaagcacgggggggaagggggaaaaaacagggggttgaactctacatacttccg aaaaccagagggaccctcacccacctccgtcggtcccgaatggacatgtgactgacctaggggcccaattgcagtccaccggccggcaa cgaccgcaaaaaggtatccgaggcggggggactgctcgtagtgtttttagctgcgagttcagtctccaccgctttgggctgtcctgatatttc tatggtccgcaaagggggaccttcgagggagcacgcgagaggacaaggctgggacggcgaatggcctatggacaggcggaaagagg gaagcccttcgcaccgcgaaagagtatcgagtgcgacatccatagagtcaagccacatccagcaagcgaggttcgacccgacacacgtg cttggggggcaagtcgggctggcgacgcggaataggccattgatagcagaactcaggttgggccattctgtgctgaatagcggtgaccgt cgtcggtgaccattgtcctaatcgtctcgctccatacatccgccacgatgtctcaagaacttcaccaccggattgatgccgatgtgatcttcttg tcataaaccatagacgcgagacgacttcggtcaatggaagcctttttctcaaccatcgagaactaggccgtttgtttggtggcgaccatcgcc accaaaaaaacaaacgttcgtcgtctaatgcgcgtctttttttcctagagttcttctaggaaactaga
[0252] SEQ ID NO: 58 gaattagtcactccgtggatagagtcgctagacagataaagcaagtaggtatcaacggactgaggggcagcacatctattgatgctatgcc ctcccgaatggtagaccggggtcacgacgttactatggcgctctgggtgcgagtggccgaggtctaaatagtcgttatttggtcggtcggc cttcccggctcgcgtcttcaccaggacgttgaaataggcggaggtaggtcagataattaacaacggcccttcgatctcattcatcaagcggt caattatcaaacgcgttgcaacaacggtaacgatgtccgtagcaccacagtgcgagcagcaaaccataccgaagtaagtcgaggccaag ggttgctagttccgctcaatgtactagggggtacaacacgttttttcgccaatcgaggaagccaggaggctagcaacagtcttcattcaaccg gcgtcacaatagtgagtaccaataccgtcgtgacgtattaagagaatgacagtacggtaggcattctacgaaaagacactgaccactcatgagttggttcagtaagactcttatcacatacgccgctggctcaacgagaacgggccgcagttatgccctattatggcgcggtgtatcgtcttga aattttcacgagtagtaaccttttgcaagaagccccgcttttgagagttcctagaatggcgacaactctaggtcaagctacattgggtgagcac gtgggttgactagaagtcgtagaaaatgaaagtggtcgcaaagacccactcgtttttgtccttccgttttacggcgttttttcccttattcccgct gtgcctttacaacttatgagtatgagaaggaaaaagttataataacttcgtaaatagtcccaataacagagtactcgcctatgtataaacttacat aaatctttttatttgtttatccccaaggcgcgtgtaaaggggcttttcacggtggaaggccactgccccagtgggtgtgacacgggtagatgc tccccatacgggagggggtacggtaggacgcagacctggaccgaccggccctggactgactgatggagtacttctaggagtggctcgc gccgatgtcgaagtggtggtgccggctcgccctttagcacgcactgtaattcctcttcgacacgatgcagcgggacctgaagctcgttctct accggtgccgacgaaggtcgaggagggacctcttctcgatgctcgacggactgccggtccagtagtggtaaccgttactcgccaaggcg acgggactccgtgagaaggtcggaaggaaggacccactcacctctgacagagggccgagacggactgtactcccaatggggagcccc gacacgacaccttcgattcaggacgggagtaaagggagagtccgtacctcaggacaccgtaggtgctttgatggaagttgaggtagtactt cacactgcacctgtaggcgtttctggacatgcggttgtgtcacgacagaccgccgtggtggtacatgggaccgtaacggctgtcctacgtct tcctctagtgacgggaccgtgggtcgtgttacttctagttccacccacagaaaggacggactcgactggacccgtccagccgacacccca ggacaccacacacccctcgacagtgtaggtcccaggagtgacggacaggggaagggaggagtctagtaacgaggaggactcgcgttc atgaggcacacctagccgccgaggtaggaccggagcgacaggtggaaggtcgtctacacctagtcgttcgtcctcatactgctcaggcc ggggaggtagcaggtggcgtttacgaagatccgcctgatactgaatcaacgcaatgtgggaaagaactgttttggattgaacgcgtcttttg ttctactctaaccgtaccgctgctcgacatgttcattccgcgcggggggggattgcaatgaccggcttcggcgaaccttattccggccacac gcaaacagatatacaataaaaggtggtataacggcagaaaaccgttacactcccgggcctttggaccgggacagaagaactgctcgtaag gatccccagaaaggggagagcggtttccttacgttccagacaacttacagcacttccttcgtcaaggagaccttcgaagaacttctgtttgttg cagacatcgctgggaaacgtccgtcgccttggggggtggaccgctgtccacggagacgccggttttcggtgcacatattctatgtggacgt ttccgccgtgttggggtcacggtgcaacactcaacctatcaacacctttctcagtttaccgagaggagttcgcataagttgttccccgacttcct acgggtcttccatggggtaacataccctagactagaccccggagccacgtgtacgaaatgtacacaaatcagctccaatttttttgcagatcc ggggggcttggtgcccctgcaccaaaaggaaactttttgtgctactattataccggtgttggtaccagagttttccgcttcttctgttgtaccgtt attagttcctcaaatacgcgaaatttcatgtataccttcccagtcatttgcctgtacttaaactctatcttccacttcctcttccatctgggatgcttc cgtgtgtttgacggttcaactttcaatgatttccgccgggggagggaaaacgaaccctgtaaaacagtggggtcaagtacataccgtcatttc gtatacaatttgtgggtcggctgtaggggctaatggagttcgagagtaaagggctcccaaaattcaccctctcccaatacttgaaacttctacc cccacatcattgtcaatgagtcctaagttcggaggtcctacccctcaaatatatatttcattttgaagctccttgattgaagggatcactacctgg acagtacgtctttttctggtacccgacccttcggagttcacttgcctacatgggcctcctgcctcggaactttccgctctatttcgtcgctaacttt aacttcctacctcctgtaatgctacgactccatttctgatggatatttcgcttctttgggcaagtcgacgggcctcgaatattgcagttatattttga actgtagtgcagggtgttgcttctgatgtgctatcacctcgtcatacttgctcgacttccggccgtgaggtgccctccatacctgctcaacatat tcctcccgtccccgtcggacgactggacgccgctgcacctcctcttggggccggggtacggcggtggagccgaagacaagaaggagg agaaggaatgtggataccttcattctggcctccttggcaaccaacactttcagctccttccactattacgtcaggaagttacagagtttccctgg tcactaccaggatgagtcgtcgagtggaccaggtcccttagaggggactttggtaaaaacttcaacagggaaccagacgggcctgacccg taagtgtactctggggaccgttagaccgaaaagtagaaattacagagggttgtttacccaccaaaaatagaaacggtcggaccgggtggat cactcttccggaccgttggcccaacctgccagttacaccttccctcgccactcgagaaagccaccttgcagtcgctaaaccccccggaaccaacgcctgactttttagctagtagtctccctggcagatcgggatcgccctttgactacaggggctttgaaatacatacccgctttctgtccgga ctttaaaccctcccgcttggcggcacggagggcgggtccctaagtgacttagttaggaacagagtcctagaatgctaccgagggccgtcat gagacaccgacaggacaccgcagggcggtctgaggcattcgtctcccggtgagaggacctgggtacacgtgggctttccgggcttcagt aacaacagtgagcttaactttctgctggcggggcgtgctctatacacccactacctttggccagacgaaaacggcgcgcgctgacgggttc tacggccgttcatgatgacagtagcccccttagagtggtacagaaaggtaaacctttattgtcgttctgggcataacaccgtgaccgaaaact cttggcccccaaccttccaaagacgacagtgagaccgaatgaactagaagacagagacaagggaacagccttaagacgtagaggtcgc ccgtgaccaagaagcttcctttgcttttgcttactggctaggatgagccgctaaaactgtagccggcgcgagggctacttgccacttccactg tcgtcagccaacctcgctcgtagggggtttcaagtgttacaccggctcctgaaactaacgtgtaacaacaaaaaaattatcagtaaggtttata ctctacgcaacaatgtccttcagggaacggtaggattttcggtggggtgaagagagattcctcttaccgggtcaggagagggttcaggtgt gtcccctccactatcgtaacgaaagcacatttaatacattacgttttaaaaaaattagaagcggaattatgaaaaaataaaacaaaataaaactt actactcggaagcacgggggggaagggggaaaaaacagggggttgaactctacatacttccgaaaaccagagggaccctcacccacct ccgtcggtcccgaatggacatgtgactgaactctggtcaacttattttcacgtgtggaatttttactccggttcacactgaaacaccacaccga cccaacccccgtcgtctcccacttgggacgtcctcccacttgggacgttttcccaccccgtcacccccggttgaacaggaatgggtctcacg tccacacacctctagggaggacggaactgtaactcgtcggaatctcccaccccctccgagtccccagtccagagacaaggacgaataac ccctcaaggaccggaccgggaagatacagaggggtccatggggtcaaaaagacccaagtgggtctcacgtctacgaactcctccaccct tccctgataaacccccacagaccgagtccacggtacggagtgaccccgaccaaccgtggacgtaaaggaccctcaccccgacagagtc ccatcgacccgtgccacaagggaactcacccccacatcacccacaaggatcgacggtgcggaaacggaagtggataccctagcaccga cagtcggaactcccagtcggaccgggtccgagggtatccgaatcctctccggcgttaaggatggacaagtaggtcttaagtgcaccccga gtggagctggtgcgaccatcgccaccaaaaaaacaaacgttcgtcgtctaatgcgcgtctttttttcctagagttcttctaggaaactagaaaa gatgaaagggtactaaggaagtataaacgtatatgctatgttccgacaatctctctattaaccttaattaaactgacatttgtgtttctataatcatg ttttatgcactgcatctttcattattaaagaacccatcaaacgtcaaaattttaatacaaaattttacctgatagtatacgaatggcattgaactttca taaagctaaagaaccgaaatatatagaacacctttcctgctttgtggcagaccggggtttaaaattaaagatgacaacatctaaccgaactga gtcctaaattaaaaaaccattatcgctactgattatgcatctacatgacggttcatcctttcagggtattccagtacatgacccgtattacggtcc gcccggtaaatggcagtaactgcagttatcccccgcatgaaccgtatactatgtgaactacatgacggttcacccgtcaaatggcatttatga ggtgggtaactgcagttacctttcacttaagcaagccgacgccgctcgccatagtcgagtgagtttccgccattatgccaataggtgtcttagt cccctattgcgtcctttcttgtacactcgttttccggtcgttttccggtccttggcatttttccggcgcaacgaccgcaaaaaggtatccgaggc ggggggactgctcgtagtgtttttagctgcgagttcagtctccaccgctttgggctgtcctgatatttctatggtccgcaaagggggaccttcg agggagcacgcgagaggacaaggctgggacggcgaatggcctatggacaggcggaaagagggaagcccttcgcaccgcgaaagag tatcgagtgcgacatccatagagtcaagccacatccagcaagcgaggttcgacccgacacacgtgcttggggggcaagtcgggctggcg acgcggaataggccattgatagcagaactcaggttgggccattctgtgctgaatagcggtgaccgtcgtcggtgaccattgtcctaatcgtc tcgctccatacatccgccacgatgtctcaagaacttcaccaccggattgatgccgatgtgatcttcttgtcataaaccatagacgcgagacga cttcggtcaatggaagcctttttctcaaccatcgagaactaggccgtttgtttggtggcgaccatcgccaccaaaaaaacaaacgttcgtcgt ctaatgcgcgtctttttttcctagagttcttctaggaaactagaaaagatgccccagactgcgagtcaccttgcttttgagtgcaattccctaaaaccagtactctaatagtttttcctagaagtggatctaggaaaatttaatttttacttcaaaatttagttagatttcatatatactcatttgaaccagactg tcaatggttac
[0253] SEQ ID NO: 59KILSAGIEAIQRNREDMTAQSGTTYIWIRSPKGDPGLAAIIGRSGREGAGSKDAIFWGAP LASRLLPGAVKDAEMWDILQQRSALTLLEGTLLKRLTTAMAVPMTTDREDNPIAENLE PEWRDLRTVHDGMNHLFATLEKPGGITTLLLNAATNDSMTIAASCLERVTMGDTLHKE TVPSYEVLDNQSYHIRRGLQEQGADIRSLVAGCLLVKFTSMMPFREEPRFSELIKGSNLD LEIYGVRAGLQDEADKVKVLTEPHAFVPLCFAAFFPILAVRFHQISM
[0254] SEQ ID NO: 60MVSKGEEDNMAIIKEFMRFKVHMEGSVNGHEFEIEGEGEGRPYEGTQTAKLKVTKGGP LPFAWDILSPQFMYGSKAYVKHPADIPDYLKLSFPEGFKWERVMNFEDGGWTVTQDS SLQDGEFIYKVKLRGTNFPSDGPVMQKKTMGWEASSERMYPEDGALKGEIKQRLKLK DGGHYDAEVKTTYKAKKPVQLPGAYNVNIKLDITSHNEDYTIVEQYERAEGRHSTGGM DELYKEGRGSLLTCGDVEENPGPMPPPRLLFFLLFLTPMEVRPEEPLWKVEEGDNAVLQCLKGTSDGPTQQLTWSRESPLKPFLKLSLGLPGLGIHMRPLAIWLFIFNVSQQMGGFYL CQPGPPSEKAWQPGWTVNVEGSGELFRWNVSDLGGLGCGLKNRSSEGPSSPSGKLMSP KLYVWAKDRPEIWEGEPPCLPPRDSLNQSLSQDLTMAPGSTLWLSCGVPPDSVSRGPLS WTHVHPKGPKSLLSLELKDDRPARDMWVMETGLLLPRATAQDAGKYYCHRGNLTMS FHLEITARPVLWHWLLRTGGWKVSAVTLAYLIFCLCSLVGILHLQRALVLRRKRKRMT DPTRRF
[0255] SEQ ID NO: 61 cggtagaaggagaccattcttggagctggaggggtgggacggtaacgtcctttccaacccagacccattatggggtcccgtaccctacag acttcaattgcagtccaccgtgaaaagcccctttacacgcgccttggggataaacaaataaaaagatttatgtaagtttatacataggcgagta ctctgttattgggactatttacgaagttattataactttttccttctcatactcataagttgtaaaggcacagcgggaataagggaaaaaacgccg taaaacggaaggacaaaaacgagtgggtctttgcgaccactttcattttctacgacttctagtcaacccacgtgctcacccaatgtagcttgac ctagagttgtcgccattctaggaactctcaaaagcggggcttcttgcaaaaggttactactcgtgaaaatttcaagacgatacaccgcgccat aatagggcataactgcggcccgttctcgttgagccagcggcgtatgtgataagagtcttactgaaccaactcatgagtggtcagtgtcttttc gtagaatgcctaccgtactgtcattctcttaatacgtcacgacggtattggtactcactattgtgacgccggttgaatgaagactgttgctagcc tcctggcttcctcgattggcgaaaaaacgtgttgtaccccctagtacattgagcggaactagcaacccttggcctcgacttacttcggtatggt ttgctgctcgcactgtggtgctacggacatcgttaccgttgttgcaacgcgtttgataattgaccgcttgatgaatgagatcgaagggccgttg ttaattatctgacctacctccgcctatttcaacgtcctggtgaagacgcgagccgggaaggccgaccgaccaaataacgactatttagacct cggccactcgcacccagagcgccatagtaacgtcgtgaccccggtctaccattcgggagggcatagcatcaatagatgtgctgcccctca gtccgttgatacctacttgctttatctgtctagcgactctatccacggagtgactaattcgtaaccattgacagtctggttcaaatgagtatatatg aaatctaactaaattttgaagtaaaaattaaattttcctagatccacttctaggaaaaactattagagtactggttttagggaattgcactcaaaagcaaggtgactcgcagtctggggcatcttttctagtttcctagaagaactctaggaaaaaaagacgcgcattagacgacgaacgtttgttttttt ggtggcgatggtcgccaccaaacaaacggcctagttctcgatggttgagaaaaaggcttccattgaccgaagtcgtctcgcgtctatggttt atgacaagaagatcacatcggcatcaatccggtggtgaagttcttgagacatcgtggcggatgtatggagcgagacgattaggacaatggt caccgacgacggtcaccgctattcagcacagaatggcccaacctgagttctgctatcaatggcctattccgcgtcgccagcccgacttgcc ccccaagcacgtgtgtcgggtcgaacctcgcttgctggatgtggcttgactctatggatgtcgcactcgatactctttcgcggtgcgaaggg cttccctctttccgcctgtccataggccattcgccgtcccagccttgtcctctcgcgtgctccctcgaaggtccccctttgcggaccatagaaa tatcaggacagcccaaagcggtggagactgaactcgcagctaaaaacactacgagcagtccccccgcctcggatacctttttgcggtcgtt gccggcttaatccgacttaaataaagacttaatgggcccacccacctatttgaacgactacgacatctagaatcggagatgtactctagtaca ccttttagactttcgtaaaatccaaggaatacaaacgttagtttattgacatgtggaaaattaaatttttcatggtactccgtgtgtgtgtgtgagc gtccttgaaaaaccgcattgttttgatcttaatctagatttcgattgacatcctgactcagataagatttgactttcggacctgtagacctcatggt ccccctctactgcacaatgcccgaaggtattttcgtcgaccgaaacttaccttcctcggttctccggtcgtgtcctcgcctaagcagcgaaagt gccggtagctcggcttggagagcgttcaggcactcggcaattcctccgggggtcagggctgggaagcggggttcggggagccccagg ggcccggaccatgaggaacggtgtgccctccccgcgccttcggccccgcctcctcctcggttggggcccgacccgactctgggcgtctc cttctgcgagatccctaaacagggcctgatcgctctaccgttccgactcctgccctccgactaactctccgcttccatgtgggattagagttat gttggaaacctcgattcggtcgttaccatctcccttctaagacgtgcagggaaggtccgccggaggggcagtggtggggggggttgggc ggggctggcctcgactctcattaagtatgttttcctgagcggggacggaaccccttagggtccctggcagcaatttgagggtgattgcatctt gggtctctagcgacgcaagggcgggggagtgggcgggcgagagcagtagtgactccacctcttctcgtacgcactccgaggccacgg gcagtcacccgtctcgcgtgtagcgggtgtcaggggctcttcaacccccctccccagccgttaacttggccacggatctctttcaccgcgcc ccatttgaccctttcactacagcacatgaccgaggcggaaaaagggctcccaccccctcttggcatatattcacgtcatcagcggcacttgc aagaaaaagcgttgcccaaacggcggtcttgtgtccattcacggcacacaccaagggcgcccggaccggagaaatgcccaataccggg aacgcacggaacttaatgaaggtgcggggaccgacgtcatgcactaagaactagggctcgaagcccaaccttcacccaccctctcaagct ccggaacgcgaattcctcggggaagcggagcacgaactcaactccggaccggacccgcgaccccggcggcgcacgcttagaccacc gtggaagcgcggacagagcgacgaaagctattcagagatcggtaaattttaaaaactactggacgacgctgcgaaaaaaagaccgttcta tcagaacatttacgcccggttctagacgtgtgaccataaagccaaaaaccccggcgcccgccgctgccccgggcacgcagggtcgcgtg tacaagccgctccgccccggacgctcgcgccggtggctcttagcctgcccccatcagagttcgaccggccggacgagaccacggaccg gagcgcggcggcacatagcggggcgggacccgccgttccgaccgggccagccgtggtcaacgcactcgcctttctaccggcgaaggg ccgggacgacgtccctcgagttttacctcctgcgccgcgagccctctcgcccgcccactcagtgggtgtgtttccttttcccggaaaggcag gagtcggcagcgaagtacactgaggtgcctcatggcccgcggcaggtccgtggagctaatcaagagctcgaaaacctcatgcagcaga aatccaacccccctccccaaaatacgctacctcaaaggggtgtgactcacccacctctgacttcaatccggtcgaaccgtgaactacattaa gaggaaccttaaacgggaaaaactcaaacctagaaccaagtaagagttcggagtctgtcaccaagtttcaaaaaaagaaggtaaagtcca cagcacttttgatggggattttcggtggtacccgggattcttcttcgccttccaaccataagtgccccacggacgccgatacttgttaccatgtt tattaaaggtcttgaagtagccgtaaagttcagaagtcttctgggagtctttacgtgaatagggatggctctgatgtgtcgtcaaataacatttttt gccctagtaatttcttctgctcgaagctcctcttttatccgtctaagagttcctataatacctgctgatgatggcacccaagtataggctttgaaac aggaggtatctgctgtagctgacctgatcggacaagctcttttacctttaagttgaattcttacccctattgtttctgtgtgattagtttctcgtccgtcttttttcctttcggtaaatattctttaagcggctactgctatctaaattcttatacaagtcacggttcgagtagagtctataggacggcctcaaaca gtatgtattattattgatgtcacggagccttttccttctcttctgagtccaatagttcgacaagagagctaaacggtgctcgaagtttctaataaag ttcttatctcggttgacaaagtcacgtctactgtaaagaaggaggaggacagtggcctatcacttactgttgcggctctataaaaaaagtttgc gtgaacatatggcttcttaacacttcttagactcgttactgctgtaattgttctaaagtccgctatatttcctgagggaatttctctactcggagctt ctttaaatatcgatactcttcatgccgctcaagtattgcgttcttccctaatcgaagatgttgctgtagacgccgtttcagttaaggaaatacttgg acatgacggtctttttatttctcttgtttttagacatgtttaactcctttgaggtgtttgtctaggacacataacgtctatggagaatgcttcagggga tatttaagctcaggctgcttctccatatggttaggcagttgccaaaagagctgttatagtcaaggtttgtgtaacacctctccgacgccttctagc cactgttaatattaccaatgttgaacctattttaaatatatcacagtttcaagatacttaggcattcggttttctgtatagccctaaccctctggtaatt atgacgtgagctttaggtaatgttattgtaggacggtccattgcccttttcattccgcctattccaatttttccgacagttcttgctggacgttttca gttaatgtctttagttactcgaccactcgttgatgttcgacacggggctgctgttatagtttcggctttggatgtatgtactttagagggtgtatga cttgttgaaactccgggtcctcgactttatattagggctctaagtggaccagctttcgcttgactttcgttcgctcgacttcttgcacgacctgca gtattacttacgtaaggtaaccacatcacataaatactggcttcttgatcaactgtttctattgttgaaaatacggcttgacctcctttagatgctgc tctagatatgacaatagtcggatatattggagcacgccttaatgcagtgagtctttggcatgtcgtgatttttttagttcgacttaaaaccttaggg gtgcaatcgtctaccgaccaggttcagatttctcatatcgttgttgcggtagtaggactacgctctgttgaatataatagagccctagaagttgc ggttcttattcgggctattcttttaataactcccgttatggtcgctcttatttcctctgatgttcttctactaaatgttggacgagggtcctggattgtt ctactaaggattccacaaagacagaaggttctgaccacacctttggatattcggtagtcgaatgtaggacctccctatgtttgttttgttcgtaga gtttagatcgttcctgaagctgtaatggaagacagtgctggaatatctgataaaattcttgacgtaacggtaagtgggtctcaccttcttgaaac ccaagctgaagagactgtgtagctgtatacttctatatagtccgaaaatggcgctccaactcgacgtccctatgttctagctgacctggatata atcgcttttcctgtaactggaggacgtcctcttcccggtcgagatagataaggtctaaatattattcctaaagtcgtttttctcaagaccattgcta ttgaacgtatggtacatggaattcttgaacaagtcgcttctcttaaactttctgtagcacgaattcgacttgcccctccgactttagaaaaaagcc ttcagtagatagttcttaggttagtaggtgttttttccttcgtagaaccacttggcctgcatgctccgccttctcttcctagtcaaacccttgtaagt ctaacacgcgttttgataaggactcttatagatagtcctcgaaatgttcatgaaattactgttcagtctgttcctcgagtcgctacttcgacgcttc gactttttacaccagcccgtagtgcttcgccggtgtttgtagcaatttctaatgtctatatgaatgctatttataaaagaggtgtacgggtagtgct agttgaaattccgattgttttgatcaaagtagttactagcgtaggatgtcatataacgttttctttttctaaatgtacactaaccgtatctgtccccac tcgcgttaaactaaatgcagagacactagctgtggacaccgttgtagcaacttgtcttctcgaaattgtagcagttacccatgctaatggtctaa tttgactttgttgtcctccctcgatctgtttaacgatctttcctcacctttctctatccattctatttcctctatttccttccgatgaacagtaatcactaa gtgctctatagcttttaccactaatttatgttgcgataatagcgatacctcctagacagcatacccaagttctttccttccaagtttcacctcgcgg ttcaaatagtcttcaagctttgttacgactatttgttcgagttaatggagcacaagttcctgtagagataatggctcttgccccctgacaactttcc gatagtcgagtggatgtagggcctcttcgaatttttacacccggtggtcacgcctacgtagaaaatacacggtcggcgaatatgttcattttag ctgggatggtgcccaaagcgattgcaagacttggagaggtttcaggctttacacctacggtagttctcgaagaagagattaaaattgctctat agtatgagtttcttcctccgggagaagttcaagtcgaagctggacctatcagactcattcttccctaagtcatcgaaacacttcaagtcatttag atttaccttacaaatatgcaagcccctcgcctagtattttggatttttatttgttccgatggccctcctgttctcttaattagactgtaaactctacttct ttgacgacttgcttatatttcagtcaaaactggaactcttattggaatagggtttggagtggagtcggttggagtttctatgaaaaaccttccttaa taagaaatagaagttctgttgggacgtcgacgccttatcgcactggttgccattccttctgcacgactagagaggtcagttcttgcgtttcccgcttaaaaagcacagaccctgtgtattattctgaaacggagtcctgacactacgcttaccccgtatgacgtagcgggacttcccggacatgctt tagttcgtctagtgactcttgacctttcttctaccattcaagagggccctgttcgatttttagagattgttcctaaccaaactgaagtatgtcttgttt gcgatagacttcgcaggacgacggtggtttttccggcctgtccgattctttttcttcactagtcggagctgacacggaagatcaacggtcggt agacaacaaacggggagggggcacggaaggaactgggaccttccacggtgagggtgacaggaaaggattattttactcctttaacgtag cgtaacagactcatccacagtaagataagaccccccaccccaccccgtcctgtcgttccccctcctaacccttctgttatcgtccgtacgaca cctacgccacccgagataccgcatgctgccagaggctagctacagagccgtgaccacatatgaggacttagtaaccttagatctctaaact accttcccccacagagataggttcttctcggttaaaggagaaagtgaaagaggaacctctcgactagaagtctcttgtgtcaccccgaacaa gacctccaggtaccatcaggtcga
[0256] SEQ ID NO: 62MSIQHFRVALIPFFAAFCLPVFAHPETLVKVKDAEDQLGARVGYIELDLNSGKILESFRP EERFPMMSTFKVLLCGAVLSRIDAGQEQLGRRIHYSQNDLVEYSPVTEKHLTDGMTVR ELCSAAITMSDNTAANLLLTTIGGPKELTAFLHNMGDHVTRLDRWEPELNEAIPNDERD TTMPVAMATTLRKLLTGELLTLASRQQLIDWMEADKVAGPLLRSALPAGWFIADKSGA GERGSRGIIAALGPDGKPSRIWIYTTGSQATMDERNRQIAEIGASLIKHW
[0257] SEQ ID NO: 63MGPKKKRKVGIHGVPAAMNNGTNNFQNFIGISSLQKTLRNALIPTETTQQFIVKNGIIKE DELRGENRQILKDIMDDYYRGFISETLSSIDDIDWTSLFEKMEIQLKNGDNKDTLIKEQA EKRKAIYKKFADDDRFKNMFSAKLISDILPEFVIHNNNYSASEKEEKTQVIKLFSRFATSF KDYFKNRANCFSADDISSSSCHRIVNDNAEIFFSNALVYRRIVKNLSNDDINKISGDIKDS LKEMSLEEIYSYEKYGEFITQEGISFYNDICGKVNSFMNLYCQKNKENKNLYKLRKLHK QILCIADTSYEVPYKFESDEEVYQSVNGFLDNISSKHIVERLRKIGDNYNGYNLDKIYIVS KFYESVSQKTYRDWETINTALEIHYNNILPGNGKSKADKVKKAVKNDLQKSITEINELV SNYKLCPDDNIKAETYIHEISHILNNFEAQELKYNPEIHLVESELKASELKNVLDVIMNAF HWCSVFMTEELVDKDNNFYAELEEIYDEIYTVISLYNLVRNYVTQKPYSTKKIKLNFGIP TLADGWSKSKEYSNNAIILMRDNLYYLGIFNAKNKPDKKIIEGNTSENKGDYKKMIYNL LPGPNKMIPKVFLSSKTGVETYKPSAYILEGYKQNKHLKSSKDFDITFCHDLIDYFKNCI AIHPEWKNFGFDFSDTSTYEDISGFYREVELQGYKIDWTYISEKDIDLLQEKGQLYLFQI YNKDFSKKSSGNDNLHTMYLKNLFSEENLKDIVLKLNGEAEIFFRKSSIKNPIIHKKGSIL VNRTYEAEEKDQFGNIQIVRKTIPENIYQELYKYFNDKSDKELSDEAAKLKNWGHHEA ATNIVKDYRYTYDKYFLHMPITINFKANKTSFINDRILQYIAKEKDLHVIGIDRGERNLIYVSVIDTCGNIVEQKSFNIVNGYDYQIKLKQQEGARQIARKEWKEIGKIKEIKEGYLSLVI HEISKMVIKYNAIIAMEDLSYGFKKGRFKVERQVYQKFETMLINKLNYLVFKDISITENG GLLKGYQLTYIPEKLKNVGHQCGCIFYVPAAYTSKIDPTTGFANVLNLSKVRNVDAIKS FFSNFNEISYSKKEALFKFSFDLDSLSKKGFSSFVKFSKSKWNVYTFGERIIKPKNKQGYREDKRINLTFEMKKLLNEYKVSFDLENNLIPNLTSANLKDTFWKELFFIFKTTLQLRNSVT NGKEDVLISPVKNAKGEFFVSGTHNKTLPQDCDANGAYCIALKGLYEIKQITENWKED GI<FSRDI<LI<ISNI<DWFDFIQNI<RYLI<RPAATI<I<AGQAI<I<I<I<
[0258] SEQ ID NO:64 (IDPTTGF) and SEQ ID NO:65 (DANGAY) are crossover points of a nuclease chimera that exchanges the Nuc domain based on ErCpfl .
[0259] Although the invention has been described with reference to the presently preferred embodiment, it should be understood that various modifications can be made without departing from the spirit of the invention. Accordingly, the invention is limited only by the following claims.
Claims
WHAT IS CLAIMED IS:
1. A nucleic acid-guided nuclease system comprising:(a) a polypeptide sequence comprising a nucleic acid-guided nuclease, wherein the polypeptide sequence has at least 90% sequence identity to:SEQ ID NO:1, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:9, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, or SEQ ID NO:31 ; or(b) a nucleic acid molecule encoding the polypeptide sequence of (a); and(c) an engineered guide nucleic acid comprising:(i) a region for complexing with the nucleic acid-guided nuclease, and(ii) a region for hybridizing with a target sequence, wherein the engineered guide nucleic acid comprises at least 95% sequence identity to SEQ ID NO: 13.
2. A nucleic acid-guided nuclease system comprising:(a) a polypeptide sequence comprising a nucleic acid-guided nuclease, wherein the polypeptide sequence has at least 90% sequence identity to:SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:9, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, or SEQ ID NO:31; and wherein the polypeptide sequence has less than 90% sequence identity to SEQ ID NO:2; or(b) a nucleic acid molecule encoding the polypeptide sequence of (a); and(c) an engineered guide nucleic acid comprising:(i) a region for complexing with the nucleic acid-guided nuclease; and(ii) a region for hybridizing with a target sequence, wherein the engineered guide nucleic acid comprises at least 95% sequence identity to SEQ ID NO: 12.
3. The nucleic acid-guided nuclease system of claim 1 or 2, wherein the target sequence encodes a protein.
4. A vector comprising at least one of:(a) a polynucleotide encoding the polypeptide sequence of claim 1(a); and(b) the engineered guide nucleic acid of claim 1(c).
5. A vector comprising at least one of:(a) a polynucleotide encoding the polypeptide sequence of claim 2(a); and(b) the engineered guide nucleic acid of claim 2(c).
6. The vector of claim 4 or 5, wherein the vector is a plasmid or a viral vector.
7. A delivery system configured to deliver one or more components of the nucleic acid-guided system of claims 1 or 2 or the vector of claim 4 or 5.
8. The delivery system of claim 7, comprising a delivery vehicle selected from a liposome, a particle, an exosome, a microvesicle or a viral vector.
9. The delivery system of claim 7, comprising a delivery method selected from electroporation, a gene-gun, calcium phosphate mediated transfer, nucleofection, sonoporation, heat shock, magneto fection, or micro injection.
10. A method of modifying a target nucleic acid comprising contacting the target nucleic acid with a nucleic acid-guided nuclease system of claim 1 or 2, wherein the guide sequence directs sequence-specific binding to the target nucleic acid sequence, whereby the target nucleic acid sequence, the expression of the target nucleic acid, or both are modified.
11. The method of claim 10, wherein modifying occurs in vitro, ex vivo, or in vivo.
12. The method of claim 10, wherein modifying the target nucleic acid comprises cleaving the target nucleic acid.
13. The method of claim 10, wherein modifying expression of the target nucleic acid comprises increasing or decreasing transcription or translation of the target nucleic acid.
14. The method of claim 10, wherein the target nucleic acid is in a prokaryotic cell.
15. The method of claim 10, wherein the target nucleic acid is in a eukaryotic cell.
16. The method of claim 15, wherein the eukaryotic cell is a mammalian cell.
17. The method of claim 16, wherein the mammalian cell is a human cell.
18. An isolated cell comprising a modified target nucleic acid of interest, wherein the target nucleic acid of interest has been modified according to the method of claim 10.
19. The cell of claim 18, wherein the modified target nucleic acid of interest results in the cell comprising altered expression of at least one gene product.
20. The cell of claim 19, wherein the expression of the at least one gene product is increased.
21. The cell of claim 19, wherein the expression of the at least one gene product is decreased.
22. The cell of claim 19, wherein the at least one gene product is modified.
23. A plant or animal model comprising one or more cells of claim 18.
24. A nucleic acid-guided nuclease system comprising:(a) a nucleic acid-guided nuclease comprising at least 90% amino acid sequence identity to a chimera nuclease based on ErCpf 1 , wherein the chimera nuclease based on ErCpf 1 comprises a nuclease (Nuc) domain substituted with a Nuc domain of Cpfl from another specie, or(b) a nucleic acid molecule encoding the nucleic acid-guided nuclease of (a); and(c) an engineered guide nucleic acid comprising:(i) a region for complexing with the nucleic acid-guided nuclease, and(ii) a region for hybridizing with a target sequence, wherein the engineered guide nucleic acid comprises at least 95% sequence identity to SEQ ID NO: 13.
25. The method of claim 24, wherein the Nuc domain substitution comprises SEQ ID NO:33 and / or SEQ ID NO:34.
Citation Information
Patent Citations
Novel engineered and chimeric nucleases
US20190359976A1
Novel crispr-associated protein and use thereof
US20210292722A1
Novel nucleic acid-guided nucleases
US20240026322A1
Cited By
Cas enzyme mutant, combined protein, nucleic acid molecule, recombinant vector, transgenic cell and application thereof
CN121653102A
A cas enzyme mutant and combination protein, nucleic acid molecule, recombinant vector, transgenic cell and application thereof
CN121653102B