Genome editing compositions and methods for treatment of alpha-1-antitrypsin deficiency
Patent Information
- Application Number
- PCT/US2026/019601
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-11-08
- Filing Date
- 2026-03-17
- Publication Date
- 2026-09-24
Smart Images

Figure US2026019601_24092026_PF_FP_ABST
Abstract
Description
WSGR Docket No. 59761-805.601GENOME EDITING COMPOSITIONS AND METHODS FOR TREATMENT OF ALPHA-1- ANTITRYPSIN DEFICIENCY CROSS-REFERENCE[1] This application claims the benefit of U. S. Provisional Application No. 63 / 773,257, filed on March 17, 2025, and U. S. Provisional Application No. 63 / 914,011, filed on November 8, 2025, each of which is incorporated herein by reference in its entirety.BACKGROUND[2] Alpha- 1 antitrypsin (AAT), also known as SERPINA1 (serine protease inhibitor, group A, member 1), is a secreted protein produced by hepatocytes and, to a lesser extent, by mononuclear phagocytes, neutrophils, and airway / intestinal epithelial cells. SERPINA1 is a circulating glycoprotein; it is water soluble and tissue diffusible. The primary role of SERPINA1 is serine protease regulation and the main site of this action is in the lungs. There, SERPINA1 binds and inactivates neutrophil elastase thereby protecting alveolar tissues from proteolytic degradation during inflammatory responses.[3] Mutations in the SERPINA1 gene can lead to inactive or defective AAT protein and result in genetic disorders such as alpha- 1 antitrypsin deficiency (AATD). In severe AATD patients, for example, patients homozygous for pathogenic mutations in the SERPINA1 gene, the SERPINA1 protein is misfolded and cannot be secreted by hepatocytes. Polymerization and aggregation of misfolded SERPINA1 proteins in the rough endoplasmic reticulum of such hepatocytes can result in severe liver damage as well as low levels of circulating functional SERPINA1, which can cause lung damage from neutrophil elastase and conditions such as pulmonary emphysema.[4] Currently there is no specific treatment for liver disease associated with AAT deficiency. AATD lung disease is often treated with one of several serum protein replacement products. However, long-term studies of the effectiveness of SERPINA1 replacement therapy are not available, and it does not reduce liver damage in AAT deficiency. There remains a need for an effective approach for treatment of AATD.SUMMARY[5] Disclosed herein are prime editing guide RNAs (PEgRNAs) and nucleic acid that encode said PEgRNAs, wherein the PEgRNAs comprise: (a) a spacer comprising a nucleotide sequence that is complementary to a search target sequence on a first strand of a SERPINA1 gene; (b) a gRNA core capable of binding to a Cas9 protein; and (c) an extension arm comprising: (i) an editing template comprising a nucleotide sequence that comprises a region of complementarity to an editing target sequence on a second strand of the SERPINA1 gene, and (ii) a primer binding site (PBS) capable of binding to a 5’ flap predicted to form following cutting of the second strand 3 nucleotides upstream of the PAM sequence associated with the search target sequence or protospacer on the first strand, wherein the first strand and second strand are complementary to each other, wherein the nucleotide sequence of the editing template encodes a wildtype amino acid sequence of an alpha-antitrypsin (AAT) protein andWSGR Docket No. 59761-805.601further encodes one or more synonymous nucleotide transversion edits compared to a wildtype SERPINA1 gene sequence.[6] Also disclosed herein are prime editing systems comprising: (a) any of the PEgRNAs disclosed herein and (b) a prime editor, or a nucleic acid encoding the prime editor, wherein the prime editor comprises: (i) a Cas9 domain; and (ii) a reverse transcriptase. The nucleic acid encoding the prime editor can be an mRNA. The systems can further comprise a ngRNA comprising (a) a ngRNA spacer complementary to a search target sequence on the second strand of the SEPRINA1 gene, and (b) a ngRNA core capable of binding to the Cas9 domain.[7] The prime editing systems disclosed herein can be encapsulated in lipid nanoparticles (LNPs). Accordingly, disclosed herein are methods of making and using LNP encapsulated prime editing systems and pharmaceutical compositions containing the LNP encapsulated prime editing systems. Such LNP encapsulated prime editing systems and pharmaceutical compositions can be useful for editing the SERPINA1 gene in cells and as a medicament for treating alpha- 1 antitrypsin disease (AATD) in subjects in need thereof.[8] Other aspects, embodiments, and features will be apparent from the following description, the drawings, and the claims.INCORPORATION BY REFERENCE[9] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The novel features of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings of which:
[0011] FIG. 1 depicts a schematic of a prime editing guide RNA (PEgRNA) binding to a doublestranded target DNA sequence.
[0012] FIG. 2 depicts a PEgRNA architectural overview in an exemplary schematic of PEgRNA designed for a prime editor.
[0013] FIG. 3 is a schematic showing the spacer and gRNA core part of an exemplary guide RNA, in two separate molecules. The rest of the PEgRNA structure is not shown.
[0014] FIG. 4 is a schematic of an exemplary PEgRNA conjugated to a nuclear localization signal (NLS) at its 3’ end. Depicted from left to right are PEgRNA spacer sequence, gRNA core secondary structure, PEgRNA extension arm, conjugation moiety, NLS peptide (represented as a coiled-ribbon schematic).WSGR Docket No. 59761-805.601
[0015] FIG. 5 is a UPLC-CAD chromatogram showing the resolution of lipid components of LNPs of the present disclosure.DETAILED DESCRIPTION
[0016] Provided herein, in some embodiments, are compositions and methods to edit the target gene SERPINA1 with prime editing. In certain embodiments, provided herein are compositions and methods for introducing nucleotide edits in the target SERPINA1 gene with prime editing. In certain embodiments, provided herein are compositions and methods for correction of mutations in the (SERPINA / ) gene associated with AATD. Compositions provided herein can comprise prime editors (PEs) that may use engineered guide polynucleotides, e.g., prime editing guide RNAs (PEgRNAs), that can direct PEs to specific DNA targets and can encode DNA edits on the target gene SERPINA1 that serve a variety of functions, including direct correction of disease-causing mutations. In some embodiments, the PEgRNA is designed to correct one or more mutations associated with AATD in the SERPINA1 gene. In certain embodiments, the PEgRNA is designed to introduce one or more mutations that improve editing outcome, e.g., efficiency of correction of the one or more AATD associated mutations in the SERPINA1 gene. In some embodiments, the PEgRNA is designed to be able to correct multiple mutations in of the SERPINA1 gene (a mutation “hotspot”).
[0017] The following description and examples illustrate embodiments of the present disclosure in detail. It is to be understood that this disclosure is not limited to the particular embodiments described herein and as such can vary. Those of skill in the art will recognize that there are numerous variations and modifications of this disclosure, which are encompassed within its scope. Although various features of the present disclosure can be described in the context of a single embodiment, the features can also be provided separately or in any suitable combination. Conversely, although the present disclosure can be described herein in the context of separate embodiments for clarity, the present disclosure can also be implemented in a single embodiment.Definitions
[0018] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art.
[0019] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent that the terms “including”, “includes”, “having”, “has”, “with”, or variants thereof as used herein mean “comprising”.
[0020] Unless otherwise specified, the words “comprising”, “comprise”, “comprises”, “having”, “have”, “has”, “including”, “includes”, “include”, “containing”, “contains” and “contain” are inclusive or open-ended and do not exclude additional, unrecited elements or method steps.
[0021] Reference to “some embodiments”, “an embodiment”, “one embodiment”, or “other embodiments” means that a particular feature or characteristic described in connection with theWSGR Docket No. 59761-805.601embodiments is included in at least one or more embodiments, but not necessarily all embodiments, of the present disclosure.
[0022] The term “about” or “approximately” in relation to a numerical means, a range of values that fall within 10% greater than or less than the value. For example, about x means x±(10% * x).
[0023] The term “between” means the range of numbers including the first and the last number in a range.
[0024] As used herein, a “cell” generally refers to a biological cell. A cell can be the basic structural, functional and / or biological unit of a living organism. A cell can originate from any organism having one or more cells.
[0025] In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. A cell may be of or derived from different tissues, organs, and / or cell types. In some embodiments, the cell is a primary cell. As used herein, the term “primary cell” means a cell isolated from an organism, e.g., a mammal, which is grown in tissue culture (i.e., in vitro) for the first time before subdivision and transfer to a subculture. In some embodiments, the cell is a primary hepatocyte.
[0026] In some embodiments, the cell is a stem cell. In some embodiments, the cell is a pluripotent cell (e.g., a pluripotent stem cell). In some embodiments, the cell (e.g., a stem cell) is an embryonic stem cell (ESC), tissue-specific stem cell, mesenchymal stem cell, or an induced pluripotent stem cell. In some embodiments, the cell is an induced pluripotent stem cell (iPSC). In some embodiments, the cell is a hepatic progenitor cell. In some embodiments, the cell is an embryonic stem cell -derived hepatocyte precursor cell.
[0027] In some embodiments, the cell is a differentiated cell. In some embodiments, the cell is a terminally differentiated cell. In some embodiments, the cell is anon-dividing cell. In some embodiments, the cell is a hepatocyte. In some embodiments, the cell is a fibroblast. In some embodiments, the cell is differentiated from an induced pluripotent stem cell.
[0028] In some non-limiting examples, mammalian cells or human cells, including primary cells and stem cells can be modified through introduction of one or more polynucleotides, polypeptide, and / or prime editing compositions (e.g., through transfection, transduction, electroporation and the like) and further passaged. Such modified cells can include epithelial cells (e.g., mammary epithelial cells, intestinal epithelial cells, hepatocytes), endothelial cells, glial cells, neural cells, formed elements of the blood (e.g., lymphocytes, bone marrow cells), precursors of any of these somatic cell types, and stem cells.
[0029] In some embodiments, a cell is not isolated from an organism but forms part of a tissue or organ of an organism, e.g., a mammal. In some non-limiting examples, the cells include muscle cells (e.g., cardiac muscle cells, smooth muscle cells, myosatellite cells), epithelial cells (e.g., mammary epithelial cells, intestinal epithelial cells, hepatocytes), endothelial cells (e.g., lung endothelial cells), pulmonary cells, glial cells, neural cells, formed elements of the blood (e.g., lymphocytes, bone marrow cells), precursors of any of these somatic cell types, and stem cells.WSGR Docket No. 59761-805.601
[0030] In some embodiments, the cell comprises a prime editor, a PEgRNA, or a prime editing composition disclosed herein. In some embodiments, the cell further comprises an ngRNA. In some embodiments, the cell is from a human subject. In some embodiments, the human subject has a disease or a condition or is at a risk of developing a disease or a condition associated with a mutation to be corrected by prime editing, for example, AATD. In some embodiments, the cell is from a human subject, and comprises a prime editor, a PEgRNA, or a prime editing composition to introduce one or more nucleotide edits into the SERPINA1 gene, e.g., one or more nucleotide edits that corrects one or more mutation associated with AATD, or one or more nucleotide edits that improves efficiency of the correction of the one or more mutations associated with AATD. In some embodiments, the cell is from a human subject and comprises a SERPINA1 gene, wherein one or more mutations associated in the SERPINA1 gene have been edited or corrected by prime editing. In some embodiments, the cell is from a human subject and comprises a SERPINA gene, wherein the SERPINA1 gene comprises one or more nucleotide edits introduced by prime editing compared to the endogenous SERPINA1 gene sequence in cell, e.g., one or more nucleotide edits that corrects one or more mutations associated with AATD, or one or more nucleotide edits that improves efficiency of the correction of the one or more mutations associated with AATD. In some embodiments, the cell is in a human subject, and comprises a prime editor, a PEgRNA, or a prime editing composition to introduce one or more nucleotide edits into the SERPINA1 gene, e.g., one or more nucleotide edits that corrects one or more mutation associated with AATD, or one or more nucleotide edits that improves efficiency of the correction of the one or more mutations associated with AATD. In some embodiments, the cell is in a human subject and comprises a SERPINA 1 gene, wherein one or more mutations associated in the SERPINA 1 gene have been edited or corrected by prime editing. In some embodiments, the cell is from a human subject and comprises a SERPINA gene, wherein the SERPINA 1 gene comprises one or more nucleotide edits introduced by prime editing compared to the endogenous SERPINA 1 gene sequence in cell, e.g., one or more nucleotide edits that corrects one or more mutations associated with AATD, or one or more nucleotide edits that improves efficiency of the correction of the one or more mutations associated with AATD.
[0031] In some embodiments, the cell edited by prime editing can be differentiated into, or give rise to recovery of a population of cells, e.g., epithelial cells or hepatocytes.
[0032] In some embodiments, the cell is in a subject, e.g., a human subject. In some embodiments, the cell is obtained from a subject prior to editing. For example, in some embodiments, the cell is obtained from a AATD patient having one or more mutations in the SERPINA 1 gene.
[0033] The term “substantially” as used herein may refer to a value approaching 100% of a given value. In some embodiments, the term may refer to an amount that may be at least about 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.9%, or 99.99% of a total amount. In some embodiments, the term may refer to an amount that may be about 100% of a total amount.
[0034] The terms “protein” and “polypeptide” can be used interchangeably to refer to a polymer of two or more amino acids joined by covalent bonds (e.g., an amide bond) that can adopt a three-dimensional conformation. In some embodiments, a protein or polypeptide comprises at least 10 amino acids, 15WSGR Docket No. 59761-805.601amino acids, 20 amino acids, 30 amino acids or 50 amino acids joined by covalent bonds (e.g., amide bonds). In some embodiments, a protein comprises at least two amide bonds. In some embodiments, a protein comprises multiple amide bonds. In some embodiments, a protein comprises an enzyme, enzyme precursor proteins, regulatory protein, structural protein, receptor, nucleic acid binding protein, a biomarker, a member of a specific binding pair (e.g., a ligand or aptamer), or an antibody. In some embodiments, a protein may be a full-length protein (e.g., a fully processed protein having certain biological function). In some embodiments, a protein may be a variant or a fragment of a full-length protein. For example, in some embodiments, a Cas9 protein domain comprises an H840A amino acid substitution compared to a naturally occurring. S', pyogenes Cas9 protein. A variant of a protein or enzyme, for example a variant reverse transcriptase, comprises a polypeptide having an amino acid sequence that is about 60% identical, about 70% identical, about 80% identical, about 90% identical, about 95% identical, about 96% identical, about 97% identical, about 98% identical, about 99% identical, about 99.5% identical, or about 99.9% identical to the amino acid sequence of a reference protein.
[0035] In some embodiments, a protein comprises one or more protein domains or subdomains. As used herein, the term “polypeptide domain”, “protein domain”, or “domain” when used in the context of a protein or polypeptide, refers to a polypeptide chain that has one or more biological functions, e.g., a catalytic function, a protein-protein binding function, or a protein-DNA function. In some embodiments, a protein comprises multiple protein domains. In some embodiments, a protein comprises multiple protein domains that are naturally occurring. In some embodiments, a protein comprises multiple protein domains from different naturally occurring proteins. For example, in some embodiments, a prime editor may be a fusion protein comprising a Cas9 protein domain of. S', pyogenes and a reverse transcriptase protein domain of Moloney murine leukemia virus. A protein that comprises amino acid sequences from different origins or naturally occurring proteins may be referred to as a fusion, or chimeric protein.
[0036] In some embodiments, a protein comprises a functional variant or functional fragment of a full-length wild-type protein. A “functional fragment” or “functional portion”, as used herein, refers to any portion of a reference protein (e.g., a wild-type protein) that encompasses less than the entire amino acid sequence of the reference protein while retaining one or more of the functions, e.g., catalytic or binding functions. For example, a functional fragment of a reverse transcriptase may encompass less than the entire amino acid sequence of a wild-type reverse transcriptase, but retains the ability under at least one set of conditions to catalyze the polymerization of a polynucleotide. When the reference protein is a fusion of multiple functional domains, a functional fragment thereof may retain one or more of the functions of at least one of the functional domains. For example, a functional fragment of a Cas9 may encompass less than the entire amino acid sequence of a wild-type Cas9, but retains its DNA binding ability and lacks its nuclease activity partially or completely.
[0037] A “functional variant” or “functional mutant”, as used herein, refers to any variant or mutant of a reference protein (e.g., a wild-type protein) that encompasses one or more alterations to the amino acid sequence of the reference protein while retaining one or more of the functions, e.g., catalytic or binding functions. In some embodiments, the one or more alterations to the amino acid sequence comprises aminoWSGR Docket No. 59761-805.601acid substitutions, insertions or deletions, or any combination thereof. In some embodiments, the one or more alterations to the amino acid sequence comprises amino acid substitutions. For example, a functional variant of a reverse transcriptase may comprise one or more amino acid substitutions compared to the amino acid sequence of a wild-type reverse transcriptase but retains the ability under at least one set of conditions to catalyze the polymerization of a polynucleotide. When the reference protein is a fusion of multiple functional domains, a functional variant thereof may retain one or more of the functions of at least one of the functional domains. For example, in some embodiments, a functional variant of a Cas9 may comprise one or more amino acid substitutions in a nuclease domain, e.g., an H840A amino acid substitution, compared to the amino acid sequence of a wild-type Cas9, but retains the DNA binding ability and lacks the nuclease activity partially or completely.
[0038] The term “function” and its grammatical equivalents as used herein may refer to a capability of operating, having, or serving an intended purpose. Functional may comprise any percent from baseline to 100% of an intended purpose. For example, functional may comprise or comprise about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or up to about 100% of an intended purpose. In some embodiments, the term functional may mean over or over about 100% of normal function, for example, 125%, 150%, 175%, 200%, 250%, 300%, 400%, 500%, 600%, 700% or up to about 1000% of an intended purpose.
[0039] In some embodiments, a protein or polypeptide includes naturally occurring amino acids (e.g., one of the twenty amino acids commonly found in peptides synthesized in nature, and known by the one letter abbreviations A, R, N, C, D, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y and V). In some embodiments, a protein or polypeptides includes non-naturally occurring amino acids (e.g., amino acids that are not one of the twenty amino acids commonly found in peptides synthesized in nature, including synthetic amino acids, amino acid analogs, and amino acid mimetics). In some embodiments, a protein or polypeptide is modified.
[0040] In some embodiments, a protein comprises an isolated polypeptide. The term “isolated” means free or removed to varying degrees from components that normally accompany it as found in the natural state or environment. For example, a polypeptide naturally present in a living animal is not isolated, and the same polypeptide partially or completely separated from the coexisting materials of its natural state is isolated.
[0041] In some embodiments, a protein is present within a cell, a tissue, an organ, or a virus particle. In some embodiments, a protein is present within a cell or a part of a cell (e.g., a bacteria cell, a plant cell, or an animal cell). In some embodiments, the cell is in a tissue, in a subject, or in a cell culture. In some embodiments, the cell is a microorganism (e.g., a bacterium, fungus, protozoan, or virus). In some embodiments, a protein is present in a mixture of analytes (e.g., a lysate). In some embodiments, the protein is present in a lysate from a plurality of cells or from a lysate of a single cell.
[0042] The terms “homologous,” “homology,” or “percent homology” as used herein refer to the degree of sequence identity between an amino acid or polynucleotide sequence and a corresponding reference sequence. “Homology” can refer to polymeric sequences, e.g., polypeptide or DNA sequences that areWSGR Docket No. 59761-805.601similar. Homology can mean, for example, nucleic acid sequences with at least about: 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity. In other embodiments, a “homologous sequence” of nucleic acid sequences may exhibit 93%, 95% or 98% sequence identity to the reference nucleic acid sequence. For example, a "region of homology to a genomic region" can be a region of DNA that has a similar sequence to a given genomic region in the genome. A region of homology can be of any length that is sufficient to promote stable binding of a spacer, primer binding site or protospacer sequence to the complementary sequence of a genomic region. For example, the region of homology can comprise at least 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, 3000, 3100 or more bases in length such that the region of homology has sufficient homology to undergo binding to the complementary sequence of a corresponding genomic region.
[0043] When a percentage of sequence homology or identity is specified, in the context of two nucleic acid sequences or two polypeptide sequences, the percentage of homology or identity generally refers to the alignment of two or more sequences across a portion of their length when compared and aligned for maximum correspondence. When a position in the compared sequence can be occupied by the same base or amino acid, then the molecules can be homologous at that position. Sequence homology or identity can be assessed over a specified length of the nucleic acid, polypeptide or portion thereof. For example, the homology or identity can be assessed over a functional portion or specified portion of the length.
[0044] Alignment of sequences for assessment of sequence homology can be conducted by algorithms known in the art, such as the Basic Local Alignment Search Tool (BLAST) algorithm, which is described in Altschul et al, J. Mol. Biol. 215:403- 410, 1990. A publicly available, internet interface, for performing BLAST analyses is accessible through the National Center for Biotechnology Information. Additional known algorithms include those published in: Smith & Waterman, “Comparison of Biosequences”, Adv. Appl. Math. 2:482, 1981; Needleman & Wunsch, “A general method applicable to the search for similarities in the amino acid sequence of two proteins” J. Mol. Biol. 48:443, 1970; Pearson & Lipman “Improved tools for biological sequence comparison”, Proc. Natl. Acad. Sci. USA 85:2444, 1988; or by automated implementation of these or similar algorithms. Global alignment programs may also be used to align similar sequences of roughly equal size. Examples of global alignment programs include NEEDLE (available at www.ebi.ac.uk / Tools / psa / emboss_needle / ) which is part of the EMBOSS package (Rice P et al., Trends Genet., 2000; 16: 276-277), and the GGSEARCH program https: / / fasta.bioch.virginia.edu / fasta_www2 / , which is part of the FASTA package (Pearson W and Lipman D, 1988, Proc. Natl. Acad. Sci. USA, 85: 2444-2448). Both of these programs are based on the Needleman-Wunsch algorithm which is used to find the optimum alignment (including gaps) of two sequences along their entire length. A detailed discussion of sequence analysis can also be found in Unit 19.3 of Ausubel et al (“Current Protocols in Molecular Biology” John Wiley & Sons Inc, 1994-1998, Chapter 15, 1998). In some embodiments, alignment between a query sequence and a reference sequenceWSGR Docket No. 59761-805.601is performed with Needleman-Wunsch alignment with Gap Costs set to Existence: 11 Extension: 1 where percent identity is calculated by dividing the number of identities by the length of the global alignment, as further described in Altschul et al. (" Gapped BLAST and PSI-BLAST: a new generation of protein database search programs", Nucleic Acids Res. 25:3389-3402, 1997) and Altschul et al, (" Protein database searches using compositionally adjusted substitution matrices", FEBS J. 272:5101-5109, 2005).
[0045] A skilled person understands that amino acid (or nucleotide) positions may be determined in homologous sequences based on alignment, for example, “H840” in a reference Cas9 sequence may correspond to H839, or another position in a Cas9 homolog.
[0046] The term “polynucleotide” or “nucleic acid molecule” can be any polymeric form of nucleotides, including DNA, RNA, a hybridization thereof, or RNA-DNA chimeric molecules. In some embodiments, a polynucleotide comprises cDNA, genomic DNA, mRNA, tRNA, rRNA, or microRNA. In some embodiments, a polynucleotide is double-stranded, e.g., a double -stranded DNA in a gene. In some embodiments, a polynucleotide is single -stranded or substantially single -stranded, e.g., single-stranded DNA or an mRNA. In some embodiments, a polynucleotide is a cell -free nucleic acid molecule. In some embodiments, a polynucleotide circulates in blood. In some embodiments, a polynucleotide is a cellular nucleic acid molecule. In some embodiments, a polynucleotide is a cellular nucleic acid molecule in a cell circulating in blood.
[0047] Polynucleotides can have any three-dimensional structure. The following are nonlimiting examples of polynucleotides: a gene or gene fragment (for example, a probe, primer, EST or SAGE tag), an exon, an intron, intergenic DNA (including, without limitation, heterochromatic DNA), messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), a ribozyme, cDNA, a recombinant polynucleotide, a branched polynucleotide, a plasmid, a vector, isolated DNA, isolated RNA, sgRNA, guide RNA, a nucleic acid probe, a primer, an snRNA, a long non-coding RNA, a snoRNA, a siRNA, a miRNA, a tRNA-derived small RNA (tsRNA), an antisense RNA, an shRNA, or a small rDNA-derived RNA (srRNA).
[0048] In some embodiments, a polynucleotide comprises deoxyribonucleotides, ribonucleotides or analogs thereof. In some embodiments, a polynucleotide comprises modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure can be imparted before or after assembly of the polynucleotide. The sequence of nucleotides can be interrupted by non-nucleotide components. A polynucleotide can be further modified after polymerization, such as by conjugation with a labeling component.
[0049] In some embodiments, a polynucleotide is composed of a specific sequence of four nucleotide bases: adenine (A); cytosine (C); guanine (G); thymine (T); and uracil (U) for thymine when the polynucleotide is RNA. In some embodiments, the polynucleotide may comprise one or more other nucleotide bases, such as inosine (I), which is read by the translation machinery as guanine (G).
[0050] In some embodiments, a polynucleotide may be modified. As used herein, the terms “modified” or “modification” refers to chemical modification with respect to the A, C, G, T and U nucleotides. In some embodiments, modifications may be on the nucleoside base and / or sugar portion of the nucleosidesWSGR Docket No. 59761-805.601that comprise the polynucleotide. In some embodiments, the modification may be on the intemucleoside linkage (e.g., phosphate backbone). In some embodiments, multiple modifications are included in the modified nucleic acid molecule. In some embodiments, a single modification is included in the modified nucleic acid molecule.
[0051] The term "complement", "complementary", or “complementarity” as used herein, refers to the ability of two polynucleotide molecules to base pair with each other. Complementary polynucleotides may base pair via hydrogen bonding, which may be Watson Crick, Hoogsteen or reversed Hoogsteen hydrogen bonding. For example, an adenine on one polynucleotide molecule will base pair to a thymine or uracil on a second polynucleotide molecule and a cytosine on one polynucleotide molecule will base pair to a guanine on a second polynucleotide molecule. Two polynucleotide molecules are complementary to each other when a first polynucleotide molecule comprising a first nucleotide sequence can base pair with a second polynucleotide molecule comprising a second nucleotide sequence. For instance, the two DNA molecules 5’-ATGC-3’ and 5'-GCAT-3’ are complementary, and the complement of the DNA molecule 5’-ATGC-3’ is 5’-GCAT-3’. A percentage of complementarity indicates the percentage of nucleotides in a polynucleotide molecule which can base pair with a second polynucleotide molecule (e.g., 5, 6, 7, 8, 9, 10 out of 10 being 50%, 60%, 70%, 80%, 90%, and 100% complementary, respectively). “Perfectly complementary” means that all the contiguous nucleotides of a polynucleotide molecule will base pair with the same number of contiguous nucleotides in a second polynucleotide molecule. " Substantially complementary" as used herein refers to a degree of complementarity that can be 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% over all or a portion of two polynucleotide molecules. In some embodiments, the portion of complementarity may be a region of 10, 15, 20, 25, 30, 35, 40, 45, 50, or more nucleotides. “Substantial complementary” can also refer to a 100% complementarity over a portion of two polynucleotide molecules. In some embodiments, the portion of complementarity between the two polynucleotide molecules is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% of the length of at least one of the two polynucleotide molecules or a functional or defined portion thereof.
[0052] As used herein, “expression” refers to the process by which polynucleotides are transcribed into mRNA and / or the process by which polynucleotides, e.g., the transcribed mRNA, are translated into peptides, polypeptides, or proteins. If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell. In some embodiments, expression of a polynucleotide, e.g., a gene or a DNA encoding a protein is determined by the amount of the protein encoded by the gene after transcription and translation of the gene. In some embodiments, expression of a polynucleotide, e.g., a gene or a DNA encoding a protein is determined by the amount of a functional form of the protein encoded by the gene after transcription and translation of the gene. In some embodiments, expression of a gene is determined by the amount of the mRNA, or transcript that is encoded by the gene after transcription of the gene. In some embodiments, expression of a polynucleotide, e.g., an mRNA, is determined by the amount of the protein encoded by the mRNA after translation of the mRNA. In some embodiments, expression of a polynucleotide, e.g., an mRNA or coding RNA, is determined by theWSGR Docket No. 59761-805.601amount of a functional form of the protein encoded by the polypeptide after translation of the polynucleotide.
[0053] The term “sequencing” as used herein, may comprise capillary sequencing, bisulfite -free sequencing, bisulfite sequencing, TET-assisted bisulfite (TAB) sequencing, ACE-sequencing, high-throughput sequencing, Maxam-Gilbert sequencing, massively parallel signature sequencing, Polony sequencing, 454 pyrosequencing, Sanger sequencing, Illumina sequencing, SOLiD sequencing, Ion Torrent semiconductor sequencing, DNA nanoball sequencing, Heliscope single molecule sequencing, single molecule real time (SMRT) sequencing, nanopore sequencing, shot gun sequencing, RNA sequencing, or any combination thereof.
[0054] The terms “equivalent” or “biological equivalent” are used interchangeably when referring to a particular molecule, or biological or cellular material, and means a molecule having minimal homology to another molecule while still maintaining a desired structure or functionality.
[0055] The term “encode” as it is applied to polynucleotides refers to a polynucleotide which is said to “encode” another polynucleotide, a polypeptide, or an amino acid if, in its native state or when manipulated by methods well known to those skilled in the art, it can be used as a polynucleotide synthesis template, e.g., transcribed into an RNA, reverse transcribed into a DNA or cDNA, and / or translated to produce an amino acid, or a polypeptide or fragment thereof. In some embodiments, a polynucleotide comprising three contiguous nucleotides form a codon that encodes a specific amino acid. In some embodiments, a polynucleotide comprises one or more codons that encode a polypeptide. In some embodiments, a polynucleotide comprising one or more codons comprises a mutation in a codon compared to a wild-type reference polynucleotide. In some embodiments, the mutation in the codon encodes an amino acid substitution in a polypeptide encoded by the polynucleotide as compared to a wild-type reference polypeptide. As it applies to a PEgRNA, and the RTT of a PEgRNA in particular, the term “encode” refers to the state of a target nucleic acid following prime editing of the target nucleic acid with the PEgRNA. Therefore, a PEgRNA comprising an RTT that encodes wild-type protein sequence means that the portion of a target gene edited by the PEgRNA will encode wild-type protein sequence following complete incorporation of the edits.
[0056] The term “mutation” as used herein refers to a change and / or alteration in an amino acid sequence of a protein or nucleic acid sequence of a polynucleotide. Such changes and / or alterations may comprise the substitution, insertion, deletion and / or truncation of one or more amino acids, in the case of an amino acid sequence, and / or nucleotides, in the case of nucleic acid sequence, compared to a reference amino acid or nucleic acid sequence. In some embodiments, the reference sequence is a wild-type sequence. In some embodiments, the reference sequence is an endogenous sequence in a target DNA sequence, e.g., in a target SERPINA1 gene in a cell or in a subject. In some embodiments, a mutation in a nucleic acid sequence of a polynucleotide encodes a mutation in the amino acid sequence of a polypeptide. In some embodiments, the mutation in the amino acid sequence of the polypeptide or the mutation in the nucleic acid sequence of the polynucleotide is a mutation associated with a disease state.WSGR Docket No. 59761-805.601
[0057] For purposes of this disclosure, the chemical elements are identified in accordance with the Periodic Table of the Elements, CAS version, Handbook of Chemistry and Physics, 75th Ed.Additionally, general principles of organic chemistry are described in " Organic Chemistry," Thomas Sorrell, University Science Books, Sausalito: 1999, and " March's Advanced Organic Chemistry," 5th Ed., Ed.: Smith, M. B. and March, J., John Wiley and Sons, New York: 2001, the entire contents of which are hereby incorporated by reference.
[0058] As described herein, compounds of the disclosure optionally may be substituted with one or more substituents, such as are illustrated generally above, or as exemplified by particular classes, subclasses, and species of the disclosure.
[0059] Unless otherwise stated, structures depicted herein also are meant to include all isomeric (e.g., enantiomeric, diastereomeric, and geometric (or conformational)) forms of the structure; for example, the R and S configurations for each asymmetric center, (Z) and (E) double bond isomers, and (Z) and (E) conformational isomers. Therefore, single stereochemical isomers as well as enantiomeric, diastereomeric, and geometric (or conformational) mixtures of the present compounds are within the scope of the disclosure. Unless otherwise stated, all tautomeric forms of the compounds of the disclosure are within the scope of the disclosure. Additionally, unless otherwise stated, structures depicted herein also are meant to include compounds that differ only in the presence of one or more isotopically enriched atoms. For example, compounds having the present structures except for the replacement of hydrogen by deuterium or tritium, or the replacement of a carbon by a 13C- or 14C-enriched carbon are within the scope of this disclosure. Such compounds are useful, for example, as analytical tools or probes in biological assays, or as therapeutic agents.
[0060] As used herein, the phrasing “-based” in the context of a chemical compound can refer to a chemical conjugation or the process of chemically linking two or more molecules to form a stable compound, often enhancing the properties or functionalities of the resulting molecule. For example, a “GalNAc -based lipid” (also referred to as “GalNAc lipid”) may comprise the conjugation of a GalNAc moiety with a lipid molecule, resulting in a hybrid compound that exhibits improved solubility, bioavailability, or targeting capabilities in a biological system.
[0061] Compounds disclosed herein may be "isomerically pure" compounds. As used herein, the term "isomerically pure" refers to an isomeric form of a compound that is substantially free from other isomeric forms of the compound (e.g., substantially free from other stereoisomers (e.g., enantiomers, diastereomers, geometric (or conformational) isomers, etc.), constitutional isomers, isotopomers, etc.). For example, an “isomerically pure” compound having at least one asymmetric center of a particular configuration (i.e., R or S configuration) is substantially free from other isomeric forms of the compound having a different configuration at the at least one asymmetric center. An “isomerically pure” compound comprises more than 75% by weight, more than 80% by weight, more than 85% by weight, more than 90% by weight, more than 91% by weight, more than 92% by weight, more than 93% by weight, more than 94% by weight, more than 95% by weight, more than 96% by weight, more than 97% by weight, more than 98% by weight, more than 98.5% by weight, more than 99% by weight, more than 99.2% byWSGR Docket No. 59761-805.601weight, more than 99.5% by weight, more than 99.6% by weight, more than 99.7% by weight, more than 99.8% by weight, or more than 99.9% by weight, of a single isomer of the compound based on the total weight of all isomers of the compound that are present.
[0062] Chemical structures and nomenclature are derived from ChemDraw, version 21.0.0, Cambridge, MA.
[0063] It also will be appreciated that certain of the compounds of the present disclosure can exist in free form for treatment, or where appropriate, as a pharmaceutically acceptable derivative (e.g., a salt) thereof. According to the present disclosure, a pharmaceutically acceptable derivative includes, but is not limited to, pharmaceutically acceptable prodrugs, salts, esters, salts of such esters, or any other adduct or derivative that upon administration to a patient in need is capable of providing, directly or indirectly, a compound as otherwise described herein, or a metabolite or residue thereof.
[0064] As used herein, the term "pharmaceutically acceptable" means approved or approvable by a regulatory agency of the Federal or a state government or the corresponding agency in countries other than the United States, or that is listed in the U. S. Pharmacopoeia or other generally recognized pharmacopoeia for use in animals, and more particularly, in humans.
[0065] As used herein, the term "pharmaceutically acceptable salt" refers to a salt of a compound of the disclosure that is pharmaceutically acceptable and that possesses the desired pharmacological activity of the parent compound. In particular, such salts are non-toxic may be inorganic or organic acid addition salts and base addition salts. Specifically, such salts include: (1) acid addition salts, formed with inorganic acids such as hydrochloric acid, hydrobromic acid, sulfuric acid, nitric acid, phosphoric acid, and the like; or formed with organic acids such as acetic acid, propionic acid, hexanoic acid, cyclopentanepropionic acid, glycolic acid, pyruvic acid, lactic acid, malonic acid, succinic acid, malic acid, maleic acid, fumaric acid, tartaric acid, citric acid, benzoic acid, 3-(4-hydroxybenzoyl) benzoic acid, cinnamic acid, mandelic acid, methanesulfonic acid, ethane sulfonic acid, 1,2-ethane-disulfonic acid, 2-hydroxyethanesulfonic acid, benzenesulfonic acid, 4 -chlorobenzene sulfonic acid, 2-naphthalenesulfonic acid, 4 -toluene sulfonic acid, camphorsulfonic acid, 4-methylbicyclo[2.2.2]-oct-2-ene-l-carboxylic acid, glucoheptonic acid, 3-phenylpropionic acid, trimethylacetic acid, tertiary butylacetic acid, lauryl sulfuric acid, gluconic acid, glutamic acid, hydroxynaphthoic acid, salicylic acid, stearic acid, muconic acid, and the like; or (2) salts formed when an acidic proton present in the parent compound either is replaced by a metal ion, e.g., an alkali metal ion, an alkaline earth ion, or an aluminum ion; or coordinates with an organic base such as ethanolamine, diethanolamine, triethanolamine, N-methylglucamine and the like. Salts further include, by way of example only, sodium, potassium, calcium, magnesium, ammonium, tetraalkylammonium, and the like; and when the compound contains a basic functionality, salts of non-toxic organic or inorganic acids, such as hydrochloride, hydrobromide, tartrate, mesylate, acetate, maleate, oxalate and the like. The term "pharmaceutically acceptable cation" refers to an acceptable cationic counter-ion of an acidic functional group. Such cations are exemplified by sodium, potassium, calcium, magnesium, ammonium, tetraalkylammonium cations, and the like. See, e.g., Berge, et al., J. Pharm. Sci. (1977) 66(1): 1-79.WSGR Docket No. 59761-805.601
[0066] The term “lipid” refers to a group of organic compounds that include, but are not limited to, esters of fatty acids and are generally characterized by being poorly soluble in water, but soluble in many organic solvents.
[0067] A “lipid nanoparticle” may refer to particles having at least one dimension in the nanometers range (e.g., 1-1,000 nm) and include the compounds described herein. In some embodiments, lipid nanoparticles (LNPs) can be included in compositions that are used to deliver at least one polynucleotide and / or construct according to this disclosure. In some embodiments, the lipid nanoparticles of the disclosure comprise at least one polynucleotide and / or one or more components of a prime editing system. In some embodiments, the lipid nanoparticles can include one or more additional lipid components. In some embodiments, the at least one polynucleotide and / or construct may be encapsulated in the lipid portion of the lipid nanoparticle or an aqueous space enveloped by some or all of the lipid portion of the lipid nanoparticle.
[0068] In some embodiments, the lipid nanoparticles can have a mean diameter of from about 30 nm to about 200 nm, from about 40 nm to about 200 nm, from about 50 nm to about 200 nm, from about 30 nm to about 150 nm, from about 40 nm to about 150 nm, from about 50 nm to about 150 nm, from about 60 nm to about 130 nm, from about 70 nm to about 110 nm, from about 70 nm to about 100 nm, from about 80 nm to about 100 nm, from about 90 nm to about 100 nm, from about 70 to about 90 nm, from about 80 nm to about 90 nm, from about 70 nm to about 80 nm, or about 30 nm, 35 nm, 40 nm, 45 nm, 50 nm, 55 nm, 60 nm, 65 nm, 70 nm, 75 nm, 80 nm, 85 nm, 90 nm, 95 nm, 100 nm, 105 nm, 110 nm, 115 nm, 120 nm, 125 nm, 130 nm, 135 nm, 140 nm, 145 nm, 150 nm, 155 nm, 160 nm, 165 nm, 170 nm, 175 nm, 180 nm, 185 nm, 190 nm, 195 nm, or 200 nm.
[0069] As used herein, the term “biodegradable” refers to materials that, when introduced into cells, are broken down by cellular machinery (e.g., enzymatic degradation) or by hydrolysis into components that cells can either reuse or dispose of without significant toxic effect(s) on the cells. In certain embodiments, components generated by breakdown of a biodegradable material do not induce inflammation and / or other adverse effects in vivo. In some embodiments, biodegradable materials are enzymatically broken down. Alternatively or additionally, in some embodiments, biodegradable materials are readily broken down and eliminated.
[0070] As used herein, “ionizable lipids,” also known as “cationic lipids,” refer to lipids that are positively charged at the head group at acidic pH and neutral at physiological pH. Ionizable lipids typically contain three domains: a hydrophilic headgroup, a hydrophobic domain, and a linker. Ionizable lipids are an amphiphilic entity with the head group having hydrophilicity and lipid tails possessing hydrophobicity.
[0071] As used herein, the term “molar ratio” (or “mol %) of lipid components (e.g., ionizable lipid, helper lipid, PEG-containing lipid, sterol, and GalNAc -based lipid) refers to the relative proportion of each lipid component in the lipid nanoparticle (LNP) formulation based on the total lipid present in the LNP, as measured after LNP formation and purification (e.g., via tangential flow filtration). The molar ratio is determined by a reversed-phase ultra-performance liquid chromatography (RP-UPLC) methodWSGR Docket No. 59761-805.601coupled with charged aerosol detection (CAD) as described in Example 40. Unless otherwise indicated, a recited molar ratio encompasses normal manufacturing variations and analytical measurement errors. A person of ordinary skill in the art would understand that the recited molar ratio represents the average molar ratio of each lipid component in the plurality of LNPs within the formulation.
[0072] As used herein, “N / P” or “N / P ratio” refers to the molar ratio of cationic (nitrogen) groups (the " N" in N / P) in the ionizable lipid or cationic lipid or polymer to the anionic (phosphate) groups (the " P" in N / P) in RNA. It is understood that a cationic group is one that is either in permanently cationic form (e.g., N+), or one that is ionizable to become cationic (e.g., under certain pH conditions). Use of a single number in an N / P ratio (e.g., an N / P ratio of about 5) is intended to refer to that number over 1, e.g., an N / P ratio of about 4 is intended to mean about 4:1.
[0073] The term “subject” and its grammatical equivalents as used herein may refer to a human or a nonhuman. A subject may be a mammal. A human subject may be male or female. A human subject may be of any age. A subject may be a human embryo. A human subject may be a newborn, an infant, a child, an adolescent, or an adult. A human subject may be up to about 100 years of age. A human subject may be in need of treatment for a genetic disease or disorder.
[0074] The terms “treatment” or “treating” and their grammatical equivalents may refer to the medical management of a subject with an intent to cure, ameliorate, or ameliorate a symptom of, a disease, condition, or disorder. Treatment may include active treatment, that is, treatment directed specifically toward the improvement of a disease, condition, or disorder. Treatment may include causal treatment, that is, treatment directed toward removal of the cause of the associated disease, condition, or disorder. In addition, this treatment may include palliative treatment, that is, treatment designed for the relief of symptoms rather than the curing of the disease, condition, or disorder. Treatment may include supportive treatment, that is, treatment employed to supplement another specific therapy directed toward the improvement of the disease, condition, or disorder. In some embodiments, a condition may be pathological. In some embodiments, a treatment may not completely cure or prevent a disease, condition, or disorder. In some embodiments, a treatment ameliorates, but does not completely cure or prevent a disease, condition, or disorder. In some embodiments, a subject may be treated for 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 2 weeks, 3 weeks, 4 weeks, 2 months, 3 months, 4 months, 5 months, 6 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, indefinitely, or life of the subject.
[0075] The term “ameliorate” and its grammatical equivalents means to decrease, suppress, attenuate, diminish, arrest, or stabilize the development or progression of a disease.
[0076] The terms “prevent” or “preventing” means delaying, forestalling, or avoiding the onset or development of a disease, condition, or disorder for a period of time. Prevent also means reducing risk of developing a disease, disorder, or condition. Prevention includes minimizing or partially or completely inhibiting the development of a disease, condition, or disorder. In some embodiments, a composition, e.g., a pharmaceutical composition, prevents a disorder by delaying the onset of the disorder for 12 hours, 24 hours, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 2 weeks, 3 weeks, 4 weeks, 2 months, 3 months, 4WSGR Docket No. 59761-805.601months, 5 months, 6 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, indefinitely, or life of a subject.
[0077] The term “effective” means having the ability to produce a biological response. For example, “effective amount” or “therapeutically effective amount” may refer to a quantity of a composition, for example a composition comprising a construct, that can be sufficient to result in a desired activity upon introduction into a subject as disclosed herein. An effective amount of the prime editing compositions can be provided to the target gene or cell, whether the cell is in vitro, ex vivo or in vivo.
[0078] An effective amount can be the amount to induce, for example, at least about a 2-fold change (increase or decrease) or more in the amount of target nucleic acid modulation (e.g., expression of SERPINA1 gene to produce functional AAT protein) observed relative to a negative control. An effective amount or dose can induce, for example, about 2-fold increase, about 3-fold increase, about 4-fold increase, about 5-fold increase, about 6-fold increase, about 7-fold increase, about 8-fold increase, about 9-fold increase, about 10-fold increase, about 25-fold increase, about 50-fold increase, about 100-fold increase, about 200-fold increase, about 500-fold increase, about 700-fold increase, about 1000-fold increase, about 5000-fold increase, or about 10,000-fold increase in target gene modulation (e.g., expression of a target SERPINA1 gene to produce functional AAT protein).
[0079] The amount of target gene modulation may be measured by any suitable method known in the art. In some embodiments, the “effective amount” or “therapeutically effective amount” is the amount of a composition that is required to ameliorate the symptoms of a disease relative to an untreated patient. In some embodiments, an effective amount is the amount of a composition sufficient to introduce an alteration in a gene of interest in a cell (e.g., a cell in vitro, ex vivo or in vivo).
[0080] An effective amount can be the amount to induce, when administered to a population of cells, a certain percentage of the population of cells to have a correction of a mutation associated with AATD. For example, in some embodiments, an effective amount can be the amount to induce, when administered to or introduced to a population of cells, installation of one or more intended nucleotide edits that correct a c.1096 G-> A (encoding E366K amino acid substitution, also referred to as E342K amino acid substitution in mature AAT protein) mutation in the SERPINA1 gene, in at least about 1%, 2%, 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 99% of the population of cells.
[0081] In some embodiments, an effective amount can be an amount to induce, when administered to a population of cells, a certain percentage of the population of cells to have a correction of one or more of mutations selected from the group consisting of c.1111del, c.1103G> A, c.1096G> A, and c.1093G> C. In some embodiments, an effective amount can be an amount to induce, when administered to a population of cells, a certain percentage of the population of cells to have a correction a c.1096G> A mutation. In some embodiments, an effective amount can be an amount to induce, when administered to a population of cells, a certain percentage of the population of cells to have a correction a c.1093G> C mutation. For example, in some embodiments, an effective amount can be the amount to induce, when administered toWSGR Docket No. 59761-805.601or introduced to a population of cells, installation of one or more intended nucleotide edits that correct a c.1111del, c.1103G> A, c.1096G> A, and / or a c.1093G> C mutation in the SERPINA1 gene, in at least about 1%, 2%, 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 99% of the population of cells.
[0082] The term “exogenous” when used in reference to a biomolecule, e.g., a polynucleotide sequence or a polypeptide sequence refers to a biomolecule that is not native to a specific biological context, e.g., a gene, a particular chromosome, a particular cell or chromosomal site of the cell, tissue, or organism, or, if from the same source, is modified from its original form or is present in a non-native location, e.g., a chromosome location.
[0083] The term “endogenous” when used in reference to a biomolecule, e.g., a polynucleotide sequence or a polypeptide sequence refers to a biomolecule that is native to or naturally occurring in a specific biological context, e.g., a gene, a particular chromosome, a particular cell or chromosomal site of the cell, tissue, or organism. For example, an endogenous sequence may be a wild-type sequence or may comprise one or more mutations compared to a wild-type sequence. In some embodiments, an endogenous sequence is mutated compared to a wild-type sequence and may cause or be associated with a disease or disorder, e.g., AATD, in a subject. As used herein, in some embodiments, a wild-type sequence, with respect to a specific gene and a specific disease, is a gene sequence found in healthy individuals, wherein the wild-type sequence does not include a mutation causative of the specific disease.
[0084] By “upstream” and “downstream” it is intended to define relative positions of at least two regions or sequences in a nucleic acid molecule oriented in a 5'-to-3' direction. For example, a first sequence is upstream of a second sequence in a DNA molecule where the first sequence is positioned 5 ’ to the second sequence. Accordingly, the second sequence is downstream, that is, 3’, of the first sequence.Prime Editing
[0085] The term “prime editing” refers to programmable editing of a target DNA using a prime editor complexed with a PEgRNA to incorporate an intended nucleotide edit (also referred to herein as a nucleotide change) into the target DNA through target-primed DNA synthesis. A target gene of prime editing may comprise a double stranded DNA molecule having two complementary strands: a first strand that may be referred to as a “target strand” or a “non-edit strand,” and a second strand that may be referred to as a “non-target strand,” or an “edit strand.” In some embodiments, in a prime editing guide RNA (PEgRNA), a spacer sequence is complementary or substantially complementary to a specific sequence on the target strand, which may be referred to as a “search target sequence.” In some embodiments, the spacer sequence anneals with the target strand at the search target sequence. The target strand may also be referred to as the “non-Protospacer Adjacent Motif (non-PAM strand).” In some embodiments, the non-target strand may also be referred to as the “PAM strand.” In some embodiments, the PAM strand comprises a protospacer sequence and optionally a protospacer adjacent motif (PAM) sequence. In prime editing using a Cas-protein-based prime editor, a PAM sequence refers to a shortWSGR Docket No. 59761-805.601DNA sequence immediately adjacent to the protospacer sequence on the PAM strand of the target gene. A PAM sequence may be specifically recognized by a programmable DNA binding protein, e.g., a Cas nickase or a Cas nuclease. In some embodiments, a specific PAM is characteristic of a specific programmable DNA binding protein, e.g., a Cas nickase or a Cas nuclease A protospacer sequence refers to a specific sequence in the PAM strand of the target gene that is complementary to the search target sequence. In a PEgRNA, a spacer sequence may have a substantially identical sequence as the protospacer sequence on the edit strand of a target gene, except that the spacer sequence may comprise Uracil (U) and the protospacer sequence may comprise Thymine (T).
[0086] In some embodiments, the double stranded target DNA comprises a nick site on the PAM strand (or non-target strand). As used herein, a “nick site” refers to a specific position in between two nucleotides or two base pairs of the double stranded target DNA. In some embodiments, the position of a nick site is determined relative to the position of a specific PAM sequence. In some embodiments, the nick site is the particular position where a nick will occur when the double stranded target DNA is contacted with a nickase, for example, a Cas nickase, that recognizes a specific PAM sequence. In some embodiments, the nick site is upstream of a specific PAM sequence on the PAM strand of the double stranded target DNA. In some embodiments, the nick site is downstream of a specific PAM sequence on the PAM strand of the double stranded target DNA. In some embodiments, the nick site is upstream of a PAM sequence recognized by a Cas9 nickase, wherein the Cas9 nickase comprises a nuclease active RuvC domain and a nuclease inactive HNH domain. In some embodiments, the nick site is 3 nucleotides upstream of the PAM sequence, and the PAM sequence is recognized by a Streptococcus pyogenes Cas9 nickase, a P. lavamentivorans Cas9 nickase, a C. diphtheriae Cas9 nickase, aN. cinerea Cas9, a S. aureus Cas9, or a N. lari Cas9 nickase. In some embodiments, the nick site is 3 base pairs upstream of the PAM sequence, and the PAM sequence is recognized by a Cas9 nickase, wherein the Cas9 nickase that comprises a nuclease active RuvC domain and a nuclease inactive HNH domain. In some embodiments, the nick site is 2 nucleotides upstream of the PAM sequence, and the PAM sequence is recognized by a S. thermophilus Cas9 nickase that comprises a nuclease active RuvC domain and a nuclease inactive HNH domain.
[0087] A “primer binding site” (also referred to as PBS or primer binding site sequence) is a singlestranded portion of the PEgRNA that comprises a region of complementarity to the PAM strand (i.e. the non-target strand or the edit strand). The PBS is complementary or substantially complementary to a sequence on the PAM strand of the double stranded target DNA that is immediately upstream of the nick site. In some embodiments, in the process of prime editing, the PEgRNA complexes with and directs a prime editor to bind the search target sequence on the target strand of the double stranded target DNA, and generates a nick at the nick site on the non-target strand of the double stranded target DNA. In some embodiments, the PBS is complementary to or substantially complementary to, and can anneal to, a free 3' end on the non-target strand of the double stranded target DNA at the nick site. In some embodiments, the PBS annealed to the free 3' end on the non-target strand can initiate target-primed DNA synthesis.WSGR Docket No. 59761-805.601
[0088] An “editing template” of a PEgRNA is a single -stranded portion of the PEgRNA that is 5' of the PBS and which encodes a single strand of DNA. The editing template may comprise a region of complementarity to the PAM strand (i.e., the non-target strand or the edit strand), and comprises one or more intended nucleotide edits compared to the endogenous sequence of the double stranded target DNA. In some embodiments, the editing template and the PBS are immediately adjacent to each other.Accordingly, in some embodiments, a PEgRNA in prime editing comprises a single -stranded portion that comprises the PBS and the editing template immediately adjacent to each other. In some embodiments, the single stranded portion of the PEgRNA comprising both the PBS and the editing template is complementary or substantially complementary to an endogenous sequence on the PAM strand (i.e., the non-target strand or the edit strand) of the double stranded target DNA except for one or more non-complementary nucleotides at the intended nucleotide edit position(s). As used herein, regardless of relative 5 '-3' positioning in other context, the relative positions as between the PBS and the editing template, and the relative positions as among elements of a PEgRNA, are determined by the 5' to 3' order of the PEgRNA as a single molecule regardless of the position of sequences in the double stranded target DNA that may have complementarity or identity to elements of the PEgRNA. In some embodiments, the editing template is complementary or substantially complementary to a sequence on the PAM strand that is immediately downstream of the nick site, except for one or more non-complementary nucleotides at the intended nucleotide edit positions. The endogenous, e.g., genomic, sequence that is complementary or substantially complementary to the editing template, except for the one or more non-complementary nucleotides at the position corresponding to the intended nucleotide edit, may be referred to as an “editing target sequence”. In some embodiments, the editing template has identity or substantial identity to a sequence on the target strand that is complementary to, or having the same position in the genome as, the editing target sequence, except for one or more insertions, deletions, or substitutions at the intended nucleotide edit positions. In some embodiments, the editing template encodes a single stranded DNA, wherein the single stranded DNA has identity or substantial identity to the editing target sequence except for one or more insertions, deletions, or substitutions at the positions of the one or more intended nucleotide edits. In some embodiments, the editing template may encode the wild-type or non-disease associated gene sequence (or its complement if the edit strand is the antisense strand of a gene). In some embodiments, the editing template may encode the wild-type or non-disease associated protein but contain one or more synonymous mutations relative to the wild-type or non-disease associated protein coding region. Such synonymous mutations may include, for example, mutations that decrease the ability of a PEgRNA to rebind to the same target sequence once the desired edit is installed in the genome (e.g., synonymous mutations that silence the endogenous PAM sequence or that edit the endogenous protospacer). “Synonymous mutation,” “synonymous edit,” “silent edit,” or “silent mutation” can refer to a nucleotide change or substitution in a protein coding sequence that results in no alteration in the encoded amino acid sequence. A synonymous mutation results in encoding the same amino acid as the one encoded by the coding sequence lacking the synonymous mutation.WSGR Docket No. 59761-805.601
[0089] In some embodiments, a PEgRNA complexes with and directs a prime editor to bind to the search target sequence of the target gene. In some embodiments, the bound prime editor generates a nick on the edit strand (PAM strand) of the target gene at the nick site. In some embodiments, a primer binding site (PBS) of the PEgRNA anneals with a free 3' end formed at the nick site, and the prime editor initiates DNA synthesis from the nick site, using the free 3' end as a primer. Subsequently, a single-stranded DNA encoded by the editing template of the PEgRNA is synthesized. In some embodiments, the newly synthesized single -stranded DNA comprises one or more intended nucleotide edits compared to the endogenous target gene sequence. Accordingly, in some embodiments, the editing template of a PEgRNA is complementary to a sequence in the edit strand except for one or more mismatches at the intended nucleotide edit positions in the editing template. The endogenous, e.g., genomic, sequence that is partially complementary to the editing template may be referred to as an “editing target sequence”. Accordingly, in some embodiments, the newly synthesized single stranded DNA has identity or substantial identity to a sequence in the editing target sequence, except for one or more insertions, deletions, or substitutions at the intended nucleotide edit positions. In some embodiments, the editing template comprises at least 4 contiguous nucleotides of complementarity with the edit strand wherein the at least 4 nucleotides contiguous are located upstream of the 5’ most edit in the editing template.
[0090] In some embodiments, the newly synthesized single -stranded DNA equilibrates with the editing target on the edit strand of the target gene for pairing with the target strand of the target gene. In some embodiments, the editing target sequence of the target gene is excised by a flap endonuclease (FEN), for example, FEN1. In some embodiments, the FEN is an endogenous FEN, for example, in a cell comprising the target gene. In some embodiments, the FEN is provided as part of the prime editor, either linked to other components of the prime editor or provided in trans. In some embodiments, the newly synthesized single stranded DNA, which comprises the intended nucleotide edit, replaces the endogenous single stranded editing target sequence on the edit strand of the target gene. In some embodiments, the newly synthesized single stranded DNA and the endogenous DNA on the target strand form a heteroduplex DNA structure at the region corresponding to the editing target sequence of the target gene. In some embodiments, the newly synthesized single-stranded DNA comprising the nucleotide edit is paired in the heteroduplex with the target strand of the target DNA that does not comprise the nucleotide edit, thereby creating a mismatch between the two otherwise complementary strands. In some embodiments, the mismatch is recognized by DNA repair machinery, e.g., an endogenous DNA repair machinery. In some embodiments, through DNA repair, the intended nucleotide edit is incorporated into the target gene. Prime Editor
[0091] The term “prime editor (PE)” refers to the polypeptide or polypeptide components involved in prime editing, or any polynucleotide(s) encoding the polypeptide or polypeptide components. In various embodiments, a prime editor includes a polypeptide domain having DNA binding activity and a polypeptide domain having DNA polymerase activity. In some embodiments, the polypeptide domain having DNA binding activity is a polypeptide domain having programmable DNA binding activity. InWSGR Docket No. 59761-805.601some embodiments, the prime editor further comprises a polypeptide domain having nuclease activity. In some embodiments, the polypeptide domain having DNA binding activity comprises a nuclease domain or nuclease activity. In some embodiments, the polypeptide domain having nuclease activity comprises a nickase, or a fully active nuclease. As used herein, the term “nickase” refers to a nuclease capable of cleaving only one strand of a double -stranded DNA target. In some embodiments, the prime editor comprises a polypeptide domain that is an inactive nuclease. In some embodiments, the polypeptide domain having programmable DNA binding activity comprises a nucleic acid guided DNA binding domain, for example, a CRISPR-Cas protein, for example, a Cas9 nickase, a Cpfl nickase, or another CRISPR-Cas nuclease. In some embodiments, the polypeptide domain having DNA polymerase activity comprises a template-dependent DNA polymerase, for example, a DNA-dependent DNA polymerase or an RNA-dependent DNA polymerase. In some embodiments, the DNA polymerase is a reverse transcriptase. In some embodiments, the prime editor comprises additional polypeptides or polypeptide domains involved in prime editing, for example, a polypeptide domain having 5 ’ endonuclease activity, e.g., a 5' endogenous DNA flap endonucleases (e.g., FEN1), for helping to drive the prime editing process towards the edited product formation. In some embodiments, the prime editor further comprises an RNA-protein recruitment polypeptide, for example, a MS2 coat protein.
[0092] A prime editor may be engineered. In some embodiments, the polypeptide components of a prime editor do not naturally occur in the same organism or cellular environment. In some embodiments, the polypeptide components of a prime editor may be of different origins or from different organisms. In some embodiments, a prime editor comprises a DNA binding domain and a DNA polymerase domain that are derived from different species. In some embodiments, a prime editor comprises a Cas polypeptide and a reverse transcriptase polypeptide that are derived from different species. For example, a prime editor may comprise a. S', pyogenes Cas9 polypeptide and a Moloney murine leukemia virus (M-MLV) reverse transcriptase polypeptide.
[0093] In some embodiments, polypeptide domains of a prime editor may be fused or linked by a peptide linker to form a fusion protein. In other embodiments, a prime editor comprises one or more polypeptide domains provided in trans as separate proteins, which are capable of being associated to each other through non-peptide linkages or through aptamers or recruitment sequences. For example, a prime editor may comprise a DNA binding domain and a reverse transcriptase domain associated with each other by an RNA-protein recruitment aptamer, e.g., an MS2 aptamer, which may be linked to a PEgRNA. Prime editor polypeptide components may be encoded by one or more polynucleotides in whole or in part. In some embodiments, a single polynucleotide, construct, or vector encodes the prime editor fusion protein. In some embodiments, multiple polynucleotides, constructs, or vectors each encode a polypeptide domain or portion of a domain of a prime editor, or a portion of a prime editor fusion protein. For example, a prime editor fusion protein may comprise an N-terminal portion fused to an intein-N and a C-terminal portion fused to an intein-C, each of which is individually encoded by an AAV vector.
[0094] The term “prime editor complex” is used interchangeably with the term “prime editing complex” and refers to a complex comprising one or more prime editor components (e.g., a polypeptide domainWSGR Docket No. 59761-805.601having DNA binding activity and a polypeptide domain having DNA polymerase activity) complexed with a PEgRNA.Prime Editor Nucleotide Polymerase Domain
[0095] In some embodiments, a prime editor comprises a nucleotide polymerase domain, e.g., a DNA polymerase domain. The DNA polymerase domain may be a wild-type DNA polymerase domain, a full-length DNA polymerase protein domain, or may be a functional mutant, a functional variant, or a functional fragment thereof. In some embodiments, the polymerase domain is a template dependent polymerase domain. For example, the DNA polymerase may rely on a template polynucleotide strand, e.g., the editing template sequence, for new strand DNA synthesis. In some embodiments, the prime editor comprises a DNA-dependent DNA polymerase. For example, a prime editor having a DNA-dependent DNA polymerase can synthesize a new single-stranded DNA using a PEgRNA editing template that comprises a DNA sequence as a template. In such cases, the PEgRNA is a chimeric or hybrid PEgRNA, and comprises an extension arm comprising a DNA strand. As used herein, an “extension arm” is a polynucleotide portion of a PEgRNA that comprises an editing template and a primer binding site sequence (PBS). In some embodiments, an extension arm further comprises additional components, for example, a 3’ modifier. The chimeric or hybrid PEgRNA may comprise an RNA portion (including the spacer and the gRNA core) and a DNA portion (the extension arm comprising the editing template that includes a strand of DNA).
[0096] The DNA polymerases can be wild-type polymerases from eukaryotic, prokaryotic, archael, or viral organisms, and / or the polymerases may be modified by genetic engineering, mutagenesis, or directed evolution-based processes. The polymerases can be a T7 DNA polymerase, T5 DNA polymerase, T4 DNA polymerase, Klenow fragment DNA polymerase, DNA polymerase III and the like. The polymerases can be thermostable, and can include Taq, Tne, Tma, Pfu, Tfl, Tth, Stoffel fragment, VENT® and DEEPVENT® DNA polymerases, KOD, Tgo, JDF3, and mutants, variants and derivatives thereof.
[0097] In some embodiments, the DNA polymerase is a bacteriophage polymerase, for example, a T4, T7, or phi29 DNA polymerase. In some embodiments, the DNA polymerase is an archaeal polymerase, for example, pol I type archaeal polymerase or a pol II type archaeal polymerase. In some embodiments, the DNA polymerase comprises a thermostable archaeal DNA polymerase. In some embodiments, the DNA polymerase comprises a eubacterial DNA polymerase, for example, Pol I, Pol II, or Pol III polymerase. In some embodiments, the DNA polymerase is a Pol I family DNA polymerase. In some embodiments, the DNA polymerase is an E.coli Pol I DNA polymerase. In some embodiments, the DNA polymerase is a Pol II family DNA polymerase. In some embodiments, the DNA polymerase is a Pyrococcus furiosus (Pfu) Pol II DNA polymerase. In some embodiments, the DNA polymerase is a Pol IV family DNA polymerase. In some embodiments, the DNA polymerase is an E.coli Pol IV DNA polymerase.
[0098] In some embodiments, the DNA polymerase comprises a eukaryotic DNA polymerase. In some embodiments, the DNA polymerase is a Pol-beta DNA polymerase, a Pol-lambda DNA polymerase, aWSGR Docket No. 59761-805.601Pol-sigma DNA polymerase, or a Pol-mu DNA polymerase. In some embodiments, the DNA polymerase is a Pol-alpha DNA polymerase. In some embodiments, the DNA polymerase is a POLA1 DNA polymerase. In some embodiments, the DNA polymerase is a POLA2 DNA polymerase. In some embodiments, the DNA polymerase is a Pol-delta DNA polymerase. In some embodiments, the DNA polymerase is a POLDI DNA polymerase. In some embodiments, the DNA polymerase is a POLD2 DNA polymerase. In some embodiments, the DNA polymerase is a human POLDI DNA polymerase. In some embodiments, the DNA polymerase is a human POLD2 DNA polymerase. In some embodiments, the DNA polymerase is a POLD3 DNA polymerase. In some embodiments, the DNA polymerase is a POLD4 DNA polymerase. In some embodiments, the DNA polymerase is a Pol-epsilon DNA polymerase. In some embodiments, the DNA polymerase is a POLE1 DNA polymerase. In some embodiments, the DNA polymerase is a POLE2 DNA polymerase. In some embodiments, the DNA polymerase is a POLE3 DNA polymerase. In some embodiments, the DNA polymerase is a Pol-eta (POLH) DNA polymerase. In some embodiments, the DNA polymerase is a Pol-iota (POLI) DNA polymerase. In some embodiments, the DNA polymerase is a Pol-kappa (POLK) DNA polymerase. In some embodiments, the DNA polymerase is a Revl DNA polymerase. In some embodiments, the DNA polymerase is a human Revl DNA polymerase. In some embodiments, the DNA polymerase is a viral DNA-dependent DNA polymerase. In some embodiments, the DNA polymerase is a B family DNA polymerases. In some embodiments, the DNA polymerase is a herpes simplex virus (HSV) UL30 DNA polymerase. In some embodiments, the DNA polymerase is a cytomegalovirus (CMV) UL54 DNA polymerase.
[0099] In some embodiments, the DNA polymerase is an archaeal polymerase. In some embodiments, the DNA polymerase is a Family B / pol I type DNA polymerase. For example, in some embodiments, the DNA polymerase is a homolog of Pfu from Pyrococcus furiosus. In some embodiments, the DNA polymerase is a pol II type DNA polymerase. For example, in some embodiments, the DNA polymerase is a homolog of / '. furiosus DP1 / DP22 -subunit polymerase. In some embodiments, the DNA polymerase lacks 5’ to 3’ nuclease activity. Suitable DNA polymerases (pol I or pol II) can be derived from archaea with optimal growth temperatures that are similar to the desired assay temperatures.
[0100] In some embodiments, the DNA polymerase comprises a thermostable archaeal DNA polymerase. In some embodiments, the thermostable DNA polymerase is isolated or derived from Pyrococcus species (furiosus, species GB-D, -woesii, abysii, horikoshii). Thermococcus species (kodakaraensis KOD1, litoralis, species 9 degrees North-7, species JDF-3, gorgonarius), Pyrodictium occultum, and Archaeoglobus fulgidus.
[0101] Polymerases may also be from eubacterial species. In some embodiments, the DNA polymerase is a Pol I family DNA polymerase. In some embodiments, the DNA polymerase is an E.coli Pol I DNA polymerase. In some embodiments, the DNA polymerase is a Pol II family DNA polymerase. In some embodiments, the DNA polymerase is a Pyrococcus furiosus (Pfu) Pol II DNA polymerase. In some embodiments, the DNA polymerase is a Pol III family DNA polymerase. In some embodiments, the DNA polymerase is a Pol IV family DNA polymerase. In some embodiments, the DNA polymerase is an E.coliWSGR Docket No. 59761-805.601Pol IV DNA polymerase. In some embodiments, the Pol I DNA polymerase is a DNA polymerase functional variant that lacks or has reduced 5' to 3' exonuclease activity.
[0102] Suitable thermostable pol I DNA polymerases can be isolated from a variety of thermophilic eubacteria, including Thermus species and Thermotoga maritima such as Thermus aquaticus (Taq), Thermus thermophilus (Tth) and Thermotoga maritima (Tma UlTma).
[0103] In some embodiments, a prime editor comprises an RNA-dependent DNA polymerase domain, for example, a reverse transcriptase (RT). A RT or an RT domain may be a wild-type RT domain, a full-length RT domain, or may be a functional mutant, a functional variant, or a functional fragment thereof. An RT or an RT domain of a prime editor may comprise a wild-type RT, or may be engineered or evolved to contain specific amino acid substitutions, truncations, or variants. An engineered RT may comprise sequences or amino acid changes different from a naturally occurring RT. In some embodiments, the engineered RT may have improved reverse transcription activity over a naturally occurring RT or RT domain. In some embodiments, the engineered RT may have improved features over a naturally occurring RT, for example, improved thermostability, reverse transcription efficiency, or target fidelity. In some embodiments, a prime editor comprising the engineered RT has improved prime editing efficiency over a prime editor having a reference naturally occurring RT.
[0104] In some embodiments, a prime editor comprises a eukaryotic RT, for example, a yeast, drosophila, rodent, or primate RT. In some embodiments, the prime editor comprises a Group II intron RT, for example, a. Geobacillus stearothermophilus Group II Intron (GsI-IIC) RT or a Eubacterium rectale group II intron (Eu.re. I2) RT. In some embodiments, the prime editor comprises a retron RT.
[0105] In some embodiments, a prime editor comprises a virus RT, for example, a retrovirus RT. Nonlimiting examples of virus RT include Moloney murine leukemia virus (M-MLV or MLVRT); human T-cell leukemia virus type 1 (HTLV-1) RT; bovine leukemia virus (BLV) RT; Rous Sarcoma Virus (RSV) RT; human immunodeficiency virus (HIV) RT, M-MFV RT, Avian Sarcoma-Leukosis Virus (ASLV) RT, Rous Sarcoma Virus (RSV) RT, Avian Myeloblastosis Virus (AMV) RT, Avian Erythroblastosis Virus (AEV) Helper Virus MCAV RT, Avian Myelocytomatosis Virus MC29 Helper Virus MCAV RT, Avian Reticuloendotheliosis Virus (REV-T) Helper Virus REV-A RT, Avian Sarcoma Virus UR2 Helper Virus (UR2AV) RT, Avian Sarcoma Virus Y73 Helper Virus YAV RT, Rous Associated Virus (RAV) RT, and Myeloblastosis Associated Virus (MAV) RT, all of which may be suitably used in the methods and composition described herein.
[0106] In some embodiments, the prime editor comprises a wild-type M-MLV RT. An exemplary sequence of a wild-type M-MLV RT is provided in SEQ ID NO: 500.
[0107] In some embodiments, the prime editor comprises a wild type M-MLV RT, a functional mutant, a functional variant, or a functional fragment thereof.
[0108] In some embodiments, the prime editor comprises a M-MLV RT comprising one or more of amino acid substitutions: P51X, S67X, E69X, L139X, T197X, D200X, H204X, F209X, E302X, T306X, F309X, W313X, T330X, L345X, L435X, N454X, D524X, E562X, D583X, H594X, L603X, E607X, or D653X as compared to a reference M-MLV RT where X is any amino acid other than the original aminoWSGR Docket No. 59761-805.601acid in the reference M-MLV RT. In some embodiments, the prime editor comprises a M-MLV RT comprising one or more of amino acid substitutions: P51L, S67K, E69K, L139P, T197A, D200N, H204R, F209N, E302K, E302R, T306K, F309N, W313F, T330P, L345G, L435G, N454K, D524G, E562Q, D583N, H594Q, L603W, E607K, or D653N as compared to a reference M-MLV RT. In some embodiments, the reference M-MLV RT is a variant M-MLV RT as set forth in SEQ ID NO: 501. In some embodiments, the reference M-MLV RT is a WT M-MLV RT as set forth in SEQ ID NO: 500.
[0109] In some embodiments, an RT variant may be a functional fragment of a reference RT that has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or up to 100, or up to 200, or up to 300, or up to 400, or up to 500 or more amino acid changes compared to the reference RT. In some embodiments, the RT variant comprises a fragment of the reference RT, such that the fragment is about 70% identical, about 80% identical, about 90% identical, about 95% identical, about 96% identical, about 97% identical, about 98% identical, about 99% identical, about 99.5% identical, or about 99.9% identical to the corresponding fragment of the reference RT. In some embodiments, the fragment is 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% of the amino acid length of a corresponding reference RT. In some embodiments, the reference M-MLV RT is a variant M-MLV RT as set forth in SEQ ID NO: 501. In some embodiments, the reference M-MLV RT is a WT M-MLV RT as set forth in SEQ ID NO: 500.
[0110] In some embodiments, the RT functional fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, or up to 600 or more amino acids in length.
[0111] In still other embodiments, the functional RT variant is truncated at the N-terminus or the C-terminus, or both, by a certain number of amino acids, which results in a truncated variant which still retains sufficient DNA polymerase function. In some embodiments, the RT truncated variant has a truncation of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, or 250 amino acids at the N-terminal end compared to a reference RT. In other embodiments, the RT truncated variant has a truncation of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, or 250 amino acids at the C-terminal end compared to a reference RT. In still other embodiments, the RT truncated variant has a truncation at the N-terminal and the C-terminal end compared to a reference RT. In some embodiments, the reference M-MLV RT is a variant M-MLV RT as set forth in SEQ ID NO: 501. In some embodiments, the reference M-MLV RT is a WT M-MLV RT as set forth in SEQ ID NO: 500.WSGR Docket No. 59761-805.601
[0112] In some embodiments, a prime editor comprises a M-MLV RT comprising one or more of amino acid substitutions D200N, T330P, L603W, T306K, orW313F as compared to a reference M-MLV RT. In some embodiments, a prime editor comprises a M-MLV RT comprising amino acid substitutions D200N, T330P, L603W, T306K, and W313F as compared to a reference M-MMLV RT. In some embodiments, a prime editor comprises a M-MLV RT comprising one or more of amino acid substitutions T128N, D200C, V223Y, T330P, T306K, or W313F as compared to a reference M-MLV RT. In some embodiments, a prime editor comprises a M-MLV RT comprising one or more of amino acid substitutions T128N, D200C, V223Y, T330P, T306K, and W313F as compared to a reference M-MLV RT. In some embodiments, a prime editor comprises a M-MLV RT comprising amino acid substitutions T128N, D200C, V223Y, T306K, W313F, and T330P, and is truncated between amino acids D497 and 1498 as compared to a reference M-MLV RT. In some embodiments, the reference M-MLV RT is a variant M-MLV RT as set forth in SEQ ID NO: 501. In some embodiments, the reference M-MLV RT is a WT M-MLV RT as set forth in SEQ ID NO: 500.
[0113] In some embodiments, an RT variant may be a functional fragment of a reference RT that has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, up to 100, up to 200, up to 300, up to 400, or up to 500 or more amino acid changes compared to a reference RT. In some embodiments, the RT variant comprises a fragment of a reference RT, such that the fragment is about 70% identical, about 80% identical, about 90% identical, about 95% identical, about 96% identical, about 97% identical, about 98% identical, about 99% identical, about 99.5% identical, or about 99.9% identical to the corresponding fragment of the reference RT. In some embodiments, the fragment is 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% identical, 96%, 97%, 98%, 99%, or 99.5% of the amino acid length of a corresponding reference RT. In some embodiments, the reference M-MLV RT is a variant M-MLV RT as set forth in SEQ ID NO: 501. In some embodiments, the reference M-MLV RT is a WT M-MLV RT as set forth in SEQ ID NO: 500.
[0114] Table 1 provides sequences of illustrative M-MLV RTs suitable for use with compositions and methods of the disclosure. In some embodiments, a prime editor comprises a M-MLV RT that comprises an amino acid sequence provided in Table 1, or a variant or functional fragment thereof. In some embodiments, the variant comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% identical to the amino acid sequence provided in Table 1. In some embodiments, the amino acid sequence of the M-MLV RT consists of the sequence provided in Table 1.
[0115] The prime editor may comprise the M-MLV RT having an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% identical to SEQ ID NO: 502 and comprising amino acid substitutions H8Y, D200N, T306K, W313F, T330P, and L603W as compared to wild-type M-MMLV RT as set forth in SEQ ID NO: 500. In some embodiments, the amino acid sequence of the M-MLV RT comprises SEQWSGR Docket No. 59761-805.601ID NO: 502. In some embodiments, the amino acid sequence of the M-MLV RT consists of SEQ ID NO: 502.
[0116] The prime editor may comprise the M-MLV RT having an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% identical to SEQ ID NO: 503 and comprising amino acid substitutions H8Y, T128N, D200C, V223Y, T306K, W313F, and T330P, and a A498-677 truncation as compared to wild-type M-MLV RT as set forth in SEQ ID NO: 500. In some embodiments, the amino acid sequence of the M-MLV RT comprises SEQ ID NO: 503. In some embodiments, the amino acid sequence of the M-MLV RT consists of SEQ ID NO: 503.
[0117] The prime editor may comprise the M-MLV RT having an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% identical to SEQ ID NO: 504 and comprising amino acid substitutions H8Y, D200N, T306K, W313F, T330P, a A498-677 truncation, and an addition of amino acids NSRLIN (SEQ ID NO: 623) at position 498 compared to wild-type M-MLV RT as set forth in SEQ ID NO: 500. In some embodiments, the amino acid sequence of the M-MLV RT comprises SEQ ID NO: 504. In some embodiments, the amino acid sequence of the M-MLV RT consists of SEQ ID NO: 504.
[0118] The prime editor may comprise the M-MLV RT having an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% identical to SEQ ID NO: 620 and comprising amino acid substitutions H8Y, T128N, D200C, V223Y, T306K, W313F, T330P, L435G, and a A498-677 truncation compared to wild-type M-MLV RT as set forth in SEQ ID NO: 500. In some embodiments, the amino acid sequence of the M-MLV RT comprises SEQ ID NO: 620. In some embodiments, the amino acid sequence of the M-MLV RT consists of SEQ ID NO: 620.
[0119] The prime editor may comprise the M-MLV RT having an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% identical to SEQ ID NO: 621 and comprising amino acid substitutions H8Y, P127Q, T128N, D200C, V223Y, T306K, W313F, T330P, L435G, a A498-677 truncation compared to wild-type M-MLV RT as set forth in SEQ ID NO: 500. In some embodiments, the amino acid sequence of the M-MLV RT comprises SEQ ID NO: 621. In some embodiments, the amino acid sequence of the M-MLV RT consists of SEQ ID NO: 621.
[0120] The prime editor may comprise the M-MLV RT having an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% identical to SEQ ID NO: 622 and comprising amino acid substitutions H8Y, T128N, D200C, V223Y, T287N, E302V, T306K, W313F, T330P, L435G, and a A498-677 truncation, compared to wild-type M-MLV RT as set forth in SEQ ID NO: 500. In some embodiments, the amino acid sequence of the M-MLV RT comprises SEQ ID NO: 622. In some embodiments, the amino acid sequence of the M-MLV RT consists of SEQ ID NO: 622.WSGR Docket No. 59761-805.601Table 1. Illustrative M-MLV RT SequencesSEQ Sequence Amino acid sequenceID descriptionNO:500 Wild-type M- TLNIEDEHRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQ MLVRT APLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSP WNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSG LPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLT WTRLPQGFKNSPTLFDEALHRDLADFRIQHPDLILLQYVDDLLLAA TSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKE GQRWLTEARKETVMGQPTPKTPRQLREFLGTAGFCRLWIPGFAEM AAPLYPLTKTGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPF ELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCL RMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSN ARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDIL AEAHGTRPDLTDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTET EVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFA TAHIHGEIYRRRGLLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPG HQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP501 Variant M-MLV TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQ RT (contains APLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSP H8Y compared WNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSG to SEQ ID LPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLT NO:500) WTRLPQGFKNSPTLFDEALHRDLADFRIQHPDLILLQYVDDLLLAA TSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKE GQRWLTEARKETVMGQPTPKTPRQLREFLGTAGFCRLWIPGFAEM AAPLYPLTKTGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPF ELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCL RMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSN ARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDIL AEAHGTRPDLTDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTET EVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFA TAHIHGEIYRRRGLLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPG HQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP502 Variant M-MLV TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQ RT APLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWSGR Docket No. 59761-805.601(contains H8Y, WNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSG D200N, T306K, LPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLT W313F, T330P, WTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAA L603W TSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKE compared to GQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEM SEQ ID AAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPF NO:500) ELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCL RMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSN ARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDIL AEAHGTRPDLTDQPLPDADHTWYTDGSSLLQEGQRKAGAAVTTET EVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFA TAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPG HQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSP503 Variant M-MLV TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQ RT (contains APLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSP H8Y, T128N, WNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPNVPNPYNLLSG D200C, V223Y, LPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLT T306K, W313F, WTRLPQGFKNSPTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAA T330P, and TSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKE A498-677 GQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEM compared to AAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPF SEQ ID NO: ELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCL 500) RMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSN ARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLD504 Variant M-MLV TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQ RT (contains APLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSP H8Y, D200N, WNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSG T306K, W313F, LPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLT T330P, A498- WTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAA 677, and TSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKE addition of GQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEM NSRLIN @498 AAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPF compared to ELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCL SEQ ID NO: RMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSN 500) ARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDNSREINWSGR Docket No. 59761-805.601620 Variant M-MLV TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQ RT (contains APLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSP H8Y, T128N, WNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPNVPNPYNLLSG D200C, V223Y, LPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLT T306K, W313F, WTRLPQGFKNSPTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAA T330P, L435G, TSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKE delta 498-677 GQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEM compared to AAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPF SEQ ID NO: ELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCL 500) RMVAAIAVLTKDAGKLTMGQPLVIGAPHAVEALVKQPPDRWLSN ARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLD621 Variant M-MLV TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQ RT (contains APLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSP H8Y, P127Q, WNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHQNVPNPYNLLSG T128N, D200C, LPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLT V223Y, T306K, WTRLPQGFKNSPTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAA W313F, T330P, TSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKE L435G, delta GQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRLFIPGFAEM 498-677 AAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPF compared to ELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCL SEQ ID NO: RMVAAIAVLTKDAGKLTMGQPLVIGAPHAVEALVKQPPDRWLSN 500) ARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLD 622 Variant M-MLV TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQ RT (contains APLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSP H8Y, T128N, WNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPNVPNPYNLLSG D200C, V223Y, LPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLT T287N, E302V, WTRLPQGFKNSPTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAA T306K, W313F, TSELDCQQGTRALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKE T330P, L435G, GQRWLTEARKENVMGQPTPKTPRQLRVFLGKAGFCRLFIPGFAEM delta 498-677 AAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPF compared to ELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCL SEQ ID NO: RMVAAIAVLTKDAGKLTMGQPLVIGAPHAVEALVKQPPDRWLSN 500) ARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDProgrammable DNA Binding Domain
[0121] In some embodiments, the DNA-binding domain of a prime editor is a programmable DNA binding domain.WSGR Docket No. 59761-805.601
[0122] A programmable DNA binding domain refers to a protein domain that is designed to bind a specific nucleic acid sequence, e.g., a target DNA or a target RNA. In some embodiments, the DNA-binding domain is a polynucleotide programmable DNA-binding domain that can associate with a guide polynucleotide (e.g., a PEgRNA) that guides the DNA-binding domain to a specific DNA sequence, e.g., a search target sequence in a target gene. In some embodiments, the DNA-binding domain comprises a Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) Associated (Cas) protein. A Cas protein may comprise any Cas protein described herein or a functional fragment or functional variant thereof. In some embodiments, a DNA-binding domain may also comprise a zinc -finger protein domain. In other cases, a DNA-binding domain comprises a transcription activator-like effector domain (TALE). In some embodiments, the DNA-binding domain comprises a DNA nuclease. For example, the DNA-binding domain of a prime editor may comprise an RNA-guided DNA endonuclease, e.g., a Cas protein. In some embodiments, the DNA-binding domain comprises a zinc finger nuclease (ZFN) or a transcription activator like effector domain nuclease (TALEN), where one or more zinc finger motifs or TALE motifs are associated with one or more nucleases, e.g., a Fok I nuclease domain.
[0123] In some embodiments, the DNA-binding domain comprises a nuclease activity. In some embodiments, the DNA-binding domain of a prime editor comprises an endonuclease domain having single strand DNA cleavage activity. For example, the endonuclease domain may comprise a FokI nuclease domain. In some embodiments, the DNA-binding domain of a prime editor comprises a nuclease having full nuclease activity. In some embodiments, the DNA-binding domain of a prime editor comprises a nuclease having modified or reduced nuclease activity as compared to a wild type endonuclease domain. For example, the endonuclease domain may comprise one or more amino acid substitutions as compared to a wild type endonuclease domain. In some embodiments, the DNA-binding domain of a prime editor has a nickase activity. In some embodiments, the DNA-binding domain of a prime editor comprises a Cas protein domain that is a nickase. In some embodiments, compared to a wild type Cas protein, the Cas nickase comprises one or more amino acid substitutions in a nuclease domain that reduces or abolishes its double strand nuclease activity but retains DNA binding activity. In some embodiments, the Cas nickase comprises an amino acid substitution in a HNH domain. In some embodiments, the Cas nickase comprises an amino acid substitution in a RuvC domain.
[0124] In some embodiments, the DNA-binding domain comprises a CRISPR associated protein (Cas protein) domain. A Cas protein may be a Class 1 or a Class 2 Cas protein. A Cas protein can be a type I, type II, type III, type IV, type V Cas protein, or a type VI Cas protein. Non-limiting examples of Cas proteins include Casl, Cas IB, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (e.g., Csnl or Csxl2), CaslO, CaslOd, Casl2a / Cpfl, Casl2b / C2cl, Casl2c / C2c3, Casl2d / CasY, Casl2e / CasX, Casl2g, Casl2h, Casl2i, Csyl, Csy2, Csy3, Csy4, Csel, Cse2, Cse3, Cse4, Cse5e, Cscl, Csc2, Csa5, Csnl, Csn2, Csml, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, CsxlS, Csxll, Csfl, Csf2, CsO, Csf4, Csdl, Csd2, Cstl, Cst2, Cshl, Csh2, Csal, Csa2, Csa3, Csa4, Csa5, Type II Cas effector proteins, Type V Cas effector proteins, Type VI Cas effector proteins, CARF, DinG, Cpfl,WSGR Docket No. 59761-805.601Casl2b / C2cl, Casl2c / C2c3, Casl2b / C2cl, Casl2c / C2c3, SpCas9(K855A), eSpCas9(l.l), SpCas9-HFl, hyper accurate Cas9 variant (HypaCas9), Cas <b. and homologues, modified or engineered variants, mutants, and / or functional fragments thereof. A Cas protein can be a chimeric Cas protein that is fused to other proteins or polypeptides. A Cas protein can be a chimera of various Cas proteins, for example, comprising domains of Cas proteins from different organisms.
[0125] A Cas protein, e.g., Cas9, can be from any suitable organism. In some aspects, the organism is Streptococcus pyogenes (S. pyogenes). In some aspects, the organism is Staphylococcus aureus (S. aureus). In some aspects, the organism is Streptococcus thermophilus (S. thermophilus). In some embodiments, the organism is Staphylococcus lugdunensis.
[0126] Non-limiting examples of suitable organism include Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinae spiralis, Streptomyces viridochromo genes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, AlicyclobacHlus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Pseudomonas aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, Acaryochloris marina, Leptotrichia shahii, and Francisella novicida. In some embodiments, the organism is Streptococcus pyogenes (S. pyogenes). In some embodiments, the organism is Staphylococcus aureus (S. aureus). In some embodiments, the organism is Streptococcus thermophilus (S. thermophilus). In some embodiments, the organism is Staphylococcus lugdunensis (S. lugdunensis).
[0127] In some embodiments, a Cas protein can be derived from a variety of bacterial species including, but not limited to, Veillonella atypical, Fusobacterium nucleatum, Filifactor alocis, Solobacterium moorei, Coprococcus catus, Treponema denticola, Peptoniphilus duerdenii, Catenibacterium mitsuokai, Streptococcus mutans, Listeria innocua, Staphylococcus pseudintermedius, Acidaminococcus intestine, Olsenella uli, Oenococcus kitaharae, Bifidobacterium bifidum, Lactobacillus rhamnosus, Lactobacillus gasseri, Finegoldia magna, Mycoplasma mobile, Mycoplasma gallisepticum, Mycoplasma ovipneumoniae, Mycoplasma canis, Mycoplasma synoviae, Eubacterium rectale, Streptococcus thermophilus, Eubacterium dolichum, Lactobacillus coryniformis subsp. Torquens, Ilyobacter polytropus, Ruminococcus albus, Akkermansia muciniphila, Acidothermus cellulolyticus, Bifidobacterium longum, Bifidobacterium dentium, Corynebacterium diphtheria, ElusimicrobiumWSGR Docket No. 59761-805.601minutum, Nitratifractor salsuginis, Sphaerochaeta globus, Fibrobacter succinogenes subsp. Succinogenes, Bacteroides fragilis, Capnocytophaga ochracea, Rhodopseudomonas palustris, Prevotella micans, Prevotella ruminicola, Flavobacterium columnare, Aminomonas paucivorans, Rhodospirillum rubrum, Candidatus Puniceispirillum marinum, Verminephrobacter eiseniae, Ralstonia syzygii, Dinoroseobacter shibae, Azospirillum, Nitrobacter hamburgensis, Bradyrhizobium, Wolinella succinogenes, Campylobacter jejuni subsp. Jejuni, Helicobacter mustelae, Bacillus cereus, Acidovorax ebreus, Clostridium perfringens, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria meningitidis, Pasteurella multocida subsp. Multocida, Sutterella wadsworthensis, proteobacterium, Legionella pneumophila, Parasutterella excrementihominis, Wolinella succinogenes, and Francisella novicida.
[0128] In some embodiments, a Cas protein, e.g., Cas9, can be a wild type or a modified form of a Cas protein. In some embodiments, a Cas protein, e.g., Cas9, can be a nuclease active variant, nuclease inactive variant, a nickase, or a functional variant or functional fragment of a wild type Cas protein. In some embodiments, a Cas protein, e.g., Cas9, can be a wild type or a modified form of a Cas protein. A Cas protein, e.g., Cas9, can be a nuclease active variant, nuclease inactive variant, a nickase, or a functional variant or functional fragment of a wild type Cas protein. In some embodiments, a Cas protein, e.g., Cas9, can comprise an amino acid change such as a deletion, insertion, substitution, fusion, chimera, or any combination thereof relative to a corresponding wild-type version of the Cas protein. In some embodiments, a Cas protein can be a polypeptide with at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or sequence similarity to a wild type exemplary Cas protein.
[0129] A Cas protein, e.g., Cas9, may comprise one or more domains. Non-limiting examples of Cas domains include, guide nucleic acid recognition and / or binding domain, nuclease domains (e.g., DNase or RNase domains, RuvC, HNH), DNA binding domain, RNA binding domain, helicase domains, proteinprotein interaction domains, and dimerization domains. In various embodiments, a Cas protein comprises a guide nucleic acid recognition and / or binding domain can interact with a guide nucleic acid, and one or more nuclease domains that comprise catalytic activity for nucleic acid cleavage.
[0130] In some embodiments, a Cas protein, e.g., Cas9, comprises one or more nuclease domains. A Cas protein can comprise an amino acid sequence having at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a nuclease domain (e.g., RuvC domain, HNH domain) of a wild-type Cas protein. In some embodiments, a Cas protein comprises a single nuclease domain. For example, a Cpfl may comprise a RuvC domain but lacks HNH domain. In some embodiments, a Cas protein comprises two nuclease domains, e.g., a Cas9 protein can comprise an HNH nuclease domain and a RuvC nuclease domain.
[0131] In some embodiments, a prime editor comprises a Cas protein, e.g., Cas9, wherein all nuclease domains of the Cas protein are active. In some embodiments, a prime editor comprises a Cas protein having one or more inactive nuclease domains. One or a plurality of the nuclease domains (e.g., RuvC, HNH) of a Cas protein can be deleted or mutated so that they are no longer functional or comprise reduced nuclease activity. In some embodiments, a Cas protein, e.g., Cas9, comprising mutations in aWSGR Docket No. 59761-805.601nuclease domain has reduced (e.g., nickase) or abolished nuclease activity while maintaining its ability to target a nucleic acid locus at a search target sequence when complexed with a guide nucleic acid, e.g., a PEgRNA.
[0132] In some embodiments, a prime editor comprises a Cas nickase that can bind to the target gene in a sequence-specific manner and generate a single-strand break at a protospacer within double -stranded DNA in the target gene, but not a double-strand break. For example, the Cas nickase can cleave the edit strand or the non-edit strand of the target gene, but may not cleave both. In some embodiments, a prime editor comprises a Cas nickase comprising two nuclease domains (e.g., Cas9), with one of the two nuclease domains modified to lack catalytic activity or deleted. In some embodiments, the Cas nickase of a prime editor comprises a nuclease inactive RuvC domain and a nuclease active HNH domain. In some embodiments, the Cas nickase of a prime editor comprises a nuclease inactive HNH domain and a nuclease active RuvC domain. In some embodiments, a prime editor comprises a Cas9 nickase having an amino acid substitution in the RuvC domain e.g., an amino acid substitution that reduces or abolishes nuclease activity of the RuvC domain. In some embodiments, the Cas9 nickase comprises a D10X amino acid substitution compared to a wild type. S', pyogenes Cas9, wherein X is any amino acid other than D. In some embodiments, a prime editor comprises a Cas9 nickase having an amino acid substitution in the HNH domain e.g., an amino acid substitution that reduces or abolishes nuclease activity of the HNH domain. In some embodiments, the Cas9 nickase comprises a H840X amino acid substitution compared to a wild type. S', pyogenes Cas9, wherein X is any amino acid other than H.
[0133] In some embodiments, a prime editor comprises a Cas protein that can bind to the target gene in a sequence-specific manner but lacks or has abolished nuclease activity and may not cleave either strand of a double stranded DNA in a target gene. Abolished activity or lacking activity can refer to an enzymatic activity less than 1%, less than 2%, less than 3%, less than 4%, less than 5%, less than 6%, less than 7%, less than 8%, less than 9%, or less than 10% activity compared to a wild-type exemplary activity (e.g., wild-type Cas9 nuclease activity). In some embodiments, a Cas protein of a prime editor completely lacks nuclease activity. A nuclease, e.g., Cas9, that lacks nuclease activity may be referred to as nuclease inactive or “nuclease dead” (abbreviated by “d”). A nuclease dead Cas protein (e.g., dCas, dCas9) can bind to a target polynucleotide but may not cleave the target polynucleotide. In some embodiments, a dead Cas protein is a dead Cas9 protein. In some embodiments, a prime editor comprises a nuclease dead Cas protein wherein all of the nuclease domains (e.g., both RuvC and HNH nuclease domains in a Cas9 protein; RuvC nuclease domain in a Cpfl protein) are mutated to lack catalytic activity, or are deleted.
[0134] A Cas protein can be modified. A Cas protein, e.g., Cas9, can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, and / or enzymatic activity. Cas proteins can also be modified to change any other activity or property of the protein, such as stability. For example, one or more nuclease domains of the Cas protein can be modified, deleted, or inactivated, or a Cas protein can be truncated to remove domains that are not essential for the function of the protein or to optimize (e.g., enhance or reduce) the activity of the Cas protein.WSGR Docket No. 59761-805.601
[0135] A Cas protein can be a fusion protein. For example, a Cas protein can be fused to a cleavage domain, an epigenetic modification domain, a transcriptional regulation domain, or a polymerase domain.A Cas protein can also be fused to a heterologous polypeptide providing increased or decreased stability. The fused domain or heterologous polypeptide can be located at the N-terminus, the C-terminus, or internally within the Cas protein.
[0136] In some embodiments, the Cas protein of a prime editor is a Class 2 Cas protein. In some embodiments, the Cas protein is a type II Cas protein. In some embodiments, the Cas protein is a Cas9 protein, a modified version of a Cas9 protein, a Cas9 protein homolog, mutant, variant, or a functional fragment thereof. As used herein, a Cas9, Cas9 protein, Cas9 polypeptide or a Cas9 nuclease refers to an RNA guided nuclease comprising one or more Cas9 nuclease domains and a Cas9 gRNA binding domain having the ability to bind a guide polynucleotide, e.g., a PEgRNA. A Cas9 protein may refer to a wild type Cas9 protein from any organism or a homolog, ortholog, or paralog from any organisms; any functional mutants or functional variants thereof; or any functional fragments or domains thereof. In some embodiments, a prime editor comprises a full-length Cas9 protein. In some embodiments, the Cas9 protein can generally comprises at least about 50%, 60%, 70%, 80%, 90%, 100% sequence identity to a wildtype reference Cas9 protein (e.g., Cas9 from. S', pyogenes). In some embodiments, the Cas9 comprises an amino acid change such as a deletion, insertion, substitution, fusion, chimera, or any combination thereof as compared to a wild type reference Cas9 protein.
[0137] In some embodiments, a Cas9 protein may comprise a Cas9 protein from Streptococcus pyogenes (Sp), Staphylococcus aureus (Sa), Streptococcus canis (Sc), Streptococcus thermophilus (St), Staphylococcus lugdunensis (Siu), Neisseria meningitidis (Nm), Campylobacter jejuni (Cj), Francisella novicida (Fn), or Treponema denticola (Td), or any Cas9 homolog or ortholog from an organism known in the art. In some embodiments, a Cas9 polypeptide is a SpCas9 polypeptide, e.g., comprising an amino acid sequence as set forth in NCBI Accession No. WP_038431314 or a fragment or variant thereof. In some embodiments, a Cas9 polypeptide is a SaCas9 polypeptide, e.g., comprising an amino acid sequence as set forth in Uniprot Accession No. J7RUA5 or a fragment or variant thereof. In some embodiments, a Cas9 polypeptide is a ScCas9 polypeptide, e.g., comprising an amino acid sequence as set forth in Uniprot Accession No. A0A3P5YA78 or a fragment or variant thereof. In some embodiments, a Cas9 polypeptide is a StCas9 polypeptide, e.g., comprising an amino acid sequence as set forth in NCBI Accession No. WP_007896501.1 or a fragment or variant thereof. In some embodiments, a Cas9 polypeptide is a SluCas9 polypeptide, e.g., comprising an amino acid sequence as set forth in any of NCBI Accession No. WP_230580236.1 or WP_250638315.1 or WP_242234150.1, WP_241435384.1, WP_002460848.1, KAK58371.1, or a fragment or variant thereof. In some embodiments, a Cas9 polypeptide is aNmCas9 polypeptide, e.g., comprising an amino acid sequence as set forth in any of NCBI Accession No. WP_002238326.1 or WP_061704949.1 or a fragment or variant thereof. In some embodiments, a Cas9 polypeptide is a CjCas9 polypeptide, e.g., comprising an amino acid sequence as set forth in any of NCBI Accession No. WP_100612036.1, WP_116882154.1, WP_116560509.1, WP_116484194.1, WP_116479303.1, WP_115794652.1, WP_100624872.1, or a fragment or variantWSGR Docket No. 59761-805.601thereof. In some embodiments, a Cas9 polypeptide is a FnCas9 polypeptide, e.g., comprising the amino acid sequence as set forth in Uniprot Accession No. A0Q5Y3 or a fragment or variant thereof. In some embodiments, a Cas9 polypeptide is a TdCas9 polypeptide, e.g., comprising the amino acid sequence as set forth in NCBI Accession No. WP_147625065.1 or a fragment or variant thereof. In some embodiments, a Cas9 polypeptide is a chimera comprising domains from two or more of the organisms described herein or those known in the art. In some embodiments, a Cas9 polypeptide is a Cas9 polypeptide from Streptococcus macacae, e.g., comprising the amino acid sequence as set forth in NCBI Accession No. WP_003079701.1 or a fragment or variant thereof. In some embodiments, a Cas9 polypeptide is a Cas9 polypeptide generated by replacing a PAM interaction domain of a SpCas9 with that of a Streptococcus macacae Cas9 (Spy -mac Cas9). Exemplary Cas9 and Cas9 nickase variants are provided in Table 2.
[0138] In some embodiments, a prime editor comprises a DNA binding domain that comprises an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one of the sequences set forth in Table 2. In some embodiments, the DNA binding domain comprises an amino acid sequence that has no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 differences e.g., mutations e.g., deletions, substitutions and / or insertions compared to any one of the amino acid sequences set forth in Table 2.
[0139] In some embodiments, a prime editor comprises a Cas9 protein that comprises an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one of the sequences set forth in Table 2. In some embodiments, a prime editor comprises a Cas9 protein is a Cas9 nickase that comprises an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one of the nickase sequences set forth in Table 2. In some embodiments, a Cas9 protein comprises an amino acid sequence that is selected from the group consisting of the sequences set forth in Table 2. In some embodiments, a prime editor comprises a Cas9 protein that comprises an amino acid sequence that lacks a N-terminus methionine relative to an amino acid sequence set forth Table 2. In some embodiments, the prime editing compositions or prime editing systems disclosed herein comprises a polynucleotide (e.g., a DNA, or an RNA, e.g., an mRNA) that encodes a Cas9 protein that comprises an amino acid sequence that is at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any one of the sequences set forth in Table 2.WSGR Docket No. 59761-805.601
[0140] In some embodiments, a Cas9 protein comprises a Cas9 protein from Streptococcus pyogenes (Sp), e.g., as according to NC_002737.2:854751-858857 or the protein encoded by UniProt Q99ZW2, e.g., as according to SEQ ID NO: 505. In some embodiments, a prime editor comprises a Cas9 protein (e.g., a SpCas9) as according to any one of the sequences set forth in SEQ ID NOs: 505-508 or a variant thereof. In some embodiments, the Cas9 protein is a SpCas9. In some embodiments, a SpCas9 can be a wild type SpCas9, a SpCas9 variant, or a nickase SpCas9. In some embodiments, the SpCas9 lacks the N-terminus methionine relative to a corresponding SpCas9 (e.g., a wild type SpCas9, a SpCas9 variant or a nickase SpCas9). In some embodiments, a prime editor comprises a Cas9 protein, having an amino acid sequence as according to SEQ ID NO: 505, not including the N-terminus methionine. In some embodiments, a wild type SpCas9 comprises an amino acid sequence set forth in SEQ ID NO: 505. In some embodiments, a prime editor comprises a Cas9 protein comprising one or more mutations (e.g., amino acid substitutions, insertions and / or deletions) relative to a corresponding wild type Cas9 protein (e.g., a wild type SpCas9). In some embodiments, the Cas9 protein comprising one or more mutations relative to a wild type Cas9 (e.g., a wild type SpCas9) protein comprises an amino acid sequence set forth in SEQ ID NO: 506, 507, or 508. Exemplary Streptococcus pyogenes Cas9 (SpCas9) amino acid sequence useful in the prime editors disclosed herein are provided in Table 2.
[0141] In some embodiments, a prime editor comprises a Cas9 protein (e.g., a SluCas9) as according to any one of the SEQ ID NOS: 509-511 or a variant thereof. In some embodiments, a prime editor comprises a Cas9 protein from Staphylococcus lugdunensis (SluCas9) e.g., as according to any one of the SEQ ID NOs: 509-511 or a variant thereof. In some embodiments, the Cas9 protein is a SluCas9. In some embodiments, a SluCas9 can be a wild type SluCas9, a SluCas9 variant, or a nickase SluCas9. In some embodiments, the SluCas9 lacks the N-terminus methionine relative to a corresponding SluCas9 (e.g., a wild type SluCas9, a SluCas9 variant or a nickase SluCas9). In some embodiments, a prime editor comprises a Cas9 protein, having an amino acid sequence as according to SEQ ID NO: 509, not including the N-terminus methionine. In some embodiments, a wild type SluCas9 comprises an amino acid sequence set forth in SEQ ID NO: 509. In some embodiments, a prime editor comprises a Cas9 protein comprising one or more mutations (e.g., amino acid substitutions, insertions and / or deletions) relative to a corresponding wild type Cas9 protein (e.g., a wild type SluCas9). In some embodiments, the Cas9 protein comprising one or mutations relative to a wild type Cas9 protein comprises an amino acid sequence set forth in SEQ ID NO: 510 or SEQ ID NO: 511. Exemplary Staphylococcus lugdunensis Cas9 (SluCas9) amino acid sequence useful in the prime editors disclosed herein are provided in Table 2.
[0142] In some embodiments, a prime editor comprises a Cas9 protein from Staphylococcus aureus (SaCas9) e.g., as according to any of the SEQ ID NOS: 512-514, or a variant thereof. In some embodiments, a prime editor comprises a Cas9 protein from Staphylococcus aureus (SaCas9) e.g., as according to any one of the SEQ ID NOS: 512-514, or a variant thereof. In some embodiments, the Cas9 protein is a SaCas9. In some embodiments, a SaCas9 can be a wild type SaCas9, a SaCas9 variant, or a nickase SaCas9. In some embodiments, the SaCas9 lacks the N-terminus methionine relative to a corresponding SaCas9 (e.g., a wild type SaCas9, a SaCas9 variant or a nickase SaCas9). In someWSGR Docket No. 59761-805.601embodiments, a prime editor comprises a Cas9 protein, having an amino acid sequence as according to SEQ ID NO: 512, not including the N-terminus methionine. In some embodiments, a wild type SaCas9 comprises an amino acid sequence set forth in SEQ ID NO: 512. In some embodiments, a prime editor comprises a Cas9 protein comprising one or more mutations (e.g., amino acid substitutions, insertions and / or deletions relative to a corresponding wild type Cas9 protein (e.g., a wild type SaCas9). In some embodiments, the Cas9 protein comprising one or more mutations relative to a wild type Cas9 protein comprises an amino acid sequence set forth in SEQ ID NO: 513 or SEQ ID NO: 514. Exemplary Staphylococcus aureus Cas9 (SaCas9) amino acid sequence useful in the prime editors disclosed herein are provided Table 2.
[0143] In some embodiments, a prime editor comprises a Cas9 protein as according to any one of the sequences set forth in SEQ ID NOs: 515-523, 530-532 or a variant thereof. In some embodiments, the Cas9 protein is a Cas9 variant, for example, a SpCas9 variant (e.g., SpCas9-NG, SpCas9-NGA, SpRY, or SpG). In some embodiments, the Cas9 protein lacks the N-terminus methionine relative to a corresponding Cas9 protein (e.g., a Cas9 variant set forth in any one of SEQ ID NOs: 515, 516, 518, 519, 521, 522, 530, or 531). In some embodiments, a prime editor comprises a Cas9 protein (e.g., a Cas9 variant), having an amino acid sequence as according to any one of SEQ ID NOs: 515, 518, 521, or 530 not including the N-terminus methionine. In some embodiments, a prime editor comprises a Cas9 protein comprising one or more mutations (e.g., amino acid substitutions, insertions and / or deletions) relative to a corresponding Cas9 protein (e.g., a Cas9 protein set forth in any one of SEQ ID NOs: 515, 518, 521, or 530). In some embodiments, the Cas9 protein comprising one or mutations relative to a corresponding Cas9 protein comprises an amino acid sequence set forth in any one of SEQ ID NOs: 516, 517, 519, 520, 522, 523, 531, or 532.
[0144] In some embodiments, a prime editor comprises a Cas9 protein as according to any one of the sequences set forth in any one of SEQ ID NOs: 533-535 or 584, 585 or a variant thereof. In some embodiments, the Cas9 protein is a Cas9 variant, for example, a SpCas9-NRTH variant. In some embodiments, the Cas9 protein lacks the N-terminus methionine relative to a corresponding Cas9 protein (e.g., a Cas9 variant set forth in any one of SEQ ID NO SEQ ID NOs: 533-535 or 584, 585. In some embodiments, a prime editor comprises a Cas9 protein (e.g., a Cas9 variant), having an amino acid sequence as according to any one of SEQ ID NOs: 533-535 or 584, 585 not including the N-terminus methionine. In some embodiments, a prime editor comprises a Cas9 protein comprising one or more mutations (e.g., amino acid substitutions, insertions and / or deletions) relative to a corresponding Cas9 protein (e.g., a Cas9 protein set forth in any one of SEQ ID NOs: 533-535 or 584, 585. In some embodiments, the Cas9 protein comprising one or mutations relative to a corresponding Cas9 protein comprises an amino acid sequence set forth in any one of SEQ ID NOs: 533-535 or 584, 585.
[0145] In some embodiments, a Cas9 protein is a chimeric Cas9, e.g., modified Cas9, e.g., synthetic RNA-guided nucleases (sRGNs), e.g., modified by DNA family shuffling, e.g., sRGN3.1, sRGN3.3. In some embodiments, the DNA family shuffling comprises, fragmentation and reassembly of parental Cas9 genes, e.g., one or more of Cas9s from Staphylococcus hyicus (Shy), Staphylococcus lugdunensis (Siu),WSGR Docket No. 59761-805.601Staphylococcus microti (Smi), and Staphylococcus pasteuri (Spa). In some embodiments, a modified sluCas9 shows increased editing efficiency and / or specificity relative to a sluCas9 that is not modified. In some embodiments, a modified Cas9, e.g., a sRGN shows at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 150%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, or at least 1000% increase in editing efficiency compared to a Cas9 that is not modified. In some embodiments, a Cas9, e.g., a sRGN shows at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 150%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, or at least 1000% increase in specificity compared to a Cas9 that is not modified. In some embodiments, a Cas9, e.g., a sRGN shows at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 150%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, or at least 1000% increase in cleavage activity compared to a Cas9 that is not modified. In some embodiments, a Cas9, e.g., a sRGN shows ability to cleave a 5'-NNGG-3' PAM-containing target. In some embodiments, a prime editor comprises a Cas9 protein (e.g., a chimeric Cas9), e.g., as according any one of the sequences set forth in SEQ ID NOs: 524-529, or a variant thereof. Exemplary amino acid sequences of Cas9 protein (e.g., sRGN) useful in the prime editors disclosed herein are provided below in SEQ ID NOs: 524-529. In some embodiments, a prime editor comprises a Cas9 protein, that lacks aN-terminus methionine relative to SEQ ID NO: 524 or SEQ ID NO: 527. In some embodiments, a prime editor comprises a Cas9 protein comprising one or more mutations (e.g., amino acid substitutions, insertions and / or deletions) relative to a corresponding Cas9 protein (e.g., a Cas9 protein set forth in SEQ ID NO: 524 or SEQ ID NO: 527). In some embodiments, the Cas9 protein comprising one or mutations relative to a corresponding Cas9 protein comprises an amino acid sequence set forth in any one of SEQ ID NOs: 525, 526, 528, or 529.
[0146] Table 2: Exemplary Cas protein sequencesSEQ Sequence Amino acid sequenceID descriptionNO:505 wild type MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKN Streptococc LIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV us Pyogenes DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRK Cas9 KLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQ (SpCas9) LVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKK NGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLA QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDE HHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFY KFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELH AILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTR KSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLL YEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKV TVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRWSGR Docket No. 59761-805.601 RRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHD DSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDE LVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELG SQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYD VDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNY WRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYK VREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVR KMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNG ETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPK RNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLK SVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELE NGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDN EQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRD KPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATL IHQSITGLYETRIDLSQLGGD506 SpCas9 MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKN H840A LIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV nickase DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRK KLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQ LVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKK NGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLA QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDE HHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFY KFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELH AILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTR KSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLL YEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKV TVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDF LDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKR RRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHD DSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDE LVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELG SQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYD VDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNY WRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYK VREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVR KMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNG ETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPK RNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLK SVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELE NGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDN EQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRD KPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATL IHQSITGLYETRIDLSQLGGD507 Met (-) DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLI SpCas9 GALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVD H840A DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK nickase LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKN GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFI KPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEWSGR Docket No. 59761-805.601 ETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYE YFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLD NEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRR YTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSL TFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELV KVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQ ILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVD AIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWR QLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHV AQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREI NNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETG EIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNS DKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSV KELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELEN GRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNE QKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDK PIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLI HQSITGLYETRIDLSQLGGD508 Met (-) DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLI CAS9 GALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVD (R221K DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK N394K LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL H840A) VQTYNQLFEENPINASGVDAKAILSARLSKSRKLENLIAQLPGEKKN nickase GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFI KPILEKMDGTEELLVKLKREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYE YFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLD NEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRR YTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSL TFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELV KVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQ ILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVD AIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWR QLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHV AQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREI NNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETG EIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNS DKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSV KELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELEN GRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNE QKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDK PIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLI HQSITGLYETRIDLSQLGGD509 wild type MNQKFILGLDIGITSVGYGLIDYETKNIIDAGVRLFPEANVENNEGR Staphylococ RSKRGSRRLKRRRIHRLERVKKLLEDYNLLDQSQIPQSTNPYAIRVK cus GLSEALSKDELVIALLHIAKRRGIHKIDVIDSNDDVGNELSTKEQLN lugdunensis KNSKLLKDKFVCQIQLERMNEGQVRGEKNRFKTADIIKEIIQLLNVQ (Slu)Cas9 KNFHQLDENFINKYIELVEMRREYFEGPGKGSPYGWEGDPKAWYETLMGHCTYFPDELRSVKYAYSADLFNALNDLNNLVIQRDGLSKLEYWSGR Docket No. 59761-805.601 HEKYHIIENVFKQKKKPTLKQIANEINVNPEDIKGYRITKSGKPQFTE FKLYHDLKSVLFDQSILENEDVLDQIAEILTIYQDKDSIKSKLTELDIL LNEEDKENIAQLTGYTGTHRLSLKCIRLVLEEQWYSSRNQMEIFTHL NIKPKKINLTAANKIPKAMIDEFILSPVVKRTFGQAINLINKIIEKYGV PEDIIIELARENNSKDKQKFINEMQKKNENTRKRINEIIGKYGNQNA KRLVEKIRLHDEQEGKCLYSLESIPLEDLLNNPNHYEVDHIIPRSVSF DNSYHNKVLVKQSENSKKSNLTPYQYFNSGKSKLSYNQFKQHILNL SKSQDRISKKKKEYLLEERDINKFEVQKEFINRNLVDTRYATRELTN YLKAYFSANNMNVKVKTINGSFTDYLRKVWKFKKERNHGYKHHA EDALIIANADFLFKENKKLKAVNSVLEKPEIESKQLDIQVDSEDNYS EMFIIPKQVQDIKDFRNFKYSHRVDKKPNRQLINDTLYSTRKKDNST YIVQTIKDIYAKDNTTLKKQFDKSPEKFLMYQHDPRTFEKLEVIMK QYANEKNPLAKYHEETGEYLTKYSKKNNGPIVKSLKYIGNKLGSHL DVTHQFKSSTKKLVKLSIKPYRFDVYLTDKGYKFITISYLDVLKKDN YYYIPEQKYDKLKLGKAIDKNAKFIASFYKNDLIKLDGEIYKIIGVNS DTRNMIELDLPDIRYKEYCELNNIKGEPRIKKTIGKKVNSIEKLTTDV LGNVFTNTQYTKPQLLFKRGN510 SluCas9 MNQKFILGLDIGITSVGYGLIDYETKNIIDAGVRLFPEANVENNEGR N582A RSKRGSRRLKRRRIHRLERVKKLLEDYNLLDQSQIPQSTNPYAIRVK nickase GLSEALSKDELVIALLHIAKRRGIHKIDVIDSNDDVGNELSTKEQLN KNSKLLKDKFVCQIQLERMNEGQVRGEKNRFKTADIIKEIIQLLNVQ KNFHQLDENFINKYIELVEMRREYFEGPGKGSPYGWEGDPKAWYE TLMGHCTYFPDELRSVKYAYSADLFNALNDLNNLVIQRDGLSKLEY HEKYHIIENVFKQKKKPTLKQIANEINVNPEDIKGYRITKSGKPQFTE FKLYHDLKSVLFDQSILENEDVLDQIAEILTIYQDKDSIKSKLTELDIL LNEEDKENIAQLTGYTGTHRLSLKCIRLVLEEQWYSSRNQMEIFTHL NIKPKKINLTAANKIPKAMIDEFILSPVVKRTFGQAINLINKIIEKYGV PEDIIIELARENNSKDKQKFINEMQKKNENTRKRINEIIGKYGNQNA KRLVEKIRLHDEQEGKCLYSLESIPLEDLLNNPNHYEVDHIIPRSVSF DNSYHNKVLVKQSEASKKSNLTPYQYFNSGKSKLSYNQFKQHILNL SKSQDRISKKKKEYLLEERDINKFEVQKEFINRNLVDTRYATRELTN YLKAYFSANNMNVKVKTINGSFTDYLRKVWKFKKERNHGYKHHA EDALIIANADFLFKENKKLKAVNSVLEKPEIESKQLDIQVDSEDNYS EMFIIPKQVQDIKDFRNFKYSHRVDKKPNRQLINDTLYSTRKKDNST YIVQTIKDIYAKDNTTLKKQFDKSPEKFLMYQHDPRTFEKLEVIMK QYANEKNPLAKYHEETGEYLTKYSKKNNGPIVKSLKYIGNKLGSHL DVTHQFKSSTKKLVKLSIKPYRFDVYLTDKGYKFITISYLDVLKKDN YYYIPEQKYDKLKLGKAIDKNAKFIASFYKNDLIKLDGEIYKIIGVNS DTRNMIELDLPDIRYKEYCELNNIKGEPRIKKTIGKKVNSIEKLTTDV LGNVFTNTQYTKPQLLFKRGN511 Met (-) NQKFILGLDIGITSVGYGLIDYETKNIIDAGVRLFPEANVENNEGRRS SluCas9 KRGSRRLKRRRIHRLERVKKLLEDYNLLDQSQIPQSTNPYAIRVKGL nickase SEALSKDELVIALLHIAKRRGIHKIDVIDSNDDVGNELSTKEQLNKN SKLLKDKFVCQIQLERMNEGQVRGEKNRFKTADIIKEIIQLLNVQKN FHQLDENFINKYIELVEMRREYFEGPGKGSPYGWEGDPKAWYETL MGHCTYFPDELRSVKYAYSADLFNALNDLNNLVIQRDGLSKLEYH EKYHIIENVFKQKKKPTLKQIANEINVNPEDIKGYRITKSGKPQFTEF KLYHDLKSVLFDQSILENEDVLDQIAEILTIYQDKDSIKSKLTELDILL NEEDKENIAQLTGYTGTHRLSLKCIRLVLEEQWYSSRNQMEIFTHLN IKPKKINLTAANKIPKAMIDEFILSPVVKRTFGQAINLINKIIEKYGVP EDIIIELARENNSKDKQKFINEMQKKNENTRKRINEIIGKYGNQNAK RLVEKIRLHDEQEGKCLYSLESIPLEDLLNNPNHYEVDHIIPRSVSFD NSYHNKVLVKQSEASKKSNLTPYQYFNSGKSKLSYNQFKQHILNLS KSQDRISKKKKEYLLEERDINKFEVQKEFINRNLVDTRYATRELTNY LKAYFSANNMNVKVKTINGSFTDYLRKVWKFKKERNHGYKHHAEDALIIANADFLFKENKKLKAVNSVLEKPEIESKQLDIQVDSEDNYSEWSGR Docket No. 59761-805.601 MFIIPKQVQDIKDFRNFKYSHRVDKKPNRQLINDTLYSTRKKDNSTY IVQTIKDIYAKDNTTLKKQFDKSPEKFLMYQHDPRTFEKLEVIMKQ YANEKNPLAKYHEETGEYLTKYSKKNNGPIVKSLKYIGNKLGSHLD VTHQFKSSTKKLVKLSIKPYRFDVYLTDKGYKFITISYLDVLKKDNY YYIPEQKYDKLKLGKAIDKNAKFIASFYKNDLIKLDGEIYKIIGVNSD TRNMIELDLPDIRYKEYCELNNIKGEPRIKKTIGKKVNSIEKLTTDVLGNVFTNTQYTKPQLLFKRGN Staphylococ MKRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGR cus aureus RSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARV Cas9 KGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQIS (SaCas9) RNSKALEEKYVAELQLERLKKDGEVRGSINRFKTSDYVKEAKQLLK VQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYE MLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLE YYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFT NLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNS ELTQEEIEQISNLKGYTGTHNLSLKAINLILDELWHTNDNQIAIFNRL KLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGL PNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAK YLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFD NSFNNKVLVKQEENSKKGNRTPFQYLSSSDSKISYETFKKHILNLAK GKGRISKTKKEYLLEERDINRFSVQKDFINRNLVDTRYATRGLMNL LRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAED ALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMPEIETEQEY KEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKG NTLIVNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLI MEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIKYYGNKLN AHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDV IKKENYYEVNSKCYEEAKKLKKISNQAEFIASFYNNDLIKINGELYR VIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIK KYSTDILGNLYEVKSKKHPQIIKKG513 SaCas9 MKRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGR N580A RSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARV nickase KGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQIS RNSKALEEKYVAELQLERLKKDGEVRGSINRFKTSDYVKEAKQLLK VQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYE MLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLE YYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFT NLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNS ELTQEEIEQISNLKGYTGTHNLSLKAINLILDELWHTNDNQIAIFNRL KLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGL PNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAK YLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFD NSFNNKVLVKQEEASKKGNRTPFQYLSSSDSKISYETFKKHILNLAK GKGRISKTKKEYLLEERDINRFSVQKDFINRNLVDTRYATRGLMNL LRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAED ALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMPEIETEQEY KEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKG NTLIVNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLI MEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIKYYGNKLN AHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDV IKKENYYEVNSKCYEEAKKLKKISNQAEFIASFYNNDLIKINGELYR VIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIK KYSTDILGNLYEVKSKKHPQIIKKG514 Met (-) KRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRR SaCas9 SKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARVKnickase GLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRWSGR Docket No. 59761-805.601 NSKALEEKYVAELQLERLKKDGEVRGSINRFKTSDYVKEAKQLLK VQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYE MLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLE YYEKFQIIENVFKQKKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFT NLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNS ELTQEEIEQISNLKGYTGTHNLSLKAINLILDELWHTNDNQIAIFNRL KLVPKKVDLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGL PNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAK YLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFD NSFNNKVLVKQEEASKKGNRTPFQYLSSSDSKISYETFKKHILNLAK GKGRISKTKKEYLLEERDINRFSVQKDFINRNLVDTRYATRGLMNL LRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYKHHAED ALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMPEIETEQEY KEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKG NTLIVNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDPQTYQKLKLI MEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIKYYGNKLN AHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDV IKKENYYEVNSKCYEEAKKLKKISNQAEFIASFYNNDLIKINGELYR VIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIK KYSTDILGNLYEVKSKKHPQIIKKG515 SpCas9-NG MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKN (VRVRFRR LIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV ) DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRK KLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQ LVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKK NGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLA QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDE HHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFY KFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELH AILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTR KSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLL YEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKV TVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDF LDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKR RRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHD DSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDE LVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELG SQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYD VDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNY WRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYK VREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVR KMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNG ETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESIRPK RNSDKLIARKKDWDPKKYGGFVSPTVAYSVLVVAKVEKGKSKKLK SVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELE NGRKRMLASARFLQKGNELALPSKYVNFLYLASHYEKLKGSPEDN EQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRD KPIREQAENIIHLFTLTNLGAPRAFKYFDTTIDRKVYRSTKEVLDATL IHQSITGLYETRIDLSQLGGD516 spCas9-NG MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKN (H840A_V LIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV RVRFRR) DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRK Nickase KLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAWSGR Docket No. 59761-805.601 QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDE HHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFY KFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELH AILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTR KSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLL YEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKV TVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDF LDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKR RRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHD DSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDE LVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELG SQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYD VDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNY WRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYK VREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVR KMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNG ETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESIRPK RNSDKLIARKKDWDPKKYGGFVSPTVAYSVLVVAKVEKGKSKKLK SVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELE NGRKRMLASARFLQKGNELALPSKYVNFLYLASHYEKLKGSPEDN EQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRD KPIREQAENIIHLFTLTNLGAPRAFKYFDTTIDRKVYRSTKEVLDATL IHQSITGLYETRIDLSQLGGD517 Met (-) DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLI SpCas9-NG GALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVD Nickase DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKN GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFI KPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYE YFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLD NEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRR YTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSL TFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELV KVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQ ILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVD AIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWR QLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHV AQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREI NNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETG EIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESIRPKRNS DKLIARKKDWDPKKYGGFVSPTVAYSVLVVAKVEKGKSKKLKSV KELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELEN GRKRMLASARFLQKGNELALPSKYVNFLYLASHYEKLKGSPEDNE QKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDK PIREQAENIIHLFTLTNLGAPRAFKYFDTTIDRKVYRSTKEVLDATLIHQSITGLYETRIDLSQLGGDWSGR Docket No. 59761-805.601518 spCas9- MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKN NGA LIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV(VRQR) DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRK KLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQ LVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKK NGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLA QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDE HHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFY KFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELH AILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTR KSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLL YEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKV TVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDF LDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKR RRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHD DSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDE LVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELG SQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYD VDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNY WRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYK VREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVR KMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNG ETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPK RNSDKLIARKKDWDPKKYGGFVSPTVAYSVLVVAKVEKGKSKKLK SVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELE NGRKRMLASARELQKGNELALPSKYVNFLYLASHYEKLKGSPEDN EQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRD KPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKQYRSTKEVLDATL IHQSITGLYETRIDLSQLGGD519 spCas9- MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKN NGA LIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV(H840A V DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRK RQR) KLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQ Nickase LVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKK NGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLA QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDE HHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFY KFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELH AILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTR KSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLL YEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKV TVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDF LDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKR RRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHD DSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDE LVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELG SQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYD VDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNY WRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYK VREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVR KMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNG ETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPK RNSDKLIARKKDWDPKKYGGFVSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELEWSGR Docket No. 59761-805.601 NGRKRMLASARELQKGNELALPSKYVNFLYLASHYEKLKGSPEDN EQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRD KPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKQYRSTKEVLDATL IHQSITGLYETRIDLSQLGGD520 Met(-) DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLI spCas9- GALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVD NGA DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKNickase LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKN GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFI KPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYE YFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLD NEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRR YTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSL TFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELV KVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQ ILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVD AIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWR QLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHV AQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREI NNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETG EIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNS DKLIARKKDWDPKKYGGFVSPTVAYSVLVVAKVEKGKSKKLKSV KELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELEN GRKRMLASARELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNE QKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDK PIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKQYRSTKEVLDATLI HQSITGLYETRIDLSQLGGD521 SpRY Cas9 MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKN LIGALLFDSGETAERTRLKRTARRRYTRRKNRICYLQEIFSNEMAKV DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRK KLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQ LVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKK NGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLA QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDE HHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFY KFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELH AILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTR KSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLL YEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKV TVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDF LDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKR RRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHD DSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDE LVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELG SQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYD VDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNY WRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRWSGR Docket No. 59761-805.601 KMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNG ETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESIRPK RNSDKLIARKKDWDPKKYGGFLWPTVAYSVLVVAKVEKGKSKKL KSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFEL ENGRKRMLASAKQLQKGNELALPSKYVNFLYLASHYEKLKGSPED NEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHR DKPIREQAENIIHLFTLTRLGAPRAFKYFDTTIDPKQYRSTKEVLDAT LIHQSITGLYETRIDLSQLGGD522 SpRY Cas9 MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKN (H840A) LIGALLFDSGETAERTRLKRTARRRYTRRKNRICYLQEIFSNEMAKV Nickase DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRK KLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQ LVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKK NGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLA QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDE HHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFY KFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELH AILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTR KSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLL YEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKV TVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDF LDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKR RRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHD DSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDE LVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELG SQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYD VDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNY WRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYK VREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVR KMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNG ETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESIRPK RNSDKLIARKKDWDPKKYGGFLWPTVAYSVLVVAKVEKGKSKKL KSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFEL ENGRKRMLASAKQLQKGNELALPSKYVNFLYLASHYEKLKGSPED NEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHR DKPIREQAENIIHLFTLTRLGAPRAFKYFDTTIDPKQYRSTKEVLDAT LIHQSITGLYETRIDLSQLGGD523 Met(-) DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLI SpRY Cas9 GALLFDSGETAERTRLKRTARRRYTRRKNRICYLQEIFSNEMAKVD Nickase DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK (H839A) LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKN GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFI KPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYE YFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLD NEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRR YTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSL TFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELV KVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDWSGR Docket No. 59761-805.601 AIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWR QLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHV AQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREI NNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETG EIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESIRPKRNS DKLIARKKDWDPKKYGGFLWPTVAYSVLVVAKVEKGKSKKLKSV KELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELEN GRKRMLASAKQLQKGNELALPSKYVNFLYLASHYEKLKGSPEDNE QKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDK PIREQAENIIHLFTLTRLGAPRAFKYFDTTIDPKQYRSTKEVLDATLIH QSITGLYETRIDLSQLGGD524 SRGN3.1 MNQKFILGLDIGITSVGYGLIDYETKNIIDAGVRLFPEANVENNEGR RSKRGSRRLKRRRIHRLERVKLLLTEYDLINKEQIPTSNNPYQIRVKG LSEILSKDELAIALLHLAKRRGIHNVDVAADKEETASDSLSTKDQIN KNAKFLESRYVCELQKERLENEGHVRGVENRFLTKDIVREAKKIIDT QMQYYPEIDETFKEKYISLVETRREYFEGPGQGSPFGWNGDLKKWY EMLMGHCTYFPQELRSVKYAYSADLFNALNDLNNLIIQRDNSEKLE YHEKYHIIENVFKQKKKPTLKQIAKEIGVNPEDIKGYRITKSGTPEFT SFKLFHDLKKVVKDHAILDDIDLLNQIAEILTIYQDKDSIVAELGQLE YLMSEADKQSISELTGYTGTHSLSLKCMNMIIDELWHSSMNQMEVF TYLNMRPKKYELKGYQRIPTDMIDDAILSPVVKRTFIQSINVINKVIE KYGIPEDIIIELARENNSDDRKKFINNLQKKNEATRKRINEIIGQTGN QNAKRIVEKIRLHDQQEGKCLYSLESIPLEDLLNNPNHYEVDHIIPRS VSFDNSYHNKVLVKQSENSKKSNLTPYQYFNSGKSKLSYNQFKQHI LNLSKSQDRISKKKKEYLLEERDINKFEVQKEFINRNLVDTRYATRE LTNYLKAYFSANNMNVKVKTINGSFTDYLRKVWKFKKERNHGYK HHAEDALIIANADFLFKENKKLKAVNSVLEKPEIETKQLDIQVDSED NYSEMFIIPKQVQDIKDFRNFKYSHRVDKKPNRQLINDTLYSTRKKD NSTYIVQTIKDIYAKDNTTLKKQFDKSPEKFLMYQHDPRTFEKLEVI MKQYANEKNPLAKYHEETGEYLTKYSKKNNGPIVKSLKYIGNKLG SHLDVTHQFKSSTKKLVKLSIKNYRFDVYLTEKGYKFVTIAYLNVF KKDNYYYIPKDKYQELKEKKKIKDTDQFIASFYKNDLIKLNGDLYK IIGVNSDDRNIIELDYYDIKYKDYCEINNIKGEPRIKKTIGKKTESIEKF TTDVLGNLYLHSTEKAPQLIFKRGL525 SRGN3.1 MNQKFILGLDIGITSVGYGLIDYETKNIIDAGVRLFPEANVENNEGR (N585A) RSKRGSRRLKRRRIHRLERVKLLLTEYDLINKEQIPTSNNPYQIRVKG Nickase LSEILSKDELAIALLHLAKRRGIHNVDVAADKEETASDSLSTKDQIN KNAKFLESRYVCELQKERLENEGHVRGVENRFLTKDIVREAKKIIDT QMQYYPEIDETFKEKYISLVETRREYFEGPGQGSPFGWNGDLKKWY EMLMGHCTYFPQELRSVKYAYSADLFNALNDLNNLIIQRDNSEKLE YHEKYHIIENVFKQKKKPTLKQIAKEIGVNPEDIKGYRITKSGTPEFT SFKLFHDLKKVVKDHAILDDIDLLNQIAEILTIYQDKDSIVAELGQLE YLMSEADKQSISELTGYTGTHSLSLKCMNMIIDELWHSSMNQMEVF TYLNMRPKKYELKGYQRIPTDMIDDAILSPVVKRTFIQSINVINKVIE KYGIPEDIIIELARENNSDDRKKFINNLQKKNEATRKRINEIIGQTGN QNAKRIVEKIRLHDQQEGKCLYSLESIPLEDLLNNPNHYEVDHIIPRS VSFDNSYHNKVLVKQSEASKKSNLTPYQYFNSGKSKLSYNQFKQHI LNLSKSQDRISKKKKEYLLEERDINKFEVQKEFINRNLVDTRYATRE LTNYLKAYFSANNMNVKVKTINGSFTDYLRKVWKFKKERNHGYK HHAEDALIIANADFLFKENKKLKAVNSVLEKPEIETKQLDIQVDSED NYSEMFIIPKQVQDIKDFRNFKYSHRVDKKPNRQLINDTLYSTRKKD NSTYIVQTIKDIYAKDNTTLKKQFDKSPEKFLMYQHDPRTFEKLEVI MKQYANEKNPLAKYHEETGEYLTKYSKKNNGPIVKSLKYIGNKLG SHLDVTHQFKSSTKKLVKLSIKNYRFDVYLTEKGYKFVTIAYLNVFKKDNYYYIPKDKYQELKEKKKIKDTDQFIASFYKNDLIKLNGDLYKWSGR Docket No. 59761-805.601 IIGVNSDDRNIIELDYYDIKYKDYCEINNIKGEPRIKKTIGKKTESIEKF TTDVLGNLYLHSTEKAPQLIFKRGL526 Met(-) NQKFILGLDIGITSVGYGLIDYETKNIIDAGVRLFPEANVENNEGRRS SRGN3.1 KRGSRRLKRRRIHRLERVKLLLTEYDLINKEQIPTSNNPYQIRVKGLS (N584A)Nic EILSKDELAIALLHLAKRRGIHNVDVAADKEETASDSLSTKDQINKN kase AKFLESRYVCELQKERLENEGHVRGVENRFLTKDIVREAKKIIDTQ MQYYPEIDETFKEKYISLVETRREYFEGPGQGSPFGWNGDLKKWYE MLMGHCTYFPQELRSVKYAYSADLFNALNDLNNLIIQRDNSEKLEY HEKYHIIENVFKQKKKPTLKQIAKEIGVNPEDIKGYRITKSGTPEFTS FKLFHDLKKVVKDHAILDDIDLLNQIAEILTIYQDKDSIVAELGQLEY LMSEADKQSISELTGYTGTHSLSLKCMNMIIDELWHSSMNQMEVFT YLNMRPKKYELKGYQRIPTDMIDDAILSPVVKRTFIQSINVINKVIEK YGIPEDIIIELARENNSDDRKKFINNLQKKNEATRKRINEIIGQTGNQ NAKRIVEKIRLHDQQEGKCLYSLESIPLEDLLNNPNHYEVDHIIPRSV SFDNSYHNKVLVKQSEASKKSNLTPYQYFNSGKSKLSYNQFKQHIL NLSKSQDRISKKKKEYLLEERDINKFEVQKEFINRNLVDTRYATREL TNYLKAYFSANNMNVKVKTINGSFTDYLRKVWKFKKERNHGYKH HAEDALIIANADFLFKENKKLKAVNSVLEKPEIETKQLDIQVDSEDN YSEMFIIPKQVQDIKDFRNFKYSHRVDKKPNRQLINDTLYSTRKKDN STYIVQTIKDIYAKDNTTLKKQFDKSPEKFLMYQHDPRTFEKLEVIM KQYANEKNPLAKYHEETGEYLTKYSKKNNGPIVKSLKYIGNKLGSH LDVTHQFKSSTKKLVKLSIKNYRFDVYLTEKGYKFVTIAYLNVFKK DNYYYIPKDKYQELKEKKKIKDTDQFIASFYKNDLIKLNGDLYKIIG VNSDDRNIIELDYYDIKYKDYCEINNIKGEPRIKKTIGKKTESIEKFTT DVLGNLYLHSTEKAPQLIFKRGL527 SRGN3.3 MNQKFILGLDIGITSVGYGLIDYETKNIIDAGVRLFPEANVENNEGR RSKRGSRRLKRRRIHRLERVKLLLTEYDLINKEQIPTSNNPYQIRVKG LSEILSKDELAIALLHLAKRRGIHNVDVAADKEETASDSLSTKDQIN KNAKFLESRYVCELQKERLENEGHVRGVENRFLTKDIVREAKKIIDT QMQYYPEIDETFKEKYISLVETRREYFEGPGQGSPFGWNGDLKKWY EMLMGHCTYFPQELRSVKYAYSADLFNALNDLNNLIIQRDNSEKLE YHEKYHIIENVFKQKKKPTLKQIAKEIGVNPEDIKGYRITKSGTPEFT SFKLFHDLKKVVKDHAILDDIDLLNQIAEILTIYQDKDSIVAELGQLE YLMSEADKQSISELTGYTGTHSLSLKCMNMIIDELWHSSMNQMEVF TYLNMRPKKYELKGYQRIPTDMIDDAILSPVVKRTFIQSINVINKVIE KYGIPEDIIIELARENNSDDRKKFINNLQKKNEATRKRINEIIGQTGN QNAKRIVEKIRLHDQQEGKCLYSLESIPLEDLLNNPNHYEVDHIIPRS VSFDNSYHNKVLVKQSENSKKSNLTPYQYFNSGKSKLSYNQFKQHI LNLSKSQDRISKKKKEYLLEERDINKFEVQKEFINRNLVDTRYATRE LTSYLKAYFSANNMDVKVKTINGSFTNHLRKVWRFDKYRNHGYK HHAEDALIIANADFLFKENKKLQNTNKILEKPTIENNTKKVTVEKEE DYNNVFETPKLVEDIKQYRDYKFSHRVDKKPNRQLINDTLYSTRMK DEHDYIVQTITDIYGKDNTNLKKQFNKNPEKFLMYQNDPKTFEKLSI IMKQYSDEKNPLAKYYEETGEYLTKYSKKNNGPIVKKIKLLGNKVG NHLDVTNKYENSTKKLVKLSIKNYRFDVYLTEKGYKFVTIAYLNVF KKDNYYYIPKDKYQELKEKKKIKDTDQFIASFYKNDLIKLNGDLYK IIGVNSDDRNIIELDYYDIKYKDYCEINNIKGEPRIKKTIGKKTESIEKF TTDVLGNLYLHSTEKAPQLIFKRGL528 sRGN3.3(N MNQKFILGLDIGITSVGYGLIDYETKNIIDAGVRLFPEANVENNEGR 585A) RSKRGSRRLKRRRIHRLERVKLLLTEYDLINKEQIPTSNNPYQIRVKG Nickase LSEILSKDELAIALLHLAKRRGIHNVDVAADKEETASDSLSTKDQIN KNAKFLESRYVCELQKERLENEGHVRGVENRFLTKDIVREAKKIIDT QMQYYPEIDETFKEKYISLVETRREYFEGPGQGSPFGWNGDLKKWY EMLMGHCTYFPQELRSVKYAYSADLFNALNDLNNLIIQRDNSEKLE YHEKYHIIENVFKQKKKPTLKQIAKEIGVNPEDIKGYRITKSGTPEFTSFKLFHDLKKVVKDHAILDDIDLLNQIAEILTIYQDKDSIVAELGQLEWSGR Docket No. 59761-805.601 YLMSEADKQSISELTGYTGTHSLSLKCMNMIIDELWHSSMNQMEVF TYLNMRPKKYELKGYQRIPTDMIDDAILSPVVKRTFIQSINVINKVIE KYGIPEDIIIELARENNSDDRKKFINNLQKKNEATRKRINEIIGQTGN QNAKRIVEKIRLHDQQEGKCLYSLESIPLEDLLNNPNHYEVDHIIPRS VSFDNSYHNKVLVKQSEASKKSNLTPYQYFNSGKSKLSYNQFKQHI LNLSKSQDRISKKKKEYLLEERDINKFEVQKEFINRNLVDTRYATRE LTSYLKAYFSANNMDVKVKTINGSFTNHLRKVWRFDKYRNHGYK HHAEDALIIANADFLFKENKKLQNTNKILEKPTIENNTKKVTVEKEE DYNNVFETPKLVEDIKQYRDYKFSHRVDKKPNRQLINDTLYSTRMK DEHDYIVQTITDIYGKDNTNLKKQFNKNPEKFLMYQNDPKTFEKLSI IMKQYSDEKNPLAKYYEETGEYLTKYSKKNNGPIVKKIKLLGNKVG NHLDVTNKYENSTKKLVKLSIKNYRFDVYLTEKGYKFVTIAYLNVF KKDNYYYIPKDKYQELKEKKKIKDTDQFIASFYKNDLIKLNGDLYK IIGVNSDDRNIIELDYYDIKYKDYCEINNIKGEPRIKKTIGKKTESIEKF TTDVLGNLYLHSTEKAPQLIFKRGL529 Met(-) NQKFILGLDIGITSVGYGLIDYETKNIIDAGVRLFPEANVENNEGRRS sRGN3.3(N KRGSRRLKRRRIHRLERVKLLLTEYDLINKEQIPTSNNPYQIRVKGLS 584A)Nicka EILSKDELAIALLHLAKRRGIHNVDVAADKEETASDSLSTKDQINKN se AKFLESRYVCELQKERLENEGHVRGVENRFLTKDIVREAKKIIDTQ MQYYPEIDETFKEKYISLVETRREYFEGPGQGSPFGWNGDLKKWYE MLMGHCTYFPQELRSVKYAYSADLFNALNDLNNLIIQRDNSEKLEY HEKYHIIENVFKQKKKPTLKQIAKEIGVNPEDIKGYRITKSGTPEFTS FKLFHDLKKVVKDHAILDDIDLLNQIAEILTIYQDKDSIVAELGQLEY LMSEADKQSISELTGYTGTHSLSLKCMNMIIDELWHSSMNQMEVFT YLNMRPKKYELKGYQRIPTDMIDDAILSPVVKRTFIQSINVINKVIEK YGIPEDIIIELARENNSDDRKKFINNLQKKNEATRKRINEIIGQTGNQ NAKRIVEKIRLHDQQEGKCLYSLESIPLEDLLNNPNHYEVDHIIPRSV SFDNSYHNKVLVKQSEASKKSNLTPYQYFNSGKSKLSYNQFKQHIL NLSKSQDRISKKKKEYLLEERDINKFEVQKEFINRNLVDTRYATREL TSYLKAYFSANNMDVKVKTINGSFTNHLRKVWRFDKYRNHGYKH HAEDALIIANADFLFKENKKLQNTNKILEKPTIENNTKKVTVEKEED YNNVFETPKLVEDIKQYRDYKFSHRVDKKPNRQLINDTLYSTRMKD EHDYIVQTITDIYGKDNTNLKKQFNKNPEKFLMYQNDPKTFEKLSII MKQYSDEKNPLAKYYEETGEYLTKYSKKNNGPIVKKIKLLGNKVG NHLDVTNKYENSTKKLVKLSIKNYRFDVYLTEKGYKFVTIAYLNVF KKDNYYYIPKDKYQELKEKKKIKDTDQFIASFYKNDLIKLNGDLYK IIGVNSDDRNIIELDYYDIKYKDYCEINNIKGEPRIKKTIGKKTESIEKF TTDVLGNLYLHSTEKAPQLIFKRGL530 SpG MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKN LIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRK KLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQ LVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKK NGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLA QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDE HHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFY KFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELH AILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTR KSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLL YEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKV TVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDF LDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKR RRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHD DSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDE LVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDWSGR Docket No. 59761-805.601 VDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNY WRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYK VREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVR KMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNG ETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPK RNSDKLIARKKDWDPKKYGGFLWPTVAYSVLVVAKVEKGKSKKL KSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFEL ENGRKRMLASAKQLQKGNELALPSKYVNFLYLASHYEKLKGSPED NEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHR DKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKQYRSTKEVLDAT LIHQSITGLYETRIDLSQLGGD531 SpG(H840A MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKN ) Nickase LIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRK KLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQ LVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKK NGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLA QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDE HHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFY KFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELH AILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTR KSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLL YEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKV TVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDF LDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKR RRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHD DSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDE LVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELG SQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYD VDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNY WRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYK VREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVR KMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNG ETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPK RNSDKLIARKKDWDPKKYGGFLWPTVAYSVLVVAKVEKGKSKKL KSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFEL ENGRKRMLASAKQLQKGNELALPSKYVNFLYLASHYEKLKGSPED NEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHR DKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKQYRSTKEVLDAT LIHQSITGLYETRIDLSQLGGD532 Met(-) DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLI SpG(H839A GALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVD ) Nickase DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKN GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFI KPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYE YFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRWSGR Docket No. 59761-805.601 YTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSL TFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELV KVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQ ILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVD AIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWR QLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHV AQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREI NNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETG EIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNS DKLIARKKDWDPKKYGGFLWPTVAYSVLVVAKVEKGKSKKLKSV KELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELEN GRKRMLASAKQLQKGNELALPSKYVNFLYLASHYEKLKGSPEDNE QKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDK PIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKQYRSTKEVLDATLI HQSITGLYETRIDLSQLGGD533 Met (-) DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLI Cas9-NRTH GALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVD (I322V, DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK S409I, LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL E427G, VQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKN R654L, GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI R753G, GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRYDEH R1114G, HQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYK D1135N, FIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGIIPHQIHLGELHAI D1180G, LRRQGDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKS G1218S, EETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYE E1219V, YFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV Q1221H, KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLD P1249S, NEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRLR E1253K, YTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSL P1321S, TFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELV D1332G, KVMGGHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQ R1335L) ILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVD HIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWR QLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHV AQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREI NNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETG EIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKGNS DKLIARKKDWDPKKYGGFNSPTVAYSVLVVAKVEKGKSKKLKSV KELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFELEN GRKRMLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNK QKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDK PIREQAENIIHLFTLTNLGASAAFKYFDTTIGRKLYTSTKEVLDATLI HQSITGLYETRIDLSQLGGD534 Cas9-NRTH MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKN Nickase LIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV (I322V, DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRK S409I, KLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQ E427G, LVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKK R654L, NGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLA R753G, QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRYDE H840A, HHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFY R1114G, KFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGIIPHQIHLGELHAD1135N, ILRRQGDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKWSGR Docket No. 59761-805.601D1180G, SEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLY G1218S, EYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT E1219V, VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFL Q1221H, DNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRL P1249S, RYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDD E1253K, SLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDEL P1321S, VKVMGGHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGS D1332G, QILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDV R1335L) DAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYW RQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKH VAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRK MIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGE TGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKG NSDKLIARKKDWDPKKYGGFNSPTVAYSVLVVAKVEKGKSKKLKS VKELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFELE NGRKRMLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDN KQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRD KPIREQAENIIHLFTLTNLGASAAFKYFDTTIGRKLYTSTKEVLDATL IHQSITGLYETRIDLSQLGGD535 Met (-) DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLI Cas9-NRTH GALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVD Nickase DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKN GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRYDEH HQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYK FIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGIIPHQIHLGELHAI LRRQGDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKS EETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYE YFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLD NEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRLR YTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSL TFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELV KVMGGHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQ ILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVD AIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWR QLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHV AQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREI NNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETG EIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKGNS DKLIARKKDWDPKKYGGFNSPTVAYSVLVVAKVEKGKSKKLKSV KELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFELEN GRKRMLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNK QKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDK PIREQAENIIHLFTLTNLGASAAFKYFDTTIGRKLYTSTKEVLDATLI HQSITGLYETRIDLSQLGGD584 variant MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKN Cas9-NRTH LIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV nickase DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRK (I322V, KLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQ R221K, LVQTYNQLFEENPINASGVDAKAILSARLSKSRKLENLIAQLPGEKKN394K, NGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAWSGR Docket No. 59761-805.601S409I, QIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRYDE E427G, HHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFY R654L, KFIKPILEKMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHA R753G, ILRRQGDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRK H840A, SEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLY R1114G, EYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVT D1135N, VKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFL D1180G, DNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRL G1218S, RYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDD E1219V, SLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDEL Q1221H, VKVMGGHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGS P1249S, QILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDV E1253K, DAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYW P1321S, RQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKH D1332G, VAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVR R1335L) EINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRK MIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGE TGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKG NSDKLIARKKDWDPKKYGGFNSPTVAYSVLVVAKVEKGKSKKLKS VKELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFELE NGRKRMLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDN KQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRD KPIREQAENIIHLFTLTNLGASAAFKYFDTTIGRKLYTSTKEVLDATL IHQSITGLYETRIDLSQLGGD585 Met- variant DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLI Cas9-NRTH GALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVD nickase DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQTYNQLFEENPINASGVDAKAILSARLSKSRKLENLIAQLPGEKKN GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRYDEH HQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYK FIKPILEKMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAI LRRQGDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKS EETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYE YFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLD NEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRLR YTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSL TFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELV KVMGGHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQ ILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVD AIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWR QLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHV AQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREI NNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETG EIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKGNS DKLIARKKDWDPKKYGGFNSPTVAYSVLVVAKVEKGKSKKLKSV KELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFELEN GRKRMLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNK QKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDK PIREQAENIIHLFTLTNLGASAAFKYFDTTIGRKLYTSTKEVLDATLI HQSITGLYETRIDLSQLGGD648 i Spy Mac DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLICas9 GALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDWSGR Docket No. 59761-805.601 DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL VQTYNQLFEENPINASGVDAKAILSARLSKSRKLENLIAQLPGEKKN GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFI KPILEKMDGTEELLVKLKREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYE YFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLD NEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRR YTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSL TFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELV KVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQ ILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVD AIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWR QLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHV AQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREI NNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETG EIVWDKGRDFATVRKVLSMPQVNIVKKTEIQTVGQNGGLFDDNPK SPLEVTPSKLVPLKKELNPKKYGGYQKPTTAYPVLLITDTKQLIPISV MNKKQFEQNPVKFLRDRGYQQVGKNDFIKLPKYTLVDIGDGIKRL WASSKEIHKGNQLVVSKKSQILLYHAHHLDSDLSNDYLQNHNQQF DVLFNEIISFSKKCKLGKEHIQKIENVYSNKKNSASIEELAESFIKLLG FTQLGATSPFNFLGVKLNQKQYKGKKDYILPCTEGTLIRQSITGLYE TRVDLSKIGED649 SpCas9(R22 DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLI IK, I322V, GALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVD N394K, DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK S409I, LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL E427G, VQTYNQLFEENPINASGVDAKAILSARLSKSRKLENLIAQLPGEKKN R654L, GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI R753G, GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRYDEH H840A, HQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYK R1114G, FIKPILEKMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAI D1135N, LRRQGDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKS V1139A, EETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYE D1180G, YFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV E1219V, KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLD Q1221H, NEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRLR A 1320V, YTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSL R1333K) TFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELV KVMGGHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQ ILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVD AIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWR QLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHV AQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREI NNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETG EIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKGNS DKLIARKKDWDPKKYGGFNSPTAAYSVLVVAKVEKGKSKKLKSV KELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFELEN GRKRMLASAGVLHKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKWSGR Docket No. 59761-805.601 PIREQAENIIHLFTLTNLGVPAAFKYFDTTIDKKRYTSTKEVLDATLI HQSITGLYETRIDLSQLGGD650 NRRH- DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLI SpCas9(H84 GALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVD 0A) nickase DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK with LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL R22IK, VQTYNQLFEENPINASGVDAKAILSARLSKSRKLENLIAQLPGEKKN N394K + GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI T1337R GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRYDEH HQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYK FIKPILEKMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAI LRRQGDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKS EETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYE YFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLD NEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRLR YTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSL TFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELV KVMGGHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQ ILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVD AIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWR QLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHV AQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREI NNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETG EIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKGNS DKLIARKKDWDPKKYGGFNSPTAAYSVLVVAKVEKGKSKKLKSV KELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFELEN GRKRMLASAGVLHKGNELALPSKYVNFLYLASHYEKLKGSPEDNE QKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDK PIREQAENIIHLFTLTNLGVPAAFKYFDTTIDKKRYRSTKEVLDATLI HQSITGLYETRIDLSQLGGD651 VRQR- DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLI SpCas9(H84 GALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVD 0A) nickase DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK with LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL R221K, VQTYNQLFEENPINASGVDAKAILSARLSKSRKLENLIAQLPGEKKN N394K GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFI KPILEKMDGTEELLVKLKREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYE YFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLD NEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRR YTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSL TFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELV KVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQ ILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVD AIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWR QLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHV AQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREI NNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSWSGR Docket No. 59761-805.601 DKLIARKKDWDPKKYGGFVSPTVAYSVLVVAKVEKGKSKKLKSV KELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELEN GRKRMLASARELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNE QKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDK PIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKQYRSTKEVLDATLI HQSITGLYETRIDLSQLGGD652 VRQR- DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLI SpCas9(H84 GALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVD 0A) nickase DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK +D1332K LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL with VQTYNQLFEENPINASGVDAKAILSARLSKSRKLENLIAQLPGEKKN R22IK, GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI N394K GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFI KPILEKMDGTEELLVKLKREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYE YFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLD NEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRR YTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSL TFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELV KVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQ ILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVD AIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWR QLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHV AQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREI NNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETG EIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNS DKLIARKKDWDPKKYGGFVSPTVAYSVLVVAKVEKGKSKKLKSV KELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELEN GRKRMLASARELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNE QKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDK PIREQAENIIHLFTLTNLGAPAAFKYFDTTIKRKQYRSTKEVLDATLI HQSITGLYETRIDLSQLGGD653 VRQR- DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLI SpCas9(H84 GALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVD 0A) nickase DSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKK +LiniR LVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQL with VQTYNQLFEENPINASGVDAKAILSARLSKSRKLENLIAQLPGEKKN R221K. GLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQI N394K GDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHH QDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFI KPILEKMDGTEELLVKLKREDLLRKQRTFDNGSIPHQIHLGELHAIL RRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSE ETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYE YFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTV KQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLD NEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRR YTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSL TFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELV KVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQ ILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVD AIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVWSGR Docket No. 59761-805.601 AQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREI NNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMI AKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETG EIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESIRPKRNS DKLIARKKDWDPKKYGGFVSPTVAYSVLVVAKVEKGKSKKLKSV KELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELEN GRKRMLASARELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNE QKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDK PIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKQYRSTKEVLDATLIHQSITGLYETRIDLSQLGGD
[0147] In some embodiments, a Cas9 protein comprises a variant Cas9 protein containing one or more amino acid substitutions. In some embodiments, a wildtype Cas9 protein comprises a RuvC domain and an HNH domain. In some embodiments, a prime editor comprises a nuclease active Cas9 protein that may cleave both strands of a double stranded target DNA sequence. In some embodiments, the nuclease active Cas9 protein comprises a functional RuvC domain and a functional HNH domain. In some embodiments, a prime editor comprises a Cas9 nickase that can bind to a guide polynucleotide and recognize a target DNA, but can cleave only one strand of a double stranded target DNA. In some embodiments, the Cas9 nickase comprises only one functional RuvC domain or one functional HNH domain. In some embodiments, a prime editor comprises a Cas9 that has a non -functional HNH domain and a functional RuvC domain. In some embodiments, the prime editor can cleave the edit strand (z.e., the PAM strand), but not the non-edit strand of a double stranded target DNA sequence. In some embodiments, a prime editor comprises a Cas9 having a non-functional RuvC domain that can cleave the target strand (i.e., the non-PAM strand), but not the edit strand of a double stranded target DNA sequence. In some embodiments, a prime editor comprises a Cas9 that has neither a functional RuvC domain nor a functional HNH domain, which may not cleave any strand of a double stranded target DNA sequence.
[0148] In some embodiments, a prime editor comprises a Cas9 having a mutation in the RuvC domain that reduces or abolishes the nuclease activity of the RuvC domain. In some embodiments, the Cas9 comprises a mutation at amino acid DIO as compared to a wild type SpCas9 as set forth in SEQ ID NO: 505, or a corresponding mutation thereof. In some embodiments, the Cas9 comprises a D10A mutation as compared to a wild type SpCas9 as set forth in SEQ ID NO: 505, or a corresponding mutation thereof. In some embodiments, the Cas9 polypeptide comprises a mutation at amino acid DIO, G12, and / or G17 as compared to a wild type SpCas9 as set forth in SEQ ID NO: 505, or a corresponding mutation thereof. In some embodiments, the Cas9 polypeptide comprises a D10A mutation, a G12A mutation, and / or a G17A mutation as compared to a wild type SpCas9 as set forth in SEQ ID NO: 505, or a corresponding mutation thereof.
[0149] In some embodiments, a prime editor comprises a Cas9 polypeptide having a mutation in the HNH domain that reduces or abolishes the nuclease activity of the HNH domain. In some embodiments, the Cas9 polypeptide comprises a mutation at amino acid H840 as compared to a wild type SpCas9 as set forth in SEQ ID NO: 505, or a corresponding mutation thereof. In some embodiments, the Cas9 polypeptide comprises a H840A mutation as compared to a wild type SpCas9 as set forth in SEQ ID NO:WSGR Docket No. 59761-805.601505, or a corresponding mutation thereof. In some embodiments, the Cas9 polypeptide comprises a mutation at amino acid E762, D839, H840, N854, N856, N863, H982, H983, A984, D986, and / or a A987 as compared to a wild type SpCas9 as set forth in SEQ ID NO: 505, or a corresponding mutation thereof. In some embodiments, the Cas9 polypeptide comprises a E762A, D839A, H840A, N854A, N856A, N863A, H982A, H983A, A984A, and / or a D986A mutation as compared to a wild type SpCas9 as set forth in SEQ ID NO: 505, or a corresponding mutation thereof. In some embodiments, the Cas9 polypeptide comprises a mutation at amino acid residue R221, N394, and / or H840 as compared to a wild type SpCas9 (e.g., SEQ ID NO: 505). In some embodiments, the Cas9 polypeptide comprises a R221K, N394L, and / or H840A mutation as compared to a wild type SpCas9 as set forth in SEQ ID NO: 505, or a corresponding mutation thereof. In some embodiments, the Cas9 polypeptide comprises a mutation at amino acid residue R220, N393, and / or H839 as compared to a wild type SpCas9 (e.g., SEQ ID NO: 505) lacking a N-terminal methionine, or a corresponding mutation thereof. In some embodiments, the Cas9 polypeptide comprises a R220K, N393K, and / or H839A mutation as compared to a wild type SpCas9 (as set forth in SEQ ID NO: 505) lacking a N-terminal methionine, or a corresponding mutation thereof.
[0150] In some embodiments, a prime editor comprises a Cas9 having one or more amino acid substitutions in both the HNH domain and the RuvC domain that reduce or abolish the nuclease activity of both the HNH domain and the RuvC domain. In some embodiments, the prime editor comprises a nuclease inactive Cas9, or a nuclease dead Cas9 (dCas9). In some embodiments, the dCas9 comprises a H840X substitution and a D10X mutation compared to a wild type SpCas9 as set forth in SEQ ID NO: 505or corresponding mutations thereof, wherein X is any amino acid other than H for the H840X substitution and any amino acid other than D for the DI OX substitution. In some embodiments, the dead Cas9 comprises a H840A and a D10A mutation as compared to a wild type SpCas9 as set forth in SEQ ID NO: 505, or corresponding mutations thereof.
[0151] In some embodiments, the N-terminal methionine is removed from the amino acid sequence of a Cas9 nickase, or from any Cas9 variant, ortholog, or equivalent disclosed or contemplated herein. For example, methionine-minus (Met (-)) Cas9 nickases include any one of the sequences set forth in SEQ ID NOs: 507, 508, 511, 514, 517, 520, 523, 526, 529, 532, 533, 535, 585, or a variant thereof having an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity thereto.
[0152] Besides dead Cas9 and Cas9 nickase variants, the Cas9 proteins used herein may also include other Cas9 variants having at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% sequence identity to any reference Cas9 protein, including any wild type Cas9, or mutant Cas9 (e.g., a dead Cas9 or Cas9 nickase), or fragment Cas9, or circular permutant Cas9, or other variant of Cas9 disclosed herein or known in the art. In some embodiments, a Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to a reference Cas9, e.g., a wild type Cas9. In some embodiments, the Cas9 variant comprisesWSGR Docket No. 59761-805.601a fragment of a reference Cas9 (e.g., a gRNA binding domain or a DNA-cleavage domain), such that the fragment is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to the corresponding fragment of a reference Cas9, e.g., a wild type Cas9. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the amino acid length of a corresponding wild type Cas9.
[0153] In some embodiments, a Cas9 fragment is a functional fragment that retains one or more Cas9 activities. In some embodiments, the Cas9 fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or at least 1300 amino acids in length.
[0154] In some embodiments, a prime editor comprises a Cas protein, e.g., a Cas9 variant, comprising modifications that allow altered PAM recognition. Exemplary Cas9 protein amino acid sequence (e.g., Cas9 variant with altered PAM recognition specificities) that are useful in the Prime editors of the disclosure are provided in Table 2. In some embodiments, a prime editor comprises a Cas protein, e.g., Cas9, containing modifications that allow altered PAM recognition. In prime editing using a Cas-protein-based prime editor, a "protospacer adjacent motif (PAM)”, PAM sequence, or PAM -like motif, may be used to refer to a short DNA sequence immediately following the protospacer sequence on the PAM strand of the target gene. In some embodiments, the PAM is recognized by the Cas nuclease in the prime editor during prime editing. In certain embodiments, the PAM is required for target binding of the Cas protein. The specific PAM sequence required for Cas protein recognition may depend on the specific type of the Cas protein. A PAM can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides in length. In some embodiments, a PAM is between 2-6 nucleotides in length. In some embodiments, the PAM can be a 5' PAM (i.e., located upstream of the 5' end of the protospacer). In other embodiments, the PAM can be a 3' PAM (i.e., located downstream of the 5' end of the protospacer). In some embodiments, the Cas protein of a prime editor recognizes a canonical PAM, for example, a SpCas9 recognizes 5'-NGG-3' PAM. In some embodiments, the Cas protein of a prime editor has altered or non-canonical PAM specificities. Exemplary PAM sequences and corresponding Cas variants are described in Table 3 below. It should be appreciated that for each of the variants provided, the Cas protein comprises one or more of the amino acid substitutions as indicated compared to a wild type Cas protein sequence, for example, the Cas9 as set forth in SEQ ID NO: 505. The PAM motifs as shown in Table 3 below are in the order of 5’ to 3 ’. In some embodiments, the Cas proteins of the disclosure can also be used to direct transcriptional control of target sequences, for example silencing transcription by sequence-specific binding to target sequences. In some embodiments, a Cas protein described herein may have one or mutations in a PAM recognition motif. In some embodiments, a Cas protein described herein may have altered PAM specificity.WSGR Docket No. 59761-805.601
[0155] As used in PAM sequences in Table 3, “N” refers to any one of nucleotides A, G, C, and T, “R” refers to nucleotide A or G, “Y” refers to nucleotide C or T, and “H” refers to any one of nucleotides A, C, or T.
[0156] Table 3: Cas protein variants and corresponding PAM sequencesVariant PAMspCas9 (wild type) NGG, NGA, NAG, NGNGA spCas9- VRVRFRR R1335V, LI 111R DI 135V, G1218R NGE1219F, A1322R, T1337RspCas9-VQR (DI 135V, R1335Q, T1337R ) NGAspCas9-EQR (DI 135E, R1335Q, T1337R) NGAspCas9-VRER (DI 135V, G1218R R1335E, T1337R) NGCGspCas9-VRQR (DI 135V, G1218R R1335Q, T1337R) NGASpCas9-VRQR(H840A) (R221K, N394K) NGASpCas9-VRQR (H840A) nickase +D1332K with R221K, NGAN394KVRQR-SpCas9(H840A) nickase +L1111R with R221K, NGAN394KCas9-NG (Lil HR DI 135V, G1218R, E1219F, A1322R NGNT1337R, R1335V)SpG Cas9 (D1135L, S1136W, G1218K, E1219Q, R1335Q, NGNT1337R)SyRY Cas9 NRN(A61R LI 111R, N1317R A1322R and R1333P)xCas9 (E480K, E543D, E1219V, K294R, Q1256K, A262T, NGNS409I, M694I)SluCas9 NNGGsRGNl, sRGN2, sRGN4, sRGN3.1, sRGN3.3 NNGGsaCas9 NNGRRT, NNGRRN saCas9-KKH (E782K, N968K, R1015H) NNNRRTspCas9-MQKSER (D1135M, S1136Q, G1218K, E1219S, NGCG / NGCN R1335E, T1337R)spCas9-LRKIQK (D1135L, S1136R, G1218K, E 12191, NGTN R1335Q, T1337K)spCas9-LRVSQK (D1135L, S1136R, G1218V, E1219S, NGTN R1335Q, T1337K)spCas9-LRVSQL(D1135L, S1136R G1218V, E1219S, NGTN R1335Q, T1337L)Cpfl TTTVSpy-Mac NAAWSGR Docket No. 59761-805.601NmCas9 NNNNGATTStCas9 NNAGAAWTdCas9 NAAAACSpCas9-NRTH (I322V, S409I, E427G, R654L, R753G, NRTHR1114G, D1135N, D1180G, GI2I8S, EI2I9V, Q1221H,P1249S, E1253K, P1321S, D1332G, R1335L)iSpyMac Cas9 NAAN SpCas9(R221K, I322V, N394K, S409I, E427G, R654L, NRRHR753G, H840A, Rl 114G, DI I35N, VI I39A, DI 180G,E1219V, Q1221H, A1320V, R1333K)SpCas9 (R221K, I322V, N394K, S409I, E427G, R654L, NRRHR753G, H840A, Rl 114G, DI 135N, V 1139A, DI 180G,E1219V, Q1221H, A1320V, R1333K, T1337R)
[0157] In some embodiments, a prime editor comprises a Cas9 polypeptide comprising one or more mutations selected from the group consisting of: A61R, LI 11R, DI 135V, R221K, A262T, I322V, R324L, N394K, S409I, S409I, E427G, E480K, M495V, N497A, Y515N, K526E, F539S, E543D, R654L, R661A, R661L, R691A, N692A, M694A, M694I, Q695A, H698A, R753G, M763I, K848A, K522N, Q926A, K1003A, R1060A, Lil HR R1114G, D1135E, D1135L, D1135N, S1136W, V1139A, D1180G, G1218K, G1218R, G1218S, E1219Q, E1219V, E1219V, Q1221H, P1249S, E1253K, N1317R, A1320V, P1321S, A1322R I1322V, D1332G, R1332N, A1332R, R1333K, R1333P, R1335L, R1335Q, R1335V, T1337N, T1337R S1338T, H1349R, and any combinations thereof as compared to a wildtype SpCas9 polypeptide as set forth in SEQ ID NO: 505.
[0158] In some embodiments, a prime editor comprises a SaCas9 polypeptide. In some embodiments, the SaCas9 polypeptide comprises one or more of mutations E782K, N968K, and R1015H as compared to a wild type SaCas9. In some embodiments, a prime editor comprises a FnCas9 polypeptide, for example, a wildtype FnCas9 polypeptide or a FnCas9 polypeptide comprising one or more of mutations E1369R E1449H, or R1556A as compared to the wild type FnCas9. In some embodiments, a prime editor comprises a Sc Cas9, for example, a wild type ScCas9 or a ScCas9 polypeptide comprises one or more of mutations I367K, G368D, I369K, H371L, T375S, T376G, and T1227K as compared to the wild type ScCas9. In some embodiments, a prime editor comprises a Stl Cas9 polypeptide, a St3 Cas9 polypeptide, or a SluCas9 polypeptide.
[0159] In some embodiments, a prime editor comprises a Cas polypeptide that comprises a circular permutant Cas variant. For example, a Cas9 polypeptide of a prime editor may be engineered such that the N-terminus and the C-terminus of a Cas9 protein (e.g., a wild type Cas9 protein, or a Cas9 nickase) are topically rearranged to retain the ability to bind DNA when complexed with a guide RNA (gRNA). An exemplary circular permutant configuration may be N-terminus-[original C-terminus]-[original N-terminus]-C-terminus. Any of the Cas9 proteins described herein, including any variant, ortholog, or naturally occurring Cas9 or equivalent thereof, may be reconfigured as a circular permutant variant.WSGR Docket No. 59761-805.601
[0160] In various embodiments, the circular permutants of a Cas protein, e.g., a Cas9, may have the following structure: N-terminus- [original C-terminus]-[optional linker]-[original N-terminus]-C-terminus. In some embodiments, a circular permutant Cas9 comprises any one of the following structures (amino acid positions as set forth in SEQ ID NO: 505):
[0161] N-terminus-[1268-1368]-[optional linker]-[l-1267]-C-terminus;
[0162] N-terminus-[l 168-1368]-[optional linker]-[l-l 167]-C-terminus;
[0163] N-terminus-[1068-1368]-[optional linker]-[l-1067]-C-terminus;
[0164] N-terminus-[968-1368]-[optional linker]-[l-967]-C-terminus;
[0165] N-terminus-[868-1368]-[optional linker]-[l-867]-C-terminus;
[0166] N-terminus-[768-1368]-[optional linker]-[l-767]-C-terminus;
[0167] N-terminus-[668-1368]-[optional linker]-[l-667]-C-terminus;
[0168] N-terminus-[568-1368]-[optional linker]-[l-567]-C-terminus;
[0169] N-terminus-[468-1368]-[optional linker]-[l-467]-C-terminus;
[0170] N-terminus-[368-1368]-[optional linker]-[l-367]-C-terminus;
[0171] N-terminus-[268-1368]-[optional linker]-[l-267]-C-terminus;
[0172] N-terminus-[168-1368]-[optional linker]-[l-167]-C-terminus;
[0173] N-terminus-[68-1368]-[optional linker]-[l-67]-C-terminus;
[0174] N-terminus-[10-1368]-[optional linker]-[l-9]-C -terminus, or the corresponding circular permutants of other Cas9 proteins (including other Cas9 orthologs, variants, etc.).
[0175] In some embodiments, a circular permutant Cas9 comprises any one of the following structures (amino acid positions as set forth in SEQ ID NO: 505):
[0176] N-terminus-[102-1368]-[optional linker]-[l-101]-C-terminus;
[0177] N-terminus-[1028-1368]-[optional linker]-[l-1027]-C-terminus;
[0178] N-terminus-[1041-1368]-[optional linker]-[l-1043]-C -terminus;
[0179] N-terminus-[1249-1368]-[optional linker]-[l-1248]-C -terminus; or
[0180] N-terminus-[1300-1368]-[optional linker]-[l-1299]-C-terminus, orthe corresponding circular permutants of other Cas9 proteins (including other Cas9 orthologs, variants, etc.).
[0181] In some embodiments, a circular permutant Cas9 comprises any one of the following structures (amino acid positions as set forth in SEQ ID NO: 505):amino acids of UniProtKB - Q99ZW2 N-terminus-[103-1368]-[optional linker]-[l-102]-C-terminus:
[0182] N-terminus-[1029-1368]-[optional linker]-[l-1028]-C-terminus;
[0183] N-terminus-[1042-1368]-[optional linker]-[ 1-104 l]-C-terminus;
[0184] N-terminus-[1250-1368]-[optional linker]-[l-1249]-C -terminus; or
[0185] N-terminus-[1301-1368]-[optional linker]-[l-1300]-C -terminus, orthe corresponding circular permutants of other Cas9 proteins (including other Cas9 orthologs, variants, etc.).
[0186] In some embodiments, the circular permutant can be formed by linking a C-terminal fragment of a Cas9 to an N-terminal fragment of a Cas9, either directly or by using a linker, such as an amino acid linker. In some embodiments, thee C-terminal fragment may correspond to the 95% or more of the C-WSGR Docket No. 59761-805.601terminal amino acids of a Cas9 (e.g., amino acids about 1300-1368 as set forth in SEQ ID No: 505 or corresponding amino acid positions thereof), or the 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% or more of the C-terminal amino acids of a Cas9 (e.g., SEQ ID NO 505 or a ortholog or a variant thereof). The N-terminal portion may correspond to 95% or more of the N-terminal amino acids of a Cas9 (e.g., amino acids about 1-1300 as set forth in SEQ ID NO: 505 or corresponding amino acid positions thereof), or 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, or 5% or more of the N terminal amino acids of a Cas9 (e.g., as set forth in SEQ ID NO: 505 or corresponding amino acid positions thereof).
[0187] In some embodiments, the circular permutant can be formed by linking a C-terminal fragment of a Cas9 to an N-terminal fragment of a Cas9, either directly or by using a linker, such as an amino acid linker. In some embodiments, the C-terminal fragment that is rearranged to the N-terminus includes or corresponds to the C-terminal 30% or less of the amino acids of a Cas9 (e.g., amino acids 1012-1368 as set forth in SEQ ID NO: 505 or corresponding amino acid positions thereof). In some embodiments, the C-terminal fragment that is rearranged to the N-terminus, includes or corresponds to the C-terminal 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% of the amino acids of a Cas9 (e.g., as set forth in SEQ ID NO: 505 or corresponding amino acid positions thereof). In some embodiments, the C-terminal fragment that is rearranged to the N-terminus, includes or corresponds to the C-terminal 410 residues or less of a Cas9 (e.g., as set forth in SEQ ID No: 505 or corresponding amino acid positions thereof). In some embodiments, the C-terminal portion that is rearranged to the N-terminus, includes or corresponds to the C-terminal 410, 400, 390, 380, 370, 360, 350, 340, 330, 320, 310, 300, 290, 280, 270, 260, 250, 240, 230, 220, 210, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 30, 20, or 10 residues of a Cas9 ( e.g., as set forth in SEQ ID NO: 505 or corresponding amino acid positions thereof). In some embodiments, the C-terminal portion that is rearranged to the N-terminus includes or corresponds to the C-terminal 357, 341, 328, 120, or 69 residues of a Cas9 (e.g., as set forth in SEQ ID NO: 505 or corresponding amino acid positions thereof).
[0188] In other embodiments, circular permutant Cas9 variants may be a topological rearrangement of a Cas9 primary structure based on the following method, which is based on. S', pyogenes Cas9 of SEQ ID NO: 505: (a) selecting a circular permutant (CP) site corresponding to an internal amino acid residue of the Cas9 primary structure, which dissects the original protein into two halves: an N-terminal region and a C-terminal region; (b) modifying the Cas9 protein sequence (e.g., by genetic engineering techniques) by moving the original C-terminal region (comprising the CP site amino acid) to precede the original N-terminal region, thereby forming a new N-terminus of the Cas9 protein that now begins with the CP site amino acid residue. The CP site can be located in any domain of the Cas9 protein, including, for example, the helical-II domain, the RuvCIII domain, or the CTD domain. For example, the CP site may be located (as set forth in SEQ ID NO: 505 or corresponding amino acid positions thereof) at original amino acid residue 181, 199, 230, 270, 310, 1010, 1016, 1023, 1029, 1041, 1247, 1249, or 1282. Thus, once relocated to the N-terminus, original amino acid 181, 199, 230, 270, 310, 1010, 1016, 1023, 1029, 1041, 1247,WSGR Docket No. 59761-805.6011249, or 1282 would become the new N-terminal amino acid. Nomenclature of these CP-Cas9 proteins may be referred to as Cas9-CP181, Cas9-CP199, Cas9-CP230, Cas9-CP270, Cas9-CP310, Cas9-CP1010, Cas9-CP1016, Cas9-CP1023, Cas9-CP1029, Cas9-CP1041, Cas9-CP1247, Cas9-CP1249, and Cas9-CP1282, respectively. This description is not meant to be limited to making CP variants from SEQ ID NO: 505, but may be implemented to make CP variants in any Cas9 sequence, either at CP sites that correspond to these positions, or at other CP sites entirely. This description is not meant to limit the specific CP sites in any way. Virtually any CP site may be used to form a CP-Cas9 variant.
[0189] In some embodiments, a prime editor comprises a Cas9 functional variant that is of smaller molecular weight than a wild type SpCas9 protein. In some embodiments, a smaller-sized Cas9 functional variant may facilitate delivery to cells, e.g., by an expression vector, nanoparticle, or other means of delivery. In certain embodiments, a smaller-sized Cas9 functional variant is a Class 2 Type II Cas protein. In certain embodiments, a smaller-sized Cas9 functional variant is a Class 2 Type V Cas protein. In certain embodiments, a smaller-sized Cas9 functional variant is a Class 2 Type VI Cas protein.
[0190] In some embodiments, a prime editor comprises a SpCas9 that is 1368 amino acids in length and has a predicted molecular weight of 158 kilodaltons. In some embodiments, a prime editor comprises a Cas9 functional variant or functional fragment that is less than 1300 amino acids, less than 1290 amino acids, than less than 1280 amino acids, less than 1270 amino acids, less than 1260 amino acid, less than 1250 amino acids, less than 1240 amino acids, less than 1230 amino acids, less than 1220 amino acids, less than 1210 amino acids, less than 1200 amino acids, less than 1190 amino acids, less than 1180 amino acids, less than 1170 amino acids, less than 1160 amino acids, less than 1150 amino acids, less than 1140 amino acids, less than 1130 amino acids, less than 1120 amino acids, less than 1110 amino acids, less than 1100 amino acids, less than 1050 amino acids, less than 1000 amino acids, less than 950 amino acids, less than 900 amino acids, less than 850 amino acids, less than 800 amino acids, less than 750 amino acids, less than 700 amino acids, less than 650 amino acids, less than 600 amino acids, less than 550 amino acids, or less than 500 amino acids, but at least larger than about 400 amino acids and retaining the one or more functions, e.g., DNA binding function, of the Cas9 protein.
[0191] In some embodiments, the Cas protein may include any CRISPR associated protein, including but not limited to, Casl2a, Casl2bl, Casl, CaslB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csxl2), CaslO, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4, homologs thereof, or modified versions thereof, and preferably comprising a nickase mutation (e.g., a mutation corresponding to the D10A mutation of the wild type Cas9 polypeptide of SEQ ID NO: 505). In various other embodiments, the napDNAbp can be any of the following proteins: a Cas9, a Cas 12a (Cpfl), a Casl2e (CasX), a Casl 2d (CasY), a Casl2bl (C2cl), a Casl3a (C2c2), a Casl2c (C2c3), a GeoCas9, a CjCas9, a Casl2g, a Casl2h, a Casl2i, a Casl3b, a Casl3c, a Casl3d, a Casl4, a Csn2, an xCas9, an SpCas9-NG, a circularly permuted Cas9, or an Argonaute (Ago) domain, or a functional variant or fragment thereof.
[0192] Exemplary Cas proteins and nomenclature are shown in Table 4 below:WSGR Docket No. 59761-805.601Table 4: Exemplary Cas proteins and nomenclatureLegacy nomenclature Current nomenclaturetype II CRISPR-Cas enzymesCas9 sametype V CRISPR-Cas enzymesCpfl Cas 12aCasX Casl2eC2cl Casl2blCasl2b2 sameC2c3 Cas 12cCasY Cas 12dC2c4 sameC2c8 sameC2c5 sameC2cl0 sameC2c9 sametype VI CRISPR-Cas enzymesC2c2 Cas 13aCas 13d sameC2c7 Cas 13cC2c6 Cas 13b
[0193] In some embodiments, prime editors described herein may also comprise Cas proteins other than Cas9. For example, in some embodiments, a prime editor as described herein may comprise a Cas 12a (Cpfl) polypeptide or functional variants thereof. In some embodiments, the Cas 12a polypeptide comprises a mutation that reduces or abolishes the endonuclease domain of the Cas 12a polypeptide. In some embodiments, the Cas 12a polypeptide is a Cas 12a nickase. In some embodiments, the Cas protein comprises an amino acid sequence that comprises at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a naturally occurring Casl2a polypeptide.
[0194] In some embodiments, a prime editor comprises a Cas protein that is a Cas 12b (C2cl) or a Cas 12c (C2c3) polypeptide. In some embodiments, the Cas protein comprises an amino acid sequence that comprises at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a naturally occurring Casl2b (C2cl) or Casl2c (C2c3) protein. In some embodiments, the Cas protein is a Casl2b nickase or a Casl2c nickase. In some embodiments, the Cas protein is a Casl2e, a Casl2d, a Cas 13, Cas 14a, Cas 14b, Cas 14c, Casl4d, Casl4e, Casl4f, Cas 14g, Casl4h, Casl4u, or a Cas<b polypeptide. In some embodiments, the Cas protein comprises an amino acid sequence that comprises at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a naturally-occurring Casl2e, Casl2d, Casl3, Casl4a, Casl4b, Casl4c, Casl4d, Casl4e, Casl4f, Casl4g, Casl4h, Casl4u, or Cas Oprotein. In some embodiments, the Cas protein is a Casl2e, Casl2d, Casl3, or Cas O nickase.
[0195] In some embodiments, a prime editor further comprises additional polypeptide components, for example, a flap endonuclease (FEN), e.g., FEN1. In some embodiments, the flap endonuclease excises the 5 ’ single-stranded DNA of the edit strand of the target gene and assists incorporation of the intendedWSGR Docket No. 59761-805.601nucleotide edit into the target gene. In some embodiments, the FEN is linked or fused to another component. In some embodiments, the FEN is provided in trans, for example, as a separate polypeptide or polynucleotide encoding the FEN.
[0196] In some embodiments, a prime editor or prime editing composition comprises a flap nuclease. In some embodiments, the flap nuclease is a FEN1, or any FEN1 functional variant, functional mutant, or functional fragment thereof. In some embodiments, the flap nuclease is a TREX2, EXO1, or any other flap nuclease known in the art, or any functional variant, functional mutant, or functional fragment thereof. In some embodiments, the flap nuclease has an amino acid sequence that is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to any of the flap nucleases described herein or known in the art.Nuclear Localization Sequences
[0197] In some embodiments, a prime editor further comprises one or more nuclear localization sequence (NLS). In some embodiments, the NLS helps promote translocation of a protein into the cell nucleus. In some embodiments, a prime editor comprises a fusion protein, e.g., a fusion protein comprising a DNA binding domain and a DNA polymerase, that comprises one or more NLSs. In some embodiments, one or more polypeptides of the prime editor are fused to or linked to one or more NLSs. In some embodiments, the prime editor comprises a DNA binding domain and a DNA polymerase domain that are provided in trans, wherein the DNA binding domain and / or the DNA polymerase domain is fused or linked to one or more NLSs.
[0198] In certain embodiments, a prime editor or prime editing complex comprises at least one NLS. In some embodiments, a prime editor or prime editing complex comprises at least two NLSs. In some embodiments, a prime editor or prime editing complex comprises at least three NLSs. In some embodiments, a prime editor or prime editing complex comprises more than 4, 5, 6, 7, 8, 9 or 10 NLSs. In embodiments with at least two NLSs, the NLSs can be the same NLS, or they can be different NLSs. In some embodiments, the one or more NLSs of a prime editor comprise bipartite NLSs.
[0199] NLSs can be expressed as part of a prime editor complex. In some embodiments, a NLS can be positioned almost anywhere in a protein's amino acid sequence, and generally comprises a short sequence of three or more or four or more amino acids. The location of the NLS fusion can be at the N-terminus, the C-terminus, or positioned anywhere within a sequence of a prime editor or a component thereof (e.g., inserted between the DNA-binding domain and the DNA polymerase domain of a prime editor fusion protein, between the DNA binding domain and a linker sequence, between a DNA polymerase and a linker sequence, between two linker sequences of a prime editor fusion protein or a component thereof, in either N-terminus to C-terminus or C-terminus to N-terminus order). In some embodiments, a prime editor is fusion protein that comprises an NLS at the N terminus. In some embodiments, a prime editor is fusion protein that comprises an NLS at the C terminus. In some embodiments, a prime editor is fusionWSGR Docket No. 59761-805.601protein that comprises at least one NLS at both the N terminus and the C terminus. In some embodiments, the prime editor is a fusion protein that comprises two NLSs at the N terminus and / or the C terminus.
[0200] Any NLSs that are known in the art are also contemplated herein. The NLSs may be any naturally occurring NLS, or any non-naturally occurring NLS (e.g., an NLS with one or more mutations relative to a wild-type NLS). In some embodiments, a nuclear localization signal (NLS) is predominantly basic. In some embodiments, the one or more NLSs of a prime editor are rich in lysine and arginine residues. In some embodiments, the one or more NLSs of a prime editor comprise proline residues.
[0201] In some embodiments, the 3 ’ motif of a PEgRNA or ngRNA can comprise a nuclear localization signal (NLS). Without wishing to be bound by any particular theory, motifs including such NLS sequence may assist trafficking and import of the PEgRNA to the nucleus, protect the PEgRNA from cytosolic RNases, promote endosomal escape, increase local PEgRNA concentration for association with the prime editor, and thereby improve editing efficiency. The NLS motif can be conjugated to the nucleotide components of the PEgRNA, e.g., at the 3’ end of the PBS, or at the 3’ end of a nucleotide 3’ motif, e.g., a UUUU motif, a hpl motif, or any other nucleotide motif described herein. In some embodiments, the PEgRNA or ngRNA comprises, or further comprises, a 3’ motif can comprise a SV40 NLS as set forth in SEQ ID NO: 549. The SV40 NLS may be conjugated to the 3’ end of the nucleotide portion of the PEgRNA or ngRNA, via a cysteine at the N terminus of the NLS (CKRTADGSEFESPKKKRKV, SEQ ID NO: 470), by any suitable moiety or linkage, for example, as described in PCT application W02023004409, incorporated herein by reference in its entirety. In some embodiments, the PEgRNA comprises amino acid SEQ ID NO: 470, which is conjugated to the 3’ most nucleotide of the PEgRNA via structure A. An exemplary structure of a PEgRNA comprising a NLS at its 3’ end via Structure A is shown in Fig. 4.sO N—(CH2)6-NH (Structure A)
[0202] As used herein, the 3’ NLS motif having the sequence of SEQ ID NO: 470 and conjugated to the 3 ’ end of a PEgRNA or ngRNA, the PEgRNA / ngRNA is connected to the NLS via conjugation as shown in Structure A, with the maleimide end connected to the 3 ’ most nucleotide of the PEgRNA or ngRNA and the thiol linkage to the NLS sequence as set forth in SEQ ID NO: 470.
[0203] Non-limiting examples of NLS sequences suitable for use with methods and compositions of the disclosure are provided in Table 5.
[0204] In some embodiments, a NLS is a monopartite NLS. For example, in some embodiments, a NLS is a SV40 large T antigen NLS comprising the sequence SEQ ID NO: 536. In some embodiments, a NLS is a bipartite NLS. In some embodiments, a bipartite NLS comprises two basic domains separated by a spacer sequence comprising a variable number of amino acids. In some embodiments, a NLS is a bipartite NLS. In some embodiments, a bipartite NLS consists of two basic domains separated by a spacer sequence comprising a variable number of amino acids. In some embodiments, the spacer amino acidWSGR Docket No. 59761-805.601sequence comprises aXenopus nucleoplasmin NLS SEQ ID NO: 554, wherein X is any amino acid. In some embodiments, the NLS comprises a nucleoplasmin NLS sequence SEQ ID NO: 553. In some embodiments, a NLS is a noncanonical sequences such as M9 of the hnRNP Al protein, the influenza virus nucleoprotein NLS, and the yeast Gal4 protein NLS.
[0205] In some embodiments, a NLS comprises an amino acid sequence that is at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence provided in Table 5. In some embodiments, a NLS comprises an amino acid sequence selected from the group consisting of the amino acid sequences provided in Table 5. In some embodiments, a prime editing composition comprises a polynucleotide that encodes a NLS that comprises an amino acid sequence that is at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence provided in Table 5. In some embodiments, a prime editing composition comprises a polynucleotide that encodes a NLS that comprises an amino acid sequence provided in Table 5.
[0206] Table 5: Exemplary nuclear localization sequencesDescription Sequence SEQ ID NO:SV40 NLS CKRTADGSEEESPKKKRKV 470 NLS of SV40 Large T- PKKKRKV536 AG NLS MKRTADGSELESPKKKRKV 537 NLS MDSLLMNRRKLLYQPKNVRWAKGRRETYLC 538 NLS of Nucleoplasmin AVKRPAATKKAGQAKKKKLD 539 NLS ofEGL-13 MSRRRKANPTKLSENAKKLAKEVEN 540 NLS ofC-Myc PAAKRVKLD 541 NLS of Tus-protein KLKIKRPVK 542 NLS of polyoma large VSRKRPRP543 T-AGNLS of Hepatitis D EGAPPAKRAR544 virus antigenNLS of Rev protein RQARRNRRRRWRERNR 545 NLS of murine p53 PPQPKKKPLDGE 546 C terminal linker and SGGSKRTADGSELEPKKKRKVNLS of an exemplary547 prime editor fusionproteinNLS KRTADGSELEPKKKRKV 548 NLS KRTADGSELESPKKKRKV 549 NLS NLSKRPAAIKKAGQAKKKK 550 NLS RQRRNELKRSL 551 NLS NQSSNLGPMKGGNLGGRSSGPYGGGGQYPAKPRNQGGY 552 nucleoplasmin NLS KRPAATKKAGQAKKKK553 sequenceXenopus KRXXXXXXXXXXKKKL, wherein X is any amino acid554 nucleoplasmin NLSNLS SGGSKRTADGSELESPKKKRKV 559NLS GSGPAAKRVKLD 560WSGR Docket No. 59761-805.601Description Sequence SEQ ID NO:Monopartite SV40 PKKKRKVGSG624 NLSN-terminal MPKKKRKVGSGMonopartite SV40 625NLS
[0207] Components of a prime editor may be connected to each other in any order. In some embodiments, the DNA binding domain and the DNA polymerase domain of a prime editor may be fused to form a fusion protein, or may be joined by a peptide or protein linker, in any order from the N terminus to the C terminus. In some embodiments, a prime editor comprises a DNA binding domain fused or linked to the C-terminal end of a DNA polymerase domain. In some embodiments, a prime editor comprises a DNA binding domain fused or linked to the N-terminal end of a DNA polymerase domain. In some embodiments, the prime editor comprises a fusion protein comprising the structure NH2-[DNA binding domain]-[polymerase]-COOH; or NH2-[polymerase]-[DNA binding domain]-COOH, wherein each instance ofindicates the presence of an optional linker sequence. In some embodiments, a prime editor comprises a fusion protein and a DNA polymerase domain provided in trans, wherein the fusion protein comprises the structure NH2-[DNA binding domain]-[RNA-protein recruitment polypeptide]-COOH. In some embodiments, a prime editor comprises a fusion protein and a DNA binding domain provided in trans, wherein the fusion protein comprises the structure NH2-[DNA polymerase domain]-[RNA-protein recruitment polypeptide]-COOH.
[0208] In some embodiments, a prime editor fusion protein, a polypeptide component of a prime editor, or a polynucleotide encoding the prime editor fusion protein or polypeptide component, may be split into an N-terminal half and a C-terminal half or polypeptides that encode the N-terminal half and the C terminal half, and provided to a target DNA in a cell separately. For example, in certain embodiments, a prime editor fusion protein may be split into a N-terminal and a C-terminal half for separate delivery in AAV vectors, and subsequently translated and colocalized in a target cell to reform the complete polypeptide or prime editor protein. In such cases, separate halves of a protein or a fusion protein may each comprise a split-intein to facilitate colocalization and reformation of the complete protein or fusion protein by the mechanism of intein facilitated trans splicing. In some embodiments, a prime editor comprises a N-terminal half fused to an intein-N, and a C-terminal half fused to an intein-C, or polynucleotides or vectors (e.g., AAV vectors) encoding each thereof. When delivered and / or expressed in a target cell, the intein-N and the intein-C can be excised via protein trans-splicing, resulting in a complete prime editor fusion protein in the target cell. In some embodiments, an exemplary protein described herein may lack a methionine residue at the N-terminus.WSGR Docket No. 59761-805.601
[0209] Table 98: Prime editor amino acid sequencesPrimeeditoraminoSequenceacidSEQ ID NO:MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGN TDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKA DLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGV DAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAK LQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSA SMVKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFY KFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDFY PFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGA SAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLS GEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDL LKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKR LRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQ KAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGGHKPENIVIEMA RENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNG RDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEE VVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHH AHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYF FYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNI562VKKTEVQTGGFSKESILPKGNSDKLIARKKDWDPKKYGGFNSPTVAYSVLVVAK VEKGKSKKLKSVKELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFE LENGRKRMLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNKQKQLFV EQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTN LGASAAFKYFDTTIGRKLYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSG GSSGGSSGGSSGGSSGGSSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSDFPQ AWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGIL VPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSH QWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTL FNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYR ASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKA GF CRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQ ALLTAP ALGLPDLT KPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIA VLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRV QFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADHTWYTDG SSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLN VYTDSRYAFATAHIHGEIYRRRGWLTSEGKEIKNKDEILALLKALFLPKRLSIIHCP GHQKGHSAEARGNRMADQAARKAAITETPDTSTLLIENSSPSGGSKRTADGSEFEP KKKRKV (SEQ ID NO: 562) MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGN TDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKA563 DLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGV DAKAILSARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDA KLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFWSGR Docket No. 59761-805.601 YKFIKPILEKMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDF YPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKG ASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFL SGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHD LLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLK RLRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDI QKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGGHKPENIVIEM ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQN GRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSE EVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHH AHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYF FYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNI VKKTEVQTGGFSKESILPKGNSDKLIARKKDWDPKKYGGFNSPTVAYSVLVVAK VEKGKSKKLKSVKELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFE LENGRKRMLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNKQKQLFV EQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTN LGASAAFKYFDTTIGRKLYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSG GSKRTADGSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSD FPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQ GILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPNVPNPYNLLSGLP PSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNS PTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAATSELDCQQGTRALLQTLGNLG YRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLG KAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLP DLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMV AAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDT DRVQFGPVVALNPATLLPLPEEGLQHNCLDSGGSKRTADGSEFESPKKKRKVGSG PAAKRVKLD (SEQ ID NO: 563) MPKKKRKVGSGDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKK NLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRL EESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLA LAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILS ARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKD TYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRY DEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILE KMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDFYPFLKDNR EKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIER MTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAI VDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDK DFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRLRYTGW GRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSG626QGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGGHKPENIVIEMARENQTT QKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYV DQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKM KNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQI LDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYL NAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMN FFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEV QTGGFSKESILPKGNSDKLIARKKDWDPKKYGGFNSPTVAYSVLVVAKVEKGKSK KLKSVKELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFELENGRKR MLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNKQKQLFVEQHKHYL DEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGASAAF KYFDTTIGRKLYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAEWSGR Docket No. 59761-805.601 TGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQS PWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPNVPNPYNLLSGLPPSHQWYT VLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFCEAL HRDLADFRIQHPDLILLQYYDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKK AQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRL FIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFEL FVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTK DAGKLTMGQPLVIGAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGP VVALNPATLLPLPEEGLQHNCLDSGGSKRTADGSEFESPKKKRKVGSGPAAKRVK LD (SEQ ID NO: 626) MPKKKRKVGSGDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKK NLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRL EESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLA LAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILS ARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKD TYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRY DEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILE KMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDFYPFLKDNR EKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIER MTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAI VDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDK DFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRLRYTGW GRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSG QGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGGHKPENIVIEMARENQTT QKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYV DQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKM KNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQI LDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYL642 NAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMN FFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEV QTGGFSKESILPKGNSDKLIARKKDWDPKKYGGFNSPTVAYSVLVVAKVEKGKSK KLKSVKELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFELENGRKR MLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNKQKQLFVEQHKHYL DEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGASAAF KYFDTTIGRKLYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSKRTAD GSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAE TGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQS PWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHQNVPNPYNLLSGLPPSHQWY TVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFCEA LHRDLADFRIQHPDLILLQYYDDLLLAATSELDCQQGTRALLQTLGNLGYRASAK KAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCR LFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFE LFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLT KDAGKLTMGQPLVIGAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFG PVVALNPATLLPLPEEGLQHNCLDSGGSKRTADGSEFESPKKKRKVGSGPAAKRV KLD (SEQ ID NO: 642) MPKKKRKVGSGDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKK NLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRL EESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLA LAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILS643 ARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKD TYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRY DEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDFYPFLKDNRWSGR Docket No. 59761-805.601 EKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIER MTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAI VDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDK DFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRLRYTGW GRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSG QGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGGHKPENIVIEMARENQTT QKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYV DQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKM KNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQI LDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYL NAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMN FFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEV QTGGFSKESILPKGNSDKLIARKKDWDPKKYGGFNSPTVAYSVLVVAKVEKGKSK KLKSVKELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFELENGRKR MLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNKQKQLFVEQHKHYL DEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGASAAF KYFDTTIGRKLYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSKRTAD GSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAE TGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQS PWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPNVPNPYNLLSGLPPSHQWYT VLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFCEAL HRDLADFRIQHPDLILLQYYDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKK AQICQKQVKYLGYLLKEGQRWLTEARKENVMGQPTPKTPRQLRVFLGKAGFCRL FIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFEL FVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTK DAGKLTMGQPLVIGAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGP VVALNPATLLPLPEEGLQHNCLDSGGSKRTADGSEFESPKKKRKVGSGPAAKRVK LD (SEQ ID NO: 643) MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGN TDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKA DLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGV DAKAILSARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDA KLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLS ASMVKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEF YKFIKPILEKMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDF YPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKG ASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFL SGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHD LLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLK RLRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDI QKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGGHKPENIVIEM657ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQN GRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSE EVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHH AHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYF FYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNI VKKTEVQTGGFSKESILPKGNSDKLIARKKDWDPKKYGGFNSPTVAYSVLVVAK VEKGKSKKLKSVKELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFE LENGRKRMLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNKQKQLFV EQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTN LGASAAFKYFDTTIGRKLYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSG GSKRTADGSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSD FPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQWSGR Docket No. 59761-805.601 PSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNS PTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAATSELDCQQGTRALLQTLGNLG YRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLG KAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLP DLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMV AAIAVLTKDAGKLTMGQPLVIGAPHAVEALVKQPPDRWLSNARMTHYQALLLDT DRVQFGPVVALNPATLLPLPEEGLQHNCLDSGGSKRTADGSEFESPKKKRKVGSG PAAKRVKLD (SEQ ID NO: 657) MPKKKRKVGSGDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKK NLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRL EESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLA LAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILS ARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKD IYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRY DEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILE KMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDFYPFLKDNR EKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIER MTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAI VDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDK DFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRLRYTGW GRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSG QGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGGHKPENIVIEMARENQTT QKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYV DQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKM KNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQI LDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYL658 NAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMN FFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEV QTGGFSKESILPKGNSDKLIARKKDWDPKKYGGFNSPTVAYSVLVVAKVEKGKSK KLKSVKELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFELENGRKR MLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNKQKQLFVEQHKHYL DEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGASAAF KYFDTTIGRKLYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSKRTAD GSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAE TGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQS PWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPNVPNPYNLLSGLPPSHQWYT VLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFCEAL HRDLADFRIQHPDLILLQYYDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKK AQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRL FIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFEL FVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTK DAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPV VALNPATLLPLPEEGLQHNCLDSGGSKRTADGSEFESPKKKRKVGSGPAAKRVKLD (SEQ ID NO: 658) MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGN TDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKA DLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGV DAKAILSARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDA1056 KLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLS ASMVKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEF YKFIKPILEKMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDF YPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLWSGR Docket No. 59761-805.601 SGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHD LLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLK RLRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDI QKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGGHKPENIVIEM ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQN GRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSE EVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHH AHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYF FYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNI VKKTEVQTGGFSKESILPKGNSDKLIARKKDWDPKKYGGFNSPTVAYSVLVVAK VEKGKSKKLKSVKELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFE LENGRKRMLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNKQKQLFV EQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTN LGASAAFKYFDTTIGRKLYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSG GSKRTADGSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSD FPQAWAETWGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQ GILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPNVPNPYNLLSGLP PSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNS PTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAATSELDCQQGTRALLQTLGNLG YRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLG KAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLP DLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMV AAIAVLTKDAGKLTMGQPLVIGAPHAVEALVKQPPDRWLSNARMTHYQALLLDT DRVQFGPVVALNPATLLPLPEEGLQHNCLDSGGSKRTADGSEFESPKKKRKVGSG PAAKRVKLD (SEQ ID NO: 1056) MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGN TDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKA DLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGV DAKAILSARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDA KLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLS ASMVKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEF YKFIKPILEKMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDF YPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKG ASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFL SGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHD LLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLK RLRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDI QKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGGHKPENIVIEM ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQN1058 GRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSE EVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHH AHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYF FYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNI VKKTEVQTGGFSKESILPKGNSDKLIARKKDWDPKKYGGFNSPTVAYSVLVVAK VEKGKSKKLKSVKELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFE LENGRKRMLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNKQKQLFV EQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTN LGASAAFKYFDTTIGRKLYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSG GSKRTADGSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSD FPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPNVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAATSELDCQQGTRALLQTLGNLGWSGR Docket No. 59761-805.601 YRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLG KAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLP DLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMV AAIAVLTKDAGKLTMGQPLVIGAPHAVEALVKQPPDRWLSNARMTHYQALLLDT DRVQFGPVVALNPATLLPLPEEGLQHNCLDSGGSKRTADGSEFESPKKKRKVGSG PAAKRVKLD (SEQ ID NO: 1058) MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGN TDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKA DLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGV DAKAILSARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDA KLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLS ASMVKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEF YKFIKPILEKMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDF YPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKG ASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFL SGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHD LLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLK RLRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDI QKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGGHKPENIVIEM ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQN GRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSE EVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHH1059 AHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYF FYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNI VKKTEVQTGGFSKESILPKGNSDKLIARKKDWDPKKYGGFNSPTVAYSVLVVAK VEKGKSKKLKSVKELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFE LENGRKRMLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNKQKQLFV EQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTN LGASAAFKYFDTTIGRKLYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSG GSKRTADGSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSD FPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQ GILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPNVPNCYNLLSGL PPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKN SPTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAATSELDCQQGTRALLQTLGNL GYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFL GKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQ ALLTAP ALGL PDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMV AAIAVLTKDAGKLTMGQPLVIGAPHAVEALVKQPPDRWLSNARMTHYQALLLDT DRVQFGPVVALNPATLLPLPEEGLQHNCLDSGGSKRTADGSEFESPKKKRKVGSG PAAKRVKLD (SEQ ID NO: 1059) MPKKKRKVGSGDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKK NLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRL EESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLA LAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILS ARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKD TYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRY DEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILE1060 KMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDFYPFLKDNR EKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIER MTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAI VDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDK DFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRLRYTGW GRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGGHKPENIVIEMARENQTTWSGR Docket No. 59761-805.601 QKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYV DQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKM KNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQI LDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYL NAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMN FFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEV QTGGFSKESILPKGNSDKLIARKKDWDPKKYGGFNSPTVAYSVLVVAKVEKGKSK KLKSVKELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFELENGRKR MLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNKQKQLFVEQHKHYL DEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGASAAF KYFDTTIGRKLYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSKRTAD GSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAE TGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQS PWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPNVPNCYNLLSGLPPSHQWYT VLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFCEAL HRDLADFRIQHPDLILLQYYDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKK AQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRL FIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFEL FVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTK DAGKLTMGQPLVIGAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGP VVALNPATLLPLPEEGLQHNCLDSGGSKRTADGSEFESPKKKRKVGSGPAAKRVK LD (SEQ ID NO: 1060) MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGN TDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKA DLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGV DAKAILSARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDA KLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLS ASMVKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEF YKFIKPILEKMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDF YPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKG ASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFL SGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHD LLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLK RLRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDI QKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGGHKPENIVIEM ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQN GRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSE EVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT1061 KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHH AHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYF FYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNI VKKTEVQTGGFSKESILPKGNSDKLIARKKDWDPKKYGGFNSPTVAYSVLVVAK VEKGKSKKLKSVKELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFE LENGRKRMLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNKQKQLFV EQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTN LGASAAFKYFDTTIGRKLYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSG GSKRTADGSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSD FPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQ GILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHQNVPNPYNLLSGL PPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKN SPTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAATSELDCQQGTRALLQTLGNL GYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFL GKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQ ALLTAP ALGL PDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVWSGR Docket No. 59761-805.601 DRVQFGPVVALNPATLLPLPEEGLQHNCLDSGGSKRTADGSEFESPKKKRKVGSG PAAKRVKLD (SEQ ID NO: 1061) MPKKKRKVGSGDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKK NLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRL EESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLA LAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILS ARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKD TYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRY DEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILE KMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDFYPFLKDNR EKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIER MTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAI VDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDK DFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRLRYTGW GRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSG QGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGGHKPENIVIEMARENQTT QKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYV DQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKM KNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQI LDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYL1063 NAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMN FFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEV QTGGFSKESILPKGNSDKLIARKKDWDPKKYGGFNSPTVAYSVLVVAKVEKGKSK KLKSVKELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFELENGRKR MLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNKQKQLFVEQHKHYL DEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGASAAF KYFDTTIGRKLYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSKRTAD GSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAE TWGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQS PWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPNVPNPYNLLSGLPPSHQWYT VLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFCEAL HRDLADFRIQHPDLILLQYYDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKK AQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRL FIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFEL FVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTK DAGKLTMGQPLVIGAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGP VVALNPATLLPLPEEGLQHNCLDSGGSKRTADGSEFESPKKKRKVGSGPAAKRVK LD (SEQ ID NO: 1063) MPKKKRKVGSGDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKK NLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRL EESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLA LAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILS ARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKD TYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMVKRY DEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILE KMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDFYPFLKDNR EKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIER1064 MTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAI VDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDK DFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRLRYTGW GRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSG QGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGGHKPENIVIEMARENQTT QKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYV DQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKM KNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLWSGR Docket No. 59761-805.601 NAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMN FFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEV QTGGFSKESILPKGNSDKLIARKKDWDPKKYGGFNSPTVAYSVLVVAKVEKGKSK KLKSVKELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFELENGRKR MLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNKQKQLFVEQHKHYL DEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGASAAF KYFDTTIGRKLYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSGGSKRTAD GSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAE TGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQS PWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPNVPNPYNLLSGLPPSHQWYT VLDLKDAFFCLRLHPTSQPLFAFKWRDPEMGISGQLTWTRLPQGFKNSPTLFCEAL HRDLADFRIQHPDLILLQYYDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKK AQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCRL FIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFEL FVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTK DAGKLTMGQPLVIGAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGP VVALNPATLLPLPEEGLQHNCLDSGGSKRTADGSEFESPKKKRKVGSGPAAKRVK LD (SEQ IDNO: 1064) MKRTADGSEFESPKKKRKVDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGN TDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKV DDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKA DLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGV DAKAILSARLSKSRKLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDA KLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLS ASMVKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEF YKFIKPILEKMDGTEELLVKLKREDLLRKQRTFDNGIIPHQIHLGELHAILRRQGDF YPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKG ASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFL SGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHD LLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLK RLRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDI QKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGGHKPENIVIEM ARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQN GRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSE EVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQIT KHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHH1065 AHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYF FYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNI VKKTEVQTGGFSKESILPKGNSDKLIARKKDWDPKKYGGFNSPTVAYSVLVVAK VEKGKSKKLKSVKELLGITIMERSSFEKNPIGFLEAKGYKEVKKDLIIKLPKYSLFE LENGRKRMLASASVLHKGNELALPSKYVNFLYLASHYEKLKGSSEDNKQKQLFV EQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTN LGASAAFKYFDTTIGRKLYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDSGGSSG GSKRTADGSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLSD FPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQ GILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPNVPNPYNLLSGLP PSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNS PTLFCEALHRDLADFRIQHPDLILLQYYDDLLLAATSELDCQQGTRALLQTLGNLG YRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKENVMGQPTPKTPRQLRVFLG KAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQALLTAPALGLP DLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLSKKLDPVAAGWPPCLRMV AAIAVLTKDAGKLTMGQPLVIGAPHAVEALVKQPPDRWLSNARMTHYQALLLDT DRVQFGPVVALNPATLLPLPEEGLQHNCLDSGGSKRTADGSEFESPKKKRKVGSGPAAKRVKLD (SEQ ID NO: 1065)WSGR Docket No. 59761-805.601
[0210] Table 99: Prime editor nucleotide full-length sequencesPrime editornucleotide SequenceSEQ ID NO:AGGCCACCAUGAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAA GAAGAAGCGGAAGGUGGACAAGAAGUACAGCAUCGGCCUGGACAUCGG CACCAACAGCGUGGGCUGGGCCGUGAUCACCGACGAGUACAAGGUGCCC AGCAAGAAGUUCAAGGUGCUGGGCAACACCGACCGGCACAGCAUCAAG AAGAACCUGAUCGGCGCCCUGCUGUUCGACAGCGGCGAGACCGCCGAGG CCACCCGGCUGAAGCGGACCGCCCGGCGGCGGUACACCCGGCGGAAGAA CCGGAUCUGCUACCUGCAGGAGAUCUUCAGCAACGAGAUGGCCAAGGU GGACGACAGCUUCUUCCACCGGCUGGAGGAGAGCUUCCUGGUGGAGGA GGACAAGAAGCACGAGCGGCACCCCAUCUUCGGCAACAUCGUGGACGAG GUGGCCUACCACGAGAAGUACCCCACCAUCUACCACCUGCGGAAGAAGC UGGUGGACAGCACCGACAAGGCCGACCUGCGGCUGAUCUACCUGGCCCU GGCCCACAUGAUCAAGUUCCGGGGCCACUUCCUGAUCGAGGGCGACCUG AACCCCGACAACAGCGACGUGGACAAGCUGUUCAUCCAGCUGGUGCAGA CCUACAACCAGCUGUUCGAGGAGAACCCCAUCAACGCCAGCGGCGUGGA CGCCAAGGCCAUCCUGAGCGCCCGGCUGAGCAAGAGCCGGAAGCUGGAG AACCUGAUCGCCCAGCUGCCCGGCGAGAAGAAGAACGGCCUGUUCGGCA ACCUGAUCGCCCUGAGCCUGGGCCUGACCCCCAACUUCAAGAGCAACUU CGACCUGGCCGAGGACGCCAAGCUGCAGCUGAGCAAGGACACCUACGAC GACGACCUGGACAACCUGCUGGCCCAGAUCGGCGACCAGUACGCCGACC UGUUCCUGGCCGCCAAGAACCUGAGCGACGCCAUCCUGCUGAGCGACAU CCUGCGGGUGAACACCGAGAUCACCAAGGCCCCCCUGAGCGCCAGCAUG GUGAAGCGGUACGACGAGCACCACCAGGACCUGACCCUGCUGAAGGCCC UGGUGCGGCAGCAGCUGCCCGAGAAGUACAAGGAGAUCUUCUUCGACC AGAGCAAGAACGGCUACGCCGGCUACAUCGACGGCGGCGCCAGCCAGGA GGAGUUCUACAAGUUCAUCAAGCCCAUCCUGGAGAAGAUGGACGGCAC CGAGGAGCUGCUGGUGAAGCUGAAGCGGGAGGACCUGCUGCGGAAGCA GCGGACCUUCGACAACGGCAUCAUCCCCCACCAGAUCCACCUGGGCGAG CUGCACGCCAUCCUGCGGCGGCAGGGCGACUUCUACCCCUUCCUGAAGG ACAACCGGGAGAAGAUCGAGAAGAUCCUGACCUUCCGGAUCCCCUACUA CGUGGGCCCCCUGGCCCGGGGCAACAGCCGGUUCGCCUGGAUGACCCGG AAGUCCGAGGAGACCAUCACCCCCUGGAACUUCGAGGAGGUGGUGGAC AAGGGCGCCAGCGCCCAGAGCUUCAUCGAGCGGAUGACCAACUUCGACA AGAACCUGCCCAACGAGAAGGUGCUGCCCAAGCACAGCCUGCUGUACGA GUACUUCACCGUGUACAACGAGCUGACCAAGGUGAAGUACGUGACCGA GGGCAUGCGGAAGCCCGCCUUCCUGAGCGGCGAGCAGAAGAAGGCCAUC GUGGACCUGCUGUUCAAGACCAACCGGAAGGUGACCGUGAAGCAGCUG AAGGAGGACUACUUCAAGAAGAUCGAGUGCUUCGACAGCGUGGAGAUC AGCGGCGUGGAGGACCGGUUCAACGCCAGCCUGGGCACCUACCACGACC UGCUGAAGAUCAUCAAGGACAAGGACUUCCUGGACAACGAGGAGAACG AGGACAUCCUGGAGGACAUCGUGCUGACCCUGACCCUGUUCGAGGACCG GGAGAUGAUCGAGGAGCGGCUCAAGACCUACGCCCACCUGUUCGACGAC AAGGUGAUGAAGCAGCUGAAGCGGCUGCGGUACACCGGCUGGGGCCGG CUGAGCCGGAAGCUGAUCAACGGCAUCCGGGACAAGCAGAGCGGCAAG ACCAUCCUGGACUUCCUCAAGAGCGACGGCUUCGCCAACCGGAACUUCA UGCAGCUGAUCCACGACGACAGCCUGACCUUCAAGGAGGACAUCCAGAA GGCCCAGGUGAGCGGCCAGGGCGACAGCCUGCACGAGCACAUCGCCAAC CUGGCCGGCAGCCCCGCCAUCAAGAAGGGCAUCCUGCAGACCGUGAAGG UGGUGGACGAGCUGGUGAAGGUGAUGGGCGGCCACAAGCCCGAGAACA UCGUGAUCGAGAUGGCCCGGGAGAACCAGACCACCCAGAAGGGCCAGA AGAACAGCCGGGAGCGGAUGAAGCGGAUCGAGGAGGGCAUCAAGGAGC UGGGCAGCCAGAUCCUGAAGGAGCACCCCGUGGAGAACACCCAGCUGCA655 GAACGAGAAGCUGUACCUGUACUACCUGCAGAACGGCCGGGACAUGUAWSGR Docket No. 59761-805.601 CGUGGACCAGGAGCUGGACAUCAACCGGCUGAGCGACUACGACGUGGA CGCCAUCGUGCCCCAGAGCUUCCUGAAGGACGACAGCAUCGACAACAAG GUGCUGACCCGGAGCGACAAGAACCGGGGCAAGAGCGACAACGUGCCCA GCGAGGAGGUGGUGAAGAAGAUGAAGAACUACUGGCGGCAGCUGCUGA ACGCCAAGCUGAUCACCCAGCGGAAGUUCGACAACCUGACCAAGGCCGA GCGGGGCGGCCUGAGCGAGCUGGACAAGGCCGGCUUCAUCAAGCGGCA GCUGGUGGAGACCCGGCAGAUCACCAAGCACGUGGCCCAGAUCCUGGAC AGCCGGAUGAACACCAAGUACGACGAGAACGACAAGCUGAUCCGGGAG GUGAAGGUGAUCACCCUCAAGAGCAAGCUGGUGAGCGACUUCCGGAAG GACUUCCAGUUCUACAAGGUGCGGGAGAUCAACAACUACCACCACGCCC ACGACGCCUACCUGAACGCCGUGGUGGGCACCGCCCUGAUCAAGAAGUA CCCCAAGCUGGAGAGCGAGUUCGUGUACGGCGACUACAAGGUGUACGA CGUGCGGAAGAUGAUCGCCAAGAGCGAGCAGGAGAUCGGCAAGGCCAC CGCCAAGUACUUCUUCUACAGCAACAUCAUGAACUUCUUCAAGACCGAG AUCACCCUGGCCAACGGCGAGAUCCGGAAGCGGCCCCUGAUCGAGACCA ACGGCGAGACCGGCGAGAUCGUGUGGGACAAGGGCCGGGACUUCGCCA CCGUGCGGAAGGUGCUGAGCAUGCCCCAGGUGAACAUCGUGAAGAAAA CCGAGGUGCAGACCGGCGGCUUCAGCAAGGAGAGCAUCCUGCCCAAGGG CAACAGCGACAAGCUGAUCGCCCGGAAGAAGGACUGGGACCCCAAGAA GUACGGCGGCUUCAACAGCCCCACCGUGGCCUACAGCGUGCUGGUGGUG GCCAAGGUGGAGAAGGGCAAGAGCAAGAAGCUCAAGAGCGUGAAGGAG CUGCUGGGCAUCACCAUCAUGGAGCGGAGCAGCUUCGAGAAGAACCCCA UCGGCUUCCUGGAGGCCAAGGGCUACAAGGAGGUGAAGAAGGACCUGA UCAUCAAGCUGCCCAAGUACAGCCUGUUCGAGCUGGAGAACGGCCGGA AGCGGAUGCUGGCCAGCGCCUCCGUGCUGCACAAGGGCAACGAGCUGGC CCUGCCCAGCAAGUACGUGAACUUCCUGUACCUGGCCAGCCACUACGAG AAGCUGAAGGGCAGCUCCGAGGACAACAAGCAGAAGCAGCUGUUCGUG GAGCAGCACAAGCACUACCUGGACGAGAUCAUCGAGCAGAUCAGCGAG UUCAGCAAGCGGGUGAUCCUGGCCGACGCCAACCUGGACAAGGUGCUG AGCGCCUACAACAAGCACCGGGACAAGCCCAUCCGGGAGCAGGCCGAGA ACAUCAUCCACCUGUUCACCCUGACCAACCUGGGCGCCAGCGCCGCCUU CAAGUACUUCGACACCACCAUCGGCCGGAAGCUGUACACCAGCACCAAG GAGGUGCUGGACGCCACCCUGAUCCACCAGAGCAUCACCGGCCUGUACG AGACCCGGAUCGACCUGAGCCAGCUGGGCGGCGACAGCGGCGGCAGCAG CGGCGGCAGCAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAAG AAGAAGCGGAAGGUGAGCGGCGGCAGCAGCGGCGGCAGCACCCUGAAC AUCGAGGACGAGUACCGGCUGCACGAGACCAGCAAGGAGCCCGACGUG AGCCUGGGCAGCACCUGGCUGAGCGACUUCCCCCAGGCCUGGGCCGAGA CCGGCGGCAUGGGCCUGGCCGUGCGGCAGGCCCCCCUGAUCAUCCCCCU GAAGGCCACCAGCACCCCCGUGAGCAUCAAGCAGUACCCCAUGAGCCAG GAGGCCCGGCUGGGCAUCAAGCCCCACAUCCAGCGGCUGCUGGACCAGG GCAUCCUGGUGCCCUGCCAGAGCCCCUGGAACACCCCCCUGCUGCCCGU GAAGAAGCCCGGCACCAACGACUACCGGCCCGUGCAGGACCUGCGGGAG GUGAACAAGCGGGUGGAGGACAUCCACCCCAACGUGCCCAACCCCUACA ACCUGCUGAGCGGCCUGCCCCCCAGCCACCAGUGGUACACCGUGCUGGA CCUGAAGGACGCCUUCUUCUGCCUGCGGCUGCACCCCACCAGCCAGCCC CUGUUCGCCUUCGAGUGGCGGGACCCCGAGAUGGGCAUCAGCGGCCAGC UGACCUGGACCCGGCUGCCCCAGGGCUUCAAGAACAGCCCCACCCUGUU CUGCGAGGCCCUGCACCGGGACCUGGCCGACUUCCGGAUCCAGCACCCC GACCUGAUCCUGCUGCAGUACUACGACGACCUGCUGCUGGCCGCCACCA GCGAGCUGGACUGCCAGCAGGGCACCCGGGCCCUGCUGCAGACCCUGGG CAACCUGGGCUACCGGGCCAGCGCCAAGAAGGCCCAGAUCUGCCAGAAG CAGGUGAAGUACCUGGGCUACCUGCUGAAGGAGGGCCAGCGGUGGCUG ACCGAGGCCCGGAAGGAGACCGUGAUGGGCCAGCCCACCCCCAAGACCC CCCGGCAGCUGCGGGAGUUCCUGGGCAAGGCCGGCUUCUGCCGGCUGUUCAUCCCCGGCUUCGCCGAGAUGGCCGCCCCCCUGUACCCCCUGACCAAGWSGR Docket No. 59761-805.601 CCCGGCACCCUGUUCAACUGGGGCCCCGACCAGCAGAAGGCCUACCAGG AGAUCAAGCAGGCCCUGCUGACCGCCCCCGCCCUGGGCCUGCCCGACCU GACCAAGCCCUUCGAGCUGUUCGUGGACGAGAAGCAGGGCUACGCCAA GGGCGUGCUGACCCAGAAGCUGGGCCCCUGGCGGCGGCCCGUGGCCUAC CUGAGCAAGAAGCUGGACCCCGUGGCCGCCGGCUGGCCCCCCUGCCUGC GGAUGGUGGCCGCCAUCGCCGUGCUGACCAAGGACGCCGGCAAGCUGAC CAUGGGCCAGCCCCUGGUGAUCCUGGCCCCCCACGCCGUGGAGGCCCUG GUGAAGCAGCCCCCCGACCGGUGGCUGAGCAACGCCCGGAUGACCCACU ACCAGGCCCUGCUGCUGGACACCGACCGGGUGCAGUUCGGCCCCGUGGU GGCCCUGAACCCCGCCACCCUGCUGCCCCUGCCCGAGGAGGGCCUGCAG CACAACUGCCUGGACAGCGGCGGCAGCAAGCGGACCGCCGACGGCAGCG AGUUCGAGAGCCCCAAGAAGAAGCGGAAGGUGGGCAGCGGCCCCGCCGC CAAGCGGGUGAAGCUGGACUGAUAGUGAGCGGCCGCUUAAUUAAGCUG CCUUCUGCGGGGCUUGCCUUCUGGCCAAGCCCUUCUUCUCUCCCUUGCA CCUGUACCUCUUGGUCUUUGAAUAAAGCCUGAGUAGGAAGAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA AAAAAAAAAAAAAAAAAAA (SEQ ID NO: 655) AGGCCACCAUGAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAA GAAGAAGCGGAAGGUGGACAAGAAGUACAGCAUCGGCCUGGACAUCGG CACCAACAGCGUGGGCUGGGCCGUGAUCACCGACGAGUACAAGGUGCCC AGCAAGAAGUUCAAGGUGCUGGGCAACACCGACCGGCACAGCAUCAAG AAGAACCUGAUCGGCGCCCUGCUGUUCGACAGCGGCGAGACCGCCGAGG CCACCCGGCUGAAGCGGACCGCCCGGCGGCGGUACACCCGGCGGAAGAA CCGGAUCUGCUACCUGCAGGAGAUCUUCAGCAACGAGAUGGCCAAGGU GGACGACAGCUUCUUCCACCGGCUGGAGGAGAGCUUCCUGGUGGAGGA GGACAAGAAGCACGAGCGGCACCCCAUCUUCGGCAACAUCGUGGACGAG GUGGCCUACCACGAGAAGUACCCCACCAUCUACCACCUGCGGAAGAAGC UGGUGGACAGCACCGACAAGGCCGACCUGCGGCUGAUCUACCUGGCCCU GGCCCACAUGAUCAAGUUCCGGGGCCACUUCCUGAUCGAGGGCGACCUG AACCCCGACAACAGCGACGUGGACAAGCUGUUCAUCCAGCUGGUGCAGA CCUACAACCAGCUGUUCGAGGAGAACCCCAUCAACGCCAGCGGCGUGGA CGCCAAGGCCAUCCUGAGCGCCCGGCUGAGCAAGAGCCGGAAGCUGGAG AACCUGAUCGCCCAGCUGCCCGGCGAGAAGAAGAACGGCCUGUUCGGCA ACCUGAUCGCCCUGAGCCUGGGCCUGACCCCCAACUUCAAGAGCAACUU CGACCUGGCCGAGGACGCCAAGCUGCAGCUGAGCAAGGACACCUACGAC GACGACCUGGACAACCUGCUGGCCCAGAUCGGCGACCAGUACGCCGACC UGUUCCUGGCCGCCAAGAACCUGAGCGACGCCAUCCUGCUGAGCGACAU CCUGCGGGUGAACACCGAGAUCACCAAGGCCCCCCUGAGCGCCAGCAUG GUGAAGCGGUACGACGAGCACCACCAGGACCUGACCCUGCUGAAGGCCC UGGUGCGGCAGCAGCUGCCCGAGAAGUACAAGGAGAUCUUCUUCGACC AGAGCAAGAACGGCUACGCCGGCUACAUCGACGGCGGCGCCAGCCAGGA GGAGUUCUACAAGUUCAUCAAGCCCAUCCUGGAGAAGAUGGACGGCAC CGAGGAGCUGCUGGUGAAGCUGAAGCGGGAGGACCUGCUGCGGAAGCA GCGGACCUUCGACAACGGCAUCAUCCCCCACCAGAUCCACCUGGGCGAG CUGCACGCCAUCCUGCGGCGGCAGGGCGACUUCUACCCCUUCCUGAAGG ACAACCGGGAGAAGAUCGAGAAGAUCCUGACCUUCCGGAUCCCCUACUA CGUGGGCCCCCUGGCCCGGGGCAACAGCCGGUUCGCCUGGAUGACCCGG AAGUCCGAGGAGACCAUCACCCCCUGGAACUUCGAGGAGGUGGUGGAC AAGGGCGCCAGCGCCCAGAGCUUCAUCGAGCGGAUGACCAACUUCGACA AGAACCUGCCCAACGAGAAGGUGCUGCCCAAGCACAGCCUGCUGUACGA GUACUUCACCGUGUACAACGAGCUGACCAAGGUGAAGUACGUGACCGA GGGCAUGCGGAAGCCCGCCUUCCUGAGCGGCGAGCAGAAGAAGGCCAUC GUGGACCUGCUGUUCAAGACCAACCGGAAGGUGACCGUGAAGCAGCUG654 AAGGAGGACUACUUCAAGAAGAUCGAGUGCUUCGACAGCGUGGAGAUCWSGR Docket No. 59761-805.601 AGCGGCGUGGAGGACCGGUUCAACGCCAGCCUGGGCACCUACCACGACC UGCUGAAGAUCAUCAAGGACAAGGACUUCCUGGACAACGAGGAGAACG AGGACAUCCUGGAGGACAUCGUGCUGACCCUGACCCUGUUCGAGGACCG GGAGAUGAUCGAGGAGCGGCUCAAGACCUACGCCCACCUGUUCGACGAC AAGGUGAUGAAGCAGCUGAAGCGGCUGCGGUACACCGGCUGGGGCCGG CUGAGCCGGAAGCUGAUCAACGGCAUCCGGGACAAGCAGAGCGGCAAG ACCAUCCUGGACUUCCUCAAGAGCGACGGCUUCGCCAACCGGAACUUCA UGCAGCUGAUCCACGACGACAGCCUGACCUUCAAGGAGGACAUCCAGAA GGCCCAGGUGAGCGGCCAGGGCGACAGCCUGCACGAGCACAUCGCCAAC CUGGCCGGCAGCCCCGCCAUCAAGAAGGGCAUCCUGCAGACCGUGAAGG UGGUGGACGAGCUGGUGAAGGUGAUGGGCGGCCACAAGCCCGAGAACA UCGUGAUCGAGAUGGCCCGGGAGAACCAGACCACCCAGAAGGGCCAGA AGAACAGCCGGGAGCGGAUGAAGCGGAUCGAGGAGGGCAUCAAGGAGC UGGGCAGCCAGAUCCUGAAGGAGCACCCCGUGGAGAACACCCAGCUGCA GAACGAGAAGCUGUACCUGUACUACCUGCAGAACGGCCGGGACAUGUA CGUGGACCAGGAGCUGGACAUCAACCGGCUGAGCGACUACGACGUGGA CGCCAUCGUGCCCCAGAGCUUCCUGAAGGACGACAGCAUCGACAACAAG GUGCUGACCCGGAGCGACAAGAACCGGGGCAAGAGCGACAACGUGCCCA GCGAGGAGGUGGUGAAGAAGAUGAAGAACUACUGGCGGCAGCUGCUGA ACGCCAAGCUGAUCACCCAGCGGAAGUUCGACAACCUGACCAAGGCCGA GCGGGGCGGCCUGAGCGAGCUGGACAAGGCCGGCUUCAUCAAGCGGCA GCUGGUGGAGACCCGGCAGAUCACCAAGCACGUGGCCCAGAUCCUGGAC AGCCGGAUGAACACCAAGUACGACGAGAACGACAAGCUGAUCCGGGAG GUGAAGGUGAUCACCCUCAAGAGCAAGCUGGUGAGCGACUUCCGGAAG GACUUCCAGUUCUACAAGGUGCGGGAGAUCAACAACUACCACCACGCCC ACGACGCCUACCUGAACGCCGUGGUGGGCACCGCCCUGAUCAAGAAGUA CCCCAAGCUGGAGAGCGAGUUCGUGUACGGCGACUACAAGGUGUACGA CGUGCGGAAGAUGAUCGCCAAGAGCGAGCAGGAGAUCGGCAAGGCCAC CGCCAAGUACUUCUUCUACAGCAACAUCAUGAACUUCUUCAAGACCGAG AUCACCCUGGCCAACGGCGAGAUCCGGAAGCGGCCCCUGAUCGAGACCA ACGGCGAGACCGGCGAGAUCGUGUGGGACAAGGGCCGGGACUUCGCCA CCGUGCGGAAGGUGCUGAGCAUGCCCCAGGUGAACAUCGUGAAGAAAA CCGAGGUGCAGACCGGCGGCUUCAGCAAGGAGAGCAUCCUGCCCAAGGG CAACAGCGACAAGCUGAUCGCCCGGAAGAAGGACUGGGACCCCAAGAA GUACGGCGGCUUCAACAGCCCCACCGUGGCCUACAGCGUGCUGGUGGUG GCCAAGGUGGAGAAGGGCAAGAGCAAGAAGCUCAAGAGCGUGAAGGAG CUGCUGGGCAUCACCAUCAUGGAGCGGAGCAGCUUCGAGAAGAACCCCA UCGGCUUCCUGGAGGCCAAGGGCUACAAGGAGGUGAAGAAGGACCUGA UCAUCAAGCUGCCCAAGUACAGCCUGUUCGAGCUGGAGAACGGCCGGA AGCGGAUGCUGGCCAGCGCCUCCGUGCUGCACAAGGGCAACGAGCUGGC CCUGCCCAGCAAGUACGUGAACUUCCUGUACCUGGCCAGCCACUACGAG AAGCUGAAGGGCAGCUCCGAGGACAACAAGCAGAAGCAGCUGUUCGUG GAGCAGCACAAGCACUACCUGGACGAGAUCAUCGAGCAGAUCAGCGAG UUCAGCAAGCGGGUGAUCCUGGCCGACGCCAACCUGGACAAGGUGCUG AGCGCCUACAACAAGCACCGGGACAAGCCCAUCCGGGAGCAGGCCGAGA ACAUCAUCCACCUGUUCACCCUGACCAACCUGGGCGCCAGCGCCGCCUU CAAGUACUUCGACACCACCAUCGGCCGGAAGCUGUACACCAGCACCAAG GAGGUGCUGGACGCCACCCUGAUCCACCAGAGCAUCACCGGCCUGUACG AGACCCGGAUCGACCUGAGCCAGCUGGGCGGCGACAGCGGCGGCAGCAG CGGCGGCAGCAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAAG AAGAAGCGGAAGGUGAGCGGCGGCAGCAGCGGCGGCAGCACCCUGAAC AUCGAGGACGAGUACCGGCUGCACGAGACCAGCAAGGAGCCCGACGUG AGCCUGGGCAGCACCUGGCUGAGCGACUUCCCCCAGGCCUGGGCCGAGA CCGGCGGCAUGGGCCUGGCCGUGCGGCAGGCCCCCCUGAUCAUCCCCCU GAAGGCCACCAGCACCCCCGUGAGCAUCAAGCAGUACCCCAUGAGCCAGGAGGCCCGGCUGGGCAUCAAGCCCCACAUCCAGCGGCUGCUGGACCAGGWSGR Docket No. 59761-805.601 GCAUCCUGGUGCCCUGCCAGAGCCCCUGGAACACCCCCCUGCUGCCCGU GAAGAAGCCCGGCACCAACGACUACCGGCCCGUGCAGGACCUGCGGGAG GUGAACAAGCGGGUGGAGGACAUCCACCCCAACGUGCCCAACCCCUACA ACCUGCUGAGCGGCCUGCCCCCCAGCCACCAGUGGUACACCGUGCUGGA CCUGAAGGACGCCUUCUUCUGCCUGCGGCUGCACCCCACCAGCCAGCCC CUGUUCGCCUUCGAGUGGCGGGACCCCGAGAUGGGCAUCAGCGGCCAGC UGACCUGGACCCGGCUGCCCCAGGGCUUCAAGAACAGCCCCACCCUGUU CUGCGAGGCCCUGCACCGGGACCUGGCCGACUUCCGGAUCCAGCACCCC GACCUGAUCCUGCUGCAGUACUACGACGACCUGCUGCUGGCCGCCACCA GCGAGCUGGACUGCCAGCAGGGCACCCGGGCCCUGCUGCAGACCCUGGG CAACCUGGGCUACCGGGCCAGCGCCAAGAAGGCCCAGAUCUGCCAGAAG CAGGUGAAGUACCUGGGCUACCUGCUGAAGGAGGGCCAGCGGUGGCUG ACCGAGGCCCGGAAGGAGACCGUGAUGGGCCAGCCCACCCCCAAGACCC CCCGGCAGCUGCGGGAGUUCCUGGGCAAGGCCGGCUUCUGCCGGCUGUU CAUCCCCGGCUUCGCCGAGAUGGCCGCCCCCCUGUACCCCCUGACCAAG CCCGGCACCCUGUUCAACUGGGGCCCCGACCAGCAGAAGGCCUACCAGG AGAUCAAGCAGGCCCUGCUGACCGCCCCCGCCCUGGGCCUGCCCGACCU GACCAAGCCCUUCGAGCUGUUCGUGGACGAGAAGCAGGGCUACGCCAA GGGCGUGCUGACCCAGAAGCUGGGCCCCUGGCGGCGGCCCGUGGCCUAC CUGAGCAAGAAGCUGGACCCCGUGGCCGCCGGCUGGCCCCCCUGCCUGC GGAUGGUGGCCGCCAUCGCCGUGCUGACCAAGGACGCCGGCAAGCUGAC CAUGGGCCAGCCCCUGGUGAUCCUGGCCCCCCACGCCGUGGAGGCCCUG GUGAAGCAGCCCCCCGACCGGUGGCUGAGCAACGCCCGGAUGACCCACU ACCAGGCCCUGCUGCUGGACACCGACCGGGUGCAGUUCGGCCCCGUGGU GGCCCUGAACCCCGCCACCCUGCUGCCCCUGCCCGAGGAGGGCCUGCAG CACAACUGCCUGGACAGCGGCGGCAGCAAGCGGACCGCCGACGGCAGCG AGUUCGAGAGCCCCAAGAAGAAGCGGAAGGUGGGCAGCGGCCCCGCCGC CAAGCGGGUGAAGCUGGACUGAUAGUGAGCGGCCGCUUAAUUAAGCUG CCUUCUGCGGGGCUUGCCUUCUGGCCAAGCCCUUCUUCUCUCCCUUGCA CCUGUACCUCUUGGUCUUUGAAUAAAGCCUGAGUAGGAAGAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA AAAAA (SEQ ID NO: 654) AGGCCACCAUGAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAA GAAGAAGCGGAAGGUGGACAAGAAGUACAGCAUCGGCCUGGACAUCGG CACCAACAGCGUGGGCUGGGCCGUGAUCACCGACGAGUACAAGGUGCCC AGCAAGAAGUUCAAGGUGCUGGGCAACACCGACCGGCACAGCAUCAAG AAGAACCUGAUCGGCGCCCUGCUGUUCGACAGCGGCGAGACCGCCGAGG CCACCCGGCUGAAGCGGACCGCCCGGCGGCGGUACACCCGGCGGAAGAA CCGGAUCUGCUACCUGCAGGAGAUCUUCAGCAACGAGAUGGCCAAGGU GGACGACAGCUUCUUCCACCGGCUGGAGGAGAGCUUCCUGGUGGAGGA GGACAAGAAGCACGAGCGGCACCCCAUCUUCGGCAACAUCGUGGACGAG GUGGCCUACCACGAGAAGUACCCCACCAUCUACCACCUGCGGAAGAAGC UGGUGGACAGCACCGACAAGGCCGACCUGCGGCUGAUCUACCUGGCCCU GGCCCACAUGAUCAAGUUCCGGGGCCACUUCCUGAUCGAGGGCGACCUG AACCCCGACAACAGCGACGUGGACAAGCUGUUCAUCCAGCUGGUGCAGA CCUACAACCAGCUGUUCGAGGAGAACCCCAUCAACGCCAGCGGCGUGGA CGCCAAGGCCAUCCUGAGCGCCCGGCUGAGCAAGAGCCGGAAGCUGGAG AACCUGAUCGCCCAGCUGCCCGGCGAGAAGAAGAACGGCCUGUUCGGCA ACCUGAUCGCCCUGAGCCUGGGCCUGACCCCCAACUUCAAGAGCAACUU CGACCUGGCCGAGGACGCCAAGCUGCAGCUGAGCAAGGACACCUACGAC GACGACCUGGACAACCUGCUGGCCCAGAUCGGCGACCAGUACGCCGACC UGUUCCUGGCCGCCAAGAACCUGAGCGACGCCAUCCUGCUGAGCGACAU CCUGCGGGUGAACACCGAGAUCACCAAGGCCCCCCUGAGCGCCAGCAUG GUGAAGCGGUACGACGAGCACCACCAGGACCUGACCCUGCUGAAGGCCC UGGUGCGGCAGCAGCUGCCCGAGAAGUACAAGGAGAUCUUCUUCGACC656 AGAGCAAGAACGGCUACGCCGGCUACAUCGACGGCGGCGCCAGCCAGGAWSGR Docket No. 59761-805.601 GGAGUUCUACAAGUUCAUCAAGCCCAUCCUGGAGAAGAUGGACGGCAC CGAGGAGCUGCUGGUGAAGCUGAAGCGGGAGGACCUGCUGCGGAAGCA GCGGACCUUCGACAACGGCAUCAUCCCCCACCAGAUCCACCUGGGCGAG CUGCACGCCAUCCUGCGGCGGCAGGGCGACUUCUACCCCUUCCUGAAGG ACAACCGGGAGAAGAUCGAGAAGAUCCUGACCUUCCGGAUCCCCUACUA CGUGGGCCCCCUGGCCCGGGGCAACAGCCGGUUCGCCUGGAUGACCCGG AAGUCCGAGGAGACCAUCACCCCCUGGAACUUCGAGGAGGUGGUGGAC AAGGGCGCCAGCGCCCAGAGCUUCAUCGAGCGGAUGACCAACUUCGACA AGAACCUGCCCAACGAGAAGGUGCUGCCCAAGCACAGCCUGCUGUACGA GUACUUCACCGUGUACAACGAGCUGACCAAGGUGAAGUACGUGACCGA GGGCAUGCGGAAGCCCGCCUUCCUGAGCGGCGAGCAGAAGAAGGCCAUC GUGGACCUGCUGUUCAAGACCAACCGGAAGGUGACCGUGAAGCAGCUG AAGGAGGACUACUUCAAGAAGAUCGAGUGCUUCGACAGCGUGGAGAUC AGCGGCGUGGAGGACCGGUUCAACGCCAGCCUGGGCACCUACCACGACC UGCUGAAGAUCAUCAAGGACAAGGACUUCCUGGACAACGAGGAGAACG AGGACAUCCUGGAGGACAUCGUGCUGACCCUGACCCUGUUCGAGGACCG GGAGAUGAUCGAGGAGCGGCUCAAGACCUACGCCCACCUGUUCGACGAC AAGGUGAUGAAGCAGCUGAAGCGGCUGCGGUACACCGGCUGGGGCCGG CUGAGCCGGAAGCUGAUCAACGGCAUCCGGGACAAGCAGAGCGGCAAG ACCAUCCUGGACUUCCUCAAGAGCGACGGCUUCGCCAACCGGAACUUCA UGCAGCUGAUCCACGACGACAGCCUGACCUUCAAGGAGGACAUCCAGAA GGCCCAGGUGAGCGGCCAGGGCGACAGCCUGCACGAGCACAUCGCCAAC CUGGCCGGCAGCCCCGCCAUCAAGAAGGGCAUCCUGCAGACCGUGAAGG UGGUGGACGAGCUGGUGAAGGUGAUGGGCGGCCACAAGCCCGAGAACA UCGUGAUCGAGAUGGCCCGGGAGAACCAGACCACCCAGAAGGGCCAGA AGAACAGCCGGGAGCGGAUGAAGCGGAUCGAGGAGGGCAUCAAGGAGC UGGGCAGCCAGAUCCUGAAGGAGCACCCCGUGGAGAACACCCAGCUGCA GAACGAGAAGCUGUACCUGUACUACCUGCAGAACGGCCGGGACAUGUA CGUGGACCAGGAGCUGGACAUCAACCGGCUGAGCGACUACGACGUGGA CGCCAUCGUGCCCCAGAGCUUCCUGAAGGACGACAGCAUCGACAACAAG GUGCUGACCCGGAGCGACAAGAACCGGGGCAAGAGCGACAACGUGCCCA GCGAGGAGGUGGUGAAGAAGAUGAAGAACUACUGGCGGCAGCUGCUGA ACGCCAAGCUGAUCACCCAGCGGAAGUUCGACAACCUGACCAAGGCCGA GCGGGGCGGCCUGAGCGAGCUGGACAAGGCCGGCUUCAUCAAGCGGCA GCUGGUGGAGACCCGGCAGAUCACCAAGCACGUGGCCCAGAUCCUGGAC AGCCGGAUGAACACCAAGUACGACGAGAACGACAAGCUGAUCCGGGAG GUGAAGGUGAUCACCCUCAAGAGCAAGCUGGUGAGCGACUUCCGGAAG GACUUCCAGUUCUACAAGGUGCGGGAGAUCAACAACUACCACCACGCCC ACGACGCCUACCUGAACGCCGUGGUGGGCACCGCCCUGAUCAAGAAGUA CCCCAAGCUGGAGAGCGAGUUCGUGUACGGCGACUACAAGGUGUACGA CGUGCGGAAGAUGAUCGCCAAGAGCGAGCAGGAGAUCGGCAAGGCCAC CGCCAAGUACUUCUUCUACAGCAACAUCAUGAACUUCUUCAAGACCGAG AUCACCCUGGCCAACGGCGAGAUCCGGAAGCGGCCCCUGAUCGAGACCA ACGGCGAGACCGGCGAGAUCGUGUGGGACAAGGGCCGGGACUUCGCCA CCGUGCGGAAGGUGCUGAGCAUGCCCCAGGUGAACAUCGUGAAGAAAA CCGAGGUGCAGACCGGCGGCUUCAGCAAGGAGAGCAUCCUGCCCAAGGG CAACAGCGACAAGCUGAUCGCCCGGAAGAAGGACUGGGACCCCAAGAA GUACGGCGGCUUCAACAGCCCCACCGUGGCCUACAGCGUGCUGGUGGUG GCCAAGGUGGAGAAGGGCAAGAGCAAGAAGCUCAAGAGCGUGAAGGAG CUGCUGGGCAUCACCAUCAUGGAGCGGAGCAGCUUCGAGAAGAACCCCA UCGGCUUCCUGGAGGCCAAGGGCUACAAGGAGGUGAAGAAGGACCUGA UCAUCAAGCUGCCCAAGUACAGCCUGUUCGAGCUGGAGAACGGCCGGA AGCGGAUGCUGGCCAGCGCCUCCGUGCUGCACAAGGGCAACGAGCUGGC CCUGCCCAGCAAGUACGUGAACUUCCUGUACCUGGCCAGCCACUACGAG AAGCUGAAGGGCAGCUCCGAGGACAACAAGCAGAAGCAGCUGUUCGUGGAGCAGCACAAGCACUACCUGGACGAGAUCAUCGAGCAGAUCAGCGAGWSGR Docket No. 59761-805.601 UUCAGCAAGCGGGUGAUCCUGGCCGACGCCAACCUGGACAAGGUGCUG AGCGCCUACAACAAGCACCGGGACAAGCCCAUCCGGGAGCAGGCCGAGA ACAUCAUCCACCUGUUCACCCUGACCAACCUGGGCGCCAGCGCCGCCUU CAAGUACUUCGACACCACCAUCGGCCGGAAGCUGUACACCAGCACCAAG GAGGUGCUGGACGCCACCCUGAUCCACCAGAGCAUCACCGGCCUGUACG AGACCCGGAUCGACCUGAGCCAGCUGGGCGGCGACAGCGGCGGCAGCAG CGGCGGCAGCAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAAG AAGAAGCGGAAGGUGAGCGGCGGCAGCAGCGGCGGCAGCACCCUGAAC AUCGAGGACGAGUACCGGCUGCACGAGACCAGCAAGGAGCCCGACGUG AGCCUGGGCAGCACCUGGCUGAGCGACUUCCCCCAGGCCUGGGCCGAGA CCGGCGGCAUGGGCCUGGCCGUGCGGCAGGCCCCCCUGAUCAUCCCCCU GAAGGCCACCAGCACCCCCGUGAGCAUCAAGCAGUACCCCAUGAGCCAG GAGGCCCGGCUGGGCAUCAAGCCCCACAUCCAGCGGCUGCUGGACCAGG GCAUCCUGGUGCCCUGCCAGAGCCCCUGGAACACCCCCCUGCUGCCCGU GAAGAAGCCCGGCACCAACGACUACCGGCCCGUGCAGGACCUGCGGGAG GUGAACAAGCGGGUGGAGGACAUCCACCCCAACGUGCCCAACCCCUACA ACCUGCUGAGCGGCCUGCCCCCCAGCCACCAGUGGUACACCGUGCUGGA CCUGAAGGACGCCUUCUUCUGCCUGCGGCUGCACCCCACCAGCCAGCCC CUGUUCGCCUUCGAGUGGCGGGACCCCGAGAUGGGCAUCAGCGGCCAGC UGACCUGGACCCGGCUGCCCCAGGGCUUCAAGAACAGCCCCACCCUGUU CUGCGAGGCCCUGCACCGGGACCUGGCCGACUUCCGGAUCCAGCACCCC GACCUGAUCCUGCUGCAGUACUACGACGACCUGCUGCUGGCCGCCACCA GCGAGCUGGACUGCCAGCAGGGCACCCGGGCCCUGCUGCAGACCCUGGG CAACCUGGGCUACCGGGCCAGCGCCAAGAAGGCCCAGAUCUGCCAGAAG CAGGUGAAGUACCUGGGCUACCUGCUGAAGGAGGGCCAGCGGUGGCUG ACCGAGGCCCGGAAGGAGACCGUGAUGGGCCAGCCCACCCCCAAGACCC CCCGGCAGCUGCGGGAGUUCCUGGGCAAGGCCGGCUUCUGCCGGCUGUU CAUCCCCGGCUUCGCCGAGAUGGCCGCCCCCCUGUACCCCCUGACCAAG CCCGGCACCCUGUUCAACUGGGGCCCCGACCAGCAGAAGGCCUACCAGG AGAUCAAGCAGGCCCUGCUGACCGCCCCCGCCCUGGGCCUGCCCGACCU GACCAAGCCCUUCGAGCUGUUCGUGGACGAGAAGCAGGGCUACGCCAA GGGCGUGCUGACCCAGAAGCUGGGCCCCUGGCGGCGGCCCGUGGCCUAC CUGAGCAAGAAGCUGGACCCCGUGGCCGCCGGCUGGCCCCCCUGCCUGC GGAUGGUGGCCGCCAUCGCCGUGCUGACCAAGGACGCCGGCAAGCUGAC CAUGGGCCAGCCCCUGGUGAUCGGCGCCCCCCACGCCGUGGAGGCCCUG GUGAAGCAGCCCCCCGACCGGUGGCUGAGCAACGCCCGGAUGACCCACU ACCAGGCCCUGCUGCUGGACACCGACCGGGUGCAGUUCGGCCCCGUGGU GGCCCUGAACCCCGCCACCCUGCUGCCCCUGCCCGAGGAGGGCCUGCAG CACAACUGCCUGGACAGCGGCGGCAGCAAGCGGACCGCCGACGGCAGCG AGUUCGAGAGCCCCAAGAAGAAGCGGAAGGUGGGCAGCGGCCCCGCCGC CAAGCGGGUGAAGCUGGACUGAUAGUGAGCGGCCGCUUAAUUAAGCUG CCUUCUGCGGGGCUUGCCUUCUGGCCAAGCCCUUCUUCUCUCCCUUGCA CCUGUACCUCUUGGUCUUUGAAUAAAGCCUGAGUAGGAAGAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA AAAAA (SEQ ID NO: 656) AGGCCACCAUGCCCAAGAAGAAGCGGAAGGUGGGCAGCGGCGACAAGA AGUACAGCAUCGGCCUGGACAUCGGCACCAACAGCGUGGGCUGGGCCGU GAUCACCGACGAGUACAAGGUGCCCAGCAAGAAGUUCAAGGUGCUGGG CAACACCGACCGGCACAGCAUCAAGAAGAACCUGAUCGGCGCCCUGCUG UUCGACAGCGGCGAGACCGCCGAGGCCACCCGGCUGAAGCGGACCGCCC GGCGGCGGUACACCCGGCGGAAGAACCGGAUCUGCUACCUGCAGGAGA UCUUCAGCAACGAGAUGGCCAAGGUGGACGACAGCUUCUUCCACCGGCU GGAGGAGAGCUUCCUGGUGGAGGAGGACAAGAAGCACGAGCGGCACCC CAUCUUCGGCAACAUCGUGGACGAGGUGGCCUACCACGAGAAGUACCCC659 ACCAUCUACCACCUGCGGAAGAAGCUGGUGGACAGCACCGACAAGGCCGWSGR Docket No. 59761-805.601 ACCUGCGGCUGAUCUACCUGGCCCUGGCCCACAUGAUCAAGUUCCGGGG CCACUUCCUGAUCGAGGGCGACCUGAACCCCGACAACAGCGACGUGGAC AAGCUGUUCAUCCAGCUGGUGCAGACCUACAACCAGCUGUUCGAGGAG AACCCCAUCAACGCCAGCGGCGUGGACGCCAAGGCCAUCCUGAGCGCCC GGCUGAGCAAGAGCCGGAAGCUGGAGAACCUGAUCGCCCAGCUGCCCGG CGAGAAGAAGAACGGCCUGUUCGGCAACCUGAUCGCCCUGAGCCUGGGC CUGACCCCCAACUUCAAGAGCAACUUCGACCUGGCCGAGGACGCCAAGC UGCAGCUGAGCAAGGACACCUACGACGACGACCUGGACAACCUGCUGGC CCAGAUCGGCGACCAGUACGCCGACCUGUUCCUGGCCGCCAAGAACCUG AGCGACGCCAUCCUGCUGAGCGACAUCCUGCGGGUGAACACCGAGAUCA CCAAGGCCCCCCUGAGCGCCAGCAUGGUGAAGCGGUACGACGAGCACCA CCAGGACCUGACCCUGCUGAAGGCCCUGGUGCGGCAGCAGCUGCCCGAG AAGUACAAGGAGAUCUUCUUCGACCAGAGCAAGAACGGCUACGCCGGC U AC A U CGACGGCGGCGCC AGCC AGGAGGAGLT U C UAC AAGUU C AU C AAG CCCAUCCUGGAGAAGAUGGACGGCACCGAGGAGCUGCUGGUGAAGCUG AAGCGGGAGGACCUGCUGCGGAAGCAGCGGACCUUCGACAACGGCAUC AUCCCCCACCAGAUCCACCUGGGCGAGCUGCACGCCAUCCUGCGGCGGC AGGGCGACUUCUACCCCUUCCUGAAGGACAACCGGGAGAAGAUCGAGA AGAUCCUGACCUUCCGGAUCCCCUACUACGUGGGCCCCCUGGCCCGGGG CAACAGCCGGUUCGCCUGGAUGACCCGGAAGUCCGAGGAGACCAUCACC CCCUGGAACUUCGAGGAGGUGGUGGACAAGGGCGCCAGCGCCCAGAGC UUCAUCGAGCGGAUGACCAACUUCGACAAGAACCUGCCCAACGAGAAG GUGCUGCCCAAGCACAGCCUGCUGUACGAGUACUUCACCGUGUACAACG AGCUGACCAAGGUGAAGUACGUGACCGAGGGCAUGCGGAAGCCCGCCU UCCUGAGCGGCGAGCAGAAGAAGGCCAU CGU GGACC UGC UGUUCAAGA CCAACCGGAAGGUGACCGUGAAGCAGCUGAAGGAGGACUACUUCAAGA AGAUCGAGUGCUUCGACAGCGUGGAGAUCAGCGGCGUGGAGGACCGGU UCAACGCCAGCCUGGGCACCUACCACGACCUGCUGAAGAUCAUCAAGGA CAAGGACUUCCUGGACAACGAGGAGAACGAGGACAUCCUGGAGGACAU CGUGCUGACCCUGACCCLTGUUCGAGGACCGGGAGAUGAUCGAGGAGCG GCUCAAGACCUACGCCCACCUGUUCGACGACAAGGUGAUGAAGCAGCUG AAGCGGCUGCGGUACACCGGCUGGGGCCGGCUGAGCCGGAAGCUGAUC AACGGCAUCCGGGACAAGCAGAGCGGCAAGACCAUCCUGGACUUCCUCA AGAGCGACGGCUUCGCCAACCGGAACUUCAUGCAGCUGAUCCACGACGA CAGCCUGACCUUCAAGGAGGACAUCCAGAAGGCCCAGGUGAGCGGCCAG GGCGACAGCCUGCACGAGCACAUCGCCAACCUGGCCGGCAGCCCCGCCA U C AAGAAGGGC AU CC UGC AGACCGUGAAGGUGGU GGACGAGC U GGU GA AGGUGAUGGGCGGCCACAAGCCCGAGAACAUCGUGAUCGAGAUGGCCC GGGAGAACCAGACCACCCAGAAGGGCCAGAAGAACAGCCGGGAGCGGA UGAAGCGGAUCGAGGAGGGCAUCAAGGAGCUGGGCAGCCAGAUCCUGA AGGAGCACCCCGUGGAGAACACCCAGCUGCAGAACGAGAAGCUGUACCU GUACUACCUGCAGAACGGCCGGGACAUGUACGUGGACCAGGAGCUGGA C AU C AAC CGGCU GAGCGAC U ACGACGUGGACGCC AU CGU GCC CC AGAGC UUCCUGAAGGACGACAGCAUCGACAACAAGGUGCUGACCCGGAGCGAC AAGAACCGGGGCAAGAGCGACAACGUGCCCAGCGAGGAGGUGGUGAAG AAGAUGAAGAACUACUGGCGGCAGCUGCUGAACGCCAAGCUGAUCACC CAGCGGAAGUUCGACAACCUGACCAAGGCCGAGCGGGGCGGCCUGAGCG AGCUGGACAAGGCCGGCUUCAUCAAGCGGCAGCUGGUGGAGACCCGGC AGAUCACCAAGCACGUGGCCCAGAUCCUGGACAGCCGGAUGAACACCAA GUACGACGAGAACGACAAGCUGAUCCGGGAGGUGAAGGUGAUCACCCU CAAGAGCAAGCUGGUGAGCGACUUCCGGAAGGACUUCCAGUUCUACAA GGUGCGGGAGAUCAACAACUACCACCACGCCCACGACGCCUACCUGAAC GCCGUGGUGGGCACCGCCCUGAUCAAGAAGUACCCCAAGCUGGAGAGCG AGUUCGUGUACGGCGACUACAAGGUGUACGACGUGCGGAAGAUGAUCG CCAAGAGCGAGCAGGAGAUCGGCAAGGCCACCGCCAAGUACUUCUUCUACAGCAACAUCAUGAACUUCUUCAAGACCGAGAUCACCCUGGCCAACGGCWSGR Docket No. 59761-805.601 GAGAUCCGGAAGCGGCCCCUGALTCGAGACCAACGGCGAGACCGGCGAGA U CGUGUGGGAC AAGGGCCGGGACUU CGCCACCGUGCGGAAGGUGC UGA GCAUGCCCCAGGUGAACAUCGUGAAGAAAACCGAGGUGCAGACCGGCG GCUUCAGCAAGGAGAGCAUCCUGCCCAAGGGCAACAGCGACAAGCUGA UCGCCCGGAAGAAGGACUGGGACCCCAAGAAGUACGGCGGCUUCAACA GCCCCACCGUGGCCUACAGCGLTGCLTGGUGGUGGCCAAGGUGGAGAAGG GCAAGAGCAAGAAGCU CAAGAGCGUGAAGGAGCUGCUGGGCAU CACCA UC AUGGAGCGGAGCAGCUUCGAGAAGAACCCC A LTCGGC LT UCCU GGAGG CCAAGGGCUACAAGGAGGUGAAGAAGGACCUGAUCAUCAAGCUGCCCA AGUACAGCCUGUUCGAGCUGGAGAACGGCCGGAAGCGGAUGCUGGCCA GCGCCUCCGUGCUGCACAAGGGCAACGAGCUGGCCCUGCCCAGCAAGUA CGUGAACLTUCCLTGUACCUGGCCAGCCACLTACGAGAAGCUGAAGGGCAGC UCCGAGGACAACAAGCAGAAGCAGCUGUUCGUGGAGCAGCACAAGCAC U ACCU GGACGAGAU C AU CGAGCAGAUCAGCGAGU U C AGCAAGCGGGUG AUCCUGGCCGACGCCAACCUGGACAAGGUGCUGAGCGCCUACAACAAGC ACCGGGACAAGCCCAUCCGGGAGCAGGCCGAGAACAUCAUCCACCUGUU CACCCUGACCAACCUGGGCGCCAGCGCCGCCUUCAAGUACUUCGACACC ACCAUCGGCCGGAAGCUGUACACCAGCACCAAGGAGGUGCUGGACGCCA CCCUGAUCCACCAGAGCAUCACCGGCCUGUACGAGACCCGGALTCGACCU GAGCCAGCUGGGCGGCGACAGCGGCGGCAGCAGCGGCGGCAGCAAGCGG ACCGCCGACGGCAGCGAGUUCGAGAGCCCCAAGAAGAAGCGGAAGGUG AGCGGCGGCAGCAGCGGCGGCAGCACCCUGAACAUCGAGGACGAGUACC GGCUGCACGAGACCAGCAAGGAGCCCGACGUGAGCCUGGGCAGCACCUG GCUGAGCGACUUCCCCCAGGCCUGGGCCGAGACCGGCGGCAUGGGCCUG GCCGUGCGGCAGGCCCCCCUGAUCAUCCCCCUGAAGGCCACCAGCACCC CCGUGAGCAUCAAGCAGUACCCCAUGAGCCAGGAGGCCCGGCUGGGCAU CAAGCCCCACAUCCAGCGGCUGCUGGACCAGGGCAUCCUGGUGCCCUGC CAGAGCCCCUGGAACACCCCCCUGCUGCCCGUGAAGAAGCCCGGCACCA ACGACUACCGGCCCGUGCAGGACCUGCGGGAGGUGAACAAGCGGGUGG AGGACAUCCACCCCAACGUGCCCAACCCCUACAACCUGCUGAGCGGCCU GCCCCCCAGCCACCAGUGGUACACCGUGCUGGACCUGAAGGACGCCUUC UUCUGCCUGCGGCUGCACCCCACCAGCCAGCCCCUGUUCGCCUUCGAGU GGCGGGACCCCGAGAUGGGCAUCAGCGGCCAGCUGACCUGGACCCGGCU GCCCCAGGGCUUCAAGAACAGCCCCACCCUGUUCUGCGAGGCCCUGCAC CGGGACCUGGCCGACUUCCGGAUCCAGCACCCCGACCUGAUCCUGCUGC AGLTACUACGACGACCUGCUGCUGGCCGCCACCAGCGAGCUGGACUGCCA GCAGGGCACCCGGGCCCUGCUGCAGACCCUGGGCAACCUGGGCUACCGG GCCAGCGCCAAGAAGGCCCAGAUCUGCCAGAAGCAGGUGAAGUACCUG GGCUACCUGCUGAAGGAGGGCCAGCGGUGGCUGACCGAGGCCCGGAAG GAGACCGUGAUGGGCCAGCCCACCCCCAAGACCCCCCGGCAGCUGCGGG AGUUCCUGGGCAAGGCCGGCUUCUGCCGGCUGUUCAUCCCCGGCUUCGC CGAGAUGGCCGCCCCCCUGUACCCCCUGACCAAGCCCGGCACCCUGUUC AACUGGGGCCCCGACCAGCAGAAGGCCUACCAGGAGAUCAAGCAGGCCC UGCUGACCGCCCCCGCCCUGGGCCUGCCCGACCUGACCAAGCCCUUCGA GCUGUUCGUGGACGAGAAGCAGGGCUACGCCAAGGGCGUGCUGACCCA GAAGCUGGGCCCCUGGCGGCGGCCCGUGGCCUACCUGAGCAAGAAGCUG GACCCCGUGGCCGCCGGCUGGCCCCCCUGCCUGCGGAUGGUGGCCGCCA UCGCCGUGCUGACCAAGGACGCCGGCAAGCUGACCAUGGGCCAGCCCCU GGUGALTCCUGGCCCCCCACGCCGUGGAGGCCCUGGUGAAGCAGCCCCCC GACCGGUGGCUGAGCAACGCCCGGAUGACCCACUACCAGGCCCUGCUGC UGGACACCGACCGGGUGCAGUUCGGCCCCGUGGUGGCCCUGAACCCCGC CACCCUGCUGCCCCUGCCCGAGGAGGGCCUGCAGCACAACUGCCUGGAC AGCGGCGGCAGCAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCA AGA AGA AGCGGA AGG U GGGC AGCGGC CC CGC C GC C A AGC GGG LT G A AGC UGGACUGAUAGUGAGCGGCCGCUUAAUUAAGCUGCCUUCUGCGGGGCUUGCCUUCUGGCCAAGCCCUUCUUCUCUCCCUUGCACCUGUACCUCUUGGWSGR Docket No. 59761-805.601 UCUUUGAAUAAAGCCUGAGUAGGAAGAAAAAAAAAAAAAAAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA (SEQ ID NO: 659) AGGCCACCAUGAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAA GAAGAAGCGGAAGGUGGACAAGAAGUACAGCAUCGGCCUGGACAUCGG CACCAACAGCGUGGGCUGGGCCGUGAUCACCGACGAGUACAAGGUGCCC AGCAAGAAGUUCAAGGUGCUGGGCAACACCGACCGGCACAGCAUCAAG AAGAACCUGAUCGGCGCCCUGCUGUUCGACAGCGGCGAGACCGCCGAGG CCACCCGGCUGAAGCGGACCGCCCGGCGGCGGUACACCCGGCGGAAGAA CCGGAUCUGCUACCUGCAGGAGAUCUUCAGCAACGAGAUGGCCAAGGU GGACGACAGCUUCUUCCACCGGCUGGAGGAGAGCUUCCUGGUGGAGGA GGACAAGAAGCACGAGCGGCACCCCAUCUUCGGCAACAUCGUGGACGAG GUGGCCUACCACGAGAAGUACCCCACCAUCUACCACCUGCGGAAGAAGC UGGUGGACAGCACCGACAAGGCCGACCUGCGGCUGAUCUACCUGGCCCU GGCCCACAUGAUCAAGUUCCGGGGCCACUUCCUGAUCGAGGGCGACCUG AACCCCGACAACAGCGACGUGGACAAGCUGUUCAUCCAGCUGGUGCAGA CCUACAACCAGCUGUUCGAGGAGAACCCCAUCAACGCCAGCGGCGUGGA CGCCAAGGCCAUCCUGAGCGCCCGGCUGAGCAAGAGCCGGAAGCUGGAG AACCUGAUCGCCCAGCUGCCCGGCGAGAAGAAGAACGGCCUGUUCGGCA ACCUGAUCGCCCUGAGCCUGGGCCUGACCCCCAACUUCAAGAGCAACUU CGACCUGGCCGAGGACGCCAAGCUGCAGCUGAGCAAGGACACCUACGAC GACGACCUGGACAACCUGCUGGCCCAGAUCGGCGACCAGUACGCCGACC UGUUCCUGGCCGCCAAGAACCUGAGCGACGCCAUCCUGCUGAGCGACAU CCUGCGGGUGAACACCGAGAUCACCAAGGCCCCCCUGAGCGCCAGCAUG GUGAAGCGGUACGACGAGCACCACCAGGACCUGACCCUGCUGAAGGCCC UGGUGCGGCAGCAGCUGCCCGAGAAGUACAAGGAGAUCUUCUUCGACC AGAGCAAGAACGGCUACGCCGGCUACAUCGACGGCGGCGCCAGCCAGGA GGAGUUCUACAAGUUCAUCAAGCCCAUCCUGGAGAAGAUGGACGGCAC CGAGGAGCUGCUGGUGAAGCUGAAGCGGGAGGACCUGCUGCGGAAGCA GCGGACCUUCGACAACGGCAUCAUCCCCCACCAGAUCCACCUGGGCGAG CUGCACGCCAUCCUGCGGCGGCAGGGCGACUUCUACCCCUUCCUGAAGG ACAACCGGGAGAAGAUCGAGAAGAUCCUGACCUUCCGGAUCCCCUACUA CGUGGGCCCCCUGGCCCGGGGCAACAGCCGGUUCGCCUGGAUGACCCGG AAGUCCGAGGAGACCAUCACCCCCUGGAACUUCGAGGAGGUGGUGGAC AAGGGCGCCAGCGCCCAGAGCUUCAUCGAGCGGAUGACCAACUUCGACA AGAACCUGCCCAACGAGAAGGUGCUGCCCAAGCACAGCCUGCUGUACGA GUACUUCACCGUGUACAACGAGCUGACCAAGGUGAAGUACGUGACCGA GGGCAUGCGGAAGCCCGCCUUCCUGAGCGGCGAGCAGAAGAAGGCCAUC GUGGACCUGCUGUUCAAGACCAACCGGAAGGUGACCGUGAAGCAGCUG AAGGAGGACUACUUCAAGAAGAUCGAGUGCUUCGACAGCGUGGAGAUC AGCGGCGUGGAGGACCGGUUCAACGCCAGCCUGGGCACCUACCACGACC UGCUGAAGAUCAUCAAGGACAAGGACUUCCUGGACAACGAGGAGAACG AGGACAUCCUGGAGGACAUCGUGCUGACCCUGACCCUGUUCGAGGACCG GGAGAUGAUCGAGGAGCGGCUCAAGACCUACGCCCACCUGUUCGACGAC AAGGUGAUGAAGCAGCUGAAGCGGCUGCGGUACACCGGCUGGGGCCGG CUGAGCCGGAAGCUGAUCAACGGCAUCCGGGACAAGCAGAGCGGCAAG ACCAUCCUGGACUUCCUCAAGAGCGACGGCUUCGCCAACCGGAACUUCA UGCAGCUGAUCCACGACGACAGCCUGACCUUCAAGGAGGACAUCCAGAA GGCCCAGGUGAGCGGCCAGGGCGACAGCCUGCACGAGCACAUCGCCAAC CUGGCCGGCAGCCCCGCCAUCAAGAAGGGCAUCCUGCAGACCGUGAAGG UGGUGGACGAGCUGGUGAAGGUGAUGGGCGGCCACAAGCCCGAGAACA UCGUGAUCGAGAUGGCCCGGGAGAACCAGACCACCCAGAAGGGCCAGA AGAACAGCCGGGAGCGGAUGAAGCGGAUCGAGGAGGGCAUCAAGGAGC UGGGCAGCCAGAUCCUGAAGGAGCACCCCGUGGAGAACACCCAGCUGCA GAACGAGAAGCUGUACCUGUACUACCUGCAGAACGGCCGGGACAUGUA1043 CGUGGACCAGGAGCUGGACAUCAACCGGCUGAGCGACUACGACGUGGAWSGR Docket No. 59761-805.601 CGCCAUCGUGCCCCAGAGCUUCCUGAAGGACGACAGCAUCGACAACAAG GUGCUGACCCGGAGCGACAAGAACCGGGGCAAGAGCGACAACGUGCCCA GCGAGGAGGUGGUGAAGAAGAUGAAGAACUACUGGCGGCAGCUGCUGA ACGCCAAGCUGAUCACCCAGCGGAAGUUCGACAACCUGACCAAGGCCGA GCGGGGCGGCCUGAGCGAGCUGGACAAGGCCGGCUUCAUCAAGCGGCA GCUGGUGGAGACCCGGCAGAUCACCAAGCACGUGGCCCAGAUCCUGGAC AGCCGGAUGAACACCAAGUACGACGAGAACGACAAGCUGAUCCGGGAG GUGAAGGUGAUCACCCUCAAGAGCAAGCUGGUGAGCGACUUCCGGAAG GACUUCCAGUUCUACAAGGUGCGGGAGAUCAACAACUACCACCACGCCC ACGACGCCUACCUGAACGCCGUGGUGGGCACCGCCCUGAUCAAGAAGUA CCCCAAGCUGGAGAGCGAGUUCGUGUACGGCGACUACAAGGUGUACGA CGUGCGGAAGAUGAUCGCCAAGAGCGAGCAGGAGAUCGGCAAGGCCAC CGCCAAGUACUUCUUCUACAGCAACAUCAUGAACUUCUUCAAGACCGAG AUCACCCUGGCCAACGGCGAGAUCCGGAAGCGGCCCCUGAUCGAGACCA ACGGCGAGACCGGCGAGAUCGUGUGGGACAAGGGCCGGGACUUCGCCA CCGUGCGGAAGGUGCUGAGCAUGCCCCAGGUGAACAUCGUGAAGAAAA CCGAGGUGCAGACCGGCGGCUUCAGCAAGGAGAGCAUCCUGCCCAAGGG CAACAGCGACAAGCUGAUCGCCCGGAAGAAGGACUGGGACCCCAAGAA GUACGGCGGCUUCAACAGCCCCACCGUGGCCUACAGCGUGCUGGUGGUG GCCAAGGUGGAGAAGGGCAAGAGCAAGAAGCUCAAGAGCGUGAAGGAG CUGCUGGGCAUCACCAUCAUGGAGCGGAGCAGCUUCGAGAAGAACCCCA UCGGCUUCCUGGAGGCCAAGGGCUACAAGGAGGUGAAGAAGGACCUGA UCAUCAAGCUGCCCAAGUACAGCCUGUUCGAGCUGGAGAACGGCCGGA AGCGGAUGCUGGCCAGCGCCUCCGUGCUGCACAAGGGCAACGAGCUGGC CCUGCCCAGCAAGUACGUGAACUUCCUGUACCUGGCCAGCCACUACGAG AAGCUGAAGGGCAGCUCCGAGGACAACAAGCAGAAGCAGCUGUUCGUG GAGCAGCACAAGCACUACCUGGACGAGAUCAUCGAGCAGAUCAGCGAG UUCAGCAAGCGGGUGAUCCUGGCCGACGCCAACCUGGACAAGGUGCUG AGCGCCUACAACAAGCACCGGGACAAGCCCAUCCGGGAGCAGGCCGAGA ACAUCAUCCACCUGUUCACCCUGACCAACCUGGGCGCCAGCGCCGCCUU CAAGUACUUCGACACCACCAUCGGCCGGAAGCUGUACACCAGCACCAAG GAGGUGCUGGACGCCACCCUGAUCCACCAGAGCAUCACCGGCCUGUACG AGACCCGGAUCGACCUGAGCCAGCUGGGCGGCGACAGCGGCGGCAGCAG CGGCGGCAGCAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAAG AAGAAGCGGAAGGUGAGCGGCGGCAGCAGCGGCGGCAGCACCCUGAAC AUCGAGGACGAGUACCGGCUGCACGAGACCAGCAAGGAGCCCGACGUG AGCCUGGGCAGCACCUGGCUGAGCGACUUCCCCCAGGCCUGGGCCGAGA CCGGCGGCAUGGGCCUGGCCGUGCGGCAGGCCCCCCUGAUCAUCCCCCU GAAGGCCACCAGCACCCCCGUGAGCAUCAAGCAGUACCCCAUGAGCCAG GAGGCCCGGCUGGGCAUCAAGCCCCACAUCCAGCGGCUGCUGGACCAGG GCAUCCUGGUGCCCUGCCAGAGCCCCUGGAACACCCCCCUGCUGCCCGU GAAGAAGCCCGGCACCAACGACUACCGGCCCGUGCAGGACCUGCGGGAG GUGAACAAGCGGGUGGAGGACAUCCACCCCAACGUGCCCAACCCCUACA ACCUGCUGAGCGGCCUGCCCCCCAGCCACCAGUGGUACACCGUGCUGGA CCUGAAGGACGCCUUCUUCUGCCUGCGGCUGCACCCCACCAGCCAGCCC CUGUUCGCCUUCGAGUGGCGGGACCCCGAGAUGGGCAUCAGCGGCCAGC UGACCUGGACCCGGCUGCCCCAGGGCUUCAAGAACAGCCCCACCCUGUU CUGCGAGGCCCUGCACCGGGACCUGGCCGACUUCCGGAUCCAGCACCCC GACCUGAUCCUGCUGCAGUACUACGACGACCUGCUGCUGGCCGCCACCA GCGAGCUGGACUGCCAGCAGGGCACCCGGGCCCUGCUGCAGACCCUGGG CAACCUGGGCUACCGGGCCAGCGCCAAGAAGGCCCAGAUCUGCCAGAAG CAGGUGAAGUACCUGGGCUACCUGCUGAAGGAGGGCCAGCGGUGGCUG ACCGAGGCCCGGAAGGAGACCGUGAUGGGCCAGCCCACCCCCAAGACCC CCCGGCAGCUGCGGGAGUUCCUGGGCAAGGCCGGCUUCUGCCGGCUGUU CAUCCCCGGCUUCGCCGAGAUGGCCGCCCCCCUGUACCCCCUGACCAAGCCCGGCACCCUGUUCAACUGGGGCCCCGACCAGCAGAAGGCCUACCAGGWSGR Docket No. 59761-805.601 AGAUCAAGCAGGCCCUGCUGACCGCCCCCGCCCUGGGCCUGCCCGACCU GACCAAGCCCUUCGAGCUGUUCGUGGACGAGAAGCAGGGCUACGCCAA GGGCGUGCUGACCCAGAAGCUGGGCCCCUGGCGGCGGCCCGUGGCCUAC CUGAGCAAGAAGCUGGACCCCGUGGCCGCCGGCUGGCCCCCCUGCCUGC GGAUGGUGGCCGCCAUCGCCGUGCUGACCAAGGACGCCGGCAAGCUGAC CAUGGGCCAGCCCCUGGUGAUCGGCGCCCCCCACGCCGUGGAGGCCCUG GUGAAGCAGCCCCCCGACCGGUGGCUGAGCAACGCCCGGAUGACCCACU ACCAGGCCCUGCUGCUGGACACCGACCGGGUGCAGUUCGGCCCCGUGGU GGCCCUGAACCCCGCCACCCUGCUGCCCCUGCCCGAGGAGGGCCUGCAG CACAACUGCCUGGACAGCGGCGGCAGCAAGCGGACCGCCGACGGCAGCG AGUUCGAGAGCCCCAAGAAGAAGCGGAAGGUGGGCAGCGGCCCCGCCGC CAAGCGGGUGAAGCUGGACUGAUAGUGAGCGGCCGCUUAAUUAAGCUG CCUUCUGCGGGGCUUGCCUUCUGGCCAAGCCCUUCUUCUCUCCCUUGCA CCUGUACCUCUUGGUCUUUGAAUAAAGCCUGAGUAGGAAGAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA(SEQ ID NO: 1043) AGGCCACCAUGAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAA GAAGAAGCGGAAGGUGGACAAGAAGUACAGCAUCGGCCUGGACAUCGG CACCAACAGCGUGGGCUGGGCCGUGAUCACCGACGAGUACAAGGUGCCC AGCAAGAAGUUCAAGGUGCUGGGCAACACCGACCGGCACAGCAUCAAG AAGAACCUGAUCGGCGCCCUGCUGUUCGACAGCGGCGAGACCGCCGAGG CCACCCGGCUGAAGCGGACCGCCCGGCGGCGGUACACCCGGCGGAAGAA CCGGAUCUGCUACCUGCAGGAGAUCUUCAGCAACGAGAUGGCCAAGGU GGACGACAGCUUCUUCCACCGGCUGGAGGAGAGCUUCCUGGUGGAGGA GGACAAGAAGCACGAGCGGCACCCCAUCUUCGGCAACAUCGUGGACGAG GUGGCCUACCACGAGAAGUACCCCACCAUCUACCACCUGCGGAAGAAGC UGGUGGACAGCACCGACAAGGCCGACCUGCGGCUGAUCUACCUGGCCCU GGCCCACAUGAUCAAGUUCCGGGGCCACUUCCUGAUCGAGGGCGACCUG AACCCCGACAACAGCGACGUGGACAAGCUGUUCAUCCAGCUGGUGCAGA CCUACAACCAGCUGUUCGAGGAGAACCCCAUCAACGCCAGCGGCGUGGA CGCCAAGGCCAUCCUGAGCGCCCGGCUGAGCAAGAGCCGGAAGCUGGAG AACCUGAUCGCCCAGCUGCCCGGCGAGAAGAAGAACGGCCUGUUCGGCA ACCUGAUCGCCCUGAGCCUGGGCCUGACCCCCAACUUCAAGAGCAACUU CGACCUGGCCGAGGACGCCAAGCUGCAGCUGAGCAAGGACACCUACGAC GACGACCUGGACAACCUGCUGGCCCAGAUCGGCGACCAGUACGCCGACC UGUUCCUGGCCGCCAAGAACCUGAGCGACGCCAUCCUGCUGAGCGACAU CCUGCGGGUGAACACCGAGAUCACCAAGGCCCCCCUGAGCGCCAGCAUG GUGAAGCGGUACGACGAGCACCACCAGGACCUGACCCUGCUGAAGGCCC UGGUGCGGCAGCAGCUGCCCGAGAAGUACAAGGAGAUCUUCUUCGACC AGAGCAAGAACGGCUACGCCGGCUACAUCGACGGCGGCGCCAGCCAGGA GGAGUUCUACAAGUUCAUCAAGCCCAUCCUGGAGAAGAUGGACGGCAC CGAGGAGCUGCUGGUGAAGCUGAAGCGGGAGGACCUGCUGCGGAAGCA GCGGACCUUCGACAACGGCAUCAUCCCCCACCAGAUCCACCUGGGCGAG CUGCACGCCAUCCUGCGGCGGCAGGGCGACUUCUACCCCUUCCUGAAGG ACAACCGGGAGAAGAUCGAGAAGAUCCUGACCUUCCGGAUCCCCUACUA CGUGGGCCCCCUGGCCCGGGGCAACAGCCGGUUCGCCUGGAUGACCCGG AAGUCCGAGGAGACCAUCACCCCCUGGAACUUCGAGGAGGUGGUGGAC AAGGGCGCCAGCGCCCAGAGCUUCAUCGAGCGGAUGACCAACUUCGACA AGAACCUGCCCAACGAGAAGGUGCUGCCCAAGCACAGCCUGCUGUACGA GUACUUCACCGUGUACAACGAGCUGACCAAGGUGAAGUACGUGACCGA GGGCAUGCGGAAGCCCGCCUUCCUGAGCGGCGAGCAGAAGAAGGCCAUC GUGGACCUGCUGUUCAAGACCAACCGGAAGGUGACCGUGAAGCAGCUG AAGGAGGACUACUUCAAGAAGAUCGAGUGCUUCGACAGCGUGGAGAUC AGCGGCGUGGAGGACCGGUUCAACGCCAGCCUGGGCACCUACCACGACC661 UGCUGAAGAUCAUCAAGGACAAGGACUUCCUGGACAACGAGGAGAACGWSGR Docket No. 59761-805.601 AGGACAUCCUGGAGGACAUCGUGCUGACCCUGACCCUGUUCGAGGACCG GGAGAUGAUCGAGGAGCGGCUCAAGACCUACGCCCACCUGUUCGACGAC AAGGUGAUGAAGCAGCUGAAGCGGCUGCGGUACACCGGCUGGGGCCGG CUGAGCCGGAAGCUGAUCAACGGCAUCCGGGACAAGCAGAGCGGCAAG ACCAUCCUGGACUUCCUCAAGAGCGACGGCUUCGCCAACCGGAACUUCA UGCAGCUGAUCCACGACGACAGCCUGACCUUCAAGGAGGACAUCCAGAA GGCCCAGGUGAGCGGCCAGGGCGACAGCCUGCACGAGCACAUCGCCAAC CUGGCCGGCAGCCCCGCCAUCAAGAAGGGCAUCCUGCAGACCGUGAAGG UGGUGGACGAGCUGGUGAAGGUGAUGGGCGGCCACAAGCCCGAGAACA UCGUGAUCGAGAUGGCCCGGGAGAACCAGACCACCCAGAAGGGCCAGA AGAACAGCCGGGAGCGGAUGAAGCGGAUCGAGGAGGGCAUCAAGGAGC UGGGCAGCCAGAUCCUGAAGGAGCACCCCGUGGAGAACACCCAGCUGCA GAACGAGAAGCUGUACCUGUACUACCUGCAGAACGGCCGGGACAUGUA CGUGGACCAGGAGCUGGACAUCAACCGGCUGAGCGACUACGACGUGGA CGCCAUCGUGCCCCAGAGCUUCCUGAAGGACGACAGCAUCGACAACAAG GUGCUGACCCGGAGCGACAAGAACCGGGGCAAGAGCGACAACGUGCCCA GCGAGGAGGUGGUGAAGAAGAUGAAGAACUACUGGCGGCAGCUGCUGA ACGCCAAGCUGAUCACCCAGCGGAAGUUCGACAACCUGACCAAGGCCGA GCGGGGCGGCCUGAGCGAGCUGGACAAGGCCGGCUUCAUCAAGCGGCA GCUGGUGGAGACCCGGCAGAUCACCAAGCACGUGGCCCAGAUCCUGGAC AGCCGGAUGAACACCAAGUACGACGAGAACGACAAGCUGAUCCGGGAG GUGAAGGUGAUCACCCUCAAGAGCAAGCUGGUGAGCGACUUCCGGAAG GACUUCCAGUUCUACAAGGUGCGGGAGAUCAACAACUACCACCACGCCC ACGACGCCUACCUGAACGCCGUGGUGGGCACCGCCCUGAUCAAGAAGUA CCCCAAGCUGGAGAGCGAGUUCGUGUACGGCGACUACAAGGUGUACGA CGUGCGGAAGAUGAUCGCCAAGAGCGAGCAGGAGAUCGGCAAGGCCAC CGCCAAGUACUUCUUCUACAGCAACAUCAUGAACUUCUUCAAGACCGAG AUCACCCUGGCCAACGGCGAGAUCCGGAAGCGGCCCCUGAUCGAGACCA ACGGCGAGACCGGCGAGAUCGUGUGGGACAAGGGCCGGGACUUCGCCA CCGUGCGGAAGGUGCUGAGCAUGCCCCAGGUGAACAUCGUGAAGAAAA CCGAGGUGCAGACCGGCGGCUUCAGCAAGGAGAGCAUCCUGCCCAAGGG CAACAGCGACAAGCUGAUCGCCCGGAAGAAGGACUGGGACCCCAAGAA GUACGGCGGCUUCAACAGCCCCACCGUGGCCUACAGCGUGCUGGUGGUG GCCAAGGUGGAGAAGGGCAAGAGCAAGAAGCUCAAGAGCGUGAAGGAG CUGCUGGGCAUCACCAUCAUGGAGCGGAGCAGCUUCGAGAAGAACCCCA UCGGCUUCCUGGAGGCCAAGGGCUACAAGGAGGUGAAGAAGGACCUGA UCAUCAAGCUGCCCAAGUACAGCCUGUUCGAGCUGGAGAACGGCCGGA AGCGGAUGCUGGCCAGCGCCUCCGUGCUGCACAAGGGCAACGAGCUGGC CCUGCCCAGCAAGUACGUGAACUUCCUGUACCUGGCCAGCCACUACGAG AAGCUGAAGGGCAGCUCCGAGGACAACAAGCAGAAGCAGCUGUUCGUG GAGCAGCACAAGCACUACCUGGACGAGAUCAUCGAGCAGAUCAGCGAG UUCAGCAAGCGGGUGAUCCUGGCCGACGCCAACCUGGACAAGGUGCUG AGCGCCUACAACAAGCACCGGGACAAGCCCAUCCGGGAGCAGGCCGAGA ACAUCAUCCACCUGUUCACCCUGACCAACCUGGGCGCCAGCGCCGCCUU CAAGUACUUCGACACCACCAUCGGCCGGAAGCUGUACACCAGCACCAAG GAGGUGCUGGACGCCACCCUGAUCCACCAGAGCAUCACCGGCCUGUACG AGACCCGGAUCGACCUGAGCCAGCUGGGCGGCGACAGCGGCGGCAGCAG CGGCGGCAGCAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAAG AAGAAGCGGAAGGUGAGCGGCGGCAGCAGCGGCGGCAGCACCCUGAAC AUCGAGGACGAGUACCGGCUGCACGAGACCAGCAAGGAGCCCGACGUG AGCCUGGGCAGCACCUGGCUGAGCGACUUCCCCCAGGCCUGGGCCGAGA CCGGCGGCAUGGGCCUGGCCGUGCGGCAGGCCCCCCUGAUCAUCCCCCU GAAGGCCACCAGCACCCCCGUGAGCAUCAAGCAGUACCCCAUGAGCCAG GAGGCCCGGCUGGGCAUCAAGCCCCACAUCCAGCGGCUGCUGGACCAGG GCAUCCUGGUGCCCUGCCAGAGCCCCUGGAACACCCCCCUGCUGCCCGUGAAGAAGCCCGGCACCAACGACUACCGGCCCGUGCAGGACCUGCGGGAGWSGR Docket No. 59761-805.601 GUGAACAAGCGGGUGGAGGACAUCCACCCCAACGUGCCCAACCCCUACA ACCUGCUGAGCGGCCUGCCCCCCAGCCACCAGUGGUACACCGUGCUGGA CCUGAAGGACGCCUUCUUCUGCCUGCGGCUGCACCCCACCAGCCAGCCC CUGUUCGCCUUCGAGUGGCGGGACCCCGAGAUGGGCAUCAGCGGCCAGC UGACCUGGACCCGGCUGCCCCAGGGCUUCAAGAACAGCCCCACCCUGUU CUGCGAGGCCCUGCACCGGGACCUGGCCGACUUCCGGAUCCAGCACCCC GACCUGAUCCUGCUGCAGUACUACGACGACCUGCUGCUGGCCGCCACCA GCGAGCUGGACUGCCAGCAGGGCACCCGGGCCCUGCUGCAGACCCUGGG CAACCUGGGCUACCGGGCCAGCGCCAAGAAGGCCCAGAUCUGCCAGAAG CAGGUGAAGUACCUGGGCUACCUGCUGAAGGAGGGCCAGCGGUGGCUG ACCGAGGCCCGGAAGGAGACCGUGAUGGGCCAGCCCACCCCCAAGACCC CCCGGCAGCUGCGGGAGUUCCUGGGCAAGGCCGGCUUCUGCCGGCUGUU CAUCCCCGGCUUCGCCGAGAUGGCCGCCCCCCUGUACCCCCUGACCAAG CCCGGCACCCUGUUCAACUGGGGCCCCGACCAGCAGAAGGCCUACCAGG AGAUCAAGCAGGCCCUGCUGACCGCCCCCGCCCUGGGCCUGCCCGACCU GACCAAGCCCUUCGAGCUGUUCGUGGACGAGAAGCAGGGCUACGCCAA GGGCGUGCUGACCCAGAAGCUGGGCCCCUGGCGGCGGCCCGUGGCCUAC CUGAGCAAGAAGCUGGACCCCGUGGCCGCCGGCUGGCCCCCCUGCCUGC GGAUGGUGGCCGCCAUCGCCGUGCUGACCAAGGACGCCGGCAAGCUGAC CAUGGGCCAGCCCCUGGUGAUCCUGGCCCCCCACGCCGUGGAGGCCCUG GUGAAGCAGCCCCCCGACCGGUGGCUGAGCAACGCCCGGAUGACCCACU ACCAGGCCCUGCUGCUGGACACCGACCGGGUGCAGUUCGGCCCCGUGGU GGCCCUGAACCCCGCCACCCUGCUGCCCCUGCCCGAGGAGGGCCUGCAG CACAACUGCCUGGACAGCGGCGGCAGCAAGCGGACCGCCGACGGCAGCG AGUUCGAGAGCCCCAAGAAGAAGCGGAAGGUGGGCAGCGGCCCCGCCGC CAAGCGGGUGAAGCUGGACUGAUAGUGAGCGGCCGCUUAAUUAAGCUG CCUUCUGCGGGGCUUGCCUUCUGGCCAAGCCCUUCUUCUCUCCCUUGCA CCUGUACCUCUUGGUCUUUGAAUAAAGCCUGAGUAGGAAGAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA(SEQ ID NO: 661) AGGCCACCAUGAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAA GAAGAAGCGGAAGGUGGACAAGAAGUACAGCAUCGGCCUGGACAUCGG CACCAACAGCGUGGGCUGGGCCGUGAUCACCGACGAGUACAAGGUGCCC AGCAAGAAGUUCAAGGUGCUGGGCAACACCGACCGGCACAGCAUCAAG AAGAACCUGAUCGGCGCCCUGCUGUUCGACAGCGGCGAGACCGCCGAGG CCACCCGGCUGAAGCGGACCGCCCGGCGGCGGUACACCCGGCGGAAGAA CCGGAUCUGCUACCUGCAGGAGAUCUUCAGCAACGAGAUGGCCAAGGU GGACGACAGCUUCUUCCACCGGCUGGAGGAGAGCUUCCUGGUGGAGGA GGACAAGAAGCACGAGCGGCACCCCAUCUUCGGCAACAUCGUGGACGAG GUGGCCUACCACGAGAAGUACCCCACCAUCUACCACCUGCGGAAGAAGC UGGUGGACAGCACCGACAAGGCCGACCUGCGGCUGAUCUACCUGGCCCU GGCCCACAUGAUCAAGUUCCGGGGCCACUUCCUGAUCGAGGGCGACCUG AACCCCGACAACAGCGACGUGGACAAGCUGUUCAUCCAGCUGGUGCAGA CCUACAACCAGCUGUUCGAGGAGAACCCCAUCAACGCCAGCGGCGUGGA CGCCAAGGCCAUCCUGAGCGCCCGGCUGAGCAAGAGCCGGAAGCUGGAG AACCUGAUCGCCCAGCUGCCCGGCGAGAAGAAGAACGGCCUGUUCGGCA ACCUGAUCGCCCUGAGCCUGGGCCUGACCCCCAACUUCAAGAGCAACUU CGACCUGGCCGAGGACGCCAAGCUGCAGCUGAGCAAGGACACCUACGAC GACGACCUGGACAACCUGCUGGCCCAGAUCGGCGACCAGUACGCCGACC UGUUCCUGGCCGCCAAGAACCUGAGCGACGCCAUCCUGCUGAGCGACAU CCUGCGGGUGAACACCGAGAUCACCAAGGCCCCCCUGAGCGCCAGCAUG GUGAAGCGGUACGACGAGCACCACCAGGACCUGACCCUGCUGAAGGCCC UGGUGCGGCAGCAGCUGCCCGAGAAGUACAAGGAGAUCUUCUUCGACC AGAGCAAGAACGGCUACGCCGGCUACAUCGACGGCGGCGCCAGCCAGGA1044 GGAGUUCUACAAGUUCAUCAAGCCCAUCCUGGAGAAGAUGGACGGCACWSGR Docket No. 59761-805.601 CGAGGAGCUGCUGGUGAAGCUGAAGCGGGAGGACCUGCUGCGGAAGCA GCGGACCUUCGACAACGGCAUCAUCCCCCACCAGAUCCACCUGGGCGAG CUGCACGCCAUCCUGCGGCGGCAGGGCGACUUCUACCCCUUCCUGAAGG ACAACCGGGAGAAGAUCGAGAAGAUCCUGACCUUCCGGAUCCCCUACUA CGUGGGCCCCCUGGCCCGGGGCAACAGCCGGUUCGCCUGGAUGACCCGG AAGUCCGAGGAGACCAUCACCCCCUGGAACUUCGAGGAGGUGGUGGAC AAGGGCGCCAGCGCCCAGAGCUUCAUCGAGCGGAUGACCAACUUCGACA AGAACCUGCCCAACGAGAAGGUGCUGCCCAAGCACAGCCUGCUGUACGA GUACUUCACCGUGUACAACGAGCUGACCAAGGUGAAGUACGUGACCGA GGGCAUGCGGAAGCCCGCCUUCCUGAGCGGCGAGCAGAAGAAGGCCAUC GUGGACCUGCUGUUCAAGACCAACCGGAAGGUGACCGUGAAGCAGCUG AAGGAGGACUACUUCAAGAAGAUCGAGUGCUUCGACAGCGUGGAGAUC AGCGGCGUGGAGGACCGGUUCAACGCCAGCCUGGGCACCUACCACGACC UGCUGAAGAUCAUCAAGGACAAGGACUUCCUGGACAACGAGGAGAACG AGGACAUCCUGGAGGACAUCGUGCUGACCCUGACCCUGUUCGAGGACCG GGAGAUGAUCGAGGAGCGGCUCAAGACCUACGCCCACCUGUUCGACGAC AAGGUGAUGAAGCAGCUGAAGCGGCUGCGGUACACCGGCUGGGGCCGG CUGAGCCGGAAGCUGAUCAACGGCAUCCGGGACAAGCAGAGCGGCAAG ACCAUCCUGGACUUCCUCAAGAGCGACGGCUUCGCCAACCGGAACUUCA UGCAGCUGAUCCACGACGACAGCCUGACCUUCAAGGAGGACAUCCAGAA GGCCCAGGUGAGCGGCCAGGGCGACAGCCUGCACGAGCACAUCGCCAAC CUGGCCGGCAGCCCCGCCAUCAAGAAGGGCAUCCUGCAGACCGUGAAGG UGGUGGACGAGCUGGUGAAGGUGAUGGGCGGCCACAAGCCCGAGAACA UCGUGAUCGAGAUGGCCCGGGAGAACCAGACCACCCAGAAGGGCCAGA AGAACAGCCGGGAGCGGAUGAAGCGGAUCGAGGAGGGCAUCAAGGAGC UGGGCAGCCAGAUCCUGAAGGAGCACCCCGUGGAGAACACCCAGCUGCA GAACGAGAAGCUGUACCUGUACUACCUGCAGAACGGCCGGGACAUGUA CGUGGACCAGGAGCUGGACAUCAACCGGCUGAGCGACUACGACGUGGA CGCCAUCGUGCCCCAGAGCUUCCUGAAGGACGACAGCAUCGACAACAAG GUGCUGACCCGGAGCGACAAGAACCGGGGCAAGAGCGACAACGUGCCCA GCGAGGAGGUGGUGAAGAAGAUGAAGAACUACUGGCGGCAGCUGCUGA ACGCCAAGCUGAUCACCCAGCGGAAGUUCGACAACCUGACCAAGGCCGA GCGGGGCGGCCUGAGCGAGCUGGACAAGGCCGGCUUCAUCAAGCGGCA GCUGGUGGAGACCCGGCAGAUCACCAAGCACGUGGCCCAGAUCCUGGAC AGCCGGAUGAACACCAAGUACGACGAGAACGACAAGCUGAUCCGGGAG GUGAAGGUGAUCACCCUCAAGAGCAAGCUGGUGAGCGACUUCCGGAAG GACUUCCAGUUCUACAAGGUGCGGGAGAUCAACAACUACCACCACGCCC ACGACGCCUACCUGAACGCCGUGGUGGGCACCGCCCUGAUCAAGAAGUA CCCCAAGCUGGAGAGCGAGUUCGUGUACGGCGACUACAAGGUGUACGA CGUGCGGAAGAUGAUCGCCAAGAGCGAGCAGGAGAUCGGCAAGGCCAC CGCCAAGUACUUCUUCUACAGCAACAUCAUGAACUUCUUCAAGACCGAG AUCACCCUGGCCAACGGCGAGAUCCGGAAGCGGCCCCUGAUCGAGACCA ACGGCGAGACCGGCGAGAUCGUGUGGGACAAGGGCCGGGACUUCGCCA CCGUGCGGAAGGUGCUGAGCAUGCCCCAGGUGAACAUCGUGAAGAAAA CCGAGGUGCAGACCGGCGGCUUCAGCAAGGAGAGCAUCCUGCCCAAGGG CAACAGCGACAAGCUGAUCGCCCGGAAGAAGGACUGGGACCCCAAGAA GUACGGCGGCUUCAACAGCCCCACCGUGGCCUACAGCGUGCUGGUGGUG GCCAAGGUGGAGAAGGGCAAGAGCAAGAAGCUCAAGAGCGUGAAGGAG CUGCUGGGCAUCACCAUCAUGGAGCGGAGCAGCUUCGAGAAGAACCCCA UCGGCUUCCUGGAGGCCAAGGGCUACAAGGAGGUGAAGAAGGACCUGA UCAUCAAGCUGCCCAAGUACAGCCUGUUCGAGCUGGAGAACGGCCGGA AGCGGAUGCUGGCCAGCGCCUCCGUGCUGCACAAGGGCAACGAGCUGGC CCUGCCCAGCAAGUACGUGAACUUCCUGUACCUGGCCAGCCACUACGAG AAGCUGAAGGGCAGCUCCGAGGACAACAAGCAGAAGCAGCUGUUCGUG GAGCAGCACAAGCACUACCUGGACGAGAUCAUCGAGCAGAUCAGCGAGUUCAGCAAGCGGGUGAUCCUGGCCGACGCCAACCUGGACAAGGUGCUGWSGR Docket No. 59761-805.601 AGCGCCUACAACAAGCACCGGGACAAGCCCAUCCGGGAGCAGGCCGAGA ACAUCAUCCACCUGUUCACCCUGACCAACCUGGGCGCCAGCGCCGCCUU CAAGUACUUCGACACCACCAUCGGCCGGAAGCUGUACACCAGCACCAAG GAGGUGCUGGACGCCACCCUGAUCCACCAGAGCAUCACCGGCCUGUACG AGACCCGGAUCGACCUGAGCCAGCUGGGCGGCGACAGCGGCGGCAGCAG CGGCGGCAGCAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAAG AAGAAGCGGAAGGUGAGCGGCGGCAGCAGCGGCGGCAGCACCCUGAAC AUCGAGGACGAGUACCGGCUGCACGAGACCAGCAAGGAGCCCGACGUG AGCCUGGGCAGCACCUGGCUGAGCGACUUCCCCCAGGCCUGGGCCGAGA CCUGGGGCAUGGGCCUGGCCGUGCGGCAGGCCCCCCUGAUCAUCCCCCU GAAGGCCACCAGCACCCCCGUGAGCAUCAAGCAGUACCCCAUGAGCCAG GAGGCCCGGCUGGGCAUCAAGCCCCACAUCCAGCGGCUGCUGGACCAGG GCAUCCUGGUGCCCUGCCAGAGCCCCUGGAACACCCCCCUGCUGCCCGU GAAGAAGCCCGGCACCAACGACUACCGGCCCGUGCAGGACCUGCGGGAG GUGAACAAGCGGGUGGAGGACAUCCACCCCAACGUGCCCAACCCCUACA ACCUGCUGAGCGGCCUGCCCCCCAGCCACCAGUGGUACACCGUGCUGGA CCUGAAGGACGCCUUCUUCUGCCUGCGGCUGCACCCCACCAGCCAGCCC CUGUUCGCCUUCGAGUGGCGGGACCCCGAGAUGGGCAUCAGCGGCCAGC UGACCUGGACCCGGCUGCCCCAGGGCUUCAAGAACAGCCCCACCCUGUU CUGCGAGGCCCUGCACCGGGACCUGGCCGACUUCCGGAUCCAGCACCCC GACCUGAUCCUGCUGCAGUACUACGACGACCUGCUGCUGGCCGCCACCA GCGAGCUGGACUGCCAGCAGGGCACCCGGGCCCUGCUGCAGACCCUGGG CAACCUGGGCUACCGGGCCAGCGCCAAGAAGGCCCAGAUCUGCCAGAAG CAGGUGAAGUACCUGGGCUACCUGCUGAAGGAGGGCCAGCGGUGGCUG ACCGAGGCCCGGAAGGAGACCGUGAUGGGCCAGCCCACCCCCAAGACCC CCCGGCAGCUGCGGGAGUUCCUGGGCAAGGCCGGCUUCUGCCGGCUGUU CAUCCCCGGCUUCGCCGAGAUGGCCGCCCCCCUGUACCCCCUGACCAAG CCCGGCACCCUGUUCAACUGGGGCCCCGACCAGCAGAAGGCCUACCAGG AGAUCAAGCAGGCCCUGCUGACCGCCCCCGCCCUGGGCCUGCCCGACCU GACCAAGCCCUUCGAGCUGUUCGUGGACGAGAAGCAGGGCUACGCCAA GGGCGUGCUGACCCAGAAGCUGGGCCCCUGGCGGCGGCCCGUGGCCUAC CUGAGCAAGAAGCUGGACCCCGUGGCCGCCGGCUGGCCCCCCUGCCUGC GGAUGGUGGCCGCCAUCGCCGUGCUGACCAAGGACGCCGGCAAGCUGAC CAUGGGCCAGCCCCUGGUGAUCGGCGCCCCCCACGCCGUGGAGGCCCUG GUGAAGCAGCCCCCCGACCGGUGGCUGAGCAACGCCCGGAUGACCCACU ACCAGGCCCUGCUGCUGGACACCGACCGGGUGCAGUUCGGCCCCGUGGU GGCCCUGAACCCCGCCACCCUGCUGCCCCUGCCCGAGGAGGGCCUGCAG CACAACUGCCUGGACAGCGGCGGCAGCAAGCGGACCGCCGACGGCAGCG AGUUCGAGAGCCCCAAGAAGAAGCGGAAGGUGGGCAGCGGCCCCGCCGC CAAGCGGGUGAAGCUGGACUGAUAGUGAGCGGCCGCUUAAUUAAGCUG CCUUCUGCGGGGCUUGCCUUCUGGCCAAGCCCUUCUUCUCUCCCUUGCA CCUGUACCUCUUGGUCUUUGAAUAAAGCCUGAGUAGGAAGAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA(SEQ ID NO: 1044) AGGCCACCAUGAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAA GAAGAAGCGGAAGGUGGACAAGAAGUACAGCAUCGGCCUGGACAUCGG CACCAACAGCGUGGGCUGGGCCGUGAUCACCGACGAGUACAAGGUGCCC AGCAAGAAGUUCAAGGUGCUGGGCAACACCGACCGGCACAGCAUCAAG AAGAACCUGAUCGGCGCCCUGCUGUUCGACAGCGGCGAGACCGCCGAGG CCACCCGGCUGAAGCGGACCGCCCGGCGGCGGUACACCCGGCGGAAGAA CCGGAUCUGCUACCUGCAGGAGAUCUUCAGCAACGAGAUGGCCAAGGU GGACGACAGCUUCUUCCACCGGCUGGAGGAGAGCUUCCUGGUGGAGGA GGACAAGAAGCACGAGCGGCACCCCAUCUUCGGCAACAUCGUGGACGAG GUGGCCUACCACGAGAAGUACCCCACCAUCUACCACCUGCGGAAGAAGC1046 UGGUGGACAGCACCGACAAGGCCGACCUGCGGCUGAUCUACCUGGCCCUWSGR Docket No. 59761-805.601 GGCCCACAUGAUCAAGUUCCGGGGCCACUUCCUGAUCGAGGGCGACCUG AACCCCGACAACAGCGACGUGGACAAGCUGUUCAUCCAGCUGGUGCAGA CCUACAA...
Claims
WSGR Docket No. 59761-805.601CLAIMS WHAT IS CLAIMED IS:
1. A prime editing guide RNA (PEgRNA) or a nucleic acid encoding the PEgRNA, wherein the PEgRNA comprises:(a) a spacer comprising a nucleotide sequence that is complementary to a search target sequence on a first strand of a SERPINA1 gene, wherein the nucleotide sequence of the spacer comprises SEQ ID NO: 1 at its 3’ end;(b) a gRNA core capable of binding to a Cas9 protein; and(c) an extension arm comprising:i. an editing template comprising a nucleotide sequence that comprises a region of complementarity to an editing target sequence on a second strand of the SERPINA 1 gene, andii. a primer binding site (PBS) comprising a nucleotide sequence that comprises a reverse complement of nucleotides 10-14 of SEQ ID NO: 1 at its 5’ end; wherein the first strand and second strand are complementary to each other; andwherein the nucleotide sequence of the editing template encodes a wildtype amino acid sequence of an alpha-antitrypsin (AAT) protein and further encodes one or more synonymous nucleotide transversion edits compared to a wildtype SERPINA 1 gene sequence.
2. The PEgRNA of claim 1, wherein the nucleotide sequence of the spacer is SEQ ID NO: 4.
3. The PEgRNA of claim 1 or 2, wherein the nucleotide sequence of the PBS is sequence number 7, 8, 9, 10, 19, or 21.
4. The PEgRNA of claim 1 or 2, wherein the nucleotide sequence of the PBS is sequence number 21.
5. The PEgRNA of any one of claims 1 -4, wherein the one or more synonymous nucleotide transversion edits are independently selected from the group consisting of: c.1086G> C, c.1086G> T, C.1089OA, C.1089OG, C.1092OA, c.1104G> C, c.1104G> T, c.1107T> G, c.1107T> G, c.1116T> G, c.1116T> G, and c. H16T> A.
6. The PEgRNA of any one of claims 1-5, wherein the nucleotide sequence of the editing template further encodes one or more synonymous nucleotide transition edits compared to a wildtype SERPINA 1 gene sequence, wherein the one or more synonymous nucleotide transitions are independently selected from the group consisting of: c,1086G> A, C.1089OT, C.1092OT, C.1095OT, c. H01A> G, c. H04G> A, c.1107T> C, c.1110A> G, and c.1113T> C.
7. The PEgRNA of any one of claims 1-6, wherein the nucleotide sequence of the editing template is 13-37 nucleotides in length, 13-22 nucleotides in length, or 16 nucleotides in length.WSGR Docket No. 59761-805.6018. The PEgRNA of any one of claims 1-4, wherein the one or more synonymous nucleotide transversion edits comprise c.lO86G> C and wherein the nucleotide sequence of the editing template comprises SEQ ID NO: 62, 22, 23, 25, 26, 28, 29, 32, 34, 35, 39, 43, 44, 47, 49, 51, 55, 57, 58, 59, 63, 66, 67, 69, 70, 75, 77, 79, 81, 85, 87, 92, 93, 94, 96, 97, 108, 109, 112, or 114 at its 3’ end.
9. The PEgRNA of claim 8, wherein the nucleotide sequence of the editing template encodes wildtype SERPINA 1 gene sequence except for:(a) c.1086G> C and c.1092OA edits and wherein the nucleotide sequence of the editing template is SEQ ID NO: 62;(b) a c.1086G> C edit and wherein the nucleotide sequence of the editing template is SEQ ID NO: 22, 23, 25, 28, 44, 47, 59, 63, or 75;(c) c.lO86G> C, C.1089OA, C.1092OT, and c.lO95C> T edits and wherein the editing template is SEQ ID NO: 35;(d) c.1086G> C and c.1089OG edits and wherein the nucleotide sequence of the editing template is SEQ ID NO: 29, 55, or 67;(e) c.1086G> C, c.1089OG, and c.1092OA edits and wherein the nucleotide sequence of the editing template is SEQ ID NO: 39;(f) c.1086G> C, c.1089OG, and c.1092OT edits and wherein the nucleotide sequence of the editing template is SEQ ID NO: 26 or 49;(g) c.1086G> C, c.1089OG, and c.1095OT edits and wherein the nucleotide sequence of the editing template is SEQ ID NO: 34 or 51;(h) c.1086G> C, c.1089OT, c.1092OT, and c.1095OT edits and wherein the nucleotide sequence of the editing template is SEQ ID NO: 32, 57, or 77;(i) c.1086G> C and c.1095OT edits and wherein the nucleotide sequence of the editing template is SEQ ID NO: 43, 58, 66, 81, 97, or 112;(j) c.lO86G> C, C.1095OT, and c.1101A> G edits and wherein the nucleotide sequence of the editing template is SEQ ID NO: 69;(k) c.1086G> C, c.1095C>T, c.1101A> G, and c. H04G> T edits and wherein the nucleotide sequence of the editing template is SEQ ID NO: 85;(l) c.lO86G> C, C.1095OT, c.l 101A> G, c. H07T> C, and c.l 110A> G edits and wherein the nucleotide sequence of the editing template is SEQ ID NO: 94;(m) c.1086G> C, c.1095C>T, c.1101A> G, and c.l 107T> G and wherein the nucleotide sequence of the editing template is SEQ ID NO: 93;WSGR Docket No. 59761-805.601(n) c.1086G> C, c.1095C>T, c.1101A> G, c. H07T> G, c. H10A> G, c. H13T> C, and c.1116T> A edits and wherein the nucleotide sequence of the editing template is SEQ ID NO: 114;(o) c.1086G> C, c.1095C>T, c.1101A> G, c. H07T> G, c. H10A> G, c. H13T> C, and c.1116T> G edits and wherein the nucleotide sequence of the editing template is SEQ ID NO: 108;(p) c.lO86G> C and c. H01A> G edits and wherein the nucleotide sequence of the editing template is SEQ ID NO: 70;(q) c.lO86G> C, c. H01A> G, and c. H04G> C and wherein the nucleotide sequence of the editing template is SEQ ID NO: 92;(r) c.lO86G> C, c. H01A> G, and c. H07T> G edits and wherein the nucleotide sequence of the editing template is SEQ ID NO: 87;(s) c.lO86G> C, c. H01A> G, c. H07T> G, and c. H10A> G edits and wherein the nucleotide sequence of the editing template is SEQ ID NO: 96;(t) c.lO86G> C, c. H01A> G, c. H07T> G, c. H10A> G, c. H13T> C, and c. H16T> G edits and wherein the nucleotide sequence of the editing template is SEQ ID NO: 109; or(u) c.1086G> C and c.1104G> C edits and wherein the nucleotide sequence of the editing template is SEQ ID NO: 79.
10. The PEgRNA of claim 1, wherein the spacer is SEQ ID NO: 4, the nucleotide sequence of the editing template is SEQ ID NO: 62, and the nucleotide sequence of the PBS is Sequence Number: 21.
11. The PEgRNA of any one of claims 1-10, wherein the nucleotide sequence of the gRNA core is SEQ ID NO: 429.
12. The PEgRNA of any one of claims 1-11, comprising, from 5’ to 3’, the spacer, the gRNA core, and the extension arm in a contiguous sequence forming a single molecule.
13. The PEgRNA of any one of claims 1-12, further comprising a 3’ motif have a nucleotide sequence that is SEQ ID NO: 440.
14. The PEgRNA of claim 1, wherein:(a) wherein the one or more synonymous nucleotide transversion edits comprise c.1086G> C;and wherein the PEgRNA has nucleotide sequence SEQ ID NO: 629, 630, 117, 118, 120, 121, 122, 123, 125, 126, 127, 128, 129, 130, 131, 132, 134, 137, 138, 140, 141, 145, 146, 147, 148, 149, 152, 154, 157, 158, 159, 161, 162, 163, 167, 168, 169, 170, 171, 173, 174, 175, 176, 180, 181, 182, 185, 186, 187, 188, 193, 194, 195, 196, 197, 200, 202, 206, 207, 660, 208, 209, 210, 211, 213, 214, 216, 217, 219, 220, 222, 224, 226, 229, 230, 231, 232, 234, 238, 239, 241, 242, 243, 246, 248, 255, 256, 257, 259, 260, 262, 263, 265, 273, 274,WSGR Docket No. 59761-805.601275, 277, 279, 280, 285, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 302, 303, 304, 306, 308, 309, 310, 312, 313, 314, 316, 317, 318, 319, 320, 321, 322, 327, 330, 332, 333, 338, 339, 340, 342, 345, 347, 349, 350, 351, 353, 354, 356, 358, 360, 361, 362, 369, 370, 372, 373, 375, 376, 378, 380, 383, 387, 389, 390, 394, 396, 403, 407, 411, 412, 413, 646, 647, 660, 663, or 664;(b) wherein the one or more synonymous nucleotide transversion edits comprise c.lO86G> C and C.1092OA; and wherein the PEgRNA has nucleotide sequence SEQ ID NO: 629, 630, 137, 162, 206, 358, 663, or 664;(c) wherein the nucleotide sequence of the editing template encodes wildtype SERPINA1 gene sequence except for a c.1086G> C edit; and wherein the PEgRNA has nucleotide sequence SEQ ID NO: 117, 118, 126, 131, 134, 309, 310, 120, 122, 129, 140, 149, 312, 314, 123, 125, 145, 167, 170, 313, 330, 121, 127, 128, 138, 148, 157, 168, 169, 174, 175, 185, 186, 209, 242, 243, 288, 289, 290, 291, 292, 293, 294, 295, 297, 298, 300, 302, 303, 304, 306, 308, 316, 319, 321, 333, 338, 351, 130, 141, 180, 200, 202, 332, 342, 147, 173, 188, 208, 210, 340, 353, 132, 154, 182, 207, 660, 211, 216, 217, 275, 296, 354, 356, 361, 646, 647, 181, 194, 219, 222, 229, 360, 376, 232, or 375; or(d) wherein the nucleotide sequence of the editing template encodes wildtype SERPINA1 gene sequence except for c.lO86G> C and C.1092OA edits; and wherein the PEgRNA has nucleotide sequence SEQ ID NO: 629, 630, 206, 358, 663, or 664.
15. The PEgRNA of claim 14, wherein the nucleotide sequence of the editing template encodes wildtype SERPINA1 gene sequence except for c.lO86G> C and C.1092OA edits; and wherein the PEgRNA has nucleotide sequence SEQ ID NO: 629 or 630.
16. The PEgRNA of any one of claims 1-15, further comprising a nuclear localization signal (NLS) connected to the 3’ end of the PEgRNA via a maleimide -thiol conjugation as shown in(CH2)6-NH (Structure A).
17. The PEgRNA of claim 16, wherein the NLS has amino acid sequence SEQ ID NO: 470.
18. The PEgRNA of any one of claims 1-17, further comprising 3’ mN*mN*mN*mN modifications at the four 3’ most nucleotides and 5’mN*mN*mN* modifications at the three 5’ most nucleotides, where m indicates that the nucleotide contains a 2’-O-Me modification and a * indicates the presence of a phosphorothioate bond.
19. A prime editing system comprising:(a) the PEgRNA of any one of claims 1-18, or a nucleic acid encoding the PEgRNA; andWSGR Docket No. 59761-805.601(b) a prime editor, or a nucleic acid encoding the prime editor, wherein the prime editor comprises:i. a Cas9 domain capable of recognizing a NRTH PAM, wherein N is A, G, C, or T, R is A or G, and H is A, C, or T; andii. a reverse transcriptase.
20. The prime editing system of claim 19, wherein the Cas9 domain is a Cas9 nickase having amino acid sequence SEQ ID NO: 585.
21. The prime editing system of claim 19 or 20, wherein the reverse transcriptase has amino acid sequence SEQ ID NO: 620.
22. The prime editing system of any one of claims 19-21, wherein the prime editor is a fusion protein having amino acid sequence SEQ ID NO: 626.
23. The prime editing system of any one of claims 19-22, comprising the nucleic acid encoding the prime editor, wherein the nucleic acid is a messenger RNA (mRNA) that comprises, from 5’ to 3’:(a) a 5’ untranslated region (5 ’UTR);(b) an open reading frame (ORF) having a nucleotide sequence encoding the prime editor; (c) a 3’ untranslated region (3’ UTR); and(d) apoly-A tail.
24. The prime editing system of claim 23, wherein the mRNA comprises: the 5’ UTR having nucleotide sequence Sequence Number 631 or 632, the nucleotide sequence of the ORF that is SEQ ID NO: 639, the 3’ UTR having nucleotide sequence SEQ ID NO: 637, and the poly-A tail that is from 90 to 110 nucleotides long.
25. The prime editing system of claim 23 or claim 24, wherein 100% of the uridines in the mRNA are N-1 methylpseudouridines.
26. The prime editing system of any one of claims 19-25, further comprising (c) a ngRNA, or a nucleic acid encoding the ngRNA, wherein the ngRNA comprises: (i) a ngRNA spacer comprising a nucleotide sequence selected from the group consisting of SEQ ID NOs: 444-452; and (ii) a ngRNA core capable of binding the Cas9 protein.
27. The prime editing system of claim 26, wherein the nucleotide sequence of the ngRNA spacer is SEQ ID NO: 449.
28. The prime editing system of claim 26 or 27, wherein the ngRNA core comprises a nucleotide sequence of SEQ ID NO: 426.
29. The prime editing system of any one of claims 26-28, wherein the ngRNA has nucleic acid sequence SEQ ID NO: 421.WSGR Docket No. 59761-805.60130. The prime editing system of any one of claims 26-29, further comprising an NLS connected to the 3’ end of the ngRNA via a maleimide -thiol conjugation as shown in—(CH2)6-NH (Structure A).
31. The prime editing system of claim 30, wherein the NLS has amino acid sequence SEQ ID NO: 470.
32. The prime editing system of any one of claims 26-31, wherein the ngRNA further comprises 3’ mN*mN*mN*mN modifications at the four 3’ most nucleotides and 5’mN*mN*mN* modifications at the three 5’ most nucleotides, where m indicates that the nucleotide contains a 2’-O-Me modification and a * indicates the presence of a phosphorothioate bond.
33. A composition comprising a plurality of lipid nanoparticles (LNPs) encapsulating an RNA mixture comprising the prime editing system of any one of claims 19-32.
34. A composition comprising:(a) an RNA mixture comprising the prime editing system of any one of claims 26-32, the RNA mixture comprising:i. 30 wt% to 90 wt% of the mRNA;ii. 5 wt% to 55 wt% of the PEgRNA; andiii. 0 wt% to 35 wt% of the ngRNAwherein wt% is based on total RNA in the RNA mixture; encapsulated in(b) a plurality of LNPs comprising:i. 25 mol% to 60 mol% of an ionizable lipid, wherein the ionizable lipid is selected from the group consisting of:(Ionizable Lipid 2), andWSGR Docket No. 59761-805.
601. M. M'0(Ionizable Lipid 5);ii. 5 mol% to 45 mol% of a helper lipid, wherein the helper lipid is selected from the group consisting of 2-distearoyl-sn-glycero-3-phosphocholine (DSPC), 1,2- dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE), 1,2-dimyristoyl-sn- glycerophosphocholine (DMPC), l,2-diphytanoyl-sn-glycero-3- phosphoethanolamine (4ME 16.0 PE), and dimethyldioctadecylammonium (bromide salt) (18:0 DDAB);iii. 0.5 mol% to 3 mol% of a PEG-containing lipid, wherein the PEG-containing lipid is selected from the group consisting of C14-PEG2K, DMG-PEG2K, DMG- PEG1K, C14-PEG550, C14-PEG750, C18-PEG2K, DSG-PEG2K, DMG-pSar25, n-tetradecyl-pSar25, and n-tetamine-pSar25;iv. 15 mol% to 50 mol% of a sterol, wherein the sterol is selected from the group consisting of cholesterol, beta sitosterol, 25 -hydroxy cholesterol, and sitosterol; andv. 0.001 mol% to 5 mol%, of a GalNAc -based lipid, wherein the GalNAc -based lipid is selected from the group consisting of(GalNAc Lipid 1),(GalNAc Lipid 2),WSGR Docket No. 59761-805.601wherein mol% is based on the total lipid present in the LNP.
35. The composition of claim 34, wherein:(a) the RNA mixture comprises:i. 47.5 wt% to 52.5 wt% of the mRNA, wherein the mRNA comprises: the 5’ UTR having nucleotide sequence Sequence Number 632, the nucleotide sequence of the ORF that is SEQ ID NO: 639, the 3’ UTR having nucleotide sequence SEQ ID NO: 637, and the poly-A tail that is from 90 to 110 nucleotides long;ii. 35 wt% to 40 wt % of the PEgRNA, wherein the PEgRNA has the nucleotide sequence of SEQ ID NO: 630 and is conjugated on its 3’ end via the maleimide -thiol conjugation to the NLS having the amino acid sequence of SEQ ID NO: 470; andiii. 10 wt% to 15 wt% of the ngRNA, wherein the ngRNA has the nucleotide sequence of SEQ ID NO: 662 and is conjugated on its 3’ end via the maleimide -thiol conjugation to the NLS having the amino acid sequence of SEQ ID NO: 470; and(b) the plurality of LNPs comprise:i. 47.5 mol% to 52.5 mol% of the ionizable lipid, wherein the ionizable lipid is Ionizable Lipid 1;ii. 7.5 mol% to 12.5 mol% of the helper lipid, wherein the helper lipid is DSPC; iii. 2 mol% to 3 mol % of the PEG-containing lipid, wherein the PEG-containing lipid is DMG-PEG2K;iv. 35 mol% to 40 mol% of the sterol, wherein the sterol is cholesterol; andWSGR Docket No. 59761-805.601v. 0.1 mol% to 0.3 mol% of the GalNAc -based lipid, wherein the GalNAc -based lipid is GalNAc Lipid 1;wherein the plurality of LNPs have a particle size of 60-80 nm and wherein the composition has an N / P ratio of 4.4-4.8.
36. The composition of claim 34 or 35, wherein the PEgRNA and the ngRNA further comprise the 3’ mN*mN*mN*mN modifications at the four 3’ most nucleotides and the 5’mN*mN*mN* modifications at the three 5’ most nucleotides, where m indicates that the nucleotide contains a 2’-O-Me modification and a * indicates the presence of a phosphorothioate bond.
37. The composition of any one of claims 34-36, wherein 100% of the uridines in the mRNA are N-l methylpseudouridine s.
38. The composition of any one of claims 34-37, wherein the mol% of each lipid component is calculated from mass values determined by reversed-phase ultra-performance liquid chromatography (RP-UPLC) coupled with charged aerosol detection (CAD).
39. A pharmaceutical composition comprising the composition of any one of claims 33-38 formulated in a buffer.
40. A method of editing a SERPINA1 gene in a cell, the method comprising contacting the cell with: (a) the prime editing system of any one of claims 19-32, (b) the composition of any one of claims 33-38, or (c) the pharmaceutical composition of claim 39.
41. The pharmaceutical composition of claim 39, for use as a medicament.
42. A method of treating Alpha- 1 Antitrypsin Deficiency (AATD) in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of the pharmaceutical composition of claim 39.
43. The method of claim 42, wherein the pharmaceutical composition is administered via intravenous infusion.
44. A method of making the composition of any one of claims 33-38, comprising:(a) preparing an aqueous solution comprising the RNA mixture;(b) preparing a lipid solution comprising the ionizable lipid, the helper lipid, the PEG- containing lipid, the sterol, and the GalNAc-based lipid in an organic solvent;(c) combining the lipid solution and the aqueous solution under mixing conditions sufficient to form lipid nanoparticles (LNPs) encapsulating the RNA mixture;(d) subjecting the LNPs to tangential flow filtration (TFF);(e) optionally adding a cryoprotectant to the LNPs;(f) optionally sterile filtering the LNPs; and(g) aseptically filling the lipid nanoparticles into a container for storage.
45. The method of claim 44, wherein:(a) preparing the aqueous solution comprising the RNA mixture comprises combining the mRNA, the pegRNA, and the ngRNA at a weight ratio of about 4:3:1 in the aqueousWSGR Docket No. 59761-805.601solution that has a pH of 5.8-6.3 and that comprises 20 mM to 60mM of a phosphate buffer, a Tris buffer, or a combination thereof;(b) the organic solvent for the lipid solution is anhydrous ethanol;(c) the lipid solution and the aqueous solution are combined using a T-mixer;(d) the LNPs are subjected to TFF using a hollow fiber membrane with a molecular weight cut-off (MWCO) ranging from 30 kDa to 1000 kDa to exchange the organic solvent with an aqueous Tris buffer having a pH of 7.0-8.0;(e) the cryoprotectant is sucrose and is added to the LNPs to a concentration of 6% to 10% (w / v);(f) the LNPs are sterile filtered through a 0.22 pm membrane filter; and(g) the container for storage is a borosilicate glass vial.
46. The method of claim 44 or 45, wherein the lipid nanoparticles are stored at a temperature of -65 °C or lower.