Systems and methods for regulating target genes

EP4713342A2Pending Publication Date: 2026-03-25EPICRISPR BIOTECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-15
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Current methods for regulating target genes lack efficiency in activating or modulating gene expression, particularly in therapeutic contexts, due to limitations in identifying effective peptide combinations and protein engineering challenges.

Method used

Development of engineered gene effectors comprising specific peptides and fusion proteins, including CRISPR/Cas endonucleases, that can activate target genes by forming complexes with guide nucleic acids and targeting specific loci, utilizing computer-implemented methods to optimize peptide sequences for enhanced activity.

Benefits of technology

The engineered gene effectors demonstrate strong activation potential, with a significant hit rate in activating genetic reporters and endogenous loci, offering a time- and cost-effective approach to discovering functional biological sequences for therapeutic applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024029475_21112024_PF_FP_ABST
    Figure US2024029475_21112024_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are compositions, methods, and systems for modulating expression of target genes (e.g., target endogenous genes). In some embodiment, engineered gene effectors that include a first peptide of 75-95 amino acids in length and a second peptide of 75-95 amino acids in length are provided. The engineered gene effectors can facilitating modulation of an expression or activity level of the target gene when brought in proximity with a target gene or target gene regulatory sequence in a complex with a targeting moiety, such as a heterologous endonuclease. Also provided are a computer-implemented methods for producing functional biological sequences, and functional biological sequences, such as an engineered gene effector made by the methods.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR REGULATING TARGET GENES REFERENCE TO PRIORITY APPLICATIONS

[0001] This application claims priority to U.S. provisional application numbers: 63 / 504661, filed May 26, 2023; 63 / 520251, filed August 17, 2023; 63 / 502891, filed May 17, 2023; 63 / 504660, filed May 26, 2023; and 63 / 504663, filed May 26, 2023. The content of each of these aforementioned applications is expressly incorporated herein by reference in its entirety. REFERENCE TO SEQUENCE LISTING

[0002] The present application is being filed along with a Sequence Listing in electronic format. The Sequence Listing is provided as a file entitled EPICR019WOSequenceListing.xml, created May 14, 2024, which is 2,506,752 bytes in size. The information in the electronic format of the Sequence Listing is incorporated by reference in its entirety. BACKGROUND

[0003] Various effectors (e.g., transcriptional regulators) can be utilized to regulate expression or activity of a target gene in the cell. For example, a heterologous gene effector can be introduced (e.g., delivered, expressed, etc.) to the cell, and the heterologous gene effector, either alone or along with an additional agent, can effect such regulation of the target gene. In some examples, the additional agent can comprise a clustered regularly interspaced short palindromic repeats (CRISPR) / CRISPR-associated protein (Cas) for specifically binding to the target gene (e.g., a target deoxyribonucleic acid (DNA) sequence or ribonucleic acid (RNA) sequence (e.g., foreign DNA sequence or RNA sequence) of the target gene), while the heterologous gene effector can regulate expression or activity level of the target gene. Such gene effectors can be utilized, e.g., as gene therapy to treat or ameliorate a condition (e.g., a disease) of a subject. SUMMARY

[0004] Provided herein is an engineered gene effector comprising a polypeptide comprising: a first peptide of 75-110 (or 75-95) amino acids in length, wherein the first peptide comprises any one of SEQ ID NOs:3-100, or a sequence at least 85% identical thereto; and a second peptide of 75-110 (or 75-95) amino acids in length and that is heterologous to the first 1peptide, wherein the second peptide comprises any one of SEQ ID NOs:3-100, or a sequence at least 85% identical thereto, optionally wherein the first peptide is different from the second peptide.

[0005] Also provided is an engineered gene effector comprising a polypeptide comprising: a first peptide comprising an amino acid sequence of 75-110 amino acids in length and based on a human or viral transcriptional modulator; and a second peptide comprising an amino acid sequence of 75-110 amino acids in length and based on a human or viral transcriptional modulator, wherein the second peptide is heterologous to the first peptide, wherein the engineered gene effector is capable of activating a target gene in a cell when the engineered gene effector is expressed therein and effectively targeted to a locus of the target gene.

[0006] Provided herein is a fusion protein comprising: the engineered gene effector of the present disclosure; and a heterologous endonuclease coupled to the polypeptide, optionally wherein the heterologous endonuclease is a Cas protein.

[0007] Also provide is a system comprising: the engineered gene effector of the present disclosure; a heterologous endonuclease coupled to the polypeptide of the engineered gene effector, optionally wherein the heterologous endonuclease is a Cas protein; and a guide nucleic acid capable of forming a complex with the heterologous endonuclease, wherein the complex exhibits specific binding to a target gene in a cell when the system is expressed therein.

[0008] Also provided is a combination of polynucleotides encoding the system of the present disclosure, wherein the combination of polynucleotides is configured to express the heterologous endonuclease coupled to the engineered gene effector and the guide nucleic acid in a cell.

[0009] Further provided is a kit that includes any of the engineered gene effector, fusion protein, combination, system, polynucleotide, vector, and / or cell of the present disclosure.

[0010] Also provided is a method of controlling a target gene in a cell, comprising contacting a cell with the engineered gene effector, the fusion protein, the polynucleotide, the vector, the system, or the combination of polynucleotides of the present disclosure. 2

[0011] Further provided is a computer-implemented method of producing functional biological sequences, comprising: (a) providing a fitness function trained on a biological dataset comprising functionally defined biological sequences with a fixed length; (b) providing, in a computer, a plurality of different sequences comprising a fixed length, each sequence associated with a temperature and a fitness based on said fitness function, wherein each sequence is associated with a different temperature of a temperature ladder; (c) by said computer, in parallel across the plurality of different sequences: (1) selecting one or more random position(s) for introducing a substitution in one or more sequences of the plurality of different sequences, optionally selecting 1-5 random positions, optionally selecting 1 random position; and for each of the one or more sequences, evaluating a first fitness change due to introduction of the substitution(s) at the one or more randomly selected position(s), and accepting or rejecting the substitution(s) based on the evaluated first fitness change, and optionally further based on the temperature associated with the sequence; and / or (2) selecting one or more pairs of said plurality of different sequences, each selected pair comprising sequences associated with consecutive temperatures of the temperature ladder, optionally selecting up to 3 pairs of said plurality of different sequences, optionally selecting one pair of said plurality of different sequences; and for each of the selected pairs: selecting one or more domains for swapping between sequences of the selected pair; and evaluating a fitness difference of the sequences of the selected pair due to swapping of the one or more domains, and accepting or rejecting the swapping of the one or more domains between said selected pair based on the fitness difference and the temperature associated with each sequence of said selected pair; and (d) performing (c) iteratively, wherein in each subsequent iteration, the accepted substitution(s) of a preceding iteration and / or the accepted swapping of domains of a preceding iteration are incorporated into the plurality of different sequences, thereby producing one or more functional sequences having a fitness at or above a desired fitness threshold.

[0012] In one aspect, a computer-implemented method of producing functional biological sequences is provided, the computer-implemented method comprising: (a) evaluating, by a computer, a sequence of a plurality of different sequences comprising a fixed length based on a fitness function trained on a biological dataset comprising functional biological sequences with the fixed length; (b) substituting, by the computer, one or more random residues in the sequence, to generate a mutated sequence; (c) evaluating, by the 3computer, the mutated sequence based on the fitness function; and (d) collecting, by the computer, functional sequences accepted by the fitness function. In some cases, the functional biological sequences comprise amino acid or nucleotide sequences of a protein or a peptide. In some cases, the protein or the peptide is an epigenetic modulator, a transcription factor, an enzyme, a nuclease, an agonist, an antagonist, a regulator, or an inhibitor. In some cases, the functional biological sequences comprise amino acid sequences or nucleotide sequences. In some cases, the functional biological sequences comprise amino acid sequences and further wherein the fixed length is at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 105, at least 110, at least 115, at least 120, at least 125, at least 130, at least 140, at least 150, or at least 200 amino acids. In some cases, the fitness function is based on one or more machine learning models, wherein the machine learning model is selected from the group consisting of: a supervised machine learning model, an unsupervised machine learning model, a reinforcement learning model, a deep learning model, a transfer learning model, and any combination thereof. In some cases, the one or more machine learning models is selected from the group consisting of: a classification model, a regression model, a convolutional neural network (CNN), a recurrent neural network (RNN), Extreme Gradient Boosting (XGBoost), long short-term memory network, generative adversarial network (GAN), an autoencoder, a transformer network, evolutionary Monte Carlo, and any combination thereof. In some cases, the computer-implemented method further comprises randomly swapping, by the computer, one or more subsequences from the mutated sequence and a different sequence of the plurality of sequences. In some cases, the fitness function comprises a threshold selected from the group consisting of: a binary threshold, a numerical threshold, a multiclass threshold, a confidence threshold, a decision threshold, and any combination thereof. In some cases, functional sequences are accepted by the fitness function when a fitness score assigned to the functional sequences by the fitness function reaches or exceeds the threshold.

[0013] In another aspect, a computer-implemented system is provided comprising a computing device comprising at least one processor and instructions executable by the at least one processor to provide an application comprising: (a) a software module configured to evaluate, by a computer, a sequence of a plurality of different sequences comprising a fixed 4length based on a fitness function trained on a biological dataset comprising functional biological sequences with the fixed length; (b) a software module configured to substitute, by the computer, one or more random residues in the sequence, to generate a mutated sequence; (c) a software module configured to evaluate, by the computer, the mutated sequence based on the fitness function; and (d) a software module configured to collect, by the computer, functional sequences accepted by the fitness function.

[0014] In another aspect, a non-transitory computer-readable medium is provided having stored thereon computer-readable instructions that, when executed by a processor, cause the processor to execute a method comprising: (a) evaluating a sequence of a plurality of different sequences comprising a fixed length based on a fitness function trained on a biological dataset comprising functional biological sequences with the fixed length; (b) substituting one or more random residues in the sequence, to generate a mutated sequence; (c) evaluating the mutated sequence based on the fitness function; and (d) collecting functional sequences accepted by the fitness function.

[0015] Provided herein is an engineered gene effector comprising a polypeptide of 85 amino acids in length that comprises any one of SEQ ID NOs: 1495, 1592, 1595, 1634, 1654, 1665, 1677, 1686, 1689, 1716, or a sequence at least 85% identical thereto.

[0016] Also provided is a computer-implemented system comprising a computing device comprising at least one processor and instructions executable by the at least one processor to provide an application, comprising: (a) a software module configured to provide, by a computer, a fitness function trained on a biological dataset comprising functionally defined biological sequences with a fixed length; (b) a software module configured to provide, by said computer, a plurality of different sequences comprising a fixed length, each sequence associated with a temperature and a fitness based on said fitness function, wherein each sequence is associated with a different temperature of a temperature ladder; (c) in parallel across the plurality of different sequences: (1) a software module configured to select, by said computer, one or more random position(s) for introducing a substitution in one or more sequences of the plurality of different sequences, optionally selecting 1-5 random positions, optionally selecting 1 random position; and for each of the one or more sequences, evaluate a first fitness change due to introduction of the substitution(s) at the one or more randomly selected position(s), and accept or reject the substitution(s) based on the evaluated first fitness 5change, and optionally further based on the temperature associated with the sequence; and / or (2) a software module configured to select, by said computer, one or more pairs of said plurality of different sequences, each selected pair comprising sequences associated with consecutive temperatures of the temperature ladder, optionally selecting up to 3 pairs of said plurality of different sequences, optionally selecting one pair of said plurality of different sequences; and for each of the selected pairs: select one or more domains for swapping between sequences of the selected pair; and evaluate a fitness difference of the sequences of the selected pair due to swapping of the one or more domains, and accept or reject the swapping of the one or more domains between said selected pair based on the fitness difference and the temperature associated with each sequence of said selected pair; and (d) a software module configured to perform, by said computer, (c) iteratively, wherein in each subsequent iteration, the accepted substitution(s) of a preceding iteration and / or the accepted swapping of domains of a preceding iteration are incorporated into the plurality of different sequences, thereby producing one or more functional sequences having a fitness at or above a desired fitness threshold.

[0017] Also provided is an engineered gene effector comprising a polypeptide comprising any one of SEQ ID NOs: 115-274, 276-287, 289-367, 370-445, 448-746, 748-777, 779-929, 931-1007, 1009-1156, 1158-1194, 1196-1288, 1290-1350, 1352-1451, or a sequence at least 85% identical thereto, or with 0-3 amino acid residue mutations, optionally wherein any mutations thereof are conservative substitutions.

[0018] Provided herein is an engineered gene effector comprising a polypeptide comprising any one of SEQ ID NOs: 1085, 122, 1084, 653, 1099, and 1107, or a sequence at least 85% identical thereto, or with 0-3 amino acid residue mutations, optionally wherein any mutations thereof are conservative substitutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The novel features of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which:

[0020] FIG. 1A demonstrates the activation of a GFP reporter via single gene modulator or combinatorial gene modulators. 6

[0021] FIG. 1B shows the activation synergy of the combinatorial gene modulators.

[0022] FIG. 2A shows a distribution of curated candidate gene modulators from the screenings.

[0023] FIG. 2B shows the activation level of the combinatorial gene modulators (weak, medium, and strong activators).

[0024] FIG. 2C shows the homology clustering of the candidate modulators.

[0025] FIG.2D shows the predicted protein domains of the candidate modulators.

[0026] FIG. 3A shows the cloning strategy used to screen combinatorial gene modulators.

[0027] FIG. 3B shows a GFP reporter design to assess gene modulation by the combinatorial gene modulators.

[0028] FIG.3C shows the validation of reversible recruitment of the combinatorial gene modulators by the reporter system.

[0029] FIG. 4 shows GFP expression of the combinatorial gene modulators.

[0030] FIG. 5A demonstrates DESEQ2 analysis of the combinatorial gene modulators.

[0031] FIG. 5B shows the activation level by the combinatorial gene modulator source (e.g., human-human vs. viral-viral).

[0032] FIG. 5C shows a heatmap indicating activation status of all combinatorial modulators.

[0033] FIG. 6A shows an analysis of biochemical and biophysical features of the combinatorial modulators by assessing electrostatic potential

[0034] FIG. 6B. shows an analysis of biochemical and biophysical features of the combinatorial modulators by assessing the B_BetaFactor.

[0035] FIGs.7A-7B depict activation of an epigenetically quiescent locus (CD45) by engineered gene effectors in HEK293. Some engineered gene effectors assayed comprise a first peptide and a second peptide. Some engineered gene effectors assayed were encoded from functional biological sequences generated using computer-implemented methods.

[0036] FIG. 8 depicts reactivation of epigenetically silenced loci by engineered gene effectors in human cells. Some engineered gene effectors assayed comprise a first peptide 7and a second peptide. Some engineered gene effectors assayed were generated using computer- implemented methods of producing functional biological sequences.

[0037] FIGs. 9A-9D provide a non-limiting example of an exemplary method of generating biologically functional sequences according to embodiments of the disclosure.

[0038] FIG.10 depicts a plot of the minimum edit distance of sequences predicted to have biological function using the methods and systems provided herein, relative to the sequences in the original training data set.

[0039] FIG. 11 depicts a representation of sequences predicted to have biological function using the methods and systems provided herein in two-dimensional space.

[0040] FIG. 12A depicts a non-limiting example of a transfer learning approach for predicting gene activators from sequence using LPLM embeddings.

[0041] FIG. 12B depicts a non-limiting example of a fitness landscape diagram representing the effective search spaces of MHMCS (pink) and EMCS (blue). MHMCS locally searches near the starting molecule for area of high fitness, while EMCS interpolates between the starting points as well as varying search speed to optimize the search.

[0042] FIG. 13 depicts Principal Component Analysis (PCA) of an original training set with novel sequences designed by EMCS and MHMCS using OneHot encoding.

[0043] FIG.14A depicts entropy change distribution for 107 MHMCS and EMCS primary iterations using default parameters. From an information perspective, EMCS explores a larger region of the fitness space per iteration compared to MHMCS.

[0044] FIG. 14B depicts: i) Iterations to convergence (f ^ 0.95) starting from random pre-defined sequences for 2361 sequences obtained by 2571 MHMCS runs at T = 2.5 × 10í3, 10í4; ii) Iterations to convergence for 1171 sequences obtained by 1171 EMCS runs under default parameters; iii) Number of iterations to arrive at convergence and / or positive hits (f ^ 0.5) for 2571 sequences from 2571 MHMCS runs at T = 2.5 × 10í3, 10í4. 210 of the 2571 sequences failed to reach convergence (but succeeded in yielding positive hits (f > 0.5)); iv) Iterations to arrive at convergence and / or positive hits for 2720 sequences obtained by 1171 EMCS runs under default parameters, yielding an average of 2.32 positive hits per EMCS run of 4 chains.

[0045] FIG. 15 depicts a schematic of fitness landscape exploration using novel EMCS, which includes domain swapping between peptide chains. The EMCS algorithm 8comprises the following components: 1) Parallel Metropolis-Hastings Monte Carlo (MHMCS) runs; 2) Temperature Ladder implementation (Parallel Tempering); and 3) Domain swapping between peptide chains run in parallel (EMCS).

[0046] FIGs. 16A-16B depict entropy change distribution and convergence times for MHMCS and EMCS iterations. PTP (Parallel Tempering) and EMC-NPT (EMCS without parallel tempering) were run for ablation studies.

[0047] FIG. 17 depicts Fluorescence Activated Cell Sorting (FACS) histograms from the engineered gene effector validation experiments. Engineered gene effectors encoded from the 4600 novel sequences were assayed for their ability to activate a synthetic genetic locus. 357 of the 4600 engineered gene effectors (7.51% hit rate) significantly activated the genetic reporter over background florescence.

[0048] FIGs. 18A-18C depict biochemical and structural property analysis of experimentally validated functional biological sequences using ESMFold.

[0049] FIGs. 19A-19B depict engineered gene effector validation assays in HEK293 cells. 10 engineered gene effectors (SEQ ID NO: 1495, 1592, 1595, 1634, 1654, 1665, 1677, 1686, 1689, 1716) were individually screened at synthetic (TRE3G) and endogenous loci (CD45). The activation potency of the engineered gene effectors was compared to the activation potency of standard activators, VP64 and vCD.

[0050] FIG. 20 depicts a non-limiting computer system that is programmed or otherwise configured to implement methods provided herein.

[0051] FIG. 21A and 21B are schematic diagrams showing non-limiting embodiments of computer-implemented methods of the present disclosure. DETAILED DESCRIPTION OVERVIEW

[0052] Various aspects of the present disclosure can provide engineered effectors (or engineered gene effectors, as used interchangeably herein) capable of regulating (e.g., activating or reducing) expression or activity level of a target gene in a cell (e.g., an endogenous target gene, a heterologous target gene, etc.), compositions, combinations, systems thereof, and methods of use thereof. Such engineered effectors can work in conjunction with a heterologous endonuclease (e.g., engineered CRISPR / Cas nuclease, or a deactivated variant thereof) to, for example, effect manipulation of the expression or activity level of the target 9gene in a cell, e.g., to treat or ameliorate a condition (e.g., a disease) of a subject. Gene expression can underpin various physiological and pathological effects in cells and tissues, contributing to many diseases and conditions, and thus compositions, combinations, systems, and methods utilizing the engineered gene effectors of the present disclosure can modulate expression of specific genes in a desirable way to have therapeutic benefit.

[0053] CRISPR-mediated transcriptional regulation has broad potential applications in synthetic biology and gene therapy. Transcriptional activation in eukaryotes depends on the complex interplay of multiple factors, including DNA-binding transcription factors, co-activators, chromatin remodelers, and basal transcriptional machinery. These multiple inputs often exhibit a synergistic relationship, whereby the transcriptional output driven by two or more factors is greater than the sum of output driven by each factor individually. Initial screening has discovered hundreds of small (85aa) peptides derived from endogenous human, viral, and archaeal proteins that activate transcription to varying degrees when fused to a programmable DNA-binding dCas protein.

[0054] Described herein is a platform to screen pairwise combinations of these activation domains in an inducible and reversible manner for improved activity as determined by both magnitude and duration of transcriptional output. The screen identified ~1400 novel combinations, and these combinatorial activators exhibit stronger activity than their constituent parts. These combinatorial activators are strongest when one of the partner tiles is of viral origin. Further described herein is a machine learning approach to identify a suite of partner tiles that are most predictive of strong combinatorial activators including canonical viral activators (e.g. VP64), novel viral activators (e.g. vIRF2 / vIRF4), and human activators (e.g. LEUTX).

[0055] Provided herein is an analysis of biochemical and biophysical features demonstrating that strong combinations have highly negative electrostatic potential and tend to have structural flexibility. In silico structural predictions of top hits showed that stabilizing intramolecular interactions between helices can stabilize the interaction interface of strong combinations and can be functionally relevant for transcriptional activation. Further provided herein is a novel toolkit of combinatorial transcriptional regulators with as-yet-unexplored potential, as well as a platform for the discovery of additional tools with applications in both basic research and therapeutics. 10

[0056] In some embodiments, engineered gene effectors of the present disclosure include a fusion of two different gene effectors that provide for regulation of gene expression, e.g., when used in conjunction with a heterologous endonuclease. In some embodiments, the regulation of gene expression achieved by the fusion of two different gene effectors provides superior effects (e.g., prolonged duration of regulated expression) relative to the effect provided by each component alone.

[0057] Designing novel protein sequences remains a slow and expensive process due to a variety of protein engineering challenges; in particular, the number of protein variants that can be experimentally tested in a given assay pales in comparison to the vastness of the overall sequence space which results in low hit rates and expensive wet lab testing cycles. Provided herein are computer-implemented methods and systems for producing a functional biological sequence, such as an engineered gene effector. In some embodiments, the computer-implemented methods and systems accelerate the discovery of new functional sequences, such as new functional amino acid sequences. In some embodiments, the computer-implemented methods and systems produce novel biologically functional sequences in a time and cost-effective fashion. TERMS

[0058] The term “heterologous,” when used herein with reference to a polypeptide sequence or a nucleic acid sequence, indicates that the polypeptide sequence or the nucleic acid sequence is (1) disposed (e.g., in an environment, such as a cell, a virus, or a fusion polypeptide molecule or a fusion polynucleotide molecule) where it is not normally found (e.g., not normally found in nature); or (2) comprises two or more subsequences that are not found in the same relationship to each other as normally found in nature. For example, a polypeptide can comprise a first polypeptide sequence and a second polypeptide sequence that are not found together in a single polypeptide in nature, and thus the first polypeptide sequence and the second polypeptide sequence can be heterologous to each other. In another example, a polynucleotide can comprise a first polynucleotide sequence and a second polynucleotide sequence that are not found together in a single polynucleotide in nature, and thus the first polynucleotide sequence and the second polynucleotide sequence can be heterologous to each other. 11

[0059] The term “cell” generally refers to a biological cell. A cell can be the basic structural, functional and / or biological unit of a living organism. A cell can originate from any organism having one or more cells. Some non-limiting examples include: a prokaryotic cell, eukaryotic cell, a bacterial cell, an archaeal cell, a cell of a single-cell eukaryotic organism, a protozoa cell, a cell from a plant (e.g. cells from plant crops, fruits, vegetables, grains, soy bean, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkin, hay, potatoes, cotton, cannabis, tobacco, flowering plants, conifers, gymnosperms, ferns, clubmosses, hornworts, liverworts, mosses), an algal cell, (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens C. Agardh, and the like), seaweeds (e.g. kelp), a fungal cell (e.g., a yeast cell, a cell from a mushroom), an animal cell, a cell from an invertebrate animal (e.g. fruit fly, cnidarian, echinoderm, nematode, etc.), a cell from a vertebrate animal (e.g., fish, amphibian, reptile, bird, mammal), a cell from a mammal (e.g., a pig, a cow, a goat, a sheep, a rodent, a rat, a mouse, a non-human primate, a human, etc.), and etcetera. Sometimes a cell is not originating from a natural organism (e.g. a cell can be a synthetically made, sometimes termed an artificial cell).

[0060] The term “nucleotide,” as used herein, generally refers to a base-sugar- phosphate combination. A nucleotide can comprise a synthetic nucleotide. A nucleotide can comprise a synthetic nucleotide analog. Nucleotides can be monomeric units of a nucleic acid sequence (e.g. deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide can include ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP) and deoxyribonucleoside triphosphates such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives can include, for example, [ĮS]dATP, 7-deaza-dGTP and 7-deaza-dATP, and nucleotide derivatives that confer nuclease resistance on the nucleic acid molecule containing them. The term nucleotide as used herein can refer to dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Illustrative examples of dideoxyribonucleoside triphosphates can include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. A nucleotide may be unlabeled or detectably labeled by well- known techniques. Labeling can also be carried out with quantum dots. Detectable labels can include, for example, radioactive isotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels and enzyme labels. Fluorescent labels of nucleotides may include but 12are not limited fluorescein, 5-carboxyfluorescein (FAM), 2ƍ7ƍ-dimethoxy-4ƍ5-dichloro-6- carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,Nƍ,Nƍ-tetramethyl-6- carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4- (4ƍdimethylaminophenylazo) benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, Cyanine and 5-(2ƍ-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS). Specific examples of fluorescently labeled nucleotides can include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G] dCTP, [TAMRA] dCTP, [JOE] ddATP, [R6G] ddATP, [FAM] ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP available from Perkin Elmer, Foster City, Calif. FluoroLink DeoxyNucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP available from Amersham, Arlington Heights, Ill.; Fluorescein-15-dATP, Fluorescein-12-dUTP, Tetramethyl-rodamine- 6-dUTP, IR770-9-dATP, Fluorescein-12-ddUTP, Fluorescein-12-UTP, and Fluorescein-15-2ƍ- dATP available from Boehringer Mannheim, Indianapolis, Ind.; and Chromosome Labeled Nucleotides, BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY- TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, fluorescein-12-UTP, fluorescein-12-dUTP, Oregon Green 488-5- dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, tetramethylrhodamine-6-UTP, tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12- dUTP available from Molecular Probes, Eugene, Oreg. Nucleotides can also be labeled or marked by chemical modification. A chemically-modified single nucleotide can be biotin- dNTP. Some non-limiting examples of biotinylated dNTPs can include, biotin-dATP (e.g., bio- N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g. biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).

[0061] The term “polynucleotide,” “oligonucleotide,” or “nucleic acid,” as used interchangeably herein, generally refers to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, either in single-, double-, or multi- stranded form. A polynucleotide can be exogenous or endogenous to a cell. A polynucleotide can exist in a cell-free environment. A polynucleotide can be a gene or fragment thereof. A polynucleotide can be DNA. A polynucleotide can be RNA. A polynucleotide can have any three dimensional structure, and can perform any function, known or unknown. A 13polynucleotide can comprise one or more analogs (e.g. altered backbone, sugar, or nucleobase). If present, modifications to the nucleotide structure can be imparted before or after assembly of the polymer. Some non-limiting examples of analogs include: 5-bromouracil, peptide nucleic acid, xeno nucleic acid, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, florophores (e.g. rhodamine or fluorescein linked to the sugar), thiol containing nucleotides, biotin linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudourdine, dihydrouridine, queuosine, and wyosine. Non- limiting examples of polynucleotides include coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. The sequence of nucleotides can be interrupted by non-nucleotide components.

[0062] The term “sequence identity” generally refers to an exact nucleotide-to- nucleotide or amino acid-to-amino acid correspondence of two polynucleotides or polypeptide sequences, respectively. Typically, techniques for determining sequence identity include determining the nucleotide sequence of a polynucleotide and / or determining the amino acid sequence encoded thereby, and comparing these sequences to a second nucleotide or amino acid sequence. Two or more sequences (polynucleotide or amino acid) can be compared by determining their “percent identity.” The percent identity of two sequences, whether nucleic acid or amino acid sequences, is the number of exact matches between two aligned sequences divided by the length of the longer sequence and multiplied by 100. Percent identity may also be determined, for example, by comparing sequence information using the advanced BLAST computer program, including version 2.2.9, available from the National Institutes of Health. The BLAST program is based on the alignment method of Karlin and Altschul, Proc. Natl. Acad. Sci. USA, 87:2264-2268 (1990) and as discussed in Altschul, et al., J. Mol. Biol., 215:403-410 (1990); Karlin And Altschul, Proc. Natl. Acad. Sci. USA, 90:5873-5877 (1993); and Altschul et al., Nucleic Acids Res., 25:3389-3402 (1997). The program may be used to 14determine percent identity over the entire length of the proteins being compared. Default parameters are provided to optimize searches with short query sequences in, for example, with the blastp program. The program also allows use of an SEG filter to mask-off segments of the query sequences as determined by the SEG program of Wootton and Federhen, Computers and Chemistry 17:149-163 (1993). Ranges of desired degrees of sequence identity are approximately 50% to 100% and integer values therebetween. In general, this disclosure encompasses sequences with at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 98% sequence identity with any sequence provided herein.

[0063] The term “gene” generally refers to a nucleic acid (e.g., DNA such as genomic DNA and cDNA) and its corresponding nucleotide sequence that is involved in encoding an RNA transcript. The term as used herein with reference to genomic DNA includes intervening, non-coding regions as well as regulatory regions and can include 5ƍ and 3ƍ ends. In some uses, the term encompasses the transcribed sequences, including 5ƍ and 3ƍ untranslated regions (5ƍ-UTR and 3ƍ-UTR), exons and introns. In some genes, the transcribed region will contain “open reading frames” that encode polypeptides. In some uses of the term, a “gene” comprises only the coding sequences (e.g., an “open reading frame” or “coding region”) necessary for encoding a polypeptide. In some cases, genes do not encode a polypeptide, for example, ribosomal RNA genes (rRNA) and transfer RNA (tRNA) genes. In some cases, the term “gene” includes not only the transcribed sequences, but in addition, also includes non- transcribed regions including upstream and downstream regulatory regions, enhancers and promoters. A gene can refer to an “endogenous gene” or a native gene in its natural location in the genome of an organism. A gene can refer to an “exogenous gene” or a non-native gene. A non-native gene can refer to a gene not normally found in the host organism, but which is introduced into the host organism by gene transfer. A non-native gene can also refer to a gene not in its natural location in the genome of an organism. A non-native gene can also refer to a naturally occurring nucleic acid or polypeptide sequence that comprises mutations, insertions and / or deletions (e.g., non-native sequence).

[0064] The term “deletion” generally refers to the removal (or loss) of one or more (or a specified number of) amino acids (e.g., contiguous or non-contiguous amino acids) from a polypeptide sequence, or the removal (or loss) one or more (or a specified number of) nucleic 15acid bases (e.g., contiguous or non-contiguous nucleic acid bases) from a polynucleotide sequence (e.g., that encodes the polypeptide sequence. The term “internal deletion” generally refers to a deletion that does not include the N- or C-terminus of a polypeptide or the 5ƍ or 3ƍ end of a polynucleotide. A deletion (e.g., an internal deletion) can be identified by comparing to a reference sequence, e.g., by specifying the start and end positions of the deletion relative to the reference sequence. A deletion (e.g., an internal deletion) is different and distinct from a substitution. For example, deletion of at least one amino acid is not followed by an insertion of at least one different amino acid at the same position as the at least one amino acid as compared to a reference polypeptide sequence, such that the size (e.g., a number of the amino acid residue(s)) of a modified (or engineered) polypeptide sequence comprising the deletion of the at least one amino acid is smaller than the reference polypeptide sequence by the size of the at least one amino acid that has been deleted.

[0065] The term “expression” generally refers to one or more processes by which a polynucleotide is transcribed from a DNA template (such as into an mRNA or other RNA transcript) and / or the process by which a transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as “gene product.” If the polynucleotide is derived from genomic DNA, expression can include splicing of the mRNA in a eukaryotic cell. “Up-regulated,” with reference to expression, generally refers to an increased expression level of a polynucleotide (e.g., RNA such as mRNA) and / or polypeptide sequence relative to its expression level in a wild-type state while “down-regulated” generally refers to a decreased expression level of a polynucleotide (e.g., RNA such as mRNA) and / or polypeptide sequence relative to its expression in a wild- type state. Expression of a transfected gene can occur transiently or stably in a cell. During “transient expression” the transfected gene is not transferred to the daughter cell during cell division. Since its expression is restricted to the transfected cell, expression of the gene is lost over time. In contrast, stable expression of a transfected gene can occur when the gene is co- transfected with another gene that confers a selection advantage to the transfected cell. Such a selection advantage may be a resistance towards a certain toxin that is presented to the cell.

[0066] The term “expression profile” generally refers to quantitative (e.g., abundance) and qualitative expression of one or more genes in a sample (e.g., a cell). The one or more genes can be expressed and ascertained in the form of a nucleic acid molecule (e.g., 16an mRNA or other RNA transcript). Alternatively or in addition to, the one or more genes can be expressed and ascertained in the form of a polypeptide (e.g., a protein measured via Western blot). An expression profile of a gene may be defined as a shape of an expression level of the gene over a time period (e.g., at least or up to about 1 hour, at least or up to about 2 hours, at least or up to about 3 hours, at least or up to about 4 hours, at least or up to about 5 hours, at least or up to about 6 hours, at least or up to about 7 hours, at least or up to about 8 hours, at least or up to about 9 hours, at least or up to about 10 hours, at least or up to about 11 hours, at least or up to about 12 hours, at least or up to about 16 hours, at least or up to about 18 hours, at least or up to about 24 hours, at least or up to about 36 hours, at least or up to about 48 hours, at least up to about 3 days, at least up to about 4 days, at least up to about 5 days, at least up to about 6 days, at least up to about 7 days, at least up to about 8 days, at least up to about 9 days, at least up to about 10 days, at least up to about 11 days, at least up to about 12 days, at least up to about 13 days, at least up to about 14 days, etc.). Alternatively, an expression profile of a gene may be defined as an expression level of the gene at a time point of interest (e.g., the expression level of the gene measured at least or up to about 1 hour, at least or up to about 2 hours, at least or up to about 3 hours, at least or up to about 4 hours, at least or up to about 5 hours, at least or up to about 6 hours, at least or up to about 7 hours, at least or up to about 8 hours, at least or up to about 9 hours, at least or up to about 10 hours, at least or up to about 11 hours, at least or up to about 12 hours, at least or up to about 16 hours, at least or up to about 18 hours, at least or up to about 24 hours, at least or up to about 36 hours, at least or up to about 48 hours, at least up to about 3 days, at least up to about 4 days, at least up to about 5 days, at least up to about 6 days, at least up to about 7 days, at least up to about 8 days, at least up to about 9 days, at least up to about 10 days, at least up to about 11 days, at least up to about 12 days, at least up to about 13 days, or at least up to about 14 days after treating a cell to induce such expression level.)

[0067] The term “peptide,” “polypeptide,” or “protein,” as used interchangeably herein, generally refers to a polymer of at least two amino acid residues joined by peptide bond(s). This term does not connote a specific length of polymer, nor is it intended to imply or distinguish whether the peptide is produced using recombinant techniques, chemical or enzymatic synthesis, or is naturally occurring. The terms apply to naturally occurring amino acid polymers as well as amino acid polymers comprising at least one modified amino acid. 17In some cases, the polymer can be interrupted by non-amino acids. The terms include amino acid chains of any length, including full length proteins, and proteins with or without secondary and / or tertiary structure (e.g., domains). The terms also encompass an amino acid polymer that has been modified, for example, by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, oxidation, and any other manipulation such as conjugation with a labeling component. The terms “amino acid” and “amino acids,” as used herein, generally refer to natural and non-natural amino acids, including, but not limited to, modified amino acids and amino acid analogues. Modified amino acids can include natural amino acids and non-natural amino acids, which have been chemically modified to include a group or a chemical moiety not naturally present on the amino acid. Amino acid analogues can refer to amino acid derivatives. The term “amino acid” includes both D-amino acids and L-amino acids.

[0068] The term “derivative,” “variant,” or “fragment,” as used herein with reference to a polypeptide, generally refers to a polypeptide related to a wild type polypeptide, for example either by amino acid sequence, structure (e.g., secondary and / or tertiary), activity (e.g., enzymatic activity) and / or function. Derivatives, variants and fragments of a polypeptide can comprise one or more amino acid variations (e.g., mutations, insertions, and deletions), truncations, modifications, or combinations thereof compared to a wild type polypeptide.

[0069] The term “engineered,” “chimeric,” or “recombinant,” as used herein with respect to a polypeptide molecule (e.g., a protein), generally refers to a polypeptide molecule having a heterologous amino acid sequence or an altered amino acid sequence as a result of the application of genetic engineering techniques to nucleic acids which encode the polypeptide molecule, as well as cells or organisms which express the polypeptide molecule. The term “engineered” or “recombinant,” as used herein with respect to a polynucleotide molecule (e.g., a DNA or RNA molecule), generally refers to a polynucleotide molecule having a heterologous nucleic acid sequence or an altered nucleic acid sequence as a result of the application of genetic engineering techniques. Genetic engineering techniques include, but are not limited to, PCR and DNA cloning technologies; transfection, transformation and other gene transfer technologies; homologous recombination; site-directed mutagenesis; and gene fusion. In some cases, an engineered or recombinant polynucleotide (e.g., a genomic DNA sequence) can be modified or altered by a gene editing moiety. For example, an heterologous 18endonuclease (e.g., an engineered Cas protein) as disclosed herein is not a naturally occurring nuclease (e.g., not a naturally occurring Cas protein). In another example, an engineered gene effector as disclosed herein is not a naturally occurring gene effector.

[0070] For example, an engineered nuclease (e.g., an engineered Cas protein) as disclosed herein is not a naturally occurring nuclease (e.g., not a naturally occurring Cas protein). The terms “engineered nuclease” and “engineered nuclease variant” may be used interchangeable herein.

[0071] The terms “engineered” and “modified” are used interchangeably herein. The terms “engineering” and “modifying” are used interchangeably herein. The terms “engineered cell” or “modified cell” are used interchangeably herein. The terms “engineered characteristic” and “modified characteristic” are used interchangeably herein.

[0072] The term “enhanced expression,” “increased expression,” or “upregulated expression” generally refers to production of a moiety of interest (e.g., a polynucleotide or a polypeptide) to a level that is above a normal level of expression of the moiety of interest in a host strain (e.g., a host cell). The normal level of expression can be substantially zero (or null) or higher than zero. The moiety of interest can comprise an endogenous gene or polypeptide construct of the host strain. The moiety of interest can comprise a heterologous gene or polypeptide construct that is introduced to or into the host strain. For example, a heterologous gene encoding a polypeptide of interest can be knocked-in (KI) to a genome of the host strain for enhanced expression of the polypeptide of interest in the host strain.

[0073] The term “enhanced activity,” “increased activity,” or “upregulated activity” generally refers to activity of a moiety of interest (e.g., a polynucleotide or a polypeptide) that is modified to a level that is above a normal level of activity of the moiety of interest in a host strain (e.g., a host cell). The normal level of activity can be substantially zero (or null) or higher than zero. The moiety of interest can comprise a polypeptide construct of the host strain. The moiety of interest can comprise a heterologous polypeptide construct that is introduced to or into the host strain. For example, a heterologous gene encoding a polypeptide of interest can be knocked-in (KI) to a genome of the host strain for enhanced activity of the polypeptide of interest in the host strain.

[0074] The term “reduced expression,” “decreased expression,” or “downregulated expression” generally refers to a production of a moiety of interest (e.g., a polynucleotide or a 19polypeptide) to a level that is below a normal level of expression of the moiety of interest in a host strain (e.g., a host cell). The normal level of expression is higher than zero. The moiety of interest can comprise an endogenous gene or polypeptide construct of the host strain. In some cases, the moiety of interest can be knocked-out or knocked-down in the host strain. In some examples, reduced expression of the moiety of interest can include a complete inhibition of such expression in the host strain.

[0075] The term “reduced activity,” “decreased activity,” or “downregulated activity” generally refers to activity of a moiety of interest (e.g., a polynucleotide or a polypeptide) that is modified to a level that is below a normal level of activity of the moiety of interest in a host strain (e.g., a host cell). The normal level of activity is higher than zero. The moiety of interest can comprise an endogenous gene or polypeptide construct of the host strain. In some cases, the moiety of interest can be knocked-out or knocked-down in the host strain. In some examples, reduced activity of the moiety of interest can include a complete inhibition of such activity in the host strain.

[0076] The term “subject,” “individual,” or “patient,” as used interchangeably herein, generally refers to a vertebrate, preferably a mammal such as a human. Mammals include, but are not limited to, murines, simians, humans, farm animals, sport animals, and pets. Tissues, cells and their progeny of a biological entity obtained in vivo or cultured in vitro are also encompassed.

[0077] The term “treatment” or “treating” generally refers to an approach for obtaining beneficial or desired results including but not limited to a therapeutic benefit and / or a prophylactic benefit. For example, a treatment can comprise administering a system or cell population disclosed herein. By therapeutic benefit is meant any therapeutically relevant improvement in or effect on one or more diseases, conditions, or symptoms under treatment. For prophylactic benefit, a composition can be administered to a subject at risk of developing a particular disease, condition, or symptom, or to a subject reporting one or more of the physiological symptoms of a disease, even though the disease, condition, or symptom may not have yet been manifested.

[0078] The term “effective amount” or “therapeutically effective amount” generally refers to the quantity of a composition, for example a composition comprising heterologous polypeptides, heterologous polynucleotides, and / or modified cells (e.g., modified 20stem cells), that is sufficient to result in a desired activity upon administration to a subject in need thereof. Within the context of the present disclosure, the term “therapeutically effective” generally refers to that quantity of a composition that is sufficient to delay the manifestation, arrest the progression, relieve or alleviate at least one symptom of a disorder treated by the methods of the present disclosure.

[0079] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than” or “greater than or equal to” applies to each of the numerical values in that series of numerical values. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.

[0080] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than,” “less than,” or “less than or equal to” applies to each of the numerical values in that series of numerical values. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.

[0081] The term “about” or “approximately” generally mean within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, “about” can mean within 1 or more than 1 standard deviation, per the practice in the art. Alternatively, “about” can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within 5-fold, and more preferably within 2- fold, of a value. Where particular values are described in the application and claims, unless otherwise stated, the term “about” meaning within an acceptable error range for the particular value should be assumed.

[0082] The use of the alternative (e.g., “or”) should be understood to mean either one, both, or any combination thereof of the alternatives. The term “and / or” should be understood to mean either one, or both of the alternatives. ENGINEERED GENE EFFECTORS

[0083] In some aspects, the present disclosure provides compositions, combinations, systems, and methods that utilize engineered gene effectors. In some 21embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-95 amino acids in length and a second peptide of 75-95 amino acids in length that is heterologous to the first peptide. In some embodiments, an engineered gene effector includes a first peptide of 75-95 amino acids in length and a second peptide of 75-95 amino acids in length that is heterologous to the first peptide, wherein the first peptide is different from the second peptide. In some embodiments, the first and / or second peptide is 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length, or optionally the first and / or second peptide is a length within a range defined by any two of the aforementioned lengths (e.g., 76-94, 77-93, 78-92, 79-91, 80-90, 82-98, etc.).

[0084] In some aspects, the present disclosure provides compositions, combinations, systems, and methods that utilize engineered gene effectors. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide. In some embodiments, an engineered gene effector includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide is different from the second peptide. In some embodiments, the first and / or second peptide is 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99¸ 100, 101, 102, 103, 104, 105, 106, 107, 108, 109 amino acids in length, or optionally a length within a range defined by any two of the aforementioned lengths (e.g., 76-94, 77-93, 78-92, 79-91, 80-90, 82- 98, 80-109, 85-108, etc.).

[0085] In some embodiments, an engineered gene effector of the present disclosure is capable of activating a target gene in a cell when the engineered gene effector is expressed therein and effectively targeted to a locus of the target gene. In some aspects, the present disclosure provides compositions, combinations, systems, and methods that utilize engineered gene effectors. In some embodiments, an engineered gene effector of the present disclosure includes a polypeptide that includes a first peptide of 75-95 (or 75-110) amino acids in length and a second peptide of 75-95 (or 75-110) amino acids in length that is heterologous to the first peptide, wherein the engineered gene effector is capable of activating a target gene in a cell when the engineered gene effector is expressed therein and effectively targeted to a locus of the target gene. In some embodiments, the first and / or second peptide are based on a human 22transcriptional modulator. In some embodiments, the first and second peptide are based on a human transcriptional modulator. In some embodiments, the first and / or second peptide are based on a viral transcriptional modulator. In some embodiments, the first and second peptide are based on a viral transcriptional modulator. In some embodiments, the first peptide is based on a human transcriptional modulator and the second peptide is based on a viral transcriptional modulator. In some embodiments, the first peptide is different from the second peptide. In some embodiments, the first peptide is a different length from the second peptide. In some embodiments, the first peptide is the same length as the second peptide.

[0086] In some embodiments, the first peptide and / or the second peptide of the engineered gene effector has a beta factor of about 30 to about 65. In some embodiments, the first peptide has a beta factor of about 30, about 35, about 40, about 45, about 50, about 55, about 60, or about 65, optionally, the beta factor of the first peptide is in a range defined by any two of the preceding values (e.g., about 30 to about 40, about 35 to about 55, about 40 to about 65). In some embodiments, the second peptide has a beta factor of about 30, about 35, about 40, about 45, about 50, about 55, about 60, or about 65, optionally, the beta factor of the second peptide is in a range defined by any two of the preceding values (e.g., about 30 to about 40, about 35 to about 55, about 40 to about 65). In some embodiments, an engineered gene effector of the present disclosure includes a first peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length and a second peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length, wherein the first and / or second peptide has a beta factor of about 30, about 35, about 40, about 45, about 50, about 55, about 60, or about 65. In some embodiments, an engineered gene effector of the present disclosure includes a first peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length and a second peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length, wherein the first and / or second peptide has a beta factor in a range defined by any two of the preceding values (e.g., about 30 to about 40, about 35 to about 55, about 40 to about 65).

[0087] In some embodiments, the first peptide and / or the second peptide of the engineered gene effector is enriched for negative electrostatic potential. In some embodiments, the first and second peptides are enriched for negative electrostatic potential. In some 23embodiments, the engineered gene effector includes a first peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length and a second peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length, wherein the first and / or second peptide is enriched for negative electrostatic potential. In some embodiments, the engineered gene effector includes a first peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length and a second peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length, wherein the first peptide is enriched for negative electrostatic potential and the second peptide is enriched for negative electrostatic potential.

[0088] In some embodiments, the first peptide and / or the second peptide of the engineered gene effector has a negative net charge. In some embodiments, the first peptide and the second peptide of the engineered gene effector has a negative net charge. In some embodiments, the engineered gene effector includes a first peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length and a second peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length, wherein the first and / or second peptide has a negative net charge. In some embodiments, the engineered gene effector includes a first peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length and a second peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length, wherein the first peptide has a negative net charge and the second peptide has a negative net charge. 24S LsecneuqeKsYQLALWFEdiRY S AAD FDASE LRPVEMNP S FDLYR T TSPP PDN SE C TLTGAILE MDSA MFGAR SDI KEDLSQVFP QTHSP THD PN P WQ ESFVEAKH SPREELLEVSGRV NNFPD LK L ANcRITa PE MA ASGE ELLSS LNLEE F C LSTH YWNGD LSAP YD L GDLYMKDTNGGIGDFKKTSI LDDG AAPoKAAAPLSLD FI LNASPL PPVVLP F LDVVRMEYSLND I D MQNALLSFID AYLFVF LTYDLGQFniL IEL PLAA ELELALN P DHTITAIVALCFPAMRDPSKVFIAVD LPMLQAVS SA K APKQ G KSLD KAHAGM RSPAIV GDDTIVDCPS MYELGSK E MNm HEaeIGQKQFDP E PSAPAQS SICI EKTLFDKRAIGK c YEP ESFDA YFRSVLWS FLARK NKDSQSAPTLK APNGTADGRLITRM PFAGSS DDSH D APSeneG LKA F ANKEGT VD PLNPIQ APSP EK HGQRL NGVCKTIIFANIS ES FNVSYGG RLWDSGVLLSKdituqIIKAGETL SYDATGRG A GL SMQGPPTHW NC TLR EGETKMMp esQSASVSNI GF P ELSGSMPDPD GSG DNATEKK QAIVTF LQ YAGL TLQGDQ L VNLSKSYNIQNNKIRRTKGD TSINF NFG GLLMQep diNDlcaa SSAIAFQVEESG KGSS EPDLF FGD L KPIQSHNV SVVDPINST ARSPSTKL CDGQ KAI GDPAK o NTGDF HEVLLT LQP TQ PFI GYSPKDE SSES DE NSVG TSTKRLPTNLQGIKGSDLLEGSNSundiVIVIIL ALAK DI VSADLMIAR PVDRS DVPE PN QNHGSGLDPDSNLNDDG AKKWGEVI SQGGim* *ADGvEGELETEPFFPMSRTKFEPLN AGRLTLET D YPN FGMAANTQSK AYG F GGG GGGVSGSGiA****A A A A A A A A A AC CD D D D D DE E EPESEVEAFCFGFG G G G G D G GdnI:O3Nel DIba QETS1 2 3 4 5 6 7 8 901112131415161718191021222324252627282920313233343536325APLAQWRPRPLASTESSVWDELAREDDSSDQDRNARKRQVRRRPQEEEEEEEEEEEEESRDMLFEDFQ V KD SHMVI CAYSTQSEMQDG CSKFGITSRLSQ RT RCLENSGELLGPRESPHCCA S DRI PDEV ARRGVPVV HLHLYEQ NERHPYP GPVI RRKGKPYLVMNAKHSVP ESQ LEP AQ REAEETSV T VHSPADLGQ IFD IL D DAPF SPLDPGRET APNSEACQDYT SGSPVF E E LS ASNVSGS I EMNANDQEVGEVIGFVGASAVYLNNYLDR P TYYPAKSND QAIAL LTYDKNIFWSGEIDEGLPLA L YNIVHILKMASQNMEGELKPN KASAGGFTLK VVD AFEPAPDNT DEN QRHD EDPVLDG GGGSGAVGNAKK REPI DGGKG VT EQYPMN CNMEERCD DA KNID NPA TL DHS VL ETD L APLK TTIARDGPQ SSVYSHDG AEIVRKP LL KLN VHI LQLNTNTESLSPKDD KSDSLSK LP S R PAP KVSF P LNQKDSNAQVD E Y QFA GVMAYGKTLI R RVE SIEPPHCCVY GDWKIIFTDGI QTELFS EP VSAS PID L YRRSLE P PSP KGDSK P KDALWD NIGLDCSRDD R AQIKK RKE EISPVDL SPHT NHAH NA APFEI RDDL IL APPLRREPL CIVPPTQYS EDEREEGRDEWHGE TMSS CVDIYR FIHFC I SDIHTMDDM LSGKL KSEE RLLPRNRFHL PVSSLFVMTGSEALVPLDSL TREQ VNDIA YS EIDS LPGSS KES SAVK ERRLHGMDSVKLDLKKISRGANEAF LP AKPKP SNR SKTVVQSA APGPAA PGSGAT RTSSDLRKSKGEGPRPNMP TKNISNMFA DM T QAPIL NNDLIM VSRTT N SN AGN GQS TMTQRNMVSRLCAWGPFFKRV LGVRADTPQELYEVCPDRKLHLGT IT HMS LKES S LYEQGGL YIWSETQYIS P MKTVS LDEV E AGLQD VP SRPS TGFI RDII D V VLVR NNYHSKPSAYGS GVGGYAAL LVGHQRFKKLPKF L EDGPHC EKNISAPAKC PHQWC GGKRG G G H HEIKDLILLLLLRLSLSLTLWLM N N N N NFPGPGPLPNPPPRPSPSPQ Q QR R RGRMRPRPR738393041424344454647484940515253545556575859506162636465666768696071727374757677726SLVPPEPLFLGESAALILSEPFFENV AQ VQVEEI I PA LGAT T PE P RKGAG HVGL P E TVFGSVKAVNGQ STAL DDTATNGLVALAI EEKNGDP EELKDD VP LLTIVAE LDT PPDVDNNMSI VEAVDQS E SEGDNVISGALDHE QSF P EL M LFDPLNKS PVL YDL LS LWVNLDEVKR LVV VLDCDPI LIGATFGFLKDD VLDASQNL IGKT C P F EKIDEG Q QMQQEFL L T SKPI HNYHD AD SSLSANND AHC TL PKL EQ KEPQIPIQGRVE EIMLVNLSQW RLAV HI EA F ADS EKT SF TIVGPDQPAILN VHLVLT L RLNSGVKVE ERT EQP TVFADP L LDLQIQVSEL DISFA N EAEKYPLPQA ASIEAYSRAGA FLVMIHPDAI YSPLSLP T L IFRIMEVCQSTGDQVAKLVRL TK GAFEVK NGD S VVAPDHPRPAE PELDASVNP AGTGAAGLSVSNKS TPRACP SLSL AE E EGVSHENT EVV MLSLE Q AEEAN AGMLVLKSVEKKEAYQPGLLHIAEYRLKMEVE SIF Q MIAV VTRLPMQ TF FHT HL TS SFSTPPLIT SEDT SDAV NE LNFGKVPAHSIDSP RITRQQSAYLTD PTSSQNG AGLNMAA NIELSPI CDPYQEVLDE SDSWMRYEPN KDTQQSGP EG SGVCYN NWE EET SIRHEQIVGTEGPSERK ERL TM L VVSQPP SMIGLGEIVTFIG SAA ILRDVQEI I S LVEMAVTNQLNITYIGRNAHEPQ S KDN EVDAY CVEPPYVVAYSVEQ GFEVAFQ TLDIS LKGLIGS P P T PNDIPNEWAGLAKRVDIPFD MC ICYRLDPIPLGDYGLVPQ EIVINS L R LHL LLHDACNKPPDPVVAKDGTSQWPGS TT S LQYTGTDEASVR SMLGQ QEAILVEKFQLSATSS SGKKL RDC RLHEMLALEPVFET SHSGYTFD MMTQE RVE HS IKGSYQIAITEAELFQD T V VE LAAYTLIWSLQEFSKERLEVHYDRNPFL Q VLHTSMA DRVSNTLSAD DVLFDR EQKKHGI P P RVHLLNVT PNDSDGA MAPNVQ TDSI RAVEQDLTVALPAR IDPS TLIFII ADKS DPEIE PIDSNQ DSGIS LDP T SRLDEVRAPS L LMSP LHE NVL HLEPDRKQYSDECGAAMNIPLVPI LGGSGPLFSL E E L L CNLIVGTVEDVMMF PAEVGSKY E KDIKL LGPYLG QTATVEYES TMNQ E RGKTEPF GQAHS P EPP SVGRYALPAP NSQS TL QVDMETDLKAQL M QNLCGTWTIVPNSMAF TTAFADVITTLINEVYGMQEGAVRKLWT TEPTF EALNQKLDS TAYAPHG T QGVY F VDA T NSPHLNI EA GV DDEERASLSPSRSSSTSVSPTTTTTWTWTAHVIKKKV V VTTMLAVTVYA V V V YPY8 9 007 7 8182838485868788898091929394959697989990127

[0089] In some embodiments, the engineered gene effector includes a polypeptide that includes a first peptide that includes any one of the sequences provided in Table 3, and a second peptide that is heterologous to the first peptide and includes any one of the sequences provided in Table 3. In some embodiments, the first peptide is 75-110 amino acids in length. In some embodiments, the second peptide is 75-110 amino acids in length. In some embodiments, the engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length, including any one of SEQ ID NOs: 3-100, and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, including any one of SEQ ID NOs: 3-100. In some embodiments, the engineered gene effector includes a polypeptide that includes a first peptide of 75-95 amino acids in length, including any one of SEQ ID NOs: 3-33 and 35-100, and a second peptide of 75-95 amino acids in length that is heterologous to the first peptide, including any one of SEQ ID NOs: 3-33 and 35-100. In some embodiments, the engineered gene effector includes a polypeptide that includes a first peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length, including any one of SEQ ID NOs: 3-33 and 35-100, and a second peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length that is heterologous to the first peptide, including any one of SEQ ID NOs: 3-33 and 35-100. In some embodiments, the first peptide is different from the second peptide. In some embodiments, the first peptide is a different length from the second peptide. In some embodiments, the first peptide is the same length as the second peptide. In some embodiments, the SEQ ID NOs of the first and second peptides are selected according to any one of the paired arrangements of SEQ ID NOs of the first peptide and the second peptide as set forth in Table 4, where the first peptide is N-terminal to the second peptide.

[0090] In some embodiments, the engineered gene effector includes a polypeptide that includes a first peptide that includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of the sequences provided in Table 3, and a second peptide that is heterologous to the first peptide and includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of the sequences provided in Table 3. In some embodiments, the first peptide includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 2894%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 3-100, or optionally, the first peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to any one of SEQ ID NOs: 3-100. In some embodiments, the second peptide includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 3-100, or optionally, the second peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to any one of SEQ ID NOs: 3-100. In some embodiments, the first peptide is different from the second peptide. In some embodiments, the first peptide has the noted percent sequence identity to any one of SEQ ID NOs: 3-100 that is different from the sequence of any one of SEQ ID NOs: 3-100 to which the second peptide has the noted percent sequence identity. In some embodiments, the first peptide is different from the second peptide. In some embodiments, the SEQ ID NOs of the first and second peptides are selected according to any one of the paired arrangements of SEQ ID NOs of the first peptide and the second peptide as set forth in Table 4, where the first peptide is N-terminal to the second peptide.

[0091] In some embodiments, the engineered gene effector includes a first peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length, and includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 3-33 and 35-100, and a second peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length that is heterologous to the first peptide, and includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 3-33 and 35-100. In some embodiments, the first peptide is different from the second peptide. In some embodiments, the first peptide has the noted percent sequence identity to any one of SEQ ID NOs: 3-33 and 35-100 that is different from the sequence of any one of SEQ ID NOs: 3-33 and 35-100 to which the second peptide has the noted percent sequence identity. In some embodiments, the first peptide is different from the second peptide. In some embodiments, the first peptide is a different length from the second peptide. In some embodiments, the first peptide is the same length as the second peptide. In some embodiments, 29the SEQ ID NOs of the first and second peptides are selected according to any one of the paired arrangements of SEQ ID NOs of the first peptide and the second peptide as set forth in Table 4, where the first peptide is N-terminal to the second peptide.

[0092] In some embodiments, the first peptide and / or the second peptide is 85 amino acids in length. In some embodiments, the first peptide and the second peptide are each 85 amino acids in length. In some embodiments, the engineered gene effector includes a polypeptide that includes a first peptide of 85 amino acids in length, including any one of SEQ ID NOs: 3-33, 35-100 and a second peptide of 85 amino acids in length that is heterologous to the first peptide, including any one of SEQ ID NOs: 3-33, 35-100. In some embodiments, the first peptide is different from the second peptide. In some embodiments, the first peptide is 85 amino acids in length and includes a sequence of, of about, of at least, or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 3-33, 35-100, or optionally, the first peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to any one of SEQ ID NOs: 3-33, 35-100. In some embodiments, the second peptide is 85 amino acids in length and includes a sequence of, of about, of at least, or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 3-33, 35-100, or optionally, the second peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to any one of SEQ ID NOs: 3-33, 35-100. In some embodiments, the engineered gene effector includes a first peptide of 85 amino acids in length including a sequence of, of about, of at least, or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 3-33, 35-100, and a second peptide of 85 amino acids in length that is heterologous to the first peptide, including a sequence of, of about, of at least, or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 3-33, 35-100. In some embodiments, the first peptide has the noted percent sequence identity to any one of SEQ ID NOs: 3-33, 35-100 that is different from the sequence of any one of SEQ ID NOs: 3-33, 35-100 to which the second peptide has the noted percent sequence identity. In some embodiments, the SEQ ID NOs of the first and second peptides are selected according to any one of the paired arrangements of 30SEQ ID NOs of the first peptide and the second peptide as set forth in Table 4, where the first peptide is N-terminal to the second peptide .

[0093] In some embodiments, the engineered gene effector includes a first peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length, and includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 3-33 and 35-100, and a second peptide that is heterologous to the first peptide, and includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 34. In some embodiments, the engineered gene effector includes a polypeptide that includes a first peptide of 75-95 amino acids in length, and has a sequence of any one of SEQ ID NOs: 3-33, 35-100 with 0, 1, 2, or 3 amino acid residue mutations thereto, and a second peptide of that is heterologous to the first peptide, and has the sequence of SEQ ID NO: 34 with 0, 1, 2, or 3 amino acid residue mutations thereto. In some embodiments, the amino acid residue mutations are conservative substitutions. In some embodiments, the engineered gene effector includes a first peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length, and includes a sequence of any one of SEQ ID NOs: 3-33 and 35- 100, and a second peptide that is heterologous to the first peptide, and includes a sequence of SEQ ID NO: 34. In some embodiments, the engineered gene effector includes a first peptide of 85 amino acids in length and includes a sequence of any one of SEQ ID NOs: 3-33 and 35- 100, and a second peptide that is heterologous to the first peptide, and includes a sequence of SEQ ID NO: 34. In some embodiments, the first peptide is N-terminal to the second peptide. In some embodiments, the second peptide is N-terminal to the first peptide.

[0094] In some embodiments, the first peptide includes a sequence of any one of SEQ ID NOs: 3-100 with 0, 1, 2, or 3 amino acid residue mutations thereto. In some embodiments, the second peptide includes a sequence of any one of SEQ ID NOs: 3-100 with 0, 1, 2, or 3 amino acid residue mutations thereto. In some embodiments, the engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length, and has a sequence of any one of SEQ ID NOs: 3-33, 35-100 with 0, 1, 2, or 3 amino acid residue mutations thereto, and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, and has any one of SEQ ID NOs: 3-100 with 0, 1, 2, or 3 31amino acid residue mutations thereto. In some embodiments, the engineered gene effector includes a polypeptide that includes a first peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length, and has a sequence of any one of SEQ ID NOs: 3-33, 35-100 with 0, 1, 2, or 3 amino acid residue mutations thereto, and a second peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length that is heterologous to the first peptide, and has any one of SEQ ID NOs: 3-33, 35-100 with 0, 1, 2, or 3 amino acid residue mutations thereto. In some embodiments, the amino acid residue mutations are conservative substitutions. In some embodiments, the engineered gene effector includes a polypeptide that includes a first peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length, and that includes any one of SEQ ID NOs: 3-33, 35-100 with 0, 1, 2, or 3 amino acid residue mutations thereto, wherein any mutations thereof are conservative substitutions, and a second peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length that is heterologous to the first peptide, and that includes any one of SEQ ID NOs: 3-33, 35-100 with 0, 1, 2, or 3 amino acid residue mutations thereto, wherein any mutations thereof are conservative substitutions. In some embodiments, the SEQ ID NOs of the first and second peptides are selected according to any one of the paired arrangements of SEQ ID NOs of the first peptide and the second peptide as set forth in Table 4, where the first peptide is N-terminal to the second peptide.

[0095] The first and second polypeptide can be at any suitable position relative to each other in the polypeptide. In some embodiments, the first peptide is N-terminal to the second peptide in the polypeptide. In some embodiments, the second peptide is N-terminal to the first peptide in the polypeptide. In some embodiments, the first peptide that is N-terminal to the second peptide includes any one of SEQ ID NOs: 3-100, and the second peptide includes any one of SEQ ID NOs: 3-100. In some embodiments, the engineered gene effector includes a polypeptide that includes a first peptide that is N-terminal to a second peptide, wherein the first peptide of 75-110 amino acids in length includes any one of SEQ ID NOs: 3-100 and the second peptide of 75-110 amino acids in length is heterologous to the first peptide includes any one of SEQ ID NOs: 3-100. In some embodiments, the engineered gene effector includes a polypeptide that includes a first peptide that is N-terminal to a second peptide, wherein the first peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 3295 amino acids in length includes any one of SEQ ID NOs: 3-33, 35-100 and the second peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length is heterologous to the first peptide includes any one of SEQ ID NOs: 3-33, 35-100. In some embodiments, the engineered gene effector includes a polypeptide that includes a first peptide that is N-terminal to a second peptide, wherein the first peptide of 85 amino acids in length includes any one of SEQ ID NOs: 3-33, 35-100 and the second peptide of 85 amino acids in length is heterologous to the first peptide and includes any one of SEQ ID NOs: 3-33, 35-100.

[0096] In some embodiments, the first peptide is N-terminal to the second peptide in the polypeptide, and the first and second peptides are according to any one of the paired arrangements of SEQ ID NOs of the first peptide and the second peptide as set forth in Table 4. In some embodiments, the engineered gene effector includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length, and where the sequence of the first and second peptides are selected according to the SEQ ID NOs in any one of the paired arrangements of SEQ ID NOs of the first peptide and the second peptide as set forth in Table 4. In some embodiments, the engineered gene effector includes a first peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length and a second peptide of 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, or 95 amino acids in length, and where the sequence of the first and second peptides are selected according to the SEQ ID NOs in any one of the paired arrangements of SEQ ID NOs of the first peptide and the second peptide as set forth in Table 4. In some embodiments, the engineered gene effector includes a first peptide of 85 or 108 amino acids in length and a second peptide of 85 or 108 amino acids in length, where the sequence of the first and second peptides are selected according to any one of the paired arrangements of SEQ ID NOs of the first peptide and the second peptide as set forth in Table 4.

[0097] In some embodiments, the engineered gene effector includes polypeptide having a first peptide N-terminal to a second peptide, wherein the first peptide includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 75, 4, 40, and 77, or optionally, the first peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) 33to any one of SEQ ID NOs: 75, 4, 40, and 77, and the second peptide includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 64, 44, 77, 17, and 96, or optionally, the second peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to any one of SEQ ID NOs: 64, 44, 77, 17, and 96. In some embodiments, the first peptide includes a sequence of any one of SEQ ID NOs: 75, 4, 40, and 77, and the second peptide includes a sequence of any one of SEQ ID NOs: 64, 44, 77, 17, and 96.

[0098] In some embodiments, the engineered gene effector includes a first peptide N-terminal to a second peptide, wherein the first peptide and the second peptide are selected according to any one of the paired arrangements of SEQ ID NOs of the first peptide and the second peptide, as set forth in Table A. In some embodiments, the first peptides is linked to the second peptide by a linker. In some embodiments, the linker comprises the sequence of any one of SEQ ID NOs: 2211-2221. In some embodiments, the linker comprises the sequence of any one of SEQ ID NO: 2211. Table A: Non-limiting examples of engineered gene effector peptide combinations

[0099] In some embodiments, the engineered gene effector includes a first peptide N-terminal to a second peptide, wherein the first peptide includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 75, or optionally, the first peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 75, and the second peptide 34includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 64, or optionally, the second peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 64. In some embodiments, the first peptide includes a sequence of SEQ ID NO: 75, and the second peptide includes a sequence of SEQ ID NO: 64. In some embodiments, the first peptides is linked to the second peptide by a linker. In some embodiments, the linker comprises the sequence of any one of SEQ ID NOs: 2211-2221. In some embodiments, the linker comprises the sequence of any one of SEQ ID NO: 2211.

[0100] In some embodiments, the engineered gene effector includes a first peptide N-terminal to a second peptide, wherein the first peptide includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 4, or optionally, the first peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 4, and the second peptide includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 44, or optionally, the second peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 44. In some embodiments, the first peptide includes a sequence of SEQ ID NO: 4, and the second peptide includes a sequence of SEQ ID NO: 44. In some embodiments, the first peptides is linked to the second peptide by a linker. In some embodiments, the linker comprises the sequence of any one of SEQ ID NOs: 2211-2221. In some embodiments, the linker comprises the sequence of any one of SEQ ID NO: 2211.

[0101] In some embodiments, the engineered gene effector includes a first peptide N-terminal to a second peptide, wherein the first peptide includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 75, or optionally, the first peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 75, and the second peptide includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 3592%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 77, or optionally, the second peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 77. In some embodiments, the first peptide includes a sequence of SEQ ID NO: 75, and the second peptide includes a sequence of SEQ ID NO: 77. In some embodiments, the first peptides is linked to the second peptide by a linker. In some embodiments, the linker comprises the sequence of any one of SEQ ID NOs: 2211-2221. In some embodiments, the linker comprises the sequence of any one of SEQ ID NO: 2211.

[0102] In some embodiments, the engineered gene effector includes a first peptide N-terminal to a second peptide, wherein the first peptide includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 40, or optionally, the first peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 40, and the second peptide includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 17, or optionally, the second peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 17. In some embodiments, the first peptide includes a sequence of SEQ ID NO: 40, and the second peptide includes a sequence of SEQ ID NO: 17. In some embodiments, the first peptides is linked to the second peptide by a linker. In some embodiments, the linker comprises the sequence of any one of SEQ ID NOs: 2211-2221. In some embodiments, the linker comprises the sequence of any one of SEQ ID NO: 2211.

[0103] In some embodiments, the engineered gene effector includes a first peptide N-terminal to a second peptide, wherein the first peptide includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 77, or optionally, the first peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 77, and the second peptide includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 96, or 36optionally, the second peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 96. In some embodiments, the first peptide includes a sequence of SEQ ID NO: 77, and the second peptide includes a sequence of SEQ ID NO: 96. In some embodiments, the first peptides is linked to the second peptide by a linker. In some embodiments, the linker comprises the sequence of any one of SEQ ID NOs: 2211-2221. In some embodiments, the linker comprises the sequence of any one of SEQ ID NO: 2211.

[0104] In some embodiments, the engineered gene effector includes a first peptide N-terminal to a second peptide, wherein the first peptide includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 77, or optionally, the first peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 77, and the second peptide includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 64, or optionally, the second peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 64. In some embodiments, the first peptide includes a sequence of SEQ ID NO: 77, and the second peptide includes a sequence of SEQ ID NO: 64. In some embodiments, the first peptides is linked to the second peptide by a linker. In some embodiments, the linker comprises the sequence of any one of SEQ ID NOs: 2211-2221. In some embodiments, the linker comprises the sequence of any one of SEQ ID NO: 2211.

[0105] The first peptide and the second peptide in the polypeptide of the engineered gene effector can be linked to each other directly or indirectly (e.g., via a spacer or linker). In some embodiments, the engineered gene effector includes a polypeptide including a first peptide and a second peptide, where the first and second peptides are linked by a spacer (e.g., peptide spacer). Any suitable spacer can be used. In some embodiments, the spacer (e.g., peptide spacer) is a flexible linker that has a sequence containing stretches of glycine and serine residues. The small size of the glycine and serine residues provides flexibility and allows for mobility of the connected functional domains. The incorporation of serine or threonine can in some embodiments maintain the stability of the spacer (e.g., peptide spacer) or linker in 37aqueous solutions by forming hydrogen bonds with the water molecules, thereby reducing unfavorable interactions between the spacer or linker and the linked moieties. Flexible spacers or linkers can also contain in some embodiments additional amino acids such as threonine and alanine to maintain flexibility, as well as polar amino acids such as lysine and glutamine to improve solubility. A rigid spacer or linker can have, for example, an alpha helix-structure. An alpha-helical rigid spacer or linker can act as a spacer between protein domains. Non- limiting examples of spacers or linkers include the sequences in Table 5, and repeats thereof, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 repeats. SEQ ID NOs: 2211-2217 provide flexible spacers or linkers or subunits thereof. SEQ ID NOs: 2218-2221 provide rigid spacers or linkers or subunits thereof. In some embodiments, the spacer (e.g., peptide spacer) or linker is or includes SEQ ID NO:2211. Table 5: Linker amino acid sequences

[0106] In some embodiments, a spacer (e.g., peptide spacer) or linker as disclosed herein can be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acid residues in length.

[0107] In some embodiments, a spacer (e.g., peptide spacer) or linker as disclosed herein can comprise at least 1, at least 2, at least 3, at least 5, at least 7, at least 9, at least 11, at least 13, at least 15, or at least 20 amino acids. In some embodiments, a linker can comprise 38at most 5, at most 7, at most 9, at most 11, at most 13, at most 15, at most 20, at most 25, at most 30, at most 40, or at most 50 amino acids.

[0108] In some embodiments, the spacer (e.g., peptide spacer) or linker includes at most about 20, at most about 15, at most about 14, at most about 13, at most about 12, at most about 11, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, at most about 2, or about 1 glycine-serine (GS) linker(s). In some embodiments, the engineered gene effector includes a polypeptide having a first peptide and a second peptide that are linked by a spacer (e.g., peptide spacer) or linker that includes at most about 20, at most about 15, at most about 14, at most about 13, at most about 12, at most about 11, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, at most about 2, or about 1 glycine-serine (GS) linker(s). In some embodiments, the engineered gene effector includes a polypeptide having a first peptide and a second peptide that are linked by a spacer (e.g., peptide spacer) or linker that includes at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about , at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 20, or more GS linker(s).

[0109] In some embodiments, the spacer (e.g., peptide spacer) or linker includes at most about 20, at most about 15, at most about 14, at most about 13, at most about 12, at most about 11, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, at most about 2, or about 1 glycine (G) linker(s). In some embodiments, the engineered gene effector includes a polypeptide having a first peptide and a second peptide that are linked by a spacer (e.g., peptide spacer) or linker that includes at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about , at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 20, or more G linker(s). In some embodiments, the engineered gene effector includes a polypeptide having a first peptide and a second peptide that are linked by a spacer (e.g., peptide spacer) or linker that includes at most about 20, at most about 15, at most about 14, at most about 13, at most about 12, at most about 11, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, at most about 2, or 39about 1 glycine (G) linker(s). In some embodiments, the engineered gene effector includes a polypeptide having a first peptide and a second peptide that are linked by a spacer (e.g., peptide spacer) or linker that includes at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about , at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 20, or more G linker(s).

[0110] In some embodiments, the engineered gene effector includes a polypeptide having a first peptide and a second peptide that are linked by a spacer (e.g., peptide spacer) or linker that includes any one of SEQ ID NOs:2211-2221, or sequence having 1-3 mutations thereto. In some embodiments, the engineered gene effector includes a polypeptide having a first peptide and a second peptide that are linked by a spacer (e.g., peptide spacer) or linker that includes SEQ ID NO:2211, or sequence having 1-3 mutations thereto.

[0111] In some embodiments, the engineered gene effector includes a polypeptide that includes a first peptide that is N-terminal to a second peptide, wherein the first peptide of 85 or 108 amino acids in length includes any one of SEQ ID NOs: 3-100 and the second peptide of 85 or 108 amino acids in length is heterologous to the first peptide and includes any one of SEQ ID NOs: 3-100, where the SEQ ID NOs of the first and second peptides are selected according to any one of the paired arrangements of SEQ ID NOs of the first peptide and the second peptide as set forth in Table 4, where the first and second peptide are linked by a spacer (e.g., peptide spacer) or linker, e.g., a peptide linker. In some embodiments, the engineered gene effector includes a polypeptide that includes a first peptide that is N-terminal to a second peptide, wherein the first peptide of 85 amino acids in length includes any one of SEQ ID NOs: 3-33, 35-100 and the second peptide of 85 amino acids in length is heterologous to the first peptide and includes any one of SEQ ID NOs: 3-33, 35-100, where the SEQ ID NOs of the first and second peptides are selected according to any one of the paired arrangements in Table 4, where the first and second peptide are linked by a spacer (e.g., peptide spacer) or linker, e.g., a peptide linker. In some embodiments, the linker is selected from any one of SEQ ID NOs: 2211-2221. In some embodiments, the spacer (e.g., peptide spacer) or linker is SEQ ID NO:2211. In some embodiments, non-peptide spacers or linkers are used. A non-peptide spacer or linker can be, for example a chemical linker. Two parts of a complex of the disclosure can be connected by a chemical linker. Each chemical linker of the disclosure can be alkylene, 40alkenylene, alkynylene, heteroalkylene, cycloalkylene, heterocycloalkylene, arylene, or heteroarylene, any of which is optionally substituted. In some embodiments, a chemical linker of the disclosure can be an ester, ether, amide, thioether, or polyethyleneglycol (PEG). In some embodiments, a spacer or linker can reverse the order of the amino acids sequence in a compound, for example, so that the amino acid sequences linked by the linked are head-to- head, rather than head-to-tail. Non-limiting examples of such spacers or linkers include diesters of dicarboxylic acids, such as oxalyl diester, malonyl diester, succinyl diester, glutaryl diester, adipyl diester, pimetyl diester, fumaryl diester, maleyl diester, phthalyl diester, isophthalyl diester, or terephthalyl diester. Non-limiting examples of such spacers or linkers include diamides of dicarboxylic acids, such as oxalyl diamide, malonyl diamide, succinyl diamide, glutaryl diamide, adipyl diamide, pimetyl diamide, fumaryl diamide, maleyl diamide, phthalyl diamide, isophthalyl diamide, or terephthalyl diamide. Non-limiting examples of such spacers or linkers include diamides of diamino linkers, such as ethylene diamine, 1,2- di(methylamino)ethane, 1,3-diaminopropane, 1,3-di(methylamino)propane, 1,4- di(methylamino)butane, 1,5-di(methylamino)pentane, 1,6-di(methylamino)hexane, and pipyrizine. Non-limiting examples of optional substituents include hydroxyl groups, sulfhydryl groups, halogens, amino groups, nitro groups, nitroso groups, cyano groups, azido groups, sulfoxide groups, sulfone groups, sulfonamide groups, carboxyl groups, carboxaldehyde groups, imine groups, alkyl groups, halo-alkyl groups, alkenyl groups, halo-alkenyl groups, alkynyl groups, halo-alkynyl groups, alkoxy groups, aryl groups, aryloxy groups, aralkyl groups, arylalkoxy groups, heterocyclyl groups, acyl groups, acyloxy groups, carbamate groups, amide groups, ureido groups, epoxy groups, or ester groups.

[0112] In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of, of about, of at least or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 115-274, 276-287, 289-367, 370-445, 448-746, 748-777, 779-929, 931-1007, 1009- 1156, 1158-1194, 1196-1288, 1290-1350, 1352-1451, or optionally, the polypeptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to any one of SEQ ID NOs: 115-274, 276- 287, 289-367, 370-445, 448-746, 748-777, 779-929, 931-1007, 1009-1156, 1158-1194, 1196- 1288, 1290-1350, 1352-1451. In some embodiments, the engineered gene effector includes a 41polypeptide that includes a sequence of any one of SEQ ID NOs: 115-274, 276-287, 289-367, 370-445, 448-746, 748-777, 779-929, 931-1007, 1009-1156, 1158-1194, 1196-1288, 1290- 1350, 1352-1451. In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of any one of SEQ ID NOs: 115-274, 276-287, 289-367, 370-445, 448-746, 748-777, 779-929, 931-1007, 1009-1156, 1158-1194, 1196-1288, 1290-1350, 1352- 1451, or a sequence with 0, 1, 2, or 3 amino acid residue mutations thereto. In some embodiments, the mutations are conservative substitutions.

[0113] In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of, of about, of at least, or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 1085, 122, 1084, 653, 1099, and 1107, or optionally, the first peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to any one of SEQ ID NOs: 1085, 122, 1084, 653, 1099, and 1107. In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of any one of SEQ ID NOs: 1085, 122, 1084, 653, 1099, and 1107 with 0, 1, 2, or 3 amino acid residue mutations thereto. In some embodiments, the mutations are conservative substitutions. In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of any one of SEQ ID NOs: 1085, 122, 1084, 653, 1099, and 1107. In some embodiments, the engineered gene effector that consists of a polypeptide having the sequence of any one of SEQ ID NOs: 1085, 122, 1084, 653, 1099, and 1107.

[0114] In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of, of about, of at least, or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 1085, 122, 1084, 653, 1099, and 1107, or optionally, the first peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90- 100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 1085, 122, 1084, 653, 1099, and 1107. In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of SEQ ID NO: 1085, 122, 1084, 653, 1099, and 1107 with 0, 1, 2, or 3 amino acid residue mutations thereto. In some embodiments, the mutations are conservative substitutions. In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of SEQ ID NO: 1085, 122, 1084, 653, 1099, and 1107. In some embodiments, the 42engineered gene effector that consists of a polypeptide having the sequence of SEQ ID NO: 1085, 122, 1084, 653, 1099, and 1107. In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of, of about, of at least, or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 1085, or optionally, the first peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95- 100%, 98-100%, etc.) to SEQ ID NO: 1085. In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of SEQ ID NO: 1085 with 0, 1, 2, or 3 amino acid residue mutations thereto. In some embodiments, the mutations are conservative substitutions. In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of SEQ ID NO: 1085. In some embodiments, the engineered gene effector that consists of a polypeptide having the sequence of SEQ ID NO: 1085.

[0115] In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of, of about, of at least, or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 122, or optionally, the first peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 122. In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of SEQ ID NO: 122 with 0, 1, 2, or 3 amino acid residue mutations thereto. In some embodiments, the mutations are conservative substitutions. In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of SEQ ID NO: 122. In some embodiments, the engineered gene effector that consists of a polypeptide having the sequence of SEQ ID NO: 122.

[0116] In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of, of about, of at least, or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 1084, or optionally, the first peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 1084. In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of SEQ ID NO: 1085 with 0, 1, 2, or 3 amino acid residue mutations thereto. In some embodiments, the mutations are conservative substitutions. In some 43embodiments, the engineered gene effector includes a polypeptide that includes a sequence of SEQ ID NO: 1084. In some embodiments, the engineered gene effector that consists of a polypeptide having the sequence of SEQ ID NO: 1084.

[0117] In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of, of about, of at least, or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 653, or optionally, the first peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 653. In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of SEQ ID NO: 653 with 0, 1, 2, or 3 amino acid residue mutations thereto. In some embodiments, the mutations are conservative substitutions. In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of SEQ ID NO: 653. In some embodiments, the engineered gene effector that consists of a polypeptide having the sequence of SEQ ID NO: 653.

[0118] In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of, of about, of at least, or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 1099, or optionally, the first peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 1099. In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of SEQ ID NO: 1099 with 0, 1, 2, or 3 amino acid residue mutations thereto. In some embodiments, the mutations are conservative substitutions. In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of SEQ ID NO: 1099. In some embodiments, the engineered gene effector that consists of a polypeptide having the sequence of SEQ ID NO: 1099.

[0119] In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of, of about, of at least, or up to 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 1107, or optionally, the first peptide includes a sequence having a percent identity in a range defined by any two of the preceding values (e.g., 85-100%, 90-100%, 95-100%, 98-100%, etc.) to SEQ ID NO: 1107. In some embodiments, the engineered gene effector includes a polypeptide that 44includes a sequence of SEQ ID NO: 1107 with 0, 1, 2, or 3 amino acid residue mutations thereto. In some embodiments, the mutations are conservative substitutions. In some embodiments, the engineered gene effector includes a polypeptide that includes a sequence of SEQ ID NO: 1107. In some embodiments, the engineered gene effector that consists of a polypeptide having the sequence of SEQ ID NO: 1107.

[0120] In some embodiments, an engineered gene effector is capable of activating a target gene in a cell when the engineered gene effector is expressed therein and effectively targeted to a locus of the target gene (e.g., in conjunction with the heterologous endonuclease). In some embodiments, the target gene is endogenous to the cell. In some embodiments, the target gene is a silenced gene, optionally wherein the silenced gene is a methylated gene. In some embodiments, an engineered gene effector is capable of repressing a target gene in a cell when the engineered gene effector is expressed therein and effectively targeted to a locus of the target gene.

[0121] The target gene may be any suitable gene. For example, without limitation, the target gene may be any one of the genes listed in Table 6.

[0122] In some embodiments, an engineered gene effector (e.g., in conjunction with the heterologous endonuclease) is capable of increasing the expression level of a target gene by, by about, or by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 250%, 300%, 400%, 500%, or more, or optionally the engineered gene effector is capable of increasing the expression level of a target gene by a percentage in a range defined by any two of the preceding values (e.g., 10-100%, 100-200%, 200-400%, 250-500%, 10-50%, or 50-100%, etc.). In some embodiments, an engineered gene effector (e.g., in conjunction with the heterologous endonuclease) is capable of increasing the expression level of a target gene by at least or about 10%, at least or about 20%, at least or about 30%, at least or about 40%, at least or about 50%, at least or about 60%, at least or about 70%, at least or about 80%, at least or about 90%, at least or about 100%, at least or about 200%, at least or about 250%, at least or about 300%, at least or about 400%, or at least or about 500%.

[0123] In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the engineered gene effector is capable of activating a synthetic reporter gene in a cell (e.g., in conjunction with the 45heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least or about 1, at least or about 2, at least or about 3, at least or about 4, at least or about 5, at least or about 6, at least or about 7, at least or about 8, at least or about 9, at least or about 10, at least or about 11, at least or about 12, at least or about 13, at least or about 14, at least or about 14, at least or about 15, or at least or about 16, when the engineered gene effector is expressed therein and effectively targeted to a locus of the synthetic reporter gene, optionally, the activation (e.g., as expressed as a log2FoldChange) is in a range defined by any two of the preceding values (e.g., about 1 to about 5, about 1 to about 3, about 3 to about 5, about 2 to about 4, about 1 to about 15, about 1 to about 10, about 5 to about 10, about 10 to about 15). In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-95 amino acids in length and a second peptide of 75-95 amino acids in length that is heterologous to the first peptide, wherein the engineered gene effector is capable of activating a synthetic reporter gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least or about 1, at least or about 2, at least or about 3, at least or about 4, at least or about 5, at least or about 6, at least or about 7, at least or about 8, at least or about 9, at least or about 10, at least or about 11, at least or about 12, at least or about 13, at least or about 14, at least or about 14, at least or about 15, or at least or about 16, when the engineered gene effector is expressed therein and effectively targeted to a locus of the synthetic reporter gene, optionally, the activation (e.g., as expressed as a log2FoldChange) is in a range defined by any two of the preceding values (e.g., about 1 to about 5, about 1 to about 3, about 3 to about 5, about 2 to about 4, about 1 to about 15, about 1 to about 10, about 5 to about 10, about 10 to about 15). In some embodiments, the engineered gene effector includes a first peptide of 85 or 108 amino acids in length and a second peptide of 85 or 108 amino acids in length, where the sequence of the first and second peptides are selected according to any one of the paired arrangements of SEQ ID NOs of the first peptide and the second peptide as set forth in Table 4. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates a synthetic reporter gene in a cell (e.g., in conjunction with the heterologous endonuclease as 46described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 1 as provided in Table 9, where the combination of SEQ ID NOs of the first peptide and the second peptide of the engineered gene effector in Table 9 are identified in Table 4 based on the barcode, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates a synthetic reporter gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 2 as provided in Table 9, where the combination of SEQ ID NOs of the first peptide and the second peptide of the engineered gene effector in Table 9 are identified in Table 4 based on the barcode, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates a synthetic reporter gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 3 as provided in Table 9, where the combination of SEQ ID NOs of the first peptide and the second peptide of the engineered gene effector in Table 9 are identified in Table 4 based on the barcode, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates a synthetic reporter gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 4 as provided in Table 9, where the combination of SEQ ID NOs of the first peptide and the 47second peptide of the engineered gene effector in Table 9 are identified in Table 4 based on the barcode, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates a synthetic reporter gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 5 as provided in Table 9, where the combination of SEQ ID NOs of the first peptide and the second peptide of the engineered gene effector in Table 9 are identified in Table 4 based on the barcode, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates a synthetic reporter gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 6 as provided in Table 9, where the combination of SEQ ID NOs of the first peptide and the second peptide of the engineered gene effector in Table 9 are identified in Table 4 based on the barcode, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates a synthetic reporter gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 7 as provided in Table 9, where the combination of SEQ ID NOs of the first peptide and the second peptide of the engineered gene effector in Table 9 are identified in Table 4 based on the barcode, optionally wherein the engineered gene effector includes the combo peptide amino 48acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates a synthetic reporter gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 8 as provided in Table 9, where the combination of SEQ ID NOs of the first peptide and the second peptide of the engineered gene effector in Table 9 are identified in Table 4 based on the barcode, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates a synthetic reporter gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 9 as provided in Table 9, where the combination of SEQ ID NOs of the first peptide and the second peptide of the engineered gene effector in Table 9 are identified in Table 4 based on the barcode, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, the first and second peptides are linked by a spacer (e.g., a peptide spacer, such as without limitation, the spacer of any one of SEQ ID NO:2211-2221, optionally the spacer of SEQ ID NO:2211. In some embodiments, the engineered gene effector having any of the noted activation level of the synthetic reporter gene in a cell is selected from SEQ ID NOs: 115-274, 276-287, 289-367, 370-445, 448-746, 748-777, 779-929, 931-1007, 1009-1156, 1158-1194, 1196-1288, 1290-1350, 1352-1451.

[0124] In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the engineered gene effector is capable of activating an endogenous CD45 gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as 49a log2FoldChange) of at least or about 1, at least or about 2, at least or about 3, at least or about 4, at least or about 5, at least or about 6, at least or about 7, at least or about 8, at least or about 9, at least or about 10, at least or about 11, at least or about 12, at least or about 13, at least or about 14, at least or about 14, at least or about 15, or at least or about 16, when the engineered gene effector is expressed therein and effectively targeted to a locus of the CD45 gene, optionally, the activation (e.g., as expressed as a log2FoldChange) is in a range defined by any two of the preceding values (e.g., about 1 to about 5, about 1 to about 3, about 3 to about 5, about 2 to about 4, about 1 to about 15, about 1 to about 10, about 5 to about 10, about 10 to about 15). In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-95 amino acids in length and a second peptide of 75-95 amino acids in length that is heterologous to the first peptide, wherein the engineered gene effector is capable of activating an endogenous CD45 gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least or about 1, at least or about 2, at least or about 3, at least or about 4, at least or about 5, at least or about 6, at least or about 7, at least or about 8, at least or about 9, at least or about 10, at least or about 11, at least or about 12, at least or about 13, at least or about 14, at least or about 14, at least or about 15, or at least or about 16, when the engineered gene effector is expressed therein and effectively targeted to a locus of the CD45 gene, optionally, the activation (e.g., as expressed as a log2FoldChange) is in a range defined by any two of the preceding values (e.g., about 1 to about 5, about 1 to about 3, about 3 to about 5, about 2 to about 4, about 1 to about 15, about 1 to about 10, about 5 to about 10, about 10 to about 15). In some embodiments, the engineered gene effector includes a first peptide of 85 or 108 amino acids in length and a second peptide of 85 or 108 amino acids in length, where the sequence of the first and second peptides are selected according to any one of the paired arrangements of SEQ ID NOs of the first peptide and the second peptide as set forth in Table 4. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates an endogenous CD45 gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 1 50as provided in Table 10, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates an endogenous CD45 gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 2 as provided in Table 10, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates an endogenous CD45 gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 3 as provided in Table 10, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates an endogenous CD45 gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 4 as provided in Table 10, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates an endogenous CD45 gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as 51expressed as a log2FoldChange) of at least 5 as provided in Table 10, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates an endogenous CD45 gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 6 as provided in Table 10, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates an endogenous CD45 gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 7 as provided in Table 10, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates an endogenous CD45 gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 8 as provided in Table 10, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates an endogenous CD45 gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an 52activation level (e.g., as expressed as a log2FoldChange) of at least 9 as provided in Table 10, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates an endogenous CD45 gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 10 as provided in Table 10, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, an engineered gene effector includes a polypeptide that includes a first peptide of 75-110 amino acids in length and a second peptide of 75-110 amino acids in length that is heterologous to the first peptide, wherein the first peptide and the second peptide are the first and second peptide of any one of the engineered gene effector that activates an endogenous CD45 gene in a cell (e.g., in conjunction with the heterologous endonuclease as described herein) at an activation level (e.g., as expressed as a log2FoldChange) of at least 11 as provided in Table 10, optionally wherein the engineered gene effector includes the combo peptide amino acid sequence as identified in Table 4 based on the barcode. In some embodiments, the first and second peptides are linked by a spacer (e.g., a peptide spacer, such as without limitation, the spacer of any one of SEQ ID NO:2211-2221, optionally the spacer of SEQ ID NO:2211. In some embodiments, the engineered gene effector having any of the noted activation level of the endogenous CD45 gene in a cell is selected from SEQ ID NOs: 115-274, 276-287, 289-367, 370-445, 448-746, 748-777, 779-929, 931-1007, 1009-1156, 1158-1194, 1196-1288, 1290-1350, 1352-1451. 53secneuqesdico IVaniYEDP IVEEEIVEDIVEFE IVE L IVEDIVE LEMIEEMGEMWEMEEMDLomGYLY GY Y WGYDY L GYDIY GYNEY GYSY GGYDF PKLIFPKLI L PLPL FPLMKI TKI SKIGnaLIKT Ei eQLIKT TLIKTGLIKT EKF LITFSLIKT LLIKTDFALNTFA NDFACNI FAGFASNV dIQVGIQVCI IQVS I I IVQVLQVGQVMIQVDLKSICSKSICLKSICDEKSNI LCLKSICHmitCapDIVCe eNVDDIDEDCIHCSNVGNVSD NITVSDCILANVLDCIDLDCIAN ANDN GN GNVD NVDDM S WPDMFS DDM SQDMGINSS GDMPVpS TLS T S T P S T S T I S T F S T SdTMPMD MLMRS MPi oNV NVQNVVNVWNVt bVIGLGLGPGPGGNV NV IQ D VII QNVII Q GVII QPVII QRGD GGP T P T PLVKNVKLVKNTSVP TKNVPKGENVII QDLVII QIML NI S IAL NIAILNILILNIHILNPpemo **L LSL**L SSL**L EP**LNS **LH**LA**LD** KG*D* KS ** KTL** KQS**IKSFpC* I C * I T *SI S *SIA*SIQ*SID*SI L * L L * LG* L P * L E * L PlairoDob ed I1 2 3 4 5 6 7 8 9 0tmitQ0ao pe EO1010101010101010111 211111nCip SNbmGC A_CCAT ATAC _ C_ T Co_yC _TT_C_ _C_ GC_A A_T A_ GA_TA_CC_CbT ybT yb CyG TyTy C yTCyG Ty C yTT yTyCC:_G_G_ Cb_ Tb_ Cb_Ab_b_ Tb_Ab_Gb_ CbC_TGAG A AG TC A A A A AGAG_A AC CTC CCAG4eeld 1G__ T1 T_1G _1 C_1_1 C_1G _2 C_C_T_ _ GT_A_T_ C _ GT_ T _T_ C2_ T2_A2_ G2T_Tboca rPAPGPGa OG OTT OC P CAPGP CAPAP C P C PGPGPGAOGOA OAOTC OA G OA A OTT OA OCTB TSAB_TS C BTS TB_TS TB_TSGB_TS TB_TS C BTS TB_TS TB_TS C BTSGB_TSATB_54PESYWGSTNKKNGPDEAYLIHSCVDMLNDNEPSDSNIKDQYSTEEAILLE CDE TWE T E TGE T E T E T E T L T L L LDL LEPMP E CELMD GEDTCGEDFSGEDSGEDLMGEDEQGEDFLGEDD DQGSALQTSAGSQTISAEQTLSAEQTSAWTKLILE PFAQKLI FA DAGIAG AV A DAGLAGHAGDAG ATAL FG L AGM VAGSAGANAFG NAVFGFNAL FGQ NAGFG NACIKSNIGFA VKNDAE EAE LAESAEAEAEAAELADGADGIADPADDADDSADWADDSGADEALGAHS GATSGAVGADENCDSIAG NCAAVQAG AGVAGFAG AVGAVPAVDAVLAG APPAGDEAP EAAEADEA AGT AFT AVT AWT AST AGDMSSL DMDSSAT LNAT RNATG E ATDLAT LV VDATNL ESATMILDEFVD IL P EFVGILFVPEPILFLEVLILFQVLTMVP L TMPGLAVSAVGT LGTHAVGT PIKD VKSAVGTAAVLAVAVGTLGTAGTDAVTAVQAVFAVDAVCAVGAVLDFA LALDAEDA DADDAN AFLAPS F NLASFLLAL FLASLL*NL ILNM AS LASSEASPAS SASGAS L SDEDE F EAE C ET*IL*KL C * IG**KDLLAAPD APS LAA G APSQAA D APTSQSAAG APSGLAA AP ESIAA H AP CSIA AAFDASDAPDAGDAQ APSDD D AEGD A G AEATD Q AEALCD AEGD DAAEIAE LAPL3141516117118119 0 1 2 3 4 5 61 1 11121212121212121G_C C_ G GT G GC GT GC G GC GGGC GT GGyC _AbT yT AGTTATGAGTCATAAAAAAA T AA ACAA AA AT_G b_ CAT ATCB_ AA GB_AA TB_AAT B_TAG AB T_AGT B T_ACB_ACAT B_ ACA TB_ACGT BAG _CABA_CTCB__T2G A_T CA A_TTCA A_ TCCAC_ CCCA A_CGCA G_C CA A_T CCAC_AA G A_CA GGC_ CCA G A_TA G G_CA G A_TT_ T2GP T_TGyAPATbTGGyTb TCGyTb CCGyTbCGyCTbTGyGTbTTGyTTbC C ybC C yb CyGyCy TC CbTTCbT CbGTTG_G_GG_GG_AG_GT G_G_T A_AA_A_A_GA_OSG O AB_TSCAGCBAC_T TA AAC_T CA ACA ACAT CAGAC_TGTAC_TCTAC_TGTAC_TCCAC_TGTTC_ C CAACTTC_GCAAGTTC_ C CAACCTC_ T CAAGTTC_GATA 55EAADDSFLGGPGFLLDEGSSGRRRSGQSGGDPGSAIGEEEAAAAAAALLQTSAE L T FKDKGKDEQGNQG QG QGLQGAFGFQ SSAAGDDT LAGDT LDATNE KADIKAPKAKAFDT EDTLDTWDTDFANGNK NDND NINLANVANLANGANKNAGFNADGLKEA GSKAMKA VEGDEGSKEA GFLKEA GEKATKADY GYFY AY QEGCEGD GDIGDI EGD D GDLY GD SLEALTALGEAA ATLCCHTLCCLTLCG LTLC T TLCGTLCIDTLCL SEYEI E SDS I SYEIIYEIGS IDS IYEI PYEINEEAGIL IT DEASKSQVPKDQVFKC LKC SKCVKC EKCAL W LELQVGQVAQVDK QVGQVDS I TKIFKI LKLILEKLIFSFVGDARILNFVGL QRDAMHC VPCQRD GHCCDLQRHCICGQRRHCWRCPQ PHCSCLQRHC QRVSCLQ HCCGL E FCVRIVDESF LVR T VSESFMVSVR DE FQVSVRGE FG VR LF A F AEAEEAAEAEAEAL EANEA KIFE IFAIFL IFVIFLLDEH QLDEDL AF PDAESEDAEDFGVSAFVDSAFN VHAFNVSAFDF S FMPKPKPD KPD KPGIAVLL A VLA VDLIGLIWLIFDLIS LI GDEAC FGYT EGQG G GTGLGLQ GLP GLGLL GLA A Q A APACG Y GEACYSE EACYGL EACYCGEACYLP EACYDF EAEL LNEAEL PNEAELDL EAEL LDEAEL RN 7282920313233 4 5 6 7 8 9 01 1 1 1 1 13131313131313141GT G G T T G GC GC GT GTAAG AAGGC G GC G GGGA AGA ACAAT CGB_ C CBC CCAB_ CACAB_CG ACBCACGB CA GBCCTTBCCTCBGTA A AA AG CTBG CGBG CAB_G CGBG CABA AC_CTT CTT CCTG_CTT_ CTA_ CC_ CC_ C_T_TA_G_GAC_ TyCGC_A T GC_ CCG A_ GG A_ TCG A_TGG G_CCTG A_TTTGC_A TCA A_TTCA A_T CA A_CGCA G_CCCA A_ TCAbC _TCC yb_ CT AyTb CCAyTbCAyTT bCAyTbTT AyTbT yGA TbT yGA TbC C ybTGC yGbTTC ybC CybT C yb TCT AC _GAAC CAGTC_ T T _AGT _AAC T _AGAGTAC_CGTAC_CCTAC_CC T _AC T _AT T _AGGAC_CCCAC_CGTAC_CT T _T G_G_G_AG_GG_AT CAAAC_CGTTC_GGT CAATC_ C CAGCCTC_ C CAGCTTC_ T CAGGGTTC_GCG 56CELVKKSNHFIDLIVDRNGYDIIFSEIVTHNYMGKLIENFSSFSKLSSFILPQGPQGLTPTPTPTPTPTPTPACDACACACAANLANMFF FIDIFF FIDL FF FIDP FF FI E FF FINEFF FI SGFF FI LDDSHDLD DSHPDSHNEDSSHGDC ESHYNCYNGDDGD DL RPETLFRP GTL S RP LTL E RP WTL T RPTLFS RPTL L RPTL F S SLSDVLNGVNEVNFSSVNLVSNWTS IEIDIL SE DPLE LTPLEV HPLEQ GPLECI PLEG LPLEM DPLED AEES L EQL EIVAEIGAEIGL EML ECL AEID AEI IYKLYI FD IGSKLIDAIWSD PAAIWSD PAIWVD AIWDED AIWLD AIWLD DAIWLATVMHSTVMVTVML TVMLDTDVMEVS VS DPPPV PPD PPGPPGIPFP DGKPGKDGKGIGKFGKGE FVE F LESEWESEPES SESQ ESES PDES P SL V LSL LDL QVRKIHSVRIALL SPLFPPPLL SPGL ELSL EPLLSPLENLSG EPRLSEP D LSP GKIL TEPKT IELLKT IGERKT IED KT IELNIPKF DVL PISMSNSMSEPMSDLLMS S LMSNLHMS L LAMSMDII EGSEDII ESDLDII ESN HDII ESLADII ESSGEPGG GE LVA GLVS LVL LVL LVLVD LVD TTP TTL TTTTD TTLE L E L L ED LEDFA MPT ED A MCGEDTA MLP EDQ A MSE ED A MSGED A MLDEVSA NP EV A NCGEVQ A NSE EVS EVTAL EALM A MFA N G A NLP14243444546474849405152 3 41 1 1 1 1 1 1 1 1 1 1515151GAGG CC C C AA T GA GGCT C CT CC C C GA GGT GG G A AT AT AC A GA T AGA AA AGGG G A A ACAA ATCACT B_CCCB_GGT B_GA TB_G GABA _GTCBAA _GGB_A GAT B_A GCCB_ATAB_ATG AB_ ATA GB_ATAT B_ ATTCB_ACC_ CyC CAC_ATTA_TTT C_ CCTTG_CTTA_TTTTA_ TCTTA_CGTT C_AT TTC_ CCTT G_CTT A_ TCTT ACTT ATTGb CCC yTbCTyGCbTTyCb CCTyCCbTTyCbTTyGCb TCTyCbCTyTCbCGyb CyCyCGbTGb T _CGyG _bCGybTGC _T A G_TC _GCAGGTTC_ TA_TA_A_GA_A_A_AA_T TAAGGTA_ CG AAGCC A_AAGGTA_ TAG AG ACAT T _TAGT _GTAT T _TAGT _ATAC T _AGGGTA A_GTA A A_GCA G A_GCTA A_GGTAC_TGTAC_TGTAC_T CGAC_TCTAC_T TA 57RFAESPSQLAEYEAPSLLDEFKYVDEPKAADDYCPPLYLSDDQLDNHIADCACAGSAG AGGAG AG AGTA YD Y Y Y YDSHVSDI DSHLSD VLAGVM LAD VSLAVV FLALVELAQVLACIVGD ADPWDSDPLWDSDIPWDE PSWWDSNPEWDS PLLNEEF VLNFTEDDSLLDTS LT T T TGT DLTLSDDSHDS SDSVDS EDSAPASGSSPAESF SPAS TSCPASFSSPAS EAELAE DNGGINGFNGSPNGANGDNGGNGDA ALAIA GA QT ITI LGSGSGSGSWGS SGSQGS SRGVRGT RGD RGL RGGVMSATVMATGKFASGRTFAD SD TFASVPTFASPPTFASLLTFAS LT ANF SGL LLH YLSLLL SALLL EGLLL LGLLLV DKTLIWGKDP KLSG SAN EHG SALEAG SAG EEG SA ENG SA EDG SA ESGA EMSQPEYVSQEWYSQEQ YSQEIYQGS E SLD EII E P T IEGT S NE L L SAQSL SADS L S PASL S SAAL S LALL S LSATL SADLQPLE PQGPL PEPQPLLENQPLE RQ LNP E LE T SDIAISVTTMV AYT EV GVFVGVCVLVSAESANQAYTAYTPAYLAYGAYPAYDF RPR SSRASSLRAHSRADL L NL EVDA NLP P GPTPTC PTE PT LPTF MSF MAF MTF MQF MAGDFAALDE FAAL LDFAALQS FAAL IQFAAL IHFAALGL FAALD DPAPDFP PAPDGL PAPDLP PAPDSE PAPDCG 55657 8 9 0 1 2 3 4 5 6 7 81 1515151616161616161616161AC AA CT CC CGCC CT CGCA C C C C GGGC GGT CTAA AA TGG A GATGTT B_ AT CAGCGA G GA G GTT GTG CB_GTGB_GT T B_ GABCGCGA GTGG GA A GGTAC TT_GT T B_ GTAB_ GT C B_ GTC B_GABGT_GGBGTBGBGBGT_GC_GG_A_G_TTGT C_A T A T A_ TCA T A_ GA TC_ CCA T A_TA T G_CCA T A_TT A TC_AAC_ CCAA_TA A_TTAA_ TCG A G_CCTybTT G TybCTT yb TCT ybC T yb CCT yGbTTT ybT T ybTGT yTbC T yCGbC T yGGbTTT yGbTGT yTGbC T yGbTT _AC T _ TG_GG_AG_GG_CG_GTG_G_TTC _GC _ _ _G_GA ACAC_TCCAC_TGTTC_AC CAGTC_ C CAACTTC_ CAAGTTC_ CAACCTC_ CAAGTTC_GAT CAATC_A AC CAGCACATAGTCC_TGTCC_TCCCC_T TACC_T CGCC_TGT58PSPRPMPWSSRPTAPPTDTGWAEPFETRPTAPPFASSDPAFSSHEEPPVVPY Y DS PDLDEIPDDEIPDPDEIEDEISDEI DLDEI NDEI D LPQWTLPQELPQF LPQLL IGPQEWSSGWSD F I F FPWFPGFPFP EFP FANCANQANSANSANFPALSPAFDLY I KEFLYLIKE LY I KT LY I KL LYGIKS LY I KFS LY I KDQST IQDQSTG QVQSTG QLQSTQVQTQLTASMASDPEL PEQPE CI PEMPEVPEGPE DP E P P L PHSP SRGD RGLGLTGLGGLDGLDGLHGLLGL LQ Q G GGDQGQSQALLYL LDLLLAQASQAVQASQ ALAAL D ALEQALQASQA AL D ALPALLQAA N GAL D GQQ GNSGNIGNPGN QL GQGGQVGQWQE FYDSQEDS EPL PEDWE EDSLE EGE F EVDQEDDEDP E E IDGE E S LTL LTL LTR LTDGDINDIDDINDI P LGDTIPPSEDLQPLEGL TPP LPPTP LLPT LPP L NTP L D PLTLGPETLRPTLLQD SQDLQDHQD EQD NYS LYSLYSYS PYS SRAASFRAMP ENP EDP E S P EPP E PPP ENPP EMVT VCVQVS VAPMPDSFPMPD MSLMS SAMLLML MAMS MHMDV EGMS SECMS SETRPFLPVPGVPSEVP FVPGL MS SEDSMS SEFP MS SEQS MS SELR FDT L L T LEIRTFLQRTFLPTRTFL LA D G A D D AL LALG AL PALG AL TAL EAL FA Q G A Q H A Q D A Q Q A QCI960 1 2 3 4 5 6 7 8 9 0 1 2171717171717171717171818181CGC CA GC GT GGGC GGGT GAAG ATATAG ACGA G GT TATA GTTTTATCTG AT TTCTTTA CGTG CAT C TAGAB_GCBTGT B TAB T B TABGT CGC_T _ T _ TC_ TT_TABCT T_ TTGB_ TTCCB_ C C B_CAB_CGB_CCA TB_CCGT B_AA_ GAC_ATA G_T TG G_CCTATG_ TTA G_ GTGC_ CCTATG_C TGC_AA C A_TT A C G_CCA C A_ TCA CC_ CCA C A_TT yGbC T yTC _AGbC yGTGb_ T yT GbTGybT yGGbCGyb CyCGb TyTCGbC T yCbTGT yCbT T yTC bC T yCb CCT yGC bTTCC _ TTCT_GTT_T_AT_GT_GT_TT _ T _GT _GT _ T _ACATAC_TCTCC_TGTCT _ TAACCCT _ TAAGTCT _GAT TAACT _ C TAACTCT _ TAAGTCT _AC TAGCT _ TAG ATA AGACAGTAC_GTAAC_GGTAC_GCGAC_GGTAC_GCC 59PAAYAQYANEPAPQPASFWSGTPATSVHYATLQGQPLPEPLVPWLV A ML GYIL G L GCGPL GD L GAIGNGGPIPI LPIF PIHPGDVG SCGDL F PANKF PD ANGF PANLCF P LANAF PLANKSF PGLANL F P LANMMLQ RYQ ML FRYNMLRYGMLD RYVLAPQGPLLA PQDF PSPV MFE PSP LMDP PDP PDP P L P PGP PDANMANDRVDIRVPSM LRVD SMSSM LRVGRVNESRME SMLMGERHMGEGDGMVE RVDKQVKRWMRVE MR PAQ VKQK VQ LQSTDQTD NT ENT ENTGNT LNTFSNTWNT FKEKE LKEK DE C LPQ QLSQL EAGDPQAANFEA L ANQEAS EAMEAEAT EADKAAKAQKA GANVANDANGANCANDHMAHMNHMFKAD L HMDD GGNLQFGN DGQDSLLVT LVLV QS LQVLQHLSLV QLLDLVL LVILVLSQL LQ DE LQATGASTATGSSTNTGTDSTGTLS LDDTIDLQDLDTIGQ LLSAQ ALSDQS PQS FQ ASLAVLADLSGIQSQ A AQDMFGLAGLSD GAAGANGAHGAG AS F PYAFYKFYLS FGLYSYSSPDW SPFPSPD LSL FSPDP FSGSPDSDFLSPDS R FSPD QFS LS DSGG LKGAG KGTGIGY YKGEKGGL EQVVDRPTFSVVDGRPFLDDVCNSDVDCLL DVECPDVN CADVCHDVNPCSDVCMR PAR PNR PERPD DSHD ESH G KSHSSH GV ALQGL TALQFK DFCGTAK GFCGT CK GFSK CGTFPF DCGT SK GFQK CGTSEFCGLTTK LFCGT L EDPA CTT E EGPA CTTVETPA CTA TEESPA CTT LDDS38485868788 9 0 1 2 3 4 51 1 1 1 18181919191919191ACACC CT C CC CT C C C C C CVTCACATA TB_ C TA AA AGA AG AGAAT C TATATCCB AA CGBACGCA BACAB AABAAB ATTBATCBC CGGB CCGCBCCGCBCX _GCIACACACC_ATT_T ATTA_T GCT T_ CT CC TT_CG_ C C_CC_TACTTAT TTATTT CACC _T_yGC _ GC T CT_GCACGGTCT CPEC bC T yT _bCAyG _bTAyC _bTAy C_b CAyG _bCAy C_b TAyT _bTAyTbC T_yGbT T_yGbT T_yG_bGT ybT _BAAACCT_TG ATA_TG ACA_GG ATA_ACG GA_AAG CA_ CG AGA_GG AGA_TG ATA_AG ATA_AG AAA_GG ACA_A_74C _GCTAC_GGTAT _GCCAT _GGTAT _GGTAT _GCTAT _GCGAT _GTAAT _GGTG A_TTCG A_TTTG A_TA G G A_T.160LFYGSRARAHEAIKSLNEFSGLLGIGRNHQSEQDECPQVRVEHCGPMRM MRMRMRNM DM AMEMWMEM GM M MEM HKRPKKKLK KREKRKRRKR IKR LKRYKRVKRWKRDKRTKQCKQEKQNKQ LFKQLKQ ME KQLQ D QAQ D QLQSQIIQFKEAAEAKAAETKAKEVKANKEASEEKAEEQKE EKEK L KAAKAGKAQKEAIIKEGKEF KAHKAGK KEAFK I KEADIHMEHMDHM HMGHM HM HM HMPHMKI HMSHMQHMTH AH DSTS SAS C S KSAS ASPSI PM M GAGR GE GT ES E SVSASNS VGTFTFALGT S T I T F TGTGVTGTGTTGETGI TGFAHGTATGTGNTGRFAMGTFANCGTG FADGTFAVEGTK FAD GT SFAD GTFASGGTFAVGTFAGGTFATSGTFAYGTFAKGYAGYHIGYLGYPGYKGYGGYEGYPGYEYTYPYV YRYSKRGPPKGL KGLGIGS G GE GGGQGGHG NGQGGHGGSGGK PP I PL KPL KP LKPHSTKP EKPGKP SKPKPVKPYKP LKPQEHDRHNRHEIRHC RHSRHRHNRHF RHRHYRHL RHRGRTAR SEAGSLS DISRS LPS D S YSVS SMS GSA TSHNSHSPPGEA THEAEAQEAEA KC T P C TAC T R C T S EA GC T F EA DC T T EAGEA NC TKC T R EA QC T L EAYEAGCTTEPPCTTGPCT P T VP T P T T P T P T P T P T P T P TQPCTTNPCTTSS697989990010203 4 5 6 7 8 91 1 1 1 2 2 202020202020202CTA C TG C TGC TA C TG C TC CGCT G CA CC CC CT CGCGCCCTC T CTCTCATBC TTCTTCGTCG TBCATCC TCCTCGGCB GCB_GGB GT BGTB GG_GABGT GABC _AB_ TB T BGBCG_ CACA_ CA_CC_ C TACG_CTB A_GT GCGA_GC _GA_GACT_CGGG _GGG_ TTGC_CGGT_A A GG_ GGT_CCGA__GCGT_ GCCGCG _GCG GC_T CGCGC_AGA_ CCTGA_G T GybATGybC TGyb AT y T y T y C T y C T yTGb CGb TGbAGb CGTGbCGybATGybATGybGTGGybT TGybGTGybT_CA_G _ AG__GAC_ CAC_ TAT_ _ _ _AG CCA GACCATAATA A A A A AACA AA A_ _ _ _ TA_AG GG G A_T CG A A_T TGTG AAG AGG AGG A G A A A A A G ATG ATA ATA ATA AT TAT CG AG A ATG ATACG G ATG AT T_ _ _ _ _ T _ _ _A_TGTA_TAT61LFYGSRARAHEAIKSLNEFSGLLGIGRNHQSEQDECPQVRVEHCGPMRM NMRMMR SMRDMRM M M M MIMEM GM MRKRAK K KPKYESKRDKRDKRKKR EKRAKRLKRL KREKRPKQLKQKKQ KKQLKQKQ DKQ EFKQ VFKQ N QRQLQLQSQTAE E E E E E E EKEKEKEKEKKNHSMCK G A HSKACGKAVMSHMHHMLKAQRL CI KAAKAKKAKKADKAKLKAQRKATYED L KAGIKAGTGHS GFIS GGS GNHSM GEH PSMDHMYHMEHMEHM HM HM HM GSS GVS GLSTS L SKS DFSSH PSMLSGTACTGTLTLGT PLTD GTITEGT S TAGTGTLGTDTNTGVTGL TGGTGTGI TGE GTI GTQ GTYGT SGTDGTKGTGFADFAFAP FAFAFAFAP FAS FAFAS FAAFADFAIFAHGYIGYVGYPGYKGYLGYMGY GYGGYEGYLGYK YLYRYVKGYKGVKGCKGK GSEGD GKGLGFTGIGG MGAGGAGGRR P L R PRPDP SS KPKP LKPAKPNKPKPYKPKPE KP LKPESEHV HP RHRHRHKRHDRHARHC RHKRHQRHRI RHS RHYR TAKSEADSEALSE LS SNSFSDS S S L S FIS T S P S GSHAPSPDPA A SC T P EAKEADEADEAE EAS EAI EAS EAT EA KC T E C TDC TYC TWC T E C T I C TMC TMC T E EACTTQCTTMCTTF P T P T I P T P T P T P T E P T Q P T P T P TDPCTTGR01112131415161718 9 0 1 2 32 2 2 2 2 2 2 2121222222222CT CC TA C TC C TGC TAC TA C TG CA CC CA CTTCG CG CGC C CTB CAC T B C C CTCT TCATBCA TCGTCC TCA TCC TCTGGB GTGAB GA_ GAB GCB GT BGAGCB_GAB GC CBCBGBCT_ CA_CC C _ CGT C T _CC_ CA_ T _ A C_CBGT_GG_GA_GCT_ GCGT_ GGA_ ATGC_AGGA _AGC_A TGC_T CAGT C_GCGCGC_AGC_ GCTGC__T CGC_ ACTGA_ACGGC_TG GybGTGybT TGGyb ATGybAT y T y T y T yGT yGGbGGbCGbAGbGbC TGyb CTTGybATAGyb ATTGybT TGyb AA_TT A_AA_ CA_A_CA_TA_AA_G A_G A_A_G_ _T_ TG AGG AAG AGTG AGCG ACG ATATAA AG AG A A AAA ACA AAA_TG A_TA A_TG A_TG A_TATA_TGTG A_TTTG A_TGTG A_TCCG A_TGCG A_TGG A A_TGCG A_TGTG A_T CA 62LFYGSRARAHEAIKSLNEFSGLLGIGRNHQSEQDECPQVRVEHCGPMRMKR EM DM M KRVKRDKRVTMEMPM M AMEM WM KRSKRGKRQKRKRGKRKKRFSMRAMRQPMRSSKQ GKQDKQ DKQ EKQFKQ SKQVQKQ L QPQK GQ TKQ QKQGKEAESKEALKE FASKEAS EVKAME P EEKALKARKERKAVK RKEALKEKKEGKAKKALK KEATK S KEASK I KEARHSMLHMPHMTHM HMPHMSHMRHMSHM HM HMLHLHSHMRTGKSTVTGPTPSLTGDCSTGD ASTGTSTGASSTGPSLS QQTGMTGLSA TGASTGGISMR SMAS RGTGR TGATGSGPGSGTGTE GTMGTVGTGT LGTVGTS GTGTP GTI GTGFAFAFAMFAE FAL FAFAE FAFAP FAE FAR FAFAS FAQGYD GYLGYLGY GYKGYAGYEGYTGYGGYDGYN YRYHYSKGW PYKGLPPPKGSPD KGMPIDKGP ENKGG PS KGEPE KGRPSIKGPPTKGPSRKGPHG QKGQG PGKGP TG PKGPGRSR R E R R R T R E R R R R R R RGEHG H H HPH H H HPH HG HSHATPAF SSEQ AASEALISCSMS SSS E SLSWF S D SES S SHNSHDPCTTSIPCTTESPCTTQEQPA CTTQEVPA CTNETWPA CTTGEPPA CTETEEEPA CT STI EGPA CTT T ENPA CTT P EQPA CTTQEDPA CTLTFESPA CTS ETMPA CTTGS4252627 8 9 0 1 2 3 4 5 6 72 2 22222223232323232323232C C C C C C C_C C C_TA C GCT C2GG3CT CTCGCCCGTCG TCC TCCTCGTCA TC_rTCCTGT _rTGTTTTT AGAC C B_ GTBA_GTBA_GTTB_GTT B_ GCGB_GekA nGAB C_GA GBC_Gek CABCT BCGB CnGG_GC _GT_ GCBT_AGGT_T CyAGCCCT_y CGAGCT_yA CGG T_ GCy TGA T_C CyAGGACT_yC Gi l CT_yGC TT_yCCG GG T_yCCGi l CT_yGATCT_y CGAAT_y C CG ACT_yACGCT_yG A GbAGb AGbCGb CGb AGbCGbGbCGbTGbGb TA GbGb G bA_GA_ CA_AA_ CA_ TA_T_ T _CA_A_GA_A_ _ C _A_ CG_AG ACG AA G AGG AGCG ACTG AGG A G ATACA A A AGA ATA AAA ACATA ATA ATA ATA AT BATACG ATACG G AT BAT CG GG G ATA AT CG A AT TATAT_ _ _ _ _ _ _ _ _ _ _G A_TG 63LFYGSRARAHEAIKSLNEFSGLLGIGRNHQSEQDECPQVRVEHCGPMRMRDM KMPM M MFM YMMM DMLM M VMTMKQLKRQRKRQSKRG QAKRQFL KRQKKRQQKRQTKRQAKRQAKRQLSKRQLKR IKRHK QKLKLK KV E DI R TQ NIQNKEAAKEAAKEAHEHPKADTKETASK KEASK KEAIK T KEALKEGKAKKESKAEKETKAMKEPKEKESDKAPKAAKASSMGHTSMAFHSMGHSMGHSMAHSMSRHSMESHSMLHSMAHMAHMPHMTPHMMIHMQ GD G G GAGWGF ESES VS S S AGT SLTAGT E TM GTGTGTATPGTPP TV GTNTGGT TNTGGTVTGGTEI TGGTLTGGT T TGGTA ATGGGT S TGGT LFAL FANFAFAFAFAVFAFAKFAFAL FAVFAP FATFAMGYPGYLGYAGYAGYNSGYFGYYFYA YAEYM YSYLYEYPKRGEGSGVGRGAGLGG YGVGGEGGDGGPSGGSGGKGGSPHKPQKPGKPGKPKPL KPKP FKP IKP FKPKPVKPHKPDSEHS RHTRHARHDRHGRHS RHKRHHRHKRHARFR R L RA AHSEARSEAS R SLSKSVSQSDSNSHLSHGS SH YSHQPCTTV APA CTTQRPCTTGEAPA CTT R EAPA CT CTI EQPA CTTGEDPA CTK TIEAEA IIPCTTYPPCTTKENPA CTTEERPA CTD TLELPA CTT P EDPA CTT T EDPA CTTD Q 83930414243 4 5 6 7 8 9 0 12 2 2 2 2424242424242425252CTG C TA C_T1C TC CC CC CG CACC CTCA CA CC CGCACAC_r CATCATCATCC TCCTCATCCTCATCT TCAT CGC B GTB GekGGB GGBGGBGT B_GBGBGBABCBABC CBC C _ CT_ CnCA_ T_A_ ACGG_ GC_ GG_GA_GT_GA_GG_GAT_AGTG _T Gi l_ GACC_CGA_T CG G_CTCGT C_AGC_T CTGA_C CGGA_ TCCGT_ACGT_ TCCGATC_TGCA_CGyAbT TGybT T y T yAGbGb GT yGT yTGbTGbC TGybATGybC TGyb CTGybGTGyCbT TGyb CTGyb ATCGyb AA_TA_A_A_A_TA_TCA_G A_GA_ CA_CA_AA_ T _ _ TG ATG AG G A G AG G ACA G AGG A G ATTAA ACATATA AGA AAA_TATA_TGC A_T BA_TGC A_TCC A_TA A_TAC A_TG A A_TCTG A_TCTG A_TTTG A_T TG A A_TAG A A_T CG 64LFYGSRARAHEAIKSLNEFSGLLGIGRNHQSEQDECPQVRVEHCGPMRM MKR TKRFLM M KRNKR LM VM VM KRVKRRKRAMWM KR TKR ST M KRDLM M M M KR EKR EKRDKRELKQENKQQKQRKQ FKQ EKQTKQ GKQ CIQPQDQQQ PTQKQ IAHAEAE E E E E E EKEPKEK K K KKLK A VKANILKAKKANKAV GKATAKADKAPKAYEGEL KAVKAVKEAAE EEKAKISMSHM HMAHMLHMYHMTGGS GKS GKS GGS GKS GPHM HMEHMAHM HM GS DHMAHMSH GAS GGS S SESSSPSISMSEGTFS TGTNSTGT I TVGTNTGGTKT TATLGTVGTQGTQTGT TGYTGLGTVGT LLTGATGVTGS GT LGT PGTYGT EFAVFAE FAFAFAFAEFASFANFARFAFADFADFAFAVGYSGYEGYDGYVGYYGYEGYLGY GYLGYPGY GYLYKYVKRGTKGFKGAKGVGDGVIGG GSLGQGVGLGPGGFL GGPPGPTPDSI PVKPQKPKP LKPT KPP KPYKPL KPQ KP PKPDEHPRHS RHRHHRHRHGRHGRHRHS RHL RHC RHP RSR PASSEAASEAYS S VLSAI S L SLP S P SDSGS VSHFSHHPCTTG NE EAVEAEA GPCTTKPCTTSFPCTTTTPCTTVSPCTT E EHPA CTT T ENPA CTT L EGPA CTE ETWPA CTTDEDPA CT ETI EHPA CTT P EAPA CTTYEKPA CTTY N 25354555657585950 1 2 3 4 52 2 2 2 2 2 2 2626262626262CTA C TT C TT C TC C TG CT TCCTTCGCA CT CT CA CCCTCT BCGCAB CACAC TTCG TCTTCC TCC TCATCC TCCTCGGT _GGB GC_ GCB GGB_GC GCGTBGTB_ T BGB C BGBTBCCCT_ CAC CG_ C C C T B CT BC C_AGA_ GA_ GG_ GG_GC_GTGT_TyGGCT_ TyTGTT_yACGGG T_yAGGCT_y CAGA T_ _yCGCT__yACGA T_TyTCTGAAT_y C CGGTT_T CyGTGG T_C CyC GA T_TyCCCGTT_ TCy TGTGT_y CA Gb_GGbG _TGb_GGbG _CGb_GGb T_TGb_GGb_GGb_AGbGGbTGGbAGbG AGbAA GATA GA A A A A A A_A_ _ _ _G AAG AGG AAG ATG AAC G G A AGAGAGACATA AAA ACA ACA_TG A_TA A_TG A_TGTA_TG A_TACG A_T TG A A_T TG A A_T CG A A_T TG A A_TGTG A_TATG A_TCCG A_T CA 65LFYGSRARAHEAIKSLNEFSGLLGIGRNHQSEQDECPQVRVEHCGPMRMLM MSMAMSMRM MDMLMFM M YM VMSKRQEKRQQPKRQGKRQLKRQNKRQQKRQLKRQPKRQQKRQTKRQDLKRQKRQMKRHKDK KFKSKRKTK M NTD Q Q LKEAF EQEH YKAQKAVKEAP EDKAGETKAQELKADKELKAAKEFKAQK KEARK KEANLK KEHASK KEAYSK KEALSSMLHTSMLHSMEKHSMSHSMWHSMAHSMDHSMIHSMEEHMDL HMDHMAFHMPPHMPGTGG G GDG GP F ES S E S S S SGTRTFAEGT S TFAAGTVFAE TGTLT R T TGTGKTGATGFASI GTFAAGTA FAAGTD FAD GTFAITGTFAVEGTMTGT TGKTGQTGLFAEGTFAMGTFAQ KGTFAQTGTFAAGYSGYSGYKSGYPGYMGYTGYLGYEGYDGYV YSY YYYGKRGSPLSLKGQPSKGPDKGPPIKGPLTKGP LAKGPA DKGLPDKGP LKGSG PCI KGEG PL KGAG PYKGP FG NKGSPFLEHRHVRHKRHEPAYSVEAPIS PISF RESHDRQSHARSHS RHKRHGGS D SS RSHC R T RAR RQSHT SHESHLF SHSSCTT ITYEAKEA DPCTQPCTTVPCTT E EDPA CTTDENPA CTG TPEEPA CTTGELPA CTT P EDPA CTY TLETPA CTTDESPA CTTTLEEPA CTVETMPA CTT R EKPA CTTH D 667686960 1 2 3 4 5 6 7 8 92 2 2 272727272727272727272CTG C TG C TG C TA CG CC CC CACTC 1 CG CG CC CACTCCCTCT TCC TCGTCATCCTCCTC_PTCG TCG TCGTCGCB GT BGTB GT BGT BGABGABGGBGABOTBT BGBCC BCC_ CA_CG_ CG_CA_CC _T_A_A_GT GA_GA_ GC _GT _GCGT_GGA_CGA_ T GA_TTGG_CTGT_C CGGA_C CGGA_TACG A_ GCCGS C_GGAC_CGA_C CGGAA_T CG A_ GTGybGTGyA bTTGy CbGTGybCTTGybT TGybGTGybC TGybT T y T y T yAT y T yGT yAGbGGbGb TGbAGbAGb GA_CA_ CA_TA_ CA_GA_AA_AA_A_AA_A_ CA_T_ _ TG ATG ACTG ATG AGG ATG ATG ACG AGC AG A A ACA AAA ACA_TCC A_TA A_TTC A_TG A_TAC A_TGC A_TCTA_TG A A_TTTG A_TB_G A_TGG A A_TACG A_T TG G A_T TA 66LFYGSRARAHEAIKSLNEFSGLLGIGRNHQSEQDECPQVRVEHCGPMRMLM M VMRGM M VMEMAM YMPM QMG DVKR LKR TKRQKSKRLKRKRKR SKRKRKRDKR PAQLAQHKQ DKQGKQ S QVQTI QWQ DIQ SIQ LIQV LQ I Q Q DGD DGSKEALKEAWKEAIIKEKEH YKAHKARK KEAQKE IKEKEDKADKADKAHK S KEALK KEATK T KEAL SGV DF SV DGDPVSMLCHTGSMA E HSMDHSMSPHSMWHSMFHSMGSHMTAHMCHMHPHMLHMAT L SDL S PGGPV VI S SVS S Q SMLLMLGGTFAETIGTFTGV FAEGTFACTGVEGT P TGFAGGTATGFASGTDTGFAEGT S TGSFAGGT T TGFAQGTDTGVFAMGT E TGFAAGT S TGFAHGTMFAPDLSLDA DDLLD EPGYHGYTGYKGYEGYVGYVGYAYQYKLYEYG YPDL SDLSFRGD PVKGRPP KGP TS KGPPS KGPQIKGCPDKGSG PS KGP LG KGG G PNKGPTEKGPMG LKGPSRFLDCG GFLDCPTSEHMRPAPSEHTAARPSEHL RAD SHFPRTSHVERSHQRHI RA HRHDRAS D S Q S NSHNRLSHGRSSHMEDG LEI LDDG LEIQSCTTLCPCTTPFPCTT T EQPA CTTQESPA CTV TSESPA CTTQESPA CTTTAESPA CTSTREEPA CTETPESPA CTTVECPA CTT ETMPA CTTQTAH D DGLAHQ D D G 0818283 4 5 6 7 8 9 0 1 2 32 2 28282828282828292929292C C C C C C C C C CC CTT T AGT G T C 2TCTCTTCT GCC TCG TCA TCCTCC TCATCG TCATC_PTCT T A TGCAC CGT BGCB_GTB_GAB GC BGGBTB CBOC C T CABT ABTABCA_CC TCCCT_G_G_GG_GT_ GT GAB GC B GT _T T_TT_GCT_TGGC_CGCG _AGC_ CCCGG_C CGGG_GCG G_ TCTGG_T CGS C_GT__ CGGT__TCGT_CAGT A_CGGTC_ CC Gyb CTTGybGTGGyb AT yCGb CT y T yCCGbGGbT TGybATGyTbC TGybTGyb GT y C T y T y yCGbAGbAbCAb CA_GA_GA_A_A_GA_GA_GA_CA A_A_A_A_T_AT_ CA A AGAGC T C TGA CCAC CAGG AG GGTG G A G A G A G A A A A A A_TG A_TA A_TA A_TGTA_TTTA_TATA_TTC A_TCCG A_TB_G A_TATG A_T CG A A_T CGCC_ACTCC_AGT67VMPLCDDLSGGLMDLDFDDLADSGLMDLDFDDLADSGLMDLDFDDLAQ AQTAQGAQLQ QTQDQN Q Q QFQQ QHQDSGDEDSGSA ADSGVDSGLDSGAA A A MDSGQ EDSGTGAA GDIA GPGDS FI DSDDPA DGSA A A TDGQLDGSG ADGHGVLDG V SQ GDW VDP GDSGVGDIVIV GGD GGDEV V V AGDAGDEGD VSVA RGD SSV TGD DSV CGD GSV SGDF SV KGD GPMLL L SLNMLP L SLL SLNMLL LMLR L S S L S L SAL SK LNMLLTEMLLVE MLL PML IL S L S L S L S L S L SL TML KSMLVMLM L ML ASMLQ KML LPD S DSD D D D D A EL LRL L L LPLDL LDLDLDHDKD D D DRD DLDDKD DLD DSD D Q DDAD DCDLFL T DLA CL FLCGLDLLLFLCCDLGFLQLCSEDLFL HLDLC L FLL LCGSDLFLGLDL LCDFLDDLQLCKFL TSDLFLQLPSDLDLFLEDL S LLFLVPIDLFLYLADLFL DLDPDGLDCIDED QDGYDYDRDDDCGDCPDC IDCYDC EDCAFLEIGDG LEIQDGILEIHDG LEID DTLEIDDG LEI LDG LEI RADGPLE DG ESSDG EEDG EQ QDG EQDG EVDG ESAHLAHEIAHDAHEAHMAHTD DPD D V D DSAHDAID HALAIHLAIHWLAIHI LAIHIN D D D D D G D DSD DED DED DR LAIML IED DPPD D D D D VC W S SHAAH Q D D N D DSV 4959697989990010203 4 5 6 72 2 2 2 2 2 3 3 30303030303TCG TCTTTT CTCAG A G G G C A A GTATTTTTCGC B_ TCGCACAC CTATACGC CTGT T T T TA GB CGCCTBCCTBCCT CGT B CATA_T TGGT BA_ TTTABG G_ TTGB_T BCGATTA G_T BATTA_T BGTA_TCTA_ T BTTA_TTA_TATA_TBTGTA_TA_ T BCT C _A yTT_ GT_ CT_C T_TGT A_CGT A_CGT A_ AGT A_G TGT A_CGT A_AGT A_CAGT A_GGT A_TAT bTGAyC _T bTT AyTbTAyTT bCAyAT bCAyTbGAyTb GT AyTbTAAyTbTAAybGyAAb CyCAb TAybAyT Ab ACC AC _GAT C _A ACC_ C C _GAACCCC_ T C _A AGTCC_GAC C _A GCC_G AC _A A_A ACC_GCAATTCC_GC _A AGCCC_GC_TAAC _AGTC_AATC_ACCTC_ACTC_AGACACC_AATCC_ACACC_G A ACC_ATACC_AACCC_ATG 68VMPLCDDLSGGLMDLDFDDLADSGLMDLDFDDLADSGLMDLDFDDLAQQAQ AQ AQI QTQ Q QY QEQSQFQLQV QEDSGA GDSGKSDSGYSP DSGIF A A IDSGSLDSGAE A DSGTA A A VDSGGI DGTDGPA DDGVALA EDGS DGGA PDGELGVLDSDVAV SGLDS EEGLDPVA V SQGLDSNGLD RVA SRGD F GV DAPGV DSPSGVADE SGV DS SV PL S L L SAL S I L SLL SDGLLDKSSVGV LDPSSSV LGLDGSSVGV LDA SVML LMLIML QMLNMLMLAMLPMLKMLLMLSMLELAL E L VDL LP DLADL TDLYL R LPLDLIRLMLILKMLGML EML ELD DED Y DRD DQD DPD D L D D DD D D DPD DSD DSDD VID DGDL E LFL HDLELLIDLLF LNDLSLLLDLGLLADLDLLR DLLP LDLALLLDL LLFDLPI LEDLDLDLFL LLDLGDLHTDC SHFDCK DFDC LDGFFDCGFDCSFLDC PEFQDC P FY DCGFA DCNFLDC FEFLK DCPI FLDCSS FLA DCIFLDCLPLEIVDGK G LEINDLEI RDGN LEIYDG LEIFSDG LEI PDGVG LEI PD EEDG EERDG EEDG EKDG EHDG EEHDG ETA AAHVAHDAHKAHNDSD DIAH V D DMRAH D DARAHA D DPLTAIHDD DT LAIHIK D DS LAID HLAIHVY D DLD DL LAID HGLAIHGLAIHRA D DTD D G DQ GK D D D D DAID D N 809001112131415 6 7 8 9 0 13 3 3 3 3 3 313131313132323TGCTC CTC CTCGA A G CC T C C C CTCA TCG TCAC A A GT T T C T C T C TCT T TCT CTTATTCGCB_TTGCB_ T GB T T BC C CGT T C CBTA_AGAC TGC _TAAT CB_TT C _T BATG_TBCTG_ T B T B TBT B T C B T CBT G_T TG_ TG_ TG_AT_GT_TGT A_ CTGT A_CGT A_CGT A_ C GT A_AGAT GAT TGG A_ T T _ T T_T TA GT_C T_ TT_ TCGT A_ GTGT A_CTGT A_ GAyAyCAyGAyAyAAyAAy CAy TAyGAy CAyAyAy T y CT bTT bCT bAT bGT b AT b C T bAT b T T b b T bGb GTbA GbAC _TA_ T C _A_AC _A_AT C _TA_ T C _A_ TGC _A_ TGC _A_AC _AC C_ACCTC_ACTC_ATTTC_TACC_TAAC _AACCAATCC ACTCC A GCC AGTCC A ACC A GCC AATCC_AGTCC_ACTCC_G A GCC_ATCCC_ATACC_ACTCC_A A A 69VMPLCDDLSGGLMDLDFDDLADSGLMDLDFDDLADSGLMDLDFDDLAQSAQFAQGAQ Q QTQG QDQD Q Q QKQLQWD I VACA DATA A AFA KA AQASGSDSGMEDSGPPDSGRSDSGDDSGGDSGPDSGLLDGLPDGNDGYDGLLDGRDGA EGVALDSAI GV LD PV STGLD TV SSGLDLSMGVELDTSVGV LD LV SSGDA V GL STSGLDC SV SGGLD PSVGSV SPGLLDSCFGLDVSV SDGLDL SV SYSGLDKSSGGV SLDPSFMLSML MML DMLLML QMLMLVML EIMLSML NML EPL L LEDLHLDTDL LLD KDL PDGDLDTRDLDEF DLH DVDLM DHDLDHLDDD LLDLDCP DLDKDLLDIM YDLAM DKDLTDRPDLPLTDLLE LDLGLLFDLLSI LDLLT LDLLR LEDLYLLADL LLVDLLPLPDL I LLL DLALLADL LLQF DLMLDL TFDCNS FDCN MFDCYF FDCPFGLDCK LFSDC T FADCTFLDCMFPDCQF C FDFIFL RI FL AADCDDCDDC IDCTDC PDGMDGNDGDD SIDGDGDGDGL DGDGIDGDGQIDGSDGPLEI LEI LEI LEI LEIELEIGLEIQLEI CEEEVE Y E E EFAHASAHW T AHDEAHGD D DLAHEY D DALAHRD DTEAHGL ID DQPAHDAHSD D D D DP LAIHPG D DI LAIHC LAIHNG D DPD DT LAIHM A D DRLEAIHA D D A D DFD DPD DSS22324252627282920 1 2 3 4 53 3 3 3 3 3 3 3333333333333T G TTT T GCGT T C A G A T TCTT C CGCT A TTT T T T T TGT TTGBGTGT_ T TBBCAC CBCGB CCT CCT B CGT CTT CTT CACCCB CGCBCT T_ T T_TBT_TTA_AGT A_C T TGTA_TAGTA_ T BTTA_TTA_ T BTTA_TBCTA_ TTAB_ TTC B_TTC _ TTTC_T AGT A_ C GTC_TCGTC_ AGTC_GGTC_G AGTC_GGTC_CGTC_CGGTC_TAGTC_ GTGTC_AGTC_CAyGAyAy CAy CAy CAyAAy TAyCAyAAyCAy yCyAyGT bCT b ATT bAT b T bGT bTT b C T bTbCb bAAbTAbGAbGC _A_ACC _A_CT C _A_GC _GA_ T C _A_GC _A_A_C CA_ C C _AGTC_TAAC _ACCTC_AATTC_AGTC_TAGC _AGCCA GCC A ACC AATCC AACCC ACCCC A ACC AACCC_A A GCC_AATCC_ATGCC_ATTCC_AGCCC_AATCC_G A A 70VMPLCDDLSGGLMDLDFDDLADSGLMDLDFDDLADSGLMDLDFDDLAQLAQDAQ QSQLQSQLQ Q Q QI QGQLQQDAAIA A A A GALA QAIATA ASGADSGFYDSGADSGIYDSGVLDSGQ ADSGG LDSGR DGTLDGNDGF DGADGNCDGAGVLD DSSGGV LDSLVA V TGLDSAGLDD VN V SVGLDSDGILDLSMGV LDFVR SV SVGD R GD DSGV DS SVSNGD ESV KL S S L S F L SNL S IGDASVHSVV VL SAGLDSAGLDSK NMLL LMLLRE MLA LAMLLCE MLL EMLLPSMLLAMLLGMLDML KMLTMLQSMLCDML SDMD D D DKD QLDL T L L L I L ELDLD SS LDALD K D DDD D VD DSDDLD D D DHD DLDD D D EDLDFLLDLLL DLLADLLT LSDLKLLSDLLALDL F LLHDLLGLDLLALDLYLLNDLNLLYDLGLLLDLYLLLDLLFTDCDFDGFDC L FDDGY VDCD EFDGEDC L FDGDDCSFTDGLSDCQ DFDCQFPDGQDGYDCG DFDCES FKF FPDGPDGPDCTDGVDCMDCG TDGGDGLFV DCTDGKFSDC SDGALEI LEI I LEIGLEI LEI LEIVLEI LEIGLEIE E K E E ENAHDLAHDAHVIAHQAHKAHDAHKAHSAHMLAIHQD DSD DED D D D DTD DID DPD DP LAIHL LAIHNK D DED DL LAIQ HLAIK D D A D DQID D VV E L A S IN D DNR H D DTL6373839304142434445 6 7 8 93 3 3 3 3 3 3 3 34343434343TATGT C TAT GGAC G C TCTCT TCCTTCBCTC TTGCTATBCTAB CTC TCBC CT T TAT T TGCABCACCCGBCGCCTCGTGC B_ TC_T C_AGCGTGCBC_TTGC_CGTG_TTG_TBTC _TCBATG_TTATT_TCBATT_TC _TCB_TGB_TGBATTGTT ATT GTT_yTT_GT_ GT_AGTC_AGTC_CGTC_TTGTC_GGTC_TGTC_ GGTC_GGTC_CGTC_CGTC_ TTAT bCAyC _T T bGAyTbTAyTb ACAyA TbGAyTb AT AyTbCAyA Tb AAyTb ATAyTbTAyAyTbGyGyG GA TbAb TAb TC AC _ T C _ACAGTCC_ T C _AAACCCC_ T C _A ATCCC_GAT C _A ACC_GAC C _A GCC_A_G_AC CAGCC_TAT CAC C_AAC _A A_ACACC _AGTTC_ATTC_ATACC_A A GCC_AGCCC_ATTCC_ACACC_AATCC_G A GCC_G A A 71VMPLCDDLSGGLMDLDFDDLADSGLMDLDFDDLADSGLMDLDFDDFDWFD DFDGFD RFDS FDFDTFDFDS FD SFDFFDNDPMLLPDPQDPEL D L KD LFDP LDP SDPDDPQDP EDPVDPSKD LND L ND LLD L VD L GD LVD LED DP PLGDPSD LLD FDPLD LAQC L L E EAQAL EAQVL EKL E L ENL E L L E E L EEI L EGL ESAQVAQRAQLAQGAQSAQNAQLLAQP L EL AQEKL EAQTIDGDDDGTDGS DGADGGTDGDDGHDGLDGYDGDGSDGLDGRSGVDGL SVDSVR SVESVSVESVQSVKSVKSVGSVASVGSVWLD LSSGD R GDVGDIGD WGDTGDVGDVGD GD Q GDSGDNGDVMLGDP L S S L SL MLHMLNL SML S R L SML SAL S P L SKL S L L SVL SGL SML LMLAML SML GMLDMLLML VPMLAML AVMLSDLGEDLHIDLVF DL LDLMDL EDL PDLWDLYDLDLGDLDLVLD GLD LLDL LDLLD LLD LLDQLDYLDDLDGPLD SLDVLD QDL LQ DL IDLL DL EI DL TDL TDLVDLGDLQDL TDL TDLVL IFLDGFLNFLS FLL FLDFL TFLL FL FFL VFLFL SFLHDFL VDCGVDCGDCKDCHDCQDC TDCGDCSDC LDCWDC SDCVDC EDGLD DGGDGGDGDGDGLEGRDGSIDGVDGF DGGDGT DGVLEIDS LEIGLEIDLEIKLEIDLEIDLEIQE ESETEPETE SAHPLLLAHSAHKAHDAHNAHDD DLAHVLAIY HKLAIHSH D D A D D Y D DS LAIHN V D DTLCAIHPD DL LAIHEH D DS LAIHSD DED D DEID D D D DQPD DRLA D DGT0515253 4 5 6 7 8 9 0 1 23 3 353535353535353636363TVTGTCT GG G C A G G C C T CX TCCTPT C CA GC TT T TATGTATGTGCCT CGC C C C CA TCA TCCTITAB_TBT B TBTTBGTA_TA_ TA_TA_T AB_T ABT GB_TAB_TCB_TCB TC BTAT CCT C_TT CCTG_GTG ATGTG_GE_ GG yT_yGGT G_CGG yTT_T GG yT T_CGG yTT_y CGT G_y TGT G_ GG y AT_y CGT G_ GGGGGGGCyCT_yCT_AT_G A TbAbCAbCAb AAbTAbAAbGAbAAbAAbTAbCAybGAybG C_BC A_T7_GT_T T_TT_GT_T T_GT_GT_GT_AT_T T_C T_G4CAGCAC CACCAT CAC CAGCAC CAACAC CAC CAT CACC _A.1CC_AGTCC_G A ACC_ACACC_AACCC_G A ACC_A A GCC_ATTCC_ACGCC_AACCC_G A ACC_AGTCC_ATT72VMPLCDDLSGGLMDLDFDDLADSGLMDLDFDDLADSGLMDLDFDDLAQQAQDAQQAQIAQSAQDQIQ QA QNQ Q QD QDSGDFDSGGSDSGRI DSGDTDSGVDSGRA DDSGHA SDSGMA A DDGQDGLA ADGLA HDGTEA A SDGY G LDGSGV V VLDSD GLDS SGGLDESPSGV VD V V LDCVP SVKSVSVP SVSVSV VIA SS GLDSA EGLDLSMGLDSVGLDSVGLD IISE GLDSKIGLDSVEGLDTSNGLDESYGLDSSFIMLL EMLML AMLTMLEMLEMLDMLTMLSV ML L A L Y LLL LLLDLALC LDLSLI LDLLSLQ ELDLQLDLM V GDM MFMSMD VDSD DLD MLIDLLDV LS DLC LDLDLLNLDSLP DLS LDEDLLQLDADLLDLDELT DLLDLYDLLDPDLVLDV VDFLDDCQFLCDTDFLCKDNFLCADFLDP DFL ICDFLDDFL FLDFLSDIDEDKDLVFLYFL NFLVFL YLDLFL RPD AD AD D QDC CDCQDCNDCDCSDCEDC LDCKI DCDDCDDGLEIQSDG LEISTDGK LEIIGSEDLEI RDG EQ VDG ED DG EEPDGD EL DG ETDG ESFDG EVCDG EII DG EDDG EDAHKAHQAHQAHEQ D D Q D D D D DS LAIHDLAIHSVLAIHS L IV D DID D V D DDSAHL L ID DASAHND DF LAIHEN D DPLLAIHMLAIHTD D D D DKLLAIHDLAIHM D DED DSPD DRF364 5 6 7 8 9 0 1 2 3 4 5 6363636363636373737373737373TGCATTT A TCT C T1_ T2TATATTTTTGTTTATTGCGGBG_TT T C CGB_TTATBCAC C C P C_P CA _TTCT BT T B TACGCAB CTBCCB CCCTB_T T _ TOTTTOTTTAB_TTA AB_TT CA_ TCC TA_TT TA_ TT BTTC TA_TA_CTG_GGT G_ TTGT G_A AGT G_TGT G_ GTGTS_ GTS_ GTT_AGTT_ GCGTT_ AGTT_G GGTT_AGTT_TGGTT_ GAyCTAyAAyGAyTCAyCAyAyAyCAyAy C y C yAy T y TT b T b T b T b T bCT b T b T bTT bA TbGA TbA GT bGAbAbGC _AG_T C _GA_ C C _CA_ C C _CA_ T C _A_GCC _A_ C _A_A_A__ CA_ T CA_ C CA_GC _AAC _AATC_GTACC _AACCAATCC ATCCC AATCC ACCCC A ACC AB_CC AB_CC ATTCC ATTCC A A GCC_ATCCC_AACCC_ATACC_A A A 73VMPLCDDLSGGLMDLDFDDLADSGLMDLDFDDLADSGLMDLDFDDLAQQAQ QEQNQSQQ QV QK Q QLQ QA QLQRDSGLADSGTT A A LDSGKI DSGLA ADSGEA A A A KDGA ADGDF DGAEDGKA A EDGADGPPAAAHARTDGFDGPDGRGVLDPV VSVSV SAGLDSQSGLDSEEGLDSGF GLDT SV SGGLDE SSKGV LDL SSDGV LD SI SSVGV LDL SVT SV SNI GLDSMGLDP SSAGV LD ASV SEGLDGSSMGV LDPSQMLAMLHMLVMLSMLDML DMLHMLYMLSLPS L A LNLGL EDLTLDLDLLALDGDLLMLD VLPDLLDV LSTDLK LDLSLDL ELDLEEDLLD LSDLLILD KLFDLGM LLDLDLLNLDPM LSRDL PM LD L DLLLS LDSM QDLLDAML EVDLD EEDFLCADGFLLDCGFL DCPDFLGDPFL SRDFLNDDFLEEDFLPS DFLCS DFL MDFLVDLT DLGDLEGFLR FLAFLED DSD HDC SDCQDCDC SDC FDC EDC EDCSDCQDCADCEDGP DGTMDGYDGDGDGSDGADGYDGDGQDGPDGDGDGELEAI EHET LEAIHLEAIN HQLEAIG HGLEIALEIGLE E E E W E T E EREGE EDTD D M D DFCAHPAHFD DLD DLAIHS LAIHKD D DED DT LAIHP LAIHP LAID HVLAIHEV D DLD DLD D K D DR LAIHAV D DS LAIHED D A DD H P L PG D DEE778797081 2 3 4 5 6 7 8 9 03 3 3 383838383838383838393TCCTGTTTTTATGT GTAGCB CACG_ TT TBCT BCTC TTCA CTCCCTCA ATCT GTATAT1_ T2_A CB_TT TC_TT TC_TTB TB CBGB CTCACreC reT TGTC_ TG_ TTG_TTGB_TTA T_ TABTC_C C TT _TT B_TT TTB_TT kTnkiTni GGTT_ TCGTT_ GCGTT_TGTT_AGTT_CCGTT_ CTGTT_ TTGTT_GGTT_CAGTT_ T GTT_G TGl_ Gl_AyGAyAAyAAyGAyGAy CAyGAyGAyGy T y C y T T y T yT b T b AT bAT bGT b T T b GT b GT b A bGAbAAb CAbAAbAbC _AC AC _ T C _A AGCCC_ C C _A AACCC_CAC C _A ACC_G AC _ATA GCC_TAT C _A GCC_ C _GCAA ACC_ CAC _TAC C_TAAC _AC TC_ATTTC_AGTC_TAC _A A GCC_ACCCC_AGTCC_ACGCC_ATACC_AGCCC_AB_CC_AB_74VMPLCDDLSGGLMDLDFDDLADSGLMDLDFDDLADSGLMDLDFDDLAQKVP LVPVP PVPVPDI VPVP SASPS S S SDSGKPQNGS PQNWTPQNLEP EQNFPQNEFPGQNL P FAQNDSQLD ASQTLP ASQLLP ASQELP AISQ FLEASQSLHGVALDA SSQNLVQNHLCIQNLQQNSLGQNL LNMNDA T QLDQL LGQEAQTAQSAQIAQCAQMDGGDQGDDGDVGSGTMLEN QV QSNV P QQDN E QVGNVLNV QVQQLQQSN QV QLN QV QA DGPYA DGPPYQ GLPYNGTPYE GDKPYCGD PYDDLDQLD SVQIVQ P VQIGQ QVQIDSQ VQIGIQ VQA I Q WVQIDFQ VQI SNELTNELSINELL E S EGE LDNLVNLHNLGDL RFLGLPFIEGLDEPFIE LLDNPFIELDLLPFIEGDRL FNPIEPDPLPFD IEDD L FLPIEGLKI CGKDMGVAI C SKPAGVAI CPAGVETKI CGVDKAI CKGVGP I CLGVFVDCDDGP E P ESEDE ENS EAE SFP SF IS PF MS PF ES PF LS PFKLEIQVNSVNLVNLVNHVN VN VND PLFNWFPNWTNWLNWQNWANWDGSNWL EAPEL SPHE L SPEE L EPME L PPPELA DRD D Q DIG DLDI CG DEDISDEG Q DIG DCDIGG D G DIDFRDLY QRGRD N DLY QRD NT LY QRD NT LY QRD NP LY QCRD NL LYAH GEI TELE E EGE E DT L IQVDK P G L P T D DD NFH 19293 4 5 6 7 8 9 0 1 2 3 43 3939393939393930404040404T3C_C C CTrTe AG GT CT CC CC CA GC GGGG GCGC GAkT CAT TATAATGATAATAATT TTATT TTTG TCTATCGniAABCT_ATBC C_ AG AB_ AA GB_AGT B_ AAT B_ ACCB_CG AB_C GBT_TBCCA_ TCTTB_ TCACB_ TCG GB_T l_GC_ CCG A_TTCG G_C CG A_ TCCG A_T CG A_CGCGC_ACC AC_C CC A_ACC GA_C CC G_ GT CC A_ AT CCC_TAyT bGyC _C b CCGyCbTGyCGyGTC bTbCGyGbTGybCGyTbC CyAb GTC yAb GCC yA Ab TC yAb CCCyAb AC yTCAbCCC _ C _ C _GCC_GCC_T CC_ACC_T_G_A_ C _G_ _G AC _AAB_GC_GAAGTGC_GATAAGC_ TAAGTGC_ACAGGC_ CAACCGC_ CAACTGC_ TGAGAGAGAGAGGATAGT TA_TGCTA_T CGTA_TG ATA_T CATA_T TGTA_T TA 75LSSTDDEGPLDYIKVFEDIEFLTSAWPPNSAGLCIQEIDYSQSDWSSVRGYVGGVGKVGPVGAVGPVGYVGVIVGV V VG VQRVERVP RVGF RVL RVRVGRYERGMRGLRGAP RGEL RWASSS RASSISHASSSLTASSSL ASSSSSAV SSSSGASSS SS AV SSSAAV SS WAV SSSSVAV SSSQPAV SSSGAV SS TSEAQ LRAQDAQ E AQLAQ D AQPSAQRAQPSASQTA D AAAAAQASAL SGQEEALVSGQMAL E SGQSRALDSGQEGAL L SGQDPAL S SGQSPALASGQRALLSGQLALKI SQGQFAL T SQQIALSSQNSW QFALQLALQGRGPNQ DRDPDPDSD D DQDD D G G GDKPYLGYL GYTGYSGYL GYGGYTGYE GYDGDA TGDWSGDRGENEQPKILANECAKLCDL PDNELDPNEGPLRNE S PLMNESPLPNELQPLNE F PLKNEKPYL ANE PYLTSNEGPYNPYTNE LNEATGSVPP E I CDKGVL GL I CG LKVS I CVRKRI CD KVP I CLVSKAI CAPKI CY VKI C ESIKI C LRKILCPAKILCA KIKILCD REFKS PFSGGDS PFGGS PFSGGS PVGFT S PFSGSVPFAGSVPFDGVES PGVFVS PFRGVPS PF TSGSVPGVFVS PFSHRLD PLGPLYE EYL PLHPLPQEERLQGL E EYVEYQQRLQRRLQS E LVPYSE LVPLAPYAEYTGRLQPRLQGRNT L L ERL PPYKERLYPYKE L RPFRYQERLVPYHERLDPYAE LHD RYIL D N N D N D G D NETQA D N G D NSFDSD N ALQA D N ALQ D NLP LQG D N ALQY D N ALQ D NIYLQ D NIN 5060708 9 0 1 2 3 4 5 6 74 4 404041414141414141414GT GG TV G TGG TC G TA G TC G TC G TG G TCG TTG TT GT GGT TCTXT T TABAAGTC TC T A T TABCG_ CIGB C_ TABT CBTABCC P C A_TC TAC A_CG_C C _T T B TGBTT BCC A_ C G_CC _ TCTBTA_CC BA_T CBCCA_GCTCCCE C CGC CGC TACGAC TGC CT C TTT CAACC CGC TACGC_yAbC C_yC __GAbC yC __B_Ab AC y_ TAbACC_yC C __AAbT C yC C __AAbC C yC __TAbGC yAC _ C _ C _ AC __AAbAC y_AAbGC y_AAbAC y_AAbT C y_CAbCCC_yG_GAbC_GGTACGA74GAAGACGATGACGATGATGACGATGACGAGGAG TG A_ATA_T.1 TA_T CATA_TA GTA_TTT TA_TG ATA_TGCTA_TTT TA_TCCTA_TG ATA_TACTA_TA GTA_TGT76LSSTDDEGPLDYIKVFEDIEFLTSAWPPNSAGLCIQEIDYSQSDWSSVRASASNS S S SQSSQ DS L SYSNSPS L S SPI SLAASQD LC ASQA LNASQD LAASQLHASQL RASQLRIASQLGASQLWASQPLAKSQLQASQLGPASQLDASQ SSAQRAQLAQIAQVAQNIAQNAQVAQAAQCA ARA QA DALRGDPGIEG V D GLGAGDGGD GNG QY Q DDG GEGQVGQAGQSGVGDS GWGDTGDRGYGMGDTG P GDNGDYSGDLGD MGDNPYLPYF PYPYPPYPYDPYYPYNPYL PYAPYPYGPYPYRNELLNELMNEQ LDNELPNEKLLNELIINELSNELLNELVNELLNEQRNEHNE EE NEGK HKFT L F P ALN L I LQLLLTI CGSVPVI C EKVPI CKVVI CVPKAI CKVL I CVSKEI CVPKQI CKVS CKC CKCKCKCKCGIVNIVHIVEPIVVIVAIVWPPEFEGS PLAP FTGMS PF DEGS PGF AS PF YSGS PFIGVS PGF QS PGF FS PDGF IS PGF AS PFSGS PFAGS PFVGS PFRARLYE E LPLPYL EYVQTERC E L PPYLS E LL PYI E LYTPL TPHEYYE L SVPEL EPKE L CPDE LALPELGPPELVEPELM NLQKD NERNLQDRD N QLQVRD N GLQYRD NF LQNRD N YLQFRD NL LY QSD NTRGLY QKRD NS LY QIRYS RYQRYGRYLD NQ N SD NYL LQ D NEKLQ D NVL LQ D NHT LQ D NTD 8191021 2 3 4 5 6 7 8 9 0 14 4 42424242424242424243434GCTTG TT G TG GA GA GC GC GA GGGCGAGC GC GGTTTGTATTT TTGTTG TTGTTT TT TTT CTT CTTATAB T CCCCABCT _C TT B_ CG GB_CCTB_CA CB_CCBT_GBGCC _CTBC_ABGCG _GBTC T_C ATB_C ACB_ TGC CT _ T T BACA_C_yGCC A_C CACG_GCCT_ TCCC_ GCTCC_GCC AA_T CCT_T CCC_ACCC_ GCCCGA _CACG_T CC A_ GCC G_CT Ab GC yAb AC yCAbT C y CAb CC yAb CT C yA AbC yG AbC yG AbGCyA AbGC yAbGC yAbGC yG AbGCy CAbAC ybTG_ CTAGG_ TTAC_GTGTATG_ T _TATTGTAGG_G T AC_A_CGTAAGAGG_AGG_TATG_CA ACG_AGG_AAG_GATA_TATA_TA A_TATA_TA A_TGC A_TA A_T TGTA_TA GTA_T CGTA_TG GTA_TAT TA_TA GTA_TA ATA_TAC 77LSSTDDEGPLDYIKVFEDIEFLTSAWPPNSAGLCIQEIDYSQSDWSSVRASAS SNSNSKS S SAS S SGSGSNS SGS TSQTLWASQELT ASQLHASQLAASQLYASQLELASQ FLLASQTLNASQLTIASQLPLASQSASQA AASQ AASQ IIIAQVAQEAQAAQGA DA YAESA MA VAIALLL AL LAL TALVG VG KGEG HGQYGQNGQTGQ Q Q HG MGLGQLGQAGQEGQDGDPYEIGD PYVFGD PYDIIGD PYTFGD PYQGD PYLFGD PYPPGD PYNGD PYKGDD PYD GD PYEGD YGGDLYAGD YAINEKILN CYNEKLK E NELFINELDINEHLSANELNNE PNESEGE F EDF P E T P EQP EKGLALSNL SNL SNLN YLANL KN ILSGSVKI CKVL I CAKVNI CDKVVI CKVF I CKVCI CKVSI CQKVAI C SKVFI I C TKVDI CKVLI CAKVAI CVIKI CAEPPEF KLGS PNGRLYP F IS PGLSP FN YSPPFRGS PKGLKP FQS PKP FFGLNS P TGCP FV RSPPFLGMSPPGF LS PL LP FCGMSPPFTGLRS PQGS PESGSVPEEP FILS P F GP F ALYDEYGE LYR EYS E LYEPE L L E LPEVE LLESE L E L E E L EQQRD N VLQLD NNRCLQSD NLRGLK Q D NQ RTLA QYRD N ALY QID NL RCLY QQRD NPS LY QSDRD N ALY QVRD NRP LY QSD ND RELY QSD NL RLLY QGD NLRGLY QQD NSRVLY QED NIK 2333435363738439430441 2 3 4 54 4 4 4 4 44444444444GTG GA GGGG GG GA GA GG GA GG GG GTTGA GCTATTATT CTTGTTG TTT TTC TTC TTT TTC T T T G TGTAC GC CB_ABCC T _ TCGCCCB_CG AB_CTAB_ CTAB_ CTBA_CCBG_CT BA_TB_ T CBCCAGC C_ TCC TTB CABTGA_CC B_C_y C CACTC_yGCGCA C_ CCy TCA C_G yTCTCA C_C CyGCCC_C CyGCAAC_y C CGCCAC_y C CCTC_yGCTCA C_yACCCCG C_CyGCC _C_yACCCTGCC_y CCA_CyGAbGAbAbGAbAAbA Ab CAbAAb AT AbGAbCAbGb bACb CG_AG_G AG_TT_A_TC_ C _G_A_A_A_CA T_GA_AA_ CTAC TATAGTAGTAGTACTGTAGAGAGAGAGAGGACGAA A_TG A_TGTA_TGTA_TATA_TAC A_TG A_T CATA_T CGTA_TA ATA_TG ATA_TCCTA_T TATA_TTT TA_TCT78LSSTDDEGPLDYIKVFEDIEFLTSAWPPNSAGLCIQEIDYSQSDWSSVRASAS L S P S S SSQ ESAS L S SLS S P SDS F SALD ASQQEALDEASQPQAALNASQLQCALNASQQSAL EASQQWALD ASQQFALHASQK QSALKASQSQ ALAQSQQ ALSEAYSQALMAFSQALDASQ NASQN L WAL EAL SGFGDYGDEGSGDTGDGAGWIG DGQQGQDPGQ KGQWGQFGDTG GSGDEG C GDLGDEGDD GDVGDE GDGDPGDVGDLPYDPYLI PYYPYGPYI PYQPYPPYE PYDPYI PYNY Y YQNEKILRCDNELHSNELGINELLLNELDENELAE T EGELETEAP EKP E L P EA GNLVNL PNLPN PLENLFINLKNLQNLVGSVLKI CKVC I CVSPKI CVGKI CVGKI CKVDI CKVAP I CVPKI CKVP I CS KVT I CKVE I CVAKCNKCAIVS IVKPPMGEFE S PVGRLVP FDS P IGLP FKS PQ GLI P FLVSPPQ GFL S P SGNP FLLL SPPGF AP SPPT GFSDSPPLFSGS PNGS PKI GS PS GS PNGN LL P FLYP FTP FEP FNSPPF SELYSEME R E L P E LSEPE LDE L P E L E F E LEE LDE LKE LEQCRLY QLNRLY QALRLY QGRLY QLRLY QERLY QLPRLY QGRLY QPPRLYY QKRLY QLDRLY QSRRLY QTYRYFNIC D N D D N Y D NPTD NLD NSD N Q D NFD N Q D N V D N K D N G D N NLQ DT H GD NTS64748494051525354555657 8 94 4 4 4 4 4 4 4 4 4 4545454G 1TT_G 2 PT _G TG GG G G G GT G G G G_C TGT G3G G TTG A AT C T GC TT G TGAC TCT _rTATT GCOT PCT COTCCGBT_CA GB_ TCTCB_TCCTCB_CCGB_TCTTTB CTBTA_CT B_TCGB TeTCA_ CkCCB_ TCGB_CSC_CTyCS_CC A_ACC GG _CCC A_TTCC A_ACC A_TCCC A_GCCCCCCACT CACC ATACnCi l CTCCACTCCTTbC y_bC yG_bT C y_T bT C y_A bT C yA_bTTC ybC C_yCC_yA AAC_yAC_y T C_y C_yGCA A A AT C_yGA A A A A GA A AbCb bGbAb bAAb TGTA_ GTA_ GTACGA_TGT TACG_A_TACTAGA_T TGATATG_A_TAT TAAG_TAA GG_AAG_AAG_AGG_AG_AAG_ATAT TATATAT T TATACTAT CATAT B TAT T TG ATB_ATB__ A _ _ A _ _ _ _ TA_TA 79LSSTDDEGPLDYIKVFEDIEFLTSAWPPNSAGLCIQEIDYSQSDWSSVRASAS S P SHS R S P S L S S S S S L SSS T SKSQLLNASQLSLASQLDASQLKASQLSLASQLD ASQFLGASQSLAASQLLQ ASQELDASQLAASQLEASQ PIASQ EAQEAQ A D Q AQRASAQPAQDAQGAQIAIA AA VAL IAL PGFGDGDSGDSG DG QGQPGLGDGGDVGDS GDEG GDVGLG HGQEEGDNGDQ GDFGQDGQDGQKGQNGDE GDLGDR GDEPYGPYGPYLPYRPYL PYQPYPYPYP PYL PYL PYDYLYNNELLNELFVNELLT NELRNELHNELGNEV LDNEK VNEQNETSNESNEYPNEAPNE CKL ERPVL L Q LALEL L L A LDI CGSVGIKI CKVKI CVLKDI CKVP I CVGKI CKVDCF KCAKC LKCKCS IVL IVEIVIVWIVKTKI CVEKI CVFKI CVETPPEF GGS PGRL RP F VS PFFGS PQGF ES PMGGS PGF LS PFDGS P IGF MS PGGFS S PGF PS PGFGS PYGFL S PAGFE S PFVLYNE L EPLDPYKQHRS EYDE L EP FYEE LAPEL LPDE LHLP PSE L LLE LASPPE LNPELDPKE LSPPELNPLE LQE D NQS LQDRD N KLQLD NARELQED NE RELYV QGRD N ALY QLD NL RCLY QID NE RELY QLRD NEI LY QQD NSRVLY QSARD N GLY QSRD NLS LYV Q D NYRLLY QSRD NQT LY QFD NTK 0616263464465466 7 8 9 0 1 2 34 4 46464646474747474GTT G TG G TG G_T2_G_T1_G TT G TA G TGG TG G TC G TG G TT GA GCTGTTTAr rACTCAT C T A T ACA GB_CTGB_CCBTeT_Ck TCek TCG AB_TCCBTG G_CABT_CTAB_TCGT B T_CTCBT_CTAB_TCTB_ TCCBA_CATCATC CACni l Cni l C C CCCTC C C T CAC T CTTGCGCC_y CCC_y CCC_y TCC_yCC_yCG C_yC CTC_y TGCG C_y TCA C_yACA C_yGCTC_yGCTC_yGC_yTTCC_yAAb TCAbGAb ATAbAbAbTAb GAb A b T bTb T bT CbACbC_ _T_ _GTACATATA GA_A GGTAGCGTATGTAAG_ _TAGTAGTATG_TACG_AC G_AC G_ACG_AT G_AC GAGG_AG A_TG A_TTC A_TGC A_T BA_T BA_TGTA_TA GTA_T CATA_T TATA_TCCTA_T TGTA_T TATA_TGCTA_TCC 80LSSTDDEGPLDYIKVFEDIEFLTSAWPPNSAGLCIQEIDYSQSDWSSVRASAS S P S R SVS S SKS P SQS E S SNS P SSQG LKLASQLKIASQLAASQLDSASQLDLASQLLKASQLYASQLC ASQLGASQLTLASQLDSASQLQ ASQLKASQVAGQSEA GQTAQVAQVAQDAQEAQVAQRFAQEA AGASA VALDDEGDAGD GFGEGTIGLGPGQALGQLGQLSGQRGQSGDEPYLGPYGEGPYKG APYLGD PYD DGD PYLLGD PYNIGD PYQ GDCPYAGD PYAGD PYMGD PYHGDTPYVGD PYLLNEIILENECL SNEVLRNESLPNELLNELQNELANELTQNEANE RE NEDLNELLNEGNEDKKKLS DARM LEL L L S LPLLGSVIS I CKVKI CKVL I CKVS I CKVDI CKVKI CVIKGI CE KVE I CKVA FI C TKVAI CD KVF I CKVPSI CKVGI CVLCPPEF E GS PVGRL EP FP S PFMGS PD GFL S PFSGGS PGGFS S PGFS S PFAGS PGFL S PFEGLS PFDGS PGFL S PVGF ES PFGLYVE LDPLL PLS PLPLAPLT PLVPAPL LPLD PLAPEP EYWEYT EYIPE LYEYKEYE EYE E LPE E L EGE LVE L IQVNPRDLQR RSRPRMR RKR D RYPRYMRYARYSRYIRYH Y D N GLQ D NP LQ D NE LQ D NL LQ D NR LQ D NL LQ D NLGLQ D NR LQ D NF LQ D NS LQ D NL LQG D N ALQD DI I D M H D D D FD N V 47576 7 8 9 0 1 2 3 4 5 6 74 4747474748484848484848484G GA G G G GTGC G TTTGA ATTGA GTGC GAGCGTTG TTGTT CTTT TTT TTC TTATT CTTC TT CTTATCTTT CCTB_C ABCABCT BCCBCAB_AB_ CB_GB AB TC BTCT T BCCCTGC C_ A _G_ C_C CB_ CACACGCG_CT_ CT _CT BC A_TC_y CACG_T CACC_TCCC A_TTCCC_ACCC_T CC AT_T CC A_ GCCCAC_C CC A_ TCCCA_CGCC A_ GCTCA_CCCCG AbC y_AAbAC y_GAbC C y_GAbC_TC yTAbC C yA _TAb_AC yAAb_ C C yAbGCyA _AAb CC yAbG CC yAbC C yAAb GC_yTAb TT C_yAb CGTTACCGTACGTATGTACGGTATGTAG G_GGAGAGGATG_ACG_ACG_AC G_AGG_AG __ TTA_TAC A_TG AT TATATA ATATAT T TG ATGTAT T TAT T TAT TATATACTA ATA AT_ G _ _ _ T _ _ C _ C _ _A_TG 81LSSTDDEGPLDYIKVFEDIEFLTSAWPPNSAGLCIQEIDYSQSDWSSVRASAS P SQS S S I SQSNS S S S S L SGK KSQ DKALP ATSQALQQEASQH HALAASQD QNALADSQQLALAASQG QKALAASQQFALVASQQNALMALSQALISASQGALE ASSQ TQDAL SPL LTNNQEEPL EQDTNWPL ETNDIGQDGGDTGDVG GDVG GDGSGDG GDFGDLFG DGQEGQ GDL GDDIGDASGQQ DDGQGIDPVFFPNSVFPNTCVF EPNFPYGPYAPYQPYPYKPYE PYT PYI PYS GY GYQYGTYI TYLNEWNEAESIEVEVEKEIEDE I P E TTP E L TLDL LDD LDTKILCAKGILCAN KILIN CYKILHN CS KILCSN SKLCLN KLRN CLWKCGN SKLCDN TKLCLN KLCAT SDL SDEGSDSASVE VAVDVPVR IVGIVVIVIIVAIVQIV QSAGIQSA Q AWPPFPF GS PAGRL EP FAS PGLAP FVS PVGLC P FP S PLGP FVGS PNGLNP FLGS PGVP F AS PL SP FSGGS PLP F S GTSPP SGFHS PMPSPTGSPTQL SSPTPGP F SP P R P PNP P PELYTEYAEYE EYEEYVE EVEAS E LQE L E L PS QSNQS SQS NQRP RLQARLQKRTLQPRSLQFLRLYV QVRLY QQIRLY QSIRLY QQLRLY QMLRLY QRTL LRHTL LRLTTL LRSA D NTD N D D NSD NFD NLD N H D N V D N D D N A D N G D N M DLQSDL LDLG 889 0 1 2 3 4 5 6 7 8 9 0 1484949494949494949494940505GTT GCGA GGGC GC GT GT GC GCTGT TT TGTCTG TT CTTA TT CTTATTA TTC TTG TTATTA TGCGCTCACCB_CGBCTB_CABGB CB_ C B TB CB T TAB CABCTB CGBCC TCC _ CGT_CA_ CGCG_ CG_CT_ C C B C T _ TG_TC_ TT_CC_CCC_T CCC_ACCC_ CCCCG_CTCC GG _CACG_C CGCG_ TCTCG_T CCT__TCCT_CACC A_ TCCC A_TTCC ATC yGC yGC y C y C y C C y C y C y CyTC y C C y T C yTCy T C_yGAbGAbTAb ACAb CAb TAbGAbGAbA AbCAbAAb bCbGbTG_TAG GG_AT ATG_TAG_ CTGTAGG_TAC _CGGTATG_GACG_GACG_CATG_AAG_A ACA C_AGA C_A AGC_TACA_TA A_TTC A_TA A_TGTA_TA A_TGT TA_TTT TA_TTCTA_TCCTA_T CATA_T CGTG_T CGTG_T TATG_TCC 82LPFVDSSMLEDIIDGSISGASSIDTASTQQLAQSRESVHQEPLNQTAMSSKK KQDLK K K KV K DKVLKPK AI DAIAIPPLDSQDD PLDQDPLDQD DPL LSVELAF SV LAWTSLAGLSVILAESLAGSSV LALESV LADFKSPLKSGPLKSPLTNGTNL TNP TND R SSR S R SMR S F RS R S SDASGASMAS EVF LVFVFLVF F C CGC CCI C C C C L CCVC CQRCCDKSK K QPNPNGSPNEPND DLDDDDLDTSDH DG DLMIVMIDLMIGTYMTYVTYQTYDHKLHKEHKHKHKSPHKVHKWVHWV WVVLSDDDLLSD DHLSDG VLSD NCGNCGNCDNCANCNCDNCACYS CYD CYDDLKGIKG KGFKWKGVKSK DSSP SSF SSSQSADSP FQ SD SAPQSADASSQ AD VLGRSVLQLSVLD DSGVLPSPVL PGSGVLLSGLVL SGH VRVH PPVRDH PDVRLPLPTDSPPT VPSP S P SPT L SPT GEKN NGEKN NGSEKLNGAEK NGNS EKENGP EK NGDEKLSNGE SG ESE S LSAE SDQPDT SLQPSG E QPSLDQPSLNEHE L E EAESE LL LMKQNKTNKDSNKNKFNKLEMHKIPHKI HKILCNKDDISDIDSDILC DRATLDS L LDRPTLSF L LDRLLL TCL LDRLD ELT IDE SLEEQT IDELLP ELT IDELGEGT IDEG LL ECT IDE PLTEQT IDELGEET IDELLDFYQ DSFLPTYQ DSLGYQ G DSLGE203 4 5 6 7 8 9 0 1 2 3 4 5505050505050505151515151515TC TGTT TA A AT A A A A A AGC CGT AA CGCC CTCCACCCCCA GCCTGG A GTGA A GA A A GCGA A GTGCGA GAT AC T B_ TA TB_TAB_ TCCB_GABCACG_GTBC C_ GABC T_ GGBC T_ GABCT_GG AB_ GCCB_TTA TB_TTAT B_ TTG AB_C_CyGCCCCC_ CyC CCCG C_CyCCTCCC_A yTG A G_ TyCT G A G_TyTTG A G_CyGG A CG_TyGGCG_ CCy CG G CG_C CyCGCTG_ATyTC CG_ CyCTCCA_CTyGCG_CyCAbAbCb bC TbC TbGTbTbT TbC TbTbCT bCG TbCG TbTC_A_ATACCTAGC_GT ATA C_TT ATA_AGCA_AGTA_AACA_TACA_AGA_GATA_TAT T _G G AT _AG AC T _GG ATG_TCTG_TGTG_TGTG_TGTGC_TGGC_TAGC_TCTGC_TCCGC_TGTGC_TGTGC_TGTA_TGTA_TCTA_TGT83MSTKEYGLYKKFETPSLDSGFVEKVEKSDKPIKVLKYRVTRSVYLYV AAIAIAIDAIH H HLH H H HAI ED AI EAI EKSAPESFSKSAP WSTKSIAPEKSSFAP DSFSPMGSMS SMSMD SMDSMESMDDGDEPLEGDL PLNGDDPLPGD LSGPGD LDPLGDDPLIGDLNTL DTIL PTSLNTLGLNTTLDIEK GKCIKLKTTWTT ETTL TT LTTTT ETT F FKEIFKIFKFWMIL MICVLWVD MIT MIDE WVSWVL I T IFI E I IGS IFIDS PSYGICSYGCSYACATQC TQSTQQTQMTQTQLTQDR TQ GS PRTM DS PLR TTWSYDLV YGI LV DYGG LLV YGGLVDLVVLVTLV VYGLYGHYGSYGLNG I NVNGLGSIND NINAHS GHS QHS PHSSNYENYLNY NYDNYSNYANYA GDSGFGVRS PRVR LVR PVRGDQGDQGDQDDFDPD DDVVLVVDVVWPE SNS PHE SNS S PESNSS PES L I SQD QVQW MNAQNAGNALLNADNAPNAPQSDP NAGSS RRLDSS RRD DLSS RRPHKI HKILHKI HKI GS LGSDIQDITDIADID GFSNGFS RGSGSGSGGSGS LNGFSD GFLGF E GF NGFLF DLLFALF NSDS SL Q DSLPLDSG LLCY DSLLDFFEG DLTFEGH D QFEGLDLCFSEGAS S S SMNLLNLNLAYQEYQLYQ QP S P P PDDPSFEGPDSPFFEGAPD GFEG DDLGEG YCGGEGDYSGGEG YGL617 8 9 0 1 2 3 4 5 6 7 8 9515151525252525252525252525CT CGCC CAT G TTTTTCT G TCTA GT G G A AC CGGTG GA GC CGCAC C C CA A AATATT TTT T TATCTAT T TGT TTGB_TTTCB_ TTGT B_ TT CCB_ CTB CABCGB CAB CAB GB CBC B CAB CGBCA GC_GG_A_T_ T_ C T_CC_AA_ AT_ AT__ TCTC A_TTTC A_TTGC C_A T G A_TT G A_ TCG G G_CG CG A_CG GGC_ CCG G A_TG GC_ACG_CCCA_CGCA_TGyTGy TGyGyAy TAyTAy TAy CAyCAyGAyTGyGyGyGT bCT bGT bTT bCT bGT bCT b T b T bCT bTT bCGbTbCbTT _ T _ T _TT _TT _ T _ T _GT _AT _ T _TT _TT _GGT _AG_TG AGCG AGTG ACG ATG AGTG AGATACAG ACATATAC TACA_TG A_TA A_TCC A_TGTA_TA A_T CG G A_TGTG A_TCTG A_TGTG A_TCCG A_TGTA A_CGTA A_CCTA A_CCC 84NIINGNSDRHELLLNSYDYQHSAFKQKAYAEVMANQEEVTSKDDNIIKAI E EAI EAI ED AI EC S S SGSLS S SA A ANTTILKWNTTTILNENTKFTIL LNTKGS TILDF IKDATDCDF DILATSDLCIFNATDE CIEATDD CIPAT ECDDII ATD CDS IGATDDL EDE L E EDDLNTDLDLNTNEDLNTDIF F S F FAIAI EAFIWAFILAFI EAFI LAFI F SKSKFSKES P CIS PGS PVS P D E S GE S F E STE SEE S F E S E SLGL S LFR TDR TLTHT L E L S E L S E L E LQE L L E LME LDGDSGDGGDLNG GINENINL RNG I NSRPNGA I NDK QGV THK QGTG L K QGCTIK QGTGK QGTTSK QGDLKGDL MEV NHMENLMETNSVGG GGIGVGSVQVVGVPVGPISISPPISI LPIS DIEPV ISID PISIAPTISIDQ FPTISIAN DNSSN P NS LN GINSADSS RLD DV NSRR SRGDVL RESRLVTHVPVTHGIGVTHG QVTHSLVTHWPNFSSRNSR PVTHDVTHSGG GS EEVPGG SEGGW EGRS E PLLFHLFPSRM SLF DDMGDM AAEAARDMLDMLDM DMDDML PGP P E PAANAADAANAALA KVEKVNKVNGLNEGT LN YLPGEGQ YSEGLEGFN YPTGLEGLYDFMH ELPLSFMHN ELLH QMH ELSLLTMH ELLLLCMH ELSLA GMH ELAA LDSMHM ELLDA LPN ENPSSA FPNH ENA SQSPNSENSA G 0313233343536537538539 0 1 2 35 5 5 5 5 53545454545GAGGT GGGA CGCT CGCT CC CC CA ATTAG ACATA TCA TG A TTA TA A TA A TA A TT C GAT AC CCCCGCCACTAC BTA _CGB_TCA TB_TCCCB_ CABCABCTB CGB CGB CAB C CBABAB GBCA GT_GG_GC_GA_T_T_ C_AT_AG_AT__TATCA_ TCACC_ CCACC_AAC_ CCA A_ TCA A_TT A G_CG CA A_TG A A_CG GAC_AG AC_ CG CA A_ TG CA A_TGy TGyTGyCGyTGyGyTGy TGyGyGGyGyTGyGyTGyGGbGGbCGbCGbCGb CCGbCGbGGbTGbTbCbCb CCbCbTT _ T _ T _ T _TT _ T _ T _ T _GT _T GT _AGT _T CT _CT _C_TA AGTA AGC G A A A ATAA_G AAG AG ATACACATAGAGTACG AC TAC TG ATG_ACA G G_ATA A G_A ATG_A ACG_A ATG_A ATA_A GTA_GCA A_C A A_C_ G _ G G G C C G GG A_GCC 85EINNSSEPIILGFSELGPLILDDFSTDCMLSDELIQQISEEMMPDPRYDQMREALA AEA VSVG VLV V V VG S GCG GDLNTDEPDLTDS EDLTE EDL DTLKSK D GNLKEKDKEKDNGNEGNPGNDI GNSK G DK NDD GNL RY I NLNRY I NDRY I NLD LLENSKLGLNSKLWTNSKF LI KS E LI KSWLI KSLLI KS E LI KSGL LI KS LGLI KSDF TNVTVDTV DENDLNDPGDG MGCGLDFEQD DID D MAIFS FMAI T FCMAE FIQMAIFL FAT MIMFMAI S FVMAID DEFVKES SVKSGS ELVKS EM GMEDMEDME L SDGSDI SDGSDS SDDSDHSDLFGFVFQNNVNNLNNENN NNL NND NNVNN NNLNN NN KI EL KIKIGNSDNSDNSGNSAS S L S S E S SDS SAS SDS SSS SAL EHEVGGSGGFG GD T S GT SGT SST SWT SFT SPT S DANGANSANSPEELLSPEEDGSPE QG ELSPESEGTDEPIGT EP QT EP LT EP PP T EPDT EVPP T EPSKGLNIIKGLNIPVKLNI DSKV VDNDLKLKVN S KVL I RDI LDI LDIMKFNKF NKF DKF NDIDS KFLD KIFGD E KIFLWE RWE PWELMNNNNNGNNLAPANAANLAN ALENSLCPENSDSPENSTLPENSDLSEDFHA QSLEDSFLA TSLEDLFLA CSLEDFAA GSLA EDFDA SSLEDPFSA FSLEDFDLVKEIH A QVKEI EAPSVKEID ALL4454647 8 9 0 1 2 3 4 5 6 75 5 54545455555555555555555ACT A CC A CGA CA CT CGCT CC CC CGCA GT GGGTCA AAGC C TABAAT B_ATB CTACAG A BAABATT AA AA AA A BAGCA BAGBAABTGG A _AAB_ACB CA G BCCG AB_CGBG_GCGC_GC_CG_CC_ CA_ CT_ CTCC T C C_ TG_ TTTA_A G G_CyCA A_ GA A_TT AC_A TCATA_C CA A_TTCG A_CCCA A_TGCA A_ GCAC_ CCC C_ACC A_ TCCCC_ CCCC G_CCC bTGy CGy TGy_GC b_C b_GC bC T yb TCT ybTGT ybT T ybTTT ybC T yb C AyTyC TbCGTyTbCG Tb CyCG TbTT T TAT T _TTG_GG_G_GTG_CG_AG_GG_TTC _GC_ _GAA_AA_ CAA_GTAA_ TA_ TAGTATATAC TATA A AGCATA GGTA GCTA G A A GGTCTGCGCT _GTACT _GGTCT _GCCCT _GCTCT _GGTCT _GGTCC_GCGCC_GGTCC_GGT86KRGDELIKKISDEISNSLLLEDFYLTRESSLLYVIDQINDLFEELKGALAIGVRYGGGAGMEF G WILFFLTLTWLT ENFE RY I NEERY I NDRY ITND R GLQTRGLS TRGL S TGT TGTGVRLC RLMRL L TRGLDFSLFFMFF SF TCFF SFQTVDTVTVS T L PAGPAGPAPAI PAPATPADSRSRI SRGN IEDN WN GNVDVDVVDLVDHVDDVDDVKE LVDSVDLYEDEDEV FED KTCED KL ED KFDTSTFDKS LVS IVSMVS STSTLFGITST SFPTSTEFGTSTFD TSTA FTSTFANIDLY DNIDEY GNIDDSIF T KIF D KIF D KIF D RRMLLRRMGR VRMPRRMQLRFRMDRW RMPPR DRMS TVF FDTVFQTVFLAEKNS EAANE E L E LGANDANARQLNIWKLNI Q KLNIFKNRE D RQRRENRQ EG ERQ ENRQ EDLRQ ENSRQGKEL SVD KP YL SVLKP YNSVLPYDDLI DSPF L FSE L PSEHR F P R FSRFAR F R FM QPSE S P ELP E P EAP ET EHAT ESHL T E LHLWEPNPWE LWEDWEGGF CGFS GF FPSGFTSLGFDS SGFGSGFDLNEGDSNEGTLNEGCN NNNNLNLDG DED D D DLD GVKN KSNKNKTAETATAT TAPTAGTACTADQYGQYPQYEE IASAVE IALTVE IA A DVE IM A DAFMIHAFQ M DAFMQSAFMLGAFMGLAFMIQAFMFDCFDEGLCFDE LGCFDE IH 859506162 3 4 5 6 7 8 9 0 15 5 5 565656565656565657575GC GGACGGGC G T A TT TT T T T A C TC T T AAGG GAGCGG GATGCT BTGA GT_ CT C B_ CTAT B_ CCBTGTAAT CAT TATTAATAATTCTAT TTTGCA_TCAC TC_ABABAABGA_ GG_TGT_AC B_AAT B_CAGT B_AC B_A AAT B_ABABCAC_AA_GyGC_TTCC A_ GCCC_A G TC_CCCA_C C C_ CCGC A_TTGC A_ GGC A_TGCC_ATTA_ GTTA_TTTTG_CCT b_ TT GyTbT_GGyTbC_GyTbGy TGy_ CTGyCGyTbTGy CGyGGyTy C y T y T_G bCbCbGb bTTbCAbAbGAbCC AC _ C CAGCCCC_GGT CAAACC_ C CAGCTCC_ TACAGGTTC_ TAC_A AGTTC_GA_AC CAGTC_GAC_A AGTTC_GAC_AACAC_ACAC_TATCC_AACCC_CAGC _GATATATC_ACTTC_ACCTC_AGTTC_ACTTC_ATATC_AGT87VMTVRKKSRMHALFNEWVLQNSNNKTYNKVTQPKLNKTVSTIDFIK DLT TIT L T S FDS F S F L S F P S F S F S FMPM MFS F LFSELFS GLFSFGIGEGGGLGGGWGFDNLDNWDNGFF SFFF FF SFFDMAEFMAFMASM AEM ALMTM DEEETELSRGSRL SRVSR DASAS SASVASASMAASAADGSGSGSMYELYETYEYE LGTLGTGGTGTQG GCIGS LGQQGQCIGQNIDLGNIDSANIDHSNIDAGF STPSGF SLPLGF SH PS GG FSVGTFSDLGTFSDEGTFSARPSG RVRPSRDERPS DRLTVKF IGTVFWTVF PVTVFDSGSVT SA GWGT SGGIGT S PPVGT SD PSGT SDFGPTSGGPTSDSDS ETDSDS ETGDS E DTFPTYRKVP KVPKVGENSYP SYGSYLSTGPPSTGGRSGPTGGSG TGLLSGD TGD SG GQLSG GGILKQ LIQID LKQLKQ DNHHPEGQT ENPHS T EE PHP T EHMNFCN FS NSNCNF E F FL TFNTFML F L FNL F LHNC PNCDLNCANCS NCQGDQGSQGAQYE EGANYGEGSYFNEGDLLVVSALVVSQLVV SSF LVVS LLVV SDLVL LVDNL LNL LNLSVS T VSLKD LKD TKD DSCFDEQQ DCFDLQPECICFDE TQY QCFD DEF SDGTFSG LLSCGTFS SLE SQGTFS PLT SQGTFSCLGSEGTFSLGSGGTFSLLPS TLGFS DLFASD GECTGEASGELTPLASGETG G 2737475767778579570581 2 3 4 55 5 5 5 5 58585858585GT C TC TT T TT TC T TT TGT G G GGAGTGCGTA GA GG GGGA GA GGGA GA GGC G GA GCGG G GTG ATGTGTAA AAGBT_GTA_ TAT BA _ A AT B_ACB GBBABCT CAC_GT_GG_GT_GAB_GAT B_ TB CBGBGTBGABCGC_GC_CA_ CC_ CT_C AyTTA_TTT C_CTT C_AA ATA ATA G_G_CGC_ CCA G G_CCA A A ATA G_ GG_ TGC_ATT G_CCTT A_TTTT A_ GC b_ CAyGC b_ TT AyCC b_ CAyTbCAyGbT yTyT AbCAb CyCAbTAybCAybT yTGAbC TCybT TCybTGTCybCCT AC _GAC CAGTC_ C CAACCTC_GCC_TAAGTTC_ TTT _A AGTCT _ CTT _A GCCCT _GT_GC TAGCT _GTT _GATTT _AACTT _AGTT _TAT C _GAT C_AGC _AACGGTCT _GGTCT _GCTCT _GTACT _GGTG G_CGTG G_CTG A G_CCT88REPLKQGYETILNPDNVNDLTIRWVASVQIVEVSSGTESMCGRNKYLLIMDIMDM M Q Q QLQEQ Q QRYIDRYIKRYDNEDNLDNEDND KADKAEKA KA KALKAD KAD S G SVISKGE FGE GGE FGEF S RDS RDS RDP S E SNS S S LDILDIFDISGSRLGS SGS S SDSLS IES L SRWSR ESRGSRD WIWI EWI LQPS TQV QGGQDLGGLGLGE LGT LGFLGL LGF SDDSDDSDNDR SRPSRHRPSRLLRPSRLDSENSDSFENLDSENQDSENCIDS SENGDSENMDSENDLSGPDL LSG DIELSG DEFSIEAD TWS ETSPDS ETGIDS EA TDNFPV YHNFPTYSNFPG YVNFPYDENFPYLLNPYDLNPDYLNNEQNNFLNNSGKLQFPIPKQ VIGI SNL F PKQGL F RKQNL FGLYSL KKP LYA EVKKEWLY P KKDES LYGLYGFILYD FLKKF LYAFN DAGFN IL TFN EQKKEGKKDKKS I SVAIILS SAIL LAI S LQGNL SQGEQGHQGMGLPGPGLGLGRGEDGEGKVDSKV KVGIKDANL PNLNLDKD SKDQKD DI TGDASP EILTDP NEG L ASEFPASESEASEL S ILTD DILTN SDILTNDILTLDILTLDKDKWDKGLLGLPGLGR DYFII PYFIIAYFPII LLYFPII LTYFPIIHYFPIIAYFPIIM DGIS LDGIS PGISN GT CGT TGTQ GT FG NSFG N G G NCG NLG NQSG NDSG NLG NLG NNSG N H 687888980591592593 4 5 6 7 8 95 5 5 595959595959595GC GGGT GA TGTC TT T TT TC TAGTGCGTGA GGCGG GTG G G ACAAAA GGG G G ATAG ATTTATA GTTTG AC GT T BG _CA TB_GA CGB_G CCCB_ CA TB_ CGT BG _CAB_CTCB_CA GB_ACAT B_ ACCCB_ C BTAGA_CGBGT_CBGG_T_T TTC_ CCTT A_ TCTTC_ACCC_ CCCC A_TCC G_CCCCA_TTCC A_ TCCC A_CGCCC_ATG_CCTA_T TA_ TCCyGbT TC _T Cyb CT yCC b TCTCyTbCCyCb CCCyGCbTCyCbTCybTGCyb TCCybCCyTbCG CybTG CyGbTG Cyb TCG AC C _G G AC _G AGCC _T_ _T_GC _ C_G ATG AGG ACG ATG AGTG AGC _AC _TCG ACG AT C _GGA_ T C_TAC C _AGG_CCC G_CGTG_C G G_CGTGT_TGTGT_TCCGT_TGTGT_TAGT_TGGT_TCTGT_TGTA AGTG A_ACCG A_ACG 89ELTLDNIRVKDDVREPNCESYGISPIKIRALYGEDTKQSNPLEFPSIV YDEGMD NDDDGDSDSLSWS S L SFS F PAC PLRYI PG LERYIGG LERYI LGLAERYIMLGGAGS LGEDQLGTDCLGEF DLGDMLGSDLG GL LGLG DEI C LEI C EQDSII CDSGDSDSD IIEIIDSIID DSGVAG AG HDSGGDIAGLSGD DGTAG SDGDAG LDGLAGDLDGLHDG DQ HD DGLWSDWDEWDWDLGGLGDL SGSGGL SSGDF SSPGGSV DGGS E SGGGSASGGSDSFGGSGSIGGSAD DVLSTD E VLV SDM D DN GL WL L GFNS SDT SDMSDDSG GVPGSGSGL GSG GQLGSGW GP GSGD GGG DSGGRGSGSGGMPG GHP LLMPGS LGL DNVNFNCNINFN NDLNFN NDLG GGGGGLGGNGGPGG GEGGDGGSGGNGGLGG AGGNG GG GLLMCGFI ELLCGDFALHAL D AL ALAGS PGS LGS LGS SGSGSHGSDLVGDL LDII SSII S EII SDFII SDGSGLGTGA GDSGQ GDDPLDL DKVPKVGKVDKVSGDKVDKQDKDKGS FGS CGS LGSGGSGSS GLDGSDDGC LGGPGG GPGLGG GE SGD LS LP GLS LGAGL PGL LGLDL L LGTGGEGGLGGCGGGGGQGGFDL LD E DGISGINIGIGGQGGIGGGGGIGGLGGDGGDGPGDGP I SG NEPG GSNSLG GSAG N D GSMG N D GSGSG Q GSHG G D GSG GLPGSGQEG GSDG G G GSGECG GSGDLG GLEGTPLG GLEH DGL00102 3 4 5 6 7 8 9 0 1 26 60606060606060606161616G G GCGG G G G G G Gb bTGTGT TAC GTGC C T A_C_ATCCAB_TTCTCBTA CABT T_C CBT CCGAB_ TACGGB T TCGTB TACGGB TACGAB_TGCGABTT AGCB_GA7 G_C7CGT C G _TGTCGC_AT CAA_AC_T AT_T CG_TC_ 4.T 4. TT C_C TA_ TTA_ GT C_AC_CG_CATATA A A ACACA1G1 CGyGy TGyGyT GyGyC G_y TG_yGG_yGG_yG_yTVCVAC b CC b C bCC bC Tb CTbT TbTbT TbC Tb TCTbCXABXABC _ CGGA_ C _G GAGC _AACC _T CAT T _ CCAGT _GCAT T _GCAGT _T CAC T _ACAC T _CAGT _TATCIT _CCCI T_ACA AGTA_ATG A A_ACTG A_AGTAC_TGTAC_TGTAC_T TAAC_TCCAC_TCTAC_T CGAC_TGTPE_yCCPE_yGC 90Q GVDSLLDLLCGEIHDVMPLCDDLSGGGLDGLDPLEQGVDSLLDLLCGP R C P C P C P S C PGSVS S SGSQSE LI CVEG EL QIC SG ELCGLICDEI CQ GLE ELG G ICM HTASMASG MADEG MALGG GDSLMAVMALGDMALDHHDD HDDDPDMLGQPGQAGQGGQGGQDGQDGQAHDD LC PDD LW S DD LLH G DD LG HDDDAGVAGWAGQAGIAGSAGFAGDY DLLLYPLYP LYLLYGLYL LYDLSVSGPKVS SVS SVSG VSDFGL SGS P SNS R S L SDSYGMGI MGV MGVPGRGPME PMNPMS PMNPMD PMLPMLPMLLGRYPLGRVPLGHLMVPGG VRMGPGDDDP SPP SPMSSMS S SPAMS LTSPSPMSHMS L SPASPM MSMSCGLMGRCGLASSS CGLSP SL GAHL GVC LYPC L LLAEGMFSPGMSGGMS LPGMQ SSGMLSCGMDSSGMDSLDDDK DDAQDDVSDDSDD QANTANLANLANEA GA GADL GKQDS LKEAL GSS LALD QL GPGED LGRDRDDSL GSGGSMQSGSMCIGSMGGSMQ DGN SMEIGN SMGLGNFSMDGDH GS L ETS LGS LGV MGDDGD PSGDARGDLD GGPLSTGGPLGPY NELGPGLSG EPEEGPGLPE TRRGPMS SNQGGAGSNQ E GAIGDSNLPGGAPGEDLGP PGPY GP P SNEGGACV GP P SNH DGGAVSNDGGAGQ GPM GP L SNDGALA GEGTGEF Y L L YA Y Y Y Y YD GYPD 31415161718 9 0 1 2 3 46 6 6 6 616162626262626b_bA_G_ATbC_AG b_T__ CTTAGb_C GATCCAT ATATGC AC CA 7C _ T_y CC_yG_yTT_y CT_yCT_yG _yT47C _7G_7T _7 G b.1GT 4.1CC4. T1T 4.1A G4.T _ Cb_TTb_Gb_ Cb G bCAbCT 1 A AG _GA_ CC AGTAG_CAT _ _G ACC ATVG VA VA VTVT1 1_ _ _ _ _ GXABXGXGBXTXC_T_ C1C_A1_ GT1_ TT 1_ T1C_TCIG_ CT BCA_ CCB C C BPATE_yCITPA_E_y T IGPGC IA_ I C _ re Grk CerkAekGrTeGreArere AE_yCTPE_yTTPE_yATnilATB_nilGTB_nilTkA CBnilAkGk k TGB_nilAB_nilATB_nilCCB91L QRSQKLEKKKKKFDDWKPKKAASEDSRGDPQFKRKKNFKNAPQFK NKMTSKMLLKMGKMDLK V MDKMHKML SYWSS R LYWFSYWWT SYWFSSYWRGSSYWE SYWDPQAPQPQVPQD PQEPQS PAH MHR LHRCHRH HRQHRDKKWSKKGI SKKDS SKKF SQ DKKGSKKP SVKKDSGPTRD P T PIPGPVPGPLLGT R SGRD GRL GRHGRVGRADAPDAGDALDADAQDAPDAG P PATPE TPL TP STPTPGR P R R R L RDR L R R LDLD DLDLGDLGDL PDLDSDLDSQVNGVNGV GVLGVNGVGGEAF EAWEAEAI EAVEAEAVSQVHQVDLQVAQVSQVEQVMSDSPSQSGS P SLSGQIAQIQIQIQI LQI PQVIMKT MG TPMTLMRMTGMTLMLD DQ DLDDSDTDSDDL S EDS ENS ENSTE S EES EDSTES LKSKCK K K GL FK DSLDSSDS SDSNDSPDSLDSMNYNYENYGNYNYP NYP NYD SSASSASS LSSHS S L S DGNTCI SGNTQSGNTEI SGNTGL SGNTL SGNT STQ GNF STDGVDS SGVGSGVTSGVQ SS SSGVF S SVC S SVLSKQ ISDGKSH GKDSD GKSGGKSGSQSGDL SWG GSWLGI SWLGPL SWE SWP GSWGGSWDGEE LK KCG QGTEFDD C D D GS PGSG GSA GDEGLGDEQ GDEG GDED GDE SG GDE IG H GDED GSD GS PGSV GS LGD P D D QD 526272829203132333435 6 7 86 6 6 6 6 6 6 6 6 636363636T_yG_y CG_AC A_CC_ yC _yG _yT _y CC_A AC A A_AT_ AT A CTTTyT _yG _yG_yT _y CC_ CyCG_C C C_A Tb_TTb TCbTbCbTb CCbCbCb TTbTb TCb CyCbT ybCA_ C_CAG_G_CA_T _A G A_ C_G C A_ G_TAG_T_ GA_T _A A_ C_CA_C_G_CAG_G_G_ GT _TT ACA AG AT3 3 3 3 3 3 3G2 2 2_2_ G _ _ G_rCC_rGT_rTT_rTC_rA_r TG_r T_rTC_rCC_rA_rGT2_r T 2_rTT 2_r TekAenilGkGeAeAe GTB_nilAkGB_nilGkAB_nilAkTB_nTeilTkC BnCe AilAkTB_nTeilCkAeAe GCBnilAkTB_nilGkTB_nTeGe GilTkC BnilAk CeAe AGB_nilAkTB_nilGk TAB_nilCCB92P ATEVVLTPYIKLMDPMFDPNAFIEKITELDKDPDAGPLGLMIHYVIKFGSESDS S S S SDS L L E LWL L L LFLILDTDTDDT LDTDT EDT PDTGW GWTGWGGWLGWSGWEGWFKDEKDLKDNKDGKDDKDLKDD AEQAE C AESAEMA A F ADLIYR P WTLIYP GLIYPEF LIYPLLIYIPELIYPELIYPFDRVLG VRVL IDRVLVRVLDREVLG LREVL LTREVLDLT TCIRT T SVRT T SGRT TM DRT TFLRT TQ GRT TD PLAS SD PEPSAS SGAS SHS APSS LDAPLPSPSSGAS SAAS SAS YILD S YLHS YLL S YL LS YLT S YLVS Y L L LPLFLIL L DE E ESE L EDE S EDE LAVPLLVPQLVPVVPDVPGVPWVP SSGDRGIKQSGDR PKVIPSGDRGKIIGSGDR FKDISGDRAI GSI GDLMLMLMP LMDLMR LMPLMGKWSDRKLS R SLDKGGLWDLGLWN SGLWG EGLWLGNGPGLALWHLWNLWKLKGKVWNVWEVWRKDLKPPK NVWVWVWDKLGT L GTLGTPGTVWMR E C RETR ESR EDSGTREQGTSGTM SREAR EDTL STL PTLTLATL NS TLL TLGLGGLLGLFGLGLEL G LLHE LHE SPHEHHEDGL SGL SHEAHELG GLG GL C EDLELPLPLGLGL LGLDGALTLGALF AQ A A AH G GAL LQ D GAL IQ H GAL LQ G GAL TQ Q GALGLQ GALQQ D GALCIQ GAL FD 9304142434445 6 7 8 9 0 1 26 6 6 6 6 64646464646565656ATG ATG ATTATCATCATTAA CT CGCGCC CT CC CAATACAG AA AA AATATA CA A GCTA CCA CA A CG AA A AC C TGTBC C_ GABCT_GABCG_GABC T_ GGBGB CC T_ GCA_ GCC B_ T B TTB TABTAB TBTGB T CBGA_C_ T_ T_G_ T_ C__TT_ CC_ TCA_CGA_TGA_CCA_ TG_CG CG_TG T G_ CG CG_CG GG_ TG CG_TGA A ACA A A A GCA G ACA A A GGC_AAy TAyC AyT Ay CAyAy TAyAy TAy TAyCAyAyTAyAyTCbGCbC CbC CbCbT CbCbCC b CbGC b C bCCbCCbTC bCG_G_G_G_AG_TG_GG_TC _GC_ C _ CC _AC _ C_TC _TA AGTG A A A AGCA ACA ACA ATA ATG ATAGAG ACAGACATA_TA A_TGTA_TG A_TCTA_TCC A_TGTA_TGTA_CGTG A_CTG A A_CGTG A_CCTG A_CCG G A_CCCG A_CGT93TQADVFPTHNALAEPIDDAMEELAVVEGHTLPTRANTRIRYTGPEFGDVEI YI E EI YIEEI YI L EI YI L EI YIDEI YD IDEI YI L RVL L RVLLRVLDRK VL R RVVLKVL L RVL LRADIR E RDP RNR S R L RD S NGS NCS NGLS NF S NS S NAS NMEIQELEEFIAQEWELT IAQELLEEIAQE ELF ESIAQEGLLEIAQELGSEIAQE FLDEPLEIVI E EPLD IVIDELPLIVIDEPPLIVE EIDI PLIVLINEEPLIVDIS EPLIVDILC F L C FCI C FQC FGC FMC FVFD R KWRK R KL R K R K RKGR KDDKPETSDKEDDKEGDKELDKEDLDKEHCSSDKL IEAYL T ISCIYLGISSYLSE IQYL ESFIR MAPSME LYLFSS IYL L ISMYL FSD GPSMV DPSMLGPISMD PPPFSME EWR VSMDSSELSADLVSLSLE EAHEAGEAT SSE LGSALLL EAD SLE LDALAD PEQ RE SREP EDL ED LED GRER EDDRE EDPREGED GLG DEG GGDESG GPDEV DG DEAG DEGG IDEDG FDEAT CA DR E NT CENAST CELDATCAENT CD AEL T CEEAPT CEMPQFDL PFDVP PG FDSLPG FDW PPG FDGPG G FDDPFDSF EGLSARF ELLTRF ELLLRF ELHRF ELARF ELSRF ELDL IKKFNSIKKFGIKKFL IKKFP IKKFRNIKKF DL IKKFLHEG HE LHE CHEQSHEDSHEFPHED HR LHREPHRDLHRNSHRH HRA HRM 35455656657658 9 0 1 2 3 4 5 66 6565666666666666666GCCGCG GCTGTGCGACGC CG GCA T T T AG AG AT T AC T AT T AC T AAAA ATA A AA ACATT T CA A G ATAGAT B_ATCB_AG AB_AA GB_AAT B_AA TB_ACCB_C TB TABTGB TGB TABTAB T CAA_TA A A_TA TAG_CA A A_ TAC C_ CT_CA_ CT_ CG_CT_ C C B_CAA_CA GAC_ CA CAC_AA A_TTCAC_ CCCA G_C CA A_T CA A_ TCCA A_C CGAC_AC yGC y T C yCC yTC y C y C yT Gy TGyGyC GyGGyT GyGyTAbTAbGAbTAbCbCb C bC TbGTb CTbT TbT TbC TbC TbCG_TTA_ CCG_ _GTA_GTGATAA_ TG_AT TAGA G_AAA CG_ACA GG_TATG_AGG_ CAGG_GATG_TACG_AGG_AACG_TATA_ACGTA_AT TA_AT TA_A ATGT TA A A A GT TGT TGTCCA GT CA A A ACA AG C G G _ _ G _ G _ _G G_TCTG_TGT94NESTMELPGPPLLPTEDAVLTPPTPAAPLSVGSPDVKLLRMSRRCGIGNERDR E R P RDIRGR RIVI S ILIGI I IKWLWP TGKTFSKWTLEKWT EWF KT LKWTWTKWTDF PHKFPK MEHML PHK MDPK PHMEPKDEHMS PHKDPKDMD HMLKHKLS PHPHQPH FVK KLGK FL KLK FGKLL PFTKHMPSKLFDKH LKL CI PKHD FD KLFDCLKFD EKICE EKFNE KE CEFKFEKLCEKFEKWCTKFGE KL CKFEKLGCSKFDE KFDRGHRGL RGVRGARGRGE RAFFF S F EQF ECF EMF EVF EDK SSKPKSKGIKSKDSKSKWK DSKFKSKG QKG SKD IKLIKGIKGIKIIKD IK K SST TST LSTVST DSTL STHISTLAKFVPKFGKFLFPFDF L FAE SAE LAED AE EAED AESPAEP R P RRRLKRPKR DKRKR GRKARKGRKS RKGRKF RKRKDP EGQNEP ENNPEHPND PEL PNNSPP ENLPENPE LAPNSPNMNESSWNESS IGNESSLLNESSQ NSSDNSSVPNSS SGRNPQNKD SRDQQN SRD L QCRN DAQRN DD QSRNLQNL PD TRDVF PVLF RL LL EL DNVFDVFNVF L EVLFGEE VLF LMEI E FK TI EQEK QI EK QE I EGK QC I EGKLKDLL YNL Y L YLL YSL YAL YPL YQP E E GE L EQ GEI EQPLEI EQDF PKIYSAPKIH Y QPKIYLCPKIYLT PKIYDS PKIYSF PKIYDL7686960 1 2 3 4 5 6 7 8 9 06 6 67676767676767676767686TCTTCTGCG TCT CT TCT CC TCC CT TATT AC AT AT A AC AGA CC CCGCC A A A GA A GG A GA AGA GTGA A A GCGTTCTA TBTG _TA TGBTA _TG TAB TA _TTGTAT T TT B_ TTAT B_ TTTCB_ T TTCCBCGB_TT_CABT ATTG_ CGBTAT TA_CTBGCTC_CABATTT_CABACT T_ C CCC T C B_CAACCA ATA GCA ATA ACA ATACA_ _C T_ CT_ TT_ GT_C T_A_Ty CCb CA_y CT A_yCTA_yG_ G _TAy CAyTTA_yTAyGCGbTA T Gyb TCA GybTA GybTAyGGbCA Gyb CAyTCGbCT_ CT bTA_GC _ CTCb_GT bGTTA_GCTTA_ T C _ TTCb_ATCb_GTCb_TGTTA_CCTTA_CCTTA_GTTTA_ TA_GCACA_G_TCCCAGG_T CA_GGCATA_G_TGT CAGG_T TA_AACACA_AGA_TATG_TCT CG_TGT CG_TGT95ADSGLMDLDFDDLADSGLMDLDFDDLADSGGLDGLDPLEQGVDSLLDLL SLVLKLKLML LQLHLQL F L L LNLGL PCDL PLLCDLAL DCLVL DCLGL DCL ELCDLV GL DCLAL DCLLLLCDL SI L DCLNL D ICL EL DCLAFLCDLL L DLK ASAA ARASAEPGSKIL CKG G G G GLGA A ASA AG AIA A AE D A DAD S DSDADGGD DGDPGDAGD CGDSGDEGD GIGDAI SEIAEI S LEIF EIEIEI SEIS EIEI FEIEIKEIEIAHSHSAH HS IL HSVHSV HSLHS LHSAIHSHSEHSISGSSDGLV ADGLADGLMLGVMGVMAVMTDL LGVMVDLV GEVMEDL EGLGAGSGN GEGTHGRHGEVMVDLIVMPEDLVMGDLVMHDL CTVMPDLVDLIVEMVVMLDLV NDLM D HVMSMPDLSTMPDA LAMPDRLSIMPDV LR MPDG LHMPDLGMPDLHSMPDSLF MPDLPTMPDLLCMPDPDMDD KMDQMDRL S LDL P LPLTLAI LHLLS L NL D LL P PLLDPLLSE PLLGCDF S CD DL CDDCD LCDE CD DS CDS CDI CDHCDPDCD DDDGDFDEE CDFDSI FDF PTFHFVC FHF FVFYFDCFQF PDPDD DD DD DDADD DD DD DD DLDLPLDLDLG VDLDLGLDLDLMRDLDL RADLDLG ADLDLVTDLDLD GDLDM LASDLD PLID DD GDLDLN QDLDLA GDLDD LEDD CDLD QLFK 18283848586 7 8 9 0 1 2 3 46 6 6 6 6868686869696969696G G G G G G G G G G G GT_CC CCCGA CCGA GA TA G G3_AACCACAC CATB CAABCAT CAACAC CAT B CTCGC C CGC reCGB_CG CBA _C AB_ CTA_ G_CBC B C B GTGCCTACT_ CC _CT _CT_AT BATBCC A_ C C_A CG AB_AABAkTCG_C niC_A yCCTCC_T CyGTCC_TyCCTTC_yGCTTA C_yGCCTA C_C CyT TA C_ACyATA C_ GyT CT A C_yACGTCC_CyGCTTGC_y C CATA C_yACTTA C_ TyC CTTlC_yAbCT AbTAbCAbGAbAAbTAbTAb GAbCAb CAb bAbCbA_ CA_AA_GTA_A_A_G_TT_ T _A_ C _A CA_GA_GA_TAGTAT TATAAATAAATAAATAAAC AAAAC AAAAAAAA A_C A A_CTC A_CAC A_C A A_C A A_CCTA_CAT TA_CTATA_CCGTA_CTGTA_CCATA_CCATA_CCGTA_CB96ADSGLMDLDFDDLADSGLMDLDFDDLADSGGLDGLDPLEQGVDSLLDLL GL L S LDLYL S LSL E L L L LNLCLSL RCDL TAL DCLDLLCDLMLCDL LGLCDLDIL DCLG TLCDL SLCDLGLCDLD L DLC LQLCDL LAL DL IDL DL FL DLKGAEAGADGADGALAIFAPAQ APALL ANAS CAE CAMCALIDAEIDS FEIDPEIDFG EIDSG EIDSAG EID AG LEIDPG TEIDCG EIDS SG EIDGG FEIDGG EIDESPG EID LLHSQ D H HSVT HSVHSEIHTSHSHS SHSG N H HS SHSQT SYDGLSVMLDGLDL DGLV GGVMAVMSDLKGV ADLV GMVVMTDLVDGLMP GHVMHVMSDLDVMPDGLEIDGNGVGLHGMHGSGVMHLVMKTDLVMSDLN T VMSDLVMLDLLKVMIMDLMD D MD P MDFMD MDYMDDMD GMDDMDYMD MDLDEDYP L P L S P L S P LHP LNP LAP LAP L F P LVP LNP LGPT M MQLCDG FLTLCDG FGLCD FFLLQLYL T LQL L LKL PSLL L PLLNPLLFI DCDCDMCDL CDDD YCD MCDCDCD PCDMCDIDD DDLDD DFDYPDFDGDFDQDFDQCDFDFDDFDPL DFDVDFG D DFDLDFDNDFDQD D NLD DDD D LLL L L L L L L LG V QL LDL L E L L C TG DL L P L LFGIL LNL LGL LAD D KD D KD D D D D D D DD D D Q D DCDLDLLPDLDLWTDLDLNT5969798999001020304 5 6 7 86 6 6 6 6 7 7 7 70707070707GCTGCCGCAGCA GCCGCTGCGGCG GCTGAGAGG GTGAAGAA AA ACG AGBACACBATBACCACCAT BCAT CAGCAGCCB_CT C ABC T_CCABCA_CB C _ TB_ C _T_ T BCG_ C TGC A C GACTGCA _CBTC T_CTC_TBTBABGCC_ CT_ C C_TCC_ AATAC T C Cy CTGC_yGT _ C T _ T T _GCTC_G ACTC_C CT A_ CCTC_GCTC_A GCTT_T CT A_TTCT A_CACTC_ GTAb_GAbC C y C y C yAC y_AAb_ TAbC_GAb_GAbT C y_CAb_ATC y C CyAb_AAb C_T C y C yGC y T Cy C yAbT_AAb_GAbGAb AT Ab CATTATATACA AT ATATATTATACCATACATAACATAGAAGAAAAAGA_AGA_AC A_AG A_CATA_CCTA_CTTA_C A A_C A A_CAC A_C G A_CAT TA_CA GTA_CTT TA_CA GTA_CTATA_CTATA_CGC 97ADSGLMDLDFDDLADSGLMDLDFDDLADSGGLDGLDPLEQGVDSLLDS LGS L S L S LNS LLD HS L S LKI S L S LVS LVS L S LVDF PLLLDPLFQLDFDELLDALFFL LDF CLELDF SLALD NFLL D F L DQL D L D L DWL D TLDEDDVDDGDDF DDSDDEDDNL F R L FDL FMP L FDL F C LDF IL DQLLDL LDALDE LDYLD PLDRDLD DDDLDVDD DDLD L DDAIDDDDNIC LD KLDPL LDAC LGGC LGLCLD LCLKLLGLL TLLNLLLKLLL LL LLL SLLLLAV AH ATAL CAIES CAVCACAAE CAPPCADDCAACAVL CAMIIDSDG SEIDQG EIDSGG AEIDSGG EID PG IEID AG PEIDAG KEID SG IEIDS PG EIDD SL GLG EIDEG EIDG NEID GSHLHSV AHGH NHSKHSAHS IHSVHGLGD DSHSDHSEIHSDIHSTDGL LDGLLAPGLGGVDLIRDGL PGLVGLYLL GLGPGLA GL EGLEMMLM MV MAMDD DD D MA M MLD GLD P MGEMED DVDVPVMAV V V V VKV V VEVMKVMKPDLL MPDQ LVMPDRLGMPDLVMPDLLYMPD LPMD DIMDFLMDPMDLQMDIKMDKSMD HLLCCDGLLCD GLCDD RLH CDVLCD GLLQPLLYPLL P PLLQPLLDGPLLDPLLSP LYDCDP CDED SCD ACD GVCDKCDLL DTDFDEIDFDRDFDRDFDTDFDEDFDVFPDDSCDFDFYDFDEDFDLDDFDNDFDSCPDFDDDLDLH DDLDLQ VDLDLA DDLDTLESDLDL T FSKDLDLAPDLDLEPDLDLKTDLDL PGDLDD LPSLLLDLDLD GDLDLKEDLDLM V 90011 2 3 4 5 6 7 8 9 0 17 71717171717171717172727GCT AGCCGCG G G G G G GVG GC ACACC CGCACTCCCCCC XCACGCGTGCAABAAB_AGB A C C A_ACB_ACBAC BAC B CG_AGBATBACIAGBAABCA_ CCC CA_AABTGC CT GT CAC CGCG_CG_CCGGCAACATCCATCC G_CA_AC TTT C CC C P CC_ C GTCA_C C E CAC C CACATTC_yC _bT C ybGTC_yT _b GC yAT _bGC yGT _bT C yC T _bC C ybCTC_yT _bGCyT _ T _GT _ T _b AC ybC yb CC ybAC yb AAA_GATA_TA_GA_ TA AGAAGA_CA ATA_TA ACA_AA_GA AAAAGA_AA_ CA ACAAAA_BA_A 7_ CA_GA_ C4AAAAAGAAG CT TA_CA GTA_CC TAC T TAC T TAC T TA ACGTAC C TAC T TAC 1 TAC T TACCGTA AG G _ G _ G _ A _ _ C _ A _ . _ C _A_C A 98ADSGLMDLDFDDLADSGLMDLDFDDLADSGGLDGLDPLEQGVDSLLDLLL DL LLDE LLD PLLD HS LLDESLLDLSLLD EI LLDC LLDKLLD TLLDC LLDLHLLD QLLD EIC LQC LVC L P C L C L C L P C LTC LGCLVC L T C LDC L P L T LNGAEIDRSKGAD EIDSFLGAA EID SGAATEIDFGALGADAEAH AA ASAEAGCAQCAY SKEIDK VEIDSG EID STG EID GG EID EIG EIDLRG EIDTG VEIDG MEID EEG EIDKHGH DHSVH QHS PHSDHSNHS PHSMHS RHSHSGHSAHSKDGL SDGLHDGLRDGLKDGLDDGLLSDGLYDGLLPDGLLDGLP GLQEGLAGLV GL LV A MKVMLLA WIFPL RDFD D D VEYMD MDSIVM MDQVM PMDYVM VM MDYMDPPVM IMDYVM VM MDCMD LVMQVM EMDGMDTVM VM VM MDGMD D MDDP LMP L E P L P LAP LGP L P LKP LDP LIP L P LKP LAP L LQL D RIL D E L DSP L DEL D F L EFLVLLLLLAL L LAL GPLLVC FTC F S C E CVC S CDCD KCDACDHCD SCDS CDCDS CDLDDSDDADFD DFDMDFDSIDFDEEDFI DIDFDFDFDKDFDLF DFDEEDFG D DFDYFVD D MEW PL L E L LA Y DINL LKL LGL L T SDSAKL L S L LQL LML L L L L S LDDSL L RDLDLS D D S D D D D D D D D D D E DD D D D D AD DGDLDLTSDLDLSS22324252627 8 9 0 1 2 3 4 57 7 7 7 7272727373737373737GCTGCAGCAGGGAGAGGGCGG GT GCG_1_ GT GG CCGCGCTC CACT C C C CC CAACCCACABA A A ACBA A ATAreCCB_CBTC_TC G_CTA_ T BAB T B T_ ABACACA_ CC_ CG_ C ATTCCCC _CG AB_C TCB_A CCBA_AkAABAGBGC n CA_CC _C_T CT A_C CT A_CGCT G_T CT A_TTCTT_ACT A_ AT CT G_ TT CT AA_C CTC_ ACiTl_CT AGCCT GCC ybAC y_bGC y_bGC y_bAC yA _T bAC y_bC C y_bAC y_b_ACyb AC yAbA AC yAbC C yGAbC_yGC_yAA GAGA AA A ATA GACAT _AbAAbGATAGATACAATAGCATACA GT ACATACA_GATAAAGAAC A_ATA_AGAAA_AGA_AAA_CATA_C G A_C A A_CAC A_CTTA_C G A_CACTA_CTGTA_CCATA_CG ATA_CCCTA_CB TA_CTT TA_CCG 99ADSGLMDLDFDDLADSGLMDLDFDDLADSGGLDGLDPLEQGVDSLLDLLALQLPL E L F L T LNL L LDLLL T LQL LCDLQLCDLSIIL DTC LDL DCLDFLCDLVL DCLQ L DCL LL D DCL LL D YCL SLCDL FSL DCL LL DLD L DLD L DLSEGAKE IGAYGAGGAYGAEGALAD AA APATATCAR CAF CAIDISE EIDDEIDSL EID L EIDK VEID APG EID EG TEID DG EIDPG QEID DG EIDLG DEID DLG EIDVG EID KTHSHSVHSHS THS EHSAHSHS SHSQHS CHS FHSMSDE SGDGLG GVMEDLC GE DLGGHDLREGSDLK GS DLAGM T DLS GG GTGMGDGE DL LDLYDLLDLDDL EHGHGD VDLVCDLKMPD QVM SMD KVM TVVM VM VM MDRMDSMDDMDLVM VMMVM MD L MD D MDFVM MDSVM DMDLVM VM VM MD S MDDD SLLLCDV FSPTLLSCDL PFD LLECDF T PALLLL PCDFYLLKPCDI PFKLLA APCDFGLL TTPCDD DF TLL L PCDDEDFDDFLLNLDCDDFFPDRLL ECDDFLPILLA E CDD DF S PDP LL CICDDF C PQMQLLDCDDFAPLLSDQCDDF RQDD DDDTDD DDVIDD DP LQT SLD NLFNDLDLQ VDLDG LRTDLDLD QDLD VLLKDLDELETDLDLDLDLDLDLDLDK LQSDLDQ LISDLDLMSDLD DLSD VDLDLK QDLD ALPH 63738393041 2 3 4 5 6 7 8 97 7 7 7 7474747474747474747GCAGCAGCG GCGGGGCGGGAGCGGGGG1_ GGGGAG AAATATCATCAGCAGCATCG ACAC BCAACAP CA ACTCAB_ TB_GB_CB_ TBAB_TB_ CB GB_ T _ CB_O GBATBCA C CGC ATC C CG_C CCCAAC C_CCACAGC T C T CG_C C_TT_ GCCTC_ACTC_GCTC_G GCT A_ TCCTT_GCT G_C CTC_A TCT A_T CT A_ACTC_ AT CTS_CT G_G CCTT_AC y C yAC yAC y C y C yGC yAC y CyGC yCC yAC y Cy T C yGAbAAbCAbTAbGAbGAbAb TAbCAbAAbCAbTAbAb b TA_AA_GA_AA_CTA_T A_AT _ C _TT_ _A_A_ _GT A_ TTAC TAT TAC TATAT TAATAGATAATAAAAAAAAAAAATA_CTTA_C A A_C A A_CCC A_CTC A_CGC A_C A A_CGTA_CTGTA_CG ATA_CGCTA_CB_ TA_CAT TA_CTG 100ADSGLMDLDFDDLADSGLMDLDFDDLADSGGLDGLDPLEQGVDSLLDLL D ES LLDSGLVL DF LLDL LLDDLLDKLLDG LLLDLQLLDL LV LD LLLDT LLDR LLD NRLLDFC LVC L C LKC LQC LYC LVS C L L C L CLAC L L C LTC LEL LKGADGARRGAEGAAGALASAG AA AAAHALAT CAGT CAYEIDAEIDSR EIDS LNEIDV SKEIDEG YEIDRG EID QG EID AEG EIDFG AEIDPG EID QG SEID AG EEIDG WEIDVHSEHSHIH NHS LHSVHS LHSKHS EHSVHSHHSLHS R SDDGL EDGLGGMIMQDL SGS MGDLSDGLSL VMPDGLNGVDLVPDGLDDGLNDGL EADGLGDGL LDGLAH DGLEPVMV VEE VMVVMFVMGVMEE VMLSVMET VMMVMMVMMVMKMDDMDGMD NMDFMD YMDLMD P MDEMDQMD MDLD D D L DAPLLPCDC PFQLLCDGPFDPLLCDCFSPELLCD TS PFALLCDL PFDLLLDS PLL T PLLNPLLTRPLL ENPLLGMSPLL FMAPLL TMDPLLA D DC FKCDFWFCDFDS CDFQCDF LVCDF T CDFN ECDFQCDFDDDVDD DD DDNDD DDGDDTDD D D D D GSDD DD DDWP D DD LT L L P L LD GRKL L T L L L L LEDDCDDMDDRDD RL LML L T L L S L LNDDYL L E L LAL L L LKD D S D D D D ND D F DD D D D D D D D I DDRDLDLCP051525354555657585950 1 2 37 7 7 7 7 7 7 7 7 767676767GCC GCCGCAGCTGCTGCCGCGGCG GCAGCTGTGT GGGGACAABAABAG ACA G AG AA ATBAAATCAACAC CG ACCATC T B_CC _CA_ CGB CT BCBC B_A_TB_ CB_ TB_BT B T BC T CTACTC CT_ CA_T A_GC GCCT C AGC C T CG_CA_ C A_TGC_ GyTCTCC_yG ATTC_yG GTCC_ TyTTTC_yGCTTG C_CyTCCTG C_G yCCTTTC_yCCCTTC_G yTCTTTC_yGCCTTC_y C CATA C_ TyC CT G C_CyTCTTCC_TyAAbCAb AAbGAbG T AbAb TAbAb GAbAAbAbAbGAbAbAA_GA_ CA_AA_ TA_G C_ C _AC_ C _G_G_AC_C_GT_ATAC TAATATAGTATATAGATAAAAAAAAAAAACAAAATA_C A A_C G A_CGTA_C A A_C A A_C A A_CACTA_CG ATA_CGCTA_CTCTA_CACTA_CCT TA_CACTA_CTT101ADSGLMDLDFDDLADSGLMDLDFDDLADSGGLDGLDPLEQGVDSLLDLL DII LLD AL LV LD RLLDPQLLDL LLDF LLDS LLDALG LDLLDSVLLDEALLDTILLDII LLDIC LFI C L C L R C L C LTC LDI C LQC LACLWC L C L T C L LDLDGAEAGANCGARGAQASAD ARIAEAA AH S AD AR CAGCATIDSNEIDSHEIDPEIDLG SGEIDAG EIDSVG EIDEG EIDAG FEID EPG EIDPG EID RG EIDWG VEIDSIG EIDASH NHAHSQHSHSWHRHSPSHS LHSFHSVHS SHSHS SHS TDGLY GVMRDL CGDDL EEGE DLASDGLPPGNDLKSDGLA GL DLAGEGPGHGAPDLTSGGGQ RDLGDLHIDLVDLADLQMD SV L MIVM VM MD YMD E MD QV S M VMKVM MDSMDQMD SVM MDPVM VMEVM VM VMSVM MD P MDPD LID QID SIDLP L P L L P L E P LVP LAP L P L E P LDP L T P LS MP LNMPM MAL DGL DVL E L P L GL TSLKL RPL ALFLGLLVPLLDT PLLQC FNC F CDF E CDF I CDF L CDCDNCD DP CD PCDCDE CD D SDDYDDKSDDEDDYDDCFGSFKFEP CFPF T FGFVFAC F RD NLDL ILQEQIDDSDDIVDLDNDLDLEEDLDLIRDLDLQEDLDLWDLDL EDD DD DD DD DD DD D QDLDLA ADLDFLASDLD QLSQDLDG LSEDLD SLSGDLD SLTD QDLD ELSV 465666768 9 0 1 2 3 4 5 6 77 7 7 767677777777777777777GCG GCC G_2_ GGGCGGGA GAGTGG GGGTGTGCACAC CAreCACCAACG ACAC CACCAGCAC CATCCCGCAC TC CB_CG TB_Ck TnBGBGBi CA_ CT_ CA_C ATB_ CBTA ACG_ CB_ CCTTC_ GCCTl_CT A_C CT A_T CT A_G TCT G_ACT A_ C CC_ABCCT C T_CCBA_AC BATBACBGCG_ CG_CT_TC_C CTC_ CCCT G_ GCT G_CGCT G_ TT CT G_TC y C yGC y C yAC yGC y T C y C yACyGC yCC y C C y Cy C yTAbGAb TAbAb TAbTAbAAbG AbAbGAbCAbGAbGbAbCA_TTATA_TATGA_ _ C _TT AATACTATACA_TAAA_CT ACA_CT ATA_A AGA_AGA_AGA_AGCA_GA_CACAATA_CGTA_C G A_CBA_C A A_CCC A_CATA_CATA_CG GTA_CG ATA_CGT TA_CGT TA_CTT TA_CTCTA_CCC 102ADSGLMDLDFDDLADSGLMDLDFDDLADSGGLDGLDPLEQGVDSLLDLLLQLLGKLGKGKGKSGKGKGGK GEG G GL DIL DLL DTKLDKL EKLDKLKLDKLKDLYLYNLYGLYDC LHGASC L C LPVPV DIVDVLVSVEE VL LYKWYKEYKLYKIEEIDS CGAAVEIDTGAPSMEIDSTDQPSELPQEDQSE EPQFDQLSE LPQGDQ NSSE EDQ GPQF SE LDQPQSEPQWDQ DTSE FWNTPQDVACWNFIVASWNM N GVADWFLVALH HPH AF LQF L F L F LSF LMF L C F LPDIDPDILPDI PDITDGLDGMDLS GP DLAIPKTGIKVKTTIKSKTVIKHKTGIKLKTDIKLKT I IKDKTDH KLRVEH RVLHDHSGRVFRVAVMMDLVM NSVM MDRMD LSETAI D ETAA IETASIPETAILEGTAI DFETAIEEA TAI DRNG DQRN DIGRND DDRN DWPLLDPLLMPLLVKC SELLKCW EPPKCVEPKCEIGKCEDKCG EQKCESN GIE LLLN NIE LLRNIL LNILPNELAELPCDDFNDEPCDF EQCDFG S Y PPI CDY PI CNY SPI CGY EPI C RY NPI CDL Y PLY ICNPI C LDIR SD VL IRD VHIRDDIRNSL L SDDTDD KYLKYKYPKYKYAKYSKYMF T FQFVS FVAD DDDLDLPPDLDLD VIL EGLCIL EA G GIL EGSFIL EH G QIL EGDSIL EGLTIL EGDLLLA YLPLLA YSELLAG Y GLLA YGL879708182783784785 6 7 8 9 0 17 7 7 787878787879797G2C_ GCTGCA CT CC C GA GA GGCT CC C GG GAGCA G GGGT G G C GC G G GAPAG ATGCGTGTGTG GA GACO AB CB TT BTTGBTTAB TTAB TTABT TBTCB TBAB ABGBCT C T _CC T_A_T_ T_G_ T_T C_TC_AC_AG_AT_AT_TS_CTT_ ACTT_ TG G G_CG C A G_TG GC_ CCG A G_ TCG A G_CG G A G_TT G GC_AGTA_TTGTA_ TCGTA_CGGTA_TC y C y T C y CCy TCyGCyC CyTy C y T yTAy TAyTAy CAyGAbAbAAb CCbCbTTCbC CbCCCb CCbGCCbCbGbCb bTA_TAA_TACCA_ TTAT_GT GA_ TG_A_ CG_A_GG_A_G_A_CGA_ CGAGG_TATA A_AGA A_AGA A_AA ACA_TACA_CB_A_C G A_C ATC GGTTC GCCTC GGTTC G GTC GCTTC_GTATC_GGTTT_CTATT_CCGTT_CCTTT_CCC 103KITEALGDWDIYKGNETEKVFKELNISGLNCSEWPLLPEEGCPNTANIKGLYDLGLYPGN LLYDVHLLGN VHLLN VHDN LIVHL LNWN VHL TVHENHFHN N GGHGDLHNPHNDYKGYKEYKF S E EFMCLFVLDIFDL IFDGIFG DLEIFGIWNSWNQWNDKI W DIHVKIIWQKI WKI WKI WIKI WSKI W HGIHL IHDL IHDIHGIHDL TM DGDTDGS TDGQTDEDGFVAVVAGVALHDHHDVHDTSHDHDEHDL HDAF L FVFGF LPDHPDVPDCS SCSCSCS DCSGC L CTIVTIVHTIVVTIVTHIRVSPHIVDRNSHIA VD TL I PGVTL IDGSTL IA GTL I FGDTL IGQ T SLIGGIT SLIDGSMYDSFMYSSPMYS DMYSSA DVRPRN DL RNSL RD GHSEPSGHSESLLHSEW SPPHSESDLHELSNHESGRHESGL GLDG LDLVG LP L SLLGLLWNILGNILNIL LNENDN NASNSSNNSNMS LMSGMS LMSPEL EEL DL ELMNPNLN NSN NLNHNMFKYAFKYEFYDFYPDIR PD VS IRVLDR RGCIVD DESF RGDELRGCDEARGDDE S RGTRGRGDPKLKNGDE LDEQDE LVNDSVNS VNLCVNSFLAF FAFAL LQPLQGLQGLQLQPLQSLQDEDEDF EDEDALYPTLLYGELLYDFLLAS TQLLASEILLASLCLLASGLLLAS LGLLAS EQLLAS FDRL PG Y GRL PYPTRL PYGERL PYGL293 4 5 6 7 8 9 0 1 2 3 4 5797979797979797080808080808GGGGT G GA GACTG CCTTCCCCCATATATG CTTCGTATCCTCG TCTTC ACGCT G G G G GTG GT TATCT TAAAGT B_GG AABG _ACCB_ TABT CAT_ TG AB_TGT B_TAT B_TTCBA _TGB_ T CCB_ CAB CABCGB CGBGT_GT_GA_GT__ CCGTG_CCGT C_A TT C_ CCATG_CACTA_TATA_CAGTA_TATTA_ TCAT C_AAA_CGAC_ CCAG_CCAA_TAyCAy TAyGyCGy TGyGGyGy TGyTGyTC y C yCCy C yGAbCb bCAbCAbAbTT AbCAbGAbCAbCAbCAbCAbTAbTT A_ATAGA_GATA A_TATG_G_GG_G_AG_G_T_CGTTT_CGTTT_CGTGAG_G GAGGTG_ TGAGGTG_ CGAGCC G_ CGAGCTG_G AGG_TATG_AACG_AGG_GATG_ACGTG A G_GCG G G_GGTA G_CCTA G_CGTA G_CGTA G_CCC 104RLSESTILDRIINKRVIAKDFKVSSRVNVFLLSKGDKDGLFEKLKGYKG VHNI EN NL E L LDL S LDLQSQSE QSQS LFGHTDWIFGNHDE IFGDDFNLQLNLWNQENQDL LLNQGNQP LD L NQIL L LENQDTSD GTDTSGT E TSLTNTSTDPDGTTFCTGFS TGD DDLK GTDKF DKDKLDKDKDKFCLGS LGGS LGMLGE LGFLDISPS LISPSWG TISPS EG FISPSLIVIDFDTIVGDFLTIVLANS S IDNS SG LNS SVNS SD NS SQ GNS S LTSNGSSDL HP GTSHPTCIHPSTGHPETQMYGSEYLY GMS GMS DDDPSEGDDPSLGDDPH SSDDP LDVD D DS DDPDDPADPALTV SHDLTSDEDLTSLLDLTSG VMLG LQ MLLIG GML SLGM ASQM AS IGM AS PVM AS FMSDAS SMSMS DD L ASWPAS SML SD RPMLRGD MLD RGIMLRDF S L F S R F S LD NLD NRD NPD ND D NLDNPD NGT VT Q T GTSKYNKYNKYML ENL ENL EGL E L L E L E L E L LDP LDLDRD LVENSLVR DENHVENDLHA ESL HA EHHA EEA P HEAHADELANLHESAMEAHE DGPGESEGPNL ESSGPSNLGEHPLSDLL P TDYLPRL PQ YSER DLPYDD FSPLPTRLD PSPLPQ RSD ESPLPSRFD PSPLP DRSD GSPLPRCD GSPLPRGD LSPLPRLG DSLQPASG FSLQLATG LSLQ AQG SSLQ ALC60708090011 2 3 4 5 6 7 8 98 8 8 8 8181818181818181818TCG TCTTGCAC G CT C T C C C CT CT GC G C CAC CAG GTTCTBTCAB TCTTCB ATTTTBATATBAC TTAB_AATTABATGTBAATTGBATTCA BACA ABTAG AA _ATBAAB AGBGC_GG_GC_TC_ TG_T T TT_TA_T_ C_ CTC C_CG_ CA_AAC_TyT AA_ TCAC_A A TT_TTTA_ TCT C_ CCTA_CGTG_CTCTA_TTGT C_A T A CC_ CCA C A_TT A C A_ TCA G_CC AbT C yGAb TC yCAyAbCAbTAyGAb TCAyAb CAyCAbCAyAbTAyAbTAyT AbC T yb Cy T yC TbGTb T CCT ybTG_G_GG_TT A_A_GA_GA_AA_GT A_ _T A_A_A_GA_GA AGTA ACA AA_GTA_ CA_A_ CA ACAAT TAGTAGTATATG_C A G_C G G_CGTATA AATA GATAGTATACTAT _AGTAT _ACCAT _AGTCT_TGTCT_T TACT_T CGCT_TGT105PPLITGGGLSLGLQSSLESQDITTLQSHGMLGSTMDTMPVDITVPPNASYQSQSQS L L L LDL L L LDLCECECECETSEGITDS D S DNLNNL ENDNLDNSNLIT T S T T LQ DEQ D WQ DLQ D IPNL L TDE TDL TDETDDEQ D GQD L D D SDS NSESSSPG SEISPSGG LISPSDF T CVFS T CVT T CVGC C L CS TVFTVTVEQT C FVDEMHPIE EMHP EFEMHPWEMT HPGLHP FDT LHPTMHPTDA MGGA ELMGCIA ELT DMGA EVMGLA MA QA ETMGEDMGEGMGEDL NQIFLLNQI SLGNQICINQIMDSTSDLTLADSDLLDDLT DDS LSLASNLKDGSISNEGSH KDSN DSPSSSNASLDWSN DDFSSNV DD SSSNAG DD VETG LKSIAVELGLLKI LVEDGLGLKI EVEDLGLKIDMTRWMR FMRD D G D QKD VKD PKDDKD LKDSAAWAAIAAAFLD PPTLDDT DSQ GPLNRQPL LQ NNPLNPQ GPLNPQPLNDQ LPLNLQPLGINLLINPIPLINGIQIA RLINLLINDGEPEDSNSGPS L LGEPL TSVFNTNHVF S TNL VF E TNP VFNS TVFATVFDLTVFMTRLYNTRLYNTRN LYS TRDLYLGSLQAG LQAG M QPYQPYT PYS PNAPNDPNLPNDR S RHR L RA A GSADSSLADLTLDSETLDLPTLDFPTLY DGLTLY DSGTLY DCGTLY DLDWLVIA GWLVIQSWLVITLWLVIDS021222324 5 6 7 8 9 0 1 2 38 8 8 828282828282838383838CC CC CAA AA AA TT T T AG AGT AGT AC TC TT TA T GC T GT T TC G GGG G A AA AA A A A GTGCG G GG GTA AA ATT AAACGT B_ ACAT B_ ACCCB_C BTG_C TBABTC_ CTT_C GT B_ C AT B_ CAB_ CCCB_AGTT B_ ATGB_ABT C_ AABT T_ACAT_TyGA C A_CGA CC_A TCA_ TCCA_TTCC_ CCTC A_T TGCA_CGTC G_CCTCC_A TGA_TGA_ TCGA_TTGA_CGAbTT _TT ybC T ybC TGyb TCTGybTGTGyb CT y T yCGbTTGbC TGybT TGybC C yGC yGbTTTCyGbCGbT C yGGbCCACAT _AAAC T _TAT T _T_TCCCT_TCTCT_TGTG AGG_T C T _G AGG G_T T T _G G AT _G AC T _AG AC T _GG AT T _TG ATG_A G_TGTG_TCC G_TCTG_TGTG_TGT CAA_ CG_AGG_AGG_AAC ACCCA_ACGCA_ATACA_ACT106SGLTCNNASLQLTPVLRIVMYSPPQQTYFNLFRKQSNIQGLSQLFNINMCT EDCE LCEE D E E E E EDESESESS ESTSDDTEMLSD EMDTPSDLDLELE PLLED LEILE NLE E LEEESE D E ELSFF SFQSFSFCLEWL EGL E L L EDPDL PDGPDGPDIHPGHPLEMFHPDFEEIE FEE EEI FF EI FS FEEITLFEEILLFEEL EI GF E FSDPGTSSPGVSPGLLSPGDENQI SVNQI EQNQIDAPQAP LAPGAPCIAPMAP SAIPPPAPPD PPLAGPPKGATALADAKDA VAKDFDFDSLFDI FDGGLHGLGGLKIVKKISKKIKI ILKIHILL GWL G L GGL GQVES VEVVEAC CAC LKC EKCDKCSKCAE LPE L L E L R E L LLKP LKLKD DDDSDDDDGIDGDFDPD DKGPKGDKGNKGNAIIAVAID AI S LLLWLGLDQLDDLDVLDSNGNNGLNGHNGSLP IAS IAINGLINL LINGNLNL PNLNLL VE LVE PVE RVE LNLNLNVEDPNLL VEGVEGL S E SAS ELCS EQS E LTTRLYEPTR LYD TRYMHNDHNNHNNHNSHNHNVL LVL SVLHVLLVLAVL EP HNLMMSA EGMA LSEGMAS MA ESEE SELPWRLVS LIFPWR L LLVILCWRDLVI LHKLD MDF CHKA K G MDFGH L MDQ FSHK E MDTK FLH P MD DFSHK G MDSVKD RLFFH P MDF LDCYCN NIRQC LYIHRN N DCLQ YRN NDE C LYLN NGL4353637383930841842843 4 5 6 78 8 8 8 8 84848484848T TGGGT TAGTGCGTGG GCGG GAACATATAGACAA G ATTA GTGATG GTGT TGATGC TGTTTATATGT TAAB AGBACBGBGBAB TBABABCBTGBTTGBTTAB TTTBTT_ TA_TC_AA_AT_AG_AC_AT_AT_AC_AT_A_G_ C_GCC_ CyCGG_CCGC_AA G A_CA C A A_TA ATA A A_CA_TT A ACA A_ GAC_ CCA AC_ATA_TAT G_CACTA_ TCAT A_TT GbG_CC yCGbT C yTGG_GGbCAybTAyGbT yTyTG_TTTT _GT_TA Tb_ CAbT yGT_GA TbC_AAyTb_CyTGyGCA Tb_ CT AbTGyC_TAbTGyC_GAb TC_ CGyGAbTC_GCA_ CA_ CA_AT TAC TAC TAGT TAC TAGTATCA_ CATA AGA AGTA AGTA AGTAT_CGTAT_CCCAT_C GAT_C AAT_CCTAT_CGTAT_CGTA ACCCA_AGT CA_ACGCA_ATA 107MFAAMATDDNDCLADFDLQAGDSLLPEHSHVAVTSQNQRTAELEGMES SESS ESP FM DI L IGI ISI I IDS EDSDDSDSDSDSDP FDVSFKQ HPDDKPG D DK P KQ KQDKQ KQEKQDKQDSDSDDSDDSD DEKD DKDLK DKSKLE D WE DLE D I E DPLPP LSFLDDPGFPPS SFDP PGLVPPAV ELLVEELWV EL LVNEL EVD IEL EVDELGLVDELDF SSSFT SSSF GSSSS EFFSSSFEEGLD DLEGL P FDDIVE IV SKKQKKT IVGC KKS IV KKFS IVFKIVIV KLKKMKKD GLEGLGEDRSRCI RKD SR RKVSR LTRSRQ G KELG VEKEL IDEKEV LHEKELG LEKELTS EKELDL EEL L RSR E RH SRSRKSSRARKSRVKNG GLKGEKGL IANGPNGMDLIRDDLRE IDL SIR PDLRL IDLAIRDLRD KIFDLA AG APA ADRDRLNQRNVRNWRNSSMEDSAR ES S ESF S EDLRLISEL RLIG EQ RLIEVPRLIEGIGRLIW EPPRLIEDRLIE SGTRLLRNTRPLRGTR PPLTR LLC LGMSA EP MYGL T SA EILILIGILDFKEDKENKEEKE RKI I D ILENENSKE LKE S SENKLSKEEPS RKNSES RKDL N NLRDCYQN NSRQCYDNN N DEANE LNL NCEANE SNL NTEANE PNA NA NAANAM N NSF ENEH N QENEA N GENENDS ENENDLK NKTNN FLPK NKSNNANNLFFPK NKFGLK NKF CG 849 0 1 2 3 4 5 6 7 8 9 0 1848585858585858585858586868ACAG AA AT AGAGAT AC AC AATTG TTG TCTTTTATTC TTTG CA G CTG CCG CG G CA GA GTATAC TAATAATAB TABT CBCGBTBABAB GBCABCCBATBAABAGBAGBAT_ T_C_A_C C_CT_ CG_ C T_CT_CC_AC_ AT_AT_ AA_TA G_CAyGT CCG_ CyCACT CG_AAG yTTA_CAyCTA A_TyT ATT CA_ CyCACTA A_ TyCATTA A_TAyGTA A_CyGAT CA_A yTTA G_TyTT CTG_ CyC TA CG_TyGTG G_CyCAbAbCAbC CbTCbGCbC CbC CbT CbCbCGbGGbGbTGbTC_AC_GC_TT_GT_ _G_G_TC_AC _TT G_G_ CG_TG_GCA_ C CA_ CA_ATA_ATA_GAA AA AA AACAA AG AGACATA ACTA AGTA AGTG AGTG ATATG_AGT TG_ACGTG_ACCTG_ACT TG_AGTA A_GTA A A_GGTA A_GCCA A_GGT108F KNTSKKQEYENPISKAGHTFDIDVRKSKQTSGSSWSHNSCSRESRKKCDS SDSEDNSDS LDLIL L LWLFL L L L FF FEGSDE SED DGSED DDDMEF D D DMMDMTD CSDLSSF FS SSSFL S IDMGDMED DMGD S DMDSAK I DL SAK I DS SK GAIDQSAKSIDV MSSFFDPLEL LTPLELDLPLELD PLELL PLELQPGLELVPLELDLYVTRSYVRLYVG RVYVRHRSRG LRSRDRSRDL D LDSFAD LDFDD LDEFGD LDLFGD LDFVD LDH FSD LDFAD PQADLD IP QIGP QIDDSSP QI PRKL RKL RKALDLDF LQL I LDL P LDR IWPR I I RILR IVSRGSRD SRG WG DGDLGDGGDSGDVGDSACPAC GACACPRAI RAF RAD E D P E D D E D E DRE D L DPD GAEAE RAE LAEGLTNGND NS I L P I L L I LNI LNI L LEI LGEI L LL N LNL D L EE RRLTR L R GTANTATAS TATAD TAE TANS SNSNS LNS PS RNES RDL TEL SS RDS SDASDL SDHSDL SDP SDMLKALKHLKLLKSMGAAGADSGATGAQGLGASGDGGG QGCGFNKHNKANK YP GYPYPLP SAPCPFAPLHDLHDSHDGHDPKNQKNDFSKNDG NF LNT LNTG NTPD N MIN MLN MLY NT EY QNTGY NTPG N M D N MIN MTY NTDQ N MFIDSVCNL IIQSVEQIVE IVTNKFSENK K C G E NNNLDSNNL IHSNNLQS26364656667 8 9 0 1 2 3 4 58 8 8 8 8686868787878787878TTTTTCTTA CC CC CGCT CT CG CATTCTTTTTTT GAG A AA AA A A AAATCTT AG AA A A G C ATCGAGG ATAGGGCAAGB_ABAT_AB AGBAABCAC_ T T_T T_ AT C B_ ATGB_ATAB_ ATA TB_ATC B_ TGBAT_TBAG_ T BAA_TABAT_TA G_ TyC TA_ GT C_AG A T A_TG A GCG ATG ATG G A_ GA_ TA_CA_CCG AC_ CCG AC_A TCA_T CA_ TCCG_CCC C_ CC Gb TGyG_ CGbCGyGbCTyAGG_AACG_TAT TbTT_T TyAC TbCTy TT_AACTbT_GTyAGTb TyT_ CTGTbTTyb CyT_GTTT_ CT bCGyGGTT_TTC bTGyTT_TCC bT_ CGyGCbTGyT _GTC b CT_ CG A_ CA_A_ C _ C _A A A AA A A A GTA ACC A AT CA_ATACA_ACGCA_AT CA_AT CA_A ATG_CCA GC CA A A G G A GCTAG C G G G C _G G_CGTG_CGT109WVKEWEEEHWDVQLSSRSNRGTWRAMLTDQDNRLMVQEEETSTPSAKSWDL NGL NAKT SKMSKS CL NGL NAL NVL N L NMGAGG GVIDC AIDD AIDDDLNESDL SNDDLNL SDL SNDDLFSNEDLSSNLDLK ND LPK DLP LK LPFK EL GP CDYVIDRDYVLRDYVLRAREPRQEDLNWRLNDLR DLNPRSLNGRD LNI RNNER NLDRS KIF SGRS KIFDP RS KIFDIRS KIFDLIIGP QI FD PQIDLSNYT LCNYGLLSNYENLYLNLYEFLNLYFS LNLYFDLFKL LLFKLLELFKL EF LFKLGSAC QR IDR ILAC DAC GRNNI RN YD NYVRNNQRNMRNL RNGRN YGNYD NYT NYL NYD HGM GQ GLG HD HHGHHT HHVAENLNL SSAELLAAE LLMFLNMAE FH GMASFVF L F S F L F L LPMADMADMAAMAGMAANVL LIDNVIVLDNVS LIANVH ISPGL SD NL SDEQE E S E F EWE I ED SEIE FS ISS IWS IK KSK DTGDGGDLY LNIYNIL LL I VYN LPL IYN LL L IDYN LD L I PY LNI GYN RL ISGYGDE E E E E EVD YGL YGP YGPHSVLPHV HVDF SNSGS L S L S L P S LNS L L IKL IKL IKP IKGLLIGSGNL LIDSADSADEADDAD ADNAD AD KEKEDKENKE ENN NNNLD DFPRLLTFPRLPSFPRLLLFPRLA DFPRL SAFPRLH QFPRLM DGPRIA DGPRILLGPRI SAGPRIPS67778 9 0 1 2 3 4 5 6 7 8 98 8787888888888888888888888TTG TCTA T CGT CGT CT TC TC TT TA GC GT GC GGGT TGATTT T T C TACTACTACTGCTT CCACCACCACCCTTAC B_TAT B_ GT CCB_GTBGABGGBGABGGBGABGCB ABGBGBABCAGC_ GT_GA_ GT_ GT_ GG_C_AT_ AA_ AT_ AT_G_TyT AC A_CAGC C_A TCA_TTC C_ CCCG_CCCA_CGCA_TAT GGC_C C C_ACT A_CGCT G_CCCT A_T CTC_ CCC bTGyT _G GC bCGybCAy TAyT _AGbGGb CAyCGbTAyGbCAyAyGbTGb TCAyTC y C y CyGC yGbCAbCAbTAbTAb CCCCT _TT A_AGA_AGA_GATA_AACA_TACA_GA_TTG_ACG_GTG_TCG_GA ATA A A A_ T _ _ _ _A AA A A A G_C A G_CCTG_CGTCC G ACC GGTCC GGTCC GCTCC GCCCC_GCGCC_GGT CA_CCT CA_CGT CA_CCCCA_CGT110I QHVLDKEGASTGGTIIIVDAIKSAEEIAEEIKDKNDGVLKHGVVNYGD NKL GPGREK LGPSP I P I P IDP ILKDL GPMP I P I P I SVSVVSVSVDAAD ANADAPAEASALPSS PSF PS CPSG AIAAEAALAALAAWAAGAADEPL EPE EPDEP ES KE RKI FNL WSRKLLIF ESI FDSEEF LSE FSLSE GLSSEELSETLSELLSEFDPP PNGEPP PGDI PP PGDLPP P EGWFKL TL KLFSL KL FAPE LA KT PEKGAEAEQAECL PKVPKGPKIAEMAEDPKD PKD LPPLFSLPPELF LPPLGLPPLTHGCI FGGFGDLQS LQL LQHLQVLQLHD HHL HHD P PAP P GP P S P P D P PE LPQPLDLPQLPAVGG L VGLT VGSVVGCIDNVI E LNVI L L LNVR I R I I R I P R IIANVWNVNVVNVS R IGRI F R ID AFNVQNVDNVS S S LASFS SAFHAFAS SSS S ES IGS I GIS I DATPPATGT P TLT L T TGGDGIGD GDPGDGE EQE EGE E SD DRADGADLADNADDADLQGGQGWQGVQGQYIGLKKNYIG KRYGGWANWANWAEWADWASWLWSFSF PSFP SFLNIKLPSPHP P PLA LP L PAPAMAE RNAE PAEGAENGE SPRILK TGEPRIHK QGEPRIMSMA DGPERGSM LGPEQ RSSM EGPESRFSM PGPERC SM GGPETRLSM PGPE DRS SMDGGPERL PDLA PEFHPQLA PE NFS PALA PEEFPPSLASPEFLT091929394895896897 8 9 0 1 2 38 8 8 898989809090909GC GGT GACCCTC G CTC CCCAATACAG AGC TCCGCC CCTAA AGCAC CAACGAT CAACATTGTAT C T TATBABCBAGBAABAABAGBATBAAB CBAABAGBAABATBC C_ ACG_ACC_ T_G_T_A_C_T_AC_CG_CT_ CT_CC_TAC_TyT TA_ TCTC_AATA_TATA_ TCAT C_ CCATG_CACTA_TTATA_CAGT C_ACG A_ TCCG A_T CGC_ CCCG A_TT AbT C yG_GAb TG_ C C yTAyGAyAbCC bTC b TGG_TC _TC _ CA Cyb CAyGC _ CC bTA C_GCybTAyC _GCbCA C_ACyTbC T yb TyGyC _TG_ C TbT Tb CT ybTGG_TG_ CG_GCAGT CAC CATG ACG G ACG A G ATAGACATA A A ACA AGA AGA_C A A_C G A_CGTA_TCC A_TG A_TGTA_TGTG A_T TG A A_TCTG A_TGTTT_CCGTT_CCCTT_CGTTT_CTA 111QRILEELKDDAPPLPGTSGPQLATMPSPSRMEQTPPLLSDRSPSPAFLALSP VSASP VSGSP VSMRKKDRKKWRKKDL RKKPLRKKE RKKGRKKDPVVHPVVAPTVVSE PVVSEP DEPL EP DPF IPFPFPFPF FPFL PFF PS NI PST PKPSEPP P SLGP GPLP PDLGPPP LP P LLG PD KSGEF R EFLKSRGTECSIKGGSKSGE SQKGS SGKGMKSGD DR EVR EGR ELREDR EDGET LEP SARGEET EPS LGET AESVFGEET ESLIP L P L E P LDR VT R VER VHR VVR VLRVLR V SGKGQPGKPGEVAGM DVGQSFS LAFGV AGF DSLKS S SVAKS SVGKSVSP SKS SVDSSKVGI SKSVDF SKSA VDFEE LSLFEEKSIFEEELSFEEKISGDDS SVS SAKA KAQKAVKA KAGKADKASPGLPGIPGNPG NKWNKLNPNLL NRNDNGTGYTGETGITGEEQGFGDDGD DQGSQGDSGPEDPN KG K KN K NGDSGDEGDDLKLQL GD G H DAGDMSES SQLI SESQSGSESQSGSESVSF DSF LSF GTTF SE TLE TPE T E T E T E TQQ QQEQQLQQVAE LAE LAE LAT FTT FST FLCT FQST FDS T FDG YGQGNGPPLAPEFAPDLA PED FLPLLA PEFMY DNNPFGGY LNNPF LGPY LNNPF FGPY TNNPFGGY ENNPFGEY QNNPFGY G GNNPF LGDA FP MPPQ GFA IIP MPP SGVAMSPP P CGSAMEPP PDGPH 40506070809900911912 3 4 5 6 79 9 9 9 9191919191919ATC AATATAATA AC A T AGAGAT AT AC AACACACACT GA A A AA AG A A GTGCG G GA GTTGGTGGTGATGGC AC T B_A CG ABA _CCCB_AGT B_ATCB_AA TB_G AABA _AGB_AAT B_ACCB_ TABTABTAB_T TB_GC_GA_TCC GGA_CGCG G_CCCGC_AAC A_TAC A_TTACC_ CCAC G_CAC A_ TCAC A_CGACC_ATC_ GTTT_ GG CTTG _GTT_CT ybC T ybT T yTbCGyGbTGybTGyb CGyCbTGyb TGybCGyTbC T yb CT yAT yGT yAG GGGT TGCGCATAbAbGAbAA_A_TAT_CA A CCTTT_TA_A CGTTT_TCT _ACCT _CAGCGTAT_CCCAT_CT T _C_C_AGTAT TAGC_AC_TAAT_CGTAT_CGTAT_CC TAC TATG_GG_AG_G_A A A ACA A A A ACGAT_CCTAT_CGTG_CGC G_CTTG_CGTG_CCA 112QPSATECLLEGLDGLDPLEQGVDSLLDLLCGEIHDVMPLCDDLGSVPSVPVPQ VPEVPDVPIVPVVPLVPIVP PLP A P QHLNS VSHLAS VFL HLD SVAHLSVAHLM SVQHLH SVD HLDLMLKS VKHS VDHS VSHLAV SVHHLGVL ESVEHVPVLIHVFIP SLP S S P SGP SWP S P SVP SEP S L P SAP SDP S ESP SDSP S LVVLVVQVVPVVQVVQEVVM VVPVVD VVGVVDVVWVVDVVDPSELPSEDPSQPS KPSPSP PS NPSF PSL PSDPST PSAPSEGTEGTGETVGETEGETHGETLGE EGEDGENGEGGEGEGEFESD ESVES LESAESVS C LTSN TSD TSKTSL TSCITSM TS TPGF PGDPGGPGT PGAEPGDDEP CEP LEPEP LEPD EP EEPDS EYS E L S EHS EDS EAS E SG EDSG E G ASVSGT SGESGE SGRFPELF E P F EQF E F EAF EDGF EEF E FEEAFEE L FEGFEL FE DTGTPQTGPPPTGVPTGRSPTGAPTGLS LDPTGTPTGDSPTGEIPTGD PETGQ PEGAPEGLGRGLGA GH A GPVSEEQ GMF L TVTMQSSSESQSLSE QSGSESHIQG SESAQG SESGL QG ESESQEQG SES LQG SESL QG SEDQG DSENQGVQG SSEE SEEVGQSL QQLQALGPPGQPQQQLVGIQ A Q QGQQQFQ MQLL QS LQS LQSQGQSSNG A AGQLDGGTGQDGQEQ Q Q I GAGTGHGCIP MP PYAMG VPP PQAMG APP PGLAMGPP PGA G GP MPPGDA EP MPPGGLVA DP MPPK GLA SP MPP LGDA FP MPPGLA HP MPP EGSA PP MPP LGPA LP MPP TGLA PP MPPGCQ 819102122 3 4 5 692927 8 9 09 9 9 929292929292939CTGC C C C CVC C1 TTC GTC G ATTTCC CTXTATACTGTCTGCTG CCCTTAT_GT CB GTB GAB GCB GGBGCIGCB GTCGGB GACB GTBGGB GPOGC_ TA_ T C _TCGG TCCG T GCTA_ TC _TPTA_ TCB_ TA_ TT_ T C_T T _TG GGGC TGEGCGGCAG GTGCAG ATG AATTGGST_AyGT_y C T_yGTT_ GT _ G T _ T _ AT _T T _TT _TT _ T T _ T _bGb A bybC T ybT T ybT ybC T ybC T yb AT yb AT ybT T y C T yG_CA_ CA A A_TG AAG_GA AGG_GA AGG_AAA T G_BA A_7_GA_TA_ TA_ TA_GAb_AAb_4G AGG ATG ACG AAG AGG AAG A CCCA G_CTA AA A A A A A GCG GC TGC CGC 1GC CGC TGC CA A A GC CGC TA AA GA _ _ G _ T _ . _ C _G _ _ G _A G_C A G_CB_113QPSATECLLEGLDGLDPLEQGVDSLLDLLCGEIHDVMPLCDDLGSVHSPVVSGPVVS WPVV SL PVVSS PVVSDPVVSPLPVVSI PVVS PVVMPVVSPVVS PVVWPD VVIPVVVP EHP E I P EDP ERS PYP PVPS LPSHPSP PS DPSPSE PSDGE TP STFGTD EP SDEGTGEP SDFGTS EP SNGERE TQ PSHGETESEP SQGETTGEP SESGEETSSHGETL ESNSGETSLGETLSESHESNGETKPGETFLGET LESKEST ESAIKSGISGP SGTSGGSG GV GVPGL PGPGP PGDPGKPGS PGSFEDFEP FEFETE AS E S EDS E S S EQS EGS E S EAS EAS E APEVPEPED PE FPE F F EDS F E F E P F EAF E F EEF E F E F E ETGQGR TGTGCGWGKPGPGAPGS PGLPGMPGTPGAPGWPG GS TGMTGR TGQTGLTGETGL TGTGGTGMTGSTGPTGEISE KQQS S SE D QSP SESLSQSESAQSESKQSELQSDSEEQSMSESAQSEM SP QSSEAQSESQEQPQESEDSE NSE AEGQKQ QGQGQ GGQDQ E GQM L Q AQLQIQQGSQ QSVQS LQSSQS SQSET GQYGQLGQDGFGQD QG Q Q AG GTGRGQAGQIAMTAMFAMAMAMAAMCAMPAMLAMAMAAMTAMGAMGMKPP PGSGPP PGYFPP PLGIQPP PD G QPP PGEVPP PGGEPP PGCQPP PGSSPP PQ G DPP PA G GPP PGTLPP PGDPPP PGLA CPP PD G K 132933934935936937938939930941942943944949CTGCTG CTGCTGCTGCTTCTC CTA CTGC _T1_ CTGC _3_ CCCCGG GTC C GAC CCrGT r TATATGGAB_ TTTB_GT TB GA_ T TABG_T TAB_ GTG ABG_TTTB_GTCTB_GT CB_GTekGT TAB_GTekGTGT BGTGCB_TA_GG TTAGG _C TAGG _ATG_CG T A_CG GTG_CG T G_ GTG T A_ GGGTC_ ACGnT ilG T GACGnT ilG_TATG T ACT y T y T y T yTT y T yCT y T y T T y T_y T_y T_y T_yGT_yGAbTAbCAb CAbTAbAAbTAb CAb GAb AAbAbAAb bTb CG_A G_A G_ CAG_GG_TG_GG_ CG_ TG_ T _ _T_A_TA_ CA A_A A A_G A AGA ATA ACA ATA AGACAAG A G ACG A G ACG AA G_C A G_CAC G_CAC G_CTGC CA A GC TA A GC CA A GA A A GCATGCATG _ _ _G G_CBG_C A G_CBG_CCC G_CCT114QPSATECLLEGLDGLDPLEQGVDSLLDLLCGEIHDVMPLCDDLGSVHSPVVSI PKVVE PVVGPVVT PVVIVPVVWCPG VVDPVVS APVVE PVVIFPVVSSPVVTIPVVHPVVDP E PSEGPSL PS QPSPSPS VP R PS FPSR PSP PSP PS APSEGE T RP S LGT SAEP S SGEGE TPSM DGEETPP SQGETW SEP SQGETD DEP SPLGEETPS EGEVE TPSQTGEETSP SGGETDGETL ESKESGSGETAEPESTGTE EG TESDG ITIESASGASGR SGLSGI SGFSGVSGD GQGL PGAPGPGSPGFPGDFEF FER FE DFESE ELE F S ELS E GS EES E L S ELS EIS ETPEAERE F EAF EVF ENF E L F EAF E I F ESF E S F ERF EAF EGTGPGQGE TGSPTGDPGAPGD PGPGDPGPPG GSENQ DTGITGE TD GITGTGATGGPGIPGAPGPG RTNPGA GVTGS TGR TGTG SLSEGQSQSES LQSESQSHSESVQCSEEQSKSEHQSLSE QSASE QSNSESYQSEVQ P QNQ ASASER SE YSEPQQSQQSQ AQTQQ Q KQSIQQTLQ HQKFQ GQSQQS RQSQAG QG GGQDSGQPG D Q Q QGSGEG Q AG QGQLGQSGQGGQSGRAMTAMGAMAMTAMAMSAMEAMAMSAMPAMTAMAAMLMGPP PGRQPP PGDPPP PG G GPP PGNSPP PA G QPP PGLSPP PGSAPP PA G GPP PGEQPP PGSFPP PGSSPP PGSLPP PGA G NPP PGDR546474849940951952 3 4 5 6 7 89 9 9 959595959595959CTACCCCC G CGC G CACCCTC C CCC T CG CCGATGABTGATGT TGATGT TGCTGGTGGT C TAT T T C TAT TBT C _ TABGBGBAB_CBABAB GGB GCB GT B_GT B GGBGT_TAT_T T_TT C TG_ TGT TG_ TC _TG_ TG_ TG_ T C T C _ TA_T_G T G TCG _GTA_CG T A G _ATG_GG TCG _ATTCG _TTT_CG GTA_ TCG TT_ TTG T G_AG CTAAG _CTA_ CG T ACC AybT T yAT yGT yAAbAbCAb GT yCT yAT yGT y T y T y T y T yCA AbTAbAbAbGAb TT y T T_yCAbGAbCAbAbGAb GG_G_A_ G _CG_AG_G_GG G_G_GCG_AG_G_AG_T_A_T_ TAG A ACAATGTGC CGTGTG GA A AA A G_CTG_CCA A A A G G_CTG_CCA A A A A A A A A G G_CAA A G GC CGC CA A GA GA A GC CGCGC A _ G _G G_CCC G_C A G_C A G_CGTG_CGC 115QPSATECLLEGLDGLDPLEQGVDSLLDLLCGEIHDVMPLCDDLGSVHSPVVPVPVR P PVVP P P PVS P PVPVPVL PVEPSA ENVPSCEE VPSENVV PSA ER VPSETIVVLPS D VVEPSSVVSPS AVPS SEVV PS ALVPS DPVA PSDVPSYVPS SGE TTGSG GGPGNGE SGETPGE EGEGEAGELGEEGENGE RP SNE TPSYTGEP SYTD EP SVTL EP S ITAEP SG TFEP SP TPEP SPTE TPSG LE TPSGE TS STMES LTS ESLFE TSPTSGLSGI SGIISGLSGSGVSGSGVSGLGT PGPGE PGNPGDFE AFES FE FFE HFE MI FE EAE EGS EAS E D S E KS EGS EGPE STGGPEQTGPIPEGSEPEGPPEGGE E F ESF EASPGKPGTPGP FPEGQF ELPGAFPE P FGVPEGT FPE FGCEFPGLSGFGKSESQ ITGI TGVTGTTGVTGVTGATGTGATGT TGGTGTGG SVSE QSR SESVTQSEEQSASE QSE SES EQSESRLQSEPQSDSE VQSP SEQQSSSESVSQSED KQSE NCQSEHQQSQQAQ HQEQQKQKSQ QQLQ GQLQPQSSQSQPIQSQVGTGLGQNGQTGHGQDGQPGQPGQPT GQGGQSGQLS GLGRAMGAMYAMAMEAMLAMAMSAMQAMAMLAMFAMAMCAMEPP PGPSPP PGGEPP PY G MPP PGNLPP PGYTPP PK GPIPP PGPEPP PGPVPP PGWFPP PGGLPP PGLDPP PGRQPP PGDIPP PGTA 950616263964965966 7 8 9 0 1 29 9 9 969696969797979CTACGCCCTCCCGCACACGCTCAC CAC GGT BTGCTGGBTGTTGATGTTGC BTCTGTGTATG TTTT TT TGC_TTGT CGGB_ T CT_ CBABTBT _GC BGAB GCB GAB GTB GT BGGBGTA_ TA_ TG_ TAATG_TG_ TT_ TA_ TC_ TA_TA_T_TTA_AG TCG _GTTGG _GTATG _TTA_ TG T A G _CTA_TCG T G_GG CTC_ ACG TTG _ATT_AG TC_CG GTC_TG AybGT yGT yAT y T yGAbTAbAbCAb AT y C T yCG AbGAbT yAAbC TAybT TAybGTAyCbT TAyG bTTAyb CTAyb AG_GG_TG_G G_GG_GG_TG_G_A G_AG_GG_A_ T _ C _ TA AAA ACA ACCA AAA AAA ATA AGAA ACATATG ATG ACG AAG_C G G_CGTG_C A G_CTC G_C A G_CTC G_CCA A G_CATA G_CACA G_CATA G_CTTA G_CTA G G_CTA G G_CCA 116QPSATECLLEGLDGLDPLEQGVDSLLDLLCGEIHDVMPLCDDLGSVHSPVVSF PVVVPA VVNPVVW VPVVMPVVMF PVVE PVVSKPVVRSE PVVEE PVVITPVVE PVVNPVVCPSE PS RPSPSVPST PSPS CPEPYPS RPSE PS PPSL PSLIGE T WP S SGET GEP S TGEVE TV QGET PSSIEP SEIGET NEP SDLGEET DPS PGENE TSP SKGECE TEP SLLGEETS SQGEETS LGEQE TE C ESGEG ETSAGTAESNERGTENESSFSGT SGGSGISGYSGG GAGG GQPGRIPGAPGSPGE PGL PGFEP FEP FEYFEFE L S EFS E S ERS EES E AS ELS E S E S E MPEAPEGEDEKE F F EIF EHF EKF EPF E E F EKF EAF EAF E ETGTGVPGPGKPGVPGEPGGPGPGSPGPGVPGF PGKPGPQGS TGTGVTGL TGKTGKTGP TGGTGTGKTGP TGL TGI TGTSE QQSV HSE EQS EQVSE C QQS E SEQSYQDSEQSAQVSE I QQS TESE L QQS PPSESQQSASE AQSLSSE D QSEESESDQSEA SP QSEV SDQSESMLGQYGQIKA GGQTGQQGQFQLQ QKQQEQ Q WQPQ AQKHG DGCGMG Q KGEGQYGQDGQDGQEP MPA T AP MPAIAP MPSLAP MPVLAP MPQAP MPKAP MPDLAP MP RIAP MPNAP MPN DAP MPGAMP RPAMP IYAMPNPGL PGE PG DPG VPG YPG DPG APGT PG KPGS PGFSPPGEPPGEPPG M 374757677978979970 1 2 3 4 5 69 9 9 989898989898989CTTC CACGC A C A CCCTCA C G CACACTCTGCTCGT TGATGATGC TGC TGATGCTGC T T TGTCTAB TGT TGAB_ TCT B_T TBC_ TGCB_GBGBABC BAB GABGABGCB GC_ GTBTCCTG_TA_T C _ TC __GG T A_CG TCGG _T T T _ TG_T C_TG_ TAC T T_ATG G _CTC_TG T A_TAG T A_ ATG TCG _ATGAG _ATT_CCG T G_TG ATA_ CCG TTG _ATA_CT yAT T yTT y T yAT yTT y T T y T yAT y T y C T y T y T y C T yAAbAbTAb AAbGAbCAbAAb AAbAbGAbAbAAbAAbGAb AG_CG_GG_ CG_G_GG_G_ CG_G G_CG_GCG_GG_C_ _ TA ACA AAA AGTA AACA ATTA AGC AGAG ACA ACATG AGG ACG_CAC G_CCTG_C A G_C G G_C A G_C A A G_CTA G G_CATA G_CATA G_CGA A G_CTTA G_CGA G G_CAA G G_CTA 117QPSATECLLEGLDGLDPLEQGVDSLLDLLCGEIHDVMPLCDDLGSVHSPVVPVE PVPVPVS PVF PVPVPAP P R P P I PPSY ENVPSEWVPS LESVPS KRVPSVVPS D VV PSMVPSA VVV PS FLVVLPSDVV PS FLVV PSH QVV PSVVV PS ALGE TP SALG ETPSVLGT FQEP S LGETQ QEP SVGETAR EP S LGESE TLP SQGET AEP SKGEGE TPSAGEKE TPSFEGETKEP SEFGET KEP SQTGEETPSQGEETM SYGET SESARSGNSGSGASGSGPSGGSGS SGVSGLSGSGQPGQPGPPGEFEC FE NFEFER FE DFEFES FER FEFEYE E S ELS EPS ETPETGHPE SQTGNPEV TGKPERGPPEGSDPEDGSPEFGILPEGSLPEG GNPEVFGDPEGE F EGF EAPGSPGQFPEGAEGAGN GNTGQTGL TGLTGL TGSECQ MTGGTGETGVTGATGQTGLSDSESKQTSESSEQSEEQSESESSIQSELQSPSE QSVSE QSL SESVQSES PQKSE QSE SES SQSETYQSELQQIQQYQEQQEQPQQEQQVQQTQ VQ QQDQQQSQS FQSMGYLG NGQFGEGQPIGHSGRPGRSGQVGQAGLG VGQNGQDAMAMAMTAMEAMEAMAMAMIAMHAMAAMGAMPAMLAMFPP PV G KPP PK G VPP PGSAPP PGEEPP PGFEPP PH G VPP PD G DPP PGPLPP PGVTPP PD G DPP PGSYPP PGIYPP PGFRPP PA G N 78 9 0 1 2 3 4 5 608 8 89999997 8 9 09 9 99999999999999901CTC C CTC _2CACGC C G CCC C T CGCCC TGC TA GCTGGT_GreTGTTATA TT C TATG TT C TCTGT CTGGT B_ T CTB_ TG TB_ T k TnTGBG_TCCB_GT T BGA A_TABG_T CGB_GT TABGA _TAB_GT TABG_TGCB_GTG GB_TC_ GCG TC_AG TC_ TTG TilG _TA_TG TTA_AG TT_CGG TC_TG CTG_GG TC_TG T A_ GCG T A_CG T A_ ATG T ATT y T yGT y T y T y C T yAT y T T y T yAT yAT y T yAT yGT_y CAbGAbTAbGAbAbAbTAbGAbCAbGAbAAbGAb TAbAbGG_T_TG_AG_T_TG_G_TCG_TG_AG_G_C_A_A_ C _A_CA AGA AAA AGA A A AGA ATA ATG ATG ATG AGG ACG AAG ACG_C A G_CBG_C G G_CTA G_CAA A G_CCA A A A GC TGC TGC TGC TA A GC TA GCG GCTTA A _ G _ T _ T _ _G G_CCT118QPSATECLLEGLDGLDPLEQGVDSLLDLLCGEIHDVMPLCDDLGSVHSPVVS DPVVST PVV SDL PVV SKPVVNPVVLP PEVVGVVSEAPVVE PVVE P PVVVSVVS TPSVVASEGP E S P EDP E P EDPSEDPSE PSEAPYPSYPSPSQP PSVFVEGE T LGTTGGSGFGLGDI GSGEGEQGEDGEDGEGGEL DKEP S LDESGTSTSKTTTS ITSS TSLITSE TS LTSI TS PTS T PGWSGL PESGWPVESGHPVESGS P SI ESGR PDESGGP IESGDPSGHESP IESGTPGDEYPGTTEPGQLEPGPPLLKTFE LPE C FEA PEEFE SPE P FEPESRFEPEWFEE SIFEE TE EES EAF E C F ESF EL SEFEE LSQFEEASTFEETPVLCVITGQTGP TGVTGTGVPTGSPGS PGVPGTPGPGSPGPG KKDGGSE EGFGPGV GASGGTGT TGDTGNTGYL TGHTGMTGAVLEQS IQHSEEQQSTSEQSGQESENQQSVSE QQSVSEQSASQSE QQSQ QSE QQSML SEQSYQFSE QQSSP SEQSGQSEPQQSSP SESAP TYLG I DQGQD QRQPQFQQQSQ Q QMQQLLA VGPGSGLGIVGIDGLG N Y QV Q DG KGYGLGQSRGSKQLNP MP PGMA PP MPP TGAA PP MPP FGPA TP MPPLGSA KP MPPGEA VP MPPGTA AP MPPA GQA SP MPPGNA EP MPPVA G KP MPPGLA DP MPPG GSA TP MPPGMA EP MPPV GGK SRPQSDLT10203 4 5 6 7 8 9 0 1 2 3 40 00000000000000 1 1 1 1 11 1 1 1 1 1 1 1010101010101CTTC C CCCTCTC C2C C C C C AGCTT GTGCTATCTGTC AT_G T PTCTCTT T ATGTA TCGTT TGAB_ GT CB GC_ TA TB_GTG ABG_T CGB_ GT TGB_GTCT BG_TOGT T BGA _T TABG_T TB_GTAB_ GT CBG _GTCB_TC_TG GTCTG _CTC_ CCG T G_CG TTG_CG T G_ TTG T G_TGTTSGCG _TT_ATT_TGCTTTGT_CTT_CGTTTTA ATT y T yGT y T y T yGT y T yTT y T yAT yGT yAT yAT_y CG_yTAb CT AbAb CAbCAbGAbAAbCAbAbAbTAbAbTAb CTbT_ _G_ C _T_ _G_CG GAATAGG GG GG GCGGG G G_ _ _ _ _ _ _A AAA AGA AG A AGA ACA ACATA G AAG ACG ACG ACG ATCAGG_C G G_C A G_CGTG_C A G_CTTG_CTCA G_CCCA G_CB_A G_CACA G_CTA A G_CACA G_CCA G G_CTAAT_CTA 119EEISILLMGGESEELGANLNRNLAKIVDADIYESFEPLIRAKRIGIINDAVTSF E LDSF ESLSF ED SF E ESF EDSF EDP ELL P ELEEP ELDP ELLDP ELDP ELE P ELDLM QSEDVPV VSVDPKIVDVLNLNNLNLDNL PNL SNLDINL L S SDILGL DEPKNDGE PK GGLDPK GEDLK L FPK GLDGPK GDFK CEKWKLK I MF CIMTCMGCMLK ECMGLK CMEKDF CMFDMEFLQLKF LLS LKLMLLKL LLTLKL SLVLKD LESEDMTGMTCI IMET S IEIEVMTQMTMIMETL IEDVD T MTDTGLVKVGV VKVG L V KVDVSV VLGVL GVD GVGVGGVD GVSGVLSVTSVK LD VK LLVKLKV LD VK LAKVH VKSKV APLPVK WLSGPIWE PH SG W SSP PWVP L PSDWS DW SAPWSASVAAT S TGI T F TWTVTL DSMPMPQ MPVMP SMP FMPWMPDSQMWYLIDLLYLIDGYLIDD DYLIDPPYLIDPYL TIDGS P GTE RS PE L TS PE P TS PELTLS PDTED S PEPTPS P GNN SPPKQKLDRLKQL RNKQLLKQALNSKQG LEKQL TLMS SSNTHS SNTSS S SGT TSE S SSD S SS L TSSSNT E LVISS SSMGGNSPQ DLK CRPQHKQ KQAKQPK FPQ D QRPDDSRPD GRPDS RDDCLSPHEQCSSPHLETCLSPHPESCFSPHLELCA CSPHEDCSSPHEACGSPHEDPLSA PPIA G 51617181910 1 2 3 4 5 6 7 802 2 2 2 2 2 2 2 2101010101010101010101010101ACTACTACACAG AACTC G CG CTCCCCCA TCGA GC C C CG G A A A G G GCGTCG AAC T CC CA GCACACT GTAGBGBGABGGBGAB_GCB BATBAAB_ABAABAGBACB GBAA_AG_AT_CT_ T C_GG_GC_GTGA_GT_GT_GC_CTT_GG_CCGA_ TCGA_A GGA_TA GGC_ CCA GC_A TATA_ TCATA_TTAT C_ CCATG_CCATA_CGATA_TAT C_AA A_TT yAbTC _GT yTAbC T yAbC T yAbTTT yCAbC T yAbC yTyTGT yTbCGbGGb CyCGbTGybCGyGbT yTT GbCAyGCbTTAT C _GC _AC _ C C _GC _ T _GT _ T _GT _GTT _AT _ T _T_A AT_CGTAT_CCACA A ATGAT_CCTAT_CCCAT_CGTAT_CGT TAA_GCTGTAA_GGTTATAA_TGGT TA T ACT ACT ATC ACA_GGT TA_GCT TA_GCCTA_GGT TA_GCC 120SPLTLSNRPTTLRSSNQSLSHLLSPSLAGSFLSSHDGDDMIADFDIQPAELLMQSSLMLMLML LMLME S E E S E E ECE LSDSQ SD GS SD Q SLL S SNEQSSSDPQSS ESEQ S DK WS S LDKD DGK L KDLKDVKDQKDDKDIKDDD DTSKDHKDG VKDLDKDDEKDLNT EE ALLDM DM DM DL AVLA A S A A A A AWTDLDG FSMEDMTDMF PWGPWAPWP PWDPWF PWGPWDT T TSGMVTGSVDVD VTGGTGQVTD GCIVTDD GD PEPKIPEPKWP PEPKVPPEPKSLPPEKDPPSVD SVSVL SVGSVD SVDEKQLPPEKSQ GADCIVALDSVAHSS LVVGVVPVV VVLVV VV VVLPHDP VALGSVAVSDVAE SGVAAQRRDLNQRDLNSQG RDLEPQRDLDLQ DLLAQ DN LSSLQDLMANS EGQMFQM QMIQMNNDNNVNNGNNSL QM NNQQMDSLHLAL S L LRL DRL TRL DVE QVS DVSPVSRVSLSLNNS GMPISQSMPIS GMPISFPMPISCMPISSGMPISLM PPISLDMTQLI L IGINIPDVIN S VI L S S E S S L S S T S SGS SGS S L S S F LNGGAGGE GGGGGGGGMVKQVKCSAPAP PAHPAL PAL PAIVK VKEIVKLVK VKDGASPPIDSSPPISFSPPIQSSPPILCSPPITLSPPIDLHA Q VDEHA Q VQEHAQQ VSHAHHADHAG Q Q V D Q V G Q VLPHA Q VDLQDQI LDTL92031323334 5 6 7 8 9 0 1 203 3 3 3 3 3 4 4 4101010101010101010101010101TC TG GTT TT TGTACCTCCCCG CTCCC G CA AGTAGT CGTG AGAGTGT GGGACGC CGACGACGT CGT GCTC ABABT T_ CTT_C BTGB TTB T CBTG_CTA_ CTC_ C C_ CA GB_ CGT B_CA TB_ CG AB_CAT B_CTCB_CCCB_ATCB_A A_CAyGCAC_ CAy CCA A_ TAy CT A G_CAyCTA A_TAyTTTAC_AG A A AyT G_ TyCG T A A G_TG yGACG_ CyCG CA G G_CG yCTA A_CG yGCA A_TG yTTAC_ACyTG A_TyTTCbCbC CbCb bGbC TbC TbT TbC Tb GTb GTbGGTbCGTbGC_A_TA_ CCTA_GC_TA_GC _GCCTA_ T CC_ C _TTAGC ATA_AGA_TACA_AGA_GATA_AACA_AGA_TATG_AGA GCTA GGTA G G A GGTA_GTATA_GGTGT_CCGGT_CCCGT_CGTGT_CGTGT_CCTGT_CTAGT_CGT TA_CTA 121IAPYIPLMASPSIIPSPTPTQPQSISAAISHTPTNSMASARGSVAAMIDDLTDL EDL LDLTDDLDDLTLDLDSAHSL EAHSLMAHSLFSAHLWT SAHLGS SAHSLFAHLDQ DENTLLDNTLNNL SNTLDNLDPNTLPQP D P P C PVPLP DE SQTAIELQTFTAEQTF LSTAGQT LLTALLQTGTAL LLTADFW S QT EQTDYFSIGW FVYFSILW FD YFG SIFLW LYFSIIFDW EYFSIFHW SYFSI TFSW AYFSILFADEVPG DSVADLADGADMAD ADQADLGD LGFLGGLGGLGPLGLGDSYDPHT PHL PHDHV HG HDLGS LGDLGI LGLGVLGWLGDSVSSANSSASNL S LPSHPSVPS LALLADAGAQLAPAPPAGPCLLVETWASP VE GIANSDTGVEFANSSTDVEPANSTVVE DTSANSANVETDG SSENE DG FLSE LNEAG FD SE RN NGN NLENGENGEFEGENPHS EFS S EPS ESGSEEM DLHDCNLMQPMQRMQD MQPMQLLM GRLR S RQR L R FSR FARF LEYLLNLNL L LGLDLQLDGCDGGDGSDGTD DHDADLDGFDGGDG PDEACGASGA GA GAEGA DLGA DMGAIG E GAIGGAI EGAIPGAIPTGAI LGAI FYRGQIAQIQIQIQIQ DCQIWG DLQ KIWC EQ D G Q DQSQ DDSQ DSFQL DG H Q KLWGQWG D Q K D Q KLWG G Q KQSWG Q KIWGD Q Q K DCRM YIH 344 5 6 7 8 0 4 604 4 4 4 494 5152535 555 5101010101010101010101010101A AGC T AC AGAT AA AA AA AG AT AC AA ATT ATCAGCGGCAGCCGAGTGGBGABGABGTBGABGGBGCBCAAGC T BA _AGB_AAT B_ AA TBC_AG AB C_ACCB_AA_C GCAT_C ACAG_CAT AC_C ATAT_C CC ATA_TACC_AGG AB_G A_T CG A_ TCCG A_CGCGC_ CCCG G_CCCGC_AG_yCTG_yGCG_y CT G_yTTG_y CCCG_yGC_TGyTTCGG_CGyGTbTGyTb TCGyTbCGyTb CyCGTbTGyTGbGbGbTbCA_GA_AA_ CGbGGbA_GA_ CGbA_ TGbC yCA_TTAbTG_TTACG_TAG_A_CGTACGTAGG_GT ATG_TT ATGATC _GGACC_ CGAGC_ CGA_ TGA_G GGA_CCGA_GA_GATA_CCC A_C G A_CCTA_CGTA_CGTA_CGTGCCTTGCCTC GCCGT CGCCA GCGCCTGCGCCCCCGCCTACA_AGT122CLACLVAKDGTNIRHFECSKCGHGPLPPCDLAFSESVAQVGHSAANQIAQED IDS EDLD F WVHVHLVHIVHFVHWVHVHD FQE SGSQE S SQ DGE S TQ DCE S LQ DFME SDYEVPLNE LY RMNEGYEY RSNERF NERSYTYEYFADGNERCEQEDGALDST DSEVPVDEVP L DEVPIDSHDS LDSD DEEVP DDS LDEVP DDS L IAELDIRVLELVIRVHEL LRVTISELLI INRVLELRVDIRGNEELRVVIRELDRVLG EPEDDPFDYADYS DYGDYGDYD DYLED LES LEALEGEGEDS EATTDSVSVP SVI SVQSVF SVDS RAFAPA AILAQ LAL LADEP DPCWPCVPCGPC LPCDPCGDRGVRGWRGGRGRGRGSAI LPHPPHP PHR PHPHDPHGAE D AEPAE P AERAELALA G TALCNPNLGLNLNLLLL P S L P SGP S P P SNP SNPESDPES LGRDEYS CNEEYP CN EYH QCNSEYL CN EYACN EYMA VGSAA VGEA SPVGNA SSVGSHASALA MWVSQVGLVGLVG GAGEAAYRGEASCYR F EAYRSE EATRL EADRS EADRLADAGSASA AA ASTAS CASDLEAGFAGGAGSEAGLAGGAGLAEGLRMLMPYCICR YTY MQ MPQCR Y DCR YLY GY M MDGCRGCFGESG GGESPTGES LEC GSQGESPLGESEIGESDFGVD YL RY DRYL RY QRYI RY DRY GRY HRY DRM G 758 9 0 1 2 3 4 805 5 6 6 6 6 6566676 69607101010101010101010101010101T T TT T T T G G G CACG GC A CGC GT GGGT GAGCCAACCACGAAC TACAACTG G G G AAACAAAG GTGA GTTAGGBGAB G B GTBGABGCBTABAB GBAB ATBAGBACBGABT T_TT_ TG_ T C_T T_ C_ T_TT_ T T_TG_ T C_TA_TC_ T_GAC_TyGGCC_ CyCGA CC_ TyCTGA C_TyTTGA C_C TyGCGCC_AG yT CA C_CG yGCCC_ CyCG CCA C_TG yGCA C_ TyCG TCA C_TyTG TCG C_CG yC CCC_AG yT CA C_CyGAbTT AbCAbCAbGAbAbCGbCGbCGbTTbCbGbTbCAbCA_CA_GA_GA_A_AA_TT_A_G G_ C _GG_G_GG_TG_ACA_ CA_ CA_ C CA_GTACAA ACA A A A A A A AGA ATA ATACA ACC A AGTA A G A A ACA_ACT CA_AGTCT_CCTCT_CGTCT_CCCCT_CCGCT_CTACT_CGTCT_CGTA A_CCT123S ALLALQRPSVADGADEGADTGAAPARGDRRADDLDHPHDPDHDPDDAQDACDA DASD GDAGPHPGGIGLT A G VGLGDQ QSQ QDQPVQP LQ QDQGPDQEQPQAQPAFWMGTGAEPVGAEPDEGAEP SAAG EPHASG EP LAGGP LAVLVPVVLVFDVLVSVLLVIVGLVGVQLVWVQDPLVSDSGCIEDDEDG D DPDIEDD GL PGLD GL LGL RGL LGL PGLGDS DT PE T SPL T PTQEPLTTWEPP TTVEPT PTGETPTSGRQGRG ERQGRL RGAQRDRQGRNRQGRNRQGRNS RQGLRWE EAGAIGTLRDEPAI EPL GTN RSAILGTP EPRNAI EPS GTG REAI R EPP GTN HAIGTLVQPQ QLQH QSM Q Q NAQMADSV F ADDSV ADLCV ADQV S ADLT V AV ADGADDLGTSE LNWVLVTVA VSRVQRD GVP GVGGVGGVEGVLPGVLGVD FASGACW GALW W PGAGGAFW PGASEWV GALDPQGP TQPQGPGL PQGPEI PQGPQPGPGCDQP LQP I PQGP FDFDES LLEGLELEL LET LEQLEFVQSVQDVQHVQVQGVQQ QD P TAEEIAE LAECIAEAEAEDLVQLVLDL E LLLEIVL L LLLGVRH M DGVG V VQR MLPGR MQEGR MSVD V QGR MECGR MDLGRLG G AGRLGGLDGVRLV G MGVRLCGPQGVRL PVGPPGRLD G YGVRLA G DMLRR PELG 17273 4 5 6 7 8 9 0 1 2 3 40 070707070707 7 8 8 8 8 81 1 1 1 1 1 101010101010101CGT CTAGGCC CGCT CA T T T T T T T GGGC GT GT GGGC GA TGGT TGTAGT CGTGGATTGCGA GA G G G GTGA GTG GTGBGTBGGBGABGBGCBGABGABGBGABGTBGGBGCB TBGA_GC_GT_GT_GG_GC_AT_AT_AA_G_ C_T_ C_ C C_CGC_CyC CA_TTCA_TGCC_ CCCA_ TCCC_A TGC_ CCGA_CGGG_CA CGA_ TA CGA_TA TGA_TA GGC_AA T G A_TT AbT C yAbT C y C yGAbTAb CC yCAb TC yCAbC C yAb CC yCAbC CyAbT C yAb TCC yAbT C y CyAyGAbTAbCbTG G_GG_G_TG_G_GG_T A_A_AA_GA_GA_A_TA_T C_A ATA AGTA ACA AG A ACA ATCAG C AC CATA AGACAT CAGA_CGTA_C A A_CCC A_CGTA_C G A_CGTG_CGTG_CCTG_CGT CG_CCGCG_CTACG_CCCCG_CGTGT_CTA 124PFMGDVAEQNADEEFGDVEVDFLDHLSIEESAESPEPFDFVYETELLGQFFIF L FFF F SHSHDSHSHSHSHDSHASMGDSLE MGGQDS EGFMGSGSMGS SMGS LMMGFSDACE SAGGC EGPALCEDAGIC E EAGWC ENAGE C EGLAE LGCGD YHRPLLDSLDG GDS TDSVDGGDG DG DS LDS DLDS DLSE S LSGME SSGE E S ESGFE SGTSCE SFSSE S SSE S FSDDRD EWENAVWE S EH AAWASE L E EADNA NAPW NAGAIW NADAHSAFW NA AD DIP DH PLDSQ IPGH PVDS LIP TH PSDSIPIPDHGSG EDIPPLHGSV LDIPHHGSPS DIP DGSPLAR FYPKGSSGSWGS VGS GGSDGSSS GD S GD S GAS GGS GSPSATRYT ELT EPT E P T E R T EDT EGR L F R S RWR RGI RGVRGDSRVFLFAD FPFGFNFLFLKIDKLI L KLI P KLI Q KLI GLIP LIS RDDEPS L FANDES FAEDEP FAH DEFAAFAM QDE DDER SDDSL RDL RSDP R S LDNR SDRKSGKSGV NRDERDLLPDTEPLL LCPSA LL GPSSLLFPSLL SEPSLLSPSD SGSGD SGNGLLL RA YRYLLRYSSGSSGSGPSGM K ARYLRHRSRVAAMLGRREEIML LRRECIMLPRRE TMLRQMLRGMLRDFKP VDSKP VCKP VGKP VTL KPY VQSKPY VFPKPY VDL RPQ RA QR EDR E L R EDR EGR EGR E L R E P R E E R E T R EDR TD 5868788898091929394 5 6 7 80 01010109 9 9 9 91 1010101010101010101TGT T GC T T GGGT T GC TCGATCCTTCCCAT TG CTTCGTG CTATTGGA ACG AT AAA AAATA ACAT TCGAABG_CGT BG_CA TB_GCA GB_GCAT BG_C CCB_ CABAT_CGBAA_CGT B_CTCB_CA GB_ CA TB_ CCCBG _GTAB_G G_CA CG A_TA GC_ CCA G A_ TCA G A_CA GGC_ACA_CGCG_CAC A_TAC A_TATCA_ TCACC_ CCACC_ATAC_TA Ay TAyGAyCAyTAy C yTGyGyCGyGGy TGyTGyCGyTyC b bTTbCbCbAbCGbCGbTGbTT GbGGbCGbCGbCGTbAC _GGATCC_CAC C _C_AGCAGC_AC_TT_CGTGT_CCCGT_CGTGT_CC CAC CATA_AA_GA_A_A_GAGGT_CCTGT_CGTG_ CGAGCTG_ TGAGGTG_ CGAGCC G_G AGA_AGA_TATT_AATGTG A G_GCG G G_GGTG G_GGT CA_TTT125SSSPEEQLESLINGLGEEWTCIDEGQLNSLTLPLGLPPPVLSQEPGYGGV AAYHQIYHYA YHSA LYHAA YHHA YHEA PYHEA A YHHYHDPA YHNA YHNSA DA A YHD YHDYHDER P R PDR P R PRR PMR P R PYR PNRPLR PMR PSR P L RPIR PGDSGRH SQPDSGRY SQDSRDSDSRPSTS NSVDRS DDRSEDSRQSISS EIDRADR ESHSESQDRNDRGDRG SESDRF SL DRSSARHGS GPQAR SAR FG ARLL G AR LG GARNCG ARTGSRGSGSE ARKARGARSGSS ARLGSL ARVGSGSARTSARGYQYP AYP VYP HYPYPDYP SYPLYPVYPQYPYPHYPYP RTR LTRTR ETR PTRL TRE TRTRL TRD TRTRGTRS TRAR RSRSRF SRKSRSRF SRT SRT SRL SRS SR ASR QSRPR T R RVRGSVRKVRVVRVE VRVVRVRNRYRLR L R L RVS RWPS R SL D L DQL DEL D A DKDVD YVDSVDLVDMVD VVDPVDPVDGP TAP TKP T P TLP TALTQ LT FLTL LTLTP LT PLTGLTLTQVASRQ VAAVAKSVAETVAVPVAEFPVAYPVAI PVADLPVAS PVAGPVAE PVANS PVASPQS RQY QD QEQFQTK YLDP PA GRRTVPPRRRTAEPRRRTKPPRRRTNLPRRTHRQPRRTKRLPQ RRTVRKPQ RRQRQ TFIPRRT C RGPQ RRTARQPQ RRT RTWPQ RRSTFRPPQ RRTGRLPQ RRTG D 99001020304050607080900 1 20 1 1 11111111 1 11 1 1 111111111111111TTGTGTGT C T A TCT TATTTGTGT G TCTCGCTGGTGTT TGTTGC TGATG GCB TGTATCTGT C TATABGTTAB_ GTAB_ GTGB_GCG ABB CB_GG_ GA_GTA_GABGGBGCBCGC_GA_ GG_G GA GB_G GA TB_G GGT BG _GCT_AA A_C TA A_CGTA A_ TTAT_GTAC_TTTACG _ATAT_ATAC_ GT TA G_C TACA_C TA GG _CTAC_ CCTA A_T TAC_GGyA TGyAGy CGyGGy CGy CGyAyCyCT yAy T yCyGT yATbT_ C Tb TTbGTbC TbTbGTbGGTbTGb GbTGb GbCGb Gb ACA_ CTT_CA_CAT_TC ATT_A_TTCCAGT_G A_T T CATAT TT_ACAGT_ATCCCAAT_ATACCAGTT_GTATT_ TAAT_AAC TT_ TAGT_TTACT_ACATGCCAT T CAT CGCATACCAT T CATCCCA ATA AT CA _ _ _ _ _ G _ _ _ G _A_TG 126SSSPEEQLESLINGLGEEWTCIDEGQLNSLTLPLGLPPPVLSQEPGYGGV AHEAHQA A ARA AASA DA GA A APA AAYPYPQ YHP LYHGYHNYH YHLYH YH YHDYHGYHSYH YHWR E R E RYR P P R PGR PGR PSR PWRPAR PAR PDR P RPDR PVDSGRSSRPDSRSHDSRN SLDSRQ SVDSRSHSPSSYDRS TDRSDRK SP DRN SVSSTDRLDRV S D SESE DR LDRGDRV EARPTG ARVG AAR FGLGDNARGARIIG AR FGSLGSDARHARKG ARNGSL AR TGSPARVGSARNL GSARAGSDAR INY DTRGYPA TRYPYPYP FYPIYPPYPKYPYPATRGTRHTR STRDTRGTRATRATR PT YPD TRFYPTRD E YPRTYPRYSR LSRSR CSRQSR ESR VSRSRASRS SRP SRLR T RGT RKVR SVRAVR FVRVVR IVR RVRMRSRGRARDS R T S RAS RKL DGL DAL D NL DAD V D DGVDEVD FSVDVDHVDMVD AVDLP THP TAT C TLT TLTKLTALTDLTLTALTL LTS LT PLTYVARVVAAPVAPI PVAGP PVAHPVAS PKVAVPVASPVAV SPVAPLPAS PVAE PVAAPADPQR RPQA ARPQLCRQQ QN Q QGR SV I LRVQ DPRR T LPYRQRR TARQGRQT RQVRQE RQT RQGRQRRTET RRTDRRTR V RMPRRTTSPRRTAPRRTDPRRTGPPRRTGPRRTESPRRTTTPRRTDPRRTVL3141516171819102122 3 4 5 6111111111111111112112112112112111TTGGTTTCCTTATCT TGT _1T _3T T T TGTCTG TTATC GBTGT_r T_r TA TBTA TTA CTGTATAGG AB_G GG CBG _GTAB_G GACB_G GCGG T_GAB_G GekG GekG GTG C_GCBG _GCB_G GTBGG A_GAB_G GGCB_T C T T C T T C C TGCT CGTAGTni Tni TGTT TTGCTATCTCA_GA_A_GA_TA_GA_ TAlAlAT TATAT TA GCA ACA GCGyAGyGTGyCyGyAy T_y_y_yG_y C_yG_yA_yG _yATbT TbTb GbGGbGGbAGb Gb GbGGb C Gb GGb T GbTGbGT_AT_AT_ C TT_GT _ T _ T _ T _ T _GT _ T T _ C T _ C T _GT _CAC CAT CACT CAATC AC TAATATATATAT TATATATAAA_TA A_TTC A_TG A_TG A_T CACA_TAT CA_T B CA_T B CA_TA GCA_T TACA_TA GCA_TG ACA_TGCCA_T CG 127SSSPEEQLESLINGLGEEWTCIDEGQLNSLTLPLGLPPPVLSQEPGYGGV AHAAHAPAHEAHDA VA AA A GA A APA AYPYPAYHIYVYDYH YH YHNYHAYHHYHAYHIIYHKYHPRARVR P I R P R P R PDS R PNR P E RP T R P S R PAR PKRPYR P TDSGRDDSRADSRVDSRDLDSRDS L SLS F S E S A SLS SNS QS E SKTS SGDR LDRS NDRSSDRSLDRSEDRSADRRDRDR PAR LGYP SARPVGSAREG PSARDGLGSGRG GG PYAR LARDARNARLARAG ARPT G ARGGST AR LGSAAAR LGSARQSITR EKYRTRSY V TRDYLTR EYP TTR LYPLTRLYPTR LYPLYPQYPVYPYP AYP NYPATRGTRKTRTRATR FTR CTR SSRT SRL SRSRSRSRC SRSRI SRII SR ASRASR ASRSRAVRGVRMVRA VRYLVRDFVRGVRKIVRGRER P RAR E RHRALPDTDLPDT LTLPD ETELPDTSPLPDTDLPDTEILPDTVL DT RV LDT SV LD ATPV LDTQV SL DV TNL D AV TCL D ITSVAK ARAMAV ADAH ADPN APAGRPQSLVQS VQIVQYVQLVQDVQAVH QVQE PVADPAL PALS PADI PAHQLPVQGV V YVTPRRTSRRPI RRRTPLPD RRTPRCPRRT L RDPRRA TERSPRRVRTMPRR DTI RYPRRQ TSREPRR QTS RVPRRTQRPPRRT L RGPQ RRQ TTRRPQ RRT L RVPQ RRT TN 72829203132333435363738 9 01 1 1 11111113 3 41 1 1 111111111111111TTGT G T C TTTGTTTTTTTATAT TTT T C T GGTTGC TGC TGCTGATGCTGAB TG GTGTCTGTA AT C T TGTTC B_GA AB_ GTTB_GTAB_ GCBT_GTAB_CTGA_ABGABGC BGCGTBCGG_GA_GG_ GT B GT_G GGBT_G GGBT_CAT_ATAC_TCTA G_ GT TAT_TGTAC_ AT TAC_GTAT_ ATA A_ TCTAT_ GCTA A_TCTAC__ATATG TTACGCTA AAGyGTb T GyT_ T TbCGyT_GTTb CT_ CGyGTbTGyT_GTb AT_ TGyATb CT_ TGy C yGTb_GGGTb T_CGyTbA _GyATbC y C_y T_yG_y_AGTb_GGTb_AGb Gb GGT _TT T _ CCATT CACAC CACT CACAATATAGTACTAATAGTATATAAA_TG A_TAC A_TA A_TA A_TGC A_TGCA_TA GCA_T CGCA_TTT CA_TAT CA_T TACA_TGCCA_TG GCA_T CG 128SSSPEEQLESLINGLGEEWTCIDEGQLNSLTLPLGLPPPVLSQEPG YGG GG GG GGGGG GGEGGVGG GG GGGGG GGLGGDGR GSGRMGRGGRLGRNGRIGR DGREGRLGRLGR AGRLGRVVRSVRWVRV VRGVRAVRH VR CVRQL VRSSVRMVR EVRAVRQAGRAGTAGV AGEAGRAGD AGKAGQAGDAGDAGPIAGFAGLYHAYHKIYHDSYHEYHKYHVYHLYHR YHLYHLYHDHLSHSR P R R PFR PVR PWR P R R PMR PSR P E R PDR PDR PDYRPQYP SDSRQDSRRDSR DSRTDSRQDSRPDSRES Y SPSFSASDRS RSGSTARQGSD RKGS ARLSGSRCIGSV RRGS LRC LGS EDRR LIGSSDRRQGS LDRR SGSDDRRDGSRMDREGSD RVGRSDRNRYPTR LASR AYPPTRAA E YPRPA DYP DA REYPRA RRYPRDDA DPGYR EA KYPRRA IPMAPLAP EAPE YRDYRAYR LYRLA P YPRGTVRASRRSITSRSTRDSRGTRQSRPTRQSRRLS LT R I T R T RPT R D T RAT RP T RDS R SL ES RPS S RVS R S S RVS RP S RWPDTAV TLPDVV TYLPDTLV SL DLV TNL DEV TE L DG TGPV LL DT EV LDTAV LDTV TVL DTGV LL DTVV EL DTLSV LDRTAVALAKRPQAVQF VAIP PAS PAE PAGE PAVPALSPAS PAMPAGPALPAM QP VQL VQE VQLQVQVPVQEVQP VQD VQVQLP VQLRRTARGPL RRRTPSPI RRRTEFPT RRRTLPPE RRRTEEPRRTDGRG VPRRRTDPPRRTKRNPRRSTFRLPRRT L RDPRRH TTRLPRRT P RQPRRT TD 1424344454647484940 1 2 31 1 1 111115 5 5 51 1 1 111111111111111TTCT T T T _ TVT TGTCAG2_TA TATATCTCTG GC TTTGABGGBGT BGT T TXT TTC TAT TAT TBGreGCGGTBGABGACBGTCBGGBGGTBGT BTC _AT CGG_GG_ GC_ Gkn GIP GC_GT _GA_GC_GT _AGA_GA_GT TTTAT TAT T i l T E T TGCTGAT TAT CATAGT CCTGCG_yA_TbGGy TA_bGGyTA_bCGyTA_bTGyA_A_A_AA_ CA_ TA_b Gyb GybAGybGGybTGybCGy CA G_y CAA G_yTTT_ATCATT_ATACT_TTACT_GTAGT_ TAT_BTA_7 T_ATAC T_CTACT_AATTT_TTbATT_ATbAAT_C Tb GAAT_ATA_TGCCA_TCCCA_TG GCA_T TACA_T B CA_4T.1 CA_T CACA_TAT CA_TTT CA_TGT CA_TA ACA_TAT CA_TAC 129SSSPEEQLESLINGLGEEWTCIDEGQLNSLTLPLGLPPPVLSQEPGYGGV AYHLAHSAHDALA A MA HA ARA AEA A ANR PILYRP SS YRP FYHRPD YHW RPQYHRPL YHRPAYHA RPGYHRPEEYHNRP IVYHTRP EYHWRP CYHNRP CYHP EDSGRDDSRPDSRDLDSREDS K SESESLS R S SKSDS ERSW SSGS S FRSEDRS DI DRDI DRNDR LDRWDRVDRPDR SDRVARDFG ARSGYP SYPPL ARQGTG AARDARAGIGSIGSYPGYPRYPTARDARFYP GYPIARKGSARQGSARQGSYPVYPAYP DAR FGSLGSF YPKARYP VARYGSL GAR LYP IYPQTR TTR STRTRD TRDTR STRATRATRATRTRE TRTR S RNSR DSRSR DSRSR RSR ISRNSR ESRE SRVSRLRNR P T RSVR CVRASVR SVR LVR SVR S RNR I RKRDRNIS RDS R I S RNLPDTMLPDTVL D LTLL DTMEL DTH HL DGV TAL DTYV LDTMV LL DTDV EL DEV TVL DT SV LD ITEV LDK TIV RL DN TKVALR SVAAPVAPEPVAVPVAI PVAS PVARSPVAL PVAEPVAC PVAGLPVAKPVAAPATPQDRQGQ QSQLIQSI LLEDKLVYRRTELPS RHRRTTSPRRT S RHPRRCTIRCPRRTNRGPRRTDRTPQ RRTGRNPQ RRETIRLPQ RRTNRDPQ RRTQRAPQ RRN TCRSPQ RRSTSRLPQ RRTYRGPQ RRTN K 4555657585950 1 2 3 4 5 6 71 1 1 1 1 16161616161616 61 1 1 1 1 1 1 1 1 1 1 11111TTGTTCTTGTT 1_ TTGTTTTTG TTG TTG TTGTTATTG TGTAGC BGAGAPT GC T TA AT TCTCGTTA_GCGB_GCCBG _GO G GCBG A_GTGBG _GTCBG _GG ABG _GABGG AGTGATAATTS TGGTGTTACTGTTG _TCGTGB_G GABGABG GGTT _CGTG_TGCTGB_G AGCTT B_A A_AA_ CA_A_A_ GA_TA_A_TA_ CA_AT_GAC_AA A ACGyTb CCGyTbCTGyATbTGyTb Gy CTbGGyTbAGy T yTbGGTb Ay C yCTGb GGbTGybGyA_yGGGbGGbTG_yGbTT_AT_ CT_T T_T_T_GT_T T_CT _ C T _GT _ T _ T _TT _ACAGCAGCATCACAG C AC CATCATATATTAATAGTACTAA A_TA A_TA A_TATA_TB_A_TGTA_TTC A_TGTA_T CACA_TG ACA_TAT CA_TGT CA_T CGCA_TGT CA_TTT130SSSPEEQLESLINGLGEEWTCIDEGQLNSLTLPLGLPPPVLSQEPGYGGV AAYHL YHA S A YHKA YHLA YHMFA YHDTA YHIA A A YHAYHGYHEA CYHPA KYHKIA YHSL ATYHIR P S R P F R PWR PES R P R PIR PVR PFRP E R P S R P R PTP L PVDSGRFLDSQGRDSGR IDSRT S DPSPSDSLS P S SVSE RS LRSM SS WSS D SPDRSNDRADRADRFDR CDRKDRRDRGDR EDRKAR EGPG GS TGSIGSPAARGARGAR PARAAR TARKAR EGSKAGSCGSTGSARAARGARVAR EGSSARDGSF ARGY VYP TYP PYPYPFYP SYPSYPLYPEYPHYP GYPLYP P STRTR PTR PTRATRIETR LTRATRGTRTRTR PTRTRYYR SSRK RNSRRASRRTSR SRT SRRKSRRRSRRE SRNSRAFSR GP SRGSRK VSR LTTSR FIV VTVSV VIVRER R RLR R R RLL D STEL DSTVL DTDPL DV TR L DT TEL DTPV RL DIV TAL DGV TVL DLV TAL DT PV PL DVV TE L DPV TDL D RV TEL DT LPV EP P P P P P P PR FV V V V V V V V PP PVPAE PV WPVSS PV A AH AG ALALAQ AEAV A ACA A AVPQTQY Q QQQD QG QEIVD V VIL VRRT S RAPRGRRTATPRRT F RYPRRPTSRPPRRTKRDPRRTARSPRRTKRDPQ RRTHRVPQ RRD TRRPPQ RRT L RAPQ RRTGRAPQ RRY TGRFPQ RRT L RYPQ RR RTPD 8696071727374757677 8 9 0 11 11111117 7 7 8 81 1111111111111111111TTTTTTTT TG TTATTA TTT TTCTTCTTATTCTT TC TTATTGTTAGG GCGTGC BGC TAA CATGT TGGTT B_GTBA_GTTBT _GBGT B_GAGA_ GC _G GGCB_G GCGB_G GCGB_G GACB_G GCG TBGACBG _GCBG C_GT BA_CAC_ TT TACG _ATA A_GTA AA_C TA A_TATA AA_C TA A_CGTA GG _ATA AC_C TA AAT TA A_CTA GTATG ACGTATGGyGGy TGyC GyGGy T yAyCyGyA_yA _yT _yA_yG_y TTb TTb CTbC TbATbAGTb AGTb Gb Gb GbCGbTGb Gb GbGT_CATGT_CACT_CAA GT_CAG_ _CTC AGCTATT_ C TAAT_CTATT_CAT TT_ TAGT_AGTT_GTACT_CTATT_AA __ AA_TA ATA ATACG ATACAT T CAT T CG ATGCAT TGCATACCAT T CATCCCA ATA AT C_ _ _ _ C _ G _ _ _ _ T _A_TA 131SSSPEEQLESLINGLGEEWTCIDEGQLNSLTLPLGLPPPVLSQEPGYGGV ACALA YA A AATACA AQAPA HA VA ADYH YHKYHVYHSYH YHLYH YHRIYH YHTYHAYH YHSYHER P LIR P R PTR PGR PKR PAR P RFR P RP S R P R PNR PNPGPADSGR EDSREEDSISLS D SLSLSVS LSS DTSVSDRS ARSY SS S LRS NDRSMDRSFDRADRDRMDRDRDRDRLDR SDRLAR FG MAR LGIG G KGSARAARDARVAR RGS QGSAR TARYGSHGSGGSQGSTSARLARWARSIARIGSARSIGSAR IHYPTR ESR PYP QTR RYPTRMIYPLTRDYPTR SSYP ETR TYPQ TREYPRPPYPRLSYPRAYPRIYYPRRWYPRDTYPRSCR T SRRKSRRGSRFR SRR SR AESRETSRQTSRPTSSR EPTSRDTSRVTSRATSSRVV VGVSV DVRVR RAR R R R R R RLPDTMLL DT SL DTTEL DTDLL DV TNL DTLV LL DTVV LD QTTV LDLV TAL D FTEV LDVV TC L D AV TSL DTV TQL DD TMVAKPARVAPKVAKPVAAPVAVPVAMPVAEDPVAYPFVAGPVATPRVAE PAVPVAQPALPQENRQ QHQ QFQ QLS P KV QI LVNRRTMPMRRRTRIPRRT L RYPRR DTS RGPRRLTLRSPRR DTF RAPRRTGRSPQ RRN TLRFPQ RRFTLRSPQ RRT T RAPQ RRTTSRLPQ RRTVREPQ RRTARQPQ RRTD N 283848586878889 0 1 2 3 4 51 1 1 1 1 1 181919191919191 1 1 1 1 1 1 1 1 1 1 1 111TTTT TTTTTC ATCTCT T T T TCT T T TTT T2GG GCGTATAT C T C TGTACTT GTA ATCTC AT_PGTT B_ GCABG CBGA_GAT B_G GG ABG _GG GBG _GA AB_G GGCB_G GCTB_G GCBC_G GTBG C_GCGB_G GCT BG _GOTAC T C_TATTAC TGC TATTAGTAATAGT CTTGT C T T TTA_AA_TA_TA_ GA_ TA_CA_CA_TATACACAA GGA G ASGyAGyAGyAGy CGy C yGyGyG_y_yG_yA _yG_yTC_yTbTbATbC TbTb T GTb GTb GTbAGTb GGbGGbCGb Gb GbT_ TT_GT_GT_AT_ CT_C T_AT_ _ T T _GT _ T _GT _CT _CACT CAGCAACAC CAGCACAG AATAC TATAGTACTATTA A_TA A_TA A_TA A_TCTA_TA A_TCT CA_TTT CA_T TGCA_T TACA_TG ACA_T TACA_TTT CA_TCCCA_TB_132SSSPEEQLESLINGLGEEWTCIDEGQLNSLTLPLGLPPPVLSQEPGYGGV AYHEAHTDP DDPSDEDDD D DL E L PDI GR PSSYRPSSD GAE PL D DPDPAEGAEWAE LGDPD AEIDP NDPGQ EAE EAED L F FG LQ FGG SLQ FLGQ EL FEGQ FL FLDGRQ SDIDR P EPL EQEPL LMEPL TCEPL SP P F PF PVE LFL E L S E LDIS S PRAGISRAVPISQPHRA ISRAL PTISM RADARYPTGSTARQPLQFYLQ GFYLDQFIQFQFL YLDYLHYLTQFGQFS YLLYLDSLDL LLDSLSSGSPDLVDLS S LADLTR LYASGVSGSG TR T LE DLE DLEESGSGLEPSLG EASLG ELSG GLEAV GSGGIV GSV GVSPGGDSV GSV GWGSDGFSR QSRVMVS SVS FVSVSVVSWVS IVSD GAGGA GAL G P AD RS RLP LALLADLAQLAP LAPLAGLAS P S R P SGP S L PAPG DPDHV TGL DTSPPTA LLPTA LDLPTALLNPTAG LE PTA LPPTA LRPTAGLLDSVAPASIDLNDSLEPDSSLDDSLNP SSDSLLARPQMLVRQRMELL ISL MELAMI S I P INS IN EL L ML S MLAMLHMILMQ H DPGQQ SPGSFQ PPGLL Q CPGAQ PGDSRRTGSPRRTMETREP CGTSRE DPSGTSRETPLEPTSREFPPETTSREPGELTSREQ PSEETSREP L FDARSSQE FQARSSQT FQARSSQGFEARSSG QLFCARSSG Q G 69798999001020304 5 6 7 8 91111111121212120120120120120120121T CT TTTTGT TAGCGG GG GCGT GGA C GT C G GGC GT CC CC A GA GAGAGTAT T T C TAT TTA ACA A AGTGTC B GA TB_GG CAB_G CAT B_G CTCB_G CA TB_G CGT B_GA CGB_G CCCB_GABCG_GA TB_GG AB_ GGT B_ GAT B_ATG__yTTCAT_CyATTAG_CCTA A_CGTA A_TTTAC_ CCTA_T TA_ TCT C_AA A_ TCCAC_ CCCA G_C CCA A_T CA A_CGTbAGTbAT ybT T ybC T ybT yGTb C AyGAyC TbTT ATTbC T yTbCAyTC bCAyCb CCAyCbTAyGCbTAyTC bCT_AT_C CT _GC_AC_C_C_C_GC_TT _GT _ T _GT _ T _ACAC CACAT TAC TAGT TAGTAC TAC TATAA_ CAGATACACA_TA A_TGAT_TGTAT_TCTAT_TAAT_TGTAT_TCCAT_TGAT_TGTA A A G A_AGTA A_AGTA A_ACCA A_ACT133G DAEEVPGGPGAEQFLESTPPPASTVRLQPSPEWSEMQEQSGRGMWSRGLQPF WGQFT F T W TETLT T G TDQMDQMQMSQMEQMIS TL FDAP SLSAS TASCPISDDLGPDLLCI PDLLQAGP SDLLMAP SFLLAS SASDDLTPLS DLVPL D AHDL L S E PDSL AES E LDNAES E SD GASL S EDEAWS E EDDIRAI RALA L A D AVAWLAW AW AAPSQPSF PSMPS S TPS S EDSVL DE DSLA DEWCS L EWDGICS E EWDGCSDS ECSDEFCSA WECSSP EWCSDSGQLGGQLSGGQLD QL GLCIGQLFLTGSGGV GSGS SEL GSEL QL SD EL LL SD ELDSD EL PP SDVE LP SD EL GKVVKV L KQDKQLLKV KV KQDKQDEKV KQSGAQ GAGS E R S E S E S EDS E S EGS EGYSGYGGYFGYGGYAP S L P S LVDNVDN SVDD VDLVDNSVDE VDMRL LR I RDR RWDSNQLSDSLM DLAHP EEQLAEE LTLALEELLAA EEDLA APADEEALEESF LEE LKF LLD DKFG LRKFD LL KFQLKLFPPFGLT Q PFGL ECRSE ECRLPECCRGECRSGEC GRL ECRPECRDG FPIYLGDL PIYNGDPIYAGDPINGDYS PIYNSARS S LRQPLAS SDQFA QA DLSYR DLSYRLAEGLSYRIA HLSYRGA LLSYCRIA QLSYTRQA SLSYRDTDPS RAC TGPS RH AQTSPS RDAS TGPS R LATTLPS RA A G 0111213141516171819102122232 2 2 2 2 212121221 1 1 1 1 12121212121C CGGGA AT AGAT AC AC AGAATTTTTCT G TCATATAG A AA A A A ATATGTAT TAGTCBGCBGAB_GTT A A BGGBGAB CT_GGBGAB_GCBAGBAABAAB_ATTBAGBC_CC_GG GC_GA_GT CGT_GTGC_GA_GG_GTCGC_GT_A A_TT AC_AA T A_ TCA T A_TT A T G_CCA T A_ GA T A_TA TC_ CCA TC_AGC G_CCGC A_ TCGC A_ GGC A_TTGC A_TAy TAyTT yTT y T T y T T y C T yGT T yCT yTy T yTy C y T yGC bT _GC bCT _TTC bCC _GC bG C_C bC_GTC bC_AC bC_TCC bC_ CCbCA _T Ab_GA Ab_ CAbAbGA_AA_GA Ab_ TTAA_GTAA_ACAGTA ACA AGCAT TA_ T TA_ C TAC TAGTACA A A A AGTCC_TGCC_TACC_TGTCC_TCTCC_TCCCC_TGTCC_TGTTTGGTTTG GTT _GCTTT _GTATT _GCC 134ALKLTQIKSLKDSPLTPIIKRLAAFAENLSQTRQRERVNALIRQSQLEEFSQMDL L L LDL L L P LKWKWTKWDKWKWDAEDQMLS LAED MIP DR TWTMIER TF MI GRT LMI IRT EMIRTGSMIRTL MITFDAI LV DAI L SAAI L EAI L LGIAI L LDS SGSDFQSPS SQDNW DLNCINWSWM WFW TDLNGNTLLNTD NLLNLNTTS LNVNWE RWDTHLNQNTGLNLQALWI SLQAAL LWIWQAAPLWIG QQAAL LWIGQAAR LWI FD AAGKLVG VHKLVLA AYYI EA GYYLA A IGYYIDYYIAASYYI PA YYVATIDYYIDNRRDL NRRPNNRRN S NRRN HNRRDLKQSKQDDDDIDDFDWDDVDSDDSKKLKKSKKLKKQKKAGYPGYDS IQIGIDIDPI P IDLIGAP CAPAAPTAPSAPDRVRYKLYKRYKDYKPYKGYKLYKLQWGQWGQWLQWEQWSKLF PLGD GKFGLALNALNAL LALAHSAHHAHAAHNSALEALAHPAHDALLM AHDLVVEI L LVVCLPVVL LQ VVDLG VVGPI EGDTYP PIYM DHKLLP I THKLIQSHKLI DSHKLA I HKSHKGL FPLLCHKL LGGRH DGGIRQGGRG L GGREGGRLDS R S TAFPPS RALQ DRS F LQ VPLRS FVEQ QRS FGQ V GRS FVLQ CR ISFVTQ QR ISFVGQ ERISFDVFG DSAVG S W MSAES WIG DSAPSWPG PSACS WPG QSA S WGL4252627282920313233343536 72 2 2 2 21212123 31 1 1 1 1212121212121TTGCTTATCG TCTTCTCTGC C CG TCTTACA TT TC T GA G GGTT TC GG GA ATGTG GA GA GCG GT T TATAGA TB_GCCB_T TBT C_ TABTG_T ABT T_ T GBT T_ TA TB_TG AB_ TCCB_ TG AB_TGT BT_TTCBT_TA GB_TTAT B_GCC GCA ATATACAT T CCTGC T CAG GCG ATG ATG ATG ACC_C C_A_ TA_CA_ GA_A_CA_ CA_ A_ C A_ A_ T A_CA_ GAyb CyTCAbCACybTGACyb TCACybCACyGbTACyb CCACybTACyTbC T yGbT T yGGbT T y TGbGT yTGbC T ybCAT _T AT _GAT _TAGGTTT _ TG_G_GGTGAGC_T TAGAGC_T CG_AG_TGGACC_TCTGACG_AGG_GATG_TATG_GATG_TACG_AG_C_TCCGC_TGTGC_TGTGC_TGTTC_TGTTC_TCCTC_T TG AGG_AATC_T CG ACGTC_TCT135AAFYWAADVPNYGNLSFLQAVKNSEEFTSANKTLYGLLPHALVMERKWHKW HMALHMI HMHMHMWHMHM GFGF F F FI LSQPAI LA DPDKGPKE PKL PKFS PKT PKE PKFHDHEGLGEGHD YSDYFDYMDYGDYCDYQDYDNSNENHNNHDNDLWIVQ WAAPNRGL I SAGMYV HMYLT MYDLMYLMYIDMYG VMYDLDSAILGLDSAILWTDSAIL EF DSAI ILEDSAIL LG REARLMK MKSMK MKLMKEMK MKAPIF PIF PIFSPIF F PIFSPNRMIPFSVP IPFAI FDIVPVFPFGVI IPFGIVQ PFDVS ILPFVD GSM ND GSCIND GSG NL GS LNT GSNVKKSKKDLVLVWLVDLVGLVL L L S LIL LIE LIL LIS LIHAP FAP LYVPYPPYDYRY YVLYVGQLD QLGQLGQLAQLSQWP Q DMNGMN MNLMNNMNNMNDMNL LAF LALAI LAWLAPL T LWFQ C E Q C N CACHC SLCLCMNPDNP QLNP GNP PNP VVVQVVDP P P SQPD QPQQPQPL QPDKSDKSKS RKS PKS PGGGR SGQGRDRVSRVARV LVQFP VQGVQS RV GVQSRVTRVC RV EVQLP VQGVQLYV VDLYVN DSYV DNYV VG DNSYDESASGGAAFV W ASSW DTSMT FVL FVFV LSQFV QTSMCITSMGTM DTSML FGTVES MI FHTVDS MFDVSHAV VDSVSHLVTV HV V LVSH VQSVSHA V GVSHPVSF83930414243 4 5 6 7 8 9 0 124 4 4 4 4 4 4 5 5121212121212121212121212121TGGT T T T T T T TTCGA TACGCAC C T CAACAACGACGT TAA AC A CAACGAT AC AGT ATACAACTACGACAACCGT B_TT CCB_GABGGBGABGABGTBGGBGCB ABTBAB GBABAT_AT_AT_AG_AC_A_ C_GT_ GC_ GG_GT_ GT_ACT_ CyCG AC_A TAC_ CCAA_TAA_CGAA_ TCAA_TA TAG_CA CAC_ATA A_CGTA A_TTTA A_ TCTA A_T TAC_ CC Gb CCT y C yGbCAb CC yGC yCAbTT AbC CAyb TCCAybT C yGAbT CAyTbCAybCAybTAyGb TyGyCAbTT Ab CC G_TAGG_TATA_GA_C_TGTTC_TGTA A A ACA_AA_A ACA AGA_TGTA_TCC A_TCTA_T CA_A_GA AGG A_T TA ATA_T A_AA_A_A AT TAA A_TGTA_TGTCC_ C TAGTAGAT _ACAT _AG GCTCC_GTACC_GCGCC_GCCCC_GGT136GEPTPEEIDLTILSLTHANVQSIIYDVCEKTSLDTQVSNGSWLPKFEQV NGF LGF I E I E I E I E I EVI E I E I E I E I E I E I EPNHDNHDDLVP GLLNV PGSLEV PGA LAVPPGKV PGDSV PGKL V PGET L VLPGAVLP GRE VLPGDVPP GSL VIPGDSAIPPI LL DFE SAILDF SPDSSSV PDDSPDL S LV PDRS LPDVS LPDSE S LPDAS LA PDDS LPDE S L PPDLS LPDL S LPDD AGSQPIFSDLGE LGL LGALGT LGALGELGL LGGLGGLDGLGGLVGL LGL LGLAGLELLGLRLLGELQLGD LELGLMLNQ IGGVLNDILGTDPLLGDPY LGDTPAGDPGPGDPSPGDPIEGDP REGDPSG EGDQG PAGDPGG VGDFG PKGDEPEL LQLA ADLAD P LEGTPLEE TP LEATAP LTEGP LEDS TP LK EITSP LET TAP LEKTP LEATEP LTED P LYTEVP LLEANPSNPS LQQLQYLQLQVLQDLQLQE LQT LQLQS LQDLQVKSLLKSGCGLCGLCQ GS CGE CGLCGEEC L CGCKC L C E C VDLDP LSDVLDVLDVEVLVLVLG VLG VLG VDGLG GYV YVLVVP LSIVLDE LVDLVP LVEVD DM VGID P D D D KD E D DKDGVSHLV SHTDGTDTD VLCVVDLPTSLPTPTSLYLPTSL L TGPD TSLGTAPD TSPLITEPD TS VLP TDPD TSM LDTFPD TSSLLTSPD TSL E TNPD TSLLLTCPD TSLATAPD TSLHT25354555657 8 9 0 1 2 3 4 52 2 2 2 25252526262626 6 61 1 1 1 1 1 1 1 1 1 1212121AT AAAAA C CG C CT C CT C CCC CA C CT C CTC CG C CGC CT CG CCC CTCG CC GTT GCTTACT CABGG AB_ GCCBA _GGB_GTAB_CGCTB_ CGCT B_CGTGB_CGTBC_ CG GGBC_GTBC_GA GB_CG GAB_CGTAB_CGGT _TGC T CGGGGT TGCAGACGATGTGGAT GCTAGT CGCGTGA A_yCA_AT_ CT_T_C T_T_ TT_C T_C T_T_ CTG_T C_TA_ GAbyTbC y yGy yTy y y yGy yCAATAGTT _GTAT _TTCb_AGTC Cb_GGCCbG _GGTT Cb_GGCCb_T GACC Cb_AGCCbG _GCCb_TTGCb_GGbTGybAGybCC C _GA TC_AC _AC AC _AGGTCC_A A A A GGTTC_CACTC_CTA A A AAA AGA A ATC_CATTC_CCTTC_C GTC_CCA ACA ATATC_CCTTC_CTA AGA A A ATA AA GTC_C ATC_CGTTC_CTTTC_C A 137LDLLCGEIHDVMPLCDDLGSVHSPVPGEPSFPTQSQGAPPALLGQTEHPFIE I E I E I E I E I E I E I E I E I E I I ISIAVLVGEV YVLVGLVGGVPSVEVTVSVEG VEGLVEGRVEGEP GS LNPL EP GGP GHPKPNP GLP GPIP GWP GAP GAP Q P APLPDEFSPDWT S LA PDNS LPDSAS LPDEE S LPDLS S LPDDS LPDIVS LPDVS LPDGL S L TPDES L IPDHS LPDR S LPDYLGS LGLGT LGE LGLLGFLGSGT LGVLGLGL LGQLGQLGNGLGGLCIGLNGLPGL LGL LGLGLGLELEI LNLAL P L T LLGD L GD D GDLGDTD Q DQDFDSGDG NDKGDQGDQGD QGDFT P L T P E PAPVGP RGPAGPVGPVGPGPVGPKGPQ GP LGPNP LL EGIP LEGTPLES TP LEATCQPP LTEKP LTEVP LE TEKP LEDTP LYTEKP LEATEP LI EI TLE PE L TL ATLG GPE PPECLGLQQLQGLQLQGLQKLQLQALQKLQI LQ GRCGLCGFSCGAPCGSCGNCGVECGE CGLC MCS LCQS LCQALQFAC NDLTD SLDLDLDVLDVELDVLVE LVLG VLG VGG VAG G VN VN VV VD ASKYL E L S LVTLVCPDHTDL TDS LKSD MIDDD L D Q D DLD PITSLQSPTSLTLPTSL T TGPD TSL P TQPD TSLMTRPD TSELFTTPD TSLDTKPD TSLDTPPD TSLQTVPD TSLLETIPD TSL S TVPD TS QLS TVPD TSLATAPD TSLLC 66768696071 2 3 4 5 6 7 8 92 2 2 2 27272727272727 7 71 1 1 1 1 1 1 1 1 1 1212121CCT C CGC CA C CA C CT C CT C CG C CCC CG C CGC CA C CG C CC CACG AGC T CTCCCCGTCATGCGC TGB_GTCB_GTB_GCGB_GCCB_CGG TBC_GTGBC_GTTB_ CGGCB_ CG GABC_GA AB_ CGTAB_CGA CB_CGTAB_GTA_ TCGTA_TTGCTTG_TGTA_TCGT C_TAGT C_ TTGTA_ T GTG_ GTGTGC_CGTG_ TTGT T_ GGTACGCT TGGT C CGyT Gy TGyGGy CGyAGyGGy CGyC GyAy y C_yA_yG_yGCbC CbGCbGCbACbGCb T bGbCbGGb ATGbAGb T Gb Gb CA_TAG_C_CCA AG_GTC_CTA AG_A_G_AA A A A A AT C _TC _GA ATA AGC_ATC_C GTC_CATTC_CATTC_C ATC_CTCTC_CCA AAC _ATC_CCA ACC _AC _ CGTC_CCA ACA ACC_AATC_CTTTC_CTA AT CA_ CACATC_CGCTC_CTG 138LDLLCGEIHDVMPLCDDLGSVHSPVPGEPSFPTQSQGAPPALLGQTEHPFIVEGDIVE I IVE R IVEGPSIVEPP IVEGIVEGIVE L I I I IPI IQVEKVE FIVE TVE TVE FVENPDS LP GAP GKPLP GKP GAP GVP G P GKP GLP GE P GP P GGP GWPDWS LK PDDS LPDRQS LPDSP S LPDYS LPDDE S LPDDS S LRPDE S LPDWI S LPDDE S L EPDSS LPDT S LDDS LDCLGKLGF LGLGSLGNLGLGLLGYLGLGFLGRGQ P GVP GDGL PGLKGLV GL LGLAGLG GL LGL S LDL T LPL L P L L E L LPGTDKD V D RRD H DLDAD D D QGDEG G GD D DTGD QSGDVGDLPKGPP LT SGPGP PGPNGPD GP LGP RGPGPGPDGPI GPDGPVL EAP LESTP L RTE PP LTEGP LTE CP LT TEGP LTEL P L ITE EP L PTE PP LR TED P LEGTP LS TEAP LEFTLP LLNCQALQR LQQLQMLQHLQALQC LQPLTL L LLS LALDL EDLGSECV GCGE CGGC AC ACGCSCQSCQMCQ Q Q Q GC I C HC IDTVDLN DVVLDVEELDVALGC LG VDVD DVP LGEDVI LGADVL LGDDVP LG DVE LGLGS LG VDVHDVHDVL LGESDVKPDSSR TPDSFLTPDSEETPDSGTPDSIYTPDASR TPDH SDTPDS SE TPDSGTPDSS TPDSV RTPDSTPTPDI TDKT LGT L L T L E T LAT L L T LGT LVT LKT LGF T LCI T L E T L T TSLEEPTSLSS08128283848586878889809192939121212121212121212121212121C_CC C_C_C 1C3_ C C2_ C1_ C CC CC C CT C CAC CGC C_C CGC CGC CA C CGC rGek CA G GB C rGek C rGek C CGGB_ CA G GB_ CCGT B_C CGAB_ CTGT B_ C PGOC TGGB_C TGGB_CCGCB_ C TAB_Gni GA_CGni Gni GT GGAC GATGTAGTGT AT T CGCGGTT l TGT l T l T C C TAC T CGTGATAC GT SGCGGAAGT TGCAG_y_Cb GyT _bCTGy_b Gy_b Gy_bGGy_b GyTGb C _y_yTGbGGbCG_y T_b Gy T_b AyTGb GT_yGT_yA CGb GbGA_ C _TA A AC C _ C _ C _TC_CBTC_CGA A A A A AT C _GC _GC _CC_AC _ C _GA A A AAA ACA AGA A A AAC _ C_GATC_CBTC_CBTC_C GTC_CGCTC_C GTC_CATTC_CATTC_CB_TC_CCA AAATC_CCA AC C _AA AGGTC_C GTC_CCG 139LDLLCGEIHDVMPLCDDLGSVHSPVPGEPSFPTQSQGAPPALLGQTEHPFIE I E I E I E I E I E I E I E I E I E I E I I IVVP GFS L EV A CP GSLGVSPP GSLYV DP GES LDV AP GDRV SLNP GPK SL IV TP GTK SL IV D FP GSLCV LP GG AVSS LNP GHV N SLMP GHVES LAP GES LDVEIP GPCVES L RP GDSL TIPDSPDQPDYPDPDGPDEPDPDI PDLPDPDEPDEPDF PDPLGKLGLGLGVLGYLGLGR LGE LGNLGT LGDLGFLGL LGAGL CGLVL GLQ GLLGLGLGE GLDGLSGL R LDLII L L LQLTGTDPGDGD HP LHGT PGP SGDTPPPGD DPIIDF GPSLGDK PAGDFPMGDPNG LGDLG PGGDPFG IGDPTG SGDTG PQGDP TSL EGP LHTEQP LEAF TPLETTP LES TP LTEKP LE TES P LEEPTPLTEAP LELTFP LEATNP LATEP LEETEP LELRCQPLLLCQVLCQKLCQPALCQEI LCQVP LCQIVLCQT LCQKI LQVLQNLQWPLQALQRG GA GQ G GVG GM CKC CPC VC PDP L L L L L L LGLGV G GYG G GTVPPDCDTVGDPDVKAADVPDVTD Y HDVWDVKDVL LKDVDDVALDVR LSDVNLSDVE L RDDVQTSLDLPTSLQTVPD TSLYTAPD TS LLS TVPD TSLNTYPD TSLYTGPD TSFLLTPPD TSL E TNPD TSA LDTIPD TS VLF THPD TSL L TGPD TSLATGPD TSL L TGPD TSLG A 49596979899 0 1 2 3 4 5 6 72 2 2 2 29203030303030 0 01 1 1 1 1 1 1 1 1 1 1313131CCC C CC C CG C CA C CC C CA C CCC CT C CT C CAC CGC CC C CTCTCACACGCTCGGCGAC CACCTGACB_GACB_GTAB_GCTB_GCBT_ CGACB_CGGBC_GTT B_CGC BA_CGG GB C_GTCBC_GGT B CA _GAB_ CGTCB_GTAG_ AyTGATGCG_y TGGTA G_CyGGAT TG_ T GyC T CG G_yGGATG G_TGGy AT TG_ TyTGGTA G_CGy AT T CG_yAGCT C_TGyTCTA_ C GyT TA_TGyGTA_ GyCGGTAA_y CACbC CbGCb TCb CCbGCbA Cb ACb A bGGb GbGGbTGb Gb AA_TAG_C_CTA AG_AA ACA_ TAT_GTC_C GTC_CACTC_CTA AC_G_ATC_CCA ACA ACA_ TACC_ATC_CTTTC_CCCTC_CTA AGC _G AA ATC _TATC_C GTC_CTA AT CA_TAC CA_ACAGA_ATG ATC_CGTTC_CCCTC_CTTTC_C A 140LDLLCGEIHDVMPLCDDLGSVHSPVPGEPSFPTQSQGAPPALLGQTEHPFIVEQIHVEGP IVEGTPIVEGPIVEQ QIVEGPSIVEGQIVESIVE L IVEQINVE E I P I S IVEVEVEDP GS LNIPS LNCPI SL IP GSL LI P GSL EPS L SS PS LGEP GSS LR P GPSL SP GFSL LP GSLTEP GMSL FP GM SLL P GSL LPDLGAPD LGESPD GK RPDLLGDPDH LGVPD LGPPD LGPCPDSLGNPD LGDLPDELGSPD LGKPD GDPPDEDDLGD P GPGL RGLYLGL LGLDGLAGLGSLAL R LNLTLVL LNLII L LLSTDPKLGDPI GDPAP L AGDP FS GDG G G PLGPGFG G GGG ADAGP PLGDPAEGDPGTGDPDGDPPPGDPKGDPAFGDPDDGGPML ELTLP LES TPP LEF TPLT TED P LTEAP LES TP LTEAP LTEWP LEETCQYL TP LEATP L ETE LP L ITE EP LESITPLEDPLCQI LCQAE LCQC LCQALCQAS LCQFL LCQR LCQLCQST LCQNILQKILQS LQVGSGKIG G GA GV AMSCTCGCTDLTDLNLML L LGA G G GV G G G G VLIVRAPDYTDADVLSDVLSDVADVADVPLPDVMLDVSELDVRLDVGLLDVELLDVALSDVV STSLQFPTSL L TYPD TSLQTTPD TSLDTEPD TSLATDPD TSG LSTTPD TSLDTRPD TS LLT TDPD TSLLTTTPD TSLLQTPPD TSLNTCPD TSLDTKPD TS SLI TDPD TSPLSF8090011 2 3 4 5 6 7 8 9 0 13 3 31313131313131 1 1 2 21 1 1 1 1 1 1 1 13131313131CCA C CG C CA C CG C CCC CC C CA C CG C CG C CA C CA CACT CACGCCCA CCA C C G CACCCG CAGAGC B_GCCGGGB_GTBCAAGT_GTBCG A_G TGGAGGCB CC_GCTGGB_ CGCGB_ CGTABGAGAC GG_CGTCGAB_ CGTBCA_GAB_CG GAB_CGTGBC_GABGAGAAGTCGTG T GA_T_T T_T_ TT_AT_T_ CT_C T_T_C T_C T TGTAATGT T TAGyC GyGGy TGyC GyGGy C yAyTT yAyG_yG_y T_yA_yCCbT CbTTCbACbCbTb T Gb C Gb Gb T GbAGbGGbAGb GbTA_G_TA_A ACA_AGA_AAC _AC _GA ATA AC C _GA AT C _GGA AT CA_AC C _GA AGC _CA AACA_AGC_GC _ACA ACA ATC CGCTC_CGTTC_CGCTC_C ATC_CTCTC_C ATC_C GTC_CACTC_C ATC_C ATC_CGTTC_C ATC_CTCTC_CTT141LDLLCGEIHDVMPLCDDLGSVHSPVPGEPSFPTQSQGAPPALLGQTEHGT EGGEGL EGDEGNEGS EGAENEGDE T E P EYE RPFLGHILAP LE HHFILIHPFLILVHDPFLL HILLAPFLDHIL CL PFLLILIHDPFLILPHGPH D HGQP LPFIL IS PFLILLAP LVHGGHGE HGPFIL LL PFLILFLPFL P P L TILHFIL AVEGDVEGD VEKVE FVEAVEQVEAVEKVEDVEAVE LVESVE PPLPLVP GLEPP GLSP GD P GRP GS P GAP GSP GNP GD P GEP GPSD DSDMSDS LDQS LDF S LDAS L F S LGS LGS L I S L E S LYS L TPLGDPLGLGPPLLLGNEPLLGLDPLGD PVPD LL LGLALGLWPDHPDLPDVPDG SLGLTLGMLGLWLGSSPD LGQPDDELGTGTD LLGDC LGD NCGDVGDQGDKG GGFGLDG QGLGGL IT GLGGPGPDDGPGPDGPAGPVGDPT GDPDIGDLGD D GD R GDEGDWP L TL E L TPLEDGTP LDTEE P L LTE PP LGTED P LR TES P LP TEAP LEDT PPLT P F T PEDFP LEVP LER T PPLEST PTP LA EECQDF LLCQLS LDLCQTVLCQPP LCQSLLCQL LCQTLCQVRLCQDLCQDLRCQS LCQNLCQPGD GGPG GLG GMGSG G GEGGY FDTVDLVGLLVQDLDTDGEDTDE LFDTVS LVL LVL LVVLVKS LDLLVLQLGF LGEDL D PED T DHDKDV ADVCDVSDVYDVTP S AP SLQP STP SL TPDSTPDR TD LYTDTDTDDTDGTDKTDRT L E T LD GT LKT LPP T LHS TSLSIPTSAPTSLQTPTSLDSPTSLQPTSLGPTSLVPTSLPT223242526272829 0 1 2 3 43 313132 3 3 3 3 31 1313131313131313131CCG CV C C C C C CCC CC C GGT G CC CG CC CG CT A XGCTB_ CGCI CAPGCB CCGA_GTCAB_ CACGCCC C CGCACAC C C GCBC_GA AB_CGTAB_ CG GAB_ CGAT B_CG GGBCA_GCBCCT_ T B C CBAGA_CGC_GT CAG_y TGT E_ GCGGCCCGAAGC TGCGGAGGACGGGGCGGTAGCTC bA Gy T_b GyAT_bCGy T_b AGyAT_bTGyCT_bCGyAT_bTGyTT_ GT_ CT_T_T_bTGybCGybTGybAGyAGyGC _ T C _BC _GC _C C _TC _GC _CC _AC _AC _GC _ACb_GCbGA AA A A_7A AGA AAA ATA ATA ACA AAA ACA ATA ACA AAA_AGTC_CGCTC_4C.1TC_CCCTC_CATTC_CATTC_CACTC_CACTC_CATTC_CCTTC_CATTC_CA GTC_CACTC_CG A 142LDLLCGEIHDVMPLCDDLGSVHSPVPGEPSFPTQSQGAPPALLGQTEHPFIVEGDIVENS IVEGNIVEQ GGIVEGF I EKI EG VGVTIVE T IVE TIIIVE L IVENIVEHIVEDIVEAPS L LP GL PAPAP NPYP GIP GNP GIP GRIP GQ P GAP GD P GWPDDF S LPDLLS LRPDP S LPDFL S L EPDWS LV PDTI S LPDVS LPDMS LV PDDS LPDVS L SPDLS LPDNS LPDL S LPDQLGDLGE LGVLGFLGLGLGMLGHLGLGMLGS LGVLGGLGKGLGLGLLGL EGLVGLNIGLKGLNGLAIGLYGLHLQL S L EGD DLGD DFGDLGD KGDLDADGD S D K DSD LGD SG IDVGD AT P T P T PHT P L T PQGPGPS GP SGP SGPP GP LGP IGPHGP TP LAL ED P LY EL P LEPP LEGP LENTSP LEMI TP LS TEFCQS LQT LQVLQNLQNLQGLQIP LQTL EQAP LLL EATQEP LEP TQP LESP TP LEYTL STLD DPE PPERLGGCGR CGECGGCGNCGSTCGLLCGCGELICQLQSLQLQVLQSGQ CLCVCPCHDL LAL L L L LML L T LGLG G GTVLMDVESPDTDSDTVEDT DVV VDVKTDVEV KDV VPVAV ACLGLHVDSDED YFDVGDVEDVE DVITSLDLPTSLLLPTSL E TNPD TSLVTHPD TSLYTNPD TSLHTLPD TSLRTPPD TSLDTAPD TS ELI TKPD TSLNTLPD TSSLFTLPD TSK LTTSPD TSPLSTFPD TS LLIN 53637383930 1 2 3 4 5 6 7 83 3 3 3 34343434343434 4 41 1 1 1 1 1 1 1 1 1 1313131CCA C CG C CT C CC C CA C CC C CA C CG C CC C CC C CAC CA C CGCGCTCTCTCACCAT CA GCACC TGCCB_GCCB_GCBA_GCGB_GCTB_ CA GAB_ CGT B_CGCBCG_GGCB_CGGCB_ CGCTBC_GTBC_ CGA TB_ CGCBA_GCGCGGTGGGGGCAGAT GATCGCAGACGAAGAGGGG C GGT_AT_T_ GT_T_T_T T_ GT_C T_T_T T_T C_AT C_C TGGGyT GyG Gy CGyA GyGTGyAGy TGyAGyGyGy T yAyC _y CCbC CbG CbGCbG CbCbCbGbTb C GbAGb GTGbCGbCGbGA_TTA_TA_CTATA_AAA_CATA_A_AAA AGC _AC _AA AAA AAC_ C C _CA AAA AAC _TA ACC _TA AGC_TA AGCA_AGC CGT C_CCCTC_CTCTC_CGTTC_CTTTC_C ATC_C ATC_C GTC_CCTTC_C GTC_C ATC_C ATC_CGTTC_CGT143LDLLCGEIHDVMPLCDDLGSVHSPVPGEPSFPTQSQGAPPALLGQTEHPFIENI E S I E P I E I EGP P P P P PDP P P PV VIVGDVGLVG Q MDQ MEQ MGQML Q MLQ MNQ MD TQD TQLP GVP GSPEP E P TS S LIS L S L L S L S L S L E S L F TLI S TLIDS LNS LGS L S L S S LGVEGVWTGVMGVEVGVFVDI S GI SPPDLGD LPD LGAPDA LGYPDQGLGDPD LGGPL STFLL STCIL STD L SQG TGL SSG TVL STSG GL STDSGSLL SGSLLEL T L SSLLGI LILQSP TSP DSPL SPVSPHSP LSPL L LML LQTDPIG RGDP IG DGDG PHDTGDLP L S GPT GPAPMS PM AGAAGE PMDPMPMSPML PMAS D SG AGFGDSGPG GDTAL TAVL EWTP LET TAP LTE CP LL TEQ P LETAALWAAG LQAALDA DAALLACQVLQLAALVA P AAGLIAASGD GALGPDDFGD PDDL S LSLQVLQS LMGAPGALGALGA GAGGA GAL TTTTSGACGTCDCHCQPQLGTMLGGLGSPWQPWQNWQAWQDWQEWQRWQTPDTPLAQNSAQSAQDAQLLPN MTS DLTSLDVVDVPDQ Q DVLDVMDVSAQ AQHAQ N ALS CS DLSLPASLP DLTSL I TD VPTSL L TAPD TSLNTDPD TSL L TGPD TSRLMTV TRSGN LTV TRTSLN PTV TRSGN GTV TRSGN ETV TRFSPN TTV TRQ SSN ETV TRSDPFTT SGDPSTT SGLC 9405152535455 6 7 8 9 0 1 23 3 3 3 3 35353535353636 61 1 1 1 1 1 1 1 1 1 1 13131CCT C C 2 C CCCC CCCTC CTCCCC CC C CT AC _PCT T ACGCTACGTT CTACTACGCCGCA TAAAAGCGGB_CGCT B C_GOCGTBCC_GA TB_CGT BT_C C B_CAT B_CG ABT_CA TB_TCA GB T_C CCBC_AAT B_CG AAB_TG_CGGTG_TGTTS_ GT TT_CGT T_CG AAA_TG A A_TG TAA_CG A G_CG AC_ CG CAA_ TG CAC_AA T A_CA T G_CGy yTy yAy T C yGC y T C yGCbGGCbCGCb Gb AGbAAbTAbGAbC CAyCbT CAyb CC yCAb TCCAyTbC CyGbC C yCbTA_TAGCA_CATA_ CA A_AC CA_ACC_CTTTC_CCCTC_CB_TC_CACTC_CC C _TGAG G_ C C _GAACC G_GAT C _AGAA G_ C C_GGAACTG_ T C _AGC _AGC _T GAT T _AGAC T _GATAGTG G_AGTG G_ACG G G_AGTAC_CCTAC_CGT144VDWTEPGAEDLLAFLSQDVDLPPPLSLLPPQAESPGESTAVDTAHVDQPTQPQT ET EPQPQPQAEAW AF A A AG ADVLVL DI LS ISDTI LI ETTI LD ITLT DD DTI LINTLI LLRQD LR TD CLR SD LRMD DLRFLD LR SD VLRD SI CD DI SI CDPLGSIE SSSWT SSS LLL F G GSSS EIFSSSDF LSG HFVLSIHFDLSG HFLLSHF LDLS THF S LSHFHLSLHFADEEVFD LEVESL L LLLCIGLLL SGLLSGLGLLLDFDFDEFDLFDTAT SAD SAVSAL SADLGDVLS LGG VLLGGVLI LGFFDA VLDLGFDSVLWLGPFD VL VLGDVLS EG NSLTEG S NSQ LG VGDS TDE TDHT L TDLKPLL PQL PGPDPPP P P PGPALE RAE R DTD G WPDGG PDSPGD PDGIG P A DD WQ LDK WQ K LNWQ LRK NWQ LLK AWQ LNKQGKQT ST TSWL E WLMNQPPT IWT SPNQSILPT TPQL TTTPVP TTTPGTRT TP SGIVL L IYL VLSIYLVLYHIVLYDISVL IYAVL P IYS VLDPLAMP PAMLTSTSNTSGTSNTSL T R C T RTT RQPSTGTGT F TY AVNPSAVDSLNSSSLS SLE SLSLMH GHLH APPS L PPS P PPEHRHR LHRPHRDF LNLNLL GF T SHPPGS T SRGL T TENIRHT TPNLRGT TQR G R C RTRATTG GTTGTLTTS T Q T D W W WN DWT TNLDWT TNIQWT TNQSWT TD N DWTDI IGLWTDI I CG 36465666768696071 2 3 4 5 63131313131313137137137137137137131CC C C CT CGTGG GTGCGCGG G GCGTACAAGTAGCAGAA TCA GCGT CG GCGACGACGC CA A GTCACAGAT B C_ATCB_CAA TB_ CA AGBC_ACCBG _GAB_GTCBA _GGB_GAT B_ GBABCBGGBGGBTACGT_GT_GC_AT_ AA_C_TA yGTA_TT A TC_ CCA T A_ TCA TC_AGTG_CGCTA_TTGTA_ TCGTA_ GGTA_TGT C_ CCGT C_ATA_T TG_CCGbTT _TC yGbTGC yb CCC yb TCC yTbCGyTbTGyTbTGyGT b TCGyTbCGyGT bTGyTT b CCGyT CyGCyTbC TTbT TbTAT _GT _GT _GT _TC _GC _ C _ C_AC _ C _ C _T_T T_GACAGC_CCCAC_CTAG AGAAC_CGTAC_CCATAAGAC_CGTG_ TAAGGTG_GGTAAA G_G ACACAGAT CAC CATGCA G G_GCTA G_GCCA G_GGTA G_GGTAT _GCCAT _GGT145EIQENDLPNDDFVFKNARPVLLHPVEAETENLVCMDVPRICCSSGTHLHVL LDL L LS HGAHG HGLHGGHG HGPHGLLQLQI C SV GSI CDV LSI CNV ESI C EV SI C LDIFTNI I TKI TKFNVFNDIFTNL IFT LNAIFTNLIFTNMRID EEDRIDEEEDDEVLDDEGMEVGDD VFDWD S DVTDVF IDNTISILNTIFIENTIGL INTIG EINTIDISNTI C IDNTIDPLVLI IE PVLI WTNS DEGSSEVEGSGEGCLESIEG DES DEPEPEP DEPE EPEL LLI SNLSDI LS PLSLSGELPSDELL PSDL SFVDL L SDCITRLN SDFELTRHN SSPELTRLN SGIELTREN SGELTRAYGE ISD ESFYGLSES E ILFYGSLILE LEYGSWTIYGL ISMYGSGIQE L CIE L E LSYGFESDE TL DP PTV SE TAP PDENQDNQ VNQ NQ Q NQS S LGS L S L S L S LDL S LVS L LQYSQ YSGP I P I P P IGP I L P IGNL LNL TSNLGNL DNLNLHNLASWSQAMDAM AMRAMNALFF LFFFFVFF EFFD FFS FFAEIPAEI LALA GENSMLP GLPALPIDLPIGLPI F LPP LPIDSNQPNQNLVA NLVP ALVHALVL ALVMNI INIWNSN Q NDNIVNEQNEQSWTD DII SN GWTDSN IIFPWTDQ N IISEWTDTNDIKGIK IILPWTDI I LAN D V KRANPN V KP IK ANLN V KL IK ANL IKDIK D V KNSAN V KLANPA V KGIKGE ANLHV K MIRVI SAHA GIRVI LATL77879708182838485 6 7 8 9 03 3 3 38 8 8 8 8 91 1 1 131313131313131313131GC T TT TC TT T TC T TCAC AGGACGCGCGGCGTGCA T GG GA GA GGT GA GGC GA A AAAGTAT BT_GA AT B_GA GB_GTCB_GCCB_TTA GB_TTGT BTG _TAB_ TTTCB_ TTAT B_ TTA TB_T TTCCB_ TCGT B TCTCB_CT_CAyGT C_ C ACTA_ T ACTA_TATT C_AGTA_ TCGTA_TGTG_CCGTA_TTGTA_CGGT C_ CCGT C_ATA__T TA_TTT bCCT yb C CyC _ C Tb T CyC TbTCyTGTbCGyCb TCGyGyCbTGCbTGy T yCbGGCbCGyb CyT ACGbC T yGAbT T ybTGAAAT_T_T _ C CAGCTAT _GCAGGTAT _GT_T_TGC CAGAT _GGT CAAAT _ TT_AGT_TACT_GATT_AGT_ACACT_ CAGT_TG ATT_TG ACT_AGGGT CA_CCGCA_CCCCA_CGT CA_CTACA_CCT CA_CGT CA_CGT TA_CCCTA_CTA 146SAIAPVFTLPNLDLDDCLIESFMEPTMLKENMNWTFADRLKQGDIEMDLQDLQD LQSLQLLQLILW FLLFVVPVVRIDPE DL RID EPLRID EGRID ENRID EDAN AEAN ANTAN WF AWMAWC AWSANGAN GAWSAWEAN A A WDDS LAD ED S IVL ELISGP LSVIE P L L P L E P L F TQVISMVISF VIDLDLTTLDDL T ILDDTLDLTLDVTQTHLDGLDDL SHFTQSHFTFLVDVL SEVDGLVDDL SVDGL SVDDLLKSGAALK GADLKE LKL LK FGAGGAGIGASP LK GAVLK GAA DCEG CVC ECTSPTPHS EPTPV DEPTP LDEPTPLLEPTPASF LLWSF LLDSF LLQL SF LLGSF LLVSF L DLS SF LL SM GLDM SGLNAQSYPSVQSYS SLQSYS FDQSYSGIQSYSDS SNPLP SNLDL SNLN SSNRLNSNPLGSNLLL SNGTNLLAELT EW HLAPAENI PGAEI LEQDAEIDLAEIGRAEIGLM DKNM AM M HMEMDM M SSDKS DDKSLDKSQDKSPDKSLDKS D VIGDLVH IGPNHQENQ PEQLNQIRI SHIRILEQANEQ QNNEQ QMQ CHIRIDSHIRIHH DGRA GQ GR SQTQ GGR LPGRSEQSQLQ GR FGR CGR LKSVICLKSCVICSA AS IRVIALKLD VS LKLVIVSKLVLVSVLKLG VSQKLV D VSPVTKLQ VSGKLVIVSDVFK K V AFPV A G V A G VQ C G ED VNRGEVNRGL19293949596979899900102 3 43 3 3 31313130 0 01 1 1 131314141414141AGAT AC AT AA CC CC C CT CGCT CATTTCAT CAA A G G G GGGG G GA G GA AAAT CACACTC CCC C TTA GTACABTT_ TCAB_TCAT B_TBT CBCCG_C C_ CGT B_CAT B_CCTCB_CA GB_ CA TB_ CG AB_C CCB_G AAB_G AGT B_AC_ CCTA G_CCTA A_ GTA A_ TCTAC_ACTA_TCTA_ GCTA_TTCTA_ TCCT C_ CCCTG_CCCT C_ACTG_CCCTA_TT yCT y T T y C T yTT yT TyGTy CTy TTyT TyC TyTyTy yGGbT_ CGb b bCbC CbTTCbCbGCbC CbC CbT CbCCTbTCTbTTTAGT_GT ATG T_AG_TACTTAGG_TCTTATG_G ACG_AG_G_G ACG AGTAGG_AGG_GATG_T_G_ATCATCACA_CGTA_CGTA_CCTA_C G A_CGTA_TCC A_TCTA_TG A A_T CG G A_TGTG A_TGTG A_TGT TA_CGT TA_CCC 147C KGCPHIFASVVRDGPKVRTVGPGVEEVIGAIEHGAIVPPPVPTKGEFV NVV V V VDVVHD VHDVH VH VH VH VHTWT LDASS NV EDAGV SLDAEV SDA SLV DA SDF EKTTP EKTTDL EKTTDIEKT ETEKT STGEKTNTE E T LTDQPQTCQPQGSHFTFS SHFTMSHFWTT SHFGTS SHFTD DRL LLTERL LTGSRL ETF RLWTTRLTLRLFK TSRLTFSDL P ISDDL PDVC EGC EDC ECI C EVC E PQP L P L L PCI L PML PGL PD EEEHMCLLLMCLLDMCLDEMCH LSMCLLAMTGGMTVTGVMTHTGT MTS TG TDEMTGD MTL TGL MTL TG TL THMLGTHMLSPGTNGIGTNFDGTNGGPTNVGTNDSQ LSS DSQ LSSSPQ LSA S Q WLSSGQ LSDSFQ LSS GIQ LSAAQQS D QGLAQVQGPAEGAEDAEQAE PAEGF SLF SVF S F SQF SDF SGF S SVN VGVHIGRVH IGLAVHLIGNVH IGG EVH IGLTTDI LTPTDIGT PTDI PTLTDINTTDID TTDI RTTDIGTLLDSTLLDEPKSNKSIKS SKS PKSMAGDAGEAGNSAGSGLGN GLVTLVSVIHVDSVI LVI SVI DALA APA ALAAA AHA M A DTS LDTS FKCQKCGKC TPKC FKC LKVR TVR F THLV NCKGTHSKAKTK D KQK D T P T PVNRSEVNR G VNRL NP NDV NFP TH V NGL TH V NLP TH V NSGTH V NSE TH V NLDTQ V QLGTQ V QTQ 5060708090011121314 5 6 7 84 41414141 1 1 1 11 1414141414141414141TTTCT T G TAATAG ACAG ACATAA A ATGGATATGTT C TTCA GCGC CGACGT CGACG GCGT T GCTT GCAGB_G AAT B_G ATCB_GA AT B_G ACCB_GG AB_ GA TB_GGT B_ GTCB_ GAT B_ GA GB_GCCB_CTB CABTC_ CTT_CTA_ TCy CCTA_CGCTA_TTCT C_ CCCC_ACC G_C CCC_ CCCCA_T CC A_TTCC A_CGCC A_ TCCCC_AA A_TT AC_ CCTb TyC_ CC bCCybT yGC b CT yTC yCCC bCAbT C yAb CC yGCy T C y C yCAbTAbGAbCAb TCC yTAbCG CybTGyGC b CTA_ GT _AT _ T _CCTA_CCTAGA_CTCATAGTC_TA_CT TAT C _GAA_CTAC_ T C _A GTAC_GC _TAC C_AGC _AACC _AGC _TAT T _AGT _ CAG GTAC_GCCC_GTAC _GT C _GCGC _A GTAC TA ACG ACCTG G G G A A C A A G _A A_CGT148SQHPSTASAVQVVQLQIHQPQQLGSASQSVPIYQIRQYDGQPTRYDTTLTIQP T L T TFT F P E PDP S P PDP P L LGLGGSLQEPFQPLSLQPMQPLSLQP EQPQSLQSQPDAPGSLQPD DDSDAPDATIWTDSTILEDSTIGDALDSD TIIEDADSDTI LDADSNTEDASIF D TIDFDNERDED NNGE RLLEDTTHSEDDTHLED THGEDLTHLEDL PTHAVDSCP PIVDQVDSMPFVDL PGVDS PVDS PD GVD DD EEIQDN EEIGILALDF LVLGI LDFDF SGFDLF STS F SVF SLF S L PGL PGGMQ MQ MQDSMQ MQSNAENAVNANANAHNALNAAQDNQ DRA WPA DALAGA HLGHLDHLDHLAHLSPHLGHL S SSL S SNQGPQGD QGQGRQGGVV VVSVVFVVWVV VVIVVDDN DNHTVLTVL TVL TVNTVLTDQLTDL TDDTDP TDVTDGTDSTQLDNVSLLDA VLLDDVLLLDH VQLLDMIDDPNIDDP L IDDPDL IDDP P IDDPPGIDDP R IGNNLNN DDP LVGP VGSEDT TAT SGDT DQLTT S SD GT LQTT S CDTQGTT SSVDQEDTQTT S L SDVSQDFVT LSDVDTLLSDV TASDVNTS SDVETPSN DVHSDVMTYL TY MAGMAQ VC TV HPVCHCV G VCDASTQT DLD V QCIV QGLV QEIV Q D V Q DLHSV G VCHGLV VCHFPV VCHSEV VCHLYA D V APPYA V AEC 9102122232425 6 7 8 9 0 1 24 42 2 2 2 2 3 3 31 1414141414141414141414141A A AT AT A T TT T TTC C AGC TC T T T T A G G AGG GCATCATCATCGTCT CCTCCGCCACCACCCCCAC TT T TGC GBC AB_ CGBCAB_CCB TB BAB_ GBAB_BC CB ATBAABT T_T T C TA_TGTC_GC_ GA_ GTGT_ GT GG_GC_ T C_TG_A A G_TGA A_ GA G_CCA A_ TCAC_A TCA A_TTCG A_CCCACA_ GCA A_T CAC_ CCCATA_C CAC_ATC A_TTTC A_ TCCybTGyT _TCC bCG T_ACCybTGyTGyT _GTC bT_ CGC bCT _T Cy TT CbGCybTCybCCyGbT yT_T C b CyAGCT_GAT CT_AACC b TyTCC bCCyCbTGCyCb TCC CT_AC CT_AGCT_GCT_TT G_GG_A A A A A A A ACA AT _ T TA AA AGA_CCC A_CCTA_CGTA_C G A_CGTG G A G_GGT TG_GCT TG_GCCTG_GGT TG_GCGTG_GGTCC_T TACC_T CG 149WAIPMFTKIENDDMQQEHVAAAAAAAAADEEGVIVEMYVEDAIEAVNLGL LGVLGGLGLGLNSNG N NEN NTNG GPD NE RTSD NRHS D NRVD NRDLD NRATP FLGTP FLSTVP FMTL DP FTLQP FF TLL P FC TLI P FD DV EISKV LVEISLLCDN EAEDN EPVEDNDEESDNDEEFDND VES R TL VPL R TVPHR TPLVDRTGVVR TTSVRTDEVRLT LAKLEFEKLEDEP IWQGP EP IGPEIL EIDEIGPGL PGPGGL DGAP ID APSPD AP FDPDPADPGDPAPDSAPAPQAPD VSEVDVIEVDLSDPSNQSDSEPQSDSDLQSD DSLQSDSMEY LGREY LVPEY LD DEY LL EYW LPY PE LLEY LGRETIIEF RETII GSDNSDNSDNLDNADNDP PNP PGP P L P P L P P P PNP P LEL LELVN AF CDLMVHMVEMVAMVDMVNSMVSLMVMEP TEPHVNTGGNN YLVTGPNN NNSNN YTVTGGVGGVGDYEI TYGTYFYRAVQYR PYRAVSAVDSYR LYRYRYRAVLAVAAVTAVDDE ASEE AFE AGE ACE AGE AL LGTYVSIAGTYVISPMACIMAQSMAHMALMAL E L EPL E L EGL E LP E ADEGWEGVYAQYAQYADYADYADLQ AG AGTAG E CL E L L E FAS P ASPV AEIV A G V A V V A G V A A Y GDEY GQSG Y GLAG D Y GIAG H Y GIAG Q Y GGLAGD Y G DPYPNPNPYPNGE3343536373839 0 1 2 3 4 5 64 43 4 4 4 4 4 4 41 1414141414141414141414141GTCAAGTGCGTT AGTC AGA AT A AC AT AC A AC CGG GACCCGTGTT BGTT CG G ACCG CA G CA GA G GC CTG TC T GAGC_ATA TB_ATAB_ ATAT B_ AT CCB_ TBTABTAB T B TGB T B T CBTGB TABCA GG_GT_GT_GA_T_C_ C_AT_ AT_C _TTyGC C_ CCTC G_CTCCA_CTGC C_A ATT T _CTC_ CCTA_CGTG_CG CTA_TG A GT _TG TTC_ACA_T C C_ CCCbTCyCb CCCyCbTCyCbCCy C yCbCAb TCC yAb CC yCAbC CyAbT C yAbT C yAbT C yTGAbCGyGCbTGyCb CG_TCCG_G_G_A_TC _GC _ C _AC_GC _TC _ C _TA_TA_ CA A G ATG ACG ATC_TCCCC_TGTCC_TGTCC_TCTCC_TGTGAG_GCGAG G_G GAGGTG_ CGAGCTG_ TGACAG ATACAG GGTG_GCCG G_GTG A G_GGTAC_TCCAC_TGT150EL EH EA E EQ L QQQ DLTDnP C PSSQ QLeFLGE PGP L PDvS EIFLQFLG FLPLFLFPHSPDDerESP LSGSDPLPD pVD V V VPVLtKVKC KG KPAaSISMSIPQSILDSIPKSD htLPILSLVSLPLSVILS SGnDC I RDVI ELI SLQI LoiQDtRDALQED RH QQ D D GLQ E QMLaCP RVMR PGRDLD GcoVGSA A VGKVD DA A VY VDL lAVPAPIKHLKMLASLLAG AFDPehVRSYKLDFKGRD RKD L tPVSVRGVDDVVGVL E tS VVRSMR R LLDRA YHRAQaL PSLKQ SL C LSL R PSLDGnMGELE TKTP SMLKEA HMGA LE D MDSLGRS MSVoLGLDSdoRSGT SMTIHSGTARTML cSFEIPE RSTGT RDL RYPRDLpPTGIGASIVMSITRSI LDLS QGPFYP MD PSR P DLLotsI SGQELSSG IKGLPSI PLL LVR L FDSILPDCTSIDaLGRGC FG GG GQSHVAD E eYASP MLPYEALDDL RQL L IHtaGPIGSP T YGC T SD D YSP R YSA D LGL L cAGR TGD ViL GdI SQPILGKGI S F IGYLI S IGGDSLI SSS IGS LI SM niSGSRSAKG G TSGSGS KSGSGLKSGPSE KGPLLGERLGRRLGMRLGE SGRSGC”D**YGCT SYGAYGD YGQL LYGD*“DGLLT S RAT S L T SET SLFSGEDG G SGF GSHDDDG EF GS FF GSDGSSLIF GSGeGdVGL SEVGASD VGSGNSGLulISKLD LGEISILKE S DSI L LV ES GV ESMcAI L L I L ni.dVE LKLLKL DKLGKL DL t iEVD VENVEVSVEVE VED a cRETP EV IE I LLERETIIE EFR TIGELR TE EV IWR TI FDhtacEPQ GEL S EIPGEL EIPMEL T EIP CIELPDsLe ielGT E TLE TDE TDE TAcncYVVGE ID YVLG EIGIYVLGVEGVEIDFYIGYIDS euunAGPSSLAGS GGDEGQEGGq eYPNLDPYPNRANPSYPANDL PSYPASLNLNPYPN Meshdti7cy484494051abde1414145141 onidomcnCCTCGACTCGCCCACG CaAeed ecTGBGAGBGT CTGTitnAA_TB_ TAT_T C B_T CB peCGCAG AT AACAATAC_ euG_yCCCbTG_y CTCG_yGC_ TC C_A TpqbCGy TGy Co esA_GCbAATA_ C CAGA_ACbACA_GCb TAGA_AT bC_TGTAC_T CmdGAC_TCTAC_T TAAC_TGTo etCsil151HETEROLOGOUS ENDONUCLEASES

[0125] In some embodiments, an engineered gene effector of the present disclosure includes a polypeptide that is coupled to a heterologous endonuclease (e.g., enzymatically active Cas protein, enzymatically deactivated Cas protein, etc.). In some embodiments, the engineered gene effector as disclosed herein, or a protein comprising the engineered gene effector (e.g., a protein comprising the engineered gene effector coupled to the heterologous endonuclease) can be referred to as an actuator moiety. The engineered gene effector and the heterologous endonuclease can be coupled to each other, e.g., directly or indirectly (e.g., via a linker). For example, the engineered gene effector and the heterologous endonuclease can be fused to each other, e.g., directly or indirectly (e.g., via the linker). In another example, the engineered gene effector and the heterologous endonuclease can be non-covalently coupled to each other, e.g., via ionic bonds, hydrogen bonds, interactions mediated by oligomerization or dimerization domains, etc. In some cases, the engineered gene effector and the heterologous endonuclease can be part of a single polypeptide molecule (e.g., a chimeric or fusion polypeptide).

[0126] In a wide variety of organisms including diverse mammals, animals, plants, microbes, and yeast, a CRISPR / Cas system (e.g., modified and / or unmodified) can be utilized as a genome engineering tool, or can be modified to direct specific binding of engineered proteins to target loci as disclosed herein. A CRISPR / Cas system can comprise a guide nucleic acid such as a guide RNA (gRNA) complexed with a Cas protein for targeted regulation of gene expression and / or activity or nucleic acid binding. An RNA-guided Cas protein (e.g., a Cas nuclease such as a Cas9 nuclease) can specifically bind a target polynucleotide (e.g., DNA) in a sequence-dependent manner. The Cas protein, if possessing nuclease activity, can cleave the DNA.

[0127] Non-limiting examples of the heterologous endonuclease as disclosed and used herein can include, but are not limited to, CRISPR-associated (Cas) proteins or Cas nucleases including type I CRISPR-associated (Cas) polypeptides, type II CRISPR-associated (Cas) polypeptides, type III CRISPR-associated (Cas) polypeptides, type IV CRISPR- associated (Cas) polypeptides, type V CRISPR-associated (Cas) polypeptides, and type VI CRISPR-associated (Cas) polypeptides; zinc finger nucleases (ZFN); transcription activator- like effector nucleases (TALEN); meganucleases; RNA-binding proteins (RBP); CRISPR- 152associated RNA binding proteins; recombinases; flippases; transposases; Argonaute (Ago) proteins (e.g., prokaryotic Argonaute (pAgo), archaeal Argonaute (aAgo), or eukaryotic Argonaute (eAgo)); or any derivative thereof; any variant thereof and any fragment thereof.

[0128] In some embodiments, the heterologous endonuclease as disclosed herein can be nuclease-deficient. In some embodiments, a heterologous endonuclease can be a nuclease-null DNA binding protein that does not induce transcriptional activation or repression of a target DNA sequence unless it is present in a complex with one or more engineered gene effectors of the disclosure. In some embodiments, the heterologous endonuclease can be a nuclease-null DNA binding protein that can induce transcriptional activation or repression of a target DNA sequence (e.g., which can be altered or augmented by the presence of an engineered gene effector as provided herein). In some cases, the Cas protein is mutated and / or modified to yield a nuclease deficient protein or a protein with decreased nuclease activity relative to a wild-type Cas protein. A nuclease deficient protein can retain the ability to bind DNA, but may lack or have reduced nucleic acid cleavage activity.

[0129] In some embodiments, the heterologous endonuclease as disclosed herein can be an RNA nuclease such as an engineered (e.g., programmable or targetable) RNA nuclease. In some embodiments, the heterologous endonuclease as disclosed herein can be a nuclease-null RNA binding protein that does not induce transcriptional activation or repression of a target RNA sequence unless it is present in a complex with one or more engineered gene effectors of the disclosure. In some embodiments, the heterologous endonuclease as disclosed herein can be a nuclease-null RNA binding protein that can induce transcriptional activation or repression of a target RNA sequence (e.g., which can be altered or augmented by the presence of an engineered gene effector as provided herein).

[0130] In some embodiments, the heterologous endonuclease can be a nucleic acid- guided targeting system. In some embodiments, the heterologous endonuclease can be a DNA- guided targeting system. In some embodiments, the heterologous endonuclease can be an RNA-guided targeting system. The nucleic acid-guided targeting system can comprise and utilize, for example, a guide nucleic acid sequence that facilitates specific binding of a CRISPR-Cas system (e.g., a nuclease deficient form thereof, such as dCas9 or dCas14) to a target gene (e.g., target endogenous gene) or target gene regulatory sequence. For example, the target gene may be any one of the genes listed in Table 6, and the target gene regulatory 153sequence may be operatively coupled to any one of the genes listed in Table 6 Binding` specificity can be determined by use of a guide nucleic acid, such as a single guide RNA (sgRNA) or a part thereof. In some embodiments, the use of different sgRNAs allows the compositions, combinations, systems, and methods of the disclosure to be used with (e.g., targeted to) different target genes (e.g., target endogenous genes) or target gene regulatory sequences.

[0131] In some embodiments, prokaryotic CRISPR-Cas (Clustered regularly interspaced short palindromic repeats-CRISPR associated) systems, for example, Class II CRISPR-Cas systems such as Cas9 and Cpfl, can be repurposed as a tool for regulation of gene expression, epigenome editing, and chromatin looping in compositions, combinations, systems, and methods of the disclosure. In some embodiments, nuclease-deactivated Cas (dCas) proteins complexed with heterologous gene effectors can allow for regulation of expression of target genes (e.g., target endogenous genes) adjacent to a site bound by the dCas.

[0132] Any suitable CRISPR / Cas system can be used. A CRISPR / Cas system can be referred to using a variety of naming systems. A CRISPR / Cas system can be a type I, a type II, a type III, a type IV, a type V, a type VI system, or any other suitable CRISPR / Cas system. A CRISPR / Cas system as used herein can be a Class 1, Class 2, or any other suitably classified CRISPR / Cas system. Class 1 or Class 2 determination can be based upon the genes encoding the effector module. Class 1 systems generally have a multi-subunit crRNA-effector complex, whereas Class 2 systems generally have a single protein, such as Cas9, Cpfl, C2c1, C2c2, C2c3 or a crRNA-effector complex. A Class 1 CRISPR / Cas system can use a complex of multiple Cas proteins to effect regulation. A Class 1 CRISPR / Cas system can comprise, for example, type I (e.g., I, IA, IB, IC, ID, IE, IF, or IU), type III (e.g., III, IIIA, IIIB, IIIC, or IIID), and type IV (e.g., IV, IVA, or IVB) CRISPR / Cas type. A Class 2 CRISPR / Cas system can use a single large Cas protein to effect regulation. A Class 2 CRISPR / Cas systems can comprise, for example, type II (e.g., II, IIA, or IIB) and type V CRISPR / Cas type. CRISPR systems can be complementary to each other, and / or can lend functional units in trans to facilitate CRISPR locus targeting.

[0133] When a heterologous endonuclease includes a Cas protein or derivative thereof, the Cas protein or derivative thereof can be a Class 1 or a Class 2 Cas protein. A Cas protein can be a type I, type II, type III, type IV, type V Cas protein, or type VI Cas protein. A 154Cas protein can comprise one or more domains. Non-limiting examples of domains include, guide nucleic acid recognition and / or binding domain, nuclease domains (e.g., DNase or RNase domains, RuvC, or HNH), DNA binding domain, RNA binding domain, helicase domains, protein-protein interaction domains, or dimerization domains. A guide nucleic acid recognition and / or binding domain can interact with a guide nucleic acid. A nuclease domain can comprise catalytic activity for nucleic acid cleavage. A nuclease domain can lack catalytic activity to prevent nucleic acid cleavage. A Cas protein can be a chimeric Cas protein or fragment thereof that is fused to other proteins or polypeptides. A Cas protein can be a chimera of various Cas proteins, for example, comprising domains from different Cas proteins.

[0134] Non-limiting examples of Cas proteins include c2c1, C2c2, c2c3, Casl, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cash, Cas6e, Cas6f, Cas7, Cas8a, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (Csnl or Csx12), Cas10, Cas10d, CasF, CasG, CasH, Cpfl, Csyl, Csy2, Csy3, Csel (CasA), Cse2 (CasB), Cse3 (CasE), Cse4 (CasC), Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csxl, Csx15, Csfl, Csf2, Csf3, Csf4, Cul966, Cas13a, Cas13b, Cas13c, Cas13d, Cas13X, or Cas13Y, and homologs or modified versions thereof.

[0135] In some embodiments, the Cas protein as disclosed herein may not and need not be Cas9 or Cas12a. The Cas protein as disclosed herein can have a smaller size as compared to Cas9 or Cas12a. The Cas protein as disclosed herein can be derived from Un1Cas12f1. In some embodiments, a heterologous endonuclease can comprise an amino acid sequence having at least or up to about 50%, at least or up to about 55%, at least or up to about 60%, at least or up to about 65%, at least or up to about 70%, at least or up to about 75%, at least or up to about 80%, at least or up to about 85%, at least or up to about 90%, at least or up to about 91%, at least or up to about 92%, at least or up to about 93%, at least or up to about 94%, at least or up to about 95%, at least or up to about 96%, at least or up to about 97%, at least or up to about 98%, at least or up to about 99%, or about 100% sequence identity to the polypeptide sequence of SEQ ID NO: 2222 (e.g., CasMini). In some embodiments, a heterologous endonuclease can comprise an amino acid sequence having at least or up to about 50%, at least or up to about 55%, at least or up to about 60%, at least or up to about 65%, at least or up to about 70%, at least or up to about 75%, at least or up to about 80%, at least or up to about 85%, at least or up to about 90%, at least or up to about 91%, at least or up to about 15592%, at least or up to about 93%, at least or up to about 94%, at least or up to about 95%, at least or up to about 96%, at least or up to about 97%, at least or up to about 98%, at least or up to about 99%, or about 100% sequence identity to the polypeptide sequence of SEQ ID NO: 2231 (e.g., dCasMini). As disclosed herein, SEQ ID NO: 2222 encodes the polypeptide sequence of Un1Cas12f1. As disclosed herein, SEQ ID NO: 2231 encodes an engineered variant of Un1Cas12f1 with reduced nuclease activity. As disclosed herein, SEQ ID NO: 2232 encodes a non-limiting examples of a Cas12f variant suitable for use in the systems, compositions, combinations, and methods of the present disclosure. In some embodiments, the Cas12f variant as disclosed herein can comprise an amino acid sequence that is at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% identical to the polypeptide sequence of SEQ ID NO: 2232. SEQ ID NO: 2222 (Un1Cas12f1) 1 MAKNTITKTL KLRIVRPYNS AEVEKIVADE KNNREKIALE KNKDKVKEAC 51 SKHLKVAAYC TTQVERNACL FCKARKLDDK FYQKLRGQFP DAVFWQEISE 101 IFRQLQKQAA EIYNQSLIEL YYEIFIKGKG IANASSVEHY LSDVCYTRAA 151 ELFKNAAIAS GLRSKIKSNF RLKELKNMKS GLPTTKSDNF PIPLVKQKGG 201 QYTGFEISNH NSDFIIKIPF GRWQVKKEID KYRPWEKFDF EQVQKSPKPI 251 SLLLSTQRRK RNKGWSKDEG TEAEIKKVMN GDYQTSYIEV KRGSKIGEKS 301 AWMLNLSIDV PKIDKGVDPS IIGGIDVGVK SPLVCAINNA FSRYSISDND 351 LFHFNKKMFA RRRILLKKNR HKRAGHGAKN KLKPITILTE KSERFRKKLI 401 ERWACEIADF FIKNKVGTVQ MENLESMKRK EDSYFNIRLR GFWPYAEMQN 451 KIEFKLKQYG IEIRKVAPNN TSKTCSKCGH LNNYFNFEYR KKNKFPHFKC 501 EKCNFKENAD YNAALNISNP KLKSTKEEP SEQ ID NO: 2231 (deactivated nuclease variant of Un1Cas12f1) 1 MAKNTITKTL KLRIVRPYNS AEVEKIVADE KNNREKIALE KNKDKVKEAC 51 SKHLKVAAYC TTQVERNACL FCKARKLDDK FYQKLRGQFP DAVFWQEISE 101 IFRQLQKQAA EIYNQSLIEL YYEIFIKGKG IANASSVEHY LSRVCYRRAA 151 ELFKNAAIAS GLRSKIKSNF RLKELKNMKS GLPTTKSDNF PIPLVKQKGG 156201 QYTGFEISNH NSDFIIKIPF GRWQVKKEID KYRPWEKFDF EQVQKSPKPI 251 SLLLSTQRRK RNKGWSKDEG TEAEIKKVMN GDYQTSYIEV KRGSKICEKS 301 AWMLNLSIDV PKIDKGVDPS IIGGIAVGVR SPLVCAINNA FSRYSISDND 351 LFHFNKKMFA RRRILLKKNR HKRAGHGAKN KLKPITILTE KSERFRKKLI 401 ERWACEIADF FIKNKVGTVQ MENLESMKRK EDSYFNIRLR GFWPYAEMQN 451 KIEFKLKQYG IEIRKVAPNN TSKTCSKCGH LNNYFNFEYR KKNKFPHFKC 501 EKCNFKENAA YNAALNISNP KLKSTKERP SEQ ID NO: 2232 (Cas12f variant) MAKNTITKTLKLRIVRPYNSAEVEKIVADEKERRKQAGGTGELDDKFYQKLR GQFPDAVFWQEISEIFRQLQKQAAEIYNQSLIELYYEIFIKGKGIANASSVE HYLSRVCYRRAAELFKNAAIAGLRSKIKSNFRLKELKNMKSGLPTTKSDNFP IPLVKQKGGQYTGFEISNHNSDFIIKIPFGRWQVKKEIDKYRPWEKFDFEQV QKSPKPISLLLSTQRRKRNKGWSKDEGTEAEIKKVMNGDYQTSYIEVKRGSK ICEKSAWMLNLSIDVPKIDKGVDPSIIGGIAVGVRSPLVCAINNAFSRYSIS DNDLFHFNKKMFARRRILLKKNRHKRAGHGAKNKLKPITILTEKSERFRKKL IERWACEIADFFIKNKVGTVQMENLESMKRKEDSYFNIRLRGFWPYAEMQNK IEFKLKQYGIEIRKVAPNNTSKTCSKCGHLNNYFNFEYRKKNKFPHFKCEKC NFKENAAYNAALNISNPKLKSTKERP

[0136] In some embodiments, the amino acid sequence of the heterologous endonuclease as disclosed herein can be mutated and / or modified to yield a nuclease deficient protein or a protein with decreased nuclease activity relative to a wild-type Cas protein. A nuclease deficient protein can retain the ability to bind a target gene (e.g., DNA), but may lack or have reduced nucleic acid cleavage activity. In some embodiments, a heterologous endonuclease can exhibit reduced nuclease activity (e.g., nuclease deficient or nuclease null) as compared to wild type Un1Cas12f1. The reduced nuclease activity can be at most about 95%, at most about 90%, at most about 80%, at most about 70%, at most about 60%, at most about 50%, at most about 40%, at most about 30%, at most about 20%, at most about 10%, at most about 5%, at most about 1%, at most about 0.5%, at most about 0.1%, or less than that of the wild type Un1Cas12f1.

[0137] In some cases, a Cas protein as provided herein may not be a Cas14 protein. 157

[0138] A Cas protein or fragment or derivative thereof can be from any suitable organism. Non-limiting examples include Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinae spiralis, Streptomyces viridochromo genes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, AlicyclobacHlus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas nap hthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Pseudomonas aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, Acaryochloris marina, Leptotrichia shahii, or Francisella novicida. In some aspects, the organism is Streptococcus pyogenes (S. pyogenes). In some aspects, the organism is Staphylococcus aureus (S. aureus). In some aspects, the organism is Streptococcus thermophilus (S. thermophilus).

[0139] A Cas protein can be derived from a variety of bacterial species including, but not limited to, Veillonella atypical, Fusobacterium nucleatum, Filifactor alocis, Solobacterium moorei, Coprococcus catus, Treponema denticola, Peptoniphilus duerdenii, Catenibacterium mitsuokai, Streptococcus mutans, Listeria innocua, Staphylococcus pseudintermedius, Acidaminococcus intestine, Olsenella uli, Oenococcus kitaharae, Bifidobacterium bifidum, Lactobacillus rhamnosus, Lactobacillus gasseri, Finegoldia magna, Mycoplasma mobile, Mycoplasma gallisepticum, Mycoplasma ovipneumoniae, Mycoplasma canis, Mycoplasma synoviae, Eubacterium rectale, Streptococcus thermophilus, Eubacterium dolichum, Lactobacillus coryniformis subsp. Torquens, Ilyobacter polytropus, Ruminococcus albus, Akkermansia muciniphila, Acidothermus cellulolyticus, Bifidobacterium longum, 158Bifidobacterium dentium, Corynebacterium diphtheria, Elusimicrobium minutum, Nitratifractorsalsuginis, Sphaerochaeta globus, Fibrobacter succinogenes subsp. Succinogenes, Bacteroides fragilis, Capnocytophaga ochracea, Rhodopseudomonas palustris, Prevotella micans, Prevotella ruminicola, Flavobacterium columnare, Aminomonas paucivorans, Rhodospirillum rubrum, Candidatus Puniceispirillum marinum, Verminephrobacter eiseniae, Ralstonia syzygii, Dinoroseobacter shibae, Azospirillum, Nitrobacter hamburgensis, Bradyrhizobium, Wolinellasuccinogenes, Campylobacter jejuni subsp. Jejuni, Helicobacter mustelae, Bacillus cereus, Acidovorax ebreus, Clostridium perfringens, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria meningitidis, Pasteurella multocida subsp. Multocida, Sutterella wadsworthensis, proteobacterium, Legionella pneumophila, Parasutterella excrementihominis, Wolinella succinogenes, or Francisella novicida.

[0140] A Cas protein as used herein can be a wildtype or a modified form of a Cas protein. A Cas protein can be an active variant, inactive variant, or fragment of a wild type or modified Cas protein. A Cas protein can comprise an amino acid change such as a deletion, insertion, substitution, variant, mutation, fusion, chimera, or any combination thereof relative to a wild-type version of the Cas protein. A Cas protein can be a polypeptide with at least or at least about 5%, at least or at least about 10%, at least or at least about 20%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 80%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or 100% sequence identity or sequence similarity to a wild type Cas protein. A Cas protein can be a polypeptide with at most or at most about 5%, at most or at most about 10%, at most or at most about 20%, at most or at most about 30%, at most or at most about 40%, at most or at most about 50%, at most or at most about 60%, at most or at most about 70%, at most or at most about 80%, at most or at most about 90%, or at most or at most about 100% sequence identity and / or sequence similarity to a wild type exemplary Cas protein. Variants or fragments can comprise at least or at least about 5%, at least or at least about 10%, at least or at least about 20%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, 159at least or at least about 60%, at least or at least about 70%, at least or at least about 80%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or 100% sequence identity or sequence similarity to a wild type or modified Cas protein or a portion thereof. Variants or fragments can be targeted to a nucleic acid locus in complex with a guide nucleic acid while lacking nucleic acid cleavage ...

Claims

WHAT IS CLAIMED IS:

1. An engineered gene effector comprising a polypeptide comprising: a first peptide of 75-110 (or 75-95) amino acids in length, wherein the first peptide comprises any one of SEQ ID NOs:3-100, or a sequence at least 85% identical thereto; and a second peptide of 75-110 (or 75-95) amino acids in length and that is heterologous to the first peptide, wherein the second peptide comprises any one of SEQ ID NOs:3-100, or a sequence at least 85% identical thereto, optionally wherein the first peptide is different from the second peptide.

2. The engineered gene effector of claim 1, wherein the first peptide and / or the second peptide is 85 or 108 amino acids in length, optionally wherein the first peptide and the second peptide are each 85 or 108 amino acids in length.

3. The engineered gene effector of claim 1 or 2, wherein the first peptide comprises the any one of SEQ ID NOs: 3-100 with 0-3 amino acid residue mutations, and / or the second peptide comprises the any one of SEQ ID NO: 3-100 with 0-3 amino acid residue mutations, optionally wherein any mutations thereof are conservative substitutions.

4. The engineered gene effector of claim 1 or 2, wherein the first peptide comprises the any one of SEQ ID NO: 3-100, and / or the second peptide comprises the any one of SEQ ID NO: 3-100.

5. The engineered gene effector of any one of the preceding claims, wherein the first peptide is N-terminal to the second peptide, wherein the any one of SEQ ID NOs: 3-100 of the first peptide comprises SEQ ID NOs: 3-100, and wherein the any one of SEQ ID NOs: 3-100 of the second peptide comprises SEQ ID NOs: 3-100.

6. The engineered gene effector of claim 5, wherein the first peptide and the second peptide are according to any one of the paired arrangements of SEQ ID NOs of the first peptide and the second peptide as set forth in Table 4, wherein the first peptide is N-terminal to the second peptide.

7. The engineered gene effector of any one of the preceding claims, wherein the first peptide and the second peptide are linked by a linker, optionally wherein the linker comprises any one or more of SEQ ID NOs: 2211-2221, optionally, wherein the linker comprises SEQ ID NO:2211.5028. The engineered gene effector of any one of the preceding claims, wherein the polypeptide comprises any one of SEQ ID NOs: 115-274, 276-287, 289-367, 370-445, 448- 746, 748-777, 779-929, 931-1007, 1009-1156, 1158-1194, 1196-1288, 1290-1350, 1352-1451, or a sequence at least 85% identical thereto.

9. The engineered gene effector of claim 7, wherein the polypeptide comprises any one of SEQ ID NOs: 115-274, 276-287, 289-367, 370-445, 448-746, 748-777, 779-929, 931-1007, 1009-1156, 1158-1194, 1196-1288, 1290-1350, 1352-1451 with 0-3 amino acid residue mutations, optionally wherein any mutations thereof are conservative substitutions.

10. The engineered gene effector of any one of the preceding claims, wherein the polypeptide comprises any one of SEQ ID NOs: 1085, 122, 1084, 653, 1099, and 1107, or a sequence at least 85% identical thereto, or with 0-3 amino acid residue mutations, optionally wherein any mutations thereof are conservative substitutions.

11. The engineered gene effector of any one of the preceding claims, wherein the engineered gene effector is capable of activating a target gene in a cell when the engineered gene effector is expressed therein and effectively targeted to a locus of the target gene, optionally wherein the target gene is endogenous to the cell.

12. The engineered gene effector of claim 11, wherein the target gene is a silenced gene, optionally wherein the silenced gene is a methylated gene.

13. The engineered gene effector of claim 11 or 12, wherein the engineered gene effector is capable of increasing the expression level of the target gene by, by at least, or by about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 250, 300, 400, 500%, optionally wherein the engineered gene effector is capable of increasing the expression level of the target gene by a percentage in a range defined by any two of the preceding values (e.g., 10-100%, 100-200%, 200-400%, 250-500%, 10-50%, 50-100%, etc.).

14. The engineered gene effector of any one of the preceding claims, wherein the polypeptide is coupled to a heterologous endonuclease, optionally wherein the heterologous endonuclease is a Cas protein.

15. The engineered gene effector of claim 14, wherein the heterologous endonuclease has a length of, of at most, or of about 450, 460, 470, 480, 490, 500, 520, 540, 560, 580, 600, 620, 640, 660, 680, 700 amino acids, optionally wherein the heterologous503endonuclease has a length in a range defined by any two of the preceding values (e.g., 450-700 amino acids, 480-600 amino acids, 500-530 amino acids, 500-600 amino acids, etc.).

16. The engineered gene effector of claim 14 or 15, wherein the heterologous endonuclease comprises the amino acid sequence of any one of SEQ ID NO: 2222-2422, or a sequence at least 85% identical thereto.

17. The engineered gene effector of any one of claims 14-16, wherein the polypeptide is fused to the heterologous endonuclease.

18. The engineered gene effector of claim 17, wherein the polypeptide is fused to the C-terminus of the heterologous endonuclease.

19. An engineered gene effector comprising a polypeptide comprising: a first peptide comprising an amino acid sequence of 75-110 amino acids in length and based on a human or viral transcriptional modulator; and a second peptide comprising an amino acid sequence of 75-110 amino acids in length and based on a human or viral transcriptional modulator, wherein the second peptide is heterologous to the first peptide, wherein the engineered gene effector is capable of activating a target gene in a cell when the engineered gene effector is expressed therein and effectively targeted to a locus of the target gene.

20. The engineered gene effector of any one of the preceding claims, wherein the first peptide and / or the second peptide has a beta factor of about 30 to about 65.

21. The engineered gene effector of any one of the preceding claims, wherein the first peptide and / or the second peptide is enriched for negative electrostatic potential.

22. The engineered gene effector of any one of the preceding claims, wherein the first peptide and / or the second peptide has a negative net charge.

23. The engineered gene effector of any one of the preceding claims, wherein the engineered gene effector is capable of activating a target gene in a cell, wherein the expression level of the target gene is activated via the engineered gene effector persists for a duration of, of about, or of at least 9, 10, 11, 12, 13, 14, 15, 16, 17, 18 days, or more, or optionally the expression level persists for a duration in a range defined by any two of the preceding values (e.g., 9-18 days, 9-14 days, 12-18 days, 14-16 days, etc.).

24. The engineered gene effector of any one of the preceding claims, wherein the engineered gene effector is capable of activating a target gene in a cell, wherein the expression504level of the target gene is increased via the engineered gene effector by at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 5, 10, 20, 30, 40, 50 fold or more compared to a control, or optionally, the expression level of the target gene is increased by a fold amount in a range defined by any two of the preceding values (e.g., 0.1-50 fold, 0.5-10 fold, 1-40 fold 2- 30 fold, etc.).

25. A fusion protein comprising: the engineered gene effector of any one of the preceding claims; and a heterologous endonuclease coupled to the polypeptide, optionally wherein the heterologous endonuclease is a Cas protein.

26. The fusion protein of claim 25, wherein the heterologous endonuclease comprises the amino acid sequence of any one of SEQ ID NO:2222-2422, or a sequence at least 85% identical thereto.

27. The fusion protein of claim 25 or 26, wherein the polypeptide is fused to the heterologous endonuclease.

28. The fusion protein of claim 25, wherein the polypeptide is fused to the C- terminus of the heterologous endonuclease.

29. A polynucleotide comprising a nucleotide sequence encoding the engineered gene effector or the fusion protein of any one of the preceding claims.

30. A vector comprising the polynucleotide of claim 29.

31. A cell comprising the polynucleotide of claim 29 or the vector of claim 30.

32. A system comprising: the engineered gene effector of any one of claims 1-24; a heterologous endonuclease coupled to the polypeptide of the engineered gene effector, optionally wherein the heterologous endonuclease is a Cas protein; and a guide nucleic acid capable of forming a complex with the heterologous endonuclease, wherein the complex exhibits specific binding to a target gene in a cell when the system is expressed therein.

33. The system of claim 32, wherein the heterologous endonuclease comprises the amino acid sequence of any one of SEQ ID NO: 2222-2422, or a sequence at least 85% identical thereto.50534. The system of claim 32 or 33, wherein the polypeptide is fused to the heterologous endonuclease.

35. The system of claim 34, wherein the polypeptide is fused to the C-terminus of the heterologous endonuclease.

36. A combination of polynucleotides encoding the system of any one of claims 32- 35, wherein the combination of polynucleotides is configured to express the heterologous endonuclease coupled to the engineered gene effector and the guide nucleic acid in a cell.

37. A kit comprising the engineered gene effector, fusion protein, combination, system, polynucleotide, vector, and / or cell of any one of the preceding claims.

38. A method of controlling a target gene in a cell, comprising contacting a cell with the engineered gene effector of any one claims 1-24, the fusion protein of any one of claims 25-28, the polynucleotide of claim 29, the vector of claim 30, the system of any one of claims 32-35, or the combination of polynucleotides of claim 36.

39. The method of claim 38, wherein the target gene is endogenous to the cell.

40. The method of claim 38 or 39, wherein the contacting is performed in vitro or ex vivo.

41. A computer-implemented method of producing functional biological sequences, comprising: (a) providing a fitness function trained on a biological dataset comprising functionally defined biological sequences with a fixed length; (b) providing, in a computer, a plurality of different sequences comprising a fixed length, each sequence associated with a temperature and a fitness based on said fitness function, wherein each sequence is associated with a different temperature of a temperature ladder; (c) by said computer, in parallel across the plurality of different sequences: (1) selecting one or more random position(s) for introducing a substitution in one or more sequences of the plurality of different sequences, optionally selecting 1-5 random positions, optionally selecting 1 random position; and for each of the one or more sequences, evaluating a first fitness change due to introduction of the substitution(s) at the one or more randomly selected position(s), and506accepting or rejecting the substitution(s) based on the evaluated first fitness change, and optionally further based on the temperature associated with the sequence; and / or (2) selecting one or more pairs of said plurality of different sequences, each selected pair comprising sequences associated with consecutive temperatures of the temperature ladder, optionally selecting up to 3 pairs of said plurality of different sequences, optionally selecting one pair of said plurality of different sequences; and for each of the selected pairs: selecting one or more domains for swapping between sequences of the selected pair; and evaluating a fitness difference of the sequences of the selected pair due to swapping of the one or more domains, and accepting or rejecting the swapping of the one or more domains between said selected pair based on the fitness difference and the temperature associated with each sequence of said selected pair; and (d) performing (c) iteratively, wherein in each subsequent iteration, the accepted substitution(s) of a preceding iteration and / or the accepted swapping of domains of a preceding iteration are incorporated into the plurality of different sequences, thereby producing one or more functional sequences having a fitness at or above a desired fitness threshold.

42. The method of claim 41, comprising at (c)(1), accepting the substitution(s) at the one or more randomly selected position(s) when the fitness of the sequence after introducing the substitution(s) is greater than the fitness of the sequence before introducing the substitution(s).

43. The method of claim 41 or 42, comprising at (c)(1), accepting or rejecting the substitution(s) at the one or more randomly selected position(s) based on a probability weighted by a ratio of the fitness of the sequence after introducing the substitution(s) to the fitness of the sequence before introducing the substitution(s).

44. The method of any one of claims 41-43, comprising at (c)(1), accepting or rejecting the substitution(s) at the one or more randomly selected position(s) based on a Boltzmann Metropolis-Hastings acceptance criterion, rmh.50745. The method of any one of claims 41-44, wherein the one or more randomly selected positions are selected uniformly across the fixed length.

46. The method of any one of claims 41-45, comprising at (c)(2), accepting swapping of the selected domains between said selected pair when the fitness of the sequence associated with a lower temperature of the consecutive temperatures after swapping is greater than the fitness of the sequence associated with a higher temperature of the consecutive temperatures after swapping, optionally wherein the selected domains comprise the entire sequence of the selected pair.

47. The method of any one of claims 41-46, comprising at (c)(2), accepting swapping of the selected domains between said selected pair based on a probability inversely proportional to a difference between the temperatures associated with each sequence of said pair, and a ratio of the fitness of the pair of sequences after swapping, optionally wherein the selected domains comprise the entire sequence of the selected pair.

48. The method of any one of claims 41-47, comprising at (c)(2), accepting or rejecting swapping of the selected domains between said pair based on a parallel tempering criterion, rre, optionally wherein the selected domains comprise the entire sequence of the selected pair.

49. The method of any one of claims 41-48, wherein (c) comprises: (3) selecting a crossover locus between one or more pairs of said plurality of different sequences; and for each of the one or more pairs in which the crossover locus is selected, evaluating a second fitness change for each sequence of the selected pair due to crossing over at the crossover locus, and accepting or rejecting crossing over at the selected crossover locus based on the second fitness changes and the temperature associated with each sequence of said selected pair.

50. The method of claim 49, comprising accepting or rejecting crossing over at the selected crossover locus based on a probability weighted by a ratio of the second fitness changes of each of the sequences of said selected pair.

51. The method of any one of claims 41-50, wherein at least one of the produced one or more functional sequences has a fitness based on the fitness function that is greater than the fitness of each of the plurality of different sequences before any iteration of (c).50852. The method of any one of claims 41-51, wherein the desired fitness threshold is based on the fitness associated with the corresponding sequence of the plurality of different sequences in (b), optionally wherein the desired fitness threshold is based on the maximum fitness among the plurality of different sequences.

53. The method of any one of claims 41-52, wherein the plurality of different sequences in (b) comprises a plurality of different naturally occurring sequences.

54. A computer-implemented method of producing functional biological sequences, comprising: (a) evaluating, by a computer, a sequence of a plurality of different sequences comprising a fixed length based on a fitness function trained on a biological dataset comprising functionally defined biological sequences with said fixed length; (b) substituting, by said computer, one or more random residues in said sequence, to generate a mutated sequence; (c) evaluating, by said computer, said mutated sequence based on said fitness function; and (d) collecting, by said computer, functional sequences accepted by said fitness function.

55. The computer-implemented method of claim 54, further comprising randomly swapping, by the computer, one or more subsequences from the mutated sequence and a different sequence of the plurality of sequences.

56. The computer-implemented method of claim 54 or 55, wherein said fitness function comprises a threshold selected from the group consisting of: a binary threshold, a numerical threshold, a multiclass threshold, a confidence threshold, a decision threshold, and any combination thereof.

57. The computer-implemented method of claim 56, wherein functional sequences are accepted by said fitness function when a fitness score assigned to said functional sequences by said fitness function exceeds said threshold.

58. The computer-implemented method of any one of claims 54-57, wherein the plurality of different sequences in (a) comprises a plurality of different naturally occurring sequences or different random sequences.

59. The computer-implemented method of any one of claims 41-58, wherein said functionally defined biological sequences comprise amino acid sequences or nucleotide509sequences of, or encoding, a protein or a peptide, optionally wherein the functionally defined biological sequences comprise transcriptional activators, further optionally wherein the functionally defined biological sequences comprise engineered gene effectors.

60. The computer-implemented method of claim 59, wherein said protein or said peptide is an epigenetic modulator, a transcription factor, an enzyme, a nuclease, an agonist, an antagonist, a regulator, or an inhibitor.

61. The computer-implemented method of any one of claims 41-60, wherein said functionally defined biological sequences comprise amino acid sequences or nucleotide sequences.

62. The computer-implemented method of any one of claims 41-61, wherein said functionally defined biological sequences comprise amino acid sequences and further wherein said fixed length is at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 105, at least 110, at least 115, at least 120, at least 125, at least 130, at least 140, at least 150, or at least 200 amino acids, or at most 500, at most 300, at most 200, at most 150, at most 120, at most 100, at most 95, at most 90, at most 85, at most 75, or at most 70 amino acids, or optionally wherein said fixed length is in a range defined by any two of the preceding values (e.g., 30-500 amino acids, 50-300 amino acids, 50-200 amino acids, 75-95 amino acids, 80-90 amino acids, or 60-150 amino acids, etc.).

63. The computer-implemented method of any one of claims 41-62, wherein said fitness function is based on one or more machine learning model, wherein said machine learning model is selected from the group consisting of: a supervised machine learning model, an unsupervised machine learning model, a reinforcement learning model, a deep learning model, a transfer learning model, and any combination thereof.

64. The computer-implemented method of claim 63, wherein said one or more machine learning models is selected from the group consisting of: a classification model, a regression model, a decision tree model, a convolutional neural network (CNN), a recurrent neural network (RNN), Extreme Gradient Boosting (XGBoost), long short-term memory network, generative adversarial network (GAN), an autoencoder, a transformer network, evolutionary Monte Carlo, and any combination thereof.51065. The computer-implemented method of claim 64, wherein the fitness function is based on an ensembled model comprising a decision tree model and a convolutional neural network, optionally wherein the ensembled model comprises CNN and XGBoost.

66. The computer-implemented method of any one of the preceding claims, comprising evaluating, by the computer, the biological dataset to generate the fitness function, the evaluating comprising: generating sequence embeddings from a large protein language model (LPLM) based on the biological dataset; and training a machine learning model with the generated sequence embeddings as input, optionally wherein the LPLM comprises an evolutionary scale modeling (ESM) language model, optionally wherein the LPLM comprises ESM-2.

67. The computer-implemented method of claim 66, wherein the machine learning model comprises an ensemble model of two or more different models.

68. The computer-implemented method of claim 67, wherein the ensemble model comprises CNN and XGBoost.

69. The computer-implemented method of any one of the preceding claims, wherein the plurality of different sequences comprise a plurality of different random sequences.

70. The computer-implemented method of any one of the preceding claims, wherein the biological dataset comprises at most or on the order of about 105biological sequences.

71. The computer-implemented method of any one of the preceding claims, wherein at most 5% of the functionally defined biological sequences comprise functional sequences.

72. A computer-implemented system comprising a computing device comprising at least one processor and instructions executable by the at least one processor to provide an application, comprising: (a) a software module configured to evaluate, by a computer, a sequence of a plurality of different sequences comprising a fixed length based on a fitness function trained on a biological dataset comprising functionally defined biological sequences with said fixed length;511(b) a software module configured to substitute, by said computer, one or more random residues in said sequence, to generate a mutated sequence; (c) a software module configured to evaluate, by said computer, said mutated sequence based on said fitness function; and (d) a software module configured to collect, by said computer, functional sequences accepted by said fitness function.

73. A non-transitory computer-readable medium having stored thereon computer- readable instructions that, when executed by a processor, cause the processor to execute a method comprising: (a) evaluating a sequence of a plurality of different sequences comprising a fixed length based on a fitness function trained on a biological dataset comprising functionally defined biological sequences with said fixed length; (b) substituting one or more random residues in said sequence, to generate a mutated sequence; (c) evaluating said mutated sequence based on said fitness function; and (d) collecting functional sequences accepted by said fitness function.

74. A computer-implemented system comprising: a computing device comprising: at least one processor and instructions executable by the at least one processor to provide an application comprising one or more software modules for performing the method of any one of claims 41-71.

75. A non-transitory computer-readable medium having stored thereon computer- readable instructions that, when executed by a processor, cause the processor to execute the method of any one of claims 41-71.

76. An engineered gene effector comprising one or more of the polypeptides produced by the method of any one of claims 41-71, or a sequence at least 85% identical thereto.

77. An engineered gene effector comprising a polypeptide of 85 amino acids in length that comprises any one of SEQ ID NOs: 1495, 1592, 1595, 1634, 1654, 1665, 1677, 1686, 1689, 1716, or a sequence at least 85% identical thereto.

78. The engineered gene effector of claim 76 or 77, wherein the engineered gene effector is capable of activating a target gene in a cell when the engineered gene effector is512expressed therein and effectively targeted to a locus of the target gene, optionally wherein the target gene is endogenous to the cell.

79. The engineered gene effector of any one of claims 76-78, wherein the polypeptide is coupled to a heterologous endonuclease, optionally wherein the heterologous endonuclease is a Cas protein.

80. The engineered gene effector of claim 79, wherein the heterologous endonuclease comprises the amino acid sequence of any one of SEQ ID NO: 2222-2422, or a sequence at least 85% identical thereto.

81. The engineered gene effector of claim 79 or 80, wherein the polypeptide is fused to the heterologous endonuclease, optionally wherein the polypeptide is fused to the C- terminus of the heterologous endonuclease.

82. A fusion protein comprising: the engineered gene effector of any one of claims 76-81; and a heterologous endonuclease coupled to the polypeptide, optionally wherein the heterologous endonuclease is a Cas protein.

83. A polynucleotide comprising a nucleotide sequence encoding the engineered gene effector or the fusion protein of any one of claims 76-81.

84. A vector comprising the polynucleotide of claim 83.

85. A cell comprising the polynucleotide of claim 83 or the vector of claim 84.

86. A system comprising: the engineered gene effector of any one of any one of claims 76-81; a heterologous endonuclease coupled to the engineered gene effector, optionally wherein the heterologous endonuclease is a Cas protein; and a guide nucleic acid capable of forming a complex with the heterologous endonuclease, wherein the complex exhibits specific binding to a target gene in a cell when the system is expressed therein.

87. A combination of polynucleotides encoding the system of claim 86, wherein the combination of polynucleotides is configured to express the heterologous endonuclease coupled to the engineered gene effector and the guide nucleic acid in a cell.

88. A method of controlling a target gene in a cell, comprising contacting a cell with the system of claim 86, or the combination of polynucleotides of claim 87.51389. The method of claim 88, wherein the target gene is endogenous to the cell.

90. The method of claim 88 or 89, wherein the contacting is performed in vitro or ex vivo.

91. A computer-implemented system comprising a computing device comprising at least one processor and instructions executable by the at least one processor to provide an application, comprising: (a) a software module configured to provide, by a computer, a fitness function trained on a biological dataset comprising functionally defined biological sequences with a fixed length; (b) a software module configured to provide, by said computer, a plurality of different sequences comprising a fixed length, each sequence associated with a temperature and a fitness based on said fitness function, wherein each sequence is associated with a different temperature of a temperature ladder; (c) in parallel across the plurality of different sequences: (1) a software module configured to select, by said computer, one or more random position(s) for introducing a substitution in one or more sequences of the plurality of different sequences, optionally selecting 1-5 random positions, optionally selecting 1 random position; and for each of the one or more sequences, evaluate a first fitness change due to introduction of the substitution(s) at the one or more randomly selected position(s), and accept or reject the substitution(s) based on the evaluated first fitness change, and optionally further based on the temperature associated with the sequence; and / or (2) a software module configured to select, by said computer, one or more pairs of said plurality of different sequences, each selected pair comprising sequences associated with consecutive temperatures of the temperature ladder, optionally selecting up to 3 pairs of said plurality of different sequences, optionally selecting one pair of said plurality of different sequences; and for each of the selected pairs: select one or more domains for swapping between sequences of the selected pair; and514evaluate a fitness difference of the sequences of the selected pair due to swapping of the one or more domains, and accept or reject the swapping of the one or more domains between said selected pair based on the fitness difference and the temperature associated with each sequence of said selected pair; and (d) a software module configured to perform, by said computer, (c) iteratively, wherein in each subsequent iteration, the accepted substitution(s) of a preceding iteration and / or the accepted swapping of domains of a preceding iteration are incorporated into the plurality of different sequences, thereby producing one or more functional sequences having a fitness at or above a desired fitness threshold.

92. An engineered gene effector comprising a polypeptide comprising any one of SEQ ID NOs: 115-274, 276-287, 289-367, 370-445, 448-746, 748-777, 779-929, 931-1007, 1009-1156, 1158-1194, 1196-1288, 1290-1350, 1352-1451, or a sequence at least 85% identical thereto, or with 0-3 amino acid residue mutations, optionally wherein any mutations thereof are conservative substitutions.

93. An engineered gene effector comprising a polypeptide comprising any one of SEQ ID NOs: 1085, 122, 1084, 653, 1099, and 1107, or a sequence at least 85% identical thereto, or with 0-3 amino acid residue mutations, optionally wherein any mutations thereof are conservative substitutions.515