Engineering of systems, methods and optimized guide compositions for sequence manipulation
Patent Information
- Application Number
- JP2024138375
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2013-06-17
- Filing Date
- 2024-08-20
- Publication Date
- 2025-05-14
- Estimated Expiration
- 2033-12-12
AI Technical Summary
There is a need for alternative and robust genome engineering techniques that facilitate targeting of multiple locations within the nuclear genome, as existing methods like designer zinc fingers and homing meganucleases are limited in versatility and scalability.
The development of CRISPR/Cas systems that utilize a short RNA molecule to program specific DNA targeting without the need for customized proteins, combined with vector systems that include regulatory elements and CRISPR enzymes for precise genome editing in eukaryotic cells.
This approach simplifies genome editing methodologies, enabling efficient and accurate targeting of genetic factors associated with biological functions and diseases, accelerating the classification and mapping of genetic factors.
Abstract
Description
[Technical Field]
[0001] Related Applications and Incorporation by Reference This application claims priority to U.S. Provisional Patent Application No. 61 / 836,127, filed June 17, 2013, entitled ENGINEERING OF SYSTEMS, METHODS AND OPTIMIZED COMPOSITIONS FOR SEQUENCE MANIPULATION. This application also claims priority to U.S. Provisional Patent Applications Nos. 61 / 758,468; 61 / 769,046; 61 / 802,174; 61 / 806,375; 61 / 814,263; 61 / 819,803 and 61 / 828,130, each entitled ENGINEERING AND OPTIMIZATION OF SYSTEMS, METHODS AND COMPOSITIONS FOR SEQUENCE MANIPULATION, filed on January 30, 2013; February 25, 2013; March 15, 2013; March 28, 2013; April 20, 2013; May 6, 2013 and May 28, 2013, respectively. Priority is also claimed to U.S. Provisional Patent Applications Nos. 61 / 736,527 and 61 / 748,427, both entitled SYSTEMS METHODS AND COMPOSITIONS FOR SEQUENCE MANIPULATION, filed December 12, 2012, and January 2, 2013, respectively. Priority is also claimed to U.S. Provisional Patent Applications Nos. 61 / 791,409 and 61 / 835,931, both entitled BI-2011 / 008 / 44790.02.2003 and BI-2011 / 008 / 44790.03.2003, filed March 15, 2013, and June 17, 2013, respectively. Reference is also made to U.S. Provisional Patent Applications Nos. 61 / 835,936, 61 / 836,101, 61 / 836,080, 61 / 836,123, and 61 / 835,973, each filed on June 17, 2013.
[0002] The above applications, and all documents cited in those applications or during their prosecution ("application cited documents"), and all documents cited or referenced in those application cited documents, as well as all documents cited or referenced herein ("herein cited documents"), and all documents cited or referenced in the herein cited documents, together with any manufacturer's instructions, manuals, product specifications, and product sheets for any products mentioned herein or in any document incorporated by reference herein, are hereby incorporated by reference and may be used in the practice of this invention. More specifically, all references are incorporated by reference to the same extent as if each individual document were individually and specifically indicated to be incorporated by reference.
[0003] The present invention relates generally to systems, methods, and compositions used in the control of gene expression, including sequence targeting, e.g., genomic perturbation or gene editing, which may use vector systems related to Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) and its components.
[0004] STATEMENT REGARDING FEDERALLY FUNDED RESEARCH This invention was made with government support funded by the National Institutes of Health, NIH Pioneer Award DP1MH100706. The U.S. government has certain rights in this invention. [Background technology]
[0005] Recent advances in genome sequencing technology and analytical methods have significantly accelerated the ability to classify and map genetic factors related to a wide range of biological functions and diseases.Accurate genome targeting technology is needed to enable the systematic reverse engineering of causal gene mutations by enabling the selective perturbation of individual genetic elements, and to advance synthetic biology, biotechnology and pharmaceutical applications.Genome editing technology, such as designer zinc finger, transcription activator-like effector (TALE), or homing meganuclease, can be used to produce targeted genome perturbations, but there is still a need for new genome engineering technology that is inexpensive, easy to set up, scalable, and easy to target multiple locations in eukaryotic genomes. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] U.S. Patent No. 4,873,316 Summary of the Invention [Problem to be solved by the invention]
[0007] There is an urgent need for alternative, robust systems and technologies for versatile sequence targeting. The present invention addresses this need and provides related advantages. CRISPR / Cas or CRISPR-Cas systems (both terms are used interchangeably throughout this application) do not require the generation of customized proteins to target specific sequences, but rather allow a single Cas enzyme to be programmed by a short RNA molecule to recognize a specific DNA target; in other words, the Cas enzyme can be recruited to a specific DNA target using the short RNA molecule. The addition of CRISPR-Cas systems to the repertoire of genome sequencing technologies and analytical methods significantly simplifies methodology and accelerates the ability to classify and map genetic factors associated with a diverse range of biological functions and diseases. To effectively utilize CRISPR-Cas systems for genome editing without adverse effects, it is important to understand the engineering and optimization aspects of these genome engineering tools, which are an embodiment of the claimed invention. [Means for solving the problem]
[0008] In one aspect, the present invention provides a vector system, comprising one or more vectors.In some embodiments, the system comprises: (a) a first regulatory element, which is operably linked to a tracr mate sequence and one or more insertion sites for inserting one or more guide sequences upstream of the tracr mate sequence (the guide sequence, when expressed, directs the sequence-specific binding of CRISPR complex to the target sequence in cells, for example, eukaryotic cells, and the CRISPR complex comprises (1) the guide sequence hybridized to the target sequence, and (2) the tracr mate sequence hybridized to the tracr sequence; and (b) a second regulatory element, which is operably linked to the enzyme coding sequence encoding the CRISPR enzyme, which comprises a nuclear localization sequence; component (a) and (b) are on the same or different vectors of the system.In some embodiments, component (a) further comprises the tracr sequence downstream of the tracr mate sequence under the control of the first regulatory element. In some embodiments, component (a) further comprises two or more guide sequences operably linked to the first regulatory element, each of which, when expressed, directs the sequence-specific binding of the CRISPR complex to a different target sequence in eukaryotic cells. In some embodiments, the system comprises a third regulatory element, such as a tracr sequence under the control of a polymerase III promoter. In some embodiments, the tracr sequence exhibits at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence complementarity along the length of the tracr mate sequence when optimally aligned. In some embodiments, the CRISPR complex comprises one or more nuclear localization sequences of sufficient strength to drive the accumulation of a detectable amount of the CRISPR complex in the nucleus of a eukaryotic cell. Without being bound by theory, nuclear localization sequences are not required for CRISPR complex activity in eukaryotic organisms, but it is believed that including such sequences improves the activity of the system, particularly with respect to targeting nucleic acid molecules in the nucleus. In some embodiments, the CRISPR enzyme is a type II CRISPR enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme.In some embodiments, the Cas9 enzyme is S. pneumoniae, S. pyogenes, or S. thermophilus Cas9, and may include mutant Cas9s from these organisms. The enzyme may be a Cas9 homolog or ortholog. In some embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In some embodiments, the CRISPR enzyme directs one or two strand cleavage at the location of the target sequence. In some embodiments, the CRISPR enzyme lacks DNA strand cleavage activity. In some embodiments, the first regulatory element is a polymerase III promoter. In some embodiments, the second regulatory element is a polymerase II promoter. In some embodiments, the guide sequence is at least 15, 16, 17, 18, 19, 20, or 25 nucleotides, or 10-30, 15-25, or 15-20 nucleotides in length. Generally, and throughout this specification, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules containing one or more free ends but no free ends (e.g., circular); nucleic acid molecules containing DNA, RNA, or both; and various other polynucleotides known in the art. One type of vector is a "plasmid," which refers to a circular double-stranded DNA loop into which additional DNA segments can be inserted, for example, by standard molecular cloning techniques. Another type of vector is a viral vector, in which viral-derived DNA or RNA sequences are present in the vector for packaging into a virus (e.g., retrovirus, replication-deficient retrovirus, adenovirus, replication-deficient adenovirus, and adeno-associated virus). Viral vectors also include polynucleotides carried by viruses for transfection into host cells. Some vectors can autonomously replicate in the host cell into which they are introduced (e.g., bacterial vectors with a bacterial origin of replication and episomal mammalian vectors).Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome. Moreover, certain vectors are capable of directing the expression of genes to which they are operatively linked. Such vectors are referred to herein as "expression vectors." Common expression vectors of utility in recombinant DNA techniques are often in the form of plasmids.
[0009] A recombinant expression vector can contain a nucleic acid of the invention in a form suitable for expression of the nucleic acid in a host cell, meaning that the recombinant expression vector can be selected based on the host cell to be used for expression and contains one or more regulatory elements operably linked to the nucleic acid sequence to be expressed. In the context of a recombinant expression vector, "operably linked" means that the nucleotide sequence of interest is linked to a regulatory element in a manner that allows for expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell).
[0010] The term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, e.g., polyadenylation signals and poly-U sequences). Such regulatory elements are described, for example, in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990). Regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells and those that direct expression of a nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can direct expression primarily in a desired target tissue, e.g., muscle, nerve cells, bone, skin, blood, specific organs (e.g., liver, pancreas), or a particular cell type (e.g., lymphocytes). The regulatory element can also direct expression in a time-dependent manner, such as a cell cycle-dependent or developmental stage-dependent manner, which may or may not be tissue- or cell type-specific. In some embodiments, the vector contains one or more PolIII promoters (e.g., 1, 2, 3, 4, 5, or more PolIII promoters), one or more PolII promoters (e.g., 1, 2, 3, 4, 5, or more PolII promoters), one or more PolI promoters (e.g., 1, 2, 3, 4, 5, or more PolI promoters), or a combination thereof. Examples of PolIII promoters include, but are not limited to, U6 and H1 promoters.Examples of pol II promoters include, but are not limited to, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with the RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with the CMV enhancer) [see, e.g., Boshart et al., Cell, 41:521-530 (1985)], the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter. The term "regulatory element" also encompasses enhancer elements, such as the WPRE; the CMV enhancer; the R-U5' segment in the LTR of HTLV-I (Mol. Cell. Biol., Vol. 8(1), p. 466-472, 1988); the SV40 enhancer; and the intron sequence between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), p. 1527-31, 1981). Those skilled in the art will recognize that the design of the expression vector may depend on factors such as the choice of host cell to be transformed, the desired expression level, and the like. The vector can be introduced into a host cell to produce a protein, including a transcript, a fusion protein, or a peptide, encoded by the nucleic acid described herein (e.g., a clustered regularly interspaced short repeat (CRISPR) transcript, protein, enzyme, mutant thereof, fusion protein thereof, etc.).
[0011] Advantageous vectors include lentiviruses and adeno-associated viruses, and such vector types can also be selected for targeting specific cell types.
[0012] In one aspect, the present invention provides a vector comprising a regulatory element operably linked to an enzyme coding sequence encoding a CRISPR enzyme comprising one or more nuclear localization sequences. In some embodiments, the regulatory element drives transcription of the CRISPR enzyme in a eukaryotic cell such that the CRISPR enzyme accumulates in detectable amounts in the nucleus of the eukaryotic cell. In some embodiments, the regulatory element is a polymerase II promoter. In some embodiments, the CRISPR enzyme is a type II CRISPR enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme. In some embodiments, In some embodiments, the Cas9 enzyme is S. pneumoniae, S. pyogenes, or S. thermophilus Cas9, and may include mutant Cas9s derived from these organisms. In some embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In some embodiments, the CRISPR enzyme directs cleavage of one or two strands at the location of the target sequence. In some embodiments, the CRISPR enzyme lacks DNA strand cleavage activity.
[0013] In one aspect, the present invention provides a CRISPR enzyme comprising one or more nuclear localization sequences of sufficient strength to drive the accumulation of a detectable amount of the CRISPR enzyme in the nucleus of a eukaryotic cell. In some embodiments, the CRISPR enzyme is a type II CRISPR enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme. In some embodiments, the Cas9 enzyme is Streptococcus pneumoniae (S. pneumoniae), Streptococcus pyogenes (S. pyogenes), or S. thermophilus (S. thermophilus) Cas9, and may include mutant Cas9s derived from these organisms. The enzyme may be a Cas9 homolog or ortholog. In some embodiments, the CRISPR enzyme lacks the ability to cleave one or more strands of the target sequence to which it binds.
[0014] In one aspect, the present invention provides a eukaryotic host cell comprising: (a) a first regulatory element operably linked to a tracr mate sequence and one or more insertion sites for inserting one or more guide sequences upstream of the tracr mate sequence (the guide sequence, when expressed, directs sequence-specific binding of a CRISPR complex to a target sequence in a eukaryotic cell, the CRISPR complex comprising a CRISPR enzyme complexed with (1) the guide sequence hybridized to the target sequence and (2) the tracr mate sequence hybridized to the tracr sequence); and / or (b) a second regulatory element operably linked to an enzyme-coding sequence encoding the CRISPR enzyme comprising a nuclear localization sequence. In some embodiments, the host cell comprises components (a) and (b). In some embodiments, components (a), (b), or (a) and (b) are stably integrated into the genome of the host eukaryotic cell. In some embodiments, component (a) further comprises a tracr sequence downstream of the tracr mate sequence under the control of the first regulatory element. In some embodiments, component (a) further comprises two or more guide sequences operably linked to the first regulatory element, each of which, when expressed, directs sequence-specific binding of the CRISPR complex to a different target sequence in the eukaryotic cell. In some embodiments, the eukaryotic host cell further comprises a third regulatory element operably linked to the tracr sequence, such as a polymerase III promoter. In some embodiments, the tracr sequence exhibits at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence complementarity along the length of the tracr mate sequence when optimally aligned. In some embodiments, the CRISPR enzyme comprises one or more nuclear localization sequences strong enough to drive the accumulation of a detectable amount of the CRISPR enzyme in the nucleus of the eukaryotic cell. In some embodiments, the CRISPR enzyme is a type II CRISPR enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme.In some embodiments, the Cas9 enzyme is S. pneumoniae, S. pyogenes, or S. thermophilus Cas9, and may include mutant Cas9s from these organisms. The enzyme may be a Cas9 homolog or ortholog. In some embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In some embodiments, the CRISPR enzyme directs one or two strand cleavage at the location of the target sequence. In some embodiments, the CRISPR enzyme lacks DNA strand cleavage activity. In some embodiments, the first regulatory element is a polymerase III promoter. In some embodiments, the second regulatory element is a polymerase II promoter. In some embodiments, the guide sequence is at least 15, 16, 17, 18, 19, 20, or 25 nucleotides, or 10-30, 15-25, or 15-20 nucleotides in length. In one aspect, the invention provides a non-human eukaryotic organism, preferably a multicellular eukaryotic organism, comprising a eukaryotic host cell according to any of the described embodiments. In another aspect, the invention provides a eukaryotic organism, preferably a multicellular eukaryotic organism, comprising a eukaryotic host cell according to any of the described embodiments. In some embodiments of these aspects, the organism may be an animal, e.g., a mammal. The organism may also be an arthropod, e.g., an insect. The organism may also be a plant. Furthermore, the organism may be a fungus.
[0015] In one aspect, the present invention provides a kit comprising one or more of the components described herein. In some embodiments, the kit comprises a vector system and instructions for use of the kit. In some embodiments, the vector system comprises: (a) a first regulatory element operably linked to a tracr mate sequence and one or more insertion sites for inserting one or more guide sequences upstream of the tracr mate sequence (the guide sequence, when expressed, directs sequence-specific binding of a CRISPR complex to a target sequence in a eukaryotic cell, the CRISPR complex comprising (1) a guide sequence hybridized to the target sequence, and (2) a CRISPR enzyme complexed with the tracr mate sequence hybridized to the tracr sequence); and / or (b) a second regulatory element operably linked to an enzyme coding sequence encoding the CRISPR enzyme comprising a nuclear localization sequence. In some embodiments, the kit comprises components (a) and (b) on the same or different vectors of the system. In some embodiments, component (a) further comprises a tracr sequence downstream of the tracr mate sequence under the control of the first regulatory element. In some embodiments, component (a) further comprises two or more guide sequences operably linked to the first regulatory element, each of which, when expressed, directs sequence-specific binding of the CRISPR complex to a different target sequence in a eukaryotic cell. In some embodiments, the system further comprises a third regulatory element operably linked to the tracr sequence, such as a polymerase III promoter. In some embodiments, the tracr sequence exhibits at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence complementarity along the length of the tracr mate sequence when optimally aligned. In some embodiments, the CRISPR enzyme comprises one or more nuclear localization sequences strong enough to drive the accumulation of a detectable amount of the CRISPR enzyme in the nucleus of a eukaryotic cell. In some embodiments, the CRISPR enzyme is a type II CRISPR system enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme.In some embodiments, the Cas9 enzyme is S. pneumoniae, S. pyogenes, or S. thermophilus Cas9, and may include mutant Cas9s from these organisms. The enzyme may be a Cas9 homolog or ortholog. In some embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In some embodiments, the CRISPR enzyme directs one or two strand cleavage at the location of the target sequence. In some embodiments, the CRISPR enzyme lacks DNA strand cleavage activity. In some embodiments, the first regulatory element is a polymerase III promoter. In some embodiments, the second regulatory element is a polymerase II promoter. In some embodiments, the guide sequence is at least 15, 16, 17, 18, 19, 20, or 25 nucleotides, or 10-30, 15-25, or 15-20 nucleotides in length.
[0016] In one aspect, the present invention provides a method for modifying a target polynucleotide in a eukaryotic cell. In some embodiments, the method comprises binding a CRISPR complex to a target polynucleotide to cause cleavage of the target polynucleotide, thereby modifying the target polynucleotide, wherein the CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence hybridized to a target sequence within the target polynucleotide, the guide sequence in turn being bound to a tracr mate sequence hybridized to a tracr sequence. In some embodiments, the cleavage comprises cleaving one or both strands at the location of the target sequence by the CRISPR enzyme. In some embodiments, the cleavage results in reduced transcription of the target gene. In some embodiments, the method further comprises repairing the cleaved target polynucleotide by homologous recombination with an exogenous template polynucleotide, wherein the repair results in a mutation comprising an insertion, deletion, or substitution of one or more nucleotides in the target polynucleotide. In some embodiments, the mutation results in one or more amino acid changes in a protein expressed from a gene comprising the target sequence. In some embodiments, the method further comprises delivering one or more vectors to the eukaryotic cell, wherein the one or more vectors drive expression of one or more of a CRISPR enzyme, a guide sequence linked to a tracr mate sequence, and a tracr sequence. In some embodiments, the vectors are delivered into the eukaryotic cell in a subject. In some embodiments, the modification is performed in the eukaryotic cell in cell culture. In some embodiments, the method further comprises isolating the eukaryotic cell from a subject before the modification. In some embodiments, the method further comprises returning the eukaryotic cell and / or cells derived therefrom to the subject.
[0017] In one aspect, the present invention provides a method for modifying the expression of polynucleotide in eukaryotic cells.In some embodiments, the method comprises: allowing a CRISPR complex to bind to polynucleotide, whereby said binding causes the expression of said polynucleotide to increase or decrease; the CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence that hybridizes to a target sequence in said polynucleotide, and said guide sequence is then linked to a tracr mate sequence that hybridizes to a tracr sequence.In some embodiments, the method further comprises delivering one or more vectors into said eukaryotic cell, and said one or more vectors drive the expression of one or more of CRISPR enzyme, the guide sequence that is linked to a tracr mate sequence, and the tracr sequence.
[0018] In one aspect, the present invention provides a method for generating a model eukaryotic cell containing a mutated disease gene. In some embodiments, the disease gene is any gene associated with an increased risk of having or developing a disease. In some embodiments, the method includes (a) introducing one or more vectors into a eukaryotic cell, where the one or more vectors drive expression of one or more of a CRISPR enzyme, a guide sequence linked to a tracr mate sequence, and a tracr sequence; and (b) binding a CRISPR complex to a target polynucleotide to cause cleavage of the target polynucleotide within the disease gene, where the CRISPR complex comprises (1) a guide sequence hybridized to a target sequence within the target polynucleotide, and (2) a CRISPR enzyme complexed with a tracr mate sequence hybridized to a tracr sequence, thereby generating a model eukaryotic cell containing a mutated disease gene. In some embodiments, the cleavage includes cleavage of one or both strands at the location of the target sequence by the CRISPR enzyme. In some embodiments, the cleavage results in decreased transcription of the target gene. In some embodiments, the method further comprises repairing the cleaved target polynucleotide by homologous recombination with an exogenous template polynucleotide, wherein the repair results in a mutation comprising an insertion, deletion, or substitution of one or more nucleotides in the target polynucleotide. In some embodiments, the mutation results in one or more amino acid changes in a protein expressed from a gene comprising the target sequence.
[0019] In one aspect, the present invention provides a method for developing a bioactive agent that modulates a cell signaling event associated with a disease gene. In some embodiments, the disease gene is any gene associated with an increased risk of having or developing a disease. In some embodiments, the method includes: (a) contacting a test compound with a model cell of any one of the described embodiments; and (b) detecting a change in a readout that indicates a reduction or increase in a cell signaling event associated with the mutation in the disease gene, thereby developing the bioactive agent that modulates the cell signaling event associated with the disease gene.
[0020] In one aspect, the present invention provides a recombinant polynucleotide comprising a guide sequence upstream of a tracr mate sequence, wherein the guide sequence, when expressed, directs the sequence-specific binding of a CRISPR complex to the corresponding target sequence present in eukaryotic cells.In some embodiments, the target sequence is a viral sequence present in eukaryotic cells.In some embodiments, the target sequence is a proto-oncogene or an oncogene.
[0021] In one aspect, the present invention provides a method for selecting one or more prokaryotic cells by introducing one or more mutations into a gene in one or more prokaryotic cells, the method comprising: introducing one or more vectors into the prokaryotic cells (the one or more vectors drive expression of one or more of a CRISPR enzyme, a guide sequence linked to a tracr mate sequence, a tracr sequence, and an editing template; the editing template comprises one or more mutations that prevent CRISPR enzyme cleavage); allowing the editing template to homologously recombine with a target polynucleotide in the cell to be selected; and allowing a CRISPR complex to bind to the target polynucleotide to cause cleavage of the target polynucleotide within the gene (the CRISPR complex comprises (1) a guide sequence hybridized to a target sequence within the target polynucleotide, and (2) a tracr mate sequence hybridized to the tracr sequence, wherein binding of the CRISPR complex to the target polynucleotide induces cell death), thereby enabling selection of one or more prokaryotic cells into which one or more mutations have been introduced. In a preferred embodiment, the CRISPR enzyme is Cas9. In another embodiment of the invention, the cells to be selected can be eukaryotic cells. This embodiment of the invention allows for the selection of defined cells without requiring a selectable marker or a two-step process that may involve a counterselection system.
[0022] In some embodiments, the invention provides a non-naturally occurring or engineered composition comprising a CRISPR-Cas system chimeric RNA (chiRNA) polynucleotide sequence, the polynucleotide sequence comprising (a) a guide sequence capable of hybridizing to a target sequence in a eukaryotic cell, (b) a tracr mate sequence, and (c) a tracr sequence, wherein (a), (b), and (c) are arranged in a 5' to 3' orientation, such that when transcribed, the guide sequence hybridizes to the tracr mate sequence vs. the tracr sequence, and the guide sequence directs sequence-specific binding of a CRISPR complex to the target sequence, the CRISPR complex comprising: (1) the guide sequence hybridized to the target sequence; and (2) the tracr mate sequence hybridized to the tracr sequence. or I. a first regulatory element operably linked to a CRISPR-Cas system chimeric RNA (chiRNA) polynucleotide sequence, the polynucleotide sequence comprising (a) one or more guide sequences capable of hybridizing to one or more target sequences in a eukaryotic cell, (b) a tracr mate sequence, and (c) one or more tracr sequences; and II. a second regulatory element operably linked to an enzyme coding sequence encoding a CRISPR enzyme comprising at least one or more nuclear localization sequences, wherein (a), (b), and (c) are located within a 5' to 3' sequence. and components I and II are arranged in a direction perpendicular to the tracr sequence, and components I and II are on the same or different vectors of the system, and when transcribed, the tracr mate sequence hybridizes to the tracr sequence and the guide sequence directs sequence-specific binding of the CRISPR complex to the target sequence, and the CRISPR complex is a CRISPR enzyme system encoded by a vector system comprising one or more vectors containing a CRISPR enzyme complexed with (1) the guide sequence hybridized to the target sequence and (2) the tracr mate sequence hybridized to the tracr sequence; or I.(a) a cell and (b) a first regulatory element operably linked to at least one or more tracr mate sequences; II. a second regulatory element operably linked to an enzyme coding sequence encoding a CRISPR enzyme; and III. a third regulatory element operably linked to a tracr sequence; wherein components I, II, and III are on the same or different vectors of the system, and when transcribed, the tracr mate sequence hybridizes to the tracr sequence, and the guide sequence directs sequence-specific binding of the CRISPR complex to the target sequence, the CRISPR complex comprising (1) the guide sequence hybridized to the target sequence, and (2) the CRISPR enzyme complexed with the tracr mate sequence hybridized to the tracr sequence; and wherein in the multiplexed system, multiple guide sequences and a single tracr sequence are used. A multiplexed CRISPR enzyme system (one or more of the guide, tracr, and tracr mate sequences are modified to improve stability) is provided, encoded by a vector system comprising one or more vectors.
[0023] In some embodiments of the present invention, the modification includes an engineered secondary structure. For example, the modification can include reducing the region of hybridization between the tracr mate sequence and the tracr sequence. For example, the modification can also include fusing the tracr mate sequence and the tracr sequence through an artificial loop. The modification can include a tracr sequence having a length of 40 to 120 bp. In embodiments of the present invention, the tracr sequence is from 40 bp to the full length of the tracr. In certain embodiments, the length of the tracRNA includes at least nucleotides 1-67 of the wild-type tracRNA, and in some embodiments, at least nucleotides 1-85. In some embodiments, nucleotides corresponding to at least nucleotides 1-67 or 1-85 of the wild-type S. pyogenes Cas9 tracRNA can be used. If the CRISPR system uses an enzyme other than Cas9 or SpCas9, the corresponding nucleotides in the relevant wild-type tracRNA can be present. In some embodiments, the length of the tracRNA includes no more than nucleotides 1-67 or 1-85 of the wild-type tracRNA. The modification can include sequence optimization. In some embodiments, sequence optimization is performed on the tracr and / or tracr mate sequences. Sequence optimization can include reducing the frequency of poly-T sequences. Sequence optimization can be combined with reducing the region of hybridization between the tracr mate sequence and the tracr sequence; for example, reducing the length of the tracr sequence.
[0024] In one aspect, the present invention provides a CRISPR-Cas system or CRISPR enzyme system, wherein the modification comprises reducing the poly-T sequence in tracr and / or tracr mate sequences. In some aspects of the present invention, one or more Ts (i.e., a stretch of more than 3, 4, 5, 6, or more consecutive T bases; in some embodiments, a stretch of 10, 9, 8, 7, 6 or less consecutive T bases) present in the poly-T sequence of the relevant wild-type sequence can be replaced with a non-T nucleotide, for example, A, thereby breaking the string into smaller stretches of T, each stretch having four or fewer (for example, three or two) consecutive Ts. Bases other than A, for example, C or G, or non-naturally occurring or modified nucleotides, can be used for substitution. When the string of Ts is involved in the formation of a hairpin (or stem loop), it is advantageous to change the complementary base of the non-T base to the complementary strand of the non-T nucleotide. For example, if the non-T base is A, its complementary strand can be changed to T, for example, to preserve or support the preservation of secondary structure. For example, 5'-TTTTT can be changed to become 5'-TTTAT and the complementary 5'-AAAAA can be changed to 5'-ATAAA.
[0025] In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system, wherein the modification comprises adding a poly-T terminator sequence. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system, wherein the modification comprises adding a poly-T terminator sequence in the tracr and / or tracr mate sequence. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system, wherein the modification comprises adding a poly-T terminator sequence in the guide sequence. The poly-T terminator sequence can comprise five consecutive T bases or more than five consecutive T bases.
[0026] In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification comprises altering the loop and / or hairpin. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification comprises providing a minimum of two hairpins in the guide sequence. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification comprises providing a hairpin formed by complementarity between the tracr and tracr mate (direct repeat) sequences. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification comprises providing one or more additional hairpins at or toward the 3' end of the tracrRNA sequence. For example, a hairpin can be formed by providing a self-complementary sequence within the tracRNA sequence connected by a loop such that a hairpin is formed upon self-folding. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification comprises providing an additional hairpin added 3' to the guide sequence. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification comprises extending the 5' end of the guide sequence. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification comprises providing one or more hairpins in the 5' end of the guide sequence. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification comprises adding the sequence (5'-AGGACGAAGTCCTAA) to the 5' end of the guide sequence. Other sequences suitable for forming hairpins are known to those skilled in the art and can be used in certain embodiments of the present invention. In some embodiments of the present invention, at least two, three, four, five, or more additional hairpins are provided. In some embodiments of the present invention, up to ten, nine, eight, seven, six, or fewer additional hairpins are provided. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification comprises two hairpins. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification comprises three hairpins.In one aspect, the invention provides a CRISPR-Cas system or CRISPR enzyme system wherein the modification comprises at most five hairpins.
[0027] In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification comprises providing a bridge or providing one or more modified nucleotides in a polynucleotide sequence. The modified nucleotides and / or bridges can be provided in any or all of the tracr, tracr mate, and / or guide sequences, and / or in the enzyme-coding sequence and / or in the vector sequence. The modification can include the inclusion of at least one non-naturally occurring nucleotide or modified nucleotide, or an analog thereof. The modified nucleotide can be modified in the ribose, phosphate, and / or base moiety. Modified nucleotides can include 2'-O-methyl analogs, 2'-deoxy analogs, or 2'-fluoro analogs. The nucleic acid backbone can be modified, for example, a phosphorothioate backbone can be used. Locked nucleic acids (LNA) or bridged nucleic acids (BNA) can also be used. Further examples of modified bases include, but are not limited to, 2-aminopurine, 5-bromo-uridine, pseudouridine, inosine, and 7-methylguanosine.
[0028] It is understood that any or all of the above modifications can be provided alone or in combination in a given CRISPR-Cas system or CRISPR enzyme system. Such a system can include one, two, three, four, five, or more of the modifications.
[0029] In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system, wherein the CRISPR enzyme is a type II CRISPR enzyme, such as a Cas9 enzyme. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system, wherein the CRISPR enzyme comprises less than 1000 amino acids, or less than 4000 amino acids. In one aspect, the invention provides a CRISPR-Cas system or CRISPR enzyme system wherein the Cas9 enzyme is StCas9 or StlCas9, or wherein the Cas9 enzyme is a Cas9 enzyme from an organism selected from the group consisting of the genus Streptococcus, Campylobacter, Nitratifractor, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, Sphaerochaeta, Lactobacillus, Eubacterium, or Corynebacter. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the CRISPR enzyme is a nuclease that directs cleavage of both strands at the location of the target sequence.
[0030] In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the first regulatory element is a polymerase III promoter. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the second regulatory element is a polymerase II promoter.
[0031] In one aspect, the present invention provides a CRISPR-Cas system or CRISPR enzyme system wherein the guide sequence comprises at least 15 nucleotides.
[0032] In one aspect, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modifications include an optimized tracr sequence and / or an optimized guide sequence RNA and / or a co-fold structure of the tracr sequence and / or the tracr mate sequence and / or a stabilized secondary structure of the tracr sequence and / or a tracr sequence and / or a tracr sequence and / or a tracr sequence fusion RNA element with regions of reduced base pairing; and / or in a multiplexed system, there are two RNAs containing a tracer and multiple guides or one RNA containing multiple chimeras.
[0033] In an embodiment of the present invention, the chimeric RNA architecture is further optimized according to the results of mutagenesis testing.In chimeric RNAs with two or more hairpins, mutations in the proximal direct repeat to stabilize the hairpin can result in the termination of CRISPR complex activity.Mutations in the distal direct repeat to shorten or stabilize the hairpin may not have any effect on CRISPR complex activity.Sequence randomization in the bulge region between the proximal and distal repeats can significantly reduce CRISPR complex activity.Single base pair changes or sequence randomization in the linker region between hairpins can result in the complete loss of CRISPR complex activity.Hairpin stabilization of the distal hairpin following the first hairpin after the guide sequence can maintain or improve CRISPR complex activity.Therefore, in a preferred embodiment of the present invention, the chimeric RNA architecture can be further optimized by generating smaller chimeric RNAs, which may be beneficial for therapeutic delivery options and other uses, and this can be achieved by modifying the distal direct repeat to shorten or stabilize the hairpin. In a further preferred embodiment of the present invention, the chimeric RNA architecture can be further optimized by stabilizing one or more of the distal hairpins. Stabilizing the hairpin can include modifying the sequence suitable for hairpin formation. In some embodiments of the present invention, at least two, three, four, five, or more additional hairpins are provided. In some embodiments of the present invention, up to ten, nine, eight, seven, six, or fewer additional hairpins are provided. In some embodiments of the present invention, stabilization can be crosslinking and other modifications. Modifications can include the inclusion of at least one non-naturally occurring nucleotide, or modified nucleotide, or analog thereof. Modified nucleotides can be modified in the ribose, phosphate, and / or base moieties. Modified nucleotides can include 2'-O-methyl, 2'-deoxy, or 2'-fluoro analogs. The nucleic acid backbone can be modified, for example, a phosphorothioate backbone can be used.The use of locked nucleic acids (LNA) or bridged nucleic acids (BNA) may also be possible. Further examples of modified bases include, but are not limited to, 2-aminopurine, 5-bromo-uridine, pseudouridine, inosine, 7-methylguanosine.
[0034] In one aspect, the present invention provides a CRISPR-Cas system or a CRISPR enzyme system in which the CRISPR enzyme is codon-optimized for expression in eukaryotic cells.
[0035] Thus, in some embodiments of the present invention, the length of the tracRNA required in a construct of the present invention, e.g., a chimeric construct, is not necessarily fixed; in some embodiments of the present invention, it can be 40-120 bp, and in some embodiments, it can be up to the full-length tracr, e.g., up to the 3' end of the tracr, where it is interrupted by a transcription termination signal in the bacterial genome. In some embodiments, the length of the tracRNA includes at least nucleotides 1-67 of the wild-type tracRNA, and in some embodiments, at least nucleotides 1-85. In some embodiments, nucleotides corresponding to at least nucleotides 1-67 or 1-85 of the wild-type S. pyogenes Cas9 tracRNA can be used. If the CRISPR system uses an enzyme other than Cas9 or SpCas9, the corresponding nucleotides in the relevant wild-type tracRNA can be present. In some embodiments, the length of the tracRNA includes no more than nucleotides 1-67 or 1-85 of the wild-type tracRNA. With respect to sequence optimization (e.g., reduction of poly-T sequences), e.g., with respect to strings of Ts within a tracr mate (direct repeat) or tracrRNA, in some embodiments of the invention, one or more Ts present in the poly-T sequence of the relevant wild-type sequence (i.e., a stretch of more than 3, 4, 5, 6, or more consecutive T bases; in some embodiments, a stretch of 10, 9, 8, 7, 6, or fewer consecutive T bases) can be replaced with a non-T nucleotide, e.g., A, thus decomposing the string into smaller stretches of Ts, each having four or fewer (e.g., three or two) consecutive Ts. When a string of Ts is involved in the formation of a hairpin (or stem-loop), it is advantageous to change the complementary base for the non-T base to the complement of the non-T nucleotide. For example, if the non-T base is A, its complement can be changed to T, e.g., to preserve or assist in preserving secondary structure. For example, 5'-TTTTT can be changed to become 5'-TTTAT and the complementary 5'-AAAAA can be changed to 5'-ATAAA.Regarding the presence of a poly-T terminator sequence, e.g., a poly-T terminator (TTTTTT or more Ts) in the tracr+tracr mate transcript, in some embodiments of the present invention, it is advantageous to add such a terminator to the end of the transcript, whether in the form of two RNAs (tracr and tracr mate) or a single guide RNA. Regarding loops and hairpins in the tracr and tracr mate transcripts, in some embodiments of the present invention, it is advantageous for a minimum of two hairpins to be present in the chimeric guide RNA. The first hairpin may be a hairpin formed by complementarity between the tracr and tracr mate (direct repeat) sequences. The second hairpin may be at the 3' end of the tracrRNA sequence, which may provide a secondary structure for interaction with Cas9. An additional hairpin may be added to the 3' end of the guide RNA, for example, to increase the stability of the guide RNA in some embodiments of the present invention. Furthermore, the 5' end of the guide RNA may be extended in some embodiments of the present invention. In some embodiments of the present invention, 20 bp in the 5' end may be considered the guide sequence. The 5' portion can be extended. One or more hairpins can be provided in the 5' portion, for example, in some embodiments of the present invention, this can also improve the stability of the guide RNA. In some embodiments of the present invention, a defined hairpin can be provided by adding the sequence (5'-AGGACGAAGTCCTAA) to the 5' end of the guide sequence, which can help improve stability in some embodiments of the present invention. Other sequences suitable for forming hairpins are known to those skilled in the art and can be used in certain embodiments of the present invention. In some embodiments of the present invention, at least 2, 3, 4, 5, or more additional hairpins are provided. In some embodiments of the present invention, 10, 9, 8, 7, 6, or fewer additional hairpins are provided. The above also provides embodiments of the present invention that include secondary structures in the guide sequence. In some embodiments of the present invention, for example, crosslinks and other modifications to improve stability can be present. Modifications can include the inclusion of at least one non-naturally occurring nucleotide, or modified nucleotide, or analog thereof.Modified nucleotides may be modified at the ribose, phosphate, and / or base moiety. Modified nucleotides may include 2'-O-methyl analogs, 2'-deoxy analogs, or 2'-fluoro analogs. The nucleic acid backbone may be modified, for example, a phosphorothioate backbone may be used. Locked nucleic acids (LNA) or bridged nucleic acids (BNA) may also be used. Further examples of modified bases include, but are not limited to, 2-aminopurine, 5-bromo-uridine, pseudouridine, inosine, and 7-methylguanosine. Such modifications or bridges may be present in the guide sequence or other sequences adjacent to the guide sequence.
[0036] Accordingly, it is not the object of the present invention to encompass any previously known products, methods of making the products, or methods of using the products, and therefore, applicants reserve the right to, and hereby disclose, a disclaimer of any previously known products, methods of making the products, or methods of using the products. It is further noted that the present invention does not encompass within its scope any products, methods, or methods of making or using the products that do not meet the description and enablement requirements of the USPTO (35 U.S.C. Section 112, first paragraph) or the EPO (European Patent Convention, Article 83), and therefore, applicants reserve the right to, and hereby disclose, a disclaimer of any previously described products, methods of making the products, or methods of using the products.
[0037] In this disclosure, and particularly in the claims and / or paragraphs, terms such as "comprises," "comprised," "comprising," and the like, may have the meaning ascribed to them in U.S. Patent Law; for example, they may mean "includes," "included," "including," and the like; and terms such as "consisting essentially of" and "consists essentially of" have the meaning ascribed to them in U.S. Patent Law, for example, it is noted that they allow for elements not expressly recited, but exclude elements found in the prior art or that affect a basic or novel characteristic of the invention. These and other embodiments are disclosed or obvious from the following detailed description and are encompassed thereby.
[0038] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings. [Brief explanation of the drawings]
[0039] [Figure 1] A schematic model of the CRISPR system is shown. The Cas9 nuclease from Streptococcus pyogenes (yellow) is targeted to genomic DNA by a synthetic guide RNA (sgRNA) consisting of a 20-nt guide sequence (blue) and a scaffold (red). The guide sequence base pairs with the DNA target (blue) immediately upstream of a required 5'-NGG protospacer adjacent motif (PAM; magenta), and Cas9 mediates a double-strand break (DSB) (red triangle) approximately 3 bp upstream of the PAM. [Figure 2A] Exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations for expression in eukaryotic cells, and results of studies assessing nuclear localization and CRISPR activity are presented. [Figure 2B]Exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations for expression in eukaryotic cells, and results of studies assessing nuclear localization and CRISPR activity are presented. [Figure 2C] Exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations for expression in eukaryotic cells, and results of studies assessing nuclear localization and CRISPR activity are presented. [Figure 2D] Exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations for expression in eukaryotic cells, and results of studies assessing nuclear localization and CRISPR activity are presented. [Figure 2E] Exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations for expression in eukaryotic cells, and results of studies assessing nuclear localization and CRISPR activity are presented. [Figure 2F] Exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations for expression in eukaryotic cells, and results of studies assessing nuclear localization and CRISPR activity are presented. [Figure 3] 1 shows exemplary expression cassettes for expression of CRISPR system elements in eukaryotic cells, the predicted structure of exemplary guide sequences, and CRISPR system activity measured in eukaryotic and prokaryotic cells. [Figure 4A] 1 shows the results of an assessment of SpCas9 specificity for exemplary targets. [Figure 4B] 1 shows the results of an assessment of SpCas9 specificity for exemplary targets. [Figure 4C] 1 shows the results of an assessment of SpCas9 specificity for exemplary targets. [Figure 4D] 1 shows the results of an assessment of SpCas9 specificity for exemplary targets. [Figure 5A] An exemplary vector system and results for its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5B] An exemplary vector system and results for its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5C] An exemplary vector system and results for its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5D] An exemplary vector system and results for its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5E] An exemplary vector system and results for its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5F] An exemplary vector system and results for its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5G] An exemplary vector system and results for its use in directing homologous recombination in eukaryotic cells are presented. [Figure 6A] A comparison of different tracrRNA transcripts for Cas9-mediated gene targeting is described. [Figure 6B] A comparison of different tracrRNA transcripts for Cas9-mediated gene targeting is described. [Figure 6C] A comparison of different tracrRNA transcripts for Cas9-mediated gene targeting is described. [Figure 7A] Exemplary CRISPR systems, exemplary adaptations for expression in eukaryotic cells, and results of tests assessing CRISPR activity are described. [Figure 7B] Exemplary CRISPR systems, exemplary adaptations for expression in eukaryotic cells, and results of tests assessing CRISPR activity are described. [Figure 7C] Exemplary CRISPR systems, exemplary adaptations for expression in eukaryotic cells, and results of tests assessing CRISPR activity are described. [Figure 7D] Exemplary CRISPR systems, exemplary adaptations for expression in eukaryotic cells, and results of tests assessing CRISPR activity are described. [Figure 8] Exemplary manipulations of CRISPR systems for targeting genomic loci in mammalian cells are described. [Figure 9] 1 illustrates the results of Northern blot analysis of crRNA processing in mammalian cells. [Figure 10A]Schematic representation of the chimeric RNA and results of the SURVEYOR assay for CRISPR system activity in eukaryotic cells are described. [Figure 10B] Schematic representation of the chimeric RNA and results of the SURVEYOR assay for CRISPR system activity in eukaryotic cells are described. [Figure 10C] Schematic representation of the chimeric RNA and results of the SURVEYOR assay for CRISPR system activity in eukaryotic cells are described. [Figure 11] 1 illustrates a graphical representation of the results of a SURVEYOR assay for CRISPR system activity in eukaryotic cells. [Figure 12] 1 illustrates the predicted secondary structure for an exemplary chimeric RNA comprising a guide sequence, a tracr mate sequence, and a tracr sequence. [Figure 13A] Phylogenetic tree of Cas genes. [Figure 13B] Phylogenetic tree of Cas genes. [Figure 13C] Phylogenetic tree of Cas genes. [Figure 13D] Phylogenetic tree of Cas genes. [Figure 14A] A phylogenetic analysis reveals five families of Cas9s, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids). [Figure 14B] A phylogenetic analysis reveals five families of Cas9s, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids). [Figure 14C] A phylogenetic analysis reveals five families of Cas9s, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids). [Figure 14D] A phylogenetic analysis reveals five families of Cas9s, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids). [Figure 14E]A phylogenetic analysis reveals five families of Cas9s, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids). [Figure 14F] A phylogenetic analysis reveals five families of Cas9s, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids). [Figure 15] 1 shows a graph illustrating the function of different optimized guide RNAs. [Figure 16] The sequences and structures of the different guide chimeric RNAs are shown. [Figure 17] Co-folded structure of tracrRNA and direct repeats is shown. [Figure 18A] 1 shows data from in vitro St1Cas9 chimeric guide RNA optimization. [Figure 18B] 1 shows data from in vitro St1Cas9 chimeric guide RNA optimization. [Figure 19] Cleavage of either unmethylated or methylated targets by SpCas9 cell lysates is shown. [Figure 20]Optimization of guide RNA architectures for SpCas9-mediated mammalian genome editing. (a) Schematic diagram of the bicistronic expression vector (PX330) for the single guide RNA (sgRNA) driven by the U6 promoter and the human codon-optimized Streptococcus pyogenes Cas9 (hSpCas9) driven by the CBh promoter, which were used in all subsequent experiments. The sgRNA consists of a 20-nt guide sequence (blue) and a scaffold (red) truncated at various positions as indicated. (b) SURVEYOR assay for SpCas9-mediated indels in the human EMX1 and PVALB loci. Arrows indicate predicted SURVEYOR fragments (n = 3). (c) Northern blot analysis of four sgRNA truncation architectures using U1 as a loading control. (d) Both wild-type (wt) and nickase mutant (D10A) SpCas9 promoted HindIII site insertion into the human EMX1 gene. Single-stranded oligonucleotides (ssODNs) were oriented in either sense or antisense orientation relative to the genomic sequence and used as homologous recombination templates. (e) Schematic diagram of the human SERPINB5 locus. sgRNAs and PAMs are indicated by colored bars above the sequence; methylcytosines (Me) are highlighted (pink) and numbered relative to the transcription start site (TSS, +1). (f) Methylation status of SERPINB5 assayed by bisulfite sequencing of 16 clones. Black circles, methylated CpGs; white circles, unmethylated CpGs. (g) Modification efficiency of three sgRNAs targeting methylated regions of SERPINB5 assayed by deep sequencing (n=2). Error bars indicate Wilson intervals (online method). [Figure 21]Further optimization of CRISPR-Cas sgRNA architectures is shown. (a) Schematic diagram of four additional sgRNA architectures I–IV. Each consists of a 20-nt guide sequence (blue) bound to a direct repeat (DR, gray) that hybridizes to tracrRNA (red). The DR-tracrRNA hybrid is truncated at +12 or +22, as indicated, and contains an artificial GAAA stem-loop. The tracrRNA truncation positions are numbered according to the previously reported transcription start sites for tracrRNA. sgRNA architectures II and IV carry mutations within their polyU tract that may function as early transcription terminators. (b) SURVEYOR assay for SpCas9-mediated indels in the human EMX1 locus for target sites 1–3. Arrows indicate predicted SURVEYOR fragments (n=3). [Figure 22] 1 illustrates the visualization of some target sites in the human genome. [Figure 23] (A) Schematic of sgRNA and (B) SURVEYOR analysis of five sgRNA variants for SaCas9 for optimal truncated architecture with maximum cleavage efficiency. DETAILED DESCRIPTION OF THE INVENTION
[0040] The drawings herein are for illustrative purposes only and are not necessarily drawn to scale.
[0041] The terms "polynucleotide," "nucleotide," "nucleotide sequence," "nucleic acid," and "oligonucleotide" are used interchangeably. They refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or their analogs. Polynucleotides can have any three-dimensional structure and can perform any function, known or unknown. The following are non-limiting examples of polynucleotides: coding or non-coding regions of a gene or gene fragment, loci defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. A polynucleotide can contain one or more modified nucleotides, such as methylated nucleotides or nucleotide analogs. Modifications to the nucleotide structure, if present, can be imparted before or after assembly of the polymer. The sequence of nucleotides can be interrupted by non-nucleotide components. A polynucleotide may be further modified after polymerization, such as by conjugation with a labeling component.
[0042] In embodiments of the present invention, the terms "chimeric RNA," "chimeric guide RNA," "guide RNA," "single guide RNA," and "synthetic guide RNA" are used interchangeably to refer to a polynucleotide sequence comprising a guide sequence, a tracr sequence, and a tracr mate sequence. The term "guide sequence" refers to an approximately 20-bp sequence within the guide RNA that defines the target site, and can be used interchangeably with the terms "guide" or "spacer." The term "tracr mate sequence" can also be used interchangeably with the term "direct repeat."
[0043] The term "wild-type" as used herein is a term of the art understood by those skilled in the art and means the typical form of an organism, strain, gene or characteristic as it occurs in nature, as distinguished from mutant or variant forms.
[0044] The term "variant" as used herein should be taken to mean a display of qualities having a pattern that deviates from that occurring in nature.
[0045] The terms "non-naturally occurring" and "engineered" are used interchangeably and refer to artificial involvement. When referring to a nucleic acid molecule or polypeptide, the term means that the nucleic acid molecule or polypeptide is at least substantially free from at least one other component with which it is naturally associated or found in the natural state.
[0046] "Complementarity" refers to the ability of a nucleic acid to form hydrogen bonds with another nucleic acid sequence, either through classical Watson-Crick base pairing or other non-classical types. The percent complementarity indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Fully complementary" means that all consecutive residues of a nucleic acid sequence will hydrogen bond with the same number of consecutive residues in a second nucleic acid sequence. As used herein, "substantially complementary" refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or refers to two nucleic acids that hybridize under stringent conditions.
[0047] As used herein, "stringent conditions" for hybridization refer to conditions under which a nucleic acid having complementarity to a target sequence will predominantly hybridize to the target sequence and will not substantially hybridize to non-target sequences. Stringent conditions are generally sequence-dependent and vary depending on numerous factors. Generally, the longer the sequence, the higher the temperature at which the sequence will specifically hybridize to its target sequence. Non-limiting examples of stringent conditions are detailed in Tijssen (1993), Laboratory Techniques In Biochemistry And Molecular Biology—Hybridization With Nucleic Acid Probes Part I, Second Chapter "Overview of principles of hybridization and the strategy of nucleic acid probe assay," Elsevier, NY.
[0048] "Hybridization" refers to a reaction in which one or more polynucleotides react to form a complex stabilized through hydrogen bonding between the bases of nucleotide residues. Hydrogen bonding can occur through Watson-Crick base pairing, Hoogsteen binding, or any other sequence-specific manner. The complex can contain two strands forming a double-stranded structure, three or more strands forming a multistranded complex, a single self-hybridizing strand, or any combination thereof. A hybridization reaction can constitute a step in a more extensive process, such as the initiation of PCR or the cleavage of a polynucleotide by an enzyme. A sequence that can hybridize to a given sequence is referred to as the "complementary strand" of the given sequence.
[0049] As used herein, "stabilization" or "increasing stability" in reference to components of a CRISPR system refers to ensuring or stabilizing the structure of a molecule. This can be achieved by introducing one or more mutations, including single or multiple base pair changes, increasing the number of hairpins, crosslinking, disrupting specific stretches of nucleotides, and other modifications. The modifications can include the inclusion of at least one non-naturally occurring nucleotide, or modified nucleotide, or their analogs. Modified nucleotides can be modified in the ribose, phosphate, and / or base moieties. Modified nucleotides can include 2'-O-methyl analogs, 2'-deoxy analogs, or 2'-fluoro analogs. The nucleic acid backbone can be modified, for example, a phosphorothioate backbone can be used. Locked nucleic acids (LNA) or bridged nucleic acids (BNA) can also be used. Further examples of modified bases include, but are not limited to, 2-aminopurine, 5-bromo-uridine, pseudouridine, inosine, and 7-methylguanosine. These modifications can be applied to any component of a CRISPR system. In a preferred embodiment, these modifications are made to the RNA component, e.g., the guide RNA or the polynucleotide sequence.
[0050] As used herein, "expression" refers to the process by which a polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcripts) and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide can be collectively referred to as a "gene product." If the polynucleotide is derived from genomic DNA, expression can include splicing of the mRNA in a eukaryotic cell.
[0051] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acids of any length. The polymer may be linear or branched, may contain modified amino acids, and may be interrupted by non-amino acids. The term also encompasses amino acid polymers that have undergone modifications, such as disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with a labeling component. As used herein, the term "amino acid" includes natural and / or unnatural or synthetic amino acids, including glycine and both D- and L-optical isomers, as well as amino acid analogs and peptidomimetics.
[0052] The terms "subject," "individual," and "patient" are used interchangeably herein to refer to a vertebrate, preferably a mammal, more preferably a human. Mammals include, but are not limited to, murines, simians, humans, livestock, sport animals, and pets. Also included are tissues, cells, and their progeny of biological entities obtained in vivo or cultured in vitro. In some embodiments, the subject may be an invertebrate, such as an insect or nematode; while the subject may be a plant or fungus.
[0053] The terms "therapeutic agent," "therapeutic drug," or "treatment agent" are used interchangeably and refer to a molecule or compound that confers some beneficial effect upon administration to a subject. Beneficial effects include enabling a diagnostic measurement; ameliorating a disease, symptom, disorder, or pathological condition; reducing or preventing the onset of a disease, symptom, disorder, or pathological condition; and generally neutralizing a disease, symptom, disorder, or pathological condition.
[0054] As used herein, "treatment" or "treating" or "alleviating" or "ameliorating" are used interchangeably. These terms refer to an approach to obtain a benefit or desired result, for example, but not limited to, a therapeutic benefit and / or a preventive benefit. A therapeutic benefit refers to any treatment-related improvement in or effect on one or more diseases, conditions, or symptoms being treated. For preventive benefit, the composition can be administered to a subject at risk of developing a particular disease, condition, or symptom, or to a subject reporting one or more physiological symptoms of a disease, even if the disease, condition, or symptom has not previously manifested.
[0055] The term "effective amount" or "therapeutically effective amount" refers to an amount of a drug sufficient to produce a benefit or desired result. The therapeutically effective amount may vary depending on one or more of the subject and condition being treated, the subject's weight and age, the severity of the condition, the mode of administration, etc., which can be easily determined by those skilled in the art. This term also applies to the dose that provides an image for detection by any one of the imaging methods described herein. The specified dose may vary depending on one or more of the specific drug selected, the dosing regimen to be followed, whether it is administered in combination with other compounds, the timing of administration, the tissue to be imaged, and the physical delivery system in which it is carried.
[0056] The practice of the present invention will employ, unless otherwise indicated, conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, which are within the skill of one in the art. See Sambrook, Fritsch, and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, 2nd edition (1989); CURRENT PROTOCOLS IN MOLECULAR BIOLOGY (F.M.A.usubel, et al. eds., (1987)); series METHODS IN ENZYMOLOGY (Academic Press, Inc.): PCR 2: A PRACTICAL APPROACH (M.J. MacPherson, B.D. Hames, and G.R. Taylor eds. (1995)), Harlow and Lane, eds. (1988), ANTIBODIES, A LABORATORY MANUAL, and ANIMAL CELL CULTURE (R.I. Freshney, ed. (1987)).
[0057] Some embodiments of the present invention relate to vector systems containing one or more vectors, or to vectors themselves.Vector can be designed for the expression of CRISPR transcripts (e.g., nucleic acid transcripts, proteins, or enzymes) in prokaryotic or eukaryotic cells.For example, CRISPR transcripts can be expressed in bacterial cells, such as Escherichia coli, insect cells (using baculovirus expression vectors), yeast cells, or mammalian cells.Suitable host cells are further discussed in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990).Alternatively, recombinant expression vectors can be transcribed and translated in vitro, for example, using T7 promoter regulatory sequences and T7 polymerase.
[0058] Vectors can be introduced into and propagated within prokaryotic cells. In some embodiments, prokaryotes are used to amplify copies of vectors to be introduced into eukaryotic cells or as intermediate vectors in the production of vectors to be introduced into eukaryotic cells (e.g., amplifying plasmids as part of a viral vector packaging system). In some embodiments, prokaryotes are used to amplify copies of vectors, express one or more nucleic acids, and provide a source of one or more proteins, for example, for delivery to a host cell or host organism. Protein expression in prokaryotes is most often carried out in Escherichia coli using vectors containing constitutive or inducible promoters directing the expression of either fusion or non-fusion proteins. Fusion vectors add multiple amino acids to the encoded protein, for example, to the amino terminus of the recombinant protein. Such fusion vectors can serve one or more purposes, for example, (i) increasing the expression of the recombinant protein; (ii) increasing the solubility of the recombinant protein; and (iii) aiding in the purification of the recombinant protein by acting as a ligand in affinity purification. In fusion expression vectors, a proteolytic cleavage site is often introduced at the junction of the fusion moiety and the recombinant protein to allow separation of the recombinant protein from the fusion moiety after purification of the fusion protein. Such enzymes and their cognate recognition sequences include factor Xa, thrombin, and enterokinase. Exemplary fusion expression vectors include pGEX (Pharmacia Biotech Inc; Smith and Johnson, 1988. Gene 67:31-40), pMAL (New England Biolabs, Beverly, Mass.), and pRIT5 (Pharmacia, Piscataway, NJ), which fuse glutathione S-transferase (GST), maltose E-binding protein, or protein A to the target recombinant protein, respectively.
[0059] Examples of suitable inducible non-fusion E. coli expression vectors include pTrc (Amrann et al., (1988) Gene 69:301-315) and pET 11d (Studier et al., GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990) 60-89).
[0060] In some embodiments, the vector is a yeast expression vector. Examples of vectors for expression in the yeast Saccharomyces cerivisae include pYepSec1 (Baldari, et al., 1987. EMBO J. 6:229-234), pMFa (Kuijan and Herskowitz, 1982. Cell 30:933-943), pJRY88 (Schultz et al., 1987. Gene 54:113-123), pYES2 (Invitrogen Corporation, San Diego, Calif.), and picZ (InVitrogen Corp, San Diego, Calif.).
[0061] In some embodiments, the vector drives protein expression in insect cells using a baculovirus expression vector. Baculovirus vectors available for protein expression in cultured insect cells (e.g., SF9 cells) include the pAc series (Smith et al., 1983. Mol. Cell. Biol. 3:2156-2165) and the pVL series (Lucklow and Summers, 1989. Virology 170:31-39).
[0062] In some embodiments, the vector may drive expression of one or more sequences in mammalian cells using a mammalian expression vector. Examples of mammalian expression vectors include pCDM8 (Seed, 1987. Nature 329:840) and pMT2PC (Kaufman, et al., 1987. EMBO J. 6:187-195). When used in mammalian cells, expression vector control functions are typically provided by one or more regulatory elements. For example, commonly used promoters are derived from polyoma, adenovirus type 2, cytomegalovirus, simian virus 40, and others disclosed herein and known in the art. For other suitable expression systems for both prokaryotic and eukaryotic cells, see, for example, Chapters 16 and 17 of Sambrook, et al., MOLECULAR CLONING: A LABORATORY MANUAL. 2nd ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.
[0063] In some embodiments, the recombinant mammalian expression vector can direct expression of the nucleic acid preferentially in a particular cell type (e.g., tissue-specific regulatory elements are used to express the nucleic acid). Tissue-specific regulatory elements are known in the art. Non-limiting examples of suitable tissue-specific promoters include the albumin promoter (liver-specific; Pinkert, et al., 1987, Genes Dev. 1:268-277), lymphoid-specific promoters (Calame and Eaton, 1988, Adv. Immunol. 43:235-275), promoters of T-cell receptors in particular (Winoto and Baltimore, 1989, EMBO J. 8:729-733) and immunoglobulins (Baneiji, et al., 1983, Cell 33:729-740; Queen and Baltimore, 1983, Cell 33:741-748), neuron-specific promoters (e.g., the neurofilament promoter; Byrne and Ruddle, 1989, Proc. Natl. Acad. Sci. USA 86:5473-5477), pancreatic-specific promoters (Edlund, et al., 1989, Proc. Natl. Acad. Sci. USA 86:5473-5477), and the like. et al., 1985, Science 230:912-916), and mammary gland-specific promoters (e.g., whey promoters; U.S. Pat. No. 4,873,316 and European Patent Application Publication No. 264,166). Developmentally regulated promoters, such as murine hox promoters (Kessel and Gruss, 1990, Science 249:374-379) and the alpha-fetoprotein promoter (Campes and Tilghman, 1989, Genes Dev. 3:537-546), are also encompassed.
[0064] In some embodiments, the regulatory element is operably linked to one or more elements of the CRISPR system, so as to drive the expression of one or more elements of the CRISPR system.Generally, CRISPR (Clustered Regularly Interspaced Short Repeats), also known as SPIDR (Spacer Interspersed Direct Repeat), constitutes a family of DNA loci that are usually specific to certain bacterial species.CRISPR loci include a distinct class of interspersed short sequence repeats (SSRs) recognized in Escherichia coli (E. coli) (Ishino et al., J. Bacteriol.,169:5429-5433
[1987] ; and Nakata et al., J. Bacteriol.,171:3553-3556
[1989] ) and related genes. Similar interspersed SSRs have been identified in Haloferax mediterranei, Streptococcus pyogenes, Anabaena, and Mycobacterium tuberculosis (see Groenen et al., Mol. Microbiol., 10:1057-1065
[1993] ; Hoe et al., Emerg. Infect. Dis., 5:254-263
[1999] ; Masepohl et al., Biochim. Biophys. Acta 1307:26-30
[1996] ; and Mojica et al., Mol. Microbiol., 17:85-93
[1995] ). CRISPR loci typically differ from other SSRs in the structure of their repeats, which are termed short regularly interspaced repeats (SRSRs) (Janssen et al., OMICS J. Integ. Biol., 6:23-33
[2002] ; and Mojica et al., Mol. Microbiol., 36:244-246
[2000] ). Generally, repeats are short elements that occur in clusters regularly spaced by unique intervening sequences of substantially constant length (Mojica et al.,
[2000] , supra). Repeat sequences are highly conserved between strains, butThe number of interspersed repeats and the sequence of the spacer region typically vary from strain to strain (van Embden et al., J. Bacteriol., 182:2393-2401
[2000] ). CRISPR loci have been identified in over 40 prokaryotes (e.g., Jansen et al., Mol. Microbiol., 43:1565-1575
[2002] ; and Mojica et al. See, e.g., et al.,
[2005] , including, but not limited to, the genera Aeropyrum, Pyrobaculum, Sulfolobus, Archaeoglobus, Halocarcula, Methanobacterium, Methanococcus, Methanosarcina, Methanopyrus, Pyrococcus, Picrophilus, Thermoplasma, Corynebacterium, Mycobacterium, Streptomyces, Aquifex, and the like. fex), Porphyromonas, Chlorobium, Thermus, Bacillus, Listeria, Staphylococcus, Clostridium, Thermoanaerobacter, Mycoplasma, Fusobacterium, Azarcus, Chromobacterium, Neisseria, Nitrosomonas, Desulfovibrio, Geobacter, Myxococcus, Campylobacter,The genera Wolinella, Acinetobacter, Erwinia, Escherichia, Legionella, Methylococcus, Pasteurella, Photobacterium, Salmonella, Xanthomonas, Yersinia, Treponema, and Thermotoga.
[0065] Generally, a "CRISPR system" collectively refers to transcripts and other elements involved in directing the expression or activity of CRISPR-associated ("Cas") genes, e.g., sequences encoding Cas genes, tracr (trans-activating CRISPR) sequences (e.g., tracrRNA or active partial tracrRNA), tracr mate sequences (including "direct repeats" and partial direct repeats processed by tracrRNA in the context of endogenous CRISPR systems), guide sequences (also referred to as "spacers" in the context of endogenous CRISPR systems), or other sequences and transcripts from a CRISPR locus. In some embodiments, one or more elements of a CRISPR system are derived from a type I, type II, or type III CRISPR system. In some embodiments, one or more elements of a CRISPR system are derived from a particular organism that contains an endogenous CRISPR system, e.g., Streptococcus pyogenes. Generally, CRISPR systems are characterized by elements that promote the formation of CRISPR complexes at target sequences (also referred to as protospacers in endogenous CRISPR systems). With regard to the formation of CRISPR complexes, "target sequence" refers to a sequence to which a guide sequence is designed to have complementarity, and hybridization between the target sequence and the guide sequence promotes the formation of a CRISPR complex. Full complementarity is not necessarily required, provided that there is sufficient complementarity to cause hybridization and promote the formation of a CRISPR complex. The target sequence can comprise any polynucleotide, for example, a DNA or RNA polynucleotide. In some embodiments, the target sequence is located in the nucleus or cytoplasm of a cell. In some embodiments, the target sequence can be present in an organelle of a eukaryotic cell, such as a mitochondria or chloroplast. A sequence or template that can be used for recombination into a targeted locus containing a target sequence is referred to as an "editing template," "editing polynucleotide," or "editing sequence." In an embodiment of the present invention, an exogenous template polynucleotide can be referred to as an editing template. In one aspect of the invention, the recombination is homologous recombination.
[0066] Typically, with respect to endogenous CRISPR systems, formation of a CRISPR complex (comprising a guide sequence hybridized to a target sequence and complexed with one or more Cas proteins) results in cleavage of one or both strands in or near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs therefrom). Without being bound by theory, the tracr sequence may comprise or consist of all or a portion of the wild-type tracr sequence (e.g., more than about or about 20, 26, 32, 45, 48, 54, 63, 67, 85, or more nucleotides of the wild-type tracr sequence), and may also form part of a CRISPR complex, for example, by hybridization along at least a portion of the tracr sequence with all or a portion of a tracr mate sequence operably linked to the guide sequence. In some embodiments, the tracr sequence has sufficient complementarity to the tracr mate sequence to hybridize and participate in the formation of a CRISPR complex. As with the target sequence, perfect complementarity is not required, provided that sufficient complementarity exists to be functional. In some embodiments, the tracr sequence has at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence complementarity along the length of the tracr mate sequence when optimally aligned. In some embodiments, one or more vectors driving the expression of one or more elements of a CRISPR system are introduced into a host cell, such that the expression of the elements of the CRISPR system directs the formation of a CRISPR complex at one or more target sites. For example, the Cas enzyme, the guide sequence linked to the tracr mate sequence, and the tracr sequence can each be operably linked to separate regulatory elements on separate vectors. Alternatively, two or more elements expressed from the same or different regulatory elements can be combined in a single vector, and one or more additional vectors providing any component of the CRISPR system are not included in the first vector.The CRISPR system elements combined in a single vector can be arranged in any suitable orientation; for example, an element can be located 5' (upstream) or 3' (downstream) relative to a second element. The coding sequence of an element can be located on the same or opposite strand as the coding sequence of a second element and can be oriented in the same or opposite direction. In some embodiments, a single promoter drives the expression of a transcript encoding a CRISPR enzyme and one or more of a guide sequence, a tracr mate sequence (optionally operably linked to a guide sequence), and a tracr sequence embedded in one or more intron sequences (e.g., each in a different intron, two or more in at least one intron, or all in a single intron). In some embodiments, the CRISPR enzyme, guide sequence, tracr mate sequence, and tracr sequence are operably linked to and expressed from the same promoter.
[0067] In some embodiments, a vector comprises one or more insertion sites, such as restriction endonuclease recognition sequences (also referred to as "cloning sites"). In some embodiments, one or more insertion sites (e.g., about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more insertion sites) are located upstream and / or downstream of one or more sequence elements of one or more vectors. In some embodiments, a vector comprises an insertion site upstream of a tracr mate sequence and optionally downstream of a regulatory element operably linked to the tracr mate sequence, such that after insertion of the guide sequence into the insertion site and upon expression, the guide sequence directs sequence-specific binding of a CRISPR complex to a target sequence in a eukaryotic cell. In some embodiments, a vector comprises two or more insertion sites, each located between two tracr mate sequences to allow insertion of a guide sequence at the respective site. In such an arrangement, the two or more guide sequences may comprise two or more copies of a single guide sequence, two or more different guide sequences, or a combination thereof. When using multiple different guide sequences, can use a single expression construct to target the CRISPR activity to multiple different corresponding target sequences in cells.For example, a single vector can comprise about or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or more guide sequences.In some embodiments, about or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more of these guide sequence-containing vectors can be provided and optionally delivered into cells.
[0068] In some embodiments, the vector comprises a regulatory element operably linked to an enzyme coding sequence that encodes a CRISPR enzyme, e.g., a Cas protein. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Csel, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified versions thereof. These enzymes are known; for example, the amino acid sequence of Streptococcus pyogenes (S. pyogenes) Cas9 protein can be found in the SwissProt database under accession number Q99ZW2. In some embodiments, the unmodified CRISPR enzyme has DNA cleavage activity, for example, Cas9. In some embodiments, the CRISPR enzyme is Cas9, and can be Cas9 from Streptococcus pyogenes (S. pyogenes) or Streptococcus pneumoniae (S. pneumoniae). In some embodiments, the CRISPR enzyme directs cleavage of one or both strands at the location of the target sequence, for example, within the target sequence and / or the complementary strand of the target sequence. In some embodiments, the CRISPR enzyme directs cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence. In some embodiments, the vector encodes a CRISPR enzyme that is mutated relative to the corresponding wild-type enzyme, such that the mutant CRISPR enzyme lacks the ability to cleave one or both strands of a target polynucleotide containing the target sequence.For example, an aspartate-to-alanine substitution (D10A) in the RuvC I catalytic domain of Cas9 from Streptococcus pyogenes (S. pyogenes) converts Cas9 from a nuclease that cleaves both strands to a nickase (cleaves a single strand). Other examples of mutations that convert Cas9 into a nickase include, but are not limited to, H840A, N854A, and N863A. In some embodiments, Cas9 nickase can be used in combination with guide sequences, e.g., two guide sequences that target the sense and antisense strands of a DNA target, respectively. This combination allows both strands to be nicked and used to induce NHEJ. Applicants have demonstrated the efficacy of two nickase targets (i.e., sgRNAs targeted to the same location but different strands of DNA) in inducing mutagenic NHEJ (data not shown). While a single nickase (Cas9-D10A with a single sgRNA) cannot induce NHEJ and create indels, we have shown that a dual nickase (Cas9-D10A and two sgRNAs targeted to different strands at the same location) can do so in human embryonic stem cells (hESCs), with approximately 50% of the efficiency of a nuclease (i.e., regular Cas9 without the D10 mutation) in hESCs.
[0069] As a further example, two or more catalytic domains of Cas9 (RuvC I, RuvC II, and RuvC III) can be mutated to produce a mutant Cas9 that substantially lacks all DNA cleavage activity. In some embodiments, the D10A mutation is combined with one or more of the H840A, N854A, or N863A mutations to produce a Cas9 enzyme that substantially lacks all DNA cleavage activity. In some embodiments, a CRISPR enzyme is considered to substantially lack all DNA cleavage activity if the DNA cleavage activity of the mutant enzyme is less than about 25%, 10%, 5%, 1%, 0.1%, 0.01%, or less of its non-mutated form. Other mutations may be useful; if the Cas9 or other CRISPR enzyme is from a species other than S. pyogenes, corresponding amino acid mutations can be made to achieve similar effects.
[0070] In some embodiments, the enzyme coding sequence encoding CRISPR enzyme is codon-optimized for expression in specific cells, for example, eukaryotic cells.Eukaryotic cells can be derived from or derived from specific organisms, for example, mammals, for example, but not limited to, humans, mice, rats, rabbits, dogs, or non-human primates.Generally, codon optimization refers to the process of modifying nucleic acid sequences to improve expression in target host cells by replacing at least one codon (for example, about or more than about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) of native sequence with the more frequently or most frequently used codon in the gene of the host cell, while maintaining the native amino acid sequence.Different species show specific bias for certain codons of specific amino acids. Codon bias (differences in codon usage between organisms) often correlates with the efficiency of messenger RNA (mRNA) translation, which in turn is thought to depend, among other things, on the characteristics of the codon being translated and the availability of specific transfer RNA (tRNA) molecules. The dominance of tRNAs selected in a cell generally reflects the codons most frequently used in peptide synthesis. Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, in "codon usage databases," and these tables can be adapted in a number of ways. See Nakamura, Y., et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000," Nucl. Acids Res. 28:292 (2000). Computer algorithms for codon-optimizing specific sequences for expression in specific host cells are also available, for example, in GeneForge (Aptagen; Jacobus, PA).In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more, or all codons) in a sequence encoding a CRISPR enzyme correspond to the most frequently used codon for a particular amino acid.
[0071] In some embodiments, the vector encodes a CRISPR enzyme that comprises one or more nuclear localization sequences (NLSs), for example, about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs. In some embodiments, the CRISPR enzyme comprises about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the amino terminus, about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the carboxy terminus, or a combination thereof (for example, one or more NLSs at the amino terminus and one or more NLSs at the carboxy terminus). When two or more NLSs are present, each can be selected independently from the others, so that a single NLS can exist in two or more copies, and / or in combination with one or more other NLSs that exist in one or more copies. In a preferred embodiment of the present invention, the CRISPR enzyme comprises at most six NLSs. In some embodiments, an NLS is considered to be near the N- or C-terminus if the nearest amino acid of the NLS is within about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, or more amino acids along the polypeptide chain from the N- or C-terminus. Typically, an NLS consists of one or more short sequences of positively charged lysines or arginines exposed on the protein surface, although other types of NLSs are known.Non-limiting examples of NLSs include the NLS of the SV40 virus large T antigen, which has the amino acid sequence PKKKRKV; an NLS from nucleoplasmin (e.g., the nucleoplasmin bipartite NLS, which has the sequence KRPAATKKAGQAKKKK); a c-myc NLS, which has the amino acid sequence PAAKRVKLD or RQRRNELKRSP; an hRNPA1 M9 NLS, which has the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY; an IBB domain from importin alpha, which has the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV; the sequences VSRKRPRP and PPKKARED of the fibroid T protein; the sequence POPKKKPL of human p53; and a mouse c-abl IV sequence SALIKKKKKMAP; influenza virus NS1 sequences DRLRR and PKQKKRK; hepatitis virus delta antigen sequence RKLKKKIKKL; mouse Mx1 protein sequence REKKKFLKRR; human poly(ADP-ribose) polymerase sequence KRKGDEVDGVDEVAKKKSKK; and steroid hormone receptor (human) glucocorticoid sequence RKCLQAGMNLEARKTKK.
[0072] Generally, one or more NLSs are strong enough to drive the accumulation of detectable amounts of CRISPR enzyme in the nucleus of eukaryotic cells.Generally, the strength of nuclear localization activity can be derived from the number of NLSs in CRISPR enzyme, the specific NLS used, or a combination of these factors.Detection of nuclear accumulation can be carried out by any suitable technique.For example, a detectable marker can be fused to CRISPR enzyme, so that intracellular localization can be visualized, for example, by combining with a means for detecting nuclear localization (for example, nuclear-specific staining, for example, DAPI).Examples of detectable markers include fluorescent proteins (for example, green fluorescent protein, or GFP; RFP; CFP) and epitope tags (HA tag, flag tag, SNAP tag).Cell nuclei can also be isolated from cells, and then their contents can be analyzed by any suitable process for detecting protein, for example, immunohistochemical analysis, Western blot, or enzyme activity assay. Accumulation in the nucleus can also be measured indirectly, for example, by assaying for the effect of CRISPR complex formation (e.g., assaying for DNA cleavage or mutation at the target sequence, or assaying for changes in gene expression activity affected by CRISPR complex formation and / or CRISPR enzymatic activity), compared to a control exposed to neither the CRISPR enzyme nor the complex, or to a CRISPR enzyme lacking one or more NLSs.
[0073] Generally, a guide sequence is any polynucleotide sequence that has sufficient complementarity with a target polynucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of a CRISPR complex to the target sequence. In some embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence is greater than about or about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more when optimally aligned using a suitable alignment algorithm. Optimal alignment can be determined using any suitable algorithm for aligning sequences, including, but not limited to, the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler Transform (e.g., Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, ELAND (Illumina, San Diego, CA), and others. Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). In some embodiments, the guide sequence is about or greater than about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, or more nucleotides in length. In some embodiments, the guide sequence is about 75, 50, 45, 40, 35, 30, 25 The length of the guide sequence is less than 1, 20, 15, 12 or less nucleotides.The ability of the guide sequence to direct the sequence-specific binding of CRISPR complex to target sequence can be evaluated by any suitable assay.For example, the components of the CRISPR system sufficient to form a CRISPR complex, for example, the guide sequence to be tested, can be provided to the host cell having the corresponding target sequence, for example, by transfection with the vector encoding the components of the CRISPR sequence, and then the preferential cleavage within the target sequence is evaluated, for example, by the Surveyor assay described herein.Similarly, cleavage of a target polynucleotide sequence can be assessed in a test tube by providing the target sequence, a component of a CRISPR complex, for example, a guide sequence to be tested and a control guide sequence that differs from the test guide sequence, and comparing the rate of binding or cleavage at the target sequence between the test and control guide sequence reactions. Other assays are contemplated and will be recognized by those skilled in the art.
[0074] The guide sequence can be selected to target any target sequence. In some embodiments, the target sequence is a sequence within the genome of a cell. Exemplary target sequences include those that are unique in the target genome. For example, for S. pyogenes Cas9, a unique target sequence in the genome can include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXGG, where NNNNNNNNNNNNXGG (N is A, G, T, or C; X can be anything) has a single occurrence in the genome. A unique target sequence in the genome can include a S. pyogenes Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXGG, where NNNNNNNNNNNXGG (N is A, G, T, or C; X can be anything) has a single occurrence in the genome. For S. thermophilus CRISPR1Cas9, unique target sequences in the genome can include Cas9 target sites of the form MMMMMMMMNNNNNNNNNNNNXXAGAAW, where N is A, G, T, or C; X is any X may be A or T; W is A or T) has a single occurrence in the genome. A unique target sequence in a genome can include an S. thermophilus CRISPR1Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXXAGAAW, where NNNNNNNNNNNXXAGAAW (N is A, G, T, or C; X can be anything; W is A or T) has a single occurrence in the genome. For S. pyogenes Cas9, a unique target sequence in a genome can include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXGGXG, where NNNNNNNNNNNNXGGXG (N is A, G, T, or C; X can be anything) has a single occurrence in the genome. A unique target sequence in a genome can include a Streptococcus pyogenes (S. pyogenes) Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXGGXG, where NNNNNNNNNNNXGGXG (N is A, G, T, or C; X can be anything) has a single occurrence in the genome. In each of these sequences, "M" can be A, G, T, or C and need not be considered in identifying the sequence as unique.
[0075] In some embodiments, the guide sequence is selected to reduce the degree of secondary structure within the guide sequence. The secondary structure can be determined by any suitable polynucleotide folding algorithm. Some programs are based on calculating the minimum Gibbs free energy. An example of such an algorithm is mFold, described by Zuker and Stiegler (Nucleic Acids Res. 9 (1981), 133-148). Another exemplary folding algorithm is the online web server RNAfold, developed by the Institute for Theoretical Chemistry at the University of Vienna, which uses a centroid structure prediction algorithm (see, for example, A.R. Gruber et al., 2008, Cell 106(1):23-24; and P.A. Carr and G.M. Church, 2009, Nature Biotechnology 27(12):1151-62). Further algorithms can be found in US Patent Application No. TBA (Broad Reference No. BI2012 / 084 44790.11.2022), which is incorporated herein by reference.
[0076] Generally, the tracr mate sequence includes any sequence that has sufficient complementarity with the tracr sequence to promote one or more of the following: (1) excision of the guide sequence flanked by the tracr mate sequence in cells containing the corresponding tracr sequence; and (2) formation of a CRISPR complex in the target sequence (the CRISPR complex includes the tracr mate sequence hybridized to the tracr sequence). Generally, the degree of complementarity is based on optimal alignment of the tracr mate sequence and the tracr sequence along the shorter length of the two sequences. Optimal alignment can be determined by any suitable alignment algorithm and can further account for secondary structures, such as self-complementarity within the tracr sequence or the tracr mate sequence. In some embodiments, the degree of complementarity between the tracr sequence and the tracr mate sequence along the shorter length of the two sequences is greater than about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99%, or more when optimally aligned. An exemplary illustration of optimal alignment between the tracr sequence and the tracr mate sequence is provided in Figures 12B and 13B. In some embodiments, the tracr sequence is about or greater than about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, or more nucleotides in length. In some embodiments, the tracr sequence and the tract mate sequence are contained within a single transcript, such that hybridization between the two produces a transcript with a secondary structure, e.g., a hairpin. A preferred loop-forming sequence used in the hairpin structure is 4 nucleotides in length and most preferably has the sequence GAAA. However, longer or shorter loop sequences can be used, as can alternative sequences. The sequence preferably includes a nucleotide triplet (e.g., AAA) and an additional nucleotide (e.g., C or G). Examples of loop-forming sequences include CAAA and AAAG. In one embodiment of the present invention, the transcript or transcribed polynucleotide sequence has at least two or more hairpins.In preferred embodiments, the transcript has two, three, four, or five hairpins. In another further embodiment of the present invention, the transcript has at most five hairpins. In some embodiments, the single transcript further comprises a transcription termination sequence; preferably, this is a poly-T sequence, e.g., six T nucleotides. An illustrative illustration of such a hairpin structure is provided in the lower portion of Figure 13B, where the final "N" and the portion of the sequence 5' upstream of the loop correspond to the tracr mate sequence, and the portion of the sequence 3' upstream of the loop corresponds to the tracr sequence. Further non-limiting examples of single polynucleotides comprising a guide sequence, a tracr mate sequence, and a tracr sequence are as follows (listed 5' to 3'), where "N" represents the base of the guide sequence, the first block of lowercase letters represents the tracr mate sequence, the second block of lowercase letters represents the tracr sequence, and the final poly-T sequence represents the transcription terminator: [ka] In some embodiments, sequences (1) through (3) are used in combination with Cas9 from S. thermophilus CRISPR1. In some embodiments, sequences (4) through (6) are used in combination with Cas9 from S. pyogenes. In some embodiments, the tracr sequence is a separate transcript from the transcript containing the tracr mate sequence (e.g., as illustrated in the top diagram of Figure 13B).
[0077] In some embodiments, a recombination template is also provided. The recombination template is a component of another vector described herein, and can be contained in a separate vector or provided as a separate polynucleotide. In some embodiments, the recombination template is designed to function as a template for homologous recombination within or near the target sequence that is nicked or cleaved by a CRISPR enzyme as part of a CRISPR complex. The template polynucleotide can be of any suitable length, for example, about or more than about 10, 15, 20, 25, 50, 75, 100, 150, 200, 500, 1000 or more nucleotides in length. In some embodiments, the template polynucleotide is complementary to a portion of the polynucleotide that comprises the target sequence. When optimally aligned, the template polynucleotide can overlap with one or more nucleotides of the target sequence (for example, about or more than about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100 or more nucleotides). In some embodiments, when polynucleotides comprising a template sequence and a target sequence are optimally aligned, the nearest neighbor nucleotide of the template polynucleotide is within about 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 1000, 5000, 10000, or more nucleotides of the target sequence.
[0078] In some embodiments, the CRISPR enzyme is part of a fusion protein containing one or more heterologous protein domains (e.g., about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more other domains of the CRISPR enzyme). The CRISPR enzyme fusion protein can include any additional protein sequence, and optionally a linker sequence between any two domains. Examples of protein domains that can be fused to the CRISPR enzyme include, but are not limited to, epitope tags, reporter gene sequences, and protein domains with one or more of the following activities: methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, and nucleic acid binding activity. Non-limiting examples of epitope tags include histidine (His) tag, V5 tag, FLAG tag, influenza hemagglutinin (HA) tag, Myc tag, VSV-G tag, and thioredoxin (Trx) tag. Examples of reporter genes include, but are not limited to, glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and autofluorescent proteins, such as blue fluorescent protein (BFP). CRISPR enzymes can be fused to gene sequences encoding proteins or protein fragments that bind to DNA molecules or other cellular molecules, such as, but not limited to, maltose binding protein (MBP), S-tags, Lex A DNA binding domain (DBD) fusions, GAL4 DNA binding domain fusions, and herpes simplex virus (HSV) BP16 protein fusions. Additional domains that can form part of fusion proteins comprising CRISPR enzymes are described in U.S. Patent Application Publication No. 20110059502, which is incorporated herein by reference. In some embodiments, tagged CRISPR enzymes are used to identify the location of the target sequence.
[0079] In some embodiments, the present invention provides methods comprising delivering one or more polynucleotides, such as, for example, one or more vectors described herein, one or more transcripts thereof, and / or one or more proteins transcribed therefrom, to a host cell. In some embodiments, the present invention further provides cells produced by such cells, and organisms (e.g., animals, plants, or fungi) comprising or produced from such cells. In some embodiments, a CRISPR enzyme in combination with (and optionally complexed with) a guide sequence is delivered to the cell. Conventional viral and non-viral gene transfer methods can be used to introduce nucleic acids into mammalian cells or target tissues. Such methods can be used to administer nucleic acids encoding components of a CRISPR system to cells in culture or into a host organism. Non-viral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of vectors described herein), naked nucleic acids, and nucleic acids complexed with a delivery vehicle, such as a liposome. Viral vector delivery systems include DNA and RNA viruses with episomal or integrated genomes after delivery to the cell.For an overview of gene therapy procedures, see Anderson, Science 256:808-813 (1992); Nabel & Felgner, TIBTECH 11:211-217 (1993); Mitani & Caskey, TIBTECH 11:162-166 (1993); Dillon, TIBTECH 11:167-175 (1993); Miller, Nature 357:455-460 (1992); Van Brunt, Biotechnology 6(10):1149-1154 (1988); Vigne, Restorative Neurology and Neuroscience 8:35-36 (1995); Kremer & Perricaudet, British Medical Bulletin 51(1):31-44 (1995); Haddada et al., in Current Topics in Microbiology and See Immunology Doerfler and Boehm (eds) (1995); and Yu et al., Gene Therapy 1:13-26 (1994).
[0080] Non-viral nucleic acid delivery methods include lipofection, nucleofection, microinjection, gene gun, virosome, liposome, immunoliposome, polycation or lipid:nucleic acid conjugate, naked DNA, artificial virion, and drug-enhanced DNA uptake.Lipofection is described, for example, in U.S. Pat. Nos. 5,049,386, 4,946,787; and 4,897,355, and lipofection reagents are commercially available (for example, Transfectam™ and Lipofectin™).Cationic and neutral lipids suitable for efficient receptor-recognition lipofection of polynucleotides include those of Felgner, WO 91 / 17424; WO 91 / 16024.Delivery can be to cells (for example, in vitro or ex vivo administration) or target tissue (for example, in vivo administration).
[0081] The preparation of lipid:nucleic acid complexes, including targeted liposomes, e.g., immunolipid complexes, is well known to those skilled in the art (see, e.g., Crystal, Science 270:404-410 (1995); Blaese et al., Cancer Gene Ther. 2:291-297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994); Gao et al., Gene Therapy 2:710-722 (1995); Ahmad et al., Cancer Res. 52:4817-4820 (1992); see U.S. Patent Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028, and 4,946,787).
[0082] The use of RNA or DNA virus-based systems for nucleic acid delivery takes advantage of a highly evolved process that targets viruses to specific cells in the body and transports the viral payload to the nucleus. Viral vectors can be administered directly to patients (in vivo) or used to treat cells in vitro, and the modified cells can then be administered to patients (ex vivo). Conventional virus-based systems include retroviral, lentiviral, adenoviral, adeno-associated, and herpes simplex virus vectors for gene transfer. Integration into the host genome is possible with retroviral, lentiviral, and adeno-associated virus gene transfer methods, often resulting in long-term expression of the inserted transgene. Furthermore, high transduction efficiencies have been observed in many different cell types and target tissues.
[0083] The tropism of retroviruses can be altered by incorporating foreign envelope proteins, expanding the potential target population of target cells. Lentiviral vectors are retroviral vectors that can transduce or infect non-dividing cells and typically produce high viral titers. Therefore, the choice of retroviral gene transfer system depends on the target tissue. Retroviral vectors contain cis-acting long terminal repeats (LTRs) that have packaging capacity for foreign sequences up to 6-10 kb. Minimal cis-acting LTRs regulate vector replication and This is sufficient for the synthesis and packaging of the vector, which is then used to integrate a therapeutic gene into target cells to provide permanent transgene expression. Widely used retroviral vectors include those based on murine leukemia virus (MuLV), gibbon ape leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), or combinations thereof (see, e.g., Buchscher et al., J. Virol. 66:2731-2739 (1992); Johann et al., J. Virol. 66:1635-1640 (1992); Sommnerfelt et al., Virol. 176:58-59 (1990); Wilson et al., J. Virol. 63:2374-2378 (1989); Miller et al., J. Virol. 65:2220-2224 (1991); PCT / US94 / 05700).
[0084] In applications where transient expression is preferred, adenovirus-based systems can be used. Adenovirus-based vectors can exhibit extremely high transduction efficiency in many cell types and do not require cell division. High titers and expression levels have been obtained with such vectors. This vector can be produced in large quantities using a relatively simple system. For example, adeno-associated virus ("AAV") vectors can also be used to transduce cells with target nucleic acids in the in vitro production of nucleic acids and peptides, and for in vivo and ex vivo gene therapy procedures (see, e.g., West et al., Virology 160:38-47 (1987); U.S. Pat. No. 4,797,368; WO 93 / 24641; Kotin, Human Gene Therapy 5:793-801 (1994); Muzyczka, J. Clin. Invest. 94:1351 (1994). The construction of recombinant AAV vectors is described in numerous publications, e.g., U.S. Pat. No. 5,173,414; Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985); Tratschin, et al. al., Mol. Cell. Biol. 4:2072-2081 (1984); Hermonat & Muzyczka, PNAS 81:6466-6470 (1984); and Samulski et al., J. Virol. 63:03822-3828 (1989).
[0085] Typically, packaging cells are used to form viral particles that can infect host cells. Such cells include 293 cells, which package adenovirus, and φ2 or PA317 cells, which package retrovirus. Viral vectors used in gene therapy are usually produced by creating cell lines that package nucleic acid vectors into viral particles. The vector typically contains the minimum viral sequences required for packaging and subsequent integration into the host, with other viral sequences replaced by an expression cassette for the polynucleotide to be expressed. Defective viral functions are typically provided in trans by the packaging cell line. For example, AAV vectors used in gene therapy typically contain only the ITR sequences from the AAV genome required for packaging and integration into the host genome. Viral DNA is packaged into a cell line containing a helper plasmid encoding other AAV genes, i.e., rep and cap, but lacking ITR sequences. The cell line can also be infected with adenovirus as a helper. Helper virus promotes the replication of AAV vector and the expression of AAV gene from helper plasmid.Helper plasmid is not packaged in significant amounts due to the lack of ITR sequence.Contamination by adenovirus can be reduced, for example, by heat treatment, to which adenovirus is more sensitive than AAV.Additional methods for delivering nucleic acid to cells are known to those skilled in the art.For example, see US Patent Application Publication No. 20030087817, which is incorporated herein by reference.
[0086] In some embodiments, host cells are transiently or non-transiently transfected with one or more vectors described herein. In some embodiments, cells are transfected as they naturally occur in a subject. In some embodiments, the transfected cells are harvested from a subject. In some embodiments, the cells are derived from cells harvested from a subject, e.g., cell lines. A wide variety of cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to, C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLa-S3, Huh1, Huh4, Huh7, HUVEC, HASMC, HEKn, HEKa, MiaPaCell, Panc1, PC-3, TF1, CTLL-2, C1R, Rat6, CV1, RPTE, A10, T24, J82, A375, ARH-77, Calu1, SW480, SW620, SKOV3, SK-UT, CaCo2, P388D1, SEM-K2, WEHI-231, HB56, TIB55, Jurkat, J45.01, LRMB, Bcl-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, and Hep G2, HeLa B, HeLa T4, COS, COS-1, COS-6, COS-M6A, BS-C-1 monkey kidney epithelium, BALB / 3T3 mouse embryonic fibroblasts, 3T3Swiss, 3T3-L1, 132-d5 human fetal fibroblasts; 10.1 mouse fibroblasts, 293-T, 3T3, 721, 9L, A2780, A2780ADR, A2780cis , A172, A20, A253, A431, A-549, ALC, B16, B35, BCP-1 cells, BEAS-2B, bEnd.3, BHK-21, BR29 3, BxPC3, C3H-10T1 / 2, C6 / 36, Cal-27, CHO, CHO-7, CHO-IR, CHO-K1, CHO-K2, CHO-T, CHO Dhfr- / -, COR-L23, COR-L23 / CPR, COR-L23 / 5010, COR-L23 / R23, COS-7, COV-434, CML T1, CMT, CT26, D17, DH82, DU145, DuCaP, EL4, EM2, EM3, EMT6 / AR1, EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA2, HEK-293, HeLa, Hepa1c1c7, HL-60, HMEC, HT-29, Jurkat, JY cells, K562 cells, Ku 812, KCL22, KG1, KYO1, LNCap, Ma-Mel1-48, MC-38, MCF-7, MCF-10A, MDA-MB-231, MDA-MB-468, MDA-MB-435, MDCK II, MDCK II, MOR / 0.2R, MONO-MAC6, MTD-1A, MyEnd, NCI-H69 / CPR, NCI-H69 / LX10, NCI-H69 / LX20, NCI-H69 / LX4, NIH-3T3, NALM-1, NW-145, OPCN / OPCT cell lines, Peer, PNT-1A / PNT2, RenCa, RIN-5F, RMA / RMAS, Saos-2 cells, Sf-9, SkBr3, T2, T-47D, T84, THP1 cell lines, U373, U87, U937, VCaP, Vero cells, WM39, WT-49, X63, YAC-1, YAR, and transgenic variants thereof. Cell lines may be of any species known to those skilled in the art. These vectors are available from various sources (see, e.g., American Type Culture Collection (ATCC), Manassus, Va.). In some embodiments, cells transfected with one or more vectors described herein are used to establish new cell lines comprising one or more vector-derived sequences. In some embodiments, cells transiently transfected with components of a CRISPR system described herein (e.g., by transient transfection of one or more vectors or transfection with RNA) and modified through the activity of a CRISPR complex are used to establish new cell lines comprising cells containing the modification but lacking any other exogenous sequences. In some embodiments, cells transiently or non-transiently transfected with one or more vectors described herein, or cell lines derived from such cells, are used in the evaluation of one or more test compounds.
[0087] In some embodiments, one or more vectors described herein are used to produce non-human transgenic animals or transgenic plants. In some embodiments, the transgenic animal is a mammal, such as a mouse, rat, or rabbit. In some embodiments, the organism or subject is a plant. In some embodiments, the organism, subject, or plant is algae. Methods for producing transgenic plants and animals are known in the art and generally begin, for example, with the cell transfection methods described herein. Transgenic plants, particularly crops and algae, as well as transgenic animals are provided. Transgenic animals or plants may be useful in applications other than providing disease models. These may include, for example, food or feed production through the expression of higher protein, carbohydrate, nutrient, or vitamin levels than normally found in wild-type plants. In this regard, transgenic plants, particularly legumes and tubers, and animals, particularly mammals, such as livestock (cattle, sheep, goats, and pigs), as well as poultry and edible insects, are preferred.
[0088] Transgenic algae or other plants, such as rapeseed, can be particularly useful, for example, in the production of vegetable oils or biofuels, such as alcohols (especially methanol and ethanol). They can be engineered to express or overexpress high levels of oils or alcohols for use in the oil or biofuel industry.
[0089] In one aspect, the present invention provides a method for modifying target polynucleotide in eukaryotic cells.In some embodiments, the method comprises: allowing CRISPR complex to bind to target polynucleotide, causing the cleavage of said target polynucleotide, thereby modifying said target polynucleotide, wherein the CRISPR complex comprises a CRISPR enzyme complexed with the guide sequence that hybridizes with the target sequence in said target polynucleotide, and said guide sequence is then bound to the tracr mate sequence that hybridizes with the tracr sequence.
[0090] In one aspect, the present invention provides a method for modifying the expression of polynucleotide in eukaryotic cells.In some embodiments, the method comprises: allowing CRISPR complex to bind to polynucleotide, and thereby causing the binding to increase or decrease the expression of said polynucleotide; CRISPR complex comprises CRISPR enzyme complexed with guide sequence that hybridizes with target sequence in said target polynucleotide, and said guide sequence is then bound to tracr mate sequence that hybridizes with tracr sequence.
[0091] With recent advances in crop genomics, the ability to use CRISPR-Cas systems to perform efficient and cost-effective gene editing and manipulation allows for the rapid selection and comparison of single and multiplexed genetic manipulations to transform such genomes for improved production and trait enhancement. In this regard, reference is made to U.S. patents and publications: U.S. Patent No. 6,603,061 - Agrobacterium-Mediated Plant Transformation Method; U.S. Patent No. 7,868,149 - Plant Genome Sequences and Uses Thereof and U.S. Patent Application Publication No. 2009 / 0100536 - Transgenic Plants with Enhanced Agronomic Traits, the entire contents and disclosures of each of which are incorporated herein by reference in their entirety. In implementing the present invention, the contents and disclosures of Morrell et al. "Crop genomics: advances and applications" Nat Rev Genet. 2011 Dec 29; 13(2): 85-96 are also incorporated herein by reference in their entirety. In an advantageous embodiment of the present invention, the CRISPR / Cas9 system is used to engineer microalgae (Example 14). Thus, references herein to animal cells may, mutatis mutandis, also apply to plant cells unless otherwise indicated.
[0092] In one aspect, the present invention provides a method for modifying a target polynucleotide in a eukaryotic cell, which may be in vivo, ex vivo, or in vitro. In some embodiments, the method includes sampling a cell or a population of cells from a human or non-human animal or plant (including microalgae) and modifying one or more cells. Culturing can be performed ex vivo at any stage. One or more cells can also be reintroduced into a non-human animal or plant (including microalgae).
[0093] In one aspect, the present invention provides a kit containing any one or more of the elements disclosed in the above methods and compositions. In some embodiments, the kit includes a vector system and instructions for use of the kit. In some embodiments, the vector system includes: (a) a first regulatory element operably linked to a tracr mate sequence and one or more insertion sites for inserting a guide sequence upstream of the tracr mate sequence (the guide sequence, when expressed, directs the sequence-specific binding of a CRISPR complex to a target sequence in a eukaryotic cell, the CRISPR complex comprising (1) a guide sequence hybridized to the target sequence, and (2) a CRISPR enzyme complexed with the tracr mate sequence hybridized to the tracr sequence); and / or (b) a second regulatory element operably linked to an enzyme coding sequence encoding the CRISPR enzyme comprising a nuclear localization sequence. The elements can be provided individually or in combination, and can be provided in any suitable container, for example, a vial, a bottle, or a tube. In some embodiments, the kit includes instructions in one or more languages, for example, in two or more languages.
[0094] In some embodiments, the kit includes one or more reagents used in a process utilizing one or more of the elements described herein. The reagents can be provided in any suitable container. For example, the kit may provide one or more reaction or storage buffers. The reagents can be provided in a form usable in a particular assay or in a form that requires the addition of one or more other components prior to use (e.g., a concentrate or lyophilized form). The buffer can be any buffer, including, but not limited to, sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In some embodiments, the buffer has a pH of about 7 to about 10. In some embodiments, the kit includes one or more oligonucleotides corresponding to the guide sequence for insertion into a vector to operably link the guide sequence and regulatory elements. In some embodiments, the kit includes a homologous recombination template polynucleotide.
[0095] In one embodiment, the present invention provides a method for using one or more elements of the CRISPR system.The CRISPR complex of the present invention provides an effective means for modifying target polynucleotide.The CRISPR complex of the present invention has a wide range of uses, for example, modifying (for example, deletion, insertion, translocation, inactivation, activation) target polynucleotide in a large number of cell types.Therefore, the CRISPR complex of the present invention has a wide range of applications, for example, in gene therapy, drug screening, disease diagnosis and prognosis.An exemplary CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence that hybridizes to a target sequence in a target polynucleotide.The guide sequence is then linked to a tracr mate sequence that hybridizes to a tract sequence.
[0096] The target polynucleotide of a CRISPR complex can be any polynucleotide that is endogenous or exogenous to a eukaryotic cell. For example, the target polynucleotide can be a polynucleotide that remains in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence that encodes a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). Without being bound by theory, it is believed that the target sequence must associate with a PAM (protospacer adjacent motif); i.e., a short sequence recognized by the CRISPR complex. While the exact sequence and length requirements for the PAM vary depending on the CRISPR enzyme used, the PAM is typically a 2-5 base pair sequence adjacent to the protospacer (i.e., the target sequence). Exemplary PAM sequences are provided in the Examples section below, and those skilled in the art can identify additional PAM sequences for use with a given CRISPR enzyme.
[0097] Target polynucleotides for CRISPR complexes can include many of the disease-associated genes and polynucleotides and signaling biochemical pathway-associated genes and polynucleotides listed in U.S. Provisional Patent Applications Nos. 61 / 736,527 and 61 / 748,427 (both entitled SYSTEMS METHODS AND COMPOSITIONS FOR SEQUENCE MANIPULATION, filed December 12, 2012 and January 2, 2013, respectively), having Broad Reference Nos. BI-2011 / 008 / WSGR Docket No. 44063-701.101 and BI-2011 / 008 / WSGR Docket No. 44063-701.102, respectively, the entire contents of which are incorporated herein by reference in their entireties.
[0098] Examples of target polynucleotides include sequences related to signaling biochemical pathways, such as signaling biochemical pathway-related genes or polynucleotides. Examples of target polynucleotides include disease-related genes or polynucleotides. "Disease-related" genes or polynucleotides refer to any gene or polynucleotide that produces transcription or translation products at abnormal levels or in abnormal forms in cells derived from diseased tissues compared with tissues or cells of non-disease control. This can be a gene that is expressed at abnormally high levels; or a gene that is expressed at abnormally low levels, and expression changes correlate with the occurrence and / or progression of disease. Disease-related genes also refer to genes that are directly responsible for the pathogenesis of disease or have mutations or genetic variations that are in linkage disequilibrium with genes that are responsible for the pathogenesis of disease. The transcription or translation products can be known or unknown, and can be at normal or abnormal levels.
[0099] Examples of disease-associated genes and polynucleotides are available from the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Md.), and are available on the World Wide Web.
[0100] Examples of disease-associated genes and polynucleotides are listed in Tables A and B. Disease-specific information is available from the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Md.), and is available on the World Wide Web. Examples of signaling biochemical pathway-associated genes and polynucleotides are listed in Table C.
[0101] The mutation in these genes and pathways can cause inappropriate protein production or inappropriate amount of protein, which affects function.More examples of genes, diseases and proteins are incorporated herein by reference from US provisional patent application No. 61 / 736,527 and No. 61 / 748,427.Such genes, proteins and pathways can be the target polynucleotide of CRISPR complex.
[0102] [Table 1]
[0103] [Table 2]
[0104] [Table 3]
[0105] [Table 4]
[0106] [Table 5]
[0107] Table 6
[0108] Table 7
[0109] Table 8
[0110] Table 9
[0111] Table 10
[0112] Table 11
[0113] Table 12
[0114] Table 13
[0115] Table 14
[0116] Table 15
[0117] [Table 16]
[0118] Embodiments of the present invention also relate to methods and compositions related to gene knockout, gene amplification, and repair of specific mutations associated with DNA repeat instability and neurological diseases (Robert D. Wells, Tetsuo Ashizawa, Genetic Instabilities and Neurological Diseases, Second Edition, Academic Press, October 13, 2011 - Medical). Tandem repeat sequences of defined nature have been found to be responsible for more than 20 human diseases (New insights into repeat instability: role of RNA-DNA hybrids. McIvor EI, Polak U, Napierala M. RNA Biol. 2010 September-October;7(5):551-8). The CRISPR-Cas system can be used to correct these abnormalities of genomic instability.
[0119] A further aspect of the present invention relates to the use of CRISPR-Cas systems to correct abnormalities in the EMP2A and EMP2B genes, which have been identified as being associated with Lafora disease. Lafora disease is an autosomal recessive condition characterized by progressive myoclonic epilepsy that may begin as epileptic seizures in adolescence. Some cases of the disease may be caused by mutations in as-yet-unidentified genes. The disease causes seizures, muscle spasms, difficulty walking, dementia, and ultimately death. Currently, no treatments have been proven effective against disease progression. Other genetic abnormalities associated with epilepsy can also be targeted using CRISPR-Cas systems, and the underlying genetics are further described in Genetics of Epilepsy and Genetic Epilepsies, edited by Giuliano Avanzini and Jeffrey L. Noebels, Mariani Foundation Paediatric Neurology: 20; 2009).
[0120] In yet another embodiment of the present invention, the CRISPR-Cas system can be used to correct eye defects resulting from several genetic mutations, as further described in Genetic Diseases of the Eye, Second Edition, edited by Elias I. Traboulsi, Oxford University Press, 2012.
[0121] Some further aspects of the present invention relate to the correction of abnormalities associated with a wide range of genetic diseases, which are further described on the National Institutes of Health website under the topic subsection Genetic Disorders. Genetic brain diseases include, but are not limited to, adrenoleukodystrophy, agenesis of the corpus callosum, Aicardi syndrome, Alpers disease, Alzheimer's disease, Barth syndrome, Batten disease, CADASIL, cerebellar degeneration, Fabry disease, Gerstmann-Straussler-Scheinker disease, Huntington's disease and other triplet repeat diseases, Leigh disease, Lesch-Nyhan syndrome, Menkes disease, mitochondrial myopathy, and NINDS colpocephaly. These diseases are further described on the National Institutes of Health website under the topic subsection Genetic Brain Disorders.
[0122] In some embodiments, the condition is neoplasia. In some embodiments, where the condition may be neoplasia, the gene to be targeted may be any of those listed in Table A (such as PTEN in this case). In some embodiments, the condition may be age-related macular degeneration. In some embodiments, the condition may be schizophrenia. In some embodiments, the condition may be a trinucleotide repeat disorder. In some embodiments, the condition may be fragile X syndrome. In some embodiments, the condition may be a secretase-associated disorder. In some embodiments, the condition may be a prion-associated disorder. In some embodiments, the condition may be ALS. In some embodiments, the condition may be drug addiction. In some embodiments, the condition may be autism. In some embodiments, the condition may be Alzheimer's disease. In some embodiments, the condition may be inflammation. In some embodiments, the condition may be Parkinson's disease.
[0123] Examples of proteins associated with Parkinson's disease include, but are not limited to, alpha-synuclein, DJ-1, LRRK2, PINK1, parkin, UCHL1, synphilin-1, and NURR1.
[0124] An example of a preference-related protein is ABAT.
[0125] Examples of inflammation-related proteins include monocyte chemoattractant protein-1 (MCP1) encoded by the Ccr2 gene, CC chemokine receptor type 5 (CCR5) encoded by the Ccr5 gene, IgG receptor IIB (FCGR2b, also known as CD32) encoded by the Fcgr2b gene, or Fc epsilon R1g (FCER1g) protein encoded by the Fcer1g gene.
[0126] Examples of cardiovascular disease-related proteins include, for example, IL1B (interleukin 1, beta), XDH (xanthine dehydrogenase), TP53 (tumor protein p53), PTGIS (prostaglandin I2 (prostacyclin) synthase), MB (myoglobin), IL4 (interleukin 4), ANGPT1 (angiopoietin 1), ABCG8 (ATP-binding cassette, subfamily G (WHITE), member 8), or CTSK (cathepsin K).
[0127] Examples of Alzheimer's disease-related proteins include, for example, the very low density lipoprotein receptor protein (VLDLR) encoded by the VLDLR gene, the ubiquitin-like modifier activating enzyme 1 (UBA1) encoded by the UBA1 gene, or the NEDD8 activating enzyme E1 catalytic subunit protein (UBE1C) encoded by the UBA3 gene.
[0128] Examples of proteins associated with autism spectrum disorders include, for example, benzodiazepine receptor (peripheral)-associated protein 1 (BZRAP1) encoded by the BZRAP1 gene, AF4 / FMR2 family member 2 protein (AFF2) encoded by the AFF2 gene (also known as MFR2), fragile X mental retardation autosomal homolog 1 protein (FXR1) encoded by the FXR1 gene, or fragile X mental retardation autosomal homolog 2 protein (FXR2) encoded by the FXR2 gene.
[0129] Examples of proteins associated with macular degeneration include the ATP-binding cassette subfamily A (ABC1) member 4 protein (ABCA4) encoded by the ABCR gene, the apolipoprotein E protein (APOE) encoded by the APOE gene, or the chemokine (CC motif) ligand 2 protein (CCL2) encoded by the CCL2 gene.
[0130] Examples of proteins associated with schizophrenia include NRG1, ErbB4, CPLX1, TPH1, TPH2, NRXN1, GSK3A, BDNF, DISC1, GSK3B, and combinations thereof.
[0131] Examples of proteins involved in tumor suppression include ATM (ataxia telangiectasia mutated), ATR (ataxia telangiectasia and Rad3 related), EGFR (epidermal growth factor receptor), ERBB2 (v-erb-b2 erythroblastic leukemia viral oncogene homolog 2), ERBB3 (v-erb-b2 erythroblastic leukemia viral oncogene homolog 3), ERBB4 (v-erb-b2 erythroblastic leukemia viral oncogene homolog 4), Notch1, Notch2, Notch3, or Notch4.
[0132] Examples of proteins associated with secretase disorders can include, for example, PSENEN (presenilin enhancer 2 homolog (C. elegans)), CTSB (cathepsin B), PSEN1 (presenilin 1), APP (amyloid beta (A4) precursor protein), APH1B (anterior pharyngeal defect 1 homolog B (C. elegans)), PSEN2 (presenilin 2 (Alzheimer's disease 4)), or BACE1 (beta-site APP cleaving enzyme 1).
[0133] Examples of proteins associated with amyotrophic lateral sclerosis include SOD1 (superoxide dismutase 1), ALS2 (amyotrophic lateral sclerosis 2), FUS (fused in sarcoma), TARDBP (TAR DNA binding protein), VAGFA (vascular endothelial growth factor A), VAGFB (vascular endothelial growth factor B), and VAGFC (vascular endothelial growth factor C), and any combination thereof.
[0134] Examples of proteins associated with prion diseases can include SOD1 (superoxide dismutase 1), ALS2 (amyotrophic lateral sclerosis 2), FUS (fused in sarcoma), TARDBP (TAR DNA binding protein), VAGFA (vascular endothelial growth factor A), VAGFB (vascular endothelial growth factor B), and VAGFC (vascular endothelial growth factor C), and any combination thereof.
[0135] Examples of proteins associated with neurodegenerative pathology in prion disorders include, for example, A2M (alpha-2-macroglobulin), AATF (apoptosis antagonistic transcription factor), ACPP (prostatic acid phosphatase), ACTA2 (aortic smooth muscle actin alpha 2), ADAM22 (ADAM metallopeptidase domain), ADORA3 (adenosine A3 receptor), or ADRA1D (alpha-1D adrenergic receptor for alpha-1D adrenoceptor).
[0136] Examples of proteins associated with immunodeficiency include, for example, A2M [alpha-2-macroglobulin]; AANAT [arylalkylamine N-acetyltransferase]; ABCA1 [ATP-binding cassette subfamily A (ABC1), member 1]; ABCA2 [ATP-binding cassette subfamily A (ABC1), member 2]; or ABCA3 [ATP-binding cassette subfamily A (ABC1), member 3].
[0137] Examples of proteins associated with trinucleotide repeat disorders include, for example, AR (androgen receptor), FMR1 (Fragile X Mental Retardation 1), HTT (Huntington's), or DMPK (Myotonic Dystrophy Protein Kinase), FXN (Frataxin), and ATXN2 (Ataxin 2).
[0138] Examples of proteins associated with impaired neurotransmission include, for example, SST (somatostatin), NOS1 (nitric oxide synthase 1 (neuronal type)), ADRA2A (adrenergic alpha-2A receptor), ADRA2C (adrenergic alpha-2C receptor), TACR1 (tachykinin receptor 1), or HTR2c (5-hydroxytryptamine (serotonin) receptor 2C).
[0139] Examples of neurodevelopment-related sequences include, for example, A2BP1 [ataxin 2-binding protein 1], AADAT [aminoadipate aminotransferase], AANAT [arylalkylamine N-acetyltransferase], ABAT [4-aminobutyrate aminotransferase], ABCA1 [ATP-binding cassette subfamily A (ABC1) member 1], or ABCA13 [ATP-binding cassette subfamily A (ABC1) member 13].
[0140] Further examples of preferred conditions treatable by the system of the present invention can be selected from the following: Alcardi-Goutières syndrome; Alexander disease; Allan-Herndon-Dudley syndrome; POLG-related disorders; alpha-mannosidosis (types II and III); Alström syndrome; Angelman syndrome; ataxia-telangiectasia; neuronal ceroid lipofuscinosis; beta-thellasamia; bilateral optic atrophy and (infantile) optic atrophy type 1; retinoblastoma (bilateral); Canavan disease; cerebro-ocular-facial-skeletal syndrome 1 [COFS1]; cerebrotendinous xanthomas Cornelia de Lange syndrome; MAPT-related disorders; inherited prion diseases; Dravet syndrome; early-onset familial Alzheimer's disease; Friedreich ataxia [FRDA]; Flins syndrome; fucosidosis; Fukuyama-type congenital muscular dystrophy; galactosialidosis; Gaucher disease; organic acidemia; hemophagocytic lymphohistiocytosis; Hutchinson-Gilford progeria syndrome; mucolipidosis type II; infantile free sialic acid storage disease; PLA2G6-associated neurodegeneration; Jervell-Lange-Nielsen syndrome; junctional epidermolysis bullosa; Huntington's disease; Krabbe disease (infantile form) ; Mitochondrial DNA-related Leigh syndrome and NARP; Lesch-Nyhan syndrome; LIS1-related lissencephaly; Lowe syndrome; Maple syrup urine disease; MECP2 duplication syndrome; ATP7A-related copper transport disorders; LAMA2-related muscular dystrophy; Arylsulfatase A deficiency; Mucopolysaccharidosis type I, II, or III; Peroxisomal biogenesis disorders, Zellweger syndrome spectrum; Neurodegeneration with cerebral iron storage; Acid sphingomyelinase deficiency; Niemann-Pick disease type C; Glycine encephalopathy; ARX-related disorders; Urea cycle disorders; COL1A1 / 2-related diaphyseal dysplasia; mitochondrial DNA deletion syndrome; PLP1-related disorders; Perry syndrome; Phelan-McDermott syndrome; glycogen storage disease type II (Pompe disease) (infantile form); MAPT-related disorders; MECP2-related disorders; rhizomelic chondrodysplasia punctata type 1; Roberts syndrome; Sandhoff disease; Schindler disease type 1; adenosine deaminase deficiency; Smith-Lemli-Opitz syndrome; spinal muscular atrophy; childhood-onset spinocerebellar ataxia; hexosaminidase A deficiency; lethal dysplasia type 1; collagen type VI-related disorders; Usher syndrome type 1; congenital muscular dystrophy;Wolf-Hirschhorn syndrome; lysosomal acid lipase deficiency; and xeroderma pigmentosum.
[0141] Chronic administration of protein therapeutics can induce unacceptable immune responses against specific proteins. The immunogenicity of protein drugs can be attributed to several immunodominant helper T lymphocyte (HTL) epitopes. Reducing the MHC binding affinity of these HTL epitopes contained within these proteins can generate drugs with lower immunogenicity (Tangri S, et al. ("Rationally engineered therapeutic proteins with reduced immunogenicity" J Immunol. 2005 Mar 15;174(6):3187-96). In the present invention, the immunogenicity of CRISPR enzymes can be reduced, particularly according to the approach first described and subsequently developed by Tangri et al. for erythropoietin. Thus, directed evolution or rational design can be used to reduce the immunogenicity of CRISPR enzymes (e.g., Cas9) in host species (humans or other species).
[0142] In plants, pathogens are often host-specific. For example, Fusarium oxysporum f.sp. lycopersici, which causes tomato wilt, attacks only tomatoes, while F. oxysporum f. dianthiii and Puccinia graminis f.sp. tritici, which cause carnation wilt, attack only wheat. Plants have pre-existing and inducible defenses to resist most pathogens. Mutation and recombination events over plant generations result in genetic variations that cause susceptibility, especially when the pathogen outgrows the plant. In plants, non-host resistance can exist, e.g., the host and pathogen are incompatible. Horizontal resistance, e.g., partial resistance to all species of pathogens, typically controlled by many genes, and vertical resistance, e.g., complete resistance to some species of pathogens but not others, typically controlled by a small number of genes, can also exist. At the gene-by-gene level, plants and pathogens evolve together, with genetic changes in some balanced by others. Thus, using natural variation, breeders combine genes that are most useful for yield, quality, uniformity, cold tolerance, and resistance. Sources of resistance genes include natural or exotic species, landraces, wild plant relatives, and induced mutations, such as treating plant material with mutagens. The present invention provides plant breeders with new tools for inducing mutations. Thus, those skilled in the art can analyze the genomes of sources of resistance genes and use the present invention to induce the appearance of resistance genes in varieties with desired characteristics or traits more precisely than traditional mutagens, thus accelerating and improving plant breeding programs.
[0143] It is clear that it is envisioned that the system of the present invention can be used to target any target polynucleotide sequence.Some examples of pathological conditions or diseases that can be usefully treated using the system of the present invention are included in the table above, and examples of genes currently associated with these pathological conditions are also provided in the table.However, the genes exemplified are not exclusive. [Example]
[0144] The following examples are given for the purpose of illustrating various embodiments of the present invention and are not meant to limit the invention in any manner. The examples, together with the methods described herein, are representative and exemplary of presently preferred embodiments and are not intended to limit the scope of the invention. Modifications thereof and other uses encompassed within the spirit of the invention as defined by the scope of the claims will occur to those skilled in the art.
[0145] Example 1: CRISPR complex activity in the nucleus of a eukaryotic cell An exemplary type II CRISPR system is the type II CRISPR locus from Streptococcus pyogenes SF370, which contains a cluster of four genes, Cas9, Cas1, Cas2, and Csn1, and two non-coding RNA elements, tracrRNA and a characteristic array of repeat sequences (direct repeats) spaced by short stretches of non-repetitive sequences (spacers, each approximately 30 bp). In this system, targeted DNA double-strand breaks (DSBs) are generated in four sequential steps (Figure 2A). First, two non-coding RNAs, the pre-crRNA array and tracrRNA, are transcribed from the CRISPR locus. Second, tracrRNA hybridizes to the direct repeats of the pre-crRNA, which are then processed into mature crRNAs containing individual spacer sequences. Third, the mature crRNA:tracrRNA complex directs Cas9 to the DNA target consisting of the protospacer and the corresponding PAM via heteroduplex formation between the spacer region of the crRNA and the protospacer DNA. Finally, Cas9 mediates cleavage of the target DNA upstream of the PAM to create a DSB within the protospacer (Figure 2A). This example describes an exemplary process for adapting this RNA-programmable nuclease system to direct CRISPR complex activity in the nucleus of a eukaryotic cell.
[0146] To improve expression of CRISPR components in mammalian cells, two genes from Streptococcus pyogenes (S. pyogenes) SF370 locus 1, Cas9 (SpCas9) and RNase III (SpRNase III), were codon-optimized. To facilitate nuclear localization, nuclear localization signals (NLSs) were included at the amino (N)- or carboxyl (C)-termini of both SpCas9 and SpRNase III (Figure 2B). To facilitate visualization of protein expression, fluorescent protein markers were also included at the N- or C-termini of both proteins (Figure 2B). A version of SpCas9 with NLSs attached to both the N- and C-termini (2xNLS-SpCas9) was also generated. Constructs containing NLS-fused SpCas9 and SpRNase III were transfected into 293FT human embryonic kidney (HEK) cells, and it was found that the relative positioning of the NLS to SpCas9 and SpRNase III affected their nuclear localization efficiency. While a C-terminal NLS was sufficient to target SpRNase III to the nucleus, attachment of a single copy of these specific NLSs to either the N- or C-terminus of SpCas9 failed to achieve proper nuclear localization in this system. In this example, the C-terminal NLS was that of nucleoplasmin (KRPAATKKAGQAKKKK), and the C-terminal NLS was that of SV40 large T antigen (PKKKRKV). Of the SpCas9 versions tested, only 2xNLS-SpCas9 exhibited nuclear localization (Figure 2B).
[0147] The tracrRNA from the CRISPR locus of Streptococcus pyogenes (S. pyogenes) SF370 has two transcription start sites, generating two transcripts of 89 nucleotides (nt) and 171 nt, which are subsequently processed into identical 75-nt mature tracrRNAs. The shorter 89-nt tracrRNA was selected for expression in mammalian cells (the expression construct illustrated in Figure 6, with functionality determined by the results of the Surveyor assay shown in Figure 6B). The transcription start site is labeled +1, and the transcription terminator and sequence probed by Northern blot are also shown. Expression of the processed tracrRNA was also confirmed by Northern blot. Figure 7C shows the results of Northern blot analysis of total RNA extracted from 293FT cells transfected with long or short tracrRNA and U6 expression constructs carrying SpCas9 and DR-EMX1(1)-DR. The left and right panels are from 293FT cells transfected without or with SpRNase III, respectively. U6 represents a loading control blotted with a probe targeting human U6 snRNA. Transfection of the short tracrRNA expression construct resulted in sufficient levels of the processed form of tracrRNA (approximately 75 bp). Very little long tracrRNA is detected on Northern blots.
[0148] To promote accurate transcription initiation, an RNA polymerase III-based U6 promoter was selected to drive tracrRNA expression (Figure 2C). Similarly, a U6 promoter-based construct was developed to express a pre-crRNA array consisting of a single spacer flanked by two direct repeats (DRs, also encompassed by the term "tracr-mate sequence"; Figure 2C). The first spacer was designed to target a 33-base pair (bp) target site (a 30-bp protospacer and a 3-bp CRISPR motif (PAM) sequence that fulfills the NGG recognition motif of Cas9) in the human EMX1 locus, a key gene in cerebral cortical development (Figure 2C).
[0149] To test whether heterologous expression of a CRISPR system (SpCas9, SpRNase III, tracrRNA, and pre-crRNA) in mammalian cells can achieve targeted cleavage of mammalian chromosomes, HEK293FT cells were transfected with a combination of CRISPR components. Because DSBs in mammalian nuclei are repaired in part by the non-homologous end joining (NHEJ) pathway, which results in the formation of indels, we used the Surveyor assay to detect potential cleavage activity at the target EMX1 locus (see, e.g., Guschin et al., 2010, Methods Mol Biol 649:247). Co-transfection of all four CRISPR components could induce cleavage of up to 5.0% of the protospacer (see Figure 2D). Cotransfection of all CRISPR components except SpRNase III also induced indels in up to 4.7% of the protospacer, suggesting the presence of endogenous mammalian RNases, such as the related Dicer and Drosha enzymes, that may assist crRNA maturation. Removal of any of the remaining three components abolished the genome cleavage activity of the CRISPR system (Figure 2D). Sanger sequencing of amplicons containing the target locus confirmed cleavage activity; five mutant alleles (11.6%) were found among 43 sequenced clones. Similar experiments using various guide sequences yielded indel rates as high as 29% (see Figures 4-8, 10, and 11). These results define a three-component system for efficient CRISPR-mediated genome modification in mammalian cells.
[0150] To optimize cleavage efficiency, we also tested whether different isoforms of tracrRNA affect cleavage efficiency and found that in this exemplary system, only the short (89 bp) transcript form could mediate cleavage of the human EMX1 genomic locus. Figure 9 provides additional Northern blot analysis of crRNA processing in mammalian cells. Figure 9A illustrates a schematic diagram showing an expression vector for a single spacer flanked by two direct repeats (DR-EMX1(1)-DR). The 30 bp spacer and direct repeat sequences targeting the human EMX1 locus protospacer 1 are shown in the lower sequence of Figure 9A. The line indicates the region where the reverse complement sequence was used to generate a Northern blot probe for EMX1(1) crRNA detection. Figure 9B shows Northern blot analysis of total RNA extracted from 293FT cells transfected with a U6 expression construct carrying DR-EMX1(1)-DR. The left and right panels are from 293FT cells transfected without or with SpRNase III, respectively. DR-EMX1(1)-DR was processed to mature crRNA only in the presence of SpCas9, whereas the short tracrRNA was not dependent on the presence of SpRNase III. The mature crRNA detected from transfected 293FT total RNA was approximately 33 bp, shorter than the 39-42 bp mature crRNA from Streptococcus pyogenes (S. pyogenes). These results demonstrate that the CRISPR system can be transplanted into eukaryotic cells and reprogrammed to promote cleavage of endogenous mammalian target polynucleotides.
[0151] Figure 2 illustrates the bacterial CRISPR system described in this example. Figure 2A illustrates a schematic diagram showing CRISPR locus 1 from Streptococcus pyogenes SF370 and the proposed mechanism of CRISPR-mediated DNA cleavage by this system. Mature crRNA processed from the direct repeat-spacer array directs Cas9 to a genomic target consisting of a complementary protospacer and protospacer-adjacent motif (PAM). Upon target-spacer base pairing, Cas9 mediates a double-strand break in the target DNA. Figure 2B illustrates the engineering of S. pyogenes Cas9 (SpCas9) and RNase III (SpRNase III) with a nuclear localization signal (NLS) to enable transport into mammalian nuclei. Figure 2C illustrates mammalian expression of SpCas9 and SpRNase III driven by the constitutive EF1a promoter and tracrRNA and pre-crRNA arrays (DR-spacer-DR) driven by the RNAPol3 promoter U6 to promote accurate transcription initiation and termination. A protospacer from the human EMX1 locus with a sufficient PAM sequence is used as the spacer in the pre-crRNA array. Figure 2D illustrates surveyor nuclease assays for SpCas9-mediated small insertions and deletions. SpCas9 was expressed with or without SpRNase III, tracrRNA, and a pre-crRNA array carrying the EMX1-targeting spacer. Figure 2E illustrates a schematic representation of base pairing between the target locus and the EMX1-targeting crRNA, as well as an example chromatogram showing the microdeletion adjacent to the SpCas9 cleavage site. Figure 2F illustrates mutant alleles identified from sequencing analysis of 43 clonal amplicons exhibiting various microinsertions and deletions. Dotted lines indicate deleted bases, unaligned or mismatched bases indicate insertions or mutations. Scale bar = 10 μm.
[0152] To further simplify the three-component system, we adapted a chimeric crRNA-tracrRNA hybrid design in which the mature crRNA (including the guide sequence) is fused to a partial tracrRNA via a stem-loop to mimic the native crRNA:tracrRNA duplex (Figure 3A).
[0153] Guide sequences can be inserted between the BbsI sites using annealed oligonucleotides. Protospacers on the sense and antisense strands are shown above and below the DNA sequence, respectively. Modification rates of 6.3% and 0.75% were achieved for the human PVALB and mouse Th loci, respectively, demonstrating the broad applicability of the CRISPR system in modifying different loci across multiple organisms. Cleavage was detected for only one of three spacers for each locus using chimeric constructs, whereas all target sequences were cleaved using the co-expressed pre-crRNA configuration. It was cleaved with an efficiency of indel generation reaching 27% (Figs. 4 and 5).
[0154] Figure 5 provides further illustration of the ability of SpCas9 to reprogram and target multiple genomic loci in mammalian cells. Figure 5A provides a schematic diagram of the human EMX1 locus, showing the location of the five protospacers indicated by the underlined sequences. Figure 5B provides a schematic diagram of the pre-crRNA / trcrRNA complex (top panel) showing hybridization between the direct repeat regions of the pre-crRNA and tracrRNA, and a schematic diagram of a chimeric RNA design (bottom panel) containing a 20-bp guide sequence and a tracr-mate and tracr sequence consisting of a partial direct repeat and tracrRNA sequence hybridized to a hairpin structure. Figure 5C illustrates the results of a Surveyor assay comparing the efficacy of Cas9-mediated cleavage at five protospacers in the human EMX1 locus. Either the processed pre-crRNA / tracrRNA complex (crRNA) or a chimeric RNA (chiRNA) was used to target each protospacer.
[0155] Because RNA secondary structure can be important for intermolecular interactions, we used a structure prediction algorithm based on minimum free energy and Boltzmann-weighted structural ensembles to compare the predicted secondary structures of all guide sequences used in our genome targeting experiments (Figure 3B) (see, e.g., Gruber et al., 2008, Nucleic Acids Research, 36:W70). The analysis revealed that, in most cases, effective guide sequences in the chimeric crRNA context were substantially free of secondary structure motifs, while ineffective guide sequences were more likely to form internal secondary structures that could interfere with base pairing with the target protospacer DNA. Therefore, it is conceivable that variability in spacer secondary structure may affect the efficiency of CRISPR-mediated interference when using chimeric crRNAs.
[0156] Figure 3 illustrates exemplary expression vectors. Figure 3A provides a schematic diagram of a synthetic crRNA-tracrRNA chimera (chimeric RNA) and a bicistronic vector for driving expression of SpCas9. The chimeric guide RNA contains a 20-bp guide sequence corresponding to the protospacer in the genomic target site. Figure 3B provides a schematic diagram showing guide sequences targeting the human EMX1, PVALB, and mouse Th loci, as well as their predicted secondary structures. The modification efficiency at each target site is shown below the RNA secondary structure diagram (EMX1, n = 216 amplicon sequencing reads; PVALB, n = 224 reads; Th, n = 265 reads). The folding algorithm produced output colored according to the probability that each base would assume the predicted secondary structure, as indicated by the rainbow scale reproduced in grayscale in Figure 3B. An additional vector design for SpCas9 is shown in Figure 3A and includes a single expression vector incorporating a U6 promoter linked to an insertion site for a guide oligo and a Cbh promoter linked to the SpCas9 coding sequence.
[0157] To test whether spacers containing secondary structures can function in prokaryotic cells where CRISPR naturally operates, we tested the transformation interference of protospacer-bearing plasmids in an Escherichia coli (E. coli) strain heterologously expressing the S. pyogenes SF370 CRISPR locus 1 (Figure 3C). The CRISPR locus was cloned into a low-copy E. coli expression vector, and the crRNA array was replaced with a single spacer flanked by a pair of DRs (pCRISPR). E. coli strains carrying different pCRISPR plasmids were transformed with challenge plasmids containing the corresponding protospacer and PAM sequences (Figure 3C). In bacterial assays, all spacers promoted efficient CRISPR interference (Figure 3C). These results suggest that there may be additional factors that affect the efficiency of CRISPR activity in mammalian cells.
[0158] To investigate the specificity of CRISPR-mediated cleavage, the effect of single-nucleotide mutations in the guide sequence on protospacer cleavage in mammalian genomes was analyzed using a series of EMX1-targeting chimeric crRNAs with single point mutations (Figure 4A). Figure 4B illustrates the results of a Surveyor nuclease assay comparing the cleavage efficiency of Cas9 when paired with different mutant chimeric RNAs. Single-base mismatches up to 12 bp 5' to the PAM virtually abolished genome cleavage by SpCas9, while spacers with mutations further upstream maintained activity against the original protospacer target. The specificity of SpCas9 was maintained (Figure 4B). In addition to the PAM, SpCas9 has single-base specificity within the last 12 bp of the spacer. Furthermore, CRISPR can mediate genome cleavage as efficiently as a pair of TALE nucleases (TALENs) targeting the same EMX1 protospacer. Figure 4C provides a schematic illustrating the design of TALENs targeting EMX1, and Figure 4D shows a Surveyor gel comparing the efficiency of TALENs and Cas9 (n=3).
[0159] Having established a set of components for achieving CRISPR-mediated gene editing in mammalian cells through the error-prone NHEJ mechanism, we tested the ability of CRISPR to stimulate homologous recombination (HR), a high-fidelity gene repair pathway, to create precise edits in the genome. Wild-type SpCas9 can mediate site-specific DSBs that can be repaired through both NHEJ and HR. Furthermore, we engineered an aspartate-to-alanine substitution (D10A) in the RuvC I catalytic domain of SpCas9 to convert the nuclease into a nickase (SpCas9n; illustrated in Figure 5A) (see, e.g., Sapranausaks et al., 2011, Cucleic Acids Research, 39:9275; Gasiunas et al., 2012, Proc. Natl. Acad. Sci. USA, 109:E2579), allowing nicked genomic DNA to undergo high-fidelity homology-directed repair (HDR). Surveyor assays confirmed that SpCas9n did not generate indels in the EMX1 protospacer target. As illustrated in Figure 5B, coexpression of an EMX1-targeting chimeric crRNA with SpCas9 generated indels in the target site, whereas coexpression with SpCas9n did not (n = 3). Furthermore, sequencing of 327 amplicons did not detect any indels induced by SpCas9n. The same locus was selected and CRISPR-mediated HR was tested by cotransfecting HEK293FT cells with a chimeric RNA targeting EMX1, hSpCas9, or hSpCas9n, and an HR template to introduce a pair of restriction sites (HindIII and NheI) near the protospacer. Figure 5C provides a schematic illustration of the HR strategy, along with the relative locations of the recombination sites and primer annealing sequences (arrows). SpCas9 and SpCas9n indeed catalyzed the integration of the HR template into the EMX1 gene.PCR amplification of the target region followed by restriction digestion with HindIII revealed cleavage products corresponding to the predicted fragment sizes (arrows in the restriction fragment length polymorphism gel analysis shown in Figure 5D), indicating that SpCas9 and SpCas9n mediated similar levels of HR efficiency. Applicants further confirmed HR using Sanger sequencing of the genomic amplicon (Figure 5E). These results demonstrate the utility of CRISPR for facilitating targeted gene insertion in mammalian genomes. Given the 14-bp target specificity of wild-type SpCas9 (12 bp from the spacer and 2 bp from the PAM), the availability of a nickase may significantly reduce the potential for off-target modifications, as single-stranded fragments are not substrates for the error-prone NHEJ pathway.
[0160] We constructed an expression construct (Figure 2A) that mimics the natural architecture of CRISPR loci with array spacers to test the feasibility of multiplexed sequence targeting. Using a single CRISPR array encoding a pair of EMX1 and PVALB targeting spacers, we detected efficient cleavage at both loci (Figure 4F, which shows both the schematic design of the crRNA array and a Surveyor blot demonstrating efficient mediation of cleavage). We also tested targeted deletion of a larger genomic region through simultaneous DSBs using spacers for two targets in EMX1, spaced 119 bp apart, and detected a deletion efficiency of 1.6% (3 out of 182 amplicons; Figure 5G). This demonstrates that the CRISPR system can mediate multiplexed editing within a single genome.
[0161] Example 2: CRISPR-based modifications and alternatives The ability to use RNA to program sequence-specific DNA cleavage defines a new class of genome engineering tools for a variety of research and industrial applications. Some aspects of the CRISPR system can be further improved to increase the efficiency and versatility of CRISPR targeting. Optimal Cas9 activity may depend on the availability of free Mg2+ at levels higher than those present in mammalian nuclei (see, e.g., Jinek et al., 2012, Science, 337:816), and the preference for NGG motifs immediately downstream of the protospacer limits targeting ability to an average of every 12 bp in the human genome. Some of these constraints can be overcome by exploiting the diversity of CRISPR loci across microbial metagenomes (see, e.g., Makarova et al., 2011, Nat Rev Microbiol, 9:467). Other CRISPR loci can be transplanted into mammalian cellular environments using methods similar to those described in Example 1. The modification efficiency at each target site is shown below the RNA secondary structure. The algorithm that generates this structure colors each base according to its probability of assuming the predicted secondary structure. RNA guide spacers 1 and 2 induced 14% and 6.4%, respectively. A statistical analysis of cleavage activity across biological replicates at these two protospacer sites is also provided in Figure 7.
[0162] Example 3: Sample target sequence selection algorithm Design a software program to identify candidate CRISPR target sequences on both strands of the input DNA sequence based on the desired guide sequence length and CRISPR motif sequence (PAM) for a given CRISPR enzyme. For example, target sites for Cas9 from Streptococcus pyogenes (S. pyogenes) can be identified by using the PAM sequence NGG to probe for 5'-Nx-NGG-3' on both the input sequence and the reverse complementary strand of the input. Similarly, target sites for Cas9 from S. thermophilus (S. thermophilus) can be identified by using the PAM sequence NNAGAAW to probe for 5'-Nx-NNAGAAW-3' on both the input sequence and the reverse complementary strand of the input. Similarly, target sites for Cas9 in S. thermophilus CRISPR3 can be identified by searching for 5'-Nx-NGGNG-3' on both the input sequence and the reverse complement of the input using the PAM sequence NGGNG. The value "x" in Nx can be fixed by the program or user-defined, e.g., 20.
[0163] Because multiple occurrences of a DNA target site in a genome can result in nonspecific genome editing, after identifying all potential sites, the program filters out sequences based on the number of times the sequence appears in the relevant reference genome. For those CRISPR enzymes whose sequence specificity is determined by a "seed" sequence, e.g., the 11-12 bp 5' from the PAM sequence, including the PAM sequence itself, the filtering step can be based on the seed sequence. Thus, to avoid editing at additional genomic loci, results are filtered based on the number of occurrences of the seed:PAM sequence in the relevant genome. The user can select the length of the seed sequence. The user can also specify the number of occurrences of the seed:PAM sequence in the genome for filtering purposes. The default is to screen for unique sequences. The filtration level can be varied by changing both the length of the seed sequence and the number of occurrences of the sequence in the genome. The program can also, or alternatively, provide a guide sequence complementary to the reported target sequence by providing the reverse complement of the identified target sequence.
[0164] Further details of methods and algorithms for optimizing sequence selection can be found in US Patent Application No. TBA (Broad Reference No. BI-2012 / 084 44790.11.2022), which is incorporated herein by reference.
[0165] Example 4: Evaluation of multiple chimeric crRNA-tracrRNA hybrids This example describes results obtained with chimeric RNAs (chiRNAs; containing guide, tracr mate, and tracr sequences in a single transcript) incorporating different lengths of wild-type tracrRNA sequences. Figure 18a illustrates a schematic diagram of the bicistronic expression vector for chimeric RNA and Cas9. Cas9 is driven by the CBh promoter, and the chimeric RNA is driven by the U6 promoter. The chimeric guide RNA consists of a 20-bp guide sequence (N) linked to truncated tracr sequences (spanning from the first "U" on the lower strand to the end of the transcript) at various positions as indicated. The guide and tracr sequences are separated by the tracr mate sequence GUUUUAGAGCUA, followed by the loop sequence GAAA. Results of SURVEYOR assays for Cas9-mediated indels at the human EMX1 and PVALB loci are illustrated in Figures 18b and 18c, respectively. Arrows indicate the predicted SURVEYOR fragments. ChiRNAs are indicated by their "+n" designation, and crRNA refers to a hybrid RNA in which the guide and tracr sequences are expressed as separate transcripts. Quantification of these results, performed in triplicate, is shown by histograms in Figures 11a and 11b, corresponding to Figures 10b and 10c, respectively ("ND" indicates no indels detected). Protospacer IDs and their corresponding genomic targets, protospacer sequences, PAM sequences, and strand localizations are provided in Table D. Guide sequences were designed to be complementary to the entire protospacer sequence in the case of separate transcripts in the hybrid system, or to only the underlined portion in the case of chimeric RNAs.
[0166] [Table 17]
[0167] Cell culture and transfection Human embryonic kidney (HEK) cell line 293FT (Life Technologies) was maintained in Dulbecco's modified Eagle's medium (DMEM) supplemented with 10% fetal bovine serum (HyClone), 2 mM GlutaMAX (Life Technologies), 100 U / mL penicillin, and 100 μg / mL streptomycin at 37°C with 5% CO2 incubation. 293FT cells were seeded onto 24-well plates (Corning) at a density of 150,000 cells per well 24 hours prior to transfection. Cells were transfected using Lipofectamine 2000 (Life Technologies) according to the manufacturer's recommended protocol. A total of 500 ng of plasmid was used for each well of the 24-well plate.
[0168] SURVEYOR assay for genome modifications 293FT cells were transfected with the above plasmid DNA. Cells were incubated at 37°C for 72 hours after transfection before genomic DNA extraction. Genomic DNA was extracted using QuickExtract DNA Extraction Solution (Epicentre) according to the manufacturer's protocol. Briefly, pelleted cells were resuspended in QuickExtract solution and incubated at 65°C for 15 minutes and 98°C for 10 minutes. Genomic regions flanking the CRISPR target site for each gene were PCR amplified (primers listed in Table E), and the products were purified using QiaQuick Spin Columns (Qiagen) according to the manufacturer's protocol. A total of 400 ng of purified PCR product was mixed with 2 μl of 10× Taq DNA Polymerase PCR buffer (Enzymatics) and brought to a final volume of 20 μl with ultrapure water. The resulting product was subjected to a reannealing process to allow heteroduplex formation: 95°C for 10 min, ramped from 95°C to 85°C at -2°C / sec, ramped from 85°C to 25°C at -0.25°C / sec, and held at 25°C for 1 min. After reannealing, the product was treated with SURVEYOR Nuclease and SURVEYOR Enhancer S (Transgenomics) according to the manufacturer's recommended protocol and analyzed on a 4-20% Novex TBE polyacrylamide gel (Life Technologies). The gel was stained with SYBR Gold DNA stain (Life Technologies) for 30 min and imaged using a Gel Doc gel imaging system (Bio-Rad). Quantitation was based on relative band intensity.
[0169] [Table 18]
[0170] Computational identification of unique CRISPR target sites To identify unique target sites for the Streptococcus pyogenes (S. pyogenes) SF370Cas9 (SpCas9) enzyme in the genomes of humans, mice, rats, zebrafish, fruit flies, and Caenorhabditis elegans (C. elegans), we developed a software package to scan both strands of a DNA sequence and identify all possible SpCas9 target sites. For this example, we operationally defined each SpCas9 target site as a 20-bp sequence followed by an NGG protospacer adjacent motif (PAM) sequence, and we identified all sequences that met this 5'-N20-NGG-3' definition on all chromosomes. To prevent nonspecific genome editing, after identifying all potential sites, we filtered all target sites based on the number of times they appeared in the relevant reference genome. To take advantage of the sequence specificity of Cas9 activity conferred by a "seed" sequence, which can be approximately 11-12 bp 5' from the PAM sequence, for example, we selected the 5'-NNNNNNNNNN-NGG-3' sequence as unique within the relevant genome. All genome sequences were downloaded from the UCSC Genome Browser (human genome hg19, mouse genome mm9, rat genome rn5, zebrafish genome danRer7, D. melanogaster genome dm4, and C. elegans genome ce10). All search results are available for viewing using the UCSC Genome Browser. An exemplary visualization of some target sites in the human genome is provided in Figure 22.
[0171] We first targeted three sites within the EMX1 locus in human HEK293FT cells. The genome modification efficiency of each chiRNA was assessed using the SURVEYOR nuclease assay, which detects mutations resulting from DNA double-strand breaks (DSBs) and their subsequent repair by the non-homologous end joining (NHEJ) DNA damage repair pathway. Constructs designated as chiRNA(+n) indicate that up to +n nucleotides of wild-type tracrRNA are included in the chimeric RNA construct, with values of 48, 54, 67, and 85 used for n. Chimeric RNAs containing longer fragments of wild-type tracrRNA (chiRNA(+67) and chiRNA(+85)) mediated DNA cleavage at all three EMX1 target sites, with chiRNA(+85) in particular demonstrating significantly higher levels of DNA cleavage than the corresponding crRNA / tracrRNA hybrids expressing the guide and tracr sequences in separate transcripts (Figures 10b and 10a). Two sites in the PVALB locus that did not produce detectable cleavage in the hybrid system (guide and tracr sequences expressed as separate transcripts) were also targeted using chiRNAs. chiRNA(+67) and chiRNA(+85) were able to mediate significant cleavage in the two PVALB protospacers (Figures 10c and 10b).
[0172] For all five targets in the EMX1 and PVALB loci, a consistent increase in genome modification efficiency was observed with increasing tracr sequence length. Without being bound by any theory, the secondary structure formed by the 3' end of tracrRNA may play a role in improving the rate of CRISPR complex formation. Figure 21 provides an illustration of the predicted secondary structure for each chimeric RNA used in this example. The secondary structure was predicted using RNAfold (http: / / rna.tbi.univie.ac.at / cgi-bin / RNAfold.cgi), which uses a minimum free energy and partition function algorithm. The pseudocolor (reproduced in grayscale) for each base indicates the probability of pairing. It is thought that chimeric RNAs with longer tracr sequences could cleave targets that are not cleaved by the natural CRISPRcrRNA / tracrRNA hybrid, allowing chimeric RNAs to be loaded onto Cas9 more efficiently than their natural hybrid counterparts. To facilitate the application of Cas9 for site-specific genome editing in eukaryotic cells and organisms, all predicted unique target sites for Streptococcus pyogenes (S. pyogenes) Cas9 were computationally identified in the human, mouse, rat, zebrafish, C. elegans, and D. melanogaster genomes. Chimeric RNAs can be designed for Cas9 enzymes from other microorganisms to expand the target space of CRISPR RNA-programmable nucleases.
[0173] Figures 11 and 21 illustrate exemplary bicistronic expression vectors for expressing chimeric RNAs containing up to +85 nucleotides of the wild-type tracrRNA sequence and SpCas9 with a nuclear localization sequence. SpCas9 is expressed from the CBh promoter and terminated by the bGH polyA signal (bGHpA). The expanded sequence illustrated directly below the schematic corresponds to the region surrounding the guide sequence insertion site and includes, from 5' to 3', the 3' portion of the U6 promoter (first shaded area), a BbsI cleavage site (arrow), a partial direct repeat (tracr mate sequence GTTTTAGAGCTA, underlined), a loop sequence GAAA, and the +85 tracr sequence (underlined sequence after the loop sequence). An exemplary guide sequence insert is illustrated below the guide sequence insertion site, with the nucleotide of the guide sequence for the selected target represented by "N."
[0174] The sequences set forth in the above examples are as follows (polynucleotide sequences are from 5' to 3'):
[0175] U6-short tracrRNA (Streptococcus pyogenes SF370): [ka]
[0176] U6-long tracrRNA (Streptococcus pyogenes SF370): [ka]
[0177] U6-DR-BbsI backbone-DR (Streptococcus pyogenes SF370): [ka]
[0178] U6-chimeric RNA-BbsI backbone (Streptococcus pyogenes SF370) [ka]
[0179] NLS-SpCas9-EGFP: [ka]
[0180] SpCas9-EGFP-NLS: [ka]
[0181] NLS-SpCas9-EGFP-NLS: [ka]
[0182] NLS-SpCas9-NLS: [ka]
[0183] NLS-mCherry-SpRNase3: [ka]
[0184] SpRNase3-mCherry-NLS: [ka]
[0185] NLS-SpCas9n-NLS (D10A nickase mutation is in lowercase): [ka]
[0186] hEMX1-HR template-HindII-NheI: [ka] [ka]
[0187] NLS-StCsn1-NLS: [ka]
[0188] U6-St_tracrRNA(7~97): [ka]
[0189] U6-DR-spacer-DR (Streptococcus pyogenes (S. pyogenes) SF370) [ka]
[0190] Chimeric RNA containing +48 tracrRNA (S. pyogenes SF370) [ka]
[0191] Chimeric RNA containing +54 tracrRNA (S. pyogenes SF370) [ka]
[0192] Chimeric RNA containing +67 tracrRNA (S. pyogenes SF370) [ka]
[0193] Chimeric RNA containing +85 tracrRNA (S. pyogenes SF370) [ka]
[0194] CBh-NLS-SpCas9-NLS [ka] [ka] [ka] [ka]
[0195] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka]
[0196] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka]
[0197] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka]
[0198] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka]
[0199] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka]
[0200] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka]
[0201] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka]
[0202] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka]
[0203] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka]
[0204] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR3Cas9 (with PAM of NGGNG) [ka]
[0205] A codon-optimized version of Cas9 from the S. thermophilus LMD-9 CRISPR3 locus (with NLS at both the 5' and 3' ends) [ka] [ka] [ka] [ka]
[0206] Example 5: Optimization of guide RNA for Streptococcus pyogenes Cas9 (referred to as SpCas9) The applicants mutated the tracrRNA and direct repeat sequences or mutated the chimeric guide RNA to improve RNA in cells.
[0207] The optimization was based on the observation that there were stretches of thymines (T) in the tracrRNA and guide RNA, which could result in premature transcription termination by the pol3 promoter. Therefore, we generated the following optimized sequences. The optimized tracrRNA and the corresponding optimized direct repeat are represented as a pair.
[0208] Optimized tracrRNA1 (mutations underlined): [ka]
[0209] Optimized direct repeat 1 (mutations are underlined): [ka]
[0210] Optimized tracrRNA2 (mutations underlined): [ka]
[0211] Optimized direct repeat 2 (mutations are underlined): [ka]
[0212] Applicants also optimized the chimeric guide RNAs for optimal activity in eukaryotic cells.
[0213] Original guide RNA: [ka]
[0214] Optimized chimeric guide RNA sequence 1: [ka]
[0215] Optimized chimeric guide RNA sequence 2: [ka]
[0216] Optimized chimeric guide RNA sequence 3: [ka]
[0217] Applicants have shown that the optimized chimeric guide RNA performs better, as shown in Figure 3. This experiment was performed by co-transfecting 293FT cells with Cas9 and U6 guide RNA DNA cassettes to express one of the four RNA forms described above. The guide RNAs target the same target site in the human Emx1 locus: "GTCACCTCCAATGACTAGGG".
[0218] Example 6: Optimization of Streptococcus thermophilus LMD-9 CRISPR1 Cas9 (referred to as St1Cas9) Applicants designed the guide chimeric RNA shown in FIG.
[0219] St1Cas9 guide RNAs can undergo the same type of optimization as SpCas9 guide RNAs by degrading stretches of polythymine (T).
[0220] Example 7: Improvement of the Cas9 system for in vivo use The applicants have conducted metagenomics search for Cas9 with small molecular weight. Most Cas9 homologs are quite large. For example, SpCas9 is about 1368 aa long, which is too large to be easily packaged into a viral vector for delivery. Some of the sequences are mis-annotated, so the exact frequency of each length is not always accurate. Nevertheless, this provides an indication of the distribution of Cas9 protein, and suggests that there are shorter Cas9 homologs.
[0221] Through computational analysis, we found that in the bacterial strain Campylobacter, there are two Cas9 proteins with less than 1,000 amino acids. The sequence for one Cas9 from Campylobacter jejuni is presented below. At this length, CjCas9 can be easily packaged into AAV, lentivirus, adenovirus, and other viral vectors for robust delivery into primary cells and in vivo in animal models.
[0222] Campylobacter jejuni Cas9 (CjCas9) [ka]
[0223] The putative tracrRNA element for this CjCas9 is: [ka]
[0224] The direct repeat sequence is: [ka]
[0225] The cofolded structure of tracrRNA and direct repeats is provided in Figure 6.
[0226] An example of a chimeric guide RNA for CjCas9 is: [ka]
[0227] Applicants also optimized Cas9 guide RNAs using in vitro methods. Figure 18 shows data from in vitro St1Cas9 chimeric guide RNA optimization.
[0228] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein can be employed in practicing the invention. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of those claims and their equivalents be covered thereby.
[0229] Example 8: Sa sgRNA optimization We designed five sgRNA variants for SaCas9 for optimal truncated architecture with maximum cleavage efficiency. Additionally, natural direct repeat:tracr duplexes were tested in parallel with the sgRNA. Guides with the indicated lengths were co-transfected with SaCas9 and tested for activity in HEK293FT cells. A total of 100 ng of sgRNA U6-PCR amplicon (or 50 ng of direct repeat and 50 ng of tracrRNA) and 400 ng of SaCas9 plasmid were co-transfected into 200,000 Hepa1-6 mouse hepatocytes, and DNA was harvested 72 hours post-transfection for SURVEYOR analysis. The results are shown in Figure 23.
[0230] References: 1. Urnov, F.D., Rebar, E.J., Holmes, M.C., Zhang, H.S. & Gregory, P.D. Genome editing with engineered zinc finger nucleases. Nat. Rev. Genet. 11, 636 - 646 (2010). 2. Bogdanove, A.J. & Voytas, D.F. TAL effectors: customizable proteins for DNA targeting. Science 333, 1843 - 1846 (2011). 3. Stoddard, B.L. Homing endonuclease structure and function. Q. Rev. Biophys. 38, 49 - 95 (2005). 4. Bae, T. & Schneewind, O. Allelic replacement in Staphylococcus aureus with inducible counter - selection. Plasmid 55, 58 - 63 (2006). 5. Sung, C.K., Li, H., Claverys, J.P. & Morrison, D.A. An rpsL cassette, janus, for gene replacement through negative selection in Streptococcus pneumoniae. Appl. Environ. Microbiol. 67, 5190 - 5196 (2001). 6. Sharan, S.K., Thomason, L.C., Kuznetsov, S.G. & Court, D.L. Recombineering: a homologous recombination - based method of genetic engineering. Nat. Protoc. 4, 206 - 223 (2009). 7.Jinek,M.et al.A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.Science 337,816-821(2012). 8.Deveau,H.,Garneau,J.E.&Moineau,S.CRISPR-Cas system and its role in phage-bacteria interactions.Annu.Rev.Microbiol.64,475-493(2010). 9.Horvath,P.&Barrangou,R.CRISPR-Cas,the immune system of bacteria and archaea.Science 327,167-170(2010). 10.Terns,M.P.&Terns,R.M.CRISPR-based adaptive immune systems.Curr.Opin.Microbiol.14,321-327(2011). 11.van der Oost,J.,Jore,M.M.,Westra,E.R.,Lundgren,M.&Brouns,S.J.CRISPR-based adaptive and heritable immunity in prokaryotes.Trends.Biochem.Sci.34,401-407(2009). 12.Brouns,S.J.et al.Small CRISPR RNAs guide antiviral defense in prokaryotes.Science 321,960-964(2008). 13.Carte,J.,Wang,R.,Li,H.,Terns,R.M.&Terns,M.P.Cas6 is an endoribonuclease that generates guide RNAs for invader defense in prokaryotes.Genes Dev.22,3489-3496(2008). 14.Deltcheva,E.et al.CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.Nature 471,602-607(2011). 15.Hatoum-Aslan,A.,Maniv,I.&Marraffini,L.A.Mature clustered,regularly interspaced,short palindromic repeats RNA(crRNA)length is measured by a ruler mechanism anchored at the precursor processing site.Proc.Natl.Acad.Sci.U.S.A.108,21218-21222(2011). 16.Haurwitz,R.E.,Jinek,M.,Wiedenheft,B.,Zhou,K.&Doudna,J.A.Sequence- and structure-specific RNA processing by a CRISPR endonuclease.Science 329,1355-1358(2010). 17.Deveau,H.et al.Phage response to CRISPR-encoded resistance in Streptococcus thermophilus.J.Bacteriol.190,1390-1400(2008). 18.Gasiunas,G.,Barrangou,R.,Horvath,P.&Siksnys,V.Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria.Proc.Natl.Acad.Sci.U.S.A.(2012). 19.Makarova,K.S.,Aravind,L.,Wolf,Y.I.&Koonin,E.V.Unification of Cas protein families and a simple scenario for the origin and evolution of CRISPR-Cas systems.Biol.Direct.6,38(2011). 20.Barrangou,R.RNA-mediated programmable DNA cleavage.Nat.Biotechnol.30,836-838(2012). 21.Brouns,S.J.Molecular biology.A Swiss army knife of immunity.Science 337,808-809(2012). 22.Carroll,D.A CRISPR Approach to Gene Targeting.Mol.Ther.20,1658-1660(2012). 23.Bikard,D.,Hatoum-Aslan,A.,Mucida,D.&Marraffini,L.A.CRISPR interference can prevent natural transformation and virulence acquisition during in vivo bacterial infection.Cell Host Microbe 12,177-186(2012). 24.Sapranauskas,R.et al.The Streptococcus thermophilus CRISPR-Cas system provides immunity in Escherichia coli.Nucleic Acids Res.(2011). 25.Semenova,E.et al.Interference by clustered regularly interspaced short palindromic repeat(CRISPR)RNA is governed by a seed sequence.Proc.Natl.Acad.Sci.U.S.A.(2011). 26.Wiedenheft,B.et al.RNA-guided complex from a bacterial immune system enhances target recognition through seed sequence interactions.Proc.Natl.Acad.Sci.U.S.A.(2011). 27.Zahner,D.&Hakenbeck,R.The Streptococcus pneumoniae beta-galactosidase is a surface protein.J.Bacteriol.182,5919-5921(2000). 28.Marraffini,L.A.,Dedent,A.C.&Schneewind,O.Sortases and the art of anchoring proteins to the envelopes of gram-positive bacteria.Microbiol.Mol.Biol.Rev.70,192-221(2006). 29.Motamedi,M.R.,Szigety,S.K.&Rosenberg,S.M.Double-strand-break repair recombination in Escherichia coli:physical evidence for a DNA replication mechanism in vivo.Genes Dev.13,2889-2903(1999). 30.Hosaka,T.et al.The novel mutation K87E in ribosomal protein S12 enhances protein synthesis activity during the late growth phase in Escherichia coli.Mol.Genet.Genomics 271,317-324(2004). 31.Costantino,N.&Court,D.L.Enhanced levels of lambda Red-mediated recombinants in mismatch repair mutants.Proc.Natl.Acad.Sci.U.S.A.100,15748-15753(2003). 32.Edgar,R.&Qimron,U.The Escherichia coli CRISPR system protects from lambda lysogenization,lysogens,and prophage induction.J.Bacteriol.192,6291-6294(2010). 33.Marraffini,L.A.&Sontheimer,E.J.Self versus non-self discrimination during CRISPR RNA-directed immunity.Nature 463,568-571(2010). 34.Fischer,S.et al.An archaeal immune system can detect multiple Protospacer Adjacent Motifs(PAMs)to target invader DNA.J.Biol.Chem.287,33351-33363(2012). 35.Gudbergsdottir,S.et al.Dynamic properties of the Sulfolobus CRISPR-Cas and CRISPR / Cmr systems when challenged with vector-borne viral and plasmid genes and protospacers.Mol.Microbiol.79,35-49(2011). 36.Wang,H.H.et al.Genome-scale promoter engineering by coselection MAGE.Nat Methods 9,591-593(2012). 37.Cong,L.et al.Multiplex Genome Engineering Using CRISPR-Cas Systems.Science In press(2013). 38.Mali,P.et al.RNA-Guided Human Genome Engineering via Cas9.Science In press(2013). 39.Hoskins,J.et al.Genome of the bacterium Streptococcus pneumoniae strain R6.J.Bacteriol.183,5709-5717(2001). 40.Havarstein,L.S.,Coomaraswamy,G.&Morrison,D.A.An unmodified heptadecapeptide pheromone induces competence for genetic transformation in Streptococcus pneumoniae.Proc.Natl.Acad.Sci.U.S.A.92,11140-11144(1995). 41.Horinouchi,S.&Weisblum,B.Nucleotide sequence and functional map of pC194,a plasmid that specifies inducible chloramphenicol resistance.J.Bacteriol.150,815-825(1982). 42.Horton,R.M.In Vitro Recombination and Mutagenesis of DNA:SOEing Together Tailor-Made Genes.Methods Mol.Biol.15,251-261(1993). 43.Podbielski,A.,Spellerberg,B.,Woischnik,M.,Pohl,B.&Lutticken,R.Novel series of plasmid vectors for gene inactivation and expression analysis in group A streptococci(GAS).Gene 177,137-147(1996). 44.Husmann,L.K.,Scott,J.R.,Lindahl,G.&Stenberg,L.Expression of the Arp protein,a member of the M protein family,is not sufficient to inhibit phagocytosis of Streptococcus pyogenes.Infection and immunity 63,345-348(1995). 45.Gibson,D.G.et al.Enzymatic assembly of DNA molecules up to several hundred kilobases.Nat Methods 6,343-345(2009). 46. Tangri S, et al. (“Rationally engineered proteins therapeutic with reduced immunogenicity” J Immunol. 2005 Mar 15;174(6):3187-96.
[0231] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein can be employed in practicing the invention. [Sequence table] SEQUENCE LISTING <110> THE BROAD INSTITUTE, INC. MASSACHUSETTS INSTITUTE OF TECHNOLOGY PRESIDENT AND FELLOWS OF HARVARD COLLEGE <120> ENGINEERING OF SYSTEMS, METHODS AND OPTIMIZED GUIDE COMPOSITIONS FOR SEQUENCE MANIPULATION <130> F50936A1 <140> PCT / US2013 / 074819 <141> 2013-12-12 <150> US 61 / 836,127 <151> 2013-06-17 <150> US 61 / 835,931 <151> 2013-06-17 <150> US 61 / 828,130 <151> 2013-05-28 <150> US 61 / 819,803 <151> 2013-05-06 <150> US 61 / 814,263 <151> 2013-04-20 <150> US 61 / 806,375 <151> 2013-03-28 <150> US 61 / 802,174 <151> 2013-03-15 <150> US 61 / 791,409 <151> 2013-03-15 <150> US 61 / 769,046 <151> 2013-02-25 <150> US 61 / 758,468 <151> 2013-01-30 <150> US 61 / 748,427 <151> 2013-01-02 <150> US 61 / 736,527 <151> 2012-12-12 <160> 264 <170> PatentIn version 3.5 <210> 1 <211> 15 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 1 aggacgaagt cctaa 15 <210> 2 <211> 7 <212> PRT <213> Simian virus 40 <400> 2 Pro Lys Lys Lys Arg Lys Val 1 5 <210> 3 <211> 16 <212> PRT <213> Unknown <220> <221> source <223> / note="Description of Unknown: Nucleoplasmin bipartite NLS sequence" <400> 3 Lys Arg Pro Ala Ala Thr Lys Lys Ala Gly Gln Ala Lys Lys Lys Lys 1 5 10 15 <210> 4 <211> 9 <212> PRT <213> Unknown <220> <221> source <223> / note="Description of Unknown: C-myc NLS sequence" <400> 4 Pro Ala Ala Lys Arg Val Lys Leu Asp 1 5 <210> 5 <211> 11 <212> PRT <213> Unknown <220> <221> source <223> / note="Description of Unknown: C-myc NLS sequence" <400> 5 Arg Gln Arg Arg Asn Glu Leu Lys Arg Ser Pro 1 5 10 <210> 6 <211> 38 <212> PRT <213> Homo sapiens <400> 6 Asn Gln Ser Ser Asn Phe Gly Pro Met Lys Gly Gly Asn Phe Gly Gly 1 5 10 15 Arg Ser Ser Gly Pro Tyr Gly Gly Gly Gly Gln Tyr Phe Ala Lys Pro 20 25 30 Arg Asn Gln Gly Gly Tyr 35 <210> 7 <211> 42 <212> PRT <213> Unknown <220> <221> source <223> / note="Description of Unknown: IBB domain from importin-alpha sequence" <400> 7 Arg Met Arg Ile Glx Phe Lys Asn Lys Gly Lys Asp Thr Ala Glu Leu 1 5 10 15 Arg Arg Arg Arg Val Glu Val Ser Val Glu Leu Arg Lys Ala Lys Lys 20 25 30 Asp Glu Gln Ile Leu Lys Arg Arg Asn Val 35 40 <210> 8 <211> 8 <212> PRT <213> Unknown <220> <221> source <223> / note="Description of Unknown: Myoma T protein sequence" <400> 8 Val Ser Arg Lys Arg Pro Arg Pro 1 5 <210> 9 <211> 8 <212> PRT <213> Unknown <220> <221> source <223> / note="Description of Unknown: Myoma T protein sequence” <400> 9 Pro Pro Lys Lys Ala Arg Glu Asp 1 5 <210> 10 <211> 8 <212> PRT <213> Homo sapiens <400> 10 Pro Gln Pro Lys Lys Lys Pro Leu 1 5 <210> 11 <211> 12 <212> PRT <213> Mus musculus <400> 11 Sister Ala Leu Ile Lys Lys Lys Lys Lys Met Ala Pro 1 5 10 <210> 12 <211> 5 <212> PRT <213> Influenza virus <400> 12 Asp Arg Leu Arg Arg 1 5 <210> 13 <211> 7 <212> PRT <213> Influenza virus <400> 13 Pro Light Gln Light Light Arg Light 1 5 <210> 14 <211> 10 <212> PRT <213> Hepatitis delta virus <400> 14 Arg Lys Leu Lys Lys Lys Ile Lys Lys Leu 1 5 10 <210> 15 <211> 10 <212> PRT <213> Mus musculus <400> 15 Arg Glu Lys Lys Lys Phe Leu Lys Arg Arg 1 5 10 <210> 16 <211> 20 <212> PRT <213> Homo sapiens <400> 16 Light Arg Light Gly Asp Glu Val Asp Gly Val Asp Glu Val Ala Light Light 1 5 10 15 Light Sees Light Light 20 <210> 17 <211> 17 <212> PRT <213> Homo sapiens <400> 17 Arg Lys Cys Leu Gln Ala Gly Met Asn Leu Glu Ala Arg Lys Thr Lys 1 5 10 15 Lys <210> 18 <211> 27 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t or g <220> <221> modified_base <222> (21)..(22) <223> a, c, t, g, unknown or other <400> 18 nnnnnnnnnn nnnnnnnnnn nnagaaw 27 <210> 19 <211> 19 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <220> <221> modified_base <222> (1)..(12) <223> a, c, t or g <220> <221> modified_base <222> (13)..(14) <223> a, c, t, g, unknown or other <400> 19 nnnnnnnnnn nnnnagaaw 19 <210> 20 <211> 27 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t or g <220> <221> modified_base <222> (21)..(22) <223> a, c, t, g, unknown or other <400> 20 nnnnnnnnnn nnnnnnnnnn nnagaaw 27 <210> 21 <211> 18 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <220> <221> modified_base <222> (1)..(11) <223> a, c, t or g <220> <221> modified_base <222> (12)..(13) <223> a, c, t, g, unknown or other <400> 21 nnnnnnnnnn nnnagaaw 18 <210> 22 <211> 137 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 22 nnnnnnnnnn nnnnnnnnnn gtttttgtac tctcaagatt tagaaataaa tcttgcagaa 60 gctacaaaga taaggcttca tgccgaaatc aacaccctgt cattttatgg cagggtgttt 120 tcgttattta atttttt 137 <210> 23 <211> 123 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 23 nnnnnnnnnn nnnnnnnnnn gtttttgtac tctcagaaat gcagaagcta caaagataag 60 gcttcatgcc gaaatcaaca ccctgtcatt ttatggcagg gtgttttcgt tatttaattt 120 ttt 123 <210> 24 <211> 110 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 24 nnnnnnnnnn nnnnnnnnnn gtttttgtac tctcagaaat gcagaagcta caaagataag 60 gcttcatgcc gaaatcaaca ccctgtcatt ttatggcagg gtgttttttt 110 <210> 25 <211> 102 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 25 nnnnnnnnnn nnnnnnnnnn gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgctttt tt 102 <210> 26 <211> 88 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 26 nnnnnnnnnn nnnnnnnnnn gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt gttttttt 88 <210> 27 <211> 76 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 27 nnnnnnnnnn nnnnnnnnnn gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcatt tttttt 76 <210> 28 <211> 12 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide” <400> 28 guuuuagagc ua 12 <210> 29 <211> 33 <212> DNA <213> Homo sapiens <400> 29 ggacatcgat gtcacctcca atgactaggg tgg 33 <210> 30 <211> 33 <212> DNA <213> Homo sapiens <400> 30 cattggaggt gacatcgatg tcctccccat tgg 33 <210> 31 <211> 33 <212> DNA <213> Homo sapiens <400> 31 ggaagggcct gagtccgagc agaagaagaa ggg 33 <210> 32 <211> 33 <212> DNA <213> Homo sapiens <400> 32 ggtggcgaga ggggccgaga ttgggtgttc agg 33 <210> 33 <211> 33 <212> DNA <213> Homo sapiens <400> 33 atgcaggagg gtggcgagag gggccgagat tgg 33 <210> 34 <211> 21 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic primary” <400> 34 aaaaccacccc ttctctctgg c 21 <210> 35 <211> 21 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic primary" <400> 35 ggagattgga gacacggaga g 21 <210> 36 <211> 20 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic primer" <400> 36 ctggaaagcc aatgcctgac 20 <210> 37 <211> 20 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic primer" <400> 37 ggcagcaaac tccttgtcct 20 <210> 38 <211> 12 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 38 gttttagagc ta 12 <210> 39 <211> 335 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <400> 39 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 aatattggaa ttaatttgac tgtaacaca agatattag tacaaatac gtgacgtaga 120 aagtaataat ttctgggta gtttgcagtt ttaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatctt gtggaagga 240 cgaaacaccg gaaccattca aaacagcata gcaagttaa aaggctag tccgttatca 300 acttgaaaaa gtggcaccga gtcggtgctt tttt 335 <210> 40 <211> 423 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <400> 40 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 aatattggaa ttaatttgac tgtaacaca agatattag tacaaatac gtgacgtaga 120 aagtaataat ttctgggta gtttgcagtt ttaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatctt gtggaagga 240 cgaaacaccg gtagtattaa gtattgtttt atggctgata aatttctttg aatttctcct 300 tgattatttg tttaaaagt tataaaataa tcttgttg accattcaa acagcatagc 360 aagttaaaat aaggctagtc cgttatcaac ttgaaaagt ggcaccgagt cggtgctttt 420 pp. 423 <210> 41 <211> 339 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <400> 41 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 aatattggaa ttaatttgac tgtaacaca agatattag tacaaatac gtgacgtaga 120 aagtaataat ttctgggta gtttgcagtt ttaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatctt gtggaagga 240 cgaaacaccg ggttttagag ctatgctgtt tgaatggtc ccaaaacggg tcttcgagaa 300 gacgttttag agctatgctg tttgaatgg tccaaac 339 <210> 42 <211> 309 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <400> 42 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 ataattggaa ttaatttgac tgtaaacaca aagatattag tacaaaatac gtgacgtaga 120 aagtaataat ttcttgggta gtttgcagtt ttaaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatatctt gtggaaagga 240 cgaaacaccg ggtcttcgag aagacctgtt ttagagctag aatagcaag ttaaaataag 300 gctagtccg 309 <210> 43 <211> 1648 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polypeptide” <400> 43 Met Asp Tyr Lys Asp His Asp Gly Asp Tyr Lys Asp His Asp Ile Asp 1 5 10 15 Tyr Lys Asp Asp Asp Asp Lys Put Here Pro Lys Lys Lys Arg Lys Val 20 25 30 Gly Ile His Gly Val Pro Ala Ala Asp Lys Lys Tyr Ser Ile Gly Leu 35 40 45 Asp Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu Tyr 50 55 60 Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg His 65 70 75 80 Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly Glu 85 90 95 Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr Thr 100 105 110 Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn Glu 115 120 125 Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser Phe 130 135 140 Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly Asn 145 150 155 160 Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr His 165 170 175 Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg Leu 180 185 190 Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe Leu 195 200 205 Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu Phe 210 215 220 Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro Ile 225 230 235 240 Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu Ser 245 250 255 Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu Lys 260 265 270 Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu Thr 275 280 285 Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu Gln 290 295 300 Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala Gln 305 310 315 320 Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu Ser 325 330 335 Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile Thr 340 345 350 Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His His 355 360 365 Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro Glu 370 375 380 Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala Gly 385 390 395 400 Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile Lys 405 410 415 Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys Leu 420 425 430 Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly Ser 435 440 445 Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg Arg 450 455 460 Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile Glu 465 470 475 480 Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala Arg 485 490 495 Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr Ile 500 505 510 Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala Gln 515 520 525 Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn Glu 530 535 540 Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val Tyr 545 550 555 560 Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys Pro 565 570 575 Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu Phe 580 585 590 Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr Phe 595 600 605 Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu Asp 610 615 620 Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile Ile 625 630 635 640 Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu Glu 645 650 655 Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile Glu 660 665 670 Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met Lys 675 680 685 Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg Lys 690 695 700 Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu Asp 705 710 715 720 Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu Ile 725 730 735 His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln Val 740 745 750 Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala Gly 755 760 765 Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val Asp 770 775 780 Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile Val Ile 785 790 795 800 Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn Ser 805 810 815 Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly Ser 820 825 830 Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln Asn Glu 835 840 845 Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr Val Asp 850 855 860 Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp His Ile 865 870 875 880 Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys Val Leu 885 890 895 Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro Ser Glu 900 905 910 Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu Asn Ala 915 920 925 Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr Lys Ala Glu Arg 930 935 940 Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe Ile Lys Arg Gln Leu 945 950 955 960 Val Glu Thr Arg Gln Ile Thr Lys His Val Ala Gln Ile Leu Asp Ser 965 970 975 Arg Met Asn Thr Lys Tyr Asp Glu Asn Asp Lys Leu Ile Arg Glu Val 980 985 990 Lys Val Ile Thr Leu Lys Ser Lys Leu Val Ser Asp Phe Arg Lys Asp 995 1000 1005 Phe Gln Phe Tyr Lys Val Arg Glu Ile Asn Asn Tyr His His Ala 1010 1015 1020 His Asp Ala Tyr Leu Asn Ala Val Val Gly Thr Ala Leu Ile Lys 1025 1030 1035 Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr Lys 1040 1045 1050 Val Tyr Asp Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu Ile 1055 1060 1065 Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile Met Asn 1070 1075 1080 Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile Arg Lys 1085 1090 1095 Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile Val Trp 1100 1105 1110 Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser Met 1115 1120 1125 Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly 1130 1135 1140 Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys Leu 1145 1150 1155 Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe 1160 1165 1170 Asp Ser Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val 1175 1180 1185 Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu 1190 1195 1200 Gly Ile Thr Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile 1205 1210 1215 Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu 1220 1225 1230 Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn Gly 1235 1240 1245 Arg Lys Arg Met Leu Ala Ser Ala Gly Glu Leu Gln Lys Gly Asn 1250 1255 1260 Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr Leu Ala 1265 1270 1275 Ser His Tyr Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn Glu Gln 1280 1285 1290 Lys Gln Leu Phe Val Glu Gln His Lys His Tyr Leu Asp Glu Ile 1295 1300 1305 Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg Val Ile Leu Ala Asp 1310 1315 1320 Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr Asn Lys His Arg Asp 1325 1330 1335 Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile Ile His Leu Phe Thr 1340 1345 1350 Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe Asp Thr 1355 1360 1365 Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu Asp 1370 1375 1380 Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr Arg 1385 1390 1395 Ile Asp Leu Ser Gln Leu Gly Gly Asp Ala Ala Ala Val Ser Lys 1400 1405 1410 Gly Glu Glu Leu Phe Thr Gly Val Val Pro Ile Leu Val Glu Leu 1415 1420 1425 Asp Gly Asp Val Asn Gly His Lys Phe Ser Val Ser Gly Glu Gly 1430 1435 1440 Glu Gly Asp Ala Thr Tyr Gly Lys Leu Thr Leu Lys Phe Ile Cys 1445 1450 1455 Thr Thr Gly Lys Leu Pro Val Pro Trp Pro Thr Leu Val Thr Thr 1460 1465 1470 Leu Thr Tyr Gly Val Gln Cys Phe Ser Arg Tyr Pro Asp His Met 1475 1480 1485 Lys Gln His Asp Phe Phe Lys Ser Ala Met Pro Glu Gly Tyr Val 1490 1495 1500 Gln Glu Arg Thr Ile Phe Phe Lys Asp Asp Gly Asn Tyr Lys Thr 1505 1510 1515 Arg Ala Glu Val Lys Phe Glu Gly Asp Thr Leu Val Asn Arg Ile 1520 1525 1530 Glu Leu Lys Gly Ile Asp Phe Lys Glu Asp Gly Asn Ile Leu Gly 1535 1540 1545 His Lys Leu Glu Tyr Asn Tyr Asn Ser His Asn Val Tyr Ile Met 1550 1555 1560 Ala Asp Lys Gln Lys Asn Gly Ile Lys Val Asn Phe Lys Ile Arg 1565 1570 1575 His Asn Ile Glu Asp Gly Ser Val Gln Leu Ala Asp His Tyr Gln 1580 1585 1590 Gln Asn Thr Pro Ile Gly Asp Gly Pro Val Leu Leu Pro Asp Asn 1595 1600 1605 His Tyr Leu Ser Thr Gln Ser Ala Leu Ser Lys Asp Pro Asn Glu 1610 1615 1620 Lys Arg Asp His Met Val Leu Leu Glu Phe Val Thr Ala Ala Gly 1625 1630 1635 Ile Thr Leu Gly Met Asp Glu Leu Tyr Lys 1640 1645 <210> 44 <211> 1625 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polypeptide" <400> 44 Met Asp Lys Lys Tyr Ser Ile Gly Leu Asp Ile Gly Thr Asn Ser Val 1 5 10 15 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile 35 40 45 Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys 65 70 75 80 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 85 90 95 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp 130 135 140 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 145 150 155 160 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 210 215 220 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 260 265 270 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 275 280 285 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 290 295 300 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 325 330 335 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 340 345 350 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 485 490 495 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys 515 520 525 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 565 570 575 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580 585 590 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830 Leu Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys 885 890 895 Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 900 905 910 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915 920 925 Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp 930 935 940 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Leu Lys Ser 945 950 955 960 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965,970,975 Glu Ile Asn Asn Tyr His Ala His Asp Ala Tyr Leu Asn Ala Val 980,985,990 Val Gly Thr Ala Leu Ile Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1010 1015 1020 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1040 1045 1050 Asn Gly Glu With Arg Lys Arg Pro Leu Glu Thr Asn Gly Glu 1055 1060 1065 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr 1085 1090 1095 Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1100 1105 1110 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 Light Light Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1190 1195 1200 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1220 1225 1230 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1265 1270 1275 Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1280 1285 1290 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn 1295 1300 1305 Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1310 1315 1320 Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser 1325 1330 1335 Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr 1340 1345 1350 Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 Ala Ala Ala Val Ser Lys Gly Glu Glu Leu Phe Thr Gly Val Val 1370 1375 1380 Pro Ile Leu Val Glu Leu Asp Gly Asp Val Asn Gly His Lys Phe 1385 1390 1395 Ser Val Ser Gly Glu Gly Glu Gly Asp Ala Thr Tyr Gly Lys Leu 1400 1405 1410 Thr Leu Lys Phe Ile Cys Thr Thr Gly Lys Leu Pro Val Pro Trp 1415 1420 1425 Pro Thr Leu Val Thr Thr Leu Thr Tyr Gly Val Gln Cys Phe Ser 1430 1435 1440 Arg Tyr Pro Asp His Met Lys Gln His Asp Phe Phe Lys Ser Ala 1445 1450 1455 Met Pro Glu Gly Tyr Val Gln Glu Arg Thr Ile Phe Phe Lys Asp 1460 1465 1470 Asp Gly Asn Tyr Lys Thr Arg Ala Glu Val Lys Phe Glu Gly Asp 1475 1480 1485 Thr Leu Val Asn Arg Ile Glu Leu Lys Gly Ile Asp Phe Lys Glu 1490 1495 1500 Asp Gly Asn Ile Leu Gly His Lys Leu Glu Tyr Asn Tyr Asn Ser 1505 1510 1515 His Asn Val Tyr Ile Met Ala Asp Lys Gln Lys Asn Gly Ile Lys 1520 1525 1530 Val Asn Phe Lys Ile Arg His Asn Ile Glu Asp Gly Ser Val Gln 1535 1540 1545 Leu Ala Asp His Tyr Gln Gln Asn Thr Pro Ile Gly Asp Gly Pro 1550 1555 1560 Val Leu Leu Pro Asp Asn His Tyr Leu Ser Thr Gln Ser Ala Leu 1565 1570 1575 Ser Lys Asp Pro Asn Glu Lys Arg Asp His Met Val Leu Leu Glu 1580 1585 1590 Phe Val Thr Ala Ala Gly Ile Thr Leu Gly Met Asp Glu Leu Tyr 1595 1600 1605 Lys Lys Arg Pro Ala Ala Thr Lys Lys Ala Gly Gln Ala Lys Lys 1610 1615 1620 Lys Lys 1625 <210> 45 <211> 1664 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polypeptide" <400> 45 Met Asp Tyr Lys Asp His Asp Gly Asp Tyr Lys Asp His Asp Ile Asp 1 5 10 15 Tyr Lys Asp Asp Asp Asp Lys Met Ala Pro Lys Lys Lys Arg Lys Val 20 25 30 Gly Ile His Gly Val Pro Ala Ala Asp Lys Lys Tyr Ser Ile Gly Leu 35 40 45 Asp Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu Tyr 50 55 60 Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg His 65 70 75 80 Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly Glu 85 90 95 Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr Thr 100 105 110 Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn Glu 115 120 125 Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser Phe 130 135 140 Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly Asn 145 150 155 160 Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr His 165 170 175 Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg Leu 180 185 190 Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe Leu 195 200 205 Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu Phe 210 215 220 Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro Ile 225 230 235 240 Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu Ser 245 250 255 Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu Lys 260 265 270 Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu Thr 275 280 285 Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu Gln 290 295 300 Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala Gln 305 310 315 320 Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu Ser 325 330 335 Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile Thr 340 345 350 Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His His 355 360 365 Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro Glu 370 375 380 Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala Gly 385 390 395 400 Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile Lys 405 410 415 Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys Leu 420 425 430 Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly Ser 435 440 445 Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg Arg 450 455 460 Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile Glu 465 470 475 480 Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala Arg 485 490 495 Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr Ile 500 505 510 Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala Gln 515 520 525 Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn Glu 530 535 540 Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val Tyr 545 550 555 560 Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys Pro 565 570 575 Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu Phe 580 585 590 Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr Phe 595 600 605 Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu Asp 610 615 620 Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile Ile 625 630 635 640 Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu Glu 645 650 655 Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile Glu 660 665 670 Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met Lys 675 680 685 Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg Lys 690 695 700 Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu Asp 705 710 715 720 Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu Ile 725 730 735 His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln Val 740 745 750 Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala Gly 755 760 765 Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val Asp 770 775 780 Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile Val Ile 785 790 795 800 Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn Ser 805 810 815 Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly Ser 820 825 830 Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln Asn Glu 835 840 845 Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr Val Asp 850 855 860 Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp His Ile 865 870 875 880 Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys Val Leu 885 890 895 Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro Ser Glu 900 905 910 Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu Asn Ala 915 920 925 Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr Lys Ala Glu Arg 930 935 940 Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe Ile Lys Arg Gln Leu 945 950 955 960 Val Glu Thr Arg Gln Ile Thr Lys His Val Ala Gln Ile Leu Asp Ser 965 970 975 Arg Met Asn Thr Lys Tyr Asp Glu Asn Asp Lys Leu Ile Arg Glu Val 980 985 990 Lys Val Ile Thr Leu Lys Ser Lys Leu Val Ser Asp Phe Arg Lys Asp 995 1000 1005 Phe Gln Phe Tyr Lys Val Arg Glu Ile Asn Asn Tyr His His Ala 1010 1015 1020 His Asp Ala Tyr Leu Asn Ala Val Val Gly Thr Ala Leu Ile Lys 1025 1030 1035 Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr Lys 1040 1045 1050 Val Tyr Asp Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu Ile 1055 1060 1065 Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile Met Asn 1070 1075 1080 Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile Arg Lys 1085 1090 1095 Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile Val Trp 1100 1105 1110 Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser Met 1115 1120 1125 Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly 1130 1135 1140 Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys Leu 1145 1150 1155 Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe 1160 1165 1170 Asp Ser Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val 1175 1180 1185 Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu 1190 1195 1200 Gly Ile Thr Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile 1205 1210 1215 Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu 1220 1225 1230 Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn Gly 1235 1240 1245 Arg Lys Arg Met Leu Ala Ser Ala Gly Glu Leu Gln Lys Gly Asn 1250 1255 1260 Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr Leu Ala 1265 1270 1275 Ser His Tyr Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn Glu Gln 1280 1285 1290 Lys Gln Leu Phe Val Glu Gln His Lys His Tyr Leu Asp Glu Ile 1295 1300 1305 Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg Val Ile Leu Ala Asp 1310 1315 1320 Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr Asn Lys His Arg Asp 1325 1330 1335 Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile Ile His Leu Phe Thr 1340 1345 1350 Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe Asp Thr 1355 1360 1365 Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu Asp 1370 1375 1380 Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr Arg 1385 1390 1395 Ile Asp Leu Ser Gln Leu Gly Gly Asp Ala Ala Ala Val Ser Lys 1400 1405 1410 Gly Glu Glu Leu Phe Thr Gly Val Val Pro Ile Leu Val Glu Leu 1415 1420 1425 Asp Gly Asp Val Asn Gly His Lys Phe Ser Val Ser Gly Glu Gly 1430 1435 1440 Glu Gly Asp Ala Thr Tyr Gly Lys Leu Thr Leu Lys Phe Ile Cys 1445 1450 1455 Thr Thr Gly Lys Leu Pro Val Pro Trp Pro Thr Leu Val Thr Thr 1460 1465 1470 Leu Thr Tyr Gly Val Gln Cys Phe Ser Arg Tyr Pro Asp His Met 1475 1480 1485 Lys Gln His Asp Phe Phe Lys Ser Ala Met Pro Glu Gly Tyr Val 1490 1495 1500 Gln Glu Arg Thr Ile Phe Phe Lys Asp Asp Gly Asn Tyr Lys Thr 1505 1510 1515 Arg Ala Glu Val Lys Phe Glu Gly Asp Thr Leu Val Asn Arg Ile 1520 1525 1530 Glu Leu Lys Gly Ile Asp Phe Lys Glu Asp Gly Asn Ile Leu Gly 1535 1540 1545 His Lys Leu Glu Tyr Asn Tyr Asn Ser His Asn Val Tyr Ile Met 1550 1555 1560 Ala Asp Lys Gln Lys Asn Gly Ile Lys Val Asn Phe Lys Ile Arg 1565 1570 1575 His Asn Ile Glu Asp Gly Ser Val Gln Leu Ala Asp His Tyr Gln 1580 1585 1590 Gln Asn Thr Pro Ile Gly Asp Gly Pro Val Leu Leu Pro Asp Asn 1595 1600 1605 His Tyr Leu Ser Thr Gln Ser Ala Leu Ser Lys Asp Pro Asn Glu 1610 1615 1620 Lys Arg Asp His Met Val Leu Leu Glu Phe Val Thr Ala Ala Gly 1625 1630 1635 Ile Thr Leu Gly Met Asp Glu Leu Tyr Lys Lys Arg Pro Ala Ala 1640 1645 1650 Thr Lys Lys Ala Gly Gln Ala Lys Lys Lys Lys 1655 1660 <210> 46 <211> 1423 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polypeptide" <400> 46 Met Asp Tyr Lys Asp His Asp Gly Asp Tyr Lys Asp His Asp Ile Asp 1 5 10 15 Tyr Lys Asp Asp Asp Asp Lys Met Ala Pro Lys Lys Lys Arg Lys Val 20 25 30 Gly Ile His Gly Val Pro Ala Ala Asp Lys Lys Tyr Ser Ile Gly Leu 35 40 45 Asp Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu Tyr 50 55 60 Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg His 65 70 75 80 Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly Glu 85 90 95 Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr Thr 100 105 110 Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn Glu 115 120 125 Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser Phe 130 135 140 Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly Asn 145 150 155 160 Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr His 165 170 175 Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg Leu 180 185 190 Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe Leu 195 200 205 Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu Phe 210 215 220 Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro Ile 225 230 235 240 Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu Ser 245 250 255 Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu Lys 260 265 270 Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu Thr 275 280 285 Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu Gln 290 295 300 Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala Gln 305 310 315 320 Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu Ser 325 330 335 Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile Thr 340 345 350 Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His His 355 360 365 Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro Glu 370 375 380 Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala Gly 385 390 395 400 Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile Lys 405 410 415 Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys Leu 420 425 430 Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly Ser 435 440 445 Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg Arg 450 455 460 Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile Glu 465 470 475 480 Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala Arg 485 490 495 Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr Ile 500 505 510 Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala Gln 515 520 525 Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn Glu 530 535 540 Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val Tyr 545 550 555 560 Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys Pro 565 570 575 Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu Phe 580 585 590 Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr Phe 595 600 605 Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu Asp 610 615 620 Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile Ile 625 630 635 640 Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu Glu 645 650 655 Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile Glu 660 665 670 Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met Lys 675 680 685 Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg Lys 690 695 700 Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu Asp 705 710 715 720 Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu Ile 725 730 735 His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln Val 740 745 750 Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala Gly 755 760 765 Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val Asp 770 775 780 Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile Val Ile 785 790 795 800 Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn Ser 805 810 815 Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly Ser 820 825 830 Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln Asn Glu 835 840 845 Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr Val Asp 850 855 860 Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp His Ile 865 870 875 880 Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys Val Leu 885 890 895 Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro Ser Glu 900 905 910 Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu Asn Ala 915 920 925 Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr Lys Ala Glu Arg 930 935 940 Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe Ile Lys Arg Gln Leu 945 950 955 960 Val Glu Thr Arg Gln Ile Thr Lys His Val Ala Gln Ile Leu Asp Ser 965 970 975 Arg Met Asn Thr Lys Tyr Asp Glu Asn Asp Lys Leu Ile Arg Glu Val 980 985 990 Lys Val Ile Thr Leu Lys Ser Lys Leu Val Ser Asp Phe Arg Lys Asp 995 1000 1005 Phe Gln Phe Tyr Lys Val Arg Glu Ile Asn Asn Tyr His His Ala 1010 1015 1020 His Asp Ala Tyr Leu Asn Ala Val Val Gly Thr Ala Leu Ile Lys 1025 1030 1035 Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr Lys 1040 1045 1050 Val Tyr Asp Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu Ile 1055 1060 1065 Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile Met Asn 1070 1075 1080 Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile Arg Lys 1085 1090 1095 Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile Val Trp 1100 1105 1110 Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser Met 1115 1120 1125 Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly 1130 1135 1140 Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys Leu 1145 1150 1155 Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe 1160 1165 1170 Asp Ser Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val 1175 1180 1185 Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu 1190 1195 1200 Gly Ile Thr Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile 1205 1210 1215 Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu 1220 1225 1230 Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn Gly 1235 1240 1245 Arg Lys Arg Met Leu Ala Ser Ala Gly Glu Leu Gln Lys Gly Asn 1250 1255 1260 Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr Leu Ala 1265 1270 1275 Ser His Tyr Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn Glu Gln 1280 1285 1290 Lys Gln Leu Phe Val Glu Gln His Lys His Tyr Leu Asp Glu Ile 1295 1300 1305 Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg Val Ile Leu Ala Asp 1310 1315 1320 Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr Asn Lys His Arg Asp 1325 1330 1335 Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile Ile His Leu Phe Thr 1340 1345 1350 Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe Asp Thr 1355 1360 1365 Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu Asp 1370 1375 1380 Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr Arg 1385 1390 1395 Ile Asp Leu Ser Gln Leu Gly Gly Asp Lys Arg Pro Ala Ala Thr 1400 1405 1410 Lys Lys Ala Gly Gln Ala Lys Lys Lys Lys 1415 1420 <210> 47 <211> 483 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polypeptide" <400> 47 Met Phe Leu Phe Leu Ser Leu Thr Ser Phe Leu Ser Ser Ser Arg Thr 1 5 10 15 Leu Val Ser Lys Gly Glu Glu Asp Asn Met Ala Ile Ile Lys Glu Phe 20 25 30 Met Arg Phe Lys Val His Met Glu Gly Ser Val Asn Gly His Glu Phe 35 40 45 Glu Ile Glu Gly Glu Gly Glu Gly Arg Pro Tyr Glu Gly Thr Gln Thr 50 55 60 Ala Lys Leu Lys Val Thr Lys Gly Gly Pro Leu Pro Phe Ala Trp Asp 65 70 75 80 Ile Leu Ser Pro Gln Phe Met Tyr Gly Ser Lys Ala Tyr Val Lys His 85 90 95 Pro Ala Asp Ile Pro Asp Tyr Leu Lys Leu Ser Phe Pro Glu Gly Phe 100 105 110 Lys Trp Glu Arg Val Met Asn Phe Glu Asp Gly Gly Val Val Thr Val 115 120 125 Thr Gln Asp Ser Ser Leu Gln Asp Gly Glu Phe Ile Tyr Lys Val Lys 130 135 140 Leu Arg Gly Thr Asn Phe Pro Ser Asp Gly Pro Val Met Gln Lys Lys 145 150 155 160 Thr Met Gly Trp Glu Ala Ser Ser Glu Arg Met Tyr Pro Glu Asp Gly 165 170 175 Ala Leu Lys Gly Glu Ile Lys Gln Arg Leu Lys Leu Lys Asp Gly Gly 180 185 190 His Tyr Asp Ala Glu Val Lys Thr Thr Tyr Lys Ala Lys Lys Pro Val 195 200 205 Gln Leu Pro Gly Ala Tyr Asn Val Asn Ile Lys Leu Asp Ile Thr Ser 210 215 220 His Asn Glu Asp Tyr Thr Ile Val Glu Gln Tyr Glu Arg Ala Glu Gly 225 230 235 240 Arg His Ser Thr Gly Gly Met Asp Glu Leu Tyr Lys Gly Ser Lys Gln 245 250 255 Leu Glu Glu Leu Leu Ser Thr Ser Phe Asp Ile Gln Phe Asn Asp Leu 260 265 270 Thr Leu Leu Glu Thr Ala Phe Thr His Thr Ser Tyr Ala Asn Glu His 275 280 285 Arg Leu Leu Asn Val Ser His Asn Glu Arg Leu Glu Phe Leu Gly Asp 290 295 300 Ala Val Leu Gln Leu Ile Ile Ser Glu Tyr Leu Phe Ala Lys Tyr Pro 305 310 315 320 Lys Lys Thr Glu Gly Asp Met Ser Lys Leu Arg Ser Met Ile Val Arg 325 330 335 Glu Glu Ser Leu Ala Gly Phe Ser Arg Phe Cys Ser Phe Asp Ala Tyr 340 345 350 Ile Lys Leu Gly Lys Gly Glu Glu Lys Ser Gly Gly Arg Arg Arg Asp 355 360 365 Thr Ile Leu Gly Asp Leu Phe Glu Ala Phe Leu Gly Ala Leu Leu Leu 370 375 380 Asp Lys Gly Ile Asp Ala Val Arg Arg Phe Leu Lys Gln Val Met Ile 385 390 395 400 Pro Gln Val Glu Lys Gly Asn Phe Glu Arg Val Lys Asp Tyr Lys Thr 405 410 415 Cys Leu Gln Glu Phe Leu Gln Thr Lys Gly Asp Val Ala Ile Asp Tyr 420 425 430 Gln Val Ile Ser Glu Lys Gly Pro Ala His Ala Lys Gln Phe Glu Val 435 440 445 Ser Ile Val Val Asn Gly Ala Val Leu Ser Lys Gly Leu Gly Lys Ser 450 455 460 Lys Lys Leu Ala Glu Gln Asp Ala Ala Lys Asn Ala Leu Ala Gln Leu 465 470 475 480 Ser Glu Val <210> 48 <211> 483 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polypeptide" <400> 48 Met Lys Gln Leu Glu Glu Leu Leu Ser Thr Ser Phe Asp Ile Gln Phe 1 5 10 15 Asn Asp Leu Thr Leu Leu Glu Thr Ala Phe Thr His Thr Ser Tyr Ala 20 25 30 Asn Glu His Arg Leu Leu Asn Val Ser His Asn Glu Arg Leu Glu Phe 35 40 45 Leu Gly Asp Ala Val Leu Gln Leu Ile Ile Ser Glu Tyr Leu Phe Ala 50 55 60 Lys Tyr Pro Lys Lys Thr Glu Gly Asp Met Ser Lys Leu Arg Ser Met 65 70 75 80 Ile Val Arg Glu Glu Ser Leu Ala Gly Phe Ser Arg Phe Cys Ser Phe 85 90 95 Asp Ala Tyr Ile Lys Leu Gly Lys Gly Glu Glu Lys Ser Gly Gly Arg 100 105 110 Arg Arg Asp Thr Ile Leu Gly Asp Leu Phe Glu Ala Phe Leu Gly Ala 115 120 125 Leu Leu Leu Asp Lys Gly Ile Asp Ala Val Arg Arg Phe Leu Lys Gln 130 135 140 Val Met Ile Pro Gln Val Glu Lys Gly Asn Phe Glu Arg Val Lys Asp 145 150 155 160 Tyr Lys Thr Cys Leu Gln Glu Phe Leu Gln Thr Lys Gly Asp Val Ala 165 170 175 Ile Asp Tyr Gln Val Ile Ser Glu Lys Gly Pro Ala His Ala Lys Gln 180 185 190 Phe Glu Val Ser Ile Val Val Asn Gly Ala Val Leu Ser Lys Gly Leu 195 200 205 Gly Lys Ser Lys Lys Leu Ala Glu Gln Asp Ala Ala Lys Asn Ala Leu 210 215 220 Ala Gln Leu Ser Glu Val Gly Ser Val Ser Lys Gly Glu Glu Asp Asn 225 230 235 240 Met Ala Ile Ile Lys Glu Phe Met Arg Phe Lys Val His Met Glu Gly 245 250 255 Ser Val Asn Gly His Glu Phe Glu Ile Glu Gly Glu Gly Glu Gly Arg 260 265 270 Pro Tyr Glu Gly Thr Gln Thr Ala Lys Leu Lys Val Thr Lys Gly Gly 275 280 285 Pro Leu Pro Phe Ala Trp Asp Ile Leu Ser Pro Gln Phe Met Tyr Gly 290 295 300 Ser Lys Ala Tyr Val Lys His Pro Ala Asp Ile Pro Asp Tyr Leu Lys 305 310 315 320 Leu Ser Phe Pro Glu Gly Phe Lys Trp Glu Arg Val Met Asn Phe Glu 325 330 335 Asp Gly Gly Val Val Thr Val Thr Gln Asp Ser Ser Leu Gln Asp Gly 340 345 350 Glu Phe Ile Tyr Lys Val Lys Leu Arg Gly Thr Asn Phe Pro Ser Asp 355 360 365 Gly Pro Val Met Gln Lys Lys Thr Met Gly Trp Glu Ala Ser Ser Glu 370 375 380 Arg Met Tyr Pro Glu Asp Gly Ala Leu Lys Gly Glu Ile Lys Gln Arg 385 390 395 400 Leu Lys Leu Lys Asp Gly Gly His Tyr Asp Ala Glu Val Lys Thr Thr 405 410 415 Tyr Lys Ala Lys Lys Pro Val Gln Leu Pro Gly Ala Tyr Asn Val Asn 420 425 430 Ile Lys Leu Asp Ile Thr Ser His Asn Glu Asp Tyr Thr Ile Val Glu 435 440 445 Gln Tyr Glu Arg Ala Glu Gly Arg His Ser Thr Gly Gly Met Asp Glu 450 455 460 Leu Tyr Lys Lys Arg Pro Ala Ala Thr Lys Lys Ala Gly Gln Ala Lys 465 470 475 480 Lys Lys Lys <210> 49 <211> 1423 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polypeptide" <400> 49 Met Asp Tyr Lys Asp His Asp Gly Asp Tyr Lys Asp His Asp Ile Asp 1 5 10 15 Tyr Lys Asp Asp Asp Asp Lys Met Ala Pro Lys Lys Lys Arg Lys Val 20 25 30 Gly Ile His Gly Val Pro Ala Ala Asp Lys Lys Tyr Ser Ile Gly Leu 35 40 45 Ala Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu Tyr 50 55 60 Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg His 65 70 75 80 Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly Glu 85 90 95 Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr Thr 100 105 110 Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn Glu 115 120 125 Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser Phe 130 135 140 Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly Asn 145 150 155 160 Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr His 165 170 175 Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg Leu 180 185 190 Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe Leu 195 200 205 Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu Phe 210 215 220 Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro Ile 225 230 235 240 Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu Ser 245 250 255 Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu Lys 260 265 270 Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu Thr 275 280 285 Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu Gln 290 295 300 Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala Gln 305 310 315 320 Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu Ser 325 330 335 Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile Thr 340 345 350 Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His His 355 360 365 Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro Glu 370 375 380 Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala Gly 385 390 395 400 Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile Lys 405 410 415 Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys Leu 420 425 430 Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly Ser 435 440 445 Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg Arg 450 455 460 Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile Glu 465 470 475 480 Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala Arg 485 490 495 Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr Ile 500 505 510 Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala Gln 515,520,525 Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn Glu 530 535 540 Lys Val Leu Pro Lys Ser Leu Leu Tyr Glu Tyr Phe Thr Val Tyr 545 550 555 560 Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys Pro 565,570,575 Gly Glu Gln Lys Lys Wing Ile Val Asp 580,585,590 Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr Phe 595,600,605 Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu Asp 610 615 620 Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile Ile 625 630 635 640 Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu Glu 645 650 655 Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile Glu 660 665 670 Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met Lys 675 680 685 Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg Lys 690 695 700 Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu Asp 705 710 715 720 Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu Ile 725 730 735 His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln Val 740 745 750 Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala Gly 755 760 765 Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val Asp 770 775 780 Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile Val Ile 785 790 795 800 Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn Ser 805 810 815 Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly Ser 820 825 830 Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln Asn Glu 835 840 845 Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr Val Asp 850 855 860 Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp His Ile 865 870 875 880 Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys Val Leu 885 890 895 Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro Ser Glu 900 905 910 Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu Asn Ala 915 920 925 Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr Lys Ala Glu Arg 930 935 940 Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe Ile Lys Arg Gln Leu 945 950 955 960 Val Glu Thr Arg Gln Ile Thr Lys His Val Ala Gln Ile Leu Asp Ser 965 970 975 Arg Met Asn Thr Lys Tyr Asp Glu Asn Asp Lys Leu Ile Arg Glu Val 980 985 990 Lys Val Ile Thr Leu Lys Ser Lys Leu Val Ser Asp Phe Arg Lys Asp 995 1000 1005 Phe Gln Phe Tyr Lys Val Arg Glu Ile Asn Asn Tyr His His Ala 1010 1015 1020 His Asp Ala Tyr Leu Asn Ala Val Val Gly Thr Ala Leu Ile Lys 1025 1030 1035 Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr Lys 1040 1045 1050 Val Tyr Asp Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu Ile 1055 1060 1065 Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile Met Asn 1070 1075 1080 Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile Arg Lys 1085 1090 1095 Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile Val Trp 1100 1105 1110 Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser Met 1115 1120 1125 Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly 1130 1135 1140 Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys Leu 1145 1150 1155 Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe 1160 1165 1170 Asp Ser Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val 1175 1180 1185 Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu 1190 1195 1200 Gly Ile Thr Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile 1205 1210 1215 Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu 1220 1225 1230 Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn Gly 1235 1240 1245 Arg Lys Arg Met Leu Ala Ser Ala Gly Glu Leu Gln Lys Gly Asn 1250 1255 1260 Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr Leu Ala 1265 1270 1275 Ser His Tyr Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn Glu Gln 1280 1285 1290 Lys Gln Leu Phe Val Glu Gln His Lys His Tyr Leu Asp Glu Ile 1295 1300 1305 Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg Val Ile Leu Ala Asp 1310 1315 1320 Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr Asn Lys His Arg Asp 1325 1330 1335 Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile Ile His Leu Phe Thr 1340 1345 1350 Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe Asp Thr 1355 1360 1365 Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu Asp 1370 1375 1380 Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr Arg 1385 1390 1395 Ile Asp Leu Ser Gln Leu Gly Gly Asp Lys Arg Pro Ala Ala Thr 1400 1405 1410 Lys Lys Ala Gly Gln Ala Lys Lys Lys Lys 1415 1420 <210> 50 <211> 2012 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <400> 50 gaatgctgcc ctcagacccg cttcctccct gtccttgtct gtccaaggag aatgaggtct 60 cactggtgga ttcggacta ccctgaggag ctggcacctg agggacaagg ccccccacct 120 gcccagctcc agcctctgat gaggggtggg agagagctac atgaggttgc taagaaagcc 180 tcccctgaag gagaccacac agtgtgtgag gttggagtct ctagcagcgg gttctgtgcc 240 cccagggata gtctggctgt ccaggcactg ctcttgatat aaacaccacc tcctagttat 300 gaaaccatgc ccattctgcc tctctgtatg gaaaagagca tggggctggc ccgtggggtg 360 gtgtccactt taggccctgt gggagatcat gggaacccac gcagtgggtc ataggctctc 420 tcatttacta ctcacatcca ctctgtgaag aagcgattat gatctctcct ctagaaactc 480 gtagagtccc atgtctgccg gcttccagag cctgcactcc tccaccttgg cttggctttg 540 ctggggctag aggagctagg atgcacagca gctctgtgac cctttgtttg agaggaacag 600 gaaaaccacc cttctctctg gcccactgtg tcctcttcct gccctgccat ccccttctgt 660 gaatgttaga cccatgggag cagctggtca gaggggaccc cggcctgggg cccctaaccc 720 tatgtagcct cagtcttccc atcaggctct cagctcagcc tgagtgttga ggccccagtg 780 gctgctctgg gggcctcctg agtttctcat ctgtgcccct ccctccctgg cccaggtgaa 840 ggtgtggttc cagaaccgga ggacaaagta caaacggcag aagctggagg aggaagggcc 900 tgagtccgag cagaagaaga agggctccca tcacatcaac cggtggcgca ttgccacgaa 960 gcaggccaat ggggaggaca tcgatgtcac ctccaatgac aagcttgcta gcggtgggca 1020 accacaaacc cacgagggca gagtgctgct tgctgctggc caggcccctg cgtgggccca 1080 agctggactc tggccactcc ctggccaggc tttggggagg cctggagtca tggccccaca 1140 gggcttgaag cccggggccg ccattgacag agggacaagc aatgggctgg ctgaggcctg 1200 ggaccacttg gccttctcct cggagagcct gcctgcctgg gcgggcccgc ccgccaccgc 1260 agcctcccag ctgctctccg tgtctccaat ctcccttttg ttttgatgca tttctgtttt 1320 aatttatttt ccaggcacca ctgtagttta gtgatcccca gtgtccccct tccctatggg 1380 aataataaaa gtctctctct taatgacacg ggcatccagc tccagcccca gagcctgggg 1440 tggtagattc cggctctgag ggccagtggg ggctggtaga gcaaacgcgt tcagggcctg 1500 ggagcctggg gtggggtact ggtggagggg gtcaagggta attcattaac tcctctcttt 1560 tgttggggga ccctggtctc tacctccagc tccacagcag gagaaacagg ctagacatag 1620 ggaagggcca tcctgtatct tgagggagga caggcccagg tctttcttaa cgtattgaga 1680 ggtgggaatc aggcccaggt agttcaatgg gagagggaga gtgcttccct ctgcctagag 1740 actctggtgg cttctccagt tgaggagaaa ccagaggaaa ggggaggatt ggggtctggg 1800 ggagggaaca ccattcacaa aggctgacgg ttccagtccg aagtcgtggg cccaccagga 1860 tgctcacctg tccttggaga accgctgggc aggttgagac tgcagagaca gggcttaagg 1920 ctgagcctgc aaccagtccc cagtgactca gggcctcctc agcccaagaa agagcaacgt 1980 gccagggccc gctgagctct tgtgttcacc tg 2012 <210> 51 <211> 1153 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polypeptide" <400> 51 Met Lys Arg Pro Ala Ala Thr Lys Lys Ala Gly Gln Ala Lys Lys Lys 1 5 10 15 Lys Ser Asp Leu Val Leu Gly Leu Asp Ile Gly Ile Gly Ser Val Gly 20 25 30 Val Gly Ile Leu Asn Lys Val Thr Gly Glu Ile Ile His Lys Asn Ser 35 40 45 Arg Ile Phe Pro Ala Ala Gln Ala Glu Asn Asn Leu Val Arg Arg Thr 50 55 60 Asn Arg Gln Gly Arg Arg Leu Ala Arg Arg Lys Lys His Arg Arg Val 65 70 75 80 Arg Leu Asn Arg Leu Phe Glu Glu Ser Gly Leu Ile Thr Asp Phe Thr 85 90 95 Lys Ile Ser Ile Asn Leu Asn Pro Tyr Gln Leu Arg Val Lys Gly Leu 100 105 110 Thr Asp Glu Leu Ser Asn Glu Glu Leu Phe Ile Ala Leu Lys Asn Met 115 120 125 Val Lys His Arg Gly Ile Ser Tyr Leu Asp Asp Ala Ser Asp Asp Gly 130 135 140 Asn Ser Ser Val Gly Asp Tyr Ala Gln Ile Val Lys Glu Asn Ser Lys 145 150 155 160 Gln Leu Glu Thr Lys Thr Pro Gly Gln Ile Gln Leu Glu Arg Tyr Gln 165 170 175 Thr Tyr Gly Gln Leu Arg Gly Asp Phe Thr Val Glu Lys Asp Gly Lys 180 185 190 Lys His Arg Leu Ile Asn Val Phe Pro Thr Ser Ala Tyr Arg Ser Glu 195 200 205 Ala Leu Arg Ile Leu Gln Thr Gln Gln Glu Phe Asn Pro Gln Ile Thr 210 215 220 Asp Glu Phe Ile Asn Arg Tyr Leu Glu Ile Leu Thr Gly Lys Arg Lys 225 230 235 240 Tyr Tyr His Gly Pro Gly Asn Glu Lys Ser Arg Thr Asp Tyr Gly Arg 245 250 255 Tyr Arg Thr Ser Gly Glu Thr Leu Asp Asn Ile Phe Gly Ile Leu Ile 260 265 270 Gly Lys Cys Thr Phe Tyr Pro Asp Glu Phe Arg Ala Ala Lys Ala Ser 275 280 285 Tyr Thr Ala Gln Glu Phe Asn Leu Leu Asn Asp Leu Asn Asn Leu Thr 290 295 300 Val Pro Thr Glu Thr Lys Lys Leu Ser Lys Glu Gln Lys Asn Gln Ile 305 310 315 320 Ile Asn Tyr Val Lys Asn Glu Lys Ala Met Gly Pro Ala Lys Leu Phe 325 330 335 Lys Tyr Ile Ala Lys Leu Leu Ser Cys Asp Val Ala Asp Ile Lys Gly 340 345 350 Tyr Arg Ile Asp Lys Ser Gly Lys Ala Glu Ile His Thr Phe Glu Ala 355 360 365 Tyr Arg Lys Met Lys Thr Leu Glu Thr Leu Asp Ile Glu Gln Met Asp 370 375 380 Arg Glu Thr Leu Asp Lys Leu Ala Tyr Val Leu Thr Leu Asn Thr Glu 385 390 395 400 Arg Glu Gly Ile Gln Glu Ala Leu Glu His Glu Phe Ala Asp Gly Ser 405 410 415 Phe Ser Gln Lys Gln Val Asp Glu Leu Val Gln Phe Arg Lys Ala Asn 420 425 430 Ser Ser Ile Phe Gly Lys Gly Trp His Asn Phe Ser Val Lys Leu Met 435 440 445 Met Glu Leu Ile Pro Glu Leu Tyr Glu Thr Ser Glu Glu Gln Met Thr 450 455 460 Ile Leu Thr Arg Leu Gly Lys Gln Lys Thr Thr Ser Ser Ser Asn Lys 465 470 475 480 Thr Lys Tyr Ile Asp Glu Lys Leu Leu Thr Glu Glu Ile Tyr Asn Pro 485 490 495 Val Val Ala Lys Ser Val Arg Gln Ala Ile Lys Ile Val Asn Ala Ala 500 505 510 Ile Lys Glu Tyr Gly Asp Phe Asp Asn Ile Val Ile Glu Met Ala Arg 515 520 525 Glu Thr Asn Glu Asp Asp Glu Lys Lys Ala Ile Gln Lys Ile Gln Lys 530 535 540 Ala Asn Lys Asp Glu Lys Asp Ala Ala Met Leu Lys Ala Ala Asn Gln 545 550 555 560 Tyr Asn Gly Lys Ala Glu Leu Pro His Ser Val Phe His Gly His Lys 565 570 575 Gln Leu Ala Thr Lys Ile Arg Leu Trp His Gln Gln Gly Glu Arg Cys 580 585 590 Leu Tyr Thr Gly Lys Thr Ile Ser Ile His Asp Leu Ile Asn Asn Ser 595 600 605 Asn Gln Phe Glu Val Asp His Ile Leu Pro Leu Ser Ile Thr Phe Asp 610 615 620 Asp Ser Leu Ala Asn Lys Val Leu Val Tyr Ala Thr Ala Asn Gln Glu 625 630 635 640 Lys Gly Gln Arg Thr Pro Tyr Gln Ala Leu Asp Ser Met Asp Asp Ala 645 650 655 Trp Ser Phe Arg Glu Leu Lys Ala Phe Val Arg Glu Ser Lys Thr Leu 660 665 670 Ser Asn Lys Lys Lys Glu Tyr Leu Leu Thr Glu Glu Asp Ile Ser Lys 675 680 685 Phe Asp Val Arg Lys Lys Phe Ile Glu Arg Asn Leu Val Asp Thr Arg 690 695 700 Tyr Ala Ser Arg Val Val Leu Asn Ala Leu Gln Glu His Phe Arg Ala 705 710 715 720 His Lys Ile Asp Thr Lys Val Ser Val Val Arg Gly Gln Phe Thr Ser 725 730 735 Gln Leu Arg Arg His Trp Gly Ile Glu Lys Thr Arg Asp Thr Tyr His 740 745 750 His His Ala Val Asp Ala Leu Ile Ile Ala Ala Ser Ser Gln Leu Asn 755 760 765 Leu Trp Lys Lys Gln Lys Asn Thr Leu Val Ser Tyr Ser Glu Asp Gln 770 775 780 Leu Leu Asp Ile Glu Thr Gly Glu Leu Ile Ser Asp Asp Glu Tyr Lys 785 790 795 800 Glu Ser Val Phe Lys Ala Pro Tyr Gln His Phe Val Asp Thr Leu Lys 805 810 815 Ser Lys Glu Phe Glu Asp Ser Ile Leu Phe Ser Tyr Gln Val Asp Ser 820 825 830 Lys Phe Asn Arg Lys Ile Ser Asp Ala Thr Ile Tyr Ala Thr Arg Gln 835 840 845 Ala Lys Val Gly Lys Asp Lys Ala Asp Glu Thr Tyr Val Leu Gly Lys 850 855 860 Ile Lys Asp Ile Tyr Thr Gln Asp Gly Tyr Asp Ala Phe Met Lys Ile 865 870 875 880 Tyr Lys Lys Asp Lys Ser Lys Phe Leu Met Tyr Arg His Asp Pro Gln 885 890 895 Thr Phe Glu Lys Val Ile Glu Pro Ile Leu Glu Asn Tyr Pro Asn Lys 900 905 910 Gln Ile Asn Glu Lys Gly Lys Glu Val Pro Cys Asn Pro Phe Leu Lys 915 920 925 Tyr Lys Glu Glu His Gly Tyr Ile Arg Lys Tyr Ser Lys Lys Gly Asn 930 935 940 Gly Pro Glu Ile Lys Ser Leu Lys Tyr Tyr Asp Ser Lys Leu Gly Asn 945 950 955 960 His Ile Asp Ile Thr Pro Lys Asp Ser Asn Asn Lys Val Val Leu Gln 965 970 975 Ser Val Ser Pro Trp Arg Ala Asp Val Tyr Phe Asn Lys Thr Thr Gly 980 985 990 Lys Tyr Glu Ile Leu Gly Leu Lys Tyr Ala Asp Leu Gln Phe Glu Lys 995 1000 1005 Gly Thr Gly Thr Tyr Lys Ile Ser Gln Glu Lys Tyr Asn Asp Ile 1010 1015 1020 Light Light Light Glu Gly Val Asp Ser Asp Ser Glu Phe Light Phe Thr 1025 1030 1035 Leu Tyr Lys Asn Asp Leu Leu Leu Val Lys Asp Thr Glu Thr Lys 1040 1045 1050 Glu Gln Gln Leu Phe Arg Phe Leu Ser Arg Thr Met Pro Lys Gln 1055 1060 1065 Lys His Tyr Val Glu Leu Lys Pro Tyr Asp Lys Gln Lys Phe Glu 1070 1075 1080 Gly Gly Glu Ala Leu Ile Lys Val Leu Gly Asn Val Ala Asn Ser 1085 1090 1095 Gly Gln Cys Lys Lys Gly Leu Gly Lys Ser Asn Ile Ser Ile Tyr 1100 1105 1110 Lys Val Arg Thr Asp Val Leu Gly Asn Gln His Ile Ile Lys Asn 1115 1120 1125 Glu Gly Asp Lys Pro Lys Leu Asp Phe Lys Arg Pro Ala Ala Thr 1130 1135 1140 Lys Lys Ala Gly Gln Ala Lys Lys Lys Lys 1145 1150 <210> 52 <211> 340 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 52 gagggcctat ttcccatgat tccttcatat ttgcatatac gatacaaggc tgttagagag 60 aatattggaa ttaatttgac tgtaacaca agatattag tacaaatac gtgacgtaga 120 aagtaataat ttctgggta gtttgcagtt ttaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatctt gtggaagga 240 cgaaacaccg ttacttaat cttgcagaag ctacaagat aaggctcat gccgaatca 300 acaccctgtc atttttggc agggtgtttt cgttatta 340 <210> 53 <211> 360 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <220> <221> modified_base <222> (288)..(317) <223> a, c, t, g, unknown or other <400> 53 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 ataattggaa ttaatttgac tgtaaacaca aagatattag tacaaaatac gtgacgtaga 120 aagtaataat ttcttgggta gtttgcagtt ttaaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatatctt gtggaaagga 240 cgaaacaccg ggttttagag ctatgctgtt ttgaatggtc ccaaaacnnnn nnnnnnnnn 300 nnnnnnnnnnn nnnnnngtt ttagagctat gctgttttga atggtcccaa aacttttttt 360 <210> 54 <211> 318 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <220> <221> modified_base <222> (250)..(269) <223> a, c, t, g, unknown or other <400> 54 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 aatattggaa ttaatttgac tgtaacaca agatattag tacaaatac gtgacgtaga 120 aagtaataat ttctgggta gtttgcagtt ttaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatctt gtggaagga 240 cgaaacaccn nnnnnnnn nnnnnnn ttttagagct agaaatagca agttaaaata 300 aggctagtcc gtttttt 318 <210> 55 <211> 325 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <220> <221> modified_base <222> (250)..(269) <223> a, c, t, g, unknown or other <400> 55 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 aatattggaa ttaatttgac tgtaacaca agatattag tacaaatac gtgacgtaga 120 aagtaataat ttctgggta gtttgcagtt ttaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatctt gtggaagga 240 cgaaacaccn nnnnnnnn nnnnnnn ttttagagct agaaatagca agttaaaata 300 aggctagtcc gttatcattt tttt 325 <210> 56 <211> 337 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <220> <221> modified_base <222> (250)..(269) <223> a, c, t, g, unknown or other <400> 56 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 aatattggaa ttaatttgac tgtaacaca agatattag tacaaatac gtgacgtaga 120 aagtaataat ttctgggta gtttgcagtt ttaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatactt gtggaagga 240 cgaaacaccn nnnnnnnn nnnnnnn ttttagagct agaaatagca agttaaaata 300 aggctagtcc gttatcact tgaaaaagtg tttttt 337 <210> 57 <211> 352 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <220> <221> modified_base <222> (250)..(269) <223> a, c, t, g, unknown or other <400> 57 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 aatattggaa ttaatttgac tgtaacaca agatattag tacaaatac gtgacgtaga 120 aagtaataat ttctgggta gtttgcagtt ttaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatctt gtggaagga 240 cgaaacaccn nnnnnnnn nnnnnnn ttttagagct agaaatagca agttaaaata 300 aggctagtcc gttatcact tgaaaagtg gcaccgagtc gtgcttttt tt 352 <210> 58 <211> 5101 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <400> 58 cgttacataa cttacggtaa atggcccgcc tggctgaccg cccaacgacc cccgcccatt 60 gacgtcaata atgacgtatg ttcccatagt aacgccaata gggactttcc attgacgtca 120 atgggtggag tatttacggt aaactgccca cttggcagta catcaagtgt atcatatgcc 180 aagtacgccc cctattgacg tcaatgacgg taaatggccc gcctggcatt atgcccagta 240 catgacctta tgggactttc ctacttggca gtacatctac gtattagtca tcgctattac 300 catggtcgag gtgagcccca cgttctgctt cactctcccc atctcccccc cctccccacc 360 cccaattttg tatttattta ttttttaatt attttgtgca gcgatggggg cggggggggg 420 gggggggcgc gcgccaggcg gggcggggcg gggcgagggg cggggcgggg cgaggcggag 480 aggtgcggcg gcagccaatc agagcggcgc gctccgaaag tttcctttta tggcgaggcg 540 gcggcggcgg cggccctata aaaagcgaag cgcgcggcgg gcgggagtcg ctgcgacgct 600 gccttcgccc cgtgccccgc tccgccgccg cctcgcgccg cccgccccgg ctctgactga 660 ccgcgttact cccacaggtg agcgggcggg acggcccttc tcctccgggc tgtaattagc 720 tgagcaagag gtaagggttt aagggatggt tggttggtgg ggtattaatg ttataattacc 780 tggagcacct gcctgaaatc actttttttc aggttggacc ggtgccacca tggactataa 840 ggaccacgac ggagactaca aggatcatga tattgattac aaagacgatg acgataagat 900 ggccccaaag aagaagcgga aggtcggtat ccacggagtc ccagcagccg acaagaagta 960 cagcatcggc ctggacatcg gcaccaactc tgtgggctgg gccgtgatca ccgacgagta 1020 caaggtgccc agcaagaaat tcaaggtgct gggcaacacc gaccggcaca gcatcaagaa 1080 gaacctgatc ggagccctgc tgttcgacag cggcgaaaca gccgaggcca cccggctgaa 1140 gagaaccgcc agaagaagat acaccagacg gaagaaccgg atctgctatc tgcaagagat 1200 cttcagcaac gagatggcca aggtggacga cagcttcttc cacagactgg aagagtcctt 1260 cctggtggaa gaggataaga agcacgagcg gcaccccatc ttcggcaaca tcgtggacga 1320 ggtggcctac cacgagaagt accccaccat ctaccacctg agaaagaaac tggtggacag 1380 caccgacaag gccgacctgc ggctgatcta tctggccctg gcccacatga tcaagttccg 1440 gggccacttc ctgatcgagg gcgacctgaa ccccgacaac agcgacgtgg acaagctggtt 1500 catccagctg gtgcagacct acaaccagct gttcgaggaa aaccccatca acgccagcgg 1560 cgtggacgcc aaggccatcc tgtctgccag actgagcaag agcagacggc tggaaaatct 1620 gatcgcccag ctgcccggcg agaagaagaa tggcctgttc ggcaacctga ttgccctgag 1680 cctgggcctg acccccaact tcaagagcaa cttcgacctg gccgaggatg ccaaactgca 1740 gctgagcaag gacacctacg aggacgacct ggacaacctg ctggcccaga tcggcgacca 1800 gtacgccgac ctgtttctgg ccgccaagaa cctgtccgac gccatcctgc tgagcgacat 1860 cctgagagtg aacaccgaga tcaccaggc cccctgagc gccctatga tcaagagata 1920 cgacgagcac caccaggacc tgaccctgct gaagctctc gtgcggcagc agctgcctga 1980 gaagtacaa gagattttct tcgaccagag caagaacggc tacgccggct acatgacgg 2040 cggagccagc caggaaggt tctacaagtt catcaagccc atcctggaaa agatggacgg 2100 caccgaggaa ctgctcgtga agctgaacag agaggacctg ctgcggaagc agcggacctt 2160 cgacaacggc agcatccccc accagatcca cctgggagag ctgcacgcca ttctgcggcg 2220 gcaggaagat ttttaccat tcctgaagga haaccgggaa aagatcgaga agatcctgac 2280 cttccgcatc ccctactacg tgggccctct ggccagggga aacagcagat tcgcctggat 2340 gaccagaaag agcgaggaaa ccatcacccc ctggaacttc gaggaagtgg tggacaaggg 2400 cgcttccgcc cagagcttca tcgagcggat gaccaacttc cataagaacc tgcccacga 2460 gaaggtgctg cccaagcaca gcctgctgta cgagtacttc accgtgtata acgagctgac 2520 caaagtgaaa tacgtgaccg agggaatgag aaagcccgcc ttcctgagcg gcgagcagaa 2580 aaaggccatc gtggacctgc tgttcaagac caaccggaaa gtgaccgtga agcagctgaa 2640 agaggactac ttcaagaaaa tcgagtgctt cgactccgtg gaaatctccg gcgtggaaga 2700 tcggttcaac gcctccctgg gcacatacca cgatctgctg aaaattatca aggacaagga 2760 cttcctggac aatgaggaaa acgaggacat tctggaagat atcgtgctga ccctgacact 2820 gtttgaggac agagagatga tcgaggaacg gctgaaaacc tatgcccacc tgttcgacga 2880 caaagtgatg aagcagctga agcggcggag atacaccggc tggggcaggc tgagccggaa 2940 gctgatcaac ggcatccggg acaagcagtc cggcaagaca atcctggatt tcctgaagtc 3000 cgacggcttc gccaacagaa acttcatgca gctgatccac gacgacagcc tgacctttaa 3060 agaggacatc cagaaagcccc aggtccgg ccaggcgat agcctgcacg agcacattgc 3120 caatctggcc ggcagccccg ccattagaa gggcatcctg cagacagtga aggtggtgga 3180 cgagctcgtg aaagtgatgg gccggcacaa gcccgagaac atcgtcg aaatggccag 3240 aggaaccag accaccaga agggagaga gaacagccgc gagagaatga agcggatcga 3300 agagggcatc aaagagctgg gcagccagat cctgaaagaa caccccgtgg aaaacaccca 3360 gctgcagaac gagaagctgt acctgtacta cctgcagaat gggcgggata tgtacgtgga 3420 ccaggaactg gatacacc ggctgtccga ctacgatgtg gaccatatcg tgcctcagag 3480 ctttctgaag gacgactcca tcgacaacaa ggtgctgacc agaagcgaca agaccgggg 3540 CAagagcgac aacgtgccct ccgagaggt cgtgagaag atgagaac actggcggca 3600 gctgctgaac gccaagctga ttaccagag aaagttcgac aatctgacca aggccgagag 3660 aggcggcctg agcgaactgg ataaggccgg cttcatcaag agacagctgg tggaaacccg 3720 gcagatcaca aagcacgtgg cacagatcct ggactcccgg atgaacacta agtacgacga 3780 gaatgacaag ctgatccggg aagtgaaagt gatcaccctg aagtccaagc tggtgtccga 3840 tttccggaag gatttccagt tttacaaagt gcgcgagatc aacaactacc accacgccca 3900 cgacgcctac ctgaacgccg tcgtgggaac cgccctgatc aaaaagtacc ctaagctgga 3960 aagcgagttc gtgtacggcg actacaaggt gtacgacgtg cggaagatga tcgccaagag 4020 cgagcaggaa atcggcaagg ctaccgccaa gtacttcttc tacagcaaca tcatgaactt 4080 tttcaagacc gagattaccc tggccaacgg cgagatccgg aagcggcctc tgatcgagac 4140 aaacggcgaa accggggaga tcgtgtggga taagggccgg gattttgcca ccgtgcggaa 4200 agtgctgagc atgccccaag tgaatatcgt gaaaaagacc gaggtgcaga caggcggctt 4260 cagcaaagag tctatcctgc ccaagaggaa cagcgataag ctgatcgcca gaaagaagga 4320 ctgggaccct aagaagtacg gcggcttcga cagccccacc gtggcctatt ctgtgctggt 4380 ggtggccaaa gtggaaaagg gcaagtccaa gaaactgaag agtgtgaaag agctgctggg 4440 gatcaccatc atggaaagaa gcagcttcga gaagaatccc atcgactttc tggaagccaa 4500 gggctacaaa gaagtgaaaa aggacctgat catcaagctg cctaagtact ccctgttcga 4560 gctggaaaac ggccggaaga gaatgctggc ctctgccggc gaactgcaga agggaaacga 4620 actggccctg ccctccaaat atgtgaactt cctgtacctg gccagccact atgagaagct 4680 gaagggctcc cccgaggata atgagcagaa acagctgttt gtggaacagc acaagcacta 4740 cctggacgag atcatcgagc agatcagcga gttctccaag agagtgatcc tggccgacgc 4800 taatctggac aaagtgctgt ccgcctacaa caagcaccgg gataagccca tcagagagca 4860 ggccgagaat atcatccacc tgtttaccct gaccaatctg ggagcccctg ccgccttcaa 4920 gtactttgac accaccatcg accggaagag gtacaccagc accaaagagg tgctggacgc 4980 caccctgatc caccagagca tcaccggcct gtacgagaca cggatcgacc tgtctcagct 5040 gggaggcgac tttctttttc ttagcttgac cagctttctt agtagcagca ggacgcttta 5100 a 5101 <210> 59 <211> 137 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 59 nnnnnnnnnn nnnnnnnnnn gtttttgtac tctcaagatt tagaaataaa tcttgcagaa 60 gctacaaaga taaggcttca tgccgaaatc aacaccctgt cattttatgg cagggtgttt 120 tcgttattta atttttt 137 <210> 60 <211> 123 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 60 nnnnnnnnnn nnnnnnnnnn gtttttgtac tctcagaaat gcagaagcta caaagataag 60 gcttcatgcc gaaatcaaca ccctgtcatt ttatggcagg gtgttttcgt tatttaattt 120 ttt 123 <210> 61 <211> 110 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 61 nnnnnnnnnn nnnnnnnnnn gtttttgtac tctcagaaat gcagaagcta caaagataag 60 gcttcatgcc gaaatcaaca ccctgtcatt ttatggcagg gtgttttttt 110 <210> 62 <211> 137 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 62 nnnnnnnnnn nnnnnnnnnn gttattgtac tctcaagatt tagaaataaa tcttgcagaa 60 gctacaaaga taaggcttca tgccgaaatc aacaccctgt cattttatgg cagggtgttt 120 tcgttattta atttttt 137 <210> 63 <211> 123 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 63 nnnnnnnnnn nnnnnnnnnn gttattgtac tctcagaaat gcagaagcta caaagataag 60 gcttcatgcc gaaatcaaca ccctgtcatt ttatggcagg gtgttttcgt tatttaattt 120 ttt 123 <210> 64 <211> 110 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 64 nnnnnnnnnn nnnnnnnnnn gttattgtac tctcagaaat gcagaagcta caaagataag 60 gcttcatgcc gaaatcaaca ccctgtcatt ttatggcagg gtgttttttt 110 <210> 65 <211> 137 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 65 nnnnnnnnnn nnnnnnnnnn gttattgtac tctcaagatt tagaaataaa tcttgcagaa 60 gctacaatga taaggcttca tgccgaaatc aacaccctgt cattttatgg cagggtgttt 120 tcgttattta atttttt 137 <210> 66 <211> 123 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 66 nnnnnnnnnn nnnnnnnnnn gttattgtac tctcagaaat gcagaagcta caatgataag 60 gcttcatgcc gaaatcaaca ccctgtcatt ttatggcagg gtgttttcgt tatttaattt 120 ttt 123 <210> 67 <211> 110 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 67 nnnnnnnnnn nnnnnnnnnn gttattgtac tctcagaaat gcagaagcta caatgataag 60 gcttcatgcc gaaatcaaca ccctgtcatt ttatggcagg gtgttttttt 110 <210> 68 <211> 107 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 68 nnnnnnnnnnn nnnnnnnnn gttttagagc tgtggaaaca cagcgagtta aaataaggct 60 tagtccgtac tcaacttgaa aaggtggcac cgattcggtg tttttt 107 <210> 69 <211> 4263 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <400> 69 atgaaaagc cggcggccac gaaaaaggcc ggccaggcaa aaaagaaaaa gaccaagcccc 60 tacagcatcg gcctggacat cggcaccaat agcgtgggct gggccgtgac caccgacaac 120 tacaaggtgc ccagcaagaa aatgaaggtg ctgggcaaca cctccaagaa gtacatcaag 180 aaaaacctgc tgggcgtgct gctgttcgac agcggcatta cagccgaggg cagacggctg 240 aagagaaccg ccagacggcg gtacacccgg cggagaaaca gaatcctgta tctgcaagag 300 atcttcagca ccgagatggc tacctggac gacgccttct tccagcggct ggacgacagc 360 ttcctggtgc ccgacgacaa gcgggacagc aagcccca tcttcggcaa cctggtggaa 420 gagaaggcct accacgacga gttccccacc atctaccacc tgagaaagta cctggccgac 480 agcaccaga aggccgacct gagactggtg tatctggccc tggcccacat gatcaagtac 540 cggggccact tcctgatcga gggcgagttc aacagcaga acacgacat ccagagaac 600 ttccaggact tcctggacac ctacaacgcc atcttcgaga gcgacctgtc cctggaaaac 660 agcaagcagc tggagagat cgtgaggac aagatcagca agctgaaaa gaggaccgc 720 atcctgaagc tgttccccgg cgagagaac agcggaatct tcagcgagtt tctgaagctg 780 atcgtgggca accaggccga cttcagaag tgctcaacc tggacgaga agccagcctg 840 cactcagca agagagcta cgacgaggac ctggaaccc tgctggata tatcggcgac 900 gactacagcg acgtgttcct gaaggccaag aagctctacg acgctatcct gctgagcggc 960 ttcctgaccg tgaccgacaa cgagacagag gccccactga gcagcgccat gattaagcgg 1020 tacaacgagc acaaagaga tctggctctg ctgaaagagt acatccggaa catcagcctg 1080 aaaacctaca atgaggtgtt caaggacgacgac accaagaacg gctacgccgg ctacatcgac 1140 ggcaagacca accaggaaga tttctatgtg tacctagaagaagctgctggc cgagttcgag 1200 ggggccgact actttctgga aaaaatcgac cgcgaggatt tcctgcggaa gcagcggacc 1260 ttcgacaacg gcagcatccc ctaccagatc catctgcagg aaatgcgggc catcctggac 1320 aagcaggcca agttctaccc attcctggcc aagaacaaag agcggatcga gaagatcctg 1380 accttccgca tcccttacta cgtgggcccc ctggccagag gcaacagcga ttttgcctgg 1440 1500 gagtccagcg ccgaggcctt catcaccgg atgaccagct tcgacctgta cctgcccgag 1560 gaaaaggtgc tgcccaagca cagcctgctg tacgagacat tcaatgtgta taacgagctg 1620 accaaagtgc gtttatcgc cgagtctatg cgggactacc agttcctgga ctccaagcag 1680 aaaaaggaca tcgtgcggct gtactcaag gatagcgga aagtgaccga taggacatc 1740 atcgagtacc tgcacgccat ctacggctac gatggcatcg agctgaaggg catcgagaag 1800 cagttcaact ccagcctgag cacataccac gacctgctga acattatca cgacaagaa 1860 tttctggacg actccagcaa cgaggccatc atcgaagaga tcatccacac cctgaccacc 1920 tttgaggacc gcgagatgat caagcagcgg ctgagcaagt tcgagaacat ctcgacaag 1980 agcgtgctga aaagctgag cagacggcac tacaccggct ggggcaagct gagcgccaag 2040 ctgatcaacg gcatccggga cgagaagtcc ggcaacacaa tcctggacta cctgatcgac 2100 gacggcatca gcaaccggaa cttcatgcag ctgatccacg acgacgccct gagcttcaag 2160 aagaagatcc agaaggccca gatcatcgggg gaxgagaca agggxacat caaaagtc 2220 gtgaagtccc tgcccggcag ccccgccatc agaagggaa tcctgcag catcagatc 2280 gtggacgagc tcgtgaaagt gatgggcggc agaagcccg agagcatcgt ggtggaaatg 2340 gctagagaga accagtacac caatcagggc agagcaca gccagcagag actgagaga 2400 ctggaaaagt ccctgaaaga gctgggcagc aagattctga aagaatat ccctgccaag 2460 ctgtccaaga tcgacaacaa cgccctgcag aacgaccggc tgtacctgta ctacctgcag 2520 aatggcagg acatgtatac aggcgacgac ctggatatcg accgcctgag acacacgac 2580 atcgaccata ttatccccca ggccttcctg aaagacaaca gcattgacaa caagtgctg 2640 gtgtcctccg ccagcaccg cggcaagtcc gatgatgtgc ccagcctgga agtcgtgaaa 2700 aagagaaaga ccttctgta tcagctgctg aaagcaagc tgattagcca gaggaagttc 2760 vakaacctga ccaaggccga gagaggcggc ctgagccctg aagataggc cggcttcatc 2820 cagagacagc tggtggaaac ccggcagatc accaagcacg tggccagact gctggatgag 2880 aagttttaca acaagaagga cgagacaac cggggccgtgc ggaccgtgaa gatcatcacc 2940 ctgaagtcca ccctgtgtc ccagttccgg aaggactcg agctgtataa agtgcgcgag 3000 atcaatgact ttcaccacgc ccacgacgcc tacctgaatg ccgtggtggc ttccgccctg 3060 ctgaagaagt accctaagct ggaacccgag ttcgtgtacg gcgactaccc caagtacaac 3120 tccttcagag agcggaagtc cgccaccgag aaggtgtact tctactccaa catcatgaat 3180 atctttaaga agtccatctc cctggccgat ggcagagtga tcgagcggcc cctgatcgaa 3240 gtgaacgaag agacaggcga gagcgtgtgg aaaaaaa gcgacctggc caccgtgcgg 3300 cgggtgctga gttatcctca agtgaatgtc gtgaagaagg tggaagaaca gaaccacggc 3360 ctggatcggg gcaagcccaa gggcctgttc aacgccaacc tgtccagcaa gcctaagccc 3420 aactccaacg agaatctcgt gggggccaaa gagtacctgg accctaagaa gtacggcgga 3480 tacgccggca tctccaatag cttcaccgtg ctcgtgaagg gcacaatcga gaagggcgct 3540 aagaaaaaga tcacaaacgt gctggaattt caggggatct ctatcctgga ccggatcaac 3600 taccggaagg ataagctgaa ctttctgctg gaaaaaggct acaaggacat tgagctgatt 3660 atcgagctgc ctaagtactc cctgttcgaa ctgagcgacg gctccagacg gatgctggcc 3720 tccatcctgt ccaccaacaa caagcggggc gagatccaca agggaaacca gatcttcctg 3780 agccagaaat ttgtgaaact gctgtaccac gccaagcgga tctccaacac catcaatgag 3840 aaccaccgga aatacgtgga aaaccacaag aaagagtttg aggaactgtt ctactacatc 3900 ctggagttca acgagaacta tgtgggagcc aagaagaacg gcaaactgct gaactccgcc 3960 ttccagagct ggcagaacca cagcatcgac gagctgtgca gctccttcat cggccctacc 4020 ggcagcgagc ggaagggact gtttgagctg acctccagag gctctgccgc cgactttgag 4080 ttcctgggag tgaagatccc ccggtacaga gactacaccc cctctagtct gctgaaggac 4140 gccaccctga tccaccagag cgtgaccggc ctgtacgaaa cccggatcga cctggctaag 4200 ctgggcgagg gaaagcgtcc tgctgctact aagaaagctg gtcaagctaa gaaaaagaaa 4260 taa 4263 <210> 70 <211> 84 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 70 ggaaccattc ataacagcat agcaagttat aataaggcta gtccgttatc aacttgaaaa 60 agtggcaccg agtcggtgct tttt 84 <210> 71 <211> 36 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 71 gttatagagc tatgctgtta tgaatggtcc caaaac 36 <210> 72 <211> 84 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 72 ggaaccattc aatacagcat agcaagttaa tataaggcta gtccgttatc aacttgaaaa 60 agtggcaccg agtcggtgct tttt 84 <210> 73 <211> 36 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 73 gtattagagc tatgctgtat tgaatggtcc caaaac 36 <210> 74 <211> 103 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 74 nnnnnnnnnn nnnnnnnnnn gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgctttt ttt 103 <210> 75 <211> 103 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 75 nnnnnnnnnn nnnnnnnnnn gtattagagc tagaaatagc aagttaatat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgctttt ttt 103 <210> 76 <211> 123 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 76 nnnnnnnnnnn nnnnnnnnn gttttagagc tatgctgttt tggaaacaaa acagcatagc 60 aagttaaaat aaggctagtc cgttatcaac ttgaaaaagt ggcaccgagt cggtgcttttt 120 ttt 123 <210> 77 <211> 123 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 77 nnnnnnnnnnn nnnnnnnnn gtattagagc tatgctgtat tggaaacaat acagcatagc 60 aagttaatat aaggctagtc cgttatcaac ttgaaaaagt ggcaccgagt cggtgctttt 120 ttt 123 <210> 78 <211> 20 <212> DNA <213> Homo sapiens <400> 78 gtcacctcca atgactaggg 20 <210> 79 <211> 984 <212> PRT <213> Campylobacter jejuni <400> 79 Met Ala Arg And Leu Ala Phe Asp And Gly And Ser Ser And Gly Trp 1 5 10 15 Ala Phe Ser Glu Asn Asp Glu Leu Lys Asp Cys Gly Val Arg Ile Phe 20 25 30 Thr Lys Val Glu Asn Pro Lys Thr Gly Glu Ser Leu Ala Leu Pro Arg 35 40 45 Arg Leu Ala Arg Ser Ala Arg Lys Arg Leu Ala Arg Arg Lys Ala Arg 50 55 60 Leu Asn His Leu Lys His Leu Ile Ala Asn Glu Phe Lys Leu Asn Tyr 65 70 75 80 Glu Asp Tyr Gln Ser Phe Asp Glu Ser Leu Ala Lys Ala Tyr Lys Gly 85 90 95 Ser Leu Ile Ser Pro Tyr Glu Leu Arg Phe Arg Ala Leu Asn Glu Leu 100 105 110 Lys Gln Asp Phe Ala Arg Val Ile Leu His Ile Ala Lys Arg 115 120 125 Arg Gly Tyr Asp Asp Ile Lys Asn Ser Asp Asp Lys Glu Lys Gly Ala 130 135 140 With Lys Only Only Only With Lys On Gln Asn Glu Glu Lys With Only On Tyr Gln 145 150 155 160 Ser Val Gly Glu Tyr Leu Tyr Lys Glu Tyr Phe Gln Lys Phe Lys Glu 165 170 175 Asn Ser Lys Glu Phe Thr Asn Val Arg Asn Lys Glu Ser Tyr Glu 180 185 190 Arg Cys contains Gln Ser with Lys Asp and Glu with Lys 195 200 205 Lys Lys Gln Arg Glu Phe Gly Phe Ser Phe Ser Lys Phe Glu Glu 210 215 220 Glu Val Leu Ser Val Ala Phe Tyr Lys Arg Ala Leu Lys Asp Phe Ser 225 230 235 240 His Leu Val Gly Asn Cys Ser Phe Phe Thr Asp Glu Lys Arg Ala Pro 245 250 255 Lys Asn Ser Pro Leu Ala Phe Met Phe Val Ala Leu Thr Arg Ile Ile 260 265 270 Asn Leu Leu Asn Asn Leu Lys Asn Thr Glu Gly Ile Leu Tyr Thr Lys 275 280 285 Asp Asp Leu Asn Ala Leu Leu Asn Glu Val Leu Lys Asn Gly Thr Leu 290 295 300 Thr Tyr Lys Gln Thr Lys Lys Leu Leu Gly Leu Ser Asp Asp Tyr Glu 305 310 315 320 Phe Lys Gly Glu Lys Gly Thr Tyr Phe Ile Glu Phe Lys Lys Tyr Lys 325 330 335 Glu Phe Ile Lys Ala Leu Gly Glu His Asn Leu Ser Gln Asp Asp Leu 340 345 350 Asn Glu Ile Ala Lys Asp Ile Thr Leu Ile Lys Asp Glu Ile Lys Leu 355 360 365 Lys Lys Ala Leu Ala Lys Tyr Asp Leu Asn Gln Asn Gln Ile Asp Ser 370 375 380 Leu Ser Lys Leu Glu Phe Lys Asp His Leu Asn Ile Ser Phe Lys Ala 385 390 395 400 Leu Lys Leu Val Thr Pro Leu Met Leu Glu Gly Lys Lys Tyr Asp Glu 405 410 415 Ala Cys Asn Glu Leu Asn Leu Lys Val Ala Ile Asn Glu Asp Lys Lys 420 425 430 Asp Phe Leu Pro Ala Phe Asn Glu Thr Tyr Tyr Lys Asp Glu Val Thr 435 440 445 Asn Pro Val Val Leu Arg Ala Ile Lys Glu Tyr Arg Lys Val Leu Asn 450 455 460 Ala Leu Leu Lys Lys Tyr Gly Lys Val His Lys Ile Asn Ile Glu Leu 465 470 475 480 Ala Arg Glu Val Gly Lys Asn His Ser Gln Arg Ala Lys Ile Glu Lys 485 490 495 Glu Gln Asn Glu Asn Tyr Lys Ala Lys Lys Asp Ala Glu Leu Glu Cys 500 505 510 Glu Lys Leu Gly Leu Lys Ile Asn Ser Lys Asn Ile Leu Lys Leu Arg 515 520 525 Leu Phe Lys Glu Gln Lys Glu Phe Cys Ala Tyr Ser Gly Glu Lys Ile 530 535 540 Lys Ile Ser Asp Leu Gln Asp Glu Lys Met Leu Glu Ile Asp His Ile 545 550 555 560 Tyr Pro Tyr Ser Arg Ser Phe Asp Asp Ser Tyr Met Asn Lys Val Leu 565 570 575 Val Phe Thr Lys Gln Asn Gln Glu Lys Leu Asn Gln Thr Pro Phe Glu 580 585 590 Ala Phe Gly Asn Asp Ser Ala Lys Trp Gln Lys Ile Glu Val Leu Ala 595 600 605 Lys Asn Leu Pro Thr Lys Lys Gln Lys Arg Ile Leu Asp Lys Asn Tyr 610 615 620 Lys Asp Lys Glu Gln Lys Asn Phe Lys Asp Arg Asn Leu Asn Asp Thr 625 630 635 640 Arg Tyr Ile Ala Arg Leu Val Leu Asn Tyr Thr Lys Asp Tyr Leu Asp 645 650 655 Phe Leu Pro Leu Ser Asp Asp Glu Asn Thr Lys Leu Asn Asp Thr Gln 660 665 670 Lys Gly Ser Lys Val His Val Glu Ala Lys Ser Gly Met Leu Thr Ser 675 680 685 Ala Leu Arg His Thr Trp Gly Phe Ser Ala Lys Asp Arg Asn Asn His 690 695 700 Leu His His Ala Ile Asp Ala Val Ile Ile Ala Tyr Ala Asn Asn Ser 705 710 715 720 Ile Val Lys Ala Phe Ser Asp Phe Lys Lys Glu Gln Glu Ser Asn Ser 725 730 735 Glu Leu Tyr Ala Lys Lys Ser Glu Leu Asp Tyr Lys Asn Lys 740,745,750 Arg Lys Phe Glu Pro Phe Ser Gly Phe Arg Gln Lys Val Leu Asp 755,760,765 Lys Ile Asp Glu Ile Phe Val Ser Lys Pro Glu Arg Lys Lys Pro Ser 770,775,780 Gly Ala Leu His Glu Glu Thr Phe Arg Lys Glu Glu Glu Phe Tyr Gln 785,790,795,800 Ser Tyr Gly Gly Lys Glu Gly Val Leu Lys Ala Leu Glu Leu Gly Lys 805 810 815 Ile Arg Lys Val Asn Gly Lys Ile Val Lys Asn Gly Asp Met Phe Arg 820 825 830 Val Asp Ile Phe Lys His Lys Lys Thr Asn Lys Phe Tyr Ala Val Pro 835 840 845 Ile Tr Thr Met Asp Phe Ala Leu Lys Val Leu Pro Asn Lys Ala Val 850 855 860 Only Arg Ser Lys Lys Gly Glu Ile Lys Asp Trp Ile Leu Met Asp Glu 865 870 875 880 Asn Tyr Glu Phe Cys Phe Ser Leu Tyr Lys Asp Ser Leu Ile Leu Ile 885,890,895 Gln Thr Lys Asp Met Gln Glu Pro Glu Phe Val Tyr Asn Ala Phe 900 905 910 His Asp Asn Lys Phe 915,920,925 Glu Thr Serves Lys Asn With Gln Lys And Has Phe Lys Asn On Glu 930,935,940 Lys Glu Val Ile Ala Lys Ser Ile Gly Ile Gln Asn Leu Lys Val Phe 945 950 955 960 Glu Lys Tyr Ile Val Ser Ala Leu Gly Glu Val Thr Lys Ala Glu Phe 965,970,975 Arg Gln Arg Glu Asp Phe Lys Lys 980 <210> 80 <211> 91 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 80 tataatctca taagaaattt aaaaagggac taaaataaag agtttgcggg actctgcggg 60 gttacaatcc cctaaaaccg cttttaaaat t 91 <210> 81 <211> 36 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 81 attttaccat aaagaaattt aaaaagggac taaaac 36 <210> 82 <211> 95 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, u, g, unknown or other <400> 82 nnnnnnnnnn nnnnnnnnnn guuuuagucc cgaaagggac uaaaauaaag aguuugcggg 60 acucugcggg guuacaaucc ccuaaaaccg cuuuu 95 <210> 83 <211> 69 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonuc...
Claims
1. An engineered CRISPR-Cas9 based chimeric RNA comprising the following nucleotide sequence: NNNNNNNNNNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCA.
2. An engineered CRISPR-Cas9 based chimeric RNA comprising the following nucleotide sequence: NNNNNNNNNNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAAGUG.
3. An engineered CRISPR-Cas9 based chimeric RNA comprising the following nucleotide sequence: NNNNNNNNNNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAAGUGGCACCGAGUCGGUGC.
4. 2. The engineered CRISPR-Cas9 based chimeric RNA of any of the preceding claims, wherein NNNNNNNNNNNNNNNNNNNNNNNNNN is a guide sequence capable of hybridizing to a target sequence in a eukaryotic cell adjacent to a protospacer adjacent motif (PAM).
5. 5. The engineered CRISPR-Cas9 chimeric RNA of claim 4, wherein the PAM is NGG.
6. 2. The engineered CRISPR-Cas9 based chimeric RNA of any of the preceding claims, wherein the chimeric RNA further comprises a polyU sequence.
7. 2. The engineered CRISPR-Cas9 based chimeric RNA of any preceding claim, wherein the chimeric RNA comprises one or more modified nucleotides.
8. 2. The engineered CRISPR-Cas9 system chimeric RNA of any preceding claim, wherein the chimeric RNA comprises one or more methylated nucleotides or nucleotide analogues.
9. Use of the engineered CRISPR-Cas9 chimeric RNA according to any of claims 1 to 8 for genome engineering, provided that said use does not comprise a step of modifying the genetic identity of the germline of a human, and that said use is not a method for the treatment of the human or animal body by surgery or therapy.
10. Use of the engineered CRISPR-Cas9 chimeric RNA according to any one of claims 1 to 8 in the production of a non-human transgenic animal or a transgenic plant.