Engineering of systems, methods, and optimization guide compositions for sequence manipulation
The CRISPR/Cas vector system enables efficient and scalable genome editing by using RNA-programmable CRISPR enzymes with nuclear localization sequences, addressing the need for cost-effective and versatile genome engineering.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- THE BROAD INST INC
- Filing Date
- 2024-08-20
- Publication Date
- 2026-04-17
AI Technical Summary
There is a need for novel genome engineering techniques that are inexpensive, easy to set up, and scalable, capable of targeting multiple locations within the eukaryotic genome, and that do not require customized proteins for sequence-specific targeting.
A vector system utilizing CRISPR/Cas systems, which can be programmed with a short RNA molecule to recognize specific DNA targets, comprising regulatory elements and CRISPR enzymes, including nuclear localization sequences to enhance targeting efficiency.
This approach simplifies genome editing methodologies and accelerates the classification and mapping of genetic factors associated with diverse biological functions and diseases, providing a robust and efficient means for sequence targeting.
Smart Images

Figure 0007847619000070 
Figure 0007847619000071 
Figure 0007847619000072
Abstract
Description
[Technical Field]
[0001] Incorporation by related applications and references This application claims priority to U.S. Provisional Patent Application No. 61 / 836,127, titled ENGINEERING OF SYSTEMS, METHODS AND OPTIMIZED COMPOSITIONS FOR SEQUENCE MANIPULATION, filed on 17 June 2013. This application also claims priority to U.S. Provisional Patent Applications No. 61 / 758,468; No. 61 / 769,046; No. 61 / 802,174; No. 61 / 806,375; No. 61 / 814,263; No. 61 / 819,803 and No. 61 / 828,130, each titled ENGINEERING AND OPTIMIZATION OF SYSTEMS, METHODS AND COMPOSITIONS FOR SEQUENCE MANIPULATION, filed on January 30, 2013; February 25, 2013; March 15, 2013; March 28, 2013; April 20, 2013; May 6, 2013 and May 28, 2013, respectively. Priority is also claimed to be the specification of U.S. Provisional Patent Application No. 61 / 736,527 and No. 61 / 748,427, both filed on December 12, 2012 and January 2, 2013, respectively, with the title SYSTEMS METHODS AND COMPOSITIONS FOR SEQUENCE MANIPULATION. Priority is also claimed to be the specification of U.S. Provisional Patent Application No. 61 / 791,409 and No. 61 / 835,931, both filed on March 15, 2013 and June 17, 2013, respectively, with the title BI-2011 / 008 / 44790.02.2003 and No. BI-2011 / 008 / 44790.03.2003. See also U.S. Provisional Patent Applications No. 61 / 835,936, No. 61 / 836,101, No. 61 / 836,080, No. 61 / 836,123, and No. 61 / 835,973, which were filed on June 17, 2013.
[0002] The above applications, as well as all documents cited in those applications or during their examination (the "application cited documents") and all documents cited or referenced in those application cited documents, and all documents cited or referenced in this specification (the "specification cited documents"), and all documents cited or referenced in the specification cited documents, together with any manufacturer's instructions, manuals, product specifications, and product sheets regarding any product mentioned in this specification or in any document incorporated herein by reference, are incorporated herein by reference and can be used in the practice of the present invention. More specifically, all the reference documents are incorporated by reference to the extent that each individual document is shown to be incorporated specifically by reference.
[0003] The present invention generally relates to sequence targeting using a vector system capable of using Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR) and its components, for example, systems, methods, and compositions for controlling gene expression, including genome perturbation or gene editing.
[0004] Description of research funded by the federal government The present invention was made with government support under National Institutes of Health, NIH Pioneer Award DP1MH100706. The United States government has certain rights in this invention. BACKGROUND OF THE INVENTION
[0005] Recent advances in genome sequencing technologies and analytical methods have significantly accelerated the ability to classify and map genetic factors associated with a diverse range of biological functions and diseases. Precise genome targeting techniques are necessary to enable the systematic reverse engineering of causal gene mutations by allowing selective perturbation of individual gene elements, and to advance synthetic biology, biotechnology, and pharmaceutical applications. While genome editing technologies, such as designer zinc fingers, transcriptional activator-like effectors (TALEs), or homing meganucleases, are available for producing targeted genome perturbations, there is still a need for novel genome engineering techniques that are inexpensive, easy to set up, scalable, and can easily target multiple locations within the eukaryotic genome. [Prior art documents] [Patent Documents]
[0006] [Patent Document 1] U.S. Patent No. 4,873,316 [Overview of the project] [Problems that the invention aims to solve]
[0007] There is an urgent need for alternative and robust systems and technologies for sequence targeting with diverse applications. This invention addresses this need and provides relevant advantages. CRISPR / Cas or CRISPR-Cas systems (both terms are used interchangeably throughout this application) do not require the generation of customized proteins to target specific sequences, but a single Cas enzyme can be programmed with a short RNA molecule to recognize a specific DNA target; in other words, the Cas enzyme can be recruited to a specific DNA target using the short RNA molecule. The addition of CRISPR-Cas systems to the repertoire of genome sequencing technologies and analytical methods significantly simplifies methodologies and accelerates the skills to classify and map genetic factors associated with a diverse range of biological functions and diseases. To effectively utilize CRISPR-Cas systems for genome editing without adverse effects, it is important to understand the engineering and optimization aspects of those genome engineering tools, which are aspects of the claimed invention. [Means for solving the problem]
[0008] In one embodiment, the present invention provides a vector system comprising one or more vectors. In some embodiments, the system comprises (a) a first regulatory element operably bound to one or more insertion sites for inserting a tracr mate sequence and one or more guide sequences upstream of the tracr mate sequence (the guide sequences, when expressed, direct sequence-specific binding of the CRISPR complex to a target sequence in a cell, e.g., in a eukaryotic cell, and the CRISPR complex comprises a CRISPR enzyme that complexes with (1) a guide sequence hybridized to the target sequence and (2) a tracr mate sequence hybridized to the tracr sequence); and (b) a second regulatory element operably bound to an enzyme coding sequence encoding the CRISPR enzyme, which comprises a nuclear localization sequence; components (a) and (b) are on the same or different vectors of the system. In some embodiments, component (a) further comprises a tracr sequence downstream of the tracr mate sequence under the control of the first regulatory element. In some embodiments, component (a) further comprises two or more guide sequences operably bound to a first regulatory element, each of which, when expressed, directs the CRISPR complex to sequence-specific binding to different target sequences in eukaryotic cells. In some embodiments, the system comprises a third regulatory element, e.g., a tracr sequence under the control of a polymerase III promoter. In some embodiments, the tracr sequence exhibits at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence complementarity along the length of the tracr mate sequence when optimally aligned. In some embodiments, the CRISPR complex comprises one or more nuclear localization sequences of sufficient strength to drive the accumulation of the CRISPR complex in a detectable amount in the nucleus of a eukaryotic cell. While not theoretically constrained, nuclear localization sequences are not required for CRISPR complex activity in eukaryotes, but it is thought that including such sequences enhances the system's activity, particularly with respect to targeting nucleic acid molecules in the nucleus. In some embodiments, the CRISPR enzyme is a type II CRISPR system enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme.In some embodiments, the Cas9 enzyme is Streptococcus pneumoniae, Streptococcus pyogenes, or S. thermophilus Cas9, and may include mutant Cas9 derived from these organisms. The enzyme may be a Cas9 homolog or orthologue. In some embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In some embodiments, the CRISPR enzyme is directed to cleave one or two strands at the localization of a target sequence. In some embodiments, the CRISPR enzyme lacks DNA strand cleavage activity. In some embodiments, the first regulatory element is the polymerase III promoter. In some embodiments, the second regulatory element is the polymerase II promoter. In some embodiments, the guide sequence is at least 15, 16, 17, 18, 19, 20, 25 nucleotides long, or 10–30, or 15–25, or 15–20 nucleotides long. In general, and throughout this specification, the term “vector” refers to a nucleic acid molecule capable of transporting another nucleic acid to which it is bound. Examples of vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules containing one or more free ends and not containing free ends (e.g., circular); nucleic acid molecules containing DNA, RNA, or both; and various other polynucleotides known in the art. One type of vector is a “plasmid,” which refers to a circular double-stranded DNA loop into which an additional DNA segment can be inserted, for example, by standard molecular cloning techniques. Another type of vector is a viral vector, in which a viral-derived DNA or RNA sequence is present in the vector for packaging into a virus (e.g., retroviruses, replication-deficient retroviruses, adenoviruses, replication-deficient adenoviruses, and adeno-associated viruses). Viral vectors also include polynucleotides carried by a virus for translocation into a host cell. Some vectors can self-replicate in the host cell into which they are introduced (e.g., bacterial vectors with bacterial origins of replication and episomal mammalian vectors).Other vectors (e.g., non-episomal mammalian vectors) are integrated into the host cell's genome upon introduction into the host cell and thereby replicate together with the host genome. Furthermore, some vectors may be directed to the expression of the gene to which they are operatively bound. Such vectors are referred to herein as “expression vectors.” Common expression vectors useful in recombinant DNA technology are often in the form of plasmids.
[0009] A recombinant expression vector may contain the nucleic acid of the present invention in a form suitable for expression of nucleic acid in a host cell, meaning that the recombinant expression vector may be selected based on the host cell to be used for expression and may contain one or more regulatory elements that are operatively bound to the nucleic acid sequence to be expressed. In a recombinant expression vector, "operatively bound" means that the target nucleotide sequence is bound to the regulatory element in such a way that it enables the expression of that nucleotide sequence (e.g., in an in vitro transcription / translation system or, if the vector is introduced into a host cell, in the host cell).
[0010] The term “regulatory element” includes promoters, enhancers, internal ribosome entry sites (IRESs), and other expression regulatory elements (e.g., transcription termination signals, e.g., polyadenylation signals and poly-U sequences). Such regulatory elements are described, for example, in Goeddel, Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, Calif. (1990). Regulatory elements include those that direct the constitutive expression of nucleotide sequences in many types of host cells and those that direct the expression of nucleotide sequences only in certain host cells (e.g., tissue-specific regulatory elements). Tissue-specific promoters may primarily direct expression in desired target tissues, e.g., muscle, nerve cells, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific cell types (e.g., lymphocytes). The regulatory elements may also be directed towards time-dependent expression, for example, cell cycle-dependent or developmental stage-dependent expression, which may or may not be tissue- or cell-type specific. In some embodiments, the vector includes one or more polIII promoters (e.g., 1, 2, 3, 4, 5, or more), one or more polII promoters (e.g., 1, 2, 3, 4, 5, or more), one or more polI promoters (e.g., 1, 2, 3, 4, 5, or more), or a combination thereof. Examples of polIII promoters include, but are not limited to, the U6 and H1 promoters.Examples of polII promoters include, but are not limited to, the retroviral Roussarcoma virus (RSV) LTR promoter (sometimes containing an RSV enhancer), the cytomegalovirus (CMV) promoter (sometimes containing a CMV enhancer) [see, e.g., Boshart et al, Cell, 41:521-530 (1985)], the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter. The term “regulatory element” also includes enhancer elements, such as WPRE; CMV enhancer; the R-U5' segment in the LTR of HTLV-I (Mol. Cell. Biol., Vol. 8(1), p. 466-472, 1988); the SV40 enhancer; and the intron sequence between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), p. 1527-31, 1981). It is recognized by those skilled in the art that the design of expression vectors may depend on factors such as the selection of host cells to be transformed and the desired expression level. By introducing a vector into host cells, it is possible to produce proteins or peptides, including transcripts, fusion proteins or peptides encoded by the nucleic acids described herein (e.g., clustered equispaced short-chain repeat (CRISPR) transcripts, proteins, enzymes, their mutants, their fusion proteins, etc.).
[0011] Lentiviruses and adeno-associated viruses are among the favorable vectors, and the type of such vector can also be selected for targeting specific types of cells.
[0012] In one embodiment, the present invention provides a vector comprising a regulatory element operably bound to an enzyme coding sequence encoding a CRISPR enzyme comprising one or more nuclear localization sequences. In some embodiments, the regulatory element drives the transcription of the CRISPR enzyme in a eukaryotic cell so that the CRISPR enzyme accumulates in a detectable amount in the nucleus of the eukaryotic cell. In some embodiments, the regulatory element is a polymerase II promoter. In some embodiments, the CRISPR enzyme is a type II CRISPR system enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme. In some embodiments... The Cas9 enzyme is Streptococcus pneumoniae, Streptococcus pyogenes, or S. thermophilus Cas9, and may include mutant Cas9 derived from these organisms. In some embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In some embodiments, the CRISPR enzyme is directed to cleave one or two strands at the localization of a target sequence. In some embodiments, the CRISPR enzyme lacks DNA strand cleavage activity.
[0013] In one embodiment, the present invention provides a CRISPR enzyme comprising one or more nuclear localization sequences of sufficient strength to drive the accumulation of the CRISPR enzyme in a detectable amount in the nucleus of a eukaryotic cell. In some embodiments, the CRISPR enzyme is a type II CRISPR system enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme. In some embodiments, the Cas9 enzyme is Streptococcus pneumoniae, Streptococcus pyogenes, or S. thermophilus Cas9, and may include mutant Cas9 derived from those organisms. The enzyme may be a Cas9 homolog or orthologue. In some embodiments, the CRISPR enzyme lacks the ability to cleave one or more strands of the target sequence to which it binds.
[0014] In one embodiment, the present invention provides a eukaryotic host cell comprising (a) a first regulatory element operably bound to one or more insertion sites for inserting a tracr mate sequence and one or more guide sequences upstream of the tracr mate sequence (the guide sequences, when expressed, direct sequence-specific binding of the CRISPR complex to a target sequence in a eukaryotic cell, and the CRISPR complex comprises a CRISPR enzyme that complexes with (1) a guide sequence hybridized to the target sequence and (2) a tracr mate sequence hybridized to the tracr sequence); and / or (b) a second regulatory element operably bound to an enzyme-coding sequence encoding the CRISPR enzyme, which comprises a nuclear localization sequence. In some embodiments, the host cell comprises components (a) and (b). In some embodiments, components (a), component (b), or components (a) and (b) are stably integrated into the genome of the host eukaryotic cell. In some embodiments, component (a) further comprises a tracr sequence downstream of the tracr mate sequence under the control of the first regulatory element. In some embodiments, component (a) further comprises two or more guide sequences operably bound to a first regulatory element, each of which, when expressed, directs the CRISPR complex to sequence-specific binding to different target sequences in a eukaryotic cell. In some embodiments, the eukaryotic host cell further comprises a third regulatory element, e.g., a polymerase III promoter, operably bound to the tracr sequence. In some embodiments, the tracr sequence exhibits at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence complementarity along the length of the tracr mate sequence when optimally aligned. In some embodiments, the CRISPR enzyme comprises one or more nuclear localization sequences of sufficient strength to drive the accumulation of the CRISPR enzyme in a detectable amount in the nucleus of a eukaryotic cell. In some embodiments, the CRISPR enzyme is a type II CRISPR system enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme.In some embodiments, the Cas9 enzyme is Streptococcus pneumoniae, Streptococcus pyogenes, or S. thermophilus Cas9, and may include mutant Cas9 derived from these organisms. The enzyme may be a Cas9 homolog or orthologue. In some embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In some embodiments, the CRISPR enzyme is directed to cleave one or two strands at the localization of a target sequence. In some embodiments, the CRISPR enzyme lacks DNA strand cleavage activity. In some embodiments, the first regulatory element is the polymerase III promoter. In some embodiments, the second regulatory element is the polymerase II promoter. In some embodiments, the guide sequence is at least 15, 16, 17, 18, 19, 20, 25 nucleotides long, or 10–30, or 15–25, or 15–20 nucleotides long. In one embodiment, the present invention provides a non-human eukaryote comprising a eukaryotic host cell according to any of the embodiments described; preferably, a multicellular eukaryote. In another embodiment, the present invention provides a eukaryote comprising a eukaryotic host cell according to any of the embodiments described; preferably, a multicellular eukaryote. In some embodiments of these aspects, the organism may be an animal; for example, a mammal. The organism may also be an arthropod, for example, an insect. The organism may also be a plant. Furthermore, the organism may be a fungus.
[0015] In one embodiment, the present invention provides a kit comprising one or more of the components described herein. In some embodiments, the kit comprises a vector system and instructions for use of the kit. In some embodiments, the vector system comprises (a) a first regulatory element operably bound to one or more insertion sites for inserting a tracr mate sequence and one or more guide sequences upstream of the tracr mate sequence (the guide sequences, when expressed, direct the sequence-specific binding of the CRISPR complex to a target sequence in a eukaryotic cell, and the CRISPR complex comprises a CRISPR enzyme that complexes with (1) a guide sequence hybridized to the target sequence and (2) a tracr mate sequence hybridized to the tracr sequence); and / or (b) a second regulatory element operably bound to an enzyme coding sequence encoding the CRISPR enzyme, which comprises a nuclear localization sequence. In some embodiments, the kit comprises components (a) and (b) on the same or different vectors of the system. In some embodiments, component (a) further comprises a tracr sequence downstream of the tracr mate sequence under the control of the first regulatory element. In some embodiments, component (a) further comprises two or more guide sequences operably bound to a first regulatory element, each of which, when expressed, directs the CRISPR complex to sequence-specific binding to different target sequences in eukaryotic cells. In some embodiments, the system further comprises a third regulatory element, e.g., a polymerase III promoter, operably bound to the tracr sequence. In some embodiments, the tracr sequence exhibits at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence complementarity along the length of the tracr mate sequence when optimally aligned. In some embodiments, the CRISPR enzyme comprises one or more nuclear localization sequences of sufficient strength to drive the accumulation of the CRISPR enzyme in a detectable amount in the nucleus of a eukaryotic cell. In some embodiments, the CRISPR enzyme is a type II CRISPR system enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme.In some embodiments, the Cas9 enzyme is Streptococcus pneumoniae, Streptococcus pyogenes, or S. thermophilus Cas9, and may include mutant Cas9 derived from these organisms. The enzyme may be a Cas9 homolog or orthologue. In some embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In some embodiments, the CRISPR enzyme is directed to cleave one or two strands at the localization of a target sequence. In some embodiments, the CRISPR enzyme lacks DNA strand cleavage activity. In some embodiments, the first regulatory element is the polymerase III promoter. In some embodiments, the second regulatory element is the polymerase II promoter. In some embodiments, the guide sequence is at least 15, 16, 17, 18, 19, 20, 25 nucleotides long, or 10–30, or 15–25, or 15–20 nucleotides long.
[0016] In one embodiment, the present invention provides a method for modifying a target polynucleotide in a eukaryotic cell. In some embodiments, the method comprises binding a CRISPR complex to a target polynucleotide to cause cleavage of the target polynucleotide, thereby modifying the target polynucleotide, the CRISPR complex comprising a CRISPR enzyme complexing with a guide sequence which hybridizes to a target sequence within the target polynucleotide, the guide sequence being bound to a tracr mate sequence which then hybridizes to a tracr sequence. In some embodiments, the cleavage comprises the CRISPR enzyme cleaving one or two strands at the localization of the target sequence. In some embodiments, the cleavage results in a reduction in the transcription of the target gene. In some embodiments, the method further comprises repairing the cleaved target polynucleotide by homologous recombination with an exogenous template polynucleotide, the repair resulting in a mutation including the insertion, deletion, or substitution of one or more nucleotides in the target polynucleotide. In some embodiments, the mutation results in a change of one or more amino acids in a protein expressed from the gene containing the target sequence. In some embodiments, the method further comprises delivering one or more vectors to the eukaryotic cells, the one or more vectors driving the expression of one or more CRISPR enzymes, guide sequences bound to tracr mate sequences, and tracr sequences. In some embodiments, the vectors are delivered into the eukaryotic cells in the subject. In some embodiments, the modification is carried out in the eukaryotic cells in a cell culture. In some embodiments, the method further comprises isolating the eukaryotic cells from the subject before the modification. In some embodiments, the method further comprises returning the eukaryotic cells and / or cells derived therefrom to the subject.
[0017] In one embodiment, the present invention provides a method for modifying the expression of a polynucleotide in a eukaryotic cell. In some embodiments, the method comprises conjugating a CRISPR complex to a polynucleotide, such that the conjugation results in an increase or decrease in the expression of the polynucleotide; the CRISPR complex comprises a CRISPR enzyme complexing with a guide sequence which hybridizes to a target sequence within the polynucleotide, the guide sequence being bound to a tracr mate sequence which then hybridizes to a tracr sequence. In some embodiments, the method further comprises delivering one or more vectors to the eukaryotic cell, the one or more vectors driving the expression of a CRISPR enzyme, a guide sequence bound to a tracr mate sequence, and one or more tracr sequences.
[0018] In one embodiment, the present invention provides a method for generating a model eukaryotic cell containing a mutated disease gene. In some embodiments, the disease gene is any gene associated with having or having an increased risk of developing a disease. In some embodiments, the method comprises (a) introducing one or more vectors into a eukaryotic cell (one or more vectors driving the expression of a CRISPR enzyme, a guide sequence bound to a tracr mate sequence, and one or more tracr sequences) and (b) binding the CRISPR complex to a target polynucleotide to cause cleavage of the target polynucleotide in the disease gene (the CRISPR complex comprises a CRISPR enzyme complexing with (1) a guide sequence hybridized to a target sequence in the target polynucleotide, and (2) a tracr mate sequence hybridized to a tracr sequence), thereby generating a model eukaryotic cell containing a mutated disease gene. In some embodiments, the cleavage comprises the CRISPR enzyme cleaving one or two strands at the localization of the target sequence. In some embodiments, the cleavage results in a reduction in the transcription of the target gene. In some embodiments, the method further comprises repairing the cleavage target polynucleotide by homologous recombination with an exogenous template polynucleotide, wherein the repair results in a mutation including the insertion, deletion, or substitution of one or more nucleotides in the target polynucleotide. In some embodiments, the mutation results in a change of one or more amino acids in a protein expressed from a gene containing the target sequence.
[0019] In one embodiment, the present invention provides a method for developing a bioactive agent that modulates cellular signaling events associated with a disease gene. In some embodiments, the disease gene is any gene associated with an increased risk of having or developing a disease. In some embodiments, the method comprises (a) contacting a test compound with a model cell of any one of the embodiments described; and (b) detecting a readout change indicating a reduction or increase in cellular signaling events associated with the mutation in the disease gene, thereby developing the bioactive agent that modulates the cellular signaling events associated with the disease gene.
[0020] In one embodiment, the present invention provides a recombinant polynucleotide comprising a guide sequence upstream of a tracr mate sequence, wherein the guide sequence, when expressed, directs the CRISPR complex to sequence-specific binding to a corresponding target sequence present in eukaryotic cells. In some embodiments, the target sequence is a viral sequence present in eukaryotic cells. In some embodiments, the target sequence is a proto-oncogene or oncogene.
[0021] In one embodiment, the present invention provides a method for selecting one or more prokaryotic cells by introducing one or more mutations in the genes of one or more prokaryotic cells, comprising: introducing one or more vectors into the prokaryotic cells (one or more vectors driving the expression of one or more CRISPR enzymes, guide sequences bound to tracr mate sequences, tracr sequences, and editing templates; the editing templates include one or more mutations that halt CRISPR enzyme cleavage); homologous recombination of the editing templates with target polynucleotides in the cells to be selected; and binding the CRISPR complex to the target polynucleotide to cause cleavage of the target polynucleotide in the gene (the CRISPR complex comprises a CRISPR enzyme complexing with (1) a guide sequence hybridized to a target sequence in the target polynucleotide, and (2) a tracr mate sequence hybridized to a tracr sequence, the binding of the CRISPR complex to the target polynucleotide induces cell death), thereby enabling the selection of one or more prokaryotic cells into which one or more mutations have been introduced. In a preferred embodiment, the CRISPR enzyme is Cas9. In another embodiment of the present invention, the cells to be selected may be eukaryotic cells. According to aspects of the present invention, the selection of a specified set of cells is possible without requiring a selection marker or a two-step process that may include a counter-selection system.
[0022] In some embodiments, the present invention comprises a CRISPR-Cas chimeric RNA (chiRNA) polynucleotide sequence, the polynucleotide sequence comprising (a) a guide sequence capable of hybridizing to a target sequence in a eukaryotic cell, (b) a tracr mate sequence, and (c) a tracr sequence, wherein (a), (b), and (c) are arranged in a 5' to 3' orientation and, upon transcription, hybridize to a tracr mate sequence vtracr sequence, and the guide sequence directs the sequence-specific binding of the CRISPR complex to the target sequence, and the CRISPR complex comprises a CRISPR enzyme that complexes with (1) a guide sequence that hybridizes to the target sequence and (2) a tracr mate sequence that hybridizes to the tracr sequence, and is not naturally occurring or an engineered composition. or I. A first regulatory element operably bound to a CRISPR-Cas system chimeric RNA (chiRNA) polynucleotide sequence (the polynucleotide sequence includes (a) one or more guide sequences that can hybridize to one or more target sequences in eukaryotic cells, (b) a tracr mate sequence, and (c) one or more tracr sequences), and II. A second regulatory element operably bound to an enzyme coding sequence encoding a CRISPR enzyme that includes at least one nuclear localization sequence, wherein (a), (b), and (c) are 5' to 3' The components I and II are arranged in a directional manner, and upon transcription, the tracr mate sequence hybridizes to the tracr sequence, and the guide sequence directs the sequence-specific binding of the CRISPR complex to the target sequence, and the CRISPR complex is encoded by a vector system comprising one or more vectors containing a CRISPR enzyme that forms a complex with (1) a guide sequence that hybridizes to the target sequence, and (2) a tracr mate sequence that hybridizes to the tracr sequence, or I.(a) cells The present invention provides a multiplexed CRISPR enzyme system encoded by a vector system comprising one or more vectors, each of which has been modified to improve stability, wherein the multiplexed system includes (b) a first regulatory element operably bound to at least one tracr mate sequence, (c) a second regulatory element operably bound to an enzyme coding sequence encoding a CRISPR enzyme, and (d) a third regulatory element operably bound to a tracr sequence, wherein components I, II, and III are on the same or different vectors of the system, and upon transcription, the tracr mate sequence hybridizes to a tracr sequence, and the guide sequence directs the CRISPR complex to sequence-specific binding to the target sequence, and the CRISPR complex comprises (1) a guide sequence hybridized to a target sequence, and (2) a CRISPR enzyme complexing with a tracr mate sequence hybridized to a tracr sequence, and in the multiplexed system, one or more vectors of which have been modified to improve stability, each of which has been modified to improve stability, and which comprises one or more vectors of which have been modified to improve stability, and which have been modified to improve stability, and which comprises one or more guide, tracr, and tracr mate sequences, and which comprises one or more vectors, in which multiple guide sequences and a single tracr sequence are used.
[0023] In aspects of the invention, the modification includes an engineered secondary structure. For example, the modification can include a reduction in the hybridization region between the tracr mate sequence and the tracr sequence. For example, the modification can also include a fusion of the tracr mate sequence and the tracr sequence through an artificial loop. The modification can include a tracr sequence having a length of 40-120 bp. In embodiments of the invention, the tracr sequence is from 40 bp to the full-length tracr. In certain embodiments, the length of the tracrRNA includes at least nucleotides 1-67 of the wild-type tracrRNA, and in some embodiments, includes at least nucleotides 1-85. In some embodiments, nucleotides corresponding to at least nucleotides 1-67 or 1-85 of the wild-type Streptococcus pyogenes (S. pyogenes) Cas9 tracrRNA can be used. When the CRISPR system uses an enzyme other than Cas9 or other than SpCas9, corresponding nucleotides in the relevant wild-type tracrRNA may be present. In some embodiments, the length of the tracrRNA includes nucleotides 1-67 or less than 1-85 of the wild-type tracrRNA. The modification can include sequence optimization. In one aspect, the sequence optimization can include a reduction in the occurrence rate of poly-T sequences in the tracr and / or tracr mate sequence The sequence optimization can include a reduction in the hybridization region between the tracr mate sequence and the tracr sequence; for example, it can be combined with a reduction in the length of the tracr sequence.
[0024] In one aspect, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification comprises a reduction of the poly-T sequence in the tracr and / or tracr mate sequences. In some aspects of the present invention, one or more Ts present in the poly-T sequence of the relevant wild-type sequence (i.e., a stretch of 3, 4, 5, 6, or more consecutive T bases; in some embodiments, a stretch of 10, 9, 8, 7, 6 or fewer consecutive T bases) can be replaced by non-T nucleotides, such as A, and thus the string is broken down into smaller stretches of Ts, each stretch having 4 or fewer (e.g., 3 or 2) consecutive Ts. Bases other than A, such as C or G, or non-natural or modified nucleotides can be used for the substitution. When a string of Ts is involved in the formation of a hairpin (or stem-loop), it is advantageous to change the complementary base for the non-T base to a non-T nucleotide in the complementary strand. For example, if the non-T base is A, its complementary strand can be changed to T, for example, to preserve or assist in preserving the secondary structure. For example, 5'-TTTTT can be changed to 5'-TTTAT, and the complementary 5'-AAAAA can be changed to 5'-ATAAA.
[0025] In one aspect, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification comprises adding a poly-T terminator sequence. In one aspect, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification comprises adding a poly-T terminator sequence in the tracr and / or tracr mate sequences. In one aspect, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification comprises adding a poly-T terminator sequence in the guide sequence. The poly-T terminator sequence can comprise five or more consecutive T bases.
[0026] In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification includes altering loops and / or hairpins. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification includes providing a minimum of two hairpins in the guide sequence. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification includes providing hairpins formed by complementarity between tracr and tracrmate (direct repeat) sequences. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification includes providing one or more additional hairpins at or toward the 3' end of a tracrRNA sequence. For example, hairpins can be formed by providing self-complementary sequences within the tracrRNA sequence that are linked by loops so that the hairpins are formed during self-folding. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the modification includes providing additional hairpins that are added to the 3' of the guide sequence. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which a modification includes extending the 5' end of a guide sequence. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which a modification includes providing one or more hairpins in the 5' end of a guide sequence. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which a modification includes adding the sequence (5'-AGGACGAAGTCCTAA) to the 5' end of a guide sequence. Other sequences suitable for hairpin formation are known to those skilled in the art and can be used in some embodiments of the present invention. In some embodiments of the present invention, at least two, three, four, five, or more additional hairpins are provided. In some embodiments of the present invention, 10, nine, eight, seven, six or fewer additional hairpins are provided. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which a modification includes two hairpins. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which a modification includes three hairpins.In one embodiment, the present invention provides a modified CRISPR-Cas system or CRISPR enzyme system comprising at most five hairpins.
[0027] In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which a modification provides a crosslink or provides one or more modified nucleotides in a polynucleotide sequence. The modified nucleotides and / or crosslinks can be provided in any or all of the tracr, tracr mate, and / or guide sequences, and / or in the enzyme coding sequence, and / or in the vector sequence. The modification may include the inclusion of at least one nucleotide that does not exist in nature, or a modified nucleotide, or an analog thereof. The modified nucleotide may be modified in the ribose, phosphate, and / or base moiety. Examples of modified nucleotides include 2'-O-methyl analogs, 2'-deoxy analogs, or 2'-fluoro analogs. The nucleic acid backbone can be modified, for example, a phosphorothioate backbone can be used. The use of locked nucleic acid (LNA) or crosslinked nucleic acid (BNA) may also be possible. Further examples of modified bases, but not limited to, include 2-aminopurine, 5-bromo-uridine, pseudouridine, inosine, and 7-methylguanosine.
[0028] It is understood that any or all of the above modifications can be provided individually or in combination within a given CRISPR-Cas system or CRISPR enzyme system. Such a system may contain one, two, three, four, five, or more of the above modifications.
[0029] In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the CRISPR enzyme is a type II CRISPR system enzyme, such as a Cas9 enzyme. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the CRISPR enzyme contains fewer than 1000 amino acids or fewer than 4000 amino acids. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the Cas9 enzyme is StCas9 or StlCas9, or the Cas9 enzyme is a Cas9 enzyme from an organism selected from the group consisting of the genera Streptococcus, Campylobacter, Nitratifractor, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, Sphaerochaeta, Lactobacillus, Eubacterium, or Corynebacter. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the CRISPR enzyme is a nuclease directed toward the cleavage of both chains at the localization of a target sequence.
[0030] In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the first regulatory element is a polymerase III promoter. In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the second regulatory element is a polymerase II promoter.
[0031] In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the guide sequence comprises at least 15 nucleotides.
[0032] In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which a modification comprises an optimized tracr sequence and / or an optimized guide sequence RNA and / or a co-fold structure of the tracr sequence and / or a tracr mate sequence and / or a stabilizing secondary structure of the tracr sequence and / or a tracr sequence fusion RNA element having a region with reduced base pairing; and / or a multiplexed system comprising a tracer and having two RNAs containing multiple guides or one RNA containing multiple chimeras.
[0033] In embodiments of the present invention, the chimeric RNA architecture is further optimized according to the results of mutagenesis testing. In chimeric RNA having two or more hairpins, mutations in the proximal direct repeats for hairpin stabilization may result in termination of CRISPR complex activity. Mutations in the distal direct repeats for shortening or stabilizing hairpins may have no effect on CRISPR complex activity. Sequence randomization in the bulge region between the proximal and distal repeats may significantly reduce CRISPR complex activity. A single base pair change or sequence randomization in the linker region between hairpins may result in complete loss of CRISPR complex activity. Hairpin stabilization of the distal hairpin following the first hairpin after the guide sequence may result in maintenance or improvement of CRISPR complex activity. Therefore, in preferred embodiments of the present invention, the chimeric RNA architecture can be further optimized by generating smaller chimeric RNAs that may be beneficial for optional therapeutic delivery and other uses, which can be achieved by modifying the distal direct repeats to shorten or stabilize hairpins. In a more preferred embodiment of the present invention, the chimeric RNA architecture can be further optimized by stabilizing one or more distal hairpins. Stabilization of hairpins may include modification of sequences suitable for hairpin formation. In some embodiments of the present invention, at least two, three, four, five, or more additional hairpins are provided. In some embodiments of the present invention, 10, nine, eight, seven, six, or fewer additional hairpins are provided. In some embodiments of the present invention, stabilization may be crosslinking and other modifications. Modifications may include the inclusion of at least one non-naturally occurring nucleotide, or a modified nucleotide, or an analog thereof. Modified nucleotides may be modified in the ribose, phosphate, and / or base moieties. Examples of modified nucleotides include 2'-O-methyl analogs, 2'-deoxy analogs, or 2'-fluoro analogs. The nucleic acid backbone can be modified, for example, a phosphorothioate backbone can be used.The use of locked nucleic acids (LNA) or cross-linked nucleic acids (BNA) may also be possible. Further examples of modified bases include, but are not limited to, 2-aminopurine, 5-bromo-uridine, pseudouridine, inosine, and 7-methylguanosine.
[0034] In one embodiment, the present invention provides a CRISPR-Cas system or CRISPR enzyme system in which the CRISPR enzyme is codon-optimized for expression in eukaryotic cells.
[0035] Therefore, in some embodiments of the present invention, the required tracRNA length in the construct of the present invention, for example, a chimeric construct, does not necessarily have to be fixed, and in some embodiments of the present invention, it may be 40 to 120 bp, and in some embodiments of the present invention, it may be at most the full length of the tracr, for example, to the 3' end of the tracr interrupted by a transcription termination signal in the bacterial genome. In some embodiments, the tracRNA length includes at least nucleotides 1 to 67 of wild-type tracRNA, and in some embodiments, it includes at least nucleotides 1 to 85. In some embodiments, nucleotides corresponding to at least nucleotides 1 to 67 or 1 to 85 of wild-type Streptococcus pyogenes (S. pyogenes) Cas9 tracRNA can be used. If the CRISPR system uses an enzyme other than Cas9 or SpCas9, corresponding nucleotides may be present in the relevant wild-type tracRNA. In some embodiments, the tracRNA length includes nucleotides 1 to 67 or 1 to 85 or less of wild-type tracRNA. With regard to sequence optimization (e.g., reduction of poly-T sequences), for example with respect to tracrmates (direct repeats) or strings of T within tracrRNA, in some embodiments of the present invention, one or more Ts present in the poly-T sequence of the relevant wild-type sequence (i.e., stretches of more than 3, 4, 5, 6, or more consecutive T bases; in some embodiments, stretches of 10, 9, 8, 7, or 6 consecutive T bases) can be substituted with non-T nucleotides, e.g., A, so that the string is broken down into smaller stretches of T, each stretch having 4, or fewer than 4 (e.g., 3 or 2) consecutive Ts. If the string of T is involved in the formation of a hairpin (or stem-loop), it is advantageous to change the complementary base for the non-T base to the complementary strand of the non-T nucleotide. For example, if the non-T base is A, its complementary strand can be changed to T, for example, to preserve or assist in the preservation of secondary structure. For example, 5'-TTTTT can be changed to 5'-TTTAT, and the complementary 5'-AAAAA can be changed to 5'-ATAAA.With regard to the presence of poly-T terminator sequences in the tracr+tracrmate transcript, e.g., poly-T terminators (TTTTT or more Ts), in some embodiments of the present invention, it is advantageous to add such sequences to the end of the transcript, whether in the form of two RNAs (tracr and tracrmate) or a single guide RNA. With regard to loops and hairpins in the tracr and tracrmate transcript, in some embodiments of the present invention, it is advantageous to have a minimum of two hairpins present in the chimeric guide RNA. The first hairpin may be a hairpin formed by complementarity between the tracr and tracrmate (direct repeat) sequences. The second hairpin may be at the 3' end of the tracrRNA sequence, which may provide a secondary structure for interaction with Cas9. Additional hairpins may be added to the 3' end of the guide RNA, for example, in some embodiments of the present invention, which can increase the stability of the guide RNA. Furthermore, the 5' end of the guide RNA may be elongated in some embodiments of the present invention. In some embodiments of the present invention, 20 bp in the 5' end may be considered a guide sequence. The 5' portion can be extended. One or more hairpins can be provided in the 5' portion, and in some embodiments of the present invention, for example, this may also improve the stability of the guide RNA. In some embodiments of the present invention, a specified hairpin can be provided by adding the sequence (5'-AGGACGAAGTCCTAA) to the 5' end of the guide sequence, and in some embodiments of the present invention, this may help improve stability. Other sequences suitable for hairpin formation are known to those skilled in the art and can be used in some embodiments of the present invention. In some embodiments of the present invention, at least two, three, four, five, or more additional hairpins are provided. In some embodiments of the present invention, 10, nine, eight, seven, six or fewer additional hairpins are provided. The above also provide embodiments of the present invention that include secondary structures in the guide sequence. In some embodiments of the present invention, for example, crosslinking and other modifications may exist to improve stability. Modifications may include the inclusion of at least one non-naturally occurring nucleotide, or modified nucleotide, or analogs thereof.Modified nucleotides may be modified in the ribose, phosphate, and / or base portions. Examples of modified nucleotides include 2'-O-methyl analogs, 2'-deoxy analogs, or 2'-fluoro analogs. The nucleic acid backbone can be modified; for example, a phosphorothioate backbone can be used. The use of locked nucleic acids (LNAs) or cross-linked nucleic acids (BNAs) may also be possible. Further examples of modified bases, but not limited to, include 2-aminopurines, 5-bromouridine, pseudouridine, inosine, and 7-methylguanosine. Such modifications or cross-links may be present in the guide sequence or other sequences adjacent to the guide sequence.
[0036] Therefore, the object of the present invention is not to include any already known product, method of producing such product, or method of using such product in the present invention, and therefore the applicants reserve the right to relinquish any already known product, method of producing such product, or method of using such product, and disclose herein this specification. The present invention does not include in the scope any product, method, or method of producing or using such product that does not satisfy the description and enablement requirements of the USPTO (Section 112 of the United States Patent Act) or the EPO (Section 83 of the European Patent Convention), and therefore the applicants reserve the right to relinquish any already described product, method of producing such product, or method of using such product, and disclose herein this specification this specification.
[0037] In this disclosure and in particular in the claims and / or paragraphs, terms such as “comprises,” “comprised,” and “comprising” may have meanings pursuant to U.S. patent law; for example, they may mean “includes,” “included,” and “including”; terms such as “consisting essentially of” and “consists essentially of” also have meanings pursuant to U.S. patent law, for example, they allow for components not expressly described, but it should be noted that they exclude components found in the prior art or that affect the basic or novel features of the present invention. These and other embodiments are disclosed or evident therefrom and encompassed by the following detailed description.
[0038] Novel features of the present invention are described in particular by the appended claims. A better understanding of the features and advantages of the present invention can be obtained by referring to the following detailed description, which describes explanatory embodiments in which the principles of the present invention are utilized, and to the accompanying drawings. [Brief explanation of the drawing]
[0039] [Figure 1] A schematic model of the CRISPR system is shown. Cas9 nuclease (yellow) from Streptococcus pyogenes is targeted to genomic DNA by a synthetic guide RNA (sgRNA) consisting of a 20nt guide sequence (blue) and a scaffold (red). The guide sequence base pair with the DNA target (blue) immediately upstream of the required 5'-NGG protospacer flanking motif (PAM; magenta), and Cas9, mediate a double-strand break (DSB) (red triangle) approximately 3bp upstream of the PAM. [Figure 2A] This paper presents exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations of expression in eukaryotic cells, and results of tests evaluating nuclear localization and CRISPR activity. [Figure 2B]This paper presents exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations of expression in eukaryotic cells, and results of tests evaluating nuclear localization and CRISPR activity. [Figure 2C] This paper presents exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations of expression in eukaryotic cells, and results of tests evaluating nuclear localization and CRISPR activity. [Figure 2D] This paper presents exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations of expression in eukaryotic cells, and results of tests evaluating nuclear localization and CRISPR activity. [Figure 2E] This paper presents exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations of expression in eukaryotic cells, and results of tests evaluating nuclear localization and CRISPR activity. [Figure 2F] This paper presents exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations of expression in eukaryotic cells, and results of tests evaluating nuclear localization and CRISPR activity. [Figure 3] This paper shows exemplary expression cassettes for the expression of CRISPR system elements in eukaryotic cells, predictive structures of exemplary guide sequences, and CRISPR system activity measured in eukaryotic and prokaryotic cells. [Figure 4A] The results of the evaluation of SpCas9 specificity for exemplary targets are shown. [Figure 4B] The results of the evaluation of SpCas9 specificity for exemplary targets are shown. [Figure 4C] The results of the evaluation of SpCas9 specificity for exemplary targets are shown. [Figure 4D] The results of the evaluation of SpCas9 specificity for exemplary targets are shown. [Figure 5A] The results of the exemplary vector system and its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5B] The results of the exemplary vector system and its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5C] The results of the exemplary vector system and its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5D] The results of the exemplary vector system and its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5E] The results of the exemplary vector system and its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5F] The results of the exemplary vector system and its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5G] The results of the exemplary vector system and its use in directing homologous recombination in eukaryotic cells are presented. [Figure 6A] This paper describes a comparison of different tracrRNA transcripts regarding Cas9-mediated gene targeting. [Figure 6B] This paper describes a comparison of different tracrRNA transcripts regarding Cas9-mediated gene targeting. [Figure 6C] This paper describes a comparison of different tracrRNA transcripts regarding Cas9-mediated gene targeting. [Figure 7A] This paper describes exemplary CRISPR systems, exemplary adaptations for expression in eukaryotic cells, and the results of tests evaluating CRISPR activity. [Figure 7B] This paper describes exemplary CRISPR systems, exemplary adaptations for expression in eukaryotic cells, and the results of tests evaluating CRISPR activity. [Figure 7C] This paper describes exemplary CRISPR systems, exemplary adaptations for expression in eukaryotic cells, and the results of tests evaluating CRISPR activity. [Figure 7D] This paper describes exemplary CRISPR systems, exemplary adaptations for expression in eukaryotic cells, and the results of tests evaluating CRISPR activity. [Figure 8] This paper describes exemplary operations of the CRISPR system for targeting genomic loci in mammalian cells. [Figure 9] This paper describes the results of Northern blot analysis of crRNA processing in mammalian cells. [Figure 10A]This paper describes the schematic representation of chimeric RNA and the results of the SURVEYOR assay for CRISPR system activity in eukaryotic cells. [Figure 10B] This paper describes the schematic representation of chimeric RNA and the results of the SURVEYOR assay for CRISPR system activity in eukaryotic cells. [Figure 10C] This paper describes the schematic representation of chimeric RNA and the results of the SURVEYOR assay for CRISPR system activity in eukaryotic cells. [Figure 11] This section explains the graphical representation of the results of the SURVEYOR assay for CRISPR system activity in eukaryotic cells. [Figure 12] This paper describes the predicted secondary structure of an exemplary chimeric RNA containing a guide sequence, a tracr mate sequence, and a tracr sequence. [Figure 13A] This is a phylogenetic tree of the Cas gene. [Figure 13B] This is a phylogenetic tree of the Cas gene. [Figure 13C] This is a phylogenetic tree of the Cas gene. [Figure 13D] This is a phylogenetic tree of the Cas gene. [Figure 14A] This paper presents a phylogenetic analysis that reveals five families of Cas9, including three groups of large Cas9 (approximately 1400 amino acids) and two groups of small Cas9 (approximately 1100 amino acids). [Figure 14B] This paper presents a phylogenetic analysis that reveals five families of Cas9, including three groups of large Cas9 (approximately 1400 amino acids) and two groups of small Cas9 (approximately 1100 amino acids). [Figure 14C] This paper presents a phylogenetic analysis that reveals five families of Cas9, including three groups of large Cas9 (approximately 1400 amino acids) and two groups of small Cas9 (approximately 1100 amino acids). [Figure 14D] This paper presents a phylogenetic analysis that reveals five families of Cas9, including three groups of large Cas9 (approximately 1400 amino acids) and two groups of small Cas9 (approximately 1100 amino acids). [Figure 14E]This paper presents a phylogenetic analysis that reveals five families of Cas9, including three groups of large Cas9 (approximately 1400 amino acids) and two groups of small Cas9 (approximately 1100 amino acids). [Figure 14F] This paper presents a phylogenetic analysis that reveals five families of Cas9, including three groups of large Cas9 (approximately 1400 amino acids) and two groups of small Cas9 (approximately 1100 amino acids). [Figure 15] The graph shows the functions of different optimization guide RNAs. [Figure 16] The sequences and structures of different guide chimeric RNAs are shown. [Figure 17] This shows the simultaneous fold structure of tracrRNA and direct repeats. [Figure 18A] Data from in vitro St1Cas9 chimeric guide RNA optimization are shown. [Figure 18B] Data from in vitro St1Cas9 chimeric guide RNA optimization are shown. [Figure 19] SpCas9 cell lysates exhibit either demethylation or cleavage of methylation targets. [Figure 20]This paper demonstrates the optimization of guide RNA architectures for SpCas9-mediated mammalian genome editing. (a) Schematic diagram of the bicistronic expression vector (PX330) for human codon-optimized Streptococcus pyogenes Cas9 (hSpCas9) driven by the U6 promoter and the CBh promoter, used in all subsequent experiments. The sgRNA consists of a 20nt guide sequence (blue) and scaffold (red) truncated at various positions shown. (b) SURVEYOR assay for SpCas9-mediated indels at the human EMX1 and PVALB gene loci. Arrows indicate predicted SURVEYOR fragments (n=3). (c) Northern blot analysis of four sgRNA truncation architectures using U1 as a loading control. (d) Both wild-type (wt) and nickase mutant (D10A) SpCas9 promoted insertion of the HindIII site into the human EMX1 gene. Single-stranded oligonucleotides (ssODNs) were oriented either sense or antisense relative to the genome sequence and used as homologous recombination templates. (e) Schematic diagram of the human SERPINB5 locus. sgRNA and PAM are shown by colored bars above the sequence; methylcytosine (Me) is highlighted (pink) and numbered relative to the transcription start site (TSS, +1). (f) Methylation status of SERPINB5 assayed by bisulfite sequencing of 16 clones. Black circles indicate methylated CpG; white circles indicate unmethylated CpG. (g) Modification efficiency (n=2) of SERPINB5 assayed by deep sequencing using three sgRNAs targeting the methylated region (SERPINB5). Error bars indicate Wilson intervals (online method). [Figure 21]Further optimization of the CRISPR-Cas sgRNA architecture is demonstrated. (a) Schematic diagrams of four additional sgRNA architectures I-IV. Each consists of a 20nt guide sequence (blue) bound to a direct repeat (DR, gray) that hybridizes to tracrRNA (red). The DR-tracrRNA hybrid is truncated at +12 or +22 as shown and has an artificial GAAA stem-loop. The tracrRNA truncation sites are numbered according to the transcription start sites for previously reported tracrRNAs. sgRNA architectures II and IV carry mutations within their poly-U tracts that can function as early transcription terminators. (b) SURVEYOR assay for SpCas9-mediated indels at the human EMX1 locus for target sites 1-3. Arrows indicate predicted SURVEYOR fragments (n=3). [Figure 22] This explains the visualization of certain target regions within the human genome. [Figure 23] (A) A schematic diagram of sgRNA and (B) SURVEYOR analysis of five sgRNA variants for SaCas9 regarding the optimal truncation architecture with maximum cleavage efficiency are shown. [Modes for carrying out the invention]
[0040] The drawings in this specification are for illustrative purposes only and are not necessarily drawn to a specific scale.
[0041] The terms “polynucleotide,” “nucleotide,” “nucleotide sequence,” “nucleic acid,” and “oligonucleotide” are used interchangeably. These refer to polymeric forms of nucleotides, deoxyribonucleotides, or ribonucleotides of any length, or analogs thereof. Polynucleotides can have any three-dimensional structure and can perform any known or unknown function. The following are non-limiting examples of polynucleotides: coding or non-coding regions of genes or gene fragments, loci (gene loci) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Polynucleotides may contain one or more modified nucleotides, e.g., methylated nucleotides or nucleotide analogs. Modifications to the nucleotide structure, if present, can be given before or after the assembly of the polymer. The sequence of nucleotides can be interrupted by non-nucleotide components. Polynucleotides can be further modified after polymerization, for example, by conjugation with a labeling component.
[0042] In aspects of the present invention, the terms “chimeric RNA,” “chimeric guide RNA,” “guide RNA,” “single guide RNA,” and “synthetic guide RNA” are used interchangeably and refer to a polynucleotide sequence comprising a guide sequence, a tracr sequence, and a tracr mate sequence. The term “guide sequence” refers to a sequence of approximately 20 bp within the guide RNA that defines the target site and can be used interchangeably with the terms “guide” or “spacer.” The term “tracr mate sequence” can also be used interchangeably with the term “direct repeat.”
[0043] As used herein, the term “wild type” is a term understood by those skilled in the art and means a typical form of an organism, strain, gene, or characteristic as it occurs in nature, distinguished from mutant or variant forms.
[0044] As used herein, the term "variant" should be interpreted as meaning a presentation of quality that has a pattern that deviates from that which occurs in the natural state.
[0045] The terms “not naturally occurring” or “engineered” are used interchangeably and indicate artificial involvement. When these terms refer to nucleic acid molecules or polypeptides, they mean that the nucleic acid molecules or polypeptides do not contain, at least substantially, at least one other component that they naturally associate with or are found in nature.
[0046] "Complementarity" refers to the ability of a nucleic acid to form hydrogen bonds with another nucleic acid sequence, either through classical Watson-Crick base pairing or other non-classical types. The complementarity percentage indicates the proportion of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (for example, 5, 6, 7, 8, 9, and 10 out of 10 have 50%, 60%, 70%, 80%, 90%, and 100% complementarity, respectively). "Perfectly complementary" means that all consecutive residues in a nucleic acid sequence can form hydrogen bonds with the same number of consecutive residues in the second nucleic acid sequence. As used herein, “substantially complementary” means a degree of complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% of a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or two nucleic acids that hybridize under stringent conditions.
[0047] As used herein, “stringent conditions” for hybridization refer to conditions under which a nucleic acid complementary to the target sequence predominantly hybridizes with the target sequence and substantially does not hybridize with the non-target sequence. Stringent conditions are generally sequence-dependent and vary depending on numerous factors. Generally, longer sequences have higher temperatures at which they specifically hybridize with their target sequence. Non-limiting examples of stringent conditions are detailed in Tijssen (1993), Laboratory Techniques In Biochemistry And Molecular Biology—Hybridization With Nucleic Acid Probes Part I, Second Chapter “Overview of principles of hybridization and the strategy of nucleic acid probe assay”, Elsevier, NY.
[0048] Hybridization refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized via hydrogen bonds between the bases of nucleotide residues. Hydrogen bonds can occur through Watson-Crick base pairing, Hoogsteen bonding, or any other sequence-specific manner. The complex may consist of two strands forming a double-stranded structure, three or more strands forming a multi-stranded complex, a single self-hybriding strand, or any combination thereof. Hybridization reactions may constitute a step in a broader process, such as the initiation of PCR or enzymatic cleavage of polynucleotides. A sequence that can hybridize with a given sequence is referred to as a "complementary strand" of the given sequence.
[0049] As used herein with respect to components of the CRISPR system, “stabilization” or “increasing stability” refers to ensuring or stabilizing the molecular structure. This can be achieved by introducing one or more mutations, including single or multiple base pair changes, an increase in hairpin number, crosslinking, disruption of specific stretches of nucleotides, and other modifications. Modifications may include the inclusion of at least one nucleotide that does not exist in nature, or a modified nucleotide, or an analog thereof. Modified nucleotides may be modified in the ribose, phosphate, and / or base moieties. Examples of modified nucleotides include 2'-O-methyl analogs, 2'-deoxy analogs, or 2'-fluoro analogs. The nucleic acid backbone can be modified, for example, a phosphorothioate backbone can be used. The use of locked nucleic acid (LNA) or crosslinked nucleic acid (BNA) may also be possible. Further examples of modified bases, but not limited to, include 2-aminopurine, 5-bromo-uridine, pseudouridine, inosine, and 7-methylguanosine. These modifications may apply to any component of the CRISPR system. In preferred embodiments, these modifications are made to RNA components, such as guide RNA or chimeric polynucleotide sequences.
[0050] As used herein, “expression” refers to the process by which polynucleotides are transcribed from a DNA template (e.g., into mRNA or other RNA transcripts) and / or the process by which the transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and the polypeptides they encode can be collectively referred to as “gene products.” If the polynucleotides originate from genomic DNA, expression may include splicing of mRNA in eukaryotic cells.
[0051] The terms “polypeptide,” “peptide,” and “protein” are used interchangeably herein to refer to polymers of amino acids of any length. The polymers may be linear or branched, and may contain modified amino acids, which may be interrupted by non-amino acids. The term also encompasses amino acid polymers that have undergone modifications, such as disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other operation, such as conjugation with a labeling component. The term “amino acid” as used herein includes natural and / or unnatural or synthetic amino acids, including glycine and both D or L optical isomers, as well as amino acid analogs and peptide mimetic compounds.
[0052] The terms “subject,” “individual,” and “patient” are used interchangeably herein to refer to vertebrates, preferably mammals, and more preferably humans. Mammals include, but are not limited to, mice, monkeys, humans, domestic animals, athletic animals, and pets. Tissues, cells, and their offspring of biological entities obtained in vivo or cultured in vitro are also included. In some embodiments, the subject may be an invertebrate, such as an insect or nematode; on the other hand, the subject may be a plant or a fungus.
[0053] The terms "therapeutic agent," "therapeutic drug," or "treatment agent" are used interchangeably to refer to molecules or compounds that, when administered to a subject, impart several beneficial effects. These beneficial effects include the availability of diagnostic measurements; improvement of a disease, symptom, disorder, or pathological condition; reduction of a disease, symptom, disorder, or pathological condition or prevention of its onset; and, generally, neutralization of a disease, symptom, disorder, or pathological condition.
[0054] As used herein, “treatment,” “to treat,” “to alleviate,” or “to improve” are interchangeable. These terms refer to an approach to obtain a benefit or desired outcome, for example, but not limited to, a therapeutic benefit and / or a preventive benefit. A therapeutic benefit means any treatment-related improvement or effect on one or more diseases, conditions, or symptoms during treatment. A preventive benefit means that a composition may be administered to subjects at risk of developing a particular disease, condition, or symptom, or to subjects reporting one or more physiological symptoms of a disease, even if the disease, condition, or symptoms have not previously manifested.
[0055] The term “effective dose” or “therapeutic effective dose” refers to the amount of drug sufficient to produce a benefit or desired outcome. The therapeutic effective dose may vary depending on one or more of the subject and condition being treated, the subject’s weight and age, the severity of the condition, and the mode of administration, which can be readily determined by those skilled in the art. This term also applies to the dose that provides an image for detection by any of the imaging methods described herein. The prescribed dose may vary depending on one or more of the specific drug selected, the administration regimen to be followed, whether or not it is administered in combination with other compounds, the timing of administration, the tissue to be imaged, and the physical delivery system on which it is carried.
[0056] The implementation of this invention will, unless otherwise specified, utilize conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, which are within the scope of the skill of those skilled in the art. See Sambrook, Fritsch and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, 2nd edition (1989); CURRENT PROTOCOLS IN MOLECULAR BIOLOGY (FMAusubel, et al. eds., (1987)); series METHODS IN ENZYMOLOGY (Academic Press, Inc.): PCR 2: A PRACTICAL APPROACH (MJ MacPherson, BD Hames and GRTaylor eds. (1995)), Harlow and Lane, eds. (1988) ANTIBODIES, A LABORATORY MANUAL, and ANIMAL CELL CULTURE (RIFreshney, ed. (1987)).
[0057] Some aspects of the present invention relate to vector systems comprising one or more vectors, or to the vectors themselves. The vectors can be designed for the expression of CRISPR transcripts (e.g., nucleic acid transcripts, proteins, or enzymes) in prokaryotic or eukaryotic cells. For example, CRISPR transcripts can be expressed in bacterial cells, e.g., Escherichia coli, insect cells (using baculovirus expression vectors), yeast cells, or mammalian cells. Suitable host cells are further discussed in Goeddel, *Gene Expression Technology: Methods in Enzymology* 185, Academic Press, San Diego, Calif. (1990). Alternatively, recombinant expression vectors can be transcribed and translated in vitro, for example, using a T7 promoter regulatory sequence and T7 polymerase.
[0058] Vectors can be introduced into prokaryotic cells and grown within them. In some embodiments, prokaryotes are used to amplify copies of vectors to be introduced into eukaryotic cells, or as intermediate vectors in the production of vectors to be introduced into eukaryotic cells (e.g., amplifying plasmids as part of a viral vector packaging system). In some embodiments, prokaryotes are used to amplify copies of vectors and express one or more nucleic acids to provide, for example, a resource of one or more proteins for delivery to host cells or host organisms. Protein expression in prokaryotes is most often carried out in Escherichia coli using vectors containing constitutive or inductive promoters directed towards the expression of either fusion or non-fusion proteins. Fusion vectors add a number of amino acids to the protein they encode, for example, to the amino terminus of a recombinant protein. Such fusion vectors can serve one or more purposes, for example, (i) increased expression of recombinant proteins; (ii) increased solubility of recombinant proteins; and (iii) assistance in the purification of recombinant proteins by acting as ligands in affinity purification. In fusion expression vectors, proteolytic cleavage sites are often introduced at the junction of the fusion region and the recombinant protein to allow for the separation of the recombinant protein from the fusion region after purification of the fusion protein. Examples of such enzymes and their cognitive recognition sequences include factor Xa, thrombin, and enterokinase. Exemplary fusion expression vectors include pGEX (Pharmacia Biotech Inc; Smith and Johnson, 1988. Gene 67:31-40), pMAL (New England Biolabs, Beverly, Mass.), and pRIT5 (Pharmacia, Piscataway, NJ), which fuse glutathione S-transferase (GST), maltose E-binding protein, or protein A to the target recombinant protein, respectively.
[0059] Suitable examples of inducible non-fusion Escherichia coli (E. coli) expression vectors include pTrc (Amrann et al., (1988) Gene 69:301-315) and pET 11d (Studier et al., GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990) 60-89).
[0060] In some embodiments, the vector is a yeast expression vector. Examples of vectors for expression in budding yeast (Saccharomyces cerivisae) include pYepSec1 (Baldari, et al., 1987. EMBO J.6:229-234), pMFa (Kuijan and Herskowitz, 1982. Cell 30:933-943), pJRY88 (Schultz et al., 1987. Gene 54:113-123), pYES2 (Invitrogen Corporation, San Diego, Calif.), and picZ (InVitrogen Corp, San Diego, Calif.).
[0061] In some embodiments, the vector uses a baculovirus expression vector to drive protein expression in insect cells. Examples of baculovirus vectors available for protein expression in cultured insect cells (e.g., SF9 cells) include the pAc series (Smith, et al., 1983. Mol. Cell. Biol. 3:2156-2165) and the pVL series (Lucklow and Summers, 1989. Virology 170:31-39).
[0062] In some embodiments, the vector may drive the expression of one or more sequences in mammalian cells using a mammalian expression vector. Examples of mammalian expression vectors include pCDM8 (Seed, 1987. Nature 329:840) and pMT2PC (Kaufman, et al., 1987. EMBO J.6:187-195). When used in mammalian cells, the expression vector regulatory function is typically provided by one or more regulatory elements. For example, commonly used promoters are derived from polyomas, adenovirus type 2, cytomegalovirus, Simianvirus 40, and others disclosed herein and known in the art. For other suitable expression systems for both prokaryotic and eukaryotic cells, see, for example, Chapters 16 and 17 of Sambrook, et al., MOLECULAR CLONING: A LABORATORY MANUAL. 2nd ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.
[0063] In some embodiments, recombinant mammalian expression vectors may preferentially direct the expression of nucleic acids in specific cell types (e.g., by using tissue-specific regulatory elements to express nucleic acids). Tissue-specific regulatory elements are known in the art. Non-limiting examples of suitable tissue-specific promoters include albumin promoters (liver-specific; Pinkert, et al., 1987. Genes Dev. 1:268-277), lymphoid-specific promoters (Calame and Eaton, 1988. Adv. Immunol. 43:235-275), particularly T cell receptor promoters (Winoto and Baltimore, 1989. EMBO J. 8:729-733) and immunoglobulins (Baneiji, et al., 1983. Cell 33:729-740; Queen and Baltimore, 1983. Cell 33:741-748), neuronal-specific promoters (e.g., neurofilament promoter; Byrne and Ruddle, 1989. Proc. Natl. Acad. Sci. USA 86:5473-5477), and pancreas-specific promoters (Edlund, et al.) Examples include mammary gland-specific promoters (e.g., al., 1985. Science 230:912-916), and mammary gland-specific promoters (e.g., whey promoter; U.S. Patent No. 4,873,316 and European Patent Application Publication No. 264,166). Developmental control promoters are also included, such as the mouse hox promoter (Kessel and Gruss, 1990. Science 249:374-379) and the α-fetoprotein promoter (Campes and Tilghman, 1989. Genes Dev.3:537-546).
[0064] In some embodiments, regulatory elements are operably bound to one or more elements of the CRISPR system to drive the expression of one or more elements of the CRISPR system. Generally, CRISPR (clustered equally spaced short repeats), also known as SPIDR (spacer interspersed direct repeats), constitute a family of DNA loci that are typically specific to certain bacterial species. CRISPR loci include distinct classes of interspersed short sequence repeats (SSRs) recognized in Escherichia coli (E. coli) (Ishino et al., J. Bacteriol., 169:5429-5433
[1987] ; and Nakata et al., J. Bacteriol., 171:3553-3556
[1989] ) and related genes. Similar scattered SSRs have been identified in Haloferax mediterranei, Streptococcus pyogenes, Anabaena species, and Mycobacterium tuberculosis (see Groenen et al., Mol. Microbiol., 10:1057-1065
[1993] ; Hoe et al., Emerg. Infect. Dis., 5:254-263
[1999] ; Masepohl et al., Biochim. Biophys. Acta 1307:26-30
[1996] ; and Mojica et al., Mol. Microbiol., 17:85-93
[1995] ). The CRISPR locus typically differs in repeat structure from other SSRs, and is referred to as a short-chain equispaced repeat (SRSR) (Janssen et al., OMICS J. Integ. Biol., 6:23-33
[2002] ; and Mojica et al., Mol. Microbiol., 36:244-246
[2000] ). Generally, repeats are short elements that occur in clusters that are equispaced by unique intervention sequences of substantially constant length (Mojica et al.,
[2000] , op. cit.). Repeat sequences are highly conserved between strains,The number of scattered repeats and the sequence of spacer regions typically differ from strain to strain (van Embden et al., J. Bacteriol., 182:2393-2401
[2000] ). The CRISPR locus has been identified in more than 40 prokaryotes (e.g., Jansen et al., Mol. Microbiol., 43:1565-1575
[2002] ; and Mojica et al.) (See al.,
[2005] ), for example, though not limited to, the genera Aeropyrum, Pyrobaculum, Sulfolobus, Archaeoglobus, Halocarcula, Methanobacterium, Methanococcus, Methanosarcina, Methanopyrus, Pyrococcus, Picrophilus, Thermoplasma, Corynebacterium, Mycobacterium, Streptomyces, and Aquifex. fex), Porphyromonas, Chlorobium, Thermus, Bacillus, Listeria, Staphylococcus, Clostridium, Thermoanaerobacter, Mycoplasma, Fusobacterium, Azarcus, Chromobacterium, Neisseria, Nitrosomonas, Desulfovibrio, Geobacter, Myxococcus, Campylobacter,These include the genera Wolinella, Acinetobacter, Erwinia, Escherichia, Legionella, Methylococcus, Pasteurella, Photobacterium, Salmonella, Xanthomonas, Yersinia, Treponema, and Thermotoga.
[0065] Generally, the “CRISPR system” refers collectively to transcripts and other elements involved in the expression or orientation of the activity of CRISPR-related (“Cas”) genes, such as the sequence encoding the Cas gene, the tracr (trans-activated CRISPR) sequence (e.g., tracrRNA or active partial tracrRNA), the tracrmate sequence (including “direct repeats” and partial direct repeats processed by tracrRNA with respect to the endogenous CRISPR system), the guide sequence (also referred to as “spacers” with respect to the endogenous CRISPR system), or other sequences and transcripts from the CRISPR locus. In some embodiments, one or more elements of the CRISPR system are derived from the type I, type II, or type III CRISPR system. In some embodiments, one or more elements of the CRISPR system are derived from certain organisms containing the endogenous CRISPR system, such as Streptococcus pyogenes. Generally, the CRISPR system is characterized by elements that promote the formation of the CRISPR complex at a target sequence (also referred to as a protospacer in relation to the endogenous CRISPR system). With respect to the formation of the CRISPR complex, the “target sequence” refers to a sequence designed to have complementarity with a guide sequence, and hybridization between the target sequence and the guide sequence promotes the formation of the CRISPR complex. Complete complementarity is not necessarily required, provided that sufficient complementarity exists to induce hybridization and promote the formation of the CRISPR complex. The target sequence may include any polynucleotide, e.g., DNA or RNA polynucleotide. In some embodiments, the target sequence is localized in the nucleus or cytoplasm of a cell. In some embodiments, the target sequence may be located in an organelle of a eukaryotic cell, e.g., mitochondria or chloroplast. A sequence or template that can be used for recombination into a targeted locus containing the target sequence is referred to as an “editing template,” “edited polynucleotide,” or “editing sequence.” In embodiments of the present invention, an exogenous template polynucleotide may be referred to as an editing template. In one embodiment of the present invention, the recombination is homologous recombination.
[0066] Typically, with respect to the endogenous CRISPR system, the formation of a CRISPR complex (including a guide sequence that hybridizes to a target sequence and complexes with one or more Cas proteins) results in the cleavage of one or both strands in or near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs from it). Although not theoretically constrained, a tracr sequence may contain, or consist of, all or part of a wild-type tracr sequence (e.g., more than approximately 20, 26, 32, 45, 48, 54, 63, 67, 85, or more nucleotides of a wild-type tracr sequence), and may also form part of a CRISPR complex by hybridization along at least part of the tracr sequence with all or part of a tracr mate sequence operably bound to a guide sequence, for example. In some embodiments, the tracr sequence has sufficient complementarity to the tracr mate sequence to hybridize and participate in the formation of the CRISPR complex. As with the target sequence, complete complementarity is not required, but sufficient complementarity may be present for it to be functional. In some embodiments, when optimally aligned, the tracr sequence has at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence complementarity along the length of the tracr mate sequence. In some embodiments, one or more vectors driving the expression of one or more elements of the CRISPR system are introduced into host cells, resulting in the expression of elements of the CRISPR system being directed towards the formation of the CRISPR complex at one or more target sites. For example, the Cas enzyme, a guide sequence bound to the tracr mate sequence, and the tracr sequence can each be operably bound to separate regulatory elements on separate vectors. Alternatively, two or more elements expressed from the same or different regulatory elements can be combined in a single vector, and one or more additional vectors providing any components of the CRISPR system are not included in the first vector.CRISPR elements combined in a single vector can be positioned in any preferred orientation; for example, one element can be localized to the 5' side (upstream) or 3' side (downstream) of a second element. The coding sequence of one element can be localized on the same or reverse strand of the coding sequence of a second element and oriented in the same or reverse direction. In some embodiments, a single promoter drives the expression of a transcript encoding a CRISPR enzyme, as well as one or more guide sequences embedded within one or more intron sequences, tracr mate sequences (optionally operably bound to the guide sequence), and tracr sequences (e.g., each in a different intron, two or more in at least one intron, or all in a single intron). In some embodiments, the CRISPR enzyme, guide sequence, tracr mate sequence, and tracr sequence are operably bound to the same promoter and expressed from there.
[0067] In some embodiments, the vector includes one or more insertion sites, e.g., restriction endonuclease recognition sequences (also referred to as “cloning sites”). In some embodiments, one or more insertion sites (e.g., about or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more) are localized upstream and / or downstream of one or more sequence elements of one or more vectors. In some embodiments, the vector includes insertion sites upstream of a tracr-mate sequence and optionally downstream of a regulatory element operably bound to the tracr-mate sequence, so that after insertion of the guide sequence into the insertion site and upon expression, the guide sequence directs the CRISPR complex to sequence-specific binding to a target sequence in eukaryotic cells. In some embodiments, the vector includes two or more insertion sites, each insertion site localized between two tracr-mate sequences to allow insertion of the guide sequence at its respective site. In such arrangements, the two or more guide sequences may include two or more copies of a single guide sequence, two or more different guide sequences, or a combination thereof. When using multiple different guide sequences, a single expression construct can be used to target CRISPR activity against multiple different corresponding target sequences within a cell. For example, a single vector may contain approximately or more than approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or more guide sequences. In some embodiments, such guide sequence-containing vectors are provided and can optionally be delivered to cells.
[0068] In some embodiments, the vector includes a regulatory element operably bound to an enzyme-coding sequence encoding a CRISPR enzyme, such as a Cas protein. Non-exclusive examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, their homologs, or modified versions thereof. These enzymes are publicly known; for example, the amino acid sequence of the *Streptococcus pyogenes* (S. pyogenes) Cas9 protein can be found in the SwissProt database under accession number Q99ZW2. In some embodiments, the unmodified CRISPR enzyme has DNA cleavage activity and is, for example, Cas9. In some embodiments, the CRISPR enzyme is Cas9 and may be Cas9 from *Streptococcus pyogenes* (S. pyogenes) or *Streptococcus pneumoniae* (S. pneumoniae). In some embodiments, the CRISPR enzyme is directed to cleave one or both strands within the target sequence and / or the complementary strand of the target sequence at the localization of the target sequence. In some embodiments, the CRISPR enzyme is directed to cleave one or both strands within approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence. In some embodiments, the vector encodes a CRISPR enzyme that is mutated from the corresponding wild-type enzyme, and as a result, the mutated CRISPR enzyme lacks the ability to cleave one or both strands of the target polynucleotide containing the target sequence.For example, the substitution of aspartic acid to alanine in the RuvC I catalytic domain of Cas9 from Streptococcus pyogenes (S. pyogenes) (D10A) converts Cas9 from a nuclease that cleaves both strands to a nickase (cleaves a single strand). Other examples of mutations that convert Cas9 to a nickase include, but are not limited to, H840A, N854A, and N863A. In some embodiments, the Cas9 nickase can be used in combination with guide sequences, e.g., two guide sequences that target the sense and antisense strands of the DNA target, respectively. This combination makes it possible to nick both strands and use them to induce an NHEJ. The applicants have demonstrated the efficacy of two nickase targets (i.e., sgRNAs that are identically localized but target different strands of DNA) in the induction of mutagenic NHEJs (data not shown). While a single nickase (Cas9-D10A with a single sgRNA) cannot induce NHEJ and create indels, we have shown that a dual nickase (Cas9-D10A and two sgRNAs targeted to different strands at the same localization) can do so in human embryonic stem cells (hESCs). The efficiency is approximately 50% of that of a nuclease (i.e., normal Cas9 without the D10 mutation) in hESCs.
[0069] As a further example, mutations in two or more catalytic domains of Cas9 (RuvC I, RuvC II, and RuvC III) can be used to produce mutant Cas9 that substantially lacks all DNA cleavage activity. In some embodiments, the D10A mutation is combined with one or more H840A, N854A, or N863A mutations to produce a Cas9 enzyme that substantially lacks all DNA cleavage activity. In some embodiments, a CRISPR enzyme is considered substantially lacking all DNA cleavage activity if the DNA cleavage activity of the mutant enzyme is less than approximately 25%, 10%, 5%, 1%, 0.1%, 0.01%, or less than that of its non-mutant form. Other mutations may be useful; if Cas9 or other CRISPR enzymes are from species other than Streptococcus pyogenes, mutations in the corresponding amino acids can be made to achieve similar effects.
[0070] In some embodiments, the enzyme coding sequence encoding a CRISPR enzyme is codon-optimized for expression in specific cells, eukaryotic cells, etc. Eukaryotic cells may be those of or derived from specific organisms, e.g., mammals, including, but not limited to, humans, mice, rats, rabbits, dogs, or non-human primates. Generally, codon optimization refers to the process of modifying a nucleic acid sequence to improve expression in a target host cell by replacing at least one codon in the native sequence (e.g., more than approximately 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more) with a more frequent or most frequent codon used in the host cell's genes, while maintaining the native amino acid sequence. Different species exhibit specific biases for certain codons of particular amino acids. Codon bias (differences in codon usage frequency between organisms) often correlates with the efficiency of messenger RNA (mRNA) translation, which is then thought to depend, in particular, on the characteristics of the translated codon and the availability of specific transfer RNA (tRNA) molecules. The dominance of tRNAs selected in a cell is generally a reflection of the codons most frequently used in peptide synthesis. Therefore, genes can be tuned for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, in "codon usage databases," and these tables can be adapted using numerous methods. See Nakamura, Y., et al. "Codon usage tabulated from the international DNA sequence databases: status for the year 2000" Nucl. Acids Res. 28:292 (2000). Computer algorithms for codon-optimizing specific sequences for expression in specific host cells are also available, for example, Gene Forge (Aptagen; Jacobus, PA).In some embodiments, one or more codons in the sequence encoding the CRISPR enzyme (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more, or all of the codons) correspond to the codons most frequently used for a particular amino acid.
[0071] In some embodiments, the vector encodes a CRISPR enzyme containing one or more nuclear localization sequences (NLSs), e.g., about or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs. In some embodiments, the CRISPR enzyme contains about or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs at or near the amino terminus, about or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs at or near the carboxy terminus, or a combination thereof (e.g., one or more NLSs at the amino terminus and one or more NLSs at the carboxy terminus). If two or more NLSs are present, each can be selected independently of the others, such that a single NLS may exist in two or more copies, and / or in combination with one or more other NLSs that exist in one or more copies. In a preferred embodiment of the present invention, the CRISPR enzyme contains at most six NLSs. In some embodiments, an NLS is considered to be located near the N or C terminus if its nearest neighbor amino acid is within approximately 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, or more amino acids along the polypeptide chain from the N or C terminus. Typically, an NLS consists of one or more short sequences of positively charged lysine or arginine exposed on the protein surface, although other types of NLSs are known.Non-limiting examples of NLS include: NLS of the SV40 virus large T antigen with the amino acid sequence PKKKRKV; NLS from nucleoplasmin (e.g., nucleoplasmin bipartite NLS with the sequence KRPAATKKAGQAKKKK); c-mycNLS with the amino acid sequence PAAKRVKLD or RQRRNELKRSP; hRNPA1 M9 NLS with the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY; IBB domain sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV from importin alpha; myoma T protein sequences VSRKRPRP and PPKKARED; human p53 sequence POPKKKPL; mouse c-abl Examples include the sequence SALIKKKKKMAP of influenza IV; the sequences DRLRR and PKQKKRK of influenza virus NS1; the sequence RKLKKKIKKL of hepatitis virus delta antigen; the sequence REKKKFLKRR of mouse Mx1 protein; the sequence KRKGDEVDGVDEVAKKKSKK of human poly(ADP-ribose) polymerase; and the NLS sequence derived from the sequence RKCLQAGMNLEARKTKK of the steroid hormone receptor (human) glucocorticoid.
[0072] Generally, one or more NLSs are strong enough to drive the accumulation of a detectable amount of CRISPR enzyme in the nucleus of eukaryotic cells. Generally, the strength of nuclear localization activity may depend on the number of NLSs in the CRISPR enzyme, the specific NLSs used, or a combination of these factors. Detection of nuclear accumulation can be carried out by any suitable technique. For example, a detectable marker can be fused to the CRISPR enzyme, and as a result, intracellular localization can be visualized in combination with means for detecting nuclear localization (e.g., nuclear-specific staining, e.g., DAPI). Examples of detectable markers include fluorescent proteins (e.g., green fluorescent protein, or GFP;RFP;CFP) and epitope tags (HA tags, flag tags, SNAP tags). Cell nuclei can also be isolated from cells, and their contents can then be analyzed by any suitable process for detecting proteins, e.g., immunohistochemical analysis, Western blotting, or enzyme activity assays. Nuclear accumulation can also be measured indirectly, for example, by assays for the effects of CRISPR complex formation (e.g., assays for DNA cleavage or mutation at target sequences, or assays for changes in gene expression activity affected by CRISPR complex formation and / or CRISPR enzyme activity), compared to controls that are not exposed to CRISPR enzymes or complexes, or are exposed to CRISPR enzymes lacking one or more NLSs.
[0073] Generally, the guide sequence is any polynucleotide sequence that hybridizes with the target sequence and has sufficient complementarity to the target polynucleotide sequence to direct the CRISPR complex to sequence-specific binding to the target sequence. In some embodiments, the degree of complementarity between the guide sequence and its corresponding target sequence exceeds approximately 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or greater, when optimally aligned using a suitable alignment algorithm. Optimal alignment can be determined using any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler Transform (e.g., Burrows Wheeler Aligner), ClustalW, ClustalX, BLAT, Novoalign (Novocraft Technologies, ELAND (Illumina, San Examples include Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). In some embodiments, the guide sequence has a nucleotide length of approximately or greater than approximately 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, or more. In some embodiments, the guide sequence has a nucleotide length of approximately 75, 50, 45, 40, 35, 30, 25 The nucleotide lengths are 20, 15, 12, or less. The ability of a guide sequence to direct sequence-specific binding of the CRISPR complex to a target sequence can be evaluated by any suitable assay. For example, components of the CRISPR system sufficient to form a CRISPR complex, such as the guide sequence to be tested, can be provided to host cells having the corresponding target sequence, for example, by transtransfer using a vector encoding components of the CRISPR sequence, and then preferential cleavage within the target sequence can be evaluated by, for example, the Surveyor assay described herein.Similarly, cleavage of a target polynucleotide sequence can be evaluated by providing the target sequence, components of the CRISPR complex, e.g., a guide sequence to be tested and a control guide sequence different from the test guide sequence, in a test tube, and comparing the ratio of binding or cleavage at the target sequence between the test and control guide sequence reactions. Other assays are conceivable and will be recognized by those skilled in the art.
[0074] The guide sequence can be selected to target any target sequence. In some embodiments, the target sequence is a sequence within the cell's genome. Exemplary target sequences include those unique within the target genome. For example, for Streptococcus pyogenes (S. pyogenes) Cas9, a unique target sequence in the genome could be the Cas9 target site of form MMMMMMMNNNNNNNNNNNNXGG, where NNNNNNNNNNNNXGG (where N is A, G, T, or C; X may be any) has a single occurrence in the genome. For S. thermophilus CRISPR1 Cas9, the unique target sequence in the genome is the Cas9 target site of form MMMMMMMMNNNNNNNNNNNNXXAGAAW, where NNNNNNNNNNNNXXAGAAW (N is A, G, T, or C; X is any of these) For example, the form MMMMMMMMMNNNNNNNNNNNXXAGAAW is a unique target sequence in the genome, and NNNNNNNNNNNNXXAGAAW (where N is A, G, T, or C; X may be any; W is A or T) has a unique target sequence in the genome. For Streptococcus pyogenes Cas9, a unique target sequence in the genome is the form MMMMMMMMNNNNNNNNNNNNXGGXG is a unique target sequence in the genome, and NNNNNNNNNNNNXGGXG (where N is A, G, T, or C; X may be any) has a unique target sequence in the genome. A unique target sequence in the genome is the Streptococcus pyogenes (S. pyogenes) Cas9 target site of form MMMMMMMMMNNNNNNNNNNNXGGXG, where NNNNNNNNNNNXGGXG (N is A, G, T, or C; X may be any of these) has a single occurrence in the genome. In each of these sequences, "M" can be A, G, T, or C, and this does not need to be considered when identifying the sequence as unique.
[0075] In some embodiments, the guide sequence is selected to reduce the degree of secondary structure within the guide sequence. The secondary structure can be determined by any suitable polynucleotide folding algorithm. Some programs are based on the calculation of the minimum Gibbs free energy. An example of such an algorithm is mFold, described by Zuker and Stiegler (Nucleic Acids Res. 9 (1981), 133-148). Another exemplary folding algorithm is RNAfold, an online web server developed by the Institute for Theoretical Chemistry at the University of Vienna, which uses a centroid structure prediction algorithm (see, e.g., ARGruber et al., 2008, Cell 106(1):23-24; and PA Carr and GM Church, 2009, Nature Biotechnology 27(12):1151-62). Further algorithms can be found in U.S. Patent Application No. TBA (Broad Reference No. BI2012 / 084 44790.11.2022), which is incorporated herein by reference.
[0076] Generally, a tracr-mate sequence includes any sequence that has sufficient complementarity with the tracr sequence to facilitate one or more of the following: (1) the excision of a guide sequence flanked by the tracr-mate sequence in a cell containing the corresponding tracr sequence; and (2) the formation of a CRISPR complex at the target sequence (the CRISPR complex includes the tracr-mate sequence which hybridizes to the tracr sequence). Generally, the degree of complementarity is based on the optimal alignment of the tracr-mate sequence and the tracr sequence along the shorter of the two sequences. The optimal alignment can be determined by any preferred alignment algorithm, which may further describe secondary structures, such as self-complementarity within the tracr sequence or tracr-mate sequence. In some embodiments, the degree of complementarity between the tracr sequence and the tracr-mate sequence along the shorter of the two sequences, when optimally aligned, exceeds approximately 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99%, or greater. Figures 12B and 13B provide an illustrative description of the optimal alignment between the tracr sequence and the tracr mate sequence. In some embodiments, the tracr sequence has a nucleotide length of approximately or greater than approximately 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, or more. In some embodiments, the tracr sequence and the tracr mate sequence are contained within a single transcript, resulting in hybridization between the two to produce a secondary structure, e.g., a transcript having a hairpin. A preferred loop-forming sequence used in a hairpin structure is 4 nucleotides long and most preferably has the sequence GAAA. However, longer or shorter loop sequences can be used, which may be alternative sequences. The sequence preferably comprises a nucleotide triplet (e.g., AAA) and an additional nucleotide (e.g., C or G). Examples of loop-forming sequences include CAAA and AAAG. In one embodiment of the present invention, the transcript or the polynucleotide sequence to be transcribed has at least two or more hairpins.In a preferred embodiment, the transcript has two, three, four, or five hairpins. In another further embodiment of the present invention, the transcript has at most five hairpins. In some embodiments, the single transcript further comprises a transcription termination sequence; preferably, this is a poly-T sequence, e.g., six T nucleotides. An exemplary description of such a hairpin structure is provided below in Figure 13B, where the last "N" and the sequence portion on the 5' side upstream of the loop correspond to the tracr mate sequence, and the sequence portion on the 3' side of the loop correspond to the tracr sequence. Further non-limiting examples of a single polynucleotide comprising a guide sequence, a tracr mate sequence, and a tracr sequence are as follows (listed from 5' to 3'), where "N" represents the base of the guide sequence, the first lowercase block represents the tracr mate sequence, the second lowercase block represents the tracr sequence, and the last poly-T sequence represents the transcription terminator. [ka] In some embodiments, sequences (1) to (3) are used in combination with Cas9 from S. thermophilus CRISPR1. In some embodiments, sequences (4) to (6) are used in combination with Cas9 from S. pyogenes. In some embodiments, the tracr sequence is a separate transcript from the transcript containing the tracr mate sequence (for example, as illustrated in the upper figure of Figure 13B).
[0077] In some embodiments, recombinant templates are also provided. Recombinant templates are components of other vectors described herein and may be contained within a separate vector or provided as separate polynucleotides. In some embodiments, recombinant templates are designed to function as templates in or near a target sequence that are nicked or cleaved by a CRISPR enzyme as part of a CRISPR complex, for example. The template polynucleotide may be of any preferred length, e.g., a nucleotide length exceeding approximately or about 10, 15, 20, 25, 50, 75, 100, 150, 200, 500, 1000, or a number greater than that. In some embodiments, the template polynucleotide is complementary to a portion of the polynucleotide containing the target sequence. When optimally aligned, the template polynucleotide may overlap with one or more nucleotides of the target sequence (e.g., more than approximately or about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, or a number greater than that). In some embodiments, when the polynucleotides containing the template sequence and the target sequence are optimally aligned, the nearest neighbor nucleotide of the template polynucleotide lies within approximately 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 1000, 5000, 10000, or more nucleotides from the target sequence.
[0078] In some embodiments, the CRISPR enzyme is part of a fusion protein containing one or more heterologous protein domains (e.g., about one or more domains other than the CRISPR enzyme, such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more). The CRISPR enzyme fusion protein may contain any additional protein sequences and, optionally, linker sequences between any two domains. Examples of protein domains that can be fused to a CRISPR enzyme include, but are not limited to, epitope tags, reporter gene sequences, and protein domains having one or more of the following activities: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, and nucleic acid binding activity. Non-exclusive examples of epitope tags include histidine (His) tags, V5 tags, FLAG tags, influenza hemagglutinin (HA) tags, Myc tags, VSV-G tags, and thioredoxin (Trx) tags. Examples of reporter genes include, but are not limited to, glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and autofluorescent proteins, such as blue fluorescent protein (BFP). CRISPR enzymes can be fused to gene sequences encoding proteins or protein fragments that bind to DNA molecules or other cellular molecules, such as, but are not limited to, maltose-binding protein (MBP), S-tags, Lex A DNA-binding domain (DBD) fusions, GAL4 DNA-binding domain fusions, and herpes simplex virus (HSV) BP16 protein fusions. Additional domains that may form part of a fusion protein containing a CRISPR enzyme are described in U.S. Patent Application Publication No. 20110059502, which is incorporated herein by reference. In some embodiments, a tagged CRISPR enzyme is used to identify the localization of a target sequence.
[0079] In some embodiments, the present invention provides methods for delivering one or more polynucleotides, e.g., or one or more vectors described herein, one or more transcripts thereof, and / or one or a protein transcribed therefrom, to a host cell. In some embodiments, the present invention further provides cells produced by such cells, and organisms (e.g., animals, plants, or fungi) comprising or produced therefrom such cells. In some embodiments, a CRISPR enzyme in combination with (and optionally complexed with) a guide sequence is delivered to the cell. Nucleic acids can be introduced into mammalian cells or target tissues using conventional viral and nonviral-based gene transfer methods. Using such methods, nucleic acids encoding components of the CRISPR system can be administered to cells in culture or into a host organism. Nonviral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of vectors described herein), naked nucleic acids, and delivery vehicles, e.g., nucleic acids complexed with liposomes. Viral vector delivery systems include DNA and RNA viruses having genomes that are episomal or integrated after delivery to cells.For an overview of the gene therapy procedure, see Anderson, Science 256:808-813 (1992); Nabel & Felgner, TIBTECH 11:211-217 (1993); Mitani & Caskey, TIBTECH 11:162-166 (1993); Dillon, TIBTECH 11:167-175 (1993); Miller, Nature 357:455-460 (1992); Van Brunt, Biotechnology 6(10):1149-1154 (1988); Vigne, Restorative Neurology and Neuroscience 8:35-36 (1995); Kremer & Perricaudet, British Medical Bulletin 51(1):31-44 (1995); Haddada et al., in Current Topics in Microbiology and See Immunology Doerfler and Boehm (eds) (1995) and Yu et al., Gene Therapy 1:13-26 (1994).
[0080] Nonviral delivery methods for nucleic acids include lipofection, nucleofection, microinjection, gene guns, virosomes, liposomes, immunoliposomes, polycations or lipids: nucleic acid conjugates, naked DNA, artificial virions, and drug-enhanced DNA uptake. Lipofection is described, for example, in U.S. Patent Nos. 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam® and Lipofectin®). Suitable cations and neutral lipids for efficient receptor-recognized lipofection of polynucleotides include those from Felgner, International Publication No. 91 / 17424; International Publication No. 91 / 16024. Delivery may be to cells (e.g., in vitro or ex vivo administration) or target tissues (e.g., in vivo administration).
[0081] The preparation of lipid-nucleic acid complexes, such as targeted liposomes, e.g., immunolipid complexes, is well known to those skilled in the art (e.g., Crystal, Science 270:404-410 (1995); Blaese et al., Cancer Gene Ther. 2:291-297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994); Gao et al., Gene Therapy 2:710-722 (1995); Ahmad et al., Cancer Res.52:4817-4820(1992); see U.S. Patent Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028, and 4,946,787).
[0082] The use of RNA or DNA virus-based systems for nucleic acid delivery utilizes highly evolved processes to target viruses to specific cells in the body and transport the viral payload to the nucleus. Viral vectors can be administered directly to patients (in vivo), or they can be used to treat cells in vitro, and in some cases, modified cells can be administered to patients (ex vivo). Conventional virus-based systems include retroviral, lentiviral, adenovirus, adeno-associated, and herpes simplex virus vectors for gene transfer. Integration into the host genome is considered for retroviral, lentiviral, and adeno-associated virus gene transfer methods, often resulting in long-term expression of the inserted transgene. Furthermore, high transduction efficiencies have been observed in many different cell types and target tissues.
[0083] The tropism of retroviruses can be altered by incorporating foreign envelope proteins, thereby expanding the potential target population of target cells. Lentiviral vectors are retroviral vectors that transduce or infect non-dividing cells and typically produce high viral titers. Therefore, the selection of retroviral gene transfer systems depends on the target tissue. Retroviral vectors contain cis-acting long-chain terminal repeats that have the ability to package foreign sequences up to 6–10 kb. The minimum cis-acting LTR is the replication and The product is sufficient for packaging and then used to integrate therapeutic genes into target cells to provide permanent transgene expression. Widely used retroviral vectors include those based on murine leukemia virus (MuLV), gibbon leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), or combinations thereof (see, for example, Buchscher et al., J.Virol.66:2731-2739 (1992); Johann et al., J.Virol.66:1635-1640 (1992); Sommnerfelt et al., Virol.176:58-59 (1990); Wilson et al., J.Virol.63:2374-2378 (1989); Miller et al., J.Virol.65:2220-2224 (1991); and PCT / US94 / 05700).
[0084] For applications where transient expression is preferred, adenovirus-based systems can be used. Adenovirus-based vectors can exhibit extremely high transduction efficiency in many cell types and do not require cell division. High titers and expression levels have been obtained with such vectors. These vectors can be produced in large quantities using relatively simple systems. For example, adeno-associated virus ("AAV") vectors can be used to transduce target nucleic acids into cells in the in vitro production of nucleic acids and peptides, as well as for in vivo and ex vivo gene therapy procedures (see, e.g., West et al., Virology 160:38-47 (1987); U.S. Patent No. 4,797,368; International Publication No. 93 / 24641; Kotin, Human Gene Therapy 5:793-801 (1994); Muzyczka, J. Clin. Invest. 94:1351 (1994). Construction of recombinant AAV vectors is described in numerous publications, e.g., U.S. Patent No. 5,173,414; Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985); Tratschin, et al. This is described in al., Mol.Cell.Biol.4:2072-2081(1984); Hermonat & Muzyczka, PNAS 81:6466-6470(1984); and Samulski et al., J.Virol.63:03822-3828(1989).
[0085] Typically, packaging cells are used to form viral particles that can infect host cells. Examples of such cells include 293 cells for packaging adenoviruses, and φ2 or PA317 cells for packaging retroviruses. Viral vectors used in gene therapy are usually produced by creating a cell line that packages nucleic acid vectors into viral particles. The vector typically contains the minimal viral sequence required for packaging and subsequent integration into the host, with other viral sequences replaced by expression cassettes for polynucleotides to be expressed. Deficient viral function is typically supplied trans by the packaging cell line. For example, AAV vectors used in gene therapy typically contain only the ITR sequence from the AAV genome required for packaging and integration into the host genome. The viral DNA is packaged into a cell line containing helper plasmids that encode other AAV genes, namely rep and cap, but lack the ITR sequence. The cell line can also be infected with adenovirus as a helper. Helper viruses facilitate the replication of AAV vectors and the expression of AAV genes from helper plasmids. Helper plasmids are not packaged in significant quantities due to the lack of ITR sequences. Adenovirus contamination can be reduced, for example, by heat treatment, in which adenoviruses are more susceptible than AAV. Additional methods for delivering nucleic acids to cells are known to those skilled in the art; see, for example, U.S. Patent Application Publication No. 20030087817, incorporated herein by reference.
[0086] In some embodiments, host cells are transfused transiently or nontransiently with one or more vectors described herein. In some embodiments, cells are transfused in their naturally occurring state within the subject. In some embodiments, the cells to be transfused are collected from the subject. In some embodiments, the cells are derived from cells collected from the subject, e.g., a cell line. A wide range of cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to, C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLa-S3, Huh1, Huh4, Huh7, HUVEC, HASMC, HEKn, HEKa, MiaPaCell, Panc1, PC-3, TF1, CTLL-2, C1R, Rat6, CV1, RPTE, A10, T24, J82, A375, ARH-77, Calu1, SW480, SW620, SKOV3, SK-UT, CaCo2, P388D1, SEM-K2, WEHI-231, HB56, TIB55, Jurkat, J45.01, LRMB, Bcl-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, Hep G2, HeLa B, HeLa T4, COS, COS-1, COS-6, COS-M6A, BS-C-1 monkey kidney epithelium, BALB / 3T3 mouse embryonic fibroblasts, 3T3Swiss, 3T3-L1, 132-d5 human fetal fibroblasts; 10.1 mouse fibroblasts, 293-T, 3T3, 721, 9L, A2780, A2780ADR, A2780cis , A172, A20, A253, A431, A-549, ALC, B16, B35, BCP-1 cells, BEAS-2B, bEnd.3, BHK-21, BR29 3, BxPC3, C3H-10T1 / 2, C6 / 36, Cal-27, CHO, CHO-7, CHO-IR, CHO-K1, CHO-K2, CHO-T, CHO Dhfr- / -, COR-L23, COR-L23 / CPR, COR-L23 / 5010, COR-L23 / R23, COS-7, COV-434, CML T1, CMT, CT26, D17, DH82, DU145, DuCaP, EL4, EM2, EM3, EMT6 / AR1, EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA2, HEK-293, HeLa, Hepa1c1c7, HL-60, HMEC, HT-29, Jurkat, JY cells, K562 cells, Ku 812, KCL22, KG1, KYO1, LNCap, Ma-Mel1-48, MC-38, MCF-7, MCF-10A, MDA-MB-231, MDA-MB-468, MDA-MB-435, MDCK II, MDCK Examples include II, MOR / 0.2R, MONO-MAC6, MTD-1A, MyEnd, NCI-H69 / CPR, NCI-H69 / LX10, NCI-H69 / LX20, NCI-H69 / LX4, NIH-3T3, NALM-1, NW-145, OPCN / OPCT cell lines, Peer, PNT-1A / PNT2, RenCa, RIN-5F, RMA / RMAS, Saos-2 cells, Sf-9, SkBr3, T2, T-47D, T84, THP1 cell lines, U373, U87, U937, VCaP, Vero cells, WM39, WT-49, X63, YAC-1, YAR, and their transgenic variants. These cell lines are species known to those skilled in the art. These are available from various resources (see, for example, the American Type Culture Collection (ATCC) (Manassus, Va.)). In some embodiments, a novel cell line containing one or more vector-derived sequences is established using cells transfected with one or more vectors described herein. In some embodiments, a novel cell line containing the modification but lacking any other exogenous sequences is established using cells transiently transfected with components of the CRISPR system described herein (e.g., transient transfecting with one or more vectors, or transfecting with RNA) and modified through the activity of the CRISPR complex. In some embodiments, cells transiently or nontransitively transfected with one or more vectors described herein, or cell lines derived from such cells, are used in the evaluation of one or more test compounds.
[0087] In some embodiments, non-human transgenic animals or transgenic plants are produced using one or more vectors described herein. In some embodiments, the transgenic animal is a mammal, e.g., a mouse, rat, or rabbit. In some embodiments, the organism or subject is a plant. In some embodiments, the organism or subject or plant is an algae. Methods for producing transgenic plants and animals are known in the art and generally start from, for example, the cell translocation methods described herein. Transgenic animals are also provided, as are transgenic plants, particularly crops and algae. Transgenic animals or plants may be useful in applications other than providing disease models. These include, for example, food or feed production through the expression of higher levels of proteins, carbohydrates, nutrients, or vitamins than normally found in the wild type. In this regard, transgenic plants, particularly legumes and tubers, as well as animals, particularly mammals, e.g., livestock (cattle, sheep, goats, and pigs), and even poultry and edible insects are preferred.
[0088] Transgenic algae or other plants, such as rapeseed, may be particularly useful in the production of vegetable oils or biofuels, such as alcohols (especially methanol and ethanol). These can be engineered to express or overexpress high levels of oil or alcohol for use in the oil or biofuel industry.
[0089] In one embodiment, the present invention provides a method for modifying a target polynucleotide in a eukaryotic cell. In some embodiments, the method comprises binding a CRISPR complex to the target polynucleotide to cause cleavage of the target polynucleotide, thereby modifying the target polynucleotide, the CRISPR complex comprising a CRISPR enzyme complexing with a guide sequence which hybridizes to a target sequence in the target polynucleotide, the guide sequence being bound to a tracr mate sequence which then hybridizes to a tracr sequence.
[0090] In one embodiment, the present invention provides a method for modifying the expression of a polynucleotide in a eukaryotic cell. In some embodiments, the method comprises conjugating a CRISPR complex to a polynucleotide such that the conjugation results in an increase or decrease in the expression of the polynucleotide; the CRISPR complex comprises a CRISPR enzyme complexing with a guide sequence which hybridizes to a target sequence in the target polynucleotide, the guide sequence being conjugated to a tracr mate sequence which then hybridizes to a tracr sequence.
[0091] With recent advances in crop genomics, the ability to perform efficient and cost-effective gene editing and manipulation using CRISPR-Cas systems enables the rapid selection and comparison of single and multiplex gene manipulations for transforming such genomes to improve production and enhance traits. In this regard, U.S. patents and publications: U.S. Patent No. 6,603,061 - Agrobacterium-Mediated Plant Transformation Method; U.S. Patent No. 7,868,149 - Plant Genome Sequences and Uses Thereof and U.S. Patent Application Publication No. 2009 / 0100536 - Transgenic Plants with Enhanced Agronomic Traits are referenced, and all of the contents and disclosures of each of these are incorporated herein by reference in their entirety. In the practice of the present invention, the contents and disclosures of Morrell et al. “Crop genomics: advances and applications” Nat Rev Genet. 2011 Dec 29;13(2):85-96 are also incorporated herein by reference in their entirety. In an advantageous embodiment of the present invention, microalgae are engineered using a CRISPR / Cas9 system (Example 14). Accordingly, references to animal cells herein may also apply to plant cells with necessary modifications, unless otherwise specified.
[0092] In one embodiment, the present invention provides a method for modifying a target polynucleotide in eukaryotic cells, which may be in vivo, ex vivo, or in vitro. In some embodiments, the method involves sampling cells or populations of cells from human or non-human animals or plants (including microalgae) and modifying one or more cells. Culturing can be performed ex vivo at any stage. One or more cells may also be reintroduced into non-human animals or plants (including microalgae).
[0093] In one embodiment, the present invention provides a kit containing one or more of the elements disclosed in the above methods and compositions. In some embodiments, the kit includes a vector system and instructions for use of the kit. In some embodiments, the vector system includes (a) a first regulatory element operably bound to one or more insertion sites for inserting a tracr mate sequence and a guide sequence upstream of the tracr mate sequence (the guide sequence, when expressed, directs the sequence-specific binding of the CRISPR complex to a target sequence in a eukaryotic cell, and the CRISPR complex includes a CRISPR enzyme that complexes with (1) a guide sequence hybridized to the target sequence and (2) a tracr mate sequence hybridized to the tracr sequence); and / or (b) a second regulatory element operably bound to an enzyme coding sequence encoding the CRISPR enzyme, including a nuclear localization sequence. The elements may be provided individually or in combination and may be provided in any suitable container, e.g., vials, bottles, or tubes. In some embodiments, the kit includes instructions in one or more languages, e.g., two or more languages.
[0094] In some embodiments, the kit includes one or more reagents used in a process utilizing one or more of the elements described herein. The reagents may be provided in any suitable container. For example, the kit may provide one or more reaction or storage buffers. The reagents may be provided in a form usable in a particular assay, or in a form requiring the addition of one or more other components before use (e.g., in concentrate or lyophilized form). The buffer may be any buffer, including, but not limited to, sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In some embodiments, the buffer has a pH of about 7 to about 10. In some embodiments, the kit includes one or more oligonucleotides corresponding to the guide sequence and the guide sequence for insertion into a vector for operably binding the regulatory element. In some embodiments, the kit includes homologous recombination template polynucleotides.
[0095] In one embodiment, the present invention provides a method using one or more elements of the CRISPR system. The CRISPR complex of the present invention provides an effective means for modifying a target polynucleotide. The CRISPR complex of the present invention has broad utility, for example, modification of a target polynucleotide (e.g., deletion, insertion, translocation, inactivation, activation) in a very large number of cell types. Therefore, the CRISPR complex of the present invention has a wide range of applications, for example, in gene therapy, drug screening, disease diagnosis, and prognosis. An exemplary CRISPR complex comprises a CRISPR enzyme complexing with a guide sequence which hybridizes to a target sequence in the target polynucleotide. The guide sequence is then bound to a tracr-mate sequence which hybridizes to a tract sequence.
[0096] The target polynucleotide of the CRISPR complex can be any polynucleotide that is endogenous or exogenous to eukaryotic cells. For example, the target polynucleotide may be a polynucleotide that remains in the nucleus of a eukaryotic cell. The target polynucleotide may be a sequence that codes for a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). Although not constrained by theory, it is conceivable that the target sequence should associate with a PAM (protospacer adjacency motif); i.e., a short sequence recognized by the CRISPR complex. The exact sequence and length requirements for the PAM vary depending on the CRISPR enzyme used, but the PAM is typically a 2-5 base pair sequence adjacent to the protospacer (i.e., the target sequence). Examples of PAM sequences are given in the Examples section below, and those skilled in the art can identify further PAM sequences used for a given CRISPR enzyme.
[0097] Examples of target polynucleotides for the CRISPR complex include numerous disease-related genes and polynucleotides, as well as signaling biochemical pathway-related genes and polynucleotides, as listed in U.S. Provisional Patent Applications No. 61 / 736,527 and No. 61 / 748,427 (both titled SYSTEMS METHODS AND COMPOSITIONS FOR SEQUENCE MANIPULATION, filed on December 12, 2012 and January 2, 2013, respectively, with Broad reference numbers BI-2011 / 008 / WSGR 44063-701.101 and BI-2011 / 008 / WSGR 44063-701.102, respectively; all of the contents of these applications are incorporated herein by reference as a whole).
[0098] Examples of targeted polynucleotides include sequences related to signaling biochemical pathways, e.g., signaling biochemical pathway-related genes or polynucleotides. Examples of targeted polynucleotides include disease-related genes or polynucleotides. A “disease-related” gene or polynucleotide refers to any gene or polynucleotide that produces a transcription or translation product at abnormal levels or in abnormal forms in cells derived from affected tissue compared to non-disease control tissue or cells. This may be a gene that becomes expressed at abnormally high levels; or it may be a gene that becomes expressed at abnormally low levels, and the change in expression correlates with the onset and / or progression of the disease. Disease-related genes also refer to genes that directly carry the pathogenesis of the disease, or genes that have mutations or gene mutations that are in linkage disequilibrium with genes that carry the pathogenesis of the disease. The transcription or translation product may be known or unknown, and may be at normal or abnormal levels.
[0099] Examples of disease-related genes and polynucleotides are available from the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Md.), and are also available on the World Wide Web.
[0100] Examples of disease-related genes and polynucleotides are listed in Tables A and B. Disease-specific information is available from the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Md.), and is also available on the World Wide Web. Examples of signaling biochemical pathway-related genes and polynucleotides are listed in Table C.
[0101] Mutations in these genes and pathways can result in the production of inappropriate proteins or inappropriate amounts of proteins that affect their function. Further examples of genes, diseases, and proteins are incorporated herein by reference from U.S. Provisional Patent Applications No. 61 / 736,527 and No. 61 / 748,427. Such genes, proteins, and pathways may be target polynucleotides of the CRISPR complex.
[0102] [Table 1]
[0103] [Table 2]
[0104] [Table 3]
[0105] [Table 4]
[0106] [Table 5]
[0107] Table 6
[0108] Table 7
[0109] Table 8
[0110] Table 9
[0111] Table 10
[0112] Table 11
[0113] Table 12
[0114] Table 13
[0115] Table 14
[0116] Table 15
[0117] [Table 16]
[0118] Embodiments of the present invention also relate to methods and compositions related to gene knockout, gene amplification, and repair of specific mutations associated with DNA repeat instability and neurological diseases (Robert D. Wells, Tetsuo Ashizawa, Genetic Instabilities and Neurological Diseases, Second Edition, Academic Press, Oct 13, 2011-Medical). Tandem repeat sequences of a defined form have been found to be responsible for more than 20 human diseases (New insights into repeat instability: role of RNA·DNA hybrids. McIvor EI, Polak U, Napierala M. RNA Biol. 2010 Sep-Oct;7(5):551-8). These abnormalities of genomic instability can be corrected using the CRISPR-Cas system.
[0119] A further aspect of the present invention relates to the use of the CRISPR-Cas system for correcting abnormalities in the EMP2A and EMP2B genes, which have been identified as being associated with Lafora disease. Lafora disease is an autosomal recessive condition characterized by progressive myoclonic epilepsy, which can begin in adolescence as epileptic seizures. Several cases of this disease may be caused by mutations in genes that have not yet been identified. The disease causes seizures, muscle spasms, difficulty walking, dementia, and ultimately death. Currently, there is no treatment that has been proven effective against disease progression. Other gene abnormalities associated with epilepsy can also be targeted using the CRISPR-Cas system, and the underlying genetics are further described in Genetics of Epilepsy and Genetic Epilepsies, edited by Giuliano Avanzini and Jeffrey L. Noebels, Mariani Foundation Paediatric Neurology: 20; 2009).
[0120] In yet another aspect of the present invention, the CRISPR-Cas system can be used to correct ocular abnormalities resulting from several gene mutations, further described in Genetic Diseases of the Eye, Second Edition, edited by Elias I. Traboulsi, Oxford University Press, 2012.
[0121] Some further aspects of the present invention relate to the correction of abnormalities associated with a wide range of genetic disorders, which are further described on the website of the National Institutes of Health under the topic subsection Genetic Disorders. Genetic brain disorders include, but are not limited to, adrenoleukodystrophy, corpus callosum agenesis, Aicardi syndrome, Alpers disease, Alzheimer's disease, Barth syndrome, Batten disease, CADASIL, cerebellar degeneration, Fabry disease, Gerstmann-Streussler-Scheinker disease, Huntington's disease and other triplet repeat diseases, Leigh disease, Lesch-Nyhan syndrome, Menkes disease, mitochondrial myopathy, and NINDS colposephary. These disorders are further described on the website of the National Institutes of Health under the subsection Genetic Brain Disorders.
[0122] In some embodiments, the pathophysiology is neoplastic. In some embodiments where the pathophysiology may be neoplastic, the gene to be targeted may be one of those listed in Table A (in this case, PTEN, etc.). In some embodiments, the pathophysiology may be age-related macular degeneration. In some embodiments, the pathophysiology may be schizophrenia. In some embodiments, the pathophysiology may be trinucleotide repeat disorder. In some embodiments, the pathophysiology may be fragile X syndrome. In some embodiments, the pathophysiology may be secretase-related disorder. In some embodiments, the pathophysiology may be prion-related disorder. In some embodiments, the pathophysiology may be ALS. In some embodiments, the pathophysiology may be drug addiction. In some embodiments, the pathophysiology may be autism. In some embodiments, the pathophysiology may be Alzheimer's disease. In some embodiments, the pathophysiology may be inflammation. In some embodiments, the pathophysiology may be Parkinson's disease.
[0123] Examples of proteins associated with Parkinson's disease, though not limited to them, include α-synuclein, DJ-1, LRRK2, PINK1, parkin, UCHL1, synphyrin-1, and NURR1.
[0124] An example of a preference-related protein is ABAT.
[0125] Examples of inflammation-related proteins include monocyte chemotactic protein-1 (MCP1) encoded by the Ccr2 gene, CC chemokine receptor type 5 (CCR5) encoded by the Ccr5 gene, IgG receptor IIB (FCGR2b, also known as CD32) encoded by the Fcgr2b gene, or Fc epsilon R1g (FCER1g) protein encoded by the Fcer1g gene.
[0126] Examples of cardiovascular disease-related proteins include IL1B (interleukin-1, beta), XDH (xanthine dehydrogenase), TP53 (tumor protein p53), PTGIS (prostaglandin I2 (prostacyclin) synthase), MB (myoglobin), IL4 (interleukin-4), ANGPT1 (angiopoietin-1), ABCG8 (ATP-binding cassette, subfamily G (WHITE), member 8), or CTSK (cathepsin K).
[0127] Examples of Alzheimer's disease-related proteins include, for example, the very low-density lipoprotein receptor protein (VLDLR) encoded by the VLDLR gene, ubiquitin-like modifier activator enzyme 1 (UBA1) encoded by the UBA1 gene, or the NEDD8 activator enzyme E1 catalytic subunit protein (UBE1C) encoded by the UBA3 gene.
[0128] Examples of proteins associated with autism spectrum disorder include, for example, benzodiazepine receptor (peripheral)-associated protein 1 (BZRAP1) encoded by the BZRAP1 gene, AF4 / FMR2 family member 2 protein (AFF2) encoded by the AFF2 gene (also known as MFR2), fragile X autosomal homolog 1 protein (FXR1) encoded by the FXR1 gene, or fragile X autosomal homolog 2 protein (FXR2) encoded by the FXR2 gene.
[0129] Examples of proteins associated with macular degeneration include, for example, ATP-binding cassette subfamily A (ABC1) member 4 protein (ABCA4) encoded by the ABCR gene, apolipoprotein E protein (APOE) encoded by the APOE gene, or chemokine (CC motif) ligand 2 protein (CCL2) encoded by the CCL2 gene.
[0130] Examples of proteins associated with schizophrenia include NRG1, ErbB4, CPLX1, TPH1, TPH2, NRXN1, GSK3A, BDNF, DISC1, GSK3B, and combinations thereof.
[0131] Examples of proteins involved in tumor suppression include ATM (ataxia telangiectasia mutation), ATR (ataxia telangiectasia and Rad3-related), EGFR (epidermal growth factor receptor), ERBB2 (v-erb-b2 erythroblastic leukemia virus oncogene homolog 2), ERBB3 (v-erb-b2 erythroblastic leukemia virus oncogene homolog 3), ERBB4 (v-erb-b2 erythroblastic leukemia virus oncogene homolog 4), Notch1, Notch2, Notch3, or Notch4.
[0132] Examples of proteins associated with secretase disorders include PSENEN (presenilin enhancer 2 homolog (nematode (C. elegans))), CTSB (cathepsin B), PSEN1 (presenilin 1), APP (amyloid beta (A4) precursor protein), APH1B (prepharyngeal deficiency 1 homolog B (nematode (C. elegans))), PSEN2 (presenilin 2 (Alzheimer's disease 4)), or BACE1 (beta-site APP cleavage enzyme 1).
[0133] Examples of proteins associated with amyotrophic lateral sclerosis (ALS) include SOD1 (superoxide dismutase 1), ALS2 (amyotrophic lateral sclerosis 2), FUS (fused in sarcoma), TARDBP (TAR DNA-binding protein), VAGFA (vascular endothelial growth factor A), VAGFB (vascular endothelial growth factor B), and VAGFC (vascular endothelial growth factor C), as well as any combination thereof.
[0134] Examples of proteins associated with prion diseases include SOD1 (superoxide dismutase 1), ALS2 (amyotrophic lateral sclerosis 2), FUS (fused in sarcoma), TARDBP (TAR DNA-binding protein), VAGFA (vascular endothelial growth factor A), VAGFB (vascular endothelial growth factor B), and VAGFC (vascular endothelial growth factor C), and any combination thereof.
[0135] Examples of proteins associated with neurodegenerative pathologies in prion disorders include, for example, A2M (alpha-2 macroglobulin), AATF (anti-apoptotic transcription factor), ACPP (prostatic acid phosphatase), ACTA2 (aortic smooth muscle actin alpha-2), ADAM22 (ADAM metallopeptidase domain), ADORA3 (adenosine A3 receptor), or ADRA1D (alpha-1D adrenergic receptor for alpha-1D adrenergic receptor).
[0136] Examples of proteins associated with immunodeficiency include, for example, A2M [alpha-2-macroglobulin]; AANAT [arylalkylamine N-acetyltransferase]; ABCA1 [ATP-binding cassette subfamily A (ABC1), member 1]; ABCA2 [ATP-binding cassette subfamily A (ABC1), member 2]; or ABCA3 [ATP-binding cassette subfamily A (ABC1), member 3].
[0137] Examples of proteins associated with trinucleotide repeat disorders include, for example, AR (androgen receptor), FMR1 (fragility x intellectual disability 1), HTT (Huntington's protein kinase), or DMPK (myotonic dysplasia protein kinase), FXN (frataxin), and ATXN2 (ataxin 2).
[0138] Examples of proteins associated with neurotransmission disorders include, for example, SST (somatostatin), NOS1 (nitric oxide synthase 1 (neuronal type)), ADRA2A (adrenergic alpha-2A receptor), ADRA2C (adrenergic alpha-2C receptor), TACR1 (tachykinin receptor 1), or HTR2c (5-hydroxytryptamine (serotonin) receptor 2C).
[0139] Examples of neurodevelopment-related sequences include, for example, A2BP1 [ataxin 2-binding protein 1], AADAT [aminoadipate aminotransferase], AANAT [arylalkylamine N-acetyltransferase], ABAT [4-aminobutyrate aminotransferase], ABCA1 [ATP-binding cassette subfamily A (ABC1) member 1], or ABCA13 [ATP-binding cassette subfamily A (ABC1) member 13].
[0140] Further examples of preferred conditions treatable by the system of the present invention can be selected from the following: Alcardi-Gutierre syndrome; Alexander disease; Alan-Herndon-Dudley syndrome; POLG-related disorder; alpha-mannosidosis (types II and III); Alström syndrome; Angelman syndrome; ataxia with telangiectasia; neuronal ceroid lipofuscinosis; beta-cerasamia; bilateral optic atrophy and (infantile) optic atrophy type 1; retinoblastoma (bilateral); Canavan disease; cerebro-ocular-facial-skeletal syndrome 1 [COFS1]; cerebral tendon xanthomas Syndrome; Cornelia de Lang syndrome; MAPT-related disorders; hereditary prion diseases; Dravet syndrome; early-onset familial Alzheimer's disease; Friedreich's ataxia [FRDA]; Flins syndrome; fucosidosis; Fukuyama-type congenital muscular dystrophy; galactosialidosis; Gaucher disease; organic acidemia; hemophagocytic lymphohistiocytosis; Hutchinson-Gilford progeria syndrome; mucolipidosis type II; infantile free sialic acid storage; PLA2G6-related neurodegeneration; Javer-Lange-Nielsen syndrome; junctional epidermolysis bullosa; Huntington's disease; Krabbe disease (infant type) Mitochondrial DNA-related Lie syndrome and NARP; Lesch-Nyhan syndrome; LIS1-related lissencephaly; Lowe syndrome; Maple syrup urine disease; MECP2 duplication syndrome; ATP7A-related copper transport disorder; LAMA2-related muscular dystrophy; Arylsulfatase A deficiency; Mucopolysaccharidosis type I, II, or III; Peroxisome dysplasia, Zellweger syndrome spectrum; Neurodegenerative diseases with cerebral iron storage; Acid sphingomyelinase deficiency; Niemann-Pick disease type C; Glycine encephalopathy; ARX-related disorders; Urea cycle disorders; COL1A1 / 2-related disorders Osteogenesis imperfecta; Mitochondrial DNA deletion syndrome; PLP1-related disorder; Perry syndrome; Phelan-McDermott syndrome; Glycogen storage disorder type II (Pompe disease) (infant form); MAPT-related disorder; MECP2-related disorder; Root chondrodysplasia type 1; Roberts syndrome; Sandhoff disease; Schindler disease type 1; Adenosine deaminase deficiency; Smith-Lemley-Opitz syndrome; Spinal muscular atrophy; Infant-onset spinocerebellar ataxia; Hexosaminidase A deficiency; Fatal dysplasia type 1; Collagen type VI-related disorder; Usher syndrome type I; Congenital muscular dystrophy;Wolff-Hirschhorn syndrome; lysosomal acid lipase deficiency; and xeroderma pigmentosum.
[0141] Chronic administration of protein-based therapeutic agents can induce unacceptable immune responses to specific proteins. The immunogenicity of protein drugs may be due to several immunodominant helper T lymphocyte (HTL) epitopes. Reducing the MHC binding affinity of these HTL epitopes contained within these proteins can produce drugs with lower immunogenicity (Tangri S, et al. "Rationally engineered therapeutic proteins with reduced immunogenicity" J Immunol. 2005 Mar 15;174(6):3187-96). In this invention, the immunogenicity of CRISPR enzymes can be reduced, in particular with respect to erythropoietin, according to an approach first described by Tangri et al. and subsequently developed. Thus, the immunogenicity of CRISPR enzymes (e.g., Cas9) in a host species (human or other species) can be reduced using directed evolution or rational design.
[0142] In plants, pathogens are often host-specific. For example, Fusarium oxysporum f.sp. lycopersici causes tomato wilt but attacks only tomatoes, while Fusarium oxysporum f. dianthii and Puccinia graminis f.sp. tritici attack only wheat. Plants have pre-existing and inducible defenses to resist most pathogens. Mutation and recombination events across plant generations result in genetic mutations that create susceptibility, especially when pathogens reproduce more frequently than the plant. Non-host resistance can exist in plants, for example, when the host and pathogen are incompatible. Horizontal resistance, such as partial and vertical resistance to all species of a pathogen typically controlled by many genes, and complete resistance to some species of a pathogen but not others, typically controlled by a small number of genes, can also exist. At the gene-to-gene level, plants and pathogens evolve together, with genetic changes in one maintaining a balance with changes in others. Therefore, using natural variation, breeders combine genes most useful for yield, quality, uniformity, cold tolerance, and resistance. Resources for resistance genes include natural or introduced species, native species, wild plant relatives, and induced mutations, such as treating plant material with mutagenic agents. The present invention provides plant breeders with a novel tool for inducing mutations. Thus, those skilled in the art can analyze the genomes of resistance gene resources and use the present invention to induce the emergence of resistance genes in varieties with desired features or traits more accurately than conventional mutagenic agents, thereby accelerating and improving plant breeding programs.
[0143] As is evident, it is assumed that any target polynucleotide sequence can be targeted using the system of the present invention. Some examples of conditions or diseases that can be usefully treated using the system of the present invention are included in the table above, and examples of genes currently associated with those conditions are also provided in the table. However, the exemplified genes are not exclusive. [Examples]
[0144] The following examples are given for the purpose of illustrating various embodiments of the present invention, and are not intended to limit the invention to any particular form. These examples, along with the methods described herein, are representative and illustrative examples of currently preferred embodiments and are not intended to limit the scope of the invention. Modifications and other uses incorporated therein into the spirit of the invention as defined by the claims will be made by those skilled in the art.
[0145] Example 1: CRISPR complex activity in the nucleus of eukaryotic cells An exemplary type II CRISPR system is a type II CRISPR locus from Streptococcus pyogenes SF370 containing a cluster of four genes, Cas9, Cas1, Cas2, and Csn1, as well as a characteristic array of repetitive sequences (direct repeats) spaced by two non-coding RNA elements, tracrRNA, and short stretches of non-repetitive sequences (spacers, each approximately 30 bp). In this system, targeted DNA double-strand breaks (DSBs) are generated in four sequential steps (Figure 2A). First, two non-coding RNAs, a pre-crRNA array, and tracrRNA are transcribed from the CRISPR locus. Second, the tracrRNA hybridizes to the direct repeats of the pre-crRNA, which is then processed into mature crRNA containing individual spacer sequences. Thirdly, the mature crRNA:tracrRNA complex directs Cas9 to a DNA target consisting of the protospacer and the corresponding PAM via heteroduplex formation between the crRNA spacer region and the protospacer DNA. Finally, Cas9 creates a double-strand break (DSB) within the protospacer by mediating the cleavage of the target DNA upstream of the PAM (Figure 2A). This example describes an exemplary process for adapting this RNA programmable nuclease system to direct CRISPR complex activity in the nucleus of a eukaryotic cell.
[0146] To improve the expression of CRISPR components in mammalian cells, two genes from the Streptococcus pyogenes (S. pyogenes) SF370 locus 1, Cas9 (SpCas9) and RNase III (SpRNase III), were codon-optimized. To promote nuclear localization, a nuclear localization signal (NLS) was included at the amino (N) or carboxyl (C) terminus of both SpCas9 and SpRNase III (Figure 2B). To facilitate the visualization of protein expression, a fluorescent protein marker was also included at the N or C terminus of both proteins (Figure 2B). A version of SpCas9 with NLS attached to both the N and C termini (2×NLS-SpCas9) was also generated. Constructs containing NLS-fused SpCas9 and SpRNase III were transfused into 293FT human embryonic kidney (HEK) cells, and it was found that the relative positioning of NLS to SpCas9 and SpRNase III affected their nuclear localization efficiency. While C-terminal NLS was sufficient to target the target SpRNase III to the nucleus, attachment of a single copy of these specific NLS to either the N or C-terminus of SpCas9 failed to achieve adequate nuclear localization in this system. In this example, the C-terminal NLS was that of nucleoplasmin (KRPAATKKAGQAKKKK), and the C-terminal NLS was that of SV40 large T antigen (PKKKRKV). Of the versions of SpCas9 tested, only 2×NLS-SpCas9 showed nuclear localization (Figure 2B).
[0147] TracrRNA from the CRISPR locus of Streptococcus pyogenes SF370 has two transcription start sites, producing two transcripts of 89 nucleotides (nt) and 171 nt, which are subsequently processed into the same 75 nt mature tracrRNA. The shorter 89 nt tracrRNA was selected for expression in mammalian cells (expression construct described in Figure 6, with functionality determined by the Surveyor assay results shown in Figure 6B). The transcription start site is labeled as +1, and sequences probed by transcription terminators and Northern blot are also shown. Expression of processed tracrRNA was also confirmed by Northern blot. Figure 7C shows the results of Northern blot analysis of total RNA extracted from 293FT cells transfected with long or short tracrRNAs, as well as U6 expression constructs carrying SpCas9 and DR-EMX1(1)-DR. The left and right panels are from 293FT cells transfused with or without SpRNase III, respectively. U6 shows the loading control blotted with a probe targeting human U6 snRNA. Transfusion with the short tracrRNA expression construct resulted in a sufficient level of processed tracrRNA (approximately 75 bp). Very small amounts of long tracrRNA were detected on the Northern blot.
[0148] To promote accurate transcription initiation, we selected an RNA polymerase III-based U6 promoter to drive tracrRNA expression (Figure 2C). Similarly, we developed a U6 promoter-based construct to express a precrRNA array consisting of a single spacer flanked by two direct repeats (DRs, also encompassed in the term "tracrmate sequence"; Figure 2C). The first spacer was designed to target a 33-base pair (bp) target site in the human EMX1 locus, a key gene in cerebral cortex development (a 30-bp protospacer and a 3-bp CRISPR motif (PAM) sequence that satisfies the Cas9 NGG recognition motif) (Figure 2C).
[0149] To investigate whether heterologous expression of the CRISPR system (SpCas9, SpRNase III, tracrRNA, and precrRNA) in mammalian cells could achieve targeted mammalian chromosome cleavage, HEK293FT cells were transfused with combinations of CRISPR components. Since double-segment breaks (DSBs) in mammalian nuclei are partially repaired by the non-homologous end joining (NHEJ) pathway, which leads to indel formation, potential cleavage activity at the target EMX1 locus was detected using a Surveyor assay (see, e.g., Guschin et al., 2010, Methods Mol Biol 649:247). Simultaneous transfusion of all four CRISPR components was able to induce cleavage of up to 5.0% of protospacers (see Figure 2D). Simultaneous translocation of all CRISPR components except SpRNase III also induced up to 4.7% of indels in the protospacer, suggesting the presence of endogenous mammalian RNases, such as the associated Dicer and Drosha enzymes, that may support crRNA maturation. Removal of any of the remaining three components halted the genome-cleaving activity of the CRISPR system (Figure 2D). Sanger sequencing of amplicons containing the target locus confirmed the cleavage activity; five mutant alleles (11.6%) were found among 43 sequenced clones. Similar experiments using various guide sequences resulted in a higher indel rate of 29% (see Figures 4-8, 10, and 11). These results define a three-component system for efficient CRISPR-mediated genome modification in mammalian cells.
[0150] To optimize cleavage efficiency, the applicants also tested whether different isoforms of tracrRNA affect cleavage efficiency and found that in this exemplary system, only the short-chain (89 bp) transcript form could mediate cleavage of the human EMX1 genomic locus. Figure 9 provides an additional Northern blot analysis of crRNA processing in mammalian cells. Figure 9A illustrates a schematic diagram showing an expression vector for a single spacer (DR-EMX1(1)-DR) flanked by two direct repeats. The 30 bp spacer and direct repeat sequences targeting human EMX1 locus protospacer 1 are shown in the sequence below Figure 9A. Lines indicate regions that generate a Northern blot probe for EMX1(1) crRNA detection using the reverse complementary sequence. Figure 9B shows a Northern blot analysis of total RNA extracted from 293FT cells transfected with a U6 expression construct carrying DR-EMX1(1)-DR. The left and right panels are from 293FT cells transfused with or without SpRNase III, respectively. DR-EMX1(1)-DR was processed into mature crRNA only in the presence of SpCas9, and short tracrRNA was independent of the presence of SpRNase III. The mature crRNA detected from transfused 293FT total RNA was approximately 33 bp, shorter than the 39–42 bp mature crRNA from Streptococcus pyogenes. These results demonstrate that the CRISPR system can be transplanted into eukaryotic cells and reprogrammed to promote the cleavage of endogenous mammalian target polynucleotides.
[0151] Figure 2 illustrates the bacterial CRISPR system described in this embodiment. Figure 2A illustrates CRISPR locus 1 from Streptococcus pyogenes SF370 and a schematic diagram showing the proposed mechanism of CRISPR-mediated DNA cleavage by this system. Mature crRNA processed from a direct repeat-spacer array directs Cas9 to a genomic target consisting of a complementary protospacer and a protospacer-adjacent motif (PAM). During target-spacer base pairing, Cas9 mediates double-strand breaks in the target DNA. Figure 2B illustrates the engineering of Streptococcus pyogenes Cas9 (SpCas9) and RNase III (SpRNase III) with nuclear localization signals (NLS) to enable transport into the mammalian nucleus. Figure 2C illustrates the mammalian expression of SpCas9 and SpRNase III driven by the constitutive EF1a promoter, as well as tracrRNA and precrRNA arrays (DR-spacer-DR) driven by the RNAPol3 promoter U6 to promote accurate transcription initiation and termination. A protospacer from the human EMX1 locus with sufficient PAM sequences is used as a spacer in the precrRNA array. Figure 2D illustrates a surveyor nuclease assay for SpCas9-mediated minor insertions and deletions. SpCas9 was expressed with or without the precrRNA array carrying SpRNase III, tracrRNA, and the EMX1-targeting spacer. Figure 2E illustrates a schematic representation of base pairing between the target locus and the EMX1-targeting crRNA, as well as an exemplary chromatogram showing a microdeletion adjacent to the SpCas9 cleavage site. Figure 2F illustrates mutant alleles identified from sequencing analysis of 43 clonal amplicons exhibiting various microinsertions and deletions. Dotted lines indicate deleted bases; unaligned or mismatched bases indicate insertions or mutations. Scale bar = 10 μm.
[0152] To further simplify the three-component system, a chimeric crRNA-tracrRNA hybrid design was adapted, mimicking a natural crRNA:tracrRNA double-strand by fusing a mature crRNA (including the guide sequence) to a partial tracrRNA via a stem-loop (Figure 3A).
[0153] Guide sequences can be inserted between BbsI sites using annealed oligonucleotides. Protospacers on the sense and antisense strands are shown above and below the DNA sequence, respectively. Modification rates of 6.3% and 0.75% were achieved for the human PVALB and mouse Th loci, respectively, demonstrating the broad applicability of the CRISPR system to modification of different loci across multiple organisms. While cleavage was detected for only one of the three spacers for each locus using a chimeric construct, all target sequences were detected when using co-expressed precrRNA placement. The indels were cleaved with an efficiency of 27% (Figures 4 and 5).
[0154] Figure 5 provides further explanation of how SpCas9 can be reprogrammed to target multiple genomic loci in mammalian cells. Figure 5A provides a schematic diagram of the human EMX1 locus showing the localization of five protospacers indicated by underlined sequences. Figure 5B provides a schematic diagram of the precrRNA / trcrRNA complex (top) showing hybridization between direct repeat regions of precrRNA and tracrRNA, a 20 bp guide sequence, and a schematic diagram of the chimeric RNA design (bottom) including a tracrmate and tracr sequence consisting of a partial direct repeat and tracrRNA sequence hybridized into a hairpin structure. Figure 5C illustrates the results of a Surveyor assay comparing the efficacy of Cas9-mediated cleavage in the five protospacers in the human EMX1 locus. Each protospacer is targeted using either the processed precrRNA / tracrRNA complex (crRNA) or chimeric RNA (chiRNA).
[0155] Because RNA secondary structure can be important for intermolecular interactions, we compared the predicted secondary structures of all guide sequences used in our genome targeting experiments using a structure prediction algorithm based on minimum free energy and Boltzmann weighted structure ensemble (Figure 3B) (see, e.g., Gruber et al., 2008, Nucleic Acids Research, 36:W70). The analysis revealed that, in most cases, effective guide sequences in the chimeric crRNA context substantially lack secondary structure motifs, while ineffective guide sequences are more likely to form internal secondary structures that can interfere with base pairing with target protospacer DNA. Therefore, variability in spacer secondary structure may affect the efficiency of CRISPR-mediated interference when using chimeric crRNA.
[0156] Figure 3 illustrates exemplary expression vectors. Figure 3A provides a schematic diagram of a bicistronic vector for driving the expression of a synthetic crRNA-tracrRNA chimera (chimeric RNA) and SpCas9. The chimeric guide RNA contains a 20 bp guide sequence corresponding to a protospacer in the genomic target site. Figure 3B provides schematic diagrams of guide sequences targeting human EMX1, PVALB, and mouse Th loci, as well as their predicted secondary structures. The modification efficiency at each target site is shown below the RNA secondary structure diagram (EMX1, n=216 amplicon sequencing reads; PVALB, n=224 reads; Th, n=265 reads). The folding algorithm produced an output colored according to the probability that each base assumes a predicted secondary structure, as shown by the rainbow scale reproduced in grayscale in Figure 3B. Further vector designs for SpCas9 are shown in Figure 3A and include a single expression vector incorporating a U6 promoter bound to the insertion site for the guide oligo and a Cbh promoter bound to the SpCas9 coding sequence.
[0157] To investigate whether spacers containing secondary structures can function in prokaryotic cells where CRISPR is naturally active, transformation interference with protospacer-supported plasmids was tested in Escherichia coli (E. coli) strains heterologously expressing the Streptococcus pyogenes (S. pyogenes) SF370 CRISPR locus 1 (Figure 3C). The CRISPR locus was cloned into a low-copy E. coli (E. coli) expression vector, and the crRNA array was replaced with a single spacer flanked by a DR pair (pCRISPR). E. coli (E. coli) strains possessing different pCRISPR plasmids were transformed with challenge plasmids containing the corresponding protospacer and PAM sequence (Figure 3C). In the bacterial assay, all spacers promoted efficient CRISPR interference (Figure 3C). These results suggest that additional factors may exist that influence the efficiency of CRISPR activity in mammalian cells.
[0158] To investigate the specificity of CRISPR-mediated cleavage, the effect of single nucleotide mutations in guide sequences on protospacer cleavage in mammalian genomes was analyzed using a series of EMX1-targeting chimeric crRNAs with single point mutations (Figure 4A). Figure 4B illustrates the results of a Surveyor nuclease assay comparing the cleavage efficiency of Cas9 when paired with different mutant chimeric RNAs. Single nucleotide mismatches up to 12 bp on the 5' end of the PAM substantially halted genome cleavage by SpCas9, while spacers with mutations further upstream activated the original protospacer target. The properties were preserved (Figure 4B). In addition to PAM, SpCas9 has single-base specificity within the last 12 bp of the spacer. Furthermore, CRISPR can mediate genome cleavage as efficiently as a pair of TALE nucleases (TALENs) targeting the same EMX1 protospacer. Figure 4C provides a schematic diagram showing the design of a TALEN targeting EMX1, and Figure 4D shows a Surveyor gel comparing the efficiency of TALEN and Cas9 (n=3).
[0159] We established a set of components to achieve CRISPR-mediated gene editing in mammalian cells via the error-prone NHEJ mechanism, and tested CRISPR's ability to stimulate homologous recombination (HR), a high-fidelity gene repair pathway for precise editing in the genome. Wild-type SpCas9 can mediate site-specific double-segment breaks that can be repaired through both NHEJ and HR. Furthermore, we engineered the aspartate-to-alanine substitution (D10A) in the RuvC I catalytic domain of SpCas9 to convert the nuclease into a nickase (SpCas9n; explained in Figure 5A) (see, e.g., Sapranausaks et al., 2011, Cucleic Acids Research, 39:9275; Gasiunas et al., 2012, Proc. Natl. Acad. Sci. USA, 109:E2579), resulting in nicked genomic DNA undergoing high-fidelity homologous recombination repair (HDR). Surveyor assays confirmed that SpCas9n does not generate indels at the EMX1 protospacer target. As illustrated in Figure 5B, co-expression of EMX1-targeting chimeric crRNA with SpCas9 induced indels at the target site, while co-expression with SpCas9n did not (n=3). Furthermore, sequencing of 327 amplicons did not detect any indels induced by SpCas9n. CRISPR-mediated HR was tested by selecting the same locus and co-transferring HEK293FT cells with an HR template to introduce EMX1-targeting chimeric RNA, hSpCas9 or hSpCas9n, and a pair of restriction sites (HindIII and NheI) near the protospacer. Figure 5C provides a schematic explanation of the HR strategy, along with the relative localization of the recombination site and primer annealing sequences (arrows). SpCas9 and SpCas9n actually catalyzed the integration of the HR template into the EMX1 gene.PCR amplification of the target region followed by restriction digestion with HindIII revealed cleavage products corresponding to predicted fragment sizes (arrows in the restriction fragment length polymorphism gel analysis shown in Figure 5D), with SpCas9 and SpCas9n mediating similar levels of HR efficiency. The applicant further confirmed HR using Sanger sequencing of genomic amplicons (Figure 5E). These results demonstrate the utility of CRISPR for facilitating targeted gene insertions in mammalian genomes. Given the 14 bp target specificity of wild-type SpCas9 (12 bp from the spacer and 2 bp from the PAM), the availability of nickase significantly reduces the possibility of off-target modifications because the single-strand degradation products are not substrates for the error-prone NHEJ pathway.
[0160] We tested the potential of multiplexed sequence targeting by constructing an expression construct (Figure 2A) that mimics the native architecture of CRISPR loci with array spacers. Using a single CRISPR array encoding a pair of EMX1 and PVALB targeting spacers, efficient cleavage was detected at both loci (Figure 4F, showing both a schematic design of the crRNA array and a Surveyor blot demonstrating efficient cleavage mediation). Targeted deletions of larger genomic regions via simultaneous DSBs using spacers for two targets within EMX1, spaced 119 bp apart, were also tested, and a deletion efficacy of 1.6% (3 out of 182 amplicons; Figure 5G) was detected. This demonstrates that the CRISPR system can mediate multiplexed editing within a single genome.
[0161] Example 2: Modified and alternative CRISPR systems The skill of using RNA to program sequence-specific DNA cleavage defines a new class of genome engineering tools for various research and industrial applications. Several embodiments of CRISPR systems can be further improved to increase the efficiency and versatility of CRISPR targeting. Optimal Cas9 activity may depend on the availability of free Mg2+ at levels higher than those present in mammalian nuclei (see, e.g., Jinek et al., 2012, Science, 337:816), and the preference for NGG motifs immediately downstream of the protospacer limits targeting ability by an average of 12 bp in the human genome. Some of these constraints can be overcome by leveraging the diversity of CRISPR loci across microbial metagenomes (see, e.g., Makarova et al., 2011, Nat Rev Microbiol, 9:467). Other CRISPR loci can be transplanted into the mammalian cell environment in a manner similar to that described in Example 1. The modification efficiency at each target site is shown below the RNA secondary structure. The algorithm that generates this structure colors each base according to its probability of assuming the predicted secondary structure. RNA guide spacers 1 and 2 induced 14% and 6.4%, respectively. A statistical analysis of cleavage activity across biological replicas at these two protospacer sites is also provided in Figure 7.
[0162] Example 3: Sample Target Sequence Selection Algorithm Design a software program to identify candidate CRISPR target sequences on both strands of an input DNA sequence based on a desired guide sequence length and CRISPR motif sequence (PAM) for a given CRISPR enzyme. For example, the target site for Cas9 from Streptococcus pyogenes can be identified by searching for 5'-Nx-NGG-3' on both the input sequence and its reverse complementary strand using the PAM sequence NGG. Similarly, the target site for Cas9 from S. thermophilus CRISPR1 can be identified by searching for 5'-Nx-NNAGAAW-3' on both the input sequence and its reverse complementary strand using the PAM sequence NNAGAAW. Similarly, the target site for Cas9 in S. thermophilus CRISPR3 can be identified by searching for 5'-Nx-NGGNG-3' on both the input sequence and the reverse complementary strand of the input using the PAM sequence NGGNG. The value "x" in Nx can be fixed by the program or specified by the user, for example, 20.
[0163] Because multiple occurrences of DNA target sites in the genome can lead to nonspecific genome editing, after identifying all potential sites, the program filters out sequences based on the number of times they appear in the relevant reference genome. For CRISPR enzymes where sequence specificity is determined by the 11-12 bp 5' end of the PAM sequence, including the PAM sequence itself, the filtering step may be based on the seed sequence. Therefore, to avoid editing at additional genomic loci, the results are filtered based on the number of occurrences of the seed:PAM sequence in the relevant genome. The user can be allowed to select the length of the seed sequence. The user can also be allowed to specify the number of occurrences of the seed:PAM sequence in the genome for the purpose of passing through the filter. The default is to screen for unique sequences. The filtration level is changed by changing both the length of the seed sequence and the number of occurrences of the sequence in the genome. The program may further or alternatively provide a guide sequence complementary to the reported target sequence by providing the inverse complementary strand of the identified target sequence.
[0164] Further details of methods and algorithms for optimizing sequence selection can be found in U.S. Patent Application No. TBA (Broad Reference No. BI-2012 / 084 44790.11.2022), which is incorporated herein by reference.
[0165] Example 4: Evaluation of multiple chimeric crRNA-tracrRNA hybrids This example describes the results obtained for chimeric RNA (chiRNA; containing a guide sequence, a tracr mate sequence, and a tracr sequence in a single transcript) having a tracr sequence that incorporates wild-type tracrRNA sequences of different lengths. Figure 18a illustrates a schematic diagram of the bicistronic expression vector for the chimeric RNA and Cas9. Cas9 is driven by the CBh promoter, and the chimeric RNA is driven by the U6 promoter. The chimeric guide RNA consists of a 20 bp guide sequence (N) bound to the truncated tracr sequence (extending from the first "U" of the lower strand to the end of the transcript) at various positions shown. The guide and tracr sequences are separated by the tracr mate sequence GUUUUAGAGCUA followed by the loop sequence GAAA. The results of SURVEYOR assays for Cas9-mediated indels at the human loci EMX1 and PVALB are described in Figures 18b and 18c, respectively. Arrows indicate predicted SURVEYOR fragments. chiRNAs are indicated by their "+n" notation, and crRNAs refer to hybrid RNAs in which the guide and tracr sequences are expressed as separate transcripts. Quantification of these results performed in triplicates is shown as histograms in Figures 11a and 11b, corresponding to Figures 10b and 10c, respectively ("ND" indicates that no indel was detected). Table D provides the protospacer IDs and their corresponding genomic targets, protospacer sequences, PAM sequences, and chain localizations. The guide sequences were designed to be complementary to the entire protospacer sequence in the case of separate transcripts in the hybrid system, or to be complementary only to the underlined portion in the case of chimeric RNA.
[0166] [Table 17]
[0167] Cell culture and translocation Human embryonic kidney (HEK) cell line 293FT (Life Technologies) was maintained at 37°C under 5% CO2 incubation in Dulbecco's Modified Eagle Medium (DMEM) supplemented with 10% fetal bovine serum (HyClone), 2 mM GlutaMAX (Life Technologies), 100 U / mL penicillin, and 100 μg / mL streptomycin. 293FT cells were seeded onto 24-well plates (Corning) at a density of 150,000 cells per well 24 hours prior to translocation. Cells were translocated using Lipofectamine 2000 (Life Technologies) according to the manufacturer's recommended protocol. A total of 500 ng of plasmid was used for each well of the 24-well plate.
[0168] SURVEYOR assay for genome modification 293FT cells were transfused with the plasmid DNA described above. The cells were incubated at 37°C for 72 hours after transfusion, and then genomic DNA was extracted. Genomic DNA was extracted using QuickExtract DNA Extraction Solution (Epicentre) according to the manufacturer's protocol. Briefly, pelletized cells were resuspended in QuickExtract solution and incubated at 65°C for 15 minutes and at 98°C for 10 minutes. Genomic regions flanking the CRISPR target sites for each gene were amplified by PCR (using the primers listed in Table E), and the products were purified using QiaQuick Spin Column (Qiagen) according to the manufacturer's protocol. A total of 400 ng of purified PCR product was mixed with 2 μl of 10× Taq DNA Polymerase PCR buffer (Enzymatics), and the mixture was diluted with ultrapure water to a final volume of 20 μl. The mixture was then subjected to a re-nealing process to enable heteroduplex formation: 95°C for 10 minutes, followed by a gradient of 95°C to 85°C at -2°C / sec, 85°C to 25°C at -0.25°C / sec, and finally maintained at 25°C for 1 minute. After re-nealing, the product was treated with SURVEYOR nuclease and SURVEYOR enhancer S (Transgenomics) according to the manufacturer's recommended protocol and analyzed on 4–20% Novex TBE polyacrylamide gel (Life Technologies). The gel was stained with SYBR Gold DNA stain (Life Technologies) for 30 minutes and imaged using a Gel Doc gel imaging system (Bio-rad). Quantification was based on relative band intensity.
[0169] [Table 18]
[0170] Computer-aided identification of unique CRISPR target sites To identify unique target sites for the Streptococcus pyogenes (S. pyogenes) SF370Cas9 (SpCas9) enzyme in human, mouse, rat, zebrafish, fruit fly, and nematode (C. elegans) genomes, the applicants developed a software package to scan both strands of DNA sequences and identify all possible SpCas9 target sites. In this example, each SpCas9 target site was operationally defined as a 20 bp sequence followed by an NGG protospacer adjacent motif (PAM) sequence, and the applicants identified all sequences on all chromosomes that satisfy this 5'-N20-NGG-3' definition. To prevent nonspecific genome editing, after identifying all potential sites, all target sites were filtered based on the number of times they appear in the associated reference genome. For example, to utilize the sequence specificity of Cas9 activity conferred by a "seed" sequence, which can be approximately 11-12 bp from the 5' end of the PAM sequence, the 5'-NNNNNNNNNN-NGG-3' sequence was selected as unique within the relevant genome. All genome sequences were downloaded from the UCSC Genome Browser (human genome hg19, mouse genome mm9, rat genome rn5, zebrafish genome danRer7, Drosophila melanogaster genome dm4, and C. elegans genome ce10). The complete search results are available for viewing using UCSC Genome Browser information. Exemplary visualizations of some target sites in the human genome are provided in Figure 22.
[0171] First, we targeted three sites within the EMX1 locus in human HEK293FT cells. The genomic modification efficiency of each chiRNA was evaluated using the SURVEYOR nuclease assay, which detects mutations resulting from subsequent repair by the DNA double-strand break (DSB) and non-homologous end-joining (NHEJ) DNA damage repair pathways. Constructs denoted as chiRNA(+n) indicate that the chimeric RNA construct contains up to +n nucleotides of wild-type tracrRNA, with values of 48, 54, 67, and 85 used for n. Chimeric RNAs containing longer fragments of wild-type tracrRNA (chiRNA(+67) and chiRNA(+85)) mediated DNA cleavage at all three EMX1 target sites, with chiRNA(+85) in particular demonstrating significantly higher levels of DNA cleavage than the corresponding crRNA / tracrRNA hybrids expressing the guide and tracr sequences in separate transcripts (Figures 10b and 10a). Two sites within the PVALB locus that did not exhibit detectable cleavage in the hybrid system (guide and tracr sequences expressed as separate transcripts) were also targeted using chiRNA. chiRNA(+67) and chiRNA(+85) were able to mediate significant cleavage in the two PVALB protospacers (Figures 10c and 10b).
[0172] A consistent increase in genome modification efficiency was observed with increasing tracr sequence length for all five targets in the EMX1 and PVALB loci. While not constrained by any theory, the secondary structure formed by the 3' end of the tracrRNA may play a role in improving the rate of CRISPR complex formation. A description of the predicted secondary structure for each of the chimeric RNAs used in this embodiment is provided in Figure 21. The secondary structures were predicted using RNAfold (http: / / rna.tbi.univie.ac.at / cgi-bin / RNAfold.cgi), which uses the minimum free energy and partition function algorithm. Pseudocolor (reproduced in grayscale) for each base indicates the probability of pair formation. Since chiRNAs with longer tracr sequences were able to cleave targets that were not cleaved by the natural CRISPRcrRNA / tracrRNA hybrid, it is conceivable that chimeric RNAs can be loaded onto Cas9 more efficiently than their natural hybrid equivalents. To facilitate the application of Cas9 for site-directed genome editing in eukaryotic cells and organisms, all predicted unique target sites for Streptococcus pyogenes (S. pyogenes) Cas9 were computer-aided in the genomes of humans, mice, rats, zebrafish, nematodes (C. elegans), and Drosophila melanogaster (D. melanogaster). Chimeric RNAs can be designed for Cas9 enzymes from other microorganisms to expand the target space for CRISPR RNA programmable nucleases.
[0173] Figures 11 and 21 illustrate an exemplary bicistronic expression vector for the expression of a chimeric RNA containing a wild-type tracrRNA sequence of up to +85 nucleotides and SpCas9 with a nuclear localization sequence. SpCas9 is expressed from the CBh promoter and terminated by a bGH polyA signal (bGHpA). The enlarged sequence described directly below the schematic diagram corresponds to the region surrounding the guide sequence insertion site and includes the 3' portion of the U6 promoter from 5' to 3' (first shaded region), the BbsI cleavage site (arrow), a partial direct repeat (tracr mate sequence GTTTTAGAGCTA, underlined), the loop sequence GAAA, and the +85 tracr sequence (underlined sequence after the loop sequence). An exemplary guide sequence insert is described below the guide sequence insertion site, with the nucleotides of the guide sequence for the selected target represented by "N".
[0174] The sequence described in the above example is as follows (the polynucleotide sequence is from 5' to 3').
[0175] U6-short tracrRNA (Streptococcus pyogenes SF370): [ka]
[0176] U6-long tracrRNA (Streptococcus pyogenes SF370): [ka]
[0177] U6-DR-BbsI skeleton-DR (Streptococcus pyogenes SF370): [ka]
[0178] U6-chimeric RNA-BbsI skeleton (Streptococcus pyogenes SF370) [ka]
[0179] NLS-SpCas9-EGFP: [ka]
[0180] SpCas9-EGFP-NLS: [ka]
[0181] NLS-SpCas9-EGFP-NLS: [ka]
[0182] NLS-SpCas9-NLS: [ka]
[0183] NLS-mCherry-SpRNase 3: [ka]
[0184] SpRNase 3-mCherry-NLS: [ka]
[0185] NLS-SpCas9n-NLS (D10A nickase mutation is lowercase): [ka]
[0186] hEMX1-HR template - HindII-NheI: [ka] [ka]
[0187] NLS-StCsn1-NLS: [ka]
[0188] U6-St_tracrRNA(7~97): [ka]
[0189] U6-DR-Spacer-DR (Streptococcus pyogenes SF370) [ka]
[0190] +48 tracrRNA-containing chimeric RNA (Streptococcus pyogenes SF370) [ka]
[0191] +54 tracrRNA-containing chimeric RNA (Streptococcus pyogenes SF370) [ka]
[0192] Chimeric RNA containing +67 tracrRNA (Streptococcus pyogenes (S.pyogenes) SF370)
Chem.
[0193] Chimeric RNA containing +85 tracrRNA (Streptococcus pyogenes (S.pyogenes) SF370)
Chem.
[0194] CBh-NLS-SpCas9-NLS
Chem.
Chem.
Chem.
Chem.
[0195] Exemplary chimeric RNA for S. thermophilus (S.thermophilus) LMD-9 CRISPR1 Cas9 (in the case of PAM of NNAGAAW)
Chem.
[0196] Exemplary chimeric RNA for S. thermophilus (S.thermophilus) LMD-9 CRISPR1 Cas9 (in the case of PAM of NNAGAAW)
Chem.
[0197] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1 Cas9 (for PAM of NNAGAAW)
Chem.
[0198] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1 Cas9 (for PAM of NNAGAAW)
Chem.
[0199] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1 Cas9 (for PAM of NNAGAAW)
Chem.
[0200] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1 Cas9 (for PAM of NNAGAAW)
Chem.
[0201] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1 Cas9 (for PAM of NNAGAAW)
Chem.
[0202] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1 Cas9 (for PAM of NNAGAAW)
Chem.
[0203] Exemplary chimeric RNA for S. thermophilus LMD-9CRISPR1Cas9 (in the case of PAM of NNAGAAW) [ka]
[0204] Exemplary chimeric RNA for S. thermophilus LMD-9CRISPR3Cas9 (in the case of PAM of NGGNG) [ka]
[0205] A codon-optimized version of Cas9 from the S. thermophilus LMD-9 CRISPR3 locus (with NLS at both the 5' and 3' ends). [ka] [ka] [ka] [ka]
[0206] Example 5: Optimization of guide RNA for Streptococcus pyogenes Cas9 (referred to as SpCas9) The applicants improved the RNA in cells by mutating tracrRNA and direct repeat sequences, or by mutating chimeric guide RNA.
[0207] The optimization is based on the observation that thymine stretches (T) were present in the tracrRNA and guide RNA, which can lead to early transcription termination by the pol3 promoter. Therefore, we generated the following optimized sequences. The optimized tracrRNA and the corresponding optimized direct repeat are represented as pairs.
[0208] Optimized tracrRNA1 (underlined indicates mutation): [ka]
[0209] Optimized Direct Repeat 1 (Underlined indicates mutation): [ka]
[0210] Optimized tracrRNA2 (underlined indicates mutation): [ka]
[0211] Optimized Direct Repeat 2 (Underlined indicates mutation): [ka]
[0212] The applicants also optimized the chimeric guide RNA for optimal activity in eukaryotic cells.
[0213] Original guide RNA: [ka]
[0214] Optimized Chimeric Guide RNA Sequence 1: [ka]
[0215] Optimized Chimeric Guide RNA Sequence 2: [ka]
[0216] Optimized Chimeric Guide RNA Sequence 3: [ka]
[0217] The applicants demonstrated that the optimized chimeric guide RNA functions better, as shown in Figure 3. This experiment was performed by simultaneously translocating 293FT cells using Cas9 and U6 guide RNA DNA cassettes to express one of the four RNA forms described above. The guide RNA targets the same target site in the human Emx1 gene locus: "GTCACCTCCAATGACTAGGG".
[0218] Example 6: Optimization of Streptococcus thermophilus LMD-9 CRISPR1 Cas9 (referred to as St1Cas9) The applicants designed the guide chimeric RNA shown in Figure 12.
[0219] St1Cas9 guide RNA can undergo the same type of optimization with respect to SpCas9 guide RNA by degrading the polythymine stretch (T).
[0220] Example 7: Improvement of Cas9 system for in vivo applications The applicants conducted a metagenomic search for Cas9 with low molecular weight. Most Cas9 homologs are quite large. For example, SpCas9 is approximately 1368 aa long, which is too large to be easily packaged in viral vectors for delivery. Some of the sequences are misannotated, and therefore the exact frequency for each length is not necessarily accurate. Nevertheless, this provides clues in the distribution of Cas9 proteins and suggests the existence of shorter Cas9 homologs.
[0221] Through computational analysis, the applicants discovered the existence of two Cas9 proteins with fewer than 1000 amino acids in bacterial strains of the genus Campylobacter. The sequence for one Cas9 from Campylobacter jejuni is presented below. At this length, CjCas9 can be readily packaged into AAV, lentivirus, adenovirus, and other viral vectors for robust in vivo delivery into primary cells and in animal models.
[0222] Campylobacter jejuni Cas9 (CjCas9) [ka]
[0223] The putative tracrRNA element for this CjCas9 is as follows: [ka]
[0224] The direct repeat sequence is as follows: [ka]
[0225] Figure 6 shows the simultaneous fold structure of tracrRNA and direct repeat.
[0226] An example of a chimeric guide RNA for CjCas9 is as follows: [ka]
[0227] The applicants also optimized the Cas9 guide RNA using an in vitro method. Figure 18 shows data from the in vitro optimization of the St1Cas9 chimeric guide RNA.
[0228] Preferred embodiments of the present invention have been shown and described herein, but it will be apparent to those skilled in the art that such embodiments are provided only as examples. Those skilled in the art can now make numerous variations, modifications, and substitutions without departing from the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein can be used in the practice of the present invention. The following claims define the scope of the present invention, and methods and structures within the scope of those claims, as well as their equivalents, are encompassed by those claims.
[0229] Example 8: Sa sgRNA optimization The applicants designed five sgRNA variants for SaCas9 for the optimal truncation architecture with maximum cleavage efficiency. Furthermore, the native direct repeat:tracr double-strand system was tested in parallel with the sgRNA. Guides of the indicated lengths were co-transplanted with SaCas9 and tested for activity in HEK293FT cells. A total of 100 ng of sgRNA U6-PCR amplicon (or 50 ng of direct repeat and 50 ng of tracrRNA) and 400 ng of SaCas9 plasmid were co-transplanted into 200,000 Hepa1-6 mouse hepatocytes, and DNA was recovered 72 hours after translocation for SURVEYOR analysis. The results are shown in Figure 23.
[0230] References: 1. Urnov, F.D., Rebar, E.J., Holmes, M.C., Zhang, H.S. & Gregory, P.D. Genome editing with engineered zinc finger nucleases. Nat. Rev. Genet. 11, 636 - 646 (2010). 2. Bogdanove, A.J. & Voytas, D.F. TAL effectors: customizable proteins for DNA targeting. Science 333, 1843 - 1846 (2011). 3. Stoddard, B.L. Homing endonuclease structure and function. Q. Rev. Biophys. 38, 49 - 95 (2005). 4. Bae, T. & Schneewind, O. Allelic replacement in Staphylococcus aureus with inducible counter - selection. Plasmid 55, 58 - 63 (2006). 5. Sung, C.K., Li, H., Claverys, J.P. & Morrison, D.A. An rpsL cassette, janus, for gene replacement through negative selection in Streptococcus pneumoniae. Appl. Environ. Microbiol. 67, 5190 - 5196 (2001). 6. Sharan, S.K., Thomason, L.C., Kuznetsov, S.G. & Court, D.L. Recombineering: a homologous recombination - based method of genetic engineering. Nat. Protoc. 4, 206 - 223 (2009). 7.Jinek,M.et al.A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.Science 337,816-821(2012). 8.Deveau,H.,Garneau,J.E.&Moineau,S.CRISPR-Cas system and its role in phage-bacteria interactions.Annu.Rev.Microbiol.64,475-493(2010). 9.Horvath,P.&Barrangou,R.CRISPR-Cas,the immune system of bacteria and archaea.Science 327,167-170(2010). 10.Terns,M.P.&Terns,R.M.CRISPR-based adaptive immune systems.Curr.Opin.Microbiol.14,321-327(2011). 11.van der Oost,J.,Jore,M.M.,Westra,E.R.,Lundgren,M.&Brouns,S.J.CRISPR-based adaptive and heritable immunity in prokaryotes.Trends.Biochem.Sci.34,401-407(2009). 12.Brouns,S.J.et al.Small CRISPR RNAs guide antiviral defense in prokaryotes.Science 321,960-964(2008). 13.Carte,J.,Wang,R.,Li,H.,Terns,R.M.&Terns,M.P.Cas6 is an endoribonuclease that generates guide RNAs for invader defense in prokaryotes.Genes Dev.22,3489-3496(2008). 14.Deltcheva,E.et al.CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.Nature 471,602-607(2011). 15.Hatoum-Aslan,A.,Maniv,I.&Marraffini,L.A.Mature clustered,regularly interspaced,short palindromic repeats RNA(crRNA)length is measured by a ruler mechanism anchored at the precursor processing site.Proc.Natl.Acad.Sci.U.S.A.108,21218-21222(2011). 16.Haurwitz,R.E.,Jinek,M.,Wiedenheft,B.,Zhou,K.&Doudna,J.A.Sequence- and structure-specific RNA processing by a CRISPR endonuclease.Science 329,1355-1358(2010). 17.Deveau,H.et al.Phage response to CRISPR-encoded resistance in Streptococcus thermophilus.J.Bacteriol.190,1390-1400(2008). 18.Gasiunas,G.,Barrangou,R.,Horvath,P.&Siksnys,V.Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria.Proc.Natl.Acad.Sci.U.S.A.(2012). 19.Makarova,K.S.,Aravind,L.,Wolf,Y.I.&Koonin,E.V.Unification of Cas protein families and a simple scenario for the origin and evolution of CRISPR-Cas systems.Biol.Direct.6,38(2011). 20.Barrangou,R.RNA-mediated programmable DNA cleavage.Nat.Biotechnol.30,836-838(2012). 21.Brouns,S.J.Molecular biology.A Swiss army knife of immunity.Science 337,808-809(2012). 22.Carroll,D.A CRISPR Approach to Gene Targeting.Mol.Ther.20,1658-1660(2012). 23.Bikard,D.,Hatoum-Aslan,A.,Mucida,D.&Marraffini,L.A.CRISPR interference can prevent natural transformation and virulence acquisition during in vivo bacterial infection.Cell Host Microbe 12,177-186(2012). 24.Sapranauskas,R.et al.The Streptococcus thermophilus CRISPR-Cas system provides immunity in Escherichia coli.Nucleic Acids Res.(2011). 25.Semenova,E.et al.Interference by clustered regularly interspaced short palindromic repeat(CRISPR)RNA is governed by a seed sequence.Proc.Natl.Acad.Sci.U.S.A.(2011). 26.Wiedenheft,B.et al.RNA-guided complex from a bacterial immune system enhances target recognition through seed sequence interactions.Proc.Natl.Acad.Sci.U.S.A.(2011). 27.Zahner,D.&Hakenbeck,R.The Streptococcus pneumoniae beta-galactosidase is a surface protein.J.Bacteriol.182,5919-5921(2000). 28.Marraffini,L.A.,Dedent,A.C.&Schneewind,O.Sortases and the art of anchoring proteins to the envelopes of gram-positive bacteria.Microbiol.Mol.Biol.Rev.70,192-221(2006). 29.Motamedi,M.R.,Szigety,S.K.&Rosenberg,S.M.Double-strand-break repair recombination in Escherichia coli:physical evidence for a DNA replication mechanism in vivo.Genes Dev.13,2889-2903(1999). 30.Hosaka,T.et al.The novel mutation K87E in ribosomal protein S12 enhances protein synthesis activity during the late growth phase in Escherichia coli.Mol.Genet.Genomics 271,317-324(2004). 31.Costantino,N.&Court,D.L.Enhanced levels of lambda Red-mediated recombinants in mismatch repair mutants.Proc.Natl.Acad.Sci.U.S.A.100,15748-15753(2003). 32.Edgar,R.&Qimron,U.The Escherichia coli CRISPR system protects from lambda lysogenization,lysogens,and prophage induction.J.Bacteriol.192,6291-6294(2010). 33.Marraffini,L.A.&Sontheimer,E.J.Self versus non-self discrimination during CRISPR RNA-directed immunity.Nature 463,568-571(2010). 34.Fischer,S.et al.An archaeal immune system can detect multiple Protospacer Adjacent Motifs(PAMs)to target invader DNA.J.Biol.Chem.287,33351-33363(2012). 35.Gudbergsdottir,S.et al.Dynamic properties of the Sulfolobus CRISPR-Cas and CRISPR / Cmr systems when challenged with vector-borne viral and plasmid genes and protospacers.Mol.Microbiol.79,35-49(2011). 36.Wang,H.H.et al.Genome-scale promoter engineering by coselection MAGE.Nat Methods 9,591-593(2012). 37.Cong,L.et al.Multiplex Genome Engineering Using CRISPR-Cas Systems.Science In press(2013). 38.Mali,P.et al.RNA-Guided Human Genome Engineering via Cas9.Science In press(2013). 39.Hoskins,J.et al.Genome of the bacterium Streptococcus pneumoniae strain R6.J.Bacteriol.183,5709-5717(2001). 40.Havarstein,L.S.,Coomaraswamy,G.&Morrison,D.A.An unmodified heptadecapeptide pheromone induces competence for genetic transformation in Streptococcus pneumoniae.Proc.Natl.Acad.Sci.U.S.A.92,11140-11144(1995). 41.Horinouchi,S.&Weisblum,B.Nucleotide sequence and functional map of pC194,a plasmid that specifies inducible chloramphenicol resistance.J.Bacteriol.150,815-825(1982). 42.Horton,R.M.In Vitro Recombination and Mutagenesis of DNA:SOEing Together Tailor-Made Genes.Methods Mol.Biol.15,251-261(1993). 43.Podbielski,A.,Spellerberg,B.,Woischnik,M.,Pohl,B.&Lutticken,R.Novel series of plasmid vectors for gene inactivation and expression analysis in group A streptococci(GAS).Gene 177,137-147(1996). 44.Husmann,L.K.,Scott,J.R.,Lindahl,G.&Stenberg,L.Expression of the Arp protein,a member of the M protein family,is not sufficient to inhibit phagocytosis of Streptococcus pyogenes.Infection and immunity 63,345-348(1995). 45.Gibson,D.G.et al.Enzymatic assembly of DNA molecules up to several hundred kilobases.Nat Methods 6,343-345(2009). 46.Tangri S, et al.(“Rationally engineered proteins therapeutic with reduced immunogenicity”J Immunol.2005 Mar 15;174(6):3187-96.
[0231] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided only as examples. Those skilled in the art can now make numerous variations, modifications, and substitutions without departing from the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein can be used in carrying out the invention. [Sequence List] SEQUENCE LISTING <110> THE BROAD INSTITUTE, INC. MASSACHUSETTS INSTITUTE OF TECHNOLOGY PRESIDENT AND FELLOWS OF HARVARD COLLEGE <120> ENGINEERING OF SYSTEMS, METHODS AND OPTIMIZED GUIDE COMPOSITIONS FOR SEQUENCE MANIPULATION <130> F50936A1 <140> PCT / US2013 / 074819 <141> 2013-12-12 <150> US 61 / 836,127 <151> 2013-06-17 <150> US 61 / 835,931 <151> 2013-06-17 <150> US 61 / 828,130 <151> 2013-05-28 <150> US 61 / 819,803 <151> 2013-05-06 <150> US 61 / 814,263 <151> 2013-04-20 <150> US 61 / 806,375 <151> 2013-03-28 <150> US 61 / 802,174 <151> 2013-03-15 <150> US 61 / 791,409 <151> 2013-03-15 <150> US 61 / 769,046 <151> 2013-02-25 <150> US 61 / 758,468 <151> 2013-01-30 <150> US 61 / 748,427 <151> 2013-01-02 <150> US 61 / 736,527 <151> 2012-12-12 <160> 264 <170> PatentIn version 3.5 <210> 1 <211> 15 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 1 aggacgaagt cctaa 15 <210> 2 <211> 7 <212> PRT <213> Simian virus 40 <400> 2 Pro Lys Lys Lys Arg Lys Val 1 5 <210> 3 <211> 16 <212> PRT <213> Unknown <220> <221> source <223> / note="Description of Unknown: Nucleoplasmin bipartite NLS sequence" <400> 3 Lys Arg Pro Ala Ala Thr Lys Lys Ala Gly Gln Ala Lys Lys Lys Lys 1 5 10 15 <210> 4 <211> 9 <212> PRT <213> Unknown <220> <221> source <223> / note="Description of Unknown: C-myc NLS sequence" <400> 4 Pro Ala Ala Lys Arg Val Lys Leu Asp 1 5 <210> 5 <211> 11 <212> PRT <213> Unknown <220> <221> source <223> / note="Description of Unknown: C-myc NLS sequence" <400> 5 Arg Gln Arg Arg Asn Glu Leu Lys Arg Ser Pro 1 5 10 <210> 6 <211> 38 <212> PRT <213> Homo sapiens <400> 6 Asn Gln Ser Ser Asn Phe Gly Pro Met Lys Gly Gly Asn Phe Gly Gly 1 5 10 15 Arg Ser Ser Gly Pro Tyr Gly Gly Gly Gly Gln Tyr Phe Ala Lys Pro 20 25 30 Arg Asn Gln Gly Gly Tyr 35 <210> 7 <211> 42 <212> PRT <213> Unknown <220> <221> source <223> / note="Description of Unknown: IBB domain from importin-alpha sequence" <400> 7 Arg Met Arg Ile Glx Phe Lys Asn Lys Gly Lys Asp Thr Ala Glu Leu 1 5 10 15 Arg Arg Arg Arg Val Glu Val Ser Val Glu Leu Arg Lys Ala Lys Lys 20 25 30 Asp Glu Gln Ile Leu Lys Arg Arg Asn Val 35 40 <210> 8 <211> 8 <212> PRT <213> Unknown <220> <221> source <223> / note="Description of Unknown: Myoma T protein sequence" <400> 8 Val Ser Arg Lys Arg Pro Arg Pro 1 5 <210> 9 <211> 8 <212> PRT <213> Unknown <220> <221> source <223> / note="Description of Unknown: Myoma T protein sequence” <400> 9 Pro Pro Lys Lys Ala Arg Glu Asp 1 5 <210> 10 <211> 8 <212> PRT <213> Homo sapiens <400> 10 Pro Gln Pro Lys Lys Lys Pro Leu 1 5 <210> 11 <211> 12 <212> PRT <213> Mus musculus <400> 11 Sister Ala Leu Ile Lys Lys Lys Lys Lys Met Ala Pro 1 5 10 <210> 12 <211> 5 <212> PRT <213> Influenza virus <400> 12 Asp Arg Leu Arg Arg 1 5 <210> 13 <211> 7 <212> PRT <213> Influenza virus <400> 13 Pro Light Gln Light Light Arg Light 1 5 <210> 14 <211> 10 <212> PRT <213> Hepatitis delta virus <400> 14 Arg Lys Leu Lys Lys Lys Ile Lys Lys Leu 1 5 10 <210> 15 <211> 10 <212> PRT <213> Mus musculus <400> 15 Arg Glu Lys Lys Lys Phe Leu Lys Arg Arg 1 5 10 <210> 16 <211> 20 <212> PRT <213> Homo sapiens <400> 16 Light Arg Light Gly Asp Glu Val Asp Gly Val Asp Glu Val Ala Light Light 1 5 10 15 Light Sees Light Light 20 <210> 17 <211> 17 <212> PRT <213> Homo sapiens <400> 17 Arg Lys Cys Leu Gln Ala Gly Met Asn Leu Glu Ala Arg Lys Thr Lys 1 5 10 15 Lys <210> 18 <211> 27 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t or g <220> <221> modified_base <222> (21)..(22) <223> a, c, t, g, unknown or other <400> 18 nnnnnnnnnn nnnnnnnnnn nnagaaw 27 <210> 19 <211> 19 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <220> <221> modified_base <222> (1)..(12) <223> a, c, t or g <220> <221> modified_base <222> (13)..(14) <223> a, c, t, g, unknown or other <400> 19 nnnnnnnnnn nnnnagaaw 19 <210> 20 <211> 27 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t or g <220> <221> modified_base <222> (21)..(22) <223> a, c, t, g, unknown or other <400> 20 nnnnnnnnnn nnnnnnnnnn nnagaaw 27 <210> 21 <211> 18 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <220> <221> modified_base <222> (1)..(11) <223> a, c, t or g <220> <221> modified_base <222> (12)..(13) <223> a, c, t, g, unknown or other <400> 21 nnnnnnnnnn nnnagaaw 18 <210> 22 <211> 137 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 22 nnnnnnnnnn nnnnnnnnnn gtttttgtac tctcaagatt tagaaataaa tcttgcagaa 60 gctacaaaga taaggcttca tgccgaaatc aacaccctgt cattttatgg cagggtgttt 120 tcgttattta atttttt 137 <210> 23 <211> 123 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 23 nnnnnnnnnn nnnnnnnnnn gtttttgtac tctcagaaat gcagaagcta caaagataag 60 gcttcatgcc gaaatcaaca ccctgtcatt ttatggcagg gtgttttcgt tatttaattt 120 ttt 123 <210> 24 <211> 110 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 24 nnnnnnnnnn nnnnnnnnnn gtttttgtac tctcagaaat gcagaagcta caaagataag 60 gcttcatgcc gaaatcaaca ccctgtcatt ttatggcagg gtgttttttt 110 <210> 25 <211> 102 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 25 nnnnnnnnnn nnnnnnnnnn gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgctttt tt 102 <210> 26 <211> 88 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 26 nnnnnnnnnn nnnnnnnnnn gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt gttttttt 88 <210> 27 <211> 76 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 27 nnnnnnnnnn nnnnnnnnnn gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcatt tttttt 76 <210> 28 <211> 12 <212> RNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide” <400> 28 guuuuagagc ua 12 <210> 29 <211> 33 <212> DNA <213> Homo sapiens <400> 29 ggacatcgat gtcacctcca atgactaggg tgg 33 <210> 30 <211> 33 <212> DNA <213> Homo sapiens <400> 30 cattggaggt gacatcgatg tcctccccat tgg 33 <210> 31 <211> 33 <212> DNA <213> Homo sapiens <400> 31 ggaagggcct gagtccgagc agaagaagaa ggg 33 <210> 32 <211> 33 <212> DNA <213> Homo sapiens <400> 32 ggtggcgaga ggggccgaga ttgggtgttc agg 33 <210> 33 <211> 33 <212> DNA <213> Homo sapiens <400> 33 atgcaggagg gtggcgagag gggccgagat tgg 33 <210> 34 <211> 21 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic primary" <400> 34 aaaaccaccc ttctctctgg c 21 <210> 35 <211> 21 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic primary” <400> 35 ggagattgga gacacggaga g 21 <210> 36 <211> 20 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic primer" <400> 36 ctggaaagcc aatgcctgac 20 <210> 37 <211> 20 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic primer" <400> 37 ggcagcaaac tccttgtcct 20 <210> 38 <211> 12 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 38 gttttagagc ta 12 <210> 39 <211> 335 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <400> 39 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 aatattggaa ttaatttgac tgtaacaca agatattag tacaaatac gtgacgtaga 120 aagtaataat ttctgggta gtttgcagtt ttaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatctt gtggaagga 240 cgaaacaccg gaaccattca aaacagcata gcaagttaa aaggctag tccgttatca 300 acttgaaaaa gtggcaccga gtcggtgctt tttt 335 <210> 40 <211> 423 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <400> 40 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 aatattggaa ttaatttgac tgtaacaca agatattag tacaaatac gtgacgtaga 120 aagtaataat ttctgggta gtttgcagtt ttaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatctt gtggaagga 240 cgaaacaccg gtagtattaa gtattgtttt atggctgata aatttctttg aatttctcct 300 tgattatttg tttaaaagt tataaaataa tcttgttg accattcaa acagcatagc 360 aagttaaaat aaggctagtc cgttatcaac ttgaaaagt ggcaccgagt cggtgctttt 420 pp. 423 <210> 41 <211> 339 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <400> 41 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 aatattggaa ttaatttgac tgtaacaca agatattag tacaaatac gtgacgtaga 120 aagtaataat ttctgggta gtttgcagtt ttaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatctt gtggaagga 240 cgaaacaccg ggttttagag ctatgctgtt tgaatggtc ccaaaacggg tcttcgagaa 300 gacgttttag agctatgctg tttgaatgg tccaaac 339 <210> 42 <211> 309 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <400> 42 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 ataattggaa ttaatttgac tgtaaacaca aagatattag tacaaaatac gtgacgtaga 120 aagtaataat ttcttgggta gtttgcagtt ttaaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatatctt gtggaaagga 240 cgaaacaccg ggtcttcgag aagacctgtt ttagagctag aatagcaag ttaaaataag 300 gctagtccg 309 <210> 43 <211> 1648 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polypeptide” <400> 43 Met Asp Tyr Lys Asp His Asp Gly Asp Tyr Lys Asp His Asp Ile Asp 1 5 10 15 Tyr Lys Asp Asp Asp Asp Lys Put Here Pro Lys Lys Lys Arg Lys Val 20 25 30 Gly Ile His Gly Val Pro Ala Ala Asp Lys Lys Tyr Ser Ile Gly Leu 35 40 45 Asp Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu Tyr 50 55 60 Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg His 65 70 75 80 Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly Glu 85 90 95 Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr Thr 100 105 110 Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn Glu 115 120 125 Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser Phe 130 135 140 Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly Asn 145 150 155 160 Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr His 165 170 175 Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg Leu 180 185 190 Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe Leu 195 200 205 Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu Phe 210 215 220 Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro Ile 225 230 235 240 Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu Ser 245 250 255 Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu Lys 260 265 270 Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu Thr 275 280 285 Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu Gln 290 295 300 Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala Gln 305 310 315 320 Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu Ser 325 330 335 Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile Thr 340 345 350 Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His His 355 360 365 Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro Glu 370 375 380 Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala Gly 385 390 395 400 Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile Lys 405 410 415 Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys Leu 420 425 430 Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly Ser 435 440 445 Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg Arg 450 455 460 Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile Glu 465 470 475 480 Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala Arg 485 490 495 Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr Ile 500 505 510 Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala Gln 515 520 525 Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn Glu 530 535 540 Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val Tyr 545 550 555 560 Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys Pro 565 570 575 Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu Phe 580 585 590 Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr Phe 595 600 605 Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu Asp 610 615 620 Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile Ile 625 630 635 640 Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu Glu 645 650 655 Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile Glu 660 665 670 Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met Lys 675 680 685 Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg Lys 690 695 700 Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu Asp 705 710 715 720 Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu Ile 725 730 735 His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln Val 740 745 750 Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala Gly 755 760 765 Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val Asp 770 775 780 Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile Val Ile 785 790 795 800 Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn Ser 805 810 815 Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly Ser 820 825 830 Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln Asn Glu 835 840 845 Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr Val Asp 850 855 860 Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp His Ile 865 870 875 880 Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys Val Leu 885 890 895 Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro Ser Glu 900 905 910 Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu Asn Ala 915 920 925 Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr Lys Ala Glu Arg 930 935 940 Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe Ile Lys Arg Gln Leu 945 950 955 960 Val Glu Thr Arg Gln Ile Thr Lys His Val Ala Gln Ile Leu Asp Ser 965 970 975 Arg Met Asn Thr Lys Tyr Asp Glu Asn Asp Lys Leu Ile Arg Glu Val 980 985 990 Lys Val Ile Thr Leu Lys Ser Lys Leu Val Ser Asp Phe Arg Lys Asp 995 1000 1005 Phe Gln Phe Tyr Lys Val Arg Glu Ile Asn Asn Tyr His His Ala 1010 1015 1020 His Asp Ala Tyr Leu Asn Ala Val Val Gly Thr Ala Leu Ile Lys 1025 1030 1035 Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr Lys 1040 1045 1050 Val Tyr Asp Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu Ile 1055 1060 1065 Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile Met Asn 1070 1075 1080 Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile Arg Lys 1085 1090 1095 Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile Val Trp 1100 1105 1110 Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser Met 1115 1120 1125 Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly 1130 1135 1140 Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys Leu 1145 1150 1155 Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe 1160 1165 1170 Asp Ser Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val 1175 1180 1185 Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu 1190 1195 1200 Gly Ile Thr Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile 1205 1210 1215 Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu 1220 1225 1230 Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn Gly 1235 1240 1245 Arg Lys Arg Met Leu Ala Ser Ala Gly Glu Leu Gln Lys Gly Asn 1250 1255 1260 Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr Leu Ala 1265 1270 1275 Ser His Tyr Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn Glu Gln 1280 1285 1290 Lys Gln Leu Phe Val Glu Gln His Lys His Tyr Leu Asp Glu Ile 1295 1300 1305 Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg Val Ile Leu Ala Asp 1310 1315 1320 Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr Asn Lys His Arg Asp 1325 1330 1335 Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile Ile His Leu Phe Thr 1340 1345 1350 Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe Asp Thr 1355 1360 1365 Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu Asp 1370 1375 1380 Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr Arg 1385 1390 1395 Ile Asp Leu Ser Gln Leu Gly Gly Asp Ala Ala Ala Val Ser Lys 1400 1405 1410 Gly Glu Glu Leu Phe Thr Gly Val Val Pro Ile Leu Val Glu Leu 1415 1420 1425 Asp Gly Asp Val Asn Gly His Lys Phe Ser Val Ser Gly Glu Gly 1430 1435 1440 Glu Gly Asp Ala Thr Tyr Gly Lys Leu Thr Leu Lys Phe Ile Cys 1445 1450 1455 Thr Thr Gly Lys Leu Pro Val Pro Trp Pro Thr Leu Val Thr Thr 1460 1465 1470 Leu Thr Tyr Gly Val Gln Cys Phe Ser Arg Tyr Pro Asp His Met 1475 1480 1485 Lys Gln His Asp Phe Phe Lys Ser Ala Met Pro Glu Gly Tyr Val 1490 1495 1500 Gln Glu Arg Thr Ile Phe Phe Lys Asp Asp Gly Asn Tyr Lys Thr 1505 1510 1515 Arg Ala Glu Val Lys Phe Glu Gly Asp Thr Leu Val Asn Arg Ile 1520 1525 1530 Glu Leu Lys Gly Ile Asp Phe Lys Glu Asp Gly Asn Ile Leu Gly 1535 1540 1545 His Lys Leu Glu Tyr Asn Tyr Asn Ser His Asn Val Tyr Ile Met 1550 1555 1560 Ala Asp Lys Gln Lys Asn Gly Ile Lys Val Asn Phe Lys Ile Arg 1565 1570 1575 His Asn Ile Glu Asp Gly Ser Val Gln Leu Ala Asp His Tyr Gln 1580 1585 1590 Gln Asn Thr Pro Ile Gly Asp Gly Pro Val Leu Leu Pro Asp Asn 1595 1600 1605 His Tyr Leu Ser Thr Gln Ser Ala Leu Ser Lys Asp Pro Asn Glu 1610 1615 1620 Lys Arg Asp His Met Val Leu Leu Glu Phe Val Thr Ala Ala Gly 1625 1630 1635 Ile Thr Leu Gly Met Asp Glu Leu Tyr Lys 1640 1645 <210> 44 <211> 1625 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polypeptide" <400> 44 Met Asp Lys Lys Tyr Ser Ile Gly Leu Asp Ile Gly Thr Asn Ser Val 1 5 10 15 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile 35 40 45 Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys 65 70 75 80 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 85 90 95 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp 130 135 140 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 145 150 155 160 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 210 215 220 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 260 265 270 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 275 280 285 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 290 295 300 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 325 330 335 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 340 345 350 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 485 490 495 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys 515 520 525 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 565 570 575 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580 585 590 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830 Leu Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys 885 890 895 Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 900 905 910 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915 920 925 Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp 930 935 940 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Leu Lys Ser 945 950 955 960 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965,970,975 Glu Ile Asn Asn Tyr His Ala His Asp Ala Tyr Leu Asn Ala Val 980,985,990 Val Gly Thr Ala Leu Ile Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1010 1015 1020 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1040 1045 1050 Asn Gly Glu With Arg Lys Arg Pro Leu Glu Thr Asn Gly Glu 1055 1060 1065 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr 1085 1090 1095 Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1100 1105 1110 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 Light Light Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1190 1195 1200 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1220 1225 1230 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1265 1270 1275 Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1280 1285 1290 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn 1295 1300 1305 Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1310 1315 1320 Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser 1325 1330 1335 Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr 1340 1345 1350 Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 Ala Ala Ala Val Ser Lys Gly Glu Glu Leu Phe Thr Gly Val Val 1370 1375 1380 Pro Ile Leu Val Glu Leu Asp Gly Asp Val Asn Gly His Lys Phe 1385 1390 1395 Ser Val Ser Gly Glu Gly Glu Gly Asp Ala Thr Tyr Gly Lys Leu 1400 1405 1410 Thr Leu Lys Phe Ile Cys Thr Thr Gly Lys Leu Pro Val Pro Trp 1415 1420 1425 Pro Thr Leu Val Thr Thr Leu Thr Tyr Gly Val Gln Cys Phe Ser 1430 1435 1440 Arg Tyr Pro Asp His Met Lys Gln His Asp Phe Phe Lys Ser Ala 1445 1450 1455 Met Pro Glu Gly Tyr Val Gln Glu Arg Thr Ile Phe Phe Lys Asp 1460 1465 1470 Asp Gly Asn Tyr Lys Thr Arg Ala Glu Val Lys Phe Glu Gly Asp 1475 1480 1485 Thr Leu Val Asn Arg Ile Glu Leu Lys Gly Ile Asp Phe Lys Glu 1490 1495 1500 Asp Gly Asn Ile Leu Gly His Lys Leu Glu Tyr Asn Tyr Asn Ser 1505 1510 1515 His Asn Val Tyr Ile Met Ala Asp Lys Gln Lys Asn Gly Ile Lys 1520 1525 1530 Val Asn Phe Lys Ile Arg His Asn Ile Glu Asp Gly Ser Val Gln 1535 1540 1545 Leu Ala Asp His Tyr Gln Gln Asn Thr Pro Ile Gly Asp Gly Pro 1550 1555 1560 Val Leu Leu Pro Asp Asn His Tyr Leu Ser Thr Gln Ser Ala Leu 1565 1570 1575 Ser Lys Asp Pro Asn Glu Lys Arg Asp His Met Val Leu Leu Glu 1580 1585 1590 Phe Val Thr Ala Ala Gly Ile Thr Leu Gly Met Asp Glu Leu Tyr 1595 1600 1605 Lys Lys Arg Pro Ala Ala Thr Lys Lys Ala Gly Gln Ala Lys Lys 1610 1615 1620 Lys Lys 1625 <210> 45 <211> 1664 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polypeptide" <400> 45 Met Asp Tyr Lys Asp His Asp Gly Asp Tyr Lys Asp His Asp Ile Asp 1 5 10 15 Tyr Lys Asp Asp Asp Asp Lys Met Ala Pro Lys Lys Lys Arg Lys Val 20 25 30 Gly Ile His Gly Val Pro Ala Ala Asp Lys Lys Tyr Ser Ile Gly Leu 35 40 45 Asp Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu Tyr 50 55 60 Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg His 65 70 75 80 Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly Glu 85 90 95 Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr Thr 100 105 110 Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn Glu 115 120 125 Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser Phe 130 135 140 Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly Asn 145 150 155 160 Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr His 165 170 175 Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg Leu 180 185 190 Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe Leu 195 200 205 Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu Phe 210 215 220 Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro Ile 225 230 235 240 Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu Ser 245 250 255 Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu Lys 260 265 270 Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu Thr 275 280 285 Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu Gln 290 295 300 Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala Gln 305 310 315 320 Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu Ser 325 330 335 Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile Thr 340 345 350 Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His His 355 360 365 Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro Glu 370 375 380 Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala Gly 385 390 395 400 Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile Lys 405 410 415 Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys Leu 420 425 430 Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly Ser 435 440 445 Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg Arg 450 455 460 Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile Glu 465 470 475 480 Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala Arg 485 490 495 Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr Ile 500 505 510 Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala Gln 515 520 525 Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn Glu 530 535 540 Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val Tyr 545 550 555 560 Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys Pro 565 570 575 Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu Phe 580 585 590 Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr Phe 595 600 605 Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu Asp 610 615 620 Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile Ile 625 630 635 640 Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu Glu 645 650 655 Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile Glu 660 665 670 Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met Lys 675 680 685 Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg Lys 690 695 700 Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu Asp 705 710 715 720 Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu Ile 725 730 735 His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln Val 740 745 750 Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala Gly 755 760 765 Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val Asp 770 775 780 Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile Val Ile 785 790 795 800 Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn Ser 805 810 815 Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly Ser 820 825 830 Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln Asn Glu 835 840 845 Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr Val Asp 850 855 860 Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp His Ile 865 870 875 880 Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys Val Leu 885 890 895 Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro Ser Glu 900 905 910 Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu Asn Ala 915 920 925 Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr Lys Ala Glu Arg 930 935 940 Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe Ile Lys Arg Gln Leu 945 950 955 960 Val Glu Thr Arg Gln Ile Thr Lys His Val Ala Gln Ile Leu Asp Ser 965 970 975 Arg Met Asn Thr Lys Tyr Asp Glu Asn Asp Lys Leu Ile Arg Glu Val 980 985 990 Lys Val Ile Thr Leu Lys Ser Lys Leu Val Ser Asp Phe Arg Lys Asp 995 1000 1005 Phe Gln Phe Tyr Lys Val Arg Glu Ile Asn Asn Tyr His His Ala 1010 1015 1020 His Asp Ala Tyr Leu Asn Ala Val Val Gly Thr Ala Leu Ile Lys 1025 1030 1035 Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr Lys 1040 1045 1050 Val Tyr Asp Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu Ile 1055 1060 1065 Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile Met Asn 1070 1075 1080 Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile Arg Lys 1085 1090 1095 Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile Val Trp 1100 1105 1110 Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser Met 1115 1120 1125 Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly 1130 1135 1140 Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys Leu 1145 1150 1155 Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe 1160 1165 1170 Asp Ser Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val 1175 1180 1185 Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu 1190 1195 1200 Gly Ile Thr Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile 1205 1210 1215 Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu 1220 1225 1230 Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn Gly 1235 1240 1245 Arg Lys Arg Met Leu Ala Ser Ala Gly Glu Leu Gln Lys Gly Asn 1250 1255 1260 Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr Leu Ala 1265 1270 1275 Ser His Tyr Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn Glu Gln 1280 1285 1290 Lys Gln Leu Phe Val Glu Gln His Lys His Tyr Leu Asp Glu Ile 1295 1300 1305 Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg Val Ile Leu Ala Asp 1310 1315 1320 Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr Asn Lys His Arg Asp 1325 1330 1335 Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile Ile His Leu Phe Thr 1340 1345 1350 Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe Asp Thr 1355 1360 1365 Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu Asp 1370 1375 1380 Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr Arg 1385 1390 1395 Ile Asp Leu Ser Gln Leu Gly Gly Asp Ala Ala Ala Val Ser Lys 1400 1405 1410 Gly Glu Glu Leu Phe Thr Gly Val Val Pro Ile Leu Val Glu Leu 1415 1420 1425 Asp Gly Asp Val Asn Gly His Lys Phe Ser Val Ser Gly Glu Gly 1430 1435 1440 Glu Gly Asp Ala Thr Tyr Gly Lys Leu Thr Leu Lys Phe Ile Cys 1445 1450 1455 Thr Thr Gly Lys Leu Pro Val Pro Trp Pro Thr Leu Val Thr Thr 1460 1465 1470 Leu Thr Tyr Gly Val Gln Cys Phe Ser Arg Tyr Pro Asp His Met 1475 1480 1485 Lys Gln His Asp Phe Phe Lys Ser Ala Met Pro Glu Gly Tyr Val 1490 1495 1500 Gln Glu Arg Thr Ile Phe Phe Lys Asp Asp Gly Asn Tyr Lys Thr 1505 1510 1515 Arg Ala Glu Val Lys Phe Glu Gly Asp Thr Leu Val Asn Arg Ile 1520 1525 1530 Glu Leu Lys Gly Ile Asp Phe Lys Glu Asp Gly Asn Ile Leu Gly 1535 1540 1545 His Lys Leu Glu Tyr Asn Tyr Asn Ser His Asn Val Tyr Ile Met 1550 1555 1560 Ala Asp Lys Gln Lys Asn Gly Ile Lys Val Asn Phe Lys Ile Arg 1565 1570 1575 His Asn Ile Glu Asp Gly Ser Val Gln Leu Ala Asp His Tyr Gln 1580 1585 1590 Gln Asn Thr Pro Ile Gly Asp Gly Pro Val Leu Leu Pro Asp Asn 1595 1600 1605 His Tyr Leu Ser Thr Gln Ser Ala Leu Ser Lys Asp Pro Asn Glu 1610 1615 1620 Lys Arg Asp His Met Val Leu Leu Glu Phe Val Thr Ala Ala Gly 1625 1630 1635 Ile Thr Leu Gly Met Asp Glu Leu Tyr Lys Lys Arg Pro Ala Ala 1640 1645 1650 Thr Lys Lys Ala Gly Gln Ala Lys Lys Lys Lys 1655 1660 <210> 46 <211> 1423 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polypeptide" <400> 46 Met Asp Tyr Lys Asp His Asp Gly Asp Tyr Lys Asp His Asp Ile Asp 1 5 10 15 Tyr Lys Asp Asp Asp Asp Lys Met Ala Pro Lys Lys Lys Arg Lys Val 20 25 30 Gly Ile His Gly Val Pro Ala Ala Asp Lys Lys Tyr Ser Ile Gly Leu 35 40 45 Asp Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu Tyr 50 55 60 Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg His 65 70 75 80 Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly Glu 85 90 95 Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr Thr 100 105 110 Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn Glu 115 120 125 Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser Phe 130 135 140 Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly Asn 145 150 155 160 Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr His 165 170 175 Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg Leu 180 185 190 Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe Leu 195 200 205 Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu Phe 210 215 220 Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro Ile 225 230 235 240 Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu Ser 245 250 255 Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu Lys 260 265 270 Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu Thr 275 280 285 Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu Gln 290 295 300 Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala Gln 305 310 315 320 Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu Ser 325 330 335 Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile Thr 340 345 350 Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His His 355 360 365 Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro Glu 370 375 380 Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala Gly 385 390 395 400 Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile Lys 405 410 415 Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys Leu 420 425 430 Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly Ser 435 440 445 Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg Arg 450 455 460 Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile Glu 465 470 475 480 Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala Arg 485 490 495 Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr Ile 500 505 510 Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala Gln 515 520 525 Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn Glu 530 535 540 Lys Val Leu Pro Lys His Ser Leu Leu Tyr Glu Tyr Phe Thr Val Tyr 545 550 555 560 Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys Pro 565 570 575 Ala Phe Leu Ser Gly Glu Gln Lys Lys Ala Ile Val Asp Leu Leu Phe 580 585 590 Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr Phe 595 600 605 Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu Asp 610 615 620 Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile Ile 625 630 635 640 Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu Glu 645 650 655 Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile Glu 660 665 670 Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met Lys 675 680 685 Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg Lys 690 695 700 Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu Asp 705 710 715 720 Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu Ile 725 730 735 His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln Val 740 745 750 Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala Gly 755 760 765 Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val Asp 770 775 780 Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile Val Ile 785 790 795 800 Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn Ser 805 810 815 Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly Ser 820 825 830 Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln Asn Glu 835 840 845 Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr Val Asp 850 855 860 Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp His Ile 865 870 875 880 Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys Val Leu 885 890 895 Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro Ser Glu 900 905 910 Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu Asn Ala 915 920 925 Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr Lys Ala Glu Arg 930 935 940 Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe Ile Lys Arg Gln Leu 945 950 955 960 Val Glu Thr Arg Gln Ile Thr Lys His Val Ala Gln Ile Leu Asp Ser 965 970 975 Arg Met Asn Thr Lys Tyr Asp Glu Asn Asp Lys Leu Ile Arg Glu Val 980 985 990 Lys Val Ile Thr Leu Lys Ser Lys Leu Val Ser Asp Phe Arg Lys Asp 995 1000 1005 Phe Gln Phe Tyr Lys Val Arg Glu Ile Asn Asn Tyr His His Ala 1010 1015 1020 His Asp Ala Tyr Leu Asn Ala Val Val Gly Thr Ala Leu Ile Lys 1025 1030 1035 Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr Lys 1040 1045 1050 Val Tyr Asp Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu Ile 1055 1060 1065 Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile Met Asn 1070 1075 1080 Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile Arg Lys 1085 1090 1095 Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile Val Trp 1100 1105 1110 Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser Met 1115 1120 1125 Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly 1130 1135 1140 Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys Leu 1145 1150 1155 Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe 1160 1165 1170 Asp Ser Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val 1175 1180 1185 Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu 1190 1195 1200 Gly Ile Thr Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile 1205 1210 1215 Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu 1220 1225 1230 Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn Gly 1235 1240 1245 Arg Lys Arg Met Leu Ala Ser Ala Gly Glu Leu Gln Lys Gly Asn 1250 1255 1260 Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr Leu Ala 1265 1270 1275 Ser His Tyr Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn Glu Gln 1280 1285 1290 Lys Gln Leu Phe Val Glu Gln His Lys His Tyr Leu Asp Glu Ile 1295 1300 1305 Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg Val Ile Leu Ala Asp 1310 1315 1320 Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr Asn Lys His Arg Asp 1325 1330 1335 Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile Ile His Leu Phe Thr 1340 1345 1350 Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe Asp Thr 1355 1360 1365 Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu Asp 1370 1375 1380 Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr Arg 1385 1390 1395 Ile Asp Leu Ser Gln Leu Gly Gly Asp Lys Arg Pro Ala Ala Thr 1400 1405 1410 Lys Lys Ala Gly Gln Ala Lys Lys Lys Lys 1415 1420 <210> 47 <211> 483 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polypeptide" <400> 47 Met Phe Leu Phe Leu Ser Leu Thr Ser Phe Leu Ser Ser Ser Arg Thr 1 5 10 15 Leu Val Ser Lys Gly Glu Glu Asp Asn Met Ala Ile Ile Lys Glu Phe 20 25 30 Met Arg Phe Lys Val His Met Glu Gly Ser Val Asn Gly His Glu Phe 35 40 45 Glu Ile Glu Gly Glu Gly Glu Gly Arg Pro Tyr Glu Gly Thr Gln Thr 50 55 60 Ala Lys Leu Lys Val Thr Lys Gly Gly Pro Leu Pro Phe Ala Trp Asp 65 70 75 80 Ile Leu Ser Pro Gln Phe Met Tyr Gly Ser Lys Ala Tyr Val Lys His 85 90 95 Pro Ala Asp Ile Pro Asp Tyr Leu Lys Leu Ser Phe Pro Glu Gly Phe 100 105 110 Lys Trp Glu Arg Val Met Asn Phe Glu Asp Gly Gly Val Val Thr Val 115 120 125 Thr Gln Asp Ser Ser Leu Gln Asp Gly Glu Phe Ile Tyr Lys Val Lys 130 135 140 Leu Arg Gly Thr Asn Phe Pro Ser Asp Gly Pro Val Met Gln Lys Lys 145 150 155 160 Thr Met Gly Trp Glu Ala Ser Ser Glu Arg Met Tyr Pro Glu Asp Gly 165 170 175 Ala Leu Lys Gly Glu Ile Lys Gln Arg Leu Lys Leu Lys Asp Gly Gly 180 185 190 His Tyr Asp Ala Glu Val Lys Thr Thr Tyr Lys Ala Lys Lys Pro Val 195 200 205 Gln Leu Pro Gly Ala Tyr Asn Val Asn Ile Lys Leu Asp Ile Thr Ser 210 215 220 His Asn Glu Asp Tyr Thr Ile Val Glu Gln Tyr Glu Arg Ala Glu Gly 225 230 235 240 Arg His Ser Thr Gly Gly Met Asp Glu Leu Tyr Lys Gly Ser Lys Gln 245 250 255 Leu Glu Glu Leu Leu Ser Thr Ser Phe Asp Ile Gln Phe Asn Asp Leu 260 265 270 Thr Leu Leu Glu Thr Ala Phe Thr His Thr Ser Tyr Ala Asn Glu His 275 280 285 Arg Leu Leu Asn Val Ser His Asn Glu Arg Leu Glu Phe Leu Gly Asp 290 295 300 Ala Val Leu Gln Leu Ile Ile Ser Glu Tyr Leu Phe Ala Lys Tyr Pro 305 310 315 320 Lys Lys Thr Glu Gly Asp Met Ser Lys Leu Arg Ser Met Ile Val Arg 325 330 335 Glu Glu Ser Leu Ala Gly Phe Ser Arg Phe Cys Ser Phe Asp Ala Tyr 340 345 350 Ile Lys Leu Gly Lys Gly Glu Glu Lys Ser Gly Gly Arg Arg Arg Asp 355 360 365 Thr Ile Leu Gly Asp Leu Phe Glu Ala Phe Leu Gly Ala Leu Leu Leu 370 375 380 Asp Lys Gly Ile Asp Ala Val Arg Arg Phe Leu Lys Gln Val Met Ile 385 390 395 400 Pro Gln Val Glu Lys Gly Asn Phe Glu Arg Val Lys Asp Tyr Lys Thr 405 410 415 Cys Leu Gln Glu Phe Leu Gln Thr Lys Gly Asp Val Ala Ile Asp Tyr 420 425 430 Gln Val Ile Ser Glu Lys Gly Pro Ala His Ala Lys Gln Phe Glu Val 435 440 445 Ser Ile Val Val Asn Gly Ala Val Leu Ser Lys Gly Leu Gly Lys Ser 450 455 460 Lys Lys Leu Ala Glu Gln Asp Ala Ala Lys Asn Ala Leu Ala Gln Leu 465 470 475 480 Ser Glu Val <210> 48 <211> 483 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polypeptide" <400> 48 Met Lys Gln Leu Glu Glu Leu Leu Ser Thr Ser Phe Asp Ile Gln Phe 1 5 10 15 Asn Asp Leu Thr Leu Leu Glu Thr Ala Phe Thr His Thr Ser Tyr Ala 20 25 30 Asn Glu His Arg Leu Leu Asn Val Ser His Asn Glu Arg Leu Glu Phe 35 40 45 Leu Gly Asp Ala Val Leu Gln Leu Ile Ile Ser Glu Tyr Leu Phe Ala 50 55 60 Lys Tyr Pro Lys Lys Thr Glu Gly Asp Met Ser Lys Leu Arg Ser Met 65 70 75 80 Ile Val Arg Glu Glu Ser Leu Ala Gly Phe Ser Arg Phe Cys Ser Phe 85 90 95 Asp Ala Tyr Ile Lys Leu Gly Lys Gly Glu Glu Lys Ser Gly Gly Arg 100 105 110 Arg Arg Asp Thr Ile Leu Gly Asp Leu Phe Glu Ala Phe Leu Gly Ala 115 120 125 Leu Leu Leu Asp Lys Gly Ile Asp Ala Val Arg Arg Phe Leu Lys Gln 130 135 140 Val Met Ile Pro Gln Val Glu Lys Gly Asn Phe Glu Arg Val Lys Asp 145 150 155 160 Tyr Lys Thr Cys Leu Gln Glu Phe Leu Gln Thr Lys Gly Asp Val Ala 165 170 175 Ile Asp Tyr Gln Val Ile Ser Glu Lys Gly Pro Ala His Ala Lys Gln 180 185 190 Phe Glu Val Ser Ile Val Val Asn Gly Ala Val Leu Ser Lys Gly Leu 195 200 205 Gly Lys Ser Lys Lys Leu Ala Glu Gln Asp Ala Ala Lys Asn Ala Leu 210 215 220 Ala Gln Leu Ser Glu Val Gly Ser Val Ser Lys Gly Glu Glu Asp Asn 225 230 235 240 Met Ala Ile Ile Lys Glu Phe Met Arg Phe Lys Val His Met Glu Gly 245 250 255 Ser Val Asn Gly His Glu Phe Glu Ile Glu Gly Glu Gly Glu Gly Arg 260 265 270 Pro Tyr Glu Gly Thr Gln Thr Ala Lys Leu Lys Val Thr Lys Gly Gly 275 280 285 Pro Leu Pro Phe Ala Trp Asp Ile Leu Ser Pro Gln Phe Met Tyr Gly 290 295 300 Ser Lys Ala Tyr Val Lys His Pro Ala Asp Ile Pro Asp Tyr Leu Lys 305 310 315 320 Leu Ser Phe Pro Glu Gly Phe Lys Trp Glu Arg Val Met Asn Phe Glu 325 330 335 Asp Gly Gly Val Val Thr Val Thr Gln Asp Ser Ser Leu Gln Asp Gly 340 345 350 Glu Phe Ile Tyr Lys Val Lys Leu Arg Gly Thr Asn Phe Pro Ser Asp 355 360 365 Gly Pro Val Met Gln Lys Lys Thr Met Gly Trp Glu Ala Ser Ser Glu 370 375 380 Arg Met Tyr Pro Glu Asp Gly Ala Leu Lys Gly Glu Ile Lys Gln Arg 385 390 395 400 Leu Lys Leu Lys Asp Gly Gly His Tyr Asp Ala Glu Val Lys Thr Thr 405 410 415 Tyr Lys Ala Lys Lys Pro Val Gln Leu Pro Gly Ala Tyr Asn Val Asn 420 425 430 Ile Lys Leu Asp Ile Thr Ser His Asn Glu Asp Tyr Thr Ile Val Glu 435 440 445 Gln Tyr Glu Arg Ala Glu Gly Arg His Ser Thr Gly Gly Met Asp Glu 450 455 460 Leu Tyr Lys Lys Arg Pro Ala Ala Thr Lys Lys Ala Gly Gln Ala Lys 465 470 475 480 Lys Lys Lys <210> 49 <211> 1423 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polypeptide" <400> 49 Met Asp Tyr Lys Asp His Asp Gly Asp Tyr Lys Asp His Asp Ile Asp 1 5 10 15 Tyr Lys Asp Asp Asp Asp Lys Met Ala Pro Lys Lys Lys Arg Lys Val 20 25 30 Gly Ile His Gly Val Pro Ala Ala Asp Lys Lys Tyr Ser Ile Gly Leu 35 40 45 Ala Ile Gly Thr Asn Ser Val Gly Trp Ala Val Ile Thr Asp Glu Tyr 50 55 60 Lys Val Pro Ser Lys Lys Phe Lys Val Leu Gly Asn Thr Asp Arg His 65 70 75 80 Ser Ile Lys Lys Asn Leu Ile Gly Ala Leu Leu Phe Asp Ser Gly Glu 85 90 95 Thr Ala Glu Ala Thr Arg Leu Lys Arg Thr Ala Arg Arg Arg Tyr Thr 100 105 110 Arg Arg Lys Asn Arg Ile Cys Tyr Leu Gln Glu Ile Phe Ser Asn Glu 115 120 125 Met Ala Lys Val Asp Asp Ser Phe Phe His Arg Leu Glu Glu Ser Phe 130 135 140 Leu Val Glu Glu Asp Lys Lys His Glu Arg His Pro Ile Phe Gly Asn 145 150 155 160 Ile Val Asp Glu Val Ala Tyr His Glu Lys Tyr Pro Thr Ile Tyr His 165 170 175 Leu Arg Lys Lys Leu Val Asp Ser Thr Asp Lys Ala Asp Leu Arg Leu 180 185 190 Ile Tyr Leu Ala Leu Ala His Met Ile Lys Phe Arg Gly His Phe Leu 195 200 205 Ile Glu Gly Asp Leu Asn Pro Asp Asn Ser Asp Val Asp Lys Leu Phe 210 215 220 Ile Gln Leu Val Gln Thr Tyr Asn Gln Leu Phe Glu Glu Asn Pro Ile 225 230 235 240 Asn Ala Ser Gly Val Asp Ala Lys Ala Ile Leu Ser Ala Arg Leu Ser 245 250 255 Lys Ser Arg Arg Leu Glu Asn Leu Ile Ala Gln Leu Pro Gly Glu Lys 260 265 270 Lys Asn Gly Leu Phe Gly Asn Leu Ile Ala Leu Ser Leu Gly Leu Thr 275 280 285 Pro Asn Phe Lys Ser Asn Phe Asp Leu Ala Glu Asp Ala Lys Leu Gln 290 295 300 Leu Ser Lys Asp Thr Tyr Asp Asp Asp Leu Asp Asn Leu Leu Ala Gln 305 310 315 320 Ile Gly Asp Gln Tyr Ala Asp Leu Phe Leu Ala Ala Lys Asn Leu Ser 325 330 335 Asp Ala Ile Leu Leu Ser Asp Ile Leu Arg Val Asn Thr Glu Ile Thr 340 345 350 Lys Ala Pro Leu Ser Ala Ser Met Ile Lys Arg Tyr Asp Glu His His 355 360 365 Gln Asp Leu Thr Leu Leu Lys Ala Leu Val Arg Gln Gln Leu Pro Glu 370 375 380 Lys Tyr Lys Glu Ile Phe Phe Asp Gln Ser Lys Asn Gly Tyr Ala Gly 385 390 395 400 Tyr Ile Asp Gly Gly Ala Ser Gln Glu Glu Phe Tyr Lys Phe Ile Lys 405 410 415 Pro Ile Leu Glu Lys Met Asp Gly Thr Glu Glu Leu Leu Val Lys Leu 420 425 430 Asn Arg Glu Asp Leu Leu Arg Lys Gln Arg Thr Phe Asp Asn Gly Ser 435 440 445 Ile Pro His Gln Ile His Leu Gly Glu Leu His Ala Ile Leu Arg Arg 450 455 460 Gln Glu Asp Phe Tyr Pro Phe Leu Lys Asp Asn Arg Glu Lys Ile Glu 465 470 475 480 Lys Ile Leu Thr Phe Arg Ile Pro Tyr Tyr Val Gly Pro Leu Ala Arg 485 490 495 Gly Asn Ser Arg Phe Ala Trp Met Thr Arg Lys Ser Glu Glu Thr Ile 500 505 510 Thr Pro Trp Asn Phe Glu Glu Val Val Asp Lys Gly Ala Ser Ala Gln 515,520,525 Ser Phe Ile Glu Arg Met Thr Asn Phe Asp Lys Asn Leu Pro Asn Glu 530 535 540 Lys Val Leu Pro Lys Ser Leu Leu Tyr Glu Tyr Phe Thr Val Tyr 545 550 555 560 Asn Glu Leu Thr Lys Val Lys Tyr Val Thr Glu Gly Met Arg Lys Pro 565,570,575 Gly Glu Gln Lys Lys Wing Ile Val Asp 580,585,590 Lys Thr Asn Arg Lys Val Thr Val Lys Gln Leu Lys Glu Asp Tyr Phe 595,600,605 Lys Lys Ile Glu Cys Phe Asp Ser Val Glu Ile Ser Gly Val Glu Asp 610 615 620 Arg Phe Asn Ala Ser Leu Gly Thr Tyr His Asp Leu Leu Lys Ile Ile 625 630 635 640 Lys Asp Lys Asp Phe Leu Asp Asn Glu Glu Asn Glu Asp Ile Leu Glu 645 650 655 Asp Ile Val Leu Thr Leu Thr Leu Phe Glu Asp Arg Glu Met Ile Glu 660 665 670 Glu Arg Leu Lys Thr Tyr Ala His Leu Phe Asp Asp Lys Val Met Lys 675 680 685 Gln Leu Lys Arg Arg Arg Tyr Thr Gly Trp Gly Arg Leu Ser Arg Lys 690 695 700 Leu Ile Asn Gly Ile Arg Asp Lys Gln Ser Gly Lys Thr Ile Leu Asp 705 710 715 720 Phe Leu Lys Ser Asp Gly Phe Ala Asn Arg Asn Phe Met Gln Leu Ile 725 730 735 His Asp Asp Ser Leu Thr Phe Lys Glu Asp Ile Gln Lys Ala Gln Val 740 745 750 Ser Gly Gln Gly Asp Ser Leu His Glu His Ile Ala Asn Leu Ala Gly 755 760 765 Ser Pro Ala Ile Lys Lys Gly Ile Leu Gln Thr Val Lys Val Val Asp 770 775 780 Glu Leu Val Lys Val Met Gly Arg His Lys Pro Glu Asn Ile Val Ile 785 790 795 800 Glu Met Ala Arg Glu Asn Gln Thr Thr Gln Lys Gly Gln Lys Asn Ser 805 810 815 Arg Glu Arg Met Lys Arg Ile Glu Glu Gly Ile Lys Glu Leu Gly Ser 820 825 830 Gln Ile Leu Lys Glu His Pro Val Glu Asn Thr Gln Leu Gln Asn Glu 835 840 845 Lys Leu Tyr Leu Tyr Tyr Leu Gln Asn Gly Arg Asp Met Tyr Val Asp 850 855 860 Gln Glu Leu Asp Ile Asn Arg Leu Ser Asp Tyr Asp Val Asp His Ile 865 870 875 880 Val Pro Gln Ser Phe Leu Lys Asp Asp Ser Ile Asp Asn Lys Val Leu 885 890 895 Thr Arg Ser Asp Lys Asn Arg Gly Lys Ser Asp Asn Val Pro Ser Glu 900 905 910 Glu Val Val Lys Lys Met Lys Asn Tyr Trp Arg Gln Leu Leu Asn Ala 915 920 925 Lys Leu Ile Thr Gln Arg Lys Phe Asp Asn Leu Thr Lys Ala Glu Arg 930 935 940 Gly Gly Leu Ser Glu Leu Asp Lys Ala Gly Phe Ile Lys Arg Gln Leu 945 950 955 960 Val Glu Thr Arg Gln Ile Thr Lys His Val Ala Gln Ile Leu Asp Ser 965 970 975 Arg Met Asn Thr Lys Tyr Asp Glu Asn Asp Lys Leu Ile Arg Glu Val 980 985 990 Lys Val Ile Thr Leu Lys Ser Lys Leu Val Ser Asp Phe Arg Lys Asp 995 1000 1005 Phe Gln Phe Tyr Lys Val Arg Glu Ile Asn Asn Tyr His His Ala 1010 1015 1020 His Asp Ala Tyr Leu Asn Ala Val Val Gly Thr Ala Leu Ile Lys 1025 1030 1035 Lys Tyr Pro Lys Leu Glu Ser Glu Phe Val Tyr Gly Asp Tyr Lys 1040 1045 1050 Val Tyr Asp Val Arg Lys Met Ile Ala Lys Ser Glu Gln Glu Ile 1055 1060 1065 Gly Lys Ala Thr Ala Lys Tyr Phe Phe Tyr Ser Asn Ile Met Asn 1070 1075 1080 Phe Phe Lys Thr Glu Ile Thr Leu Ala Asn Gly Glu Ile Arg Lys 1085 1090 1095 Arg Pro Leu Ile Glu Thr Asn Gly Glu Thr Gly Glu Ile Val Trp 1100 1105 1110 Asp Lys Gly Arg Asp Phe Ala Thr Val Arg Lys Val Leu Ser Met 1115 1120 1125 Pro Gln Val Asn Ile Val Lys Lys Thr Glu Val Gln Thr Gly Gly 1130 1135 1140 Phe Ser Lys Glu Ser Ile Leu Pro Lys Arg Asn Ser Asp Lys Leu 1145 1150 1155 Ile Ala Arg Lys Lys Asp Trp Asp Pro Lys Lys Tyr Gly Gly Phe 1160 1165 1170 Asp Ser Pro Thr Val Ala Tyr Ser Val Leu Val Val Ala Lys Val 1175 1180 1185 Glu Lys Gly Lys Ser Lys Lys Leu Lys Ser Val Lys Glu Leu Leu 1190 1195 1200 Gly Ile Thr Ile Met Glu Arg Ser Ser Phe Glu Lys Asn Pro Ile 1205 1210 1215 Asp Phe Leu Glu Ala Lys Gly Tyr Lys Glu Val Lys Lys Asp Leu 1220 1225 1230 Ile Ile Lys Leu Pro Lys Tyr Ser Leu Phe Glu Leu Glu Asn Gly 1235 1240 1245 Arg Lys Arg Met Leu Ala Ser Ala Gly Glu Leu Gln Lys Gly Asn 1250 1255 1260 Glu Leu Ala Leu Pro Ser Lys Tyr Val Asn Phe Leu Tyr Leu Ala 1265 1270 1275 Ser His Tyr Glu Lys Leu Lys Gly Ser Pro Glu Asp Asn Glu Gln 1280 1285 1290 Lys Gln Leu Phe Val Glu Gln His Lys His Tyr Leu Asp Glu Ile 1295 1300 1305 Ile Glu Gln Ile Ser Glu Phe Ser Lys Arg Val Ile Leu Ala Asp 1310 1315 1320 Ala Asn Leu Asp Lys Val Leu Ser Ala Tyr Asn Lys His Arg Asp 1325 1330 1335 Lys Pro Ile Arg Glu Gln Ala Glu Asn Ile Ile His Leu Phe Thr 1340 1345 1350 Leu Thr Asn Leu Gly Ala Pro Ala Ala Phe Lys Tyr Phe Asp Thr 1355 1360 1365 Thr Ile Asp Arg Lys Arg Tyr Thr Ser Thr Lys Glu Val Leu Asp 1370 1375 1380 Ala Thr Leu Ile His Gln Ser Ile Thr Gly Leu Tyr Glu Thr Arg 1385 1390 1395 Ile Asp Leu Ser Gln Leu Gly Gly Asp Lys Arg Pro Ala Ala Thr 1400 1405 1410 Lys Lys Ala Gly Gln Ala Lys Lys Lys Lys 1415 1420 <210> 50 <211> 2012 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <400> 50 gaatgctgcc ctcagacccg cttcctccct gtccttgtct gtccaaggag aatgaggtct 60 cactggtgga ttcggacta ccctgaggag ctggcacctg agggacaagg ccccccacct 120 gcccagctcc agcctctgat gaggggtggg agagagctac atgaggttgc taagaaagcc 180 tcccctgaag gagaccacac agtgtgtgag gttggagtct ctagcagcgg gttctgtgcc 240 cccagggata gtctggctgt ccaggcactg ctcttgatat aaacaccacc tcctagttat 300 gaaaccatgc ccattctgcc tctctgtatg gaaaagagca tggggctggc ccgtggggtg 360 gtgtccactt taggccctgt gggagatcat gggaacccac gcagtgggtc ataggctctc 420 tcatttacta ctcacatcca ctctgtgaag aagcgattat gatctctcct ctagaaactc 480 gtagagtccc atgtctgccg gcttccagag cctgcactcc tccaccttgg cttggctttg 540 ctggggctag aggagctagg atgcacagca gctctgtgac cctttgtttg agaggaacag 600 gaaaaccacc cttctctctg gcccactgtg tcctcttcct gccctgccat ccccttctgt 660 gaatgttaga cccatgggag cagctggtca gaggggaccc cggcctgggg cccctaaccc 720 tatgtagcct cagtcttccc atcaggctct cagctcagcc tgagtgttga ggccccagtg 780 gctgctctgg gggcctcctg agtttctcat ctgtgcccct ccctccctgg cccaggtgaa 840 ggtgtggttc cagaaccgga ggacaaagta caaacggcag aagctggagg aggaagggcc 900 tgagtccgag cagaagaaga agggctccca tcacatcaac cggtggcgca ttgccacgaa 960 gcaggccaat ggggaggaca tcgatgtcac ctccaatgac aagcttgcta gcggtgggca 1020 accacaaacc cacgagggca gagtgctgct tgctgctggc caggcccctg cgtgggccca 1080 agctggactc tggccactcc ctggccaggc tttggggagg cctggagtca tggccccaca 1140 gggcttgaag cccggggccg ccattgacag agggacaagc aatgggctgg ctgaggcctg 1200 ggaccacttg gccttctcct cggagagcct gcctgcctgg gcgggcccgc ccgccaccgc 1260 agcctcccag ctgctctccg tgtctccaat ctcccttttg ttttgatgca tttctgtttt 1320 aatttatttt ccaggcacca ctgtagttta gtgatcccca gtgtccccct tccctatggg 1380 aataataaaa gtctctctct taatgacacg ggcatccagc tccagcccca gagcctgggg 1440 tggtagattc cggctctgag ggccagtggg ggctggtaga gcaaacgcgt tcagggcctg 1500 ggagcctggg gtggggtact ggtggagggg gtcaagggta attcattaac tcctctcttt 1560 tgttggggga ccctggtctc tacctccagc tccacagcag gagaaacagg ctagacatag 1620 ggaagggcca tcctgtatct tgagggagga caggcccagg tctttcttaa cgtattgaga 1680 ggtgggaatc aggcccaggt agttcaatgg gagagggaga gtgcttccct ctgcctagag 1740 actctggtgg cttctccagt tgaggagaaa ccagaggaaa ggggaggatt ggggtctggg 1800 ggagggaaca ccattcacaa aggctgacgg ttccagtccg aagtcgtggg cccaccagga 1860 tgctcacctg tccttggaga accgctgggc aggttgagac tgcagagaca gggcttaagg 1920 ctgagcctgc aaccagtccc cagtgactca gggcctcctc agcccaagaa agagcaacgt 1980 gccagggccc gctgagctct tgtgttcacc tg 2012 <210> 51 <211> 1153 <212> PRT <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polypeptide" <400> 51 Met Lys Arg Pro Ala Ala Thr Lys Lys Ala Gly Gln Ala Lys Lys Lys 1 5 10 15 Lys Ser Asp Leu Val Leu Gly Leu Asp Ile Gly Ile Gly Ser Val Gly 20 25 30 Val Gly Ile Leu Asn Lys Val Thr Gly Glu Ile Ile His Lys Asn Ser 35 40 45 Arg Ile Phe Pro Ala Ala Gln Ala Glu Asn Asn Leu Val Arg Arg Thr 50 55 60 Asn Arg Gln Gly Arg Arg Leu Ala Arg Arg Lys Lys His Arg Arg Val 65 70 75 80 Arg Leu Asn Arg Leu Phe Glu Glu Ser Gly Leu Ile Thr Asp Phe Thr 85 90 95 Lys Ile Ser Ile Asn Leu Asn Pro Tyr Gln Leu Arg Val Lys Gly Leu 100 105 110 Thr Asp Glu Leu Ser Asn Glu Glu Leu Phe Ile Ala Leu Lys Asn Met 115 120 125 Val Lys His Arg Gly Ile Ser Tyr Leu Asp Asp Ala Ser Asp Asp Gly 130 135 140 Asn Ser Ser Val Gly Asp Tyr Ala Gln Ile Val Lys Glu Asn Ser Lys 145 150 155 160 Gln Leu Glu Thr Lys Thr Pro Gly Gln Ile Gln Leu Glu Arg Tyr Gln 165 170 175 Thr Tyr Gly Gln Leu Arg Gly Asp Phe Thr Val Glu Lys Asp Gly Lys 180 185 190 Lys His Arg Leu Ile Asn Val Phe Pro Thr Ser Ala Tyr Arg Ser Glu 195 200 205 Ala Leu Arg Ile Leu Gln Thr Gln Gln Glu Phe Asn Pro Gln Ile Thr 210 215 220 Asp Glu Phe Ile Asn Arg Tyr Leu Glu Ile Leu Thr Gly Lys Arg Lys 225 230 235 240 Tyr Tyr His Gly Pro Gly Asn Glu Lys Ser Arg Thr Asp Tyr Gly Arg 245 250 255 Tyr Arg Thr Ser Gly Glu Thr Leu Asp Asn Ile Phe Gly Ile Leu Ile 260 265 270 Gly Lys Cys Thr Phe Tyr Pro Asp Glu Phe Arg Ala Ala Lys Ala Ser 275 280 285 Tyr Thr Ala Gln Glu Phe Asn Leu Leu Asn Asp Leu Asn Asn Leu Thr 290 295 300 Val Pro Thr Glu Thr Lys Lys Leu Ser Lys Glu Gln Lys Asn Gln Ile 305 310 315 320 Ile Asn Tyr Val Lys Asn Glu Lys Ala Met Gly Pro Ala Lys Leu Phe 325 330 335 Lys Tyr Ile Ala Lys Leu Leu Ser Cys Asp Val Ala Asp Ile Lys Gly 340 345 350 Tyr Arg Ile Asp Lys Ser Gly Lys Ala Glu Ile His Thr Phe Glu Ala 355 360 365 Tyr Arg Lys Met Lys Thr Leu Glu Thr Leu Asp Ile Glu Gln Met Asp 370 375 380 Arg Glu Thr Leu Asp Lys Leu Ala Tyr Val Leu Thr Leu Asn Thr Glu 385 390 395 400 Arg Glu Gly Ile Gln Glu Ala Leu Glu His Glu Phe Ala Asp Gly Ser 405 410 415 Phe Ser Gln Lys Gln Val Asp Glu Leu Val Gln Phe Arg Lys Ala Asn 420 425 430 Ser Ser Ile Phe Gly Lys Gly Trp His Asn Phe Ser Val Lys Leu Met 435 440 445 Met Glu Leu Ile Pro Glu Leu Tyr Glu Thr Ser Glu Glu Gln Met Thr 450 455 460 Ile Leu Thr Arg Leu Gly Lys Gln Lys Thr Thr Ser Ser Ser Asn Lys 465 470 475 480 Thr Lys Tyr Ile Asp Glu Lys Leu Leu Thr Glu Glu Ile Tyr Asn Pro 485 490 495 Val Val Ala Lys Ser Val Arg Gln Ala Ile Lys Ile Val Asn Ala Ala 500 505 510 Ile Lys Glu Tyr Gly Asp Phe Asp Asn Ile Val Ile Glu Met Ala Arg 515 520 525 Glu Thr Asn Glu Asp Asp Glu Lys Lys Ala Ile Gln Lys Ile Gln Lys 530 535 540 Ala Asn Lys Asp Glu Lys Asp Ala Ala Met Leu Lys Ala Ala Asn Gln 545 550 555 560 Tyr Asn Gly Lys Ala Glu Leu Pro His Ser Val Phe His Gly His Lys 565 570 575 Gln Leu Ala Thr Lys Ile Arg Leu Trp His Gln Gln Gly Glu Arg Cys 580 585 590 Leu Tyr Thr Gly Lys Thr Ile Ser Ile His Asp Leu Ile Asn Asn Ser 595 600 605 Asn Gln Phe Glu Val Asp His Ile Leu Pro Leu Ser Ile Thr Phe Asp 610 615 620 Asp Ser Leu Ala Asn Lys Val Leu Val Tyr Ala Thr Ala Asn Gln Glu 625 630 635 640 Lys Gly Gln Arg Thr Pro Tyr Gln Ala Leu Asp Ser Met Asp Asp Ala 645 650 655 Trp Ser Phe Arg Glu Leu Lys Ala Phe Val Arg Glu Ser Lys Thr Leu 660 665 670 Ser Asn Lys Lys Lys Glu Tyr Leu Leu Thr Glu Glu Asp Ile Ser Lys 675 680 685 Phe Asp Val Arg Lys Lys Phe Ile Glu Arg Asn Leu Val Asp Thr Arg 690 695 700 Tyr Ala Ser Arg Val Val Leu Asn Ala Leu Gln Glu His Phe Arg Ala 705 710 715 720 His Lys Ile Asp Thr Lys Val Ser Val Val Arg Gly Gln Phe Thr Ser 725 730 735 Gln Leu Arg Arg His Trp Gly Ile Glu Lys Thr Arg Asp Thr Tyr His 740 745 750 His His Ala Val Asp Ala Leu Ile Ile Ala Ala Ser Ser Gln Leu Asn 755 760 765 Leu Trp Lys Lys Gln Lys Asn Thr Leu Val Ser Tyr Ser Glu Asp Gln 770 775 780 Leu Leu Asp Ile Glu Thr Gly Glu Leu Ile Ser Asp Asp Glu Tyr Lys 785 790 795 800 Glu Ser Val Phe Lys Ala Pro Tyr Gln His Phe Val Asp Thr Leu Lys 805 810 815 Ser Lys Glu Phe Glu Asp Ser Ile Leu Phe Ser Tyr Gln Val Asp Ser 820 825 830 Lys Phe Asn Arg Lys Ile Ser Asp Ala Thr Ile Tyr Ala Thr Arg Gln 835 840 845 Ala Lys Val Gly Lys Asp Lys Ala Asp Glu Thr Tyr Val Leu Gly Lys 850 855 860 Ile Lys Asp Ile Tyr Thr Gln Asp Gly Tyr Asp Ala Phe Met Lys Ile 865 870 875 880 Tyr Lys Lys Asp Lys Ser Lys Phe Leu Met Tyr Arg His Asp Pro Gln 885 890 895 Thr Phe Glu Lys Val Ile Glu Pro Ile Leu Glu Asn Tyr Pro Asn Lys 900 905 910 Gln Ile Asn Glu Lys Gly Lys Glu Val Pro Cys Asn Pro Phe Leu Lys 915 920 925 Tyr Lys Glu Glu His Gly Tyr Ile Arg Lys Tyr Ser Lys Lys Gly Asn 930 935 940 Gly Pro Glu Ile Lys Ser Leu Lys Tyr Tyr Asp Ser Lys Leu Gly Asn 945 950 955 960 His Ile Asp Ile Thr Pro Lys Asp Ser Asn Asn Lys Val Val Leu Gln 965 970 975 Ser Val Ser Pro Trp Arg Ala Asp Val Tyr Phe Asn Lys Thr Thr Gly 980 985 990 Lys Tyr Glu Ile Leu Gly Leu Lys Tyr Ala Asp Leu Gln Phe Glu Lys 995 1000 1005 Gly Thr Gly Thr Tyr Lys Ile Ser Gln Glu Lys Tyr Asn Asp Ile 1010 1015 1020 Light Light Light Glu Gly Val Asp Ser Asp Ser Glu Phe Light Phe Thr 1025 1030 1035 Leu Tyr Lys Asn Asp Leu Leu Leu Val Lys Asp Thr Glu Thr Lys 1040 1045 1050 Glu Gln Gln Leu Phe Arg Phe Leu Ser Arg Thr Met Pro Lys Gln 1055 1060 1065 Lys His Tyr Val Glu Leu Lys Pro Tyr Asp Lys Gln Lys Phe Glu 1070 1075 1080 Gly Gly Glu Ala Leu Ile Lys Val Leu Gly Asn Val Ala Asn Ser 1085 1090 1095 Gly Gln Cys Lys Lys Gly Leu Gly Lys Ser Asn Ile Ser Ile Tyr 1100 1105 1110 Lys Val Arg Thr Asp Val Leu Gly Asn Gln His Ile Ile Lys Asn 1115 1120 1125 Glu Gly Asp Lys Pro Lys Leu Asp Phe Lys Arg Pro Ala Ala Thr 1130 1135 1140 Lys Lys Ala Gly Gln Ala Lys Lys Lys Lys 1145 1150 <210> 52 <211> 340 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <400> 52 gagggcctat ttcccatgat tccttcatat ttgcatatac gatacaaggc tgttagagag 60 aatattggaa ttaatttgac tgtaacaca agatattag tacaaatac gtgacgtaga 120 aagtaataat ttctgggta gtttgcagtt ttaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatctt gtggaagga 240 cgaaacaccg ttactttaat cttgcagaag ctacaagat aaggctcat gccgaatca 300 acaccctgtc atttttggc agggtgtttt cgttatta 340 <210> 53 <211> 360 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <220> <221> modified_base <222> (288)..(317) <223> a, c, t, g, unknown or other <400> 53 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 ataattggaa ttaatttgac tgtaaacaca aagatattag tacaaaatac gtgacgtaga 120 aagtaataat ttcttgggta gtttgcagtt ttaaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatatctt gtggaaagga 240 cgaaacaccg ggttttagag ctatgctgtt ttgaatggtc ccaaaacnnnn nnnnnnnnn 300 nnnnnnnnnnn nnnnnngtt ttagagctat gctgttttga atggtcccaa aacttttttt 360 <210> 54 <211> 318 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <220> <221> modified_base <222> (250)..(269) <223> a, c, t, g, unknown or other <400> 54 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 aatattggaa ttaatttgac tgtaacaca agatattag tacaaatac gtgacgtaga 120 aagtaataat ttctgggta gtttgcagtt ttaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatctt gtggaagga 240 cgaaacaccn nnnnnnnn nnnnnnn ttttagagct agaaatagca agttaaaata 300 aggctagtcc gtttttt 318 <210> 55 <211> 325 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <220> <221> modified_base <222> (250)..(269) <223> a, c, t, g, unknown or other <400> 55 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 aatattggaa ttaatttgac tgtaacaca agatattag tacaaatac gtgacgtaga 120 aagtaataat ttctgggta gtttgcagtt ttaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatctt gtggaagga 240 cgaaacaccn nnnnnnnn nnnnnnn ttttagagct agaaatagca agttaaaata 300 aggctagtcc gttatcatttt tttt 325 <210> 56 <211> 337 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <220> <221> modified_base <222> (250)..(269) <223> a, c, t, g, unknown or other <400> 56 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 aatattggaa ttaatttgac tgtaacaca agatattag tacaaatac gtgacgtaga 120 aagtaataat ttctgggta gtttgcagtt ttaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatctt gtggaagga 240 cgaaacaccn nnnnnnnn nnnnnnn ttttagagct agaaatagca agttaaaata 300 aggctagtcc gttatcact tgaaaaagtg tttttt 337 <210> 57 <211> 352 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <220> <221> modified_base <222> (250)..(269) <223> a, c, t, g, unknown or other <400> 57 gaggggcctat ttcccatgat tccttcatat tgcatatac gatacaggc tgttagagag 60 aatattggaa ttaatttgac tgtaacaca agatattag tacaaatac gtgacgtaga 120 aagtaataat ttctgggta gtttgcagtt ttaaattat gttttaaaat ggactatcat 180 atgcttaccg taacttgaaa gtatttcgat ttcttggctt tatatctt gtggaagga 240 cgaaacaccn nnnnnnnn nnnnnnn ttttagagct agaaatagca agttaaaata 300 aggctagtcc gttatcact tgaaaagtg gcaccgagtc gtgcttttt tt 352 <210> 58 <211> 5101 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <400> 58 cgttacataa cttacggtaa atggcccgcc tggctgaccg cccaacgacc cccgcccatt 60 gacgtcaata atgacgtatg ttcccatagt aacgccaata gggactttcc attgacgtca 120 atgggtggag tatttacggt aaactgccca cttggcagta catcaagtgt atcatatgcc 180 aagtacgccc cctattgacg tcaatgacgg taaatggccc gcctggcatt atgcccagta 240 catgacctta tgggactttc ctacttggca gtacatctac gtattagtca tcgctattac 300 catggtcgag gtgagcccca cgttctgctt cactctcccc atctcccccc cctccccacc 360 cccaattttg tatttattta ttttttaatt attttgtgca gcgatggggg cggggggggg 420 gggggggcgc gcgccaggcg gggcggggcg gggcgagggg cggggcgggg cgaggcggag 480 aggtgcggcg gcagccaatc agagcggcgc gctccgaaag tttcctttta tggcgaggcg 540 gcggcggcgg cggccctata aaaagcgaag cgcgcggcgg gcgggagtcg ctgcgacgct 600 gccttcgccc cgtgccccgc tccgccgccg cctcgcgccg cccgccccgg ctctgactga 660 ccgcgttact cccacaggtg agcgggcggg acggcccttc tcctccgggc tgtaattagc 720 tgagcaagag gtaagggttt aagggatggt tggttggtgg ggtattaatg ttataattacc 780 tggagcacct gcctgaaatc actttttttc aggttggacc ggtgccacca tggactataa 840 ggaccacgac ggagactaca aggatcatga tattgattac aaagacgatg acgataagat 900 ggccccaaag aagaagcgga aggtcggtat ccacggagtc ccagcagccg acaagaagta 960 cagcatcggc ctggacatcg gcaccaactc tgtgggctgg gccgtgatca ccgacgagta 1020 caaggtgccc agcaagaaat tcaaggtgct gggcaacacc gaccggcaca gcatcaagaa 1080 gaacctgatc ggagccctgc tgttcgacag cggcgaaaca gccgaggcca cccggctgaa 1140 gagaaccgcc agaagaagat acaccagacg gaagaaccgg atctgctatc tgcaagagat 1200 cttcagcaac gagatggcca aggtggacga cagcttcttc cacagactgg aagagtcctt 1260 cctggtggaa gaggataaga agcacgagcg gcaccccatc ttcggcaaca tcgtggacga 1320 ggtggcctac cacgagaagt accccaccat ctaccacctg agaaagaaac tggtggacag 1380 caccgacaag gccgacctgc ggctgatcta tctggccctg gcccacatga tcaagttccg 1440 gggccacttc ctgatcgagg gcgacctgaa ccccgacaac agcgacgtgg acaagctggtt 1500 catccagctg gtgcagacct acaaccagct gttcgaggaa aaccccatca acgccagcgg 1560 cgtggacgcc aaggccatcc tgtctgccag actgagcaag agcagacggc tggaaaatct 1620 gatcgcccag ctgcccggcg agaagaagaa tggcctgttc ggcaacctga ttgccctgag 1680 cctgggcctg acccccaact tcaagagcaa cttcgacctg gccgaggatg ccaaactgca 1740 gctgagcaag gacacctacg aggacgacct ggacaacctg ctggcccaga tcggcgacca 1800 gtacgccgac ctgtttctgg ccgccaagaa cctgtccgac gccatcctgc tgagcgacat 1860 cctgagagtg aacaccgaga tcaccaggc cccctgagc gccctatga tcaagagata 1920 cgacgagcac caccaggacc tgaccctgct gaagctctc gtgcggcagc agctgcctga 1980 gaagtacaa gagattttct tcgaccagag caagaacggc tacgccggct acatgacgg 2040 cggagccagc caggaaggt tctacaagtt catcaagccc atcctggaaa agatggacgg 2100 caccgaggaa ctgctcgtga agctgaacag agaggacctg ctgcggaagc agcggacctt 2160 cgacaacggc agcatccccc accagatcca cctgggagag ctgcacgcca ttctgcggcg 2220 gcaggaagat ttttaccat tcctgaagga haaccgggaa aagatcgaga agatcctgac 2280 cttccgcatc ccctactacg tgggccctct ggccagggga aacagcagat tcgcctggat 2340 gaccagaaag agcgaggaaa ccatcacccc ctggaacttc gaggaagtgg tggacaaggg 2400 cgcttccgcc cagagcttca tcgagcggat gaccaacttc cataagaacc tgcccacga 2460 gaaggtgctg cccaagcaca gcctgctgta cgagtacttc accgtgtata acgagctgac 2520 caaagtgaaa tacgtgaccg agggaatgag aaagcccgcc ttcctgagcg gcgagcagaa 2580 aaaggccatc gtggacctgc tgttcaagac caaccggaaa gtgaccgtga agcagctgaa 2640 agaggactac ttcaagaaaa tcgagtgctt cgactccgtg gaaatctccg gcgtggaaga 2700 tcggttcaac gcctccctgg gcacatacca cgatctgctg aaaattatca aggacaagga 2760 cttcctggac aatgaggaaa acgaggacat tctggaagat atcgtgctga ccctgacact 2820 gtttgaggac agagagatga tcgaggaacg gctgaaaacc tatgcccacc tgttcgacga 2880 caaagtgatg aagcagctga agcggcggag atacaccggc tggggcaggc tgagccggaa 2940 gctgatcaac ggcatccggg acaagcagtc cggcaagaca atcctggatt tcctgaagtc 3000 cgacggcttc gccaacagaa acttcatgca gctgatccac gacgacagcc tgacctttaa 3060 agaggacatc cagaaagcccc aggtccgg ccaggcgat agcctgcacg agcacattgc 3120 caatctggcc ggcagccccg ccattagaa gggcatcctg cagacagtga aggtggtgga 3180 cgagctcgtg aaagtgatgg gccggcacaa gcccgagaac atcgtcg aaatggccag 3240 aggaaccag accaccaga agggagaga gaacagccgc gagagaatga agcggatcga 3300 agagggcatc aaagagctgg gcagccagat cctgaaagaa caccccgtgg aaaacaccca 3360 gctgcagaac gagaagctgt acctgtacta cctgcagaat gggcgggata tgtacgtgga 3420 ccaggaactg gatacacc ggctgtccga ctacgatgtg gaccatatcg tgcctcagag 3480 ctttctgaag gacgactcca tcgacaacaa ggtgctgacc agaagcgaca agaccgggg 3540 CAagagcgac aacgtgccct ccgagaggt cgtgagaag atgagaac actggcggca 3600 gctgctgaac gccaagctga ttaccagag aaagttcgac aatctgacca aggccgagag 3660 aggcggcctg agcgaactgg ataaggccgg cttcatcaag agacagctgg tggaaacccg 3720 gcagatcaca aagcacgtgg cacagatcct ggactcccgg atgaacacta agtacgacga 3780 gaatgacaag ctgatccggg aagtgaaagt gatcaccctg aagtccaagc tggtgtccga 3840 tttccggaag gatttccagt tttacaaagt gcgcgagatc aacaactacc accacgccca 3900 cgacgcctac ctgaacgccg tcgtgggaac cgccctgatc aaaaagtacc ctaagctgga 3960 aagcgagttc gtgtacggcg actacaaggt gtacgacgtg cggaagatga tcgccaagag 4020 cgagcaggaa atcggcaagg ctaccgccaa gtacttcttc tacagcaaca tcatgaactt 4080 tttcaagacc gagattaccc tggccaacgg cgagatccgg aagcggcctc tgatcgagac 4140 aaacggcgaa accggggaga tcgtgtggga taagggccgg gattttgcca ccgtgcggaa 4200 agtgctgagc atgccccaag tgaatatcgt gaaaaagacc gaggtgcaga caggcggctt 4260 cagcaaagag tctatcctgc ccaagaggaa cagcgataag ctgatcgcca gaaagaagga 4320 ctgggaccct aagaagtacg gcggcttcga cagccccacc gtggcctatt ctgtgctggt 4380 ggtggccaaa gtggaaaagg gcaagtccaa gaaactgaag agtgtgaaag agctgctggg 4440 gatcaccatc atggaaagaa gcagcttcga gaagaatccc atcgactttc tggaagccaa 4500 gggctacaaa gaagtgaaaa aggacctgat catcaagctg cctaagtact ccctgttcga 4560 gctggaaaac ggccggaaga gaatgctggc ctctgccggc gaactgcaga agggaaacga 4620 actggccctg ccctccaaat atgtgaactt cctgtacctg gccagccact atgagaagct 4680 gaagggctcc cccgaggata atgagcagaa acagctgttt gtggaacagc acaagcacta 4740 cctggacgag atcatcgagc agatcagcga gttctccaag agagtgatcc tggccgacgc 4800 taatctggac aaagtgctgt ccgcctacaa caagcaccgg gataagccca tcagagagca 4860 ggccgagaat atcatccacc tgtttaccct gaccaatctg ggagcccctg ccgccttcaa 4920 gtactttgac accaccatcg accggaagag gtacaccagc accaaagagg tgctggacgc 4980 caccctgatc caccagagca tcaccggcct gtacgagaca cggatcgacc tgtctcagct 5040 gggaggcgac tttctttttc ttagcttgac cagctttctt agtagcagca ggacgcttta 5100 a 5101 <210> 59 <211> 137 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 59 nnnnnnnnnn nnnnnnnnnn gtttttgtac tctcaagatt tagaaataaa tcttgcagaa 60 gctacaaaga taaggcttca tgccgaaatc aacaccctgt cattttatgg cagggtgttt 120 tcgttattta atttttt 137 <210> 60 <211> 123 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 60 nnnnnnnnnn nnnnnnnnnn gtttttgtac tctcagaaat gcagaagcta caaagataag 60 gcttcatgcc gaaatcaaca ccctgtcatt ttatggcagg gtgttttcgt tatttaattt 120 ttt 123 <210> 61 <211> 110 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 61 nnnnnnnnnn nnnnnnnnnn gtttttgtac tctcagaaat gcagaagcta caaagataag 60 gcttcatgcc gaaatcaaca ccctgtcatt ttatggcagg gtgttttttt 110 <210> 62 <211> 137 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 62 nnnnnnnnnn nnnnnnnnnn gttattgtac tctcaagatt tagaaataaa tcttgcagaa 60 gctacaaaga taaggcttca tgccgaaatc aacaccctgt cattttatgg cagggtgttt 120 tcgttattta atttttt 137 <210> 63 <211> 123 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 63 nnnnnnnnnn nnnnnnnnnn gttattgtac tctcagaaat gcagaagcta caaagataag 60 gcttcatgcc gaaatcaaca ccctgtcatt ttatggcagg gtgttttcgt tatttaattt 120 ttt 123 <210> 64 <211> 110 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 64 nnnnnnnnnn nnnnnnnnnn gttattgtac tctcagaaat gcagaagcta caaagataag 60 gcttcatgcc gaaatcaaca ccctgtcatt ttatggcagg gtgttttttt 110 <210> 65 <211> 137 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 65 nnnnnnnnnn nnnnnnnnnn gttattgtac tctcaagatt tagaaataaa tcttgcagaa 60 gctacaatga taaggcttca tgccgaaatc aacaccctgt cattttatgg cagggtgttt 120 tcgttattta atttttt 137 <210> 66 <211> 123 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 66 nnnnnnnnnn nnnnnnnnnn gttattgtac tctcagaaat gcagaagcta caatgataag 60 gcttcatgcc gaaatcaaca ccctgtcatt ttatggcagg gtgttttcgt tatttaattt 120 ttt 123 <210> 67 <211> 110 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 67 nnnnnnnnnn nnnnnnnnnn gttattgtac tctcagaaat gcagaagcta caatgataag 60 gcttcatgcc gaaatcaaca ccctgtcatt ttatggcagg gtgttttttt 110 <210> 68 <211> 107 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 68 nnnnnnnnnnn nnnnnnnnn gttttagagc tgtggaaaca cagcgagtta aaataaggct 60 tagtccgtac tcaacttgaa aaggtggcac cgattcggtg tttttt 107 <210> 69 <211> 4263 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <400> 69 atgaaaagc cggcggccac gaaaaaggcc ggccaggcaa aaaagaaaaa gaccaagcccc 60 tacagcatcg gcctggacat cggcaccaat agcgtgggct gggccgtgac caccgacaac 120 tacaaggtgc ccagcaagaa aatgaaggtg ctgggcaaca cctccaagaa gtacatcaag 180 aaaaacctgc tgggcgtgct gctgttcgac agcggcatta cagccgaggg cagacggctg 240 aagagaaccg ccagacggcg gtacacccgg cggagaaaca gaatcctgta tctgcaagag 300 atcttcagca ccgagatggc tacctggac gacgccttct tccagcggct ggacgacagc 360 ttcctggtgc ccgacgacaa gcgggacagc aagcccca tcttcggcaa cctggtggaa 420 gagaaggcct accacgacga gttccccacc atctaccacc tgagaaagta cctggccgac 480 agcaccaga aggccgacct gagactggtg tatctggccc tggcccacat gatcaagtac 540 cggggccact tcctgatcga gggcgagttc aacagcaga acacgacat ccagagaac 600 ttccaggact tcctggacac ctacaacgcc atcttcgaga gcgacctgtc cctggaaaac 660 agcaagcagc tggagagat cgtgaggac aagatcagca agctgaaaa gaggaccgc 720 atcctgaagc tgttccccgg cgagagaac agcggaatct tcagcgagtt tctgaagctg 780 atcgtgggca accaggccga cttcagaag tgctcaacc tggacgaga agccagcctg 840 cactcagca aagagagcta cgacgaggac ctggaaccc tgctgggata tatcggcgac 900 gactacagcg acgtgttcct gaaggccaag aagctctacg acgctatcct gctgagcggc 960 ttcctgaccg tgaccgacaa cgagacagag gccccactga gcagcgccat gattaagcgg 1020 tacaacgagc acaaagaga tctggctctg ctgaaagagt acatccggaa catcagcctg 1080 aaaacctaca atgaggtgtt caaggacgacgac accaagaacg gctacgccgg ctacatcgac 1140 ggcaagacca accaggaaga tttctatgtg tacctagaagaagctgctggc cgagttcgag 1200 ggggccgact actttctgga aaaaatcgac cgcgaggatt tcctgcggaa gcagcggacc 1260 ttcgacaacg gcagcatccc ctaccagatc catctgcagg aaatgcgggc catcctggac 1320 aagcaggcca agttctaccc attcctggcc aagaacaaag agcggatcga gaagatcctg 1380 accttccgca tcccttacta cgtgggcccc ctggccagag gcaacagcga ttttgcctgg 1440 1500 gagtccagcg ccgaggcctt catcaccgg atgaccagct tcgacctgta cctgcccgag 1560 gaaaaggtgc tgcccaagca cagcctgctg tacgagacat tcaatgtgta taacgagctg 1620 accaaagtgc gtttatcgc cgagtctatg cgggactacc agttcctgga ctccaagcag 1680 aaaaaggaca tcgtgcggct gtactcaag gatagcgga aagtgaccga taggacatc 1740 atcgagtacc tgcacgccat ctacggctac gatggcatcg agctgaaggg catcgagaag 1800 cagttcaact ccagcctgag cacataccac gacctgctga acattatca cgacaagaa 1860 tttctggacg actccagcaa cgaggccatc atcgaagaga tcatccacac cctgaccacc 1920 tttgaggacc gcgagatgat caagcagcgg ctgagcaagt tcgagaacat ctcgacaag 1980 agcgtgctga aaagctgag cagacggcac tacaccggct ggggcaagct gagcgccaag 2040 ctgatcaacg gcatccggga cgagaagtcc ggcaacacaa tcctggacta cctgatcgac 2100 gacggcatca gcaaccggaa cttcatgcag ctgatccacg acgacgccct gagcttcaag 2160 aagaagatcc agaaggccca gatcatcgggg gaxgagaca agggxacat caaaagtc 2220 gtgaagtccc tgcccggcag ccccgccatc agaagggaa tcctgcag catcagatc 2280 gtggacgagc tcgtgaaagt gatgggcggc agaagcccg agagcatcgt ggtggaaatg 2340 gctagagaga accagtacac caatcagggc agagcaca gccagcagag actgagaga 2400 ctggaaaagt ccctgaaaga gctgggcagc aagattctga aagaatat ccctgccaag 2460 ctgtccaaga tcgacaacaa cgccctgcag aacgaccggc tgtacctgta ctacctgcag 2520 aatggcagg acatgtatac aggcgacgac ctggatatcg accgcctgag acacacgac 2580 atcgaccata ttatccccca ggccttcctg aaagacaaca gcattgacaa caagtgctg 2640 gtgtcctccg ccagcaccg cggcaagtcc gatgatgtgc ccagcctgga agtcgtgaaa 2700 aagagaaaga ccttctgta tcagctgctg aaagcaagc tgattagcca gaggaagttc 2760 vakaacctga ccaaggccga gagaggcggc ctgagccctg aagataggc cggcttcatc 2820 cagagacagc tggtggaaac ccggcagatc accaagcacg tggccagact gctggatgag 2880 aagttttaca acaagaagga cgagacaac cggggccgtgc ggaccgtgaa gatcatcacc 2940 ctgaagtcca ccctgtgtc ccagttccgg aaggactcg agctgtataa agtgcgcgag 3000 atcaatgact ttcaccacgc ccacgacgcc tacctgaatg ccgtggtggc ttccgccctg 3060 ctgaagaagt accctaagct ggaacccgag ttcgtgtacg gcgactaccc caagtacaac 3120 tccttcagag agcggaagtc cgccaccgag aaggtgtact tctactccaa catcatgaat 3180 atctttaaga agtccatctc cctggccgat ggcagagtga tcgagcggcc cctgatcgaa 3240 gtgaacgaag agacaggcga gagcgtgtgg aaaaaaa gcgacctggc caccgtgcgg 3300 cgggtgctga gttatcctca agtgaatgtc gtgaagaagg tggaagaaca gaaccacggc 3360 ctggatcggg gcaagcccaa gggcctgttc aacgccaacc tgtccagcaa gcctaagccc 3420 aactccaacg agaatctcgt gggggccaaa gagtacctgg accctaagaa gtacggcgga 3480 tacgccggca tctccaatag cttcaccgtg ctcgtgaagg gcacaatcga gaagggcgct 3540 aagaaaaaga tcacaaacgt gctggaattt caggggatct ctatcctgga ccggatcaac 3600 taccggaagg ataagctgaa ctttctgctg gaaaaaggct acaaggacat tgagctgatt 3660 atcgagctgc ctaagtactc cctgttcgaa ctgagcgacg gctccagacg gatgctggcc 3720 tccatcctgt ccaccaacaa caagcggggc gagatccaca agggaaacca gatcttcctg 3780 agccagaaat ttgtgaaact gctgtaccac gccaagcgga tctccaacac catcaatgag 3840 aaccaccgga aatacgtgga aaaccacaag aaagagtttg aggaactgtt ctactacatc 3900 ctggagttca acgagaacta tgtgggagcc aagaagaacg gcaaactgct gaactccgcc 3960 ttccagagct ggcagaacca cagcatcgac gagctgtgca gctccttcat cggccctacc 4020 ggcagcgagc ggaagggact gtttgagctg acctccagag gctctgccgc cgactttgag 4080 ttcctgggag tgaagatccc ccggtacaga gactacaccc cctctagtct gctgaaggac 4140 gccaccctga tccaccagag cgtgaccggc ctgtacgaaa cccggatcga cctggctaag 4200 ctgggcgagg gaaagcgtcc tgctgctact aagaaagctg gtcaagctaa gaaaaagaaa 4260 taa 4263 <210> 70 <211> 84 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 70 ggaaccattc ataacagcat agcaagttat aataaggcta gtccgttatc aacttgaaaa 60 agtggcaccg agtcggtgct tttt 84 <210> 71 <211> 36 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 71 gttatagagc tatgctgtta tgaatggtcc caaaac 36 <210> 72 <211> 84 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 72 ggaaccattc aatacagcat agcaagttaa tataaggcta gtccgttatc aacttgaaaa 60 agtggcaccg agtcggtgct tttt 84 <210> 73 <211> 36 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic oligonucleotide" <400> 73 gtattagagc tatgctgtat tgaatggtcc caaaac 36 <210> 74 <211> 103 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 74 nnnnnnnnnn nnnnnnnnnn gttttagagc tagaaatagc aagttaaaat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgctttt ttt 103 <210> 75 <211> 103 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 75 nnnnnnnnnn nnnnnnnnnn gtattagagc tagaaatagc aagttaatat aaggctagtc 60 cgttatcaac ttgaaaaagt ggcaccgagt cggtgctttt ttt 103 <210> 76 <211> 123 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide" <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 76 nnnnnnnnnnn nnnnnnnnn gttttagagc tatgctgttt tggaaacaaa acagcatagc 60 aagttaaaat aaggctagtc cgttatcaac ttgaaaaagt ggcaccgagt cggtgcttttt 120 ttt 123 <210> 77 <211> 123 <212> DNA <213> Artificial Sequence <220> <221> source <223> / note="Description of Artificial Sequence: Synthetic polynucleotide” <220> <221> modified_base <222> (1)..(20) <223> a, c, t, g, unknown or other <400> 77 nnnnnnnnnnn nnnnnnnnn gtattagagc tatgctgtat tggaaacaat acagcatagc 60 aagttaatat aaggctagtc cgttatcaac ttgaaaaagt ggcaccgagt cggtgctttt 120 ttt 123 <210> 78 <211> 20 <212> DNA <213> Homo sapiens <400> 78 gtcacctcca atgactaggg 20 <210> 79 <211> 984 <212> PRT <213> Campylobacter jejuni <400> 79 Met Ala Arg And Leu Ala Phe Asp And Gly And Ser Ser And Gly Trp 1 5 10 15 Ala Phe Ser Glu Asn Asp Glu Leu Lys Asp Cys Gly Val Arg Ile Phe 20 25 30 Thr Lys Val Glu Asn Pro Lys Thr Gly Glu Ser Leu Ala Leu Pro Arg 35 40 45 Arg Leu Ala Arg Ser Ala Arg Lys Arg Leu Ala Arg Arg Lys Ala Arg 50 55 60 Leu Asn His Leu Lys His Leu Ile Ala Asn Glu Phe Lys Leu Asn Tyr 65 70 75 80 Glu Asp Tyr Gln Ser Phe Asp Glu Ser Leu Ala Lys Ala Tyr Lys Gly 85 90 95 Ser Leu Ile Ser Pro Tyr Glu Leu Arg Phe Arg Ala Leu Asn Glu Leu 100 105 110 Lys Gln Asp Phe Ala Arg Val Ile Leu His Ile Ala Lys Arg 115 120 125 Arg Gly Tyr Asp Asp Ile Lys Asn Ser Asp Asp Lys Glu Lys Gly Ala 130 135 140 With Lys Only Only Only With Lys On Gln Asn Glu Glu Lys With Only On Tyr Gln 145 150 155 160 Ser Val Gly Glu Tyr Leu Tyr Lys Glu Tyr Phe Gln Lys Phe Lys Glu 165 170 175 Asn Ser Lys Glu Phe Thr Asn Val Arg Asn Lys Glu Ser Tyr Glu 180 185 190 Arg Cys contains Gln Ser with Lys Asp and Glu with Lys 195 200 205 Lys Lys Gln Arg Glu Phe Gly Phe Ser Phe Ser Lys Phe Glu Glu 210 215 220 Glu Val Leu Ser Val Ala Phe Tyr Lys Arg Ala Leu Lys Asp Phe Ser 225 230 235 240 His Leu Val Gly Asn Cys Ser Phe Phe Thr Asp Glu Lys Arg Ala Pro 245 250 255 Lys Asn Ser Pro Leu Ala Phe Met Phe Val Ala Leu Thr Arg Ile Ile 260 265 270 Asn Leu Leu Asn Asn Leu Lys Asn Thr Glu Gly Ile Leu Tyr Thr Lys 275 280 285 Asp Asp Leu Asn Ala Leu Leu Asn Glu Val Leu Lys Asn Gly Thr Leu 290 295 300 Thr Tyr Lys Gln Thr Lys Lys Leu Leu Gly Leu Ser Asp Asp Tyr Glu 305 310 315 320 Phe Lys Gly Glu Lys Gly Thr Tyr Phe Ile Glu Phe Lys Lys Tyr Lys 325 330 335 Glu Phe Ile Lys Ala Leu Gly Glu His Asn Leu Ser Gln Asp Asp Leu 340 345 350 Asn Glu Ile Ala Lys Asp Ile Thr Leu Ile Lys Asp Glu Ile Lys Leu 355 360 365 Lys Lys Ala Leu Ala Lys Tyr Asp Leu Asn Gln Asn Gln Ile Asp Ser 370 375 380 Leu Ser Lys Leu Glu Phe Lys Asp His Leu Asn Ile Ser Phe Lys Ala 385 390 395 400 Leu Lys Leu Val Thr Pro Leu Met Leu Glu Gly Lys Lys Tyr Asp Glu 405 410 415 Ala Cys Asn Glu Leu Asn Leu Lys Val Ala Ile Asn Glu Asp Lys Lys 420 425 430 Asp Phe Leu Pro Ala Phe Asn Glu Thr Tyr Tyr Lys Asp Glu Val Thr 435 440 445 Asn Pro Val Val Leu Arg Ala Ile Lys Glu Tyr Arg Lys Val Leu Asn 450 455 460 Ala Leu Leu Lys Lys Tyr Gly Lys Val His Lys Ile Asn Ile Glu Leu 465 470 475 480 Ala Arg Glu Val Gly Lys Asn His Ser Gln Arg Ala Lys Ile Glu Lys 485 490 495 Glu Gln Asn Glu Asn Tyr Lys Ala Lys Lys Asp Ala Glu Leu Glu Cys 500 505 510 Glu Lys Leu Gly Leu Lys Ile Asn Ser Lys Asn Ile Leu Lys Leu Arg 515 520 525 Leu Phe Lys Glu Gln Lys Glu Phe Cys Ala Tyr Ser Gly Glu Lys Ile 530 535 540 Lys Ile Ser Asp Leu Gln Asp Glu Lys Met Leu Glu Ile Asp His Ile 545 550 555 560 Tyr Pro Tyr Ser Arg Ser Phe Asp Asp Ser Tyr Met Asn Lys Val Leu 565 570 575 Val Phe Thr Lys Gln Asn Gln Glu Lys Leu Asn Gln Thr Pro Phe Glu 580 585 590 Ala Phe Gly Asn Asp Ser Ala Lys Trp Gln Lys Ile Glu Val Leu Ala 595 600 605 Lys Asn Leu Pro Thr Lys Lys Gln Lys Arg Ile Leu Asp Lys Asn Tyr 610 615 620 Lys Asp Lys Glu Gln Lys Asn Phe Lys Asp Arg Asn Leu Asn Asp Thr 625 630 635 640 Arg Tyr Ile Ala Arg Leu Val Leu Asn Tyr Thr Lys Asp Tyr Leu Asp 645 650 655 Phe Leu Pro Leu Ser Asp Asp Glu Asn Thr Lys Leu Asn Asp Thr Gln 660 665 670 Lys Gly Ser Lys Val His Val Glu Ala Lys Ser Gly Met Leu Thr Ser 675 680 685 Ala Leu Arg His Thr Trp Gly Phe Ser Ala Lys Asp Arg Asn Asn His 690 695 700 Leu His His Ala Ile Asp Ala Val Ile Ile Ala Tyr Ala Asn Asn Ser 705 710 715 720 Ile Val Lys Ala Phe Ser Asp Phe Lys Lys Glu Gln Glu Ser Asn Ser 725 730 735 Glu Leu Tyr Ala Lys Lys Ser Glu Leu Asp Tyr Lys Asn Lys 740,745,750 Arg Lys Phe Phe Glu Pro Phe Ser Gly Phe Arg Gln Lys Val Leu Asp 755,760,765 Lys Ile Asp Glu Ile Phe Val Ser Lys Pro Glu Arg Lys Lys Pro Ser 770,775,780 Gly Ala Leu His Glu Glu Thr Phe Arg Lys Glu Glu Glu Phe Tyr Gln 785,790,795,800 Ser Tyr Gly Gly Lys Glu Gly Val Leu Lys Ala Leu Glu Leu Gly Lys 805 810 815 Ile Arg Lys Val Asn Gly Lys Ile Val Lys Asn Gly Asp Met Ph...
Claims
1. An engineered CRISPR-Cas9 chimeric RNA containing the following nucleotide sequence: NNNNNNNNNNNNNNNNNNNNNNGUUUUUAGAGCUAGAAAUAGCAAGUUAAAAAAAGGCUAGUCCGUUAUCA.
2. Engineered CRISPR-Cas9 chimeric RNA containing the following nucleotide sequence: NNNNNNNNNNNNNNNNNNNNNNGUUUUUAGCUAGAAAUAGCAAGUUAAAAUAAGCUAGUCCGUUAUCAACUUGAAAAGUG.
3. The engineered CRISPR-Cas9 chimeric RNA according to claim 1 or 2, wherein NNNNNNNNNNNNNNNNNNNNNN is a guide sequence that can hybridize to a target sequence in a eukaryotic cell adjacent to a protospacer adjacent motif (PAM).
4. The engineered CRISPR-Cas9 chimeric RNA according to claim 3, wherein PAM is NGG.
5. The engineered CRISPR-Cas9-based chimeric RNA according to claim 1 or 2, wherein the chimeric RNA further comprises a poly-U sequence.
6. An engineered CRISPR-Cas9-based chimeric RNA according to claim 1 or 2, wherein the chimeric RNA comprises one or more modified nucleotides.
7. The engineered CRISPR-Cas9-based chimeric RNA according to claim 1 or 2, wherein the chimeric RNA comprises one or more methylated nucleotides or nucleotide analogs.
8. Use of the engineered CRISPR-Cas9 chimeric RNA according to claim 1 or 2 for genome engineering, provided that such use does not involve the modification of the genetic identity of a human germline, and such use is not a method for surgical or therapeutic treatment of the human or animal body.
9. Use of the engineered CRISPR-Cas9 chimeric RNA according to claim 1 or 2 in the production of a non-human transgenic animal or transgenic plant.
Citation Information
Patent Citations
Engineering of systems, methods and optimized guide compositions for sequence manipulation
JP2016093196A
crispr-cas systems and methods for altering expression of gene products
JP2016500262A
Systems, methods and optimization guide compositions for sequence engineering
JP2023040015A
Isolation of exogenous recombinant proteins from the milk of transgenic mammals
US4873316A