Fusion Molecules Targeting VEGF and Angiopoietin and Their Uses
By developing a polypeptide that combines VEGF and angiopoietin domains and using gene therapy technology, the problem of repeated injection of existing drugs for treating abnormal angiogenic diseases has been solved, achieving longer-term and safer therapeutic effects.
Patent Information
- Application Number
- CN202280031671.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-31
- Filing Date
- 2022-03-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-03-30
AI Technical Summary
Existing drugs for treating abnormal angiogenesis-related diseases require repeated injections, which may lead to an increased risk of inflammation and infection, and the patient's compliance and persistence are low, resulting in poor treatment results.
A polypeptide was developed that combines the domain of VEGF and angiopoietin for the treatment of gene therapy. The polypeptide comprises the domains of VEGF receptor-1 and VEGFR-2, as well as the domain of angiopoietin, and is administered by gene therapy techniques.
This polypeptide can be expressed for a long time through gene therapy, reducing the frequency of drug injections, improving the persistence and effectiveness of treatment, reducing the risk of inflammation and infection, and improving patient compliance and treatment effectiveness.
Smart Images

Figure BDA0004517268230000201 
Figure BDA0004517268230000211 
Figure BDA0004517268230000251
Abstract
Description
[0001] Related Applications
[0002] This invention claims the benefit of PCT Patent Application No. PCT / CN2021 / 084559, filed on March 31, 2021, which is incorporated herein by reference in its entirety.
[0003] Reference to an Electronically Submitted Sequence Listing
[0004] This patent application contains a sequence listing that has been electronically submitted via EFS-Web as an ASCII formatted sequence listing having a file name of "14652-025-228_SEQ_LISTING.txt", a creation date of March 18, 2022, and a size of 217,975 bytes. The sequence listing submitted via EFS-Web is part of this specification and is incorporated herein by reference in its entirety. 1. Technical Field
[0005] The present disclosure relates to fusion molecules (e.g., polypeptides) that bind to both vascular endothelial growth factor (VEGF) and angiopoietins (e.g., angiopoietin 1 and angiopoietin 2), gene therapies based on these fusion molecules, and methods of use thereof. 2. Background Art
[0006] The VEGF / VEGFR and angiopoietin / Tie-2 signal transduction pathways are important in the process of vascular endothelial growth (angiogenesis) and in the maintenance of angiogenesis-related blood vessels. See Biela and Siemann, Cancer Lett. 380(2):525–533 (2016). Aberrant angiogenesis is involved in a number of conditions such as diabetic retinopathy, psoriasis, exudative or “wet” age-related macular degeneration (“wAMD”), rheumatoid arthritis and other inflammatory diseases and most cancers. The diseased tissue or tumor associated with these conditions typically expresses abnormally high levels of VEGF and exhibits high levels of angiogenesis or vascular permeability. For example, wAMD is an angiogenic disease characterized by choroidal neovascularization in one or both eyes in elderly individuals and is a leading cause of blindness. See Gehrs et al., Ann Med. 38(7):450–471 (2006). There are a number of therapeutic strategies for inhibiting aberrant angiogenesis, targeting VEGF or angiopoietin. See Biela and Siemann, Cancer Lett. 380(2):525–533 (2016). However, these treatments typically require repeated injections, which can increase the risk of inflammation, infection and other adverse effects in some patients. Repeated injections are also associated with patient compliance, and difficulty in adherence and non-compliance can lead to vision loss and worsening of eye diseases or conditions. The rate of non-compliance and non-adherence for treatment regimens that require repeated or frequent visits to a medical office is particularly high in most elderly patients affected by AMD. In the art, there is a need for improved therapeutic molecules, specifically for use in gene therapy for treating diseases such as wAMD. 3. Summary of the Invention
[0007] In one aspect, the present disclosure provides polypeptides comprising: (i) a first domain that binds to VEGF; (ii) a second domain that binds to VEGF; and (iii) a third domain that binds to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2). In some embodiments, the third domain is optionally at the N-terminus of the first and second domains.
[0008] In some embodiments, the first domain is derived from vascular endothelial growth factor receptor-1 (VEGFR-1 or FLT-1). In some embodiments, the first domain comprises domain 2 of VEGFR-1 or a variant thereof. In some embodiments, the first domain comprises the amino acid sequence shown in SEQ ID NO: 1 or consists of said amino acid sequence. In some embodiments, the first domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity to SEQ ID NO: 1 or consists of said amino acid sequence.
[0009] In some embodiments, the second domain is derived from vascular endothelial growth factor receptor-2 (VEGFR-2 or Flk-1). In some embodiments, the second domain comprises domain 3 of VEGFR-2 or a variant thereof. In some embodiments, the second domain comprises the amino acid sequence shown in SEQ ID NO: 2 or consists of said amino acid sequence. In some embodiments, the second domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity to SEQ ID NO: 2 or consists of said amino acid sequence.
[0010] In some embodiments, the third domain (ABD) comprises one or two repeats of the amino acid sequence shown in SEQ ID NO: 3. In some embodiments, the third domain comprises one or two repeats of the amino acid sequence shown in SEQ ID NO: 51. In some embodiments, the third domain comprises two repeats of the amino acid sequence shown in SEQ ID NO: 3. In some embodiments, the third domain comprises two repeats of the amino acid sequence shown in SEQ ID NO: 51. In some embodiments, the third domain comprises the amino acid sequence shown in SEQ ID NO: 3 and the amino acid sequence shown in SEQ ID NO: 51.
[0011] In some embodiments, the third domain comprises the amino acid sequence shown in SEQ ID NO: 4 or consists of said amino acid sequence, or the third domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity to SEQ ID NO: 4.
[0012] In some embodiments, the third domain comprises or consists of the amino acid sequence shown in SEQ ID NO: 52, or the third domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity to SEQ ID NO: 52.
[0013] In some embodiments, the third domain comprises or consists of the amino acid sequence shown in SEQ ID NO: 53, or the third domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity to SEQ ID NO: 53.
[0014] In some embodiments, the third domain comprises or consists of the amino acid sequence shown in SEQ ID NO: 54, or the third domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity to SEQ ID NO: 54.
[0015] In another aspect, the present disclosure provides a polypeptide comprising (i) a first domain that binds to VEGF, the first domain comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 1; (ii) a second domain that binds to VEGF, the second domain comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 2; and (iii) a third domain that binds to angiopoietin, the third domain comprising two amino acid sequences, each sequence comprising an amino acid sequence having at least 80%, 85%, 90% or 100% identity to SEQ ID NO: 3.
[0016] In another aspect, the present disclosure provides polypeptides that comprise (i) a first domain that binds to VEGF, the first domain comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 1; (ii) a second domain that binds to VEGF, the second domain comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 2; and (iii) a third domain that binds to angiopoietin, the third domain comprising two amino acid sequences, each sequence having at least 80%, 85%, 90% or 100% identity to SEQ ID NO: 51.
[0017] In another aspect, the present disclosure provides polypeptides that comprise (i) a first domain that binds to VEGF, the first domain comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 1; (ii) a second domain that binds to VEGF, the second domain comprising an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 2; and (iii) a third domain that binds to angiopoietin, the third domain comprising an amino acid sequence having at least 80%, 85%, 90% or 100% identity to SEQ ID NO: 3 and an amino acid sequence having at least 80%, 85%, 90% or 100% identity to SEQ ID NO: 51.
[0018] In some embodiments, the polypeptide further comprises the Fc region of an antibody. In some embodiments, the Fc region comprises the amino acid sequence shown in SEQ ID NO: 5.
[0019] In some embodiments, the polypeptide further comprises a signal peptide. In some embodiments, the signal peptide comprises the amino acid sequence shown in SEQ ID NO: 6.
[0020] In some embodiments, the polypeptide further comprises one or more linkers.
[0021] In some embodiments, the present disclosure provides polypeptides that comprise the amino acid sequences set forth in SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, or SEQ ID NO:10, or amino acid sequences having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity to SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, or SEQ ID NO:10.
[0022] In some embodiments, the present disclosure provides polypeptides that comprise the amino acid sequences set forth in SEQ ID NO:58, SEQ ID NO:59, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:62, SEQ ID NO:63, SEQ ID NO:64, SEQ ID NO:65, or SEQ ID NO:66, or amino acid sequences having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity to SEQ ID NO:58, SEQ ID NO:59, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:62, SEQ ID NO:63, SEQ ID NO:64, SEQ ID NO:65, or SEQ ID NO:66.
[0023] In some embodiments, the polypeptides provided herein further comprise a VEGFC-binding domain. In some embodiments, the VEGFC-binding domain is derived from VEGFR-2. In other embodiments, the VEGFC-binding domain is derived from VEGFR-3.
[0024] In some embodiments, the VEGFC-binding domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO:55.
[0025] In some embodiments, the VEGFC-binding domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO:56.
[0026] In some embodiments, the VEGFC binding domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 57.
[0027] In some embodiments, the polypeptide is genetically fused or chemically conjugated to a reagent.
[0028] In another aspect, the present disclosure provides an isolated nucleic acid comprising a nucleic acid sequence encoding a polypeptide provided herein.
[0029] In another aspect, the present disclosure provides a vector comprising the isolated nucleic acid provided herein. In some embodiments, the vector is a viral vector. In some embodiments, the viral vector is an adeno-associated virus (AAV) vector. In some embodiments, the AAV vector is derived from AAV1, AAV2, AAV2i8, AAV3, AAV3-B, AAV4, AAV5, AAV6, AAV7, AAV8, AAVrh8, AAVrh8R, AAV9, AAV10, AAVrh10, AAV11, AAV12, AAV13, AAV-DJ, AAVLK03, AAVrh74, AAV44-9 or a combination or variant thereof. In some embodiments, the vector is a recombinant AAV2 (rAAV2) vector, a recombinant AAV8 (rAAV8) vector or a variant thereof.
[0030] In another aspect, the present disclosure provides a recombinant AAV (rAAV) vector comprising a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a first domain derived from VEGFR-1; (ii) a second domain derived from VEGFR-2; and (iii) a third domain capable of binding to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2), wherein the rAAV vector comprises inverted terminal repeats (ITRs) from: AAV1, AAV2, AAV2i8, AAV3, AAV3-B, AAV4, AAV5, AAV6, AAV7, AAV8, AAVrh8, AAVrh8R, AAV9, AAV10, AAVrh10, AAV11, AAV12, AAV13, AAV-DJ, AAVLK03, AAVrh74, AAV44-9 or a combination or variant thereof. In some embodiments, the ITRs are from AAV2. In other embodiments, the ITRs are from AAV8. In some embodiments, the third domain is at the N-terminus of the first and second domains.
[0031] In another aspect, the present disclosure provides recombinant AAV (rAAV) vectors comprising a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a first domain derived from VEGFR-1; (ii) a second domain derived from VEGFR-2; (iii) a third domain capable of binding to angiopoietin; and (iv) a fourth domain capable of binding to VEGFC, wherein the rAAV vector comprises inverted terminal repeats (ITRs) from: AAV1, AAV2, AAV2i8, AAV3, AAV3-B, AAV4, AAV5, AAV6, AAV7, AAV8, AAVrh8, AAVrh8R, AAV9, AAV10, AAVrh10, AAV11, AAV12, AAV13, AAV-DJ, AAV LK03, AAVrh74 or AAV44-9. In some embodiments, the ITRs are from AAV2. In other embodiments, the ITRs are from AAV8. In some embodiments, the third domain is at the N-terminus of the first and second domains.
[0032] In some embodiments, the present disclosure provides vectors comprising a nucleic acid sequence set forth in SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22 or SEQ ID NO: 23, or a nucleic acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity to SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22 or SEQ ID NO: 23.
[0033] In some embodiments, provided herein are vectors comprising a nucleic acid sequence set forth in SEQ ID NO: 67, SEQ ID NO: 68, SEQ ID NO: 69, SEQ ID NO: 70, SEQ ID NO: 71, SEQ ID NO: 72, SEQ ID NO: 73, SEQ ID NO: 74, or SEQ ID NO: 75, or a nucleic acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity to SEQ ID NO: 67, SEQ ID NO: 68, SEQ ID NO: 69, SEQ ID NO: 70, SEQ ID NO: 71, SEQ ID NO: 72, SEQ ID NO: 73, SEQ ID NO: 74, or SEQ ID NO: 75.
[0034] In another aspect, provided herein are recombinant AAV (rAAV) particles comprising (a) a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a first domain derived from VEGFR-1; (ii) a second domain derived from VEGFR-2; and (iii) a third domain capable of binding to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2); and (b) a capsid protein selected from: AAV1, AAV2, AAV2i8, AAV3, AAV3-B, AAV4, AAV5, AAV6, AAV7, AAV8, AAVrh8, AAVrh8R, AAV9, AAV10, AAVrh10, AAV11, AAV12, AAV13, AAV-DJ, AAV LK03, AAVrh74, AAV44-9, or variants thereof. In some embodiments, the third domain is at the N-terminus of the first and second domains. In some embodiments, the capsid protein is an AAV2 capsid protein. In some embodiments, the capsid protein is an AAV8 capsid protein. In some embodiments, the capsid protein is a variant of the AAV2 capsid protein comprising the amino acid sequence set forth in SEQ ID NO: 48, wherein the variant comprises the amino acid substitutions Y444F, R487G, T491V, Y500F, R585S, R588T, and Y730F in the capsid protein VP1 of AAV2.
[0035] In another aspect, the present disclosure provides recombinant AAV (rAAV) particles comprising (a) a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a first domain derived from VEGFR-1; (ii) a second domain derived from VEGFR-2; and (iii) a third domain capable of binding to angiopoietin; and (iv) a fourth domain capable of binding to VEGFC; and (b) a capsid protein selected from: AAV1, AAV2, AAV2i8, AAV3, AAV3-B, AAV4, AAV5, AAV6, AAV7, AAV8, AAVrh8, AAVrh8R, AAV9, AAV10, AAVrh10, AAV11, AAV12, AAV13, AAV-DJ, AAV LK03, AAVrh74, AAV44-9 or variants thereof.
[0036] A pharmaceutical composition comprising the polypeptide, vector or rAAV vector or rAAV particles provided herein and a pharmaceutically acceptable excipient.
[0037] In another aspect, the present disclosure provides a method of treating a disease or disorder in a subject, comprising administering to the subject the polypeptide, vector or pharmaceutical composition provided herein. In some embodiments, the disease or disorder is an angiogenesis or neovascular disease or disorder. In some embodiments, the disease or disorder is an inflammatory disease, an ocular disease, an autoimmune disease or cancer. In some embodiments, the disease or disorder is an ocular disease or disorder. In some embodiments, the ocular disease or disorder is selected from uveitis, retinitis pigmentosa, neovascular glaucoma, diabetic retinopathy (DR) (including proliferative diabetic retinopathy), ischemic retinopathy, intraocular neovascularization, age-related macular degeneration (AMD), retinal neovascularization, diabetic macular edema (DME), diabetic retinal ischemia, diabetic retinal edema, retinal vein occlusion (including central retinal vein occlusion and branch retinal vein occlusion), macular edema and macular edema after retinal vein occlusion (RVO). In some embodiments, the disease or disorder is age-related macular degeneration (AMD). In some embodiments, the AMD is wet AMD (wAMD).
[0038] In some embodiments, the method comprises administration by intravitreal or subretinal injection into the eye of the subject. Sequence Listing <110> Hangzhou Geneinno Biotech Co., Ltd. <120> Fusion Molecules Targeting VEGF and Angiopoietin and Their Uses <130> 14652-025-228 <140> <141> <150> PCT / CN2021 / 084559 <151> 2021-03-31 <160> 75 <170> PatentIn version 3.5 <210> 1 <211> 103 <212> PRT <213> Artificial Sequence <220> <223> Exemplary First Domain Derived from VEGFR-1 (D2) <400> 1 Ser Asp Thr Gly Arg Pro Phe Val Glu Met Tyr Ser Glu Ile Pro Glu 1 5 10 15 Ile Ile His Met Thr Glu Gly Arg Glu Leu Val Ile Pro Cys Arg Val 20 25 30 Thr Ser Pro Asn Ile Thr Val Thr Leu Lys Lys Phe Pro Leu Asp Thr 35 40 45 Leu Ile Pro Asp Gly Lys Arg Ile Ile Trp Asp Ser Arg Lys Gly Phe 50 55 60 Ile Ile Ser Asn Ala Thr Tyr Lys Glu Ile Gly Leu Leu Thr Cys Glu 65 70 75 80 Ala Thr Val Asn Gly His Leu Tyr Lys Thr Asn Tyr Leu Thr His Arg 85 90 95 Gln Thr Asn Thr Ile Ile Asp 100 <210> 2 <211> 101 <212> PRT <213> Artificial sequence <220> <223> Exemplary second domain derived from VEGFR-2 (D3) <400> 2 Val Val Leu Ser Pro Ser His Gly Ile Glu Leu Ser Val Gly Glu Lys 1 5 10 15 Leu Val Leu Asn Cys Thr Ala Arg Thr Glu Leu Asn Val Gly Ile Asp 20 25 30 Phe Asn Trp Glu Tyr Pro Ser Ser Lys His Gln His Lys Lys Leu Val 35 40 45 Asn Arg Asp Leu Lys Thr Gln Ser Gly Ser Glu Met Lys Lys Phe Leu 50 55 60 Ser Thr Leu Thr Ile Asp Gly Val Thr Arg Ser Asp Gln Gly Leu Tyr 65 70 75 80 Thr Cys Ala Ala Ser Ser Gly Leu Met Thr Lys Lys Asn Ser Thr Phe 85 90 95 Val Arg Val His Glu 100 <210> 3 <211> 14 <212> PRT <213> Artificial sequence <220> <223> Sequence of 14 amino acids in the exemplary third domain (ABD), ABD (Con4) <400> 3 Gln Glu Glu Cys Glu Trp Asp Pro Trp Thr Cys Glu His Met 1 5 10 <210> 4 <211> 59 <212> PRT <213> Synthetic sequence <220> <223> Exemplary third domain (ABD), ABD(1)(2xCon4) <400> 4 Gly Gly Gly Gly Gly Ala Gln Gln Glu Glu Cys Glu Trp Asp Pro Trp 1 5 10 15 Thr Cys Glu His Met Gly Ser Gly Ser Ala Thr Gly Gly Ser Gly Ser 20 25 30 Thr Ala Ser Ser Gly Ser Gly Ser Ala Thr His Gln Glu Glu Cys Glu 35 40 45 Trp Asp Pro Trp Thr Cys Glu His Met Leu Glu 50 55 <210> 5 <211> 227 <212> PRT <213> Synthetic sequence <220> <223> Exemplary IgG-Fc(1) <400> 5 Asp Lys Thr His Thr Cys Pro Pro Cys Pro Ala Pro Glu Leu Leu Gly 1 5 10 15 Gly Pro Ser Val Phe Leu Phe Pro Pro Lys Pro Lys Asp Thr Leu Met 20 25 30 Ile Ser Arg Thr Pro Glu Val Thr Cys Val Val Val Asp Val Ser His 35 40 45 Glu Asp Pro Glu Val Lys Phe Asn Trp Tyr Val Asp Gly Val Glu Val 50 55 60 His Asn Ala Lys Thr Lys Pro Arg Glu Glu Gln Tyr Asn Ser Thr Tyr 65 70 75 80 Arg Val Val Ser Val Leu Thr Val Leu His Gln Asp Trp Leu Asn Gly 85 90 95 Lys Glu Tyr Lys Cys Lys Val Ser Asn Lys Ala Leu Pro Ala Pro Ile 100 105 110 Glu Lys Thr Ile Ser Lys Ala Lys Gly Gln Pro Arg Glu Pro Gln Val 115 120 125 Tyr Thr Leu Pro Pro Ser Arg Asp Glu Leu Thr Lys Asn Gln Val Ser 130 135 140 Leu Thr Cys Leu Val Lys Gly Phe Tyr Pro Ser Asp Ile Ala Val Glu 145 150 155 160 Trp Glu Ser Asn Gly Gln Pro Glu Asn Asn Tyr Lys Thr Thr Pro Pro 165 170 175 Val Leu Asp Ser Asp Gly Ser Phe Phe Leu Tyr Ser Lys Leu Thr Val 180 185 190 Asp Lys Ser Arg Trp Gln Gln Gly Asn Val Phe Ser Cys Ser Val Met 195 200 205 His Glu Ala Leu His Asn His Tyr Thr Gln Lys Ser Leu Ser Leu Ser 210 215 220 Pro Gly Lys 225 <210> 6 <211> 26 <212> PRT <213> Artificial sequence <220> <223> Flt-1 signal <400> 6 Met Val Ser Tyr Trp Asp Thr Gly Val Leu Leu Cys Ala Leu Leu Ser 1 5 10 15 Cys Leu Leu Leu Thr Gly Ser Ser Ser Gly 20 25 <210> 7 <211> 517 <212> PRT <213> Artificial sequence <220> <223> Exemplary polypeptide 1 <400> 7 Met Val Ser Tyr Trp Asp Thr Gly Val Leu Leu Cys Ala Leu Leu Ser 1 5 10 15 Cys Leu Leu Leu Thr Gly Ser Ser Ser Gly Ser Asp Thr Gly Arg Pro 20 25 30 Phe Val Glu Met Tyr Ser Glu Ile Pro Glu Ile Ile His Met Thr Glu 35 40 45 Gly Arg Glu Leu Val Ile Pro Cys Arg Val Thr Ser Pro Asn Ile Thr 50 55 60 Val Thr Leu Lys Lys Phe Pro Leu Asp Thr Leu Ile Pro Asp Gly Lys 65 70 75 80 Arg Ile Ile Trp Asp Ser Arg Lys Gly Phe Ile Ile Ser Asn Ala Thr 85 90 95 Tyr Lys Glu Ile Gly Leu Leu Thr Cys Glu Ala Thr Val Asn Gly His 100 105 110 Leu Tyr Lys Thr Asn Tyr Leu Thr His Arg Gln Thr Asn Thr Ile Ile 115 120 125 Asp Val Val Leu Ser Pro Ser His Gly Ile Glu Leu Ser Val Gly Glu 130 135 140 Lys Leu Val Leu Asn Cys Thr Ala Arg Thr Glu Leu Asn Val Gly Ile 145 150 155 160 Asp Phe Asn Trp Glu Tyr Pro Ser Ser Lys His Gln His Lys Lys Leu 165 170 175 Val Asn Arg Asp Leu Lys Thr Gln Ser Gly Ser Glu Met Lys Lys Phe 180 185 190 Leu Ser Thr Leu Thr Ile Asp Gly Val Thr Arg Ser Asp Gln Gly Leu 195 200 205 Tyr Thr Cys Ala Ala Ser Ser Gly Leu Met Thr Lys Lys Asn Ser Thr 210 215 220 Phe Val Arg Val His Glu Lys Asp Lys Thr His Thr Cys Pro Pro Cys 225 230 235 240 Pro Ala Pro Glu Leu Leu Gly Gly Pro Ser Val Phe Leu Phe Pro Pro 245 250 255 Lys Pro Lys Asp Thr Leu Met Ile Ser Arg Thr Pro Glu Val Thr Cys 260 265 270 Val Val Val Asp Val Ser His Glu Asp Pro Glu Val Lys Phe Asn Trp 275 280 285 Tyr Val Asp Gly Val Glu Val His Asn Ala Lys Thr Lys Pro Arg Glu 290 295 300 Glu Gln Tyr Asn Ser Thr Tyr Arg Val Val Ser Val Leu Thr Val Leu 305 310 315 320 His Gln Asp Trp Leu Asn Gly Lys Glu Tyr Lys Cys Lys Val Ser Asn 325 330 335 Lys Ala Leu Pro Ala Pro Ile Glu Lys Thr Ile Ser Lys Ala Lys Gly 340 345 350 Gln Pro Arg Glu Pro Gln Val Tyr Thr Leu Pro Pro Ser Arg Asp Glu 355 360 365 Leu Thr Lys Asn Gln Val Ser Leu Thr Cys Leu Val Lys Gly Phe Tyr 370 375 380 Pro Ser Asp Ile Ala Val Glu Trp Glu Ser Asn Gly Gln Pro Glu Asn 385 390 395 400 Asn Tyr Lys Thr Thr Pro Pro Val Leu Asp Ser Asp Gly Ser Phe Phe 405 410 415 Leu Tyr Ser Lys Leu Thr Val Asp Lys Ser Arg Trp Gln Gln Gly Asn 420 425 430 Val Phe Ser Cys Ser Val Met His Glu Ala Leu His Asn His Tyr Thr 435 440 445 Gln Lys Ser Leu Ser Leu Ser Pro Gly Lys Gly Gly Gly Gly Gly Ala 450 455 460 Gln Gln Glu Glu Cys Glu Trp Asp Pro Trp Thr Cys Glu His Met Gly 465 470 475 480 Ser Gly Ser Ala Thr Gly Gly Ser Gly Ser Thr Ala Ser Ser Gly Ser 485 490 495 Gly Ser Ala Thr His Gln Glu Glu Cys Glu Trp Asp Pro Trp Thr Cys 500 505 510 Glu His Met Leu Glu 515 <210> 8 <211> 523 <212> PRT <213> Artificial sequence <220> <223> Exemplary polypeptide 2 <400> 8 Met Val Ser Tyr Trp Asp Thr Gly Val Leu Leu Cys Ala Leu Leu Ser 1 5 10 15 Cys Leu Leu Leu Thr Gly Ser Ser Ser Gly Ser Asp Thr Gly Arg Pro 20 25 30 Phe Val Glu Met Tyr Ser Glu Ile Pro Glu Ile Ile His Met Thr Glu 35 40 45 Gly Arg Glu Leu Val Ile Pro Cys Arg Val Thr Ser Pro Asn Ile Thr 50 55 60 Val Thr Leu Lys Lys Phe Pro Leu Asp Thr Leu Ile Pro Asp Gly Lys 65 70 75 80 Arg Ile Ile Trp Asp Ser Arg Lys Gly Phe Ile Ile Ser Asn Ala Thr 85 90 95 Tyr Lys Glu Ile Gly Leu Leu Thr Cys Glu Ala Thr Val Asn Gly His 100 105 110 Leu Tyr Lys Thr Asn Tyr Leu Thr His Arg Gln Thr Asn Thr Ile Ile 115 120 125 Asp Val Val Leu Ser Pro Ser His Gly Ile Glu Leu Ser Val Gly Glu 130 135 140 Lys Leu Val Leu Asn Cys Thr Ala Arg Thr Glu Leu Asn Val Gly Ile 145 150 155 160 Asp Phe Asn Trp Glu Tyr Pro Ser Ser Lys His Gln His Lys Lys Leu 165 170 175 Val Asn Arg Asp Leu Lys Thr Gln Ser Gly Ser Glu Met Lys Lys Phe 180 185 190 Leu Ser Thr Leu Thr Ile Asp Gly Val Thr Arg Ser Asp Gln Gly Leu 195 200 205 Tyr Thr Cys Ala Ala Ser Ser Gly Leu Met Thr Lys Lys Asn Ser Thr 210 215 220 Phe Val Arg Val His Glu Lys Gly Gly Gly Gly Gly Ser Gly Gly Gly 225 230 235 240 Gly Gly Ala Gln Gln Glu Glu Cys Glu Trp Asp Pro Trp Thr Cys Glu 245 250 255 His Met Gly Ser Gly Ser Ala Thr Gly Gly Ser Gly Ser Thr Ala Ser 260 265 270 Ser Gly Ser Gly Ser Ala Thr His Gln Glu Glu Cys Glu Trp Asp Pro 275 280 285 Trp Thr Cys Glu His Met Leu Glu Asp Lys Thr His Thr Cys Pro Pro 290 295 300 Cys Pro Ala Pro Glu Leu Leu Gly Gly Pro Ser Val Phe Leu Phe Pro 305 310 315 320 Pro Lys Pro Lys Asp Thr Leu Met Ile Ser Arg Thr Pro Glu Val Thr 325 330 335 Cys Val Val Val Asp Val Ser His Glu Asp Pro Glu Val Lys Phe Asn 340 345 350 Trp Tyr Val Asp Gly Val Glu Val His Asn Ala Lys Thr Lys Pro Arg 355 360 365 Glu Glu Gln Tyr Asn Ser Thr Tyr Arg Val Val Ser Val Leu Thr Val 370 375 380 Leu His Gln Asp Trp Leu Asn Gly Lys Glu Tyr Lys Cys Lys Val Ser 385 390 395 400 Asn Lys Ala Leu Pro Ala Pro Ile Glu Lys Thr Ile Ser Lys Ala Lys 405 410 415 Gly Gln Pro Arg Glu Pro Gln Val Tyr Thr Leu Pro Pro Ser Arg Asp 420 425 430 Glu Leu Thr Lys Asn Gln Val Ser Leu Thr Cys Leu Val Lys Gly Phe 435 440 445 Tyr Pro Ser Asp Ile Ala Val Glu Trp Glu Ser Asn Gly Gln Pro Glu 450 455 460 Asn Asn Tyr Lys Thr Thr Pro Pro Val Leu Asp Ser Asp Gly Ser Phe 465 470 475 480 Phe Leu Tyr Ser Lys Leu Thr Val Asp Lys Ser Arg Trp Gln Gln Gly 485 490 495 Asn Val Phe Ser Cys Ser Val Met His Glu Ala Leu His Asn His Tyr 500 505 510 Thr Gln Lys Ser Leu Ser Leu Ser Pro Gly Lys 515 520 <210> 9 <211> 529 <212> PRT <213> Artificial Sequence <220> <223> Exemplary Polypeptide 3 <400> 9 Met Val Ser Tyr Trp Asp Thr Gly Val Leu Leu Cys Ala Leu Leu Ser 1 5 10 15 Cys Leu Leu Leu Thr Gly Ser Ser Ser Gly Gly Gly Gly Gly Gly Ser 20 25 30 Gly Gly Gly Gly Gly Ala Gln Gln Glu Glu Cys Glu Trp Asp Pro Trp 35 40 45 Thr Cys Glu His Met Gly Ser Gly Ser Ala Thr Gly Gly Ser Gly Ser 50 55 60 Thr Ala Ser Ser Gly Ser Gly Ser Ala Thr His Gln Glu Glu Cys Glu 65 70 75 80 Trp Asp Pro Trp Thr Cys Glu His Met Leu Glu Gly Gly Gly Gly Gly 85 90 95 Ser Ser Asp Thr Gly Arg Pro Phe Val Glu Met Tyr Ser Glu Ile Pro 100 105 110 Glu Ile Ile His Met Thr Glu Gly Arg Glu Leu Val Ile Pro Cys Arg 115 120 125 Val Thr Ser Pro Asn Ile Thr Val Thr Leu Lys Lys Phe Pro Leu Asp 130 135 140 Thr Leu Ile Pro Asp Gly Lys Arg Ile Ile Trp Asp Ser Arg Lys Gly 145 150 155 160 Phe Ile Ile Ser Asn Ala Thr Tyr Lys Glu Ile Gly Leu Leu Thr Cys 165 170 175 Glu Ala Thr Val Asn Gly His Leu Tyr Lys Thr Asn Tyr Leu Thr His 180 185 190 Arg Gln Thr Asn Thr Ile Ile Asp Val Val Leu Ser Pro Ser His Gly 195 200 205 Ile Glu Leu Ser Val Gly Glu Lys Leu Val Leu Asn Cys Thr Ala Arg 210 215 220 Thr Glu Leu Asn Val Gly Ile Asp Phe Asn Trp Glu Tyr Pro Ser Ser 225 230 235 240 Lys His Gln His Lys Lys Leu Val Asn Arg Asp Leu Lys Thr Gln Ser 245 250 255 Gly Ser Glu Met Lys Lys Phe Leu Ser Thr Leu Thr Ile Asp Gly Val 260 265 270 Thr Arg Ser Asp Gln Gly Leu Tyr Thr Cys Ala Ala Ser Ser Gly Leu 275 280 285 Met Thr Lys Lys Asn Ser Thr Phe Val Arg Val His Glu Lys Asp Lys 290 295 300 Thr His Thr Cys Pro Pro Cys Pro Ala Pro Glu Leu Leu Gly Gly Pro 305 310 315 320 Ser Val Phe Leu Phe Pro Pro Lys Pro Lys Asp Thr Leu Met Ile Ser 325 330 335 Arg Thr Pro Glu Val Thr Cys Val Val Val Asp Val Ser His Glu Asp 340 345 350 Pro Glu Val Lys Phe Asn Trp Tyr Val Asp Gly Val Glu Val His Asn 355 360 365 Ala Lys Thr Lys Pro Arg Glu Glu Gln Tyr Asn Ser Thr Tyr Arg Val 370 375 380 Val Ser Val Leu Thr Val Leu His Gln Asp Trp Leu Asn Gly Lys Glu 385 390 395 400 Tyr Lys Cys Lys Val Ser Asn Lys Ala Leu Pro Ala Pro Ile Glu Lys 405 410 415 Thr Ile Ser Lys Ala Lys Gly Gln Pro Arg Glu Pro Gln Val Tyr Thr 420 425 430 Leu Pro Pro Ser Arg Asp Glu Leu Thr Lys Asn Gln Val Ser Leu Thr 435 440 445 Cys Leu Val Lys Gly Phe Tyr Pro Ser Asp Ile Ala Val Glu Trp Glu 450 455 460 Ser Asn Gly Gln Pro Glu Asn Asn Tyr Lys Thr Thr Pro Pro Val Leu 465 470 475 480 Asp Ser Asp Gly Ser Phe Phe Leu Tyr Ser Lys Leu Thr Val Asp Lys 485 490 495 Ser Arg Trp Gln Gln Gly Asn Val Phe Ser Cys Ser Val Met His Glu 500 505 510 Ala Leu His Asn His Tyr Thr Gln Lys Ser Leu Ser Leu Ser Pro Gly 515 520 525 Lys <210> 10 <211> 316 <212> PRT <213> Artificial Sequence <220> <223> Exemplary Polypeptide 4 <400> 10 Met Val Ser Tyr Trp Asp Thr Gly Val Leu Leu Cys Ala Leu Leu Ser 1 5 10 15 Cys Leu Leu Leu Thr Gly Ser Ser Ser Gly Ser Asp Thr Gly Arg Pro 20 25 30 Phe Val Glu Met Tyr Ser Glu Ile Pro Glu Ile Ile His Met Thr Glu 35 40 45 Gly Arg Glu Leu Val Ile Pro Cys Arg Val Thr Ser Pro Asn Ile Thr 50 55 60 Val Thr Leu Lys Lys Phe Pro Leu Asp Thr Leu Ile Pro Asp Gly Lys 65 70 75 80 Arg Ile Ile Trp Asp Ser Arg Lys Gly Phe Ile Ile Ser Asn Ala Thr 85 90 95 Tyr Lys Glu Ile Gly Leu Leu Thr Cys Glu Ala Thr Val Asn Gly His 100 105 110 Leu Tyr Lys Thr Asn Tyr Leu Thr His Arg Gln Thr Asn Thr Ile Ile 115 120 125 Asp Val Val Leu Ser Pro Ser His Gly Ile Glu Leu Ser Val Gly Glu 130 135 140 Lys Leu Val Leu Asn Cys Thr Ala Arg Thr Glu Leu Asn Val Gly Ile 145 150 155 160 Asp Phe Asn Trp Glu Tyr Pro Ser Ser Lys His Gln His Lys Lys Leu 165 170 175 Val Asn Arg Asp Leu Lys Thr Gln Ser Gly Ser Glu Met Lys Lys Phe 180 185 190 Leu Ser Thr Leu Thr Ile Asp Gly Val Thr Arg Ser Asp Gln Gly Leu 195 200 205 Tyr Thr Cys Ala Ala Ser Ser Gly Leu Met Thr Lys Lys Asn Ser Thr 210 215 220 Phe Val Arg Val His Glu Lys Asp Lys Thr His Thr Cys Pro Pro Cys 225 230 235 240 Pro Ala Pro Glu Leu Leu Gly Gly Pro Ser Val Gly Gly Gly Gly Gly 245 250 255 Ser Gly Gly Gly Gly Gly Ala Gln Gln Glu Glu Cys Glu Trp Asp Pro 260 265 270 Trp Thr Cys Glu His Met Gly Ser Gly Ser Ala Thr Gly Gly Ser Gly 275 280 285 Ser Thr Ala Ser Ser Gly Ser Gly Ser Ala Thr His Gln Glu Glu Cys 290 295 300 Glu Trp Asp Pro Trp Thr Cys Glu His Met Leu Glu 305 310 315 <210> 11 <211> 3672 <212> DNA <213> Artificial Sequence <220> <223> EXG102-02 <400> 11 ttggccactc cctctctgcg cgctcgctcg ctcactgagg ccgggcgacc aaaggtcgcc 60 cgacgcccgg gctttgcccg ggcggcctca gtgagcgagc gagcgcgcag agagggagtg 120 gccaactcca tcactagggg ttcctcagat ctgaattcgg tacctagtta ttaatagtaa 180 tcaattacgg ggtcattagt tcatagccca tatatggagt tccgcgttac ataacttacg 240 gtaaatggcc cgcctggctg accgcccaac gacccccgcc cattgacgtc aataatgacg 300 tatgttccca tagtaacgcc aatagggact ttccattgac gtcaatgggt ggagtattta 360 cggtaaactg cccacttggc agtacatcaa gtgtatcata tgccaagtac gccccctatt 420 gacgtcaatg acggtaaatg gcccgcctgg cattatgccc agtacatgac cttatgggac 480 tttcctactt ggcagtacat ctacgtatta gtcatcgcta ttaccatggt cgaggtgagc 540 cccacgttct gcttcactct ccccatctcc cccccctccc cacccccaat tttgtattta 600 tttatttttt aattattttg tgcagcgatg ggggcggggg gggggggggg gcgcgcgcca 660 ggcggggcgg ggcggggcga ggggcggggc ggggcgaggc ggagaggtgc ggcggcagcc 720 aatcagagcg gcgcgctccg aaagtttcct tttatggcga ggcggcggcg gcggcggccc 780 tataaaaagc gaagcgcgcg gcgggcggga gtcgctgcgc gctgccttcg ccccgtgccc 840 cgctccgccg ccgcctcgcg ccgcccgccc cggctctgac tgaccgcgtt actcccacag 900 gtgagcgggc gggacggccc ttctcctccg ggctgtaatt agcgcttggt ttaatgacgg 960 cttgtttctt ttctgtggct gcgtgaaagc cttgaggggc tccgggaggg ccctttgtgc 1020 ggggggagcg gctcgggggg tgcgtgcgtg tgtgtgtgcg tggggagcgc cgcgtgcggc 1080 tccgcgctgc ccggcggctg tgagcgctgc gggcgcggcg cggggctttg tgcgctccgc 1140 agtgtgcgcg aggggagcgc ggccgggggc ggtgccccgc ggtgcggggg gggctgcgag 1200 gggaacaaag gctgcgtgcg gggtgtgtgc gtgggggggt gagcaggggg tgtgggcgcg 1260 tcggtcgggc tgcaaccccc cctgcacccc cctccccgag ttgctgagca cggcccggct 1320 tcggtcgggc tgcaaccccc cctgcacccc cctccccgag ttgctgagca cggcccggct 1320 tcgggtgcgg ggctccgtac ggggcgtggc gcggggctcg ccgtgccggg cggggggtgg 1380 tcgggtgcgg ggctccgtac ggggcgtggc gcggggctcg ccgtgccggg cggggggtgg 1380 cggcaggtgg gggtgccggg cggggcgggg ccgcctcggg ccggggaggg ctcgggggag 1440 cggcaggtgg gggtgccggg cggggcgggg ccgcctcggg ccggggaggg ctcgggggag 1440 gggcgcggcg gcccccggag cgccggcggc tgtcgaggcg cggcgagccg cagccattgc 1500 gggcgcggcg gcccccggag cgccggcggc tgtcgaggcg cggcgagccg cagccattgc 1500 cttttatggt aatcgtgcga gagggcgcag ggacttcctt tgtcccaaat ctgtgcggag 1560 cttttatggt aatcgtgcga gagggcgcag ggacttcctt tgtcccaaat ctgtgcggag 1560 ccgaaatctg ggaggcgccg ccgcaccccc tctagcgggc gcggggcgaa gcggtgcggc 1620 ccgaaatctg ggaggcgccg ccgcaccccc tctagcgggc gcggggcgaa gcggtgcggc 1620 gccggcagga aggaaatggg cggggagggc cttcgtgcgt cgccgcgccg ccgtcccctt 1680 gccggcagga aggaaatggg cggggagggc cttcgtgcgt cgccgcgccg ccgtcccctt 1680 ctccctctcc agcctcgggg ctgtccgcgg ggggacggct gccttcgggg gggacggggc 1740 ctccctctcc agcctcgggg ctgtccgcgg ggggacggct gccttcgggg gggacggggc 1740 agggcggggt tcggcttctg gcgtgtgacc ggcggctcta gagcctctgc taaccatgtt 1800 agggcggggt tcggcttctg gcgtgtgacc ggcggctcta gagcctctgc taaccatgtt 1800 catgccttct tctttttcct acagctcctg ggcaacgtgc tggttattgt gctgtctcat 1860 catgccttct tctttttcct acagctcctg ggcaacgtgc tggttattgt gctgtctcat 1860 cattttggca aagaattcct cgaagatcta ggcaacgcgt ctcgaggcgg ccgccaccat 1920 cattttggca aagaattcct cgaagatcta ggcaacgcgt ctcgaggcgg ccgccaccat 1920 ggtgtcttat tgggatactg gcgtgctgct ctgtgccctc ctgagttgcc tgctcctgac 1980 ggtgtcttat tgggatactg gcgtgctgct ctgtgccctc ctgagttgcc tgctcctgac 1980 tggttcttct tctgggtccg atactgggcg ccccttcgtg gagatgtact ccgagatccc 2040 tgaaatcatt cacatgactg agggtcggga actggtcatc ccatgccgcg tgacctctcc 2100 caacattact gtgaccctga agaaattccc tctggacacc ctcatcccag atgggaagag 2160 gatcatttgg gactcaagaa agggttttat catcagcaac gctacataca aggagattgg 2220 cctgctcacc tgcgaagcaa cagtgaacgg acacctgtac aagactaatt atctcaccca 2280 tagacagaca aacactatca ttgatgtggt cctgtcacca agccacggca tcgagctcag 2340 cgtcggtgaa aagctggtgc tcaattgtac agcccggact gagctgaacg tgggcattga 2400 cttcaattgg gaatacccca gctccaagca ccagcataag aaactggtga accgcgatct 2460 caaaacccag tccggatctg agatgaagaa atttctgagc accctcacaa tcgacggcgt 2520 gacacgatcc gatcagggac tgtatacttg cgccgcttct agtggcctga tgaccaagaa 2580 aaatagcaca ttcgtcaggg tgcacgaaaa ggacaaaact catacctgcc caccttgtcc 2640 agcaccagag ctgctcggag gaccatccgt gttcctgttt ccacccaagc ccaaagatac 2700 tctgatgatt tcacgcacac ccgaagtcac ttgcgtggtc gtggacgtgt cccacgagga 2760 ccccgaagtc aagtttaact ggtacgtgga cggcgtcgag gtgcataatg ctaagacaaa 2820 accccgagag gaacagtaca actctaccta tagggtcgtg agtgtcctga cagtgctcca 2880 ccaggattgg ctgaacggaa aggagtataa gtgcaaagtg tctaataagg cactgcctgc 2940 cccaatcgag aaaacaatta gtaaggccaa agggcagccc agagaacctc aggtgtacac 3000 tctgcctcca tctcgggacg agctcactaa gaaccaggtc agtctgacct gtctcgtgaa 3060 agggttctat cctagtgata tcgctgtgga gtgggaatca aatggtcagc cagagaacaa 3120 ttacaagacc acaccccctg tcctggacag cgatggctcc ttctttctgt attccaagct 3180 caccgtggac aaatctcgat ggcagcaggg aaacgtcttt agttgttcag tgatgcacga 3240 agccctccat aaccactaca ctcagaaaag cctcagcctc agccctggga aatgataagc 3300 ggccgcgcgg atccagacat gataagatac attgatgagt ttggacaaac cacaactaga 3360 atgcagtgaa aaaaatgctt tatttgtgaa atttgtgatg ctattgcttt atttgtaacc 3420 attataagct gcaataaaca agttaacaac aacaattgca ttcattttat gtttcaggtt 3480 cagggggagg tgtgggaggt tttttagtcg actggggaga gatctgagga acccctagtg 3540 atggagttgg ccactccctc tctgcgcgct cgctcgctca ctgaggccgc ccgggcaaag 3600 cccgggcgtc gggcgacctt tggtcgcccg gcctcagtga gcgagcgagc gcgcagagag 3660 ggagtggcca ac 3672 <210> 12 <211> 4348 <212> DNA <213> Artificial Sequence <220> <223> EXG102-03-01 <400> 12 ttggccactc cctctctgcg cgctcgctcg ctcactgagg ccgggcgacc aaaggtcgcc 60 cgacgcccgg gctttgcccg ggcggcctca gtgagcgagc gagcgcgcag agagggagtg 120 gccaactcca tcactagggg ttcctcagat ctgaattcgg tacctagtta ttaatagtaa 180 tcaattacgg ggtcattagt tcatagccca tatatggagt tccgcgttac ataacttacg 240 gtaaatggcc cgcctggctg accgcccaac gacccccgcc cattgacgtc aataatgacg 300 tatgttccca tagtaacgcc aatagggact ttccattgac gtcaatgggt ggagtattta 360 cggtaaactg cccacttggc agtacatcaa gtgtatcata tgccaagtac gccccctatt 420 gacgtcaatg acggtaaatg gcccgcctgg cattatgccc agtacatgac cttatgggac 480 tttcctactt ggcagtacat ctacgtatta gtcatcgcta ttaccatggt cgaggtgagc 540 cccacgttct gcttcactct ccccatctcc cccccctccc cacccccaat tttgtattta 600 tttatttttt aattattttg tgcagcgatg ggggcggggg gggggggggg gcgcgcgcca 660 ggcggggcgg ggcggggcga ggggcggggc ggggcgaggc ggagaggtgc ggcggcagcc 720 aatcagagcg gcgcgctccg aaagtttcct tttatggcga ggcggcggcg gcggcggccc 780 tataaaaagc gaagcgcgcg gcgggcggga gtcgctgcgc gctgccttcg ccccgtgccc 840 cgctccgccg ccgcctcgcg ccgcccgccc cggctctgac tgaccgcgtt actcccacag 900 gtgagcgggc gggacggccc ttctcctccg ggctgtaatt agcgcttggt ttaatgacgg 960 cttgtttctt ttctgtggct gcgtgaaagc cttgaggggc tccgggaggg ccctttgtgc 1020 ggggggagcg gctcgggggg tgcgtgcgtg tgtgtgtgcg tggggagcgc cgcgtgcggc 1080 tccgcgctgc ccggcggctg tgagcgctgc gggcgcggcg cggggctttg tgcgctccgc 1140 agtgtgcgcg aggggagcgc ggccgggggc ggtgccccgc ggtgcggggg gggctgcgag 1200 gggaacaaag gctgcgtgcg gggtgtgtgc gtgggggggt gagcaggggg tgtgggcgcg 1260 tcggtcgggc tgcaaccccc cctgcacccc cctccccgag ttgctgagca cggcccggct 1320 tcgggtgcgg ggctccgtac ggggcgtggc gcggggctcg ccgtgccggg cggggggtgg 1380 cggcaggtgg gggtgccggg cggggcgggg ccgcctcggg ccggggaggg ctcgggggag 1440 gggcgcggcg gcccccggag cgccggcggc tgtcgaggcg cggcgagccg cagccattgc 1500 cttttatggt aatcgtgcga gagggcgcag ggacttcctt tgtcccaaat ctgtgcggag 1560 ccgaaatctg ggaggcgccg ccgcaccccc tctagcgggc gcggggcgaa gcggtgcggc 1620 gccggcagga aggaaatggg cggggagggc cttcgtgcgt cgccgcgccg ccgtcccctt 1680 ctccctctcc agcctcgggg ctgtccgcgg ggggacggct gccttcgggg gggacggggc 1740 agggcggggt tcggcttctg gcgtgtgacc ggcggctcta gagcctctgc taaccatgtt 1800 catgccttct tctttttcct acagctcctg ggcaacgtgc tggttattgt gctgtctcat 1860 cattttggca aagaattcct cgaagatcta ggcaacgcgt ctcgagagaa ttcgccacca 1920 tggtgagcta ttgggataca ggcgtgctgc tgtgcgccct gctgagctgc ctgctgctga 1980 ccggcagcag cagcggcagc gacaccggca gacctttcgt ggagatgtac agcgagatcc 2040 ctgagatcat ccacatgacc gagggcagag agctggtgat cccttgcaga gtgaccagcc 2100 ctaatatcac cgtgaccctc aagaagttcc ctctggatac cctgatccct gacggcaaga 2160 gaatcatctg ggacagcaga aagggcttca tcatcagcaa tgccacctac aaggagatcg 2220 gcctgctgac ctgcgaggcc accgtgaatg gccacctgta caagaccaat tacctgaccc 2280 acagacagac caataccatc atcgacgtgg tgctgagccc tagccacggc atcgagctga 2340 gcgtgggcga gaagctggtg ctgaattgca ccgccagaac cgagctgaat gtgggcatcg 2400 acttcaattg ggagtaccct agcagcaagc accagcacaa gaagctggtg aatagagacc 2460 tgaagaccca gagcggcagc gagatgaaga aattcctgag caccctgacc atcgacggcg 2520 tgaccagaag cgaccagggc ctgtacacct gcgctgccag cagcggcctg atgaccaaga 2580 agaatagcac cttcgtgaga gtgcacgaga aggacaagac ccacacctgc cctccttgcc 2640 ctgcccctga gctgctgggc ggccctagcg tgttcctgtt ccctcctaag cctaaggaca 2700 ccctcatgat cagcagaacc cctgaggtga cctgcgtggt ggtggacgtg agccacgagg 2760 accctgaggt gaagttcaat tggtacgtgg acggcgtgga ggtgcacaat gccaagacca 2820 agcctagaga ggagcagtac aatagcacct acagagtggt gagcgtgctg accgtgctgc 2880 accaggactg gctgaatggc aaggagtaca agtgcaaggt gagcaataag gccctgcctg 2940 cccctatcga gaagaccatc agcaaggcca agggccagcc tagagagcct caggtgtaca 3000 ccctgcctcc tagcagagac gagctgacca agaatcaggt gagcctgacc tgcctggtga 3060 agggcttcta ccctagcgac atcgccgtgg agtgggagag caatggccag cctgagaata 3120 attacaagac cacccctcct gtgctggaca gcgacggcag cttcttcctg tacagcaagc 3180 tgaccgtgga caagagcaga tggcagcagg gcaatgtgtt cagctgcagc gtgatgcacg 3240 agaatagcac cttcgtgaga gtgcacgaga aggacaagac ccacacctgc cctccttgcc 2640 ctgcccctga gctgctgggc ggccctagcg tgttcctgtt ccctcctaag cctaaggaca 2700 ccctcatgat cagcagaacc cctgaggtga cctgcgtggt ggtggacgtg agccacgagg 2760 accctgaggt gaagttcaat tggtacgtgg acggcgtgga ggtgcacaat gccaagacca 2820 agcctagaga ggagcagtac aatagcacct acagagtggt gagcgtgctg accgtgctgc 2880 accaggactg gctgaatggc aaggagtaca agtgcaaggt gagcaataag gccctgcctg 2940 cccctatcga gaagaccatc agcaaggcca agggccagcc tagagagcct caggtgtaca 3000 ccctgcctcc tagcagagac gagctgacca agaatcaggt gagcctgacc tgcctggtga 3060 agggcttcta ccctagcgac atcgccgtgg agtgggagag caatggccag cctgagaata 3120 attacaagac cacccctcct gtgctggaca gcgacggcag cttcttcctg tacagcaagc 3180 tgaccgtgga caagagcaga tggcagcagg gcaatgtgtt cagctgcagc gtgatgcacg 3240 aggccctgca caatcactac acacagaaga gcctgagcct gagccctggc aagtgataag 3300 gatatcaaga tcttgcggcc gctcgataat caacctctgg attacaaaat ttgtgaaaga 3360 ttgactggta ttcttaacta tgttgctcct tttacgctat gtggatacgc tgctttaatg 3420 cctttgtatc atgctattgc ttcccgtatg gctttcattt tctcctcctt gtataaatcc 3480 tggttgctgt ctctttatga ggagttgtgg cccgttgtca ggcaacgtgg cgtggtgtgc 3540 actgtgtttg ctgacgcaac ccccactggt tggggcattg ccaccacctg tcagctcctt 3600 tccgggactt tcgctttccc cctccctatt gccacggcgg aactcatcgc cgcctgcctt 3660 gcccgctgct ggacaggggc tcggctgttg ggcactgaca attccgtggt gttgtcgggg 3720 aaatcatcgt cctttccttg gctgctcgcc tgtgttgcca cctggattct gcgcgggacg 3780 tccttctgct acgtcccttc ggccctcaat ccagcggacc ttccttcccg cggcctgctg 3840 ccggctctgc ggcctcttcc gcgtcttcgc cttcgccctc agacgagtcg gatctccctt 3900 tgggccgcct ccccgcatcg ataccgtcga ggaagcttaa gctagagctc gctgatcagc 3960 ctcgactgtg ccttctagtt gccagccatc tgttgtttgc ccctcccccg tgccttcctt 4020 gaccctggaa ggtgccactc ccactgtcct ttcctaataa aatgaggaaa ttgcatcgca 4080 ttgtctgagt aggtgtcatt ctattctggg gggtggggtg gggcaggaca gcaaggggga 4140 ggattgggaa gacaatagca ggcatgctgg ggagagtcta gagtcgactg gggagagatc 4200 tgaggaaccc ctagtgatgg agttggccac tccctctctg cgcgctcgct cgctcactga 4260 ggccgcccgg gcaaagcccg ggcgtcgggc gacctttggt cgcccggcct cagtgagcga 4320 gcgagcgcgc agagagggag tggccaac 4348 <210> 13 <211> 4348 <212> DNA <213> Artificial Sequence <220> <223> EXG102-03-2 <400> 13 ttggccactc cctctctgcg cgctcgctcg ctcactgagg ccgggcgacc aaaggtcgcc 60 cgacgcccgg gctttgcccg ggcggcctca gtgagcgagc gagcgcgcag agagggagtg 120 gccaactcca tcactagggg ttcctcagat ctgaattcgg tacctagtta ttaatagtaa 180 tcaattacgg ggtcattagt tcatagccca tatatggagt tccgcgttac ataacttacg 240 gtaaatggcc cgcctggctg accgcccaac gacccccgcc cattgacgtc aataatgacg 300 tatgttccca tagtaacgcc aatagggact ttccattgac gtcaatgggt ggagtattta 360 cggtaaactg cccacttggc agtacatcaa gtgtatcata tgccaagtac gccccctatt 420 gacgtcaatg acggtaaatg gcccgcctgg cattatgccc agtacatgac cttatgggac 480 tttcctactt ggcagtacat ctacgtatta gtcatcgcta ttaccatggt cgaggtgagc 540 cccacgttct gcttcactct ccccatctcc cccccctccc cacccccaat tttgtattta 600 tttatttttt aattattttg tgcagcgatg ggggcggggg gggggggggg gcgcgcgcca 660 ggcggggcgg ggcggggcga ggggcggggc ggggcgaggc ggagaggtgc ggcggcagcc 720 aatcagagcg gcgcgctccg aaagtttcct tttatggcga ggcggcggcg gcggcggccc 780 tataaaaagc gaagcgcgcg gcgggcggga gtcgctgcgc gctgccttcg ccccgtgccc 840 cgctccgccg ccgcctcgcg ccgcccgccc cggctctgac tgaccgcgtt actcccacag 900 gtgagcgggc gggacggccc ttctcctccg ggctgtaatt agcgcttggt ttaatgacgg 960 cttgtttctt ttctgtggct gcgtgaaagc cttgaggggc tccgggaggg ccctttgtgc 1020 ggggggagcg gctcgggggg tgcgtgcgtg tgtgtgtgcg tggggagcgc cgcgtgcggc 1080 tccgcgctgc ccggcggctg tgagcgctgc gggcgcggcg cggggctttg tgcgctccgc 1140 agtgtgcgcg aggggagcgc ggccgggggc ggtgccccgc ggtgcggggg gggctgcgag 1200 gggaacaaag gctgcgtgcg gggtgtgtgc gtgggggggt gagcaggggg tgtgggcgcg 1260 tcggtcgggc tgcaaccccc cctgcacccc cctccccgag ttgctgagca cggcccggct 1320 tcgggtgcgg ggctccgtac ggggcgtggc gcggggctcg ccgtgccggg cggggggtgg 1380 cggcaggtgg gggtgccggg cggggcgggg ccgcctcggg ccggggaggg ctcgggggag 1440 gggcgcggcg gcccccggag cgccggcggc tgtcgaggcg cggcgagccg cagccattgc 1500 cttttatggt aatcgtgcga gagggcgcag ggacttcctt tgtcccaaat ctgtgcggag 1560 ccgaaatctg ggaggcgccg ccgcaccccc tctagcgggc gcggggcgaa gcggtgcggc 1620 gccggcagga aggaaatggg cggggagggc cttcgtgcgt cgccgcgccg ccgtcccctt 1680 ctccctctcc agcctcgggg ctgtccgcgg ggggacggct gccttcgggg gggacggggc 1740 agggcggggt tcggcttctg gcgtgtgacc ggcggctcta gagcctctgc taaccatgtt 1800 catgccttct tctttttcct acagctcctg ggcaacgtgc tggttattgt gctgtctcat 1860 cattttggca aagaattcct cgaagatcta ggcaacgcgt ctcgagagaa ttcgccacca 1920 tggtgagcta ttgggataca ggcgtgctgc tgtgcgccct gctgagctgc ctgctgctga 1980 ccggcagcag cagcggcagc gacaccggca gacctttcgt ggagatgtac agcgagatcc 2040 ctgagatcat ccacatgacc gagggcagag agctggtgat cccttgcaga gtgaccagcc 2100 ctaatatcac cgtgaccctc aagaagttcc ctctggatac cctgatccct gacggcaaga 2160 gaatcatctg ggacagcaga aagggcttca tcatcagcaa tgccacctac aaggagatcg 2220 gcctgctgac ctgcgaggcc accgtgaatg gccacctgta caagaccaat tacctgaccc 2280 acagacagac caataccatc atcgacgtgg tgctgagccc tagccacggc atcgagctga 2340 gcgtgggcga gaagctggtg ctgaattgca ccgccagaac cgagctgaat gtgggcatcg 2400 acttcaattg ggagtaccct agcagcaagc accagcacaa gaagctggtg aatagagacc 2460 tgaagaccca gagcggcagc gagatgaaga aattcctgag caccctgacc atcgacggcg 2520 tgaccagaag cgaccagggc ctgtacacct gcgctgccag cagcggcctg atgaccaaga 2580 agaatagcac cttcgtgaga gtgcacgaga aggacaagac ccacacctgc cctccttgcc 2640 ctgcccctga gctgctgggc ggccctagcg tgttcctgtt ccctcctaag cctaaggaca 2700 ccctcatgat cagcagaacc cctgaggtga cctgcgtggt ggtggacgtg agccacgagg 2760 accctgaggt gaagttcaat tggtacgtgg acggcgtgga ggtgcacaat gccaagacca 2820 agcctagaga ggagcagtac aatagcacct acagagtggt gagcgtgctg accgtgctgc 2880 accaggactg gctgaatggc aaggagtaca agtgcaaggt gagcaataag gccctgcctg 2940 cccctatcga gaagaccatc agcaaggcca agggccagcc tagagagcct caggtgtaca 3000 ccctgcctcc tagcagagac gagctgacca agaatcaggt gagcctgacc tgcctggtga 3060 agggcttcta ccctagcgac atcgccgtgg agtgggagag caatggccag cctgagaata 3120 attacaagac cacccctcct gtgctggaca gcgacggcag cttcttcctg tacagcaagc 3180 tgaccgtgga caagagcaga tggcagcagg gcaatgtgtt cagctgcagc gtgatgcacg 3240 aggccctgca caatcactac acacagaaga gcctgagcct gagccctggc aagtgataag 3300 gatatcaaga tcttgcggcc gctcgataat caacctctgg attacaaaat ttgtgaaaga 3360 ttgactggta ttcttaacta tgttgctcct tttacgctat gtggatacgc tgctttaatg 3420 cctttgtatc atgctattgc ttcccgtatg gctttcattt tctcctcctt gtataaatcc 3480 tggttgctgt ctctttatga ggagttgtgg cccgttgtca ggcaacgtgg cgtggtgtgc 3540 actgtgtttg ctgacgcaac ccccactggt tggggcattg ccaccacctg tcagctcctt 3600 tccgggactt tcgctttccc cctccctatt gccacggcgg aactcatcgc cgcctgcctt 3660 gcccgctgct ggacaggggc tcggctgttg ggcactgaca attccgtggt gttgtcgggg 3720 aaatcatcgt cctttccttg gctgctcgcc tgtgttgcca cctggattct gcgcgggacg 3780 tccttctgct acgtcccttc ggccctcaat ccagcggacc ttccttcccg cggcctgctg 3840 ccggctctgc ggcctcttcc gcgtcttcgc cttcgccctc agacgagtcg gatctccctt 3900 tgggccgcct ccccgcatcg ataccgtcga ggaagcttaa gctagagctc gctgatcagc 3960 ctcgactgtg ccttctagtt gccagccatc tgttgtttgc ccctcccccg tgccttcctt 4020 gaccctggaa ggtgccactc ccactgtcct ttcctaataa aatgaggaaa ttgcatcgca 4080 ttgtctgagt aggtgtcatt ctattctggg gggtggggtg gggcaggaca gcaaggggga 4140 ggattgggaa gacaatagca ggcatgctgg ggagagtcta gagtcgactg gggagagatc 4200 tgaggaaccc ctagtgatgg agttggccac tccctctctg cgcgctcgct cgctcactga 4260 ggccgcccgg gcaaagcccg ggcgtcgggc gacctttggt cgcccggcct cagtgagcga 4320 gcgagcgcgc agagagggag tggccaac 4348 <210> 14 <211> 2755 <212> DNA <213> Artificial Sequence <220> <223> EXG102-04 <400> 14 ctgcgcgctc gctcgctcac tgaggccgcc cgggcaaagc ccgggcgtcg ggcgaccttt 60 ctgcgcgctc gctcgctcac tgaggccgcc cgggcaaagc ccgggcgtcg ggcgaccttt 60 ggtcgcccgg cctcagtgag cgagcgagcg cgcagagagg gagtggaatg cacgcgtgga 120 ggtcgcccgg cctcagtgag cgagcgagcg cgcagagagg gagtggaatg cacgcgtgga 120 tctgagttca attcacgcgt ggtacctctg gtcgttacat aacttacggt aaatggcccg 180 tctgagttca attcacgcgt ggtacctctg gtcgttacat aacttacggt aaatggcccg 180 cctggctgac cgcccaacga cccccgccca ttgacgtcaa taatgacgta tgttcccata 240 cctggctgac cgcccaacga cccccgccca ttgacgtcaa taatgacgta tgttcccata 240 gtaacgccaa tagggacttt ccattgacgt caatgggtgg agtatttacg gtaaactgcc 300 gtaacgccaa tagggacttt ccattgacgt caatgggtgg agtatttacg gtaaactgcc 300 cacttggcag tacatcaagt gtatcatatg ccaagtacgc cccctattga cgtcaatgac 360 cacttggcag tacatcaagt gtatcatatg ccaagtacgc cccctattga cgtcaatgac 360 ggtaaatggc ccgcctggca ttatgcccag tacatgacct tatgggactt tcctacttgg 420 ggtaaatggc ccgcctggca ttatgcccag tacatgacct tatgggactt tcctacttgg 420 cagtacatct actcgaggcc acgttctgct tcactctccc catctccccc ccctccccac 480 cagtacatct actcgaggcc acgttctgct tcactctccc catctccccc ccctccccac 480 ccccaatttt gtatttattt attttttaat tattttgtgc agcgatgggg gcgggggggg 540 ccccaatttt gtatttattt attttttaat tattttgtgc agcgatgggg gcgggggggg 540 ggggggggcg cgcgccaggc ggggcggggc ggggcgaggg gcggggcggg gcgaggcgga 600 ggggggggcg cgcgccaggc ggggcggggc ggggcgaggg gcggggcggg gcgaggcgga 600 gaggtgcggc ggcagccaat cagagcggcg cgctccgaaa gtttcctttt atggcgaggc 660 gaggtgcggc ggcagccaat cagagcggcg cgctccgaaa gtttcctttt atggcgaggc 660 ggcggcggcg gcggccctat aaaaagcgaa gcgcgcggcg ggcgggagcg ggatcagcca 720 ggcggcggcg gcggccctat aaaaagcgaa gcgcgcggcg ggcgggagcg ggatcagcca 720 ccgcggtggc ggcctagagt cgacgaggaa ctgaaaaacc agaaagttaa ctggtaagtt 780 tagtcttttt gtcttttatt tcaggtcccg gatccggtgg tggtgcaaat caaagaactg 840 ctcctcagtg gatgttgcct ttacttctag gcctgtacgg aagtgttact tctgctctaa 900 aagctgcgga attgtacccg cggccgatcc accggtccgg aattcgccac catggtgtct 960 tattgggata ctggcgtgct gctctgtgcc ctcctgagtt gcctgctcct gactggttct 1020 tcttctgggt ccgatactgg gcgccccttc gtggagatgt actccgagat ccctgaaatc 1080 attcacatga ctgagggtcg ggaactggtc atcccatgcc gcgtgacctc tcccaacatt 1140 actgtgaccc tgaagaaatt ccctctggac accctcatcc cagatgggaa gaggatcatt 1200 tgggactcaa gaaagggttt tatcatcagc aacgctacat acaaggagat tggcctgctc 1260 acctgcgaag caacagtgaa cggacacctg tacaagacta attatctcac ccatagacag 1320 acaaacacta tcattgatgt ggtcctgtca ccaagccacg gcatcgagct cagcgtcggt 1380 gaaaagctgg tgctcaattg tacagcccgg actgagctga acgtgggcat tgacttcaat 1440 tgggaatacc ccagctccaa gcaccagcat aagaaactgg tgaaccgcga tctcaaaacc 1500 cagtccggat ctgagatgaa gaaatttctg agcaccctca caatcgacgg cgtgacacga 1560 tccgatcagg gactgtatac ttgcgccgct tctagtggcc tgatgaccaa gaaaaatagc 1620 acattcgtca gggtgcacga aaaggacaaa actcatacct gcccaccttg tccagcacca 1680 gagctgctcg gaggaccatc cgtgttcctg tttccaccca agcccaaaga tactctgatg 1740 atttcacgca cacccgaagt cacttgcgtg gtcgtggacg tgtcccacga ggaccccgaa 1800 gtcaagttta actggtacgt ggacggcgtc gaggtgcata atgctaagac aaaaccccga 1860 gaggaacagt acaactctac ctatagggtc gtgagtgtcc tgacagtgct ccaccaggat 1920 tggctgaacg gaaaggagta taagtgcaaa gtgtctaata aggcactgcc tgccccaatc 1980 gagaaaacaa ttagtaaggc caaagggcag cccagagaac ctcaggtgta cactctgcct 2040 ccatctcggg acgagctcac taagaaccag gtcagtctga cctgtctcgt gaaagggttc 2100 tatcctagtg atatcgctgt ggagtgggaa tcaaatggtc agccagagaa caattacaag 2160 accacacccc ctgtcctgga cagcgatggc tccttctttc tgtattccaa gctcaccgtg 2220 gacaaatctc gatggcagca gggaaacgtc tttagttgtt cagtgatgca cgaagccctc 2280 cataaccact acactcagaa aagcctcagc ctcagccctg ggaaatgata aggatatcaa 2340 gatctacaaa gcttatcgat accgtcgact agagctcgct gatcagcctc gactgtgcct 2400 tctagttgcc agccatctgt tgtttgcccc tcccccgtgc cttccttgac cctggaaggt 2460 gccactccca ctgtcctttc ctaataaaat gaggaaattg catcgcattg tctgagtagg 2520 tgtcattcta ttctgggggg tggggtgggg caggacagca agggggagga ttgggaagtc 2580 tagagcaggc atgctgggga gagatcgatc tgaggaaccc ctagtgatgg agttggccac 2640 tccctctctg cgcgctcgct cgctcactga ggccgggcga ccaaaggtcg cccgacgccc 2700 gggctttgcc cgggcggcct cagtgagcga gcgagcgcgc agagagggag tggcc 2755 <210> 15 <211> 2755 <212> DNA <213> Artificial Sequence <220> <223> EXG102-05 <400> 15 ctgcgcgctc gctcgctcac tgaggccgcc cgggcaaagc ccgggcgtcg ggcgaccttt 60 ctgcgcgctc gctcgctcac tgaggccgcc cgggcaaagc ccgggcgtcg ggcgaccttt 60 ggtcgcccgg cctcagtgag cgagcgagcg cgcagagagg gagtggaatg cacgcgtgga 120 ggtcgcccgg cctcagtgag cgagcgagcg cgcagagagg gagtggaatg cacgcgtgga 120 tctgagttca attcacgcgt ggtacctctg gtcgttacat aacttacggt aaatggcccg 180 tctgagttca attcacgcgt ggtacctctg gtcgttacat aacttacggt aaatggcccg 180 cctggctgac cgcccaacga cccccgccca ttgacgtcaa taatgacgta tgttcccata 240 cctggctgac cgcccaacga cccccgccca ttgacgtcaa taatgacgta tgttcccata 240 gtaacgccaa tagggacttt ccattgacgt caatgggtgg agtatttacg gtaaactgcc 300 gtaacgccaa tagggacttt ccattgacgt caatgggtgg agtatttacg gtaaactgcc 300 cacttggcag tacatcaagt gtatcatatg ccaagtacgc cccctattga cgtcaatgac 360 cacttggcag tacatcaagt gtatcatatg ccaagtacgc cccctattga cgtcaatgac 360 ggtaaatggc ccgcctggca ttatgcccag tacatgacct tatgggactt tcctacttgg 420 ggtaaatggc ccgcctggca ttatgcccag tacatgacct tatgggactt tcctacttgg 420 cagtacatct actcgaggcc acgttctgct tcactctccc catctccccc ccctccccac 480 cagtacatct actcgaggcc acgttctgct tcactctccc catctccccc ccctccccac 480 ccccaatttt gtatttattt attttttaat tattttgtgc agcgatgggg gcgggggggg 540 ccccaatttt gtatttattt attttttaat tattttgtgc agcgatgggg gcgggggggg 540 ggggggggcg cgcgccaggc ggggcggggc ggggcgaggg gcggggcggg gcgaggcgga 600 ggggggggcg cgcgccaggc ggggcggggc ggggcgaggg gcggggcggg gcgaggcgga 600 gaggtgcggc ggcagccaat cagagcggcg cgctccgaaa gtttcctttt atggcgaggc 660 gaggtgcggc ggcagccaat cagagcggcg cgctccgaaa gtttcctttt atggcgaggc 660 ggcggcggcg gcggccctat aaaaagcgaa gcgcgcggcg ggcgggagcg ggatcagcca 720 ggcggcggcg gcggccctat aaaaagcgaa gcgcgcggcg ggcgggagcg ggatcagcca 720 ccgcggtggc ggcctagagt cgacgaggaa ctgaaaaacc agaaagttaa ctggtaagtt 780 tagtcttttt gtcttttatt tcaggtcccg gatccggtgg tggtgcaaat caaagaactg 840 ctcctcagtg gatgttgcct ttacttctag gcctgtacgg aagtgttact tctgctctaa 900 aagctgcgga attgtacccg cggccgatcc accggtccgg aattcgccac catggtgtct 960 tattgggata ctggcgtgct gctctgtgcc ctcctgagtt gcctgctcct gactggttct 1020 tcttctgggt ccgatactgg gcgccccttc gtggagatgt actccgagat ccctgaaatc 1080 attcacatga ctgagggtcg ggaactggtc atcccatgcc gcgtgacctc tcccaacatt 1140 actgtgaccc tgaagaaatt ccctctggac accctcatcc cagatgggaa gaggatcatt 1200 tgggactcaa gaaagggttt tatcatcagc aacgctacat acaaggagat tggcctgctc 1260 acctgcgaag caacagtgaa cggacacctg tacaagacta attatctcac ccatagacag 1320 acaaacacta tcattgatgt ggtcctgtca ccaagccacg gcatcgagct cagcgtcggt 1380 gaaaagctgg tgctcaattg tacagcccgg actgagctga acgtgggcat tgacttcaat 1440 tgggaatacc ccagctccaa gcaccagcat aagaaactgg tgaaccgcga tctcaaaacc 1500 cagtccggat ctgagatgaa gaaatttctg agcaccctca caatcgacgg cgtgacacga 1560 tccgatcagg gactgtatac ttgcgccgct tctagtggcc tgatgaccaa gaaaaatagc 1620 acattcgtca gggtgcacga aaaggacaaa actcatacct gcccaccttg tccagcacca 1680 gagctgctcg gaggaccatc cgtgttcctg tttccaccca agcccaaaga tactctgatg 1740 atttcacgca cacccgaagt cacttgcgtg gtcgtggacg tgtcccacga ggaccccgaa 1800 gtcaagttta actggtacgt ggacggcgtc gaggtgcata atgctaagac aaaaccccga 1860 gaggaacagt acaactctac ctatagggtc gtgagtgtcc tgacagtgct ccaccaggat 1920 tggctgaacg gaaaggagta taagtgcaaa gtgtctaata aggcactgcc tgccccaatc 1980 gagaaaacaa ttagtaaggc caaagggcag cccagagaac ctcaggtgta cactctgcct 2040 ccatctcggg acgagctcac taagaaccag gtcagtctga cctgtctcgt gaaagggttc 2100 tatcctagtg atatcgctgt ggagtgggaa tcaaatggtc agccagagaa caattacaag 2160 accacacccc ctgtcctgga cagcgatggc tccttctttc tgtattccaa gctcaccgtg 2220 gacaaatctc gatggcagca gggaaacgtc tttagttgtt cagtgatgca cgaagccctc 2280 cataaccact acactcagaa aagcctcagc ctcagccctg ggaaatgata aggatatcaa 2340 gatctacaaa gcttatcgat accgtcgact agagctcgct gatcagcctc gactgtgcct 2400 tctagttgcc agccatctgt tgtttgcccc tcccccgtgc cttccttgac cctggaaggt 2460 gccactccca ctgtcctttc ctaataaaat gaggaaattg catcgcattg tctgagtagg 2520 tgtcattcta ttctgggggg tggggtgggg caggacagca agggggagga ttgggaagtc 2580 tagagcaggc atgctgggga gagatcgatc tgaggaaccc ctagtgatgg agttggccac 2640 tccctctctg cgcgctcgct cgctcactga ggccgggcga ccaaaggtcg cccgacgccc 2700 gggctttgcc cgggcggcct cagtgagcga gcgagcgcgc agagagggag tggcc 2755 <210> 16 <211> 3002 <212> DNA <213> Artificial Sequence <220> <223> EXG102-06 <400> 16 ctgcgcgctc gctcgctcac tgaggccgcc cgggcaaagc ccgggcgtcg ggcgaccttt 60 ctgcgcgctc gctcgctcac tgaggccgcc cgggcaaagc ccgggcgtcg ggcgaccttt 60 ggtcgcccgg cctcagtgag cgagcgagcg cgcagagagg gagtggaatg cacgcgtgga 120 ggtcgcccgg cctcagtgag cgagcgagcg cgcagagagg gagtggaatg cacgcgtgga 120 tctgagttca attcacgcgt ggtacctctg gtcgttacat aacttacggt aaatggcccg 180 tctgagttca attcacgcgt ggtacctctg gtcgttacat aacttacggt aaatggcccg 180 cctggctgac cgcccaacga cccccgccca ttgacgtcaa taatgacgta tgttcccata 240 cctggctgac cgcccaacga cccccgccca ttgacgtcaa taatgacgta tgttcccata 240 gtaacgccaa tagggacttt ccattgacgt caatgggtgg agtatttacg gtaaactgcc 300 gtaacgccaa tagggacttt ccattgacgt caatgggtgg agtatttacg gtaaactgcc 300 cacttggcag tacatcaagt gtatcatatg ccaagtacgc cccctattga cgtcaatgac 360 cacttggcag tacatcaagt gtatcatatg ccaagtacgc cccctattga cgtcaatgac 360 ggtaaatggc ccgcctggca ttatgcccag tacatgacct tatgggactt tcctacttgg 420 ggtaaatggc ccgcctggca ttatgcccag tacatgacct tatgggactt tcctacttgg 420 cagtacatct actcgaggcc acgttctgct tcactctccc catctccccc ccctccccac 480 cagtacatct actcgaggcc acgttctgct tcactctccc catctccccc ccctccccac 480 ccccaatttt gtatttattt attttttaat tattttgtgc agcgatgggg gcgggggggg 540 ccccaatttt gtatttattt attttttaat tattttgtgc agcgatgggg gcgggggggg 540 ggggggggcg cgcgccaggc ggggcggggc ggggcgaggg gcggggcggg gcgaggcgga 600 ggggggggcg cgcgccaggc ggggcggggc ggggcgaggg gcggggcggg gcgaggcgga 600 gaggtgcggc ggcagccaat cagagcggcg cgctccgaaa gtttcctttt atggcgaggc 660 gaggtgcggc ggcagccaat cagagcggcg cgctccgaaa gtttcctttt atggcgaggc 660 ggcggcggcg gcggccctat aaaaagcgaa gcgcgcggcg ggcgggagcg ggatcagcca 720 ggcggcggcg gcggccctat aaaaagcgaa gcgcgcggcg ggcgggagcg ggatcagcca 720 ccgcggtggc ggcctagagt cgacgaggaa ctgaaaaacc agaaagttaa ctggtaagtt 780 tagtcttttt gtcttttatt tcaggtcccg gatccggtgg tggtgcaaat caaagaactg 840 ctcctcagtg gatgttgcct ttacttctag gcctgtacgg aagtgttact tctgctctaa 900 aagctgcgga attgtacccg cggccgatcc accggtccgg aattcgccac catggtgtct 960 tattgggata ctggcgtgct gctctgtgcc ctcctgagtt gcctgctcct gactggttct 1020 tcttctgggt ccgatactgg gcgccccttc gtggagatgt actccgagat ccctgaaatc 1080 attcacatga ctgagggtcg ggaactggtc atcccatgcc gcgtgacctc tcccaacatt 1140 actgtgaccc tgaagaaatt ccctctggac accctcatcc cagatgggaa gaggatcatt 1200 tgggactcaa gaaagggttt tatcatcagc aacgctacat acaaggagat tggcctgctc 1260 acctgcgaag caacagtgaa cggacacctg tacaagacta attatctcac ccatagacag 1320 acaaacacta tcattgatgt ggtcctgtca ccaagccacg gcatcgagct cagcgtcggt 1380 gaaaagctgg tgctcaattg tacagcccgg actgagctga acgtgggcat tgacttcaat 1440 tgggaatacc ccagctccaa gcaccagcat aagaaactgg tgaaccgcga tctcaaaacc 1500 cagtccggat ctgagatgaa gaaatttctg agcaccctca caatcgacgg cgtgacacga 1560 tccgatcagg gactgtatac ttgcgccgct tctagtggcc tgatgaccaa gaaaaatagc 1620 acattcgtca gggtgcacga aaaggacaaa actcatacct gcccaccttg tccagcacca 1680 gagctgctcg gaggaccatc cgtgttcctg tttccaccca agcccaaaga tactctgatg 1740 atttcacgca cacccgaagt cacttgcgtg gtcgtggacg tgtcccacga ggaccccgaa 1800 gtcaagttta actggtacgt ggacggcgtc gaggtgcata atgctaagac aaaaccccga 1860 gaggaacagt acaactctac ctatagggtc gtgagtgtcc tgacagtgct ccaccaggat 1920 tggctgaacg gaaaggagta taagtgcaaa gtgtctaata aggcactgcc tgccccaatc 1980 gagaaaacaa ttagtaaggc caaagggcag cccagagaac ctcaggtgta cactctgcct 2040 ccatctcggg acgagctcac taagaaccag gtcagtctga cctgtctcgt gaaagggttc 2100 tatcctagtg atatcgctgt ggagtgggaa tcaaatggtc agccagagaa caattacaag 2160 accacacccc ctgtcctgga cagcgatggc tccttctttc tgtattccaa gctcaccgtg 2220 gacaaatctc gatggcagca gggaaacgtc tttagttgtt cagtgatgca cgaagccctc 2280 cataaccact acactcagaa aagcctcagc ctcagccctg ggaaatgata aggatatcaa 2340 gatctataat caacctctgg attacaaaat ttgtgaaaga ttgactggta ttcttaacta 2400 tgttgctcct tttacgctat gtggatacgc tgctttaatg cctttgtatc atgctattgc 2460 ttcccgtatg gctttcattt tctcctcctt gtataaatcc tggttagttc ttgccacggc 2520 ggaactcatc gccgcctgcc ttgcccgctg ctggacaggg gctcggctgt tgggcactga 2580 caattccgtg gtgttaagct tatcgatacc gtcgactaga gctcgctgat cagcctcgac 2640 tgtgccttct agttgccagc catctgttgt ttgcccctcc cccgtgcctt ccttgaccct 2700 ggaaggtgcc actcccactg tcctttccta ataaaatgag gaaattgcat cgcattgtct 2760 gagtaggtgt cattctattc tggggggtgg ggtggggcag gacagcaagg gggaggattg 2820 ggaagtctag agcaggcatg ctggggagag atcgatctga ggaaccccta gtgatggagt 2880 tggccactcc ctctctgcgc gctcgctcgc tcactgaggc cgggcgacca aaggtcgccc 2940 gacgcccggg ctttgcccgg gcggcctcag tgagcgagcg agcgcgcaga gagggagtgg 3000 cc 3002 <210> 17 <211> 3179 <212> DNA <213> Artificial Sequence <220> <223> EXG102-07 <400> 17 ctgcgcgctc gctcgctcac tgaggccgcc cgggcaaagc ccgggcgtcg ggcgaccttt 60 ggtcgcccgg cctcagtgag cgagcgagcg cgcagagagg gagtggaatg cacgcgtgga 120 tctgagttca attcacgcgt ggtacctctg gtcgttacat aacttacggt aaatggcccg 180 cctggctgac cgcccaacga cccccgccca ttgacgtcaa taatgacgta tgttcccata 240 gtaacgccaa tagggacttt ccattgacgt caatgggtgg agtatttacg gtaaactgcc 300 cacttggcag tacatcaagt gtatcatatg ccaagtacgc cccctattga cgtcaatgac 360 ggtaaatggc ccgcctggca ttatgcccag tacatgacct tatgggactt tcctacttgg 420 cagtacatct actcgaggcc acgttctgct tcactctccc catctccccc ccctccccac 480 ccccaatttt gtatttattt attttttaat tattttgtgc agcgatgggg gcgggggggg 540 ggggggggcg cgcgccaggc ggggcggggc ggggcgaggg gcggggcggg gcgaggcgga 600 gaggtgcggc ggcagccaat cagagcggcg cgctccgaaa gtttcctttt atggcgaggc 660 ggcggcggcg gcggccctat aaaaagcgaa gcgcgcggcg ggcgggagcg ggatcagcca 720 ccgcggtggc ggcctagagt cgacgaggaa ctgaaaaacc agaaagttaa ctggtaagtt 780 tagtcttttt gtcttttatt tcaggtcccg gatccggtgg tggtgcaaat caaagaactg 840 ctcctcagtg gatgttgcct ttacttctag gcctgtacgg aagtgttact tctgctctaa 900 aagctgcgga attgtacccg cggccgatcc accggtccgg aattcgccac catggtgtct 960 tattgggata ctggcgtgct gctctgtgcc ctcctgagtt gcctgctcct gactggttct 1020 tcttctgggt ccgatactgg gcgccccttc gtggagatgt actccgagat ccctgaaatc 1080 attcacatga ctgagggtcg ggaactggtc atcccatgcc gcgtgacctc tcccaacatt 1140 actgtgaccc tgaagaaatt ccctctggac accctcatcc cagatgggaa gaggatcatt 1200 tgggactcaa gaaagggttt tatcatcagc aacgctacat acaaggagat tggcctgctc 1260 tgggactcaa gaaagggttt tatcatcagc aacgctacat acaaggagat tggcctgctc 1260 acctgcgaag caacagtgaa cggacacctg tacaagacta attatctcac ccatagacag 1320 acctgcgaag caacagtgaa cggacacctg tacaagacta attatctcac ccatagacag 1320 acaaacacta tcattgatgt ggtcctgtca ccaagccacg gcatcgagct cagcgtcggt 1380 acaaacacta tcattgatgt ggtcctgtca ccaagccacg gcatcgagct cagcgtcggt 1380 gaaaagctgg tgctcaattg tacagcccgg actgagctga acgtgggcat tgacttcaat 1440 gaaaagctgg tgctcaattg tacagcccgg actgagctga acgtgggcat tgacttcaat 1440 tgggaatacc ccagctccaa gcaccagcat aagaaactgg tgaaccgcga tctcaaaacc 1500 tgggaatacc ccagctccaa gcaccagcat aagaaactgg tgaaccgcga tctcaaaacc 1500 cagtccggat ctgagatgaa gaaatttctg agcaccctca caatcgacgg cgtgacacga 1560 cagtccggat ctgagatgaa gaaatttctg agcaccctca caatcgacgg cgtgacacga 1560 tccgatcagg gactgtatac ttgcgccgct tctagtggcc tgatgaccaa gaaaaatagc 1620 tccgatcagg gactgtatac ttgcgccgct tctagtggcc tgatgaccaa gaaaaatagc 1620 acattcgtca gggtgcacga aaaggacaaa actcatacct gcccaccttg tccagcacca 1680 acattcgtca gggtgcacga aaaggacaaa actcatacct gcccaccttg tccagcacca 1680 gagctgctcg gaggaccatc cgtgttcctg tttccaccca agcccaaaga tactctgatg 1740 gagctgctcg gaggaccatc cgtgttcctg tttccaccca agcccaaaga tactctgatg 1740 atttcacgca cacccgaagt cacttgcgtg gtcgtggacg tgtcccacga ggaccccgaa 1800 atttcacgca cacccgaagt cacttgcgtg gtcgtggacg tgtcccacga ggaccccgaa 1800 gtcaagttta actggtacgt ggacggcgtc gaggtgcata atgctaagac aaaaccccga 1860 gtcaagttta actggtacgt ggacggcgtc gaggtgcata atgctaagac aaaaccccga 1860 gaggaacagt acaactctac ctatagggtc gtgagtgtcc tgacagtgct ccaccaggat 1920 gaggaacagt acaactctac ctatagggtc gtgagtgtcc tgacagtgct ccaccaggat 1920 tggctgaacg gaaaggagta taagtgcaaa gtgtctaata aggcactgcc tgccccaatc 1980 gagaaaacaa ttagtaaggc caaagggcag cccagagaac ctcaggtgta cactctgcct 2040 ccatctcggg acgagctcac taagaaccag gtcagtctga cctgtctcgt gaaagggttc 2100 tatcctagtg atatcgctgt ggagtgggaa tcaaatggtc agccagagaa caattacaag 2160 accacacccc ctgtcctgga cagcgatggc tccttctttc tgtattccaa gctcaccgtg 2220 gacaaatctc gatggcagca gggaaacgtc tttagttgtt cagtgatgca cgaagccctc 2280 cataaccact acactcagaa aagcctcagc ctcagccctg ggaaaggcgg cggcggaggc 2340 gcccagcaag aagagtgcga gtgggatccc tggacctgcg agcacatggg atccggcagc 2400 gccaccggag gatccggaag caccgcctcc agcggctccg gcagcgccac ccaccaggag 2460 gagtgtgagt gggacccctg gacctgcgaa cacatgctgg agtgataagg atatcaagat 2520 ctataatcaa cctctggatt acaaaatttg tgaaagattg actggtattc ttaactatgt 2580 tgctcctttt acgctatgtg gatacgctgc tttaatgcct ttgtatcatg ctattgcttc 2640 ccgtatggct ttcattttct cctccttgta taaatcctgg ttagttcttg ccacggcgga 2700 actcatcgcc gcctgccttg cccgctgctg gacaggggct cggctgttgg gcactgacaa 2760 ttccgtggtg ttaagcttat cgataccgtc gactagagct cgctgatcag cctcgactgt 2820 gccttctagt tgccagccat ctgttgtttg cccctccccc gtgccttcct tgaccctgga 2880 aggtgccact cccactgtcc tttcctaata aaatgaggaa attgcatcgc attgtctgag 2940 taggtgtcat tctattctgg ggggtggggt ggggcaggac agcaaggggg aggattggga 3000 agtctagagc aggcatgctg gggagagatc gatctgagga acccctagtg atggagttgg 3060 ccactccctc tctgcgcgct cgctcgctca ctgaggccgg gcgaccaaag gtcgcccgac 3120 gcccgggctt tgcccgggcg gcctcagtga gcgagcgagc gcgcagagag ggagtggcc 3179 <210> 18 <211> 2932 <212> DNA <213> Artificial Sequence <220> <223> EXG102-08 <400> 18 ctgcgcgctc gctcgctcac tgaggccgcc cgggcaaagc ccgggcgtcg ggcgaccttt 60 ggtcgcccgg cctcagtgag cgagcgagcg cgcagagagg gagtggaatg cacgcgtgga 120 tctgagttca attcacgcgt ggtacctctg gtcgttacat aacttacggt aaatggcccg 180 cctggctgac cgcccaacga cccccgccca ttgacgtcaa taatgacgta tgttcccata 240 gtaacgccaa tagggacttt ccattgacgt caatgggtgg agtatttacg gtaaactgcc 300 cacttggcag tacatcaagt gtatcatatg ccaagtacgc cccctattga cgtcaatgac 360 ggtaaatggc ccgcctggca ttatgcccag tacatgacct tatgggactt tcctacttgg 420 cagtacatct actcgaggcc acgttctgct tcactctccc catctccccc ccctccccac 480 ccccaatttt gtatttattt attttttaat tattttgtgc agcgatgggg gcgggggggg 540 ggggggggcg cgcgccaggc ggggcggggc ggggcgaggg gcggggcggg gcgaggcgga 600 gaggtgcggc ggcagccaat cagagcggcg cgctccgaaa gtttcctttt atggcgaggc 660 ggcggcggcg gcggccctat aaaaagcgaa gcgcgcggcg ggcgggagcg ggatcagcca 720 ccgcggtggc ggcctagagt cgacgaggaa ctgaaaaacc agaaagttaa ctggtaagtt 780 tagtcttttt gtcttttatt tcaggtcccg gatccggtgg tggtgcaaat caaagaactg 840 ctcctcagtg gatgttgcct ttacttctag gcctgtacgg aagtgttact tctgctctaa 900 aagctgcgga attgtacccg cggccgatcc accggtccgg aattcgccac catggtgtct 960 tattgggata ctggcgtgct gctctgtgcc ctcctgagtt gcctgctcct gactggttct 1020 tcttctgggt ccgatactgg gcgccccttc gtggagatgt actccgagat ccctgaaatc 1080 attcacatga ctgagggtcg ggaactggtc atcccatgcc gcgtgacctc tcccaacatt 1140 actgtgaccc tgaagaaatt ccctctggac accctcatcc cagatgggaa gaggatcatt 1200 tgggactcaa gaaagggttt tatcatcagc aacgctacat acaaggagat tggcctgctc 1260 acctgcgaag caacagtgaa cggacacctg tacaagacta attatctcac ccatagacag 1320 acaaacacta tcattgatgt ggtcctgtca ccaagccacg gcatcgagct cagcgtcggt 1380 gaaaagctgg tgctcaattg tacagcccgg actgagctga acgtgggcat tgacttcaat 1440 tgggaatacc ccagctccaa gcaccagcat aagaaactgg tgaaccgcga tctcaaaacc 1500 cagtccggat ctgagatgaa gaaatttctg agcaccctca caatcgacgg cgtgacacga 1560 tccgatcagg gactgtatac ttgcgccgct tctagtggcc tgatgaccaa gaaaaatagc 1620 acattcgtca gggtgcacga aaaggacaaa actcatacct gcccaccttg tccagcacca 1680 gagctgctcg gaggaccatc cgtgttcctg tttccaccca agcccaaaga tactctgatg 1740 atttcacgca cacccgaagt cacttgcgtg gtcgtggacg tgtcccacga ggaccccgaa 1800 gtcaagttta actggtacgt ggacggcgtc gaggtgcata atgctaagac aaaaccccga 1860 gaggaacagt acaactctac ctatagggtc gtgagtgtcc tgacagtgct ccaccaggat 1920 tggctgaacg gaaaggagta taagtgcaaa gtgtctaata aggcactgcc tgccccaatc 1980 gagaaaacaa ttagtaaggc caaagggcag cccagagaac ctcaggtgta cactctgcct 2040 ccatctcggg acgagctcac taagaaccag gtcagtctga cctgtctcgt gaaagggttc 2100 tatcctagtg atatcgctgt ggagtgggaa tcaaatggtc agccagagaa caattacaag 2160 accacacccc ctgtcctgga cagcgatggc tccttctttc tgtattccaa gctcaccgtg 2220 gacaaatctc gatggcagca gggaaacgtc tttagttgtt cagtgatgca cgaagccctc 2280 cataaccact acactcagaa aagcctcagc ctcagccctg ggaaaggcgg cggcggaggc 2340 gcccagcaag aagagtgcga gtgggatccc tggacctgcg agcacatggg atccggcagc 2400 gccaccggag gatccggaag caccgcctcc agcggctccg gcagcgccac ccaccaggag 2460 gagtgtgagt gggacccctg gacctgcgaa cacatgctgg agtgataagg atatcaagat 2520 ctacaaagct tatcgatacc gtcgactaga gctcgctgat cagcctcgac tgtgccttct 2580 agttgccagc catctgttgt ttgcccctcc cccgtgcctt ccttgaccct ggaaggtgcc 2640 actcccactg tcctttccta ataaaatgag gaaattgcat cgcattgtct gagtaggtgt 2700 cattctattc tggggggtgg ggtggggcag gacagcaagg gggaggattg ggaagtctag 2760 agcaggcatg ctggggagag atcgatctga ggaaccccta gtgatggagt tggccactcc 2820 ctctctgcgc gctcgctcgc tcactgaggc cgggcgacca aaggtcgccc gacgcccggg 2880 ctttgcccgg gcggcctcag tgagcgagcg agcgcgcaga gagggagtgg cc 2932 <210> 19 <211> 3937 <212> DNA <213> Artificial Sequence <220> <223> EXG102-09 <400> 19 ttggccactc cctctctgcg cgctcgctcg ctcactgagg ccgggcgacc aaaggtcgcc 60 cgacgcccgg gctttgcccg ggcggcctca gtgagcgagc gagcgcgcag agagggagtg 120 gccaactcca tcactagggg ttcctcagat ctgaattcgg tacctagtta ttaatagtaa 180 tcaattacgg ggtcattagt tcatagccca tatatggagt tccgcgttac ataacttacg 240 gtaaatggcc cgcctggctg accgcccaac gacccccgcc cattgacgtc aataatgacg 300 tatgttccca tagtaacgcc aatagggact ttccattgac gtcaatgggt ggagtattta 360 cggtaaactg cccacttggc agtacatcaa gtgtatcata tgccaagtac gccccctatt 420 gacgtcaatg acggtaaatg gcccgcctgg cattatgccc agtacatgac cttatgggac 480 tttcctactt ggcagtacat ctacgtatta gtcatcgcta ttaccatggt cgaggtgagc 540 cccacgttct gcttcactct ccccatctcc cccccctccc cacccccaat tttgtattta 600 tttatttttt aattattttg tgcagcgatg ggggcggggg gggggggggg gcgcgcgcca 660 ggcggggcgg ggcggggcga ggggcggggc ggggcgaggc ggagaggtgc ggcggcagcc 720 aatcagagcg gcgcgctccg aaagtttcct tttatggcga ggcggcggcg gcggcggccc 780 tataaaaagc gaagcgcgcg gcgggcggga gtcgctgcgc gctgccttcg ccccgtgccc 840 cgctccgccg ccgcctcgcg ccgcccgccc cggctctgac tgaccgcgtt actcccacag 900 gtgagcgggc gggacggccc ttctcctccg ggctgtaatt agcgcttggt ttaatgacgg 960 cttgtttctt ttctgtggct gcgtgaaagc cttgaggggc tccgggaggg ccctttgtgc 1020 ggggggagcg gctcgggggg tgcgtgcgtg tgtgtgtgcg tggggagcgc cgcgtgcggc 1080 tccgcgctgc ccggcggctg tgagcgctgc gggcgcggcg cggggctttg tgcgctccgc 1140 agtgtgcgcg aggggagcgc ggccgggggc ggtgccccgc ggtgcggggg gggctgcgag 1200 gggaacaaag gctgcgtgcg gggtgtgtgc gtgggggggt gagcaggggg tgtgggcgcg 1260 tcggtcgggc tgcaaccccc cctgcacccc cctccccgag ttgctgagca cggcccggct 1320 tcgggtgcgg ggctccgtac ggggcgtggc gcggggctcg ccgtgccggg cggggggtgg 1380 cggcaggtgg gggtgccggg cggggcgggg ccgcctcggg ccggggaggg ctcgggggag 1440 gggcgcggcg gcccccggag cgccggcggc tgtcgaggcg cggcgagccg cagccattgc 1500 cttttatggt aatcgtgcga gagggcgcag ggacttcctt tgtcccaaat ctgtgcggag 1560 ccgaaatctg ggaggcgccg ccgcaccccc tctagcgggc gcggggcgaa gcggtgcggc 1620 gccggcagga aggaaatggg cggggagggc cttcgtgcgt cgccgcgccg ccgtcccctt 1680 ctccctctcc agcctcgggg ctgtccgcgg ggggacggct gccttcgggg gggacggggc 1740 agggcggggt tcggcttctg gcgtgtgacc ggcggctcta gagcctctgc taaccatgtt 1800 catgccttct tctttttcct acagctcctg ggcaacgtgc tggttattgt gctgtctcat 1860 cattttggca aagaattcct cgaagatcta ggcaacgcgt ctcgaacgcg tctcgagaga 1920 attcgccacc atggtgtctt attgggatac tggcgtgctg ctctgtgccc tcctgagttg 1980 cctgctcctg actggttctt cttctgggtc cgatactggg cgccccttcg tggagatgta 2040 ctccgagatc cctgaaatca ttcacatgac tgagggtcgg gaactggtca tcccatgccg 2100 ctccgagatc cctgaaatca ttcacatgac tgagggtcgg gaactggtca tcccatgccg 2100 cgtgacctct cccaacatta ctgtgaccct gaagaaattc cctctggaca ccctcatccc 2160 cgtgacctct cccaacatta ctgtgaccct gaagaaattc cctctggaca ccctcatccc 2160 agatgggaag aggatcattt gggactcaag aaagggtttt atcatcagca acgctacata 2220 agatgggaag aggatcattt gggactcaag aaagggtttt atcatcagca acgctacata 2220 caaggagatt ggcctgctca cctgcgaagc aacagtgaac ggacacctgt acaagactaa 2280 caaggagatt ggcctgctca cctgcgaagc aacagtgaac ggacacctgt acaagactaa 2280 ttatctcacc catagacaga caaacactat cattgatgtg gtcctgtcac caagccacgg 2340 ttatctcacc catagacaga caaacactat cattgatgtg gtcctgtcac caagccacgg 2340 catcgagctc agcgtcggtg aaaagctggt gctcaattgt acagcccgga ctgagctgaa 2400 catcgagctc agcgtcggtg aaaagctggt gctcaattgt acagcccgga ctgagctgaa 2400 cgtgggcatt gacttcaatt gggaataccc cagctccaag caccagcata agaaactggt 2460 cgtgggcatt gacttcaatt gggaataccc cagctccaag caccagcata agaaactggt 2460 gaaccgcgat ctcaaaaccc agtccggatc tgagatgaag aaatttctga gcaccctcac 2520 gaaccgcgat ctcaaaaccc agtccggatc tgagatgaag aaatttctga gcaccctcac 2520 aatcgacggc gtgacacgat ccgatcaggg actgtatact tgcgccgctt ctagtggcct 2580 aatcgacggc gtgacacgat ccgatcaggg actgtatact tgcgccgctt ctagtggcct 2580 gatgaccaag aaaaatagca cattcgtcag ggtgcacgaa aagggcggcg gcggaggcag 2640 gatgaccaag aaaaatagca cattcgtcag ggtgcacgaa aagggcggcg gcggaggcag 2640 cggcggcggc ggaggcgccc agcaagaaga gtgcgagtgg gatccctgga cctgcgagca 2700 cggcggcggc ggaggcgccc agcaagaaga gtgcgagtgg gatccctgga cctgcgagca 2700 catgggatcc ggcagcgcca ccggaggatc cggaagcacc gcctccagcg gctccggcag 2760 catgggatcc ggcagcgcca ccggaggatc cggaagcacc gcctccagcg gctccggcag 2760 cgccacccac caggaggagt gtgagtggga cccctggacc tgcgaacaca tgctggagga 2820 caaaactcat acctgcccac cttgtccagc accagagctg ctcggaggac catccgtgtt 2880 cctgtttcca cccaagccca aagatactct gatgatttca cgcacacccg aagtcacttg 2940 cgtggtcgtg gacgtgtccc acgaggaccc cgaagtcaag tttaactggt acgtggacgg 3000 cgtcgaggtg cataatgcta agacaaaacc ccgagaggaa cagtacaact ctacctatag 3060 ggtcgtgagt gtcctgacag tgctccacca ggattggctg aacggaaagg agtataagtg 3120 caaagtgtct aataaggcac tgcctgcccc aatcgagaaa acaattagta aggccaaagg 3180 gcagcccaga gaacctcagg tgtacactct gcctccatct cgggacgagc tcactaagaa 3240 ccaggtcagt ctgacctgtc tcgtgaaagg gttctatcct agtgatatcg ctgtggagtg 3300 ggaatcaaat ggtcagccag agaacaatta caagaccaca ccccctgtcc tggacagcga 3360 tggctccttc tttctgtatt ccaagctcac cgtggacaaa tctcgatggc agcagggaaa 3420 cgtctttagt tgttcagtga tgcacgaagc cctccataac cactacactc agaaaagcct 3480 cagcctcagc cctgggaaat gataagcggc cgcaagctta agagcttaag ctagagctcg 3540 ctgatcagcc tcgactgtgc cttctagttg ccagccatct gttgtttgcc cctcccccgt 3600 gccttccttg accctggaag gtgccactcc cactgtcctt tcctaataaa atgaggaaat 3660 tgcatcgcat tgtctgagta ggtgtcattc tattctgggg ggtggggtgg ggcaggacag 3720 caagggggag gattgggaag acaatagcag gcatgctggg gagagtctag agtcgactgg 3780 ggagagatct gaggaacccc tagtgatgga gttggccact ccctctctgc gcgctcgctc 3840 gctcactgag gccgcccggg caaagcccgg gcgtcgggcg acctttggtc gcccggcctc 3900 agtgagcgag cgagcgcgca gagagggagt ggccaac 3937 <210> 20 <211> 3224 <212> DNA <213> Artificial Sequence <220> <223> EXG102-10 <400> 20 ctgcgcgctc gctcgctcac tgaggccgcc cgggcaaagc ccgggcgtcg ggcgaccttt 60 ggtcgcccgg cctcagtgag cgagcgagcg cgcagagagg gagtggaatg cacgcgtgga 120 tctgagttca attcacgcgt ggtacctctg gtcgttacat aacttacggt aaatggcccg 180 cctggctgac cgcccaacga cccccgccca ttgacgtcaa taatgacgta tgttcccata 240 gtaacgccaa tagggacttt ccattgacgt caatgggtgg agtatttacg gtaaactgcc 300 cacttggcag tacatcaagt gtatcatatg ccaagtacgc cccctattga cgtcaatgac 360 ggtaaatggc ccgcctggca ttatgcccag tacatgacct tatgggactt tcctacttgg 420 cagtacatct actcgaggcc acgttctgct tcactctccc catctccccc ccctccccac 480 ccccaatttt gtatttattt attttttaat tattttgtgc agcgatgggg gcgggggggg 540 ggggggggcg cgcgccaggc ggggcggggc ggggcgaggg gcggggcggg gcgaggcgga 600 gaggtgcggc ggcagccaat cagagcggcg cgctccgaaa gtttcctttt atggcgaggc 660 ggcggcggcg gcggccctat aaaaagcgaa gcgcgcggcg ggcgggagcg ggatcagcca 720 ccgcggtggc ggcctagagt cgacgaggaa ctgaaaaacc agaaagttaa ctggtaagtt 780 tagtcttttt gtcttttatt tcaggtcccg gatccggtgg tggtgcaaat caaagaactg 840 ctcctcagtg gatgttgcct ttacttctag gcctgtacgg aagtgttact tctgctctaa 900 aagctgcgga attgtacccg cggccgatcc accggtccgg aattacgcgt ctcgagagaa 960 ttcgccacca tggtgtctta ttgggatact ggcgtgctgc tctgtgccct cctgagttgc 1020 ctgctcctga ctggttcttc ttctgggtcc gatactgggc gccccttcgt ggagatgtac 1080 tccgagatcc ctgaaatcat tcacatgact gagggtcggg aactggtcat cccatgccgc 1140 gtgacctctc ccaacattac tgtgaccctg aagaaattcc ctctggacac cctcatccca 1200 gatgggaaga ggatcatttg ggactcaaga aagggtttta tcatcagcaa cgctacatac 1260 aaggagattg gcctgctcac ctgcgaagca acagtgaacg gacacctgta caagactaat 1320 tatctcaccc atagacagac aaacactatc attgatgtgg tcctgtcacc aagccacggc 1380 atcgagctca gcgtcggtga aaagctggtg ctcaattgta cagcccggac tgagctgaac 1440 gtgggcattg acttcaattg ggaatacccc agctccaagc accagcataa gaaactggtg 1500 aaccgcgatc tcaaaaccca gtccggatct gagatgaaga aatttctgag caccctcaca 1560 atcgacggcg tgacacgatc cgatcaggga ctgtatactt gcgccgcttc tagtggcctg 1620 atgaccaaga aaaatagcac attcgtcagg gtgcacgaaa agggcggcgg cggaggcagc 1680 ggcggcggcg gaggcgccca gcaagaagag tgcgagtggg atccctggac ctgcgagcac 1740 atgggatccg gcagcgccac cggaggatcc ggaagcaccg cctccagcgg ctccggcagc 1800 gccacccacc aggaggagtg tgagtgggac ccctggacct gcgaacacat gctggaggac 1860 aaaactcata cctgcccacc ttgtccagca ccagagctgc tcggaggacc atccgtgttc 1920 ctgtttccac ccaagcccaa agatactctg atgatttcac gcacacccga agtcacttgc 1980 gtggtcgtgg acgtgtccca cgaggacccc gaagtcaagt ttaactggta cgtggacggc 2040 gtcgaggtgc ataatgctaa gacaaaaccc cgagaggaac agtacaactc tacctatagg 2100 gtcgtgagtg tcctgacagt gctccaccag gattggctga acggaaagga gtataagtgc 2160 aaagtgtcta ataaggcact gcctgcccca atcgagaaaa caattagtaa ggccaaaggg 2220 cagcccagag aacctcaggt gtacactctg cctccatctc gggacgagct cactaagaac 2280 caggtcagtc tgacctgtct cgtgaaaggg ttctatccta gtgatatcgc tgtggagtgg 2340 gaatcaaatg gtcagccaga gaacaattac aagaccacac cccctgtcct ggacagcgat 2400 ggctccttct ttctgtattc caagctcacc gtggacaaat ctcgatggca gcagggaaac 2460 gtctttagtt gttcagtgat gcacgaagcc ctccataacc actacactca gaaaagcctc 2520 agcctcagcc ctgggaaatg ataagcggcc gcaagcttaa gagatctata atcaacctct 2580 ggattacaaa atttgtgaaa gattgactgg tattcttaac tatgttgctc cttttacgct 2640 atgtggatac gctgctttaa tgcctttgta tcatgctatt gcttcccgta tggctttcat 2700 tttctcctcc ttgtataaat cctggttagt tcttgccacg gcggaactca tcgccgcctg 2760 ccttgcccgc tgctggacag gggctcggct gttgggcact gacaattccg tggtgttaag 2820 cttatcgata ccgtcgacta gagctcgctg atcagcctcg actgtgcctt ctagttgcca 2880 gccatctgtt gtttgcccct cccccgtgcc ttccttgacc ctggaaggtg ccactcccac 2940 tgtcctttcc taataaaatg aggaaattgc atcgcattgt ctgagtaggt gtcattctat 3000 tctggggggt ggggtggggc aggacagcaa gggggaggat tgggaagtct agagcaggca 3060 tgctggggag agatcgatct gaggaacccc tagtgatgga gttggccact ccctctctgc 3120 gcgctcgctc gctcactgag gccgggcgac caaaggtcgc ccgacgcccg ggctttgccc 3180 gggcggcctc agtgagcgag cgagcgcgca gagagggagt ggcc 3224 <210> 21 <211> 3937 <212> DNA <213> Artificial sequence <220> <223> EXG102-11 <400> 21 ttggccactc cctctctgcg cgctcgctcg ctcactgagg ccgggcgacc aaaggtcgcc 60 cgacgcccgg gctttgcccg ggcggcctca gtgagcgagc gagcgcgcag agagggagtg 120 gccaactcca tcactagggg ttcctcagat ctgaattcgg tacctagtta ttaatagtaa 180 tcaattacgg ggtcattagt tcatagccca tatatggagt tccgcgttac ataacttacg 240 gtaaatggcc cgcctggctg accgcccaac gacccccgcc cattgacgtc aataatgacg 300 tatgttccca tagtaacgcc aatagggact ttccattgac gtcaatgggt ggagtattta 360 cggtaaactg cccacttggc agtacatcaa gtgtatcata tgccaagtac gccccctatt 420 gacgtcaatg acggtaaatg gcccgcctgg cattatgccc agtacatgac cttatgggac 480 tttcctactt ggcagtacat ctacgtatta gtcatcgcta ttaccatggt cgaggtgagc 540 cccacgttct gcttcactct ccccatctcc cccccctccc cacccccaat tttgtattta 600 tttatttttt aattattttg tgcagcgatg ggggcggggg gggggggggg gcgcgcgcca 660 ggcggggcgg ggcggggcga ggggcggggc ggggcgaggc ggagaggtgc ggcggcagcc 720 aatcagagcg gcgcgctccg aaagtttcct tttatggcga ggcggcggcg gcggcggccc 780 tataaaaagc gaagcgcgcg gcgggcggga gtcgctgcgc gctgccttcg ccccgtgccc 840 cgctccgccg ccgcctcgcg ccgcccgccc cggctctgac tgaccgcgtt actcccacag 900 gtgagcgggc gggacggccc ttctcctccg ggctgtaatt agcgcttggt ttaatgacgg 960 cttgtttctt ttctgtggct gcgtgaaagc cttgaggggc tccgggaggg ccctttgtgc 1020 ggggggagcg gctcgggggg tgcgtgcgtg tgtgtgtgcg tggggagcgc cgcgtgcggc 1080 tccgcgctgc ccggcggctg tgagcgctgc gggcgcggcg cggggctttg tgcgctccgc 1140 agtgtgcgcg aggggagcgc ggccgggggc ggtgccccgc ggtgcggggg gggctgcgag 1200 gggaacaaag gctgcgtgcg gggtgtgtgc gtgggggggt gagcaggggg tgtgggcgcg 1260 tcggtcgggc tgcaaccccc cctgcacccc cctccccgag ttgctgagca cggcccggct 1320 tcgggtgcgg ggctccgtac ggggcgtggc gcggggctcg ccgtgccggg cggggggtgg 1380 cggcaggtgg gggtgccggg cggggcgggg ccgcctcggg ccggggaggg ctcgggggag 1440 gggcgcggcg gcccccggag cgccggcggc tgtcgaggcg cggcgagccg cagccattgc 1500 cttttatggt aatcgtgcga gagggcgcag ggacttcctt tgtcccaaat ctgtgcggag 1560 ccgaaatctg ggaggcgccg ccgcaccccc tctagcgggc gcggggcgaa gcggtgcggc 1620 gccggcagga aggaaatggg cggggagggc cttcgtgcgt cgccgcgccg ccgtcccctt 1680 ctccctctcc agcctcgggg ctgtccgcgg ggggacggct gccttcgggg gggacggggc 1740 agggcggggt tcggcttctg gcgtgtgacc ggcggctcta gagcctctgc taaccatgtt 1800 catgccttct tctttttcct acagctcctg ggcaacgtgc tggttattgt gctgtctcat 1860 cattttggca aagaattcct cgaagatcta ggcaacgcgt ctcgaacgcg tctcgagaga 1920 attcgccacc atggtgtctt attgggatac tggcgtgctg ctctgtgccc tcctgagttg 1980 cctgctcctg actggttctt cttctggggg cggcggcgga ggcgcccagc aagaagagtg 2040 cgagtgggat ccctggacct gcgagcacat gggatccggc agcgccaccg gaggatccgg 2100 aagcaccgcc tccagcggct ccggcagcgc cacccaccag gaggagtgtg agtgggaccc 2160 ctggacctgc gaacacatgc tggagggcgg cggcggaggc agctccgata ctgggcgccc 2220 cttcgtggag atgtactccg agatccctga aatcattcac atgactgagg gtcgggaact 2280 ggtcatccca tgccgcgtga cctctcccaa cattactgtg accctgaaga aattccctct 2340 ggacaccctc atcccagatg ggaagaggat catttgggac tcaagaaagg gttttatcat 2400 cagcaacgct acatacaagg agattggcct gctcacctgc gaagcaacag tgaacggaca 2460 cctgtacaag actaattatc tcacccatag acagacaaac actatcattg atgtggtcct 2520 gtcaccaagc cacggcatcg agctcagcgt cggtgaaaag ctggtgctca attgtacagc 2580 ccggactgag ctgaacgtgg gcattgactt caattgggaa taccccagct ccaagcacca 2640 gcataagaaa ctggtgaacc gcgatctcaa aacccagtcc ggatctgaga tgaagaaatt 2700 tctgagcacc ctcacaatcg acggcgtgac acgatccgat cagggactgt atacttgcgc 2760 cgcttctagt ggcctgatga ccaagaaaaa tagcacattc gtcagggtgc acgaaaagga 2820 caaaactcat acctgcccac cttgtccagc accagagctg ctcggaggac catccgtgtt 2880 cctgtttcca cccaagccca aagatactct gatgatttca cgcacacccg aagtcacttg 2940 cgtggtcgtg gacgtgtccc acgaggaccc cgaagtcaag tttaactggt acgtggacgg 3000 cgtcgaggtg cataatgcta agacaaaacc ccgagaggaa cagtacaact ctacctatag 3060 ggtcgtgagt gtcctgacag tgctccacca ggattggctg aacggaaagg agtataagtg 3120 caaagtgtct aataaggcac tgcctgcccc aatcgagaaa acaattagta aggccaaagg 3180 gcagcccaga gaacctcagg tgtacactct gcctccatct cgggacgagc tcactaagaa 3240 ccaggtcagt ctgacctgtc tcgtgaaagg gttctatcct agtgatatcg ctgtggagtg 3300 ggaatcaaat ggtcagccag agaacaatta caagaccaca ccccctgtcc tggacagcga 3360 tggctccttc tttctgtatt ccaagctcac cgtggacaaa tctcgatggc agcagggaaa 3420 cgtctttagt tgttcagtga tgcacgaagc cctccataac cactacactc agaaaagcct 3480 cagcctcagc cctgggaaat gataagcggc cgcaagctta agagcttaag ctagagctcg 3540 ctgatcagcc tcgactgtgc cttctagttg ccagccatct gttgtttgcc cctcccccgt 3600 gccttccttg accctggaag gtgccactcc cactgtcctt tcctaataaa atgaggaaat 3660 tgcatcgcat tgtctgagta ggtgtcattc tattctgggg ggtggggtgg ggcaggacag 3720 caagggggag gattgggaag acaatagcag gcatgctggg gagagtctag agtcgactgg 3780 ggagagatct gaggaacccc tagtgatgga gttggccact ccctctctgc gcgctcgctc 3840 gctcactgag gccgcccggg caaagcccgg gcgtcgggcg acctttggtc gcccggcctc 3900 agtgagcgag cgagcgcgca gagagggagt ggccaac 3937 <210> 22 <211> 2941 <212> DNA <213> Artificial Sequence <220> <223> EXG102-12 <400> 22 ctgcgcgctc gctcgctcac tgaggccgcc cgggcaaagc ccgggcgtcg ggcgaccttt 60 ggtcgcccgg cctcagtgag cgagcgagcg cgcagagagg gagtggaatg cacgcgtgga 120 tctgagttca attcacgcgt ggtacctctg gtcgttacat aacttacggt aaatggcccg 180 cctggctgac cgcccaacga cccccgccca ttgacgtcaa taatgacgta tgttcccata 240 gtaacgccaa tagggacttt ccattgacgt caatgggtgg agtatttacg gtaaactgcc 300 cacttggcag tacatcaagt gtatcatatg ccaagtacgc cccctattga cgtcaatgac 360 ggtaaatggc ccgcctggca ttatgcccag tacatgacct tatgggactt tcctacttgg 420 cagtacatct actcgaggcc acgttctgct tcactctccc catctccccc ccctccccac 480 ccccaatttt gtatttattt attttttaat tattttgtgc agcgatgggg gcgggggggg 540 ggggggggcg cgcgccaggc ggggcggggc ggggcgaggg gcggggcggg gcgaggcgga 600 gaggtgcggc ggcagccaat cagagcggcg cgctccgaaa gtttcctttt atggcgaggc 660 ggcggcggcg gcggccctat aaaaagcgaa gcgcgcggcg ggcgggagcg ggatcagcca 720 ccgcggtggc ggcctagagt cgacgaggaa ctgaaaaacc agaaagttaa ctggtaagtt 780 tagtcttttt gtcttttatt tcaggtcccg gatccggtgg tggtgcaaat caaagaactg 840 ctcctcagtg gatgttgcct ttacttctag gcctgtacgg aagtgttact tctgctctaa 900 aagctgcgga attgtacccg cggccgatcc accggtccgg aattcgccac catggtgtct 960 tattgggata ctggcgtgct gctctgtgcc ctcctgagtt gcctgctcct gactggttct 1020 tcttctgggg gcggcggcgg aggcgcccag caagaagagt gcgagtggga tccctggacc 1080 tgcgagcaca tgggatccgg cagcgccacc ggaggatccg gaagcaccgc ctccagcggc 1140 tccggcagcg ccacccacca ggaggagtgt gagtgggacc cctggacctg cgaacacatg 1200 ctggagggcg gcggcggagg cagctccgat actgggcgcc ccttcgtgga gatgtactcc 1260 gagatccctg aaatcattca catgactgag ggtcgggaac tggtcatccc atgccgcgtg 1320 acctctccca acattactgt gaccctgaag aaattccctc tggacaccct catcccagat 1380 gggaagagga tcatttggga ctcaagaaag ggttttatca tcagcaacgc tacatacaag 1440 gagattggcc tgctcacctg cgaagcaaca gtgaacggac acctgtacaa gactaattat 1500 ctcacccata gacagacaaa cactatcatt gatgtggtcc tgtcaccaag ccacggcatc 1560 gagctcagcg tcggtgaaaa gctggtgctc aattgtacag cccggactga gctgaacgtg 1620 ggcattgact tcaattggga ataccccagc tccaagcacc agcataagaa actggtgaac 1680 cgcgatctca aaacccagtc cggatctgag atgaagaaat ttctgagcac cctcacaatc 1740 gacggcgtga cacgatccga tcagggactg tatacttgcg ccgcttctag tggcctgatg 1800 accaagaaaa atagcacatt cgtcagggtg cacgaaaagg acaaaactca tacctgccca 1860 ccttgtccag caccagagct gctcggagga ccatccgtgt tcctgtttcc acccaagccc 1920 aaagatactc tgatgatttc acgcacaccc gaagtcactt gcgtggtcgt ggacgtgtcc 1980 cacgaggacc ccgaagtcaa gtttaactgg tacgtggacg gcgtcgaggt gcataatgct 2040 aagacaaaac cccgagagga acagtacaac tctacctata gggtcgtgag tgtcctgaca 2100 gtgctccacc aggattggct gaacggaaag gagtataagt gcaaagtgtc taataaggca 2160 ctgcctgccc caatcgagaa aacaattagt aaggccaaag ggcagcccag agaacctcag 2220 gtgtacactc tgcctccatc tcgggacgag ctcactaaga accaggtcag tctgacctgt 2280 ctcgtgaaag ggttctatcc tagtgatatc gctgtggagt gggaatcaaa tggtcagcca 2340 gagaacaatt acaagaccac accccctgtc ctggacagcg atggctcctt ctttctgtat 2400 tccaagctca ccgtggacaa atctcgatgg cagcagggaa acgtctttag ttgttcagtg 2460 atgcacgaag ccctccataa ccactacact cagaaaagcc tcagcctcag ccctgggaaa 2520 tgataagcgg ccgcaagctt atcgataccg tcgactagag ctcgctgatc agcctcgact 2580 gtgccttcta gttgccagcc atctgttgtt tgcccctccc ccgtgccttc cttgaccctg 2640 gaaggtgcca ctcccactgt cctttcctaa taaaatgagg aaattgcatc gcattgtctg 2700 agtaggtgtc attctattct ggggggtggg gtggggcagg acagcaaggg ggaggattgg 2760 gaagtctaga gcaggcatgc tggggagaga tcgatctgag gaacccctag tgatggagtt 2820 ggccactccc tctctgcgcg ctcgctcgct cactgaggcc gggcgaccaa aggtcgcccg 2880 acgcccgggc tttgcccggg cggcctcagt gagcgagcga gcgcgcagag agggagtggc 2940 c 2941 <210> 23 <211> 2603 <212> DNA <213> Artificial sequence <220> <223> EXG102-13 <400> 23 ctgcgcgctc gctcgctcac tgaggccgcc cgggcaaagc ccgggcgtcg ggcgaccttt 60 ggtcgcccgg cctcagtgag cgagcgagcg cgcagagagg gagtggaatg cacgcgtgga 120 tctgagttca attcacgcgt ggtacctctg gtcgttacat aacttacggt aaatggcccg 180 cctggctgac cgcccaacga cccccgccca ttgacgtcaa taatgacgta tgttcccata 240 gtaacgccaa tagggacttt ccattgacgt caatgggtgg agtatttacg gtaaactgcc 300 cacttggcag tacatcaagt gtatcatatg ccaagtacgc cccctattga cgtcaatgac 360 ggtaaatggc ccgcctggca ttatgcccag tacatgacct tatgggactt tcctacttgg 420 cagtacatct actcgaggcc acgttctgct tcactctccc catctccccc ccctccccac 480 ccccaatttt gtatttattt attttttaat tattttgtgc agcgatgggg gcgggggggg 540 ggggggggcg cgcgccaggc ggggcggggc ggggcgaggg gcggggcggg gcgaggcgga 600 gaggtgcggc ggcagccaat cagagcggcg cgctccgaaa gtttcctttt atggcgaggc 660 ggcggcggcg ggcggccctat aaaaagcgaa gcgcgcggcg ggcgggagcg ggatcagcca 720 ccgcggtggc ggcctagagt cgacgaggaa ctgaaaaacc agaaagttaa ctggtaagtt 780 tagtcttttt gtcttttatt tcaggtcccg gatccggtgg tggtgcaaat caaagaactg 840 ctcctcagtg gatgttgcct ttacttctag gcctgtacgg aagtgttact tctgctctaa 900 aagctgcgga attgtacccg cggccgatcc accggtccgg aattacgcgt ctcgagagaa 960 ttcgccacca tggtgtctta ttgggatact ggcgtgctgc tctgtgccct cctgagttgc 1020 ctgctcctga ctggttcttc ttctgggtcc gatactgggc gccccttcgt ggagatgtac 1080 tccgagatcc ctgaaatcat tcacatgact gagggtcggg aactggtcat cccatgccgc 1140 gtgacctctc ccaacattac tgtgaccctg aagaaattcc ctctggacac cctcatccca 1200 gatgggaaga ggatcatttg ggactcaaga aagggtttta tcatcagcaa cgctacatac 1260 aaggagattg gcctgctcac ctgcgaagca acagtgaacg gacacctgta caagactaat 1320 tatctcaccc atagacagac aaacactatc attgatgtgg tcctgtcacc aagccacggc 1380 atcgagctca gcgtcggtga aaagctggtg ctcaattgta cagcccggac tgagctgaac 1440 gtgggcattg acttcaattg ggaatacccc agctccaagc accagcataa gaaactggtg 1500 aaccgcgatc tcaaaaccca gtccggatct gagatgaaga aatttctgag caccctcaca 1560 atcgacggcg tgacacgatc cgatcaggga ctgtatactt gcgccgcttc tagtggcctg 1620 atgaccaaga aaaatagcac attcgtcagg gtgcacgaaa aggacaaaac tcatacctgc 1680 ccaccttgtc cagcaccaga gctgctcgga ggaccatccg tgggcggcgg cggaggcagc 1740 ggcggcggcg gaggcgccca gcaagaagag tgcgagtggg atccctggac ctgcgagcac 1800 atgggatccg gcagcgccac cggaggatcc ggaagcaccg cctccagcgg ctccggcagc 1860 gccacccacc aggaggagtg tgagtgggac ccctggacct gcgaacacat gctggagtga 1920 taagcggccg caagcttaag agatctataa tcaacctctg gattacaaaa tttgtgaaag 1980 attgactggt attcttaact atgttgctcc ttttacgcta tgtggatacg ctgctttaat 2040 gcctttgtat catgctattg cttcccgtat ggctttcatt ttctcctcct tgtataaatc 2100 ctggttagtt cttgccacgg cggaactcat cgccgcctgc cttgcccgct gctggacagg 2160 ggctcggctg ttgggcactg acaattccgt ggtgttaagc ttatcgatac cgtcgactag 2220 agctcgctga tcagcctcga ctgtgccttc tagttgccag ccatctgttg tttgcccctc 2280 ccccgtgcct tccttgaccc tggaaggtgc cactcccact gtcctttcct aataaaatga 2340 ggaaattgca tcgcattgtc tgagtaggtg tcattctatt ctggggggtg gggtggggca 2400 ggacagcaag ggggaggatt gggaagtcta gagcaggcat gctggggaga gatcgatctg 2460 aggaacccct agtgatggag ttggccactc cctctctgcg cgctcgctcg ctcactgagg 2520 ccgggcgacc aaaggtcgcc cgacgcccgg gctttgcccg ggcggcctca gtgagcgagc 2580 gagcgcgcag agagggagtg gcc 2603 <210> 24 <211> 5 <212> PRT <213> Synthetic Sequence <220> <223> Exemplary Peptide Linker <400> 24 Asp Gly Gly Gly Ser 1 5 <210> 25 <211> 5 <212> PRT <213> Synthetic Sequence <220> <223> Exemplary Peptide Linker <400> 25 Thr Gly Glu Lys Pro 1 5 <210> 26 <211> 4 <212> PRT <213> Synthetic Sequence <220> <223> Exemplary Peptide Linker <400> 26 Gly Gly Arg Arg 1 <210> 27 <211> 24 <212> PRT <213> Synthetic Sequence <220> <223> Exemplary Peptide Linker <400> 27 Gly Gly Gly Gly Ser Gly Gly Gly Gly Ser Gly Gly Gly Gly Gly Gly 1 5 10 15 Ser Gly Ser Gly Gly Gly Gly Ser 20 <210> 28 <211> 34 <212> PRT <213> Artificial Sequence <220> <223> Exemplary Peptide Linker <400> 28 Gly Gly Gly Gly Ser Gly Gly Gly Gly Ser Gly Gly Gly Gly Gly Gly 1 5 10 15 Ser Gly Ser Gly Gly Gly Gly Ser Gly Gly Gly Gly Ser Gly Gly Gly 20 25 30 Gly Ser <210> 29 <211> 14 <212> PRT <213> Artificial Sequence <220> <223> Exemplary Peptide Linker <400> 29 Glu Gly Lys Ser Ser Gly Ser Gly Ser Glu Ser Lys Val Asp 1 5 10 <210> 30 <211> 16 <212> PRT <213> Artificial Sequence <220> <223> Exemplary Peptide Linker <400> 30 Lys Glu Ser Gly Ser Val Ser Ser Glu Gln Leu Ala Gln Phe Arg Ser 1 5 10 15 <210> 31 <211> 8 <212> PRT <213> Artificial Sequence <220> <223> Exemplary Peptide Linker <400> 31 Gly Gly Arg Arg Gly Gly Gly Ser 1 5 <210> 32 <211> 9 <212> PRT <213> Artificial sequence <220> <223> Exemplary peptide linker <400> 32 Leu Arg Gln Arg Asp Gly Glu Arg Pro 1 5 <210> 33 <211> 12 <212> PRT <213> Artificial sequence <220> <223> Exemplary peptide linker <400> 33 Leu Arg Gln Lys Asp Gly Gly Gly Ser Glu Arg Pro 1 5 10 <210> 34 <211> 16 <212> PRT <213> Artificial sequence <220> <223> Exemplary peptide linker <400> 34 Leu Arg Gln Lys Asp Gly Gly Gly Ser Gly Gly Gly Ser Glu Arg Pro 1 5 10 15 <210> 35 <211> 16 <212> PRT <213> Artificial sequence <220> <223> Exemplary peptide linker <400> 35 Gly Ser Thr Ser Gly Ser Gly Lys Pro Gly Ser Gly Glu Gly Ser Thr 1 5 10 15 <210> 36 <211> 14 <212> PRT <213> Artificial sequence <220> <223> Exemplary peptide linker <400> 36 Gly Ser Thr Ser Gly Ser Gly Lys Ser Ser Glu Gly Lys Gly 1 5 10 <210> 37 <211> 18 <212> PRT <213> Artificial sequence <220> <223> Exemplary peptide linker <400> 37 Lys Glu Ser Gly Ser Val Ser Ser Glu Gln Leu Ala Gln Phe Arg Ser 1 5 10 15 Leu Asp <210> 38 <211> 2 <212> PRT <213> Artificial sequence <220> <223> Exemplary peptide linker <220> <221> MISC_FEATURE <222> (1)..(2) <223> GS can be repeated n times, where n is an integer, including, for example, 1, 2, 3, 4, 5 and 6 <400> 38 Gly Ser 1 <210> 39 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Exemplary peptide linker <220> <221> MISC_FEATURE <222> (1)..(5) <223> GSGGS can be repeated n times, where n is an integer, including, for example, 1, 2, 3, 4, 5, and 6 <400> 39 Gly Ser Gly Gly Ser 1 5 <210> 40 <211> 4 <212> PRT <213> Artificial sequence <220> <223> Exemplary peptide linker <220> <221> MISC_FEATURE <222> (1)..(4) <223> GGGS can be repeated n times, where n is an integer, including, for example, 1, 2, 3, 4, 5, and 6 <400> 40 Gly Gly Gly Ser 1 <210> 41 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Exemplary peptide linker <220> <221> MISC_FEATURE <222> (1)..(5) <223> GGGGS can be repeated n times, where n is an integer, including, for example, 1, 2, 3, 4, 5, and 6 <400> 41 Gly Gly Gly Gly Ser 1 5 <210> 42 <211> 6 <212> PRT <213> Artificial sequence <220> <223> Exemplary peptide linker <220> <221> MISC_FEATURE <222> (1)..(6) <223> GGGGGS can be repeated n times, where n is an integer, including, for example, 1, 2, 3, 4, 5, and 6 <400> 42 Gly Gly Gly Gly Gly Ser 1 5 <210> 43 <211> 735 <212> PRT <213> Artificial sequence <220> <223> AAV2 (UniProt: P03135-1) <400> 43 Met Ala Ala Asp Gly Tyr Leu Pro Asp Trp Leu Glu Asp Thr Leu Ser 1 5 10 15 Glu Gly Ile Arg Gln Trp Trp Lys Leu Lys Pro Gly Pro Pro Pro Pro 20 25 30 Lys Pro Ala Glu Arg His Lys Asp Asp Ser Arg Gly Leu Val Leu Pro 35 40 45 Gly Tyr Lys Tyr Leu Gly Pro Phe Asn Gly Leu Asp Lys Gly Glu Pro 50 55 60 Val Asn Glu Ala Asp Ala Ala Ala Leu Glu His Asp Lys Ala Tyr Asp 65 70 75 80 Arg Gln Leu Asp Ser Gly Asp Asn Pro Tyr Leu Lys Tyr Asn His Ala 85 90 95 Asp Ala Glu Phe Gln Glu Arg Leu Lys Glu Asp Thr Ser Phe Gly Gly 100 105 110 Asn Leu Gly Arg Ala Val Phe Gln Ala Lys Lys Arg Val Leu Glu Pro 115 120 125 Leu Gly Leu Val Glu Glu Pro Val Lys Thr Ala Pro Gly Lys Lys Arg 130 135 140 Pro Val Glu His Ser Pro Val Glu Pro Asp Ser Ser Ser Gly Thr Gly 145 150 155 160 Lys Ala Gly Gln Gln Pro Ala Arg Lys Arg Leu Asn Phe Gly Gln Thr 165 170 175 Gly Asp Ala Asp Ser Val Pro Asp Pro Gln Pro Leu Gly Gln Pro Pro 180 185 190 Ala Ala Pro Ser Gly Leu Gly Thr Asn Thr Met Ala Thr Gly Ser Gly 195 200 205 Ala Pro Met Ala Asp Asn Asn Glu Gly Ala Asp Gly Val Gly Asn Ser 210 215 220 Ser Gly Asn Trp His Cys Asp Ser Thr Trp Met Gly Asp Arg Val Ile 225 230 235 240 Thr Thr Ser Thr Arg Thr Trp Ala Leu Pro Thr Tyr Asn Asn His Leu 245 250 255 Tyr Lys Gln Ile Ser Ser Gln Ser Gly Ala Ser Asn Asp Asn His Tyr 260 265 270 Phe Gly Tyr Ser Thr Pro Trp Gly Tyr Phe Asp Phe Asn Arg Phe His 275 280 285 Cys His Phe Ser Pro Arg Asp Trp Gln Arg Leu Ile Asn Asn Asn Trp 290 295 300 Gly Phe Arg Pro Lys Arg Leu Asn Phe Lys Leu Phe Asn Ile Gln Val 305 310 315 320 Lys Glu Val Thr Gln Asn Asp Gly Thr Thr Thr Ile Ala Asn Asn Leu 325 330 335 Thr Ser Thr Val Gln Val Phe Thr Asp Ser Glu Tyr Gln Leu Pro Tyr 340 345 350 Val Leu Gly Ser Ala His Gln Gly Cys Leu Pro Pro Phe Pro Ala Asp 355 360 365 Val Phe Met Val Pro Gln Tyr Gly Tyr Leu Thr Leu Asn Asn Gly Ser 370 375 380 Gln Ala Val Gly Arg Ser Ser Phe Tyr Cys Leu Glu Tyr Phe Pro Ser 385 390 395 400 Gln Met Leu Arg Thr Gly Asn Asn Phe Thr Phe Ser Tyr Thr Phe Glu 405 410 415 Asp Val Pro Phe His Ser Ser Tyr Ala His Ser Gln Ser Leu Asp Arg 420 425 430 Leucine Methionine Asparagine Proline Leucine Isoleucine Aspartic acid Glutamine Tyrosine Leucine Tyrosine Tyrosine Leucine Serine Arginine Threonine 435 440 445 Asparagine Threonine Proline Serine Glycine Threonine Threonine Threonine Glutamine Serine Arginine Leucine Glutamine Phenylalanine Serine Glutamine 450 455 460 Alanine Glycine Alanine Serine Aspartic acid Isoleucine Arginine Aspartic acid Glutamine Serine Arginine Asparagine Tryptophan Leucine Proline Glycine 465 470 475 480 Proline Cysteine Tyrosine Arginine Glutamine Glutamine Arginine Valine Serine Lysine Threonine Serine Alanine Aspartic acid Asparagine Asparagine 485 490 495 Asparagine Serine Glutamic acid Tyrosine Serine Tryptophan Threonine Glycine Alanine Threonine Lysine Tyrosine Histidine Leucine Asparagine Glycine 500 505 510 Arginine Aspartic acid Serine Leucine Valine Asparagine Proline Glycine Proline Alanine Methionine Alanine Serine Histidine Lysine Aspartic acid 515 520 525 Aspartic acid Glutamic acid Glutamic acid Lysine Phenylalanine Phenylalanine Proline Glutamine Serine Glycine Valine Leucine Isoleucine Phenylalanine Glycine Lysine 530 535 540 Glutamine Glycine Serine Glutamic acid Lysine Threonine Asparagine Valine Aspartic acid Isoleucine Glutamic acid Lysine Valine Methionine Isoleucine Threonine 545 550 555 560 Aspartic acid Glutamic acid Glutamic acid Glutamic acid Isoleucine Arginine Threonine Threonine Asparagine Proline Valine Alanine Threonine Glutamic acid Glutamine Tyrosine 565 570 575 Glycine Serine Valine Serine Threonine Asparagine Leucine Glutamine Arginine Glycine Asparagine Arginine Glutamine Alanine Alanine Threonine 580 585 590 Ala Asp Val Asn Thr Gln Gly Val Leu Pro Gly Met Val Trp Gln Asp 595 600 605 Arg Asp Val Tyr Leu Gln Gly Pro Ile Trp Ala Lys Ile Pro His Thr 610 615 620 Asp Gly His Phe His Pro Ser Pro Leu Met Gly Gly Phe Gly Leu Lys 625 630 635 640 His Pro Pro Pro Gln Ile Leu Ile Lys Asn Thr Pro Val Pro Ala Asn 645 650 655 Pro Ser Thr Thr Phe Ser Ala Ala Lys Phe Ala Ser Phe Ile Thr Gln 660 665 670 Tyr Ser Thr Gly Gln Val Ser Val Glu Ile Glu Trp Glu Leu Gln Lys 675 680 685 Glu Asn Ser Lys Arg Trp Asn Pro Glu Ile Gln Tyr Thr Ser Asn Tyr 690 695 700 Asn Lys Ser Val Asn Val Asp Phe Thr Val Asp Thr Asn Gly Val Tyr 705 710 715 720 Ser Glu Pro Arg Pro Ile Gly Thr Arg Tyr Leu Thr Arg Asn Leu 725 730 735 <210> 44 <211> 738 <212> PRT <213> Artificial sequence <220> <223> AAV8 (Uniprot Q8JQF8_9VIRU) <400> 44 Met Ala Ala Asp Gly Tyr Leu Pro Asp Trp Leu Glu Asp Asn Leu Ser 1 5 10 15 Glu Gly Ile Arg Glu Trp Trp Ala Leu Lys Pro Gly Ala Pro Lys Pro 20 25 30 Lys Ala Asn Gln Gln Lys Gln Asp Asp Gly Arg Gly Leu Val Leu Pro 35 40 45 Gly Tyr Lys Tyr Leu Gly Pro Phe Asn Gly Leu Asp Lys Gly Glu Pro 50 55 60 Val Asn Ala Ala Asp Ala Ala Ala Leu Glu His Asp Lys Ala Tyr Asp 65 70 75 80 Gln Gln Leu Gln Ala Gly Asp Asn Pro Tyr Leu Arg Tyr Asn His Ala 85 90 95 Asp Ala Glu Phe Gln Glu Arg Leu Gln Glu Asp Thr Ser Phe Gly Gly 100 105 110 Asn Leu Gly Arg Ala Val Phe Gln Ala Lys Lys Arg Val Leu Glu Pro 115 120 125 Leu Gly Leu Val Glu Glu Gly Ala Lys Thr Ala Pro Gly Lys Lys Arg 130 135 140 Pro Val Glu Pro Ser Pro Gln Arg Ser Pro Asp Ser Ser Thr Gly Ile 145 150 155 160 Gly Lys Lys Gly Gln Gln Pro Ala Arg Lys Arg Leu Asn Phe Gly Gln 165 170 175 Thr Gly Asp Ser Glu Ser Val Pro Asp Pro Gln Pro Leu Gly Glu Pro 180 185 190 Pro Ala Ala Pro Ser Gly Val Gly Pro Asn Thr Met Ala Ala Gly Gly 195 200 205 Gly Ala Pro Met Ala Asp Asn Asn Glu Gly Ala Asp Gly Val Gly Ser 210 215 220 Ser Ser Gly Asn Trp His Cys Asp Ser Thr Trp Leu Gly Asp Arg Val 225 230 235 240 Ile Thr Thr Ser Thr Arg Thr Trp Ala Leu Pro Thr Tyr Asn Asn His 245 250 255 Leu Tyr Lys Gln Ile Ser Asn Gly Thr Ser Gly Gly Ala Thr Asn Asp 260 265 270 Asn Thr Tyr Phe Gly Tyr Ser Thr Pro Trp Gly Tyr Phe Asp Phe Asn 275 280 285 Arg Phe His Cys His Phe Ser Pro Arg Asp Trp Gln Arg Leu Ile Asn 290 295 300 Asn Asn Trp Gly Phe Arg Pro Lys Arg Leu Ser Phe Lys Leu Phe Asn 305 310 315 320 Ile Gln Val Lys Glu Val Thr Gln Asn Glu Gly Thr Lys Thr Ile Ala 325 330 335 Asn Asn Leu Thr Ser Thr Ile Gln Val Phe Thr Asp Ser Glu Tyr Gln 340 345 350 Leu Pro Tyr Val Leu Gly Ser Ala His Gln Gly Cys Leu Pro Pro Phe 355 360 365 Pro Ala Asp Val Phe Met Ile Pro Gln Tyr Gly Tyr Leu Thr Leu Asn 370 375 380 Asn Gly Ser Gln Ala Val Gly Arg Ser Ser Phe Tyr Cys Leu Glu Tyr 385 390 395 400 Phe Pro Ser Gln Met Leu Arg Thr Gly Asn Asn Phe Gln Phe Thr Tyr 405 410 415 Thr Phe Glu Asp Val Pro Phe His Ser Ser Tyr Ala His Ser Gln Ser 420 425 430 Leu Asp Arg Leu Met Asn Pro Leu Ile Asp Gln Tyr Leu Tyr Tyr Leu 435 440 445 Ser Arg Thr Gln Thr Thr Gly Gly Thr Ala Asn Thr Gln Thr Leu Gly 450 455 460 Phe Ser Gln Gly Gly Pro Asn Thr Met Ala Asn Gln Ala Lys Asn Trp 465 470 475 480 Leu Pro Gly Pro Cys Tyr Arg Gln Gln Arg Val Ser Thr Thr Thr Gly 485 490 495 Gln Asn Asn Asn Ser Asn Phe Ala Trp Thr Ala Gly Thr Lys Tyr His 500 505 510 Leu Asn Gly Arg Asn Ser Leu Ala Asn Pro Gly Ile Ala Met Ala Thr 515 520 525 His Lys Asp Asp Glu Glu Arg Phe Phe Pro Ser Asn Gly Ile Leu Ile 530 535 540 Phe Gly Lys Gln Asn Ala Ala Arg Asp Asn Ala Asp Tyr Ser Asp Val 545 550 555 560 Met Leu Thr Ser Glu Glu Glu Ile Lys Thr Thr Asn Pro Val Ala Thr 565 570 575 Glu Glu Tyr Gly Ile Val Ala Asp Asn Leu Gln Gln Gln Asn Thr Ala 580 585 590 Pro Gln Ile Gly Thr Val Asn Ser Gln Gly Ala Leu Pro Gly Met Val 595 600 605 Trp Gln Asn Arg Asp Val Tyr Leu Gln Gly Pro Ile Trp Ala Lys Ile 610 615 620 Pro His Thr Asp Gly Asn Phe His Pro Ser Pro Leu Met Gly Gly Phe 625 630 635 640 Gly Leu Lys His Pro Pro Pro Gln Ile Leu Ile Lys Asn Thr Pro Val 645 650 655 Pro Ala Asp Pro Pro Thr Thr Phe Asn Gln Ser Lys Leu Asn Ser Phe 660 665 670 Ile Thr Gln Tyr Ser Thr Gly Gln Val Ser Val Glu Ile Glu Trp Glu 675 680 685 Leu Gln Lys Glu Asn Ser Lys Arg Trp Asn Pro Glu Ile Gln Tyr Thr 690 695 700 Ser Asn Tyr Tyr Lys Ser Thr Ser Val Asp Phe Ala Val Asn Thr Glu 705 710 715 720 Gly Val Tyr Ser Glu Pro Arg Pro Ile Gly Thr Arg Tyr Leu Thr Arg 725 730 735 Asn Leu <210> 45 <211> 736 <212> PRT <213> Artificial Sequence <220> <223> AAVrh8 (Uniprot Q808Y3_9VIRU) <400> 45 Met Ala Ala Asp Gly Tyr Leu Pro Asp Trp Leu Glu Asp Asn Leu Ser 1 5 10 15 Glu Gly Ile Arg Glu Trp Trp Asp Leu Lys Pro Gly Ala Pro Lys Pro 20 25 30 Lys Ala Asn Gln Gln Lys Gln Asp Asp Gly Arg Gly Leu Val Leu Pro 35 40 45 Gly Tyr Lys Tyr Leu Gly Pro Phe Asn Gly Leu Asp Lys Gly Glu Pro 50 55 60 Val Asn Ala Ala Asp Ala Ala Ala Leu Glu His Asp Lys Ala Tyr Asp 65 70 75 80 Gln Gln Leu Lys Ala Gly Asp Asn Pro Tyr Leu Arg Tyr Asn His Ala 85 90 95 Asp Ala Glu Phe Gln Glu Arg Leu Gln Glu Asp Thr Ser Phe Gly Gly 100 105 110 Asn Leu Gly Arg Ala Val Phe Gln Ala Lys Lys Arg Val Leu Glu Pro 115 120 125 Leu Gly Leu Val Glu Glu Gly Ala Lys Thr Ala Pro Gly Lys Lys Arg 130 135 140 Pro Val Glu Gln Ser Pro Gln Glu Pro Asp Ser Ser Ser Gly Ile Gly 145 150 155 160 Lys Thr Gly Gln Gln Pro Ala Lys Lys Arg Leu Asn Phe Gly Gln Thr 165 170 175 Gly Asp Ser Glu Ser Val Pro Asp Pro Gln Pro Leu Gly Glu Pro Pro 180 185 190 Ala Ala Pro Ser Gly Leu Gly Pro Asn Thr Met Ala Ser Gly Gly Gly 195 200 205 Ala Pro Met Ala Asp Asn Asn Glu Gly Ala Asp Gly Val Gly Asn Ser 210 215 220 Ser Gly Asn Trp His Cys Asp Ser Thr Trp Leu Gly Asp Arg Val Ile 225 230 235 240 Thr Thr Ser Thr Arg Thr Trp Ala Leu Pro Thr Tyr Asn Asn His Leu 245 250 255 Tyr Lys Gln Ile Ser Asn Gly Thr Ser Gly Gly Ser Thr Asn Asp Asn 260 265 270 Thr Tyr Phe Gly Tyr Ser Thr Pro Trp Gly Tyr Phe Asp Phe Asn Arg 275 280 285 Phe His Cys His Phe Ser Pro Arg Asp Trp Gln Arg Leu Ile Asn Asn 290 295 300 Asn Trp Gly Phe Arg Pro Lys Arg Leu Asn Phe Lys Leu Phe Asn Ile 305 310 315 320 Gln Val Lys Glu Val Thr Thr Asn Glu Gly Thr Lys Thr Ile Ala Asn 325 330 335 Asn Leu Thr Ser Thr Val Gln Val Phe Thr Asp Ser Glu Tyr Gln Leu 340 345 350 Pro Tyr Val Leu Gly Ser Ala His Gln Gly Cys Leu Pro Pro Phe Pro 355 360 365 Ala Asp Val Phe Met Val Pro Gln Tyr Gly Tyr Leu Thr Leu Asn Asn 370 375 380 Gly Ser Gln Ala Leu Gly Arg Ser Ser Phe Tyr Cys Leu Glu Tyr Phe 385 390 395 400 Pro Ser Gln Met Leu Arg Thr Gly Asn Asn Phe Gln Phe Ser Tyr Thr 405 410 415 Phe Glu Asp Val Pro Phe His Ser Ser Tyr Ala His Ser Gln Ser Leu 420 425 430 Asp Arg Leu Met Asn Pro Leu Ile Asp Gln Tyr Leu Tyr Tyr Leu Val 435 440 445 Arg Thr Gln Thr Thr Gly Thr Gly Gly Thr Gln Thr Leu Ala Phe Ser 450 455 460 Gln Ala Gly Pro Ser Ser Met Ala Asn Gln Ala Arg Asn Trp Val Pro 465 470 475 480 Gly Pro Cys Tyr Arg Gln Gln Arg Val Ser Thr Thr Thr Asn Gln Asn 485 490 495 Asn Asn Ser Asn Phe Ala Trp Thr Gly Ala Ala Lys Phe Lys Leu Asn 500 505 510 Gly Arg Asp Ser Leu Met Asn Pro Gly Val Ala Met Ala Ser His Lys 515 520 525 Asp Asp Asp Asp Arg Phe Phe Pro Ser Ser Gly Val Leu Ile Phe Gly 530 535 540 Lys Gln Gly Ala Gly Asn Asp Gly Val Asp Tyr Ser Gln Val Leu Ile 545 550 555 560 Thr Asp Glu Glu Glu Ile Lys Ala Thr Asn Pro Val Ala Thr Glu Glu 565 570 575 Tyr Gly Ala Val Ala Ile Asn Asn Gln Ala Ala Asn Thr Gln Ala Gln 580 585 590 Thr Gly Leu Val His Asn Gln Gly Val Ile Pro Gly Met Val Trp Gln 595 600 605 Asn Arg Asp Val Tyr Leu Gln Gly Pro Ile Trp Ala Lys Ile Pro His 610 615 620 Thr Asp Gly Asn Phe His Pro Ser Pro Leu Met Gly Gly Phe Gly Leu 625 630 635 640 Lys His Pro Pro Pro Gln Ile Leu Ile Lys Asn Thr Pro Val Pro Ala 645 650 655 Asp Pro Pro Leu Thr Phe Asn Gln Ala Lys Leu Asn Ser Phe Ile Thr 660 665 670 Gln Tyr Ser Thr Gly Gln Val Ser Val Glu Ile Glu Trp Glu Leu Gln 675 680 685 Lys Glu Asn Ser Lys Arg Trp Asn Pro Glu Ile Gln Tyr Thr Ser Asn 690 695 700 Tyr Tyr Lys Ser Thr Asn Val Asp Phe Ala Val Asn Thr Glu Gly Val 705 710 715 720 Tyr Ser Glu Pro Arg Pro Ile Gly Thr Arg Tyr Leu Thr Arg Asn Leu 725 730 735 <210> 46 <211> 736 <212> PRT <213> Artificial Sequence <220> <223> AAV9 (Uniprot Q6JC40_9VIRU) <400> 46 Met Ala Ala Asp Gly Tyr Leu Pro Asp Trp Leu Glu Asp Asn Leu Ser 1 5 10 15 Glu Gly Ile Arg Glu Trp Trp Ala Leu Lys Pro Gly Ala Pro Gln Pro 20 25 30 Lys Ala Asn Gln Gln His Gln Asp Asn Ala Arg Gly Leu Val Leu Pro 35 40 45 Gly Tyr Lys Tyr Leu Gly Pro Gly Asn Gly Leu Asp Lys Gly Glu Pro 50 55 60 Val Asn Ala Ala Asp Ala Ala Ala Leu Glu His Asp Lys Ala Tyr Asp 65 70 75 80 Gln Gln Leu Lys Ala Gly Asp Asn Pro Tyr Leu Lys Tyr Asn His Ala 85 90 95 Asp Ala Glu Phe Gln Glu Arg Leu Lys Glu Asp Thr Ser Phe Gly Gly 100 105 110 Asn Leu Gly Arg Ala Val Phe Gln Ala Lys Lys Arg Leu Leu Glu Pro 115 120 125 Leu Gly Leu Val Glu Glu Ala Ala Lys Thr Ala Pro Gly Lys Lys Arg 130 135 140 Pro Val Glu Gln Ser Pro Gln Glu Pro Asp Ser Ser Ala Gly Ile Gly 145 150 155 160 Lys Ser Gly Ala Gln Pro Ala Lys Lys Arg Leu Asn Phe Gly Gln Thr 165 170 175 Gly Asp Thr Glu Ser Val Pro Asp Pro Gln Pro Ile Gly Glu Pro Pro 180 185 190 Ala Ala Pro Ser Gly Val Gly Ser Leu Thr Met Ala Ser Gly Gly Gly 195 200 205 Ala Pro Val Ala Asp Asn Asn Glu Gly Ala Asp Gly Val Gly Ser Ser 210 215 220 Ser Gly Asn Trp His Cys Asp Ser Gln Trp Leu Gly Asp Arg Val Ile 225 230 235 240 Thr Thr Ser Thr Arg Thr Trp Ala Leu Pro Thr Tyr Asn Asn His Leu 245 250 255 Tyr Lys Gln Ile Ser Asn Ser Thr Ser Gly Gly Ser Ser Asn Asp Asn 260 265 270 Ala Tyr Phe Gly Tyr Ser Thr Pro Trp Gly Tyr Phe Asp Phe Asn Arg 275 280 285 Phe His Cys His Phe Ser Pro Arg Asp Trp Gln Arg Leu Ile Asn Asn 290 295 300 Asn Trp Gly Phe Arg Pro Lys Arg Leu Asn Phe Lys Leu Phe Asn Ile 305 310 315 320 Gln Val Lys Glu Val Thr Asp Asn Asn Gly Val Lys Thr Ile Ala Asn 325 330 335 Asn Leu Thr Ser Thr Val Gln Val Phe Thr Asp Ser Asp Tyr Gln Leu 340 345 350 Pro Tyr Val Leu Gly Ser Ala His Glu Gly Cys Leu Pro Pro Phe Pro 355 360 365 Ala Asp Val Phe Met Ile Pro Gln Tyr Gly Tyr Leu Thr Leu Asn Asp 370 375 380 Gly Ser Gln Ala Val Gly Arg Ser Ser Phe Tyr Cys Leu Glu Tyr Phe 385 390 395 400 Pro Ser Gln Met Leu Arg Thr Gly Asn Asn Phe Gln Phe Ser Tyr Glu 405 410 415 Phe Glu Asn Val Pro Phe His Ser Ser Tyr Ala His Ser Gln Ser Leu 420 425 430 Asp Arg Leu Met Asn Pro Leu Ile Asp Gln Tyr Leu Tyr Tyr Leu Ser 435 440 445 Lys Thr Ile Asn Gly Ser Gly Gln Asn Gln Gln Thr Leu Lys Phe Ser 450 455 460 Val Ala Gly Pro Ser Asn Met Ala Val Gln Gly Arg Asn Tyr Ile Pro 465 470 475 480 Gly Pro Ser Tyr Arg Gln Gln Arg Val Ser Thr Thr Val Thr Gln Asn 485 490 495 Asn Asn Ser Glu Phe Ala Trp Pro Gly Ala Ser Ser Trp Ala Leu Asn 500 505 510 Gly Arg Asn Ser Leu Met Asn Pro Gly Pro Ala Met Ala Ser His Lys 515 520 525 Glu Gly Glu Asp Arg Phe Phe Pro Leu Ser Gly Ser Leu Ile Phe Gly 530 535 540 Lys Gln Gly Thr Gly Arg Asp Asn Val Asp Ala Asp Lys Val Met Ile 545 550 555 560 Thr Asn Glu Glu Glu Ile Lys Thr Thr Asn Pro Val Ala Thr Glu Ser 565 570 575 Tyr Gly Gln Val Ala Thr Asn His Gln Ser Ala Gln Ala Gln Ala Gln 580 585 590 Thr Gly Trp Val Gln Asn Gln Gly Ile Leu Pro Gly Met Val Trp Gln 595 600 605 Asp Arg Asp Val Tyr Leu Gln Gly Pro Ile Trp Ala Lys Ile Pro His 610 615 620 Thr Asp Gly Asn Phe His Pro Ser Pro Leu Met Gly Gly Phe Gly Met 625 630 635 640 Lys His Pro Pro Pro Gln Ile Leu Ile Lys Asn Thr Pro Val Pro Ala 645 650 655 Asp Pro Pro Thr Ala Phe Asn Lys Asp Lys Leu Asn Ser Phe Ile Thr 660 665 670 Gln Tyr Ser Thr Gly Gln Val Ser Val Glu Ile Glu Trp Glu Leu Gln 675 680 685 Lys Glu Asn Ser Lys Arg Trp Asn Pro Glu Ile Gln Tyr Thr Ser Asn 690 695 700 Tyr Tyr Lys Ser Asn Asn Val Glu Phe Ala Val Asn Thr Glu Gly Val 705 710 715 720 Tyr Ser Glu Pro Arg Pro Ile Gly Thr Arg Tyr Leu Thr Arg Asn Leu 725 730 735 <210> 47 <211> 738 <212> PRT <213> Artificial Sequence <220> <223> AAVrh10 (Uniprot Q808W5_9VIRU) <400> 47 Met Ala Ala Asp Gly Tyr Leu Pro Asp Trp Leu Glu Asp Asn Leu Ser 1 5 10 15 Glu Gly Ile Arg Glu Trp Trp Asp Leu Lys Pro Gly Ala Pro Lys Pro 20 25 30 Lys Ala Asn Gln Gln Lys Gln Asp Asp Gly Arg Gly Leu Val Leu Pro 35 40 45 Gly Tyr Lys Tyr Leu Gly Pro Phe Asn Gly Leu Asp Lys Gly Glu Pro 50 55 60 Val Asn Ala Ala Asp Ala Ala Ala Leu Glu His Asp Lys Ala Tyr Asp 65 70 75 80 Gln Gln Leu Lys Ala Gly Asp Asn Pro Tyr Leu Arg Tyr Asn His Ala 85 90 95 Asp Ala Glu Phe Gln Glu Arg Leu Gln Glu Asp Thr Ser Phe Gly Gly 100 105 110 Asn Leu Gly Arg Ala Val Phe Gln Ala Lys Lys Arg Val Leu Glu Pro 115 120 125 Leu Gly Leu Val Glu Glu Gly Ala Lys Thr Ala Pro Gly Lys Lys Arg 130 135 140 Pro Val Glu Pro Ser Pro Gln Arg Ser Pro Asp Ser Ser Thr Gly Ile 145 150 155 160 Gly Lys Lys Gly Gln Gln Pro Ala Lys Lys Arg Leu Asn Phe Gly Gln 165 170 175 Thr Gly Asp Ser Glu Ser Val Pro Asp Pro Gln Pro Ile Gly Glu Pro 180 185 190 Pro Ala Gly Pro Ser Gly Leu Gly Ser Gly Thr Met Ala Ala Gly Gly 195 200 205 Gly Ala Pro Met Ala Asp Asn Asn Glu Gly Ala Asp Gly Val Gly Ser 210 215 220 Ser Ser Gly Asn Trp His Cys Asp Ser Thr Trp Leu Gly Asp Arg Val 225 230 235 240 Ile Thr Thr Ser Thr Arg Thr Trp Ala Leu Pro Thr Tyr Asn Asn His 245 250 255 Leu Tyr Lys Gln Ile Ser Asn Gly Thr Ser Gly Gly Ser Thr Asn Asp 260 265 270 Asn Thr Tyr Phe Gly Tyr Ser Thr Pro Trp Gly Tyr Phe Asp Phe Asn 275 280 285 Arg Phe His Cys His Phe Ser Pro Arg Asp Trp Gln Arg Leu Ile Asn 290 295 300 Asn Asn Trp Gly Phe Arg Pro Lys Arg Leu Asn Phe Lys Leu Phe Asn 305 310 315 320 Ile Gln Val Lys Glu Val Thr Gln Asn Glu Gly Thr Lys Thr Ile Ala 325 330 335 Asn Asn Leu Thr Ser Thr Ile Gln Val Phe Thr Asp Ser Glu Tyr Gln 340 345 350 Leu Pro Tyr Val Leu Gly Ser Ala His Gln Gly Cys Leu Pro Pro Phe 355 360 365 Pro Ala Asp Val Phe Met Ile Pro Gln Tyr Gly Tyr Leu Thr Leu Asn 370 375 380 Asn Gly Ser Gln Ala Val Gly Arg Ser Ser Phe Tyr Cys Leu Glu Tyr 385 390 395 400 Phe Pro Ser Gln Met Leu Arg Thr Gly Asn Asn Phe Glu Phe Ser Tyr 405 410 415 Gln Phe Glu Asp Val Pro Phe His Ser Ser Tyr Ala His Ser Gln Ser 420 425 430 Leu Asp Arg Leu Met Asn Pro Leu Ile Asp Gln Tyr Leu Tyr Tyr Leu 435 440 445 Ser Arg Thr Gln Ser Thr Gly Gly Thr Ala Gly Thr Gln Gln Leu Leu 450 455 460 Phe Ser Gln Ala Gly Pro Asn Asn Met Ser Ala Gln Ala Lys Asn Trp 465 470 475 480 Leu Pro Gly Pro Cys Tyr Arg Gln Gln Arg Val Ser Thr Thr Leu Ser 485 490 495 Gln Asn Asn Asn Ser Asn Phe Ala Trp Thr Gly Ala Thr Lys Tyr His 500 505 510 Leu Asn Gly Arg Asp Ser Leu Val Asn Pro Gly Val Ala Met Ala Thr 515 520 525 His Lys Asp Asp Glu Glu Arg Phe Phe Pro Ser Ser Gly Val Leu Met 530 535 540 Phe Gly Lys Gln Gly Ala Gly Lys Asp Asn Val Asp Tyr Ser Ser Val 545 550 555 560 Met Leu Thr Ser Glu Glu Glu Ile Lys Thr Thr Asn Pro Val Ala Thr 565 570 575 Glu Gln Tyr Gly Val Val Ala Asp Asn Leu Gln Gln Gln Asn Ala Ala 580 585 590 Pro Ile Val Gly Ala Val Asn Ser Gln Gly Ala Leu Pro Gly Met Val 595 600 605 Trp Gln Asn Arg Asp Val Tyr Leu Gln Gly Pro Ile Trp Ala Lys Ile 610 615 620 Pro His Thr Asp Gly Asn Phe His Pro Ser Pro Leu Met Gly Gly Phe 625 630 635 640 Gly Leu Lys His Pro Pro Pro Gln Ile Leu Ile Lys Asn Thr Pro Val 645 650 655 Pro Ala Asp Pro Pro Thr Thr Phe Ser Gln Ala Lys Leu Ala Ser Phe 660 665 670 Ile Thr Gln Tyr Ser Thr Gly Gln Val Ser Val Glu Ile Glu Trp Glu 675 680 685 Leu Gln Lys Glu Asn Ser Lys Arg Trp Asn Pro Glu Ile Gln Tyr Thr 690 695 700 Ser Asn Tyr Tyr Lys Ser Thr Asn Val Asp Phe Ala Val Asn Thr Asp 705 710 715 720 Gly Thr Tyr Ser Glu Pro Arg Pro Ile Gly Thr Arg Tyr Leu Thr Arg 725 730 735 Asn Leu <210> 48 <211> 735 <212> PRT <213> Artificial Sequence <220> <223> AAV2v <400> 48 Met Ala Ala Asp Gly Tyr Leu Pro Asp Trp Leu Glu Asp Thr Leu Ser 1 5 10 15 Glu Gly Ile Arg Gln Trp Trp Lys Leu Lys Pro Gly Pro Pro Pro Pro 20 25 30 Lys Pro Ala Glu Arg His Lys Asp Asp Ser Arg Gly Leu Val Leu Pro 35 40 45 Gly Tyr Lys Tyr Leu Gly Pro Phe Asn Gly Leu Asp Lys Gly Glu Pro 50 55 60 Val Asn Glu Ala Asp Ala Ala Ala Leu Glu His Asp Lys Ala Tyr Asp 65 70 75 80 Arg Gln Leu Asp Ser Gly Asp Asn Pro Tyr Leu Lys Tyr Asn His Ala 85 90 95 Asp Ala Glu Phe Gln Glu Arg Leu Lys Glu Asp Thr Ser Phe Gly Gly 100 105 110 Asn Leu Gly Arg Ala Val Phe Gln Ala Lys Lys Arg Val Leu Glu Pro 115 120 125 Leu Gly Leu Val Glu Glu Pro Val Lys Thr Ala Pro Gly Lys Lys Arg 130 135 140 Pro Val Glu His Ser Pro Val Glu Pro Asp Ser Ser Ser Gly Thr Gly 145 150 155 160 Lys Ala Gly Gln Gln Pro Ala Arg Lys Arg Leu Asn Phe Gly Gln Thr 165 170 175 Gly Asp Ala Asp Ser Val Pro Asp Pro Gln Pro Leu Gly Gln Pro Pro 180 185 190 Ala Ala Pro Ser Gly Leu Gly Thr Asn Thr Met Ala Thr Gly Ser Gly 195 200 205 Ala Pro Met Ala Asp Asn Asn Glu Gly Ala Asp Gly Val Gly Asn Ser 210 215 220 Ser Gly Asn Trp His Cys Asp Ser Thr Trp Met Gly Asp Arg Val Ile 225 230 235 240 Thr Thr Ser Thr Arg Thr Trp Ala Leu Pro Thr Tyr Asn Asn His Leu 245 250 255 Tyr Lys Gln Ile Ser Ser Gln Ser Gly Ala Ser Asn Asp Asn His Tyr 260 265 270 Phe Gly Tyr Ser Thr Pro Trp Gly Tyr Phe Asp Phe Asn Arg Phe His 275 280 285 Cys His Phe Ser Pro Arg Asp Trp Gln Arg Leu Ile Asn Asn Asn Trp 290 295 300 Gly Phe Arg Pro Lys Arg Leu Asn Phe Lys Leu Phe Asn Ile Gln Val 305 310 315 320 Lys Glu Val Thr Gln Asn Asp Gly Thr Thr Thr Ile Ala Asn Asn Leu 325 330 335 Thr Ser Thr Val Gln Val Phe Thr Asp Ser Glu Tyr Gln Leu Pro Tyr 340 345 350 Val Leu Gly Ser Ala His Gln Gly Cys Leu Pro Pro Phe Pro Ala Asp 355 360 365 Val Phe Met Val Pro Gln Tyr Gly Tyr Leu Thr Leu Asn Asn Gly Ser 370 375 380 Gln Ala Val Gly Arg Ser Ser Phe Tyr Cys Leu Glu Tyr Phe Pro Ser 385 390 395 400 Gln Met Leu Arg Thr Gly Asn Asn Phe Thr Phe Ser Tyr Thr Phe Glu 405 410 415 Asp Val Pro Phe His Ser Ser Tyr Ala His Ser Gln Ser Leu Asp Arg 420 425 430 Leu Met Asn Pro Leu Ile Asp Gln Tyr Leu Tyr Phe Leu Ser Arg Thr 435 440 445 Asn Thr Pro Ser Gly Thr Thr Thr Gln Ser Arg Leu Gln Phe Ser Gln 450 455 460 Ala Gly Ala Ser Asp Ile Arg Asp Gln Ser Arg Asn Trp Leu Pro Gly 465 470 475 480 Pro Cys Tyr Arg Gln Gln Gly Val Ser Lys Val Ser Ala Asp Asn Asn 485 490 495 Asn Ser Glu Phe Ser Trp Thr Gly Ala Thr Lys Tyr His Leu Asn Gly 500 505 510 Arg Asp Ser Leu Val Asn Pro Gly Pro Ala Met Ala Ser His Lys Asp 515 520 525 Asp Glu Glu Lys Phe Phe Pro Gln Ser Gly Val Leu Ile Phe Gly Lys 530 535 540 Gln Gly Ser Glu Lys Thr Asn Val Asp Ile Glu Lys Val Met Ile Thr 545 550 555 560 Asp Glu Glu Glu Ile Arg Thr Thr Asn Pro Val Ala Thr Glu Gln Tyr 565 570 575 Gly Ser Val Ser Thr Asn Leu Gln Ser Gly Asn Thr Gln Ala Ala Thr 580 585 590 Ala Asp Val Asn Thr Gln Gly Val Leu Pro Gly Met Val Trp Gln Asp 595 600 605 Arg Asp Val Tyr Leu Gln Gly Pro Ile Trp Ala Lys Ile Pro His Thr 610 615 620 Asp Gly His Phe His Pro Ser Pro Leu Met Gly Gly Phe Gly Leu Lys 625 630 635 640 His Pro Pro Pro Gln Ile Leu Ile Lys Asn Thr Pro Val Pro Ala Asn 645 650 655 Pro Ser Thr Thr Phe Ser Ala Ala Lys Phe Ala Ser Phe Ile Thr Gln 660 665 670 Tyr Ser Thr Gly Gln Val Ser Val Glu Ile Glu Trp Glu Leu Gln Lys 675 680 685 Glu Asn Ser Lys Arg Trp Asn Pro Glu Ile Gln Tyr Thr Ser Asn Tyr 690 695 700 Asn Lys Ser Val Asn Val Asp Phe Thr Val Asp Thr Asn Gly Val Tyr 705 710 715 720 Ser Glu Pro Arg Pro Ile Gly Thr Arg Phe Leu Thr Arg Asn Leu 725 730 735 <210> 49 <211> 736 <212> PRT <213> Artificial Sequence <220> <223> AAV44-9 <400> 49 Met Ala Ala Asp Gly Tyr Leu Pro Asp Trp Leu Glu Asp Asn Leu Ser 1 5 10 15 Glu Gly Ile Arg Glu Trp Trp Asp Leu Lys Pro Gly Ala Pro Lys Pro 20 25 30 Lys Ala Asn Gln Gln Lys Gln Asp Asp Gly Arg Gly Leu Val Leu Pro 35 40 45 Gly Tyr Lys Tyr Leu Gly Pro Phe Asn Gly Leu Asp Lys Gly Glu Pro 50 55 60 Val Asn Ala Ala Asp Ala Ala Ala Leu Glu His Asp Lys Ala Tyr Asp 65 70 75 80 Gln Gln Leu Lys Ala Gly Asp Asn Pro Tyr Leu Arg Tyr Asn His Ala 85 90 95 Asp Ala Glu Phe Gln Glu Arg Leu Gln Glu Asp Thr Ser Phe Gly Gly 100 105 110 Asn Leu Gly Arg Ala Val Phe Gln Ala Lys Lys Arg Val Leu Glu Pro 115 120 125 Leu Gly Leu Val Glu Glu Gly Ala Lys Thr Ala Pro Gly Lys Lys Arg 130 135 140 Pro Val Glu Gln Ser Pro Gln Glu Pro Asp Ser Ser Ser Gly Ile Gly 145 150 155 160 Lys Thr Gly Gln Gln Pro Ala Lys Lys Arg Leu Asn Phe Gly Gln Thr 165 170 175 Gly Asp Thr Glu Ser Val Pro Asp Pro Gln Pro Leu Gly Glu Pro Pro 180 185 190 Ala Ala Pro Ser Gly Leu Gly Pro Asn Thr Met Ala Ser Gly Gly Gly 195 200 205 Ala Pro Met Ala Asp Asn Asn Glu Gly Ala Asp Gly Val Gly Asn Ser 210 215 220 Ser Gly Asn Trp His Cys Asp Ser Thr Trp Leu Gly Asp Arg Val Ile 225 230 235 240 Thr Thr Ser Thr Arg Thr Trp Ala Leu Pro Thr Tyr Asn Asn His Leu 245 250 255 Tyr Lys Gln Ile Ser Asn Gly Thr Ser Gly Gly Ser Thr Asn Asp Asn 260 265 270 Thr Tyr Phe Gly Tyr Ser Thr Pro Trp Gly Tyr Phe Asp Phe Asn Arg 275 280 285 Phe His Cys His Phe Ser Pro Arg Asp Trp Gln Arg Leu Ile Asn Asn 290 295 300 Asn Trp Gly Phe Arg Pro Lys Arg Leu Asn Phe Lys Leu Phe Asn Ile 305 310 315 320 Gln Val Lys Glu Val Thr Thr Asn Glu Gly Thr Lys Thr Ile Ala Asn 325 330 335 Asn Leu Thr Ser Thr Val Gln Val Phe Thr Asp Ser Glu Tyr Gln Leu 340 345 350 Pro Tyr Val Leu Gly Ser Ala His Gln Gly Cys Leu Pro Pro Phe Pro 355 360 365 Ala Asp Val Phe Met Val Pro Gln Tyr Gly Tyr Leu Thr Leu Asn Asn 370 375 380 Gly Ser Gln Ala Leu Gly Arg Ser Ser Phe Tyr Cys Leu Glu Tyr Phe 385 390 395 400 Pro Ser Gln Met Leu Arg Thr Gly Asn Asn Phe Gln Phe Ser Tyr Thr 405 410 415 Phe Glu Asp Val Pro Phe His Ser Ser Tyr Ala His Ser Gln Ser Leu 420 425 430 Asp Arg Leu Met Asn Pro Leu Ile Asp Gln Tyr Leu Tyr Tyr Leu Val 435 440 445 Arg Thr Gln Thr Thr Gly Thr Gly Gly Thr Gln Thr Leu Ala Phe Ser 450 455 460 Gln Ala Gly Pro Ser Asn Met Ala Ser Gln Ala Arg Asn Trp Val Pro 465 470 475 480 Gly Pro Ser Tyr Arg Gln Gln Arg Val Ser Thr Thr Thr Asn Gln Asn 485 490 495 Asn Asn Ser Asn Phe Ala Trp Thr Gly Ala Ala Lys Phe Lys Leu Asn 500 505 510 Gly Arg Asp Ser Leu Met Asn Pro Gly Val Ala Met Ala Ser His Lys 515 520 525 Asp Asp Glu Asp Arg Phe Phe Pro Ser Ser Gly Val Leu Ile Phe Gly 530 535 540 Lys Gln Gly Ala Gly Asn Asp Gly Val Asp Tyr Ser Gln Val Leu Ile 545 550 555 560 Thr Asp Glu Glu Glu Ile Lys Ala Thr Asn Pro Val Ala Thr Glu Glu 565 570 575 Tyr Gly Ala Val Ala Ile Asn Asn Gln Ala Ala Asn Thr Gln Ala Gln 580 585 590 Thr Gly Leu Val His Asn Gln Gly Val Ile Pro Gly Met Val Trp Gln 595 600 605 Asn Arg Asp Val Tyr Leu Gln Gly Pro Ile Trp Ala Lys Ile Pro His 610 615 620 Thr Asp Gly Asn Phe His Pro Ser Pro Leu Met Gly Gly Phe Gly Leu 625 630 635 640 Lys His Pro Pro Pro Gln Ile Leu Ile Lys Asn Thr Pro Val Pro Ala 645 650 655 Asp Pro Pro Leu Thr Phe Asn Gln Ala Lys Leu Asn Ser Phe Ile Thr 660 665 670 Gln Tyr Ser Thr Gly Gln Val Ser Val Glu Ile Glu Trp Glu Leu Gln 675 680 685 Lys Glu Asn Ser Lys Arg Trp Asn Pro Glu Ile Gln Tyr Thr Ser Asn 690 695 700 Tyr Tyr Lys Ser Thr Asn Val Asp Phe Ala Val Asn Thr Glu Gly Val 705 710 715 720 Tyr Ser Glu Pro Arg Pro Ile Gly Thr Arg Tyr Leu Thr Arg Asn Leu 725 730 735 <210> 50 <211> 228 <212> PRT <213> Artificial Sequence <220> <223> Exemplary IgG-Fc(2) <400> 50 Lys Asp Lys Thr His Thr Cys Pro Pro Cys Pro Ala Pro Glu Leu Leu 1 5 10 15 Gly Gly Pro Ser Val Phe Leu Phe Pro Pro Lys Pro Lys Asp Thr Leu 20 25 30 Met Ile Ser Arg Thr Pro Glu Val Thr Cys Val Val Val Asp Val Ser 35 40 45 His Glu Asp Pro Glu Val Lys Phe Asn Trp Tyr Val Asp Gly Val Glu 50 55 60 Val His Asn Ala Lys Thr Lys Pro Arg Glu Glu Gln Tyr Asn Ser Thr 65 70 75 80 Tyr Arg Val Val Ser Val Leu Thr Val Leu His Gln Asp Trp Leu Asn 85 90 95 Gly Lys Glu Tyr Lys Cys Lys Val Ser Asn Lys Ala Leu Pro Ala Pro 100 105 110 Ile Glu Lys Thr Ile Ser Lys Ala Lys Gly Gln Pro Arg Glu Pro Gln 115 120 125 Val Tyr Thr Leu Pro Pro Ser Arg Asp Glu Leu Thr Lys Asn Gln Val 130 135 140 Ser Leu Thr Cys Leu Val Lys Gly Phe Tyr Pro Ser Asp Ile Ala Val 145 150 155 160 Glu Trp Glu Ser Asn Gly Gln Pro Glu Asn Asn Tyr Lys Thr Thr Pro 165 170 175 Pro Val Leu Asp Ser Asp Gly Ser Phe Phe Leu Tyr Ser Lys Leu Thr 180 185 190 Val Asp Lys Ser Arg Trp Gln Gln Gly Asn Val Phe Ser Cys Ser Val 195 200 205 Met His Glu Ala Leu His Asn His Tyr Thr Gln Lys Ser Leu Ser Leu 210 215 220 Ser Pro Gly Lys 225 <210> 51 <211> 20 <212> PRT <213> Artificial Sequence <220> <223> L1 <400> 51 Lys Phe Asn Pro Leu Asp Glu Leu Glu Glu Thr Leu Tyr Glu Gln Phe 1 5 10 15 Thr Phe Gln Gln 20 <210> 52 <211> 57 <212> PRT <213> Artificial sequence <220> <223> Exemplary ABD(2)(2xL1) <400> 52 Gly Gly Gly Gly Gly Ala Gln Lys Phe Asn Pro Leu Asp Glu Leu Glu 1 5 10 15 Glu Thr Leu Tyr Glu Gln Phe Thr Phe Gln Gln Gly Gly Gly Gly Gly 20 25 30 Gly Gly Gly Lys Phe Asn Pro Leu Asp Glu Leu Glu Glu Thr Leu Tyr 35 40 45 Glu Gln Phe Thr Phe Gln Gln Leu Glu 50 55 <210> 53 <211> 51 <212> PRT <213> Artificial sequence <220> <223> Exemplary ABD(3)(Con4-L1) <400> 53 Gly Gly Gly Gly Gly Ala Gln Gln Glu Glu Cys Glu Trp Asp Pro Trp 1 5 10 15 Thr Cys Glu His Met Gly Gly Gly Gly Gly Gly Gly Gly Lys Phe Asn 20 25 30 Pro Leu Asp Glu Leu Glu Glu Thr Leu Tyr Glu Gln Phe Thr Phe Gln 35 40 45 Gln Leu Glu 50 <210> 54 <211> 70 <212> PRT <213> Artificial Sequence <220> <223> Exemplary ABD(4)(L1 - Con4) <400> 54 Gly Gly Gly Gly Gly Ala Gln Gly Ser Gly Ser Ala Thr Gly Gly Ser 1 5 10 15 Gly Ser Thr Ala Ser Ser Gly Ser Gly Ser Ala Thr His Lys Phe Asn 20 25 30 Pro Leu Asp Glu Leu Glu Glu Thr Leu Tyr Glu Gln Phe Thr Phe Gln 35 40 45 Gln Gly Gly Gly Gly Gly Gln Glu Glu Cys Glu Trp Asp Pro Trp Thr 50 55 60 Cys Glu His Met Leu Glu 65 70 <210> 55 <211> 105 <212> PRT <213> Artificial Sequence <220> <223> Exemplary Trap C(1) Derived from VEGFR - 2 <400> 55 Val Tyr Val Gln Asp Tyr Arg Ser Pro Phe Ile Ala Ser Val Ser Asp 1 5 10 15 Gln His Gly Val Val Tyr Ile Thr Glu Asn Lys Asn Lys Thr Val Val 20 25 30 Ile Pro Cys Leu Gly Ser Ile Ser Asn Leu Asn Val Ser Leu Cys Ala 35 40 45 Arg Tyr Pro Glu Lys Arg Phe Val Pro Asp Gly Asn Arg Ile Ser Trp 50 55 60 Asp Ser Lys Lys Gly Phe Thr Ile Pro Ser Tyr Met Ile Ser Tyr Ala 65 70 75 80 Gly Met Val Phe Cys Glu Ala Lys Ile Asn Asp Glu Ser Tyr Gln Ser 85 90 95 Ile Met Tyr Ile Val Val Val Val Gly 100 105 <210> 56 <211> 202 <212> PRT <213> Artificial Sequence <220> <223> Exemplary Trap C(2) Derived from VEGFR-3 <400> 56 Tyr Ser Met Thr Pro Pro Thr Leu Asn Ile Thr Glu Glu Ser His Val 1 5 10 15 Ile Asp Thr Gly Asp Ser Leu Ser Ile Ser Cys Arg Gly Gln His Pro 20 25 30 Leu Glu Trp Ala Trp Pro Gly Ala Gln Glu Ala Pro Ala Thr Gly Asp 35 40 45 Lys Asp Ser Glu Asp Thr Gly Val Val Arg Asp Cys Glu Gly Thr Asp 50 55 60 Ala Arg Pro Tyr Cys Lys Val Leu Leu Leu His Glu Val His Ala Gln 65 70 75 80 Asp Thr Gly Ser Tyr Val Cys Tyr Tyr Lys Tyr Ile Lys Ala Arg Ile 85 90 95 Glu Gly Thr Thr Ala Ala Ser Ser Tyr Val Phe Val Arg Asp Phe Glu 100 105 110 Gln Pro Phe Ile Asn Lys Pro Asp Thr Leu Leu Val Asn Arg Lys Asp 115 120 125 Ala Met Trp Val Pro Cys Leu Val Ser Ile Pro Gly Leu Asn Val Thr 130 135 140 Leu Arg Ser Gln Ser Ser Val Leu Trp Pro Asp Gly Gln Glu Val Val 145 150 155 160 Trp Asp Asp Arg Arg Gly Met Leu Val Ser Thr Pro Leu Leu His Asp 165 170 175 Ala Leu Tyr Leu Gln Cys Glu Thr Thr Trp Gly Asp Gln Asp Phe Leu 180 185 190 Ser Asn Pro Phe Leu Val His Ile Thr Gly 195 200 <210> 57 <211> 202 <212> PRT <213> Artificial sequence <220> <223> Exemplary Trap C (3) derived from VEGFR-3 <400> 57 Tyr Ser Met Thr Pro Pro Thr Leu Asn Ile Thr Glu Glu Ser His Val 1 5 10 15 Ile Asp Thr Gly Asp Ser Leu Ser Ile Ser Cys Arg Gly Gln His Pro 20 25 30 Leu Glu Trp Ala Trp Pro Gly Ala Gln Glu Ala Pro Ala Thr Gly Asp 35 40 45 Lys Asp Ser Glu Asp Thr Gly Val Val Arg Asp Cys Glu Gly Thr Asp 50 55 60 Ala Arg Pro Tyr Cys Lys Val Leu Leu Leu His Glu Val His Ala Asn 65 70 75 80 Asp Thr Gly Ser Tyr Val Cys Tyr Tyr Lys Tyr Ile Lys Ala Arg Ile 85 90 95 Glu Gly Thr Thr Ala Ala Ser Ser Tyr Val Phe Val Arg Asp Phe Glu 100 105 110 Gln Pro Phe Ile Asn Lys Pro Asp Thr Leu Leu Val Asn Arg Lys Asp 115 120 125 Ala Met Trp Val Pro Cys Leu Val Ser Ile Pro Gly Leu Asn Val Thr 130 135 140 Leu Arg Ser Gln Ser Ser Val Leu Trp Pro Asp Gly Gln Glu Val Val 145 150 155 160 Trp Asp Asp Arg Arg Gly Met Leu Val Ser Thr Pro Leu Leu His Asp 165 170 175 Ala Leu Tyr Leu Gln Cys Glu Thr Thr Trp Gly Asp Gln Asp Phe Leu 180 185 190 Ser Asn Pro Phe Leu Val His Ile Thr Gly 195 200 <210> 58 <211> 634 <212> PRT <213> Artificial Sequence <220> <223> Exemplary Polypeptide 5 (such as in EXG102-24) <400> 58 Met Val Ser Tyr Trp Asp Thr Gly Val Leu Leu Cys Ala Leu Leu Ser 1 5 10 15 Cys Leu Leu Leu Thr Gly Ser Ser Ser Gly Gly Gly Gly Gly Gly Ala 20 25 30 Gln Gln Glu Glu Cys Glu Trp Asp Pro Trp Thr Cys Glu His Met Gly 35 40 45 Ser Gly Ser Ala Thr Gly Gly Ser Gly Ser Thr Ala Ser Ser Gly Ser 50 55 60 Gly Ser Ala Thr His Gln Glu Glu Cys Glu Trp Asp Pro Trp Thr Cys 65 70 75 80 Glu His Met Leu Glu Gly Gly Gly Gly Gly Ser Ser Asp Thr Gly Arg 85 90 95 Pro Phe Val Glu Met Tyr Ser Glu Ile Pro Glu Ile Ile His Met Thr 100 105 110 Glu Gly Arg Glu Leu Val Ile Pro Cys Arg Val Thr Ser Pro Asn Ile 115 120 125 Thr Val Thr Leu Lys Lys Phe Pro Leu Asp Thr Leu Ile Pro Asp Gly 130 135 140 Lys Arg Ile Ile Trp Asp Ser Arg Lys Gly Phe Ile Ile Ser Asn Ala 145 150 155 160 Thr Tyr Lys Glu Ile Gly Leu Leu Thr Cys Glu Ala Thr Val Asn Gly 165 170 175 His Leu Tyr Lys Thr Asn Tyr Leu Thr His Arg Gln Thr Asn Thr Ile 180 185 190 Ile Asp Val Val Leu Ser Pro Ser His Gly Ile Glu Leu Ser Val Gly 195 200 205 Glu Lys Leu Val Leu Asn Cys Thr Ala Arg Thr Glu Leu Asn Val Gly 210 215 220 Ile Asp Phe Asn Trp Glu Tyr Pro Ser Ser Lys His Gln His Lys Lys 225 230 235 240 Leu Val Asn Arg Asp Leu Lys Thr Gln Ser Gly Ser Glu Met Lys Lys 245 250 255 Phe Leu Ser Thr Leu Thr Ile Asp Gly Val Thr Arg Ser Asp Gln Gly 260 265 270 Leu Tyr Thr Cys Ala Ala Ser Ser Gly Leu Met Thr Lys Lys Asn Ser 275 280 285 Thr Phe Val Arg Val His Glu Lys Gly Gly Gly Gly Gly Ser Val Tyr 290 295 300 Val Gln Asp Tyr Arg Ser Pro Phe Ile Ala Ser Val Ser Asp Gln His 305 310 315 320 Gly Val Val Tyr Ile Thr Glu Asn Lys Asn Lys Thr Val Val Ile Pro 325 330 335 Cys Leu Gly Ser Ile Ser Asn Leu Asn Val Ser Leu Cys Ala Arg Tyr 340 345 350 Pro Glu Lys Arg Phe Val Pro Asp Gly Asn Arg Ile Ser Trp Asp Ser 355 360 365 Lys Lys Gly Phe Thr Ile Pro Ser Tyr Met Ile Ser Tyr Ala Gly Met 370 375 380 Val Phe Cys Glu Ala Lys Ile Asn Asp Glu Ser Tyr Gln Ser Ile Met 385 390 395 400 Tyr Ile Val Val Val Val Gly Asp Lys Thr His Thr Cys Pro Pro Cys 405 410 415 Pro Ala Pro Glu Leu Leu Gly Gly Pro Ser Val Phe Leu Phe Pro Pro 420 425 430 Lys Pro Lys Asp Thr Leu Met Ile Ser Arg Thr Pro Glu Val Thr Cys 435 440 445 Val Val Val Asp Val Ser His Glu Asp Pro Glu Val Lys Phe Asn Trp 450 455 460 Tyr Val Asp Gly Val Glu Val His Asn Ala Lys Thr Lys Pro Arg Glu 465 470 475 480 Glu Gln Tyr Asn Ser Thr Tyr Arg Val Val Ser Val Leu Thr Val Leu 485 490 495 His Gln Asp Trp Leu Asn Gly Lys Glu Tyr Lys Cys Lys Val Ser Asn 500 505 510 Lys Ala Leu Pro Ala Pro Ile Glu Lys Thr Ile Ser Lys Ala Lys Gly 515 520 525 Gln Pro Arg Glu Pro Gln Val Tyr Thr Leu Pro Pro Ser Arg Asp Glu 530 535 540 Leu Thr Lys Asn Gln Val Ser Leu Thr Cys Leu Val Lys Gly Phe Tyr 545 550 555 560 Pro Ser Asp Ile Ala Val Glu Trp Glu Ser Asn Gly Gln Pro Glu Asn 565 570 575 Asn Tyr Lys Thr Thr Pro Pro Val Leu Asp Ser Asp Gly Ser Phe Phe 580 585 590 Leu Tyr Ser Lys Leu Thr Val Asp Lys Ser Arg Trp Gln Gln Gly Asn 595 600 605 Val Phe Ser Cys Ser Val Met His Glu Ala Leu His Asn His Tyr Thr 610 615 620 Gln Lys Ser Leu Ser Leu Ser Pro Gly Lys 625 630 <210> 59 <211> 628 <212> PRT <213> Artificial Sequence <220> <223> Exemplary Polypeptide 6 (such as in EXG102-25) <400> 59 Met Val Ser Tyr Trp Asp Thr Gly Val Leu Leu Cys Ala Leu Leu Ser 1 5 10 15 Cys Leu Leu Leu Thr Gly Ser Ser Ser Gly Ser Asp Thr Gly Arg Pro 20 25 30 Phe Val Glu Met Tyr Ser Glu Ile Pro Glu Ile Ile His Met Thr Glu 35 40 45 Gly Arg Glu Leu Val Ile Pro Cys Arg Val Thr Ser Pro Asn Ile Thr 50 55 60 Val Thr Leu Lys Lys Phe Pro Leu Asp Thr Leu Ile Pro Asp Gly Lys 65 70 75 80 Arg Ile Ile Trp Asp Ser Arg Lys Gly Phe Ile Ile Ser Asn Ala Thr 85 90 95 Tyr Lys Glu Ile Gly Leu Leu Thr Cys Glu Ala Thr Val Asn Gly His 100 105 110 Leu Tyr Lys Thr Asn Tyr Leu Thr His Arg Gln Thr Asn Thr Ile Ile 115 120 125 Asp Val Val Leu Ser Pro Ser His Gly Ile Glu Leu Ser Val Gly Glu 130 135 140 Lys Leu Val Leu Asn Cys Thr Ala Arg Thr Glu Leu Asn Val Gly Ile 145 150 155 160 Asp Phe Asn Trp Glu Tyr Pro Ser Ser Lys His Gln His Lys Lys Leu 165 170 175 Val Asn Arg Asp Leu Lys Thr Gln Ser Gly Ser Glu Met Lys Lys Phe 180 185 190 Leu Ser Thr Leu Thr Ile Asp Gly Val Thr Arg Ser Asp Gln Gly Leu 195 200 205 Tyr Thr Cys Ala Ala Ser Ser Gly Leu Met Thr Lys Lys Asn Ser Thr 210 215 220 Phe Val Arg Val His Glu Lys Gly Gly Gly Gly Gly Ser Val Tyr Val 225 230 235 240 Gln Asp Tyr Arg Ser Pro Phe Ile Ala Ser Val Ser Asp Gln His Gly 245 250 255 Val Val Tyr Ile Thr Glu Asn Lys Asn Lys Thr Val Val Ile Pro Cys 260 265 270 Leu Gly Ser Ile Ser Asn Leu Asn Val Ser Leu Cys Ala Arg Tyr Pro 275 280 285 Glu Lys Arg Phe Val Pro Asp Gly Asn Arg Ile Ser Trp Asp Ser Lys 290 295 300 Lys Gly Phe Thr Ile Pro Ser Tyr Met Ile Ser Tyr Ala Gly Met Val 305 310 315 320 Phe Cys Glu Ala Lys Ile Asn Asp Glu Ser Tyr Gln Ser Ile Met Tyr 325 330 335 Ile Val Val Val Val Gly Asp Lys Thr His Thr Cys Pro Pro Cys Pro 340 345 350 Ala Pro Glu Leu Leu Gly Gly Pro Ser Val Phe Leu Phe Pro Pro Lys 355 360 365 Pro Lys Asp Thr Leu Met Ile Ser Arg Thr Pro Glu Val Thr Cys Val 370 375 380 Val Val Asp Val Ser His Glu Asp Pro Glu Val Lys Phe Asn Trp Tyr 385 390 395 400 Val Asp Gly Val Glu Val His Asn Ala Lys Thr Lys Pro Arg Glu Glu 405 410 415 Gln Tyr Asn Ser Thr Tyr Arg Val Val Ser Val Leu Thr Val Leu His 420 425 430 Gln Asp Trp Leu Asn Gly Lys Glu Tyr Lys Cys Lys Val Ser Asn Lys 435 440 445 Ala Leu Pro Ala Pro Ile Glu Lys Thr Ile Ser Lys Ala Lys Gly Gln 450 455 460 Pro Arg Glu Pro Gln Val Tyr Thr Leu Pro Pro Ser Arg Asp Glu Leu 465 470 475 480 Thr Lys Asn Gln Val Ser Leu Thr Cys Leu Val Lys Gly Phe Tyr Pro 485 490 495 Ser Asp Ile Ala Val Glu Trp Glu Ser Asn Gly Gln Pro Glu Asn Asn 500 505 510 Tyr Lys Thr Thr Pro Pro Val Leu Asp Ser Asp Gly Ser Phe Phe Leu 515 520 525 Tyr Ser Lys Leu Thr Val Asp Lys Ser Arg Trp Gln Gln Gly Asn Val 530 535 540 Phe Ser Cys Ser Val Met His Glu Ala Leu His Asn His Tyr Thr Gln 545 550 555 560 Lys Ser Leu Ser Leu Ser Pro Gly Lys Gly Gly Gly Gly Gly Ala Gln 565 570 575 Gln Glu Glu Cys Glu Trp Asp Pro Trp Thr Cys Glu His Met Gly Ser 580 585 590 Gly Ser Ala Thr Gly Gly Ser Gly Ser Thr Ala Ser Ser Gly Ser Gly 595 600 605 Ser Ala Thr His Gln Glu Glu Cys Glu Trp Asp Pro Trp Thr Cys Glu 610 615 620 His Met Leu Glu 625 <210> 60 <211> 639 <212> PRT <213> Artificial Sequence <220> <223> Exemplary Polypeptide 7 (such as in EXG102 - 26) <400> 60 Met Val Ser Tyr Trp Asp Thr Gly Val Leu Leu Cys Ala Leu Leu Ser 1 5 10 15 Cys Leu Leu Leu Thr Gly Ser Ser Ser Gly Ser Asp Thr Gly Arg Pro 20 25 30 Phe Val Glu Met Tyr Ser Glu Ile Pro Glu Ile Ile His Met Thr Glu 35 40 45 Gly Arg Glu Leu Val Ile Pro Cys Arg Val Thr Ser Pro Asn Ile Thr 50 55 60 Val Thr Leu Lys Lys Phe Pro Leu Asp Thr Leu Ile Pro Asp Gly Lys 65 70 75 80 Arg Ile Ile Trp Asp Ser Arg Lys Gly Phe Ile Ile Ser Asn Ala Thr 85 90 95 Tyr Lys Glu Ile Gly Leu Leu Thr Cys Glu Ala Thr Val Asn Gly His 100 105 110 Leu Tyr Lys Thr Asn Tyr Leu Thr His Arg Gln Thr Asn Thr Ile Ile 115 120 125 Asp Val Val Leu Ser Pro Ser His Gly Ile Glu Leu Ser Val Gly Glu 130 135 140 Lys Leu Val Leu Asn Cys Thr Ala Arg Thr Glu Leu Asn Val Gly Ile 145 150 155 160 Asp Phe Asn Trp Glu Tyr Pro Ser Ser Lys His Gln His Lys Lys Leu 165 170 175 Val Asn Arg Asp Leu Lys Thr Gln Ser Gly Ser Glu Met Lys Lys Phe 180 185 190 Leu Ser Thr Leu Thr Ile Asp Gly Val Thr Arg Ser Asp Gln Gly Leu 195 200 205 Tyr Thr Cys Ala Ala Ser Ser Gly Leu Met Thr Lys Lys Asn Ser Thr 210 215 220 Phe Val Arg Val His Glu Lys Gly Gly Gly Gly Gly Ser Val Tyr Val 225 230 235 240 Gln Asp Tyr Arg Ser Pro Phe Ile Ala Ser Val Ser Asp Gln His Gly 245 250 255 Val Val Tyr Ile Thr Glu Asn Lys Asn Lys Thr Val Val Ile Pro Cys 260 265 270 Leu Gly Ser Ile Ser Asn Leu Asn Val Ser Leu Cys Ala Arg Tyr Pro 275 280 285 Glu Lys Arg Phe Val Pro Asp Gly Asn Arg Ile Ser Trp Asp Ser Lys 290 295 300 Lys Gly Phe Thr Ile Pro Ser Tyr Met Ile Ser Tyr Ala Gly Met Val 305 310 315 320 Phe Cys Glu Ala Lys Ile Asn Asp Glu Ser Tyr Gln Ser Ile Met Tyr 325 330 335 Ile Val Val Val Val Gly Asp Lys Thr His Thr Cys Pro Pro Cys Pro 340 345 350 Ala Pro Glu Leu Leu Gly Gly Pro Ser Val Phe Leu Phe Pro Pro Lys 355 360 365 Pro Lys Asp Thr Leu Met Ile Ser Arg Thr Pro Glu Val Thr Cys Val 370 375 380 Val Val Asp Val Ser His Glu Asp Pro Glu Val Lys Phe Asn Trp Tyr 385 390 395 400 Val Asp Gly Val Glu Val His Asn Ala Lys Thr Lys Pro Arg Glu Glu 405 410 415 Gln Tyr Asn Ser Thr Tyr Arg Val Val Ser Val Leu Thr Val Leu His 420 425 430 Gln Asp Trp Leu Asn Gly Lys Glu Tyr Lys Cys Lys Val Ser Asn Lys 435 440 445 Ala Leu Pro Ala Pro Ile Glu Lys Thr Ile Ser Lys Ala Lys Gly Gln 450 455 460 Pro Arg Glu Pro Gln Val Tyr Thr Leu Pro Pro Ser Arg Asp Glu Leu 465 470 475 480 Thr Lys Asn Gln Val Ser Leu Thr Cys Leu Val Lys Gly Phe Tyr Pro 485 490 495 Ser Asp Ile Ala Val Glu Trp Glu Ser Asn Gly Gln Pro Glu Asn Asn 500 505 510 Tyr Lys Thr Thr Pro Pro Val Leu Asp Ser Asp Gly Ser Phe Phe Leu 515 520 525 Tyr Ser Lys Leu Thr Val Asp Lys Ser Arg Trp Gln Gln Gly Asn Val 530 535 540 Phe Ser Cys Ser Val Met His Glu Ala Leu His Asn His Tyr Thr Gln 545 550 555 560 Lys Ser Leu Ser Leu Ser Pro Gly Lys Gly Gly Gly Gly Gly Ala Gln 565 570 575 Gly Ser Gly Ser Ala Thr Gly Gly Ser Gly Ser Thr Ala Ser Ser Gly 580 585 590 Ser Gly Ser Ala Thr His Lys Phe Asn Pro Leu Asp Glu Leu Glu Glu 595 600 605 Thr Leu Tyr Glu Gln Phe Thr Phe Gln Gln Gly Gly Gly Gly Gly Gln 610 615 620 Glu Glu Cys Glu Trp Asp Pro Trp Thr Cys Glu His Met Leu Glu 625 630 635 <210> 61 <211> 731 <212> PRT <213> Artificial Sequence <220> <223> Exemplary Polypeptide 8 (such as in EXG102 - 27) <400> 61 Met Val Ser Tyr Trp Asp Thr Gly Val Leu Leu Cys Ala Leu Leu Ser 1 5 10 15 Cys Leu Leu Leu Thr Gly Ser Ser Ser Gly Gly Gly Gly Gly Gly Ala 20 25 30 Gln Gln Glu Glu Cys Glu Trp Asp Pro Trp Thr Cys Glu His Met Gly 35 40 45 Ser Gly Ser Ala Thr Gly Gly Ser Gly Ser Thr Ala Ser Ser Gly Ser 50 55 60 Gly Ser Ala Thr His Gln Glu Glu Cys Glu Trp Asp Pro Trp Thr Cys 65 70 75 80 Glu His Met Leu Glu Gly Gly Gly Gly Gly Ser Ser Asp Thr Gly Arg 85 90 95 Pro Phe Val Glu Met Tyr Ser Glu Ile Pro Glu Ile Ile His Met Thr 100 105 110 Glu Gly Arg Glu Leu Val Ile Pro Cys Arg Val Thr Ser Pro Asn Ile 115 120 125 Thr Val Thr Leu Lys Lys Phe Pro Leu Asp Thr Leu Ile Pro Asp Gly 130 135 140 Lys Arg Ile Ile Trp Asp Ser Arg Lys Gly Phe Ile Ile Ser Asn Ala 145 150 155 160 Thr Tyr Lys Glu Ile Gly Leu Leu Thr Cys Glu Ala Thr Val Asn Gly 165 170 175 His Leu Tyr Lys Thr Asn Tyr Leu Thr His Arg Gln Thr Asn Thr Ile 180 185 190 Ile Asp Val Val Leu Ser Pro Ser His Gly Ile Glu Leu Ser Val Gly 195 200 205 Glu Lys Leu Val Leu Asn Cys Thr Ala Arg Thr Glu Leu Asn Val Gly 210 215 220 Ile Asp Phe Asn Trp Glu Tyr Pro Ser Ser Lys His Gln His Lys Lys 225 230 235 240 Leu Val Asn Arg Asp Leu Lys Thr Gln Ser Gly Ser Glu Met Lys Lys 245 250 255 Phe Leu Ser Thr Leu Thr Ile Asp Gly Val Thr Arg Ser Asp Gln Gly 260 265 270 Leu Tyr Thr Cys Ala Ala Ser Ser Gly Leu Met Thr Lys Lys Asn Ser 275 280 285 Thr Phe Val Arg Val His Glu Lys Gly Gly Gly Gly Gly Ser Tyr Ser 290 295 300 Met Thr Pro Pro Thr Leu Asn Ile Thr Glu Glu Ser His Val Ile Asp 305 310 315 320 Thr Gly Asp Ser Leu Ser Ile Ser Cys Arg Gly Gln His Pro Leu Glu 325 330 335 Trp Ala Trp Pro Gly Ala Gln Glu Ala Pro Ala Thr Gly Asp Lys Asp 340 345 350 Ser Glu Asp Thr Gly Val Val Arg Asp Cys Glu Gly Thr Asp Ala Arg 355 360 365 Pro Tyr Cys Lys Val Leu Leu Leu His Glu Val His Ala Gln Asp Thr 370 375 380 Gly Ser Tyr Val Cys Tyr Tyr Lys Tyr Ile Lys Ala Arg Ile Glu Gly 385 390 395 400 Thr Thr Ala Ala Ser Ser Tyr Val Phe Val Arg Asp Phe Glu Gln Pro 405 410 415 Phe Ile Asn Lys Pro Asp Thr Leu Leu Val Asn Arg Lys Asp Ala Met 420 425 430 Trp Val Pro Cys Leu Val Ser Ile Pro Gly Leu Asn Val Thr Leu Arg 435 440 445 Ser Gln Ser Ser Val Leu Trp Pro Asp Gly Gln Glu Val Val Trp Asp 450 455 460 Asp Arg Arg Gly Met Leu Val Ser Thr Pro Leu Leu His Asp Ala Leu 465 470 475 480 Tyr Leu Gln Cys Glu Thr Thr Trp Gly Asp Gln Asp Phe Leu Ser Asn 485 490 495 Pro Phe Leu Val His Ile Thr Gly Asp Lys Thr His Thr Cys Pro Pro 500 505 510 Cys Pro Ala Pro Glu Leu Leu Gly Gly Pro Ser Val Phe Leu Phe Pro 515 520 525 Pro Lys Pro Lys Asp Thr Leu Met Ile Ser Arg Thr Pro Glu Val Thr 530 535 540 Cys Val Val Val Asp Val Ser His Glu Asp Pro Glu Val Lys Phe Asn 545 550 555 560 Trp Tyr Val Asp Gly Val Glu Val His Asn Ala Lys Thr Lys Pro Arg 565 570 575 Glu Glu Gln Tyr Asn Ser Thr Tyr Arg Val Val Ser Val Leu Thr Val 580 585 590 Leu His Gln Asp Trp Leu Asn Gly Lys Glu Tyr Lys Cys Lys Val Ser 595 600 605 Asn Lys Ala Leu Pro Ala Pro Ile Glu Lys Thr Ile Ser Lys Ala Lys 610 615 620 Gly Gln Pro Arg Glu Pro Gln Val Tyr Thr Leu Pro Pro Ser Arg Asp 625 630 635 640 Glu Leu Thr Lys Asn Gln Val Ser Leu Thr Cys Leu Val Lys Gly Phe 645 650 655 Tyr Pro Ser Asp Ile Ala Val Glu Trp Glu Ser Asn Gly Gln Pro Glu 660 665 670 Asn Asn Tyr Lys Thr Thr Pro Pro Val Leu Asp Ser Asp Gly Ser Phe 675 680 685 Phe Leu Tyr Ser Lys Leu Thr Val Asp Lys Ser Arg Trp Gln Gln Gly 690 695 700 Asn Val Phe Ser Cys Ser Val Met His Glu Ala Leu His Asn His Tyr 705 710 715 720 Thr Gln Lys Ser Leu Ser Leu Ser Pro Gly Lys 725 730 <210> 62 <211> 725 <212> PRT <213> Artificial Sequence <220> <223> Exemplary Polypeptide 9 (such as in EXG102-28) <400> 62 Met Val Ser Tyr Trp Asp Thr Gly Val Leu Leu Cys Ala Leu Leu Ser 1 5 10 15 Cys Leu Leu Leu Thr Gly Ser Ser Ser Gly Ser Asp Thr Gly Arg Pro 20 25 30 Phe Val Glu Met Tyr Ser Glu Ile Pro Glu Ile Ile His Met Thr Glu 35 40 45 Gly Arg Glu Leu Val Ile Pro Cys Arg Val Thr Ser Pro Asn Ile Thr 50 55 60 Val Thr Leu Lys Lys Phe Pro Leu Asp Thr Leu Ile Pro Asp Gly Lys 65 70 75 80 Arg Ile Ile Trp Asp Ser Arg Lys Gly Phe Ile Ile Ser Asn Ala Thr 85 90 95 Tyr Lys Glu Ile Gly Leu Leu Thr Cys Glu Ala Thr Val Asn Gly His 100 105 110 Leu Tyr Lys Thr Asn Tyr Leu Thr His Arg Gln Thr Asn Thr Ile Ile 115 120 125 Asp Val Val Leu Ser Pro Ser His Gly Ile Glu Leu Ser Val Gly Glu 130 135 140 Lys Leu Val Leu Asn Cys Thr Ala Arg Thr Glu Leu Asn Val Gly Ile 145 150 155 160 Asp Phe Asn Trp Glu Tyr Pro Ser Ser Lys His Gln His Lys Lys Leu 165 170 175 Val Asn Arg Asp Leu Lys Thr Gln Ser Gly Ser Glu Met Lys Lys Phe 180 185 190 Leu Ser Thr Leu Thr Ile Asp Gly Val Thr Arg Ser Asp Gln Gly Leu 195 200 205 Tyr Thr Cys Ala Ala Ser Ser Gly Leu Met Thr Lys Lys Asn Ser Thr 210 215 220 Phe Val Arg Val His Glu Lys Gly Gly Gly Gly Gly Ser Tyr Ser Met 225 230 235 240 Thr Pro Pro Thr Leu Asn Ile Thr Glu Glu Ser His Val Ile Asp Thr 245 250 255 Gly Asp Ser Leu Ser Ile Ser Cys Arg Gly Gln His Pro Leu Glu Trp 260 265 270 Ala Trp Pro Gly Ala Gln Glu Ala Pro Ala Thr Gly Asp Lys Asp Ser 275 280 285 Glu Asp Thr Gly Val Val Arg Asp Cys Glu Gly Thr Asp Ala Arg Pro 290 295 300 Tyr Cys Lys Val Leu Leu Leu His Glu Val His Ala Gln Asp Thr Gly 305 310 315 320 Ser Tyr Val Cys Tyr Tyr Lys Tyr Ile Lys Ala Arg Ile Glu Gly Thr 325 330 335 Thr Ala Ala Ser Ser Tyr Val Phe Val Arg Asp Phe Glu Gln Pro Phe 340 345 350 Ile Asn Lys Pro Asp Thr Leu Leu Val Asn Arg Lys Asp Ala Met Trp 355 360 365 Val Pro Cys Leu Val Ser Ile Pro Gly Leu Asn Val Thr Leu Arg Ser 370 375 380 Gln Ser Ser Val Leu Trp Pro Asp Gly Gln Glu Val Val Trp Asp Asp 385 390 395 400 Arg Arg Gly Met Leu Val Ser Thr Pro Leu Leu His Asp Ala Leu Tyr 405 410 415 Leu Gln Cys Glu Thr Thr Trp Gly Asp Gln Asp Phe Leu Ser Asn Pro 420 425 430 Phe Leu Val His Ile Thr Gly Asp Lys Thr His Thr Cys Pro Pro Cys 435 440 445 Pro Ala Pro Glu Leu Leu Gly Gly Pro Ser Val Phe Leu Phe Pro Pro 450 455 460 Lys Pro Lys Asp Thr Leu Met Ile Ser Arg Thr Pro Glu Val Thr Cys 465 470 475 480 Val Val Val Asp Val Ser His Glu Asp Pro Glu Val Lys Phe Asn Trp 485 490 495 Tyr Val Asp Gly Val Glu Val His Asn Ala Lys Thr Lys Pro Arg Glu 500 505 510 Glu Gln Tyr Asn Ser Thr Tyr Arg Val Val Ser Val Leu Thr Val Leu 515 520 525 His Gln Asp Trp Leu Asn Gly Lys Glu Tyr Lys Cys Lys Val Ser Asn 530 535 540 Lys Ala Leu Pro Ala Pro Ile Glu Lys Thr Ile Ser Lys Ala Lys Gly 545 550 555 560 Gln Pro Arg Glu Pro Gln Val Tyr Thr Leu Pro Pro Ser Arg Asp Glu 565 570 575 Leu Thr Lys Asn Gln Val Ser Leu Thr Cys Leu Val Lys Gly Phe Tyr 580 585 590 Pro Ser Asp Ile Ala Val Glu Trp Glu Ser Asn Gly Gln Pro Glu Asn 595 600 605 Asn Tyr Lys Thr Thr Pro Pro Val Leu Asp Ser Asp Gly Ser Phe Phe 610 615 620 Leu Tyr Ser Lys Leu Thr Val Asp Lys Ser Arg Trp Gln Gln Gly Asn 625 630 635 640 Val Phe Ser Cys Ser Val Met His Glu Ala Leu His Asn His Tyr Thr 645 650 655 Gln Lys Ser Leu Ser Leu Ser Pro Gly Lys Gly Gly Gly Gly Gly Ala 660 665 670 Gln Gln Glu Glu Cys Glu Trp Asp Pro Trp Thr Cys Glu His Met Gly 675 680 685 Ser Gly Ser Ala Thr Gly Gly Ser Gly Ser Thr Ala Ser Ser Gly Ser 690 695 700 Gly Ser Ala Thr His Gln Glu Glu Cys Glu Trp Asp Pro Trp Thr Cys 705 710 715 720 Glu His Met Leu Glu 725 <210> 63 <211> 736 <212> PRT <213> Artificial Sequence <220> <223> Exemplary Polypeptide 10 (such as in EXG102 - 29) <400> 63 Met Val Ser Tyr Trp Asp Thr Gly Val Leu Leu Cys Ala Leu Leu Ser 1 5 10 15 Cys Leu Leu Leu Thr Gly Ser Ser Ser Gly Ser Asp Thr Gly Arg Pro 20 25 30 Phe Val Glu Met Tyr Ser Glu Ile Pro Glu Ile Ile His Met Thr Glu 35 40 45 Gly Arg Glu Leu Val Ile Pro Cys Arg Val Thr Ser Pro Asn Ile Thr 50 55 60 Val Thr Leu Lys Lys Phe Pro Leu Asp Thr Leu Ile Pro Asp Gly Lys 65 70 75 80 Arg Ile Ile Trp Asp Ser Arg Lys Gly Phe Ile Ile Ser Asn Ala Thr 85 90 95 Tyr Lys Glu Ile Gly Leu Leu Thr Cys Glu Ala Thr Val Asn Gly His 100 105 110 Leu Tyr Lys Thr Asn Tyr Leu Thr His Arg Gln Thr Asn Thr Ile Ile 115 120 125 Asp Val Val Leu Ser Pro Ser His Gly Ile Glu Leu Ser Val Gly Glu 130 135 140 Lys Leu Val Leu Asn Cys Thr Ala Arg Thr Glu Leu Asn Val Gly Ile 145 150 155 160 Asp Phe Asn Trp Glu Tyr Pro Ser Ser Lys His Gln His Lys Lys Leu 165 170 175 Val Asn Arg Asp Leu Lys Thr Gln Ser Gly Ser Glu Met Lys Lys Phe 180 185 190 Leu Ser Thr Leu Thr Ile Asp Gly Val Thr Arg Ser Asp Gln Gly Leu 195 200 205 Tyr Thr Cys Ala Ala Ser Ser Gly Leu Met Thr Lys Lys Asn Ser Thr 210 215 220 Phe Val Arg Val His Glu Lys Gly Gly Gly Gly Gly Ser Tyr Ser Met 225 230 235 240 Thr Pro Pro Thr Leu Asn Ile Thr Glu Glu Ser His Val Ile Asp Thr 245 250 255 Gly Asp Ser Leu Ser Ile Ser Cys Arg Gly Gln His Pro Leu Glu Trp 260 265 270 Ala Trp Pro Gly Ala Gln Glu Ala Pro Ala Thr Gly Asp Lys Asp Ser 275 280 285 Glu Asp Thr Gly Val Val Arg Asp Cys Glu Gly Thr Asp Ala Arg Pro 290 295 300 Tyr Cys Lys Val Leu Leu Leu His Glu Val His Ala Gln Asp Thr Gly 305 310 315 320 Ser Tyr Val Cys Tyr Tyr Lys Tyr Ile Lys Ala Arg Ile Glu Gly Thr 325 330 335 Thr Ala Ala Ser Ser Tyr Val Phe Val Arg Asp Phe Glu Gln Pro Phe 340 345 350 Ile Asn Lys Pro Asp Thr Leu Leu Val Asn Arg Lys Asp Ala Met Trp 355 360 365 Val Pro Cys Leu Val Ser Ile Pro Gly Leu Asn Val Thr Leu Arg Ser 370 375 380 Gln Ser Ser Val Leu Trp Pro Asp Gly Gln Glu Val Val Trp Asp Asp 385 390 395 400 Arg Arg Gly Met Leu Val Ser Thr Pro Leu Leu His Asp Ala Leu Tyr 405 410 415 Leu Gln Cys Glu Thr Thr Trp Gly Asp Gln Asp Phe Leu Ser Asn Pro 420 425 430 Phe Leu Val His Ile Thr Gly Asp Lys Thr His Thr Cys Pro Pro Cys 435 440 445 Pro Ala Pro Glu Leu Leu Gly Gly Pro Ser Val Phe Leu Phe Pro Pro 450 455 460 Lys Pro Lys Asp Thr Leu Met Ile Ser Arg Thr Pro Glu Val Thr Cys 465 470 475 480 Val Val Val Asp Val Ser His Glu Asp Pro Glu Val Lys Phe Asn Trp 485 490 495 Tyr Val Asp Gly Val Glu Val His Asn Ala Lys Thr Lys Pro Arg Glu 500 505 510 Glu Gln Tyr Asn Ser Thr Tyr Arg Val Val Ser Val Leu Thr Val Leu 515 520 525 His Gln Asp Trp Leu Asn Gly Lys Glu Tyr Lys Cys Lys Val Ser Asn 530 535 540 Lys Ala Leu Pro Ala Pro Ile Glu Lys Thr Ile Ser Lys Ala Lys Gly 545 550 555 560 Gln Pro Arg Glu Pro Gln Val Tyr Thr Leu Pro Pro Ser Arg Asp Glu 565 570 575 Leu Thr Lys Asn Gln Val Ser Leu Thr Cys Leu Val Lys Gly Phe Tyr 580 585 590 Pro Ser Asp Ile Ala Val Glu Trp Glu Ser Asn Gly Gln Pro Glu Asn 595 600 605 Asn Tyr Lys Thr Thr Pro Pro Val Leu Asp Ser Asp Gly Ser Phe Phe 610 615 620 Leu Tyr Ser Lys Leu Thr Val Asp Lys Ser Arg Trp Gln Gln Gly Asn 625 630 635 640 Val Phe Ser Cys Ser Val Met His Glu Ala Leu His Asn His Tyr Thr 645 650 655 Gln Lys Ser Leu Ser Leu Ser Pro Gly Lys Gly Gly Gly Gly Gly Ala 660 665 670 Gln Gly Ser Gly Ser Ala Thr Gly Gly Ser Gly Ser Thr Ala Ser Ser 675 680 685 Gly Ser Gly Ser Ala Thr His Lys Phe Asn Pro Leu Asp Glu Leu Glu 690 695 700 Glu Thr Leu Tyr Glu Gln Phe Thr Phe Gln Gln Gly Gly Gly Gly Gly 705 710 715 720 Gln Glu Glu Cys Glu Trp Asp Pro Trp Thr Cys Glu His Met Leu Glu 725 730 735 <210> 64 <211> 534 <212> PRT <213> Artificial Sequence <220> <223> Exemplary Polypeptide 11 (such as in EXG102 - 30) <400> 64 Met Val Ser Tyr Trp Asp Thr Gly Val Leu Leu Cys Ala Leu Leu Ser 1 5 10 15 Cys Leu Leu Leu Thr Gly Ser Ser Ser Gly Gly Gly Gly Gly Gly Ala 20 25 30 Gln Gly Ser Gly Ser Ala Thr Gly Gly Ser Gly Ser Thr Ala Ser Ser 35 40 45 Gly Ser Gly Ser Ala Thr His Lys Phe Asn Pro Leu Asp Glu Leu Glu 50 55 60 Glu Thr Leu Tyr Glu Gln Phe Thr Phe Gln Gln Gly Gly Gly Gly Gly 65 70 75 80 Gln Glu Glu Cys Glu Trp Asp Pro Trp Thr Cys Glu His Met Leu Glu 85 90 95 Gly Gly Gly Gly Gly Ser Ser Asp Thr Gly Arg Pro Phe Val Glu Met 100 105 110 Tyr Ser Glu Ile Pro Glu Ile Ile His Met Thr Glu Gly Arg Glu Leu 115 120 125 Val Ile Pro Cys Arg Val Thr Ser Pro Asn Ile Thr Val Thr Leu Lys 130 135 140 Lys Phe Pro Leu Asp Thr Leu Ile Pro Asp Gly Lys Arg Ile Ile Trp 145 150 155 160 Asp Ser Arg Lys Gly Phe Ile Ile Ser Asn Ala Thr Tyr Lys Glu Ile 165 170 175 Gly Leu Leu Thr Cys Glu Ala Thr Val Asn Gly His Leu Tyr Lys Thr 180 185 190 Asn Tyr Leu Thr His Arg Gln Thr Asn Thr Ile Ile Asp Val Val Leu 195 200 205 Ser Pro Ser His Gly Ile Glu Leu Ser Val Gly Glu Lys Leu Val Leu 210 215 220 Asn Cys Thr Ala Arg Thr Glu Leu Asn Val Gly Ile Asp Phe Asn Trp 225 230 235 240 Glu Tyr Pro Ser Ser Lys His Gln His Lys Lys Leu Val Asn Arg Asp 245 250 255 Leu Lys Thr Gln Ser Gly Ser Glu Met Lys Lys Phe Leu Ser Thr Leu 260 265 270 Thr Ile Asp Gly Val Thr Arg Ser Asp Gln Gly Leu Tyr Thr Cys Ala 275 280 285 Ala Ser Ser Gly Leu Met Thr Lys Lys Asn Ser Thr Phe Val Arg Val 290 295 300 His Glu Lys Asp Lys Thr His Thr Cys Pro Pro Cys Pro Ala Pro Glu 305 310 315 320 Leu Leu Gly Gly Pro Ser Val Phe Leu Phe Pro Pro Lys Pro Lys Asp 325 330 335 Thr Leu Met Ile Ser Arg Thr Pro Glu Val Thr Cys Val Val Val Asp 340 345 350 Val Ser His Glu Asp Pro Glu Val Lys Phe Asn Trp Tyr Val Asp Gly 355 360 365 Val Glu Val His Asn Ala Lys Thr Lys Pro Arg Glu Glu Gln Tyr Asn 370 375 380 Ser Thr Tyr Arg Val Val Ser Val Leu Thr Val Leu His Gln Asp Trp 385 390 395 400 Leu Asn Gly Lys Glu Tyr Lys Cys Lys Val Ser Asn Lys Ala Leu Pro 405 410 415 Ala Pro Ile Glu Lys Thr Ile Ser Lys Ala Lys Gly Gln Pro Arg Glu 420 425 430 Pro Gln Val Tyr Thr Leu Pro Pro Ser Arg Asp Glu Leu Thr Lys Asn 435 440 445 Gln Val Ser Leu Thr Cys Leu Val Lys Gly Phe Tyr Pro Ser Asp Ile 450 455 460 Ala Val Glu Trp Glu Ser Asn Gly Gln Pro Glu Asn Asn Tyr Lys Thr 465 470 475 480 Thr Pro Pro Val Leu Asp Ser Asp Gly Ser Phe Phe Leu Tyr Ser Lys 485 490 495 Leu Thr Val Asp Lys Ser Arg Trp Gln Gln Gly Asn Val Phe Ser Cys 500 505 510 Ser Val Met His Glu Ala Leu His Asn His Tyr Thr Gln Lys Ser Leu 515 520 525 Ser Leu Ser Pro Gly Lys 530 <210> 65 <211> 742 <212> PRT <213> Artificial sequence <220> <223> Exemplary polypeptide 12 (such as in EXG102 - 31) <400> 65 Met Val Ser Tyr Trp Asp Thr Gly Val Leu Leu Cys Ala Leu Leu Ser 1 5 10 15 Cys Leu Leu Leu Thr Gly Ser Ser Ser Gly Gly Gly Gly Gly Gly Ala 20 25 30 Gln Gly Ser Gly Ser Ala Thr Gly Gly Ser Gly Ser Thr Ala Ser Ser 35 40 45 Gly Ser Gly Ser Ala Thr His Lys Phe Asn Pro Leu Asp Glu Leu Glu 50 55 60 Glu Thr Leu Tyr Glu Gln Phe Thr Phe Gln Gln Gly Gly Gly Gly Gly 65 70 75 80 Gln Glu Glu Cys Glu Trp Asp Pro Trp Thr Cys Glu His Met Leu Glu 85 90 95 Gly Gly Gly Gly Gly Ser Ser Asp Thr Gly Arg Pro Phe Val Glu Met 100 105 110 Tyr Ser Glu Ile Pro Glu Ile Ile His Met Thr Glu Gly Arg Glu Leu 115 120 125 Val Ile Pro Cys Arg Val Thr Ser Pro Asn Ile Thr Val Thr Leu Lys 130 135 140 Lys Phe Pro Leu Asp Thr Leu Ile Pro Asp Gly Lys Arg Ile Ile Trp 145 150 155 160 Asp Ser Arg Lys Gly Phe Ile Ile Ser Asn Ala Thr Tyr Lys Glu Ile 165 170 175 Gly Leu Leu Thr Cys Glu Ala Thr Val Asn Gly His Leu Tyr Lys Thr 180 185 190 Asn Tyr Leu Thr His Arg Gln Thr Asn Thr Ile Ile Asp Val Val Leu 195 200 205 Ser Pro Ser His Gly Ile Glu Leu Ser Val Gly Glu Lys Leu Val Leu 210 215 220 Asn Cys Thr Ala Arg Thr Glu Leu Asn Val Gly Ile Asp Phe Asn Trp 225 230 235 240 Glu Tyr Pro Ser Ser Lys His Gln His Lys Lys Leu Val Asn Arg Asp 245 250 255 Leu Lys Thr Gln Ser Gly Ser Glu Met Lys Lys Phe Leu Ser Thr Leu 260 265 270 Thr Ile Asp Gly Val Thr Arg Ser Asp Gln Gly Leu Tyr Thr Cys Ala 275 280 285 Ala Ser Ser Gly Leu Met Thr Lys Lys Asn Ser Thr Phe Val Arg Val 290 295 300 His Glu Lys Gly Gly Gly Gly Gly Ser Tyr Ser Met Thr Pro Pro Thr 305 310 315 320 Leu Asn Ile Thr Glu Glu Ser His Val Ile Asp Thr Gly Asp Ser Leu 325 330 335 Ser Ile Ser Cys Arg Gly Gln His Pro Leu Glu Trp Ala Trp Pro Gly 340 345 350 Ala Gln Glu Ala Pro Ala Thr Gly Asp Lys Asp Ser Glu Asp Thr Gly 355 360 365 Val Val Arg Asp Cys Glu Gly Thr Asp Ala Arg Pro Tyr Cys Lys Val 370 375 380 Leu Leu Leu His Glu Val His Ala Asn Asp Thr Gly Ser Tyr Val Cys 385 390 395 400 Tyr Tyr Lys Tyr Ile Lys Ala Arg Ile Glu Gly Thr Thr Ala Ala Ser 405 410 415 Ser Tyr Val Phe Val Arg Asp Phe Glu Gln Pro Phe Ile Asn Lys Pro 420 425 430 Asp Thr Leu Leu Val Asn Arg Lys Asp Ala Met Trp Val Pro Cys Leu 435 440 445 Val Ser Ile Pro Gly Leu Asn Val Thr Leu Arg Ser Gln Ser Ser Val 450 455 460 Leu Trp Pro Asp Gly Gln Glu Val Val Trp Asp Asp Arg Arg Gly Met 465 470 475 480 Leu Val Ser Thr Pro Leu Leu His Asp Ala Leu Tyr Leu Gln Cys Glu 485 490 495 Thr Thr Trp Gly Asp Gln Asp Phe Leu Ser Asn Pro Phe Leu Val His 500 505 510 Ile Thr Gly Asp Lys Thr His Thr Cys Pro Pro Cys Pro Ala Pro Glu 515 520 525 Leu Leu Gly Gly Pro Ser Val Phe Leu Phe Pro Pro Lys Pro Lys Asp 530 535 540 Thr Leu Met Ile Ser Arg Thr Pro Glu Val Thr Cys Val Val Val Asp 545 550 555 560 Val Ser His Glu Asp Pro Glu Val Lys Phe Asn Trp Tyr Val Asp Gly 565 570 575 Val Glu Val His Asn Ala Lys Thr Lys Pro Arg Glu Glu Gln Tyr Asn 580 585 590 Ser Thr Tyr Arg Val Val Ser Val Leu Thr Val Leu His Gln Asp Trp 595 600 605 Leu Asn Gly Lys Glu Tyr Lys Cys Lys Val Ser Asn Lys Ala Leu Pro 610 615 620 Ala Pro Ile Glu Lys Thr Ile Ser Lys Ala Lys Gly Gln Pro Arg Glu 625 630 635 640 Pro Gln Val Tyr Thr Leu Pro Pro Ser Arg Asp Glu Leu Thr Lys Asn 645 650 655 Gln Val Ser Leu Thr Cys Leu Val Lys Gly Phe Tyr Pro Ser Asp Ile 660 665 670 Ala Val Glu Trp Glu Ser Asn Gly Gln Pro Glu Asn Asn Tyr Lys Thr 675 680 685 Thr Pro Pro Val Leu Asp Ser Asp Gly Ser Phe Phe Leu Tyr Ser Lys 690 695 700 Leu Thr Val Asp Lys Ser Arg Trp Gln Gln Gly Asn Val Phe Ser Cys 705 710 715 720 Ser Val Met His Glu Ala Leu His Asn His Tyr Thr Gln Lys Ser Leu 725 730 735 Ser Leu Ser Pro Gly Lys 740 <210> 66 <211> 736 <212> PRT <213> Artificial Sequence <220> <223> Exemplary polypeptide 13 (such as in EXG102-32) <400> 66 Met Val Ser Tyr Trp Asp Thr Gly Val Leu Leu Cys Ala Leu Leu Ser 1 5 10 15 Cys Leu Leu Leu Thr Gly Ser Ser Ser Gly Ser Asp Thr Gly Arg Pro 20 25 30 Phe Val Glu Met Tyr Ser Glu Ile Pro Glu Ile Ile His Met Thr Glu 35 40 45 Gly Arg Glu Leu Val Ile Pro Cys Arg Val Thr Ser Pro Asn Ile Thr 50 55 60 Val Thr Leu Lys Lys Phe Pro Leu Asp Thr Leu Ile Pro Asp Gly Lys 65 70 75 80 Arg Ile Ile Trp Asp Ser Arg Lys Gly Phe Ile Ile Ser Asn Ala Thr 85 90 95 Tyr Lys Glu Ile Gly Leu Leu Thr Cys Glu Ala Thr Val Asn Gly His 100 105 110 Leu Tyr Lys Thr Asn Tyr Leu Thr His Arg Gln Thr Asn Thr Ile Ile 115 120 125 Asp Val Val Leu Ser Pro Ser His Gly Ile Glu Leu Ser Val Gly Glu 130 135 140 Lys Leu Val Leu Asn Cys Thr Ala Arg Thr Glu Leu Asn Val Gly Ile 145 150 155 160 Asp Phe Asn Trp Glu Tyr Pro Ser Ser Lys His Gln His Lys Lys Leu 165 170 175 Val Asn Arg Asp Leu Lys Thr Gln Ser Gly Ser Glu Met Lys Lys Phe 180 185 190 Leu Ser Thr Leu Thr Ile Asp Gly Val Thr Arg Ser Asp Gln Gly Leu 195 200 205 Tyr Thr Cys Ala Ala Ser Ser Gly Leu Met Thr Lys Lys Asn Ser Thr 210 215 220 Phe Val Arg Val His Glu Lys Gly Gly Gly Gly Gly Ser Tyr Ser Met 225 230 235 240 Thr Pro Pro Thr Leu Asn Ile Thr Glu Glu Ser His Val Ile Asp Thr 245 250 255 Gly Asp Ser Leu Ser Ile Ser Cys Arg Gly Gln His Pro Leu Glu Trp 260 265 270 Ala Trp Pro Gly Ala Gln Glu Ala Pro Ala Thr Gly Asp Lys Asp Ser 275 280 285 Glu Asp Thr Gly Val Val Arg Asp Cys Glu Gly Thr Asp Ala Arg Pro 290 295 300 Tyr Cys Lys Val Leu Leu Leu His Glu Val His Ala Asn Asp Thr Gly 305 310 315 320 Ser Tyr Val Cys Tyr Tyr Lys Tyr Ile Lys Ala Arg Ile Glu Gly Thr 325 330 335 Thr Ala Ala Ser Ser Tyr Val Phe Val Arg Asp Phe Glu Gln Pro Phe 340 345 350 Ile Asn Lys Pro Asp Thr Leu Leu Val Asn Arg Lys Asp Ala Met Trp 355 360 365 Val Pro Cys Leu Val Ser Ile Pro Gly Leu Asn Val Thr Leu Arg Ser 370 375 380 Gln Ser Ser Val Leu Trp Pro Asp Gly Gln Glu Val Val Trp Asp Asp 385 390 395 400 Arg Arg Gly Met Leu Val Ser Thr Pro Leu Leu His Asp Ala Leu Tyr 405 410 415 Leu Gln Cys Glu Thr Thr Trp Gly Asp Gln Asp Phe Leu Ser Asn Pro 420 425 430 Phe Leu Val His Ile Thr Gly Asp Lys Thr His Thr Cys Pro Pro Cys 435 440 445 Pro Ala Pro Glu Leu Leu Gly Gly Pro Ser Val Phe Leu Phe Pro Pro 450 455 460 Lys Pro Lys Asp Thr Leu Met Ile Ser Arg Thr Pro Glu Val Thr Cys 465 470 475 480 Val Val Val Asp Val Ser His Glu Asp Pro Glu Val Lys Phe Asn Trp 485 490 495 Tyr Val Asp Gly Val Glu Val His Asn Ala Lys Thr Lys Pro Arg Glu 500 505 510 Glu Gln Tyr Asn Ser Thr Tyr Arg Val Val Ser Val Leu Thr Val Leu 515 520 525 His Gln Asp Trp Leu Asn Gly Lys Glu Tyr Lys Cys Lys Val Ser Asn 530 535 540 Lys Ala Leu Pro Ala Pro Ile Glu Lys Thr Ile Ser Lys Ala Lys Gly 545 550 555 560 Gln Pro Arg Glu Pro Gln Val Tyr Thr Leu Pro Pro Ser Arg Asp Glu 565 570 575 Leu Thr Lys Asn Gln Val Ser Leu Thr Cys Leu Val Lys Gly Phe Tyr 580 585 590 Pro Ser Asp Ile Ala Val Glu Trp Glu Ser Asn Gly Gln Pro Glu Asn 595 600 605 Asn Tyr Lys Thr Thr Pro Pro Val Leu Asp Ser Asp Gly Ser Phe Phe 610 615 620 Leu Tyr Ser Lys Leu Thr Val Asp Lys Ser Arg Trp Gln Gln Gly Asn 625 630 635 640 Val Phe Ser Cys Ser Val Met His Glu Ala Leu His Asn His Tyr Thr 645 650 655 Gln Lys Ser Leu Ser Leu Ser Pro Gly Lys Gly Gly Gly Gly Gly Ala 660 665 670 Gln Gly Ser Gly Ser Ala Thr Gly Gly Ser Gly Ser Thr Ala Ser Ser 675 680 685 Gly Ser Gly Ser Ala Thr His Lys Phe Asn Pro Leu Asp Glu Leu Glu 690 695 700 Glu Thr Leu Tyr Glu Gln Phe Thr Phe Gln Gln Gly Gly Gly Gly Gly 705 710 715 720 Gln Glu Glu Cys Glu Trp Asp Pro Trp Thr Cys Glu His Met Leu Glu 725 730 735 <210> 67 <211> 1905 <212> DNA <213> Artificial Sequence <220> <223> Nucleic acid of exemplary polypeptide 5 (such as in EXG102 - 24) <400> 67 atggtgtctt attgggatac tggcgtgctg ctctgtgccc tcctgagttg cctgctcctg 60 actggttctt cttctggggg cggcggcgga ggcgcccagc aagaagagtg cgagtgggat 120 ccctggacct gcgagcacat gggatccggc agcgccaccg gaggatccgg aagcaccgcc 180 tccagcggct ccggcagcgc cacccaccag gaggagtgtg agtgggaccc ctggacctgc 240 gaacacatgc tggagggcgg cggcggaggc agctccgata ctgggcgccc cttcgtggag 300 atgtactccg agatccctga aatcattcac atgactgagg gtcgggaact ggtcatccca 360 tgccgcgtga cctctcccaa cattactgtg accctgaaga aattccctct ggacaccctc 420 atcccagatg ggaagaggat catttgggac tcaagaaagg gttttatcat cagcaacgct 480 acatacaagg agattggcct gctcacctgc gaagcaacag tgaacggaca cctgtacaag 540 actaattatc tcacccatag acagacaaac actatcattg atgtggtcct gtcaccaagc 600 cacggcatcg agctcagcgt cggtgaaaag ctggtgctca attgtacagc ccggactgag 660 ctgaacgtgg gcattgactt caattgggaa taccccagct ccaagcacca gcataagaaa 720 ctggtgaacc gcgatctcaa aacccagtcc ggatctgaga tgaagaaatt tctgagcacc 780 ctcacaatcg acggcgtgac acgatccgat cagggactgt atacttgcgc cgcttctagt 840 ggcctgatga ccaagaaaaa tagcacattc gtcagggtgc acgaaaaggg cggcggcgga 900 ggcagcgtgt acgtgcaaga ctacagaagc cccttcatcg ctagcgtgag cgatcagcac 960 ggcgtggtgt acatcaccga gaacaagaac aagaccgtgg tgatcccctg cctgggcagc 1020 atcagcaacc tgaacgtgag cctgtgcgct agataccccg agaagagatt cgtgcccgac 1080 ggcaacagaa tcagctggga cagcaagaag ggcttcacca tccctagcta catgatcagc 1140 tacgccggca tggtgttctg cgaggccaag atcaacgacg agagctatca gagcatcatg 1200 tacatcgtgg tcgtggtggg cgacaaaact catacctgcc caccttgtcc agcaccagag 1260 ctgctcggag gaccatccgt gttcctgttt ccacccaagc ccaaagatac tctgatgatt 1320 tcacgcacac ccgaagtcac ttgcgtggtc gtggacgtgt cccacgagga ccccgaagtc 1380 aagtttaact ggtacgtgga cggcgtcgag gtgcataatg ctaagacaaa accccgagag 1440 gaacagtaca actctaccta tagggtcgtg agtgtcctga cagtgctcca ccaggattgg 1500 ctgaacggaa aggagtataa gtgcaaagtg tctaataagg cactgcctgc cccaatcgag 1560 aaaacaatta gtaaggccaa agggcagccc agagaacctc aggtgtacac tctgcctcca 1620 tctcgggacg agctcactaa gaaccaggtc agtctgacct gtctcgtgaa agggttctat 1680 cctagtgata tcgctgtgga gtgggaatca aatggtcagc cagagaacaa ttacaagacc 1740 acaccccctg tcctggacag cgatggctcc ttctttctgt attccaagct caccgtggac 1800 aaatctcgat ggcagcaggg aaacgtcttt agttgttcag tgatgcacga agccctccat 1860 aaccactaca ctcagaaaag cctcagcctc agccctggga aatga 1905 <210> 68 <211> 1887 <212> DNA <213> Artificial Sequence <220> <223> Nucleic acid of exemplary polypeptide 6 (such as in EXG102 - 25) <400> 68 atggtgtctt attgggatac tggcgtgctg ctctgtgccc tcctgagttg cctgctcctg 60 actggttctt cttctgggtc cgatactggg cgccccttcg tggagatgta ctccgagatc 120 cctgaaatca ttcacatgac tgagggtcgg gaactggtca tcccatgccg cgtgacctct 180 cccaacatta ctgtgaccct gaagaaattc cctctggaca ccctcatccc agatgggaag 240 cccaacatta ctgtgaccct gaagaaattc cctctggaca ccctcatccc agatgggaag 240 aggatcattt gggactcaag aaagggtttt atcatcagca acgctacata caaggagatt 300 aggatcattt gggactcaag aaagggtttt atcatcagca acgctacata caaggagatt 300 ggcctgctca cctgcgaagc aacagtgaac ggacacctgt acaagactaa ttatctcacc 360 ggcctgctca cctgcgaagc aacagtgaac ggacacctgt acaagactaa ttatctcacc 360 catagacaga caaacactat cattgatgtg gtcctgtcac caagccacgg catcgagctc 420 catagacaga caaacactat cattgatgtg gtcctgtcac caagccacgg catcgagctc 420 agcgtcggtg aaaagctggt gctcaattgt acagcccgga ctgagctgaa cgtgggcatt 480 agcgtcggtg aaaagctggt gctcaattgt acagcccgga ctgagctgaa cgtgggcatt 480 gacttcaatt gggaataccc cagctccaag caccagcata agaaactggt gaaccgcgat 540 gacttcaatt gggaataccc cagctccaag caccagcata agaaactggt gaaccgcgat 540 ctcaaaaccc agtccggatc tgagatgaag aaatttctga gcaccctcac aatcgacggc 600 ctcaaaaccc agtccggatc tgagatgaag aaatttctga gcaccctcac aatcgacggc 600 gtgacacgat ccgatcaggg actgtatact tgcgccgctt ctagtggcct gatgaccaag 660 gtgacacgat ccgatcaggg actgtatact tgcgccgctt ctagtggcct gatgaccaag 660 aaaaatagca cattcgtcag ggtgcacgaa aagggcggcg gcggaggcag cgtgtacgtg 720 aaaaatagca cattcgtcag ggtgcacgaa aagggcggcg gcggaggcag cgtgtacgtg 720 caagactaca gaagcccctt catcgctagc gtgagcgatc agcacggcgt ggtgtacatc 780 caagactaca gaagcccctt catcgctagc gtgagcgatc agcacggcgt ggtgtacatc 780 accgagaaca agaacaagac cgtggtgatc ccctgcctgg gcagcatcag caacctgaac 840 accgagaaca agaacaagac cgtggtgatc ccctgcctgg gcagcatcag caacctgaac 840 gtgagcctgt gcgctagata ccccgagaag agattcgtgc ccgacggcaa cagaatcagc 900 gtgagcctgt gcgctagata ccccgagaag agattcgtgc ccgacggcaa cagaatcagc 900 tgggacagca agaagggctt caccatccct agctacatga tcagctacgc cggcatggtg 960 ttctgcgagg ccaagatcaa cgacgagagc tatcagagca tcatgtacat cgtggtcgtg 1020 gtgggcgaca aaactcatac ctgcccacct tgtccagcac cagagctgct cggaggacca 1080 tccgtgttcc tgtttccacc caagcccaaa gatactctga tgatttcacg cacacccgaa 1140 gtcacttgcg tggtcgtgga cgtgtcccac gaggaccccg aagtcaagtt taactggtac 1200 gtggacggcg tcgaggtgca taatgctaag acaaaacccc gagaggaaca gtacaactct 1260 acctataggg tcgtgagtgt cctgacagtg ctccaccagg attggctgaa cggaaaggag 1320 tataagtgca aagtgtctaa taaggcactg cctgccccaa tcgagaaaac aattagtaag 1380 gccaaagggc agcccagaga acctcaggtg tacactctgc ctccatctcg ggacgagctc 1440 actaagaacc aggtcagtct gacctgtctc gtgaaagggt tctatcctag tgatatcgct 1500 gtggagtggg aatcaaatgg tcagccagag aacaattaca agaccacacc ccctgtcctg 1560 gacagcgatg gctccttctt tctgtattcc aagctcaccg tggacaaatc tcgatggcag 1620 cagggaaacg tctttagttg ttcagtgatg cacgaagccc tccataacca ctacactcag 1680 aaaagcctca gcctcagccc tgggaaaggc ggcggcggag gcgcccagca agaagagtgc 1740 gagtgggatc cctggacctg cgagcacatg ggatccggca gcgccaccgg aggatccgga 1800 agcaccgcct ccagcggctc cggcagcgcc acccaccagg aggagtgtga gtgggacccc 1860 tggacctgcg agcacatgct ggagtga 1887 <210> 69 <211> 1920 <212> DNA <213> Artificial Sequence <220> <223> Nucleic acid of exemplary polypeptide 7 (such as in EXG102-26) <400> 69 atggtgtctt attgggatac tggcgtgctg ctctgtgccc tcctgagttg cctgctcctg 60 actggttctt cttctgggtc cgatactggg cgccccttcg tggagatgta ctccgagatc 120 cctgaaatca ttcacatgac tgagggtcgg gaactggtca tcccatgccg cgtgacctct 180 cccaacatta ctgtgaccct gaagaaattc cctctggaca ccctcatccc agatgggaag 240 aggatcattt gggactcaag aaagggtttt atcatcagca acgctacata caaggagatt 300 ggcctgctca cctgcgaagc aacagtgaac ggacacctgt acaagactaa ttatctcacc 360 catagacaga caaacactat cattgatgtg gtcctgtcac caagccacgg catcgagctc 420 agcgtcggtg aaaagctggt gctcaattgt acagcccgga ctgagctgaa cgtgggcatt 480 gacttcaatt gggaataccc cagctccaag caccagcata agaaactggt gaaccgcgat 540 ctcaaaaccc agtccggatc tgagatgaag aaatttctga gcaccctcac aatcgacggc 600 gtgacacgat ccgatcaggg actgtatact tgcgccgctt ctagtggcct gatgaccaag 660 aaaaatagca cattcgtcag ggtgcacgaa aagggcggcg gcggaggcag cgtgtacgtg 720 caagactaca gaagcccctt catcgctagc gtgagcgatc agcacggcgt ggtgtacatc 780 accgagaaca agaacaagac cgtggtgatc ccctgcctgg gcagcatcag caacctgaac 840 gtgagcctgt gcgctagata ccccgagaag agattcgtgc ccgacggcaa cagaatcagc 900 tgggacagca agaagggctt caccatccct agctacatga tcagctacgc cggcatggtg 960 ttctgcgagg ccaagatcaa cgacgagagc tatcagagca tcatgtacat cgtggtcgtg 1020 gtgggcgaca aaactcatac ctgcccacct tgtccagcac cagagctgct cggaggacca 1080 tccgtgttcc tgtttccacc caagcccaaa gatactctga tgatttcacg cacacccgaa 1140 gtcacttgcg tggtcgtgga cgtgtcccac gaggaccccg aagtcaagtt taactggtac 1200 gtggacggcg tcgaggtgca taatgctaag acaaaacccc gagaggaaca gtacaactct 1260 acctataggg tcgtgagtgt cctgacagtg ctccaccagg attggctgaa cggaaaggag 1320 tataagtgca aagtgtctaa taaggcactg cctgccccaa tcgagaaaac aattagtaag 1380 gccaaagggc agcccagaga acctcaggtg tacactctgc ctccatctcg ggacgagctc 1440 actaagaacc aggtcagtct gacctgtctc gtgaaagggt tctatcctag tgatatcgct 1500 gtggagtggg aatcaaatgg tcagccagag aacaattaca agaccacacc ccctgtcctg 1560 gacagcgatg gctccttctt tctgtattcc aagctcaccg tggacaaatc tcgatggcag 1620 cagggaaacg tctttagttg ttcagtgatg cacgaagccc tccataacca ctacactcag 1680 aaaagcctca gcctcagccc tgggaaaggg ggcgggggcg gcgctcaagg cagcgggagc 1740 gctaccggcg gctccgggag cacagctagc agcggcagcg gctccgctac ccacaagttc 1800 aaccccctgg acgagctgga ggagaccctc tacgagcagt tcaccttcca acaaggcggg 1860 ggcgggggcc aagaggagtg cgagtgggac ccctggacct gcgagcacat gctggagtga 1920 <210> 70 <211> 2196 <212> DNA <213> Artificial Sequence <220> <223> Nucleic acid of exemplary polypeptide 8 (such as in EXG102 - 27) <400> 70 atggtgtctt attgggatac tggcgtgctg ctctgtgccc tcctgagttg cctgctcctg 60 actggttctt cttctggggg cggcggcgga ggcgcccagc aagaagagtg cgagtgggat 120 ccctggacct gcgagcacat gggatccggc agcgccaccg gaggatccgg aagcaccgcc 180 tccagcggct ccggcagcgc cacccaccag gaggagtgtg agtgggaccc ctggacctgc 240 gaacacatgc tggagggcgg cggcggaggc agctccgata ctgggcgccc cttcgtggag 300 atgtactccg agatccctga aatcattcac atgactgagg gtcgggaact ggtcatccca 360 tgccgcgtga cctctcccaa cattactgtg accctgaaga aattccctct ggacaccctc 420 atcccagatg ggaagaggat catttgggac tcaagaaagg gttttatcat cagcaacgct 480 acatacaagg agattggcct gctcacctgc gaagcaacag tgaacggaca cctgtacaag 540 actaattatc tcacccatag acagacaaac actatcattg atgtggtcct gtcaccaagc 600 cacggcatcg agctcagcgt cggtgaaaag ctggtgctca attgtacagc ccggactgag 660 ctgaacgtgg gcattgactt caattgggaa taccccagct ccaagcacca gcataagaaa 720 ctggtgaacc gcgatctcaa aacccagtcc ggatctgaga tgaagaaatt tctgagcacc 780 ctcacaatcg acggcgtgac acgatccgat cagggactgt atacttgcgc cgcttctagt 840 ggcctgatga ccaagaaaaa tagcacattc gtcagggtgc acgaaaaggg cggcggcgga 900 ggcagctact ccatgacccc cccgaccctg aacatcaccg aggagagcca cgtgatcgac 960 accggcgaca gcctgagcat cagttgccgg ggccaacacc ccctggaatg ggcatggcct 1020 ggggcacaag aggcccccgc aacgggcgac aaggacagcg aggacaccgg cgtggtgaga 1080 gactgcgagg gcaccgacgc tagaccctac tgcaaggtgc tgctgctgca cgaggtgcac 1140 gcccaagaca ccggtagcta cgtgtgctac tataagtaca tcaaggctag aatcgagggc 1200 accaccgccg ctagcagcta cgtgtttgtg agagacttcg agcagccctt catcaacaag 1260 cccgacaccc tgctggtgaa cagaaaggac gccatgtggg tgccctgcct ggtgagcatc 1320 cccggcctga acgtgaccct gagatctcag agcagcgtgc tgtggcccga cggccaagag 1380 gtggtgtggg acgacagaag aggcatgctg gtgagcaccc ccctgctgca cgacgccctg 1440 tacctgcagt gcgagaccac ctggggcgac caagacttcc tgagcaaccc cttcctggtg 1500 cacatcacag gcgacaaaac tcatacctgc ccaccttgtc cagcaccaga gctgctcgga 1560 ggaccatccg tgttcctgtt tccacccaag cccaaagata ctctgatgat ttcacgcaca 1620 cccgaagtca cttgcgtggt cgtggacgtg tcccacgagg accccgaagt caagtttaac 1680 tggtacgtgg acggcgtcga ggtgcataat gctaagacaa aaccccgaga ggaacagtac 1740 aactctacct atagggtcgt gagtgtcctg acagtgctcc accaggattg gctgaacgga 1800 aaggagtata agtgcaaagt gtctaataag gcactgcctg ccccaatcga gaaaacaatt 1860 agtaaggcca aagggcagcc cagagaacct caggtgtaca ctctgcctcc atctcgggac 1920 gagctcacta agaaccaggt cagtctgacc tgtctcgtga aagggttcta tcctagtgat 1980 atcgctgtgg agtgggaatc aaatggtcag ccagagaaca attacaagac cacaccccct 2040 gtcctggaca gcgatggctc cttctttctg tattccaagc tcaccgtgga caaatctcga 2100 tggcagcagg gaaacgtctt tagttgttca gtgatgcacg aagccctcca taaccactac 2160 actcagaaaa gcctcagcct cagccctggg aaatga 2196 <210> 71 <211> 2178 <212> DNA <213> Artificial Sequence <220> <223> Nucleic acid of exemplary polypeptide 9 (such as in EXG102-28) <400> 71 atggtgtctt attgggatac tggcgtgctg ctctgtgccc tcctgagttg cctgctcctg 60 actggttctt cttctgggtc cgatactggg cgccccttcg tggagatgta ctccgagatc 120 cctgaaatca ttcacatgac tgagggtcgg gaactggtca tcccatgccg cgtgacctct 180 cccaacatta ctgtgaccct gaagaaattc cctctggaca ccctcatccc agatgggaag 240 aggatcattt gggactcaag aaagggtttt atcatcagca acgctacata caaggagatt 300 ggcctgctca cctgcgaagc aacagtgaac ggacacctgt acaagactaa ttatctcacc 360 catagacaga caaacactat cattgatgtg gtcctgtcac caagccacgg catcgagctc 420 agcgtcggtg aaaagctggt gctcaattgt acagcccgga ctgagctgaa cgtgggcatt 480 gacttcaatt gggaataccc cagctccaag caccagcata agaaactggt gaaccgcgat 540 ctcaaaaccc agtccggatc tgagatgaag aaatttctga gcaccctcac aatcgacggc 600 gtgacacgat ccgatcaggg actgtatact tgcgccgctt ctagtggcct gatgaccaag 660 aaaaatagca cattcgtcag ggtgcacgaa aagggcggcg gcggaggcag ctactccatg 720 acccccccga ccctgaacat caccgaggag agccacgtga tcgacaccgg cgacagcctg 780 agcatcagtt gccggggcca acaccccctg gaatgggcat ggcctggggc acaagaggcc 840 cccgcaacgg gcgacaagga cagcgaggac accggcgtgg tgagagactg cgagggcacc 900 gacgctagac cctactgcaa ggtgctgctg ctgcacgagg tgcacgccca agacaccggt 960 agctacgtgt gctactataa gtacatcaag gctagaatcg agggcaccac cgccgctagc 1020 agctacgtgt ttgtgagaga cttcgagcag cccttcatca acaagcccga caccctgctg 1080 gtgaacagaa aggacgccat gtgggtgccc tgcctggtga gcatccccgg cctgaacgtg 1140 accctgagat ctcagagcag cgtgctgtgg cccgacggcc aagaggtggt gtgggacgac 1200 agaagaggca tgctggtgag cacccccctg ctgcacgacg ccctgtacct gcagtgcgag 1260 accacctggg gcgaccaaga cttcctgagc aaccccttcc tggtgcacat caccggcgac 1320 aaaactcata cctgcccacc ttgtccagca ccagagctgc tcggaggacc atccgtgttc 1380 ctgtttccac ccaagcccaa agatactctg atgatttcac gcacacccga agtcacttgc 1440 gtggtcgtgg acgtgtccca cgaggacccc gaagtcaagt ttaactggta cgtggacggc 1500 gtcgaggtgc ataatgctaa gacaaaaccc cgagaggaac agtacaactc tacctatagg 1560 gtcgtgagtg tcctgacagt gctccaccag gattggctga acggaaagga gtataagtgc 1620 aaagtgtcta ataaggcact gcctgcccca atcgagaaaa caattagtaa ggccaaaggg 1680 cagcccagag aacctcaggt gtacactctg cctccatctc gggacgagct cactaagaac 1740 caggtcagtc tgacctgtct cgtgaaaggg ttctatccta gtgatatcgc tgtggagtgg 1800 gaatcaaatg gtcagccaga gaacaattac aagaccacac cccctgtcct ggacagcgat 1860 ggctccttct ttctgtattc caagctcacc gtggacaaat ctcgatggca gcagggaaac 1920 gtctttagtt gttcagtgat gcacgaagcc ctccataacc actacactca gaaaagcctc 1980 agcctcagcc ctgggaaagg cggcggcgga ggcgcccagc aagaagagtg cgagtgggat 2040 ccctggacct gcgagcacat gggatccggc agcgccaccg gaggatccgg aagcaccgcc 2100 tccagcggct ccggcagcgc cacccaccag gaggagtgtg agtgggaccc ctggacctgc 2160 gaacacatgc tggagtga 2178 <210> 72 <211> 2211 <212> DNA <213> Artificial Sequence <220> <223> Nucleic acid of exemplary polypeptide 10 (such as in EXG102-29) <400> 72 atggtgtctt attgggatac tggcgtgctg ctctgtgccc tcctgagttg cctgctcctg 60 actggttctt cttctgggtc cgatactggg cgccccttcg tggagatgta ctccgagatc 120 cctgaaatca ttcacatgac tgagggtcgg gaactggtca tcccatgccg cgtgacctct 180 cccaacatta ctgtgaccct gaagaaattc cctctggaca ccctcatccc agatgggaag 240 aggatcattt gggactcaag aaagggtttt atcatcagca acgctacata caaggagatt 300 ggcctgctca cctgcgaagc aacagtgaac ggacacctgt acaagactaa ttatctcacc 360 catagacaga caaacactat cattgatgtg gtcctgtcac caagccacgg catcgagctc 420 agcgtcggtg aaaagctggt gctcaattgt acagcccgga ctgagctgaa cgtgggcatt 480 gacttcaatt gggaataccc cagctccaag caccagcata agaaactggt gaaccgcgat 540 ctcaaaaccc agtccggatc tgagatgaag aaatttctga gcaccctcac aatcgacggc 600 gtgacacgat ccgatcaggg actgtatact tgcgccgctt ctagtggcct gatgaccaag 660 aaaaatagca cattcgtcag ggtgcacgaa aagggcggcg gcggaggcag ctactccatg 720 acccccccga ccctgaacat caccgaggag agccacgtga tcgacaccgg cgacagcctg 780 agcatcagtt gccggggcca acaccccctg gaatgggcat ggcctggggc acaagaggcc 840 cccgcaacgg gcgacaagga cagcgaggac accggcgtgg tgagagactg cgagggcacc 900 gacgctagac cctactgcaa ggtgctgctg ctgcacgagg tgcacgccca agacaccggt 960 agctacgtgt gctactataa gtacatcaag gctagaatcg agggcaccac cgccgctagc 1020 agctacgtgt ttgtgagaga cttcgagcag cccttcatca acaagcccga caccctgctg 1080 gtgaacagaa aggacgccat gtgggtgccc tgcctggtga gcatccccgg cctgaacgtg 1140 accctgagat ctcagagcag cgtgctgtgg cccgacggcc aagaggtggt gtgggacgac 1200 agaagaggca tgctggtgag cacccccctg ctgcacgacg ccctgtacct gcagtgcgag 1260 accacctggg gcgaccaaga cttcctgagc aaccccttcc tggtgcacat caccggcgac 1320 aaaactcata cctgcccacc ttgtccagca ccagagctgc tcggaggacc atccgtgttc 1380 ctgtttccac ccaagcccaa agatactctg atgatttcac gcacacccga agtcacttgc 1440 gtggtcgtgg acgtgtccca cgaggacccc gaagtcaagt ttaactggta cgtggacggc 1500 gtcgaggtgc ataatgctaa gacaaaaccc cgagaggaac agtacaactc tacctatagg 1560 gtcgtgagtg tcctgacagt gctccaccag gattggctga acggaaagga gtataagtgc 1620 aaagtgtcta ataaggcact gcctgcccca atcgagaaaa caattagtaa ggccaaaggg 1680 cagcccagag aacctcaggt gtacactctg cctccatctc gggacgagct cactaagaac 1740 caggtcagtc tgacctgtct cgtgaaaggg ttctatccta gtgatatcgc tgtggagtgg 1800 gaatcaaatg gtcagccaga gaacaattac aagaccacac cccctgtcct ggacagcgat 1860 ggctccttct ttctgtattc caagctcacc gtggacaaat ctcgatggca gcagggaaac 1920 gtctttagtt gttcagtgat gcacgaagcc ctccataacc actacactca gaaaagcctc 1980 agcctcagcc ctgggaaagg gggcgggggc ggcgctcaag gcagcgggag cgctaccggc 2040 ggctccggga gcacagctag cagcggcagc ggctccgcta cccacaagtt caaccccctg 2100 gacgagctgg aggagaccct ctacgagcag ttcaccttcc aacaaggcgg gggcgggggc 2160 caagaggagt gcgagtggga cccctggacc tgcgagcaca tgctggagtg a 2211 <210> 73 <211> 1605 <212> DNA <213> Artificial sequence <220> <223> Nucleic acid of exemplary polypeptide 11 (such as in EXG102 - 30) <400> 73 atggtgtctt attgggatac tggcgtgctg ctctgtgccc tcctgagttg cctgctcctg 60 actggttctt cttctggggg gggcgggggc ggcgctcaag gcagcgggag cgctaccggc 120 ggctccggga gcacagctag cagcggcagc ggctccgcta cccacaagtt caaccccctg 180 gacgagctgg aggagaccct ctacgagcag ttcaccttcc aacaaggcgg gggcgggggc 240 caagaggagt gcgagtggga cccctggacc tgcgagcaca tgctggaggg cggcggcgga 300 ggcagctccg atactgggcg ccccttcgtg gagatgtact ccgagatccc tgaaatcatt 360 cacatgactg agggtcggga actggtcatc ccatgccgcg tgacctctcc caacattact 420 gtgaccctga agaaattccc tctggacacc ctcatcccag atgggaagag gatcatttgg 480 gactcaagaa agggttttat catcagcaac gctacataca aggagattgg cctgctcacc 540 tgcgaagcaa cagtgaacgg acacctgtac aagactaatt atctcaccca tagacagaca 600 aacactatca ttgatgtggt cctgtcacca agccacggca tcgagctcag cgtcggtgaa 660 aagctggtgc tcaattgtac agcccggact gagctgaacg tgggcattga cttcaattgg 720 gaatacccca gctccaagca ccagcataag aaactggtga accgcgatct caaaacccag 780 tccggatctg agatgaagaa atttctgagc accctcacaa tcgacggcgt gacacgatcc 840 gatcagggac tgtatacttg cgccgcttct agtggcctga tgaccaagaa aaatagcaca 900 ttcgtcaggg tgcacgaaaa ggacaaaact catacctgcc caccttgtcc agcaccagag 960 ctgctcggag gaccatccgt gttcctgttt ccacccaagc ccaaagatac tctgatgatt 1020 tcacgcacac ccgaagtcac ttgcgtggtc gtggacgtgt cccacgagga ccccgaagtc 1080 aagtttaact ggtacgtgga cggcgtcgag gtgcataatg ctaagacaaa accccgagag 1140 gaacagtaca actctaccta tagggtcgtg agtgtcctga cagtgctcca ccaggattgg 1200 ctgaacggaa aggagtataa gtgcaaagtg tctaataagg cactgcctgc cccaatcgag 1260 aaaacaatta gtaaggccaa agggcagccc agagaacctc aggtgtacac tctgcctcca 1320 tctcgggacg agctcactaa gaaccaggtc agtctgacct gtctcgtgaa agggttctat 1380 cctagtgata tcgctgtgga gtgggaatca aatggtcagc cagagaacaa ttacaagacc 1440 acaccccctg tcctggacag cgatggctcc ttctttctgt attccaagct caccgtggac 1500 aaatctcgat ggcagcaggg aaacgtcttt agttgttcag tgatgcacga agccctccat 1560 aaccactaca ctcagaaaag cctcagcctc agccctggga aatga 1605 <210> 74 <211> 2229 <212> DNA <213> Artificial Sequence <220> <223> Nucleic acid of exemplary polypeptide 12 (such as in EXG102-31) <400> 74 atggtgtctt attgggatac tggcgtgctg ctctgtgccc tcctgagttg cctgctcctg 60 actggttctt cttctggggg gggcgggggc ggcgctcaag gcagcgggag cgctaccggc 120 ggctccggga gcacagctag cagcggcagc ggctccgcta cccacaagtt caaccccctg 180 gacgagctgg aggagaccct ctacgagcag ttcaccttcc aacaaggcgg gggcgggggc 240 caagaggagt gcgagtggga cccctggacc tgcgagcaca tgctggaggg cggcggcgga 300 ggcagctccg atactgggcg ccccttcgtg gagatgtact ccgagatccc tgaaatcatt 360 cacatgactg agggtcggga actggtcatc ccatgccgcg tgacctctcc caacattact 420 gtgaccctga agaaattccc tctggacacc ctcatcccag atgggaagag gatcatttgg 480 gactcaagaa agggttttat catcagcaac gctacataca aggagattgg cctgctcacc 540 tgcgaagcaa cagtgaacgg acacctgtac aagactaatt atctcaccca tagacagaca 600 aacactatca ttgatgtggt cctgtcacca agccacggca tcgagctcag cgtcggtgaa 660 aagctggtgc tcaattgtac agcccggact gagctgaacg tgggcattga cttcaattgg 720 gaatacccca gctccaagca ccagcataag aaactggtga accgcgatct caaaacccag 780 tccggatctg agatgaagaa atttctgagc accctcacaa tcgacggcgt gacacgatcc 840 gatcagggac tgtatacttg cgccgcttct agtggcctga tgaccaagaa aaatagcaca 900 ttcgtcaggg tgcacgaaaa gggcggcggc ggaggcagct actccatgac ccccccgacc 960 ctgaacatca ccgaggagag ccacgtgatc gacaccggcg acagcctgag catcagttgc 1020 cggggccaac accccctgga atgggcatgg cctggggcac aagaggcccc cgcaacgggc 1080 gacaaggaca gcgaggacac cggcgtggtg agagactgcg agggcaccga cgctagaccc 1140 tactgcaagg tgctgctgct gcacgaggtg cacgccaacg acaccggtag ctacgtgtgc 1200 tactataagt acatcaaggc tagaatcgag ggcaccaccg ccgctagcag ctacgtgttt 1260 gtgagagact tcgagcagcc cttcatcaac aagcccgaca ccctgctggt gaacagaaag 1320 gacgccatgt gggtgccctg cctggtgagc atccccggcc tgaacgtgac cctgagatct 1380 cagagcagcg tgctgtggcc cgacggccaa gaggtggtgt gggacgacag aagaggcatg 1440 ctggtgagca cccccctgct gcacgacgcc ctgtacctgc agtgcgagac cacctggggc 1500 gaccaagact tcctgagcaa ccccttcctg gtgcacatca caggcgacaa aactcatacc 1560 tgcccacctt gtccagcacc agagctgctc ggaggaccat ccgtgttcct gtttccaccc 1620 aagcccaaag atactctgat gatttcacgc acacccgaag tcacttgcgt ggtcgtggac 1680 gtgtcccacg aggaccccga agtcaagttt aactggtacg tggacggcgt cgaggtgcat 1740 aatgctaaga caaaaccccg agaggaacag tacaactcta cctatagggt cgtgagtgtc 1800 ctgacagtgc tccaccagga ttggctgaac ggaaaggagt ataagtgcaa agtgtctaat 1860 aaggcactgc ctgccccaat cgagaaaaca attagtaagg ccaaagggca gcccagagaa 1920 cctcaggtgt acactctgcc tccatctcgg gacgagctca ctaagaacca ggtcagtctg 1980 acctgtctcg tgaaagggtt ctatcctagt gatatcgctg tggagtggga atcaaatggt 2040 cagccagaga acaattacaa gaccacaccc cctgtcctgg acagcgatgg ctccttcttt 2100 ctgtattcca agctcaccgt ggacaaatct cgatggcagc agggaaacgt ctttagttgt 2160 tcagtgatgc acgaagccct ccataaccac tacactcaga aaagcctcag cctcagccct 2220 gggaaatga 2229 <210> 75 <211> 2211 <212> DNA <213> Artificial Sequence <220> <223> Nucleic acid of exemplary polypeptide 13 (such as in EXG102 - 32) <400> 75 atggtgtctt attgggatac tggcgtgctg ctctgtgccc tcctgagttg cctgctcctg 60 actggttctt cttctgggtc cgatactggg cgccccttcg tggagatgta ctccgagatc 120 cctgaaatca ttcacatgac tgagggtcgg gaactggtca tcccatgccg cgtgacctct 180 cccaacatta ctgtgaccct gaagaaattc cctctggaca ccctcatccc agatgggaag 240 aggatcattt gggactcaag aaagggtttt atcatcagca acgctacata caaggagatt 300 ggcctgctca cctgcgaagc aacagtgaac ggacacctgt acaagactaa ttatctcacc 360 catagacaga caaacactat cattgatgtg gtcctgtcac caagccacgg catcgagctc 420 agcgtcggtg aaaagctggt gctcaattgt acagcccgga ctgagctgaa cgtgggcatt 480 gacttcaatt gggaataccc cagctccaag caccagcata agaaactggt gaaccgcgat 540 ctcaaaaccc agtccggatc tgagatgaag aaatttctga gcaccctcac aatcgacggc 600 gtgacacgat ccgatcaggg actgtatact tgcgccgctt ctagtggcct gatgaccaag 660 aaaaatagca cattcgtcag ggtgcacgaa aagggcggcg gcggaggcag ctactccatg 720 acccccccga ccctgaacat caccgaggag agccacgtga tcgacaccgg cgacagcctg 780 agcatcagtt gccggggcca acaccccctg gaatgggcat ggcctggggc acaagaggcc 840 cccgcaacgg gcgacaagga cagcgaggac accggcgtgg tgagagactg cgagggcacc 900 gacgctagac cctactgcaa ggtgctgctg ctgcacgagg tgcacgccaa cgacaccggt 960 agctacgtgt gctactataa gtacatcaag gctagaatcg agggcaccac cgccgctagc 1020 agctacgtgt ttgtgagaga cttcgagcag cccttcatca acaagcccga caccctgctg 1080 gtgaacagaa aggacgccat gtgggtgccc tgcctggtga gcatccccgg cctgaacgtg 1140 accctgagat ctcagagcag cgtgctgtgg cccgacggcc aagaggtggt gtgggacgac 1200 agaagaggca tgctggtgag cacccccctg ctgcacgacg ccctgtacct gcagtgcgag 1260 accacctggg gcgaccaaga cttcctgagc aaccccttcc tggtgcacat cacaggcgac 1320 aaaactcata cctgcccacc ttgtccagca ccagagctgc tcggaggacc atccgtgttc 1380 ctgtttccac ccaagcccaa agatactctg atgatttcac gcacacccga agtcacttgc 1440 gtggtcgtgg acgtgtccca cgaggacccc gaagtcaagt ttaactggta cgtggacggc 1500 gtcgaggtgc ataatgctaa gacaaaaccc cgagaggaac agtacaactc tacctatagg 1560 gtcgtgagtg tcctgacagt gctccaccag gattggctga acggaaagga gtataagtgc 1620 aaagtgtcta ataaggcact gcctgcccca atcgagaaaa caattagtaa ggccaaaggg 1680 cagcccagag aacctcaggt gtacactctg cctccatctc gggacgagct cactaagaac 1740 caggtcagtc tgacctgtct cgtgaaaggg ttctatccta gtgatatcgc tgtggagtgg 1800 gaatcaaatg gtcagccaga gaacaattac aagaccacac cccctgtcct ggacagcgat 1860 ggctccttct ttctgtattc caagctcacc gtggacaaat ctcgatggca gcagggaaac 1920 gtctttagtt gttcagtgat gcacgaagcc ctccataacc actacactca gaaaagcctc 1980 agcctcagcc ctgggaaagg gggcgggggc ggcgctcaag gcagcgggag cgctaccggc 2040 ggctccggga gcacagctag cagcggcagc ggctccgcta cccacaagtt caaccccctg 2100 gacgagctgg aggagaccct ctacgagcag ttcaccttcc aacaaggcgg gggcgggggc 2160 caagaggagt gcgagtggga cccctggacc tgcgagcaca tgctggagtg a 2211 4. Description of the Drawings
[0039] Figure 1 Shows the exemplary fusion protein construct provided herein (Flt1 signal: signal peptide sequence of VEGFR1 (Flt1); VEGFR1 D2: IgG-like domain 2 of VEGF receptor 1; VEGFR2 D3: IgG-like domain 3 of VEGF receptor 2; Ang BD: angiopoietin binding domain; IgG Fc: Fc fragment of human IgG).
[0040] Figure 2 Shows the exemplary rAAV vector provided herein. TR represents the AAV2 terminal inverted repeat, CBA is the 1.68 kb chicken β-actin promoter, CB is the 0.78 kb chick β-actin promoter, D2 represents the IgG-like domain 2 of VEGFR-1, D3 represents the IgG-like domain 3 of VEGFR-2, ABD (or Ang BD) represents the angiopoietin binding domain, IgG Fc is the Fc fragment of human IgG, Fc-hinge is the N-terminal 21 amino acids of Fc plus a 6 amino acid GS linker, WPRE represents the woodchuck hepatitis virus post-transcriptional regulatory element (600 bp), mini mWPRE is WPRE (240 bp), SV40pA is the simian virus 40 polyadenylation signal and bGHpA is the bovine growth hormone polyadenylation signal.
[0041] Figures 3A - 3B Shows the experimental design and assay for testing the expression and binding affinity of the constructs provided herein.
[0042] Figures 4A - 4B Shows that the position of Ang BD at the C-terminus does not affect the binding ability to rVEGF, but reduces the transgene expression level.
[0043] Figures 5A - 5B Shows that the position of Ang BD in the middle between the VEGF binding domain and the Fc region does not affect the binding ability to rVEGF and transgene expression.
[0044] Figures 6A - 6B Shows that the position of Ang BD in the middle between the VEGF binding domain and the Fc region inhibits the binding ability to Ang2. In Figure 6B In the upper panel, the right column of each of the two pairs represents the binding to EXG102-09, and the left column of each of the two pairs represents the binding to EXG102-04.
[0045] Figures 7A - 7D Shows that the construct with Ang BD at the N-terminus strongly binds to both VEGF and Ang2 while maintaining a high transgene expression level. In Figure 7DAmong them, for each EXG construct, the left column represents binding to rVEGF, the middle column represents binding to Ang2, and the right column represents the protein expression level.
[0046] Figure 8 Shows a schematic design of other exemplary fusion protein constructs provided herein that contain a VEGF-C binding domain (Trap-C1, C2, or C3) and an angiopoietin binding domain (ABD or ABD2). TR, terminal repeat; CBA, chimeric CMV - chicken β-actin promoter; ABD, angiopoietin binding domain; D2, IgG-like domain 2 of VEGF receptor 1; D3, IgG-like domain 3 of VEGF receptor 2; Trap C, VEGFC binding domain; Fc, crystallizable fragment of IgG1; bGHpA, bovine growth hormone polyadenylation signal; VEGF, vascular endothelial growth factor; IgG1, immunoglobulin G1.
[0047] Figures 9A - 9I Show EXG102-24, EXG102-25, EXG102-26, EXG102-27, EXG102-28, EXG102-29, EXG102-30, EXG102-31, and EXG102-32, respectively. The sequences are shown and numbered ( Figures 9A - 9I The fusion proteins shown in contain SEQ ID NOs: 58-66, respectively. The signal peptide, ABD, D2, D3, Trap C, and Fc constructs are also shown in the figure. The GS linker is highlighted. TR, terminal repeat; CBA, chimeric CMV - chicken β-actin promoter; ABD, angiopoietin binding domain; D2, IgG-like domain 2 of VEGF receptor 1; D3, IgG-like domain 3 of VEGF receptor 2; Trap C, VEGFC binding domain; Fc, crystallizable fragment of IgG1; bGHpA, bovine growth hormone polyadenylation signal; VEGF, vascular endothelial growth factor; IgG1, immunoglobulin G1.
[0048] Figures 10A - 10F Show that constructs containing ABD2 are comparable to constructs with an ABD domain in terms of protein expression, VEGF-A binding, or Ang2 binding.
[0049] Figure 11Shows the mean FFA scores in a laser-induced choroidal neovascularization (CNV) mouse model. On days 29 and 36 after injection, the mean FFA scores were measured in eyes treated with AAV-GFP (negative control), EXG102-02 (positive control), EXG102-04, EXG102-09, EXG102-10, or EXG102-11. As shown below, angiograms were graded: score 0, no staining; score 1, slight staining; score 2, moderate staining; and score 3, strong staining. The two constructs with the highest CNV inhibition were highlighted by arrows. Bars represent the mean with SD.
[0050] Figure 12 Shows the proportion of grade 3 CNV lesions in a laser-induced CNV mouse model. The mean FFA scores were measured in eyes treated with AAV-GFP (negative control), EXG102-02 (positive control), EXG102-04, EXG102-09, EXG102-10, or EXG102-11. As shown below, angiograms were graded: score 0, no staining; score 1, slight staining; score 2, moderate staining; and score 3, strong staining. Bars represent the mean with SD.
[0051] Figures 13A - 13C Shows the in vitro VEGF-A, Ang2, or VEGF-C binding affinities of transgenic products expressed from HEK293T cells transfected with plasmid DNA. HEK293T cells were transfected with plasmid DNA of pEXG102-02, pEXG102-30, or pEXG102-31 and the expressed target proteins were affinity-purified from cell lysates. The binding abilities of EXG102-02, EXG102-30, or EXG102-31 to VEGF-A, Ang2, or VEGF-C, respectively, were measured by ELISA. Bars represent the mean with SD.
[0052] Figure 14 Shows the mean FFA scores in a laser-induced CNV mouse model. The mean FFA scores were measured in eyes treated with vehicle control (black), EXG102-02 (medium gray), EXG102-30 (dark gray), or EXG102-31 (light gray). Before fluorescein angiography, animals were given an intraperitoneal injection of sodium fluorescein (100 mg / ml, 30 μL / animal). FFA images were acquired 3 minutes after injection. As shown below, angiograms were graded: score 0, no staining; score 1, slight staining; score 2, moderate staining; and score 3, strong staining. The proportion of lesions with score 3 and the mean score were calculated. Bars represent the mean with SD.
[0053] Figure 15Shows the proportion of grade 3 CNV lesions in a laser-induced CNV mouse model. The proportion of grade 3 CNV lesions was measured in eyes treated with vehicle control, EXG102-02, EXG102-30, or EXG102-31. Prior to fluorescein angiography, animals were given an intraperitoneal injection of sodium fluorescein (100 mg / ml, 30 μL / animal). FFA images were acquired 3 minutes after injection. The angiograms were graded as follows: score 0, unstained; score 1, slightly stained; score 2, moderately stained; and score 3, strongly stained. The proportion of lesions with a score of 3 and the mean score were calculated. Bars represent the mean with SD. 5. DETAILED DESCRIPTION OF THE INVENTION
[0054] The disclosure of the present invention is based in part on novel fusion proteins comprising domains from VEGFR-1, VEGFR-2, and / or VEGFR-3 and an angiopoietin-binding domain, AAV vectors comprising nucleic acids encoding these fusion proteins, and their improved properties.
[0055] 5.1. Definitions
[0056] The techniques and procedures described or referenced herein include the use of conventional methods, those well understood and / or commonly used by those of ordinary skill in the art, such as (e.g.) Sambrook et al., Molecular Cloning: A Laboratory Manual (3rd ed. 2001); Current Protocols in Molecular Biology (Ausubel et al., eds., 2003); Therapeutic Monoclonal Antibodies: From Bench to Clinic (An, ed., 2009); Monoclonal Antibodies :Methods and Protocols (Albitar, ed., 2010); and Antibody Engineering widely used methods described in Volumes 1 and 2 (Kontermann and Dübel, eds., 2nd ed. 2010).
[0057] Unless defined otherwise herein, technical and scientific terms used in the description of the present invention have the meanings commonly understood by those of ordinary skill in the art. For the purposes of interpreting this specification, the following description of terms will apply, and where appropriate, terms used in the singular will also include the plural and vice versa. If any description of a term conflicts with any document incorporated herein by reference, the description of the term set forth below shall control.
[0058] In this text, the terms "polypeptide", "peptide", and "protein" are used interchangeably and denote polymers of amino acids of any length. The polymers can be linear or branched, they can contain modified amino acids, and they can be interrupted by non-amino acids. The terms also encompass amino acid polymers that have been modified either naturally or by intervention; for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification. For example, polypeptides containing one or more analogs of amino acids, including (but not limited to) non-natural amino acids and other modified polypeptides known in the art, are also included within this definition.
[0059] The term "binds" or "binding" denotes an interaction between molecules, including (for example) to form a complex. The interaction can be (for example) a non-covalent interaction, including hydrogen bonds, ionic bonds, hydrophobic interactions, and / or van der Waals interactions. A complex can also include the association of two or more molecules held together by covalent or non-covalent bonds, interactions, or forces. The strength of the total non-covalent interaction between two molecules is the affinity of one molecule for the other. The dissociation rate (k off ) of a binding molecule from another molecule, relative to the association rate (k on ), is the dissociation constant K off / k on ), which is inversely related to the affinity. The smaller the K D value, the greater the affinity. Based on both k D and k on and k off ), the value of K D varies for different complexes. The dissociation constant K D can be determined using any method provided herein or any other method known to those skilled in the art.
[0060] In connection with the binding molecules described herein, terms such as "binds to", "specifically binds to", and similar terms are also used interchangeably herein and denote that a binding domain specifically binds to another molecule, such as a polypeptide. A binding molecule or binding domain that binds to or specifically binds to another molecule can be identified (for example) by immunoassays, or other techniques known to those skilled in the art. In some embodiments, as determined using experimental techniques such as radioimmunoassay (RIA) and enzyme-linked immunosorbent assay (ELISA), a binding molecule or binding domain binds to or specifically binds to a molecule when it binds to the molecule with an affinity higher than that of any cross-reactive molecule. Typically, a specific or selective reaction will be at least twice the background signal or noise and can be greater than 10-fold the background. For a discussion of binding specificity, see, for example, Fundamental Immunology332 - 36 (edited by Paul, 2nd edition, 1989). In certain embodiments, the degree of binding of a binding molecule or binding domain to a "non - target" protein is less than about 10% of the binding of the binding molecule or binding domain to its specific target, e.g., as determined by fluorescence - activated cell sorting (FACS) analysis or RIA. For terms such as "specifically binds", "specifically binds to", or "is specific for" denote binding that is measurably different from non - specific interactions. Specific binding can be measured, for example, by determining the binding of a molecule compared to the binding of a control molecule, which is typically a molecule of similar structure that does not have binding activity. For example, specific binding can be determined by competition with a control molecule similar to the target, e.g., an excess of unlabeled target. In this case, if the binding of a labeled target to a probe is competitively inhibited by an excess of unlabeled target, it indicates specific binding. A binding molecule or binding domain that binds to a molecule includes a binding molecule or binding domain that can bind to the molecule with an affinity sufficient to render the binding molecule useful, e.g., as a diagnostic or therapeutic agent in targeting the molecule. In certain embodiments, the binding molecule or binding domain that binds to a target molecule has a dissociation constant (K D ) less than or equal to 1 μM, 800 nM, 600 nM, 550 nM, 500 nM, 300 nM, 250 nM, 100 nM, 50 nM, 10 nM, 5 nM, 4 nM, 3 nM, 2 nM, 1 nM, 0.9 nM, 0.8 nM, 0.7 nM, 0.6 nM, 0.5 nM, 0.4 nM, 0.3 nM, 0.2 nM, or 0.1 nM.
[0061] "Binding affinity" generally refers to the strength of the sum of non - covalent interactions between a single binding site of a molecule and its binding partner. Unless otherwise stated, as used herein, "binding affinity" refers to the intrinsic binding affinity, which reflects the 1:1 interaction between the members of a binding pair. Generally, the affinity of a binding molecule X for its binding partner Y can be represented by the dissociation constant (K D ). Affinity can be measured by common methods known in the art, including those described herein. A variety of methods for measuring binding affinity are known in the art, and any of them can be used for the purposes disclosed in the present invention. Specific illustrative embodiments include the following. In one embodiment, "K D " or "K D value" can be measured by assays known in the art, e.g., by a binding assay. It can also be measured by using biolayer interferometry (BLI) or surface plasmon resonance (SPR) assays, by using, for example, of the Red96 system or by using, for example, TM-2000 or of TM-3000 Measure K D or K D value. The “binding rate” (on-rate or rate of association or association rate) or “kon” can also be determined by the same biolayer interferometry (BLI) or surface plasmon resonance (SPR) techniques described above, using, for example, Red96, TM-2000 or TM-3000 system.
[0062] The terms “antibody,” “immunoglobulin,” or “Ig” are used interchangeably herein and are used in the broadest sense and specifically cover, for example, monoclonal antibodies (including agonists, antagonists, neutralizing antibodies, full-length or intact monoclonal antibodies), antibody compositions having multi-epitope or mono-epitope specificity, polyclonal or monovalent antibodies, multivalent antibodies, multispecific antibodies formed from at least two intact antibodies (e.g., bispecific antibodies, provided they exhibit the desired biological activity), single-chain antibodies and fragments thereof (e.g., domain antibodies), as described below. Antibodies can be human, humanized, chimeric, and / or affinity matured, as well as antibodies from other species, e.g., mice, rabbits, llamas, etc. The term “antibody” is intended to include the polypeptide products of B cells within the immunoglobulin type of polypeptides that are capable of binding to a specific molecular antigen and are composed of the same pair of polypeptide chains, where each pair has one heavy chain (about 50-70 kDa) and one light chain (about 25 kDa), and each amino-terminal portion of each chain includes a variable region of about 100 to about 130 or more amino acids and each carboxyl-terminal portion of each chain includes a constant region. See, for example, Antibody Engineering (edited by Borrebaeck, 2nd ed. 1995); and Kuby, Immunology(3rd Edition 1997). In a specific embodiment, the antibodies provided herein can include polypeptides or epitope-binding specific molecular antigens. Antibodies also include (but are not limited to) synthetic antibodies, recombinantly produced antibodies, single-domain antibodies, including antibodies from species of the genus Camelus (e.g., llama or alpaca) or their humanized variants, intracellular antibodies, anti-idiotypic (anti-Id) antibodies, and any functional fragments of any of the above (e.g., antigen-binding fragments), which refer to a portion of an antibody heavy or light chain polypeptide that retains some or all of the binding activity of the antibody from which the fragment is derived. Non-limiting examples of functional fragments (e.g., antigen-binding fragments) include single-chain Fv (scFv) (e.g., including monospecific, bispecific, etc.), Fab fragments, F(ab') fragments, F(ab) 2 fragments, F(ab') 2 fragments, disulfide-linked Fv (sdFv), Fd fragments, Fv fragments, diabodies, triabodies, tetra-bodies, and microantibodies. Specifically, the antibodies provided herein include immunoglobulin molecules and immunologically active portions of immunoglobulin molecules, e.g., antigen-binding domains or molecules containing antigen-binding sites that bind to an antigen (e.g., one or more CDRs of an antibody). Such antibody fragments can be found in, e.g., Harlow and Lane, Antibodies:A Laboratory Manual (1989); Mol.Biology and Biotechnology:A Comprehensive Desk Reference (ed. Myers, 1995); Huston et al., 1993, Cell Biophysics 22:189-224; Plückthun and Skerra, 1989, Meth. Enzymol. 178:497-515; and Day, Advanced Immunochemistry (2nd Edition 1990). The antibodies provided herein can be of any type of immunoglobulin molecule (e.g., IgG, IgE, IgM, IgD, and IgA) or any subclass (e.g., IgG1, IgG2, IgG3, IgG4, IgA1, and IgA2). The antibodies can be agonistic antibodies or antagonistic antibodies. The antibodies can be neither agonistic nor antagonistic.
[0063] As used herein, the term "Fc region" is used to define the C-terminal region of an immunoglobulin heavy chain, which includes, for example, a native sequence Fc region, a recombinant Fc region, and a variant Fc region. Although the boundaries of the Fc region of an immunoglobulin heavy chain may vary, the human IgG heavy chain Fc region is generally defined to extend from the amino acid residue at position Cys226 or Pro230 to its carboxyl terminus. For example, the C-terminal lysine of the Fc region (residue 447 according to the EU numbering system) can be removed during the production or purification of an antibody or by recombinant engineering of the nucleic acid encoding the antibody heavy chain. Thus, a composition of intact antibodies can include a population of antibodies in which all K447 residues have been removed, a population of antibodies in which the K447 residues have not been removed, and a population of antibodies that is a mixture of antibodies with and without the K447 residue. A "functional Fc region" has the "effector functions" of a native sequence Fc region. Exemplary "effector functions" include C1q binding; CDC; Fc receptor binding; ADCC; phagocytosis; downregulation of cell surface receptors (e.g., B cell receptors), etc. These effector functions generally require the Fc region to be combined with a binding region or domain (e.g., an antibody variable region or domain) and can be evaluated using a variety of assays known to those of skill in the art. A "variant Fc region" includes an amino acid sequence that differs from the amino acid sequence of a native sequence Fc region by virtue of at least one amino acid modification (e.g., substitution, addition, or deletion). In certain embodiments, the variant Fc region has at least one amino acid substitution in the native sequence Fc region or in the Fc region of a parental polypeptide, for example, from about 1 to about 10 amino acid substitutions, or from about 1 to about 5 amino acid substitutions, compared to the native sequence Fc region or the Fc region of the parental polypeptide. As used herein, a variant Fc region can have at least about 80% homology with a native sequence Fc region and / or the Fc region of a parental polypeptide, or at least about 90% homology therewith, for example, at least about 95% homology therewith.
[0064] "Polynucleotide" or "nucleic acid", as used interchangeably herein, refers to a polymer of nucleotides of any length and includes DNA and RNA. The nucleotides can be deoxyribonucleotides, ribonucleotides, modified nucleotides or bases and / or their analogs, or any substrate that can be introduced into the polymer by DNA or RNA polymerase or by a synthetic reaction. Polynucleotides can include modified nucleotides, such as methylated nucleotides and their analogs. As used herein, "oligonucleotide" refers to short, usually single-stranded, synthetic polynucleotides, which are usually, but not necessarily, less than about 200 nucleotides in length. The terms "oligonucleotide" and "polynucleotide" are not mutually exclusive. The above description of polynucleotides applies equally and fully to oligonucleotides. The cells that produce the binding molecules of the present disclosure can include parental hybridoma cells and bacterial and eukaryotic host cells into which nucleic acids encoding the polypeptides have been introduced. Unless otherwise specifically stated, the left end of any single-stranded polynucleotide sequence disclosed herein is the 5' end; the left-hand direction of a double-stranded polynucleotide sequence is called the 5' direction. The 5' to 3' addition direction of nascent RNA transcripts is called the transcription direction; the sequence region in the 5' direction on the DNA strand having the same sequence as the RNA transcript and located at the 5' end of the RNA transcript is called the "upstream sequence"; the sequence region in the 3' direction on the DNA strand having the same sequence as the RNA transcript and located at the 3' end of the RNA transcript is called the "downstream sequence".
[0065] "Isolated nucleic acid" is a nucleic acid, e.g., RNA, DNA or a mixed nucleic acid, which is generally substantially separated from other genomic DNA sequences and proteins or complexes, such as ribosomes and polymerases, that naturally accompany the native sequence. An "isolated" nucleic acid molecule is a nucleic acid molecule that is separated from other nucleic acid molecules present in the natural source of the nucleic acid molecule. In addition, an "isolated" nucleic acid molecule, such as a cDNA molecule, can be substantially free of other cellular material or medium when produced by recombinant techniques, or can be substantially free of chemical precursors or other chemicals when chemically synthesized. In a specific embodiment, one or more nucleic acid molecules encoding a fusion protein as described herein are isolated or purified. The term encompasses nucleic acid sequences that have been removed from their natural environment, and includes recombinant or cloned DNA isolates and chemically synthesized analogs or analogs biosynthesized by heterologous systems. Substantially pure molecules can include the isolated form of the molecule. Specifically, an "isolated" nucleic acid molecule encoding a polypeptide as described herein is a nucleic acid molecule that has been identified and separated from at least one contaminating nucleic acid molecule that is normally associated with it in the environment in which it is produced.
[0066] As used herein in connection with a polypeptide or protein, the term "variant" refers to a polypeptide having a specified percentage sequence identity to a reference polypeptide, e.g., a polypeptide having at least about 80% amino acid sequence identity to a reference polypeptide, e.g., the corresponding full-length native sequence. These polypeptide variants include, e.g., polypeptides in which one or more amino acid residues have been added or deleted. In embodiments, the variant has at least about 80% amino acid sequence identity, at least about 81% amino acid sequence identity, at least about 82% amino acid sequence identity, at least about 83% amino acid sequence identity, at least about 84% amino acid sequence identity, at least about 85% amino acid sequence identity, at least about 86% amino acid sequence identity, at least about 87% amino acid sequence identity, at least about 88% amino acid sequence identity, at least about 89% amino acid sequence identity, at least about 90% amino acid sequence identity, alternatively at least about 91% amino acid sequence identity, at least about 92% amino acid sequence identity, at least about 93% amino acid sequence identity, at least about 94% amino acid sequence identity, at least about 95% amino acid sequence identity, at least about 96% amino acid sequence identity, at least about 97% amino acid sequence identity, at least about 98% amino acid sequence identity or at least about 99% amino acid sequence identity to the reference polypeptide, e.g., the corresponding full-length native sequence. In embodiments, the variant polypeptide is at least about 10 amino acids in length, at least about 20 amino acids in length, at least about 30 amino acids in length, at least about 40 amino acids in length, at least about 50 amino acids in length, at least about 60 amino acids in length, at least about 70 amino acids in length, at least about 80 amino acids in length, at least about 90 amino acids in length, at least about 100 amino acids in length, at least about 150 amino acids in length, at least about 200 amino acids in length, at least about 300 amino acids in length or longer. Variants include conservative or non-conservative substitutions in nature. For example, the polypeptide of interest may include up to about 5-10 conservative or non-conservative amino acid substitutions or even up to about 15-25 or 50 conservative or non-conservative amino acid substitutions, or any number of substitutions between 5-50.
[0067] As used herein, the term "homology" refers to the percentage of identity between two polynucleotides or two polypeptide moieties. Two DNA or two polypeptide sequences are "substantially homologous" to each other when the sequences show at least about 50%, at least about 75%, at least about 80%-85%, at least about 90%, at least about 95%-98% sequence identity, at least about 99% or any percentage therebetween over a defined length of the molecule. As used herein, substantially homologous also refers to a sequence that shows complete identity to the indicated DNA or polypeptide sequence.
[0068] As used herein, the term "identity" refers to the exact nucleotide-to-nucleotide or amino acid-to-amino acid correspondence of two polynucleotide or polypeptide sequences, respectively. Methods for determining the percentage of identity are well known in the art. For example, the percentage of identity can be determined by directly comparing the sequence information between two molecules, by sequence alignment, counting the exact number of matches between the two aligned sequences, dividing by the length of the shorter sequence and multiplying the result by 100. Readily available computer programs can be used to assist in the analysis, such as ALIGN, Dayhoff, M.O. in Atlas of Protein Sequence and Structure, M.O. Dayhoff, ed., 5 Suppl. 3:353-358, National Biomedical Research Foundation, Washington, D.C., which uses the local homology algorithm of Smith and Waterman, Advances in Appl. Math. 2:482-489, 1981 for peptide analysis. Programs for determining nucleotide sequence identity are available in the Wisconsin Sequence Analysis Package, version 8 (obtained from Genetics Computer Group, Madison, Wis.), for example, the BESTFIT, FASTA, and GAP programs, which also rely on the Smith and Waterman algorithm. These programs are easily used with the default parameters recommended by the manufacturer and described in the Wisconsin Sequence Analysis Package mentioned above. For example, the percentage of identity of a particular nucleotide sequence with a reference sequence can be determined using the homology algorithm of Smith and Waterman with a default scoring table and a gap penalty of 6 nucleotide positions. In the context of the present invention, another method for establishing the percentage of identity is to use the MPSRCH program package copyrighted by the University of Edinburgh, developed by John F. Collins and Shane S. Sturrok, and distributed by IntelliGenetics, Inc. (Mountain View, Calif.). According to this software package, the Smith-Waterman algorithm can be used, where the default parameters are used for the scoring table (e.g., a gap opening penalty of 12, a gap extension penalty of 1, and a gap of 6). From the resulting data, the "match" value reflects the "sequence identity". Other programs suitable for calculating the percentage of identity or similarity between sequences are generally known in the art. For example, another alignment program is BLAST, which uses default parameters.For example, using the following default parameters: genetic code = standard; filter = none; strands = both; cutoff = 60; expect = 10; Matrix = BLOSUM62; description = 50 sequences; sort by = high score; database = non-redundant, GenBank+EMBL+DDBJ+PDB+GenBank CDS translations+Swiss protein+Spupdate+PIR, BLASTN and BLASTP can be used. The details of these programs are well known in the art. As an alternative, homology can be determined by polynucleotide hybridization under conditions that form stable double helices between homologous regions, followed by digestion with a single-strand-specific nuclease and determination of the size of the digestion fragments. Substantially homologous DNA sequences can be identified in Southern hybridization experiments under, for example, stringent conditions as defined for that particular system. Defining appropriate hybridization conditions is within the skill of the art. See, e.g., Sambrook et al., supra; DNA Cloning, supra; Nucleic Acid Hybridization, supra.
[0069] As used herein, the term "vector" refers to a substance for carrying or including a nucleic acid sequence, e.g., for introducing the nucleic acid sequence into a host cell. Suitable vectors include, e.g., expression vectors, plasmids, phage vectors, viral vectors, episomes, and artificial chromosomes, which may include selectable sequences or markers operable for stable integration into the host cell chromosome. Additionally, the vector may include one or more selectable marker genes and appropriate expression control sequences. Selectable marker genes that may be included provide, e.g., antibiotic or toxin resistance, complementation of auxotrophs, or provision of essential nutrients not present in the medium. Expression control sequences may include constitutive and inducible promoters, transcriptional enhancers, transcriptional terminators, etc., well known in the art. When co-expressing two or more nucleic acid molecules, the two nucleic acid molecules may be inserted, e.g., into a single expression vector or different expression vectors. Methods well known in the art can be used to confirm the introduction of the nucleic acid molecule into the host cell. These methods include, e.g., nucleic acid analysis such as RNA blotting of mRNA or polymerase chain reaction (PCR) amplification, immunoblotting for expression of the gene product, or other suitable analytical methods to test for expression of the introduced nucleic acid sequence or its corresponding gene product. The term "vector" includes cloning and expression vehicles as well as viral vectors. In certain embodiments, the vectors provided herein are recombinant AAV vectors.
[0070] As used herein, the term "recombinant AAV vector (rAAV vector)" refers to a polynucleotide vector that contains a nucleic acid sequence from AAV and one or more heterologous sequences (i.e., nucleic acid sequences not of AAV origin). In some embodiments, the one or more heterologous sequences are flanked by at least one, and in certain embodiments, two AAV inverted terminal repeats (ITRs). In some embodiments, these rAAV vectors can be replicated and packaged into infectious virus capsid particles, for example, when present in a host cell that has been infected with a suitable helper virus (or expresses suitable helper virus functions) and expresses AAV rep and cap gene products (i.e., AAV Rep and capsid proteins). The rAAV vector can be introduced into a larger polynucleotide (e.g., a chromosome or another vector such as a plasmid used for cloning or transfection), and the rAAV vector can be "rescued" by replication and encapsidation in the presence of AAV packaging functions and appropriate helper virus functions. The rAAV vector can be in any of several forms, including (but not limited to) plasmids, linearized artificial chromosomes, complexed with lipids, encapsulated within liposomes, and encapsidated within virus capsid particles, specifically AAV particles. The rAAV vector can be packaged into an AAV capsid to produce "recombinant adeno-associated virus capsid particles (rAAV particles)".
[0071] As used herein with respect to nucleic acid sequences such as coding sequences and control sequences, the term "heterologous" refers to sequences that are not normally joined together and / or not normally associated with a particular cell. Thus, a "heterologous" region of a nucleic acid construct or vector is a nucleic acid fragment that is within or joined to another nucleic acid molecule that is not found in nature in association with the other molecule. For example, a heterologous region of a nucleic acid construct can include a coding sequence flanked by sequences that are not found in nature in association with the coding sequence. Another example of a heterologous coding sequence is a construct in which the coding sequence itself does not exist in nature (e.g., a synthetic sequence having codons different from a native gene).
[0072] As used herein with respect to a sequence flanked by other elements, the term "flanked" refers to the presence of one or more flanking elements that are upstream and / or downstream, i.e., 5' and / or 3', relative to the sequence. The term "flanked" is not intended to mean that the sequences must be contiguous. For example, there can be intervening sequences between a nucleic acid encoding a transgene and the flanking elements. A sequence (e.g., a transgene) "flanked" by two other elements (e.g., TRs) means that one element is located 5' of the sequence and the other is located 3' of the sequence; however, there can be intervening sequences between them.
[0073] As used herein, the term "terminal inverted repeat" or "ITR" sequence refers to relatively short sequences located at the ends of a viral genome that are in opposite orientations. "AAV terminal inverted repeat (ITR)" sequences are well known in the art and are typically sequences of about 145 nucleotides that are present at both ends of the native single-stranded AAV genome. The outermost 125 nucleotides of the ITR can be present in either of two alternative orientations, which results in heterogeneity between different AAV genomes and between the two ends of a single AAV genome. The outermost 125 nucleotides also contain several shorter regions with self-complementarity (designated regions A, A', B, B', C, C', and D), which allow for intrastrand base pairing within this portion of the ITR.
[0074] A "coding sequence" or sequence that "encodes" a selected polypeptide is a nucleic acid molecule that, when placed under the control of appropriate regulatory sequences, is transcribed (in the case of DNA) and translated (in the case of mRNA) into a polypeptide. The boundaries of the coding sequence are determined by a start codon at the 5' (amino) terminus and a translation stop codon at the 3' (carboxyl) terminus. A transcription termination sequence can be located 3' of the coding sequence.
[0075] The term "control sequence" refers to DNA sequences that are necessary for the expression of an operably linked coding sequence in a particular host organism. Control sequences suitable for prokaryotes (for example) include a promoter, optionally an operator sequence, and a ribosome binding site. It is known that eukaryotic cells use promoters, polyadenylation signals, and enhancers.
[0076] When used to refer to nucleic acids or amino acids, as used herein, the term "operably linked" and like phrases (e.g., gene fusion) denote the operative joining of nucleic acid sequences or amino acid sequences that are arranged in a functional relationship with each other. For example, operably linked promoter, enhancer element, open reading frame, 5' and 3' UTR, and terminator sequences result in the accurate production of a nucleic acid molecule (e.g., RNA). In some embodiments, operably linked nucleic acid elements result in the transcription of an open reading frame and ultimately the production of a polypeptide (i.e., the expression of the open reading frame). As another example, an operably linked peptide is one in which functional domains are arranged at an appropriate distance from each other to confer the intended function on each domain.
[0077] As used herein in its ordinary meaning, the term "promoter" refers to a nucleotide region that includes a DNA regulatory sequence, where the regulatory sequence is derived from a gene capable of binding RNA polymerase and initiating transcription of a downstream (3'-direction) coding sequence. Transcriptional promoters can include "inducible promoters" (where the expression of a polynucleotide sequence operably linked to the promoter is induced by an analyte, cofactor, regulatory protein, etc.), "repressible promoters" (where the expression of a polynucleotide sequence operably linked to the promoter is repressed by an analyte, cofactor, regulatory protein, etc.), and "constitutive promoters".
[0078] As used herein in a broad sense, the term "transgene" refers to the introduction of a viral vector, e.g., any heterologous nucleotide sequence for expression in a target cell and it can be associated with an expression control sequence, such as a promoter. Those skilled in the art will understand that the expression control sequence will be selected based on the ability to promote the expression of the transgene in the target cell. Examples of transgenes are nucleic acids encoding a therapeutic polypeptide or a detectable marker.
[0079] As used herein, the term "AAV capsid" or "AAV capsid protein" or "AAV cap" refers to the proteins encoded by the AAV capsid (cap) gene (e.g., VP1, VP2, and VP3) or variants thereof. For example, the term includes (but is not limited to) capsid proteins derived from any AAV serotype, such as AAV1, AAV2, AAV2i8, AAV3, AAV3-B, AAV4, AAV5, AAV6, AAV7, AAV8, AAVrh8, AAVrh8R, AAV9, AAV10, AAVrh10, AAV11, AAV12, AAV13, AAV-DJ, AAV-2 / 1, AAV 2 / 6, AAV 2 / 7, AAV 2 / 8, AAV 2 / 9, AAV LK03, AAVrh10, AAVrh74, AAV44-9 or variants thereof. The term also includes capsid proteins expressed by recombinant AAVs, such as chimeric AAVs or derived from them.
[0080] As used herein, the term "AAV capsid particle" or "AAV particle" includes at least one AAV capsid protein (e.g., VP1 protein, VP2 protein, VP3 protein or variants thereof) and optionally encapsulates a nucleic acid from the AAV genome or a nucleic acid derived from the AAV genome.
[0081] Based on capsid protein sequences and capsid structures, the term "serotype" for a vector or viral capsid is defined by different immunological profiles.
[0082] As used herein, the term "chimeric" with respect to a viral capsid or particle means that the capsid or particle comprises sequences from different parvoviruses, preferably different AAV serotypes, as described in Rabinowitz et al., U.S. Patent No. 6,491,907, the disclosure of which is incorporated herein by reference in its entirety.
[0083] The term "recombinant" means a genetic entity that is different from what is normally found in nature. As applied to a polynucleotide or gene, this means that the polynucleotide is the product of a combination of cloning, restriction, and / or ligation steps, and other procedures that result in the production of a construct that is different from the polynucleotides found in nature.
[0084] As used herein, the term "recombinant virus" refers to a virus that has been genetically altered, for example, by the addition or insertion of a heterologous nucleic acid construct into the particle. For example, as used herein, the term "recombinant AAV particle" or "rAAV" refers to an AAV that has been genetically altered, for example, by the deletion or other mutation of an endogenous AAV gene and / or the addition or insertion of a heterologous nucleic acid construct into the polynucleotide of the AAV particle.
[0085] As used herein, the term "transfection" or "transformation" or "transduction" refers to the process of transferring or introducing exogenous nucleic acid into a host cell. A "transfected" or "transformed" or "transduced" cell is a cell that has been transfected, transformed, or transduced with exogenous nucleic acid. For example, the term "transfection" is used to denote the uptake of exogenous DNA by a cell, and a cell has been "transfected" when the exogenous DNA has been introduced inside the cell membrane. Some transfection techniques are well known in the art. See, for example, Graham et al. (1973) Virology, 52:456, Sambrook et al. (1989) Molecular Cloning, a laboratory manual, Cold Spring Harbor Laboratories, New York, Davis et al. (1986) Basic Methods in Molecular Biology, Elsevier, and Chu et al. (1981) Gene 13:197. These techniques can be used to introduce one or more exogenous molecules into a suitable host cell. "Transduction" of a cell by a virus means the transfer of nucleic acid, such as DNA or RNA, from the viral particle into the cell.
[0086] As used herein, the term "host cell" refers to a particular cell that can be transfected with a nucleic acid molecule and the progeny or potential progeny of such cell. A host cell can be a bacterial cell, a yeast cell, an insect cell, or a mammalian cell.
[0087] The term "purified" refers to the isolation of a substance (compound, polynucleotide, protein, polypeptide, polypeptide composition) such that the substance of interest is the majority of the sample in which it is present. Typically, in a sample, a substantially purified component comprises 50%, 80%-85%, 90-99% of the sample, such as at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%. Techniques for purifying polynucleotides and polypeptides of interest are well known in the art and include, for example, ion exchange chromatography, affinity chromatography, and sedimentation according to density.
[0088] As used herein, the term "pharmaceutically acceptable" means approved by a regulatory agency of the federal or state government or listed in the U.S. Pharmacopeia, European Pharmacopeia, or other generally recognized pharmacopeia for use in animals and more particularly in humans.
[0089] In one embodiment, each component is "pharmaceutically acceptable" in the sense of being compatible with the other ingredients of a pharmaceutical formulation and suitable for use in contact with the tissues or organs of humans and animals without excessive toxicity, irritation, allergic response, immunogenicity, or other problems or complications commensurate with a reasonable benefit / risk ratio. See, e.g., Lippincott Williams & Wilkins: Philadelphia, PA, 2005; Handbook of Pharmaceutical Excipients, 6th Edition; edited by Rowe et al.; The Pharmaceutical Press and the American Pharmaceutical Association: 2009; Handbook of Pharmaceutical Additives, 3rd Edition; edited by Ash and Ash; Gower Publishing Company: 2007; Pharmaceutical Preformulation and Formulation, 2nd Edition; edited by Gibson; CRC Press LLC: Boca Raton, FL, 2009. In some embodiments, a pharmaceutically acceptable excipient is non-toxic to the cells or mammals to which it is exposed at the dosages and concentrations used. In some embodiments, a pharmaceutically acceptable excipient is an aqueous solution buffered to a pH.
[0090] As used herein, the terms “treat,” “treatment,” and “treating” refer to a decrease or improvement in the development, severity, and / or duration of a disease or condition resulting from the administration of one or more therapies. Treatment can be determined by observing an improvement in the patient by assessing a decrease, alleviation, and / or remission of one or more symptoms associated with the underlying condition, even though the patient may still be afflicted with the underlying condition. The term “treatment” encompasses both the control and the improvement of a disease. The term “manage,” “managing,” and “management” refer to the beneficial effects obtained by a subject from a therapy that does not necessarily result in a cure of the disease. Treatment includes: (1) preventing a disease in a subject who may be exposed to or is susceptible to the disease but has not yet experienced or exhibited symptoms of the disease, i.e., preventing the development of the disease or causing the disease to occur at a lower intensity, (2) inhibiting the disease, i.e., arresting development, preventing or delaying development, or reversing the disease state, (3) alleviating the symptoms of the disease, i.e., reducing the number of symptoms experienced by the subject, and (4) reducing, preventing, or delaying the development of the disease or its symptoms. The term “prevent,” “preventing,” and “prevention” refer to reducing the likelihood of the onset (or recurrence) of a disease, disorder, condition, or related symptoms.
[0091] As used herein, “administer,” “administration,” or “administering” refers to the act of injecting or otherwise physically delivering a substance (e.g., a conjugate or pharmaceutical composition provided herein) to a subject or patient (e.g., a human), such as by oral, mucosal, topical, intradermal, parenteral, intravenous, intravitreal, intra-articular, subretinal, intramuscular, intrathecal delivery, and / or any other physical delivery method described herein or known in the art. In a particular embodiment, administration is by intravenous infusion. The conjugate or composition provided herein can be delivered systemically or to a particular tissue.
[0092] As used herein, the terms “effective amount” or “therapeutically effective amount” refer to an amount of a therapeutic agent (e.g., a conjugate or pharmaceutical composition provided herein) sufficient to treat, diagnose, prevent, delay the onset of, reduce, and / or improve the severity and / or duration of a given condition, disorder, or disease and / or its associated symptoms. These terms also encompass an amount necessary to reduce, slow down, or improve the progression or development of a given disease, reduce, slow down, or improve the recurrence, development, or onset of a given disease, and / or improve or enhance the prophylactic or therapeutic effect of another therapy or serve as a bridge to connect another therapy. In some embodiments, “effective amount” as used herein also refers to an amount of a conjugate described herein that achieves the specified result. As used herein, the terms “subject” and “patient” are used interchangeably.
[0093] As used herein, a subject is a mammal, such as a non - primate (e.g., cow, pig, horse, cat, dog, goat, rabbit, rat, mouse, etc.) or a primate (e.g., monkey and human), e.g., human. In certain embodiments, the subject is a mammal diagnosed with a disease or disorder provided herein, e.g., human. In another embodiment, the subject is a mammal at risk of developing a disease or disorder provided herein, e.g., human. In a specific embodiment, the subject is human.
[0094] As used herein, the term "therapy" can mean any protocol, method, composition, preparation, and / or reagent that can be used to prevent, treat, control, or improve a disease or disorder or its symptoms (e.g., a disease or disorder provided herein or one or more symptoms or conditions associated therewith). In certain embodiments, the term "therapy" means a pharmaceutical therapy, adjunctive therapy, radiation, surgery, biological therapy, supportive therapy, and / or other therapy useful for treating, controlling, preventing, or improving a disease or disorder or one or more of its symptoms. In certain embodiments, the term "therapy" refers to a therapy other than the conjugate or its pharmaceutical composition described herein.
[0095] As used herein, the term "disease or disorder associated with angiogenesis" refers to a disease or disorder that involves angiogenesis (including abnormal angiogenesis), e.g., as a symptom or as a direct or indirect cause. The term includes diseases or disorders whose development involves angiogenesis, including (e.g.) cancer and eye diseases or disorders.
[0096] The terms "about" and "approximately" mean within 20%, within 15%, within 10%, within 9%, within 8%, within 7%, within 6%, within 5%, within 4%, within 3%, within 2%, within 1% or less of a given value or range.
[0097] Unless the context clearly dictates otherwise, as used in the disclosure and claims of the present invention, the singular forms "a" and "the" include the plural forms.
[0098] It should be understood that wherever embodiments are described herein using the term "comprising", similar embodiments are also provided by the terms "consisting of" and / or "consisting essentially of". It should also be understood that wherever embodiments are described herein using the phrase "consisting essentially of", similar embodiments are also provided by the term "consisting of".
[0099] The term "between" as used in phrases such as "between A and B" or "A - B" refers to a range that includes both A and B.
[0100] As used herein, the term "and / or" as used in a phrase such as "A and / or B" is intended to include both A and B; A or B; A alone; and B alone. Similarly, the term "and / or" as used in a phrase such as "A, B, and / or C" is intended to cover each of the following embodiments: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A alone; B alone; and C alone.
[0101] 5.2. Fusion proteins and variants thereof
[0102] 5.2.1. Multispecific fusion proteins
[0103] In one aspect, the present disclosure provides multispecific fusion proteins (or polypeptides) that include multiple VEGF-binding domains (e.g., two or three VEGF-binding domains) and a domain that binds to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2). In some embodiments, the fusion proteins provided herein include at least three binding domains—a first domain that binds to VEGF, a second domain that binds to VEGF, and a third domain that binds to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2). In some embodiments, the multispecific fusion protein further includes a fourth domain that binds to VEGF, and thus in some embodiments, the present disclosure provides a fusion protein that includes at least four binding domains—a first domain that binds to VEGF, a second domain that binds to VEGF, a third domain that binds to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2), and a fourth domain that binds to VEGF (more specifically VEGFC).
[0104] Vascular endothelial growth factor (VEGF) belongs to the PDGF supergene family characterized by eight conserved cysteines and functions as a homodimeric structure and is produced by a variety of cell types, including tumor cells, macrophages, platelets, keratinocytes, and renal mesangial cells. VEGF-A regulates angiogenesis and vascular permeability by activating two receptors, VEGFR-1 (Flt-1) and VEGFR-2 (KDR / Flk1). In addition to its role in the vasculature, VEGF plays a role in normal physiological functions such as bone formation, hematopoiesis, wound healing, and development. For example, the formation of new blood vessels caused by the overproduction of growth factors such as VEGF is a key component of diseases such as tumor growth, age-related macular degeneration (AMD), and proliferative diabetic retinopathy (PDR). The VEGF-VEGFR system is an important target for anti-angiogenic therapy in cancer and is also an attractive system for pro-angiogenic therapy in the treatment of neuronal degeneration and ischemic diseases. Shibuya, Genes Cancer, 2(12):1097–1105 (2011). Involves the binding of VEGF-C to VEGFR-3 and this binding results in most of the biological effects of VEGFR-3. The discovery of the soluble form of VEGFR-3 (sVEGFR-3) and experiments on transgenic mice expressing this gene led to the conclusion that sVEGFR-3 inhibits lymphatic vessel development and causes edema, thereby inhibiting VEGF-C- and VEGF-D-mediated signaling.
[0105] Angiopoietins are a family of growth factors that includes the glycoproteins angiopoietin 1 (ANGPT1 or Ang1) and angiopoietin 2 (ANGPT2 or Ang2) and orthologs 3 (in mice) and 4 (in humans). Angiopoietins are involved in embryonic blood vessel development. Angiopoietin 1 is expressed by a variety of cell types, while angiopoietin 2 is almost restricted to endothelial cells. Both act on the TIE2 receptor tyrosine kinase, which is mainly present on endothelial cells and hematopoietic stem cells. These proteins are important regulators of angiogenesis and the maintenance of vascular integrity. As used herein, "angiopoietin" refers to any angiopoietin peptide, including angiopoietin 1 and angiopoietin 2; and "domain that binds to angiopoietin" refers to a domain that binds to angiopoietin 1 and / or angiopoietin 2.
[0106] In some embodiments, the first and second domains provided herein bind to human VEGF. In some embodiments, the third domain provided herein binds to human angiopoietin (e.g., angiopoietin 1 and angiopoietin 2). In some embodiments, the fourth domain provided herein binds to human VEGF (e.g., VEGFC).
[0107] In some embodiments, the fusion proteins provided herein modulate the activity of one or more VEGFs. In some embodiments, the fusion proteins provided herein modulate the activity of one or more angiopoietins (e.g., angiopoietin 1 and angiopoietin 2).
[0108] In some embodiments, the first domain provided herein binds to VEGF (e.g., human VEGF) with the following dissociation constant (K D ): ≤ 1 μM, ≤ 100 nM, ≤ 10 nM, ≤ 1 nM, ≤ 0.1 nM, ≤ 0.01 nM, or ≤ 0.001 nM (e.g., 10 - 8 M or less, e.g., 10 -8 M to 10 -13 M, e.g., 10 -9 M to 10 -13 M).
[0109] In some embodiments, the second domain provided herein binds to VEGF (e.g., human VEGF) with the following dissociation constant (K D ): ≤ 1 μM, ≤ 100 nM, ≤ 10 nM, ≤ 1 nM, ≤ 0.1 nM, ≤ 0.01 nM, or ≤ 0.001 nM (e.g., 10 - 8 M or less, e.g., 10 -8 M to 10 -13 M, e.g., 10 -9 M to 10 -13 M).
[0110] In some embodiments, the third domain provided herein binds to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2) with the following dissociation constant (K D ): ≤ 1 μM, ≤ 100 nM, ≤ 10 nM, ≤ 1 nM, ≤ 0.1 nM, ≤ 0.01 nM, or ≤ 0.001 nM (e.g., 10 -8 M or less, e.g., 10 -8 M to 10 -13 M, e.g., 10 -9 M to 10 -13 M).
[0111] In some embodiments, the fourth domain provided herein binds to VEGF (e.g., human VEGFC) with the following dissociation constant (K D ): ≤ 1 μM, ≤ 100 nM, ≤ 10 nM, ≤ 1 nM, ≤ 0.1 nM, ≤ 0.01 nM, or ≤ 0.001 nM (e.g., 10 -8 M or less, e.g., 10 -8 M to 10-13 M, for example, 10 -9 M to 10 -13 M).
[0112] K can be measured using techniques known in the art and the methods described in Section 5.1 above D .
[0113] In some embodiments, the first domain is derived from VEGF receptor-1 (VEGFR-1 or FLT-1). In some embodiments, the first domain comprises the IgG-like domain 2 of VEGFR-1 (or domain 2 or D2) or a variant thereof. Thus, in the figures of the present invention, for example, in Figure 1 , Figure 2 , Figure 8 and Figures 9A - 9I this domain is referred to as D2. In some embodiments, the first domain comprises the amino acid sequence shown in SEQ ID NO: 1 or consists of said amino acid sequence. In some embodiments, the first domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity to SEQ ID NO: 1 or consists of said amino acid sequence.
[0114] In some embodiments, the second domain is derived from VEGF receptor-2 (VEGFR-2 or Flk-1). In some embodiments, the second domain comprises the IgG-like domain 3 of VEGFR-2 (or domain 3 or D3) or a variant thereof. Thus, in the figures of the present invention, for example, in Figure 1 , Figure 2 , Figure 8 and Figures 9A - 9I this domain is referred to as D3. In some embodiments, the second domain comprises the amino acid sequence shown in SEQ ID NO: 2 or consists of said amino acid sequence. In some embodiments, the second domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity to SEQ ID NO: 2 or consists of said amino acid sequence.
[0115] In some embodiments, the third domain (ABD) comprises one or more repeats of the amino acid sequence shown in SEQ ID NO: 3. In some embodiments, the third domain (ABD) comprises one or more repeats of the amino acid sequence shown in SEQ ID NO: 51. In some embodiments, the third domain comprises two repeats of the amino acid sequence shown in SEQ ID NO: 3. In some embodiments, the third domain comprises two repeats of the amino acid sequence shown in SEQ ID NO: 51. In other embodiments, the third domain comprises the amino acid sequence shown in SEQ ID NO: 3 and the amino acid sequence shown in SEQ ID NO: 51. In cases where the ABD in the fusion protein of the present invention comprises the amino acid sequence shown in SEQ ID NO: 3 and the amino acid sequence shown in SEQ ID NO: 51, the two sequences can be in any order. For example, the amino acid sequence shown in SEQ ID NO: 3 can be located at the N-terminus or C-terminus of the amino acid sequence shown in SEQ ID NO: 51.
[0116] In some specific embodiments, the third domain (ABD) comprises or consists of the amino acid sequence shown in SEQ ID NO: 4. In some embodiments, the third domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity to SEQ ID NO: 4 and is capable of binding to angiopoietin (e.g., angiopoietin 2 or Ang2).
[0117] In some embodiments, the third domain (ABD) comprises or consists of the amino acid sequence shown in SEQ ID NO: 52. In some embodiments, the third domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity to SEQ ID NO: 52 and is capable of binding to angiopoietin (e.g., angiopoietin 2 or Ang2).
[0118] In some embodiments, the third domain (ABD) comprises or consists of the amino acid sequence shown in SEQ ID NO: 53. In some embodiments, the third domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity to SEQ ID NO: 53 and is capable of binding to angiopoietin (e.g., angiopoietin 2 or Ang2).
[0119] In some embodiments, the third domain (ABD) comprises the amino acid sequence shown in SEQ ID NO: 54 or consists of said amino acid sequence. In some embodiments, the third domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identity to SEQ ID NO: 54 and is capable of binding to angiopoietin (e.g., angiopoietin 2 or Ang2).
[0120] In certain embodiments, the third domain that binds to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2) is located at the N-terminus of the first domain or the second domain. In a specific embodiment, the polypeptide comprises, from the N-terminus to the C-terminus, the third domain, the first domain, and the second domain. In another specific embodiment, the polypeptide comprises, from the N-terminus to the C-terminus, the third domain, the second domain, and the first domain.
[0121] In cases where the fourth domain, i.e., the VEGFC binding domain (also referred to as Trap C), is present in the fusion protein, in some embodiments, the VEGFC binding domain is derived from VEGFR-2 (e.g., domain 2). In other embodiments, the VEGFC binding domain is derived from VEGFR-3 (e.g., domain 1 and / or domain 2).
[0122] In some embodiments, the VEGFC binding domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 55.
[0123] In some embodiments, the VEGFC binding domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 56.
[0124] In some embodiments, the VEGFC binding domain comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 57.
[0125] In some embodiments, the polypeptide further comprises the Fc region of an antibody. In some embodiments, the Fc region is derived from human IgG. In some embodiments, the Fc region comprises the amino acid sequence shown in SEQ ID NO: 5. In some embodiments, the Fc region is located at the C-terminus of the polypeptide. In other embodiments, the Fc region is not located at the C-terminus of the polypeptide.
[0126] The domains in the fusion protein of the present invention can be present in any order. In some embodiments, D2, D3, ABD, the Fc region, and / or Trap C are in the order as shown in Figure 1 , Figure 2 , Figure 8 and Figures 9A - 9I . In some more specific embodiments, D2, D3, ABD, the Fc region, and / or Trap C are in the order as shown in Figure 9A . In some more specific embodiments, D2, D3, ABD, the Fc region, and / or Trap C are in the order as shown in Figure 9B . In some more specific embodiments, D2, D3, ABD, the Fc region, and / or Trap C are in the order as shown in Figure 9C . In some more specific embodiments, D2, D3, ABD, the Fc region, and / or Trap C are in the order as shown in Figure 9D . In some more specific embodiments, D2, D3, ABD, the Fc region, and / or Trap C are in the order as shown in Figure 9E . In some more specific embodiments, D2, D3, ABD, the Fc region, and / or Trap C are in the order as shown in Figure 9F . In some more specific embodiments, D2, D3, ABD, the Fc region, and / or Trap C are in the order as shown in Figure 9G . In some more specific embodiments, D2, D3, ABD, the Fc region, and / or Trap C are in the order as shown in Figure 9H . In some more specific embodiments, D2, D3, ABD, the Fc region, and / or Trap C are in the order as shown in Figure 9I .
[0127] In some embodiments, the polypeptide further comprises a signal peptide located at the N-terminus of the polypeptide. Generally, a signal peptide is a peptide sequence that targets a polypeptide to a desired site in a cell. In some embodiments, the signal peptide targets an effector molecule to the secretory pathway of the cell and will allow the effector molecule to integrate and anchor in the lipid bilayer. Signal peptides compatible for use in the fusion proteins described herein include the signal sequences of naturally occurring proteins or synthetic, non-naturally occurring signal sequences, which will be apparent to those skilled in the art. In some specific embodiments, the signal peptide comprises the amino acid sequence shown in SEQ ID NO: 6.
[0128] In some embodiments, the polypeptide further comprises one or more linkers located between the above-described domains. The domains described herein may be fused to each other via a peptide linker. In some embodiments, certain domains are directly fused to each other without any peptide linker. The peptide linkers connecting different domains may be the same or different. In some embodiments, the polypeptides provided herein comprise peptide linkers between certain domains but do not comprise other domains therein.
[0129] Each peptide linker in the polypeptides provided herein may have the same or different length and / or sequence. Each peptide linker can be independently selected and optimized. The length, flexibility, and / or other properties of the peptide linkers used in the fusion proteins of the present invention may have an impact on the properties of one or more specific target molecules, including (but not limited to) affinity, specificity, or avidity. In some embodiments, the peptide linker comprises flexible residues (such as glycine and serine) so that adjacent domains can move freely relative to each other. For example, a glycine-serine dyad may be a suitable peptide linker.
[0130] The peptide linker can have any suitable length. In some embodiments, the length of the peptide linker is at least any one of about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50 or more amino acids. In some embodiments, the length of the peptide linker is no more than any one of about 100, 75, 50, 40, 35, 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5 or fewer amino acids. In some embodiments, the length of the peptide linker is any of the following: about 1 amino acid to about 10 amino acids, about 1 amino acid to about 20 amino acids, about 1 amino acid to about 30 amino acids, about 5 amino acids to about 15 amino acids, about 10 amino acids to about 25 amino acids, about 5 amino acids to about 30 amino acids, about 10 amino acids to about 30 amino acids long or about 30 amino acids to about 50 amino acids.
[0131] The peptide linker can have a naturally occurring sequence or a non-naturally occurring sequence. For example, a sequence derived from the hinge region of a heavy-chain only antibody can be used as a linker. See, for example, WO1996 / 34103. In some embodiments, the peptide linker is a flexible linker. In some embodiments, the peptide linker provided herein is a (GxS)n linker, where x and n can independently be an integer between 3 and 12, including 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or greater. Exemplary flexible linkers include (but are not limited to) glycine polymers (G) n , glycine-serine polymers (including, for example, (GS) n , (GSGGS) n , (GGGS) n and (GGGGS) n , where n is an integer of at least 1), glycine-alanine polymers, alanine-serine polymers, and other flexible linkers known in the art. The following table lists exemplary peptide linkers.
[0132] Table 1. Exemplary Peptide Linkers
[0133]
[0134]
[0135] The fusion proteins disclosed herein can include a hinge domain between the domains described above. A hinge domain is an amino acid segment that is typically present between two domains of a protein and can allow flexibility of the protein and movement of one or both of the domains relative to each other.
[0136] The hinge domain may contain about 10 - 100 amino acids, for example, any one of about 15 - 75 amino acids, 20 - 50 amino acids, or 30 - 60 amino acids. In some embodiments, the hinge domain can be 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, or 75 amino acids in length, at least about any one of these.
[0137] In some embodiments, the hinge domain is the hinge domain of a naturally occurring protein. In some embodiments, the hinge domain is at least a portion of the hinge domain of a naturally occurring protein. The hinge domains of antibodies, such as IgG, IgA, IgM, IgE, or IgD antibodies, are also compatible for use in the fusion proteins described herein. In some embodiments, the hinge domain is the hinge domain that connects the constant domains CH1 and CH2 of an antibody. In some embodiments, the hinge domain has an antibody and includes the hinge domain of the antibody and one or more constant regions of the antibody. In some embodiments, the hinge domain includes the hinge domain of the antibody and the CH3 constant region of the antibody. In some embodiments, the hinge domain includes the hinge domain of the antibody and the CH2 and CH3 constant regions of the antibody. In some embodiments, the antibody is an IgG, IgA, IgM, IgE, or IgD antibody. In some embodiments, the antibody is an IgG antibody. In some embodiments, the antibody is an IgG1, IgG2, IgG3, or IgG4 antibody. In some embodiments, the hinge region includes the hinge region and the CH2 and CH3 constant regions of an IgG1 antibody. In some embodiments, the hinge region includes the hinge region and the CH3 constant region of an IgG1 antibody.
[0138] Non - naturally occurring peptides can also be used as hinge domains for the fusion proteins described herein. In some embodiments, the linker is a self - cleavable linker.
[0139] For example, other linkers known in the art as described in WO2016014789, WO2015158671, WO2016102965, US20150299317, WO2018067992, US7741465, Colcher et al., J. Nat. Cancer Inst. 82:1191 - 1197 (1990), and Bird et al., Science 242:423 - 426 (1988) can also be included in the fusion proteins provided herein, and the disclosure of each of the above - mentioned documents is incorporated herein by reference.
[0140] In certain embodiments, Figure 1 the polypeptides provided herein are shown.
[0141] In some specific embodiments, polypeptides comprising the amino acid sequence shown in SEQ ID NO: 7 are provided herein.
[0142] In some specific embodiments, polypeptides comprising the amino acid sequence shown in SEQ ID NO: 8 are provided herein.
[0143] In some specific embodiments, polypeptides comprising the amino acid sequence shown in SEQ ID NO: 9 are provided herein.
[0144] In some specific embodiments, polypeptides comprising the amino acid sequence shown in SEQ ID NO: 10 are provided herein.
[0145] In certain embodiments, Figure 8 the polypeptides provided herein are shown.
[0146] In some specific embodiments, polypeptides comprising the amino acid sequence shown in SEQ ID NO: 58 are provided herein.
[0147] In some specific embodiments, polypeptides comprising the amino acid sequence shown in SEQ ID NO: 59 are provided herein.
[0148] In some specific embodiments, polypeptides comprising the amino acid sequence shown in SEQ ID NO: 60 are provided herein.
[0149] In some specific embodiments, polypeptides comprising the amino acid sequence shown in SEQ ID NO: 61 are provided herein.
[0150] In some specific embodiments, polypeptides comprising the amino acid sequence shown in SEQ ID NO: 62 are provided herein.
[0151] In some specific embodiments, polypeptides comprising the amino acid sequence shown in SEQ ID NO: 63 are provided herein.
[0152] In some specific embodiments, polypeptides comprising the amino acid sequence shown in SEQ ID NO: 64 are provided herein.
[0153] In some specific embodiments, polypeptides comprising the amino acid sequence shown in SEQ ID NO: 65 are provided herein.
[0154] In some specific embodiments, polypeptides comprising the amino acid sequence shown in SEQ ID NO: 66 are provided herein.
[0155] In certain embodiments, the fusion proteins described herein comprise an amino acid sequence having a specified percentage identity to any of the polypeptides described in Sections 6 below.
[0156] Determination of the percentage identity between two sequences, such as, for example, an amino acid sequence or a nucleic acid sequence, can be accomplished using a mathematical algorithm. A preferred, non-limiting example of a mathematical algorithm for comparing two sequences is the algorithm of Karlin and Altschul, Proc. Natl. Acad. Sci. U.S.A. 87:2264–2268 (1990), as modified in Karlin and Altschul, Proc. Natl. Acad. Sci. U.S.A. 90:5873–5877 (1993). This algorithm is incorporated into the NBLAST and XBLAST programs of Altschul et al., J. Mol. Biol. 215:403 (1990). The NBLAST nucleotide program parameters can be set, for example, to score = 100, wordlength = 12, for BLAST nucleotide searches to obtain nucleotide sequences identical to the nucleic acid molecules described herein. The XBLAST program parameters can be set, for example, to score = 50, wordlength = 3, for BLAST protein searches to obtain protein sequences identical to the protein molecules described herein. For purposes of comparison, to obtain gapped alignments, Gapped BLAST can be used as described in Altschul et al., Nucleic Acids Res. 25:3389–3402 (1997). Alternatively, PSI BLAST can be used to perform an iterative search that detects distant relationships between molecules (ibid.). When using the BLAST, Gapped BLAST, and PSI Blast programs, the default parameters of the respective programs (e.g., XBLAST and NBLAST) can be used (see, e.g., the National Center for Biotechnology Information (NCBI) at world wide web ncbi.nlm.nih.gov). Another non-limiting example of a mathematical algorithm for sequence comparison is the algorithm of Myers and Miller, CABIOS 4:11–17 (1998). This algorithm is incorporated into the ALIGN program (version 2.0) as part of the GCG sequence alignment software package. When using the ALIGN program to compare amino acid sequences, the PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4 can be used.
[0157] The percentage identity between two sequences can be determined using techniques similar to those described above, with or without allowing gaps. In calculating the percentage identity, only exact matches are typically counted.
[0158] In some embodiments, polypeptides are provided that have at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to an amino acid sequence selected from SEQ ID NOs: 7-10 and 58-66. In some embodiments, polypeptides having at least about 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to a reference sequence contain substitutions (e.g., conservative substitutions), insertions or deletions, but two domains or three domains within the polypeptide containing the sequence retain the ability to bind to VEGF and one domain within the polypeptide retains the ability to bind to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2).
[0159] In some embodiments, a total of 1 to 10 amino acids have been substituted, inserted and / or deleted in an amino acid sequence selected from SEQ ID NOs: 7-10 and 58-66.
[0160] In certain embodiments, the fusion proteins described herein contain an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the amino acid sequence shown in SEQ ID NO: 7, wherein the first and second domains retain the ability to bind to VEGF and the third domain retains the ability to bind to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2).
[0161] In certain embodiments, the fusion proteins described herein contain an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the amino acid sequence shown in SEQ ID NO: 8, wherein the first and second domains retain the ability to bind to VEGF and the third domain retains the ability to bind to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2).
[0162] In certain embodiments, the fusion proteins described herein comprise an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the amino acid sequence set forth in SEQ ID NO: 9, wherein the first and second domains retain the ability to bind to VEGF and the third domain retains the ability to bind to angiopoietin.
[0163] In certain embodiments, the fusion proteins described herein comprise an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the amino acid sequence set forth in SEQ ID NO: 10, wherein the first and second domains retain the ability to bind to VEGF and the third domain retains the ability to bind to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2).
[0164] In certain embodiments, the fusion proteins described herein comprise an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the amino acid sequence set forth in SEQ ID NO: 58, wherein the first, second and fourth domains retain the ability to bind to VEGF and the third domain retains the ability to bind to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2).
[0165] In certain embodiments, the fusion proteins described herein comprise an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the amino acid sequence set forth in SEQ ID NO: 59, wherein the first, second and fourth domains retain the ability to bind to VEGF and the third domain retains the ability to bind to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2).
[0166] In certain embodiments, the fusion proteins described herein comprise an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the amino acid sequence set forth in SEQ ID NO: 60, wherein the first, second, and fourth domains retain the ability to bind to VEGF and the third domain retains the ability to bind to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2).
[0167] In certain embodiments, the fusion proteins described herein comprise an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the amino acid sequence set forth in SEQ ID NO: 61, wherein the first, second, and fourth domains retain the ability to bind to VEGF and the third domain retains the ability to bind to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2).
[0168] In certain embodiments, the fusion proteins described herein comprise an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the amino acid sequence set forth in SEQ ID NO: 62, wherein the first, second, and fourth domains retain the ability to bind to VEGF and the third domain retains the ability to bind to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2).
[0169] In certain embodiments, the fusion proteins described herein comprise an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the amino acid sequence set forth in SEQ ID NO: 63, wherein the first, second, and fourth domains retain the ability to bind to VEGF and the third domain retains the ability to bind to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2).
[0170] In certain embodiments, the fusion proteins described herein comprise an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the amino acid sequence shown in SEQ ID NO: 64, wherein the first domain, the second domain and the fourth domain retain the ability to bind to VEGF and the third domain retains the ability to bind to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2).
[0171] In certain embodiments, the fusion proteins described herein comprise an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the amino acid sequence shown in SEQ ID NO: 65, wherein the first domain, the second domain and the fourth domain retain the ability to bind to VEGF and the third domain retains the ability to bind to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2).
[0172] In certain embodiments, the fusion proteins described herein comprise an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to the amino acid sequence shown in SEQ ID NO: 66, wherein the first domain, the second domain and the fourth domain retain the ability to bind to VEGF and the third domain retains the ability to bind to angiopoietin (e.g., angiopoietin 1 and angiopoietin 2).
[0173] Other exemplary fusion proteins provided herein are described in more detail in the following sections. In some embodiments, the fusion proteins according to any of the above embodiments can incorporate any features, alone or in combination, as described in Sections 5.2.2 to 5.2.4 below.
[0174] 5.2.2. Variants and Modifications of Polypeptides
[0175] In some embodiments, amino acid sequence modifications of the fusion proteins described herein are contemplated. For example, it may be desirable to optimize the binding affinity and / or other biological properties of the fusion proteins, including (but not limited to) specificity, thermal stability, expression level, glycosylation, reduced immunogenicity, or solubility. Accordingly, in addition to the fusion proteins described herein, variants of the fusion proteins described herein are contemplated. For example, peptide variants can be prepared by introducing appropriate nucleotide changes into the encoding DNA and / or by synthesis of the desired polypeptide. Those skilled in the art who understand amino acid changes can alter the post-translational processes of the peptide.
[0176] The changes can be substitutions, deletions, or insertions of one or more codons encoding the polypeptide that result in an amino acid sequence change as compared to the original polypeptide.
[0177] Amino acid substitutions can be due to the replacement of one amino acid with another amino acid having similar structure and / or chemical properties, such as the replacement of leucine with serine, for example, conservative amino acid substitutions. Standard techniques known to those skilled in the art can be used to introduce mutations in the nucleotide sequences encoding the molecules provided herein, including (for example) site-directed mutagenesis and PCR-mediated mutagenesis that result in amino acid substitutions. Insertions or deletions can optionally be in the range of about 1 to 5 amino acids. In certain embodiments, the substitutions, deletions, or insertions include fewer than 25 amino acid substitutions, fewer than 20 amino acid substitutions, fewer than 15 amino acid substitutions, fewer than 10 amino acid substitutions, fewer than 5 amino acid substitutions, fewer than 4 amino acid substitutions, fewer than 3 amino acid substitutions, or fewer than 2 amino acid substitutions relative to the original molecule. In a particular embodiment, the substitutions are conservative amino acid substitutions made at one or more predicted non-essential amino acid residues. Permissible changes can be determined by systematically making insertions, deletions, or substitutions of amino acids in the sequence and testing the activity exhibited by the resulting variants for the parental peptide.
[0178] Amino acid sequence insertions include amino- and / or carboxyl-terminal fusions, which range in length from 1 residue to polypeptides containing multiple residues, as well as insertions within the sequence of a single amino acid residue or multiple amino acid residues. Examples of terminal insertions include polypeptides having an N-terminal methionyl residue.
[0179] The disclosure of the present invention includes fusion proteins produced by conservative amino acid substitutions. In conservative amino acid substitutions, an amino acid residue is replaced with an amino acid residue having a side chain with a similar charge. As described above, families of amino acid residues having side chains with similar charges are defined in the art. These families include amino acids having basic side chains (e.g., lysine, arginine, histidine), acidic side chains (e.g., aspartic acid, glutamic acid), uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine), nonpolar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), β-branched side chains (e.g., threonine, valine, isoleucine), and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine). Alternatively, mutations can be introduced randomly along all or a portion of the coding sequence, such as by saturation mutagenesis, and the resulting mutants can be screened for biological activity to identify mutants that retain activity. After mutagenesis, the encoded protein can be expressed and the protein activity can be determined. Conservative substitutions (e.g., within amino acid groups having similar properties and / or side chains) can be made to maintain or not significantly alter properties. The following table shows exemplary substitutions.
[0180] Table 2. Amino Acid Substitutions
[0181]
[0182]
[0183] Amino acids can be grouped according to the similarity of the properties of their side chains (see, e.g., Lehninger, Biochemistry pp. 73-75 (2nd ed., 1975)): (1) nonpolar: Ala (A), Val (V), Leu (L), Ile (I), Pro (P), Phe (F), Trp (W), Met (M); (2) uncharged polar: Gly (G), Ser (S), Thr (T), Cys (C), Tyr (Y), Asn (N), Gln (Q); (3) acidic: Asp (D), Glu (E); and (4) basic: Lys (K), Arg (R), His (H). Alternatively, naturally occurring residues can be grouped based on common side chain properties: (1) hydrophobic: norleucine, Met, Ala, Val, Leu, Ile; (2) neutral hydrophilic: Cys, Ser, Thr, Asn, Gln; (3) acidic: Asp, Glu; (4) basic: His, Lys, Arg; (5) residues affecting chain orientation: Gly, Pro; and (6) aromatic: Trp, Tyr, Phe.
[0184] For example, any cysteine residue that does not participate in maintaining the correct conformation of the polypeptides provided herein can also be replaced with, e.g., another amino acid such as alanine or serine to improve the oxidative stability of the molecule and prevent aberrant cross-linking.
[0185] Non-conservative substitutions would require replacement of a member of one of these classes with a member of another class.
[0186] Amino acid sequence insertions include amino- and / or carboxyl-terminal fusions ranging in length from one residue to polypeptides of 100 or more residues, as well as insertions within the sequence of single amino acid residues or multiple amino acid residues. Examples of terminal insertions include polypeptides having an N-terminal methionyl residue.
[0187] Changes can be made using methods known in the art such as oligonucleotide-mediated (site-directed) mutagenesis, alanine scanning, and PCR mutagenesis. Site-directed mutagenesis (see, e.g., Carter, Biochem J. 237:1-7 (1986); and Zoller et al., Nucl. Acids Res. 10:6487-500 (1982)), cassette mutagenesis (see, e.g., Wells et al., Gene 34:315-23 (1985)), or other known techniques can be implemented on cloned DNA to generate polypeptide variant DNA.
[0188] Covalent modification of the fusion proteins provided herein is included within the scope of the present disclosure. Covalent modification includes reacting targeted amino acid residues of the polypeptides provided herein with an organic derivatizing reagent capable of reacting with the selected side chains or N- or C-terminal residues of the polypeptide. Other modifications include deamidation of glutaminyl and asparaginyl residues to glutamyl and aspartyl residues, hydroxylation of proline and lysine, phosphorylation of the hydroxyl groups of serine or threonyl residues, methylation of the α-amino groups of lysine, arginine, and histidine side chains (see, e.g., Creighton, Proteins:Structure and Molecular Properties 79-86 (1983)), acetylation of the N-terminal amine, and amidation of any C-terminal carboxyl.
[0189] In some embodiments, for example, the fusion proteins provided herein are chemically modified by covalent attachment of any type of molecule to the fusion protein. Polypeptide derivatives can include polypeptides that have been chemically modified by, for example, glycosylation, acetylation, PEGylation, phosphorylation, amidation, derivatization by known protecting / blocking groups, proteolytic cleavage, conjugation to a cell ligand or other protein, or conjugation to one or more immunoglobulin domains (e.g., Fc or a portion of Fc). Any of a variety of chemical modifications can be carried out by known techniques, including, but not limited to, specific chemical cleavage, acetylation, formulation, metabolic synthesis with tunicamycin, etc. Additionally, the polypeptide can contain one or more non-canonical amino acids.
[0190] In some embodiments, the fusion proteins provided herein are altered to increase or decrease the degree of glycosylation of the fusion protein. Addition or deletion of glycosylation sites in the polypeptide can be conveniently achieved by altering the amino acid sequence such that one or more glycosylation sites are created or removed.
[0191] When the fusion proteins provided herein are fused to an Fc region, the carbohydrate attached thereto can be altered. Native antibodies produced by mammalian cells typically contain branched, biantennary oligosaccharides that are generally attached via an N-linkage to Asn297 in the CH2 domain of the Fc region. See, e.g., Wright et al. TIBTECH 15:26-32 (1997). The oligosaccharide can include a variety of carbohydrates, e.g., mannose, N-acetylglucosamine (GlcNAc), galactose, and sialic acid, as well as fucose attached to GlcNAc in the "stem" of the biantennary oligosaccharide structure. In some embodiments, the oligosaccharides in the binding molecules provided herein can be modified to produce variants having certain improved properties.
[0192] In molecules containing an Fc region, one or more amino acid modifications can be introduced into the Fc region, thereby generating Fc region variants. Fc region variants can contain a human Fc region sequence (e.g., a human IgG1, IgG2, IgG3, or IgG4 Fc region) that contains amino acid modifications (e.g., substitutions) at one or more amino acid positions.
[0193] In some embodiments, the present invention contemplates variants having some, but not all, effector function, making it a desirable candidate for applications where the in vivo half-life of the binding molecule is important, but certain effector functions (such as complement and ADCC) are not necessary or are detrimental. In vitro and / or in vivo cytotoxicity assays can be performed to confirm reduced / depleted CDC and / or ADCC activity. For example, an Fc receptor (FcR) binding assay can be performed to ensure that the binding molecule lacks FcγR binding (and thus likely lacks ADCC activity), but retains the ability to bind FcRn. Non-limiting examples of in vitro assays for evaluating the ADCC activity of a molecule of interest are described in U.S. Patent No. 5,500,362 (see, e.g., Hellstrom, I. et al. Proc. Nat’l Acad. Sci. USA 83:7059-7063 (1986)) and Hellstrom, I et al., Proc. Nat’l Acad. Sci. USA 82:1499-1502 (1985); 5,821,337 (see Bruggemann, M. et al., J. Exp. Med. 166:1351-1361 (1987)). Alternatively, non-radioactive assay methods can be used (see, e.g., ACTI for flow cytometry TM Non-radioactive cytotoxicity assay (Cell Technology, Inc. Mountain View, CA); and CytoTox Non-radioactive cytotoxicity assays (Promega, Madison, WI). Effector cells useful for these assays include peripheral blood mononuclear cells (PBMC) and natural killer (NK) cells. As an alternative or in addition, ADCC activity of a molecule of interest can be evaluated in vivo, for example, in an animal model as disclosed in Clynes et al., Proc. Nat’l Acad. Sci. USA 95:652-656 (1998). A C1q binding assay can also be performed to confirm that the binding molecule does not bind C1q and thus lacks CDC activity. See, for example, the C1q and C3c binding ELISAs in WO 2006 / 029879 and WO 2005 / 100402. To evaluate complement activation, a CDC assay can be performed (see, for example, Gazzano-Santoro et al., J. Immunol. Methods 202:163 (1996); Cragg, M. S. et al., Blood 101:1045-1052 (2003); and Cragg, M. S. and M. J. Glennie, Blood 103:2738-2743 (2004)). FcRn binding and in vivo clearance / half-life determination can also be performed using methods known in the art (see, for example, Petkova, S. B. et al., Int’l. Immunol. 18(12):1759-1769 (2006)).
[0194] Binding molecules with reduced effector factor function include those having substitutions at one or more of Fc region residues 238, 265, 269, 270, 297, 327, and 329 (U.S. Patent No. 6,737,056). These Fc mutants include Fc mutants having substitutions at two or more of amino acid positions 265, 269, 270, 297, and 327, including the so-called "DANA" Fc mutant having residues 265 and 297 replaced with alanine (U.S. Patent No. 7,332,581).
[0195] Certain variants having improved or reduced binding to FcR are described. (See, for example, U.S. Patent No. 6,737,056; WO 2004 / 056312 and Shields et al., J. Biol. Chem. 9(2):6591-6604 (2001)).
[0196] In some embodiments, the variant comprises an Fc region having one or more amino acid substitutions that improve ADCC, for example, substitutions at positions 298, 333, and / or 334 of the Fc region (EU residue numbering).
[0197] In some embodiments, changes are made in the Fc region, which result in altered (i.e., improved or reduced) C1q binding and / or complement-dependent cytotoxicity (CDC), e.g., as described in U.S. Patent No. 6,194,551, WO 99 / 51642, and Idusogie et al., J. Immunol. 164:4178-4184 (2000).
[0198] Binding molecules having an increased half-life and improved binding to the neonatal Fc receptor (FcRn) result in the transfer of maternal IgG to the fetus (Guyer et al., J. Immunol. 117:587 (1976) and Kim et al., J. Immunol. 24:249 (1994)), and such binding molecules are described in US2005 / 0014934A1 (Hinton et al.). Those molecules contain an Fc region having one or more substitutions therein that improve the binding of the Fc region to FcRn. These Fc variants include those having substitutions at one or more Fc region residues: 238, 256, 265, 272, 286, 303, 305, 307, 311, 312, 317, 340, 356, 360, 362, 376, 378, 380, 382, 413, 424, or 434, e.g., substitution of Fc region residue 434 (U.S. Patent No. 7,371,826). For other examples of Fc region variants, see also Duncan & Winter, Nature 322:738-40 (1988); U.S. Patent No. 5,648,260; U.S. Patent No. 5,624,821; and WO 94 / 29351.
[0199] In some embodiments, it may be desirable to generate cysteine-engineered polypeptides in which one or more residues of the polypeptide are replaced with cysteine residues. In some embodiments, the residues at which replacement can occur are accessible sites of the peptide. By replacing those residues with cysteine, a reactive thiol group is thereby positioned at an accessible site of the peptide and the reactive thiol group can be used to conjugate the peptide to other moieties, such as a drug moiety or a linker-drug moiety, to generate an immunoconjugate, as further described herein.
[0200] 5.2.3. Preparation of Fusion Proteins
[0201] Also provided herein are methods for preparing the various fusion proteins provided herein.
[0202] In a specific embodiment, the fusion proteins provided herein are recombinantly expressed. Recombinant expression of the fusion proteins provided herein may require the construction of an expression vector containing a polynucleotide encoding the protein or a fragment thereof. Once a polynucleotide encoding the protein or a fragment thereof provided herein is obtained, vectors for producing the molecule can be generated by DNA recombination techniques using techniques well known in the art. Accordingly, methods for preparing proteins by expressing polynucleotides containing the encoding nucleotide sequences are described herein. Methods well known to those skilled in the art can be used to construct expression vectors containing the coding sequence and appropriate transcriptional and translational control signals. These methods include, for example, in vitro recombinant DNA techniques, synthetic techniques, and in vivo genetic recombination. Also provided are replicable vectors containing a nucleotide sequence encoding the fusion protein or a fragment thereof provided herein operably linked to a promoter.
[0203] The expression vector can be transferred into a host cell by conventional techniques, and then the transfected cells can be cultured by conventional techniques to produce the fusion protein provided herein. Accordingly, host cells containing a polynucleotide encoding the fusion protein or a fragment thereof provided herein operably linked to a heterologous promoter are also provided.
[0204] A variety of host-expression vector systems can be used to express the fusion proteins provided herein. These host-expression systems represent the vehicles by which the coding sequences of interest can be produced and subsequently purified, and also represent the cells that can express in situ the fusion proteins provided herein when transformed or transfected with the appropriate nucleotide coding sequences. These include (but are not limited to) microorganisms such as bacteria (e.g., Escherichia coli and Bacillus subtilis) transformed with recombinant phage DNA, plasmid DNA, or cosmid DNA expression vectors containing the coding sequences; yeast (e.g., Saccharomyces pichia) transformed with recombinant yeast expression vectors containing the coding sequences; insect cell systems infected with recombinant viral expression vectors (e.g., baculovirus) containing the coding sequences; plant cell systems infected with recombinant viral expression vectors (e.g., cauliflower mosaic virus, CaMV; tobacco mosaic virus, TMV) or transformed with recombinant plasmid expression vectors (e.g., Ti plasmid) containing the coding sequences; or mammalian cell systems (e.g., COS, CHO, BHK, 293, NS0, and 3T3 cells) having recombinant expression constructs with promoters derived from mammalian cell genomes (e.g., metallothionein promoter) or promoters derived from mammalian viruses (e.g., adenovirus late promoter; vaccinia virus 7.5K promoter). Bacterial cells, such as Escherichia coli, or eukaryotic cells, particularly eukaryotic cells used for the expression of intact recombinant molecules, can be used for the expression of recombinant fusion proteins. For example, in combination with vectors from the human cytomegalovirus, such as the major immediate early gene promoter element, mammalian cells such as Chinese hamster ovary cells (CHO) are efficient expression systems for antibodies or variants thereof. In a specific embodiment, the expression of the nucleotide sequences encoding the fusion proteins provided herein is regulated by a constitutive promoter, an inducible promoter, or a tissue-specific promoter.
[0205] In bacterial systems, some expression vectors can be advantageously selected based on the intended use of the fusion protein to be expressed. For example, when large amounts of such fusion protein are to be produced, for the production of a pharmaceutical composition of the fusion protein, a vector that directs high-level expression of a fusion protein product that is easy to purify may be desired. These vectors include (but are not limited to) the E. coli expression vector pUR278 (Ruther et al., EMBO 12:1791 (1983)), in which the coding sequence can be ligated separately into the vector in frame with the lac Z coding region, thereby producing a fusion protein; pIN vectors (Inouye & Inouye, Nucleic Acids Res. 13:3101-3109 (1985); Van Heeke & Schuster, J. Biol. Chem. 24:5503-5509 (1989)); etc. pGEX vectors can also be used to express foreign polypeptides as fusion proteins with glutathione S-transferase (GST). Generally, these fusion proteins are soluble and can be easily purified from lysed cells by adsorption and binding to the matrix glutathione agarose beads and then elution in the presence of free glutathione. The pGEX vectors are designed to include thrombin or factor Xa protease cleavage sites so that the cloned target gene product can be released from the GST moiety.
[0206] In mammalian host cells, some virus-based expression systems can be used. If adenovirus is used as an expression vector, the coding sequence of interest can be ligated to an adenovirus transcriptional / translational control complex, e.g., a late promoter and tripartite leader sequence. Then, this chimeric gene can be inserted into the adenovirus genome by in vitro or in vivo recombination. Insertion in a non-essential region of the viral genome (e.g., the E1 or E3 region) will result in the production of a recombinant virus that is viable in the infected host and capable of expressing the fusion protein (e.g., see Logan & Shenk, Proc. Natl. Acad. Sci. USA 81:355-359 (1984)). For efficient translation of the inserted coding sequence, specific initiation signals may also be required. These signals include the ATG initiation codon and adjacent sequences. In addition, the initiation codon must be in phase with the reading frame of the desired coding sequence to ensure translation of the entire insert. These foreign translational control signals and initiation codons can be of multiple origins, i.e., both natural and synthetic origins. Expression efficiency can be enhanced by including appropriate transcriptional enhancer elements, transcriptional terminators, etc. (see, e.g., Bittner et al., Methods in Enzymol. 153:51-544 (1987)).
[0207] Alternatively, it is possible to modulate the expression of the inserted sequence, or to alter and process the host cell line of the gene product in a desired specific manner. These modifications (e.g., glycosylation) and processing (e.g., cleavage) of the protein product can be important for the function of the protein. Different host cells have characteristic and specific mechanisms for the post-translational processing and modification of proteins and gene products. A suitable cell line or host system can be selected to ensure the correct modification and processing of the expressed foreign protein. For this purpose, eukaryotic host cells with cellular mechanisms for the appropriate processing of primary transcripts, glycosylation, and phosphorylation of gene products can be used. These mammalian host cells include (but are not limited to) CHO, VERY, BHK, Hela, COS, MDCK, 293, 3T3, W138, BT483, Hs578T, HTB2, BT2O, and T47D, NS0 (a murine myeloma cell line that does not endogenously produce any immunoglobulin chains), CRL7O3O, and HsS78Bst cells.
[0208] For the long-term, high-yield production of recombinant proteins, stable expression can be used. For example, cell lines that stably express fusion proteins can be engineered. Without using an expression vector containing a viral origin of replication, host cells can be transformed with DNA controlled by suitable expression control elements (e.g., promoters, enhancers, sequences, transcription terminators, polyadenylation sites, etc.) and selectable markers. After the introduction of the foreign DNA, the engineered cells can be allowed to grow in enriched medium for 1-2 days and then transferred to selective medium. The selectable marker in the recombinant plasmid confers selection tolerance and allows the cells to stably integrate the plasmid into their chromosomes and grow to form foci that can in turn be cloned and expanded into cell lines. This method can be advantageously used to engineer cell lines that express fusion proteins. These engineered cell lines can be particularly useful in the screening and evaluation of compositions that directly or indirectly interact with binding molecules.
[0209] A number of selection systems can be used, including (but not limited to) the herpes simplex virus thymidine kinase (Wigler et al., Cell 11:223 (1977)), hypoxanthine-guanine phosphoribosyltransferase (Szybalska & Szybalski, Proc. Natl. Acad. Sci. USA 48:202 (1992)), and adenine phosphoribosyltransferase (Lowy et al., Cell 22:8-17 (1980)) genes, which can be used in tk-, hgprt-, or aprt- cells, respectively. Additionally, antimetabolite resistance can be used as the basis for selection of the following genes: dhfr, which confers resistance to methotrexate (Wigler et al., Natl. Acad. Sci. USA 77:357 (1980); O'Hare et al., Proc. Natl. Acad. Sci. USA 78:1527 (1981)); gpt, which confers resistance to mycophenolic acid (Mulligan & Berg, Proc. Natl. Acad. Sci. USA 78:2072 (1981)); neo, which confers resistance to the aminoglycoside G-418 (Wu and Wu, Biotherapy 3:87-95 (1991); Tolstoshev, Ann. Rev. Pharmacol. Toxicol. 32:573-596 (1993); Mulligan, Science 260:926-932 (1993); and Morgan and Anderson, Ann. Rev. Biochem. 62:191-217 (1993); May, TIB TECH 11(5):l55-2 15(1993)); and hygro, which confers resistance to hygromycin (Santerre et al., Gene 30:147 (1984)). Methods commonly known in the art of DNA recombination can be routinely applied to select the desired recombinant clones, and these methods are described in, for example, Ausubel et al. (eds.), Current Protocols in Molecular Biology , John Wiley & Sons, NY (1993); Kriegler, Gene Transfer and Expression , A Laboratory Manual, Stockton Press, NY (1990); and in Chapters 12 and 13, Dracopoli et al. (eds.), Current Protocols in Human Genetics, John Wiley & Sons, NY (1994); Colberre - Garapin et al., J. Mol. Biol. 150:1 (1981), the above - mentioned literature is incorporated herein by reference in its entirety.
[0210] The expression level of the fusion protein can be increased by vector amplification (for a review, see Bebbington and Hentschel, The use of vectors based on gene amplification for the expression of cloned genes in mammalian cells in DNA cloning, Volume 3 (Academic Press, New York, 1987)). When the marker in the vector system expressing the fusion protein is amplifiable, an increase in the level of inhibitor present in the host cell culture will increase the number of copies of the marker gene. Since the amplified region is related to the fusion protein gene, the production of the fusion protein will also increase (Crouse et al., Mol. Cell. Biol. 3:257 (1983)).
[0211] Host cells can be co - transfected with multiple expression vectors provided herein. The vectors can contain the same selectable marker, which enables equal expression of each encoded polypeptide. As an alternative, a single vector can be used, which encodes and is capable of expressing multiple polypeptides. The coding sequence can comprise cDNA or genomic DNA.
[0212] Once the fusion proteins provided herein are produced by recombinant expression, they can be purified by any method known in the art for purifying polypeptides (e.g., immunoglobulin molecules), e.g., by chromatography (e.g., ion - exchange chromatography, affinity chromatography, specifically affinity chromatography for a specific antigen after protein A, size - exclusion column chromatography, and kappa - selective affinity chromatography), centrifugation, differential solubility, or by any other standard technique for protein purification. In addition, the fusion protein molecules provided herein can be fused to heterologous polypeptide sequences described herein or otherwise known in the art to aid in purification.
[0213] The following describes in more detail various aspects of the recombinant production of the fusion proteins provided herein in the context of prokaryotic or eukaryotic cells.
[0214] Recombinant production in prokaryotic cells
[0215] Polynucleotide sequences encoding the fusion proteins disclosed by the present invention can be obtained using standard recombinant techniques. For example, polynucleotides can be synthesized using a nucleotide synthesizer or PCR techniques. Once obtained, the sequence encoding the polypeptide is inserted into a recombinant vector capable of replicating and expressing heterologous polynucleotides in a prokaryotic host. A variety of vectors available and known in the art can be used for the purposes disclosed by the present invention. The choice of the appropriate vector will mainly depend on the size of the nucleic acid to be inserted into the vector and the specific host cell to be transformed with the vector. Each vector contains a variety of components based on its function (amplification or expression of the heterologous polynucleotide or both) and its compatibility with the specific host cell in which it is present. The vector components generally include (but are not limited to): an origin of replication, a selectable marker gene, a promoter, a ribosome binding site (RBS), a signal sequence, a heterologous nucleic acid insert, and a transcription termination sequence.
[0216] Generally, plasmid vectors containing replicons and control programs derived from species compatible with these hosts are used in combination with these hosts. The vector usually has a replication site and a marker sequence capable of providing phenotypic selection in the transformed cells. For example, pBR322, a plasmid derived from the Escherichia coli (E. coli) species, is commonly used to transform Escherichia coli (E. coli). pBR322 contains genes encoding ampicillin (Amp) and tetracycline (Tet) resistance and thus provides an easy way to identify transformed cells. pBR322, its derivatives, or other microbial plasmids or phages may also contain or be modified to contain promoters that can be used by the microorganism to express endogenous proteins. Examples of pBR322 derivatives for expressing specific antibodies are described in detail in Carter et al., U.S. Patent No. 5,648,237.
[0217] In addition, phage vectors containing replicons and control sequences compatible with the host microorganism can be used as transformation vectors related to these hosts. For example, phages such as GEM TM -11 can be used to prepare recombinant vectors that can be used to transform susceptible host cells such as Escherichia coli (E. coli) LE392.
[0218] The expression vectors of the present invention application can contain two or more promoter-cistron pairs, each of which encodes one of the polypeptide components. A promoter is an untranslated regulatory sequence located upstream (5') of the cistron that regulates the expression of the cistron. Prokaryotic promoters are generally divided into two categories, inducible and constitutive. An inducible promoter is a promoter that responds to changes in culture conditions, such as the presence or absence of nutrients or temperature changes, and initiates an increased transcriptional level of the cistron under its control.
[0219] A large number of promoters recognized by a variety of potential host cells are well known. By removing the promoter from the source DNA via restriction enzyme digestion and inserting the isolated promoter sequence into the vector of the present application, the selected promoter can be operably linked to the cistronic DNA encoding the fusion protein of the present invention. Both native promoter sequences and a variety of heterologous promoters can be used to direct the amplification and / or expression of the target gene. In some embodiments, heterologous promoters are used because they generally permit higher transcription and higher yields of the expressed target gene compared to native target polypeptide promoters.
[0220] Promoters suitable for prokaryotic hosts include the PhoA promoter, the β-galactosidase and lactose promoter systems, the tryptophan (trp) promoter system, and hybrid promoters such as the tac or trc promoters. However, other promoters functional in bacteria (such as other known bacterial or phage promoters) are also suitable. Their nucleic acid sequences have been published, whereby those skilled in the art can operably link them to the cistron encoding the target peptide using linkers or adaptors providing any desired restriction sites (Siebenlist et al. Cell 20:269 (1980)).
[0221] In one aspect, each cistron within the recombinant vector contains a secretory signal sequence component that directs the transmembrane transport of the expressed polypeptide. Generally, the signal sequence can be a component of the vector, or it can be part of the target polypeptide DNA inserted into the vector. The signal sequence selected for the purposes of the present invention should be a signal sequence recognized and processed by the host cell (i.e., cleaved by signal peptidase). For prokaryotic host cells that do not recognize and process native signal sequences of heterologous polypeptides, the signal sequence is replaced with a prokaryotic signal sequence selected from, for example, alkaline phosphatase, penicillinase, Ipp, or the heat-stable enterotoxin II (STII) leader sequence, LamB, PhoE, PelB, OmpA, and MBP. In some embodiments of the present invention, the signal sequence used in both cistrons of the expression system is the STII signal sequence or a variant thereof.
[0222] In some embodiments, the production of the fusion protein disclosed according to the present invention can occur in the cytoplasm of the host cell and thus does not require the presence of a secretory signal sequence within each cistron. Certain host strains (e.g., the E. coli trxB - strain) provide cytoplasmic conditions favorable for disulfide bond formation, thereby allowing for the correct folding and assembly of the expressed protein subunits.
[0223] Prokaryotic host cells suitable for expressing the fusion proteins disclosed by the present invention include archaebacteria and eubacteria, such as Gram-negative or Gram-positive organisms. Examples of useful bacteria include Escherichia (e.g., Escherichia coli), Bacilli (e.g., Bacillus subtilis), Enterobacteria, Pseudomonas species (e.g., Pseudomonas aeruginosa), Salmonella typhimurium, Serratia marcescens, Klebsiella, Proteus, Shigella, Rhizobia, Vitreoscilla, or Paracoccus. In some embodiments, Gram-negative cells are used. In one embodiment, Escherichia coli cells are used as the host. Examples of Escherichia coli strains include strain W3110 (Bachmann, Cellular and Molecular Biology, Volume 2 (Washington, D.C.: American Society for Microbiology, 1987), pp. 1190-1219; ATCC Deposit No. 27,325) and its derivatives, including strain 33D3, which has the genotype W3110 ΔfhuA (ΔtonA) ptr3 lacIq lacL8 ΔompT Δ(nmpc-fepE) degP41 kan R(U.S. Patent No. 5,639,635). Other strains and their derivatives, such as Escherichia coli (E. coli) 294 (ATCC 31,446), Escherichia coli (E. coli) B, Escherichia coli (E. coli) 1776 (ATCC 31,537), and Escherichia coli (E. coli) RV308 (ATCC 31,608) are also suitable. These examples are illustrative rather than limiting. Methods for constructing derivatives of any of the above bacteria with a defined genotype are known in the art and described, for example, in Bass et al., Proteins, 8:309-314 (1990). It is generally necessary to consider the replicability of the replicon in the bacterial cell to select an appropriate bacterium...
Claims
1. A polypeptide, comprising: (i) a first domain that binds to VEGF, wherein the first domain consists of the amino acid sequence of SEQ ID NO: 1; (ii) a second domain that binds to VEGF, wherein the second domain consists of the amino acid sequence of SEQ ID NO: 2; and (iii) a third domain that binds to angiopoietin, wherein the third domain consists of the amino acid sequence of SEQ ID NO: 54; and (iv) a VEGFC-binding domain, wherein the VEGFC-binding domain consists of the amino acid sequence of SEQ ID NO: 57, wherein, the polypeptide sequentially comprises, from the N-terminus to the C-terminus, the third domain, the first domain, the second domain, and the VEGFC-binding domain.
2. The polypeptide according to claim 1, further comprising the Fc region of an antibody or a variant thereof.
3. The polypeptide according to claim 2, wherein the Fc region comprises the amino acid sequence shown in SEQ ID NO:
5.
4. The polypeptide according to claim 1, further comprising a signal peptide.
5. The polypeptide according to claim 4, wherein the signal peptide comprises the amino acid sequence shown in SEQ ID NO:
6.
6. The polypeptide according to claim 1, further comprising one or more linkers.
7. A polypeptide, comprising the amino acid sequence shown in SEQ ID NO:
65.
8. An isolated nucleic acid, comprising a nucleic acid sequence encoding the polypeptide according to claim 1.
9. A vector, comprising the isolated nucleic acid according to claim 8.
10. The vector according to claim 9, wherein the vector is an adeno-associated virus (AAV) vector.
11. The vector according to claim 10, wherein the AAV vector is derived from AAV1, AAV2, AAV2i8, AAV3, AAV3-B, AAV4, AAV5, AAV6, AAV7, AAV8, AAVrh8, AAVrh8R, AAV9, AAV10, AAVrh10, AAV11, AAV12, AAV13, AAV-DJ, AAV LK03, AAVrh74, AAV44-9 or a combination or variant thereof.
12. A recombinant AAV (rAAV) vector, comprising a nucleic acid encoding a polypeptide, wherein the polypeptide comprises: (i) a first domain derived from VEGFR-1, wherein the first domain consists of the amino acid sequence of SEQ ID NO: 1; (ii) a second domain derived from VEGFR-2, wherein the second domain consists of the amino acid sequence of SEQ ID NO: 2; (iii) a third domain capable of binding to angiopoietin, wherein the third domain consists of the amino acid sequence of SEQ ID NO: 54; and (iv) a fourth domain capable of binding to VEGFC, wherein the fourth domain consists of the amino acid sequence of SEQ ID NO: 57, wherein, The polypeptide sequentially comprises a third domain, a first domain, a second domain, and a fourth domain from the N-terminus to the C-terminus. The rAAV vector comprises inverted terminal repeats (ITRs) from the following: AAV1, AAV2, AAV2i8, AAV3, AAV3-B, AAV4, AAV5, AAV6, AAV7, AAV8, AAVrh8, AAVrh8R, AAV9, AAV10, AAVrh10, AAV11, AAV12, AAV13, AAV-DJ, AAV LK03, AAVrh74, or AAV44-9.
13. A vector comprising the nucleic acid sequence shown in SEQ ID NO:
74.
14. A recombinant AAV (rAAV) particle comprising (a) a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a first domain derived from VEGFR-1, wherein the first domain consists of the amino acid sequence of SEQ ID NO: 1; (ii) a second domain derived from VEGFR-2, wherein the second domain consists of the amino acid sequence of SEQ ID NO: 2; and (iii) a third domain capable of binding to angiopoietin, wherein the third domain consists of the amino acid sequence of SEQ ID NO: 54; and (iv) a fourth domain capable of binding to VEGFC, wherein the fourth domain consists of the amino acid sequence of SEQ ID NO: 57; and (b) a capsid protein of AAV1, AAV2, AAV2i8, AAV3, AAV3-B, AAV4, AAV5, AAV6, AAV7, AAV8, AAVrh8, AAVrh8R, AAV9, AAV10, AAVrh10, AAV11, AAV12, AAV13, AAV-DJ, AAV LK03, AAVrh74, AAV44-9 or a variant thereof, wherein the polypeptide sequentially comprises a third domain, a first domain, a second domain, and a fourth domain from the N-terminus to the C-terminus.
15. The rAAV particle according to claim 14, wherein the capsid protein is a variant of the AAV2 capsid protein comprising the amino acid sequence shown in SEQ ID NO: 48, and the variant of the AAV2 capsid protein comprises amino acid substitutions Y444F, R487G, T491V, Y500F, R585S, R588T, and Y730F in the capsid protein VP1 of AAV2.
16. A pharmaceutical composition comprising the polypeptide according to any one of claims 1 to 7, the vector or rAAV vector according to any one of claims 9 to 13, or the rAAV particle according to claim 14 or 15 and a pharmaceutically acceptable excipient. Use of the polypeptide according to any one of claims 1 to 7, the vector or rAAV according to any one of claims 9 to 13, or the rAAV particle according to claim 14 or 15, in the preparation of a medicament for treating a disease or disorder in a subject, wherein the disease or disorder is selected from uveitis, retinitis pigmentosa, neovascular glaucoma, diabetic retinopathy (DR), ischemic retinopathy, intraocular neovascularization, age-related macular degeneration (AMD), diabetic macular edema, retinal vein occlusion, macular edema.
18. The use according to claim 17, wherein the diabetic retinopathy (DR) is proliferative diabetic retinopathy.
19. The use according to claim 17, wherein the intraocular neovascularization is retinal neovascularization.
20. The use according to claim 17, wherein the macular edema is diabetic macular edema (DME) or macular edema after retinal vein occlusion (RVO).
21. The use according to claim 17, wherein the retinal vein occlusion is central retinal vein occlusion or branch retinal vein occlusion.
22. The use according to claim 17, wherein the ischemic retinopathy is diabetic retinal ischemia.
23. The use according to claim 17, wherein the disease or disorder is age-related macular degeneration (AMD).
24. The use according to claim 23, wherein the AMD is wet AMD (wAMD).
Citation Information
Patent Citations
Altered antibodies
EP0307434A1
Modified interleukin-2 and production thereof
EP0367166A1
Chimaeric CD4-immunoglobulin polypeptides
EP0394827A1
Cytotoxic agents comprising maytansinoids and their therapeutic use
EP0425235B1
X-ray sensor system for intraoral tomography
US11911197B2