Fusion proteins and products for hydroxylated amino acids
By expressing fusion proteins in yeast, the problem of low tetramer formation efficiency of prolyl 4-hydroxylase in existing technologies has been solved, achieving efficient protein hydroxylation and compound production.
Patent Information
- Application Number
- CN201980052539.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-08-17
- Filing Date
- 2019-08-16
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2039-12-17
AI Technical Summary
Existing technologies struggle to efficiently prepare fusion proteins for commercial applications, particularly in yeast and bacteria where expression and effective formation of prolyl 4-hydroxylase tetramers can be challenging, impacting compound production efficiency.
By expressing a fusion protein in yeast, which contains the α and β subunits of prolyl 4-hydroxylase, and using overlap extension PCR technology to link DNA fragments, a functional hydroxylase is formed, reducing the number of transfection reactions and improving the efficiency of enzyme tetramer formation.
This method enables the efficient formation of prolyl 4-hydroxylase tetramers in yeast, improving protein hydroxylation efficiency and production efficiency while simplifying the preparation process.
Smart Images

Figure CN112566927B_ABST
Abstract
Description
[0001] SEQUENCE LISTING
[0002] The instant application contains a Sequence Listing which has been submitted electronically in ASCII format and is hereby incorporated by reference in its entirety. TECHNICAL FIELD
[0003] Described herein are engineered proteins and their use in fermentation, methods for producing proteins, and methods for in vitro and in vivo hydroxylation of proteins. BACKGROUND
[0004] There is an entire industry that uses microorganisms to make compounds for commercial applications. Microorganisms are often engineered with DNA necessary to make these compounds. Examples of these microorganisms include yeast and bacteria. Compounds made include pharmaceuticals, fragrances, flavors, proteins, and the like.
[0005] Fusion proteins are produced by joining two or more genes that originally made separate proteins. One purpose in producing fusion proteins in drug development is to impart the properties of each "parent" protein to the resulting fusion protein. SUMMARY
[0006] In some embodiments, the present disclosure provides a fusion protein comprising: a prolyl 4-hydroxylase alpha subunit; and a soluble protein chaperone. In some embodiments, the present disclosure provides a fusion protein encoded by: a DNA sequence encoding a prolyl 4-hydroxylase alpha subunit; and a DNA sequence encoding a soluble protein chaperone.
[0007] In some embodiments, the prolyl 4-hydroxylase alpha subunit is selected from the group consisting of: prolyl 4-hydroxylase alpha subunit-1, prolyl 4-hydroxylase alpha subunit-2, and prolyl 4-hydroxylase alpha subunit-3. In some embodiments, the soluble protein chaperone is selected from the group consisting of: prolyl 4-hydroxylase beta subunit, maltose binding protein, small ubiquitin-like modifier, calmodulin binding protein, and glutathione S-transferase. In certain embodiments, the prolyl 4-hydroxylase alpha subunit is from a species selected from the group consisting of: bovine, human, rat, mouse, bacteria, virus, fish, and C. elegans.
[0008] In some embodiments, the present disclosure provides a fusion protein comprising: a prolyl 4-hydroxylase alpha subunit-1; and a prolyl 4-hydroxylase beta subunit. In some embodiments, the present disclosure provides a fusion protein comprising: a DNA sequence encoding a prolyl 4-hydroxylase alpha subunit; and a DNA sequence encoding a prolyl 4-hydroxylase beta subunit. In certain embodiments, the prolyl 4-hydroxylase alpha subunit-1 is located at the N-terminus of the fusion protein. In particular embodiments, the prolyl 4-hydroxylase beta subunit is located at the C-terminus of the fusion protein.
[0009] In some embodiments, the present disclosure provides a fusion protein comprising: a prolyl 4-hydroxylase alpha subunit-1; and a prolyl 4-hydroxylase beta subunit, wherein the prolyl 4-hydroxylase alpha subunit-1 is located at the N-terminus of the fusion protein and the prolyl 4-hydroxylase beta subunit is located at the C-terminus of the fusion protein.
[0010] In certain embodiments, the prolyl 4-hydroxylase alpha subunit is from a species selected from the group consisting of: bovine, human, rat, mouse, bacteria, virus, fish, and Caenorhabditis elegans. In some embodiments, the prolyl 4-hydroxylase alpha subunit-1 is encoded by the nucleic acid of SEQ ID NO: 1 and the prolyl 4-hydroxylase beta subunit is encoded by the nucleic acid of SEQ ID NO: 2.
[0011] In some embodiments, the present disclosure provides a microorganism comprising any of the fusion proteins disclosed herein. In some embodiments, the present disclosure provides a microorganism comprising: a fusion protein comprising a prolyl 4-hydroxylase alpha subunit-1 and a prolyl 4-hydroxylase beta subunit. In some embodiments, the present disclosure provides a microorganism comprising: a fusion protein comprising a prolyl 4-hydroxylase alpha subunit-1 located at the N-terminus and a prolyl 4-hydroxylase beta subunit located at the C-terminus. In some embodiments, the present disclosure provides a microorganism comprising:
[0012] In some embodiments, the present disclosure provides a microorganism comprising: a fusion protein comprising a prolyl 4-hydroxylase alpha subunit-1 and a prolyl 4-hydroxylase beta subunit; and a second protein to be hydroxylated. In certain embodiments, the microorganism is selected from the group consisting of: Bacillus, Escherichia coli, and filamentous fungi. In some embodiments, the microorganism is a yeast. In particular embodiments, the second protein is selected from the group consisting of: collagen, recombinant collagen, collagen-like proteins, and the like. In some embodiments, the prolyl 4-hydroxylase alpha subunit-1 is encoded by the nucleic acid of SEQ ID NO: 1 and the prolyl 4-hydroxylase beta subunit is encoded by the nucleic acid of SEQ ID NO: 2.
[0013] In some embodiments, the present disclosure provides a method for providing a skin care benefit to the skin of an individual, the method comprising: applying a fusion protein disclosed herein to the skin. In certain embodiments, the fusion protein is formulated into a composition selected from the group consisting of a cream, a lotion, an ointment, a gel, a serum, and combinations thereof. In some embodiments, the skin care benefit is selected from the group consisting of anti-wrinkle, improving skin pigmentation, hydration, reducing acne, preventing acne, reducing blackheads, preventing blackheads, reducing stretch marks, preventing stretch marks, preventing cellulite, reducing cellulite, and combinations thereof. In certain embodiments, the fusion protein is combined with other skin care benefit ingredients selected from the group consisting of salicylic acid, retinol, benzoyl peroxide, vitamin C, glycerin, alpha-hydroxy acids, hydroquinone, kojic acid, hyaluronic acid, and combinations thereof.
[0014] In some embodiments, the present disclosure provides an in vitro method for hydroxylating a protein, the method comprising: providing a microorganism containing a protein to be hydroxylated; providing a fusion protein disclosed herein; lysing the microorganism to produce a lysate; adding a specific concentration of the fusion protein to the lysate; and incubating the lysate and the fusion protein under reaction conditions that promote hydroxylation of the protein by the fusion protein. In some embodiments, the lysate is purified prior to adding the fusion protein. In certain embodiments, the concentration of the fusion protein ranges from about 0.05 uM to about 5 uM based on about 1 uM of the protein to be hydroxylated. In particular embodiments, the hydroxylation is performed at a pH ranging from about 5 to about 12. In some embodiments, the hydroxylation is performed at a temperature ranging from about 16 °C to about 40 °C. In certain embodiments, the hydroxylation is performed for greater than about 30 minutes to about 1 hour.
[0015] In some embodiments, the present disclosure provides a method for preparing a hydroxylated protein, the method comprising: providing a microorganism disclosed herein; and growing the microorganism in a culture medium for a time sufficient to hydroxylate a second protein. In certain embodiments, the microorganism is a yeast. In a particular embodiment, the yeast is Pichia pastoris. In some embodiments, the microorganism is grown for about 50 hours to about 72 hours.
[0016] In some embodiments, the present disclosure provides a microorganism comprising: a DNA sequence encoding a prolyl 4-hydroxylase alpha subunit; and a DNA sequence encoding a soluble protein chaperone.
[0017] Additional aspects and embodiments exist in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1MMV-130 is shown, which was used to generate Pichia pastoris strain PP153 as described in Example 1.
[0019] Figure 2 MMV156 is shown, which was used to generate Pichia pastoris strain PP154 as described in Example 3.
[0020] Figure 3 MMV-191, which was used to generate Pichia pastoris strain PP268 as described in Example 3.
[0021] Figure 4 MMV-290 vector is shown, which was generated and transformed into Pichia pastoris strain PP153 as described in Example 1 to generate Pichia pastoris strain PP336 and express a fusion protein having P4HA1 at the N-terminus, P4HB at the C-terminus, and a linker sequence "GSGSGS".
[0022] Figure 5 MMV-289 vector is shown, which was generated and transformed into Pichia pastoris strain PP153 as described in Example 2 to generate Pichia pastoris strain PP335 and express a fusion protein having P4HB at the N-terminus and P4HA1 at the C-terminus.
[0023] Figure 6 MMV-400 vector is shown as described in Example 4 and contains the DNA sequence for the AB fusion protein (i.e., the fusion protein having P4HA1 at the N-terminus and P4HB at the C-terminus as described in Example 3).
[0024] Figure 7 MMV-502 vector is shown as described in Example 5 and contains the DNA sequence for the AB fusion protein, a nucleotide sequence representing six consecutive amino acids of histidine (His tag), two stop codons, and an AOX1 transcription terminator.
[0025] Figure 8 MMV-503 vector is shown as described in Example 5 and contains the C-terminus of the P4HB subunit protein, a nucleotide sequence representing six consecutive amino acids of histidine (His tag), two stop codons, and an AOX1 transcription terminator.
[0026] Figure 9 MMV411 vector used in Example 7 is shown.
[0027] Figure 10 Vector MMV-644 is shown as described in Example 1. DETAILED DESCRIPTION
[0028] All methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, and the practice or testing of the present disclosure. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including the definitions herein, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting.
[0029] In some embodiments, the present disclosure provides a fusion protein encoded by a DNA sequence of a prolyl 4-hydroxylase alpha subunit and a DNA sequence of a soluble protein chaperone. In certain embodiments, the fusion protein comprises a prolyl 4-hydroxylase alpha subunit-1 (P4HA1) and a prolyl 4-hydroxylase beta subunit (P4HB). In certain embodiments, a monomeric prolyl 4-hydroxylase alpha subunit can be used in any of the embodiments in place of the fusion proteins disclosed herein.
[0030] The P4HA gene and the P4HB gene encode the composition of prolyl-4-hydroxylase, a key enzyme in collagen synthesis, which consists of two identical alpha subunits and two beta subunits (heterotetramer). The P4HA encodes one of several different types of alpha subunits and provides the major part of the catalytic site of the active enzyme. See, e.g., Crit Rev Biochem Mol Biol. 45(2): 106-124 (2010). P4HA comprises three domains: a dimerization domain, a substrate binding domain, and a catalytic domain. In some embodiments, the prolyl 4-hydroxylase alpha subunit is from a species selected from the group consisting of bovine, human, rat, mouse, bacteria, virus, fish, and Caenorhabditis elegans. In certain embodiments, the monomeric prolyl 4-hydroxylase alpha subunit is from a species selected from the group consisting of bacteria, virus, fungus, and algae. In certain embodiments, the monomeric prolyl 4-hydroxylase alpha subunit is from a mycovirus (DNA sequence: SEQ ID NO: 15; protein sequence: SEQ ID NO: 16). See, e.g., Rutschmann et al., Appl. Microbiol Biotechnol. 98:4445-4455 (2014) and Shi et al., Protein J. 36:322-331 (2017). In collagen and related proteins, prolyl 4-hydroxylase catalyzes the formation of 4-hydroxyproline, which is important for the proper three-dimensional folding of newly synthesized procollagen chains. The P4HB protein is also known as disulfide isomerase. It is an enzyme that in humans is encoded by the P4HB gene. The human P4HB gene is located in chromosome 17q25. This protein is multifunctional, unlike other prolyl 4-hydroxylase family proteins, and acts as an oxidoreductase for disulfide formation, breakage, and isomerization. The activity of P4HB is tightly regulated, and both dimer dissociation and substrate binding can enhance its enzymatic activity during the catalytic process. In some embodiments, the P4HB is from a species selected from the group consisting of bovine, human, rat, mouse, bacteria, virus, fish, and Caenorhabditis elegans.
[0031] The DNA sequence of P4HA (NCBI Reference Number: XP_005226443.1; UNIPROT: Q1RMU3), the DNA sequence of P4HB (GENBANK: AAI46272.1; UNIPROT: P05307), the DNA sequence of P4HA3 (UNIPROT: P4HA3), and the DNA sequence of P4HA2 (UNIPROT: G3N2F2) are known and commercially available. In some embodiments, the fusion protein is prepared by removing the stop codon from the cDNA sequence encoding the first protein, and then appending the DNA sequence of the second protein in frame by ligation or overlap extension polymerase chain reaction (PCR). The DNA sequence of the fusion protein will then be expressed by the cell as a single protein.
[0032] One technique for making fusion proteins is ligation, which is the joining of two nucleic acid fragments under the action of an enzyme. DNA fragments are joined together to create a recombinant DNA molecule, such as when a foreign DNA fragment is inserted into a plasmid. The ends of the DNA fragments are joined together by the formation of a phosphodiester bond between the 3'-hydroxyl of one DNA end and the 5'-phosphoryl of the other DNA end. Another technique for making fusion proteins is overlap extension PCR, also known as splicing by overlap extension. Overlap extension PCR is used to insert specific mutations at specific points in a sequence or to splice smaller DNA fragments into larger polynucleotides. A secretion signal sequence, such as the Saccharomyces cerevisiae alpha mating factor signal, can be placed in front of the monomeric prolyl 4-hydroxylase alpha subunit to secrete the protein from the host into the production media.
[0033] In some embodiments, the fusion proteins disclosed herein can be encoded by a combination of a DNA sequence of prolyl 4-hydroxylase alpha subunit-1 (P4HA1) or prolyl 4-hydroxylase alpha subunit-2 (P4HA2) or prolyl 4-hydroxylase alpha subunit-3 (P4HA3) and a DNA sequence of prolyl 4-hydroxylase beta subunit (P4HB); and a DNA sequence of prolyl 4-hydroxylase alpha subunit-1 (P4HA1) or prolyl 4-hydroxylase alpha subunit-2 (P4HA2) or prolyl 4-hydroxylase alpha subunit-3 (P4HA3) and a DNA sequence of a soluble chaperone selected from the group consisting of prolyl 4-hydroxylase beta subunit (P4HB), maltose binding protein, small ubiquitin-like modifier, calmodulin binding protein, glutathione S-transferase, and the like. The active prolyl-4-hydroxylase complex can include P4H subunits from species such as bovine, human, rat, mouse, Caenorhabditis elegans, and the like. In one embodiment, the fusion protein comprises P4HA1 and P4HB.
[0034] When preparing the fusion protein described herein, proteins with P4HA or P4HB at the N-terminus can be prepared. Surprisingly, we found that the fusion protein with P4HA at the N-terminus forms a functional hydroxylase in yeast in the presence of free proline, while the fusion protein with P4HB at the N-terminus does not form a functional hydroxylase in yeast. In some embodiments, the fusion protein has P4HA at the N-terminus and a second protein at the C-terminus. In some embodiments, the fusion protein has P4HA at the N-terminus and P4HB at the C-terminus.
[0035] The DNA encoding the fusion protein of P4HA1 and P4HB, or the DNA of the monomeric prolyl 4-hydroxylase α subunit, can be transformed or transfected into an organism. Suitable organisms include yeast, bacteria, fungi, etc. In some embodiments, the bacteria may be Bacillus or Escherichia coli. In some embodiments, the microorganism may be a filamentous fungus. In some embodiments, the organism may be yeast. In some embodiments, the yeast may be Pichia pastoris. Typically, multiple transfection / conversion reactions are required for the hydroxylase to function. The fusion protein described herein enables a more efficient process. The fusion protein described herein reduces the number of conversion reactions to one instead of two (e.g., one for P4HA1 and another for P4HB). If the enzymes are converted separately, they will undergo three reactions to form a tetramer in order to become an effective enzyme. The tetramer consists of, for example, two P4HA subunits and two P4HB subunits. The three reactions are as follows: 1) the first P4HA and the first P4HB combine to form a first dimer, 2) the second P4HA and the second P4HB combine to form a second dimer, and 3) the two dimers form a tetramer. When the enzyme is converted separately, not all P4HA and P4HB react to form a tetramer. The fusion protein will need to react once with another fusion protein to form an efficient tetramer. A beneficial effect of this disclosure is that the fusion protein (two molecules) forms a tetramer more efficiently than the separated proteins (four proteins). Two fusion proteins will form one tetramer. Therefore, the fusion protein described herein provides a more efficient and effective hydroxylase. In some embodiments, the fusion protein can be used in methods for the in vitro hydroxylation of proteins. In some embodiments, the fusion protein can be used in methods for the in vivo hydroxylation of proteins.
[0036] In some embodiments, the fusion protein described herein can be used to hydroxylate proteins in vitro. Microorganisms containing proteins (such as collagen) can be lysed to produce lysates. The lysates can be processed to produce purified proteins. The fusion protein can be added to a purified protein sample or to the lysates. In some embodiments, the cofactor for the hydroxylation reaction can include one or more of the following: ascorbic acid, sodium ascorbate, or iron(II), such as FeSO4. In some embodiments, the substrate for the hydroxylation reaction can be selected from: AKG, molecular collagen, and molecular oxygen. In some embodiments, bovine serum albumin and / or catalase can be added to the reaction to facilitate efficient hydroxylation. The hydroxylation reaction can be carried out at a temperature in the range of about 16°C to about 40°C (e.g., about 32°C). In some embodiments, the hydroxylation reaction can be carried out at approximately 16°C, approximately 17°C, approximately 18°C, approximately 19°C, approximately 20°C, approximately 21°C, approximately 22°C, approximately 23°C, approximately 24°C, approximately 25°C, approximately 26°C, approximately 27°C, approximately 28°C, approximately 29°C, approximately 30°C, approximately 31°C, approximately 32°C, approximately 33°C, approximately 34°C, approximately 35°C, approximately 36°C, approximately 37°C, approximately 38°C, approximately 39°C, or approximately 40°C. Based on the 1 μM protein to be hydroxylated, the amount of fusion protein added to the lysate can range from approximately 0.05 μM to approximately 5 μM, for example, approximately 2.5 μM. In some embodiments, based on the 1 μM protein to be hydroxylated, the amount of fusion protein added to the lysate can be about 0.05 μM, about 0.1 μM, about 0.15 μM, about 0.2 μM, about 0.25 μM, about 0.3 μM, about 0.35 μM, about 0.4 μM, about 0.5 μM, about 0.6 μM, about 0.7 μM, about 0.8 μM, about 0.9 μM, about 1.0 μM, about 1.1 μM, about 1.2 μM, about 1.3 μM, about 1.4 μM, about 1.5 μM, about 1.6 μM, about 1.7 μM, about 1.8 μM, about 1.9 μM, about 2.0 μM, about 2.5 μM, about 3.0 μM, about 3.5 μM, about 4.0 μM, about 4.5 μM, or about 5 μM. Based on the 1 μM protein to be hydroxylated, the amount of fusion protein added to the purified protein can range from 0.05 μM to 5 μM, for example, 2.5 μM.In some embodiments, based on the 1 μM protein to be hydroxylated, the amount of fusion protein added to the purified protein can be about 0.05 μM, about 0.1 μM, about 0.15 μM, about 0.2 μM, about 0.25 μM, about 0.3 μM, about 0.35 μM, about 0.4 μM, about 0.5 μM, about 0.6 μM, about 0.7 μM, about 0.8 μM, about 0.9 μM, about 1.0 μM, about 1.1 μM, about 1.2 μM, about 1.3 μM, about 1.4 μM, about 1.5 μM, about 1.6 μM, about 1.7 μM, about 1.8 μM, about 1.9 μM, about 2.0 μM, about 2.5 μM, about 3.0 μM, about 3.5 μM, about 4.0 μM, about 4.5 μM, or about 5 μM. In some embodiments, hydroxylation is performed at a pH in the range of about 5 to about 12 (e.g., about 7.5). In some embodiments, the pH can be about 5.0, about 5.5, about 6, about 6.5, about 7, about 7.5, about 8, about 8.5, about 9.0, about 9.5, about 10.0, about 10.5, about 11, about 11.5, or about 12. In some embodiments, hydroxylation is performed for more than about 30 minutes to about 5 hours, for example, about 1 hour. In some embodiments, hydroxylation is performed for more than 30 minutes, about 45 minutes, about 1 hour, about 1.5 hours, about 2 hours, about 2.5 hours, about 3 hours, about 3.5 hours, about 4 hours, about 4.5 hours, or about 5 hours. After the reaction, the fusion protein can be inactivated by adding acid to lower the pH of the solution to 4 or by adding 50% to 80% methanol. In the implementation scheme, in vitro hydroxylation can be performed using any of the methods disclosed in U.S. Patent No. 7,932,053, the entire contents of which are incorporated herein by reference.
[0037] Alternatively, the DNA sequence of the fusion protein can be transfected into microorganisms and used for intracellular / in vivo protein hydroxylation. The transfected microorganisms can be grown in a culture medium suitable for the specific microorganism under conditions well known to those skilled in the art. In some embodiments, suitable culture media may be, for example, LB (lysozyme broth) for *Escherichia coli*, BMGY (buffered glycerol complex medium) for *Pichia*, YPD (yeast extract peptone dextran) for *Pichia*, or HMP (sodium hexametaphosphate) for *Pichia*. The temperature of the culture medium can be in the range of about 16°C to 42°C. In some embodiments, the temperature of the culture medium can be about 16°C, about 18°C, about 20°C, about 22°C, about 24°C, about 26°C, about 28°C, about 29°C, about 30°C, about 31°C, about 32°C, about 33°C, about 34°C, about 35°C, about 36°C, about 38°C, about 40°C, or about 42°C. In some embodiments, the microorganism is *Pichia pastoris*, and the temperature of the culture medium can be in the range of about 28°C to about 36°C, for example, about 32°C. In some embodiments, the temperature of the culture medium can be about 28°C, about 29°C, about 30°C, about 31°C, about 32°C, about 33°C, about 34°C, about 35°C, or about 36°C. The microorganism can grow for a period of about 50 hours to about 72 hours (e.g., about 68 hours). In some embodiments, the microorganism can grow for about 50 hours, about 51 hours, about 52 hours, about 53 hours, about 54 hours, about 55 hours, about 56 hours, about 57 hours, about 58 hours, about 59 hours, about 60 hours, about 61 hours, about 62 hours, about 63 hours, about 64 hours, about 65 hours, about 66 hours, about 67 hours, about 68 hours, about 69 hours, about 70 hours, about 71 hours, or about 72 hours. In some implementations, the substrate used for the hydroxylation reaction may be selected from the group consisting of AKG, molecular collagen, and molecular oxygen.
[0038] In some embodiments, the DNA sequence of the fusion protein may be placed in the vector along with the following sequences: the DNA sequence of the fusion protein promoter; the DNA sequence of the fusion protein terminator; a selectable DNA sequence, the DNA sequence of the selectable promoter; the DNA sequence of the selectable terminator; a DNA sequence of the origin of replication, one of which is a bacterial origin of replication and the other is a yeast origin of replication; and / or a DNA sequence homologous to the yeast genome (optionally used to improve efficiency when transforming into yeast). In some embodiments, the vector has been inserted into an organism (or has become an appendage thereto). In some embodiments, the vector can then be transformed into a microorganism by methods known in the art, such as electroporation.
[0039] DNA encoding a fusion protein of prolyl 4-hydroxylase α subunit-1 (P4HA1) and prolyl 4-hydroxylase β subunit (P4HB), as well as DNA encoding a second protein to be hydroxylated, can be transformed into microorganisms. Hydroxylation modification can be performed on a variety of amino acids, including but not limited to proline, lysine, asparagine, aspartic acid, and histidine. Suitable proteins that can be hydroxylated include collagen. In any embodiment, method, and / or reaction described herein, the monomeric prolyl 4-hydroxylase α subunit can be used instead of the fusion protein.
[0040] In some embodiments, the DNA sequence of the fusion protein may be placed in the vector along with the following sequences: the DNA sequence of the fusion protein promoter; the DNA sequence of the fusion protein terminator; a selectable DNA sequence, the DNA sequence of the selectable promoter; the DNA sequence of the selectable terminator; a DNA sequence of the origin of replication, one of which is a bacterial origin of replication and the other is a yeast origin of replication; and / or a DNA sequence homologous to the host organism's genome. In some embodiments, the DNA sequence of the second protein to be hydroxylated may be placed in the vector along with the following sequences: the DNA sequence of the second protein promoter; the DNA sequence of the second protein terminator; a selectable DNA sequence, the DNA sequence of the selectable promoter; the DNA sequence of the selectable terminator; a DNA sequence of the origin of replication, one of which is a bacterial origin of replication and the other is a yeast origin of replication; and / or a DNA sequence homologous to the host organism's genome. In some embodiments, both vectors are then transformed into microorganisms using methods known in the art, such as electroporation.
[0041] Alternatively, in some embodiments, an all-in-one vector may be used, wherein the DNA of the fusion protein, including a promoter and a terminator; the DNA of the second protein, including a promoter and a terminator; the DNA of a selection marker, including a promoter and a terminator; and / or DNA homologous to the genome of an organism for integration into the genome is contained in the all-in-one vector. The all-in-one vector can then be transformed into microorganisms using methods known in the art, such as electroporation.
[0042] Promoters are known in the art to increase protein yield. A promoter is a DNA sequence contained in a vector. Suitable promoters used in this disclosure include, but are not limited to, the AOXl methanol-induced promoter, the pDF derepressor promoter, the pCAT derepressor promoter, the Dasl-Das2 methanol-induced bidirectional promoter, the pHTXl constitutive bidirectional promoter, the pGCW14-pGAP1 constitutive bidirectional promoter, and combinations thereof.
[0043] Each open reading frame utilized in a vector bound to yeast requires a terminator at its end. In some embodiments, the DNA sequence of the terminator can be inserted into the vector.
[0044] The origin of replication is essential for initiating replication. In some implementations, the DNA sequence of the origin of replication can be inserted into a vector.
[0045] When yeast is a microorganism, it is essential to have a DNA sequence that is homologous to the yeast genome and can be incorporated into a vector.
[0046] Selection markers are used to select organisms that have been successfully transformed. These markers are sometimes associated with antibiotic resistance. They can also be associated with the ability to grow with or without certain amino acids (auxotrophic markers). Suitable auxotrophic markers include, but are not limited to, ADE, HIS, URA, LEU, LYS, TRP, and combinations thereof. In some embodiments, the DNA sequence of the selection marker can be incorporated into a vector. This disclosure includes methods for growing cells expressing fusion proteins, expressing fusion proteins, isolating and purifying fusion proteins. This disclosure also includes uses of the fusion proteins as described herein.
[0047] Specifically, the fusion protein described herein can be used in personal care compositions. In the case of personal care compositions, the fusion protein can be applied to the skin. For this purpose, the fusion protein can be wholly or partially isolated or purified (e.g., at least 25% purification, at least 50% purification, at least 65% purification, at least 75% purification, at least 85% purification, at least 90% purification, at least 95% purification, at least 96% purification, at least 97% purification, at least 98% purification, at least 99% purification, or 100% purification). In other words, the fusion protein can be added to personal care products as a purified protein, or as part of a fraction from which the protein is found. The fusion protein can be formulated into creams, lotions, ointments, gels, serums, etc.
[0048] Personal care compositions can provide formulations suitable for topical application to the skin. The composition may also contain a cosmetically acceptable carrier. The cosmetically acceptable carrier may comprise from about 50% to about 99% by weight of the composition (e.g., from about 80% to about 95% by weight of the composition). In some embodiments, the carrier may comprise about 50% by weight, about 55% by weight, about 60% by weight, about 65% by weight, about 70% by weight, about 75% by weight, about 80% by weight, about 85% by weight, about 90% by weight, about 95% by weight, about 96% by weight, about 97% by weight, about 98% by weight, or about 99% by weight of the composition. These compositions can be formulated into a wide variety of product types, including but not limited to liquid compositions such as lotions, creams, gels, sticks, sprays, shaving creams, ointments, makeup removers, and solid sticks, pastes, powders, mousses, masks, peels, cosmetics, and wipes. These product types may contain several types of cosmetically acceptable carriers, including but not limited to solutions, emulsions (e.g., microemulsions and nanoemulsions), gels, solids, and liposomes. The following are non-limiting examples of such carriers. Other carriers can be formulated by those skilled in the art.
[0049] Topical compositions that can be used in this disclosure can be formulated as solutions. Solutions typically contain an aqueous solvent (e.g., about 50% to about 99% or about 90% to about 95% of a cosmetically acceptable aqueous solvent). In some embodiments, the solution may have about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, or about 99% of a cosmetically acceptable aqueous solvent. Topical compositions can be formulated as solutions containing emollients. Such compositions preferably contain about 2% to about 50% emollient. In some embodiments, the composition may comprise about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 12%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, or about 50% of an emollient. As used herein, an emollient means a material used to prevent or relieve dryness and to protect the skin. A wide variety of suitable emollients are known and can be used in personal care compositions. See the International Cosmetic Ingredient Dictionary and Handbook, edited by Wenninger and McEwen, (The Cosmetic, Toiletry, and Fragrance Assoc., Washington, DC, 7th edition, 1997) (hereinafter referred to as the “CTFAs Handbook”), which contains many examples of suitable materials.
[0050] Emulsions can be made from such solutions. Emulsions typically contain about 1% to about 20% (e.g., about 5% to about 10%) of emollients and about 50% to about 90% (e.g., about 60% to about 80%) of water. In some embodiments, the emulsion may have about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, or about 20% of emollients. In some embodiments, the emulsion may have about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, or about 80% water.
[0051] Another type of product that can be formulated from a solution is a cream. Creams typically contain about 5% to about 50% (e.g., about 10% to about 20%) of an emollient and about 45% to about 85% (e.g., about 50% to about 75%) of water. In some embodiments, the cream may have about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, or about 50% of an emollient. In some embodiments, the cream may have about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, or about 85% water.
[0052] Another class of products that can be formulated from solutions can be ointments. Ointments may contain a simple base of animal or vegetable oils or semi-solid hydrocarbons. Ointments may contain about 2% to about 10% emollients, plus about 0.1% to about 2% thickeners. In some embodiments, the ointment may have about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, or about 10% emollients. In some embodiments, the ointment may have about 0.1%, about 0.2%, about 0.3%, about 0.4%, about 0.6%, about 0.8%, about 1.0%, about 1.2%, about 1.4%, about 1.6%, about 1.8%, or about 2.0% thickeners. A more complete disclosure of thickeners or viscous agents that may be used herein can be found in the CTFA manual.
[0053] These personal care compositions can be formulated as emulsions. If the carrier is an emulsion, then about 1% to about 10% (e.g., about 2% to about 5%) of the carrier contains an emulsifier. In some embodiments, about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, or about 10% of the carrier contains an emulsifier. The emulsifier can be a nonionic emulsifier, anionic emulsifier, or cationic emulsifier. Suitable emulsifiers are disclosed, for example, in the CTFA manual.
[0054] Milks and creams can be formulated as emulsions. Typically, such milks contain 0.5% to about 5% emulsifier. Such creams will typically contain about 1% to about 20% (e.g., about 5% to about 10%) emollient; about 20% to about 80% (e.g., 30% to about 70%) water; and about 1% to about 10% (e.g., about 2% to about 5%) emulsifier.
[0055] Single-phase emulsion skincare compositions (such as lotions and creams) of water-in-oil and oil-in-water types are well-known in the beauty industry and are used in personal care compositions. Multiphase emulsion compositions (such as water-in-oil-in-water types) are also available. Generally, these single-phase or multiphase emulsions contain water, emollients, and emulsifiers as basic ingredients.
[0056] The personal care compositions disclosed herein can also be formulated as gels (e.g., hydrogels using suitable gelling agents). Suitable gelling agents for hydrogels include, but are not limited to, natural gums, polymers and copolymers of acrylic acids, polymers and copolymers of acrylates, and cellulose derivatives (e.g., hydroxymethylcellulose and hydroxypropylcellulose). Suitable gelling agents for oils (such as mineral oils) include, but are not limited to, hydrogenated butene / ethylene / styrene copolymers and hydrogenated ethylene / propylene / styrene copolymers. Such gels typically contain between about 0.1% and 5% by weight of such gelling agents. In some embodiments, the gel contains about 0.1% by weight, about 0.2% by weight, about 0.3% by weight, about 0.4% by weight, about 0.5% by weight, about 1.0% by weight, about 1.5% by weight, about 2.0% by weight, about 2.5% by weight, about 3.0% by weight, about 3.5% by weight, about 4.0% by weight, about 4.5% by weight, or about 5.0% by weight of such gelling agents.
[0057] In addition to the aforementioned components, the personal care compositions that may be used in this disclosure may also contain a variety of additional oil-soluble and / or water-soluble materials, which are conventionally used in compositions applied to the skin in amounts established in the field.
[0058] Personal care compositions may be applied to or applied to the skin as needed and / or as part of a routine regimen, ranging from once a week to once or more daily (e.g., twice daily). Dosage will vary depending on factors such as the end-user's age and physical condition, treatment duration, the specific compound, product, or composition used, and the specific cosmetically acceptable carrier employed.
[0059] The fusion proteins described in this article can be used in personal care applications to achieve beneficial skincare effects, such as anti-wrinkle, improving skin pigmentation, hydration, reducing acne, preventing acne, reducing blackheads, preventing blackheads, reducing stretch marks, preventing stretch marks, preventing cellulite, and reducing cellulite. Improving skin pigmentation refers to evening out or reducing skin pigmentation to provide a fairer complexion.
[0060] The fusion proteins described in this article can also be combined with other skincare-beneficial ingredients, such as, but not limited to, salicylic acid, retinol, benzoyl peroxide, vitamin C, glycerin, alpha-hydroxy acids, hydroquinone, kojic acid, and hyaluronic acid.
[0061] Collagen prolyl 4-hydroxylase contains conserved domains similar to those of prolyl hydroxylase domain proteins (PHDs) (including PHD1, PHD2, PHD3, PHD4, etc.). These PHDs play a key role in regulating the hydroxylation of hypoxia-inducible factor (HIF). HIF is a DNA-binding transcription factor that interacts with specific nuclear cofactors under hypoxic conditions. HIF transactivates a series of hypoxia-related genes to trigger adaptive responses. Due to its role in cells, HIF is associated with many cellular functions, such as homeostasis, angiogenesis, and anaerobic metabolism. Upregulation and downregulation of HIF in cells can induce angiogenesis or proliferation of cancer cells; therefore, HIF and prolyl hydroxylase are increasingly studied due to their therapeutic potential. Therefore, the fusion protein described in this article can be applied to prolyl hydroxylase domain proteins.
[0062] In the context of this specification, unless otherwise specified, all publications, patent applications, patents and other references mentioned herein are expressly incorporated herein by reference in their entirety for all purposes, as if fully explained, and their entirety shall be considered part of this disclosure.
[0063] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. In case of any discrepancy, this specification and its included definitions shall prevail.
[0064] When quantities, concentrations, or other values or parameters are given as ranges or by listing upper and lower limits, they should be understood to specifically disclose all ranges formed by any pair of any upper and lower limits, regardless of whether the range is disclosed individually. When a range of values is referenced herein, unless otherwise specified, the range is intended to include its endpoints, as well as all integers and fractions within that range. When a range is defined, it is not intended to limit the scope of this disclosure to the specific values listed.
[0065] Furthermore, unless otherwise expressly stated to the contrary, when one or more ranges or lists of items are provided, this shall be understood to expressly disclose any single specified value or item in such ranges or lists, and any combination thereof with any other single value or item in the same or any other list.
[0066] As used herein, the terms “comprising,” “including,” “having,” or any other variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, article of manufacture, or apparatus that includes a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to such process, method, article of manufacture, or apparatus.
[0067] Furthermore, unless explicitly stated to the contrary, “or” and “and / or” are inclusive rather than exclusive. For example, any of the following conditions satisfy either A or B, or A and / or B: A is true (or exists) and B is false (or does not exist), A is false (or does not exist) and B is true (or exists), and both A and B are true (or exist).
[0068] The use of “a” or “an” to describe the various elements and components of this document is merely for convenience and to give the general meaning of this disclosure. This description should be interpreted as including one or more than one, and the singular includes the plural, unless explicitly stated otherwise.
[0069] The foregoing written description provides manners and procedures for preparing and using the description, enabling any person skilled in the art to prepare and use the description, and this practicability is specifically provided with respect to the subject matter of the appended claims, which form part of the original description.
[0070] As used in this article, phrases such as “selected from the group consisting of”, “selected from”, etc., include mixtures of specified materials.
[0071] When a feature or element is referred to herein as being "on" another feature or element, it may be directly on top of the other feature or element, or there may be intervening features and / or elements. In contrast, when a feature or element is referred to as being "directly on" another feature or element, there are no intervening features or elements. It should also be understood that when a feature or element is referred to as being "connected," "attached," or "joined" to another feature or element, it may be directly connected, attached, or joined to the other feature or element, or there may be intervening features or elements. In contrast, when a feature or element is referred to as being "directly connected," "directly attached," or "directly joined" to another feature or element, there are no intervening features or elements. Although described or illustrated relative to one embodiment, the features and elements thus described or illustrated are applicable to other embodiments. Those skilled in the art will also recognize that a structure or feature referred to as being "adjacent" to another feature may have portions that overlap with or are located below the adjacent feature.
[0072] For ease of description, spatial relative terms such as “below,” “under,” “lower,” “above,” “upper,” etc., may be used herein to describe the relationship between one element or feature and another, as illustrated in the accompanying drawings. It should be understood that, in addition to the orientations depicted in the drawings, spatial relative terms are also intended to cover different orientations of the device in use or operation. For example, if the device in the drawings is inverted, an element described as “below” or “below” the other element or feature would be oriented “above” the other element or feature. Thus, the exemplary term “below” can cover both orientations of “above” and “below.” The device may be oriented in other ways (rotated 90 degrees or otherwise), and the spatial relative descriptive terms used herein will be interpreted accordingly. Similarly, the terms “up,” “down,” “vertical,” “horizontal,” etc., are used herein for illustrative purposes only, unless otherwise expressly stated.
[0073] Although the terms “first” and “second” may be used herein to describe various features / elements, these features / elements should not be limited by these terms unless the context otherwise indicates. These terms may be used to distinguish one feature / element from another. Thus, the first feature / element discussed below may be referred to as the second feature / element, and similarly, the second feature / element discussed below may be referred to as the first feature / element, without departing from the teachings of this disclosure.
[0074] When the term "about" is used, it is used to indicate that a certain effect or result can be obtained within a certain tolerance, and that a person skilled in the art knows how to obtain the tolerance. When the term "about" is used to describe the endpoints of a value or range, this disclosure should be understood to include the specific value or endpoint mentioned. In embodiments, "about" may refer to a range of up to 10% (i.e., ±10%).
[0075] Any range of values listed in this article is intended to include all subranges contained therein.
[0076] The embodiments and examples included herein illustrate, by way of illustration and not limitation, specific implementations in which the subject matter can be practiced. As mentioned, other embodiments may be used and derived therefrom, allowing for structural and logical substitutions and changes without departing from the scope of this disclosure. Such embodiments of the subject matter of the invention may be mentioned individually or collectively herein only for convenience and are not intended to automatically limit the scope of this application to any single inventive concept (if more than one is disclosed in fact). Therefore, although specific embodiments have been illustrated and described herein, any arrangement intended to achieve the same purpose may replace the specific embodiments shown. This disclosure is intended to cover any and all modifications or variations of the various embodiments. Combinations of the above embodiments and other embodiments not specifically described herein will become apparent to those skilled in the art upon review of the foregoing description.
[0077] The above description is presented to enable those skilled in the art to prepare and use all the fusion proteins disclosed herein, and is provided in the context of a particular application and its requirements. Various modifications to the preferred embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of this disclosure. Therefore, this disclosure is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0078] This disclosure has been generally described, and further understanding can be obtained by referring to certain specific embodiments, which are provided herein for illustrative purposes only and are not intended to be limiting unless otherwise indicated.
[0079] Example
[0080] Example 1 :
[0081] DNA sequences of bovine P4HA1 (SEQ ID NO:1) and bovine P4HB (SEQ ID NO:2) were obtained from DNA 2.0. Using these DNA sequences as templates, polymerase chain reaction (PCR) was performed using primers MM-1090 (SEQ ID NO:3), MM-750 (SEQ ID NO:4), MM-0782 (SEQ ID NO:5), MM-0783 (SEQ ID NO:6), MM-0784 (SEQ ID NO:7), and MM-0785 (SEQ ID NO:8). The resulting DNA was then assembled by Gibson into the vector MMV290 (SEQ ID NO:9) (Gibson DG, Young L, Chuang RY, Venter JC, Hutchison CA, Smith HO. Enzymatic assembly of DNA molecules up to several hundred kilobases. NatMethods. 2009; 6:343–5.). The final vector MMV290 ( Figure 4 The strain was confirmed by sequencing and transformed into Pichia pastoris strain PP153 to generate strain PP336, which has P4HA1 at the N-terminus and P4HB at the C-terminus.
[0082] MMV-130 was digested with Pme I ( Figure 1 It is then converted into PP1 to generate PP153. PP153 contains wild-type collagen driven by the pDF promoter.
[0083] The DNA sequence of monomeric prolyl 4-hydroxylase α was obtained from IDT (SEQ ID NO:15). Using the DNA sequence as a template, polymerase chain reaction (PCR) was performed using primers MM-0579 (SEQ ID NO:18), MM-0580 (SEQ ID NO:19), MM-1569 (SEQ ID NO:20), MM-1570 (SEQ ID NO:21), and MM-0784 (SEQ ID NO:7), followed by Gibson assembly into the vector MMV-644 (SEQ ID NO:17). The final vector MMV-644 (SEQ ID NO:17) Figure 10 The strain was confirmed by sequencing and transformed into Pichia pastoris strain PP97 to generate strain PP765.
[0084] MMV-644 was digested with Swa I ( Figure 10It is then converted into PP97 to generate PP765. PP765 contains a monomeric prolyl 4-hydroxylase with a 6X His tag at the C-terminus driven by the pDF promoter and a secretory signal from the α-mating factor of Saccharomyces cerevisiae.
[0085] Example 2 :
[0086] DNA sequences of bovine P4HA1 and bovine P4HB were obtained from DNA 2.0. Using these DNA sequences as templates, polymerase chain reaction (PCR) was performed using primers MM-1090, MM-750, MM-779, MM-780, MM-781, and MM-369, followed by Gibson assembly into the vector MMV289 (SEQ ID NO: 10). The final vector MMV289 (… Figure 5 The strain was confirmed by sequencing and transformed into yeast strain PP153 to generate strain PP335, which has P4HB at the N-terminus and P4HA1 at the C-terminus.
[0087] Example 3 :
[0088] The PP336 strain was inoculated into 24-well plates containing 2 mL of BMGY medium and grown at 30 °C with shaking at 900 rpm for 48 h. Cells were rapidly centrifuged and then lysed using a Qiagen tissue lysator in 800 μL of lysis buffer. The lysis buffer was prepared with the following components: 2.5 mL 1 M HEPES; 50 mM, 438.3 mg NaCl; 150 mM, 5 mL glycerol; 10% final concentration, 0.5 mL Triton X-100; 1% final concentration, and 42 mL Lilipure water. The supernatant containing the fusion protein (AB fusion protein) with P4HA1 at the N-terminus and P4HB at the C-terminus was loaded onto an SDS-PAGE gel and transferred to a PVDF membrane. The fusion protein was detected in a Western blot using a P4HB antibody.
[0089] The PP765 strain was inoculated into 24-well plates containing 2 mL of BMGY medium and grown at 30°C with shaking at 900 rpm for 48 hours. Cells were rapidly centrifuged and the medium was collected. The supernatant containing monomeric prolyl 4-hydroxylase was loaded onto an SDS-PAGE gel and transferred to a PVDF membrane. The fusion protein was detected in a Western blot using a His-tagged antibody.
[0090] The same procedure described above was performed using strain PP335 to generate a fusion protein (BA fusion protein) with P4HB at the N-terminus and P4HA1 at the C-terminus.
[0091] For the AB22 fusion protein, we detected a fusion protein with a molecular weight of approximately 120 kDa using both Coomassie staining and Western blotting. For the BA fusion protein, we could not detect the fusion protein using both methods simultaneously.
[0092] Strain PP336 was inoculated into 24-well plates containing 2 mL of BMGY fermentation medium and grown at 30°C with shaking at 900 rpm for 48 hours. Simultaneously, the baseline yeast strain PP268, containing DNA sequences of collagen, P4HA, and P4HB respectively, was grown under the same conditions.
[0093] MMV156 was digested with Bam HI ( Figure 2 ) and then converted to PP153 to generate PP154, which in turn generates PP268, and then MMV-191 was digested with Bam HI. Figure 3 And then convert it into PP154 to generate PP268.
[0094] The following procedure was followed to analyze samples PP336 and PP268 by pepsin assay to assess the sensitivity of collagen trimers to pepsin. PP336 will exhibit similar pepsin tolerance to PP268.
[0095] The proline hydroxylation of PP336 and PP268 was analyzed by amino acid analysis. PP336 will exhibit similar or better proline hydroxylation as observed for PP268.
[0096] The pepsin assay was performed using the following procedure:
[0097] 1. Prior to pepsin treatment, dioctanedinic acid (BCA) was determined according to the Thermo Scientific protocol to obtain total protein for each sample. For all samples, total protein was normalized to the lowest concentration.
[0098] 2. Place 100 μL of the lysate into a microcentrifuge tube.
[0099] 3. A main mixture containing the following substances is produced:
[0100] a. 37% HCl (containing 0.6 mL of acid per 100 mL) and
[0101] b. Pepsin (the stock solution in deionized water is 1 mg / mL, and the final addition of pepsin should be at a ratio of 1:25 pepsin:total protein (weight:weight).
[0102] c. Based on step #1, namely the standardization of total protein, the amount of pepsin will vary with the final addition, and adjustments should be made using the created spreadsheet.
[0103] 4. After adding pepsin, mix three times with a pipette, and then incubate the sample at room temperature for one hour to allow for the pepsin reaction.
[0104] 5. After one hour, add a 1:1 volume of LDS loading buffer containing β-mercaptoethanol to each sample and incubate at 70°C for 7 minutes.
[0105] 6. Then rotate at 14,000 rpm for 1 minute to remove turbidity.
[0106] Example 4 :
[0107] Yeast strain PP97, lacking both collagen and the fusion protein DNA, was grown overnight in YPD medium with 80 mM proline to produce a culture. 5 mL of the culture was inoculated with 20 mL of YPD medium and 80 mM proline and incubated at 30°C for 1 hour at 300 rpm. Cells were harvested by centrifugation at 5000 rpm for 5 minutes at 4°C, washed twice with sterile water, and then mixed with 10 mL of transformation buffer and incubated at 25°C with 10 mM DDT for 25 minutes. Cells were harvested, washed twice with cold sorbitol, and then transformed into MMV400 (SEQ ID NO: 11 and SEQ ID NO: 12) containing the AB fusion protein DNA via electroporation. Figure 6 Cells were plated after incubating for three hours on bleomycin 500 plates containing 80 mM proline throughout the entire duration. The plates were incubated at 30°C for two days, and colonies were then screened according to the procedure described in Example 3. The results showed that the fusion protein was transformed into empty host cells in the presence of proline.
[0108] In the absence of proline in YPD medium, no colonies or only a few colonies formed. When these colonies were analyzed by Western blotting, all colonies were negative for the AB fusion protein. In experiments where 80 mM proline was added to YPD medium, Western blotting analysis of 6 / 6 colonies showed that all colonies were positive for the AB fusion protein.
[0109] Example 5 :
[0110] The vector MMV290 was packaged using BglII and MluI. Figure 4 (SEQ ID NO:9) was digested and then assembled with an insert sequence (SEQ ID NO:12) encompassing the C-terminus of the AB fusion protein, a six-amino acid sequence representing histidine (His tag), two stop codons, and the AOX1 transcription terminator, thereby generating the vector MMV502. Figure 7).
[0111] The vector MMV156 was injected using BglII and MluI. Figure 2 (SEQ ID NO:13) was digested and then assembled with the insert sequence (SEQ ID NO:12) using Gibson assembly. The insert sequence covers the C-terminus of the P4HB subunit protein, a nucleotide sequence of six consecutive amino acids representing histidine (His tag), two stop codons, and the AOX1 transcription terminator, thereby generating the vector MMV503. Figure 8 ).
[0112] MMV502 was transformed into PP153 to generate strain PP548. This strain was cultured, lysed, and protein content was determined using various methods, including Western blotting and Coomassie staining gel electrophoresis. Western blotting confirmed the presence of the AB fusion protein. Coomassie staining gel electrophoresis confirmed the molecular weight (119 kDa) of the His-tagged AB fusion protein. High-expression variants of strain PP548 were grown in shake flasks and fermenters. Once confluent, cells were centrifuged to form a pellet and washed. Cells were then lysed using a Qiagen tissue lysator in 800 μL of lysis buffer. The lysis buffer was prepared with the following components: 2.5 mL 1 M HEPES; 50 mM final concentration, 438.3 mg NaCl; 150 mM final concentration, 5 mL glycerol; 10% final concentration, 0.5 mL Triton X-100; 1% final concentration, and 42 mL Millipure water. The lysate was centrifuged, and the soluble fraction was incubated with nickel-NTA agarose beads. The clarified lysate-bead mixture was applied to a column retaining the beads. The nickel-NTA beads were then washed with different concentrations of imidazole (potentially including other chemicals such as 1,10-phenanthroline and EDTA). The AB fusion protein encoded by plasmid MMV502 and tagged with His was then eluted by washing with 300 mM imidazole. These eluates were combined or kept separate, and then buffer-exchanged using an Amico Ultra-15 filter column to remove residual imidazole. The AB fusion protein was then used for subsequent assays.
[0113] MMV503 was transformed into PP153 to generate strain PP549. This strain was cultured, lysed, and protein content was determined using various methods, including Western blotting and Coomassie staining gel. Western blotting confirmed the presence of P4HA and P4HB enzymes. Coomassie staining gel confirmed the molecular weight of P4HA (61 kDa) and P4HB (57 kDa). High-expression variants of PP549 were grown in shake flasks and fermenters. Once confluent, cells were centrifuged to form a pellet and washed. Cells were then lysed using a Qiagen tissue lysator in 800 μL of lysis buffer. The lysis buffer was prepared with the following components: 2.5 mL 1 M HEPES; 50 mM final concentration, 438.3 mg NaCl; 150 mM final concentration, 5 mL glycerol; 10% final concentration, 0.5 mL Triton X-100; 1% final concentration, and 42 mL Millipure water. The lysate was centrifuged, and the soluble fraction was incubated with nickel-NTA agarose beads. The clarified lysate-bead mixture was applied to a column retaining the beads. The nickel-NTA beads were then washed with different concentrations of imidazole (potentially including other chemicals such as 1,10-phenanthroline and EDTA). P4HA and P4HB, encoded by plasmid MMV503 and tagged with His, were then eluted with 300 mM imidazole. The eluates were combined or kept separate, and then buffer-exchanged using an Amico Ultra-15 filter column to remove residual imidazole. The P4HA and P4HB proteins were then used for subsequent assays.
[0114] Example 6 :
[0115] By improving the method based on α-ketoglutarate hydroxylation coupling decarboxylation, the activity of the fusion protease from PP548 (Kivirikko, KI and Myllyla) was confirmed. ·· , R. (1982) Post-translational enzymes in the biosynthesis of collagen: intracellular enzymes. Methods Enzymol., 82, 245–304; Kivirikko, KI and Myllyla ··R. (1997) Characterization of the iron-and 2-oxoglutarate-binding sites of human prolyl 4-hydroxylase. The EMBO Journal, 16, 1173-1180. This activity measurement is based on the presence of the AB fusion protein at (Pro-Pro-Gly) sites. 10 The consumption of α-ketoglutarate over time was induced by hydroxylation of the peptide model substrate. The amount of AB fusion protein ranged from 0.12 nmol to 0.4 nmol per reaction. In 96-well plates, the reaction was terminated by mixing 50 μl of selected samples at different time points from 0 to 10 minutes into 150 μl of 30 mM o-phenylenediamine in 0.5 M HCl solution. The plate was placed on a heating block set to 95 °C for 10 minutes to stop color formation, and then cooled on ice for 2 minutes. 50 μl of the sample was then mixed with 30 μl of 1.25 M NaOH in a black 96-well plate. The fluorescence of the sample was read on a microplate reader at emission 420°C and excitation 340°C. The α-ketoglutarate concentration was derived from α-ketoglutarate standards measured under the same assay conditions. The consumption of α-ketoglutarate was calculated by subtracting the sample concentration from the zero concentration.
[0116] The P4HA and P4HB enzyme activities from PP548 were confirmed by the same assays described above.
[0117] The results showed that the sample containing the AB fusion protein contained less α-ketoglutarate compared to samples containing native P4HA and P4HB proteins. This indicates that the AB fusion protein has greater activity than native P4HA and P4HB proteins.
[0118] Example 7 :
[0119] MMV411 (SEQ ID NO:14 and) were digested with Pme I. Figure 9 And it is converted into PP97 to generate PP434.
[0120] A single colony was inoculated into 50 mL of BMGY medium and incubated overnight at 30°C with constant shaking at 250 rpm. The next day, the overnight culture in a 1 L Erlenmeyer flask was inoculated into 500 mL of fresh BMGY medium and incubated for 2 days at 30°C with constant shaking at 250 rpm.
[0121] PP434 cells were resuspended (1 g wet cell weight (wcw)) in 5.667 ml phosphate-buffered saline (50 mM, pH 7.4). Cell lysis was performed using Matrix D beads in a bead mill for 5 cycles (cooling for 1 min between each cycle) to generate whole cell lysates. The whole cell lysates were then placed in several 1.5 ml microcentrifuge tubes and heated at 70 °C for 30 min, gently mixing every 5 min. The whole cell lysates were then rapidly centrifuged at 21000*g for 5 min at 4 °C. The supernatant was placed on ice for 10 min. Ni-NTA resin (for 1 g wcw, 0.5 ml bed volume) was equilibrated three times with deionized water, and ethanol was removed by centrifugation at 800*g for 2 min at 4 °C. The clarified lysate was added to the equilibrated Ni-NTA resin and incubated at 4 °C with inverted rotation for 60 min. The supernatant was collected by centrifugation at 800*g for 5 min at 4 °C. The resin was washed with 10 column volumes of 50 mM phosphate-buffered saline (pH 7.4) and 20 mM imidazole by centrifugation at 800 μg for 2 minutes at 4°C. The resin was then washed again with 10 column volumes of 50 mM phosphate-buffered saline (pH 7.4) and 250 mM imidazole by centrifugation at 800 μg for 2 minutes at 4°C. After incubating the protein with the elution buffer at 4°C for 5 minutes (inverted rotation), the protein was eluted three times with 5 mL of 50 mM phosphate-buffered saline (pH 7.4) and 500 mM imidazole by centrifugation at 800 μg for 2 minutes at 4°C. The sample (both supernatant and precipitate, along with whole-cell lysate) was analyzed on SDSPAGE. The sample was then dialyzed at least once in 50 mM Tris (pH 8.0) and 100 mM NaCl (dialyzed in at least 100 sample volumes).
[0122] - Preparation of PP547
[0123] Vector MMV363 was modified to include a 22kD small Pre-Pro-Col3 domain, along with associated promoter pDF and terminator AOX1TT, a Flag tag and a HA tag, a DNA sequence for labeling expression, and associated promoter and terminator, a DNA sequence for the origin of replication in bacteria and yeast, and a DNA sequence homologous to the yeast genome for integration. Vector MMV88 is the source DNA for the Pre-Pro-Col3 domain. Vector MMV130 is the source DNA for the Col3A1 domain with added HA and Flag tags. The total length of the Col3A1 polypeptide is 190 amino acids (aa). The three fragments were assembled using Gibson to obtain plasmid MMV383.
[0124] Integration was performed using an Aox landing pad, transforming MMV383 into PP97. The resulting Pichia pastoris strain was PP414. Subsequent Western blotting revealed the secretion of small 22kD Col3 molecules.
[0125] Convert PP414 using MMV502 (the His-tagged version of MMV290) to generate PP547.
[0126] - Preparation of PP635 and PP636
[0127] Single colonies of PP97 were inoculated into 15 ml of YPD medium containing 80 mM proline and incubated overnight at 30°C with shaking (250 rpm). The next day, the volume of medium was doubled with fresh YPD containing 80 mM proline and incubated for another hour at 30°C with shaking (250 rpm). Cells were centrifuged at 3,500 g for 5 minutes; washed twice with sterile water, and resuspended in 10 ml of transformation buffer (10 mM Tris-Cl (pH 7.5), 100 mM LiAc, 0.6 M sorbitol), with 10 mM dithiothreitol (DTT) added and mixed thoroughly. The resuspended solution was incubated at room temperature for 30 minutes. Cells were centrifuged at 3,500 × g for 5 minutes, and the pellet was resuspended in 5 ml of ice-cold 1 M sorbitol, and centrifuged again at 3,500 × g for 5 minutes. Washing was repeated twice with 5 ml of 1 M sorbitol. The washed precipitate was resuspended in 500 μl of ice-cold 1M sorbitol, and 100 μl of this resuspended solution was aliquoted into pre-cooled 0.2 cm electroporation cuvettes. The linearized DNA sequence of MMV502 ( Figure 7 ) and linearized DNA sequence of MMV503 ( Figure 8 The DNA was added to the cells (in separate cuvettes) and mixed by pipetting. A negative control was also set up in which water, instead of the linearized DNA sequence, was added to the cell mixture. The mixture was incubated on ice for 10 minutes. After incubation, electroporation was performed via pulsed method using the Pichia pastoris-WU protocol (1500 v, 25 u F, 200 W) with Bio-Rad Gene Pulser Xcell. TM Electroporation was performed. Cells were immediately transferred to a 500 μl mixture of YPD and 1 M sorbitol (1:1) and incubated at 30 °C for 2 hours. 100 μl aliquots of this 2-hour incubated culture were plated onto G418 antibiotic plates containing 750 μg / ml. The plates were incubated at 30 °C for two days.
[0128] Colonies appearing on the plate after 2 days of incubation were picked and inoculated into BMGY medium containing 500 μg / ml G418. Inoculation was performed in 2 ml wells, with each well containing 24 cells. The plates were incubated at 30°C with shaking (900 rpm) for 2 days. Each 2 ml culture was rapidly centrifuged, and 100 mg of the precipitate was resuspended in 1 ml lysis buffer (50 mM sodium phosphate, 5% glycerol, and 1% EDTA, pH 7.5). Lysis was performed using a tissue lysator and Y matrix beads for 15 min. The lysate was mixed with SDS-Licor-loaded dye at a 5:1 ratio, heated at 90°C for 10 min, and then loaded onto a 4% to 12% Bis-Tris gel. The gel was transferred to a PVDF membrane. Western blot analysis was performed using anti-His and anti-collagen antibodies. Because P4H carries a His tag, the fusion P4H appears as a 110 kDa protein in the red channel, while the bidirectionally expressed P4HA / B appears at 59 kDa in the red channel. No collagen bands were observed in the blot, confirming that the P4H plasmid has been transformed and therefore lacks collagen. The clone showing high expression of the fusion P4H was identified as PP635, and the clone showing high expression of bidirectional P4H was identified as PP636.
[0129] Individual colonies of each strain were inoculated into 50 mL of BMGY medium and incubated overnight at 30°C with constant shaking at 250 rpm. The next day, the overnight culture in a 1 L Erlenmeyer flask was inoculated into 500 mL of fresh BMGY medium and incubated for 2 days at 30°C with constant shaking at 250 rpm.
[0130] Cells (0.45 g wet cell weight) were resuspended in 0.65 ml of lysis buffer (25 mM Tris (pH 7.5), 50 mM NaCl, 20 mM imidazole) to obtain a 45% suspension. Cells were lysed for 5 cycles using Matrix D beads in a bead mill (cooling for 1 min between each cycle) to generate lysates. The lysates were rapidly centrifuged to clarify the supernatant and precipitate (4 °C, 10 min, 16000*g). The clarified lysates were removed and placed on ice. The precipitate was resuspended in lysis buffer at 2 times the volume of the wet cell weight and centrifuged at 16000*g for 10 min to collect a more clarified lysate. The clarified lysates were combined. Ni-NTA resin (approximately 0.025 ml bed volume for 1 g wet cell weight, scaled up appropriately) was equilibrated three times in deionized water, and ethanol was removed by centrifugation at 800*g for 2 min at 4 °C. The clarified lysate was added to equilibrated Ni-NTA resin and incubated overnight at 4°C with inverted rotation. The supernatant was collected by centrifugation at 800 μg for 5 min at 4°C. The resin was washed with 10 column volumes of lysis buffer containing 50 mM imidazole by centrifugation at 800 μg for 2 min at 4°C. The resin was then washed with 10 column volumes of 50 mM phosphate buffer (pH 7.4) and 250 mM imidazole by centrifugation at 800 μg for 5 min at 4°C. After incubating the protein with elution buffer at 4°C for 5 min (inverted rotation), the protein was eluted with 5 mL of lysis buffer containing 300 mM imidazole by centrifugation at 800 μg for 5 min at 4°C. Two more elutions were performed (a total of 3). The sample (both supernatant and precipitate, along with whole-cell lysate) was analyzed on SDSPAGE. The sample was dialyzed at least once in 50 mM Tris (pH 8.0) and 100 mM NaCl (in at least 100 times the sample volume) to produce purified collagen lysate.
[0131] - In vitro hydroxylation reaction with purified collagen monomers
[0132] 1) Prepare reaction mixtures for 40 reactions (250 μL for each reaction) according to the table below.
[0133]
[0134]
[0135] 2) For 250 μL of reactant, divide 20 μL of the above mixture into each tube (three portions of each reactant).
[0136] 3) Add 1 g / L BSA, 0.1 g / L catalase, and water to bring the final volume to 250 μL.
[0137] 4) Add 5uM fusion protein
[0138] 5) Add 2 μM collagen sample
[0139] 6) Incubate the reactants at 32°C for 2 minutes.
[0140] 7) Add 2.5 μL of 0.4 M 2-glutaric acid and mix thoroughly.
[0141] 8) Incubate at 32℃ for 1 hour
[0142] 9) Transfer 100 μL of each reactant to a new tube and transfer the sample for hydroxyproline determination.
[0143] - Hydroxyproline assay
[0144] 1. Prepare the following solutions:
[0145] A. Citrate / acetate buffer (for 100 mL)
[0146]
[0147] Make up to 100 mL with Milli-Q water.
[0148] B. Chloroamine T (for 20 mL)
[0149] 1.41g chloramine T
[0150] 10mL isopropanol
[0151] 10mL Milli-Q water
[0152] C. Ehrlich's solution (for 20 mL)
[0153] 4g p-Dimethylbenzaldehyde (DMAB)
[0154] 6mL hydrochloric acid
[0155] 14mL isopropanol
[0156] D. Chloroamine T / citrate-acetate solution (for 20 mL)
[0157] 4 mL of chloramine T (from above)
[0158] 16 mL citrate / acetate buffer (from above)
[0159] 2. Sample preparation :
[0160] a. Place 100 μL of the in vitro hydroxylation reaction product containing collagen into an amber glass vial.
[0161] b. Add 500 μL of concentrated HCl and tighten the cap of the vial.
[0162] c. Incubate the vial in a heating block at 125°C for at least 18 hours.
[0163] d. Dry the sample using rapid vacuum.
[0164] e. Resuspend the dried sample in a vial using 225 μL of Milli-Q water.
[0165] f. Centrifuge the sample at 10,000 x g for 5 minutes to remove precipitates and debris, and collect the supernatant for determination.
[0166] 3. Standard curve preparation :
[0167] a. Preparation of 1000ug / mL hydroxyproline stock solution
[0168] b. Use this stock solution to prepare the highest standard concentration of 50 μg / mL
[0169] c. Using a 50 μg / mL solution, prepare the standard curve with the following concentrations: 25 μg / mL, 18.75 μg / mL, 12.5 μg / mL, 6.25 μg / mL, and 3.125 μg / mL.
[0170] d.0ug / mL = water
[0171] e. Place these standards in wells A1 to A7 of a 96-well plate, and place parallel samples in wells B1 to B7.
[0172] 4. Internal reference :
[0173] a. Follow steps 2a to 2d, but use 400 μL of type III collagen (Abcam, ab7528) instead of the collagen-containing in vitro hydroxylation reactant.
[0174] b. Resuspended in 400 μL Milli-Q water
[0175] c. Place the internal control in wells A8 and B8 of a 96-well plate.
[0176] 5. Internal reference quantification :
[0177] a. Take 50 μL aliquots from the type III collagen stock solution vial for testing on qSDS.
[0178] b. Use the concentration obtained from qSDS to calculate the hydroxylation percentage of the internal control.
[0179] 6. Hydroxyproline assay :
[0180] a. Add 50 μL of the standard and take four samples (two parallel samples will be blank samples without chloramine T).
[0181] b. For each reactant to be analyzed (including the standard curve wells), add 100 μL of chloramine T / citrate-acetate solution.
[0182] c. For Blank sample ,Add to 100 ul water / citrate-acetate solution (Oxidation should not occur in these samples)
[0183] d. Seal the plate and incubate it at 30°C with shaking for 25 minutes.
[0184] e. Add 100 μL of Ehrlich solution and mix thoroughly in each well until the well is clear.
[0185] f. Seal the plate and incubate it at 65°C with shaking for 25 minutes.
[0186] g. Remove the plate from the heat source and measure the absorbance of all samples / blank samples at 560 nm.
[0187] h. The percentage of hydroxylation is calculated by obtaining the molecular weight of the collagen used. The number of hydroxyproline sites and proline residues in the helical regions of the collagen used are also required.
[0188] i. Exemplary calculation of the percentage of hydroxyproline (%):
[0189] The molecular weight of PP685 collagen is 94,752 g / mol.
[0190] The molecular weight of hydroxyproline is 131.13 g / mol.
[0191] The number of hydroxyproline sites in the helical region is 145.
[0192] The number of proline sites in the helical region is 246.
[0193] The concentration of PP685 collagen in the IVOH reaction was 0.084 g / L.
[0194] a. Hydroxyproline concentration obtained from the standard curve of the IVOH reaction
[0195] 3.91ug / mL
[0196] Corrected using a multiplication factor: 3.1 × 3.91 ug / mL = 12.1 ug / mL
[0197] b. Hydroxyproline concentration expressed in micrograms (µg)
[0198] Use 50 µL of sample per well
[0199] (50uL × 12.1ug / mL) divided by 1000 = 0.607ug hydroxyproline
[0200] c. The number of micrograms of collagen used in the IVOH reaction
[0201] Use 50 µL of sample per well
[0202] (50uL × 0.084g / L) multiplied by 1 × 10 6 =4.2ug
[0203] d.PP685 collagen nmol
[0204] ·(4.2ug / 1×10 6 ug)×1g=4.2×10 -6 g
[0205] ·(4.2×10 -6 (g) / (94752.76g / mol of PP685 collagen)=4.4×10 -11
[0206] mol
[0207] 4.4×10 -11 mol×(1×10 9 nmol / 1mol) = 0.044 nmol
[0208] e. hydroxyproline nmol
[0209] (0.607ug / 131.13g / mol)×1000=4.6nmol hydroxyproline
[0210] f. nmol of proline
[0211] 0.044 nmol collagen × 246 = 10.8 nmol proline
[0212] g. Percentage of hydroxyproline
[0213] • (4.6 nmol / 10.8 nmol) × 100 = 42% hydroxylation
[0214] - Results
[0215] Non-hydroxylated collagen strain Hydroxylase Hydroxyproline % PP434 PP547 (fusion protein) 34.8% PP434 PP635 (fusion protein) 26.6% PP434 PP636 (bidirectional P4HA / B) 0.9%
[0216] The results showed that, in the presence of necessary cofactors and appropriate reaction conditions (temperature and pH), the fusion proteins in both strains (PP547 and PP635) were able to hydroxylate collagen substrates to a higher percentage than those in strain PP636, which contained the non-fusion protein. The difference between PP547 and PP635 lies in the presence of a small fragment of collagen in the former, which was initially considered essential for the stability of the strain and the protein. This indicates that the fusion protein is stable and can function as a better dioxygenase in vitro compared to the non-fusion protein, thus providing an advantage over the non-fusion counterpart. The fusion of P4HA and P4HB produced a stoichiometric amount of protein, resulting in a functional tetramer that contributes to the protein's structure and stability. The % hydroxylation results were confirmed by mass spectrometry.
[0217] Example 8
[0218] - In vitro hydroxylation of cell lysates at pH 12 was performed using NaPO4 buffer, followed by mixing with 0.1 mM FeSO4, 2 mM ascorbic acid, 25 mM DTT, and 25 mM α-ketoglutarate. The mixture was adjusted to pH 7.5 and incubated with shaking at 32°C for 3 hours to allow the reaction to proceed. After the reaction was complete, the pH was lowered to 4, and the reaction was incubated overnight (approximately 18 hours) at 25°C, followed by centrifugation at approximately 7,000 x g to collect the supernatant. The supernatant was dialyzed against water or buffer and used for hydroxyproline assay.
[0219] Example 9
[0220] For in vivo hydroxylation, collagen is synthesized in the rough endoplasmic reticulum (ER) with the aid of several molecular chaperones and enzymes. The collagen folding mechanism is assisted by a protein disulfide isomerase (PDI), part of the P4HA-B fusion protein present in the strain used in this study. PDI facilitates the proper formation of disulfide bonds at the non-collagenous N-terminus and C-terminus of the protein, followed by hydroxylation of proline residues by the P4HA moiety of the fusion protein. Cofactors involved in the hydroxylation reaction are present in the ER, making it a crucial organelle for in vivo hydroxylation. Once collagen is synthesized, it is stabilized by molecular chaperones present in the ER and hydroxylated by the P4HA-B fusion protein, where the B subunit further stabilizes and / or facilitates trimerization, while the A subunit uses its dioxygenase activity to hydroxylate proline residues.
[0221] Based on the foregoing teachings, numerous modifications and variations of this disclosure are possible. Therefore, it should be understood that, within the scope of the appended claims, this disclosure may be practiced in ways other than those specifically described herein. sequence list <110> Modern Meadow, Inc. <120> Fusion proteins and products for hydroxylated amino acids <130> 514761WO <160> twenty one <170> PatentIn version 3.5 <210> 1 <211> 1612 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotides <400> 1 atgatttggt atatcctagt cgttggtatt ttgttgccac agtcactggc tcacccaggc 60 ttcttcactt ctataggaca gatgactgat ttgattcaca cagaaaaaga cctagttaca 120 agccttaaag actatatcaa agctgaagag gataagttgg agcaaatcaa aaagtgggca 180 gagaaactcg atagattgac tagtactgca acaaaagatc ctgagggttt tgtgggtcac 240 ccagtgaatg ctttcaagct gatgaagaga cttaatacag agtggtcaga attggaaaac 300 ttggtactta aagatatgag tgatggattc atttctaact taacaattca aagacaatac 360 tttccaaacg atgaggacca agtaggagca gcaaaagctt tgttgcgatt gcaggacaca 420 tacaatttgg acaccgacac gatatcgaag ggtgatttac ctggtgtgaa gcataagtcc 480 ttcctcactg tggaagattg ttttgaattg ggaaaagtcg catatacaga agccgactac 540 tatcacacag aattatggat ggagcaagct ctgcgtcagt tggacgaagg tgaagtttct 600 accgttgata aggtttcagt tttggattac ttatcatacg ctgtttacca gcaaggtgat 660 ctggacaaag ctctactttt aactaaaaag ttgttggagc tggacccgga gcatcaaaga 720 gctaacggta atctgaaata ctttgaatac atcatggcta aggaaaagga cgcaaataag 780 tccgtccg atgaccaatc cgatcaaaag accactctga aaaaaaagg tgcagctgtt 840 gactacctcc cagagagaca aaagtatgaa atgctgtgta gaggagaggg tatcaagatg 900 actccaagga gacagaaaaa gctgttctgt agatatcatg atgggaaccg taacccaaaa 960 ttcattcttg ctccagcgaa acaggaagat gaatgggaca agcctagaat cattcgtttt 1020 catgacatca tctccgatgc agaatagag gttgtgaaag acttggccaa accaagattg 1080 agtagggcta ccgtccatga ccctgagact ggaaaattga ctaccgcaca atatcgtgtc 1140 tctaaatcag catggttgtc cggttacgag aatcccgtgg tcagccgtat caatatgcgt 1200 attcaagatt tgactggtct tgacgtaagc actgctgagg aactacaagt tgccaactat 1260 ggtgtgggcg gtcagtatga accccacttt gatttcgcca gaaaggacga gcctgatgct 1320 tttaaggagc taggtactgg aaatagaatc gcaacgtggt tgttctatat gtccgatgtg 1380 cttgctggag gagccacagt tttccctgag gtaggtgctt ctgtttggcc taaaaagggc 1440 acggccgtat tttggtacaa tctgtttgca tctggagaag gtgattacag cactagacat 1500 gctgcttgtc ccgtcttagt cggtaataag tgggtttcca ataagtggct gcatgagaga 1560 ggtcaagagt ttaggaggcc atgcacattg tcagaattag aatgataatt tt 1612 <210> 2 <211> 1750 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Nucleotide <400> 2 aaaatgagat tcccatctat tttcaccgct gtcttgttcg ctgcctcctc tgcattggct 60 gcccctgtta acactaccac tgaagacgag actgctcaaa ttccagctga agcagttatc 120 ggttactctg accttgaggg tgatttcgac gtcgctgttt tgcctttctc taactccact 180 aacaacggtt tgttgttcat taacaccact atcgcttcca ttgctgctaa ggaagagggt gtctctctcg agaaagaga ggccgaagct gcacccgatg aggaagatca tgttttagta 360. ttgcataaag gaaatttcga tgaagctttg gccgctcaca aatatctgct cgtcgagttt tacgctccct ggtgcggtca ttgtaaggcc cttgcaccag agtacgccaa ggcagctggt 420 aagttaaagg ccgaaggttc agagatcaga ttagcaaaag ttgatgctac agaagagtcc gatcttgctc aacaatacgg ggttcgagga tacccaacaa ttaagttttt caaaaatggt gatactgctt ccccaaagga atatactgct ggtagagagg cagacgacat agtcaactgg ctcaaaaaga gaacgggccc agctgcgtct acattaagcg acggagcagc agccgaagct cttgtggaat ctagtgaagt tgctgtaatc ggtttcttta aggacatgga atctgattca 720 gctaaacagt tccttttagc agctgaagca atcgatgaca tccctttcgg aatcacctca 780 840. aatagtgacg tgttcagcaa gtaccaactt gacaaagatg gagtggtctt gttcaaaaag tttgacgaag gcagaaacaa tttcgagggt gaggttacaa aggagaaact gcttgatttc attaaacata accaactacc cttagttatc gattcactg aacaaactgc tcctaagatt 960 ttcggtggag aaatcaaac acatcttg ttgttttgc CAagtccgt atcggattat 1020 gaaggtaac tctccaattt CAAAAAggcc gctgagagct ttaagggca gatttgttc 1080 atctttattg actcagacca cacagacaat cagaggattt tggagtttt cggtttgaaa 1140 aaggaggaat gtccagcagt ccgtttgatc accttggagg aggagatgac caatacaa 1200 ccagagtcgg atgagttgac tgccgagaag atacagaat ttgtcacag atttctggaa 1260 ggtaagatca agcctcatct tatgtctca gagttgcctg atgactggga taagcaacca 1320 gttaaagtat tggtgggtaa aactttgag gaagtggcct tcgacgagaa aaaaatgtc 1380 ttgttgaat tctatgctcc gtggtgtggt cactgtaagc agctgcacc aatttgggat 1440 aaactgggtg aaacttacaa agatcacgaa aacattgtta ttgcaagat ggacagtact 1500 gctaacgaag tggaggt gaagttcac tccttcccta cgctgaagtt ctttcctgca 1560 tctgctgaca gaactgttat cgactataat ggagagagga cattggatgg ttttaaaag 1620 tttcttgaat ccggaggtca agacggagct ggtgacgacg atgatttgga agatctggag 1680 gaggctgagg aacctgatct tgaggaggat gacgaccaga aggcagtcaa agatgaactg 1740 tgataagggg 1750 <210> 3 <211> 58 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotides <400> 3 ctcaattgtt gtttatatca ttgctattta aatcaggtga acccacctaa ctattttt 58 <210> 4 <211> 30 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotides <400> 4 ttttgttgtt gagtgaagcg agtgacggaa 30 <210> 5 <211> 60 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotides <400> 5 ttccgtcact cgcttcactc aacaacaaaa atgatttggt atatcctagt cgttggtatt 60 <210> 6 <211> 30 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotides <400> 6 ttctaattct gacaatgtgc atggcctcct 30 <210> 7 <211> 84 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotides <400> 7 aggaggccat gcacattgtc agaattagaa ggttctggct ctggttctgg ctctatgaga 60 ttcccatcta ttttcaccgc tgtc 84 <210> 8 <211> 69 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotides <400> 8 ctgcaacaaa agaaacaaga cattactgaa gggccggccg cacaaacgaa ggtctcactt 60 aatcttctg 69 <210> 9 <211> 10109 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotides <400> 9 ggatccttca gtaatgtctt gtttcttttg ttgcagtggt gagccatttt gacttcgtga 60 aagtttcttt agaatagttg tttccagagg ccaaacattc cacccgtagt aaagtgcaag 120 cgtaggaaga ccaagactgg cataaatcag gtataagtgt cgagcactgg caggtgatct 180 tctgaaagtt tctactagca gataagatcc agtagtcatg catatggcaa caatgtaccg 240 tgtggatcta agaacgcgtc ctactaacct tcgcattcgt tggtccagtt tgttgttac 300 gatcaacgtg acaaggttgt cgattccgcg taagcatgca tacccaagga cgcctgttgc 360 aattccaagt gagccagttc caacaatctt tgtaatatta gagcacttca ttgtgttgcg 420 cttgaaagta aaatgcgaac aaattaagag ataatctcga aaccgcgact tcaaacgcca 480 atatgatgtg cggcacacaa taagcgttca tatccgctgg gtgactttct cgctttaaaa 540 aattatccga aaaaatttc ctctagaatg ggtaaggaaa agactcacgt ttcgaggccg 600 cgattaaatt ccaacatgga tgctgattta tatgggtata aatgggctcg cgataatgtc 660 gggcaatcag gtgcgacaat ctatcgattg tatgggaagc ccgatgcgcc agagttgttt 720 ctgaaacatg gcaaaggtag cgttgccaat gatgttacag atgagatggt cagactaaac 780 tggctgacgg aattatgcc tcttccgacc atcaagcatt ttccgtac tcctgatgat 840 gcatggttac tcaccactgc gatccccggc aaaacagcat tccaggtatt agaagaatat 900 cctgattcag gtgaaaatat tgttgatgcg ctggcagtgt tcctgcgccg gttgcattcg 960 attcctgttt gtaattgtcc ttttaacagc gatcgcgtat ttcgtctcgc tcaggcgcaa 1020 tcacgaatga ataacggttt ggttgatgcg agtgattttg atgacgagcg taatggctgg 1080 cctgttgaac aagtctggaa agaaatgcat aagcttttgc cattctcacc ggattcagtc 1140 gtcactcatg gtgatttctc acttgataac cttatttttg acgaggggaa attaataggt 1200 tgtattgatg ttggacgagt cggaatcgca gaccgatacc aggatcttgc catcctatgg 1260 aactgcctcg gtgagttttc tccttcatta cagaaacggc tttttcaaaa atatggtatt 1320 gataatcctg atatgaataa attgcagttt catttgatgc tcgatgagtt tttctaaaat 1380 tgacacctta cgattattta gagagtattt attagtttta ttgtatgtat acggatgttt 1440 tattatctat ttatgccctt atattctgta actatccaaa agtcctatct tatcaagcca 1500 gcaatctatg tccgcgaacg tcaactaaaa ataagctttt tatgctgttc tctctttttt 1560 tcccttcggt ataattatac cttgcatcca cagattctcc tgccaaattt tgcataatcc 1620 tttacaacat ggctatatgg gagcacttag cgccctccaa aacccatatt gcctacgcat 1680 gtataggtgt tttttccaca atattttctc tgtgctctct ttttattaaa gagaagctct 1740 atatcggaga agcttctgtg gccgttatat tcggccttat cgtgggacca cattgcctga 1800 attggtttgc cccggaagat tggggaaact tggatctgat taccttagct gcatcagaat 1860 tggttaattg gttgtaacac tgacccctat ttgtttattt ttctaaatac attcaaatat 1920 gtatccgctc atgagacaat aaccctgata aatgcttcaa taatattgaa aaaggaagaa 1980 tatgagtatt caacatttcc gtgtcgccct tattcccttt tttgcggcat tttgccttcc 2040 tgtttttgct cacccagaaa cgctggtgaa agtaaaagat gctgaagatc agttgggtgc 2100 acgagtgggt tacatcgaac tggatctcaa cagcggtaag atccttgaga gttttcgccc 2160 cgaagaacgt tttccaatga tgagcacttt taaagttctg ctatgtggcg cggtattatc 2220 ccgtattgac gccgggcaag agcaactcgg tcgccgcata cactattctc agaatgactt 2280 ggttgagtac tcaccagtca cagaaaagca tcttacggat ggcatgacag taagagaatt 2340 atgcagtgct gccataacca tgagtgataa cactgcggcc aacttactc tgacaacgat 2400 cggaggaccg aaggagctaa ccgcttttt gcacacatg ggggatcatg taactcgcct 2460 tgatcgttgg gaaccggagc tgaatgaagc cataccaac gacgagcgtg acaccacgat 2520 gcctgtagcg atggcaaca cgttgcgcaa actattact ggcgaactac ttactctagc 2580 ttcccggcaa cattaatag actggatgga ggcggataa gttgcaggac cactctgcg 2640 ctcggcctt ccggctggct gttttattgc tgataaatcc ggagccggtg agcgtggttc 2700 tcgcggtatc atcgcagcgc tggggccaga tggtaagccc tccgtatcg tagttatcta 2760 cacgacgggg agtcaggcaa ctatggatga acgaaataga cagatcgctg agataggtgc 2820 ctcactgatt aagcattggt aactgcagga aaagggtacc actgagcgtc agaccccgta 2880 gaaaagagatca aaggatctc ttgagatcct tttttctgc gcgtaatctg ctgcttgcaa 2940 aaaaaaaac caccgctacc agcggtggtt tgtttgccgg atcagagct accactctt 3000 tttccgaagg taactggctt cagcagagcg cagataccaa atactgttct tctagtgtag 3060 ccgtagttag gccaccactt caagaactct gtagcaccgc ctacatacct cgctctgcta 3120 atcctgttac cagtggctgc tgccagtggc gataagtcgt gtcttaccgg gttggaccca 3180 agacgatagt taccggataa ggcgcagcgg tcgggctgaa cggggggttc gtgcacacag 3240 cccagcttgg agcgaacgac ctacaccgaa ctgagatacc tacagcgtga gctatgagaa 3300 agcgccacgc ttcccgaagg gagaaaggcg gacaggtatc cggtaagcgg cagggtcgga 3360 acaggagagc gcacgaggga gcttccaggg ggaaacgcct ggtatcttta tagtcctgtc 3420 gggtttcgcc acctctgact tgagcgtcga tttttgtgat gctcgtcagg ggggcggagc 3480 ctatggaaaa acgccagcaa cgcggccttt ttacggttcc tggccttttg ctggcctttt 3540 gctcacatgt tttgttcgat tattctccag ataaaatcaa caatagttgt ttgtaagtaa 3600 acgaatcaag atactgaaaa tagtttcaaa agcagatcat ctgggattta tatatcaggc 3660 atcctgcttt agttcttttt tgaacccaaa ggctatctga tgaaaagttg atataggtat 3720 gaagaccaga atttgcctag aggctaaccg agacctgagg ctaaaaaagg caggaggaaa 3780 agtcctgcca aagataggta ttgaacttg ttcgaaaag gcggaagtttt aaacacatgg 3840 ttggagcaag cggcggaata gcggagggat gatacgcagc aaggctggga tcattcgagt 3900 ttcaggaac gttagctcaa cattcattga ctggtaagcg acactggtt tcatctggggt 3960 ggagttagtc tggtgttggg atgctagttg ttcccacaa tgaggcca gatgaggagg 4020 atggtgtggt gataagagat gcaacagat ggttatggcc ttttgagaac aaagtagacc 4080 tgtcactcaa tgttgttta tatcattgct atttaaataa tgtatctaa cgcaactcc 4140 gaggctggaaa atgttaccg gcgatgcgcg gatattag aggcggcgat caagaacac 4200 ctgctgggcg agcagtctgg agcacagtct tcgatgggcc cgagatccca ccgcgttcct 4260 gggtaccggg acgtgaggca gcgcgacatc catcaatat accaggcgcc aaccgagtgt 4320 ctcggaaaac agcttctgga tatctccgc tggcggcgca acgacgaata atagtccctg 4380 gaggtgacgg atatatatg tgtggaggt aaatctgaca gggtgtagca aaggtatatat 4440 tttcctaaaa catgcaatcg gctgccccgc aacgggaaa agaatgactt tgcactt 4500 caccagagtg gggtgtcccg ctcgtgtgtg caaataggct cccactggtc accccggatt 4560 ttgcagaaaa acagcaagtt ccggggtgtc tcactggtgt ccgccaataa gaggagccgg 4620 caggcacgga gtttacatca agctgtctcc gatacactcg actaccatcc gggtctctca 4680 gagaggggaa tggcactata aataccgcct ccttgcgctc tctgccttca tcaatcaaat 4740 catgctgagg actcgaattc gacctctgtt gcctctttgt tggacgaacc attcaccggt 4800 gtcttgtact taaagggcag tggtatcact gaagacttcc agtccctaaa gggtaagaag 4860 atcggttacg ttggtgactt cggtaagatc caaatcgatg aattgaccaa gcactacggt 4920 atgaagccag aagactacac cgccgtcaga tgtggtatga atgtcgccaa gtacatcatc 4980 gaaggtaaga ttgatgccgg tattggtatc gaatgtatgc aacaagtcga attggaagag 5040 tacttggcca agcaaggcag accagcttct gatgctaaaa tgttgagaat tgacaagttg 5100 gcttgcttgg gttgctgttg cttctgtacc gttctttaca tctgcaacga tgaatttttg 5160 aagaagaacc ctgaaaaggt cagaaagttc ttgaaagcca tcaagaaggc aaccgactac 5220 gttctagccg accctgtgaa ggcttggaaa gaatacatcg acttcaagcc tcaattgaac 5280 aacgatctat cctacaagca ataccaaaga tgttacgctt acttctcttc atctttgtac 5340 aatgttcacc gtgactggaa gaaggttacc ggttacggta agagattagc catcttgcca 5400 ccagactatg tctcgaacta cactaatgaa tacttgtcct ggccagaacc agaagaggtt 5460 tctgatcctt tggaagctca aagattgatg gctattcatc aagaaaaatg cagacaggaa 5520 ggtactttca agagattggc tcttccagct taagcggccg cgagtcgtga gtaatcaaga 5580 ggatgtcaga atgccatttg cctgagagat gcaggcttca tttttgatac ttttttattt 5640 gtaacctata tagtatagga ttttttttgt cattttgttt cttctcgtac gagcttgctc 5700 ctgatcagcc tatctcgcag ctgatgaata tcttgtggta ggggtttggg aaaatcattc 5760 gagtttgatg tttttcttgg tatttcccac tcctcttcag agtacagaag attaagtgag 5820 acgttcgttt gtgctccgga caggtgaacc cacctaacta tttttaactg ggatccagtg 5880 agctcgctgg gtgaaagcca accatctttt gtttcgggga accgtgctcg ccccgtaaag 5940 ttaattttt ttcccgcgc agctttaatc ttcggcaga gaaggcgttt tcatcgtagc 6000 gtgggaacag aataatcagt tcatgtgcta tacaggcaca tggcagcagt cactattttg 6060 cttttaacc ttaaagtcgt tcatcaatca ttaactgacc aatcagattt tttgcatttg 6120 ccacttatct aaaatactt ttgtatctcg cagatacgtt cagtggtttc caggacaaca 6180 cccaaaaaaa ggtatcaatg ccactaggca gtcggtttta ttttggtca cccacgcaaa 6240 gaagcaccca cctcttttag gttttaagtt gtgggaacag taacaccgcc tagagcttca 6300 ggaaaaacca gtacctgtga ccgcaattca ccatgatgca gaatgttaat ttaaacgagt 6360 gccaaatcaa gatttcaaca gacaaatcaa tcgatccata gttacccatt ccagcctttt 6420 cgtcgtcgag cctgcttcat tcctgcctca ggtgcataac tttgcatgaa aagtccagat 6480 tagggcagat tttgagttta aaataggaaa tataaacaaa tataccgcga aaaaggtttg 6540 tttatagctt ttcgcctggt gccgtacggt ataaatacat actctcctcc cccccctggt 6600 tctctttttc ttttgttact tacattttac cgttccgtca ctcgcttcac tcaacaaaa 6660 aaatgattg gtatatccta gtcgttggta ttttgttgcc acagtcactg gctcacccag 6720 gcttcttcac ttctatagga cagatgactg atttgattca cacagaaaaa gacctagtta 6780 caagccttaa agactatatc aaagctgaag aggataagtt ggagcaaatc aaaaagtggg 6840 cagagaaact cgatagattg actagtactg caacaaaaga tcctgagggt tttgtgggtc 6900 acccagtgaa tgctttcaag ctgatgaaga gacttaatac agagtggtca gaattggaaa 6960 acttggtact taaagatatg agtgatggat tcatttctaa cttaacaatt caaagacaat 7020 actttccaaa cgatgaggac caagtaggag cagcaaaagc tttgttgcga ttgcaggaca 7080 catacaattt ggacaccgac acgatatcga agggtgattt acctggtgtg aagcataagt 7140 ccttcctcac tgtggaagat tgttttgaat tgggaaaagt cgcatataca gaagccgact 7200 actatcacac agaattatgg atggagcaag ctctgcgtca gttggacgaa ggtgaagtttt 7260 ctaccgttga taaggtttca gttttggatt acttatcata cgctgtttac cagcaaggtg 7320 atctggacaa agctctactt ttaactaaaa agttgttgga gctggacccg gagcatcaaa 7380 gagctaacgg taatctgaaa tactttgaat acatcatggc taaggaaaag gacgcaaata 7440 agtcctcgtc cgatgaccaa tccgatcaaa agaccactct gaaaaaaaaa ggtgcagctg 7500 7560 tgactccaag gagacagaaa aagctgttct gtagatatca tgatgggaac cgtaacccaa 7620 7680 ttcatgacat catctccgat gcagaaatag aggttgtgaa agacttggcc aaaccaagat 7740 tgagtagggc taccgtccat gaccctgaga ctggaaaatt gactaccgca caatatcgtg 7800 tctctaaatc agcatggttg tccggttacg agaatcccgt ggtcagccgt atcaatatgc 7860 gtattcaaga tttgactggt cttgacgtaa gcactgctga ggaactacaa gttgccaact 7920 atggtgtggg cggtcagtat gaacccact ttgatttcgc cagaaggac gagcctgatg 7980 cttttaagga gctaggtact ggaatagaa tcgcaacgtg gttgttctat atgtccgatg 8040 tgcttgctgg aggagccaca gttttccctg aggtaggtgc ttctgtttgg cctaaaaagg 8100 gcaccggccgt attttggtac aatctgtttg catctggaga aggtgattac agcactagac 8160 8220 gaggtcaaga gtttaggagg ccatgcacat tgtcagaatt agaaggttct ggctctggtt 8280 ctggctctat gagattccca tctattttca ccgctgtctt gttcgctgcc tcctctgcat 8340 tggctgcacc cgatgaggaa gatcatgttt tagtattgca taaggaaat ttcgatgaag 8400 ctttggccgc tcaacaatat ctgctcgtcg agttttacgc tccctggtgc ggtcattgta 8460 aggcccttgc accagagtac gccaaggcag ctggtaagtt aaaggccgaa ggttcagaga 8520 tcagattagc aaaagttgat gctacagaag agtccgatct tgctcaacaa tacggggttc 8580 gaggataccc aacaattaag ttttcaaaa atggtgatac tgcttcccca aaagaatata 8640 ctgctggtag agaggcagac gacatagtca actggctcaa aaagagaacg ggcccagctg 8700 cgtctacatt aagcgacgga gcagcagccg aagctcttgt ggaatctagt gaagttgctg 8760 taatcggttt ctttaaggac atggaatctg attcagctaa acagttcctt ttagcagctg 8820 aagcaatcga tgacatccct ttcggaatca cctcaaatag tgacgtgttc agcaagtacc 8880 aacttgacaa agatggagtg gtcttgttca aaaagtttga cgaaggcaga aacaatttcg 8940 agggtgaggt tacaaaggag aaactgcttg atttcattaa acataaccaa ctacccttag 9000 ttatcgaatt cactgaacaa actgctccta agattttcgg tggagaaatc aaaacacata 9060 tcttgttgtt tttgccaaag tccgtatcgg attatgaagg taaactctcc aatttcaaaa 9120 aggccgctga gagctttaag ggcaagattt tgttcatctt tattgactca gaccacacag 9180 acaatcagag gattttggag tttttcggtt tgaaaaagga ggaatgtcca gcagtccgtt 9240 tgatcacctt ggaggaggag atgaccaaat acaaaccaga gtcggatgag ttgactgccg 9300 agaagataac agaattttgt cacagatttc tggaaggtaa gatcaagcct catcttatgt 9360 ctcaagagtt gcctgatgac tgggataagc aaccagttaa agtattggtg ggtaaaaact 9420 ttgaggaagt ggccttcgac gagaaaaaaa atgtctttgt tgaattctat gctccgtggt 9480 gtggtcactg taagcagctg gcaccaattt gggataaact gggtgaaact tacaaagatc 9540 acgaaaacat tgttattgca aagatggaca gtactgctaa cgaagtggag gctgtgaaag 9600 ttcactcctt ccctacgctg aagttctttc ctgcatctgc tgacagaact gttatcgact 9660 ataatggaga gaggacattg gatggtttta aaaagtttct tgaatccgga ggtcaagacg 9720 gagctggtga cgacgatgat ttggaagatc tggaggaggc tgaggaacct gatcttgagg 9780 aggatgacga ccagaaggca gtcaaagatg aactgtgata aggggtcaag aggatgtcag 9840 aatgccattt gcctgagaga tgcaggcttc atttttgata cttttttatt tgtaacctat 9900 atagtatagg attttttttg tcattttgtt tcttctcgta cgagcttgct cctgatcagc 9960 ctatctcgca gcagatgaat atcttgtggt aggggtttgg gaaaatcatt cgagtttgat 10020 gtttttcttg gtatttccca ctcctcttca gagtacagaa gattaagtga gaccttcgtt 10080 tgtgcggttc tggctctggt tctggctct 10109 <210> 10 <211> 10075 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Nucleotides <400> 10 ggatccttca gtaatgtctt gtttctttg ttgcagtggt gagccatttt gacttcgtga 60 aagttcttt agaatagttg tttccagagg ccaaacattc cacccgtagt aaagtgcaag 120 cgtaggaaga ccaagactgg cataaatcag gtataagtgt cgagcactgg caggtgatct 180 tctgaaagtt tctactagca gataagatcc agtagtcatg catatggcaa caatgtaccg 240 tgtggatcta agaacgcgtc ctactaacct tcgcattcgt tggtccagtt tgttgttac 300 gatcaacgtg acaaggttgt cgattccgcg taagcatgca tacccaagga cgcctgttgc 360 aattccaagt gagccagttc caacaatctt tgtaatatta gagcacttca ttgtgttgcg 420 cttgaaagta aaatgcgaac aaattaagag ataatctcga aaccgcgact tcaaacgcca 480 atatgatgtg cggcacacaa taagcgttca tatccgctgg gtgactttct cgctttaaaa 540 aattatccga aaaaatttc ctctagaatg ggtaaggaaa agactcacgt ttcgaggccg 600 cgattaaatt ccaacatgga tgctgattta tatgggtata aatgggctcg cgataatgtc 660 gggcaatcag gtgcgacaat ctatcgattg tatgggaagc ccgatgcgcc agagttgttt 720 ctgaaacatg gcaaaggtag cgttgccaat gatgttacag atgagatggt cagactaaac 780 tggctgacgg aattatgcc tcttccgacc atcaagcatt ttccgtac tcctgatgat 840 gcatggttac tcaccactgc gatccccggc aaaacagcat tccaggtatt agaagaatat 900 cctgattcag gtgaaaatat tgttgatgcg ctggcagtgt tctgcgccg gttgcattcg 960 attcctgttt gtaattgtcc ttttaacagc gatcgcgtat ttcgtctcgc tcaggcgcaa 1020 tcacgaatga ataacggttt ggttgatgcg agtgattttg atgacgagcg taatggctgg 1080 cctgttgaac aagtctggaa agaaatgcat aagcttttgc cattctcacc ggattcagtc 1140 gtcactcatg gtgatttctc acttgataac cttatttttg acgaggggaa attaataggt 1200 tgtattgatg ttggacgagt cggaatcgca gaccgatacc aggatcttgc catcctatgg 1260 aactgcctcg gtgagttttc tccttcatta cagaaacggc tttttcaaaa atatggtatt 1320 gataatcctg atatgaataa attgcagttt catttgatgc tcgatgagtt tttctaaaat 1380 tgacacctta cgattattta gagagtattt attagtttta ttgtatgtat acggatgttt 1440 tattatctat tttgccctt atattctgta actatccaaa agtcctatct tatcaagcca 1500 gcaatctatg tccgcgaacg tcaactaaaa ataagctttt tatgctgttc tctcttttt 1560 tcccttcggt aatatatac cttgcatcca cagattctcc tgccaaattt tgcataatcc 1620 tttacaacat ggctatatgg gagcacttag cgccctccaa aacccatatt gcctacgcat 1680 gtataggtgt ttttccaca atattttctc tgtgctctct ttttattaaa gagaagctct 1740 atatcggaga agcttctgtg gccgttatat tcggccttat cgtgggacca cattgcctga 1800 attggttgc cccggaagat tggggaaact tggatctgat taccttagct gcatcagaat 1860 tggttaattg gttgtaacac tgacccctat ttgtttatttt ttctaaatac attcaaatat 1920 gtatccgctc atgagacaat aaccctgata aatgcttcaa tatattgaa aaaggaagaa 1980 tatgagtatt caacatttcc gtgcgccct tatcccttt tttgcggcat tttgccttcc 2040 tgtttttgct cacccagaaa cgctggtgaa agtaaaagat gctgaagatc agttgggtgc 2100 acgagtgggt tacatcgaac tggatctcaa cagcggtaag atccttgaga gttttcgccc 2160 cgaagaacgt tttccaatga tgagcacttt taaagttctg ctatgtggcg cggtattatc 2220 ccgtattgac gccgggcaag agcaactcgg tcgccgcat cactattctc agaatgactt 2280 ggttgagtac tcaccagtca cagaaaagca tcttacggat ggcatgacag taagagaatt 2340 atgcagtgct gccataacca tgagtgataa cactgcggcc aacttacttc tgacaacgat 2400 cggaggaccg aaggagctaa ccgcttttt gcacaacatg ggggatcatg taactcgcct 2460 tgatcgttgg gaaccggagc tgaatgaagc cataccaaac gacgagcgtg acaccacgat 2520 gcctgtagcg atggcaacaa cgttgcgcaa actattaact ggcgaactac ttactctagc 2580 ttcccggcaa caattaatag actggatgga ggcggataaa gttgcaggac cacttctgcg 2640 ctcggccctt ccggctggct ggtttattgc tgataaatcc ggagccggtg agcgtggttc 2700 tcgcggtatc atcgcagcgc tggggccaga tggtaagccc tcccgtatcg tagttatcta 2760 cacgacgggg agtcaggcaa ctatggatga agaataga cagatcgctg agataggtgc 2820 ctcactgatt aagcattggt aactgcagga aaagggtacc actgagcgtc agaccccgta 2880 gaaaagatca aaggatcttc ttgagatcct ttttttctgc gcgtaatctg ctgcttgcaa 2940 acaaaaaaac caccgctacc agcggtggtt tgtttgccgg atcaagagct accaactctt 3000 tttccgaagg taactggctt cagcagagcg cagataccaa atactgttct tctagtgtag 3060 ccgtagttag gccaccactt caagaactct gtagcaccgc ctacatacct cgctctgcta 3120 atcctgttac cagtggctgc tgccagtggc gataagtcgt gtcttaccgg gttggaccca 3180 agacgatagt taccggataa ggcgcagcgg tcgggctgaa cggggggttc gtgcacacag 3240 cccagcttgg agcgaacgac ctacaccgaa ctgagatacc tacagcgtga gctatgagaa 3300 agcgccacgc ttcccgaagg gagaaaggcg gacaggtatc cggtaagcgg cagggtcgga 3360 acaggagagc gcacgaggga gcttccaggg ggaaacgcct ggtatcttta tagtcctgtc 3420 gggttcgcc acctctgact tgagcgtcga ttttgtgat gctcgtcagg ggggcggagc 3480 ctatggaaaa acgccagcaa cgcggccttt ttacggttcc tggccttttg ctggcctttt 3540 gctcacatgt tttgttcgat tattctccag ataaaatcaa caatagttgt ttgtaagtaa 3600 acgaatcaag attackgaaaa tagttcaa agcagatcat ctgggattta tatatcaggc 3660 atcctgcttt agttctttt tgaacccaaa ggctatctga tgaaagttg atataggtat 3720 gagaccaga atttgcctag aggctaccg agacctgagg ctaaaaaagg caggaggaaa 3780 agtcctgcca aagataggta ttgaacttg ttcgaaaag gcggaagtttt aaacacatgg 3840 ttggagcaag cggcggaata gcggagggat gatacgcagc aaggctggga tcattcgagt 3900 ttcaggaac gttagctcaa cattcattga ctggtaagcg acactggtt tcatctggggt 3960 ggagttagtc tggtgttggg atgctagttg ttcccacaa tgaggcca gatgaggagg 4020 atggtgtggt gataagagat gcaacagat ggttatggcc ttttgagaac aaagtagacc 4080 tgtcactcaa tgttgttta tatcattgct atttaaataa tgtatctaa cgcaactcc 4140 gaggctggaaa atgttaccg gcgatgcgcg gatattag aggcggcgat caagaacac 4200 ctgctgggcg agcagtctgg agcacagtct tcgatgggcc cgagatccca ccgcgttcct 4260 gggtaccggg acgtgaggca gcgcgacatc catcaatat accaggcgcc aaccgagtgt 4320 ctcggaaaac agcttctgga tatcttccgc tggcggcgca acgacgaata atagtccctg 4380 gaggtgacgg aatatatatg tgtggagggt aaatctgaca gggtgtagca aaggtaatat 4440 tttcctaaaa catgcaatcg gctgccccgc aacgggaaaa agaatgactt tggcactctt 4500 caccagagtg gggtgtcccg ctcgtgtgtg caaataggct cccactggtc accccggatt 4560 ttgcagaaaa acagcaagtt ccggggtgtc tcactggtgt ccgccaataa gaggagccgg 4620 caggcacgga gtttacatca agctgtctcc gatacactcg actaccatcc gggtctctca 4680 gagaggggaa tggcactata aataccgcct ccttgcgctc tctgccttca tcaatcaaat 4740 catgctgagg actcgaattc gacctctgtt gcctctttgt tggacgaacc attcaccggt 4800 gtcttgtact taaagggcag tggtatcact gaagacttcc agtccctaaa gggtaagaag 4860 atcggttacg ttggtgactt cggtaagatc caaatcgatg aattgaccaa gcactacggt 4920 atgaagccag aagactacac cgccgtcaga tgtggtatga atgtcgccaa gtacatcatc 4980 gaaggtaaga ttgatgccgg tattggtatc gaatgtatgc aacaagtcga attggaagag 5040 tacttggcca agcaaggcag accagcttct gatgctaaaa tgttgagaat tgacaagttg 5100 gcttgcttgg gttgctgttg cttctgtacc gttctttaca tctgcaacga tgaatttttg 5160 aagaagaacc ctgaaaaggt cagaaagttc ttgaaagcca tcaagaaggc aaccgactac 5220 gttctagccg accctgtgaa ggcttggaaa gaatacatcg acttcaagcc tcaattgaac 5280 aacgatctat cctacaagca ataccaaaga tgttacgctt acttctcttc atctttgtac 5340 aatgttcacc gtgactggaa gaaggttacc ggttacggta agagattagc catcttgcca 5400 ccagactatg tctcgaacta cactaatgaa tacttgtcct ggccagaacc agaagaggtt 5460 tctgatcctt tggaagctca aagattgatg gctattcatc aagaaaaatg cagacaggaa 5520 ggtactttca agagattggc tcttccagct taagcggccg cgagtcgtga gtaatcaaga 5580 ggatgtcaga atgccatttg cctgagagat gcaggcttca tttttgatac ttttttattt 5640 gtaacctata tagtatagga ttttttttgt cattttgttt cttctcgtac gagcttgctc 5700 ctgatcagcc tatctcgcag ctgatgaata tcttgtggta ggggtttggg aaaatcattc 5760 gagtttgatg ttttcttgg tatttcccac tcctcttcag agtacagaag attaagtgag 5820 acgttcgttt gtgctccgga caggtgaacc cacctaacta ttttaactg ggatccagtg 5880 agctcgctgg gtgaaagcca accatcttttt gttcgggga accgtgctcg ccccgtaaag 5940 ttaattttt ttcccgcgc agctttaatc ttcggcaga gaaggcgttt tcatcgtagc 6000 gtgggaacag aataatcagt tcatgtgcta tacaggcaca tggcagcagt cactattttg 6060 cttttaacc ttaaagtcgt tcatcaatca ttaactgacc aatcagattt tttgcatttg 6120 ccacttatct aaaatactt ttgtatctcg cagatacgtt cagtggtttc caggacaaca 6180 cccaaaaaaa ggtatcaatg ccactaggca gtcggtttta ttttggtca cccacgcaaa 6240 gaagcaccca cctcttttag gttttaagtt gtgggaacag taacaccgcc tagagcttca 6300 ggaaaaacca gtacctgtga ccgcaattca ccatgatgca gaatgttaat ttaaacgagt 6360 gccaaatcaa gatttcaaca gacaaatcaa tcgatccata gttacccatt ccagcctttt 6420 cgtcgtcgag cctgcttcat tcctgcctca ggtgcataac tttgcatgaa aagtccagat 6480 tagggcagat tttgagttta aaataggaaa tataaacaaa tataccgcga aaaaggtttg 6540 tttatagctt ttcgcctggt gccgtacggt ataaatacat actctcctcc cccccctggt 6600 tctctttttc ttttgttact tacattttac cgttccgtca ctcgcttcac tcaacaaaa 6660 aaatgagatt cccatctatt ttcaccgctg tcttgttcgc tgcctcctct gcattggctg 6720 cacccgatga ggaagatcat gttttagtat tgcataaagg aaatttcgat gaagctttgg 6780 ccgctcacaa atactgctc gtcgagtttt acgctccctg gtgcggtcat tgtaaggccc 6840 ttgcaccaga gtacgccaag gcagctggta agttaaaggc cgaaggttca gagatcagat 6900 tagcaaaagt tgatgctaca gaagagtccg atcttgctca acaatacggg gttcgaggat 6960 acccaacaat taagtttttc aaaaatggtg atactgcttc cccaaaggaa tatactgctg 7020 gtagagaggc agacgacata gtcaactggc tcaaaaagag aacgggccca gctgcgcta 7080 cattaagcga cggagcagca gccgaagctc ttgtggaatc tagtgaagtt gctgtaatcg 7140 gttctttaa ggacatggaa tctgattcag ctaaacagtt ccttttagca gctgaagcaa 7200 tcgatgacat ccctttcgga atcacctcaa atagtgacgt gttcagcag taccacttg 7260 aaaagatgg agtggtcttg ttcaaaagt tcgaggagg cagaaacaat tcgagggtg 7320 aggttacaa ggagaaactg cttgatttca ttaacataa ccaactacc ttagttatcg 7380 aattcactga acaactgct cctagattt tcggtggaga atcaaaca catatcttgt 7440 tgtttttgcc aaagtccgta tcggattatg aaggtaact ctccaatttc aaaaaggccg 7500 ctgagagctt taagggcaag atttgttca tctttattga ctcagaccac acagacaatc 7560 agaggatttt ggagttttc ggtttgaaaa aggaggaatg tccagcagtc cgttgatca 7620 ccttggagga ggagatgacc aaatacaaac cagagtcgga tgagttgact gccgagaga 7680 taacagaatt ttgtcacaga tttctggaag gtaagatcaa gcctcatctt atgtctcaag 7740 agttgcctga tgactgggat aagcaccag ttaaagtatt ggtgggtaa aactttgagg 7800 aagtggcctt cgacgagaaa aaaatgtct ttgttgaatt ctagctccg tggtgtggtc 7860 actgtaagca gctggcacca atttgggata aactgggtga aacttacaa gatcacgaaa 7920 acattgttat tgcaaagatg gacagtactg ctaacgaagt ggaggctgtg aaagttcact 7980 ccttccctac gctgaagttc tttcctgcat ctgctgacag aactgttatc gactataatg 8040 gagagaggac attggatggt tttaaaaagt ttcttgaatc cggaggtcaa gacggagctg 8100 gtgacgacga tgatttggaa gatctggagg aggctgagga acctgatctt gaggaggatg 8160 acgaccagaa ggcagtcaaa gatgaactgg gttctggctc tggttctggc tctatgattt 8220 ggtatatcct agtcgttggt attttgttgc cacagtcact ggctcaccca ggcttcttca 8280 cttctatagg acagatgact gatttgattc acacagaaaa agacctagtt acaagcctta 8340 aagactatat caaagctgaa gaggataagt tggagcaaat caaaaagtgg gcagagaaac 8400 tcgatagatt gactagtact gcaacaaaag atcctgaggg ttttgtgggt cacccagtga 8460 atgctttcaa gctgatgaag agacttaata cagagtggtc agaattggaa aacttggtac 8520 ttaaagatat gagtgatgga ttcatttcta acttaacaat tcaaagacaa tactttccaa 8580 acgatgagga ccaagtagga gcagcaaaag ctttgttgcg attgcaggac acatacaatt 8640 tggacaccga cacgatatcg aagggtgatt tacctggtgt gaagcataag tccttcctca 8700 ctgtggaaga ttgttttgaa ttgggaaaag tcgcatatac agaagccgac tactatcaca 8760 cagaattatg gatggagcaa gctctgcgtc agttggacga aggtgaagtt tctaccgttg 8820 ataaggtttc agttttggat tacttatcat acgctgttta ccagcaaggt gatctggaca 8880 aagctctact tttaactaaa aagttgttgg agctggaccc ggagcatcaa agagctaacg 8940 gtaatctgaa atactttgaa tacatcatgg ctaaggaaaa ggacgcaaat aagtcctcgt 9000 ccgatgacca atccgatcaa aagaccactc tgaaaaaaaa aggtgcagct gttgactacc 9060 tcccagagag acaaaagtat gaaatgctgt gtagaggaga gggtatcaag atgactccaa 9120 ggagacagaa aaagctgttc tgtagatatc atgatgggaa ccgtaaccca aaattcattc 9180 ttgctccagc gaaacaggaa gatgaatggg acaagcctag aatcattcgt tttcatgaca 9240 tcatctccga tgcagaaata gaggttgtga aagacttggc caaaccaaga ttgagtaggg 9300 ctaccgtcca tgaccctgag actggaaaat tgactaccgc acaatatcgt gtctctaaat 9360 cagcatggtt gtccggttac gagaatcccg tggtcagccg tatcaatatg cgtattcaag 9420 atttgactgg tcttgacgta agcactgctg aggaactaca agttgccaac tatggtgtgg 9480 gcggtcagta tgaaccccac tttgattcg ccagaaagga cgagcctgat gcttttaagg 9540 agctaggtac tggaataga atcgcaacgt ggttgttcta tatgtccgat gtgcttgctg 9600 gaggagccac agttttccct gaggtaggtg cttctgtttg gcctaaaaag ggcacggccg 9660 tattttggta caatctgttt gcatctggag aaggtgatta cagcactaga catgctgctt 9720 gtcccgtctt agtcggtaat aagtgggttt ccaataagtg gctgcatgag agaggtcaag 9780 agtttaggag gccatgcaca ttgtcagaat tagaatgata atttacggg aagtctttac 9840 agttttagtt aggagccctt atatatgaca gtaatgctag tacgttttgt tttgtttaat 9900 tataactta gtttatgtta gcctagtata gactccatca attttttttg ttattacgta 9960 agccgcgatg ataatatctg atgaaaaatt cctatcagaa aataatttat caaaagtttc 10020 atgcgatatg agactaagta gaataggac tcccaaagtg tcagtcacaa gggtc 10075 <210> 11 <211> 8413 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotide <400> 11 ggatccttca gtaatgtctt gtttcttttg ttgcagtggt gagccatttt gacttcgtga 60 aagtttcttt agaatagttg tttccagagg ccaaacattc cacccgtagt aaagtgcaag 120 cgtaggaaga ccaagactgg cataaatcag gtataagtgt cgagcactgg caggtgatct 180 tctgaaagtt tctactagca gataagatcc agtagtcatg catatggcaa caatgtaccg 240 tgtggatcta agaacgcgtc ctactaacct tcgcattcgt tggtccagtt tgttgttatc 300 gatcaacgtg acaaggttgt cgattccgcg taagcatgca tacccaagga cgcctgttgc 360 aattccaagt gagccagttc caacaatctt tgtaatatta gagcacttca ttgtgttgcg 420 cttgaaagta aaatgcgaac aaattaagag ataatctcga aaccgcgact tcaaacgcca 480 atatgatgtg cggcacacaa taagcgttca tatccgctgg gtgactttct cgctttaaaa 540 aattatccga aaaaattttc ctctagaatg ggtaaggaaa agactcacgt ttcgaggccg 600 cgattaaatt ccaacatgga tgctgattta tatgggtata aatgggctcg cgataatgtc 660 gggcaatcag gtgcgacaat ctatcgattg tatgggaagc ccgatgcgcc agagttgttt 720 ctgaaacatg gcaaaggtag cgttgccaat gatgttacag atgagatggt cagactaaac 780 tggctgacgg aatttatgcc tcttccgacc atcaagcatt ttatccgtac tcctgatgat 840 gcatggttac tcaccactgc gatccccggc aaaacagcat tccaggtatt agaagaatat 900 cctgattcag gtgaaaatat tgttgatgcg ctggcagtgt tcctgcgccg gttgcattcg 960 attcctgttt gtaattgtcc ttttaacagc gatcgcgtat ttcgtctcgc tcaggcgcaa 1020 tcacgaatga ataacggttt ggttgatgcg agtgattttg atgacgagcg taatggctgg 1080 cctgttgaac aagtctggaa agaaatgcat aagcttttgc cattctcacc ggattcagtc 1140 gtcactcatg gtgatttctc acttgataac cttatttttg acgaggggaa attaataggt 1200 tgtattgatg ttggacgagt cggaatcgca gaccgatacc aggatcttgc catcctatgg 1260 aactgcctcg gtgagttttc tccttcatta cagaaacggc tttttcaaaa atatggtatt 1320 gataatcctg atatgaataa attgcagttt catttgatgc tcgatgagtt tttctaaaat 1380 tgacacctta cgattattta gagagtattt attagtttta ttgtatgtat acggatgttt 1440 tattatctat ttatgccctt atattctgta actatccaaa agtcctatct tatcaagcca 1500 gcaatctatg tccgcgaacg tcaactaaaa ataagctttt tatgctgttc tctctttttt 1560 tcccttcggt ataattatac cttgcatcca cagattctcc tgccaaattt tgcataatcc 1620 tttacaacat ggctatatgg gagcacttag cgccctccaa aacccatatt gcctacgcat 1680 gtataggtgt tttttccaca atattttctc tgtgctctct ttttattaaa gagaagctct 1740 atatcggaga agcttctgtg gccgttatat tcggccttat cgtgggacca cattgcctga 1800 attggtttgc cccggaagat tggggaaact tggatctgat taccttagct gcattaccaa 1860 tgcttaatca gtgaggcacc tatctcagcg atctgtctat ttcgttcatc catagttgcc 1920 tgactccccg tcgtgtagat aactacgata cgggagggct taccatctgg ccccagcgct 1980 gcgatgatac cgcgagaacc acgctcaccg gctccggatt tatcagcaat aaaccagcca 2040 gccggaaggg ccgagcgcag aagtggtcct gcaactttat ccgcctccat ccagtctatt 2100 aattgttgcc gggaagctag agtaagtagt tcgccagtta atagtttgcg caacgttgtt 2160 gccatcgcta caggcatcgt ggtgtcacgc tcgtcgtttg gtatggcttc attcagctcc 2220 ggttcccaac gatcaaggcg agttacatga tcccccatgt tgtgcaaaaa agcggttagc 2280 tccttcggtc ctccgatcgt tgtcagaagt aagttggccg cagtgttatc actcatggtt 2340 atggcagcac tgcataattc tcttactgtc atgccatccg taagatgctt ttctgtgact 2400 ggtgagtact caaccaagtc attctgagaa tagtgtatgc ggcgaccgag ttgctcttgc 2460 ccggcgtcaa tacgggataa taccgcgcca catagcagaa ctttaaaagt gctcatcatt 2520 ggaaaacgtt cttcggggcg aaaactctca aggatcttac cgctgttgag atccagttcg 2580 atgtaaccca ctcgtgcacc caactgatct tcagcatctt ttactttcac cagcgtttct 2640 gggtgagcaa aaacaggaag gcaaaatgcc gcaaaaaagg gaataagggc gacacggaaa 2700 tgttgaatac tcatattctt cctttttcaa tattattgaa gcatttatca gggttattgt 2760 ctcatgagcg gatacatatt tgaatgtatt tagaaaaata aacaaatagg ggtcagtgtt 2820 acaaccaatt aaccaattct gaaaggaaga atctgcagga aaagggtacc actgagcgtc 2880 agaccccgta gaaaagatca aaggatcttc ttgagatcct ttttttctgc gcgtaatctg 2940 ctgcttgcaa acaaaaaaac caccgctacc agcggtggtt tgtttgccgg atcaagagct 3000 accaactctt tttccgaagg taactggctt cagcagagcg cagataccaa atactgttct 3060 tctagtgtag ccgtagttag gccaccactt caagaactct gtagcaccgc ctacatacct 3120 cgctctgcta atcctgttac cagtggctgc tgccagtggc gataagtcgt gtcttaccgg 3180 gttggaccca agacgatagt taccggataa ggcgcagcgg tcgggctgaa cggggggttc 3240 gtgcacacag cccagcttgg agcgaacgac ctacaccgaa ctgagatacc tacagcgtga 3300 gctatgagaa agcgccacgc ttcccgaagg gagaaaggcg gacaggtatc cggtaagcgg 3360 cagggtcgga acaggagagc gcacgaggga gcttccaggg ggaaacgcct ggtatcttta 3420 tagtcctgtc gggtttcgcc acctctgact tgagcgtcga tttttgtgat gctcgtcagg 3480 ggggcggagc ctatggaaaa acgccagcaa cgcggccttt ttacggttcc tggccttttg 3540 ctggcctttt gctcacatgt tttgttcgat tattctccag ataaaatcaa caatagttgt 3600 ttgtaagtaa acgaatcaag atactgaaaa tagtttcaaa agcagatcat ctgggattta 3660 tatatcaggc atcctgcttt agttctttt tgaacccaaa ggctatctga tgaaaagttg 3720 atataggtat gaagaccaga atttgcctag aggctaaccg agacctgagg ctaaaaaagg 3780 caggaggaaa agtcctgcca aagataggta tttgaacttg ttcgaaaaag gcggaagtttt 3840 aaacacatgg ttggagcaag cggcggaata gcggagggat gatacgcagc aaggctggga 3900 tcattcgagt ttcaaggaac gttagctcaa cattcattga ctggtaagcg acaactggtt 3960 tcatctgggt ggagttagtc tggtgttggg atgctagttg ttccccacaa ttgaaggcca 4020 gatgaggagg atggtgtggt gataagagat gcaaacagat ggttatggcc ttttgagaac 4080 aaagtagacc tgtcactcaa ttgttgtta tatcattgct atttaaatca ggtgaaccca 4140 cctaactatt tttaactggc atccagtgag ctcgctgggt gaaagccaac catctttgtgt 4200 ttcggggaac cgtgctcgcc ccgtaaagtt aatttttttt tcccgcgcag ctttaatctt 4260 tcggcagaga aggcgttttc atcgtagcgt gggaacagaa taatcagttc atgtgctata 4320 caggcacatg gcagcagtca ctattttgct ttttaacctt aaagtcgttc atcaatcatt 4380 aactgaccaa tcagattttt tgcatttgcc acttatctaa aaatactttt gtatctcgca 4440 gatacgttca gtggtttcca ggacaacacc caaaaaaagg tatcaatgcc actaggcagt 4500 cggttttatt tttggtcacc cacgcaaaga agcacccacc tcttttaggt tttaagttgt 4560 gggaacagta acaccgccta gagcttcagg aaaaaccagt acctgtgacc gcaattcacc 4620 atgatgcaga atgttaattt aaacgagtgc caaatcaaga tttcaacaga caaatcaatc 4680 gatccatagt tacccattcc agccttttcg tcgtcgagcc tgcttcattc ctgcctcagg 4740 tgcataactt tgcatgaaaa gtccagatta gggcagattt tgagtttaaa ataggaaata 4800 taaacaaata taccgcgaaa aaggtttgtt tatagctttt cgcctggtgc cgtacggtat 4860 aaatacatac tctcctcccc cccctggttc tctttttctt ttgttactta cattttaccg 4920 ttccgtcact cgcttcactc aacaacaaaa atgttctctc caattttgtc cttggaaatt 4980 attttagctt tggctacttt gcaatctgtc ttcgctcacc caggcttctt cactttata 5040 ggacagatga ctgatttgat tcacacagaa aaagacctag ttacaagcct taaagactat 5100 atcaaagctg aagaggataa gttggagcaa atcaaaaagt gggcagagaa actcgataga 5160 5220 5280 atgagtgatg gattcatttc taacttaaca attcaaagac aatactttcc aaacgatgag 5340 gaccaagtag gagcagaaa agctttgttg cgattgcagg acacatacaa tttggacacc 5400 gacacgatat cgaagggtga tttacctggt gtgaagcata agtccttcct cactgtggaa 5460 gattgttttg aattgggaaa agtcgcatat agaagccg actactatca cacagaatta 5520 tggatggagc aagctctgcg tcagttggac gaaggtgaag tttctaccgt tgataaggtt 5580 tcagttttgg attacttatc atacgctgtt taccagcaag gtgatctgga caaagctcta 5640 cttttaacta aaaagttgtt ggagctggac ccggagcatc aaagagctaa cggtaatctg 5700 aaatactttg aatacatcat ggctaaggaa aaggacgcaa ataagtcctc gtccgatgac 5760 caatccgatc aaaagaccac tctgaaaaaa aaaggtgcag ctgttgacta cctcccagag 5820 agacaaaagt atgaaatgct gtgtagagga gagggtatca agatgactcc aaggagacag 5880 aaaaagctgt tctgtagata tcatgatggg aaccgtaacc caaaattcat tcttgctcca 5940 gcgaaacagg aagatgaatg ggacaagcct agaatcattc gttttcatga catcatctcc 6000 gatgcagaaa tagaggttgt gaaagacttg gccaaaccaa gattgagtag ggctaccgtc 6060 catgaccctg agactggaaa attgactacc gcacaatatc gtgtctctaa atcagcatgg 6120 ttgtccggtt acgagaatcc cgtggtcagc cgtatcaata tgcgtattca agatttgact 6180 ggtcttgacg taagcactgc tgaggacta caagttgcca actatggtgt gggcggtcag 6240 tatgaacccc actttgattt cgccagaaag gacgagcctg atgcttttaa ggagctaggt 6300 actggaaata gaatcgcaac gtggttgttc tatatgtccg atgtgcttgc tggaggagcc 6360 acagttttcc ctgaggtagg tgcttctgtt tggcctaaaa agggcacggc cgtattttgg 6420 tacaatctgt ttgcatctgg agaaggtgat tacagcacta gacatgctgc ttgtcccgtc 6480 ttagtcggta ataagtgggt ttccaataag tggctgcatg agagaggtca agagtttagg 6540 aggccatgca cattgtcaga attagaaggt tctggctctg gttctggctc tatgagattc 6600 ccatctattt tcaccgctgt cttgttcgct gcctcctctg cattggctgc acccgatgag 6660 gaagatcatg ttttagtatt gcataaagga aatttcgatg aagctttggc cgctcacaaa 6720 tatctgctcg tcgagtttta cgctccctgg tgcggtcatt gtaaggccct tgcaccagag 6780 tacgccaagg cagctggtaa gttaaaggcc gaaggttcag agatcagatt agcaaaagtt 6840 gatgctacag aagagtccga tcttgctcaa caatacgggg ttcgaggata cccaacaatt 6900 aagtttttca aaaatggtga tactgcttcc ccaaaggaat atactgctgg tagagaggca 6960 gacgacatag tcaactggct caaaaagaga acgggcccag ctgcgtctac attaagcgac 7020 ggagcagcag ccgaagctct tgtggaatct agtgaagttg ctgtaatcgg tttctttaag 7080 gatgaatgaat ctgattcagc window cttttagcag ctgaagcaat cgatgacatc 7140 cctttcggaa tcacctcaaa tagtgacgtg ttcagcagt accaacttga caagatgga 7200 gtggtcttgt tcaaaagtt tgacgaaggc agaacaatt tcgagggtga ggttacaaag 7260 gagaaactgc ttgatttcat taaacatac caactaccct tagttatcga attcactgaa 7320 caaactgctc ctaagatttt cggtggagaa atcaaacac atatcttgtt gttttgcca 7380 aagtccgtat cggattatga agtaaacctc tccaatttca aaaagccgc tgagagcttt 7440 aagggcaaga tttgttcat ctttattgac tcagaccaca cagacaatca gaggatttg 7500 gagtttttcg gtttgaaaaa ggaggaatgt ccagcagtcc gttgatcac ctggaggag 7560 gagatgacca atacaacc agagtcggat gagttgactg ccgagagat aacagaattt 7620 tgtcacagat ttctggagg tagatcaag cctcatctta tgtctcaaga gttgcctgat 7680 gactgggata agcaccagt taagtattg gtgggtaaaa actttgagga agtggccttc 7740 gacgagaaaa aaatgtctt tgttgaatttc tatgctccgt ggtgtggtca ctgtaagcag 7800 ctggcaccaa tttgggataa actgggtgaa acttacaaag atcacgaaaa cattgttatt 7860 gcaaagatgg acagtactgc taacgaagtg gaggctgtga aagttcactc cttccctacg 7920 ctgaagttct ttcctgcatc tgctgacaga actgttatcg actataatgg agagaggaca 7980 ttggatggtt ttaaaaagtt tcttgaatcc ggaggtcaag acggagctgg tgacgacgat 8040 gatttggaag atctggagga ggctgaggaa cctgatcttg aggaggatga cgaccagaag 8100 gcagtcaaag atgaactgtg ataagggggg ccgcgagtcg tgagtaatca agaggatgtc 8160 agaatgccat ttgcctgaga gatgcaggct tcatttttga tactttttta tttgtaacct 8220 atatagtata ggattttttt tgtcattttg tttcttctcg tacgagcttg ctcctgatca 8280 gcctatctcg cagctgatga atatcttgtg gtaggggttt gggaaaatca ttcgagtttg 8340 atgtttttct tggtatttcc cactcctctt cagagtacag aagattaagt gagacgttcg 8400 tttgtgctcc gga 8413 <210> 12 <211> 714 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotide <400> 12 gttcttgaa tccggaggtc aagacggagc tggtgacgac gatgatttgg aagatctgga 60 ggaggctgag gaacctgatc ttgaggagga tgacgaccag aaggcagtca aagatgaact 120 gcatcatcat catcatcatt gataaggggt caagaggatg tcagaatgcc atttgcctga 180 gagatgcagg cttcattttt gatacttttt tatttgtaac ctatatagta taggattttt 240 tttgtcattt tgtttcttct cgtacgagct tgctcctgat cagcctatct cgcagcagat 300 gaatatcttg tggtaggggt ttgggaaaat cattcgagtt tgatgttttt cttggtattt 360 cccactcctc ttcagagtac agaagattaa gtgagacctt cgtttgtgcg gttctggctc 420 tggttctggc tctggatcct tcagtaatgt cttgtttctt ttgttgcagt ggtgagccat 480 ttgacttcg tgaaagtttc tttagaatag ttgtttccag aggccaaaca ttccacccgt 540 agtaaagtgc aagcgtagga agaccaagac tggcataaat caggtataag tgtcgagcac 600 tggcaggtga tcttctgaaa gtttctacta gcagataaga tccagtagtc atgcatatgg 660 caacaatgta ccgtgtggat ctaagaacgc gtcctactaa ccttcgcatt cgtt 714 <210> 13 <211> 7605 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotide <400> 13 tgcaggtacc actgagcgtc agaccccgta gaaaagatca aaggatcttc ttgagatcct 60 ttttttctgc gcgtaatctg ctgcttgcaa acaaaaaaac caccgctacc agcggtggtt 120 tgtttgccgg atcaagagct accaactctt tttccgaagg taactggctt cagcagagcg 180 cagataccaa atactgttct tctagtgtag ccgtagttag gccaccactt caagaactct 240 gtagcaccgc ctacatacct cgctctgcta atcctgttac cagtggctgc tgccagtggc 300 gataagtcgt gtcttaccgg gttggactca agacgatagt taccggataa ggcgcagcgg 360 tcgggctgaa cggggggttc gtgcacacag cccagcttgg agcgaacgac ctacaccgaa 420 ctgagatacc tacagcgtga gctatgagaa agcgccacgc ttcccgaagg gagaaaggcg 480 gacaggtatc cggtaagcgg cagggtcgga acaggagagc gcacgaggga gcttccaggg 540 ggaaacgcct ggtatcttta tagtcctgtc gggtttcgcc acctctgact tgagcgtcga 600 ttttgtgat gctcgtcagg ggggcggagc ctatggaaaa acgccagcaa cgcggccttt 660 ttacggttcc tggccttttg ctggcttt gctcacatgt tctttcctgc ggtacccaga 720 tccaattccc gctttgactg cctgaaatct ccatcgccta caatgatgac atttggattt 780 ggttgactca tgttggtatt gtgaaataga cgcagatcgg gaacactgaa aaatacacag 840 attattca ttcagaagc gatgagaga ctgcgctaag cattaatgag attattttg 900 agcattcgtc aatcaatacc aaacaagaca aacggtagc cgactttgg aagttcttt 960 ttgaccaact ggccgttagc atttcaacga accaactta gttcatcttg gatgagatca 1020 cgctttgtc atattaggtt ccagacagc gtttaactg tcagttttgg gccattttggg 1080 gaacatgaaa ctatttgacc ccacactcag aaagccctca tctggagtga tgttcggtg 1140 taatgcggag cttgttgcat tcggaaataa aaacatga acctcgccag gggggccagg 1200 atagacaggc taataagtc atggtgttag tagcctaata gaggaatttg gaataatga 1260 cccttgtgac tgacacttg ggagtcccta tttacttag tctcatatcg catgaaactt 1320 ttgataaatt attttctgat aggaattttt catcagatat tatcatcgcg gcttacgtaa 1380 taacaaaaaa aattgatgga gtctatacta ggctaacata aactaagtta ttaattaaac 1440 aaaacaaaac gtactagcat tactgtcata tataagggct cctaactaaa actgtaaaga 1500 cttcccgtaa aattatcatt ctaattctga caatgtgcat ggcctcctaa actcttgacc 1560 tctctcatgc agccacttat tggaaaccca cttattaccg actaagacgg gacaagcagc 1620 atgtctagtg ctgtaatcac cttctccaga tgcaaacaga ttgtaccaaa atacggccgt 1680 gcccttttta ggccaaacag aagcacctac ctcagggaaa actgtggctc ctccagcaag 1740 cacatcggac atatagaaca accacgttgc gattctattt ccagtaccta gctccttaaa 1800 agcatcaggc tcgtcctttc tggcgaaatc aaagtggggt tcatactgac cgcccacacc 1860 atagttggca acttgtagtt cctcagcagt gcttacgtca agaccagtca aatcttgaat 1920 acgcatattg atacggctga ccacgggatt ctcgtaaccg gacaaccatg ctgatttaga 1980 gacacgatat tgtgcggtag tcaattttcc agtctcaggg tcatggacgg tagccctact 2040 caatcttggt ttggccaagt ctttcacaac ctctatttct gcatcggaga tgatgtcatg 2100 aaaacgaatg attctaggct tgtcccattc atcttcctgt ttcgctggag caagaatgaa 2160 ttttgggtta cggttcccat catgatatct acagaacagc tttttctgtc tccttggagt 2220 catcttgata ccctctcctc tacacagcat ttcatacttt tgtctctctg ggaggtagtc 2280 aacagctgca cctttttttt tcagagtggt cttttgatcg gattggtcat cggacgagga 2340 cttatttgcg tccttttcct tagccatgat gtattcaaag tatttcagat taccgttagc 2400 tctttgatgc tccgggtcca gctccaacaa ctttttagtt aaaagtagag ctttgtccag 2460 atcaccttgc tggtaaacag cgtatgataa gtaatccaaa actgaaacct tatcaacggt 2520 agaaacttca ccttcgtcca actgacgcag agcttgctcc atccataatt ctgtgtgata 2580 gtagtcggct tctgtatatg cgacttttcc caattcaaaa caatcttcca cagtgaggaa 2640 ggacttatgc ttcacaccag gtaaatcacc cttcgatatc gtgtcggtgt ccaaattgta 2700 tgtgtcctgc aatcgcaaca aagcttttgc tgctcctact tggtcctcat cgtttggaaa 2760 gtattgtctt tgaattgtta agttagaat gatccatca ctcatatctt taagtaccaa 2820 gttttccaat tctgaccact ctgtattaag tcttcatc agcttgaag cattcactgg 2880 gtgacccaca aaaccctcag gatctttgt tgcagtacta gtcaatctat cgagtttctc 2940 dress ttgattgct sunglasses ctcttcat tgattgct ctgattgct 3000 tgtaactagg tctttttctg tgtgaatcaa atcagtcatc tgtcctatag aagtgaagaa 3060 gcctgggtga gccagtgact gtggcaacaa ataccaacg actaggatat accaaatcat 3120 gcggcctgtt gtagttttaa tatagtttga gtagagatg gaactcagaa cgaggaatt 3180 atcaccagtt tatatattct gaggaaaggg tgtgtcctaa attggacagt cacgatggca 3240 aaacgctc agccaatcag aatgcaggag ccataaattg ttgttatt gctgcaat 3300 ttatgtgggt tcacattcca ctgaatggtt ttcactgtag aattggtgtc ctagttgtta 3360 tgtttcgaga tgttttcaag aaaaactaaa atgcacaac tgaccaataa tgtgccgtcg 3420 cgcttgtac aaacgtcagg attgccacca cttttcgc actctgtac aaagttcgc 3480 acttccact cgtatgtaac gaaaacaga gcagtctatc gaacgaga caaattagcg 3540 cgtactgtcc cattccataa ggtatcatag gaacgagag tcctcccccc atcacgtata 3600 tataaacaca ctgatatccc acatccgctt gtcaccaaac taatacatcc agttcaagtt 3660 acctaaaaa atcaaagcat gagattccca tctattttca ccgctgtctt gttcgctgcc 3720 tcctctgcat tggctgcacc cgatgaggaa gatcatgttt tagtattgca taaaggaat 3780 ttcgatgaag ctttggccgc tcaacaatat ctgctcgtcg agttttacgc tccctggtgc 3840 ggtcattgta aggcccttgc accagagtac gccaaggcag ctggtaagtt aaaggccgaa 3900 ggttcagaga tcagattagc aaaagttgat gctacagaag agtccgatct tgctcaacaa 3960 tacggggttc gaggataccc aacaattaag ttttcaaaa atggtgatac tgcttcccca 4020 aaagaatata ctgctggtag agaggcagac gacatagtca actggctcaa aaagagaacg 4080 ggcccagctg cgtctacatt aagcgacgga gcagcagccg aagctcttgt ggaatctagt 4140 gaagttgctg taatcggttt ctttaaggac atggaatctg attcagctaa acagttcctt 4200 ttagcagctg aagcaatcga tgacatccct ttcggaatca cctcaaatag tgacgtgttc 4260 agcaagtacc aacttgacaa agatggagtg gtcttgttca aaaagtttga cgaaggcaga 4320 aacaatttcg agggtgaggt tacaaaggag aaactgcttg atttcattaa acataaccaa 4380 ctacccttag ttatcgaatt cactgaacaa actgctccta agattttcgg tggagaaatc 4440 aaaacacata tcttgttgtt tttgccaaag tccgtatcgg attatgaagg taaactctcc 4500 aatttcaaaa aggccgctga gagctttaag ggcaagattt tgttcatctt tattgactca 4560 gaccacacag acaatcagag gattttggag tttttcggtt tgaaaaagga ggaatgtcca 4620 gcagtccgtt tgatcacctt ggaggaggag atgaccaaat acaaaccaga gtcggatgag 4680 ttgactgccg agaagataac agaattttgt cacagatttc tggaaggtaa gatcaagcct 4740 catcttatgt ctcaagagtt gcctgatgac tgggataagc aaccagttaa agtattggtg 4800 ggtaaaaact ttgaggaagt ggccttcgac gagaaaaaaa atgtctttgt tgaattctat 4860 gctccgtggt gtggtcactg taagcagctg gcaccaattt gggataaact gggtgaaact 4920 4980. 4980. cgaaagatc acgaaaacat tgttattgca aagatggaca gtactgctaa cgaagtggag gctgtgaaag ttcactcctt ccctacgctg aagttctttc ctgcatctgc tgacagaact gttatcgact ataatggaga gaggacattg gatggtttta aaaagtttct tgaatccgga ggtcaagacg gagctggtga cgacgatgat ttggaagatc tggaggaggc tgaggaacct gatcttgagg aggatgacga ccagaaggca gtcaaagatg aactgtgata aggggtcaag aggatgtcag aatgccattt gcctgagaga tgcaggcttc atttttgata cttttttatt 5280 tgtaacctat atagtatagg attttttttg tcattttgtt tcttctcgta cgagcttgct 5340 cctgatcagc ctatctcgca gcagatgaat atcttgtggt aggggtttgg gaaaatcatt cgagtttgat gtttttcttg gtatttccca ctcctcttca gagtacagaa gattagtga gaccttcgtt tgtgcggatc cttcagtaat gtcttgtttc ttttgttgca gtggtgagcc 5520 attttgactt cgtgaaagtt tctttagaat agttgtttcc agaggccaaa cattccaccc gtagtaaagt gcaagcgtag gaagaccaag actggcataa atcaggtata agtgtcgagc actggcaggt gatcttctga aagtttctac tagcagataa gatccagtag tcatgcatat 5700 ggcaacaatg taccgtgtgg atctaagaac gcgtcctact aaccttcgca ttcgttggtc 5760 cagtttgttg ttatcgatca acgtgacaag gttgtcgatt ccgcgtaagc atgcataccc 5820 aaggacgcct gttgcaattc caagtgagcc agttccaaca atctttgtaa tattagagca 5880 cttcattgtg ttgcgcttga aagtaaaatg cgaacaaatt aagagataat ctcgaaaccg 5940 cgacttcaaa cgccaatatg atgtgcggca cacaataagc gttcatatcc gctgggtgac 6000 tttctcgctt taaaaaatta tccgaaaaaa ttttctagag tgttgacact ttatacttcc 6060 ggctcgtata atacgacaag gtgtaaggag gactaaacca tgggtaaaa gcctgaactc 6120 accgcgacgt ctgtcgagaa gtttctgatc gaaaagttcg acagcgtctc cgacctgatg 6180 cagctctcgg agggcgaaga atctcgtgct ttcagcttcg atgtaggagg gcgtggatat 6240 gtcctgcggg taaatagctg cgccgatggt ttctacaaag atcgttatgt ttatcggcac 6300 tttgcatcgg ccgcgctccc gattccggaa gtgcttgaca ttggggaatt cagcgagagc 6360 ctgacctatt gcatctcccg ccgtgcacag ggtgtcacgt tgcaagacct gcctgaaacc 6420 gaactgcccg ctgttctgca gccggtcgcg gaggccatgg atgcgatcgc tgcggccgat 6480 cttagccaga cgagcgggtt cggcccattc ggaccgcaag gaatcggtca atacactaca 6540 tggcgtgatt tcatatgcgc gattgctgat ccccatgtgt atcactggca aactgtgatg 6600 gacgacaccg tcagtgcgtc cgtcgcgcag gctctcgatg agctgatgct ttgggccgag 6660 gactgccccg aagtccggca cctcgtgcac gcggatttcg gctccaacaa tgtcctgacg 6720 gacaatggcc gcataacagc ggtcattgac tggagcgagg cgatgttcgg ggattcccaa 6780 tacgaggtcg ccaacatctt cttctggagg ccgtggttgg cttgtatgga gcagcagacg 6840 cgctacttcg agcggaggca tccggagctt gcaggatcgc cgcggctccg ggcgtatatg 6900 ctccgcattg gtcttgacca actctatcag agcttggttg acggcaattt cgatgatgca 6960 gcttgggcgc agggtcgatg cgacgcaatc gtccgatccg gagccgggac tgtcgggcgt 7020 acacaaatcg cccgcagaag cgcggccgtc tggaccgatg gctgtgtaga agtactcgcc 7080 gatagtggaa accgacgccc cagcactcgt ccgagggcaa aggaataaca attgacacct 7140 tacgattatt tagagagtat ttattagttt tattgtatgt atacggatgt tttattatct 7200 atttatgccc ttatattctg taactatcca aaagtcctat cttatcaagc cagcaatcta 7260 tgtccgcgaa cgtcaactaa aaataagctt tttatgctct tctctctttt tttcccttcg 7320 gtataattat accttgcatc cacagattct cctgccaaat tttgcataat cctttacaac 7380 atggctatat gggagcactt agcgccctcc aaaacccata ttgcctacgc atgtataggt 7440 gttttttcca caatattttc tctgtgctct ctttttatta aagagaagct ctatatcgga 7500 gaagcttctg tggccgttat attcggcctt atcgtgggac cacattgcct gaattggttt 7560 gccccggaag attggggaaa cttggatctg attaccttag ctgca 7605 <210> 14 <211> 7377 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotide <400> 14 ggatccttca gtaatgtctt gtttcttttg ttgcagtggt gagccatttt gacttcgtga 60 aagtttcttt aagatagttg tttccagagg ccaaacattc cacccgtagt aaagtgcaag 180. cgtaggaaga ccaagactgg cataaatcag gtataagtgt cgagcactgg caggtgatct tctgaaagtt tctactagca gataagatcc agtagtcatg catatggcaa caatgtaccg tgtggatcta agaacgcgtc ctactaacct tcgcattcgt tggtccagtt tgttgttatc 300 gatcaacgtg acaaggttgt cgattccgcg tagcatgca tacccaagga cgcctgttgc 360 aattccaagt gagccagttc caacaatctt tgtaatatta gagcacttca ttgtgttgcg cttgaaagta aaatgcgaac aaattaagag fatherctcga aaccgcgact tcaaacgcca atatgatgtg cggcacacaa tagcgttca tatccgctgg gtgactttct cgctttaaaa aattatccga aaaaatttc taggtgttg ttactttata cttccggctc gtataatacg acaaggtgta aggaggacta aaccatggct aaactcacct ctgctgttcc agtcctgact 660 gctcgtgatg ttgctggtgc tgttgagttc tggactgata ggctcggttt ctcccgtgac 720 ttcgtagagg acgactttgc cggtgttgta cgtgacgacg ttaccctgtt catctccgca 780 gttcaggacc aggttgtgcc agacaacact ctggcatggg tatgggttcg tggtctggac 840 gaactgtacg ctgagtggtc tgaggtcgtg tctaccaact tccgtgatgc atctggtcca 900 gctatgaccg agatcggtga acagccctgg ggtcgtgagt ttgcactgcg tgatccagct 960 ggtaactgcg tgcatttcgt cgcagaagag caggactaac aattgacacc ttacgattat 1020 ttagagagta tttattagtt ttattgtatg tatacggatg ttttattatc tatttatgcc 1080 cttatattct gtaactatcc aaaagtccta tcttatcaag ccagcaatct atgtccgcga 1140 acgtcaacta aaaataagct ttttatgctc ttctctcttt ttttcccttc ggtataatta 1200 taccttgcat ccacagattc tcctgccaaa ttttgcataa tcctttacaa catggctata 1260 tgggagcact tagcgccctc caaaacccat attgcctacg catgtatagg tgttttttcc 1320 acaatatttt ctctgtgctc tctttttatt aaagagaagc tctatatcgg agaagcttct 1380 gtggccgtta tattcggcct tatcgtggga ccacattgcc tgaattggtt tgccccggaa 1440 gattggggaa acttggatct gattacctta gctgcagaaa agggtaccac tgagcgtcag 1500 accccgtaga aaagatcaaa ggatcttctt gagatcctttt ttttctgcgc gtaatctgct gcttgcaaac aaaaaaacca cggctaccag cggtggtttg tttgccggat caagagctac caactctttt tccgaaggta actggcttca gcagagcgca throw actgttcttc tagtgtagcc gtagttaggc caccacttca agaactctgt agcaccgcct acatacctcg ctctgctaat cctgttacca gtggctgctg ccagtggcga tagtcgtgt cttaccgggt tggacccaag acgatagtta ccggataagg cgcagcggtc gggctgaacg gggggttcgt gcacacagcc cagcttggag cgaacgacct acaccgaact gagataccta cagcgtgagc tatgagaaag cgccacgctt cccgaagggga gaaggcgga caggtatccg gtaagcggca gggtcggac aggagagcgc acgagggagc ttccaggggg aaacgcctgg tatcttata gtcctgtcgg gtttcgccac ctctgacttg agcgtcgatt tttgtgatgc tcgtcagggg 2100 ggcggagcct atggaaaaac gccagcaacg cggcctttttt acggttcctg gccttttgct ggccttttgc tcacatgttt cagaagcgat agagagactg cgctaagcat taatgagatt atttttgagc attcgtcaat caataccaaa caagacaaac ggtatgccga cttttggaag 2280 tttctttttg accaactggc cgttagcatt tcaacgaacc aaacttagtt catcttggat 2340 gagatcacgc ttttgtcata ttaggttcca agacagcgtt taaactgtca gttttgggcc 2400 atttggggaa catgaaacta tttgacccca cactcagaaa gccctcatct ggagtgatgt 2460 tcgggtgtaa tgcggagctt gttgcattcg gaaataaaca aacatgaacc tcgccagggg 2520 ggccaggata gacaggctaa taaagtcatg gtgttagtag cctaatagaa ggaattggaa 2580 ataatgtatc taaacgcaaa ctccgagctg gaaaaatgtt accggcgatg cgcggacaat 2640 ttagaggcgg cgatcaagaa acacctgctg ggcgagcagt ctggagcaca gtcttcgatg 2700 ggcccgagat cccaccgcgt tcctgggtac cgggacgtga ggcagcgcga catccatcaa 2760 atataccagg cgccaaccga gtctctcgga aaacagcttc tggatatctt ccgctggcgg 2820 cgcaacgacg aataatagtc cctggaggtg acggaatata tatgtgtgga gggtaaatct 2880 gacagggtgt agcaaaggta atattttcct aaaacatgca atcggctgcc ccgcaacggg 2940 aaaaagaatg actttggcac tcttcaccag agtggggtgt cccgctcgtg tgtgcaaata 3000 ggctcccact ggtcaccccg gattttgcag aaaaacagca agttccgggg tgtctcactg 3060 gtgtccgcca ataagaggag ccggcaggca cggagtctac atcaagctgt ctccgataca 3120 ctcgactacc atccgggtct ctcagagagg ggaatggcac tataaatacc gcctccttgc 3180 gctctctgcc ttcatcaatc aaatcatgtt ctctccaatt ttgtccttgg aaattatttt 3240 agctttggct actttgcaat ctgtcttcgc tcaacaggaa gcagtagatg gtggttgctc 3300 acatttaggt caatcttacg cagatagaga tgtatggaaa cctgaaccat gtcaaatttg 3360 cgtgtgtgac tcaggttcag tgctctgcga cgatatcata tgtgacgacc aggaattgga 3420 ctgtccaaac ccagagatac cattcggtga atgttgtgct gtttgtccac agccaccaac 3480 tgctcctaca agacctccaa acggtcaagg tccacaaggt cctaaaggtg atccgggtcc 3540 acctggtatt cctggtagaa atggtgaccc tggacctccc ggttccccag gtagcccagg 3600 atcacctggg cctcctggaa tatgtgaatc ctgcccaact ggtggtcaga actatagccc 3660 acaatacgag gcctacgacg tcaaatctgg tgttgctgga ggaggtattg caggctaccc 3720 tggtcccgca gggcccccag gtccgccggg tccgcccgga acatcaggtc atcccggagc 3780 ccctggtgca ccaggttatc agggaccgcc cggagagcct ggacaagctg gtcccgctgg 3840 accccctggt ccaccaggtg ctattggacc aagtggtcct gccggaaaag acggtgaatc 3900 cggtagacct ggtagacccg gcgaaagggg tttcccaggt cctcccggaa tgaagggtcc 3960 agccggtatg cccggtttc ctgggatgaa gggtcacaga ggatttgatg gtagaaacgg 4020 agaagaaggc gaaaccggtg ctcccggact gaagggtgaa aacggtgtcc ctggtgagaa 4080 cggcgctcct ggacctatgg gtccacgtgg tgctccagga gaaagaggca gaccaggatt 4140 gcctggtgca gctggtgcta gaggtaacga tggtgcccgt ggttccgatg gacaacccgg 4200 gccacccggc cctccaggta ccgctggatt tcctggaagc cctggtgcta agggggaggt 4260 tggtccggct ggtagtcccg gaagtagcgg tgccccaggt caaagaggcg aaccaggccc 4320 tcagggtcac gcaggagcac ctggaccgcc tggtcctcct ggttcgaatg gttcgcctgg 4380 aggaaaaggt gaaatggggc ccgcaggaat ccccggtgcg cctggtctta ttggtgccag 4440 gggtcctcca ggcccgccag gtacaaatgg tgtacccgga cagcgaggag cagctggtga 4500 acctggtaaa aacggtgcca aaggagatcc aggtcctcgt ggagagcgtg gtgaagctgg 4560 ctctcccggt atcgccggtc caaaaggtga ggacggtaag gacggttccc ctggtgagcc 4620 aggtgcgaac ggactgccag gtgcagccgg agagcgaga gtcccaggat tcagggggacc 4680 agccggtgct aacggcttgc ctggtgaaaa agggccccct ggtgataggg gaggacccgg 4740 tccagcaggc cctcgtggag ttgctggtga gcctggacgt gacggtttac caggagggcc 4800 aggtttgagg ggtattcccg ggtcccctgg cggtcctgga tcggatggaa aaccagggcc 4860 accaggttcg cagggtgaaa caggacgtcc aggcccaccc ggctcacctg gtccaagggg 4920 tcagcctggt gtcatgggtt tccccggtcc aaagggtaat gacggagcac cgggataaaaa 4980 tggtgaacgt ggtggcccag gtggtccagg accccaaggt ccagctggaa aaaacggtga 5040 gacaggtcct caaggacctc caggacctac cggtcctagc ggagataagg gagatacggg 5100 accgccagga cctcaaggat tgcaaggttt gcctggtaca tctggccctc ccggagaaaa 5160 tggtaagcct ggagagccag gaccaaaagg cgaagctgga gccccaggta tccccggagg 5220 taagggagac tcaggtgctc cgggtgagcg tggtcctccg ggtgccggtg gtccacctgg 5280 acctagaggt ggtgccgggc cgccaggtcc tgaaggtggt aaaggtgctg ctggtccacc 5340 gggaccgcct ggctctgctg gtactcctgg cttgcaggga atgccaggag agagaggtgg 5400 acctggaggt cccggtccga agggtgataa aggggagcca ggatcatccg gtgttgacgg 5460 cgcacctggt aaagacggac caaggggacc aacgggtcca atcggaccac caggacccgc 5520 tggccagcca ggagataaag gcgagtccgg agcacccggt gttcctggta tagctggacc 5580 caggggtggt cccggtgaaa gaggtgaaca gggcccaccg ggtcccgccg gtttccctgg 5640 cgcccctggt caaaatggag aaccaggtgc aaagggcgag agaggagccc caggagaaaa 5700 gggtgaggga ggaccacccg gtgctgccgg tccagctggg ggttcaggtc ctgctggacc 5760 accaggtcca cagggcgtta aaggtgagag aggaagtcca ggtggtcctg gagctgctgg 5820 attcccaggt ggccgtggac ctcctggtcc ccctggatcg aatggtaatc ctggtccgcc 5880 aggtagttcg ggtgctcctg ggaggacgg tccacctggc cccccaggta gtaacggtgc 5940 acctggtagt ccaggtatat ccggacctaa aggagattcc ggtccaccag gcgaaagagg ggccccaggc ccacagggtc ccccggtcct ctgggtattg ctggtcttac 6060. tggtgcacgt ggactggccg gtccacccgg aatgcctgga gcaagaggtt cacctggacc 6120 acaaggtatt aaaggagaga acggtaaacc tggaccttcc ggtcaaaacg gagagcgggg acccccaggc ccccaaggtc tgccaggact agctggtacc gcaggggac caggagaga 6240 tggaatcca ggttcagacg gactacccgg tagagatggt gcaccggggg ccaagggcga caggggtgag aatggatctc ctggtgcgcc aggggcacca ggccacccag gtcccccagg 6360 tcctgtgggc cctgctggaa agtcaggtga caggggagag acaggcccgg ctggtccatc 6420. tggcgcaccc ggaccagctg gttccagagg cccacctggt ccgcaaggcc ctagaggtga 6480 6540. caagggagag actggagac gaggtgctat gggtatcaag ggtcatagg gttttccggg taatcccggc gccccaggtt ctcctggtcc agctggccat caaggtgcag tcggatcgcc 6600 cggcccagcc ggtcccaggg gccctgttgg tccatccggt cctccaggaa aggatggtgc 6660 ttctggacac ccaggaccta tcggacctcc gggtcctaga ggtaatagag gagaacgtgg 6720 atccgagggt agtcctggtc accctggtca acctggccca ccagggcctc caggtgcacc 6780 cggtccatgt tgtggtgcag gcggtgtggc tgcaattgct ggtgtgggtg ctgaaaaggc 6840 cggcggtttc gctccatatt atggtgatgg ttacattcct gaagctccta gagacggaca 6900 agcatacgtt agaaaggacg gtgagtgggt gttgctgtcc accttcttag gttctggttc 6960 tggttctgat tacaaggatg acgacgataa gggatcgtgt tgcccgggct gctgtggcaa 7020 accaatacct aaccctttac tgggccttga cagtacgtat ccgtatgatg tgccggatta 7080 tgcgcatcac catcatcacc atagatctta atcaagagga tgtcagaatg ccatttgcct 7140 gagagatgca ggcttcattt ttgatacttt tttatttgta acctatatag tataggattt 7200 tttttgtcat tttgtttctt ctcgtacgag cttgctcctg atcagcctat ctcgcagctg 7260 atgaatatct tgtggtaggg gtttgggaaa atcattcgag tttgatgttt ttcttggtat 7320 ttcccactcc tcttcagagt acagaagatt aagtgagacg ttcgtttgtg ctccgga 7377 <210> 15 <211> 951 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotide <400> 15 atgagattcc catctatttt caccgctgtc ttgttcgctg cctcctctgc attggctgcc 60 cctgttaaca ctaccactga agacgagact gctcaaattc cagctgaagc agttatcggt 120 tactctgacc ttgagggtga tttcgacgtc gctgttttgc ctttctctaa ctccactaac 180 aacggtttgt tgttcattaa caccactatc gcttccattg ctgctaagga agagggtgtc 240 tctctcgaga aaagagaggc cgaagctgtg ctgtcaaagt cctgtgtcag tcactttaga 300 aatgttggat ccttgaatag tagggatgtc aatctgaaag atgacttttc ctatgctaat 360 attgatgatc cctataacaa gcctttcgtc ctaaataacc taataaaccc taccaagtgt 420 caagagatca tgcaatttgc caatggcaag ttgtttgact cccaagtcct gagtggcacg 480 gacaagaaca tacgtaactc tcaacaaatg tggatatcca agaacaaccc tatggtaaaa 540 cccattttcg agaacatatg caggcagttt aacgtaccct ttgataatgc cgaggaccta 600 caggtcgtcc gttacttgcc taatcaatat tataatgagc atcatgactc atgctgtgac 660 tcctccaagc aatgcagtga atttatagag aggggcggtc agaggattct gaccgtttta 720 atttacctaa acaacgagtt ctcagatgga cacacgtact ttcctaattt aaaccaaaag 780 ttcaagccca agactggtga tgctttggtt ttttaccctt tagccaacaa ctctaataaa 840 tgtcacccat acagtctaca cgcaggtatg cccgtcacgt caggagagaa gtggattgct 900 aatctgtggt ttcgtgagcg taagttctcc caccaccacc accaccacta a 951 <210> 16 <211> 316 <212> PRT <213> Artificial Sequence <220> <223> Synthetic Peptide <400> 16 Met Arg Phe Pro Ser Ile Phe Thr Ala Val Leu Phe Ala Ala Ser Ser 1 5 10 15 Ala Leu Ala Ala Pro Val Asn Thr Thr Thr Glu Asp Glu Thr Ala Gln 20 25 30 Ile Pro Ala Glu Ala Val Ile Gly Tyr Ser Asp Leu Glu Gly Asp Phe 35 40 45 Asp Val Ala Val Leu Pro Phe Ser Asn Ser Thr Asn Asn Gly Leu Leu 50 55 60 Phe Ile Asn Thr Thr Ile Ala Ser Ile Ala Ala Lys Glu Glu Gly Val 65 70 75 80 Ser Leu Glu Lys Arg Glu Ala Glu Ala Val Leu Ser Lys Ser Cys Val 85 90 95 Ser His Phe Arg Asn Val Gly Ser Leu Asn Ser Arg Asp Val Asn Leu 100 105 110 Lys Asp Asp Phe Ser Tyr Ala Asn Ile Asp Asp Pro Tyr Asn Lys Pro 115 120 125 Phe Val Leu Asn Asn Leu Ile Asn Pro Thr Lys Cys Gln Glu Ile Met 130 135 140 Gln Phe Ala Asn Gly Lys Leu Phe Asp Ser Gln Val Leu Ser Gly Thr 145 150 155 160 Asp Lys Asn Ile Arg Asn Ser Gln Gln Met Trp Ile Ser Lys Asn Asn 165 170 175 Pro Met Val Lys Pro Ile Phe Glu Asn Ile Cys Arg Gln Phe Asn Val 180 185 190 Pro Phe Asp Asn Ala Glu Asp Leu Gln Val Val Arg Tyr Leu Pro Asn 195 200 205 Gln Tyr Tyr Asn Glu His His Asp Ser Cys Cys Asp Ser Ser Lys Gln 210 215 220 Cys Ser Glu Phe Ile Glu Arg Gly Gly Gln Arg Ile Leu Thr Val Leu 225 230 235 240 Ile Tyr Leu Asn Asn Glu Phe Ser Asp Gly His Thr Tyr Phe Pro Asn 245 250 255 Leu Asn Gln Lys Phe Lys Pro Lys Thr Gly Asp Ala Leu Val Phe Tyr 260 265 270 Pro Leu Ala Asn Asn Ser Asn Lys Cys His Pro Tyr Ser Leu His Ala 275 280 285 Gly Met Pro Val Thr Ser Gly Glu Lys Trp Ile Ala Asn Leu Trp Phe 290 295 300 Arg Glu Arg Lys Phe Ser His His His His His His 305 310 315 <210> 17 <211> 4029 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotide <400> 17 ggatccttca gtaatgtctt gtttcttttg ttgcagtggt gagccatttt gacttcgtga aagtttcttt aagatagttg tttccagagg ccaaacattc cacccgtagt aaagtgcaag 180. cgtaggaaga ccaagactgg cataaatcag gtataagtgt cgagcactgg caggtgatct tctgaaagtt tctactagca gataagatcc agtagtcatg catatggcaa caatgtaccg tgtggatcta agaacgcgtc ctactaacct tcgcattcgt tggtccagtt tgttgttatc 300 gatcaacgtg acaaggttgt cgattccgcg tagcatgca tacccaagga cgcctgttgc 360 aattccaagt gagccagttc caacaatctt tgtaatatta gagcacttca ttgtgttgcg cttgaaagta aaatgcgaac aaattaagag fatherctcga aaccgcgact tcaaacgcca atatgatgtg cggcacacaa tagcgttca tatccgctgg gtgactttct cgctttaaaa aattatccga aaaaatttc taggtgttg ttactttata cttccggctc gtataatacg acaaggtgta aggaggacta aaccatggct aaactcacct ctgctgttcc agtcctgact 660 gctcgtgatg ttgctggtgc tgttgagttc tggactgata ggctcggttt ctcccgtgac 720 ttcgtagagg acgactttgc cggtgttgta cgtgacgacg ttaccctgtt catctccgca 780 gttcaggacc aggttgtgcc agacaacact ctggcatggg tatgggttcg tggtctggac 840 gaactgtacg ctgagtggtc tgaggtcgtg tctaccaact tccgtgatgc atctggtcca 900 gctatgaccg agatcggtga acagccctgg ggtcgtgagt ttgcactgcg tgatccagct 960 ggtaactgcg tgcatttcgt cgcagaagag caggactaac aattgacacc ttacgattat 1020 ttagagagta tttattagtt ttattgtatg tatacggatg ttttattatc tatttatgcc 1080 cttatattct gtaactatcc aaaagtccta tcttatcaag ccagcaatct atgtccgcga 1140 acgtcaacta aaaataagct ttttatgctc ttctctcttt ttttcccttc ggtataatta 1200 taccttgcat ccacagattc tcctgccaaa ttttgcataa tcctttacaa catggctata 1260 tgggagcact tagcgccctc caaaacccat attgcctacg catgtatagg tgttttttcc 1320 acaatatttt ctctgtgctc tctttttatt aaagagaagc tctatatcgg agaagcttct 1380 gtggccgtta tattcggcct tatcgtggga ccacattgcc tgaattggtt tgccccggaa 1440 gattggggaa acttggatct gattacctta gctgcagaaa agggtaccac tgagcgtcag 1500 accccgtaga aaagatcaaa ggatcttctt gagatccttt ttttctgcgc gtaatctgct 1560 gcttgcaaac aaaaaaacca ccgctaccag cggtggtttg tttgccggat caagagctac 1620 caactctttt tccgaaggta actggcttca gcagagcgca gataccaaat actgttcttc 1680 tagtgtagcc gtagttaggc caccacttca agaactctgt agcaccgcct acatacctcg 1740 ctctgctaat cctgttacca gtggctgctg ccagtggcga taagtcgtgt cttaccgggt 1800 tggacccaag acgatagtta ccggataagg cgcagcggtc gggctgaacg gggggttcgt 1860 gcacacagcc cagcttggag cgaacgacct acaccgaact gagataccta cagcgtgagc 1920 tatgagaaag cgccacgctt cccgaaggga gaaaggcgga caggtatccg gtaagcggca 1980 gggtcggaac aggagagcgc acgagggagc ttccaggggg aaacgcctgg tatctttata 2040 gtcctgtcgg gtttcgccac ctctgacttg agcgtcgatt tttgtgatgc tcgtcagggg 2100 ggcggagcct atggaaaaac gccagcaacg cggccttttt acggttcctg gccttttgct 2160 ggccttttgc tcacatgtat ttaaataatg tatctaaacg caaactccga gctggaaaa 2220 tgttaccggc gatgcgcgga caatttagag gcggcgatca agaaacacct gctgggcgag 2280 cagtctggag cacagtcttc gatgggcccg agatcccacc gcgttcctgg gtaccgggac 2340 gtgaggcagc gcgacatcca tcaaatatac caggcgccaa ccgagtgtct cggaaaacag 2400 cttctggata tcttccgctg gcggcgcaac gacgaataat agtccctgga ggtgacggaa 2460 tatatatgtg tggagggtaa atctgacagg gtgtagcaaa ggtaatattt tcctaaaaca 2520 tgcaatcggc tgccccgcaa cgggaaaaag aatgactttg gcactcttca ccagagtggg 2580 gtgtcccgct cgtgtgtgca aataggctcc cactggtcac cccggatttt gcagaaaaac 2640 agcaagttcc ggggtgtctc actggtgtcc gccaataaga ggagccggca ggcacggagt 2700 ttacatcaag ctgtctccga tacactcgac taccatccgg gtctctcaga gaggggaatg 2760 gcactataaa taccgctcc ttgcgctctc tgccttcatc aatcaaatca tgagattccc 2820 atctattttc accgctgtct tgttcgctgc ctcctctgca ttggctccc ctgttaacac 2880 taccactgaa gacgagactg ctcaaattcc agctgaagca gttatcggtt actctgacct 2940 tgagggtgat ttcgacgtcg ctgttttgcc tttctctaac tccactaaca acggtttgtt 3000 gttcattaac accactatcg cttccattgc tgctaaggaa gagggtgtct ctctcgagaa 3060 aagagaggcc gaagctgtgc tgtcaaagtc ctgtgtcagt cactttagaa atgttggatc 3120 cttgaatagt agggatgtca atctgaaaga tgacttttcc tatgctaata ttgatgatcc 3180 ctataacaag cctttcgtcc taataacct aataaaccct accaagtgtc aagagatcat 3240 gcaatttgcc aatggcaagt tgtttgactc ccaagtcctg agtggcacgg aaagaacat 3300 acgtaactct caacaaatgt ggatatccaa gaacaaccct atggtaaaac ccatttttcga 3360 gaacatatgc aggcagttta acgtaccctt tgataatgcc gaggacctac aggtcgtccg 3420 ttacttgcct aatcaatatt ataatgagca tcatgactca tgctgtgact cctccaagca 3480 3540 caacgagttc tcagatggac acacgtactt tcctaattta aaccaaaagt tcaagcccaa 3600 gactggtgat gctttggttt tttacccttt agccaacaac tctaataaat gtcacccata 3660 cagtctacac gcaggtatgc ccgtcacgtc aggagagaag tggattgcta atctgtggtt 3720 tcgtgagcgt aagttctccc accaccacca ccaccactaa taatcaagag gatgtcagaa 3780 tgccatttgc ctgagagatg caggcttcat ttttgatact tttttatttg taacctatat 3840 agtataggat tttttttgtc attttgtttc ttctcgtacg agcttgctcc tgatcagcct 3900 atctcgcagc tgatgaatat cttgtggtag gggtttggga aaatcattcg agtttgatgt 3960 ttttcttggt atttcccact cctcttcaga gtacagaaga ttaagtgaga cgttcgtttg 4020 tgctccgga 4029 <210> 18 <211> 50 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotide <400> 18 ctctgccttc atcaatcaaa tcatgagatt cccatctatt ttcaccgctg 50 <210> 19 <211> 25 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotide <400> 19 agcttcggcc tctcttttct cgaga 25 <210> 20 <211> 55 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotides <400> 20 tctcgagaaa agagaggccg aagctgtgct gtcaaagtcc tgtgtcagtc acttt 55 <210> twenty one <211> 60 <212> DNA <213> Artificial sequence <220> <223> Synthetic nucleotides <400> twenty one gcaaatggca ttctgacatc ctcttgatta gtggtggtgg tggtggtggg agaacttacg 60
Claims
1. A fusion protein consisting of: a prolyl 4-hydroxylase alpha subunit-1; and a prolyl 4-hydroxylase beta subunit, wherein the prolyl 4-hydroxylase alpha subunit-1 is at the N-terminus of the fusion protein and the prolyl 4-hydroxylase beta subunit is at the C-terminus of the fusion protein, and wherein the prolyl 4-hydroxylase alpha subunit-1 is encoded by SEQ ID NO: 1 and the prolyl 4-hydroxylase beta subunit is encoded by SEQ ID NO:
2.
2. A microorganism comprising: the fusion protein of claim 1, wherein the microorganism is Pichia pastoris.
3. The microorganism of claim 2, further comprising a second protein to be hydroxylated.
4. The microorganism of claim 3, wherein the second protein is recombinant collagen.
5. An in vitro method for hydroxylating a protein, the in vitro method comprising: lysing a microorganism containing a protein to be hydroxylated to produce a lysate; adding a specific concentration of the fusion protein of claim 1 to the lysate; and incubating the lysate and the fusion protein under reaction conditions that promote the protein to be hydroxylated being hydroxylated by the fusion protein.
6. The method of claim 5, wherein the concentration of the fusion protein is in the range of 0.05 μΜ to 5 μΜ based on 1 μΜ of the protein to be hydroxylated.
7. The method of claim 5, wherein the hydroxylation is performed at a pH in the range of 5 to 12.
8. The method of claim 5, wherein the hydroxylation is performed at a temperature in the range of 16 °C to 40 °C.
9. A method for preparing a hydroxylated protein, the method comprising: growing the microorganism of claim 3 in a culture medium for a time sufficient to hydroxylate the second protein.
10. The method of claim 9, wherein the microorganism is grown for 50 hours to 72 hours.
Citation Information
Patent Citations
Method for preparing modified collagen outside a host cell
US7932053B2
Method for expressing recombinant human prolyl-4-hydroxylase (rhP4H) in pichia pastoris system and application
CN110172470A
Method of detecting effect of controlling synoviolin activity
US20070134720A1
Recombinant spider silk-reinforced collagen proteins produced in plants and the use thereof
WO2023194333A1