Bacterial host strain
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ALDEVRON LLC
- Filing Date
- 2025-11-20
- Publication Date
- 2026-06-03
Smart Images

Figure 00000133_0000 
Figure 00000133_0001 
Figure 00000133_0002
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims priority to U.S. Provisional Patent Application No. 62 / 988,223, entitled “Bacterial Host Strains,” filed on 11 March 2020, the entire contents of which are incorporated herein by reference.
[0002] Sequence List This application includes a sequence listing, which is submitted electronically in ASCII format and is incorporated herein by reference in its entirety. The ASCII copy, created on March 11, 2021, is named 85535-334987_SL.txt and has a size of 112,796 bytes.
[0003] Built-in by reference WO2008 / 153733, WO2014 / 035457, and WO2019 / 183248 are incorporated herein by reference in their entirety. Furthermore, all publications, patents, and patent application publications referenced herein are incorporated herein by reference in their entirety. [Background technology]
[0004] Escherichia coli (E. coli) plasmids have long been a vital source of recombinant DNA molecules used by researchers and industry. Today, plasmid DNA is becoming increasingly important as next-generation biotechnology products (e.g., gene therapies and DNA vaccines) progress to clinical trials and eventually enter the pharmaceutical market. Plasmid DNA vaccines can be applied as preventive vaccines for viral, bacterial, or parasitic diseases, as immunizers for the preparation of high-titer immunoglobulin products, as therapeutic vaccines for infectious diseases, or as cancer vaccines. Plasmids are also used in gene therapy or gene replacement applications, where the desired gene product is expressed from the plasmid after administration to the patient. Plasmids are also used in non-viral transposon vectors (e.g., Sleeping Beauty, PiggyBac, TCBuster, etc.) for gene therapy or gene replacement applications, where the desired gene product is expressed from the genome after transposition and integration from the plasmid. Plasmids are also used in gene editing (e.g., homologous recombination repair (HDR) / CRISPR-Cas9) and as non-viral vectors for gene therapy or gene replacement applications, where the desired gene product is expressed from the genome after excision from the plasmid and genome integration. Plasmids are also used in viral vectors (e.g., AAV, lentivirus, retroviral vectors) for gene therapy or gene replacement applications, where the desired gene product is packaged into transduction virus particles after transfection of the producing cell line and then expressed from the virus in target cells after viral introduction.
[0005] Non-viral and viral vector plasmids typically contain origins of replication derived from pMB1, ColE1, or pBR322. Common high-copy-number derivatives have mutations that affect copy number regulation, such as ROP (primer gene repressor) deletions and second-site mutations that increase copy number (e.g., a G-to-A point mutation in pMB1 pUC, or ColE1 pMM1). Selective plasmid amplification with pUC and pMM1 origins of replication can be induced using higher temperatures (42°C).
[0006] WO2014 / 035457 discloses a miniaturized vector (Nanoplasmid®) that utilizes RNA-OUT antibiotic-free selection to replace a large 1000 bp pUC origin with a novel 300 bp R6K origin. Reducing the spacer region by linking the 5' and 3' ends of the transgene expression cassette to <500 bp with the R6K origin-RNA-OUT backbone improves expression levels compared to conventional minicircle DNA vectors.
[0007] U.S. Patent No. 7,943,377, which is incorporated herein by reference in its entirety, describes a method for fed-batch fermentation in which plasmid-containing E. coli cells are grown at a reduced temperature during a portion of the fed-batch phase, during which the growth rate is limited, followed by a temperature increase shift, during which the cells are continuously grown at a high temperature to accumulate plasmids. The temperature shift at a limited growth rate improved plasmid yield and purity. This fermentation process is referred herein to as the HyperGRO fermentation process. Other fermentation processes for plasmid production are described in Carnes AE2005 BioProcess Intl 3:36-44, which is incorporated herein by reference in its entirety.
[0008] WO2014 / 035457 also discloses a host strain for R6K origin vector production in the HyperGRO fermentation process.
[0009] Along with Schnodt et al., (2016) Mol Ther-Nucleic Acids 5 e355, Chadeuf et al., (2005) Molecular Therapy 12:744-53 and Gray, 2017. WO2017 / 066579 teach that AAV helper plasmid antibiotic resistance markers are packaged on viral particles, indicating the need to remove antibiotic markers from AAV helper plasmids and AAV vectors. Antibiotic-free Nanoplasmid® vectors disclosed in WO2014 / 035457 do not have antibiotic marker transcription.
[0010] Viral vectors such as AAV contain palindromic inverted terminal repeat (ITR) DNA sequences at their ends.
[0011] Paraindrome and reverse repeats are inherently unstable in high-yield E. coli production hosts such as DH1, DH5α, JM107, JM108, JM109, and XL1Blue.
[0012] Growth of AAV ITR-containing vectors is recommended in the multiply mutant sbcC knockout cell line SURE (a recB derivative of SRB) or SURE2.
[0013] The SURE cell line has the following genotype: F'[proAB + lacI q lacZΔM15 Tn10(Tet R ]endA1 glnV44 thi-1 gyrA96 relA1 lac recB recJ sbcC umuC::Tn5 Kan R uvrC e14 - (mcrA - )Δ(mcrCB-hsdSMR-mrr)171(where the SURE stabilizing mutations are recB recJ umuC uvrC - (mcrA - (Includes sbcC in combination with mcrBC-hsd-mrr).
[0014] The SRB cell line has the following genotype: F’[proAB + lacI q lacZΔM15 endA1 glnV44 thi-1 gyrA96 relA1 lac recJ sbcC umuC::Tn5(Kan R uvrC e14 - (mcrA - )Δ(mcrCB-hsdSMR-mrr)171 (where the SRB stabilization mutation includes sbcC in combination with recJ umuC uvrC - (mcrA - )mcrBC-hsd-mrr).
[0015] The SURE2 cell line has the following genotype: endA1 glnV44 thi-1 gyrA96 relA1 lac recB recJ sbcC umuC::Tn5 Kan R uvrC e14-Δ(mcrCB-hsdSMR-mrr)171 F’[proAB + lacI q lacZΔM15 Tn10(Tet R )Amy Cm R (where the SURE2 stabilization mutation includes sbcC in combination with recB recJ uvrC - (mcrA - )mcrBC-hsd-mrr).<s
[0016] SbcCD is a nuclease that cleaves palindromic DNA sequences and contributes to palindromic instability in E. coli (Chalker AF, Leach DR, Lloyd RG. 1988 Gene 71:201-5). Palindromes such as shRNA or AAV ITR are more stable in SbcC knockout strains such as SURE cells than DH5α, as indicated in Gray SJ, Choi, VW, Asokan, A, Haberman RA, McCown TJ, Samulski RJ (2011) Curr Protoc Neurosci Chapter 4: Unit 4.17, which states that "AAV ITR is unstable in E. coli, and plasmids that lose ITR have a replication advantage in transformed cells." For these reasons, bacteria containing ITR plasmids should not be grown for more than 12-14 hours, and any recovered plasmids should be evaluated for ITR retention…DH10B competent cells (or other equivalent high-efficiency strains) can be used to transform the ligation reaction for ITR-containing plasmid cloning. After screening clones positive for ITR integrity, good clones should be transformed into SURE or SURE2 cells (Agilent Technologies) for plasmid and glycerol stock production. SURE cells are engineered to maintain an irregular DNA structure but have lower transformation efficiency compared to DH10B. Furthermore, Siew SM, 2014, describes a recombinant AAV-mediated gene therapy approach for treating progressive familial intrahepatic cholestasis type 3. The University of Sydney paper, uploaded on December 3, 2014, teaches that "SURE2 cells are a commonly used sbcC mutant strain for propagating plasmids containing the palindromic AAV ITR." Therefore, it is generally understood that SURE or SURE2 sbcC mutant strains are preferred for propagating plasmids containing the palindrom AAV ITR.
[0017] However, SURE and SURE2 cell lines have limitations. For example, SURE and SURE2 are kan R Therefore, it cannot be used to generate kanamycin-resistant plasmids (as opposed to ampicillin-resistant plasmids) which are typically used in cGMP production. Furthermore, the art teaches that sbcC knockout stabilization of palindromic sequences requires further mutations in other genes such as recB, recJ, uvrC, mcrA, or mcrBC-hsd-mrr. Doherty JP, Lindeman R, Trent RJ, Graham MW, Woodcock DM. 1993. Gene 124:29-35 reports that not all palindromic sequences are stabilized in SURE (or related SRB cell lines). They recommended that an additional mutation (recC) is necessary for palindromic sequence stabilization, stating: "However, while palindromic sequence-containing phages were plated with reasonable efficiency on SURE(recB sbcC recJ umuC uvrC) and SRB(sbcC recJ umuC uvrC), the majority of phages recovered from these strains no longer required the sbcC host for subsequent plating." These two strains also resulted in poorer titers in low-yield phage clones from the human Prader-Willi chromosome region. The optimal phage host appears to be a combination of mcrA delta (mcrBC-hsd-mrr) and sbcC plus a recBC or recD mutation."
[0018] In line with this, other SbcC host strains also exhibit similar behavior, for example, PMC103:mcrA Δ(mcrBC-hsdRMS-mrr)102 recD sbcC (where the PMC103 stabilizing mutation is recD(mcrA-)mcrBC-hsd - (including sbcC combined with mrr), and PMC107:mcrAΔ(mcrBC-hsdRMS-mrr)102 recB21 recC22 recJ154 sbcB15 sbcC201 (where PMC107 stabilizing mutations are recB recJ sbcB(mcrA- (Including sbcC combined with mcrBC-hsd-mrr.)
[0019] Therefore, the art teaches that sbcC knockout stabilization of the palindrom requires mutations in sbcB, recB, recD, and recJ, and possibly further mutations in uvrC, mcrA, and / or mcrBC-hsd-mrr. This teaches, apart from the application of sbcC knockout, to improve palindromic stability in standard E. coli plasmid-producing strains such as DH1, DH5α, JM107, JM108, JM109, and XL1Blue that do not contain these additional mutations.
[0020] For example, the genotypes of some standard E. coli plasmid-producing strains are as follows: DH1:F - λ - endA1 recA1 relA1 gyrA96 thi-1 glnV44 hsdR17(r K - m K - ) DH5α:F-φ80lacZΔM15 Δ(lacZYA-argF)U169 recA1 endA1 hsdR17(r k -,m k +)gal-phoA supE44 λ-thi-1 gyrA96 relA1 JM107:endA1 glnV44 thi-1 relA1 gyrA96 Δ(lac-proAB)[F' traD36 proAB + lacI q lacZΔM15]hsdR17(R K - m K + )λ - JM108:endA1 recA1 gyrA96 thi-1 relA1 glnV44 Δ(lac-proAB)hsdR17(r K - mK + ) JM109:endA1 glnV44 thi-1 relA1 gyrA96 recA1 mcrB + Δ(lac-proAB)e14-[F' traD36 proAB + lacI q lacZΔM15]hsdR17(r K - m K + ) MG1655 K-12 F - λ - ilvG - RFB-50 RPH-1 XL1Blue:endA1 gyrA96(nal R )thi-1 recA1 relA1 lac glnV44 F'[::Tn10 proAB + lacI q Δ(lacZ)M15]hsdR17(r K - m K + )
[0021] Standard E. coli plasmid-producing strains are endA, recA. However, standard producing strains do not contain any of the necessary mutations in sbcB, recB recD, and recJ, and in some cases, they do not contain uvrC, mcrA, or mcrBC-hsd-mrr. Therefore, knockout of sbcC is not expected to effectively stabilize palindromes or reverse repeats in the absence of these additional mutations.
[0022] However, the presence of multiple mutations in the SURE and SURE2 cell lines reduces cell line viability and productivity in the E. coli fermentation plasmid production process. For example, Table 1 summarizes the HyperGRO fermentation plasmid yield and quality in SURE2 or XL1Blue (exemplary high-yield E. coli production hosts). All three plasmids were prone to multimerization and produced in low yields in SURE2, but were produced in high yields (2-4x) and of high quality (low multimerization) in XL1Blue. [Table 1]
[0023] Reduced viability and productivity are common characteristics of polymutant "stabilized hosts," such as Stbl2, Stbl3, and Stbl4, which are used to stabilize vectors containing direct repeats, such as lentiviral vectors, but do not contain SbcC knockouts. The genotypes of Stbl2, Stbl3, and Stbl4 are shown below. Stbl2:F-endA1 glnV44 thi-1 recA1 gyrA96 relA1 Δ(lac-proAB)mcrA Δ(mcrBC-hsdRMS-mrr)λ - Stbl2 stabilizing mutation = mcrA Δ(mcrBC-hsdRMS-mrr) (Trinh, T., Jessee, J., Bloom, FR, and Hirsch, V. (1994) FOCUS 16,78.) Stbl3:F-mcrB mrr hsdS20(rB-,mB-)recA13 supE44 ara-14 galK2 lacY1 proA2 rpsL20(Strr)xyl-5 -leu mtl-1 Stbl3 stabilizing mutation = mcrBC - mrr Stbl4:endA1 glnV44 thi-1 recA1 gyrA96 relA1 Δ(lac-proAB)mcrA Δ(mcrBC-hsdRMS-mrr)λ - gal F'[proAB + lackIq [lacZΔM15 Tn10] Stbl4 stabilizing mutation = mcrA Δ(mcrBC-hsdRMS-mrr)
[0024] Therefore, there is a need for high-yield E. coli strains for the high-yield production of palindrome and reverse repeat-containing vectors without ITR deletions or rearrangements that are not plagued by low stability or low viability. [Overview of the project]
[0025] This disclosure relates to host bacterial strains, methods for producing such host bacterial strains, and methods for improving plasmid production using such host bacterial strains.
[0026] In some embodiments, engineered E. coli host cells are provided that have SbcC, SbcD, or both knockouts, but do not have specific additional mutations.
[0027] In some embodiments, methods for preparing manipulated E. coli host cells are provided.
[0028] In some embodiments, methods for replicating a vector in manipulated E. coli host cells of the present disclosure are provided. [Brief explanation of the drawing]
[0029] For a more complete understanding of the present invention and its advantages, refer to the following description in conjunction with the accompanying drawings.
[0030] [Figure 1A] Display pKD4 SbcCD targeting PCR fragments. [Figure 1B] Display the SbcCD gene locus. [Figure 1C] Display the integrated pKD4 PCR product that knocks out SbcCD. [Figure 1D] Displaying scars after FRT-mediated excision with pKD4 kanR markers. [Modes for carrying out the invention]
[0031] This disclosure provides a bacterial host strain, a method for modifying a bacterial host strain, and a manufacturing method that can improve the yield and quality of plasmids.
[0032] The bacterial host strains and methods of this disclosure can enable the improved production of vectors such as nonviral transposon vectors (transposase vectors, Sleeping Beauty transposase vectors, Sleeping Beauty transposase vectors, PiggyBac transposase vectors, PiggyBac transposase vectors, expression vectors, etc.) or nonviral gene editing (e.g., homologous recombination repair (HDR) / CRISPR-Cas9) vectors, as well as viral vectors (e.g., AAV vectors, AAV rep cap vectors, AAV helper vectors, Ad helper vectors, lentiviral vectors, lentiviral envelope vectors, lentiviral packaging vectors, retroviral vectors, retroviral envelope vectors, retroviral packaging vectors, etc.) for cell therapy, gene therapy, or gene replacement applications.
[0033] Improved plasmid production may include improved plasmid stability (e.g., reduced plasmid deletion, reversed, or other recombination products), and / or improved plasmid quality (e.g., reduced cleaved, linear, or dimerized products), and / or improved plasmid supercoiling (e.g., reduced supercoiled topological isoforms), compared to plasmid production using alternative host strains known in the art. All references cited herein should be understood to be incorporated by reference in their entirety.
[0034] definition As used herein, the singular forms "a," "an," and "the" include plural referents unless otherwise specified by the context.
[0035] The use of the term “or” in the claims and this disclosure shall mean “and / or” unless expressly indicated to refer only to the alternatives, or unless the alternatives are mutually exclusive.
[0036] The use of the term "approximately" when used with a number is intended to include a + / - 10% range. For example, if the number of amino acids is specified as approximately 200, this would include 180 to 220 (plus or minus 10%).
[0037] As used herein, “AAV vector” refers to an adeno-associated virus vector or an episomal virus vector. Examples of “AAV vectors” include, but are not limited to, self-complementary adeno-associated virus vectors (scAAV) and single-stranded adeno-associated virus vectors (ssAAV).
[0038] As used herein, "amp" refers to ampicillin.
[0039] As used herein, "ampR" refers to the ampicillin resistance gene.
[0040] As used herein, “bacterial region” refers to the region of a vector, such as a plasmid, required for prorogation and selection in a bacterial host.
[0041] When used herein, "Cat R This refers to the chloramphenicol resistance gene.
[0042] As used herein, “ccc” or “CCC” means “covalently closed circular” unless used in relation to a nucleotide or amino acid sequence.
[0043] As used herein, "cI" means lambda repressor.
[0044] As used herein, "cITs857" refers to a lambda repressor that further incorporates a C-to-T (Ala-to-Thr) mutation that confers temperature sensitivity. cITs857 is a functional repressor at 28–30°C but is nearly inactive at 37–42°C. It is also known as cI857 or cI857ts.
[0045] As used herein, "cmv" or "CMV" refers to cytomegalovirus.
[0046] As used herein, “copy cutter host strain” refers to an R6K origin-producing strain containing a phage φ80 attachment site chromosome-integrated copy of the arabinose-inducible CI857ts gene. Addition of arabinose to a plate or culture medium (e.g., up to a final concentration of 0.2–0.4%) induces pARA-mediated CI857ts repressor expression that reduces the copy number at 30°C through CI857ts-mediated downregulation of the R6K Rep protein expressing the pL promoter [i.e., additional CI857ts mediating a more effective downregulation of the pL(OL1-G~T) promoter at 30°C]. Copy number induction after a temperature shift to 37–42°C is not impaired because the CI857ts repressor is inactivated at these high temperatures. The copy cutter host strain increases the R6K vector temperature upshift copy number induction ratio by reducing the copy number at 30°C. This is advantageous for the production of large, toxic, or easily dimerizable R6K-based vectors.
[0047] As used herein, "dcm methylation" refers to methylation by E. coli methyltransferase that methylates the sequence CC(A / T)GG at the C5 position of the second cytosine.
[0048] As used herein, "derived from" means that the cells are descendants of a particular cell line. For example, "derived from DH5α" means that the cells are made from DH5α or its descendants. Thus, derivative cells may include polymorphisms and other changes that occur in the cell line as it is cultured.
[0049] As used herein, "EGFP" refers to highly sensitive green fluorescent protein.
[0050] As used herein, “manipulated E. coli strain” should be understood to mean the E. coli strain of this disclosure having a gene knockout (or knockdown) in SbcC, SbcD, or both, produced by human intervention.
[0051] As used herein, “manipulated mutation” should be understood as a mutation that does not occur naturally but is instead the product of direct human intervention.
[0052] As used herein, “eukaryotic expression vector” refers to a vector for expressing mRNA, protein antigens, protein therapeutics, shRNA, RNA, or microRNA genes in a target eukaryote using RNA polymerase I, II, or III promoters.
[0053] As used herein, “eukaryotic region” refers to the region of a plasmid that codes for eukaryotic sequences and / or sequences necessary for plasmid function in the target organism. This includes regions of plasmid vectors necessary for the expression of one or more transgenes in the target organism, including RNA PolII enhancers, promoters, transgenes, and polyA sequences. This also includes regions of plasmid vectors necessary for the expression of one or more transgenes in the target organism using RNA PolI or RNA PolIII promoters, RNA PolI or RNA PolIII expressing transgenes, or RNA. The eukaryotic region may optionally include other functional sequences such as eukaryotic transcription termination factors, supercoil-induced DNA double-strand destabilization (SIDD) structures, S / MARs, and boundary elements. In lentiviral or retroviral vectors, the eukaryotic region contains adjacent direct repeats (LTRs); in AAV vectors, the eukaryotic region contains adjacent reverse-terminus repeats; and in transposon vectors, the eukaryotic region contains adjacent transposon reverse-terminus repeats or IR / DR termins (e.g., Sleeping Beauty). In genome integration vectors, the eukaryotic region can encode homology arms to direct targeted integration.
[0054] As used herein, “expression vector” refers to a vector for the expression of mRNA, protein antigens, protein therapeutics, shRNA, RNA, or microRNA genes in a target organism.
[0055] As used herein, “target gene” refers to a gene expressed in a target organism. This includes mRNA genes encoding protein or peptide antigens, mRNA, shRNA, RNA, or microRNA encoding protein or peptide therapeutics, and mRNA, shRNA, RNA, or microRNA encoding RNA therapeutics, as well as mRNA, shRNA, RNA, or microRNA encoding RNA vaccines.
[0056] As used herein, “genome” in relation to Rep proteins and promoters refers to the nucleic acid sequence into which RNA-IN, including RNA-IN-regulated selectable markers, antibiotic resistance markers, and lambda repressors, is incorporated into a bacterial host strain.
[0057] As used herein, “high-yield plasmid-producing host” refers to sbcB, recB, recD, and recJ, as well as recA-, endA- cell lines that do not contain viability or yield-reducing mutations in uvrC, mcrA, and / or mcrBC-hsd-mrr, e.g., DH1, DH5α, JM107, JM108, JM109, MG1655, and XL1Blue.
[0058] As used herein, “HyperGRO fermentation process” refers to fed-batch fermentation, in which plasmid-containing E. coli cells are grown at a reduced temperature during a portion of the fed-batch phase, during which the growth rate is limited, followed by a temperature increase shift, during which the cells are continuously grown at a high temperature to accumulate plasmids. The temperature shift at a limited growth rate improved plasmid yield and purity.
[0059] As used herein, a “reverse repeat” refers to a single-stranded sequence of nucleotides followed downstream by its reverse complement. The intervening nucleotide sequence between the initial sequence and the reverse complement can be of any length, including zero. When the intervening length is zero, the composite sequence is a palindrome. It should be understood that reverse repeats can occur within double-stranded DNA, while other reverse repeats can occur within intervening sequences.
[0060] As used herein, "IR / DR" refers to a reverse repeat that is directly repeated twice. For example, the Sleeping Beauty transposon IR / DR repeat.
[0061] As used herein, “iteron” refers to a DNA sequence that is directly repeated at the origin of replication and is necessary for replication initiation. R6K origin iteron repeats are 22 bp long, such as sequence numbers 19-23 of WO2019 / 183248 (aaacatgaga gcttagtacg tg, aaacatgaga gcttagtacg tt, agccatgaga gcttagtacg tt, agccatgagg gtttagttcg tt, and aaacatgaga gcttagtacg ta, respectively).
[0062] As used herein, "ITR" refers to an inverted terminal repeat.
[0063] As used herein, "kan" refers to kanamycin.
[0064] As used herein, "kanR" refers to the kanamycin resistance gene.
[0065] As used herein, “knockdown” refers to the disruption of a gene that results in a decrease in the expression of the gene product and / or a decrease in the activity of the gene product.
[0066] As used herein, “knockout” refers to the disruption of a gene resulting in ablation of gene expression from that gene, and / or that the expressed gene product is non-functional.
[0067] As used herein, “Kozak sequence” refers to the optimized consensus DNA sequence gccRccATG (R=G or A) immediately upstream of the ATG start codon, which ensures efficient translation initiation. The SalI site (GTCGAC) immediately upstream of the ATG start codon (GTCGACATG) is a valid Kozak sequence.
[0068] As used herein, “lentiviral vector” refers to an embedded viral vector capable of infecting both dividing and non-dividing cells. It is also called a lentiviral transfer plasmid. The plasmid encodes a lentiviral LTR flanking expression unit. The transfer plasmid, along with the lentiviral envelope and packaging plasmids necessary for producing viral particles, is transfected into producing cells.
[0069] As used herein, “lentiviral envelope vector” refers to a plasmid encoding an envelope glycoprotein.
[0070] As used herein, “lentiviral packaging vector” refers to one or two plasmids that express the gag, pol, and Rev gene functions required for a lentiviral packaging vector.
[0071] As used herein, “minicircle” refers to a covalently bound closed cyclic plasmid derivative in which the bacterial region has been removed from the parent plasmid by in vivo or in vitro site-specific recombination or in vitro restriction digestion / ligation. Minicircle vectors are incapable of replicating in bacterial cells.
[0072] As used herein, "mSEAP" refers to mouse-secreted alkaline phosphatase.
[0073] As used herein, “Nanoplasmid® vector” refers to a vector that combines an RNA-selectable marker with R6K, ColE2, or a ColE2-associated origin of replication. Examples include the NTC9385C, NTC9685C, NTC9385R, NTC9685R vectors, and the modified versions described in WO2014 / 035457.
[0074] As used herein, “mutation” may refer to any type of mutation, such as substitution, addition, or deletion.
[0075] As used herein, "non-functional" with respect to the SbcCD complex refers to an SbcCD complex that is unable to cleave palindromic sequences.
[0076] As used herein, the “NTC8 series” refers to vectors such as the NTC8385, NTC8485, and NTC8685 plasmids, which are antibiotic-free pUC-derived vectors containing short RNA (RNA-OUT) selectable markers instead of antibiotic resistance markers such as kanR. The preparation and application of these RNA-OUT-based antibiotic-free vectors are described in WO2008 / 153733.
[0077] As used herein, “NTC9385R” refers to the NTC9385R Nanoplasmid® vector described in WO2014 / 035457, which has a spacer region encoding an NheI-trpA terminator-R6K origin RNA-OUT-KpnI bacterial region and is linked to a eukaryotic region via adjacent NheI and KpnI sites.
[0078] When used herein, "OD 600 " refers to the optical density at 600 nm.
[0079] As used herein, PCR refers to "polymerase chain reaction."
[0080] As used herein, “pDNA” refers to plasmid DNA.
[0081] As used herein, “piggyback transposon” refers to a transposon system that incorporates an ITR-adjacent PB transposon into the genome via a simple cleavage and paste mechanism mediated by a PB transposase. Transposon vectors typically contain a promoter-transgene-polyA expression cassette between PB ITRs that is excised and incorporated into the genome.
[0082] As used herein, "pINT pR pL vector" means "pINT pR pL att HK022 This refers to an integrated expression vector, which is described in Luke et al., 2011 Mol Biotechnol 47:43 and is incorporated herein by reference. The target gene to be expressed is cloned downstream of the pL promoter. The vector encodes a temperature-inducible cI857 repressor, enabling heat-inducible target gene expression.
[0083] When used in this specification, "P L "Promoter" refers to the lambda promoter on the left. L This is a potent promoter that is suppressed by cI repressor binding to the OL1, OL2, and OL3 repressor binding sites. The temperature-sensitive cI857 repressor is functional at 30°C, suppressing gene expression, but is inactivated at 37-42°C, allowing gene expression to occur, thus enabling heat-induced control of gene expression.
[0084] When used in this specification, "P L The term "(OL1G~T) promoter" refers to the left-hand lambda promoter that has a mutation from OL1G to T. L This is a potent promoter that is repressed by cI repressor binding to the OL1, OL2, and OL3 repressor binding sites. The temperature-sensitive cI857 repressor is functional at 30°C, repressing gene expression, but is inactivated at 37–42°C, allowing gene expression to occur, thus enabling heat-induced control of gene expression. cI repressor binding to OL1 is reduced by the OL1G to T mutation, as described in WO2014 / 035457, resulting in increased promoter activity at 30°C and 37–42°C.
[0085] As used herein, “plasmid” refers to an extra chromosomal DNA molecule isolated from chromosomal DNA that can replicate independently of chromosomal DNA.
[0086] As used herein, “plasmid copy number” refers to the number of plasmid copies per cell. An increase in plasmid copy number indicates an increase in plasmid production yield.
[0087] As used herein, "Pol" refers to polymerase.
[0088] As used herein, "PolI" refers to E. coli DNA polymerase I.
[0089] As used herein, "PolIII" refers to E. coli DNA polymerase III.
[0090] As used herein, “PolIII-dependent origin of replication” refers to an origin of replication that does not require PolI, such as the rep protein-dependent R6K gamma origin of replication. Many additional PolIII-dependent origins of replication are known in the art, many of which are summarized in del Solar et al., 1998, which is incorporated herein by reference.
[0091] As used herein, "poly(A)" refers to a polyadenylation signal or site. Polyadenylation is the addition of a poly(A) tail to an RNA molecule. Polyadenylation signals contain a sequence motif recognized by RNA cleavage complexes. Most human polyadenylation signals contain the AAUAAA motif and its conserved 5' and 3' sequences. Commonly used poly(A) signals are derived from rabbit β-globin, bovine growth hormone, early SV40, or late SV40 poly(A) signals.
[0092] As used herein, “poly-A repeat” refers to a sequence of adenine nucleotides as a direct repeat. Similarly, “poly-G repeat” refers to a sequence of guanine nucleotides as a direct repeat, “poly-C repeat” refers to a sequence of cytosine nucleotides as a direct repeat, and “poly-T repeat” refers to a sequence of thymine nucleotides as a direct repeat. “mRNA vector” contains poly-A repeats.
[0093] As used herein, “pUC origin” refers to a replication origin derived from pBR322 that has a G-to-A transposition that increases in copy number at elevated temperatures and deletes a ROP-negative regulator.
[0094] As used herein, "pUC-free" refers to a plasmid that does not contain a pUC origin.
[0095] As used herein, "pUC plasmid" refers to a plasmid containing a pUC origin.
[0096] As used herein, “R6K plasmid” refers to plasmids having an origin of replication of R6K or an origin of replication of R6K, such as the NTC9385R, NTC9685R, NTC9385R2-O1, NTC9385R2-O2, NTC9385R2a-O1, NTC9385R2a-O2, NTC9385R2b-O1, NTC9385R2b-O2, NTC9385Ra-O1, NTC9385Ra-O2, NTC9385RaF, and NTC9385RbF vectors, as well as modified and surrogate vectors containing an R6K origin of replication as described in WO2014 / 035457 and WO2019 / 183248. Known alternative R6K vectors in this field include, but are not limited to, the pCOR vector (Gencell), the pCpG-free vector (Invivogen), and the Oxford University CpG-free vector, including pGM169.
[0097] As used herein, “R6K replication origin” refers to a region specifically recognized by the R6K Rep protein to initiate DNA replication, including, but not limited to, the R6K gamma replication origin sequences disclosed in WO2019 / 183248 as SEQ ID NOs. 1, 2, 4, and 18 (SEQ ID NOs. 43-44, 46, and 60, respectively). It also includes the CpG-free variant described in Drocourt et al., U.S. Patent No. 7244609 (SEQ ID NO: 3), which is incorporated herein by reference (SEQ ID NO: 63).
[0098] As used herein, “R6K replication origin-RNA-OUT bacterial origin” comprises an R6K replication origin for propagation and RNA-OUT selectable markers disclosed in WO2019 / 183248 (SEQ ID NOs. 50-59, respectively) (e.g., SEQ ID NOs. 8, 9, 10, 11, 12, 13, 14, 15, 16, 17).
[0099] As used herein, “Rep protein-dependent plasmid” refers to a plasmid whose replication depends on a replication (Rep) protein provided in trans. Examples include the R6K origin of replication, the ColE2-P9 origin of replication, and ColE2-associated origin of replication plasmids from which the Rep protein is expressed from the host strain genome. Numerous additional Rep protein-dependent plasmids are known in the art, many of which are summarized in del Solar et al., 1998, Mol. Biol. Rev. 62:44-464, which are incorporated herein by reference.
[0100] As used herein, “retroviral vector” refers to an embedded viral vector capable of infecting dividing cells. It is also called a transfer plasmid. The plasmid encodes a retroviral LTR flanking expression unit. The transfer plasmid, along with the envelope and packaging plasmids necessary to produce the viral particle, is transfected into the producing cell.
[0101] As used herein, “retroviral envelope vector” refers to a plasmid encoding an envelope glycoprotein.
[0102] As used herein, “retroviral packaging vector” refers to plasmids encoding retroviral gag and pol genes necessary for packaging a retroviral transfer vector.
[0103] As used herein, “RNA-IN” refers to an insertion sequence 10 (IS10) that codes for RNA-IN, which is RNA-complementary and antisense to a portion of RNA's RNA-OUT. When RNA-IN is cloned into the untranslated reader of mRNA, the annealing of RNA-IN to RNA-OUT reduces the translation of the gene encoded downstream of RNA-IN.
[0104] As used herein, “RNA-IN regulated selective marker” refers to a genome-expressed RNA-IN regulated selective marker. In the presence of plasmid-borne RNA-OUT antisense repressor RNA (e.g., SEQ ID NO: 6 disclosed in WO2019 / 183248 (SEQ ID NO: 48)), the expression of a protein encoded downstream of RNA-IN (e.g., having the sequence gccaaaaatcaataatcagacaacaagatg) is repressed. RNA-IN regulated selective markers are configured such that RNA-IN regulates either the protein itself or a toxic substance (e.g., SacB) that is lethal or toxic to the cell, or 2) the transcription of a gene essential for the proliferation of the bacterial cell (e.g., the murA essential gene regulated by the RNA-IN tetR repressor gene) that is lethal or toxic to the cell. For example, a genome-expressing RNA-IN-SacB cell line for RNA-OUT plasmid selection / propagation is described in WO2008 / 153733. Alternative selection markers described in the art may be replaced with SacB.
[0105] As used herein, “RNA-OUT” refers to the RNA-OUT encoded in insertion sequence 10 (IS10), which is an antisense RNA that hybridizes to a transposon gene expressed downstream of RNA-IN and reduces translation. The RNA-OUT RNA (SEQ ID NO: 6) disclosed in WO2019 / 183248 (SEQ ID NO: 48) and the sequences of the RNA-IN-SacB cell line expressed genomically in complementary RNA-IN SacB may be modified to incorporate alternative functional RNA-IN / RNA-OUT binding pairs, such as those described in Mutalik et al., 2012 Nat Chem Biol 8:447, including the RNA-OUT A08 / RNA-IN S49 pair, the RNA-OUT A08 / RNA-IN S08 pair, and RNA-OUT
[0106] This includes, but is not limited to, CpG-free modifications of RNA-OUT A08 that modify the CG in the TIFF2026048673000002.tif930 sequence to a non-CpG sequence. CpG-free RNA-OUTs may also be constructed using a number of alternative substitutions to remove two CpG motifs (mutating each CpG to either CpA, CpC, CpT, ApG, GpG, or TpG).
[0107] As used herein, “RNA-OUT selectable marker” refers to an RNA-OUT selectable marker DNA fragment containing an E. coli transcription promoter and terminator sequence adjacent to an RNA-OUT RNA. RNA-OUT selectable markers utilizing RNA-OUT promoter and terminator sequences adjacent to DraIII and KpnI restriction enzyme sites, as well as designer RNA-IN-SacB cell lines expressed genome-wise for RNA-OUT plasmid propagation, are described in WO2008 / 153733 and are incorporated herein by reference. The RNA-OUT promoter and terminator sequences adjacent to the RNA-OUT RNA may be replaced with heterologous promoter and terminator sequences. For example, the RNA-OUT promoter may be replaced with a CpG-free promoter known in the art, such as the I-EC2K promoter, or the P5 / 6 5 / 6 or P5 / 6 6 / 6 promoter described in WO2008 / 153733 and incorporated herein by reference. A 2CpG RNA-OUT selectable marker, in which two CpG motifs within the RNA-OUT promoter were removed, was given as Sequence ID 7 of WO2019 / 183248 (SEQ ID NO: 49). Vectors incorporating the CpG-free RNA-OUT selectable marker can be selected for sucrose tolerance using the RNA-IN-SacB cell line described in WO2008 / 153733, or any cell line having RNA-IN-SacB as described in WO2008 / 153733. Alternatively, the RNA-IN sequence in these cell lines may be modified to incorporate a 1 bp change necessary to perfectly match the CpG-free RNA-OUT region complementary to RNA-IN.
[0108] As used herein, “RNA-selectable marker” refers to plasmid-retained non-coding RNA that modulates a chromosomally expressed target gene to result in selection. This may be a nonsense repressor tRNA that modulates a nonsense-repressible selectable chromosomal target, as described in U.S. Patent No. 6,977,174, 2005 by Crouzet J and Soubrier F, included herein by reference. This may also be plasmid-borne antisense repressor RNAs, and an unspecified list included herein by reference includes RNA-OUT (WO2008 / 153733), which represses the RNA-IN regulatory target; pMB1 plasmid originating code RNAI (Grabherr R, Pfaffenzeller I. 2006 U.S. Patent Application No. 2006 / 0063232; Cranenburgh RM. 2009; U.S. Patent No. 7,611,883), IncB plasmid originating code pMU720, which represses the RNA-II regulatory target (Wilson IW, Siemering KR, Prazkier J, Pittard AJ. 1997. J Bacteriol 179:742-53); ParB locus Sok of plasmid R1, which represses the Hok regulatory target; and FlmB of F plasmid, which represses the flmA regulatory target (Morsey Examples include MA, 1999 (U.S. Patent No. 5922583). RNA-selectable markers may also be other natural antisense repressor RNAs known in the art, such as those described in Wagner EGH, Altuvia S, Romby P. 2002. Avd Genet 46:361-98 and Franch T, and Gerdes K. 2000. Current Opin Microbiol 3:159-64. RNA-selectable markers may also be engineered repressor RNAs, such as synthetic small RNAs expressed in SgrS, MicC, or MicF scaffolds, as described in Na D, Yoo SM, Chung H, Park H, Park JH, Lee SY. 2013. Nat Biotechnol 31:170-4.RNA-selectable markers may also be engineered repressor RNAs as part of a selectable marker that represses target RNA fused to a regulated target gene such as SacB, as described in US2015 / 0275221.
[0109] As used herein, "SacB" refers to the structural gene encoding Bacillus subtilus levansucase. Expression of SacB in Gram-negative bacteria is toxic in the presence of sucrose.
[0110] As used herein, "SEAP" refers to secreted alkaline phosphatase.
[0111] As used herein, “selectable marker” or “selection marker” refers to a selectable marker, such as a kanamycin resistance gene or an RNA selectable marker.
[0112] As used herein, the term “sequence identity” refers to the degree of identity between any given query sequence and subject sequence. A subject sequence may have, for example, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with a given query sequence. To determine the sequence identity percentage, the query sequence (e.g., a nucleic acid sequence) is aligned to one or more subject sequences (global alignment) using any suitable sequence alignment program known in the art, for example, the computer program ClustalW (version 1.83, default parameters), which allows the alignment of nucleic acid sequences to be performed over their entire length. Chema et al., 2003 Nucleic Acids Res., 31:3497-500. In a preferred method, a sequence alignment program (e.g., ClustalW) calculates the best match between a query sequence and one or more subject sequences and aligns them so that identity, similarity, and difference can be determined. One or more nucleotide gaps can be inserted into the query sequence, subject sequences, or both to maximize the sequence alignment. For fast pairwise alignment of nucleic acid sequences, preferred default parameters suitable for a particular alignment program can be selected. The output is a sequence alignment reflecting the relationships between sequences. To further determine the identity percentage of the subject nucleic acid sequence to the query sequence, align the sequences using the alignment program, divide the number of identical matches in the alignment by the length of the query sequence, and multiply the result by 100. Note that the identity percentage value can be rounded to two decimal places. For example, 78.11, 78.12, 78.13, and 78.14 are rounded down to 78.1, while 78.15, 78.16, 78.17, 78.18, and 78.19 are rounded up to 78.2.
[0113] As used herein, "shRNA" refers to short hairpin RNA.
[0114] As used herein, “S / MAR” refers to a scaffold / matrix attachment region containing a eukaryotic sequence that mediates DNA attachment to the nuclear matrix.
[0115] As used herein, “Sleeping Beauty transposon” refers to a transposon system that incorporates IR / DR-adjacent SB transposons into the genome via a simple cleavage and paste mechanism mediated by SB transposases. Transposon vectors typically contain a promoter-transgene-polyA expression cassette between IR / DR that is excised and incorporated into the genome.
[0116] As used herein, the “spacer region” refers to the region connecting the 5' and 3' ends of a eukaryotic region sequence. Since the 5' and 3' ends of the eukaryotic region are typically separated by the bacterial origin of replication and bacterial selectable marker in the plasmid vector (bacterial region), many spacer regions consist of the bacterial region. In the PolIII-dependent origin of replication vector of the present invention, this spacer region is preferably less than 1000 bp.
[0117] As used herein, “structured DNA sequence” refers to a DNA sequence capable of forming a replication-inhibiting secondary structure (Mirkin and Mirkin, 2007. Microbiology and Molecular Biology Reviews 71:13-35). This includes, but is not limited to, reverse repeats, palindromes, direct repeats, IR / DR, homopolymer repeats containing eukaryotic promoter enhancers or repeats containing eukaryotic promoter enhancers, or repeats containing eukaryotic origins of replication.
[0118] As used herein, “SV40 origin” refers to Simian virus 40 genomic DNA containing the origin of replication.
[0119] As used herein, “SV40 enhancer” refers to Simian virus 40 genomic DNA containing 72 bp and optionally 21 bp enhancer repeats.
[0120] As used herein, "TE buffer" refers to a solution containing approximately 10 mM Tris pH 8 and 1 mM EDTA.
[0121] As used herein, "TetR" refers to the tetracycline resistance gene.
[0122] As used herein, “transcriptional terminator” means (1) in the context of bacteria, a DNA sequence that marks the end of a gene or operon for transcription. This may be an intrinsic transcription termination factor or a Rho-dependent transcription termination factor. In the case of an intrinsic terminator, such as the trpA terminator, a hairpin structure is formed within the transcript that disrupts the mRNA-DNA-RNA polymerase ternary complex. Alternatively, a Rho-dependent transcriptional terminator disrupts the Rho factor, RNA helicase protein complex, or (2) in the context of eukaryotes, the polyA signal is not a “terminator,” but rather an internal cleavage at the polyA site leaves the untapped 5' end on the 3'UTR RNA for nuclease digestion. The nuclease catches up with RNA PolII, causing termination. Termination may be facilitated within a short region of the polyA site by the introduction of an RNA PolII pause site (eukaryotic transcriptional terminator). The pause of RNA PolII allows nucleases introduced into 3'UTR mRNA after PolyA cleavage to catch up with RNA PolII at the pause site. A non-exclusive list of eukaryotic transcriptional terminators known in the art includes C2x4 and gastrin terminators. Eukaryotic transcriptional terminators can increase mRNA levels by enhancing proper 3' end processing of mRNA.
[0123] As used herein, “transfection” means methods for delivering nucleic acids to cells, as known in the art and included herein by reference [e.g., poly(lactide-coglycolide) (PLGA), ISCOM, liposomes, niosomes, viromosomes, block copolymers, pluronic block copolymers, chitosan, and other biodegradable polymers, microparticles, microspheres, calcium phosphate nanoparticles, nanoparticles, nanocapsules, nanospheres, poloxamine nanospheres, electroporation, nucleofection, piezoelectric permeabilization, sonoporation, iontophoresis, ultrasound, SQZ fast cell deformation-mediated membrane disruption, corona plasma, plasma-assisted delivery, tissue-resistant plasma, laser microporation, shock wave energy, magnetic field, non-contact magnetic permeabilization, gene cancer, microneedles, microdermabrasion, hydrodynamic delivery, high-pressure tail vein injection, etc.]. The transfection of DNA into E. coli, commonly referred to as transformation, is typically carried out using chemically competent or electrocompetent E. coli cells, employing standard methodologies known in the art and incorporated herein by reference.
[0124] As used herein, “transgene” refers to the gene of interest that is cloned into a vector for expression in a target organism.
[0125] As used herein, “transposase vector” refers to a vector that encodes a transposase.
[0126] As used herein, “transposon vector” refers to a vector encoding a transposon, which is a substrate for transposase-mediated gene integration.
[0127] As used herein, "ts" means temperature sensitivity.
[0128] As used herein, “UTR” refers to the untranslated region of mRNA (5' or 3' relative to the coding region).
[0129] As used herein, “vector” refers to gene delivery vehicles including viral (e.g., alphaviruses, poxviruses, lentiviruses, retroviruses, adenoviruses, adenovirus-associated viruses, etc.) and non-viral (e.g., plasmids, MIDGE, transcriptionally active PCR fragments, minicircles, bacteriophages, Nanoplasmid®, etc.) vectors. These are well known in the art and are incorporated herein by reference.
[0130] As used herein, “vector skeleton” refers to the eukaryotic and bacterial regions of a vector that do not contain the transgene or target antigen coding region.
[0131] In some embodiments, the engineered Escherichia coli (E.coli) host cells include a gene knockout of at least one gene selected from the group consisting of SbcC and SbcD, and the engineered E.coli host cells do not contain any engineered viability or yield reduction mutations in any of sbcB, recB, recD, and recJ, and optionally, at least one of uvrC, mcrA, mcrBC-hsd-mrr and their combinations. In some embodiments, the engineered E.coli host cells do not contain any engineered mutations in any of sbcB, recB, recD, and recJ, and optionally, at least one of uvrC, mcrA, mcrBC-hsd-mrr and their combinations. In some embodiments, the manipulated E. coli host cells are free from any mutations in sbcB, recB, recD, and recJ, and optionally, in at least one of uvrC, mcrA, mcrBC-hsd-mrr, and combinations thereof.
[0132] It should be understood that engineered E. coli host cells comprising gene knockout (or knockdown) of at least one gene selected from the group consisting of SbcC and SbcD are within the scope of this disclosure, and that engineered E. coli host cells do not contain engineered viability or yield reduction mutations in at least one of sbcB, recB, recD, recJ, uvrC, mcrA, and mcrBC-hsd-mrr, or in some embodiments, do not contain engineered mutations or any mutations at all. It should also be understood that engineered E. coli host cells comprising gene knockout of at least one gene selected from the group consisting of SbcC and SbcD are within the scope of this disclosure, and that engineered E. coli host cells do not contain engineered viability or yield reduction mutations in at least one of sbcB, recB, recD, and recJ, or in some embodiments, do not contain engineered mutations or any mutations at all. In some embodiments, the manipulated E. coli host cells include a gene knockout of at least one gene selected from the group consisting of SbcC and SbcD, but in mcrA, they do not contain a survival or yield reduction mutation, or in some embodiments, they do not contain a manipulated mutation or any mutation at all. In some embodiments, the manipulated E. coli host cells include a gene knockout of at least one gene selected from the group consisting of SbcC and SbcD, and the manipulated E. coli host cells do not contain a manipulated survival or yield reduction mutation in any of sbcB, recB, recD, and recJ, or in some embodiments, they do not contain a manipulated mutation or any mutation at all.
[0133] In other embodiments, the manipulated E. coli host cells include a gene knockout of at least one gene selected from the group consisting of SbcC and SbcD, and do not contain any manipulated viability or yield reduction mutations in at least one of sbcB, recB, recD, recJ, uvrC, mcrA, and mcrBC-hsd-mrr. In other embodiments, the manipulated E. coli host cells include a gene knockout of at least one gene selected from the group consisting of SbcC and SbcD, and do not contain any manipulated mutations in at least one of sbcB, recB, recD, recJ, uvrC, mcrA, and mcrBC-hsd-mrr. In other embodiments, the manipulated E. coli host cells include a gene knockout of at least one gene selected from the group consisting of SbcC and SbcD, and do not contain any mutations in at least one of sbcB, recB, recD, recJ, uvrC, mcrA, and mcrBC-hsd-mrr. In some embodiments, the manipulated E. coli host cells contain a gene knockout of at least one gene selected from the group consisting of SbcC and SbcD, and are free from any mutations in sbcB, recB, recD, recJ, and uvrC. In some embodiments, the manipulated E. coli host cells contain a gene knockout of at least one gene selected from the group consisting of SbcC and SbcD, and are free from any mutations in mcrA.
[0134] In some embodiments, engineered E. coli host cells are provided that include a gene knockout of at least one gene selected from the group consisting of SbcC and SbcD, wherein the engineered E. coli host cells do not contain any engineered viability or yield reduction mutations in any of sbcB, recB, recD, and recJ. In any of the embodiments described above, the engineered E. coli host cells cannot contain any engineered mutations in sbcB, recB, recD, and recJ. In any of the embodiments described above, the engineered E. coli host cells cannot contain any mutations in any of sbcB, recB, recD, and recJ. In some embodiments, engineered E. coli host cells are provided, comprising a gene knockout of at least one gene selected from the group consisting of SbC and SbcD, wherein the E. coli host cells are isogenic to the strain from which they originate, which is selected from the group consisting of DH5α, DH1, JM107, JM108, JM109, MG1655, and XL1Blue. In some embodiments, engineered E. coli host cells are provided, comprising a gene knockout of at least one gene selected from the group consisting of SbC and SbcD, wherein the E. coli host cells are isogenic to the strain from which they originate, which is selected from the group consisting of DH5α(dcm-), NTC4862, NTC4862-HF, NTC1050811, NTC1050811-HF, NTC1050811-HF(dcm-), HB101, TG1, and NEB Turbo.
[0135] To the extent that is not inconsistent with any of the embodiments described above, the manipulated E. coli cells may further be free from manipulated viability or yield reduction mutations in at least one of uvrC, mcrA, mcrBC-hsd-mrr, and combinations thereof. In any of the embodiments described above, the manipulated E. coli host cells may further be free from any manipulated mutations in at least one of uvrC, mcrA, mrBC-hsd-mrr, and combinations thereof. In any of the embodiments described above, the manipulated E. coli host cells may further be free from any mutations in at least one of uvrC, mcrA, mrBC-hsd-mrr, and combinations thereof. Therefore, in some embodiments, the manipulated E. coli host cells may further be free from manipulated viability or yield reduction mutations, manipulated mutations, or any mutations in uvrC. In other embodiments, the manipulated E. coli host cells may further be free of manipulated viability or yield reduction mutations, manipulated mutations, or any mutations in mcrA. Furthermore, in other embodiments, the manipulated E. coli host cells may further be free of manipulated viability or yield reduction mutations, manipulated mutations, or any mutations in mcrBC-hsd-mrr. In yet another embodiment, the manipulated E. coli host cells may further be free of manipulated viability or yield reduction mutations, manipulated mutations, or any mutations in mcrA and mrBC-hsd-mrr. Throughout this disclosure, it should be understood that mrBC-hsd-mrr refers to sequences including sequences 16-21.
[0136] In any of the embodiments described above, the manipulated E. coli host cells may contain non-functional SbcCD complexes, or in other words, may not contain functional SbcCD complexes. Alternatively, in some embodiments, the manipulated E. coli host cells may not contain SbcCD complexes.
[0137] In any of the embodiments described above, the gene knockout of the manipulated E. coli host cells may be a knockout of SbcC. Alternatively, in some embodiments, the gene knockout of the manipulated E. coli host cells may be a knockout of SbcD. In any of the embodiments described above, the gene knockout of the manipulated E. coli host cells may be a knockout of both SbcC and SbcD.
[0138] In any of the embodiments described above, the manipulated E. coli host cells may be derived from cell lines selected from the group consisting of DH5α, DH1, JM107, JM108, JM109, MG1655, and XL1Blue. In any of the embodiments described above, the manipulated E. coli host cells may be derived from DH5α(dcm-), NTC4862, NTC4862-HF, NTC1050811, NTC1050811-HF, or NTC1050811-HF(dcm-). In any of the embodiments described above, the manipulated E. coli host cells may be derived from cell lines selected from the group consisting of HB101, TG1, and NEB Turbo. The genotypes of these cell lines are as follows: DH5α(dcm-):DH5α dcm- NTC4862:DH5α att λ ::P c -RNA-IN-SacB,catR NTC4862-HF:DH5α att λ ::P c -RNA-IN-SacB,catR;att φ80 ::pARA-CI857ts P c -RNA-IN-SacB,tetR NTC1050811:DH5α att λ ::P c -RNA-IN-SacB,catR;att HK022 ::pL(OL1-G to T)P42L-P106I-F107S P113S(P3-),SpecR StrepR;att φ80 ::pARA-CI857ts,tetR NTC1050811-HF:DH5α att λ ::P c -RNA-IN-SacB,catR;att HK022 ::pL(OL1-G to T)P42L-P106I-F107S P113S(P3-),SpecR StrepR;att φ80 ::pARA-CI857ts P c -RNA-IN-SacB,tetR NTC1050811-HF(dcm-):DH5α dcm-att λ ::P c -RNA-IN-SacB,catR;att HK022 ::pL(OL1-G to T)P42L-P106I-F107S P113S(P3-),SpecR StrepR;att φ80 ::pARA-CI857ts P c -RNA-IN-SacB,tetR HB101:F - mcrB expression of hsdS20(r B - m B - )recA13 leuB6 ara-14 proA2 lacY1 galK2 xyl-5 mtl-1 rpsL20(Sm R )glnV44 λ - TG1:K-12 glnV44 thi-1 Δ(lac-proAB)Δ(mcrB-hsdSM)5(r K - m K - )F′[traD36 proAB + lacI q lacZΔM15] NEB Turbo:F' proA + B + lacI q ΔlacZM15 / fhuA2 Δ(lac-proAB)glnV galK16 galE15 R(zgb-210::Tn10)Tet S endA1 thi-1 Δ(hsdS-mcrB)5
[0139] In any of the embodiments described above, the manipulated E. coli host cells may further include a genomic antibiotic resistance marker. For example, but not limited to, the genomic antibiotic resistance marker may be a kanR containing a sequence having at least 90%, at least 95%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 23 (kanR, 795 bp). As a further example, but not limited to, the genomic antibiotic resistance marker may be a kanR containing a sequence encoding a protein having at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 36 (kanR). As yet another example, the genomic antibiotic resistance marker may be a chloramphenicol resistance marker, a gentamicin resistance marker, a kanamycin resistance marker, a spectinomycin and streptomycin resistance marker, a trimethoprim resistance marker, or a tetracycline resistance marker. Alternatively, in any of the embodiments described above, the E. coli host cells may not contain a genomic antibiotic resistance marker.
[0140] In any of the embodiments described above, the manipulated E. coli host cells may further contain Rep proteins suitable for culturing Rep protein-dependent plasmids. For example, but not limited to, the manipulated E. coli host cells may contain genomic nucleic acid sequences having at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with sequences selected from the group consisting of SEQ ID NO: 26 (P42L-P106I-F107S-P113S, 918bp), SEQ ID NO: 27 (P42L-Δ106-107-P113S, 912bp), SEQ ID NO: 28 (P42L-P106L-F107S, 918bp), and SEQ ID NO: 29 (P42L-P113S, 918bp). As a further example, but not limited to, engineered E. coli host cells may contain genomic nucleic acid sequences encoding Rep proteins that have at least 90%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acid sequences selected from the group consisting of SEQ ID NO: 39 (P42L-P106I-F107S-P113S), SEQ ID NO: 40 (P42L-Δ106-107-P113S), SEQ ID NO: 42 (P42L-P106L-F107S), SEQ ID NO: 41 (P42L-P113S), SEQ ID NO: 34 (ColE2 wild-type), and SEQ ID NO: 35 (ColE2 mutant G194D). For example, but not limited to, manipulated E. coli host cells may contain Rep proteins having at least 90%, at least 95%, at least 98%, at least 99%, or 100% identity with amino acid sequences selected from the group consisting of SEQ ID NO: 39 (P42L-P106I-F107S-P113S), SEQ ID NO: 40 (P42L-Δ106-107-P113S), SEQ ID NO: 42 (P42L-P106L-F107S, 305aa), SEQ ID NO: 41 (P42L-P113S, 305aa), SEQ ID NO: 34 (ColE2 wild-type), and SEQ ID NO: 35 (ColE2 mutant G194D). L It may be under the control of the promoter, and such P LIt should be understood that if the promoter has a genomically present lambda repressor such as cITs857, it may enable temperature-sensitive expression of the Rep protein. For example, but not limited to, P L The promoter is ttgacataaa taccactggc ggtgatact(P L promoter(-35~-10)),ttgacataaa taccactggc gtgatact(P L Promoter OL1-G(-35~-10)), or ttgacataaa taccactggc gttgatact(P L The sequence may have at least 95%, at least 98%, at least 99%, or 100% sequence identity with the promoter OL1-G~T(-35~-10). It should be further understood that if the Rep protein is an R6K Rep protein such as SEQ ID NOs. 39~42, the vector transfected into engineered E. coli host cells may contain an R6K origin of replication, or if the Rep protein is a ColE2 Rep protein, the vector transfected into engineered E. coli host cells may contain a ColE2 origin of replication.
[0141] In any of the embodiments described above, the engineered E. coli host cell may further include a genomic nucleic acid sequence encoding a genome-expressed RNA-IN-regulated selectable marker. For example, but not limited to, the engineered E. coli host cell may include a genomic nucleic acid sequence (encoding a selectable marker) having at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 25 (SacB, 1422 bp). For example, but not limited to, the engineered E. coli host cell may include a genomic nucleic acid sequence encoding a selectable marker having an amino acid sequence having at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 38 (SacB). As a further example, but not limited to, engineered E. coli host cells may include RNA-IN-modified selectable markers having an amino acid sequence with at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 38 (SacB). In any of the embodiments described above, the RNA-IN-modified selectable marker may be downstream of the RNA-IN having the sequence gccaaaaatcaataatcagacaacaagatg, and in embodiments using this RNA-IN, the corresponding RNA-OUT in the vector may be SEQ ID NO: 6 of WO2019 / 183248 (SEQ ID NO: 48). Therefore, with respect to SacB, the RNA-IN SacB sequence is,
[0142] This may be TIFF2026048673000003.tif110150. Any suitable RNA-IN-modulated selection marker and RNA-IN can be used, and it should be understood that these are known in the art.
[0143] In any of the embodiments described above, the manipulated E. coli host cell may further contain a genomic nucleic acid sequence encoding a temperature-sensitive lambda repressor. For example, the temperature-sensitive lambda repressor may be cITs857. For example, but not limited to, engineered E. coli host cells may contain a genomic nucleic acid sequence (encoding a temperature-sensitive lambda repressor) having at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 24 (cITs857, 714 bp). For further example, but not limited to, engineered E. coli host cells may further contain a genomic nucleic acid sequence encoding cITs857 having an amino acid sequence with at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 37 (cITs857). As a further example, but not limited to, the engineered E. coli host cell may further include a temperature-sensitive lambda repressor having an amino acid sequence having at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 37 (cITs857). In any of the embodiments described above, if the engineered E. coli host cell further includes a genomic nucleic acid sequence encoding a temperature-sensitive lambda repressor, the temperature-sensitive lambda repressor may be a phage φ80 binding site chromosomal integration copy of the arabinose-inducible CITs857 gene. As an example, but not limited to, the cITs857 gene is under the control of the pBAD promoter and is arabinose-inducible (pBAD promoter,
[0144] We can provide TIFF2026048673000004.tif98148).
[0145] In some embodiments, the following genotype is used: F-φ80lacZΔM15 Δ(lacZYA-argF)U169 recA1 endA1 hsdR17(r k -,m k Manipulated E. coli host cells having +)gal-phoA supE44 λ-thi-1 gyrA96 relA1 ΔSbcDC::kanR are provided.
[0146] In some embodiments, the following genotype is used: F-φ80lacZΔM15 Δ(lacZYA-argF)U169 recA1 endA1 hsdR17(r k -,m k Manipulated E. coli host cells are provided, possessing +)gal-phoA supE44 λ-thi-1 gyrA96 relA1 ΔSbcDC.
[0147] In some embodiments, the following genotype: DH5α att HK022 Manipulated E. coli host cells are provided, possessing ::pL(OL1-G to T)P42L-P106I-F107S P113S(P3-),SpecR StrepR;ΔSbcDC::kanR.
[0148] In some embodiments, the following genotype: DH5α att HK022 Manipulated E. coli host cells are provided, possessing ::pL(OL1-G to T)P42L-P106I-F107S P113S(P3-),SpecR StrepR;ΔSbcDC.
[0149] In some embodiments, the following genotype is used: F-φ80lacZΔM15 Δ(lacZYA-argF)U169 recA1 endA1 hsdR17(r k -,m k Manipulated E. coli host cells having +)gal-phoA supE44 λ-thi-1 gyrA96 relA1;ΔSbcDC::kanR are provided.
[0150] In some embodiments, manipulated E. coli host cells having the following genotype:DH5α dcm-;ΔSbcDC are provided.
[0151] In some embodiments, manipulated E. coli host cells having the following genotype:DH5α dcm-;ΔSbcDC::kanR are provided.
[0152] In some embodiments, the following genotype: DH5α att λ ::P c Engineered E. coli host cells having RNA-IN-SacB,catR;ΔSbcDC are provided.
[0153] In some embodiments, the following genotype: DH5α att λ ::P c Manipulated E. coli host cells having -RNA-IN-SacB,catR;ΔSbcDC::kanR are provided.
[0154] In some embodiments, the following genotype: DH5α att λ ::P c -RNA-IN-SacB,catR;att φ80 ::pARA-CI857ts P c Engineered E. coli host cells having RNA-IN-SacB,tetR;ΔSbcDC are provided.
[0155] In some embodiments, the following genotype: DH5α att λ ::P c -RNA-IN-SacB,catR;att φ80 ::pARA-CI857ts P c Manipulated E. coli host cells having -RNA-IN-SacB,tetR;ΔSbcDC::kanR are provided.
[0156] In some embodiments, the following genotype: DH5α att λ ::P c -RNA-IN-SacB,catR;att HK022 ::pLOL1-G to T)P42L-P106I-F107S P113S(P3-),SpecR StrepR;att φ80 Manipulated E. coli host cells having ::pARA-CI857ts,tetR;ΔSbcDC are provided.
[0157] In some embodiments, the following genotype: DH5α att λ ::P c -RNA-IN-SacB,catR;att HK022 ::pL(OL1-G to T)P42L-P106I-F107S P113S(P3-),SpecR StrepR;att φ80 Manipulated E. coli host cells having ::pARA-CI857ts,tetR;ΔSbcDC::kanR are provided.
[0158] In some embodiments, the following genotype: DH5α att λ ::P c -RNA-IN-SacB,catR;att HK022 ::pL(OL1-G to T)P42L-P106I-F107S P113S(P3-),SpecR StrepR;att φ80 ::pARA-CI857ts P c Engineered E. coli host cells having RNA-IN-SacB,tetR;ΔSbcDC are provided.
[0159] In some embodiments, the following genotype: DH5α att λ ::P c -RNA-IN-SacB,catR;att HK022 ::pL(OL1-G to T)P42L-P106I-F107S P113S(P3-),SpecR StrepR;att φ80 ::pARA-CI857ts P c Manipulated E. coli host cells having -RNA-IN-SacB,tetR;ΔSbcDC::kanR are provided.
[0160] In some embodiments, the following genotype is used: DH5α dcm-att λ ::P c -RNA-IN-SacB,catR;att HK022 ::pL(OL1-G to T)P42L-P106I-F107S P113S(P3-),SpecR StrepR;att φ80::pARA-CI857ts P c Engineered E. coli host cells having RNA-IN-SacB,tetR;ΔSbcDC are provided.
[0161] In some embodiments, the following genotype is used: DH5α dcm-att λ ::P c -RNA-IN-SacB,catR;att HK022 ::pL(OL1-G to T)P42L-P106I-F107S P113S(P3-),SpecR StrepR;att φ80 ::pARA-CI857ts P c Manipulated E. coli host cells having -RNA-IN-SacB,tetR;ΔSbcDC::kanR are provided.
[0162] In any of the embodiments described above, the SbcC gene may include a sequence having at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 9. In any of the embodiments described above, the SbcD gene may include a sequence having at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 10. It should be understood that this can be applied before or after knockout or knockdown, i.e., to genes in manipulated E. coli host cells. For reference, the wild-type sequence of SbcC from NCBI for E. coli K12 (reference sequence: WP_206061808.1) is:
[0163] The SbcD wild-type sequence (AAB18122.1) for E. coli K12, given by TIFF2026048673000005.tif80150, is from GenBank.
[0164] These are given by TIFF2026048673000006.tif36164. It should be understood that these amino acid sequences are illustrative, and those skilled in the art can identify the SbcC and SbcD genes, as well as proteins containing the complex, in other strains and cell lines based on homology.
[0165] In any of the embodiments described above, the sbcB gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 11. In any of the embodiments described above, the recB gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 12. In any of the embodiments described above, the recD gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 13. In any of the embodiments described above, the recJ gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 65.
[0166] In any of the embodiments described above, the uvrC gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 14. In any of the embodiments described above, the mcrA gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 15. In any of the embodiments described above, the mcrBC-hsd-mrr gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NOs: 16-21.
[0167] In any of the embodiments described above, the manipulated E. coli host cells may further contain a vector. For example, but not limited to, the vector may be a nonviral transposon vector, e.g., a transposase vector, Sleeping Beauty transposon vector, Sleeping Beauty transposase vector, PiggyBac transposon vector, PiggyBac transposase vector, expression vector, etc.; a nonviral gene editing vector, e.g., a homologous recombination repair (HDR) / CRISPR-Cas9 vector, etc.; or a viral vector, e.g., an AAV vector, AAV rep cap vector, AAV helper vector, Ad helper vector, lentiviral vector, lentiviral envelope vector, lentiviral packaging vector, retroviral vector, retroviral envelope vector, retroviral packaging vector, mRNA vector, etc.
[0168] In any of the above embodiments, in which E. coli host cells further include a vector, the vector may include a palindrome-containing nucleic acid sequence. A palindrome sequence can be understood as a nucleic acid sequence within a double-stranded DNA molecule in which a reading in a particular direction on the single strand coincides with a sequence reading in the opposite direction on the complementary strand, thereby creating complementary regions along the single strand with no intervening sequences between the complementary regions. For example, and not limited to, the complementary sequences of palindromes are approximately 10-200 base pairs, 15-200 base pairs, 20-200 base pairs, 25-200 base pairs, 30-200 base pairs, 40-200 base pairs, 50-200 base pairs, 75-200 base pairs, 100-200 base pairs, 15-200 base pairs, 10-150 base pairs, 15-150 base pairs, 20-150 base pairs, 25-150 base pairs, 30-150 base pairs, 30-150 base pairs, 40-150 base pairs, 50-150 base pairs, 100-150 base pairs, and 10-140 base pairs. Base pairs, approximately 15 to approximately 140 base pairs, approximately 20 to approximately 140 base pairs, approximately 25 to approximately 140 base pairs, approximately 30 to approximately 140 base pairs, approximately 30 to approximately 140 base pairs, approximately 40 to approximately 140 base pairs, approximately 50 to approximately 140 base pairs, approximately 100 to approximately 140 base pairs, approximately 10 to approximately 100 base pairs, approximately 15 to approximately 100 base pairs, approximately 20 to approximately 100 base pairs, approximately It may contain 25 to approximately 100 base pairs, approximately 30 to approximately 100 base pairs, approximately 40 to approximately 100 base pairs, approximately 50 to approximately 100 base pairs, or approximately 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 base pairs.
[0169] In any of the above embodiments in which an E. coli host cell further comprises a vector, the vector may comprise a nucleic acid sequence having at least one direct repeat. For example, but not limited to, the at least one direct repeat may comprise about 40 to 150 nucleotides, about 60 to about 120 nucleotides, or about 90 nucleotides. For example, but not limited to, the at least one direct repeat may be a simple repeat comprising a short sequence of DNA consisting of multiple repeats of a single base, such as a poly-A repeat, poly-T repeat, poly-C repeat, or poly-G repeat, and the simple repeat comprises about 40 to about 150 consecutive repeats of the same base, about 60 to about 120 consecutive repeats of the same base, or about 90 consecutive repeats of the same base. For example, but not limited to, the poly-A repeat may comprise 40 to 150 consecutive adenine nucleotides, 60 to 120 consecutive adenine nucleotides, or about 90 adenine nucleotides.
[0170] In any of the above embodiments in which E. coli host cells further comprise the vector, the vector may comprise a reverse repeat sequence, a direct repeat sequence, a homopolymer repeat sequence, a eukaryotic origin of replication, and a eukaryotic promoter-enhancer sequence. As a further example, the vector may comprise a sequence selected from the group consisting of polyA repeats, SV40 origin of replication, viral LTRs, lentiviral LTRs, retroviral LTRs, transposon IR / DR repeats, Sleeping Beauty transposon IR / DR repeats, AAV ITRs, CMV enhancers, and SV40 enhancers. For example, but not limited to, an AAV vector may comprise an AAV ITR. In some embodiments in which E. coli host cells further comprise the vector, the vector may comprise a nucleic acid sequence having at least one reverse repeat sequence, which may also be a reverse terminal repeat such as an AAV ITR, for example, but not limited to. Thus, in any of the above embodiments, the vector may comprise an AAV ITR. It should be understood that a reverse repeat sequence is a single-stranded sequence of nucleotides followed downstream by its reverse complement. It should be further understood that the single-stranded sequence may be part of a double-stranded vector. The nucleotide intervening sequence between the initial sequence and the reverse complement can be of any length, including zero. If the intervening length is zero, the composite sequence is a palindrome. If the intervening length is greater than zero, the composite sequence is a reverse repeat. In any of the embodiments described above, the intervening sequence may be 1 to about 2000 base pairs long.For example, although not limited to them, possible reverse repeats are approximately 1 to 2000 base pairs, 5 to 2000 base pairs, 10 to 2000 base pairs, 25 to 2000 base pairs, 50 to 2000 base pairs, 100 to 2000 base pairs, 250 to 2000 base pairs, 500 to 2000 base pairs, 750 to 2000 base pairs, 1000 to 2000 base pairs, 1250 to 2000 base pairs, 1500 to 2000 base pairs, 1750 to 2000 base pairs, and approximately 1 to 2000 base pairs. They can be separated by intervening sequences containing 100 base pairs, approximately 1 to 50 base pairs, approximately 1 to 25 base pairs, approximately 1 to 20 base pairs, approximately 1 to 10 base pairs, approximately 1 to 5 base pairs, or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 25, 50, 75, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, or 2000 base pairs. For example, though not limited to, the complementary parts of reverse repeats are approximately 10-200 base pairs, 15-200 base pairs, 20-200 base pairs, 25-200 base pairs, 30-200 base pairs, 40-200 base pairs, 50-200 base pairs, 75-200 base pairs, 100-200 base pairs, 15-200 base pairs, 10-150 base pairs, 15-150 base pairs, 20-150 base pairs, 25-150 base pairs, 30-150 base pairs, 30-150 base pairs, 40-150 base pairs, 50-150 base pairs, 100-150 base pairs, and 10-140 base pairs. Pairs, approximately 15-140 base pairs, approximately 20-140 base pairs, approximately 25-140 base pairs, approximately 30-140 base pairs, approximately 30-140 base pairs, approximately 40-140 base pairs, approximately 50-140 base pairs, approximately 100-140 base pairs, approximately 10-100 base pairs, approximately 15-100 base pairs, approximately 20-100 base pairs, approximately It may contain 25 to approximately 100 base pairs, approximately 30 to approximately 100 base pairs, approximately 40 to approximately 100 base pairs, approximately 50 to approximately 100 base pairs, or approximately 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 base pairs.For example, though not limited to, at least one reverse iteration is ttggccactccctctctgcgcgctcgctcgctcactgaggccgggcgaccaaaggtcgcccgacgcccgggctttgcccgggcggcct cagtgagcgagcgagcgcgcagagagggagtggccaactccatcactaggggttcct(5'AAV ITR) and aggaacccctagtgatggagttggccactccctctgcgcgctcgctcactgaggccgggcgaccaaaggtcgcccgacgccc gggctttgcccgggcggcctcagtgagcgagcgagcgcgcagagagggagtggccaa(3'AAV AAV iterators may include iterator repeats containing sequences that have at least 95%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with the iterator repeat.
[0171] Alternatively, in any of the above embodiments in which E. coli host cells further include a vector, the vector may not contain a nucleic acid sequence having a palindromic, direct repeat, or reverse repeat.
[0172] In any of the embodiments described above, the vector may be an AAV vector. In some embodiments where the vector is an AAV vector, the AAV vector includes an AAV ITR. In other embodiments, the vector may be a lentiviral vector, a lentiviral envelope vector, or a lentiviral packaging vector. In yet another embodiment, the vector may be a retroviral vector, a retroviral envelope vector, or a retroviral packaging vector. In yet another embodiment, the vector may be a transposase vector or a transposon vector. In yet another embodiment, the vector may be an mRNA vector. For example, but not limited to, the mRNA vector may include the poly-A repeats described herein.
[0173] In any of the embodiments described above, the vector may be a plasmid. In any of the embodiments described above, the vector may be a Rep protein-dependent plasmid.
[0174] In any of the embodiments described above, the vector may further include an RNA-selectable marker. For example, but not limited to, the RNA-selectable marker may be RNA-OUT. As further examples, though not limited to them, RNA-OUT includes sequence numbers 5 (gtagaattgg taaagagagt cgtgtaaaat atcgagttcg cacatcttgt tgtctgatta ttgatttttg gcgaaaccat ttgatcatat gacaagatgt gtatctacct taacttaatg attttgataa aaatcatta) and 7 (gtagaattgg taaagagagt tgtgtaaaat attgagttcg cacatcttgt tgtctgatta ttgatttttg gcgaaaccat ttgatcatat gacaagatgt gtatctacct taacttaatg attttgataa) in WO2019 / 183248 (sequence numbers 47 and 49, respectively). Sequences selected from the group consisting of aaatcatta) can have at least 95%, at least 98%, at least 99%, or 100% sequence identity. In some embodiments, the manipulated E. coli host cells may contain corresponding RNA-IN sequences to enable the regulation of downstream markers by RNA-OUT, where the RNA-OUT sequence corresponds to the RNA-IN.
[0175] In any of the embodiments described above, the vector may further include RNA-OUT antisense repressor RNA. For example, but not limited to, the RNA-OUT antisense repressor RNA may have a sequence that has at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with sequence number 6 of WO2019 / 183248 (sequence number 48).
[0176] In any of the embodiments described above, the vector may further include a bacterial origin of replication. For example, but not limited to, recent origins of replication may be selected from the group consisting of R6K, pUC, and ColE2. As further examples, though not limited to them, bacterial replication origins include sequence number 1 (ggcttgttgt ccacaaccgt taaaccttaa aagctttaaa agccttatat attctttttt ttcttataaa acttaaaacc ttagaggcta tttaagttgc tgatttatat taattttatt gttcaaacat gagagcttag tacgtgaaac atgagagctt agtacgttag ccatgagagc ttagtacgtt agccatgagg gtttagttcg ttaaacatga gagcttagta cgttaaacat gagagcttag tacgtactat caacaggttg aactgctgat c) and sequence number 2 (ggcttgttgt ccacaaccat taaaccttaa aagctttaaa agccttatat attctttttt ttcttataaa acttaaaacc ttagaggcta tttaagttgc tgatttatat taattttatt gttcaaacat gagagcttag tacgtgaaac atgagagctt agtacattag ccatgagagc ttagtacatt agccatgagg gtttagttca ttaaacatga gagcttagta cattaaacat gagagcttag tacatactat caacaggttg aactgctgat c), sequence number 3 (aaaccttaaa acctttaaaa gccttatata ttcttttttt tcttataaaa cttaaaacct tagaggctat ttaagttgct gatttatatt aattttattg ttcaaacatg agagcttagt acatgaaaca tgagagctta gtacattagc catgagagct tagtacatta gccatgagggtttagttcat taaacatgag agcttagtac attaaacatg agagcttagt acatactatc aacaggttga actgctgatc), SEQ ID NO: 4 (tgtcagccgt taagtgttcc tgtgtcactg aaaattgctt tgagaggctc taagggcttc tcagtgcgtt acatccctgg cttgttgtcc acaaccgtta aaccttaaaa gctttaaaag ccttatatat tctttttttt cttataaaac ttaaaacctt agaggctatt taagttgctg atttatatta attttattgt tcaaacatga gagcttagta cgtgaaacat gagagcttag tacgttagcc atgagagctt agtacgttag ccatgagggt ttagttcgtt aaacatgaga gcttagtacg ttaaacatga gagcttagta cgtgaaacat gagagcttag tacgtactat caacaggttg aactgctgat cttcagatc), and SEQ ID NO: 18 (ggcttgttgt ccacaaccgt taaaccttaa aagctttaaa agccttatat attctttttt ttcttataaa acttaaaacc ttagaggcta tttaagttgc tgatttatat taattttatt gttcaaacat gagagcttag tacgtgaaac atgagagctt agtacgttag ccatgagagc ttagtacgtt agccatgagg gtttagttcg ttaaacatga gagcttagta cgttaaacat gagagcttag tacgttaaac atgagagctt agtacgtact atcaacaggt tgaactgctgThis could be an R6K gamma replication origin having at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with atc), such as SEQ ID NO: 30 (ColE2 origin (+7), 45 bp), SEQ ID NO: 31 (ColE2 origin (+7, CpG-free), 45 bp), SEQ ID NO: 32 (ColE2 origin (Min), 38 bp), SEQ ID NO: 33 (ColE2 origin (+16), 60 bp), and SEQ ID NO: 22 (pUC, 784 bp).
[0177] In any of the embodiments described above, the engineered E. coli host cell may further include a eukaryotic pUC-free minicircle expression vector comprising: (i) a eukaryotic region sequence encoding the gene of interest and having 5' and 3' ends; and (ii) a spacer region having less than 1000, preferably less than 500, base pairs in length, ligating the 5' and 3' ends of the eukaryotic region sequence and containing an R6K bacterial origin and RNA-OUT selectable marker. For example, but not limited to, the R6K bacterial origin and RNA-OUT selectable marker may have sequences described in this disclosure and known in the art. Alternatively, in any of the embodiments described above, the engineered E. coli cell may further include a covalently bound closed circular plasmid having a backbone comprising a backbone containing a PolIII-dependent R6K origin and RNA-OUT selectable marker, having a backbone less than 1000 bp, preferably less than 500 bp, and an insert containing a structured DNA sequence. For example, but not limited to, structured DNA sequences may include sequences selected from the group consisting of reversed repeat sequences, direct repeat sequences, homopolymer repeat sequences, eukaryotic origins of replication, and eukaryotic promoter-enhancer sequences. Further examples may include sequences selected from the group consisting of polyA repeats, SV40 origins of replication, viral LTRs, lentiviral LTRs, retroviral LTRs, transposon IR / DR repeats, Sleeping Beauty transposon IR / DR repeats, AAV ITRs, CMV enhancers, and SV40 enhancers. For example, but not limited to, inserts may be transposase vectors, AAV vectors, or lentiviral vectors. For example, but not limited to, polIII-dependent R6K replication origins may have sequences that have at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NOs. 43, 44, 45, 46, and 60 (from SEQ ID NOs. 1-4 and 18 of WO2019 / 183248).As an example, but not limited to, an RNA-OUT selectable marker may be an RNA-IN regulated RNA-OUT functional variant having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 47 or SEQ ID NO: 49 (from SEQ ID NOs: 5 and 7 of WO2019 / 183248). As a further example, an RNA-OUT selectable marker may be an RNA-OUT antisense repressor RNA. As an example, but not limited to, an RNA-OUT antisense repressor RNA may have a sequence having at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 6 of WO2019 / 183248 (SEQ ID NO: 48).
[0178] It should be understood that a viability or yield reduction mutation refers to a mutation that reduces the viability or yield of the cell line from which the mutated cell line originates, under the same culture conditions. It should also be understood that such mutations can be manipulated or occur spontaneously.
[0179] Methods for knocking out or knocking down genes, as disclosed herein, are known in the art and include, but are not limited to, the methods disclosed in the examples herein (recombinant), as well as P1 phage transduction, genomic material transfer, and CRISPR / Cas9. It should be understood that gene knockout can result in either the cessation of protein expression or the expression of a non-functional protein. Thus, the SbcCD complex may or may not be present in the bacterial host strains of this disclosure, but if present, it will be non-functional in the case of knockout, or its activity as a nuclease will be reduced in the case of knockdown. It should be understood that embodiments of this disclosure may include knockout or knockdown of SbcC, SbcD, or both.
[0180] While not theoretically bound, knockout of SbcC or SbcD alone is expected to be sufficient to achieve the desired effects of the present invention, since both proteins are essential subunits of the SbcCD nuclease (Connelly JC and Leach DR, Genes Cells 1:285, 1996). The sbcC and sbcD genes of E. coli encode nucleases involved in palindromic inactivation and recombination (Connelly JC and Leach DR, Genes Cells 1:285, 1996).
[0181] Within this disclosure, it should be understood that manipulated E. coli host cells may contain vectors such as those described herein. The vectors may include any suitable vectors, including those described in the references incorporated herein by reference. For example, in some cases, the vector may contain a structured DNA sequence. In other cases, the vector may not contain a structured DNA sequence.
[0182] In some embodiments, the engineered E. coli host cells may further include vectors as understood herein. Such vectors may occur spontaneously or be engineered. The vectors included in the engineered E. coli host cells of this disclosure may include any of the features considered herein and in documents incorporated by reference. The vectors included in the engineered E. coli host cells of this disclosure may not include, for example, at least one reversed repeat, direct repeat, or any of the aforementioned structured DNA sequences, such as reversed terminal repeats or palindromes.
[0183] Method for producing manipulated E. coli host cells In some embodiments, a method for producing engineered E.coli host cells is provided, comprising the step of obtaining engineered E.coli cells by knocking out at least one gene selected from the group consisting of SbcC and SbcD in starting E.coli cells that do not contain any engineered viability or yield reduction mutations in any of sbcB, recB, recD, and recJ. In some embodiments, a method for producing engineered E.coli host cells is provided, comprising the step of obtaining engineered E.coli host cells by knocking out at least one gene selected from the group consisting of SbcC and SbcD in starting E.coli cells that do not contain any engineered mutations in any of sbcB, recB, recD, and recJ. In some embodiments, a method for producing engineered E.coli host cells is provided, comprising the step of obtaining engineered E.coli host cells by knocking out at least one gene selected from the group consisting of SbcC and SbcD in starting E.coli cells that do not contain any mutations in any of sbcB, recB, recD, and recJ.
[0184] In any of the embodiments described above, the starting E. coli cells may further be free of any manipulated viability or yield reduction mutations in at least one of uvrC, mcrA, mcrBC-hsd-mrr, and combinations thereof. In any of the embodiments described above, the starting E. coli cells may further be free of any mutations in at least one of uvrC, mcrA, mcrBC-hsd-mrr, and combinations thereof. In any of the embodiments described above, the starting E. coli cells may further be free of any mutations in at least one of uvrC, mcrA, mcrBC-hsd-mrr, and combinations thereof.
[0185] In any of the embodiments described above, the step of knocking out at least one gene does not result in any mutations in sbcB, recB, recD, and recJ. In any of the embodiments described above, the step of knocking out at least one gene does not result in any mutations in at least one of uvrC, mcRA, mcrBC-hsd-mrr, and any combination thereof.
[0186] In any of the embodiments described above, the manipulated E. coli cells may further be free from the manipulated viability or yield reduction mutation in at least one of uvrC, mcrA, mcrBC-hsd-mrr, and combinations thereof. In any of the embodiments described above, the manipulated E. coli host cells may further be free from the manipulated mutation in at least one of uvrC, mcrA, mcrBC-hsd-mrr, and combinations thereof. In any of the embodiments described above, the manipulated E. coli host cells may be free from any mutation in at least one of uvrC, mcrA, mcrBC-hsd-mrr, and combinations thereof.
[0187] In any of the embodiments described above, the manipulated E. coli host cells cannot contain the manipulated viability or yield reduction mutations in sbcB, recB, recD, and recJ. In any of the embodiments described above, the manipulated E. coli host cells cannot contain the manipulated mutations in sbcB, recB, recD, and recJ. In any of the embodiments described above, the manipulated E. coli host cells cannot contain any mutations in sbcB, recB, recD, and recJ.
[0188] In any of the embodiments described above, the manipulated E. coli host cells do not contain the functional SbcCD complex. In any of the embodiments described above, the manipulated E. coli host cells do not produce the SbcCD complex. Alternatively, in some embodiments, the manipulated E. coli host cells produce the non-functional SbcCD complex.
[0189] It should be understood that in any embodiment of the method described above, the manipulated E. coli host cell may be any of the E. coli host cells described herein.
[0190] In any of the embodiments described above, the SbcC gene may include a sequence having at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 9. In any of the embodiments described above, the SbcD gene may include a sequence having at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 10. It should be understood that this can be applied before or after knockout or knockdown, i.e., to genes in manipulated E. coli host cells.
[0191] In any of the embodiments described above, the sbcB gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 11. In any of the embodiments described above, the recB gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 12. In any of the embodiments described above, the recD gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 13. In any of the embodiments described above, the recJ gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 65.
[0192] In any of the embodiments described above, the uvrC gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 14. In any of the embodiments described above, the mcrA gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 15. In any of the embodiments described above, the mcrBC-hsd-mrr gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NOs: 16-21.
[0193] Methods for vector production In some embodiments, an improved method for vector production is provided, comprising the steps of transfecting engineered E. coli host cells with a vector to obtain transfected host cells, and incubating the transfected host cells under conditions sufficient to replicate the vector, wherein the E. coli host cells do not contain any engineered viability or yield reduction mutations in any of sbcB, recB, recD, and recJ. It should be understood that the vector used to transfect the engineered E. coli host cells may be any vector described in this disclosure, including embodiments disclosed in which the engineered E. coli host cells include the vector.
[0194] In some embodiments, a method for improved vector production is provided, comprising the steps of: incubating transfected host cells which are engineered E. coli host cells that include a vector and do not contain engineered viability or yield reduction mutations in any of sbcB, recB, recD, and recJ which include a vector; and incubating the transfected host cells under conditions sufficient to replicate the vector.
[0195] It should be understood that in any of the embodiments described above, the manipulated E. coli host cell may be any of the manipulated E. coli host cells of this disclosure.
[0196] In any of the embodiments described above, the method may further include isolating the vector from the transfected host cells.
[0197] In any of the embodiments described above, the step of incubating the transfected host cells is carried out by fed-batch fermentation, which is performed by feeding the transfected cells or by feeding the vector after transfection, and the fed-batch fermentation includes growing the engineered E. coli host cells by growing them at a reduced temperature (which may be under growth-limiting conditions) during a first part of the fed-batch phase, and then increasing the temperature to a higher temperature during a second part of the fed-batch phase. For example, the reduced temperature may be about 28–30°C, and the higher temperature may be about 37–42°C. For example, the first part may be about 12 hours, and the second part may be about 8 hours. When fed-batch fermentation with temperature increase is used, the engineered E. coli host cells may be regulated by a temperature-sensitive lambda repressor. L It should be understood that lambda repressors and Rep proteins can be present under the control of the promoter.
[0198] In any of the embodiments described above, the plasmid yield after incubation of host cells transfected under conditions sufficient to replicate the vector can be higher than that of cell lines derived from manipulated E. coli cells treated under the same conditions.
[0199] In any of the embodiments described above, the SbcC gene may include a sequence having at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 9. In any of the embodiments described above, the SbcD gene may include a sequence having at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 10. It should be understood that this can be applied before or after knockout or knockdown, i.e., to genes in manipulated E. coli host cells.
[0200] In any of the embodiments described above, the sbcB gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 11. In any of the embodiments described above, the recB gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 12. In any of the embodiments described above, the recD gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 13. In any of the embodiments described above, the recJ gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 65.
[0201] In any of the embodiments described above, the uvrC gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 14. In any of the embodiments described above, the mcrA gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 15. In any of the embodiments described above, the mcrBC-hsd-mrr gene may include a sequence having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NOs: 16-21.
[0202] In any of the embodiments described above, it should be understood that the vector transfected into the manipulated E. coli host cells may be any of the vectors described herein.
[0203] It should be understood that in any of the embodiments described above, the manipulated E. coli host cells may include knockdown of SbcC, SbcD, or both, rather than knockout. Knockdown may result in reduced expression and / or activity of the SbcCD complex. The reduction may be at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or more.
[0204] The bacterial host strains and methods of this disclosure are described here with reference to the following non-limiting examples. [Examples]
[0205] Most therapeutic plasmids utilize pUC origins (closely related to ColE1 origins), which are high-copy derivatives of pMB1 origins. For pMB1 replication, plasmid DNA synthesis is unidirectional and does not require plasmid-retaining initiator proteins. pUC origins are copy-up derivatives of pMB1 origins that delete accessory ROP(rom) proteins and possess additional temperature-sensitive mutations that destabilize RNAI / RNAII interactions. Plasmid copy number increases by shifting cultures containing these origins from 30°C to 42°C. pUC plasmids can be produced in a wide variety of E. coli cell lines.
[0206] In the following examples, a proprietary plasmid + shaking culture medium was used for shaking flask production. Seed cultures were started from glycerol stocks or colonies and streaked onto LB agar plates containing 50 μg / mL of antibiotic (ampR or kanR selective plasmid) or 6% sucrose (RNA-OUT selective plasmid). The plates were grown at 30–32°C, and the cells were resuspended in medium at approximately 2.5 OD. 600 The inoculum was used to prepare a 500 mL plasmid-shaking flask. This flask contained 50 μg / mL of antibiotic for ampR or kanR selection plasmids, or 0.5% sucrose for selection for RNA-OUT plasmids. The flask was grown with shaking until saturated at the growth temperature as shown.
[0207] In the following examples, HyperGRO fermentation was performed in a New Brunswick BioFlo 110 bioreactor using proprietary fed-batch medium (NTC3019, HyperGRO medium) as described (the entire text is incorporated herein by reference, U.S. Patent No. 7,943,377). Seed cultures were started from glycerol stocks or colonies and streaked onto LB agar plates containing 50 μg / mL of antibiotic (ampR or kanR selective plasmid) or 6% sucrose (RNA-OUT selective plasmid). Plates were grown at 30–32°C, and cells were resuspended in medium and used to provide approximately 0.1% inoculum for fermentation, which contained 50 μg / mL of antibiotic for ampR or kanR selective plasmids, or 0.5% sucrose for RNA-OUT plasmids. HyperGRO temperature shifts were as shown.
[0208] In the following examples, culture samples were taken at key points and regular intervals throughout all fermentation. The samples were immediately treated with biomass (OD). 600The plasmid yield was analyzed. When plasmid yield was determined, the analysis was performed by quantifying plasmids obtained from Qiagen Spin Miniprep Kit preparations, as described in U.S. Patent No. 7,943,377. In short, cells were alkaline lysed, clarified, the plasmids were column-purified, eluted, and then quantified. Plasmid quality was determined by agarose gel electrophoresis (AGE), performed on 0.8–1% Tris / acetate / EDTA (TAE) gels, as described in U.S. Patent No. 7,943,377.
[0209] The strains used in the following examples included the following:
[0210] Background of RNA-QUT antibiotic-free selectable markers: Antibiotic-free selection is performed in E. coli strains containing phage-lambda binding site chromosome-integrated pCAH63-CAT RNA-IN-SacB(P5 / 6 6 / 6), such as NTC4862, as described, for example, in WO2008 / 153733. SacB (Bacillus subtilis levans sucrase) is a reverse-selectable marker that is lethal to E. coli cells in the presence of sucrose. Translation of SacB from RNA-IN-SacB transcripts is inhibited by plasmid-coding RNA-OUT. This promotes plasmid selection by inhibiting SacB-mediated lethality in the presence of sucrose.
[0211] Background of R6K origin vector replication: The R6K gamma plasmid replication origin requires a single plasmid replication protein π that binds to multiple repeat "iteron" sites (seven core repeats containing the TGAGNG consensus) as replication initiation monomers, and to the repression site (TGAGNG) and a reduced-affinity iterone as replication inhibitory dimers. Replication requires multiple host factors, including IHF, DnaA, and primosome assembly proteins DnaB, DnaC, and DnaG (Abhyankar et al., 2003 J Biol Chem 278:45476-45484). The R6K core origin contains DnaA and IHF binding sites that affect plasmid replication, as π, IHF, and DnaA interact to initiate replication.
[0212] Different versions of the R6Kγ replication origin have been utilized in various eukaryotic expression vectors, such as the pCOR vector (Soubrier et al., 1999, Gene Therapy 6:1482-88), the pCpGfree vector (Invivogen, San Diego CA), and the CpG-free version in pGM169 (University of Oxford). Six highly minimized iteron-R6Kgamma-derived replication origins containing the core sequence necessary for replication (including the DnaA box and stb1-3 sites; Wu et al., 1995, J Bacteriol 177:6338-6345), with the upstream π-dimer repressor binding site and downstream π-promoter deleted (by removing one copy of iteron), are described in WO2014 / 035457 and are incorporated herein by reference (SEQ ID NO: 1 from WO2019 / 183248 (SEQ ID NO: 43)). This R6K origin contains six tandem direct repeat iterones. The NTC9385R Nanoplasmid™ vector, containing this minimized R6K origin and an RNA-OUT AF (antibiotic-free) selectable marker in the spacer region, is described in WO2014 / 035457 and is incorporated herein by reference. R6K origins containing seven tandem direct repeat iterones and R6K origins containing six tandem direct repeat iterones and a single CpG residue are described in WO2019 / 183248 and are incorporated herein by reference. The use of conditional replication origins such as R6Kγ, which require cell lines specialized for propagation, adds a safety margin because the vector will not replicate if it migrates to the patient's endogenous microbiota.
[0213] Typical R6K-producing strains express from the genome a π protein derivative PIR116 that contains the P106L substitution that increases copy number (by reducing π dimerization; the π monomer is activating, while the π dimer is inhibitory). Fermentation results with pCOR (Soubrier et al. (supra) 1999) and pCpG plasmids (Hebei HL, Cai Y, Davies LA, Hyde SC, Pringle IA, Gill DR. 2008. Mol Ther 16:S110) were low, about 100 mg / L in the PIR116 cell line.
[0214] Mutagenesis of the pir-116 replication protein and selection for increased copy number have been used to generate new producing strains. For example, the TEX2pir42 strain contains a combination of P106L and P42L. The P42L mutation impedes DNA loop replication repression. The TEX2pir42 cell line improved copy number and fermentation yield with the pCOR plasmid, which was reported to have a yield of 205 mg / L (see Soubrier F. 2004. International Patent Application No. W02004 / 033664).
[0215] Other combinations of π copy number variants that improve copy number include "P42L and P113S" and "P42L, P106L and F107S" (Abhyankar et al., 2004. J Biol Chem 279:6711-6719).
[0216] WO2014 / 035457 describes a host strain that expresses a phage HK022 binding site integrated pL promoter heat-inducible πP42L, P106L and F107S high copy variant replication (Rep) protein for selection and propagation of an R6K origin Nanoplasmid™ vector.
[0217] Propagation and fermentation of the RNA-OUT selectable marker-R6K plasmid described in WO2014 / 035457 was in the DH5α host strain NTC711772 = DH5α dcm-att λ ::Pc -RNA-IN-SacB,catR;att HK022 The study was conducted using thermoinducible "P42L, P106L, and F107S" π copy number mutant cell lines such as ::pL(OL1-G to T)P42L-P106L-F107S(P3-),SpecR StrepR. A maximum production yield of 695 mg / L was reported.
[0218] Further R6K-based "copy cutter" host cell lines were created and disclosed in Williams 2019 VIRAL AND NON-VIRAL NANOPLASMID VECTORS WITH IMPROVED PRODUCTION, International Patent Application No. WO2019 / 183248, which is NTC1050811 DH5α att λ ::P c -RNA-IN-SacB,catR;att HK022 ::pL(OL1-G to T)P42L-P106I-F107S P113S(P3-),SpecR StrepR;att φ80::Contains pARA-CI857ts, tetR = pARA-CI857ts (a derivative of NTC940211). This "copy cutter" host strain contains a chromosomal integrated copy of the phage φ80 attachment site of the arabinose-inducible CI857ts gene. Addition of arabinose to the plate or medium (e.g., up to a final concentration of 0.2 - 0.4%) induces the expression of the pL promoter. Induces pARA-mediated CI857ts repressor expression that reduces the copy number at 30°C through CI857ts-mediated downregulation of the Rep protein [i.e., additional CI857ts that mediates more effective downregulation of the pL(OL1-G~T) promoter at 30°C]. Copy number induction after a temperature shift to 37 - 4°C is not impaired because the CI857ts repressor is inactivated at these elevated temperatures. The dcm-derivative (NTC1050811 dcm-) is used when dcm methylation is not desired. NTC1050811-HF is a derivative of the NTC1050811 cell line that contains a second copy of the RNA-IN-SacB expression cassette and has no mutations in sbcB, recB, recD, recJ, uvrC, mcrA or mcrBC-hsd-mrr.
[0219] In each case, both strains (NTC1050811 and NTC1050811-HF) contain a phage φ80 attachment site chromosome-integrated copy of the arabinose-inducible CI857ts gene. Addition of arabinose to the plate or culture medium (e.g., up to a final concentration of 0.2–0.4%) induces the expression of the pL promoter. This induces pARA-mediated CI857ts repressor expression, which reduces the copy number at 30°C through CI857ts-mediated downregulation of the Rep protein [i.e., additional CI857ts mediating a more effective downregulation of the pL(OL1-G~T) promoter at 30°C]. Copy number induction after a temperature shift to 37–42°C is not impaired because the CI857ts repressor is inactivated at these high temperatures. These “copy cutter host strains” increase the R6K vector temperature upshift copy number induction ratio by reducing the copy number at 30°C. This is advantageous for the production of large, toxic, or easily dimerizable R6K-based vectors.
[0220] Nanoplasmid® production yields are improved in the quadruple mutant heat-inducible pL(OL1-G to T)P42L-P106I-F107S P113S(P3-) described in WO2019 / 183248 compared to the triple mutant heat-inducible pL(OL1-G to T)P42L-P106L-F107S(P3-) described in WO2014 / 035457. Yields exceeding 2 g / L of Nanoplasmid® have been obtained in the quadruple mutant NTC1050811 cell line (WO2019 / 183248).
[0221] The use of conditional origins of replication, such as R6K origins, which require cell lines specifically designed for propagation, adds a safety margin because the vector will not replicate if it migrates to the patient's endogenous microbiome.
[0222] The RNA-OUT-producing host described in WO2019 / 183248 has been modified to create an HF host. SacB (Bacillus subtilis levans sucrase) is a reverse-selectable marker that is lethal to E. coli cells in the presence of sucrose. Translation of SacB from the RNA-IN-SacB transcript is inhibited by plasmid-coding RNA-OUT. This promotes plasmid selection by inhibiting SacB-mediated lethality in the presence of sucrose. Mutations in the chromosomal copy of the RNA-IN-SacB expression cassette that eliminate SacB expression are sucrose-resistant (in the absence of the plasmid). The presence of a second copy of the RNA-IN-SacB expression cassette dramatically reduces the number of sucrose-resistant colonies (in the absence of the plasmid), as each individual RNA-IN-SacB expression cassette copy mediates sucrose lethality in the absence of the plasmid, and both RNA-IN-SacB expression cassette chromosomal copies mediate sucrose lethality in the absence of very rare mutations in the plasmid.
[0223] NTC1011592 Stbl4 attk::P c -RNA-IN-SacB,catR(WO2019 / 183248) was also used.
[0224] The following examples include unmodified production strains: DH5α, Sure2, Stbl2, Stbl3, or Stbl4.
[0225] Example 1: Preparation of SbcCD knockout strain SbcCD knockout strains were generated using Red Gam recombinant cloning, as described in Datsenko and Wanner, PNAS USA 97:6640-6645 (2000). The pKD4 plasmid (Datsenko and Wanner, 2000) was PCR amplified with the following primers to introduce SbcC and SbcD targeting the homology arm. Sequence ID 1 (SbccR-pKD4): CCCTCTGTATTCATTATCCTGCTGAATAGTTATTTCACTGCAAACGTACTCATATGAATATCCTCCTTAG Sequence ID 2 (SbcdF-pKD4): TCTGTTTGGGTATAATCGCGCCCATGCTTTTTCGCCAGGGAACCGTTATGTGTAGGCTGGAGCTGCTTCG
[0226] 1.6kb PCR product (SEQ ID NO: 5,
[0227] TIFF2026048673000007.tif131165) (Figure 1A) was purified, and DpnI was digested (to eliminate the template plasmid). Host lines in which the SbcCD gene was knocked out were transformed with the pKD46-RecApa recombinant plasmid (WO2008 / 153731, the whole plasmid is incorporated herein by reference) and transformants selected for ampicillin resistance. Electrocompetent cells of the transformed cell lines were given an OD of approximately 0.05. 600 Cells were prepared by growth in LB medium containing 50 μg / mL ampicillin, recombinant gene expression was induced by adding 0.2% arabinose, and the cells were grown to an intermediate logarithmic phase. Electrocompetent cells were then prepared by centrifugation and resuspension in 10% glycerol at 1 / 200 of the original volume. 5 μL of DpnI digestion-purified PCR product was electroporated into 25 μL of electrocompetent cells, followed by the addition of 1 mL of SOC medium. The cells were grown at 30°C for 2 hours, plated on LB agar plates containing 20 μg kanamycin, and grown overnight at 37°C. Individual kanR colonies were screened for ΔSbcDC::kanR using SbcDF and SbcCR primers as described below. Sequence ID 3 (SbcDF primer): cgtctcgccatgatttgccctg Sequence ID 4 (SbcCR primer): cgttatgcgccagctccgtgag Host: Products of SbcDF and SbcCR primers = 4.8kb (Figure 1B) (SEQ ID NO: 6)
[0228] TIFF2026048673000008.tif102149
[0229] TIFF2026048673000009.tif219149
[0230] TIFF2026048673000010.tif168149) Host ΔSbcDC::kanR:SbcDF and SbcCR primer products = 1.9kb (Figure 1C) (SEQ ID NO: 7)
[0231] TIFF2026048673000011.tif36150
[0232] TIFF2026048673000012.tif138150)
[0233] The temperature-sensitive pKD46-recApa plasmid was cured from cell lines by growing it at 37–42°C. Ampicillin sensitivity of individual kanR colonies was also examined.
[0234] For host lines of antibiotic-resistant plasmids (e.g., pUC origin; antibiotic selection; R6K origin; antibiotic selection), the kanR chromosome marker was removed from ΔSbcDC::kanR using FRT recombination, as described (Datsenko and Wanner, (above), 2000). Briefly, ΔSbcDC::kanR cell lines were grown with the pCP20 FRT plasmid (Datsenko and Wanner, (above), 2000) at 30°C and transformed with transformants selected for ampicillin resistance. Individual colonies were streaked (without ampicillin) on LB medium plates and grown at 43°C to cure the temperature-sensitive pCP20 plasmid. Single colonies from 43°C LB plates were streaked onto LB amp and LB kan plates to verify the loss of ampR pCP20 plasmid and kanR excision, respectively. Individual amp and kan-sensitive colonies were screened for ΔSbcDC by PCR using SbcDF and SbcCR primers (Figure 1D). The PCR products for SbcDF and SbcCR primers were 0.53 kb in size, as shown in Figure 1D (SEQ ID NO: 8).
[0235] Regarding DH5α, the starting strain has the following genotype: F-φ80lacZΔM15 Δ(lacZYA-argF)U169 recA1 endA1 hsdR17(r k -,m k +)gal-phoA supE44 λ-thi-1 gyrA96 relA1 was present. After SbcCD knockout and kanR excision, the knockout strain (DH5α[SbcCD-]) had the following genotype: F-φ80lacZΔM15 Δ(lacZYA-argF)U169 recA1 endA1 hsdR17(r k -,m k It has +)gal-phoA supE44 λ-thi-1 gyrA96 relA1 ΔSbcDC.
[0236] As described in WO2014 / 035457, by integrating the thermoinducible R6K Rep protein cassette (att HK022 ::pL(OL1-G to T)P42L-P106I-F107S P113S(P3-),SpecR StrepR) into the host genome, additional strains were produced from DH5α[SbcD-] to obtain a new strain, DH5α R6K Rep[SbcCD-], which has the genotype: DH5α att HK022 ::pL(OL1-G to T)P42L-P106I-F107S P113S(P3-),SpecR StrepR;ΔSbcDC. This strain can be used for the production of plasmids having the R6K bacterial replication origin.
[0237] R6K replication origin by RNA-OUT selection. Furthermore, the genotype DH5α att disclosed in WO2019 / 183248 λ ::P c -RNA-IN-SacB,catR;att HK022 ::pL(OL1-G to T)P42L-P106I-F107S P113S(P3-),SpecR StrepR;att φ80 ::pARA-CI857ts,tetR having NTC1050811 was also treated in the same way as knocking out SbcDC, but without kanR excision, DH5α att λ ::P c -RNA-IN-SacB,catR;att HK022 ::pL(OL1-G to T)P42L-P106I-F107S P113S(P3-),SpecR StrepR;att φ80We obtained NTC1300441 (DH5αΔSbcDC) having the genotype ::pARA-CI857ts,tetR ΔSbcDC::kanR (SbcCD knockout copy cutter host strain derivative). NTC1050811-HF, a derivative of NTC1050811 containing a second copy of the RNA-IN-SacB expression cassette without mutations in sbcB, recB, recD, recJ, uvrC, and mcrA, was also used by the same method to generate a knockout strain, obtaining NTC1050811-HF[SbcCD-] which does not excise kanR.
[0238] pUC replication origins were determined by RNA-OUT selection. Furthermore, using NTC4862-HF, a derivative of NTC4862 disclosed in WO2008 / 153733 that contains a second copy of the RNA-IN-SacB expression cassette and has no mutations in sbcB, recB, recD, recJ, uvrC, and mcrA, a knockout strain was generated by the same method to obtain NTC4862-HF[SbcCD-] which does not excise kanR.
[0239] Example 2: Performance of SbcCD knockout strains using large palindromic vectors The performance of SbcCD knockout strains was evaluated using large palindromic vectors, including assessment of shaking flasks and HyperGRO production.
[0240] NTC1011641 (Genotype: Stbl4 att) λ ::P c -RNA-IN-SacB,catR;att HK022 ::pL P42L-P106L-F107S(P3-)SpecR StrepR (disclosed in WO2019 / 183248) and NTC1300441 (genotype: DH5α att λ ::P c -RNA-IN-SacB,catR;att HK022 ::pL(OL1-G to T)P42L-P106I-F107S P113S(P3-),SpecR StrepR;att φ80::pARA-CI857ts,tetR ΔSbcDC::kanR) was transformed with the AAV vector pAAV-GFP Nanoplasmid® (pAAV-GFP NP), which includes a spacer region containing a palindromic AAV ITR and a pAAV-GFP mini-intron plasmid (pAAV-GFP MIP), as well as an intronic R6K bacterial origin and RNA-OUT selection, and 140 base pair reverse repeats with a 4 base pair intervening sequence.
[0241] Lu J, Williams JA, Luke J, Zhang F, Chu K, and Kay MA. 2017. Human Gene Therapy 28:125-34 discloses antibiotic-free miniintron plasmid (MIP) AAV vectors, suggesting that MIP intron AAV vectors can be modified to produce shorter AAV vectors by removing the vector skeleton. Attempts to create mini-circle-like spacer regions in miniintron plasmid AAV vectors with an intron R6K origin and an RNA-OUT selection marker (intron nanoplasmid vector) were presumed to be toxic due to the creation of long 140 bp reverse repeats by such close juxtaposition of AAV ITRs (e.g., pAAV-GFP MIP; see Table 2). In contrast, pAAV-GFP MIP was recoverable in the DH5α ΔSbcDC host strain and exhibited excellent shaking flask production yields (see Table 2). Each AAV ITR contained a 26 bp palindromic sequence separated at 43 bp. [Table 2]
[0242] This recovery of viability in the DH5α ΔSbcDC host strain is not limited to Nanoplasmid® vectors. This is demonstrated by the robust proliferation and HyperGRO plasmid production of a pUC-originating kanR-selective AAV helper plasmid containing an 85 bp reverse repeat with a 17 base pair intercalation sequence in DH5α ΔSbcDC, but not in DH5α (Table 3). [Table 3]
[0243] Example 3: Performance of SbcCD knockout strain using AAV ITR vector: ITR stability and shaking flask production The application of DH5α ΔSbcDC host lines to stabilize AAV ITR-containing vectors was evaluated by next-generation sequence verification of AAV vector-transformed cell lines and production lots.
[0244] AAV ITR allows for accurate sequencing using next-generation sequencing, compared to conventional sequencing (Doherty et al., (above), 1993) (Saveliev A Liu J, Li M, Hirata L, Latshaw C, Zhang J, Wilson JM. 2018. Accurate and rapid sequence analysis of Adeno-Associated virus plasmid by Illumina Next Generation Sequencing. Hum Gene Ther Methods 29:201-211).
[0245] To evaluate the DH5α ΔSbcDC host strain and stabilize AAV ITR, nine different AAV ITR nanoplasmid vectors ranging from 2.4 to 5.4 kb were transformed into NTC1050811-HF[SbcCD-]. Individual colonies were screened for intact ITR by Small digestion, and a single correct clone was then submitted to the Mass General Hospital (MGH) CCIB DNA Core (Cambridge MA) for complete plasmid sequencing by next-generation sequencing. The results are summarized in Table 4 below, showing ITR stability during transformation (25 / 26 colonies screened correct by Small digestion; 9 / 10 of these (one from each of the nine nanoplasmid vectors) were demonstrated to be correct by complete plasmid sequencing). ITR stability was maintained during preparation in a shaking flask (5 / 5 preparations correct by complete plasmid sequencing). This demonstrates that the DH5α ΔSbcDC host strain stabilizes AAV ITR during transformation and production. [Table 4]
[0246] Next, we evaluated the application of DH5α ΔSbcDC host strains to improve AAV ITR-containing vector production using standardized GFP AAV2 EGFP transgene vectors with different bacterial skeletons. pUC-initiated antibiotic-selective AAV vectors (Table 5); pUC-starting RNA-OUT selective AAV vector (Table 6); or R6K-starting RNA-OUT selective AAV nanoplasmid vector (Table 7) [Table 5] [Table 6] [Table 7]
[0247] An additional panel of three larger 4.8–5.2 kb AAV nanoplasmid vectors was evaluated in Stbl4 compared to the DH5α SbcCD NP host (Table 8). Dramatic improvements in yield and quality were observed using the DH5α SbcCD host. [Table 8]
[0248] Summary: DH5α SbCD hosts showed improved plasmid production and / or plasmid quality compared to Stbl4 hosts with AAV ITR vectors, particularly compared to Stbl4 hosts with larger therapeutic transgenes encoding AAV ITR vectors (Table 8).
[0249] Example 4: Performance of SbcCD knockout strain using AAV ITR vector: HyperGRO fermentation Next, the improvement in AAV ITR-containing vector production by applying the DH5α ΔSbcDC host strain was evaluated in HyperGRO fermentation using a 3.3kb AAV2 EGFP transgene R6K origin-RNA-OUT marker Nanoplasmid vector pAAV-GFP Nanoplasmid (evaluated in the shaking flask of Example 3) in the DH5α ΔSbcDC Nanoplasmid host compared to the Stbl4 Nanoplasmid host, and a 12kb pUC origin-kanR AAV vector in DH5α ΔSbcDC compared to stbl3. The results are summarized in Tables 9 and 10. [Table 9] [Table 10]
[0250] Summary: The DH5α SbcCD host showed improved plasmid production and / or plasmid quality compared to Stbl3 or Stbl4 hosts with the AAV ITR vector, particularly compared to Stbl3 or Stbl4 hosts with larger therapeutic transgenes encoding the AAV ITR vector (Table 10).
[0251] Example 5: Performance of SbcCD knockout strains using non-palindromic vectors DH5α[SbcCD-] was evaluated in comparison to DH5α in terms of the production yield of a standard vector (12kb pHelper vector, pUC origin-kanR selected). The results showed that DH5α[SbcCD-] was superior to DH5α in terms of the production of a standard plasmid. [Table 11]
[0252] This was unexpected, as while SbcCD knockout can stabilize the palindrom, it is not expected to improve the yield of standard plasmids that do not contain the palindrom.
[0253] Example 6: Improved plasmid poly(A) repeat stability in DH5α TSbcD-1 compared to Stbl4. The pUC-AmpR plasmid vector encoding A90 repeats was transformed into Stbl4 or DH5α[SbcCD-], and the stability of the A90 repeats in four individual colonies from each transformation was determined by sequencing. All four Stbl4 colonies were deleting at least 20 bps of A90 repeats (i.e., all four colonies were<A70)であったが、DH5α[SbcCD-]コロニーは、> The result was A70, with 2 / 4 having intact A90 repeats. This demonstrates that DH5α[SbcCD-] stabilizes simple sequence repeats compared to stabilizing hosts in the art. This was unexpected, as SbcCD knockout is not expected to stabilize simple repeats.
[0254] Plasmid vectors encoding A117 repeats were transformed into DH5α[SbcCD-] and NTC1050811-HF[SbcCD-], and the stability of the A117 repeats was determined by sequencing. Cells were cultured at 30°C for 12 hours, ramped to 37°C at 24 EFT until OD decreased or lysis was observed, and then held at 25°C under HyperGro conditions as in Example 4. All transformed cell lines (2DH5α[SbcCD-], 2NTC1050811-HF[SbcCD-]) had intact A117 repeats and high yields, as shown in Table 12 below. This was unexpected, as SbcCD knockout is not expected to stabilize simple repeats. [Table 12]
[0255] The same procedure was used for plasmid vectors encoding A98-100 and A99-100 repeats, with DH5α[SbcCD-], NTC4862-HF[SbcCD-], and NTC1050811-HF[SbcCD-]. All transformed cell lines had intact repeats. All transformed cell lines had intact repeats and high yields. This was unexpected, as SbcCD knockouts are not expected to stabilize simple repeats. [Table 13]
[0256] Example 7: Cell line The above-described examples may be repeated using DH1, JM107, JM108, JM109, MG1655, XL1Blue and similar cell lines, or using SURE, SURE2, Stbl2, Stbl3, Stbl4 and non-SbcC, SbcD and / or SbcCD knockout strains.
[0257] All references, including publications, patent applications, and patents, are incorporated herein by reference to the same extent that each reference is incorporated by reference, to the same extent that the whole is incorporated herein, as is indicated individually and specifically.
[0258] The terms “includes,” “have,” “contains,” and “contains” should be interpreted as open-ended terms (i.e., “includes but not limited to”) unless otherwise stated herein. The enumeration of value ranges herein is intended solely as a simplification of each individual value included within the range, unless otherwise indicated herein, and each individual value is incorporated herein as if it were individually enumerated herein. All methods described herein may be performed in any preferred order unless otherwise indicated herein or unless it is clearly inconsistent with the context. Any and all examples provided herein, or the use of exemplary language (e.g., “for example”), are intended solely to better illustrate the invention and not to limit the scope of the invention unless specifically requested otherwise. No language herein should be interpreted as indicating that any unclaimed element is essential to the practice of the invention.
[0259] Preferred embodiments of the Invention, including the best mode known to the inventors for carrying out the Invention, are described herein. Variations of these preferred embodiments may become apparent to those skilled in the art by reading the foregoing description. The inventors expect that such variations will be appropriately used by those skilled in the art, and they intend that the Invention may be carried out in ways other than those specifically described herein. Accordingly, the Invention includes all modifications and equivalents of the subject matter enumerated in the claims appended herein, as permitted by applicable law. Furthermore, unless otherwise indicated herein or unless it is clearly inconsistent with the context, any combination of the elements described above in all possible variations is encompassed by the Invention.
[0260] SEQUENCE LISTING <110> Aldevron, L.L.C. <120> BACTERIAL HOST STRAINS <130> PA25-505 <150> 62 / 988,223 <151> 2020-03-11 <160> 65 <170> PatentIn version 3.5 <210> 1 <211> 70 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic primer <400> 1 ccctctgtat tcattatcct gctgaatagt tatttcactg caaacgtact catatgaata 60 tcctccttag 70 <210> 2 <211> 70 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic primer <400> 2 tctgtttggg tataatcgcg cccatgcttt ttcgccaggg aaccgttatg tgtaggctgg 60 agctgcttcg 70 <210> 3 <211> 22 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic primer <400> 3 cgtctcgcca tgatttgccc tg 22 <210> 4 <211> 22 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic primer <400> 4 cgttatgcgc cagctccgtg ag 22 <210> 5 <211> 1576 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 5 tctgtttgg tataatcgcg cccatgcttt ttcgccaggg aaccgttatg tgtaggctgg 60 120 atcccctcac gctgccgcaa gcactcaggg cgcaagggct gctaaaggaa gcggaacagg 180 tagaaagcca gtccgcagaa acggtgctga ccccggatga atgtcagcta ctgggctatc 240 tggacaaggg aaaacgcaag cgcaaagaga aagcaggtag cttgcagtgg gcttacatgg 300 cgatagctag actgggcggt tttatggaca gcaagcgaac cggaattgcc agctggggcg 360 ccctctggta aggttgggaa gccctgcaaa gtaaactgga tggctttctt gccgccaagg 420 atctgatggc gcaggggatc aagatctgat caagacag gatgaggatc gtttcgcatg 480 attgaacaag atggattgca cgcaggttct ccggccgctt gggtggagag gctattcggc 540 tatgactggg cacaagac aatcggctgc tctgatgccg ccgtgttccg gctgtcagcg 600 caggggcgcc cggttctttt tgtcaagacc gacctgtccg gtgccctgaa tgaactgcag 660 gacgaggcag cgcggctatc gtggctggcc acgacgggcg ttccttgcgc agctgtgctc 720 gacgttgtca ctgaagcggg aagggactgg ctgctattgg gcgaagtgcc ggggcaggat 780 ctcctgtcat ctcaccttgc tcctgccgag aaagtatcca tcatggctga tgcaatgcgg 840 cggctgcata cgcttgatcc ggctacctgc ccattcgacc accaagcga acatcgcatc gagcgagcac gtactcggat ggagccggt cttgtcgatc aggatgatct ggacgaagag catcaggggc tcgcgccagc cgaactgttc gccaggctca aggcgcgcat gcccgacggc 1020 gaggatctcg tcgtgaccca tggcgatgcc tgcttgccga atatcatggt ggaaaatggc 1080 cgcttttctg gattcatcga ctgtggccgg ctgggtgtgg cggaccgcta tcaggacata 1140 gcgttggcta cccgtgatat tgctgaagag cttggcggcg aatgggctga ccgcttcctc gtgctttacg gtatcgccgc tcccgattcg cagcgcatcg ccttctatcg ccttcttgac 1260 gagttcttct gagcgggact ctggggttcg aaatgaccga ccaagcgacg cccaacctgc 1320 catcacgaga tttcgattcc accgccgcct tctatgaaag gttgggcttc ggaatcgttt 1380 tccgggacgc cggctggatg atcctccagc gcggggatct catgctggag ttcttcgccc 1440 accccagctt caaaagcgct ctgaagttcc tatactttct agagaatagg aacttcggaa 1500 taggaactaa ggaggatatt catatgagta cgtttgcagt gaaataacta ttcagcagga 1560 taatgaatac agaggg 1576 <210> 6 <211> 5403 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 6 cgtctcgcca tgatttgccc tgttgtaata aataggttgc gatcattaat gcgacgtcat 60 tatgcgtcag atttatgaca gatttatgaa aagctcgtcg cacatatctt caggttattg 120 atttccgtgg cgcagaaaaa agcaaatggc acatctgttt gggtataatc gcgcccatgc 180 tttttcgcca gggaaccgtt atgcgcatcc ttcacacctc agactggcat ctcggccaga 240 acttctacag taaaagccgc gaagctgaac atcaggcttt tcttgactgg ctgctggaga 300 cagcacaaac ccatcaggtg gatgcgatta ttgttgccgg tgatgttttc gataccggct 360 cgccgcccag ttacgcccgc acgttataca accgttttgt tgtcaattta cagcaaactg 420 gctgtcatct ggtggtactg gcaggaaacc atgactcggt cgccacgctg aatgaatcgc 480 gcgatatcat ggcgttcctc aatactaccg tggtcgccag cgccggacat gcgccgcaaa 540 tcttgcctcg tcgcgacggg acgccaggcg cagtgctgtg ccccattccg tttttacgtc 600 cgcgtgacat tattaccagc caggcggggc ttaacggtat tgaaaaacag cagcatttac 660 tggcagcgat taccgattat taccaacaac actatgccga tgcctgcaaa ctgcgcggcg 720 atcagcctct gcccatcatc gccacgggac atttaacgac cgtgggggcc agtaaaagtg 780 acgccgtgcg tgacatttat attggcacgc tggacgcgtt tccggcacaa aactttccac 840 cagccgacta catcgcgctc gggcatattc accgcgcaca gattattggc ggcatggaac 900 atgttcgcta ttgcggctcc cccattccac tgagttttga tgaatgcggt aagagtaaat 960 atgtccatct ggtgacattt tcaaacggca aattagagag cgtggaaaac ctgaacgtac 1020 cggtaacgca acccatggca gtgctgaaag gcgatctggc gtcgattacc gcacagctgg 1080 aacagtggcg cgatgtatcg caggagccac ctgtctggct ggatatcgaa atcactactg 1140 atgagtatct gcatgatatt cagcgcaaaa tccaggcatt aaccgaatca ttgcctgtcg 1200 aagtattgct ggtacgtcgg agtcgtgaac agcgcgagcg tgtgttagcc agccaacagc 1260 gtgaaaccct cagcgaactc agcgtcgaag aggtgttcaa tcgccgtctg gcactggaag 1320 aactggatga atcgcagcag caacgtctgc agcatctttt caccacgacg ttgcataccc 1380 tcgccggaga acacgaagca tgaaaattct cagcctgcgc ctgaaaaacc tgaactcatt 1440 aaaaggcgaa tggaagattg atttcacccg cgagccgttc gccagcaacg ggctgtttgc 1500 tattaccggc ccaacaggtg cggggaaaac caccctgctg gacgccattt gtctggcgct 1560 gtatcacgaa actccgcgtc tctctaacgt ttcacaatcg caaaatgatc tcatgacccg 1620 cgataccgcc gaatgtctgg cggaggtgga gtttgaagtg aaaggtgaag cgtaccgtgc 1680 attctggagc cagaatcggg cgcgtaacca acccgacggt aatttgcagg tgccacgcgt 1740 agagctggcg cgctgcgccg acggcaaaat tctcgccgac aaagtgaaag ataagctgga 1800 actgacagcg acgttaaccg ggctggatta cgggcgcttc acccgttcga tgctgctttc 1860 gcaggggcaa tttgctgcct tcctgaatgc caaacccaaa gaacgcgcgg aattgctcga 1920 ggagttaacc ggcactgaaa tctacgggca aatctcggcg atggtttttg agcagcacaa 1980 atcggcccgc acagagctgg agaagctgca agcgcaggcc agcggcgtca cgttgctcac 2040 gccggaacaa gtgcaatcgc tgacagcgag tttgcaggta cttactgacg aagaaaaaca 2100 gttaattacc gcgcagcagc aagaacaaca atcgctaaac tggttaacgc gtcaggacga 2160 attgcagcaa gaagccagcc gccgtcagca ggccttgcaa caggcgttag ccgaagaaga 2220 aaaagcgcaa cctcaactgg cggcgcttag tctggcacaa ccggcacgaa atcttcgtcc 2280 acactgggaa cgcatcgcag aacacagcgc ggcgctggcg catattcgcc agcagattga 2340 agaagtaat actcgcttac agagcaat ggcgcttcgc gcgagcattc gccaccacgc 2400 ggcgaagcag tcagcagaat tacagcagca gcaacaaagc ctgaatacct ggttacagga 2460 acacgaccgc ttccgtcagt ggaacaacga accggcgggt tggcgtgcgc agttctccca 2520 acaaaccagc gatcgcgagc atctgcggca atggcagcaa cagttaaccc atgctgagca 2580 aaaacttaat gcgcttgcgg cgatcacgtt gacgttaacc gccgatgaag ttgctaccgc 2640 cctggcgcaa catgctgagc aacgccccact gcgtcagcac ctggtcgcgc tgcatggaca 2700 gattgttccc caacaaaaac gtctggcgca gttacaggtc gctatccaga atgtcacgca 2760 agaacagacg caacgtaacg ccgcacttaa cgaaatgcgc cagcgttata aagaaaagac 2820 gcagcaactt gccgatgtga aaaccatttg cgagcaggaa gcgcgcatca aaacgctgga 2880 agctcaacgt gcacagttac aggcgggtca gccttgccca ctttgtggtt ccaccagcca 2940 cccggcggtc gaggcgtatc aggcgctgga gcctggcgtt aatcagtctc gattactggc 3000 gctggaaaac gaagttaaaa agctcggtga agaaggtgcg acgctacgtg ggcaactgga 3060 cgccataaca aagcagcttc agcgtgatga aaacgaagcg caaagcctcc gacaagatga 3120 gcaagcactt actcaacaat ggcaagccgt cacggccagc ctcaatatca ccttgcagcc 3180 actggacgat attcaaccgt ggctggatgc acaagatgag cacgaacgcc agctgcggtt 3240 actcagccaa cggcatgaat tacaagggca gattgccgcg cataatcagc aaattatcca 3300 gtatcaacag caaattgaac aacgccagca actactttta acgacattga cgggttatgc 3360 actgacattg ccacaggaag atgaagaaga gagctggttg gcgacacgtc agcaagaagc 3420 gcagagctgg cagcaacgcc agaacgaatt aaccgcgctg caaaaccgta ttcagcagct 3480 gacgccgatt ctggaaacgt tgccgcaaag tgatgaactc ccgcactgcg aagaaactgt 3540 ggtattggaa aactggcggc aggtacatga acaatgtctc gcattacaca gccagcagca 3600 gacgttacag caacaggatg ttctggcggc gcaaagtctg caaaaagccc aggcgcagtt 3660 tgacaccgcg ctacaggcca gcgtctttga cgatcagcag gcgttccttg cggcgctaat 3720 ggatgaacaa acactaacgc agctggaaca gctcaagcag aatctggaaa accagcgccg 3780 tcaggcgcaa actctggtca ctcagacagc agaaacgctg gcacagcatc aacaacaccg 3840 acctgacgac gggttggctc tcactgtgac ggtggagcag attcagcaag agttagcgca 3900 aactcaccaa aagttgcgtg aaaacaccac gagtcaaggc gagattcgcc agcagctgaa 3960 gcaggatgca gataaccgtc agcaacaaca aaccttaatg cagcaaattg ctcaaatgac 4020 gcagcaggtt gaggactggg gatatctgaa ttcgctaata ggttccaaag agggcgataa 4080 attccgcaag tttgcccagg ggctgacgct ggataattta gtccatctcg ctaatcagca 4140 acttacccgg ctgcacgggc gctatctgtt acagcgcaaa gccagcgagg cgctggaagt 4200 cgaggttgtt gatacctggc aggcagatgc ggtacgcgat acccgtaccc tttccggcgg 4260 cgaaagtttc ctcgttagtc tggcgctggc gctggcgctt tcggatctgg tcagccataa 4320 aacacgtatt gactcgctgt tccttgatga aggttttggc acgctggata gcgaaacgct 4380 ggataccgcc cttgatgcgc tggatgccct gaacgccagt ggcaaaacca tcggtgtgat 4440 tagccacgta gaagcgatga aagagcgtat tccggtgcag atcaaagtga aaaagatcaa 4500 cggcctgggc tacagcaaac tggaaagtac gtttgcagtg aaataactat tcagcaggat 4560 aatgaataca gaggggcgaa ttatctcttg gccttgctgg tcgttatcct gcaagctatc 4620 actttattgg ctacggtgat tggtagccgt tctggtggtt gtgatggtgg tatgaaaaaa 4680 gtcattttat ctttggctct gggcacgttt ggtttgggga tggccgaatt tggcattatg 4740 ggcgtgctca cggagctggc gcataacgta ggaatttcga ttcctgccgc cgggcatatg 4800 atctcgtatt atgcactggg ggtggtggtc ggtgcgccaa tcatcgcact cttttccagc 4860 cgctactcac tcaaacatat cttgttgttt ctggtggcgt tgtgcgtcat tggcaacgcc 4920 atgttcacgc tctcttcgtc ttacctgatg ctcgccattg gtcggctggt atccggcttt 4980 ccgcatggcg cattttttgg cgtcggagcg atcgtgttat caaaaattat caaacccgga 5040 aaagtcaccg ccgccgtggc ggggatggtt tccgggatga cagtcgccaa tttgctgggc 5100 attccgctgg gaacgtattt aagtcaggaa tttagctggc gttacacctt tttattgatc 5160 gctgttttta atattgcggt gatggcatcg gtctattttt gggtgccaga tattcgcgac 5220 gaggcgaaag gaaatctgcg cgaacaattt cactttttgc gcagcccggc cccgtggtta 5280 attttcgccg ccacgatgtt tggcaacgca ggtgtgtttg cctggttcag ctacgtaaag 5340 ccatacatga tgtttatttc cggtttttcg gaaacggcga tgacctttat tatgatgtta 5400 gtt 5403 <210> 7 <211> 1922 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 7 cgtctcgcca tgatttgccc tgttgtaata aataggttgc gatcattaat gcgacgtcat 60 tatgcgtcag atttatgaca gatttatgaa aagctcgtcg cacatatctt caggttattg 120 atttccgtgg cgcagaaaaa agcaaatggc acatctgttt gggtataatc gcgcccatgc 180 tttttcgcca gggaaccgtt atgtgtaggc tggagctgct tcgaagttcc tatactttct 240 agagaatagg aacttcggaa taggaacttc aagatcccct cacgctgccg caagcactca 300 gggcgcaagg gctgctaaag gaagcggaac acgtagaaag ccagtccgca gaaacggtgc 360 tgaccccgga tgaatgtcag ctactgggct atctggacaa gggaaaacgc aagcgcaaag 420 agaaagcagg tagcttgcag tgggcttaca tggcgatagc tagactgggc ggttttatgg 480 acagcaagcg aaccggaatt gccagctggg gcgccctctg gtaaggttgg gaagccctgc 540 aaagtaaact ggatggcttt cttgccgcca aggatctgat ggcgcagggg atcaagatct 600 gatcaagaga caggatgagg atcgtttcgc atgattgaac aagatggatt gcacgcaggt 660 tctccggccg cttgggtgga gaggctattc ggctatgact gggcacaaca gacaatcggc 720 tgctctgatg ccgccgtgtt ccggctgtca gcgcaggggc gcccggttct ttttgtcaag 780 accgacctgt ccggtgccct gaatgaactg caggacgagg cagcgcggct atcgtggctg 840 gccacgacgg gcgttccttg cgcagctgtg ctcgacgttg tcactgaagc gggaagggac 900 tggctgctat tgggcgaagt gccggggcag gatctcctgt catctcacct tgctcctgcc 960 gagaaagtat ccatcatggc tgatgcaatg cggcggctgc atacgcttga tccggctacc 1020 tgcccattcg accaccaagc gaaacatcgc atcgagcgag cacgtactcg gatggaagcc 1080 ggtcttgtcg atcaggatga tctggacgaa gagcatcagg ggctcgcgcc agccgaactg 1140 ttcgccaggc tcaaggcgcg catgcccgac ggcgaggatc tcgtcgtgac ccatggcgat 1200 gcctgcttgc cgaatatcat ggtggaaaat ggccgctttt ctggattcat cgactgtggc 1260 cggctgggtg tggcggaccg ctatcaggac atagcgttgg ctacccgtga tattgctgaa 1320 gagcttggcg gcgaatgggc tgaccgcttc ctcgtgcttt acggtatcgc cgctcccgat 1380 tcgcagcgca tcgcttcta tcgcttctt gacgagttct tctgagcggg actctggggt 1440 tcgaaatgac cgaccaagcg acgcccaacc tgccatcacg agatttcgat tccaccgccg 1500 ccttctatga aaggttgggc ttcggaatcg ttttccggga cgccggctgg atgatcctcc 1560 agcgcgggga tctcatgctg gagttcttcg cccaccccag cttcaaaagc gctctgaagt 1620 tcctatactt tctagagaat aggaacttcg gaataggaac tagggaggat attcatatga 1680 gtacgtttgc agtgaaataa ctattcagca ggataatgaa tacagagggg cgaattatct 1740 cttggccttg ctggtcgtta tcctgcaagc tatcacttta ttggctacgg tgattggtag 1800 ccgttctggt ggttgtgatg gtggtatgaa aaaagtcatt ttatctttgg ctctgggcac 1860 gtttggtttg gggatggccg aatttggcat tatgggcgtg ctcacggagc tggcgcataa 1920 cg 1922 <210> 8 <211> 529 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 8 cgtctcgcca tgatttgccc tgttgtaata aataggttgc gatcattaat gcgacgtcat 60 tatgcgtcag atttatgaca gatttatgaa aagctcgtcg cacatatctt caggttattg 120 atttccgtgg cgcagaaaaa agcaaatggc acatctgttt gggtataatc gcgcccatgc 180 tttttcgcca gggaaccgtt atgtgtaggc tggagctgct tcgaagttcc tatactttct 240 agagaatagg aacttcggaa taggaactaa ggaggatatt catatgagta cgtttgcagt 300 gaaataacta ttcagcagga taatgaatac agaggggcga attatctctt ggccttgctg 360 gtcgttatcc tgcaagctat cactttattg gctacggtga ttggtagccg ttctggtggt 420 tgtgatggtg gtatgaaaaa agtcatttta tctttggctc tgggcacgtt tggtttgggg 480 atggccgaat ttggcattat gggcgtgctc acggagctgg cgcataacg 529 <210> 9 <211> 3147 <212> DNA <213> Escherichia coli <400> 9 atgaaaattc tcagcctgcg cctgaaaaac ctgaactcat taaaaggcga atggaagatt 60 gatttcaccc gcgagccgtt cgccagcaac gggctgtttg ctattaccgg cccaacaggt 120 gcggggaaaa ccaccctgct ggacgccatt tgtctggcgc tgtatcacga aactccgcgt 180 ctctctaacg tttcacaatc gcaaaatgat ctcatgaccc gcgataccgc cgaatgtctg 240 gcggaggtgg agtttgaagt gaaaggtgaa gcgtaccgtg cattctggag ccagaatcgg 300 gcgcgtaacc aacccgacgg taatttgcag gtgccacgcg tagagctggc gcgctgcgcc 360 gacggcaaaa ttctcgccga caaagtgaaa gataagctgg aactgacagc gacgttaacc 420 gggctggatt acgggcgctt cacccgttcg atgctgcttt cgcaggggca atttgctgcc 480 ttcctgaatg ccaaacccaa agaacgcgcg gaattgctcg aggagttaac cggcactgaa 540 atctacgggc aaatctcggc gatggttttt gagcagcaca aatcggcccg cacagagctg 600 gagaagctgc aagcgcaggc cagcggcgtc acgttgctca cgccggaaca agtgcaatcg 660 ctgacagcga gtttgcaggt acttactgac gaaaaaac agttaattac cgcgcagcag 720 aagaacaac aatcgctaa ctggttaacg cgtcaggacg aattgcagca agaagccagc 780 cgccgtcagc aggccttgca acaggcgtta gccgaagaag aaaaagcgca acctcaactg 840 gcggcgctta gtctggcaca accggcacga aatcttcgtc cacactggga acgcatcgca 900 gaacacagcg cggcgctggc gcatattcgc cagcagattg aagaagtaaa tactcgctta 960 cagagcacaa tggcgcttcg cgcgagcatt cgccaccacg cggcgaagca gtcagcagaa 1020 ttacagcagc agcaacaaag cctgaatacc tggttacagg aacacgaccg cttccgtcag 1080 tggaacaacg aaccggcggg ttggcgtgcg cagttctccc aacaaaccag cgatcgcgag 1140 catctgcggc aatggcagca acagttaacc catgctgagc aaaaacttaa tgcgcttgcg 1200 gcgatcacgt tgacgttaac cgccgatgaa gttgctaccg ccctggcgca acatgctgag 1260 caacgcccac tgcgtcagca cctggtcgcg ctgcatggac agattgttcc ccaacaaaaa 1320 cgtctggcgc agttacaggt cgctatccag aatgtcacgc aagaacagac gcaacgtaac 1380 gccgcactta acgaaatgcg ccagcgttat aaagaaaaga cgcagcaact tgccgatgtg 1440 aaaaccattt gcgagcagga agcgcgcatc aaaacgctgg aagctcaacg tgcacagtta 1500 caggcgggtc agccttgccc actttgtggt tccaccagcc acccggcggt cgaggcgtat 1560 caggcgctgg agcctggcgt taatcagtct cgattactgg cgctggaaaa cgaagttaaa 1620 aagctcggtg aagaaggtgc gacgctacgt gggcaactgg acgccataac aaagcagctt 1680 cagcgtgatg aaaacgaagc gcaaagcctc cgacaagatg agcaagcact tactcaacaa 1740 tggcaagccg tcacggccag cctcaatatc accttgcagc cactggacga tattcaaccg 1800 tggctggatg cacaagatga gcacgaacgc cagctgcggt tactcagcca acggcatgaa 1860 ttacaagggc agattgccgc gcataatcag caaattatcc agtatcaaca gcaaattgaa 1920 caacgccagc aactactttt aacgacattg acgggttatg cactgacatt gccacaggaa 1980 gatgaagaag agagctggtt ggcgacacgt cagcaagaag cgcagagctg gcagcaacgc 2040 cagaacgaat taaccgcgct gcaaaaccgt attcagcagc tgacgccgat tctggaaacg 2100 ttgccgcaaa gtgatgaact cccgcactgc gaagaactg tggtattgga aaactggcgg 2160 caggtacatg aacaatgtct cgcattacac agccagcagc agacgttaca gcaacaggat 2220 gttctggcgg cgcaaagtct gcaaaaagcc caggcgcagt ttgacaccgc gctacaggcc 2280 agcgtctttg acgatcagca ggcgttcctt gcggcgctaa tggatgaaca aacactaacg 2340 cagctggaac agctcaagca gaatctggaa aaccagcgcc gtcaggcgca aactctggtc 2400 actcagacag cagaaacgct ggcacagcat caacaacacc gacctgacga cgggttggct 2460 ctcactgtga cggtggagca gattcagcaa gagttagcgc aaactcacca aaagttgcgt 2520 gaaaacacca cgagtcaagg cgagattcgc cagcagctga agcaggatgc agataaccgt 2580 cagcaacaac aaaccttaat gcagcaaatt gctcaaatga cgcagcaggt tgaggactgg 2640 ggatatctga attcgctaat aggttccaaa gagggcgata aattccgcaa gtttgcccag 2700 gggctgacgc tggataattt agtccatctc gctaatcagc aacttacccg gctgcacggg 2760 cgctatctgt tacagcgcaa agccagcgag gcgctggaag tcgaggttgt tgatacctgg 2820 caggcagatg cggtacgcga tacccgtacc ctttccggcg gcgaaagttt cctcgttagt 2880 ctggcgctgg cgctggcgct ttcggatctg gtcagccata aaacacgtat tgactcgctg 2940 ttccttgatg aaggttttgg cacgctggat agcgaaacgc tggataccgc ccttgatgcg 3000 ctggatgccc tgaacgccag tggcaaaacc atcggtgtga ttagccacgt agaagcgatg 3060 aaagagcgta ttccggtgca gatcaaagtg aaaaagatca acggcctggg ctacagcaaa 3120 ctggaaagta cgtttgcagt gaaataa 3147 <210> 10 <211> 1227 <212> DNA <213> Escherichia coli <400> 10 atgctttttc gccagggaac cgttatgcgc atccttcaca cctcagactg gcatctcggc 60 cagaacttct acagtaaaag ccgcgaagct gaacatcagg cttttcttga ctggctgctg 120 gagacagcac aaacccatca ggtggatgcg attattgttg ccggtgatgt tttcgatacc 180 ggctcgccgc ccagttacgc ccgcacgtta tacaaccgtt ttgttgtcaa tttacagcaa 240 actggctgtc atctggtggt actggcagga aaccatgact cggtcgccac gctgaatgaa 300 tcgcgcgata tcatggcgtt cctcaatact accgtggtcg ccagcgccgg acatgcgccg 360 caaatcttgc ctcgtcgcga cgggacgcca ggcgcagtgc tgtgccccat tccgttttta 420 cgtccgcgtg acattattac cagccaggcg gggcttaacg gtattgaaaa acagcagcat 480 ttactggcag cgattaccga ttattaccaa caacactatg ccgatgcctg caaactgcgc 540 ggcgatcagc ctctgcccat catcgccacg ggacatttaa cgaccgtggg ggccagtaaa 600 agtgacgccg tgcgtgacat ttatattggc acgctggacg cgtttccggc acaaaacttt 660 ccaccagccg actacatcgc gctcgggcat attcaccgcg cacagattat tggcggcatg 720 gaacatgttc gctattgcgg ctcccccatt ccactgagtt ttgatgaatg cggtaagagt 780 aaatatgtcc atctggtgac attttcaaac ggcaaattag agagcgtgga aaacctgaac gtaccggtaa cgcaacccat ggcagtgctg aaaggcgatc tggcgtcgat taccgcacag 900 ctggaacagt ggcgcgatgt atcgcaggag ccacctgtct ggctggatat cgaaatcact 960 actgatgagt atctgcatga tattcagcgc aaaatccagg cattaaccga atcattgcct 1020 gtcgaagtt tgctggtacg tcggagtcgt gaacagcgcg agcgtgtgtt agccagccaa 1080 cagcgtgaaa ccctcagcga actcagcgtc gaagaggtgt tcaatcgccg tctggcactg 1140 gaagaactgg atgaatcgca gcagcaacgt ctgcagcatc ttttcaccac gacgttgcat 1200 accctcgccg gagaacga agcatga 1227 <210> 11 <211> 1428 <212> DNA <213> Escherichia coli <400> 11 atgatgaatg acggtaagca acaatctacc tttttgtttc acgattacga aacctttggc acgcaccccg cgttagatcg ccctgcacag ttcgcagcca ttcgcaccga tagcgaattc aatgtcatcg gcgaacccga agtcttttac tgcaagcccg ctgatgacta tttaccccag 180. ccaggagccg tattaattac cggtattacc ccgcaggaag cacgggcgaa aggagaaac gaagccgcgt ttgccgcccg tattcactcg ctttttaccg taccgaagac ctgtattctg 300 ggctacaaca atgtgcgttt cgacgacga gtcacacgca acatttttta tcgtaatttc tacgatcctt acgcctggag ctggcagcat gataactcgc gctgggattt actggatgtt 420 atgcgtgcct gttatgccct gcgcccgga ggaataact ggcctgaaaa tgatgacggt 480 ctaccgagct ttcgccttga gcatttaacc aaagcgaatg gtattgaaca tagcaacgcc cacgatgcga tggctgatgt gtacgccact attgcgatgg caaagctggt aaaaacgcgt 600 cagccacgcc tgtttgatta tctctttacc catcgtaata aacacaaact gatggcgttg 660 attgatgttc cgcagatgaa acccctggtg cacgtttccg gaatgtttgg agcatggcgc 720 ggcaatacca gctgggtggc accgctggcg tggcatcctg aaaatcgcaa tgccgtaatt 780 atggtggatt tggcaggaga catttcgcca ttactggaac tggatagcga cacattgcgc 840 gagcgtttat ataccgcaaa aaccgatctt ggcgataacg ccgccgttcc ggttaagctg 900 gtgcatatca ataaatgtcc ggtgctggcc caggcgaata cgctacgccc ggaagatgcc 960 gaccgactgg gaattaatcg tcagcattgc ctcgataacc tgaaaattct gcgtgaaaat 1020 ccgcaagtgc gcgaaaaagt ggtggcgata ttcgcggaag ccgaaccgtt tacgcttca 1080 gataacgtgg atgcacagct ttataacggc tttttcagtg acgcagatcg tgcagcaatg 1140 aaaattgtgc tggaaaccga gccgcgtaat ttaccggcac tggatatcac ttttgttgat 1200 aaacggattg aaaagctgtt gttcaattat cgggcacgca acttcccggg gacgctggat 1260 tatgccgagc agcaacgctg gctggagcac cgtcgccagg tcttcacgcc agagtttttg 1320 cagggttatg ctgatgaatt gcagatgctg gtacaacaat atgccgatga caaagagaaa 1380 gtggcgctgt taaaagcact ttggcagtac gcggaagaga ttgtctaa 1428 <210> 12 <211> 3543 <212> DNA <213> Escherichia coli <400> 12 atgagtgatg tcgccgagac actagatcct ttgcgcttgc ccttacaggg tgagcgcctg 60 attgaagcct ctgccggcac aggcaaaacc tttacgattg cggcgctcta tttgcgcctg 120 ttacttggac taggcggttc cgccgccttt ccccgcccgc tgaccgttga agaactgctg 180 gtggtcacct ttaccgaggc tgccacggca gaattgcgcg gtcgtatccg tagcaatatc 240 cacgagttgc gcatcgcctg tctgcgtgaa accaccgaca atccactgta cgaacgcctg 300 ctggaagaga tcgacgataa agcgcaagcc gcgcagtggt tgttgttagc cgaacggcag 360 atggatgaag cggcagtctt tactattcac ggctttgcc agcgcatgct caacctgaat 420 gcctttgaat ccggcatgct gtttgagcag cagctgattg aagatgagtc tctgctacgc 480 taccaggcct gcgccgattt ctggcgtcgc cactgctacc cgctgccgcg tgaaatagcc 540 caggtcgtct ttgaaacctg gaaagggccg caggcgttgc tgcgcgatat taatcgttat 600 ctgcaagggcg aagcgccggt tatcaaagca ccgccgccg atgatgaaac gctggcttcc 660 cgtcacgcgc aaattgtggc gcgtattgat acggtaaaac agcagtggcg cgacgcagtg 720 ggtgaactgg atgcgctgat cgaatctttct ggtattgatc gacgcaagtt taaccgtagc 780 aatcaggcta aatggatcga caagatcagc gcctgggcag aagaagagac aaacgttat 840 cagttgccgg agtcgctgga aaaattctcc cagcgtttct tagagaatcg cacgaaggcc 900 gggggggaaa ccccgcgaca tccactgttt gaggcgatcg atcaactgct tgcagaacca 960 ttgtcgatcc gcgatctggt gatcacccgc gcattggctg agatccgcga aacagtagcg 1020 cgtgaaaaac gccgccgtgg cgaattgggt tttgatgaca tgttaagtcg gctcgattcc 1080 gcgctgcgta gcgaaagcgg tgaggtgttg gcagcggcga tccgtacgcg attcccggtg 1140 gcaatgatcg atgaatttca ggataccgac ccccagcagt accgaatttt tcgccgtatc 1200 tggcaccatc agccggaaac cgcattgttg ctaattggcg acccgaagca ggccatatat 1260 gcattccggg gtgcggatat cttcacttat atgaaggcgc gtagcgaagt tcacgcccac 1320 tacactttag acaccaactg gcgttccgca ccaggaatgg tgaacagcgt gaataagctt 1380 ttcagccaga ctgatgacgc gttcatgttt cgcgaaatac cgtttattcc agtgaaatca 1440 gccgggaaaa atcaggcgtt acgttttgta tttaaaggtg aaacacagcc tgcgatgaaa 1500 atgtggctga tggaaggcga aagctgcggc gttggcgatt atcaaagtac catggcgcag 1560 gtatgtgctg cgcaaatccg cgactggcta caagccggac agcggggcga agcgttgctg 1620 atgaacggcg acgacgcgcg tccggtgcgt gcttcggaca tcagtgtgct ggtgcgcagc 1680 cgccaggagg ccgcccaggt gcgcgatgcc ttaacgttgc tggaaatccc ttccgtttac 1740 ctttcgaacc gcgacagtgt ttttgaaact ctggaagcgc aggaaatgct ttggttgttg 1800 caggcggtga tgacgcccga acgtgagaac accctgcgta gtgcgctggc aacgtcaatg 1860 atggggctga acgcgctgga tatcgaaacg ctgaacaatg acgaacatgc gtgggatgtg 1920 gtagtcgaag agttcgatgg ttatcggcaa atctggcgca aacgtggcgt tatgccgatg 1980 ctgcgggcgc tgatgtcggc gcgtaacatt gctgaaaact tgctggcaac ggcaggcggt 2040 gagcggcgtc ttaccgatat cttgcatatc agcgaactgc tacaagaagc cggaacgcag 2100 ctggaaagtg aacatgcgct ggtacgctgg ttatcgcaac atatcctcga gccagacagt 2160 aatgcctcca gccaacaaat gcgtctcgaa agtgataaac atctggtgca gattgtcacg 2220 atccacaaat cgaaagggct ggaatatcca ttggtctggc tgccgtttat caccaatttc 2280 cgcgtccagg agcaggcgtt ttatcacgat cgccactcgt ttgaggcagt tctggatctt 2340 aatgctgcgc cagaaagcgt cgacctcgcg gaggccgaac gtctggcgga agatctgcgt 2400 ttgctttacg tggcgctgac acgttcggtt tggcattgca gtctcggcgt tgcaccgctg 2460 gtgcgccgtc gtggcgataa aaaaggtgac accgacgtcc accaaagtgc gctcgggcgt 2520 ttgctgcaaa aaggggaacc gcaagatgcg gcagggcttc gcacctgtat tgaagcgtta 2580 tgcgatgatg atattgcctg gcaaacggca caaactggtg ataaccaacc ctggcaggtt 2640 aatgatgttt ctacagcaga gctgaatgcg aagacgttac aacgattgcc cggcgataac 2700 tggcgcgtca ccagctactc tggtttgcaa cagcgtggtc acggtatcgc ccaggatttg 2760 atgcctcggc tggatgtcga tgctgcaggc gttgccagcg tcgttgaaga accgacgtta 2820 acaccacatc agtttccgcg cggtgcgtca ccggggacgt tcttgcacag tttgtttgaa 2880 gacctggatt ttacccagcc ggttgacccg aactgggtgc gggaaaaact ggaactcggc 2940 ggctttgaat cgcagtggga accggtattg accgagtgga tcacggctgt cctccaggca 3000 cctctcaatg aaaccggcgt aagcctgagt caactttccg cccgcaataa acaggtggag 3060 atggagtttt atctgccgat tagtgaaccg cttatcgcca gtcagcttga tacgttaatc 3120 cgccagtttg acccgctatc cgcaggctgc ccgccgctgg agttcatgca ggtacgtggc 3180 atgttaaaag gctttatcga cctggtgttc cgccacgaag ggcgttatta cctgctcgac 3240 tataaatcca actggttggg tgaagacagt tcggcttaca cccaacaggc tatggcagcg 3300 gcaatgcagg cacaccgcta tgatctgcaa tatcagcttt ataccctggc gctgcatcgt 3360 tatctgcgcc atcgcattgc tgattacgac tatgagcacc actttggcgg cgttatttat 3420 ctgttcctgc gtggcgttga taaagaacat ccgcaacagg ggatttacac aacccgaccc 3480 aacgccgggt tgattgccct gatggatgag atgtttgccg gtatgaccct ggaggaggcg 3540 taa 3543 <210> 13 <211> 1827 <212> DNA <213> Escherichia coli <400> 13 atgaaattgc aaaagcaatt actggaagct gtggagcaca aacagctacg cccgctggat 60 gtgcaatttg ccctgaccgt ggcgggagat gaacatcctg ccgtcaccct cgcggcggca 120 ctgttaagtc atgatgccgg agagggacac gtttgtttgc cgctttcacg actggaaaat 180 aacgaggcgt cgcatccgct gttggcgacc tgtgtcagtg aaatcggtga gctacaaaat 240 tgggaagaat gcttgctggc ttctcaagcg gtcagcaggg gagatgaacc cacgccgatg 300 atcctctgtg gcgatcgtct ttatttgaat cgcatgtggt gtaacgagcg cacagtggca 360 cgctttttca acgaagtgaa tcatgccatt gaggttgatg aagctctact ggcgcaaacc 420 ctggacaaac tttttccagt aagcgatgaa attaactggc aaaaagttgc ggcggcagtg 480 gcgctgacgc ggcggatctc ggtgatttcc ggcggccctg gcaccggtaa aacgaccacc 540 gtagcgaagt tgctggcagc gttaattcaa atggccgacg gcgaacgctg ccgtatccgt 600 ctggctgcac caacgggtaa agctgccgcg cgcttaaccg aatctctcgg caaggctttg 660 cgacagttac cgctgaccga tgaacaaaag aaacgcattc cggaagatgc cagcactttg 720 caccgattgc tgggcgcgca gccgggtagc cagcgtttac gtcatcatgc cggtaacccg 780 ctgcatcttg atgtgctggt ggtagatgaa gcgtcaatga tcgatctgcc tatgatgtcg 840 agactgatcg acgccttgcc cgatcatgcg cgagtgatct ttctcggcga tcgtgatcaa 900 ctggcctcgg ttgaggctgg ggctgtgctg ggcgatatct gcgcttatgc caacgcgggc 960 tttaccgccg agcgtgccag gcagctaagc cgcctgacgg ggactcacgt tccggcagga 1020 accggcacag aagcggcatc tttgcgcgac agtctctgcc tgctgcaaaa aagctatcgt 1080 ttcggcagcg attctggcat tggtcagtta gctgcggcga tcaaccgtgg tgataaaacg 1140 gcagtgaaaa ccgtttttca gcaggatttt actgatatcg aaaaacggct tttacagagc 1200 ggcgaagatt atattgcgat gcttgaggaa gctcttgcgg gttacggacg ttatctggat 1260 ctgctgcaag cgcgtgccga gccggattta atcattcagg cgttcaatga gtaccagctt 1320 ttgtgcgccc tgcgggaagg gccgtttggc gtggctggac tgaatgagcg aattgagcag 1380 tttatgcaac agaagcgcaa aattcatcgt catccgcact ctcgttggta cgaaggtcga 1440 ccggtgatga ttgcccgtaa tgacagcgcg cttgggttgt ttaatggcga tatcggtatt 1500 gcgctggatc gcgggcaggg gacgcgcgtc tggtttgcga tgccggacgg caatattaag 1560 tctgtgcaac cgagtcgctt gccagagcac gaaactacgt gggcgatgac ggtacataaa 1620 tcgcagggat cggagttcga ccatgcggcg ttgattttgc cgagccaacg cacgccggta 1680 gtaacgcgag agctggttta taccgcggtg acccgcgcgc gtcgccgtct gtcgctgtat 1740 gccgatgagc gcatattaag tgcggcaatc gccactcgta ctgagcggcg cagtggtctg 1800 gcggcgttgt ttagttcacg ggaataa 1827 <210> 14 <211> 1833 <212> DNA <213> Escherichia coli <400> 14 gtgagtgatc agtttgacgc aaaagcgttt ttaaaaaccg taaccagcca gccaggcgtt 60 tatcgcatgt acgatgctgg tggtacggtt atctatgtcg gcaaagcgaa agacctgaaa 120 aaacggcttt ccagctattt ccgtagcaac ctcgcttcgc gcaaaaccga agcgctggtc 180 gcccagatcc agcaaattga tgtaacggtt actcacacag aaaccgaagc gctgttgctg 240 gaacacaact acatcaaact ctatcagccg cgttacaacg ttttgctacg cgatgataaa 300 tcatatcctt ttatcttcct gagtggtgat acccacccgc gtctggcgat gcatcgtggt 360 gcgaagcatg ccaaaggtga atatttcggc ccgttcccga atggctatgc cgtacgtgaa 420 acactggcgc tactgcaaaa gattttcccc attcgccagt gcgaaaatag tgtttatcgc 480 aatcgctcgc gtccgtgtct gcaataccag atagggcgct gtctgggacc gtgcgttgaa 540 ggactggtga gtgaagaaga atacgctcag caggtcgagt atgtgcgcct gtttttgtct 600 ggcaaagatg atcaggtgct tacgcaactc attagtcgta tggaaactgc cagccagaat 660 ctggagtttg aagaagctgc acgtattcgc gaccaaattc aggcggtgcg acgcgtcacc 720 gaaaaacaat tcgtttccaa taccggcgac gacctcgacg ttattggtgt ggcgttcgat 780 gcgggcatgg cttgtgtcca cgtattgttc attcgtcagg gcaaagtgct cggcagccgc 840 agctatttcc cgaaagtgcc tggcggtacg gaactgagcg aggtggtaga aaccttcgta 900 ggccagttct atttacaagg cagccagatg cgcaccttac cgggtgagat cctgctcgat 960 tttaatctta gcgataaaac gctgctcgcc gattcccttt cagaactggc gggacgcaag 1020 attaatgttc aaaccaaacc tcgcggcgat agggcgcgtt atctgaaact cgcgcgcacc 1080 aatgcggcga cggccttaac cagcaaactt tcgcagcaat ctaccgttca ccagcgactg 1140 accgcgcttg ccagcgtgtt gaaattgccg gaagtgaagc ggatggagtg ctttgacatc 1200 agccatacca tgggcgaaca aaccgtcgct tcctgtgtgg tgtttgatgc taacggcccg 1260 ctgcgtgcgg agtatcggcg ctataacatt acagcatca cgccgggcga tgatttagcg 1320 gcgatgaatc aggtgctgcg tcggcgttat ggtaaagcca ttgacgacag taagatcccg 1380 gatgtgatcc tttcgacgg cggcaaggc cagcttgcgc aggcgaaaa tgtcttcgcc 1440 gaactggatg tctcatggga taaaaatcat ccgctgctac tggcgttgc CAAggagca 1500 gatcgtagg ctggactgga aacgctgttc ttgagccgg aggtgaggg atttagttg 1560 ccgccagatt cacccgcgct gcatgttatc cagcatattc gcgatgaatc acatgatcac 1620 gcgattggcg ggcaccgtaa aaacggggcg aaggtcaaa attackacttc cctggaaacc 1680 attgaaggcg tcggccaaa acgtcggcaa atgttgttga atatatggg cggtttgcaa 1740 ggtttacgta acgccagcgt cgaggaatt gcaaagtgc cgggtatttc gcaaggtctg 1800 gcagaaaaga tcttctggtc gttgaaacat tga 1833 <210> 15 <211> 834 <212> DNA <213> Escherichia coli <400> 15 atgcatgttt ttgataataa tggaattgaa ctgaaagctg agtgttcgat aggtgaagag 60 gatggtgttt atggtctaat ccttgagtcg tgggggccgg gtgacagaaa caaagattac 120 aatatcgctc ttgattatat cattgaacgg ttggttgatt ctggtgtatc ccaagtcgta 180 gtatatctgg cgtcatcatc agtcagaaaa catatgcatt ctttggatga aagaaaaatc 240 catcctggtg aatattttac tttgattggt atagccccc gcgatatacg cttgaagatg 300 tgtggttatc aggcttattt tagtcgtacg gggagaaagg aaattccttc cggcaataga 360 acgaaacgaa tattgataaa tgttccaggt atttatagtg acagtttttg ggcgtctata 420 atacgtggag aactatcaga gctttcacag cctacagatg atgaatcgct tctgaatatg 480 agggttagta attaattaa gaaaacgttg agtcaacccg agggctccag gaaaccagtt 540 gaggtagaaa gactacaaaa agtttatgtc cgagacccga tggtaaagc ttggatttta 600 cagcaagta aaggtatatg tgaaaactgt ggtaaaaatg ctccgtttta tttaaatgat 660 ggaaacccat atttggaagt acatcatgta attccctgt cttcaggtgg tgctgataca 720 acagataact gtgttgccct ttgtccgaat tgccatagag aattgcacta tagtaaaaat 780 gcaaagaac taatcgagat gctttacgtt ataataacc gattacaga ata 834 <210> 16 <211> 1380 <212> DNA <213> Escherichia coli <400> 16 atggaatcta ttcaccctg gattgaaaaa ttttagc aagcacagca acacgttcg 60 caatccacta aagattatcc aacgtcttac cgtaacctgc gagtaaaatt gagtttcggt 120 tatggtatt ttacgtctat tccctgttt gcatttcttg gagaaggtca ggaagcttct 180 aacggtatat atcccgttat tctttatat aaagatttg atgagttggt ttggcttat 240 ggtataagcg acacgaatga accacatgcc caatggcagt tctcttcaga catacctaaa 300 acaatcgcag agtattttca ggcaacttcg ggtgtatatc ctaaaaaata cggacagtcc 360 tattacgcct gttcccaaaa agtctcacag ggtattgatt acacccgatt tgcctctatg 420 ctggacaaca taatcaacga ctataaatta atatttaatt ctggcaagag tgttattcca 480 cctatgtcaa aaactgaatc atactgtctg gaagatgcgt taaatgattt gtttatccct 540 gaaaccacaa tagagacgat actcaaacga ttaaccatca aaaaaatat tatcctccag 600 gggccgcccg gcgttggaaa aacctttgtt gcacgccgtc tggcttactt gctgacagga 660 gaaaaggctc cgcaacgcgt caatatggtt cagttccatc aatcttatag ctatgaggat 720 tttatacagg gctatcgtcc gaatggcgtc ggcttccgac gtaaagacgg catattttac 780 aatttttgtc agcaagctaa agagcagcca gagaaaagt atattttt tatagatgaa 840 atcaatcgtg ccaatctcag taaagtattt ggcgaagtga tgatgttaat ggaacatgat 900 aaacgaggtg aaaactggtc tgttccccta acctactccg aaaacgatga agaacgattc 960 tatgtcccgg agaatgttta tatcatcggt ttaatgaata ctgccgatcg ctctctggcc 1020 gttgttgact atgccctacg cagacgattt tctttcatag atattgagcc aggttttgat 1080 acaccacagt tccggaattt tttactgaat aaaaaagcag aaccttcatt tgttgagtct 1140 ttatgccaaa aaatgaacga gttgaaccag gaaatcagca aagaggccac tatccttggg 1200 aaaggattcc gcattgggca tagttacttc tgctgtgggt tggaagatgg cacctctccg 1260 gatacgcaat ggcttaatga aattgtgatg acggatatcg cccctttat cgaagaatat 1320 ttctttgatg acccctataa acaacagaaa tggaccaaca aattattagg ggactcatag 1380 <210> 17 <211> 1047 <212> DNA <213> Escherichia coli <400> 17 gtggaacagc ccgtgatacc tgtccgtaat atctattaca tgcttaccta tgcatggggt 60 tattacagg aattaagca ggcaacctt gaagccatac ccggtacaa tctcttgat 120 atcctggggt atgtattaaa taaaggggtt ttacagcttt cacgccgagg gcttgagctt 180 gattacaatc ctacaccga gatcattcct ggcatcaag ggcgaataga gtttgctaaa 240 acaatacgcg gcttccatct taatcatggg aaaaccgtca gtactttga tatgcttaat 300 gagacacgc tggctaccg aattataaaa agcacattag ccatattaat taagcatgaa 360 aagttaaatt caactatcag agatgaagct cgttcacttt atagaaattt accggcatt 420 agcactctc atttaaccc gcagcatttc agctatctga atggcggaaa aaatacgcgt 480 tattataat tcgttatcag tgtctgcaa ttcatcgtca atattctat tccaggtca 540 aaaaaggac actaccgttt ctatgatttt gaagaaacg aaaaagagat gtcattactt 600 tatcaaagt ttctttatga atttgccgt cgtgaattaa cgtctgcaaa cacaacccgc 660 tcttatttaa atgggatgc atcgagtata tcggatcagt cacttaattt gttacctcga 720 atggaaactg acatcaccat tcgctcatca gaaaaaatac ttatcgttga cgccaaatac 780 tataagagca tttttcacg acgaatggga acagaaaaat ttcattcgca aaatctttat 840 siactgatga attacttag gtcgttaaag cctgaaaatg gcgaaaacat aggggtta 900 ttaatatatc cccacgtaga taccgcagtg aaacatcgtt aataaattaa tggcttcgat 960 attggcttgt gtaccgtcaa tttaggtcag gatggccgt gtatacatca agaattactc 1020 gatattttcg atgaatct cataa 1047 <210> 18 <211> 1395 <212> DNA <213> Escherichia coli <400> 18 atgagtgcgg ggaaattgcc ggaggggtgg gttatcgccc cagtatcc ggtcacaact 60 ctaatccgag gagtaacgta taaaaaagag caggcaataa atatctaaa atgattat 120 ttgcctctta tccgtgcgaa caatattcag aatggcaagt ttgatactac ggacttggtt 180 tttgttccta aaatctctgt taaagaagt caaaaatat ctcctgaaga tattgttatt 240 gcaatgtcat cagggagcaa atccgtagtt ggtaaatccg cacatcagca tctaccattt 300 gatgtagtt tcggcgcatt ttgcggtgta ttacgtcctg aaaaacttat atttctggt 360 tttattgctc atttcacaa atctctctt tatcgaaca aaatttcatc actttctgct 420 ggtgcaaata ttaatatat taagccggca agctttgatt tgataatat accaatccca 480 ccacttgccg aaaaaaat catcgctgaa aactcgata cgctgctgc gcaggtagac 540 agcaccaaag cacgttttga gcaatccca caaatcctga aacgttttcg tcaagcggta 600 ttggggggcg cagttaatgg aaaattgaca gaaaaatggc gtattttga gccgcacat 660 tctgtattta agaagttaaa ttttgaatct atcttaactg aattacgtaa tggtctttca 720 tcaaagccaa atgaaagtgg tgttggtcat ccaatactac gcattagttc tgtacgtgct 780 ggccaatgtag atcaaaacga tattcggttt ctagaatgtt cagaaagtga actaaaccgc 840 caaaattac aagatggagaga tcttttattt actcgctata acggaagttt agaatttgtt 900 ggtgtttgtg ggttattgaa aaaattacaa catcaaaatt tgctatatcc tgataaactt 960 attcgagctc gattaaccaa agatgcttta ccagaatata tcgaaatatt tttttcatcc 1020 ccctcagcac gaaatgcaat gatgaactgc gtgaaaacaa cttctggtca aaaaggtatt 1080 tcaggaaaag atatcaaatc ccaagttgtt ttattacctc cattaaaaga acaagccgaa 1140 atcgttcgcc gcgtcgagca actcttcgcc tacgccgaca ccatagaaaa acaggtcaac 1200 aacgccttag cccgcgtcaa caacctgacg caatccatcc tggcaaaagc gttccgtggt 1260 gaacttaccg cccagtggcg ggccgaaaac ccggatttga tcagcggaga aaacagcgcc 1320 gccgcgttgc tggaaaaaat caaagctgaa cgcgcagcta gcgggggtaa aaaagcctca 1380 cgtaaaaaat cctga 1395 <210> 19 <211> 1590 <212> DNA <213> Escherichia coli <400> 19 atgaacaata acgatctggt cgcgaagctg tggaagctgt gcgacaacct gcgcgatggc 60 ggcgtttcct atcaaaacta cgtcaatgaa ctcgcctcgc tgctgttttt gaaaatgtgt 120 aaagagaccg gtcaggaagc ggaatacctg ccggaaggtt accgctggga tgacctgaaa 180 tcccgcatcg gccaggagca gttgcagttc taccgaaaaa tgctcgtgca tttaggcgaa 240 gatgacaaaa agctggtaca ggcagttttt cataatgtta gtaccaccat caccgagccg 300 aaacaaataa ccgcactggt cagcaatatg gattcgctgg actggtacaa cggcgcgcac 360 ggtaagtcgc gcgatgactt cggcgatatg tacgaagggc tgttgcagaa gaacgcgaat 420 gaaaccaagt ctggtgcagg ccagtacttc accccgcgtc cgctgattaa aaccattatt 480 catctgctga aaccgcagcc gcgtgaagtg gtgcaggacc cggcggcagg tacggcgggc 540 tttttgattg aagccgaccg ctatgttaag tcgcaaacca atgatctgga cgaccttgat 600 ggcgacacgc aggatttcca gatccaccgc gcgtttatcg gcctcgaact ggtgcccggc 660 acccgtcgtc tggcactgat gaactgcctg ctgcacgata ttgaaggcaa cctcgaccac 720 ggcggcgcaa tccgtctggg caacactctg ggtagcgacg gtgaaaacct gccgaaggcg 780 catattgtcg ccactaaccc gccgtttggc agcgccgcag gcaccaacat tacccgcacc 840 tttgttcacc cgaccagcaa caaacagttg tgctttatgc agcatattat cgaaacgctg 900 catcccggcg gtcgtgcggc ggtggtggtg ccggataacg tgctgtttga aggcggcaaa 960 ggcaccgaca ttcgtcgtga cctgatggat aagtgtcatc tgcacaccat tctgcgtctg 1020 ccgaccggta tttttacgc tcagggcgtg aagaccaacg tgctgttctt taccaaaggg 1080 acggtggcga acccgaatca ggataagaac tgtaccgatg atgtgtgggt gtatgacctg 1140 cgtaccaata tgccgagttt cggcaagcgc acaccgttta ccgacgagca tttgcagccg 1200 tttgagcgcg tgtatggcga agacccgcac ggtttaagcc cgcgcactga aggtgaatgg 1260 agttttaacg ccgaagagac ggaagttgcc gacagcgaag agaacaaaaa caccgaccag 1320 catcttgcta ccagccgctg gcgcaagttc agccgtgagt ggatccgcac cgcaaaatcc 1380 gattcgctgg atactcctg gctgaaagat aaagacagta ttgatgccga cagcctgccg 1440 gagccggatg tattagcggc agaagcgatg ggcgaactgg tacaggcgct gtctgaactg 1500 gatgcgctga tgcgtgaact gggggcgagc gatgaggccg atttgcagcg tcagttgctg 1560 gaagaagcgt ttggtggggt gaaggaatga 1590 <210> 20 <211> 3513 <212> DNA <213> Escherichia coli <400> 20 atgatgaata aatccaattt tgaattcctg aagggcgtca acgacttcac ttatgccatc 60 gcctgtgcgg cggaaaataa ctacccggat gatcccaaca cgacgctgat taaaatgcgt 120 atgtttggcg aagccacagc gaaacatctt ggtctgttac tcaacatccc cccttgtgag 180 aatcaacacg atctcctgcg tgaactcggc aaaatcgcct ttgttgatga caacatcctc 240 tctgtatttc acaaattacg ccgcattggt aaccaggcgg tgcacgaata tcataacgat 300 ctcaacgatg cccagatgtg cctgcgactc gggttccgcc tggctgtctg gtactaccgt 360 ctggtcacta aagattatga cttcccggtg ccggtgtttg tgttgccgga acgtggtgaa 420 aacctctatc accaggaagt gctgacgcta aaacaacagc ttgaacagca ggtgcgagaa 480 aaagcgcaga ctcaggcaga agtcgaagcg caacagcaga agctggttgc cctgaacggc 540 tatatcgcca ttctggaagg caaacagcag gaaaccgaag cgcaaaccca ggctcgcctt 600 gcggcactgg aagcacagct cgccgagaag aacgcggaac tggcaaaaca gaccgaacag 660 gaacgtaagg cttaccacaa agaaattacc gatcaggcca tcaagcgcac actcaacctt 720 agcgaagaag agagtcgctt cctgattgat gcgcaactgc gtaaagcagg ctggcaggcc 780 gacagcaaaa ccctgcgctt ctccaaaggc gcacgtccgg aacccggcgt caataaagcc 840 attgccgaat ggccgaccgg aaaagatgaa acgggtaatc agggctttgc ggattatgtg 900 ctgtttgtcg gcctcaaacc catcgcggtg gtagaggcga aacgtaacaa tatcgacgtt 960 cccgccaggc tcaatgagtc gtatcgctac agtaaatgtt tcgataatgg cttcctgcgg 1020 gaaaccttgc ttgagcacta ctcaccggat gaagtgcatg aagcagtgcc agagtatgaa 1080 accagctggc aggacaccag cggcaaacaa cggtttaaaa tccccttctg ctactcgacc 1140 aacgggcgcg aataccgcgc aacaatgaag accaaaagcg gcatctggta tcgcgacgtg 1200 cgtgataccc gcaatatgtc gaaagcctta cccgagtggc accgcccgga agagctgctg 1260 gaaatgctcg gcagcgaacc gcaaaaacag aatcagtggt ttgccgataa ccctggcatg 1320 agcgagctgg gcctgcgtta ttatcaggaa gatgccgtcc gcgcggttga aaaggcaatc 1380 gtcaaggggc aacaagagat cctgctggcg atggcgaccg gtaccggtaa aacccgtacg 1440 gcaatcgcca tgatgttccg cctgatccag tcccagcgtt ttaaacgcat tctcttcctt 1500 gtcgaccgcc gttctcttgg cgaacaggcg ctgggcgcgt ttgaagatac gcgtattaac 1560 ggcgacacct tcaacagcat tttcgacatt aaagggctga cggataaatt cccggaagac 1620 agcaccaaaa ttcacgttgc caccgtacag tcgctggtga aacgcaccct gcaatcagat 1680 gaaccgatgc cggtggcccg ttacgactgt atcgtcgttg acgaagcgca tcgcggctat 1740 attctcgata aagagcagac cgaaggcgaa ctgcagttcc gcagccagct ggattacgtc 1800 tctgcctacc gtcgcattct cgatcacttc gatgcggtaa aaatcgctct caccgccacc 1860 ccggcgctac atactgtgca gattttcggc gagccggttt accgttatac ctaccgtacc 1920 gcggttatcg acggttttct gatcgaccag gatccgccta ttcagatcat cacccgcaac 1980 gcgcaggagg gggtttatct ctccaaaggc gagcaggtag agcgcatcag cccgcaggga 2040 gaagtgatca atgacaccct ggaagacgat caggattttg aagtcgccga ctttaaccgt 2100 ggcctggtga tcccggcgtt taaccgcgcc gtctgtaacg aactcaccaa ttatcttgac 2160 ccgaccggat cgcaaaaaac gctggtcttc tgcgtcacca atgcccatgc cgatatggtg 2220 gtggaagagc tgcgtgccgc gttcaagaaa aagtatccgc aactggagca cgacgcgatc 2280 atcaagatca ccggtgatgc cgataaagac gcgcgcaaag tgcagaccat gatcacccgc 2340 ttcaataaag agcggctgcc caatatcgtg gtaaccgtcg acctgctgac gaccggcgtc 2400 gatattccgt cgatctgtaa tatcgtgttc ctgcgtaaag tacgcagccg cattctgtac 2460 gaacagatga aaggccgcgc cacgcgctta tgcccggagg tgaataaaac cagctttaag 2520 atttttgact gtgtcgatat ctacagcacg ctggagagcg tcgacaccat gcgtccggtg 2580 gtggtgcgcc cgaaggtgga actgcaaacg ctggtcaatg aaattaccga ttcagaaacc 2640 tataaaatca ccgaagcgga tggccgcagt tttgccgagc acagccatga acaactggtg 2700 gcgaagctcc agcgtatcat cggtctggcc acgtttaacc gtgaccgcag cgaaacgata 2760 gataaacagg tgcgtcgtct ggatgagcta tgccaggacg cggcgggcgt gaactttaac 2820 ggcttcgcct cgcgcctgcg ggaaaaaggg ccgcactgga gcgccgaagt ctttaacaaa 2880 ctgcctggct ttatcgcccg tctggaaaag ctgaaaacgg acatcaacaa cctgaatgat 2940 gcgccgatct tcctcgatat cgacgatgaa gtggtgagtg taaatcgct gtacggtgat 3000 tacgacacgc cgcaggattt cctcgaagcc ttgactcgc tggtgcaacg ttccccgaac 3060 gcgcaaccgg cattgcaggc agttattaat cgcccgcgg atctcaccg taaagggctg 3120 gtcgagctac aggagtggtt tgaccgccag cactttgagg aatcttccct gcgcaaagca 3180 tggaaagaga cgcgcaatga agatatcgcc gcccggctga ttggtcatat tcgccgcgct 3240 gcggtgggcg atgcgctgaa accgtttgag gaacgtgtcg atcacgcgct gacgcgcatt 3300 aagggcgaaa acgactggag cagcgagcaa ttaagctggc tcgatcgttt agcgcaggcg 3360 ctgaaagaga aagtggtgct cgacgacgat gtcttcaaaa ccggcaactt ccaccgtcgc 3420 ggcgggaagg cgatgctgca aagaaccttt gacgataatc tcgataccct gctgggcaaa 3480 ttcagcgatt atatctggga cgagctggcc is 3513 <210> 21 <211> 915 <212> DNA <213> Escherichia coli <400> 21 atgacggttc ctacctatga caaatttatt gaacctgttc tgcgttatct ggcaacaaaa 60 ccggaaggtg cagccgcgcg tgatgttcat gaggctgccg cggatgcatt aggactggat 120 gacagccagc gagcgaaagt cattaccagc ggacaacttg tttataaaaa tcgtgcaggc 180 tgggcgcatg accgtttaaa acgtgccggg ttgtcgcaaa gtttgtcgcg tggcaaatgg 240 tgcctgactc ctgcgggttt tgactgggtt gcgtctcatc cccagccaat gacggagcag 300 gagacgaacc atctggcctt cgcttttgtg aatgtcaaac ttaagtcacg gccggatgcc 360 gtcgatttag atccgaaagc cgactctccc gatcatgaag aacttgcaaa gagcagcccg 420 gacgatcggt tagatcaggc gctaaaagag cttcgtgatg cggtggctga tgaggttctg 480 gaaaacttat tgcaggtttc tccttcgcgc tttgaagtca ttgttctgga tgttttgcat 540 cgcctggggt atggcggcca ccgtgatgat ttgcagcgtg ttggcggtac tggagatggt 600 ggcatcgatg gtgtgatatc gcttgataaa cttggcctgg agaaagttta tgttcaggca 660 aaacgttggc agaatactgt aggcaggcca gaattacagg cattttacgg cgcactggct 720 gggcaaaaag cgaaacgtgg ggtgtttatt accacttctg gatttacttc tcaggcgcgt 780 gactttgccc aatccgtcga gggtatggtg ttggttgatg gggaacgcct ggtgcactta 840 atgatcgaaa acgaagtagg ggtttcttca cgtttgttga aggtgccgaa actggatatg 900 gactattttg agtga 915 <210> 22 <211> 784 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 22 aagaacatgt gagcaaaagg ccagcaaaag gccaggaacc gtaaaaaggc cgcgttgctg 60 gcgtttttcc ataggctccg cccccctgac gagcatcaca aaaatcgacg ctcaagtcag 120 aggtggcgaa acccgacagg actataaaga taccaggcgt ttccccctgg aagctccctc 180 gtgcgctctc ctgttccgac cctgccgctt accggatacc tgtccgcctt tctcccttcg 240 ggaagcgtgg cgctttctca tagctcacgc tgtaggtatc tcagttcggt gtaggtcgtt 300 cgctccaagc tgggctgtgt gcacgaaccc cccgttcagc ccgaccgctg cgccttatcc 360 ggtaactatc gtcttgagtc caacccggta agacacgact tatcgccact ggcagcagcc 420 actggtaaca ggattagcag agcgaggtat gtaggcggtg ctacagagtt cttgaagtgg 480 tggcctaact acggctcac tagagaaca gtatttggta tctgcgctct gctgaagcca 540 gttaccttcg gaaaaagagt tggtagctct tgatccggca aacaaaccac cgctggtagc 600 ggtggttttt ttgtttgcaa gcagcagatt acgcgcagaa aaaaaggatc tcaagaagat 660 cctttgatct tttctacggg gtctgacgct cagtggaacg aaaactcacg ttaagggatt 720 ttggtcatga gattatcaaa aaggatcttc acctagatcc ttttaaatta aaaatgaagt 780 ttta 784 <210> 23 <211> 795 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 23 atgattgaac aagatggatt gcacgcaggt tctccggccg cttgggtgga gaggctattc 60 ggctatgact gggcacaaca gacaatcggc tgctctgatg ccgccgtgtt ccggctgtca 120 gcgcaggggc gcccggttct ttttgtcaag accgacctgt ccggtgccct gaatgaactg 180 caggacgagg cagcgcggct atcgtggctg gccacgacgg gcgttccttg cgcagctgtg 240 ctcgacgttg tcactgaagc gggaagggac tggctgctat tgggcgaagt gccggggcag 300 gatctcctgt catctcacct tgctcctgcc gagaaagtat ccatcatggc tgatgcaatg 360 cggcggctgc atacgcttga tccggctacc tgcccattcg accaccaagc gaaacatcgc 420 atcgagcgag cacgtactcg gatggaagcc ggtcttgtcg atcaggatga tctggacgaa 480 gagcatcagg ggctcgcgcc agccgaactg ttcgccaggc tcaaggcgcg catgcccgac 540 ggcgaggatc tcgtcgtgac ccatggcgat gcctgcttgc cgaatatcat ggtggaaaat 600 ggccgctttt ctggattcat cgactgtggc cggctgggtg tggcggaccg ctatcaggac 660 atagcgttgg ctacccgtga tattgctgaa gagcttggcg gcgaatgggc tgaccgcttc 720 ctcgtgcttt acggtatcgc cgctcccgat tcgcagcgca tcgccttcta tcgccttctt 780 gacgagttct tctga 795 <210> 24 <211> 714 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 24 atgagcacaa aaaagaaacc attaacacaa gagcagcttg aggacgcacg tcgccttaaa 60 gcaatttatg aaaaaaagaa aaatgaactt ggcttatccc aggaatctgt cgcagacaag 120 atggggatgg ggcagtcagg cgttggtgct ttatttaatg gcatcaatgc attaaatgct 180 tataacgccg cattgcttac aaaaattctc aaagttagcg ttgaagaatt tagcccttca 240 atcgccagag aaatctacga gatgtatgaa gcggttagta tgcagccgtc acttagaagt 300 gagtatgagt accctgtttt ttctcatgtt caggcaggga tgttctcacc taagcttaga 360 acctttacca aaggtgatgc ggagagatgg gtaagcacaa ccaaaaaagc cagtgattct 420 gcattctggc ttgaggttga aggtaattcc atgaccgcac caacaggctc caagccaagc 480 tttcctgacg gaatgttaat tctcgttgac cctgagcagg ctgttgagcc aggtgatttc 540 tgcatagcca gacttggggg tgatgagttt accttcaaga aactgatcag ggatagcggt 600 caggtgtttt tacaaccact aaacccacag tacccaatga tcccatgcaa tgagagttgt 660 tccgttgtgg ggaaagttat cgctagtcag tggcctgaag agacgtttgg ctga 714 <210> 25 <211> 1422 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 25 atgaacatca aaaagtttgc aaaacaagca acagtattaa cctttactac cgcactgctg 60 gcaggaggcg caactcaagc gtttgcgaaa gaaacgaacc aaaagccata taaggaaaca 120 tacggcattt cccatattac acgccatgat atgctgcaaa tccctgaaca gcaaaaaaat 180 gaaaaatata aagttcctga gttcgattcg tccacaatta aaaatatctc ttctgcaaaa 240 ggcctggacg tttgggacag ctggccatta caaaacactg acggcactgt cgcaaactat 300 cacggctacc acatcgtctt tgcattagcc ggagatccta aaaatgcgga tgacacatcg 360 atttacatgt tctatcaaaa agtcggcgaa acttctattg acagctggaa aaacgctggc 420 cgcgtcttta aagacagcga caaattcgat gcaaatgatt ctatcctaaa agaccaaaca 480 caagaatggt caggttcagc cacatttaca tctgacggaa aaatccgttt attctacact 540 gatttctccg gtaaacatta cggcaaacaa acactgacaa ctgcacaagt taacgtatca 600 gcatcagaca gctctttgaa catcaacggt gtagaggatt ataaatcaat ctttgacggt 660 gacggaaaaa cgtatcaaaa tgtacagcag ttcatcgatg aaggcaacta cagctcaggc 720 gacaaccata cgctgagaga tcctcactac gtagaagata aaggccacaa atacttagta 780 tttgaagcaa acactggaac tgaagatggc taccaaggcg aagaatcttt atttaacaaa 840 gcatactatg gcaaaagcac atcattcttc cgtcaagaaa gtcaaaaact tctgcaaagc 900 gataaaaaac gcacggctga gttagcaaac ggcgctctcg gtatgattga gctaaacgat 960 gattacacac tgaaaaaagt gatgaaaccg ctgattgcat ctacacagt aacagatgaa 1020 attgaacgcg cgaacgtctt taaaatgaac ggcaatggt acctgttcac tgactcccgc 1080 ggatcaaaaa tgacgattga cggcattacg tctaacgata tttacatgct tggttatgtt 1140 tctaattctt taactggccc atacaagccg ctgaacaaaa ctggccttgt gttaaaaatg 1200 gatcttgatc ctaacgatgt aaccttact tactcacact tcgctgtacc tcaagcgaaa 1260 ggaaaacaatg tcgtgattac aagctatatg aaaacagag gattctacgc agaaaaaaaaaa 1320 tcaacgtttg cgcctagctt cctgctgaac atcaaggca agaaaacatc tgttgtcaa 1380 gapcatcc ttgacaagg acattaca gttaacaat aa 1422 <210> 26 <211> 918 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 26 atgagactca aggtcatgat ggacgtgaac aaaaaaacga aaattcgcca ccgaaacgag 60 ctaaatcaca ccctggctca acttcctttg cccgcaaagc gagtgatgta tatggcgctt 120 gctctcattg atagcaaaga acctcttgaa cgagggcgag ttttcaaaat tagggctgaa 180 gaccttgcag cgctcgccaa aatcacccca tcgcttgctt atcgacaatt aaaagagggt 240 ggtaaattac ttggtgccag caaaatttcg ctaagagggg atgatatcat tgctttagct 300 aaagagctta acctgatctc tactgctaaa aactccagcg aagagttaga tcttaacatt 360 attgagtgga tagcttattc aaatgatgaa ggatacttgt ctttaaaatt caccagaacc 420 atagaaccat atatctctag ccttattggg aaaaaaaata aattcacaac gcaattgtta 480 acggcaagct tacgcttaag tagccagtat tcatcttctc tttatcaact tatcaggaag 540 cattactcta attttaagaa gaaaaattat tttattattt ccgttgatga gttaaaggaa 600 gagttaatag cttatacttt tgataaagat ggaaatattg agtacaaata ccctgacttt 660 cctattttta aaagggatgt gttaaataaa gccattgctg aaattaaaaa gaaaacagaa 720 atatcgtttg ttggcttcac tgttcatgaa aaagaaggga gaaaattag taagctgaag 780 ttcgaatttg tcgttgatga agatgaattt tctggcgata aagatgatga agcttttttt 840 atgaatttat ctgaagctga tgcagctttt ctcaaggtat ttgatgaaac cgtacctccc 900 aaaaaagcta aggggtga 918 <210> 27 <211> 912 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 27 atgagactca aggtcatgat ggacgtgaac aaaaaaacga aaattcgcca ccgaaacgag 60 ctaaatcaca ccctggctca acttcctttg cccgcaaagc gagtgatgta tatggcgctt 120 gctctcattg atagcaaaga acctcttgaa cgagggcgag ttttcaaaat tagggctgaa 180 gaccttgcag cgctcgccaa aatcacccca tcgcttgctt atcgacaatt aaaagagggt 240 ggtaaattac ttggtgccag caaaatttcg ctaagagggg atgatatcat tgctttagct 300 aaagagctta acctgactgc taaaaactcc agcgaagagt tagatcttaa cattattgag 360 tggatagctt attcaaatga tgaaggatac ttgtctttaa aattcaccag aaccatagaa 420 ccatatatct ctagccttat tgggaaaaaa aataaattca caacgcaatt gttaacggca 480 agcttacgct taagtagcca gtattcatct tctctttatc aacttatcag gaagcattac 540 tctaatttta agaagaaaaa ttattttatt atttccgttg atgagttaaa ggaagagtta 600 atagcttata cttttgataa agatggaaat attgagtaca aataccctga ctttcctatt 660 tttaaaaggg atgtgttaaa taaagccatt gctgaaatta aaaagaaaac agaaatatcg 720 tttgttggct tcactgttca tgaaaaagaa gggagaaaaa ttagtaagct gaagttcgaa 780 tttgtcgttg atgaagatga attttctggc gataaagatg atgaagcttt ttttatgaat 840 ttatctgaag ctgatgcagc ttttctcaag gtatttgatg aaaccgtacc tcccaaaaaa 900 gctaaggggt ga 912 <210> 28 <211> 918 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 28 atgagactca aggtcatgat ggacgtgaac aaaaaaacga aaattcgcca ccgaaacgag 60 ctaaatcaca ccctggctca acttcctttg cccgcaaagc gagtgatgta tatggcgctt 120 gctctcattg atagcaaaga acctcttgaa cgagggcgag ttttcaaaat tagggctgaa 180 gaccttgcag cgctcgccaa aatcacccca tcgcttgctt atcgacaatt aaaagagggt 240 ggtaaattac ttggtgccag caaaatttcg ctaagagggg atgatatcat tgctttagct 300 aaagagctta acctgctgtc tactgctaaa aactcccctg aagagttaga tcttaacatt 360 attgagtgga tagcttattc aaatgatgaa ggatacttgt ctttaaaatt caccagaacc 420 atagaaccat atatctctag ccttattggg aaaaaaaata aattcacaac gcaattgtta 480 acggcaagct tacgcttaag tagccagtat tcatcttctc tttatcaact tatcaggaag 540 cattactcta attttaagaa gaaaaattat tttattattt ccgttgatga gttaaaggaa 600 gagttaatag cttatacttt tgataaagat ggaaatattg agtacaaata ccctgacttt 660 cctattttta aaagggatgt gttaaataaa gccattgctg aaattaaaaa gaaaacagaa 720 atatcgtttg ttggcttcac tgttcatgaa aaagaaggga gaaaaattag taagctgaag 780 ttcgaatttg tcgttgatga agatgaattt tctggcgata aagatgatga agcttttttt 840 atgaatttat ctgaagctga tgcagctttt ctcaaggtat ttgatgaaac cgtacctccc 900 aaaaaagcta aggggtga 918 <210> 29 <211> 918 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 29 atgagactca aggtcatgat ggacgtgaac aaaaaaacga aaattcgcca ccgaaacgag 60 ctaaatcaca ccctggctca acttcctttg cccgcaaagc gagtgatgta tatggcgctt 120 gctctcattg atagcaaaga acctcttgaa cgagggcgag ttttcaaaat tagggctgaa 180 gaccttgcag cgctcgccaa aatcacccca tcgcttgctt atcgacaatt aaaagagggt 240 ggtaaattac ttggtgccag caaaatttcg ctaagagggg atgatatcat tgctttagct 300 aaagagctta acctgccctt tactgctaaa aactccagcg aagagttaga tcttaacatt 360 attgagtgga tagcttattc aaatgatgaa ggatacttgt ctttaaaatt caccagaacc 420 atagaaccat atatctctag ccttattggg aaaaaaaata aattcacaac gcaattgtta 480 acggcaagct tacgcttaag tagccagtat tcatcttctc tttatcaact tatcaggaag 540 cattactcta attttaagaa gaaaaattat tttattattt ccgttgatga gttaaaggaa 600 gagttaatag cttatacttt tgataaagat ggaaatattg agtacaaata ccctgacttt 660 cctattttta aaagggatgt gttaaataaa gccattgctg aaattaaaaa gaaaacagaa 720 atatcgtttg ttggcttcac tgttcatgaa aaagaaggga gaaaaattag taagctgaag 780 ttcgaatttg tcgttgatga agatgaattt tctggcgata aagatgatga agcttttttt 840 atgaatttat ctgaagctga tgcagctttt ctcaaggtat ttgatgaaac cgtacctccc 900 aaaaaagcta aggggtga 918 <210> 30 <211> 45 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic oligonucleotide <400> 30 caaaagggcg ctgttatctg ataaggctta tctggtctca ttttg 45 <210> 31 <211> 45 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic oligonucleotide <400> 31 caaaaggggg ctgttatctg ataaggctta tctggtctca ttttg 45 <210> 32 <211> 38 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic oligonucleotide <400> 32 ggcgctgtta tctgataagg cttatctggt ctcatttt 38 <210> 33 <211> 60 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic oligonucleotide <400> 33 ctgctcaaaa agacgccaaa agggcgctgt tatctgataa ggcttatctg gtctcatttt 60 <210> 34 <211> 297 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polypeptide <400> 34 Met Ser Ala Val Leu Gln Arg Phe Arg Glu Lys Leu Pro His Lys Pro 1 5 10 15 Tyr Cys Thr Asn Asp Phe Ala Tyr Gly Val Arg Ile Leu Pro Lys Asn 20 25 30 Ile Ala Ile Leu Ala Arg Phe Ile Gln Gln Asn Gln Pro His Ala Leu 35 40 45 Tyr Trp Leu Pro Phe Asp Val Asp Arg Thr Gly Ala Ser Ile Asp Trp 50 55 60 Ser Asp Arg Asn Cys Pro Ala Pro Asn Ile Thr Val Lys Asn Pro Arg 65 70 75 80 Asn Gly His Ala His Leu Leu Tyr Ala Leu Ala Leu Pro Val Arg Thr 85 90 95 Ala Pro Asp Ala Ser Ala Ser Ala Leu Arg Tyr Ala Ala Ala Ile Glu 100 105 110 Arg Ala Leu Cys Glu Lys Leu Gly Ala Asp Val Asn Tyr Ser Gly Leu 115 120 125 Ile Cys Lys Asn Pro Cys His Pro Glu Trp Gln Glu Val Glu Trp Arg 130 135 140 Glu Glu Pro Tyr Thr Leu Asp Glu Leu Ala Asp Tyr Leu Asp Leu Ser 145 150 155 160 Ala Ser Ala Arg Arg Ser Val Asp Lys Asn Tyr Gly Leu Gly Arg Asn 165 170 175 Tyr His Leu Phe Glu Lys Val Arg Lys Trp Ala Tyr Arg Ala Ile Arg 180 185 190 Gln Gly Trp Pro Val Phe Ser Gln Trp Leu Asp Ala Val Ile Gln Arg 195 200 205 Val Glu Met Tyr Asn Ala Ser Leu Pro Val Pro Leu Ser Pro Ala Glu 210 215 220 Cys Arg Ala Ile Gly Lys Ser Ile Ala Lys Tyr Thr His Arg Lys Phe 225 230 235 240 Ser Pro Glu Gly Phe Ser Ala Val Gln Ala Ala Arg Gly Arg Lys Gly 245 250 255 Gly Thr Lys Ser Lys Arg Ala Ala Val Pro Thr Ser Ala Arg Ser Leu 260 265 270 Lys Pro Trp Glu Ala Leu Gly Ile Ser Arg Ala Thr Tyr Tyr Arg Lys 275 280 285 Leu Lys Cys Asp Pro Asp Leu Ala Lys 290 295 <210> 35 <211> 297 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polypeptide <400> 35 Met Ser Ala Val Leu Gln Arg Phe Arg Glu Lys Leu Pro His Lys Pro 1 5 10 15 Tyr Cys Thr Asn Asp Phe Ala Tyr Gly Val Arg Ile Leu Pro Lys Asn 20 25 30 Ile Ala Ile Leu Ala Arg Phe Ile Gln Gln Asn Gln Pro His Ala Leu 35 40 45 Tyr Trp Leu Pro Phe Asp Val Asp Arg Thr Gly Ala Ser Ile Asp Trp 50 55 60 Ser Asp Arg Asn Cys Pro Ala Pro Asn Ile Thr Val Lys Asn Pro Arg 65 70 75 80 Asn Gly His Ala His Leu Leu Tyr Ala Leu Ala Leu Pro Val Arg Thr 85 90 95 Ala Pro Asp Ala Ser Ala Ser Ala Leu Arg Tyr Ala Ala Ala Ile Glu 100 105 110 Arg Ala Leu Cys Glu Lys Leu Gly Ala Asp Val Asn Tyr Ser Gly Leu 115 120 125 Ile Cys Lys Asn Pro Cys His Pro Glu Trp Gln Glu Val Glu Trp Arg 130 135 140 Glu Glu Pro Tyr Thr Leu Asp Glu Leu Ala Asp Tyr Leu Asp Leu Ser 145 150 155 160 Ala Ser Ala Arg Arg Ser Val Asp Lys Asn Tyr Gly Leu Gly Arg Asn 165 170 175 Tyr His Leu Phe Glu Lys Val Arg Lys Trp Ala Tyr Arg Ala Ile Arg 180 185 190 Gln Asp Trp Pro Val Phe Ser Gln Trp Leu Asp Ala Val Ile Gln Arg 195 200 205 Val Glu Met Tyr Asn Ala Ser Leu Pro Val Pro Leu Ser Pro Ala Glu 210 215 220 Cys Arg Ala Ile Gly Lys Ser Ile Ala Lys Tyr Thr His Arg Lys Phe 225 230 235 240 Ser Pro Glu Gly Phe Ser Ala Val Gln Ala Ala Arg Gly Arg Lys Gly 245 250 255 Gly Thr Lys Ser Lys Arg Ala Ala Val Pro Thr Ser Ala Arg Ser Leu 260 265 270 Lys Pro Trp Glu Ala Leu Gly Ile Ser Arg Ala Thr Tyr Tyr Arg Lys 275 280 285 Leu Lys Cys Asp Pro Asp Leu Ala Lys 290 295 <210> 36 <211> 264 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polypeptide <400> 36 Met Ile Glu Gln Asp Gly Leu His Ala Gly Ser Pro Ala Ala Trp Val 1 5 10 15 Glu Arg Leu Phe Gly Tyr Asp Trp Ala Gln Gln Thr Ile Gly Cys Ser 20 25 30 Asp Ala Ala Val Phe Arg Leu Ser Ala Gln Gly Arg Pro Val Leu Phe 35 40 45 Val Lys Thr Asp Leu Ser Gly Ala Leu Asn Glu Leu Gln Asp Glu Ala 50 55 60 Ala Arg Leu Ser Trp Leu Ala Thr Thr Gly Val Pro Cys Ala Ala Val 65 70 75 80 Leu Asp Val Val Thr Glu Ala Gly Arg Asp Trp Leu Leu Leu Gly Glu 85 90 95 Val Pro Gly Gln Asp Leu Leu Ser Ser His Leu Ala Pro Ala Glu Lys 100 105 110 Val Ser Ile Met Ala Asp Ala Met Arg Arg Leu His Thr Leu Asp Pro 115 120 125 Ala Thr Cys Pro Phe Asp His Gln Ala Lys His Arg Ile Glu Arg Ala 130 135 140 Arg Thr Arg Met Glu Ala Gly Leu Val Asp Gln Asp Asp Leu Asp Glu 145 150 155 160 Glu His Gln Gly Leu Ala Pro Ala Glu Leu Phe Ala Arg Leu Lys Ala 165 170 175 Arg Met Pro Asp Gly Glu Asp Leu Val Val Thr His Gly Asp Ala Cys 180 185 190 Leu Pro Asn Ile Met Val Glu Asn Gly Arg Phe Ser Gly Phe Ile Asp 195 200 205 Cys Gly Arg Leu Gly Val Ala Asp Arg Tyr Gln Asp Ile Ala Leu Ala 210 215 220 Thr Arg Asp Ile Ala Glu Glu Leu Gly Gly Glu Trp Ala Asp Arg Phe 225 230 235 240 Leu Val Leu Tyr Gly Ile Ala Ala Pro Asp Ser Gln Arg Ile Ala Phe 245 250 255 Tyr Arg Leu Leu Asp Glu Phe Phe 260 <210> 37 <211> 237 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polypeptide <400> 37 Met Ser Thr Lys Lys Lys Pro Leu Thr Gln Glu Gln Leu Glu Asp Ala 1 5 10 15 Arg Arg Leu Lys Ala Ile Tyr Glu Lys Lys Lys Asn Glu Leu Gly Leu 20 25 30 Ser Gln Glu Ser Val Ala Asp Lys Met Gly Met Gly Gln Ser Gly Val 35 40 45 Gly Ala Leu Phe Asn Gly Ile Asn Ala Leu Asn Ala Tyr Asn Ala Ala 50 55 60 Leu Leu Thr Lys Ile Leu Lys Val Ser Val Glu Glu Phe Ser Pro Ser 65 70 75 80 Ile Ala Arg Glu Ile Tyr Glu Met Tyr Glu Ala Val Ser Met Gln Pro 85 90 95 Ser Leu Arg Ser Glu Tyr Glu Tyr Pro Val Phe Ser His Val Gln Ala 100 105 110 Gly Met Phe Ser Pro Lys Leu Arg Thr Phe Thr Lys Gly Asp Ala Glu 115 120 125 Arg Trp Val Ser Thr Thr Lys Lys Ala Ser Asp Ser Ala Phe Trp Leu 130 135 140 Glu Val Glu Gly Asn Ser Met Thr Ala Pro Thr Gly Ser Lys Pro Ser 145 150 155 160 Phe Pro Asp Gly Met Leu Ile Leu Val Asp Pro Glu Gln Ala Val Glu 165 170 175 Pro Gly Asp Phe Cys Ile Ala Arg Leu Gly Gly Asp Glu Phe Thr Phe 180 185 190 Lys Lys Leu Ile Arg Asp Ser Gly Gln Val Phe Leu Gln Pro Leu Asn 195 200 205 Pro Gln Tyr Pro Met Ile Pro Cys Asn Glu Ser Cys Ser Val Val Gly 210 215 220 Lys Val Ile Ala Ser Gln Trp Pro Glu Glu Thr Phe Gly 225 230 235 <210> 38 <211> 473 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polypeptide <400> 38 Met Asn Ile Lys Lys Phe Ala Lys Gln Ala Thr Val Leu Thr Phe Thr 1 5 10 15 Thr Ala Leu Leu Ala Gly Gly Ala Thr Gln Ala Phe Ala Lys Glu Thr 20 25 30 Asn Gln Lys Pro Tyr Lys Glu Thr Tyr Gly Ile Ser His Ile Thr Arg 35 40 45 His Asp Met Leu Gln Ile Pro Glu Gln Gln Lys Asn Glu Lys Tyr Lys 50 55 60 Val Pro Glu Phe Asp Ser Ser Thr Ile Lys Asn Ile Ser Ser Ala Lys 65 70 75 80 Gly Leu Asp Val Trp Asp Ser Trp Pro Leu Gln Asn Thr Asp Gly Thr 85 90 95 Val Ala Asn Tyr His Gly Tyr His Ile Val Phe Ala Leu Ala Gly Asp 100 105 110 Pro Lys Asn Ala Asp Asp Thr Ser Ile Tyr Met Phe Tyr Gln Lys Val 115 120 125 Gly Glu Thr Ser Ile Asp Ser Trp Lys Asn Ala Gly Arg Val Phe Lys 130 135 140 Asp Ser Asp Lys Phe Asp Ala Asn Asp Ser Ile Leu Lys Asp Gln Thr 145 150 155 160 Gln Glu Trp Ser Gly Ser Ala Thr Phe Thr Ser Asp Gly Lys Ile Arg 165 170 175 Leu Phe Tyr Thr Asp Phe Ser Gly Lys His Tyr Gly Lys Gln Thr Leu 180 185 190 Thr Thr Ala Gln Val Asn Val Ser Ala Ser Asp Ser Ser Leu Asn Ile 195 200 205 Asn Gly Val Glu Asp Tyr Lys Ser Ile Phe Asp Gly Asp Gly Lys Thr 210 215 220 Tyr Gln Asn Val Gln Gln Phe Ile Asp Glu Gly Asn Tyr Ser Ser Gly 225 230 235 240 Asp Asn His Thr Leu Arg Asp Pro His Tyr Val Glu Asp Lys Gly His 245 250 255 Lys Tyr Leu Val Phe Glu Ala Asn Thr Gly Thr Glu Asp Gly Tyr Gln 260 265 270 Gly Glu Glu Ser Leu Phe Asn Lys Ala Tyr Tyr Gly Lys Ser Thr Ser 275 280 285 Phe Phe Arg Gln Glu Ser Gln Lys Leu Leu Gln Ser Asp Lys Lys Arg 290 295 300 Thr Ala Glu Leu Ala Asn Gly Ala Leu Gly Met Ile Glu Leu Asn Asp 305 310 315 320 Asp Tyr Thr Leu Lys Lys Val Met Lys Pro Leu Ile Ala Ser Asn Thr 325 330 335 Val Thr Asp Glu Ile Glu Arg Ala Asn Val Phe Lys Met Asn Gly Lys 340 345 350 Trp Tyr Leu Phe Thr Asp Ser Arg Gly Ser Lys Met Thr Ile Asp Gly 355 360 365 Ile Thr Ser Asn Asp Ile Tyr Met Leu Gly Tyr Val Ser Asn Ser Leu 370 375 380 Thr Gly Pro Tyr Lys Pro Leu Asn Lys Thr Gly Leu Val Leu Lys Met 385 390 395 400 Asp Leu Asp Pro Asn Asp Val Thr Phe Thr Tyr Ser His Phe Ala Val 405 410 415 Pro Gln Ala Lys Gly Asn Asn Val Val Ile Thr Ser Tyr Met Thr Asn 420 425 430 Arg Gly Phe Tyr Ala Asp Lys Gln Ser Thr Phe Ala Pro Ser Phe Leu 435 440 445 Leu Asn Ile Lys Gly Lys Lys Thr Ser Val Val Lys Asp Ser Ile Leu 450 455 460 Glu Gln Gly Gln Leu Thr Val Asn Lys 465 470 <210> 39 <211> 305 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polypeptide <400> 39 Met Arg Leu Lys Val Met Met Asp Val Asn Lys Lys Thr Lys Ile Arg 1 5 10 15 His Arg Asn Glu Leu Asn His Thr Leu Ala Gln Leu Pro Leu Pro Ala 20 25 30 Lys Arg Val Met Tyr Met Ala Leu Ala Leu Ile Asp Ser Lys Glu Pro 35 40 45 Leu Glu Arg Gly Arg Val Phe Lys Ile Arg Ala Glu Asp Leu Ala Ala 50 55 60 Leu Ala Lys Ile Thr Pro Ser Leu Ala Tyr Arg Gln Leu Lys Glu Gly 65 70 75 80 Gly Lys Leu Leu Gly Ala Ser Lys Ile Ser Leu Arg Gly Asp Asp Ile 85 90 95 Ile Ala Leu Ala Lys Glu Leu Asn Leu Ile Ser Thr Ala Lys Asn Ser 100 105 110 Ser Glu Glu Leu Asp Leu Asn Ile Ile Glu Trp Ile Ala Tyr Ser Asn 115 120 125 Asp Glu Gly Tyr Leu Ser Leu Lys Phe Thr Arg Thr Ile Glu Pro Tyr 130 135 140 Ile Ser Ser Leu Ile Gly Lys Lys Asn Lys Phe Thr Thr Gln Leu Leu 145 150 155 160 Thr Ala Ser Leu Arg Leu Ser Ser Gln Tyr Ser Ser Ser Leu Tyr Gln 165 170 175 Leu Ile Arg Lys His Tyr Ser Asn Phe Lys Lys Lys Asn Tyr Phe Ile 180 185 190 Ile Ser Val Asp Glu Leu Lys Glu Glu Leu Ile Ala Tyr Thr Phe Asp 195 200 205 Lys Asp Gly Asn Ile Glu Tyr Lys Tyr Pro Asp Phe Pro Ile Phe Lys 210 215 220 Arg Asp Val Leu Asn Lys Ala Ile Ala Glu Ile Lys Lys Lys Thr Glu 225 230 235 240 Ile Ser Phe Val Gly Phe Thr Val His Glu Lys Glu Gly Arg Lys Ile 245 250 255 Ser Lys Leu Lys Phe Glu Phe Val Val Asp Glu Asp Glu Phe Ser Gly 260 265 270 Asp Lys Asp Asp Glu Ala Phe Phe Met Asn Leu Ser Glu Ala Asp Ala 275 280 285 Ala Phe Leu Lys Val Phe Asp Glu Thr Val Pro Pro Lys Lys Ala Lys 290 295 300 Gly 305 <210> 40 <211> 303 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polypeptide <400> 40 Met Arg Leu Lys Val Met Met Asp Val Asn Lys Lys Thr Lys Ile Arg 1 5 10 15 His Arg Asn Glu Leu Asn His Thr Leu Ala Gln Leu Pro Leu Pro Ala 20 25 30 Lys Arg Val Met Tyr Met Ala Leu Ala Leu Ile Asp Ser Lys Glu Pro 35 40 45 Leu Glu Arg Gly Arg Val Phe Lys Ile Arg Ala Glu Asp Leu Ala Ala 50 55 60 Leu Ala Lys Ile Thr Pro Ser Leu Ala Tyr Arg Gln Leu Lys Glu Gly 65 70 75 80 Gly Lys Leu Leu Gly Ala Ser Lys Ile Ser Leu Arg Gly Asp Asp Ile 85 90 95 Ile Ala Leu Ala Lys Glu Leu Asn Leu Thr Ala Lys Asn Ser Ser Glu 100 105 110 Glu Leu Asp Leu Asn Ile Ile Glu Trp Ile Ala Tyr Ser Asn Asp Glu 115 120 125 Gly Tyr Leu Ser Leu Lys Phe Thr Arg Thr Ile Glu Pro Tyr Ile Ser 130 135 140 Ser Leu Ile Gly Lys Lys Asn Lys Phe Thr Thr Gln Leu Leu Thr Ala 145 150 155 160 Ser Leu Arg Leu Ser Ser Gln Tyr Ser Ser Ser Leu Tyr Gln Leu Ile 165 170 175 Arg Lys His Tyr Ser Asn Phe Lys Lys Lys Asn Tyr Phe Ile Ile Ser 180 185 190 Val Asp Glu Leu Lys Glu Glu Leu Ile Ala Tyr Thr Phe Asp Lys Asp 195 200 205 Gly Asn Ile Glu Tyr Lys Tyr Pro Asp Phe Pro Ile Phe Lys Arg Asp 210 215 220 Val Leu Asn Lys Ala Ile Ala Glu Ile Lys Lys Lys Thr Glu Ile Ser 225 230 235 240 Phe Val Gly Phe Thr Val His Glu Lys Glu Gly Arg Lys Ile Ser Lys 245 250 255 Leu Lys Phe Glu Phe Val Val Asp Glu Asp Glu Phe Ser Gly Asp Lys 260 265 270 Asp Asp Glu Ala Phe Phe Met Asn Leu Ser Glu Ala Asp Ala Ala Phe 275 280 285 Leu Lys Val Phe Asp Glu Thr Val Pro Pro Lys Lys Ala Lys Gly 290 295 300 <210> 41 <211> 305 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polypeptide <400> 41 Met Arg Leu Lys Val Met Met Asp Val Asn Lys Lys Thr Lys Ile Arg 1 5 10 15 His Arg Asn Glu Leu Asn His Thr Leu Ala Gln Leu Pro Leu Pro Ala 20 25 30 Lys Arg Val Met Tyr Met Ala Leu Ala Leu Ile Asp Ser Lys Glu Pro 35 40 45 Leu Glu Arg Gly Arg Val Phe Lys Ile Arg Ala Glu Asp Leu Ala Ala 50 55 60 Leu Ala Lys Ile Thr Pro Ser Leu Ala Tyr Arg Gln Leu Lys Glu Gly 65 70 75 80 Gly Lys Leu Leu Gly Ala Ser Lys Ile Ser Leu Arg Gly Asp Asp Ile 85 90 95 Ile Ala Leu Ala Lys Glu Leu Asn Leu Pro Phe Thr Ala Lys Asn Ser 100 105 110 Ser Glu Glu Leu Asp Leu Asn Ile Ile Glu Trp Ile Ala Tyr Ser Asn 115 120 125 Asp Glu Gly Tyr Leu Ser Leu Lys Phe Thr Arg Thr Ile Glu Pro Tyr 130 135 140 Ile Ser Ser Leu Ile Gly Lys Asn Lys Phe Thr Thr Gln Leu Leu 145 150 155 160 Thr Ala Serving Leu Arg Serving Leu Serving Gln Tyr Serving Serving Leu Serving Tyr Gln 165 170 175 Leu Ile Arg Lys His Tyr Ser Asn Phe Lys Lys Asn Tyr Phe Ile 180 185 190 Val Asp Glu Lys Glu Glu Leu Ile Tyr Thr Phe Asp 195 200 205 Lys Asp Gly Asn Ile Glu Tyr Lys Tyr Pro Asp Phe Pro Ile Phe Lys 210 215 220 Arg Asp Val Leu Asn Lys Lys Ala Ile Ala Glu Ile Lys Lys Lys Thr Glu 225 230 235 240 Ile Ser Phe Val Gly Phe Thr Val His Glu Lys Glu Gly Arg Lys Ile 245 250 255 Ser Lys Leu Lys Phe Glu Phe Val Val Asp Glu Asp Glu Phe Ser Gly 260 265 270 Asp Lys Asp Asp Glu Ala Phe Phe Met Asn Leu Ser Glu Ala Asp Ala 275 280 285 Ala Phe Leu Lys Val Phe Asp Glu Thr Val Pro Pro Lys Lys Ala Lys 290 295 300 Gly 305 <210> 42 <211> 305 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polypeptide <400> 42 Met Arg Leu Lys Val Met Met Asp Val Asn Lys Lys Thr Lys Ile Arg 1 5 10 15 His Arg Asn Glu Leu Asn His Thr Leu Ala Gln Leu Pro Leu Pro Ala 20 25 30 Lys Arg Val Met Tyr Met Ala Leu Ala Leu Ile Asp Ser Lys Glu Pro 35 40 45 Leu Glu Arg Gly Arg Val Phe Lys Ile Arg Ala Glu Asp Leu Ala Ala 50 55 60 Leu Ala Lys Ile Thr Pro Ser Leu Ala Tyr Arg Gln Leu Lys Glu Gly 65 70 75 80 Gly Lys Leu Leu Gly Ala Ser Lys Ile Ser Leu Arg Gly Asp Asp Ile 85 90 95 Ile Ala Leu Ala Lys Glu Leu Asn Leu Leu Ser Thr Ala Lys Asn Ser 100 105 110 Pro Glu Glu Leu Asp Leu Asn Ile Ile Glu Trp Ile Ala Tyr Ser Asn 115 120 125 Asp Glu Gly Tyr Leu Ser Leu Lys Phe Thr Arg Thr Ile Glu Pro Tyr 130 135 140 Ile Ser Ser Leu Ile Gly Lys Lys Asn Lys Phe Thr Thr Gln Leu Leu 145 150 155 160 Thr Ala Ser Leu Arg Leu Ser Ser Gln Tyr Ser Ser Ser Leu Tyr Gln 165 170 175 Leu Ile Arg Lys His Tyr Ser Asn Phe Lys Lys Asn Tyr Phe Ile 180 185 190 Val Asp Glu Lys Glu Glu Leu Ile Tyr Thr Phe Asp 195 200 205 Lys Asp Gly Asn Ile Glu Tyr Lys Tyr Pro Asp Phe Pro Ile Phe Lys 210 215 220 Arg Asp Val Leu Asn Lys Lys Ala Ile Ala Glu Ile Lys Lys Lys Thr Glu 225 230 235 240 Ile Ser Phe Val Gly Phe Thr Val His Glu Lys Glu Gly Arg Lys Ile 245 250 255 Ser Lys Leu Lys Phe Glu Phe Val Val Asp Glu Asp Glu Phe Ser Gly 260 265 270 Asp Lys Asp Asp Glu Ala Phe Phe Met Asn Leu Ser Glu Ala Asp Ala 275 280 285 Ala Phe Leu Lys Val Phe Asp Glu Thr Val Pro Pro Lys Lys Ala Lys 290,295,300 Gly 305 <210> 43 <211> 281 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 43 ggcttgttgt ccacaaccgt taaaccttaa aagctttaaa agccttatat attctttttt 60 ttcttataaa acttaaaacc ttagaggcta tttaagttgc tgatttatat taattttatt 120 gttcaaacat gagagcttag tacgtgaaac atgagagctt agtacgttag ccatgagagc 180 ttagtacgtt agccatgagg gtttagttcg ttaaacatga gagcttagta cgttaaacat 240 gagagcttag tacgtactat caacaggttg aactgctgat c 281 <210> 44 <211> 281 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 44 ggcttgttgt ccacaaccat taaccttaa aagctttaa agccttatat attctttttt ttcttataaa acttaaaacc ttgaggcta tttaagttgc tgatttatat taattttatt gttcaaacat gagagcttag tacgtgaaac atgagagctt agtacattag ccatgagagc ttagtacatt agccatgagg gtttagttca ttaacatga gagcttagta cattaaacat gagagcttag tacatactat caacaggttg aactgctgat c <210> 45 <211> 260 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 45 aaaccttaaa acctttaaaa gccttatata ttcttttttt tcttataaa cttaaaacct 120. tagggctat ttagttgct gatttatatt aattttattg ttcaaacatg agagcttagt acatgaaaca tgagagctta gtacattagc catgagagct tagtacatta gccatgaggg 180 tttagttcat taaacatgag agcttagtac attaaacatg agagcttagt acatactatc 240 aacaggttga actgctgatc 260 <210> 46 <211> 389 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 46 tgtcagccgt taagtgttcc tgtgtcactg aaaattgctt tgagaggctc taagggcttc 60 tcagtgcgtt acatccctgg cttgttgtcc acaaccgtta aaccttaaaa gctttaaaag 120 ccttatatat tctttttttt cttataaaac ttaaaacctt agaggctatt taagttgctg 180 atttatatta attttattgt tcaaacatga gagcttagta cgtgaaacat gagagcttag 240 tacgttagcc atgagagctt agtacgttag ccatgagggt ttagttcgtt aaacatgaga 300 gcttagtacg ttaaacatga gagcttagta cgtgaaacat gagagcttag tacgtactat 360 caacaggttg aactgctgat cttcagatc 389 <210> 47 <211> 139 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 47 gtagaattgg taaagagagt cgtgtaaaat atcgagttcg cacatcttgt tgtctgatta 60 ttgatttttg gcgaaaccat ttgatcatat gacaagatgt gtatctacct taacttaatg 120 attttgataa aaatcatta 139 <210> 48 <211> 69 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic oligonucleotide <400> 48 tcgcacatct tgttgtctga ttattgattt ttggcgaaac catttgatca tatgacaaga 60 tgtgtatct 69 <210> 49 <211> 139 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 49 gtagaattgg taaagagagt tgtgtaaaat attgagttcg cacatcttgt tgtctgatta 60 ttgatttttg gcgaaaccat ttgatcatat gacaagatgt gtatctacct taacttaatg 120 attttgataa aaatcatta 139 <210> 50 <211> 466 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 50 gctagcccgc ctaatgagcg ggcttttttt tggcttgttg tccacaaccg ttaaacctta 60 aaagctttaa aagccttata tattcttttt tttcttataa aacttaaaac cttagaggct 120 atttaagttg ctgatttata ttaattttat tgttcaaaca tgagagctta gtacgtgaaa 180 catgagagct tagtacgtta gccatgagag cttagtacgt tagccatgag ggtttagttc 240 gttaaacatg agagcttagt acgttaaaca tgagagctta gtacgtacta tcaacaggtt 300 gaactgctga tccacgttgt ggtagaattg gtaaagagag tcgtgtaaaa tatcgagttc 360 gcacatcttg ttgtctgatt attgattttt ggcgaaacca tttgatcata tgacaagatg 420 tgtatctacc ttaacttaat gattttgata aaaatcatta ggtacc 466 <210> 51 <211> 439 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 51 gctagctggc ttgttgtcca caaccattaa accttaaaag ctttaaaagc cttatatatt 60 cttttttttc ttataaaact taaaacctta gaggctattt aagttgctga tttatattaa 120 ttttattgtt caaacatgag agcttagtac gtgaaacatg agagcttagt acattagcca 180 tgagagctta gtacattagc catgagggtt tagttcatta aacatgagag cttagtacat 240 taaacatgag agcttagtac atactatcaa caggttgaac tgctgatctg tacagtagaa 300 ttggtaaaga gagttgtgta aaatattgag ttcgcacatc ttgttgtctg attattgatt 360 tttggcgaaa ccatttgatc atatgacaag atgtgtatct accttaactt aatgattttg 420 ataaaaatca ttaggtacc 439 <210> 52 <211> 565 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 52 gaattcgagc tcggtacctc gcgaatgcat ctaggggacg gccgctagcc cgcctaatga 60 gcgggcttt ttttggcttg ttgtccacaa ccgttaaacc ttaaaagctt taaaagcctt 120 atatattctt ttttttctta taaaacttaa aaccttagag gctatttaag ttgctgattt 180 atattaattt tattgttcaa acatgagagc ttagtacgtg aaacatgaga gcttagtacg 240 ttagccatga gagcttagta cgttagccat gagggtttag ttcgttaaac atgagagctt 300 agtacgttaa acatgagagc ttagtacgta ctatcaacag gttgaactgc tgatccacgt 360 tgtggtagaa ttggtaaaga gagtcgtgta aaatatcgag ttcgcacatc ttgttgtctg 420 attattgatt tttggcgaaa ccatttgatc atatgacaag atgtgtatct accttaactt 480 aatgattttg ataaaaatca ttaggagcta gcattgggtc atcggatccc gggcccgtcg 540 actgcagagg cctgcatgca agctt 565 <210> 53 <211> 584 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 53 gaattcgagc tcggtacctc gcgaatgcat ctaggggacg gccgctagcc cgcctaatga 60 gcgggcttt ttttggcttg ttgtccacaa ccgttaaacc ttaaaagctt taaaagcctt 120 atatattctt ttttttctta taaaacttaa aaccttagag gctatttaag ttgctgattt 180 atattaattt tattgttcaa acatgagagc ttagtacgtg aaacatgaga gcttagtacg 240 ttagccatga gagcttagta cgttagccat gagggtttag ttcgttaaac atgagagctt 300 agtacgttaa acatgagagc ttagtacgta ctatcaacag gttgaactgc tgatccacgt 360 tgtggtagaa ttggtaaaga gagtcgtgta aaatatcgag ttcgcacatc ttgttgtctg 420 attattgatt tttggcgaaa ccatttgatc atatgacaag atgtgtatct accttaactt 480 aatgatttg ataaaaatca ttaggactag tcccgggcgc tagttattaa tattgggtca 540 tcggatcccg ggcccgtcga ctgcagaggc ctgcatgcaa gctt 584 <210> 54 <211> 557 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 54 gaattcgagc tcggtacctc gcgaatgcat ctaggggacg gccgctagcc cgcctaatga 60 gcgggctttt ttttggcttg ttgtccacaa ccgttaaacc ttaaaagctt taaaagcctt 120 atatattctt ttttttctta taaaacttaa aaccttagag gctatttaag ttgctgattt 180 atattaattt tattgttcaa acatgagagc ttagtacgtg aaacatgaga gcttagtacg 240 ttagccatga gagcttagta cgttagccat gagggtttag ttcgttaaac atgagagctt 300 agtacgttaa acatgagagc ttagtacgta ctatcaacag gttgaactgc tgatccacgt 360 tgtggtagaa ttggtaaaga gagtcgtgta aaatatcgag ttcgcacatc ttgttgtctg 420 attattgatt tttggcgaaa ccatttgatc atatgacaag atgtgtatct accttaactt 480 aatgattttg ataaaaatca ttaggtaccg agctcggatc ccgggcccgt cgactgcaga 540 ggcctgcatg caagctt 557 <210> 55 <211> 557 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 55 aagcttgcat gcaggcctct gcagtcgacg ggcccgggat ccgagctcgg tacctcgcga 60 atgcatctag gggacggccg ctagcccgcc taatgagcgg gctttttttt ggcttgttgt 120 ccacaaccgt taaaccttaa aagctttaaa agccttatat attctttttt ttcttataaa 180 acttaaaacc ttagaggcta tttaagttgc tgatttatat taattttatt gttcaaacat 240 gagagcttag tacgtgaaac atgagagctt agtacgttag ccatgagagc ttagtacgtt 300 agccatgagg gtttagttcg ttaaacatga gagcttagta cgttaaacat gagagcttag 360 tacgtactat caacaggttg aactgctgat ccacgttgtg gtagaattgg taaagagagt 420 cgtgtaaaat atcgagttcg cacatcttgt tgtctgatta ttgatttttg gcgaaaccat 480 ttgatcatat gacaagatgt gtatctacct taacttaatg atttgataa aaatcattag 540 gtaccgagct cgaattc 557 <210> 56 <211> 503 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 56 ggcgccgcta gccgcctaa tgagcgggct ttttttggc ttgttgtcca caaccgttaa 60 accttaaaag cttaaaagc cttatatatt cttttttttc ttaaaact taaaacctta 120 gaggctattt aagttgctga tttatattaa ttttattgtt caaacatgag agcttagtac 180 gtgaaacatg agagcttagt acgttagcca tgagagctta gtacgttagc catgagggtt 240 tagttcgtta aacatgagag cttagtacgt taaacatgag agcttagtac gtactatcaa 300 caggttgaac tgctgatcca cgttgtggta gaattggtaa agagagtcgt gtaaaatatc 360 gagttcgcac atcttgttgt ctgattattg atttttggcg aaaccatttg atcatatgac 420 aagatgtgta tctaccttaa cttaatgatt ttgataaaaa tcattaggta ccacatgtcc 480 tgcagaggcc tgcatgcaag ctt 503 <210> 57 <211> 539 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 57 gaattctgca gatatccatc acactggcgg ccgctagccc gcctaatgag cgggcttttt 60 tttggcttgt tgtccacaac cgttaaacct taaaagcttt aaaagcctta tatattcttt 120 tttttcttat aaaacttaaa accttagagg ctatttaagt tgctgattta tattaatttt 180 attgttcaaa catgagagct tagtacgtga aacatgagag cttagtacgt tagccatgag 240 agcttagtac gttagccatg agggtttagt tcgttaaaca tgagagctta gtacgttaaa 300 catgagagct tagtacgtac tatcaacagg ttgaactgct gatccacgtt gtggtagaat 360 tggtaaagag agtcgtgtaa aatatcgagt tcgcacatct tgttgtctga ttattgattt 420 ttggcgaaac catttgatca tatgacaaga tgtgtatcta ccttaactta atgattttga 480 taaaaatcat taggtaccgg gccccccctc gatcgaggtc gacggtatcg ggggagctc 539 <210> 58 <211> 500 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 58 gcgcgcagcc ttaattaagc tagcccgcct aatgagcggg cttttttttg gcttgttgtc 60 cacaaccgtt aaaccttaaa agctttaaaa gccttatata ttcttttttt tcttataaaa 120 cttaaaacct tagaggctat ttaagttgct gatttatatt aattttattg ttcaaacatg 180 agagcttagt acgtgaaaca tgagagctta gtacgttagc catgagagct tagtacgtta 240 gccatgaggg tttagttcgt taaacatgag agcttagtac gttaaacatg agagcttagt 300 acgtactatc aacaggttga actgctgatc cacgttgtgg tagaattggt aaagagagtc 360 gtgtaaaata tcgagttcgc acatcttgtt gtctgattat tgatttttgg cgaaaccatt 420 tgatcatatg acaagatgtg tatctacctt aacttaatga ttttgataaa aatcattagg 480 taccttaatt aactgcgcgc 500 <210> 59 <211> 530 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 59 aagcttgcat gcaggcctct gcagtcgacg ggccctctag actcgagctg gcttgttgtc 60 cacaaccatt aaaccttaaa agctttaaaa gccttatata ttcttttttt tcttataaaa 120 cttaaaacct tagaggctat ttaagttgct gatttatatt aattttattg ttcaaacatg 180 agagcttagt acgtgaaaca tgagagctta gtacattagc catgagagct tagtacatta 240 gccatgaggg tttagttcat taaacatgag agcttagtac attaaacatg agagcttagt 300 acatactatc aacaggttga actgctgatc tgtacagtag aattggtaaa gagagttgtg 360 taaaatattg agttcgcaca tcttgttgtc tgattattga tttttggcga aaccatttga 420 tcatatgaca agatgtgtat ctaccttaac ttaatgattt tgataaaaat cattaggtac 480 cgctagcggc cgtcccctag atgcattcgc gaggtaccga gctcgaattc 530 <210> 60 <211> 303 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <400> 60 ggcttgttgt ccacaaccgt taaaccttaa aagctttaaa agccttatat attctttttt 60 ttcttataaa acttaaaacc ttagaggcta tttaagttgc tgatttatat taattttatt 120 gttcaaacat gagagcttag tacgtgaaac atgagagctt agtacgttag ccatgagagc 180 ttagtacgtt agccatgagg gtttagttcg ttaaacatga gagcttagta cgttaaacat 240 gagagcttag tacgttaaac atgagagctt agtacgtact atcaacaggt tgaactgctg 300 atc 303 <210> 61 <211> 304 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polypeptide <400> 61 Met Arg Leu Lys Val Met Met Asp Val Asn Lys Lys Thr Lys Ile Arg 1 5 10 15 His Arg Asn Glu Leu Asn His Thr Leu Ala Gln Leu Pro Leu Pro Ala 20 25 30 Lys Arg Val Met Tyr Met Ala Leu Ala Leu Ile Asp Ser Lys Glu Pro 35 40 45 Leu Glu Arg Gly Arg Val Phe Lys Ile Arg Ala Glu Asp Leu Ala Ala 50 55 60 Leu Ala Lys Ile Thr Pro Ser Leu Ala Tyr Arg Gln Leu Lys Glu Gly 65 70 75 80 Gly Lys Leu Leu Gly Ala Ser Lys Ile Ser Leu Arg Gly Asp Asp Ile 85 90 95 Ile Ala Leu Ala Lys Glu Leu Asn Leu Leu Ser Thr Ala Lys Asn Ser 100 105 110 Ser Glu Glu Leu Asp Leu Asn Ile Ile Glu Trp Ile Ala Tyr Ser Asn 115 120 125 Asp Glu Gly Tyr Leu Ser Leu Lys Phe Thr Arg Thr Ile Glu Pro Tyr 130 135 140 Ile Ser Ser Leu Ile Gly Lys Asn Lys Phe Thr Thr Gln Leu Leu 145 150 155 160 Thr Ala Serving Leu Arg Serving Leu Serving Gln Tyr Serving Serving Leu Serving Tyr Gln 165 170 175 Leu Ile Arg Lys His Tyr Ser Asn Phe Lys Lys Asn Tyr Phe Ile 180 185 190 Val Asp Glu Lys Glu Glu Leu Ile Tyr Thr Phe Asp 195 200 205 Lys Asp Gly Asn Ile Glu Tyr Lys Tyr Pro Asp Phe Pro Ile Phe Lys 210 215 220 Arg Asp Val Leu Asn Lys Lys Ala Ile Ala Glu Ile Lys Lys Lys Thr Glu 225 230 235 240 Ile Ser Phe Val Gly Phe Thr Val His Glu Lys Glu Gly Arg Lys Ile 245 250 255 Ser Lys Leu Lys Phe Glu Phe Val Val Asp Glu Asp Glu Phe Ser Gly 260 265 270 Asp Lys Asp Asp Glu Ala Phe Phe Met Asn Leu Ser Glu Ala Asp Ala 275 280 285 Ala Phe Leu Lys Val Phe Asp Glu Thr Val Pro Pro Lys Lys Ala Lys 290 295 300 <210> 62 <211> 304 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polypeptide <400> 62 Met Arg Leu Lys Val Met Met Asp Val Asn Lys Lys Thr Lys Ile Arg 1 5 10 15 His Arg Asn Glu Leu Asn His Thr Leu Ala Gln Leu Pro Leu Pro Ala 20 25 30 Lys Arg Val Met Tyr Met Ala Leu Ala Leu Ile Asp Ser Lys Glu Pro 35 40 45 Leu Glu Arg Gly Arg Val Phe Lys Ile Arg Ala Glu Asp Leu Ala Ala 50 55 60 Leu Ala Lys Ile Thr Pro Ser Leu Ala Tyr Arg Gln Leu Lys Glu Gly 65 70 75 80 Gly Lys Leu Leu Gly Ala Ser Lys Ile Ser Leu Arg Gly Asp Asp Ile 85 90 95 Ile Ala Leu Ala Lys Glu Leu Asn Leu Leu Ser Thr Ala Lys Asn Ser 100 105 110 Pro Glu Glu Leu Asp Leu Asn Ile Ile Glu Trp Ile Ala Tyr Ser Asn 115 120 125 Asp Glu Gly Tyr Leu Ser Leu Lys Phe Thr Arg Thr Ile Glu Pro Tyr 130 135 140 Ile Ser Ser Leu Ile Gly Lys Lys Asn Lys Phe Thr Thr Gln Leu Leu 145 150 155 160 Thr Ala Ser Leu Arg Leu Ser Ser Gln Tyr Ser Ser Ser Leu Tyr Gln 165 170 175 Leu Ile Arg Lys His Tyr Ser Asn Phe Lys Lys Asn Tyr Phe Ile 180 185 190 Val Asp Glu Lys Glu Glu Leu Ile Tyr Thr Phe Asp 195 200 205 Lys Asp Gly Asn Ile Glu Tyr Lys Tyr Pro Asp Phe Pro Ile Phe Lys 210 215 220 Arg Asp Val Leu Asn Lys Lys Ala Ile Ala Glu Ile Lys Lys Lys Thr Glu 225 230 235 240 Ile Ser Phe Val Gly Phe Thr Val His Glu Lys Glu Gly Arg Lys Ile 245 250 255 Ser Lys Leu Lys Phe Glu Phe Val Val Asp Glu Asp Glu Phe Ser Gly 260 265 270 Asp Lys Asp Asp Glu Ala Phe Phe Met Asn Leu Ser Glu Ala Asp Ala 275 280 285 Ala Phe Leu Lys Val Phe Asp Glu Thr Val Pro Pro Lys Lys Ala Lys 290,295,300 <210> 63 <211> 1040 <212> DNA <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polynucleotide <220> <221> CDS <222> (3)..(1025) <400> 63 tc atg aag aaa cct gaa ctc aca gca act tct gtt gag aag ttt ctc 47 Met Lys Lys Pro Glu Leu Thr Ala Thr Ser Val Glu Lys Phe Leu 1 5 10 15 att gaa aaa ttt gat tct gtt tct gat ctc atg cag ctg tct gaa ggt 95 Glu Lys Phe Asp Ser Val Ser Asp Leu Met Gln Leu Glu Gly Ser 20 25 30 gaa gaa agc aga gcc ttt tct tt gat gtt gga gga aga ggt tat gtt 143 Glu Glu Ser Arg Ala Phe Ser Phe Asp Val Gly Gly Arg Gly Tyr Val 35 40 45 ctg agg gtc aat tct tgt gct gat ggt ttt tac aaa gac aga tat gtt 191 Leu Arg Val Is Cys Ala Asp Gly Phe Tyr Lys Asp Arg Tyr Val 50 55 60 tac aga cac ttt gcc tct gct gct ctg cca att cca gaa gtt ctg gac 239 Tyr Arg His Phe Ala Ser Ala Ala Leu Pro Ile Pro Glu Val Leu Asp 65 70 75 att gga gaa ttt tct gaa tct ctc acc tac tgc atc agc aga aga gca It contains Gly Glu Phe containing Glu Ser Leu Thr Tyr Cys containing Arg Arg Ala 80 85 90 95 caa gga gtc act ctc cag gat ctc cct gaa act gag ctg cca gct gtt 335 Gln Gly Val Thr Leu Gln Asp Leu Pro Glu Thr Glu Leu Pro Ala Val 100 105 110 ctg caa cct gtt gct gaa gca atg gat gcc att gca gca gct gat ctg 383 Leu Gln Pro Val Glu Al-Ala Met Asp L-A Ile L-A-A Asp Leu 115 120 125 agc caa acc tct gga ttt ggt cct ttt ggt ccc caa ggc att ggt cag 431 Ser Gln Thr Ser Gly Phe Gly Pro Phe Gly Pro Gln Gly Ile Gly Gln 130 135 140 tac acc act tgg agg gat ttc att tgt gcc att gct gat cct cat gtc 479 Tyr Thr Thr Trp Arg Asp Phe Ile Cys Ala Ile Ala Asp Pro His Val 145 150 155 tat cac tgg cag act gtg atg gat gac aca gtt tct gct tct gtt gct 527 Tyr His Trp Gln Thr Val Met Asp Asp Thr Val Ser Ala Ser Val Ala 160 165 170 175 cag gca ctg gat gaa ctc atg ctg tgg gca gaa gat tgt cct gaa gtc 575 Gln Ala Leu Asp Glu Leu Met Leu Trp Ala Glu Asp Cys Pro Glu Val 180 185 190 aga cac ctg gtc cat gct gat ttt gga agc aac aat gtt ctg aca gac 623 Arg His Leu Val His Ala Asp Phe Gly Ser Asn Asn Val Leu Thr Asp 195 200 205 aat ggc aga atc act gca gtc att gac tgg tct gaa gcc atg ttt gga 671 Asn Gly Arg Ile Thr Ala Val Ile Asp Trp Ser Glu Ala Met Phe Gly 210 215 220 gat tct caa tat gag gtt gcc aac att ttt ttt tgg aga cct tgg ctg 719 Asp Ser Gln Tyr Glu Val Ala Asn Ile Phe Phe Trp Arg Pro Trp Leu 225 230 235 gct tgc atg gaa caa caa aca aga tat ttt gaa aga aga cac cca gaa 767 Ala Cys Met Glu Gln Gln Thr Arg Tyr Phe Glu Arg Arg His Pro Glu 240 245 250 255 ctg gct ggt tcc ccc aga ctg aga gcc tac atg ctc aga to ggc ctg 815 Leu Ala Gly Ser Pro Arg Leu Arg Ala Tyr Met Leu Arg Ile Gly Leu 260 265 270 gac caa ctg tat caa tct ctg gtt gat gga aac ttt gat gat gct gct 863 Asp Gln Leu Tyr Gln Ser Leu Val Asp Gly Asn Phe Asp Asp Ala Ala 275 280 285 tgg gca caa gga aga tgt gat gcc att gtg agg tct ggt gct gga act 911 Trp Ala Gln Gly Arg Cys Asp Ala Ile Val Arg Ser Gly Ala Gly Thr 290 295 300 gtt gga aga act caa att gca aga agg tct gct gct gtt tgg act gat 959 Val Gly Arg Thr Gln Ile Ala Arg Arg Ser Ala Ala Val Trp Thr Asp 305 310 315 gga tgt gtt gaa gtt ctg gct gac tct gga aac agg aga ccc tcc aca 1007 Gly Cys Val Glu Val Leu Ala Asp Ser Gly Asn Arg Arg Pro Ser Thr 320 325 330 335 aga ccc aga gcc aag gaa tgaatattag ctagc 1040 Arg Pro Arg Ala Lys Glu 340 <210> 64 <211> 341 <212> PRT <213> Artificial Sequence <220> <223> Description of Artificial Sequence: Synthetic polypeptide <400> 64 Met Lys Lys Pro Glu Leu Thr Ala Thr Ser Val Glu Lys Phe Leu Ile 1 5 10 15 Glu Lys Phe Asp Ser Val Ser Asp Leu Met Gln Leu Ser Glu Gly Glu 20 25 30 Glu Ser Arg Ala Phe Ser Phe Asp Val Gly Gly Arg Gly Tyr Val Leu 35 40 45 Arg Val Asn Ser Cys Ala Asp Gly Phe Tyr Lys Asp Arg Tyr Val Tyr 50 55 60 Arg His Phe Ala Ser Ala Ala Leu Pro Ile Pro Glu Val Leu Asp Ile 65 70 75 80 Gly Glu Phe Ser Glu Ser Leu Thr Tyr Cys Ile Ser Arg Arg Ala Gln 85 90 95 Gly Val Thr Leu Gln Asp Leu Pro Glu Thr Glu Leu Pro Ala Val Leu 100 105 110 Gln Pro Val Ala Glu Ala Met Asp Ala Ile Ala Ala Ala Asp Leu Ser 115 120 125 Gln Thr Ser Gly Phe Gly Pro Phe Gly Pro Gln Gly Ile Gly Gln Tyr 130 135 140 Thr Thr Trp Arg Asp Phe Ile Cys Ala Ile Ala Asp Pro His Val Tyr 145 150 155 160 His Trp Gln Thr Val Met Asp Asp Thr Val Ser Ala Ser Val Ala Gln 165 170 175 Ala Leu Asp Glu Leu Met Leu Trp Ala Glu Asp Cys Pro Glu Val Arg 180 185 190 His Leu Val His Ala Asp Phe Gly Ser Asn Asn Val Leu Thr Asp Asn 195 200 205 Gly Arg Ile Thr Ala Val Ile Asp Trp Ser Glu Ala Met Phe Gly Asp 210 215 220 Ser Gln Tyr Glu Val Ala Asn Ile Phe Phe Trp Arg Pro Trp Leu Ala 225 230 235 240 Cys Met Glu Gln Gln Thr Arg Tyr Phe Glu Arg Arg His Pro Glu Leu 245 250 255 Ala Gly Ser Pro Arg Leu Arg Ala Tyr Met Leu Arg Ile Gly Leu Asp 260 265 270 Gln Leu Tyr Gln Ser Leu Val Asp Gly Asn Phe Asp Asp Ala Ala Trp 275 280 285 Ala Gln Gly Arg Cys Asp Ala Ile Val Arg Ser Gly Ala Gly Thr Val 290 295 300 Gly Arg Thr Gln Ile Ala Arg Arg Ser Ala Ala Val Trp Thr Asp Gly 305 310 315 320 Cys Val Glu Val Leu Ala Asp Ser Gly Asn Arg Arg Pro Ser Thr Arg 325 330 335 Pro Arg Ala Lys Glu 340 <210> 65 <211> 1734 <212> DNA <213> Escherichia coli <400> 65 gtgaaacaac agatacaact tcgtcgccgt gaagtcgatg aaacggcaga cttgcccgct 60 gaattgcctc ccttgctgcg ccgtttatac gccagccggg gagtacgcag tgcgcaagaa 120 ctggaacgca gtgttaaagg tatgctgccc tggcagcaac tgagcggcgt cgaaaaggcc 180 gttgagatcc tttacaacgc ttttcgcgaa ggaacgcgga ttattgtggt cggtgatttc 240 gacgccgacg gcgcgaccag cacggctcta agcgtgctgg cgatgcgctc gcttggttgc 300 agcaatatcg actacctggt accaaaccgt ttcgaagacg gttacggctt aagcccggaa 360 gtggtcgatc aggcccatgc ccgtggcgcg cagttaattg tcacggtgga taacggtatt 420 tcctcccatg cgggggttga gcacgctcgc tcgttgggca tcccggttat tgttaccgat 480 caccatttgc caggcgacac attacccgca gcggaagcga tcattaaccc taacttgcgc 540 gactgtaatt tcccgtcgaa atcactggca ggcgtgggtg tggcgtttta tctgatgctg 600 gcgctgcgca cctttttgcg cgatcagggc tggtttgatg agcgtaacat cgcaattcct 660 aacctggcag aactgctgga tctggtcgcg ctggggacag tggcggacgt cgtgccgctg 720 gacgctaata atcgcattct gacctggcag gggatgagtc gcatccgagc cggaaagtgc 780 cgtccgggga ttaaagcgct gcttgaagtg gcaaaccgtg atgcacaaaa actcgccgcc 840 agcgatttag gttttgcgct ggggccacgt ctcaatgctg ccggacgact ggacgatatg 900 tccgtcggtg tggcgctgtt gttgtgcgac aacatcggcg aagcgcgcgt gctggcaaat 960 gaactcgatg cgctaaacca gacgcgaaaa gagatcgaac aaggaatgca aattgaagcc 1020 ctgaccctgt gcgagaaact ggagcgcagc cgtgacacgc tacccggcgg gctggcaatg 1080 tatcaccccg aatggcatca gggcgttgtc ggtattctgg cttcgcgcat caaagagcgt 1140 tttcaccgtc cggttatcgc gtttgcgcca gcaggtgacg gtacgctgaa aggttccggt 1200 cgctccattc aggggctgca tatgcgtgat gcgctggagc gattagacac actctaccct 1260 ggcatgatgc tgaagtttgg cggtcatgcg atggcggcgg gtttgtcgct ggaagaggat 1320 aaattcaaac tctttcaaca acggtttggc gaactggtta ctgagtggct ggacccttcg 1380 ctattgcaag gcgaagtggt atcagacggt ccgttaagcc cggccgaaat gaccatggaa 1440 gtggcgcagc tgctgcgcga tgctggcccg tgggggcaga tgttcccgga gccgctgttt 1500 gacggtcatt tccgtctgct gcaacagcgg ctggtgggcg aacgtcattt gaaggtgatg 1560 gtcgaaccgg tcggcggcgg tccactgctg gatggtattg cttttaatgt cgataccgcc 1620 ctctggccgg ataacggcgt gcgcgaagtg caactggctt ataagctcga tatcaacgag 1680 tttcgcggca accgcagcct gcaaattatc atcgacaata tctggccaat ttag 1734
Claims
1. Engineered Escherichia coli (E. coli) host cells, wherein the engineered E. coli host cells include a gene knockout of at least one gene selected from the group consisting of SbcC and SbcD, the engineered E. coli host cells do not contain any engineered viability or yield reduction mutations in any of sbcB, recB, recD, and recJ, and the engineered E. coli host cells do not contain or produce the SbcCD complex.
2. The manipulated E. coli host cell according to claim 1, wherein the manipulated E. coli host cell contains no mutations or manipulated mutations in any of sbcB, recB, recD, and recJ.
3. The manipulated E. coli host cell according to claim 1 or 2, wherein the gene knockout comprises a knockout of SbcC.
4. The manipulated E. coli host cell according to any one of claims 1 to 3, wherein the gene knockout comprises a knockout of SbcD.
5. The manipulated E. coli host cell according to any one of claims 1 to 4, wherein the manipulated E. coli host cell is derived from a cell line selected from the group consisting of DH5α, DH1, JM107, JM108, JM109, MG1655, and XL1Blue.
6. The manipulated E. coli host cell according to any one of claims 1 to 5, wherein the manipulated E. coli host cell further comprises a genomic antibiotic resistance marker.
7. A modified E. coli host cell according to any one of claims 1 to 6, further comprising a genomic nucleic acid sequence encoding a temperature-sensitive lambda repressor.
8. The manipulated E. coli host cell according to claim 7, wherein the temperature-sensitive lambda repressor is cITs857.
9. The manipulated E. coli host cell according to claim 7 or 8, wherein the temperature-sensitive lambda repressor is a chromosomal integration copy of the phage φ80 binding site of the arabinose-inducible CITs857 gene.
10. The engineered E. coli host cell according to any one of claims 1 to 9, further comprising a genomic nucleic acid sequence encoding a genome-expressed RNA-IN-regulated selectable marker.
11. The manipulated E. coli host cell according to any one of claims 1 to 10, further comprising a vector.
12. The manipulated E. coli host cell according to claim 11, wherein the vector comprises a nucleic acid sequence having reverse repeats.
13. The engineered E. coli host cell according to claim 11, wherein the vector comprises a nucleic acid sequence having at least one direct repeat or at least one reverse repeat, or does not comprise a palindrome, direct repeat, or reverse repeat.
14. The aforementioned vector, Optionally, AAV vectors including AAV ITR, Lentiviral vectors, lentiviral envelope vectors, or lentiviral packaging vectors, Retroviral vectors, retroviral envelope vectors, or retroviral packaging vectors, and plasmid A manipulated E. coli host cell according to any one of claims 11 to 13, selected from the above.
15. The engineered E. coli host cell according to any one of claims 11 to 14, wherein the vector further comprises an RNA-selectable marker comprising an RNA-OUT sequence or a functional variant thereof having at least 95%, at least 98%, at least 99%, or 100% sequence identity with a sequence selected from the group consisting of SEQ ID NO: 47 and SEQ ID NO: 49, and optionally the RNA-OUT antisense repressor RNA having at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO:
48.
16. The aforementioned vector, R6K, pUC, and Cole2, or Sequences having at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with sequences selected from the group consisting of SEQ ID NOs: 43, 44, 45, 46, 30, 31, 32, 33, and 22. The manipulated E. coli host cell according to any one of claims 11 to 15, further comprising a bacterial replication origin selected from.
17. The engineered E. coli host cell according to any one of claims 11 to 16, wherein the vector is a Rep protein-dependent plasmid.
18. The engineered E. coli host cell according to any one of claims 11 to 17, wherein the vector is a eukaryotic pUC-free minicircle expression vector comprising (i) a eukaryotic region sequence encoding a gene of interest and having 5' and 3' ends, and (ii) a spacer region having less than 1000, preferably less than 500, base pairs in length, ligating the 5' and 3' ends of the eukaryotic region sequence and containing an R6K bacterial origin of replication and an RNA selectable marker.
19. The engineered E. coli host cell according to any one of claims 11 to 18, wherein the vector is a covalently bound closed circular plasmid having a backbone comprising a PolIII-dependent R6K replication origin and an RNA-OUT selectable marker, the backbone being less than 1000 bp, and an insert comprising a structured DNA sequence.
20. The manipulated E. coli host cell according to claim 19, wherein the structured DNA sequence is selected from the group consisting of inverted repetitive sequences, direct repetitive sequences, homopolymerized repetitive sequences, eukaryotic origins of replication, and eukaryotic promoter-enhancer sequences.
21. The manipulated E. coli host cell according to claim 19, wherein the structured DNA sequence is selected from the group consisting of polyA repeats, SV40 origin of replication, viral LTRs, lentiviral LTRs, retroviral LTRs, transposon IR / DR repeats, Sleeping Beauty transposon IR / DR repeats, AAV ITRs, CMV enhancers, and SV40 enhancers.
22. The manipulated E. coli host cell according to any one of claims 19 to 21, wherein the PolIII-dependent R6K replication origin has at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with a sequence selected from the group consisting of SEQ ID NO: 43, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 46, and SEQ ID NO:
60.
23. The manipulated E. coli host cell according to any one of claims 19 to 22, wherein the RNA-OUT selectable marker is an RNA-IN regulated RNA-OUT functional variant having at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 47 or SEQ ID NO:
49.
24. The manipulated E. coli host cell according to any one of claims 19 to 23, wherein the RNA-OUT antisense repressor RNA may have a sequence having at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO:
48.
25. A method for producing manipulated Escherichia coli (E. coli) cells, A method comprising knocking out at least one gene selected from the group consisting of SbcC and SbcD in starting E. coli cells, wherein none of sbcB, recB, recD, and recJ are engineered viability or yield reduction mutations, wherein the engineered E. coli host cells are free from or do not produce the SbcCD complex, or do not contain a functional SbcCD complex.
26. The method according to claim 25, wherein the starting E. coli cells do not contain any mutations or manipulated mutations in any of sbcB, recB, recD, and recJ.
27. The starting E. coli cells are free from any manipulated mutations or manipulated viability or yield reduction mutations in at least one of the following: uvrC, mcrA, mcrBC-hsd-mrr, and combinations thereof. The method according to claim 25 or 26, wherein the step of knocking out at least one gene does not result in any mutation in at least one of uvrC, mcrA, mcrBC-hsd-mrr, and combinations thereof.