cssDNA-producing hosts and phagemids
A phagemid-integrated production strain optimizes cssDNA production by reducing reliance on helper virions and plasmids, addressing inefficiencies in existing methods to achieve high-yield, high-fidelity ssDNA and cssDNA production.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- カノ セラピューティクス インコーポレイテッド
- Filing Date
- 2024-03-29
- Publication Date
- 2026-05-01
AI Technical Summary
Existing methods struggle to produce large-scale, high-fidelity circular single-stranded DNA (cssDNA) and single-stranded DNA (ssDNA) exceeding 1 kb with enzymatic or chemical processes, leading to inefficiencies and metabolic waste.
A production strain comprising a phagemid and integrated phage coding sequences, which reduces reliance on helper virions and plasmids, eliminates antibiotic resistance genes, and optimizes phage protein expression through genomic modifications and differential regulation, enabling cssDNA production without batch-to-batch inconsistency.
The solution enhances cssDNA production efficiency by minimizing replication-competent phages and metabolic burden, allowing for cost-effective, high-yield production of ssDNA and cssDNA.
Smart Images

Figure 2026513790000001_ABST
Abstract
Description
Technical Field
[0001] Cross - References to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 493,492, filed Mar. 31, 2023, and U.S. Provisional Patent Application No. 63 / 605,225, filed Dec. 1, 2023, the entire contents of each of which are incorporated herein by reference.
Background Art
[0002] Background It is difficult to produce large - scale high - fidelity cssDNA (circular single - stranded DNA) and ssDNA (single - stranded DNA) of custom arrays exceeding 1 kb by enzymatic or chemical processes.
[0003] Reference to Electronic Sequence Listing The content of the electronic sequence listing (304832000340SEQLIST.xml; size: 61,620 bytes; and creation date: Mar. 28, 2024) is incorporated herein by reference in its entirety.
Summary of the Invention
Means for Solving the Problems
[0004] Summary This specification describes improvements in the development and production of ssDNA (single - stranded DNA) and cssDNA (circular single - stranded DNA).
[0005] In particular, this disclosure describes a production strain comprising a phagemid and at least two phage coding sequences incorporated into the production strain's genome. Some production strains may include at least three, four, five, six, seven, eight, nine, ten, or eleven phage coding sequences in their genome. Incorporation of coding sequences into the production strain's genome offers improvements, including a reduction in the likelihood of replication-competent phages being generated and / or mitigation of batch-to-batch inconsistency issues thanks to reduced reliance on helper virions and helper plasmids. Alternatively or additionally, this approach can reduce the metabolic burden on the production host by eliminating the need to maintain extrachromosomal replication helper plasmids, which have traditionally been maintained by the addition of antibiotics to the culture medium.
[0006] cssDNA production systems that do not contain non-endogenous antibiotic resistance genes are also described herein. In some embodiments, neither the production strain nor the phagemid contained within the production strain contains non-endogenous antibiotic resistance coding sequences. Those skilled in the art understand that including antibiotic resistance coding sequences in recombinant DNA constructs aids many molecular biological techniques. However, this disclosure recognizes certain drawbacks of such approaches. For example, including such sequences may be undesirable for reasons of environmental and human safety. There are also metabolic costs to the production host from expressing antibiotic resistance sequences, and the cost of exposure to antibiotics even if the strain is resistant to them. In certain embodiments, the cssDNA production systems provided by this disclosure do not utilize antibiotic resistance coding sequences in the production strain.
[0007] In certain embodiments, the production strain provided herein includes a phagemide comprising a packaging signal and a designed sequence and at least one selectable sequence selected from a sequence encoding a sequence that expresses a tRNA that associates with a non-natural amino acid, rather than an antibiotic resistance gene, that inhibits the expression of a gene that delays or halts bacterial growth when expressed in the absence of a nutrient requirement marker, an antitoxin, an RNA that inhibits the expression of a gene that delays or halts bacterial growth when expressed in the absence of a transcription factor, a transcription factor repressor that inhibits the expression of a gene that delays or halts bacterial growth when expressed in the absence of a transcription factor, a transcription activator that activates such a repressor, or a sequence that produces a sequence that associates with a non-natural amino acid as required by the production strain. In many cases, the production strain itself is genomically modified to complement the activity of the selectable sequence in the phagemide, resulting in the production strain being unable to survive without the presence of the phagemide.
[0008] Suitable production strains include those that can be produced, support phagemid replication, and are susceptible to infection by single-stranded filamentous bacteriophages from the Monodnaviria realm. Bacteria from the families Enterobacteriaceae, Pseudomonadaceae, Spirillaceae, Xanthomonadaceae, Clostridium, and Propionibacterium are potentially useful production strains. Those skilled in the art will understand that E. coli strains are particularly useful as production strains.
[0009] Exemplary phages that may be useful as sources of phage proteins include M13, Ff, Fd, Enterobacteriaceae phage F1 [EF068134], Enterobacteriaceae phage ID2, Enterobacteriaceae phage NL95 [AF059243], Enterobacteriaceae phage SP [X07489], Enterobacteriaceae phage TW28, Enterobacteriaceae phage Qbeta, Enterobacteriaceae phage Qβ [AY099114], Enterobacteriaceae phage M11 [AF059242], Enterobacteriaceae phage ST, Enterobacteriaceae phage TW18 [FJ483840], and Enterobacteriaceae phage VK, or their functional equivalents. Phage-derived packaging signals and other regulatory sequences can be used to produce phagemids as described herein. In some embodiments, the genome of a specific phage, excluding the packaging signal and phage origin of replication (ori), can be incorporated into the genome of the production strain. In such embodiments, helper plasmids or helper phages are not required, and since the phage genome lacks a phage ori and packaging signal, the phage genome is not packaged within the phage capsid.
[0010] In some embodiments, coding sequences integrated into the genome of phage proteins can be characterized by their activity compared to their intrinsic activity when expressed from regulatory sequences found in wild-type phages. If the regulatory sequences originate from a source different from the source of the coding sequence operably ligated in the expression cassette, the regulatory sequences may be referred to as non-intrinsic regulatory sequences, non-endogenous regulatory sequences, or synthesized regulatory sequences. Regulatory sequences found operably ligated to the coding sequence in the organism that is the source of the coding sequence may be referred to as intrinsic regulatory sequences or endogenous regulatory sequences. Those skilled in the art will understand that the activity of a protein can be altered by changing its amino acid sequence, for example, by sequence shortening or amino acid residue substitution. Activity can also be altered by increasing the expression level of a protein through any means known in the art for overexpression. Phage protein sequences integrated into the genome can be synthesized to be expressed differentially compared to each other and compared to their intrinsic expression. Differential expression includes not only increasing or decreasing the activity of a protein, but also optimizing cssDNA production by temporally expressing one phage protein at different times than another phage protein.
[0011] In some embodiments, the production strain includes phage coding sequences integrated in the genome that encode at least two of the following M13 phage proteins (or analogues): p1, p2, p3, p4, p5, p6, p7, p8, p9, p10, and p11 (see Figure 1, panels A and B). In some embodiments, the sequences integrated in the genome can be overexpressed compared to their native expression and / or the activity of one or more of the proteins can be altered. In certain embodiments, protein overexpression may be achieved by using a promoter, e.g., a synthetic promoter, an inducible promoter system, or a promoter known to result in high levels of expression.
[0012] In some embodiments, the activity of one or more phage proteins can be reduced compared to their native activity. For example, in some embodiments, the production strain may contain p3 and / or p5 phage proteins having lower activity compared to their native expression. In one exemplary embodiment described herein, the reduction in activity can be achieved through the use of p3 or p5 shortening.
[0013] As further described herein, the selection sequence can be a sequence that is either in RNA form or translated into a protein form, and / or a sequence that compensates for a defect or susceptibility in the production strain. In some cases, for the selection sequence in the phagemid to function, the production strain described herein needs to be genetically engineered to be either deficient or susceptible. In some embodiments, the modification of the production strain described herein can be achieved using the genomic modification lambda red recombineering system described in Example 1 below.
[0014] In certain embodiments, if the selection sequence is a toxin / antitoxin system, the selection sequence can be selected from ccdB / ccdA, hokA / sokA, pemK / pemI, mazF / mazE, ChpBK ChpBI, relE / relB, parE / parD, hipA / hipB, or other toxin / antitoxin systems in which the toxin is expressed from the host genome and the antitoxin is expressed from a phagemide. In certain embodiments, if the selection sequence is an RNAi that downregulates a counter-selectable sequence in the production strain, the RNAi sequence can be selected to interfere with the translation of HSVtk, Ura3, tetA, sacB, rpsL, pheS, pheS*, pheS**, thyA, lacY, gata-1, ccdB, hokA, pemK, mazF, chpBK, relE, parE, hipA, or other counter-selectable markers or toxins. Alternatively or additionally, in some embodiments, the selection sequence may include, or can be selected from, transcriptional repressors that can be used to downregulate the transcription of harmful sequences in the production strain. In some embodiments, the transcriptional repressor sequence may be selected from tetR, araC, lacI, xylS, or other sequences that reduce the expression of counterselectable markers or toxins. Similarly, in some embodiments, the selection sequence may be selected from any transcriptional activator known in the art. In some embodiments, transcriptional activators can be used to increase the transcription levels of genes necessary for the survival of the production strain. Examples of transcriptional activators include araC and xylR.
[0015] Particularly useful select sequences used in the methods and production systems described herein include nutritional requirement sequences, which are sequences encoding proteins required by the production strain host for proper cellular function. In some embodiments utilizing such sequences, the production strain host is created to require the product of the select sequence for optimal growth. A particularly useful select sequence is the pyrF gene, which can be used to support the growth of production strains having an inactivated pyrF gene (pyrF minus). In some embodiments, the production strain is cultured in a medium supplemented with uracil before transformation of the pyrF minus production strain with a phagemide containing the pyrF select sequence. After transformation with a phagemide containing the pyrF selectable sequence, the production strain is cultured in a medium without uracil supplementation. In some embodiments, the pyrF gene contains a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 98%, at least 99%, or at least 99.5% sequence identity with the sequence described in SEQ ID NO: 34. In some embodiments, the pyrF gene contains the sequence of SEQ ID NO: 34. Alternatively or additionally, in some embodiments, the nutrient-requiring selectable sequence may be adenine, cytosine, guanine, thymine, alanine, arginine, asparagine, aspartate, cysteine, glutamine, glutamate, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine, biotin, uridine-5'-monophosphate, pantothenate, xanthine, spermidine, para-aminobenzoate, lipoate (lipoa Genes may be selected from those that compensate for deficiencies in the production strain in one or more anabolic processes, such as te), nicotinamide riboside, nicotinamide mononucleotide, D-glucosamine, thiamine, shikimate, aminoethyl-phosphonate, beta-alanine, s-methyl-methionine, ornithine, indole, indoleacetic acid, L-threonine, L-threonine O-3-phosphate, nicotinate, ribosylnicotinamide, and agmatine, and / or for the synthesis of other essential compounds.
[0016] In some embodiments, the production strain includes, or achieves, variable expression of phage genes (e.g., phage genes expressed at levels different from those expressed in their native genome). Variable expression includes overexpression and reduced expression compared to the expression levels of phage genes in their native genome. In some embodiments, the ratio of gene expression between different phage genes is altered (e.g., altered from the native ratio of expression of native phage genes in their native genomes).
[0017] Those skilled in the art will be familiar with a variety of promoters that may be useful for altering the expression of a single or multiple phage genes. In some embodiments, a promoter can be selected to cause overexpression of a phage gene compared to the gene as it would be expressed from its native promoter (e.g., when incorporated and expressed from its native promoter, or when expressed from its native promoter in the native bacteriophage genome). Conversely, in some embodiments, a promoter can be selected to cause reduced expression compared to the expression of the native gene (e.g., when incorporated and expressed from its native promoter, or when expressed from its native promoter in the native bacteriophage genome). Alternatively or additionally, with respect to any of the foregoing, in some embodiments, a promoter can be selected to achieve a specific timing, pattern, or responsiveness of expression (e.g., to environmental or applied stimuli). In some embodiments, cssDNA production can be improved (e.g., optimized) for a given production strain and phagemid combination through the use of such promoters (e.g., heterologous promoters) and / or through the incorporation of some or all relevant phage genes in the production strain genome. Promoters particularly useful in certain embodiments include, for example, the standard T7 promoter, or mutant T7 promoters under the control of inducible T7 polymerase, the lacI promoter, the lacIq promoter, the araBAD promoter, the tet promoter, the temperature-sensitive promoter, the stress-responsive promoter, the quorum-sensing promoter, the photosensitive promoter, or other inducible or repressive promoters.
[0018] In some embodiments, the production strain provided in accordance with this disclosure may include a genomically integrated phage protein coding sequence incorporated at two or more loci in the production strain genome. The selection of loci in the production strain genome can affect the expression level of the phage coding sequence and can be selected to optimize the performance of the production strain. In some embodiments, the phage protein coding sequence may be incorporated genomically at at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 distinct loci in the genome.
[0019] Those skilled in the art will understand that, in addition to or as an alternative to manipulation of the regulatory sequences that control the expression of phage proteins, variants of naturally occurring phage proteins may be utilized in some embodiments. In some embodiments, the phage protein variant includes a sequence modified from its naturally occurring sequence. In some embodiments, the phage protein variant has modified expression, activity, and / or sequence compared to the naturally occurring phage protein. Those skilled in the art will be familiar with various techniques for modifying protein sequences (and / or the sequences of the nucleic acids that encode them), and will further understand that in some embodiments, the activity of proteins known in the art can be used to generate desired phage proteins with desired characteristics. Exemplary methods include random mutagenesis, rational design, assisted laboratory evolution, directional evolution, or other means. In some embodiments, modifications to proteins may be performed, followed by testing to determine whether the desired effects are achieved, for example, by measuring the yield, quality, and / or fidelity of cssDNA.
[0020] Alternatively or additionally, in some embodiments, the selection sequence may be, or may include, a synthetic tRNA sequence capable of recognizing (and, in some embodiments, necessary for recognizing) a recoded codon encoding a non-natural amino acid in at least one essential gene in a production strain as described herein. In such embodiments, the production strain cannot grow without the presence of a synthetic tRNA sequence that recognizes the non-natural amino acid and the recoded codon of the non-natural amino acid. For example, E. coli can recode in its genome so that the UAG codon becomes a dedicated codon for the non-natural amino acid L-4,4'-biphenylalanine (bipA, also known as the bipA chemical) when bipA aminoacyl-tRNA synthetase is expressed intracellularly. (Mandell et al., Biocontainment of genetically modified organisms by synthetic protein design. Nature 518, 55-60 (2015)). The UAG codon can be inserted into three essential genes so that these genes cannot be translated in the absence of either the bipA chemical or bipA tRNA. For example, the UAG codon can be inserted into the essential genes adenylate kinase (adk.d6), tyrosyl-tRNA synthetase (tyrS.d8), and BipA-dependent aminoacyl-tRNA synthetase for BipA aminoacylation (BipARS.d6). Kunjapur et al., Synthetic auxotrophy remains stable after continuous evolution and in co-culture with mammalian cells. Sci Adv. 2021 Jul 2;7(27):eabf5851. In this scenario, by expressing bipA tRNA from phagemids and adding the bipA chemical to the culture medium, a condition can be created in which only E. coli cells possessing phagemids can survive, ensuring that virtually all E. coli cells in a population contain phagemids without selection by antibiotics. Furthermore, these E. coli cells cannot survive outside of the specific culture medium to which the bipA chemical has been added, because they cannot utilize bipA in that environment. This prevents the produced production host from leaking out of the laboratory or fermentation facility environment.This also prevents the production host from contaminating fermentation vessels that may be used for other strains, for example, in a contract manufacturing organization (CMO), which is a concern considering that the production host may express phage particles that could theoretically infect other bacterial strains in the CMO.
[0021] In some embodiments, one or more phage proteins expressed from a helper plasmid, helper virus, or the genome of a production strain may further comprise a tag. The tag may be useful for identifying, quantifying, and / or separating phage particles. Preferred tags include fluorescent tags, luminescent tags, chromogenic tags, and affinity tags. Particularly useful affinity tags include biotin, his, myc, flag, CBP, GST, HA, HBH, MBP, S, and V5, or other affinity tags that assist in the purification of phage particles from production broth.
[0022] In particular, this disclosure provides a specific method for producing cssDNA. In certain embodiments, such a method includes culturing a production strain in a culture medium, the production strain genome containing at least two phage protein coding sequences integrated in the genome, and introducing a phagemide into the production strain. Those skilled in the art will be familiar with various techniques for introducing a phagemide into a production strain. Exemplary such techniques include, for example, conjugation, transduction, electroporation, chemical transformation, or by utilizing a natural phage infection mechanism. The resulting phage particles can be collected, and cssDNA can be collected. Exemplary separation means include centrifugation, gel electrophoresis, membrane separation, tangential flow filtration, and liquid chromatography. In some embodiments, the means for separating cssDNA includes substantially separating any dsDNA from the cssDNA. Those skilled in the art will be familiar with various techniques for quantifying cssDNA, such as absorption spectroscopy, fluorescent intercalator dyes, UV detection, gel electrophoresis, and microscopy. In embodiments in which the phage particles further include a tag, the tag can be used to assist in the separation and / or quantification steps.
[0023] Additional methods for producing cssDNA are also provided, for example, using a cssDNA production system. In some embodiments, the provided cssDNA production system uses a production strain that already contains a phage particle production sequence. An exemplary such method includes culturing the cssDNA system in a culture medium, and optionally inducing the production of one or more of the phage genes, collecting the phage particles, and separating the phage particles. In some embodiments, the collection and separation steps may be simultaneous, and in particular, when affinity tags are used, separation of the phage particle coat from the cssDNA is possible in a single step. [Brief explanation of the drawing]
[0024] [Figure 1]Figures 1A - B show a schematic diagram of an M13 phage particle showing the structure of the coat proteins around the ssDNA genome (Figure 1A). Figure 1B shows a map of the M13 phage genome identifying all genes, intergenic regions, promoters (uppercase P) and terminators (uppercase T). The coding region of each M13 protein is designated as "p#".
[0025] [Figure 2] Figure 2 shows a phagemid containing an F1 origin of replication; a selectable sequence; a user - defined sequence; and, optionally, a plasmid origin of replication.
[0026] [Figure 3] Figure 3 shows a production strain having at least two phage - encoded sequences integrated into the genome and a phagemid containing a phage origin of replication, a packaging signal, a selection sequence and a designed sequence.
[0027] [Figure 4] Figure 4 shows a map of plasmid M13KO7 containing bacteriophage M13 genes I - XI, a p15A origin of replication, an M13 origin of replication and a packaging signal, and a kanamycin resistance marker.
[0028] [Figure 5] Figure 5 shows a plasmid expressing a lambda red recombinase system containing beta, exo and gamma from an arabinose - inducible promoter (catalog number CAS9BAC1P, Sigma Aldrich, Burlington, Massachusetts).
[0029] [Figure 6]Figure 6 shows a schematic diagram of the accumulation of further cssDNA templates in E. coli cells, which leads to increased production of packaged cssDNA (Lee et al., Optimizing protein V untranslated region sequence in M13 phage for increased production of single-stranded DNA for origami. Nucleic Acids Res. 2021 Jun 21;49(11):6596-6603. This is incorporated herein by reference in its entirety).
[0030] [Figure 7] Figure 7 shows PCR primers designed to amplify the first transcription unit so that the 5' UTR (untranslated region) of gene V changes from wild-type TCACA to GAGGT (Figure 7, Panel A). Panel B of Figure 7 shows a schematic diagram of PCR primers designed to amplify the remainder of transcription unit 1 via fusion PCR so that the mutated 5' UTR of gene V is incorporated into transcription unit 1, creating a 2117 bp variation of transcription unit 1 containing the altered 5' UTR sequence of gene V.
[0031] [Figure 8] Figures 8A-8B show phagemide maps containing pUC replication origins (Figure 8A), as well as portions of the pUC ori sequences and primers used to generate inc1 and inc2 mutations.
[0032] [Figure 9] Figure 9 presents results demonstrating the effect of phagemid replication origins on cssDNA production.
[0033] [Figure 10] Figure 10 presents the results of larger-scale culture experiments demonstrating that inc1 and inc2 mutations increase cssDNA yield.
[0034] [Figure 11] Figure 11 presents the results of qPCR experiments to analyze the relative cssDNA yield using helper plasmids with different origins of replication.
[0035] [Figure 12] Figure 12 shows a comparison of the final OD of various potential bacterial production strains.
[0036] [Figure 13] Figure 13A shows verification of cssDNA production in the ungenerated BW25113 parent strain. Figure 13B shows results demonstrating the induction of pLac:T7 polymerase expression in the generated bacterial strain bac058 containing the incorporated pLac:T7 construct and emGFP reporter.
[0037] [Figure 14A] Figure 14A shows the results of screening small combinatorial libraries of embedded M13 TU1 and TU2 driven by promoters of different strengths.
[0038] [Figure 14B] Figure 14B shows the results of comparing the selected strains with the parental strains from which they were not produced.
[0039] [Figure 15A] Figure 15A shows a comparison between the best-produced production host (bac105) and its parental strain and conventional production strain DH5α, both using a conventional helper plasmid system.
[0040] [Figure 15B] Figure 15B shows the results regarding the production of various cssDNA sequences from the generated host strains.
[0041] [Figure 15C]Figure 15C shows a comparison of cssDNA production by the resulting bac105-producing strain after transformation with RFP / ampicillin phagemide cdsDNA111, with cssDNA production by its parent strain bac016 and the conventional production host bac001, as well as bac105 cssDNA production after transformation with RFP / pyrF phagemide cdsDNA117. [Modes for carrying out the invention]
[0042] Detailed explanation The production of ssDNA of a certain length is cumbersome, partly because the success rate of adding subsequent bases to the polymer decreases as the length of the DNA molecule increases during chemosynthesis. Similarly, the production of longer ssDNA polymers using PCR-based methods or via plasmids leads to the need to remove any region of dsDNA, thus complicating the procedure and resulting in metabolic waste. For example, in some cases, double-stranded DNA is produced at metabolic cost to the host organism using a typical plasmid amplification scheme, after which the undesirable strand is enzymatically removed. This post-assembly modification method, given the inaccuracies of the subsequent enzymatic removal step, leads to loss of cssDNA products and a wide range of ssDNA lengths. Also, synthesizing large amounts of dsDNA just to degrade one of the two strands is metabolically wasteful, meaning that half of the synthesized and polymerized deoxyribonucleotide triphosphate (dNTP) molecules are wasted.
[0043] Bacteriophages have been used to produce ssDNA and cssDNA, and the production of these products has been reported (Bush et al., Synthesis of DNA Origami Scaffolds: Current and Emerging Strategies, Molecules 2020, 25, 3386). However, with synthesized bacteriophages containing DNA not derived from natural bacteriophages, such as fluorescent protein tags, the cssDNA yield is significantly reduced compared to the production of natural bacteriophages. This creates room for improving the efficiency of such systems for producing ssDNA and cssDNA.
[0044] ssDNA and cssDNA are desirable for several known applications, and more applications are expected to be found as technologies related to DNA data preservation and DNA nanotechnology (i.e., DNA origami) grow. For example, dsDNA derived from salmon eggs has been shown to be an effective transparent type of sunscreen (Gasperini, et al., Non-ionising UV light increases the optical density of hygroscopic self assembled DNA crystal films 2017 Nature Scientific Reports 7: 6631). However, given the costs associated with its production, its commercial use is likely not practical. Such DNA can be produced in a cost-effective manner by combining synthetic biogenesis with efficient biotechnological production methods.
[0045] Definition: Unless otherwise specified, technical terms are used according to their conventional usage. Definitions of common terms in molecular biology can be found, for example, in Benjamin Lewin, Genes XII, Jones & Bartlett Learning; 12th edition (March 16, 2017) and other similar references. As used herein, the singular forms “a,” “an,” and “the” refer to both singular and plural unless the context explicitly indicates otherwise. For example, the term “a peptide” can be considered equivalent to the phrase “at least one peptide,” including one or more antigens. As used herein, the term “comprises” means “includes.” Thus, “comprising a protein” means “including a protein” without excluding other elements. Unless otherwise indicated, it should be further understood that any base size or amino acid size, and all molecular weight or molecular mass values given for nucleic acids or polypeptides are approximations and are provided for descriptive purposes. Many methods and materials similar to or equivalent to those described herein can be used, but certain preferred methods and materials are described herein. In case of any conflict, this specification shall prevail, including the definition of terms. In addition, materials, methods, and examples are illustrative and not intended to limit the scope. To facilitate an overview of the various embodiments, the following definitions of terms are provided.
[0046] As used herein, the term “equivalent” means two or more sets of actors, entities, situations, states, etc., that do not have to be identical to one another but are similar enough that comparison between them is possible, and therefore a person skilled in the art will understand that conclusions can be reasonably drawn based on observed differences or similarities. In some embodiments, equivalent sets of states, environments, individuals, or populations are characterized by several substantially identical features and one or a few varied features. A person skilled in the art will understand, in context, what degree of identity is required for two or more such sets of actors, entities, situations, states, etc., to be considered equivalent in any given environment. For example, a person skilled in the art will understand that sets of environments, individuals, or populations are equivalent to one another if they are characterized by a sufficient number and variety of substantially identical features to guarantee a reasonable conclusion that differences in results or observed phenomena under or using different sets of environments, individuals, or populations are caused by or exhibit variations in those varied features. For example, in some embodiments, a regulatory sequence is used to express a coding sequence in the resulting nucleic acid sequence, but the regulatory sequence used is not the same as the regulatory sequences found in natural genes expressing proteins in microorganisms. The use of a regulatory sequence different from those found in nature can be referred to as non-endogenous in comparison to the coding sequence. In such cases, the activity of a natural gene expression cassette can be compared to the resulting version of the gene created using a non-endogenous regulatory sequence. In such a comparison, if both expression cassettes are expressed under substantially the same experimental conditions, the non-endogenous regulatory sequence can be found to be more active or inactive compared to the natural gene expression cassette.
[0047] A “regulatory sequence” refers to a nucleic acid sequence that regulates the expression of a nucleic acid sequence to which it is operatively ligated. A regulatory sequence is operatively ligated to a nucleic acid sequence when it controls and regulates the transcription of that nucleic acid sequence, and, if necessary, translation. Therefore, the term “regulatory sequence” can refer to elements such as promoters, enhancers, transcriptional terminators, ribosome binding sites and start codons (ATGs) preceding protein-coding genes, intron splicing signals, maintenance of the correct reading frame of that gene to enable correct mRNA translation, and stop codons. The term “regulatory sequence” is intended to include at least components whose presence can influence expression, and may also include additional components whose presence is advantageous, such as leader sequences and fusion partner sequences. A regulatory sequence may include a promoter.
[0048] The term “designed sequence,” as used herein, refers to a nucleic acid sequence created to be included in a template cssDNA, template ssDNA, phagemide, or host cell genome. In some embodiments, the designed sequence, once introduced into a production strain, is intended to be produced by the production strain described herein. The designed sequence may include genome editing sequences, such as structural sequences for use in DNA origami, or any other desired nucleic acid sequence.
[0049] Generally, the terms “created” or “synthetic” refer to an artificially manipulated form. For example, a polynucleotide is considered “created” if two or more sequences that are not linked together in order in nature are artificially manipulated to be directly linked to one another in the created polynucleotide, and / or if certain residues in the polynucleotide are not naturally present and / or are artificially linked to entities or parts that are not linked in nature. For example, in some embodiments described and / or utilized herein, a created polynucleotide includes a regulatory sequence that is found in nature to be operationally related to a first coding sequence but not to a second coding sequence, and this is artificially linked to be operationally related to the second coding sequence. To the same extent, a polypeptide can be considered “created” if it is encoded by or expressed from a created polynucleotide, and / or produced by means other than natural expression in cells. Similarly, a cell or organism is considered “created” if it is subjected to manipulation such that its genetic, epigenetic, and / or phenotypic identity is altered compared to a suitable reference cell, e.g., a cell otherwise identical and not manipulated in this way. In some embodiments, the manipulation is or includes genetic manipulation such that its genetic information is altered (e.g., new genetic material that was not previously present is introduced, e.g., by transformation, mating, somatic hybridization, transfection, transduction or other mechanisms, or previously present genetic material is altered or removed, e.g., by substitution or deletion mutations or by a mating protocol). In some embodiments, the created cell is one that has been manipulated to contain and / or express a particular active agent of interest (e.g., a protein, nucleic acid, and / or a specific form thereof) in altered amounts and / or depending on the timing of the alteration, compared to such a suitable reference cell.In common practice, and as will be understood by those skilled in the art, a produced polynucleotide or cellular offspring is typically still referred to as “produced,” even if the actual operation was performed on a previous entity. Another example of the use of the term “synthetic” is synthetic nucleases, which are nucleases that further include domains directly or indirectly related to the coexisting sequence.
[0050] As used herein, "promoter" refers to the minimum sequence sufficient to direct transcription. A promoter is also part of a group of nucleic acid sequences commonly referred to as regulatory sequences. It also includes promoter elements sufficient to enable promoter-dependent gene expression to be controlled in a cell-type, tissue-specific manner, or to be induced by an external signal or activator; such elements may be located in the 5' or 3' region of the gene. Both constitutive and inductive promoters are included (see, e.g., Bitter et al., Methods in Enzymology 153:516-544, 1987). For example, in bacterial cloning, inductive promoters such as the bacteriophage lambda pL, plac, ptrp, and ptac (ptrp-lac hybrid promoter) can be used.
[0051] As used herein, "cssDNA" refers to a circular single-stranded deoxyribonucleic acid. The cssDNA described herein can be packaged in phages or phage particles.
[0052] A "functionally inactivated gene / protein / polypeptide" refers to a gene / protein that does not exhibit its native biological function or activity. Therefore, for example, a functionally inactivated gene is one that cannot be transcribed and / or translated into the protein or polypeptide that would otherwise exhibit the biological function or activity of the protein or polypeptide it encodes. Gene inactivation can be achieved in many ways, including but not limited to deletions, substitutions, insertions, mutations, and the introduction of stop codons, frameshifts, and shortenings. A functionally inactivated protein can also be described as having reduced activity or no activity at all.
[0053] When used herein, "genome editing sequence" refers to a specific type of designed sequence. Genome editing sequences are intended to cause alterations in the host cell genome. In some cases, a genome editing sequence may include one or more regions homologous to the host cell genome. In some cases, a genome editing sequence may include an entire gene expression cassette, e.g., a promoter, open reading frame, and terminator. In other embodiments, a genome editing sequence may include a gene, or a portion of a gene, e.g., an intron, an exon. Genome editing sequences in phagemids can be designed to alter the host cell genome so that the performance of endogenous nucleic acid sequences in the host cell is altered. In some cases, a genome editing sequence may be designed to delete, downregulate, or reduce the activity of endogenous genomic products. In yet another case, a genome editing sequence may be designed to upregulate or increase the expression of endogenous genomic products. In yet another case, a genome editing sequence may be designed to repair undesirable sequences found in the host intracellular genomic genome, e.g., immature stop codons. Examples of genome editing sequences include cssDNA molecules that, when incorporated into the host cell genome, produce siRNA, ribozymes, antisense sequences, RNAi, genes, corrective sequences for various genetic diseases, or beneficial genes or other sequences.
[0054] Where used herein, "introduced" means the transfer of nucleic acid sequences or proteins into a cell, for example, through phage infection (transduction), conjugation, or transformation using molecular biological techniques, such as electroporation or heat shock of chemically competent cells.
[0055] As used herein, "non-endogenous" refers to sequences, proteins, etc., that are not endogenous to a particular host cell (e.g., not naturally present in their relevant environment). For example, in some embodiments, a non-endogenous sequence is introduced into the production strain genome, and that non-endogenous sequence is not naturally found in the production strain genome before the genetic engineering.
[0056] "Nucleic acid," "polynucleotide," and "oligonucleotide" are used interchangeably and refer to deoxyribonucleotide or ribonucleotide polymers in linear or cyclic conformations, in single-stranded or double-stranded forms. For the purposes of this disclosure, these terms should not be construed as limitations relating to the length of the polymer. These terms may encompass known analogues of natural nucleotides, as well as nucleotides with modified base, sugar, and / or phosphate moieties (e.g., phosphorothioate backbone, locked nucleic acids). In general, unless otherwise specified, analogues of a particular nucleotide have the same base-pairing specificity; i.e., analogues of adenine base-pair with thymine. Where double-stranded DNA is described, the DNA may be described as A-DNA, B-DNA, or Z-DNA, according to the conformation that the helical DNA takes. B-DNA, described by James Watson and Francis Crick, is considered dominant in cells and elongates approximately 34 Å per 10 bp sequence; A-DNA elongates approximately 23 Å per 10 bp sequence; and Z-DNA elongates approximately 38 Å per 10 bp sequence.
[0057] In some cases, nucleotide sequences are provided using the letter representations recommended by the International Union of Pure and Applied Chemistry (IUPAC) or a subset thereof. The IUPAC nucleotide codes used herein include: A=adenine, C=cytosine, G=guanine, T=thymine, U=uracil, R=A or G, Y=C or T, S=G or C, W=A or T, K=G or T, M=A or C, B=C or G or T, D=A or G or T, H=A or C or T, V=A or C or G, N=any base, "." or "-" = gap. In some embodiments, the set of letters is (A, C, G, T, U) for adenosine, cytidine, guanosine, thymidine, and uridine, respectively.
[0058] A nucleotide refers to a molecule containing a base moiety, a sugar moiety, and a phosphate moiety. Nucleotides can be linked together via their phosphate and sugar moieties to create nucleoside linkages. The base moiety of a nucleotide may be adenine-9-yl (A), cytosine-1-yl (C), guanine-9-yl (G), uracil-1-yl (U), and thymine-1-yl (T). The sugar moiety of a nucleotide is ribose or deoxyribose. The phosphate moiety of a nucleotide is a pentavalent phosphate. Non-limiting examples of nucleotides are 3'-AMP (3'-adenosine monophosphate) or 5'-GMP (5'-guanosine monophosphate). There are many types of molecules of these kinds that are available in the art and available herein.
[0059] "Oligonucleotide" or "polynucleotide" refers to a synthetic or isolated nucleic acid polymer containing multiple nucleotide subunits.
[0060] When used herein, "phagemide" includes a bacteriophage replication origin, a selection sequence, a designed sequence, and a packaging signal. In some embodiments, the phagemide may also include a plasmid replication origin and / or an antibiotic resistance marker. The phagemide sequence does not contain a phage protein coding sequence, and as a result, the cell will not produce further phage particles upon infection with the phagemide unless it also contains the coding sequence for the phage protein required by the cell.
[0061] As used herein, "phage particle" refers to a cssDNA sequence that is coated with a phage protein but does not contain the coding sequence necessary to produce the phage protein.
[0062] As used herein, “phage protein” refers to a protein encoded by a native bacteriophage that is required for phage replication in a host. See, for example, Figures 1A and 1B, which show the M13 phage protein and the M13 genome structure, respectively. In some embodiments, the phage protein includes a phage protein having nucleic acid sequence modifications compared to the native sequence. Such modifications include, for example, altering the nucleic acid sequence encoding the phage protein to optimize it based on the codon usage of a particular production host. In some embodiments, such modifications result in changes to the activity of the phage protein. It is also possible to modify the nucleic acid sequence encoding the phage protein to produce a phage protein having alternative amino acids compared to the native phage protein. In some cases, the phage protein may contain at least 1, 2, 3, 4, 5, 10, 15, 20, 25, or 30 amino acids that differ from the native phage sequence. Further modifications that may be included in the phage protein include the addition of tags, such as fluorescent markers, to amino acids. The phage proteins described herein, such as native or modified phage proteins, continue to function in production strains. For example, a phage protein containing a tag may have looser binding to other phage proteins if the tag is present, but is still considered a phage protein. The phage proteins specifically described herein include the M13 phage protein and nucleic acid sequence provided in the sequence listing. These proteins are designated p1, p2, p3, p4, p5, p6, p7, p8, p9, p10, and p11, and the genes are designated using their respective Roman numerals (I, II, III, IV, V, VI, VII, VIII, IX, X, and XI). If a phage protein is modified, the modified phage protein can be characterized by comparison with the native phage protein.
[0063] The terms "polypeptide," "peptide," and "protein" are used interchangeably to refer to polymers of amino acid residues. These terms also apply to amino acid polymers in which one or more amino acids are chemical analogues or modified derivatives of the corresponding naturally occurring amino acids.
[0064] Where used herein, “production strain” refers to a bacterial cell used for the production of ssDNA (e.g., cssDNA). In some embodiments, the production strain includes at least two phage proteins incorporated into the genome of the production strain. In some embodiments, the at least two phage proteins may be expressed differentially compared to the expression levels that would result if endogenous regulatory sequences derived from native bacteriophages were used. For example, one or more of the at least two phage protein sequences may be under the control of an inducible promoter. In some embodiments, the production strain is an E. coli strain further produced to reduce the presence of the 3-deoxy-d-manno-octa-2-urosonic acid (Kdo) component in the cell capsule, e.g., a strain of non-toxic E. coli described in U.S. Patent No. 8,303,964 (incorporated herein by reference).
[0065] A “selectable sequence” or “selectable sequence,” as used herein, is a sequence of nucleic acids or amino acids that, when included in a phagemide sequence, enables the retention of a phagemide within a production strain population. Examples of selectable sequences include antibiotic resistance sequences, generated tRNA sequences, nutritional requirement markers, antitoxin genes, or sequences that alter transcription or translation. Examples of selectable sequences include, but are not limited to, genes encoding proteins that increase or decrease resistance or sensitivity to antibiotics (e.g., ampicillin resistance genes, kanamycin resistance genes, neomycin resistance genes, tetracycline resistance genes, chloramphenicol resistance genes, and spectinomycin resistance genes) or other compounds. Further examples of selectable sequences include, but are not limited to, genes encoding proteins that enable cell growth in media lacking otherwise essential nutrients, such as uracil, leucine, adenine, histidine, arginine, lysine, tryptophan, methionine, or other essential metabolites. In some cases, production hosts are induced to have reduced or no production of essential nutrients.
[0066] "Template cssDNA," as used herein, refers to circular single-stranded DNA used in the production of cssDNA. In some embodiments, the template cssDNA comprises a packaging sequence, a designed sequence, and a phage ori, which, when introduced into a bacterial cell, e.g., E. coli cells, is replicated into a replicated form using bacterial cell proteins and nucleic acid sequences, and subsequently packaged into phage particles.
[0067] As used herein, the term “variant” refers to an entity that exhibits significant structural identity with a reference entity but is structurally different from the reference entity in terms of the presence or level of one or more chemical parts. In many embodiments, a variant is also functionally different from its reference entity. Generally, whether a particular entity is appropriately considered a “variant” of a reference entity depends on the degree of its structural identity with the reference entity. As will be understood by those skilled in the art, every biological or chemical reference entity has certain characteristic structural elements. A variant is, by definition, a distinct chemical entity that shares one or more such characteristic structural elements. To give some examples, a small molecule may have a characteristic core structural element (e.g., a macrocyclic core) and / or one or more characteristic pendant portions, and as a result, a variant of this small molecule may share the core structural element and characteristic pendant portions but differ in other pendant portions and / or the type of bonds present in the core (e.g., single bond vs. double bond, E vs. Z); a polypeptide may have a characteristic sequence element consisting of multiple amino acids that have designated positions relative to each other in linear or three-dimensional space and / or contribute to a specific biological function; and a nucleic acid may have a characteristic sequence element consisting of multiple nucleotide residues that have designated positions relative to each other in linear or three-dimensional space. For example, a variant polypeptide may differ from a reference polypeptide as a result of one or more differences in the amino acid sequence and / or one or more differences in the chemical moieties covalently bonded to the polypeptide backbone (e.g., carbohydrates, lipids, etc.). In some embodiments, the variant polypeptide exhibits overall sequence identity with the reference polypeptide, being at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 99%. Alternatively or additionally, in some embodiments, the variant polypeptide does not share at least one characteristic sequence element with the reference polypeptide. In some embodiments, the reference polypeptide has one or more biological activities.In some embodiments, the variant polypeptide shares one or more of the biological activities of the reference polypeptide. In some embodiments, the variant polypeptide lacks one or more of the biological activities of the reference polypeptide. In some embodiments, the variant polypeptide exhibits a reduced level of one or more biological activities compared to the reference polypeptide. In many embodiments, the polypeptide of interest is considered a “variant” of the parent or reference polypeptide if it has an amino acid sequence that is identical to the parent’s amino acid sequence except for a few sequence changes at specific positions. Typically, compared to the parent, less than 20%, less than 15%, less than 10%, less than 9%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, or less than 2% of the residues in the variant are substituted. In some embodiments, the variant has 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 substituted residue compared to the parent. In many cases, variants have a very small number of substituted functional residues (i.e., residues involved in specific biological activity) (e.g., fewer than 5, 4, 3, 2, or 1). Furthermore, variants typically have 5 or fewer, 4 or fewer, 3 or fewer, 2 or fewer, or 1 or fewer additions or deletions compared to the parent, and often have neither additions nor deletions. Furthermore, any additions or deletions are typically less than approximately 25 residues, less than approximately 20 residues, less than approximately 19 residues, less than approximately 18 residues, less than approximately 17 residues, less than approximately 16 residues, less than approximately 15 residues, less than approximately 14 residues, less than approximately 13 residues, less than approximately 10 residues, less than approximately 9 residues, less than approximately 8 residues, less than approximately 7 residues, less than approximately 6 residues, and generally less than approximately 5 residues, less than approximately 4 residues, less than approximately 3 residues, or less than approximately 2 residues. In some embodiments, the parent polypeptide or reference polypeptide is one found in nature. As those skilled in the art will understand, multiple variants of a particular polypeptide of interest are commonly found in nature.
[0068] cssDNA production This disclosure provides, in particular, a method for constructing cssDNA. Those skilled in the art will understand, by reading this disclosure, that the designed sequence contained in the phagemide may be of any length necessary to achieve the desired function of the designed sequence. For example, if the desired sequence is a three-dimensional DNA structure (i.e., DNA origami), the length is selected to fit the size of the desired final product. Similarly, if the designed sequence is for application in cell therapy or gene therapy, for example, for final use in CAR-T cell therapy, the sequence may be the sum of the length of the gene targeted for insertion and the length of the homology arm necessary to target the gene to a specific locus. Specific uses of cssDNA include use in combination with genome editing techniques such as homologous recombination and CRISPR / Cas-related techniques. Considering the various uses of cssDNA, the designed sequences are expected to be in the ranges of approximately 10 bp to 10,000 bp, approximately 100 bp to 10,000 bp, approximately 200 bp to 9,000 bp, approximately 300 bp to 9,000 bp, approximately 500 bp to 9,000 bp, approximately 500 bp to 8,000 bp, approximately 1,000 bp to 10,000 bp, approximately 5,000 bp to 15,000 bp, and approximately 10,000 bp to 30,000 bp. The cssDNA produced by the methods described herein is consistent in some embodiments, for example, more than 95%, more than 96%, more than 97%, more than 98%, more than 99%, more than 99.5%, or more than 99.9% of the produced molecules are mutation-free compared to the template. In some embodiments, the molecule may also be longer than 10,000 bp, 20,000 bp, or 30,000 bp.
[0069] In some embodiments, cssDNA production occurs when a dsDNA phagemid, which encodes a phage replication origin and packaging signal but lacks a phage gene, is replicated as cssDNA by a trans-expressed phage protein in the host strain and subsequently packaged into a phage particle by the phage protein. When produced as described above, the phage particle does not contain any phage protein coding sequence and therefore cannot replicate itself even if it infects another bacterial cell. This differs from wild-type phages, which encode the genes necessary for their own replication in their genome and can replicate themselves, infect other bacterial cells, and generate new self-replicating phage particles that create even more self-replicating phages. The phagemids produced herein may contain a phage replication origin and a plasmid replication origin. However, in some cases, the phagemid contains only a phage replication origin and lacks a plasmid replication origin.
[0070] In some embodiments, the phagemids described herein include one or more selectable sequences. The selectable sequences function to link the phagemid to the production strain such that the production strain will not reproduce without the phagemid encoding the selectable sequence. Examples of selectable sequences include sequences encoding antibiotic resistance proteins, sequences encoding nutrient-dependent amino acids; tRNAs that compensate for defects introduced into the production strain; RNA sequences that alter transcription or translation; or antitoxin amino acids (if the production strain produces toxins). Those skilled in the art will understand that there are many specific examples of such amino acid and nucleic acid sequences that can be used to create a relationship between the phagemid and the production strain such that the production strain cannot continue to function fully without the phagemid. In preferred embodiments, sequences encoding genes or proteins that could have adverse effects on the final product in which the cssDNA is used are avoided. For example, antibiotic resistance sequences and toxin expression sequences are undesirable in cssDNA when it is used for gene therapy or for the production of microorganisms released into the environment.
[0071] Those skilled in the art will understand that several methods exist for producing phagemids that do not contain plasmid dsDNA replication origins. This can be useful when it is undesirable for the plasmid replication origin to be present in the final cssDNA produced. For example, a desired phagemid sequence can be cloned into a plasmid having further regulatory sequences, such as a T7 promoter sequence, a designed sequence, a selection sequence, and a packaging signal sequence. An in vitro transcription reaction can be carried out, and the resulting RNA can then be reverse transcribed to the desired ssDNA sequence. The resulting ssDNA can be circularized to cssDNA using any method known in the art, for example, by annealing a sprint (a short DNA with complementary regions at both ends of the ssDNA) and subjecting the reaction to a ligation reaction to anneal the ends and form cssDNA. Iyer et al., Efficient Homology-directed Repair with Circular ssDNA Donors, CRISPR J, Oct;5(5):685-701, 2022. Additionally, linear ssDNA can be converted to cssDNA using ssDNA ligase.
[0072] Alternatively, desired components of ssDNA can be produced using a combination of synthetic DNA synthesis and standard cloning and PCR techniques. The sequence can be generated in a double-stranded form, and then the DNA strands can be selectively removed using lambda exonuclease.
[0073] Further techniques include the use of streptavidin-coated beads and biotin-conjugated PCR primers, where one primer is biotinylated and the other is not. After PCR is complete, one strand remains biotinylated while the other is not. Exonuclease is applied to degrade the non-biotinylated strand. The resulting biotinylated ssDNA product is bound to streptavidin-coated beads, the beads are physically separated from the solution, and then the ssDNA is recovered by eluting the DNA from the beads by resuspending them in elution buffer. Avci-Adali M, Paul A, Wilhelm N, Ziemer G, Wendel HP. Upgrading SELEX technology by using lambda exonuclease digestion for single-stranded DNA generation. Molecules. 2009 Dec 24;15(1):1-11.
[0074] The ssDNA, which is ultimately packaged into phage particles, can also be cloned into the host cell genome as part of an expression cassette that enables the direct production of RNA transcripts containing the desired phagemide components. Transcription can be induced at any desired time, and once transcribed, the resulting transcript can be reverse transcribed to allow the packaging signal to reach the rest of the phage genes required for packaging. The phage genes encoding the coat proteins may be integrated in the genome or supplied via a helper virus or helper plasmid. Those skilled in the art will understand that it may be necessary to further express the desired reverse transcriptase.
[0075] Phagemids can also be produced without a plasmid origin using any standard in vitro dsDNA cloning method, including restriction / ligation, isothermal assembly, Golden Gate assembly, or any other method used for assembling DNA fragments. Once produced in vitro, the circular dsDNA can then be transformed into any bacterial strain that expresses the phage mechanisms necessary to replicate and package cssDNA from a dsDNA template containing a phage origin and phage packaging signals.
[0076] In cases where the phagemide contains both a phage origin and a plasmid origin, standard plasmid cloning techniques can be used to generate the desired template sequence for producing the cssDNA product. While including the plasmid origin is convenient for many genetic engineering techniques, it also increases the likelihood that the plasmid origin sequence will be present in the cssDNA product, which may be undesirable in some cases.
[0077] In some embodiments, the phagemide is a high copy number phagemide. In some embodiments, the phagemide copy number in production cell lines grown to late logarithmic phase is at least 1,000, and optionally, the phagemide copy number in production cell lines grown to late logarithmic phase is at least 2,000, at least 4,000, at least 7,000, at least 8,000, or at least 15,000. As shown in Example 6, in some embodiments, a phagemide containing one or more mutations that increase copy number increases cssDNA yield. In some embodiments, the phagemide contains the inc1 mutation. In some embodiments, the phagemide contains the inc2 mutation. In some embodiments, the phagemide contains both the inc1 and inc2 mutations. In some embodiments, the phagemide contains pST19, pDHA29, pDHA30, pDHK29, pDHK30, or a runaway R1 origin of replication or its derivatives.
[0078] In some embodiments, the phagemide includes a pUC origin (ori). In some embodiments, the phagemide ori includes a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 27. In some embodiments, the phagemide includes an inc1 mutant ori derived from the pUC ori. In some embodiments, the inc1 mutant ori includes a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 28. In some embodiments, the phagemide includes an inc1 mutant ori derived from the pUC ori. In some embodiments, the inc1 mutant ori includes a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 31. In some embodiments, the inc1 mutant is a C59T mutant relative to a standard pUC origin, and the nucleotide numbering is based on the sequence of the standard pUC origin shown in SEQ ID NO: 27. In some embodiments, the phagemide includes an inc2 mutant ori derived from the pUC ori. In some embodiments, the inc2 mutant is a C92T mutant relative to the standard pUC origin, and the nucleotide numbering is based on the sequence of the standard pUC origin shown in SEQ ID NO: 27. In some embodiments, the inc2 mutant ori includes a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 29. In some embodiments, the inc2 mutant ori includes a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 32. In some embodiments, the phagemide includes inc1 and inc2 mutant ori derived from the pUC ori. In some embodiments, the inc1&2 ori (also referred to as inc1 / inc2 or inc1 / 2) includes C59T and C92T mutants relative to the standard pUC origin, and the nucleotide numbering is based on the sequence of the standard pUC origin shown in SEQ ID NO: 27.In some embodiments, the inc1 and inc2 mutant ori (inc1&2 ori) contain sequences having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to SEQ ID NO: 30. In some embodiments, the inc1 and inc2 mutant ori (inc1&2 ori) contain sequences having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to SEQ ID NO: 33. In some embodiments, the phagemide contains the inc3 mutant. In some embodiments, the phagemide contains the inc5 mutant. In some embodiments, the phagemide copy number in production cells grown to the late logarithmic phase is at least 1,000. In some embodiments, the phagemide copy number in production cells grown to the late logarithmic phase is at least 2,000. In some embodiments, the phagemide copy number in production cells grown to the late logarithmic phase is at least 4,000. In some embodiments, culturing a cssDNA production strain or production system containing phagemids with inc1&2 ori yields at least 10 times higher yields compared to culturing a cssDNA production strain or production system containing phagemids with wild-type pUC ori. In some embodiments, once the phagemids are produced, they can be introduced into the production strain, for example, by any method known in the art, including electroporation, or by using chemically competent cells, such as those prepared with divalent ions such as calcium chloride, magnesium chloride, or rubidium chloride solution.
[0079] Production of production stocks This disclosure exemplifies production strains and methods for producing them. The production strains described are useful, among other things, for producing cssDNA. Those skilled in the art will be familiar with various techniques that may be useful for generating such production strains. Production strains described herein include bacterial strains that can produce cssDNA phage particles after transformation with phagemids or nucleic acid sequences including phage packaging signals, designed sequences, and phage replication origins. Bacteria may naturally contain intracellular mechanisms necessary for producing cssDNA from template dsDNA, ssDNA, or cssDNA, or bacteria may be produced to contain such mechanisms. Mechanisms, as used herein, refer to proteins necessary for cssDNA creation, rolling circle replication (sometimes referred to herein as ssDNA production), or ssDNA packaging (sometimes referred to herein as phage particle packaging). Production strains as described herein can be natural targets for infection by phages containing template cssDNA, or they can be easily produced to have the correct support protein production for phage particle production upon introduction of template cssDNA or dsDNA using standard transformation techniques. Exemplary bacteria that can be used to produce production strains include E. coli, as well as other Gram-negative bacterial species capable of replicating cssDNA phage genomes of cssDNA phages in the Monodnaviria region, particularly including phages Ff, Fd, F1, and M13.
[0080] In some embodiments, the production strain includes a copy of a native gene derived from a bacteriophage, such as the M13 bacteriophage, which is integrated into the genome at one or more loci. For example, a gene encoding a phage protein can be amplified by PCR from the M13KO7 helper phage genome and integrated into the genome of a production strain as described herein. These production strains can produce cssDNA when template dsDNA, cssDNA, or ssDNA is introduced, as described herein. A production strain containing at least one or at least two copies of a gene encoding a phage protein may also contain one or more further copies of the phage protein gene. For example, phage genes I, II, III, IV, V, VI, VII, VIII, IX, X, and XI can be incorporated into the genome using their endogenous regulatory sequences as described herein, and any further copy of any one of genes I, II, III, IV, V, VI, VII, VIII, IX, X, and XI can be incorporated into the genome or expressed from a helper plasmid using either the gene's native regulatory sequence or a heterologous regulatory sequence, such as an inducible promoter.
[0081] In some embodiments, a circular single-stranded DNA (cssDNA) production system comprising a production strain and a phagemide is provided herein, wherein the production system comprises a packaging signal, a designed sequence, and at least one selectable sequence.
[0082] The phagemide includes a pUC replication origin or its derivatives, and the phagemide includes the inc1 mutation and / or the inc2 mutation. In some embodiments, the phagemide includes the inc1 mutation and the inc2 mutation. In some embodiments, the production strain includes two or more phage genes selected from the group consisting of genes I, II, III, IV, V, VI, VII, VIII, IX, X, and XI, which are integrated in the genome of the production strain cells. In some embodiments, the phage genes are expressed from a helper plasmid (in addition to or instead of the integrated phage genes). In some embodiments, the production strain includes a helper plasmid derived from the helper phage M13KO7 by removing the F1 origin and packaging signal. In some embodiments, the helper plasmid copy number in the production strain cells grown to late logarithmic phase is at least 1,000. In some embodiments, the helper plasmid copy number in the production cell line grown to the late logarithmic phase is at least 2,000, at least 4,000, at least 7,000, at least 8,000, or at least 15,000. In some embodiments, the helper plasmid includes a pUC replication origin.
[0083] In some embodiments, the helper plasmid contains the inc1 mutation. In some embodiments, the helper plasmid contains the inc2 mutation. In some embodiments, the helper plasmid contains both the inc1 and inc2 mutations. In some embodiments, the helper plasmid contains pST19, pDHA29, pDHA30, pDHK29, pDHK30, or a runaway R1 origin of replication or its derivatives.
[0084] In some embodiments, the helper plasmid includes a pUC origin (ori). In some embodiments, the helper plasmid ori includes a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 27. In some embodiments, the helper plasmid includes an inc1 mutant ori derived from the pUC ori. In some embodiments, the inc1 ori includes a C59T mutation to the standard pUC origin, and the nucleotide numbering is based on the sequence of the standard pUC origin shown in SEQ ID NO: 27. In some embodiments, the inc1 mutant ori includes a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 28. In some embodiments, the inc1 mutant ori includes a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 31. In some embodiments, the helper plasmid includes an inc2 mutant ori derived from the pUC ori. In some embodiments, the inc2 ori includes a C92T mutation relative to a standard pUC origin, and the nucleotide numbering is based on the sequence of the standard pUC origin shown in SEQ ID NO: 27. In some embodiments, the inc2 mutant ori includes a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 29. In some embodiments, the inc2 mutant ori includes a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 32. In some embodiments, the helper plasmid includes inc1 and inc2 mutant ori derived from the pUC ori. In some embodiments, the inc1&2 ori (also referred to as inc1 / inc2 or inc1 / 2) includes C59T and C92T mutations relative to a standard pUC origin, and the nucleotide numbering is based on the sequence of the standard pUC origin shown in SEQ ID NO: 27.In some embodiments, the inc1 and inc2 mutant ori (inc1&2 ori) contain sequences having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to SEQ ID NO: 30. In some embodiments, the inc1 and inc2 mutant ori (inc1&2 ori) contain sequences having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with respect to SEQ ID NO: 33. In some embodiments, the helper plasmid contains the inc3 mutant. In some embodiments, the helper plasmid contains the inc5 mutant. In some embodiments, the helper plasmid copy number in production cells grown to the late logarithmic phase is at least 1,000. In some embodiments, the helper plasmid copy number in production cells grown to the late logarithmic phase is at least 2,000. In some embodiments, the helper plasmid copy number in production cells grown to the late logarithmic phase is at least 4,000.
[0085] In some embodiments, the production strains described herein may also contain a shortened version of gene III that reduces the production strain's resistance to subsequent infection by phage particles. Shortenings may include deletions of amino acid residues 21–273, 93–121, 93–141, or insertions of stop codons at codons 20, 25, 28, and 30 within the p3 gene. See U.S. Patent No. 8,227,242, which is incorporated herein by reference in its entirety.
[0086] In some production strains, the expression of the phage protein p5 is altered (e.g., decreased) to increase cssDNA production. Decreasing the activity of p5 protein production can be achieved using any method known in the art, including the methods already described. P5 activity can also be altered through mutations in the untranslated region of the p5 gene, as described in Lee et al., Optimizing protein V untranslated region sequence in M13 phage for increased production of single-stranded DNA for origami. Nucleic Acids Research, Volume 49, Issue 11, 21 June 2021 (see, for example, Figure 6).
[0087] Gene overexpression can be achieved using any method known in the art. Generally, overexpression refers to an increase in the amount of protein produced compared to the amount normally expressed when the protein is expressed in its native genomic environment. For this reason, overexpression is sometimes described in relation to the second protein being expressed. Specific examples of how overexpression can be achieved include inserting multiple copies of a gene or switching a promoter or other regulatory sequence to increase expression. The production strains described herein may contain a copy of the M13 genome containing the M13 regulatory sequence. For the purposes of clarity, the term overexpression is used to refer to the production of M13 protein in amounts greater than the amount expressed when the M13 protein is under the control of its native regulatory sequence. In other words, if two copies of the M13 genome are included in the production strain, all phage proteins are considered to be overexpressed, but if only a single copy of the M13 genome is included with its native regulatory sequence, no phage proteins are considered to be overexpressed.
[0088] The production strains described herein can also be described by the ratio of one protein to another. The production ratio of individual phage proteins relative to each other can increase the efficiency of cssDNA product production. For example, expression of p2:p5 in ratios of 1:1, 2:1, 3:1, or 10:1 during the stationary or logarithmic phase of proliferation may be beneficial for phage particle production.
[0089] Overexpression of p2 and p10 through the use of additional copies of native genes, coding sequences placed under the control of a constitutive promoter, coding sequences placed under the control of an inductive promoter, or some combination thereof, is thought to increase phage particle production. See Behler, et al., Phage-free production of artificial ssDNA with Escherichia coli, Biotechnol Bioeng. 2022 Oct;119(10):2878-2889.
[0090] The p8 protein constitutes a large portion of the phage coat, based on its overall amino acid contribution. Indeed, the wild-type M13 phage capsid contains 2700 copies of the p8 protein. In synthesized phage capsids containing a larger genome than the wild-type 7.2kb genome, the capsid incorporates even more p8 protein, as the length of the capsid is proportional to the size of the packaged DNA. Therefore, overexpression of the p8 protein is particularly useful when the designed sequence is longer. Overexpression of p8 through the use of additional copies of the native gene, coding sequences placed under the control of a constitutive promoter, coding sequences placed under the control of an inductive promoter, or some combination thereof, is thought to increase phage particle production. In certain embodiments, the production strain contains an overexpressed p8 protein and a designed sequence that is at least 2kb in length. In yet another embodiment, the production strain includes an overexpressed p8 protein and a designed sequence that is at least 3kb, 4kb, 5kb, 6kb, 7kb, 8kb, 9kb, 10kb, 15kb, 20kb, 25kb, or 30kb in length.
[0091] In some embodiments, the p8 protein is overexpressed from a standard T7 promoter or a variant thereof. Overexpression of genes II, X, IV, I, XI, and especially VIII from a standard T7 promoter increases the production of phage particles and cssDNA. Decreasing the expression of gene V also increases its production.
[0092] Synthetic promoters can be characterized by the amount of RNA transcription they can produce. Promoters that produce a large amount of RNA transcript are characterized as high, promoters that produce somewhat less can be described as medium, and promoters that produce relatively less can be described as low. Examples of T7 promoters that fall into these categories include: high [ka] Medium [ka] low [ka]
[0093] The production strain may contain at least two coding sequences that produce phage proteins that promote cssDNA production upon introduction of a phagemid. The at least two phage proteins incorporated into the genome of the production strain are selected to optimize the yield and / or fidelity of the resulting cssDNA.
[0094] In some embodiments, the production strain is p1 phage protein [ka] The genome includes a coding sequence integrated into the genome. In some embodiments, the p1 coding sequence is under the control of an inducible promoter (e.g., a temperature-sensitive promoter or a chemosensitive promoter). In some embodiments, the second of at least two coding sequences is selected from the following phage proteins: p2, p3, p4, p5, p6, p7, p8, p9, p10, and p11. In further embodiments, at least two coding sequences are integrated into the genome at distinct loci. This reduces the likelihood of recombinant and replication-competent phage formation.
[0095] In further embodiments, the p1 coding sequence and a second of at least two coding sequences selected from the phage proteins p2, p3, p4, p5, p6, p7, p8, p9, p10, and p11 are placed under the control of the same promoter so that they are expressed substantially simultaneously at production. In some embodiments, p1 and a second of at least two coding sequences selected from the phage proteins p2, p3, p4, p5, p6, p7, p8, p9, p10, and p11 are integrated in the genome at the same locus, while a third phage protein coding sequence is integrated in the genome at a separate locus. In some embodiments, a helper plasmid or helper virus is used to deliver any additional phage proteins required for cssDNA production.
[0096] In some embodiments, the production strain is a p2 phage protein [ka] The genome includes a coding sequence integrated into the genome. In some embodiments, the p2 coding sequence is under the control of an inducible promoter (e.g., a temperature-sensitive promoter or a chemosensitive promoter). In some embodiments, the second of at least two coding sequences is selected from the following phage proteins: p1, p3, p4, p5, p6, p7, p8, p9, p10, and p11. In further embodiments, at least two coding sequences are integrated into the genome at distinct loci. This reduces the likelihood of recombinant and replication-competent phage formation.
[0097] In further embodiments, the p2 coding sequence and a second of at least two coding sequences selected from the phage proteins p1, p3, p4, p5, p6, p7, p8, p9, p10, and p11 are placed under the control of the same promoter so that they are expressed substantially simultaneously at production. In some embodiments, p2 and a second of at least two coding sequences selected from the phage proteins p1, p3, p4, p5, p6, p7, p8, p9, p10, and p11 are integrated in the genome at the same locus, while a third phage protein coding sequence is integrated in the genome at a separate locus. In some embodiments, a helper plasmid or helper virus is used to deliver any additional phage proteins required for cssDNA production.
[0098] In some embodiments, the production strain is a p3 phage protein [ka] The genome includes a coding sequence integrated into the genome. In some embodiments, the p3 coding sequence is under the control of an inducible promoter (e.g., a temperature-sensitive promoter or a chemosensitive promoter). In some embodiments, the second of at least two coding sequences is selected from the following phage proteins: p1, p2, p4, p5, p6, p7, p8, p9, p10, and p11. In further embodiments, at least two coding sequences are integrated into the genome at distinct loci. This reduces the likelihood of recombinant and replication-competent phage formation.
[0099] In further embodiments, the p3 coding sequence and a second of at least two coding sequences selected from the phage proteins p1, p2, p4, p5, p6, p7, p8, p9, p10, and p11 are placed under the control of the same promoter so that they are expressed substantially simultaneously at production. In some embodiments, p3 and a second of at least two coding sequences selected from the phage proteins p1, p2, p4, p5, p6, p7, p8, p9, p10, and p11 are integrated in the genome at the same locus, while a third phage protein coding sequence is integrated in the genome at a separate locus. In some embodiments, a helper plasmid or helper virus is used to deliver any additional phage proteins required for cssDNA production.
[0100] In some embodiments, the production strain is a p4 phage protein [ka] The genome includes a coding sequence integrated into the genome. In some embodiments, the p4 coding sequence is under the control of an inducible promoter (e.g., a temperature-sensitive promoter or a chemosensitive promoter). In some embodiments, the second of at least two coding sequences is selected from the following phage proteins: p1, p2, p3, p5, p6, p7, p8, p9, p10, and p11. In further embodiments, at least two coding sequences are integrated into the genome at distinct loci. This reduces the likelihood of recombinant and replication-competent phage formation.
[0101] In further embodiments, the p4 coding sequence and a second of at least two coding sequences selected from the phage proteins p1, p2, p3, p5, p6, p7, p8, p9, and p10 are placed under the control of the same promoter so that they are expressed substantially simultaneously at production. In some embodiments, p4 and a second of at least two coding sequences selected from the phage proteins p1, p2, p3, p5, p6, p7, p8, p9, p10, and p11 are integrated in the genome at the same locus, while a third phage protein coding sequence is integrated in the genome at a separate locus. In some embodiments, a helper plasmid or helper virus is used to deliver any additional phage proteins required for cssDNA production.
[0102] In some embodiments, the production strain is p5 phage protein [ka] The genome includes a coding sequence integrated into the genome. In some embodiments, the p5 coding sequence is under the control of an inducible promoter (e.g., a temperature-sensitive promoter or a chemosensitive promoter). In some embodiments, the second of at least two coding sequences is selected from the following phage proteins: p1, p2, p3, p4, p6, p7, p8, p9, p10, and p11. In further embodiments, at least two coding sequences are integrated into the genome at distinct loci. This reduces the likelihood of recombinant and replication-competent phage formation.
[0103] In further embodiments, the p5 coding sequence and a second of at least two coding sequences selected from the phage proteins p1, p2, p3, p4, p6, p7, p8, p9, p10, and p11 are placed under the control of the same promoter so that they are expressed substantially simultaneously at production. In some embodiments, p5 and a second of at least two coding sequences selected from the phage proteins p1, p2, p3, p4, p6, p7, p8, p9, p10, and p11 are integrated in the genome at the same locus, while a third phage protein coding sequence is integrated in the genome at a separate locus. In some embodiments, a helper plasmid or helper virus is used to deliver any additional phage proteins required for cssDNA production.
[0104] In some embodiments, the production strain is a p6 phage protein [ka] The genome includes a coding sequence integrated into the genome. In some embodiments, the p6 coding sequence is under the control of an inducible promoter (e.g., a temperature-sensitive promoter or a chemosensitive promoter). In some embodiments, the second of at least two coding sequences is selected from the following phage proteins: p1, p2, p3, p4, p5, p7, p8, p9, p10, and p11. In further embodiments, at least two coding sequences are integrated into the genome at distinct loci. This reduces the likelihood of recombinant and replication-competent phage formation.
[0105] In further embodiments, the p6 coding sequence and a second of at least two coding sequences selected from the phage proteins p1, p2, p3, p4, p5, p7, p8, p9, and p10 are placed under the control of the same promoter so that they are expressed substantially simultaneously at production. In some embodiments, p6 and a second of at least two coding sequences selected from the phage proteins p1, p2, p3, p4, p5, p7, p8, p9, p10, and p11 are integrated in the genome at the same locus, while a third phage protein coding sequence is integrated in the genome at a separate locus. In some embodiments, a helper plasmid or helper virus is used to deliver any additional phage proteins required for cssDNA production.
[0106] In some embodiments, the production strain includes a coding sequence for the p7 phage protein (SEQ ID NO: 10:MEQVADFDTIYQAMIQISVVLCFALGIIAGGQR) integrated into the genome. In some embodiments, the p7 coding sequence is under the control of an inducible promoter (e.g., a temperature-sensitive promoter or a chemosensitive promoter). In some embodiments, the second of at least two coding sequences is selected from the following phage proteins: p1, p2, p3, p4, p5, p6, p8, p9, and p10. In further embodiments, at least two coding sequences are integrated into the genome at distinct loci. This reduces the likelihood of recombinant and replication-competent phage formation.
[0107] In further embodiments, the p7 coding sequence and a second of at least two coding sequences selected from the phage proteins p1, p2, p3, p4, p5, p6, p8, p9, p10, and p11 are placed under the control of the same promoter so that they are expressed substantially simultaneously at production. In some embodiments, p7 and a second of at least two coding sequences selected from the phage proteins p1, p2, p3, p4, p5, p6, p8, p9, p10, and p11 are integrated in the genome at the same locus, while a third phage protein coding sequence is integrated in the genome at a separate locus. In some embodiments, a helper plasmid or helper virus is used to deliver any additional phage proteins required for cssDNA production.
[0108] In some embodiments, the production strain is p8 phage protein [ka] The genome includes a coding sequence integrated into the genome. In some embodiments, the p8 coding sequence is under the control of an inducible promoter (e.g., a temperature-sensitive promoter or a chemosensitive promoter). In some embodiments, the second of at least two coding sequences is selected from the following phage proteins: p1, p2, p3, p4, p5, p6, p7, p9, p10, and p11. In further embodiments, at least two coding sequences are integrated into the genome at distinct loci. This reduces the likelihood of recombinant and replication-competent phage formation.
[0109] In further embodiments, the p8 coding sequence and a second of at least two coding sequences selected from the phage proteins p1, p2, p3, p4, p5, p6, p7, p9, p10, and p11 are placed under the control of the same promoter so that they are expressed substantially simultaneously at production. In some embodiments, p8 and a second of at least two coding sequences selected from the phage proteins p1, p2, p3, p4, p5, p6, p7, p9, p10, and p11 are integrated in the genome at the same locus, while a third phage protein coding sequence is integrated in the genome at a separate locus. In some embodiments, a helper plasmid or helper virus is used to deliver any additional phage proteins required for cssDNA production.
[0110] In some embodiments, the production strain includes a coding sequence for the p9 phage protein (SEQ ID NO: 12: MSVLVYSFASFVLGWCLRSGITYFTRLMETSS) integrated into the genome. In some embodiments, the p9 coding sequence is under the control of an inducible promoter (e.g., a temperature-sensitive promoter or a chemosensitive promoter). In some embodiments, the second of at least two coding sequences is selected from the following phage proteins: p1, p2, p3, p4, p5, p6, p7, p8, p10, and p11. In further embodiments, at least two coding sequences are integrated into the genome at distinct loci. This reduces the likelihood of recombinant and replication-competent phage formation.
[0111] In further embodiments, the p9 coding sequence and a second of at least two coding sequences selected from the phage proteins p1, p2, p3, p4, p5, p6, p7, p8, p10, and p11 are placed under the control of the same promoter so that they are expressed substantially simultaneously at production. In some embodiments, p9 and a second of at least two coding sequences selected from the phage proteins p1, p2, p3, p4, p5, p6, p7, p8, p10, and p11 are integrated in the genome at the same locus, while a third phage protein coding sequence is integrated in the genome at a separate locus. In some embodiments, a helper plasmid or helper virus is used to deliver any additional phage proteins required for cssDNA production.
[0112] In some embodiments, the production strain is p10 phage protein [ka] The genome includes a coding sequence integrated into the genome. In some embodiments, the p10 coding sequence is under the control of an inducible promoter (e.g., a temperature-sensitive promoter or a chemosensitive promoter). In some embodiments, the second of at least two coding sequences is selected from the following phage proteins: p1, p2, p3, p4, p5, p6, p7, p8, p9, and p11. In further embodiments, at least two coding sequences are integrated into the genome at distinct loci. This reduces the likelihood of recombinant and replication-competent phage formation.
[0113] In further embodiments, the p10 coding sequence and a second of at least two coding sequences selected from the phage proteins p1, p2, p3, p4, p5, p6, p7, p8, p9, and p11 are placed under the control of the same promoter so that they are expressed substantially simultaneously at production. In some embodiments, p10 and a second of at least two coding sequences selected from the phage proteins p1, p2, p3, p4, p5, p6, p7, p8, p9, and p11 are integrated in the genome at the same locus, while a third phage protein coding sequence is integrated in the genome at a separate locus. In some embodiments, a helper plasmid or helper virus is used to deliver any additional phage proteins required for cssDNA production.
[0114] In some embodiments, the production strain is the p11 phage protein [ka] The genome includes a coding sequence integrated into the genome. In some embodiments, the p11 coding sequence is under the control of an inducible promoter (e.g., a temperature-sensitive promoter or a chemosensitive promoter). In some embodiments, the second of at least two coding sequences is selected from the following phage proteins: p1, p2, p3, p4, p5, p6, p7, p8, p9, and p10. In further embodiments, at least two coding sequences are integrated into the genome at distinct loci. This reduces the likelihood of recombinant and replication-competent phage formation.
[0115] In further embodiments, the p11 coding sequence and a second of at least two coding sequences selected from the phage proteins p1, p2, p3, p4, p5, p6, p7, p8, p9, and p10 are placed under the control of the same promoter so that they are expressed substantially simultaneously at production. In some embodiments, p11 and a second of at least two coding sequences selected from the phage proteins p1, p2, p3, p4, p5, p6, p7, p8, p9, and p10 are integrated in the genome at the same locus, while a third phage protein coding sequence is integrated in the genome at a separate locus. In some embodiments, a helper plasmid or helper virus is used to deliver any additional phage proteins required for cssDNA production. [Examples]
[0116] (Example 1) Incorporation of the entire M13 genome into production strains The M13 genome is composed of two distinct transcription units. The first consists of genes II, X, V, VII, IX, and VIII, which are expressed by strong promoters. The second consists of genes III, VI, I, XI, and IV, which are expressed by weak promoters. Two of these genes overlap with the open reading frames of two other genes (for example, the gene for p10 is entirely contained within the coding region of p2). PA, PB, and PH are strongly active promoters (see Figure 1B). Two additional weak promoters are denoted as PZ and PW. Two terminators (T) are strongly rho-independent terminators, and a third, weaker terminator (T(weak)) is rho-dependent.
[0117] Smeal et al., Simulation of the M13 life cycle I: Assembly of a genetically-structured deterministic chemical kinetic simulation, Virology, Volume 500, 2017, Pages 259-274.
[0118] To create a stable cell line capable of producing M13 phage particles while being highly unlikely to recombine with phagemids to generate competent phages, each transcription unit of the M13 bacteriophage, which contains protein-coding genes and lacks phage replication origins and phage packaging signals, was cloned from the publicly available helper phage M13KO7 (www.neb.com / products / n0315-M13KO7-helper-phage) (SEQ ID NO: 24) and incorporated into two loci of the E. coli MG1655 genome.
[0119] The first transcription unit is cloned into a circular double-stranded plasmid along with homology arms for integration into the E. coli MG1655 genome at the pyrF locus, a spectinomycin selection marker to enable integrant selection, and a p15A origin for replication in E. coli. The selection marker is flanked by loxP sites to allow the marker to be recycled during cre gene expression. The insert containing the homology arms, the M13 protein-coding gene, and the selection marker is flanked by restriction enzyme sites, thus allowing the circular double-stranded DNA to be linearized before transformation of E. coli. The plasmid is digested with restriction enzymes and subsequently electrophoresed on a 1% agarose tris-acetate-ethylenediaminetetraacetic acid (TAE) gel to separate the desired integration cassette from the unwanted plasmid backbone containing the p15A E. coli replication origin. The bands containing the embedded cassette are purified using the Qiagen Gel Extraction Kit (catalog number 28706, Qiagen, Hilden, Germany) according to the manufacturer's recommended protocol. Since the p15A E. coli replication origin required for plasmid replication in E. coli is selected and eliminated using agarose gel electrophoresis, it can be confident that all E. coli cells that survive antibiotic selection after transformation with the embedded cassette have likely undergone a genomic integration event rather than transformation with the replication plasmid.
[0120] Creation of a reconvening competent host The E. coli host strain MG1655 was initially transformed using a publicly available reconvening plasmid expressing a lambda red reconvening system containing beta, exo, and gamma from an arabinose-inducible promoter (catalog number CAS9BAC1P, Sigma Aldrich, Burlington, Massachusetts) (see Figure 5).
[0121] The plasmid also expresses the kanamycin antibiotic marker, which allows for plasmid selection and maintenance in a population of bacterial cells, as well as the temperature-sensitive pSC101 replication factor repA101ts, which allows for plasmid elimination when expression of the lambda red repair system is no longer required. Transformants are selected by plating onto Lysogeny Broth (LB) agar (Bertani, G. 1952. "Studies on Lysogenesis. I. The mode of phage liberation by lysogenic Escherichia coli." J. Bacteriology, 62:293-300) supplemented with 50 ug / ml kanamycin. Colonies are collected and inoculated into 15 ml of LB liquid medium supplemented with 50 ug / ml kanamycin to maintain plasmid retention. The culture is incubated overnight at 30°C. A 1 ml overnight culture is mixed with 1 ml of 50% glycerol / 50% aqueous solution and stored at -80°C to create a permanent bank of a new strain, referred to as the MG1655 reconciling strain. Plasmid DNA is isolated from the remaining 14 ml of culture and sequenced via Oxford Nanopore technology to confirm the presence and sequence identity of the reconciling plasmid.
[0122] Induction of the expression of reconvening genes Once transformation of E. coli MG1655 strain by the reconvening plasmid is confirmed, the reconvening strain is streaked onto LB agar supplemented with 50 ug / ml kanamycin and incubated overnight at 30°C. Subsequently, a single colony is inoculated into LB Lennox (low-salt formulation) supplemented with 50 ug / ml kanamycin and incubated overnight at 30°C. Next, this culture is diluted to OD 0.01 in LB liquid medium supplemented with 50 ug / ml kanamycin. The culture is incubated at 30°C until it reaches OD 0.30. At this point, it is diluted 1:2 with preheated LB broth supplemented with 4% L-arabinose to create a final concentration of 2% L-arabinose to induce expression of the reconvening genes beta, exo, and gamma. Subsequently, the culture is incubated for a further 1 hour at 30°C to accumulate sufficient amounts of reconvening proteins beta, exo, and gamma in the cells.
[0123] Electroporation using a phage genome embedded cassette After 1 hour of reconciliation gene expression, cells from 1 ml of culture are harvested by centrifugation at 7000 RCF for 1 minute, followed by washing three times with ice-cold deionized water to remove salt. The cells are then resuspended in 49 μl of ice-cold deionized water, as well as in 1 μl of gel-purified linearized integration cassettes consisting of two pyrF homology arms, all protein coding sequences derived from the first transcription unit of M13KO7, and a selection marker flanked by the loxP site.
[0124] The cell-DNA mixture is electroporated using a Bio-Rad Gene Pulser with a 1 mm cuvette, 2.1 kV, 100 ohm resistance, and 25 microfarad capacitance. Immediately after applying the electrical pulse, the cells are resuspended in 1 ml of LB Lennox medium and incubated overnight at 30°C. After this overnight recovery period, the cells are plated onto LB agar supplemented with 50 ug / ml spectinomycin to select for integration events.
[0125] Identification of cells with correct embedded events Incubate the plate overnight to allow colonies to form. Screen the colonies by PCR to identify the correct colonies formed from the correct integration events. Use three pairs of PCR primers. The first checks the upstream end of the integration. The forward primer binds to the genomic region upstream of the sequence found in the upstream homology arm and points to the direction of the integration cassette. The reverse primer binds to the M13 gene in the integration cassette and points to the direction of the E. coli genomic region upstream of the desired integration event.
[0126] The second pair of primers checks the downstream end of the integration. The forward primer binds to the M13 gene in the integration cassette and points to the E. coli genomic region downstream of the desired integration. The reverse primer binds to the E. coli genomic region downstream of the downstream homology arm sequence and points to the integration cassette. If both the first and second pairs of primers produce PCR products of the predicted size, it can be assumed that some cells in the screened colonies have the desired integration event.
[0127] The third primer pair consists of a forward (genomic) primer from the first pair and a reverse (genomic) primer from the second pair. If these primers produce an 8kb PCR product containing the pyrF homology arm, the protein-coding region of M13KO7, and the spectinomycin marker, it can be assumed that some cells in the colony have the correct integration event. If these primers produce a 2kb product containing only the pyrF homology arm and the pyrF gene, it can be assumed that some cells in the colony retain the wild-type pyrF sequence.
[0128] Correct colonies produce PCR products using all three primer pairs, but there is no evidence of a smaller 2kb PCR product produced by the third primer pair. A subset of these correct colonies is inoculated into LB supplemented with 50ug / ml spectinomycin and grown overnight. 1 ml of the overnight culture is mixed with 1 ml of 50% glycerol / 50% aqueous solution and stored at -80°C to create a permanent bank of new strains. Genomic DNA is extracted from another 1 ml of the overnight culture using the Sigma GenElute Genomic Prep Kit (catalog number NA2120, Sigma Aldrich, Burlington, Massachusetts). This purified genomic DNA is sequenced by Azenta, Inc. (Chelmsford, Massachusetts) using standard Illumina next-generation sequencing equipment and protocols. The fastQ files created by Azenta are aligned to both the expected sequence at the embedded locus and the wild-type sequence at the same locus using the embedded mapping algorithm in Geneious Prime (Copyright (C) 2005-2022 Biomatters Ltd.). Samples in which the reads do not align to the wild-type pyrF gene, and some reads align to all of the protein-coding sequences of the first transcription unit of M13KO7 without any sequence mismatches, are considered correct. A glycerol stock associated with one of these correct samples is selected for future steps, and the others are discarded.
[0129] Removal of reconvening plasmids The resulting E. coli strain, possessing the precise incorporation of all protein-coding genes from the first transcription unit of M13 ko7 and the deletion of pyrF, was streaked onto LB agar supplemented with 50 ug / ml spectinomycin and incubated overnight at 37°C. Notably, the LB agar did not contain kanamycin, which was used to maintain the reconciling plasmid. A single colony was inoculated into LB broth supplemented with 50 ug / ml spectinomycin and incubated overnight at 42°C to prevent the temperature-sensitive pSC101 replication factor from initiating plasmid replication. A dilution series of this culture was plated onto LB agar supplemented with 50 ug / ml spectinomycin to obtain colonies constructed by single cells. These plates were incubated overnight at 37°C, and the resulting colonies were patched onto both LB agar supplemented with 50 ug / ml spectinomycin and LB agar supplemented with both 50 ug / ml spectinomycin and 50 ug / ml kanamycin. Colonies that grow only on the former plate and not on the latter plate are presumed to lack the reconciling plasmid. One of these plasmid-free colonies is inoculated into LB broth supplemented with 50 ug / ml spectinomycin and incubated overnight at 37°C. The glycerol stock is stored at -80°C.
[0130] Recycling of antibiotic markers Dilute the overnight culture without the plasmid described above to an OD of 0.01 and incubate at 37°C until an OD of 0.30 is achieved. Then, take 1 ml of the culture and wash it three times in ice-cold deionized water by centrifugation at 7000 RCF for 1 minute. Next, resuspend the cells in 49 μl of ice-cold deionized water and mix with 1 μl of plasmid DNA encoding the cre gene, pSC101ts origin of replication, and kanamycin marker. Next, electroporate the washed cell-DNA mixture as described above, resuspend it in 1 ml of LB broth, and incubate at 37°C for 1 hour. Then, plate the cells onto LB supplemented with 50 ug / ml kanamycin to select cells transformed with the cre expression plasmid. Incubate the plate overnight at 37°C. The resulting colonies are transferred to both LB supplemented with 50 ug / ml kanamycin and LB supplemented with 50 ug / ml kanamycin and 50 ug / ml spectinomycin. Colonies that grow on the former plate but not on the latter can be assumed to have recycled the spectinomycin marker via cre-inducible recombination at the loxP marker adjacent to the spectinomycin gene. This can be verified by colony PCR using a primer that binds to the integration cassette outside the loxP recombination site and points in the direction of the region where spectinomycin was present. A 1kb amplicon indicates that the spectinomycin gene was still in place, while a shorter amplicon indicates that it was lost by recombination. The lambda red recombinase system and the cre / lox marker recycling system are described in Tuntufye and Goddeeris, FEMS Microbiol Lett 325 (2011) 140-147, which is incorporated herein by reference in its entirety.
[0131] Exclusion of CRE expression plasmids One colony that showed no growth on a spectinomycin plate and was confirmed to have lost the spectinomycin resistance gene by both colony PCR was inoculated into antibiotic-free LB broth and incubated overnight at 42°C to prevent the temperature-sensitive pSC101 replication factor from initiating replication of the cre expression plasmid. A dilution series of this culture was plated onto antibiotic-free LB agar and incubated overnight at 37°C to obtain colonies constructed by single cells. The colonies were transferred to both antibiotic-free LB and LB supplemented with 50 ug / ml kanamycin and checked for loss of the cre expression plasmid. Colonies that grew on the former plate but not on the latter plate could be presumed to have rejected the plasmid. One such colony was inoculated into antibiotic-free LB and incubated overnight at 37°C. The resulting culture was mixed 1:1 with 50% glycerol / 50% aqueous solution, labeled as KTx00001, and stored at -80°C.
[0132] Incorporation of the second transfer unit Repeat the above steps for the incorporation of the second transcription unit. The only change is in the sequence of the transcription unit, homology arm, and primer used to verify incorporation. The second transcription unit is incorporated at the neutral incorporation site, yhiN in E. coli MG1655. Bernhards et al., ACS Synth. Biol. 2022, 11, 1681-1685. This new strain is denoted KTx00002 and stored at -80°C.
[0133] Testing of new bacterial strains Confirmation that KTx00002 can produce cssDNA when infected with a phagemide (Figure 2) containing morphologically selectable sequences for the F1 ori, packaging signal, and ampicillin resistance gene is achieved by electroporating the phagemide into KTx00002 and then quantifying the cssDNA as follows.
[0134] Phagemids are transformed into KTx00002 using a standard electroporation protocol. Briefly, cells are grown to an OD of 0.3, followed by collection of 1 ml of culture by centrifugation at 6000 RCF for 1 minute. The pellet is washed three times with ice-cold deionized water, and the cells are resuspended in 49 μl of deionized water. 1 μl of phagemid DNA is added to the cells, followed by mixing by gently tapping the tube. The cells and DNA are transferred to a 1 mm cuvette and electroporated using a Bio Rad Gene Pulser at 2.1 kv / cm, 100 ohms, and 25 μF. After recovering the cells in SOC medium with shaking at 37°C for 1 hour, they are washed and plated on uracil-free standard medium. (Bacto CD Supreme Fermentation Production Medium (FPM) catalog number A49737-01, Thermo Fisher, Waltham, MA, USA). Incubate the plate overnight at 37°C. Collect a single colony and inoculate it into 4 ml of uracil-free FPM, then grow it overnight at 30°C. The next morning, measure the optical density (OD) and dilute the culture to OD 0.05 in 750 ml of uracil-free FPM. Then, incubate this culture at 30°C for 8 hours.
[0135] The culture is centrifuged to remove all cells, and phage particles are collected in the supernatant. The phage particles are purified from the fermentation medium by precipitation of polyethylene glycol and sodium chloride (3 wt / vol%), followed by centrifugation at 5000 g. Subsequently, the phage particle pellet is dissolved using a solution of 1% sodium dodecyl sulfate and 200 mM sodium hydroxide. The dissolved phage particle fragments are pelletized by centrifugation at 12,000 g, and the supernatant containing cssDNA is retained. Endotoxin lipopolysaccharides are removed by adding triton® X-100 surfactant, mixing thoroughly, and then discarding the resulting surfactant layer. Subsequently, cssDNA is precipitated by adding 100% ethanol to this solution at -20°C, incubating on ice for 30 minutes, and then pelletizing by centrifugation at 12,000 g. Subsequently, the cssDNA pellet is washed with 75% ethanol at 4°C, incubated on ice for 20 minutes, and centrifuged again at 12,000 g to remove residual salts. The obtained cssDNA pellet is dried to remove residual ethanol, and then resuspended in tris-EDTA(TE) buffer solution or water without nuclease.
[0136] cssDNA is quantified using Qubit's commercially available assay kits and instruments according to the manufacturer's recommended protocol. Nucleic acid purity is quantified using an Agilent Bioanalyzer. Endotoxin levels are measured using an Endosafe® nexgen-PTS® portable spectrophotometer and cartridge assay (catalog number MCS150K, Charles River Laboratories, Wilmington, MA, USA) and USP BET. <85> Quantification is performed according to the manufacturer's suggested protocol using the horseshoe crab mebosite lysate kinetics assay, in accordance with EP BET<2.6.14>criver.com / sites / default / files / resources / Endosafe-PTS-Regulatory-Requirements-USPBET85-EPBET2.6.14.pdf.
[0137] For comparison, the parent strain of KTx0002 (E. coli MG1655) is similarly transformed with the same phagemide and helper plasmid, but this helper plasmid is derived from helper phage M13KO7 by amplifying the kanamycin marker and all protein-coding genes derived from M13KO7 using primers designed to eliminate the F1 origin of replication and packaging signal, and then removing the F1 origin of replication and packaging signal by re-circulating the amplified DNA using an in vitro reaction of DNA ligase. KTx00002 is expected to produce cssDNA with comparable yield and fidelity to the parent control.
[0138] (Example 2) The effect of induced gene expression on KTx00002 Strain KTx00002 contains an integrated copy of two transcription units of M13KO7 and, when transformed with phagemide, achieves cssDNA production comparable to that of its parent strain MG1655 transformed with the M13KO7 helper plasmid and the same phagemide. To achieve high cssDNA production, certain genes derived from M13KO7 are overexpressed.
[0139] To overexpress the M13KO7 gene in E. coli, the T7 polymerase coding sequence is first incorporated into the neutral integration locus yjhV in the E. coli genome. This is achieved in the same manner as the two transcription units of M13KO7 were incorporated. Briefly, the T7 polymerase gene is codon-optimized for expression in E. coli and synthesized by Genscript (No. 28 Yongxi Road, Jiangning District, Nanjing, Jiangsu, China 211100). This is cloned into a plasmid under the control of the lactose-inducible promoter pLac. Downstream of this pLac-T7 polymerase expression cassette are the pheA transcriptional terminator and the spectinomycin marker flanked by the loxP site. Both the expression cassette and the spectinomycin marker are flanked by homology arms and restriction enzyme sites. This plasmid is extracted from the E. coli host using the QIAprep Spin Miniprep Kit (catalog number 27106 Qiagen, Hilden, Germany), digested with restriction enzymes to linearize the vector, and an integration cassette is created separated from the bacterial origin of replication. KTx00002 is transformed with the above-mentioned reconciliation plasmid and grown to an OD of 0.3, at which point the expression of the reconciliation gene is induced with L-arabinose. The cells are incubated for a further 1 hour to express the reconciliation gene. Subsequently, the cells are harvested, washed three times by centrifugation, mixed with 1 μl of the linearized integration cassette, and then electroporated. The cells are allowed to recover overnight, then plated on LB agar supplemented with 50 ug / ml spectinomycin to select the integration cassette. Integration is verified by colony PCR and sequencing as described above. The reconciliation plasmid is excluded as described above. The cells are transformed with the cre expression plasmid as described above, and the markers are recycled as described above. Finally, the cre expression plasmid is removed, and the new strain, labeled KTx00003, is stored at -80°C.
[0140] There is evidence that overexpression of a specific M13 phage gene leads to increased cssDNA production in E. coli. Behler et al., 2022 Oct;119(10):2878-2889. The open reading frames of genes 2 and 10 are involved in phage gene expression, the open reading frames of genes 1 and 11 are involved in the creation of pores for phage secretion in the bacterial membrane, gene 4 is also involved in the creation of secretory pores, and gene 8 is the major coat protein. However, this evidence does not stem from the integrated expression of phage genes in the genome.
[0141] Individual and combined overexpression of phage genes is achieved by cloning coding sequences under the control of the T7 promoter into plasmids containing a transcription terminator downstream of one or more genes, followed by a spectinomycin selection marker flanked by loxP sites. Both the phage gene expression cassette and the spectinomycin marker are flanked by homology arms for integration into the E. coli genome at neutral integration loci. The homology arms are flanked by restriction enzyme recognition sequences, which consequently allow the vector to be linearized before integration into the E. coli genome.
[0142] These plasmids are prepared and digested as described above and sequentially incorporated into the E. coli KTx00003 genome using the above-described reconvening and cre / lox system to create a novel E. coli strain containing a novel overexpression cassette containing two wild-type transcription units of M13KO7, plus T7 polymerase under the control of the lac promoter, and the T7 promoter that drives the expression of phage genes as shown in Table 1. The novel strain is named using the strain nomenclature KTx00004.XY, where "X" is 1, the code refers to an expression cassette that overexpresses p1; "X" is 2, the code refers to an expression cassette that overexpresses p2, and so on for each of genes I-XI. Similarly, if two different phage genes are overexpressed, "Y" is used to identify the second phage gene to be overexpressed. The following table presents various combinations with the identified XY. [Table 1]
[0143] Overexpression of phage genes from the T7 promoter is verified by reverse transcription quantitative polymerase chain reaction (rt-qPCR). Cultures of new strains and parental control strains are grown overnight, then diluted to OD 0.01, and grown with or without 100 μM IPTG. After 5 hours of growth, cells are collected by centrifugation. mRNA is prepared using the Qiagen RNeasy Mini kit (catalog no. 74104, Qiagen Hilden, Germany) according to the manufacturer's suggested protocol. cDNA is reverse transcribed using TaqMan® Reverse Transcription Reagents (catalog no. N8080234, ThermoFisher Scientific, Waltham, MA) according to the manufacturer's recommended protocol. The cDNA preparations are normalized to equal concentrations and amplified in multiplexed reactions using primers targeting each of the M13 genes and a primer targeting the housekeeping control gene rpoD. Taqman probes (ThermoFisher Scientific, Waltham, MA) are added to the reactants to quantify the amount of each PCR product produced. For the M13 gene, the probe has a 5' fluorescein reporter dye and a 3' Iowa Black® FQ quencher. For the rpoD housekeeping control gene, the probe has a 5' cyanine-5 reporter dye and a 3' Iowa Black® RQ quencher. The reaction is performed using a QuantStudio 7 Flex (Thermofisher Scientific, Waltham, MA). Overexpression is determined by comparing cycles in which the fluorescein signal crosses an arbitrary threshold to cycles in which the cyanine-5 signal crosses the same threshold (delta cycle threshold or Δct), compared to wild-type parent and uninducible control (delta cycle threshold or ΔΔct). Strains expressing the M13 gene from the T7 promoter cross over an arbitrary threshold (normalized for cyanine-5 signaling) after fewer cycles than strains expressing the phage gene from only their native promoters.
[0144] The cssDNA from KTx00004 will be compared to the cssDNA production from its parent strain KTx00002, which contains only the two wild-type transcription units of M13KO7. For this comparison, each strain will be transformed with the same phagemid. Cultures of the phagemid-containing strains will be grown in parallel, and cssDNA will be extracted and quantified.
[0145] (Example 3) Differential expression of phage genes in the production host For experimental purposes, each phage gene is cloned into an integration cassette under the control of high, medium, and low-intensity T7 promoters (see above). Each gene is integrated in parallel into the E. coli KTx0003 genome using the lambda red reconciliation system described above to create 33 new strains, each expressing one phage gene from one T7 promoter. The genes examined using plasmids include phage genes I, II, III, IV, V, VI, VII, VIII, IX, X, and XI. The 33 new strains are tested for their ability to produce cssDNA.
[0146] In this experiment, all genes except the test gene (non-test genes) are placed under the control of their native promoters, which can produce substantially the same levels of non-test gene expression as wild-type phage expression levels. A single test gene is induced by the addition of IPTG to activate T7 polymerase expression, followed by expression of the test gene from the T7 promoter. IPTG is added at various time points in the growth of the production host. The amount of cssDNA produced at the time of IPTG introduction is quantified. For clarity, one example involves placing gene VIII under the control of the T7 promoter (SEQ ID NO: 1), and the remaining genes under the control of their native promoters. T7 polymerase expression is induced using lactose or its analogue IPTG, followed by transcription of gene VIII from the T7 promoter at OD 0.05, 0.5, 1.0, and 3.0. Qubit is used with a Bioanalyzer and commercially available assay kits and protocols designed to work with these instruments to determine the yield and purity of cssDNA.
[0147] To quantify the cssDNA produced from these test strains and the parental control strain KTx0003, the production host was transformed with the same phagemide and cultured in Bacto CD Supreme Fermentation Production Medium (FPM) (catalog no. A4973702 Thermo Fisher Scientific, Waltham, MA USA) using sterile medium as a blank until the indicated OD was reached. At the indicated OD, isopropyl β-d-1-thiogalactopyranoside (IPTG) was added, and the culture was returned to shaking at 200 RPM in a New Brunswick Innova shaking incubator. The culture was continued to grow until the total growth time reached 8 hours. The culture was centrifuged to remove all cells, and phage particles were collected in the supernatant. The phage particles were purified from the fermentation medium by polyethylene glycol and sodium chloride precipitation (3 wt / vol%), followed by centrifugation at 5000 g. Subsequently, the phage particle pellet was dissolved using a solution of 1% sodium dodecyl sulfate and 200 mM sodium hydroxide. The dissolved phage particle fragments are pelletized by centrifugation at 12000g, and the supernatant containing cssDNA is retained. Subsequently, 100% ethanol is added to this solution at -20°C, incubated on ice for 30 minutes, and then pelletized by centrifugation at 12000g to precipitate the cssDNA. The cssDNA pellet is then washed with 75% 4°C ethanol, incubated on ice for 20 minutes, and precipitated again by centrifugation at 12000g to remove residual salts. The resulting cssDNA pellet is dried to remove the ethanol and resuspended in TE buffer solution or nuclease-free water.
[0148] cssDNA is quantified using Qubit's commercially available assay kits and instruments according to the manufacturer's recommended protocol. Nucleic acid purity is quantified using an Agilent Bioanalyzer. Endotoxin levels are measured using an Endosafe® nexgen-PTS® portable spectrophotometer and cartridge assay (catalog number MCS150K, Charles River Laboratories, Wilmington, MA, USA) and USP BET. <85> Quantification is performed according to the manufacturer's suggested protocol using the horseshoe crab mebosite lysate kinetics assay in accordance with EP BET<2.6.14>www.criver.com / sites / default / files / resources / Endosafe-PTS-Regulatory-Requirements-USPBET85-EPBET2.6.14.pdf. [Table 2] [Table 3-1] [Table 3-2] [Table 3-3] [Table 3-4]
[0149] (Example 4) Production of host cells with reduced p5 activity The production strain KTx00003 will be further developed to reduce the relative production of p5 (KTx00006). p5 is an ssDNA-binding protein. By altering five bases in the 5' UTR of gene V, the interaction between gene V mRNA and ribosomes is inhibited, reducing the amount of p5 protein translated by ribosomes.
[0150] Phagemids produce cssDNA through rolling circle amplification using their replicative dsDNA morphology as a template. After cssDNA is produced within the bacterial host, two different things can happen. If p5 levels are low, p5 does not bind to cssDNA, thus allowing the bacterial replication mechanism to polymerize the complementary strand and create more replicative dsDNA. This replicative dsDNA is then used as a further rolling circle amplification template, leading to the creation of more cssDNA. However, if p5 levels are high, p5 binds to cssDNA, preventing the bacterial host mechanism from polymerizing the complementary strand and thus preventing it from being used as a rolling circle amplification template. P5 binding to cssDNA instead causes the cssDNA to be packaged in the phage capsid and effluxed from the host cell. By reducing p5 translation, the cell is able to produce more phage cssDNA templates before switching to packaging cssDNA into the phage capsid and effluxing the phage capsid from the cell. Further accumulation of cssDNA templates within E. coli cells leads to increased production of packaged cssDNA. A schematic of this strategy is shown in Figure 6 (see Lee et al., Optimizing protein V untranslated region sequence in M13 phage for increased production of single-stranded DNA for origami. Nucleic Acids Res. 2021 Jun 21;49(11):6596-6603, which is incorporated herein by reference in its entirety).
[0151] Prior to cloning the first transcription unit as described in Example 1 above, PCR primers for amplifying the first transcription unit were designed so that the 5' UTR (untranslated region) of gene V was changed from wild-type TCACA to GAGGT (as shown in Figure 7, Panel A).
[0152] Furthermore, the PCR primers are designed to amplify the remainder of transcription unit 1 as described in Example 1 by fusion PCR, so that the mutated 5' UTR of gene V is incorporated into transcription unit 1, creating a 2117 bp variation of transcription unit 1 containing the modified 5' UTR sequence of gene V (as shown in Figure 7, panel B).
[0153] Identical E. coli-producing host strains are created using two surrogate versions of transcription unit 1, one containing a mutation at the 5' UTR and the other lacking the mutation. Both strains are transformed with the same phagemide and cultured side-by-side to allow comparison of cssDNA production between the two. While we do not wish to be constrained by any theory, strain KTx00006, which has a mutation at the 5' UTR of gene V, is expected to produce more cssDNA than the control strain lacking the mutation (KTx00003). Analysis of the cssDNA produced from these strains can be performed as previously described.
[0154] (Example 5) Creation of circular single-stranded DNA phagemids with plasmid ori The minimal phagemide skeleton is cloned from the M13KO7 template by amplifying the M13 origin of replication, including the packaging signal, using primers containing 30nt homology tails to the other fragments. The pyrF gene is amplified from the E. coli MG1655 genome, and the synthetic promoter and terminator sequences are attached using primer tails. The pUC19 origin of replication is amplified from pUC19 (catalog number N3041S, New England Biolabs, Ipswich, MA, USA) which has primer tails homologous to the other fragments. These three fragments are assembled by Gibson isothermal assembly, Gibson et al., Nat Methods. 2009 May;6(5):343-5. The Golden Gate Assembly site (Engler et al., PLoS One. 2008;3(11):e3647) is inserted into this phagemide by including these in the primer tails used for fragment amplification and assembly. The Golden Gate Assembly site includes BsaI, BsmBI, and PaqCI. The M13 origin of replication is included in the design so that the phagemide can replicate as cssDNA within E. coli. A packaging signal is included so that the M13 protein packages cssDNA into the phage particle. The pUC19 origin of replication is included so that the phagemide can replicate as dsDNA within E. coli. The pyrF gene is used as a selectable nutrient requirement marker for use in KTx00001 and its offspring, which are uracil-requiring strains resulting from disruption of the pyrF gene genome copy by insertion of M13KO7 transcription unit 1 (see Example 1). The Golden Gate Assembly site is included to facilitate insertion of user-defined sequences. User-defined sequences can be synthesized or amplified using primers so that they are flanked by BsaI, BsmBI, or PaqCI restriction sites.Next, the insert and minimal phagemid are incubated with the appropriate restriction enzyme, buffer, and ligase, and the temperature is cycled between 37°C for restriction enzyme cleavage and 16°C for ligation. Because the restriction enzyme recognition sequence is asymmetric and several base pairs away from the cleavage site, the recognition sequence is inserted into the phagemid and insert so that it is cleaved after the initial digestion. If these fragments are subsequently assembled in the desired manner to insert a user-defined sequence into the phagemid, the recognition sequence is excised and no further cleavage can occur. However, if the fragment is ligated back to the recognition sequence to reform the initial input, the recognition sequence is recreated, and another round of cleavage may follow. In this way, the reaction proceeds in one direction until almost all restriction recognition sites have disappeared and almost all inserts have been correctly inserted into the phagemid skeleton.
[0155] Notably, the pyrF nutritional requirement marker is included in the phagemid instead of conventional antibiotic markers to enhance the safety of the cssDNA produced from the system. Since one application of the cssDNA produced from this system is human therapeutic gene and cell therapy, including antibiotic marker sequences is undesirable. The exclusion of antibiotic marker sequences prevents the system from spreading antibiotic resistance to microorganisms that could potentially infect human patients with antibiotic-resistant strains.
[0156] (Example 6) Increase in cssDNA yield associated with increased phagemide copy number Considering the interrelationships between multiple factors, including both the copy number of the helper plasmid (or the copy number of the incorporated phage gene), phage gene expression, and the phagemid copy number for fermentation, it was unpredictable how changes in phagemid copy number would affect cssDNA yield in fermentation-based production methods. This embodiment provides results demonstrating that an increase in phagemid copy number increased cssDNA production from bacterial production strains containing phagemid and helper plasmids.
[0157] As shown in Table E1 below, a series of phagemids containing different origins of replication were constructed. A map of phagemids containing the pUC19 origin is shown in Figure 8A, and the primers used to generate inc1 and inc2 mutations at the pUC19 origin are shown in Figure 8B. The phagemids were tested in combination with two different helper plasmids: KHP0 and KHP1 (each containing the p15A origin of replication). KHP0 was derived from helper phage M13KO7 by removing the F1 origin and packaging signal. This was achieved by amplifying the kanamycin marker and all of the protein-coding genes from M13KO7 using primers designed to eliminate the F1 origin and packaging signal, followed by re-circularizing the amplified DNA using an in vitro reaction of DNA ligase. KHP1 was similarly derived from helper phage M13KO7, but further includes gene V attenuation created by mutating the 5 base pairs immediately upstream of gene V to reduce the translation rate of gene V. Gene V is involved in regulating the transition from replicating double-stranded DNA to non-replicating single-stranded DNA within bacterial cells. Reduced expression of gene V has been shown to result in longer replication periods as double-stranded DNA, thus leading to greater accumulation of dsDNA templates and higher cssDNA yield. [Table E1]
[0158] Standard NEB5alpa chemically competent E. coli cells were co-transformed with Kano Helper Plasmid 1 (KHP1) and one of the four phagemide variants shown in Figure 9. Individual colonies were selected on LB agar plates supplemented with 50 ug / ml kanamycin (helper plasmid) and 100 ug / ml carbenicillin (phagemide). Six individual colonies were collected from each plate and inoculated into starter cultures in 2x YT medium supplemented with 50 ug / ml kanamycin and 100 ug / ml carbenicillin. The starter cultures were grown overnight until saturated, then diluted to approximately OD600 0.04 in the same medium. The cultures were then grown at 30°C for 18 hours with shaking. The cultures were collected, and bacterial cells were removed from the medium by centrifugation. The supernatant containing phage particles, and therefore cssDNA, was pipetted into a new container, and the concentration of cssDNA in each sample was analyzed by qPCR using primers and probes targeting the F1 origin of replication on cssDNA, as well as a FAM reporter and an NFQ-MGB quencher.
[0159] As shown in Figure 9, the inc2 and inc1&2 phagemide mutations resulted in higher cssDNA yields compared to the "standard" pUC19 phagemide. Phagemids with reduced copy numbers yielded significantly lower yields compared to standard pUC19, inc2, or inc1&2. The inc1 mutation is a C59T mutation relative to the standard pUC origin, and the nucleotide numbering is based on the sequence of the standard pUC origin shown in SEQ ID NO: 27. The inc2 mutation is a C92T mutation relative to the standard pUC origin, and the nucleotide numbering is based on the sequence of the standard pUC origin shown in SEQ ID NO: 27. The inc2 and inc1 / inc2 phagemides used in these experiments also contained the G62A mutation compared to the standard pUC origin shown in SEQ ID NO: 27, although this mutation is not expected to affect the function of the origin.
[0160] These results suggest that an increase in phagemid copy number can lead to an increase in the yield of cssDNA produced from phagemids. Decreased copy number origins yield an estimated 15 copies per cell, while standard pUC19 origins yield an estimated 1000 copies per cell, and inc1&2 mutant pUC19 origins yield an estimated 4000 copies per cell. These copy number estimates vary significantly depending on the growth medium and growth stage.
[0161] A significant increase in cssDNA yield with phagemids containing inc1&2 mutations was also observed in larger cultures. Standard NEB5alpa chemically competent E. coli cells were co-transformed with Kano Helper Plasmid 1 (KHP1) and one of two phagemid variants (standard (pUC19) or inc1 / inc2), as shown in Figure 10. Individual colonies were selected on LB agar plates supplemented with 50 ug / ml kanamycin (helper plasmid) and 100 ug / ml carbenicillin (phagemid). Fifteen individual colonies were collected from each plate and inoculated into starter cultures in 2xYT medium supplemented with 50 ug / ml kanamycin and 100 ug / ml carbenicillin. Starter cultures were grown overnight to saturation, then diluted to approximately OD600 0.04 in the same medium in either 100 ml (9 cultures) or 500 ml (6 cultures) volumes. The cultures were then grown at 30°C for 18 hours with shaking. The cultures were harvested, and bacterial cells were removed from the medium by centrifugation. The supernatant containing phage particles, and therefore cssDNA, was pipetted into a new container, and the cssDNA concentration in each sample was analyzed by qPCR using primers and probes targeting the F1 origin of replication on the cssDNA, as well as a FAM reporter and NFQ-MGB quencher. As shown in Figure 10, inc1&2 mutations in the phagemid resulted in at least a 5-fold increase in cssDNA yield compared to the yield obtained with the wild-type pUC19 origin. The increased cssDNA yield was reproducible for both 100 mL and 500 mL culture production.
[0162] The results above demonstrate that the inc mutation in the pUC19 phagemid replication origin for cssDNA production confers a significant increase in cssDNA yield compared to the pUC19 replication origin (for example, yields increased approximately 5 to 10 times compared to the wild-type pUC19 origin, and more than 100 times compared to the p15A origin).
[0163] (Example 7) Increase in cssDNA yield associated with increased helper plasmid copy number This example provides results demonstrating that increasing the copy number of a helper plasmid, including in combination with high copy number phagemids (inc1&2 phagemids with an estimated 4000 copies per cell), can also increase cssDNA yield.
[0164] As shown in Table E2 below, a series of different helper plasmids were constructed. [Table E2]
[0165] E. coli-producing host cells were transformed with the helper plasmids shown in Table E2, combined with phagemide 163 as described in Example 6. Individual colonies of double-transformed E. coli were collected, grown in starter cultures, then diluted and grown for small-scale production analysis. After 18 hours of culture, the bacteria were lysed and relative cssDNA yield was analyzed by qPCR. As shown in Figure 11, the pUC19 origin helper plasmid (148) resulted in a higher cssDNA yield from the phagemide at 18 hours than the lower copy number helper plasmids 81 (pSC101 ori) and 114 (p15A ori).
[0166] (Example 8) Creation of production stocks Host strain selection The first step was to select a background strain to create the production mechanism. CssDNA production has typically been carried out in E. coli cloning host strains, e.g., DH5α, JM109, XL-1 Blue, or M1061. These common laboratory strains have several advantages because they have been bred to optimize plasmid DNA cloning and preparation. They lack recA to reduce recombination events in the desired cssDNA product. They lack endA to reduce DNA degradation. They lack host nucleases to increase transformation efficiency. They have F pili eliminated to avoid cell reinfection by infectious phage particles. However, they are also archaic strains that have been adapted over many years and bred or mutated to have other deletions that may or may not be beneficial for cssDNA production. They generally have slower growth rates and lower biomass yields than wild-type strains with many mutations.
[0167] The inventors hypothesized that wild-type strains may serve as better production hosts due to their overall improved fitness, including improved growth rate and biomass yield per unit of culture medium compared to highly developed laboratory strains. To test this hypothesis, the inventors obtained a small library of E. coli background strains, including the common laboratory strains mentioned above and other strains with lower levels of adaptation. These included BW25113, K12, and MG1655. A detailed list of the strains is included in Table 4. [Table 4-1] [Table 4-2] The inventors chose to use a standard medium for cssDNA production because it reduces the cost of scaling up and down production, increases the reproducibility of production runs, allows the use of nutritional requirement markers instead of antibiotic markers, and therefore enhances the safety of the final cssDNA product. The inventors chose to use a standard medium recipe based on Riesenberg's 1991 recipe because it is well-established and has been commonly used since its publication. Riesenberg, et al. High cell density cultivation of Escherichia coli at controlled specific growth rate. J Biotechnol. 1991 Aug;20(1):17-27. The Reisenberg recipe has also been used to successfully produce large quantities of cssDNA in a fed-batch reactor system. Kick, et al. Efficient Production of Single-Stranded Phage DNA as Scaffolds for DNA Origami. Nano Lett. 2015 Jul 8;15(7):4672-6. Hereinafter, the inventors refer to this medium as Riesenberg medium. However, other bacterial cell culture media can also be used.
[0168] The inventors first screened their collection of bacterial strains for growth in Riesenberg medium. Single colonies were inoculated in triplicate in LB medium and grown overnight as pre-cultures. The following morning, the cultures were diluted in Riesenberg medium to an OD500 of 0.05. These cultures were incubated at 30°C for 24 hours with shaking, and the final OD600 was recorded. Cells from each strain were streaked onto antibiotic-free LB agar to obtain single colonies.
[0169] Figure 12 shows the results of comparing the final OD for the tested strains. As shown in Figure 12, we found that various isolates of DH5α, XL-1 Blue, and MC1061 reached the lowest final OD. Strains with a lower degree of production, such as MG1655 and K12, reached higher ODs (Figure 12). Various isolates of BW25113 reached the highest OD (Figure 12). These strains have a slightly reduced genome with deletions of lactose, arabinose, and rhamnose operons, which may offer growth benefits because the cells have less DNA to replicate and less protein to produce. These alternative sugar operons are not required in fermentation cultures where glucose is provided as the sole carbon source. Furthermore, since BW25113 was the parent strain of the Keio knockout collection, single-gene deletion strains already existed in this background for genes essential for the synthesis of certain amino acids and nucleosides. In fact, when we included three of these deletion strains—bac012(ΔpyrF), bac013(ΔpyrF), bac016(ΔpyrF, Δkan), and bac021(ΔleuB)—in our screening, all performed near the top. The best performing strain was the BW25113 parent of the single-gene deletion strains mentioned above, but strains lacking either pyrF or leuB did not perform significantly worse, and conveniently, they were already suitable for use in selecting and maintaining plasmids with nutritional requirement markers. Therefore, we selected bac016 for further development because it grows vigorously in our preferred culture medium, already possesses deletions for pyrF and kanamycin resistance markers, and is the most advanced in terms of the production process. The other strains can also be used for cssDNA production by any of the methods disclosed herein.
[0170] After identifying bac016 as a strain that grows well in a standard culture medium selected by the inventors, the inventors then verified that cssDNA could be produced using helper plasmids and phagemids before any genome production step. bac016 was co-transformed with a standard DH5α-producing host bac001 using a helper plasmid (cdsDNA114) and phagemid (cdsDNA111) system. In another experiment, bac016 was co-transformed with the cdsDNA114 helper plasmid and a novel cdsDNA117 phagemid that expresses a pyrF nutrient requirement marker instead of an antibiotic resistance gene.
[0171] The cdsDNA114 helper plasmid expresses all M13 genes and kanamycin markers and possesses the p15A origin of replication. It lacks the M13 origin of replication and packaging signal, and therefore cannot replicate as an infectious phage particle. The cdsDNA111 phagemide has a pUC origin of replication and expresses RFP and ampicillin resistance markers. Unlike the helper plasmid, it contains the M13 origin of replication and packaging signal, making it possible to produce it as cssDNA and package it in a phage capsid. These capsids lack the necessary M13 genes and therefore cannot replicate within a new bacterial host. The cdsDNA117 phagemide is identical to the cds111 phagemide except that it expresses the pyrF nutrient requirement marker instead of the ampicillin antibiotic marker.
[0172] Colonies from each of the three transformation reactions described above were inoculated in triplicate in Riesenberg medium and grown overnight at 37°C as a pre-culture with appropriate antibiotics and, if necessary, uracil supplementation to complement pyrF nutritional requirements. Subsequently, the culture was diluted to OD600 0.05 and incubated at 30°C for 48 hours to test cssDNA production. The broth was collected and centrifuged to isolate E. coli cells from the supernatant containing secreted phage particles. Subsequently, the supernatant was transferred to a clean tube and heated at 99°C for 10 minutes to lyse the phage particles containing cssDNA. This material was then used as a template for qPCR targeting the F1 origin of replication present on the cssDNA.
[0173] This experiment demonstrated that BW25113 can produce cssDNA when transformed by the conventional two-plasmid helper plasmid-phagemid production mechanism. When transformed with an ampicillin-resistant phagemid, it produced the same amount of cssDNA as a conventional DH5α-producing host transformed with the same phagemid. When transformed with a novel cdsDNA117 phagemid expressing the pyrF nutrient requirement marker, it produced significantly more cssDNA.
[0174] Since DH5α does not have a pyrF deletion, cssDNA production using phagemids expressing the pyrF nutrient requirement marker was not tested with this strain. However, strains like DH5α can be induced to have a deletion in pyrF for compatibility with nutrient requirement selection.
[0175] Figure 13A shows the verification of cssDNA production in the unproduced BW25113 parent strain. The conventional production host bac001 (DH5α) and the newly proposed production host bac016 (BW25113 ΔpyrF, Δkan) were transformed using the conventional production mechanism consisting of the cdsDNA114 helper plasmid and the cdsDNA111 phagemide, respectively. bac016 was further co-transformed with the same cdsDNA114 helper plasmid and a novel cdsDNA117 phagemide expressing the pyrF nutritional requirement marker. Colonies were inoculated in triplicate in Riesenberg medium supplemented with appropriate antibiotics and uracil, plasmids were selected, grown overnight in pre-culture, then diluted to OD600 0.05 in the same medium and incubated at 30°C for a further 48 hours to produce cssDNA. cssDNA was quantified using qPCR targeting the F1 origin of replication present on the cssDNA backbone. The data confirm that bac016, a non-producing host, produced the same amount of cssDNA as a standard bac001 DH5α-producing host, and indeed produced significantly more cssDNA when the pyrF-containing cdsDNA117 phagemide was used as a cssDNA template.
[0176] Incorporation of the M13 transfer unit The M13 genome is composed of two distinct transcription units. The first consists of genes II, X, V, VII, IX, and VIII, which are expressed by strong promoters. The second consists of genes III, VI, I, XI, and IV, which are expressed by weak promoters. Two of these genes overlap with the open reading frames of the other two genes (i.e., the gene encoding p10 is entirely contained within the coding region of p2, and the gene encoding p11 is entirely contained within the open reading frame of pI). PA, PB, and PH are strongly active promoters (see Figure 1B). Two additional weak promoters are denoted as PZ and PW. Two terminators (T) are strongly rho-independent terminators, and a third, weaker terminator (T(weak)) is rho-dependent. Smeal et al., Simulation of the M13 life cycle I: Assembly of a genetically-structured deterministic chemical kinetic simulation, Virology, Volume 500, 2017, Pages 259-274.
[0177] To create a stable cell line capable of producing M13 phage particles while being highly unlikely to recombine with phagemids to produce competent phages, each transcription unit of the M13 bacteriophage, containing a protein-coding gene and lacking a phage replication origin and phage packaging signal, was cloned from the publicly available helper phage M13KO7 (catalog number N0315S, NEB, Ipswich, MA) (SEQ ID NO: 24) and incorporated at two loci into a ΔpyrF version of the E. coli BW25113 genome produced as part of the Keio knockout collection called isolate JW1273-1. Baba, et al, Construction of Escherichia coli K-12 in-frame, single-gene knockout mutants: the Keio collection. Mol Syst Biol. 2006;2:2006.0008. To modulate M13 protein production, the gene encoding T7 bacteriophage RNA polymerase was also incorporated into the host strain at a different locus under the control of the E. coli lac operator and promoter. This allowed for the inducible expression of T7 polymerase by the addition of isopropyl β-d-1-thiogalactopyranoside (IPTG), thereby enabling the inducible expression of the M13 phage gene. The JW1273-1(bac012) host strain, an E. coli BW25113 ΔpyrF Keio isolate, initially contained a spectinomycin resistance marker flanked by FRT recombination sites instead of the pyrF gene of wild-type E. coli BW25113. This was genetically deleted to create a strain lacking the spectinomycin resistance marker.
[0178] Incorporation of T7 polymerase cassette To create strains exhibiting inductive expression of the M13 phage gene, the T7 polymerase gene, under the control of an inductive promoter (IPTG-inducible lac promoter), was inserted at the endA locus using reconciliation and CRISPR-negative selection. The endA locus was selected as the insertion site because the EndA protein is known to degrade DNA, and deletion of endA is known to improve the yield and quality of DNA produced in E. coli. (Lin, JJ (1992) Endonuclease A degrades chromosomal and plasmid DNA of Escherichia coli present in most preparations of single stranded DNA from phagemids. Proc. Natl. Sci. Counc. Repub. China B 16 1-5). Correctly identified the correct insertions were determined by colony PCR and Oxford Nanopore sequencing.
[0179] After incorporating an inducible T7 polymerase expression cassette, the inventors then verified that the novel strain bac035 could actually induce T7 polymerase expression and that this polymerase could transcribe genes from the T7 promoter. To test this, a pT7:emGFP expression cassette (containing a nucleotide sequence for emGFP expression driven by the pT7 promoter) was incorporated into the strain. Specifically, strain bac036 was created by incorporating pLac:T7 polymerase from E. coli BL21(DE3) into the endA locus of the BW25113 ΔpyrF background strain. Subsequently, strain bac058 was created by inserting a cassette expressing emGFP from the T7 promoter into yhaV. This strain was grown in IPTG at various concentrations to induce the expression of T7 polymerase and, as a substitute, emGFP. The cultures were grown in a shaking incubator at 30°C for 48 hours. Absorbance at 600 nm (OD600) and GFP fluorescence were measured at the completion of the growth period using a Biotek Synergy H1 plate reader. The obtained arbitrary fluorescence units were divided by OD600 and compared with the parent strain bac036 lacking the GFP expression plasmid (Figure 13B). As shown in Figure 13B, the results demonstrated induction of pLac:T7 polymerase expression.
[0180] Incorporation of the first transfer unit The first transcription unit (TU1) of the M13 genome was incorporated into strain bac036 using reconciliation and antibiotic selection. TU1 was cloned into four circular double-stranded plasmids, cdsDNA206, cdsDNA207, cdsDNA208, and cdsDNA209, along with homology arms for integration into the E. coli BW25113 genome at the intA locus, and a spectinomycin selection marker to enable selection of the integrater. TU1 contains M13 genes II, V, VII, VIII, and IX. Since the optimal expression levels of the integrated M13 genes were not known beforehand, each of the four plasmids had a promoter of different strengths expressing mRNA at different levels. cdsDNA206 had a moderate-strength T7 promoter, cdsDNA207 had a weak T7 promoter, cdsDNA208 had a wild-type M13 TU1 promoter, and cdsDNA209 had a strong T7 promoter. Based on data from Komura et al., High-throughput evaluation of T7 promoter variants using biased randomization and DNA barcoding, Plos, 2018, the T7 promoter was mutated from the wild-type T7 promoter to achieve these different expression levels. This plasmid contains a p15A origin for replication in E. coli and an ampicillin resistance marker for plasmid selection in E. coli. A spectinomycin selection marker was flanked by an FRT recombination site to recycle the marker during Flp recombinase expression.
[0181] bac090~bac092 expressed T7 polymerase from the lac promoter and M13 TU1 from three different promoters (integrated into the E. coli BW25113 genome at the intA locus); a strong T7 promoter, a weak T7 promoter, and a native M13 TU1 promoter, respectively.
[0182] Incorporation of the second transfer unit Next, a second transcription unit (TU2) was incorporated into bacterial strains bac090-092 at the intZ locus to create small combinatorial libraries with different combinations of TU1 and TU2 expression levels. TU2 contains M13 genes III, VI, I, and IV.
[0183] Four different promoters were tested to drive TU2 expression. In cdsDNA215, TU2 was driven by a strong T7 promoter; in cdsDNA216, TU2 was driven by a moderate T7 promoter; in cdsDNA217, TU2 was driven by a weak T7 promoter; and in cdsDNA218, TU2 was driven by its wild-type M13 TU2 promoter. The intZ gene was selected as the integration locus because it is a prophage-associated integrase that does not have any known beneficial function for E. coli. The resulting new strains were denoted as bac105-110, and glycerol stocks were stored at -80°C.
[0184] By incorporating an IPTG-inducible T7 polymerase cassette, as well as a small library of promoters driving both M13 TU1 and M13 TU2, into the same strain, we hereby obtained a set of strains containing all the genes of the M13 genome, as well as the T7 polymerase and promoter system.
[0185] Assay of a new strain for cssDNA production After creating a small library of strains expressing M13 TU1 and M13 TU2 from promoters of different strengths under the control of IPTG-inducible T7 polymerase, the inventors then attempted to assay the novel strains for their ability to produce cssDNA.
[0186] First, strains bac105-bac110 were transformed by electroporation using phagemide cdsDNA111 containing an F1 origin of replication, an F1 packaging signal, and an ampicillin resistance gene and RFP expression cassette. Triple colonies were collected and inoculated into 500 µl LB supplemented with 50 µg / ml kanamycin, and grown overnight at 37°C. The following morning, the optical density (OD600) at 600 nm was measured, and the culture was diluted to OD 0.05 in 200 µl Riesenberg standard medium. Subsequently, the culture was incubated at 30°C until saturated, which took approximately 48 hours at 30°C.
[0187] The culture was centrifuged at 4000 RCF to remove E. coli cells containing double-stranded plasmid DNA, and the supernatant containing phage particles was collected for analysis. Subsequently, the phage particles were lysed by heating the solution at 99°C for 10 minutes. The supernatant containing cssDNA was retained for use as a qPCR template. cssDNA was quantified using qPCR calibrated with a standard curve for cdsDNA111 DNA.
[0188] Figure 14A shows the results of screening small combinatorial libraries of integrated TU1 and TU2 driven by promoters of different strengths. Strains possessing integrated T7 polymerase, M13 TU1 and M13 TU2 were transformed with the phagemide cdsDNA111 to create hosts theoretically capable of producing cssDNA. Single colonies were collected in quintuples and inoculated into 500 μl of LB Miller supplemented with carbenicillin to select for phagemides, and grown overnight as pre-cultures. Subsequently, these cultures were diluted to OD600 0.05 and incubated on plates for approximately 48 hours until saturated. The supernatant was collected, and phage particles were lysed to release their cssDNA for use as qPCR templates. CssDNA was quantified using qPCR primers targeting the F1 ori present on the cssDNA and standard curve cdsDNA111 DNA. bac105, expressing both M13 TU1 and TU2 from a strong T7 promoter, produced the most cssDNA. Other strains with weaker promoter combinations did not produce significantly more DNA than negative controls lacking phagemids that could be used as templates for cssDNA production.
[0189] Next, the inventors attempted to reproduce the above data by comparing them with other production hosts lacking the incorporated M13 gene. As a control, the parent strain bac016 (E. coli BW25113 ΔpyrF) and the standard DH5α cssDNA production host bac001 were transformed with the helper plasmid (cdsDNA114) derived from the helper phage M13KO7 by amplifying the kanamycin marker and all protein-coding genes using the same cdsDNA111 phagemide and primers designed to eliminate the F1 origin and packaging signal, and then removing the F1 origin and packaging signal by re-circulating the amplified DNA using an in vitro reaction of DNA ligase.
[0190] Figure 14B shows the results of comparing the selected and produced strains with their non-produced parental strains. The produced host strains expressing the M13 phage gene from their genomes were transformed with cdsDNA111 phagemide. The non-produced parental strain bac016 was transformed with the helper plasmid cdsDNA114 ("HP" in Figure 14B) to express the M13 phage gene and the same cdsDNA111 phagemide. Transformed colonies were inoculated in quintuples and grown for 48 hours at 30°C in 200 μL of medium in a 96-well microplate containing appropriate antibiotics for plasmid maintenance. The supernatant was then collected by centrifugation to remove E. coli, and the phage particles were lysed by boiling at 99°C for 10 minutes. This was then used as a template for qPCR to quantify cssDNA production. A standard curve was used for qPCR to measure the presence of F1 Ori in solution as a surrogate for cssDNA concentration. Similar to previous experiments, bac105, expressing M13 TU1 and TU2 from the strongest promoter pair, outperformed the other three strains tested (bac106, bac108, and bac110) (all of which had weaker promoters driving M13 gene expression). Notably, bac105 outperformed its parent strain bac016, which utilized a conventional two-plasmid production system rather than an integrated phage gene like bac105. In all three negative control strains lacking a complete cssDNA production system, little to no cssDNA was detected.
[0191] Figure 15A shows a comparison of the best-produced production host (bac105) using a conventional helper plasmid system with its parent strain and the conventional production strain DH5α. All three strains were transformed with the phagemide cdsDNA111 to serve as cssDNA production templates. The unproduced DH5α and BW25113 strains bac001 and bac016 were also transformed with the helper plasmid cdsDNA114 to express the M13 phage gene. The strains were grown in 200 μL of medium in a microplate containing appropriate antibiotics for plasmid maintenance at 30°C for 48 hours. The supernatant was collected by centrifugation to remove E. coli, and the phage particles were lysed by boiling at 99°C for 10 minutes. This was then used as a template for qPCR to quantify cssDNA production. A standard curve was used for qPCR to measure the presence of F1 Ori in solution as a surrogate for cssDNA concentration. The generated bac105 strain, expressing M13 integrated copies of TU1 and TU2 derived from a strong T7 promoter, produced more cssDNA than either of the two strains expressing the M13 phage gene derived from a helper plasmid. Notably, since bac016 was the parent strain of bac105, the difference in cssDNA production is likely due to differences in M13 gene expression rather than differences in strain background. Also noteworthy is bac001, which has similar helper plasmids and phagemids, and is a common cssDNA-producing strain in industrial and academic settings.
[0192] Integration of pT7 that drives individual M13 gene VIII The innate stoichiometry of the various M13 genes present in the wild-type M13 genome may not be optimal for high levels of cssDNA production, particularly for cssDNA constructs whose qualities, such as length and GC content, differ significantly from those of the M13 genome. For example, the M13 shaft is composed of thousands of copies of the gene VIII protein product. The shaft lengthens to accommodate longer cssDNA constructs and contracts to accommodate shorter ones. The production of cssDNA constructs, which differ significantly in length from the 7KB of the M13 wild-type genome, may be improved by different ratios of gene VIII expression to other M13 genes.
[0193] We generated two strains: bac132 (in addition to bac105, gene II is integrated at the mazF locus) and bac135 (in addition to bac105, gene VIII is integrated at the mazF locus). The mazF locus was selected as the integration site because it is the toxic portion of the mazE / mazF toxin / antitoxin system. Deletion of the toxic gene mazF was thought to potentially have beneficial effects on these strains.
[0194] Figure 15B shows the results of assays for cssDNA production across the generated strains incorporating additional individual phage genes. bac001 (DH5α not produced), bac016 (BW25113 not produced), bac105 (best produced strain to date), bac132 (bac105+ gene II), and bac135 (bac105+ gene VIII) were transformed with cdsDNA111 RFP phagemide and grown for 48 hours in Riesenberg standard medium supplemented with uracil. Broth was collected and centrifuged to remove E. coli cells from the culture. The supernatant was transferred to a clean tube and heated at 99°C for 10 minutes to lyse the phage particles and release cssDNA into the medium. This solution was used as a template for qPCR using primers targeting the RFP gene and a standard curve for cdsDNA111 RFP phagemide. The newly created strain bac105 performed better than the control strain that had not been created. Its daughter strain bac132, which incorporated an extra copy of gene II, performed poorly, but another daughter strain bac135, which inserted an extra copy of gene VIII, performed slightly better than its parent bac105. Because the difference was small, the inventors chose to compare the two strains bac105 and bac135 with each other in future experiments.
[0195] After identifying strain bac135, containing T7 polymerase, TU1, TU2, and individual copies of gene VIII, as the best-producing strain among two strains incorporating individual M13 genes, the inventors then aimed to test the production of alternative cssDNA sequences other than the RFP sequence from cdsDNA111. One hypothesis was that adding an extra copy of gene VIII would allow cells to produce more cssDNA, particularly if the cssDNA sequence was significantly longer than the 7kb genome of the wild-type M13 genome. This is because the shaft of the M13 phage capsid, composed of thousands of copies of p8 (the product of gene VIII), elongates and contracts to accommodate cssDNA sequences of different lengths. Therefore, since longer cssDNA sequences require more copies of the p8 protein to be packaged in the capsid, it is logical that more gene VIII expression would promote the production of longer cssDNA sequences.
[0196] (Example 9) Production of cssDNA encoding eutrophin and dystrophin This example demonstrates the ability of the produced strain described in Example 8 to produce cssDNA containing long gene sequences encoding dystrophin and eutrophin. The produced cssDNA system advantageously enables the production of long cssDNA templates useful for gene therapy.
[0197] The strains were assayed by transforming them with phagemids cdsDNA100 (eutrophin) and cdsDNA101 (dystrophin). Control strains bac001 and bac016, which were not produced, were also transformed with helper plasmid cdsDNA114 to express the M13 phage gene. Transformed colonies were inoculated in triplicates and grown for 48 hours at 30°C in 200 μL of medium in a 96-well microplate containing appropriate antibiotics for plasmid maintenance. Subsequently, the broth was collected by centrifugation to remove E. coli cells, and the supernatant containing phage particles was collected and boiled at 99°C for 10 minutes to lyse the phage particles. This was then used as a template for qPCR to quantify cssDNA production. A standard curve was used for qPCR to measure the presence of AAVS1 homology arm sequences in solution as a surrogate for cssDNA concentration. Figure 15B shows the results for the production of various cssDNA sequences from the produced host strains.
[0198] bac001 (conventional DH5α), bac016 (unproduced parent), bac105 (best produced strain to date), and bac135 (bac105 + extra copy of gene VIII) were assayed. The two produced strains, bac105 and bac135, produced more cssDNA than their unproduced comparison strains, and such production was more consistent with smaller standard deviations between biological repeats (bioreps). Notably, both eutrophin and dystrophin constructs are therapeutically associated with the treatment of Duchenne muscular dystrophy, and each construct is longer than 13kb when produced as cssDNA. Eutrophin can be produced at low levels using conventional helper plasmid systems, but dystrophin cannot, while each can be produced at moderate levels using the produced strains. Incorporation of further copies of gene VIII appeared to increase eutrophin cssDNA yield somewhat.
[0199] (Example 10) Production of cssDNA using phagemids with nutritional requirement markers Standard phagemide cdsDNA111 was amplified using PCR primers designed to amplify the entire plasmid except for the ampicillin-resistant cassette. The pyrF gene cassette was amplified from the E. coli DH5α genome along with its 75 bp native promoter. In both PCR reactions, the primers were designed to have a tail that overlapped with the other fragments by 30 nt. These two linear double-stranded DNA fragments were assembled using the NEBuilder HiFi DNA assembly enzyme mix (catalog number E2621L, NEB, Ipswich, MA). The sequence of this novel plasmid was confirmed using Oxford Nanopore technology and stored as cdsDNA117.
[0200] An F1 replication origin is included in the design to enable the phagemide to replicate as cssDNA within E. coli. A packaging signal is included so that the M13 protein packages the cssDNA into the phage particle. A pUC19 replication origin is included so that the phagemide can replicate as dsDNA within E. coli. The pyrF gene is used as a selectable nutrient requirement marker for use in the produced product and its offspring, which are uracil-requiring strains resulting from disruption of the pyrF gene's genomic copy. The Golden Gate Assembly® site (Engler et al., PLoS One. 2008;3(11):e3647) is included in this phagemide. The Golden Gate Assembly® site can be recognized by BsmBI type II restriction enzymes, allowing for easy insertion of novel user-defined sequences into the pyrF phagemide backbone. User-defined sequences can be synthesized or amplified using primers so that they are sandwiched between BsmBI restriction sites. Next, the insert and phagemide are incubated with BsmBI restriction enzyme, buffer, and ligase, and the temperature is cycled between 37°C for restriction enzyme cleavage and 16°C for ligation. Because the restriction enzyme recognition sequence is asymmetric and several base pairs away from the cleavage site, the recognition sequence is inserted into the phagemide and insert so that it is cleaved after the initial digestion. If these fragments are subsequently assembled in the desired manner to insert a user-defined sequence into the phagemide, the recognition sequence is excised and no further cleavage occurs. However, if the fragment is ligated back to the recognition sequence to reassemble the initial input, the recognition sequence is recreated, and another round of cleavage occurs. In this way, the reaction proceeds in one direction until almost all restriction recognition sites have disappeared and almost all inserts have been correctly inserted into the phagemide backbone.
[0201] Notably, the pyrF nutritional requirement marker is included in the phagemid instead of conventional antibiotic markers to enhance the safety of the cssDNA produced from the system. Since one application of the cssDNA produced from this system is human therapeutic gene and cell therapy, including antibiotic marker sequences is undesirable. The exclusion of antibiotic marker sequences prevents the system from spreading antibiotic resistance to microorganisms (which could infect human patients with antibiotic-resistant strains).
[0202] The generated production host, bac105, was transformed with phagemide cdsDNA111 (RFP with an ampicillin resistance marker) and cdsDNA117 (RFP with a pyrF nutritional requirement marker). Control strains bac001 (DH5α) and bac016 (BW25113), which had not been generated, were transformed with phagemide cdsDNA111 and helper plasmid cdsDNA114 to express the M13 phage gene. The strains were grown in 200 μL of standard medium in a microplate containing appropriate antibiotics or nucleoside dropout for plasmid maintenance at 30°C for 48 hours. The generated strains possessing phagemides with nutritional requirement markers were grown in standard medium without uracil. Other strains were grown in standard medium containing uracil. Broth was collected, and the supernatant containing phage particles was clarified by centrifugation to remove E. coli. The supernatant was heated at 99°C for 10 minutes to lyse phage particles containing cssDNA, and this was used as a template for qPCR to quantify cssDNA production. A standard curve for cdsDNA111 containing the RFP sequence was used. qPCR primers targeted the RFP gene insert in the cssDNA. Parental strain bac016, containing only the helper plasmid cdsDNA114 and only the phagemide cdsDNA111, was included as a negative control.
[0203] The results are shown in Figure 15C. When transformed with RFP / ampicillin phagemide cdsDNA111, bac105 was superior to its parent strain bac016 and the conventional production host bac001 when transformed with the same phagemide (Figure 15C). Notably, bac105 was also superior to these two strains when transformed with RFP / pyrF phagemide cdsDNA117 (Figure 15C). A direct comparison of the developed strain bac105 with the conventional production host bac001 (DH5α), which has not produced a superior result, was not performed because this strain does not require uracil and there is no way to select a pyrF marker.
[0204] conclusion This high level of cssDNA production from strains created using phagemide templates with nutritional requirement markers establishes bac105 as the safest available cssDNA-producing host. Because the M13 replication origin and packaging signal reside on phagemides lacking the genes necessary for their own replication, this produces more cssDNA than conventional helper plasmids or helper-phage systems without generating infectious phage particles. Furthermore, the fact that the M13 gene necessary for phage replication is stably integrated into two separate loci in the E. coli genome significantly reduces the likelihood of infectious phages evolving from this production system. This means that two separate recombination events are required for the two sets of M13 genes to recombine with each other and with the phagemide, recombining all M13 genes with the M13 replication origin and packaging signal. Furthermore, the ability of this strain to produce cssDNA from phagemid templates with nutrient requirement markers means that the entire strain does not contain any antibiotic markers during production, and there is no need to add antibiotics to the production medium. This reduces the likelihood of antibiotic-resistant bacteria evolving, either during the cssDNA production process or after the cssDNA has been administered to a patient. Finally, the resulting production host grows well in standardized media, thus avoiding the need for expensive, nutrient-rich media that can produce inconsistent results with each run. [Table 5-1] [Table 5-2] [Table 5-3] [Table 5-4] [Table 5-5] [Table 5-6] This disclosure is not intended to be limited to the scope of specific disclosed embodiments provided, for example, to illustrate various aspects of this disclosure. Various modifications to the compositions and methods described will become apparent from the description and teaching herein. Such modifications can be implemented without departing from the true scope and spirit of this disclosure and are intended to be within the scope of this disclosure. [Table 6-1] [Table 6-2] [Table 6-3] [Table 6-4] [Table 6-5] [Table 6-6] [Table 6-7] [Table 6-8]
Claims
1. A production strain comprising a phagemide and at least two phage protein coding sequences incorporated into the genome of the production strain, wherein the production strain can produce cssDNA when the phagemide is introduced.
2. A circular single-stranded DNA (cssDNA) production system comprising at least two phage protein coding sequences incorporated into the genome of a production strain and a phagemide, wherein the production strain can produce cssDNA when the phagemide is introduced into the production strain.
3. A circular single-stranded DNA (cssDNA) production system that does not contain non-endogenous antibiotic resistance genes, Production stock; and A phagemide comprising a packaging signal, a designed sequence, a nutritional requirement marker, an antitoxin, RNA which inhibits the expression of a gene that delays or halts bacterial growth when expressed in the absence of the RNA, a transcription factor repressor which inhibits the expression of a gene that delays or halts bacterial growth when expressed in the absence of the transcription factor repressor, a transcription activator that activates the transcription repressor, a sequence that expresses tRNA that associates with a non-natural amino acid created as required by the production strain, and at least one selectable sequence encoding one or more combinations thereof. Production systems, including those mentioned above.
4. The production system according to claim 3, wherein the genome of the production strain includes at least two phage protein coding sequences incorporated into the genome of the production strain.
5. The production strain or production system according to any one of claims 1 to 4, wherein the production strain is susceptible to infection by single-stranded filamentous bacteriophages from the Monodnaviria region, and further, the production strain is a Gram-negative bacterium belonging to a family selected from the group consisting of Enterobacteriaceae, Pseudomonadaceae, Spirillaceae, Xanthomonadaceae, Clostridium, and Propionibacterium.
6. The production strain or production system according to claim 5, wherein the production strain is susceptible to infection by a single-strand filamentous bacteriophage selected from the group consisting of Ff, Fd, F1, and M13.
7. The production strain or production system according to claim 6, wherein the production strain is susceptible to infection by M13 bacteriophage.
8. The production strain or production system according to any one of claims 1 to 7, wherein the production strain is a strain of E. coli.
9. The production strain or production system according to any one of claims 1 to 2 or 4 to 8, wherein at least two phage protein coding sequences are expressed differentially by comparing them with each other.
10. The production strain or production system according to any one of claims 1 to 2 or 4 to 8, wherein the at least two phage protein coding sequences are expressed differentially compared to the expression of the natural phage genome.
11. The production strain or production system according to claim 9 or 10, wherein the at least two phage protein coding sequences code for at least two phage proteins selected from the group consisting of p1, p2, p8, p10, and p11.
12. The production strain or production system according to claim 11, wherein at least one of the at least two phage proteins is selected from the group consisting of p3 and p5.
13. The production strain or production system according to claim 12, wherein the p3 activity is reduced compared to the natural phage genome activity.
14. The production strain or production system according to claim 12, wherein the p5 activity is reduced compared to the natural phage genome activity.
15. The production strain or production system according to claim 1 or 2, wherein the at least two phage protein coding sequences are encoded by filamentous bacteriophages in the Monodnaviria region.
16. The production strain or production system according to claim 15, wherein the filamentous bacteriophage is selected from the group consisting of M13, Ff, Fd, enterobacteriaceae phage F1 [EF068134], enterobacteriaceae phage ID2, enterobacteriaceae phage NL95 [AF059243], enterobacteriaceae phage SP [X07489], enterobacteriaceae phage TW28, enterobacteriaceae phage Qbeta, enterobacteriaceae phage Qβ [AY099114], enterobacteriaceae phage M11 [AF059242], enterobacteriaceae phage ST, enterobacteriaceae phage TW18 [FJ483840], and enterobacteriaceae phage VK, or functional equivalents thereof.
17. The production strain or production system according to claim 15 or claim 16, wherein the at least two phage protein coding sequences include at least 3, 4, 5, 6, 7, 8, 9, 10, or 11 phage protein coding sequences.
18. The production strain or production system according to claim 1 or 2, wherein the at least two phage protein coding sequences include sequences that encode one or more bacteriophage M13 proteins selected from the group consisting of p1, p2, p3, p4, p5, p6, p7, p8, p9, p10, and p11.
19. The production strain or production system according to claim 18, wherein the sequence of the bacteriophage M13 protein comprises an M13 bacteriophage gene selected from the group consisting of I, II, III, IV, V, VI, VII, VIII, IX, X, and XI.
20. The production strain or production system according to claim 1 or claim 2, wherein the at least two phage protein coding sequences comprise at least three phage protein coding sequences.
21. The CSS DNA production system according to claim 3, wherein the production strain comprises at least two phage protein coding sequences incorporated in the genome.
22. The cssDNA production system according to claim 21, wherein the at least two phage protein coding sequences incorporated in the genome are selected from sequences derived from one or more filamentous bacteriophages in the Monodnaviria region.
23. The cssDNA production system according to claim 22, wherein one or more filamentous bacteriophages are selected from the group consisting of M13, Ff, Fd, enterobacteriaceae phage F1 [EF068134], enterobacteriaceae phage ID2, enterobacteriaceae phage NL95 [AF059243], enterobacteriaceae phage SP [X07489], enterobacteriaceae phage TW28, enterobacteriaceae phage Qbeta, enterobacteriaceae phage Qβ [AY099114], enterobacteriaceae phage M11 [AF059242], enterobacteriaceae phage ST, enterobacteriaceae phage TW18 [FJ483840], enterobacteriaceae phage VK, and functional equivalents thereof.
24. The cssDNA production system according to claim 3, wherein the selectable sequence is an antitoxin sequence derived from a toxin / antitoxin system.
25. The ccsDNA production system according to claim 24, wherein the toxin / antitoxin system is selected from the group consisting of ccdb / ccdA, hokaA / sokaA, pemK / pemI, mazF / mazE, ChpBK / ChpBI, relE / relB, parE / parD, hipA / hipB, and other toxin / antitoxin systems in which the toxin is expressed from the host genome and the antitoxin is expressed from the phagemide.
26. The cssDNA production system according to claim 3, wherein the selectable sequence is selected from the group consisting of HSVtk, Ura3, tetA, sacB, rpsL, pheS, pheS*, pheS**, thyA, lacY, gata-1, ccdB, hokA, pemK, mazF, chpBK, relE, parE, hipA, and other toxins, and produces at least one RNA molecule that downregulates the expression of a counter-selectable sequence.
27. The CSS DNA production system according to claim 3, wherein the selectable sequence encodes a transcription factor repressor.
28. The cssDNA production system according to claim 27, wherein the transcription factor repressor is selected from the group consisting of tetR, araC, lacI, xylS, and other sequences that reduce the expression of a counterselectable marker or toxin.
29. The CSS DNA production system according to claim 3, wherein the transcription activator is selected from the group consisting of araC and xylR.
30. The aforementioned nutritional requirement markers include uracil, adenine, cytosine, guanine, thymine, alanine, arginine, asparagine, aspartate, cysteine, glutamine, glutamate, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine, biotin, uridine-5'-monophosphate, pantothenate, xanthine, spermidine, para-aminobenzoate, lipoate, nicotinamide riboside, nicotinamide mononucleotide, D-glucosamine, and thia. The cssDNA production system according to claim 3, selected from the group consisting of mine, sikimate, aminoethyl-phosphonate, beta-alanine, s-methyl-methionine, ornithine, indole, indoleacetic acid, L-threonine, L-threonine O-3-phosphate, nicotinate, ribosylnicotinamide, pyrimidine, agmatine, purine, and genes for the synthesis of other essential compounds expressed from the phagemide to compensate for naturally occurring deficiencies or synthetic deletions of genes with the same or similar functions in the host genome.
31. The production strain according to claim 1, further comprising the phagemide, wherein the phagemide comprises a packaging signal, a designed sequence, and at least one selectable sequence, wherein the selectable sequence is optionally selected from the group consisting of a sequence encoding a nutritional requirement marker, an antitoxin, RNA which inhibits the expression of a gene that delays or halts bacterial growth when expressed in the absence of the RNA, a transcription factor repressor which inhibits the expression of a gene that delays or halts bacterial growth when expressed in the absence of a transcription factor, a transcription activator which activates such a repressor, or a sequence which expresses tRNA that associates with a non-natural amino acid as required by the production strain.
32. The aforementioned nutritional requirement markers include uracil, adenine, cytosine, guanine, thymine, alanine, arginine, asparagine, aspartate, cysteine, glutamine, glutamate, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine, biotin, uridine-5'-monophosphate, pantothenate, xanthine, spermidine, para-aminobenzoate, lipoate, nicotinamide riboside, nicotinamide mononucleotide, D-glucosamine, The production strain according to claim 31, selected from the group consisting of genes for the synthesis of thiamine, sikimate, aminoethyl-phosphonate, beta-alanine, s-methyl-methionine, ornithine, indole, indoleacetic acid, L-threonine, L-threonine O-3-phosphate, nicotinate, ribosylnicotinamide, pyrimidine, agmatine, purine, and other essential compounds expressed from the phagemide to compensate for naturally occurring deficiencies or synthetic deletions of genes with the same or similar functions in the host genome.
33. The production strain or production system according to claim 1 or 2, wherein the at least two phage protein coding sequences comprise at least two phage genes, and the at least two phage genes are operably linked to a synthetic promoter for optimizing expression for cssDNA production, comprising one or more standard T7 promoters and mutant T7 promoters under the control of inducible T7 polymerase, lacI, lacIq, araBAD, tet, temperature-sensitive promoter, stress-responsive promoter, quorum-sensing promoter, photosensitive promoter, and other inducible or repressive promoters.
34. The production strain or production system according to claim 1 or 2, wherein the at least two phage protein coding sequences are incorporated into the genome of the production strain at at least two distinct loci.
35. The production strain or production system according to claim 1 or 2, wherein at least one of the at least two phage protein coding sequences is modified from their endogenous sequence through random mutagenesis, rational design, assisted laboratory evolution, directional evolution, or a combination thereof, to create a protein that produces a higher yield or a higher purity yield of cssDNA.
36. The cssDNA production system according to claim 2, wherein the phagemide expresses tRNA necessary for translating re-encoded codons of non-natural amino acids, the production strain is unable to incorporate the non-natural amino acids into its protein in the absence of the phagemide, and neither the phagemide nor the production strain can be amplified or replicated in the absence of the non-natural amino acids.
37. The production strain or production system according to claim 1 or claim 2, wherein at least one of the at least two phage protein coding sequences further comprises a tag.
38. The production strain according to claim 37, wherein the tag is selected from affinity tags or detection tags.
39. The production strain according to claim 38, wherein the detection tag is selected from the group consisting of fluorescent tags, luminescent tags, color-developing tags, and other tags that enable rapid quantification of the number of phage particles in a solution.
40. The production strain according to claim 38, wherein the affinity tag is selected from the group consisting of biotin, his, myc, flag, CBP, GST, HA, HBH, MBP, S, V5, and another affinity tag that assists in the purification of phage particles from the production broth.
41. The production strain or production system according to any one of claims 1 to 40, wherein the phagemid includes a pUC replication origin or its derivatives.
42. The production strain or production system according to any one of claims 1 to 40, wherein the phagemide includes pST19, pDHA29, pDHA30, pDHK29, pDHK30, or a runaway R1 replication origin or a derivative thereof.
43. The production strain or production system according to any one of claims 1 to 42, wherein the phagemide comprises an inc1 mutation and / or an inc2 mutation, and optionally the phagemide comprises an inc1 mutation and an inc2 mutation.
44. The production strain or production system according to any one of claims 1 to 43, wherein the phagemid contains the inc3 mutation.
45. The production strain or production system according to any one of claims 1 to 44, wherein the phagemid contains the inc5 mutation.
46. A production strain or production system according to any one of claims 1 to 45, wherein the phagemide copy number in the production strain cells grown to the late logarithmic phase is at least 1,000, and optionally, the phagemide copy number in the production strain cells grown to the late logarithmic phase is at least 2,000, at least 4,000, at least 7,000, at least 8,000, or at least 15,000.
47. A circular single-stranded DNA (cssDNA) production system, Production stock; and A phagemide comprising a packaging signal, a designed sequence, and at least one selectable sequence. Includes, A production system in which the phagemide includes a pUC replication origin or its derivatives, the phagemide includes an inc1 mutation and / or an inc2 mutation, and optionally the phagemide includes an inc1 mutation and an inc2 mutation.
48. The production system according to claim 47, wherein the production strain comprises at least two phage protein coding sequences incorporated into the genome of the production strain.
49. A method for producing cssDNA, A step of culturing the production strain according to any one of claims 1, 5 to 20, 31 to 35, or 37 to 46 in a culture medium; Steps to introduce phagemids; and Steps to collect phage particles; and Step of separating the cssDNA Methods that include...
50. The method according to claim 49, further comprising the step of modifying at least one of at least two phage protein coding sequences.
51. The method according to claim 50, further comprising the step of comparing the collected cssDNA with cssDNA collected from the parent strain from which the production strain was produced, to determine whether the modified sequence improves the titer or quality of the produced cssDNA.
52. The method according to claim 49, wherein the production strain further comprises at least one phage protein including a tag.
53. The method according to claim 52, wherein the tag is used to isolate the phage.
54. A method for producing cssDNA, A step of culturing the cssDNA production system according to any one of claims 2 to 48 in a culture medium; Steps to collect phage particles; and Step of separating the cssDNA Methods that include...
55. The method according to claim 54, wherein the production strain further comprises at least one phage protein including a tag.
56. The method according to claim 55, wherein the tag is used to isolate the phage.
57. The production system according to claim 3 or the method according to claim 54, wherein the production strain comprises a helper plasmid derived from helper phage M13KO7 by removing the F1 origin and the packaging signal, and optionally the number of helper plasmid copies in the production strain cells grown to late logarithmic phase is at least 1,000, and optionally the number of helper plasmid copies in the production strain cells grown to late logarithmic phase is at least 2,000, at least 4,000, at least 7,000, at least 8,000, or at least 15,000.
58. The production strain, production system, or method according to claim 57, wherein the helper plasmid includes a pUC replication origin.
59. The production strain, production system, or method according to any one of the claims, wherein the production strain is a variant of the BW25113 strain.
60. The production strain, production system, or method according to any one of claims 33 to 59, wherein the nucleotide sequence encoding the T7 polymerase is incorporated into the genome of the production strain at the endA locus.
61. The production strain, production system, or method according to claim 60, wherein the nucleotide sequence encoding the T7 polymerase is under the control of an IPTG-inducible lac promoter.
62. The production strain, production system, or method according to any one of the claims, wherein the production strain comprises M13 genes II, V, VII, VIII, and IX, which are incorporated into the genome of the production strain at the intA locus.
63. The production strain, production system, or method according to any one of the claims, wherein the production strain comprises M13 genes III, VI, I, and IV integrated into the genome of the production strain at the intZ locus.
64. The production strain, production system, or method according to any one of the claims, further comprising the M13 gene II incorporated into the genome of the production strain at the mazF locus.
65. The production strain, production system, or method according to any one of the claims, further comprising the M13 gene VIII, which is incorporated into the genome of the production strain at the mazF locus.