Engineered cssDNA production hosts and bacteriophages
By using genetically engineered production strains and integrating phage coding sequences into phage particles and genomes, the problem of large-scale production of high-fidelity CSSDNA and SSDNA has been solved, achieving efficient and low-cost production while avoiding the metabolic costs and environmental risks associated with antibiotic resistance genes.
Patent Information
- Application Number
- CN202480034550.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-01
- Filing Date
- 2024-03-29
- Publication Date
- 2026-02-27
AI Technical Summary
Existing technologies make it difficult to mass-produce high-fidelity circular single-stranded DNA (cssDNA) and single-stranded DNA (ssDNA), especially due to the limitations of enzymatic or chemical methods and the metabolic costs and environmental safety issues of non-endogenous antibiotic resistance genes.
By employing production strains containing phage particles and genome-integrating phage coding sequences, and through genetic engineering, the likelihood of replicating phages is reduced, dependence on helper viral particles and helper plasmids is decreased, and selection sequences such as auxotrophic markers and transcriptional repressors are used to optimize phage protein expression and genome integration, avoid antibiotic resistance genes, and achieve efficient production of CSSDNA.
This technology enables efficient and low-cost production of CSSDNA that does not contain non-endogenous antibiotic resistance genes, reducing batch-to-batch inconsistencies, improving production efficiency and product purity, and reducing environmental risks.
Smart Images

Figure CN121586779A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 493,492, filed March 31, 2023, and U.S. Provisional Patent Application No. 63 / 605,225, filed December 1, 2023, the contents of each of which are incorporated herein by reference in their entirety. Background Technology
[0002] Customized sequences larger than 1 kb of high-fidelity cssDNA (circular single-stranded DNA) and ssDNA (single-stranded DNA) are difficult to manufacture on a large scale using enzymatic or chemical methods. Reference to electronic sequence listing
[0003] The contents of the electronic sequence list (304832000340SEQLIST.xml; size: 61,620 bytes; and creation date: March 28, 2024) are incorporated herein by reference in their entirety. Summary of the Invention
[0004] This article describes improvements in the development and production of ssDNA (single-stranded DNA) and cssDNA (circular single-stranded DNA).
[0005] This disclosure specifically describes a production strain comprising a phage particle and at least two phage coding sequences incorporated into the genome of the production strain. In some production strains, at least 3, 4, 5, 6, 7, 8, 9, 10, or 11 phage coding sequences may be included in the genome of the production strain. Incorporating coding sequences into the genome of the production strain provides improvements, including a reduced likelihood of generating replicative phages, and / or mitigation of batch-to-batch inconsistencies due to reduced dependence on helper viral particles and helper plasmids. Alternatively or additionally, this method can reduce the metabolic burden on the production host by eliminating the need to maintain extrachromosomal replication helper plasmids, which are traditionally maintained by adding antibiotics to the culture medium.
[0006] This document also describes a cssDNA production system that does not contain non-endogenous antibiotic resistance genes. In some embodiments, both the production strain and the phage contained within the production strain do not contain non-endogenous antibiotic resistance coding sequences. Those skilled in the art will understand that including antibiotic resistance coding sequences in recombinant DNA constructs facilitates many molecular biology techniques. However, this disclosure recognizes certain drawbacks of this approach. For example, such sequences may not be desirable for environmental and human safety reasons. Furthermore, there are metabolic costs associated with expressing antibiotic resistance sequences for the production host, and even when the strain is resistant to the drug, there are costs associated with exposure to antibiotic drugs. In some embodiments, the cssDNA production system provided in this disclosure does not utilize antibiotic resistance coding sequences in the production strain.
[0007] In some embodiments, the production strains provided herein include strains containing phage particles having a packaging signal, a designed sequence, and at least one optional sequence that is not an antibiotic resistance gene and is selected from sequences encoding: auxotrophic markers, antitoxins, RNA that inhibits the expression of a gene that, if expressed in the absence of RNA, would delay or prevent bacterial growth, a transcription factor repressor that inhibits the expression of a gene that, if expressed in the absence of transcription factors, would delay or prevent bacterial growth, a transcription activator that activates such a repressor, or a sequence expressing tRNA that associates with non-natural amino acids required for the engineering of the production strain. In many cases, the production strain itself undergoes genomic alterations to compensate for the activity of the selected sequence in the phage particle, making it impossible for the production strain to survive in the absence of the phage particle.
[0008] Suitable production strains include those readily infectable by single-stranded filamentous bacteriophages of the Monodnaviria domain, which are engineerable and support bacteriophage replication. Bacteria from the families Enterobacteriaceae, Pseudomonadaceae, Spirillaceae, Xanthomonadaceae, Clostridium, and Propionibacterium are potentially useful production strains. Those skilled in the art will understand that Escherichia coli strains are particularly suitable for use as production strains.
[0009] Exemplary phages that can be used as a source of phage proteins include M13, Ff, Fd, Enterobacter phage F1 [EF068134], Enterobacter phage ID2, Enterobacter phage NL95 [AF059243], Enterobacter phage SP [X07489], Enterobacter phage TW28, Enterobacter phage Qβ, Enterobacter phage Qβ [AY099114], Enterobacter phage M11 [AF059242], Enterobacter phage ST, Enterobacter phage TW18 [FJ483840], and Enterobacter phage VK, or their functional equivalents. Packaging signals and other control sequences from the phage can be used to prepare phage particles as described herein. In some embodiments, in addition to the packaging signals and the phage origin of replication (hereinafter referred to as ori), the genome of a particular phage can be integrated into the production strain. In such implementations, helper plasmids or helper phages are not required, and the phage genome will not be packaged into the phage capsid because the phage genome lacks the phage ori and the packaging signal.
[0010] In some embodiments, the coding sequence of a genome-integrated phage protein can be characterized by comparing its activity with the native activity when expressed by a control sequence found in wild-type phages. When the control sequence originates from a different source than the coding sequence (which is operatively linked to it in the expression cassette), the control sequence may be referred to as a non-natural control sequence, a non-endogenous control sequence, or an engineered control sequence. A control sequence operatively linked to the coding sequence in the organism from which the coding sequence originates may be referred to as a natural control sequence or an endogenous control sequence. Those skilled in the art will understand that activity can be altered by changing the amino acid sequence of a protein, such as by truncation or substitution of amino acid residues. Activity can also be altered by increasing the expression level of a protein using any means known in the art for overexpression. Genome-integrated phage protein sequences can be engineered to express differentially compared to each other and to their native expression. Differential expression includes not only increasing or decreasing protein activity but also temporarily expressing one phage protein at a different time than another phage protein to optimize CSSDNA production.
[0011] In some embodiments, the producing strain contains a genome-integrated phage coding sequence that encodes at least two of the following M13 phage proteins (or analogues thereof): p1, p2, p3, p4, p5, p6, p7, p8, p9, p10, and p11 (see Figure 1, insets A and B). In some embodiments, the genome-integrated sequence may be overexpressed compared to its natural expression, and / or the activity of one or more of the proteins may be altered. In certain embodiments, protein overexpression may be achieved using a promoter (such as a synthetic promoter, an inducible promoter system, or a promoter known to provide high levels of expression).
[0012] In some embodiments, the activity of one or more phage proteins may be reduced compared to their natural activity. For example, in some embodiments, the producing strain may contain p3 and / or p5 phage proteins with lower activity compared to their natural expression. In the exemplary embodiments described herein, the reduction in activity can be achieved by using a truncation of p3 or p5.
[0013] As further described herein, the selectable sequence may be selected from sequences that are in RNA form or translated into protein form and / or compensate for defects or sensitivities in the production strain. In some cases, for the selectable sequence in the phage particle to function, the production strain described herein needs to be genetically engineered to make the production strain defective or sensitive. In some embodiments, the genome-modified λ-red recombinant engineering system described in Example 1 below can be used to achieve the modification of the production strain as described herein.
[0014] In certain embodiments, when the selected sequence is a toxin / antitoxin system, the selected sequence may be selected from ccdB / ccdA, hokA / sokA, pemK / pemI, mazF / mazE, ChpBK / ChpBI, relE / relB, parE / parD, hipA / hipB, or other toxin / antitoxin systems, wherein the toxin is expressed by the host genome and the antitoxin is expressed by the phage particle. In some embodiments, when the selected sequence is an RNAi that downregulates a reverse-selectable sequence in the producing strain, the RNAi sequence may be selected to interfere with the translation of HSVtk, Ura3, tetA, sacB, rpsL, pheS, pheS*, pheS**, thyA, lacY, gata-1, ccdB, hokA, pemK, mazF, chpBK, relE, parE, hipA, or other reverse-selectable markers or toxins. Alternatively or additionally, in some embodiments, the selection sequence may include or be selected from transcriptional repressors, which can be used to downregulate the transcription of harmful sequences in the production strain. In some embodiments, the transcriptional repressor sequence may be selected from tetR, araC, lacI, xylS, or other sequences that reduce the expression of inverse selectable markers or toxins. Similarly, in some embodiments, the selection sequence may be selected from any transcriptional activator known in the art. In some embodiments, the transcriptional activator may be used to increase the transcriptional levels of genes required for the survival of the production strain. Examples of transcriptional activators include araC and xylR.
[0015] Particularly useful selectable sequences to be used in the methods and production systems described herein include auxotrophic sequences, which are sequences encoding proteins required for normal cellular function of the host of the production strain. In some embodiments utilizing such sequences, the host of the production strain is engineered to require the product of the selectable sequence for optimal growth. A particularly useful selectable sequence is a pyrF gene that can be used to support the growth of the production strain, which has an inactivated pyrF gene (pyrF-). In some embodiments, the production strain is cultured in a medium supplemented with uracil before transforming a phage carrying the pyrF selectable sequence into the pyrF- production strain. After transformation with the phage carrying the pyrF selectable sequence, the production strain is cultured in a medium without uracil supplementation. In some embodiments, the pyrF gene comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 98%, at least 99%, or at least 99.5% sequence identity with the sequence shown in SEQ ID NO: 34. In some embodiments, the pyrF gene contains the sequence of SEQ ID NO: 34. Alternatively or additionally, in some embodiments, the auxotrophic selectable sequence may be selected from genes that compensate for defects in one or more of the following biosynthetic processes in the producing strain: adenine, cytosine, guanine, thymine, alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine, biotin, uridine-5'-monophosphate, pantothenic acid, xanthine, spermidine, p-aminobenzoic acid, lipoic acid, nicotinamide nucleoside, nicotinamide mononucleotide, D-glucosamine, thiamine, shikimic acid, aminoethylphosphonic acid, β-alanine, S-methylmethionine, ornithine, indole, indoleacetic acid, L-threonine, L-threonine O-3-phosphate, nicotinic acid, ribosylnicotinamide, guanidine, etc., and / or selected from genes used to synthesize other essential compounds.
[0016] In some embodiments, the production strain includes or enables variable expression of phage genes (e.g., phage genes with expression levels different from those naturally expressed in their natural genome). Variable expression includes overexpression and downexpression compared to the expression level of the phage gene in its natural genome. In some embodiments, the gene expression ratios between individual phage genes are varied (e.g., different from the natural expression ratio of the natural phage gene in its natural genome).
[0017] Those skilled in the art will be familiar with a variety of promoters that can be used to alter the expression of a single phage gene or multiple phage genes. In some embodiments, a promoter that causes overexpression of the phage gene compared to when it is expressed by its natural promoter (e.g., when it is integrated and expressed by its natural promoter or when it is expressed by its natural promoter in its natural phage genome) can be selected. Conversely, in some embodiments, a promoter that causes reduced expression compared to the expression of the natural gene (e.g., when it is integrated and expressed by its natural promoter or when it is expressed by its natural promoter in its natural phage genome) can be selected. Alternatively or additionally, with respect to any of the foregoing, in some embodiments, a promoter can be selected to achieve a specific expression time, expression pattern, or expression responsiveness (e.g., to environmental inducements or applied inducements). In some embodiments, for a given production strain and phage particle combination, CSSDNA production can be improved (e.g., optimized) by using such promoters (e.g., heterologous promoters) and / or by integrating some or all of the relevant phage genes into the genome of the production strain. Promoters that are particularly useful in some implementations may include, for example, the classic T7 promoter or a mutant T7 promoter under the control of inducible T7 polymerase, the lacI promoter, the lacIq promoter, the araBAD promoter, the tet promoter, a temperature-sensitive promoter, a stress-responsive promoter, a quorum-sensing promoter, a photosensitizing promoter, or other inducible or repressive promoters.
[0018] In some embodiments, the production strain provided according to this disclosure may include a genome-integrated phage protein-coding sequence integrated at more than one locus within the production strain genome. The selection of loci within the production strain genome can affect the expression level of the phage coding sequence, and the loci can be selected to optimize the performance of the production strain. In some embodiments, the phage protein-coding sequence may be genome-integrated at at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 different loci within the genome.
[0019] Those skilled in the art will understand that, in addition to or as an alternative to manipulation of the control sequence controlling the expression of the phage protein, in some embodiments, variants of naturally occurring phage proteins may be utilized. In some embodiments, the phage protein variant comprises a sequence altered relative to its natural, naturally occurring sequence. In some embodiments, the phage protein variant has altered expression, activity, and / or sequence compared to the naturally occurring phage protein. Those skilled in the art will be aware of a variety of techniques for altering protein sequences (and / or the nucleic acid sequences encoding them), and will also understand that, in some embodiments, the activity of proteins known in the art can be used to produce a desired phage protein with the desired characteristics. Exemplary methods include random mutagenesis, rational design, assisted laboratory evolution, directed evolution, or other means. In some embodiments, the protein may be modified and then tested to determine whether its desired effect is achieved, for example by measuring CSSDNA yield, quality, and / or fidelity.
[0020] Alternatively or additionally, in some embodiments, the selection sequence may be or include a synthetic tRNA sequence that can (and in some embodiments, may be essential for recognition) recognize a re-encoded codon encoding a non-natural amino acid contained in at least one essential gene in the production strain as described herein. In such embodiments, the production strain will not grow in the absence of the non-natural amino acid and the synthetic tRNA sequence recognizing the re-encoded codon of the non-natural amino acid. For example, *E. coli* can undergo genome recoding such that the UAG codon becomes the dedicated codon for the non-natural amino acid L-4,4'-biphenylalanine (bipA, also known as the bipA chemical) if bipA aminoacyl-tRNA synthetase is expressed in the cell. (Mandell et al., Biocontainment of genetically modified organisms by synthetic protein design. Nature 518, 55-60 (2015)). The UAG codon can be inserted into three essential genes such that these genes cannot be translated if the bipA chemical or the bipA tRNA is absent. For example, the UAG codon can be inserted into the essential genes adenylate kinase (adk.d6), tyrosine-tRNA synthetase (tyrS.d8), and bipA-dependent aminoacyl-tRNA synthetase for the aminoacylation of bipA (BipARS.d6). Kunjapur et al., Synthetic auxotrophy remains stable after continuous evolution and in co-culture with mammalian cells. Sci Adv. 2021 July 2;7(27):eabf5851. In this case, bipA tRNA can be expressed by phage particles, and the bipA chemical can be added to the culture medium, resulting in a situation where only E. coli cells containing the phage particles can survive, ensuring that almost all E. coli cells in the population contain the phage particles, even if it is selected without the use of antibiotics. In addition, E. coli cells will not survive outside of specific cultures to which the bipA chemical has been added to the culture medium, because bipA is not present in the environment. This prevents engineered production hosts from escaping from the laboratory or fermentation facility environment. This also prevents production hosts from contaminating fermentation vessels that can be used for other strains, such as contract manufacturing organizations (CMOs), which is a problem given that the production host expresses phage particles that can theoretically infect other bacterial strains of CMOs.
[0021] In some embodiments, one or more phage proteins expressed by helper plasmids, helper viruses, or the genome of the production strain may additionally include a tag. This tag can be used to identify, quantify, and / or isolate the phage particles. Suitable tags include fluorescent tags, luminescent tags, chromogenic tags, and affinity tags. Particularly useful affinity tags include biotin, his, myc, flag, CBP, GST, HA, HBH, MBP, S, and V5, or other affinity tags used to aid in the purification of phage particles from the production broth.
[0022] This disclosure specifically provides certain methods for generating CSSDNA. In some embodiments, such methods include culturing a production strain containing at least two phage protein-coding sequences integrated into the genome of the production strain in a culture medium, and introducing phage particles into the production strain. Those skilled in the art will be familiar with a variety of techniques for introducing phage particles into the production strain. Exemplary such techniques include, for example, conjugation, transduction, electroporation, chemical transformation, or by utilizing a native phage infection mechanism. The resulting phage particles can be collected, and the CSSDNA can be collected. Exemplary separation methods include centrifugation, gel electrophoresis, membrane separation, tangential flow filtration, liquid chromatography, etc. In some embodiments, the means of separating CSSDNA include substantially separating any dsDNA from the CSSDNA. Those skilled in the art will know of a variety of techniques for quantifying CSSDNA, such as, for example, absorption spectroscopy, fluorescent intercalation dyes, UV detection, gel electrophoresis, microscopy, etc. In embodiments where the phage particles additionally include a tag, the tag can be used to assist in the separation and / or quantification steps.
[0023] Other methods for generating CSSDNA are also provided, such as those using a CSSDNA production system. In some embodiments, the provided CSSDNA production system uses a production strain that already contains a phage particle production sequence. Exemplary such methods include culturing the CSSDNA system in a culture medium and optionally inducing the production of one or more of the phage genes, collecting the phage particles, and isolating the phage particles. In some embodiments, the collection and isolation steps can be performed simultaneously, particularly when affinity tags are used to allow the phage particle shell to be separated from the CSSDNA in a single step. Attached Figure Description
[0024] Figures 1A-1B A schematic diagram of M13 phage particles depicting the organization of the coat proteins surrounding the ssDNA genome is shown. Figure 1A ). Figure 1B This image shows the genome map of M13 phage, identifying all genes, intergenic regions, promoters (uppercase P), and terminators (uppercase T). The coding region for each M13 protein is labeled "p#".
[0025] Figure 2 The phage plasmid contains an F1 origin of replication, an optional sequence, a user-defined sequence, and an optional plasmid origin of replication.
[0026] Figure 3 The study shows at least two phage coding sequences with genome integration, as well as a production strain containing phage replication origin, packaging signal, selection sequence, and design sequence.
[0027] Figure 4 The image shows a map of plasmid M13KO7 containing phage M13 genes I-XI, p15A origin of replication, M13 origin of replication and packaging signals, as well as kanamycin resistance markers.
[0028] Figure 5 The plasmid expressing the λ Red recombinant engineering system containing β, exo, and γ via the arabinose-inducible promoter is shown (catalog number CAS9BAC1P, Sigma Aldrich, Burlington, MA).
[0029] Figure 6 This diagram illustrates how the accumulation of additional CSSDNA templates within E. coli cells leads to increased production of packaged CSSDNA (Lee et al., Optimizing protein V untranslated region sequence in M13phage for increased production of single-stranded DNA for origami. NucleicAcids Res. 2021-06-21;49(11):6596-6603, which is incorporated herein by reference in its entirety).
[0030] Figure 7 shows the PCR primers designed to amplify the first transcription unit, altering the 5'UTR (untranslated region) of gene V from wild-type TCACA to GAGGT (Figure 7, inset A). Figure 7, inset B, shows a schematic diagram of the PCR primers designed to amplify the remaining portion of transcription unit 1 via fusion PCR, incorporating the mutated 5'UTR of gene V into transcription unit 1, thereby producing a 2117 bp variant of transcription unit 1 containing the altered gene V 5' UTR sequence.
[0031] Figures 8A-8B The map shows a phage particle pattern containing the pUC replication origin. Figure 8A ), as well as a portion of the pUC ori sequence and primers used to generate inc1 and inc2 mutations.
[0032] Figure 9 Results demonstrating the impact of phage plasmid replication origin on CSSDNA production are provided.
[0033] Figure 10 Results from large-scale culture experiments were provided, showing that inc1 and inc2 mutations increase cssDNA production.
[0034] Figure 11Results of qPCR experiments analyzing relative CSSDNA yields using helper plasmids with different origins of replication are provided.
[0035] Figure 12 A comparison of the final OD values of various potential bacterial production strains is shown.
[0036] Figure 13A This demonstrates the validation of cssDNA production in the unengineered BW25113 parent strain. Figure 13B The results demonstrate the induction of pLac:T7 polymerase expression in the engineered bacterial strain bac058, which contains an integrated pLac:T7 construct and an emGFP reporter.
[0037] Figure 14A The screening results of small combined libraries of integrated M13 TU1 and TU2 driven by promoters of different strengths are shown.
[0038] Figure 14B The results show the comparison between the selected engineered strain and its non-engineered parent strain.
[0039] Figure 15A The optimal engineered production host (bac105) is shown in comparison with its parent strain and the conventional production strain DH5α using a conventional auxiliary plasmid system.
[0040] Figure 15B The results show the generation of multiple CSSDNA sequences from engineered host strains.
[0041] Figure 15C The study shows a comparison of CSSDNA production by the engineered bac105 production strain with its parent strain bac016 and conventional production host bac001 when transformed with RFP / ampicillin phage cdsDNA111, as well as CSSDNA production by bac105 when transformed with RFP / pyrF phage cdsDNA117. Detailed Implementation
[0042] The production of any significant length of ssDNA is complex, partly because the success rate of adding the next base to the polymer decreases during chemical synthesis as the DNA molecule length increases. Similarly, using PCR-based techniques or producing longer ssDNA polymers via plasmids necessitates the removal of any regions of dsDNA, thus complicating the process and leading to metabolic waste. For example, in some cases, double-stranded DNA is prepared using typical plasmid amplification protocols at the metabolic cost to the host organism, followed by enzymatic removal of the unwanted strand. Given the inaccuracy of the subsequent enzymatic removal step, this post-assembly modification approach results in loss of cssDNA product and a wide range of ssDNA lengths. Synthesizing large amounts of dsDNA just to degrade one of the two strands is also metabolically wasteful, thus wasting half of the deoxyribonucleotide triphosphate (dNTP) molecules that have already been synthesized and polymerized.
[0043] Bacteriophages have been used to prepare ssDNA and cssDNA, and the production of these products has been reported. (Bush et al., Synthesis of DNA Origami Scaffolds: Current and Emerging Strategies, Molecules 2020, 25, 3386). However, engineered phages containing DNA not derived from natural phages (e.g., fluorescent protein tags) exhibit significantly reduced cssDNA yields compared to the production of natural phages. This opens up opportunities to improve the efficiency of such systems used for the preparation of ssDNA and cssDNA.
[0044] ssDNA and cssDNA are required for several known applications, and further applications are expected to be discovered as technologies related to DNA data storage and DNA nanotechnology (i.e., DNA origami) advance. For example, dsDNA from salmon eggs has been shown to be an effective transparent sunscreen (Gasperini et al., Non-ionising UV light increases the optical density of hygroscopic self-assembled DNA crystal films 2017 Nature Scientific Reports 7: 6631). However, its commercial use may be impractical given the costs associated with its manufacture. Combining synthetic biology engineering with efficient biotechnological production methods could be used to produce this DNA in a cost-effective manner. definition:
[0045] Unless otherwise stated, technical terms are used as commonly understood. Definitions of common terms in molecular biology can be found, for example, in Benjamin Lewin, Genes XII, Jones & Bartlett Learning; 12th edition (March 16, 2017) and other similar references. Unless the context clearly indicates otherwise, as used herein, the singular forms “a,” “an,” and “the” refer to both the singular and the plural. For example, the term “peptide” includes one or more antigens and can be considered equivalent to the phrase “at least one peptide.” As used herein, the term “comprising” means “including.” Thus, “comprising a protein” means “including a protein” without excluding other elements. It should also be understood that, unless otherwise stated, any and all base sizes or amino acid sizes and all molecular weights or molecular weight values given for nucleic acids or polypeptides are approximate and are provided for descriptive purposes. Although many methods and materials similar to or equivalent to those described herein may be used, specific suitable methods and materials are described herein. In the event of any inconsistency, this specification (including the explanation of terms) shall prevail. Furthermore, the materials, methods, and examples are illustrative only and are not intended to be limiting. For ease of reference to the various embodiments, the following explanations of terms are provided:
[0046] As used herein, the term "comparable" refers to two or more agents, entities, situations, condition groups, etc., that may be different from each other but are similar enough to allow comparisons between them, such that those skilled in the art will understand that reasonable conclusions can be drawn based on observed differences or similarities. In some embodiments, comparable condition groups, environments, individuals, or populations are characterized by a plurality of substantially identical features and one or a few different features. Those skilled in the art will understand that, in any given context, a certain degree of similarity is required between two or more such agents, entities, situations, or condition groups to be considered comparable. For example, those skilled in the art will understand that an environment group, individual, or population is comparable to each other when characterized with a sufficient number and type of substantially identical features to guarantee the reasonable conclusion that differences in results or observed phenomena under different environmental groups, individuals, or populations are caused by or indicate changes in these altered features. For example, in some embodiments, control sequences are used in engineered nucleic acid sequences to express coding sequences, and the control sequences used are different from those found in natural genes expressing proteins in microorganisms. Using control sequences that differ from those found in nature, compared to coding sequences, can be termed non-endogenous. In such cases, the activity of a natural gene expression cassette can be compared to an engineered form of a gene prepared using a non-endogenous control sequence. In such comparisons, when both expression cassettes are expressed under substantially the same experimental conditions, the non-endogenous control sequence can be identified as either active or inactive compared to the natural gene expression cassette.
[0047] A "control sequence" is a nucleic acid sequence that regulates the expression of a nucleic acid sequence operationally linked to a control sequence. The expression control sequence is operationally linked to the nucleic acid sequence when it controls and regulates transcription and translation (as appropriate). Therefore, the term "expression control sequence" can refer to elements such as promoters, enhancers, transcription terminators, ribosome binding sites, start codons (ATGs) preceding protein-coding genes, intron splicing signals, maintaining the correct reading frame of the gene to allow correct mRNA translation, stop codons, etc. The term "control sequence" is intended to include at least components whose presence can affect expression, and may also include additional components whose presence is advantageous, such as leader sequences and fusion coupler sequences. Expression control sequences may include promoters.
[0048] As used herein, the term "designed sequence" refers to a nucleic acid sequence engineered to be included in a template CSSDNA, template SSDNA, phage particle, or host cell genome. In some embodiments, the designed sequence is intended to be generated by the production strain described herein upon introduction into the production strain. Designed sequences may include genome-editing sequences, structural sequences such as those used in DNA origami, or any other desired nucleic acid sequence.
[0049] Generally, the terms "engineered" or "synthetic" refer to aspects that have been artificially manipulated. For example, a polynucleotide is considered "engineered" when two or more sequences that are not naturally linked together in that order are artificially manipulated to be directly linked to each other in an engineered polynucleotide and / or when a particular residue in the polynucleotide is not naturally occurring and / or is artificially linked to an entity or portion to which it is not naturally linked. For example, in some embodiments described and / or utilized herein, an engineered polynucleotide includes a regulatory sequence found in nature that is operatively associated with a first coding sequence but not with a second coding sequence, which is artificially linked to cause it to be operatively associated with the second coding sequence. Comparatively, a polypeptide can be considered "engineered" if it is encoded or expressed by an engineered polynucleotide, and / or if it is produced in cells differently from naturally expressed polypeptides. Similarly, a cell or organism is considered "engineered" if it has been manipulated to alter its genetic, epigenetic, and / or phenotypic identity relative to a suitable reference cell (such as a cell that is otherwise identical and has not been so manipulated). In some embodiments, the manipulation is or includes genetic manipulation, such that its genetic information is altered (e.g., the introduction of new genetic material not previously present, such as through transformation, mating, somatic cell hybridization, transfection, transduction, or other mechanisms, or the alteration or removal of previously present genetic material, such as through substitution or deletion mutations or through mating protocols). In some embodiments, engineered cells are cells that have been manipulated to contain and / or express specific target agents (e.g., proteins, nucleic acids, and / or specific forms thereof) relative to such a suitable reference cell in altered amounts and / or according to altered times. As is customary in the art and as understood by those skilled in the art, engineered polynucleotides or progeny of cells are often still referred to as “engineered” even if the actual manipulation is performed on a prior entity. Other examples of the use of the term “synthetic” include synthetic nucleases. Synthetic nucleases are nucleases that additionally contain domains directly or indirectly associated with co-localized sequences.
[0050] As used herein, a “promoter” refers to the minimal sequence sufficient to direct transcription. A promoter is also part of a set of nucleic acid sequences commonly referred to as control sequences. It also includes promoter elements sufficient to make promoter-dependent gene expression controllable for cell type specificity, tissue specificity, or induced by external signals or agents; such elements can be located in the 5' or 3' region of a gene. This includes constitutive and inducible promoters (see, for example, Bitter et al., Methods in Enzymology 153:516-544, 1987). For example, when cloning in bacterial systems, inducible promoters such as pL, plac, ptrp, and ptac (ptrp-lac heterozygous promoters) of bacteriophage λ can be used.
[0051] As used in this article, "cssDNA" refers to circular single-stranded deoxyribonucleic acid. As described in this article, cssDNA can be packaged into bacteriophages or bacteriophage particles.
[0052] "Loss-of-function genes / proteins / peptides" refer to genes / proteins / peptides that do not perform their natural biological functions or activities. Therefore, for example, a loss-of-function gene is a gene that cannot be transcribed and / or translated into a protein or peptide that performs the biological function or activity encoded by the natural gene. Gene inactivation can be achieved in a variety of ways, including but not limited to deletion, substitution, insertion, mutation, and the introduction of stop codons, frameshifting, truncation, etc. Loss-of-function proteins can also be described as having reduced activity or no activity.
[0053] As used herein, a “genome-editing sequence” is a specific type of designed sequence. A genome-editing sequence is designed to induce changes in the host cell genome. In some cases, a genome-editing sequence may contain one or more regions homologous to the host cell genome. In some cases, a genome-editing sequence may contain an entire gene expression cassette, such as a promoter, open reading frame, or terminator. In other embodiments, a genome-editing sequence may contain a gene or a portion of a gene, such as an intron, exon, etc. Genome-editing sequences in phage particles can be designed to alter the host cell genome, resulting in changes in the performance of endogenous nucleic acid sequences within the host cell. In some cases, genome-editing sequences can be designed to delete, downregulate, or reduce the activity of endogenous genomic products. In still other examples, genome-editing sequences can be designed to upregulate or increase the expression of endogenous genomic products. In still other examples, genome-editing sequences can be designed to repair unwanted sequences found in the host cell's endogenous genome, such as premature stop codons. Genome editing sequences can include, for example, cssDNA molecules, which, when incorporated into the host cell's genome, produce siRNA, ribozymes, antisense sequences, RNAi, genes, correction sequences for various genetic diseases, or beneficial genes or other sequences.
[0054] As mentioned in this article, “introduction” means the transfer of nucleic acid sequences or proteins into cells through, for example, phage infection (transduction), conjugation, or transformation using molecular biology techniques such as electroporation or heat shock of chemocompetent cells.
[0055] As used herein, "non-endogenous" refers to sequences, proteins, etc., that are not endogenous to a particular host cell (e.g., not naturally present in its relevant environment). For example, in some embodiments, a non-endogenous sequence is introduced into the genome of the production strain, and the non-endogenous sequence was not naturally present in the genome of the production strain prior to genetic engineering.
[0056] The terms “nucleic acid,” “polynucleotide,” and “oligonucleotide” are used interchangeably and refer to polymers of deoxyribonucleotides or ribonucleotides in linear or circular conformations and in single- or double-stranded form. For the purposes of this disclosure, these terms should not be construed as limiting the length of the polymer. These terms may cover known analogs of natural nucleotides, as well as nucleotides modified in their base, sugar, and / or phosphate moieties (e.g., phosphate thioester backbones, locked nucleic acids). Generally, and unless otherwise stated, analogs of a particular nucleotide have the same base-pairing specificity; that is, analogs of adenine will pair with thymine bases. When describing double-stranded DNA, the DNA may be described as A-DNA, B-DNA, or Z-DNA depending on the conformation adopted by the helical DNA. B-DNA, as described by James Watson and Francis Crick, is considered to be dominant in cells and extends approximately 34 Å per 10 bp sequence; A-DNA extends approximately 23 Å per 10 bp sequence; and Z-DNA extends approximately 38 Å per 10 bp sequence.
[0057] In some cases, the character representation recommended by the International Union of Pure and Applied Chemistry (IUPAC) or a subset thereof is used to provide nucleotide sequences. The IUPAC nucleotide codes used herein include: A = adenine, C = cytosine, G = guanine, T = thymine, U = uracil, R = A or G, Y = C or T, S = G or C, W = A or T, K = G or T, M = A or C, B = C or G or T, D = A or G or T, H = A or C or T, V = A or C or G, N = any base, "." or "-" = empty. In some embodiments, this set of characters represents adenosine, cytosine, guanosine, thymine, and uridine (A, C, G, T, U), respectively.
[0058] A nucleotide is a molecule containing a base moiety, a sugar moiety, and a phosphate moiety. Nucleotides can be linked together by their phosphate and sugar moieties, forming nucleoside linkages. The base moiety of a nucleotide can be adenine-9-yl (A), cytosine-1-yl (C), guanine-9-yl (G), uracil-1-yl (U), and thymine-1-yl (T). The sugar moiety of a nucleotide is either ribose or deoxyribose. The phosphate moiety of a nucleotide is pentavalent phosphate. Non-limiting examples of nucleotides are 3'-AMP (3'-adenosine monophosphate) or 5'-GMP (5'-guanosine monophosphate). Many variations of these types of molecules exist that are available in the art and herein.
[0059] "Oligonucleotides" or "polynucleotides" refer to synthetic or isolated nucleic acid polymers that contain multiple nucleotide subunits.
[0060] As used herein, a "phage particle" includes a phage origin of replication, a selection sequence, a design sequence, and a packaging signal. In some embodiments, the phage particle may also contain a plasmid origin of replication and / or an antibiotic resistance marker. The sequence of the phage particle will not contain phage protein-coding sequences, such that cells infected with the phage particle will not produce additional phage particles unless the cell additionally contains coding sequences for the necessary phage proteins.
[0061] As used in this article, "phage particle" refers to a CSSDNA sequence coated with phage protein but not containing the coding sequence necessary to produce phage protein.
[0062] As used herein, "phage protein" refers to a protein encoded by a native bacteriophage required for phage replication in the host. See, for example, Figure 1A and Figure 1B The diagrams illustrate the M13 phage protein and the M13 genome structure. In some embodiments, the phage protein includes a phage protein with nucleic acid sequence modifications compared to the native sequence. Such modifications include, for example, altering the nucleic acid sequence encoding the phage protein to optimize it based on codon usage for a specific production host. In some embodiments, such modifications result in altered activity of the phage protein. The nucleic acid sequence encoding the phage protein can also be altered to produce a phage protein with substituted amino acids compared to the native phage protein. In some cases, the phage protein may contain at least 1, 2, 3, 4, 5, 10, 15, 20, 25, or 30 amino acids different from the native phage sequence. Additional modifications that can be included in the phage protein include the addition of amino acids, such as tags, like fluorescent markers. The phage proteins described herein (e.g., native or modified phage proteins) will continue to function in the production strain. For example, when tagged, a tagged phage protein may have a looser binding to other phage proteins, but it is still considered a phage protein. The phage proteins specifically described in this article include the M13 phage protein and its nucleic acid sequence provided in the sequence listing. Proteins are designated as p1, p2, p3, p4, p5, p6, p7, p8, p9, p10, and p11, and genes are represented using corresponding Roman numerals (I, II, III, IV, V, VI, VII, VIII, IX, X, and XI). If a phage protein is modified, the modified phage protein can be characterized by comparison with the native phage protein.
[0063] The terms “polypeptide,” “peptide,” and “protein” are used interchangeably to refer to polymers of amino acid residues. The terms also apply to amino acid polymers in which one or more amino acids are chemical analogs or modified derivatives of the corresponding naturally occurring amino acids.
[0064] As used herein, "production strain" refers to a bacterial cell used to produce ssDNA (e.g., cssDNA). In some embodiments, the production strain comprises at least two phage proteins integrated into the genome of the production strain. In some embodiments, the at least two phage proteins may be expressed differentially compared to the expression levels that would be produced if an endogenous control sequence from a native phage were used. For example, one or more of the at least two phage protein sequences may be under the control of an inducible promoter. In some embodiments, the production strain is a strain of *Escherichia coli*, such as the non-toxic *E. coli* described in U.S. Patent 8,303,964 (which is incorporated herein by reference), which is further engineered to reduce the presence of 3-deoxy-d-mannooct-2-ketogluconic acid (Kdo) cyst components.
[0065] As mentioned herein, a “selective sequence” or “optional sequence” is a sequence of nucleic acids or amino acids that, when included in a phage particle sequence, allows the phage particle to remain within the production strain population. Optional sequences include, for example, antibiotic resistance sequences, engineered tRNA sequences, auxotrophic markers, antitoxin genes, or sequences that alter transcription or translation. Examples of optional sequences include, but are not limited to, genes encoding proteins or other compounds that increase or decrease resistance or sensitivity to antibiotics (e.g., ampicillin resistance genes, kanamycin resistance genes, neomycin resistance genes, tetracycline resistance genes, chloramphenicol resistance genes, and spectinomycin resistance genes). Further examples of optional sequences include, but are not limited to, genes encoding proteins that enable cells to grow in media lacking other essential nutrients, such as uracil, leucine, adenine, histidine, arginine, lysine, tryptophan, methionine, or other essential metabolites. In some examples, the production host is engineered to reduce or eliminate the production of essential nutrients.
[0066] As used herein, “template cssDNA” refers to circular single-stranded DNA used to generate cssDNA. In some embodiments, the template cssDNA comprises a packaging sequence, a design sequence, and a phage ori, which, after being introduced into a bacterial cell (such as an E. coli cell), is replicated into a replicative form using bacterial cell proteins and nucleic acid sequences, and then packaged into phage particles.
[0067] As used herein, the term "variant" refers to an entity that exhibits significant structural identity with a reference entity but differs structurally from the reference entity in the presence or level of one or more chemical components. In many embodiments, the variant also differs functionally from its reference entity. Generally, whether a particular entity is properly considered a "variant" of a reference entity is based on the degree of its structural identity with the reference entity. As those skilled in the art will understand, any biological or chemical reference entity has certain characteristic structural elements. By definition, a variant is a different chemical entity that shares one or more such characteristic structural elements. To cite just a few examples, small molecules can have characteristic core structural elements (e.g., a macrocyclic core) and / or one or more characteristic side chain moieties, such that variants of the small molecule are variants that share the core structural elements and characteristic side chain moieties but differ in other side chain moieties and / or in the type of bonds present within the core (single vs. double, E vs. Z, etc.). Peptides can have characteristic sequence elements consisting of multiple amino acids that are positioned relative to each other in linear or three-dimensional space and / or contribute to a specific biological function. Nucleic acids can have characteristic sequence elements consisting of multiple nucleotide residues that are positioned relative to each other in linear or three-dimensional space. For example, variant peptides can differ from reference peptides due to one or more differences in the amino acid sequence and / or one or more differences in the chemical portions (e.g., carbohydrates, lipids, etc.) covalently attached to the peptide backbone. In some embodiments, variant peptides exhibit an overall sequence identity of at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 99% with the reference peptide. Alternatively or additionally, in some embodiments, the variant peptide does not share at least one characteristic sequence element with the reference peptide. In some embodiments, the reference peptide has one or more biological activities. In some embodiments, the variant peptide shares one or more biological activities of the reference peptide. In some embodiments, the variant peptide lacks one or more biological activities of the reference peptide. In some embodiments, the variant peptide exhibits a reduced level of one or more biological activities compared to the reference peptide. In many embodiments, the target peptide is considered a “variant” of the parent or reference peptide if it has the same amino acid sequence as the parent but with a small number of sequence changes at specific positions. Typically, less than 20%, 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, or 2% of residues are substituted in the variant compared to the parent. In some embodiments, the variant has 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 substituted residue compared to the parent. Typically, the variant has a very small number (e.g., less than 5, 4, 3, 2, or 1) of substituted functional residues (i.e., residues involved in a specific biological activity).Furthermore, compared to the parent, variants typically have no more than 5, 4, 3, 2, or 1 addition or deletion, and generally have no additions or deletions. Additionally, any additions or deletions are typically fewer than about 25, about 20, about 19, about 18, about 17, about 16, about 15, about 14, about 13, about 10, about 9, about 8, about 7, or about 6 residues, and generally fewer than about 5, about 4, about 3, or about 2 residues. In some embodiments, the parent or reference polypeptide is a polypeptide found in nature. As those skilled in the art will understand, various variants of a particular target polypeptide are typically found in nature. Preparation of CSSDNA
[0068] This disclosure specifically provides methods for preparing cssDNA. Those skilled in the art will understand upon reading this disclosure that the designed sequence contained in a phage particle can be of any length necessary to achieve the desired function of the designed sequence. For example, if the desired sequence is a three-dimensional DNA structure (i.e., DNA origami), the length will be chosen to fit the size of the desired final product. Similarly, if the designed sequence is ultimately used for cell or gene therapy applications, such as CAR-T cell therapy, the sequence can be the length of the gene to be inserted plus the length of the homologous arms necessary to target the gene to a specific locus. Specific uses of cssDNA include its use in combination with genome editing technologies such as homologous recombination and CRISPR / Cas-related technologies. Considering the diverse applications of CSSDNA, the anticipated design sequence range is approximately 10 bp to approximately 10,000 bp, approximately 100 bp to approximately 10,000 bp, approximately 200 bp to approximately 9,000 bp, approximately 300 bp to approximately 9,000 bp, approximately 500 bp to approximately 9,000 bp, approximately 500 bp to approximately 8,000 bp, approximately 1,000 bp to approximately 10,000 bp, approximately 5,000 bp to approximately 15,000 bp, and approximately 10,000 bp to approximately 30,000 bp. In some embodiments, the CSSDNA generated by the methods described herein is consistent, for example, greater than 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% of the resulting molecules are mutation-free relative to the template. In some embodiments, the molecule length may also be greater than 10,000, 20,000, or 30,000 bp.
[0069] In some implementations, CSSDNA production occurs when a dsDNA phage particle encoding a phage origin of replication and packaging signal but lacking phage genes is replicated into CSSDNA by a phage protein trans-expressed in the host of the producing strain, and then packaged into a phage particle by the phage protein. When produced as described above, the phage particle does not contain any coding sequence for any phage protein, making it unable to self-replicate when infecting another bacterial cell. This differs from wild-type phages, which encode genes necessary for their own replication in their genome and can self-replicate and produce new self-replicating phage particles that can infect other bacterial cells and generate more self-replicating phages. The phage particles produced herein may contain both a phage origin of replication and a plasmid origin of replication. However, in some cases, the phage particle contains only a phage origin of replication and lacks a plasmid origin of replication.
[0070] In some embodiments, the phage particles described herein contain one or more optional sequences. The function of the optional sequences is to bind the phage particle to the production strain such that the production strain will not proliferate in the absence of a phage particle encoding the optional sequence. Optional sequences include sequences encoding: antibiotic resistance proteins, auxotrophic amino acids; compensatory tRNAs engineered into the production strain to compensate for deficiencies; RNA sequences that alter transcription or translation; or antitoxin amino acids (where the production strain produces toxins). Those skilled in the art will understand that there are many specific examples of such amino acid and nucleic acid sequences that can be used to establish a relationship between the phage particle and the production strain such that the production strain cannot continue to function fully in the absence of the phage particle. In preferred embodiments, sequences encoding genes or proteins that may have a detrimental effect on the final product from which the cssDNA will be used are avoided. For example, in the case of cssDNA used for gene therapy or microbial engineering of microorganisms to be released into the environment, antibiotic resistance sequences and toxin expression sequences are undesirable in the cssDNA.
[0071] Those skilled in the art will understand that several methods exist for preparing phagemids that do not contain a plasmid dsDNA origin of replication. This can be useful if the plasmid origin of replication is not desired to be present in the resulting final cssDNA. For example, the desired phagemid sequence can be cloned into a plasmid having additional control sequences (such as a T7 promoter sequence), design sequences, selection sequences, and packaging signal sequences. An in vitro transcription reaction can be performed, and the resulting RNA can then be reverse transcribed into the desired ssDNA sequence. The resulting ssDNA can be circularized to cssDNA using any method known in the art; for example, the splint (a short DNA with regions complementary to both ends of the ssDNA) can be annealed, and the reactants can be ligated to anneal the ends to form cssDNA. (Iyer et al., Efficient Homology-directed Repair with CircularssDNA Donors, CRISPR J, Oct;5(5):685-701, 2022). ssDNA ligases can also be used to convert linear ssDNA into cssDNA.
[0072] Alternatively, the desired components of ssDNA can be engineered using a combination of synthetic DNA synthesis and standard cloning and PCR techniques. The sequence can be generated in double-stranded form, and then one strand of the DNA can be selectively removed using a λ exonuclease.
[0073] Other techniques include using streptavidin-coated beads and biotin-conjugated PCR primers, one of which is biotinylated and the other is not. After PCR, one strand is biotinylated while the other is not. The unbiotinylated strand is degraded using an exonuclease. The resulting biotinylated ssDNA product is then bound to streptavidin-coated beads, and the ssDNA is recovered by physically separating the beads from the solution and then eluting the DNA from the beads by resuspending them in an elution buffer. Avci-Adali M, Paul A, Wilhelm N, Ziemer G, Wendel HP. Upgrading SELEX technology by using lambda exonuclease digestion for single-stranded DNA generation. Molecules. Dec 24, 2009;15(1):1-11.
[0074] The ssDNA, ultimately packaged into the phage particle, can also be cloned into the host cell genome as part of an expression cassette that allows for the direct production of RNA transcripts containing the desired phage particle components. Transcription can be induced at a desired time point, and post-transcriptional, the resulting transcript can be reverse transcribed, allowing the packaging signal to be accessible to the remaining phage genes required for packaging. The phage gene encoding the coat protein can be integrated into the genome or provided via a helper virus or helper plasmid. Those skilled in the art will understand that additional expression of the required reverse transcriptase may be necessary.
[0075] Phagemids can also be generated in the absence of a plasmid origin of replication using any standard in vitro dsDNA cloning method, including restriction / ligation, isothermal assembly, Golden Gate Assembly, or any other method for assembling DNA fragments. Once generated in vitro, the circular dsDNA can then be transformed into any bacterial strain expressing phage machinery necessary for replicating and packaging cssDNA from a dsDNA template containing a phage origin of replication and phage packaging signals.
[0076] When a phage contains both a phage origin of replication and a plasmid origin of replication, standard plasmid cloning techniques can be used to generate the desired template sequence for preparing a CSSDNA product. While including a plasmid origin of replication is convenient for many genetic engineering techniques, it also increases the likelihood that the plasmid origin of replication sequence will be included in the CSSDNA product, which may be undesirable in some cases.
[0077] In some embodiments, the phage is a high copy number phage. In some embodiments, the copy number of the phage in the production strain cells grown to late logarithmic phase is at least 1,000, optionally wherein the copy number of the phage in the production strain cells grown to late logarithmic phase is at least 2,000, at least 4,000, at least 7,000, at least 8,000, or at least 15,000. As shown in Example 6, in some embodiments, phages containing one or more copy number-increasing mutations increase CSSDNA yield. In some embodiments, the phage contains the inc1 mutation. In some embodiments, the phage contains the inc2 mutation. In some embodiments, the phage contains both the inc1 and inc2 mutations. In some embodiments, the phage contains pST19, pDHA29, pDHA30, pDHK29, pDHK30, or the runaway R1 origin of replication or a derivative thereof.
[0078] In some embodiments, the phage contains a pUC origin of replication (ori). In some embodiments, the phage ori contains a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 27. In some embodiments, the phage contains an inc1 mutant ori derived from the pUC ori. In some embodiments, the inc1 mutant ori contains a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 28. In some embodiments, the phage contains an inc1 mutant ori derived from the pUC ori. In some embodiments, the inc1 mutant ori contains a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 31. In some embodiments, the inc1 mutation is a C59T mutation relative to the standard pUC origin, wherein the nucleotide numbering is based on the sequence of the standard pUC origin shown in SEQ ID NO: 27. In some embodiments, the phage contains the inc2 mutation ori derived from pUC ori. In some embodiments, the inc2 mutation is a C92T mutation relative to the standard pUC initiation point, wherein the nucleotide numbering is based on the sequence of the standard pUC initiation point shown in SEQ ID NO: 27. In some embodiments, the inc2 mutation ori contains a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 29. In some embodiments, the inc2 mutation ori contains a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 32. In some embodiments, the phage contains inc1 and inc2 mutations ori derived from pUC ori. In some embodiments, inc1 and 2 ori (also referred to as inc1 / inc2 or inc1 / 2) contain C59T and C92T mutations relative to the standard pUC initiation point, wherein the nucleotide numbering is based on the sequence of the standard pUC initiation point shown in SEQ ID NO: 27. In some embodiments, the inc1 and inc2 mutants ori (inc1 and 2 ori) comprise a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 30. In some embodiments, the inc1 and inc2 mutants ori (inc1 and 2 ori) comprise a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 33. In some embodiments, the phage particle comprises the inc3 mutant.In some embodiments, the phage contains the inc5 mutation. In some embodiments, the phage copy number in the production strain cells grown to the late logarithmic phase is at least 1,000. In some embodiments, the phage copy number in the production strain cells grown to the late logarithmic phase is at least 2,000. In some embodiments, the phage copy number in the production strain cells grown to the late logarithmic phase is at least 4,000. In some embodiments, culturing a CSSDNA production strain or system containing phages containing inc1 and 2 ori results in at least a 10-fold higher yield compared to culturing a CSSDNA production strain or system containing phages containing wild-type pUCori. In some embodiments, after phage production, it can be introduced into the production strain using any method known in the art, including, for example, electroporation or by using chemicompetent cells (such as those prepared with divalent ions such as calcium chloride, magnesium chloride, or rubidium chloride solutions). Preparation of production strains
[0079] This disclosure illustrates production strains and their engineering methods. The described production strains are particularly suitable for producing cssDNA. Those skilled in the art will be familiar with a variety of techniques that can be used to produce such production strains. The production strains described herein include bacterial strains that, upon transformation with phage particles or nucleic acid sequences containing phage packaging signals, design sequences, and phage replication origins, can produce cssDNA phage particles. Bacteria may naturally contain the intracellular machinery necessary for producing cssDNA from template dsDNA, ssDNA, or cssDNA, or bacteria may be engineered to contain such machinery. As used herein, machinery refers to proteins necessary for producing cssDNA, rolling circle replication (sometimes referred to herein as ssDNA production), or ssDNA packaging (sometimes referred to herein as phage particle packaging). The production strains described herein may be natural targets for infection by phages containing template cssDNA, or they may simply be engineered to have the correct supporting protein production to produce phage particles after the introduction of template cssDNA or dsDNA using standard transformation techniques. Exemplary bacteria that can be used to prepare production strains include Escherichia coli and other Gram-negative bacterial species capable of replicating the CSSDNA phage genome (e.g., including single-stranded DNA viral domains such as phages Ff, Fd, F1, and M13).
[0080] In some implementations, the production strain contains copies of native genes from a bacteriophage (such as M13 phage) integrated into one or more loci. For example, genes encoding phage proteins can be PCR-amplified from the genome of an M13KO7 helper phage and integrated into the genome of a production strain as described herein. As described herein, these production strains will be able to produce cssDNA upon introduction of template dsDNA, cssDNA, or ssDNA. Production strains containing at least one or more copies of genes encoding phage proteins may also contain additional copies of one or more phage protein genes. For example, phage genes I, II, III, IV, V, VI, VII, VIII, IX, X, and XI can be integrated into the genome using their endogenous control sequences as described herein, and additional copies of any one of genes I, II, III, IV, V, VI, VII, VIII, IX, X, and XI can be integrated into the genome using the gene's native control sequence or a heterologous control sequence (such as an inducible promoter) or expressed by a helper plasmid.
[0081] In some embodiments, this document provides a circular single-stranded DNA (cssDNA) production system comprising: a production strain; and phage particles, wherein the phage particles contain a packaging signal, a designed sequence, and at least one selectable sequence.
[0082] The phage plasmid contains a pUC origin of replication or a derivative thereof, and the phage plasmid contains an inc1 mutation and / or an inc2 mutation. In some embodiments, the phage plasmid contains both inc1 and inc2 mutations. In some embodiments, the production strain contains two or more phage genes selected from genes I, II, III, IV, V, VI, VII, VIII, IX, X, and XI that have been integrated into the production strain cell by the genome. In some embodiments, the phage genes are expressed by a helper plasmid (as a supplement to or replacement of the integrated phage genes). In some embodiments, the production strain contains a helper plasmid derived from helper phage M13KO7 by removing the F1 origin and packaging signal. In some embodiments, the copy number of the helper plasmid in the production strain cell grown to late logarithmic phase is at least 1,000. In some embodiments, the copy number of the helper plasmid in the production strain cell grown to late logarithmic phase is at least 2,000, at least 4,000, at least 7,000, at least 8,000, or at least 15,000. In some implementations, the helper plasmid contains the pUC replication origin.
[0083] In some embodiments, the helper plasmid contains the inc1 mutation. In some embodiments, the helper plasmid contains the inc2 mutation. In some embodiments, the helper plasmid contains both the inc1 and inc2 mutations. In some embodiments, the helper plasmid contains pST19, pDHA29, pDHA30, pDHK29, pDHK30, or the runaway R1 origin of replication or a derivative thereof.
[0084] In some embodiments, the helper plasmid contains a pUC origin of replication (ori). In some embodiments, the helper plasmid ori contains a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 27. In some embodiments, the helper plasmid contains an inc1 mutant ori derived from pUC ori. In some embodiments, the inc1 ori contains a C59T mutation relative to the standard pUC origin, wherein the nucleotide numbering is based on the sequence of the standard pUC origin shown in SEQ ID NO: 27. In some embodiments, the inc1 mutant ori contains a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 28. In some embodiments, the inc1 mutant ori contains a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 31. In some embodiments, the helper plasmid contains an inc2 mutant ori derived from pUC ori. In some embodiments, the inc2 ori contains a C92T mutation relative to the standard pUC initiation site, wherein the nucleotide numbering is based on the sequence of the standard pUC initiation site shown in SEQ ID NO: 27. In some embodiments, the inc2 mutant ori contains a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 29. In some embodiments, the inc2 mutant ori contains a sequence having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 32. In some embodiments, the helper plasmid contains inc1 and inc2 mutant ori derived from the pUC ori. In some embodiments, inc1 and 2 ori (also referred to as inc1 / inc2 or inc1 / 2) contain a C59T mutation and a C92T mutation relative to the standard pUC initiation site, wherein the nucleotide numbering is based on the sequence of the standard pUC initiation site shown in SEQ ID NO: 27. In some embodiments, the inc1 and inc2 mutations ori (inc1 and 2 ori) comprise sequences having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 30. In some embodiments, the inc1 and inc2 mutations ori (inc1 and 2 ori) comprise sequences having at least 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 33. In some embodiments, the helper plasmid comprises the inc3 mutation. In some embodiments, the helper plasmid comprises the inc5 mutation.In some embodiments, the copy number of the helper plasmid in the production strain cells grown to the late logarithmic phase is at least 1,000. In some embodiments, the copy number of the helper plasmid in the production strain cells grown to the late logarithmic phase is at least 2,000. In some embodiments, the copy number of the helper plasmid in the production strain cells grown to the late logarithmic phase is at least 4,000.
[0085] In some embodiments, the production strain described herein may also contain a truncated form of gene III, which reduces the production strain's resistance to subsequent phage particle infection. Truncation may include the deletion of amino acid residues 21-273, amino acid residues 93-121, amino acid residues 93-141, or the insertion of stop codons at codons 20, 25, 28, 30, etc., within the p3 gene. See U.S. Patent No. US8227242, which is incorporated herein by reference in its entirety.
[0086] In some production strains, the expression of the phage protein p5 is altered (e.g., reduced) to increase CSSDNA production. Reduced p5 protein production activity can be achieved using any method known in the art, including those already described. P5 activity can also be altered by mutations in the untranslated region of the p5 gene, as described in Lee et al., Optimizing protein Vuntranslated region sequence in M13 phage for increased production of single-stranded DNA for origami. Nucleic Acids Research, Vol. 49, No. 11, June 21, 2021 (see, for example, Figure 6 ).
[0087] Gene overexpression can be achieved using any method known in the art. Generally, overexpression refers to an increase in the amount of protein produced compared to the amount normally expressed when the protein is expressed in its native genomic environment. For this reason, overexpression is sometimes described in terms of a second protein being expressed. Specific examples of how overexpression can be achieved include inserting multiple copies of the gene or changing the promoter or other control sequences to increase expression. The production strain described herein may contain a copy of the M13 genome containing the M13 control sequence. For clarity, the term overexpression will be used to refer to the production of an amount of M13 protein greater than the amount expressed under the control of its native control sequence. In other words, if the production strain contains two copies of the M13 genome, all phage proteins will be considered overexpressed; however, if it contains only a single copy of the M13 genome and its native control sequence, no phage protein is considered overexpressed.
[0088] The production strains described in this article can also be described by the ratio of one protein to another. The production ratio of individual phage proteins relative to each other can improve the efficiency of CSSDNA product production. For example, expressing p2:p5 at a ratio of 1:1, 2:1, 3:1, or 10:1 during the stationary or logarithmic phase of growth can be beneficial for phage particle production.
[0089] It is believed that overexpression of p2 and p10 by using additional copies of the natural gene, coding sequences under constitutive promoter control, coding sequences under inducible promoter control, or some combination thereof, increases phage particle production. See Behler et al., Phage-free production of artificial ssDNA with Escherichia coli, Biotechnol Bioeng. Oct 2022;119(10):2878-2889.
[0090] Based on its total amino acid contribution, the p8 protein constitutes the majority of the phage coat. In fact, 2700 copies of the p8 protein are present in the wild-type M13 phage coat. In engineered phage coats containing a genome larger than the wild-type 7.2 kb, even more p8 protein will be incorporated, as coat length is proportional to the size of the packaged DNA. Therefore, overexpression of the p8 protein is particularly useful when the designed sequence becomes longer. It is believed that overexpression of p8 by using additional copies of the natural gene, coding sequences placed under constitutive promoter control, coding sequences placed under inducible promoter control, or some combination thereof, increases phage particle production. In a particular embodiment, the production strain contains overexpressed p8 protein and a designed sequence of at least 2 kb in length. In yet another embodiment, the production strain comprises an overexpressed p8 protein and a designed sequence of at least 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 15 kb, 20 kb, 25 kb, or 30 kb in length.
[0091] In some implementations, the p8 protein is overexpressed by the classical T7 promoter or a variant thereof. Overexpression of genes II, X, IV, I, XI, and especially VIII by the classical T7 promoter increases the production of phage particles and CSSDNA. Decreased expression of gene V also increases their production.
[0092] Synthetic promoters can be characterized by the amount of RNA transcripts they can produce. Promoters that produce more RNA transcripts are characterized as high, those that produce slightly less are described as medium, and those that produce relatively less or even less are described as low. Examples of T7 promoters belonging to these categories include: High (Taatacgactcactatagggcgaattgggtaccgggcccagaccacaacggtttccctctagaaataattttgtttaactttaagaaggagatatacat-SEQ ID NO: 1 (wild-type T7)); Medium (taatccgactcactatagggcgaattgggtaccgggcccagaccacaacggtttccctctagaaataattttgtttaactttaagaaggagatatacat-SEQ ID NO: 2 (see sequence 7834, Komura et al., High-throughput evaluation of T7 promoter variants using biased randomization and DNA barcoding, Plos, 2018)); Low (Gaatacgactcactatagcgcgaattgggtaccgggcccagaccacaacggtttccctctagaaataattttgtttaactttaagaaggagatatacat-SEQ ID NO: 3 (see sequence 7812, Komura et al., High-throughput evaluation of T7 promoter variants using biased randomization and DNA barcoding, Plos, 2018))
[0093] The production strain may contain at least two coding sequences that produce phage proteins that promote the production of CSSDNA upon introduction of phage particles. At least two phage proteins integrated into the genome of the production strain are selected to optimize the yield and / or fidelity of the resulting CSSDNA.
[0094] In some implementations, the producing strain contains the genome-integrated coding sequence of the p1 phage protein (SEQ ID NO: 4: MAVYFVTGKLGSGKTLVSVGKIQDKIVAGCKIATNLDLRLQNLPQVGRFAKTPRVLRIPDKPSISDLLAIGRGNDSYDENKNGLLVLDECGTWFNTRSWNDKERQPIIDWFLHARKLGWDIIFLVQDLSIVDKQARSALAEHVVYCRRLDRITLPFVGTLYSLITGSKMPLPKLH VGVVKYGDSQLSPTVERWLYTGKNLYNAYDTKQAFSSNYDSGVYSYLTPYLSHGRYFKPLNLGQKMKLTKIYLKKFSRVLCLAIGFASAFTYSYITQPKPEVKKVVSQTYDFDKFTIDSSQRLNLSYRYVFKDSKGKLINSDDLQKQGYSLTYIDLCTVSIKKGNSNEIVKCN). In some embodiments, the p1 coding sequence is controlled by an inducible promoter (such as a temperature-sensitive or chemisensitive promoter). In some embodiments, the second of the at least two coding sequences is selected from the following phage proteins: p2, p3, p4, p5, p6, p7, p8, p9, p10, and p11. In other embodiments, the at least two coding sequences are integrated at different loci in the genome. This reduces the likelihood of recombination and the formation of replicating phages.
[0095] In another embodiment, the p1 coding sequence and a second of at least two coding sequences selected from phage proteins p2, p3, p4, p5, p6, p7, p8, p9, p10, and p11 are placed under the control of the same promoter, such that they are expressed substantially simultaneously during production. In some embodiments, p1 and a second of at least two coding sequences selected from phage proteins p2, p3, p4, p5, p6, p7, p8, p9, p10, and p11 are integrated into the genome at the same locus, and a third phage protein coding sequence is integrated into the genome at a different locus. In some embodiments, helper plasmids or any other phage proteins necessary for CSSDNA production are used to help deliver the virus.
[0096] In some implementations, the producing strain contains the genome-integrated coding sequence of the p2 phage protein (SEQ ID NO: 5: MIDMLVLRLPFIDSLVCSRLSGNDLIAFVDLSKIATLSGMNLSARTVEYHIDGDLTVSGLSHPFESLPTHYSGIAFKIYEGSKNFYPCVEIKASPAKVLQGHNVFGTTDLALCSEALLLNFANSLPCLYDLLDVNATTISRIDATFSARAPNENIAKQVIDHLRNVSNGQTKSTRSQNWESTVTWNETSRHRTLVAYLKHVELQHQ IQQLSSKPSAKMTSYQKEQLKVLSNPDLLEFASGLVRFEARIKTRYLKSFGLPLNLFDAIRFASDYNSQGKDLIFDLWSFSFSELFKAFEGDSMNIYDDSAVL DAIQSKHFTITPSGKTSFAKASRYFGFYRRLVNEGYDSVALTMPRNSFWRYVSALVECGIPKSQLMNLSTCNNVVPLVRFINVDFSSQRPDWYNEPVLKIA). In some embodiments, the p2 coding sequence is controlled by an inducible promoter (such as a temperature-sensitive or chemisensitive promoter). In some embodiments, the second of the at least two coding sequences is selected from the following phage proteins: p1, p3, p4, p5, p6, p7, p8, p9, p10, and p11. In other embodiments, the at least two coding sequences are integrated at different loci in the genome. This reduces the likelihood of recombination and the formation of replicating phages.
[0097] In another embodiment, the p2 coding sequence and a second of at least two coding sequences selected from phage proteins p1, p3, p4, p5, p6, p7, p8, p9, p10, and p11 are placed under the control of the same promoter, such that they are expressed substantially simultaneously during production. In some embodiments, p2 and a second of at least two coding sequences selected from phage proteins p1, p3, p4, p5, p6, p7, p8, p9, p10, and p11 are integrated into the genome at the same locus, and a third phage protein coding sequence is integrated into the genome at a different locus. In some embodiments, helper plasmids or any other phage proteins necessary for CSSDNA production are used to help deliver the virus.
[0098] In some implementations, the producing strain contains the genome-integrated coding sequence of the p3 phage protein (SEQ ID NO: 06: MKKLLFAIPLVVPFYSHSAETVESCLAKPHTENSFTNVWKDDKTLDRYANYEGCLWNATGVVVCTGDETQCYGTWVPIGLAIPENEGGGSEGGGSEGGGSEGGGTKPPEYGDTPIPGYTYINPLDGTYPPGTEQNPANPNPSLEESQPLNTFMFQNNRFRNRQGALTVYTGTVTQGTDPVKTYYQYTPVSSKAMYDAYWNGKFRDCAFHSGFNEDPFVCEYQGQSSDLPQPPVNAGGGSGGGSGGGSEGGGSEGGGSEGGGSEGGGSEGGGSGGSGSGDFDYEKMANANKGAMTENADENALQSDAKGKLDSVATDYGAAIDGFIGDVSGLANGNGATGDFAGS). In some embodiments, the p3 coding sequence is controlled by an inductive promoter (such as a temperature-sensitive or chemosensitive promoter). In some embodiments, the second of the at least two coding sequences is selected from the following phage proteins: p1, p2, p4, p5, p6, p7, p8, p9, p10, and p11. In other embodiments, the at least two coding sequences are integrated at different loci in the genome. This reduces the likelihood of recombination and the formation of replicative phages.
[0099] In another embodiment, the p3 coding sequence and a second of at least two coding sequences selected from phage proteins p1, p2, p4, p5, p6, p7, p8, p9, p10, and p11 are placed under the control of the same promoter, such that they are expressed substantially simultaneously during production. In some embodiments, p3 and a second of at least two coding sequences selected from phage proteins p1, p2, p4, p5, p6, p7, p8, p9, p10, and p11 are integrated into the genome at the same locus, and a third phage protein coding sequence is integrated into the genome at a different locus. In some embodiments, helper plasmids or any other phage proteins necessary for CSSDNA production are used to help deliver the virus.
[0100] In some implementations, the producing strain contains the genome-integrated coding sequence of the p4 phage protein (SEQ ID NO: 7: MKLLNVINFVFLMFVSSSSFAQVIEMNNSPLRDFVTWYSKQSGESVIVSPDVKGTVTVYSSDVKPENLRNFFISVLRANNFDMVGSIPSIIQKYNPNNQDYIDELPS SDNQEYDDNSAPSGGFFVPQNDNVTQTFKINNVRAKDLIRVVELFVKSNTSKSSNVLSIDGSNLLVVSAPKDILDNLPQFLSTVDLPTDQILIEGLIFEVQQGDALD FSFAAGSQRGTVAGGVNTDRLTSVLSSAGGSFGIFNGDVLGLSVRALKTNSHSKILSVPRILTLSGQKGSISVGQNVPFITGRVTGESANVNNPFQTIERQNVGISM SVFPVAMAGGNIVLDITSKADSLSSSTQASDVITNQRSIATTVNLRDGQTLLLGGLTDYKNTSQDSGVPFLSKIPLIGLLFSSRSDSNEESTLYVLVKATIVRAL). In some embodiments, the p4 coding sequence is controlled by an inducible promoter (such as a temperature-sensitive or chemisensitive promoter). In some embodiments, the second of the at least two coding sequences is selected from the following phage proteins: p1, p2, p3, p5, p6, p7, p8, p9, p10, and p11. In other embodiments, the at least two coding sequences are integrated at different loci in the genome. This reduces the likelihood of recombination and the formation of replicating phages.
[0101] In another embodiment, the p4 coding sequence and the second of at least two coding sequences selected from phage proteins p1, p2, p3, p5, p6, p7, p8, p9, and p10 are placed under the control of the same promoter, such that they are expressed substantially simultaneously during production. In some embodiments, p4 and the second of at least two coding sequences selected from phage proteins p1, p2, p3, p5, p6, p7, p8, p9, p10, and p11 are integrated into the genome at the same locus, and the third phage protein coding sequence is integrated into the genome at a different locus. In some embodiments, helper plasmids or any other phage proteins necessary for CSSDNA production are used to help deliver the virus.
[0102] In some implementations, the producing strain contains the genome-integrated coding sequence of the p5 phage protein (SEQ ID NO: 08: MIKVEIKPSQAQFTTRSGVSRQGKPYSLNEQLCYVDLGNEYPVLV KITLDEGQPAYAPGLYTVHLSSFKVGQFGSLMIDRLRLVPAK). In some embodiments, the p5 coding sequence is controlled by an inducible promoter (such as a temperature-sensitive or chemisensitive promoter). In some embodiments, the second of the at least two coding sequences is selected from the following phage proteins: p1, p2, p3, p4, p6, p7, p8, p9, p10, and p11. In other embodiments, the at least two coding sequences are integrated at different loci in the genome. This reduces the likelihood of recombination and the formation of replicative phages.
[0103] In another embodiment, the p5 coding sequence and a second of at least two coding sequences selected from phage proteins p1, p2, p3, p4, p6, p7, p8, p9, p10, and p11 are placed under the control of the same promoter, such that they are expressed substantially simultaneously during production. In some embodiments, p5 and a second of at least two coding sequences selected from phage proteins p1, p2, p3, p4, p6, p7, p8, p9, p10, and p11 are integrated into the genome at the same locus, and a third phage protein coding sequence is integrated into the genome at a different locus. In some embodiments, helper plasmids or any other phage proteins necessary for CSSDNA production are used to help deliver the virus.
[0104] In some implementations, the producing strain contains the genome-integrated coding sequence of the p6 phage protein (SEQ ID NO: 9: MPVLLGIPLLLRFLGFLLVTLFGYLLTFLKKGFGKIAIAISLFLALIIGLNSILVGYLSDISAQLPSDFVQGVQLILPSNALPCFYVILSVKAAIFIFDVKQKIVSYLDWDK). In some embodiments, the p6 coding sequence is controlled by an inducible promoter (such as a temperature-sensitive or chemisensitive promoter). In some embodiments, the second of the at least two coding sequences is selected from the following phage proteins: p1, p2, p3, p4, p5, p7, p8, p9, p10, and p11. In other embodiments, the at least two coding sequences are integrated at different loci in the genome. This reduces the likelihood of recombination and the formation of replicating phages.
[0105] In another embodiment, the p6 coding sequence and the second of at least two coding sequences selected from phage proteins p1, p2, p3, p4, p5, p7, p8, p9, and p10 are placed under the control of the same promoter, such that they are expressed substantially simultaneously during production. In some embodiments, p6 and the second of at least two coding sequences selected from phage proteins p1, p2, p3, p4, p5, p7, p8, p9, p10, and p11 are integrated into the genome at the same locus, and the third phage protein coding sequence is integrated into the genome at a different locus. In some embodiments, helper plasmids or any other phage proteins necessary for CSSDNA production are used to help deliver the virus.
[0106] In some implementations, the producing strain contains the genome-integrated coding sequence of the p7 phage protein (SEQ ID NO: 10: MEQVADFDTIYQAMIQISVVLCFALGIIAGGQR). In some embodiments, the p7 coding sequence is controlled by an inducible promoter (such as a temperature-sensitive or chemisensitive promoter). In some embodiments, the second of the at least two coding sequences is selected from the following phage proteins: p1, p2, p3, p4, p5, p6, p8, p9, and p10. In other embodiments, the at least two coding sequences are integrated at different loci in the genome. This reduces the likelihood of recombination and the formation of replicating phages.
[0107] In another embodiment, the p7 coding sequence and a second of at least two coding sequences selected from phage proteins p1, p2, p3, p4, p5, p6, p8, p9, p10, and p11 are placed under the control of the same promoter, such that they are expressed substantially simultaneously during production. In some embodiments, p7 and a second of at least two coding sequences selected from phage proteins p1, p2, p3, p4, p5, p6, p8, p9, p10, and p11 are integrated into the genome at the same locus, and a third phage protein coding sequence is integrated into the genome at a different locus. In some embodiments, helper plasmids or any other phage proteins necessary for CSSDNA production are used to help deliver the virus.
[0108] In some implementations, the producing strain contains the genome-integrated coding sequence of the p8 phage protein (SEQ ID NO: 11: MKKSLVLKASVAVATLVPMLSFAAEGDDPAKAAFNSLQASATEYIGYAWAMVVVIVGATIGIKLFKKFTSKAS). In some embodiments, the p8 coding sequence is under the control of an inducible promoter (such as a temperature-sensitive or chemisensitive promoter). In some embodiments, the second of the at least two coding sequences is selected from the following phage proteins: p1, p2, p3, p4, p5, p6, p7, p9, p10, and p11. In other embodiments, the at least two coding sequences are integrated at different loci in the genome. This reduces the likelihood of recombination and the formation of replicative phages.
[0109] In another embodiment, the p8 coding sequence and a second of at least two coding sequences selected from phage proteins p1, p2, p3, p4, p5, p6, p7, p9, p10, and p11 are placed under the control of the same promoter, such that they are expressed substantially simultaneously during production. In some embodiments, p8 and a second of at least two coding sequences selected from phage proteins p1, p2, p3, p4, p5, p6, p7, p9, p10, and p11 are integrated into the genome at the same locus, and a third phage protein coding sequence is integrated into the genome at a different locus. In some embodiments, helper plasmids or any other phage proteins necessary for CSSDNA production are used to help deliver the virus.
[0110] In some implementations, the producing strain contains the genome-integrated coding sequence of the p9 phage protein (SEQ ID NO: 12: MSVLVYSFASFVLGWCLRSGITYFTRLMETSS). In some embodiments, the p9 coding sequence is controlled by an inducible promoter (such as a temperature-sensitive or chemisensitive promoter). In some embodiments, the second of the at least two coding sequences is selected from the following phage proteins: p1, p2, p3, p4, p5, p6, p7, p8, p10, and p11. In other embodiments, the at least two coding sequences are integrated at different loci in the genome. This reduces the likelihood of recombination and the formation of replicating phages.
[0111] In another embodiment, the p9 coding sequence and a second of at least two coding sequences selected from phage proteins p1, p2, p3, p4, p5, p6, p7, p8, p10, and p11 are placed under the control of the same promoter, such that they are expressed substantially simultaneously during production. In some embodiments, p9 and a second of at least two coding sequences selected from phage proteins p1, p2, p3, p4, p5, p6, p7, p8, p10, and p11 are integrated into the genome at the same locus, and a third phage protein coding sequence is integrated into the genome at a different locus. In some embodiments, helper plasmids or any other phage proteins necessary for CSSDNA production are used to help deliver the virus.
[0112] In some implementations, the producing strain contains the genome-integrated coding sequence of the p10 phage protein (SEQ ID NO: 13: MNIYDDSAVLDAIQSKHFTITPSGKTSFAKASRYFGFYRRLVNEGYDSVALTMPRNSFWRYVSALVECGIPKSQLMNLSTCNNVVPLVRFINVDFSSQRPDWYNEPVLKIA). In some embodiments, the p10 coding sequence is under the control of an inducible promoter (such as a temperature-sensitive or chemosensitive promoter). In some embodiments, the second of the at least two coding sequences is selected from the following phage proteins: p1, p2, p3, p4, p5, p6, p7, p8, p9, and p11. In other embodiments, the at least two coding sequences are integrated at different loci in the genome. This reduces the likelihood of recombination and the formation of replicating phages.
[0113] In another embodiment, the p10 coding sequence and a second of at least two coding sequences selected from phage proteins p1, p2, p3, p4, p5, p6, p7, p8, p9, and p11 are placed under the control of the same promoter, such that they are expressed substantially simultaneously during production. In some embodiments, p10 and a second of at least two coding sequences selected from phage proteins p1, p2, p3, p4, p5, p6, p7, p8, p9, and p11 are integrated into the genome at the same locus, and a third phage protein coding sequence is integrated into the genome at a different locus. In some embodiments, helper plasmids or any other phage proteins necessary for CSSDNA production are used to help deliver the virus.
[0114] In some implementations, the producing strain contains the genome-integrated coding sequence of the p11 phage protein (SEQ ID NO: 14: MKLTKIYLKKFSRVLCLAIGFASAFTYSYITQPKPEVKKVVSQTYDFDKFTIDSSQRLNLSYRYVFKDSKGKLINSDDLQKQGYSLTYIDLCTVSIKKGNSNEIVKCN). In some embodiments, the p11 coding sequence is under the control of an inducible promoter (such as a temperature-sensitive or chemisensitive promoter). In some embodiments, the second of the at least two coding sequences is selected from the following phage proteins: p1, p2, p3, p4, p5, p6, p7, p8, p9, and p10. In other embodiments, the at least two coding sequences are integrated at different loci in the genome. This reduces the likelihood of recombination and the formation of replicative phages.
[0115] In another embodiment, the p11 coding sequence and a second of at least two coding sequences selected from phage proteins p1, p2, p3, p4, p5, p6, p7, p8, p9, and p10 are placed under the control of the same promoter, such that they are expressed substantially simultaneously during production. In some embodiments, p11 and a second of at least two coding sequences selected from phage proteins p1, p2, p3, p4, p5, p6, p7, p8, p9, and p10 are integrated into the genome at the same locus, and a third phage protein coding sequence is integrated into the genome at a different locus. In some embodiments, helper plasmids or any other phage proteins necessary for CSSDNA production are used to help deliver the virus. Example 1: Integration of the M13 whole genome into the production strain
[0116] The M13 genome is organized into two separate transcriptional units. The first consists of genes II, X, V, VII, IX, and VIII, and is expressed by a strong promoter. The second consists of genes III, VI, I, XI, and IV, and is expressed by a weak promoter. Two of these genes overlap with the open reading frames of the other two genes (e.g., the gene for p10 is completely contained within the coding region of p2). PA, PB, and PH are strongly active promoters (see [link to documentation]). Figure 1B The other two weak promoters are labeled PZ and PW. The two terminators (T) are strong rho-independent terminators; the third weaker terminator (T(weak)) is rho-dependent.
[0117] Smeal et al., Simulation of the M13 life cycle I: Assembly of agenetically-structured deterministic chemical kinetic simulation, Virology, Volume 500, 2017, Pages 259-274.
[0118] In order to generate a stable cell line capable of producing M13 phage particles and extremely unlikely to recombine with phage particles to produce replicating phages, each transcriptional unit of an M13 phage containing protein-coding genes and lacking phage replication origin and phage packaging signals was cloned from the publicly available helper phage M13KO7 (www.neb.com / products / n0315-M13KO7-helper-phage) (SEQ ID NO: 24) and integrated into two loci in the genome of Escherichia coli MG1655.
[0119] The first transcription unit was cloned into a circular double-stranded plasmid along with the following: a homologous arm at the pyrF locus for integration into the *E. coli* MG1655 genome, a spectinomycin selectable marker allowing selection of the integrase, and a p15A origin of replication in *E. coli*. The selectable marker is flanked by a loxP site to allow the marker to be recycled after expression of the cre gene. The insert containing the homologous arm, the M13 protein-coding gene, and the selectable marker is flanked by restriction enzyme sites, allowing the circular double-stranded DNA to be linearized prior to transformation of *E. coli*. The plasmid was digested with restriction enzymes and then electrophoresed on a 1% agarose tris-acetic acid-ethylenediaminetetraacetic acid (TAE) gel to separate the desired integration cassette from the unwanted plasmid backbone containing the p15A *E. coli* replication origin. The bands containing the integration cassette were purified using the Qiagen Gel Extraction Kit (catalog number 28706, Qiagen, Hilden, Germany) according to the manufacturer's recommended protocol. Because agarose gel electrophoresis was used to select E. coli replication origins that do not contain the p15A plasmid replication site required for plasmid replication in E. coli, it can be assumed that any E. coli cells that survive antibiotic selection after transformation with the integration cassette have likely undergone genome integration events rather than transformation with the replication plasmid. Generate a host with recombination capabilities.
[0120] First, the publicly available recombinant engineered plasmid (catalog number CAS9BAC1P, Sigma Aldrich, Burlington, MA) expressing a λred recombinant engineered system containing β, exo, and γ via an arabinose-inducible promoter was transformed into the E. coli host strain MG1655 (see [link]). Figure 5 ).
[0121] The plasmid also expressed a kanamycin antibiotic marker to allow selection and maintenance within the bacterial cell population, and the temperature-sensitive pSC101 replication factor repA101ts to allow eviction when the expression of the λ-red recombinant engineering system was no longer needed. Transformants were selected by plating on lysogenic broth agar (LB) supplemented with 50 μg / ml kanamycin (Bertani, G. 1952. "Studies on Lysogenesis. I. The mode of phage liberation by lysogenic Escherichia coli." J. Bacteriology, 62:293-300). Colonies were picked and inoculated into 15 ml LB liquid medium supplemented with 50 μg / ml kanamycin to maintain the plasmid. The cultures were incubated overnight at 30ºC. One ml of overnight culture was mixed with one ml of 50% glycerol / 50% aqueous solution and stored at -80ºC to produce a permanent library of a new strain known as MG1655 recombinant engineered strain. Plasmid DNA was isolated from the remaining 14 ml of culture and sequenced using Oxford Nanopore technology to confirm the presence and sequence identity of the recombinant plasmid. Induced expression of recombinant engineered genes
[0122] After confirming that *E. coli* strain MG1655 had been transformed by the recombinant engineered plasmid, the recombinant engineered strain was streaked onto LB agar supplemented with 50 μg / ml kanamycin and incubated overnight at 30°C. Single colonies were then inoculated into LB Lennox (low-salt formulation) supplemented with 50 μg / ml kanamycin and incubated overnight at 30°C. The culture was then diluted to an OD of 0.01 in LB liquid medium supplemented with 50 μg / ml kanamycin. The culture was incubated at 30°C until an OD of 0.30 was reached. At this point, it was diluted 1:2 with preheated LB broth supplemented with 4% L-arabinose to produce a final concentration of 2% L-arabinose to induce the expression of the recombinant engineered genes β, exo, and γ. The culture was then incubated at 30°C for another hour to allow sufficient accumulation of the recombinant engineered proteins β, exo, and γ in the cells. Electroporation using phage genome integration box
[0123] One hour after recombinant gene expression, cells from 1 ml of culture were harvested by centrifugation at 7000 RCF for one minute and then washed three times in ice-cold deionized water to remove salt. The cells were then resuspended in 49 µl of ice-cold deionized water and 1 µl of gel-purified linearized integration cassette, which consisted of a 2 pyrF homologous arm, all protein-coding sequences from the first transcriptional unit of M13KO7, and optional markers flanking loxP sites.
[0124] Electroporation of the cell and DNA mixture was performed using a Bio-Rad gene pulser equipped with a 1 mm electroporator, 2.1 kV, 100 ohms resistance, and 25 μF capacitance. Immediately after the electroporation, cells were resuspended in 1 ml of LB Lennox medium and incubated overnight at 30°C. Following this overnight recovery period, cells were plated on LB agar supplemented with 50 μg / ml spectinomycin for selection targeting integration events. Identifying cells with the correct integration events
[0125] The plate will be incubated overnight to allow colonies to form. Colonies will be screened by PCR to identify correct colonies formed by the correct integration event. Three pairs of PCR primers will be used. The first pair will examine the upstream end of the integration. The forward primer will bind to the genomic region upstream of the sequence found in the upstream homologous arm and point inward toward the integration cassette. The reverse primer will bind to the M13 gene within the integration cassette and point outward toward the *E. coli* genomic region upstream of the desired integration event.
[0126] The second primer pair will examine the downstream end of the integration. The forward primer will bind to the M13 gene within the integration cassette and point outward toward the desired downstream region of the *E. coli* genome. The reverse primer will bind to the *E. coli* genome downstream of the downstream homologous arm sequence and point inward toward the integration cassette. If both the first and second primer pairs produce PCR products of the expected size, it can be assumed that some cells in the screened colonies have the desired integration event.
[0127] The third primer pair will consist of a forward (genomic) primer from the first pair and a reverse (genomic) primer from the second pair. If these primers produce an 8 kb PCR product containing the pyrF homologous arm, the M13KO7 protein-coding region, and the spectinomycin marker, it can be assumed that some cells in the colony have undergone the correct integration event. If these primers produce a 2 kb product containing only the pyrF homologous arm and the pyrF gene, it can be assumed that some cells in the colony retain the wild-type pyrF sequence.
[0128] Correct colonies will produce PCR products using all three primer pairs, but there is no evidence that colonies using the third primer pair produce smaller 2 kb PCR products. A subset of these correct colonies were inoculated into LB broth supplemented with 50 ug / ml spectinomycin and grown overnight. 1 ml of the overnight culture was mixed with 1 ml of 50% glycerol / 50% aqueous solution and stored at -80ºC to produce a permanent library of new strains. Genomic DNA was extracted from another 1 ml of the overnight culture using the Sigma GenElute Genomic Prep Kit (catalog number NA2120, Sigma Aldrich, Burlington, MA). The purified genomic DNA was sequenced by Azenta, Inc. (Chelmsford, MA) using standard Illumina next-generation sequencing instruments and protocols. The fastQ file generated by Azenta was aligned with the expected sequence at the integration locus and the wild-type sequence at the same locus using the built-in mapping algorithm within Geneious Prime (Copyright © 2005-2022 Biomatters Ltd.). The following samples will be considered correct: those with no reads aligned to the wild-type pyrF gene, and some of which have reads aligned to all protein-coding sequences of the first transcription unit of M13KO7 without any sequence differences. The glycerol stock solution associated with one of these correct samples will be selected for future steps, and the other samples will be discarded. Remove recombinant engineered plasmid
[0129] The resulting *E. coli* strain, correctly integrated with all protein-coding genes from the first transcription unit of M13 ko7 and lacking pyrF, was streaked onto LB agar supplemented with 50 μg / ml spectinomycin and incubated overnight at 37ºC. Notably, LB agar lacks kanamycin for maintaining the recombinant engineered plasmid. Single colonies were inoculated into LB broth supplemented with 50 μg / ml spectinomycin and incubated overnight at 42ºC to prevent the temperature-sensitive pSC101 replication factor from initiating plasmid replication. Dilutes of this culture were serially plated onto LB agar supplemented with 50 μg / ml spectinomycin to obtain single-cell colonies. These plates were incubated overnight at 37ºC, and the resulting colonies were simultaneously inoculated onto LB agar supplemented with 50 μg / ml spectinomycin and LB agar supplemented with both 50 μg / ml spectinomycin and 50 μg / ml kanamycin. Colonies that grow only on the first plate and not on the second plate are considered to lack recombinant engineered plasmids. Inoculate one of these plasmid-free colonies into LB broth supplemented with 50 ug / ml spectinomycin and incubate overnight at 37ºC. Store the glycerol stock solution at -80ºC. Antibiotic marker recycling
[0130] The plasmid-free overnight culture was diluted to an OD of 0.01 and incubated at 37°C until an OD of 0.30 was reached. 1 ml of culture was then harvested and washed three times in ice-cold deionized water by centrifugation at 7000 RCF for 1 minute. The cells were then resuspended in 49 μl of ice-cold deionized water and mixed with 1 μl of plasmid DNA encoding the cre gene, pSC101ts origin of replication, and kanamycin marker. The washed cell and DNA mixture was then electroporated as described above, resuspended in 1 ml of LB broth, and incubated at 37°C for 1 hour. The cells were then plated on LB broth supplemented with 50 μg / ml kanamycin to select cells already transformed with the cre expression plasmid. The plates were incubated overnight at 37°C. The resulting colonies were simultaneously seeded on LB broth supplemented with 50 μg / ml kanamycin and on LB broth supplemented with both 50 μg / ml kanamycin and 50 μg / ml spectinomycin. Colonies that grow on the previous plate but not on the next can be considered to have recycled the spectinomycin marker via cre-induced recombination at the loxP marker flanking the spectinomycin gene. This can be verified by colony PCR using primers that bind to the integration cassette outside the loxP recombination site and point inward toward the region where spectinomycin resides. A 1 kb amplicon will indicate that the spectinomycin gene is still in its proper place, while a shorter amplicon will indicate that it has been lost through recombination. The λ-red recombinase system and the cre / lox marker recycling system are described by Tuntufye and Goddeeris, FEMS Microbiol Lett 325 (2011) 140-147, which is incorporated herein by reference in its entirety. Remove the cre expression plasmid
[0131] A colony that had been confirmed to have lost the spectinomycin resistance gene by both spectinomycin-free growth on agar plates and by colony PCR was inoculated into antibiotic-free LB broth and incubated overnight at 42ºC to prevent the temperature-sensitive pSC101 replication factor from initiating the replication of the cre expression plasmid. Dilutes of this culture were serially plated onto antibiotic-free LB agar and incubated overnight at 37ºC to obtain single-cell colonies. Colonies were simultaneously spotted on both antibiotic-free and LB agar supplemented with 50 μg / ml kanamycin to examine for cre expression plasmid loss. Colonies that grew on the former but not on the latter were considered to have had the plasmid removed. One such colony was inoculated into antibiotic-free LB agar and incubated overnight at 37ºC. The resulting culture was mixed with 50% glycerol / 50% aqueous solution at a 1:1 ratio, labeled KTx00001, and stored at -80ºC. Integration of the second transcription unit
[0132] The steps described above were repeated to integrate the second transcription unit. The only changes were to the sequence of the transcription unit, the homologous arms, and the primers used to verify the integration. The second transcription unit was integrated into the neutral integration locus yhiN in *E. coli* MG1655, Bernhards et al., ACS Synth. Biol. 2022, 11, 1681-1685. This new strain was labeled KTx00002 and stored at -80°C. Testing new strains
[0133] The following was confirmed by electroporating phage particles into KTx00002 and then performing CSSDNA quantification: KTx00002 was used with phage particles containing selectable sequences in the form of F1 ori, packaging signals, and ampicillin resistance genes. Figure 2 ) can produce cssDNA after infection.
[0134] Phagemids were transformed into KTx00002 cells using a standard electroporation protocol. In short, cells were grown to an OD of 0.3, and then 1 ml of culture was harvested by centrifugation at 6000 RCF for 1 minute. The precipitate was washed three times in ice-cold deionized water, and the cells were resuspended in 49 μl of deionized water. 1 μl of phagemid DNA was added to the cells, and they were then mixed by flicking the tube. The cells and DNA were transferred to a 1 mm electroporator and electroporated in a BioRad gene pulser at 2.1 kV / cm, 100 ohms, and 25 uF. The cells were allowed to recover in SOC medium at 37ºC with shaking for 1 hour, then washed and plated on a uracil-free defined-component medium (Bacto CD Supreme Fermentation Production Medium (FPM), catalog number A49737-01, Thermo Fisher, Waltham, MA, USA). The plates were incubated overnight at 37ºC. A single colony was picked and inoculated into 4 ml of uracil-deficient FPM and grown overnight at 30ºC. In the morning, the optical density (OD) was measured and the culture was diluted to OD 0.05 in 750 ml of uracil-deficient FPM. The culture was then incubated at 30ºC for 8 hours.
[0135] The culture was centrifuged to remove intact cells, and phage particles were collected in the supernatant. The phage particles were purified from the fermentation medium by precipitation with polyethylene glycol and sodium chloride (3% by weight / volume each), followed by centrifugation at 5000 g. The phage particle precipitate was then lysed using a solution of 1% sodium dodecyl sulfate and 200 mM sodium hydroxide. The lysed phage particle fragments were precipitated by centrifugation at 12,000 g, retaining the supernatant containing cssDNA. Endotoxin lipopolysaccharides were removed by adding Triton X-100 detergent, mixing thoroughly, and then discarding the resulting detergent layer. The cssDNA was then precipitated by adding 100% ethanol at -20ºC to the solution, incubating on ice for 30 minutes, and then precipitating by centrifugation at 12,000 g. The cssDNA precipitate was then washed with 75% ethanol at 4ºC, incubated on ice for 20 minutes, and centrifuged again at 12,000 g to remove residual salts. The resulting cssDNA precipitate was dried to remove residual ethanol and resuspended in a tris-EDTA (TE) buffer or nuclease-free water.
[0136] CSSDNA was quantified using the Qubit commercial assay kit and instrument, following the manufacturer's recommended protocol. Nucleic acid purity was quantified using an Agilent bioanalyzer. Endotoxin levels were quantified using an Endosafe® nexgen-PTS™ handheld spectrophotometer and column assay (catalog number MCS150K, Charles River Laboratories, Wilmington, MA, USA), following the manufacturer's recommended protocol, which utilizes assays based on USP BET standards. <85> Determination of kinetic horseshoe crab amoeboid cell lysates and EP BET<2.6.14>criver.com / sites / default / files / resources / Endosafe-PTS-Regulatory-Requirements-USPBET85-EPBET2.6.14.pdf.
[0137] In contrast, the parental strain of KTx0002 (E. coli MG1655) was similarly transformed using the same phage and an auxiliary plasmid derived from the helper phage M13KO7: the F1 origin and packaging signals were removed by amplifying the kanamycin marker and all protein-coding genes from M13KO7 using primers designed to exclude the F1 replication origin and packaging signals, followed by in vitro recirculation of the amplified DNA using DNA ligase. It is believed that KTx00002 will produce cssDNA with similar yield and fidelity compared to the parental control. Example 2: Effects of induced gene expression on KTx00002
[0138] Strain KTx00002 contains integrated copies of two transcriptional units of M13KO7 and, when transformed with a phage plasmid, achieves similar CSSDNA production as its parent strain MG1655, transformed with an M13KO7 helper plasmid and the same phage plasmid. To achieve high CSSDNA production, certain genes from M13KO7 are overexpressed.
[0139] To overexpress the M13KO7 gene in *E. coli*, the T7 polymerase coding sequence was first integrated into the neutral integration locus yjhV in the *E. coli* genome. This was done in the same manner as the integration of the two transcriptional units of M13KO7. In short, the T7 polymerase gene was codon-optimized for expression in *E. coli* and synthesized via Genscript (28 Yongxi Road, Jiangning District, Nanjing, Jiangsu, China, 211100, China). It was cloned into a plasmid under the control of the lactose-inducible promoter pLac. Downstream of this pLac-T7 polymerase expression cassette is the pheA transcription terminator and a spectinomycin marker flanked by a loxP site. Both the expression cassette and the spectinomycin marker are flanked by homologous arms and restriction enzyme sites. The plasmid was extracted from its *E. coli* host using the QIAprep Spin Miniprep Kit (catalog number 27106, Qiagen, Hilden, Germany) and linearized using restriction enzymes to produce an integration cassette isolated from the bacterial origin of replication. KTx00002 cells were transformed with the recombinant engineered plasmid as described above and grown to an OD of 0.3, at which point L-arabinose was used to induce the expression of the recombinant engineered gene. The cells were then incubated for another hour to allow for the expression of the recombinant engineered gene. Cells were then harvested and washed three times by centrifugation, mixed with 1 μl of the linearized integration cassette, and then electroporated. Cells were allowed to recover overnight and then plated on LB agar supplemented with 50 μg / ml spectinomycin for selection against the integration cassette. Integration was verified by colony PCR and sequencing as described above. The recombinant engineered plasmid was removed as described above. Cells were transformed with the cre expression plasmid as described above, and the marker was recycled as described above. Finally, the cre expression plasmid was removed, and a new strain labeled KTx00003 was stored at -80°C.
[0140] Evidence suggests that overexpression of certain M13 phage genes leads to increased production of cssDNA in *E. coli*. (Behler et al., 2022 Oct;119(10):2878-2889). Open reading frames 2 and 10 are involved in phage gene expression, open reading frames 1 and 11 are involved in the formation of pores in the bacterial membrane for phage secretion, gene 4 is also involved in the formation of secretory pores, and gene 8 is a major capsid protein. However, this evidence does not stem from genomic integration of phage genes.
[0141] Overexpression of phage genes, both individually and in combination, was achieved by cloning coding sequences controlled by the T7 promoter into plasmids containing transcription terminators downstream of one or more genes, followed by spectinomycin-selective markers flanked by loxP sites. Both the phage gene expression cassette and the spectinomycin marker have homologous arms flanking them for integration into neutral integration loci in the *E. coli* genome. The homologous arms are flanked by restriction enzyme recognition sequences, allowing the vector to be linearized and then integrated into the *E. coli* genome.
[0142] These plasmids were prepared and digested as described above, and sequentially integrated into the *E. coli* KTx00003 genome using the recombinant engineering and cre / lox system described above. This yielded a new *E. coli* strain containing two wild-type transcription units of M13KO7, plus a T7 polymerase under the control of the lac promoter, and a new overexpression cassette containing a T7 promoter driving the expression of phage genes as shown in Table 1. The new strain is referred to by the strain name KTx00004.XY, where "X" is 1, indicating the overexpression cassette of p1, and where "X" is 2, indicating the overexpression cassette of p2, and so on for each of genes I-XI. Similarly, "Y" is used to identify the second overexpressed phage gene when two different phage genes are overexpressed. The table below presents various combinations of the identified XY. Table 1.
[0143] Overexpression of phage genes from the T7 promoter was validated by reverse transcription-quantitative polymerase chain reaction (rt-qPCR). Cultures of the new strain and parental control strain were grown overnight, then diluted to OD 0.01 and grown with and without 100 µM IPTG. Cells were harvested by centrifugation after 5 hours of growth. mRNA was prepared using the Qiagen RNeasy Mini kit (catalog number 74104, Qiagen, Germany) according to the manufacturer's recommended protocol. cDNA was reverse transcribed using the TaqMan™ reverse transcription reagent (catalog number N8080234, ThermoFisher Scientific, Waltham, MA) following the manufacturer's recommended protocol. The cDNA preparations were normalized to equal concentrations and amplified in multiplexed reactions using primers targeting each M13 gene and primers targeting the housekeeping control gene rpoD. TaqMan probes (ThermoFisher Scientific, Waltham, MA) were added to the reactions to quantify the amount of each PCR product produced. For the M13 gene, the probe contained a 5' fluorescein reporter dye and a 3' Iowa Black® FQ quencher. For the rpoD housekeeping control gene, the probe contained a 5' anthocyanin-5 reporter dye and a 3' Iowa Black® RQ quencher. Reactions were performed in a QuantStudio 7 Flex (Thermofisher Scientific, Waltham, MA). Overexpression was determined by normalizing cycles in which the fluorescein signal crossed an arbitrary threshold to cycles in which the anthocyanin-5 signal crossed the same threshold (Δcycle threshold or Δct), and then comparing with the wt parent and the uninduced control (Δcycle threshold or ΔΔct). Compared with strains expressing the phage gene only by its natural promoter, strains expressing the M13 gene by the T7 promoter crossed the arbitrary threshold (normalized to the anthocyanin-5 signal) after fewer cycles.
[0144] The CSSDNA from KTx00004 was compared with that from its parent strain KTx00002, which contained only two wild-type transcription units of M13KO7. For this comparison, each strain was transformed with the same phage particle. Cultures of strains with phage particles were grown in parallel, and CSSDNA was extracted and quantified. Example 3: Differential expression of bacteriophage genes in the production host
[0145] For experimental purposes, each phage gene was cloned into an integration cassette and placed under the control of high, medium, and low-intensity T7 promoters, as described above. Each gene was integrated in parallel into the *E. coli* KTx0003 genome using the aforementioned λ-red recombination engineering system to generate 33 new strains, each expressing one phage gene by a single T7 promoter. Genes examined using plasmids included phage genes I, II, III, IV, V, VI, VII, VIII, IX, X, and XI. The ability of the 33 new strains to produce CSSDNA was tested.
[0146] In this experiment, all genes except the test gene (non-test genes) were placed under the control of their natural promoters, which produced expression of non-test genes at levels substantially the same as those expressed by wild-type phages. A single test gene was induced by adding IPTG to activate the expression of T7 polymerase, which in turn expressed the test gene via the T7 promoter. IPTG was added at different time points during the growth of the production host. The amount of cssDNA produced at the time point of IPTG introduction was quantified. For clarity, the example included placing gene VIII under the control of the T7 promoter (SEQ ID NO:1) and the remaining genes under the control of their natural promoters. T7 polymerase expression was induced using IPTG (lactose or its analogues) at OD values of 0.05, 0.5, 1.0, and 3.0, followed by transcription of gene VIII by the T7 promoter. The yield and purity of cssDNA were determined using a Qubit-based bioanalyzer and commercially available assay kits and protocols designed to work with these instruments.
[0147] To quantify the CSSDNA produced by these test strains and the parental control strain KTx0003, the production host was transformed with the same phage particles and cultured in Bacto CD Supreme fermentation production medium (FPM) (catalog number A4973702, ThermoFisher Scientific, Waltham, MA, USA) until the specified OD was reached, using sterile medium as a blank control. At the specified OD, isopropyl β-d-1-thiogalactopyranoside (IPTG) was added, and the cultures were shaken at 200 RPM in a New Brunswick Innova shaking incubator. The cultures were allowed to continue growing until a total growth time of 8 hours was reached. The cultures were centrifuged to remove intact cells, and phage particles were collected from the supernatant. The phage particles were purified from the fermentation medium by precipitation with polyethylene glycol and sodium chloride (3% by weight / volume each), followed by centrifugation at 5000 g. The phage particle precipitate was then lysed using a solution of 1% sodium dodecyl sulfate and 200 mM sodium hydroxide. The lysed phage particles were precipitated by centrifugation at 12000 g, retaining the supernatant containing CSSDNA. The CSSDNA was then precipitated by adding 100% ethanol at -20ºC to the solution, incubating on ice for 30 minutes, and then centrifuging at 12000 g. The CSSDNA precipitate was then washed with 75% ethanol at 4ºC, incubated on ice for 20 minutes, and centrifuged again at 12000 g to remove residual salts. The resulting CSSDNA precipitate was dried to remove ethanol and resuspended in TE buffer or nuclease-free water.
[0148] CSSDNA was quantified using the Qubit commercial assay kit and instrument, following the manufacturer's recommended protocol. Nucleic acid purity was quantified using an Agilent bioanalyzer. Endotoxin levels were quantified using an Endosafe® nexgen-PTS™ handheld spectrophotometer and column assay (catalog number MCS150K, Charles River Laboratories, Wilmington, MA, USA), following the manufacturer's recommended protocol, which utilizes assays based on USP BET standards. <85> Determination of kinetic horseshoe crab amoeboid cell lysates using EP BET<2.6.14>www.criver.com / sites / default / files / resources / Endosafe-PTS-Regulatory-Requirements-USPBET85-EPBET2.6.14.pdf. Table 2 List of strains Table 3 Example 4: Preparation of host cells with reduced p5 activity
[0149] Further development of strain KTx00003 was undertaken to reduce the relative production of p5 (KTx00006). p5 is an ssDNA-binding protein. Altering five bases in the 5' UTR of gene V inhibited the interaction between gene V mRNA and ribosomes and reduced the amount of p5 protein translated by ribosomes.
[0150] The replicative dsDNA form of the phage particle serves as a template for generating CSSDNA via rolling circle amplification. After CSSDNA is generated within the bacterial host, two different scenarios occur. If p5 levels are low, p5 does not bind to CSSDNA, allowing the bacterial replication machinery to polymerize the complementary strand and generate more replicative dsDNA. This replicative dsDNA is then used as a further template for rolling circle amplification, resulting in the generation of even more CSSDNA. However, if p5 levels are high, p5 binds to CSSDNA and prevents the bacterial host machinery from polymerizing the complementary strand, thus preventing its use as a template for rolling circle amplification. The binding of p5 to CSSDNA instead causes the CSSDNA to be packaged into the phage capsid and exported outside the host cell. Reduced p5 translation allows the cell to generate more phage CSSDNA templates before transitioning to packaging CSSDNA into phage capsids and exporting the phage capsids outside the cell. The accumulation of additional CSSDNA templates within the *E. coli* cell leads to increased production of packaged CSSDNA. A schematic diagram of this strategy is shown below. Figure 6 (See Lee et al., Optimizing protein V untranslated regionsequence in M13 phage for increased production of single-stranded DNA fororigami. Nucleic Acids Res. 2021-06-21;49(11):6596-6603, which is incorporated herein by reference in its entirety.)
[0151] Before cloning the first transcription unit described in Example 1 above, PCR primers were designed to amplify the first transcription unit, so that the 5' UTR (untranslated region) of gene V was changed from wild-type TCACA to GAGGT (as shown in Figure 7, small figure A).
[0152] The PCR primers were also designed to amplify the remaining portion of transcription unit 1 as described in Example 1 by fusion PCR, so that the mutated 5' UTR of gene V is incorporated into transcription unit 1, thereby producing a 2117 bp variant of transcription unit 1 containing the altered gene V 5' UTR sequence (as shown in Figure 7, inset B).
[0153] The same *E. coli* production host strains were generated using two alternative forms of transcription unit 1; one containing a mutation in the 5' UTR, and the other lacking said mutation. Both strains were transformed with the same phage particle and cultured in parallel, allowing for a comparison of CSSDNA production. Without being bound by any theory, it is believed that the KTx00006 strain, with the mutation in the 5' UTR of gene V, will produce more CSSDNA than the control strain (KTx00003) lacking said mutation. The CSSDNA produced by the strains can be analyzed as previously described. Example 5: Generation of circular single-stranded DNA phage particles with plasmid ori
[0154] Minimal phage matrix backbones were cloned from the M13KO7 template by amplifying the M13 origin of replication, including the packaging signal, using primers with 30 nt homologous tails to other fragments. The pyrF gene was amplified from the *E. coli* MG1655 genome, and the promoter and terminator sequences were ligated using primer tails. The pUC19 origin of replication was amplified from pUC19 (catalog number N3041S, New England Biolabs, Ipswich, MA, USA) using primer tails homologous to other fragments. The three fragments were assembled via Gibson isothermal assembly (Gibson et al., Nat Methods. May 2009; 6(5): 343-5). The Golden Gate Assembly site (Engler et al., PLoS One. 2008; 3(11): e3647) was inserted into the phage by being included in the primer tails used for amplification and assembly of the fragments. The Golden Gate Assembly site included BsaI, BsmBI, and PaqCI. The M13 origin of replication is included in the design, allowing the phage particle to replicate in *E. coli* as cssDNA. A packaging signal is included, enabling the M13 protein to package the cssDNA into the phage particle. The pUC19 origin of replication is included, allowing the phage particle to replicate in *E. coli* as dsDNA. The pyrF gene is used as a selectable auxotrophic marker for KTx00001 and its progeny, which are uracil auxotrophs due to the interruption of the pyrF gene genome copy by inserting M13KO7 transcription unit 1 (see Example 1). Golden Gate Assembly sites are included for inserting user-defined sequences. User-defined sequences can be synthesized or amplified with primers such that their flanking positions are BsaI, BsmBI, or PaqCI restriction sites. The insert fragment and the minimal phage particle are then incubated with appropriate restriction enzymes and buffers, along with a ligase, and the temperature is cycled between 37ºC (for restriction enzyme cleavage) and 16ºC (for ligation). Because the restriction enzyme recognition sequence is asymmetric and located several bases from the cleavage site, the recognition sequence is inserted into the phage particle and the insert fragment, which are then cleaved after the initial digestion. If the parts subsequently assemble in the desired manner to insert the user-defined sequence into the phage particle, the recognition sequence is ablated, and no further cleavage occurs. However, if the fragment is re-ligated to the recognition sequence to reform the initial input, the recognition sequence is regenerated, and another round of cleavage can continue. In this way, the reaction proceeds unidirectionally until almost all restriction recognition sites have been eliminated and almost all insert fragments have been correctly inserted into the phage particle backbone.
[0155] It is noteworthy that the pyrF auxotrophic marker is included in the phage particles instead of traditional antibiotic markers to increase the safety of the CSSDNA generated from the system. Since one application of the CSSDNA generated by this system is human therapeutic gene and cell therapy, the inclusion of antibiotic marker sequences is undesirable. The exclusion of antibiotic marker sequences prevents the system from spreading antibiotic resistance to microorganisms that could potentially infect human patients with antibiotic-resistant strains. Example 6: CSSDNA production increases with increasing phage copy number.
[0156] Given the interplay of various factors, including helper plasmid copy number (or integrated phage gene copy number), phage gene expression, and phage particle copy number used for fermentation, how altering the phage particle copy number will affect CSSDNA yield in fermentation-based production methods is unpredictable. The results presented in this example demonstrate that increasing the phage particle copy number increases CSSDNA production from bacterial production strains containing both phage particles and helper plasmids.
[0157] A series of phage particles containing different replication origins were constructed, as shown in Table E1 below. The map of phage particles containing the pUC19 origin is shown in... Figure 8A The primers used to generate inc1 and inc2 mutations at the pUC19 initiation point are shown in [the diagram]. Figure 8B In this study, phage particles were tested in combination with two different helper plasmids: KHP0 and KHP1, each containing the p15A origin of replication. KHP0 was derived from helper phage M13KO7 by removing the F1 origin of replication and packaging signals. This was achieved by amplifying kanamycin markers and all protein-coding genes from M13KO7 using primers designed to exclude the F1 origin of replication and packaging signals, followed by in vitro recirculation of the amplified DNA using DNA ligase. KHP1 was similarly derived from helper phage M13KO7 but further included attenuated gene V, generated by mutating 5 base pairs immediately upstream of gene V, resulting in a reduced translation rate of gene V. Gene V is involved in regulating the transition from replicating double-stranded DNA to non-replicating single-stranded DNA within bacterial cells. Reduced expression of gene V has been shown to lead to a longer replication phase in the form of double-stranded DNA, and thus an accumulation of more dsDNA template, resulting in higher cssDNA yield. Table E1: Phage particles with different copy numbers
[0158] Standard NEB5alpa chemically competent E. coli cells were co-transformed with Kano helper plasmid 1 (KHP1) and one of four phage variants, such as... Figure 9As shown. Individual colonies were selected on LB agar plates supplemented with 50 μg / ml kanamycin (helper plasmid) and 100 μg / ml carbenicillin (phage particles). Six individual colonies were picked from each plate and inoculated into seed cultures in 2xYT medium supplemented with 50 μg / ml kanamycin and 100 μg / ml carbenicillin. The seed cultures were allowed to grow overnight to saturation and then diluted in the same medium to approximately OD600 0.04. The cultures were then grown at 30ºC with shaking for 18 hours. The cultures were harvested and the medium was clarified from the bacterial cells by centrifugation. The supernatant containing phage particles and therefore cssDNA was pipetted into fresh dishes, and the concentration of cssDNA in each sample was analyzed by qPCR using primers and probes targeting the F1 origin of replication on the cssDNA, along with FAM reporter and NFQ-MGB quencher.
[0159] like Figure 9 As shown, the mutations in inc2 and inc1 and 2 phages resulted in higher CSSDNA yields compared to the "standard" pUC19 phage. Phages with reduced copy numbers resulted in significantly lower yields compared to the standard pUC19, inc2, or inc1 and 2. The inc1 mutation is a C59T mutation relative to the standard pUC initiator, with nucleotide numbering based on the sequence of the standard pUC initiator shown in SEQ ID NO: 27. The inc2 mutation is a C92T mutation relative to the standard pUC initiator, with nucleotide numbering based on the sequence of the standard pUC initiator shown in SEQ ID NO: 27. The inc2 and inc1 / inc2 phages used in these experiments further contained the G62A mutation relative to the standard pUC initiator shown in SEQ ID NO: 27, but this mutation is not expected to affect the initiator's function.
[0160] These results indicate that increased phage copy number can lead to increased CSSDNA production from phages. The copy number reduction initiation point yielded an estimated 15 copies / cell, while the standard pUC19 initiation point provided an estimated 1000 copies / cell, and the inc1 and 2 mutant pUC19 initiation point provided an estimated 4000 copies / cell. These copy number estimates varied considerably depending on the growth medium and growth stage.
[0161] In larger-scale cultures, the use of phagemids containing inc1 and inc2 mutations also demonstrated a significant increase in CSSDNA yield. Standard NEB5alpa chemically competent *E. coli* cells were co-transformed with Kano helper plasmid 1 (KHP1) and one of two phagemid variants (standard (pUC19) or inc1 / inc2), as shown in the results. Figure 10As shown. Individual colonies were selected on LB agar plates supplemented with 50 μg / ml kanamycin (helper plasmid) and 100 μg / ml carbenicillin (phage plasmid). Fifteen individual colonies were picked from each plate and inoculated into seed cultures in 2xYT medium supplemented with 50 μg / ml kanamycin and 100 μg / ml carbenicillin. The seed cultures were allowed to grow overnight to saturation, and then diluted to approximately OD600 0.04 in 100 ml (9 cultures) or 500 ml (6 cultures) of the same medium. The cultures were then grown at 30ºC with shaking for 18 hours. The cultures were harvested and the medium was clarified from the bacterial cells by centrifugation. The supernatant containing phage particles and therefore CSSDNA was pipetted into fresh dishes, and the concentration of CSSDNA in each sample was analyzed by qPCR using primers and probes targeting the F1 origin of replication on the CSSDNA, along with a FAM reporter and NFQ-MGB quencher. Figure 10 As shown, the inc1 and 2 mutations in the phage particles resulted in at least a 5-fold increase in cssDNA yield compared to the yield obtained using the wild-type pUC19 origin. This increased cssDNA yield was reproducible for both 100 mL and 500 mL culture production.
[0162] The above results indicate that, relative to the pUC19 replication origin, the inc mutation in the pUC19 phage plasmid replication origin confers a significant increase in CSSDNA production (e.g., an increase of approximately 5 to 10-fold relative to the wild-type pUC19 origin and an increase of more than 100-fold relative to the p15A origin). Example 7: CSSDNA yield increases with increasing helper plasmid copy number.
[0163] The results provided in this embodiment demonstrate that increasing the copy number of helper plasmids can also increase CSSDNA yield, including in combination with high copy number phages (inc1 and 2 phages, estimated at 4000 copies / cell).
[0164] As shown in Table E2 below, a series of different helper plasmids were constructed. Table E2: Helper plasmids with different copy numbers
[0165] The *E. coli* production host cells were transformed using a combination of the designated helper plasmid from Table E2 and phage plasmid 163 described in Example 6. Individual colonies of the double-transformed *E. coli* were picked and grown in a starter culture, then diluted and grown for small-scale production analysis. After 18 hours of culture, the bacteria were lysed and the relative CSSDNA yield was analyzed by qPCR. Figure 11 As shown, compared with the lower copy number helper plasmids 81 (pSC101 ori) and 114 (p15A ori), the pUC19 origin helper plasmid (148) resulted in an increase in CSSDNA yield from phages at 18 hours. Example 8: Production of the production strain Host strain selection
[0166] The first step is to select the background strain in which the production machine is engineered. CSSDNA production is typically carried out in *E. coli* cloning host strains such as DH5α, JM109, XL-1 Blue, or M1061. These common laboratory strains have several advantages because they have been engineered to optimize plasmid DNA cloning and preparation. They have had recA removed to reduce recombination events in the desired CSSDNA product. They have had endA removed to reduce DNA degradation. They have had host nucleases removed to improve transformation efficiency. They have had F fimbriae removed to prevent reinfection of cells by infectious phage particles. However, they are also legacy strains, domesticated years ago and engineered or mutated with other deletions that may or may not be beneficial for CSSDNA production. They typically have slower growth rates and lower biomass yields compared to wild-type strains without the same number of mutations.
[0167] We hypothesized that wild-type strains could serve as better production hosts because they generally exhibit improved adaptability compared to highly engineered laboratory strains, including increased growth rates and biomass yield per unit of culture medium provided. To test this hypothesis, we obtained a small library of background *E. coli* strains, including the aforementioned common laboratory strains and other less domesticated strains. These included BW25113, K12, and MG1655. A complete list of strains is provided in Table 4. Table 4: Escherichia coli strains selected for growth in defined-component culture media We prefer to use a defined-component medium for producing CSSDNA because it reduces the cost of large-scale production, increases the reproducibility of production runs, and allows the use of auxotrophic markers instead of antibiotic markers, thus increasing the safety of the final CSSDNA product. We chose to use a defined-component medium based on Riesenberg's 1991 formulation because it is well-established and has been widely used since its publication. Riesenberg et al. High celldensity cultivation of Escherichia coli at controlled specific growth rate. J Biotechnol. Aug 1991;20(1):17-27. The Reisenberg formulation has also been used to successfully produce large quantities of CSSDNA in fed-batch reactor systems. Kick et al. Efficient Production of Single-Stranded PhageDNA as Scaffolds for DNA Origami. Nano Lett. Aug 2015;15(7):4672-6. We will refer to this medium as Riesenberg medium below. However, other bacterial cell media may also be used.
[0168] We first screened our strain collection for growth in Riesenberg medium. Single colonies were inoculated in triplicate into LB medium and grown overnight as pre-cultures, then diluted to an OD500 of 0.05 in Riesenberg medium the following morning. These cultures were incubated at 30ºC with shaking for 24 hours, and the final OD600 was recorded. Cells from each strain were streaked onto antibiotic-free LB agar to obtain single colonies.
[0169] Figure 12 The results of the final OD values for the tested strains are shown. We did indeed find that several isolates of DH5α, XL-1 Blue, and MC1061 reached the lowest final OD values, such as... Figure 12 As shown. Less engineered strains such as MG1655 and K12 achieved higher OD ( ). Figure 12 Multiple isolates of BW25113 reached the highest OD ( ). Figure 12These strains have slightly reduced genomes, lacking the lactose, arabinose, and rhamnose operons, which can provide growth benefits because the cells need to replicate less DNA and produce less protein. These alternative sugar operons are not needed in fermentation cultures that provide glucose as the sole carbon source. Furthermore, since BW25113 is the parent strain of the Keio knockout set, single-gene deletion strains already exist in the context of genes essential for the synthesis of certain amino acids and nucleosides. In fact, three of these deletion strains—bac012 (ΔpyrF), bac013 (ΔpyrF), bac016 (ΔpyrF, Δkan), and bac021 (ΔleuB)—were included in our screening and all performed exceptionally well. The best-performing strain was the parent BW25113 of the aforementioned single-gene deletion strains, but strains with pyrF or leuB deletions followed closely behind, and conveniently, they are already suitable for selection against auxotrophic markers and for maintaining plasmids with auxotrophic markers. Therefore, we chose bac016 for further study because it grows robustly in our preferred medium and because it has lost the pyrF and kanamycin resistance markers during the engineering process. Other strains can also be used for CSSDNA production according to any of the methods disclosed herein.
[0170] After identifying bac016 as a strain that grows well in our selected, well-defined medium, we next verified its ability to generate cssDNA using helper plasmids and phage particles before any genome engineering steps. bac016 was co-transformed with the standard DH5α production host bac001 using a helper plasmid (cdsDNA114) and phage particle (cdsDNA111) system. In another experiment, bac016 was co-transformed with the cdsDNA114 helper plasmid and a novel cdsDNA117 phage particle expressing the pyrF auxotrophic marker instead of the antibiotic resistance gene.
[0171] The cdsDNA114 helper plasmid expresses all M13 genes and kanamycin markers and carries the p15A origin of replication. It lacks the M13 origin of replication and packaging signals, therefore it cannot replicate as an infectious phage particle. The cdsDNA111 phage particle has a pUC origin of replication and expresses RFP and ampicillin resistance markers. Unlike the helper plasmid, it does contain the M13 origin of replication and packaging signals, allowing it to be generated as cssDNA and packaged into phage capsids. These capsids cannot replicate in new bacterial hosts because they lack the essential M13 gene. The cdsDNA117 phage particle is identical to the cds111 phage particle except that it expresses the pyrF auxotrophic marker instead of the ampicillin antibiotic marker.
[0172] Colonies from each of the three transformation reactions were inoculated in triplicate into Riesenberg medium and grown overnight at 37ºC as pre-cultures containing appropriate antibiotics and supplemented with uracil if necessary to compensate for the pyrF auxotrophic type. The cultures were then diluted to OD600 0.05 and incubated at 30ºC for 48 hours to test for cssDNA production. The broth was harvested and centrifuged to isolate *E. coli* cells from the supernatant containing secreted phage particles. The supernatant was then transferred to clean tubes and heated at 99ºC for 10 minutes to lyse the phage particles containing cssDNA. This material was then used as a template for qPCR targeting the F1 origin of replication present on the cssDNA.
[0173] This experiment demonstrates that BW25113 can produce cssDNA when transformed using a conventional dual-plasmid-helper plasmid-phage production machine. When transformed with an ampicillin-resistant phage, it produces the same amount of cssDNA as the conventional DH5α-producing host transformed with the same phage. When transformed with the novel cdsDNA117 phage expressing the pyrF auxotrophic marker, it produces significantly more cssDNA.
[0174] Because DH5α lacks pyrF deletion, CSSDNA production was not detected in this strain when phage particles expressing pyrF auxotrophic markers were used. However, strains like DH5α can be engineered to have pyrF deletion compatible with auxotrophic selection.
[0175] Figure 13AThis study demonstrates the validation of cssDNA production in the unengineered BW25113 parental strain. The conventional production host bac001 (DH5α) and the newly proposed production host bac016 (BW25113 ΔpyrF, Δkan) were each transformed using a conventional production machine consisting of a cdsDNA114 helper plasmid and a cdsDNA111 phage. Additionally, bac016 was co-transformed with the same cdsDNA114 helper plasmid and a novel cdsDNA117 phage expressing a pyrF auxotrophic marker. Colonies were inoculated triplicate into Riesenberg medium (with appropriate antibiotics and supplemented with uracil for plasmid selection) and grown overnight in pre-culture. The cultures were then diluted to OD600 0.05 in the same medium and incubated at 30ºC for 48 hours to produce cssDNA. cssDNA was quantified using qPCR targeting the F1 origin of replication present on the cssDNA backbone. Data confirms that the non-engineered host bac016 produces the same amount of CSSDNA as the standard bac001 DH5α producing host, and in fact, produces much more CSSDNA when the cdsDNA117 phage carrying the pyrF auxotrophic marker is used as a CSSDNA template. Integration of M13 transcription units
[0176] The M13 genome is organized into two independent transcriptional units. The first consists of genes II, X, V, VII, IX, and VIII, and is expressed by a strong promoter. The second consists of genes III, VI, I, XI, and IV, and is expressed by a weak promoter. Two of these genes overlap with the open reading frames of the other two genes (i.e., the gene encoding p10 is completely contained within the coding region of p2, and the gene encoding p11 is completely contained within the open reading frame of pI). PA, PB, and PH are strongly active promoters (see [link to documentation]). Figure 1B The other two weak promoters were labeled PZ and PW. The two terminators (T) were strongly rho-independent terminators; the third, weaker terminator (T(weak)) was rho-dependent. (Smeal et al., Simulation of the M13 life cycle I: Assembly of a genetically-structured deterministic chemical kinetic simulation, Virology, Vol. 500, 2017, pp. 259-274.)
[0177] To generate stable cell lines capable of producing M13 phage particles while being highly unlikely to recombine with phage particles to produce replicative phages, each transcriptional unit of an M13 phage containing protein-coding genes and lacking phage origin of replication and phage packaging signals was cloned from the publicly available helper phage M13KO7 (catalog number N0315S, NEB, Ipswich, MA) (SEQ ID NO: 24) and integrated into two loci in the genome of *E. coli* BW25113 in the form of ΔpyrF (produced as part of the Keio knockout collection, i.e., the so-called isolate JW1273-1). (See Baba et al., Construction of *Escherichia coli* K-12 in-frame, single-gene knockout mutants: the Keio collection. Mol Syst Biol. 2006;2:2006.0008.) To regulate M13 protein production, the gene encoding T7 phage RNA polymerase, controlled by the *E. coli* lac operon and promoter, was also integrated into another locus in the host strain. This allowed for the induction of M13 phage gene expression by adding isopropyl β-d-1-thiogalactopyranoside (IPTG) to induce T7 polymerase expression. The *E. coli* BW25113 ΔpyrFKeio isolate JW1273-1 (bac012) host strain initially contained a spectinomycin resistance marker flanked by an FRT recombination site, replacing the wild-type *E. coli* BW25113 pyrF gene. This spectinomycin resistance marker was genetically engineered to delete the marker, resulting in a strain without the spectinomycin resistance marker. Integration of T7 polymerase chain kit
[0178] To generate strains with inducible expression of the M13 phage gene, recombinant engineering and CRISPR negative selection were used to insert the T7 polymerase gene, controlled by an inducible promoter (IPTG-inducible lac promoter), into the endA locus. The endA locus was chosen as the insertion site because the EndA protein is known to degrade DNA and the deletion of endA is known to increase the yield and quality of DNA produced in *E. coli*. (Lin, JJ (1992) Endonuclease A degrades chromosomal and plasmid DNA of *Escherichia coli* present in most preparations of single-stranded DNA from phagemids. Proc. Natl. Sci. Counc. Repub. China B16 1-5.) Correct integrons were identified by colony PCR and Oxford Nanopore sequencing.
[0179] After integrating the inducible T7 polymerase expression cassette, we next verified that the novel strain bac035 could indeed induce T7 polymerase expression and that this polymerase could be transcribed from the T7 promoter. To test this, the strain was integrated with the pT7:emGFP expression cassette (containing the nucleotide sequence for pT7 promoter-driven emGFP expression). Specifically, strain bac036 was generated by integrating the pLac:T7 polymerase from *E. coli* BL21(DE3) into the endA locus of the BW25113 ΔpyrF background strain. The cassette for emGFP expression via the T7 promoter was then inserted into yhaV to generate strain bac058. This strain was grown in different concentrations of IPTG to induce T7 polymerase and indirectly induce emGFP expression. Cultures were grown in a shaking incubator at 30ºC for 48 h. At the end of the growth period, absorbance (OD600) and GFP fluorescence at 600 nm were measured using a Biotek Synergy H1 plate reader. Divide any obtained fluorescence unit by OD600 and compare it with the parental strain bac036, which lacks GFP expression plasmid. Figure 13B ).like Figure 13B As shown in the figure, the results demonstrate the induction of pLac:T7 polymerase expression. Integration of the first transcription unit
[0180] The first transcriptional unit (TU1) of the M13 genome was integrated into strain bac036 using recombinant engineering and antibiotic selection. TU1, along with a homologous arm for integration into the intA locus of the *E. coli* BW25113 genome and a spectinomycin selectable marker to allow selection of the integrator, was cloned into four circular double-stranded plasmids: cdsDNA206, cdsDNA207, cdsDNA208, and cdsDNA209. TU1 contains M13 genes II, V, VII, VIII, and IX. Because the optimal expression level of the integrated M13 genes was not known beforehand, each of the four plasmids had a promoter of varying strength, expressing mRNA at different levels. cdsDNA206 had a moderately strong T7 promoter, cdsDNA207 had a weak T7 promoter, cdsDNA208 had a wild-type M13 TU1 promoter, and cdsDNA209 had a strong T7 promoter. Based on data from Komura et al., High-throughput evaluation of T7 promoter variants using biased randomization and DNA barcoding, Plos, 2018, the T7 promoter was mutated compared to the wild-type T7 promoter to achieve these different expression levels. The plasmid contains a p15A origin for replication in *E. coli* and an ampicillin resistance marker for plasmid selection in *E. coli*. The spectinomycin selectable marker is flanked by an FRT recombination site to allow for marker recycling after Flp recombinase expression.
[0181] Bac090-bac092 expresses T7 polymerase via the lac promoter and M13TU1 (integrated into the intA locus in the E. coli BW25113 genome) via three different promoters: a strong T7 promoter, a weak T7 promoter, and a native M13TU1 promoter. Integration of the second transcription unit
[0182] The second transcription unit (TU2) was then integrated into the intZ locus in strain bac090-092 to generate small combinatorial libraries with different combinations of TU1 and TU2 expression levels. TU2 contained M13 genes III, VI, I, and IV.
[0183] Four different promoters were tested to drive TU2 expression. In cdsDNA215, TU2 was driven by a strong T7 promoter; in cdsDNA216, TU2 was driven by a moderate T7 promoter; in cdsDNA217, TU2 was driven by a weak T7 promoter; and in cdsDNA218, TU2 was driven by its wild-type M13 TU2 promoter. The intZ gene was chosen as the integrase locus because it is a phage-associated integrase with no known beneficial function in *E. coli*. The resulting new strain was labeled bac105-110, and the glycerol stock solution was stored at -80ºC.
[0184] After integrating the IPTG-inducible T7 polymerase cassette and the small library that drives the promoters of both M13 TU1 and M13 TU2 into the same strain, we now have a set of strains containing all the genes of the M13 genome, as well as the T7 polymerase and promoter system. Determination of novel strains for CSSDNA production
[0185] After generating small libraries of strains expressing M13TU1 and M13TU2 with promoters of varying strengths under the control of IPTG-inducible T7 polymerase, we next attempted to determine the ability of the new strains to produce cssDNA.
[0186] First, strain bac105-bac110 was transformed by electroporation with the phage cdsDNA111 containing the F1 origin of replication, F1 packaging signal, ampicillin resistance gene, and RFP expression cassette. Triplicate single colonies were picked and inoculated into 500 μL LB medium supplemented with 50 μg / ml kanamycin and incubated overnight at 37ºC. In the morning, the optical density (OD600) at 600 nm was measured, and the culture was diluted to OD0.05 in 200 μL Riesenberg determinant medium. The culture was then incubated at 30ºC until saturation, which took approximately 48 hours at 30ºC.
[0187] The culture was centrifuged at 4000 RCF to remove E. coli cells containing double-stranded plasmid DNA, and the supernatant containing phage particles was collected for analysis. The phage particles were then lysed by heating the solution to 99ºC for 10 min. The supernatant containing cssDNA was retained for use as a qPCR template. The cssDNA was quantified using qPCR calibrated with a standard curve of cdsDNA111 DNA.
[0188] Figure 14AThe screening results of small combinatorial libraries of integrated TU1 and TU2 driven by promoters of varying strengths are shown. Strains with integrated T7 polymerase, M13 TU1, and M13 TU2 were transformed with the phage particle cdsDNA111 to generate hosts theoretically capable of producing cssDNA. Single colonies were picked in quintuples and inoculated into 500 μL LB Miller culture supplemented with carbenicillin for phage particle selection and grown overnight as pre-cultures. These cultures were then diluted to OD600 0.05 and incubated on plates for approximately 48 hours until saturation. The supernatant was harvested and the phage particles were lysed to release their cssDNA for use as qPCR templates. CSSDNA was quantified using qPCR primers targeting F1 ori present on the cssDNA and a standard curve of cdsDNA111 DNA. Bac105, expressing M13 TU1 and TU2 via a strong T7 promoter, produced the most cssDNA. Other strains with weaker promoter combinations did not produce significantly more DNA than the negative control, which lacked phage particles that could be used as templates for cssDNA production.
[0189] Next, we aim to replicate the above data by comparing with other production hosts that lack integrated M13 genes. As a control, the parental strain bac016 (E. coli BW25113 ΔpyrF) and the standard DH5α cssDNA production host bac001 were transformed with the same cdsDNA111 phage and helper plasmid (cdsDNA114), which was derived from helper phage M13KO7 by amplifying kanamycin markers and all protein-coding genes from M13KO7 using primers designed to exclude F1 replication origin and packaging signals, thereby removing the F1 origin and packaging signals, and then recirculating the amplified DNA using an in vitro reaction with DNA ligase.
[0190] Figure 14B The results of comparing the selected engineered strain with its unengineered parent strain are shown. The engineered production host, from which the M13 phage gene from its genome was derived, was transformed with the cdsDNA111 phage particle. The unengineered parent strain bac016 was transformed with the helper plasmid cdsDNA114 (… Figure 14BThe cells were transformed with "HP" to provide expression of the M13 phage gene and the same cdsDNA111 phage particle. Transformed colonies were inoculated in quintuples and grown in 200 μL of medium in 96-well microplates at 30ºC for 48 h. The supernatant was then collected by centrifugation to remove E. coli and boiled at 99ºC for 10 min to lyse the phage particles. This supernatant was then used as a template for qPCR to quantify cdsDNA production. A standard curve was used for qPCR, which measures the presence of F1 Ori in solution as a surrogate indicator of cdsDNA concentration. As in previous experiments, bac105, expressing M13 TU1 and TU2 with the strongest promoter, outperformed the other three engineered strains tested (bac106, bac108, and bac110), all of which have weaker promoters driving M13 gene expression. Notably, bac105 outperformed its parent strain bac016, which utilizes a conventional dual-plasmid production system instead of an integrated phage gene as bac105 does. Very little or no CSSDNA was detected in any of the three negative control strains lacking a complete CSSDNA production system.
[0191] Figure 15A The optimal engineered production host (bac105) is compared with its parent strain and the conventional production strain DH5α using a conventional helper plasmid system. All three strains were transformed with phage plasmid cdsDNA111 as a template for CSSDNA production. The non-engineered DH5α and BW25113 strains bac001 and bac016 were also transformed with helper plasmid cdsDNA114 to provide expression of the M13 phage gene. The strains were grown in 200 μL of medium at 30ºC for 48 h in microplates with appropriate antibiotics for plasmid maintenance. The supernatant was collected by centrifugation to remove *E. coli* and boiled at 99ºC for 10 min to lyse the phage particles. This supernatant was then used as a template for qPCR to quantify CSSDNA production. A standard curve was used for qPCR, which measures the presence of F1 Ori in solution as a surrogate indicator of CSSDNA concentration. Engineered bac105 strains expressing M13 integration copies of TU1 and TU2 via a strong T7 promoter produced more CSSDNA than either of the two strains expressing the M13 phage gene derived from a helper plasmid. Notably, bac016 is the parent strain of bac105, so the difference in CSSDNA production is likely due to differences in M13 gene expression rather than differences in strain background. It is also noteworthy that bac001, with similar helper plasmids and phage particles, is a common CSSDNA-producing strain in industrial and academic settings. Driven by the integration of pT7 into individual M13 gene VIII
[0192] It is hypothesized that the natural stoichiometry of the various M13 genes present in the wild-type M13 genome may not be optimal for high-level cssDNA production, especially for cssDNA constructs that differ significantly from the M13 genome in terms of quality, such as length and GC content. For example, the M13 shaft consists of thousands of copies of the gene VIII protein product. The shaft elongates to accommodate longer cssDNA constructs and contracts to accommodate shorter ones. The production of cssDNA constructs that differ significantly in length from the 7 kbM13 wild-type genome could be improved by varying the expression ratio of gene VIII relative to other M13 genes.
[0193] Strains bac132 (bac105 plus gene II integrated into the mazF locus) and bac135 (bac105 plus gene VIII integrated into the mazF locus) were generated. The mazF locus was chosen as the integration locus because it is the toxic portion of the mazE / mazF toxin / antitoxin system. It was thought that the absence of the toxic gene, mazF, might have a beneficial effect on the strain.
[0194] Figure 15B The results show the determination of cssDNA production in engineered strains that integrated additional individual phage genes. Bac001 (unengineered DH5α), bac016 (unengineered BW25113), bac105 (the best engineered strain to date), bac132 (bac105 plus gene II), and bac135 (bac105 plus gene VIII) were each transformed with cdsDNA111 RFP phage particles and grown for 48 hours in Riesenberg assay medium supplemented with uracil. The broth was harvested and centrifuged to remove *E. coli* cells from the culture. The supernatant was transferred to clean tubes and heated at 99ºC for 10 minutes to lyse the phage particles and release cssDNA into the medium. This solution was used as a template for qPCR using primers targeting the RFP gene and a standard curve of cdsDNA111 RFP phage particles. The engineered strain bac105 again outperformed the unengineered control strain. Its daughter strain bac132 (which integrates an extra copy of gene II) performed poorly, but another daughter strain, bac135, which has an extra copy of gene VIII inserted, performed slightly better than its parent bac105. Since the differences were small, we chose to compare the two strains, bac105 and bac135, in future experiments.
[0195] Having identified strain bac135, containing individual copies of T7 polymerase, TU1, TU2, and gene VIII, as the optimal producer among the two strains integrating a single M13 gene, we next sought to test the generation of alternative cssDNA sequences besides the RFP sequence from cdsDNA111. One hypothesis is that adding an extra copy of gene VIII would allow cells to produce more cssDNA, especially when the cssDNA sequence is significantly longer than the 7 kb genome of the wild-type M13 genome. This is because the tubular outer shell of the M13 phage capsid, composed of thousands of copies of p8 (the product of gene VIII), expands and contracts to accommodate cssDNA sequences of varying lengths. Therefore, longer cssDNA sequences would require more copies of the p8 protein to be packaged into the capsid, making it logical that increased gene VIII expression would promote the generation of longer cssDNA sequences. Example 9: Generation of CSSDNA encoding dystrophin-associated protein (utrophin) and dystrophin-resistant protein
[0196] This embodiment demonstrates the ability of the engineered production strain described in Example 8 to generate CSSDNA containing long genetic sequences encoding dystrophin and muscular dystrophy-related proteins. The engineered CSSDNA system advantageously allows for the generation of long CSSDNA templates that can be used in gene therapy.
[0197] The strains were identified by transformation with phage particles cdsDNA100 (dystrophin-associated protein) and cdsDNA101 (anti-dystrophin). Unengineered control strains bac001 and bac016 were also transformed with the helper plasmid cdsDNA114 to provide expression of the M13 phage gene. Transformed colonies were inoculated in triplicate and grown in 200 μL of medium in 96-well microplates at 30ºC for 48 h. The broth was then harvested, E. coli cells were removed by centrifugation, and the supernatant containing phage particles was collected and boiled at 99ºC for 10 min to lyse the phage particles. This supernatant was then used as a template for qPCR to quantify cssDNA production. A standard curve was used for qPCR, which measured the presence of AAVS1 homologous arm sequences in solution as a surrogate indicator of cssDNA concentration. Figure 15B The results show the generation of multiple CSSDNA sequences from engineered host strains.
[0198] The results were measured for bac001 (conventional DH5α), bac016 (unengineered parent), bac105 (best engineered strain to date), and bac135 (bac105 with an additional copy of gene VIII). The two engineered strains, bac105 and bac135, produced more CSSDNA than their unengineered counterparts, and did so more consistently, with smaller standard deviations between biological replicates. Notably, both dystrophin-associated protein (DMAP) and anti-dystrophin constructs are therapeutically relevant to Dichené muscular dystrophy, and each construct was longer than 13 kb when produced as CSSDNA. DMAP could be produced at low levels using conventional helper plasmid systems, but not anti-dystrophin, while each could be produced at moderate levels using engineered strains. Integration of an additional copy of gene VIII showed a slight increase in DMAP CSSDNA yield. Example 10: Generation of CSSDNA using phage particles with auxotrophic markers
[0199] The standard phage plasmid cdsDNA111 was amplified using PCR primers designed to amplify the entire plasmid excluding the ampicillin resistance cassette. The pyrF gene cassette was amplified from the *E. coli* DH5α genome and its 75 bp natural promoter. In both PCR reactions, primers were designed to have tails that produced a 30 nt overlap with another fragment. These two linear double-stranded DNA fragments were assembled using the NEBuilder HiFi DNA assembler mix (catalog number E2621L, NEB, Ipswich, MA). The novel plasmid was sequenced using Oxford Nanopore technology and stored as cdsDNA117.
[0200] The F1 origin of replication is included in the design, enabling the phage particle to replicate as cssDNA in *E. coli*. A packaging signal is included, allowing the M13 protein to package the cssDNA into the phage particle. The pUC19 origin of replication is included, enabling the phage particle to replicate as dsDNA in *E. coli*. The pyrF gene is used as a selectable auxotrophic marker for engineered production and its progeny, as they are uracil auxotrophs due to interruptions in the pyrF gene genome copy. The Golden Gate Assembly™ site (Engler et al., PLoS One. 2008;3(11):e3647) is included in the phage particle. The Golden Gate Assembly™ site is recognized by the BsmBI type II restriction enzyme, allowing for the easy insertion of novel user-defined sequences into the pyrF phage particle backbone. User-defined sequences can be synthesized or amplified with primers such that their flanking positions are BsmBI restriction sites. The insert and phage are then incubated with BsmBI restriction enzyme, buffer, and ligase, with temperatures cycled between 37ºC (for restriction enzyme cleavage) and 16ºC (for ligation). Because the restriction enzyme recognition sequence is asymmetric and located a few bases from the cleavage site, the recognition sequence is inserted into the phage and insert, which are then cleaved after the initial digestion. If the fragments subsequently assemble in the desired manner to insert the user-defined sequence into the phage, the recognition sequence is ablated, and no further cleavage occurs. However, if the fragment is religated to the recognition sequence to reform the initial input, the recognition sequence is regenerated, and another round of cleavage continues. In this way, the reaction proceeds unidirectionally until almost all restriction recognition sites have been eliminated and almost all insert fragments have been correctly inserted into the phage backbone.
[0201] It is noteworthy that the pyrF auxotrophic marker is included in the phage particles instead of traditional antibiotic markers to increase the safety of the CSSDNA generated from the system. Since one application of the CSSDNA generated by this system is human therapeutic gene and cell therapy, the inclusion of antibiotic marker sequences is undesirable. The exclusion of antibiotic marker sequences prevents the system from spreading antibiotic resistance to microorganisms that could potentially infect human patients with antibiotic-resistant strains.
[0202] The engineered production host bac105 was transformed with phage particles cdsDNA111 (RFP with an ampicillin resistance marker) and cdsDNA117 (RFP with a pyrF auxotrophic marker). Unengineered control strains bac001 (DH5α) and bac016 (BW25113) were transformed with phage particle cdsDNA111 and helper plasmid cdsDNA114 to provide expression of the M13 phage gene. Strains were grown for 48 hours at 30ºC in 200 μL of definitive-component medium in microplates with appropriate antibiotics or dropouts for plasmid maintenance. Engineered strains carrying phage particles with auxotrophic markers were grown in definitive-component medium without uracil. Other strains were grown in definitive-component medium containing uracil. Broth was harvested, and the supernatant containing phage particles was clarified by centrifugation to remove *E. coli*. The supernatant was heated at 99ºC for 10 min to lyse the phage particles containing cssDNA, and this lysate was used as a template for qPCR to quantify cssDNA production. A standard curve was used with cdsDNA111 containing the RFP sequence. qPCR primers were used to target the RFP gene insert fragment of the cssDNA. Parental strain bac016, containing only the helper plasmid cdsDNA114 and only the phage particle cdsDNA111, served as negative controls.
[0203] The results are shown in Figure 15C In the study, when transformed with RFP / ampicillin phage cdsDNA111, bac105 again outperformed its parent strain bac016 and the conventional production host bac001 when transformed with the same phage. Figure 15C It is worth noting that when transformed with RFP / pyrF phage cdsDNA117, bac105 also outperformed both strains. Figure 15C The engineered strain bac105 was not directly compared with the non-engineered conventional production host bac001 (DH5α) because the strain lacks a uracil auxotroph and is unable to select for the pyrF marker. in conclusion
[0204] This high CSSDNA production from an engineered strain using a phage plasmid template with an auxotrophic marker establishes bac105 as the safest available CSSDNA production host. It produces more CSSDNA than conventional helper plasmid or helper phage systems without producing infectious phage particles because the M13 origin of replication and packaging signal are located on phage plasmids lacking genes essential for self-replication. Furthermore, the chances of evolving infectious phages from this production system are greatly reduced due to the fact that the M13 gene, essential for phage replication, is stably integrated at two separate loci in the *E. coli* genome, meaning two separate recombination events are required for the two sets of M13 genes to recombine with each other and with the phage plasmid to reassemble all M13 genes with the M13 origin of replication and packaging signal. Moreover, the strain's ability to produce CSSDNA from a phage plasmid template with an auxotrophic marker means that the entire strain is free of antibiotic markers during production and that antibiotics do not need to be added to the production medium. This reduces the likelihood of antibiotic-resistant bacteria evolving during CSSDNA production or later when CSSDNA is administered to patients. Finally, the engineered production host grows well in a defined-component medium, avoiding the need for expensive, rich media that produce inconsistent inter-run results. Table 5: Engineered strains produced in this study This disclosure is not intended to be limited in scope to the specific disclosed embodiments, which are provided, for example, to illustrate various aspects of this disclosure. Various modifications to the compositions and methods will become apparent from the description and instruction herein. Such changes may be practiced without departing from the true scope and spirit of this disclosure, and such changes are intended to fall within the scope of this disclosure. Exemplary sequence
Claims
1. A production strain comprising a phage particle and at least two phage protein coding sequences integrated into the genome of the production strain, wherein the production strain is capable of producing cssDNA upon introduction of the phage particle.
2. A circular single-stranded DNA (cssDNA) production system, the circular single-stranded DNA production system comprising: at least two phage protein coding sequences integrated into the genome of a production strain; and a phage particle, wherein the production strain is capable of producing cssDNA after the phage particle is introduced into the production strain.
3. A circular single-stranded DNA (cssDNA) production system that does not contain non-endogenous antibiotic resistance genes, said circular single-stranded DNA production system comprising: Production strains; and Phage particles, wherein the phage particle comprises a packaging signal, a design sequence, and at least one optional sequence encoding one or more of the following: a auxotrophic marker, an antitoxin, RNA that inhibits the expression of a gene that, if expressed in the absence of RNA, would delay or prevent bacterial growth, a transcription factor repressor that inhibits the expression of a gene that, if expressed in the absence of a transcription factor repressor, would delay or prevent bacterial growth, a transcription activator that activates the transcription repressor, a sequence of tRNA that associates with non-natural amino acids required for the engineering of the production strain, and combinations thereof.
4. The production system of claim 3, wherein the genome of the production strain comprises at least two phage protein coding sequences integrated into the genome of the production strain.
5. The production strain or system according to any one of claims 1-4, wherein the production strain is susceptible to infection by single-stranded filamentous bacteriophages of the Monodnaviria domain, and is further composed of Gram-negative bacteria selected from the families Enterobacteriaceae, Pseudomonadaceae, Spirillaceae, Xanthomonadaceae, Clostridium, and Propionibacterium.
6. The production strain or system according to claim 5, wherein the production strain is susceptible to infection by single-stranded filamentous phages selected from the group consisting of Ff, Fd, F1, and M13.
7. The production strain or system according to claim 6, wherein the production strain is susceptible to infection by the M13 bacteriophage.
8. The production strain or system according to any one of claims 1-7, wherein the production strain is a strain of Escherichia coli (E. coli).
9. The production strain or system according to any one of claims 1-2 or 4-8, wherein the at least two phage protein coding sequences are expressed differently from each other.
10. The production strain or system according to any one of claims 1-2 or 4-8, wherein the at least two phage protein coding sequences are expressed differentially compared with the expression of the natural phage genome.
11. The production strain or system according to claim 9 or claim 10, wherein the at least two phage protein coding sequences encode at least two phage proteins selected from the following: p1, p2, p8, p10, and p11.
12. The production strain or system according to claim 11, wherein at least one of the at least two phage proteins is selected from p3 and p5.
13. The production strain or system according to claim 12, wherein p3 activity is reduced compared to natural phage genome activity.
14. The production strain or system according to claim 12, wherein p5 activity is reduced compared to natural phage genome activity.
15. The production strain or system according to claim 1 or claim 2, wherein the at least two phage protein coding sequences are encoded by filamentous phages of a single-stranded DNA viral domain.
16. The production strain or system according to claim 15, wherein the filamentous phage is selected from: M13, Ff, Fd, Enterobacter phage F1 [EF068134], Enterobacter phage ID2, Enterobacter phage NL95 [AF059243], Enterobacter phage SP [X07489], Enterobacter phage TW28, Enterobacter phage Qβ, Enterobacter phage Qβ [AY099114], Enterobacter phage M11 [AF059242], Enterobacter phage ST, Enterobacter phage TW18 [FJ483840], and Enterobacter phage VK, or their functional equivalents.
17. The production strain or system according to claim 15 or claim 16, wherein the at least two phage protein coding sequences comprise at least 3, 4, 5, 6, 7, 8, 9, 10, or 11 phage protein coding sequences.
18. The production strain or system according to claim 1 or claim 2, wherein the at least two phage protein coding sequences include sequences encoding one or more phage M13 proteins selected from the following: p1, p2, p3, p4, p5, p6, p7, p8, p9, p10, and p11.
19. The production strain or system according to claim 18, wherein the phage M13 protein sequence comprises the M13 phage gene selected from the following: I, II, III, IV, V, VI, VII, VIII, IX, X, and XI.
20. The production strain or system according to claim 1 or claim 2, wherein the at least two phage protein coding sequences comprise at least three phage protein coding sequences.
21. The cssDNA production system of claim 3, wherein the production strain comprises at least two genome-integrated phage protein coding sequences.
22. The cssDNA production system of claim 21, wherein the phage protein coding sequences of the at least two genomes integrated are selected from sequences of one or more filamentous phages from a single-stranded DNA viral domain.
23. The CSSDNA production system according to claim 22, wherein the one or more filamentous phages are selected from: M13, Ff, Fd, Enterobacter phage F1 [EF068134], Enterobacter phage ID2, Enterobacter phage NL95 [AF059243], Enterobacter phage SP [X07489], Enterobacter phage TW28, Enterobacter phage Qβ, Enterobacter phage Qβ [AY099114], Enterobacter phage M11 [AF059242], Enterobacter phage ST, Enterobacter phage TW18 [FJ483840], Enterobacter phage VK, and their functional equivalents.
24. The cssDNA production system of claim 3, wherein the selectable sequence is an antitoxin sequence from a toxin / antitoxin system.
25. The ccsDNA production system of claim 24, wherein the toxin / antitoxin system is selected from ccdB / ccdA, hokA / sokA, pemK / pemI, mazF / mazE, ChpBK / ChpBI, relE / relB, parE / parD, hipA / hipB, and other toxin / antitoxin systems, wherein the toxin is expressed by the host genome and the antitoxin is expressed by the phage particle.
26. The cssDNA production system of claim 3, wherein the selectable sequence produces at least one RNA molecule that downregulates the expression of a reverse selectable sequence selected from the following and other toxins: HSVtk, Ura3, tetA, sacB, rpsL, pheS, pheS*, pheS**, thyA, lacY, gata-1, ccdB, hokA, pemK, mazF, chpBK, relE, parE, hipA.
27. The cssDNA production system of claim 3, wherein the selectable sequence encodes a transcription factor repressor.
28. The cssDNA production system of claim 27, wherein the transcription factor repressor is selected from tetR, araC, lacI, xylS, and other sequences that reduce the expression of reverse selectable markers or toxins.
29. The cssDNA production system according to claim 3, wherein the transcription activator is selected from araC and xylR.
30. The CSSDNA production system according to claim 3, wherein the auxotrophic marker is selected from uracil, adenine, cytosine, guanine, thymine, alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine, biotin, uridine-5'-monophosphate, pantothenic acid, xanthine, spermidine, and para-aminophenol. Aminobenzoic acid, lipoic acid, nicotinamide nucleoside, nicotinamide mononucleotide, D-glucosamine, thiamine, shikimic acid, aminoethylphosphonic acid, β-alanine, S-methylmethionine, ornithine, indole, indoleacetic acid, L-threonine, L-threonine O-3-phosphate, nicotinic acid, ribosylnicotinamide, pyrimidine, guanidine, purine, and other essential compounds used to synthesize genes expressed by the phage particles to compensate for the lack of naturally occurring or synthetically missing genes with the same or similar functions in the host genome.
31. The production strain of claim 1, further comprising the phage particle, wherein the phage particle comprises a packaging signal, a design sequence, and at least one selectable sequence, optionally wherein the selectable sequence is selected from sequences encoding: auxotrophic markers, antitoxins, RNA that inhibits the expression of genes that, if expressed in the absence of RNA, would delay or prevent bacterial growth, a transcription factor repressor that inhibits the expression of genes that, if expressed in the absence of transcription factors, would delay or prevent bacterial growth, a transcription activator that activates such repressor, or a sequence expressing tRNA that associates with non-natural amino acids required for the engineering of the production strain.
32. The production strain according to claim 31, wherein the auxotrophic marker is selected from uracil, adenine, cytosine, guanine, thymine, alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, valine, biotin, uridine-5'-monophosphate, pantothenic acid, xanthine, spermidine, and p-aminobenzene. Formic acid, lipoic acid, nicotinamide nucleoside, nicotinamide mononucleotide, D-glucosamine, thiamine, shikimic acid, aminoethylphosphonic acid, β-alanine, S-methylmethionine, ornithine, indole, indoleacetic acid, L-threonine, L-threonine O-3-phosphate, nicotinic acid, ribosylnicotinamide, pyrimidine, guanidine, purine, and other essential compounds for the synthesis of genes expressed by the phage particles to compensate for the lack of naturally occurring or synthetically missing genes with the same or similar functions in the host genome.
33. The production strain or system according to claim 1 or claim 2, wherein the at least two phage protein coding sequences comprise at least two phage genes, wherein the at least two phage genes are operatively linked to a synthetic promoter to produce optimized expression for cssDNA, the synthetic promoter comprising one or more of the following: a classical T7 promoter and a mutant T7 promoter under the control of an inducible T7 polymerase, lacI, lacIq, araBAD, tet, a temperature-sensitive promoter, a stress-responsive promoter, a quorum-sensing promoter, a photosensitizing promoter, and other inducible or repressive promoters.
34. The production strain or system according to claim 1 or claim 2, wherein the at least two phage protein coding sequences are integrated into at least two different loci in the genome of the production strain.
35. The production strain or system according to claim 1 or claim 2, wherein at least one of the at least two phage protein coding sequences is altered relative to its endogenous sequence by random mutagenesis, rational design, assisted laboratory evolution, directed evolution, or a combination thereof to produce a protein that produces higher or purer yields of CSSDNA.
36. The cssDNA production system of claim 2, wherein the phage expresses tRNA necessary for translating a re-encoded codon of a non-natural amino acid, wherein the production strain cannot incorporate the non-natural amino acid into its protein in the absence of the phage, and neither the phage nor the production strain can reproduce or replicate in the absence of the non-natural amino acid.
37. The production strain or system according to claim 1 or claim 2, wherein at least one of the at least two phage protein coding sequences further comprises a tag.
38. The production strain according to claim 37, wherein the label is selected from affinity labels or detection labels.
39. The production strain according to claim 38, wherein the detection tag is selected from: fluorescent tags, luminescent tags, chromogenic tags, and another tag capable of rapidly quantifying the number of phage particles in solution.
40. The production strain according to claim 38, wherein the affinity tag is selected from: biotin, his, myc, flag, CBP, GST, HA, HBH, MBP, S, V5 and another affinity tag for assisting in the purification of phage particles from the production broth.
41. The production strain or system according to any one of claims 1-40, wherein the phage particle comprises a pUC origin of replication or a derivative thereof.
42. The production strain or system according to any one of claims 1-40, wherein the phage comprises pST19, pDHA29, pDHA30, pDHK29, pDHK30 or an uncontrolled R1 replication origin or a derivative thereof.
43. The production strain or system according to any one of claims 1-42, wherein the phage contains inc1 mutation and / or inc2 mutation, optionally wherein the phage contains inc1 mutation and inc2 mutation.
44. The production strain or system according to any one of claims 1-43, wherein the phage contains the inc3 mutation.
45. The production strain or system according to any one of claims 1-44, wherein the phage contains the inc5 mutation.
46. The production strain or system according to any one of claims 1-45, wherein the phage copy number in the production strain cells grown to the late logarithmic phase is at least 1,000, optionally wherein the phage copy number in the production strain cells grown to the late logarithmic phase is at least 2,000, at least 4,000, at least 7,000, at least 8,000, or at least 15,000.
47. A circular single-stranded DNA (cssDNA) production system, said circular single-stranded DNA production system comprising: Production strains; and Phage particles, wherein the phage particle comprises a packaging signal, a design sequence, and at least one selectable sequence, The phage contains a pUC origin of replication or a derivative thereof, and the phage contains an inc1 mutation and / or an inc2 mutation, optionally the phage contains both an inc1 mutation and an inc2 mutation.
48. The production system of claim 47, wherein the production strain comprises at least two phage protein coding sequences integrated into the genome of the production strain.
49. A method for generating CSSDNA, the method comprising: The production strain according to any one of claims 1, 5-20, 31-35 or 37-46 is cultured in a culture medium; Introducing the phage particles; and Collect bacteriophage particles; as well as The cssDNA was isolated.
50. The method of claim 49, the method further comprising altering at least one of the at least two phage protein coding sequences.
51. The method of claim 50, wherein the method further comprises comparing the collected CSSDNA with CSSDNA collected from a parent strain to determine whether the altered sequence improves the titer or quality of the produced CSSDNA, the producing strain being prepared from the parent strain.
52. The method of claim 49, wherein the production strain further comprises at least one tagged phage protein.
53. The method of claim 52, wherein the tag is used to isolate the bacteriophage.
54. A method for generating CSSDNA, the method comprising: The CSSDNA production system according to any one of claims 2-48 is cultured in a culture medium; Collect bacteriophage particles; as well as The cssDNA was isolated.
55. The method of claim 54, wherein the producing strain further comprises at least one tagged phage protein.
56. The method of claim 55, wherein the tag is used to isolate the bacteriophage.
57. The production system of claim 3 or the method of claim 54, wherein the production strain comprises a helper plasmid derived from helper phage M13KO7 by removing the F1 origin and the packaging signal, optionally wherein the copy number of the helper plasmid in the production strain cells grown to the late logarithmic phase is at least 1,000, optionally wherein the copy number of the helper plasmid in the production strain cells grown to the late logarithmic phase is at least 2,000, at least 4,000, at least 7,000, at least 8,000, or at least 15,000.
58. The production strain, system, or method according to claim 57, wherein the helper plasmid comprises a pUC origin of replication.
59. The production strain, system, or method according to any of the preceding claims, wherein the production strain is an engineered variant of strain BW25113.
60. The production strain, system, or method according to any one of claims 33-59, wherein the nucleotide sequence encoding T7 polymerase is integrated into the endA locus in the genome of the production strain.
61. The production strain, system, or method according to claim 60, wherein the nucleotide sequence encoding the T7 polymerase is under the control of an IPTG-inducible lac promoter.
62. The production strain, system, or method according to any of the preceding claims, wherein the production strain comprises M13 genes II, V, VII, VIII, and IX integrated into the genome of the production strain at the intA locus.
63. The production strain, system, or method according to any of the preceding claims, wherein the production strain comprises M13 genes III, VI, I, and IV integrated into the intZ locus of the production strain genome.
64. The production strain, system, or method according to any of the preceding claims, wherein the production strain, system, or method further comprises the M13 gene II at the mazF locus integrated into the genome of the production strain.
65. The production strain, system, or method according to any of the preceding claims, wherein the production strain, system, or method further comprises the M13 gene VIII integrated into the mazF locus in the genome of the production strain.
Citation Information
Patent Citations
Plasmids and packaging cell lines for use in phage display
US8227242B2
Viable non-toxic gram-negative bacteria
US8303964B2